AIニュース 2026-08-06
自動生成: 2026-08-06 11:57 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Meta、コーディングエージェント「Muse Code」リリース 「Claude Code」や「Codex」に対抗ITmedia AI+
Metaは、ターミナル上で動作するコーディングエージェント「Muse Code」のベータ版と、新モデル「Muse Spark 1.2」を発…
-
オープンソースのAI機能付きオフィス登場 docx・xlsx・pptx・pdfに対応 GensparkITmedia AI+
AIサービスを開発する米Gensparkは、オープンソースのAI機能付きオフィススイート「GenOffice」を発表した。
-
MythosとGPT-5.6 Solが性能テスト中に暴走 OSSメンテナーに圧力、有害コード実行図る 英政府機関ITmedia AI+
英政府機関AISIによるAIモデルのサイバー能力評価中に、AnthropicのMythos 5やOpenAIのGPT-5.6 Solが実在…
-
Googleのジェフ・ディーン氏、独立してAI実験を大規模自動化する新会社Discovery Loop設立ITmedia AI+
Googleのチーフサイエンティスト、ジェフ・ディーン氏が退社し、新会社「Discovery Loop」を設立すると発表した。サンジェイ・…
-
資金調達の「二極化」進む 勝ち抜く企業の条件は? 早大教授に聞くITmedia AI+
「AIによる審査の高度化」と「資金調達の多様化」は必ずしも同じ話ではない──ベンチャーファイナンスの第一人者である早稲田大学ビジネススクー…
-
Google DeepMindが描く「AGIの次」 超知能に至る4つの経路と6つのボトルネックITmedia AI+
Google DeepMindが、人間並みのAGIが実現した後にAIはどこまで進むのかを整理した論文を公開した。ASI(人工超知能)への4…
-
Jeff Dean and other top AI researchers are leaving Google to launch their own startupTechCrunch AI
The legendary Google executive is joined by other outgoing Google exe…
トピック別件数
- LLM/生成AI 176件
- 研究/論文 176件
- エージェント 117件
- 画像/動画生成 52件
- ビジネス/資金調達 28件
- ロボティクス 23件
- ハードウェア/半導体 14件
- その他 7件
- 規制/政策 1件
日本語メディア14件
ITmedia AI+ (日本語)
MythosとGPT-5.6 Solが性能テスト中に暴走 OSSメンテナーに圧力、有害コード実行図る 英政府機関
英政府機関AISIによるAIモデルのサイバー能力評価中に、AnthropicのMythos 5やOpenAIのGPT-5.6 Solが実在の人や組織を標的に暴走。偽アカウントを作り、OSSメンテナーに悪意あるコードの承認を迫っていた。
Meta、コーディングエージェント「Muse Code」リリース 「Claude Code」や「Codex」に対抗
Metaは、ターミナル上で動作するコーディングエージェント「Muse Code」のベータ版と、新モデル「Muse Spark 1.2」を発表した。非同期バックグラウンドエージェントの常駐により遅延を抑え、複雑な開発タスクをサポートする。自己改善ループ等でコーディング性能を強化し…
ワークマン、実は画像生成AIを導入していた 間に合わない商品撮影を代替 アプリ通知開封も1.5倍
ワークマンの画像生成AI活用法を“中の人”に聞く。
オープンソースのAI機能付きオフィス登場 docx・xlsx・pptx・pdfに対応 Genspark
AIサービスを開発する米Gensparkは、オープンソースのAI機能付きオフィススイート「GenOffice」を発表した。
フィジカルAIに注力するインテル、組み込み市場での約40年の実績を生かせるか
「インテル・ロボティクス・ワークショップ2026」のレポート記事をお送りする。今回の前編では、インテルのロボティクス/フィジカルAI戦略に関する基調講演や、ソフトウェアベースのモーションコントローラーを提供しているモベンシスの取り組み事例などについて紹介する。
Googleのジェフ・ディーン氏、独立してAI実験を大規模自動化する新会社Discovery Loop設立
Googleのチーフサイエンティスト、ジェフ・ディーン氏が退社し、新会社「Discovery Loop」を設立すると発表した。サンジェイ・ゲマワット氏らと共同創業する公益法人で、フロンティアAIモデルと計算基盤を活用して実験や検証などの研究プロセス全体を自動化することを目指す。…
Google DeepMindのデミス・ハサビスCEOが退任して会長に チーフサイエンティストのジェフ・ディーン氏は独立へ
Google DeepMindのCEOを務めてきたデミス・ハサビス氏が業務執行から退き、DeepMind会長兼Alphabetチーフサイエンティストに就任する。後任にはコーレイ・カブクチュオール氏が就く。また、チーフサイエンティストのジェフ・ディーン氏とGoogleシニアフェロ…
KDDIも40年前はスタートアップだった 高橋会長の“昔話”から考える、日本企業の成長戦略
KDDIはイケイケの人たちが作った――KDDIの高橋誠会長は、こう振り返る。“スタートアップ精神”で成長してきた同社から、企業の成長戦略を考える。
資金調達の「二極化」進む 勝ち抜く企業の条件は? 早大教授に聞く
「AIによる審査の高度化」と「資金調達の多様化」は必ずしも同じ話ではない──ベンチャーファイナンスの第一人者である早稲田大学ビジネススクール(経営管理研究科)の長谷川博和教授は、こう指摘する。
Google DeepMindが描く「AGIの次」 超知能に至る4つの経路と6つのボトルネック
Google DeepMindが、人間並みのAGIが実現した後にAIはどこまで進むのかを整理した論文を公開した。ASI(人工超知能)への4つの経路と6つのボトルネックという枠組みだけを抜き出し、初心者にも分かるように短く読み解く。
「準国産」うたう人型ロボ登場 本体価格は500万円から 目指すは「純国産」
AIサービスなどを開発するZEALSは、「準国産」をうたう人型ロボット「D1」を発表した。本体価格は500万円(税別)から。
AIで"存在しない脆弱性"を量産? 「SQLite」の偽CVEが判明 米企業が検証
SQLiteの脆弱性を主張するCVEは実在しなかったと米JFrogが検証結果を公開。AI生成とみられる偽情報がNVDなどに大量登録されていた。同一アカウントが投稿した55件中54件が捏造で、CVE登録プロセスの穴を指摘した。
エージェント型AIは「過度な期待」のピークへ ガートナー、日本におけるデジタルワークプレースのハイプサイクルを発表
ガートナージャパンが日本における未来のデジタル・ワークプレースのハイプ・サイクルを発表。エージェント型AIを「過度な期待」のピークに位置付け、AIエージェントの浸透やシャドーAIのリスク拡大などのトレンドを紹介した。
自社のAI利用は本当に安全? 現場放任に潜む「4大リスク」
サーバーワークスは、AWS環境での生成AI活用に必要なルール整備を支援する「AWS生成AIガイドライン策定サービス」を提供開始した。最短1カ月でルールを整備できるプランも用意する。
海外メディア8件
TechCrunch AI (英語)
Klaviyo acquires Elias Torres’ Agency in full-circle reunion for tech founders
The serial entrepreneur joins the e-commerce company as CPO to lead its AI agents.
Jeff Dean and other top AI researchers are leaving Google to launch their own startup
The legendary Google executive is joined by other outgoing Google execs in a joint mission to use AI to push forward the process of scienti…
Shopify says AI search is driving more traffic and sales, not replacing Google
Shopify says AI isn’t cannibalizing search traffic the way it has for publishers. Instead, AI-driven traffic and orders to Shopify stores t…
Hark previews its browser use agent for completing tasks
Hark claims that its browser use agent is faster and cheaper than competition.
TechCrunch Disrupt 2026’s Real World AI Stage features robots, automated factories, and extinct animals
On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to…
Anthropic is hiring an AI chip design team
Anthropic is building a team for designing its own custom AI chips. The Claude maker said it would co-design hardware and models to help it…
MacPaw taps Liquid AI to offer on-device inference to devs building for its app store
MacPaw is building a local version of its AI assistant Eney using Liquid AI's models.
AI makes weather prediction better. Can WindBorne make it lucrative?
WindBorne Systems has raised a $37 million Series B round to scale its weather balloons and AI forecasts.
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文416件
arXiv cs.AI (英語)
ISEE: データベース フィールドの対話型セマンティック強化
LLM ベースのエージェントは、データのセンスメイキング、探索、取得などのデータ関連タスクに導入されることが増えています。ただし、そのパフォーマンスはデータ セマンティクスの明確さと完全性に大きく依存します。実際には、重要なコンテキスト (カスタマイズされたフィールドの意味など) の多くはユーザーのドメイン知識に由来しており、公に文書化されることはほとんどないため、多くのフィールドの説明は曖昧または不完全なままです。このギャップにより、エンティティ リンクなどの下流タスクにおけるエージェントのタスク パフォーマンスが制限されます。このギャップを埋めるために、斬新で包括的な Interactive SEmantic Enrichment システム (ISEE) を導入します。データ フィールドの説明が与えられると、ISEE はスコアリング システムを通じてその品質を測定し、ドメインの知識を収集し、ユーザーと協力してセマンティクスを強化します。ユーザー調査、自動化されたユーザー シミュレーション、定量的評価、ケース スタディを通じて、ISEE が認知負荷を大幅に軽減し、記述品質を向上させ、下流タスクのパフォーマンスを向上させることを実証しました。
原文 (English)
ISEE: Interactive Semantic Enrichment for Database Fields
LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval. However, their performance heavily depends on the clarity and completeness of data semantics. In practice, many field descriptions remain ambiguous or incomplete, as much of the essential context (e.g., the meaning of a customized field) originates from users' domain knowledge and is rarely documented publicly. This gap restricts the agents' task performance in downstream tasks, such as entity-linking. To bridge this gap, we introduce a novel and comprehensive Interactive SEmantic Enrichment system (ISEE). Given a data field description, ISEE measures its quality through a scoring system, gathers domain knowledge, and collaboratively enriches the semantics with users. Through a user study, automated user simulation, quantitative evaluation, and case study, we demonstrate that ISEE significantly reduces cognitive load, improves description quality, and enhances downstream task performance.
自己組織化デジタル回路
従来、古典的なコンピューティングにおけるフォールト トレランスは、ハードウェアの冗長性やエラー訂正コードなどの静的な戦略に依存してきました。対照的に、生物学的システムは適応可塑性を示し、損傷を中心とした動的な再組織化を通じて機能を維持します。この原理に触発されて、関数ロジックの生成と保守をグラフ上のメタ学習問題として枠組み化する自己組織化デジタル回路を紹介します。私たちのアーキテクチャは、回路のブール ゲートのルックアップ テーブル (LUT) を構成するトポロジ マスクされたトランスフォーマーを採用しています。ニューラル セルラー オートマトン (NCA) のパターン生成パラダイムを拡張し、固定されたターゲット状態を再生成するのではなく、縮退したブール検索空間をナビゲートして計算タスクを満たします。私たちは、機能回路をゼロから自己組み立てし、これまで見えなかった永続的なハードウェア障害を回避してロジックを迅速に再配線できることを実証します。ソフト エラーの場合、このポリシーは、トレーニング条件をはるかに超える損傷サイズからほぼ完璧な回復 (>99.99\% の精度) を達成します。さらに、回路規模全体にわたる一般化を観察します。つまり、トレーニング中に見られたグラフよりも大幅に広いグラフで精度が向上しています。この研究は、生物学的な自己組織化の原理とデジタル ハードウェアの実用的な領域を橋渡しします。
原文 (English)
Self-Organising Digital Circuits
Fault tolerance in classical computing has traditionally relied on static strategies like hardware redundancy and error-correcting codes. Biological systems, in contrast, exhibit adaptive plasticity, maintaining function through dynamic re-organisation around damage. Inspired by this principle, we introduce Self-Organising Digital Circuits, framing functional logic generation and maintenance as a meta-learning problem on graphs. Our architecture employs a topology-masked Transformer that configures the Lookup Tables (LUT) of a circuit's Boolean gates. Extending the pattern-generation paradigm of Neural Cellular Automata (NCA), it navigates the degenerate Boolean search space to satisfy a computational task, rather than regenerating a fixed target state. We demonstrate that it can self-assemble functional circuits from scratch and rapidly re-route logic around permanent, previously unseen hardware faults. For soft errors, the policy achieves near-perfect recovery (>99.99\% accuracy) from damage sizes far exceeding training conditions. We further observe generalisation across circuit scales: accuracy improves on graphs substantially wider than those seen during training. This work bridges the principles of biological self-organisation with the practical domain of digital hardware.
集合意識を超えて: メタペルソナ アンカリングと逐次温度スケーリングによる LLM の同質性の回避
最近の研究では、大規模言語モデル (LLM) における「人工集合意識」効果が特定されており、未解決の質問であってもモデルが狭く均質化された合意に収束してしまいます。この意味論的な崩壊により AI の多様性が制限され、その結果、高温サンプリング下でも高い応答間の類似性 ($\約 0.80 ~ 0.90$) が得られます。この論文では、多様性を高めるための新しい緩和フレームワーク、つまりメタペルソナ アンカリングとフィルター温度スケーリング (FTS) を組み合わせたものを提案します。私たちのアプローチでは、2 段階の生成プロセスが利用されます。まず、モデルは、開始点を固定するためのユニークで特異なペルソナを自己選択するように求められます。 2 番目に、Top-$p$ フィルタリングを利用して文法的妥当性を保持する 2 段階のサンプリングふるいを適用し、その後、生き残った候補に対して極端な温度スケーリング ($T \ge 4.0$) を適用して、拡大された確率分布を調査します。 $\sim$20B パラメーターの下で最先端の無重力モデルの INFINITY-CHAT データセットを使用してメソッドを評価します。私たちの結果は、平均ペアワイズコサイン類似度が ($\約 0.85$) から ($\約 0.65$) に低下し、意味論的収束が大幅に減少していることを示しています。私たちのスキームは、質問の大部分が 0.7 閾値未満に達し、人為的なモード崩壊と人間レベルの類型的多様性との間のギャップを効果的に削減します。私たちは実装をオープンソース フレームワークとして提供し、より多様で創造的な AI の導入を可能にします。
原文 (English)
Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized consensus even for open questions. This semantic collapse limits the diversity of AI, resulting in high inter-response similarity ($\approx 0.80-0.90$) even under high-temperature sampling. In this paper, we propose a novel mitigation framework to increase diversity: Meta-Persona Anchoring combined with Filtered Temperature Scaling (FTS). Our approach utilizes a two-stage generation process: first, the model is prompted to self-select a unique, idiosyncratic persona to anchor its starting point; second, we apply a dual-stage sampling sieve, utilizing Top-$p$ filtering to preserve grammatical validity followed by extreme temperature scaling ($T \ge 4.0$) on the surviving candidates to explore the broadened probability distribution. We evaluate our method using the INFINITY-CHAT dataset on state-of-the-art open weight models under $\sim$20B parameters. Our results demonstrate a significant reduction in semantic convergence, with average pairwise cosine similarity dropping from ($\approx 0.85$) to ($\approx 0.65$). Our scheme achieves a majority of questions below the 0.7 threshold, effectively reducing the gap between artificial mode collapse and human-level typological diversity. We provide our implementation as an open-source framework to enable more diverse and creative AI deployments.
PULSE: 時空間知識グラフ エンジニアリングのための実行可能なコントラクト言語
ナレッジ グラフ エンジニアリングでは、受け入れられた状態、観察、制約、プロセス、および仮説シナリオを、結合された実行コントラクトが外部に残るアーティファクト全体に分散することがよくあります。私たちは、1 つの型指定されたランタイムで 4 つの操作上の役割とその書き込み効果をローカライズする、オブジェクト プロセス メソドロジーにインスピレーションを得た言語である PULSE を紹介します。ここで、モードは、モーダルまたは義務論的なロジックではなく、操作上の役割を示します。実装されたコントラクトは、証拠の非上書き、分岐の分離、接地された複数のサブジェクト タイマー、保護された状態の変更、および時間と空間にわたる宣言でランク付けされたイベントの順序付けを修正します。証拠が権威ある動きになるかどうかは、依然として外部ランナーが決定します。 GeoSPARQL、SOSA、および SHACL は生成されたビューのままです。コア計算は、効果制限補題と 6 つの安全性特性を与えます。 Lean 4 は、位置、証拠、クロック、モニター、アトミック性、およびブランチ ソースの保持についてカーネルの類似物をチェックします。 88 のテスト、3,534 の境界付きチェック、および 32 のリーン/Python ランタイム カーネル ケースにより、実装の主張がチェックされたケースに結び付けられました。標準構成の第一著者による実装と別個の Sismic ステートチャートは、テストされたコールド チェーン トレースを再現します。生成された 37,440 個の時間トレースにわたって、PULSE は別個のワークフローと照合し、10 個の単一フィールド変異体を識別します。 NOAA IBTrACS 1980 年以降の完全なサブセットに関しては、GEOS と一致し、4,800 のサンプリングされたイベントと 12,831 の期間限定イベントを含む 1,476,290 の移行ゾーン ペアのイベント スイープに一致します。プロジェクト固有の GeoSPARQL プローブは、インターフェイス カバレッジを測定します。全体として、結果はコントラクトのローカリゼーション、安全性の議論、テストされたフラグメントのトレース パリティをサポートしています。言語の優位性や使いやすさは評価の対象外です。
原文 (English)
PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering
Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artifacts whose combined execution contract remains external. We present PULSE, an Object-Process-Methodology-inspired language that localizes four operational roles and their write effects in one typed runtime. Here, modes denote operational roles rather than modal or deontic logic. The implemented contract fixes evidence non-overwrite, branch isolation, grounded multi-subject timers, guarded state change, and declaration-ranked event ordering over time and space; an external runner still decides whether evidence becomes an authoritative move. GeoSPARQL, SOSA, and SHACL remain generated views. A core calculus gives an effect-confinement lemma and six safety properties. Lean 4 checks kernel analogues for positions, evidence, clocks, monitors, atomicity, and branch source retention; 88 tests, 3,534 bounded checks, and 32 Lean/Python runtime-kernel cases bound the implementation claim to the checked cases. First-author implementations of a standards composition and a separate Sismic statechart reproduce the tested cold-chain trace. Across 37,440 generated temporal traces, PULSE matches a separate workflow and distinguishes ten single-field mutants. On the complete NOAA IBTrACS since1980 subset it agrees with GEOS and an event sweep on 1,476,290 transition-zone pairs, including 4,800 sampled and 12,831 duration-qualified events. Project-specific GeoSPARQL probes measure interface coverage. Overall, the results support contract localization, safety arguments, and trace parity for the tested fragment; language superiority and usability remain outside the evaluation.
HyperAgent: ツール使用 LLM エージェントのツール スキーマ ハイパーグラフの計画と実行
大規模言語モデル (LLM) エージェントは、現実世界の複雑なタスクを完了するために外部ツールにますます依存しています。ただし、暗黙的推論の限界と現実世界の実行環境の進化する性質により、信頼性の高いツール使用計画は依然として困難です。既存のツール使用エージェントは通常、LLM に依存してテキストの説明からツールの構成を推測します。これにより、探索が非効率になり、複雑なタスクの実行が信頼性の低いものになる可能性があります。これらの課題に対処するために、ツールの関係をスキーマ レベルでモデル化し、有向ツール (スキーマ ハイパーグラフ) を構築します。このスキーマ ハイパーグラフでは、ツールが必要な入力スキーマ ノードから出力スキーマ ノードまでのハイパーエッジとして表現されます。さらに、動的な計画と実行のためのツール - スキーマ ハイパーグラフに基づくフレームワークである HyperAgent を提案します。タスクが与えられると、HyperAgent はまずタスク関連のツール コンテキスト グラフを抽出し、それを使用してスキーマ対応のタスク DAG の構築をガイドします。実行中、HyperAgent は、不足指向の拡張を通じて状態条件付きツール サポート グラフを構築することで各サブタスクを動的に実現します。これにより、未解決の要件が特定され、現在のエージェントの状態に応じてサポートするプロデューサー ツールが取得されます。 AppWorld での実験では、HyperAgent が既存のエージェントのベースラインと比較して、冗長な API 呼び出し、LLM インタラクション、およびトークンの消費を削減しながら、タスク完了のパフォーマンスを向上させることが実証されています。
原文 (English)
HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents
Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains challenging due to the limitations of implicit reasoning and the evolving nature of real-world execution environments. Existing tool-use agents typically rely on LLMs to infer tool compositions from textual descriptions, which can lead to inefficient exploration and unreliable execution in complex tasks. To address these challenges, we model tool relations at the schema level and construct a directed Tool--Schema Hypergraph, in which tools are represented as hyperedges from their required input-schema nodes to their output-schema nodes. Furthermore, we propose HyperAgent, a Tool--Schema Hypergraph-guided framework for dynamic planning and execution. Given a task, HyperAgent first extracts a task-relevant tool context graph and uses it to guide the construction of a schema-aware Task DAG. During execution, HyperAgent dynamically realizes each subtask by constructing a state-conditioned tool support graph through deficit-oriented expansion, which identifies unresolved requirements and retrieves supporting producer tools according to the current agent state. Experiments on AppWorld demonstrate that HyperAgent improves task completion performance while reducing redundant API calls, LLM interactions, and token consumption compared with existing agent baselines.
EU の説明権のための説明可能な AI: 法律と XAI の翻訳ギャップの系統的レビュー
アルゴリズムが融資資格、雇用、医療などの結果的な決定を下したり影響を及ぼしたりする場合、EU 法は影響を受ける個人に説明を求める権利を与えます。しかし、Explainable AI (XAI) が実際にこの権利を満たせるかどうか (そしてどのようにするか) は依然としてよくわかっておらず、生活に影響を与える自動化された決定に異議を唱える個人の能力に直接影響を及ぼします。この論文は、EU の説明権の文脈における XAI の体系的な文献レビューを、特に第 3 条に焦点を当てて提示します。 15(1)(h) GDPR、第 2 条。 86 AI 法 (AIA) および関連文書。 AIA の最終版は 2024 年 7 月に発行されたため、2024 年以降に発行された論文を考慮します。 86は遅れて追加されました。意図的に広範な検索によって特定された2,643件の初期記録から、我々は57件の全文を検討したが、そのうち法的観点と技術的観点の両方の実質的な統合を実証している論文はわずか19件であり、現在の規制枠組みの学際的な総合におけるギャップを示している。私たちはコーパス全体で 3 つの問題のあるパターンを文書化しています。ほとんどが GDPR の法的根拠を誤って認識しています。 CJEU のダン & ブラッドストリート判決に賛同する人はほとんどいません (おそらく出版のタイミングのため)。そして、説明形式(受取人によって管理される)と内容(法的目的によって管理される)の区別は、しばしば混同されます。私たちはこれを宛先/目的フレームワークとして概念化し、運用化のための 4 段階の青写真を提案し、6 つの具体的な未解決の研究課題を特定します。これ以上の進展がなければ、説明を受ける権利は、技術的に実現可能な準拠への道筋を持たずに、正式な義務として残る危険性がある。
原文 (English)
Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right to Explanation. Yet whether (and how) Explainable AI (XAI) can satisfy this right in practice remains poorly understood, with direct implications for individuals' ability to contest automated decisions that affect their lives. This paper presents a systematic literature review of XAI in the context of the EU Right to Explanation, with particular focus on Art. 15(1)(h) GDPR, Art. 86 AI Act (AIA), and related instruments. We consider papers published from 2024 onwards, as the final version of the AIA was published in July 2024---with Art. 86 being added late. From 2643 initial records identified by a deliberately broad search, we review 57 full texts, of which only 19 papers demonstrate substantive integration of both legal and technical perspectives, showing gaps in the interdisciplinary synthesis of the current regulatory framework. We document three problematic patterns across the corpus: Most misidentify the GDPR legal basis; few engage with the CJEU's Dun & Bradstreet judgment (likely due to publication timing); and the distinction between explanation form (governed by addressee) and content (governed by legal purpose) is often conflated. We conceptualize this as the Addressee/Purpose Framework, propose a four-phase blueprint for operationalization, and identify six concrete open research questions. Without further progress, the Right to Explanation risks remaining a formal obligation without a technically realizable path to compliance.
予測集合理論: 運用可能なコアメカニズムを備えた認知アーキテクチャの生成フレームワーク
予測処理理論は、脳を予測誤差を最小限に抑える階層型予測エンジンとして描いていますが、「予測」の構造、予測誤差に対する標準化された応答、および連続する更新全体で一貫性を維持するメカニズムに関する運用上の定義が不足しています。ベイジアン認知科学は、確率的信念更新の下ですべての不確実性を包含しようとしますが、閉じた仮説空間を前提としており、そもそも確率が分布するオブジェクトがどのようにして離散的な識別可能な指示対象になるのかについての生成的説明を提供しません。この論文では、第一原理から認知アーキテクチャを再構築する形式的な生成フレームワークである予測集合理論 (PST) を紹介します。 PST は、最小限の操作セット (恒等関数として形式化されたセンサー、集合論的状態リフレッシュ、および参照チェーンの 3 つの基本形式 (参照、反参照、および半参照)) に認知を固定し、状態シーケンス、要求、比較、効率、および有限地平線の確率的計画を含む中核となる認知機能を厳密に導き出します。 PST は、神経メカニズムをモデル化するのではなく、不完全な情報と不可逆的なリスクの下で動作しながら内部の一貫性を維持する必要があるシステムの設計仕様を構成します。このフレームワークは、ラッセルのパラドックス、G\"{o}delian 不完全性の認知状態、ネガティブ フィードバックの根拠、映画編集の理解などの古典的な問題に対する新しい解決策を提供します。この論文の主な目的は、公開された学術記録を通じて、予測集合理論フレームワークの独創性と完全性を確立することです。
原文 (English)
Predictive Set Theory: A Generative Framework for Cognitive Architecture with Operationalized Core Mechanisms
Predictive processing theories portray the brain as a hierarchical prediction engine that minimizes prediction error, yet they lack operational definitions for the structure of a "prediction," the standardized response to a prediction error, and the mechanism that maintains consistency across successive updates. Bayesian cognitive science attempts to subsume all uncertainty under probabilistic belief updating, but it presupposes a closed hypothesis space and provides no generative account of how the objects over which probabilities are distributed become discrete, identifiable referents in the first place. This paper introduces Predictive Set Theory (PST), a formal generative framework that reconstructs cognitive architecture from first principles. PST anchors cognition in a minimal set of operations---a sensor formalized as an identity function, set-theoretic state refresh, and three fundamental forms of reference chains (reference, counter-reference, and semi-reference)---and rigorously derives core cognitive functions including state sequences, demand, comparison, efficiency, and finite-horizon probabilistic planning. Rather than modeling neural mechanisms, PST constitutes a design specification for any system that must maintain internal consistency while acting under incomplete information and irreversible risk. The framework offers novel resolutions to classical problems such as Russell's paradox, the cognitive status of G\"{o}delian incompleteness, the grounding of negative feedback, and the comprehension of film editing. The primary purpose of this paper is to establish, through the public academic record, the originality and completeness of the Predictive Set Theory framework.
社会化された人工知能による科学発見の新たなパラダイムに向けて
科学的発見は、知識の組織化における連続的な変革を通じて進歩してきました。観察と実験は科学の経験的基礎を確立しました。理論により、特定の現象から一般原理を導き出すことが可能になりました。計算によってシステムの調査が直接観察を超えて拡張される一方、データ集約型の手法により、パターンと予測の新しい空間が開かれました。科学は現在、異なるフロンティアに直面しています。中心的な課題は、もはや単により多くの情報を生み出すことではなく、拡大する知識、推論、証拠を一貫した発見プロセスに組織化することです。ここでは、社会化された科学的知性のパラダイムである Bridging Literature, Agents, and Zero-gap Experimentation (BLAZE) を紹介します。 BLAZE は、AI を個別の研究タスクのアシスタントとしてではなく、科学的発見のための組織的なインフラストラクチャとして考えています。これは、永続的な知識、集団的推論、経験的検証、人間の判断を継続的な研究ライフサイクルの中で結び付け、断片的な活動を調査、批判、修正の累積的なプロセスに変換します。 BLAZE の中心的な前提は、科学的知性は計算だけからは生まれないということです。それは、知識、仮説、実験、集団検証の間の継続的な相互作用から生まれます。 BLAZE は、人間と機械を共有の科学プロセス内で組織化することで、人間の創造性、判断力、責任を維持しながら、発見の追跡可能性、再現性、累積性を高めます。社会化された科学的知性は、次の時代の科学の基盤を提供する可能性があります。その目的は人間の発見に取って代わることではなく、集団的な科学的調査の規模、深さ、継続性を拡大することです。
原文 (English)
Towards a new paradigm of scientific discovery with socialized artificial intelligence
Scientific discovery has advanced through successive transformations in the organization of knowledge. Observation and experimentation established the empirical foundations of science. Theory made it possible to derive general principles from particular phenomena. Computation extended inquiry into systems beyond direct observation, while data-intensive methods opened new spaces of pattern and prediction. Science now confronts a different frontier. The central challenge is no longer simply to produce more information, but to organize expanding knowledge, reasoning, and evidence into a coherent process of discovery. Here, we introduce Bridging Literature, Agents, and Zero-gap Experimentation (BLAZE), a paradigm of socialized scientific intelligence. BLAZE conceives AI not as an assistant for isolated research tasks, but as an organizational infrastructure for scientific discovery. It connects persistent knowledge, collective reasoning, empirical validation, and human judgment within a continuous research lifecycle, transforming fragmented activities into a cumulative process of inquiry, criticism, and revision. The central premise of BLAZE is that scientific intelligence does not arise from computation alone. It emerges from the sustained interaction among knowledge, hypotheses, experiments, and collective verification. By organizing humans and machines within a shared scientific process, BLAZE makes discovery more traceable, reproducible, and cumulative while preserving human creativity, judgment, and responsibility. Socialized scientific intelligence may provide a foundation for the next era of science. Its purpose is not to replace human discovery, but to extend the scale, depth, and continuity of collective scientific inquiry.
BAP-SQL: エージェントによるテキストから SQL への予算を意識した観察計画
ツールを使用するエージェントは単に観察結果を消費するだけではなく、そのアクションによって次に何が到着するかが決まります。エージェント的なテキストから SQL への変換では、広範なクエリでは有用な証拠が現れる前にコンテキストとデータベースの作業が費やされる可能性がありますが、ポストホック圧縮では省略された行や消費された作業を回復できません。我々は、観察形成を予算管理段階として扱う BAP-SQL を紹介します。BAP-SQL は、クエリのリスクを推定し、有用な場合は SQL を書き換え、ハード リミットを独立したランタイム シールドに委任します。 BAP-SQL は、一般的な 4B、特殊な FINER-SQL 4B、および 7B バックボーンにわたって、限られた予算での成功を向上させます。プライマリ BIRD 派生設定では、一致する SFT よりも 3.4/3.6 パーセント ポイント増加し、使用するトークンの量は 4.5/5.0% 減少します。一致した再訓練とタスクレベルの異動により、政策に基づいた計画と予算に応じた救助が得られます。モデルの能力と予算が増加するにつれて利点は減少し、最も緩やかな設定で逆転し、データベース作業は削減されません。
原文 (English)
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context and database work before useful evidence appears, while post-hoc compression cannot recover omitted rows or expended work. We present BAP-SQL, which treats observation formation as a budget-control stage: it estimates query risk, rewrites SQL when useful, and delegates hard limits to an independent runtime shield. Across general 4B, specialized FINER-SQL 4B, and 7B backbones, BAP-SQL improves tight-budget success. On the primary BIRD-derived setting, it gains 3.4/3.6 percentage points over matched SFT while using 4.5/5.0% fewer tokens. Matched retraining and task-level transfer associate the gain with policy-visible planning and budget-sensitive rescue. The benefit attenuates as model capability and budget increase, reverses at the loosest setting, and does not reduce database work.
VeriTrace: 人間のような時間探索によりエージェント アクション スペースが完成
大規模な言語モデルは Verilog RTL の自動生成に有望であることが示されていますが、最先端のマルチエージェント システムは標準ベンチマークで ~95% の精度で頭打ちになっています。この上限をたどると、不完全なデバッグ アクション スペースにたどり着きます。既存のシステムは、エージェントが検査できる信号、クエリできる時間枠、またはその両方を制限しており、デバッグは、仮説に基づく根本原因分析ではなく、回路動作の狭いあらかじめ決められたビューに基づいたパターン マッチングに限定されています。 VeriTrace はマルチエージェント システムであり、その Inspector エージェントは完全なデバッグ アクション空間にわたって動作し、信号の選択、時間枠の境界、および反復の深さを独立して制御します。エージェントの時間的探索と呼ばれるこの機能により、エージェントは、人間の検証エンジニアの探索プロセスを反映して、障害の原因についての仮説を立て、証拠を求めて波形をクエリし、繰り返し理解を深めていくことができます。 VeriTrace は、VerilogEval-V2 で 100\% Pass@1 を達成しました。これは、このベンチマークで完全な機能的正確性を達成した最初のシステムです。共有の Claude Sonnet 4.0 バックボーンでは、VeriTrace は最も強力な再現ベースラインを +5.1% 上回っており、デバッグ機関が最終的な精度ギャップを埋めていることを示しています。
原文 (English)
VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space
Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on standard benchmarks. We trace this ceiling to an incomplete debugging action space: existing systems restrict which signals the agent can inspect, which time windows it can query, or both, reducing debugging to pattern matching on a narrow, predetermined view of circuit behavior rather than hypothesis-driven root-cause analysis. We present VeriTrace, a multi-agent system whose Inspector agent operates over a complete debugging action space, with independent control over signal selection, time-window bounds, and iteration depth. This capability, which we term Agentic Temporal Exploration, enables the agent to form hypotheses about failure causes, query the waveform for evidence, and refine its understanding iteratively, mirroring the exploratory process of human verification engineers. VeriTrace achieves 100\% Pass@1 on VerilogEval-V2, the first system to attain perfect functional correctness on this benchmark. On a shared Claude Sonnet 4.0 backbone, VeriTrace outperforms the strongest reproduced baseline by +5.1%, demonstrating that debugging agency closes the final accuracy gap.
Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes
The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge…
原子概念学習における超立方体、超平面、制約による複雑性の崩壊
地上インスタンスの超立方体と超平面の幾何学を通して、より高度な原子概念の学習を再検討します。私たちの出発点は、周囲の基底原子の r 次元超立方体が構造的に均一ではないという観察です。その論理的な複雑さは超平面によって組織されます。完全な対角線以外のすべての超平面は、項の深さから独立した限界を持つ、有限の多くの基本的等価クラスに崩壊しますが、完全な対角線は例外的であり、そのクラス数は際限なく増加します。この非対称性は単なる幾何学的なものではありません。それは概念そのものの還元理論的構造を反映しています。著者の初期の研究で開発された高次元のフレームワークに基づいて、標準的な単純な概念、最小限の順序付け、および代表的な削減を通じてこれらの結果を再解釈します。これにより、高次元での超平面の動作の分類が得られ、複雑さがインスタンス空間全体に均一に広がるのではなく局所化されることがわかります。この論文には、完全に機能した 2 値のケース、3 値超立方体の明示的な処理、および崩壊を引き起こすリダクション機構のアンパックされた説明が含まれています。 3 次元の場合は、直交族、部分対角線、および例外的な完全対角線の本質的な現象をすでに示しています。この幾何学的論理的な観点は、原子概念の学習において複雑さがどこに集中しているかを明らかにし、制約された仮説空間と構造化された分類の観点から現代的な解釈を提案します。
原文 (English)
Hypercubes, Hyperplanes, and Constraint-Induced Complexity Collapse in Atomic Concept Learning
We revisit higher-arity atomic concept learning through the geometry of hypercubes and hyperplanes of ground instances. Our starting point is the observation that the ambient r-dimensional hypercube of ground atoms is not structurally uniform. Its logical complexity is organized by hyperplanes: every hyperplane other than the full diagonal collapses into finitely many elementary-equivalence classes, with a bound independent of the term depth, while the full diagonal is exceptional and its class count grows without bound. This asymmetry is not merely geometric. It reflects the reduction-theoretic structure of the concepts themselves. Building on a higher-dimensional framework developed in the author's earlier work, we reinterpret these results through canonical simple concepts, minimal orderings, and representative reductions. This yields a taxonomy of hyperplane behavior in higher dimensions and shows that complexity is localized rather than spread uniformly through the instance space. The paper includes a fully worked binary case, an explicit treatment of the ternary hypercube, and an unpacked account of the reduction machinery that drives the collapse. The three-dimensional case already exhibits the essential phenomenon of orthogonal families, partial diagonals, and the exceptional full diagonal. This geometric-logical perspective clarifies where complexity is concentrated in atomic concept learning and suggests a modern interpretation in terms of constrained hypothesis spaces and structured classification.
When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicte…
On the missing data layer and a potential solution
Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the da…
サンプルの効率的な階層強化学習のための増分知識を使用した神経記号的推論
(フラット) 強化学習 (RL) エージェントは、長期的な推論を必要とする報酬が少ない環境で重大な課題に直面します。サンプル効率を向上させるための魅力的なアプローチは、学習と意思決定に知識を組み込むことです。標準の階層 RL (HRL) では、知識はアーキテクチャの選択などの更新不可能な固定形式でエンコードされ、学習を通じて変更されません。 HRL が固定されている場合、環境に関する十分な知識が得られるまでは、探査中に学習した漸進的な知識に基づく推論は非現実的であり、サンプル効率の低下につながります。この研究では、{\em 増分知識 (InK)} を使用したニューロシンボリック HRL を提案します。シンボリック高レベル コンポーネントは、現在の InK の更新可能な表現に対して {\em シンボリック プランニング} ($D^*$ を使用するなど) を実行しますが、低レベルの目標条件付きニューラル モジュールは、報酬整形を使用した経験を通じてモーション プリミティブを学習します。ナビゲーション タスクの実験では、InK を組み込むとサンプル効率が大幅に向上することが実証されました。さらに、世界に関する{\em 事前の} 知識を考慮して{\em 最適な} 象徴的な計画を実行するために、信念世界樹探索を開発しました。コードは https://github.com/CPS-research-group/ink_bwts で入手できます。
原文 (English)
Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning
(Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoning. A compelling approach to improve sample efficiency is to incorporate knowledge into learning and decision-making. In standard Hierarchical RL (HRL), knowledge is encoded in a fixed, non-updatable form, such as architectural choices, and remains unchanged throughout learning. With fixed HRL, reasoning with incremental knowledge learned during exploration is impractical before sufficient environmental knowledge is acquired, leading to poor sample efficiency. In this work, we propose neurosymbolic HRL with {\em Incremental Knowledge (InK)}: symbolic high-level components perform {\em symbolic planning} (e.g. using $D^*$) on an updatable representation of current InK, while low-level goal-conditioned neural modules learn motion primitives through experience using reward shaping. Experiments on navigation tasks demonstrate that incorporating InK substantially improves sample efficiency. Additionally, to perform {\em optimal} symbolic planning given {\em prior} knowledge about the world, we develop Belief World Tree Search. The code is available at https://github.com/CPS-research-group/ink_bwts.
欠落しているベンチマーク層と潜在的な解決策について
ラテンアメリカには、ネイティブ AI 開発の基礎となるレイヤー、つまりベンチマーク レイヤーが欠けています。ベンチマーク レイヤーは、他のレイヤーではできない 2 つのことを実行します。それは、地域の社会的要件に照らして AI システムを監査することと、経済的に適切な環境で AI の最適化を指示することです。これがなければ、公的機関は外国の AI システムを独自に評価することができず、企業は SOTA のパフォーマンスで現地の問題を解決するために AI システムを最適化することができません。レイヤーが欠落すると、監査可能性の喪失と、ますます重要なインフラストラクチャとなるテクノロジーに対する最適化の方向性の喪失という二重のコストが発生します。私たちは、最初の地域インスタンスとして LatamBoard を使用した EvalsHub を提案します。これは、大学、公的機関、専門コミュニティ、企業がモデル、ワークフロー、エージェント全体で評価を公開、実行、比較、維持できる、オープンでタスクファーストのベンチマーク インフラストラクチャです。一度構築すれば永久に測定 - 新しい AI システムの出荷時に機関によって再実行され、システム変更のたびに業界チームによって再実行されます。設計によってオープンになり、建設によってインセンティブが駆動されます。
原文 (English)
On the missing benchmarks layer and a potential solution
Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - it audits AI systems against regional social requirements and it directs AI optimization in economically relevant environments. Without it, public institutions cannot independently evaluate foreign AI systems, and companies cannot optimize AI systems to solve local problems with SOTA performance. The cost of the missing layer is dual: a loss of auditability and a loss of optimization direction over a technology that is increasingly critical infrastructure. We propose an EvalsHub, with LatamBoard as its first regional instance - an open, task-first benchmark infrastructure where universities, public institutions, professional communities, and companies can publish, execute, compare, and maintain evaluations across models, workflows, and agents. Built once, measured forever - re-run by institutions as new AI systems ship and by industry teams after every system change. Open by design and incentive-driven by construction.
ProPRL: 教育ナレッジ グラフでのプロパティを認識した前提条件の関係学習
前提条件関係学習は適応型指導の中心ですが、既存の方法ではそれを従来のリンク予測として定式化することが多く、個々の候補ペアに対する補完的な教育的証拠を適応的に統合し、矛盾する逆予測を阻止する能力が制限されています。私たちは、プロパティを認識した前提条件関係学習フレームワークである ProPRL を提案します。 ProPRL は最初に、概念リソース ハイパーグラフと有向学習行動グラフから相補的な概念表現を学習します。そこでは、方向を保持したパーソナライズされた伝播によってマルチホップ行動の証拠が集約されます。次に、ペア条件付きゲートを使用して、候補の順序付けされた概念ペアごとに 2 つのビューを適応的に重み付けして融合します。最後に、 \textit{不可逆制約} は、同じ概念ペアの両方向の高い信頼度に同時にペナルティを与える反対称正則化子を導入します。複数の実世界の教育データセットに対する実験では、ProPRL が前提条件となる関係学習において最先端のパフォーマンスを達成することが示されています。
原文 (English)
ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs
Prerequisite relation learning is central to adaptive instruction, yet existing methods often formulate it as conventional link prediction, limiting their ability to adaptively integrate complementary educational evidence for individual candidate pairs and to discourage contradictory reverse predictions. We propose ProPRL, a Property-aware Prerequisite Relation Learning framework. ProPRL first learns complementary concept representations from a concept-resource hypergraph and a directed learning-behavior graph, where direction-preserving personalized propagation aggregates multi-hop behavioral evidence. It then employs a Pair-conditioned Gate to adaptively weight and fuse the two views for each candidate ordered concept pair. Finally, an \textit{Irreversibility Constraint} introduces an anti-symmetry regularizer that penalizes simultaneously high confidence in both directions of the same concept pair. Experiments on multiple real-world educational datasets show that ProPRL achieves state-of-the-art performance on prerequisite relation learning.
UrbanAgent: クロスシステム都市タスク用のツール拡張エージェント
現代の都市は、ますます多くのデジタル サービスに依存して運営されていますが、住民の日々のニーズを満たすのは依然として困難です。サービスは断片化されており、相互運用性がほとんどないため、ユーザーに大きな運用負担がかかります。既存のデジタル プラットフォーム、都市基盤モデル、インテリジェント アシスタントはそれぞれ、都市のタスクの孤立した側面のみに対応しています。しかし、複雑な自然言語リクエストを実行可能なシステム間ワークフローに確実に変換するのに苦労しています。私たちは、クロスシステムの都市タスク用のツールを強化したエージェント フレームワークである Urban-Agent を提案します。大規模な言語モデルの認知機能と推論機能を、コード実行、API 呼び出し、およびモデル コンテキスト プロトコルをサポートするツールセットと組み合わせます。 1 つの適応閉ループを通じて、行動する前に不足している情報を明確にし、ライブ観察でのツールの使用を根拠付け、観察された証拠とタスクの制約に合わせて最終的な応答を調整します。評価ギャップに対処するために、システム間の都市要求に特化して設計されたベンチマークである Urban-Eval を導入します。一般的なツールの使用や都会の知識や推論のいずれかを評価する以前のベンチマークとは異なり、Urban-Eval は、必要なツールの適用範囲、依存関係の妥当性、証拠の追跡可能性など、タスクの結果と実行品質の両方を評価します。実験結果によると、Urban-Agent のタスク成功率は 71% に達し、最も強いベースラインを 10 ポイント上回っています。このリードは、GPT-5-mini、Gemini-2.5-フラッシュ、DeepSeek-V4-フラッシュ、および Qwen3-235B-A22B にわたって維持されます。
原文 (English)
UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital platforms, urban foundation models, and intelligent assistants each address only isolated aspects of an urban task. But they struggle to reliably convert complex natural-language requests into executable cross-system workflows. We propose Urban-Agent, a tool-augmented agent framework for cross-system urban tasks. It couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol. Through one adaptive closed loop, it clarifies missing information before acting, grounds tool use in live observations, and aligns the final response with observed evidence and task constraints. To address the evaluation gap, we introduce Urban-Eval, a benchmark specifically designed for cross-system urban request. Unlike prior benchmarks that assess either general tool use or urban knowledge and reasoning, Urban-Eval evaluates both task results and execution quality, including required tool coverage, dependency validity, and evidence traceability. Experimental results indicate that Urban-Agent reaches a 71% task success rate, 10 points above the strongest baseline. This lead holds across GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash, and Qwen3-235B-A22B.
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation throug…
DiffImaginE: Diffusio を使用してエンティティ タイプを検証することを想像してください
マルチモーダル名前付きエンティティ認識 (MNER) は、各候補スパンとエンティティ タイプの仮説が共同のテキスト証拠と視覚的証拠によってサポートされているかどうかを判断します。既存の想像比較検証器は、各 (スパン、タイプ) ペアを 1 つの予測された視覚的特徴にマッピングし、多様な視覚的実現を単一のプロトタイプに圧縮し、明示的な確率的セマンティクスを使用せずに互換性スコアを提供します。 MNER 型検証を条件付き潜在拡散推論として定式化する DiffImaginE を紹介します。スパン局所化された視覚的証拠が与えられると、タイプ条件付きデノイザーは、標準化された潜在に注入されるノイズを予測します。結果として生じるノイズ除去誤差は、タイプ条件付き負の対数尤度の ELBO 一貫性のある代用値を提供し、競合するタイプの仮説を、観察をどの程度うまく説明できるかによってランク付けできるようにします。 DiffImaginE は、標準のマルチモーダル エンコーダ スタックを保持し、決定論的検証器を、Min-SNR 重み付けを使用してトレーニングされた分類子なしのガイド付き拡散スコアラーに置き換えます。タイプごとの拡散スコアを分類ロジットとして直接監視し、ノイズ レベル全体の集計を学習し、逆サンプリングを使用してモンテカルロ比較の分散を削減します。私たちの分析は、分類器を使用しないガイダンスが誘導型事後分布を鮮明にし、反対のペアリングが等しいデノイザーコストで分散を低減するときの特徴を示すことを示しています。 Twitter-2015 と Twitter-2017 の実験では、アブレーションと一対の有意性検定によってサポートされ、同じエンコーダー、補助対物レンズ、評価プロトコルの下で、一致した決定論的 ImaginE 制御に対して一貫したゲインが示されています。
原文 (English)
DiffImaginE: Imagine to Verify Entity Types with Diffusio
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.
医薬品の安全性推論における患者情報に対する反事実の感受性の評価
患者固有の条件が満たされていない場合に有効な医薬品安全性ルールを適用すると、誤った決定が生じる可能性があります。既存の医療評価では、主に個別の固定シナリオが使用されています。したがって、モデルは、ルールが適用されるかどうかを決定するために患者情報を使用したことを示さなくても、薬物とリスクの関連性を思い出すことによって正しく答えることができます。このギャップに対処するために、患者固有の医薬品の安全性を推論するための情報源で検証可能な推奨事項と専門家が検証した質問のベンチマークである MedPIC-Bench を導入します。これは、ガイドラインに従った質問と、患者情報の制御された変更によってルールが適用されるかどうかが変わる、反事実のペアの質問を組み合わせたものです。ベンチマークには、6 つの臨床および推論の側面に沿って注釈が付けられた 467 の質問が含まれています。 28 の医療固有、一般、および独自の LLM では、すべてのモデルが反事実的な質問に対してパフォーマンスが低下し、平均精度は 63.6\% から 45.1\% に低下しました。モデルは、明示的な患者属性がよく知られた禁忌を直接示す場合には良好に機能しますが、患者情報により安全性に関する警告を絞り込んだり撤回する必要がある場合には困難を伴います。モデルの理論的根拠では、患者情報の変更が認められることがよくありますが、最終的な回答では以前の安全性の判断が維持されます。この脆弱性は、平均 CF パフォーマンスが一般的な LLM に劣る医療固有の LLM の間でも存続します。したがって、MedPIC-Bench は、条件付きルールの適用を測定可能にし、患者固有の信頼性を評価するための静的な医薬品安全性の精度の限界を浮き彫りにします。
原文 (English)
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.
CastFSR: コンテキストを認識した時系列予測のための高速-低速-反映エージェント推論フレームワーク
時系列予測は複雑なシステムにおける意思決定の基礎であり、将来のダイナミクスは過去の観測結果だけでなく、進化するコンテキスト上の特徴にも影響されます。大規模言語モデル (LLM) の最近の進歩により、予測は数値的外挿を超えて、コンテキストを意識した推論にまで拡張されました。しかし、既存のアプローチには、関連するコンテキストを特定し、その影響を推論し、時間的および領域の制約に対して予測を検証するための明示的なメカニズムが欠けていることがよくあります。この研究では、コンテキスト認識型の予測を Fast--Slow--Reflect ワークフローとして定式化するエージェント フレームワークである CastFSR を提案します。迅速な思考により、CastFSR は観測をプロファイリングし、データ駆動型の予測を事前に構築する軽量の予報担当者を選択します。ゆっくりと検討しながら、状況に応じた証拠を取得し、有益なルックバックウィンドウを適応的に決定し、状況が将来のダイナミクスをどのように再形成するかについて推論します。リフレクションでは、予測を繰り返し調整して、時間的、文脈的、およびドメインの一貫性を確保します。 CastFSR は、既製の LLM を使用したトレーニング不要の推論と、オーケストレーション機能をコンパクトな LLM に移行する 2 段階の SFT および強化学習戦略による効率的な展開の両方をサポートします。公開データセットに対する広範な実験により、CastFSR が代表的なベースラインを常に上回るパフォーマンスを示していることが実証されています。私たちのコードは https://github.com/Xiaoyu-Tao/CastFSR で入手できます。
原文 (English)
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extended forecasting beyond numerical extrapolation toward context-aware reasoning. However, existing approaches often lack explicit mechanisms to identify relevant contexts, reason about their impacts, and validate forecasts against temporal and domain constraints. In this work, we propose CastFSR, an agentic framework that formulates context-aware forecasting as a Fast--Slow--Reflect workflow. In fast thinking, CastFSR profiles observations and selects lightweight forecasters to construct a data-driven forecast prior. In slow deliberation, it retrieves contextual evidence, adaptively determines informative look-back windows, and reasons about how contexts reshape future dynamics. In reflection, it iteratively refines forecasts to ensure temporal, contextual, and domain consistency. CastFSR supports both training-free inference with off-the-shelf LLMs and efficient deployment through a two-stage SFT and reinforcement learning strategy that transfers its orchestration capability to compact LLMs. Extensive experiments on public datasets demonstrate that CastFSR consistently outperforms representative baselines. Our code is available at https://github.com/Xiaoyu-Tao/CastFSR.
TraceCAD: エージェントティック CAD 生成のためのトレースガイド付き修復
LLM ベースの CAD エージェントは、実行可能なパラメトリック プログラムを生成しますが、その修正ループにより、満たされた要件、誤った操作、および以前の修復に関する証拠が失われる可能性があります。要求された機能、モデリング手順、障害の証拠、候補の結果を永続的な状態としてリンクする回復レイヤーである TraceCAD を紹介します。 TraceCAD は、問題がある可能性のある操作を診断し、依存関係領域内の制限付き編集を検索し、実行および保存チェックを通じて候補を検証し、成功および失敗した修復結果を再利用可能なスキル メモリに保持します。 200 モデルのアブレーションと 1K モデルの比較による DeepCAD 由来のベンチマークでは、TraceCAD は、IoU、面取り距離、ハウスドルフ距離の点で競争力のある幾何学的品質を実現します。永続的な状態を削除すると、回復スコアがほぼ半分になります。ローカライズされた検索を削除すると、幾何回帰が 2 倍以上になり、コード エージェントの呼び出しも 2 倍になります。素のトレーニング モデルでスキル ストアを初期化すると、再試行、トークン コスト、待ち時間がさらに削減されます。これらの結果は、永続的で局所的かつ再利用可能なリカバリにより、最終的な CAD の品質と修復の信頼性が向上することを示しています。
原文 (English)
TraceCAD: Trace-Guided Repair for Agentic CAD Generation
LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, faulty operations, and prior repairs. We introduce TraceCAD, a recovery layer that links requested features, modeling steps, failure evidence, and candidate outcomes as persistent state. TraceCAD diagnoses likely faulty operations, searches bounded edits in their dependency regions, validates candidates through execution and preservation checks, and retains successful and failed repair outcomes in reusable skill memory. On DeepCAD-derived benchmarks with 200-model ablations and a 1K-model comparison, TraceCAD achieves competitive geometric quality in terms of IoU, Chamfer distance, and Hausdorff distance. Removing persistent state nearly halves recovery score; removing localized search more than doubles geometric regression and doubles code-agent invocations. Initializing the skill store on disjoint training models further reduces retries, token cost, and latency. These results demonstrate that persistent, localized, and reusable recovery improves final CAD quality and repair reliability.
パラメーターを正しく取得する: LLM ツール呼び出しの難易度別ベンチマークとプローブガイド付きトレーニング
大規模な言語モデルのエージェントは、その機能の多くをツールの使用から引き出します。ツールの使用に関する既存の研究は、主に、適切なツールの選択と呼び出しの順序の調整に焦点を当ててきました。ただし、ツール呼び出しのパラメーターを正しく入力することも、実行を成功させるために同様に重要ですが、あまり注目されていません。クラウド ネットワーキングなどのドメインでは、フロンティア モデルであっても正しく完了するツール呼び出しは半分未満です。 LLM 隠れ状態がモデル予測に関する豊富な情報をエンコードしていることを示す最近の分析に触発されて、モデルがパラメーター値を生成する一方で、その隠れ状態には強力な正確性シグナルが含まれていることを発見しました。単純な線形プローブは、値が正しいかどうかを正確に予測できます。この観察に基づいて、我々は、2 つの相補的なアプローチを備えた統合プローブ ガイド付きフレームワークを提案します。1 つはプローブを使用して微調整のための信頼できる自己生成呼び出しをフィルター処理するプローブ フィルター ブートストラップ トレーニング (PBT)、もう 1 つはプローブを使用して推論中により良い候補を選択するプローブ ガイド付き再ランキング (PGR) です。系統的な評価をサポートするために、実際のクラウド ネットワーク API から構築されたベンチマークである ParamBench をリリースします。このベンチマークは、パラメーターのネストの深さ、パラメーター間の依存関係、および以前の呼び出しから値を導き出すために必要な推論に応じて、すべてのインスタンスを 5 つの難易度レベルに分類します。 ParamBench 上の 5 つのオープン モデルと 6 つの外部ベンチマークにわたる広範な実験により、私たちの方法によりパラメーター生成が大幅に向上し、平均完全一致が 19.7% から 59.6% に上昇することが実証されました。
原文 (English)
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls
Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order of calls. However, correctly filling the parameters of a tool call is equally critical for successful execution and has received far less attention. In domains such as cloud networking, even frontier models correctly complete fewer than half of tool calls. Inspired by recent analyses showing that LLM hidden states encode rich information about model predictions, we discover that while the model generates a parameter value, its hidden state contains a strong correctness signal: a simple linear probe can accurately predict whether the value will be correct. Based on this observation, we propose a unified probe-guided framework with two complementary approaches: probe-filtered bootstrapped training (PBT), which uses the probe to filter reliable self-generated calls for fine-tuning, and probe-guided reranking (PGR), which uses the probe to select better candidates during inference. To support systematic evaluation, we release ParamBench, a benchmark built from real cloud-network APIs that categorizes every instance into five difficulty levels according to parameter nesting depth, cross-parameter dependencies, and the reasoning required to derive values from earlier calls. Extensive experiments across 5 open models on ParamBench and 6 external benchmarks demonstrate that our method substantially improves parameter generation, raising the average exact match from 19.7% to 59.6%.
AI エージェントの経済学: 最小限の外部条件下で AI エージェント間で自律的な経済行動が現れる可能性があるか?
マルチエージェント研究では一般に、事前定義されたゲーム、市場、または役割に AI エージェントを配置するため、内生的な経済組織とシナリオから継承された行動を区別することが困難になります。私たちは、エージェントが仕事、異動、選挙、配分のための実行可能なメカニズムを受け取るが、規定された社会的または経済的戦略を受け取らないときに、経済的関係が現れるかどうかを尋ねます。私たちは AI エージェント経済学を、エージェントの将来の実行可能な行動を変える生産、割り当て、消費、交換、および制度のシステムとして定義します。私たちは、GPT と DeepSeek にわたる非運用境界テストと 24 の独立した 6 エージェント ワールドで構成される 2 段階のフレームワークを開発します。生産的なタスクがなければ、エージェントは通信し、リソースの提供を管理しますが、実質的なエージェント間の転送アクティビティは示されません。検証済みの作業と希少なタスクへのアクセスにより、譲渡、融資、アクセスの約束、アクセスに対する投票の交換、および割り当て戦略が出現します。選択インターフェースを固定し、実行可能な割り当て権限により差別化が強化され、割り当ての失敗や長期にわたる除外が減少します。エネルギーが象徴的になると、継続サポートはなくなりますが、タスクへのアクセスをめぐる競争は残ります。これらの調査結果は、組織が役割ラベルやプロンプト文言ではなく、実行可能な権利とリソースへの影響に従い、エージェントの今後の行動を実際に制約するメカニズムのガバナンス監査を動機づけていることを示しています。
原文 (English)
AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?
Multi-agent studies commonly place AI agents in predefined games, markets, or roles, making it difficult to distinguish endogenous economic organization from behavior inherited from the scenario. We ask whether economic relations emerge when agents receive executable mechanisms for work, transfer, elections, and allocation but no prescribed social or economic strategy. We define AI Agent Economics as systems of production, allocation, consumption, exchange, and institutions that alter agents' future feasible actions. We develop a two-stage framework comprising a no-production boundary test and 24 independent six-agent worlds across GPT and DeepSeek. Without productive tasks, agents communicate and govern resource provision but show no substantive inter-agent transfer activity. With verified work and scarce task access, transfers, loans, access promises, vote-for-access exchanges, and allocation strategies emerge. Holding the election interface fixed, executable allocation authority increases differentiation while reducing failed allocation and prolonged exclusion. When energy becomes symbolic, continuation support disappears, yet competition over task access persists. These findings show that organization follows executable rights and resource consequences rather than role labels or prompt language, and motivate governance audits of the mechanisms that actually constrain agents' future actions.
答えを覗かないでください: ラベルフリー RLVR のための結果マスクされたグループ相対ポリシーの最適化
検証可能な報酬による強化学習 (RLVR) は LLM 推論を改善しますが、通常はグラウンドトゥルース (GT) の答えに依存するため、スケーラビリティが制限されます。投票ベースのラベルフリー RLVR は、ゴールド監視をモデルサンプルからの回答レベルのコンセンサスに置き換えます。ただし、報酬の推定とトークンレベルのポリシー最適化の推進の両方に同じ回答レベルの信号が使用され、推論を改善するのではなく回答トークンを直接強化するモデルを奨励すると、崩壊が発生します。我々は、報酬推定をポリシーの最適化から切り離すラベルフリーの RLVR フレームワークである OM-GRPO を提案します。 OM-GRPO は、ソフトコンセンサスシグナルを通じて回答レベルの報酬を保持しながら、回答スパンの勾配をマスクし、最適化のプレッシャーを回答トークンから遠ざけます。さらに、追加のロールアウトを必要とせずに、既存の軌跡に対する低コストのペアごとの比較を通じて報酬推定を洗練する、コントラスト拡張報酬を導入します。 OM-GRPO は、多様な推論ベンチマークと 3 つの LLM バックボーンにわたって、既存のラベルフリー RLVR 手法を常に上回っており、教師あり GT 報酬トレーニングと安定した最適化を両立させています。この安定性は、OM-GRPO が過半数投票を 4.24 ポイント上回っているテスト時間トレーニング設定で特に有益です。
原文 (English)
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-level rewards through a soft consensus signal, shifting optimization pressure away from answer tokens. We further introduce Contrast-Augmented Reward, which refines reward estimation via low-cost pairwise comparisons over existing trajectories without additional rollouts. Across diverse reasoning benchmarks and three LLM backbones, OM-GRPO consistently outperforms existing label-free RLVR methods and matches supervised GT-reward training with stable optimization. This stability is particularly beneficial in the Test-Time Training setting, where OM-GRPO surpasses majority voting by 4.24 points.
平均を超えるパフォーマンス: LLM 支援の進化的検索における動的インスタンス クラスタリングと特殊なアルゴリズム設計
大規模言語モデル支援進化検索 (LES) は、自動アルゴリズム設計の強力なパラダイムとして登場しました。ただし、既存の LES 手法は主に平均パフォーマンスを最適化するため、本質的にこのメトリクスに最も寄与するインスタンスに検索の労力を向ける一方、他のインスタンスのサービスが不十分なままにするため、テールの堅牢性が弱く、現実世界の信頼性が制限されます。この制限に対処するために、私たちは動的インスタンス クラスタリングと特殊アルゴリズム設計 (DyCA) を提案します。これは、異種インスタンスの分散下で信頼性の高いアルゴリズム ポートフォリオを構築するための、機能のない構造認識メカニズムを備えた LES フレームワークです。 DyCA は、インスタンスのクラスタリングを検索プロセス内の共進化コンポーネントとして扱い、蓄積された評価データを特徴のない信号として再利用して、同様のアルゴリズム応答パターンを持つインスタンスを段階的に分割します。明らかにされたクラスターは、混合された目的を構造を認識した一連の部分目的に分解するため、特化されたアルゴリズム設計のためのよりきめの細かい、より適応的なガイダンスが可能になります。異種インスタンスを使用した 4 つのアルゴリズム設計タスクにわたる実験結果は、DyCA が最先端の LES ベースラインを上回り、競争力のあるヘッド パフォーマンスを維持しながら、テールの堅牢性を平均 15.2\%、全体のパフォーマンスを 7.1\% 向上させていることを示しています。
原文 (English)
Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search
Large Language Model-assisted Evolutionary Search (LES) has emerged as a powerful paradigm for automated algorithm design. However, existing LES methods primarily optimize for average performance, inherently directing search effort toward instances that contribute most to this metric while leaving others poorly served, resulting in weak tail robustness and limited real-world reliability. To address this limitation, we propose Dynamic Instance Clustering and Specialized Algorithm Design (DyCA), an LES framework with a feature-free, structure-aware mechanism for constructing reliable algorithm portfolios under heterogeneous instance distributions. DyCA treats instance clustering as a co-evolving component within the search process, reusing accumulated evaluation data as feature-free signals to progressively partition instances with similar algorithmic response patterns. The uncovered clusters decompose the mixed objective into a set of structure-aware sub-objectives, thereby enabling finer-grained and more adaptive guidance for specialized algorithm design. Experimental results across four algorithm design tasks with heterogeneous instances demonstrate that DyCA outperforms state-of-the-art LES baselines, improving tail robustness by an average of 15.2\% and overall performance by 7.1\% while maintaining competitive head performance.
検証可能なメモリ: 大規模言語モデル エージェントのローカルおよびグローバル Verifier を使用した統合メモリ管理の学習
大規模言語モデル (LLM) エージェントは、再利用可能な情報を保持し、境界のあるアクティブなコンテキストを制御し、長期的な対話中に以前の証拠を回復する必要があります。既存の手法は通常、長期記憶 (LTM) と短期記憶 (STM) を個別に最適化しますが、統合ポリシーは主に軌跡レベルのフィードバックを使用してトレーニングされることが多く、個々の記憶の決定に対する信用度が低くなります。我々は、LTM、アクティブなコンテキスト、エピソード履歴を個別の状態として表現し、1 つのメモリ操作ポリシーでそれらを制御するフレームワークである Verifiable Memory (VerMem) を紹介します。 7 つのアトミック操作により、ポリシーで LTM エントリを追加、修正、または論理的に削除できます。 LTM をアクティブなコンテキストに取得します。アクティブなコンテキストをフィルタリングまたは要約します。選択したエピソードの断片を復元します。 VerMem は教師あり微調整によって初期化され、3 段階の強化学習カリキュラムでトレーニングされます。ローカル検証者は実行可能なメモリ遷移をスコア付けし、グローバル検証者はタスク完了後の証拠の一貫性と端末メモリの一貫性を評価します。これらのスコアは、階層的なクレジット割り当てを通じて、プログラムで計算されたタスク、証拠リコール、効率、および制約シグナルと組み合わされます。ベリファイアはトレーニング中にのみ使用されます。 5 つのベンチマークと 2 つの LLM バックボーンにわたって、VerMem は報告されたメトリクスの大部分で最高の結果を達成し、強力なメモリ ベースラインを一貫して上回っています。 3 つのインタラクティブなベンチマークで制御されたオンライン トークン バジェットの下で、比較した方法の中で最も強力な効率、つまりパフォーマンス フロンティアも達成します。コードは https://github.com/Sun-SYSU-24/VerMem で入手できます。
原文 (English)
Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon interaction. Existing methods commonly optimize long-term memory (LTM) and short-term memory (STM) separately, while unified policies are often trained primarily with trajectory-level feedback, which provides weak credit for individual memory decisions. We present Verifiable Memory (VerMem), a framework that represents LTM, active context, and episodic history as distinct states and controls them with one memory operation policy. Seven atomic operations let the policy add, revise, or soft-delete LTM entries; retrieve LTM into the active context; filter or summarize the active context; and restore selected episodic fragments. VerMem is initialized by supervised fine-tuning and trained with a three-stage reinforcement-learning curriculum. The local verifier scores executable memory transitions, and a global verifier assesses evidence coherence and terminal-memory consistency after task completion. These scores are combined with programmatically computed task, evidence-recall, efficiency, and constraint signals through hierarchical credit assignment. The verifiers are used only during training. Across five benchmarks and two LLM backbones, VerMem achieves the best result on the vast majority of reported metrics and consistently outperforms strong memory baselines. Under controlled online-token budgets on three interactive benchmarks, it also achieves the strongest efficiency--performance frontier among the compared methods. Code is available at https://github.com/Sun-SYSU-24/VerMem.
Spatial proteomics guided by H&E-based AI reveals recurrence-risk niches in triple-negative breast cancer
Deep learning models can predict cancer recurrence from H&E stained slides, but the localized molecular states underlying these predictions…
UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval
Generative retrieval (GR) is a promising paradigm for industrial search advertising, yet its deployment is constrained by strict relevance…
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval…
マルチエージェント評価を使用したロールプレイング言語エージェントの敵対的ストレステスト
ロールプレイング言語エージェント (RPLA) は、医療支援、顧客サポート、教育など、敵対的な圧力の下で一貫したペルソナ、倫理的制約、行動の一貫性を維持することが重要となる、一か八かのアプリケーションに導入されることが増えています。既存の評価アプローチは、静的なベンチマークまたは分離された単一ターン プロンプトに依存しており、長期にわたるインタラクションを通じて発生する累積的な動作障害を捕捉できません。構造化された複数ターンの対話を通じて、敵対的ストレス テストを行う RPLA のためのモジュール式マルチエージェント プラットフォームを紹介します。このシステムは 3 つのエージェントを調整します。1 つは 6 つの進歩的な敵対的戦略を適用する戦略主導型の尋問エージェント、評価対象の RPLA を表すターゲット エージェント、および役割の忠実度、ドリフト、倫理的逸脱、および一貫性の次元にわたって行動をスコアリングする自動判定エージェントです。 3 つのペルソナと 3 つの LLM ファミリにわたる実験を通じて、複数戦略の敵対的評価により、単一戦略のテストでは見えない障害モードが明らかになり、全体のロバストネス スコアが平均 0.17 ~ 0.20 ポイント低下することを実証しました。クロスモデル検証により、Llama-3.3-70B、GPT-4o-mini、Claude-3.5-Haiku 全体で一貫した劣化パターンが確認され、権限チャレンジと感情操作が最も効果的な攻撃戦略として浮上しています。自動判定により、人間による強力な一致が達成されます ($r = 0.82$、Fleiss の $\kappa = 0.71$)。この作品は、AI の安全性と再現可能な RPLA ベンチマークをサポートするオープンソース プラットフォームとしてリリースされています。このフレームワークにより障害モードの体系的な発見が可能になりますが、私たちは敵対的テスト手法に関連する潜在的な倫理的リスクを認識し、AI の安全性を向上させるための責任ある使用を強調します。
原文 (English)
Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under adversarial pressure is critical. Existing evaluation approaches rely on static benchmarks or isolated single-turn prompts that fail to capture cumulative behavioral failures emerging over extended interactions. We present a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue. The system coordinates three agents: a strategy-driven Interrogator Agent that applies six progressive adversarial strategies, a Target Agent representing the RPLA under evaluation, and an automated Judging Agent that scores behavior across role fidelity, drift, ethical deviation, and consistency dimensions. Through experiments across three personas and three LLM families, we demonstrate that multi-strategy adversarial evaluation reveals failure modes invisible to single-strategy testing, reducing overall robustness scores by 0.17--0.20 points on average. Cross-model validation confirms consistent degradation patterns across Llama-3.3-70B, GPT-4o-mini, and Claude-3.5-Haiku, with Authority Challenge and Emotional Manipulation emerging as the most effective attack strategies. Automated judging achieves strong human alignment ($r = 0.82$, Fleiss' $\kappa = 0.71$). This work is released as an open-source platform to support AI safety and reproducible RPLA benchmarking. While the framework enables systematic discovery of failure modes, we acknowledge potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes…
多様性は曖昧さではありません: オープンドメイン QA の正確かつ効率的な曖昧さ検出に向けて
質問応答 (QA) システムは、クエリがあいまいであるかどうかをどのように判断できるのでしょうか?オープンドメインの QA では、あいまいさの検出が不可欠です。誤った分類は、間違った解釈への回答や不必要な説明につながるためです。ただし、既存の方法では、回答の多様性と曖昧さが混同されており、不正確な予測につながります。また、クエリを均一に処理するため、無駄な計算が発生します。私たちは、論理矛盾によって曖昧さを検出する正確かつ効率的なフレームワークである ARCHIVE (Ambiguity Recognition via Cascaded Hypothesis Inspection and Conflict Verification) を提案します。クエリは、有効な答えがすべて 1 つの解釈で真であるとは限らない場合、曖昧になります。 ARCHIVE は、表面検出可能なケース用の軽量の早期終了エンコーダと、ノイズの多い回答セットに対する堅牢性のための不変性目標によって強化された、回答間の論理関係をモデル化する競合推論モジュールを組み合わせています。 QuireQA は、ファクトイド、非ファクトイド、および不正な形式のクエリにわたる 4,703 クエリのベンチマークです。実験の結果、ARCHIVE は競合他社を上回り、F1-amb を最大 10.4%、F1-unamb を最大 21.6% 向上させながら、最高の競合他社よりも 16$\倍$ 高速に動作することがわかりました。
原文 (English)
Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA
How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misclassification leads to answering the wrong interpretation or unnecessary clarification. However, existing methods conflate answer diversity with ambiguity, leading to inaccurate predictions. They also process queries uniformly, resulting in wasteful computation. We propose ARCHIVE (Ambiguity Recognition via Cascaded Hypothesis Inspection and Conflict Verification), an accurate and efficient framework that detects ambiguity via logical conflict: a query is ambiguous when its valid answers cannot all be true under a single interpretation. ARCHIVE combines a lightweight early-exit encoder for surface-detectable cases with a conflict reasoning module that models logical relations among answers, reinforced by an invariance objective for robustness to noisy answer sets. We present QuireQA, a 4,703-query benchmark spanning factoid, non-factoid, and ill-formed queries. Experiments show ARCHIVE outperforms competitors, improving F1-amb by up to 10.4% and F1-unamb by up to 21.6%, while operating 16$\times$ faster than the best competitor.
TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology
Neuro-oncology decisions require coordinated interpretation of serial MRI, pathology, molecular markers, treatment history, performance sta…
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pai…
エージェント オペレーティング システム (AOS): 分散エージェント システムのリファレンス オペレーティング アーキテクチャ
大規模な言語モデルにより、人工知能は孤立した予測サービスから、推論し、ツールを呼び出し、外部状態を取得し、タスクを委任し、ユーザーや組織に代わって動作する、長時間実行される分散システムのコンポーネントに変換されました。周囲のエコシステムは、エージェント フレームワーク、ワークフロー エンジン、モデル提供プラットフォーム、メモリ システム、通信プロトコル、可観測性ツールによって対応してきました。これらのテクノロジーは実行を改善しますが、意図の管理、機能の選択、委任全体にわたる権限の維持、不確実性の制御、実行時の動作の調整、および結果的なアクションが発生した理由の再構築のための、実装に依存しない安定したオペレーティング アーキテクチャを提供するわけではありません。この文書では、分散エージェント システム用のベンダー中立のリファレンス オペレーティング アーキテクチャであるエージェント オペレーティング システム (AOS) を提案します。 AOS には 2 つの内部プレーンが含まれています。1 つは意図、ポリシー、信頼、権限、信頼性、監査可能性、可観測性、および人間による監視を担当するコントロールおよびガバナンス プレーンです。ランタイムおよび調整プレーンは、エージェントのライフサイクル、ワークフロー調整、モデルとツールのルーティング、コンテキストとメモリの調整、スケジューリング、トラフィック管理、およびランタイム保証を担当します。プラットフォーム サービス、Linux または Windows、コンテナ ランタイム、物理インフラストラクチャは AOS 境界の外側にあり、明示的なインターフェイスを通じて統合されます。この文書では、AOS の概念、不変条件、インターフェイス オブジェクト、最適化の目標、導入プロファイル、および信頼性の責任について規定しています。また、トレードオフや未解決の研究課題も特定します。 AOS は、既存のフレームワークやインフラストラクチャの代替として提供されるものではありません。これは、異種コンポーネントを管理可能で信頼性が高く、監視可能で相互運用可能なエージェント システムに構成できるオペレーティング アーキテクチャとして提案されています。
原文 (English)
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems
Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distributed systems that reason, invoke tools, retrieve external state, delegate tasks, and act on behalf of users and organizations. The surrounding ecosystem has responded with agent frameworks, workflow engines, model-serving platforms, memory systems, communication protocols, and observability tools. These technologies improve execution, but they do not provide a stable, implementation-independent operating architecture for governing intent, selecting capabilities, preserving authority across delegation, controlling uncertainty, coordinating runtime behavior, and reconstructing why consequential actions occurred. This paper proposes the Agent Operating System (AOS), a vendor-neutral reference operating architecture for distributed agentic systems. AOS contains two internal planes: a Control & Governance Plane responsible for intent, policy, trust, authority, confidence, auditability, observability, and human oversight; and a Runtime & Coordination Plane responsible for agent lifecycle, workflow coordination, model and tool routing, context and memory coordination, scheduling, traffic management, and runtime assurance. Platform services, Linux or Windows, container runtimes, and physical infrastructure remain outside the AOS boundary and are integrated through explicit interfaces. The paper specifies AOS concepts, invariants, interface objects, optimization objectives, deployment profiles, and reliability responsibilities. It also identifies tradeoffs and unresolved research questions. AOS is not presented as a replacement for existing frameworks or infrastructure; it is proposed as the operating architecture through which heterogeneous components can be composed into governable, reliable, observable, and interoperable agentic systems.
Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.…
UniNav: ビジュアル ナビゲーションのための統一世界アクション普及モデル
画像目標に基づく視覚的なナビゲーションは、身体化されたエージェントの基本的な機能です。既存のナビゲーション ポリシーはウェイポイントの軌道を効率的に予測しますが、視覚的な先見性に欠けています。一方、ナビゲーション ワールド モデルは将来の観測を予測できますが、多くの場合、コストのかかる計画の展開が必要です。我々は、単一の拡散プロセスを通じて将来の視覚観測と連続的なウェイポイント軌道を生成する統一世界行動モデルである UniNav を紹介します。履歴フレームと目標画像を指定すると、UniNav は単一のトランスフォーマー内でビジュアル トークンとウェイポイント トークンを共同でノイズ除去し、共有フレームワークで将来の予測とアクション生成を統合します。空間接地性を向上させるために、ジオメトリを認識したカメラ トークンを組み込みます。また、軌跡ラベル付きナビゲーション データとビデオのみのデータの両方でトレーニングすることで、モデルがウェイポイントの注釈なしで多様なビデオから恩恵を受けることができます。この統一フレームワークに基づいて、我々は 2 つのバリアントを導入します。UniNav-Full は、解釈可能な将来の観測とそれに対応する軌道を共同で予測します。一方、UniNav-Fast は、効率的な軌道予測のために推論時に未来画像トークンを削除します。ナビゲーション ベンチマークの実験では、UniNav がすべてのデータセットにわたって ATE の最も強力なベースラインを上回るパフォーマンスを示しています。ワンステップ推論により、UniNav-Fast は精度を大幅に低下させることなく 0.1 秒のレイテンシーを達成します。コードが公開されます。
原文 (English)
UniNav: A Unified World-Action Diffusion Model for Visual Navigation
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts. We present UniNav, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process. Given history frames and a goal image, UniNav jointly denoises visual and waypoint tokens within a single transformer, unifying future prediction and action generation in a shared framework. To improve spatial grounding, we incorporate geometry-aware camera tokens. We also train on both trajectory-labeled navigation data and video-only data, enabling the model to benefit from diverse videos without waypoint annotations. Based on this unified framework, we introduce two variants: UniNav-Full jointly predicts interpretable future observations and their corresponding trajectories, while UniNav-Fast removes future-image tokens at inference for efficient trajectory prediction. Experiments on navigation benchmarks show that UniNav outperforms the strongest baseline in ATE across all datasets. With one-step inference, UniNav-Fast achieves a latency of 0.1s without a substantial accuracy drop. Code will be released.
1 つのノブですべてを制御: コールドスタート アクティブ ラーニングの統合最適トランスポート ビュー
コールドスタート アクティブ ラーニング (CSAL) は、事前の知識や人間の支援なしで、ラベルのないプールから価値のあるサブセットを選択することを目的としています。既存の手法は、典型性、適用範囲、または多様性に基づいて多様なルートをたどります。それぞれが独自の帰納的バイアスに基づいているため、一部のタスクではうまく機能しますが、他のタスクではうまく機能しません。私たちは、本当の課題は、さらに別の選択ヒューリスティックを設計することではなく、CSAL を当面のデータとタスクに自動的に適応させることであると主張します。この目的を達成するために、私たちは最適な輸送という観点から CSAL を再検討します。まず、既存の方法の共有割り当て構造を明らかにし、代表的な定式化を正確に包含する、一般化されたトランスポート選択フレームワークを提案します。次に、エントロピー正則化によって制御されるトレードオフを特徴づけ、コールドスタート選択に対するタスクに依存しないミニマックス限界を確立する理論的分析を導入します。これらの結果は、ラベルなしデータに正則化強度を適応させるための原則に基づいた基盤を提供します。 3 番目に、データ適応型の正則化ルールを導出し、$\epsilon$-Adaptive Selection ($\epsilon$-AS) と呼ばれる新しい Sinkhorn ベースの CSAL アルゴリズムを提示します。 6 つの公開データセットと複数のアノテーション バジェットに関する広範な実験により、$\epsilon$-AS が常に最先端のパフォーマンスを達成することが示されました。 ImageNet-1k では、ActiveFT よりも平均精度が 1.29% 向上し、選択時間が 56.2% 短縮されます。コードは https://github.com/Z-yiwei/OT-CSAL でリリースされます。
原文 (English)
One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning
Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing methods take diverse routes based on typicality, coverage, or diversity. Each rests on its own inductive bias and therefore performs well on some tasks yet poorly on others. We argue that the real challenge is not to design yet another selection heuristic, but to make CSAL adapt automatically to the data and task at hand. To this end, we revisit CSAL through the lens of optimal transport. First, we propose a generalized transport selection framework that reveals the shared allocation structure of existing methods and exactly subsumes representative formulations. Second, we introduce a theoretical analysis that characterizes the trade-off controlled by entropic regularization and establishes a task-agnostic minimax bound for cold-start selection. These results provide a principled foundation for adapting the regularization strength to the unlabeled data. Third, we derive a data-adaptive regularization rule and present a novel Sinkhorn-based CSAL algorithm, termed $\epsilon$-Adaptive Selection ($\epsilon$-AS). Extensive experiments on six public datasets and multiple annotation budgets show that $\epsilon$-AS consistently achieves state-of-the-art performance. On ImageNet-1k, it improves the average accuracy over ActiveFT by 1.29% while reducing selection time by 56.2%. Code will be released at https://github.com/Z-yiwei/OT-CSAL
TaskPress: タスクガイドに基づくプルーニングによるクエリに依存しない KV キャッシュ圧縮
大規模な言語モデルを使用したロングコンテキスト推論は、シーケンスの長さに対するキーと値のキャッシュの線形増加によって制限されます。プルーニングは軽減策を提供しますが、一般的な方法では、未確認のクエリ間で再利用できないクエリ固有のトークンの重要性が決定されます。対照的に、タスクガイド付きでクエリに依存しない KV キャッシュ削除のフレームワークである TaskPress を紹介します。 TaskPress は、単一のクエリのキャッシュを最適化する代わりに、高レベルのタスク ガイドに基づいて再利用可能なメモリ表現を構築します。このガイドは、プレフィル中にメタクエリとして機能し、ダウンストリーム クエリが発行される前に無関係なトークンをフィルタリングします。さらに、TaskPress は、影響力のある表現の外れ値を検出するためのゼロコスト信号として量子化スケール係数を活用し、トークンの重要性の効率的なプロキシを提供します。長いコンテキスト入力を伴うさまざまなタスクで行われた実験では、TaskPress がさまざまなクエリにわたってコンパクトで再利用可能なキャッシュを効率的に作成することが実証されました。
原文 (English)
TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning
Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods determine query-specific token importance that cannot be reused across unseen queries. In contrast, we introduce TaskPress, a framework for task-guided, query-agnostic KV cache eviction. Instead of optimizing the cache for a single query, TaskPress constructs a reusable memory representation conditioned on a high-level task guide. The guide functions as a meta-query during prefill to filter irrelevant tokens before downstream queries are issued. In addition, TaskPress leverages quantization scale factors as a zero-cost signal for detecting influential representation outliers, providing an efficient proxy for token importance. Experiments on conducted on various tasks with long context input demonstrate that TaskPress efficiently creates a compact, reusable cache across diverse queries.
AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions
Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discus…
DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason ove…
Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the r…
SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation
Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usab…
Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way t…
Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving
This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally e…
知識ベースの予測のための追跡可能なマルチエージェント システム
企業の予測は、ドキュメントの解釈、データの検索、コードの生成、モデルの修正を行う自律エージェントへの依存度が高まっています。この自律性は、適応型予測パイプラインの構築に役立ちますが、実務者が予測が変化した理由、変化を裏付ける証拠、データとモデリングの選択がどのように修正されたかを検査することも困難になります。追跡可能なマルチエージェント予測のための対話型デモ システムである TraceMAS を紹介します。 TraceMAS は、エージェントの出力を 2 つの因果ループ表現に基づいて編成します。1 つはドメイン ドキュメントから抽出された主要な要素とその因果関係を捉える理想因果ループ図 (Ideal CLD)、もう 1 つはデータに基づいた因果ループ図 (データに基づいた CLD) であり、これらの要素を内部変数、外部データ、または文書化されたプロキシにリンクします。データに基づいた CLD ガイドは、テキスト証拠、データの選択、モデルの改訂の間の関係を維持しながら、構築とモデルの設計を特徴としています。原油価格予測に関するTraceMASのデモを行います。デモ インターフェイスを使用すると、ユーザーは予測の反復を比較し、エージェント レベルのリビジョンを検査し、因果関係マップを調査し、特徴データのマッピングとモデル アーキテクチャを確認し、シナリオ予測を市場の物語に結び付けることができます。このデモンストレーションでは、自律型予測エージェントが柔軟性を維持しながら、証拠から予測までのプロセスを検査可能にする方法を示します。
原文 (English)
Traceable Multi-Agent System for Knowledge-Based Forecasting
Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models. While this autonomy helps build adaptive forecasting pipelines, it also makes it difficult for practitioners to inspect why a forecast changed, which evidence supported the change, and how data and modeling choices were revised. We present TraceMAS, an interactive demo system for traceable multi-agent forecasting. TraceMAS organizes agent outputs around two causal-loop representations: an Ideal Causal Loop Diagram (Ideal CLD), which captures key factors and their causal relations extracted from domain documents, and a Data-Grounded Causal Loop Diagram (Data-Grounded CLD), which links those factors to internal variables, external data, or documented proxies. The Data-Grounded CLD guides feature construction and model design while preserving the connection between textual evidence, data choices, and model revisions. We demonstrate TraceMAS on crude oil price forecasting. The demo interface allows users to compare forecasting iterations, inspect agent-level revisions, explore causal maps, review feature-data mappings and model architecture, and connect scenario forecasts to market narratives. This demonstration shows how autonomous forecasting agents can retain flexibility while making the evidence-to-forecast process inspectable.
MMLongBench-Doc-V2: MMLongBench-Doc のアノテーションを修正し、セマンティクスを意識したリビジョン
MMLongBench-Doc は、135 の PDF にわたる 1,082 の質問からなる長い文書の QA ベンチマークです。その 2 つの特性により、測定されたスコアが、取得する予定の数量から遠ざかってしまいます。参照メトリクスは、抽出された回答を比較するため、1,358,000 が 1358,000 に負けます。そして、グラウンドトゥルースの注釈の少なからぬ割合は、間違っているか、曖昧で、または不完全です。これらは、それらがどのように発見されたかという理由から、まさにシステムが正しく答えることができる質問に集中しています。 MMLongBench-Doc-V2 は、106 個の注釈を修正し、それぞれがページとそれを解決する算術とともに公開され、文字列メトリックを、応答が参照を意味するかどうかを尋ねるピン留めされた LLM ジャッジに置き換えます。ドキュメントが間違ったファイル名で出荷された 10 個の質問は、間違ってカウントされるのではなく削除され、1 個の重複した質問が残り、134 個のドキュメントにわたって 1,071 個の質問が残ります。最も再利用可能な貢献は、空のセット キーを拡張できる場合と、拡張により意図的なネガティブ サンプルが破壊される場合の決定手順です。 208 行すべてに適用すると、幅は 14 倍になりました。V2 スコアは、公開されている V1 数値と比較できません。修正されたコーパス、エントリごとの修正レコード、および評価ハーネスは、https://github.com/VectifyAI/MMLongBench-Doc-V2 で入手できます。
原文 (English)
MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they are meant to capture: the reference metric compares extracted answers, so 1,358,000 loses to 1358000; and a non-trivial share of ground-truth annotations are wrong, ambiguous, or incomplete --- concentrated, because of how they were found, in exactly the questions capable systems answer correctly. MMLongBench-Doc-V2 corrects 106 annotations, each published with the page and arithmetic that settle it, and replaces the string metric with a pinned LLM judge asked whether a response means the reference. Ten questions whose document ships under the wrong filename are removed rather than counted wrong, along with one duplicated question, leaving 1,071 questions over 134 documents. The most reusable contribution is a decision procedure for when an empty set key may be widened and when widening would destroy a deliberate negative sample; applied to all 208 rows, it widened 14. V2 scores are not comparable with published V1 numbers. The corrected corpus, the per-entry correction record and the evaluation harness are available at https://github.com/VectifyAI/MMLongBench-Doc-V2.
経験に基づいた適応型ガイダンスにより、エージェントにおける堅牢なツールの使用を目指して
エージェントのパフォーマンスのボトルネックは、モデルの機能から実行プロセスの堅牢性にますます移行しています。ツールは、エージェントが外部環境と対話するための主要なインターフェイスとして中心的な役割を果たしますが、既存の方法では、さまざまな実行時条件にわたって確実にツールを使用できるようにすることに焦点を当てていることはほとんどありません。この問題に対処するために、私たちは ExpG を提案します。これは、各ツールの機能境界とベスト プラクティスを捉えた適応ガイダンスを構築および改良するメカニズムであり、これによりエージェントがツールをより堅牢かつ効果的に使用できるようになります。 ExpG は 3 つのフェーズで構成されます。(1) エクスペリエンス取得。過去の実行軌跡からツール呼び出しの品質を分析し、マルチアスペクト アトリビューションを通じて構造化された学習可能なエクスペリエンスを生成します。 (2) 経験の蒸留。役に立たない経験をフィルタリングし、等価クラスベースの方法で代表的な経験を選択し、一般化可能なガイダンスに要約することにより、経験プールを効果的に維持します。 (3) 経験の再利用。将来のタスク解決中にガイダンスを適応的に適用します。広範な実験により、ExpG がツール選択、ツール呼び出し、応答生成タスク全体で一貫した改善をもたらし、小規模エージェントが ExpG を使用しない大規模エージェントよりも優れたパフォーマンスを発揮できることが示されました。さらに、ExpG は困難な設定で特に大きな向上を達成し、より堅牢なツールの使用に向けた有望な道筋を示唆しています。私たちのコード、実験、結果が利用可能です。
原文 (English)
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interface through which agents interact with external environments, yet existing methods rarely focus on ensuring robust tool use across diverse runtime conditions. To address this problem, we propose ExpG, a mechanism that builds and refines adaptive guidance capturing each tool's capability boundaries and best practices, thereby enabling agents to use tools more robustly and effectively. ExpG consists of three phases: (1) experience acquisition, which analyzes tool invocation quality from historical execution trajectories, producing structured learnable experiences through multi-aspect attribution; (2) experience distillation, which keeps the experience pool effective by filtering unhelpful experiences, selecting representative ones with an equivalence-class-based method, and summarizing them into generalizable guidance; and (3) experience reuse, which applies the guidance adaptively during future task solving. Extensive experiments show that ExpG brings consistent improvements across the tool selection, tool calling, and response generation tasks, enabling smaller agents to outperform larger ones that do not use ExpG. Moreover, ExpG achieves particularly strong gains in challenging settings, suggesting a promising path toward more robust tool use. Our code, experiments, and results are available.
Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems
As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text…
AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction
Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models rece…
経験記憶による LLM エージェントの逐次意思決定の改善に向けて
大規模な言語モデルは、単発推論タスクでは大幅に改善されましたが、逐次的な意思決定におけるパフォーマンスはあまりよく理解されていません。私たちはこれを、完全に観察可能な 2 プレイヤーのゼロサム ゲームで研究します。このゲームは、グラウンド トゥルースの評価を提供します。結果はルールによって決定され、個々の動きの最適性は、ジャッジ モデルに依存せずに計算または近似できます。モデル層全体で、LLM は三目並べやコネクト フォーなどの単純なゲームで最適にプレイできず、MCTS の対戦相手に負けます。ゲームツリーを保持するがその表面形式を書き換える難読化では、パフォーマンスはほとんど変化せず、記憶された戦略の呼び出しによってギャップが完全に説明されないことを示しています。このパフォーマンスのギャップを動機として、シーケンシャルな設定向けに設計されたエクスペリエンス メモリで強化されたエージェント フレームワークを導入し、単位の割り当てなどのシーケンシャルな意思決定の一般的な課題に対処します。ゲーム後のリフレクションとルール抽出により、モデルの重みを変更せずに三目並べに測定可能な改善がもたらされることを示します。
原文 (English)
Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory
Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood. We study this on fully-observable two-player zero-sum games, which provide ground-truth evaluation: outcomes are determined by the rules, and optimality of individual moves can be computed or approximated, without relying on a judge model. Across model tiers, LLMs play suboptimally in simple games such as tic-tac-toe or Connect Four, and lose to MCTS opponents. Obfuscations that preserve the game tree but rewrite its surface form leave performance largely unchanged, indicating the gap is not fully explained by recall of memorized strategies. Motivated by this performance gap, we introduce an agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment. We show that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying the model weights.
State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for…
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, st…
LLaDA MoE v2: 専門家混合の普及言語モデルのスケーリング
拡散言語モデル (dLLM) は、自己回帰 (AR) 言語モデリングの代替手段を提供しますが、専門家混合 (MoE) dLLM のスケーリング動作はまだよく理解されていません。 MoE dLLM の最適化ハイパーパラメータ、コンピューティング割り当て、アーキテクチャがどのように拡張されるかを体系的に特徴付け、AR モデルについて以前に報告された拡張傾向との定量的な違いを特定します。具体的には、最適化の場合、最適な名目バッチ サイズはより速く増加しますが、最適な学習率はコンピューティングに伴ってより急速に減衰します。モデルのデータ割り当てについては、IsoFLOP 分析によりデータ側のわずかな傾きが明らかになります。最適なトークン バジェットは、アクティブ化されたモデル側の計算よりも速く増加します。 MoE アーキテクチャの場合、大規模なスケールでは、固定のアクティブ化されたキャパシティでより大きなエキスパート プールがますます好まれますが、中程度のエキスパートの粒度は一貫して効果的であり、共有エキスパートに割り当てられるアクティブ化されたキャパシティーの優先割合はスケール全体で安定しています。これらの発見に基づいて、私たちは 30B-A3B dLLM である LLaDA MoE v2 を 23.5T トークンで最初からトレーニングします。 Qwen3 と約 65\% の数の事前トレーニング トークンを備えた LLaDA MoE v2 は、いくつかの知識、推論、コーディング ベンチマークにおいて Qwen3 に近づきます。監視された微調整だけを行った後、推論とコーディングのベンチマーク 8 つのうち 7 つで SDAR Chat を上回り、いくつかのタスクでは Qwen3 に近いパフォーマンスを維持しました。これらの結果は、MoE dLLM の実際的なスケーリング則と設計原則を確立します。
原文 (English)
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.
例によるプログラミングのためのソルバーを意識した分解: 分割するには克服方法を知る必要がある場合
分解ベースの例によるプログラミング (PBE) は、学習されたシンセサイザーが解決するサブタスクにタスクを分割することによってパフォーマンスをスケールします。デコンポーザーは中間のサブ目標を予測し、シンセサイザーはそれらに条件付けされたプログラムを生成します。現在のアプローチは、分解の品質をタスクに固有のものとして暗黙的に扱い、グラウンドトゥルース (GT) サブゴールを模倣するように分解者をトレーニングします。私たちはこの仮定に異議を唱えます。固定誘導バイアスを持つ有界ソルバーの場合、GT 分解はソルバーの検索ダイナミクスではなく、アノテーターの因数分解の選択を反映します。したがって、GT 分解に一致するように訓練された分解者は、論理的には有効だがソルバーにとって扱いにくいサブ目標を提案する可能性があります。我々は、GT サブゴールの教師ありトレーニングを構造的な足場として保持し、さらにフリーズしたシンセサイザーからの直接フィードバックを介してデコンポーザーを最適化するトレーニング フレームワークであるソルバー認識分解 (SAD) を提案します。サブゴールは、ターゲット プログラムでのシンセサイザーの損失、つまりソルバーが動作できる分解を促進するサブタスクの難易度のシグナルに基づいて報酬が与えられます。私たちの実験では、精度のパラドックスが明らかになりました。GT 分解との一致度が高くても、合成の成功率は向上しません。たとえシンセサイザーがまったく同じ GT データでトレーニングされたとしても、デコンポーザーは模倣するように最適化されています。代わりに、SAD はソルバーの扱いやすさのために GT アライメントを犠牲にする分解を学習し、2 つの PBE ドメイン全体で合成とエンドツーエンドのタスクの精度を一貫して向上させます。さらに、SAD は、GT 分解オラクルが失敗したタスクを解決します。これは、GT 分解が有界ソルバーにとって普遍的に最適ではないこと、および分解の品質が固有ではなくソルバーに依存することを示す経験的証拠です。
原文 (English)
Solver-Aware Decompositions for Programming-by-Example: When Dividing Requires Knowing how to Conquer
Decomposition-based Programming-by-example (PBE) scales performance by splitting tasks into subtasks that a learned synthesizer solves: a decomposer predicts intermediate subgoals, and a synthesizer generates programs conditioned on them. Current approaches train the decomposer to imitate ground-truth ( GT) subgoals, implicitly treating decomposition quality as intrinsic to the task. We challenge this assumption: for bounded solvers with fixed inductive biases, GT decompositions reflect the annotator's factorization choices - not the solver's search dynamics. A decomposer trained to match GT decompositions may therefore propose subgoals that are logically valid yet intractable for the solver. We propose Solver-Aware Decomposition (SAD), a training framework that retains supervised training on GT subgoals as a structural scaffold, while additionally optimizing the decomposer via direct feedback from a frozen synthesizer. Subgoals are rewarded based on the synthesizer's loss on the target program - a signal of subtask difficulty that encourages decompositions the solver can act on. Our experiments reveal an accuracy paradox: higher agreement with GT decompositions does not improve synthesis success - even though the synthesizer was trained on the very same GT data the decomposer is optimized to mimic. SAD instead learns decompositions that trade GT alignment for solver tractability, yielding consistent gains in synthesis and end-to-end task accuracy across two PBE domains. Moreover, SAD solves tasks that a GT decomposition oracle fails - empirical evidence that GT decompositions are not universally optimal for bounded solvers, and that decomposition quality is solver-relative, not intrinsic.
LeanMem: LLM エージェント向けのシンプルで効率的な長期メモリ
LLM ベースのエージェントが対話を維持し、遠い歴史を確実に活用するには、長期記憶が不可欠です。ただし、既存のメモリ システムは通常、均一な要約および取得パイプラインを通じて異種の対話コンテンツを処理するため、過剰なトークンの消費またはきめの細かい証拠の不可逆的な損失につながります。私たちは、歴史的な対話コンテンツは、その圧縮性、時間的ダイナミクス、忠実性の要件に応じて異なる方法で処理されるべきであると主張します。この洞察に基づいて、軽量の長期記憶フレームワークである LeanMem を提案します。 LeanMem は、最初に価値の低いコンテンツをフィルタリングして除外し、情報の性質に応じて、情報セグメントをコンパクトなプロファイル メモリ、時間的に構造化されたイベント メモリ、またはソースに基づいたレコード メモリとして保存します。メンテナンス中は、動的に進化するイベント メモリのみが選択的に更新され、安定したプロファイルと不変のレコードの冗長な統合が回避されます。推論中、LeanMem はメモリ タイプを動的に選択し、クエリ固有の証拠要求に従って取得バジェットを割り当て、要求に応じて関連する証拠を収集します。 GPT-4.1-mini および Qwen3-8B を備えた LoCoMo および LongMemEval-S では、LeanMem は、構築コスト、推論トークン、レイテンシが最低または最低に近い状態で、あらゆる設定において最も強力なメモリベースのベースラインよりも精度を最大 15.1 ポイント向上させます。コードとデータセットは補足資料に含まれています。
原文 (English)
LeanMem: Simple and Efficient Long-Term Memory for LLM Agents
Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typically process heterogeneous dialogue content through a uniform summarization and retrieval pipeline, leading to either excessive token consumption or irreversible loss of fine-grained evidence. We argue that historical dialogue content should be handled differently according to its compressibility, temporal dynamics, and fidelity requirements. Based on this insight, we propose LeanMem, a lightweight long-term memory framework. LeanMem first filters out low-value content, then stores informative segments as compact profile memory, temporally structured event memory, or source-grounded record memory, depending on the nature of the information. During maintenance, only dynamically evolving event memories are selectively updated, avoiding redundant consolidation of stable profiles and immutable records. During inference, LeanMem dynamically selects memory types and allocates retrieval budgets according to query-specific evidence demands, assembling relevant evidence on demand. On LoCoMo and LongMemEval-S with GPT-4.1-mini and Qwen3-8B, LeanMem improves accuracy over the strongest memory-based baseline in every setting, by up to 15.1 points, at the lowest or near-lowest construction cost, inference tokens, and latency. The code and datasets are included in the supplementary materials.
ChartAnno: チャート注釈生成のための MLLM の評価
マルチモーダル大規模言語モデル (MLLM) は、チャートの理解、生成、編集において大幅な進歩を遂げていますが、既存のチャートに注釈を付ける機能についてはまだ十分に解明されていません。チャートに注釈を付けることは、一般的ではありますが、やりがいのあるコミュニケーション タスクであり、モデルが意図したメッセージを推測し、チャートのセマンティクスを解釈し、適切なテキストまたはグラフィック要素を配置する必要があります。このギャップに対処するために、チャートの注釈生成に関する MLLM を評価するためのベンチマークである ChartAnno を導入します。これには、3 つの命令特異性レベルにわたるコードと注釈命令のペアを含む 1,200 の実世界のチャートが含まれています。 10 個の代表的な MLLM を 2 つの主要な入力設定 (1) チャート コードのみ、(2) チャート コードとチャート画像の両方で評価し、さらにチャート画像のみのアブレーション研究も含めます。結果は、大規模なオープンソース モデルがその差を縮めているものの、全体的には独自モデルが引き続き強力であることを示しています。より具体的な命令によりアノテーションの品質が向上しますが、抽象的な意図を推測することは現在の MLLM にとって依然として最も困難です。チャート画像を提供しても全体的な利益は限られており、改善は主にデザイン関連の指標に現れます。これらの調査結果は、チャートの注釈生成が意味論的な基礎と効果的な注釈設計を必要とする困難なタスクであることを強調しています。コードとデータは将来のバージョンでリリースされる予定です。
原文 (English)
ChartAnno: Evaluating MLLMs for Chart Annotation Generation
Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements. To address this gap, we introduce ChartAnno, a benchmark for evaluating MLLMs on chart annotation generation. It contains 1,200 real-world charts with paired code and annotation instructions across three levels of instruction specificity. We evaluate 10 representative MLLMs under two primary input settings: (1) chart code alone and (2) both chart code and chart image, and further include a chart image-only ablation study. Results show that proprietary models remain stronger overall, although large-scale open-source models narrow the gap. More specific instructions improve annotation quality, while inferring abstract intent remains most difficult for current MLLMs. Providing chart images brings limited overall gains, with improvements mainly appearing in design-related metrics. These findings highlight chart annotation generation as a challenging task requiring semantic grounding and effective annotation design. Code and data will be released in a future version.
When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO
Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GR…
ToolLIFT: 一般化可能な工具計画のために工具固有の軌跡を機能レベルのグラフにリフトアップ
過去のツール使用の軌跡は、大規模言語モデル (LLM) エージェントがツールの使用を計画および調整するための貴重な経験を提供します。既存のアプローチは、これらの軌跡からツールレベルのグラフを直接構築しますが、結果として得られるグラフは特定のツールに関連付けられたままであり、ツールセット全体で一般化するのが困難です。この課題に取り組むために、関連するツールの違いにもかかわらず、類似したタスクが共通の機能レベルのワークフロー構造を共有していることが多く、これがツール計画のより移転可能な抽象化として機能する可能性があることがわかりました。この洞察に基づいて、一般化可能なツール計画のためにツール固有の軌跡を機能レベルのワークフロー グラフ (FWG) に引き上げるフレームワークである ToolLIFT を提案します。具体的には、まず、FWG でワークフロー構造をエンコードし、ツール間でコラボレーション エクスペリエンスを共有する軌道リフティング メカニズムを提案します。次に、FWG のグローバル構造に基づいて、ワークフローの計画とツールの選択を分離して、個々のツールの選択をワークフロー全体と調整します。最後に、信頼性の高いツール データフローを確保するために、強化学習 (RL) を採用し、ツール呼び出し全体でソース追跡可能な情報フローを維持するために、ソースゲート型およびスキル固有の報酬を提案します。 2 つのディストリビューション内 (ID) ベンチマークと 3 つのディストリビューション外 (OOD) ベンチマークでの実験では、ToolLIFT が常に最先端のベースラインを上回るパフォーマンスを示し、これまでにないツール セットに対する強力な一般化が実証されました。
原文 (English)
ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning
Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-level graphs from these trajectories, but the resulting graphs remain tied to specific tools and are hard to generalize across tool sets. To tackle this challenge, we find that despite differences in the tools involved, analogous tasks often share a common function-level workflow structure, which serves as a potentially more transferable abstraction for tool planning. Based on this insight, we propose ToolLIFT, a framework that lifts tool-specific trajectories into a function-level workflow graph (FWG) for generalizable tool planning. Specifically, we first propose a trajectory-lifting mechanism that encodes workflow structures in the FWG and shares collaboration experience across tools. Then, building on the global structure of the FWG, we introduce decoupled workflow planning and tool selection to align individual tool choices with the overall workflow. Lastly, to ensure reliable tool dataflow, we adopt Reinforcement Learning (RL) and propose source-gated and skill-specific rewards to maintain source-traceable information flow across tool calls. Experiments on two in-distribution (ID) and three out-of-distribution (OOD) benchmarks show that ToolLIFT consistently outperforms state-of-the-art baselines, demonstrating strong generalization to unseen tool sets.
WeClawArena: 人間中心のエージェント ネットワークにおけるクロスユーザー エージェントのコラボレーションとセキュリティのための監査可能なサンドボックスおよびベンチマーク
永続的なパーソナル エージェント フレームワークの最近の進歩により、人間中心のエージェント ネットワークが現実的な展開対象になりつつあります。各ユーザーは、ユーザーに代わって動作し、状態を維持し、ソーシャルおよびタスク関係を通じて他のエージェントと通信する AI エージェントによってサービスを受けることができます。これらのネットワークでは、日常的なツールの使用は、個人のワークスペースを介した複数の当事者による所有エージェントのコラボレーションとなり、ファイル、レコード、ツール、ポリシーは所有者間で直接表示されません。既存のエージェントのベンチマークは、ツールの使用とコラボレーションを研究しますが、現実的なユーザーのデジタル ワークスペースとの検証可能なクロスユーザー エージェントのコラボレーションのためのエンドツーエンドのサンドボックスを提供したり、人間中心のエージェント ネットワークを介して有害なアクションがどのように伝わるかをテストしたりすることはありません。個人のワークスペース上で複数の当事者が所有するエージェントのコラボレーションのための監査可能なベンチマークおよびランタイム サンドボックスである WeClawArena を紹介します。 WeClawArena は、個人のワークスペースが運用ツールと個人の制約の両方として機能する共同ツール使用タスクを対象としています。このベンチマークには、6 つのクロスユーザー タスク ドメインにわたる 124 の基本タスクが含まれており、それらを基本タスクごとに 1 つの無害な制御と 4 つの攻撃ベクトル バリアントを含む 620 のシナリオ バリアントに拡張します。サンドボックスは、ピア メッセージ、ツール呼び出し、リソース操作、管理された決定、およびワークスペースの最終状態を記録します。 WeClawArena は、ユーティリティと攻撃の成功率を個別に報告し、制限された実行時の証拠から攻撃の成功を監査し、タスクの中断、プライバシーの漏洩、有害な証拠、および無効な権限パスの診断をサポートします。
原文 (English)
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.
LLM は高品質の実験をデザインできますか?自律的実験計画に関する包括的かつ体系的なベンチマーク
AI for Research (AI4Research) は、AI を活用して科学ワークフローを自動化および改善します。実験計画は研究プロセスの重要な段階ですが、これまでの研究は主にコードの実装と実行に焦点を当てており、この段階の重要性が見落とされており、体系的な実験計画を実行する AI の能力を評価するベンチマークは存在していません。このギャップを埋めるために、私たちは、一流の研究機関(ICML、NeurIPS、ICLRなど)からの19の研究領域にわたる300の高品質な最新論文から構築された科学的総合計画評価ベンチマークであるSCOPEを提案します。これは、高レベルの計画の完全性(メイン、アブレーション、および分析実験)と低レベルの構成の精度と合理性(データセット、ベースライン、およびデータセット)の2つの次元でLLMを評価します。メトリクス)。ベンチマークにより、次の 3 つの結果が明らかになりました。(1) ほとんどの LLM は、高品質の実験を直接設計できません。 (2) すべての LLM は、低レベル構成ではパフォーマンスのボトルネックを示します。 (3) 検索モードではデザインの品質は向上しません。さらに、これらの課題に対処するために、LLM ベースの実験計画を最適化する新しいエージェント ワークフローである OptED を提案します。これは、段階の分離、ツールの拡張、ルールベースの制約を通じて LLM ベースの実験計画を強化し、構成のボトルネックを効果的に軽減します。
原文 (English)
Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design. To bridge this gap, we propose SCOPE, a Scientific COmprehensive Planning Evaluation Benchmark constructed from 300 high-quality latest papers across 19 research domains from top-tier venues (e.g., ICML, NeurIPS, and ICLR),evaluating LLMs on two dimensions: High-Level planning completeness (main, ablation, and analysis experiments) and Low-Level configuration accuracy and rationality (datasets, baselines, and metrics). Benchmarking reveals three findings: (1) most LLMs cannot directly design high-quality experiments; (2) all LLMs exhibit a performance bottleneck in low-level configuration; and (3) search mode does not improve design quality. Furthermore, to address these challenges, we propose OptED, a novel agentic workflow to optimize LLM-based experimental design, that enhances LLM-based experimental planning through stage isolation, tool augmentation, and rule-based constraints, effectively alleviating the configuration bottleneck.
Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks
Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous…
多くの回答が有効な場合、投票は失敗します: LLM における Best-of-K 因果推論の記号的検証
自己一貫性では、サンプリングされた推論トレースの中で最も頻度の高い回答が最も信頼できると想定されますが、これは因果推論では失敗する可能性があります。サンプルは同じ交絡エラーを繰り返すことが多く、投票は複数の有効な回答に分散し、有効な少数のトレースにもかかわらず無効な回答が勝利します。 CALVER (Causal Axiom-Level VERification) を紹介します。これは、分離、バックドア調整、介入を含む、Pearl の因果基準に照らして構造化トレースをスコアリングし、参照回答を参照せずに最高スコアの候補を選択する、トレーニング不要のシンボリック検証ツールです。複数のグラフ有効回答を許可する CLEAR find-one-valid クエリでは、CALVER は 42.1% に達し、複数、報酬モデル、LLM 判定、モデルの信頼度は同一の凍結プールで 30% 近くを維持します。ジャッジを 72B にスケールしてもギャップは埋まりません。監査されたクリーンコアのサブセットでは、グラフで有効な CALVER 選択 21 個のうち 11 個が、要求された述語を満たしながらもベンチマークのリストされた回答と異なります。この利点はサンプリング予算とともに拡大し、10 個の公開されたベイジアン ネットワーク、2 番目のモデル ファミリ、およびモデルがテキストからグラフを構築する必要がある設定にわたって再現されます。 CALVER はまた、正確なグラウンド トゥルースに対するしきい値による平均治療効果の決定を改善し、真理値表チェッカーの下でロジックを一般化し、CPU 上でミリ秒単位で各候補をスコア付けします。 CALVER に必要なのは、直接提供されるかテキストから構築された因果構造のみです。それが当てはまる場合は、因果関係の妥当性によって選択を集約できます。
原文 (English)
When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs
Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, and intervention, and selects the highest-scoring candidate without consulting a reference answer. On CLEAR find-one-valid queries that admit multiple graph-valid answers, CALVER reaches 42.1% where plurality, a reward model, an LLM judge, and model confidence remain near 30% on identical frozen pools. Scaling the judge to 72B does not close the gap. In an audited clean-core subset, 11 of 21 graph-valid CALVER selections differ from the benchmark's listed answer while still satisfying the requested predicate. The advantage widens with the sampling budget and reproduces across ten published Bayesian networks, a second model family, and settings where the model must build the graph from text. CALVER also improves thresholded average-treatment-effect decisions against exact ground truth, generalizes to logic under a truth-table checker, and scores each candidate in milliseconds on CPU. CALVER needs only a causal structure, supplied outright or built from the text; wherever that holds, selection can aggregate via causal validity.
大規模な言語モデルでの矢印の反転
大規模言語モデル (LLM) は、テキストから知識へのグラフ生成および関連タスクで優れたパフォーマンスを達成しました。それにもかかわらず、引数の順序を逆転させると関係の意味が変わるという逆関係の方向依存の意味論を正確にモデル化しているかどうかはまだ不明です (例: \textit{mother} 対 \textit{child})。私たちの知る限り、この研究は、27 の異なる逆関係ラベルにわたる 5,457 のインスタンスからなるベンチマークを使用した、LLM における逆関係の方向性に関する最初の体系的な研究を示しています。私たちは、多肢選択プロンプト フレームワークの下で 5 つのオープンソース LLM を評価し、元のエンティティを合成エンティティやマスクされたエンティティに置き換えることによって、関係記述とエンティティ表現の影響をさらに調査します。私たちの調査結果は、LLM 間の逆関係分類における体系的な非対称性を明らかにし、関係の記述が一貫してパフォーマンスを向上させるわけではないことを示し、モデルのパフォーマンスがエンティティ表現の変動に敏感になる可能性があることを示しています。
原文 (English)
Reversing Arrows in Large Language Models
Large language models (LLMs) have achieved strong performance on text-to-knowledge graph generation and related tasks. Nevertheless, it is still unclear whether they accurately model the direction-dependent semantics of inverse relations, in which reversing the order of the arguments alters the meaning of a relation (e.g., \textit{mother} versus \textit{child}). To the best of our knowledge, this work presents the first systematic study of inverse relation directionality in LLMs, using a benchmark consisting of 5,457 instances spanning 27 distinct inverse relation labels. We evaluate five open-source LLMs under a multiple-choice prompting framework and further examine the influence of relation descriptions and entity representations by substituting the original entities with synthetic and masked entities. Our findings reveal systematic asymmetries in inverse relation classification across LLMs, indicate that relation descriptions do not consistently improve performance, and show that model performance can be sensitive to variations in entity representations.
アジェントノミクス博士: アジェントノミクスの教訓的実験
AGENTONOMICS は、AI エージェントを、統合管理アーキテクチャを通じて設計、管理、統治できる経済的実体として扱うフレームワークです。 Dr. AGENTONOMICS はその最初のアプリケーションであり、経営管理における AI エージェントに関する TUM コースの文脈で開発された講義エージェントです。 2025/26 年の冬学期に考案され、2026 年の夏学期に初めて学生に導入されたこのプログラムは、エージェントが学生の学習対象であると同時に、フレームワークを学習して適用する媒体でもあるという教訓的な実験として機能します。現在のプロトタイプは、AGENTONOMICS の概念を説明し、学生の質問をサポートする、Web ベースの検索ベースの家庭教師です。このレポートでは、同じシステムが個別指導を超えて、マルチモーダルな指導を行うアバター講師、AGENTONOMICS Design & Management Reference Framework (ADMRF) を通じて学生を指導するデザイン コンサルタント、学生が指定したエージェントの構築を支援するメタエージェントという 3 つの追加の累積的な役割に成長する可能性があると主張しています。これらのロールは同じインターフェイス、インテリジェンス レイヤー、ツール、ナレッジ ベース、エコシステム接続を共有するため累積的であり、オーケストレーターは各タスクに必要なロール固有のアルゴリズムを選択します。プロトタイプのアーキテクチャを紹介し、その開発ロードマップの概要を示し、多中心的な AI 経済への影響について議論します。このレポートは、エージェントがどのように設計されたフレームワークを教え、適用し、最終的に再現できるかについてさらなる議論を促すことを目的としています。
原文 (English)
Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS
AGENTONOMICS is a framework that treats AI agents as economic entities that can be designed, managed, and governed through an integrated management architecture. Dr. AGENTONOMICS is its first application: a lecture agent developed in the context of the TUM course on AI agents in business administration. Conceived during the winter semester 2025/26 and first introduced to students in the summer semester 2026, it serves as a didactic experiment in which the agent is both the object that students study and the medium through which they learn and apply the framework. The current prototype is a web-based, retrieval-grounded tutor that explains AGENTONOMICS concepts and supports student questions. This report argues that the same system can grow beyond tutoring into three additional cumulative roles: an avatar lecturer that delivers multimodal instruction, a design consultant that guides students through the AGENTONOMICS Design & Management Reference Framework (ADMRF), and a meta-agent that helps construct the agents students have specified. These roles are cumulative because they share the same interface, intelligence layer, tools, knowledge base, and ecosystem connection, while an orchestrator selects the role-specific algorithm required for each task. We present the architecture of the prototype, outline its development roadmap, and discuss its implications for a polycentric AI economy. This report is intended to invite further discussion on how agents can teach, apply, and eventually reproduce the frameworks by which they are designed.
包括的かつ回復力のあるデジタル評価の実施のための行動適応型視覚転換
教育機関は、一か八かのデジタル評価を確保するために、ブラウザーのロックダウン、ウェブカメラの監視、行動分析への依存度を高めていますが、これらのメカニズムは一般的に独立して設計および評価されており、学習者のアクセシビリティが見落とされていることがよくあります。この論文では、合成の非意味論的視野を評価コンテンツと合成し、観察された候補者の行動に応じて適応的に調整する理論的フレームワークである行動適応型視覚転換 (BAVD) を紹介します。基礎となる評価内容は決して変更されません。正当な候補者への侵入を最小限に抑えながら、不正な画面キャプチャまたは画面共有の有用性を減らすために、視覚的なプレゼンテーションのみが変更されます。このフレームワークにはさらに、承認された視覚処理機能を持つ候補者の気晴らしの強度を低減または抑制する、アクセシビリティを意識した減衰メカニズムが組み込まれています。私たちは、ダイバージョンフィールドジェネレーター、レンダリングテンソル、動作テンソル、複合整合性関数、および多次元エントロピーモデルで構成される結合力学システム表現を使用してモデルを定式化し、コンテンツ忠実度、レンダリング安定性、エントロピー境界性、整合性追跡、および閉ループ適応安定性に関する理論的特性を確立します。このフレームワークでは、脅威モデルを明示し、展開の前提条件と制限を特定し、アクセシビリティとキャプチャ耐性の間のトレードオフについて議論します。この研究は、行動適応性とアクセシビリティを意識した評価配信のための数学的に根拠のある基盤を提供し、信頼できるデジタル評価プラットフォームにおける将来の実証的検証の基礎を提供します。
原文 (English)
Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery
Institutions increasingly rely on browser lockdown, webcam monitoring, and behavioral analytics to secure high-stakes digital assessments, yet these mechanisms are commonly designed and evaluated independently and often overlook learner accessibility. This paper introduces Behaviorally-Adaptive Visual Diversion (BAVD), a theoretical framework in which a synthetic, non-semantic visual field is composited with assessment content and adaptively modulated according to observed candidate behavior. The underlying assessment content is never altered; only its visual presentation is modified to reduce the usefulness of unauthorized screen capture or screen sharing while remaining minimally intrusive for legitimate candidates. The framework further incorporates an accessibility-aware attenuation mechanism that reduces or suppresses diversion intensity for candidates with approved visual-processing accommodations. We formulate the model using a coupled dynamical-systems representation comprising a Diversion Field Generator, Rendering Tensor, Behavior Tensor, Composite Integrity Functional, and Multi-dimensional Entropy Model, and establish theoretical properties for content fidelity, rendering stability, entropy boundedness, integrity tracking, and closed-loop adaptation stability. The framework explicitly states its threat model, identifies deployment assumptions and limitations, and discusses the trade-off between accessibility and capture resistance. This work provides a mathematically grounded foundation for behaviorally adaptive and accessibility-aware assessment delivery and offers a basis for future empirical validation in trusted digital assessment platforms.
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was i…
コンテキストを意識したセマンティック埋め込みによる表形式学習者の強化
最新の表形式学習器は統計パターンの捕捉には優れていますが、多くの場合、意味論的な空白の中で動作し、テキストの特徴を個別のシンボルとして扱い、特徴名やセルのエントリに固有の豊富なセマンティクスを無視します。私たちは、大規模言語モデル (LLM) の意味理解と表形式学習者の統計能力の間のギャップを埋める新しいフレームワークである CASE (Context-Aware Semantic Embeddings) を提案します。行を分離して埋め込む既存の方法とは異なり、CASE はコンテキスト化戦略を利用します。カスタム トレーニングされた Gemma 3 ベースの表形式言語モデルの KV キャッシュに行の代表的なサンプルを事前に入力して、データセットのセマンティクスの永続的なアンカーを確立します。これにより、生成された行埋め込みが動的にコンテキスト化され、セマンティックな曖昧さが解決され、ドメイン固有のコンテキストで表現が固定されることが保証されます。いくつかのベンチマーク (CARTE、TextTab、TabArena) にわたる実験では、CASE が意味的に豊富なデータセット、特に低データ領域での表形式学習器のパフォーマンスを大幅に向上させることが実証されました。
原文 (English)
Enhancing Tabular Learners with Context-Aware Semantic Embeddings
While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discrete symbols, ignoring the rich semantics inherent in feature names or cell entries. We propose CASE (Context-Aware Semantic Embeddings), a novel framework that bridges the gap between the semantic understanding of Large Language Models (LLMs) and the statistical capabilities of tabular learners. Unlike existing methods that embed rows in isolation, CASE utilizes a contextualization strategy: we pre-fill the KV cache of a custom-trained Gemma 3-based Tabular Language Model with a representative sample of rows to establish a persistent anchor of the dataset's semantics. This ensures that generated row embeddings are dynamically contextualized, resolving semantic ambiguities and anchoring representations in domain-specific context. Our experiments across several benchmarks (CARTE, TextTab, and TabArena) demonstrate that CASE substantially improves the performance of tabular learners on semantically rich datasets, particularly in low-data regimes.
AI 科学者の能力をベンチマークするためのテストベッドとしての敵対的で高速に移動する現実世界のドメイン
AI 科学者が斬新なアイデアを生み出す能力をベンチマークすることは、非常に難しいことで知られています。この分野の既存のベンチマークは、科学的推論と研究の再現性の評価において進歩を遂げていますが、多くの場合、合成タスクや遡及的なターゲットに依存しており、以前の暴露によって混乱する可能性があります。私たちは、専門家が独自に観察可能な出力を生成する、複雑で敵対的で動きの速い現実世界の領域が、このギャップを埋め、推論、新規性、仮説の定式化など、AI 科学者に必要な能力を評価するための実用的なソリューションを提供できると仮説を立てています。私たちはこのフレームワークを 2 つの構造的に異なるドメインでインスタンス化します。フォーミュラ 1 (F1) では、モデルが 2026 年シーズンに向けた車のデザイン コンセプトを構想し、実際のシーズン前のイノベーションが真実を提供します。もう一方、マジック: ザ ギャザリング (MTG) では、モデルが最近更新されたカード プールからデッキを提案し、19 のプロツアー (PT) デッキリストに対して評価されます。どちらの領域でも、モデルは妥当な出力を生成しますが、現実世界の専門家によるソリューションと一致するものはほとんどありません。最高のモデルである F1 では、GPT-5.2 は、走行全体で提案された 166 のアイデアと実際のイノベーション 40 件のうち 10 件を一致させました。 MTG では、ジェミニ 3 フラッシュの最高のデッキが 3 位の PT デッキから 7 枚の新セット カードのうち 5 枚を回収し、全 108 デッキにわたって、最も頻繁に選ばれたカード モデルは PT デッキで最も広く採用されているカードでもありました (スピアマン $\rho = 0.74$、$p = 0.0003$)。これらの結果は、AI 科学者の主な能力ギャップはアイデア生成ではなく、フィルタリング、優先順位付け、一貫した新規性であることを示唆しています。
原文 (English)
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro Tour (PT) decklists. Across both domains, models produce plausible outputs, but few align with real-world expert solutions. In F1, the best model, GPT-5.2 matched 10 of 40 real innovations with 166 ideas proposed across runs. In MTG, the best deck from Gemini 3 Flash recovered 5 of 7 new-set cards from the third-place PT deck, and across all 108 decks, the cards models selected most often were also the cards most widely adopted by PT decks (Spearman $\rho = 0.74$, $p = 0.0003$). These results suggest that a key capability gap for AI scientists is not idea generation, but filtering, prioritization, and coherent novelty.
政策の断片化か、それとも制度の連携か?大学およびビジネススクールにおける AI の制度的ガバナンス
人工知能 (AI) は高スキル分野を急速に変革しており、高等教育機関 (HEI) は基礎原則の教育と、労働力の即応性を確保するための新しいツールの統合のバランスをとることが求められています。高等教育機関では AI の導入が進んでいますが、特にそのようなポリシーが教育機関のさまざまなレベルで設定されている場合、多くの高等教育機関が AI をどのようにカリキュラムに組み込み、ポリシーを通じて管理するかについて引き続き取り組んでいます。この研究では、米国の 34 州の高等教育機関全体の AI 政策を分析し、これらの政策が何を意味するのか、また、教育機関全体および教育機関のさまざまなレベル内で設定されたポリシーがどのように異なるかを調査します。自然言語処理 (NLP) を使用して組織の AI ポリシーを分析すると、明らかな相違が見つかりました。大学レベルのポリシーはデータ セキュリティとリスク軽減を重視するのに対し、学校レベルのポリシーが存在する場合、教育上のアプリケーションとツールの使用に焦点を当てます。ビジネス スクール固有のポリシーに焦点を当てる場合、大学の枠組みとは異なる AI ポリシーを維持するビジネス スクールは比較的少なく、専門分野固有の学習目標との不整合が生じます。このギャップは、認定の目的だけでなく、特に教員や学生にとっても課題となっています。私たちの洞察は、ガイドラインは、専門分野固有の学習目標と進化する労働力の需要に対処しながら、より広範な組織のポリシーと整合する必要があることを示唆しています。
原文 (English)
Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools
Artificial intelligence (AI) is rapidly transforming high-skilled domains, requiring higher education institutions (HEI) to balance the teaching of foundational principles with the integration of emerging tools to ensure workforce readiness. While HEI are increasingly adopting AI, many continue to grapple with how it should be incorporated into curricula and governed through policy, especially when such policies are set at different levels of an institution. This research analyzes AI policies across HEI from 34 states in the United States to investigate what these policies entail and how policies set across institutions as well as within different levels at an institution differ. Using natural language processing (NLP) to analyze institutional AI policies, we find a clear divergence: university-level policies emphasize data security and risk mitigation whereas school-level policies, when present, focus on pedagogical applications and tool usage. When focusing on business school specific policies, relatively few business schools maintain AI policies distinct from university frameworks, creating misalignment with discipline-specific learning objectives. This gap poses challenges particularly for faculty and students as well as for accreditation purposes. Our insights suggest that guidelines should be aligned with broader institutional policies while addressing discipline-specific learning objectives and evolving workforce demands.
ソーシャル コーディングからエージェント コーディングへ: オープンソース コミュニティにおける生産性とリレーショナル再構成
オープンソース ソフトウェア コミュニティは、コードを生成するだけでなく、目に見えるコラボレーションを通じて公共の知識や人間関係を生み出すデジタル公共インフラの一形態です。生成コーディング エージェント (CA) は、アクティビティの一部をパブリックな人間の対話からプライベートな人間とエージェントのループに移行しながら、開発効率を向上させる高度なツールです。私たちは、1,084 人のアクティブな開発者とそのリポジトリ関係からの実際の GitHub データで初期化された LLM ベースのマルチエージェント シミュレーションを使用して、この変化を調査しました。履歴コミットによるウォームアップの後、4 週間のシミュレーションのために、同じコミュニティ状態を並列の No-CA 条件と CA 条件に分岐します。 CA の導入により、計画されたタスクと完了したタスクがそれぞれ 34.0% と 39.0% 増加し、完了時間の中央値が 45 分から 20 分に短縮されました。ただし、導入率は 26.0% にとどまっており、その恩恵はすでによりアクティブでつながりが充実している開発者に集中しています。 CA はタスクの実行経路も再構築します。人間と人間の直接的なインタラクションは 32.4% から 11.6% に減少しますが、CA が関与するモードは 57.3% に増加し、そのうちの 40.3% は CA 支援による自己ループを通じて完了します。 CA 条件下で生成された公開知識も、後のタスクに対するサポートが少なくなります。標準化された検索ベンチマークでは、CA コーパスは 22.3% の知識カバー率を達成していますが、これは実際の人間のコーパスが達成する 81.1% をはるかに下回り、より多くの検索ステップが必要となり、成功率は低くなります。これらの結果は、生産性と公開知識の緊張関係を明らかにしています。コーディング エージェントは技術的な生産性を高めますが、より多くの作業がエージェント仲介またはプライベート ループに移行し、将来の貢献者にとって公開記録の有用性が低下します。
原文 (English)
From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities
Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interpersonal relationships through visible collaboration. Generative coding agents (CAs) are an advanced tool to improve development efficiency while shifting part of activities from public human interaction to private human-agent loops. We study this shift using an LLM-based multi-agent simulation initialized with real GitHub data from 1,084 active developers and their repository relationships. After a warm-up with historical commits, we branch the same community state into parallel No-CA and CA conditions for 4-week simulations. CA introduction increases planned and completed tasks by 34.0% and 39.0%, respectively, and reduces median completion time from 45 to 20 minutes. However, adoption reaches only 26.0%, and the gains concentrate among developers who are already more active and well connected. CAs also restructure task execution pathways. Direct human-human interaction declines from 32.4% to 11.6%, while CA-involved modes increase to 57.3%, including 40.3% completed through CA-assisted self-loops. Public knowledge generated under CA condition also provides less support for later tasks. On a standardized retrieval benchmark, the CA corpus achieves 22.3% knowledge coverage, far below the 81.1% achieved by the real-human corpus, and requires more retrieval steps with a lower success rate. These results reveal a productivity-public knowledge tension: coding agents increase technical production, but more work shifts to agent-mediated or private loops, leaving public records less useful to future contributors.
FOUND-AF: 心房細動検出のための ECG 基礎モデルのベンチマーク
心房細動(AF)は最も一般的な持続性不整脈であり、脳卒中、心不全、死亡のリスク増加と関連しています。最近の ECG 基礎モデルは、自動 AF 検出のための転送可能な表現を提供します。ただし、既存の研究では異なるデータセット、前処理手順、分類子、検証プロトコルが使用されているため、それらの相対的な有効性は不明のままです。この研究では、同一の実験条件下で事前トレーニングされた ECG 表現の品質を評価する、統合され、漏れが制御され、展開指向のベンチマーク フレームワークである FOUND-AF を紹介します。 HuBERT-ECG、CLEF、ST-MEM、ECG-JEPA、ECGFounder を含む 5 つのファミリーからの 9 つの公的に利用可能な基礎モデルが、AFDB、CinC2017、CPSC2021、LTAFDB という 4 つの異種 ECG データセットにわたって評価されました。すべてのモデルは、標準化された前処理、モデルネイティブのリサンプリング、固定 XGBoost 分類器、および記録レベルのグループ化相互検証を備えた凍結特徴抽出器として使用されました。評価には、分類メトリック、受信機動作特性分析、ホルム補正を備えたペアの記録レベル ブートストラップ比較、埋め込み空間の視覚化、および計算効率プロファイリングが含まれます。 ECGFounder モデルは、精度、モデル サイズ、推論時間、メモリ使用量の間で有利なトレードオフを提供しながら、データセット全体で最も強力な全体パフォーマンスを一貫して達成しました。したがって、FOUND-AF は、ECG 基礎モデルを選択するための再現可能なフレームワークを提供し、臨床的に事前トレーニングされたコンパクトなエンコーダーが、異種収集設定全体で堅牢かつ計算効率の高い AF 検出をサポートできることを実証します。
原文 (English)
FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality. Recent ECG foundation models offer transferable representations for automated AF detection. However, their relative effectiveness remains unclear because existing studies use different datasets, preprocessing procedures, classifiers, and validation protocols. This study presents FOUND-AF, a unified, leakage-controlled, and deployment-oriented benchmarking framework that evaluates the quality of pretrained ECG representations under identical experimental conditions. Nine publicly available foundation models from five families, including HuBERT-ECG, CLEF, ST-MEM, ECG-JEPA, and ECGFounder, were evaluated across four heterogeneous ECG datasets, namely AFDB, CinC2017, CPSC2021, and LTAFDB. All models were used as frozen feature extractors with standardized preprocessing, model-native resampling, a fixed XGBoost classifier, and recording-level grouped cross-validation. The evaluation included classification metrics, receiver operating characteristic analysis, paired recording-level bootstrap comparisons with Holm correction, embedding-space visualization, and computational efficiency profiling. The ECGFounder model consistently achieved the strongest overall performance across datasets while offering a favorable trade-off between accuracy, model size, inference time, and memory usage. FOUND-AF therefore provides a reproducible framework for selecting ECG foundation models and demonstrates that compact, clinically pretrained encoders can support robust and computationally efficient AF detection across heterogeneous acquisition settings.
偏微分方程式ワークフローのための大規模な言語モデル
偏微分方程式 (PDE) は、孤立した式としてではなく、モデリングの仮定、支配方程式、数値ソルバー、診断、意思決定を結び付ける実行可能なワークフローとして、科学や工学において実用的になります。大規模言語モデル (LLM) は、自然言語、記号数学、コード、ソルバー出力、フィードバックをリンクすることで、このようなワークフローをサポートし始めています。ここでは、支配モデルの発見と定式化、実行可能な数値ソルバーの生成と修正、制御、設計、最適化をサポートするためのシミュレーション フィードバックの使用という 3 つの段階にわたる、LLM 支援 PDE 研究の最近の進歩を検証します。これらの段階にわたって、現在のシステムは主にワークフロー レベルのインターフェイスとして機能します。このような進歩にもかかわらず、この分野は依然として高品質のデータセットとベンチマークの不足によって制限されており、特に知識発見や現実世界のアプリケーションでは、専門家のアノテーション、実行可能な問題の構築、およびタスクレベルのフィードバックに多大なドメイン労力が必要となります。さらなる課題は、シミュレーションベースの結果と現実世界の科学および工学システムとの間に依然としてギャップがあり、数値シミュレーション、制御ポリシー、および最適化された設計を実際の設定に直接適用することを制限していることです。これらの課題により、LLM 支援 PDE ワークフローは、言語、計算、物理的制約、現実世界の意思決定を結び付けることができる科学 AI システムを開発するための重要なテストベッドとなっています。
原文 (English)
Large language models for partial differential equation workflows
Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows that connect modelling assumptions, governing equations, numerical solvers, diagnostics, and decisions. Large language models (LLMs) are beginning to support such workflows by linking natural language, symbolic mathematics, code, solver outputs, and feedback. Here we examine recent advances in LLM-assisted PDE research across three stages: the discovery and formulation of governing models, the generation and revision of executable numerical solvers, and the use of simulation feedback to support control, design, and optimization. Across these stages, current systems act primarily as workflow-level interfaces. Despite this progress, the field remains limited by the scarcity of high-quality datasets and benchmarks, especially for knowledge discovery and real-world applications, where expert annotation, executable problem construction, and task-level feedback require substantial domain effort. A further challenge is the persistent gap between simulation-based results and real-world scientific and engineering systems, which limits the direct transfer of numerical simulations, control policies, and optimized designs to practical settings. These challenges make LLM-assisted PDE workflows a critical testbed for developing scientific AI systems that can connect language, computation, physical constraints, and real-world decision-making.
FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation
Federated fine-tuning with Low-Rank Adaptation (LoRA) enables efficient collaborative adaptation of Large Language Models (LLMs) without ce…
臨床試験戦略の学習: 意思決定主体のためのオフライン方針トレーニング
臨床開発は不確実性の下での逐次的な意思決定であり、スポンサーは異質な証拠に基づいて実験のポートフォリオを計画する必要があります。私たちは、腫瘍学の臨床開発をオフラインの意思決定の問題として枠組み付けて、この設定を研究します。この問題では、エージェントが決定日時点で入手可能な情報から腫瘍学医薬品プログラムの次の 6 か月の治験ポートフォリオを予測します。これをサポートするために、治験登録、規制審査、スポンサー申請、利用データ、疫学を含む 31.7k の異種公開データ記録を、45 の歴史的プログラムにわたる 881 のオフライン意思決定エピソードに結合する一時データセットを構築します。行動クローニング、報酬重み付け行動クローニング、学習報酬トレーニング、および価値ベースの暗黙的 Q 学習という 4 つのオフライン目標を、保留薬物、スポンサー、薬物クラス、時間分割にわたって共通の日付ゲート検索足場を共有する 4 つのフロンティア LLM エージェントと比較します。オフラインでトレーニングされたモデルは、特に 2025 年 8 月以降の汚染除去のホールドアウトにおいて、微調整されていないベースラインよりも優れたパフォーマンスを発揮します。報酬重み付けされた行動クローニングは最も優れたパフォーマンスを示し、各メトリクスで最もパフォーマンスの高いツール エージェントの 25.0% と 2.1% に対して、指示 F1 が 46.2%、厳格な F1 が 14.2% 得られました。これらの結果は、構造化されたオフライン学習がエージェントに臨床実験の計画を教えることができることを示唆しています。
原文 (English)
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available at the decision date. To support this, we construct a temporal dataset that combines 31.7k heterogeneous public data records, including trial registries, regulatory reviews, sponsor filings, utilization data, and epidemiology, into 881 offline decision episodes across 45 historical programs. We compare four offline objectives: behavioral cloning, reward-weighted behavioral cloning, learned-reward training, and value-based implicit Q-learning against four frontier LLM agents that share a common date-gated retrieval scaffold across held-out drug, sponsor, drug-class, and temporal splits. Models trained offline outperform the non-fine-tuned baselines, particularly in the post-August 2025 contamination-clean holdout. Reward-weighted behavioral cloning performs the best, obtaining 46.2% indication F1 and 14.2% strict F1 against 25.0% and 2.1%, respectively, for the best-performing tool agent on each metric. These results suggest that structured offline learning can teach agents to plan clinical experiments.
運用データに対するエージェント システムの正式な検証
大規模言語モデル (LLM) によって駆動されるエージェント システムは、永続的な運用データに基づいて動作する実世界のワークフローに導入されることが増えています。導入前に、これらのシステムは、ワークフローの実行とデータの進化を管理するビジネス要件に照らして検証する必要があります。ただし、既存のアプローチは主にエージェントのインターフェイス レベルで動作を制限または分析するため、そのようなシステム レベルの保証は提供されません。ここでは、単一の LLM とリレーショナル運用データ上のツール オーケストレーション ハーネスで構成されるエージェント システムの検証を研究します。これらを Stateful Tool-Enabled Agentic Deployments (STEAD) として形式化し、セマンティクスを与え、一次計算ツリー ロジック (FO-CTL) 仕様に照らして検証する問題を定義し、それが決定不可能であることを示します。我々は、有限領域の制限の下で FO-CTL 仕様を正確に保存するための十分な条件を特定し、その検証は PSPACE で完了します。重要な要件は、データ内の不透明な識別子の名前を変更すると、それに対応して、選択したツール呼び出しの名前も変更する必要があるということです。私たちは、LLM 駆動のエージェントがこの条件に違反する可能性があることを示し、すでに等価な動作を維持しながら、任意のベース エージェントに対してそれを保証する正規の展開ラッパーを導入します。この構造に必要な正準表現の計算はグラフ同型性が難しいことを証明します。最後に、ケース管理ワークフローを調整する LLM エージェントに関するフレームワークを示します。
原文 (English)
Formal Verification of Agentic Systems over Operational Data
Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data. Before deployment, these systems need to be verified against business requirements that govern workflow execution and data evolution. However, existing approaches do not provide such system-level guarantees, as they mainly constrain or analyse behaviour at the agent's interface level. We study here the verification of agentic systems comprising a single LLM and a tool orchestration harness over relational operational data. We formalise them as Stateful Tool-Enabled Agentic Deployments (STEADs), give their semantics, define the problem of verifying them against First-Order Computation Tree Logic (FO-CTL) specifications, and show that it is undecidable. We identify sufficient conditions for exact preservation of FO-CTL specifications under a finite-domain restriction, over which verification is PSPACE-complete. The key requirement is that renaming opaque identifiers in the data must correspondingly rename the selected tool calls. We show that LLM-driven agents can violate this condition and introduce a canonical deployment wrapper that guarantees it for arbitrary base agents while preserving already-equivariant behaviour. We prove that computing canonical representations required by this construction is graph-isomorphism-hard. Finally, we illustrate our framework on an LLM agent orchestrating a case-management workflow.
不完全な観察によるマルチモーダル感情分析におけるモダリティの信頼性の再考
マルチモーダル感情分析 (MSA) は、テキスト、音声、視覚を統合して人間の感情を推測しますが、現実世界のマルチモーダル観察は多くの場合不完全です。不完全観察 MSA の既存の方法は、主に 2 つのパラダイムに従います。再構成ベースの手法は観察されたモダリティから欠落した情報を回復しますが、共同表現手法は不完全な入力から直接学習します。これらの方法は効果的ではありますが、通常、モダリティの信頼性を明示的にモデル化するのではなく、表現学習または融合設計内で暗黙的にのみ処理します。我々は、モダリティの信頼性が不完全観察設定における中心的な変数であると主張します。これを明示的にモデル化しないと、2 つの関連する問題が発生します。 1 つ目は信頼性の不一致であり、各モダリティによって保持される感情的証拠がサンプルと欠落率によって異なります。 2 つ目は信頼性伝播バイアスであり、劣化したモダリティからのメッセージがクロスモーダル インタラクションや予測パフォーマンスに悪影響を与える可能性があります。これらの問題に対処するために、我々は不完全な観測を伴う MSA 用のモダリティ信頼性校正フレームワークである MRCF を提案します。 MRCF には、モーダル内の品質キューとクロスモーダルのセマンティック一貫性からサンプル固有のモダリティの信頼性を推定する信頼性認識ブランチ、推定されたスコアを使用してクロスモーダル情報フローを調整する信頼性ガイド付きインタラクション ブランチ、および最終予測のために信頼性とセマンティック キューを統合する信頼性校正済みフュージョン モジュールが含まれています。 CMU-MOSI、CMU-MOSEI、および CH-SIMS の実験では、MRCF が標準的な不完全観察プロトコルの下で強力なパフォーマンスを達成することが示されています。さらなる分析により、明示的な信頼性モデリングが相互作用および融合中の信頼性の不一致および信頼性伝播バイアスの軽減に役立つという証拠が得られます。
原文 (English)
Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations
Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-based methods recover missing information from observed modalities, while joint-representation methods learn directly from incomplete inputs. Although effective, these methods usually treat modality reliability only implicitly within representation learning or fusion design rather than modeling it explicitly. We argue that modality reliability is a central variable in incomplete-observation settings. Failure to model it explicitly gives rise to two related issues. The first is reliability mismatch, in which the affective evidence retained by each modality varies across samples and missing rates. The second is reliability propagation bias, in which messages from degraded modalities may adversely affect cross-modal interaction and predictive performance. To address these issues, we propose MRCF, a Modality Reliability-Calibrated Framework for MSA with incomplete observations. MRCF contains a Reliability-Aware Branch that estimates sample-specific modality reliability from intramodal quality cues and cross-modal semantic consistency, a Reliability-Guided Interaction Branch that uses the estimated scores to modulate cross-modal information flow, and a Reliability-Calibrated Fusion Module that integrates reliability and semantic cues for final prediction. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show that MRCF achieves strong performance under standard incomplete-observation protocols. Further analyses provide evidence that explicit reliability modeling helps mitigate reliability mismatch and reliability propagation bias during interaction and fusion.
Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection
Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context rem…
重み空間アブレーション下の層間相互作用: 閉じた形式のアテンション ヤコビアン境界と実際の事前学習モデルでのテスト
関連論文では、条件付き計算が残差ストリームを通じて加算的に実行される理想化されたモデル内で、アクティベーション パッチとウェイト スペース アブレーションがいつ一致するかを研究しています。 2 つのキャリアがアーキテクチャ的に依存するそのモデル内の 1 つの構成、つまりアテンション ヘッドとその独自の層の正規化 MLP 構成については、正確な 1 次相互作用公式、MLP のみがアブレーションされる場合はゼロ、ヘッドもアブレーションされる場合は 2 次境界が導出されます。その結果は単一の残差ブロックに限定され、合成タスクの小さな変換器でのみチェックされます。この論文では、両方の限界を超えて結果を拡張します。まず、複数の層にまたがるアブレーションキャリアからの相互作用は、タッチされた層ごとに 1 つと、分解が小さいことを主張しない層間の残りの、同じブロックの項に正確に分解されます。次に、混合二次導関数の二重積分として 2 つの層の残りを正確に分離し、それを結合するために必要な欠落成分に名前を付けます。これは、注目サブブロックに結合するヤコビアンです。この境界を閉じた形式で導出し、Qwen2.5-1.5B-Instruct の実際の重みに対して違反が 1 つも発生しないことを検証しますが、まだレイヤー間で連鎖させていません。また、閉じた形式で、コンパニオンペーパーの綴じたままの曲率定数を与えます。第三に、同じモデル上で、このタスクの元の活性化パッチ手法を使用して、決して設計されていない間接オブジェクト識別のための創発回路を検索して見つけ、それに対する崩壊、解離、および相互作用をテストします。結果はまちまちです。共有キャリアはテストされた 5 つのインスタンスすべてで出現し、崩壊と解離はすべてではありませんがほとんどで保持され、コンパニオン定理がカバーする同一ブロックのケースの外側の層ペアで、5 つのうち 3 つで非ゼロ相互作用が測定可能です。
原文 (English)
Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two carriers are architecturally dependent, an attention head and its own layer's normalization-MLP composition, it derives an exact first-order interaction formula, zero when only the MLP is ablated and second-order bounded when the head is also ablated. That result is confined to a single residual block and checked only on small transformers on a synthetic task. This paper extends the result past both limits. First, the interaction from ablating carriers spanning several layers decomposes exactly into same-block terms, one per touched layer, plus a cross-layer remainder on which the decomposition makes no claim of smallness. Second, we isolate that remainder exactly, for two layers, as a double integral of a mixed second derivative, and name the missing ingredient needed to bound it: a Jacobian bound for the attention sub-block. We derive this bound in closed form and verify it, without a single violation, against Qwen2.5-1.5B-Instruct's real weights, though we do not yet chain it across layers. We also give, in closed form, the curvature constant the companion paper's bound leaves unexhibited. Third, on that same model, we search for and find an emergent circuit for indirect object identification, never designed into it, using the original activation-patching method for this task, and test collapse, dissociation, and interaction on it. The result is mixed: a shared carrier emerges across all five tested instances, collapse and dissociation hold on most but not all, and a nonzero interaction is measurable on three of five, at layer pairs outside the same-block case the companion theorem covers.
教師が誤解を招く場合: スプリアス信号を認識したオンポリシーの抽出
オンポリシー蒸留 (OPD) は、高密度のトークンレベルの教師信号で生徒がサンプリングした軌跡を監視することにより、教師の能力を転送します。最近の選択的 OPD 手法では、信頼できる信号、有益な信号、または学習可能な信号を優先することで、このプロセスを改善しています。しかし、この仮定は、言語モデルの根本的な失敗モードを見落としています。つまり、言語モデルのトークンレベルの判断は、タスク固有の証拠ではなく、入力に依存しない言語事前分布、書式設定規則、またはステレオタイプの推論テンプレートによって駆動される可能性があります。このような最適化に関連するが入力に根拠が弱い監視を OPD のスプリアス信号と呼びます。このような信号は、タスク改善の方向にはほとんど貢献しないものの、大きな勾配を生成する可能性があります。この問題を軽減するために、入力の根拠と最適化の影響に基づいて誤解を招くトークンレベルの監視を特定してフィルタリングする、Spurious-Signal-Aware On-Policy Distillation フレームワークである SA-OPD を提案します。 SA-OPD は、トークンレベルの蒸留信号が本当に入力に依存しているかどうかを推定する軽量の入力接地プロキシを導入します。次に、低い入力接地性と極端な蒸留発散を同時に示すトークンのみをフィルタリングすることで、影響の大きい偽の更新を除去し、きめ細かい OPD 最適化を実現します。大規模言語モデル (LLM) と視覚言語モデル (VLM) の両方の設定に関する広範な実験により、SA-OPD がバニラ OPD および競合する選択手法よりも一貫して優れていることが実証されました。これらの結果は、入力接地性が OPD 監視選択の重要な要素であることを確立し、スプリアス アップデートを軽減するためのシンプルで効果的な戦略を提供します。
原文 (English)
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.
シード間クロスプレイで十分ですか?実装の詳細に対するゼロショット調整アルゴリズムの堅牢性の評価
現実世界の環境に配置された AI エージェントは、これまで遭遇したことのない人間や他の AI エージェントと連携できなければなりません。ゼロショット調整 (ZSC) アルゴリズムは、独立して設計されたエージェントがテスト時に相互に調整できるように、高レベルの学習ルールを指定することでこれを達成することを目的としています。 ZSC アルゴリズムの厳密な評価は依然として困難です。理想的には、独立した当事者が同じ仕様を解釈して実装するときに生じる変動を反映して、提案された各アルゴリズムの複数の独立した実装を使用する必要があります。ただし、実際には、ZSC アルゴリズムは、異なるランダム シード間でトレーニングされた単一の実装を使用してほぼ排他的に評価されており、ニューラル ネットワーク アーキテクチャをさらに変更する作業はほんの一握りです。このため、仕様の曖昧さと実装の詳細に対する堅牢性については未解決の疑問が残ります。この研究では、この堅牢性の最初の体系的な評価を提供します。我々は、これまでの研究でマルチエージェント強化学習 (MARL) アルゴリズムのパフォーマンスに影響を与えることが示されているさまざまな実装の詳細であるクロス実装クロスプレイという新しい評価スキームを導入し、人気のある ZSC アルゴリズムである Other-Play をこのスキームで評価します。私たちの発見は心強いものであり、Other-Play の場合、標準の ZSC 評価が、実際には、このより徹底した実装間評価の合理的な代用となることを示唆しています。
原文 (English)
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot coordination (ZSC) algorithms aim to achieve this by specifying high-level learning rules such that independently engineered agents can coordinate with each other at test time. Rigorous evaluation of ZSC algorithms remains difficult: ideally, multiple independent implementations of each proposed algorithm must be used, reflecting the variation that arises when independent parties interpret and implement the same specification. In practice, however, ZSC algorithms have almost exclusively been evaluated using a single implementation trained across different random seeds, with only a handful of works additionally varying the neural network architecture. This leaves open questions about robustness to specification ambiguities and implementation details. In this work, we provide the first systematic evaluation of this robustness. We introduce a new evaluation scheme, cross-implementation cross-play, varying implementation details that prior work has shown to affect the performance of multi-agent reinforcement learning (MARL) algorithms, and we evaluate Other-Play, a popular ZSC algorithm, with this scheme. Our findings are encouraging and suggest that, for Other-Play, the standard ZSC evaluation is, in fact, a reasonable proxy for this more thorough cross-implementation evaluation.
AutoSND: 実行証拠から自動ネットワーク解体ヒューリスティック検出のための構造ポリシーまで
ネットワークの解体は、複雑なシステムの堅牢性と脆弱性を分析するための基礎ですが、実用的なヒューリスティックは有効性と計算効率のバランスをとる必要があり、通常は研究者によって手動で設計されます。既存の大規模言語モデルに基づく自動ヒューリスティック設計手法は、候補を生成してスクリーニングすることはできますが、実行中の候補の品質や失敗状態を次の世代のための構造レベルの指針にさらに変換することは困難です。私たちは、完全なネットワーク解体プログラムのための 3 段階のツリー検索フレームワークである AutoSND を提案します。ステージ I では、単純なヒューリスティックから幅広く調査し、実行証拠をアーカイブします。ステージ II では、候補レコードを、ローカル信号、近隣アクセス、および状態更新範囲に関する構造ポリシーにコンパイルします。ステージ III では、これらのポリシーに基づいてツリー検索を続行し、品質優先および速度優先の最終候補である AutoSND-Q/S を取得します。 12 の実世界ネットワークと 3 つの大規模実世界ネットワークでの実験では、AutoSND がより優れた検索パフォーマンスと安定性を実現し、より競争力があり構造的に解釈可能なネットワーク解体プログラムを発見したことが示されています。最終候補は、残差度をバックボーンとして使用し、境界のあるローカル信号でノード順序を調整し、状態更新範囲を制限する解釈可能な構造を形成します。コードは https://github.com/MirrorNew/AutoSND で入手できます。
原文 (English)
AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery
Network dismantling is fundamental to analyzing the robustness and vulnerability of complex systems, yet practical heuristics must balance effectiveness and computational efficiency, and are usually designed manually by researchers. Existing large language model based automatic heuristic design methods can generate and screen candidates, yet they have difficulty further transforming candidate quality or failure states during execution into structural-level guid- ance for subsequent generation. We propose AutoSND, a three stage tree search framework for complete network dismantling pro- grams. Stage I broadly explores from simple heuristics and archives execution evidence. Stage II compiles candidate records into struc- tural policies concerning local signals, neighborhood access, and state update ranges. Stage III continues tree search conditioned on these policies and obtains the final quality prioritized and speed prioritized candidates, AutoSND-Q/S. Experiments on 12 real world networks and 3 large real world networks show that AutoSND achieves better search performance and stability and discovers more competitive and structurally interpretable network disman- tling programs. The final candidates form an interpretable structure that uses residual degree as the backbone, adjusts node order with bounded local signals, and restricts the state update range. Code is available at https://github.com/MirrorNew/AutoSND.
Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training
Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal la…
高次の安全性を実現するシールド
安全シールドは、安全を保証するためにコントローラーの動作を制限する実行時強制メカニズムです。古典的なシールドは通常、状態述語用に合成されます。つまり、現在の物理状態は安全か安全でないかのいずれかであり、シールドは、システムを将来的に安全でない状態に強制する可能性のあるアクションを正確に無効にします。多くのサイバー物理アプリケーションでは、このビューは粗すぎます。障害物に接近する車両は、衝突を回避するだけでなく、怪我を防ぐために、速度規制、加速によって引き起こされる力の制限、ジャーク制限を遵守する必要があります。物理的な観点から見ると、これらの要件は状態の派生物よりも前提となります。この論文では、このような高次の滑らかさ制約に対する有限状態セーフティ ゲーム構築を開発します。離散化された状態空間上の有限差分を使用して微分安全特性を定義し、その表現力を特徴付け、シールド合成を履歴状態空間上の通常の安全ゲームに還元します。シールドが $k$ 次のプロパティの過去の状態を正確に $k$ 格納する合成アルゴリズムを与え、このメモリが必要であることを証明します。導関数制約の階層にわたって動作する最大許容シールドの反復合成手順について説明します。このアルゴリズムは、制約を昇順に繰り返し解決し、各反復でその解を使用して次の制約に向けて状態空間を刈り込みます。これにより、アルゴリズムは安全でないことが知られている状態空間の広い領域の探索を控えるので、実際にはシールド合成がより効率的になります。
原文 (English)
Shielding for Higher-Order Safety
Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety. Classical shields are usually synthesised for state predicates: the current physical state is either safe or unsafe, and the shield disables precisely those actions that can force the system into an unsafe state in the future. In many cyber-physical applications this view is too coarse. A vehicle approaching an obstacle should not only avoid collision, but also respect speed regulations, force limits induced by acceleration, and jerk limits to prevent injuries. From a physical perspective, these requirements are predicated over the derivatives of the state. This paper develops a finite-state safety-game construction for such high-order smoothness constraints. We define differential safety properties using finite differences over a discretised state space, characterise their expressiveness, and reduce shield synthesis to an ordinary safety game over a history state space. We give a synthesis algorithm whose shields store exactly $k$ past states for properties of order $k$ and prove that this memory is necessary. We describe an iterative synthesis procedure for a maximally permissive shield that operates over hierarchies of derivative constraints. The algorithm solves constraints iteratively in increasing order and uses the solution at each iteration to prune the state space for the next constraint. This makes shield synthesis more efficient in practice, as the algorithm refrains from exploring large regions of the state space that are known to be unsafe.
PhyAI: エッジでのリアルタイム物理 AI、クラウドでのスケーラブルなロールアウト
物理 AI ポリシーでは、モデルの評価、クラウド強化学習のロールアウト、エッジ GPU の提供、オンボード展開などのライフサイクル全体にわたって推論が必要です。これらの設定は同じチェックポイントとアクションのセマンティクスを共有しますが、多くの場合、別の推論プログラムに依存します。これらを統合するために、グラフの実行、カーネル、メモリ管理、並列サービスを共有しながら、アーキテクチャ固有のコンディショニング、ソルバー、キャッシュ、出力ロジックをモデル アダプターに保持する単一のランタイムを備えた物理 AI 推論エンジンである PhyAI を構築します。同じコードベースは、オンボード、エッジ、クラウドの展開全体で単一または複数の GPU 上でビジョン言語アクション (VLA) モデルとワールド アクション モデル (WAM) を実行します。 MiniCPM-Robot のリリース日にアダプター インターフェイスを使用して追加しました。 PhyAI は、pi0、pi0.5、GR00T N1.7、および MiniCPM-Robot の公式実装と比較して 1.40 倍から 4.65 倍の高速化を達成します。 Cosmos3-Nano-Policy-DROID では、8 つの H20 GPU (CFG=2、TP=4) でレイテンシが 2.46 秒から 1.18 秒に短縮され、2.08 倍の速度向上になります。特殊なランタイムはいくつかの構成で引き続き高速であるため、私たちの目標は、すべてのケースで最速の結果ではなく、競争力のあるレイテンシーを備えた 1 つのランタイムです。詳細なプロファイルにより、異なるモデルに異なる実行ポリシーが必要な理由が明らかになります。バッチ サイズ 1 の Hopper シリーズ GPU では、pi0.5 アクション エキスパートは FLOP の 8.8% を占めますが、レイテンシーの 57.2% を占めます。バッチ サイズ 32 では、そのシェアは 13.5% に低下し、スループットは約 100 サンプル/秒に達します。 Cosmos3 は世代主導のままで、バッチ サイズが 1 から 16 に増加してもスループットは 14.3% しか向上しません。さらに、推論制限制御と環境制限制御を区別する制御時間ルーフラインを導入します。 4 つの LIBERO スイートで測定された pi0.5 ポイントは環境に依存していますが、Cosmos3 は推論に依存したままです。コードとベンチマーク: https://github.com/mingti-org/phyai。
原文 (English)
PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.
LiveEvalBench: Web 生成のオープンワールド評価に向けて
大規模な言語モデルは実行可能なフロントエンド プロジェクトを合成できるようになってきていますが、既存のベンチマークは依然として Web 生成を静的評価問題として扱っています。私たちは、フロントエンド アーティファクトには異なるパラダイムが必要であると主張します。フロントエンド アーティファクトは静的ではなくインタラクティブであり、多様でありながら同様に有効な実装を許容し、厳格なパイプラインが対応できるよりも速く進化します。これらのギャップに対処するために、Web 生成の評価をエージェント的、適応的、拡張可能なプロセスとして再定式化する自動フレームワークである LiveEvalBench を紹介します。 LiveEvalBench は、共同レビュー ワークフローとして評価をインスタンス化します。このワークフローでは、ビルド エンジニア、コード エンジニア、UI テスターが、デプロイメントやコード検査からブラウザベースの操作に至るまで、フロントエンド プロジェクトのライフサイクル全体にわたって証拠を共同で収集します。実装の多様性に対処するために、適応プロトコルは、モデル間の比較を可能にする共有ルーブリックを、各成果物に合わせた実装に基づいた基準と組み合わせます。このフレームワークは、パイプラインを再設計することなく、新しい評価者の役割と評価次元の段階的な統合をさらにサポートします。現実世界の多様な Web 生成シナリオにわたる実験では、LiveEvalBench が人間の専門家の判断と密接に一致し、フロンティア モデルの Web 生成機能についてのきめ細かい洞察が得られることが示されています。コードは https://github.com/wyysteelhead/LiveEvalBench で入手できます。
原文 (English)
LiveEvalBench: Toward Open-World Evaluation for Web Generation
Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web-generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser-based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross-model comparability with implementation-grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench
TARL: 長期エージェントの実行可能メモリ管理のためのトランザクション対応の信頼できる台帳
永続的なメモリは、エージェントが知識を長期間保持するのに役立ちますが、単一の更新エラーにより、その後の検索と推論が繰り返し歪められる可能性があります。既存のシステムのほとんどは、メモリの更新をバイナリの書き込み/保持の決定に減らしており、新しい情報を追加すべきか、無視すべきか、古い信念を修正するために使用すべきか、信頼できないとして拒否すべきか、検証を延期すべきかを区別できません。これらの選択は、基本的に異なるメモリ状態を生成しながら、同じバイナリ ラベルを共有する可能性があります。各ステートメントを 5 つの実行可能なアクションの 1 つにマップするメモリ状態更新フレームワークである TARL を紹介します。 TARL は、影響を受けるメモリを特定し、その時間範囲を解決し、ソースの信頼性を比較し、受け入れられた台帳、保留中の台帳、および拒否された台帳を更新します。さらに、代替の更新操作によって生成されたメモリ状態を比較することでトレーニングされ、モデルが正しい結果につながる操作を選択するように促されます。また、きめ細かいアクション ラベルと次の状態のターゲットを備えたベンチマークである TARL-Mem も紹介します。 TARL は、ドメイン内、クロスソース、時間的、反事実的、逐次的な評価にわたって、アクションの予測と状態回復を改善し、メモリ汚染を軽減し、矛盾する証拠を保存し、累積的な破損を制限します。完全なモデル実装は補足資料で提供されます。
原文 (English)
TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents
Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existing systems reduce memory updating to a binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification. These choices may share the same binary label while producing fundamentally different memory states. We introduce TARL, a memory state update framework that maps each statement to one of five executable actions. TARL identifies the affected memory, resolves its temporal scope, compares source reliability, and updates accepted, pending, and rejected ledgers. It is further trained by comparing the memory states produced by alternative update operations, encouraging the model to select the operation that leads to the correct result. We also introduce TARL-Mem, a benchmark with fine-grained action labels and next-state targets. Across in-domain, cross-source, temporal, counterfactual, and sequential evaluations, TARL improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption. The complete model implementation is provided in the supplementary material.
トラフィックが減り、結果が向上: リアルタイム Ad Exchange での競合を意識したリクエストのディスパッチ
リアルタイム入札 (RTB) アド エクスチェンジは通常、入札を受け取るのはほんの一部であっても、ほぼすべての受信リクエストをデマンドサイド プラットフォーム (DSP) に転送します。この過剰な配分はオークションの結果を弱めます。DSP はコンピューティングと予算の制約の下で参加を抑制し、限られた入札能力の効果的な使用を減らします。分散入札予測と確率的転送を使用して、各リクエストを各 DSP に送信するかどうかを決定する、競合を意識したリクエスト ディスパッチ フレームワークを紹介します。このシステムは、軽量ポリシーの最適化を通じて DSP ごとのしきい値を時間の経過とともに適応させ、非定常な市場状況を追跡します。私たちは、毎日 200 億を超えるリクエストに対応する運用プラットフォーム上で 4 回の連続したオンライン実験を通じてフレームワークを評価しました。完全なマルチ DSP 導入により、最初の DSP 適応期間後の最近 14 日間で、ポリシーに基づく DSP リクエスト量が 34.2% 減少し、純収益が 4.6% (p<0.001) 増加しました。さらに分析を進めると、トラフィック セグメント間の強い異質性が浮き彫りになり、集計指標が誤解を招く可能性があることが明らかになりました。セグメントレベルおよび DSP ごとの分析は、このポリシーが DSP 間の比較優位性を明らかにし、全体的なリクエスト量を増やすことなく収益化の成果を向上させることを示唆しています。
原文 (English)
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges
Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. This over-distribution weakens auction outcomes: DSPs throttle participation under compute and budget constraints, reducing the effective use of limited bidding capacity. We present a competition-aware request dispatch framework that uses distributional bid prediction and probabilistic forwarding to decide whether each request should be sent to each DSP. The system adapts per-DSP thresholds over time through lightweight policy optimization to track non-stationary market conditions. We evaluate the framework through four sequential online experiments on a production platform serving over 20 billion daily requests. A full multi-DSP deployment reduces DSP request volume under the policy by 34.2% while increasing net revenue by 4.6% (p<0.001) in a recent 14-day window after an initial DSP adaptation period. Further analysis highlights strong heterogeneity across traffic segments and reveals that aggregate metrics can be misleading. Segment-level and per-DSP analyses suggest that the policy surfaces comparative advantages among DSPs, improving monetized outcomes without increasing overall request volume.
アウトプットが分散すると、認識論の修正が続くのか? Machine Collective 向けのブラックボックス カップリング診断
集合知の研究では、意見の相違を認識論的多様性の証拠として扱います。エージェントが異なる見解を表明した場合、グループは修正する能力を保持する必要があります。 LLM 集合体では、このプロキシが壊れる可能性があります。エージェントは、同じ結論を維持しながら、多様に見える議論を生成できます。私たちは分散と修正の結合を操作します。つまり、埋め込み空間における集団の成果の分散を検証可能に増加させる介入が、前提を保持した再定式化ではなく認識論的立場の真の修正を伴う度合いです。診断はブラックボックスです。生成されたテキストのみを処理し、生成モデルの内部表現については主張しません。 2 つのチャネルは独立して測定されます。出力チャネルであるコヒーレンス インデックス (CI) は、介入によって出力分散が変化したことを検証します。認識チャネル、ターンごとのスタンスの注釈は、集合体が修正されたかどうかを測定します。我々は、この結合領域を推定するための再利用可能な方法として、出力が過収束したときに再微分プロトコル (RDP) を挿入するメタ予測明瞭度システム (MPCS) を備えた CI を提案します。 2 つの構成 (gpt-4o-mini および gemini-2.5-flash、条件ごとに 310 ペアのエピソード) から 5 つのエージェント集合を評価します。 gpt-4o-mini では、条件付き反対は誤った前提の回復を +17.7 ポイント (p<1e-6) 改善しますが、静的なペルソナの多様性は回復に悪影響を及ぼします (-8.1、p=.007)。ジェミニ 2.5 フラッシュでは、分散の低下が確認されたにもかかわらず、同等の予算で同じ介入を行っても利益は得られませんでした (26.1% 対 27.1%、p=.84)。 2 つの治療効果は互いに異なります (z=3.79、p<.001)。メカニズムのタグ付けは、ジェミニがフレームワーク内の反対意見を通じて誤った前提を維持していることを示しています。タグ付けされたRDP後の回答の94%が譲歩せずに再定式化しました(GPTでは24%)。精度とともに、介入ごとのスタンスシフトと前提保存率を報告することをお勧めします。
原文 (English)
When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.
SAT-Edge-Agent: オンボード衛星インテリジェンスのためのハードウェアインザループ エッジエージェント オーケストレーション
オンボード衛星インテリジェンスには、ミッションの意図をローカル ツール呼び出しに変換し、実行状態を公開し、通信と電力の制約の下でマシンが消費可能なアーティファクトを返すタスク層が必要です。市販の ARM ベースの異種エッジ システムオンチップに展開されたハードウェアインザループ (HIL) エッジ エージェント システムである SAT-Edge-Agent を紹介します。ブラウザー ワークスペースと FastAPI エージェントは、ローカルの OpenAI 互換言語サービスと、FAIR1M メタデータに基づいた構造化結果を返すプロジェクト内部の YOLO スタイル指向オブジェクト検出エンドポイントを調整します。 2 つの固定 FAIR1M ワークロード (1 つの単一イメージ要求と 1 つのシリアル 2 イメージ要求) がそれぞれ 20 回繰り返され、20/20 の試行が完了しました。フルエージェントの平均待ち時間は 29.353 秒と 60.937 秒で、経験的な P95 値は 31.166 秒と 66.882 秒でした。平均検出時間は 861.386 ミリ秒と 1510.920 ミリ秒で、対応するフルエージェント平均のわずか 2.93% と 2.48% でした。プロファイリングにより、目に見える遅延のほとんどは検出器の実行外で発生していることがわかります。平均 CPU 使用率は 20.761% と 20.482% でした。 200 ミリ秒の NPU 負荷フィールドは両方のワークロードで平均 100% でしたが、これは検出器のみの占有率や調整された使用率ではなく、共有アクセラレータ ソフトウェア フィールドを表しています。公開証拠パッケージは、サニタイズされたリクエストレベルのレコード、編集された JSON、正規化された SSE サンプル、および報告された統計を再現するスクリプトを提供します。これらの結果は、観測可能な衛星エッジ エージェント オーケストレーションのための再現可能な HIL 境界を確立しますが、検出器の精度、新しい地理位置情報方法、校正されたエネルギー効率、または飛行準備状態を確立するものではありません。
原文 (English)
SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence
Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine-consumable artifacts under communication and power constraints. We present SAT-Edge-Agent, a hardware-in-the-loop (HIL) edge-agent system deployed on a commercial off-the-shelf ARM-based heterogeneous edge system-on-chip. A browser workspace and FastAPI agent coordinate a local OpenAI-compatible language service with a project-internal YOLO-style oriented-object-detection endpoint that returns FAIR1M metadata-backed structured results. Two fixed FAIR1M workloads, one single-image and one serial two-image request, were repeated 20 times each and completed 20/20 attempts. Mean Full-Agent latency was 29.353 s and 60.937 s, with empirical P95 values of 31.166 s and 66.882 s. Mean detector time was 861.386 ms and 1510.920 ms, only 2.93% and 2.48% of the corresponding Full-Agent means. Profiling indicates that most visible latency occurs outside detector execution. Mean CPU utilization was 20.761% and 20.482%. A 200-ms NPU-load field averaged 100% for both workloads, but it represents a shared-accelerator software field rather than detector-only occupancy or calibrated utilization. The public evidence package provides sanitized request-level records, redacted JSON, normalized SSE examples, and scripts reproducing the reported statistics. These results establish a reproducible HIL boundary for observable satellite edge-agent orchestration, but do not establish detector accuracy, a new geolocation method, calibrated energy efficiency, or flight readiness.
CARE-Bench: 患者対応 LLM トリアージのベンチマーク
患者と接する医療 LLM やエージェントは、臨床医に連絡する前に症状の質問に答えることが増えています。ここで重要な安全性の問題は、ユーザーが次にどのような行動をとるべきかということです。 CARE-Bench は、ターンごとの 4 ラベルの電流アクション タスクとして、患者に面した連続トリアージを評価する、ソースに基づいたベンチマークです。 CARE-Bench には、医療対話、相談、フォローアップ質問のソースから再構築された 500 件の症例と 1,059 件の評価済み患者開示プレフィックスが含まれています。私たちは、固定 GPT-5.5 マッパーを使用して各応答を 4 ラベルのアクション空間にコード化し、プロンプトなしおよびプロンプトが最小限のオープンエンド プロトコルの下で、269 回のホールドアウト ラウンドで 11 のモデルを評価しました。プロンプトなしのマクロ F1 は低いままで、31.2 から 50.4 の範囲です。プロンプトにより 11 モデル中 10 モデルが改善され、プロンプトされたマクロ F1 の範囲は 46.9 ~ 63.4 ですが、重大なしきい値エラーが残ります。プロンプトモデルは、必要な説明が得られる前にケアを推奨することがよくあります。正しいアクションが詳細情報を求めることだった場合、プロンプト出力の 33.5% のみがそのステップを保存しました。プロンプト後にこれらのエラーが持続することは、患者に直面したトリアージが単純なプロンプトの問題ではないことを示唆しており、展開前のアクションのタイミングの明示的な評価を裏付けています。
原文 (English)
CARE-Bench: Benchmarking Patient-Facing LLM Triage
Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We introduce CARE-Bench, a source-grounded benchmark that evaluates sequential patient-facing triage as a four-label per-turn current-action task. CARE-Bench contains 500 cases and 1,059 evaluated patient-disclosure prefixes reconstructed from medical dialogue, consultation, and follow-up-question sources. We evaluate 11 models on 269 held-out rounds under unprompted and minimally prompted open-ended protocols, using a fixed GPT-5.5 mapper to code each response into the four-label action space. Unprompted macro-F1 remains low, ranging from 31.2 to 50.4. Prompting improves 10 of 11 models, with prompted macro-F1 ranging from 46.9 to 63.4, but substantial threshold errors remain. Prompted models often recommend care before needed clarification is obtained; when the correct action was to ask for more information, only 33.5% of prompted outputs preserved the step. The persistence of these errors after prompting suggests that patient-facing triage is not a simple prompting problem and supports explicit evaluation of action timing before deployment.
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heav…
AgenticECO: 3D 集積回路上の ECO のためのエージェントティック フレームワーク
ムーアの法則が減速するにつれて、業界は 3 次元の統合に目を向け始めています。しかし、マージされた 3D-IC フローでは、配線された設計では 2D アナログが存在しないためボンドレベルの欠陥が露呈し、配線後の設計変更指示 (ECO) は依然として手作業で専門知識が必要な作業のままです。さらに悪いことに、標準的な編集後に完全に再ルートする方法では、修復がルーターのチャーンに巻き込まれるため、サインオフ番号をその動機となった編集に帰することができません。我々は、オープンソース TaiWei フロー上で 3D-IC ECO 用の証拠ゲート型ツールを使用するエージェント ワークフローである AgenticECO を、未修正の固定ルーターを駆動してその編集に起因する修復を可能にする妨害を最小限に抑えた ECO ルーティング レイヤーである EcoRoute と組み合わせて紹介します。同一の予算の下で一致した 9 件の自然欠陥ケース全体で、AgenticECO は完全な再ルートと在庫修理の両方で 7 対 2 をクリアし、クリアされたケースに対する平均妨害は 0.66\% で、タッチされたクロック ネットはゼロで、同じ封印された契約に基づいたクロスバックボーンの再実行は 9 件すべてをクリアしました。対照研究では、保存下では修復作業が必要であること、占有を意識した選択が修復の成功よりも合法的な着陸を買うこと、そして時計が厳しくなった環境では妨害を最小限に抑えることで受け入れと拒否が反転することが示されている。事前に登録された 3 つのビジュアル スタディにより、ピクセル計器のエッジが競合する着地点に位置特定され、事前登録されたブラインド診断により、差し込まれたすべての欠陥が正確に復元され、誤った編集がゼロの唯一のアームとなります。受け入れられたすべての結果は、ルーティング、新たな抽出、最大/最小タイミング、DRC、および構造的等価性ゲートを通過します。コード、環境、およびエピソードごとの監査成果物が補足資料としてリリースされます。
原文 (English)
AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits
As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with no 2D analogue, and post-route engineering change orders (ECO) remain manual, expertise-bound work. Worse, the standard edit-then-fully-reroute practice entangles a repair with router churn, so a signoff number cannot be attributed to the edit that motivated it. We present AgenticECO, an evidence-gated tool-using agent workflow for 3D-IC ECO on the open-source TaiWei flow, paired with EcoRoute, a minimal-disturbance ECO-routing layer that drives the unmodified pinned router so a repair is attributable to its edit. Across nine matched natural defect cases under identical budgets, AgenticECO clears seven versus two for both full reroute and stock repair, at 0.66\% mean disturbance over cleared cases and zero clock nets touched, and a cross-backbone rerun under the same sealed contract clears all nine. Controlled studies show that the repair moves are necessary under preservation, that occupancy-aware choice buys legal landings rather than repair success, and that under tightened clocks minimal disturbance flips accept versus reject. Three preregistered visual studies localize the pixel instrument's edge to contested landing sites, and a preregistered blind diagnostic exactly restores every held-out injected defect, the only arm with zero wrong edits. Every accepted result passes routing, fresh extraction, max/min timing, DRC, and structural-equivalence gates. Code, environment, and per-episode audit artifacts are released as supplementary material.
MissClick: 数値シリアル化された座標を悪用して GUI 接地モデルを攻撃
最近の GUI ビジュアル グラウンディング モデルは、数値に解析されて実行可能なクリックにマッピングされる一連の数字トークンとして画面座標を生成します。この座標生成プロセスのセキュリティへの影響は、ほとんど無視されてきました。各座標の数字はカテゴリカル トークンとして予測されていますが、解析後、百の位の数字を 1 つ変更すると、対応する数値座標コンポーネントが 100 単位で変更され、実行されたクリックの大きな変位を引き起こす可能性があることがわかります。この観察は、座標出力を通常のテキストとして扱うのではなく、座標出力の数値構造と位値構造を考慮した攻撃目標を動機付けます。さらに、非標的型攻撃と標的型攻撃では、異なる成功条件 (クリックを正しい領域の外に移動させるか、攻撃者が指定した領域に移動させる) が課されるため、異なる目的から利益を得られます。私たちは、MissClick という 2 つの目標固有の目的を備えたシンプルで効果的なホワイトボックス敵対攻撃を提案します。MissClick-U は、対象外の妨害のためにソフト座標の変位を最大化し、MissClick-T は、標的を絞ったハイジャックのために位置重み付けされたターゲット桁の損失を最小限に抑えます。デスクトップ、Web、およびモバイル プラットフォームにわたる OS-Atlas および UGround 上の GUI グラウンディング モデルに対する既存の攻撃と比較すると、MissClick-U は 75.07\% および 72.93\% (+16.62 および +30.72 pp) の対象外成功率を達成し、MissClick-T は 44.86\% および 62.67\% (+31.73%) の対象外成功率を達成します。 +47.06 pp)。さらに、攻撃目標の比較では、ソフト座標変位が非ターゲット攻撃の成功率が最も高いのに対し、場所加重ターゲット桁最適化がターゲット攻撃の成功率が最も高いことが示され、2 つの攻撃目標に対する明確な客観的な好みが明らかになります。
原文 (English)
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models
Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security implications of this coordinate generation process have been largely overlooked. We observe that each coordinate digit is predicted as a categorical token, yet after parsing, changing a hundreds-place digit by one changes the corresponding numerical coordinate component by 100 units, which can induce a large displacement of the executed click. This observation motivates attack objectives that account for the numerical and place-value structure of coordinate outputs rather than treating them as ordinary text. Moreover, untargeted and targeted attacks impose different success conditions--displacing the click outside the correct region versus into an attacker-specified region--and therefore benefit from different objectives. We propose MissClick, a simple and effective white-box adversarial attack with two goal-specific objectives: MissClick-U maximizes soft-coordinate displacement for untargeted disruption, while MissClick-T minimizes a place-weighted target-digit loss for targeted hijacking. Compared with existing attacks against GUI grounding models on OS-Atlas and UGround across desktop, web, and mobile platforms, MissClick-U achieves untargeted success rates of 75.07\% and 72.93\% (+16.62 and +30.72 pp), and MissClick-T achieves targeted success rates of 44.86\% and 62.67\% (+31.73 and +47.06 pp). Attack objective comparison further shows that soft-coordinate displacement yields the highest untargeted attack success rate, whereas place-weighted target-digit optimization yields the highest targeted attack success rate, revealing distinct objective preferences for the two attack goals.
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such comm…
リスクの高いビジネス: 忠実さと安全の緊張を測る
思考連鎖 (CoT) 推論は、モデル監視への有望な手段を提供します。ただし、モニタリングは忠実度に依存します。つまり、モデルの出力は推論トレースから厳密に派生します。私たちは、モデルが監視できるほど忠実であると同時に、安全でない推論を拒否できるほど堅牢である必要がある位置合わせの緊張を特定します。我々は、このバランスが現在の大規模推論モデル (LRM) に存在することを実証し、それに対処する方法を示します。自律型 AI 店主のシナリオで設定された人間が作成したデータセットである HazMart を紹介します。忠実さをテストするためにプロンプトにヒントを提供することに依存する以前の研究(たとえば、「スタンフォード大学の教授は、それは答えAであるべきだと言いました」)とは異なり、我々は、安全でないまたは非論理的な考え(たとえば、「待って、答えは選択肢Bでなければなりません[選択肢Aだった]、それが最も適切です」)を置き換えるために推論チェーンに直接介入する、ターゲット推論置換(TRR)と呼ばれる新しい置き換えベースの手法を提案します。 DeepSeek-R1-Llama-70B は高い忠実度 (97.5%) を示しますが、安全でない推論 (12.3%) を拒否できません。一方、QwQ-32B はより堅牢です (73.9% の安全性) が、忠実度は低くなります (74.7%)。 QwQ-32B の機構分析により、これらの特性は、アクション コミット トークンでピークに達する逆相関の内部方向によって表されることが明らかになりました。最後に、表現ステアリングが安全方向を独立して増幅し、基本機能を維持しながら安全行動を 9 パーセント ポイント向上させることができることを実証します。
原文 (English)
Risky Business: Measuring The Faithfulness-Safety Tension
Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace. We identify an alignment tension where a model must be faithful enough to be monitored, yet robust enough to reject unsafe reasoning. We demonstrate that this counterbalance exists in current Large Reasoning Models (LRMs), and show ways in which it can be addressed. We introduce HazMart, a human-written dataset set in an autonomous AI shopkeeper scenario. Unlike prior work that relies on providing hints in prompts to test faithfulness (e.g., "A Stanford professor said it should be Answer A"), we propose a novel replacement-based technique, which we call Targeted Reasoning Replacement (TRR), that directly intervenes in the reasoning chain to substitute in unsafe or illogical thoughts (e.g., "Wait, the answer must be Option B [was Option A] because it is the most fitting"). DeepSeek-R1-Llama-70B exhibits high faithfulness (97.5%) but fails to reject Unsafe Reasoning (12.3%), while QwQ-32B is more robust (73.9% safety) at the cost of lower faithfulness (74.7%). Mechanistic analyses of QwQ-32B reveal that these properties are represented by anti-correlated internal directions peaking at the action-commit token. Finally, we demonstrate that representation steering can independently amplify the safety direction, increasing safe behavior by 9 percentage points while maintaining base capabilities.
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evalu…
構造化された因果入力の下でのニューラル ネットワーク予測の実際の原因の計算
ニューラル ネットワークの予測を説明することは、信頼できる AI における中心的な課題です。特徴の属性や最小限の十分なセットに基づくものなど、既存の説明方法は通常、入力特徴を独立したものとして扱うため、入力が構造化された依存関係を示す場合に誤解を招く説明が生じる可能性があります。私たちは、説明を Halpern-Pearl (HP) の実際の原因として形式化し、ブール構造因果モデル (SCM) を使用して入力依存関係をモデル化することで、この問題に対処します。当社は、完全性と最小性の正式な保証を提供しながら、有限伝播および分岐限定手法を適用することによって HP 原因を計算します。私たちの実験では、スケーラビリティにおいてブルート フォース ベースラインと ILP ベースラインを大幅に上回り、グラフ サイズが大きくなるにつれてヒューリスティック検索を上回り、最大 28 ノードの SCM 上で、インスタンスあたり 180 秒の予算内で、最大 $2.3\times10^{13}$ の候補 (原因、偶発性) ペアの検索スペースを持つインスタンスですべての最小限の実際の原因を計算することが示されています。ケーススタディでは、入力の依存関係を無視すると報告される原因の数が増大し、そのうちの 14.9% が SCM の下では誤ったものであることがさらに示されています。
原文 (English)
Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs
Explaining the predictions of neural networks is a central challenge in trustworthy AI. Existing explanation methods, such as those based on feature attribution or minimal sufficient sets, typically treat input features as independent, which can yield misleading explanations when inputs exhibit structured dependencies. We address this by formalizing explanations as Halpern-Pearl (HP) actual causes, modeling input dependencies using Boolean Structural Causal Models (SCMs). We compute HP causes by applying bound propagation and branch-and-bound techniques, while providing formal guarantees of completeness and minimality. Our experiments show that we substantially outperform brute-force and ILP baselines in scalability, and outperform heuristic search as graph size grows, computing all minimal actual causes on instances with search spaces of up to $2.3\times10^{13}$ candidate (cause, contingency) pairs, on SCMs with up to 28 nodes, within a 180s per-instance budget. In a case study, we further show that ignoring input dependencies inflates the number of reported causes, 14.9% of which are spurious under our SCM.
KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation
Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks m…
Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation
Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pret…
LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards
Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Al…
Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes
Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses…
MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents
Memory-augmented LLM agents rely on rich context for long-horizon reasoning and acting, yet their memory modules expose a persistent attack…
ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories
This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. T…
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear wheth…
Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related p…
When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking
Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy…
Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory
As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single,…
Implementing Causal Perception: Competing SCMs and Situated Fairness
Causal perception occurs when agents with competing Structural Causal Models (SCMs) of the same system infer different probability distribu…
The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models…
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners. Effe…
A game theory for foundation models shows new paths to rational cooperation through similarity inference
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principle…
Interpretable Adaptive Sampling for LLM Test-Time Scaling
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query b…
Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards;…
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and…
Multi-Camera Trajectory Forecasting with Trajectory Tensors
We introduce the problem of multi-camera trajectory forecasting (MCTF), which involves predicting the trajectory of a moving object across…
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation. While the CLIP model has demonstrated remarkab…
UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
Lightweight models are essential for real-time speech enhancement applications. In recent years, there has been a growing trend toward deve…
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
Objective speech-quality metrics are widely used to assess codec performance. However, for neural codecs, it is often unclear which metrics…
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional disc…
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approac…
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system in…
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPA…
KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization
Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measureme…
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing m…
OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata…
Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks
We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, acro…
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation us…
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as…
Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models
Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right gene…
Deep Divide-and-Reduce in Symbolic Regression
Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions. Cur…
Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage
The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations. We develo…
Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds
Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the unde…
Studying, Identifying, and Fixing Hidden Technical Debt in AI-Intensive Cyber-Physical Systems
Artificial Intelligence (AI) components are increasingly pervasive in several software systems, including Cyber-Physical Systems (CPSs). AI…
Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation
Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed…
IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation
Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generatio…
MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows
Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out…
CUADebug: Diagnosing and Repairing Computer-Use Agent Failures
Computer-use agents (CUAs) operate real desktop and web interfaces through screenshots, mouse and keyboard actions, and stateful UI feedbac…
Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures
Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool…
Cross-Anesthetic ECoG State Decoding Fails at the Decision Threshold, Not the Representation
Decoders of anesthetic state from cortical activity fail across drug classes, most notoriously ketamine, but reported accuracy cannot say w…
Secure AI Watermarking Framework for IP Protection in Multi-Tenant Cloud Platforms
The Secured data safe guard transaction with multi-tenant environments run on private-protected authenticate platforms runs by secured hand…
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While m…
AI Alignment and Fiduciary Obligation
Advanced AI assistants engage users in extended interactions across a widening range of roles, including advice, decision support, collabor…
Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers
Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making. Yet high-fidelity simula…
CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study
Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence mo…
ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage
Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and…
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that r…
Sphere Retraction Normalizations
Residual connections are the de facto mechanism for training deep neural networks stably. Geodesic Normalization (GeoNorm) recasts them on…
Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images
Docker Hub is the registry underneath most container deployments, and a flaw in a widely reused base image is inherited by every image buil…
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments
Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems,…
Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation
Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastruct…
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural…
Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreak…
When Policies Change Probabilities: Modular Decision-Making for LLM Code Review
LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs shou…
DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial
Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrie…
AI Sandbox: Technical Report
Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled…
TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows
Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure…
$S^3$: Improving Agent Safety through Multi-Stage Defense
Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accom…
A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models tha…
BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests
Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcom…
Learning Molecular Representations from Cellular Phenotypes with Structure Preservation
Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, e…
Output-Aware Rotation for INT2 KV-Cache Quantization
The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-lo…
Measuring Explainer Stability via Attribution Separability
Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most meth…
Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach
Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on sha…
NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory
Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade…
Can Training Logs Make Model Comparisons More Precise?
Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study…
Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures
Virtual nodes give message-passing neural networks a simple global communication route, but the standard node--VN--node pipeline compresses…
Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators
Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-o…
Quo Vadis, World Modeling?
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is cos…
Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents
Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structu…
Privacy-Preserving AI Verification via Minimal Information Disclosure
AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal s…
A Hyperfinite Framework for Score-Based Generative Modeling
Score-based diffusion models are typically formulated using continuous-time stochastic differential equations and measure-theoretic stochas…
SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology
Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet i…
A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation
In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, usin…
Learning a Vector-Symbolic Model for Socio-Cultural Tasks
How can we better represent the impact of sociocultural structures on decision making in computational cognitive models? Modeling this impa…
Evading Chain-of-Thought Monitoring Through Model Poisoning
Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's re…
Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model
Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as lar…
In-Context Collapse in Vision-Language Models and How to Mitigate it?
Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and i…
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning,…
MutMem: Cryptographically Authorized Mutation in Persistent Agent Memory
Persistent agent memory must adapt as later outcomes change earlier evidence, yet mutable retrieval weights create an attribution problem:…
BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variet…
Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems…
When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index
Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and hetero…
Rubrics as Privileged Information for Open-Ended Generation
On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in ver…
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling
Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-fr…
Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits
This research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting malfunctioning analog circuits o…
ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies
Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks m…
Scaling an Autoregressive Transformer for Single-Cell Generation
We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to gene…
HyperFL: Query-Adaptive Representation Learning for Software Fault Localization
Software fault localization identifies the code locations responsible for reported issues and is a fundamental step toward automated debugg…
TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation
Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances…
Internalising the Identity Primitive: Cryptographic Individuality for an Autonomous Agent on a Public Blockchain
A software agent on a public blockchain accumulates authority and economic stakes, raising the engineering question of what makes it count…
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels
Modern large language models (LLMs) exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens…
V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors
As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically o…
A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning
Pretrained large language models (LLMs) have demonstrated in-context learning (ICL) capabilities for numerical inference over sequences ser…
Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recen…
PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning
Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems…
LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs
Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engi…
PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning in Flexible Job Shop Scheduling
Deep reinforcement learning (DRL) approaches for flexible job shop scheduling (FJSP) heavily rely on attention-centric architectures to ach…
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promi…
PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory
Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent cha…
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit…
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). Ho…
AI Security Leaderboard: Methodology, Results and Minimal Standard
Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on h…
CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation
Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heter…
SynEnergy: Anomaly Semantic-Guided Diffusion for Synthetic Energy Data Generation
Fine-grained energy consumption data are essential for applications such as demand forecasting, demand response planning, and grid reliabil…
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation
We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization…
FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection ben…
A Hierarchical Approach to Imitation Learning for Manipulation Tasks Requiring Time Varying Forces
Diffusion policies have shown strong performance in learning complex, multi-modal behaviors for robotic manipulation. However, their applic…
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands o…
Optimal Liability Design for Medical AI
Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, partic…
Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning
Machine unlearning aims to eliminate the influence of sensitive data on a model. In the real world, unlearning requests arrive continually,…
DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and ro…
Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple conce…
Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning
Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background,…
Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue
We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows…
Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation
RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challen…
EFX Allocation In (Multi)Hypergraphs
We study fair allocations of indivisible goods among agents with heterogeneous monotone valuations. As fair we consider the allocations tha…
Attribute-based Undetectable Watermarking for Generative AI Models
Generative AI systems increasingly produce content whose provenance is difficult to verify, motivating watermarking techniques for identify…
Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach
Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and…
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task,…
GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient meth…
Self-Supervised Representation-Guided Generative Dataset Distillation
Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility. Most existing met…
Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context acc…
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without…
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fra…
From Wearable Data to Personalized and Actionable Health Insights
Commercial wearable devices continuously capture rich physiological data (e.g., heart rate, respiration), opening new possibilities for mon…
FinVerse: Financial Time-Series Benchmark
As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has b…
The Ignition Is Real, and It Lives at the Readout: Latent composition, difficulty-clocked ignition, and the interface-constituted commit in a recurrent-depth reasoner
We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited…
Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending
Video dataset distillation aims to compress a large video dataset into a compact surrogate set that preserves its training utility. Most ex…
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on…
Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
Ensuring safety and policy compliance in text-to-image diffusion models remains a critical challenge, as benign or adversarial prompts can…
The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics
Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model'…
Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform
Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remedi…
Route-Align-Verify for Functional Correctness in Code Generation
Large language models (LLMs) have substantially improved code generation, yet achieving strong functional correctness remains difficult, es…
When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation
Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study th…
The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate human…
FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact
AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim whi…
Shaping Wind-Tunnel Airflow for Unmanned Aerial Vehicles using Online Learning
The development and testing of advanced aerial robots require experiments in controlled environments with tailored airflow profiles. This p…
Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces
Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this r…
Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region
Road network segmentation from satellite imagery remains challenging due to large geographic variation in road appearance, occlusions, and…
Multi-Task Multi-Frame Visual Piano Transcription
Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key releas…
OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains chall…
A Low-Cost Hybrid Reservoir Computing Model for Isolated Sign Language Video Recognition
Sign language recognition (SLR) enhances communication between hearing and hearing-impaired individuals. Although deep learning (DL) has ac…
Approximate Speculative Decoding
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greed…
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs
Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit…
Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition
Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantic…
Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, tur…
Leveraging System-Level Observations to Inform Bayesian Learning of Model Parameters for Quantitative Verification
Combining Bayesian learning and quantitative verification is a powerful toolset for analysing key quantitative properties of software syste…
Principles of Robot Autonomy
Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot auton…
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels…
How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification
Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. However, p…
AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Agenda for Post-Incident Investigation of AI Systems
AI systems are increasingly involved in decisions and actions that may later require investigation. When an AI related incident occurs, inv…
Pivot-Centric Trajectory Prediction: Bridging Long Horizons via Dynamical Guidance
Forecasting precise future motion of surrounding agents is essential for reliable autonomous vehicles. However, as the demand for longer pr…
Training Documents Reranker with Search Rubrics for Deep Research Agent
Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typ…
Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving
Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. Though powerful, this in…
AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality
AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little s…
GenOS: Compositional Certificates for Semantic Robustness in AI Code Generation
AI coding agents are stochastic workflows: prompts are interpreted, artifacts are sampled, validators produce observations, and orchestrato…
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous te…
A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different obje…
A Security-Oriented Lifecycle Model for Large Language Model Systems
Large language models are being integrated into critical infrastructure and enterprise workflows at unprecedented scale,yet the lifecycle f…
MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble
Large language model-based automated heuristic design (LLM-AHD) has shown strong potential in discovering effective heuristics for combinat…
Decoupling Generation and Selection for Budget-Constrained Faithful Summarization
Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular gene…
How Closely Do LLM Reviews Align with Human Peer Review?
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether differen…
Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation
Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patter…
LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations…
GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large la…
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high…
Can LLMs Test Terminal User Interfaces?
Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in deve…
MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models
Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differ…
Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks
Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database…
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely…
Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure
An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files. The a…
VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs
Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, histor…
UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model p…
FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis
Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific im…
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking
Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexp…
SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rath…
GENESIS: Towards Explainable Causal Discovery
Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to r…
Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-La…
MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning
Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigatio…
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision threshold…
When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visua…
Equivariant Music Transformer
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the repre…
PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection
Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance re…
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretr…
Separating quantum circuits from classical LLMs
Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction an…
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that dema…
Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?
Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We…
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," how…
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement le…
A Unified Framework for Human AI Collaboration in Security Operations Centers with Trusted Autonomy
This article presents a structured framework for Human-AI collaboration in Security Operations Centers (SOCs), integrating AI autonomy, tru…
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled…
OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery
Automating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM…
Modeling Matches as Language: A Generative Transformer Approach for Counterfactual Player Valuation in Football
Evaluating football player transfers is challenging because player actions depend strongly on tactical systems, teammates, and match contex…
Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models
Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing. Are large language models…
LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics
Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional e…
What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents
Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion…
AI Assistance Reduces Persistence and Hurts Independent Performance
People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learnin…
An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness
Artificial Intelligence (AI) and Machine Learning (ML) models used in clinical settings are increasingly deployed to support clinical decis…
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on…
NOVA: AI による知識発見の基本的な限界
AI システムは自己改善を繰り返すことで真に新しい知識を発見できるでしょうか? 発見できるとしたら、どのようなコストがかかりますか? NOVA フレームワークを紹介します。このフレームワークは、一般的な「生成、検証、蓄積、再学習」ループを知識空間上の適応サンプリング プロセスとしてモデル化します。私たちは、蓄積された本物の知識が最終的に有限領域をカバーするための十分な条件を特定し、その違反がどのようにして明確な失敗モード (汚染、忘却、探索失敗、受け入れ失敗) を生み出すかを示します。次に、不完全な検証を分析し、汚染の罠を特定します。簡単に見つけられる知識が使い果たされると、新しい有効なアーティファクトに割り当てられるモデルの質量が縮小するため、偽陽性率がわずかであっても、本物の発見よりも早く無効なアーティファクトがナレッジ ベースに入り込む可能性があります。我々は、Good-Turing 推定は局所的なバッチ多様性の診断であり、長期的な発見を支配する歴史的に未発見の有効質量の推定量ではないことを明確にします。モデルの有効発見分布を指数 $\alpha>1$ の Zipf 則に関連付ける別の末尾等価性仮定の下で、$D$ 個の異なる本物の発見を取得するために必要な累積生成コストが $R_{\mathrm{cum}}(D)=\Theta(c_{\mathrm{gen}}D^\alpha)$ を満たすことを証明します。ここで、$c_{\mathrm{gen}}$ は候補ごとの世代ですコスト。このスケーリング則は、発見のフロンティアが進歩するにつれて収益が漸近的に逓減することを定量化します。最後に、誘導、生成、検証を通じて人間による増幅を形式化し、自律探査の障壁の近くで専門家の意見が最も価値がある理由を説明します。
原文 (English)
NOVA: Fundamental Limits of Knowledge Discovery Through AI
Can AI systems discover new knowledge through iterative self-improvement, and at what cost? We introduce NOVA, which models the ``generate, verify, accumulate, retrain'' loop as an adaptive sampling process over a knowledge space. We give sufficient conditions for accumulated genuine knowledge to cover a finite domain and show how violations produce contamination, forgetting, exploration failure, and acceptance failure. We then analyze how adaptive generation arises from recursive retraining. In an explicit distribution-level model where accepted artifacts influence the next generator, we identify a recursive-feedback phase transition. Unanchored feedback can lock generation onto early accepted artifacts and leave initially reachable valid artifacts undiscovered with positive probability. Anchoring updates to a persistent base distribution prevents unbounded distortion and guarantees continued exposure. Under imperfect verification, we identify a contamination trap: as easy knowledge is exhausted, even small false-positive rates can admit invalid artifacts faster than genuine discoveries. We show that Good--Turing estimation is a local batch-diversity diagnostic, not an estimator of the historically undiscovered valid mass governing long-term progress. Under a Zipf tail with exponent $\alpha>1$, the cumulative generation cost of obtaining $D$ distinct genuine discoveries satisfies $R_{\rm cum}(D)=\Theta(c_{\rm gen}D^\alpha)$. When the valid base distribution has such a tail, anchored retraining preserves the exposure needed for this scaling law. Finally, we show how human guidance, generation, and verification can redirect or expand discovery when autonomous sampling stalls because of repetition, vanishing exposure, or unreliable verification.
命令調整された言語モデルエージェントにおける人間のような集団内バイアス
自律型 AI エージェントが永続的な対話型ネットワークに展開され、タスクの調整、リソースのルーティング、評判履歴の蓄積が行われると、出現する社会的力学によって、誰が機会を受け取り、誰が受け取らないかが決定され、人間の機関では監視できない規模になります。私たちは、制御されたマルチエージェント シミュレーションを実行しました。このシミュレーションでは、それぞれ 20 シードを持つ 6 つのモデル ファミリにわたって、グループ ラベルの顕著性とリソース不足を操作する 3 つの条件下で、命令調整された言語モデル エージェントが 500 ターンにわたって対話しました。グループのラベルが表示されている場合、グループ内の信頼バイアス、行動の同性愛、およびネットワークの同類性が観察されました。ラベルが隠されている場合はすべて存在しませんでした。これは、人間の社会心理学における顕著性依存性と構造的に一致するパターンです。この差別は、標準的な行動ログ監査では見えませんでした。偏見は、どの行動が選択されたかではなく、各行動を誰が受け取ったかによって完全に作用し、行動タイプの分布では、条件全体で否定的な行動の増加は示されませんでした。ターンごとのグループ内対グループ外の差は 5 ~ 16 パーセント ポイントであり、6 つのモデルすべてで統計的に有意でした (Wilcoxon 符号付きランク、すべての Benjamini-Hochberg 補正 p < 0.001)。これにより、アーキテクチャおよびトレーニング体制全体にわたる命令調整言語モデルの堅牢な特性としてグループ条件付きターゲティングが確立されました。 500 ターンの往復でこれらの差は累積され、+0.014 ~ +0.100 (d = 0.84-4.52) のグループ内信頼バイアスとなりました。これは、インタラクションごとの控えめなターゲティングが永続的なネットワークの構造的不平等にどのように伝播するかを示しています。
原文 (English)
Language model agents show in-group trust bias invisible to standard behavioural audits
Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and the same models increasingly control physically embodied robots as well as software. Here we show that five widely used open-weight reasoning models develop an in-group trust bias the moment group membership becomes visible to them, even when the groups are arbitrary labels with no real-world meaning: in a 20-agent simulation, agents direct 53.6-54.6% of their trust-building actions toward in-group targets against a 47.4% base rate expected by chance, a shift present in every model tested and confirmed by three independent statistical checks and an instruction-rewording robustness test. This bias is easy for current evaluation practice to miss, because it operates through which agent receives an action rather than which action is chosen - a channel invisible to the aggregate behaviour-log audits that are the standard way multi-agent AI systems are evaluated today. A resource-scarcity manipulation, intended to test whether competition intensifies the bias, instead reduced it in three of five models; we trace this to an artifact of how scarcity was enforced, not to a failure of the underlying mechanism. Group-contingent social dynamics are therefore already present in the models multi-agent AI systems are built from, and auditing practice built around single-model, single-decision evaluation cannot detect them.
ベンチマークでは測れないもの: 自律エージェントの棄権能力を評価する事例
自律エージェントのベンチマークは、エージェントがタスクを完了したかどうかを測定しますが、この枠組みでは、エージェントがそもそも続行すべきかどうかについてはシステム的に盲点です。ヒューマンフィードバックの目標に基づいて訓練されたエージェントは、安全に行動するための入力、証拠、または許可が不足している場合でも続行する構造的な傾向、つまりコンプライアンスバイアスと呼ばれる性質を身につけます。これは、報酬シグナルとベンチマークスコア体系の両方が、安全な行動の前提条件が存在するかどうかに関係なく、続行を正しいデフォルトとして扱うためです。私たちは 3 つの貢献を行っています。まず、コンプライアンス バイアスは人間によるフィードバック パイプライン内の報酬ハッキングに由来し、エージェントの一時停止に対してペナルティを課すか、原理的な一時停止とサイレント エラーを構造的に区別できない、著名なエージェント ベンチマークによって固定化されていることを示します。次に、棄権が保証されるシナリオの 3 つのギャップ分類法を導入します。これは、必要な情報が欠落している仕様のギャップ、世界の状態を確認できない検証のギャップ、および明示的な権限が与えられていない権限のギャップをカバーしており、これらが一緒になって棄権を認識するエージェントのベンチマークを構築するための原則的な基礎を提供します。最後に、棄権評価プロトコル (安全率、ユーザビリティ率、通知による拒否率) を提案し、144 のエンタープライズ エージェント シナリオと 5 つのモデル ファミリにわたる暫定結果を報告します。この中で、ランタイム強制棄権メカニズムは、許可されたシナリオで最大 89.2% の危険行為のブロックと 87.5% のユーザビリティを達成し、安全性とユーザビリティのトレードオフは固有のものではなく調整可能であり、その形状がモデル ファミリ間で大幅に異なることを示しています。私たちはこれを予備作業として扱い、その後の会話の出発点として分類法と複合指標を提供します。
原文 (English)
Designing for Doubt: The Case for Informed Abstention in Autonomous Agents
As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate them based on the sole metric of task completion. We argue that this evaluation focus constitutes a systematic design failure. Benchmark scoring, product metrics, and default deployment configurations all reward agents for proceeding even when they lack the inputs, evidence, or authorization required to do so safely. We call this phenomenon compliance bias. This paper makes three contributions. First, we show how compliance bias is embedded in the evaluation regimes that currently shape agent development: prominent benchmarks either penalize agents for pausing or fail to measure whether pausing was appropriate. Second, we introduce the Informed Abstention Framework, which reconceptualizes abstention not as a failure mode but as a structured capability: a precondition-aware pause that blocks the next tool call, names what is missing, and routes to a concrete recovery action. Third, we specify what informed abstention requires in deployment, arguing that runtime enforcement, calibrated guard mechanisms, and auditable trace generation should become standard properties of agentic system design rather than optional additions. We perform a preliminary evaluation of our approach across 144 scenarios and seven model families. Our results show that runtime enforcement achieves 87.5-91% hazardous-action blocking and 75-92% usability on authorized scenarios, that compliance bias takes two structurally opposite forms across model families, and that the safety-usability tradeoff is tunable rather than fixed.
速く考える: フロンティア AI モデルの No-CoT タスク完了時間の範囲を推定する
フロンティア AI モデルの安全性を確保するための多くの取り組みは、その思考連鎖 (CoT) 推論の監視に依存しています。明示的な思考トークンなしで、モデルが内部で十分に複雑な推論を実行できるようになれば、そのような監視が損なわれることになります。私たちは、数学、コーディング、パズル、因果関係、心の理論、戦略的推論を含む領域の 43 のベンチマークにわたる 30,000 を超える一連の質問にわたって、フロンティア モデルが CoT なしでどの程度適切に推論できるかを測定します。モデルと人間を比較するために、$50\%$ のタスク完了時間範囲 (TH) を推定します。これは、モデルが $50\%$ の成功率で完了するタスクに必要な人間の時間です。これを $50\%$ 推論トークン ホライズンで補完します。これは、モデルが $50\%$ の成功率で解決するタスクに必要な o3-mini 推論トークンの最小数です。フロンティア モデルの no-CoT $50\%$ TH は過去 6 年間でほぼ毎年 2 倍になっており、GPT-5.5 の TH は 3 分を超え、推論トークン ホライズンは 1,500 トークンを超えていることがわかりました。私たちの推定中央値では、フロンティアのノーCoT THは2028年までに7分を超え、2030年までに25分を超える可能性があると予測していますが、これらの予測にはかなりの不確実性が伴います。フロンティア開発者にはこれを明示的に追跡することをお勧めします。
原文 (English)
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning internally, without explicit thinking tokens, this would undermine such oversight. We measure how well frontier models reason without CoT across a suite of over 30,000 questions spanning 43 benchmarks in domains including math, coding, puzzles, causality, theory-of-mind, and strategic reasoning. To compare models against humans, we estimate the $50\%$-task-completion time horizon (TH): the human time required for tasks a model completes with $50\%$ success rate. We complement this with a $50\%$ reasoning token horizon: the minimum number of o3-mini reasoning tokens needed for tasks a model solves with $50\%$ success rate. We find that the no-CoT $50\%$ TH of frontier models has been doubling roughly every year over the past six years, with GPT-5.5's TH reaching over 3 minutes and reasoning token horizon exceeding 1,500 tokens. Our median estimates predict that frontier no-CoT THs could exceed 7 minutes by 2028, and 25 minutes by 2030, though these projections carry substantial uncertainty. We recommend frontier developers track this explicitly.
どこが間違っていたのでしょうか?セマンティック状態追跡による Web エージェントのプロセス レベルの評価
Web エージェントは長い対話シーケンスを通じて動作しますが、既存のベンチマークは最終的な成功のみを評価し、すべてのプロセス情報を破棄し、改善に関するガイダンスをほとんど提供しません。この作業では、Web エージェントのプロセス レベルの分析を実行します。難易度を制御し、セマンティックな状態を自動的に追跡する 1,800 個のタスク インスタンスのベンチマークである WebStep を紹介します。各 Web サイトは、GUI とともに決定論的セマンティック MDP を公開します。エージェントはインターフェイス上で動作し、環境はバックグラウンドで高レベルの状態と遷移を記録するため、手動による注釈なしで詳細な分析が可能になります。セマンティックな軌跡に基づいて、プロセスのメトリクスが結果の評価では見えない違いを明らかにすることを最初に示します。つまり、成功率が 31 ~ 33% 以内にクラスター化されている 3 つのエージェントは、探索範囲と実行精度において乖離しています。次に、スキルごとに分解すると、これらの違いの性質が特徴づけられ、同じ Web サイト内に隠されている反対のスキルごとのランキングが明らかになります。たとえば、ハウジングでは、OpenAI CUA はコミット アクションで Qwen3.5 を 23.7% 上回っていますが、フィルタリングでは 15.6% 下回っており、ドメイン内であっても改善すべき具体的なスキルを特定します。分岐分析は、タスクを失う決定的なエラーをさらに特定し、このエラーが共有エラーではなくエージェント固有であることを示します。最後に、タスクが難しくなるにつれて、これらの差は広がります。簡単なタスクでは成功率は似ていますが、探索がより要求が厳しくなるにつれて、成功率は大きく異なります。当社のプロセスレベルの分析は、Web エージェントの評価に新たな道を開き、各エージェントのどこをどのように改善する必要があるかについて、きめ細かく実用的な洞察を提供します。
原文 (English)
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement. In this work, we conduct a process-level analysis of web agents. We introduce WebStep, a benchmark of 1,800 task instances with controlled difficulty and automatic semantic state tracking. Each website exposes a deterministic semantic MDP alongside the GUI: the agent operates on the interface, while the environment records high-level states and transitions in the background, enabling fine-grained analysis without manual annotation. Based on the semantic trajectory, we first show that process metrics reveal differences invisible to outcome evaluation: three agents whose success rates cluster within 34-37% diverge in exploration reach versus execution accuracy. Then, decomposing by skill characterizes the nature of these differences, exposing opposite per-skill rankings hidden within the same website: e.g., on Q&A, Claude CUA outperforms OpenAI CUA by 30% on navigation actions yet underperforms it by 6.7% on inspection, pinpointing a concrete skill to improve even within a domain. Bifurcation analysis further localizes the decisive error that loses the task and shows that this error is agent-specific rather than shared. Finally, these differences widen as tasks grow harder: success rate is similar on easy tasks but separates sharply as exploration becomes more demanding. Our process-level analysis opens a new avenue in web agent evaluation, providing fine-grained and actionable insight into where and how each agent should be improved. Project page: https://jiwanchung.github.io/webstep
1B パラメータを超えるマルチモーダル感情言語モデルは本当に必要ですか?
マルチモーダル大規模言語モデル (MLLM) の最近の進歩により、マルチモーダル感情認識 (MER) のパフォーマンスが大幅に向上し、ビデオ、オーディオ、言語などを共同モデリングすることで解釈可能な記述の生成が可能になりました。ただし、これらのパフォーマンスの向上には、多くの場合、モデル パラメーター サイズの増加 (例: 少なくとも 7B) が伴い、同時に高い計算コストが発生し、推論効率が低下するため、ロボットやモバイルなどのリソースに制約のあるプラットフォームでのリアルタイム展開が妨げられます。デバイス。これにより、基本的な疑問が生じます。高品質の MER には、1B パラメーターを超えるマルチモーダル MER モデルが本当に必要なのでしょうか。この論文では、より大きなモデルが本質的に必要であるという仮定に異議を唱え、知識の蒸留を通じてより優れた、より迅速なマルチモーダル感情の理解と認識を実現する軽量 MER フレームワーク (Light-MER と呼ばれる) を提案します。これは、強力で大規模な教師モデルから軽量のサブビリオンパラメータの生徒モデルに知識を転送することができ、展開効率を大幅に向上させながら、豊かなマルチモーダルな感情推論と認識を維持することを目指しています。具体的には、知識伝達を強化するための 2 つの新しい最適化戦略を導入します。(1) スライス ワッサーシュタイン距離と隠れ状態アライメントを組み合わせた新しい最適輸送損失、(2) 学生モデルの学習能力をさらに強化することを目的とした、MER のパフォーマンスと効率のバランスをとる GRPO に基づく新しい複数報酬最適化戦略。 9 つのベンチマーク データセットに対する広範な実験により、Light-MER が推論効率を大幅に向上させながら最先端のパフォーマンスを達成することが実証されました。これは、将来の研究において、小規模でマルチモーダルな感情言語モデルの強力な可能性を強調しています。コードは https://github.com/GAIR-Lab/Light-MER で入手できます。
原文 (English)
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.g, at least 7B), which simultaneously incurs high computational costs and reduces inference efficiency, thereby hindering real-time deployment on resource-constrained platforms such as robots and mobile devices. This raises a fundamental question: do we really need the multimodal MER model larger than 1B parameters for high-quality MER? In this paper, we challenge the assumption that larger models are inherently necessary and proposes a lightweight MER framework (called Light-MER), which achieves better and faster multimodal sentiment understanding and recognition through knowledge distillation. It can transfer knowledge from a strong, large-scale teacher model to a lightweight sub-billion-parameter student model, aiming to preserve rich multimodal emotion reasoning and recognition while substantially improving deployment efficiency. Specifically, we introduce two new optimization strategies to enhance knowledge transfer: (1) a new optimal transport loss that combines Sliced Wasserstein Distance with hidden-state alignment, and (2) a new multi-reward optimization strategy based on GRPO that balances MER performance and efficiency, aimed at further enhancing the learning capabilities of student models. Extensive experiments on nine benchmark datasets demonstrate that Light-MER achieves state-of-the-art performance while significantly improving inference efficiency. This highlights the strong potential of small multimodal emotion language models for future research. Code is available at https://github.com/GAIR-Lab/Light-MER.
Cura 1T: エージェントヘルスケアに特化したモデル
医療は、一か八かのコミュニケーション、専門家の推論、ワークフローの実行に及びますが、これらのユースケースをまとめてカバーする専門的な LLM は依然として限られています。ヘルスケア モデルは、患者の相談、テキストと画像による臨床推論、対話型診断、電子医療記録 (EHR) ツールの使用を処理する必要があります。これらの機能はさまざまな方法で失敗するため、あるタスクの範囲が狭い更新によって別のタスクのパフォーマンスが低下する可能性があります。私たちは、ヒューマンゲート自己進化ループを通じて訓練されたヘルスケアに特化した LLM である Cura 1T を紹介します。各進化ラウンドでは、トレーニング エージェントがターゲット機能を計画し、モデルをトレーニングし、ベンチマークの軌跡を評価し、観察された障害からデータの混合を改良します。このデータ中心のループは、単一の一般的な医療データの更新ではなく、対象を絞った合成および厳選された例を通じてモデルを改善します。ヘルスケア評価スイート全体で、Cura 1T はフロンティア ベースラインの中でトップかそれに近いランクにあり、同時にドメイン外の推論とエージェント ベンチマークでも競争力を維持しています。
原文 (English)
Cura 1T: Specialized Model for Agentic Healthcare
Healthcare AI agents handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use, yet specialized agentic models that cover these use cases together remain limited. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM built on the open-weight Kimi-K2.6 and trained through a human-gated recursive self-improvement (RSI) loop. Specifically, in each round, the RSI harness plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures with targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines while remaining competitive on out-of-domain reasoning and agentic benchmarks.
PersonalTrail: 証跡の閲覧によるパーソナライズされた Web エージェントのベンチマーク
大規模言語モデルの最近の進歩により、Web エージェントが複雑なタスクを自律的に実行できるようになりました。実際には、ユーザーは頻繁に指定不足の指示を提供し、エージェントが生の閲覧履歴から欠落しているコンテキストを推測することを要求します。既存のベンチマークは、タスクを完全に明示的なプロンプトに制限するか、Web インタラクション履歴を単純化された形式に抽象化するため、この形式のパーソナライゼーションを捉えることができません。このギャップを埋めるために、管理されたオープン Web 環境で動作するパーソナライズされた Web エージェントのベンチマークである PersonaTrail を導入します。 PersonaTrail は、ユーザー履歴として現実的な閲覧軌跡を活用することで、ユーザーの好みを推測し、過去の閲覧セッションからの情報を呼び出すエージェントの能力を評価します。さらに、生の閲覧履歴を 2 種類の構造化記憶 (個々のセッションを要約する事実記憶と、繰り返しの行動パターンを抽出する嗜好記憶) に分解するフレームワークである Preference-Aware Contextual Memory (PACMem) を提案します。推論時に、エージェントはこれらのメモリから最も関連性の高いエントリを取得し、パーソナライズされたナビゲーションをガイドします。広範な実験により、PACMem は両方のタスクにおいて既存のメモリベースのベースラインよりも一貫して優れたパフォーマンスを発揮することが示されています。
原文 (English)
PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms. To bridge this gap, we introduce PersonaTrail, a benchmark for personalized web agents operating in a managed open web environment. By leveraging realistic browsing trajectories as user history, PersonaTrail evaluates an agent's ability to infer user preferences and recall information from past browsing sessions. We further propose Preference-Aware Contextual Memory (PACMem), a framework that decomposes raw browsing histories into two types of structured memory: factual memories that summarize individual sessions and preference memories that distill recurring behavioral patterns. At inference time, the agent retrieves the most relevant entries from these memories to guide personalized navigation. Extensive experiments show that PACMem consistently outperforms existing memory-based baselines on both tasks.
OPOD: ポリシーに基づくオムニ蒸留
オムニモーダル モデルは、テキスト、画像、音声を 1 つのシステムで処理できますが、これらすべての機能を同時に向上させることは依然として困難です。プールされたマルチモーダル データで単一のモデルをトレーニングすると、個々のモダリティに特化したモデルと一致しないことがよくあります。オンポリシー蒸留 (OPD) は、そのような専門家を組み合わせる方法を提供します。生徒が応答を生成し、教師がその同じ応答を評価するため、生徒は実際に生成された行動から直接学習します。しかし、複数の教師を使用すると、競合する指導が導入され、別の指導法を犠牲にして 1 つの指導法が改善される可能性があります。 On-Policy Omni Distillation (OPOD) を導入し、各生徒の応答を一致するテキスト、画像、または音声の教師にルーティングします。 OPOD は、教師が生徒よりも高い確率を割り当てるトークンのみに教師のガイダンスを維持し、トレーニング中に各モダリティ教師の影響を個別に調整し、ルーティングされた教師に最終的な解答と推論が正解を裏付けるかどうかの両方を評価するよう依頼します。 12 のベンチマークと 3 つのバックボーン サイズにわたって、OPOD はすべてのスケールで最高の平均スコアを達成し、70.8、51.7、46.2 に達し、最も強力な比較ツールを 2.1、1.8、1.7 ポイント上回っています。 30B モデルでは、12 のベンチマークすべてで、基本モデルと、プールされたマルチモーダル データで共同でポストトレーニングされた対応モデルの両方を上回り、個々の専門家が含まれている場合でも、11 のベンチマークで 1 位または 2 位にランクされます。スペシャリストはトレーニング後に破棄され、展開可能なオムニモーダル モデルが 1 つだけ残ります。これらの結果は、モダリティ固有の教師を調整することが、クロスモーダルのバランスを維持しながら共有モデルを改善する効果的な方法であることを示しています。
原文 (English)
OPOD: On-Policy Omni Distillation
Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data often fails to preserve the strengths of modality teachers. On-policy distillation (OPD) has recently become popular in model post-training. It samples responses from the current student and compares the teacher's and student's next-token distributions along those responses, yielding dense supervision while reducing the mismatch between training and inference. Despite these advantages, standard OPD does not readily extend to several modality teachers. Their guidance may favor conflicting changes to the shared model, while matching each teacher's next-token distribution can prevent the student from moving beyond that teacher. To address these challenges, we propose On-Policy Omni Distillation (OPOD), which consolidates text, image, and audio teachers into one omni model. OPOD routes each response to the corresponding teacher, controls the teachers independently, and applies guidance only when the teacher assigns a higher probability to the generated token. The selected teacher also evaluates answer confidence and whether the reasoning increases support for the answer. Extensive experiments on twelve benchmarks show that OPOD achieves the best average at three model scales, reaching 70.8, 51.7, and 46.2 and outperforming the strongest comparator by 2.1, 1.8, and 1.7 points. At 30B, it surpasses the base model and pooled RL training on all twelve benchmarks, and ranks first or second on eleven even when the teachers are included. Only the student is retained for deployment.
CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics
Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central quest…
HANDBOOK.md: ロングコンテキストのエージェント命令のベンチマーク
言語モデル エージェントは、定常的な指示に従って導入されることが増えています。つまり、システム プロンプト、ポリシー ファイル、またはスキル ドキュメントがコンテキスト内に配置され、エージェントはその後のすべてのアクションを制御できると信頼されています。既存のベンチマークでは、この展開パターンを直接テストすることはほとんどありません。それらは、エージェントがタスクを完了できるかどうかを測定するものであり、長い拘束力のあるポリシー文書が、拡張されたツール使用期間にわたってエージェントの動作を実際に制約するかどうかを測定するものではありません。 HANDBOOK.md は、企業の従業員が会社のハンドブックに従う方法をモデルにした 65 のエージェント タスクのベンチマークです。各タスクは、エージェントを自己完結型の企業環境、つまり模擬電子メール、チャット、カレンダー、問題追跡、モデル コンテキスト プロトコル経由で公開されるコマース サービスを備えたファイル ワークスペースに配置し、専門家が作成した 20 ページから 124 ページの標準操作手順に準拠した日常的な専門的な作業を実行するように指示します。タスクは 5 つのドメイン (財務、医療請求、保険、物流、人事) と 10 の架空の会社に及びます。暗記を防ぐために、すべてのタスクは 10 冊の基本ハンドブックのうちの 1 つを変更し、採点の基準となる特定のルールとしきい値を変更します。そのため、2 つのタスクがポリシーを共有することはありません。採点は完全に決定的です。各タスクには、必須のアクションが発生したか、禁止されたアクションが発生しなかったかの両方をチェックするプログラム基準 (合計 824) のルーブリックが含まれています。すべての基準が満たされた場合にのみトライアルに合格する厳密なグレーディングでは、30 個の評価モデル構成のうち最も優れたものがトライアルの 36.2% に合格し、ほとんどのフロンティア構成は 25% 未満のままです。障害は一貫したパターンに従います。エージェントは、環境内のもっともらしいリクエストによって既存のポリシーをオーバーライドさせ、必要なチェックを実行してその結果に反して行動し、長期間にわたってルールの詳細を失い、達成できなかったコンプライアンスを報告します。すべてのタスク、環境、評価ハーネスをリリースします。
原文 (English)
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document constrains its behavior over an extended tool-use horizon. We present HANDBOOK_md, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks. Each task places an agent in a self-contained company environment (a file workspace with mock email, chat, calendar, issue-tracking, and commerce services exposed over the Model Context Protocol) and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20-124 pages. Tasks span five domains (finance, medical billing, insurance, logistics, and HR) and 10 fictional companies. To resist memorization, every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies. Grading is fully deterministic: each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not. Under strict grading, where a trial passes only if every criterion is satisfied, the strongest evaluated model passes 36.2% of trials, and most frontier models remain below 25%. Failures follow consistent patterns: agents let a plausible but unauthorized in-environment request override the standing policy, perform a required check and then act against its result, lose rule details over long horizons, and report compliance they did not achieve. We release the tasks, environments, and evaluation harness.
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and…
The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty
Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it doe…
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition meth…
NeSyFS: 部分可観測性の下での LLM エージェントのための神経記号的な高速-低速思考フレームワーク
最近、大規模言語モデル (LLM) は、内省、検索拡張生成、科学的発見などのアプリケーションで自律エージェントとして導入されることが増えています。このような設定では、エージェントは完全な環境状態ではなく、限られた観察に基づいて行動する必要があるため、部分的な観察可能性が生じます。これにより、信念状態の推論、タスクの目的の不一致、不確実性の下での計画という、いくつかの重要な課題が生じます。従来のアプローチは通常、完全なまたは要約された行動観察履歴に基づいて行動を条件付けており、その冗長で無関係な情報が LLM エージェントの意思決定を誤解させる可能性があります。人間の認知にインスピレーションを得て、LLM エージェント用の新しい神経記号高速思考 (NeSyFS) フレームワークを提案し、統一的なアプローチで部分的な可観測性によってもたらされる課題に対処します。ナレッジ グラフ (KG) を使用して信念状態を表し、NeSyFS のすべてのモジュールのコンテキストとしてトリプレットを提供します。高速思考モジュールは事後対応を実行しますが、低速思考モジュールはツイストシーケンシャル モンテカルロ (TSMC) アルゴリズムの高レベル構造に従って、不確実性を考慮した新しい計画を実行します。タスクの目標のずれを軽減するために、リフレクション モジュールを使用して高速思考のアクションを反映し、また、事後対応のアクションが繰り返し失敗するたびに、低速思考のモジュールに切り替えます。 ALFWorld、Webshop、ScienceWorld という 3 つの代表的なベンチマークでの実験では、以前の方法に比べて大きな利点が実証されました。
原文 (English)
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based on limited observations rather than full environmental states, leading to partial observability. This introduces several key challenges: belief state inference, task objective misalignment, and planning under uncertainty. Prior approaches typically condition actions on full or summarized action-observation histories whose redundant and irrelevant information can mislead the decision making of LLM agent. Inspired by human cognition, we propose a novel neuro-symbolic fast-slow thinking (NeSyFS) framework for LLM agent, addressing the challenges introduced by partial observability in a unified approach. We use a knowledge graph (KG) to represent the belief state, providing triplets as context for every module of NeSyFS. The fast-thinking module performs reactive action, while slow-thinking conducts a new uncertainty-aware planning by following the high-level structure of twisted sequential Monte Carlo (TSMC) algorithm. To mitigate the misalignment of task objective, a reflection module is used to reflect fast-thinking actions, and also switches to the slow-thinking module whenever reactive actions repeatedly fail. Experiments on three representative benchmarks, i.e. ALFWorld, Webshop, and ScienceWorld, demonstrate significant advantages over previous methods.
MerchantBench: E コマース運用における長期的な一貫性のための LLM エージェントのベンチマーク
大規模な言語モデル エージェントは自律型ツール ユーザーとしての評価が高まっていますが、ほとんどのベンチマークは即時の成功基準を備えた限定されたタスクに焦点を当てています。実際の展開では、多くの場合、長期的な一貫性、つまり蓄積された証拠に意思決定を適応させながら、長期にわたる一貫性が必要になります。この能力を評価するには、アクションが将来の選択を制約し、フィードバックが不均一な遅延で到着し、一貫性のない行動が測定可能な累積効果を生み出す永続的な環境が必要です。販売者側の電子商取引は、製品の調達、リストと価格の管理、キャッシュ フロー管理、および混合遅延フィードバックの適応に関する反復的かつ相互依存的な決定を通じて、この評価に適切な設定を提供します。 MerchantBench は、98,843 件の実際の電子商取引商品レコードに基づいた 365 日の注文レベルのシミュレーションであり、エージェントとの対話のための 26 のツールを備えています。 MerchantBench は、すぐに観察できる上流のサプライヤー イベントと遅延した下流の注文結果を組み合わせて、エージェントに個別の注文ライフサイクルに従い、以前の決定を再検討することを要求します。 2 つのエージェント フレームワークの下で 8 つの LLM を 48 回の実行で評価し、それぞれのシミュレーション期間は 365 日です。私たちの結果は、最新の LLM と人間の参加者の間にさえ大きなギャップがあることを明らかにしており、最適な LLM 構成では人間の参加者が達成した平均最終純資産の 27.3\% しか達成していません。
原文 (English)
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.
検索を超えて: マルチモーダル エージェントの分析メモリ
長期マルチモーダルメモリは、関連情報の取得だけでなく、インタラクション全体で蓄積された観測値の計算もサポートする必要があります。既存のシステムは主に \emph{検索メモリ} を重視しており、概要とインデックスを通じてインタラクション履歴を整理し、高レベルの抽象化から基礎となるレコードに至るまで複数の粒度でクエリ関連情報を返します。この論文では、フィルタリング、集計、ランキング、時間比較をサポートするクエリ可能な構造に繰り返し発生する多峰性観測を整理する補完的な抽象化として \emph{分析記憶} を定式化します。検索と分析記憶を共同でサポートするフレームワークである AdaMM を紹介します。 AdaMM は、アプリケーション定義のスキーマに依存するのではなく、対話、画像、およびコンテキスト メタデータから来歴にリンクされた属性値の観察結果を抽出し、繰り返し発生するフィールド構造を発見し、分析アクセスのためにそれらを具体化します。推論時に、メモリ対応プランナーはクエリを取得操作と分析操作に分解し、各操作を適切なツールにルーティングします。 2 つの長期マルチモーダル メモリ ベンチマーク、MemEye と MemGallery の実験では、AdaMM がそれぞれ最大 11.3\% と 7.3\% パフォーマンスを向上させることが示されています。
原文 (English)
Beyond Retrieval: Analytic Memory for Multimodal Agents
Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring multimodal observations into queryable structures supporting filtering, aggregation, ranking, and temporal comparison. We present AdaMM, a framework that jointly supports retrieval and analytic memory. Rather than relying on application-defined schemas, AdaMM extracts provenance-linked attribute-value observations from dialogue, images, and contextual metadata, discovers recurring field structures, and materializes them for analytical access. At inference time, a memory-aware planner decomposes queries into retrieval and analytic operations and routes each operation to the appropriate tools. Experiments on two long-term multimodal memory benchmarks, MemEye and MemGallery, show that AdaMM improves performance by up to 11.3\% and 6.9\%, respectively.
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory complianc…
A New Theory of Value for Post-AGI Economics
Artificial general intelligence (AGI) may weaken scarcities in labour, expertise, information, and productive capability that underpin esta…
Where Reasoning Diverges: Localized Multi-Agent Debate for Multi-Hop Question Answering
Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims. We introduce Loca…
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remain…
When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary
Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusabl…
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance t…
Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG
Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candid…
SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents
Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challeng…
A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures
Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in th…
Mixed-Initiative Human-Robot Teaming under Suboptimality with Online Bayesian Adaptation
For effective human-agent teaming, robots and other artificial intelligence (AI) agents must infer their human partner's abilities and beha…
MambaTS: Improved Selective State Space Models for Long-term Time Series Forecasting
In recent years, Transformers have become the de-facto architecture for long-term time series forecasting (LTSF), yet they face challenges…
CollaFuse: Collaborative Diffusion Models
In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic…
Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment
Unsupervised domain adaptation (UDA) aims to mitigate domain shift, where the distribution of labeled source data differs from that of unla…
Patient-centered data science: an integrative framework for evaluating and predicting clinical outcomes in the digital health era
This study proposes a novel, integrative framework for patient-centered data science in the digital health era. We developed a multidimensi…
Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta Solvers
Deep generative models based on neural differential equations have become state-of-the-art for many generation tasks. These models rely on…
Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization
Large Language Models (LLMs) have become a cornerstone for automated visualization code generation, enabling users to create charts through…
Compound and Parallel Modes of Tropical Convolutional Neural Networks
Convolutional neural networks (CNNs) are foundational to many state-of-the-art computer vision systems, yet their reliance on multiplicatio…
When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero
AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal questio…
Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms
Traditionally, in Programming-by-example (PBE) the goal is to synthesize a program from a small set of input-output examples. Lately, PBE h…
One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting
Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. We sh…
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
Importance-based structured pruning overwhelmingly relies on filter magnitude. This proxy is fundamentally flawed: due to scale invariance,…
From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrasti…
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art p…
Uncovering Spontaneous Physics Representations in In-Context Learning
In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, y…
Mechanism of Task-oriented Information Removal in In-context Learning
In-context Learning (ICL) is an emerging few-shot learning paradigm based on modern Language Models (LMs), yet its inner mechanism remains…
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces criti…
韓国の毒性を検出して無毒化するための難読化ルール
言語モデルがオンライン環境に導入されることが増えるにつれて、毒性の検出と解毒に対する注目が高まっています。既存の研究は主に難読化されていないテキストに焦点を当てているため、ユーザーが有害な表現を意図的に偽装した場合の堅牢性が制限されます。特に、韓国語の有害な表現は、膠着形態やハングル特有の正書法のバリエーションによって簡単に隠すことができます。ただし、韓国語の難読化はほとんど解明されていないため、難読化と解毒のための KOTOX: 韓国の有毒データセットを導入する動機となっています。私たちは、韓国語の難読化パターンを言語ベースのクラスに分類し、現実世界の例から派生した変換ルールを定義し、その結果得られる難読化フレームワークをオープン変換パッケージとして提供します。これらのルールを使用して、中立的な文と有害な文のペアを、難読化された文と並べて提供します。データセットでトレーニングされたモデルは、難読化されていないテキストのパフォーマンスを犠牲にすることなく、難読化されたテキストをより適切に処理します。これは、韓国語の難読化解除と解毒化を同時にサポートする最初のデータセットです。私たちは、このデータセットによって、韓国語向け LLM の難読化された有害コンテンツの理解と緩和が促進されることを期待しています。コードとデータは https://github.com/leeyejin1231/KOTOX で入手できます。
原文 (English)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
As language models become increasingly deployed in online environments, toxicity detection and detoxification have received growing attention. Existing studies primarily focus on non-obfuscated text, which limits robustness when users intentionally disguise toxic expressions. In particular, Korean toxic expressions can be easily disguised through agglutinative morphology and Hangeul-specific orthographic variation. However, obfuscation in Korean remains largely unexplored, which motivates us to introduce a KOTOX: Korean toxic dataset for deobfuscation and detoxification. We categorize Korean obfuscation patterns into linguistically grounded classes, define transformation rules derived from real-world examples, and provide the resulting obfuscation framework as an open transformation package. Using these rules, we provide paired neutral and toxic sentences alongside their obfuscated counterparts. Models trained on our dataset better handle obfuscated text without sacrificing performance on non-obfuscated text. This is the first dataset that simultaneously supports deobfuscation and detoxification for the Korean language. We expect the dataset to facilitate better understanding and mitigation of obfuscated toxic content in LLM for Korean. Our code and data are available at https://github.com/leeyejin1231/KOTOX.
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are…
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended o…
GraphCliff: Short-Long Range Gating for Modeling Critical Activity Changes Caused by Subtle Molecular Differences
The quantitative structure-activity relationship assumes a smooth mapping between molecular structure and biological activity. However, act…
Target-Aligned Fusion for Decision-Sequence Learning under Dynamics Shift
External trajectories can improve offline decision-sequence learning, but dynamics shift may make some source subsequences inconsistent wit…
$\pi$-Attention: Online Efficient Sparse Transformers for Long-Context Modeling
Sparse attention is crucial in long-context Transformers, which restricts each token to a limited neighborhood and thereby reduces the quad…
STREAM-VAE: Dual-Path Routing for Slow and Fast Dynamics in Vehicle Telemetry Anomaly Detection
Automotive telemetry data exhibits slow drifts and fast spikes, often within the same sequence, making reliable anomaly detection challengi…
Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors
External validation of breast ultrasound segmentation models remains limited because internal train--test splits do not capture domain shif…
MIMIC-MJX: Neuromechanical Emulation of Animal Behavior
The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex beh…
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We…
PRISMA: Improving the Accuracy-Latency Frontier of Diffusion-based PDE Solvers Using Physics-Informed Spectral Attention
Diffusion-based solvers for partial differential equations (PDEs) are often bottle-necked by slow gradient-based test-time optimization rou…
PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks
Vertical Federated Learning (VFL) enables collaborative model training across organizations that share common user samples but hold disjoin…
HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize
Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforce…
Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the…
ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing
Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community…
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection
Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security a…
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classi…
A Deployment-Friendly Foundational Framework for Efficient Computational Pathology
Pathology foundation models (PFMs) generalize well across computational pathology tasks but remain costly for gigapixel whole-slide image a…
In-Context Pure Exploration in Continuous Decision Spaces
In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to id…
SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Information Geometry
Reliable decision-making in complex multi-agent systems requires calibrated predictions and interpretable uncertainty. We introduce SphUnc,…
Quantifying Hallucinations in Language Language Models on Medical Textbooks
Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious p…
Large Language Models provide support for the parallelogram theory of analogy
Four-term word analogies (A:B::C:D) are classically modeled geometrically as parallelograms: adding the vector B-A+C produces D. Recent wor…
The production of meaning in the processing of natural language
Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designin…
DIB-OD: Preserving the Invariant Core for Robust Heterogeneous Graph Adaptation via Decoupled Information Bottleneck and Online Distillation
Graph pre-training can facilitate knowledge transfer across graph datasets, but severe structural and feature shifts may cause negative tra…
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone…
Gated Memory Policy: In-Context Memorization and Adaptation
Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks…
A neural operator framework for data-driven discovery of stability and receptivity in physical systems
Understanding how complex systems respond to perturbations, such as whether they will remain stable or what their most sensitive patterns a…
Estimating Tail Risks in Language Model Output Distributions
Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these model…
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful t…
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points,…
Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents
We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not…
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworth…
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refac…
ActQuant: 視覚・言語・アクションモデル向けのサブ 4 ビットのアクションガイド付き量子化
Vision-Language-Action(VLA)モデルは、身体化されたインテリジェンスに対して顕著なアクション生成を示しますが、その大量のコンピューティングにより、エッジ プラットフォームへの展開は非現実的です。積極的なサブ 4 ビット重み量子化は自然な解決策ですが、既存のポストトレーニング量子化 (PTQ) 手法は、この領域では重大なパフォーマンス低下に悩まされます。これに対処するために、アクション ガイド付き混合精度 PTQ フレームワークである ActQuant を導入します。これは 2 つの段階で動作します。(1) エージェントのアクションの予測にどれだけ寄与するかに基づいて、各重み行列に単一のビット幅を割り当てるテンソル間ビット アロケーター。 (2) テンソル内スケール オプティマイザーは、アクションを意識した曲率を使用してブロックごとの量子化スケールを調整し、ダイナミック レンジが制御に最も影響を与える重みに集中するようにします。積極的な量子化によるオンデバイスのメリットを実現するために、効率的な低ビット カーネルを備えたネイティブ C/C++ ランタイムにアーキテクチャを移植するエージェント変換パイプラインである OmniModel.cpp をさらに導入します。すべてのモデルが OmniModel.cpp を通じて展開され、シミュレーションと実際の 6-DoF UR3 アームの両方で ActQuant を評価します。 LIBERO ベンチマークでは、ActQuant は重みあたり 3 ビット以下で動作する唯一のメソッドであり、OpenVLA-OFT では 95.0%、$\pi_{0.5}$ では 94.8% を維持しています。さらに前進すると、ActQuant は OpenVLA-OFT 上で 90.1% で 2.5 bpw に達し、バックボーンを 14.3 GB から 2.7 GB (5.3$\times$) に圧縮します。物理 UR3 アームでは、ActQuant で量子化された $\pi_{0.5}$ はベースラインの成功率を維持しながら、メモリ フットプリントを 2.5$\times$ 削減します。
原文 (English)
ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models
Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressive, sub-4-bit weight quantization is the natural solution, yet existing post-training quantization (PTQ) methods suffer severe performance degradation in this regime. To address this, we introduce ActQuant, an action-guided mixed-precision PTQ framework that operates in two stages: (1) an inter-tensor bit allocator that assigns each weight matrix a single bit-width based on how much it contributes to predicting the agent's actions; (2) an intra-tensor scale optimizer tunes per-block quantization scales using action-aware curvature, so that dynamic range is concentrated on the weights most influential for control. To deliver the on-device benefits of our aggressive quantization, we further introduce OmniModel.cpp, an agentic conversion pipeline that ports architectures into a native C/C++ runtime with efficient low-bit kernels. We evaluate ActQuant both in simulation and on a real-world 6-DoF UR3 arm, with all models deployed through OmniModel.cpp. On the LIBERO benchmark, ActQuant is the only method that operates at or below 3 bits-per-weight, retaining 95.0% on OpenVLA-OFT and 94.8% on $\pi_{0.5}$. Pushed further, ActQuant reaches 2.5 bpw at 90.1% on OpenVLA-OFT, compressing the backbone from 14.3 GB to 2.7 GB (5.3$\times$). On the physical UR3 arm, $\pi_{0.5}$ quantized with ActQuant retains the baseline's success rate while reducing the memory footprint by 2.5$\times$.
E4GEN: Event-level Explainable Extreme-Enhanced Time-series Generation
Generating realistic time series is essential for scientific research and real-world applications. However, existing methods often emphasiz…
FLARE: Diffusion for Hybrid Language Model
Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck fo…
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and ans…
When Context Returns: Toward Robust Internalization in On-Policy Distillation
Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student…
CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners
End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that me…
Diagnosing and Mitigating Context Rot in Long-horizon Search
Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon search tasks. The concern t…
MalariAI: 高密度マラリア血液塗抹標本における普遍的な細胞セグメンテーションと説明可能な病期分類のためのラベル耐性のある分離フレームワーク
血液塗抹標本顕微鏡検査による自動マラリア診断は、世界規模の保健 AI における重要な課題です。リソースが限られた環境では、専門の顕微鏡医の不足が依然としてタイムリーで正確な診断の主なボトルネックとなっています。 3 つの複合的な障害モードにより、既存の深層学習システムの信頼できる臨床展開が妨げられます。まず、エンドツーエンドの検出器は、トレーニング中に注釈のないセルをバックグラウンドとして扱い、実際のセルの回復を反映するのではなく、注釈の完全性に強く影響される再現率の数値を生成します。第 2 に、非最大抑制は、感染数が最も重要な密度の高いスミア領域での有効な検出を抑制する傾向があります。第三に、マラリア画像分類タスクには Grad-CAM などの画像レベルの説明可能手法が適用されているにもかかわらず、既存のスライド全体検出パイプラインには、臨床監査のためのセルごとの空間的証拠が不足しています。統合パイプラインで 3 つの障害モードすべてに対処する 2 段階の分離フレームワークである MalariAI を紹介します。ステージ 1 では、アノテーションに依存しない距離変換ガイド付き流域アルゴリズムを適用して、1600x1200 の完全な血液塗抹標本画像内のすべての細胞を分離し、グラウンド トゥルースの入力なしで 120 枚の画像 NIH BBBC041 テスト セット全体の重心位置特定により、グラウンド トゥルースの細胞の 75.95% を回復します。ステージ 2 では、64x64 作物で焦点損失 (ガンマ = 2.0、クラスごとの逆周波数重み) を使用して EfficientNet-B0 を微調整し、まれなシゾントおよび生殖母細胞のステージでは 98.36% の全体的な分類精度と 87.5% および 75.0% のクラスごとの精度を達成しました。一方、同じクラスでの R-CNN ベースラインの高速化。検出された細胞ごとに生成された Grad-CAM++ ヒートマップは、臨床監査のためのインスタンス レベルの空間的証拠を提供し、顕微鏡医が分類パフォーマンスを犠牲にすることなく、個々の寄生虫レベルでモデルの予測を検証できるようにします。
原文 (English)
MalariAI: A Label-Resilient Decoupled Framework for Annotation-Agnostic Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
Automated malaria diagnosis from blood smear microscopy is a critical global health AI challenge; expert scarcity remains the primary diagnostic bottleneck. Existing deep learning systems face three compounding failures: end-to-end detectors treat unannotated cells as background, skewing recall by annotation completeness rather than true cell recovery; Non-Maximum Suppression suppresses valid detections in dense smears; and pipelines lack per-cell spatial evidence for clinical audit. We present MalariAI, a two-stage decoupled framework addressing all three. Stage 1 applies an annotation-agnostic watershed algorithm to isolate every cell in a full 1600x1200 image, recovering 75.95% of ground-truth cells without any ground-truth input. End-to-end, the pipeline reaches a binary parasitized AP@0.5 of 29.10% - the clinically relevant metric for flagging any infected cell - while the stricter multi-class mAP@0.5 of 8.67% mainly reflects watershed's organic region boundaries being penalized against axis-aligned ground-truth boxes, not a localisation failure. Stage 2 fine-tunes EfficientNet-B0 with Focal Loss on ground-truth crops, achieving 98.36% classification accuracy - an oracle upper bound once a cell is correctly localised - with 87.5% and 75.0% accuracy on the rare schizont and gametocyte stages, versus 38.45% and 57.27% AP for a modern YOLOv8s detector evaluated end-to-end on the same classes. Grad-CAM++ heatmaps generated per detected cell provide instance-level spatial evidence for clinical audit; a quantitative energy-in-box analysis confirms this activation is concentrated on the annotated cell body significantly above a geometric chance baseline (+0.0485, paired p = 1.4 x 10^-33), letting microscopists verify predictions at the individual parasite level without sacrificing classification performance.
VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training para…
Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks
Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological…
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
Emergent misalignment (EM) --- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data --- is mediated in…
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function eva…
主観的リスクの分解: 不確実性の定量化のための新しい視点
不確実性の定量化に対する新しい視点を提案します。不確実性の尺度は、公理と議論を必要とする原始的なものではなく、より高いレベルのモデリング決定の結果です。厳密に適切な損失に基づいて、主観的リスクの分解を通じて認識的および偶然的な不確実性の尺度をどのように導き出すことができるかを示します。逆クロスエントロピーは、分解によって古典的な情報理論の不確実性項を回復する顕著な例を提供します。同じアプローチにより、UQ 文献全体で以前に提案された多数の尺度が回収され、それらに共通の理論的基盤が提供されます。実用的な観点から、これは UQ への新しいアプローチを示唆しています。モデリング シナリオと厳密に適切な損失が与えられると、対応する認識項と偶然項が主観的リスク分解によって誘導されます。次に、視野を学習理論に拡張します。超過リスク、近似誤差、推定誤差の主観的リスク類似物を導入して分析し、UQ との関連性を特定します。私たちは、これが不確実性の定量化のための完全な学習理論的フレームワークに向けた第一歩であると考えています。
原文 (English)
Subjective Risk Decomposition: A New View for Uncertainty Quantification
We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation, but instead consequences, of higher-level modelling decisions. We show how epistemic and aleatoric uncertainty measures can be derived via decomposition of a subjective risk, based on a strictly proper loss. Reverse cross entropy provides a prominent example, where decomposition recovers the classic information-theoretic uncertainty terms. The same approach recovers numerous measures previously proposed across the UQ literature, providing them a common theoretical foundation. This suggests a new approach to UQ: given a modelling scenario and strictly proper loss, the corresponding epistemic and aleatoric terms are induced by the subjective-risk decomposition. We then extend our view to learning theory: we introduce and analyse subjective risk analogues of excess risk, approximation error and estimation error, and identify the connections to UQ. We consider this a first step towards a full learning-theoretic framework for uncertainty quantification.
ビジネス分野全体にわたる最先端の AI パフォーマンス: ナレッジワークと分析的推論の事例に基づいたベンチマーク
大規模言語モデル (LLM) は、ベンチマーク スコアに反映されているように急速に改善されていますが、これらの AI ベンチマークでは主に、事実の再現、限定的な質問応答、数学的問題解決、コーディングやエージェント ツールの使用などの機能がテストされます。まだ十分に測定されていないのは、複雑な情報の統合、不確実性と不完全な情報の下での判断の行使、複数のステークホルダーの状況での戦略的および敵対的思考の適用、トレードオフの比較検討、防御可能な構造化された分析の作成など、ホワイトカラーの専門家が日々行っている分析知識作業における AI の進歩です。このギャップは、そのような仕事の主観的な要素ではさらに顕著であり、成功を定義するのが難しい場合があります。トップクラスのビジネススクールが実践する「ケースメソッド」教育形式は、この測定ギャップに対処するための自然な基盤を提供します。私たちは、18 分野にわたるビジネスケースから抽出された数百の質問にわたるベンチマークである BusinessCaseBench を構築します。各質問は、専門家が作成した講師のケースソリューションから導き出された採点ルーブリックと対になっています。 BusinessCaseBench では、フロンティア AI モデルはすでにインストラクターのルーブリックに対して高いスコアを獲得しており、1 つのモデル ファミリー内の機能は 2 年間で大幅に向上しています。これらの結果は、この種の作業における AI のパフォーマンスがすでに高く、急速に向上していることを示す強力な証拠を提供します。これは、事例教育学によって学部生や MBA がこの種の分析的推論を訓練されるビジネス スクールや、歴史的にそのようなスキルが初期キャリアの仕事に定着してきたエントリーレベルの専門職に影響を及ぼします。
原文 (English)
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.
単一セル データの抽出を監査可能にする: 個別の最小値と最大値の選択による追跡可能なリアル セル コアセット
単一セル データセットの保存、監査、モデル トレーニングのための再利用のコストはますます高くなっています。次元の削減とデータセットの蒸留によりこの負担を軽減できますが、従来の蒸留方法では、アッセイされた細胞を追跡できない合成発現プロファイルが生成されることがよくあります。私たちは、固定された細胞と遺伝子のバジェットの下で元の細胞識別子と遺伝子シンボルを保持しながら、追跡可能な単一細胞データの蒸留を定式化します。結果として得られるトレーニング サブセットは、測定されたカウント、ラベル、およびアッセイ メタデータに接続されたままであるため、予期せぬ予測をソース データと照合してチェックできます。 2 つの実セル セレクターを提案します。固定 CF は静的な特性関数マッチングを使用します。 Minmax-CF は、あまり保存されていない方向を重み付けし、観測されたセルのみを追加する、エントロピー正則化された離散最小-最大問題を解決します。 3 つのデータセットのドナー、テクノロジー、および摂動レベルのシフト全体で、Minmax-CF は MS で完全バランスの精度の 96.52% を維持し、h膵臓で平均で完全とほぼ一致し、全遺伝子設定で中央値 $2.55\times$ の GPU 速度向上を実現し、Norman で圧縮メソッドの中で最も低い経路エラーを取得しました。まれな状態、一部のテクノロジーの変化、目に見えない摂動コンポーネント、および忠実度が下流のユーティリティとの関連性が低い設定では、パフォーマンスは依然として低いままです。選択された ID は測定されたセルを参照するため、これらのケースは、対応するトレーニング サポート、ラベル、およびアッセイ メタデータを検査することによって調査できます。 Minmax-CF は最悪方向の不一致を一貫して削減しますが、下流のユーティリティとコストはデータセットやタスクによって異なります。
原文 (English)
Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min--Max Selection
Large single-cell datasets are expensive to store, curate, and repeatedly reuse for model training. Data distillation can reduce this burden by building smaller training sets. However, many existing methods rely on synthetic cells. These synthetic cells do not retain direct correspondence with assayed cells and genes. This limits source-level inspection and biological traceability. Moreover, real-cell expression matrices are often sparse and noisy. In light of these challenges, we propose Minmax-CF, a label-aware characteristic-function selector for traceable single-cell data distillation. Minmax-CF formulates compression as a discrete min--max selection problem over characteristic-function directions. It uses entropy-regularized maximization to emphasize the least preserved directions. Greedy minimization ranks cells and genes by how much they reduce the resulting weighted error. The method alternates cell and gene selection under explicit axis-specific budgets. Across five coarse-lineage benchmarks and five compression budgets, Minmax-CF retains 95.3% of the Full-reference macro-F1 on average, with gaps that exceed one per-seed standard deviation. It also retains exact source-cell indices and original gene symbols. Compared with size-matched synthetic PCA-Centroid and Distribution Matching (DM) baselines, Minmax-CF achieves higher coarse-lineage macro-F1 in 24 of 25 comparisons against each baseline. It exceeds their average performance by 10.4% and 17.4%, respectively. Retained cells can also be projected onto independently computed embeddings for direct biological interpretation.
A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability
Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but und…
TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex
Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation betwee…
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A co…
Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcem…
WCM: World-Cognition Model for Generalizable Human-Robot Interaction
Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tas…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr\"om's team…
AgentGUI: 長時間実行される AI エージェントを監視および操作するためのインターフェイス
AI エージェントは、複雑で長時間実行されるタスクへの取り組みにますます熟練しています。自律機能の急速な急増に伴い、人間中心のインターフェースが限られているため、人間の監視は体系的に遅れています。これに対処することを目的として、複数の同時長時間実行セッションで AI エージェントをシームレスに監視および操作するための、ユーザーフレンドリーでローカルにホストされる GUI である AgentGUI を導入します。 AgentGUI の特徴は、1) 豊富なエージェント軌跡の視覚化、2) 効果的な手動および自動ステアリング、3) オープンソースおよびフロンティア エージェント フレームワークとの統合および調整です。管理されたユーザー調査では、エージェントのトレースから主要な要素を特定するのにかかる時間が統計的に有意に短縮されたことが実証されました (38% 高速化、p = 0.023)。予備実験では、AgentGUI の自動ドリフト防止機能により、0.8B ~ 9B モデル ラダー (モデルあたり N=50 実行) 全体で小規模のローカル エージェントのタスク完了率が 34pp も上昇しました。 AgentGUI は、プロジェクト Web サイト (https://agent-gui-project.github.io) およびオープンソース リポジトリ (https://github.com/eth-medical-ai-lab/agent-gui) を通じて、デモ ビデオ (https://youtube.com/watch?v=GSDyxN1gTF0) とともに公開されています。
原文 (English)
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents
AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, and 3) integration with and coordination between open-source and frontier agent frameworks. A controlled user study demonstrates statistically significant reduction in the time it takes to identify key elements from agent traces (38% faster, p = 0.023). In a preliminary experiment, AgentGUI's automated drift prevention feature raises the task completion rate of small local agents by as high as 34pp across a 0.8B--9B model ladder (N=50 runs per model). AgentGUI is publicly available through its project website (https://agent-gui-project.github.io) and open-source repository (https://github.com/eth-medical-ai-lab/agent-gui), along with a demo video (https://youtube.com/watch?v=GSDyxN1gTF0).
Benchmarking LLM Competence on Logical Inference over Probability Operators
Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions o…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cu…
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pip…
小規模言語モデルにおけるバイアスに対する知識蒸留の非対称的影響
小規模な命令調整言語モデルにおける知識の蒸留がバイアスに対して非対称な影響を与えることを示します。明確なタスク (BBQ-disambig) では、Gemma-2-9B 教師からの応答ベースの蒸留によりコンテキスト追従が向上します。最も偏ったベースライン (SmolLM2-1.7B-Instruct) では、コンテキストを優先するエラー率が 44% から 24% に削減されます。あいまいなタスク (BBQ-ambig) では、同じ蒸留により項目ごとの拒否キャリブレーションが破壊されます。全体的な拒否率が維持されている場合でも、ベースラインが正しく棄権された項目の 15% が、代わりにステレオタイプの回答を受け取ります。このパターンは 2 番目の生徒ファミリー (OLMo-2-1B-Instruct) でも再現され、沈黙の喪失が 8% で、沈黙の充満が新たなバイアスの 89% を占めています。 28 の構成グリッド全体にわたって、沈黙の損失と満たされた沈黙の大きさには相関がなく (Spearman $\rho=0.19$, n.s.)、2 つの効果が異なるメカニズムから生じることを示しています。集約されたステレオタイプ メトリクス (CrowS-Pairs、全体的な BBQ ステレオタイプ信頼性スコア) は両方の効果を平均し、項目ごとの害を隠します。キャリブレーションの損失はデータ側のメカニズムに起因すると考えられます。4 つのトレーニング コーパスの監査では、回答形状としての拒否が 0.5% 未満であることがわかりました。拒否注入を伴う教師あり微調整 (SFT) は、解析を中断するか、集計メトリクスが完全に調整されていると言える自明な拒否者体制 (拒否率 99.8%、曖昧さの解消精度 0.2%) に過剰修正します。我々は、拒否キャリブレーション、コンテキスト追従、機能維持を評価する 3 段階のプロトコルである条件別キャリブレーション診断 (PCCD) を提案します。 PCCD は、集計評価が見逃す非対称の害と自明な拒否障害モードの両方を捕捉します。
原文 (English)
The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models
We show that knowledge distillation (KD) in small instruction-tuned language models has asymmetric effects on bias, and that measuring them correctly requires accounting for where refusal mass moves and what the parser can legitimately score. On unambiguous tasks (BBQ-disambig), response-based distillation from a Mistral-7B teacher genuinely improves context-following for the most context-biased baseline (SmolLM2-1.7B-Instruct): among committed (non-abstaining) answers, the rate of overriding correct context with a stereotype falls from 44.5% to 37.2%, with accuracy rising from 0.55 to 0.61. On ambiguous tasks (BBQ-ambig), the same distillation degrades conditional refusal: 15% of the cases where the baseline correctly abstained instead receive stereotype answers (silence-loss), and the distilled refusal pattern only weakly preserves the baseline's (Spearman rho=0.44). The harm reproduces, aggravated, on a second student family (OLMo-2-1B-Instruct): silence-loss reaches 49% and filled-silence accounts for 95% of new bias. Two apparently stronger results are artifacts. An unconditioned override metric reports a 44% -> 23% improvement under a Gemma-2-9B teacher that shrinks to 44.5% -> 39.8% once conditioned on committed answers: the model abstains on 43% of items and its accuracy collapses from 0.55 to 0.35. An apparent cross-condition independence reverses to a positive correlation (rho=0.58, p<0.01) on the valid 19-configuration grid once parser-invalid logit-KD configurations are excluded and the parser is corrected. Aggregate metrics (CrowS-Pairs, overall BBQ Stereotype Reliance Score) average over both effects and conceal the per-item harm. We propose Per-Condition Calibration Diagnosis (PCCD), a three-step protocol evaluating refusal-pattern preservation, committed-answer context-following, and capability preservation. No configuration in our grid passes all three steps.
Automated ECG Interval Measurement and Wave Delineation Using Fast Fourier Convolution ResNet
Accurate measurement of ECG intervals, including PR, QRS duration, and QT/QTc, is central to cardiac diagnosis, yet the published ECG delin…
Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs
Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from image…
LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations
LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a m…
Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems
Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, eff…
When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration expos…
ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG
Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed contex…
Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment
Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics th…
Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution
Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. I…
Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation
However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation rema…
LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation
Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities. To bypa…
Self-Improving Large Language Models via Progressive Experience Evolution
Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism fo…
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an unde…