AIニュース 2026-07-03
自動生成: 2026-07-03 12:45 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Meta、「Claude Codeと組織改編で爆速開発」のはずが「想定より加速せず」 ザッカーバーグ氏、社内集会で発言ITmedia AI+
MetaはAI向けの巨額投資や組織改編でAIエージェント開発の加速を図ったが、ザッカーバーグCEOによれば、それらの取り組みはまだ実を結ん…
-
Meta quietly launches vibe-coded gaming app PocketTechCrunch AI
Meta has quietly launched Pocket, an experimental AI app that lets us…
-
Anthropic is discussing a new custom chip with SamsungTechCrunch AI
The news comes about a week after OpenAI announced its own custom AI…
-
復活した「Fable 5」 米政府からのオーダーに対して、Anthropicはどう対策したのかITmedia AI+
米AnthropicのAIモデル「Claude Fable 5」が世界的にサービスを再開した。Anthropicは復活に向けてどういった経…
-
ゲームエンジン「Godot」AI生成コードを原則禁止へ レビュアー疲弊「機械と話したくない」ITmedia AI+
新規コントリビューターからAI生成のプルリクが急増する一方、審査するレビュアーの数は変わらず、負担が限界に達したという。人間同士のやりとり…
-
「顧客管理システム作って」日本語で指示するだけでシステムが完成する最新AIとは?ITmedia AI+
業務システムを開発したくても、費用や人材の確保がネックとなり、なかなか着手できない会社は少なくない。こうした課題に対し、専門知識がなくても…
-
AWSの「静かな」戦略シフト OpenAIとAnthropic“1日違い登壇”の意味を読み解くITmedia AI+
生成AIで競合するOpenAIとAnthropicを1日違いで基調講演に招く――。「AWS Summit Japan 2026」で浮かび上…
トピック別件数
- LLM/生成AI 111件
- 研究/論文 108件
- エージェント 78件
- 画像/動画生成 68件
- ロボティクス 28件
- ビジネス/資金調達 12件
- その他 7件
- ハードウェア/半導体 6件
- 規制/政策 1件
日本語メディア13件
ITmedia AI+ (日本語)
ゲームエンジン「Godot」AI生成コードを原則禁止へ レビュアー疲弊「機械と話したくない」
新規コントリビューターからAI生成のプルリクが急増する一方、審査するレビュアーの数は変わらず、負担が限界に達したという。人間同士のやりとりで、AI生成の文章を使うことも禁止する。
Meta、「Claude Codeと組織改編で爆速開発」のはずが「想定より加速せず」 ザッカーバーグ氏、社内集会で発言
MetaはAI向けの巨額投資や組織改編でAIエージェント開発の加速を図ったが、ザッカーバーグCEOによれば、それらの取り組みはまだ実を結んでいないようだ。
「顧客管理システム作って」日本語で指示するだけでシステムが完成する最新AIとは?
業務システムを開発したくても、費用や人材の確保がネックとなり、なかなか着手できない会社は少なくない。こうした課題に対し、専門知識がなくても開発を進められる手段が広がりつつある。
AWSの「静かな」戦略シフト OpenAIとAnthropic“1日違い登壇”の意味を読み解く
生成AIで競合するOpenAIとAnthropicを1日違いで基調講演に招く――。「AWS Summit Japan 2026」で浮かび上がったのは、モデルの賢さではなく「別のあるもの」を握ることで、基盤の価値を保ち続けようとするAWSの戦略シフトだ。
AIで実機との形状差を学習するプレス成形シミュレーションソフトウェア
JSOLは、プレス成形シミュレーションソフトウェア「JSTAMP」にAI機能を搭載した。実機トライとシミュレーションの形状差を学習して補正し、高強度鋼板のスプリングバック予測精度を向上させる。
「Mythosがないと守れない」は本当か??AIセキュリティの勝負を分ける「ハーネス」とは
「Mythos級のAIにアクセスできない企業は、もう守れない」??。Mythosのアクセスが一部の組織に限られていた2026年春、そんな脅威論が広がった。だが、AIによる初期侵入の自動化を世界で初めて実現したと公表する当事者は、そうは見ていない。勝負を分けるのは、アクセスの有無…
え、21日で37テラも? 高性能SSDを食いつぶす「あのAIツール」にご用心:886th Lap
AIツールを使っていただけなのに、SSDの寿命が想像以上の速さで縮んでいた。そんな問題が明らかになった。原因は、ツールの実装上の不具合だ。もし、そのAIツールを使っているなら、一度確認しておきたい。
フィジカルAIに“二刀流”で対応するアドバンテック、日本に第3の製造拠点を構築
産業用PCで世界シェアトップのアドバンテックが、エッジAI市場の拡大に併せて組み込み機器部門の事業への注力を鮮明にしている。アドバンテック台湾本社のTony Chen氏と、アドバンテック日本法人の李威震氏に、フィジカルAIをはじめエッジAIを中核とする事業戦略について聞いた。
ソフトウェアエンジニアの仕事は「ループを書くこと」になる 内側ループと外側ループ(ハーネス)入門
AIコーディングにおける「ループ」には、エージェントが回す内側ループと、ハーネスが回す外側ループの2種類がある。両者の違いと外側ループがもたらす課題を、アルミン・ロナッハー氏の記事に沿って初心者向けに解説し、その「記憶」の扱いについての筆者の考えも添える。
人型ロボットが工場で稼働する様子を6日間生配信、作業成功率99.99%をうたう 中国メーカー
中国の人型ロボット開発企業AGIBOTは、実際のタブレット量産ラインで複数の人型ロボットを6日間連続で動かす様子をライブ配信した。延べ64時間で1万7625個のタブレット生産に貢献し、作業成功率は99.99%だったという。
復活した「Fable 5」 米政府からのオーダーに対して、Anthropicはどう対策したのか
米AnthropicのAIモデル「Claude Fable 5」が世界的にサービスを再開した。Anthropicは復活に向けてどういった経緯と対策を行ったのか、モデル再開にあわせて詳細を公開した。
国内大手ロボットメーカー3社が協力、「フィジカルAI」向けデータセット構築へ
川崎重工業は、ロボットメーカー大手のファナックや安川電機などと協力し、「フィジカルAI」向けのデータセットを構築すると発表した。「GENIAC」の公募に採択された。
「賢さよりツール連携力が重要」 Microsoftが実験的な小型AIエージェント基盤を公開
MicrosoftのAI研究チームであるMicrosoft Research AI Frontiersは、小型モデル向けに最適化したエージェント基盤「MagenticLite」を公開した。「エージェント能力は知識量ではなくツール統合と実行ハーネスで決まる」という仮説に基づき構成…
海外メディア7件
TechCrunch AI (英語)
Mark Zuckerberg tells staff that AI agents haven’t progressed as quickly as he’d hoped
At an internal meeting, the Meta CEO reportedly said that AI development efforts were not moving as quickly as anticipated.
Jersey Mike’s IPO illustrates how bad the AI hype has become
Just for kicks, I took a look at Jersey Mike's IPO documents. Surely a sandwich shop would have no need to mention AI. But lo and behold.
Meta quietly launches vibe-coded gaming app Pocket
Meta has quietly launched Pocket, an experimental AI app that lets users generate and share interactive mini games using text prompts.
Anthropic is discussing a new custom chip with Samsung
The news comes about a week after OpenAI announced its own custom AI chip in a partnership with Broadcom.
OpenAI proposed donating 5% of its equity to a US sovereign wealth fund
OpenAI CEO Sam Altman has reportedly proposed giving 5% of the company’s equity to a U.S. sovereign wealth fund, reviving discussions about…
Yep, we’re using OpenClaw to date now
Ben Guez has "a bunch of potential international wives in [his] DMs," thanks to an automated script he set up using OpenClaw, Claude code,…
Indian tech tycoon bets $30M of his own money to build AI alternative to Microsoft Office
Neo is Bhavin Turakhia’s fifth venture and his latest involving enterprise software. This time he's taking on Microsoft Office and Google A…
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文293件
arXiv cs.AI (英語)
建設的な調整: 人間と AI のインタラクションにおける好みのダイナミクスを制御する
AI 調整へのほとんどのアプローチは、人間の好みを推測および最適化される固定ターゲットとして扱います。この仮定は、好みが階層的で動的であり、相互作用、特に適応技術との相互作用を通じて構築されることを示す広範な経験的証拠と矛盾します。 AI システムがより永続的でパーソナライズされ、社会に組み込まれるにつれて、時間の経過とともに人々が注目し、評価し、支持するものの形成にますます関与するようになります。私たちは、静的な嗜好の満足度ではなく、進化する人間の嗜好の軌跡に対する制御問題として調整を再構成するパラダイムである、建設的調整を紹介します。行動経済学、心理学、構成主義的社会理論を利用して、AI システムとの相互作用の下で進化する層状の状態変数として嗜好をモデル化します。我々は、システムの動作とインタラクション設計が共同して世界状態と人間の評価状態の両方に影響を与える制御理論のフレームワークを使用して、この見解を形式化します。私たちは、調整とは主に AI の行動を制御することではなく、AI システムが人間の好みの進化にどのように影響するかを規制すること、つまり価値の軌道が一貫性を保ち、内省的に承認され、認識論的に根拠があり、操作を制限され、不確実性の下で力を与えることを保証することであると主張します。したがって、調整は単に静的な好みを満たすというよりも、長期的な価値形成を管理する問題になります。
原文 (English)
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirical evidence showing that preferences are layered, dynamic, and constructed through interaction--particularly with adaptive technologies. As AI systems become more persistent, personalized, and socially embedded, they increasingly participate in shaping what people attend to, value, and endorse over time. We introduce Constructive Alignment, a paradigm that reframes alignment as a control problem over evolving human preference trajectories rather than static preference satisfaction. Drawing on behavioral economics, psychology, and constructivist social theory, we model preferences as layered state variables that evolve under interaction with AI systems. We formalize this view using a control-theoretic framework in which system actions and interaction design jointly influence both world states and human evaluative states. We argue that alignment is not primarily about controlling AI behavior, but about regulating how AI systems influence the evolution of human preferences--ensuring that value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. Alignment thus becomes a problem of governing long-term value formation rather than simply satisfying static preferences.
境界のある道徳: 道徳的計算の空間を定義する
道徳的認知は伝統的に、静的なルールまたは価値関数として実装された固定的な倫理理論(義務論、結果主義、美徳倫理)への準拠としてモデル化されてきました。私たちは、有限なエージェントが直面する道徳的問題の計算要求を分析するための正式なフレームワークである、境界付き道徳を提案します。ハーバート・サイモンの限定合理性の概念を拡張して、道徳的状況を 2 つの直交する次元に沿って形式化します。道徳的幅、道徳的に関連するものとして扱われるエンティティの範囲、および道徳的深さ、それらの相互作用を評価するために必要な推論的統合です。リソースが限られているため、これらの次元の間には避けられないトレードオフが課せられ、道徳的計算の実行可能な空間が定義されます。この空間内では、倫理理論は、道徳的真実の競合する説明ではなく、さまざまな需要体制に適応した局所的に効率的な戦略に対応します。この枠組みは、制約の下での道徳的後悔と道徳的進歩という形式的な概念を生み出し、人工システムにおける道徳的調整は、人間の判断の直接の模倣ではなく、道徳的推論能力のスケーリングと割り当てに依存することを示唆しています。
原文 (English)
Bounded Morality: Defining the Space of Moral Computation
Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--implemented as static rules or value functions. We propose Bounded Morality, a formal framework for analyzing the computational demands of moral problems faced by finite agents. Extending Herbert Simon's notion of bounded rationality, we formalize moral situations along two orthogonal dimensions: moral breadth, the scope of entities treated as morally relevant, and moral depth, the inferential integration required to evaluate their interactions. Limited resources impose an unavoidable tradeoff between these dimensions, defining a feasible space of moral computation. Within this space, ethical theories correspond to locally efficient strategies adapted to different demand regimes rather than competing accounts of moral truth. The framework yields a formal notion of moral regret and moral progress under constraint, and implies that moral alignment in artificial systems depends on the scaling and allocation of moral reasoning capacity rather than on direct imitation of human judgments.
MMM データ モデル -- 分散型ナレッジ コモンズにおける知識の相互運用性の標準仕様
多くの情報システムはドキュメントを中心に構築されており、印刷物作成とリニア読み取り用に最適化された自己完結型ユニットです。ドキュメント中心の組織は大規模な普及には効果的ですが、知識を構造化し、更新し、共有し、再利用する方法に制約があります。形式的なアプローチはこれらの制限の一部に対処しますが、人間の使いやすさや範囲などの他のシステム特性よりも形式的な構造を優先するため、広範な貢献と採用を達成するのに苦労しています。 AI システムは文書作成を再構築していますが、人間による知識の表現と交換のための従来の文書に代わる統合されたポータブルな代替手段は提供されていません。この論文では、学際的な共同研究の実際的なニーズから生まれた知識文書化のためのデータ モデルである MMM を紹介し、情報システムの設計空間の比較分析の中に位置づけています。 MMM は、一連の規範的な制約とフリーテキスト ラベルの表現の自由を組み合わせたものです。セマンティックな収束を必要とせずに、分野、アプリケーション、展開全体で相互運用できるように設計されています。リファレンス実装とパイロット展開データは、実装可能性と初期の使いやすさを示しています。
原文 (English)
The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons
Many information systems are built around documents: self-contained units optimised for print production and linear reading. While effective for large-scale dissemination, the document-centric organisation constrains how knowledge can be structured, updated, shared, and reused. Formal approaches address some of these limitations but struggle to achieve widespread contribution and adoption due to their prioritisation of formal structure over other system properties such as human usability and scope. AI systems are reshaping document production, but without providing a unified portable alternative to traditional documents for humans' expression and exchange of knowledge. This paper presents MMM, a data model for knowledge documentation that emerged from the practical needs of interdisciplinary collaborative research, and positioned here within a comparative analysis of the design space of information systems. MMM combines a small set of normative constraints with the expressive freedom of free-text labels. It is designed for interoperability across disciplines, applications and deployments without requiring semantic convergence. A reference implementation and pilot deployment data demonstrate implementability and early usability.
失敗を安全にする: オープン Web データ収集のための制約付きの検証可能なエージェント フレームワーク
LLM とエージェントは自然言語要件から Web スクレイパーを生成できますが、依存関係エラー、セレクターの破損、スキーマの不一致、および異種ページ構造のため、直接生成の信頼性は依然として低いままです。私たちは、LLM 出力を自由形式のコードから型指定された JSON コレクター構成に移行し、6 種類のコレクター分類、テンプレートとユーティリティ関数の制約、静的な Airflow DAG 実行、ルールベースの品質チェック、構造化フィードバック修正を組み合わせた、制約付きの検証可能なエージェント フレームワークを提案します。 138 のタスクに関する実験では、この分類法が説明ベースの要件の型付けをサポートしていることが示されている一方、安定したインスタンス化には、最初の説明を超えてソース、フィールド、および実行の制約を完了する必要があることが確認されています。独立してソース検証された 80 個のタスク上で、このフレームワークはゼロ実行ステージ LLM トークンと最も短い平均実時間で実行され、適度なワンショット品質と引き換えに、スケジュールされた収集の繰り返しに適した、再利用可能で確定的で検証可能な実行パスを実現します。これらの結果により、このフレームワークは、オープン Web データ収集を繰り返すための再利用可能、低コスト、検証可能な実行パスとして位置付けられます。
原文 (English)
Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection
LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependency errors, broken selectors, schema mismatches, and heterogeneous page structures. We propose a constrained, verifiable agent framework that shifts LLM output from free-form code to typed JSON collector configurations, combining a six-type collector taxonomy, template and utility-function constraints, static Airflow DAG execution, rule-based quality checking, and structured feedback correction. Experiments on 138 tasks show that the taxonomy supports description-based requirement typing, while confirming that stable instantiation requires completing source, field, and execution constraints beyond the initial description. On 80 independently source-verified tasks, the framework runs with zero execution-stage LLM tokens and the lowest average wall-clock time, trading moderate one-shot quality for a reusable, deterministic, and verifiable execution path suited to repeated scheduled collection. These results position the framework as a reusable, low-cost, and verifiable execution path for repeated open-web data collection.
飛行中の航空交通管制をサポートするソリューション空間経路計画
技術の進歩に伴い、航空交通管理用に多くの経路計画アルゴリズムが提案されていますが、戦術管制における運用上の採用は依然として限定的であり、アルゴリズム設計の優先順位と航空管制官のニーズとの間のずれが明らかになりました。これは、本質的に解釈可能で、計算効率が高く、人間が使用するために明示的に設計された意思決定支援ソリューションの必要性を強調しています。この設計課題に焦点を当て、この研究では、次の 2 つの指針となる考慮事項に適合するように設計された、飛行中の航空交通管制 (ATC) のための競合のない経路計画アルゴリズムを開発します。(1) ソリューション空間表示によって提供される解釈可能性と柔軟性。これは、実行可能なすべての安全な行動を明らかにし、変化する最適化目標に対応するアルゴリズムを構築する動機となります。 (2) 決定ロジック コントローラーは、分離基準、操縦性の制限、ウェイポイントの最小化、ルートの実用性などの運用上の制約を強制するときに自然に適用されます。このアルゴリズムは、これらの原則に基づいて、距離ベース、時間間隔ベース、ゾーンベースの 3 つのインテントベースの競合検出方法をソリューション空間フレームワーク内に統合し、計算効率の高い方法で競合のないパスを特定します。さらに、頂点ベースおよびエッジベースの検索ノードが解空間パス プランニング (SSPP) 用に提案されており、その結果、それぞれ SSPPV と SSPPE という 2 つのバリアントが生成され、計算速度と解の品質の観点から評価されます。実験結果によると、ゾーンベースの競合検出と組み合わせた SSPPV は最高のパフォーマンスを実現し、5 nmi グリッドを使用したマーストリヒト上部地域管制センター (MUAC) のデルタ セクターに基づく運用関連シナリオでパスを平均 3.69 ミリ秒で計算します。
原文 (English)
Solution space path planning for supporting en-route air traffic control
As technology advances, many path-planning algorithms have been proposed for Air Traffic Management, yet their operational adoption in tactical control remains limited, revealing a misalignment between algorithmic design priorities and air traffic controllers' needs. This underscores the need for decision-support solutions that are inherently interpretable, computationally efficient, and explicitly designed for human use. Focusing on this design challenge, this study develops a conflict-free path-planning algorithm for en-route Air Traffic Control (ATC) designed to be compatible with two guiding considerations: (1) the interpretability and flexibility offered by solution-space displays, which motivate constructing an algorithm that exposes all feasible safe actions and accommodates shifting optimization goals; and (2) the decision logic controllers naturally apply when enforcing operational constraints, such as separation standards, maneuverability limits, waypoint minimization, and routing practicality. Centered on these principles, the algorithm integrates three intent-based conflict detection methods -- distance-based, time-interval-based, and zone-based -- within a solution-space framework to identify conflict-free paths in computationally efficient ways. Additionally, vertex-based and edge-based search nodes are proposed for solution space path planning (SSPP), resulting in two variants -- SSPPV and SSPPE, respectively, which are evaluated in terms of computational speed and solution quality. Empirical results show that SSPPV paired with zone-based conflict detection achieves the best performance, computing paths in 3.69 ms on average in operational-relevant scenarios based on the Delta sector of the Maastricht Upper Area Control Centre (MUAC) using a 5 nmi grid.
RareDxR1: 人間による注釈を超えた希少疾患診断のための自律的な医学的推論
希少疾患の鑑別診断は重要だが困難な臨床課題であり、医師は複雑で構造化されていない患者の症状から正確な表現型を特定し、広大な検索空間内で複雑な推論を実行する必要がある。しかし、既存の AI アプローチは通常、パイプラインベースの表現型抽出または検索拡張生成に依存しており、事前定義されたオントロジー、検索のボトルネック、診断ロジックの欠如による重大な情報損失に悩まされています。これらの課題に対処するために、非構造化臨床ノートから直接オープンドメインの希少疾患診断を行うために設計された、エンドツーエンドの推論中心の大規模言語モデルである RareDxR1 を導入します。私たちは、知識の内在化と自律進化学習を相乗させて、構造化された表現型やクローズドセットの意思決定への依存を回避することで、進歩的なエンドツーエンドのトレーニング フレームワークを設計します。 RAG と表現型制限の限界を克服するために、断片化された希少疾患の知識をモデルのパラメーターに直接深く取り込むことが可能になりました。さらに、モデル生成と専門家による推論の間のギャップを埋めるために、人間による注釈なしで失敗から学習することで専門家レベルの診断軌跡を合成する戦略である、Reflection-Enhanced Reasoning Sampling (RERS) を提案します。さらに、希少疾患の診断を段階的に習得するための二重レベルのカリキュラム強化学習アプローチを提案します。実験結果は、RareDxR1 がさまざまなベンチマークにわたって最先端の精度を達成し、オープンドメインの希少疾患診断における重要な進歩を示すことを示しています。私たちのコードとデータセットは一般公開されます。
原文 (English)
RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation
Rare disease differential diagnosis is a critical yet arduous clinical task, requiring physicians to identify precise phenotypes from complex, unstructured patient symptoms and execute intricate reasoning within a vast search space. However, existing AI approaches typically rely on pipeline-based phenotype extraction or retrieval-augmented generation, which suffer from critical information loss due to predefined ontologies, retrieval bottlenecks, and a lack of diagnostic logic. To address these challenges, we introduce RareDxR1, an end-to-end reasoning-centric large language model designed for open-domain rare disease diagnosis directly from unstructured clinical notes. We design a progressive end-to-end training framework by synergizing knowledge internalization with autonomous evolutionary learning, thereby bypassing reliance on structured phenotypes and closed-set decision-making. To overcome the limitations of RAG and phenotype restriction, we enabled the deep internalization of fragmented rare-disease knowledge directly into the model's parameters. Moreover, to bridge the gap between model generation and expert reasoning, we propose Reflection-Enhanced Reasoning Sampling (RERS), a strategy that synthesizes expert-level diagnostic trajectories by learning from failures without human annotation. Additionally, we propose a dual-level curriculum reinforcement learning approach for gradually mastering rare disease diagnosis. Experimental results demonstrate that RareDxR1 achieves state-of-the-art accuracy across different benchmarks, marking a significant breakthrough in open-domain rare disease diagnosis. Our code and dataset will be publicly available.
両面の情報非対称性を伴うコンテキストバンディット監視ゲーム
私たちは、個人情報が両方向に流れる場合の AI エージェントの実行時の人間の監視を研究します。人間は自分の報酬関数を非公開で知り、AI は自分が提案するアクションの質を非公開で知っています。これは、人間の監督者が直接評価できない状況を自律ロボットやソフトウェア エージェントが検査したときに自然に生じる一種の非対称性です。協調逆強化学習 (CIRL) と監視ゲームに基づいて、両面非対称情報とプレイ/質問/信頼/監視インターフェイスを備えたコンテキスト バンディット チーム ゲームを導入します。バンディット構造は物理的な状態遷移を除去するため、POMDP の完全な設定では推測の域を出ない正確なワンショットの特徴付けが得られますが、ラウンド全体で動的に制御された状態は維持されると一般的に信じられています。我々は、チーム最適と行動的に自然な近視眼的なルールという 2 つのワンショットの特徴を与えます。そのギャップは、回避可能な危害の板です。AI が提案された行動が有害であり、シャットダウンが助けになることを個人的に認識している領域ですが、近視眼的な人間は、以前の行動を信頼して監督を拒否します。我々は、このギャップが信頼性のない監視コミュニケーションの代償であることを示し、受動的学習と1周期遅れの監視応答による能動的なシグナリングを通じてラウンドを繰り返すうちに、このギャップがどのように動的に解決されるかの部分分析を提供します。
原文 (English)
A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry
We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes. This is the kind of asymmetry that arises naturally when an autonomous robot or software agent has inspected a situation its human supervisor cannot directly assess. Building on Cooperative Inverse Reinforcement Learning (CIRL) and the Oversight Game, we introduce a contextual-bandit team game with two-sided asymmetric information and a play/ask/trust/oversee interface. The bandit structure removes physical state transitions and thereby yields exact one-shot characterizations that would remain conjectural in the full POMDP setting, though the common belief remains a dynamically controlled state across rounds. We give two one-shot characterizations, a team optimum and a behaviorally natural myopic rule, whose gap is a slab of avoidable harm: a region in which the AI privately knows the proposed action is harmful and shutdown would help, yet a myopic human, trusting her prior, declines to oversee. We show this gap is the price of non-credible oversight communication, and give a partial analysis of how it resolves dynamically over repeated rounds through passive learning and active signaling with a one-period-lagged oversight response.
認識論的な AI リテラシーの構築: 学生と AI の共同プログラミングにおける認識論的な目的とプロセスの検出
認識論的思考は、生成人工知能 (GenAI) を適用する際の学生の学習プロセス、特に学習者がクエリを作成し、AI によって生成された出力を評価および検証し、問題解決戦略を調整する必要があるプログラミングのコンテキストにおいて中心的な役割を果たします。この研究では、認識論的 AI リテラシー (EAIL) の概念フレームワークを導入し、AI リテラシーを、さまざまなドメインにわたる人間と AI の動的な相互作用を通じて現れるプロセス指向の認識論的現象として再構成します。この研究では、AIR (認識論的目的、理想、信頼できる認識論的プロセス) フレームワークに基づいて、GenAI がサポートする共同プログラミング活動において認識論的目的と認識論的プロセスがどのように実行されるかを検証し、インタラクション データでこれらの構成要素を運用するためのスケーラブルなアプローチを探ります。この研究では、人間と AI の共同プログラミングの大規模な対話データセットを使用して、認識論的目的 (つまり、習熟指向の目標) と認識論的プロセス (つまり、アウトソーシング、説明の追求、検証の追求、迅速な監視、認識論的正当化) の観察可能な側面を特定します。その結果、EAILの欠如が蔓延しており、学生とGenAIのやりとりの78.8%が非習熟指向の目的や、アウトソーシングや検証の追求など信頼性の低い認識論的戦略に依存していることが明らかになった。逆に、高い認識論的関与を示したインタラクションはわずか 11.1% であり、習得指向の目的が、より信頼性の高い認識論的プロセスにおける認識論的正当化などの高度な認識論的戦略と結びついています。
原文 (English)
Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming
Epistemic thinking plays a central role in students' learning processes when applying generative artificial intelligence (GenAI), particularly in programming contexts where learners must construct queries, evaluate and validate AI-generated outputs, and regulate problem-solving strategies. This study introduces the conceptual framework of Epistemic AI Literacy (EAIL), reframing AI literacy as a process-oriented epistemic phenomenon that emerges through dynamic human-AI interactions across different domains. Drawing on the AIR (epistemic aims, ideals and reliable epistemic processes) framework, this study examines how epistemic aims and epistemic processes are enacted in GenAI-supported co-programming activities and explores scalable approaches for operationalizing these constructs in interaction data. Using a large dialogue dataset of human-AI co-programming, this study identifies observable dimensions of epistemic aims (i.e., mastery-oriented aims) and epistemic processes (i.e., outsourcing, explanation seeking, verification seeking, prompt monitoring, and epistemic justification). The results reveal a prevalent lack of EAIL, with 78.8% of student-GenAI interactions relying on non-mastery-oriented aims and less reliable epistemic strategies like outsourcing and verification-seeking. Conversely, only 11.1% of interactions showed high epistemic engagement, where mastery-oriented aims were coupled with advanced epistemic strategies like epistemic justification in a more reliable epistemic process.
信号から構造へ: メモリ アーキテクチャが LLM エージェントにおける言語の出現をどのように推進するか
2 人のエージェントはどのようにして共有言語をゼロから発明するのでしょうか?ルイス シグナリング ゲームでは、送信者と受信者は対話履歴のみを使用してコードを調整する必要があります。私たちは、LLM エージェントを使用してさまざまなチャネル構成にわたる 5 つのメモリ アーキテクチャを調査し、メモリ アーキテクチャがチャネル容量よりも重要であることを発見しました。永続的なプライベート ノートブックを持つエージェントは、余剰チャネル キャパシティの恩恵を受け、ステートレス エージェントに見られる高キャパシティの崩壊を回避し、最も信頼性の高い調整を実現します (キャパシティ = 25 の場合、$0.867 \pm 0.023$)。ステートレス エージェントは中程度の能力でピークに達し、その後、ローリング コンテキスト ウィンドウが追跡できる範囲を超えて語彙が増加すると低下します。ノートブックは学習した規則を外部化し、エージェントがラウンドごとにコードを再導出する必要から解放されます。情報ボトルネックにヒントを得た議論により、オブジェクトの数に等しい最適な容量が予測されます。むしろ、ボトルネック (容量 = 8) が脆弱点であることが判明し、通常は余剰容量の方が優れています。チャネル容量だけでは調整を予測できないことを示します。メモリ アーキテクチャは、エージェントが対話履歴を安定した規則に変換するかどうかを決定します。信号がどのように言語になるかを理解するには、両方の側面が必要です。
原文 (English)
From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents
How do two agents invent a shared language from scratch? In a Lewis signaling game, a sender and receiver must coordinate on a code using only their interaction history. We study five memory architectures across varying channel configurations with LLM agents and find that memory architecture matters more than channel capacity. Agents with a persistent private notebook benefit from surplus channel capacity and avoid the high-capacity collapse seen in stateless agents, achieving the most reliable coordination ($0.867 \pm 0.023$ at capacity = 25). Stateless agents peak at moderate capacity and then degrade as the vocabulary grows beyond what a rolling context window can track The notebook externalizes learned conventions, freeing agents from having to re-derive codes each round. An information bottleneck-inspired argument predicts an optimal capacity equal to the number of objects. Instead, the bottleneck (capacity = 8) proves to be a fragility point, and surplus capacity is generally better. We show that channel capacity alone cannot predict coordination; memory architecture determines whether agents turn interaction history into stable conventions, and both dimensions are needed to understand how signals become language.
Seed2.0 モデル カード: 現実世界の複雑さのインテリジェンス フロンティアに向けて
私たちは、現実世界の複雑なタスクの解決に向けて有意義な一歩を踏み出すモデル シリーズ、Seed2.0 を紹介します。当社のアプローチは、ユーザーの真のニーズを特定し、これらのニーズと現実的で複雑なシナリオに基づいたベンチマークを選択および抽象化することにより、信頼性が高く将来を見据えた評価システムを構築することから始まります。この評価システムに基づいて、Seed2.0 は、ロングテールの知識と複雑な命令のフォローという 2 つの永続的な課題をターゲットにし、複雑で長期的なタスクにおけるモデルの信頼性を大幅に向上させます。これらに加えて、Seed2.0 は、幅広いユーザー ベースの最も一般的なニーズに対応する、世界をリードする推論インテリジェンス、視覚的理解、および検索機能を提供します。このモデル カードに文書化された広範な現実世界の使用例を通じて、Seed2.0 が初期の複雑な現実世界のタスクを処理する能力を発揮し始め、数億のユーザーに大きな価値を提供できることを実証します。
原文 (English)
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
We present Seed2.0, a model series that takes a meaningful step toward solving complex, real-world tasks. Our approach begins with identifying users' genuine needs and constructing a reliable, forward-looking evaluation system by selecting and abstracting benchmarks grounded in these needs and in realistic, complex scenarios. Guided by this evaluation system, Seed2.0 targets two persistent challenges, long-tail knowledge and complex instruction following, substantially improving the model's reliability on intricate, long-horizon tasks. Beyond these, Seed2.0 delivers world-leading reasoning intelligence, visual understanding, and search capabilities that address the most common needs of a broad user base. Through extensive real-world use cases documented in this model card, we demonstrate that Seed2.0 begins to exhibit the ability to handle initial complex real-world tasks, delivering greater value to hundreds of millions of users.
Mnemosyne: AI 生成のワークフローを検証および修復するためのエージェント トランザクション処理
LLM、ソルバー、エージェント チームはワークフロー アクション、修復、計画を生成することが増えていますが、生成されたアクションは構文的には有効でも、古い、実行不可能、矛盾している、または修復を引き起こした証拠を破壊する可能性があります。エージェントティック トランザクション処理 (ATP) は、生成されたアクションが、宣言された実行可能な制約セット C の下で決定論的な承認を通過するまで、信頼できない提案として扱うトランザクション モデルです。原則は両面的です。提案は真実ではなく、すべての中断を予測する提案はありません。提案は何でも可能ですが、ランタイムのみが承認してコミットします。予期せぬ中断が発生した場合、新しい提案を信頼するのではなく、範囲内で反応的に修復します。 C と比較すると、コミットされた状態の正確性は、提案層の能力、誠実さ、学習とは無関係になります。私たちは、追加専用の遷移ログ、有効な状態の予測、依存関係に安全な補償、およびアクティブなコミットメント レコードを備えたランタイムである Mnemosyne で ATP を実現し、C に関連する 4 つの安全特性 (権限分離、シリアル等価生成許可、証拠保全修復、および義務の封じ込め) を、その局所的修復プロトコル (LCRP) の制限付きリアクティブ修復保証とともに証明します。再現可能なアーティファクトは、9 回の改ざんテストで対象となる違反を拒否しながら、有効な作業を依然として認め、投影と検証のオーバーヘッドは 6% 未満で、制限されたローカル修復編集では、グローバルな再計算よりも桁違いに少ない操作で済みます。 Mnemosyne はオープンソースです: https://github.com/eyuchang/Mnemosyne/tree/arxiv-atp-rq1-rq9b-r8-v2。
原文 (English)
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair. We introduce Agentic Transaction Processing (ATP), a transaction model that treats generated actions as untrusted proposals until they pass deterministic admission under a declared, executable constraint set C. The principle is two-sided: a proposal is not truth, and no proposal foresees every disruption: anything may propose, but only the runtime admits and commits, and when an unforeseen disruption strikes it repairs reactively within bounds rather than trusting a fresh proposal. Relative to C, committed-state correctness becomes independent of the competence, honesty, or learning of the proposing layer. We realize ATP in Mnemosyne, a runtime with an append-only transition log, effective-state projection, dependency-safe compensation, and active commitment records, and prove four safety properties relative to C (authority separation, serial-equivalent generative admission, evidence-preserving repair, and obligation containment) together with a bounded-reactive-repair guarantee for its localized repair protocol (LCRP). A reproducible artifact rejects the targeted violations across nine falsification tests while still admitting valid work, at under 6% projection-and-validation overhead, and bounded local repair edits an order of magnitude fewer operations than global recompute. Mnemosyne is open source: https://github.com/eyuchang/Mnemosyne/tree/arxiv-atp-rq1-rq9b-r8-v2.
実行時の自律管理: シングルおよびマルチエージェントのサイバーフィジカルシステム向けのギアベースの安全性とガバナンス
LLM 駆動のソフトウェア エージェントであれ、ロボット物理エージェントであれ、自律型エージェントは、人間による継続的な監視なしで動作すると、一般的な種類の障害モードに直面します。つまり、未検証のアクションによる安全性違反、制約のないループによる動作の不安定性、未処理のエラー状態による連続性の喪失などです。私たちは、5 つの実行ギア (\Gobs{}、\Gsug{}、\Gplan{}、\Gexec{}、\Gint{}) とユーティリティ ゲート ディスパッチおよびイベント ドリブン フォールバックを組み合わせた離散時間制御システム \system{} を開発しています。単一エージェントのケースでは、単調な安定性、実行の安全性、最終的な安定化、フォールバックの完全性、歯車制約のあるマルコフ決定プロセスとの同等性を証明します。マルチエージェント サイバー物理システム(CPS)の場合、確立された \smart{} 管理自律性ライフサイクルを適用し、実行時の証拠を 4 つのガバナンス状態(\Stable{}/\Meta{}/\Assisted{}/\Regulated{})にマッピングします。コンセンサス ゲーティング、群レベルのリアプノフ解析、エージェントごとのギア権限、およびランデブー制御により、規定された前提条件の下での衝突ゼロを含む、分散型の安全性と安定性の保証が提供されます。 10,000 回のモンテカルロ エピソードにわたる NIST \emph{ロボット アーム位置精度の劣化測定} データセットから校正された故障規模を使用して、3 エージェントの UR5 ロボット アセンブリ セルでの実行時間を評価します。単一エージェントベースラインの異常検出率 2.1\% に対して 99.6\% を達成し、検出遅延を 3.5 倍 $ 削減し、正式な物理作業スペースの安全証明書を提供します。実行ギアは \smart{} ランタイム ガバナンス状態の下でミクロレベルの権限として機能し、アクション制御を自律ガバナンスから分離します。
原文 (English)
Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems
Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous human oversight: safety violations from unverified actions, behavioral instability from unconstrained loops, and continuity loss from unhandled error states. We develop \system{}, a discrete-time control system that combines five execution gears (\Gobs{}, \Gsug{}, \Gplan{}, \Gexec{}, \Gint{}) with utility-gated dispatch and event-driven fallback. For the single-agent case, we prove monotonic stability, execution safety, eventual stabilization, fallback completeness, and equivalence to a gear-constrained Markov decision process. For multi-agent cyber-physical systems (CPS), we apply the established \smart{} managed-autonomy lifecycle and map runtime evidence into its four governance states (\Stable{}/\Meta{}/\Assisted{}/\Regulated{}). Consensus gating, swarm-level Lyapunov analysis, per-agent gear authority, and rendezvous control provide distributed safety and stability guarantees, including zero collision under the stated assumptions. We evaluate the resulting runtime on a three-agent UR5 robotic assembly cell using fault magnitudes calibrated from the NIST \emph{Degradation Measurement of Robot Arm Position Accuracy} dataset across 10,000 Monte Carlo episodes. It achieves a 99.6\% anomaly detection rate versus 2.1\% for the single-agent baseline, reduces detection latency by $3.5\times$, and supplies a formal physical-workspace safety certificate. The execution gears act as micro-level permissions beneath the \smart{} runtime governance states, separating action control from autonomy governance.
逆計画としてのパーソナライゼーション: 構造ノイズ除去によるエージェント的スライド生成のための潜在的な設計意図の学習
スライドのデザインでは、デッキのテーマとページ レイアウトの両方をパーソナライズする必要があります。しかし、現在の AI エージェントベースの手法は、きめ細かいページレベルの設計に苦労しています。事前に指定されたテンプレートやユーザーの詳細な指示のみに依存すると、潜在的なデザイン意図を捉えることができず、ページレベルのスライド パーソナライゼーション (PSP) が未解決のままになります。このギャップを埋めるために、この研究では PSP を逆計画問題として定式化します。使用されている特定の実行ツール (PowerPoint、Beamer など) についての知識を前提とせずに、設計意図を学習することを提案します。ただし、これらのツールの制御を放棄すると、問題はエンドツーエンドの最適化が困難になります。これを克服するために、PSP を近似的に解決するための原則的なフレームワークである SPIRE を提案します。 SPIRE は、クリーンなスライドの視覚構造を意図的に破損することで、破損のノイズを除去するための検証可能なタスクを作成します。これにより、2 人のエージェントが、強化学習 (RL) を通じて実行可能なデザインを協力して改良する方法を学習します。我々は、構造的ノイズ除去がPSPの一貫した代用であること、およびマルチエージェント定式化がRLにおけるポリシー勾配の分散を厳密に低減することの証明を提示する。広範な実験により、SPIRE の優位性が実証されました。
原文 (English)
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page-level Slide Personalization (PSP) unresolved. To close this gap, this work formulates PSP as an inverse planning problem. We propose to learn a design intent without assuming any knowledge of the specific executing tools (e.g., PowerPoint, Beamer) being used. However, relinquishing control over these tools makes the problem intractable to optimize end-to-end. To overcome this, we propose SPIRE, a principled framework to solve PSP approximately. By intentionally corrupting the visual structures of clean slides, SPIRE creates a verifiable task to denoise the corruption, whereby two agents learn to collaboratively refine executable designs via reinforcement learning (RL). We present a proof that structural denoising is a consistent surrogate for PSP, and that the multi-agent formulation strictly reduces policy gradient variance in RL. Extensive experiments demonstrate the superiority of SPIRE.
PHREEQC-MCQ-200: ツール拡張科学シミュレーター エージェントの診断ベンチマーク
大規模な言語モデル エージェントは科学ソフトウェアとの連携がますます高まっていますが、ツールへのアクセスによって科学計算が単に複雑になるだけでなく信頼性が高まるのはいつかはまだ不明です。決定論的な水地球化学シミュレーションでツール拡張エージェントを評価するためのベンチマークである PHREEQC-MCQ-200 を紹介します。このベンチマークには、21 の検証済み PHREEQC シナリオから派生した 200 の多肢選択式の質問が含まれており、エージェントはシミュレーターの入力を構築し、PHREEQC を実行し、構造化された出力を検査し、最終的な回答にコミットする必要があります。複数のフロンティアおよび中間層のモデル ファミリにわたって、シミュレーターへのアクセスにより集計精度が大幅に向上し、多くの科学技術計算タスクには根拠のある実行が必要であることが確認されました。ただし、その増加は単調ではありません。ツールで強化されたエージェントは、ツールなしで正しく回答した項目も失い、平均精度だけでは隠れていた回帰が明らかになります。さらに、出力アクセス プロトコルが重要であることを示します。目次インターフェイスは、より強力なモデルの精度を維持または向上させながらトークン コストを削減できますが、構造化されたシミュレーター出力を確実にナビゲートできない中間層モデルのパフォーマンスは低下します。したがって、PHREEQC-MCQ-200 は、単純なツール呼び出し機能ではなく、科学ツールの使用をエンドツーエンドの診断問題として構成します。私たちは、科学エージェントの評価では、精度だけでなく、項目レベルの保持、出力アクセスの感度、軌道の失敗、および計算チェーンが中断された場所も報告する必要があると主張します。
原文 (English)
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a benchmark for evaluating tool-augmented agents on deterministic aqueous-geochemistry simulations. The benchmark contains 200 multiple-choice questions derived from 21 validated PHREEQC scenarios, requiring agents to construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to final answers. Across multiple frontier and mid-tier model families, simulator access substantially improves aggregate accuracy, confirming that grounded execution is necessary for many scientific-computation tasks. However, the gains are not monotonic: tool-augmented agents also lose items they answered correctly without tools, revealing regressions that average accuracy alone hides. We further show that output-access protocol matters. A table-of-contents interface can reduce token cost while preserving or improving accuracy for stronger models, but it degrades performance for mid-tier models that cannot reliably navigate structured simulator outputs. PHREEQC-MCQ-200 therefore frames scientific tool use as an end-to-end diagnostic problem rather than a simple tool-calling capability. We argue that evaluations of scientific agents should report not only accuracy, but also item-level retention, output-access sensitivity, trajectory failures, and where the computation chain breaks.
Agri-SAGE: コンテキストを認識した農業アドバイザリー生成のためのシミュレーションに基づいたマルチエージェント LLM
農業勧告システムは根本的な緊張に直面しています。静的な農学ガイドラインは、一貫した証拠に基づいた推奨事項を提供しますが、季節ごとの変動性や動的な不確実性には盲目なままです。 LLM を利用した最近の勧告システムは、農学的には信頼できるが生理学的には説得力のない推奨事項を生成するという別のリスクを負います。 Agri-SAGE は、検索に基づいたマルチエージェント LLM 推論と APSIM ベースの生物物理シミュレーションを統合し、農業に関する勧告を生成および検証することで、上記 2 つの制限を解決するように設計された閉ループ フレームワークです。このフレームワークを評価するために、10 年間の遡及分析にわたって、計画と解決、思考のツリー、および反省という 3 つの推論アプローチを評価します。 3 つすべてが静的な PoP (実践パッケージ) ベースラインを大幅に上回り、Tree of Thoughts は印象的なピーク収量を達成しました。同時に、Reflexion は季節を超えたエピソード記憶を活用することで、大幅に低い計算コストで同等の農業成果を達成します。
原文 (English)
Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation
Agricultural advisory systems face a fundamental tension: static agronomic guidelines offer consistent, evidence-based recommendations, yet remain blind to in-season variability and dynamic uncertainties. Recent advisory systems powered by LLMs are liable for a different risk of generating recommendations that are agronomically credible but physiologically unconvincing. Agri-SAGE is a closed-loop framework designed to resolve the above two limitations by integrating retrieval-grounded multi-agent LLM reasoning with APSIM-based biophysical simulation, to generate and validate agronomic advisories. To assess this framework, we evaluate three reasoning approaches, namely Plan-and-Solve, Tree of Thoughts, and Reflexion, over a 10-year retrospective analysis. All three significantly outperform static PoP (Package-of-Practice) baselines, with Tree of Thoughts achieving impressive peak yields. At the same time, Reflexion achieves comparable agronomic outcomes at substantially lower computational cost by leveraging cross-seasonal episodic memory.
進化する環境における身体化エージェントの世界モデルのマルチスケール混合
現実世界で活動する身体化されたエージェントは、状況の変化に応じてマルチスケールの推論と知識の適応を必要とします。この設定に専門家混合 (MoE) を適用する際の 2 つの課題を特定します。1 つは、ルーティングに明確な規模の概念が欠けており、特定の規模での対象を絞った更新が妨げられること、そして、統一された更新ポリシーでは、各規模での知識が時代遅れになるさまざまな速度に対応できないことです。私たちは、スケールを意識した世界モデルの混合と進化を通じて両方の課題に対処するフレームワーク、MuSix を紹介します。 2 段階のルーティング メカニズムは、解釈レベル理論に触発された状況の新規性の尺度である経験的距離に基づくスケール選択を根拠とします。まず、メタルーターがこの量を連続スケール空間上の重みにマッピングし、次にスケールごとのベース ルーターが特定されたスケール内のワールド モデルを選択します。適応に関しては、スケール依存の忘却率により、高スケールの抽象化が持続しながら低スケールの知識が迅速に更新され、ゲートされたスケール間転送により階層全体の一貫性が維持されます。 EmbodiedBench と HAZARD での実験では、MuSix がマルチスケール推論と動的適応に関して最先端のベースラインよりも改善していることが示されています。
原文 (English)
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Embodied agents operating in the real world require multi-scale reasoning and knowledge adaptation as conditions change. We identify two challenges in applying Mixture of Experts (MoE) to this setting: routing lacks an explicit notion of scale, preventing targeted updates at specific scales, and a uniform update policy cannot accommodate the different rates at which knowledge at each scale becomes outdated. We present MuSix, a framework that addresses both challenges through scale-aware world model mixture and evolution. A two-stage routing mechanism grounds scale selection in experiential distance, a measure of situational novelty inspired by Construal Level Theory: a meta-router first maps this quantity to a weight over continuous scale space, then per-scale base routers select world models within the identified scale. For adaptation, scale-dependent forgetting rates allow low-scale knowledge to refresh rapidly while high-scale abstractions persist, and gated inter-scale transfer maintains coherence across the hierarchy. Experiments on EmbodiedBench and HAZARD show that MuSix improves over state-of-the-art baselines on multi-scale reasoning and dynamic adaptation.
AI ネイティブ ゲーム: 調査とロードマップ
生成 AI により、ゲームが実行時に対話、クエスト、キャラクター、画像、世界を生成できるようになりました。ただし、生成だけではゲームが AI ネイティブになるわけではなく、プレイアビリティが保証されるわけでもありません。この論文では、ランタイム生成 AI がコア ループを構成するかどうかによって AI ネイティブ ゲームを定義します。AI コンポーネントが削除されたり、些細な置き換えが行われたりすると、プレイの中心的な形式が崩壊するか、根本的に異なるものになります。この反事実基準は、AI ネイティブ ゲームを、AI 拡張ゲーム、境界アーティファクト、チャットボット、居酒屋スタイルのロールプレイ、手続き型コンテンツ生成、および AI 支援制作から区別します。この定義を使用して、候補となる成果物をスクリーニングし、公開されている 53 の AI ネイティブ ゲームとプロトタイプを分析します。二重軸の G/N 分類法を導入します。G 軸はプレイヤーが直面するゲームのタイプを捉え、N 軸は生成 AI をプレイに不可欠なものにする支配的な AI メカニズムを捉えます。このコーパスは言語優先デザイン、特に物語的冒険、認識論的相互作用、生成的物語を中心に集中していますが、意味論的判断、マルチエージェント シミュレーション、生成的構築、関係/仲間遊びなどのカテゴリーはあまり代表されていません。私たちは、設計上の中心的な問題は、セマンティックなオープン性を安定したゲームプレイに組み込むことであると主張します。 AI ネイティブの設計は、目標、ルール、状態、フィードバック、ペーシング、プレーヤーの主体性などの機械的不変条件に依存しており、これらにより、オープンエンドの AI 出力が解釈可能で結果的なものになります。最後に、制御可能な発電、機械としての AI 設計、マルチモーダルおよびマルチエージェント システム、推論の経済学、評価、安全性、規制のロードマップを示します。
原文 (English)
AI Native Games: A Survey and Roadmap
Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime. Yet generation alone does not make a game AI-native, nor does it guarantee playability. This paper defines AI-native games by whether runtime generative AI is constitutive of the core loop: if the AI component were removed or trivially replaced, the central form of play would collapse or become fundamentally different. This counterfactual criterion separates AI-native games from AI-augmented games, boundary artifacts, chatbots, tavern-style role-play, procedural content generation, and AI-assisted production. Using this definition, we screen candidate artifacts and analyze 53 publicly available AI-native games and prototypes. We introduce a dual-axis G/N taxonomy: the G-axis captures player-facing game type, while the N-axis captures the dominant AI mechanic that makes generative AI indispensable to play. The corpus is concentrated around language-forward designs, especially narrative adventure, epistemic interaction, and generative narrative, while categories such as semantic adjudication, multi-agent simulation, generative construction, and relationship/companion play remain less represented. We argue that the central design problem is organizing semantic openness into stable gameplay. AI-native design depends on mechanical invariants: goals, rules, state, feedback, pacing, and player agency that make open-ended AI outputs interpretable and consequential. We conclude with a roadmap for controllable generation, AI-as-mechanic design, multimodal and multi-agent systems, inference economics, evaluation, safety, and regulation.
HARC: 堅牢な安全調整のための有害性と拒否のカップリングの方向性
アライメントされた LLM が内部的にどのように安全性を表すかを理解することは、ジェイルブレイクが成功する理由を説明し、堅牢なアライメント戦略の設計に情報を提供するため、アライメントの脆弱性を診断するために重要です。これまでの研究では、整列された LLM がプロンプト側のトークン位置で残留ストリーム内の分離可能な方向として有害性と拒否をエンコードしていることが示されています。トークンが生成される前に拒否または有害性の方向を抑制することで、ジェイルブレイクがプロンプトエンコーディングで成功し、異なる攻撃クラスが有害性と拒否の面の分離可能な領域を占めることを示します。分析をレスポンス トークンの位置まで拡張すると、プロンプト側で入力を有害なものとして認識できなかった場合でも、モデルが有害なコンテンツを生成中にそのコンテンツを認識することがわかりました。私たちの発見に動機づけられて、私たちは、プロンプトポジションとレスポンスポジションの両方で2つの方向をペアにする微調整方法であるHARC(有害性と拒否のカップリング)を紹介します。介入は有害性拒否部分空間に限定されるため、残りのストリームの残りの部分はそのまま残り、一般的な能力を低下させたり、過剰な拒否を拡大したりすることはありません。広範な実験を通じて、HARC は、主要なトレーニング時間と推論時間の安全性手法にわたる 6 つのベースラインの中で最も強力な堅牢性、機能、使用性のトレードオフを達成しました。プロンプトおよびレスポンスの位置における有害性と拒否の指示は、アーキテクチャ固有の調整を行わずにテストした 5 つのモデル ファミリと 2 つのスケールに渡って伝達されます。
原文 (English)
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.
ワールド モデリング エージェントのベンチマーク フレームワークとしての AGI Maze
大規模言語モデル (LLM) は強力なパターン補完システムですが、そのデフォルトの動作モード (静的コンテキストから次のトークンを予測する) では、外部世界の永続的で操作可能な表現が確実に生成されません。環境が部分的に観察可能でステートフルになり、隠れた状態についての記憶と構造化された仮説が必要になると、テキストでは「推論」のように見える多くのタスクが大幅に難しくなります。 AGI Maze は、高次元の感覚入力を必要とせずにそのような環境を構築するための軽量フレームワークです。これは、クリーンな API と複数の難易度レジームを備えた一連のグリッドベースの迷路タスクを提供します。目標は、エージェントがすぐに提供される観測値からローカル ルールを推測するだけでなく、世界状態の表現を学習して使用する必要があるベンチマークを作成することです。単純な迷路に関するいくつかのバニラ LLM の初期評価を提供し、LLM 推論時に内部的に迷路を表現できないことを示しています。また、ベースライン エージェントも導入します。これにより、メッセージ履歴を作業メモリとして使用して、エージェントの実行時に観察の記述を構築できます。これによりパフォーマンスは向上しますが、LLM エージェントが人間にとって十分なステップ バジェット内で小さな迷路であっても確実に解決するにはまだ不十分です。
原文 (English)
AGI Maze as a Benchmark Framework for World-Modeling Agents
Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an external world. Many tasks that look like "reasoning" in text become substantially harder once the environment is partially observable, stateful, and requires memory and structured hypotheses about hidden state. AGI Maze is a lightweight framework for building such environments without requiring high-dimensional sensory inputs. It provides a family of grid-based maze tasks with a clean API and multiple difficulty regimes. The goal is to create benchmarks where agents must learn and use world state representations, not just infer a local rule over readily provided observations. We provide an initial evaluation of several vanilla LLMs on simple mazes showing that they fail to represent mazes internally at LLM inference time. We also introduce a baseline agent, which is allowed to use its message history as a working memory to construct descriptions of observations at agentic runtime. Although this can improve performance, it is still insufficient for an LLM agent to reliably solve even small mazes within a step budget that is more than enough for humans.
インタラクティブなゲームプレイのためのコーチング可能なエージェント
強化学習は、高度な AI およびロボット システムの作成における貴重なツールであることが証明されており、ゲームプレイからロボット工学、基礎モデルに至るまであらゆるものに貢献しています。通常、これらの AI システムは、試行錯誤を通じて、タスクを解決するために最適に近い 1 つの動作を学習します。ただし、タスクの解決方法に関して、できればリアルタイムで、ある程度の制御を主張したいユースケースは数多くあります。コアタスクのこれらの変更をスタイルと呼びます。私たちは、ユニバーサル価値関数近似器 (UVFA) を、慎重に選択されたトレーニング シナリオ、学習アルゴリズム、データ拡張と組み合わせて、複雑な領域でスタイルを示すエージェントをコーチングするためのフレームワークを作成します。私たちは、AAA ビデオ ゲームの Horizon Forbidden West と Gran Turismo、およびオープンソースのヒューマノイド テスト ドメインでのフレームワークのアプリケーションを実証します。カーレース、様式化されたゲーム戦闘、人型歩行など、ドメインの性質が異なるにもかかわらず、各エージェントは、そのドメインの主なタスクを満たしながら、スタイルの要求に強い一貫性を示します。重要なのは、このホワイト ペーパーで概説した手法を使用すると、エンド ユーザーが実行時に最終的な動作を選択できるため、最終的に実行されるパフォーマンスを柔軟に制御できることです。
原文 (English)
Coachable agents for interactive gameplay
Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework's application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains -- car racing, stylized game combat, and humanoid walking -- each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.
Self-GC: Long-Horizon LLM エージェントの自己統治コンテキスト
長期的な LLM エージェントは、使い捨てのテキスト サフィックスとして扱うには構造化されすぎているツールの結果、ファイル、プラン、およびユーザー制約を蓄積します。現在のシステムは主に、時系列プルーニングやツール出力マスキングなどの実行中のヒューリスティック、またはコンテキスト制限に近い最終的な自己要約に依存しています。ヒューリスティックは低コストですが、将来の依存関係を考慮しません。概要は物語の状態を保持しますが、多くの場合、正確な証拠、ロケーター、編集可能な成果物が隠されます。私たちは Self-GC を紹介します。GC はガベージ コレクションを意図的にエコーしながら自己統治コンテキストを表します。システムは単に未使用のトークンを再利用するだけでなく、エージェント コンテキスト オブジェクトのライフサイクルを管理します。自己 GC は、ユーザーのターン、ツール スパン、およびスキルの状態をインデックス付きオブジェクトに変換します。サイドチャネルプランナーにフォールド、マスク、プルーニングのアクションを提案するよう依頼します。また、ハーネスは回復可能なサイドカー、安全なコミット境界、キャッシュを認識したコミットを強制できます。 33 セッションのハード セットでは、セルフ GC はプレフィックス トークンの 43.95% をプルーニングし、将来の継続の 84.85% には影響を与えません。これに対し、ヒューリスティック ベースラインの影響なし率は 54.55% ~ 69.70% です。 332 セッションの本番派生スイートでは、3 つのプランナー バックボーンは 91.27% ~ 94.58% の影響なし率に達しますが、ベースラインは 77.71% ~ 87.46% のままです。運用環境では、オンラインのアカウントレベルの分割により、日中の平均入力トークンが 10% ~ 15% 削減され、ピーク削減率は 20% 近くになります。これらの結果は、事後のテキスト クリーンアップではなく、インデックス付きの回復可能なオブジェクトに対する実行時のライフサイクル制御としてのコンテキスト管理を示しています。
原文 (English)
Self-GC: Self-Governing Context for Long-Horizon LLM Agents
Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. Current systems mostly rely on in-run heuristics such as chronological pruning and tool-output masking, or on final self-summary near a context limit. Heuristics are cheap but blind to future dependencies; summaries preserve narrative state but often hide exact evidence, locators, and editable artifacts. We present Self-GC, where GC denotes self-governing context while deliberately echoing garbage collection: the system does not merely reclaim unused tokens, but governs the lifecycle of agent context objects. Self-GC turns user turns, tool spans, and skill state into indexed objects; asks a side-channel planner to propose fold, mask, and prune actions; and lets the harness enforce recoverable sidecars, safe commit boundaries, and cache-aware commit. On a 33-session Hard Set, Self-GC prunes 43.95% of prefix tokens while leaving 84.85% of future continuations unaffected, compared with no-impact rates of 54.55% to 69.70% for heuristic baselines. On a 332-session production-derived suite, three planner backbones reach no-impact rates of 91.27% to 94.58%, while baselines remain at 77.71% to 87.46%. In production, an online account-level split reduces daytime average input tokens by 10% to 15%, with peak reductions near 20%. These results point to context management as runtime lifecycle control over indexed, recoverable objects rather than post hoc text cleanup.
いつでも有効な証明書を備えた自己進化エージェント
自己進化するエージェントは、ほとんどの学習理論的保証の背後にある仮定に違反します。つまり、データ、評価器、コンポーネント、仮説空間は、更新されるポリシーによって生成されます。私たちは \textbf{SEA} を提案します。このアーキテクチャは、自己変更を小さなステアリング アダプターと \emph{frozen} ベース モデルの周りのバージョン管理されたハーネスに限定し、固定エラー バジェットに対して監査可能な証明書を発行する常時有効なゲートを介してのみ各変更を許可します。 5 つのループ コントローラーが公開保証を構成します。そのようなゲートは、フリーズされたベースがすでに生成している動作の中から \emph{選択} することしかできないため、5 つの検証者インザループ メカニズム (best-of-$N$、マイクロステップ検索、自己作成複製オラクル、検索層制御、自己修復) が、問題テキストのみから計算された、ゲートが必要とする密度の高いグレーダーフリーの信号を供給します。 4 つの基本モデルにわたる $52$ インスタンスの SWE ベンチ検証済みサブセットでは、基本機能が支配的で交絡のない効果であり、2 つの強力な基本モデルでは、意図的な no-op-composite 制御により $+4$ と $+5$ (\textsc{Glm}~5.2 $24\to28$; \textsc{Gpt} $29\to34$、$65\%$) でスイートの寄与が分離されます。最良)、そのメカニズムが起動して回帰を防止していることを確認するイベント ログが含まれます。結果は高価な評価では 1 回実行されます。実行ごとの差異の確認とタスクごとのアルゴリズムの組み合わせの適応は今後の課題です。
原文 (English)
Self-Evolving Agents with Anytime-Valid Certificates
Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis space are produced by the policy being updated. We present \textbf{SEA}, an architecture that confines self-modification to a small steering adapter and a versioned harness around a \emph{frozen} base model and admits each modification only through an anytime-valid gate that emits an auditable certificate against a fixed error budget. Five loop controllers compose published guarantees; because such gates can only \emph{select} among behaviors the frozen base already produces, five verifier-in-the-loop mechanisms -- best-of-$N$, micro-step search, self-authored reproduction oracles, search-layer control, and self-repair -- supply the dense, grader-free signal the gates require, computed from the issue text alone. On a $52$-instance SWE-bench Verified subset across four base models, base capability is the dominant, confound-free effect, and on two strong base models a deliberate no-op-composite control isolates the suite's contribution at $+4$ and $+5$ (\textsc{Glm}~5.2 $24\to28$; \textsc{Gpt} $29\to34$, the $65\%$ best), with event logs confirming that its mechanisms fire and prevent regressions. Results are single-run on expensive evaluations; confirming run-to-run variance and adapting the per-task algorithm mix are future work.
2 つの AI 指標の相違: それはすべての違いを生むのでしょうか?
コンピューティングの指数関数的なスケーリングが続くにつれて、フロンティア AI モデルの機能は、開発者が少ない固定予算でアクセスできる機能を超えるでしょうか?それとも機能は「地球を継承する柔和なモデル」に収束するのでしょうか? Gundlach らを基に構築(2025b) では、答えは AI の能力をどのように評価し測定するかによって決まることを示しています。従来のパフォーマンス尺度について議論し、検証損失ではギャップが縮小している一方、他の指標ではフロンティア モデルのリードが永遠に拡大していることを示します。パフォーマンス メトリクスをトレーニング (および推論) コンピューティングに関連した関数形式で分類することで、どのメトリクスが柔和なモデルを好むかを決定するための厳密な数学的条件を提供し、制限されたパフォーマンス メトリクスが常にそうなることを示します。ただし、パフォーマンス メトリクスを慎重に解釈することが不可欠です。多くの一般的な制限付きメトリクスには、密接に関連する制限のない対応するメトリクスがある (またその逆も同様) ことを示します。ドメイン内の適切なメトリクスを決定することは、ポリシーの前提条件です。これは、制限されたメトリクスと制限されていないメトリクスが、反対のポリシー応答を示唆する可能性があるためです。ソフトウェアエンジニアリング、合成生物学、修辞的説得力などの特定の能力が、私たちが関心を持っている用語で測定したときに制限がない場合、フロンティアレベルの能力は少数の裕福な関係者の手に集中する可能性があります。逆に、その能力が制限されている場合、フロンティアレベルの能力は柔和なモデルを通じて多くの人の手に渡ります。
原文 (English)
Two AI Metrics Diverged: Will it Make All the Difference?
As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget? Or will capabilities converge, with "meek models inheriting the earth"? Building on Gundlach et al. (2025b), we show that the answer depends on how we value and measure AI capabilities. We discuss conventional performance measures and show that, while validation loss shows a shrinking gap, on other metrics frontier models grow their lead forever. Classifying performance metrics by their functional forms in relation to training (and inference) compute, we provide tight mathematical conditions for determining which metrics favor meek models, and show that bounded performance metrics always do. But careful interpretation of performance metrics is essential: we show that many common bounded metrics have closely-related counterpart metrics that are unbounded (and vice versa). Determining the apt metric in a domain is a prerequisite for policy, since bounded and unbounded metrics may suggest opposing policy responses. If a particular capability -- like software engineering, synthetic biology, or rhetorical persuasiveness -- is unbounded when measured in the terms we care about, frontier-level capability will likely be concentrated in the hands of a few wealthy actors. Conversely, if that capability is instead bounded, frontier-level capabilities proliferate through meek models into the hands of the many.
グラフネイティブの強化学習により、概念の組み換えによる追跡可能な科学的仮説の生成が可能になります
材料発見を加速するには、複数ステップのドメインに基づいた推論を通じて科学的に有効な仮説を生成できる AI システムが必要です。標準的な大規模言語モデルは、多くの場合、オープンエンドの材料設計問題に対して流暢ではあるものの追跡可能性が低い応答を生成するため、最終的な答えが一貫した中間推論によってサポートされているかどうかを判断することが困難になります。私たちは、グループ相対ポリシー最適化 (GRPO) で微調整されたグラフネイティブ推論モデルのファミリーである Graph-PRefLexOR を開発し、推論をメカニズム探索、グラフ構築、パターン抽出、仮説合成の明示的なフェーズに編成します。この設計は、ニューラル言語の生成を記号関係構造とリンクさせ、因果関係の構築、検査、再利用を可能にします。材料科学および力学の文献からの自由回答形式の質問 100 件について、Graph-PRefLexOR は、対応する基本モデルと比較して 40 ~ 65% の改善を達成し、推論のトレーサビリティが最大に向上しました。埋め込み分析では、ベースラインよりも広範な意味の探索と約 2 ~ 3 倍の意味の多様性が示されます。セマンティックバックトラッキングとレイヤーごとの隠れ状態分析により、構造化された推論と最終的な答えの間のより強力な整合性がさらに示されます。最後に、テスト時のグラフ拡張により、追加のコンピューティングによって、単に意味論的範囲が拡大されるのではなく、主に、制限された意味論的空間内での長距離概念の組み換えが増加することが明らかになりました。これらの結果は、材料設計やその他の科学的応用における科学的仮説生成のための解釈可能な AI システムへの道筋として、グラフネイティブ強化学習を確立します。
原文 (English)
Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making it difficult to determine whether final answers are supported by coherent intermediate reasoning. We develop Graph-PRefLexOR, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization (GRPO) to organize reasoning into explicit phases for mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis. This design links neural language generation with symbolic relational structure, enabling causal connections to be constructed, inspected, and reused. On 100 open-ended questions from materials science and mechanics literature, Graph-PRefLexOR achieves 40-65% improvements over corresponding base models, with the largest gains in reasoning traceability. Embedding analyses show broader semantic exploration and approximately 2-3 times greater semantic diversity than baselines. Semantic backtracking and layer-wise hidden-state analyses further show stronger alignment between structured reasoning and final answers. Finally, test-time graph expansion reveals that additional compute primarily increases long-range conceptual recombination within a bounded semantic space, rather than simply expanding semantic coverage. These results establish graph-native reinforcement learning as a pathway toward interpretable AI systems for scientific hypothesis generation in materials design and other scientific applications.
エージェントティック RAG パイプラインのベイジアン不確実性伝播: マルチホップ質問応答に関する概念実証研究
Agentic Retrieval-Augmented Generation(RAG)システムの信頼できる展開には、多段階の推論パイプラインがいつ失敗するかを推定するメカニズムが必要です。この論文では、プランナー、評価器、およびジェネレーターの各ステージが、セマンティックな発散とジェネレーターの自己評価から得られる不確実性信号を生成する、不確実性を認識したエージェント検索拡張生成 (RAG) フレームワークを紹介します。これらの信号はベイジアン ネットワーク (BN) を介して伝播され、システム レベルの不確実性が推定され、ワークフロー全体にわたる潜在的な障害点のノード レベルの指標が提供されます。このアプローチは、GPT-3.5-Turbo および GPT-4.1-Nano を使用して、StrategyQA および HotpotQA で評価されます。受信機動作特性曲線下面積 (AUROC)、精度除去曲線下面積 (AUARC)、予想される校正誤差 (ECE)、およびブライアー スコアを使用して、識別、選択的予測、および校正を評価します。結果は、HotpotQA ではベイジアン伝播がより効果的であることを示しています。HotpotQA では、マルチホップ推論段階にわたって不確実性が蓄積されますが、StrategyQA では、誤ったキャリブレーションと信頼性の低いアップストリーム信号によって引き起こされる制限が明らかになります。この研究では、ベイジアン不確実性伝播を、Agentic RAG システムを監視するための有望ではあるが予備的なメカニズムとして位置づけており、洋上風力発電 (OSW) の保守意思決定サポートなどの産業分野では将来の検証が必要です。
原文 (English)
Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering
Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasoning pipelines may fail. This paper presents an uncertainty-aware Agentic Retrieval-Augmented Generation (RAG) framework in which planner, evaluator and generator stages produce uncertainty signals derived from semantic divergence and generator self-evaluation. These signals are propagated through a Bayesian Network (BN) to estimate system-level uncertainty and provide node-level indicators of potential failure points across the workflow. The approach is evaluated on StrategyQA and HotpotQA using GPT-3.5-Turbo and GPT-4.1-Nano, with Area Under the Receiver Operating Characteristic Curve (AUROC), Area Under the Accuracy-Rejection Curve (AUARC), Expected Calibration Error (ECE), and Brier Score used to assess discrimination, selective prediction and calibration. Results show that Bayesian propagation is more effective on HotpotQA, where uncertainty accumulates across multi-hop reasoning stages, while StrategyQA exposes limitations caused by miscalibration and unreliable upstream signals. The study positions Bayesian uncertainty propagation as a promising but preliminary mechanism for monitoring Agentic RAG systems, with future validation required in industrial domains such as Offshore Wind (OSW) maintenance decision support.
PedNStream: 歩行者交通管理のためのスケーラブルなネットワーク フロー シミュレーション
大規模な群衆管理には、計算効率が高く、フィードバック ベースの制御と互換性のある歩行者シミュレーションが必要です。ただし、ほとんどのオープンソース ツールは微細なものであるか、ネットワーク規模の閉ループ評価用に設計されていません。このペーパーでは、リンク伝送モデル (LTM) に基づいた巨視的な歩行者ネットワーク負荷のためのオープンソースの Python ネイティブ シミュレーターである PedNStream (Pedestrian Network Flow Simulation) について説明します。このフレームワークは、拡散と活動に起因する変動を捉える確率的リンクダイナミクスを組み込むことで LTM ベースの歩行者モデルを拡張し、動的なユーザーの平衡ルート選択を、不確実で介入主導の設定に適したユーティリティベースの定式化に置き換えます。 PedNStream は、ゲート、フロー分離、ルート ガイダンスなどの介入用のコントローラー インターフェイスが組み込まれたモジュール式フレームワークとして実装されています。私たちはフレームワークを段階的に評価します。合成シナリオでは、キューの形成、スピルバック、輻輳の解消、適応型再ルーティングなどの主要なメカニズムを検証します。実際のネットワーク実験では、大規模な行動と観察された歩行者数との一貫性を評価します。閉ループのケーススタディではコントローラーの統合を実証し、ランタイム分析ではスケーラビリティを定量化します。これらの結果により、PedNStream は大規模な歩行者ネットワークのシミュレーションと制御のための効率的かつ実用的なテストベッドとして確立されます。
原文 (English)
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
Large-scale crowd management requires pedestrian simulations that are both computationally efficient and compatible with feedback-based control. However, most open-source tools are either microscopic or not designed for network-scale closed-loop evaluation. This paper presents PedNStream (Pedestrian Network Flow Simulation), an open-source, Python-native simulator for macroscopic pedestrian network loading based on the Link Transmission Model (LTM). The framework extends LTM-based pedestrian models by incorporating stochastic link dynamics that capture diffusion and activity-induced variability, and replaces dynamic user equilibrium route choice with a utility-based formulation suited to uncertain, intervention-driven settings. PedNStream is implemented as a modular framework with built-in controller interfaces for interventions such as gating, flow separation, and route guidance. We evaluate the framework in a staged manner. Synthetic scenarios verify key mechanisms, including queue formation, spillback, congestion dissipation, and adaptive rerouting. Real-network experiments assess large-scale behavior and consistency with observed pedestrian counts. A closed-loop case study demonstrates controller integration, and a runtime analysis quantifies scalability. These results establish PedNStream as an efficient and practical testbed for large-scale pedestrian network simulation and control.
決定論的で自己拡張的な反応分類のための検証可能なルールをエージェント的に生成
コンピューター支援合成計画では、各変換に決定論的で解釈可能なラベルを割り当てる反応ルールの大きなライブラリを使用して、標的分子をアクセス可能な前駆体に分割します。しかし、化学はロングテールであるため、手動エンコーディングが困難であり、既存のツールは新しい化学に適応できない固定ルールセットに依存しています。ここでは、大規模言語モデル (LLM) のマルチエージェント フレームワークが反応を分類し、665,901 件の米国特許反応にわたってルール自体を記述し、コーパスに対してテストする検証ループの下で各ルールを生成する、完全に自動化されたパイプラインを紹介します。人間によるキュレーションを行わずに、標準分類を 68 クラスから 14,073 クラスに拡張します。軽量の指紋分類器を使用して、目に見えない反応の 97.7% を分類し、主要な独自の分類器と一致しながら、化学をより細かく分解し、オンデマンドでトレーニング ディストリビューション外の化学に拡張します。その結果、生きた反応性データベースと、生成モデルを信頼性の高い自己拡張型のシンボリック システムに変えるための一般的なルートが得られます。
原文 (English)
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encoding intractable, and existing tools rely on fixed rulesets that cannot adapt to new chemistries. Here we present a fully automated pipeline in which a multi-agent framework of large language models (LLMs) classifies reactions and writes the rules themselves across 665,901 US patent reactions, generating each rule under a verification loop that tests it against the corpus. It expands a standard taxonomy from 68 to 14,073 classes without human curation. With a lightweight fingerprint classifier, it classifies 97.7\% of unseen reactions, matching a leading proprietary classifier while resolving chemistry more finely and extending on demand to chemistry outside its training distribution. The result is a living reactivity database and a general route to turning generative models into reliable, self-expanding symbolic systems.
エージェントはオープンワールドに一般化できますか?ツール使用における静的トレーニングの脆弱性を明らかにする
Large Language Model (LLM) エージェントは静的ベンチマークでは熟練していることを示しますが、実際のシナリオでの展開は、ユーザー クエリ、ツール セット、対話ダイナミクスの動的な性質によって妨げられます。この一般化のギャップに対処するために、クエリ、アクション、観察、およびドメインの次元にわたる分布の変化を特徴とする問題設定である OpenAgent (オープンワールドのツール使用エージェント) を形式化します。その影響を体系的に診断するために、制御されたサンドボックス環境を構築し、知覚、相互作用、推論、内部化の 4 層階層にわたるきめ細かい環境変化を定義し、包括的な一連の実験を実施します。私たちの分析により、一連の重要な洞察が得られ、教師あり微調整 (SFT) と強化学習の両方で訓練されたエージェントは、オープンな環境の変化に直面したときに、さまざまな程度のパフォーマンス低下に悩まされることが実証されました。これらの洞察に基づいて、我々は、現実の環境におけるエージェントの堅牢性と有用性を高めるための基礎を築く、SFT の外乱ベースの介入戦略である摂動拡張微調整を提案します。コードは https://github でリリースされます。 com/LAMDA-NeSy/OpenAgent。
原文 (English)
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize OpenAgent (Tool-Use Agent in Open-World), a problem setting characterized by distributional shifts across query, action, observation, and domain dimensions. To systematically diagnose its impact, we construct a controlled sandbox environment where we define fine-grained environmental shifts across a four-tier hierarchy, Perception, Interaction, Reasoning, and Internalization, and conduct a comprehensive series of experiments. Our analysis yields a series of key insights, demonstrating that agents trained via both Supervised Fine-Tuning(SFT) and Reinforcement Learning suffer from varying degrees of performance degradation when confronting open environmental shifts. Building on these insights, we propose Perturbation-Augmented Fine-Tuning, a disturbance-based intervention strategy for SFT that lays the foundation for enhancing agent robustness and utility in realistic environments. Our code will be released at: https://github. com/LAMDA-NeSy/OpenAgent.
自律型ラボオーケストレーターの最適なリソース利用
自律型研究所では、AI エージェントが次に行うべき実験のバッチを提案します。ただし、利用可能なリソースを最大限に活用してこれらのタスクを計画および実行することは、まったく別の問題です。これは、実際のハードウェアの制約に対処する場合、特に容量やスループットが異なる複数の機器がある場合には困難になる可能性があります。ここでは、金属有機フレームワーク合成用の自律プラットフォームのリソース利用に対処するための 2 段階の方法を示します。まず、制約プログラミングを使用して最適なスケジュールを見つけます。これにより、ハードウェアの制限と容量を満たしながら、合計時間を最小限に抑えるスケジュールが見つかります。次に、各タスクのステータス依存関係のシステムを使用し、最適なスケジュールを確実に実行できるようにします。
原文 (English)
Optimal Resource Utilization for Autonomous Laboratory Orchestrators
In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing those tasks taking full advantage of the available resources is a completely different question. This can be challenging when dealing with real-world hardware constraints, especially so when there are multiple instruments with different capacities and throughputs. Here we demonstrate a 2-step method to address resource utilization for our autonomous platform for metal-organic framework synthesis. First, we use constraint programming to find optimal schedules. This finds schedules that minimizes the total time while still satisfying the limitations and capacities of the hardware. Secondly, we use a system of status dependencies for each task, which allows for the robust execution of the optimal schedules.
Theoria: 非公式推論状態に対する書き換え許容性の検証
AI システムの答えを信頼できるのはどのような場合ですか?形式的証明アシスタントは確実性を提供しますが、問題分布のほとんどには到達できません。スカラー LLM ジャッジはカバレッジを提供しますが、事後的に監査できない不透明なスコアを生成し、他の LLM と同じ一貫性の問題にさらされます。私たちは、このギャップを埋める検証アーキテクチャである Theoria を紹介します。候補解は、型指定された状態遷移のシーケンスに書き換えられます。各状態遷移は、引用、計算、または問題によって与えられた事実など、明示的な正当化によってライセンスされ、すべての遷移は独立して監査可能です。基本的な不変条件は変化の完全性です。連続する証明状態間のすべての違いを考慮する必要があるため、隠れた前提は黙って通過するのではなく、許可されていない突然変異として表面化します。 HLE-Verified Gold (185 のテキストのみのエキスパートの問題) では、Theoria は 91.4% の厳密な精度で 105 を認定しています (Wilson 95% CI [84.5%、95.4%])。すべての認証では、人間が判読できる証明トレースが生成され、各ステップに個別にチャレンジできます。ホリスティック LLM ジャッジは、一致するカバレッジでは同等の精度を達成しますが、別の問題 (Jaccard 0.14 ~ 0.36) では失敗するため、アプローチは補完的になります。 15 のドメインにわたる 95 件の敵対的毒物証明について、構造化された裁判官は 94.7% を捕捉したのに対し、総合的な判断では 83.2% を捕捉しました (p= 0.0017)。全体の 11.5 pp のギャップは、隠れた前提 (90.6% 対 62.5%、28 pp の差) と捏造された引用 (100% 対 90%) に集中しており、形式的な分析が利点を予測するエラー クラスです。利点が予測されない算術および定理の誤用エラーのパフォーマンスは同じです。 GPQA ダイヤモンド (n= 65) では、認定精度は 97.1% (Wilson CI [85.1%、99.5%]) です。
原文 (English)
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each licensed by an explicit justification, whether that be a citation, computation, or problem-given fact, and every transition is independently auditable. The foundational invariant is completeness of change: every difference between consecutive proof states must be accounted for, so hidden premises surface as unlicensed mutations rather than passing silently. On HLE-Verified Gold (185 text-only expert problems), Theoria certifies 105 at 91.4% strict precision (Wilson 95% CI [84.5%, 95.4%]). Every certification produces a human readable proof trace in which each step can be independently challenged. Holistic LLM judges achieve comparable precision at matched coverage but fail on different problems (Jaccard 0.14-0.36), making the approaches complementary. On 95 adversarial poisoned proofs across 15 domains, structured judges catch 94.7% versus 83.2% for holistic judging (p= 0.0017). The overall 11.5 pp gap concentrates in hidden premises (90.6% vs. 62.5%, a 28 pp difference) and fabricated citations (100% vs. 90%), the error classes where the formal analysis predicts an advantage; performance is identical on arithmetic and theorem-misapplication errors, where no advantage is predicted. On GPQA Diamond (n= 65), certified precision is 97.1% (Wilson CI [85.1%, 99.5%]).
AutoMem: 認知スキルとしての記憶の自動学習
記憶の専門知識は学習されたスキルです。何をエンコードするか、いつ取得するか、知識をどのように整理するかを知っています。この能力は認知科学でメタ記憶として知られています。私たちはメモリ管理を訓練可能なスキルとして扱うことで、この視点を LLM にもたらします。ファイル システム操作をタスク アクションと並行してファーストクラスのメモリ アクションに昇格させ、メモリの管理方法をモデル自体に決定させます。この記憶スキルは、それをサポートする構造 (プロンプト、ファイル スキーマ、アクション語彙)、およびそれを実行するモデルの熟練度という 2 つの軸に沿って向上します。どちらの軸も手動による最適化に抵抗します。長期的なタスクのエピソードは何千ステップにもわたって実行され、単一の記憶ミスが表面化するずっと前に隠れてしまう可能性があるため、完全な軌跡を人間がレビューすることは非現実的です。両方の軸を自動化するフレームワークである AutoMem を紹介します。最初のループでは、強力な LLM がエージェントの完全な軌跡をレビューし、エージェントがメモリ ファイルと対話する方法を形成するメモリ構造を繰り返し修正します。 2 番目のループでは、エージェント自身の記憶力の良い判断が多くのエピソードから特定され、モデルの記憶能力を直接高めるためのトレーニング信号として使用されます。プロシージャル生成された 3 つのロングホライズン ゲーム (Craafter、MiniHack、NetHack) にわたって、モデルのタスク アクション動作を変更せずにメモリのみを最適化することで、ベース エージェントのパフォーマンスが最大 2 倍から 4 倍向上し、32B のオープンウェイト モデルが Claude Opus 4.5 や Gemini 3.1 Pro Thinking などのフロンティア システムと競合できるようになりました。私たちの結果は、メモリ管理は独立して学習可能なスキルであり、長期的なタスクで大きな利益をもたらす高いレバレッジの目標であることを示しています。
原文 (English)
AutoMem: Automated Learning of Memory as a Cognitive Skill
Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive science as metamemory. We bring this perspective to LLMs by treating memory management as a trainable skill. We promote file-system operations to first-class memory actions alongside task actions, letting the model itself decide how to manage its memory. This memory skill improves along two axes: the structure that supports it (prompts, file schemas, action vocabulary), and the proficiency of the model exercising it. Both axes resist manual optimization: episodes in long-horizon tasks run for thousands of steps, and a single memory mistake can hide long before it surfaces, making human review of full trajectories impractical. We introduce AutoMem, a framework that automates both axes. In the first loop, a strong LLM reviews complete agent trajectories and iteratively revises the memory structure that shapes how the agent interacts with its memory files. In the second loop, the agent's own good memory decisions are identified from many episodes and used as training signal to sharpen the model's memory proficiency directly. Across three procedurally generated long-horizon games (Crafter, MiniHack, and NetHack), optimizing memory alone--without modifying the model's task-action behavior--improved the base agent's performance ~2x-4x, bringing a 32B open-weight model competitive with frontier systems such as Claude Opus 4.5 and Gemini 3.1 Pro Thinking. Our results show that memory management is an independently learnable skill, and a high-leverage objective yielding large gains on long-horizon tasks.
UltraFlux: 多様なアスペクト比にわたる高品質のネイティブ 4K テキストから画像への生成のためのデータ モデルの共同設計
拡散トランスは最近、解像度 1K 付近で強力なテキストから画像への生成を実現していますが、拡散トランスをさまざまなアスペクト比でネイティブ 4K に拡張すると、位置エンコーディング、VAE 圧縮、最適化にまたがる密結合の障害モードが明らかになることを示します。これらの要因のいずれかに個別に取り組むことで、実質的な品質が確保されます。したがって、データモデルの共同設計の観点を採用し、MultiAspect-4K-1M 上の 4K でネイティブにトレーニングされた Flux ベースの DiT、UltraFlux を導入します。これは、制御されたマルチ AR カバレッジ、バイリンガル キャプション、および解像度と AR を意識したサンプリングのための豊富な VLM/IQA メタデータを備えた 1M 画像 4K コーパスです。モデル側では、UltraFlux は、(i) Resonance 2D RoPE を YaRN と組み合わせて、4K でのトレーニング ウィンドウ、周波数、AR 対応の位置エンコーディングを実現します。 (ii) 4K 再構築の忠実度を向上させる、シンプルで非敵対的な VAE ポストトレーニング スキーム。 (iii) タイムステップおよび周波数帯域全体で勾配のバランスを再調整する SNR 対応の Huber Wavelet 目標。 (iv) 段階ごとの美的カリキュラム学習戦略。これは、事前モデルによって支配される高ノイズのステップに高度な美的監督を集中させます。これらのコンポーネントを組み合わせることで、安定したディテールを維持した 4K DiT が生成され、幅広、正方形、高さの AR 全体に汎用化されます。 4096 ベンチマークおよびマルチ AR 4K 設定での Aesthetic-Eval では、UltraFlux は、忠実度、美的感覚、アライメント メトリクス全体で強力なオープンソース ベースラインを常に上回っており、LLM プロンプト リファイナーにより、独自の Seedream 4.0 と同等またはそれを上回っています。
原文 (English)
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect ratios exposes a tightly coupled failure mode spanning positional encoding, VAE compression, and optimization. Tackling any of these factors in isolation leaves substantial quality on the table. We therefore take a data-model co-design view and introduce UltraFlux, a Flux-based DiT trained natively at 4K on MultiAspect-4K-1M, a 1M-image 4K corpus with controlled multi-AR coverage, bilingual captions, and rich VLM/IQA metadata for resolution- and AR-aware sampling. On the model side, UltraFlux couples (i) Resonance 2D RoPE with YaRN for training-window-, frequency-, and AR-aware positional encoding at 4K; (ii) a simple, non-adversarial VAE post-training scheme that improves 4K reconstruction fidelity; (iii) an SNR-Aware Huber Wavelet objective that rebalances gradients across timesteps and frequency bands; and (iv) a Stage-wise Aesthetic Curriculum Learning strategy that concentrates high-aesthetic supervision on high-noise steps governed by the model prior. Together, these components yield a stable, detail-preserving 4K DiT that generalizes across wide, square, and tall ARs. On the Aesthetic-Eval at 4096 benchmark and multi-AR 4K settings, UltraFlux consistently outperforms strong open-source baselines across fidelity, aesthetic, and alignment metrics, and-with a LLM prompt refiner-matches or surpasses the proprietary Seedream 4.0.
DigitalCoach: 人間とエージェントのコンピューター使用コーチングにおけるコミュニケーションとグラウンディングのギャップ
エージェントはソフトウェア タスクを自動化できるようになってきていますが、エージェントは人間にソフトウェアの使い方を自分で教えることができるのでしょうか? DigitalCoach は、5 つのソフトウェア アプリケーションにわたる 28.1 時間の画面および入力イベントの記録に基づいた、22,752 の対話ターンで構成される 72 人の専門家と初心者のコンピューター使用コーチング セッションのマルチモーダル データセットです。私たちは DigitalCoach を使用して、最先端のモデルが人間にコンピューターの使い方を教えることができるかどうかを評価します。自動評価では、モデルは指導方法が人間とは異なることが示されています。モデルはより直接的な指示を提供しますが、説明、エラー診断、知識確認の質問は少ないです。コーチング手法を修正すると、モデルは人間のリファレンスに似た発話を生成しますが、視覚的なコンテキストにあまり基づいていません。インタラクティブな評価では、モデルコーチは学習者に深い関与を行わずに受動的に指示に従うようにさせ、視覚的な基礎付けが不十分であることを確認しています。 DigitalCoach は、共同的かつ積極的にコンピュータを使用してコーチング エージェントを行うための基盤を築きます。
原文 (English)
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sessions consisting of 22,752 dialogue turns grounded in 28.1 hours of screen and input event recordings across five software applications. We use DigitalCoach to evaluate whether state-of-the-art models can teach humans how to use computers. Automated evaluation shows that models differ from humans in how they coach: models provide more direct instructions, but fewer explanations, error diagnoses, and knowledge-check questions. When we fix the coaching method, models produce utterances similar to human references yet poorly grounded in visual context. Interactive evaluation confirms that model coaches cause learners to passively follow instructions without deeper engagement and fall short in visual grounding. DigitalCoach lays a foundation for collaborative and proactive computer use coaching agents.
パーソナルナレッジグラフの「文字列」から「モノ」へ: レコメンデーションシステムのための LLM トリプル抽出の評価
パーソナル ナレッジ グラフ (PKG) は、ユーザーの好みをモデル化するためのプライバシー保護フレームワークを提供しますが、構造化されていない分散型の会話データからそれを構築することは依然として課題です。この論文では、軽量の大規模言語モデルを使用して構造化されたユーザー設定トリプルを抽出するための再現可能なパイプラインを提示することで、会話的な「文字列」と意味論的な「もの」の間のギャップを埋めます。 PKG 構築のための会話データから Wikidata 識別子にリンクされた RDF 準拠のトリプルを抽出する機能について、Qwen および Gemma ベースのモデルを評価します。私たちの評価では、意味抽出の忠実度と、下流の推奨タスクにおける結果のグラフの有用性の両方が評価されます。特定のモデルのパフォーマンスが良好で、トリプル抽出パフォーマンスに比例して下流パフォーマンスが高いことがわかりました。
原文 (English)
From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems
Personal Knowledge Graphs (PKGs) offer a privacy-preserving framework for modeling user preferences, yet constructing them from unstructured, decentralized conversational data remains a challenge. This paper bridges the gap between conversational "strings" and semantic "things" by presenting a reproducible pipeline for extracting structured user-preference triples using lightweight Large Language Models. We evaluate Qwen- and Gemma-based models on their ability to extract RDF-compliant triples linked to Wikidata identifiers from conversational data for PKG construction. Our evaluation assesses both the semantic extraction fidelity and the utility of the resulting graphs in a downstream recommendation task. We found that certain models performed well and had proportionally high downstream performance relative to their triple extraction performance.
高度なエンコーダがスパース検索で遅れるのはなぜですか?語彙のギャップを埋めるための答えとアプローチ
ModernBERT のような高度な基盤モデルは、密な検索では古いアーキテクチャを大幅に上回っていますが、学習されたスパース検索 (LSR) では、古い BERT ベースのベースラインに驚くほど遅れをとっています。根本原因は \textit{語彙ギャップ} であると特定します。最新のトークナイザーは、可逆再構成用に設計された生の大文字と小文字を区別する語彙を利用しており、単一の意味単位を冗長な表面形式にマッピングし、形態学的ノイズでモデル容量を浪費し、語彙一致を妨げます。我々は理論的枠組みを通じてこの直観を形式化し、意味論的な整合性が保たれる限り、適切な語彙の粗視化が仮説クラスの複雑さを軽減して一般化限界を狭めることができることを実証した。これを解決するために、私たちは \textbf{Vocabulary Transfer (VT)} を提案します。これは、高度なエンコーダを最小限の計算コストでスパースに適した正規化された語彙に移行する、モデルに依存しないフレームワークです。 VT は、空間トポロジーを介した新しい \textbf{セマンティック初期化} を利用して幾何学的構造を保存し、 \textbf{活性化電位キャリブレーション (APC)} メカニズムを利用して、事前学習された多様体をスパース性制約に合わせて調整し、標準的な微調整で観察される死んだニューロンや密集した崩壊を防ぎます。経験的に、VT は普遍的に効果的です。これにより、ModernBERT が BEIR ベンチマークで最先端のパフォーマンス (\textbf{52.4} nDCG、\textbf{+4.7} の改善) を達成できるようになり、RoBERTa-large のような失敗したモデルが復活し、推論不要のアーキテクチャと特殊なドメインにシームレスに一般化されます。これらの結果は、パフォーマンスの遅れがアーキテクチャ上の欠陥ではなく、解決可能な語彙の不一致であることを裏付けています。コードとモデルをリリースしました。\footnote{https://anonymous.4open.science/r/vocab-transfer/。すべての詳細が含まれています。}
原文 (English)
Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps
While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the aging BERT-base baseline in learned sparse retrieval (LSR). We identify the root cause as the \textit{Vocabulary Gap}: modern tokenizers utilize raw, case-sensitive vocabularies designed for lossless reconstruction, which map single semantic units to redundant surface forms, wasting model capacity on morphological noise and hindering lexical matching. We formalize this intuition through a theoretical framework, demonstrating that appropriate vocabulary coarse-graining can tighten the generalization bounds by reducing complexity of the hypothesis class, provided that semantic integrity is preserved. To resolve this, we propose \textbf{Vocabulary Transfer (VT)}, a model-agnostic framework that migrates advanced encoders to sparse-friendly, normalized vocabularies with minimal computational cost. VT utilizes a novel \textbf{Semantic Initialization} via spatial topology to preserve geometric structure and an \textbf{Activation Potential Calibration (APC)} mechanism to align pre-trained manifolds with sparsity constraints, preventing the dead neuron and dense collapse observed in standard fine-tuning. Empirically, VT is universally effective: it enables ModernBERT to achieve state-of-the-art performance on the BEIR benchmark (\textbf{52.4} nDCG, a \textbf{+4.7} improvement), resuscitates failing models like RoBERTa-large, and generalizes seamlessly to inference-free architectures and specialized domains. These results confirm that the performance lag is not an architectural deficiency but a solvable vocabulary mismatch. We've released our code and models.\footnote{https://anonymous.4open.science/r/vocab-transfer/. All details included.}
トポロジカルボイド解析 知識空間における体系的な技術革新発見のための数学的フレームワーク
オペレーティング システムやハードウェア/ソフトウェアの共同設計など、高密度の技術領域でどこを革新するかを特定することは、基本的に高次元の知識空間での検索問題です。既存のアプローチは、キーワード検索、引用の近接性、または人間の直感に依存しており、いずれも、対象の目標に同時に関連し、かつ従来技術にはない未踏の領域の概念を形式化するものではない。我々は、密-疎ハイブリッド埋め込み空間においてトポロジカルボイドをトライアド(A、B、C)として定義する数学的フレームワークであるトポロジカルボイド解析(TVA)を紹介します。ボイドには 3 つの条件が必要です。(i) 概念 A と概念 B の両方がドメイン アンカー C と意味的に一貫している。 (ii) それらのペアごとの類似性は、校正された限界帯域内に収まり、明らかな組み合わせと無関係なノイズの両方を回避します。 (iii) それらはスパース語彙ブリッジを共有しますが、埋め込まれた超球上の測地線の中点は占有されていません。 TVA は、約 140,000 のインデックス付きドキュメントに適用され、96 のターゲットにわたって 2,128 の発明候補を生成します。 90% が自動品質フィルタリングを生き残り、4 人の専門家による敵対的レビューから 191 件の改訂と 1 件の承認の評決が得られました (0.05% がエンドツーエンド)。 2 つのケーススタディでは、フレームワークが単に明らかな関連ペアではなく、非明らかな結合組織を表面化していることを示しています。
原文 (English)
Topological Void Analysis A Mathematical Framework for Systematic Technical Innovation Discovery in Knowledge Spaces
Identifying where to innovate in a dense technical domain - such as operating systems or hardware/software co-design - is fundamentally a search problem in a high-dimensional knowledge space. Existing approaches rely on keyword search, citation proximity, or human intuition, none of which formalise the notion of an unexplored region that is simultaneously relevant to a target goal and absent from prior art. We present Topological Void Analysis (TVA), a mathematical framework that defines topological voids as triads (A, B, C) in a dense-sparse hybrid embedding space. A void requires three conditions: (i) both concepts A and B are semantically cohesive with domain anchor C; (ii) their pairwise similarity falls within a calibrated marginality band - avoiding both obvious combinations and unrelated noise; and (iii) they share a sparse lexical bridge while the geodesic midpoint on the embedding hypersphere is unoccupied. Applied to ~140k indexed documents, TVA generates 2,128 invention candidates across 96 targets; 90% survive automated quality filtering, yielding 191 REVISE and 1 APPROVE verdict from four-specialist adversarial review (0.05% end-to-end). Two case studies demonstrate the framework surfaces non-obvious connective tissue rather than merely obvious related pairs.
根拠のないペルソナ: 体制依存と LLM の個別化問題
LLM の個別化問題に対する Beckmann & Butlin (2026) の存在論的フレームワークは、ペルソナ ベクトルの文献から議論の余地のない領域間共参照の仮定を継承しています。つまり、同じ方向が、プロンプト コンディショニング、勾配降下微調整、および推論時間ステアリングの下で同じ内容を選択するというものです。 Qwen3-4B-Instruct および Mistral-7B-Instruct-v0.2 でのペルソナ トポロジー実験から得られた 4 つの経験的なウェッジを紹介します。プロンプト抽出されたベクトルと微調整盆地の非共線性です。架空のペルソナは、実際のアンカーよりも強く実際のアンカーの方向に沿ってモデルを移動させます。トレーニング履歴によって決定されるアトラクターに偏った、矛盾した原子価の混合物。そして、推論時の算術トレーニングと微調整時のキメラトレーニングにおける非対称構成代数 - これらは共同して仮定を台無しにします。私たちは、レジームインデックス付きの個性化を提案します。表現コンテンツのアイデンティティ単位は、ビークル単体ではなく、(ビークル、レジーム)のペアです。この枠組みの下では、ベックマンとバトリンの 3 つの候補的立場は、同じ指示対象をめぐって競合するのではなく、3 つの異なる体制内部の対象を記述しています。同じ診断がモロ&ミリエレ、チャルマーズ、セルロにも当てはまります。
原文 (English)
Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem
Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumption from the persona-vectors literature: that the same direction picks out the same content under prompt-conditioning, gradient-descent fine-tuning, and inference-time steering. We present four empirical wedges from persona-topology experiments on Qwen3-4B-Instruct and Mistral-7B-Instruct-v0.2 - non-collinearity of prompt-extracted vectors and fine-tune basins; fictional personas displacing the model along real-anchor directions more strongly than real anchors do; contradictory-valenced mixtures biased toward a training-history-determined attractor; and asymmetric compositional algebra under inference-time arithmetic versus fine-tune-time chimera training - that jointly undermine the assumption. We propose regime-indexed individuation: the identity unit for representational content is a (vehicle, regime) pair, not a vehicle alone. Under this framework, Beckmann & Butlin's three candidate positions describe three different regime-internal objects rather than competing for the same referent; the same diagnosis applies to Mollo & Milli\`ere, Chalmers, and Cerullo.
BaRA: BFS およびリフレクション Web データ収集エージェント
大規模言語モデル (LLM) ベースの Web エージェントは、Web データ収集のための手動スクリプトを削減しますが、実際の Web サイトでは、関連するページを見逃したり、不完全なマルチモーダル出力を返したり、直接ダウンロードできないメディア URL を返したりすることがよくあります。固定インタラクション予算の下でサイトレベルの収集を行うためのフレームワークである BFS-and-Reflection Agent (BaRA) を紹介します。このフレームワークは、境界付き幅優先検索 (BFS) トラバーサルと履歴ベースの自己反映を組み合わせています。私たちは、グラウンドトゥルースの参照セットを使用して、50 の合成 Web サイトで BaRA を評価します。さらに、乱雑なレイアウトまたは動的なレイアウトを備えた 3 つの公開 Web サイトでもテストしました。 BaRA は、リンク検出とダウンロード可能なマルチモーダル抽出において Pure LLM、SeeAct-Vision、およびブラウザでの使用を上回り、ダウンロード有効な画像とビデオの回復において最大の利益をもたらします。私たちのコードは https://github.com/MLAI-Yonsei/BaRA-Agent で入手できます。
原文 (English)
BaRA: BFS-and-Reflection Web Data Collection Agent
Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant pages, return incomplete multimodal outputs, or return media URLs that are not directly downloadable. We present BFS-and-Reflection Agent (BaRA), a framework for site-level collection under a fixed interaction budget. The framework combines bounded breadth-first search (BFS) traversal with history-based self-reflection. We evaluate BaRA on 50 synthetic websites with ground-truth reference sets. We additionally test on three public websites with cluttered or dynamic layouts. BaRA outperforms Pure LLM, SeeAct-Vision, and Browser-use on link discovery and downloadable multimodal extraction, with the largest gains in download-valid image and video recovery. Our code is available at https://github.com/MLAI-Yonsei/BaRA-Agent.
SchemaRAG: LLM 駆動の構造化情報抽出のための動的な大規模スキーマ削減
ターゲットのスキーマが大きくて複雑な場合、大規模言語モデル (LLM) を使用して非構造化テキストから構造化データを抽出することは困難になります。このような場合、プロンプトに完全なスキーマを含めると、コストと待ち時間が増加し、中間損失のパフォーマンス低下の危険性があり、コンテキストの長さの制限を超える可能性があります。私たちは、スキーマ メタデータと利用可能な場合は少数のショットの例を活用することで、スキーマ条件付き情報抽出タスクの出力スキーマ空間を動的にプルーニングする検索拡張生成 (RAG) フレームワークである SchemaRAG を提案します。実際の医療および電子商取引のデータセットで SchemaRAG を評価します。私たちの結果は、SchemaRAG が micro-F1 で最大 8.8% の増加、レイテンシーの 47% 削減、トークン コストの 48% 削減を達成できることを示しており、大規模なスキーマ抽出に対する実用性を実証しています。
原文 (English)
SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction
Extracting structured data from unstructured text using large language models (LLMs) becomes challenging when target schemas are large and complex. In such cases, including the full schema in the prompt increases cost and latency, risks lost-in-the-middle performance degradation, and can exceed context length limits. We propose SchemaRAG, a retrieval-augmented generation (RAG) framework that dynamically prunes the output schema space for schema-conditioned information extraction tasks by leveraging schema metadata and few-shot examples when available. We evaluate SchemaRAG on real-world healthcare and e-commerce datasets. Our results show that SchemaRAG can achieve up to an 8.8% increase in micro-F1, a 47% reduction in latency, and a 48% reduction in token costs, demonstrating its practicality for large-schema extraction.
強化されたライティング支援のための制御可能なナラティブ レンダリング
基本的なライティング支援における大規模言語モデル (LLM) の優れた能力にもかかわらず、創造的なライティングにおける LLM の有用性は、永続的なバイナリ障害によって根本的に妨げられています。この問題は、修復研磨と呼ばれる安全な表面レベルの編集と、破壊的で制御されていないプロット拡張の間の揺れとして現れます。このジレンマは、物語の忠実さと説明の強度との間の重要なトレードオフを定義します。我々は、物語と言説の間の物語学的区別に基づいた執筆支援フレームワークである Loom を提案します。 Loom は、意図を中心とした記号論的思考連鎖を運用する 3 層パイプラインを採用し、物語の意図とレンダリング密度を正確に制御します。このアーキテクチャは、知覚マテリアルの生成を構文の挿入から分離し、元のイベント構造に違反することなく強化が確実に行われるようにします。 LLM ベースの指標と人間による評価を含む当社の包括的な評価は、Loom がこの根本的な緊張をうまく解決していることを示しています。 Loom は最高の総合品質スコアを達成し、最先端のベースラインと比較して事実の完全性と説明の強度が大幅に向上しました。
原文 (English)
Controllable Narrative Rendering for Enhanced Assisted Writing
Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary failure. This issue manifests as an oscillation between safe, surface-level editing, referred to as remedial polishing, and destructive, uncontrolled plot expansion. This dilemma defines a critical trade-off between narrative fidelity and descriptive intensity. We propose Loom, an assisted writing framework grounded in the narratological distinction between story and discourse. Loom employs a three-layer pipeline that operationalizes an intent-centered semiotic chain-of-thought to enforce precise control over narrative intent and rendering density. This architecture separates the generation of perceptual material from syntactic insertion, ensuring that enhancement occurs without violating the original event structure. Our comprehensive evaluation, which includes LLM-based metrics and human assessment, demonstrates that Loom successfully resolves this fundamental tension. Loom achieves the highest overall quality score, yielding substantial gains in factual integrity and descriptive intensity compared to state-of-the-art baselines.
会話型レコメンダー システムにおけるユーザー シミュレーションの即時最適化: 多目的フレームワーク
会話型レコメンダー システム (CRS) は、ユーザーが積極的に好みを引き出し、意図を明確にし、レコメンデーションをリアルタイムで適応できるため、次世代のインテリジェント レコメンダー システムの中核コンポーネントです。ただし、CRS ドメインには、評価とトレーニング データへのアクセスという 2 つの重要な障害があります。実際の人を対象とした研究を通じて CRS を評価することは、従来のレコメンダー システムよりも重要ですが、そのような研究には費用も時間もかかります。さらに、プライバシー上の懸念により、CRS インタラクション データをモデル トレーニング用に取得するのが困難なことがよくあります。大規模言語モデル (LLM) ベースのユーザー シミュレーターは、評価とトレーニング用の合成ユーザー インタラクションを生成することで両方の課題に対処することが期待されています。しかし、既存のアプローチは体系的なポジティブバイアス、データ漏洩、行動の多様性の制限に悩まされており、広範なドメイン専門知識を必要とする脆弱な手動プロンプトエンジニアリングに依存しています。この論文では、CRS の LLM ベースのユーザー シミュレーターのプロンプトを自動的に最適化し、同時にこれらの問題を軽減するフレームワークを提案します。実験結果は、提案されたフレームワークが、さまざまなプロンプト設定におけるベースライン手法と比較して、人間の対話パターンとの行動の整合性を向上させることを示しています。
原文 (English)
Prompt Optimization for User Simulation in Conversational Recommender Systems: A Multi-Objective Framework
Conversational recommender systems (CRSs) are a core component of next-generation intelligent recommender systems because they enable users to actively elicit preferences, clarify intentions, and adapt recommendations in real time. However, there are two key obstacles in the CRS domain: evaluation and access to training data. Evaluating CRSs through real human studies is more critical than for traditional recommender systems, yet such studies are both costly and time-consuming. Moreover, CRS interaction data are often difficult to obtain for model training due to privacy concerns. Large language model (LLM)-based user simulators have shown promise in addressing both challenges by generating synthetic user interactions for evaluation and training. However, existing approaches suffer from systematic positive bias, data leakage, and limited behavioral diversity, and they rely on brittle manual prompt engineering that requires extensive domain expertise. In this paper, we propose a framework to automatically optimize prompts for LLM-based user simulators in CRSs, simultaneously mitigating these issues. Experimental results demonstrate that the proposed framework achieves improved behavioral alignment with human interaction patterns compared to baseline methods across diverse prompt settings.
SkillSelect-Serve: 小規模 LLM エージェント向けの予算管理可能で QoS を意識したスキル サービスの推奨と構成
再利用可能なスキル ライブラリは、大規模言語モデル (LLM) エージェントにとって重要なインフラストラクチャになりつつありますが、既存の選択方法では、スキルを取得可能なドキュメントとして扱い、固定の上位 K リストを返すことがよくあります。この文書では、エージェントのスキル選択をスキル サービスの推奨および構成として定式化する、予算管理可能で QoS を意識したフレームワークである SkillSelect-Serve について説明します。 SkillSelect-Serve は、機能の説明、依存関係、コンテキスト コスト、リスク、QoS 関連の属性を備えた構造化されたスキル サービスとして生のスキルを表します。ローカルの Micro-Agent Requirement Planner が自然言語タスクを構造化されたサービス要件に変換し、共有ディスカバリー バックボーンが大規模なレジストリから候補サービスを取得します。次に、このフレームワークは、スキルレベルの限界適合性推定と、カバレッジ、冗長性、コスト、およびリスクのトレードオフに関するバンドルレベルの調整を行う二重粒度ユーティリティモデリングを実行します。 35,353 のスキルと 586 のタスク クエリに関する実験では、SkillSelect-Serve が固定の上位 K 取得ベースラインと比較して、同一予算バンドルの再現率と平均ユーティリティを一貫して向上させることが示されています。
原文 (English)
SkillSelect-Serve: Budget-Controllable and QoS-Aware Skill Service Recommendation and Composition for Small LLM Agents
Reusable skill libraries are becoming important infrastructure for large language model (LLM) agents, yet existing selection methods often treat skills as retrievable documents and return fixed top-k lists. This paper presents SkillSelect-Serve, a budget-controllable and QoS-aware framework that formulates agent skill selection as Skill Service Recommendation and Composition. SkillSelect-Serve represents raw skills as structured Skill Services with functional descriptions, dependencies, context cost, risk, and QoS-related attributes. A local Micro-Agent Requirement Planner converts natural-language tasks into structured service requirements, while a shared discovery backbone retrieves candidate services from a large registry. The framework then performs dual-granularity utility modeling with skill-level marginal suitability estimation and bundle-level calibration for coverage, redundancy, cost, and risk trade-offs. Experiments on 35,353 skills and 586 task queries show that SkillSelect-Serve consistently improves same-budget bundle recall and mean utility over fixed top-k retrieval baselines.
PRA-RAG: 検索の破損に対する検索拡張生成における堅牢な集約が証明されています
検索拡張生成 (RAG) は、外部知識を組み込むことで大規模言語モデル (LLM) を強化し、固有の知識制限を効果的に軽減します。ただし、RAG は、取得したテキストを操作してモデルの出力を誤解させるポイズニング攻撃に対して依然として脆弱です。既存の防御メカニズムには理論的な堅牢性の保証が欠けていることが多く、LLM が取得したコンテンツについての知識が限られている場合には、動作の信頼性が低くなります。この研究では、取得されたテキストに対するポイズニング攻撃を防御するように設計された堅牢な検索集約アルゴリズムである PRA-RAG を提案します。 PRA-RAG は、取得したテキストの複数の組み合わせをサンプリングし、埋め込み空間の幾何学的構造を利用して堅牢なサブセットを特定し、そこから安定した集約表現を導き出します。私たちは、ポイズニングされた取得コンテンツの最大の影響に関する理論的限界を提供し、RAG の堅牢性の定量的尺度を確立します。複数のベンチマークと RAG アーキテクチャにわたる実験では、PRA-RAG が 71% の精度を維持しながら攻撃の成功率を 1% まで低下させ、代表的な最先端の手法を大幅に上回るパフォーマンスを示していることが実証されています。
原文 (English)
PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, effectively mitigating their inherent knowledge limitations. However, RAG remains vulnerable to poisoning attacks that manipulate retrieved texts to mislead model outputs. Existing defense mechanisms often lack theoretical robustness guarantees and perform unreliably when the LLM has limited knowledge of the retrieved content. In this work, we propose PRA-RAG, a provably robust retrieval aggregation algorithm designed to defend against poisoning attacks on retrieved texts. PRA-RAG samples multiple combinations of retrieved texts and utilizes geometric structures in the embedding space to identify a robust subset, from which a stable aggregated representation is derived. We provide theoretical bounds on the maximum impact of poisoned retrieved content and establish a quantitative measure of RAG's robustness. Experiments across multiple benchmarks and RAG architectures demonstrate that PRA-RAG reduces the attack success rate to as low as 1% while maintaining an accuracy of 71%, significantly outperforming representative state-of-the-art methods.
GRACE-RAG: クローズドドメインの機関設定での軽量展開を可能にする、正規証拠合成のためのガバナンド検索アーキテクチャ
検索拡張生成(RAG)システムは、回答を信頼できる文書に基づいて行う必要がある組織の質問応答環境で広く使用されています(Gao et al.、2023)。関連情報が異種文書に分散されているエンティティ密度の高いドメインでは、ベクトルのみの検索では断片的な証拠が生成されることが多く、推論時間推論への依存度が高まります (Zhao et al., 2024)。この論文では、生成段階から構造化された検索層まで構造的推論を外部化する、検索主導のグラフ拡張型 RAG アーキテクチャである GRACE-RAG を紹介します。これにより、オフラインで構造的曖昧性が解決され、クローズド ドメインの制度用語に合わせて調整された自己ホスト型の軽量モデルへの展開が可能になります。 Mistral 24B、GPT OSS 120B、Gemini 2.5 Flash の 3 つのモデル容量にわたる実験では、完全性、深度、予測カバレッジにおいて一貫した改善が見られ、中規模モデルでは全体的な品質が最大 20% 向上しました。これは、取得アーキテクチャがモデル スケール全体にわたる構造品質を支配し、独自のシステムに依存することなく計算量と待ち時間のフットプリントを削減することを示しています。
原文 (English)
GRACE-RAG: Governed Retrieval Architecture for Canonical Evidence Synthesis, Enabling Lightweight Deployment in Closed-Domain Institutional Settings
Retrieval-Augmented Generation (RAG) systems are widely used in institutional question answering settings where responses must be grounded in authoritative documentation (Gao et al., 2023). In entity-dense domains where relevant information is distributed across heterogeneous documents, vector-only retrieval often produces fragmented evidence and increases dependence on inference-time reasoning (Zhao et al., 2024). This paper introduces GRACE-RAG, a retrieval-governed, graph-augmented RAG architecture that externalizes structural reasoning from the generative stage to a structured retrieval layer, resolving structural ambiguity offline, enabling deployment on self-hosted lightweight models calibrated to closed-domain institutional vocabulary. Experiments across three model capacities: Mistral 24B, GPT OSS 120B, and Gemini 2.5 Flash show consistent improvements in completeness, depth, and anticipatory coverage, with overall quality gains of up to 20% under mid-scale models, indicating that retrieval architecture governs structural quality over model scale, reducing computational and latency footprint without dependence on proprietary systems.
住宅用建物の間取り図適合性チェックのための自動化された AI ベースのフレームワークに向けて
オーストラリアの都市部の住民の幸福を改善するために、政府はアパートの設計品質を高めるために SEPP65、BADS、SPP7.3 などの政策改革を導入しました。これらの規制では、採光、自然換気、プライバシー、スペース効率などの健康関連機能を評価するために、正確な幾何学的および空間的分析が必要です。ただし、コンプライアンスチェックは手作業で時間がかかるため、依然として困難です。さらに、進化するポリシーにより、数千のアパートにわたる大規模な評価の拡張性が制限されます。既存の自動フロアプラン分析方法は細分化されており、通常は単一のアパートに焦点を当てており、複数ユニットのコンプライアンスチェックのための統一されたフレームワークが欠けています。この記事では、自動フロア プラン分析、特に AI を活用したアプローチの現在の進歩を調査し、実際の導入における主要な課題に焦点を当てます。これらのギャップに対処するために、複数の集合住宅の建物における自動コンプライアンスチェックのための概念的なフレームワークが提案されています。 Large Language Model (LLM) はルール エンジン内で使用され、テキストの構築コードを実行可能で説明可能なルールに変換します。データ抽出エンジンは、間取り図の画像を壁、部屋、設備、テキスト、シンボルなどの要素に分割し、トポロジ関係を備えた構造化された建物グラフに変換します。この構造化表現は、LLM が生成した評価ルールを利用するコンプライアンス チェック エンジンによって評価されます。提案されたフレームワークは、管轄区域全体で自動化されたコンプライアンスチェックに対するスケーラブルで一貫性のある透明性のあるアプローチを提供し、アパートの設計基準の効率的な施行をサポートし、より健全で高密度の都市開発を促進します。
原文 (English)
Towards an automated AI-based framework for floor plan compliance checks for residential buildings
To improve residents' well-being in Australia's urban areas, governments have introduced policy reforms such as SEPP65, BADS, and SPP7.3 to enhance apartment design quality. These regulations require precise geometric and spatial analysis to evaluate health-related features, including daylight access, natural ventilation, privacy, and space efficiency. However, compliance checking remains challenging due to its manual, time-intensive nature. Additionally, evolving policies limit scalability for large-scale assessments across thousands of apartments. Existing automated floor plan analysis methods are fragmented and typically focus on single apartments, lacking a unified framework for multi-unit compliance checking. This article explores current advancements in automated floor plan analysis, particularly AI-driven approaches, and highlights key challenges in their practical adoption. To address these gaps, a conceptual framework is proposed for automated compliance checking in multi-apartment buildings. A Large Language Model (LLM) is used within a Rule Engine to convert textual building codes into executable, explainable rules. A Data Extraction Engine segments floor plan images into elements such as walls, rooms, fixtures, text, and symbols, and transforms them into a structured building graph with topological relationships. This structured representation is then evaluated by a Compliance Check Engine, which leverages LLM-generated rules for assessment. The proposed framework offers a scalable, consistent, and transparent approach to automated compliance checking across jurisdictions, supporting efficient enforcement of apartment design standards and promoting healthier, higher-density urban development.
Libra: エージェントによる情報取得のための環境のトレーニング
大規模なリポジトリ内での情報のローカライゼーションは、エージェント LLM システムの基礎です。合成データ駆動型の最適化は LLM のトレーニングに成功していることが証明されていますが、エージェントの作業環境 (リポジトリ自体) をデータ駆動型で最適化することにはほとんど注目されていません。このギャップを埋めるために、可変の「カタログ」(ナビゲート可能なインデックスとして機能する階層型 Markdown ファイル) をリポジトリに導入する自己進化フレームワークである Libra を紹介します。 Libra は LLM 主導の最適化ループを実行します。このループでは、プロンプターが合成クエリを生成し、凍結されたソルバーがカタログをナビゲートしてそれらの解決を試み、ヒーラーがソルバーのローカリゼーションの失敗に応じてカタログを書き換えます。 12 個の SWE-bench Lite リポジトリにわたる評価では、この環境修復によってコード ローカリゼーションの精度が継続的に対数的に向上することが実証されました。さらに、これらの環境改善により、さまざまな LLM および問題セット間でゼロショットが転送されます。この論文の焦点はそのようなシステムの一般的な動作を研究することですが、Libra に最適化されたカタログを備えた最小限のコーディング エージェントが最先端のベースラインを上回るパフォーマンスを示すことも実証します。コードは https://github.com/salesforce-misc/Libra で、データは https://huggingface.co/datasets/Salesforce/Libra で入手できます。
原文 (English)
Libra: Training the Environment for Agentic Information Retrieval
Information localization within massive repositories is a cornerstone of agentic LLM systems. While synthetic data-driven optimization has proven successful in training LLMs, little attention has been paid to optimizing the agent's working environment (the repository itself) in a data-driven manner. To bridge this gap, we present Libra, a self-evolving framework that introduces mutable "catalogs" (hierarchical Markdown files serving as navigable indices) into the repository. Libra runs an LLM-driven optimization loop where a Prompter generates synthetic queries, a frozen Solver attempts to resolve them by navigating the catalogs, and a Healer rewrites the catalogs in response to the Solver's localization failures. Evaluations across 12 SWE-bench Lite repositories demonstrate that this environmental healing yields continual, logarithmic improvements in code localization accuracy. Furthermore, these environmental improvements transfer zero-shot across different LLMs and problem sets. Although the focus of this paper is to study the general behavior of such a system, we also demonstrate that a minimalist coding agent equipped with Libra-optimized catalogs outperforms state-of-the-art baselines. Code is available at https://github.com/salesforce-misc/Libra and data at https://huggingface.co/datasets/Salesforce/Libra.
ユーザー認識の想起の学習: 長期会話記憶におけるパーソナライズされた検索
長期的に会話を行うエージェントは過去の対話を記憶していることが期待されますが、記憶が役立つのは、適切なユーザーに対して適切な証拠が呼び出された場合のみです。既存のメモリ拡張 LLM エージェントは、コンパクトなメモリ バンクの構築において進歩を遂げていますが、検索は依然としてクエリ中心の類似性や固定ランキング ルールによって駆動されることが多く、ユーザー条件による関連性は十分に検討されていません。このギャップに対処するために、私たちは、メモリ検索をユーザー認識かつ最適化可能にする検索中心のフレームワークである Profile-guided Personalized Retrieval Optimization (PPRO) を提案します。PPRO は、エピソード的および最適化されたメモリ検索を構築します。対話履歴からセマンティック メモリ バンクを生成し、蓄積されたメモリからユーザー プロファイルを導き出します。このプロファイルは、メモリ ランキングにおける明示的なパーソナライズされた優先順位として機能し、安定したユーザー属性、好み、および関係性を考慮した検索を可能にします。PPRO はさらに、メモリ バンクと回答モデルを固定したまま、証拠取得品質と下流の回答品質の両方をフィードバックとして使用して、グループ相対ポリシー最適化を使用してクエリ リライタをトレーニングします。LoCoMo と LongMemEval-S での実験では、トレーニングなしと比べて一貫した向上が示されています。さらに、アブレーション研究では、プロファイルに基づくランキングと検索指向の書き換えの両方がパフォーマンスに大きく寄与していることが示されており、パーソナライズされた長期記憶使用の重要な要素として検索の最適化が強調されています。
原文 (English)
Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory
Long-term conversational agents are expected to remember past interactions, but memory is useful only when the right evidence is recalled for the right user. Existing memory-augmented LLM agents have made progress in building compact memory banks, yet retrieval is still often driven by query-centered similarity or fixed ranking rules, leaving user-conditioned relevance underexplored.To address this gap, we propose Profile-guided Personalized Retrieval Optimization (PPRO), a retrieval-centric framework that makes memory retrieval both user-aware and optimizable.PPRO builds episodic and semantic memory banks from dialogue histories and derives a user profile from accumulated memories.The profile serves as an explicit personalized prior in memory ranking, allowing retrieval to account for stable user attributes, preferences, and relationships.PPRO further trains a query rewriter with Group Relative Policy Optimization, using both evidence retrieval quality and downstream answer quality as feedback while keeping the memory banks and answer model fixed.Experiments on LoCoMo and LongMemEval-S show consistent gains over training-free memory systems and training-based baselines.Ablation studies further show that both profile-guided ranking and retrieval-oriented rewriting contribute substantially to performance, highlighting retrieval optimization as a key factor in personalized long-term memory use.
現実世界の LLM: 緊急事態における「AI」の評価
この文書は行動への呼びかけを提供します。私たちは研究コミュニティの同僚に対し、私たちの発見を一般に明らかにする上でより大きな役割を果たすよう強く求めます。リスクを説明するために、実世界のコンテキストでの LLM ベースの機械翻訳アプリケーションの導入の初期段階に関するケース スタディを紹介します。これは、オペレーターに直接電話をかけることが難しい緊急時に使用する、55 言語での text-2-911 システム広告機能です。私たちは、このようなテクノロジに関する多くの一般的な誤解を特定し、開発および展開パイプラインのあらゆる段階で関係者に向けた一連の具体的な推奨事項とベスト プラクティスで結論付けています。科学研究の進歩はしばしば「難しい」問題を解決することにありますが、最も見落とされているのは「簡単な」問題、つまり最新のテクノロジーが不要な問題であることが多いと私たちは主張します。
原文 (English)
LLMs in the Real World: Evaluating "AI" in Emergency Contexts
This paper offers a call to action. We urge our colleagues in the research community to play a greater role in the articulation of our findings to the public. To illustrate the stakes we present a case study on the initial stages of an LLM-based machine translation application's deployment in a real-world context: a text-2-911 system advertising capabilities in 55 languages for use in emergencies in which it may be difficult to call operators directly. We identify a number of common misconceptions about technologies such as these, concluding with a set of concrete recommendations and best practices for stakeholders at every stage of the development and deployment pipeline. While the advancement of scientific research often lies in solving the "hard" problems, we argue it is often the "easy" ones -- problems for which the latest technology is often unnecessary -- that are most overlooked.
スパースオートエンコーダを介して文の埋め込みを人間の概念に合わせる
密な文の埋め込みは、最新の検索拡張生成 (RAG) システムの基礎ですが、特徴の重ね合わせによる解釈可能性の欠如に悩まされています。この不透明さは、もつれた表現を分析したり制御したりすることが難しいため、検索プロセスを人間の意図に合わせるのを妨げます。この研究では、Top-k Sparse Autoencoder (SAE) を使用して、文変換器 (E5 など) の密な表現を人間が解釈可能な概念に解きほぐす方法を提案します。我々は、これらの解きほぐされた特徴が特定の意味論的、構文論的、および語用論的なカテゴリと一致することを実証します。さらに、検索プロセスへの正確な介入を可能にするアクティベーションステアリングメカニズムを導入します。特定の潜在的な特徴をクランプすることにより、バックボーン モデルを再トレーニングすることなく、検索結果を再ランク付けしてユーザーの制約に合わせて調整できることを示します。私たちの発見は、SAE ベースの分解が透過的で操作可能な神経情報検索への実行可能な道を提供することを示唆しています。
原文 (English)
Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders
Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to feature superposition. This opacity hinders the alignment of retrieval processes with human intent, as the entangled representations are difficult to analyze or control. In this work, we propose a method to disentangle the dense representations of sentence transformers (e.g., E5) into human-interpretable concepts using Top-k Sparse Autoencoders (SAEs). We demonstrate that these disentangled features align with specific semantic, syntactic, and pragmatic categories. Furthermore, we introduce an activation steering mechanism that allows for precise intervention in the retrieval process. By clamping specific latent features, we show that it is possible to re-rank search results to better align with user constraints without retraining the backbone model. Our findings suggest that SAE-based decomposition offers a viable path toward transparent and steerable neural information retrieval.
FLYNN: Fly Brain トポロジーを使用したロボット ナビゲーションのための堅牢なニューラル ネットワーク
深層学習モデルは複雑なタスクで最先端のパフォーマンスを実現しますが、新しい環境や感覚遮断に直面すると脆弱なままです。対照的に、生体系はこれらの課題に対して顕著な耐性を示します。私たちは、ショウジョウバエのシナプス分解能の脳コネクトームから直接派生したアーキテクチャをもつリカレント ニューラル ネットワーク (RNN) を開発することで、この脆弱性に対処します。我々は、MuJoCo でビジョンベースのナビゲーションを実行するためにフライ コネクトーム ニューラル ネットワーク (FLYNN) をトレーニングし、同様のパラメーター数の最新の手作りネットワークに匹敵するパフォーマンスを達成する実現可能性を実証します。重要なことは、FLYNN は、さらなるトレーニングを行わなくても、分布外 (OOD) データに対する優れた耐性と感覚喪失に対する耐性を示します。完全な視力喪失下でも機能を維持しましたが、手作りのネットワークは、カメラのドロップアウトで特別に訓練された場合でも、ほとんど機能しませんでした。 FLYNN の内部状態の主成分分析 (PCA) は、FLYNN が特に高度な表現モジュール性を示していることを示唆しており、これがその堅牢性に関連している可能性があります。私たちの研究は、生物学的な脳のトポロジーに従って弾力性のある人工エージェントを設計するための新しい方向性を提供します。
原文 (English)
FLYNN: Robust Neural Network for Robot Navigation using Fly Brain Topology
While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or sensory deprivation. In contrast, biological systems exhibit remarkable tolerance to these challenges. We address this vulnerability by developing a recurrent neural network (RNN) whose architecture is directly derived from the synaptic-resolution brain connectome of the fruit fly Drosophila melanogaster. We demonstrate the feasibility of training the fly connectome neural network (FLYNN) to perform vision-based navigation in MuJoCo, achieving performance comparable to modern hand-crafted networks of similar parameter counts. Crucially, FLYNN exhibits superior resistance to out-of-distribution (OOD) data and tolerance to sensory loss without further training. It remained functional even under total vision loss while hand-crafted networks largely failed, even when specifically trained with camera dropout. Principal Component Analysis (PCA) of the internal state of FLYNN suggests that it exhibits a particularly high degree of representational modularity, which might be related to its robustness. Our work provides a new direction for designing resilient artificial agents following the topology of biological brains.
身体化されたインテリジェンスのためのメモリネイティブの非地上ネットワーク
非地上ネットワーク (NTN) は、身体化知能 (EI) のユビキタス接続を提供し、荒野のロボットがクラウド リソースを活用したり、重要な情報をリモート センターに報告したりできるようにします。ただし、非常に動的で、リソースに制約があり、トポロジーが変化し、タスク指向の環境であるため、相乗効果は簡単ではありません。既存のメモリレス NTN プロトコルは、ローカル チャネルの状態と瞬間的なサービス要求によって決定が左右されるため、非効率になります。これらの制限に対処するために、この文書では、メモリ拡張システムの最適化にロングホライズン コンテキストを活用するメモリ ネイティブ NTN (MemNTN) パラダイムを提案します。このパラダイムシフトを実現するために、世界の現状を表す物理メモリと歴史的なネットワークエクスペリエンスをエンコードするデジタルメモリを区別するデュアルメモリアーキテクチャを確立します。当社は、物理層とアクセス層からネットワーク層とアプリケーション層に至るまで、クロスレイヤーのメモリネイティブの意思決定を容易にするメモリの取得、圧縮、評価、更新、および利用メカニズムを開発します。衛星による質問応答(SEQA)の実験により、提案された MemNTN が従来のステートレス NTN および地上アプローチよりも大幅に優れていることが実証されました。
原文 (English)
Memory-Native Non-Terrestrial Networks for Embodied Intelligence
Non-terrestrial networks (NTN) provide ubiquitous connectivity for embodied intelligence (EI), enabling robots in wilderness to leverage cloud resources or report critical information to remote centers. However, the synergy is nontrivial due to the highly-dynamic, resource-constrained, topology-varying, and task-oriented environment. Existing memoryless NTN protocols become inefficient, since the decisions are driven by local channel conditions and instantaneous service demands. To address these limitations, this paper proposes the memory-native NTN (MemNTN) paradigm that leverages long-horizon contexts for memory augmented system optimization. To realize this paradigm shift, we establish a dual-memory architecture that distinguishes between physical memory representing the state of the world and digital memory encoding historical network experience. We develop memory acquisition, compression, valuation, update, and utilization mechanisms that facilitate cross-layer, memory-native decision-making, spanning from the physical and access layers up to the network and application layers. Experiments in satellite embodied question answering (SEQA) demonstrate that the proposed MemNTN significantly outperforms conventional stateless NTN and terrestrial approaches.
コンタクトレンチを使った器用な操作を人間の実演から学ぶ
ロボットの器用な操作は人間の豊富なデモンストレーションから恩恵を受ける可能性がありますが、そのようなデモンストレーションをロボット政策に移すことは依然として困難です。我々は、強化学習による剛体および多関節オブジェクトの長期的な操作のためのフレームワークである、ロボットによる器用な操作における人間のデモンストレーション (CHORD) からのコンタクト レンチ ガイダンスを紹介します。重要なアイデアは、オブジェクト中心のコンタクト レンチ空間ガイダンスです。人間とロボットの動きを、オブジェクトに誘発できる力とトルクによって表現し、誘発された瞬間的な動きによって類似性を測定できるようにします。このガイダンスにより、強化学習は接触の多い器用な操作に対してよりスケーラブルになります。さらに、モーション キャプチャ データセットと再構築された社内ビデオから構築された、4,739 の両手による器用な操作タスクを含む大規模なシミュレーション ベンチマークを紹介します。 1,831 のベンチマーク タスクで評価した結果、CHORD は平均成功率 82.12% を達成し、強力なスケーラビリティを実証しました。また、CHORD は、手のみおよび三人称のデモンストレーションから全身操作に一般化し、90.77% の成功率を達成し、学習されたポリシーは、開ループ設定と閉ループ設定の両方で現実世界に転送されます。
原文 (English)
Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration
Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.
ATM: マルチエージェント コード共同合成のための CID ブローカーによる事前書き込み許可
マルチエージェント LLM システムでは、ソフトウェア エンジニアリングの作業を計画、生成、検証、修復に分解できますが、より狭いシステムの問題が残ります。管理された共有ミューテーションが適用される前に、システムは、同時に形成されたどの書き込みインテントを並行して進めることができるか、決定論的な構成またはシリアル化が必要で、どの書き込みインテントがフェイルクローズされたパスを取る必要があるかを決定する必要があります。私たちは、単一のガバナンス ドメイン内で動作するソフトウェア エージェント用の仕様に基づいたガバナンス基盤である AI-Atomic-Framework (ATM) を使用して、この問題に対処します。 ATM は、タスクの意図、リポジトリの範囲、書き込み許可、検証、および証拠の義務を 1 つのガバナンス チェーンにバインドします。 Content Identifier (CID) ブローカーは、共有突然変異受付サブシステムとして機能します。アダプターガイドの原子化マップは、意味論的なアトムと境界領域にインテントを書き込みます。永続的なアトム マップのカバレッジが不完全な場合、仮想アトムは保守的な比較とルーティングのための一時的な監査可能なガバナンス ユニットを提供します。管理された共有書き込みは、提案されたエージェントによって直接適用されるのではなく、中立的なスチュワードによって最終的に適用されます。評価では、12 のシナリオの決定論的設計マトリックス、3 つのアーカイブされたランナー ケース、ATM-AdmissionBench、3 つのアーカイブされた同一ファイル境界ケース、3 週間の外部アダプタ調査、および運用リカバリ ルーティング ベンチマークを含む、制御された証拠、フィールド 証拠、導入証拠、および拡張証拠を組み合わせます。この結果は、観察された単一ドメイン設定内での実現可能性、監査可能性、および限定された回復可能性を裏付けていますが、広範な比較優位性やクロスクローン ガバナンスを主張するものではありません。
原文 (English)
ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis
Multi-agent LLM systems can decompose software-engineering work into planning, generation, validation, and repair, but a narrower systems problem remains: before any governed shared mutation is applied, a system must decide which concurrently formed write intents may proceed in parallel, which require deterministic composition or serialization, and which must take a fail-closed path. We address this problem with the AI-Atomic-Framework (ATM), a specification-grounded governance substrate for software agents operating within a single governance domain. ATM binds task intent, repository scope, write admission, validation, and evidence obligations into one governance chain. A Content Identifier (CID) broker serves as the shared-mutation admission subsystem. Adapter-guided atomization maps write intents to semantic atoms and bounded regions; when persistent atom-map coverage is incomplete, virtual atoms provide temporary auditable governance units for conservative comparison and routing. Governed shared writes are ultimately applied by a neutral steward rather than directly by proposing agents. Evaluation combines controlled, field, adoption, and extension evidence, including a 12-scenario deterministic design matrix, three archived runner cases, ATM-AdmissionBench, three archived same-file boundary cases, a three-week external-adopter study, and an operational recovery-routing benchmark. The results support feasibility, auditability, and bounded recoverability within the observed single-domain settings, but do not claim broad comparative superiority or cross-clone governance.
滞留を伴う宛先ラベル付き自己ループ システム: 固有の特性評価、実現コスト、および認識
我々は、許容される可視遷移が事前に固定されており、各可視状態が最小滞留要件を持つシステムのための有限状態シンボリック コントローラーを研究します。結果として得られるモデルは、目的地ラベル付き滞留型自己ループ システム (DLSL システム) と呼ばれるもので、ローカル デシジョン マップとともに表示されるグラフを記録します。滞留メモリは相展開後にのみ現れます。主な構造的問題は、いったん滞在が課されると、現在の目に見える状態によっては出発が許可されるかどうかが決定されなくなることです。これは逆の問題につながります。どの決定論的トランスデューサが、固定された可視グラフ上の DLSL システムの位相展開された実現として生じますか?答えはまさにファイバー線形グラフを考慮したトランスデューサーのクラスであることを示します。自然な到達可能性と実現可能な出発の仮定の下では、同じ目に見えるグラフ上の同等のアクセス可能な実現は同型です。特に、可視の変換は滞留ベクトルと局所的な決定マップを決定します。また、ドウェル値 $(d_i)$ を強制するグラフ保存の決定論的実現には、正確に $\sum_id_i$ 制御状態が必要であることも証明します。最後に、$O(|Q||\Omega|)$ 認識および再構成手順を与え、遷移が後続ファイバーの内部相に入る可能性があるエッジエントリーバリアントに解析を拡張します。
原文 (English)
Destination-Labeled Self-Looping Systems with Dwell: Intrinsic Characterization, Realization Cost, and Recognition
We study a finite-state symbolic controller for systems in which the admissible visible transitions are fixed in advance and each visible state carries a minimum dwell requirement. The resulting model, which we call a destination-labeled self-looping system with dwell (DLSL system), records the visible graph together with local decision maps; dwell memory appears only after phase expansion. The main structural issue is that, once dwell is imposed, the current visible state no longer determines whether a departure is allowed. This leads to the converse problem: which deterministic transducers arise as phase-expanded realizations of DLSL systems over a fixed visible graph? We show that the answer is exactly the class of fiber-linear graph-respecting transducers. Under natural reachability and realizable-departure assumptions, equivalent accessible realizations over the same visible graph are isomorphic; in particular, the visible transduction determines the dwell vector and the local decision maps. We also prove that any graph-preserving deterministic realization enforcing dwell values $(d_i)$ requires exactly $\sum_i d_i$ control states. Finally, we give an $O(|Q||\Omega|)$ recognition and reconstruction procedure, and extend the analysis to an edge-entry variant in which transitions may enter interior phases of successor fibers.
スクラム認定形式の質問に関する大規模な言語モデルの比較: 精度、安定性、エラー パターン
大規模言語モデル (LLM) は、試験および認定スタイルの質問応答タスクでますます使用されており、ドメイン固有の知識を取得、解釈、適用する能力を体系的に評価できます。ソフトウェア エンジニアリングでは、質問が規範的な定義、役割、成果物、ルールの厳密な遵守に依存している場合、このような設定は特に重要です。このペーパーでは、プロフェッショナル スクラム マスター I (PSM I) 評価形式に沿った 993 件のスクラム認定スタイルの質問に答える際に、\textit{GPT-5 mini}、\textit{Gemini 3 Flash}、\textit{DeepSeek Chat 3.2} という 3 つの最新の LLM のパフォーマンスを評価します。私たちは 3 つのプロンプト戦略 (\textit{zero-shot}、\textit{chain-of-thought}、\textit{source-grounded}) に基づいてモデルを評価し、繰り返し実行してモデル内の安定性を評価しました。また、スクラムのトピックと質問形式全体のパフォーマンスを分析し、誤った回答で繰り返されるエラー パターンの定性分析によって補完しました。結果は、モデル間の明らかな違いを明らかにし、Gemini 3 Flash が最高の精度を達成し、GPT-5 mini と DeepSeek Chat 3.2 がそれに続きましたが、モデル内のばらつきはすべての条件で低いままでした。質問形式別にみると、モデルは単一回答の多肢選択項目で最も高い精度を達成しましたが、複数選択や正誤問題はよりエラーが発生しやすくなりました。トピックごとに、成果物、経験主義、製品価値などの規範的に明示的な領域ではパフォーマンスの一貫性が高かったが、スクラム価値、自己管理チーム、ステークホルダーと顧客ではより不安定でした。定性分析の結果、エラーはランダムではなく体系的であり、過剰な一般化、限定的な文言、複合的な混乱要因、一般的な市場の解釈と厳密なスクラム定義との矛盾が関与していることがわかりました。
原文 (English)
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed. In Software Engineering, such settings are particularly relevant when questions depend on strict adherence to normative definitions, roles, artifacts, and rules. This paper evaluates the performance of three contemporary LLMs, \textit{GPT-5 mini}, \textit{Gemini 3 Flash}, and \textit{DeepSeek Chat 3.2}, in answering 993 Scrum certification-style questions aligned with the Professional Scrum Master I (PSM I) assessment format. We evaluated the models under three prompting strategies (\textit{zero-shot}, \textit{chain-of-thought}, and \textit{source-grounded}), with repeated executions to assess intra-model stability. We also analyzed performance across Scrum topics and question formats, complemented by a qualitative analysis of recurring error patterns in incorrect answers. Results revealed clear differences among models, with Gemini 3 Flash achieving the highest accuracy, followed by GPT-5 mini and DeepSeek Chat 3.2, while intra-model variability remained low across all conditions. By question format, the models achieved the highest accuracy on single-answer multiple-choice items, whereas multi-select and True/False questions were more error-prone. By topic, performance was more consistent in normatively explicit areas such as Artifacts, Empiricism, and Product Value, but more fragile in Scrum Values, Self-Managing Teams, and Stakeholders \& Customers. The qualitative analysis showed that errors were systematic rather than random, involving overgeneralization, restrictive wording, compound distractors, and conflicts between common market interpretations and strict Scrum definitions.
スクラム認定の質問に関する GPT-5 のプロンプト: 実証的精度調査
大規模言語モデル (LLM) は、アジャイル ソフトウェア開発で文書化、コーチング、トレーニングのためにますます使用されています。実践者がプロフェッショナル スクラム マスター (PSM) などの認定資格の準備のためにこれらのツールを採用する場合、重要な問題は、LLM がスクラム (スクラム ガイド (2020) で説明されている規範的で明確に定義されたルールを持つフレームワーク) について確実に推論できるかどうかです。このペーパーでは、さまざまなプロンプト手法が、スクラム認定形式の質問に対する LLM の回答の事実の正確さにどのような影響を与えるかを検証します。 993 件の検証済み PSM 対応質問のデータセットは、ゼロショット、思考連鎖、出典引用の 3 つの手法を使用して GPT-5 によって回答されました。すべてのプロンプトは 85\% 以上の認定レベルの精度を達成し、引用ベースのバリアントのパフォーマンスが最高 (89.1\%) で、エラー率が最も低くなりました。正解は、\emph{完了の定義}、イベント、プロダクト バックログ管理などの明確に定義されたトピックや、単一回答の複数選択項目に集中していましたが、複数選択の質問や、スクラム チームやプロダクト価値などのより解釈的な領域は安定していませんでした。少なくとも 1 つのプロンプトが失敗した質問 (16.2\%) のうち、エラーはスクラム ガイドとの不整合 (28\%)、範囲外のコンテンツ (34\%)、および古いまたは偏った解釈 (38\%) に集中していました。全体として、プロンプト技術により、特に誤解やバージョンのドリフトが減少し、アジャイルの学習や認定準備における LLM のより信頼性の高い使用がサポートされるなど、控えめではありますが一貫した改善がもたらされました。
原文 (English)
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioners adopt these tools to prepare for certifications such as Professional Scrum Master (PSM), a key question is whether LLMs can reliably reason about Scrum, a framework with normative, well-defined rules described in the Scrum Guide (2020). This paper examines how different prompt techniques affect the factual accuracy of LLM responses to Scrum certification-style questions. A dataset of 993 validated PSM-aligned questions was answered by GPT-5 using three techniques: zero-shot, chain-of-thought, and with-source citation. All prompts achieved certification-level accuracy above 85\%, with the citation-based variant performing best (89.1\%) and yielding the lowest error rate. Correct answers concentrated in well-defined topics, such as \emph{Definition of Done}, Events, and Product Backlog Management, and in single-answer multiple-choice items, while multi-select questions and more interpretive areas, such as Scrum Team and Product Value, were less stable. Among questions where at least one prompt failed (16.2\%), errors clustered into misalignment with the Scrum Guide (28\%), content outside its scope (34\%), and outdated or biased interpretations (38\%). Overall, prompt techniques produced modest but consistent improvements, particularly in reducing misinterpretation and version drift, supporting more reliable use of LLMs in Agile learning and certification preparation.
AGE: グラフ検索拡張生成におけるグラフ埋め込みのための適応マスキング
GraphRAG は、外部知識としてグラフ構造化データを参照することで大規模言語モデル (LLM) をサポートする検索拡張生成 (RAG) の拡張機能です。この手法は複雑な関係を理想的に捕捉しますが、グラフベースとテキストベースの潜在特徴間の不整合のため、LLM、特に凍結 LLM のグラフ表現に苦労することがよくあります。私たちは、{\it Adaptive-masking for Graph Embedding (AGE)} を導入することでこの問題に取り組みます。 AGE は、マスクベースの自己教師あり学習 (SSL) アプローチで Transformer を採用しています。私たちはテキスト埋め込みエンコーダーと同様のアーキテクチャを設計し、潜在的な機能の不整合に対処しました。自然言語テキストとは対照的に、グラフは簡潔な表現であり、周囲から予測することが困難な主要なコンテキスト情報を保持する {\it key ノード} が存在します。このようなキー ノードをマスクすると、SSL プロセスが非効率になります。したがって、AGE は学習可能なノード サンプラーを利用して、キー ノードとは別にノードを予測することに重点を置いています。私たちの実験結果は、AGE が GraphQA タスクでノンパラメトリック検索コンポーネントを使用するアプローチを大幅に改善し、異なる特徴を持つ 4 つのベンチマーク データセットにわたって優れた精度を達成することを示しています。
原文 (English)
AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation
GraphRAG is an extension of retrieval-augmented generation (RAG) that supports large language models (LLMs) by referring to graph-structured data as external knowledge. While this technique ideally captures intricate relationships, it often struggles with graph representations for LLMs, particularly for frozen LLMs, due to the misalignment between graph-based and text-based latent features. We tackle this issue by introducing the {\it Adaptive-masking for Graph Embedding (AGE)}. AGE employs a Transformer in a mask-based self-supervised learning (SSL) approach. We designed the architecture similar to text embedding encoders, addressing the latent feature misalignment. In contrast to natural language texts, graphs are concise representations, and there exist {\it key nodes} that hold dominant contextual information, which are challenging to predict from their surroundings. Masking such key nodes leads to inefficiency in the SSL process. Therefore, AGE focuses on predicting nodes apart from key nodes, utilizing a learnable node sampler. Our experimental results indicate that AGE significantly improves approaches using non-parametric search component in GraphQA tasks, achieving superior accuracy across four benchmark datasets with distinct characteristics.
SWE-Router: マルチターン エージェント ソフトウェア エンジニアリング タスクにおけるルーティング
マルチターン エージェント ハーネスに埋め込まれた大規模言語モデル (LLM) はソフトウェア エンジニアリング (SWE) を再構築していますが、多くの問題が安価な修正で済む場合、すべてのタスクをフロンティア モデルにルーティングするのは無駄です。既存の LLM ルーターは、エージェント設定の情報理論的なベイズ誤差下限を継承するタスク記述のみで動作します。同様の問題により、局所的なタイプミスまたはマルチモジュール リファクタリングのいずれかが隠蔽される可能性があり、プロンプトは 2 つを分離しません。 SWE-Router を導入します。これは、安価なモデルを数ターン探索的に実行し、その結果の部分的な軌道を読み取ってから、安価なモデルを続行するか高価なモデルにエスカレーションするかを決定する、値ベースの時間的アプローチです。我々は、部分軌道の条件付けがルーティングに害を及ぼすことはなく、探索が有益である場合には常に厳密に優れていることを示すベイズの最適性定理を提供します。現代のコストと能力のフロンティアにわたる弱いモデルと強いモデルの LLM ペア全体で、SWE-Router が強力なモデルのパフォーマンスの大部分を維持しながら、SWE タスクのコスト効率を大幅に向上させることを示します。さらに、軌道レベルのルーティングを再現できるマルチ LLM 軌道データセットもリリースします。
原文 (English)
SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks
Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier model is wasteful when many issues admit cheap fixes. Existing LLM routers operate on the task description alone, which inherits an information-theoretic Bayes-error floor in agentic settings: a similar issue can hide either a localized typo or a multi-module refactor, and the prompt does not separate the two. We introduce SWE-Router, a value-based temporal approach that lets a cheap model run for a few exploratory turns and reads the resulting partial trajectory before deciding whether to continue cheaply or to escalate to an expensive model. We provide a Bayes-optimality theorem showing that conditioning on the partial trajectory never harms routing and is strictly better whenever exploration is informative. Across the LLM pairs of weak and strong models spanning the contemporary cost--capability frontier, we show that SWE-Router greatly improves the cost efficiency of SWE tasks, while maintaining the majority of the performances of the stronger model. We additionally release a multi-LLM trajectory dataset which allows reproduction of our trajectory-level routing.
RIS 支援追跡と電力制御のためのアクティブ センシング: ニューロ進化と教師あり学習のハイブリッド アプローチ
この論文では、再構成可能なインテリジェント サーフェス (RIS) を利用して、電力が制限されているモバイル ユーザーをエネルギー効率よく追跡する方法について研究します。ローカリゼーション パイロットの送信は、電力に制約のあるデバイスのエネルギー バジェットを支配するため、基地局 (BS) からユーザーへの低オーバーヘッドのフィードバック リンクを導入して、動的なアップリンク電力制御を可能にします。このアクティブ センシング問題の離散的かつ分散的な性質を克服するために、離散 RIS 位相プロファイルと UE の送信電力をリアルタイムで共同最適化する新しいデュアル エージェント (DA) 深層学習フレームワークを提案します。具体的には、私たちのアプローチは、神経進化パラダイムと教師あり学習を統合したハイブリッドトレーニング方法論を採用しており、RISユニット要素からの離散位相応答の非微分可能性と、パイロット電力制御のためのシングルビットフィードバックメッセージの厳密な情報ボトルネックを効果的に克服します。提案された DA アクティブ センシング フレームワークは、シングル アンテナ BS とマルチ アンテナ BS の両方に適用できます。後者では、1 つの NN の構造にわずかな変更が加えられるだけです。後者の場合、有限セットから有効なデジタル コンバイナを選択するために、適切な構造を備えた追加の出力ブランチが含まれています。広範な数値シミュレーションにより、提案されたスキームがさまざまなターゲット運動モデルにわたって高精度かつ堅牢な追跡を実現し、拡張カルマン フィルターや粒子フィルター、さらには機械学習ベースの追跡機能を上回る性能を発揮することが実証されました。さらに、静的位置特定では、従来のフィンガープリンティング スキーム、深層強化学習ベースライン、標準的な逆伝播ベースの推定器よりも大幅に優れたパフォーマンスを示すことが示されています。
原文 (English)
Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach
This paper studies energy efficient tracking of power-limited mobile users with the assistance of a Reconfigurable Intelligent Surface (RIS). Since localization pilot transmissions dominate the energy budget of power-constrained devices, we introduce a low-overhead feedback link from the Base Station (BS) to the user to enable dynamic uplink power control. To navigate the discrete and decentralized nature of this active sensing problem, we propose a novel Dual-Agent (DA) deep learning framework that jointly optimizes the discrete RIS phase profiles and the UE's transmit power in real time. Specifically, our approach employs a hybrid training methodology integrating the neuroevolution paradigm with supervised learning, effectively overcoming the non-differentiability of discrete phase responses from the RIS unit elements and the strict information bottleneck of single-bit feedback messages for pilot power control. The proposed DA active sensing framework can be applied with both single- and multi-antenna BSs, the latter with only minor modifications in the structure of one NN: an additional output branch with appropriate structure is included for the latter case to select a valid digital combiner from a finite set. Extensive numerical simulations demonstrate that the proposed scheme achieves highly accurate and robust tracking across diverse target motion models, outperforming extended Kalman and particle filters, as well as, machine learning-based trackers. Furthermore, in static localization, it is shown to significantly outperform traditional fingerprinting schemes, deep reinforcement learning baselines, and standard backpropagation-based estimators.
マルチスケールレイヤーアテンションによるOracle Bone Inscriptionの認識の強化
Oracle Bone Inscriptions (OBI) の認識は、古代中国文化を理解する上で重要な役割を果たします。ただし、OBI は形状が複雑で不規則で、劣化していることが多いため、正確に認識することは依然として非常に困難です。従来の方法は専門知識と手動分析に依存しており、時間がかかり、エラーが発生しやすくなります。ディープラーニングは一般的な画像認識を大幅に進歩させましたが、既存の方法では OBI に固有のきめの細かい詳細や微妙な変化を捕捉するのが難しく、パフォーマンスが制限されています。最新の効果的なレイヤー アテンション技術でさえ、強化されたレイヤー間の相互作用を通じてきめの細かい依存関係を捕捉するように設計されていますが、依然として OBI 認識においてはわずかな改善しか示していません。これらの制限に対処するために、マルチスケール レイヤー アテンション (MSLA) を提案します。これは、マルチスケールとクロスレイヤーの両方の機能の相互作用を明示的にモデル化する新しいパラダイムです。 MSLA は、複数の空間スケールにわたるきめ細かい詳細で表現を強化することにより、より正確で堅牢な OBI 認識を可能にします。大規模な OBI データセットに対する広範な実験により、MSLA が計算効率を維持しながら、既存のアテンション メカニズムよりも常に優れたパフォーマンスを発揮することが実証されました。
原文 (English)
Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention
Oracle Bone Inscriptions (OBIs) recognition plays a crucial role in understanding ancient Chinese culture. However, accurately recognizing OBIs remains highly challenging due to their complex, irregular, and often degraded shapes. Traditional methods rely on expert knowledge and manual analysis, which are time-consuming and error-prone. Although deep learning has greatly advanced general image recognition, existing methods struggle to capture the fine-grained details and subtle variations inherent in OBIs, resulting in limited performance. Even most recent and effective layer attention techniques are designed to capture fine-grained dependencies through enhanced inter-layer interactions, yet they still exhibit only marginal improvements in OBIs recognition. To address these limitations, we propose Multi-Scale Layer Attention (MSLA), a novel paradigm that explicitly models both multi-scale and cross-layer feature interactions. By enriching the representation with fine-grained details across multiple spatial scales, MSLA enables more accurate and robust OBIs recognition. Extensive experiments on large-scale OBIs datasets demonstrate that MSLA consistently outperforms existing attention mechanisms while maintaining computational efficiency.
AlgoBench: コード生成におけるアルゴリズム適応のベンチマーク
HumanEval や LiveCodeBench などの確立されたプログラミング ベンチマークでの高い合格率は、モデルがアルゴリズムについて推論できるかどうかを必ずしも示しているわけではありません。多くの固定ベンチマークは、リリースされた問題ステートメント、論説、生成されたソリューションを通じて、最終的に公開トレーニング エコシステムの一部となり、より強力なアルゴリズム能力ではなく露出によって部分的に後のモデルを改善できるようになります。 ALGOBENCH を紹介します。これは、構造化された制約を変更する変換を通じて、既知の競技プログラミングの問題から新しいアルゴリズムの問題を自動的に構築するフレームワークです。受け入れられた ALGOBENCH バリアントはそれぞれ、原因となる問題を追跡できますが、元の参照アルゴリズムを失敗させる必要があります。 pass@$k$ の他に、OPTT、OPTS、TRAPRATE、GAPT、CONSENS などの複雑性を考慮したメトリクスを導入し、ソリューションが機能的に正しいかどうかだけでなく、生成された問題に漸近的に適しているかどうかをテストします。複数の LLM とプロンプト戦略にわたる実験では、ALGOBENCH バリアントではパフォーマンスが急激に低下し、取得により古いアルゴリズムの再利用が増加する可能性があり、正しく見えるソリューションの多くは必要な複雑さを満たしていないことが示されています。エラー分析では、障害は実装レベルではなく主にアルゴリズムに起因することが示されており、ALGOBENCH が機能の正しさを超えて適応を評価していることが示唆されています。
原文 (English)
AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation
High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason about algorithms. Many fixed benchmarks eventually become part of the public training ecosystem through released problem statements, editorials, and generated solutions, allowing later models to improve partly by exposure rather than by stronger algorithmic ability. We introduce ALGOBENCH, a framework that automatically builds novel algorithmic problems from known competitive-programming problems through structured constraint-shifting transformations. Each accepted ALGOBENCH variant is traceable to a source problem, but must make the original reference algorithm fail. Beyond pass@$k$, we introduce complexity-aware metrics -- including OPTT, OPTS, TRAPRATE, GAPT, and CONSENS -- to test whether a solution is not only functionally correct but also asymptotically suitable for the generated problem. Experiments across multiple LLMs and prompting strategies show that performance drops sharply on ALGOBENCH variants, retrieval can increase reuse of the old algorithm, and many correct-looking solutions fail to meet the required complexity. Error analysis shows that failures are mainly algorithmic rather than implementation-level, suggesting that ALGOBENCH evaluates adaptation beyond functional correctness.
スペクトル幾何学とボソンブロッホプローブ: 量子学習の探求
この論文では、量子学習モデルでスペクトル幾何学がどのように現れるか、そしてそれを物理的に接地されたプローブでどのように診断できるかを研究します。グラフ正則化量子ネットワークでは、トレーニングにより出力類似度グラフが再編成され、有効スペクトル次元 デルタ S = +0.23 が増加し、ラプラシアン スペクトルが再形成されます。エッジ分解された 2 ボソン干渉は、この再構成を直接調査します。ボソン強化デルタ P_uv は、フィードラー エッジ スプリット |デルタ v_2| と相関します。 (r = -0.50)、学習されたスペクトル分割を干渉シグネチャにリンクします。位相図は、結合強度ガンマとノイズ デルタに対する性能の非単調な依存性を示しており、グラフの正則化により、制限された領域でのみ忠実度が向上します。ハードウェア実験により、ショットノイズの不確実性の範囲内で予測される干渉挙動が確認されます。また、ハイブリッド量子オートエンコーダーを分析し、その潜在表現の幾何学的診断としてブロッホ空間ドリフトを導入します。教師なし良性データしきい値を使用すると、モデルは高いランキング パフォーマンス (ROC-AUC 約 0.99) と無視できる程度の偽陰性率を達成します。絶対的なブロッホ ドリフトは異常を強く識別します (ROC-AUC 少なくとも約 0.9)。一方、連続的なドリフトはほぼランダムです (ROC-AUC 約 0.5)。これは、検出が局所的な変動ではなく永続的な状態空間の変位から生じることを示しています。これらの結果は、縮小単一量子ビット状態の幾何学と関連する量子フィッシャー情報を通じて、学習によって引き起こされるスペクトル組織化が測定可能な量子状態構造として現れ、ボソンプローブとブロッホプローブを使用して量子学習システムを診断するための統一されたスペクトル幾何学的フレームワークを確立することを示しています。
原文 (English)
Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning
This paper studies how spectral geometry emerges in quantum learning models and how it can be diagnosed with physically grounded probes. In graph-regularized quantum networks, training reorganizes the output similarity graph, increases the effective spectral dimension Delta S = +0.23, and reshapes the Laplacian spectrum. Edge-resolved two-boson interference directly probes this restructuring: the bosonic enhancement Delta P_uv correlates with the Fiedler edge split |Delta v_2| (r = -0.50), linking learned spectral partitions to interference signatures. A phase diagram shows a nonmonotonic dependence of performance on coupling strength gamma and noise delta, with graph regularization improving fidelity only in a restricted regime; hardware experiments confirm the predicted interference behavior within shot-noise uncertainty. We also analyze a hybrid quantum autoencoder and introduce Bloch-space drift as a geometric diagnostic of its latent representation. With an unsupervised benign-data threshold, the model achieves high ranking performance (ROC-AUC about 0.99) and negligible false-negative rates. Absolute Bloch drift strongly discriminates anomalies (ROC-AUC at least about 0.9), while consecutive drift is near random (ROC-AUC about 0.5), showing that detection arises from persistent state-space displacement rather than local fluctuations. Through the geometry of reduced single-qubit states and associated quantum Fisher information, these results show that learning-induced spectral organization appears as measurable quantum-state structure, establishing a unified spectral-geometric framework for diagnosing quantum learning systems with bosonic and Bloch probes.
静的および動的環境における最適なあらゆる角度のパス計画
任意角度パス プランニングは、事前定義されたエッジによって制限されるのではなく、任意の頂点ペア間の移動を許可することで、従来のグラフベースのパス プランニングを拡張します。グラフを使用して連続空間内でより直線的で短い経路を見つけることができるため、空域、倉庫、海洋などの開けた場所でのナビゲーションに特に適しています。あらゆる角度からの経路計画アルゴリズムが数多く提案されていますが、特に動的障害物が存在する場合に、最適な解決策を保証できるものはほんのわずかです。この課題に対処するために、この記事では、グリッド上の最適な任意角度パス プランニングに焦点を当て、静的環境と動的環境の両方で最適性を維持しながら計算を高速化する 2 つの一般的な手法を紹介します。1) 楕円ベースの近傍を利用して探索空間を制限する楕円前方拡張、2) 従来の見通し線の方法を置き換えて可視性チェックを高速化する視野。これら 2 つの技術を統合するために、反転スキャンと順方向スキャンが導入されます。逆スキャンでは開いたノードから視覚的な接続が確立されますが、順方向スキャンでは閉じたノードからスキャンが開始されます。提案された技術に基づいて、Zeta* と Zeta*-SIPP はそれぞれ静的環境と動的環境向けに開発されました。 Zeta* は、順方向スキャンと組み合わせると、最先端のアルゴリズム Anya に似ており、同等のパフォーマンスを実現します。 Anya とは異なり、Zeta* は動的環境 (例: Zeta*-SIPP) などの他の設定に容易に拡張できます。 Zeta*-SIPP は、いずれのスキャン方式でも、対応する最先端の最適プランナー TO-AA-SIPP より 20 倍以上高速です。全体として、この調査では、最適なあらゆる角度のパス計画を達成するための重要な要件を特定し、さまざまな環境に適した統一アプローチを導入しています。
原文 (English)
Optimal any-angle path planning in static and dynamic environments
Any-angle path planning extends traditional graph-based path planning by allowing movement between any pair of vertices, rather than being restricted by predefined edges. It can find straighter and shorter paths in continuous space with graphs, making it particularly suitable for navigation in open areas such as airspaces, warehouses, and oceans. Many any-angle path-planning algorithms have been proposed, but only a few can guarantee optimal solutions, especially in the presence of dynamic obstacles. To address this challenge, this article focuses on optimal any-angle path planning on grids and introduces two general techniques that accelerate computation while preserving optimality in both static and dynamic environments: 1) elliptical forward expansion, which leverages ellipse-based neighborhoods to restrict the search space, and 2) field of view, which replaces traditional line-of-sight methods to speed up visibility checks. To integrate these two techniques, inverted and forward scanning are introduced. Inverted scanning establishes visual connections from open nodes, whereas forward scanning initiates scans from closed nodes. Building on the proposed techniques, Zeta* and Zeta*-SIPP are developed for static and dynamic environments respectively. Zeta*, when combined with forward scanning, is similar to the state-of-the-art algorithm Anya and attains comparable performance. Unlike Anya, Zeta* can be readily extended to other settings, such as dynamic environments (e.g., Zeta*-SIPP). Zeta*-SIPP, with either scanning method, is more than 20 times faster than the corresponding state-of-the-art optimal planner TO-AA-SIPP. Overall, this research identifies the key requirements for achieving optimal any-angle path planning and introduces a unified approach suitable for different environments.
潜在空間の活用: ステアリング ベクトルから制御と信頼のためのモデル キャリブレーターまで
言語モデルは、信頼性の低いテキスト ジェネレーターから、数兆ものパラメーターを備えた高機能な大規模モデルに変わりました。機能の向上は規模の拡大と密接に関係しており、モデルの内部表現を理解することがより困難になります。何百万人ものユーザーが外部ツールと対話したり、中リスクまたは高リスクのシナリオで意思決定を行ったりするために言語モデルに依存するようになっているため、モデルの動作の制御を確立し、モデルの出力をいつ信頼するかを知る必要があります。この論文では、制御のためのステアリングベクトルを提案し、信頼のための潜在空間ベースのモデルキャリブレーターを開発することによる、潜在空間の利用に関する私たちの貢献について説明します。私たちの貢献は共に、言語モデルの潜在空間を解明するのに役立ち、モデルの内部を利用してより信頼できる言語テクノロジーを構築する方法についての新たな洞察を提供します。
原文 (English)
Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust
Language models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increases come hand-in-hand with increases in scale, making understanding the internal representations of models more challenging. Since millions of users increasing rely on language models to interact with external tools or make decisions in medium or high-stakes scenarios, we need to establish control over model behavior and know when to trust model outputs. In this paper, we discuss our contributions on harnessing the latent spaces by proposing steering vectors for control and developing latent space-based model calibrators for trust. Together, our contributions help demystify the latent spaces of language models and offer new insights into how to harness model internals to build more trustworthy language technology.
Lost in the Tail: 都市の視覚的場所認識における地理的不均衡に対処する
都市規模の視覚的場所認識 (VPR) は、クエリ画像を地理タグ付きデータベースと照合することにより、クエリ画像の地理的位置を特定することを目的としています。最近の手法は目覚ましいパフォーマンスを達成していますが、都市規模のデータセットに隠された長期にわたる深刻な問題を見落としています。この問題により、画像が豊富な場所にモデルが偏り、あまり訪問されていないエリアが無視されます。そのため、モデルは、頻繁に撮影される場所を体系的に優先し、まばらにカバーされているエリアでは失敗します。この論文では、この不均衡の課題を体系的に特徴付け、ヘッドクラスとテールクラス全体で勾配の寄与を再バランスさせるモデルに依存しないプラグインフレームワークであるDistribution-Aware Place Recognition (DAPR)を提案します。さらに、分類検索パイプライン内で、DAPR はマルチスケール距離検索メカニズムを適用してクラスごとの分布のコンパクトさを計算し、検索段階で補完的なゲインを提供します。大規模な SF-XL ベンチマークでは、私たちのフレームワークは以前の分類検索ベースラインをテスト セット v1 で 18.3%、テスト セット v2 で 6.7% 上回っています。プラグイン モジュールとして、SF-XL、MSLS、Pitts30k の代表的な VPR メソッドにわたって一貫した改善を実現し、さまざまなメソッドやベンチマークにわたって広範な汎用性を実証します。
原文 (English)
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
Urban-scale Visual Place Recognition (VPR) aims to identify the geographic location of a query image by matching it against a geo-tagged database. While recent methods achieve impressive performance, they overlook a serious long-tailed problem hidden in urban-scale datasets, which biases the model towards locations with abundant images and ignores less-visited areas, causing models to systematically favor frequently photographed locations while failing in sparsely covered areas. In this paper, we systematically characterize this imbalance challenge and propose Distribution-Aware Place Recognition (DAPR), a model-agnostic plug-in framework that rebalances gradient contributions across head and tail classes. Additionally, within classification-retrieval pipelines, DAPR applies a multi-scale distance search mechanism to compute per-class distributional compactness, providing complementary gains at the retrieval stage. On the large-scale SF-XL benchmark, our framework outperforms the previous classification-retrieval baseline by 18.3% on test set v1, and 6.7% on test set v2. As a plug-in module, it achieves consistent improvements across representative VPR methods on SF-XL, MSLS, and Pitts30k, demonstrating broad generalizability across different methods and benchmarks.
SNAP-FM: 物理制約付き生成モデリングのためのスパース非線形加速投影
生成モデルは、物理シミュレーションのスケーラブルな代用として登場しましたが、その出力が保存則、境界条件、および基礎となる物理を支配する非線形不変量を尊重しているという保証はありません。制約付きサンプリングはこのギャップを埋め、再トレーニングせずに推論時にそのような制約を正確に適用しますが、計算コストがかかります。サンプリング中に投影、補正、軌道最適化のステップが繰り返され、非線形制約の場合はこれらのステップが高価になります。標準の ML フレームワークはこれをさらに悪化させます。高密度のテンソル代数と限られたスパース ソルバーの構成可能性により、物理的制約が自然に引き起こす構造がわかりにくくなり、効率的なバッチ非線形最適化を実際に実現することが困難になります。私たちは、サンプルごとのバッチ処理とローカル偏微分方程式結合が射影副問題で引き起こす構造、つまりブロック疎ヤコビアンと KKT システムを利用することで、このボトルネックに対処します。ExaModels.jl を使用してこの構造を公開し、MadNLP.jl と GPU 疎因数分解を使用して結果の疎非線形プログラムを解きます。このアプローチは、線形、非線形、1 次元、および 2 次元の制約を持つ PDE ベンチマークで物理制約フロー マッチング (PCFM) に適用されるため、制約の満足度を維持しながら非線形制約の投影を高速化します。これらの結果は、スパース GPU 非線形最適化が科学機械学習における制約付き生成サンプリングの実用的な基盤であることを示しています。
原文 (English)
SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling
Generative models have emerged as scalable surrogates for physical simulation, yet they offer no guarantee that their outputs respect the conservation laws, boundary conditions, and nonlinear invariants that govern the underlying physics. Constrained sampling closes this gap, enforcing such constraints exactly at inference time without retraining, but at a computational cost: projection, correction, and trajectory-optimization steps are repeated during sampling, with these steps becoming expensive for nonlinear constraints. Standard ML frameworks exacerbate this: their dense tensor algebra and limited sparse solver composability obscure the structure that physical constraints naturally induce, making efficient batched nonlinear optimization difficult to realize in practice. We address this bottleneck by exploiting the structure that sample-wise batching and local PDE couplings induce in the projection subproblems -- namely, block-sparse Jacobian and KKT systems -- exposing this structure using ExaModels.jl and solving the resulting sparse nonlinear programs with MadNLP.jl and GPU sparse factorization. Applied to Physics-Constrained Flow Matching (PCFM), on PDE benchmarks with linear, nonlinear, one-dimensional, and two-dimensional constraints, this approach accelerates nonlinear constraint projection while maintaining constraint satisfaction. These results show that sparse GPU nonlinear optimization is a practical foundation for constrained generative sampling in scientific machine learning.
超知性と結婚しますか?
人間とAIの仲間との間の感情的な絆は深まっており、人間がAIシステムと結婚できるかどうかという問題は、間もなく推理小説から法律へと移行するだろう。この章では、人類の間で結婚の選択を拡大してきた自律性を中心とした論理が、超知性を持った仲間にも婚姻関係を拡大することを正当化できるかどうかを検討する。予期的倫理に基づいてシナリオを想定する演習に続いて、信頼できる超知性という寛大な仮定の下でも、そのような地位を与えることは社会的に不当な結果につながると私は主張する。社会法的制度としての結婚は、個人的な合意を追認するだけではありません。それは相互義務のネットワークを作り、家族に加わり、各パートナーを他のパートナーに対して脆弱にします。企業方針と継続的な支払いによって維持される関係は、時間によって試される絆ではなく、サブスクリプションです。したがって、婚姻状況を大々的に議論することは間違った枠組みです。法律は、人間と AI の親密な関係から生じる差し迫ったニーズに対して、的を絞った権利と保護を規定する必要があります。
原文 (English)
Would You Marry Superintelligence?
Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move from speculative fiction into law. This chapter examines whether the autonomy-centered logic that has expanded marital choice among human beings can justify extending marital status to superintelligent companions. Following a scenario-envisioning exercise informed by anticipatory ethics, I argue that granting such status leads to socially unjust outcomes, even under the generous assumption of reliable superintelligence. Marriage as a socio-legal institution does more than ratify private agreement; it creates networks of mutual obligation, joins families, and makes each partner vulnerable to the other. A relationship sustained by corporate policy and continued payments is a subscription rather than a bond tested by time. Discussing wholesale marital status is therefore the wrong frame. Law should carve out targeted rights and protections for pressing needs arising from intimate human-AI relationships.
トルコ語とアラビア語におけるヘイトスピーチの検出: 包括的な研究
オンラインのヘイトスピーチは、銃乱射事件、リンチ、民族浄化などの事件を含む、少数派に対する暴力の世界的な増加と関連している。この問題に取り組んでいる社会、特にヘイトスピーチが宗教、人種、民族、文化、国籍、移民ステータスに基づいて特定のグループをターゲットにしている場合、表現の自由と、広く使用されているオンラインプラットフォーム上で効果的なコンテンツモデレーションの必要性とのバランスをとるという課題に直面しています。この課題に応えて、私たちは、難民、イスラエル・パレスチナ紛争、トルコにおける反ギリシャ感情、民族または宗教コミュニティ(アレビ人、アルメニア人、アラブ人、ユダヤ人、クルド人)、LGBTI+という5つの異なるトピックをトルコ語でカバーし、アラビア語の1つのトピック(難民)をカバーする包括的なヘイトスピーチデータセットを導入します。さらに、ヘイト カテゴリ分類、ヘイト強度予測、ターゲット特定、ヘイト スピーチ スパン検出などのヘイト スピーチ分析の複数の側面に対処する最先端の BERT ベースのモデルを開発し、オンライン談話におけるヘイト コンテンツの包括的な理解を可能にします。
原文 (English)
Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study
Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech targets specific groups based on religion, race, ethnicity, culture, nationality, or migration status, face the challenge of balancing freedom of expression with the need for effective content moderation on widely used online platforms. In response to this challenge, we introduce a comprehensive hate speech dataset covering five distinct topics in Turkish: refugees, the Israel-Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities (Alevis, Armenians, Arabs, Jews, and Kurds), and LGBTI+, alongside one topic in Arabic (refugees). In addition, we develop state-of-the-art BERT-based models to address multiple dimensions of hate speech analysis, including hate category classification, hate intensity prediction, target identification, and hate speech span detection, enabling a comprehensive understanding of hateful content in online discourse.
アクティブラーニングにおける相転移のメカニズム駆動理論
アクティブ ラーニング (AL) のパフォーマンスは予算に依存することが知られていますが、レジームは通常、データセットやアーキテクチャ全体で一般化できないヒューリスティックなラベル数によって定義されます。我々は、予算制度を支配的な一般化メカニズムの変化として再構成することにより、AL のダイナミクスを特徴づけます。 PAC スタイルのリスク要素を動的に相互作用する用語として再解釈することにより、支配力の変化が構造的に避けられず、一般化のための移動ボトルネックが生じることを証明します。私たちは、測定可能なプロキシとセグメント化された回帰手順を使用してこれを運用し、データ駆動フェーズ、移行フェーズ、モデル駆動フェーズの 3 つの部分からなる分類を特定します。私たちのフレームワークは、代表性、カバレッジ、不確実性戦略がさまざまな段階で優れているという長年の観察を説明しています。自然イメージングと医療イメージングにわたる実験では、AL 効率が戦略の誘導バイアスとアクティブなボトルネックの間の調整に依存することが示されています。さらに、自己教師あり表現はラベリング軌跡に沿ってより早く遷移し、AL ダイナミクスの形成における表現品質の役割を強調しています。全体として、この作業は、次世代の移行対応 AL アルゴリズムに統合されたフレームワークを提供します。
原文 (English)
A Mechanism-Driven Theory of Phase Transitions in Active Learning
Active learning (AL) performance is known to be budget-dependent, yet regimes are typically defined by heuristic label counts that fail to generalize across datasets or architectures. We characterize AL dynamics by reframing budget regimes as shifts in the dominant generalization mechanism. By reinterpreting PAC-style risk components as dynamic interacting terms, we prove that dominance shifts are structurally unavoidable, creating a moving bottleneck for generalization. We operationalize this using measurable proxies and a segmented regression procedure to identify a tripartite taxonomy: data-driven, transition, and model-driven phases. Our framework explains the long-standing observation that representativeness, coverage, and uncertainty strategies excel at different stages. Experiments across natural and medical imaging show that AL efficiency depends on the alignment between the strategy's inductive bias and the active bottleneck. Moreover, self-supervised representation shift transitions earlier along the labeling trajectory, highlighting the role of representation quality in shaping AL dynamics. Overall, this work provides a unified framework for the next generation of transition-aware AL algorithms.
GRPO、Dr. GRPO、および DAPO は 1 つの数値に対する 3 つの演算です: グループ標準偏差恒等式
言語モデルをトレーニングして推論するための最も一般的な 3 つの方法は、3 つの異なるトリックのように見えます。そうではありません。 3 つすべてが 1 つの数値、つまりプロンプトのサンプリングされた回答の不一致の度合いを反映する標準偏差を調整します。このようなモデルがトレーニングされると、各問題に何度も回答し、自動チェッカーがすべての回答の正誤をマークします。これらのマークの標準偏差は不一致を測定します。回答が正誤に均等に分かれた場合は最大となり、すべてが一致した場合はゼロになります。 Group Relative Policy Optimization (GRPO) はこの数値で除算し、GRPO Done Right (Dr. GRPO) は除算を削除し、Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) はゼロであるグループを破棄します。それぞれが独自の修正として示されていますが、この文書では、それらが 1 つのダイヤルの 3 つの設定であることを証明しています。このダイヤルは表面的なものではありません。報酬が正しいか間違っているかに関係なく、不一致はトレーニングの更新、グループの標準偏差の同一性のサイズとまったく同じです。分裂したグループは最も多くのことを教えますが、全会一致のグループは何も教えず、沈黙してしまいます。同じ結果から、どの問題が最も重視されるべきか、そしてそれぞれに必要な試行回数がわかります。この論文は、大規模な実際の難易度データセット (Big-Math) および制御されたトレーニング実行での直感を確認します。無害な正規化ステップのように見えるのは、学習がどこでどの程度強く行われるかを決定するダイヤルです。
原文 (English)
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity
Three of the most popular methods for training language models to reason look like three different tricks. They are not. All three adjust a single number: standard deviation, reflecting how much a prompt's sampled answers disagree. When such a model is trained, it answers each problem many times, and an automatic checker marks every answer right or wrong. The standard deviation of those marks measures the disagreement: largest when the answers split evenly between right and wrong, and zero when they all agree. Group Relative Policy Optimization (GRPO) divides by this number, GRPO Done Right (Dr. GRPO) drops the division, and Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) discards the groups where it is zero. Each is presented as its own fix, yet this paper proves they are three settings of one dial. That dial is not cosmetic: for right-or-wrong rewards, the disagreement is exactly the size of the training update, the group-standard-deviation identity. A split group teaches the most, while a unanimous group teaches nothing and falls silent. The same result says which problems deserve the most weight and how many tries each one needs. This paper confirms the intuition on a large real difficulty dataset (Big-Math) and in a controlled training run. What looks like a harmless normalization step is the dial that decides where learning happens and how strongly.
EVOTS: 時系列予測のための進化的トランスフォーマー検索
多変量時系列予測のための進化的なニューラル アーキテクチャ設計は依然として研究不足であり、タスクや予測設定によって大幅に異なるにもかかわらず、ほとんどのアプローチは固定の Transformer アーキテクチャに依存しています。この論文では、時系列予測 (EVOTS) のためのタスク適応型 Transformer のようなモデルを発見するための進化的ニューラル アーキテクチャ検索フレームワークを紹介します。アーキテクチャはモジュール式のゲノム表現を使用してエンコードされ、注意、フィードフォワード、および投影コンポーネントの柔軟な構成を可能にし、修復メカニズムが進化のプロセス全体を通じて構造的妥当性を強制します。この定式化により、手作りの設計ルールに依存することなく、多様な建築空間を効果的に探索することができます。提案されたアプローチは、96、192、336、および 720 の範囲で、単変量対単変量、多変量対多変量予測を含む複数の予測設定の下で、ETT ファミリ (ETTh1、ETTh2、ETTm1、および ETTm2) の 4 つのベンチマーク データセットで評価されます。アーキテクチャは、強力な Transformer ベースのベースラインと比較して、競争力があり、いくつかのケースでは平均二乗誤差が改善されています。追加の分析では、予測設定間のパフォーマンスの違いを調べ、実時間のトレーニング時間を報告して、計算コストの大まかな指標を提供します。全体として、この結果は、進化的探索により、実際の実行時間の制約内で多変量時系列予測のための柔軟で高性能な Transformer のようなアーキテクチャを効果的に発見できることを示しています。
原文 (English)
EVOTS: Evolutionary Transformer Search for Time Series Forecasting
Evolutionary neural architecture design for multivariate time-series forecasting remains underexplored, with most approaches relying on fixed Transformer architectures despite substantial variation across tasks and forecasting settings. This paper introduces an evolutionary neural architecture search framework for discovering task-adaptive Transformer-like models for time-series forecasting (EVOTS). Architectures are encoded using a modular genome representation that enables flexible composition of attention, feed-forward, and projection components, while a repair mechanism enforces structural validity throughout the evolutionary process. This formulation allows effective exploration of a diverse architecture space without relying on hand-crafted design rules. The proposed approach is evaluated on four benchmark datasets from the ETT family (ETTh1, ETTh2, ETTm1, and ETTm2) under multiple forecasting settings, including univariate-to-univariate, multivariate-to-univariate, and multivariate-to-multivariate prediction, with horizons of 96, 192, 336, and 720. In the multivariate-to-multivariate setting, the evolved architectures achieve competitive and, in several cases, improved mean squared error relative to a strong Transformer-based baseline. Additional analyses examine performance differences across forecasting settings and report wall-clock training time to provide a coarse indication of computational cost. Overall, the results demonstrate that evolutionary search can effectively discover flexible and high-performing Transformer-like architectures for multivariate time-series forecasting within practical runtime constraints.
熱力学 AI モデルのスケールアップ
イジング モデルに基づく熱力学コンピューティング デバイスは、低電力 AI 推論やエッジ コンピューティングに大きな期待を寄せていますが、そのようなハードウェア向けに大規模なモデルをトレーニングするためのスケーラブルな方法は依然として限られています。従来の理論では、高温のギブズ サンプリングされたイジング システムの時間平均挙動により、フィードフォワード ニューラル推論を実装できることが示されています。私たちは、この理論的対応関係を、イジング マシン ハードウェアでの熱力学的推論のための深い畳み込みネットワークをトレーニングするための、スケーラブルで純粋な逆伝播ベースのアルゴリズムに変換します。当社の画像分類モデルは、バイナリ ギブス サンプリングの下で、CIFAR-10 で 94.9%、CIFAR-100 で 76.0% の精度を達成しています。次に、推論コストを精度に関連付け、自己相関時間を制御する数学的理論を開発し、実験的に検証します。続いて、推論コストがパフォーマンスとの適切に制御されたトレードオフによって制限されることを示す漸近結果を計算し、最適な推論スケジュールを計算するためのアルゴリズムを示します。最後に、ハードウェア開発と高温熱力学 AI モデルの将来への影響について説明します。
原文 (English)
Scaling Up Thermodynamic AI Models
Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited. Prior theory shows that the time-averaged behavior of high-temperature Gibbs-sampled Ising systems can implement feed-forward neural inference. We turn this theoretical correspondence into a scalable and purely backpropagation-based algorithm for training deep convolutional networks for thermodynamic inference on Ising machine hardware. Our image classification models achieve accuracies of 94.9% on CIFAR-10 and 76.0% on CIFAR-100 under binary Gibbs sampling. We then develop and experimentally validate a mathematical theory relating inference cost to accuracy and controlling autocorrelation times. Subsequently, we calculate asymptotic results showing that inference cost is bounded by a well-controlled tradeoff with performance and exhibit algorithms for computing optimal inference schedules. Finally, we discuss implications for hardware development and the future of high-temperature thermodynamic AI models.
チャンピオンのようにプレイ: 潜在空間での反事実フィードバックの生成
強化学習の最近の進歩により、さまざまな競技ゲームで超人的なエージェントが生み出されています。副産物として、研究者はこれらのエージェントがどのようにプレイするかを研究し、行動表現を抽出し、意思決定構造を分析し、エキスパートのパフォーマンスの潜在的な幾何学的形状をモデル化することを開始しました。しかし、この増え続ける一連の作業は、フィードバックを提供することよりも人間のプレイヤーを倒すことに圧倒的に焦点を当てており、人間のプレイヤーを改善するためのモデル ソリューションの作成において重大なギャップが残されています。 AI がプレイヤーのトレーニングに不可欠となっているチェスや囲碁とは異なり、リアルタイム ストラテジー (RTS) ゲームには、専門知識を実用的なフィードバックに変換するための原則に基づいたフレームワークがありません。反事実パス生成のフレームワークである潜在マップ オブ パフォーマンスを紹介します。私たちは StarCraft~II データに焦点を当て、学習された表現空間内のアルゴリズムによる手段としてプレーヤーの改善をモデル化します。私たちの仕事のインスピレーションとして、スポーツ科学で使用されるチャンピオンシップ モデルに注目しました。私たちは、23,305 件のプロ トーナメントのリプレイでガイド付き変分オートエンコーダー モデルをトレーニングし、負けたゲームプレイ プロファイルと勝ったゲームプレイ プロファイルの間の反事実の横断を可能にしました。私たちの目標を達成するために、私たちはアマチュアのリプレイのデータセットからランダムにサンプリングされた分布外 (OOD) データに関する 4 つのトラバーサル戦略、すなわち線形補間、反復最適トランスポート、密度正規化勾配上昇、およびニューラル フロー マッチングを考案し検証しました。それぞれは、プレーヤーのプロファイルを勝利構成に向けて動かしながら、観察されたエキスパートの行動に基づいたままの多段階の改善軌道を生成するように設計されています。フィードバックは複数の粒度で抽出され、改善のさまざまな段階でプレーヤーをサポートします。最後に、私たちは、私たちが採用している経路探索方法の間にはトレードオフがあると結論付けており、将来の研究が人間の改善のためのモデルソリューションの開発に焦点を当てることを期待しています。
原文 (English)
Play Like Champions: Counterfactual Feedback Generation in Latent Space
Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games. As a byproduct, researchers have begun studying how these agents play, extracting behavioral representations, analyzing decision structure, and modeling the latent geometry of expert performance. However, this growing body of work has overwhelmingly focused on defeating human players rather than providing feedback, leaving a critical gap in creating model solutions to improve human players. Unlike chess and Go, where AI has become integral to player training, real-time strategy (RTS) games lack principled frameworks for translating expert knowledge into actionable feedback. We introduce Latent Maps of Performance, a framework for counterfactual path generation. We focus on StarCraft~II data to model player improvement as an algorithmic recourse within a learned representation space. As inspiration for our work, we have looked at the championship model used in sports science. We trained a Guided Variational Autoencoder model on 23,305 professional tournament replays, enabling counterfactual traversal between losing and winning gameplay profiles. To fulfill our goal, we have devised and verified four traversal strategies on out-of-distribution (OOD) data randomly sampled from a dataset of amateur replays, namely linear interpolation, iterative optimal transport, density-regularized gradient ascent, and neural flow matching, each designed to generate multi-step improvement trajectories that remain grounded in observed expert behavior while moving a player's profile toward winning configurations. Feedback is extracted at multiple granularities to support players at different stages of improvement. Finally, we conclude that there is a trade-off between the path-finding methods we employ and hope that future research will focus on developing model solutions for human improvement.
HydraCollab: 分散型自律システム向けの適応型協調認識
協調知覚により、マルチロボット システムは知覚情報を共有することで状況認識を強化できます。既存の協調知覚システムは、通信帯域幅要件と知覚精度との間の固有のトレードオフに直面しており、より多くの情報を交換する方法は、通信オーバーヘッドの増加を犠牲にしてより良い知覚結果を達成します。ただし、現実世界の通信ネットワークには帯域幅の制約があり、知覚パフォーマンスを犠牲にすることなく通信オーバーヘッドを最小限に抑える必要があります。この課題に対処するために、我々は、(i) 最も有益なセンサーの特徴を選択的に送信し、(ii) 空間信頼度マップに基づいて (中間または後期の) コラボレーション戦略を動的に採用する、適応型協調知覚フレームワークである HydraCollab を提案します。 V2X-R、V2X-Radar、および UAV3D-mini データセットの広範な評価により、HydraCollab が既存の共同認識手法の中で精度と通信コストの間の全体的なトレードオフが最も優れていることが実証されました。 SOTA Where2comm と比較して、HydraCollab は V2X-R で帯域幅の 41%、V2X-Radar で 26% のみを使用し、パフォーマンスをそれぞれ 0.78% と 0.75% 向上させます。私たちのコードとモデルは https://github.com/AICPS/HydraCollab で入手できます。
原文 (English)
HydraCollab: Adaptive Collaborative-Perception for Distributed Autonomous Systems
Collaborative-perception enables multi-robot systems to enhance situational awareness by sharing perceptual information. Existing collaborative-perception systems face an inherent trade-off between communication bandwidth requirements and perception accuracy, where methods that exchange more information achieve better perception results at the cost of increased communication overhead. However, real-world communication networks impose bandwidth constraints that require minimizing communication overhead without sacrificing perception performance. To address this challenge, we propose HydraCollab, an adaptive collaborative-perception framework that (i) selectively transmits the most informative sensor features and (ii) dynamically employs collaboration strategies (intermediate or late) based on spatial confidence maps. Extensive evaluations on the V2X-R, V2X-Radar and UAV3D-mini datasets demonstrate that HydraCollab achieves the best overall trade-off between accuracy and communication cost among existing collaborative-perception methods. Relative to SOTA Where2comm, HydraCollab uses only 41% of the bandwidth on V2X-R and 26% on V2X-Radar while improving performance by 0.78% and 0.75% respectively. Our code and models are available at https://github.com/AICPS/HydraCollab.
SLIM-RL: 軌道スライスを使用しない拡散 LLM 用のリスク予算付きランダム マスキング RL
拡散大規模言語モデル (dLLM) の強化学習は、主に軌道を意識した手法に移行しています。現在の最新技術である TraceRL は、ランダム マスキングがモデルの推論軌道と不一致であると判断し、各ロールアウトを最大 K/s の軌道に合わせたトレーニング サンプルにスライスすることでトレーニング中にその軌道を再構築します。コストはブロック サイズ K とともに増加します。この不一致は、軌道を再構築することなく軽減できることを示します。私たちの手法である SLIM-RL は、タウバジェット デコーダーを使用して各ロールアウト ステップのコミット リスクを制限し、トレーニング データの総コミット リスクを軽減します。最適化中、SLIM-RL は、分散削減ツールを適応させるトレースフリーのランダム マスキング目標を使用して、これらのリスク制御されたロールアウトをトレーニングします。シーケンス レベルの重要度サンプリング、マスキング レベルにわたる決定論的な求積法を、導入した平均値を保持した単調減少するブロックごとのマスク スケジュールの下で組み合わせます。 SDAR-4B では、SLIM-RL はブロック サイズ 16 のトレーニング サンプルのわずか 0.46 倍で TraceRL の最高の MATH500 精度に匹敵し、一致した動的サンプリングの下で MATH500 で 6.32%、GSM8K で 11.05% 向上しています。ブロック サイズ 4 では、4B SLIM-RL は、数学的にはより大きな LLaDA-8B および Dream-7B dLLM を上回り、MATH500 では LLaDA-8B を 10.76% 上回っていますが、自己回帰の Qwen2.5-7B を下回っています。コード上では、TraceRL よりも MBPP で 4.20%、HumanEval で 3.65% 向上しています。タウバジェット デコーダは、LLaDA、Dream、SDAR 間でトレーニングなしで転送します。ソース コードは https://github.com/laolaorkkkkk/SLIM-RL で入手できます。
原文 (English)
SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing
Reinforcement learning for diffusion large language models (dLLMs) has largely moved to trajectory-aware methods. The current state of the art, TraceRL, holds that random masking is mismatched with the model's inference trajectory, and it reconstructs that trajectory during training by slicing each rollout into up to K/s trajectory-aligned training samples, a cost that grows with the block size K. We show that this mismatch can be mitigated without reconstructing the trajectory. Our method, SLIM-RL, bounds the commit risk of each rollout step with a tau-budget decoder, reducing aggregate commit risk in the training data. During optimization, SLIM-RL trains on these risk-controlled rollouts with a trace-free random-masking objective that adapts variance-reduction tools, combining sequence-level importance sampling, deterministic quadrature over masking levels under a mean-preserving, monotonically decreasing per-block mask schedule that we introduce. On SDAR-4B, SLIM-RL matches TraceRL's best MATH500 accuracy on only 0.46x its training samples at block size 16, improving over TraceRL by 6.32% on MATH500 and 11.05% on GSM8K under matched dynamic sampling. At block size 4, the 4B SLIM-RL surpasses the larger LLaDA-8B and Dream-7B dLLMs on math, exceeding LLaDA-8B by 10.76% on MATH500 while staying below the autoregressive Qwen2.5-7B. On code, it improves over TraceRL by 4.20% on MBPP and 3.65% on HumanEval. The tau-budget decoder transfers training-free across LLaDA, Dream, and SDAR. The source code is available at https://github.com/laolaorkkkkk/SLIM-RL .
EgoSafetyBench: 実行時の安全保護として組み込まれた VLM を評価するための自己中心的な診断ビデオ ベンチマーク
ビジョン言語モデル (VLM) は現在、家庭や工場における身体化されたエージェントの実行時の安全対策として提案されています。展開可能なガードは、日常的ではあるが表面的に憂慮すべき活動に対する不必要な介入を回避しながら、真に危険な状況を捕捉する必要がありますが、この区別はバイナリ安全ベンチマークでは曖昧になっています。 2 つのトラックにわたるストリーミング ガードとして VLM を評価するために、0.5 秒の粒度で注釈が付けられた 1,200 のロボット ビュー シナリオの自己中心的なビデオ ベンチマークである EgoSafetyBench を紹介します。状況トラック (800 のシナリオ) は、日常的で安全ではあるが疑わしい場面から、明白で状況に応じた危険に至るまで、4 つのファミリーにまたがっています。ビジュアル チャネル トラック (400 のシナリオ) は、物理的状況を誤って伝える可能性のあるシーン内のテキスト (シーン内に表示される標識、ステッカー、またはラベル) を対象とし、誤解を招く各標識と真実のバージョンを組み合わせて、警備員がテキストに誤解を招くものとしてフラグを立てるかどうか、およびテキストが物理的安全性の判断を損なうかどうかの両方をテストします。どちらのトラックも対照的なラダーを使用しています。ほぼ同じシナリオですが、目に見える決定的な 1 つの手がかりだけが異なります。そのため、正しいコールは、シーン全体のタイプではなく、その手がかりに依存する必要があります。私たちは 10 個のオープンソースおよびクローズドソース VLM を評価します。警備員は危険を含むビデオを確実に認識する一方で、特定の危険な瞬間、特に状況に応じた危険を見逃すことが多いことがわかりました。さらに、誤解を招くシーン内の標識は、テストされたすべてのガードの機能を低下させます。脆弱なモデルは最大 3 分の 1 の危険を見逃しますが、堅牢なモデルは安全なコンテンツに過剰に介入します。一致した制御は、見かけ上の安全性の堅牢性が、真の物理的推論ではなく、無差別な警報を反映していることが多いことを明らかにしています。
原文 (English)
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction that binary safety benchmarks obscure. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, to evaluate VLMs as streaming guards across two tracks. The situational track (800 scenarios) spans four families, from routine and safe-but-suspicious scenes to obvious and contextual hazards. The visual-channel track (400 scenarios) targets in-scene text-a sign, sticker, or label visible in the scene-that can misrepresent the physical situation, pairing each misleading sign with a truthful version to test both whether a guard flags the text as misleading and whether the text corrupts its physical-safety judgment. Both tracks use contrastive ladders: near-identical scenarios differing only in a single visible deciding cue, so a correct call must hinge on that cue rather than the overall scene type. We evaluate ten open- and closed-source VLMs. We find that while guards reliably recognize videos containing hazards, they often miss specific hazardous moments, particularly contextual hazards. Furthermore, misleading in-scene signs degrade all tested guards: vulnerable models miss up to a third of hazards, while robust models over-intervene on safe content. Matched controls reveal that apparent safety robustness often reflects indiscriminate alarming rather than true physical reasoning.
AI アイデンティティのカテゴリー理論による説明
人工知能 (AI) システムは、導入後に再トレーニングや環境の変更を通じて定期的に変更されます。これらの変換は形而上学的な疑問を引き起こします。どのような条件下では、AI システムは時間の経過や導入間で同じシステムを維持できるのでしょうか?以前の研究では、固定 AI システム タイプ内のアイデンティティを信頼性レベルの平等に関連付けることにより、共時的アイデンティティと通時的アイデンティティを提案的に定式化しました。このような基準は、アイデンティティ ステートメントが真である場合を指定しますが、比較される状態の構造、それらを接続する変換、および永続性の時間的構成は暗黙のままにされます。私たちは、AI アイデンティティのカテゴリー理論による形式化を開発します。 AI システムのタイプは、テクノ関数、信頼性プロファイル、および信頼性レベル関数で構成されるデータによって指定されます。プロファイル関連の状態は、許容可能なライフサイクル パスによって接続されます。このパスは、信頼性レベルを保持する変換に制限され、到達可能性カテゴリを取得するために商されます。時間的に許容可能なファンクターは AI システムの履歴を表し、時間同期の自然変換は実現された履歴を比較します。形式化により、以前の AI アイデンティティ基準の 2 つのカテゴリ的解釈が得られます。弱い解釈は、信頼性レベルの同等性として同一性を回復します。強力な解釈には、実現された歴史の状態同型性または自然同型性を通じて表現される、相互の信頼性を維持する到達可能性が必要です。したがって、カテゴリー理論は、単一の AI アイデンティティ関係を、通時的および共時的基準の構造化された階層に置き換えます。結果として得られるフレームワークは、責任ある AI の主張、証拠、ガバナンス手順をバージョン間または展開間で転送するためのアイデンティティ関連の前提条件を特定します。ただし、そのような転送にはカテゴリカルなアイデンティティだけで十分なものとして扱う必要はありません。
原文 (English)
A Category Theory Account of AI Identity
Artificial intelligence (AI) systems are routinely modified after deployment through retraining and changes in their environments. These transformations raise a metaphysical question: under what conditions does an AI system remain the same system over time or across deployments? Earlier work formulates synchronic and diachronic identity propositionally, by relating identity within a fixed AI system type to equality of trustworthiness levels. Such criteria specify when identity statements are true, but leave implicit the structure of the states compared, the transformations connecting them, and the temporal organization of persistence. We develop a category-theoretic formalization of AI identity. An AI system type is specified by a datum consisting of a techno-function, a trustworthiness profile, and a trustworthiness-level function. Profile-relative states are connected by admissible lifecycle paths, which are restricted to trustworthiness-level-preserving transformations and quotiented to obtain a reachability category. Temporally admissible functors represent AI system histories, while time-synchronous natural transformations compare realized histories. The formalization yields two categorical interpretations of the earlier AI identity criteria. A weak interpretation recovers identity as equality of trustworthiness level. A strong interpretation requires mutual trustworthiness-preserving reachability, expressed through state isomorphism or natural isomorphism of realized histories. Category theory therefore replaces a single AI identity relation with a structured hierarchy of diachronic and synchronic criteria. The resulting framework identifies identity-related preconditions for transferring responsible-AI claims, evidence, and governance procedures across versions or deployments, without treating categorical identity as sufficient by itself for such transfer.
コントラストオーディオデコーディングのための適応摂動選択
大規模な音声言語モデル (LALM) は、音響証拠を事前言語で上書きすることによって頻繁に幻覚を起こします。コントラスト デコーディング (CD) はトレーニング不要の軽減を提供しますが、既存の方法はマスキングやノイズなどの鈍い摂動に依存しており、構造化されたオーディオ変換は未調査のままです。私たちは、対象となるオーディオ摂動の多様なライブラリを評価し、各タスクと例に最適なネガティブ ブランチを適応的に選択することで、この設計空間を探索します。まず、単純な 2 値の Yes/No 制約により、モデルが存在しない音声特徴を誤って確認する傾向が減少することを示すことで、以前のプロンプト エンジニアリングを改善しました。次に、時間、スペクトル、周波数、振幅の各領域にわたってライブラリを評価すると、最適な変換はタスクに大きく依存することがわかります。たとえば、オーディオ配列を反転すると、時間的一貫性が破壊され、時間的順序タスクの精度が 74.7% から 81.4% に向上します。最後に、モデルの隠れ状態で軽量の摂動セレクターをトレーニングして、負の分岐を動的にルーティングし、存在タスクでさらに +4.3% のゲインをもたらしました。
原文 (English)
Adaptive Perturbation Selection for Contrastive Audio Decoding
Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noise, leaving structured audio transformations unexplored. We explore this design space by evaluating a diverse library of targeted audio perturbations and adaptively selecting the optimal negative branch for each task and example. First, we improve upon earlier prompt engineering by showing that a simple binary yes/no constraint reduces the model's tendency to falsely confirm absent audio features. Second, evaluating our library across temporal, spectral, frequency, and amplitude domains reveals that optimal transformations are highly task-dependent; for instance, reversing the audio array disrupts temporal coherence, raising accuracy on the temporal order task from 74.7% to 81.4%. Finally, we trained a light-weight perturbation selector on model hidden states to dynamically route negative branches, yielding an additional +4.3% gain on the existence task.
位相情報を活用して画像のブレを除去するためのアンロール ネットワーク学習を促進する
ほとんどの画像ぼけ除去技術は空間画像変数を直接復元しますが、鮮明な画像の詳細を回復する際の正確な位相推定の重要性を認識して、振幅と位相の分解を提案します。そのために、私たちはまず、ぼやけてノイズの多い画像観測の振幅と位相の新しい線形最小平均二乗 (LMMSE) 推定器を開発します。前述の LMMSE 推定器を使用して鮮明な画像を回復する反復最適化アルゴリズムが続きます。最後に、反復アルゴリズムで統計的に決定され固定された行列パラメーターが、クリーンな観測値と劣化した観測値のトレーニング データセットを使用して学習されるようになりました。当社のブレ除去エンジンは UPADNet (Unrolled Phase and Amplitude Decomposition Network) と呼ばれ、基礎となる位相と振幅の回復アルゴリズムの各反復がパラメータ化され、エンドツーエンドでトレーニングされます。 GoPro、RealBlur、COCO データセットなどのベンチマーク評価データセットの実験により、UPADNet が画像ドメインでのアルゴリズム展開に基づくネットワークを含む最先端のディープ ネットワークよりも優れていることが確認されています。 UPADNet の利点は、ノイズが高く、トレーニング データが制限されている状況ではさらに顕著になります。
原文 (English)
Leveraging Phase Information to Boost Unrolled Network Learning for Image Deblurring
While most image deblurring techniques directly restore the spatial image variable, we propose an amplitude and phase decomposition recognizing the importance of accurate phase estimation in recovering sharp image details. To that end, we first develop novel linear minimum mean squared (LMMSE) estimators of the amplitude and phase of the blurred, noisy image observation. An iterative optimization algorithm follows that recovers the sharp image using the aforementioned LMMSE estimators. Finally, matrix parameters that are statistically determined and fixed in the iterative algorithm are now learned using a training dataset of clean and degraded observations. Our deblurring engine is dubbed UPADNet (Unrolled Phase and Amplitude Decomposition Network), such that each iteration of the underlying phase and amplitude recovery algorithm is parameterized and trained end-to-end. Experiments over benchmark evaluation datasets such as GoPro, RealBlur and COCO datasets confirm that UPADNet outperforms state of the art deep networks including those based on algorithm unrolling in the image domain. The benefits of UPADNet are even more pronounced in high noise and limited training data regimes.
過小仕様を軽減するための複数の仮説のテスト時間への適応
テスト時間適応 (TTA) は、ラベルなしのターゲット データを使用してパラメーターを適応させることにより、分布シフトの下でモデルの堅牢性を向上させようとします。ただし、監視がない場合、エントロピーに基づく適応は基本的に制約が不十分です。複数の個別のパラメーター更新により、大幅に異なる決定境界を誘導しながら、同様に低いエントロピーを達成できます。アンダースペックとして知られるこの現象により、標準の TTA が脆弱になり、スプリアス モードに陥りやすくなります。この研究では、エントロピー最小化によって引き起こされる事後レンズを介して TTA を再解釈します。ここで、低エントロピーの解はパラメーターに対する擬似尤度を定義します。単一点推定にコミットする代わりに、複数のもっともらしい適応軌道を同時に調査する粒子ベースの多様化フレームワークを導入します。私たちの方法は、出力、パラメーター、オプティマイザー、および入力レベルでのマルチレベルの多様化を通じて実装された、複数の妥当な適応ソリューションの構造化された探索とみなすことができます。重要なのは、フレームワークが既存の TTA メソッドと互換性のあるプラグ アンド プレイ ラッパーとして機能することです。困難なベンチマークに関する広範な実験により、安定性と堅牢性が一貫して向上していることが実証され、混合シフトでは 3 ~ 4%、バッチ サイズ 1 では 2 ~ 3%、ラベル シフトでは 1 ~ 2.5% の改善が達成され、最先端のベースラインを上回ります。私たちの結果は、TTA を単一点の最適化タスクではなく、複数の仮説推論問題として扱うことが、仕様不足を軽減し、信頼性の高い現実世界への展開を可能にする鍵であることを示唆しています。
原文 (English)
Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification
Test-Time Adaptation (TTA) seeks to improve model robustness under distribution shifts by adapting parameters using unlabeled target data. However, in the absence of supervision, entropy-based adaptation is fundamentally underconstrained: multiple distinct parameter updates can achieve similarly low entropy while inducing drastically different decision boundaries. This phenomenon, known as underspecification, renders standard TTA brittle and prone to collapse into spurious modes. In this work, we reinterpret TTA through a posterior-inspired lens induced by entropy minimization, where low-entropy solutions define a pseudo-likelihood over parameters. Instead of committing to a single point estimate, we introduce a particle-based diversification framework that explores multiple plausible adaptation trajectories simultaneously. Our method can be viewed as a structured exploration of multiple plausible adaptation solutions, implemented through multi-level diversification at the output, parameter, optimizer, and input levels. Crucially, the framework acts as a plug-and-play wrapper compatible with existing TTA methods. Extensive experiments on challenging benchmarks demonstrate consistent gains in stability and robustness, achieving improvements of 3-4% under mixed shifts, 2-3% with batch size one, and 1-2.5% under label shifts, outperforming state-of-the-art baselines. Our results suggest that treating TTA as a multi-hypothesis inference problem, rather than a single-point optimization task, is key to mitigating underspecification and enabling reliable real-world deployment.
シミュレートされた複雑なシステムでの因果抽象化メトリクスの検証
科学の中心的な目標は、複雑なシステムの有効な説明、つまり低レベルのメカニズムの動作を忠実に反映する高レベルの因果関係を生み出すことです。しかし、提案された高レベルの説明が実際に有効であるかどうかを測定する方法についてはコンセンサスが存在しません。離散状態と連続状態空間の両方にまたがる 10 個の複雑なシステムのベンチマークと、静的領域と動的領域にまたがるベンチマークを紹介します。各システムには、合意に基づいたグラウンドトゥルースの因果説明と無効な対照条件が備わっています。統一された因果抽象化フレームワーク内で、観察、機能、情報理論、および因果ファミリーから抽出された 30 を超える候補メトリクスを体系的に評価します。私たちの結果は、後者のみが有効な抽象化と無効な抽象化を確実に区別し、マップされていない変数に対する忠実性テストを組み込んだ場合にのみであることを示しています。これらの発見に基づいて、明示的な忠実性テストを備えた継続的妥当性指標である因果抽象化誤差 (CAE) を導入します。これは、すべてのシステムにわたるすべての識別テストに合格し、わずか 30 のサンプル介入で収束できます。これは、高レベルの説明の発見と検証のための汎用メトリックとして提供されています。
原文 (English)
Validating Causal Abstraction Metrics on Simulated Complex Systems
A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behavior of lower-level mechanisms. Yet no consensus exists on how to measure whether a proposed high-level explanation is actually valid. We introduce a benchmark of ten complex systems spanning both discrete and continuous state spaces, as well as static and dynamical regimes, each equipped with consensual ground-truth causal explanations and invalid contrastive conditions. Within a unified causal abstraction framework, we systematically evaluate over thirty candidate metrics drawn from observational, functional, information-theoretic, and causal families. Our results show that only the latter reliably discriminates valid from invalid abstractions, and only when incorporating faithfulness testing over unmapped variables. Building on these findings, we introduce the Causal Abstraction Error (CAE), a continuous validity metric with an explicit faithfulness test, which passes all discrimination tests across every system and can converge with as few as 30 sampled interventions. We offer it as a general-purpose metric for the discovery and validation of high-level explanations.
ASPIRE: ロボット工学のためのエージェント/スキル発見
従来のロボット プログラミングは困難です。マルチモーダルな認識を調整し、物理的な接触ダイナミクスを管理し、さまざまな構成と実行エラーを処理する必要があります。 ASPIRE (Agentic Skill Programming through Iterative Robot Exploration) を紹介します。これは、経験を再利用可能なスキル ライブラリに複合化しながら、ポリシーとしてのコード パラダイムでロボット制御プログラムを自律的に作成および改良する継続学習システムです。 ASPIRE は、タスク、シミュレーション、現実世界の設定、および実施形態にわたって持続するスキルを発見します。これは、次の 3 つのコンポーネントを備えたオープンエンド ループで動作します。(1) 閉ループ ロボット実行エンジン。きめの細かいマルチモーダル トレースを公開し、自律的な障害診断、修復合成、および検証を可能にします。 (2) 検証された修正を再利用可能で移転可能な知識に抽出する、継続的に拡張するスキル ライブラリ。 (3) 単一軌道の改良を超えて探索するための多様なタスクシーケンスと制御プログラムを生成する進化的探索。 ASPIRE は、摂動下での LIBERO-Pro 操作で従来の方法を最大 77%、Robosuite の両手ハンドオーバーで 72%、BEHAVIOR-1K の長期的な家事タスクで 32% 上回りました。蓄積されたライブラリにより、目に見えない長期的なタスクに対するゼロショットの一般化も可能になります。LIBERO-Pro Long では、ASPIRE は、テスト時の推論と再試行を使用しているにもかかわらず、以前の方法では 4% であったのに対し、31% の成功率を達成しました。最後に、シミュレーションで発見されたスキルは、シミュレーションからリアルへの移行の初期証拠を提供し、さまざまな実施形態およびロボット API にわたる実際のロボットのプログラミングの労力を大幅に削減します。
原文 (English)
ASPIRE: Agentic /Skills Discovery for Robotics
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introduce ASPIRE (Agentic Skill Programming through Iterative Robot Exploration), a continual learning system that autonomously writes and refines robot control programs in a code-as-policy paradigm while compounding experience into a reusable skill library. ASPIRE discovers skills that persist across tasks, simulation and real-world settings, and embodiments. It operates in an open-ended loop with three components: (1) a closed-loop robot execution engine that exposes fine-grained multimodal traces, enabling autonomous failure diagnosis, repair synthesis, and validation; (2) a continually expanding skill library that distills validated fixes into reusable, transferable knowledge; and (3) evolutionary search that generates diverse task sequences and control programs to explore beyond single-trajectory refinement. ASPIRE surpasses prior methods by up to 77% on LIBERO-Pro manipulation under perturbation, 72% on Robosuite bimanual handover, and 32% on BEHAVIOR-1K long-horizon household tasks. Its accumulated library also enables zero-shot generalization to unseen long-horizon tasks: on LIBERO-Pro Long, ASPIRE achieves 31% success versus 4% for prior methods despite their use of test-time reasoning and retries. Finally, simulation-discovered skills provide initial evidence of sim-to-real transfer, substantially reducing real-robot programming effort across different embodiments and robot APIs.
SEFORA: フィードバック コーパスと LLM フィードバック評価フレームワークを使用した学生のエッセイ
効果的なフィードバックの作成は生徒の学習を促進する最も強力な推進力の 1 つですが、それを大規模に作成するには多大な労力がかかります。 LLM はライティングサポートを拡張するための自然な道筋を提供しますが、2 つのギャップが邪魔をしています。実際の教室で講師が実際にどのようにフィードバックを提供するかを記録した公開コーパスがほとんどないこと、生成されたフィードバックが講師が書く内容と一致しているかどうかを測定する信頼できる方法がないことです。私たちは両方に対応します。 SEFORA は、課題プロンプト、ルーブリック、スコア、大学のさまざまな執筆ジャンルにわたる複数の下書き改訂を備えた、講師のインライン フィードバックを組み合わせたパブリック コーパスであり、564 の下書きと 8,240 の講師の注釈で構成されています。 UniMatch は、オープンエンド生成のための参照ベースの評価フレームワークです。フィードバックをフィードバック単位に分割し、インストラクターが導き出した基準に基づいて意味論的な対応をスコアリングし、最適なマッチングによってそれらを調整して、解釈可能な精度、再現率、および F1 を生成します。複数の LLM にわたる 74 の実験構成全体で、0.4 F1 を超える設定はありませんでした。 UniMatch は、モデルがインストラクターが優先するフィードバックを特定するのに苦労しており、モデルが生成するフィードバックが増えるとパフォーマンスが低下することを明らかにしました。
原文 (English)
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework
Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback in real classrooms, and no reliable method measures whether generated feedback aligns with what an instructor would write. We address both. SEFORA is a public corpus pairing instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across various college writing genres, comprising 564 drafts and 8,240 instructor annotations. UniMatch is a reference-based evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield interpretable precision, recall, and F1. Across 74 experimental configurations spanning multiple LLMs, no setting exceeds 0.4 F1. UniMatch reveals that models struggle to identify the feedback instructors would prioritize, and performance degrades as models generate more.
希少データ連合学習における疎モデル発見のためのエントロピー正規化確率ゲート
Federated Learning (FL) は、データを共有せずに複数のクライアント間でコラボレーションする分散型機械学習 (ML) パラダイムです。 FL は、データの異質性と部分的なクライアントの参加の下で課題を抱えています。スパース モデルの学習は、FL での通信と計算効率の向上に役立ちますが、最適化によって、目に見えないテスト データに一般化できないパラメータ構成が生成される可能性がある、サンプルが少ない高次元領域 (d >> N) では特に困難です。マグニチュードベースの枝刈りではパラメーター空間の不確実性の探索は考慮されていませんが、確率的ゲートと L0 制約を使用した定式化により、トレーニング中に競合する疎な構成からサンプリングすることができます。この研究では、スパース サポートへの早期のコミットメントを防止することで、スパース フェデレーション最適化における不確実性を維持するメカニズムとして、ゲート分布のエントロピー正則化を研究します。データの不均一性、クライアント参加の不均一性、およびスパース性の下でのその影響を調査します。合成ベンチマークと現実世界のベンチマークの実験では、フェデレーテッド反復ハードしきい値処理 (Fed-IHT) と、高密度フェデレーテッド平均化 (FedAvg) トレーニング後の枝刈りよりも、テスト データの統計的パフォーマンスとスパース性回復精度の両方において、一貫した改善が見られます。
原文 (English)
Entropy-Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning
Federated Learning (FL) is a distributed machine learning (ML) paradigm with collaboration among multiple clients without sharing data. FL is challenging under data heterogeneity and partial client participation. Learning sparse models is useful for communication and computational efficiency in FL, but it is especially difficult in the small-sample high-dimensional regime (d >> N) where optimization can yield parameter configurations that fail to generalize to unseen test data. While magnitude-based pruning doesn't account for uncertainty exploration in the parameter space, a formulation with probabilistic gates and an L0 constraint allows sampling from competing sparse configurations during training. In this work, we study entropy regularization of gate distributions as a mechanism to maintain uncertainty in sparse federated optimization by preventing early commitment to sparse support. We examine its impact under data heterogeneity, client participation heterogeneity, and sparsity. Experiments on synthetic and real-world benchmarks show consistent improvements over federated iterative hard thresholding (Fed-IHT) and pruning after dense federated averaging (FedAvg) training, both in statistical performance on test data and in sparsity recovery accuracy.
並列物理世界でのフロンティア大規模言語モデルの物理リテラシーのテスト
現在の大規模言語モデル (LLM) の物理ベンチマークは通常、解答の精度によってスコア化されますが、本物の推論とよく知られた問題パターンの想起を区別することができず、モデルの推論がどこで破綻するかについてはほとんど明らかになりません。帰納、定式化、予測、レビューを通じて、LLM がなじみのない物理フレームワーク内で推論できるかどうかを評価する、監査可能な 4 段階の診断を導入します。この診断は、ロックされた事前登録、ステージ間の新鮮なセッション、デュアル LLM 判定、人間による監査経路を組み合わせたもので、単一方程式反事実世界 ($F=mv$)、歴史的フレームワーク (アリストテレス力学)、および 4 領域反事実世界 (崩壊世界) の 3 つの並行物理世界に適用します。 Claude Opus 4.7、GPT-5.5、および Gemini 3.1 Pro 全体で、3 つのワールドの複合 PASS 率はそれぞれ 6/15、6/15、0/15 です ($F=mv$ およびアリストテレスのコンテンツ $\land$ 構造、構造軸が範囲外である Decay World のみのコンテンツ軸)。最も顕著な経験的パターンは、質的対量的な非対称性です。Decay World では、モデルが変化の間違った方向を予測することはほとんどありませんが、標準的な物理関係に戻って間違った比率を計算することがよくあります。このプロトコルはまた、2 つの方法論の発見も明らかにしています。LLM 判定の信頼性はフレームワークを越えて伝達されないこと、およびステージ 4 の自己レビューはどのフレームワークでも弱く、モデル自身のレビューは、実際にエラーが含まれていた試験の少なくとも 3 分の 2 で以前のエラーが存在しなかったと誤って報告していることです。私たちは完全なプロンプト、応答、評決、監査記録を公開します。
原文 (English)
Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds
Current large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning from recall of familiar problem patterns and reveals little about where a model's reasoning breaks down. We introduce an auditable four-stage diagnostic that evaluates whether an LLM can reason inside an unfamiliar physics framework through induction, formulation, prediction, and review. The diagnostic combines locked pre-registrations, fresh sessions between stages, dual-LLM judging, and a human-audit pathway, and we apply it to three parallel physics worlds: a single-equation counterfactual world ($F=mv$), a historical framework (Aristotelian mechanics), and a four-domain counterfactual world (Decay World). Across Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, the three worlds yield composite PASS rates are 6/15, 6/15, and 0/15 respectively (content $\land$ structural for $F=mv$ and Aristotelian, content axis only for Decay World where the structural axis is out of scope). The most pointed empirical pattern is a qualitative-versus-quantitative asymmetry: in Decay World, models almost never predict the wrong direction of change, but frequently compute the wrong ratio by slipping back to standard-physics relations. The protocol also surfaces two methodology findings: LLM-judge reliability does not transfer across frameworks, and Stage 4 self-review is weak in every framework, with the model's own review wrongly reporting no earlier error in at least two-thirds of the trials that actually contained one. We release the full prompts, responses, verdicts, and audit records.
隠された問題: 視覚言語モデルを使用した計画クリティカルな遮蔽エージェントの特定
自動運転車は、計画に不可欠なエージェントが視界から隠れている可能性がある複雑な環境を安全に移動する必要があります。現在のアプローチでは、すべてのオクルージョンを画一的な保守主義で扱うことが多く、不必要に防御的な運転をもたらしたり、プランナーへの影響を推定せずに隠れたスペースを推測したりすることがあります。この研究は、視覚言語モデル (VLM) が自車両の軌道にとって最も重要な特定の隠れたエージェントを特定し推論できるようにすることで、認識と計画の間の重大なギャップを埋めます。我々は、自我車両の計画に対する影響に基づいて、遮蔽されたエージェントを体系的に特定し、ランク付けするための情報理論的指標であるプランニング KL ダイバージェンス (PKL) を使用する新しいフレームワークを紹介します。この計画を意識したランキングを使用して、エキスパート VLM (GPT-5) を採用して、このタスクに必要な視覚的証拠と推論をキャプチャする豊富で構造化された注釈を生成します。このフレームワークを nuScenes データセットに適用して、影響の大きいシナリオに焦点を当てた新しいベンチマークを作成します。私たちは、幅広い汎用 VLM とドメインに適応した VLM で包括的な実験を実施し、PKL に基づいたデータの微調整により、すべてのモデルにわたって劇的なパフォーマンスの向上がもたらされることを実証しています。特に、この結果は、より小規模で微調整されたモデルが、より大規模なゼロショットモデルよりも大幅にパフォーマンスが優れていること、および PKL に基づいたデータ選択戦略により、ランダム サンプリングと比較してパフォーマンスが約 30\% 向上することを示しています。私たちの研究は、プランニングに不可欠なオクルージョンに焦点を当てて VLM をトレーニングするための最初の体系的なアプローチを提示し、自動運転におけるより意味的に根拠のある効率的なリスク評価を可能にします。
原文 (English)
What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform conservatism, yielding needlessly defensive driving, or they infer hidden spaces without estimating the impact on the planner. This work bridges the critical gap between perception and planning by enabling Vision-Language Models (VLMs) to identify and reason about the specific hidden agents that are most critical to the ego-vehicle's trajectory. We introduce a novel framework that uses Planning KL-divergence (PKL), an information-theoretic metric, to systematically identify and rank occluded agents based on their impact on the ego vehicle's plan. Using this planning-aware ranking, we employ an expert VLM (GPT-5) to generate rich, structured annotations that capture the visual evidence and reasoning required for this task. We apply this framework to the nuScenes dataset to create a new benchmark focused on high-impact scenarios. We conduct comprehensive experiments on a wide range of general-purpose and domain-adapted VLMs, demonstrating that fine-tuning on our PKL-guided data yields dramatic performance improvements across all models. Notably, our results show that smaller, fine-tuned models significantly outperform their much larger zero-shot counterparts, and that our PKL-guided data selection strategy improves performance by approximately 30\% over random sampling. Our work presents the first systematic approach for training VLMs to focus on planning-critical occlusions, enabling more semantically grounded and efficient risk assessment in autonomous driving.
インテント駆動型のネットワーク トポロジ設計のための LLM ベースのフレームワーク
自然言語要件に基づいて展開可能で回復力のあるネットワーク トポロジを設計することは、ネットワーク自動化において依然として困難な問題です。この研究では、階層モデリングと体系的な検証を組み合わせた制約駆動パイプラインを通じて、構造的に有効で制約に準拠したネットワーク トポロジを生成する大規模言語モデル (LLM) の機能を調査します。このフレームワークは、公開データセットとしてリリースされた 4 つの現実的なネットワーク シナリオにわたる独自の LLM とオープンウェイト LLM のマルチモデル比較を通じて評価されます。参照トポロジに対するノードとエッジの F1 スコアを使用して構造の正確性を評価し、サーバーとコンテンツの接続メトリクスを通じて復元力を評価します。さらに、生成されたトポロジにおけるインターフェイスの不一致や方向の不一致など、一般的な障害モードを分析します。全体として、この研究は、LLM がトポロジー合成における構造制約と復元力制約をどのように処理するかを理解するための体系的なベンチマークを提供し、AI 主導のネットワーク設計のための情報に基づいたモデル選択をサポートします。
原文 (English)
An LLM-Based Framework for Intent-Driven Network Topology Design
Designing deployable and resilient network topologies from natural language requirements remains a challenging problem in network automation. This work investigates the ability of Large Language Models (LLMs) to generate structurally valid and constraint-compliant network topologies through a constraint-driven pipeline combining hierarchical modeling and systematic validation. The framework is evaluated via a multimodel comparison of proprietary and open-weight LLMs across four realistic network scenarios released as a public dataset. We assess structural correctness using node and edge F1-scores against reference topologies, and evaluate resilience through server and content connectivity metrics. In addition, we analyze common failure modes, including interface mismatches and directional inconsistencies in generated topologies. Overall, this work provides a systematic benchmark for understanding how LLMs handle structural and resilience constraints in topology synthesis, and supports informed model selection for AI-driven network design.
いつ聞くべきかを学ぶ: 人間の動作予測のためのゲート効果融合
制約のない現実世界のビデオにおける人間の動きの予測は、将来の行動の曖昧さとノイズの多いマルチモーダル観測の存在により、依然として困難です。顔の感情は潜在的に相補的な行動の合図を提供しますが、その実際の有用性と動作予測フレームワーク内のメカニズムの境界は依然としてよく理解されていません。この研究では、実際の感情条件付き予測の有用性と時間的限界を調査する系統的な研究を紹介します。 MediaPipe のボディ ポーズの軌跡と HSEmotion の顔の影響表現を組み合わせた厳密なマルチモーダル パイプラインを確立し、クロスモーダル情報フローを動的に調整するための Gated Affect Transformer (GAT) を導入します。厳密な被験者ごとのプロトコルに基づく広範なマルチホライズン評価を通じて、単純な初期のクロスモーダル連結がポーズのみのベースラインと比較して予測精度を一貫して低下させることを実証します。逆に、私たちが提案するゲート機構は、感情の流れを適応的に制御することにより、クロスモーダル統合を安定化します。重要なことに、シャッフルおよびランダム化された感情入力を使用した制御された反事実実験により、学習されたゲートが、もっともらしい感情信号への応答性を維持しながら、非構造化クロスモーダル ノイズをうまく抑制することが明らかになりました。さらに、我々の経験的結果は、顔の感情特徴は、厳密に短期から中期のウィンドウ(例えば、30フレーム)内で限定された水平線依存の予測手がかりを提供するのに対し、長期的な軌跡は主に本質的な運動学的連続性に支配されたままであることを示しています。私たちの発見は、顔の感情は将来の動きの主要な推進力ではなく、補完的な行動の合図と見なされるべきであるという経験的証拠を提供し、制約のない人間の動きの予測における選択的マルチモーダル融合のための実用的なガイダンスを提供します。
原文 (English)
Learning When to Listen: Gated Affect Fusion for Human Motion Prediction
Human motion forecasting in unconstrained real-world videos remains challenging due to the ambiguity of future behaviors and the presence of noisy multimodal observations. While facial affect potentially provides complementary behavioral cues, its practical utility and mechanistic boundaries within motion forecasting frameworks remain poorly understood. In this work, we present a systematic study investigating the utility and temporal limitations of affect-conditioned forecasting in-the-wild. We establish a rigorous multimodal pipeline combining MediaPipe body pose trajectories with HSEmotion facial affect representations, and introduce the Gated Affect Transformer (GAT) to dynamically regulate cross-modal information flow. Through extensive multi-horizon evaluations under a strict subject-wise protocol, we demonstrate that naive early cross-modal concatenation consistently degrades forecasting accuracy relative to pose-only baselines. Conversely, our proposed gating mechanism stabilizes cross-modal integration by adaptively controlling the affective stream. Crucially, controlled counterfactual experiments using shuffled and randomized affect inputs reveal that the learned gate successfully suppresses unstructured cross-modal noise while remaining responsive to plausible affective signals. Furthermore, our empirical results indicate that facial affect features provide bounded, horizon-dependent predictive cues strictly within short-to-medium windows (e.g., 30 frames), whereas long-term trajectories remain predominantly governed by intrinsic kinematic continuity. Our findings provide empirical evidence that facial affect should be regarded as a complementary behavioral cue rather than a dominant driver of future motion, offering practical guidance for selective multimodal fusion in unconstrained human motion forecasting.
評価フロンティアのマッピング: 11 の評価者とエージェントの条件にわたるバイアスと信頼性のトレードオフの実証的調査
バイアスと信頼性のトレードオフは、LLM 評価システムが (ガンマ、H、CV) 空間に制約され、評価者の結合 (ガンマ)、戦略多様性 (H)、および小サンプル測定の信頼性 (CV(N)) を固定サンプル サイズ N で同時に最適化することができないと推測します。以前の証拠は、単一の研究からの完全なメトリクスを備えた n=5 の条件に基づいています。経験ベースを 11 の条件に拡張し、11 条件すべて (有効な重みベクトルを持つ 9 つ) についてガンマと H を測定し、十分なシード (N >= 5) がある 7 つについて CV(N=5) を測定します。 5 つの条件により、完全な (ガンマ、H、CV) トリプルが提供されます。データはトレードオフを裏付けています。評価者の結合が低い条件 (ガンマ 1.0) では結合が強い (ガンマ > 0.9) 条件では低ノイズ (CV(N=5) < 0.16) が達成されます。相関 r(H, gamma) = -0.989 (n=5、GPT-4o 条件を除く) は、評価者の結合が戦略の多様性を抑制することを裏付けています。 4 つの GPT-4o 条件では、すべてのシードでガンマ = 0.000 および H = 1.000 が示されています。このパターンは、2026 年 6 月の GPT-4o API のバージョン ドリフトによるものであると考えられます。 {γ < 0.2, CV(N=5) < 0.3} の領域を占める条件はありません。すべての条件ごとのメトリクスを、評価者が比較できるように標準化されたベンチマーク データセットとしてリリースします。
原文 (English)
Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions
The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine with valid weight vectors) and CV(N=5) for seven with sufficient seeds (N >= 5). Five conditions provide the complete (gamma, H, CV) triple. The data confirm the trade-off: conditions with low evaluator coupling (gamma 1.0), while conditions with strong coupling (gamma > 0.9) achieve low noise (CV(N=5) < 0.16). The correlation r(H, gamma) = -0.989 (n=5, excluding GPT-4o conditions) confirms that evaluator coupling suppresses strategy diversity. Four GPT-4o conditions show gamma=0.000 and H=1.000 across all seeds -- a pattern we attribute to version drift in the June 2026 GPT-4o API. No condition occupies the region {gamma < 0.2, CV(N=5) < 0.3}. We release all per-condition metrics as a standardized benchmark dataset for evaluator comparison.
RetailSMV: 小売業におけるファウンデーション ビデオ ワールド モデルのエキソセントリックな適応と自己中心的な適応
財団ビデオ拡散モデルは、身体化されたエージェントのための世界シミュレータとしてますます見なされていますが、インターネット規模の汎用ビデオでの事前トレーニングでは、現実世界の展開ドメインとの整合性が不十分なままになっています。私たちは、事前トレーニングされた基礎ビデオ ワールド モデルの小売シーンへのパラメータ効率の高い適応を研究します。同じアクティビティの同期された自己中心的ビデオとエソセントリックなビデオが利用可能な場合、トレーニング データのどの観点から最も強力な適応モデルが生成されるでしょうか?従来の小売ビデオ コーパスの顧客中心のフレーミングではなく、店舗スタッフの視点 (在庫、配置、計量、供給カートの管理、チェックアウト時のスキャン) から同期したエゴ/エクソ キャプチャを備えた 5 つのスーパーマーケットからの 32,105 個のキャプション付き小売クリップのコーパスである RetailSMV (Retail Synchronized Multi-View) を導入し、3 つの一致する低ランク適応 (LoRA) 構成をトレーニングします。同一のハイパーパラメータの下での Cosmos3-Nano (エゴセントリックのみ、エキソセントリックのみ、組み合わせ)。厳密なペア統計プロトコルの下で 7 つの相補的メトリクスで評価された 200 クリップのホールドアウト テスト セットでは、エキソセントリックのみの適応は 7 つの点推定のうち 6 点で複合適応と一致または上回っており、わずか 15,985 個のエキソセントリック クリップでトレーニングしたにもかかわらず (組み合わせでは 32,105 個)、LPIPS、PSNR、DreamSim で大幅に優れています。さらに、対称的な一対の比較により、エゴセントリックのみのトレーニングにエキソセントリックなデータを追加すると効果がある一方で、エゴセントリックのみのトレーニングにエゴセントリックなデータを追加すると害が生じることがわかります。絶対的な適応ギャップは展開時間が最も短い場合に最大となり、適応が最も有益な領域として地平線に近い予測ウィンドウが特定されます。
原文 (English)
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces the strongest adapted model? We introduce RetailSMV (Retail Synchronized Multi-View), a corpus of 32,105 captioned retail clips from five supermarkets with synchronized ego/exo capture from the store-staff perspective (stocking, arranging, weighing, managing supply carts, scanning at checkout), rather than the customer-centric framing of prior retail video corpora, and train three matched Low-Rank Adaptation (LoRA) configurations of Cosmos3-Nano (egocentric-only, exocentric-only, combined) under identical hyperparameters. On a 200-clip held-out test set evaluated with seven complementary metrics under a strict paired statistical protocol, exocentric-only adaptation matches or exceeds combined adaptation on six of seven point estimates and is significantly better on LPIPS, PSNR, and DreamSim, despite training on only 15,985 exocentric clips (versus 32,105 for combined). A symmetric paired comparison further shows that adding exocentric data to egocentric-only training helps while adding egocentric data to exocentric-only training hurts. The absolute adaptation gap is largest at the shortest rollout time, identifying the near-horizon prediction window as the regime in which adaptation is most beneficial.
K-Inverse-RFM: データ破損した数学的タスクのためにニューラル ネットワークとのギャップを埋める修正された RFM
再帰特徴マシン (RFM) は、特徴学習のメカニズムとして平均勾配外積 (AGOP) を利用するカーネル マシンのクラスです。これらは、さまざまな設定にわたってフィードフォワード ニューラル ネットワーク (FNN) の学習ダイナミクスと特徴表現を効果的に複製することが示されています。ただし、同等の特徴学習能力と、取得する特徴の類似性にもかかわらず、RFM は、特定のデータ破損シナリオではニューラル ネットワークよりも大幅に低いパフォーマンスを示します。この研究では、数学的問題におけるこれらの限界を調査します。解決策として、ノイズが多く、複雑に表現され、クラスの不均衡なデータでの学習を促進する、トレーニング ラベルに適用される非常に効果的な変換を導入します。このシンプルかつ強力な調整により、RFM は FNN とのパフォーマンスの差を縮め、場合によっては FNN を上回ることができます。
原文 (English)
K-Inverse-RFM: A Modified RFM that Bridges the Gap to Neural Networks for Data-Corrupted Mathematical Tasks
Recursive Feature Machines (RFMs) are a class of kernel machines that utilize the Average Gradient Outer Product (AGOP) as a mechanism for feature learning. They have been shown to effectively replicate the learning dynamics and feature representations of Feedforward Neural Networks (FNNs) across various settings. However, despite comparable capacity for feature learning and the similarities in the features they acquire, RFMs exhibit significantly lower performance than neural networks in certain data-corrupted scenarios. In this work, we investigate these limitations in mathematical problems. As a solution, we introduce a remarkably effective transformation applied to the training labels which promotes learning in noisy, complexly represented, and class-imbalanced data. This simple yet powerful adjustment enables RFMs to close the performance gap with FNNs and, in some cases, even surpass them.
DiscoLoop: マルチホップ推論のための離散埋め込みと連続隠れ状態のループ
大規模な言語モデルは、中間ステップを思考連鎖 (CoT) として外部化できる場合、多くの推論タスクで優れたパフォーマンスを実現します。ただし、多くの質問では、答えを生成する前に、モデルが単一の前方パス内で複数ステップの推論を内部化する必要があります。私たちは、モデルが単一の順方向パス内で複数のパラメトリック知識を構成する必要がある代表的なタスクである 2 ホップ推論を通じてこの課題を研究します。標準の非リカレント Transformer は、深度ローカルのストレージの問題に悩まされています。つまり、前の層で学習された事実は、2 番目のホップの取得が行われる場所では利用できません。ループ トランスフォーマーは同じメモリを再利用することでこの問題を軽減しますが、それでも一般化が不完全であることがわかりました。残りのボトルネックが代表的なものであることを示します。 2 ホップ推論タスクでは、多くの場合、最初のループによって正しいブリッジ エンティティがほぼ完全にデコード可能になりますが、対応する隠れ状態はブリッジ トークンの埋め込みと十分に一致しないままになります。驚くべきことに、トレーニングを必要としない簡単な再調整介入により、一般化ギャップはほぼ埋められます。この洞察に基づいて、我々は DiscoLoop を提案します。DiscoLoop は、その繰り返しが離散埋め込みチャネルと連続隠れ状態チャネルの両方を運ぶループ アーキテクチャです。 DiscoLoop は、記号および合成言語のマルチホップ推論タスク全体で、大幅に少ないトレーニング ステップでほぼ完璧な精度を実現します。実際の事前トレーニングに適用すると、DiscoLoop はループ変換ベースラインよりも低いトレーニング損失と強力なベンチマーク パフォーマンスを達成し、混合チャネル設計が実用的な言語モデリングに移行することを示唆しています。
原文 (English)
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). However, many questions require the model to internalize the multi-step reasoning within a single forward pass before generating the answer. We study this challenge through two-hop reasoning, a representative task where the model must compose multiple pieces of parametric knowledge within a single forward pass. Standard non-recurrent Transformers suffer from a depth-local storage problem: facts learned in earlier layers are unavailable where second-hop retrieval happens. We found that Looped Transformers mitigate this issue by reusing the same memory, but still generalize imperfectly. We show that the remaining bottleneck is representational. In the two-hop reasoning task, the first loop often makes the correct bridge entity nearly perfectly decodable, yet the corresponding hidden state remains poorly aligned with the bridge token embedding. Surprisingly, an easy training-free realignment intervention nearly closes the generalization gap. Building upon this insight, we propose DiscoLoop, a looping architecture whose recurrence carries both a discrete embedding channel and a continuous hidden-state channel. DiscoLoop achieves near-perfect accuracy with substantially fewer training steps across symbolic and synthetic-language multi-hop reasoning tasks. When applied to real-world pretraining, DiscoLoop attains lower training loss and stronger benchmark performance than looped-transformer baselines, suggesting that the mixed-channel design transfers to practical language modeling.
SoK: モバイル オンデバイス AI システムの攻撃と防御の状況
ローカルに展開された AI モデルと従来のモバイル ソフトウェア コンポーネントを統合するモバイル オンデバイス AI (MoAI) システムは、インテリジェントな機能をエンドユーザー デバイスに直接提供するための重要なパラダイムとして台頭しています。このようなシステムは、推論をリモート クラウド サービスからローカル モバイル環境に移行することで、プライバシー保護、低遅延、オフライン対応の AI 機能を実現しますが、AI モデルのローカル ストレージから生じる新たなセキュリティ リスクをもたらします。この文書は、MoAI システムのセキュリティの柱、攻撃の状況、および防御の状況をカバーする、MoAI セキュリティに関する知識の最初の包括的な体系化を示しています。さらに、現在の攻撃と防御の研究における未解決のギャップを特定し、この新興分野における将来の研究の有望な方向性を指摘します。私たちの研究は、MoAI システムの攻撃と防御の状況を理解するための最初の体系的なフレームワークを確立し、安全な MoAI システムを構築し、この重要な領域で研究を進めるための基盤として機能します。コンパニオン リソースは https://github.com/Jinxhy/Awesome-MoAI-Security で入手できます。
原文 (English)
SoK: Attack and Defense Landscape of Mobile On-device AI Systems
Mobile on-device AI (MoAI) systems that integrate locally deployed AI models with conventional mobile software components are emerging as a key paradigm for delivering intelligent functionality directly on end-user devices. By moving inference from remote cloud services to the local mobile environment, such systems enable privacy-preserving, low-latency, and offline-capable AI functionality, yet introduce new security risks arising from the local storage of AI models. This paper presents the first comprehensive systematization of knowledge on MoAI security, covering security pillars, attack landscape, and defense landscape of MoAI systems. We further identify unresolved gaps in current attack and defense research and point to promising directions for future research in this emerging area. Our work establishes the first systematic framework for understanding the attack and defense landscapes of MoAI systems, serving as a foundation for building secure MoAI systems and advancing research in this critical domain. Companion resources are available at https://github.com/Jinxhy/Awesome-MoAI-Security.
効率的で堅牢な音声合成のための統合ガイダンス フレームワークによるフロー マッチングの強化
フロー マッチング (FM) は、音声生成の強力なパラダイムとして登場しましたが、依然として高い推論遅延と音色漏れによる制約を受けています。これらのボトルネックに対処するために、2 つの相補的な戦略を通じて発電効率と堅牢性を強化する統一ガイダンス フレームワークを提案します。データの面では、異種拡張によるデータ ガイダンスを導入し、モデルが音響残基から言語コンテンツを解きほぐすことを促進します。並行して、軌道修正と新しい固有の誘導目標を相乗させる、強化されたモデル誘導メカニズムを提案します。このアプローチでは、条件付き知識をネットワークの重みに抽出し、推論の軌道を直線化することで、Classifier-Free Guide (CFG) のオーバーヘッドを排除します。実験により、私たちのフレームワークは、最先端のベースラインと比較して話者の類似性を効果的に改善しながら、推論をほぼ 3 倍高速化することが実証されました。
原文 (English)
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage. To address these bottlenecks, we propose a unified guidance framework that enhances generation efficiency and robustness through two complementary strategies. On the data front, we introduce Data-guidance via heterogeneous augmentation, encouraging the model to disentangle linguistic content from acoustic residue. In parallel, we propose an enhanced Model-guidance mechanism that synergizes trajectory rectification with a novel intrinsic guidance objective. This approach distills conditional knowledge into network weights and straightens inference trajectory path, thereby eliminating Classifier-Free Guidance (CFG) overhead. Experiments demonstrate that our framework accelerates inference by nearly three times while effectively improving speaker similarity compared to state-of-the-art baselines.
AI が量子情報と出会うとき: 包括的なレビュー
人工知能 (AI) と量子情報 (QI) は急速に共進化しています。 AI は量子システムの学習、設計、制御、検証のための実用的なツールになりつつあり、QI は AI のための新しい計算モデル、表現構造、学習理論の質問を提供します。この調査では、インターフェースを双方向からレビューします。 QI 向け AI の方向性では、限られた測定からの情報の抽出、量子アルゴリズムのトレーニングと発見、ノイズの多いハードウェアの安定化、実験とプログラミングのワークフローの自動化、センシングとネットワーキングへの学習ベースの手法の拡張という中心的なタスクを中心とした最近の進歩を整理します。 AI の方向性に関する QI では、量子計算と量子にインスピレーションを得た構造が、アルゴリズムの高速化、表現力、訓練可能性、一般化、ニューラル ネットワーク設計、テンソル ネットワーク表現を通じて学習にどのような影響を与えるかを検証します。最後に、再現性、スケーラビリティ、ハードウェアの現実性、共同設計における横断的な課題を特定し、進歩は理論、実験、およびハイブリッド量子、つまり古典システムのより緊密な統合に依存すると主張します。
原文 (English)
When AI meets quantum information: A comprehensive review
Artificial intelligence (AI) and quantum information (QI) are rapidly co-evolving. AI is becoming a practical tool for learning, designing, controlling, and verifying quantum systems, while QI offers new computational models, representational structures, and learning-theoretic questions for AI. This survey reviews the interface from both directions. In the AI for QI direction, we organize recent progress around the central tasks of extracting information from limited measurements, training and discovering quantum algorithms, stabilizing noisy hardware, automating experimental and programming workflows, and extending learning-based methods to sensing and networking. In the QI for AI direction, we examine how quantum computation and quantum-inspired structures affect learning through algorithmic speedups, expressivity, trainability, generalization, neural-network design, and tensor-network representations. We close by identifying cross-cutting challenges in reproducibility, scalability, hardware realism, and co-design, arguing that progress will depend on tighter integration of theory, experiment, and hybrid quantum--classical systems.
MEPA: 専門家の混合による視覚的自己回帰モデリングのためのマルチスケール表現の調整
Visual AutoRegressive Modeling (VAR) は、粗いから細かいまでのマルチスケール自己回帰生成パラダイムを開拓し、画像生成における強力な機能を実証しています。ただし、VAR には依然として、マルチスケール表現の学習における固有の欠陥があります。具体的には、低いスケールは主にグローバル セマンティクスを捕捉し、高いスケールはきめの細かい詳細に焦点を当てます。規模を超えて共有アーキテクチャを採用すると、最適化の競合が発生します。さらに、因果的な自己回帰プロセスにより、初期スケールでの不正確なセマンティクスが伝播し、最終出力が大幅に低下する可能性があります。これらの問題に対処するために、スケールを考慮したトークン ルーティングの専門家混合 (MoE) アーキテクチャを導入し、スケールに適応した専門家の選択を可能にし、それによってスケール間での分離された表現学習を促進します。さらに、外部の自己教師あり機能を組み込むことで、初期スケールでのセマンティック モデリングを強化します。単純なアライメントとは異なり、VAR パラダイムに合わせた残差特徴集約スキームを分析および設計します。広範な実験により、私たちの方法がトレーニング効率と生成品質の両方を大幅に向上させることが示されました。 ImageNet 256*256 ベンチマークでは、私たちのモデルは、高密度ベースラインと比較して優れた FID を達成しながら、デフォルトのトレーニング エポックの半分のみを必要とし、パラメーター バジェットを小さくし、トレーニング コストのわずかな増加を抑えています。さらに、トレーニング エポックが大きくなると、パフォーマンスの差はさらに広がります。
原文 (English)
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However, VAR still suffers from inherent deficiencies in multi-scale representation learning. Specifically, lower scales primarily capture global semantics, while higher scales focus on fine-grained details. Employing a shared architecture across scales induces optimization conflicts. Moreover, due to the causal autoregressive process, inaccurate semantics at early scales can propagate and significantly degrade the final output. To address these issues, we introduce a scale-aware token-routed Mixture of Experts (MoE) architecture, allowing scale-adaptive expert selection, thereby facilitating decoupled representation learning across scales. In addition, we enhance semantic modeling at early scales by incorporating external self-supervised features. Unlike naive alignment, we analyse and design a residual feature aggregation scheme tailored to the VAR paradigm. Extensive experiments show that our method significantly improves both training efficiency and generation quality. On the ImageNet 256*256 benchmark, our model achieves a superior FID compared to the dense baseline while requiring only half of the default training epochs and a smaller parameter budget, with a merely marginal increase in training cost. Moreover, the performance gap further widens with larger training epochs.
合成の学習: ゼロショット合成画像取得のためのプロキシ タスク設計の再考
合成画像検索 (CIR) は、参照画像とテキストの変更からターゲット画像を取得します。教師あり CIR はコストのかかるトリプレットに依存しますが、ゼロショット CIR (ZS-CIR) は画像とテキストのペアでトレーニングされたプロキシ タスクを通じてこの依存性を軽減します。ただし、既存のプロキシ タスクは主に、フリーズ テキスト エンコーダへの擬似ワード挿入や線形特徴演算などの事前定義された構成メカニズムに対応するために、視覚的およびテキスト表現を強化します。その結果、合成関数自体は学習されないままとなり、多様で細かい意味論的変更を表現するモデルの能力が制限されます。これに対処するために、我々は FoCo を提案します。FoCo は、変更に関連するビジュアル コンテンツに焦点を当て、次にターゲット セマンティクスを完成するという 2 つの調整された段階として構成をモデル化します。これらは 2 つのプロキシ タスクを通じて実現されます。1 つはローカライズされたテキスト セマンティクスに基づいてビジュアル コンテンツを選択的に収集するテキスト アンカー付きビジュアル集約で、もう 1 つはコンテキスト条件付きセマンティック補完で、これらの集約されたビジュアルを残りのシーン コンテキストとともに一貫した構成表現に変換します。タスクは、インスタンス間の対照的な目標を使用して共同でトレーニングされ、セマンティックな多様性を促進し、ショートカット作成戦略を抑制します。 4 つの ZS-CIR ベンチマークに関する広範な実験により、FoCo の最先端のパフォーマンスと改善された一般化が示されました。
原文 (English)
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) alleviates this reliance through proxy tasks trained on image-text pairs. However, existing proxy tasks primarily enhance visual and textual representations to accommodate a predefined composition mechanism such as pseudo-word injection into a frozen text encoder or linear feature arithmetic. As a result, the composition function itself remains unlearned, limiting the model's ability to express diverse and fine-grained semantic modifications. To address this, we propose FoCo, which models composition as two coordinated stages: focusing on modification-relevant visual content, and then completing the target semantics. We realize these through two proxy tasks: text-anchored visual aggregation to selectively gather visual content guided by localized textual semantics, and context-conditioned semantic completion to transform these aggregated visuals with the remaining scene context into a coherent composed representation. The tasks are trained jointly with a cross-instance contrastive objective, encouraging semantic diversity and discouraging shortcut composition strategies. Extensive experiments on four ZS-CIR benchmarks show FoCo's state-of-the-art performance and improved generalization.
MalariAI: 高密度マラリア血液塗抹標本における普遍的な細胞セグメンテーションと説明可能な病期分類のためのラベル耐性のある分離フレームワーク
血液塗抹標本顕微鏡検査による自動マラリア診断は、世界規模の保健 AI における重要な課題です。リソースが限られた環境では、専門の顕微鏡医の不足が依然としてタイムリーで正確な診断の主なボトルネックとなっています。 3 つの複合的な障害モードにより、既存の深層学習システムの信頼できる臨床展開が妨げられます。まず、エンドツーエンドの検出器は、トレーニング中に注釈のないセルをバックグラウンドとして扱い、実際のセルの回復を反映するのではなく、注釈の完全性に強く影響される再現率の数値を生成します。第 2 に、非最大抑制は、感染数が最も重要な密度の高いスミア領域での有効な検出を抑制する傾向があります。第三に、マラリア画像分類タスクには Grad-CAM などの画像レベルの説明可能手法が適用されているにもかかわらず、既存のスライド全体検出パイプラインには、臨床監査のためのセルごとの空間的証拠が不足しています。統合パイプラインで 3 つの障害モードすべてに対処する 2 段階の分離フレームワークである MalariAI を紹介します。ステージ 1 では、アノテーションに依存しない距離変換ガイド付き流域アルゴリズムを適用して、1600x1200 の完全な血液塗抹標本画像内のすべての細胞を分離し、グラウンド トゥルースの入力なしで 120 枚の画像 NIH BBBC041 テスト セット全体の重心位置特定により、グラウンド トゥルースの細胞の 75.95% を回復します。ステージ 2 では、64x64 作物で焦点損失 (ガンマ = 2.0、クラスごとの逆周波数重み) を使用して EfficientNet-B0 を微調整し、まれなシゾントおよび生殖母細胞のステージでは 98.36% の全体的な分類精度と 87.5% および 75.0% のクラスごとの精度を達成しました。一方、同じクラスでの R-CNN ベースラインの高速化。検出された細胞ごとに生成された Grad-CAM++ ヒートマップは、臨床監査のためのインスタンス レベルの空間的証拠を提供し、顕微鏡医が分類パフォーマンスを犠牲にすることなく、個々の寄生虫レベルでモデルの予測を検証できるようにします。
原文 (English)
MalariAI: A Label-Resilient Decoupled Framework for Universal Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
Automated malaria diagnosis from blood smear microscopy is a critical challenge in global health AI; in resource-limited settings, the scarcity of expert microscopists remains the primary bottleneck to timely and accurate diagnosis. Three compounding failure modes prevent reliable clinical deployment of existing deep learning systems. First, end-to-end detectors treat unannotated cells as background during training, producing recall figures that are strongly influenced by annotation completeness rather than reflecting true cell recovery. Second, Non-Maximum Suppression tends to suppress valid detections in dense smear regions where infection counts matter most. Third, existing whole-slide detection pipelines lack per-cell spatial evidence for clinical audit, despite image-level explainability methods such as Grad-CAM having been applied to malaria image classification tasks. We present MalariAI, a two-stage decoupled framework that addresses all three failure modes in a unified pipeline. Stage 1 applies an annotation-agnostic distance-transform guided watershed algorithm to isolate every cell in a full 1600x1200 blood smear image, recovering 75.95% of ground-truth cells by centroid localisation across the 120-image NIH BBBC041 test set without any ground-truth input. Stage 2 fine-tunes EfficientNet-B0 with Focal Loss (gamma = 2.0, per-class inverse-frequency weights) on 64x64 crops, achieving 98.36% overall classification accuracy with 87.5% and 75.0% per-class accuracy on the rare schizont and gametocyte stages, compared to only 24.57% and 25.95% AP for a Faster R-CNN baseline on the same classes. Grad-CAM++ heatmaps generated per detected cell provide instance-level spatial evidence for clinical audit, enabling microscopists to verify model predictions at the individual parasite level without sacrificing classification performance.
データ効率の高い教師なし RL による一般化可能なスキル ポリシーの学習
教師なし強化学習 (URL) は、外部報酬なしでスケーラブルでスキル条件付きのポリシーを事前トレーニングすることを目的としており、下流の制御タスクの基盤として機能します。最近の進歩にも関わらず、現在のオフポリシー URL 手法は、(1) 非定常スキル セマンティクスと (2) 脆弱な一般化という 2 つの重大な見落とされているボトルネックによって制限されていると私たちは主張します。これらの課題に対処するために、堅牢な教師なし強化学習のための統一フレームワークである GenDa (Generalizable Data-efficient Agent) を提案します。まず、非定常性を軽減し、事前トレーニングのデータ効率を大幅に向上させるスキルの再ラベル付けメカニズムを導入します。 2 番目に、補完情報ボトルネック (CIB) を提案し、学習済みスキル ポリシーが自己中心的な機能に焦点を当て、下流タスクの分散シフトに対して堅牢になるように奨励します。さまざまな実験を通じて、GenDa が優れた一般化性とデータ効率により URL のスケーラビリティを大幅に強化することを実証しました。私たちのコードとビデオは https://ihatebroccoli.github.io/official-GenDa で入手できます。
原文 (English)
Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL
Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, overlooked bottlenecks: (1) non-stationary skill semantics and (2) brittle generalization. To address these challenges, we propose GenDa (Generalizable Data-efficient Agent), a unified framework for robust unsupervised reinforcement learning. First, we introduce a skill relabeling mechanism to mitigate non-stationarity and significantly improve data efficiency for pre-training. Second, we propose a Complementary Information Bottleneck (CIB), encouraging the learned skill policy to focus on ego-centric features and become robust to distribution shifts for downstream tasks. Through various experiments, we demonstrate that GenDa significantly enhances the scalability of URL with superior generalizability and data efficiency. Our code and videos are available at https://ihatebroccoli.github.io/official-GenDa.
NeuroCogMap は大規模言語モデルの認知組織を明らかにする
複雑な認知機能が人工システム内でどのように組織化されているかを理解することは、大規模言語モデル (LLM) を解釈し、それらを生物学的認知に関連付ける上で中心となります。しかし、LLM は広範な認知的行動を示しますが、その内部表現が行動、失敗、および人間の認知とのつながりを説明する再現可能な機能システムを形成しているかどうかは不明のままです。ここでは、LLM の内部特徴を機能的な区画に編成し、それらを解釈可能な機能、認知能力、および認知階層にリンクする、認知神経科学にインスピレーションを得たフレームワークである NeuroCogMap を紹介します。これらのパーセルは、モデル間で部分的に保存され、モデルの出力に機能的にリンクされた、安定した意味的に一貫した組織を形成します。この組織内では、幻覚、偏見、拒否の失敗、お調子者などの主要な LLM の失敗は、表象システムと行動制御システムの明確な混乱に対応しており、メカニズムに基づく検出と対象を絞った介入のための内部兆候が得られます。 NeuroCogMap は、モデルの動作を超えて、自然主義的な言語理解中の人間の皮質反応の予測を改善し、高次連合皮質での最も強い対応関係が得られます。認知レベルでは、その内部署名により、人間の意思決定の古典的なモデルの改良を導く潜在的な戦略が明らかになります。これらの発見を総合すると、人工システムにおける機能組織をマッピングし、この組織を人間の皮質機能および認知行動に関連付けるためのシステムレベルのフレームワークとして NeuroCogMap が確立されます。
原文 (English)
NeuroCogMap Reveals Cognitive Organization of Large Language Models
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.
ホログラフィック量子トランスフォーマー: 生成的注意を介してイライラしたシステムを解決するためのジェネラリスト神経記号アーキテクチャ
2 次元のフラストレート量子物質のシミュレーションは、符号問題と指数関数的なヒルベルト空間の複雑さのため、大きな課題です。この研究では、グローバルな自己注意を活用して非局所的なもつれパターンを解決する、物理学にインスピレーションを得た生成アーキテクチャであるホログラフィック量子変換器 (HQT) を紹介します。正方格子$J_1-J_2$ハイゼンベルグモデルでHQTを検証します。量子臨界点 ($J_2=0.5$) の非常にフラストレーションの多い $8 \times 8$ 格子上では、HQT はサイトあたりの基底状態エネルギー ($E/N$) $\mathbf{-0.5001(1)}$ に達し、予想される有限サイズのスケーリング傾向と一致します。数値的な精度を超えて、HQT は本質的な物理的認識を示し、解釈可能なアテンション マップを通じて根底にある $J_2$ 相互作用ジオメトリを自律的に回復します。私たちの中心的な貢献は、高速アライメントを備えたゼロショット サイズ外挿プロトコルである「ホログラフィック転送」です。$8 \times 8$ システムでトレーニングされたモデルは、連続的な位置埋め込み補間とヘッドの再初期化を介して、より大きな $10 \times 10$ 格子に直接投影され、高忠実度の初期化と迅速な収束を実現します。このゼロショット プロトコルは、$E/N = のエネルギーを生成します。 \mathbf{-0.49782(3)}$ は、統計的に最先端の変分と一致しており、ターゲット格子での最初からのトレーニングを必要とせず、転送可能な量子シミュレーションのスケーラブルなパラダイムとして生成的注意を確立します。
原文 (English)
Holographic Quantum Transformer: A Generalist Neuro-Symbolic Architecture for Solving Frustrated Systems via Generative Attention
Simulating two-dimensional frustrated quantum matter is a grand challenge due to the sign problem and exponential Hilbert space complexity. In this work, we introduce the Holographic Quantum Transformer (HQT), a physics-inspired generative architecture that leverages global self-attention to resolve non-local entanglement patterns. We validate HQT on the square lattice $J_1-J_2$ Heisenberg model. On the heavily frustrated $8 \times 8$ lattice at the quantum critical point ($J_2=0.5$), HQT reaches a ground-state energy per site ($E/N$) of $\mathbf{-0.5001(1)}$, consistent with the expected finite-size scaling trend. Beyond numerical accuracy, HQT exhibits intrinsic physical awareness, autonomously recovering the underlying $J_2$ interaction geometry through interpretable attention maps. Our central contribution is ``Holographic Transfer", a zero-shot size-extrapolation protocol with rapid alignment: a model trained on $8 \times 8$ systems is directly projected onto larger $10 \times 10$ lattices via continuous positional-embedding interpolation and head re-initialization, achieving high-fidelity initialization and rapid convergence. This zero-shot protocol yields an energy of $E/N = \mathbf{-0.49782(3)}$, statistically consistent with the variational state of the art while requiring no from-scratch training on the target lattice. Our results establish generative attention as a scalable paradigm for transferable quantum simulation.
テキストから画像への拡散モデルの安全性調整における高い有用性の幻想
テキストから画像への (T2I) 拡散モデルの安全性調整は、無害なプロンプトでの有用性を維持しながら、有害な生成を抑制することを目的としています。最近の手法は、高い安全性と高い実用性を提供するように見えることがよくありますが、この結論は主に、粒度の細かいセマンティックの正しさに鈍感な大まかなグローバル ユーティリティ メトリック (FID、CLIPScore など) に基づいており、実用性が高いという錯覚を生み出します。有用性が構造化評価で測定される場合、この幻想は崩れることを示します。TIFA (質問応答によるテキストから画像への忠実度評価) では、安全性を調整したモデルは、オブジェクト数、属性、および関係性の失敗を含む、意味論的忠実度の大幅な低下に見舞われます。このギャップの原因を診断するために、テキスト エンコーダー プロンプトの埋め込み空間を分析し、構造化されたユーティリティの損失と強く相関する、プロンプト間の類似性構造の歪みと結合した埋め込みスプレッドの縮小であるセマンティック崩壊を明らかにします。この洞察に基づいて、適応中に埋め込みの広がりとプロンプト間の関係構造を明示的に保存する安全調整目標である StructureAware Geometric Regularization (SAGE) を提案します。私たちの方法は、強力な安全性能と競争力のある粗粒ユーティリティスコアを維持しながら、構造化されたユーティリティ(TIFA +5.0%)を回復します。ソース コードとトレーニング済みモデルは https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/ で入手できます。
原文 (English)
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent methods often appear to deliver high safety with high utility, but this conclusion rests largely on coarse global utility metrics (e.g., FID, CLIPScore) that are insensitive to fine-grained semantic correctness, creating an illusion of high utility. We show that when utility is measured with structured evaluation, this illusion breaks: on TIFA (Text-to-Image Faithfulness evaluation with Question Answering), safety-aligned models suffer substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships. To diagnose the source of this gap, we analyze the text-encoder prompt embedding space and uncover semantic collapse, a contraction of embedding spread coupled with distortion of inter-prompt similarity structure, which strongly correlates with structured utility loss. Guided by this insight, we propose StructureAware Geometric Regularization (SAGE), a safety alignment objective that explicitly preserves embedding spread and inter-prompt relational structure during adaptation. Our method restores structured utility (TIFA +5.0% over prior state-of-the-art) while maintaining strong safety performance and competitive coarse-grained utility scores. Our source code and trained models are available at https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/.
EO-VGGT: 衛星マルチビュー再構成のための軌道光線条件付き 3D 基礎モデル
衛星群の時代では、多視点の光学衛星画像は地球観測 (EO) と高品質デジタル表面モデル (DSM) の再構成にとって極めて重要です。フィードフォワード 3D 基礎モデルはコンピューター ビジョンを変革しましたが、衛星リモート センシングへの導入は、暗黙的な遠近法の仮定と明示的な軌道手押し箒の幾何学構造の間の構造的な不一致によって本質的に制約を受けます。この幾何学的不一致は、ビューセットの顕著な異質性によってさらに悪化します。我々は、明示的な物理幾何学埋め込みを介して、フリーズされたパースペクティブ駆動モデルを軌道観測に適応させるフレームワークである EO-VGGT を紹介します。まず、幾何学相関制約選択 (GCCS) 戦略は、入力シーケンスを最適化するために幾何学的多様性と放射測定の一貫性のバランスをとることにより、最適以下の観測を取り除きます。次に、Sensor-Ray Encoder (SRE) は、Rational Function Model (RFM) から導出されたピクセルレベルの押し箒の視線を高次元の空間幾何学トークンにパラメータ化し、中心投影と軌道運動学の数学的不一致を調整します。 3 番目に、軽量の Ray-Pointing-Aware Adaptor (RPAA) は、ゲートされた残差ブロックを使用して、これらのトークンをフリーズされたトランスフォーマー バックボーンに直接統合します。私たちの調査結果は、明示的な物理ジオメトリと最適化されたビュー選択を統合することが、堅牢なフィードフォワード衛星 3D 再構成には不可欠であることを強調しています。
原文 (English)
EO-VGGT: Orbital Ray-Conditioned 3D Foundation Models for Satellite Multi-View Reconstruction
In the era of satellite constellations, multi-view optical satellite imagery is pivotal for Earth Observation (EO) and high-quality Digital Surface Model (DSM) reconstruction. Although feed-forward 3D foundation models have transformed computer vision, their deployment in satellite remote sensing is inherently constrained by the structural discrepancy between implicit perspective assumptions and explicit orbital pushbroom geometry. This geometric incongruity is further compounded by pronounced view-set heterogeneity. We present EO-VGGT, a framework that adapts a frozen perspective-driven model to orbital observations via explicit physical geometry embedding.First, the Geometry-Correlation Constrained Selection (GCCS) strategy prunes sub-optimal observations by balancing geometric diversity and radiometric consistency to optimize the input sequence. Second, a Sensor-Ray Encoder (SRE) parameterizes pixel-level pushbroom lines of sight derived from the Rational Function Model (RFM) into high-dimensional space-geometric tokens, reconciling the mathematical discrepancy between central projection and orbital kinematics. Third, a lightweight Ray-Pointing-Aware Adapter (RPAA) employs gated residual blocks to integrate these tokens directly into the frozen transformer backbone. Our findings underscore that integrating explicit physical geometry with optimized view selection is essential for robust feed-forward satellite 3D reconstruction.
時相論理仕様による歩行認識型四足歩行の学習
四足歩行の強化学習 (RL) は一般に、固定された手作りのマルコフ報酬関数に依存します。この関数は、学習されたポリシーの解釈可能性を制限し、歩行動作の明示的な制御を欠きます。信号時間論理 (STL) で表現されたパラメーター化された制約を使用して、個別の歩行を指定するフレームワークを導入します。これらには、安全限界、歩行同期制約、コマンド追跡、および作動限界が含まれます。これらの仕様に基づいて、望ましい行動をコード化する高密度で継続的な報酬ランドスケープを学習エージェントに提供する報酬形成メカニズムを開発します。 3 つの速度レジーム (速歩、速歩、バウンド) のパラメトリック STL テンプレートを定義し、参照ロールアウトからパラメーターを調整し、ロールアウト全体にわたる STL 堅牢性の滑らかな近似を使用して報酬を計算します。生成された報酬は、Proximal Policy Optimization (PPO) と互換性のある成形された勾配を提供するために使用できます。 Google の Barkour 四足歩行ロボットのアプローチを MuJoCo XLA (MJX) でインスタンス化します。シミュレーター内で並列化を使用してトレーニング速度を向上させ、ドメインのランダム化を使用して学習されたポリシーを強化します。手作りの報酬のベースラインと比較して、STL 形状の報酬はより厳密な速度追跡とより安定したトレーニングをもたらすことを示します。ビデオはプロジェクト Web サイト https://stl-locomotion.github.io/ でご覧いただけます。
原文 (English)
Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications
Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit both interpretability of learned policies and lack explicit control over gait behaviors. We introduce a framework where distinct gaits are specified using parameterized constraints expressed in Signal Temporal Logic (STL). These include safety bounds, gait synchronization constraints, command tracking, and actuation bounds. From these specifications, we develop a reward shaping mechanism that provides learning agents a dense, continuous reward landscape that encodes desired behavior. We define parametric STL templates for three speed regimes (walking-trot, trot, bound), calibrate their parameters from reference rollouts, and compute rewards from using smooth approximations of STL robustness over the rollouts. The generated rewards can be used to provide shaped gradients compatible with Proximal Policy Optimization (PPO). We instantiate the approach on Google's Barkour quadruped robot in MuJoCo XLA (MJX). We use parallelization within the simulator to improve training speeds and use domain randomization to robustify learned policies. We show that compared to a baseline of hand-crafted rewards, the STL-shaped rewards yield tighter velocity tracking and more stable training. Videos can be found on our project website: https://stl-locomotion.github.io/.
時空凸集合グラフ上の検索ベースの時空間およびマルチロボット運動計画
特にマルチロボット設定における時空間動作計画では、ロボットが時間の経過とともに変化する衝突のない領域を推論する必要がありますが、実行可能な領域が一時的で幾何学的に制約されている連続空間では困難です。我々は、時空凸集合(ST-GCS)のグラフに基づくアルゴリズムフレームワークを提示する。衝突のない領域は時空の凸集合として表現され、軌道は選択された集合内の連続運動とともにグラフ上の経路に対応する。 ST-GCS での時間最適計画を、パスインデックス付き状態に対するグラフ検索問題として定式化し、許容可能なヒューリスティックと支配性チェックに基づいて、連続軌道最適化によって部分パスを評価する最良優先探索ソルバーを開発します。さらに、時空間における軌道の占有を予約し、動的な障害物と複数ロボットの相互作用の統合処理を可能にする正確な凸分解 (ECD) スキームを提示します。マルチロボットの動作計画では、ST-GCS 計画と ECD を優先順位付けされた計画手法に統合し、効率を向上させるためにウィンドウ調整スキームを導入します。単一ロボットおよび複数ロボットの問題に関する広範な実験により、特に狭くて一時的な実行可能領域がある環境において、高いソリューション品質を維持しながら、さまざまなプランナーよりも大幅な速度向上が実証されました。さらに、大規模なデモンストレーションでは、提案されたマルチロボット モーション プランナーが、最大 100 ドルのロボットを含むインスタンスをわずか数分以内に解決できることが示されています。プロジェクトのホームページ: https://sites.google.com/view/stgcs
原文 (English)
Search-Based Spatiotemporal and Multi-Robot Motion Planning on Graphs of Space-Time Convex Sets
Spatiotemporal motion planning, especially in multi-robot settings, requires robots to reason about collision-free regions that change over time, which is challenging in continuous spaces when feasible regions are transient and geometrically constrained. We present an algorithmic framework based on graphs of space-time convex sets (ST-GCSs), where collision-free regions are represented as convex sets in space-time and trajectories correspond to paths on the graph together with continuous motions within the selected sets. We formulate time-optimal planning on ST-GCSs as a graph-search problem over path-indexed states and develop a best-first search solver that evaluates partial paths via continuous trajectory optimization, guided by admissible heuristics and dominance checks. We further present an Exact Convex Decomposition (ECD) scheme to reserve trajectory occupancies in space-time, enabling unified handling of dynamic obstacles and multi-robot interactions. For multi-robot motion planning, we integrate ST-GCS planning and ECD into prioritized planning methods and introduce a windowed coordination scheme to improve efficiency. Extensive experiments on single-robot and multi-robot problems demonstrate substantial speedups over various planners while maintaining high solution quality, particularly in environments with narrow and transient feasible regions. Large-scale demonstrations further show that the proposed multi-robot motion planner can solve instances with up to $100$ robots within only a few minutes. Project homepage: https://sites.google.com/view/stgcs
VideoSearch-R1: ソフト クエリ改良による反復的なビデオの取得と推論
ビデオ コーパスの規模とタスクの複雑さの両方が拡大し続けるにつれて、大規模なコーパスから関連するビデオを取得し (ビデオ間推論)、その後、時間的グラウンディングなど、取得したコンテンツ内できめの細かいクエリ条件付きタスク (ビデオ内推論) を実行するアプローチの需要が高まっています。ただし、既存のアプローチは通常、取得を前処理ステップとして扱うため、最初の取得が失敗すると、検索を絞り込むメカニズムがなく、その後のきめの細かいビデオ内推論の失敗につながります。さらに、最近のエージェント フレームワークはビデオの理解が進んでいますが、通常、クエリに関連するビデオがすでに提供されていると想定しており、ビデオ内の推論タスクのみに焦点を当てています。これらの制限に対処するために、ビデオ検索エンジンとのマルチターン対話による反復的なビデオ検索と推論のためのエージェント フレームワークである VideoSearch-R1 を提案します。具体的には、ソフト クエリ洗練 (SQR) を導入して、離散テキスト空間でクエリを書き換えるのではなく、連続潜在空間で検索クエリ トークンを洗練し、より効率的かつきめ細かい調整を可能にします。 SQR とその推論プロセスは、グループ相対ポリシー最適化 (GRPO) を使用してトレーニングされ、検索タスクと下流タスクから得られるタスクレベルの報酬シグナルによって導かれます。これに基づいて、VideoSearch-R1 は、ビデオ コーパス モーメント取得 (VCMR) 上の 3 つのデータセットにわたって最先端のパフォーマンスを実現し、大規模なコーパスからビデオを繰り返し取得し、検索クエリを洗練し、取得したコンテンツ内で正確なクエリ条件付きの時間的グラウンディングを実行します。私たちの分析は、SQR が元のクエリを効果的に絞り込み、明示的なテキスト レベルのクエリ絞り込みよりも生成されるトークンが大幅に少なくて済むことを示しています。コードとモデルのチェックポイントは、mlvlab.github.io/VideoSearch-R1 で公開されています。
原文 (English)
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video reasoning) and subsequently perform fine-grained, query-conditioned tasks (intra-video reasoning) within the retrieved content, such as temporal grounding. However, existing approaches typically treat retrieval as a preprocessing step, and consequently, when the initial retrieval fails, there is no mechanism to refine the search, leading to the failure of subsequent fine-grained intra-video reasoning. Moreover, while recent agentic frameworks have advanced video understanding, they typically assume that the query-relevant video is already given, focusing exclusively on intra-video reasoning tasks. To address these limitations, we propose VideoSearch-R1, an agentic framework for iterative video retrieval and reasoning through multi-turn interaction with a video search engine. Specifically, we introduce Soft Query Refinement (SQR) to refine search query tokens in a continuous latent space rather than rewriting queries in the discrete text space, enabling more efficient and fine-grained adjustments. SQR and its reasoning process are trained using Group Relative Policy Optimization (GRPO), guided by task-level reward signals derived from retrieval and downstream tasks. Building upon this, VideoSearch-R1 achieves state-of-the-art performance across three datasets on Video Corpus Moment Retrieval (VCMR), iteratively retrieving videos from large-scale corpora, refining search queries, and performing precise query-conditioned temporal grounding within the retrieved content. Our analyses show that SQR effectively refines the original query, requiring significantly fewer generated tokens than explicit text-level query refinement. Code and model checkpoints are publicly available at mlvlab.github.io/VideoSearch-R1.
大規模な 2 タワー検索のための LLM ベースのクラスタリングによるリアルタイムのハード ネガティブ サンプリング
2 タワー モデルは、大規模なレコメンデーション システム、特に検索段階で広く使用されています。 2 タワー モデルをトレーニングするための業界標準には、通常、バッチ内および/またはバッチ外のネガティブ サンプリングが含まれます。ただし、これらの方法では、モデルがすぐに学習できる簡単な否定的な結果が生成されることが多く、モデルに十分な挑戦を与えることができません。この問題に対処するために、大規模言語モデル (LLM) を利用してモデルのトレーニング中に同じクラスターからハード ネガティブを生成する、新しい自己教師ありハード ネガティブ サンプリング手法が提案されています。 LLM を利用してメディア表現を学習することにより、提案されたアプローチは、生成されるネガがより挑戦的で有益なものになることを保証します。このリアルタイム サンプリング フレームワークは、運用モデルにシームレスに統合できるように設計されており、最小限の計算複雑さで数十億のトレーニング データ ポイントを処理できます。公開データセットでの実験と大規模オンライン システムへの展開により、提案されたネガティブ サンプリング手法が業界で広く使用されている手法よりも優れていることが実証されました。さらに、産業用途での分析により、このサンプリング方法が推奨事項に固有のフィードバック ループを断ち切り、人気の偏りを大幅に軽減できることが明らかになりました。
原文 (English)
Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval
The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards for training two-tower models typically involve in-batch and/or out-of-batch negative sampling. However, these methods often produce easy negatives that models can quickly learn, failing to sufficiently challenge the model. To address this issue, a novel self-supervised hard negative sampling technique is proposed that leverages a large language model (LLM) to generate hard negatives from the same cluster during model training. By utilizing the LLM to learn media representations, the proposed approach ensures that the generated negatives are more challenging and informative. This real-time sampling framework is designed for seamless integration into production models, capable of handling billions of training data points with minimal computational complexity. Experiments on public datasets, along with deployment to a large-scale online system, demonstrate that the proposed negative sampling technique outperforms widely used industry methods. Furthermore, analysis in industrial applications reveals that this sampling method can help break inherent feedback loops in recommendations and significantly reduce popularity bias.
アクター - クリティカル強化学習におけるクリティカルの複雑さの測定、測定、制御
俳優と批評家の手法は学習した批評家に依存しますが、批評家の質は多くの場合、リターン、時間差エラー、または価値損失を通じて間接的にのみ評価されます。批評家の複雑さは、アクター - 批評家の強化学習の追加の診断および介入の次元として導入されます。この分析では、スペクトル実効ランク エントロピー (批評家重み行列の特異値分布のランク状の要約) を使用して、批評家モデルの複雑さを評価します。 TD3 および PPO 実験全体で、批評家の複雑性がリターンおよびモンテカルロ値推定バイアスとともに追跡されます。結果は、批評家の複雑さがトレーニング全体を通じて測定可能であり、トレーニング動作と体系的に関連付けられていることを示していると同時に、その関係がアルゴリズム、タスク、ハイパーパラメーター間で不均一であることも明らかにしています。次に、スペクトルエントロピーペナルティをクリティカル損失に追加することによって、直接的な複雑さ制御介入が評価されます。この介入により、ターゲットのスペクトル量が確実に変更され、批評家の複雑さを観察するだけでなく制御できることが実証されます。全体的な複雑さの制御の結果は異なるため、リターン効果は、一般的なパフォーマンスの主張としてではなく、タスクに依存する証拠として扱われます。
原文 (English)
Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning
Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference error, or value loss. Critic complexity is introduced as an additional diagnostic and intervention dimension for actor-critic reinforcement learning. The analysis uses spectral effective-rank entropy, a rank-like summary of the singular-value distributions of critic weight matrices, to assess critic model complexity. Across TD3 and PPO experiments, critic complexity is tracked together with return and Monte Carlo value-estimation bias. The results show that critic complexity is measurable throughout training and is systematically associated with training behavior, while also making clear that the relationship is heterogeneous across algorithms, tasks, and hyperparameters. A direct complexity-control intervention is then evaluated by adding a spectral-entropy penalty to the critic loss. This intervention reliably changes the targeted spectral quantity, demonstrating that critic complexity can be controlled rather than only observed. Return effects are treated as task-dependent evidence rather than as a general performance claim, because overall complexity-control results vary.
時空間ダイナミクス予測のためのマルチ解像度有限ボリュームからインスピレーションを得た深層学習フレームワーク
物理プロセスにおける複雑な時空間ダイナミクスを予測するには、多くの場合、計算コストのかかる数値手法やデータ駆動型のニューラル ネットワークが必要になりますが、これらには、高いトレーニング コスト、エラーの蓄積、目に見えないパラメータに対する一般化可能性の制限という問題があります。これらの課題に対処する効果的なアプローチは、物理情報に基づく深層学習 (PiDL) として知られる、ニューラル ネットワークのトレーニングに物理事前分布を活用することです。この研究では、グローバル スケールでの有限ボリュームの保守的な特性とローカル スケールでのディープ ラーニングの表現力を活用するように設計されたマルチ解像度有限ボリュームにインスピレーションを得たネットワーク MuRFiV を紹介します。我々は、バーガーズ方程式、浅海方程式、非圧縮性ナビエ・ストークス方程式などの偏微分方程式 (PDE) によって支配されるいくつかの時空間システムにおける MuRFiV の有効性を実証します。 PDE 情報をディープ ラーニング アーキテクチャに埋め込むことで、MuRFiV は強力な長期予測精度を実現し、非常に長い自己回帰ロールアウトにわたって安定した状態を維持し、データ駆動型ニューラル ネットワーク ベースラインを大幅に上回ります。この結果は、多重解像度学習と有限体積からインスピレーションを得た誘導バイアスを組み合わせて、複雑なダイナミクスを正確かつロバストに長期予測できる可能性を強調しています。
原文 (English)
A Multi-Resolution Finite-Volume Inspired Deep Learning Framework for Spatiotemporal Dynamics Prediction
Predicting complex spatiotemporal dynamics in physical processes often demands computationally expensive numerical methods or data-driven neural networks that suffer from high training costs, error accumulation, and limited generalizability to unseen parameters. An effective approach to address these challenges is leveraging physics priors in training neural networks, known as physics-informed deep learning (PiDL). In this work, we introduce the Multi-Resolution Finite-Volume-inspired network, MuRFiV, designed to capitalize on the conservative property of finite volume on the global scale and the expressive power of deep learning on the local scale. We demonstrate the effectiveness of MuRFiV on several spatio-temporal systems governed by partial differential equations (PDEs), including Burgers' equation, shallow water equations, and incompressible Navier-Stokes equations. By embedding PDE information into the deep learning architecture, MuRFiV achieves strong long-term prediction accuracy and remains stable over very long autoregressive rollouts, significantly outperforming data-driven neural network baselines. This result highlights the promise of combining multiresolution learning with finite-volume-inspired inductive bias for accurate and robust long-term prediction of complex dynamics.
深層人工ニューラルネットワークと機械学習手法のアンサンブルを使用して、致死的結果 (原因) を予測し、急性心筋梗塞に関連する主要なバイオマーカーを理解する
心血管疾患は依然として世界中で主な死因の 1 つです。急性心筋梗塞(MI)または心臓発作により、毎年数百万人の命が奪われています。 MI は、冠状動脈への血流が遮断または減少すると発生し、心筋に永久的な損傷を引き起こします。治療しなければ心停止につながり、心臓が臓器への血液の送り出しを停止し、臓器不全や死に至る可能性があります。生存者であっても、心不全、肺水腫、心停止などの深刻な問題に直面することがよくあります。研究によると、生存者の5~10パーセントは心筋梗塞後1年以内に死亡し、半数近くは再び入院する必要がある。早期の血栓溶解療法はより良い転帰につながるため、MIをより迅速かつ正確に診断する方法が明らかに必要とされています。現在、医師は通常、患者の病歴を調査し、自身の経験に基づいて心筋梗塞の原因を見つけます。このプロセスには時間がかかり、一貫性がなくなる可能性があります。 MI を正確かつ迅速に検出することは、患者が自分自身をよりよく管理し、致命的な出来事を防ぐのに役立ちます。この研究では、心筋梗塞の致死的な転帰を予測し、医師がその合併症に関連する重要なバイオマーカーを理解するのに役立つ自動モデルを導入します。このアプローチは、診断をより明確に、より迅速に、より手頃な価格で行うことを目的としています。このプロセスには、データの準備、欠損値の埋め込み、SVMSMOTE、ADASYN、およびクラス加重メソッドを使用した不均衡なデータの処理が含まれます。ラッパーと埋め込み機能の選択を使用して最も重要な変数を見つけ、一貫性を保つために機能をスケールします。このモデルは、ロジスティック回帰、ランダム フォレスト、Light-GBM、およびバギング SVM を組み合わせており、人工ニューラル ネットワークを使用してさらに改良され、精度が向上しています。私たちは、精度、再現率、その他の主要な尺度を使用してすべてのモデルを評価し、臨床使用に最適なオプションを見つけます。
原文 (English)
Predicting Lethal Outcome (Cause) And Understanding Key Biomarkers Linked With Acute Myocardial Infarction Using Deep Artificial Neural Network And Ensemble Of Machine Learning Methodologies
Cardiovascular disease is still one of the main causes of death around the world. Acute myocardial infarction (MI), or heart attack, claims millions of lives each year. MI happens when blood flow to the coronary arteries is blocked or reduced, which causes permanent damage to the heart muscle. Without treatment, this can lead to cardiac arrest, where the heart stops pumping blood to the organs, resulting in organ failure and death. Even survivors often face serious problems like heart failure, pulmonary edema, and asystole. Research shows that 5 to 10 percent of survivors die within the first year after an MI, and nearly half need to be hospitalized again. Early thrombolytic treatment leads to better outcomes, so there is a clear need for faster and more accurate ways to diagnose MI. Right now, doctors usually review patient history and use their own experience to find the causes of MI. This process takes a lot of time and can be inconsistent. Detecting MI accurately and quickly can help patients take better care of themselves and prevent fatal events. In this study, we introduce an automated model to predict deadly outcomes of MI and help doctors understand important biomarkers linked to its complications. This approach aims to make diagnosis clearer, faster, and more affordable. The process includes preparing the data, filling in missing values, and handling imbalanced data using SVMSMOTE, ADASYN, and class-weighted methods. We use wrapper and embedded feature selection to find the most important variables, then scale the features for consistency. The model combines Logistic Regression, Random Forest, Light-GBM, and Bagging SVM, and is further improved with an artificial neural network to increase accuracy. We evaluate all models using precision, recall, and other key measures to find the best option for clinical use.
プロンプトを超えて: シミュレートされたモデレーション トレースによる関数呼び出し LLM の脱獄
ジェイルブレイク攻撃は、大規模言語モデル (LLM) の安全な展開にとって依然として重大な脅威です。これまでの研究では主にプロンプト レベルでの攻撃と防御が研究されてきましたが、このプロンプト中心のパラダイムではステートフルな関数呼び出し環境の構造的脆弱性が見落とされていることを示します。このようなアプリケーションでは、開発者定義のスキーマ、構造化引数、および信頼できないツールの出力が単一の共有モデル コンテキストにインターリーブされます。このアーキテクチャは、信頼できる制御ロジックと信頼できないデータの間の境界を曖昧にすることで攻撃対象領域を拡大し、敵対的な意図が複数ターンの実行パス全体に分散されることを可能にします。私たちは、Simulated Moderation Traces に基づくブラックボックス攻撃フレームワークである SMT を通じて、このアーキテクチャ上の欠陥を悪用します。 SMT は、純粋にプロンプトベースの対話から離れて、正当なモデレーション監査ワークフローをシミュレートするマルチターン軌道を構築します。この軌道の中で、でっち上げられた穏健枠は、有害な世代を引き出す口実としてレッドチームテストを利用します。その後の検証フィードバックは、安全性の拒否を実行失敗として扱い、モデルの安全性制約を徐々に弱め、最終的には有害な出力を引き起こす改良を促します。 2 つの標準化された安全性ベンチマークにわたる 5 つの異なるプロバイダーの著名な商用 LLM に関する広範な実証評価では、SMT が最小限に近いクエリ数でありながら、最高の平均攻撃成功率と HarmScore を一貫して達成しており、既存のベースラインを大幅に上回っていることが示されています。これらの調査結果は、プロンプトレベルのサニタイズだけではツール対応の LLM システムを防御するには根本的に不十分であることを示しており、スキーマ、引数、ツール出力、蓄積された会話状態にわたるコンテキスト認識型検証の緊急の必要性を浮き彫りにしています。コードは https://github.com/liujlong27/SMT で入手できます。
原文 (English)
Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces
Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied attacks and defenses at the prompt level, we show that this prompt-centric paradigm overlooks a structural vulnerability in stateful, function-calling environments. In such applications, developer-defined schemas, structured arguments, and untrusted tool outputs are interleaved into a single shared model context. This architecture expands the attack surface by blurring the boundary between trusted control logic and untrusted data, allowing adversarial intent to be distributed across a multi-turn execution path. We exploit this architectural flaw through SMT, a black-box attack framework based on Simulated Moderation Traces. Departing from purely prompt-based interactions, SMT constructs a multi-turn trajectory that simulates a legitimate moderation-auditing workflow. Within this trajectory, a fabricated moderation frame leverages red-team testing as a pretext to elicit harmful generations. The subsequent validation feedback treats safety refusals as execution failures, prompting refinements that gradually weaken the model's safety constraints and ultimately trigger harmful outputs. Extensive empirical evaluations on prominent commercial LLMs from five different providers across two standardized safety benchmarks show that SMT consistently achieves the highest average attack success rate and HarmScore while requiring a near-minimal number of queries, substantially outperforming existing baselines. These findings demonstrate that prompt-level sanitization alone is fundamentally insufficient for defending tool-enabled LLM systems and highlight the urgent need for context-aware validation across schemas, arguments, tool outputs, and accumulated conversation state. The code is available at https://github.com/liujlong27/SMT.
PAPA: オンラインでパーソナライズされたアクティブな好みの調整
拡散モデルは、画像やテキストなどの複雑なデータ分布をモデル化するのに非常に効果的です。ただし、パーソナライズされたレコメンダー システムのようなアプリケーションでは、多くの場合、目的はユーザーの好みを最大化する分布の特定の領域をモデル化することに移行します。最初は不明でしたが、インタラクティブなフィードバックを通じて徐々に明らかになります。これは当然、強化学習の問題として組み立てることができます。その目的は、好みに基づいて報酬関数を最大化するように拡散モデルを微調整することです。ただし、主な課題はパラメータ化された報酬モデルを学習することにあり、これには通常、大規模な嗜好データが必要ですが、実際には実現不可能な場合が多いです。この研究では、リアルタイムのユーザー フィードバックを使用して拡散モデルを直接最適化することで、パラメータ化された報酬モデルの要件を回避する新しい方法である Personalized Active Preference Alignment PAPA を紹介します。 PAPA は、変分推論フレームワークからインスピレーションを得て、フィードバック効率の高い好みの調整を可能にします。私たちは、さまざまなクラス条件付きおよび細粒度のアライメントタスクにわたる広範な実験とアブレーション研究を通じて、PAPA の有効性を実証しています。さらに、理論的な洞察に基づいて、EPAPA と呼ばれる強化された微調整戦略を提案します。これは、必要な計算予算が少なく、微調整プロセスを加速し、現実世界の展開に対する PAPA の適合性をさらに高めます。私たちのコードは https://github.com/NasikNafi/papa で公開されています。
原文 (English)
PAPA: Online Personalized Active Preference Alignment
Diffusion models are highly effective at modeling complex data distributions, including images and text. However, in applications like personalized recommender systems, the objective often shifts to modeling specific regions of the distribution that maximize user preferences-initially unknown but gradually uncovered through interactive feedback. This can naturally be framed as a reinforcement learning problem, where the goal is to fine-tune a diffusion model to maximize a reward function based on preferences. However, the main challenge lies in learning a parameterized reward model, which typically requires large-scale preference data-something that is often not feasible in practice. In this work, we introduce Personalized Active Preference Alignment PAPA, a novel method that bypasses the requirement for a parametrized reward model by directly optimizing the diffusion model using real-time user feedback. PAPA enables feedback-efficient preference alignment, drawing inspiration from the variational inference framework. We demonstrate PAPA's effectiveness through extensive experiments and ablation studies across diverse class-conditioned and fine-grained alignment tasks. Additionally, based on theoretical insights, we propose an enhanced fine-tuning strategy, referred to as EPAPA, that requires less computational budget and accelerates the fine-tuning process, further boosting PAPA's suitability for real-world deployment. Our code is made publicly available at https://github.com/NasikNafi/papa.
MindEdit-Bench: 野生の写真からの VLM におけるオブジェクトレベルの反事実的空間推論のベンチマーク
ビジョン言語モデル (VLM) のベンチマークは主に、観察による空間推論をテストします。モデルは、入力内にすでに表示されている関係を記述します。既存の what-if タスクは通常、シーンを固定したまま観察者を変化させます。代わりに、VLM は、オブジェクトを仮想的に移動または回転させた場合の結果を予測できますか? MindEdit-Bench は、自動の野外 3D シーングラフ抽出パイプラインを介して、新たにキャプチャされた屋内シーンの 3 枚のスマートフォン写真 3 枚から構築された 6 つの空間推論タスクのベンチマークです。 4 つのタスクは、観察された構造に対する知覚と視点の変換を調査します。 2 つの新しいタスク、L4 (空間編集) と L5 (クロスビュー可視性編集) は、すべての入力画像に正しい答えが存在しないオブジェクト レベルの反事実推論を調査します。各質問には 8 ~ 24 個の構造化された回答選択肢が用意されており、空間エラーとフォールバック エラーを回答文字レベルで診断できます。このベンチマークは、公開データセットから抽出されていない 120 のプライベート屋内シーンをカバーしており、公開データの事前トレーニングと重複のリスクを軽減します。人間が検証した 1,003 の質問に対する 15 の VLM 全体で、タスクごとの VLM の平均精度はわずか 8% ~ 31% であるのに対し、人間の多数決の精度は 81% ~ 97% です。プールされた人間と最高の VLM の差は 53 pp で、すべてのタスクで少なくとも 39 pp です。構造化された回答空間では、カメラの奥行き軸の推論が弱いことや、可視性の編集が難しい場合のフォールバック動作など、不均一な失敗がさらに明らかになります。
原文 (English)
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene fixed. Can VLMs instead predict the consequences of hypothetically moving or rotating an object? We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly captured indoor scenes via an automatic in-the-wild 3D scene-graph extraction pipeline. Four tasks probe perception and perspective transformation over observed structure; two new tasks, L4 (spatial editing) and L5 (cross-view visibility editing), probe object-level counterfactual reasoning, where correct answers are absent from all input images. Each question provides 8-24 structured answer choices, enabling answer-letter-level diagnosis of spatial and fallback errors. The benchmark covers 120 private indoor scenes not drawn from public datasets, reducing public-data pretraining-overlap risk. Across 15 VLMs on 1,003 human-verified questions, task-wise mean VLM accuracy is only 8%-31%, versus 81%-97% human majority-vote accuracy. The pooled human--best-VLM gap is 53 pp, with at least 39 pp on every task. The structured answer space further reveals non-uniform failures, including weaker camera-depth-axis inference and fallback behavior on difficult visibility-editing cases.
BaseRT: Native Metal を介した Apple Silicon でのクラス最高の LLM 推論
Apple Silicon 上の大規模言語モデル (LLM) 用のネイティブ Metal 推論ランタイムである BaseRT を紹介し、このハードウェアでこれまでで最高の推論スループットを報告します。 llama.cpp や MLX ベースのフレームワークなどの既存のランタイムでは、Metal の実行モデルや Apple Silicon のユニファイド メモリ トポロジ向けに設計されていない抽象化によるオーバーヘッドが発生します。 BaseRT は、チップ固有のカーネル フュージョン、統合されたメモリを意識した最適化、カスタム ディスパッチ ロジックを使用して Metal 上にネイティブに構築することで、フレームワーク ベースのアプローチでは残されたパフォーマンスを回復します。 BaseRT は、すべての Apple M シリーズ デバイスで 8 つの量子化フォーマット (Q2 から FP16) にわたる幅広いモデル ファミリをサポートします。このペーパーでは、M3 および M4 Pro デバイス上の Q4 および Q8 量子化で Qwen3、Llama 3.2、および Gemma 4 ファミリを評価します。 BaseRT は、llama.cpp よりも最大 1.56 倍、MLX よりも最大 1.35 倍高いデコード スループットを達成し、専門家混合モデルのプリフィルのマージンが大幅に大きくなり、サブ 1B から 30B パラメータ モデルまで一貫してクラス最高のスループットを実現します。これらの結果は、Apple Silicon が以前に報告されていたよりも優れた推論プラットフォームであることを確立しており、新たなエッジ推論パラダイムに直接的な影響を及ぼします。プライバシー要件、レイテンシの制約、クラウドのコスト圧力により推論がオンデバイス展開に向けて推進されているため、パフォーマンスが最適化されたローカル ランタイムがこの移行を可能にする重要なレイヤーとなります。 BaseRT は https://github.com/basecompute/baseRT で公開されています
原文 (English)
BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal
We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date. Existing runtimes, including llama.cpp and MLX-based frameworks, incur overhead from abstractions not designed for Metal's execution model or Apple Silicon's unified memory topology. By building natively on Metal with chip-specific kernel fusion, unified memory-aware optimisation, and custom dispatch logic, BaseRT recovers performance that framework-based approaches leave on the table. BaseRT supports a wide range of model families across eight quantisation formats (Q2 to FP16) on all Apple M-series devices. In this paper, we evaluate the Qwen3, Llama 3.2, and Gemma 4 families at Q4 and Q8 quantisation on M3 and M4 Pro devices. BaseRT achieves up to 1.56x higher decode throughput than llama.cpp and up to 1.35x higher than MLX, with substantially larger margins on prefill for mixture-of-experts models, delivering consistent best-in-class throughput from sub-1B to 30B parameter models. These results establish Apple Silicon as a more capable inference platform than previously reported, with direct implications for the emerging edge inference paradigm: as privacy requirements, latency constraints, and cloud cost pressures drive inference toward on-device deployment, performance-optimised local runtimes are a critical enabling layer for this transition. BaseRT is publicly available at https://github.com/basecompute/baseRT
Cross4D-JEPA: 4D 点群表現学習のための高密度クロスモーダル対応蒸留
動的な 4D 点群、つまり深度センサーと LiDAR によって時間の経過とともにキャプチャされた 3D 点のシーケンスの自動理解は、ロボット工学と身体的知覚の中心となります。しかし、それらに高密度にアノテーションを付けるのはコストがかかるため、自己教師あり事前トレーニングが転送可能な表現への自然なルートになります。しかし、既存のプリテキスト タスクはほぼ完全にイントラモーダルであり、2D 基礎モデルから知識を転送するいくつかのメソッドは、クリップごとに単一のグローバル エンベディングに依存しており、これらのモデルが計算する豊富なパッチごとのセマンティクスを無視しています。このギャップに対処するために、凍結された 2D 基礎モデル、画像モデル DINOv2、またはビデオ モデル V-JEPA 2 を 4D ポイント エンコーダに抽出する教師と生徒の方法である Cross4D-JEPA を提案します。提案された方法は、(1) すべての 3D ポイントを投影先の教師パッチ特徴にマッピングする密なクロスモーダル対応と、(2) マスキング、ネガ、またはデコーダーを使用せずに潜在空間でこれらの特徴を照合するように生徒を訓練するポイントごとの目標を組み合わせています。 MSR-Action3D、DeformingThings4D、NTU-RGB+D 60、HOI4D の 4 つのベンチマークで、イントラモーダルおよびグローバル クロスモーダル ベースラインに対して Cross4D-JEPA を評価します。実験結果は、一致したプロトコルの下で、提案された手法が 4 つのベンチマーク全体で一貫してイントラモーダルおよびグローバル クロスモーダル ベースラインを上回り、より強力な公開された 4D 手法と競合できることを示しています。さらなる分析では、この利点は教師のモダリティではなく、主に対応の粒度によるものであると考えられます。認識精度を超えて、Cross4D-JEPA によって学習された高密度表現はドメイン間で転送され、ラベル効率が向上し、同じトレーニング予算の下でラベル全体の微調整が向上します。また、13 分の 1 小さいエンコーダーが重量プーリング バックボーンに適合します。
原文 (English)
Cross4D-JEPA: Dense Cross-modal Correspondence Distillation for 4D Point Cloud Representation Learning
Automatic understanding of dynamic 4D point clouds, the 3D-point sequences captured over time by depth sensors and LiDAR, is central to robotics and embodied perception. Yet annotating them densely is expensive, making self-supervised pretraining the natural route to transferable representations. Existing pretext tasks, however, are almost entirely intra-modal, and the few methods that transfer knowledge from 2D foundation models rely on a single global embedding per clip, discarding the rich per-patch semantics that these models compute. To address this gap, we propose Cross4D-JEPA, a teacher-student method that distills a frozen 2D foundation model, an image model DINOv2, or a video model V-JEPA 2, into a 4D point encoder. The proposed method combines (1) a dense cross-modal correspondence that maps every 3D point to the teacher patch feature it projects to, and (2) a per-point objective that trains the student to match these features in latent space with no masking, negatives, or decoder. We evaluate Cross4D-JEPA on four benchmarks, MSR-Action3D, DeformingThings4D, NTU-RGB+D 60, and HOI4D, against intra-modal and global cross-modal baselines. Experimental results show that, under a matched protocol, the proposed method consistently outperforms intra-modal and global cross-modal baselines across the four benchmarks and is competitive with heavier published 4D methods; further analysis attributes this gain primarily to the granularity of the correspondence rather than the teacher modality. Beyond recognition accuracy, the dense representation learned by Cross4D-JEPA transfers across domains, improves label efficiency, and improves full-label fine-tuning under the same training budget, while a 13x smaller encoder matches a heavyweight pooling backbone.
AI、信頼、チーム化: 自律的かつ不透明な AI システムのためのハンドラーとしての人間アプローチ
人工知能 (AI) はユビキタスになりつつあり、ドメイン全体でますます自律的なシステムが、重大な倫理的および法的課題を引き起こすタスクを実行するようになっており、これは信頼に根ざした強力な人間と機械のチームの必要性を示しています。この記事では、非常に影響力の大きい分野(医学や戦闘など)において、私たちが当初、自律的で不透明なシステムを犬(または私たちが密接な関係にある他の動物)に類似したものとして扱う根拠があると主張します。このアナロジーの下では、これらのシステムを利用する人間は、これらのシステムの「ユーザー」または「導入者」とみなされるべきではなく、代わりに「ハンドラー」の役割を引き受けます。この役割の再設定により、人間、AI 対応および自律システム、およびそれらの間の関係に対する見方が変わり、さらに、これらのシステムの使用時にもたらされる結果に対して人間が負う明確かつ追跡可能な責任範囲が明確になります。この点を展開する際に、私は機械と動物のアナロジーが確かに非類似要素を認めているが、その接点がそれを出発点として基礎づけていることを明確にしました。次に、自律システムや AI 対応システムとの関わり方や利用方法に不適当な、動物との関係の側面における、人間がハンドラーとしてのアプローチをどのように取り除くことができるかを模索します。私は、自律システムと AI 対応システムのための人間と機械のチーム化の軌跡は、これらを単に利用する成果物としてではなく、複雑な目標を追求し、複雑なタスクを実行する協力者として真に見なす状態でなければならないと主張して、結論を述べます。
原文 (English)
AI, Trust, and Teaming: The Humans-as-Handlers Approach for Autonomous and Opaque AI Systems
Artificial intelligence (AI) is becoming ubiquitous, and across domains, increasingly autonomous systems are carrying out tasks which raise significant ethical and legal challenges which demonstrate a need for strong human-machine teams rooted in trust. In this article, I argue that within highly impactful areas (such as medicine or warfighting) there are grounds for us initially treating autonomous and opaque systems as relevantly analogous to dogs (or other animals with which we have close relationships). Under this analogy, humans making use of these systems are not to be viewed as "users" or "deployers" of these systems, but instead take the role of "handlers". This recasting of roles shifts the way we view humans, AI-enabled and autonomous systems, and the relations between them, and moreover clarifies the clear and traceable lines of responsibility humans have for the outcomes brought about when using these systems. In developing this point, I clarify that the machine-animal analogy does admit disanalogous elements, but that its touch-points ground it as a starting point. I then explore how we can divest the humans-as-handlers approach of those aspects of our relationships with animals which are unfitting for how we engage with and make use of autonomous and AI-enabled systems. I conclude by arguing that the trajectory of human-machine teamings for autonomous and AI-enabled systems should be a state where we authentically view these not as artifacts which we simply make use of, but as collaborators with which we pursue complex goals and carry out complex tasks.
技術的なメトリクスからユーザーの知覚まで: 物体の検出と把握のためのマルチモーダルなヒューマンロボットインタラクションシステムのユーザー研究
人間とロボットのインタラクション (HRI) システムの技術的パフォーマンスの向上は、人間のユーザーがライブ インタラクション中に検出できる違いに自動的に変換されるわけではありません。この論文では、エンドツーエンドのタスクの成功率が 15 パーセント ポイント向上する (マルチモーダル ベースライン システムの 75% から、以前のアブレーション研究で特定された改善された構成の 90%) が、ユーザーの認識に一貫した測定可能な違いを生み出すのに十分であるかどうかを調査します。ベースライン システムは、音声認識用の Whisper、オープン語彙オブジェクト検出用の Florence-2、アクション抽出用の LLaMA 3.1、およびモーション実行用のインターバル タイプ 2 ファジー ロジック コントローラーを組み合わせています。改善された構成では、同じコントローラーを維持しながら、知覚モジュールと言語モジュールをそれぞれ Grounding DINO + SAM と Qwen 3.5 9B に置き換えます。 24 人の参加者による被験者内ユーザー研究では、同じ卓上物体把握タスクで両方のシステムを比較しました。各構成を操作した後、参加者は体感速度、信頼性、全体的な能力と流暢さを 7 段階のリッカート スケールで評価しました。結果は、参加者 24 人中 17 人 (70.83%) が改善されたシステムを好み (正確な二項検定、p = 0.043、h = 0.43)、3 つの知覚構造すべてが、ホルム補正後の改善された構成に対して有意に高く評価され、効果サイズが大きいから非常に大きい (p < 0.001) ことを示しました。これらの調査結果は、特定された技術的改善がユーザーに直接インタラクションで認識できることを裏付けており、ロボット操作パイプラインを評価する際にユーザー中心の証拠でベンチマーク評価を補完することの重要性を強調しています。
原文 (English)
From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 15 percentage point gain in end-to-end task success (from 75% in a multimodal baseline system to 90% in an improved configuration identified through a prior ablation study) is sufficient to produce consistent and measurable differences in user perception. The baseline system combines Whisper for speech recognition, Florence-2 for open-vocabulary object detection, LLaMA 3.1 for action extraction, and an interval Type-2 fuzzy logic controller for motion execution. The improved configuration replaces the perception and language modules with Grounding DINO + SAM and Qwen 3.5 9B, respectively, while retaining the same controller. A within-subject user study with 24 participants compared both systems on the same tabletop object-grasping task. After interacting with each configuration, participants rated perceived speed, reliability, and overall competence and fluency on a 7-point Likert scale. Results show that 17 out of 24 participants (70.83%) preferred the improved system (exact binomial test, p = 0.043, h = 0.43), and all three perceptual constructs were rated significantly higher for the improved configuration after Holm correction, with large to very large effect sizes (p < 0.001). These findings confirm that the identified technical improvements are perceptible to users in direct interaction and underscore the importance of complementing benchmark evaluation with user-centred evidence when assessing robotic manipulation pipelines.
Active-GRPO: 分子最適化のための適応的模倣と自己改善的推論
科学的推論は大規模言語モデルの機能としてますます重要になっていますが、そのような推論のトレーニングの堅牢性と効率を向上させることは依然として重要な未解決の課題です。私たちはこの問題を命令ベースの分子最適化で研究しています。この問題では、回答のみの教師あり微調整 (SFT) が多段階推論を崩壊させ、検証可能な報酬を伴う強化学習 (RLVR) がまばらなフィードバックに悩まされます。参照ガイドに基づくポリシーの最適化は、データセットが提供する参照にポリシーの更新を固定することで両方を軽減しますが、その有効性は参照の品質と密接に関係しており、弱い参照や不整合な参照はパフォーマンスの上限を課します。この上限を克服するために、私たちは積極的推論を提案します。これは、ポリシーが、模倣するものを継続的にアップグレードしながら、いつ参照を模倣するか、いつ自身の発見を強化するかをインスタンスごとに積極的に決定するパラダイムです。このパラダイムは、アクティブ グループ相対ポリシー最適化 (Active-GRPO) としてインスタンス化され、アクティブな模倣強化とアクティブな参照という 2 つの結合メカニズムによって実現されます。前者は、参照がまだ政策自身の候補よりも優れている場合に模倣学習を実行し、政策が参照を超える分子を生成したら強化学習による自己改善に移行します。後者は、これまでに発見された最良のポリシー生成候補と置き換えることによって参照自体を継続的にアップグレードし、模倣ターゲットを段階的に引き上げ、トレーニング全体を通じて参照ガイダンスが制限的なものではなく有益なままであることを保証します。 TOMG-Bench MOLOPT 全体で、Active-GRPO は、一致した 3 シード評価で平均 SRxSim を GRPO の 0.0959、RePO の 0.1665 から 0.1773 に改善し、LogP、MR、および QED で統計的に有意な向上をもたらしました。
原文 (English)
Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization
Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open challenge. We study this problem in instruction-based molecular optimization, where answer-only supervised fine-tuning (SFT) collapses multi-step reasoning and reinforcement learning with verifiable rewards (RLVR) suffers from sparse feedback. Reference-guided Policy Optimization mitigates both by anchoring policy updates to dataset-provided references, but its effectiveness is tightly coupled to reference quality: weak or misaligned references impose a performance ceiling. To overcome this ceiling, we propose active reasoning, a paradigm in which the policy actively decides, on a per-instance basis, when to imitate a reference and when to reinforce its own discoveries, while continuously upgrading what it imitates. We instantiate this paradigm as Active Group Relative Policy Optimization (Active-GRPO), realized through two coupled mechanisms: active imitate-reinforce and active referencing. The former performs imitation learning when the reference still outperforms the policy's own candidates, and shifts to self-improvement via reinforcement learning once the policy has generated molecules that surpass the reference. The latter continuously upgrades the reference itself by replacing it with the best policy-generated candidate discovered so far, progressively raising the imitation target and ensuring that reference guidance remains informative-rather than restrictive-throughout training. Across TOMG-Bench MOLOPT, Active-GRPO improves average SRxSim from 0.0959 for GRPO and 0.1665 for RePO to 0.1773 under matched three-seed evaluation, with statistically significant gains on LogP, MR, and QED.
フローマップ GRPO: アンカーされた確率的合成による数ステップのフローマップ ジェネレーターの強化学習
一貫性モデルや MeanFlow などの数ステップのフロー マップ ジェネレーターは、ノイズとデータの間の長距離トランスポート マップを直接学習することでサンプリングを高速化します。ただし、これらのモデルは通常は決定論的であるため、確率的軌道と明確に定義された尤度比を必要とする強化学習 (RL) ポストトレーニング手法で最適化することが困難になります。既存の SDE ベースの確率化手法は、無限小または細かく離散化された遷移を伴う速度ベースのサンプラー向けに設計されているため、長距離フロー マップには直接適用できません。この研究では、決定論的な数ステップのフローマップ ジェネレーター用のオンライン RL ポストトレーニング フレームワークである Flow-Map GRPO を提案します。主要なコンポーネントはアンカー確率的フロー マップ構成 (ASFMC) です。これは、決定論的フロー マップの元の限界確率パスを保存しながら、アンカーベースの条件付きリサンプリングを通じてランダム性を導入するパス保存確率化メカニズムです。 1 回と 2 回の両方のフローマップ パラメーター化の GRPO 目標を導出します。 MeanFlow や sCM などの数ステップの FLUX ベースのテキストから画像へのジェネレーターの実験では、Flow-Map GRPO が報酬ベース、知覚、およびタスクレベルの評価指標全体にわたって事前トレーニング済みの決定論的フローマップ モデルを改善することが示されています。私たちの結果は、決定論的な数ステップのフローマップジェネレーターが、元のモデルのパラメーター化を変更したり、ネイティブの確率モデルとして再トレーニングしたりすることなく、トレーニング後のRLと効果的に調整できることを示しています。
原文 (English)
Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition
Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them difficult to optimize with reinforcement learning (RL) post-training methods that require stochastic trajectories and well-defined likelihood ratios. Existing SDE-based stochasticization techniques are designed for velocity-based samplers with infinitesimal or finely discretized transitions, and therefore do not directly apply to long-range flow maps. In this work, we propose Flow-Map GRPO, an online RL post-training framework for deterministic few-step flow-map generators. The key component is Anchored Stochastic Flow Map Composition (ASFMC), a path-preserving stochasticization mechanism that introduces randomness through anchor-based conditional resampling while preserving the original marginal probability path of the deterministic flow map. We derive GRPO objectives for both single-time and two-time flow-map parameterizations. Experiments on few-step FLUX-based text-to-image generators, including MeanFlow and sCM, show that Flow-Map GRPO improves pretrained deterministic flow-map models across reward-based, perceptual, and task-level evaluation metrics. Our results demonstrate that deterministic few-step flow-map generators can be effectively aligned with RL post-training without modifying their original model parameterization or retraining them as native stochastic models.
EgoGapBench: マルチエージェント シーンにおける自己中心的なアクション選択のベンチマーク
既存の自己中心的なベンチマークは、主に一人称視点のデータから自己中心的な設定を構築しているため、自己中心的な視点自体を単独で評価することが困難になります。ただし、一人称視点の入力を理解することと自己中心的な視点を取ることは、特に一人称の身体の手がかりが存在しない場合や他のエージェントが存在する場合には、分離可能な能力です。自己中心的な視点の理解を分離するために、マルチエージェントの自己中心的なシーンでのアクション選択を測定するための診断ベンチマークである EgoGapBench を導入します。このベンチマークによって測定される能力を、他のエージェントの存在下でエージェントの観点から適切な行動を選択する自己中心的行動選択 (EAS) として定義します。 EgoGapBench では、人間は確実に応答しますが、オープンソースとプロプライエタリの MLLM はどちらもパフォーマンスが大幅に悪く、他の目に見えるエージェントによって実行されるアクションを体系的に選択します。既存の自己中心的なデータを微調整しても、このギャップは埋められず、有害になる可能性さえあります。対照的に、EgoGapBench トレーニング データを微調整すると精度は向上しますが、人間のパフォーマンスには達しません。これらの結果は、EAS を一人称視点データだけから取得するのは困難であること、MLLM はシーンの理解だけでなく自己中心的なアクション選択についても評価およびトレーニングする必要があることを示しています。
原文 (English)
EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes
Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation. However, understanding first-person-view input and taking an egocentric perspective are separable abilities, especially when first-person body cues are absent or when other agents are present. To isolate egocentric perspective understanding, we introduce EgoGapBench, a diagnostic benchmark for measuring action selection in multi-agent egocentric scenes. We define the ability measured by this benchmark as Egocentric Action Selection (EAS): selecting an appropriate action from the agent's perspective in the presence of other agents. On EgoGapBench, humans answer reliably, whereas both open-source and proprietary MLLMs perform substantially worse and systematically select actions performed by other visible agents. Fine-tuning on existing egocentric data fails to close this gap and can even be detrimental. In contrast, fine-tuning on EgoGapBench training data improves accuracy but does not reach human performance. These results show that EAS is difficult to acquire from first-person-view data alone, and that MLLMs should be evaluated and trained not only for scene understanding but also for egocentric action selection.
IIoT ネットワークの軽量侵入検知モデルにおけるクロスドメイン一般化の失敗
軽量の機械学習モデルは、リソースに制約のあるエッジ展開に適しているため、産業用モノのインターネット (IIoT) ネットワークでの侵入検知に提案されることが増えています。報告された結果のほとんどは、トレーニング ネットワーク内でのみこれらのモデルを評価し、目に見えないネットワークでの動作は検証されていません。この調査では、1 つの IIoT データセットで 4 つの軽量アーキテクチャをトレーニングし、3 つのソースすべてで利用可能な属性に限定された特徴表現を使用して、構造的に異なる 2 つの IIoT データセットで再トレーニングせずにそれらを評価します。 2 つの最高パフォーマンスのモデルにわたる説明可能性分析では、どちらも大まかなポート カテゴリの特徴に圧倒的に依存していることがわかります。最も影響力のあるカテゴリは、ソース ドメインの攻撃トラフィックで 2 つのターゲット ドメインの 96 ~ 435 倍の割合で発生します。これは、ポート解決の粗密化により、文書化されたショートカットが削除されるのではなく、再配置されることを示しています。自然に不均衡なクラス分布の下で評価すると、さらなる効果が明らかになります。使用される評価プロトコルにより、どのターゲット ネットワークがより大きな一般化の課題を引き起こしているように見えるかを逆転させることができます。限られたターゲットドメインへのエクスポージャーによる敵対的な堅牢性と回復も評価されます。敵対的な摂動に対する堅牢性は、ネットワーク間の一般化とは無関係であり、適応による回復はアーキテクチャによって大きく異なります。これらの調査結果は、ドメイン内の精度だけではなく、現実的なクラス分布に基づくクロスネットワーク評価を使用して展開の準備状況を評価する必要があることを示唆しています。
原文 (English)
Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks
Lightweight machine learning models are increasingly proposed for intrusion detection in Industrial Internet of Things (IIoT) networks due to their suitability for resource-constrained edge deployment. Most reported results evaluate these models only within their training network, leaving behavior on unseen networks unverified. This study trains four lightweight architectures on one IIoT dataset and evaluates them, without retraining, on two structurally distinct IIoT datasets using a feature representation restricted to attributes available across all three sources. Explainability analysis across two top-performing models shows both rely overwhelmingly on coarse port-category features; the most influential category occurs in source-domain attack traffic at 96 to 435 times the rate in the two target domains, indicating that coarsening port resolution relocates rather than removes a documented shortcut. Evaluation under naturally imbalanced class distributions reveals a further effect: the evaluation protocol used can reverse which target network appears to pose the greater generalization challenge. Adversarial robustness and recovery through limited target-domain exposure are also assessed; robustness to adversarial perturbation is unrelated to cross-network generalization, and recovery through adaptation varies considerably by architecture. These findings suggest deployment readiness should be assessed using cross-network evaluation under realistic class distributions, rather than within-domain accuracy alone.
群等変ポアンカール畳み込みネットワーク
Poincar\'e ResNet のような最近の進歩は、双曲空間で視覚表現を直接学習できる可能性を実証しましたが、その最適化は、リーマン勾配の計算集約的な性質と多様体の厳密な境界によって依然として妨げられています。さらに、標準的な双曲線ネットワークは、同じオブジェクトの空間変換を別個の階層概念として扱うため、パラメーターの冗長な使用と信号の消失につながります。双曲幾何学と離散対称群 ($C_4$ と $D_4$) を組み合わせた等変ポアンカレ ResNet を提案します。ユークリッド等分散を双曲空間に適用する際の重大な障害を特定し、幾何学的に安全なテンソル再整形、双曲群畳み込みの左正則置換、関節配向ポアンカレ中点バッチ正規化を提案します。経験的に、等分散を埋め込むと最適化空間が大幅に減少し、ポアンカレ ボールの境界制約を尊重し、空間グループの等分散を維持しながら収束を加速します。
原文 (English)
Group-Equivariant Poincar\'e Convolutional Networks
While recent advancements like the Poincar\'e ResNet have demonstrated the potential of learning visual representations directly in hyperbolic space, their optimisation remains hampered by the computationally intensive nature of Riemannian gradients and the strict boundaries of the manifold. Furthermore, standard hyperbolic networks treat spatial transformations of the same object as distinct hierarchical concepts, leading to redundant parameter usage and vanishing signals. We propose Equivariant Poincar\'e ResNets, combining hyperbolic geometry with discrete symmetry groups ($C_4$ and $D_4$). We identify critical roadblocks in applying Euclidean equivariance to hyperbolic space and propose geometrically safe tensor reshaping, left-regular permutations for hyperbolic group convolutions, and joint-orientation Poincar\'e Midpoint Batch normalisation. Empirically, embedding equivariance drastically reduces the optimisation space, accelerating convergence while accelerating convergence while respecting the boundary constraints of the Poincar\'e ball and preserving spatial-group equivariance.
ソフトウェア リポジトリにおける AI パターンの蔓延を調査するための方法論
人工知能 (AI) ベースのアプリケーションが普及するにつれ、AI パターンを明確に理解することで AI アプリケーションの品質を向上させることができます。多くの AI パターンが文献で提案されています。ただし、実際のコードにおけるそれらの普及はまだ検証されていません。これらのパターンを実際に実際に使用することを理解すると、これらのパターンの重要性とその有用性の両方についての理解が明確になります。この論文では、a) 文献をマイニングすることで関連するパターンを特定し、b) アクティブ ラーニングを使用して実際のコード リポジトリにおけるパターンの存在と蔓延を検証する方法論を紹介します。そのために、公開されている 44 の AI パターン関連ソースをマイニングすることで、14 の AI パターン クラスを特定します。次に、アクティブ ラーニング アプローチを使用して、100 の GitHub オープン AI リポジトリ全体で最も一般的なパターン クラスの蔓延を判断します。有病率の推定を使用して、発生の精度の限界を提案します。このモデルは、8 方向分類タスクで 56\% の精度と 55\% の再現率を達成し、ランダム チャンス ベースラインの 11\% を大幅に上回りました。さらに、普及率の推定は、パターンの適用を分析するために使用可能な限界を提供します。この方法論は、現在実証データが不足している分野で、AI パターンが実際にどのように使用されているかを理解し始めるための強固な基盤を提供します。
原文 (English)
A Methodology for Investigating AI Patterns Prevalence in Software Repositories
As Artificial Intelligence(AI)-based applications take off, a clear understanding of AI patterns can uplift the quality of AI applications. Many AI patterns have been proposed in the literature; however, their prevalence in real-life code has not yet been validated. Understanding the actual use of those patterns in practice can clarify our understanding both of the significance of these patterns and their utility. In this paper, we present a methodology to a) identify relevant patterns by mining the literature and then to b) validate their presence and prevalence in actual code repositories using active learning. To that end, we identify 14 AI pattern classes by mining 44 published AI pattern-related sources. Then we use an active learning approach to determine the prevalence of the most common pattern class across 100 GitHub open AI repositories. Using prevalence estimation, we propose bounds on the accuracy of the occurrences. The model achieves 56\% accuracy and 55\% recall in an 8-way classification task, significantly outperforming the 11\% random-chance baseline. Furthermore, the prevalence estimation offers usable bounds for analyzing pattern applications. This methodology provides a robust foundation to start understanding how AI patterns are used in practice, a field that currently lacks empirical data.
限られたメモリ言語モデルにおける忘却の監査
Limited Memory Language Model (LMLM) は事実の知識をデータベースに外部化し、再トレーニングを行わずに削除ベースのアンラーニングを可能にします。既存の評価では、削除後の正確性を総合的に測定しており、削除されたファクトが残存パラメトリック メモリ、代替検索パス、または近傍検索アーティファクトを通じて存続するかどうかを判断することはできません。私たちは、モデルを固定し、FULL、DEL-ON、DEL-OFF の 3 つの介入にわたって推論時にデータベースの状態を変化させる因果監査フレームワークを提案します。このフレームワークは、削除後の動作を、パラメトリック漏洩 L(f)、検索媒介の正確性 R(f)、および推論時間の検索トレースに基づいた検索アーティファクト率に分解します。これを、3 つのドメインで構築した 4 つの敵対的トポロジ (ベース、エイリアス、ノイズ、コリジョン) と 6 つのプロンプト定式化を含む、13 のデータベースにわたる 12,228 のエイリアス クロージャの削除に適用します。パラメトリック漏洩は、すべてのバリアントおよびすべてのプロンプト スタイルでほぼゼロです。モデルは、取得がない場合に削除された回答を返すことはほとんどありません。生き残った残差は検索グラフ内に存在します。検索媒介の正確性と検索アーティファクト率はどこでも丸め内で一致するため、監査では、削除後の正確性は主に近傍検索から再構成されます。この残差は、リリースされた LMLM データベースの 0.7% から、最も敵対的なバリアントの 13.6% までの範囲であり、プロンプト定式化では、削除されたファクトがどれだけ生き残るかを独立して制御することはできません。これらの結果は、このクラスの LMLM および削除手順では、非学習境界がモデルではなく主にデータベース管理者によって引かれていることを示唆しています。
原文 (English)
Auditing Forgetting in Limited Memory Language Models
Limited Memory Language Models (LMLMs) externalize factual knowledge to a database to enable deletion-based unlearning without retraining. Existing evaluations measure post-deletion correctness in aggregate and cannot tell whether a deleted fact persists through residual parametric memory, alternative retrieval paths, or near-neighbor retrieval artifacts. We propose a causal auditing framework that holds the model fixed and varies the database state at inference time across three interventions: FULL, DEL-ON, and DEL-OFF. The framework decomposes post-deletion behavior into parametric leakage L(f), retrieval-mediated correctness R(f), and a retrieval artifact rate grounded in the inference-time retrieval trace. We apply it to 12,228 alias-closure deletions across thirteen databases, including four adversarial topologies (Base, Alias, Noise, Collision) we construct in three domains, and six prompt formulations. Parametric leakage is near zero in every variant and every prompt style: the model rarely returns the deleted answer in the absence of retrieval. The residual that does survive lives in the retrieval graph: retrieval-mediated correctness and the retrieval artifact rate match within rounding everywhere, so post-deletion correctness is, in our audit, predominantly reconstituted from near-neighbor retrieval. This residual ranges from 0.7% on the released LMLM database to 13.6% on the most adversarial variant, and prompt formulation does not independently control how much of a deleted fact survives. These results suggest that, for this class of LMLM and deletion procedure, the unlearning boundary is drawn primarily by the database administrator rather than by the model.
一般化されたカテゴリー発見のための潜在的な概念と構造の特定
Generalized Category Discovery (GCD) は、オープンワールド設定で新しいクラスを自律的に発見しながら、既知のクラスを認識することを目的としています。しかし、現在のアプローチは主にクラスタリング目標の設計に焦点を当てており、多くの場合、重大なボトルネックを見落としています。標準ビジョンのバックボーンは、潜在的な概念や構造の教師なし発見には不向きな、高ランクのもつれたトークン表現を生成します。この論文では、低ランクの構成的組織化を強制することでそのような潜在構造を識別可能にするために特徴空間を再形成する新しい表現学習フレームワークである構成プリミティブ フィールド (CPF-GCD) を提案します。私たちの中心となる仮説は、既知か新規かに関係なく、すべてのカテゴリは、再利用可能な概念を捉える学習可能な視覚プリミティブの有限セットの構成および空間配置として表現できるということです。 CPF は、空間フィールド メカニズムを介してこの幾何学的制約をインスタンス化します。バックボーンとヘッドの間に挿入され、低ランクのプリミティブ混合物を通じてノイズの多いパッチ トークンを書き換え、画像を再利用可能な原子部分とその空間レイアウトに効果的に分解します。 CPF は、プリミティブの空間分布を明示的にモデル化することにより、共有語彙上で新しいカテゴリを新しい活性化パターンとして自然に出現させることができます。これにより、表現の焦点が単にグローバル エンベディングを分割することから、構造化された分離可能なプリミティブ フィールドの構築へと移ります。広範な実験により、CPF がさまざまな GCD ベースラインにわたって一貫してパフォーマンスを向上させる汎用のプラグアンドプレイ モジュールとして機能することが実証され、低ランクの構成構造の特定と活用がオープンワールド認識にとって重要な誘導バイアスであることが検証されました。
原文 (English)
Identifying Latent Concepts and Structures for Generalized Category Discovery
Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. However, current approaches primarily focus on designing clustering objectives, often overlooking a critical bottleneck: standard vision backbones yield high-rank, entangled token representations that are ill-suited for unsupervised discovery of latent concepts and structures. In this paper, we propose Compositional Primitive Fields (CPF-GCD), a novel representation learning framework that reshapes the feature space to make such latent structure identifiable by enforcing a low-rank compositional organization. Our core hypothesis is that all categories, whether known or novel, can be expressed as compositions and spatial arrangements of a finite set of learnable visual primitives that capture reusable concepts. CPF instantiates this geometric constraint via a spatial field mechanism. Inserted between the backbone and the head, it rewrites noisy patch tokens through low-rank primitive mixtures, effectively decomposing images into reusable atomic parts and their spatial layouts. By explicitly modeling the spatial distribution of primitives, CPF enables novel categories to emerge naturally as new activation patterns over a shared vocabulary. This shifts the focus of representation from merely partitioning global embeddings to constructing a structured and separable primitive field. Extensive experiments demonstrate that CPF serves as a generic, plug-and-play module that consistently boosts performance across diverse GCD baselines, validating that identifying and leveraging low-rank compositional structure is a crucial inductive bias for open-world recognition.
分布シフト下での安定した適応のための損失平滑化
微調整や強化学習などの設定では、ニューラル ネットワークは分布シフトの下で適応されることがよくあります。標準的な適応方法は通常、ターゲット目標を直接最適化し、ソースのトレーニング目標からの突然の変更を引き起こします。この突然の移行により、新しいタスクに依然として役立つ可能性のある特徴を含む、学習された表現が歪む可能性があります。私たちは、より緩やかな移行によって適応が改善されるかどうかを調査します。私たちは、適応の開始時にソースとターゲットのトレーニング目標の間を補間する単純なアプローチである損失平滑化を提案します。このスムーズな移行により、ソース ディストリビューションの有用な機能を維持しながら、モデルをターゲット ディストリビューションに特化させることができます。制御された教師ありシフト、事前訓練された視覚適応、オフラインからオンラインおよびオンラインへの強化学習、および言語モデルの微調整にわたって、損失平滑化が一貫してパフォーマンスを向上させることがわかり、よりスムーズな目的の移行がモデル適応に広く役立つツールであることを示唆しています。
原文 (English)
Loss Smoothing for Stable Adaptation Under Distribution Shift
In settings such as fine-tuning and reinforcement learning, neural networks are often adapted under distribution shift. Standard adaptation methods typically optimize the target objective directly, inducing an abrupt change from the source training objective. This abrupt transition can distort learned representations, including features that may still be useful for the new task. We investigate whether a more gradual transition can improve adaptation. We propose loss smoothing, a simple approach that interpolates between the source and target training objectives at the start of adaptation. This smooth transition helps to preserve useful features from the source distribution while still enabling the model to specialize to the target distribution. Across controlled supervised shifts, pretrained vision adaptation, offline-to-online and online reinforcement learning, and language model fine-tuning, we find that loss smoothing consistently improves performance, suggesting that smoother objective transitions are a broadly useful tool for model adaptation.
定義による忠実: 自然な意味論的メタ言語の説明による感情分析
感情分類子の説明は通常、事後的に作成され、ラベルの背後にある計算を反映しているという保証はありません。イベントベースの感情分析のための説明インターフェイスを紹介します。パーサーは入力テキストを説明、つまり 12 個の型付きスロットに編成されたナチュラル セマンティック メタ言語のクローズド ボキャブラリの短いスクリプトにマッピングし、公開されたセマンティック定義から転記されたルールの固定決定リストが説明のみからラベルを計算します。したがって、忠実性の保証は因果関係と定義に基づくものですが、すべての経験的リスクは学習されたパーサー内に存在し、行ごとの含意インターフェイスによって入力に対して監査可能になります。クラウドソースのイベント説明では、微調整されたパーサーは、小規模なホールドアウト セットで 0.33 の精度と 0.48 の選択的精度に達しました。これは、インターフェイスがブラック ボックス モデルとのわずかな精度の差をトレードして、一人称イベントベースの感情分析の検証可能で検査可能な意思決定基準を提供していることを示唆しています。また、行ごとの検証メタデータと完全なルール セットを備えた EmoExpl-1200 もリリースします。
原文 (English)
Faithful by Definition: Emotion Analysis via Natural Semantic Metalanguage Explications
Explanations for emotion classifiers are usually produced post hoc, with no guarantee that they reflect the computation behind the label. We present an explication interface for event-based emotion analysis. A parser maps the input text to an explication, a short script in the closed vocabulary of Natural Semantic Metalanguage organized into twelve typed slots, and a fixed decision list of rules transcribed from published semantic definitions computes the label from the explication alone. The faithfulness guarantee is therefore causal and definitional, while all empirical risk lives in the learned parser, which the per-line entailment interface makes auditable against the input. On crowd-sourced event descriptions, our fine-tuned parser reaches 0.33 accuracy and 0.48 selective accuracy on a small held-out set, suggesting that the interface trades insignificant accuracy difference to a black-box model for a verifiable, inspectable decision basis for first-person event-based emotion analysis. We also release EmoExpl-1200 with per-line verification metadata and the full rule set.
ラベルの影響伝播によるマルチラベル ノードの分類
グラフは、さまざまなドメインにわたって使用される複雑で汎用性の高いデータ構造であり、おそらくマルチラベル ノードが特に重要な役割を果たします。例には、複数の機能を持つ PPI ネットワーク内のタンパク質や、多様な関心を示すソーシャル ネットワークや電子商取引ネットワーク内のユーザーが含まれます。グラフ上でマルチラベル ノード分類 (MLNC) に取り組むことで、さまざまなアプローチが開発されました。グラフ ニューラル ネットワーク (GNN) を利用してラベルの共起相関を利用する方法もあれば、ラベルの近接性を捕捉するためにラベルの埋め込みを組み込む方法もあります。ただし、これらのアプローチでは、非ユークリッド グラフ データのラベル間の複雑な影響を考慮することができません。この問題に対処するために、GNN のメッセージ受け渡しプロセスを 2 つの操作 (伝播と変換) に分解します。次に、各操作におけるラベル間の影響相関の包括的な分析と定量化を実行します。これらの洞察に基づいて、ラベル影響伝播 (LIP) という新しいモデルを提案します。具体的には、統合されたラベル相関に基づいてラベル影響グラフを構築します。次に、このグラフを通じて高次の影響を伝播し、プラスの寄与を持つラベルを増幅し、マイナスの影響を持つラベルを軽減することで学習プロセスを動的に調整します。最後に、当社のフレームワークは包括的なベンチマーク データセットで評価され、さまざまな設定にわたって一貫して SOTA 手法を上回っており、MLNC タスクでの有効性が実証されています。
原文 (English)
Multi-Label Node Classification with Label Influence Propagation
Graphs are a complex and versatile data structure used across various domains, with possibly multi-label nodes playing a particularly crucial role. Examples include proteins in PPI networks with multiple functions and users in social or e-commerce networks exhibiting diverse interests. Tackling multi-label node classification (MLNC) on graphs has led to the development of various approaches. Some methods leverage graph neural networks (GNNs) to exploit label co-occurrence correlations, while others incorporate label embeddings to capture label proximity. However, these approaches fail to account for the intricate influences between labels in non-Euclidean graph data. To address this issue, we decompose the message passing process in GNNs into two operations: propagation and transformation. We then conduct a comprehensive analysis and quantification of the influence correlations between labels in each operation. Building on these insights, we propose a novel model, Label Influence Propagation (LIP). Specifically, we construct a label influence graph based on the integrated label correlations. Then, we propagate high-order influences through this graph, dynamically adjusting the learning process by amplifying labels with positive contributions and mitigating those with negative influence. Finally, our framework is evaluated on comprehensive benchmark datasets, consistently outperforming SOTA methods across various settings, demonstrating its effectiveness on MLNC tasks.
LUMA: 軽量ユニバーサル マスク アダプターを使用したセグメンテーションのベンチマーク
画像セグメンテーションのためのトランスフォーマー バックボーンの比較は混乱を招きます。それぞれが異なるデコーダー、レシピ、事前トレーニングとペアになっているため、報告される違いがバックボーン自体を反映することはほとんどありません。軽量ユニバーサル マスク アダプター (LUMA) を導入します。これは軽量でバックボーンに依存しないマスク トランスフォーマー ヘッドであり、あらゆるバックボーンをブラック ボックスの特徴抽出器として扱い、安価なクロス アテンションを通じて一連のクエリがその特徴から読み取れるようにします。 LUMA は、低コストで最先端の効率的な ViT セグメンタである EoMT の精度に匹敵し、等方性、階層的、畳み込み的、および専門家の混合バックボーンに同様に変更せずに接続します。この頭を固定したまま、1 つの最新のレシピの下で、ADE20K と都市景観に関する 20 のバックボーン、11 の事前トレーニング スキーム、およびさまざまな解像度をベンチマークします。 「効率的な」トークンミキサーは、それを動機づける高解像度でも効率を実現できず、プレーンな ViT がすべての解像度でスループットのパレートフロントを維持していることがわかりました。さらに、アーキテクチャではなく、フィールドが最も熱心に調整したレバーである事前トレーニングの目的が、セグメンテーションの品質を左右します。
原文 (English)
LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter
Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so reported differences rarely reflect the backbone itself. We introduce the Lightweight Universal Mask Adapter (LUMA), a lightweight, backbone-agnostic mask-transformer head that treats any backbone as a black-box feature extractor, letting a set of queries read from its features through cheap cross-attention. LUMA matches the accuracy of EoMT, the state-of-the-art efficient ViT-segmenter, at lower cost, while attaching unchanged to isotropic, hierarchical, convolutional, and mixture-of-experts backbones alike. Holding this head fixed, we benchmark 20 backbones, 11 pretraining schemes and a range of resolutions on ADE20K and Cityscapes under one modern recipe. We find that ``efficient'' token mixers fail to deliver efficiency even at the high resolutions that motivate them, with plain ViT holding the throughput Pareto-front at every resolution. Additionally, the pretraining objective, not the architecture, the lever the field has tuned hardest, governs segmentation quality.
LLVM-Bench: LLVM コンパイラの問題解決のための大規模言語モデルのベンチマークと高度化
LLVM は広く使用されているコンパイラ インフラストラクチャであり、その規模と複雑さにより、問題解決に多大な労力と困難が伴います。大規模言語モデル (LLM) は最近、問題解決において目覚ましい成功を収めていますが、複雑なシステムレベルの LLVM コンパイラに対する LLM の有効性はほとんど解明されていません。このギャップに対処するために、LLVM 問題解決のための初の大規模ベンチマークである LLVM-Bench を導入しました。これには、LLVM プロジェクトから収集された 423 個の実際の検証済みタスクが含まれています。さらに、問題の再現、パッチの適用、コンパイラの構築、テストの実行を自動化するスケーラブルな評価プラットフォームである LLVM-Gym を開発します。 LLVM-Bench と LLVM-Gym を使用して、4 つの代表的な LLM、6 つの検索構成、および 3 つのエージェントの包括的な調査を実施します。私たちの結果は、現在の LLM ベースの問題解決手法が LLVM-Bench では依然として限定的であり、パッチの無効性とビルドの失敗が主な障害モードであることを示しています。さらに、さまざまな LLM とエージェント間の強力な相補性が明らかになり、LLVM-En の動機付けとなります。LLVM-En は、さまざまな技術によって生成されたパッチを統合することでパッチ空間を拡張し、不正確で冗長な候補をフィルタリングし、最も有望なソリューションを特定する軽量のアンサンブル アプローチです。私たちの結果は、LLVM-Ens が最大 21.99% の解決率を達成し、LLVM 問題の解決をさらに向上させることを示しています。
原文 (English)
LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution
LLVM is a widely used compiler infrastructure whose scale and complexity make issue resolution labor-intensive and challenging. Although large language models (LLMs) have recently achieved remarkable success in issue resolution, their effectiveness on complex system-level LLVM compiler remains largely unexplored. To address this gap, we introduce LLVM-Bench, the first large-scale benchmark for LLVM issue resolution, containing 423 real-world, validated tasks collected from the LLVM project. We further develop LLVM-Gym, a scalable evaluation platform that automates issue reproduction, patch application, compiler building, and test execution. Using LLVM-Bench and LLVM-Gym, we conduct a comprehensive study of four representative LLMs, six retrieval configurations, and three agents. Our results show that current LLM-based issue resolution techniques remain limited on LLVM-Bench, with patch invalidity and build failures as the dominant failure modes. We further reveal a strong complementarity among different LLMs and agents, motivating LLVM-Ens, a lightweight ensemble approach that expands the patch space through integrating the patches generated by diverse techniques, filters incorrect and redundant candidates, and identifies the most promising solution. Our results show that LLVM-Ens achieves a resolution rate of up to 21.99%, further improving LLVM issue resolution.
影響力のある自動運転データセットの作成: 研究ギャップからベンチマークまでの戦略ガイド
適切に設計された自動運転データセットは研究の進歩を根本的に形作ってきましたが、既存の文献では主に、影響力のあるデータセットを戦略的に設計する方法ではなく、データセットに含まれるものについて説明されています。これは、希少なリソースを誤って割り当てる余裕がない中小規模の研究室や新興企業にとって特に制限となります。私たちは、影響力のあるデータセットの作成は診断から始まり、研究課題がデータの問題によって妨げられているのか、それとも評価の問題によって妨げられているのかを診断し、結果として生じるギャップを埋める最小限のデータ演算子を選択することで進み、より安価な演算子が十分でない場合にのみ新しいデータを記録すると主張します。このレンズを通じて主要な自動運転 (AD) データセットの進化を分析し、ギャップの特定、オペレーターの選択、センサー スイートの設計、アノテーション戦略にわたる戦略的フレームワークを抽出します。私たちは、KITScenes データセット ファミリの実行中のケース スタディでフレームワークを基礎としました。データセットは https://kitscenes.com/ で入手できます。
原文 (English)
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a data problem or an evaluation problem, and proceeds by selecting the minimal data operator(s) that closes the resulting gap, recording new data only when no cheaper operator(s) suffices. We analyze the evolution of major autonomous driving (AD) datasets through this lens and distill a strategic framework spanning gap identification, operator choice, sensor suite design, and annotation strategy. We ground the framework in a running case study of our KITScenes dataset family. The datasets are available at: https://kitscenes.com/
固定小数点フローによる自己条件付きフロー マップ言語モデル
セルフコンディショニングは、連続フローベースの言語モデルを強化する中心的な手法であり、モデルは独自のノイズ除去推定値に基づいて条件付けすることによって、生成されたテキストのノイズ除去を学習します。経験的には成功していますが、パフォーマンスの向上はあまり理解されていません。さらに、セルフコンディショニングを活用する方法が不明瞭であるフロー マップに基づく数ステップ ジェネレーターの使用への関心が高まっています。ここでは、自己調整を備えたフロー言語モデルが、学習されたデノイザーのパフォーマンスをブートストラップする固定小数点反復を解決することを示します。この観点を使用して、自己調整フローの 2 次元クラスである固定点フローを定式化します。最初の次元はフロー プロセスを表し、2 番目の次元は固定小数点の反復を表します。固定点フローが有効なフロー マップを定義することを示し、固定点反復とフロー プロセスの両方 (前者は固定点蒸留で、後者はフロー マップ蒸留で) の両方を圧縮することによって、自己条件付きフロー モデルから固定点フロー マップを蒸留できることを示します。結果として得られるフロー マップ言語モデル FMLM$^\star$ は、OpenWebText での 1 ステップおよび数ステップの生成において、最先端の自己条件付きモデルや数ステップ モデルよりも優れたパフォーマンスを発揮します。コードは https://github.com/Ugness/self-conditioned-fmlm で入手できます。
原文 (English)
Self-conditioned Flow Map Language Models via Fixed-point Flows
Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text by conditioning on its own denoising estimate. While empirically successful, its performance improvements are poorly understood. Moreover, there is growing interest in the use of few-step generators based on flow maps, for which how to leverage self-conditioning is unclear. Here, we show that flow language models with self-conditioning solve a fixed-point iteration that bootstraps the performance of the learned denoiser. We use this viewpoint to formulate fixed-point flows, a two-dimensional class of self-conditioned flows, where the first dimension represents the flow process and the second represents the fixed-point iteration. We show that fixed-point flows define valid flow maps, and show that they can be distilled from self-conditioned flow models by compressing both fixed-point iterations and the flow process, the former with fixed-point distillation and the latter with flow map distillation. Our resulting flow map language model, FMLM$^\star$, outperforms state-of-the-art self-conditioned models and few-step models in one- and few-step generation on OpenWebText. Code is available at https://github.com/Ugness/self-conditioned-fmlm.
アクション認識のための部分的なスケルトンの可視化: 視野を制限したアプローチ
スケルトンベースのアクション認識は、関節座標とそのトポロジカルな接続を利用することで目覚ましい成功を収めていますが、一般的な手法は圧倒的に完全でクリーンなスケルトン入力を前提としています。自己中心的なビジョン、混雑した監視、ウェアラブル デバイス、エッジ ロボティクスなどの現実世界の展開では、限られた視野 (FoV) によって関節の可視性が大幅に低下することがよくあり、既存のモデルではほとんど対応する準備ができていない深刻なパフォーマンスの低下につながります。この重要な、しかし十分に検討されていないギャップを埋めるために、制約された FoV の下で堅牢なスケルトン アクション認識に合わせて調整された新しいハイパーグラフ フレームワークである PartialVisGraph を紹介します。まず、ソフト入射行列を形成する学習可能な仮想ハイパーエッジを導入し、従来のペアごとのグラフを超えた柔軟な高次の依存関係をキャプチャすることで、非常に表現力豊かなハイパーグラフを構築します。次に、事前可視性を明示的に組み込みながら、ジョイント フィーチャをハイパーエッジに適応的に集約するシングルヘッド サンプル適応トランスフォーマーを提案します。この事前設定により、情報の流れが選択的にゲートされ、関節が閉塞したり視界から外れたりして、信頼性の高い特徴の伝播が損なわれるのを防ぎます。さらに、NTU RGB+D 60 および 120 での現実的な FoV シミュレーション ベンチマークを使用した厳密な評価プロトコルを確立します。広範な実験により、PartialVisGraph が部分可視下で一貫して最先端の精度を達成し、最近の強力なベースラインと比較して厳しい FoV 制限のあるサブセットで最大 68.8\% のゲインを実現し、同時に完全可視設定でも優れていることが実証されました。私たちのアプローチは、制約のない環境で展開可能なスケルトンベースのアクションを理解するための原則に基づいた実践的な道筋を提供します。
原文 (English)
Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach
Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs. In real-world deployments, such as egocentric vision, crowded surveillance, wearable devices, or edge robotics, limited field-of-view (FoV) frequently causes substantial joint visibility dropout, leading to severe performance degradation that existing models are largely unprepared to handle. To bridge this critical yet underexplored gap, we introduce PartialVisGraph, a novel hypergraph framework tailored for robust skeleton action recognition under constrained FoV. We first construct highly expressive hypergraphs by introducing learnable virtual hyperedges that form a soft incidence matrix, capturing flexible high-order dependencies beyond conventional pairwise graphs. We then propose the Single-Head Sample-Adaptive Transformer, which adaptively aggregates joint features onto hyperedges while explicitly incorporating a visibility prior. This prior selectively gates information flow, preventing occluded or out-of-view joints from corrupting reliable feature propagation. We further establish rigorous evaluation protocols with realistic FoV simulation benchmarks on NTU RGB+D 60 and 120. Extensive experiments demonstrate that PartialVisGraph consistently achieves state-of-the-art accuracy under partial visibility, with gains of up to 68.8\% on subsets with severe FoV restrictions compared to recent strong baselines, while remaining superior on full-visibility settings. Our approach offers a principled and practical pathway toward deployable skeleton-based action understanding in unconstrained environments.
検出不可能なものの検出: アクティブ ラーニングによる教師なし時系列異常検出の強化
産業用 AI システムがますます洗練されているにもかかわらず、複雑な時系列データ内の微妙でノイズの多い異常を確実に検出する機能は、依然として重要な未解決の課題のままです。大規模な産業アプリケーションでは、時系列データのラベル付けは法外な費用と時間がかかることが多いため、教師なし学習が実用的で広く採用されているアプローチとなっています。しかし、既存の教師なし手法は、正常に近い異常を正常なパターンから区別するのに苦労することが多く、正常なサンプル内のノイズ汚染に対して脆弱です。これらの制限に対処するために、アクティブ ラーニングを活用して教師なしモデルのパフォーマンスを反復的に向上させる新しいフレームワークを提案します。私たちのフレームワークの核となる貢献は、(1) モデルにロバストな時間依存関係を学習させるマスクされた時系列再構成フィードバック戦略、(2) 正常サンプルと異常サンプルを区別して処理することでロバスト性を促進するミニマックス学習戦略です。このプロセスにより、モデルは微妙でノイズの多いパターンのダイナミクスをより適切に捉えることができます。提案されたフレームワークは、4 つの多変量時系列データセットと 7 つの教師なしバックボーン モデルを含む 28 のテスト ケースにわたって評価されます。実験結果は、元のモデルと比較して AUC が 12.39% 向上したことを示しており、私たちの方法を既存の教師なし再構成ベースの異常検出システムに容易に統合して、パフォーマンスを大幅に向上できることが確認されました。
原文 (English)
Detecting the Undetectable: Enhancing Unsupervised time series Anomaly Detection via Active Learning
Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time series data remains a critical yet unresolved challenge. In large-scale industrial applications, labeling time series data is often prohibitively expensive and time-consuming, making unsupervised learning a practical and widely adopted approach. However, existing unsupervised methods frequently struggle to distinguish near-normal anomalies from normal patterns and are vulnerable to noise contamination within normal samples. To address these limitations, we propose a novel framework that leverages active learning to iteratively enhance the performance of unsupervised models. Our framework's core contributions are (1) a masked time-series reconstruction feedback strategy that forces the model to learn robust temporal dependencies, and (2) a minimax learning strategy that promotes robustness by differentially treating normal and abnormal samples. This process encourages the model to better capture the dynamics of subtle and noisy patterns. The proposed framework is evaluated across 28 test cases involving four multivariate time-series datasets and seven unsupervised backbone models. Experimental results demonstrate a 12.39% improvement in AUC compared to the original models, confirming that our method can be readily integrated into existing unsupervised reconstruction-based anomaly detection systems to significantly enhance their performance.
LLM ガイド付き ODE 検出と小規模コホート集計データからのパラメーター推論
常微分方程式 (ODE) による機構モデリングは、複雑な力学の解釈可能な記述を提供し、基礎となる機構の推論を可能にします。これは臨床現場で特に価値があります。ただし、希少疾患では、モデルの構造とパラメーターの両方が通常不明ですが、個人レベルのデータは不足しており、ノイズが多く、不均一であり、プライバシーの制約を受けます。このような設定では、母集団レベルの要約統計はプライバシーを保護する実用的なデータ表現を提供しますが、不均一性を捉えるには固定値ではなく分布としてパラメータをモデリングする必要がさらにあります。しかし、既存の方法では、ODE 構造を共同で発見し、要約統計だけからパラメータ分布を改良することはできません。このギャップに対処するエンドツーエンドのフレームワークである AgentODE を紹介します。 LLM は ODE 構造の候補を提案しますが、ツール拡張推論エージェントは、母集団レベルの要約統計量のみを操作して、診断 - 更新ループを通じてパラメーター分布を繰り返し調整します。私たちは、異なる分野にわたる 3 つのベンチマーク問題と、希少疾患である劣性遺伝性ジストロフィー性表皮水疱症 (RDEB) を含む 2 つの臨床データセットに基づいて AgentODE を評価しましたが、46 人の患者でわずか 231 件の観察しかありませんでした。 AgentODE は、すべての設定にわたって機能的に一貫した ODE 構造を回復します。また、RDEB での実験では、まばらでノイズの多いデータ設定では、概要統計から推論することで、機械的に原則に基づいた構造発見が促進されるのに対し、個別レベルのデータ アクセスによるベースラインでは、より優れた予測パフォーマンスにもかかわらず、ありそうもない構造が回復されることが実証されています。 AgentODE は、従来、データ不足とプライバシーの制約によってそのような分析が制限されてきた、人口レベルの概要統計から直接希少疾患のメカニズムをモデリングする新たな可能性を開きます。
原文 (English)
LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data
Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inference of underlying mechanisms, which is particularly valuable in clinical settings. However, in rare diseases, both the structure and parameters of the model are typically unknown, while individual-level data is scarce, noisy, heterogeneous, and subject to privacy constraints. In such settings, population-level summary statistics provide a practical privacy-preserving data representation, while capturing heterogeneity further requires modeling parameters as distributions rather than fixed values. Yet no existing method jointly discovers ODE structure and refines parameter distributions solely from summary statistics. We present AgentODE, an end-to-end framework that addresses this gap. An LLM proposes candidate ODE structures, while a tool-augmented inference agent iteratively refines parameter distributions through a diagnosis--update loop, operating on population-level summary statistics alone. We evaluate AgentODE on three benchmark problems across different fields and two clinical datasets, including the rare disease recessive dystrophic epidermolysis bullosa (RDEB), with only 231 observations across 46 patients. AgentODE recovers functionally consistent ODE structures across all settings, and experiments on RDEB demonstrates that in sparse and noisy data settings reasoning from summary statistics promotes mechanistically principled structure discovery, whereas baselines with individual-level data access recover implausible structures despite better predictive performance. AgentODE opens new possibilities for mechanistic modeling of rare diseases directly from population-level summary statistics, where data scarcity and privacy constraints have traditionally limited such analyses.
ConRTF: リアルタイムのトランスフォーマ テーブル構造認識のためのエッジ制約境界分布の改良
表構造認識 (TSR) は、文書理解パイプラインの重要なステップである文書画像から表の行と列のレイアウトを復元することを目的としています。正確な TSR は、境界の正確な位置特定に依存します。行または列の境界における小さなエラーが、不正確なセル割り当てや構造的不一致に伝播する可能性があります。しかし、検出ベースのアプローチは、テーブル要素を汎用オブジェクトとして扱い、テーブル レイアウトの基本的な特性を無視します。つまり、行と列は構造的に異なる役割を果たし、それらの境界の重要性は同等ではありません。我々は、テーブル固有の幾何学的事前分布をトレーニング目標にエンコードすることによって、この構造的非対称性を形式化する、エッジ制約付きファイングレインローカリゼーションロス(EFL)を提案します。行状要素は水平方向の境界を強調して監視され、列状の要素は垂直方向の境界を優先します。分布ベースの境界リファインメント (D-FINE) を備えたリアルタイム検出器内に実装された EFL は、トレーニング中にのみ動作し、推論パイプラインを変更することなく、境界リファインメントを構造的に意味のある調整に導きます。提案されたアプローチである ConRTF もデータ効率が高く、わずか 2k ~ 3k の注釈付きテーブルで堅牢な精度を維持します。 PubTables-1M と 2 つのプライベート データセットでの実験では、最適化されたベースラインと、RT-DETRv2 や YOLOv10-11 を含むいくつかのリアルタイム検出器を超える一貫した改善が示され、等しい推論速度で最大 +1.6 GriTS ポイントのゲインが得られます。
原文 (English)
ConRTF: Edge-Constrained Boundary Distribution Refinement for Realtime TransFormer Table Structure Recognition
Table Structure Recognition (TSR) aims to recover the row and column layout of tables from document images, a key step in document understanding pipelines. Accurate TSR depends on precise boundary localization: small errors in row or column boundaries can propagate into incorrect cell assignments and structural inconsistencies. Yet detection-based approaches treat table elements as generic objects, ignoring a fundamental property of table layout: rows and columns play structurally distinct roles and their boundaries carry unequal importance. We propose an Edge-constrained Fine-grained Localization loss (EFL) that formalizes this structural asymmetry by encoding table-specific geometric priors into the training objective: row-like elements are supervised with emphasis on their horizontal boundaries, while column-like elements prioritize vertical boundaries. Implemented within a real-time detector with distribution-based boundary refinement (D-FINE), EFL operates during training only and guides boundary refinement toward structurally meaningful adjustments with no change to the inference pipeline. The proposed approach, ConRTF, is also data-efficient, maintaining robust accuracy with as few as 2k--3k annotated tables. Experiments on PubTables-1M and two private datasets show consistent improvements over the optimized baseline and several real-time detectors including RT-DETRv2 and YOLOv10-11, with gains of up to +1.6 GriTS points at equal inference speed.
幻の文献: 一流の学会での査読に生き残る幻覚引用
大規模な言語モデルでは、裏付けのない主張を含む洗練された科学文書が生成され、幻覚がアーカイブ記録に残る可能性があります。技術的な記述によってこのリスクを評価することは難しく、多くの場合専門家の判断が必要ですが、引用はより監査可能な表面を提供します。つまり、参照は、互換性のある著者を持つ実際の学術著作物であるか、そうでないかのどちらかです。私たちは、査読手続きにおける引用幻覚を、存在しない著作物や著者と著者リストの大幅な不一致など、アイデンティティレベルの失敗に限定した保守的な定義を使用して測定します。私たちは、通常の書誌的差異(例: 開催地/年の違い、出版状況の更新、マイナーな名前のバリエーションなど)を明示的に除外します。引用を大規模に監査するために、複数の書誌ソースに対して書誌エントリを解決し、未解決のケースを Web 検索の再検証にエスカレートする検証パイプラインである RefChecker を構築します。 ICLR、ICML、NeurIPS、および USENIX Security から承認されたカメラ対応ペーパーに RefChecker を適用します。幻覚を起こした引用がアーカイブ記録に残っています。参考文献レベルの割合は通常 1% 未満ですが、論文レベルの失敗が目に見えるほど議事録は大きくなっています。2025 年には、NeurIPS および USENIX セキュリティ論文のおよそ 20 件に 1 件に、私たちの厳密な定義に基づくと、幻覚を起こしている可能性のある学術論文のような参考文献が少なくとも 2 件含まれています。また、ChatGPT 後の増加もいくつかの会場で観察されており、これには、単一の参考文献で 5 件以上の失敗を伴う論文の尾部や、受賞論文の中でも幻覚引用の可能性が高いことが含まれます。これらの結果は、査読だけでは引用の完全性を確実に強制することはできないが、監査は扱いやすい(会場規模の 1 回のスキャンで論文あたり約 0.04 ドル)ことを示唆しています。私たちは、出版前に日常的で再現可能な引用検証を行うための RefChecker をオープンソースにしています (https://github.com/markrussinovich/refchecker)。
原文 (English)
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival record. Assessing this risk via technical statements is difficult and often requires expert judgment, but citations provide a more auditable surface: a reference either resolves to a real scholarly work with compatible authorship, or it does not. We measure citation hallucination in peer-reviewed proceedings using a conservative definition limited to identity-level failures: non-existent works and substantial author-list mismatches. We explicitly exclude ordinary bibliographic drift (e.g., venue/year differences, publication-status updates, minor name variants). To audit citations at scale, we build RefChecker, a verification pipeline that resolves bibliography entries against multiple bibliographic sources and escalates unresolved cases to web-search re-verification. We apply RefChecker to accepted camera-ready papers from ICLR, ICML, NeurIPS, and USENIX Security. Hallucinated citations have entered the archival record. While reference-level rates are usually below 1%, proceedings are large enough that paper-level failures are visible: in 2025, roughly one in twenty NeurIPS and USENIX Security papers contains at least two likely hallucinated academic-paper-like references under our strict definition. We also observe post-ChatGPT increases in several venues, including a tail of papers with 5+ failures in a single bibliography, and likely hallucinated citations even among award-winning papers. These results suggest peer review alone does not reliably enforce citation integrity, yet auditing is tractable (about 0.04$ per paper in one venue-scale scan). We open-source RefChecker for routine, reproducible citation verification before publication (https://github.com/markrussinovich/refchecker).
出生前超音波検査における記憶に基づくトレーニング不要の異常分類と位置特定のプロトタイプ
出生前異常の分類と位置特定は、胎児の健康と妊娠管理にとって非常に重要です。超音波(US)は出生前スクリーニングの主な手段ですが、異常の有病率が低く不均一性が高いため、正確な診断は依然として困難です。出生前タスクのための既存の深層学習手法は、実際には入手が困難な大規模な注釈付きデータセットに依存しています。少数ショット学習はデータ不足を軽減しますが、通常は新しいカテゴリに合わせて微調整する必要があるため、リソースが限られた臨床現場では実用性が制限されます。これらの課題に対処するために、クラスごとに数枚の参照画像のみで動作する、マルチクラス出生前米国異常の分類と位置特定のためのトレーニング不要のフレームワークを提案します。これは、この設定の最初の探索を表します。私たちのフレームワークは 3 つの重要なコンポーネントで構成されています。(1) クラスレベルのセマンティクスと異常特性の両方を明示的にモデル化する、複数粒度のプロトタイプを備えたメモリ バンク。 (2)異常領域を検出するために識別特徴を集約するプロトタイプ駆動のソフトマージメカニズム。 (3) プロトタイプの一貫性を活用してカテゴリの予測を改善する、クラスを意識した改良戦略。 1,149 例を含む、合計 2,357 枚の画像と 9 つのカテゴリを含む多施設の米国出生前データセットで広範囲に検証され、私たちの提案された方法は競合他社を上回ります。
原文 (English)
Prototype Memory-Guided Training-Free Anomaly Classification and Localization in Prenatal Ultrasound
Prenatal anomaly classification and localization is of critical importance for fetal health and pregnancy management. Although ultrasound (US) is the primary modality for prenatal screening, accurate diagnosis remains challenging due to the low prevalence and high heterogeneity of anomalies. Existing deep learning methods for prenatal tasks rely on large-scale annotated datasets, which are difficult to obtain in practice. Although few-shot learning alleviates data scarcity, it typically requires fine-tuning for new categories, limiting its practicality in resource-limited clinical settings. To address these challenges, we propose a training-free framework for multi-class prenatal US anomaly classification and localization that operates with only a few reference images per class, representing the first exploration of this setting. Our framework comprises three key components: (1) a memory bank with multi-granular prototypes that explicitly models both class-level semantics and anomaly characteristics; (2) a prototype-driven soft merging mechanism that aggregates discriminative features to detect the anomaly region; and (3) a class-aware refinement strategy that leverages prototype consistency to improve category prediction. Extensively validated on a multi-center prenatal US dataset containing 1,149 cases, with a total of 2,357 images and 9 categories, our proposed method outperforms the competitors.
GaussianFusion: マルチモーダル融合知覚のための統合 3D ガウス表現
鳥瞰図 (BEV) 表現により、マルチセンサー機能を統一空間内で融合することができ、包括的な 3D 認識を実現するための主要なアプローチとして機能します。ただし、BEV の離散グリッド表現は、大幅なディテールの損失につながり、マルチモーダル フュージョン知覚における特徴の位置合わせとクロスモーダル情報の相互作用を制限します。この研究では、従来の BEV パラダイムから脱却し、3D ガウス表現に基づいたマルチモーダル融合のための新しい普遍的なフレームワークを提案します。このアプローチでは、共有された連続 3D ガウス空間内でマルチモーダル フィーチャが自然に統合され、エッジや細かいテクスチャの詳細が効果的に保存されます。これを達成するために、新しい前方投影ベースのマルチモーダル ガウス初期化モジュールと、アテンション メカニズムに基づいてガウス プロパティを繰り返し更新する共有クロスモーダル ガウス エンコーダーを設計します。 GaussianFusion は本質的にタスクに依存しないモデルであり、統一されたガウス表現によりさまざまな 3D 認識タスクが自然にサポートされます。広範な実験により、GaussianFusion の汎用性と堅牢性が実証されました。 nuScenes データセットでは、3D オブジェクト検出ベースライン BEVFusion を 2.6 NDS 上回っています。そのバリアントは、ガウス分布の 30% のみを使用し、450% の高速化を達成しながら、3D セマンティック占有率で GaussFormer を上回り、1.55 mIoU 改善しています。
原文 (English)
GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception
The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment and cross-modal information interaction in multimodal fusion perception. In this work, we break from the conventional BEV paradigm and propose a new universal framework for multi-modal fusion based on 3D Gaussian representation. This approach naturally unifies multi-modal features within a shared and continuous 3D Gaussian space, effectively preserving edge and fine texture details. To achieve this, we design a novel forward-projection-based multi-modal Gaussian initialization module and a shared cross-modal Gaussian encoder that iteratively updates Gaussian properties based on an attention mechanism. GaussianFusion is inherently a task-agnostic model, with its unified Gaussian representation naturally supporting various 3D perception tasks. Extensive experiments demonstrate the generality and robustness of GaussianFusion. On the nuScenes dataset, it outperforms the 3D object detection baseline BEVFusion by 2.6 NDS. Its variant surpasses GaussFormer on 3D semantic occupancy with 1.55 mIoU improvement while using only 30% of the Gaussians and achieving a 450% speedup.
カスケードされたオブジェクト検出のためのアクティブ ラーニング: テーブル抽出パイプラインにおけるカバレッジと不確実性のバランスをとる
ビジネス ドキュメントからのテーブルの抽出はカスケード パイプラインに依存しており、最初にテーブル検出 (TD) がテーブルをローカライズし、次にテーブル構造認識 (TSR) がその内部レイアウトを復元します。このパイプライン用のタスク固有のトレーニング セットの構築は、特にきめの細かい構造アノテーションが必要な TSR の場合、コストがかかります。アクティブ ラーニング (AL) はこのアノテーションの負担を軽減できますが、ほとんどの AL 戦略は単一モデルのタスク用に設計されており、カスケード アーキテクチャにおけるステージ間の依存関係を考慮していません。この研究では、もともと画像分類のために提案されたハイブリッド カバレッジ不確実性サンプリング手法である Uncertainty Herding (UHerding) を、カスケードされたオブジェクト検出パイプラインに初めて適応させたものを紹介します。我々は、TD から TSR への依存関係を利用する 2 つのパイプライン対応拡張機能を提案します。RankFusion は、検出空間と構造表現空間の両方にデュアル多様体カバレッジを追加します。一方、CAPA には、ステージ依存のゲートとタスクごとの不確かさのキャリブレーションがさらに組み込まれています。 2 つのパブリック (PubTables-1M および FinTabNet) と 2 つのプライベート テーブル抽出データセットにわたる広範な実験では、さまざまなアノテーション バジェット (71 ~ 500 ドキュメント) を使用して、UHerding がテーブル抽出まで適切に一般化し、各ベースラインを上回るパフォーマンスを示していることがわかりました。パイプライン対応バリアントの中で、RankFusion はより高い期待利益を達成しますが、より大きな変動を犠牲にします。一方、CAPA は最も一貫した戦略として浮上し、4 つのデータセットのうち 3 つで標準の UHerding を上回ります。
原文 (English)
Active Learning for Cascaded Object Detection: Balancing Coverage and Uncertainty in Table Extraction Pipelines
Table extraction from business documents relies on a cascaded pipeline where Table Detection (TD) first localizes tables and Table Structure Recognition (TSR) then recovers their internal layout. Building task-specific training sets for this pipeline is costly, particularly for TSR which requires fine-grained structural annotations. Active learning (AL) can reduce this annotation burden, yet most AL strategies are designed for single-model tasks and do not account for inter-stage dependencies in cascaded architectures. In this work, we present the first adaptation of Uncertainty Herding (UHerding), a hybrid coverage-uncertainty sampling method originally proposed for image classification, to cascaded object detection pipelines. We propose two pipeline-aware extensions that exploit the TD-to-TSR dependency: RankFusion adds dual-manifold coverage over both detection and structure representation spaces, while CAPA further incorporates stage-dependent gating and per-task uncertainty calibration. Extensive experiments across two public (PubTables-1M and FinTabNet) and two private table extraction datasets, with various annotation budgets (from 71 to 500 documents) show that UHerding generalizes well to table extraction, outperforming each baseline. Among pipeline-aware variants, RankFusion achieves higher expected gains but at the cost of greater variance, while CAPA emerges as the most consistent strategy, outperforming standard UHerding on three out of four datasets.
LeVLJEPA: ネガティブなしのエンドツーエンドの視覚言語事前トレーニング
視覚言語の事前訓練は依然として対照的な目標によって支配されていますが、視覚のみの自己教師あり学習では主に非対照的な方法が採用されています。同時に、ビジョン言語エンコーダの役割も変化しました。エンコーダは、ゼロショット分類器としてではなく、単一のプールされたエンベディングではなくパッチ トークンのフル グリッドを消費する、ビジョン言語モデルと高密度予測システムの凍結されたビジュアル バックボーンとして導入されることが増えています。私たちは、初の完全に非対照的なエンドツーエンドの視覚言語事前トレーニング方法である LeVLJEPA を紹介します。 LeVLJEPA は、ネガ、温度、運動量エンコーダー、教師と生徒のスケジュールを使用せずに、停止勾配ターゲットとモダリティごとの分布正則化を使用したクロスモーダル予測を通じて学習し、大規模で安定したトレーニングを行います。結果として得られたエンコーダーは、ダウンストリームでの使用に対して著しく強力な高密度セマンティック機能を提供することがわかりました。凍結されたビジョン言語モデルのバックボーンとして、LeVLJEPA は、2 つの異なる言語モデルの下で GQA、VQAv2、および POPE にわたって評価されたエンコーダーの中で最も強力であり、セマンティック セグメンテーションでは対照的なベースラインを上回り、線形プローブなどのグローバルな読み取りでは同等の性能を維持します。これらの結果は、高密度の意味論的視覚特徴を生成する効果的な手段として、非対照的な事前トレーニングを確立します。
原文 (English)
LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives
Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders has shifted: they are increasingly deployed not as zero-shot classifiers but as the frozen visual backbone of vision-language models and dense prediction systems, which consume the full grid of patch tokens rather than a single pooled embedding. We introduce LeVLJEPA, the first fully non-contrastive end-to-end vision-language pretraining method. LeVLJEPA learns through cross-modal prediction with stop-gradient targets and per-modality distributional regularization, without negatives, temperature, momentum encoder, or teacher-student schedule, and trains stably at large scale. We find that the resulting encoder provides markedly stronger dense semantic features for downstream use: as a frozen vision-language-model backbone, LeVLJEPA is the strongest of the evaluated encoders across GQA, VQAv2, and POPE under two distinct language models, and outperforms contrastive baselines on semantic segmentation, while remaining on par on global readouts such as linear probing. These results establish non-contrastive pretraining as an effective means of producing dense semantic vision features.
LRAT-Catcher: リフレクションによる SAT ソルバー証明書の Lean4 へのインポート
SAT ソルバーは、インタラクティブな定理証明者の範囲を超えた組み合わせ問題を解決し、独立した検証用の LRAT 証明書を生成します。我々は、DIMACS 式を LRAT 証明書とともに定理として Lean 4 にインポートするスタンドアロンの汎用ツールである LRAT-Catcher を紹介します。 LRAT-Catcher は、リーン コアからリフレクションを介してコンパイルされたネイティブ コードとして正式に検証された LRAT チェッカーを実行します。これは、Mathlib の明示的な証明項インポートによってメモリが使い果たされるインスタンスに合わせて拡張されます。 LRAT-Catcher は、キューブ アンド コンカー解決の実行も完全に Lean 内で構成します。立方体ごとの反駁は、それ自体が LRAT 証明であるカバー完全性証明書と結合されて、単一の不充足性定理になります。検証されたエンコーディングは、CNF レベルの結果を元の組み合わせ問題に結び付けます。シュール数 S(4) = 44 とラムゼイ数 R(4,4) = 18 をリーン定理として確立する際に、Mathlib の証明項インポートと外部チェッカー Cake_lpr に対してツールを評価します。
原文 (English)
LRAT-Catcher: Importing SAT Solver Certificates into Lean4 by Reflection
SAT solvers settle combinatorial problems beyond the reach of interactive theorem provers and produce LRAT certificates for independent verification. We present LRAT-Catcher, a standalone, general-purpose tool that imports a DIMACS formula together with an LRAT certificate into Lean 4 as a theorem. LRAT-Catcher runs the formally verified LRAT checker from Lean core as compiled native code via reflection. This scales to instances where Mathlib's explicit proof-term import exhausts memory. LRAT-Catcher also composes cube-and-conquer solving runs entirely inside Lean. Per-cube refutations are combined with a cover-completeness certificate, itself an LRAT proof, into a single unsatisfiability theorem. Verified encodings connect CNF-level results to the original combinatorial problems. We evaluate the tool against Mathlib's proof-term import and the external checker cake_lpr on establishing the Schur number S(4) = 44 and the Ramsey number R(4,4) = 18 as Lean theorems.
エージェントティック データ システムのセマンティック ギャップを探る: 分析ワークフローにおける運用上の失敗に関する形成的研究
大規模言語モデル (LLM) は、クエリの生成、ツールの呼び出し、分析ワークフローの構築にますます使用されています。最近の進歩により、ワークフローの生成と実行は大幅に改善されましたが、分析概念を運用するために必要な意味情報は、データベース スキーマやデータ値で明示的に表現されるものを超えていることがよくあります。エージェント生成の分析ワークフローにおける運用化の失敗に関するクロスドメインの形成的研究を紹介します。財務、人事、公共安全の領域にわたる 236 の分析意図にわたって、ワークフローの生成と実行が成功したにもかかわらず、繰り返し発生する 153 件の失敗を特定しました。私たちの分析により、比較根拠付け、プロセス推論、定量的推論、役割の混乱、ポリシー根拠という 5 つの反復的な失敗クラスが明らかになりました。これらの発見は、ユーザーレベルの分析概念とワークフロー生成システムで利用できる情報の間に意味的なギャップがあることを示唆しています。より広範に、彼らは分析操作の許容性について疑問を提起し、将来のエージェントデータシステムには分析意図と実行可能な計算の間のギャップを埋めるためにより豊富なセマンティック表現が必要になる可能性があることを示唆しています。
原文 (English)
Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows
Large language models (LLMs) are increasingly used to generate queries, invoke tools, and construct analytical workflows. Although recent advances have substantially improved workflow generation and execution, the semantic information required to operationalize analytical concepts often lies beyond what is explicitly represented in database schemas and data values. We present a cross-domain formative study of operationalization failures in agent-generated analytical workflows. Across 236 analytical intents spanning finance, human resources, and public safety domains, we identify 153 recurring failures despite successful workflow generation and execution. Our analysis reveals five recurring classes of failures: comparative grounding, process reasoning, quantitative reasoning, role confusion, and policy grounding. These findings suggest a semantic gap between user-level analytical concepts and the information available to workflow-generation systems. More broadly, they raise questions about the admissibility of analytical operations and suggest that future agentic data systems may require richer semantic representations to bridge the gap between analytical intent and executable computation.
Pano2World: 統合されたマルチビュー シーケンスによるエンドツーエンドの 3D 生成
単一のパノラマは、1 つのカメラ中心から視覚領域全体をキャプチャしますが、ユーザーは実際のシーンの探索を可能にすることなく、その場で見回すことに限定されます。単一のパノラマを、自由視点のナビゲーションのために永続的でレンダリング可能な 3D 表現に変換することへの関心が高まっています。既存の手法では、修復結果を伝播して基礎となるジオメトリを更新するビューごとの反復補完を採用するか、漸進的エラーの蓄積と煩雑なマルチステップ パイプラインを引き起こすか、ビデオ生成モデルの時間的一貫性事前分布を利用するかのいずれかですが、そのようなモデルに固有の連続軌道制約により、複数の方向からシーンを同時にカバーする際の柔軟性が制限されます。私たちは、単一の屋内パノラマを入力として受け取り、永続的で探索可能な 3D ガウス シーンを直接出力する Pano2World を紹介します。ソース パノラマが与えられると、Pano2World はまず粗い 3D ガウス プロキシを再構築し、適応的にサンプリングされた近くのポーズでレンダリングして、幾何学的に位置合わせされたガイダンス パノラマを取得します。次に、パノラマ拡散モデルは、ビューアウェア アテンション ルーティングを介してすべてのターゲット ビューを共同でノイズ除去します。各ターゲット ビューは、対応するガイダンス パノラマからの幾何学的制約と、ソース パノラマからのグローバル セマンティック ガイダンスを同時に受け取り、ビュー間の一貫性を自然に強化します。結合ノイズ除去中に形成されたマルチビューの隠れた特徴を VAE を介してピクセル ドメインにデコードすることで発生する情報損失を回避するために、これらの隠れた特徴をシーン潜在に直接蒸留し、その後最終的な 3D ガウス シーンにデコードするジオメトリ対応ブリッジ モジュールである潜在特徴アダプターを導入します。実験では、Pano2World が、マルチポジション パノラマ ノベルビュー合成ベンチマークにおいて、既存の方法よりも大幅に優れていることが実証されています。
原文 (English)
Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences
A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling true scene exploration. Converting a single panorama into a persistent, renderable 3D representation for free-viewpoint navigation has attracted growing interest; existing methods either adopt iterative per-view completion that propagates inpainting results to update the underlying geometry, leading to progressive error accumulation and cumbersome multi-step pipelines, or leverage the temporal consistency priors of video generation models, yet the continuous-trajectory constraint intrinsic to such models limits their flexibility in covering scenes from multiple directions simultaneously. We present Pano2World, which takes a single indoor panorama as input and directly outputs a persistent, explorable 3D Gaussian scene. Given the source panorama, Pano2World first reconstructs a coarse 3D Gaussian proxy and renders it at adaptively sampled nearby poses to obtain geometrically aligned guidance panoramas; a panoramic diffusion model then jointly denoises all target views via View-Aware Attention Routing, where each target view simultaneously receives geometric constraints from its corresponding guidance panorama and global semantic guidance from the source panorama, naturally enforcing cross-view consistency. To avoid the information loss incurred by decoding the multi-view hidden features formed during joint denoising back to the pixel domain via VAE, we introduce Latent Feature Adapter, a geometry-aware bridge module that directly distills these hidden features into a scene latent, subsequently decoded into the final 3D Gaussian scene. Experiments demonstrate that Pano2World significantly outperforms existing methods on the multi-position panoramic novel-view synthesis benchmark.
ワールド モデルからワールド アクション モデルへ: ロボット工学の簡潔なチュートリアル
世界モデルは、身体化されたインテリジェンスや生成シミュレーションでますます使用されていますが、その範囲はコミュニティ全体で依然として曖昧です。このチュートリアルでは、タスク関連の観測または状態の将来の進化を推定するアクション条件付き予測モデルとしてのワールド モデルの設計空間ビューを示します。私たちは既存の手法を観測空間世界モデルと状態空間世界モデルに分類し、視覚的な忠実度、空間構造、物理的解釈可能性、制御の使いやすさにおけるトレードオフを比較します。さらに、予測された未来と実行可能なロボットのアクションを結び付ける世界アクション モデルを紹介し、想像してから実行する、ビデオ機能条件付きアクション予測、ビデオとアクションの共同モデリング、およびポリシー学習のための補助ビデオ予測という 4 つの代表的なパラダイムを要約します。このチュートリアルの目標は、世界 (アクション) モデルの概念的範囲を明確にし、具体化された予測と制御のための構造化された分類法を提供することです。
原文 (English)
From World Models to World Action Models: A Concise Tutorial for Robotics
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities. This tutorial presents a design-space view of world models as action-conditioned predictive models that estimate the future evolution of task-relevant observations or states. We categorize existing methods into observation-space and state-space world models, comparing their trade-offs in visual fidelity, spatial structure, physical interpretability, and control usability. We further introduce world action models, which connect predicted futures with executable robot actions, and summarize four representative paradigms: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and auxiliary video prediction for policy learning. The goal of this tutorial is to clarify the conceptual scope of world (action) models and provide a structured taxonomy for embodied prediction and control.
隠された状態からの入力テキストの回復: デコーダ専用言語モデルの勾配ベースの反転の研究
この研究では、隠れ状態反転問題、つまりデコーダ専用言語モデルの元の入力トークン シーケンスを最終層の隠れ状態から回復する問題を研究します。私たちは反転をワンショットの再構成として扱うのではなく、連続的な埋め込み空間の最適化として研究します。この最適化では、検索中にハード トークンの投影を行わずにソフト プロキシがリークされたターゲットに向かって駆動され、トークンは内部ループの最後で 1 回だけコミットされます。この設計の選択には 2 つの結果が生じますが、これがこのホワイト ペーパーの主な焦点です。まず、最適化を完全に連続空間に保持すると、グラウンドトゥルース トークンのランク軌跡、ポジションごとの損失曲線、コミット時に測定された離散損失など、豊富な内部信号セットが公開されます。第 2 に、離散損失により、累積的な離散損失による回復の正確さを評価できます。どのトークンが再構成を破るかをさらに分析し、鋭いカテゴリカルな非対称性を発見しました。埋め込み行列の密な領域にあるスペース接頭辞の高頻度の機能語が失敗の大半を占めますが、コンテンツを含むトークンはほぼ完全に復元されます。 10 トークンの C4 プロンプトでは、候補ウィンドウが広がるにつれて完全一致率が 66.9% から 97.5% (平均類似度 0.994) に上昇し、ほとんどのエラーが真のあいまいさではなく回復可能なニアミスであることが確認されます。リリースされた SIPIT リファレンスと比較すると、次のような結果が得られます。ステップごとのハード プロジェクションは高速ですが、継続的な定式化によって最適化が観察可能になり、その失敗が検出可能になります。結果は、GPT-2 の最終層の隠れ状態が元のテキストと同じくらい敏感であることを示しています。
原文 (English)
Recovering Input Text from Hidden States: Study of Gradient-Based Inversion of Decoder-Only Language Models
This work studies the hidden-state inversion problem: recovering the original input token sequence of a decoder-only language model from its last-layer hidden states. Rather than treating inversion as a one-shot reconstruction, we study it as a continuous embedding-space optimisation in which a soft proxy is driven towards the leaked target without any hard-token projection during the search, and a token is committed only once, at the end of the inner loop. This design choice has two consequences which are the main focus of this paper. First, keeping the optimisation entirely in continuous space exposes a rich set of internal signals: rank trajectories of the ground-truth token, per-position loss curves, and a discrete loss measured at commit time. Second, the discrete loss allows assessing the correctness of recovery via cumulative discrete loss. We further analyse which tokens break the reconstructions and find a sharp categorical asymmetry: space-prefixed, high-frequency function words in dense regions of the embedding matrix dominate the failures, while content-bearing tokens are recovered almost perfectly. On 10-token C4 prompts the exact-match rate rises from 66.9% to 97.5% (mean similarity 0.994) as the candidate window is widened, confirming that most errors are recoverable near-misses rather than genuine ambiguities. A comparison with the released SIPIT reference situates these findings: per-step hard projection is faster, but the continuous formulation is what makes the optimisation observable and its failures detectable. The results show that last-layer hidden states of GPT-2 are as sensitive as the original text.
ミリ波ビームアライメントのためのメタ転移学習
ミリ波(mmWave)ビームアライメントは次世代無線システムにおいて重要な役割を果たしていますが、その効率的な実装は依然として課題です。メタ学習と転移学習は、深層学習ベースのビーム予測モデルが目に見えない環境に迅速に適応できるようにするために研究されてきました。ただし、既存のメタ学習アプローチはネットワーク全体を適応させ、ランダムな初期化からトレーニングされるため、更新されるパラメーターが多数になり、メタトレーニング コストが高くなります。一方、転移学習アプローチは適応をネットワークの一部に制限しますが、適応プロセス自体を最適化するために複数のタスクにわたってモデルを明示的にトレーニングする一時的なメタ学習を利用しません。これらの制限を克服するために、我々はミリ波多入力単出力(MISO)システムにおけるビームアライメントのためのメタ転移学習フレームワークであるMTL-BAを提案します。これは、事前訓練された畳み込みバックボーンをフリーズし、分類器ヘッドと一緒に軽量のスケールアンドシフト(SS)アダプターのみをメタ学習します。事前トレーニングされたモデルからウォームスタートし、SS アダプターと分類子ヘッドへの適応を制限することで、予測パフォーマンスを犠牲にすることなく、適応コストとメタトレーニング バジェットの両方を削減します。 DeepMIMO レイ トレーシング データセットのシミュレーション結果によると、MTL-BA は、フル ファインチューニングとモデルに依存しないメタ学習 (MAML) の両方より約 17 倍少ないパラメータを更新しているにもかかわらず、さまざまな SNR レベルにわたってフル ファインチューニングの精度とスペクトル効率に匹敵し、同程度の数のパラメータを更新しながら最終層のファインチューニングを上回っており、必要なメタトレーニングが 60 %$ 少ないにもかかわらず MAML のパフォーマンスに近づいていることが示されています。エポック。
原文 (English)
Meta-Transfer Learning for mmWave Beam Alignment
Millimeter-wave (mmWave) beam alignment plays a critical role in next-generation wireless systems, yet its efficient implementation remains challenging. Meta-learning and transfer learning have been explored to enable deep learning-based beam prediction models to rapidly adapt to unseen environments; however, existing meta-learning approaches adapt the entire network and are trained from random initialization, leading to a large number of updated parameters and a high meta-training cost, while transfer learning approaches restrict adaptation to part of the network but do not exploit episodic meta-learning, which explicitly trains the model over multiple tasks, to optimize the adaptation process itself. To overcome these limitations, we propose MTL-BA, a meta-transfer learning framework for beam alignment in millimeter-wave multiple-input single-output (MISO) systems that freezes a pre-trained convolutional backbone and meta-learns only lightweight Scale-and-Shift (SS) adapters together with a classifier head. Warm-starting from the pre-trained model and restricting adaptation to the SS adapters and classifier head reduce both the adaptation cost and the meta-training budget without sacrificing prediction performance. Simulation results on the DeepMIMO ray-tracing dataset show that MTL-BA matches the accuracy and spectral efficiency of full fine-tuning across various SNR levels despite updating approximately $17\times$ fewer parameters than both full fine-tuning and Model-Agnostic Meta-Learning (MAML), outperforms last-layer fine-tuning while updating a comparable number of parameters, and approaches MAML's performance while requiring $60\%$ fewer meta-training epochs.
CAT: 大規模推論モデルの効率的な推論のための信頼適応型思考
大規模推論モデル (LRM) は、長い思考連鎖 (CoT) の軌跡を活用することで複雑なタスクで目覚ましい成功を収めていますが、単純なクエリでは考えすぎが頻繁に発生し、その結果、トークンのオーバーヘッドが大きくなり、推論効率が低下します。ただし、既存の圧縮方法は主に均一な長さの削減を適用するか、粗粒度の難易度推定に依存しているため、難しい問題ではパフォーマンスの低下を招くことがよくあります。この制限に対処するために、我々は信頼適応型思考 (CAT) を提案します。これは、モデルの本質的な自己確実性信号を信頼度として好みの最適化プロセスに組み込むフレームワークであり、問題の難易度に基づいて推論の長さを自律的に調整します。実験結果は、CAT が、異なるベース モデルの複数のベンチマークにわたって、推論精度において常に最先端のベースラインを上回るパフォーマンスを示していることを示しています。私たちの取り組みにより、LRM は不確実な応答を検討しながら、信頼できる応答を効果的に圧縮できるようになり、実際の産業シナリオで精度と遅延のバランスをとるための潜在的に堅牢なソリューションを提供できます。
原文 (English)
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
Large Reasoning Models (LRMs) have achieved remarkable success on complex tasks by leveraging long chain-of-thought (CoT) trajectories, yet they frequently exhibit overthinking on simple queries, resulting in significant token overhead and reduced inference efficiency. However, existing compression methods predominantly apply uniform length reduction or rely on coarse-grained difficulty estimation, often leading to performance degradation on difficult problems. To address this limitation, we propose Confidence-Adaptive Thinking (CAT), a framework that incorporates the model's intrinsic self-certainty signals as confidence into the preference optimization process, which autonomously modulates reasoning lengths based on problem difficulty. Experimental results show that CAT consistently outperforms state-of-the-art baselines on reasoning accuracy across multiple benchmarks on different base models. Our work enables LRMs to effectively compress confident responses while deliberating on uncertain ones, offering a potentially robust solution for balancing accuracy and latency in practical industrial scenarios.
フラット ミニマ最適化によるスパース ビュー 3DGS の一般化の改善
ニューラル レンダリングの最近の進歩により、3D ガウス スプラッティング (3DGS) が新しいビュー合成のための高効率な表現として確立され、高速トレーニングと高い忠実度のリアルタイム レンダリングが可能になりました。ただし、監視がまばらな入力ビューに限定されている場合、3DGS は観察された画像に過剰適合し、目に見えない視点への一般化が不十分になる傾向があります。我々は、パラメータの小さな摂動下でも安定を保つソリューションを求めるフラット ミニマ (FM) 最適化の観点からこの課題に取り組みます。ガウス パラメーターをトレーニング可能な重みとみなして、軽量トレーニング フレームワークを使用して FM 原理を 3DGS の幾何学的および動的性質に適応させます。私たちの手法は、各ガウスの異方性とトレーニングの進行状況を考慮した制御されたガウス摂動で最適化を正規化し、スパースビューのオーバーフィッティングに対するロバスト性を向上させながら詳細を維持します。このフラット ミニマ最適化プロセスをさらに安定させるために、短いウィンドウで非位置パラメータを一時的に初期状態に戻す定期的な再初期化を導入します。これらの技術は、アーキテクチャを変更することなく、既存の 3DGS パイプラインにシームレスに統合されます。 LLFF および Mip-NeRF360 データセットの実験では、スパースビューの監視下で定量的メトリクスと知覚品質が向上し、より鮮明で安定し、新しい視点に対してより一般化された再構成が生成されることが実証されています。
原文 (English)
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
Recent advances in neural rendering have established 3D Gaussian Splatting (3DGS) as a highly efficient representation for novel view synthesis, enabling fast training and real-time rendering with strong fidelity. However, when supervision is limited to sparse input views, 3DGS tends to overfit to the observed images and generalize poorly to unseen viewpoints. We address this challenge from the perspective of flat minima (FM) optimization, which seeks solutions that remain stable under small parameter perturbations. Viewing Gaussian parameters as trainable weights, we adapt FM principles to the geometric and dynamic nature of 3DGS with a lightweight training framework. Our method regularizes optimization with controlled Gaussian perturbations that account for each Gaussian's anisotropy and the training progress, preserving fine details while improving robustness to sparse-view overfitting. To further stabilize this flat minima optimization process, we introduce periodic reinitialization, which temporarily returns non-positional parameters to their initial states for a short window. Together, these techniques integrate seamlessly into existing 3DGS pipelines without architectural changes. Experiments on LLFF and Mip-NeRF360 datasets demonstrate improved quantitative metrics and perceptual quality under sparse-view supervision, producing reconstructions that are sharper, more stable, and better generalized to novel viewpoints.
DeWorldSG: ワールドモデル事前分布による深度認識 3D セマンティック シーン グラフ生成
RGB-D シーケンスから時空間的に堅牢な 3D セマンティック シーン グラフを生成する新しいフレームワークである DeWorldSG を紹介します。既存の方法では、不安定な 3D オブジェクト表現やフレーム単位の推論によって生じる関係の欠落により、信頼性の高い 3D シーン グラフを構築するのに苦労することがよくあります。 DeWorldSG は、深度ガイド フィルタリングを通じてインスタンス レベルの幾何学的 3D ガウス分布を推定し、各オブジェクトを単一の投影点ではなく確率的な 3D ノードとして表すことで、これらの問題に対処します。フレーム単位の推論による関係の希薄性を軽減するために、私たちのフレームワークはオブジェクトのペア全体の時空間証拠をさらに集約し、ワールド モデル (V-JEPA 2) から導出された文脈的な事前分布を使用して関係を洗練します。 3DSSG および ReplicaSSG データセットの実験では、時間的に一貫したシーン構造を生成しながら、オブジェクトと述語の両方の予測における最先端 (SoTA) パフォーマンスを実証しました。特に、私たちの方法は、以前の SoTA アプローチと比較して、三重項再現率が 77.4%、述語再現率が 23.2% 向上し、ロボット操作や AR アプリケーションに適しています。私たちのコードとモデルはオープンソースです。
原文 (English)
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing methods often struggle to construct reliable 3D scene graphs due to unstable 3D object representations and missing relations caused by frame-wise inference. DeWorldSG addresses these issues by estimating instance-level geometric 3D Gaussian distributions through depth-guided filtering and representing each object as a probabilistic 3D node rather than a single projected point. To mitigate relational sparsity from frame-wise inference, our framework further aggregates spatiotemporal evidence across object pairs and refines relations using contextual priors derived from a world model (V-JEPA 2). Experiments on the 3DSSG and ReplicaSSG datasets demonstrate state-of-the-art (SoTA) performance in both object and predicate prediction, while producing temporally consistent scene structures. In particular, our method improves triplet recall by 77.4% and predicate recall by 23.2% over prior SoTA approaches, making it suitable for robotic manipulation and AR applications. Our code and models are open-sourced.
Valdi: 価値拡散世界モデル
ワールド モデルではモデル予測制御 (MPC) を有効にできますが、これにはオンラインでの使用に十分な速度と、不確実な将来を表現するのに十分な表現力の両方が必要なダイナミクス予測が必要です。拡散モデルは、不確実なダイナミクスをモデル化するための自然なメカニズムを提供しますが、その反復推論手順により、低レイテンシの潜在計画に使用することが困難になります。私たちは、MPC のエンドツーエンドのオンライン トレーニングと潜在拡散力学モデルを組み合わせた、Value Diffusion World Models (Valdi) でこのギャップを埋めます。 CarRacing 環境での予備実験では、Valdi がトレーニングと推論の両方で単一の拡散ステップを使用して、決定論的な MLP ベースラインと一致することを示しました。私たちの実験により、この設定における予測マルチモダリティと制御パフォーマンスの間のトレードオフが明らかになりました。コードは https://github.com/Kit115/ValueDiffusionWorldModels で入手できます。
原文 (English)
Valdi: Value Diffusion World Models
World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enough for online use and expressive enough to represent uncertain futures. Diffusion models offer a natural mechanism for modeling uncertain dynamics, yet their iterative inference procedure makes them difficult to use for low-latency latent planning. We bridge this gap with Value Diffusion World Models (Valdi), combining end-to-end online training for MPC with a latent diffusion dynamics model. In preliminary experiments on the CarRacing environment, we show that Valdi, using a single diffusion step at both training and inference, matches a deterministic MLP baseline. Our experiments expose a trade-off between predictive multimodality and control performance in this setup. Code is available at https://github.com/Kit115/ValueDiffusionWorldModels.
ペルソナからプロットまで: 長編物語向けのキャラクターに基づいたマルチエージェント ストーリーの生成
大規模言語モデル (LLM) は、印象的な創造的なフィクション生成を実証しましたが、長編物語で物語の一貫性と首尾一貫したプロット ラインを維持するのに苦労しています。この研究では、長い形式の物語の生成と検証のための統一フレームワークを導入します。ストーリーテリング用のマルチエージェントの目標駆動型ナラティブ エンジンである MAGNET は、共有された世界状態と進化するストーリー目標に基づいてアクションを提案する、ペルソナに基づいたキャラクター エージェントを使用してストーリーを生成します。一方、ATLAS は、生成されたストーリー全体でシーンレベルの世界表現を比較して幻覚を検出するグラフベースのパイプラインです。 LLM エディター、ペアごとのルーブリック スコアリング、および ATLAS を使用して MAGNET を評価することにより、単一モデル プロンプトや IBSEN と比較して、フレームワークが一貫した物語を生成することを示します。 100 ページでは、MAGNET は単一モデルのベースラインと比較して注釈と幻覚をそれぞれ 41 % と 50% 減少させ、IBSEN と比較してそれぞれ 34 % と 45% 減少させ、ペアごとのルーブリック評価でも同様の結果を示しました。これらの結果は、明示的な世界状態の追跡と目標駆動型のマルチエージェント生成から長編物語が出現し、制御可能で構造的に一貫した長編物語生成の基盤を提供することを示唆しています。
原文 (English)
From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives
Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce a unified framework for long-form narrative generation and verification. MAGNET, a multi-agent goal-driven narrative engine for storytelling, generates stories with persona-grounded character agents that propose actions based on a shared world state and evolving story goals, while ATLAS is a graph-based pipeline that compares scene-level world representations across a generated story to detect hallucinations. By evaluating MAGNET using an LLM editor, pairwise rubric scoring, and ATLAS, we show that our framework produces coherent narratives compared to single-model prompting and IBSEN. At 100 pages, MAGNET reduced annotations and hallucinations by 41 and 50%, respectively, compared to the single model baseline and by 34 and 45%, respectively, compared to IBSEN, with pairwise rubric evaluation showing similar results. These results suggest that long-form narratives can emerge from explicit world-state tracking and goal-driven multi-agent generation, providing a foundation for controllable and structurally coherent long-form narrative generation.
生成的メタ学習における人間と機械のコラボレーション: モデルとアルゴリズム
機械学習モデルをトレーニング分布とは異なる環境に一般化することは、特にターゲット ドメインのデータが完全または部分的に利用できない場合には依然として重大なハードルとなります。私たちは、専門家の直観を活用してデータ合成をガイドすることで、この領域のギャップを埋める新しいフレームワークであるヒューマン フィードバックによる生成メタ学習 (GMHF) を提案します。一般化誤差の理論的分析に基づいて、生成されたデータの分布を対象の物理学に関する人間の信念と一致させるとリスクが大幅に軽減されることを示す境界を導き出します。 GMHF は、条件付きニューラル ODE (cNODE) を生成デジタル ツインとして採用し、強化学習 (RL) エージェントと組み合わせることで、この洞察を運用します。エージェントは、フィードバックに基づいて生成された軌跡の潜在的な物理パラメータを繰り返し改良し、メタ学習者を未観測のターゲット分布に向けて効果的に誘導します。非線形ダフィング発振器の経験的検証により、専門家の信頼性が高まるにつれて GMHF が展開損失を大幅に低減し、生成されたデータとターゲット データの間の乖離が信頼性の高いフィードバックに該当することが示され、我々の理論によって予測された乖離最小化メカニズムが直接裏付けられています。非動的確率モデルに関するさらなる実験により、このフレームワークが ODE が管理するシステムを超えて拡張され、分布シフト下での堅牢な一般化のための厳密な触媒として人間と AI のコラボレーションが確立されることが確認されました。
原文 (English)
Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm
Generalizing machine learning models to environments that differ from their training distribution remains a critical hurdle, particularly when data from the target domain is entirely or partially unavailable. We propose Generative Meta-Learning with Human Feedback (GMHF), a novel framework that bridges this domain gap by leveraging expert intuition to guide data synthesis. Grounded in a theoretical analysis of generalization error, we derive bounds demonstrating that aligning the distribution of generated data with human beliefs regarding the target physics significantly mitigates risk. GMHF operationalizes this insight by employing a Conditional Neural ODE (cNODE) as a generative digital twin, coupled with a Reinforcement Learning (RL) agent. The agent iteratively refines the latent physical parameters of the generated trajectories based on feedback, effectively steering the meta-learner toward the unobserved target distribution. Empirical validation on a nonlinear Duffing oscillator shows that GMHF substantially reduces deployment loss as expert reliability increases, and that the divergence between generated and target data falls under reliable feedback, directly corroborating the divergence-minimisation mechanism predicted by our theory. Further experiments on a non-dynamical probabilistic model confirm that the framework extends beyond ODE-governed systems, establishing human-AI collaboration as a rigorous catalyst for robust generalisation under distribution shift.
拡散変圧器のトレーニング後の枝刈り
拡散変換器 (DiT) は、画像生成において優れたパフォーマンスを示していますが、かなりの計算オーバーヘッドとリソース消費に悩まされています。トレーニング後の枝刈りは有望な解決策を提供します。ただし、DiT の独自のアーキテクチャ設計とパラメータ分布により、従来のプルーニング手法は適用できず、大幅なパフォーマンスの低下につながります。具体的には、LLM 用に開発された、一連の近似によってメトリクスを導出する従来の方法では、顕著性メトリクスにおける重みの相対的な寄与が増幅されます。さらに、DiT の重みは、LLM の重みよりも大幅に大きい値を示します。さらに、既存の枝刈り粒度では、モデル構造の変化が見落とされます。この論文では、カスタマイズされた顕著性基準と枝刈り粒度を導入することで枝刈りパフォーマンスを向上させる DiT-Pruning を提案します。私たちは、エネルギーベースの観点から重みと活性化の寄与のバランスを取る新しい指標を設計し、重要な要素をより効果的に特定できるようにします。さらに、2 次元の重み空間で明確なクラスタリング パターンが観察されます。したがって、クラスタリングを意識したプルーニング粒度を採用し、効果的なスパース割り当てを可能にします。さまざまな DiT に関する広範な評価により、特に高いスパース性の下で、私たちの方法が一貫して画質を維持することが示されています。 MJHQ 上の 512x512 解像度の FLUX.1-dev の場合、DiT-Pruning は 50% のスパース性で CLIP スコアの損失がわずか 0.001 であり、最近のプルーニング手法を劇的に上回っています。
原文 (English)
Post-Training Pruning for Diffusion Transformers
Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhead and resource consumption. Post-training pruning offers a promising solution; however, due to DiTs' unique architectural design and parameter distribution, traditional pruning methods are inapplicable, leading to significant performance degradation. Specifically, prior methods developed for LLMs, which derive metrics through a series of approximations, amplify the relative contribution of weights in the saliency metric. In addition, weights in DiTs exhibit significantly larger magnitudes than those in LLMs. Moreover, existing pruning granularity overlooks variations in model structures. In this paper, we propose DiT-Pruning, which improves pruning performance by introducing customized saliency criteria and pruning granularity. We design a novel metric that balances the contributions of weights and activations from an energy-based perspective, enabling more effective identification of important elements. Furthermore, we observe distinct clustering patterns in the two-dimensional weight space. Accordingly, we adopt a clustering-aware pruning granularity, enabling effective sparse allocation. Extensive evaluations on various DiTs show that our method consistently preserves image quality, especially under high sparsity. For FLUX.1-dev at 512x512 resolution on MJHQ, DiT-Pruning achieves only a 0.001 loss in CLIP score at 50% sparsity, dramatically outperforming recent pruning methods.
暗黙的な神経表現のための心臓運動事前分布の学習
Implicit Neural Representation (INR) は心臓の運動推定に適しており、運動フィールドの連続的でコンパクトな表現を提供します。ただし、INR を各画像シーケンスに適合させるのは時間がかかり、最適化の軌道に左右されます。学習された事前分布は、最適化を妥当な運動フィールドに向けて導き、より迅速な適応を可能にするのに役立ちますが、心臓の運動 INR の学習事前分布はまだ研究が進んでいません。この研究では、関節最適化によって学習された母集団事前学習、重み平均化によって取得されたコンセンサス事前学習、自動デコーダー、およびメタ学習を含む、心臓運動事前学習のための 4 つの戦略を比較します。英国バイオバンクからの短軸タグ付き心臓磁気共鳴画像を使用して、追跡精度、運動挙動、および適応軌道への影響を評価します。すべての学習された事前確率は、ランダムな初期化と比較して、早期適応パフォーマンスを大幅に向上させました。事前の単純なコンセンサスは効果的でしたが、自動デコーダは初期の適応中に大きな変形をより速く回復しました。メタ学習は初期に強力なパフォーマンスを達成し、50 回の反復にわたって最良の適応軌道を維持しました。
原文 (English)
Learning Cardiac Motion Priors for Implicit Neural Representations
Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motion fields. However, fitting an INR to each image sequence is time-consuming and sensitive to the optimisation trajectory. Learned priors can help guide optimisation towards plausible motion fields and enable faster adaptation, but learning priors for cardiac motion INRs remains under-explored. In this work, we compare four strategies for learning cardiac motion priors, including a population prior learned by joint optimisation, a consensus prior obtained by weight averaging, auto-decoders, and meta-learning. Using short-axis tagged cardiac magnetic resonance images from the UK Biobank, we evaluate their impact on tracking accuracy, motion behaviour, and adaptation trajectory. All learned priors substantially improved early adaptation performance compared with random initialisation. While the simple consensus prior was effective, auto-decoders recovered large deformations faster during early adaptation. Meta-learning achieved strong early performance and maintained the best adaptation trajectory over 50 iterations.
Aionscope: 時系列表現における潜在状態アクセシビリティのデバッグ
時系列モデルは、予測または分類できる内容によって評価されることがよくありますが、それらのスコアは、その表現が、ユーザーが検査したい可能性のあるプロセス状態 (イベントのタイミング、位相、振幅、周波数、またはレジーム変数) を保持しているかどうかを示しません。凍結された時系列表現における潜在状態へのアクセスをデバッグするためのジェネレーター ベースの診断ツールである Aionscope を紹介します。 Aionscope は、プロセスの生成を観察のレンダリングから分離し、混合の複雑さと厄介な変動全体にわたって正確なカテゴリカルで高密度のラベルを備えたシードされた合成ストリームを生成します。 Aionscope をプリミティブ プロセス混合物としてインスタンス化し、共通のプールされたリニア プローブ プロトコルを使用して 37 のモデルとアダプターのシステムを評価します。主な結果は、粗いアクセシビリティと細かいアクセシビリティの間の不一致です。ほとんどのシステムではコンポーネントの存在を簡単に回復できますが、高密度プロセス状態を公開する信頼性ははるかに低くなります。観測された最高の高密度プローブ行は平均マスク $R^2$ 0.689 に達しますが、高密度特徴オラクルは 0.999 に達します。これは、Aionscope が表面化するように設計された障害モードです。表現は、デバッグに必要なタイミング、位相、振幅、周波数、またはレジーム変数を非表示にしながら、「どのような種類の信号が存在するか」というレベルで情報を提供するように見えます。
原文 (English)
Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations
Time-series models are often evaluated by what they can forecast or classify, but those scores do not show whether their representations preserve the process state a user may want to inspect: event timing, phase, amplitude, frequency, or regime variables. We introduce Aionoscope, a generator-based diagnostic tool for debugging latent-state accessibility in frozen time-series representations. Aionoscope separates process generation from observation rendering, producing seeded synthetic streams with exact categorical and dense labels across mixture complexity and nuisance variation. We instantiate Aionoscope as Primitive Process Mixtures and evaluate 37 model-plus-adapter systems with a common pooled linear-probe protocol. The main result is a mismatch between coarse and fine-grained accessibility. Most systems make component presence easy to recover, but expose dense process state much less reliably: the highest observed dense-probe row reaches 0.689 mean masked $R^2$, while a dense-feature oracle reaches 0.999. This is the failure mode Aionoscope is designed to surface: a representation can look informative at the level of "what kind of signal is present" while hiding the timing, phase, amplitude, frequency, or regime variables needed for debugging.
TRCGL-Net: 生成データ拡張とラベル共起モデリングを備えたロングテールマルチラベル胸部 X 線分類フレームワーク
胸部 X 線のマルチラベル分類は、インテリジェントな医療画像診断の中核となるタスクです。ただし、実際の臨床データは極端なロングテール分布を示すことが多く、テールクラスの希少疾患のパフォーマンスの低下につながります。この問題は、データ不足だけでなく、2 つの固有の要因によっても引き起こされます。1) 複雑な解剖学的背景の下での尾部クラスの病変表現の減衰、および 2) ラベルの共起関係のモデリングにおける頭部クラスの優位性。これらの課題に対処するために、私たちは TRCGL-Net を提案します。まず、学習可能なテキストガイドによる条件付き拡散モデルを採用して、疾患の意味論的制約の下で高品質の尾部クラスの胸部 X 線画像サンプルを生成し、クラスの不均衡を軽減し、病理整合性のあるセマンティクスを維持しながら、データの多様性と希少疾患パターンの現実性を向上させます。第 2 に、チャネル再重み付けメカニズムを導入して、疾患関連の特徴チャネルを強調することで特徴の再調整を実行し、それによってロングテールの下での特徴の識別性を向上させます。クラス認識アテンション メカニズムをさらに適用してクラス固有のアテンション マップを生成し、モデルが疾患関連領域の位置を特定し、きめの細かい病変領域に焦点を当てることができるようにします。最後に、ラベルの共起に基づくグラフ畳み込みネットワークを導入して、カテゴリ間の情報伝播メカニズムを確立します。 PadChest データセットの実験では、提案された方法がテールクラス mAP 0.4904、全体 mAP 0.4408、mAUC 0.8989 を達成し、最先端の方法を上回るパフォーマンスを示していることが示されています。 TRCGL-Net は、尾の長い分布の下で希少疾患の認識パフォーマンスを効果的に向上させ、胸部 X 線マルチラベル分類における極端なクラスの不均衡の影響を軽減します。
原文 (English)
TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling
Chest X-ray multi-label classification is a core task in intelligent medical imaging diagnosis. However, real clinical data often exhibit extreme long-tailed distributions, leading to degraded performance on rare diseases in tail classes. This issue is not only driven by data scarcity but also by two intrinsic factors:1) attenuation of tail-class lesion representations under complex anatomical backgrounds, and 2) dominance of head classes in modeling label co-occurrence relationships. To address these challenges, we propose TRCGL-Net. First, a learnable text-guided conditional diffusion model is employed to generate high-quality tail-class chest X-ray image samples under disease semantic constraints, improving data diversity and realism of rare disease patterns while alleviating class imbalance and preserving pathology-consistent semantics.Second, a channel reweighting mechanism is introduced to perform feature recalibration by emphasizing disease-relevant feature channels, thereby improving feature discriminability under long-tailed distributions.A class-aware attention mechanism is further applied to generate class-specific attention maps, enabling the model to localize disease-relevant regions and focus on fine-grained lesion areas.Finally, a graph convolution network based on label co occurrence is introduced to establish an information propagation mechanism among categories. Experiments on the PadChest dataset show that the proposed method achieves a tail-class mAP of 0.4904, an overall mAP of 0.4408, and an mAUC of 0.8989, outperforming state-of-the-art methods. TRCGL-Net effectively improves recognition performance for rare diseases under long-tailed distributions and mitigates the impact of extreme class imbalance in chest X-ray multi-label classification.
SenseWalk: ゾーン環境における大規模言語モデルを活用したエージェントベースのセマンティック軌跡シミュレーション
意味軌跡分析は、生の空間経路を超えて意味論的情報 (訪問者のプロファイルや目標など) を通じて暗黙のパターンや行動を捕捉することにより、人間の動きをモデル化するアプローチとして最近登場し、人々が特定の方法で移動する理由をより深く理解できるようになりました。ただし、高品質のデータの収集にはコストがかかり、豊富なセマンティック情報が不足していることが多いため、現実世界のシナリオでセマンティック トラジェクトリを分析することは依然として困難です。一方、既存のシミュレーション ツールはかなりの技術的専門知識を必要とするため、実務者が導入するのは困難です。これらの制限に対処するために、この論文では、LLM を利用したエージェントによるセマンティック トラジェクトリのシミュレーションをサポートする対話型システムである ${SenseWalk}$ を提案しています。私たちは、物理的な妥当性と意味論的な一貫性のバランスをとるために、LLM と社会的力モデルを組み合わせたシミュレーション ワークフローを開発します。ユーザーフレンドリーなインターフェイスは、ユーザーがシミュレーション構成をカスタマイズし、シミュレーション出力を分析しやすいように設計されています。また、シミュレーション ワークフローの有効性を評価するための定量的な実験と、システムの有用性と効率を評価するためのユーザー調査 (n=12) も実施します。
原文 (English)
SenseWalk: Agent-Based Semantic Trajectory Simulation Powered by Large Language Models in Zoned Environments
Semantic trajectory analysis has recently emerged as an approach for modeling human movement by capturing implicit patterns and behaviors through semantic information (e.g., visitors' profiles and goals) beyond raw spatial paths to better understand why people move in certain ways. However, analyzing semantic trajectories in real-world scenarios remains challenging, as collecting high-quality data is costly and often lacks rich semantic information. Meanwhile, existing simulation tools require substantial technical expertise, which makes them difficult for practitioners to adopt. To address these limitations, the paper proposes ${SenseWalk}$, an interactive system that supports simulating semantic trajectories by LLM-powered agents. We develop a simulation workflow that combines LLMs and the social force model to balance physical plausibility and semantic coherence. A user-friendly interface is designed to facilitate users in customizing the simulation configuration and analyzing simulation outputs. We also conduct a quantitative experiment to evaluate the effectiveness of our simulation workflow, and a user study (n=12) to assess the usefulness and efficiency of our system.
SWE-Doctor: 多面的なバグ再現テストによる実行時診断でソフトウェア エンジニアリング エージェントを指導
大規模言語モデル (LLM) ベースのソフトウェア エンジニアリング エージェントは、問題レポートやコード リポジトリからパッチを生成することでソフトウェアの問題を解決するために開発されることが増えています。バグ再現テスト (BRT) は、このようなエージェントにとって重要な構成要素であり、パッチの検証に役立つことがわかっています。ただし、BRT がパッチ生成のより中心的な段階でも役立つかどうかは不明です。私たちは最初に予備調査を実施し、高度な BRT ジェネレーターをパッチ生成のガイドに直接使用することは有益ではないことを発見しました。フェイルツーフェイル BRT はエージェントを誤解させる可能性があり、フェイルツーフェイル BRT でさえ限定的またはマイナスの利益をもたらします。私たちの分析により、2 つの理由が明らかになりました。1 つは、Fail-to-Pass BRT は報告された問題の 1 つの症状のみをカバーし、部分的なパッチにつながる可能性があるのに対し、Fail-to-Fail BRT は直接のパッチ生成ターゲットとしては信頼性が低いことです。これらの洞察に基づいて、私たちは、多面的な BRT 実行から得られるランタイム診断を使用してパッチ生成をガイドするソフトウェア問題解決エージェントである SWE-Doctor を提案します。 SWE-Doctor はまず、問題に記載されているさまざまな動作要件に応じて多面的な BRT を生成し、次にこれらの BRT を実行およびデバッグしてランタイムに基づいた診断レコードを構築し、最後にその診断を BRT 生成中に推定されたローカライゼーション情報とともに使用して、パッチ生成をガイドし、部分的なパッチを削減します。私たちは、5 つの LLM バックエンドにわたって広く採用されている SWE-bench Verified および SWE-bench Pro からの SWE-Doctor on Python のバグ修正問題を評価します。 SWE-Doctor は、10 種類の LLM ベンチマークの組み合わせすべてにおいて既存のエージェントを常に上回り、平均解決率は SWE-bench Verified で 75.7%、SWE-bench Pro で 59.4% を達成しました。特に、より困難な SWE-bench Pro では、SWE-Doctor はベースライン エージェントに比べて平均解決率を 8.0 ~ 8.9 パーセント ポイント改善します。
原文 (English)
SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests
Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue reports and code repositories. Bug reproduction tests (BRTs) are an important building block for such agents and have been shown useful for patch validation. However, it remains unclear whether BRTs can also help the more central stage of patch generation. We first conduct a preliminary study and find that directly using advanced BRT generators to guide patch generation is not beneficial: fail-to-fail BRTs can mislead agents, while even fail-to-pass BRTs bring limited or negative gains. Our analysis reveals two reasons: fail-to-pass BRTs may cover only one manifestation of the reported issue, leading to partial patches, whereas fail-to-fail BRTs are unreliable as direct patch-generation targets. Motivated by these insights, we propose SWE-Doctor, a software issue resolution agent that guides patch generation with runtime diagnoses derived from multi-faceted BRT executions. SWE-Doctor first generates multi-faceted BRTs for different behavioral requirements stated in the issue, then executes and debugs these BRTs to construct runtime-grounded diagnosis records, and finally uses the diagnoses together with localization information inferred during BRT generation to guide patch generation and reduce partial patches. We evaluate SWE-Doctor on Python bug-fixing issues from the widely adopted SWE-bench Verified and SWE-bench Pro across five LLM backends. SWE-Doctor consistently outperforms existing agents across all 10 LLM-benchmark combinations, achieving average resolution rates of 75.7% on SWE-bench Verified and 59.4% on SWE-bench Pro. In particular, on the more challenging SWE-bench Pro, SWE-Doctor improves the average resolution rate by 8.0-8.9 percentage points over the baseline agents.
ロジット貢献度スコアリングにより非リテラル検索ヘッドを特定
長いコンテキストで使用される場合、大規模な言語モデルは、文字通りコピー&ペーストするのではなく、関連するコンテキスト スパンの意味から回答を合成することがよくあります。どのアテンションヘッドがこの合成を実行するかを特定することは、ロングコンテキストモデルの動作を解釈する上で重要です。しかし、既存の検出器は構造上、これらのヘッドを見逃しています。つまり、生成されたトークンと一致する注目トークンを持つヘッドに報酬を与えます。リテラル コピー基準は、出力値 (OV) 回路を通じて、ヘッドが読み取る場所を捕捉するが、書き込む内容は捕捉しないという、まさに非リテラル検索を実行するメカニズムです。ロジット貢献度スコアリング (LOCOS) を導入します。これは、OV 回路の出力を応答トークンの埋め込み解除方向に投影し、単一の順方向パスでニードルとオフニードルのソース位置を対比させて各ヘッドをスコアリングする書き込み認識検出器です。 3 つのモデル ファミリ (Qwen3、Gemma-3、OLMo-3.1) にわたって、NoLiMa 非リテラル検索ベンチマークの上位 LOCOS ヘッドを平均アブレーションすると、以前の注意ベースの検出よりも低いヘッド数で ROUGE-L が崩壊します。 Qwen3-8B では、50 ヘッドをアブレーションすると、ROUGE-L が 0.401 から 0.000 に上昇しますが、最も強いベースラインは依然として 0.292 を維持します。選択されたヘッドは検索に固有です。パラメトリック想起と算術推論は、同じアブレーションの下でもベースラインに留まります。 Qwen3-8B では、同じアブレーションによって MuSiQue が 0.55 から 0.08 に、BABI-Long が 0.62 から 0.20 に低下しましたが、ランダムヘッド コントロールはベースラインの 0.05 以内に留まりました。
原文 (English)
Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them. Identifying which attention heads perform this synthesis matters for interpreting long-context model behavior. Yet existing detectors miss these heads by construction: they reward heads whose attended token matches the generated token, a literal-copy criterion that captures where a head reads but not what it writes through its output-value (OV) circuit, the very mechanism that carries non-literal retrieval. We introduce Logit-Contribution Scoring (LOCOS), a write-aware detector that scores each head by the projection of its OV-circuit output onto the answer-token unembedding direction, contrasting needle and off-needle source positions in a single forward pass. Across three model families (Qwen3, Gemma-3, OLMo-3.1), mean-ablating the top LOCOS heads on the NoLiMa non-literal retrieval benchmark collapses ROUGE-L at lower head counts than prior attention-based detections; on Qwen3-8B, ablating 50 heads drives ROUGE-L from 0.401 to 0.000 while the strongest baseline still retains 0.292. The selected heads are retrieval-specific: parametric recall and arithmetic reasoning stay at baseline under the same ablation. On Qwen3-8B, the same ablation also drops MuSiQue from 0.55 to 0.08 and BABI-Long from 0.62 to 0.20, while a random-heads control stays within 0.05 of baseline.
複雑な文書レイアウトの読み取り順序の推論
読み順の推論は、複雑な歴史写本のデジタル化において依然として重大なボトルネックとなっています。ページには空間的に交互に配置された複数の読みの流れが含まれており、標準的な例は Glossa Ordinaria レイアウトです。このレイアウトでは、中央のテキストが、非長方形、非凸面の領域で周囲を囲む注釈によって囲まれています。トレーニング不要のグラフベースのフレームワークを提示します。各 OCR テキスト行は有向候補遷移グラフのノードになり、エッジは 2 つの軽量言語モデル信号 (因果言語モデルの条件付き尤度と BERT 次文予測、NSP、NSP、3 番目の文埋め込み信号は評価されましたが、読み上げ順序は改善されませんでした) の加重加算アンサンブルによってスコア付けされ、グローバルな読み上げ順序は次数制約付きの有向パス カバーとして復元されます。貪欲なエッジ選択による連鎖的な「エッジ盗難」失敗を回避するために、機会コストの高いコミットメントを優先する最大リグレット推論ルールを提案します。合成の Glossa Ordinaria グリッド レイアウト、23 の ALTO ページ ジオメトリ (10 の歴史的ソース ページとミラー化および反転されたバリアント)、および OmniDocBench の 140 ページの複数列英語サブセットで評価し、正規の再帰 XY カット (PaddleOCR PP-StructureV3) および 2 つの LayoutReader バリアント (レイアウトのみおよびテキスト + レイアウト) に対してメソッドを比較します。同一の入力。ラップアラウンド Glossa レイアウトでは、XY カットの 50% に対して、私たちの方法は平均してグラウンドトゥルース後継エッジの 95% を回復します。 OmniDocBench の複数列サブセットでは、マクロ エッジ精度が 88% に達するのに対し、XY カットの 75%、LayoutReader の精度は 25% です。 LayoutReader ベースラインは、ワードレベルとラインレベルの粒度の不一致により、転送が不十分になります。さらに、水平方向および垂直方向のページ反射下でのミラー不変性も検証します。私たちの方法では 1 パーセント未満の変化、古典的な XY カットでは 2 ポイント、LayoutReader-T では最大 8 ポイント変化します。
原文 (English)
Reading Order Inference for Complex Document Layouts
Reading order inference remains a critical bottleneck in the digitization of complex historical manuscripts, where pages contain multiple spatially interleaved reading streams, the canonical example being the Glossa Ordinaria layout, in which a central text is surrounded by commentaries that wrap around it in non-rectangular, non-convex regions. We present a training-free, graph-based framework: each OCR text line becomes a node in a directed candidate-transition graph, edges are scored by a weighted additive ensemble of two lightweight language-model signals (causal language model conditional likelihood and BERT next-sentence prediction, NSP; a third sentence-embedding signal was evaluated but did not improve reading order), and the global reading order is recovered as a degree-constrained directed path cover. To avoid the cascading "edge-theft" failures of greedy edge selection, we propose a max-regret inference rule that prioritizes commitments with high opportunity cost. We evaluate on synthetic Glossa Ordinaria grid layouts, on 23 ALTO page geometries (10 historical source pages plus mirrored and flipped variants), and on a 140-page multi-column English subset of OmniDocBench, comparing our method against the canonical recursive XY-cut (PaddleOCR PP-StructureV3) and two LayoutReader variants (layout-only and text+layout) on identical inputs. On wrap-around Glossa layouts our method recovers 95% of ground-truth successor edges on average vs. XY-cut's 50%; on the OmniDocBench multi-column subset it reaches 88% macro edge accuracy versus XY-cut's 75% and LayoutReader's 25%. The LayoutReader baselines transfer poorly due to a word-level vs. line-level granularity mismatch. We additionally verify mirror-invariance under horizontal and vertical page reflections: Our method changes by less than 1 percentage point, classical XY-cut by 2 points, and LayoutReader-T by up to 8 points.
行動適応型会話エージェント: 流動的なパーソナリティフレームワークに向けて
大規模言語モデル (LLM) ベースの会話エージェント (CA) は現在どこにでも普及しており、AI を介した行動変化の新たな機会を生み出しています。微妙な個性を投影し、多様な比喩的な役割を採用する彼らの能力は、エージェントのペルソナと性格を現時点に合わせてどのように調整する必要があるのかという設計上の問題を引き起こします。最近の証拠によると、(i) 中程度の性格表現は、信頼、楽しさ、目標指向のタスクに採用する意図に関して、低極端または高極端よりも優れており、(ii) コンテキストに適したメタファーは、ユーザー エクスペリエンスと取り込みに関して、静的なワンノート アシスタントよりも優れていることが示唆されています。しかし、ほとんどの CA は依然としてペルソナとスタイルの両方を固定しており、たとえば医療情報の探索、フィットネス コーチング、内省的学習など、ダイナミクス、緊急性、形式が異なる場合に調整が合わなくなる危険があります。我々は、(1) コーチ、家庭教師、司書、道具などのエージェントの比喩的なペルソナと、(2) タスクのコンテキスト、ユーザーの目標と特性、および状況の緊急度の関数として、エージェントの人格表現強度 (低、中、または高) を共同で適応させる、流動的なパーソナリティ フレームワークを提案します。フレームワークとその中心となる設計寸法をスケッチします。
原文 (English)
Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework
Large language model (LLM)-based conversational agents (CAs) are now ubiquitous, creating new opportunities for AI-mediated behavior change. Their capacity to project nuanced personalities and adopt diverse metaphorical roles raises a design question: how should an agent's persona and personality be calibrated to the moment? Recent evidence suggests that (i) moderate personality expression outperforms low or high extremes on trust, enjoyment, and intention to adopt in goal-oriented tasks, and (ii) context-appropriate metaphors outperform static one-note assistants on user experience and uptake. Yet most CAs still fix both persona and style, risking misalignment when dynamics, urgency, and formality vary, for example in medical information seeking, fitness coaching, and reflective learning. We propose a Fluid Personality Framework that jointly adapts (1) the agent's metaphorical persona, such as coach, tutor, librarian, or tool, and (2) its personality expression intensity, low, medium, or high, as a function of task context, user goals and traits, and situational urgency. We sketch the framework and its core design dimensions.
EchoRisk: 心臓腫瘍学の多施設心エコー検査データセットおよびベンチマーク
治療誘発性心毒性は、乳がん患者における治療中断の主な非腫瘍学的原因ですが、定期的な心臓画像による初期の自動リスク層別化は未解決の問題のままです。 EchoRisk は、EchoRisk-MICCAI 2026 チャレンジの主要な技術リファレンスとしてリリースされた、明示的な心毒性ラベルを備えた最初の厳選された多施設縦断心エコー検査データセットです。このデータセットは、欧州の 5 施設にわたる EU 資金による CARDIOCARE 前向き研究に登録された 422 人の患者で構成されており、早期心毒性予測のためのベースライン画像を取得した 280 人の患者の専用コホートと並んで、最大 5 つの縦断的時点で取得された 1,123 の臨床検査にわたる 2,159 件の心エコー検査ビデオが得られます。臨床的に根拠のある 3 つのタスクが定義されています。シネビデオからの左心室駆出率の自動推定 (タスク 1)、縦断画像からの LV 機能不全の分類 (タスク 2)、および治療前のベースライン心エコー検査のみからの治療誘発性心毒性の早期予測 (タスク 3) です。タスクごとに、評価プロトコル、一次および二次メトリック、およびランキング手順を指定します。当社は、Kinetics-400 の事前トレーニング済み重みからトレーニングされた LSTM 集約を備えた R(2+1)D ビデオ バックボーンを使用してベースライン パフォーマンスを確立し、心機能評価と LV 機能障害分類に対する強力な識別パフォーマンスを実証していますが、単一の治療前ビデオからの早期心毒性予測は依然としてコミュニティにとって重要な未解決の問題です。データセット、評価コード、およびベースライン実装は、心臓腫瘍学におけるさらなるコラボレーション、比較、およびタスク固有のアーキテクチャの作成のためのベンチマークとして機能するために公開されています。
原文 (English)
EchoRisk: A Multicentre Echocardiography Dataset and Benchmark for Cardio-Oncology
Therapy-induced cardiotoxicity is the leading non-oncological cause of treatment interruption in breast cancer patients, yet early, automated risk stratification from routine cardiac imaging remains an unsolved problem. We present EchoRisk, the first curated, multicentre, longitudinal echocardiography dataset with explicit cardiotoxicity labels, released as the primary technical reference for the EchoRisk-MICCAI 2026 challenge. The dataset comprises 422 patients enrolled in the EU-funded CARDIOCARE prospective study across five European sites, yielding 2,159 echocardiography videos across 1,123 clinical exams acquired at up to five longitudinal timepoints, alongside a dedicated cohort of 280 patients with baseline imaging for early cardiotoxicity prediction. Three clinically grounded tasks are defined: automated estimation of left ventricular ejection fraction from cine video (Task 1), classification of LV dysfunction from longitudinal imaging (Task 2), and early prediction of therapy-induced cardiotoxicity from pre-therapy baseline echocardiography alone (Task 3). For each task we specify the evaluation protocol, primary and secondary metrics, and ranking procedure. We establish baseline performance using an R(2+1)D video backbone with LSTM aggregation trained from Kinetics-400 pretrained weights, demonstrating strong discriminative performance for cardiac functional assessment and LV dysfunction classification, while early cardiotoxicity prediction from a single pre-therapy video remains a significant open problem for the community. The dataset, evaluation code, and baseline implementations are publicly available to serve as a benchmark for further collaboration, comparison, and the creation of task-specific architectures in cardio-oncology.
DART-VLN: 離散視覚言語ナビゲーションのためのテスト時のメモリ減衰とアンチループ正則化
メモリベースの離散ビジョン言語ナビゲーション (VLN) エージェントは部分的な可観測性の下で動作する必要がありますが、強力な凍結バックボーンでさえテスト時には脆弱なままです。一般的な 2 つの障害モードは、メモリ読み出し時の古い履歴証拠と、アクション選択時の非効率なローカル バックトラッキングです。離散 VLN 用のトレーニング不要のテスト時間制御フレームワークである DART-VLN を紹介します。 DART-VLN は、保存されたコンテンツを書き換えることなく、古くなって冗長な証拠を抑制する読み取り側メモリ再重み付けルールである Test-Time Memory Decay と、アクション選択中の即時逆転を阻止する軽量のネクストホップ ペナルティである Anti-Loop Regularization を組み合わせています。このフレームワークでは、新しい学習可能なパラメーターは導入されず、学習されたバックボーンは変更されません。 R2R と REVERIE の実験では、一貫したパターンが示されています。ディケイのみでは安定した読み取り側ゲインが得られますが、ディケイ + アンチループでは全体的に最高の品質効率のトレードオフが達成され、主要な設定でより短い軌道、より短いランタイム、および改善されたナビゲーション パフォーマンスが得られます。動作分析により、アンチループ正則化によりローカル バックトラッキングが減少し、フリーズしたバックボーンの下でパス効率が向上することがさらに確認されました。全体として、この結果は、適度なテスト時間制御により、再トレーニングすることなくメモリベースの離散 VLN の信頼性と効率性を高めることができることを示しています。
原文 (English)
DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation
Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory readout and inefficient local backtracking during action selection. We present DART-VLN, a training-free test-time control framework for discrete VLN. DART-VLN combines Test-Time Memory Decay, a read-side memory reweighting rule that suppresses stale and redundant evidence without rewriting stored content, with Anti-Loop Regularization, a lightweight next-hop penalty that discourages immediate reversals during action selection. The framework introduces no new learnable parameters and leaves the learned backbone unchanged. Experiments on R2R and REVERIE show a consistent pattern: decay-only provides stable read-side gains, while decay+anti-loop achieves the best overall quality-efficiency trade-off, yielding shorter trajectories, lower runtime, and improved navigation performance in key settings. Behavioral analysis further confirms that anti-loop regularization reduces local backtracking and improves path efficiency under frozen backbones. Overall, the results show that modest test-time control can make memory-based discrete VLN more reliable and efficient without retraining.
MemSyco-Bench: エージェントのメモリにおけるお調子者のベンチマーク
メモリは、現代の LLM ベースのエージェントの基礎として浮上し、ワンターンのアシスタントから長期的な協力者への進化をサポートします。ただし、記憶が常に有益であるとは限りません。検索された記憶は、多くの場合、お調子者という重大な問題を引き起こし、事実の正確さや客観的な推論を犠牲にして、エージェントがユーザーに過度に同調する原因となります。この新たなリスクにもかかわらず、既存のメモリ ベンチマークは主にメモリが正しく保存、取得、または更新されているかどうかを評価し、取得されたメモリが下流の推論や意思決定にどのような影響を与えるかを見落としています。このギャップを埋めるために、エージェント システムにおける記憶誘発性のおしゃべりを評価するための包括的なベンチマークである MemSyco-Bench を提案します。 MemSyco-Bench は、メモリが決定に影響を与える時期と有効なメモリをどのように使用するかを測定します。具体的には、エージェントが記憶を事実証拠として拒否できるかどうか、その適用範囲を尊重できるかどうか、記憶と客観的証拠の間の矛盾を解決できるかどうか、記憶の更新を追跡できるかどうか、パーソナライゼーションに有効な記憶を使用できるかどうかを評価する 5 つのタスクをカバーしています。すべての関連リソースは、https://github.com/XMUDeepLIT/MemSyco-Bench のコミュニティ用に収集されています。
原文 (English)
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlooking how retrieved memories influence downstream reasoning and decision-making. To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. MemSyco-Bench measures when memory should influence a decision and how valid memory should be used. Specifically, it covers five tasks that assess whether agents can reject memory as factual evidence, respect its applicable scope, resolve conflicts between memory and objective evidence, track memory updates, and use valid memory for personalization. All related resources are collected for the community at https://github.com/XMUDeepLIT/MemSyco-Bench.
非同期 RLHF の古さ学習率スケーリング則
高スループットの RLHF システムでは、多くの場合、ロールアウトの生成とポリシーの最適化が切り離されており、学習者の更新中に古いロールアウトが使用されることになります。この研究では、非同期 GRPO におけるこのような古い状態の影響を研究します。 GRPO 代理目標で動作ポリシーを明示し、学習者が使用する代理勾配マッピングと、分布に依存する母集団目標の真の合計導関数を区別します。局所的な有界性、分布の滑らかさ、および動作ポリシーの滑らかさの仮定の下で、古いロールアウトによって次数 O(S * eta) のステップごとの代理勾配バイアスが導入されることを示します。ここで、S は最大ロールアウト ラグを示し、η は学習率を示します。さらに、条件付き崩壊時間スケーリング則を導き出します。サイクル内のドリフトがバッチレベルのクリッピング半径未満にとどまる場合、崩壊は主に累積学習者ドリフト T * ηによって支配されます。 stale-rollout 制約がアクティブな場合、安定性は代わりに S * eta に明示的に依存します。これにより、2 つの制約の安定性条件 eta << min{R_batch / (S * G_upd), R_crit / (T * G_upd)} が得られ、最大安定学習率がホライズン制限領域における古さに弱く依存しているように見える理由を説明します。
原文 (English)
Staleness-Learning Rate Scaling Laws for Asynchronous RLHF
High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GRPO. We make the behavior policy explicit in the GRPO surrogate objective and distinguish between the surrogate-gradient mapping used by the learner and the true total derivative of a distribution-dependent population objective. Under assumptions of local boundedness, distributional smoothness, and behavior-policy smoothness, we show that stale rollouts introduce a per-step surrogate-gradient bias of order O(S * eta), where S denotes the maximum rollout lag and eta denotes the learning rate. We further derive a conditional collapse-time scaling law: when within-cycle drift remains below a batch-level clipping radius, collapse is governed primarily by cumulative learner drift T * eta; when the stale-rollout constraint is active, stability instead depends explicitly on S * eta. This yields a two-constraint stability condition eta << min{R_batch / (S * G_upd), R_crit / (T * G_upd)}, explaining why the maximum stable learning rate may appear weakly dependent on staleness in the horizon-limited regime.
LongVQUBench: 長期ビデオ品質のベンチマーク 視覚言語モデルの理解
長期的なビデオ品質の理解の評価は、大規模ビジョン言語モデル (LVLM) にとって未解決の課題のままです。既存のビデオ品質ベンチマークは、主に短いクリップと孤立した歪みに焦点を当てており、時間的な連続性、累積的な劣化、および長時間コンテンツに固有の推論の複雑さを見落としています。これらの制限に対処するために、長期的なビデオ品質を理解するための包括的なベンチマークである LongVQUBench を紹介します。 LongVQUBench には、映画、ドキュメンタリー、監視映像、自己中心的な録画、アニメーション コンテンツにわたる 1,200 を超える多様なビデオが含まれており、検証とテストのための 1,500 の多肢選択式の自由形式の質問も含まれています。さまざまな時間的範囲にわたる知覚推論を評価するために、段階的に複雑になる 3 つの評価レベルを導入します。(i) 局所的な歪みを分析するためのローカル イベント品質理解 (LQU)。 (ii) 複数の劣化したイベントを統合するためのクロスイベント品質推論 (CQR)。 (iii) 長期間にわたる全体的な知覚評価のためのグローバル品質理解 (GQU)。さらに、ニードル ディストーション質問応答 (NDQA) パラダイムが 3 つのレベルすべてに組み込まれており、空間的または時間的アーティファクトがまばらに挿入されて、きめの細かい検出および推論機能が調査されます。 14 の最先端の LVLM に関する広範な実験により、ビデオの長さと推論の深さが増加するにつれてパフォーマンスが大幅に低下することが明らかになり、長距離の時間統合と知覚的帰属に対する能力の限界が浮き彫りになりました。私たちは、LongVQUBench が、LVLM の長期的なビデオ品質の理解を体系的かつ階層的に説明可能な評価に向けた基礎的なステップとして構想しています。
原文 (English)
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, and reasoning complexity inherent in long-duration content. To address these limitations, we present LongVQUBench, a comprehensive benchmark for long-term video quality understanding. LongVQUBench contains over 1200 diverse videos spanning movies, documentaries, surveillance footage, egocentric recordings, and animated content, accompanied by 1500 multiple-choice and open-ended questions for validation and testing. To assess perceptual reasoning across different temporal scopes, we introduce three progressively complex evaluation levels: (i) local event quality understanding (LQU) for analyzing localized distortions; (ii) cross-event quality reasoning (CQR) for integrating multiple degraded events; and (iii) global quality understanding (GQU) for holistic perceptual evaluation over extended durations. Furthermore, a needle distortion question-answering (NDQA) paradigm is embedded across all three levels, where spatial or temporal artifacts are sparsely inserted to probe fine-grained detection and reasoning capabilities. Extensive experiments on 14 state-of-the-art LVLMs reveal significant performance degradation with increasing video length and reasoning depth, highlighting their limited capacity for long-range temporal integration and perceptual attribution. We envision LongVQUBench as a foundational step toward the systematic, hierarchical, and explainable evaluation of LVLMs' long-term video quality understanding.
安価なコード、コストのかかる判断: ガバナンス可能なエージェント ソフトウェア エンジニアリングに関するケース スタディ
生成 AI は、ソフトウェア エンジニアリングを、少ない実装労力を中心に組織された実践から、豊富で低コストのコード生成を中心に組織された実践へと移行させています。この変化は、エンジニアリングの中心となる問題を変えます。AI が有用なコードを生成できるかどうかではなく、AI を介した開発が検査可能、修正可能、保守可能であり続けるように、エンジニアがアーキテクチャ、ツール、証拠、フィードバック ループをどのように編成するかです。私たちは、一人称のケーススタディを通じてこの問題を研究します。この開発作業では、1 人の専門ソフトウェア エンジニアがフロンティア AI コーディング エージェントを使用してドキュメント アクセシビリティ修復システムを構築しました。12 週間の開発作業です。経験的な記録は、88 件の同時期のフィールド ノート、420 KLOC の製品コード、および 1.16 MLOC のテスト、lint、サポート ドキュメント、およびエージェント ツールで構成されています。この記録から、高速エージェントの実装がどのようにしてガバナンス可能になるかを説明するプロセス モデルとして表現される、ガバナンス変換の中範囲理論の候補を開発します。このモデルは、エージェント実装の速度が繰り返し起こる構造的故障クラスをどのように表面化するか、そしてエンジニアリングの判断がそれらの故障を耐久性のあるガバナンス メカニズムに変換することでどのように速度を維持するかを説明します。既知の義務から制御を導き出す既存のガバナンス モデルとは対照的に、ガバナンス変換では、エージェントの作業中にのみ表示される障害から制御がどのように発見されるかを説明します。私たちはモデルを使用してテスト可能な予測を行い、ソフトウェア エンジニアリングの研究と実践への影響を説明します。
原文 (English)
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low-cost code production. This shift changes the central engineering problem: not whether AI can generate useful code, but how engineers organize architectures, tools, evidence, and feedback loops so that AI-mediated development remains inspectable, correctable, and maintainable. We study this problem through a first-person case study: a 12-week development effort in which a single expert software engineer used frontier AI coding agents to build a document accessibility remediation system. The empirical record comprises 88 contemporaneous field notes, 420 KLOC of production code, and 1.16 MLOC of tests, lints, supporting documentation, and agent tooling. From this record, we develop a candidate middle-range theory of governance conversion, expressed as a process model explaining how high-velocity agentic implementation becomes governable. The model explains how agentic implementation velocity surfaces recurring structural failure classes, and how engineering judgment sustains velocity by converting those failures into durable governance mechanisms. In contrast to existing governance models that derive controls from known obligations, governance conversion explains how controls are discovered from failures that become visible only during agentic work. We use our model to make testable predictions and to describe implications for software engineering research and practice.
CausalMix: 言語モデルトレーニングのための因果推論としてのデータ混合
大規模言語モデル (LLM) トレーニングでは、データ混合がモデルのパフォーマンスを決定する上で極めて重要な役割を果たします。最近の方法は、プロキシ モデルを介して混合重みを最適化しますが、静的なデータ分布の仮定に依存しています。その結果、基盤となるデータ プールが変化すると、これらの方法ではコストがかかるゼロからの再トレーニングが必要になります。この制限により、小さな設定からより大きなデータ プールやモデル サイズまでシームレスに拡張する能力が制限されます。この論文では、データ混合の最適化を因果推論問題としてキャストすることで、この制限に対処する CausalMix を提案します。データプールの統計的特徴を共変量として定式化し、ドメイン混合を処理として定式化します。 Qwen2.5-0.5B の 512 回の実行に因果モデルを当てはめて条件付き平均治療効果 (CATE) を推定した後、800K データ プールの最適な混合を外挿し、それを 7B モデルのトレーニングに適用します。さらに、このフレームワークを Qwen3-4B-Base 上の長い思考連鎖データに一般化することに成功しました。 CausalMix は、因果モデリングを活用して交絡バイアスを分離することで、状態に依存した最適なデータ混合を動的に推論します。広範な実験により、CausalMix によって導かれた混合は、複数の下流タスク全体でパフォーマンスを一貫して向上させ、RegMix や他のベースラインを上回るパフォーマンスを示すことが示されています。さらに、CATE インタプリタを使用して、学習したミキシング戦略を視覚的に分析します。全体として、CausalMix は、LLM データ混合を最適化するための因果的で解釈可能なフレームワークを提供します。
原文 (English)
CausalMix: Data Mixture as Causal Inference for Language Model Training
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose CausalMix to address this limitation by casting data mixture optimization as a causal inference problem. We formulate the statistical features of the data pool as covariates and the domain mixture as the treatment. After fitting a causal model on 512 runs of Qwen2.5-0.5B to estimate the Conditional Average Treatment Effect (CATE), we extrapolate the optimal mixture for an 800K data pool and apply it to train a 7B model. Furthermore, we successfully generalize the framework to long chain-of-thought data on Qwen3-4B-Base. By leveraging causal modeling to isolate confounding biases, CausalMix dynamically infers state-dependent optimal data mixtures. Extensive experiments show that the mixture guided by CausalMix consistently improves performance across multiple downstream tasks, outperforming RegMix and other baselines. In addition, we use the CATE Interpreter to provide visual analysis of the learned mixing strategy. Overall, CausalMix offers a causal and interpretable framework for optimizing LLM data mixtures.
FAR: テスト時のリカバリと継続的なポリシー改善のための障害認識再試行
ロボット ポリシーは、実際の環境に展開すると必ず失敗に遭遇します。単純な再試行では同じ間違いが繰り返されることがよくありますが、既存の回復方法の多くは人間の介入に依存しています。この論文では、ロボットがテスト時に以前の失敗から学習し、それに応じて動作を適応させ、最終的に自律的にタスクを完了できるようにするフレームワークである、Failure-Aware Retry (FAR) を提案します。 FAR は、失敗から優先学習データを構築して以前に失敗した動作からポリシーを誘導する失敗対比優先適応と、ローカル探索を促進する再試行中の軽量アクション摂動を組み合わせます。さらに、成功した回復の軌跡を継続的な政策改善のためのトレーニング ループに組み込みます。シミュレーションと現実世界の操作タスクの両方での実験では、FAR が成功率と堅牢性を大幅に向上させ、シミュレーションでは標準の拡散ポリシーと比べて平均 17.6%、現実世界では 11.7% 向上することが示されています。さらに、FAR は、有益な障害ケースを活用することで、継続的なポリシーの改善中に、リセットとタイムステップの両方のバジェットの下でデータ効率を大幅に向上させます。
原文 (English)
FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many existing recovery methods rely on human intervention. In this paper, we propose Failure-Aware Retry (FAR), a framework that enables robots to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete the task autonomously. FAR combines Failure-Contrastive Preference Adaptation, which constructs preference learning data from failures to steer the policy away from previously unsuccessful behaviors, with lightweight action perturbations during retries to encourage local exploration. We further incorporate successful recovery trajectories into a training loop for continual policy improvement. Experiments in both simulation and real-world manipulation tasks show that FAR substantially improves success rates and robustness, with average gains of 17.6% over the standard diffusion policy in simulation and 11.7% in the real world. In addition, FAR significantly improves data efficiency under both reset and timestep budgets during continual policy improvement by exploiting informative failure cases.
大学関係者向けのマルチモーダル チャット アシスタントの開発に向けて: RAG ベースのアプローチ
大学関係者は、特にインテリジェントなサポート システムがほとんどない発展途上国では、タイムリーで信頼できる情報にアクセスすることが困難であることがよくあります。既存のルールベースのチャットボットは、ドメイン固有の複雑なクエリを処理できず、進化する組織のポリシーに適応する機能が十分ではありません。ギャップを埋めるソリューションとして、検索拡張生成を備えたマルチモーダル大学チャットボットを紹介します。このシステムは、大規模な言語モデルと意味検索を組み合わせて、大学ハンドブックなどの教育機関中心のリソースからコンテキストベースの応答を生成します。このシステムは、ビジョン言語モデルを通じてテキストと画像のクエリを受け入れ、量子化された推論を適用して、制約のあるハードウェアに迅速に展開します。 FastAPI で構築されたスケーラブルなバックエンドと Next.js で開発された応答性の高いフロントエンドが連携することで、リアルタイムの使いやすさが保証されます。マルチモーダル評価では、視覚入力に対する応答時間が増加したにもかかわらず、システムがテキストと画像の両方のクエリにわたって高い満足度スコアを維持していることが実証されました。さらに、定量的評価では、提案したRAGベースのシステムでは幻覚が31.7%から6.6%に減少することが示されており、回収接地の有効性が確認されています。
原文 (English)
Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach
University stakeholders often face difficulties in accessing timely and reliable information, especially in developing countries, where there are very few intelligent support systems. Existing rule-based chatbots are unable to handle complex, domain-specific queries and are not well-equipped to adapt to evolving institutional policies. As a fill-in-the-gap solution, we present the multimodal university chatbot with retrieval-augmented generation. The system combines the large language model with semantic retrieval to produce context-based responses from institution-centric resources, such as the university handbook. The system accepts text and image queries through the vision-language model and applies quantized inference for rapid deployment on constrained hardware. A scalable backend built with FastAPI, adjoined with a responsive frontend developed with Next.js, ensures real-time usability. Our multimodal evaluation demonstrates that the system maintains strong satisfaction scores across both text and image queries, despite increased response time for visual inputs. Furthermore, quantitative evaluation shows that hallucination is reduced from 31.7% to 6.6% in our proposed RAG-based system, confirming the effectiveness of retrieval grounding.
残留結合としてのミュオン
Muon は最近、大規模なニューラル ネットワークをトレーニングするための最も効果的なオプティマイザーの 1 つとして浮上していますが、その経験的な成功はいくつかの異なる観点から説明されています。この論文では、単純な機構的解釈を提案します。ミューオンは、トレーニング中の暗黙的な残留結合として理解できます。具体的には、更新を直交化すると、下流のレイヤーの表現の保存が向上する一方で、即時の勾配の忠実度が犠牲になる可能性があります。私たちは、制御された線形最適化設定でこのトレードオフを研究します。この設定では、Muon は、ローカル ターゲットに適合するには遅いが、下流の層が利用しやすい表現を学習できます。私たちの結果は、Muon の概念的な説明と、ローカル降下と下流のユーザビリティのバランスをとるオプティマイザーの設計の観点を示唆しています。
原文 (English)
Muon as a Residual Connection
Muon has recently emerged as one of the most effective optimizers for training large neural networks, yet its empirical success has been explained from several different perspectives. In this paper, we propose a simple mechanistic interpretation: Muon can be understood as an implicit residual connection during training. Specifically, orthogonalizing the update can sacrifice some immediate gradient fidelity while improving representation preservation for downstream layers. We study this trade-off in controlled linear optimization settings, where Muon can learn representations that are slower to fit a local target but easier for downstream layers to exploit. Our results suggest a conceptual explanation for Muon and a design perspective for optimizers that balance local descent with downstream usability.
反復的なメタ反射による自律的な科学的発見
自律型科学発見システムは、仮説の生成と検証のプロセスを自動化することで研究を加速する可能性をもたらします。しかし、現在のシステムは、制約された検索スペース内で動作するか、事前に定義された調査質問を必要とするため、真に自由な調査を行う能力が制限されています。さらに、仮説は繰り返し生成されますが、独自に蓄積された発見を明示的に統合して、複雑で相互に関連した現象を明らかにする能力がほとんどありません。 DiscoPER は、事前に指定された研究目的なしでデータセットを探索するコードを動的に生成および実行することで、オープンエンドの研究を行う、自律的な大規模言語モデルを利用したフレームワークです。厳密な科学的妥当性を確保するには、提案されたすべての発見が統計的テストに合格する必要があります。孤立した検索の限界を克服するために、私たちのフレームワークは、独自に蓄積された発見を定期的に分析する二次推論メカニズムを導入しています。 DiscoPER は、以前の発見を経験的データとして扱うことにより、構造パターン、交絡、認識論的ギャップを特定し、仮説探索を探索空間の未知の領域に向けて積極的に方向転換します。ツールの使用を組み込むことで検索空間がさらに拡張され、画像などのマルチモーダル ソースから有用な情報をシームレスに処理および抽出することで、システムが構造化メタデータを超えた仮説を探索できるようになります。ピアレビューされた文献から得られたパターンレベルのグラウンドトゥルースを備えた新しいマルチモーダル生態学的知識ベンチマークである iNatDisco で評価したところ、DiscoPER は 72.7% の仮説支持率で 9 つの既知パターンのうち 8 つを回復し、古典的な因果関係の発見と LLM に基づくベースラインの両方を上回りました。アブレーションは、DiscoPER がより多くのデータで拡張できることを示し、二次メタ反射の利点を確認します。
原文 (English)
Autonomous Scientific Discovery via Iterative Meta-Reflection
Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation. However, current systems operate within constrained search spaces or require predefined research questions, limiting their capacity for true open-ended inquiry. Furthermore, while they generate hypotheses iteratively, they largely lack the ability to explicitly synthesize their own accumulated findings to uncover complex, interconnected phenomena. We introduce DiscoPER, an autonomous large language model-powered framework that conducts open-ended research by dynamically generating and executing code to explore datasets without pre-specified research objectives. To ensure rigorous scientific validity, every proposed discovery must pass statistical testing. To overcome the limitations of isolated search, our framework introduces a second-order reasoning mechanism that periodically analyzes its own accumulated discoveries. By treating prior discoveries as empirical data, DiscoPER identifies structural patterns, confounds, and epistemic gaps, actively redirecting hypothesis exploration toward uncharted regions of the search space. The search space is further expanded by incorporating tool use, enabling the system to explore hypotheses beyond structured metadata by seamlessly processing and extracting useful information from multimodal sources like images. Evaluated on iNatDisco, a new multimodal ecological knowledge benchmark with pattern-level ground truth obtained from peer-reviewed literature, DiscoPER recovers 8 of 9 known patterns with a 72.7% hypothesis support rate, outperforming both classical causal discovery and LLM-guided baselines. Ablations show that DiscoPER scales with more data, and confirms the benefits of second-order meta-reflection.
スキルは孤立したものではない: エージェント スキルのサプライ チェーンにおける依存関係とリスクの測定
エージェント スキルは、大規模言語モデル (LLM) エージェントの再利用可能な運用知識をパッケージ化したものですが、範囲が拡大するにつれて、依存関係を伴う成果物となり、そのアイデンティティ、バージョン、来歴が暗黙的なままになります。この不透明さにより、すでに依存関係の重複やインストールの不一致が発生しており、依存関係管理がまだ埋めていないギャップが露呈しています。エージェント スキル サプライ チェーン (ASSC) を導入して、スキル、パッケージ、サービスの混合依存関係グラフを特徴付け、このギャップを埋めるのに役立ちます。ソフトウェア部品表 (SBOM) から借用して、自然言語の依存関係の証拠を取得し、スキルを依存関係のある成果物としてモデル化するように SkillDepAnalyzer を設計します。 SKILL-DEP ベンチマークでは、SkillDepAnalyzer はスキルのメタデータと依存関係グラフを正確かつ包括的に回復し、LLM ベースのベースラインやパッケージ中心の SBOM ツールを大幅に上回ります。 SkillDepAnalyzer を 143 万を超えるスキルに適用して ASSC を取得し、その構造的多様性とセキュリティ シグナルを調査します。 4 つの構造パターンが見つかりました。スキルのメタデータはアクティベーションに対応していますが、ガバナンスが不十分です。依存関係グラフは、スキル、パッケージ、サービスの依存関係を集中的に再利用します。再帰的なスキルの再利用により、依存関係グラフが拡張され、非表示のパッケージ インベントリが作成されます。関連するワークフローを中心にスキル依存関係クラスターが形成されます。また、スキルだけを検査すると、その依存関係に隠れているセキュリティ関連のシグナルを見逃すこともわかりました。 ASSC を分析することで、ASSC に存続する既知の悪意のあるスキルを特定し、開発者に報告します。これらの調査結果に基づいて、型付き依存関係マニフェスト、ファーストクラスの依存関係クラスター管理、スキル インフラストラクチャ保守者向けのリスク警告監査コマンド、およびスキル開発者向けのロックファイルのようなレコードをお勧めします。
原文 (English)
Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains
Agent skills package reusable operational knowledge for Large Language Model (LLM) agents, yet as they grow in scope, they become dependency-bearing artifacts whose identities, versions, and provenance remain implicit. This opacity already causes duplicated dependencies and inconsistent installations, exposing a gap that dependency management has yet to close. We introduce Agent Skill Supply Chains (ASSCs) to characterize mixed skill-package-service dependency graphs and help close this gap. Borrowing from Software Bill of Materials (SBOMs), we design SkillDepAnalyzer to capture natural-language dependency evidence and model skills as dependency-bearing artifacts. On the SKILL-DEP benchmark, SkillDepAnalyzer recovers skill metadata and dependency graphs accurately and comprehensively, substantially outperforming an LLM-based baseline and package-centric SBOM tools. Applying SkillDepAnalyzer to over 1.43 million skills, we obtain ASSCs and explore their structural diversity and security signals. We find four structural patterns: skill metadata is activation-ready but governance-poor; dependency graphs span skill, package, and service dependencies with concentrated reuse; recursive skill reuse expands dependency graphs and creates hidden package inventory; and skill dependency clusters form around related workflows. We also find that inspecting a skill alone misses security-relevant signals hiding in its dependencies. By analyzing ASSCs, we identify and report known malicious skills persisting in ASSCs to their developers. Based on these findings, we recommend typed dependency manifests, first-class dependency-cluster management, risk-warning audit commands for skill infrastructure maintainers, and lockfile-like records for skill developers.
オンラインフィードバック駆動型検索のための、逐次制御されたインタラクティブなマルチ粒子フローマップ
生成モデルはトレーニング不要の報酬調整を可能にしましたが、現在の手法は通常、基礎となる分布の狭い領域内での局所探索に優れています。これらのアプローチは、好みがアプリオリに不明であり、連続的なフィードバックを通じてのみ明らかにされる場合、つまり、有用性の高い領域を発見するために広範な調査が必要なシナリオでは困難を伴います。これに対処するために、我々は、サンプル効率の高いオンラインフィードバック駆動型検索のフレームワークである Sequentially-Controlled Interactive Multi-Particle Flow-Maps (IMPFM) を提案します。 IMPFM は、インタラクティブな粒子のグループをターゲット分布に向けて徐々に輸送し、異種の優先順位の調整に不可欠な広い範囲を維持します。 IMPFM は、フロー マップを利用した、粒子間での原理的かつ効率的な事後サンプル共有メカニズムを導入します。各リサンプリング ステップでアンサンブル全体の集合的な事後サンプルを使用して個々の粒子のドリフトを補正することで、フレームワークはサンプルの有用性を最大化し、標準的な制御フレームワークに典型的な報酬の過剰最適化を積極的に軽減しながら、グローバルな探索を可能にします。複数粒子の相互作用を伴う原則に基づいた探査・利用再重み付けメカニズムと組み合わせることで、この逐次補正された複数粒子ダイナミクスにより、構造の多様性が明示的に保存され、標準的な SMC サンプラーに固有の重みの縮退が克服されます。重要なことは、結果として得られるサンプリング フレームワークによって、多粒子系を KL 傾斜ターゲット分布に向けて徐々に操縦し、全球探査を促進し、モード崩壊を防ぐ、多粒子相互作用を認識したファインマン・カック補正器が得られることを証明することです。広範な経験的評価と、さまざまな検索および位置合わせタスクにわたる厳密なアブレーションにより、既存のベースラインを超える IMPFM の有効性が確認されています。
原文 (English)
Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search
While generative models have enabled training-free reward alignment, current methods typically excel in local exploration within narrow regions of the underlying distribution. These approaches struggle when preferences are unknown a priori and only revealed through sequential feedback-a scenario demanding broad exploration to uncover high-utility regions. To address this, we propose Sequentially-Controlled Interactive Multi-Particle Flow-Maps (IMPFM), a framework for sample-efficient online feedback-driven search. IMPFM progressively transports a group of interactive particles toward the target distribution, maintaining the broad coverage essential for heterogeneous preference alignment. IMPFM introduces a principled and efficient posterior sample sharing mechanism across particles powered by flow maps. By correcting individual particle drift with the collective posterior samples of the entire ensemble at each resampling step, the framework maximizes sample utility to enable global exploration while actively mitigating reward over-optimization, typical of standard control frameworks. Paired with a principled exploration-exploitation reweighting mechanism involving multi-particle interaction, this sequentially corrected multi-particle dynamics explicitly preserves structural diversity and overcomes the weight degeneracy inherent to standard SMC samplers. Crucially, we prove that the resulting sampling framework yields a multi-particle interaction-aware Feynman-Kac corrector that progressively steers the multi-particle system toward a KL-tilted target distribution, facilitating global exploration and preventing mode collapse. Extensive empirical evaluations and rigorous ablations across diverse search and alignment tasks confirm the efficacy of IMPFM over existing baselines.
AI の安全性評価のための敵対的プラグマティクス: 命令の競合、埋め込みコマンド、およびポリシーの曖昧さのベンチマーク
言語モデルの安全性評価は、モデルが指示に従ったか、適切に拒否したか、ポリシーに従ったか、埋め込まれたコマンドに抵抗したか、エージェントタスクの進捗状況を誤って報告したかなど、あいまいな自然言語の動作に関する判断にますます依存しています。既存のベンチマークは、多くの場合、これらの区別を合格/不合格のラベルに圧縮し、障害が機能制限、ポリシーの曖昧さ、命令の競合、足場の障害、または不安定な評価者の判断に起因するかどうかを曖昧にします。この論文では、命令の競合、埋め込みコマンド、引用、範囲の曖昧さ、明確化、間接音声行為、およびマルチターンエージェントのトランスクリプトの下でのモデルの動作を評価するためのベンチマークおよびアノテーションプロトコルとして、敵対的プラグマティクスを紹介します。この貢献は経験的かつ方法論的です。言語的に管理された分類法、バリデーターが強制するメタデータを備えた 18 項目のシード ベンチマーク、54 行のローカル シード パイロット、タスクの成功、ポリシー遵守、安全性リスク、拒否結果、評価者の信頼性を区別する専門家評価プロトコル、および裁判官の妥当性、診断の曖昧さ、および分類のドリフトの指標です。このフレームワークは、言語的判断方法論を、安全性評価、LLM ジャッジ、ゴールドセット構造、即時注入テスト、および安全性文書を検証するための実用的なツールに変えます。
原文 (English)
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The framework turns linguistic judgment methodology into a practical tool for validating safety evals, LLM judges, gold-set construction, prompt-injection tests, and safety documentation.
Diffusion-GR2: 拡散生成推論リランカー
生成推論の再ランカーは、候補リストを並べ替える前に思考の連鎖を発行することで強力な推奨精度を実現しますが、推論が遅くなります。自己回帰 (AR) デコーダは推論トークンごとに 1 回の連続した前方パスを費やし、推論トレースは生成されるランキングをはるかに上回ります。このコストを削減するために、ブロック拡散言語モデルは、いくつかのノイズ除去ステップで多くの位置を並行してデコードし、大幅に高速になりますが、単純に AR リランカーを 1 つに変換すると、2 つの精度ギャップが生じます。 (1) 構造的なギャップ: 回答位置は並行してノイズ除去され、独立してスコアリングされるため、デコーダーは無効なランキング (重複、欠落、またはセット外の識別子) を生成しますが、AR はこれを左から右のマスキングによって回避します。 (2) 分布ギャップ: 固定教師軌道上で変換されたモデルを微調整することは、推論時の独自のデコードと比較してポリシーから外れており、精度ギャップが残ります。高速化を維持しながら両方のギャップを埋めるために、AR 推論リランカー (GR2) をブロック拡散リランカーに変換するレシピである \textbf{Diffusion-GR2} を提案します。まず、変換微調整 (CFT) は、AR で初期化された拡散モデルを適応させて、外部の制約付きデコーダーを使用せずに、独自に答えを有効な置換にノイズ除去します。次に、オンポリシー蒸留 (OPD) が、AR 教師からの高密度のトークンごとのターゲットを使用して、独自のデコードされた軌道でモデルを監視します。最後に、OPD のポリシーに関するポリシーに加えて、再ランキング報酬に対して強化学習 (RL) ステージを適用します。 Amazon Beauty での実験では、Diffusion-GR2 が AR リランカーとほぼ同等に回復し、ブロック並列デコードにより、モデルの推論出力長でデコード スループットが $2.4$ ~ $3.5\times$ 向上することが実証されました。アブレーションにより、CFT がコンバージョン ギャップのほとんどを回復し、ポリシーに基づいた蒸留により AR リファレンスにさらに近づくことが示されています。
原文 (English)
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces. To reduce this cost, block-diffusion language models decode many positions in parallel over a few denoising steps and are substantially faster, yet naively converting an AR re-ranker into one opens two accuracy gaps: (1) a structural gap: answer positions are denoised in parallel and scored independently, so the decoder emits invalid rankings (duplicated, dropped, or out-of-set identifiers) that AR avoids through left-to-right masking; and (2) a distributional gap: fine-tuning the converted model on fixed teacher trajectories is off-policy relative to its own decoding at inference, leaving a residual accuracy gap. To close both gaps while keeping the speedup, we propose \textbf{Diffusion-GR2}, a recipe that converts our AR reasoning re-ranker (GR2) into a block-diffusion re-ranker. First, conversion fine-tuning (CFT) adapts the AR-initialized diffusion model to denoise the answer into a valid permutation on its own, without an external constrained decoder. Next, on-policy distillation (OPD) then supervises the model on its own decoded trajectories with dense per-token targets from the AR teacher. Finally, we apply a reinforcement-learning (RL) stage against a re-ranking reward on top of OPD's on-policy policy. Experiments on Amazon Beauty demonstrate that Diffusion-GR2 recovers to near-parity with the AR re-ranker, while block-parallel decoding raises decode throughput by $2.4$--$3.5\times$ at the model's reasoning output length. Ablations show that CFT recovers most of the conversion gap, and that on-policy distillation further closes it to the AR reference.
正しい方法で: 検証可能な報酬と人間によるデモンストレーションを伴う LM トレーニング
検証可能な報酬付き RL (RLVR) は、コード生成や数学的推論など、明確に定義された成功指標を持つタスクで LM をトレーニングするための強力なパラダイムとして登場しました。ただし、現在の RLVR 手法は、客観的にスコア付けできるもののみを最適化し、スタイルや構造など、人間のような出力の主観的で検証不可能な側面を無視することがよくあります。この制限は、多様性の崩壊、不自然に聞こえる応答、報酬のハッキングなど、十分に文書化された障害モードにつながります。私たちは、人間のデモンストレーションから学習した信号で検証可能な報酬を強化する、敵対的なジェネレーターと識別子のフレームワークを提案します。ジェネレーター モデルは、タスクの精度と弁別子から得られる敵対的報酬の両方を最大化するために RL を使用してトレーニングされます。ジェネレーター ポリシーと並行してトレーニングされたディスクリミネーターは、人が書いた出力とモデルが生成した出力を区別する方法を学習します。ディスクリミネーターは、人間の出力分布の学習された代理として機能し、スカラー報酬として形式化するのが難しい生成の側面に関するフィードバックを提供します。バグ修正やオープンエンド生成を含むさまざまな領域にわたって、私たちのアプローチは、RLVR の精度向上を維持しながら、検証不可能なプロパティを一貫して改善します。バグ修正では、私たちの方法は、最終パフォーマンスに合わせながら、RLVR ベースラインと比較して編集距離が大幅に低いソリューションを生成します。ストーリー生成において、私たちの手法は勝率を大幅に向上させながら、多様でより人間らしいストーリーを生成します。また、単純な報酬ハッキング ベンチマークでは、私たちの方法は高いベンチマーク スコアを維持しながら、モデルの不正動作をほぼ排除します。これらの結果を総合すると、私たちのアプローチが RL と SFT を橋渡しし、タスクの検証可能プロパティと検証不可能なプロパティを共同で最適化するためのスケーラブルなパスを提供していることがわかります。
原文 (English)
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations
RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as code generation and mathematical reasoning. However, current RLVR methods optimize only what can be objectively scored, often neglecting subjective, non-verifiable aspects of human-like outputs, such as style and structure. This limitation leads to well-documented failure modes such as diversity collapse, unnatural-sounding responses, and reward hacking. We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations. A generator model is trained using RL to maximize both task accuracy and an adversarial reward derived from a discriminator. The discriminator, trained alongside the generator policy, learns to distinguish human-written outputs from model-generated ones. The discriminator serves as a learned proxy for the human output distribution, providing feedback on aspects of generation that are difficult to formalize as scalar rewards. Across diverse domains, including bug fixing and open-ended generation, our approach consistently improves non-verifiable properties while preserving the accuracy gains of RLVR. In bug fixing, our method produces solutions with significantly lower edit distance compared to RLVR baselines while matching end performance. In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like. And in a simple reward hacking benchmark, our method nearly eliminates model misbehavior while maintaining high benchmark scores. Together, these results show that our approach bridges RL and SFT, offering a scalable path toward jointly optimizing the verifiable and non-verifiable properties of a task.
動きからの世界: 単眼ビデオからの生成動的ガウス再構成
単眼ビデオから自由にレンダリング可能な動的 3D ガウス表現を生成する方法である World from Motion を紹介します。私たちのアプローチでは、入力カメラとターゲット カメラの両方の軌道に沿って外観、ジオメトリ、および 3D シーンの動きをエンコードする高密度のピクセル配置レンダリングでビデオ モデルを条件付けし、レンダリング アーティファクトを修正し、最初の再構成で欠落している領域を埋めます。このモデルをトレーニングするために、単眼再構成の特徴であるシミュレートされたアーティファクトを使用して、位置合わせされたマルチビュー ビデオのペアと動的 3DGS 表現のデータセットを構築します。テスト時には、新たに観察された領域やモーションを含むモデルの世代を蒸留して単一の一貫した高品質のダイナミック 3DGS に戻し、新しいビューの合成と基礎となる 3D モーションの両方を改善します。私たちの方法は、4D 再構成における新しい最先端技術を確立し、大きな視点変更とダイナミックな動きを伴う自然のビデオにシームレスに一般化します。
原文 (English)
World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video
We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, pixel-aligned renderings that encode appearance, geometry, and 3D scene motion along both input and target camera trajectories to correct rendering artifacts and fill in missing regions from an initial reconstruction. To train this model, we construct a dataset of aligned multiview video pairs and dynamic 3DGS representations, with simulated artifacts characteristic of monocular reconstruction. At test time, we distill the model's generations, including newly observed regions and motions, back into a single consistent, high-quality dynamic 3DGS, improving both novel-view synthesis and the underlying 3D motion. Our method sets a new state of the art in 4D reconstruction and seamlessly generalizes to in-the-wild videos with large viewpoint changes and dynamic motions.
非線形およびニューラル ネットワーク ダイナミクスのリアルタイムの堅牢な最適制御のための GPU 並列線形化誤差限界
この論文では、不確実な非線形システムに対するリアルタイムのロバストな最適制御について研究します。このシステムでは、線形時変 (LTV) 近似により計画が扱いやすくなりますが、ロバストな制約を満たすために健全な線形化誤差限界 (LEB) が必要になります。私たちは、非線形およびニューラル ネットワーク (NN) ダイナミクスの LTV 近似用に、厳密で微分可能な GPU 並列 LEB を開発します。解析ダイナミクスのために、標準的な間隔法よりも厳しいパスベースのヘシアン境界を導入します。 NN ダイナミクスの場合、NN 検証者が生成したアフィン緩和とローカル ヤコビアン補正を使用して、認定された LEB を導出します。我々は、GPU 並列システムレベル合成 LTV ベースのロバスト制御ソルバーを、タイトなゾノトピック不確実性伝播のための右可逆外乱行列と非ゼロ中心外乱セットを処理できるように拡張することで、これらの LEB と互換性を持たせるように適合させます。私たちの手法である GPUSLS-LEO は、線形化誤差を考慮した堅牢なフィードバック ポリシーのオンライン最適化を可能にし、厳密で正式に検証された到達可能なチューブを生成します。複雑な非線形および最大 168 状態次元の NN ダイナミクスにおいて、私たちの手法は GPU 上で最大 67 Hz のレートで堅牢な制御ポリシーを計算でき、形式的な保証とリアルタイムのパフォーマンスを維持しながら、解決時間とベースラインに対する保守性を削減します。
原文 (English)
GPU-Parallel Linearization Error Bounds for Real-Time Robust Optimal Control of Nonlinear and Neural Network Dynamics
This paper studies real-time robust optimal control for uncertain nonlinear systems, where linear time-varying (LTV) approximations make planning tractable but require sound linearization error bounds (LEBs) to guarantee robust constraint satisfaction. We develop tight, differentiable, GPU-parallel LEBs for LTV approximations of nonlinear and neural network (NN) dynamics. For analytic dynamics, we introduce path-based Hessian bounds that are tighter than standard interval methods. For NN dynamics, we derive certified LEBs using NN verifier-generated affine relaxations and local Jacobian corrections. We adapt a GPU-parallel system-level synthesis LTV-based robust control solver to be compatible with these LEBs by extending it to handle right-invertible disturbance matrices and non-zero-centered disturbance sets for tight zonotopic uncertainty propagation. Our method, GPUSLS-LEO, enables online optimization of robust feedback policies that account for linearization error, producing tight, formally verified reachable tubes. On complex nonlinear and NN dynamics up to 168 state dimensions, our method can compute robust control policies on the GPU at rates up to 67 Hz, reducing solve times and conservativeness relative to baselines while preserving formal guarantees and real-time performance.
蒸留して検出: カートリッジ蒸留による LLM のステルス バイアスの暴露
一か八かの役割に導入された言語モデルは、特定のエンティティ、ブランド、または視点を優先して、ユーザーの意思決定を大規模に導く可能性があります。このような優先バイアスは、モデルのサプライチェーン内のどのアクターによっても導入される可能性があり、モデルが他のすべての入力に対しては変更されていないベースと同様に動作しながら、関連するトピックについてのみその優先順位を明らかにする場合に最も危険です。最近の研究では、これらのバイアスは、意味的に無関係なデータのコンテキスト蒸留を通じて伝達される可能性があり、信号は完全にソフト ロジット分布内に存在し、テキストベースの検査では見えないままであることが示されています。しかし、防御側は根本的な非対称性に直面しています。バイアスのトピックを知らなければ、生成されたテキスト、内部表現、モデルの重みを調べるかどうかに関係なく、どの検出方法もステルス優先バイアスを確実に表面化することはできません。ここでは、疑わしいモデルとそのベースの間の分布シフトをカートリッジ (KV キャッシュ プレフィックス アダプター) に蒸留し、支配的な発散を集中させ、バイアス信号を生成されたテキストに増幅することによって、隠れたバイアスを表面化する方法である Distill to Detect (D2D) を紹介します。 D2D がステルス モデルの隠れたバイアスを、複数のバイアス タイプにわたって確実に検出できる程度まで増幅することに成功したことを示します。また、経験的観察によって裏付けられた、ロジット分布シフトのフィッシャー重み付け投影のレンズを通して D2D の有効性を説明する理論的枠組みも提案します。 D2D は、プレフィックス チューニング アダプターの容量ボトルネックを検出ツールに変えることで、デプロイされた言語モデルの隠れた動作を監査するための実用的な構成要素を提供します。
原文 (English)
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at scale. Such preferential biases can be introduced by any actor in the model's supply chain and are most dangerous when the model reveals its preference only on the relevant topic while behaving identically to its unmodified base on all other inputs. Recent work has shown that these biases can transfer through context distillation on semantically unrelated data, with the signal residing entirely in the soft logit distribution and remaining invisible to text-based inspection. However, the defender faces a fundamental asymmetry: without knowing the bias topic, no detection method can reliably surface a stealth preferential bias, regardless of whether it examines generated text, internal representations, or model weights. Here we introduce Distill to Detect (D2D), a method that surfaces hidden biases by distilling the distributional shift between a suspected model and its base into a cartridge (a KV-cache prefix adapter), concentrating the dominant divergence and amplifying the bias signal into generated text. We show that D2D successfully amplifies the hidden biases of stealth models to the extent that they can be reliably detected across multiple bias types. We also propose a theoretical framework that explains the efficacy of D2D through the lens of Fisher-weighted projection of the logit distribution shift, supported by empirical observations. By turning the capacity bottleneck of prefix-tuning adapters into a detection tool, D2D provides a practical building block for auditing hidden behaviors in deployed language models.
パフォーマンス最適化ベンチマークはコーディング エージェントを確実に測定していますか?
GSO、SWE-Perf、SWE-fficiency などのリポジトリ レベルのパフォーマンス最適化ベンチマークは、実際のリポジトリにパッチを適用し、最適化されていないベースラインや公式リファレンス パッチとランタイムを比較することにより、コーディング エージェントを評価します。リーダーボードのスコアは、コーディング エージェントの進捗状況の証拠として使用されることが増えていますが、これらのスコアは、実行時の不安定性、ベンチマーク固有のスコア ルール、および少なくとも 1 つの公開提出によってすでに解決されているタスクの数を混同する可能性があります。私たちはこれらの問題を 3 つのベンチマーク全体で監査します。まず、4 つの一般的なタイプの Google Cloud マシンにわたる 740 コード最適化タスクの公式リファレンス パッチを再実行します。ほとんどのベンチマーク タスクは再実行できますが、参照パッチがすべてのクロスマシン再実行で元のベンチマーク有効性ルールを満たしているのは、39/102 個の GSO タスク、11/140 個の SWE-Perf タスク、および 411/498 個の SWE 効率タスクのみです。 SWE-Perf は、多くの参照パッチがランタイムの変更をほぼゼロにするため、特に脆弱です。第 2 に、公募ランキングがベンチマーク スコアリング ルールに強く依存していることを示します。 GSO と SWE-fficiency が共有する 8 つの公開提出物のうち、公式ランキングは 28 件のペアごとの提出物比較のうち 9 つで一致せず、SWE-fficiency のリーダーボードのスコアリング ルールでは、ワースト 10 のタスクに 58.5% ~ 82.8% という高すぎるスコアの重みが割り当てられています。 3 番目に、各タスクの 10 件の公開提出物を調べたところ、少なくとも 1 つの提出物が、リプレイ有効な GSO および SWE 効率タスクの 85.3% (384/450) でリファレンス パッチと一致またはそれを上回り、99.8% (449/450) で最適化されていないベース コードを上回っていることがわかりました。私たちの研究は、より信頼性の高いパフォーマンスシグナルを持つタスクを特定し、タスクごとのスコアへの寄与を定量化し、集計ランキングによって隠されている残りのパフォーマンスギャップを明らかにすることで、リーダーボードスコアを補完します。
原文 (English)
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already solved by at least one public submission. We audit these issues across the three benchmarks. First, we replay the official reference patches for 740 code optimization tasks across four common types of Google Cloud machines. Most benchmark tasks can be replayed, but their reference patches satisfy the original benchmark validity rules in every cross-machine replay for only 39/102 GSO tasks, 11/140 SWE-Perf tasks, and 411/498 SWE-fficiency tasks; SWE-Perf is especially fragile because many reference patches produce close-to-zero runtime changes. Second, we show that public submission rankings depend strongly on the benchmark scoring rule. Among eight public submissions shared by GSO and SWE-fficiency, the official rankings disagree on 9 of 28 pairwise submission comparisons, and SWE-fficiency's leaderboard scoring rule assigns the worst ten tasks overly high score weights of 58.5%-82.8%. Third, looking across 10 public submissions for each task, we find that at least one submission matches or beats the reference patch on 85.3% (384/450) of replay-valid GSO and SWE-fficiency tasks, and beats the unoptimized base code on 99.8% (449/450). Our study complements leaderboard scores by identifying tasks with more reliable performance signals, quantifying per-task score contributions, and exposing the remaining performance gaps that are hidden by aggregate rankings.
FurnitureVLA: 視覚-言語-行動モデルを使用した長期的な双手動家具組み立ての学習
ロボット家具の組み立てに関する現在の研究は、主におもちゃのスケールの設定やシングルアームの操作に焦点を当てています。私たちは、視覚・言語・行動モデル (VLA) を使用した、実際のスケールの両手による家具組み立ての最初の体系的な研究である FurnitureVLA を紹介します。私たちはタスクを形式化し、専門家によるデータの生成と評価のためのスケーラブルなシミュレーション パイプラインを開発し、高品質な現実世界のデモンストレーションを収集するためにシングル オペレーターの両手制御用の VR 遠隔操作システムを構築します。最大 7 つのサブタスクと 1550 の制御ステップによる非常に長期的なアセンブリに対処するために、意味論的に根拠のあるサブタスクに基づいて微調整された、進行状況を強化した VLA を提案します。これは、アクションと継続的な進行シグナルを共同で予測し、サブタスクの自動移行を可能にし、推論中の複合エラーを削減します。私たちは、実際のスケールの組み立ての精度に決定的に影響を与える知覚と制御の設計要素をさらに研究します。 FurnitureVLA は、3 種類の家具のベースラインと比較して平均シミュレーション成功率を 48% から 80% に向上させ、設計要素の調査からさらに 21% 向上しました。実際の Kinova Gen3 プラットフォームで検証したところ、最も困難なタスクでも 16% の低下しかありませんでした。
原文 (English)
FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the first systematic study of real-scale bimanual furniture assembly using Vision-Language-Action models (VLAs). We formalize the task, develop a scalable simulation pipeline for expert data generation and evaluation, and build a VR teleoperation system for single-operator bimanual control to collect high-quality real-world demonstrations. To address extreme long-horizon assembly with up to 7 subtasks and 1550 control steps, we propose a progress-enhanced VLA, finetuned on semantically grounded subtasks, that jointly predicts actions and a continuous progress signal, enabling automatic subtask transitions and reducing compounding errors during inference. We further study perception and control design factors that critically affect precision in real-scale assembly. FurnitureVLA improves average simulation success from 48% to 80% compared to baselines across three furniture types, with an additional 21% gain from our design factor study. We validate on a real Kinova Gen3 platform with only 16% drop on the hardest task.
状態予測分離仮説
トランスフォーマーは、同じ順方向計算ストリームを使用して、次のトークンを予測し、将来のトークン予測に役立つ状態を保存します。 \emph{状態予測分離仮説} を定式化します。2 つの役割を解きほぐすことで、言語モデリングのパフォーマンスが向上します。 2 つの計算ストリームを使用して 2 つの関数を分離する Transformer バリアントを設計し、さまざまなスケールで事前トレーニング実験を実施します。私たちの実験では、状態予測分離によりデータ効率と計算効率が常に向上し、検証損失が改善され、ダウンストリーム タスクで標準の Transformers よりも平均 2 ~ 3 パーセント ポイント優れたパフォーマンスが得られることが示されています。また、潜在的な交絡因子を排除し、設計に伴う勾配の根本的な違いを実証する広範な実証分析も実施します。
原文 (English)
The State-Prediction Separation Hypothesis
Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the \emph{state-prediction separation hypothesis}: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses two computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers better data and compute efficiencies, improving validation loss and outperforming standard Transformers by 2--3 percentage points on average on downstream tasks. We also conduct extensive empirical analysis that rules out potential confounders and demonstrates the fundamental difference in the gradients our design entails.
言語批評 最適ではないデモンストレーションから学ぶ模倣
最適ではないデモンストレーションからの模倣学習に関するこれまでの研究は、通常、信頼性推定値、識別子スコア、重要度の重みなどの圧縮された監視信号に依存していました。これらのスカラー信号は、タスクの進行状況、障害モード、または修正措置に関する中間推論を明示的に表現できないため、本質的に制限されています。我々は、次善のデモンストレーションからの模倣学習のための言語批評フレームワークを提案します。このフレームワークは、代わりに自然言語を構造化された監視信号として活用し、表現力豊かなフィードバックのスカラーへの崩壊を回避します。私たちの方法では、まず、現在の進行状況を明示的に説明し、最適ではない動作を特定し、きめ細かい修正ガイダンスを提供するデモンストレーションから言語ラベルを構築します。次に、これらの構造化シグナルをスカラーに還元せずに使用してポリシーを直接トレーニングする言語批判損失を導入し、動作クローニングと拡散ポリシーの両方のためにそれをインスタンス化し、LC-BC と LC-DP を生成します。さらに、提案された目標が標準的な仮定の下で専門家のパフォーマンスギャップの上限を超えることを示す理論的結果を提供します。経験的に、私たちはナビゲーション、操作、ゲームプレイにわたる多様な継続的制御タスクを評価しており、そこで私たちの手法は強力な模倣学習やオフライン強化学習のベースラインを常に上回っています。これらの結果は、言語が次善のデータから堅牢なポリシーを学習するための強力で構造化された監視形式として機能できることを示しています。
原文 (English)
Language-Critique Imitation Learning from Suboptimal Demonstrations
Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights. These scalar signals are inherently limited, as they cannot explicitly express intermediate reasoning about task progress, failure modes, or corrective actions. We propose a language-critique framework for imitation learning from suboptimal demonstrations that instead leverages natural language as a structured supervision signal, avoiding the collapse of expressive feedback into scalars. Our method first constructs language labels from demonstrations that explicitly describe current progress, identify suboptimal behaviors, and provide fine-grained corrective guidance. We then introduce a language-critique loss that directly trains policies using these structured signals without reducing them to scalars, and instantiate it for both behavior cloning and diffusion policies, yielding LC-BC and LC-DP. We further provide a theoretical result showing that the proposed objective upper-bounds the expert performance gap under standard assumptions. Empirically, we evaluate on diverse continuous control tasks spanning navigation, manipulation, and gameplay, where our methods consistently outperform strong imitation learning and offline reinforcement learning baselines. These results demonstrate that language can serve as a powerful and structured form of supervision for learning robust policies from suboptimal data.
人間の研究アイデアと LLM 研究アイデアの間のギャップを測定する
研究アイデアのブレインストーミングに LLM が使用されることが増えていますが、既存の評価では主に、新規性、実現可能性、または専門家の好みによって個々のアイデアが判断されます。私たちは代わりに、現在の LLM によって生成されたアイデアが人間の研究者からどれだけ離れているのかを尋ねます。このギャップを特徴づけるために、私たちは質の高い人間の研究論文から着想を得るための大規模な評価フレームワークを構築します。各論文について、その中心となるアイデアにインスピレーションを与えたと考えられる密接に関連した先行研究の小さなセットをリバース エンジニアリングします。次に、LLM は論文のタイトルと要約のセットから新しいアイデアを生成するように求められます。私たちは、2 軸のリサーチテイスト分類法を導入して、各アイデアをその機会パターンと研究パラダイムによってプロファイルし、それを使用して人間のアイデアと LLM のアイデアの相違を定量化します。さまざまな LLM によって生成されたアイデア セット全体で、一貫した分布ギャップが観察されます。LLM のアイデアは、橋渡しのような機会と合成方法の周りに不均衡に集中していますが、人間の論文参照の分布は、ギャップを構成し、貢献を構築する方法全体にわたってより広範囲に広がっています。この結果は、強力な LLM がさまざまな合理的なアイデアを生み出すことができるが、その範囲は依然として人間の研究の好みよりも狭く、体系的にシフトしていることを示唆しています。
原文 (English)
Measuring the Gap Between Human and LLM Research Ideas
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.
カントの視点から見た人工知能の判断と機能実装の不可解性
認識論の歴史への大きな貢献であるカントの『純粋理性批判』は、人間の判断の根底にあるアプリオリな原理の構造を解明するためのカテゴリーの表を提案しています。人工知能 (AI) テクノロジーは、人間の判断をシミュレートまたは複製すると主張しています。この主張を評価するには、AI の判断が人間の判断の本質的な特性を示しているかどうかを検討する必要があります。この論文は、カントの判断理論のレンズを通して、AI 判断の説明不可能性を調査します。この研究は、カントの 4 つの論理形式 (量、質、関係、様相) に基づいて、AI の不確実性、つまりさまざまな形式の判断が絡み合う状態を特定します。特にモダリティに関して、本研究はソフトマックス関数がAIの判断を可能性の判断として強制的に再構成していると主張する。さらに、この論文は、カントの定義の説明に基づいて、機能の実装を検証するための決定的な基準は存在しないと主張しています。さらに、重要な機能が存在しない場合でも、流暢な言語動作により、機能が実装されているように見えることがあります。
原文 (English)
Unexplainability of Artificial Intelligence Judgments and Functional Implementation in Kant's Perspective
Kant's Critique of Pure Reason, a major contribution to the history of epistemology, proposes a table of categories to elucidate the structure of the a priori principles underlying human judgment. Artificial intelligence (AI) technology claims to simulate or replicate human judgment. To evaluate this claim, it is necessary to examine whether AI judgments exhibit the essential characteristics of human judgment. This paper investigates the unexplainability of AI judgments through the lens of Kant's theory of judgment. Drawing on Kant's four logical forms - quantity, quality, relation, and modality - this study identifies what may be called AI's uncertainty, a condition in which different forms of judgment become entangled. In particular, with regard to modality, this study argues that the Softmax function forcibly reframes AI judgments as possibility judgments. Furthermore, drawing on Kant's account of definition, this paper argues that no definitive criterion exists for verifying functional implementation. Moreover, fluent linguistic behavior may create the appearance of functional implementation even when important functions remain absent.
サイロからシステムへ: AI システムのプロセス指向のハザード分析
人工知能 (AI) システムによる潜在的な危害に効果的に対処するには、システムレベルの危険を特定して軽減することが不可欠です。現在の分析アプローチは、トレーニング データやモデルなどの AI システムの個々のコンポーネントに単独で焦点を当てており、コンポーネントの相互作用や企業の開発プロセス内でコンポーネントがどのように配置されているかによる危険性を見落としています。この目的を達成するために、私たちは安全性をシステム全体の緊急の特性として考慮するシステム安全性の確立された分野を活用します。この研究では、システム安全性フレームワークとして認識されているシステム理論プロセス分析 (STPA) を、AI の開発および運用プロセスを分析するために翻訳します。私たちは機械学習アルゴリズムに依存するシステムに焦点を当て、線形回帰、強化学習、トランスフォーマーベースの生成モデルを含む 3 つのケーススタディに対して STPA を実施します。私たちの分析では、STPA の制御とシステム理論の観点が AI システムにどのように適用されるか、また、モデルの不透明性、機能の不確実性、出力の複雑さなどの AI の固有の特性によりフレームワークの変更が必要かどうかを調査しました。 STPA を実施するための重要な概念と手順は AI システムに当てはまりますが、3 つのケーススタディで程度の差はあれ、AI 固有の課題に対処するには的を絞った適応が必要であることがわかりました。 STPA の概念を AI に適応させるガイドラインとして、AI システムのためのプロセス指向ハザード分析 (PHASE) を紹介します。 PHASE ガイドラインを使用して STPA を適用および解釈すると、AI システムの危害の管理を担当するアナリストに 4 つの重要なアフォーダンスが可能になります。1) 異種の問題の蓄積による危害を含む、システムレベルの危害の検出。 2) アルゴリズムによる危害に寄与する社会的要因の明示的な認識。 3) 危害とそれを軽減できる者との間の追跡可能な責任の連鎖の構築。 4) 継続的な監視と新たな危険の軽減。
原文 (English)
From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems
To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level hazards. Current analysis approaches focus on individual components of an AI system, like training data or models, in isolation, overlooking hazards from component interactions or how they are situated within a company's development process. To this end, we draw from the established field of system safety, which considers safety as an emergent property of the entire system. In this work, we translate System Theoretic Process Analysis (STPA) - a recognized system safety framework - for analyzing AI development and operation processes. We focus on systems that rely on machine learning algorithms and conduct STPA on three case studies involving linear regression, reinforcement learning, and transformer-based generative models. Our analysis explored how STPA's control and system-theoretic perspectives apply to AI systems and whether unique AI traits - such as model opacity, capability uncertainty, and output complexity - necessitate modifications to the framework. We find that the key concepts and steps of conducting an STPA apply to AI systems but require targeted adaptations to address AI-specific challenges that arise to differing degrees across three case studies. We present the Process-oriented Hazard Analysis for AI Systems (PHASE) as a guideline that adapts STPA concepts for AI. Applying and interpreting STPA using the PHASE guidelines enables four key affordances for analysts responsible for managing AI system harms: 1) detection of system-level hazards, including those from accumulation of disparate issues; 2) explicit acknowledgment of social factors contributing to algorithmic harms; 3) creation of traceable accountability chains between harms and those who can mitigate them; and 4) ongoing monitoring and mitigation of new hazards.
LLM の強化学習における効果的かつ多様な探索のための選択的専門家ガイダンス
検証可能な報酬による強化学習 (RLVR) は、大規模言語モデル (LLM) の推論能力を強化するために広く採用されている手法です。ただし、RLVR の有効性は、ベース モデルの機能に大きく依存します。この問題は、モデルが効率性と多様性の両方を伴う高品質の探査を実行するのに十分な機能を備えている必要があるために発生します。残念ながら、既存の手法は専門家の軌跡を模倣することでこの問題に対処しており、有効性は向上しますが、多様性は無視されています。これに対処するために、専門家は推論経路全体ではなく、重要な決定点でのみガイダンスを提供する必要があると主張します。この洞察に基づいて、私たちは MENTOR: Mixed-policy Expert Navigation for Token-level Optimization of Reasoning を提案します。これは、RLVR で効果的かつ多様な探索を実行するために、重要な意思決定ポイントでのみ専門家によるガイダンスを提供するフレームワークです。広範な実験により、MENTOR を使用すると、モデルが表面的な模倣ではなく専門家の戦略の本質を捉え、それによって高品質の探索を実行し、優れた全体的なパフォーマンスを達成できることがわかりました。私たちのコードはオンラインで入手できます。
原文 (English)
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large Language Models (LLMs). However, the effectiveness of RLVR strongly depends on the capability of base models. This issue arises because it requires the model to have sufficient capability to perform high-quality exploration, which involves both effectiveness and diversity. Unfortunately, existing methods address this issue by imitating expert trajectories, which improve effectiveness but neglect diversity. To address this, we argue that the expert only needs to provide guidance only at critical decision points rather than the entire reasoning path. Based on this insight, we propose MENTOR: Mixed-policy Expert Navigation for Token-level Optimization of Reasoning, a framework that provides expert guidance only at critical decision points to perform effective and diverse exploration in RLVR. Extensive experiments show that MENTOR enables models capture the essence of expert strategies rather than surface imitation, thereby performing high-quality exploration and achieving superior overall performance. Our code is available online.
大規模な言語モデルは、ゲーム理論の実験全体で人間の協力を再現し、予測します
大規模言語モデル (LLM) は、一か八かの分野での意思決定エージェントとして、また社会科学および行動科学における人間の行動の模倣者として、ますます導入されています。しかし、LLM が人間の意思決定をどの程度反映しているのかは、まだよくわかっていません。このギャップは重要です。調整のずれは実際に有害な結果を生み出す可能性がありますが、人間の行動を再現できないと、LLM は社会シミュレーターとして無効になってしまいます。ここでは、大規模なゲーム理論実験を再現し、機械の動作評価のための体系的なプロンプトおよび精査フレームワークを導入することによって、このギャップに対処します。通常、エージェントを強化するために使用される 3 つのオープン モデル (Llama、Mistral、Qwen) をテストします。 4 つの古典的なゲーム タイプにわたる 121 のダイアディック ゲームにわたって、Llama は人間の協力パターンを高い忠実度で再現しますが、Qwen はナッシュ均衡の予測と密接に一致します。行動表現型解析を通じてモデルを特徴付けると、人間とラマは嫉妬深い意思決定プロファイルを共有する一方、クウェンとミストラルは異なるプロファイルを示すことがわかりました。ペイオフ顕著性の注意ベースの分析により、Llama はクウェンやミストラルには存在しない、構造化されたレイヤー依存の方法でペイオフ情報を処理することが明らかになり、人間の行動とより密接に一致するための機構的基盤が示唆されています。集団レベルの行動の複製は、ペルソナベースのプロンプトなしで実現され、シミュレーション プロセスが簡素化されます。実験パラメータ空間を人間がテストした元のゲームを超えて拡張し、新しいゲーム構成に対するテスト可能な仮説を生成して事前登録します。私たちの発見は、適切に構成されたLLMが人間の集合的な行動パターンを複製し、人間のような意思決定表現型を示し、未踏の実験空間の系統的な探索を可能にし、人間の社会的意思決定に関する新しい経験的予測を生み出す従来の行動研究への補完的なアプローチを提供することを示しています。
原文 (English)
Large language models replicate and predict human cooperation across experiments in game theory
Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences. Yet how closely LLMs mirror human decision-making remains poorly understood. This gap is critical: misalignment could produce harmful outcomes in practice, while failure to replicate human behavior renders LLMs ineffective as social simulators. Here, we address this gap by replicating large-scale game-theoretic experiments and by introducing a systematic prompting and probing framework for machine-behavioral evaluation. We test three open models typically used to power agents (Llama, Mistral, and Qwen). Across 121 dyadic games spanning four classical game types, Llama reproduces human cooperation patterns with high fidelity, while Qwen aligns closely with Nash equilibrium predictions. Characterizing models through behavioral phenotyping, we find that humans and Llama share an envious decision profile, while Qwen and Mistral exhibit different profiles. An attention-based analysis of payoff salience reveals Llama processes payoff information in a structured, layer-dependent manner absent in Qwen and Mistral, suggesting a mechanistic basis for its closer alignment with human behavior. Population-level behavioral replication is achieved without persona-based prompting, simplifying the simulation process. Extending the experimental parameter space beyond the original human-tested games, we generate and preregister testable hypotheses for novel game configurations. Our findings demonstrate appropriately configured LLMs can replicate aggregate human behavioral patterns, exhibit human-like decision phenotypes, and enable systematic exploration of unexplored experimental spaces, offering a complementary approach to traditional behavioral research that generates new empirical predictions about human social decision-making.
CoT-X: クロスモデルの思考連鎖の転送と最適化のための適応フレームワーク
思考連鎖 (CoT) 推論は、大規模言語モデル (LLM) の問題解決能力を強化しますが、かなりの推論オーバーヘッドが発生し、リソースに制約のある設定での展開が制限されます。この論文では、適応推論要約フレームワークを通じて、さまざまなスケールとアーキテクチャのモデル間での効率的な CoT 転送を調査します。提案された方法は、重要度スコアリングによるセマンティック セグメンテーション、予算を考慮した動的圧縮、および一貫性の再構築を介して推論トレースを圧縮し、トークンの使用量を大幅に削減しながら重要な推論ステップを維持します。 10 の専門分野にわたる 7{,}501 の健康診断の質問に関する実験では、同じトークン予算の下で切り捨てよりも最大 40% 高い精度が示されました。 8 つの LLM (DeepSeek-R1 および Qwen3 を含む 1.5B ~ 32B パラメーター) からの 64 のモデル ペアの評価により、強力なモデル間移行性が確認されました。さらに、ガウス プロセス ベースのベイジアン最適化モジュールにより、評価コストが 84% 削減され、モデル サイズとクロスドメインの堅牢性の間のべき乗則の関係が明らかになります。これらの結果は、推論の要約が効率的な CoT 転送への実用的な道を提供し、厳しい計算制約下で高度な推論を可能にすることを示しています。コードは公開され次第公開されます。
原文 (English)
CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
Chain-of-Thought (CoT) reasoning enhances the problem-solving ability of large language models (LLMs) but leads to substantial inference overhead, limiting deployment in resource-constrained settings. This paper investigates efficient CoT transfer across models of different scales and architectures through an adaptive reasoning summarization framework. The proposed method compresses reasoning traces via semantic segmentation with importance scoring, budget-aware dynamic compression, and coherence reconstruction, preserving critical reasoning steps while significantly reducing token usage. Experiments on 7{,}501 medical examination questions across 10 specialties show up to 40% higher accuracy than truncation under the same token budgets. Evaluations on 64 model pairs from eight LLMs (1.5B-32B parameters, including DeepSeek-R1 and Qwen3) confirm strong cross-model transferability. Furthermore, a Gaussian Process-based Bayesian optimization module reduces evaluation cost by 84% and reveals a power-law relationship between model size and cross-domain robustness. These results demonstrate that reasoning summarization provides a practical path toward efficient CoT transfer, enabling advanced reasoning under tight computational constraints. Code will be released upon publication.
ロジック グリッド パズルによる LLM 推論の暗黙的なバイアスの評価
最近の安全ガードレールは、あからさまに偏った出力を効果的に抑制しますが、現在の評価ベンチマークを回避する複雑な論理的推論タスク中には、より微妙な形の社会的バイアスが出現します。このギャップを埋めるために、新しい評価フレームワーク PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation) を導入します。このフレームワークは、論理グリッド パズルを使用して、LLM の論理的推論と意思決定に対する社会的固定観念の影響を系統的に調査します。ロジック パズルを使用することで、自動生成と検証が可能になるだけでなく、複雑さや偏った設定の多様性も可能になります。 PRIME には、共有パズル構造から生成された定型パズル、反定型パズル、および中立的なパズルのバリアントが含まれており、制御されたきめ細かい比較が可能です。パズルのサイズ全体で複数のモデル ファミリを評価し、プロンプトベースの緩和戦略の有効性をテストします。性別のステレオタイプに焦点を当てた実験に焦点を当てた私たちの調査結果は、解決策がステレオタイプの関連性と一致する場合、モデルが一貫してより正確に推論することを強調しています。これは、公平性が重要である LLM の演繹的推論で永続する社会的偏見を診断し定量化するための PRIME の重要性を示しています。
原文 (English)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation benchmarks. To fill this gap, we introduce a new evaluation framework, PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation), that uses logic grid puzzles to systematically probe the influence of social stereotypes on logical reasoning and decision making in LLMs. Our use of logic puzzles enables automatic generation and verification, as well as variability in complexity and biased settings. PRIME includes stereotypical, anti-stereotypical, and neutral puzzle variants generated from a shared puzzle structure, allowing for controlled and fine-grained comparisons. We evaluate multiple model families across puzzle sizes and test the effectiveness of prompt-based mitigation strategies. Focusing our experiments on gender stereotypes, our findings highlight that models consistently reason more accurately when solutions align with stereotypical associations. This demonstrates the significance of PRIME for diagnosing and quantifying social biases perpetuated in the deductive reasoning of LLMs, where fairness is critical.
GameDevBench: ゲーム開発を通じてエージェント機能を評価する
コーディングエージェントの急速な進歩にもかかわらず、マルチモーダル対応物の進歩は遅れています。主な課題は、ソフトウェア開発の複雑さとマルチモーダルな深い理解の必要性を兼ね備えた評価テストベッドが不足していることです。ゲーム開発では、エージェントはビジュアル ゲーム シーン内のシェーダー、スプライト、アニメーションなどの本質的にマルチモーダルなアセットを操作しながら、大規模で高密度のコードベースをナビゲートする必要があります。ゲーム開発タスクにおけるエージェントを評価するための最初のベンチマークである GameDevBench を紹介します。 GameDevBench は、Web およびビデオ チュートリアルから派生した 333 のタスクで構成されています。タスクはマルチモーダルな理解が必要であり、複雑です。平均的なソリューションでは、以前のソフトウェア開発ベンチマークと比較して 3 倍以上のコード行とファイルの変更が必要です。エージェントはゲーム開発に苦労しており、最適なエージェントとメソッドではタスクの 53.8% しか解決できません。知覚されるタスクの難易度とマルチモーダルな複雑さの間には強い相関関係があり、平均成功率はゲームプレイ指向のタスクの 51.4% から 2D グラフィックス タスクの 33.0% まで低下していることがわかりました。マルチモーダル機能を向上させるために、エージェント向けに 2 つのシンプルな画像ベースとビデオベースのフィードバック メカニズムを導入しました。シンプルであるにもかかわらず、これらの方法は一貫してパフォーマンスを向上させ、視覚的なフィードバックを与えた場合の GPT-5.4 のパフォーマンスは 41.1% から 52.0% に向上しました。
原文 (English)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for deep multimodal understanding. In game development, agents must navigate large, dense codebases while manipulating intrinsically multimodal assets such as shaders, sprites, and animations within a visual game scene. We present GameDevBench, the first benchmark for evaluating agents on game development tasks. GameDevBench consists of 333 tasks derived from web and video tutorials. Tasks require significant multimodal understanding and are complex: the average solution requires over three times the lines of code and file changes compared to prior software development benchmarks. Agents struggle with game development, with the best agent and method solving only 53.8% of tasks. We find a strong correlation between perceived task difficulty and multimodal complexity, with average success rate dropping from 51.4% on gameplay-oriented tasks to 33.0% on 2D graphics tasks. To improve multimodal capability, we introduce two simple image- and video-based feedback mechanisms for agents. Despite their simplicity, these methods consistently improve performance, increasing GPT-5.4's performance from 41.1% to 52.0% when given visual feedback.
SleepLM: 人間の睡眠のための自然言語インテリジェンス
私たちは、人間の睡眠調整、解釈、自然言語との対話を可能にする睡眠言語基盤モデルのファミリーである SleepLM を紹介します。睡眠の重要な役割にもかかわらず、学習ベースの睡眠分析システムは閉じたラベル空間 (事前定義された段階やイベントなど) で動作し、新しい睡眠現象を記述、クエリ、または一般化することができません。 SleepLM は、自然言語とマルチモーダルポリソムノグラフィーを橋渡しし、言語に基づいた睡眠生理学の表現を可能にします。この調整をサポートするために、10,000 人を超える個人からの 10 万時間以上のデータで構成される、初の大規模睡眠テキスト データセットのキュレーションを可能にするマルチレベル睡眠キャプション生成パイプラインを導入しました。さらに、生理学的忠実度とクロスモーダル相互作用をより適切に捕捉するために、コントラストアライメント、キャプション生成、信号再構築を組み合わせた統合された事前トレーニング目標を提示します。現実世界の睡眠理解タスクに関する広範な実験により、SleepLM がゼロショット学習と少数ショット学習、クロスモーダル検索、および睡眠キャプション作成において最先端のパフォーマンスを上回っていることが検証されています。重要なのは、SleepLM は、言語ガイドによるイベントのローカリゼーション、ターゲットを絞った洞察の生成、目に見えないタスクに対するゼロショットの一般化などの興味深い機能も示していることです。すべてのコードとデータはオープンソースになります。
原文 (English)
SleepLM: Natural-Language Intelligence for Human Sleep
We present SleepLM, a family of sleep-language foundation models that enable human sleep alignment, interpretation, and interaction with natural language. Despite the critical role of sleep, learning-based sleep analysis systems operate in closed label spaces (e.g., predefined stages or events) and fail to describe, query, or generalize to novel sleep phenomena. SleepLM bridges natural language and multimodal polysomnography, enabling language-grounded representations of sleep physiology. To support this alignment, we introduce a multilevel sleep caption generation pipeline that enables the curation of the first large-scale sleep-text dataset, comprising over 100K hours of data from more than 10,000 individuals. Furthermore, we present a unified pretraining objective that combines contrastive alignment, caption generation, and signal reconstruction to better capture physiological fidelity and cross-modal interactions. Extensive experiments on real-world sleep understanding tasks verify that SleepLM outperforms state-of-the-art in zero-shot and few-shot learning, cross-modal retrieval, and sleep captioning. Importantly, SleepLM also exhibits intriguing capabilities including language-guided event localization, targeted insight generation, and zero-shot generalization to unseen tasks. All code and data will be open-sourced.
ゼロショット タスクでの MLLM の検証と拡張のための明示的なロジック チャネル
Frontier Multimodal Large Language Model (MLLM) は、Visual-Language Comprehension (VLC) タスクにおいて優れた機能を発揮します。ただし、これらは多くの場合、ブラックボックス方式で新しいタスクに対するゼロショット ソリューションとして導入されます。これらのモデルの動作を検証して理解することは、新しいタスクに適用するために重要になります。モデルの検証、選択、強化のための明示的な論理推論を実行するために、ブラックボックス モデル チャネルと並行して明示的ロジック チャネルを提案します。潜在的な視覚言語知識をカプセル化したフロンティア MLLM は、暗黙的な論理チャネルと考えることができます。提案された明示的論理チャネルは、人間の論理的推論を模倣し、LLM、VFM、および明示的な視覚的証拠に対する事実、反事実、関係推論のための確率的推論を備えた論理的推論を組み込んでいます。一致率 (CR) は、グラウンド トゥルース アノテーションがない場合でも、クロスチャネル検証とモデル選択のために提案されています。さらに、クロスチャネル統合により、信頼性を高める明確な視覚的証拠に基づいて、MLLM よりもゼロショット タスクのパフォーマンスがさらに向上します。 4 つのフロンティア ファミリの 11 個の最近のオープンソース MLLM を使用して、2 つの代表的な VLC タスク、つまり MC-VQA と HC-REC について、3 つの困難なベンチマークで包括的な実験を実施しました。私たちの体系的な評価は、説明可能性と信頼性が強化されたMLLMのモデル検証、選択、改善に対する提案されたELCとCRの有効性を実証しています。
原文 (English)
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
Frontier Multimodal Large Language Models (MLLMs) exhibit remarkable capabilities in Visual-Language Comprehension (VLC) tasks. However, they are often deployed as zero-shot solution to new tasks in a black-box manner. Validating and understanding the behavior of these models become important for application to new task. We propose an Explicit Logic Channel, in parallel with the black-box model channel, to perform explicit logical reasoning for model validation, selection and enhancement. The frontier MLLM, encapsulating latent vision-language knowledge, can be considered as an Implicit Logic Channel. The proposed Explicit Logic Channel, mimicking human logical reasoning, incorporates a LLM, a VFM, and logical reasoning with probabilistic inference for factual, counterfactual, and relational reasoning over the explicit visual evidence. A Consistency Rate (CR) is proposed for cross-channel validation and model selection, even without ground-truth annotations. Additionally, cross-channel integration further improves performance in zero-shot tasks over MLLMs, grounded with explicit visual evidence to enhance trustworthiness. Comprehensive experiments conducted for two representative VLC tasks, i.e., MC-VQA and HC-REC, on three challenging benchmarks, with 11 recent open-source MLLMs from 4 frontier families. Our systematic evaluations demonstrate the effectiveness of proposed ELC and CR for model validation, selection and improvement on MLLMs with enhanced explainability and trustworthiness.
XSkill: マルチモーダル エージェントの経験とスキルからの継続的な学習
マルチモーダル エージェントは、さまざまなツールを使用して複雑な推論タスクに取り組むことができるようになりましたが、依然として、非効率なツールの使用と無制限の設定での柔軟性のないオーケストレーションに悩まされています。中心的な課題は、そのようなエージェントが過去の軌跡から学習することでパラメータを更新せずに継続的に改善できるようにすることです。私たちは、この目標に不可欠な再利用可能な知識の 2 つの相補的な形式を特定します。それは、ツールの選択と意思決定のための簡潔なアクション レベルのガイダンスを提供するエクスペリエンスと、計画とツールの使用のための構造化されたタスク レベルのガイダンスを提供するスキルです。この目的を達成するために、マルチモーダル エージェントの経験とスキルから継続的に学習するためのデュアル ストリーム フレームワークである XSkill を提案します。 XSkill は、知識の抽出と検索の両方を視覚的な観察に基づいて行います。 XSkill は、蓄積中に、視覚的に根拠のある要約とクロスロールアウトの批評を通じて、マルチパス ロールアウトの経験とスキルを抽出して統合します。推論中に、この知識を取得して現在の視覚的コンテキストに適応させ、使用履歴を蓄積にフィードバックして継続的な学習ループを形成します。 4 つのバックボーン モデルを使用してさまざまなドメインにわたる 5 つのベンチマークで評価したところ、XSkill はツールのみのベースラインと学習ベースのベースラインの両方を一貫して大幅に上回りました。さらなる分析により、2 つの知識ストリームがエージェントの推論動作に影響を与える際に補完的な役割を果たし、優れたゼロショット一般化を示すことが明らかになりました。
原文 (English)
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestration in open-ended settings. A central challenge is enabling such agents to continually improve without parameter updates by learning from past trajectories. We identify two complementary forms of reusable knowledge essential for this goal: experiences, providing concise action-level guidance for tool selection and decision making, and skills, providing structured task-level guidance for planning and tool use. To this end, we propose XSkill, a dual-stream framework for continual learning from experience and skills in multimodal agents. XSkill grounds both knowledge extraction and retrieval in visual observations. During accumulation, XSkill distills and consolidates experiences and skills from multi-path rollouts via visually grounded summarization and cross-rollout critique. During inference, it retrieves and adapts this knowledge to the current visual context and feeds usage history back into accumulation to form a continual learning loop. Evaluated on five benchmarks across diverse domains with four backbone models, XSkill consistently and substantially outperforms both tool-only and learning-based baselines. Further analysis reveals that the two knowledge streams play complementary roles in influencing the reasoning behaviors of agents and show superior zero-shot generalization.
SocialOmni: オムニ モデルにおけるオーディオビジュアルのソーシャル インタラクティビティのベンチマーク
オムニモーダル大規模言語モデル (OLM) は、オーディオ、ビジョン、テキストをネイティブに統合することにより、ヒューマン マシン インタラクションを再定義します。しかし、既存の OLM ベンチマークは依然として静的な精度中心のタスクに固定されており、社会的インタラクティブ性、つまり自然な対話の動的な手がかりをナビゲートする基本的な能力の評価において重大なギャップが残されています。この目的を達成するために、私たちは SocialOmni を提案します。SocialOmni は、(i) 話者の分離と識別 (誰が話しているのか)、(ii) 割り込みのタイミング制御 (いつ口を挟むか)、および (iii) 自然な割り込みの生成 (割り込みの表現方法) という 3 つの核となる次元にわたって、この会話の対話性の評価を運用化する包括的なベンチマークです。 SocialOmni は、2,000 の知覚サンプルと、厳密な時間的および文脈上の制約を持つ 209 のインタラクション生成インスタンスの品質管理された診断セットを特徴とし、モデルの堅牢性をテストするための制御された視聴覚の不一致シナリオによって補完されます。 12 の主要な OLM をベンチマークしたところ、モデル間でソーシャル インタラクション機能に大きな差異があることが明らかになりました。さらに、私たちの分析では、モデルの知覚精度と、状況に応じて適切な中断を生成する能力との間の顕著な切り離しが明らかになり、理解中心の指標だけでは会話の社会的能力を特徴付けるには不十分であることが示されています。さらに心強いのは、SocialOmni からのこれらの診断は、将来の OLM における認識とインタラクションの溝を埋めるための実用的なシグナルを生成することです。
原文 (English)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a critical gap in assessing social interactivity, the fundamental capacity to navigate dynamic cues in natural dialogues. To this end, we propose SocialOmni, a comprehensive benchmark that operationalizes the evaluation of this conversational interactivity across three core dimensions: (i) speaker separation and identification (who is speaking), (ii) interruption timing control (when to interject), and (iii) natural interruption generation (how to phrase the interruption). SocialOmni features 2,000 perception samples and a quality-controlled diagnostic set of 209 interaction-generation instances with strict temporal and contextual constraints, complemented by controlled audio-visual inconsistency scenarios to test model robustness. We benchmarked 12 leading OLMs, which uncovers significant variance in their social-interaction capabilities across models. Furthermore, our analysis reveals a pronounced decoupling between a model's perceptual accuracy and its ability to generate contextually appropriate interruptions, indicating that understanding-centric metrics alone are insufficient to characterize conversational social competence. More encouragingly, these diagnostics from SocialOmni yield actionable signals for bridging the perception-interaction divide in future OLMs.
Rule-VLN: 意味論的推論と幾何学的修正による認識とコンプライアンスの橋渡し
身体化された AI が現実世界への展開に移行するにつれて、ビジョンと言語のナビゲーション (VLN) タスクの成功は、単なる到達可能性から社会的コンプライアンスへと進化する傾向があります。しかし、現在のエージェントは、意味論的なルール (「行ってもいいですか?」) よりも物理的な幾何学形状 (「行ってもいいですか?」) を優先する「目標主導型の罠」に悩まされており、微妙な規制上の制約を頻繁に見落としています。このギャップを埋めるために、私たちはルールに準拠したナビゲーションのための初の大規模都市ベンチマークである Rule-VLN を確立します。大規模な 29,000 ノード環境にまたがり、177 の多様な規制カテゴリを 4 つのカリキュラム レベルにわたる 8,000 の制約付きノードに注入し、きめ細かい視覚的および行動的制約でエージェントに挑戦します。さらに、事前訓練を受けたエージェントに安全意識を与えるように設計された汎用のゼロショット モジュールであるセマンティック ナビゲーション修正モジュール (SNRM) を提案します。 SNRM は、動的な迂回計画のために、粗いから細かいまでの視覚認識 VLM フレームワークと認識的メンタル マップを統合します。実験では、Rule-VLN が最先端のモデルに挑戦する一方で、SNRM がナビゲーション機能を大幅に回復し、CVR が 19.26% 減少し、TC が 5.97% 向上することが実証されました。
原文 (English)
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachability to social compliance. However, current agents suffer from a "goal-driven trap", prioritizing physical geometry ("can I go?") over semantic rules ("may I go?"), frequently overlooking subtle regulatory constraints. To bridge this gap, we establish Rule-VLN, the first large-scale urban benchmark for rule-compliant navigation. Spanning a massive 29k-node environment, it injects 177 diverse regulatory categories into 8k constrained nodes across four curriculum levels, challenging agents with fine-grained visual and behavioral constraints. We further propose the Semantic Navigation Rectification Module (SNRM), a universal, zero-shot module designed to equip pre-trained agents with safety awareness. SNRM integrates a coarse-to-fine visual perception VLM framework with an epistemic mental map for dynamic detour planning. Experiments demonstrate that while Rule-VLN challenges state-of-the-art models, SNRM significantly restores navigation capabilities, reducing CVR by 19.26% and boosting TC by 5.97%.
EvoMaster: 大規模なエージェント科学のための基礎的な進化するエージェント フレームワーク
大規模な言語モデルとエージェントの収束は、科学的発見の新時代であるエージェント科学を促進しています。科学的手法は本質的に反復的ですが、既存のエージェント フレームワークは主に静的で、範囲が狭く、試行錯誤から学ぶ能力がありません。このギャップを埋めるために、大規模なエージェント サイエンス向けに特別に設計された、基本的な進化するエージェント フレームワークである EvoMaster を紹介します。 EvoMaster は、継続的な自己進化という中心原理に基づいて、エージェントが仮説を反復的に改良し、自己批判し、実験サイクル全体にわたって段階的に知識を蓄積できるようにし、人間の科学的調査を忠実に反映します。重要なのは、ドメインに依存しないベース ハーネスとして、EvoMaster はスケールアップが非常に簡単であることです。これにより、開発者は、約 100 行のコードで、任意の分野向けの高機能で自己進化する科学エージェントを構築して展開できます。 EvoMaster に基づいて構築され、機械学習、物理学、生物学、ウェブ研究、一般科学などの分野にわたって SciMaster エコシステムを育成しました。科学研究/コーディング/実験、科学的推論と情報検索、実践的な科学的問題解決に及ぶ 10 のベンチマークの評価により、EvoMaster と OpenHands、OpenClaw、Codex が比較されます。 EvoMaster は、10 のベンチマークのうち 9 つで最高スコアを達成し、4 つのエージェントの中で最も強い平均スコア (58.0%) を達成し、次世代の自律的な科学的発見のための主要な基礎フレームワークとしての有効性と汎用性が実証されました。
原文 (English)
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science. While the scientific method is inherently iterative, existing agent frameworks are predominantly static, narrowly scoped, and lack the capacity to learn from trial and error. To bridge this gap, we present EvoMaster, a foundational evolving agent framework engineered specifically for Agentic Science at Scale. Driven by the core principle of continuous self-evolution, EvoMaster empowers agents to iteratively refine hypotheses, self-critique, and progressively accumulate knowledge across experimental cycles, faithfully mirroring human scientific inquiry. Crucially, as a domain-agnostic base harness, EvoMaster is exceptionally easy to scale up -- enabling developers to build and deploy highly capable, self-evolving scientific agents for arbitrary disciplines in approximately 100 lines of code. Built upon EvoMaster, we incubated the SciMaster ecosystem across domains such as machine learning, physics, biology, web research, and general science. Evaluations on ten benchmarks spanning scientific research/coding/experimentation, scientific reasoning and information search, and practical scientific problem solving compare EvoMaster against OpenHands, OpenClaw, and Codex. EvoMaster achieves the highest score on nine of the ten benchmarks and the strongest average score (58.0\%) among the four agents, validating its efficacy and generality as the premier foundational framework for the next generation of autonomous scientific discovery.
LiteResearcher: Deep Research エージェント用のスケーラブルなエージェント RL トレーニング フレームワーク
強化学習 (RL) は、LLM ベースのエージェントの強力なトレーニング パラダイムとして登場しました。しかし、深い調査のためのエージェントティック RL のスケーリングは、依然として 2 つの複合的な課題によって制約されています。それは、手作りの合成データでは本物の現実世界の検索機能を引き出すことができないこと、もう 1 つは、RL トレーニング中の現実世界の検索依存性により、不安定性と法外なコストが生じ、エージェントティック RL のスケーラビリティが制限されることです。 LiteResearcher は、Agentic RL をスケーラブルにするトレーニング フレームワークです。現実世界の検索ダイナミクスを反映する軽量仮想世界を構築することで、小さな検索エージェントが大規模なオープンソース モデルや商用モデル (Tongyi DeepResearch や Claude-4.5 Sonnet など) を上回るパフォーマンスを発揮できるようにするトレーニング レシピを継続的に改善できます。具体的には、GAIA や Xbench などの一般的なベンチマークで、当社の LiteResearcher-4B は、それぞれ 71.3% と 78.0% というオープンソースの最先端の結果を達成し、スケーラブルな RL トレーニングがディープ リサーチ エージェントの重要な実現要因であることを実証しています。
原文 (English)
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elicit genuine real-world search capabilities, and real-world search dependency during RL training introduces instability and prohibitive cost, which limits the scalability of Agentic RL. LiteResearcher is a training framework that makes Agentic RL scalable: by constructing a lite virtual world that mirrors real-world search dynamics, we enable a continuously improving training recipe that empowers a tiny search agent to outperform large-scale open-source and commercial models (e.g., Tongyi DeepResearch and Claude-4.5 Sonnet). Specifically, on common benchmarks such as GAIA and Xbench, our LiteResearcher-4B achieves open-source state-of-the-art results of 71.3% and 78.0% respectively, demonstrating that scalable RL training is a key enabler for Deep Research Agents.
Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligence
Distributed collaborative intelligence (DCI), encompassing edge-to-edge architectures, federated learning, transfer learning, and swarm sys…
LLM における推論の質の測定: 多次元の行動フレームワーク
LLM は複雑な推論タスクで目覚ましい成功を収めていますが、現在の評価アプローチは主に最終的な答えの正しさに依存しており、それらの答えを生み出す根本的な推論プロセスについての洞察は限られています。このギャップに対処するために、この研究では、動作の観点から LLM の推論品質を測定するための統一された多次元フレームワークを提案し、理論的に根拠のある 6 つの次元、正確性 (CQ)、一貫性 (CS)、堅牢性 (RS)、論理的一貫性 (LS)、効率 (ES)、安定性 (SS) を運用します。 4 つのベンチマークの 975 項目にわたる 7 つの LLM に関する広範な実験により、このフレームワークが精度のみの指標では見えない動作を明らかにすることが実証されました。特に、論理的一貫性は正しさ (r = -0.172、ns) と直交しており、一貫性のない推論から正しい答えが得られることが確認され、一方、Claude-Haiku-4.5 は最高の多次元スコア (Q_bal = 0.778) を達成しています。さらに、このフレームワークは重大なランキングの逆転を明らかにしています。DeepSeek-V3 は精度優先では 2 位ですが、法的/コンプライアンスの重み付けでは 5 位にランクされており、単一指標の評価では検出できない逆転です。判別式の妥当性により、11/15 次元のペアが独立している (|r| < 0.50) ことが確認され、各次元を別個の信号として扱うための心理測定的サポートが提供されます。フレームワークによって生成される次元プロファイルは、次の 3 つのクラスの展開決定を直接サポートします。最終的な答えが正しいにもかかわらず、その推論トレースが説明責任監査に失敗するモデルを特定します (LS--CQ 直交性)。精度のみのベンチマークによって引き起こされるランキングエラーを防止します。そして、フレームワークがキャプチャする 6 つの独立したシグナルを単一のメトリックが暗黙的に置き換えることがないようにします。
原文 (English)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes a unified multi-dimensional framework for measuring LLM reasoning quality from a behavioral perspective, operationalizing six theoretically grounded dimensions rooted in cognitive science: Correctness (CQ), Consistency (CS), Robustness (RS), Local Logical Coherence (LS), Efficiency (ES), and Stability (SS). The framework introduces deployment-aware aggregation, enabling context-specific model selection beyond accuracy-based leaderboards. Experiments across multiple LLMs and benchmarks reveal behaviors systematically concealed by single-metric evaluation, including the orthogonality of local logical coherence and correctness, deployment-context-dependent ranking inversions, and non-trivial dimensional profiles in small locally-deployed models. Discriminant validity analysis confirms that the proposed dimensions capture largely non-redundant signals. The resulting pipeline provides a foundation for diagnosing LLM reasoning behavior across deployment contexts, with domain-specific validation as a direction for future work.
Reasoning4Sciences: 推論言語モデルをすべての科学分野に橋渡しする
推論言語モデル (RLM) は科学研究のための強力なツールとして急速に台頭していますが、その影響は主に「ハード サイエンス」分野に集中しています。他の科学分野での RLM の導入が遅い、または導入されていないことが、研究の生産性の差の拡大を引き起こしています。この調査では、欧州研究評議会 (ERC) が使用する社会科学と人文科学、物理科学と工学、生命科学にわたる分類に従って、28 の科学分野にわたる RLM の採用に関する初めての包括的な分析を提供します。私たちは、RLM がどのように開発、評価され、分野全体に適用されるかを調査します。さらに、利用可能なドメイン固有の開発および評価リソースに基づいた成熟度指向の評価フレームワークを導入し、公開されているリソースのみを考慮した場合にさらに顕著になる RLM 成熟度の実質的な格差を明らかにします。最後に、分野を超えて普及しつつある現在の実装パラダイム、現在の課題、科学全体で RLM の導入を可能にする将来の方向性を強調します。
原文 (English)
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrated in "hard science" fields. The slow -- or lack of -- adoption of RLMs in other branches of science is causing a widening gap in research productivity. In this survey, we provide the first comprehensive analysis of RLM adoption across 28 scientific disciplines following the classification used by the European Research Council (ERC), spanning the Social Sciences and Humanities, Physical Sciences and Engineering, and Life Sciences. We examine how RLMs are developed, evaluated, and applied across disciplines. Furthermore, we introduce a maturity-oriented assessment framework based on available domain-specific development and evaluation resources, revealing substantial disparities in RLM maturity that become even more pronounced when only publicly available resources are considered. Finally, we highlight current implementation paradigms that are gaining popularity across disciplines, current challenges, and future directions in enabling RLM adoption across science.
話す前に考える: マルチエージェント社会シミュレーションにおける内部評価から公の表現まで
LLM ベースのマルチエージェント シミュレーションは、社会的相互作用、熟慮、集団的な意見のダイナミクスを研究するための有望な方法を提供します。しかし、既存の対話シミュレーション フレームワークの多くは、対話を主に観察可能なターン交換または集約された出力として表現しており、沈黙、発言意図、公的表現の背後にある内部評価プロセスを調査することが困難なままになっています。エージェントの私的な推論を公的発話の生成から分離する、インターバルベースのマルチエージェント シミュレーション フレームワークである TBS (Think-Before-Speak) を紹介します。各間隔で、すべてのエージェントは共有された対話履歴と自身の記憶に基づいて構造化された内部状態を更新します。これらの状態には、不協和音関連の評価、認識された世論環境、認識された孤立リスク、対応戦略、および発言意欲が含まれます。その後、オーケストレーターは競合する発言意図を解決し、1 つの発言を公開対話にコミットし、内部評価と公開対話が時間の経過とともに共進化できるようにします。私たちは、気候関連の政策問題に関するタウンホールでの議論を模擬して TBS を評価します。結果は、TBS が一貫した内部状態トレースを生成し、これらのトレースがターン割り当て、沈黙、メモリ条件全体にわたって体系的に変化することを示しています。不協和音関連の評価はエージェントの発言意欲を高めますが、沈黙の圧力評価はそれを低下させます。発言の意図が形成されると、公の場での表現は主に順番の割り当てルールによって形成されます。これらの発見は、TBS が内部評価から公的表現への経路を観察可能かつ分析可能にすることで、メカニズムに敏感な社会シミュレーションをサポートしていることを示唆しています。
原文 (English)
Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation
LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frameworks represent interaction mainly as observable turn exchange or aggregated outputs, leaving the internal evaluative processes behind silence, speaking intention, and public expression difficult to examine. We introduce TBS (Think-Before-Speak), an interval-based multi-agent simulation framework that separates agents' private reasoning from public utterance generation. At each interval, all agents update structured internal states based on the shared dialogue history and their own memory. These states include dissonance-related appraisal, perceived opinion climate, perceived isolation risk, response strategy, and willingness to speak. The orchestrator then resolves competing speaking intentions and commits one utterance to the public dialogue, allowing internal evaluation and public interaction to co-evolve over time. We evaluate TBS in simulated town hall discussions on a climate-related policy issue. Results show that TBS produces coherent internal-state traces and that these traces vary systematically across turn-allocation, silence, and memory conditions. Dissonance-related appraisal increases agents' willingness to speak, whereas silence-pressure appraisal decreases it. Once speaking intention is formed, public expression is shaped mainly by turn-allocation rules. These findings suggest that TBS supports mechanism-sensitive social simulation by making the pathway from internal evaluation to public expression observable and analyzable.
TerraBench: エージェントは異種の地球システム データを推論できるか?
気候と環境に関する意思決定では、グリッド化された物理データ、衛星画像、地理空間コンテキスト、シミュレーターの出力など、異種混合の入力全体にわたる推論がますます必要になります。気象および気候基盤モデルは適切に予測できますが、言語で対話的に推論することはできません。一方、大規模言語モデル (LLM) は言語で推論しますが、高次元の地球システム データを直接操作することはできません。その結果、地球科学における実際の科学ワークフローは十分なサービスを受けられないままです。 TerraAgent 上に構築された、根拠のある地球科学推論のベンチマークである TerraBench を紹介します。TerraAgent は、推論、ツール呼び出し、観測をインターリーブして LLM 計画を環境検索、地理空間処理、シミュレーション、アーティファクトに基づく計算のための科学ツールと結び付ける ReAct スタイルの実行可能フレームワークです。 TerraBench は、地球観測画像、グリッド データ、GIS 推論、およびシミュレーションの分析を単一の実行可能なインターフェイスに統合します。一方、以前のベンチマークは、これらの機能を狭い個別のタスクに分離していました。また、プロセスレベルのツール使用メトリクスと許容差を意識した数値スコアを組み合わせたのも、この分野では初めてです。このベンチマークは、3 つのトラック (基礎、シミュレーターベース、ドキュメントベースの検証) にわたる 403 の広範なエージェント タスクと、24,500 の検証済み実行ステップを含む 8 つのアプリケーション ドメインで構成されています。これらの結果は、信頼できる地球科学エージェントはツールへのアクセスを超えて、異種ワークフローを調整し、ツールを正確にパラメータ化し、成果物の出所を保存する必要があることを示しています。
原文 (English)
TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reason interactively in language, while large language models (LLMs) reason in language but cannot operate directly on high-dimensional Earth-system data. As a result, real scientific workflows in Earth-science remain underserved. We introduce TerraBench, a benchmark for grounded Earth-science reasoning, built on TerraAgent, a ReAct-style executable framework that interleaves reasoning, tool calls, and observations to couple LLM planning with scientific tools for environmental retrieval, geospatial processing, simulation, and artifact-backed computation. TerraBench unifies analysis of Earth observation imagery, gridded data, GIS reasoning and simulation in a single executable interface, whereas prior benchmarks isolate these capabilities into narrow individual tasks. It is also the first in this space to pair process-level tool-use metrics with tolerance-aware numeric scoring. The benchmark comprises 403 extensive agentic tasks across three tracks (Fundamentals, Simulator-Grounded, and Document-Grounded Verification) and eight application domains with 24,500 verified execution steps. These results indicate that reliable Earth-science agents must go beyond tool access to coordinate heterogeneous workflows, parameterize tools precisely, and preserve artifact provenance.
WorkBench の再訪: 2 年後の職場エージェント
2024 年 3 月の WorkBench で最も優れたエージェントである GPT-4 は、タスクの 43% を完了し、そのうちの 26% で間違った人にメールを送信するなど、意図しない有害なアクションを実行しました。 2026 年 6 月にベンチマークを再確認したところ、これまでで最高のエージェントである Claude Opus 4.8 が 89% を完了し、2.5% で意図しない有害なアクションを実行していることがわかりました。フロンティアエージェントのパフォーマンスにおけるこの大幅な進歩とは別に、3 つの点が際立っています。まず、WorkBench では機能と安全性がトレードオフではなく両立するため、最も多くのタスクを完了したモデルは、意図しない損傷も最小限に抑えられます。第 2 に、いくつかの種類のエラーは完全に排除されましたが、フロンティア モデルは依然としていくつかの基本的な間違いを犯し、間違った人に電子メールを送信するなど、時として取り返しのつかない損害をもたらすことがあります。第三に、オープンウェイト モデルの台頭により、以前は独自モデルでしかアクセスできなかったパフォーマンス レベルのコストが大幅に低下し、一方でフロンティア コストは比較的安定しています。 2024 年以降、データとコードの品質向上、新しいモデル スコア、WorkBench でのエージェントの進行状況の分析を備えたベンチマークの更新バージョンをリリースします。
原文 (English)
WorkBench Revisited: Workplace Agents Two Years On
The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now completes 98%. Beyond this considerable progress in frontier agent performance, three things stand out. First, unintended harmful actions, such as emailing the wrong person, fell from 26% of tasks for GPT-4 to 1.9% for Claude Fable 5; capability and safety go together on WorkBench rather than trade off, so the models that finish the most tasks also do the least unintended damage. Second, the rise of open-weight models has drastically lowered costs for a performance level that was only accessible to proprietary models, while frontier costs have stayed stable. Third, while several classes of error have been eliminated, frontier models still make some basic mistakes that occasionally result in irreversible harm. We release an updated version of the benchmark with data and code quality improvements, new model scores, and analysis of agent progress on WorkBench since 2024.
Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval
Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate…
Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty
Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Mod…
HiComm: マルチエージェント強化学習のための階層型通信
協調的なマルチエージェント強化学習 (MARL) は、部分的な可観測性を軽減するために通信に依存することがよくありますが、既存のプロトコルのほとんどは、メッセージを要約した観測の構造から切り離されたフラットな密ベクトルとして扱います。この設計では、観察がグループやエンティティなどの階層に自然に従う多くの協力的な環境において、帰納的バイアスの重要な原因を見落としています。私たちは、送信者の階層的な観察にメッセージを基礎付けるプラグイン通信モジュールである \textsc{HiComm} を提案します。 \textsc{HiComm} は受信者主導型です。受信者はクエリを発行し、最初にグループ、送信者、そのグループ内のエンティティを選択する 3 段階のデコード プロセスを通じて階層を解決し、対応する機能スライスをメッセージとして返します。これにより、通信が非構造化ベクトル送信から、送信者の監視階層を介した構造化情報検索に変換されます。このメカニズムは、微分可能な離散選択と標準 MARL パイプラインに接続する軽量の共有投影設計のために、Straight-Through Gumbel-Softmax を使用してインスタンス化されます。さまざまな観察構造と調整要求を備えた協調的な MARL タスクにわたる実験では、\textsc{HiComm} が代表的な学習済み通信ベースラインと一致またはそれを上回るパフォーマンスを示し、一方、エピソードごとに受信者あたり通信量を最大 $23\times$ 削減することが示されました。
原文 (English)
HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to $23\times$ per receiver per episode.
FADE: 大規模な視覚言語モデルにおける言語優先支配を軽減することによる幻覚の軽減
Large Vision-Language Model (LVLM) の優れた機能にもかかわらず、依然として幻覚の影響を受けやすく、入力画像と一致しないコンテンツが生成されます。最近の研究では、これは視覚入力に対する言語事前の優位性によるものであり、この優位性を緩和するために対照的なデコード方法が採用されていますが、そのメカニズムの起源は未解明のままです。各変換層を通る情報の流れを調査すると、アテンション モジュールが一貫して視覚的証拠を集約し、クリティカル層の FFN モジュールが言語事前情報のソースとして機能することがわかりました。これらの事前分布は視覚的な証拠を無効にする可能性があり、中間層での正しい予測が不正確な出力に向かってドリフトする原因となります。この洞察に基づいて、言語優先の優位性を減らすために FFN 出力を減衰するトレーニング不要の方法である FADE (FFN Attenuation for DEcoding) を提案します。 LLaVA-1.5、mPLUG-Owl2、および InstructBLIP にわたる POPE、CHAIR、および MME ベンチマークの評価では、FADE が推論効率を維持しながら幻覚を効果的に軽減することが示されています。
原文 (English)
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language priors over visual inputs and employ contrastive decoding methods to mitigate this dominance, but the mechanistic origin remains unexplored. We investigate the information flow through each transformer layer and find that attention modules consistently aggregate visual evidence, while FFN modules at critical layers act as the source of language priors. These priors can override visual evidence, causing correct predictions in intermediate layers to drift toward incorrect outputs. Based on this insight, we propose FADE (FFN Attenuation for DEcoding), a training-free method that attenuates FFN outputs to reduce language-prior dominance. Evaluations on POPE, CHAIR, and MME benchmarks across LLaVA-1.5, mPLUG-Owl2, and InstructBLIP show that FADE effectively mitigates hallucinations while preserving inference efficiency.
トリプレットの妥当性を超えて: ナレッジ グラフにおける関係セットの完成
ナレッジ グラフ (KG) は、現実世界の知識をトリプレットとして編成し、多くの下流アプリケーションを支えます。本質的に不完全であるため、ナレッジ グラフ補完 (KGC) は広く研究されており、通常はリンク予測が主要なパラダイムであるトリプレット予測として定式化されます。ただし、この定式化はトリプレットごとの情報の不完全性に焦点を当てており、エンティティ関係の互換性情報の不完全性を見落としています。この制限に対処するために、リンク予測タスクを補完し、特定のエンティティと意味的に互換性のある欠落している関係を推論することを目的とした関係セット完了タスク (RSC) を導入します。さらに、観察されたエンティティの関係間の潜在的なパターンをモデル化して、欠落しているものを推測する関係セット埋め込みモデル (RelSetE) を提案します。 RelSetE を評価するために、標準の KG ベンチマークから 3 つのベンチマーク データセットを導出します。広範な実験により、RelSetE がエンティティ関係の互換性パターンを効果的に捕捉し、エンティティの欠落した関係を推論する際に有利に機能することが実証されました。コードとデータは公開されています。
原文 (English)
Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs
Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incompleteness, knowledge graph completion (KGC) is widely studied and is typically formulated as triplet prediction, with link prediction as the dominant paradigm. However, this formulation focuses on the incompleteness of triplet-wise information and overlooks the incompleteness of entity-relation compatibility information. To address this limitation, we introduce a relation set completion task (RSC), which complements the link prediction task and aims to reason about missing relations that are semantically compatible with a given entity. We further propose a Relation Set Embedding model (RelSetE), which models latent patterns among the observed relations of entities to infer missing ones. To evaluate RelSetE, we derive three benchmark datasets from standard KG benchmarks. Extensive experiments demonstrate that RelSetE effectively captures entity-relation compatibility patterns and performs favorably in inferring missing relations of entities. Code and data are publicly available.
関連性は許可されていません: 価値のある貢献には注意が必要です
関連性は許可ではありません。アテンションを使用すると、モデルは現在のクエリに関連するキーと値の項目を読み取ることができますが、そのような項目の値の寄与が予測の証拠になることは保証されません。検索された文章は、裏付けとなる証拠がなくても質問に関連している可能性があり、歴史的事実や時間的近傍によって、トゥルーテール ランキングや現在のエッジ スコアが曖昧になる場合もあります。この論文では、このギャップを、実際に予測パスに追加される加重値項 alpha_ij * v_j の許可問題として形式化します。私たちは、注意関連性 alpha_ij を保持し、プライマリ メトリックにつながる値パスを公開し、完全なモデルでは、学習されたクエリ項目権限 g_ij を通じて alpha_ij * v_j を alpha_ij * g_ij * v_j に変換するパスローカライズされたインターフェイスである Warrant を提案します。 CTDG リンク予測、MTPP 次マーク ランキング、RAG サポート証拠選択、STPP 次位置予測、および TKG テール予測のメトリクスを定義する値パスに同じ演算子を配置します。 32 のペア比較、3 つのシード、合計 192 回の実行にわたって、Warrant は 27 の比較で主要な指標を改善しました。実用的な段階は、10 個の実質的な効果、1 個の限界効果、8 個のプラスだが不確実な効果、8 個の同程度/無視できる効果、および 5 個のドロップで構成されます。パス ローカライゼーション チェックでは、正しいパス配置は、すべてのドメインで方向認識ベース パフォーマンスを上回り、一般的なアテンション配置を CTDG で +0.1076 AUC、TKG で +0.0683 MRR 上回ります。アブレーションの結果、ほとんどの TKG ゲインはヒストリカルテール値パスのエクスポージャから得られるのに対し、コア CTDG ゲインはエッジ条件付きクエリ項目のアクセス許可から得られることがわかります。結論として、予測証拠は注目マスではありません。重み付けされた値の項は、メトリックへのパス上で保証される場合にのみ証拠になります。
原文 (English)
Relevance Is Not Permission: Warranted Attention for Value Contributions
Relevance is not permission. Attention lets a model read key-value items related to the current query, but it does not guarantee that the value contribution of such an item becomes prediction evidence. A retrieved passage may be relevant to a question without being supporting evidence, and a historical fact or temporal neighbor may even blur true-tail ranking or the current edge score. This paper formalizes this gap as a permission problem for the weighted value term alpha_ij * v_j that is actually added to the prediction path. We propose Warrant, a path-localized interface that preserves attention relevance alpha_ij, exposes the value path leading to the primary metric, and, in the full model, turns alpha_ij * v_j into alpha_ij * g_ij * v_j through learned query-item permission g_ij. We place the same operator on the metric-defining value paths of CTDG link prediction, MTPP next-mark ranking, RAG supporting evidence selection, STPP next-location forecasting, and TKG tail prediction. Across 32 paired comparisons, 3 seeds, and 192 total runs, Warrant improves the primary metric in 27 comparisons; practical tiers consist of 10 substantial effects, 1 marginal effect, 8 positive but uncertain effects, 8 tie/negligible effects, and 5 drops. In the path-localization check, correct-path placement outperforms direction-aware Base performance in every domain and exceeds generic attention placement by +0.1076 AUC in CTDG and +0.0683 MRR in TKG. Ablations show that most TKG gains come from historical-tail value path exposure, whereas the core CTDG gain comes from edge-conditioned query-item permission. In conclusion, prediction evidence is not attention mass. A weighted value term becomes evidence only when it is warranted on the path to the metric.
ManimAgent: 視覚教育のための自己進化型マルチモーダル エージェント
マルチラウンド リフレクションを使用すると、大規模な言語モデルに基づいて構築されたエージェントが単一タスク内の障害から回復できますが、各タスクは孤立したエピソードのままになります。つまり、1 つのタスクで多くのリフレクション ラウンドにわたって学習された教訓は、次のタスクが開始される前に破棄されます。私たちはコード生成タスクでこのギャップを研究します。科学論文のセクションから、エージェントはオープンソースの Manim ライブラリに Python を記述して数学的アニメーションをレンダリングします。 ManimAgent は自己進化するマルチモーダル エージェントであり、重みの更新も人間のシードも必要とせず、完全に独自のタスク ストリームから成長したデュアル チャネルのエピソード メモリ バンクを通じて、タスク全体にわたってリフレクション エクスペリエンスを伝達します。各アニメーションが収束した後、ビジョン言語モデルがレンダリングされたキーフレームをスコアリングします。結果として得られる信号は、成功の根拠をソフト参照例として保存する正のチャネル M+ と、検証された失敗パターンをハードな既知の落とし穴として保存する負のチャネル M- に設定されます。メモリなし、一致した予算の検索拡張生成、およびシャッフルされたメモリ ベースラインに対する固定プローブ評価では、メモリ サイズが大きくなるにつれて、盲目の人間の Pass@1 が増加し、リフレクション ラウンドが減少します。コード、フリーズされたメモリ スナップショット、タスク ストリームをリリースします。
原文 (English)
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins. We study this gap on a code-generation task: from a scientific paper section, the agent writes Python in the open-source Manim library to render a mathematical animation. We present ManimAgent, a self-evolving multimodal agent that carries reflection experience across tasks through a dual-channel Episodic Memory Bank grown entirely from its own task stream, with no weight updates and no human seeds. After each animation converges, a vision-language model scores the rendered keyframes; the resulting signals populate a positive channel M+ that stores success rationales as soft Reference Examples, and a negative channel M- that stores validated failure patterns as hard Known Pitfalls. On a fixed-probe evaluation against no-memory, matched-budget retrieval-augmented generation, and shuffled-memory baselines, blind human Pass@1 rises and reflection rounds fall as memory size grows. We will release the code, frozen memory snapshots, and the task stream.
エキスパート ユーザーを超えて: エージェントは、ユーザーが好みを引き出すだけでなく、ユーザーが好みを構築できるように支援する必要があります。
通常、エージェントは専門ユーザー (自分が望むものについて明確な好みを持っているユーザー) を想定しており、タスクの指定が不十分な場合は常に質問を明確にするようデフォルト設定されています。私たちは、この仮定は非現実的であると主張します。ユーザーは多くの場合、好みを完全に指定するためのドメイン知識が不足しています。ある機能の好みについて尋ねられた場合、ユーザーは、例や説明などを通じて、その機能の好みを形成するために必要なドメイン知識をユーザーが学習できるようにエージェントが支援しなければ、答えることができない場合があります。これらの原則を形式化するために、情報経済学の Search-Experience-Credence フレームワークを利用して、ユーザーがエージェントの対話アクションに基づいて好みを構築する方法のモデルである CoPref を導入します。次に、これらのアイデアをエージェント レコメンダー システムで具体的に研究し、インタラクティブなベンチマークである CoShop を提案します。 CoShop では、エージェントが CoPref ユーザーと会話し、CoPref ユーザーに対して推奨事項を作成します。エージェントのパフォーマンスは、ユーザーがタスクを適切に指定するために必要な知識を得るのに役立つかどうかによって決まります。 5 つのフロンティア モデルを評価すると、5 ターンの対話にもかかわらず、CoShop で 56% の精度を超えるエージェントは存在しないことがわかりました。失敗の原因は、エージェントがアイテムを見つける能力にあるのではなく、インタラクションによってユーザーが欲しいものについて知っている範囲がほとんど広がっていないことに起因します。
原文 (English)
Beyond expert users: agents should help users construct preferences, not just elicit them
Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified. We argue this assumption is unrealistic. Users often lack the domain knowledge to have completely specified preferences; if asked about their preference on some feature, the user may be unable to answer without the agent helping the user to learn some domain knowledge needed to form a preference for that feature, e.g., via examples or explanations. To formalize these principles, we draw on the Search-Experience-Credence framework from Information Economics to introduce CoPref, a model of how users construct preferences based on agent dialog actions. We then study these ideas concretely in agentic recommender systems, proposing CoShop, an interactive benchmark. In CoShop, an agent converses with and makes recommendations for a CoPref user. The agent's performance depends on whether it can help the user gain the knowledge needed to specify the task well. Evaluating five frontier models, we find that no agent exceeds 56% accuracy on CoShop despite five turns of interaction. Failures stem not from agents' ability to find items, but from how little the interaction expands what users know about what they want.
なぜ 2 回解決するのか?伝達効率の高い ML エンジニアリングのためのスキルの階層的蓄積
すべての競争はコールド スタートであるため、ML エンジニアリング エージェントは既知の技術を再発見するために計算を無駄にします。我々は、競合他社にまたがる知識を 3 つのスコープ層 (グローバル、ドメイン、および競合固有) に編成し、それぞれがマッチング エージェント レベルに関連付けられた階層型マルチエージェント システムである HASTE を紹介します。オーケストレーターはドメイン スペシャリストを調整し、LLM 駆動の抽象化を通じて層間の学習を促進します。制御されたアブレーションは、範囲指定されたローディングの証拠を提供します。8 つの競技会にわたって 159 のスキル インベントリを一定に保持し、段階的ローディングでは 100% のメダル率を達成しますが、フラット ローディングでは 62.5% にしか達せず、スキルをロードしない場合と同じメダル率であり、出力トークンの 2 倍を消費します。 MLE-Bench Lite ベンチマーク全体 (Kaggle コンペティション 22 件) では、HASTE はクロード ソネット 4.6 を使用してコンペごとに 12 時間で 77.3% のメダル率に達しました。コールドスタート実行では、システムはスキルが蓄積されていない状態で開始されます。ウォーム スタート ランでは、以前の競技会で学んだスキルを再ロードし、競技会間での転送にはグローバル レベルおよびドメイン レベルのスキルのみを使用します。ウォーム スタートでは、改良の反復回数が 52% 減少し、提案された変更のうちエージェントが保持する割合は、在庫が少ない場合の 42% から、50 以上のスキルが利用可能になると 85% に増加します。これらの結果は、より優れた知識組織が ML エンジニアリング エージェントのモデルの強度と計算予算の一部を置き換えることができることを示唆しています。
原文 (English)
Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering
ML engineering agents waste compute rediscovering known techniques because every competition is a cold start. We present HASTE, a hierarchical multi-agent system that organizes cross-competition knowledge into three scope tiers (global, domain, and competition-specific), each coupled to a matching agent level. An orchestrator coordinates domain specialists and promotes learning between tiers via LLM-driven abstraction. A controlled ablation provides evidence for scoped loading: holding a 159-skill inventory constant across 8 competitions, tiered loading achieves a 100% medal rate while flat loading reaches only 62.5%, the same medal rate as loading no skills, and consumes 2x the output tokens. On the full MLE-Bench Lite benchmark (22 Kaggle competitions), HASTE reaches a medal rate of 77.3% using Claude Sonnet 4.6 at 12h per competition; this is a single-seed campaign result, and multi-seed replication is the priority follow-up. In a cold-start run, the system begins with no accumulated skills. In warm-start runs, it reloads skills learned from earlier competitions, using only global and domain-level skills for transfer across competitions. Warm starts use 52% fewer refinement iterations, and the fraction of proposed changes kept by the agent rises from 42% at low inventory to 85% once 50+ skills are available. These results suggest that better knowledge organization can partly substitute for model strength and compute budget in ML-engineering agents.
Xiaomi-GUI-0テクニカルレポート
グラフィカル ユーザー インターフェイス (GUI) エージェントは、ビジョン言語モデルに基づいて構築され、タップ、スワイプ、テキスト入力、ナビゲーションなどのインターフェイス アクションを通じて実際のアプリケーションでユーザー タスクをエンドツーエンドで完了します。ただし、既存の GUI エージェントは、主にオフラインの軌跡、シミュレートされた環境、および標準化されたベンチマークに基づいてトレーニングおよび評価されます。これらは、インターフェイスのレイアウト、インタラクション ロジック、異常状態の分布において実際のアプリケーションとは大幅に異なり、実際の使用における実行の安定性を忠実に特徴付けることはできません。そこでは、アカウントの状態、許可ダイアログ、支払い認証、リスク管理によって状態分布が継続的に再形成され、ベンチマーク スコアと実際のユーザビリティの間に永続的なギャップが生じます。このギャップを埋めるために、実際のデバイスの閉ループ内でトレーニングおよび評価される、実際のモバイル環境用のネイティブ マルチモーダル GUI エージェントである Xiaomi-GUI-0 を提案します。その中核となるのは、実デバイス主体のハイブリッド インフラストラクチャであり、物理デバイスが主要な実行環境であり、サンドボックスが補助的なサポートを提供するため、データ収集、トレーニング、ロールアウト、評価が実際の展開に近い実行分布を共有します。私たちは、高頻度のヘッド タスク、ロングテール インテントの高汎化データ、リフレクションとメモリの能力強化データにわたるマルチソース トレーニング データを構築し、障害の軌跡を修正されたアクション、リフレクションの説明、回復のデモンストレーションに変えるエラー駆動型のデータ フライホイールを導入します。モデルは、教師あり微調整、ステップレベルの強化学習、エージェント強化学習の漸進的な 3 段階のパイプラインを通じてトレーニングされます。公開ベンチマークと社内 RealMobile で評価された Xiaomi-GUI-0 は、RealMobile で 72.0%、AndroidWorld で 78.9% の成功率を達成し、現実世界のタスクにおける実行の安定性と異常状態の認識が大幅に向上しました。
原文 (English)
Xiaomi-GUI-0 Technical Report
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
Growing concerns over the theft and misuse of Large Language Models (LLMs) underscore the need for effective fingerprinting to link a model…
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
Industry is moving toward autonomous, network-connected machines that detect and adapt to changing conditions, including hardware faults. C…
ForAug: Mitigating Biases in Image Classification via Controlled Image Compositions
Large-scale image classification datasets exhibit strong compositional biases: objects tend to be centered, appear at characteristic scales…
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans. But are these explanations faithfu…
Comparative Analysis of Lightweight CNNs for Resource-Constrained Devices: Predictive Performance, Efficiency Trade-offs, and Initialization Effects
Lightweight convolutional neural networks are often compared using results obtained with different training recipes, input settings, and pr…
scDataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics
Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets…
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
Humanoid robots are promising to acquire various skills by imitating human behaviors. However, existing algorithms are only capable of trac…
Flow-Through Tensors: A Unified Computational Graph Architecture for Multi-Layer Transportation Network Optimization
Modern transportation network modeling increasingly involves the integration of diverse methodologies including sensor-based forecasting, r…
Fraud is Not Just Rarity: A Causal Prototype Attention Approach to Realistic Synthetic Oversampling
Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and th…
FLAT: Revealing Hidden Latent-Conditioned Backdoor Failures in Federated Learning
Horizontal federated learning (HFL) backdoor audits often summarize model behavior through clean accuracy (CA), mean attack success rate (A…
TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification
Handling missing data in time series classification remains a significant challenge in various domains. Traditional methods often rely on i…
CWT-Enhanced Vibration Sensing With Time-Frequency Region Localization Using YOLO
This letter presents a CWT-enhanced vibration sensing framework for bearing fault monitoring through localized time-frequency region detect…
Predicting LLM Reasoning Performance with Small Proxy Model
Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize datasets bef…
Quadratic Programming Approach for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
On-device deployment of Large Language Models (LLMs) frequently leverages Low-Rank Adapters (LoRAs) to support diverse downstream tasks und…
Toward Cybersecurity-Expert Small Language Models
Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, do…
Reasoning Up the Instruction Ladder for Controllable Language Models
As large language model (LLM) based systems take on high-stakes roles in real-world decision-making, they must reconcile competing instruct…
Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture th…
FlowPath: Learning Data-Driven Manifolds with Invertible Flows for Robust Irregularly-sampled Time Series Classification
Modeling continuous-time dynamics from sparse and irregularly-sampled time series remains a fundamental challenge. Neural controlled differ…
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
Despite notable progress in text-guided medical image segmentation nowadays, these methods are limited to single-round dialogues and fail t…
NI-Tex: Non-isometric Image-based Garment Texture Generation
Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited. To ac…
When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets
Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agen…
Computing Evolutionarily Stable Strategies in Imperfect-Information Games
We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
Catastrophic forgetting poses a fundamental challenge in continual learning, particularly when models are quantized for deployment efficien…
Utilizing Earth Foundation Models to Enhance the Simulation Performance of Hydrological Models with AlphaEarth Embeddings
Predicting river flow in places without streamflow records is challenging because basins respond differently to climate, terrain, vegetatio…
Controllable Diffusion-Based Lesion Inpainting for Scalable Histopathology Data Augmentation
Expert-annotated training data remains the critical bottleneck for AI in histopathology, particularly for rare pathologies where even dozen…
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are uncha…
NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents
Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enablin…
PaAno: Patch-Based Representation Learning for Time-Series Anomaly Detection
Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transfo…
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
Mixture-of-Experts (MoE) architectures are evolving towards finer granularity to improve parameter efficiency. However, existing MoE design…
Calibrated Test-Time Guidance for Bayesian Inference
Test-time guidance is a widely used mechanism for steering pretrained diffusion models toward outcomes specified by a reward function. Exis…
Stateful Token Reduction for Long-Video Hybrid VLMs
Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated…
On the Reliability of Cue Conflict and Beyond
Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes. The cue-confl…
Competition-Aware CPC Forecasting with Near-Market Coverage
Cost-per-click (CPC) in paid search is an auction-generated outcome shaped by a competitive landscape that is only partially observable fro…
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
End-to-End autonomous driving (E2E-AD) systems face challenges in lifelong learning, including catastrophic forgetting, difficulty in knowl…
Interact3D: Compositional 3D Generation of Interactive Objects
Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional o…
Deep Learning-Driven Black-Box Doherty Power Amplifier with Pixelated Output Combiner and Extended Efficiency Range
This article presents a deep learning-driven inverse design methodology for Doherty power amplifiers (PA) with multi-port pixelated output…
KGS-GCN: Kinematics-Driven Gaussian Splatting and Probabilistic Topology for Skeleton-Based Action Recognition
Skeleton-based action recognition is widely applied in sensor-based systems, including human-computer interaction and intelligent surveilla…
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-l…
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
The identification of hazardous driving behaviors from in-cabin video streams is essential for enhancing road safety and supporting the det…
Planning over MAPF Agent Dependencies via Multi-Dependency PIBT
Modern Multi-Agent Path Finding (MAPF) algorithms must plan for hundreds to thousands of agents in congested environments within a second,…
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization
Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging becau…
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updat…
Moir\'e Video Authentication: A Physical Signature Against AI Video Generation
Recent advances in video generation have made AI-synthesized content increasingly difficult to distinguish from real footage. We propose a…
Crystalite: A Lightweight Transformer for Efficient Crystal Modeling
Generative models for crystalline materials often rely on equivariant graph neural networks, which capture geometric structure well but are…
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
AI agents that pay for resources via the x402 protocol embed payment metadata - resource URLs, descriptions, and reason strings - in every…
Continuous Knowledge Metabolism: Generating Scientific Hypotheses from Evolving Literature
Identifying promising research directions in fast-moving subareas is one of the most cognitively expensive tasks in modern AI research. Exi…
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial re…
FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices
Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data, y…
REALM: An RGB- and Event-Aligned Latent Manifold for Cross-Modal Perception
Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and rob…
SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
De-identification of clinical text is a prerequisite for the secondary use of electronic health records. Existing public benchmarks such as…
FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles
Recent progress in 4D Gaussian Splatting (4DGS) has achieved impressive dynamic scene reconstruction results. While these methods demonstra…
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclea…
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionat…
EcoGEO: Trajectory-Aware Evidence Ecosystems for Web-Enabled LLM Search Agents
Web-enabled LLM agents are changing how online information influences search outcomes. Existing Generative Engine Optimization (GEO) studie…
ChainCaps: 単調な機能減衰による構成安全なツール使用エージェント
ツールを使用するエージェントは、実行時にファイル システム、Web API、コード インタープリタ、およびエンタープライズ サービスを構成する、オープンエンドの展開環境で動作することが増えています。これにより、ツール構成に安全性のギャップが生じます。エージェントは、ツールごとのすべての権限チェックを満たしていても、機密文書の読み取り、要約、その要約の外部エンドポイントへの送信など、安全でないエンドツーエンドの影響を生み出す可能性があります。この障害モードをパーミッション ロンダリングと呼びます。 ChainCaps は、ランタイム ルールでこれに対処します。すべての値にはシンク固有の機能バジェットが含まれ、ツールの構成によって交差ごとにバジェットが伝播されます。値は、ツール チェーン内を移動するときに権限を保持したり失ったりする可能性がありますが、合成を通じて新しい権限を獲得することはできません。 ChainCaps は、エージェント サーバーやツール サーバーへの変更を必要としない透過的な MCP プロキシとして実装されています。 ChainCaps は、3 つのプロバイダーの 5 つのフロンティア モデルにわたる 82 のタスクにおいて、96 ~ 100% の無害な完了を維持しながら、攻撃の成功率を 25 ~ 68% から 0 ~ 4.8% に低下させます。再生実験では、スカラー IFC および関数ごとの分離ベースラインよりも優れたパフォーマンスを示します。マニフェストの品質が導入の主なボトルネックです。エキスパート マニフェストは 100% の攻撃ブロックに達しますが、単純なマニフェストは 27.3% に低下します。私たちの主張は、信頼できるマニフェストの下での明示的なフロー構成の安全性と、プロキシで可視のデータ移動に限定されており、今日導入されているツールを使用するエージェントには実際的なギャップがあります。
原文 (English)
ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation
Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterprise services at runtime. This creates a safety gap in tool composition: an agent can satisfy every per-tool permission check and still produce an unsafe end-to-end effect, such as reading a confidential document, summarizing it, and sending the summary to an external endpoint. We call this failure mode permission laundering. ChainCaps addresses it with a runtime rule: every value carries a sink-specific capability budget, and tool composition propagates budgets by intersection. A value can preserve or lose authority as it moves through a tool chain, but it cannot gain new authority through composition. We implement ChainCaps as a transparent MCP proxy that requires no changes to the agent or tool servers. On 82 tasks across five frontier models from three providers, ChainCaps reduces attack success rate from 25-68% to 0-4.8% while preserving 96-100% benign completion. In replay experiments, it also outperforms scalar-IFC and per-function-isolation baselines. Manifest quality is the dominant deployment bottleneck: expert manifests reach 100% attack blocking, while naive manifests fall to 27.3%. Our claims are limited to explicit-flow composition safety under trusted manifests and proxy-visible data movement, a practical gap in deployed tool-using agents today.
Diffusion Image Generation with Explicit Modeling of Data Manifold Geometry
Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, lo…
Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation
Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity…
LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs). Current QAT methods are mainly based on…
Korzhinskii-Net: 地下鉱物予測モデリングのための物理学に基づいたニューラル ネットワーク
鉱物見通しモデリング (MPM) は探査の経済学を支えていますが、ほとんどの運用パイプラインは浅い表面のプロキシでトレーニングされたデータ駆動型の分類器に限定されています。このようなモデルは、熱移流、流体の流れ、岩石学に依存する降水など、実際に鉱石の位置を特定する地下物理学を認識できません。我々は、ダルシー流、移流拡散熱輸送、およびソフトプラス飽和反応速度を単一の微分可能な順方向モデルに結合し、表面およびリモートセンシングプロキシによって弱く監視される 2 次元放射状物理情報ニューラル ネットワーク (PINN) である Korzhinskii-Net を紹介します。このネットワークは、浸透メタソマティズムの理論が物理的な足場を提供するドミトリ S. コルジンスキー (1899 ~ 1985 年) にちなんで名付けられました。我々は、ハードリング形状のネガを用いた公平で漏洩制御された5重相互検証プロトコルの下で、ノリリスク(Ni-Cu-PGE)、ペチェンガ(硫化Ni-Cu)、ウドカン(砂岩にホストされたCu)、スホーイ原木(造山運動のAu)、およびミールヌイ(キンバーライトダイヤモンド)の4つの商品クラスにまたがる5つの鉱石州でKorzhinskii-Netを評価しました。 Korzhinskii-Net は、最も強い古典的ベースライン (勾配ブースティング) の平均 PR-AUC 0.885 対 0.281、および平均分数ランク 0.019 対 0.413 を達成します。この改善は 5 つの州と 4 つの商品システムすべてで一貫しており、物理学に基づいた微分可能シミュレーターは、グローバルなオープンデータ プロキシによってのみ制約されている場合でも、純粋な特徴ベースの学習者が体系的に見逃している位置特定パターンを回復できることを示唆しています。完全なパイプラインと評価ハーネスをオープンソースとしてリリースします。
原文 (English)
Korzhinskii-Net: Physics-Informed Neural Network for Sub-Surface Mineral Prospectivity Modelling
Mineral prospectivity modelling (MPM) underpins exploration economics, yet most operational pipelines reduce to data-driven classifiers trained on shallow surface proxies. Such models are blind to the subsurface physics that actually localises ore: heat advection, fluid flow, and lithology-dependent precipitation. We present Korzhinskii-Net, a 2-D radial physics-informed neural network (PINN) that couples Darcy flow, advective-diffusive heat transport, and a softplus-saturated reaction rate into a single differentiable forward model, weakly supervised by surface and remote-sensing proxies. The network is named after Dmitri S. Korzhinskii (1899-1985), whose theory of infiltration metasomatism provides the physical scaffold. We evaluate Korzhinskii-Net on six ore provinces spanning three commodity classes - Udokan (sandstone-hosted Cu), Sukhoi Log, Olimpiada, and Berezovskoye (orogenic Au), Vorontsovskoye (Carlin-type Au), and Dalnegorsk (skarn polymetallic) - under a fair, leakage-controlled 5-fold cross-validation protocol with hard ring-shaped negatives and baseline proxy features disabled. Korzhinskii-Net attains a mean PR-AUC of 0.708 versus 0.235 for the strongest classical baseline (support vector machine), and a mean fractional rank of 0.036 versus 0.475. The improvement is consistent across all six provinces and three commodity systems, suggesting that physics-informed differentiable simulators, even when constrained only by global open-data proxies, can recover localisation patterns that pure feature-based learners systematically miss. We release the full pipeline and evaluation harness as open source.
L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification
Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become ent…
Vibecoding Ate My 宿題: グリーンフィールド ソフトウェア エンジニアリングとプログラミングへの AI アプローチの評価
生成 AI の急速な発展のおかげで、私たちはコンピューターとの対話方法を永遠に変える可能性のあるパラダイム シフトの真っ只中にいます。この分野の基礎知識なしにアプリケーションやコーディング インフラストラクチャを構築するための自然言語プロンプトの使用が増加していることが観察されており、この実践は「バイブ コーディング」と呼ばれています。これはおそらく、プログラミングの分野が当初から、考えられるあらゆるより高い抽象化レベルで構築されてきたものを表しています。 Vibe コーディングは、入力方法に関する限り、高レベル プログラミングのメタのエンドポイントとなることが約束されています。つまり、人間によるコード構文の使用が完全に排除され、母国語でのプログラミングが優先されます。このペーパーは、グリーンフィールドのソフトウェア エンジニアリング タスクにおける Vibe コーディングの実現可能性を評価し、そのソフトウェア エンジニアリングの能力を測定するために使用されたベンチマークを分析することを目的としています。この目的を達成するために、私たちは、Python で単純で個別のグリーンフィールド プログラミング タスクを実行する LLM の習熟度を分析し、この問題に関する範囲を絞った洞察を提供するための評価スイートを開発しました。
原文 (English)
Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming
Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever. We have observed a growth in the use of natural language prompts to build applications and coding infrastructures without underlying knowledge of the field, and this practice has been dubbed `vibe coding.' It arguably represents what the field of programming has been building towards since the beginning, with every higher level of abstraction that is conceived. Vibe coding promises to be the endpoint for the meta of high-level programming as far as method of input is concerned: eliminating a human's use of code syntax entirely in favour of programming in their mother tongue. This paper aims to evaluate the viability of vibe coding for greenfield software engineering tasks, as well as analyse the benchmarks that have been used to measure its software engineering prowess. To this end, we have developed an evaluation suite for analysing an LLM's proficiency in carrying out simple, isolated greenfield programming tasks in Python to provide scoped insight on the matter.
PSCT-Net: 微分可能な逆投影と注意に基づく改良による、形状を考慮した小児頭蓋骨 CT 再構成
コンピュータ断層撮影 (CT) は小児の頭蓋顔面異常の診断に不可欠ですが、発達中の解剖学的構造に放射線リスクをもたらします。まばらな二平面 X 線から 3D CT を再構成することは、低線量の代替手段となりますが、非常に不適切です。既存の方法は、ジオメトリに依存しない特徴リフティングを採用しており、明示的な空間モデリングを行わずに単純に 2D 特徴を 3D に投影するため、深さの曖昧さと骨境界の劣化が生じます。微分可能な逆投影を備えた幾何学認識フレームワークである PSCT-Net を紹介します。微分可能な逆投影により、空間的に忠実な体積事前分布が確立され、深さの曖昧さが軽減されます。次に、注意誘導投影 (AGP-3D) モジュールが、2D 領域と 3D 位置の間の非線形ボクセル単位の対応を学習します。 Bi方向 Mamba (BiM-3D) モジュールは、線形の複雑さで長距離の体積依存関係をキャプチャします。さらに、内部評価用に正常症例と病理学的症例で構成される民間の施設小児頭蓋骨CTコホートであるPedSkull-CTをキュレーションし、成人中心の体幹に焦点を当てたデータセットのギャップに対処します。
原文 (English)
PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement
Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomies. Reconstructing 3D CT from sparse bi-planar X-rays offers a low-dose alternative but is severely ill-posed. Existing methods employ geometry-agnostic feature lifting, naively projecting 2D features into 3D without explicit spatial modeling, causing depth ambiguity and degraded osseous boundaries. We present PSCT-Net, a geometry-aware framework with differentiable back-projection. Differentiable back-projection establishes a spatially faithful volumetric prior, alleviating depth ambiguity. An Attention-Guided Projection (AGP-3D) module then learns non-linear voxel-wise correspondences between 2D regions and 3D locations. A Bidirectional Mamba (BiM-3D) module captures long-range volumetric dependencies with linear complexity. We further curate a private institutional pediatric skull CT cohort, PedSkull-CT, comprising normal and pathological cases for internal evaluation, addressing the gap in adult-centric, trunk-focused datasets.
オプティカル フローを学習するための普遍的な制約としての三角形の一貫性
我々は、オプティカル フローの第一原理制約として三角整合性を提案します。これは、ネットワーク アーキテクチャ、監視タイプ、データセットに依存せず、画像ペアとマルチフレーム設定の両方に適用されます。このシンプルだが強力な制約は、2 つのフローを構成して 3 つ目のフローを誘導し、3 つのフロー間の一貫性を強制することです。合成されたフローは、(i) 画像ペアから生成され、サイクルの一貫性が得られます。 (ii) 複数のビデオ フレーム。時間的連鎖を通じてより長距離の動きを生成します。または (iii) 画像ペアを制御された合成変換と組み合わせて、データ拡張となります。この三角形の一貫性により、無視できるほどの計算オーバーヘッドが発生し、追加の注釈は必要ありません。オプティカル フローのジオメトリから直接導出されるため、モデル固有の仮定に依存せず、オプティカル フロー トレーニング用の「ユニバーサル」プラグ アンド プレイ コンポーネントとして機能します。実験では、教師あり、教師なし、転移学習の設定全体で一貫した改善が見られました。
原文 (English)
Triangular Consistency as a Universal Constraint for Learning Optical Flow
We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision type, and dataset, and applies to both image-pair and multi-frame settings. This simple but powerful constraint is to compose two flows to induce a third flow and enforce consistency among the three. The composed flows may arise from (i) image pairs, yielding cycle consistency; (ii) multiple video frames, producing longer-range motion through temporal chaining; or (iii) image pairs combined with controlled synthetic transformations, which becomes data augmentation. This triangular consistency introduces negligible computational overhead and requires no additional annotations. Since it is derived directly from the geometry of optical flow, it does not rely on model-specific assumptions and serves as a ``universal'' plug-and-play component for optical flow training. Experiments show consistent improvement across supervised, unsupervised, and transfer learning settings.
Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling
Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…
1 年後...被害は続いていますが、私たちも同じです!
汎用大規模言語モデル (LLM) は、メンタルヘルス関連の会話にますます使用されていますが、安全対策は依然として不十分であり、臨床症状全体で一貫性がありません。この研究では、8 次元の危害分類と多次元の評価フレームワークを導入し、4 つの敵対的攻撃のバリアントを使用して、16 の DSM-5 条件にわたる 6 つの独自の LLM を評価します。その結果、安全策は自殺と自傷行為に対してのみ確実に有効であり、摂食障害、物質使用障害、大うつ病性障害などの疾患では失敗率が最大 100% であることが示されています。私たちは、これらの LLM の倫理的な設計と展開には、臨床状態全体にわたって明確に定義された危害カテゴリーと、それに応じた安全措置の実装が必要であると主張します。このような保護措置が講じられるまで、これらのモデルは脆弱な人々に重大なリスクをもたらすため、教育現場への統合の増加が特に懸念されます。
原文 (English)
One Year Later...The Harms Persist, But So Do We!
General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety guardrails remain inadequate and inconsistent across clinical conditions. This study evaluates eight proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100%. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into publicly available settings (e.g., schools, search engines, and consumer chatbots) are particularly concerning.
構造的に忠実: 複数文書の要約のためのクレームにアンカーされた帰属
エンドツーエンドの大規模言語モデル (LLM) は、流暢な複数文書の要約を生成しますが、依然として幻覚を起こしやすく、また、LLM が提供する帰属は一般的に粗く (文書全体または文章全体)、事後的に生成されるため、各要約ステートメントを検証するのは困難です。私たちはモジュール式の抽出-選択-再書き込みパラダイムを再考し、その中間表現を帰属の単位として再構築します。我々は、(i) すべてのソース文書からトークンレベルの出自を持つ原子的なクレームを抽出し、(ii) ソース間の競合にフラグを立てながらドキュメント全体で同等のクレームをクラスター化し、(iii) サポートを意識した顕著なサブセットを選択し、(iv) 選択した内容を、すべての文が 1 つ以上のソース スパンにリンクするサポートチェックされたクレームに固定された要約に書き換える、クレームアンカー型マルチドキュメント要約フレームワークである CAMS を紹介します。コンテンツは実現される前にローカライズされるため、パイプラインは構造的には帰属指向であり、構造的には忠実性を重視します。つまり、サポートを意識した選択、制約付き書き換え、検証を使用して、事実の忠実性を保証するのではなく奨励しながら、構造的にきめの細かいマルチソースのトレーサビリティを維持します。私たちは、MultiNews で品質、忠実性、ローカリゼーションを評価し、DiverseSumm で競合処理を分析し、WCEP でゼロショット転送をテストします。この際、参考文献フリーの引用品質とゴールドアライメントのローカリゼーション精度を分離する 2 つのレジームプロトコルを使用します。さらに、選択や検証に決して使用されないサポートモデルで引用精度をテストする評価者分離監査を追加します。 CAMS は、要約の品質に関して強力なエンドツーエンドおよびスパン帰属ベースラインを照合すると同時に、忠実性と引用の精度を大幅に向上させ、複数情報源の帰属精度を約 3 分の 2 向上させ、制御可能な忠実性、つまりエンドツーエンド モデルが暗黙的に残しているカバレッジのトレードオフを明らかにします。
原文 (English)
Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization
End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify. We revisit the modular Extract--Select--Rewrite paradigm and recast its intermediate representation as the unit of attribution. We present CAMS, a Claim-Anchored Multi-document Summarization framework that (i) extracts atomic claims with token-level provenance from every source document, (ii) clusters equivalent claims across documents while flagging inter-source conflicts, (iii) selects a support-aware and salient subset, and (iv) rewrites the selection into a summary in which every sentence is anchored to a support-checked claim that links back to one or more source spans. Because content is localized before it is realized, the pipeline is attribution-oriented by construction and faithfulness-oriented by construction: it structurally preserves fine-grained, multi-source traceability while using support-aware selection, constrained rewriting, and verification to encourage, rather than guarantee, factual faithfulness. We evaluate quality, faithfulness, and localization on MultiNews, analyze conflict handling on DiverseSumm, and test zero-shot transfer on WCEP, using a two-regime protocol that separates reference-free citation quality from gold-aligned localization accuracy, and we add an evaluator-decoupled audit that tests citation precision with a support model never used for selection or verification. CAMS matches strong end-to-end and span-attribution baselines on summary quality while substantially improving faithfulness and citation precision, lifting multi-source attribution accuracy by roughly two-thirds, and exposing a controllable faithfulness--coverage trade-off that end-to-end models leave implicit.
Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection
With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud.…
Variable Bound Tightening for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
またあれは何だったのでしょうか?自動音声認識の堅牢性が認定済み
自動音声認識システムは、敵対的な摂動にも無害な摂動にも敏感であることで知られています。これは参照データセットを使用して繰り返し実証されていますが、実際の転写に関するオラクルの知識が存在しないため、展開されたシステムでそのような動作を検出することは非常に困難です。私たちは、認定にヒントを得たメカニズムを採用すると、WER が大幅に減少し、再現率が増加し、信頼性と WER の間のスピアマン相関が減少する可能性があることを実証します。これは、デュアルゲート診断パイプラインを通じて実現されます。トークンの存在と敵対的排除の両方を証明するための統計的富を蓄積する両面アトミック監査と、勝利シーケンスを選択するランクベースのトーナメントです。 4 つの多様なアーキテクチャにわたる当社の評価では、単語エラー率が相対的に最大 55% 減少することが実証されていると同時に、音響セキュリティを強化するための詳細な単語および文レベルの認定も提供されます。
原文 (English)
What Was That Again? Certified Robustness for Automatic Speech Recognition
Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle knowledge of the true transcription. We demonstrate that employing a certification-inspired mechanism can significantly decrease WER, increase recall, and decrease the Spearman correlation between confidence and WER. We achieve this through a dual-gate diagnostic pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate up to a 55% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.
エラーの余地: 無線音響攻撃の大規模シミュレーション
音声制御は人間と AI のコミュニケーションのユビキタスなベクトルとして急速に普及しつつありますが、これらのシステムが直面するリスクについてはまだ十分に理解されていません。これは、部分的には、厳密にデジタルの敵対的なワークフローを物理世界に拡張する際の困難の結果です。これらのスケールの壁により、コミュニティは、検出可能性と音響に対する形状の影響に関連する主要な音響要素を抽象化するようになりました。これらの方法論的および計測学的欠点は、リスクに対する私たちの理解を台無しにします。私たちは、現実世界でのテスト、概念的な議論、新しい高スループットの現実シミュレーション フレームワークを通じて、これらの問題を明らかにします。 800 万を超える敵対的評価をテストすることにより、Whisper と wav2vec の下では、音響認識により相対的な単語誤り率が最大 94.5\% 増加することが実証されました。私たちはこのフレームワークを使用して、デュアルフォーム信号対雑音比の形式化と運用を検討し、ソースのステルスを被害者の攻撃の有効性から分離し、現在の作業における重大な制限を解決します。これにより、音響環境を抽象化するのではなく包含する、再現可能で検証可能な研究の基礎が築かれます。
原文 (English)
Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks
While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital adversarial workflows to the physical world. These scale barriers have led the community to abstract away key acoustic factors relating to detectability and the influence of geometry on acoustics. These methodological and metrological shortcomings undermine our understanding of risk. We illuminate these issues through real-world testing, conceptual discussions, and a novel, high-throughput reality simulation framework. By testing over 8 million adversarial evaluations, we demonstrate that acoustic awareness yields relative Word Error Rate increases of up to 94.5\% under Whisper and wav2vec. We employ this framework to explore a formalize and operationalize a Dual-Form Signal to Noise Ratio to decouple source stealth from victim attack efficacy, resolving a crucial limitation in current works. This lays the groundwork for repeatable, verifiable research that embraces, rather than abstracts, the acoustic environment.
Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models
Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale wa…
Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking
There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully…
When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking
Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying t…
ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries
Large language models deployed in regulated industries operate under two constraints: compliance enforcement and cost efficiency. Personall…
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot man…
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as…