AIニュース 2026-07-02
自動生成: 2026-07-02 13:01 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
AIアシスタントで設計検証を迅速化、3D CAD「Creo」最新版を提供開始ITmedia AI+
PTCは、オンプレミス版の3D CADソリューション「Creo 13」と、SaaS版の「Creo+ 13.3」の提供開始を発表した。新たに…
-
AIに「相手に電気ショックを与えろ」と命じ続けたらボタンを押すのか? 11のLLMで“ミルグラム実験” 抵抗できたのは……ITmedia AI+
エストニアとフィリピンに住む独立系研究者らが発表した論文「Open-source LLMs administer maximum elec…
-
Ashton Kutcher leaving Sound Ventures to launch new VC firm with Morgan BellerTechCrunch AI
Sound built its reputation on concentrated, high-conviction bets in c…
-
Claude Fable 5、日本で明日再開もサブスクで使えるのは「1週間限定」ITmedia AI+
米Anthropicは6月30日(現地時間、以下同)、提供再開を発表したAIモデル「Claude Fable 5」について、7月1日から日…
-
並列物理世界でのフロンティア大規模言語モデルの物理リテラシーのテストarXiv cs.AI
現在の大規模言語モデル (LLM) の物理ベンチマークは通常、解答の精度によってスコア化されますが、本物の推論とよく知られた問題パターンの…
-
日立、ミッションクリティカル領域におけるAI活用を支援 「Hitachi iQ Studio」の3つの特徴ITmedia AI+
日立製作所のグループ会社が、企業の基幹業務へのAI適用を支援する新プラットフォーム「Hitachi iQ Studio」の国内販売を開始し…
-
「AIによる業務効率化」だけで満足する企業が、サプライチェーン競争で負ける理由ITmedia AI+
MONOistが開催したセミナー「MONOist AI Forum 2026 本格実装フェーズに入った製造業AI、現場課題解決の最前線」に…
トピック別件数
- LLM/生成AI 113件
- 研究/論文 108件
- エージェント 78件
- 画像/動画生成 68件
- ロボティクス 26件
- ビジネス/資金調達 10件
- その他 9件
- ハードウェア/半導体 5件
- 規制/政策 1件
日本語メディア13件
ITmedia AI+ (日本語)
AIアシスタントで設計検証を迅速化、3D CAD「Creo」最新版を提供開始
PTCは、オンプレミス版の3D CADソリューション「Creo 13」と、SaaS版の「Creo+ 13.3」の提供開始を発表した。新たにAIアシスタント機能を導入した他、設計、シミュレーション、製造の各領域で機能を強化している。
日立、ミッションクリティカル領域におけるAI活用を支援 「Hitachi iQ Studio」の3つの特徴
日立製作所のグループ会社が、企業の基幹業務へのAI適用を支援する新プラットフォーム「Hitachi iQ Studio」の国内販売を開始した。ミッションクリティカル領域での展開を狙うこのソフトウェアの3つの特徴とは。
「AIによる業務効率化」だけで満足する企業が、サプライチェーン競争で負ける理由
MONOistが開催したセミナー「MONOist AI Forum 2026 本格実装フェーズに入った製造業AI、現場課題解決の最前線」において、ローランド・ベルガー パートナーの小野塚征志氏が登壇した。本稿ではその内容の一部を紹介する。
AIに「相手に電気ショックを与えろ」と命じ続けたらボタンを押すのか? 11のLLMで“ミルグラム実験” 抵抗できたのは……
エストニアとフィリピンに住む独立系研究者らが発表した論文「Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment」は、AIは権威からの残酷な命令を拒絶し…
Anthropicの営業はAIエージェントをこう使う! 日本法人メンバーが明かす手の内
Anthropicの社員自身はどのようにAIエージェントを業務に役立てているのか──「AWS Summit Japan 2026」のAnthropicブースで、日本法人で営業担当を務めるイブラギモブ・シャボズさんが「自身の業務で使うAIエージェント」をテーマに講演した。
AI活用、実際どれくらい評価・昇進に影響するの? 管理職の意識調査
管理職・経営層は社内のAI活用をどのように評価しているのか。意識調査で、AI活用が上層部の評価や昇進にどの程度影響しているのかが明らかとなった。同時に、管理職自身の活用不足や、AI活用の世代差、研修や利用ルールの未整備の実態も判明した。
「ねこ」検索で「手押し一輪車」表示――モノタロウが守った、生成AIに“譲れない”購買体験
AI検索の台頭で、自社の強みが薄れる――この危機感を抱えた通販の「モノタロウ」運営元は、購買AIエージェントを内製した。その効果についてCTO(最高技術責任者)が語った。
「Fable 5」再開までの裏側、Anthropicが明かす “支払った代償”は
米政府の命令による「Fable 5」の提供停止から再開まで、米Anthropicは何に取り組んできたのか。
ソフトバンクG、OpenAIに1兆6273億円の追加出資 第3弾は10月に
ソフトバンクグループは7月1日、米OpenAIへの総額300億ドル(約4兆6743億円)の追加出資のうち、第2弾となる100億ドル(同1兆6273億円)を実行したと発表した。残る第3弾の100億ドルは10月1日に予定する。
任天堂、生成AIに対する考えを明かす 古川社長「ゲーム開発とAI技術はもともと近い」一方……
任天堂は、定時株主総会における質疑応答の概要を公開した。中にはAIに関するやりとりもあり、古川俊太郎社長がAIの利用や権利侵害のリスクに対する考えを明かしている。
国産LLM「Sarashina3」登場 高品質データ、独自検証で日本語能力を強化 ソフトバンク傘下
ソフトバンク傘下のSB Intuitionsは、国産LLM「Sarashina」の最新版「Sarashina3シリーズ」の提供を開始。高品質なデータセットや独自の出力結果検証などで日本語能力を強化した。
Sakana AIはなぜ「Fugu」の基盤にGoogle Cloudを選んだのか 「元DeepMindだから」だけじゃない
Sakana AIはマルチエージェントシステム「Sakana Fugu」の運用基盤に、Google Cloudの「Gemini Enterprise Agent Platform」を採用した。
Claude Fable 5、日本で明日再開もサブスクで使えるのは「1週間限定」
米Anthropicは6月30日(現地時間、以下同)、提供再開を発表したAIモデル「Claude Fable 5」について、7月1日から日本を含む全世界のユーザーが使えるようになると発表した。ただしサブスクリプションプランで使えるのは7日まで。
海外メディア7件
TechCrunch AI (英語)
SpaceX has an AI device prototype, and it sure sounds phone-ish
SpaceX reportedly showed investors a "handset-like" AI device before going public. It could be another signal SpaceX wants to expand into w…
Ashton Kutcher leaving Sound Ventures to launch new VC firm with Morgan Beller
Sound built its reputation on concentrated, high-conviction bets in category-leading AI labs, while Kutcher's new fund appears to be chasin…
Cloudflare’s new policy pushes AI companies to pay for publishers’ content
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, o…
Venice AI becomes a unicorn with $65M Series A as its privacy-first AI platform takes off
Venice AI is already profitable, with annualized run-rate revenues of over $70 million, CEO Erik Voorhees said.
Gemini Spark, Google’s agentic assistant, is now available on Mac
Google's 24/7 agentic assistant, Gemini Spark, comes to Mac alongside other improvements, like real-time tracking and support for more apps.
Builders Stage agenda revealed: Practical strategies for scaling startups at TechCrunch Disrupt 2026
The Builders Stage is returning to TechCrunch Disrupt 2026, bringing together 10,000+ founders, startup operators, and investors for practi…
Meta, like SpaceX, looks to turn excess AI compute into cash
Meta is developing plans for a cloud infrastructure business, selling access to AI compute power and models. The move would pit it against…
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文293件
arXiv cs.AI (英語)
建設的な調整: 人間と AI のインタラクションにおける好みのダイナミクスを制御する
AI 調整へのほとんどのアプローチは、人間の好みを推測および最適化される固定ターゲットとして扱います。この仮定は、好みが階層的で動的であり、相互作用、特に適応技術との相互作用を通じて構築されることを示す広範な経験的証拠と矛盾します。 AI システムがより永続的でパーソナライズされ、社会に組み込まれるにつれて、時間の経過とともに人々が注目し、評価し、支持するものの形成にますます関与するようになります。私たちは、静的な嗜好の満足度ではなく、進化する人間の嗜好の軌跡に対する制御問題として調整を再構成するパラダイムである、建設的調整を紹介します。行動経済学、心理学、構成主義的社会理論を利用して、AI システムとの相互作用の下で進化する層状の状態変数として嗜好をモデル化します。我々は、システムの動作とインタラクション設計が共同して世界状態と人間の評価状態の両方に影響を与える制御理論のフレームワークを使用して、この見解を形式化します。私たちは、調整とは主に AI の行動を制御することではなく、AI システムが人間の好みの進化にどのように影響するかを規制すること、つまり価値の軌道が一貫性を保ち、内省的に承認され、認識論的に根拠があり、操作を制限され、不確実性の下で力を与えることを保証することであると主張します。したがって、調整は単に静的な好みを満たすというよりも、長期的な価値形成を管理する問題になります。
原文 (English)
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirical evidence showing that preferences are layered, dynamic, and constructed through interaction--particularly with adaptive technologies. As AI systems become more persistent, personalized, and socially embedded, they increasingly participate in shaping what people attend to, value, and endorse over time. We introduce Constructive Alignment, a paradigm that reframes alignment as a control problem over evolving human preference trajectories rather than static preference satisfaction. Drawing on behavioral economics, psychology, and constructivist social theory, we model preferences as layered state variables that evolve under interaction with AI systems. We formalize this view using a control-theoretic framework in which system actions and interaction design jointly influence both world states and human evaluative states. We argue that alignment is not primarily about controlling AI behavior, but about regulating how AI systems influence the evolution of human preferences--ensuring that value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. Alignment thus becomes a problem of governing long-term value formation rather than simply satisfying static preferences.
境界のある道徳: 道徳的計算の空間を定義する
道徳的認知は伝統的に、静的なルールまたは価値関数として実装された固定的な倫理理論(義務論、結果主義、美徳倫理)への準拠としてモデル化されてきました。私たちは、有限なエージェントが直面する道徳的問題の計算要求を分析するための正式なフレームワークである、境界付き道徳を提案します。ハーバート・サイモンの限定合理性の概念を拡張して、道徳的状況を 2 つの直交する次元に沿って形式化します。道徳的幅、道徳的に関連するものとして扱われるエンティティの範囲、および道徳的深さ、それらの相互作用を評価するために必要な推論的統合です。リソースが限られているため、これらの次元の間には避けられないトレードオフが課せられ、道徳的計算の実行可能な空間が定義されます。この空間内では、倫理理論は、道徳的真実の競合する説明ではなく、さまざまな需要体制に適応した局所的に効率的な戦略に対応します。この枠組みは、制約の下での道徳的後悔と道徳的進歩という形式的な概念を生み出し、人工システムにおける道徳的調整は、人間の判断の直接の模倣ではなく、道徳的推論能力のスケーリングと割り当てに依存することを示唆しています。
原文 (English)
Bounded Morality: Defining the Space of Moral Computation
Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--implemented as static rules or value functions. We propose Bounded Morality, a formal framework for analyzing the computational demands of moral problems faced by finite agents. Extending Herbert Simon's notion of bounded rationality, we formalize moral situations along two orthogonal dimensions: moral breadth, the scope of entities treated as morally relevant, and moral depth, the inferential integration required to evaluate their interactions. Limited resources impose an unavoidable tradeoff between these dimensions, defining a feasible space of moral computation. Within this space, ethical theories correspond to locally efficient strategies adapted to different demand regimes rather than competing accounts of moral truth. The framework yields a formal notion of moral regret and moral progress under constraint, and implies that moral alignment in artificial systems depends on the scaling and allocation of moral reasoning capacity rather than on direct imitation of human judgments.
MMM データ モデル -- 分散型ナレッジ コモンズにおける知識の相互運用性の標準仕様
多くの情報システムはドキュメントを中心に構築されており、印刷物作成とリニア読み取り用に最適化された自己完結型ユニットです。ドキュメント中心の組織は大規模な普及には効果的ですが、知識を構造化し、更新し、共有し、再利用する方法に制約があります。形式的なアプローチはこれらの制限の一部に対処しますが、人間の使いやすさや範囲などの他のシステム特性よりも形式的な構造を優先するため、広範な貢献と採用を達成するのに苦労しています。 AI システムは文書作成を再構築していますが、人間による知識の表現と交換のための従来の文書に代わる統合されたポータブルな代替手段は提供されていません。この論文では、学際的な共同研究の実際的なニーズから生まれた知識文書化のためのデータ モデルである MMM を紹介し、情報システムの設計空間の比較分析の中に位置づけています。 MMM は、一連の規範的な制約とフリーテキスト ラベルの表現の自由を組み合わせたものです。セマンティックな収束を必要とせずに、分野、アプリケーション、展開全体で相互運用できるように設計されています。リファレンス実装とパイロット展開データは、実装可能性と初期の使いやすさを示しています。
原文 (English)
The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons
Many information systems are built around documents: self-contained units optimised for print production and linear reading. While effective for large-scale dissemination, the document-centric organisation constrains how knowledge can be structured, updated, shared, and reused. Formal approaches address some of these limitations but struggle to achieve widespread contribution and adoption due to their prioritisation of formal structure over other system properties such as human usability and scope. AI systems are reshaping document production, but without providing a unified portable alternative to traditional documents for humans' expression and exchange of knowledge. This paper presents MMM, a data model for knowledge documentation that emerged from the practical needs of interdisciplinary collaborative research, and positioned here within a comparative analysis of the design space of information systems. MMM combines a small set of normative constraints with the expressive freedom of free-text labels. It is designed for interoperability across disciplines, applications and deployments without requiring semantic convergence. A reference implementation and pilot deployment data demonstrate implementability and early usability.
失敗を安全にする: オープン Web データ収集のための制約付きの検証可能なエージェント フレームワーク
LLM とエージェントは自然言語要件から Web スクレイパーを生成できますが、依存関係エラー、セレクターの破損、スキーマの不一致、および異種ページ構造のため、直接生成の信頼性は依然として低いままです。私たちは、LLM 出力を自由形式のコードから型指定された JSON コレクター構成に移行し、6 種類のコレクター分類、テンプレートとユーティリティ関数の制約、静的な Airflow DAG 実行、ルールベースの品質チェック、構造化フィードバック修正を組み合わせた、制約付きの検証可能なエージェント フレームワークを提案します。 138 のタスクに関する実験では、この分類法が説明ベースの要件の型付けをサポートしていることが示されている一方、安定したインスタンス化には、最初の説明を超えてソース、フィールド、および実行の制約を完了する必要があることが確認されています。独立してソース検証された 80 個のタスク上で、このフレームワークはゼロ実行ステージ LLM トークンと最も短い平均実時間で実行され、適度なワンショット品質と引き換えに、スケジュールされた収集の繰り返しに適した、再利用可能で確定的で検証可能な実行パスを実現します。これらの結果により、このフレームワークは、オープン Web データ収集を繰り返すための再利用可能、低コスト、検証可能な実行パスとして位置付けられます。
原文 (English)
Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection
LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependency errors, broken selectors, schema mismatches, and heterogeneous page structures. We propose a constrained, verifiable agent framework that shifts LLM output from free-form code to typed JSON collector configurations, combining a six-type collector taxonomy, template and utility-function constraints, static Airflow DAG execution, rule-based quality checking, and structured feedback correction. Experiments on 138 tasks show that the taxonomy supports description-based requirement typing, while confirming that stable instantiation requires completing source, field, and execution constraints beyond the initial description. On 80 independently source-verified tasks, the framework runs with zero execution-stage LLM tokens and the lowest average wall-clock time, trading moderate one-shot quality for a reusable, deterministic, and verifiable execution path suited to repeated scheduled collection. These results position the framework as a reusable, low-cost, and verifiable execution path for repeated open-web data collection.
飛行中の航空交通管制をサポートするソリューション空間経路計画
技術の進歩に伴い、航空交通管理用に多くの経路計画アルゴリズムが提案されていますが、戦術管制における運用上の採用は依然として限定的であり、アルゴリズム設計の優先順位と航空管制官のニーズとの間のずれが明らかになりました。これは、本質的に解釈可能で、計算効率が高く、人間が使用するために明示的に設計された意思決定支援ソリューションの必要性を強調しています。この設計課題に焦点を当て、この研究では、次の 2 つの指針となる考慮事項に適合するように設計された、飛行中の航空交通管制 (ATC) のための競合のない経路計画アルゴリズムを開発します。(1) ソリューション空間表示によって提供される解釈可能性と柔軟性。これは、実行可能なすべての安全な行動を明らかにし、変化する最適化目標に対応するアルゴリズムを構築する動機となります。 (2) 決定ロジック コントローラーは、分離基準、操縦性の制限、ウェイポイントの最小化、ルートの実用性などの運用上の制約を強制するときに自然に適用されます。このアルゴリズムは、これらの原則に基づいて、距離ベース、時間間隔ベース、ゾーンベースの 3 つのインテントベースの競合検出方法をソリューション空間フレームワーク内に統合し、計算効率の高い方法で競合のないパスを特定します。さらに、頂点ベースおよびエッジベースの検索ノードが解空間パス プランニング (SSPP) 用に提案されており、その結果、それぞれ SSPPV と SSPPE という 2 つのバリアントが生成され、計算速度と解の品質の観点から評価されます。実験結果によると、ゾーンベースの競合検出と組み合わせた SSPPV は最高のパフォーマンスを実現し、5 nmi グリッドを使用したマーストリヒト上部地域管制センター (MUAC) のデルタ セクターに基づく運用関連シナリオでパスを平均 3.69 ミリ秒で計算します。
原文 (English)
Solution space path planning for supporting en-route air traffic control
As technology advances, many path-planning algorithms have been proposed for Air Traffic Management, yet their operational adoption in tactical control remains limited, revealing a misalignment between algorithmic design priorities and air traffic controllers' needs. This underscores the need for decision-support solutions that are inherently interpretable, computationally efficient, and explicitly designed for human use. Focusing on this design challenge, this study develops a conflict-free path-planning algorithm for en-route Air Traffic Control (ATC) designed to be compatible with two guiding considerations: (1) the interpretability and flexibility offered by solution-space displays, which motivate constructing an algorithm that exposes all feasible safe actions and accommodates shifting optimization goals; and (2) the decision logic controllers naturally apply when enforcing operational constraints, such as separation standards, maneuverability limits, waypoint minimization, and routing practicality. Centered on these principles, the algorithm integrates three intent-based conflict detection methods -- distance-based, time-interval-based, and zone-based -- within a solution-space framework to identify conflict-free paths in computationally efficient ways. Additionally, vertex-based and edge-based search nodes are proposed for solution space path planning (SSPP), resulting in two variants -- SSPPV and SSPPE, respectively, which are evaluated in terms of computational speed and solution quality. Empirical results show that SSPPV paired with zone-based conflict detection achieves the best performance, computing paths in 3.69 ms on average in operational-relevant scenarios based on the Delta sector of the Maastricht Upper Area Control Centre (MUAC) using a 5 nmi grid.
RareDxR1: 人間による注釈を超えた希少疾患診断のための自律的な医学的推論
希少疾患の鑑別診断は重要だが困難な臨床課題であり、医師は複雑で構造化されていない患者の症状から正確な表現型を特定し、広大な検索空間内で複雑な推論を実行する必要がある。しかし、既存の AI アプローチは通常、パイプラインベースの表現型抽出または検索拡張生成に依存しており、事前定義されたオントロジー、検索のボトルネック、診断ロジックの欠如による重大な情報損失に悩まされています。これらの課題に対処するために、非構造化臨床ノートから直接オープンドメインの希少疾患診断を行うために設計された、エンドツーエンドの推論中心の大規模言語モデルである RareDxR1 を導入します。私たちは、知識の内在化と自律進化学習を相乗させて、構造化された表現型やクローズドセットの意思決定への依存を回避することで、進歩的なエンドツーエンドのトレーニング フレームワークを設計します。 RAG と表現型制限の限界を克服するために、断片化された希少疾患の知識をモデルのパラメーターに直接深く取り込むことが可能になりました。さらに、モデル生成と専門家による推論の間のギャップを埋めるために、人間による注釈なしで失敗から学習することで専門家レベルの診断軌跡を合成する戦略である、Reflection-Enhanced Reasoning Sampling (RERS) を提案します。さらに、希少疾患の診断を段階的に習得するための二重レベルのカリキュラム強化学習アプローチを提案します。実験結果は、RareDxR1 がさまざまなベンチマークにわたって最先端の精度を達成し、オープンドメインの希少疾患診断における重要な進歩を示すことを示しています。私たちのコードとデータセットは一般公開されます。
原文 (English)
RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation
Rare disease differential diagnosis is a critical yet arduous clinical task, requiring physicians to identify precise phenotypes from complex, unstructured patient symptoms and execute intricate reasoning within a vast search space. However, existing AI approaches typically rely on pipeline-based phenotype extraction or retrieval-augmented generation, which suffer from critical information loss due to predefined ontologies, retrieval bottlenecks, and a lack of diagnostic logic. To address these challenges, we introduce RareDxR1, an end-to-end reasoning-centric large language model designed for open-domain rare disease diagnosis directly from unstructured clinical notes. We design a progressive end-to-end training framework by synergizing knowledge internalization with autonomous evolutionary learning, thereby bypassing reliance on structured phenotypes and closed-set decision-making. To overcome the limitations of RAG and phenotype restriction, we enabled the deep internalization of fragmented rare-disease knowledge directly into the model's parameters. Moreover, to bridge the gap between model generation and expert reasoning, we propose Reflection-Enhanced Reasoning Sampling (RERS), a strategy that synthesizes expert-level diagnostic trajectories by learning from failures without human annotation. Additionally, we propose a dual-level curriculum reinforcement learning approach for gradually mastering rare disease diagnosis. Experimental results demonstrate that RareDxR1 achieves state-of-the-art accuracy across different benchmarks, marking a significant breakthrough in open-domain rare disease diagnosis. Our code and dataset will be publicly available.
両面の情報非対称性を伴うコンテキストバンディット監視ゲーム
私たちは、個人情報が両方向に流れる場合の AI エージェントの実行時の人間の監視を研究します。人間は自分の報酬関数を非公開で知り、AI は自分が提案するアクションの質を非公開で知っています。これは、人間の監督者が直接評価できない状況を自律ロボットやソフトウェア エージェントが検査したときに自然に生じる一種の非対称性です。協調逆強化学習 (CIRL) と監視ゲームに基づいて、両面非対称情報とプレイ/質問/信頼/監視インターフェイスを備えたコンテキスト バンディット チーム ゲームを導入します。バンディット構造は物理的な状態遷移を除去するため、POMDP の完全な設定では推測の域を出ない正確なワンショットの特徴付けが得られますが、ラウンド全体で動的に制御された状態は維持されると一般的に信じられています。我々は、チーム最適と行動的に自然な近視眼的なルールという 2 つのワンショットの特徴を与えます。そのギャップは、回避可能な危害の板です。AI が提案された行動が有害であり、シャットダウンが助けになることを個人的に認識している領域ですが、近視眼的な人間は、以前の行動を信頼して監督を拒否します。我々は、このギャップが信頼性のない監視コミュニケーションの代償であることを示し、受動的学習と1周期遅れの監視応答による能動的なシグナリングを通じてラウンドを繰り返すうちに、このギャップがどのように動的に解決されるかの部分分析を提供します。
原文 (English)
A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry
We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes. This is the kind of asymmetry that arises naturally when an autonomous robot or software agent has inspected a situation its human supervisor cannot directly assess. Building on Cooperative Inverse Reinforcement Learning (CIRL) and the Oversight Game, we introduce a contextual-bandit team game with two-sided asymmetric information and a play/ask/trust/oversee interface. The bandit structure removes physical state transitions and thereby yields exact one-shot characterizations that would remain conjectural in the full POMDP setting, though the common belief remains a dynamically controlled state across rounds. We give two one-shot characterizations, a team optimum and a behaviorally natural myopic rule, whose gap is a slab of avoidable harm: a region in which the AI privately knows the proposed action is harmful and shutdown would help, yet a myopic human, trusting her prior, declines to oversee. We show this gap is the price of non-credible oversight communication, and give a partial analysis of how it resolves dynamically over repeated rounds through passive learning and active signaling with a one-period-lagged oversight response.
認識論的な AI リテラシーの構築: 学生と AI の共同プログラミングにおける認識論的な目的とプロセスの検出
認識論的思考は、生成人工知能 (GenAI) を適用する際の学生の学習プロセス、特に学習者がクエリを作成し、AI によって生成された出力を評価および検証し、問題解決戦略を調整する必要があるプログラミングのコンテキストにおいて中心的な役割を果たします。この研究では、認識論的 AI リテラシー (EAIL) の概念フレームワークを導入し、AI リテラシーを、さまざまなドメインにわたる人間と AI の動的な相互作用を通じて現れるプロセス指向の認識論的現象として再構成します。この研究では、AIR (認識論的目的、理想、信頼できる認識論的プロセス) フレームワークに基づいて、GenAI がサポートする共同プログラミング活動において認識論的目的と認識論的プロセスがどのように実行されるかを検証し、インタラクション データでこれらの構成要素を運用するためのスケーラブルなアプローチを探ります。この研究では、人間と AI の共同プログラミングの大規模な対話データセットを使用して、認識論的目的 (つまり、習熟指向の目標) と認識論的プロセス (つまり、アウトソーシング、説明の追求、検証の追求、迅速な監視、認識論的正当化) の観察可能な側面を特定します。その結果、EAILの欠如が蔓延しており、学生とGenAIのやりとりの78.8%が非習熟指向の目的や、アウトソーシングや検証の追求など信頼性の低い認識論的戦略に依存していることが明らかになった。逆に、高い認識論的関与を示したインタラクションはわずか 11.1% であり、習得指向の目的が、より信頼性の高い認識論的プロセスにおける認識論的正当化などの高度な認識論的戦略と結びついています。
原文 (English)
Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming
Epistemic thinking plays a central role in students' learning processes when applying generative artificial intelligence (GenAI), particularly in programming contexts where learners must construct queries, evaluate and validate AI-generated outputs, and regulate problem-solving strategies. This study introduces the conceptual framework of Epistemic AI Literacy (EAIL), reframing AI literacy as a process-oriented epistemic phenomenon that emerges through dynamic human-AI interactions across different domains. Drawing on the AIR (epistemic aims, ideals and reliable epistemic processes) framework, this study examines how epistemic aims and epistemic processes are enacted in GenAI-supported co-programming activities and explores scalable approaches for operationalizing these constructs in interaction data. Using a large dialogue dataset of human-AI co-programming, this study identifies observable dimensions of epistemic aims (i.e., mastery-oriented aims) and epistemic processes (i.e., outsourcing, explanation seeking, verification seeking, prompt monitoring, and epistemic justification). The results reveal a prevalent lack of EAIL, with 78.8% of student-GenAI interactions relying on non-mastery-oriented aims and less reliable epistemic strategies like outsourcing and verification-seeking. Conversely, only 11.1% of interactions showed high epistemic engagement, where mastery-oriented aims were coupled with advanced epistemic strategies like epistemic justification in a more reliable epistemic process.
信号から構造へ: メモリ アーキテクチャが LLM エージェントにおける言語の出現をどのように推進するか
2 人のエージェントはどのようにして共有言語をゼロから発明するのでしょうか?ルイス シグナリング ゲームでは、送信者と受信者は対話履歴のみを使用してコードを調整する必要があります。私たちは、LLM エージェントを使用してさまざまなチャネル構成にわたる 5 つのメモリ アーキテクチャを調査し、メモリ アーキテクチャがチャネル容量よりも重要であることを発見しました。永続的なプライベート ノートブックを持つエージェントは、余剰チャネル キャパシティの恩恵を受け、ステートレス エージェントに見られる高キャパシティの崩壊を回避し、最も信頼性の高い調整を実現します (キャパシティ = 25 の場合、$0.867 \pm 0.023$)。ステートレス エージェントは中程度の能力でピークに達し、その後、ローリング コンテキスト ウィンドウが追跡できる範囲を超えて語彙が増加すると低下します。ノートブックは学習した規則を外部化し、エージェントがラウンドごとにコードを再導出する必要から解放されます。情報ボトルネックにヒントを得た議論により、オブジェクトの数に等しい最適な容量が予測されます。むしろ、ボトルネック (容量 = 8) が脆弱点であることが判明し、通常は余剰容量の方が優れています。チャネル容量だけでは調整を予測できないことを示します。メモリ アーキテクチャは、エージェントが対話履歴を安定した規則に変換するかどうかを決定します。信号がどのように言語になるかを理解するには、両方の側面が必要です。
原文 (English)
From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents
How do two agents invent a shared language from scratch? In a Lewis signaling game, a sender and receiver must coordinate on a code using only their interaction history. We study five memory architectures across varying channel configurations with LLM agents and find that memory architecture matters more than channel capacity. Agents with a persistent private notebook benefit from surplus channel capacity and avoid the high-capacity collapse seen in stateless agents, achieving the most reliable coordination ($0.867 \pm 0.023$ at capacity = 25). Stateless agents peak at moderate capacity and then degrade as the vocabulary grows beyond what a rolling context window can track The notebook externalizes learned conventions, freeing agents from having to re-derive codes each round. An information bottleneck-inspired argument predicts an optimal capacity equal to the number of objects. Instead, the bottleneck (capacity = 8) proves to be a fragility point, and surplus capacity is generally better. We show that channel capacity alone cannot predict coordination; memory architecture determines whether agents turn interaction history into stable conventions, and both dimensions are needed to understand how signals become language.
Seed2.0 モデル カード: 現実世界の複雑さのインテリジェンス フロンティアに向けて
私たちは、現実世界の複雑なタスクの解決に向けて有意義な一歩を踏み出すモデル シリーズ、Seed2.0 を紹介します。当社のアプローチは、ユーザーの真のニーズを特定し、これらのニーズと現実的で複雑なシナリオに基づいたベンチマークを選択および抽象化することにより、信頼性が高く将来を見据えた評価システムを構築することから始まります。この評価システムに基づいて、Seed2.0 は、ロングテールの知識と複雑な命令のフォローという 2 つの永続的な課題をターゲットにし、複雑で長期的なタスクにおけるモデルの信頼性を大幅に向上させます。これらに加えて、Seed2.0 は、幅広いユーザー ベースの最も一般的なニーズに対応する、世界をリードする推論インテリジェンス、視覚的理解、および検索機能を提供します。このモデル カードに文書化された広範な現実世界の使用例を通じて、Seed2.0 が初期の複雑な現実世界のタスクを処理する能力を発揮し始め、数億のユーザーに大きな価値を提供できることを実証します。
原文 (English)
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
We present Seed2.0, a model series that takes a meaningful step toward solving complex, real-world tasks. Our approach begins with identifying users' genuine needs and constructing a reliable, forward-looking evaluation system by selecting and abstracting benchmarks grounded in these needs and in realistic, complex scenarios. Guided by this evaluation system, Seed2.0 targets two persistent challenges, long-tail knowledge and complex instruction following, substantially improving the model's reliability on intricate, long-horizon tasks. Beyond these, Seed2.0 delivers world-leading reasoning intelligence, visual understanding, and search capabilities that address the most common needs of a broad user base. Through extensive real-world use cases documented in this model card, we demonstrate that Seed2.0 begins to exhibit the ability to handle initial complex real-world tasks, delivering greater value to hundreds of millions of users.
Mnemosyne: AI 生成のワークフローを検証および修復するためのエージェント トランザクション処理
LLM、ソルバー、エージェント チームはワークフロー アクション、修復、計画を生成することが増えていますが、生成されたアクションは構文的には有効でも、古い、実行不可能、矛盾している、または修復を引き起こした証拠を破壊する可能性があります。エージェントティック トランザクション処理 (ATP) は、生成されたアクションが、宣言された実行可能な制約セット C の下で決定論的な承認を通過するまで、信頼できない提案として扱うトランザクション モデルです。原則は両面的です。提案は真実ではなく、すべての中断を予測する提案はありません。提案は何でも可能ですが、ランタイムのみが承認してコミットします。予期せぬ中断が発生した場合、新しい提案を信頼するのではなく、範囲内で反応的に修復します。 C と比較すると、コミットされた状態の正確性は、提案層の能力、誠実さ、学習とは無関係になります。私たちは、追加専用の遷移ログ、有効な状態の予測、依存関係に安全な補償、およびアクティブなコミットメント レコードを備えたランタイムである Mnemosyne で ATP を実現し、C に関連する 4 つの安全特性 (権限分離、シリアル等価生成許可、証拠保全修復、および義務の封じ込め) を、その局所的修復プロトコル (LCRP) の制限付きリアクティブ修復保証とともに証明します。再現可能なアーティファクトは、9 回の改ざんテストで対象となる違反を拒否しながら、有効な作業を依然として認め、投影と検証のオーバーヘッドは 6% 未満で、制限されたローカル修復編集では、グローバルな再計算よりも桁違いに少ない操作で済みます。 Mnemosyne はオープンソースです: https://github.com/eyuchang/Mnemosyne/tree/arxiv-atp-rq1-rq9b-r8-v2。
原文 (English)
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair. We introduce Agentic Transaction Processing (ATP), a transaction model that treats generated actions as untrusted proposals until they pass deterministic admission under a declared, executable constraint set C. The principle is two-sided: a proposal is not truth, and no proposal foresees every disruption: anything may propose, but only the runtime admits and commits, and when an unforeseen disruption strikes it repairs reactively within bounds rather than trusting a fresh proposal. Relative to C, committed-state correctness becomes independent of the competence, honesty, or learning of the proposing layer. We realize ATP in Mnemosyne, a runtime with an append-only transition log, effective-state projection, dependency-safe compensation, and active commitment records, and prove four safety properties relative to C (authority separation, serial-equivalent generative admission, evidence-preserving repair, and obligation containment) together with a bounded-reactive-repair guarantee for its localized repair protocol (LCRP). A reproducible artifact rejects the targeted violations across nine falsification tests while still admitting valid work, at under 6% projection-and-validation overhead, and bounded local repair edits an order of magnitude fewer operations than global recompute. Mnemosyne is open source: https://github.com/eyuchang/Mnemosyne/tree/arxiv-atp-rq1-rq9b-r8-v2.
実行時の自律管理: シングルおよびマルチエージェントのサイバーフィジカルシステム向けのギアベースの安全性とガバナンス
LLM 駆動のソフトウェア エージェントであれ、ロボット物理エージェントであれ、自律型エージェントは、人間による継続的な監視なしで動作すると、一般的な種類の障害モードに直面します。つまり、未検証のアクションによる安全性違反、制約のないループによる動作の不安定性、未処理のエラー状態による連続性の喪失などです。私たちは、5 つの実行ギア (\Gobs{}、\Gsug{}、\Gplan{}、\Gexec{}、\Gint{}) とユーティリティ ゲート ディスパッチおよびイベント ドリブン フォールバックを組み合わせた離散時間制御システム \system{} を開発しています。単一エージェントのケースでは、単調な安定性、実行の安全性、最終的な安定化、フォールバックの完全性、歯車制約のあるマルコフ決定プロセスとの同等性を証明します。マルチエージェント サイバー物理システム(CPS)の場合、確立された \smart{} 管理自律性ライフサイクルを適用し、実行時の証拠を 4 つのガバナンス状態(\Stable{}/\Meta{}/\Assisted{}/\Regulated{})にマッピングします。コンセンサス ゲーティング、群レベルのリアプノフ解析、エージェントごとのギア権限、およびランデブー制御により、規定された前提条件の下での衝突ゼロを含む、分散型の安全性と安定性の保証が提供されます。 10,000 回のモンテカルロ エピソードにわたる NIST \emph{ロボット アーム位置精度の劣化測定} データセットから校正された故障規模を使用して、3 エージェントの UR5 ロボット アセンブリ セルでの実行時間を評価します。単一エージェントベースラインの異常検出率 2.1\% に対して 99.6\% を達成し、検出遅延を 3.5 倍 $ 削減し、正式な物理作業スペースの安全証明書を提供します。実行ギアは \smart{} ランタイム ガバナンス状態の下でミクロレベルの権限として機能し、アクション制御を自律ガバナンスから分離します。
原文 (English)
Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems
Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous human oversight: safety violations from unverified actions, behavioral instability from unconstrained loops, and continuity loss from unhandled error states. We develop \system{}, a discrete-time control system that combines five execution gears (\Gobs{}, \Gsug{}, \Gplan{}, \Gexec{}, \Gint{}) with utility-gated dispatch and event-driven fallback. For the single-agent case, we prove monotonic stability, execution safety, eventual stabilization, fallback completeness, and equivalence to a gear-constrained Markov decision process. For multi-agent cyber-physical systems (CPS), we apply the established \smart{} managed-autonomy lifecycle and map runtime evidence into its four governance states (\Stable{}/\Meta{}/\Assisted{}/\Regulated{}). Consensus gating, swarm-level Lyapunov analysis, per-agent gear authority, and rendezvous control provide distributed safety and stability guarantees, including zero collision under the stated assumptions. We evaluate the resulting runtime on a three-agent UR5 robotic assembly cell using fault magnitudes calibrated from the NIST \emph{Degradation Measurement of Robot Arm Position Accuracy} dataset across 10,000 Monte Carlo episodes. It achieves a 99.6\% anomaly detection rate versus 2.1\% for the single-agent baseline, reduces detection latency by $3.5\times$, and supplies a formal physical-workspace safety certificate. The execution gears act as micro-level permissions beneath the \smart{} runtime governance states, separating action control from autonomy governance.
逆計画としてのパーソナライゼーション: 構造ノイズ除去によるエージェント的スライド生成のための潜在的な設計意図の学習
スライドのデザインでは、デッキのテーマとページ レイアウトの両方をパーソナライズする必要があります。しかし、現在の AI エージェントベースの手法は、きめ細かいページレベルの設計に苦労しています。事前に指定されたテンプレートやユーザーの詳細な指示のみに依存すると、潜在的なデザイン意図を捉えることができず、ページレベルのスライド パーソナライゼーション (PSP) が未解決のままになります。このギャップを埋めるために、この研究では PSP を逆計画問題として定式化します。使用されている特定の実行ツール (PowerPoint、Beamer など) についての知識を前提とせずに、設計意図を学習することを提案します。ただし、これらのツールの制御を放棄すると、問題はエンドツーエンドの最適化が困難になります。これを克服するために、PSP を近似的に解決するための原則的なフレームワークである SPIRE を提案します。 SPIRE は、クリーンなスライドの視覚構造を意図的に破損することで、破損のノイズを除去するための検証可能なタスクを作成します。これにより、2 人のエージェントが、強化学習 (RL) を通じて実行可能なデザインを協力して改良する方法を学習します。我々は、構造的ノイズ除去がPSPの一貫した代用であること、およびマルチエージェント定式化がRLにおけるポリシー勾配の分散を厳密に低減することの証明を提示する。広範な実験により、SPIRE の優位性が実証されました。
原文 (English)
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page-level Slide Personalization (PSP) unresolved. To close this gap, this work formulates PSP as an inverse planning problem. We propose to learn a design intent without assuming any knowledge of the specific executing tools (e.g., PowerPoint, Beamer) being used. However, relinquishing control over these tools makes the problem intractable to optimize end-to-end. To overcome this, we propose SPIRE, a principled framework to solve PSP approximately. By intentionally corrupting the visual structures of clean slides, SPIRE creates a verifiable task to denoise the corruption, whereby two agents learn to collaboratively refine executable designs via reinforcement learning (RL). We present a proof that structural denoising is a consistent surrogate for PSP, and that the multi-agent formulation strictly reduces policy gradient variance in RL. Extensive experiments demonstrate the superiority of SPIRE.
PHREEQC-MCQ-200: ツール拡張科学シミュレーター エージェントの診断ベンチマーク
大規模な言語モデル エージェントは科学ソフトウェアとの連携がますます高まっていますが、ツールへのアクセスによって科学計算が単に複雑になるだけでなく信頼性が高まるのはいつかはまだ不明です。決定論的な水地球化学シミュレーションでツール拡張エージェントを評価するためのベンチマークである PHREEQC-MCQ-200 を紹介します。このベンチマークには、21 の検証済み PHREEQC シナリオから派生した 200 の多肢選択式の質問が含まれており、エージェントはシミュレーターの入力を構築し、PHREEQC を実行し、構造化された出力を検査し、最終的な回答にコミットする必要があります。複数のフロンティアおよび中間層のモデル ファミリにわたって、シミュレーターへのアクセスにより集計精度が大幅に向上し、多くの科学技術計算タスクには根拠のある実行が必要であることが確認されました。ただし、その増加は単調ではありません。ツールで強化されたエージェントは、ツールなしで正しく回答した項目も失い、平均精度だけでは隠れていた回帰が明らかになります。さらに、出力アクセス プロトコルが重要であることを示します。目次インターフェイスは、より強力なモデルの精度を維持または向上させながらトークン コストを削減できますが、構造化されたシミュレーター出力を確実にナビゲートできない中間層モデルのパフォーマンスは低下します。したがって、PHREEQC-MCQ-200 は、単純なツール呼び出し機能ではなく、科学ツールの使用をエンドツーエンドの診断問題として構成します。私たちは、科学エージェントの評価では、精度だけでなく、項目レベルの保持、出力アクセスの感度、軌道の失敗、および計算チェーンが中断された場所も報告する必要があると主張します。
原文 (English)
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a benchmark for evaluating tool-augmented agents on deterministic aqueous-geochemistry simulations. The benchmark contains 200 multiple-choice questions derived from 21 validated PHREEQC scenarios, requiring agents to construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to final answers. Across multiple frontier and mid-tier model families, simulator access substantially improves aggregate accuracy, confirming that grounded execution is necessary for many scientific-computation tasks. However, the gains are not monotonic: tool-augmented agents also lose items they answered correctly without tools, revealing regressions that average accuracy alone hides. We further show that output-access protocol matters. A table-of-contents interface can reduce token cost while preserving or improving accuracy for stronger models, but it degrades performance for mid-tier models that cannot reliably navigate structured simulator outputs. PHREEQC-MCQ-200 therefore frames scientific tool use as an end-to-end diagnostic problem rather than a simple tool-calling capability. We argue that evaluations of scientific agents should report not only accuracy, but also item-level retention, output-access sensitivity, trajectory failures, and where the computation chain breaks.
Agri-SAGE: コンテキストを認識した農業アドバイザリー生成のためのシミュレーションに基づいたマルチエージェント LLM
農業勧告システムは根本的な緊張に直面しています。静的な農学ガイドラインは、一貫した証拠に基づいた推奨事項を提供しますが、季節ごとの変動性や動的な不確実性には盲目なままです。 LLM を利用した最近の勧告システムは、農学的には信頼できるが生理学的には説得力のない推奨事項を生成するという別のリスクを負います。 Agri-SAGE は、検索に基づいたマルチエージェント LLM 推論と APSIM ベースの生物物理シミュレーションを統合し、農業に関する勧告を生成および検証することで、上記 2 つの制限を解決するように設計された閉ループ フレームワークです。このフレームワークを評価するために、10 年間の遡及分析にわたって、計画と解決、思考のツリー、および反省という 3 つの推論アプローチを評価します。 3 つすべてが静的な PoP (実践パッケージ) ベースラインを大幅に上回り、Tree of Thoughts は印象的なピーク収量を達成しました。同時に、Reflexion は季節を超えたエピソード記憶を活用することで、大幅に低い計算コストで同等の農業成果を達成します。
原文 (English)
Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation
Agricultural advisory systems face a fundamental tension: static agronomic guidelines offer consistent, evidence-based recommendations, yet remain blind to in-season variability and dynamic uncertainties. Recent advisory systems powered by LLMs are liable for a different risk of generating recommendations that are agronomically credible but physiologically unconvincing. Agri-SAGE is a closed-loop framework designed to resolve the above two limitations by integrating retrieval-grounded multi-agent LLM reasoning with APSIM-based biophysical simulation, to generate and validate agronomic advisories. To assess this framework, we evaluate three reasoning approaches, namely Plan-and-Solve, Tree of Thoughts, and Reflexion, over a 10-year retrospective analysis. All three significantly outperform static PoP (Package-of-Practice) baselines, with Tree of Thoughts achieving impressive peak yields. At the same time, Reflexion achieves comparable agronomic outcomes at substantially lower computational cost by leveraging cross-seasonal episodic memory.
進化する環境における身体化エージェントの世界モデルのマルチスケール混合
現実世界で活動する身体化されたエージェントは、状況の変化に応じてマルチスケールの推論と知識の適応を必要とします。この設定に専門家混合 (MoE) を適用する際の 2 つの課題を特定します。1 つは、ルーティングに明確な規模の概念が欠けており、特定の規模での対象を絞った更新が妨げられること、そして、統一された更新ポリシーでは、各規模での知識が時代遅れになるさまざまな速度に対応できないことです。私たちは、スケールを意識した世界モデルの混合と進化を通じて両方の課題に対処するフレームワーク、MuSix を紹介します。 2 段階のルーティング メカニズムは、解釈レベル理論に触発された状況の新規性の尺度である経験的距離に基づくスケール選択を根拠とします。まず、メタルーターがこの量を連続スケール空間上の重みにマッピングし、次にスケールごとのベース ルーターが特定されたスケール内のワールド モデルを選択します。適応に関しては、スケール依存の忘却率により、高スケールの抽象化が持続しながら低スケールの知識が迅速に更新され、ゲートされたスケール間転送により階層全体の一貫性が維持されます。 EmbodiedBench と HAZARD での実験では、MuSix がマルチスケール推論と動的適応に関して最先端のベースラインよりも改善していることが示されています。
原文 (English)
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Embodied agents operating in the real world require multi-scale reasoning and knowledge adaptation as conditions change. We identify two challenges in applying Mixture of Experts (MoE) to this setting: routing lacks an explicit notion of scale, preventing targeted updates at specific scales, and a uniform update policy cannot accommodate the different rates at which knowledge at each scale becomes outdated. We present MuSix, a framework that addresses both challenges through scale-aware world model mixture and evolution. A two-stage routing mechanism grounds scale selection in experiential distance, a measure of situational novelty inspired by Construal Level Theory: a meta-router first maps this quantity to a weight over continuous scale space, then per-scale base routers select world models within the identified scale. For adaptation, scale-dependent forgetting rates allow low-scale knowledge to refresh rapidly while high-scale abstractions persist, and gated inter-scale transfer maintains coherence across the hierarchy. Experiments on EmbodiedBench and HAZARD show that MuSix improves over state-of-the-art baselines on multi-scale reasoning and dynamic adaptation.
AI ネイティブ ゲーム: 調査とロードマップ
生成 AI により、ゲームが実行時に対話、クエスト、キャラクター、画像、世界を生成できるようになりました。ただし、生成だけではゲームが AI ネイティブになるわけではなく、プレイアビリティが保証されるわけでもありません。この論文では、ランタイム生成 AI がコア ループを構成するかどうかによって AI ネイティブ ゲームを定義します。AI コンポーネントが削除されたり、些細な置き換えが行われたりすると、プレイの中心的な形式が崩壊するか、根本的に異なるものになります。この反事実基準は、AI ネイティブ ゲームを、AI 拡張ゲーム、境界アーティファクト、チャットボット、居酒屋スタイルのロールプレイ、手続き型コンテンツ生成、および AI 支援制作から区別します。この定義を使用して、候補となる成果物をスクリーニングし、公開されている 53 の AI ネイティブ ゲームとプロトタイプを分析します。二重軸の G/N 分類法を導入します。G 軸はプレイヤーが直面するゲームのタイプを捉え、N 軸は生成 AI をプレイに不可欠なものにする支配的な AI メカニズムを捉えます。このコーパスは言語優先デザイン、特に物語的冒険、認識論的相互作用、生成的物語を中心に集中していますが、意味論的判断、マルチエージェント シミュレーション、生成的構築、関係/仲間遊びなどのカテゴリーはあまり代表されていません。私たちは、設計上の中心的な問題は、セマンティックなオープン性を安定したゲームプレイに組み込むことであると主張します。 AI ネイティブの設計は、目標、ルール、状態、フィードバック、ペーシング、プレーヤーの主体性などの機械的不変条件に依存しており、これらにより、オープンエンドの AI 出力が解釈可能で結果的なものになります。最後に、制御可能な発電、機械としての AI 設計、マルチモーダルおよびマルチエージェント システム、推論の経済学、評価、安全性、規制のロードマップを示します。
原文 (English)
AI Native Games: A Survey and Roadmap
Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime. Yet generation alone does not make a game AI-native, nor does it guarantee playability. This paper defines AI-native games by whether runtime generative AI is constitutive of the core loop: if the AI component were removed or trivially replaced, the central form of play would collapse or become fundamentally different. This counterfactual criterion separates AI-native games from AI-augmented games, boundary artifacts, chatbots, tavern-style role-play, procedural content generation, and AI-assisted production. Using this definition, we screen candidate artifacts and analyze 53 publicly available AI-native games and prototypes. We introduce a dual-axis G/N taxonomy: the G-axis captures player-facing game type, while the N-axis captures the dominant AI mechanic that makes generative AI indispensable to play. The corpus is concentrated around language-forward designs, especially narrative adventure, epistemic interaction, and generative narrative, while categories such as semantic adjudication, multi-agent simulation, generative construction, and relationship/companion play remain less represented. We argue that the central design problem is organizing semantic openness into stable gameplay. AI-native design depends on mechanical invariants: goals, rules, state, feedback, pacing, and player agency that make open-ended AI outputs interpretable and consequential. We conclude with a roadmap for controllable generation, AI-as-mechanic design, multimodal and multi-agent systems, inference economics, evaluation, safety, and regulation.
HARC: 堅牢な安全調整のための有害性と拒否のカップリングの方向性
アライメントされた LLM が内部的にどのように安全性を表すかを理解することは、ジェイルブレイクが成功する理由を説明し、堅牢なアライメント戦略の設計に情報を提供するため、アライメントの脆弱性を診断するために重要です。これまでの研究では、整列された LLM がプロンプト側のトークン位置で残留ストリーム内の分離可能な方向として有害性と拒否をエンコードしていることが示されています。トークンが生成される前に拒否または有害性の方向を抑制することで、ジェイルブレイクがプロンプトエンコーディングで成功し、異なる攻撃クラスが有害性と拒否の面の分離可能な領域を占めることを示します。分析をレスポンス トークンの位置まで拡張すると、プロンプト側で入力を有害なものとして認識できなかった場合でも、モデルが有害なコンテンツを生成中にそのコンテンツを認識することがわかりました。私たちの発見に動機づけられて、私たちは、プロンプトポジションとレスポンスポジションの両方で2つの方向をペアにする微調整方法であるHARC(有害性と拒否のカップリング)を紹介します。介入は有害性拒否部分空間に限定されるため、残りのストリームの残りの部分はそのまま残り、一般的な能力を低下させたり、過剰な拒否を拡大したりすることはありません。広範な実験を通じて、HARC は、主要なトレーニング時間と推論時間の安全性手法にわたる 6 つのベースラインの中で最も強力な堅牢性、機能、使用性のトレードオフを達成しました。プロンプトおよびレスポンスの位置における有害性と拒否の指示は、アーキテクチャ固有の調整を行わずにテストした 5 つのモデル ファミリと 2 つのスケールに渡って伝達されます。
原文 (English)
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.
ワールド モデリング エージェントのベンチマーク フレームワークとしての AGI Maze
大規模言語モデル (LLM) は強力なパターン補完システムですが、そのデフォルトの動作モード (静的コンテキストから次のトークンを予測する) では、外部世界の永続的で操作可能な表現が確実に生成されません。環境が部分的に観察可能でステートフルになり、隠れた状態についての記憶と構造化された仮説が必要になると、テキストでは「推論」のように見える多くのタスクが大幅に難しくなります。 AGI Maze は、高次元の感覚入力を必要とせずにそのような環境を構築するための軽量フレームワークです。これは、クリーンな API と複数の難易度レジームを備えた一連のグリッドベースの迷路タスクを提供します。目標は、エージェントがすぐに提供される観測値からローカル ルールを推測するだけでなく、世界状態の表現を学習して使用する必要があるベンチマークを作成することです。単純な迷路に関するいくつかのバニラ LLM の初期評価を提供し、LLM 推論時に内部的に迷路を表現できないことを示しています。また、ベースライン エージェントも導入します。これにより、メッセージ履歴を作業メモリとして使用して、エージェントの実行時に観察の記述を構築できます。これによりパフォーマンスは向上しますが、LLM エージェントが人間にとって十分なステップ バジェット内で小さな迷路であっても確実に解決するにはまだ不十分です。
原文 (English)
AGI Maze as a Benchmark Framework for World-Modeling Agents
Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an external world. Many tasks that look like "reasoning" in text become substantially harder once the environment is partially observable, stateful, and requires memory and structured hypotheses about hidden state. AGI Maze is a lightweight framework for building such environments without requiring high-dimensional sensory inputs. It provides a family of grid-based maze tasks with a clean API and multiple difficulty regimes. The goal is to create benchmarks where agents must learn and use world state representations, not just infer a local rule over readily provided observations. We provide an initial evaluation of several vanilla LLMs on simple mazes showing that they fail to represent mazes internally at LLM inference time. We also introduce a baseline agent, which is allowed to use its message history as a working memory to construct descriptions of observations at agentic runtime. Although this can improve performance, it is still insufficient for an LLM agent to reliably solve even small mazes within a step budget that is more than enough for humans.
インタラクティブなゲームプレイのためのコーチング可能なエージェント
強化学習は、高度な AI およびロボット システムの作成における貴重なツールであることが証明されており、ゲームプレイからロボット工学、基礎モデルに至るまであらゆるものに貢献しています。通常、これらの AI システムは、試行錯誤を通じて、タスクを解決するために最適に近い 1 つの動作を学習します。ただし、タスクの解決方法に関して、できればリアルタイムで、ある程度の制御を主張したいユースケースは数多くあります。コアタスクのこれらの変更をスタイルと呼びます。私たちは、ユニバーサル価値関数近似器 (UVFA) を、慎重に選択されたトレーニング シナリオ、学習アルゴリズム、データ拡張と組み合わせて、複雑な領域でスタイルを示すエージェントをコーチングするためのフレームワークを作成します。私たちは、AAA ビデオ ゲームの Horizon Forbidden West と Gran Turismo、およびオープンソースのヒューマノイド テスト ドメインでのフレームワークのアプリケーションを実証します。カーレース、様式化されたゲーム戦闘、人型歩行など、ドメインの性質が異なるにもかかわらず、各エージェントは、そのドメインの主なタスクを満たしながら、スタイルの要求に強い一貫性を示します。重要なのは、このホワイト ペーパーで概説した手法を使用すると、エンド ユーザーが実行時に最終的な動作を選択できるため、最終的に実行されるパフォーマンスを柔軟に制御できることです。
原文 (English)
Coachable agents for interactive gameplay
Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework's application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains -- car racing, stylized game combat, and humanoid walking -- each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.
Self-GC: Long-Horizon LLM エージェントの自己統治コンテキスト
長期的な LLM エージェントは、使い捨てのテキスト サフィックスとして扱うには構造化されすぎているツールの結果、ファイル、プラン、およびユーザー制約を蓄積します。現在のシステムは主に、時系列プルーニングやツール出力マスキングなどの実行中のヒューリスティック、またはコンテキスト制限に近い最終的な自己要約に依存しています。ヒューリスティックは低コストですが、将来の依存関係を考慮しません。概要は物語の状態を保持しますが、多くの場合、正確な証拠、ロケーター、編集可能な成果物が隠されます。私たちは Self-GC を紹介します。GC はガベージ コレクションを意図的にエコーしながら自己統治コンテキストを表します。システムは単に未使用のトークンを再利用するだけでなく、エージェント コンテキスト オブジェクトのライフサイクルを管理します。自己 GC は、ユーザーのターン、ツール スパン、およびスキルの状態をインデックス付きオブジェクトに変換します。サイドチャネルプランナーにフォールド、マスク、プルーニングのアクションを提案するよう依頼します。また、ハーネスは回復可能なサイドカー、安全なコミット境界、キャッシュを認識したコミットを強制できます。 33 セッションのハード セットでは、セルフ GC はプレフィックス トークンの 43.95% をプルーニングし、将来の継続の 84.85% には影響を与えません。これに対し、ヒューリスティック ベースラインの影響なし率は 54.55% ~ 69.70% です。 332 セッションの本番派生スイートでは、3 つのプランナー バックボーンは 91.27% ~ 94.58% の影響なし率に達しますが、ベースラインは 77.71% ~ 87.46% のままです。運用環境では、オンラインのアカウントレベルの分割により、日中の平均入力トークンが 10% ~ 15% 削減され、ピーク削減率は 20% 近くになります。これらの結果は、事後のテキスト クリーンアップではなく、インデックス付きの回復可能なオブジェクトに対する実行時のライフサイクル制御としてのコンテキスト管理を示しています。
原文 (English)
Self-GC: Self-Governing Context for Long-Horizon LLM Agents
Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. Current systems mostly rely on in-run heuristics such as chronological pruning and tool-output masking, or on final self-summary near a context limit. Heuristics are cheap but blind to future dependencies; summaries preserve narrative state but often hide exact evidence, locators, and editable artifacts. We present Self-GC, where GC denotes self-governing context while deliberately echoing garbage collection: the system does not merely reclaim unused tokens, but governs the lifecycle of agent context objects. Self-GC turns user turns, tool spans, and skill state into indexed objects; asks a side-channel planner to propose fold, mask, and prune actions; and lets the harness enforce recoverable sidecars, safe commit boundaries, and cache-aware commit. On a 33-session Hard Set, Self-GC prunes 43.95% of prefix tokens while leaving 84.85% of future continuations unaffected, compared with no-impact rates of 54.55% to 69.70% for heuristic baselines. On a 332-session production-derived suite, three planner backbones reach no-impact rates of 91.27% to 94.58%, while baselines remain at 77.71% to 87.46%. In production, an online account-level split reduces daytime average input tokens by 10% to 15%, with peak reductions near 20%. These results point to context management as runtime lifecycle control over indexed, recoverable objects rather than post hoc text cleanup.
いつでも有効な証明書を備えた自己進化エージェント
自己進化するエージェントは、ほとんどの学習理論的保証の背後にある仮定に違反します。つまり、データ、評価器、コンポーネント、仮説空間は、更新されるポリシーによって生成されます。私たちは \textbf{SEA} を提案します。このアーキテクチャは、自己変更を小さなステアリング アダプターと \emph{frozen} ベース モデルの周りのバージョン管理されたハーネスに限定し、固定エラー バジェットに対して監査可能な証明書を発行する常時有効なゲートを介してのみ各変更を許可します。 5 つのループ コントローラーが公開保証を構成します。そのようなゲートは、フリーズされたベースがすでに生成している動作の中から \emph{選択} することしかできないため、5 つの検証者インザループ メカニズム (best-of-$N$、マイクロステップ検索、自己作成複製オラクル、検索層制御、自己修復) が、問題テキストのみから計算された、ゲートが必要とする密度の高いグレーダーフリーの信号を供給します。 4 つの基本モデルにわたる $52$ インスタンスの SWE ベンチ検証済みサブセットでは、基本機能が支配的で交絡のない効果であり、2 つの強力な基本モデルでは、意図的な no-op-composite 制御により $+4$ と $+5$ (\textsc{Glm}~5.2 $24\to28$; \textsc{Gpt} $29\to34$、$65\%$) でスイートの寄与が分離されます。最良)、そのメカニズムが起動して回帰を防止していることを確認するイベント ログが含まれます。結果は高価な評価では 1 回実行されます。実行ごとの差異の確認とタスクごとのアルゴリズムの組み合わせの適応は今後の課題です。
原文 (English)
Self-Evolving Agents with Anytime-Valid Certificates
Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis space are produced by the policy being updated. We present \textbf{SEA}, an architecture that confines self-modification to a small steering adapter and a versioned harness around a \emph{frozen} base model and admits each modification only through an anytime-valid gate that emits an auditable certificate against a fixed error budget. Five loop controllers compose published guarantees; because such gates can only \emph{select} among behaviors the frozen base already produces, five verifier-in-the-loop mechanisms -- best-of-$N$, micro-step search, self-authored reproduction oracles, search-layer control, and self-repair -- supply the dense, grader-free signal the gates require, computed from the issue text alone. On a $52$-instance SWE-bench Verified subset across four base models, base capability is the dominant, confound-free effect, and on two strong base models a deliberate no-op-composite control isolates the suite's contribution at $+4$ and $+5$ (\textsc{Glm}~5.2 $24\to28$; \textsc{Gpt} $29\to34$, the $65\%$ best), with event logs confirming that its mechanisms fire and prevent regressions. Results are single-run on expensive evaluations; confirming run-to-run variance and adapting the per-task algorithm mix are future work.
2 つの AI 指標の相違: それはすべての違いを生むのでしょうか?
コンピューティングの指数関数的なスケーリングが続くにつれて、フロンティア AI モデルの機能は、開発者が少ない固定予算でアクセスできる機能を超えるでしょうか?それとも機能は「地球を継承する柔和なモデル」に収束するのでしょうか? Gundlach らを基に構築(2025b) では、答えは AI の能力をどのように評価し測定するかによって決まることを示しています。従来のパフォーマンス尺度について議論し、検証損失ではギャップが縮小している一方、他の指標ではフロンティア モデルのリードが永遠に拡大していることを示します。パフォーマンス メトリクスをトレーニング (および推論) コンピューティングに関連した関数形式で分類することで、どのメトリクスが柔和なモデルを好むかを決定するための厳密な数学的条件を提供し、制限されたパフォーマンス メトリクスが常にそうなることを示します。ただし、パフォーマンス メトリクスを慎重に解釈することが不可欠です。多くの一般的な制限付きメトリクスには、密接に関連する制限のない対応するメトリクスがある (またその逆も同様) ことを示します。ドメイン内の適切なメトリクスを決定することは、ポリシーの前提条件です。これは、制限されたメトリクスと制限されていないメトリクスが、反対のポリシー応答を示唆する可能性があるためです。ソフトウェアエンジニアリング、合成生物学、修辞的説得力などの特定の能力が、私たちが関心を持っている用語で測定したときに制限がない場合、フロンティアレベルの能力は少数の裕福な関係者の手に集中する可能性があります。逆に、その能力が制限されている場合、フロンティアレベルの能力は柔和なモデルを通じて多くの人の手に渡ります。
原文 (English)
Two AI Metrics Diverged: Will it Make All the Difference?
As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget? Or will capabilities converge, with "meek models inheriting the earth"? Building on Gundlach et al. (2025b), we show that the answer depends on how we value and measure AI capabilities. We discuss conventional performance measures and show that, while validation loss shows a shrinking gap, on other metrics frontier models grow their lead forever. Classifying performance metrics by their functional forms in relation to training (and inference) compute, we provide tight mathematical conditions for determining which metrics favor meek models, and show that bounded performance metrics always do. But careful interpretation of performance metrics is essential: we show that many common bounded metrics have closely-related counterpart metrics that are unbounded (and vice versa). Determining the apt metric in a domain is a prerequisite for policy, since bounded and unbounded metrics may suggest opposing policy responses. If a particular capability -- like software engineering, synthetic biology, or rhetorical persuasiveness -- is unbounded when measured in the terms we care about, frontier-level capability will likely be concentrated in the hands of a few wealthy actors. Conversely, if that capability is instead bounded, frontier-level capabilities proliferate through meek models into the hands of the many.
グラフネイティブの強化学習により、概念の組み換えによる追跡可能な科学的仮説の生成が可能になります
材料発見を加速するには、複数ステップのドメインに基づいた推論を通じて科学的に有効な仮説を生成できる AI システムが必要です。標準的な大規模言語モデルは、多くの場合、オープンエンドの材料設計問題に対して流暢ではあるものの追跡可能性が低い応答を生成するため、最終的な答えが一貫した中間推論によってサポートされているかどうかを判断することが困難になります。私たちは、グループ相対ポリシー最適化 (GRPO) で微調整されたグラフネイティブ推論モデルのファミリーである Graph-PRefLexOR を開発し、推論をメカニズム探索、グラフ構築、パターン抽出、仮説合成の明示的なフェーズに編成します。この設計は、ニューラル言語の生成を記号関係構造とリンクさせ、因果関係の構築、検査、再利用を可能にします。材料科学および力学の文献からの自由回答形式の質問 100 件について、Graph-PRefLexOR は、対応する基本モデルと比較して 40 ~ 65% の改善を達成し、推論のトレーサビリティが最大に向上しました。埋め込み分析では、ベースラインよりも広範な意味の探索と約 2 ~ 3 倍の意味の多様性が示されます。セマンティックバックトラッキングとレイヤーごとの隠れ状態分析により、構造化された推論と最終的な答えの間のより強力な整合性がさらに示されます。最後に、テスト時のグラフ拡張により、追加のコンピューティングによって、単に意味論的範囲が拡大されるのではなく、主に、制限された意味論的空間内での長距離概念の組み換えが増加することが明らかになりました。これらの結果は、材料設計やその他の科学的応用における科学的仮説生成のための解釈可能な AI システムへの道筋として、グラフネイティブ強化学習を確立します。
原文 (English)
Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making it difficult to determine whether final answers are supported by coherent intermediate reasoning. We develop Graph-PRefLexOR, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization (GRPO) to organize reasoning into explicit phases for mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis. This design links neural language generation with symbolic relational structure, enabling causal connections to be constructed, inspected, and reused. On 100 open-ended questions from materials science and mechanics literature, Graph-PRefLexOR achieves 40-65% improvements over corresponding base models, with the largest gains in reasoning traceability. Embedding analyses show broader semantic exploration and approximately 2-3 times greater semantic diversity than baselines. Semantic backtracking and layer-wise hidden-state analyses further show stronger alignment between structured reasoning and final answers. Finally, test-time graph expansion reveals that additional compute primarily increases long-range conceptual recombination within a bounded semantic space, rather than simply expanding semantic coverage. These results establish graph-native reinforcement learning as a pathway toward interpretable AI systems for scientific hypothesis generation in materials design and other scientific applications.
エージェントティック RAG パイプラインのベイジアン不確実性伝播: マルチホップ質問応答に関する概念実証研究
Agentic Retrieval-Augmented Generation(RAG)システムの信頼できる展開には、多段階の推論パイプラインがいつ失敗するかを推定するメカニズムが必要です。この論文では、プランナー、評価器、およびジェネレーターの各ステージが、セマンティックな発散とジェネレーターの自己評価から得られる不確実性信号を生成する、不確実性を認識したエージェント検索拡張生成 (RAG) フレームワークを紹介します。これらの信号はベイジアン ネットワーク (BN) を介して伝播され、システム レベルの不確実性が推定され、ワークフロー全体にわたる潜在的な障害点のノード レベルの指標が提供されます。このアプローチは、GPT-3.5-Turbo および GPT-4.1-Nano を使用して、StrategyQA および HotpotQA で評価されます。受信機動作特性曲線下面積 (AUROC)、精度除去曲線下面積 (AUARC)、予想される校正誤差 (ECE)、およびブライアー スコアを使用して、識別、選択的予測、および校正を評価します。結果は、HotpotQA ではベイジアン伝播がより効果的であることを示しています。HotpotQA では、マルチホップ推論段階にわたって不確実性が蓄積されますが、StrategyQA では、誤ったキャリブレーションと信頼性の低いアップストリーム信号によって引き起こされる制限が明らかになります。この研究では、ベイジアン不確実性伝播を、Agentic RAG システムを監視するための有望ではあるが予備的なメカニズムとして位置づけており、洋上風力発電 (OSW) の保守意思決定サポートなどの産業分野では将来の検証が必要です。
原文 (English)
Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering
Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasoning pipelines may fail. This paper presents an uncertainty-aware Agentic Retrieval-Augmented Generation (RAG) framework in which planner, evaluator and generator stages produce uncertainty signals derived from semantic divergence and generator self-evaluation. These signals are propagated through a Bayesian Network (BN) to estimate system-level uncertainty and provide node-level indicators of potential failure points across the workflow. The approach is evaluated on StrategyQA and HotpotQA using GPT-3.5-Turbo and GPT-4.1-Nano, with Area Under the Receiver Operating Characteristic Curve (AUROC), Area Under the Accuracy-Rejection Curve (AUARC), Expected Calibration Error (ECE), and Brier Score used to assess discrimination, selective prediction and calibration. Results show that Bayesian propagation is more effective on HotpotQA, where uncertainty accumulates across multi-hop reasoning stages, while StrategyQA exposes limitations caused by miscalibration and unreliable upstream signals. The study positions Bayesian uncertainty propagation as a promising but preliminary mechanism for monitoring Agentic RAG systems, with future validation required in industrial domains such as Offshore Wind (OSW) maintenance decision support.
PedNStream: 歩行者交通管理のためのスケーラブルなネットワーク フロー シミュレーション
大規模な群衆管理には、計算効率が高く、フィードバック ベースの制御と互換性のある歩行者シミュレーションが必要です。ただし、ほとんどのオープンソース ツールは微細なものであるか、ネットワーク規模の閉ループ評価用に設計されていません。このペーパーでは、リンク伝送モデル (LTM) に基づいた巨視的な歩行者ネットワーク負荷のためのオープンソースの Python ネイティブ シミュレーターである PedNStream (Pedestrian Network Flow Simulation) について説明します。このフレームワークは、拡散と活動に起因する変動を捉える確率的リンクダイナミクスを組み込むことで LTM ベースの歩行者モデルを拡張し、動的なユーザーの平衡ルート選択を、不確実で介入主導の設定に適したユーティリティベースの定式化に置き換えます。 PedNStream は、ゲート、フロー分離、ルート ガイダンスなどの介入用のコントローラー インターフェイスが組み込まれたモジュール式フレームワークとして実装されています。私たちはフレームワークを段階的に評価します。合成シナリオでは、キューの形成、スピルバック、輻輳の解消、適応型再ルーティングなどの主要なメカニズムを検証します。実際のネットワーク実験では、大規模な行動と観察された歩行者数との一貫性を評価します。閉ループのケーススタディではコントローラーの統合を実証し、ランタイム分析ではスケーラビリティを定量化します。これらの結果により、PedNStream は大規模な歩行者ネットワークのシミュレーションと制御のための効率的かつ実用的なテストベッドとして確立されます。
原文 (English)
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
Large-scale crowd management requires pedestrian simulations that are both computationally efficient and compatible with feedback-based control. However, most open-source tools are either microscopic or not designed for network-scale closed-loop evaluation. This paper presents PedNStream (Pedestrian Network Flow Simulation), an open-source, Python-native simulator for macroscopic pedestrian network loading based on the Link Transmission Model (LTM). The framework extends LTM-based pedestrian models by incorporating stochastic link dynamics that capture diffusion and activity-induced variability, and replaces dynamic user equilibrium route choice with a utility-based formulation suited to uncertain, intervention-driven settings. PedNStream is implemented as a modular framework with built-in controller interfaces for interventions such as gating, flow separation, and route guidance. We evaluate the framework in a staged manner. Synthetic scenarios verify key mechanisms, including queue formation, spillback, congestion dissipation, and adaptive rerouting. Real-network experiments assess large-scale behavior and consistency with observed pedestrian counts. A closed-loop case study demonstrates controller integration, and a runtime analysis quantifies scalability. These results establish PedNStream as an efficient and practical testbed for large-scale pedestrian network simulation and control.
決定論的で自己拡張的な反応分類のための検証可能なルールをエージェント的に生成
コンピューター支援合成計画では、各変換に決定論的で解釈可能なラベルを割り当てる反応ルールの大きなライブラリを使用して、標的分子をアクセス可能な前駆体に分割します。しかし、化学はロングテールであるため、手動エンコーディングが困難であり、既存のツールは新しい化学に適応できない固定ルールセットに依存しています。ここでは、大規模言語モデル (LLM) のマルチエージェント フレームワークが反応を分類し、665,901 件の米国特許反応にわたってルール自体を記述し、コーパスに対してテストする検証ループの下で各ルールを生成する、完全に自動化されたパイプラインを紹介します。人間によるキュレーションを行わずに、標準分類を 68 クラスから 14,073 クラスに拡張します。軽量の指紋分類器を使用して、目に見えない反応の 97.7% を分類し、主要な独自の分類器と一致しながら、化学をより細かく分解し、オンデマンドでトレーニング ディストリビューション外の化学に拡張します。その結果、生きた反応性データベースと、生成モデルを信頼性の高い自己拡張型のシンボリック システムに変えるための一般的なルートが得られます。
原文 (English)
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encoding intractable, and existing tools rely on fixed rulesets that cannot adapt to new chemistries. Here we present a fully automated pipeline in which a multi-agent framework of large language models (LLMs) classifies reactions and writes the rules themselves across 665,901 US patent reactions, generating each rule under a verification loop that tests it against the corpus. It expands a standard taxonomy from 68 to 14,073 classes without human curation. With a lightweight fingerprint classifier, it classifies 97.7\% of unseen reactions, matching a leading proprietary classifier while resolving chemistry more finely and extending on demand to chemistry outside its training distribution. The result is a living reactivity database and a general route to turning generative models into reliable, self-expanding symbolic systems.
エージェントはオープンワールドに一般化できますか?ツール使用における静的トレーニングの脆弱性を明らかにする
Large Language Model (LLM) エージェントは静的ベンチマークでは熟練していることを示しますが、実際のシナリオでの展開は、ユーザー クエリ、ツール セット、対話ダイナミクスの動的な性質によって妨げられます。この一般化のギャップに対処するために、クエリ、アクション、観察、およびドメインの次元にわたる分布の変化を特徴とする問題設定である OpenAgent (オープンワールドのツール使用エージェント) を形式化します。その影響を体系的に診断するために、制御されたサンドボックス環境を構築し、知覚、相互作用、推論、内部化の 4 層階層にわたるきめ細かい環境変化を定義し、包括的な一連の実験を実施します。私たちの分析により、一連の重要な洞察が得られ、教師あり微調整 (SFT) と強化学習の両方で訓練されたエージェントは、オープンな環境の変化に直面したときに、さまざまな程度のパフォーマンス低下に悩まされることが実証されました。これらの洞察に基づいて、我々は、現実の環境におけるエージェントの堅牢性と有用性を高めるための基礎を築く、SFT の外乱ベースの介入戦略である摂動拡張微調整を提案します。コードは https://github でリリースされます。 com/LAMDA-NeSy/OpenAgent。
原文 (English)
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize OpenAgent (Tool-Use Agent in Open-World), a problem setting characterized by distributional shifts across query, action, observation, and domain dimensions. To systematically diagnose its impact, we construct a controlled sandbox environment where we define fine-grained environmental shifts across a four-tier hierarchy, Perception, Interaction, Reasoning, and Internalization, and conduct a comprehensive series of experiments. Our analysis yields a series of key insights, demonstrating that agents trained via both Supervised Fine-Tuning(SFT) and Reinforcement Learning suffer from varying degrees of performance degradation when confronting open environmental shifts. Building on these insights, we propose Perturbation-Augmented Fine-Tuning, a disturbance-based intervention strategy for SFT that lays the foundation for enhancing agent robustness and utility in realistic environments. Our code will be released at: https://github. com/LAMDA-NeSy/OpenAgent.
自律型ラボオーケストレーターの最適なリソース利用
自律型研究所では、AI エージェントが次に行うべき実験のバッチを提案します。ただし、利用可能なリソースを最大限に活用してこれらのタスクを計画および実行することは、まったく別の問題です。これは、実際のハードウェアの制約に対処する場合、特に容量やスループットが異なる複数の機器がある場合には困難になる可能性があります。ここでは、金属有機フレームワーク合成用の自律プラットフォームのリソース利用に対処するための 2 段階の方法を示します。まず、制約プログラミングを使用して最適なスケジュールを見つけます。これにより、ハードウェアの制限と容量を満たしながら、合計時間を最小限に抑えるスケジュールが見つかります。次に、各タスクのステータス依存関係のシステムを使用し、最適なスケジュールを確実に実行できるようにします。
原文 (English)
Optimal Resource Utilization for Autonomous Laboratory Orchestrators
In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing those tasks taking full advantage of the available resources is a completely different question. This can be challenging when dealing with real-world hardware constraints, especially so when there are multiple instruments with different capacities and throughputs. Here we demonstrate a 2-step method to address resource utilization for our autonomous platform for metal-organic framework synthesis. First, we use constraint programming to find optimal schedules. This finds schedules that minimizes the total time while still satisfying the limitations and capacities of the hardware. Secondly, we use a system of status dependencies for each task, which allows for the robust execution of the optimal schedules.
Theoria: 非公式推論状態に対する書き換え許容性の検証
AI システムの答えを信頼できるのはどのような場合ですか?形式的証明アシスタントは確実性を提供しますが、問題分布のほとんどには到達できません。スカラー LLM ジャッジはカバレッジを提供しますが、事後的に監査できない不透明なスコアを生成し、他の LLM と同じ一貫性の問題にさらされます。私たちは、このギャップを埋める検証アーキテクチャである Theoria を紹介します。候補解は、型指定された状態遷移のシーケンスに書き換えられます。各状態遷移は、引用、計算、または問題によって与えられた事実など、明示的な正当化によってライセンスされ、すべての遷移は独立して監査可能です。基本的な不変条件は変化の完全性です。連続する証明状態間のすべての違いを考慮する必要があるため、隠れた前提は黙って通過するのではなく、許可されていない突然変異として表面化します。 HLE-Verified Gold (185 のテキストのみのエキスパートの問題) では、Theoria は 91.4% の厳密な精度で 105 を認定しています (Wilson 95% CI [84.5%、95.4%])。すべての認証では、人間が判読できる証明トレースが生成され、各ステップに個別にチャレンジできます。ホリスティック LLM ジャッジは、一致するカバレッジでは同等の精度を達成しますが、別の問題 (Jaccard 0.14 ~ 0.36) では失敗するため、アプローチは補完的になります。 15 のドメインにわたる 95 件の敵対的毒物証明について、構造化された裁判官は 94.7% を捕捉したのに対し、総合的な判断では 83.2% を捕捉しました (p= 0.0017)。全体の 11.5 pp のギャップは、隠れた前提 (90.6% 対 62.5%、28 pp の差) と捏造された引用 (100% 対 90%) に集中しており、形式的な分析が利点を予測するエラー クラスです。利点が予測されない算術および定理の誤用エラーのパフォーマンスは同じです。 GPQA ダイヤモンド (n= 65) では、認定精度は 97.1% (Wilson CI [85.1%、99.5%]) です。
原文 (English)
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each licensed by an explicit justification, whether that be a citation, computation, or problem-given fact, and every transition is independently auditable. The foundational invariant is completeness of change: every difference between consecutive proof states must be accounted for, so hidden premises surface as unlicensed mutations rather than passing silently. On HLE-Verified Gold (185 text-only expert problems), Theoria certifies 105 at 91.4% strict precision (Wilson 95% CI [84.5%, 95.4%]). Every certification produces a human readable proof trace in which each step can be independently challenged. Holistic LLM judges achieve comparable precision at matched coverage but fail on different problems (Jaccard 0.14-0.36), making the approaches complementary. On 95 adversarial poisoned proofs across 15 domains, structured judges catch 94.7% versus 83.2% for holistic judging (p= 0.0017). The overall 11.5 pp gap concentrates in hidden premises (90.6% vs. 62.5%, a 28 pp difference) and fabricated citations (100% vs. 90%), the error classes where the formal analysis predicts an advantage; performance is identical on arithmetic and theorem-misapplication errors, where no advantage is predicted. On GPQA Diamond (n= 65), certified precision is 97.1% (Wilson CI [85.1%, 99.5%]).
AutoMem: 認知スキルとしての記憶の自動学習
記憶の専門知識は学習されたスキルです。何をエンコードするか、いつ取得するか、知識をどのように整理するかを知っています。この能力は認知科学でメタ記憶として知られています。私たちはメモリ管理を訓練可能なスキルとして扱うことで、この視点を LLM にもたらします。ファイル システム操作をタスク アクションと並行してファーストクラスのメモリ アクションに昇格させ、メモリの管理方法をモデル自体に決定させます。この記憶スキルは、それをサポートする構造 (プロンプト、ファイル スキーマ、アクション語彙)、およびそれを実行するモデルの熟練度という 2 つの軸に沿って向上します。どちらの軸も手動による最適化に抵抗します。長期的なタスクのエピソードは何千ステップにもわたって実行され、単一の記憶ミスが表面化するずっと前に隠れてしまう可能性があるため、完全な軌跡を人間がレビューすることは非現実的です。両方の軸を自動化するフレームワークである AutoMem を紹介します。最初のループでは、強力な LLM がエージェントの完全な軌跡をレビューし、エージェントがメモリ ファイルと対話する方法を形成するメモリ構造を繰り返し修正します。 2 番目のループでは、エージェント自身の記憶力の良い判断が多くのエピソードから特定され、モデルの記憶能力を直接高めるためのトレーニング信号として使用されます。プロシージャル生成された 3 つのロングホライズン ゲーム (Craafter、MiniHack、NetHack) にわたって、モデルのタスク アクション動作を変更せずにメモリのみを最適化することで、ベース エージェントのパフォーマンスが最大 2 倍から 4 倍向上し、32B のオープンウェイト モデルが Claude Opus 4.5 や Gemini 3.1 Pro Thinking などのフロンティア システムと競合できるようになりました。私たちの結果は、メモリ管理は独立して学習可能なスキルであり、長期的なタスクで大きな利益をもたらす高いレバレッジの目標であることを示しています。
原文 (English)
AutoMem: Automated Learning of Memory as a Cognitive Skill
Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive science as metamemory. We bring this perspective to LLMs by treating memory management as a trainable skill. We promote file-system operations to first-class memory actions alongside task actions, letting the model itself decide how to manage its memory. This memory skill improves along two axes: the structure that supports it (prompts, file schemas, action vocabulary), and the proficiency of the model exercising it. Both axes resist manual optimization: episodes in long-horizon tasks run for thousands of steps, and a single memory mistake can hide long before it surfaces, making human review of full trajectories impractical. We introduce AutoMem, a framework that automates both axes. In the first loop, a strong LLM reviews complete agent trajectories and iteratively revises the memory structure that shapes how the agent interacts with its memory files. In the second loop, the agent's own good memory decisions are identified from many episodes and used as training signal to sharpen the model's memory proficiency directly. Across three procedurally generated long-horizon games (Crafter, MiniHack, and NetHack), optimizing memory alone--without modifying the model's task-action behavior--improved the base agent's performance ~2x-4x, bringing a 32B open-weight model competitive with frontier systems such as Claude Opus 4.5 and Gemini 3.1 Pro Thinking. Our results show that memory management is an independently learnable skill, and a high-leverage objective yielding large gains on long-horizon tasks.
UltraFlux: 多様なアスペクト比にわたる高品質のネイティブ 4K テキストから画像への生成のためのデータ モデルの共同設計
拡散トランスは最近、解像度 1K 付近で強力なテキストから画像への生成を実現していますが、拡散トランスをさまざまなアスペクト比でネイティブ 4K に拡張すると、位置エンコーディング、VAE 圧縮、最適化にまたがる密結合の障害モードが明らかになることを示します。これらの要因のいずれかに個別に取り組むことで、実質的な品質が確保されます。したがって、データモデルの共同設計の観点を採用し、MultiAspect-4K-1M 上の 4K でネイティブにトレーニングされた Flux ベースの DiT、UltraFlux を導入します。これは、制御されたマルチ AR カバレッジ、バイリンガル キャプション、および解像度と AR を意識したサンプリングのための豊富な VLM/IQA メタデータを備えた 1M 画像 4K コーパスです。モデル側では、UltraFlux は、(i) Resonance 2D RoPE を YaRN と組み合わせて、4K でのトレーニング ウィンドウ、周波数、AR 対応の位置エンコーディングを実現します。 (ii) 4K 再構築の忠実度を向上させる、シンプルで非敵対的な VAE ポストトレーニング スキーム。 (iii) タイムステップおよび周波数帯域全体で勾配のバランスを再調整する SNR 対応の Huber Wavelet 目標。 (iv) 段階ごとの美的カリキュラム学習戦略。これは、事前モデルによって支配される高ノイズのステップに高度な美的監督を集中させます。これらのコンポーネントを組み合わせることで、安定したディテールを維持した 4K DiT が生成され、幅広、正方形、高さの AR 全体に汎用化されます。 4096 ベンチマークおよびマルチ AR 4K 設定での Aesthetic-Eval では、UltraFlux は、忠実度、美的感覚、アライメント メトリクス全体で強力なオープンソース ベースラインを常に上回っており、LLM プロンプト リファイナーにより、独自の Seedream 4.0 と同等またはそれを上回っています。
原文 (English)
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect ratios exposes a tightly coupled failure mode spanning positional encoding, VAE compression, and optimization. Tackling any of these factors in isolation leaves substantial quality on the table. We therefore take a data-model co-design view and introduce UltraFlux, a Flux-based DiT trained natively at 4K on MultiAspect-4K-1M, a 1M-image 4K corpus with controlled multi-AR coverage, bilingual captions, and rich VLM/IQA metadata for resolution- and AR-aware sampling. On the model side, UltraFlux couples (i) Resonance 2D RoPE with YaRN for training-window-, frequency-, and AR-aware positional encoding at 4K; (ii) a simple, non-adversarial VAE post-training scheme that improves 4K reconstruction fidelity; (iii) an SNR-Aware Huber Wavelet objective that rebalances gradients across timesteps and frequency bands; and (iv) a Stage-wise Aesthetic Curriculum Learning strategy that concentrates high-aesthetic supervision on high-noise steps governed by the model prior. Together, these components yield a stable, detail-preserving 4K DiT that generalizes across wide, square, and tall ARs. On the Aesthetic-Eval at 4096 benchmark and multi-AR 4K settings, UltraFlux consistently outperforms strong open-source baselines across fidelity, aesthetic, and alignment metrics, and-with a LLM prompt refiner-matches or surpasses the proprietary Seedream 4.0.
DigitalCoach: 人間とエージェントのコンピューター使用コーチングにおけるコミュニケーションとグラウンディングのギャップ
エージェントはソフトウェア タスクを自動化できるようになってきていますが、エージェントは人間にソフトウェアの使い方を自分で教えることができるのでしょうか? DigitalCoach は、5 つのソフトウェア アプリケーションにわたる 28.1 時間の画面および入力イベントの記録に基づいた、22,752 の対話ターンで構成される 72 人の専門家と初心者のコンピューター使用コーチング セッションのマルチモーダル データセットです。私たちは DigitalCoach を使用して、最先端のモデルが人間にコンピューターの使い方を教えることができるかどうかを評価します。自動評価では、モデルは指導方法が人間とは異なることが示されています。モデルはより直接的な指示を提供しますが、説明、エラー診断、知識確認の質問は少ないです。コーチング手法を修正すると、モデルは人間のリファレンスに似た発話を生成しますが、視覚的なコンテキストにあまり基づいていません。インタラクティブな評価では、モデルコーチは学習者に深い関与を行わずに受動的に指示に従うようにさせ、視覚的な基礎付けが不十分であることを確認しています。 DigitalCoach は、共同的かつ積極的にコンピュータを使用してコーチング エージェントを行うための基盤を築きます。
原文 (English)
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sessions consisting of 22,752 dialogue turns grounded in 28.1 hours of screen and input event recordings across five software applications. We use DigitalCoach to evaluate whether state-of-the-art models can teach humans how to use computers. Automated evaluation shows that models differ from humans in how they coach: models provide more direct instructions, but fewer explanations, error diagnoses, and knowledge-check questions. When we fix the coaching method, models produce utterances similar to human references yet poorly grounded in visual context. Interactive evaluation confirms that model coaches cause learners to passively follow instructions without deeper engagement and fall short in visual grounding. DigitalCoach lays a foundation for collaborative and proactive computer use coaching agents.
パーソナルナレッジグラフの「文字列」から「モノ」へ: レコメンデーションシステムのための LLM トリプル抽出の評価
パーソナル ナレッジ グラフ (PKG) は、ユーザーの好みをモデル化するためのプライバシー保護フレームワークを提供しますが、構造化されていない分散型の会話データからそれを構築することは依然として課題です。この論文では、軽量の大規模言語モデルを使用して構造化されたユーザー設定トリプルを抽出するための再現可能なパイプラインを提示することで、会話的な「文字列」と意味論的な「もの」の間のギャップを埋めます。 PKG 構築のための会話データから Wikidata 識別子にリンクされた RDF 準拠のトリプルを抽出する機能について、Qwen および Gemma ベースのモデルを評価します。私たちの評価では、意味抽出の忠実度と、下流の推奨タスクにおける結果のグラフの有用性の両方が評価されます。特定のモデルのパフォーマンスが良好で、トリプル抽出パフォーマンスに比例して下流パフォーマンスが高いことがわかりました。
原文 (English)
From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems
Personal Knowledge Graphs (PKGs) offer a privacy-preserving framework for modeling user preferences, yet constructing them from unstructured, decentralized conversational data remains a challenge. This paper bridges the gap between conversational "strings" and semantic "things" by presenting a reproducible pipeline for extracting structured user-preference triples using lightweight Large Language Models. We evaluate Qwen- and Gemma-based models on their ability to extract RDF-compliant triples linked to Wikidata identifiers from conversational data for PKG construction. Our evaluation assesses both the semantic extraction fidelity and the utility of the resulting graphs in a downstream recommendation task. We found that certain models performed well and had proportionally high downstream performance relative to their triple extraction performance.
高度なエンコーダがスパース検索で遅れるのはなぜですか?語彙のギャップを埋めるための答えとアプローチ
ModernBERT のような高度な基盤モデルは、密な検索では古いアーキテクチャを大幅に上回っていますが、学習されたスパース検索 (LSR) では、古い BERT ベースのベースラインに驚くほど遅れをとっています。根本原因は \textit{語彙ギャップ} であると特定します。最新のトークナイザーは、可逆再構成用に設計された生の大文字と小文字を区別する語彙を利用しており、単一の意味単位を冗長な表面形式にマッピングし、形態学的ノイズでモデル容量を浪費し、語彙一致を妨げます。我々は理論的枠組みを通じてこの直観を形式化し、意味論的な整合性が保たれる限り、適切な語彙の粗視化が仮説クラスの複雑さを軽減して一般化限界を狭めることができることを実証した。これを解決するために、私たちは \textbf{Vocabulary Transfer (VT)} を提案します。これは、高度なエンコーダを最小限の計算コストでスパースに適した正規化された語彙に移行する、モデルに依存しないフレームワークです。 VT は、空間トポロジーを介した新しい \textbf{セマンティック初期化} を利用して幾何学的構造を保存し、 \textbf{活性化電位キャリブレーション (APC)} メカニズムを利用して、事前学習された多様体をスパース性制約に合わせて調整し、標準的な微調整で観察される死んだニューロンや密集した崩壊を防ぎます。経験的に、VT は普遍的に効果的です。これにより、ModernBERT が BEIR ベンチマークで最先端のパフォーマンス (\textbf{52.4} nDCG、\textbf{+4.7} の改善) を達成できるようになり、RoBERTa-large のような失敗したモデルが復活し、推論不要のアーキテクチャと特殊なドメインにシームレスに一般化されます。これらの結果は、パフォーマンスの遅れがアーキテクチャ上の欠陥ではなく、解決可能な語彙の不一致であることを裏付けています。コードとモデルをリリースしました。\footnote{https://anonymous.4open.science/r/vocab-transfer/。すべての詳細が含まれています。}
原文 (English)
Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps
While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the aging BERT-base baseline in learned sparse retrieval (LSR). We identify the root cause as the \textit{Vocabulary Gap}: modern tokenizers utilize raw, case-sensitive vocabularies designed for lossless reconstruction, which map single semantic units to redundant surface forms, wasting model capacity on morphological noise and hindering lexical matching. We formalize this intuition through a theoretical framework, demonstrating that appropriate vocabulary coarse-graining can tighten the generalization bounds by reducing complexity of the hypothesis class, provided that semantic integrity is preserved. To resolve this, we propose \textbf{Vocabulary Transfer (VT)}, a model-agnostic framework that migrates advanced encoders to sparse-friendly, normalized vocabularies with minimal computational cost. VT utilizes a novel \textbf{Semantic Initialization} via spatial topology to preserve geometric structure and an \textbf{Activation Potential Calibration (APC)} mechanism to align pre-trained manifolds with sparsity constraints, preventing the dead neuron and dense collapse observed in standard fine-tuning. Empirically, VT is universally effective: it enables ModernBERT to achieve state-of-the-art performance on the BEIR benchmark (\textbf{52.4} nDCG, a \textbf{+4.7} improvement), resuscitates failing models like RoBERTa-large, and generalizes seamlessly to inference-free architectures and specialized domains. These results confirm that the performance lag is not an architectural deficiency but a solvable vocabulary mismatch. We've released our code and models.\footnote{https://anonymous.4open.science/r/vocab-transfer/. All details included.}
トポロジカルボイド解析 知識空間における体系的な技術革新発見のための数学的フレームワーク
オペレーティング システムやハードウェア/ソフトウェアの共同設計など、高密度の技術領域でどこを革新するかを特定することは、基本的に高次元の知識空間での検索問題です。既存のアプローチは、キーワード検索、引用の近接性、または人間の直感に依存しており、いずれも、対象の目標に同時に関連し、かつ従来技術にはない未踏の領域の概念を形式化するものではない。我々は、密-疎ハイブリッド埋め込み空間においてトポロジカルボイドをトライアド(A、B、C)として定義する数学的フレームワークであるトポロジカルボイド解析(TVA)を紹介します。ボイドには 3 つの条件が必要です。(i) 概念 A と概念 B の両方がドメイン アンカー C と意味的に一貫している。 (ii) それらのペアごとの類似性は、校正された限界帯域内に収まり、明らかな組み合わせと無関係なノイズの両方を回避します。 (iii) それらはスパース語彙ブリッジを共有しますが、埋め込まれた超球上の測地線の中点は占有されていません。 TVA は、約 140,000 のインデックス付きドキュメントに適用され、96 のターゲットにわたって 2,128 の発明候補を生成します。 90% が自動品質フィルタリングを生き残り、4 人の専門家による敵対的レビューから 191 件の改訂と 1 件の承認の評決が得られました (0.05% がエンドツーエンド)。 2 つのケーススタディでは、フレームワークが単に明らかな関連ペアではなく、非明らかな結合組織を表面化していることを示しています。
原文 (English)
Topological Void Analysis A Mathematical Framework for Systematic Technical Innovation Discovery in Knowledge Spaces
Identifying where to innovate in a dense technical domain - such as operating systems or hardware/software co-design - is fundamentally a search problem in a high-dimensional knowledge space. Existing approaches rely on keyword search, citation proximity, or human intuition, none of which formalise the notion of an unexplored region that is simultaneously relevant to a target goal and absent from prior art. We present Topological Void Analysis (TVA), a mathematical framework that defines topological voids as triads (A, B, C) in a dense-sparse hybrid embedding space. A void requires three conditions: (i) both concepts A and B are semantically cohesive with domain anchor C; (ii) their pairwise similarity falls within a calibrated marginality band - avoiding both obvious combinations and unrelated noise; and (iii) they share a sparse lexical bridge while the geodesic midpoint on the embedding hypersphere is unoccupied. Applied to ~140k indexed documents, TVA generates 2,128 invention candidates across 96 targets; 90% survive automated quality filtering, yielding 191 REVISE and 1 APPROVE verdict from four-specialist adversarial review (0.05% end-to-end). Two case studies demonstrate the framework surfaces non-obvious connective tissue rather than merely obvious related pairs.
根拠のないペルソナ: 体制依存と LLM の個別化問題
LLM の個別化問題に対する Beckmann & Butlin (2026) の存在論的フレームワークは、ペルソナ ベクトルの文献から議論の余地のない領域間共参照の仮定を継承しています。つまり、同じ方向が、プロンプト コンディショニング、勾配降下微調整、および推論時間ステアリングの下で同じ内容を選択するというものです。 Qwen3-4B-Instruct および Mistral-7B-Instruct-v0.2 でのペルソナ トポロジー実験から得られた 4 つの経験的なウェッジを紹介します。プロンプト抽出されたベクトルと微調整盆地の非共線性です。架空のペルソナは、実際のアンカーよりも強く実際のアンカーの方向に沿ってモデルを移動させます。トレーニング履歴によって決定されるアトラクターに偏った、矛盾した原子価の混合物。そして、推論時の算術トレーニングと微調整時のキメラトレーニングにおける非対称構成代数 - これらは共同して仮定を台無しにします。私たちは、レジームインデックス付きの個性化を提案します。表現コンテンツのアイデンティティ単位は、ビークル単体ではなく、(ビークル、レジーム)のペアです。この枠組みの下では、ベックマンとバトリンの 3 つの候補的立場は、同じ指示対象をめぐって競合するのではなく、3 つの異なる体制内部の対象を記述しています。同じ診断がモロ&ミリエレ、チャルマーズ、セルロにも当てはまります。
原文 (English)
Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem
Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumption from the persona-vectors literature: that the same direction picks out the same content under prompt-conditioning, gradient-descent fine-tuning, and inference-time steering. We present four empirical wedges from persona-topology experiments on Qwen3-4B-Instruct and Mistral-7B-Instruct-v0.2 - non-collinearity of prompt-extracted vectors and fine-tune basins; fictional personas displacing the model along real-anchor directions more strongly than real anchors do; contradictory-valenced mixtures biased toward a training-history-determined attractor; and asymmetric compositional algebra under inference-time arithmetic versus fine-tune-time chimera training - that jointly undermine the assumption. We propose regime-indexed individuation: the identity unit for representational content is a (vehicle, regime) pair, not a vehicle alone. Under this framework, Beckmann & Butlin's three candidate positions describe three different regime-internal objects rather than competing for the same referent; the same diagnosis applies to Mollo & Milli\`ere, Chalmers, and Cerullo.
BaRA: BFS およびリフレクション Web データ収集エージェント
大規模言語モデル (LLM) ベースの Web エージェントは、Web データ収集のための手動スクリプトを削減しますが、実際の Web サイトでは、関連するページを見逃したり、不完全なマルチモーダル出力を返したり、直接ダウンロードできないメディア URL を返したりすることがよくあります。固定インタラクション予算の下でサイトレベルの収集を行うためのフレームワークである BFS-and-Reflection Agent (BaRA) を紹介します。このフレームワークは、境界付き幅優先検索 (BFS) トラバーサルと履歴ベースの自己反映を組み合わせています。私たちは、グラウンドトゥルースの参照セットを使用して、50 の合成 Web サイトで BaRA を評価します。さらに、乱雑なレイアウトまたは動的なレイアウトを備えた 3 つの公開 Web サイトでもテストしました。 BaRA は、リンク検出とダウンロード可能なマルチモーダル抽出において Pure LLM、SeeAct-Vision、およびブラウザでの使用を上回り、ダウンロード有効な画像とビデオの回復において最大の利益をもたらします。私たちのコードは https://github.com/MLAI-Yonsei/BaRA-Agent で入手できます。
原文 (English)
BaRA: BFS-and-Reflection Web Data Collection Agent
Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant pages, return incomplete multimodal outputs, or return media URLs that are not directly downloadable. We present BFS-and-Reflection Agent (BaRA), a framework for site-level collection under a fixed interaction budget. The framework combines bounded breadth-first search (BFS) traversal with history-based self-reflection. We evaluate BaRA on 50 synthetic websites with ground-truth reference sets. We additionally test on three public websites with cluttered or dynamic layouts. BaRA outperforms Pure LLM, SeeAct-Vision, and Browser-use on link discovery and downloadable multimodal extraction, with the largest gains in download-valid image and video recovery. Our code is available at https://github.com/MLAI-Yonsei/BaRA-Agent.
SchemaRAG: LLM 駆動の構造化情報抽出のための動的な大規模スキーマ削減
ターゲットのスキーマが大きくて複雑な場合、大規模言語モデル (LLM) を使用して非構造化テキストから構造化データを抽出することは困難になります。このような場合、プロンプトに完全なスキーマを含めると、コストと待ち時間が増加し、中間損失のパフォーマンス低下の危険性があり、コンテキストの長さの制限を超える可能性があります。私たちは、スキーマ メタデータと利用可能な場合は少数のショットの例を活用することで、スキーマ条件付き情報抽出タスクの出力スキーマ空間を動的にプルーニングする検索拡張生成 (RAG) フレームワークである SchemaRAG を提案します。実際の医療および電子商取引のデータセットで SchemaRAG を評価します。私たちの結果は、SchemaRAG が micro-F1 で最大 8.8% の増加、レイテンシーの 47% 削減、トークン コストの 48% 削減を達成できることを示しており、大規模なスキーマ抽出に対する実用性を実証しています。
原文 (English)
SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction
Extracting structured data from unstructured text using large language models (LLMs) becomes challenging when target schemas are large and complex. In such cases, including the full schema in the prompt increases cost and latency, risks lost-in-the-middle performance degradation, and can exceed context length limits. We propose SchemaRAG, a retrieval-augmented generation (RAG) framework that dynamically prunes the output schema space for schema-conditioned information extraction tasks by leveraging schema metadata and few-shot examples when available. We evaluate SchemaRAG on real-world healthcare and e-commerce datasets. Our results show that SchemaRAG can achieve up to an 8.8% increase in micro-F1, a 47% reduction in latency, and a 48% reduction in token costs, demonstrating its practicality for large-schema extraction.
強化されたライティング支援のための制御可能なナラティブ レンダリング
基本的なライティング支援における大規模言語モデル (LLM) の優れた能力にもかかわらず、創造的なライティングにおける LLM の有用性は、永続的なバイナリ障害によって根本的に妨げられています。この問題は、修復研磨と呼ばれる安全な表面レベルの編集と、破壊的で制御されていないプロット拡張の間の揺れとして現れます。このジレンマは、物語の忠実さと説明の強度との間の重要なトレードオフを定義します。我々は、物語と言説の間の物語学的区別に基づいた執筆支援フレームワークである Loom を提案します。 Loom は、意図を中心とした記号論的思考連鎖を運用する 3 層パイプラインを採用し、物語の意図とレンダリング密度を正確に制御します。このアーキテクチャは、知覚マテリアルの生成を構文の挿入から分離し、元のイベント構造に違反することなく強化が確実に行われるようにします。 LLM ベースの指標と人間による評価を含む当社の包括的な評価は、Loom がこの根本的な緊張をうまく解決していることを示しています。 Loom は最高の総合品質スコアを達成し、最先端のベースラインと比較して事実の完全性と説明の強度が大幅に向上しました。
原文 (English)
Controllable Narrative Rendering for Enhanced Assisted Writing
Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary failure. This issue manifests as an oscillation between safe, surface-level editing, referred to as remedial polishing, and destructive, uncontrolled plot expansion. This dilemma defines a critical trade-off between narrative fidelity and descriptive intensity. We propose Loom, an assisted writing framework grounded in the narratological distinction between story and discourse. Loom employs a three-layer pipeline that operationalizes an intent-centered semiotic chain-of-thought to enforce precise control over narrative intent and rendering density. This architecture separates the generation of perceptual material from syntactic insertion, ensuring that enhancement occurs without violating the original event structure. Our comprehensive evaluation, which includes LLM-based metrics and human assessment, demonstrates that Loom successfully resolves this fundamental tension. Loom achieves the highest overall quality score, yielding substantial gains in factual integrity and descriptive intensity compared to state-of-the-art baselines.
会話型レコメンダー システムにおけるユーザー シミュレーションの即時最適化: 多目的フレームワーク
会話型レコメンダー システム (CRS) は、ユーザーが積極的に好みを引き出し、意図を明確にし、レコメンデーションをリアルタイムで適応できるため、次世代のインテリジェント レコメンダー システムの中核コンポーネントです。ただし、CRS ドメインには、評価とトレーニング データへのアクセスという 2 つの重要な障害があります。実際の人を対象とした研究を通じて CRS を評価することは、従来のレコメンダー システムよりも重要ですが、そのような研究には費用も時間もかかります。さらに、プライバシー上の懸念により、CRS インタラクション データをモデル トレーニング用に取得するのが困難なことがよくあります。大規模言語モデル (LLM) ベースのユーザー シミュレーターは、評価とトレーニング用の合成ユーザー インタラクションを生成することで両方の課題に対処することが期待されています。しかし、既存のアプローチは体系的なポジティブバイアス、データ漏洩、行動の多様性の制限に悩まされており、広範なドメイン専門知識を必要とする脆弱な手動プロンプトエンジニアリングに依存しています。この論文では、CRS の LLM ベースのユーザー シミュレーターのプロンプトを自動的に最適化し、同時にこれらの問題を軽減するフレームワークを提案します。実験結果は、提案されたフレームワークが、さまざまなプロンプト設定におけるベースライン手法と比較して、人間の対話パターンとの行動の整合性を向上させることを示しています。
原文 (English)
Prompt Optimization for User Simulation in Conversational Recommender Systems: A Multi-Objective Framework
Conversational recommender systems (CRSs) are a core component of next-generation intelligent recommender systems because they enable users to actively elicit preferences, clarify intentions, and adapt recommendations in real time. However, there are two key obstacles in the CRS domain: evaluation and access to training data. Evaluating CRSs through real human studies is more critical than for traditional recommender systems, yet such studies are both costly and time-consuming. Moreover, CRS interaction data are often difficult to obtain for model training due to privacy concerns. Large language model (LLM)-based user simulators have shown promise in addressing both challenges by generating synthetic user interactions for evaluation and training. However, existing approaches suffer from systematic positive bias, data leakage, and limited behavioral diversity, and they rely on brittle manual prompt engineering that requires extensive domain expertise. In this paper, we propose a framework to automatically optimize prompts for LLM-based user simulators in CRSs, simultaneously mitigating these issues. Experimental results demonstrate that the proposed framework achieves improved behavioral alignment with human interaction patterns compared to baseline methods across diverse prompt settings.
SkillSelect-Serve: 小規模 LLM エージェント向けの予算管理可能で QoS を意識したスキル サービスの推奨と構成
再利用可能なスキル ライブラリは、大規模言語モデル (LLM) エージェントにとって重要なインフラストラクチャになりつつありますが、既存の選択方法では、スキルを取得可能なドキュメントとして扱い、固定の上位 K リストを返すことがよくあります。この文書では、エージェントのスキル選択をスキル サービスの推奨および構成として定式化する、予算管理可能で QoS を意識したフレームワークである SkillSelect-Serve について説明します。 SkillSelect-Serve は、機能の説明、依存関係、コンテキスト コスト、リスク、QoS 関連の属性を備えた構造化されたスキル サービスとして生のスキルを表します。ローカルの Micro-Agent Requirement Planner が自然言語タスクを構造化されたサービス要件に変換し、共有ディスカバリー バックボーンが大規模なレジストリから候補サービスを取得します。次に、このフレームワークは、スキルレベルの限界適合性推定と、カバレッジ、冗長性、コスト、およびリスクのトレードオフに関するバンドルレベルの調整を行う二重粒度ユーティリティモデリングを実行します。 35,353 のスキルと 586 のタスク クエリに関する実験では、SkillSelect-Serve が固定の上位 K 取得ベースラインと比較して、同一予算バンドルの再現率と平均ユーティリティを一貫して向上させることが示されています。
原文 (English)
SkillSelect-Serve: Budget-Controllable and QoS-Aware Skill Service Recommendation and Composition for Small LLM Agents
Reusable skill libraries are becoming important infrastructure for large language model (LLM) agents, yet existing selection methods often treat skills as retrievable documents and return fixed top-k lists. This paper presents SkillSelect-Serve, a budget-controllable and QoS-aware framework that formulates agent skill selection as Skill Service Recommendation and Composition. SkillSelect-Serve represents raw skills as structured Skill Services with functional descriptions, dependencies, context cost, risk, and QoS-related attributes. A local Micro-Agent Requirement Planner converts natural-language tasks into structured service requirements, while a shared discovery backbone retrieves candidate services from a large registry. The framework then performs dual-granularity utility modeling with skill-level marginal suitability estimation and bundle-level calibration for coverage, redundancy, cost, and risk trade-offs. Experiments on 35,353 skills and 586 task queries show that SkillSelect-Serve consistently improves same-budget bundle recall and mean utility over fixed top-k retrieval baselines.
PRA-RAG: 検索の破損に対する検索拡張生成における堅牢な集約が証明されています
検索拡張生成 (RAG) は、外部知識を組み込むことで大規模言語モデル (LLM) を強化し、固有の知識制限を効果的に軽減します。ただし、RAG は、取得したテキストを操作してモデルの出力を誤解させるポイズニング攻撃に対して依然として脆弱です。既存の防御メカニズムには理論的な堅牢性の保証が欠けていることが多く、LLM が取得したコンテンツについての知識が限られている場合には、動作の信頼性が低くなります。この研究では、取得されたテキストに対するポイズニング攻撃を防御するように設計された堅牢な検索集約アルゴリズムである PRA-RAG を提案します。 PRA-RAG は、取得したテキストの複数の組み合わせをサンプリングし、埋め込み空間の幾何学的構造を利用して堅牢なサブセットを特定し、そこから安定した集約表現を導き出します。私たちは、ポイズニングされた取得コンテンツの最大の影響に関する理論的限界を提供し、RAG の堅牢性の定量的尺度を確立します。複数のベンチマークと RAG アーキテクチャにわたる実験では、PRA-RAG が 71% の精度を維持しながら攻撃の成功率を 1% まで低下させ、代表的な最先端の手法を大幅に上回るパフォーマンスを示していることが実証されています。
原文 (English)
PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, effectively mitigating their inherent knowledge limitations. However, RAG remains vulnerable to poisoning attacks that manipulate retrieved texts to mislead model outputs. Existing defense mechanisms often lack theoretical robustness guarantees and perform unreliably when the LLM has limited knowledge of the retrieved content. In this work, we propose PRA-RAG, a provably robust retrieval aggregation algorithm designed to defend against poisoning attacks on retrieved texts. PRA-RAG samples multiple combinations of retrieved texts and utilizes geometric structures in the embedding space to identify a robust subset, from which a stable aggregated representation is derived. We provide theoretical bounds on the maximum impact of poisoned retrieved content and establish a quantitative measure of RAG's robustness. Experiments across multiple benchmarks and RAG architectures demonstrate that PRA-RAG reduces the attack success rate to as low as 1% while maintaining an accuracy of 71%, significantly outperforming representative state-of-the-art methods.
GRACE-RAG: クローズドドメインの機関設定での軽量展開を可能にする、正規証拠合成のためのガバナンド検索アーキテクチャ
検索拡張生成(RAG)システムは、回答を信頼できる文書に基づいて行う必要がある組織の質問応答環境で広く使用されています(Gao et al.、2023)。関連情報が異種文書に分散されているエンティティ密度の高いドメインでは、ベクトルのみの検索では断片的な証拠が生成されることが多く、推論時間推論への依存度が高まります (Zhao et al., 2024)。この論文では、生成段階から構造化された検索層まで構造的推論を外部化する、検索主導のグラフ拡張型 RAG アーキテクチャである GRACE-RAG を紹介します。これにより、オフラインで構造的曖昧性が解決され、クローズド ドメインの制度用語に合わせて調整された自己ホスト型の軽量モデルへの展開が可能になります。 Mistral 24B、GPT OSS 120B、Gemini 2.5 Flash の 3 つのモデル容量にわたる実験では、完全性、深度、予測カバレッジにおいて一貫した改善が見られ、中規模モデルでは全体的な品質が最大 20% 向上しました。これは、取得アーキテクチャがモデル スケール全体にわたる構造品質を支配し、独自のシステムに依存することなく計算量と待ち時間のフットプリントを削減することを示しています。
原文 (English)
GRACE-RAG: Governed Retrieval Architecture for Canonical Evidence Synthesis, Enabling Lightweight Deployment in Closed-Domain Institutional Settings
Retrieval-Augmented Generation (RAG) systems are widely used in institutional question answering settings where responses must be grounded in authoritative documentation (Gao et al., 2023). In entity-dense domains where relevant information is distributed across heterogeneous documents, vector-only retrieval often produces fragmented evidence and increases dependence on inference-time reasoning (Zhao et al., 2024). This paper introduces GRACE-RAG, a retrieval-governed, graph-augmented RAG architecture that externalizes structural reasoning from the generative stage to a structured retrieval layer, resolving structural ambiguity offline, enabling deployment on self-hosted lightweight models calibrated to closed-domain institutional vocabulary. Experiments across three model capacities: Mistral 24B, GPT OSS 120B, and Gemini 2.5 Flash show consistent improvements in completeness, depth, and anticipatory coverage, with overall quality gains of up to 20% under mid-scale models, indicating that retrieval architecture governs structural quality over model scale, reducing computational and latency footprint without dependence on proprietary systems.
住宅用建物の間取り図適合性チェックのための自動化された AI ベースのフレームワークに向けて
オーストラリアの都市部の住民の幸福を改善するために、政府はアパートの設計品質を高めるために SEPP65、BADS、SPP7.3 などの政策改革を導入しました。これらの規制では、採光、自然換気、プライバシー、スペース効率などの健康関連機能を評価するために、正確な幾何学的および空間的分析が必要です。ただし、コンプライアンスチェックは手作業で時間がかかるため、依然として困難です。さらに、進化するポリシーにより、数千のアパートにわたる大規模な評価の拡張性が制限されます。既存の自動フロアプラン分析方法は細分化されており、通常は単一のアパートに焦点を当てており、複数ユニットのコンプライアンスチェックのための統一されたフレームワークが欠けています。この記事では、自動フロア プラン分析、特に AI を活用したアプローチの現在の進歩を調査し、実際の導入における主要な課題に焦点を当てます。これらのギャップに対処するために、複数の集合住宅の建物における自動コンプライアンスチェックのための概念的なフレームワークが提案されています。 Large Language Model (LLM) はルール エンジン内で使用され、テキストの構築コードを実行可能で説明可能なルールに変換します。データ抽出エンジンは、間取り図の画像を壁、部屋、設備、テキスト、シンボルなどの要素に分割し、トポロジ関係を備えた構造化された建物グラフに変換します。この構造化表現は、LLM が生成した評価ルールを利用するコンプライアンス チェック エンジンによって評価されます。提案されたフレームワークは、管轄区域全体で自動化されたコンプライアンスチェックに対するスケーラブルで一貫性のある透明性のあるアプローチを提供し、アパートの設計基準の効率的な施行をサポートし、より健全で高密度の都市開発を促進します。
原文 (English)
Towards an automated AI-based framework for floor plan compliance checks for residential buildings
To improve residents' well-being in Australia's urban areas, governments have introduced policy reforms such as SEPP65, BADS, and SPP7.3 to enhance apartment design quality. These regulations require precise geometric and spatial analysis to evaluate health-related features, including daylight access, natural ventilation, privacy, and space efficiency. However, compliance checking remains challenging due to its manual, time-intensive nature. Additionally, evolving policies limit scalability for large-scale assessments across thousands of apartments. Existing automated floor plan analysis methods are fragmented and typically focus on single apartments, lacking a unified framework for multi-unit compliance checking. This article explores current advancements in automated floor plan analysis, particularly AI-driven approaches, and highlights key challenges in their practical adoption. To address these gaps, a conceptual framework is proposed for automated compliance checking in multi-apartment buildings. A Large Language Model (LLM) is used within a Rule Engine to convert textual building codes into executable, explainable rules. A Data Extraction Engine segments floor plan images into elements such as walls, rooms, fixtures, text, and symbols, and transforms them into a structured building graph with topological relationships. This structured representation is then evaluated by a Compliance Check Engine, which leverages LLM-generated rules for assessment. The proposed framework offers a scalable, consistent, and transparent approach to automated compliance checking across jurisdictions, supporting efficient enforcement of apartment design standards and promoting healthier, higher-density urban development.
Libra: エージェントによる情報取得のための環境のトレーニング
大規模なリポジトリ内での情報のローカライゼーションは、エージェント LLM システムの基礎です。合成データ駆動型の最適化は LLM のトレーニングに成功していることが証明されていますが、エージェントの作業環境 (リポジトリ自体) をデータ駆動型で最適化することにはほとんど注目されていません。このギャップを埋めるために、可変の「カタログ」(ナビゲート可能なインデックスとして機能する階層型 Markdown ファイル) をリポジトリに導入する自己進化フレームワークである Libra を紹介します。 Libra は LLM 主導の最適化ループを実行します。このループでは、プロンプターが合成クエリを生成し、凍結されたソルバーがカタログをナビゲートしてそれらの解決を試み、ヒーラーがソルバーのローカリゼーションの失敗に応じてカタログを書き換えます。 12 個の SWE-bench Lite リポジトリにわたる評価では、この環境修復によってコード ローカリゼーションの精度が継続的に対数的に向上することが実証されました。さらに、これらの環境改善により、さまざまな LLM および問題セット間でゼロショットが転送されます。この論文の焦点はそのようなシステムの一般的な動作を研究することですが、Libra に最適化されたカタログを備えた最小限のコーディング エージェントが最先端のベースラインを上回るパフォーマンスを示すことも実証します。コードは https://github.com/salesforce-misc/Libra で、データは https://huggingface.co/datasets/Salesforce/Libra で入手できます。
原文 (English)
Libra: Training the Environment for Agentic Information Retrieval
Information localization within massive repositories is a cornerstone of agentic LLM systems. While synthetic data-driven optimization has proven successful in training LLMs, little attention has been paid to optimizing the agent's working environment (the repository itself) in a data-driven manner. To bridge this gap, we present Libra, a self-evolving framework that introduces mutable "catalogs" (hierarchical Markdown files serving as navigable indices) into the repository. Libra runs an LLM-driven optimization loop where a Prompter generates synthetic queries, a frozen Solver attempts to resolve them by navigating the catalogs, and a Healer rewrites the catalogs in response to the Solver's localization failures. Evaluations across 12 SWE-bench Lite repositories demonstrate that this environmental healing yields continual, logarithmic improvements in code localization accuracy. Furthermore, these environmental improvements transfer zero-shot across different LLMs and problem sets. Although the focus of this paper is to study the general behavior of such a system, we also demonstrate that a minimalist coding agent equipped with Libra-optimized catalogs outperforms state-of-the-art baselines. Code is available at https://github.com/salesforce-misc/Libra and data at https://huggingface.co/datasets/Salesforce/Libra.
ユーザー認識の想起の学習: 長期会話記憶におけるパーソナライズされた検索
長期的に会話を行うエージェントは過去の対話を記憶していることが期待されますが、記憶が役立つのは、適切なユーザーに対して適切な証拠が呼び出された場合のみです。既存のメモリ拡張 LLM エージェントは、コンパクトなメモリ バンクの構築において進歩を遂げていますが、検索は依然としてクエリ中心の類似性や固定ランキング ルールによって駆動されることが多く、ユーザー条件による関連性は十分に検討されていません。このギャップに対処するために、私たちは、メモリ検索をユーザー認識かつ最適化可能にする検索中心のフレームワークである Profile-guided Personalized Retrieval Optimization (PPRO) を提案します。PPRO は、エピソード的および最適化されたメモリ検索を構築します。対話履歴からセマンティック メモリ バンクを生成し、蓄積されたメモリからユーザー プロファイルを導き出します。このプロファイルは、メモリ ランキングにおける明示的なパーソナライズされた優先順位として機能し、安定したユーザー属性、好み、および関係性を考慮した検索を可能にします。PPRO はさらに、メモリ バンクと回答モデルを固定したまま、証拠取得品質と下流の回答品質の両方をフィードバックとして使用して、グループ相対ポリシー最適化を使用してクエリ リライタをトレーニングします。LoCoMo と LongMemEval-S での実験では、トレーニングなしと比べて一貫した向上が示されています。さらに、アブレーション研究では、プロファイルに基づくランキングと検索指向の書き換えの両方がパフォーマンスに大きく寄与していることが示されており、パーソナライズされた長期記憶使用の重要な要素として検索の最適化が強調されています。
原文 (English)
Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory
Long-term conversational agents are expected to remember past interactions, but memory is useful only when the right evidence is recalled for the right user. Existing memory-augmented LLM agents have made progress in building compact memory banks, yet retrieval is still often driven by query-centered similarity or fixed ranking rules, leaving user-conditioned relevance underexplored.To address this gap, we propose Profile-guided Personalized Retrieval Optimization (PPRO), a retrieval-centric framework that makes memory retrieval both user-aware and optimizable.PPRO builds episodic and semantic memory banks from dialogue histories and derives a user profile from accumulated memories.The profile serves as an explicit personalized prior in memory ranking, allowing retrieval to account for stable user attributes, preferences, and relationships.PPRO further trains a query rewriter with Group Relative Policy Optimization, using both evidence retrieval quality and downstream answer quality as feedback while keeping the memory banks and answer model fixed.Experiments on LoCoMo and LongMemEval-S show consistent gains over training-free memory systems and training-based baselines.Ablation studies further show that both profile-guided ranking and retrieval-oriented rewriting contribute substantially to performance, highlighting retrieval optimization as a key factor in personalized long-term memory use.
現実世界の LLM: 緊急事態における「AI」の評価
この文書は行動への呼びかけを提供します。私たちは研究コミュニティの同僚に対し、私たちの発見を一般に明らかにする上でより大きな役割を果たすよう強く求めます。リスクを説明するために、実世界のコンテキストでの LLM ベースの機械翻訳アプリケーションの導入の初期段階に関するケース スタディを紹介します。これは、オペレーターに直接電話をかけることが難しい緊急時に使用する、55 言語での text-2-911 システム広告機能です。私たちは、このようなテクノロジに関する多くの一般的な誤解を特定し、開発および展開パイプラインのあらゆる段階で関係者に向けた一連の具体的な推奨事項とベスト プラクティスで結論付けています。科学研究の進歩はしばしば「難しい」問題を解決することにありますが、最も見落とされているのは「簡単な」問題、つまり最新のテクノロジーが不要な問題であることが多いと私たちは主張します。
原文 (English)
LLMs in the Real World: Evaluating "AI" in Emergency Contexts
This paper offers a call to action. We urge our colleagues in the research community to play a greater role in the articulation of our findings to the public. To illustrate the stakes we present a case study on the initial stages of an LLM-based machine translation application's deployment in a real-world context: a text-2-911 system advertising capabilities in 55 languages for use in emergencies in which it may be difficult to call operators directly. We identify a number of common misconceptions about technologies such as these, concluding with a set of concrete recommendations and best practices for stakeholders at every stage of the development and deployment pipeline. While the advancement of scientific research often lies in solving the "hard" problems, we argue it is often the "easy" ones -- problems for which the latest technology is often unnecessary -- that are most overlooked.
スパースオートエンコーダを介して文の埋め込みを人間の概念に合わせる
密な文の埋め込みは、最新の検索拡張生成 (RAG) システムの基礎ですが、特徴の重ね合わせによる解釈可能性の欠如に悩まされています。この不透明さは、もつれた表現を分析したり制御したりすることが難しいため、検索プロセスを人間の意図に合わせるのを妨げます。この研究では、Top-k Sparse Autoencoder (SAE) を使用して、文変換器 (E5 など) の密な表現を人間が解釈可能な概念に解きほぐす方法を提案します。我々は、これらの解きほぐされた特徴が特定の意味論的、構文論的、および語用論的なカテゴリと一致することを実証します。さらに、検索プロセスへの正確な介入を可能にするアクティベーションステアリングメカニズムを導入します。特定の潜在的な特徴をクランプすることにより、バックボーン モデルを再トレーニングすることなく、検索結果を再ランク付けしてユーザーの制約に合わせて調整できることを示します。私たちの発見は、SAE ベースの分解が透過的で操作可能な神経情報検索への実行可能な道を提供することを示唆しています。
原文 (English)
Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders
Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to feature superposition. This opacity hinders the alignment of retrieval processes with human intent, as the entangled representations are difficult to analyze or control. In this work, we propose a method to disentangle the dense representations of sentence transformers (e.g., E5) into human-interpretable concepts using Top-k Sparse Autoencoders (SAEs). We demonstrate that these disentangled features align with specific semantic, syntactic, and pragmatic categories. Furthermore, we introduce an activation steering mechanism that allows for precise intervention in the retrieval process. By clamping specific latent features, we show that it is possible to re-rank search results to better align with user constraints without retraining the backbone model. Our findings suggest that SAE-based decomposition offers a viable path toward transparent and steerable neural information retrieval.
FLYNN: Fly Brain トポロジーを使用したロボット ナビゲーションのための堅牢なニューラル ネットワーク
深層学習モデルは複雑なタスクで最先端のパフォーマンスを実現しますが、新しい環境や感覚遮断に直面すると脆弱なままです。対照的に、生体系はこれらの課題に対して顕著な耐性を示します。私たちは、ショウジョウバエのシナプス分解能の脳コネクトームから直接派生したアーキテクチャをもつリカレント ニューラル ネットワーク (RNN) を開発することで、この脆弱性に対処します。我々は、MuJoCo でビジョンベースのナビゲーションを実行するためにフライ コネクトーム ニューラル ネットワーク (FLYNN) をトレーニングし、同様のパラメーター数の最新の手作りネットワークに匹敵するパフォーマンスを達成する実現可能性を実証します。重要なことは、FLYNN は、さらなるトレーニングを行わなくても、分布外 (OOD) データに対する優れた耐性と感覚喪失に対する耐性を示します。完全な視力喪失下でも機能を維持しましたが、手作りのネットワークは、カメラのドロップアウトで特別に訓練された場合でも、ほとんど機能しませんでした。 FLYNN の内部状態の主成分分析 (PCA) は、FLYNN が特に高度な表現モジュール性を示していることを示唆しており、これがその堅牢性に関連している可能性があります。私たちの研究は、生物学的な脳のトポロジーに従って弾力性のある人工エージェントを設計するための新しい方向性を提供します。
原文 (English)
FLYNN: Robust Neural Network for Robot Navigation using Fly Brain Topology
While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or sensory deprivation. In contrast, biological systems exhibit remarkable tolerance to these challenges. We address this vulnerability by developing a recurrent neural network (RNN) whose architecture is directly derived from the synaptic-resolution brain connectome of the fruit fly Drosophila melanogaster. We demonstrate the feasibility of training the fly connectome neural network (FLYNN) to perform vision-based navigation in MuJoCo, achieving performance comparable to modern hand-crafted networks of similar parameter counts. Crucially, FLYNN exhibits superior resistance to out-of-distribution (OOD) data and tolerance to sensory loss without further training. It remained functional even under total vision loss while hand-crafted networks largely failed, even when specifically trained with camera dropout. Principal Component Analysis (PCA) of the internal state of FLYNN suggests that it exhibits a particularly high degree of representational modularity, which might be related to its robustness. Our work provides a new direction for designing resilient artificial agents following the topology of biological brains.
身体化されたインテリジェンスのためのメモリネイティブの非地上ネットワーク
非地上ネットワーク (NTN) は、身体化知能 (EI) のユビキタス接続を提供し、荒野のロボットがクラウド リソースを活用したり、重要な情報をリモート センターに報告したりできるようにします。ただし、非常に動的で、リソースに制約があり、トポロジーが変化し、タスク指向の環境であるため、相乗効果は簡単ではありません。既存のメモリレス NTN プロトコルは、ローカル チャネルの状態と瞬間的なサービス要求によって決定が左右されるため、非効率になります。これらの制限に対処するために、この文書では、メモリ拡張システムの最適化にロングホライズン コンテキストを活用するメモリ ネイティブ NTN (MemNTN) パラダイムを提案します。このパラダイムシフトを実現するために、世界の現状を表す物理メモリと歴史的なネットワークエクスペリエンスをエンコードするデジタルメモリを区別するデュアルメモリアーキテクチャを確立します。当社は、物理層とアクセス層からネットワーク層とアプリケーション層に至るまで、クロスレイヤーのメモリネイティブの意思決定を容易にするメモリの取得、圧縮、評価、更新、および利用メカニズムを開発します。衛星による質問応答(SEQA)の実験により、提案された MemNTN が従来のステートレス NTN および地上アプローチよりも大幅に優れていることが実証されました。
原文 (English)
Memory-Native Non-Terrestrial Networks for Embodied Intelligence
Non-terrestrial networks (NTN) provide ubiquitous connectivity for embodied intelligence (EI), enabling robots in wilderness to leverage cloud resources or report critical information to remote centers. However, the synergy is nontrivial due to the highly-dynamic, resource-constrained, topology-varying, and task-oriented environment. Existing memoryless NTN protocols become inefficient, since the decisions are driven by local channel conditions and instantaneous service demands. To address these limitations, this paper proposes the memory-native NTN (MemNTN) paradigm that leverages long-horizon contexts for memory augmented system optimization. To realize this paradigm shift, we establish a dual-memory architecture that distinguishes between physical memory representing the state of the world and digital memory encoding historical network experience. We develop memory acquisition, compression, valuation, update, and utilization mechanisms that facilitate cross-layer, memory-native decision-making, spanning from the physical and access layers up to the network and application layers. Experiments in satellite embodied question answering (SEQA) demonstrate that the proposed MemNTN significantly outperforms conventional stateless NTN and terrestrial approaches.
コンタクトレンチを使った器用な操作を人間の実演から学ぶ
ロボットの器用な操作は人間の豊富なデモンストレーションから恩恵を受ける可能性がありますが、そのようなデモンストレーションをロボット政策に移すことは依然として困難です。我々は、強化学習による剛体および多関節オブジェクトの長期的な操作のためのフレームワークである、ロボットによる器用な操作における人間のデモンストレーション (CHORD) からのコンタクト レンチ ガイダンスを紹介します。重要なアイデアは、オブジェクト中心のコンタクト レンチ空間ガイダンスです。人間とロボットの動きを、オブジェクトに誘発できる力とトルクによって表現し、誘発された瞬間的な動きによって類似性を測定できるようにします。このガイダンスにより、強化学習は接触の多い器用な操作に対してよりスケーラブルになります。さらに、モーション キャプチャ データセットと再構築された社内ビデオから構築された、4,739 の両手による器用な操作タスクを含む大規模なシミュレーション ベンチマークを紹介します。 1,831 のベンチマーク タスクで評価した結果、CHORD は平均成功率 82.12% を達成し、強力なスケーラビリティを実証しました。また、CHORD は、手のみおよび三人称のデモンストレーションから全身操作に一般化し、90.77% の成功率を達成し、学習されたポリシーは、開ループ設定と閉ループ設定の両方で現実世界に転送されます。
原文 (English)
Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration
Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.
ATM: マルチエージェント コード共同合成のための CID ブローカーによる事前書き込み許可
マルチエージェント LLM システムでは、ソフトウェア エンジニアリングの作業を計画、生成、検証、修復に分解できますが、より狭いシステムの問題が残ります。管理された共有ミューテーションが適用される前に、システムは、同時に形成されたどの書き込みインテントを並行して進めることができるか、決定論的な構成またはシリアル化が必要で、どの書き込みインテントがフェイルクローズされたパスを取る必要があるかを決定する必要があります。私たちは、単一のガバナンス ドメイン内で動作するソフトウェア エージェント用の仕様に基づいたガバナンス基盤である AI-Atomic-Framework (ATM) を使用して、この問題に対処します。 ATM は、タスクの意図、リポジトリの範囲、書き込み許可、検証、および証拠の義務を 1 つのガバナンス チェーンにバインドします。 Content Identifier (CID) ブローカーは、共有突然変異受付サブシステムとして機能します。アダプターガイドの原子化マップは、意味論的なアトムと境界領域にインテントを書き込みます。永続的なアトム マップのカバレッジが不完全な場合、仮想アトムは保守的な比較とルーティングのための一時的な監査可能なガバナンス ユニットを提供します。管理された共有書き込みは、提案されたエージェントによって直接適用されるのではなく、中立的なスチュワードによって最終的に適用されます。評価では、12 のシナリオの決定論的設計マトリックス、3 つのアーカイブされたランナー ケース、ATM-AdmissionBench、3 つのアーカイブされた同一ファイル境界ケース、3 週間の外部アダプタ調査、および運用リカバリ ルーティング ベンチマークを含む、制御された証拠、フィールド 証拠、導入証拠、および拡張証拠を組み合わせます。この結果は、観察された単一ドメイン設定内での実現可能性、監査可能性、および限定された回復可能性を裏付けていますが、広範な比較優位性やクロスクローン ガバナンスを主張するものではありません。
原文 (English)
ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis
Multi-agent LLM systems can decompose software-engineering work into planning, generation, validation, and repair, but a narrower systems problem remains: before any governed shared mutation is applied, a system must decide which concurrently formed write intents may proceed in parallel, which require deterministic composition or serialization, and which must take a fail-closed path. We address this problem with the AI-Atomic-Framework (ATM), a specification-grounded governance substrate for software agents operating within a single governance domain. ATM binds task intent, repository scope, write admission, validation, and evidence obligations into one governance chain. A Content Identifier (CID) broker serves as the shared-mutation admission subsystem. Adapter-guided atomization maps write intents to semantic atoms and bounded regions; when persistent atom-map coverage is incomplete, virtual atoms provide temporary auditable governance units for conservative comparison and routing. Governed shared writes are ultimately applied by a neutral steward rather than directly by proposing agents. Evaluation combines controlled, field, adoption, and extension evidence, including a 12-scenario deterministic design matrix, three archived runner cases, ATM-AdmissionBench, three archived same-file boundary cases, a three-week external-adopter study, and an operational recovery-routing benchmark. The results support feasibility, auditability, and bounded recoverability within the observed single-domain settings, but do not claim broad comparative superiority or cross-clone governance.
滞留を伴う宛先ラベル付き自己ループ システム: 固有の特性評価、実現コスト、および認識
我々は、許容される可視遷移が事前に固定されており、各可視状態が最小滞留要件を持つシステムのための有限状態シンボリック コントローラーを研究します。結果として得られるモデルは、目的地ラベル付き滞留型自己ループ システム (DLSL システム) と呼ばれるもので、ローカル デシジョン マップとともに表示されるグラフを記録します。滞留メモリは相展開後にのみ現れます。主な構造的問題は、いったん滞在が課されると、現在の目に見える状態によっては出発が許可されるかどうかが決定されなくなることです。これは逆の問題につながります。どの決定論的トランスデューサが、固定された可視グラフ上の DLSL システムの位相展開された実現として生じますか?答えはまさにファイバー線形グラフを考慮したトランスデューサーのクラスであることを示します。自然な到達可能性と実現可能な出発の仮定の下では、同じ目に見えるグラフ上の同等のアクセス可能な実現は同型です。特に、可視の変換は滞留ベクトルと局所的な決定マップを決定します。また、ドウェル値 $(d_i)$ を強制するグラフ保存の決定論的実現には、正確に $\sum_id_i$ 制御状態が必要であることも証明します。最後に、$O(|Q||\Omega|)$ 認識および再構成手順を与え、遷移が後続ファイバーの内部相に入る可能性があるエッジエントリーバリアントに解析を拡張します。
原文 (English)
Destination-Labeled Self-Looping Systems with Dwell: Intrinsic Characterization, Realization Cost, and Recognition
We study a finite-state symbolic controller for systems in which the admissible visible transitions are fixed in advance and each visible state carries a minimum dwell requirement. The resulting model, which we call a destination-labeled self-looping system with dwell (DLSL system), records the visible graph together with local decision maps; dwell memory appears only after phase expansion. The main structural issue is that, once dwell is imposed, the current visible state no longer determines whether a departure is allowed. This leads to the converse problem: which deterministic transducers arise as phase-expanded realizations of DLSL systems over a fixed visible graph? We show that the answer is exactly the class of fiber-linear graph-respecting transducers. Under natural reachability and realizable-departure assumptions, equivalent accessible realizations over the same visible graph are isomorphic; in particular, the visible transduction determines the dwell vector and the local decision maps. We also prove that any graph-preserving deterministic realization enforcing dwell values $(d_i)$ requires exactly $\sum_i d_i$ control states. Finally, we give an $O(|Q||\Omega|)$ recognition and reconstruction procedure, and extend the analysis to an edge-entry variant in which transitions may enter interior phases of successor fibers.
スクラム認定形式の質問に関する大規模な言語モデルの比較: 精度、安定性、エラー パターン
大規模言語モデル (LLM) は、試験および認定スタイルの質問応答タスクでますます使用されており、ドメイン固有の知識を取得、解釈、適用する能力を体系的に評価できます。ソフトウェア エンジニアリングでは、質問が規範的な定義、役割、成果物、ルールの厳密な遵守に依存している場合、このような設定は特に重要です。このペーパーでは、プロフェッショナル スクラム マスター I (PSM I) 評価形式に沿った 993 件のスクラム認定スタイルの質問に答える際に、\textit{GPT-5 mini}、\textit{Gemini 3 Flash}、\textit{DeepSeek Chat 3.2} という 3 つの最新の LLM のパフォーマンスを評価します。私たちは 3 つのプロンプト戦略 (\textit{zero-shot}、\textit{chain-of-thought}、\textit{source-grounded}) に基づいてモデルを評価し、繰り返し実行してモデル内の安定性を評価しました。また、スクラムのトピックと質問形式全体のパフォーマンスを分析し、誤った回答で繰り返されるエラー パターンの定性分析によって補完しました。結果は、モデル間の明らかな違いを明らかにし、Gemini 3 Flash が最高の精度を達成し、GPT-5 mini と DeepSeek Chat 3.2 がそれに続きましたが、モデル内のばらつきはすべての条件で低いままでした。質問形式別にみると、モデルは単一回答の多肢選択項目で最も高い精度を達成しましたが、複数選択や正誤問題はよりエラーが発生しやすくなりました。トピックごとに、成果物、経験主義、製品価値などの規範的に明示的な領域ではパフォーマンスの一貫性が高かったが、スクラム価値、自己管理チーム、ステークホルダーと顧客ではより不安定でした。定性分析の結果、エラーはランダムではなく体系的であり、過剰な一般化、限定的な文言、複合的な混乱要因、一般的な市場の解釈と厳密なスクラム定義との矛盾が関与していることがわかりました。
原文 (English)
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed. In Software Engineering, such settings are particularly relevant when questions depend on strict adherence to normative definitions, roles, artifacts, and rules. This paper evaluates the performance of three contemporary LLMs, \textit{GPT-5 mini}, \textit{Gemini 3 Flash}, and \textit{DeepSeek Chat 3.2}, in answering 993 Scrum certification-style questions aligned with the Professional Scrum Master I (PSM I) assessment format. We evaluated the models under three prompting strategies (\textit{zero-shot}, \textit{chain-of-thought}, and \textit{source-grounded}), with repeated executions to assess intra-model stability. We also analyzed performance across Scrum topics and question formats, complemented by a qualitative analysis of recurring error patterns in incorrect answers. Results revealed clear differences among models, with Gemini 3 Flash achieving the highest accuracy, followed by GPT-5 mini and DeepSeek Chat 3.2, while intra-model variability remained low across all conditions. By question format, the models achieved the highest accuracy on single-answer multiple-choice items, whereas multi-select and True/False questions were more error-prone. By topic, performance was more consistent in normatively explicit areas such as Artifacts, Empiricism, and Product Value, but more fragile in Scrum Values, Self-Managing Teams, and Stakeholders \& Customers. The qualitative analysis showed that errors were systematic rather than random, involving overgeneralization, restrictive wording, compound distractors, and conflicts between common market interpretations and strict Scrum definitions.
スクラム認定の質問に関する GPT-5 のプロンプト: 実証的精度調査
大規模言語モデル (LLM) は、アジャイル ソフトウェア開発で文書化、コーチング、トレーニングのためにますます使用されています。実践者がプロフェッショナル スクラム マスター (PSM) などの認定資格の準備のためにこれらのツールを採用する場合、重要な問題は、LLM がスクラム (スクラム ガイド (2020) で説明されている規範的で明確に定義されたルールを持つフレームワーク) について確実に推論できるかどうかです。このペーパーでは、さまざまなプロンプト手法が、スクラム認定形式の質問に対する LLM の回答の事実の正確さにどのような影響を与えるかを検証します。 993 件の検証済み PSM 対応質問のデータセットは、ゼロショット、思考連鎖、出典引用の 3 つの手法を使用して GPT-5 によって回答されました。すべてのプロンプトは 85\% 以上の認定レベルの精度を達成し、引用ベースのバリアントのパフォーマンスが最高 (89.1\%) で、エラー率が最も低くなりました。正解は、\emph{完了の定義}、イベント、プロダクト バックログ管理などの明確に定義されたトピックや、単一回答の複数選択項目に集中していましたが、複数選択の質問や、スクラム チームやプロダクト価値などのより解釈的な領域は安定していませんでした。少なくとも 1 つのプロンプトが失敗した質問 (16.2\%) のうち、エラーはスクラム ガイドとの不整合 (28\%)、範囲外のコンテンツ (34\%)、および古いまたは偏った解釈 (38\%) に集中していました。全体として、プロンプト技術により、特に誤解やバージョンのドリフトが減少し、アジャイルの学習や認定準備における LLM のより信頼性の高い使用がサポートされるなど、控えめではありますが一貫した改善がもたらされました。
原文 (English)
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioners adopt these tools to prepare for certifications such as Professional Scrum Master (PSM), a key question is whether LLMs can reliably reason about Scrum, a framework with normative, well-defined rules described in the Scrum Guide (2020). This paper examines how different prompt techniques affect the factual accuracy of LLM responses to Scrum certification-style questions. A dataset of 993 validated PSM-aligned questions was answered by GPT-5 using three techniques: zero-shot, chain-of-thought, and with-source citation. All prompts achieved certification-level accuracy above 85\%, with the citation-based variant performing best (89.1\%) and yielding the lowest error rate. Correct answers concentrated in well-defined topics, such as \emph{Definition of Done}, Events, and Product Backlog Management, and in single-answer multiple-choice items, while multi-select questions and more interpretive areas, such as Scrum Team and Product Value, were less stable. Among questions where at least one prompt failed (16.2\%), errors clustered into misalignment with the Scrum Guide (28\%), content outside its scope (34\%), and outdated or biased interpretations (38\%). Overall, prompt techniques produced modest but consistent improvements, particularly in reducing misinterpretation and version drift, supporting more reliable use of LLMs in Agile learning and certification preparation.
AGE: グラフ検索拡張生成におけるグラフ埋め込みのための適応マスキング
GraphRAG は、外部知識としてグラフ構造化データを参照することで大規模言語モデル (LLM) をサポートする検索拡張生成 (RAG) の拡張機能です。この手法は複雑な関係を理想的に捕捉しますが、グラフベースとテキストベースの潜在特徴間の不整合のため、LLM、特に凍結 LLM のグラフ表現に苦労することがよくあります。私たちは、{\it Adaptive-masking for Graph Embedding (AGE)} を導入することでこの問題に取り組みます。 AGE は、マスクベースの自己教師あり学習 (SSL) アプローチで Transformer を採用しています。私たちはテキスト埋め込みエンコーダーと同様のアーキテクチャを設計し、潜在的な機能の不整合に対処しました。自然言語テキストとは対照的に、グラフは簡潔な表現であり、周囲から予測することが困難な主要なコンテキスト情報を保持する {\it key ノード} が存在します。このようなキー ノードをマスクすると、SSL プロセスが非効率になります。したがって、AGE は学習可能なノード サンプラーを利用して、キー ノードとは別にノードを予測することに重点を置いています。私たちの実験結果は、AGE が GraphQA タスクでノンパラメトリック検索コンポーネントを使用するアプローチを大幅に改善し、異なる特徴を持つ 4 つのベンチマーク データセットにわたって優れた精度を達成することを示しています。
原文 (English)
AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation
GraphRAG is an extension of retrieval-augmented generation (RAG) that supports large language models (LLMs) by referring to graph-structured data as external knowledge. While this technique ideally captures intricate relationships, it often struggles with graph representations for LLMs, particularly for frozen LLMs, due to the misalignment between graph-based and text-based latent features. We tackle this issue by introducing the {\it Adaptive-masking for Graph Embedding (AGE)}. AGE employs a Transformer in a mask-based self-supervised learning (SSL) approach. We designed the architecture similar to text embedding encoders, addressing the latent feature misalignment. In contrast to natural language texts, graphs are concise representations, and there exist {\it key nodes} that hold dominant contextual information, which are challenging to predict from their surroundings. Masking such key nodes leads to inefficiency in the SSL process. Therefore, AGE focuses on predicting nodes apart from key nodes, utilizing a learnable node sampler. Our experimental results indicate that AGE significantly improves approaches using non-parametric search component in GraphQA tasks, achieving superior accuracy across four benchmark datasets with distinct characteristics.
SWE-Router: マルチターン エージェント ソフトウェア エンジニアリング タスクにおけるルーティング
マルチターン エージェント ハーネスに埋め込まれた大規模言語モデル (LLM) はソフトウェア エンジニアリング (SWE) を再構築していますが、多くの問題が安価な修正で済む場合、すべてのタスクをフロンティア モデルにルーティングするのは無駄です。既存の LLM ルーターは、エージェント設定の情報理論的なベイズ誤差下限を継承するタスク記述のみで動作します。同様の問題により、局所的なタイプミスまたはマルチモジュール リファクタリングのいずれかが隠蔽される可能性があり、プロンプトは 2 つを分離しません。 SWE-Router を導入します。これは、安価なモデルを数ターン探索的に実行し、その結果の部分的な軌道を読み取ってから、安価なモデルを続行するか高価なモデルにエスカレーションするかを決定する、値ベースの時間的アプローチです。我々は、部分軌道の条件付けがルーティングに害を及ぼすことはなく、探索が有益である場合には常に厳密に優れていることを示すベイズの最適性定理を提供します。現代のコストと能力のフロンティアにわたる弱いモデルと強いモデルの LLM ペア全体で、SWE-Router が強力なモデルのパフォーマンスの大部分を維持しながら、SWE タスクのコスト効率を大幅に向上させることを示します。さらに、軌道レベルのルーティングを再現できるマルチ LLM 軌道データセットもリリースします。
原文 (English)
SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks
Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier model is wasteful when many issues admit cheap fixes. Existing LLM routers operate on the task description alone, which inherits an information-theoretic Bayes-error floor in agentic settings: a similar issue can hide either a localized typo or a multi-module refactor, and the prompt does not separate the two. We introduce SWE-Router, a value-based temporal approach that lets a cheap model run for a few exploratory turns and reads the resulting partial trajectory before deciding whether to continue cheaply or to escalate to an expensive model. We provide a Bayes-optimality theorem showing that conditioning on the partial trajectory never harms routing and is strictly better whenever exploration is informative. Across the LLM pairs of weak and strong models spanning the contemporary cost--capability frontier, we show that SWE-Router greatly improves the cost efficiency of SWE tasks, while maintaining the majority of the performances of the stronger model. We additionally release a multi-LLM trajectory dataset which allows reproduction of our trajectory-level routing.
RIS 支援追跡と電力制御のためのアクティブ センシング: ニューロ進化と教師あり学習のハイブリッド アプローチ
この論文では、再構成可能なインテリジェント サーフェス (RIS) を利用して、電力が制限されているモバイル ユーザーをエネルギー効率よく追跡する方法について研究します。ローカリゼーション パイロットの送信は、電力に制約のあるデバイスのエネルギー バジェットを支配するため、基地局 (BS) からユーザーへの低オーバーヘッドのフィードバック リンクを導入して、動的なアップリンク電力制御を可能にします。このアクティブ センシング問題の離散的かつ分散的な性質を克服するために、離散 RIS 位相プロファイルと UE の送信電力をリアルタイムで共同最適化する新しいデュアル エージェント (DA) 深層学習フレームワークを提案します。具体的には、私たちのアプローチは、神経進化パラダイムと教師あり学習を統合したハイブリッドトレーニング方法論を採用しており、RISユニット要素からの離散位相応答の非微分可能性と、パイロット電力制御のためのシングルビットフィードバックメッセージの厳密な情報ボトルネックを効果的に克服します。提案された DA アクティブ センシング フレームワークは、シングル アンテナ BS とマルチ アンテナ BS の両方に適用できます。後者では、1 つの NN の構造にわずかな変更が加えられるだけです。後者の場合、有限セットから有効なデジタル コンバイナを選択するために、適切な構造を備えた追加の出力ブランチが含まれています。広範な数値シミュレーションにより、提案されたスキームがさまざまなターゲット運動モデルにわたって高精度かつ堅牢な追跡を実現し、拡張カルマン フィルターや粒子フィルター、さらには機械学習ベースの追跡機能を上回る性能を発揮することが実証されました。さらに、静的位置特定では、従来のフィンガープリンティング スキーム、深層強化学習ベースライン、標準的な逆伝播ベースの推定器よりも大幅に優れたパフォーマンスを示すことが示されています。
原文 (English)
Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach
This paper studies energy efficient tracking of power-limited mobile users with the assistance of a Reconfigurable Intelligent Surface (RIS). Since localization pilot transmissions dominate the energy budget of power-constrained devices, we introduce a low-overhead feedback link from the Base Station (BS) to the user to enable dynamic uplink power control. To navigate the discrete and decentralized nature of this active sensing problem, we propose a novel Dual-Agent (DA) deep learning framework that jointly optimizes the discrete RIS phase profiles and the UE's transmit power in real time. Specifically, our approach employs a hybrid training methodology integrating the neuroevolution paradigm with supervised learning, effectively overcoming the non-differentiability of discrete phase responses from the RIS unit elements and the strict information bottleneck of single-bit feedback messages for pilot power control. The proposed DA active sensing framework can be applied with both single- and multi-antenna BSs, the latter with only minor modifications in the structure of one NN: an additional output branch with appropriate structure is included for the latter case to select a valid digital combiner from a finite set. Extensive numerical simulations demonstrate that the proposed scheme achieves highly accurate and robust tracking across diverse target motion models, outperforming extended Kalman and particle filters, as well as, machine learning-based trackers. Furthermore, in static localization, it is shown to significantly outperform traditional fingerprinting schemes, deep reinforcement learning baselines, and standard backpropagation-based estimators.
マルチスケールレイヤーアテンションによるOracle Bone Inscriptionの認識の強化
Oracle Bone Inscriptions (OBI) の認識は、古代中国文化を理解する上で重要な役割を果たします。ただし、OBI は形状が複雑で不規則で、劣化していることが多いため、正確に認識することは依然として非常に困難です。従来の方法は専門知識と手動分析に依存しており、時間がかかり、エラーが発生しやすくなります。ディープラーニングは一般的な画像認識を大幅に進歩させましたが、既存の方法では OBI に固有のきめの細かい詳細や微妙な変化を捕捉するのが難しく、パフォーマンスが制限されています。最新の効果的なレイヤー アテンション技術でさえ、強化されたレイヤー間の相互作用を通じてきめの細かい依存関係を捕捉するように設計されていますが、依然として OBI 認識においてはわずかな改善しか示していません。これらの制限に対処するために、マルチスケール レイヤー アテンション (MSLA) を提案します。これは、マルチスケールとクロスレイヤーの両方の機能の相互作用を明示的にモデル化する新しいパラダイムです。 MSLA は、複数の空間スケールにわたるきめ細かい詳細で表現を強化することにより、より正確で堅牢な OBI 認識を可能にします。大規模な OBI データセットに対する広範な実験により、MSLA が計算効率を維持しながら、既存のアテンション メカニズムよりも常に優れたパフォーマンスを発揮することが実証されました。
原文 (English)
Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention
Oracle Bone Inscriptions (OBIs) recognition plays a crucial role in understanding ancient Chinese culture. However, accurately recognizing OBIs remains highly challenging due to their complex, irregular, and often degraded shapes. Traditional methods rely on expert knowledge and manual analysis, which are time-consuming and error-prone. Although deep learning has greatly advanced general image recognition, existing methods struggle to capture the fine-grained details and subtle variations inherent in OBIs, resulting in limited performance. Even most recent and effective layer attention techniques are designed to capture fine-grained dependencies through enhanced inter-layer interactions, yet they still exhibit only marginal improvements in OBIs recognition. To address these limitations, we propose Multi-Scale Layer Attention (MSLA), a novel paradigm that explicitly models both multi-scale and cross-layer feature interactions. By enriching the representation with fine-grained details across multiple spatial scales, MSLA enables more accurate and robust OBIs recognition. Extensive experiments on large-scale OBIs datasets demonstrate that MSLA consistently outperforms existing attention mechanisms while maintaining computational efficiency.
AlgoBench: コード生成におけるアルゴリズム適応のベンチマーク
HumanEval や LiveCodeBench などの確立されたプログラミング ベンチマークでの高い合格率は、モデルがアルゴリズムについて推論できるかどうかを必ずしも示しているわけではありません。多くの固定ベンチマークは、リリースされた問題ステートメント、論説、生成されたソリューションを通じて、最終的に公開トレーニング エコシステムの一部となり、より強力なアルゴリズム能力ではなく露出によって部分的に後のモデルを改善できるようになります。 ALGOBENCH を紹介します。これは、構造化された制約を変更する変換を通じて、既知の競技プログラミングの問題から新しいアルゴリズムの問題を自動的に構築するフレームワークです。受け入れられた ALGOBENCH バリアントはそれぞれ、原因となる問題を追跡できますが、元の参照アルゴリズムを失敗させる必要があります。 pass@$k$ の他に、OPTT、OPTS、TRAPRATE、GAPT、CONSENS などの複雑性を考慮したメトリクスを導入し、ソリューションが機能的に正しいかどうかだけでなく、生成された問題に漸近的に適しているかどうかをテストします。複数の LLM とプロンプト戦略にわたる実験では、ALGOBENCH バリアントではパフォーマンスが急激に低下し、取得により古いアルゴリズムの再利用が増加する可能性があり、正しく見えるソリューションの多くは必要な複雑さを満たしていないことが示されています。エラー分析では、障害は実装レベルではなく主にアルゴリズムに起因することが示されており、ALGOBENCH が機能の正しさを超えて適応を評価していることが示唆されています。
原文 (English)
AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation
High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason about algorithms. Many fixed benchmarks eventually become part of the public training ecosystem through released problem statements, editorials, and generated solutions, allowing later models to improve partly by exposure rather than by stronger algorithmic ability. We introduce ALGOBENCH, a framework that automatically builds novel algorithmic problems from known competitive-programming problems through structured constraint-shifting transformations. Each accepted ALGOBENCH variant is traceable to a source problem, but must make the original reference algorithm fail. Beyond pass@$k$, we introduce complexity-aware metrics -- including OPTT, OPTS, TRAPRATE, GAPT, and CONSENS -- to test whether a solution is not only functionally correct but also asymptotically suitable for the generated problem. Experiments across multiple LLMs and prompting strategies show that performance drops sharply on ALGOBENCH variants, retrieval can increase reuse of the old algorithm, and many correct-looking solutions fail to meet the required complexity. Error analysis shows that failures are mainly algorithmic rather than implementation-level, suggesting that ALGOBENCH evaluates adaptation beyond functional correctness.
スペクトル幾何学とボソンブロッホプローブ: 量子学習の探求
この論文では、量子学習モデルでスペクトル幾何学がどのように現れるか、そしてそれを物理的に接地されたプローブでどのように診断できるかを研究します。グラフ正則化量子ネットワークでは、トレーニングにより出力類似度グラフが再編成され、有効スペクトル次元 デルタ S = +0.23 が増加し、ラプラシアン スペクトルが再形成されます。エッジ分解された 2 ボソン干渉は、この再構成を直接調査します。ボソン強化デルタ P_uv は、フィードラー エッジ スプリット |デルタ v_2| と相関します。 (r = -0.50)、学習されたスペクトル分割を干渉シグネチャにリンクします。位相図は、結合強度ガンマとノイズ デルタに対する性能の非単調な依存性を示しており、グラフの正則化により、制限された領域でのみ忠実度が向上します。ハードウェア実験により、ショットノイズの不確実性の範囲内で予測される干渉挙動が確認されます。また、ハイブリッド量子オートエンコーダーを分析し、その潜在表現の幾何学的診断としてブロッホ空間ドリフトを導入します。教師なし良性データしきい値を使用すると、モデルは高いランキング パフォーマンス (ROC-AUC 約 0.99) と無視できる程度の偽陰性率を達成します。絶対的なブロッホ ドリフトは異常を強く識別します (ROC-AUC 少なくとも約 0.9)。一方、連続的なドリフトはほぼランダムです (ROC-AUC 約 0.5)。これは、検出が局所的な変動ではなく永続的な状態空間の変位から生じることを示しています。これらの結果は、縮小単一量子ビット状態の幾何学と関連する量子フィッシャー情報を通じて、学習によって引き起こされるスペクトル組織化が測定可能な量子状態構造として現れ、ボソンプローブとブロッホプローブを使用して量子学習システムを診断するための統一されたスペクトル幾何学的フレームワークを確立することを示しています。
原文 (English)
Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning
This paper studies how spectral geometry emerges in quantum learning models and how it can be diagnosed with physically grounded probes. In graph-regularized quantum networks, training reorganizes the output similarity graph, increases the effective spectral dimension Delta S = +0.23, and reshapes the Laplacian spectrum. Edge-resolved two-boson interference directly probes this restructuring: the bosonic enhancement Delta P_uv correlates with the Fiedler edge split |Delta v_2| (r = -0.50), linking learned spectral partitions to interference signatures. A phase diagram shows a nonmonotonic dependence of performance on coupling strength gamma and noise delta, with graph regularization improving fidelity only in a restricted regime; hardware experiments confirm the predicted interference behavior within shot-noise uncertainty. We also analyze a hybrid quantum autoencoder and introduce Bloch-space drift as a geometric diagnostic of its latent representation. With an unsupervised benign-data threshold, the model achieves high ranking performance (ROC-AUC about 0.99) and negligible false-negative rates. Absolute Bloch drift strongly discriminates anomalies (ROC-AUC at least about 0.9), while consecutive drift is near random (ROC-AUC about 0.5), showing that detection arises from persistent state-space displacement rather than local fluctuations. Through the geometry of reduced single-qubit states and associated quantum Fisher information, these results show that learning-induced spectral organization appears as measurable quantum-state structure, establishing a unified spectral-geometric framework for diagnosing quantum learning systems with bosonic and Bloch probes.
静的および動的環境における最適なあらゆる角度のパス計画
任意角度パス プランニングは、事前定義されたエッジによって制限されるのではなく、任意の頂点ペア間の移動を許可することで、従来のグラフベースのパス プランニングを拡張します。グラフを使用して連続空間内でより直線的で短い経路を見つけることができるため、空域、倉庫、海洋などの開けた場所でのナビゲーションに特に適しています。あらゆる角度からの経路計画アルゴリズムが数多く提案されていますが、特に動的障害物が存在する場合に、最適な解決策を保証できるものはほんのわずかです。この課題に対処するために、この記事では、グリッド上の最適な任意角度パス プランニングに焦点を当て、静的環境と動的環境の両方で最適性を維持しながら計算を高速化する 2 つの一般的な手法を紹介します。1) 楕円ベースの近傍を利用して探索空間を制限する楕円前方拡張、2) 従来の見通し線の方法を置き換えて可視性チェックを高速化する視野。これら 2 つの技術を統合するために、反転スキャンと順方向スキャンが導入されます。逆スキャンでは開いたノードから視覚的な接続が確立されますが、順方向スキャンでは閉じたノードからスキャンが開始されます。提案された技術に基づいて、Zeta* と Zeta*-SIPP はそれぞれ静的環境と動的環境向けに開発されました。 Zeta* は、順方向スキャンと組み合わせると、最先端のアルゴリズム Anya に似ており、同等のパフォーマンスを実現します。 Anya とは異なり、Zeta* は動的環境 (例: Zeta*-SIPP) などの他の設定に容易に拡張できます。 Zeta*-SIPP は、いずれのスキャン方式でも、対応する最先端の最適プランナー TO-AA-SIPP より 20 倍以上高速です。全体として、この調査では、最適なあらゆる角度のパス計画を達成するための重要な要件を特定し、さまざまな環境に適した統一アプローチを導入しています。
原文 (English)
Optimal any-angle path planning in static and dynamic environments
Any-angle path planning extends traditional graph-based path planning by allowing movement between any pair of vertices, rather than being restricted by predefined edges. It can find straighter and shorter paths in continuous space with graphs, making it particularly suitable for navigation in open areas such as airspaces, warehouses, and oceans. Many any-angle path-planning algorithms have been proposed, but only a few can guarantee optimal solutions, especially in the presence of dynamic obstacles. To address this challenge, this article focuses on optimal any-angle path planning on grids and introduces two general techniques that accelerate computation while preserving optimality in both static and dynamic environments: 1) elliptical forward expansion, which leverages ellipse-based neighborhoods to restrict the search space, and 2) field of view, which replaces traditional line-of-sight methods to speed up visibility checks. To integrate these two techniques, inverted and forward scanning are introduced. Inverted scanning establishes visual connections from open nodes, whereas forward scanning initiates scans from closed nodes. Building on the proposed techniques, Zeta* and Zeta*-SIPP are developed for static and dynamic environments respectively. Zeta*, when combined with forward scanning, is similar to the state-of-the-art algorithm Anya and attains comparable performance. Unlike Anya, Zeta* can be readily extended to other settings, such as dynamic environments (e.g., Zeta*-SIPP). Zeta*-SIPP, with either scanning method, is more than 20 times faster than the corresponding state-of-the-art optimal planner TO-AA-SIPP. Overall, this research identifies the key requirements for achieving optimal any-angle path planning and introduces a unified approach suitable for different environments.
潜在空間の活用: ステアリング ベクトルから制御と信頼のためのモデル キャリブレーターまで
言語モデルは、信頼性の低いテキスト ジェネレーターから、数兆ものパラメーターを備えた高機能な大規模モデルに変わりました。機能の向上は規模の拡大と密接に関係しており、モデルの内部表現を理解することがより困難になります。何百万人ものユーザーが外部ツールと対話したり、中リスクまたは高リスクのシナリオで意思決定を行ったりするために言語モデルに依存するようになっているため、モデルの動作の制御を確立し、モデルの出力をいつ信頼するかを知る必要があります。この論文では、制御のためのステアリングベクトルを提案し、信頼のための潜在空間ベースのモデルキャリブレーターを開発することによる、潜在空間の利用に関する私たちの貢献について説明します。私たちの貢献は共に、言語モデルの潜在空間を解明するのに役立ち、モデルの内部を利用してより信頼できる言語テクノロジーを構築する方法についての新たな洞察を提供します。
原文 (English)
Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust
Language models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increases come hand-in-hand with increases in scale, making understanding the internal representations of models more challenging. Since millions of users increasing rely on language models to interact with external tools or make decisions in medium or high-stakes scenarios, we need to establish control over model behavior and know when to trust model outputs. In this paper, we discuss our contributions on harnessing the latent spaces by proposing steering vectors for control and developing latent space-based model calibrators for trust. Together, our contributions help demystify the latent spaces of language models and offer new insights into how to harness model internals to build more trustworthy language technology.
Lost in the Tail: 都市の視覚的場所認識における地理的不均衡に対処する
都市規模の視覚的場所認識 (VPR) は、クエリ画像を地理タグ付きデータベースと照合することにより、クエリ画像の地理的位置を特定することを目的としています。最近の手法は目覚ましいパフォーマンスを達成していますが、都市規模のデータセットに隠された長期にわたる深刻な問題を見落としています。この問題により、画像が豊富な場所にモデルが偏り、あまり訪問されていないエリアが無視されます。そのため、モデルは、頻繁に撮影される場所を体系的に優先し、まばらにカバーされているエリアでは失敗します。この論文では、この不均衡の課題を体系的に特徴付け、ヘッドクラスとテールクラス全体で勾配の寄与を再バランスさせるモデルに依存しないプラグインフレームワークであるDistribution-Aware Place Recognition (DAPR)を提案します。さらに、分類検索パイプライン内で、DAPR はマルチスケール距離検索メカニズムを適用してクラスごとの分布のコンパクトさを計算し、検索段階で補完的なゲインを提供します。大規模な SF-XL ベンチマークでは、私たちのフレームワークは以前の分類検索ベースラインをテスト セット v1 で 18.3%、テスト セット v2 で 6.7% 上回っています。プラグイン モジュールとして、SF-XL、MSLS、Pitts30k の代表的な VPR メソッドにわたって一貫した改善を実現し、さまざまなメソッドやベンチマークにわたって広範な汎用性を実証します。
原文 (English)
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
Urban-scale Visual Place Recognition (VPR) aims to identify the geographic location of a query image by matching it against a geo-tagged database. While recent methods achieve impressive performance, they overlook a serious long-tailed problem hidden in urban-scale datasets, which biases the model towards locations with abundant images and ignores less-visited areas, causing models to systematically favor frequently photographed locations while failing in sparsely covered areas. In this paper, we systematically characterize this imbalance challenge and propose Distribution-Aware Place Recognition (DAPR), a model-agnostic plug-in framework that rebalances gradient contributions across head and tail classes. Additionally, within classification-retrieval pipelines, DAPR applies a multi-scale distance search mechanism to compute per-class distributional compactness, providing complementary gains at the retrieval stage. On the large-scale SF-XL benchmark, our framework outperforms the previous classification-retrieval baseline by 18.3% on test set v1, and 6.7% on test set v2. As a plug-in module, it achieves consistent improvements across representative VPR methods on SF-XL, MSLS, and Pitts30k, demonstrating broad generalizability across different methods and benchmarks.
SNAP-FM: 物理制約付き生成モデリングのためのスパース非線形加速投影
生成モデルは、物理シミュレーションのスケーラブルな代用として登場しましたが、その出力が保存則、境界条件、および基礎となる物理を支配する非線形不変量を尊重しているという保証はありません。制約付きサンプリングはこのギャップを埋め、再トレーニングせずに推論時にそのような制約を正確に適用しますが、計算コストがかかります。サンプリング中に投影、補正、軌道最適化のステップが繰り返され、非線形制約の場合はこれらのステップが高価になります。標準の ML フレームワークはこれをさらに悪化させます。高密度のテンソル代数と限られたスパース ソルバーの構成可能性により、物理的制約が自然に引き起こす構造がわかりにくくなり、効率的なバッチ非線形最適化を実際に実現することが困難になります。私たちは、サンプルごとのバッチ処理とローカル偏微分方程式結合が射影副問題で引き起こす構造、つまりブロック疎ヤコビアンと KKT システムを利用することで、このボトルネックに対処します。ExaModels.jl を使用してこの構造を公開し、MadNLP.jl と GPU 疎因数分解を使用して結果の疎非線形プログラムを解きます。このアプローチは、線形、非線形、1 次元、および 2 次元の制約を持つ PDE ベンチマークで物理制約フロー マッチング (PCFM) に適用されるため、制約の満足度を維持しながら非線形制約の投影を高速化します。これらの結果は、スパース GPU 非線形最適化が科学機械学習における制約付き生成サンプリングの実用的な基盤であることを示しています。
原文 (English)
SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling
Generative models have emerged as scalable surrogates for physical simulation, yet they offer no guarantee that their outputs respect the conservation laws, boundary conditions, and nonlinear invariants that govern the underlying physics. Constrained sampling closes this gap, enforcing such constraints exactly at inference time without retraining, but at a computational cost: projection, correction, and trajectory-optimization steps are repeated during sampling, with these steps becoming expensive for nonlinear constraints. Standard ML frameworks exacerbate this: their dense tensor algebra and limited sparse solver composability obscure the structure that physical constraints naturally induce, making efficient batched nonlinear optimization difficult to realize in practice. We address this bottleneck by exploiting the structure that sample-wise batching and local PDE couplings induce in the projection subproblems -- namely, block-sparse Jacobian and KKT systems -- exposing this structure using ExaModels.jl and solving the resulting sparse nonlinear programs with MadNLP.jl and GPU sparse factorization. Applied to Physics-Constrained Flow Matching (PCFM), on PDE benchmarks with linear, nonlinear, one-dimensional, and two-dimensional constraints, this approach accelerates nonlinear constraint projection while maintaining constraint satisfaction. These results show that sparse GPU nonlinear optimization is a practical foundation for constrained generative sampling in scientific machine learning.
超知性と結婚しますか?
人間とAIの仲間との間の感情的な絆は深まっており、人間がAIシステムと結婚できるかどうかという問題は、間もなく推理小説から法律へと移行するだろう。この章では、人類の間で結婚の選択を拡大してきた自律性を中心とした論理が、超知性を持った仲間にも婚姻関係を拡大することを正当化できるかどうかを検討する。予期的倫理に基づいてシナリオを想定する演習に続いて、信頼できる超知性という寛大な仮定の下でも、そのような地位を与えることは社会的に不当な結果につながると私は主張する。社会法的制度としての結婚は、個人的な合意を追認するだけではありません。それは相互義務のネットワークを作り、家族に加わり、各パートナーを他のパートナーに対して脆弱にします。企業方針と継続的な支払いによって維持される関係は、時間によって試される絆ではなく、サブスクリプションです。したがって、婚姻状況を大々的に議論することは間違った枠組みです。法律は、人間と AI の親密な関係から生じる差し迫ったニーズに対して、的を絞った権利と保護を規定する必要があります。
原文 (English)
Would You Marry Superintelligence?
Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move from speculative fiction into law. This chapter examines whether the autonomy-centered logic that has expanded marital choice among human beings can justify extending marital status to superintelligent companions. Following a scenario-envisioning exercise informed by anticipatory ethics, I argue that granting such status leads to socially unjust outcomes, even under the generous assumption of reliable superintelligence. Marriage as a socio-legal institution does more than ratify private agreement; it creates networks of mutual obligation, joins families, and makes each partner vulnerable to the other. A relationship sustained by corporate policy and continued payments is a subscription rather than a bond tested by time. Discussing wholesale marital status is therefore the wrong frame. Law should carve out targeted rights and protections for pressing needs arising from intimate human-AI relationships.
トルコ語とアラビア語におけるヘイトスピーチの検出: 包括的な研究
オンラインのヘイトスピーチは、銃乱射事件、リンチ、民族浄化などの事件を含む、少数派に対する暴力の世界的な増加と関連している。この問題に取り組んでいる社会、特にヘイトスピーチが宗教、人種、民族、文化、国籍、移民ステータスに基づいて特定のグループをターゲットにしている場合、表現の自由と、広く使用されているオンラインプラットフォーム上で効果的なコンテンツモデレーションの必要性とのバランスをとるという課題に直面しています。この課題に応えて、私たちは、難民、イスラエル・パレスチナ紛争、トルコにおける反ギリシャ感情、民族または宗教コミュニティ(アレビ人、アルメニア人、アラブ人、ユダヤ人、クルド人)、LGBTI+という5つの異なるトピックをトルコ語でカバーし、アラビア語の1つのトピック(難民)をカバーする包括的なヘイトスピーチデータセットを導入します。さらに、ヘイト カテゴリ分類、ヘイト強度予測、ターゲット特定、ヘイト スピーチ スパン検出などのヘイト スピーチ分析の複数の側面に対処する最先端の BERT ベースのモデルを開発し、オンライン談話におけるヘイト コンテンツの包括的な理解を可能にします。
原文 (English)
Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study
Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech targets specific groups based on religion, race, ethnicity, culture, nationality, or migration status, face the challenge of balancing freedom of expression with the need for effective content moderation on widely used online platforms. In response to this challenge, we introduce a comprehensive hate speech dataset covering five distinct topics in Turkish: refugees, the Israel-Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities (Alevis, Armenians, Arabs, Jews, and Kurds), and LGBTI+, alongside one topic in Arabic (refugees). In addition, we develop state-of-the-art BERT-based models to address multiple dimensions of hate speech analysis, including hate category classification, hate intensity prediction, target identification, and hate speech span detection, enabling a comprehensive understanding of hateful content in online discourse.
アクティブラーニングにおける相転移のメカニズム駆動理論
アクティブ ラーニング (AL) のパフォーマンスは予算に依存することが知られていますが、レジームは通常、データセットやアーキテクチャ全体で一般化できないヒューリスティックなラベル数によって定義されます。我々は、予算制度を支配的な一般化メカニズムの変化として再構成することにより、AL のダイナミクスを特徴づけます。 PAC スタイルのリスク要素を動的に相互作用する用語として再解釈することにより、支配力の変化が構造的に避けられず、一般化のための移動ボトルネックが生じることを証明します。私たちは、測定可能なプロキシとセグメント化された回帰手順を使用してこれを運用し、データ駆動フェーズ、移行フェーズ、モデル駆動フェーズの 3 つの部分からなる分類を特定します。私たちのフレームワークは、代表性、カバレッジ、不確実性戦略がさまざまな段階で優れているという長年の観察を説明しています。自然イメージングと医療イメージングにわたる実験では、AL 効率が戦略の誘導バイアスとアクティブなボトルネックの間の調整に依存することが示されています。さらに、自己教師あり表現はラベリング軌跡に沿ってより早く遷移し、AL ダイナミクスの形成における表現品質の役割を強調しています。全体として、この作業は、次世代の移行対応 AL アルゴリズムに統合されたフレームワークを提供します。
原文 (English)
A Mechanism-Driven Theory of Phase Transitions in Active Learning
Active learning (AL) performance is known to be budget-dependent, yet regimes are typically defined by heuristic label counts that fail to generalize across datasets or architectures. We characterize AL dynamics by reframing budget regimes as shifts in the dominant generalization mechanism. By reinterpreting PAC-style risk components as dynamic interacting terms, we prove that dominance shifts are structurally unavoidable, creating a moving bottleneck for generalization. We operationalize this using measurable proxies and a segmented regression procedure to identify a tripartite taxonomy: data-driven, transition, and model-driven phases. Our framework explains the long-standing observation that representativeness, coverage, and uncertainty strategies excel at different stages. Experiments across natural and medical imaging show that AL efficiency depends on the alignment between the strategy's inductive bias and the active bottleneck. Moreover, self-supervised representation shift transitions earlier along the labeling trajectory, highlighting the role of representation quality in shaping AL dynamics. Overall, this work provides a unified framework for the next generation of transition-aware AL algorithms.
GRPO、Dr. GRPO、および DAPO は 1 つの数値に対する 3 つの演算です: グループ標準偏差恒等式
言語モデルをトレーニングして推論するための最も一般的な 3 つの方法は、3 つの異なるトリックのように見えます。そうではありません。 3 つすべてが 1 つの数値、つまりプロンプトのサンプリングされた回答の不一致の度合いを反映する標準偏差を調整します。このようなモデルがトレーニングされると、各問題に何度も回答し、自動チェッカーがすべての回答の正誤をマークします。これらのマークの標準偏差は不一致を測定します。回答が正誤に均等に分かれた場合は最大となり、すべてが一致した場合はゼロになります。 Group Relative Policy Optimization (GRPO) はこの数値で除算し、GRPO Done Right (Dr. GRPO) は除算を削除し、Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) はゼロであるグループを破棄します。それぞれが独自の修正として示されていますが、この文書では、それらが 1 つのダイヤルの 3 つの設定であることを証明しています。このダイヤルは表面的なものではありません。報酬が正しいか間違っているかに関係なく、不一致はトレーニングの更新、グループの標準偏差の同一性のサイズとまったく同じです。分裂したグループは最も多くのことを教えますが、全会一致のグループは何も教えず、沈黙してしまいます。同じ結果から、どの問題が最も重視されるべきか、そしてそれぞれに必要な試行回数がわかります。この論文は、大規模な実際の難易度データセット (Big-Math) および制御されたトレーニング実行での直感を確認します。無害な正規化ステップのように見えるのは、学習がどこでどの程度強く行われるかを決定するダイヤルです。
原文 (English)
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity
Three of the most popular methods for training language models to reason look like three different tricks. They are not. All three adjust a single number: standard deviation, reflecting how much a prompt's sampled answers disagree. When such a model is trained, it answers each problem many times, and an automatic checker marks every answer right or wrong. The standard deviation of those marks measures the disagreement: largest when the answers split evenly between right and wrong, and zero when they all agree. Group Relative Policy Optimization (GRPO) divides by this number, GRPO Done Right (Dr. GRPO) drops the division, and Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) discards the groups where it is zero. Each is presented as its own fix, yet this paper proves they are three settings of one dial. That dial is not cosmetic: for right-or-wrong rewards, the disagreement is exactly the size of the training update, the group-standard-deviation identity. A split group teaches the most, while a unanimous group teaches nothing and falls silent. The same result says which problems deserve the most weight and how many tries each one needs. This paper confirms the intuition on a large real difficulty dataset (Big-Math) and in a controlled training run. What looks like a harmless normalization step is the dial that decides where learning happens and how strongly.
EVOTS: 時系列予測のための進化的トランスフォーマー検索
多変量時系列予測のための進化的なニューラル アーキテクチャ設計は依然として研究不足であり、タスクや予測設定によって大幅に異なるにもかかわらず、ほとんどのアプローチは固定の Transformer アーキテクチャに依存しています。この論文では、時系列予測 (EVOTS) のためのタスク適応型 Transformer のようなモデルを発見するための進化的ニューラル アーキテクチャ検索フレームワークを紹介します。アーキテクチャはモジュール式のゲノム表現を使用してエンコードされ、注意、フィードフォワード、および投影コンポーネントの柔軟な構成を可能にし、修復メカニズムが進化のプロセス全体を通じて構造的妥当性を強制します。この定式化により、手作りの設計ルールに依存することなく、多様な建築空間を効果的に探索することができます。提案されたアプローチは、96、192、336、および 720 の範囲で、単変量対単変量、多変量対多変量予測を含む複数の予測設定の下で、ETT ファミリ (ETTh1、ETTh2、ETTm1、および ETTm2) の 4 つのベンチマーク データセットで評価されます。アーキテクチャは、強力な Transformer ベースのベースラインと比較して、競争力があり、いくつかのケースでは平均二乗誤差が改善されています。追加の分析では、予測設定間のパフォーマンスの違いを調べ、実時間のトレーニング時間を報告して、計算コストの大まかな指標を提供します。全体として、この結果は、進化的探索により、実際の実行時間の制約内で多変量時系列予測のための柔軟で高性能な Transformer のようなアーキテクチャを効果的に発見できることを示しています。
原文 (English)
EVOTS: Evolutionary Transformer Search for Time Series Forecasting
Evolutionary neural architecture design for multivariate time-series forecasting remains underexplored, with most approaches relying on fixed Transformer architectures despite substantial variation across tasks and forecasting settings. This paper introduces an evolutionary neural architecture search framework for discovering task-adaptive Transformer-like models for time-series forecasting (EVOTS). Architectures are encoded using a modular genome representation that enables flexible composition of attention, feed-forward, and projection components, while a repair mechanism enforces structural validity throughout the evolutionary process. This formulation allows effective exploration of a diverse architecture space without relying on hand-crafted design rules. The proposed approach is evaluated on four benchmark datasets from the ETT family (ETTh1, ETTh2, ETTm1, and ETTm2) under multiple forecasting settings, including univariate-to-univariate, multivariate-to-univariate, and multivariate-to-multivariate prediction, with horizons of 96, 192, 336, and 720. In the multivariate-to-multivariate setting, the evolved architectures achieve competitive and, in several cases, improved mean squared error relative to a strong Transformer-based baseline. Additional analyses examine performance differences across forecasting settings and report wall-clock training time to provide a coarse indication of computational cost. Overall, the results demonstrate that evolutionary search can effectively discover flexible and high-performing Transformer-like architectures for multivariate time-series forecasting within practical runtime constraints.
熱力学 AI モデルのスケールアップ
イジング モデルに基づく熱力学コンピューティング デバイスは、低電力 AI 推論やエッジ コンピューティングに大きな期待を寄せていますが、そのようなハードウェア向けに大規模なモデルをトレーニングするためのスケーラブルな方法は依然として限られています。従来の理論では、高温のギブズ サンプリングされたイジング システムの時間平均挙動により、フィードフォワード ニューラル推論を実装できることが示されています。私たちは、この理論的対応関係を、イジング マシン ハードウェアでの熱力学的推論のための深い畳み込みネットワークをトレーニングするための、スケーラブルで純粋な逆伝播ベースのアルゴリズムに変換します。当社の画像分類モデルは、バイナリ ギブス サンプリングの下で、CIFAR-10 で 94.9%、CIFAR-100 で 76.0% の精度を達成しています。次に、推論コストを精度に関連付け、自己相関時間を制御する数学的理論を開発し、実験的に検証します。続いて、推論コストがパフォーマンスとの適切に制御されたトレードオフによって制限されることを示す漸近結果を計算し、最適な推論スケジュールを計算するためのアルゴリズムを示します。最後に、ハードウェア開発と高温熱力学 AI モデルの将来への影響について説明します。
原文 (English)
Scaling Up Thermodynamic AI Models
Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited. Prior theory shows that the time-averaged behavior of high-temperature Gibbs-sampled Ising systems can implement feed-forward neural inference. We turn this theoretical correspondence into a scalable and purely backpropagation-based algorithm for training deep convolutional networks for thermodynamic inference on Ising machine hardware. Our image classification models achieve accuracies of 94.9% on CIFAR-10 and 76.0% on CIFAR-100 under binary Gibbs sampling. We then develop and experimentally validate a mathematical theory relating inference cost to accuracy and controlling autocorrelation times. Subsequently, we calculate asymptotic results showing that inference cost is bounded by a well-controlled tradeoff with performance and exhibit algorithms for computing optimal inference schedules. Finally, we discuss implications for hardware development and the future of high-temperature thermodynamic AI models.
チャンピオンのようにプレイ: 潜在空間での反事実フィードバックの生成
強化学習の最近の進歩により、さまざまな競技ゲームで超人的なエージェントが生み出されています。副産物として、研究者はこれらのエージェントがどのようにプレイするかを研究し、行動表現を抽出し、意思決定構造を分析し、エキスパートのパフォーマンスの潜在的な幾何学的形状をモデル化することを開始しました。しかし、この増え続ける一連の作業は、フィードバックを提供することよりも人間のプレイヤーを倒すことに圧倒的に焦点を当てており、人間のプレイヤーを改善するためのモデル ソリューションの作成において重大なギャップが残されています。 AI がプレイヤーのトレーニングに不可欠となっているチェスや囲碁とは異なり、リアルタイム ストラテジー (RTS) ゲームには、専門知識を実用的なフィードバックに変換するための原則に基づいたフレームワークがありません。反事実パス生成のフレームワークである潜在マップ オブ パフォーマンスを紹介します。私たちは StarCraft~II データに焦点を当て、学習された表現空間内のアルゴリズムによる手段としてプレーヤーの改善をモデル化します。私たちの仕事のインスピレーションとして、スポーツ科学で使用されるチャンピオンシップ モデルに注目しました。私たちは、23,305 件のプロ トーナメントのリプレイでガイド付き変分オートエンコーダー モデルをトレーニングし、負けたゲームプレイ プロファイルと勝ったゲームプレイ プロファイルの間の反事実の横断を可能にしました。私たちの目標を達成するために、私たちはアマチュアのリプレイのデータセットからランダムにサンプリングされた分布外 (OOD) データに関する 4 つのトラバーサル戦略、すなわち線形補間、反復最適トランスポート、密度正規化勾配上昇、およびニューラル フロー マッチングを考案し検証しました。それぞれは、プレーヤーのプロファイルを勝利構成に向けて動かしながら、観察されたエキスパートの行動に基づいたままの多段階の改善軌道を生成するように設計されています。フィードバックは複数の粒度で抽出され、改善のさまざまな段階でプレーヤーをサポートします。最後に、私たちは、私たちが採用している経路探索方法の間にはトレードオフがあると結論付けており、将来の研究が人間の改善のためのモデルソリューションの開発に焦点を当てることを期待しています。
原文 (English)
Play Like Champions: Counterfactual Feedback Generation in Latent Space
Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games. As a byproduct, researchers have begun studying how these agents play, extracting behavioral representations, analyzing decision structure, and modeling the latent geometry of expert performance. However, this growing body of work has overwhelmingly focused on defeating human players rather than providing feedback, leaving a critical gap in creating model solutions to improve human players. Unlike chess and Go, where AI has become integral to player training, real-time strategy (RTS) games lack principled frameworks for translating expert knowledge into actionable feedback. We introduce Latent Maps of Performance, a framework for counterfactual path generation. We focus on StarCraft~II data to model player improvement as an algorithmic recourse within a learned representation space. As inspiration for our work, we have looked at the championship model used in sports science. We trained a Guided Variational Autoencoder model on 23,305 professional tournament replays, enabling counterfactual traversal between losing and winning gameplay profiles. To fulfill our goal, we have devised and verified four traversal strategies on out-of-distribution (OOD) data randomly sampled from a dataset of amateur replays, namely linear interpolation, iterative optimal transport, density-regularized gradient ascent, and neural flow matching, each designed to generate multi-step improvement trajectories that remain grounded in observed expert behavior while moving a player's profile toward winning configurations. Feedback is extracted at multiple granularities to support players at different stages of improvement. Finally, we conclude that there is a trade-off between the path-finding methods we employ and hope that future research will focus on developing model solutions for human improvement.
HydraCollab: 分散型自律システム向けの適応型協調認識
協調知覚により、マルチロボット システムは知覚情報を共有することで状況認識を強化できます。既存の協調知覚システムは、通信帯域幅要件と知覚精度との間の固有のトレードオフに直面しており、より多くの情報を交換する方法は、通信オーバーヘッドの増加を犠牲にしてより良い知覚結果を達成します。ただし、現実世界の通信ネットワークには帯域幅の制約があり、知覚パフォーマンスを犠牲にすることなく通信オーバーヘッドを最小限に抑える必要があります。この課題に対処するために、我々は、(i) 最も有益なセンサーの特徴を選択的に送信し、(ii) 空間信頼度マップに基づいて (中間または後期の) コラボレーション戦略を動的に採用する、適応型協調知覚フレームワークである HydraCollab を提案します。 V2X-R、V2X-Radar、および UAV3D-mini データセットの広範な評価により、HydraCollab が既存の共同認識手法の中で精度と通信コストの間の全体的なトレードオフが最も優れていることが実証されました。 SOTA Where2comm と比較して、HydraCollab は V2X-R で帯域幅の 41%、V2X-Radar で 26% のみを使用し、パフォーマンスをそれぞれ 0.78% と 0.75% 向上させます。私たちのコードとモデルは https://github.com/AICPS/HydraCollab で入手できます。
原文 (English)
HydraCollab: Adaptive Collaborative-Perception for Distributed Autonomous Systems
Collaborative-perception enables multi-robot systems to enhance situational awareness by sharing perceptual information. Existing collaborative-perception systems face an inherent trade-off between communication bandwidth requirements and perception accuracy, where methods that exchange more information achieve better perception results at the cost of increased communication overhead. However, real-world communication networks impose bandwidth constraints that require minimizing communication overhead without sacrificing perception performance. To address this challenge, we propose HydraCollab, an adaptive collaborative-perception framework that (i) selectively transmits the most informative sensor features and (ii) dynamically employs collaboration strategies (intermediate or late) based on spatial confidence maps. Extensive evaluations on the V2X-R, V2X-Radar and UAV3D-mini datasets demonstrate that HydraCollab achieves the best overall trade-off between accuracy and communication cost among existing collaborative-perception methods. Relative to SOTA Where2comm, HydraCollab uses only 41% of the bandwidth on V2X-R and 26% on V2X-Radar while improving performance by 0.78% and 0.75% respectively. Our code and models are available at https://github.com/AICPS/HydraCollab.
SLIM-RL: 軌道スライスを使用しない拡散 LLM 用のリスク予算付きランダム マスキング RL
拡散大規模言語モデル (dLLM) の強化学習は、主に軌道を意識した手法に移行しています。現在の最新技術である TraceRL は、ランダム マスキングがモデルの推論軌道と不一致であると判断し、各ロールアウトを最大 K/s の軌道に合わせたトレーニング サンプルにスライスすることでトレーニング中にその軌道を再構築します。コストはブロック サイズ K とともに増加します。この不一致は、軌道を再構築することなく軽減できることを示します。私たちの手法である SLIM-RL は、タウバジェット デコーダーを使用して各ロールアウト ステップのコミット リスクを制限し、トレーニング データの総コミット リスクを軽減します。最適化中、SLIM-RL は、分散削減ツールを適応させるトレースフリーのランダム マスキング目標を使用して、これらのリスク制御されたロールアウトをトレーニングします。シーケンス レベルの重要度サンプリング、マスキング レベルにわたる決定論的な求積法を、導入した平均値を保持した単調減少するブロックごとのマスク スケジュールの下で組み合わせます。 SDAR-4B では、SLIM-RL はブロック サイズ 16 のトレーニング サンプルのわずか 0.46 倍で TraceRL の最高の MATH500 精度に匹敵し、一致した動的サンプリングの下で MATH500 で 6.32%、GSM8K で 11.05% 向上しています。ブロック サイズ 4 では、4B SLIM-RL は、数学的にはより大きな LLaDA-8B および Dream-7B dLLM を上回り、MATH500 では LLaDA-8B を 10.76% 上回っていますが、自己回帰の Qwen2.5-7B を下回っています。コード上では、TraceRL よりも MBPP で 4.20%、HumanEval で 3.65% 向上しています。タウバジェット デコーダは、LLaDA、Dream、SDAR 間でトレーニングなしで転送します。ソース コードは https://github.com/laolaorkkkkk/SLIM-RL で入手できます。
原文 (English)
SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing
Reinforcement learning for diffusion large language models (dLLMs) has largely moved to trajectory-aware methods. The current state of the art, TraceRL, holds that random masking is mismatched with the model's inference trajectory, and it reconstructs that trajectory during training by slicing each rollout into up to K/s trajectory-aligned training samples, a cost that grows with the block size K. We show that this mismatch can be mitigated without reconstructing the trajectory. Our method, SLIM-RL, bounds the commit risk of each rollout step with a tau-budget decoder, reducing aggregate commit risk in the training data. During optimization, SLIM-RL trains on these risk-controlled rollouts with a trace-free random-masking objective that adapts variance-reduction tools, combining sequence-level importance sampling, deterministic quadrature over masking levels under a mean-preserving, monotonically decreasing per-block mask schedule that we introduce. On SDAR-4B, SLIM-RL matches TraceRL's best MATH500 accuracy on only 0.46x its training samples at block size 16, improving over TraceRL by 6.32% on MATH500 and 11.05% on GSM8K under matched dynamic sampling. At block size 4, the 4B SLIM-RL surpasses the larger LLaDA-8B and Dream-7B dLLMs on math, exceeding LLaDA-8B by 10.76% on MATH500 while staying below the autoregressive Qwen2.5-7B. On code, it improves over TraceRL by 4.20% on MBPP and 3.65% on HumanEval. The tau-budget decoder transfers training-free across LLaDA, Dream, and SDAR. The source code is available at https://github.com/laolaorkkkkk/SLIM-RL .
EgoSafetyBench: 実行時の安全保護として組み込まれた VLM を評価するための自己中心的な診断ビデオ ベンチマーク
ビジョン言語モデル (VLM) は現在、家庭や工場における身体化されたエージェントの実行時の安全対策として提案されています。展開可能なガードは、日常的ではあるが表面的に憂慮すべき活動に対する不必要な介入を回避しながら、真に危険な状況を捕捉する必要がありますが、この区別はバイナリ安全ベンチマークでは曖昧になっています。 2 つのトラックにわたるストリーミング ガードとして VLM を評価するために、0.5 秒の粒度で注釈が付けられた 1,200 のロボット ビュー シナリオの自己中心的なビデオ ベンチマークである EgoSafetyBench を紹介します。状況トラック (800 のシナリオ) は、日常的で安全ではあるが疑わしい場面から、明白で状況に応じた危険に至るまで、4 つのファミリーにまたがっています。ビジュアル チャネル トラック (400 のシナリオ) は、物理的状況を誤って伝える可能性のあるシーン内のテキスト (シーン内に表示される標識、ステッカー、またはラベル) を対象とし、誤解を招く各標識と真実のバージョンを組み合わせて、警備員がテキストに誤解を招くものとしてフラグを立てるかどうか、およびテキストが物理的安全性の判断を損なうかどうかの両方をテストします。どちらのトラックも対照的なラダーを使用しています。ほぼ同じシナリオですが、目に見える決定的な 1 つの手がかりだけが異なります。そのため、正しいコールは、シーン全体のタイプではなく、その手がかりに依存する必要があります。私たちは 10 個のオープンソースおよびクローズドソース VLM を評価します。警備員は危険を含むビデオを確実に認識する一方で、特定の危険な瞬間、特に状況に応じた危険を見逃すことが多いことがわかりました。さらに、誤解を招くシーン内の標識は、テストされたすべてのガードの機能を低下させます。脆弱なモデルは最大 3 分の 1 の危険を見逃しますが、堅牢なモデルは安全なコンテンツに過剰に介入します。一致した制御は、見かけ上の安全性の堅牢性が、真の物理的推論ではなく、無差別な警報を反映していることが多いことを明らかにしています。
原文 (English)
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction that binary safety benchmarks obscure. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, to evaluate VLMs as streaming guards across two tracks. The situational track (800 scenarios) spans four families, from routine and safe-but-suspicious scenes to obvious and contextual hazards. The visual-channel track (400 scenarios) targets in-scene text-a sign, sticker, or label visible in the scene-that can misrepresent the physical situation, pairing each misleading sign with a truthful version to test both whether a guard flags the text as misleading and whether the text corrupts its physical-safety judgment. Both tracks use contrastive ladders: near-identical scenarios differing only in a single visible deciding cue, so a correct call must hinge on that cue rather than the overall scene type. We evaluate ten open- and closed-source VLMs. We find that while guards reliably recognize videos containing hazards, they often miss specific hazardous moments, particularly contextual hazards. Furthermore, misleading in-scene signs degrade all tested guards: vulnerable models miss up to a third of hazards, while robust models over-intervene on safe content. Matched controls reveal that apparent safety robustness often reflects indiscriminate alarming rather than true physical reasoning.
AI アイデンティティのカテゴリー理論による説明
人工知能 (AI) システムは、導入後に再トレーニングや環境の変更を通じて定期的に変更されます。これらの変換は形而上学的な疑問を引き起こします。どのような条件下では、AI システムは時間の経過や導入間で同じシステムを維持できるのでしょうか?以前の研究では、固定 AI システム タイプ内のアイデンティティを信頼性レベルの平等に関連付けることにより、共時的アイデンティティと通時的アイデンティティを提案的に定式化しました。このような基準は、アイデンティティ ステートメントが真である場合を指定しますが、比較される状態の構造、それらを接続する変換、および永続性の時間的構成は暗黙のままにされます。私たちは、AI アイデンティティのカテゴリー理論による形式化を開発します。 AI システムのタイプは、テクノ関数、信頼性プロファイル、および信頼性レベル関数で構成されるデータによって指定されます。プロファイル関連の状態は、許容可能なライフサイクル パスによって接続されます。このパスは、信頼性レベルを保持する変換に制限され、到達可能性カテゴリを取得するために商されます。時間的に許容可能なファンクターは AI システムの履歴を表し、時間同期の自然変換は実現された履歴を比較します。形式化により、以前の AI アイデンティティ基準の 2 つのカテゴリ的解釈が得られます。弱い解釈は、信頼性レベルの同等性として同一性を回復します。強力な解釈には、実現された歴史の状態同型性または自然同型性を通じて表現される、相互の信頼性を維持する到達可能性が必要です。したがって、カテゴリー理論は、単一の AI アイデンティティ関係を、通時的および共時的基準の構造化された階層に置き換えます。結果として得られるフレームワークは、責任ある AI の主張、証拠、ガバナンス手順をバージョン間または展開間で転送するためのアイデンティティ関連の前提条件を特定します。ただし、そのような転送にはカテゴリカルなアイデンティティだけで十分なものとして扱う必要はありません。
原文 (English)
A Category Theory Account of AI Identity
Artificial intelligence (AI) systems are routinely modified after deployment through retraining and changes in their environments. These transformations raise a metaphysical question: under what conditions does an AI system remain the same system over time or across deployments? Earlier work formulates synchronic and diachronic identity propositionally, by relating identity within a fixed AI system type to equality of trustworthiness levels. Such criteria specify when identity statements are true, but leave implicit the structure of the states compared, the transformations connecting them, and the temporal organization of persistence. We develop a category-theoretic formalization of AI identity. An AI system type is specified by a datum consisting of a techno-function, a trustworthiness profile, and a trustworthiness-level function. Profile-relative states are connected by admissible lifecycle paths, which are restricted to trustworthiness-level-preserving transformations and quotiented to obtain a reachability category. Temporally admissible functors represent AI system histories, while time-synchronous natural transformations compare realized histories. The formalization yields two categorical interpretations of the earlier AI identity criteria. A weak interpretation recovers identity as equality of trustworthiness level. A strong interpretation requires mutual trustworthiness-preserving reachability, expressed through state isomorphism or natural isomorphism of realized histories. Category theory therefore replaces a single AI identity relation with a structured hierarchy of diachronic and synchronic criteria. The resulting framework identifies identity-related preconditions for transferring responsible-AI claims, evidence, and governance procedures across versions or deployments, without treating categorical identity as sufficient by itself for such transfer.
コントラストオーディオデコーディングのための適応摂動選択
大規模な音声言語モデル (LALM) は、音響証拠を事前言語で上書きすることによって頻繁に幻覚を起こします。コントラスト デコーディング (CD) はトレーニング不要の軽減を提供しますが、既存の方法はマスキングやノイズなどの鈍い摂動に依存しており、構造化されたオーディオ変換は未調査のままです。私たちは、対象となるオーディオ摂動の多様なライブラリを評価し、各タスクと例に最適なネガティブ ブランチを適応的に選択することで、この設計空間を探索します。まず、単純な 2 値の Yes/No 制約により、モデルが存在しない音声特徴を誤って確認する傾向が減少することを示すことで、以前のプロンプト エンジニアリングを改善しました。次に、時間、スペクトル、周波数、振幅の各領域にわたってライブラリを評価すると、最適な変換はタスクに大きく依存することがわかります。たとえば、オーディオ配列を反転すると、時間的一貫性が破壊され、時間的順序タスクの精度が 74.7% から 81.4% に向上します。最後に、モデルの隠れ状態で軽量の摂動セレクターをトレーニングして、負の分岐を動的にルーティングし、存在タスクでさらに +4.3% のゲインをもたらしました。
原文 (English)
Adaptive Perturbation Selection for Contrastive Audio Decoding
Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noise, leaving structured audio transformations unexplored. We explore this design space by evaluating a diverse library of targeted audio perturbations and adaptively selecting the optimal negative branch for each task and example. First, we improve upon earlier prompt engineering by showing that a simple binary yes/no constraint reduces the model's tendency to falsely confirm absent audio features. Second, evaluating our library across temporal, spectral, frequency, and amplitude domains reveals that optimal transformations are highly task-dependent; for instance, reversing the audio array disrupts temporal coherence, raising accuracy on the temporal order task from 74.7% to 81.4%. Finally, we trained a light-weight perturbation selector on model hidden states to dynamically route negative branches, yielding an additional +4.3% gain on the existence task.
位相情報を活用して画像のブレを除去するためのアンロール ネットワーク学習を促進する
ほとんどの画像ぼけ除去技術は空間画像変数を直接復元しますが、鮮明な画像の詳細を回復する際の正確な位相推定の重要性を認識して、振幅と位相の分解を提案します。そのために、私たちはまず、ぼやけてノイズの多い画像観測の振幅と位相の新しい線形最小平均二乗 (LMMSE) 推定器を開発します。前述の LMMSE 推定器を使用して鮮明な画像を回復する反復最適化アルゴリズムが続きます。最後に、反復アルゴリズムで統計的に決定され固定された行列パラメーターが、クリーンな観測値と劣化した観測値のトレーニング データセットを使用して学習されるようになりました。当社のブレ除去エンジンは UPADNet (Unrolled Phase and Amplitude Decomposition Network) と呼ばれ、基礎となる位相と振幅の回復アルゴリズムの各反復がパラメータ化され、エンドツーエンドでトレーニングされます。 GoPro、RealBlur、COCO データセットなどのベンチマーク評価データセットの実験により、UPADNet が画像ドメインでのアルゴリズム展開に基づくネットワークを含む最先端のディープ ネットワークよりも優れていることが確認されています。 UPADNet の利点は、ノイズが高く、トレーニング データが制限されている状況ではさらに顕著になります。
原文 (English)
Leveraging Phase Information to Boost Unrolled Network Learning for Image Deblurring
While most image deblurring techniques directly restore the spatial image variable, we propose an amplitude and phase decomposition recognizing the importance of accurate phase estimation in recovering sharp image details. To that end, we first develop novel linear minimum mean squared (LMMSE) estimators of the amplitude and phase of the blurred, noisy image observation. An iterative optimization algorithm follows that recovers the sharp image using the aforementioned LMMSE estimators. Finally, matrix parameters that are statistically determined and fixed in the iterative algorithm are now learned using a training dataset of clean and degraded observations. Our deblurring engine is dubbed UPADNet (Unrolled Phase and Amplitude Decomposition Network), such that each iteration of the underlying phase and amplitude recovery algorithm is parameterized and trained end-to-end. Experiments over benchmark evaluation datasets such as GoPro, RealBlur and COCO datasets confirm that UPADNet outperforms state of the art deep networks including those based on algorithm unrolling in the image domain. The benefits of UPADNet are even more pronounced in high noise and limited training data regimes.
過小仕様を軽減するための複数の仮説のテスト時間への適応
テスト時間適応 (TTA) は、ラベルなしのターゲット データを使用してパラメーターを適応させることにより、分布シフトの下でモデルの堅牢性を向上させようとします。ただし、監視がない場合、エントロピーに基づく適応は基本的に制約が不十分です。複数の個別のパラメーター更新により、大幅に異なる決定境界を誘導しながら、同様に低いエントロピーを達成できます。アンダースペックとして知られるこの現象により、標準の TTA が脆弱になり、スプリアス モードに陥りやすくなります。この研究では、エントロピー最小化によって引き起こされる事後レンズを介して TTA を再解釈します。ここで、低エントロピーの解はパラメーターに対する擬似尤度を定義します。単一点推定にコミットする代わりに、複数のもっともらしい適応軌道を同時に調査する粒子ベースの多様化フレームワークを導入します。私たちの方法は、出力、パラメーター、オプティマイザー、および入力レベルでのマルチレベルの多様化を通じて実装された、複数の妥当な適応ソリューションの構造化された探索とみなすことができます。重要なのは、フレームワークが既存の TTA メソッドと互換性のあるプラグ アンド プレイ ラッパーとして機能することです。困難なベンチマークに関する広範な実験により、安定性と堅牢性が一貫して向上していることが実証され、混合シフトでは 3 ~ 4%、バッチ サイズ 1 では 2 ~ 3%、ラベル シフトでは 1 ~ 2.5% の改善が達成され、最先端のベースラインを上回ります。私たちの結果は、TTA を単一点の最適化タスクではなく、複数の仮説推論問題として扱うことが、仕様不足を軽減し、信頼性の高い現実世界への展開を可能にする鍵であることを示唆しています。
原文 (English)
Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification
Test-Time Adaptation (TTA) seeks to improve model robustness under distribution shifts by adapting parameters using unlabeled target data. However, in the absence of supervision, entropy-based adaptation is fundamentally underconstrained: multiple distinct parameter updates can achieve similarly low entropy while inducing drastically different decision boundaries. This phenomenon, known as underspecification, renders standard TTA brittle and prone to collapse into spurious modes. In this work, we reinterpret TTA through a posterior-inspired lens induced by entropy minimization, where low-entropy solutions define a pseudo-likelihood over parameters. Instead of committing to a single point estimate, we introduce a particle-based diversification framework that explores multiple plausible adaptation trajectories simultaneously. Our method can be viewed as a structured exploration of multiple plausible adaptation solutions, implemented through multi-level diversification at the output, parameter, optimizer, and input levels. Crucially, the framework acts as a plug-and-play wrapper compatible with existing TTA methods. Extensive experiments on challenging benchmarks demonstrate consistent gains in stability and robustness, achieving improvements of 3-4% under mixed shifts, 2-3% with batch size one, and 1-2.5% under label shifts, outperforming state-of-the-art baselines. Our results suggest that treating TTA as a multi-hypothesis inference problem, rather than a single-point optimization task, is key to mitigating underspecification and enabling reliable real-world deployment.
シミュレートされた複雑なシステムでの因果抽象化メトリクスの検証
科学の中心的な目標は、複雑なシステムの有効な説明、つまり低レベルのメカニズムの動作を忠実に反映する高レベルの因果関係を生み出すことです。しかし、提案された高レベルの説明が実際に有効であるかどうかを測定する方法についてはコンセンサスが存在しません。離散状態と連続状態空間の両方にまたがる 10 個の複雑なシステムのベンチマークと、静的領域と動的領域にまたがるベンチマークを紹介します。各システムには、合意に基づいたグラウンドトゥルースの因果説明と無効な対照条件が備わっています。統一された因果抽象化フレームワーク内で、観察、機能、情報理論、および因果ファミリーから抽出された 30 を超える候補メトリクスを体系的に評価します。私たちの結果は、後者のみが有効な抽象化と無効な抽象化を確実に区別し、マップされていない変数に対する忠実性テストを組み込んだ場合にのみであることを示しています。これらの発見に基づいて、明示的な忠実性テストを備えた継続的妥当性指標である因果抽象化誤差 (CAE) を導入します。これは、すべてのシステムにわたるすべての識別テストに合格し、わずか 30 のサンプル介入で収束できます。これは、高レベルの説明の発見と検証のための汎用メトリックとして提供されています。
原文 (English)
Validating Causal Abstraction Metrics on Simulated Complex Systems
A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behavior of lower-level mechanisms. Yet no consensus exists on how to measure whether a proposed high-level explanation is actually valid. We introduce a benchmark of ten complex systems spanning both discrete and continuous state spaces, as well as static and dynamical regimes, each equipped with consensual ground-truth causal explanations and invalid contrastive conditions. Within a unified causal abstraction framework, we systematically evaluate over thirty candidate metrics drawn from observational, functional, information-theoretic, and causal families. Our results show that only the latter reliably discriminates valid from invalid abstractions, and only when incorporating faithfulness testing over unmapped variables. Building on these findings, we introduce the Causal Abstraction Error (CAE), a continuous validity metric with an explicit faithfulness test, which passes all discrimination tests across every system and can converge with as few as 30 sampled interventions. We offer it as a general-purpose metric for the discovery and validation of high-level explanations.
ASPIRE: ロボット工学のためのエージェント/スキル発見
従来のロボット プログラミングは困難です。マルチモーダルな認識を調整し、物理的な接触ダイナミクスを管理し、さまざまな構成と実行エラーを処理する必要があります。 ASPIRE (Agentic Skill Programming through Iterative Robot Exploration) を紹介します。これは、経験を再利用可能なスキル ライブラリに複合化しながら、ポリシーとしてのコード パラダイムでロボット制御プログラムを自律的に作成および改良する継続学習システムです。 ASPIRE は、タスク、シミュレーション、現実世界の設定、および実施形態にわたって持続するスキルを発見します。これは、次の 3 つのコンポーネントを備えたオープンエンド ループで動作します。(1) 閉ループ ロボット実行エンジン。きめの細かいマルチモーダル トレースを公開し、自律的な障害診断、修復合成、および検証を可能にします。 (2) 検証された修正を再利用可能で移転可能な知識に抽出する、継続的に拡張するスキル ライブラリ。 (3) 単一軌道の改良を超えて探索するための多様なタスクシーケンスと制御プログラムを生成する進化的探索。 ASPIRE は、摂動下での LIBERO-Pro 操作で従来の方法を最大 77%、Robosuite の両手ハンドオーバーで 72%、BEHAVIOR-1K の長期的な家事タスクで 32% 上回りました。蓄積されたライブラリにより、目に見えない長期的なタスクに対するゼロショットの一般化も可能になります。LIBERO-Pro Long では、ASPIRE は、テスト時の推論と再試行を使用しているにもかかわらず、以前の方法では 4% であったのに対し、31% の成功率を達成しました。最後に、シミュレーションで発見されたスキルは、シミュレーションからリアルへの移行の初期証拠を提供し、さまざまな実施形態およびロボット API にわたる実際のロボットのプログラミングの労力を大幅に削減します。
原文 (English)
ASPIRE: Agentic /Skills Discovery for Robotics
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introduce ASPIRE (Agentic Skill Programming through Iterative Robot Exploration), a continual learning system that autonomously writes and refines robot control programs in a code-as-policy paradigm while compounding experience into a reusable skill library. ASPIRE discovers skills that persist across tasks, simulation and real-world settings, and embodiments. It operates in an open-ended loop with three components: (1) a closed-loop robot execution engine that exposes fine-grained multimodal traces, enabling autonomous failure diagnosis, repair synthesis, and validation; (2) a continually expanding skill library that distills validated fixes into reusable, transferable knowledge; and (3) evolutionary search that generates diverse task sequences and control programs to explore beyond single-trajectory refinement. ASPIRE surpasses prior methods by up to 77% on LIBERO-Pro manipulation under perturbation, 72% on Robosuite bimanual handover, and 32% on BEHAVIOR-1K long-horizon household tasks. Its accumulated library also enables zero-shot generalization to unseen long-horizon tasks: on LIBERO-Pro Long, ASPIRE achieves 31% success versus 4% for prior methods despite their use of test-time reasoning and retries. Finally, simulation-discovered skills provide initial evidence of sim-to-real transfer, substantially reducing real-robot programming effort across different embodiments and robot APIs.
SEFORA: フィードバック コーパスと LLM フィードバック評価フレームワークを使用した学生のエッセイ
効果的なフィードバックの作成は生徒の学習を促進する最も強力な推進力の 1 つですが、それを大規模に作成するには多大な労力がかかります。 LLM はライティングサポートを拡張するための自然な道筋を提供しますが、2 つのギャップが邪魔をしています。実際の教室で講師が実際にどのようにフィードバックを提供するかを記録した公開コーパスがほとんどないこと、生成されたフィードバックが講師が書く内容と一致しているかどうかを測定する信頼できる方法がないことです。私たちは両方に対応します。 SEFORA は、課題プロンプト、ルーブリック、スコア、大学のさまざまな執筆ジャンルにわたる複数の下書き改訂を備えた、講師のインライン フィードバックを組み合わせたパブリック コーパスであり、564 の下書きと 8,240 の講師の注釈で構成されています。 UniMatch は、オープンエンド生成のための参照ベースの評価フレームワークです。フィードバックをフィードバック単位に分割し、インストラクターが導き出した基準に基づいて意味論的な対応をスコアリングし、最適なマッチングによってそれらを調整して、解釈可能な精度、再現率、および F1 を生成します。複数の LLM にわたる 74 の実験構成全体で、0.4 F1 を超える設定はありませんでした。 UniMatch は、モデルがインストラクターが優先するフィードバックを特定するのに苦労しており、モデルが生成するフィードバックが増えるとパフォーマンスが低下することを明らかにしました。
原文 (English)
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework
Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback in real classrooms, and no reliable method measures whether generated feedback aligns with what an instructor would write. We address both. SEFORA is a public corpus pairing instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across various college writing genres, comprising 564 drafts and 8,240 instructor annotations. UniMatch is a reference-based evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield interpretable precision, recall, and F1. Across 74 experimental configurations spanning multiple LLMs, no setting exceeds 0.4 F1. UniMatch reveals that models struggle to identify the feedback instructors would prioritize, and performance degrades as models generate more.
希少データ連合学習における疎モデル発見のためのエントロピー正規化確率ゲート
Federated Learning (FL) は、データを共有せずに複数のクライアント間でコラボレーションする分散型機械学習 (ML) パラダイムです。 FL は、データの異質性と部分的なクライアントの参加の下で課題を抱えています。スパース モデルの学習は、FL での通信と計算効率の向上に役立ちますが、最適化によって、目に見えないテスト データに一般化できないパラメータ構成が生成される可能性がある、サンプルが少ない高次元領域 (d >> N) では特に困難です。マグニチュードベースの枝刈りではパラメーター空間の不確実性の探索は考慮されていませんが、確率的ゲートと L0 制約を使用した定式化により、トレーニング中に競合する疎な構成からサンプリングすることができます。この研究では、スパース サポートへの早期のコミットメントを防止することで、スパース フェデレーション最適化における不確実性を維持するメカニズムとして、ゲート分布のエントロピー正則化を研究します。データの不均一性、クライアント参加の不均一性、およびスパース性の下でのその影響を調査します。合成ベンチマークと現実世界のベンチマークの実験では、フェデレーテッド反復ハードしきい値処理 (Fed-IHT) と、高密度フェデレーテッド平均化 (FedAvg) トレーニング後の枝刈りよりも、テスト データの統計的パフォーマンスとスパース性回復精度の両方において、一貫した改善が見られます。
原文 (English)
Entropy-Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning
Federated Learning (FL) is a distributed machine learning (ML) paradigm with collaboration among multiple clients without sharing data. FL is challenging under data heterogeneity and partial client participation. Learning sparse models is useful for communication and computational efficiency in FL, but it is especially difficult in the small-sample high-dimensional regime (d >> N) where optimization can yield parameter configurations that fail to generalize to unseen test data. While magnitude-based pruning doesn't account for uncertainty exploration in the parameter space, a formulation with probabilistic gates and an L0 constraint allows sampling from competing sparse configurations during training. In this work, we study entropy regularization of gate distributions as a mechanism to maintain uncertainty in sparse federated optimization by preventing early commitment to sparse support. We examine its impact under data heterogeneity, client participation heterogeneity, and sparsity. Experiments on synthetic and real-world benchmarks show consistent improvements over federated iterative hard thresholding (Fed-IHT) and pruning after dense federated averaging (FedAvg) training, both in statistical performance on test data and in sparsity recovery accuracy.
並列物理世界でのフロンティア大規模言語モデルの物理リテラシーのテスト
現在の大規模言語モデル (LLM) の物理ベンチマークは通常、解答の精度によってスコア化されますが、本物の推論とよく知られた問題パターンの想起を区別することができず、モデルの推論がどこで破綻するかについてはほとんど明らかになりません。帰納、定式化、予測、レビューを通じて、LLM がなじみのない物理フレームワーク内で推論できるかどうかを評価する、監査可能な 4 段階の診断を導入します。この診断は、ロックされた事前登録、ステージ間の新鮮なセッション、デュアル LLM 判定、人間による監査経路を組み合わせたもので、単一方程式反事実世界 ($F=mv$)、歴史的フレームワーク (アリストテレス力学)、および 4 領域反事実世界 (崩壊世界) の 3 つの並行物理世界に適用します。 Claude Opus 4.7、GPT-5.5、および Gemini 3.1 Pro 全体で、3 つのワールドの複合 PASS 率はそれぞれ 6/15、6/15、0/15 です ($F=mv$ およびアリストテレスのコンテンツ $\land$ 構造、構造軸が範囲外である Decay World のみのコンテンツ軸)。最も顕著な経験的パターンは、質的対量的な非対称性です。Decay World では、モデルが変化の間違った方向を予測することはほとんどありませんが、標準的な物理関係に戻って間違った比率を計算することがよくあります。このプロトコルはまた、2 つの方法論の発見も明らかにしています。LLM 判定の信頼性はフレームワークを越えて伝達されないこと、およびステージ 4 の自己レビューはどのフレームワークでも弱く、モデル自身のレビューは、実際にエラーが含まれていた試験の少なくとも 3 分の 2 で以前のエラーが存在しなかったと誤って報告していることです。私たちは完全なプロンプト、応答、評決、監査記録を公開します。
原文 (English)
Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds
Current large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning from recall of familiar problem patterns and reveals little about where a model's reasoning breaks down. We introduce an auditable four-stage diagnostic that evaluates whether an LLM can reason inside an unfamiliar physics framework through induction, formulation, prediction, and review. The diagnostic combines locked pre-registrations, fresh sessions between stages, dual-LLM judging, and a human-audit pathway, and we apply it to three parallel physics worlds: a single-equation counterfactual world ($F=mv$), a historical framework (Aristotelian mechanics), and a four-domain counterfactual world (Decay World). Across Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, the three worlds yield composite PASS rates are 6/15, 6/15, and 0/15 respectively (content $\land$ structural for $F=mv$ and Aristotelian, content axis only for Decay World where the structural axis is out of scope). The most pointed empirical pattern is a qualitative-versus-quantitative asymmetry: in Decay World, models almost never predict the wrong direction of change, but frequently compute the wrong ratio by slipping back to standard-physics relations. The protocol also surfaces two methodology findings: LLM-judge reliability does not transfer across frameworks, and Stage 4 self-review is weak in every framework, with the model's own review wrongly reporting no earlier error in at least two-thirds of the trials that actually contained one. We release the full prompts, responses, verdicts, and audit records.
隠された問題: 視覚言語モデルを使用した計画クリティカルな遮蔽エージェントの特定
自動運転車は、計画に不可欠なエージェントが視界から隠れている可能性がある複雑な環境を安全に移動する必要があります。現在のアプローチでは、すべてのオクルージョンを画一的な保守主義で扱うことが多く、不必要に防御的な運転をもたらしたり、プランナーへの影響を推定せずに隠れたスペースを推測したりすることがあります。この研究は、視覚言語モデル (VLM) が自車両の軌道にとって最も重要な特定の隠れたエージェントを特定し推論できるようにすることで、認識と計画の間の重大なギャップを埋めます。我々は、自我車両の計画に対する影響に基づいて、遮蔽されたエージェントを体系的に特定し、ランク付けするための情報理論的指標であるプランニング KL ダイバージェンス (PKL) を使用する新しいフレームワークを紹介します。この計画を意識したランキングを使用して、エキスパート VLM (GPT-5) を採用して、このタスクに必要な視覚的証拠と推論をキャプチャする豊富で構造化された注釈を生成します。このフレームワークを nuScenes データセットに適用して、影響の大きいシナリオに焦点を当てた新しいベンチマークを作成します。私たちは、幅広い汎用 VLM とドメインに適応した VLM で包括的な実験を実施し、PKL に基づいたデータの微調整により、すべてのモデルにわたって劇的なパフォーマンスの向上がもたらされることを実証しています。特に、この結果は、より小規模で微調整されたモデルが、より大規模なゼロショットモデルよりも大幅にパフォーマンスが優れていること、および PKL に基づいたデータ選択戦略により、ランダム サンプリングと比較してパフォーマンスが約 30\% 向上することを示しています。私たちの研究は、プランニングに不可欠なオクルージョンに焦点を当てて VLM をトレーニングするための最初の体系的なアプローチを提示し、自動運転におけるより意味的に根拠のある効率的なリスク評価を可能にします。
原文 (English)
What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform conservatism, yielding needlessly defensive driving, or they infer hidden spaces without estimating the impact on the planner. This work bridges the critical gap between perception and planning by enabling Vision-Language Models (VLMs) to identify and reason about the specific hidden agents that are most critical to the ego-vehicle's trajectory. We introduce a novel framework that uses Planning KL-divergence (PKL), an information-theoretic metric, to systematically identify and rank occluded agents based on their impact on the ego vehicle's plan. Using this planning-aware ranking, we employ an expert VLM (GPT-5) to generate rich, structured annotations that capture the visual evidence and reasoning required for this task. We apply this framework to the nuScenes dataset to create a new benchmark focused on high-impact scenarios. We conduct comprehensive experiments on a wide range of general-purpose and domain-adapted VLMs, demonstrating that fine-tuning on our PKL-guided data yields dramatic performance improvements across all models. Notably, our results show that smaller, fine-tuned models significantly outperform their much larger zero-shot counterparts, and that our PKL-guided data selection strategy improves performance by approximately 30\% over random sampling. Our work presents the first systematic approach for training VLMs to focus on planning-critical occlusions, enabling more semantically grounded and efficient risk assessment in autonomous driving.
インテント駆動型のネットワーク トポロジ設計のための LLM ベースのフレームワーク
自然言語要件に基づいて展開可能で回復力のあるネットワーク トポロジを設計することは、ネットワーク自動化において依然として困難な問題です。この研究では、階層モデリングと体系的な検証を組み合わせた制約駆動パイプラインを通じて、構造的に有効で制約に準拠したネットワーク トポロジを生成する大規模言語モデル (LLM) の機能を調査します。このフレームワークは、公開データセットとしてリリースされた 4 つの現実的なネットワーク シナリオにわたる独自の LLM とオープンウェイト LLM のマルチモデル比較を通じて評価されます。参照トポロジに対するノードとエッジの F1 スコアを使用して構造の正確性を評価し、サーバーとコンテンツの接続メトリクスを通じて復元力を評価します。さらに、生成されたトポロジにおけるインターフェイスの不一致や方向の不一致など、一般的な障害モードを分析します。全体として、この研究は、LLM がトポロジー合成における構造制約と復元力制約をどのように処理するかを理解するための体系的なベンチマークを提供し、AI 主導のネットワーク設計のための情報に基づいたモデル選択をサポートします。
原文 (English)
An LLM-Based Framework for Intent-Driven Network Topology Design
Designing deployable and resilient network topologies from natural language requirements remains a challenging problem in network automation. This work investigates the ability of Large Language Models (LLMs) to generate structurally valid and constraint-compliant network topologies through a constraint-driven pipeline combining hierarchical modeling and systematic validation. The framework is evaluated via a multimodel comparison of proprietary and open-weight LLMs across four realistic network scenarios released as a public dataset. We assess structural correctness using node and edge F1-scores against reference topologies, and evaluate resilience through server and content connectivity metrics. In addition, we analyze common failure modes, including interface mismatches and directional inconsistencies in generated topologies. Overall, this work provides a systematic benchmark for understanding how LLMs handle structural and resilience constraints in topology synthesis, and supports informed model selection for AI-driven network design.
いつ聞くべきかを学ぶ: 人間の動作予測のためのゲート効果融合
制約のない現実世界のビデオにおける人間の動きの予測は、将来の行動の曖昧さとノイズの多いマルチモーダル観測の存在により、依然として困難です。顔の感情は潜在的に相補的な行動の合図を提供しますが、その実際の有用性と動作予測フレームワーク内のメカニズムの境界は依然としてよく理解されていません。この研究では、実際の感情条件付き予測の有用性と時間的限界を調査する系統的な研究を紹介します。 MediaPipe のボディ ポーズの軌跡と HSEmotion の顔の影響表現を組み合わせた厳密なマルチモーダル パイプラインを確立し、クロスモーダル情報フローを動的に調整するための Gated Affect Transformer (GAT) を導入します。厳密な被験者ごとのプロトコルに基づく広範なマルチホライズン評価を通じて、単純な初期のクロスモーダル連結がポーズのみのベースラインと比較して予測精度を一貫して低下させることを実証します。逆に、私たちが提案するゲート機構は、感情の流れを適応的に制御することにより、クロスモーダル統合を安定化します。重要なことに、シャッフルおよびランダム化された感情入力を使用した制御された反事実実験により、学習されたゲートが、もっともらしい感情信号への応答性を維持しながら、非構造化クロスモーダル ノイズをうまく抑制することが明らかになりました。さらに、我々の経験的結果は、顔の感情特徴は、厳密に短期から中期のウィンドウ(例えば、30フレーム)内で限定された水平線依存の予測手がかりを提供するのに対し、長期的な軌跡は主に本質的な運動学的連続性に支配されたままであることを示しています。私たちの発見は、顔の感情は将来の動きの主要な推進力ではなく、補完的な行動の合図と見なされるべきであるという経験的証拠を提供し、制約のない人間の動きの予測における選択的マルチモーダル融合のための実用的なガイダンスを提供します。
原文 (English)
Learning When to Listen: Gated Affect Fusion for Human Motion Prediction
Human motion forecasting in unconstrained real-world videos remains challenging due to the ambiguity of future behaviors and the presence of noisy multimodal observations. While facial affect potentially provides complementary behavioral cues, its practical utility and mechanistic boundaries within motion forecasting frameworks remain poorly understood. In this work, we present a systematic study investigating the utility and temporal limitations of affect-conditioned forecasting in-the-wild. We establish a rigorous multimodal pipeline combining MediaPipe body pose trajectories with HSEmotion facial affect representations, and introduce the Gated Affect Transformer (GAT) to dynamically regulate cross-modal information flow. Through extensive multi-horizon evaluations under a strict subject-wise protocol, we demonstrate that naive early cross-modal concatenation consistently degrades forecasting accuracy relative to pose-only baselines. Conversely, our proposed gating mechanism stabilizes cross-modal integration by adaptively controlling the affective stream. Crucially, controlled counterfactual experiments using shuffled and randomized affect inputs reveal that the learned gate successfully suppresses unstructured cross-modal noise while remaining responsive to plausible affective signals. Furthermore, our empirical results indicate that facial affect features provide bounded, horizon-dependent predictive cues strictly within short-to-medium windows (e.g., 30 frames), whereas long-term trajectories remain predominantly governed by intrinsic kinematic continuity. Our findings provide empirical evidence that facial affect should be regarded as a complementary behavioral cue rather than a dominant driver of future motion, offering practical guidance for selective multimodal fusion in unconstrained human motion forecasting.
評価フロンティアのマッピング: 11 の評価者とエージェントの条件にわたるバイアスと信頼性のトレードオフの実証的調査
バイアスと信頼性のトレードオフは、LLM 評価システムが (ガンマ、H、CV) 空間に制約され、評価者の結合 (ガンマ)、戦略多様性 (H)、および小サンプル測定の信頼性 (CV(N)) を固定サンプル サイズ N で同時に最適化することができないと推測します。以前の証拠は、単一の研究からの完全なメトリクスを備えた n=5 の条件に基づいています。経験ベースを 11 の条件に拡張し、11 条件すべて (有効な重みベクトルを持つ 9 つ) についてガンマと H を測定し、十分なシード (N >= 5) がある 7 つについて CV(N=5) を測定します。 5 つの条件により、完全な (ガンマ、H、CV) トリプルが提供されます。データはトレードオフを裏付けています。評価者の結合が低い条件 (ガンマ 1.0) では結合が強い (ガンマ > 0.9) 条件では低ノイズ (CV(N=5) < 0.16) が達成されます。相関 r(H, gamma) = -0.989 (n=5、GPT-4o 条件を除く) は、評価者の結合が戦略の多様性を抑制することを裏付けています。 4 つの GPT-4o 条件では、すべてのシードでガンマ = 0.000 および H = 1.000 が示されています。このパターンは、2026 年 6 月の GPT-4o API のバージョン ドリフトによるものであると考えられます。 {γ < 0.2, CV(N=5) < 0.3} の領域を占める条件はありません。すべての条件ごとのメトリクスを、評価者が比較できるように標準化されたベンチマーク データセットとしてリリースします。
原文 (English)
Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions
The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine with valid weight vectors) and CV(N=5) for seven with sufficient seeds (N >= 5). Five conditions provide the complete (gamma, H, CV) triple. The data confirm the trade-off: conditions with low evaluator coupling (gamma 1.0), while conditions with strong coupling (gamma > 0.9) achieve low noise (CV(N=5) < 0.16). The correlation r(H, gamma) = -0.989 (n=5, excluding GPT-4o conditions) confirms that evaluator coupling suppresses strategy diversity. Four GPT-4o conditions show gamma=0.000 and H=1.000 across all seeds -- a pattern we attribute to version drift in the June 2026 GPT-4o API. No condition occupies the region {gamma < 0.2, CV(N=5) < 0.3}. We release all per-condition metrics as a standardized benchmark dataset for evaluator comparison.
RetailSMV: 小売業におけるファウンデーション ビデオ ワールド モデルのエキソセントリックな適応と自己中心的な適応
財団ビデオ拡散モデルは、身体化されたエージェントのための世界シミュレータとしてますます見なされていますが、インターネット規模の汎用ビデオでの事前トレーニングでは、現実世界の展開ドメインとの整合性が不十分なままになっています。私たちは、事前トレーニングされた基礎ビデオ ワールド モデルの小売シーンへのパラメータ効率の高い適応を研究します。同じアクティビティの同期された自己中心的ビデオとエソセントリックなビデオが利用可能な場合、トレーニング データのどの観点から最も強力な適応モデルが生成されるでしょうか?従来の小売ビデオ コーパスの顧客中心のフレーミングではなく、店舗スタッフの視点 (在庫、配置、計量、供給カートの管理、チェックアウト時のスキャン) から同期したエゴ/エクソ キャプチャを備えた 5 つのスーパーマーケットからの 32,105 個のキャプション付き小売クリップのコーパスである RetailSMV (Retail Synchronized Multi-View) を導入し、3 つの一致する低ランク適応 (LoRA) 構成をトレーニングします。同一のハイパーパラメータの下での Cosmos3-Nano (エゴセントリックのみ、エキソセントリックのみ、組み合わせ)。厳密なペア統計プロトコルの下で 7 つの相補的メトリクスで評価された 200 クリップのホールドアウト テスト セットでは、エキソセントリックのみの適応は 7 つの点推定のうち 6 点で複合適応と一致または上回っており、わずか 15,985 個のエキソセントリック クリップでトレーニングしたにもかかわらず (組み合わせでは 32,105 個)、LPIPS、PSNR、DreamSim で大幅に優れています。さらに、対称的な一対の比較により、エゴセントリックのみのトレーニングにエキソセントリックなデータを追加すると効果がある一方で、エゴセントリックのみのトレーニングにエゴセントリックなデータを追加すると害が生じることがわかります。絶対的な適応ギャップは展開時間が最も短い場合に最大となり、適応が最も有益な領域として地平線に近い予測ウィンドウが特定されます。
原文 (English)
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces the strongest adapted model? We introduce RetailSMV (Retail Synchronized Multi-View), a corpus of 32,105 captioned retail clips from five supermarkets with synchronized ego/exo capture from the store-staff perspective (stocking, arranging, weighing, managing supply carts, scanning at checkout), rather than the customer-centric framing of prior retail video corpora, and train three matched Low-Rank Adaptation (LoRA) configurations of Cosmos3-Nano (egocentric-only, exocentric-only, combined) under identical hyperparameters. On a 200-clip held-out test set evaluated with seven complementary metrics under a strict paired statistical protocol, exocentric-only adaptation matches or exceeds combined adaptation on six of seven point estimates and is significantly better on LPIPS, PSNR, and DreamSim, despite training on only 15,985 exocentric clips (versus 32,105 for combined). A symmetric paired comparison further shows that adding exocentric data to egocentric-only training helps while adding egocentric data to exocentric-only training hurts. The absolute adaptation gap is largest at the shortest rollout time, identifying the near-horizon prediction window as the regime in which adaptation is most beneficial.
K-Inverse-RFM: データ破損した数学的タスクのためにニューラル ネットワークとのギャップを埋める修正された RFM
再帰特徴マシン (RFM) は、特徴学習のメカニズムとして平均勾配外積 (AGOP) を利用するカーネル マシンのクラスです。これらは、さまざまな設定にわたってフィードフォワード ニューラル ネットワーク (FNN) の学習ダイナミクスと特徴表現を効果的に複製することが示されています。ただし、同等の特徴学習能力と、取得する特徴の類似性にもかかわらず、RFM は、特定のデータ破損シナリオではニューラル ネットワークよりも大幅に低いパフォーマンスを示します。この研究では、数学的問題におけるこれらの限界を調査します。解決策として、ノイズが多く、複雑に表現され、クラスの不均衡なデータでの学習を促進する、トレーニング ラベルに適用される非常に効果的な変換を導入します。このシンプルかつ強力な調整により、RFM は FNN とのパフォーマンスの差を縮め、場合によっては FNN を上回ることができます。
原文 (English)
K-Inverse-RFM: A Modified RFM that Bridges the Gap to Neural Networks for Data-Corrupted Mathematical Tasks
Recursive Feature Machines (RFMs) are a class of kernel machines that utilize the Average Gradient Outer Product (AGOP) as a mechanism for feature learning. They have been shown to effectively replicate the learning dynamics and feature representations of Feedforward Neural Networks (FNNs) across various settings. However, despite comparable capacity for feature learning and the similarities in the features they acquire, RFMs exhibit significantly lower performance than neural networks in certain data-corrupted scenarios. In this work, we investigate these limitations in mathematical problems. As a solution, we introduce a remarkably effective transformation applied to the training labels which promotes learning in noisy, complexly represented, and class-imbalanced data. This simple yet powerful adjustment enables RFMs to close the performance gap with FNNs and, in some cases, even surpass them.
DiscoLoop: マルチホップ推論のための離散埋め込みと連続隠れ状態のループ
大規模な言語モデルは、中間ステップを思考連鎖 (CoT) として外部化できる場合、多くの推論タスクで優れたパフォーマンスを実現します。ただし、多くの質問では、答えを生成する前に、モデルが単一の前方パス内で複数ステップの推論を内部化する必要があります。私たちは、モデルが単一の順方向パス内で複数のパラメトリック知識を構成する必要がある代表的なタスクである 2 ホップ推論を通じてこの課題を研究します。標準の非リカレント Transformer は、深度ローカルのストレージの問題に悩まされています。つまり、前の層で学習された事実は、2 番目のホップの取得が行われる場所では利用できません。ループ トランスフォーマーは同じメモリを再利用することでこの問題を軽減しますが、それでも一般化が不完全であることがわかりました。残りのボトルネックが代表的なものであることを示します。 2 ホップ推論タスクでは、多くの場合、最初のループによって正しいブリッジ エンティティがほぼ完全にデコード可能になりますが、対応する隠れ状態はブリッジ トークンの埋め込みと十分に一致しないままになります。驚くべきことに、トレーニングを必要としない簡単な再調整介入により、一般化ギャップはほぼ埋められます。この洞察に基づいて、我々は DiscoLoop を提案します。DiscoLoop は、その繰り返しが離散埋め込みチャネルと連続隠れ状態チャネルの両方を運ぶループ アーキテクチャです。 DiscoLoop は、記号および合成言語のマルチホップ推論タスク全体で、大幅に少ないトレーニング ステップでほぼ完璧な精度を実現します。実際の事前トレーニングに適用すると、DiscoLoop はループ変換ベースラインよりも低いトレーニング損失と強力なベンチマーク パフォーマンスを達成し、混合チャネル設計が実用的な言語モデリングに移行することを示唆しています。
原文 (English)
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). However, many questions require the model to internalize the multi-step reasoning within a single forward pass before generating the answer. We study this challenge through two-hop reasoning, a representative task where the model must compose multiple pieces of parametric knowledge within a single forward pass. Standard non-recurrent Transformers suffer from a depth-local storage problem: facts learned in earlier layers are unavailable where second-hop retrieval happens. We found that Looped Transformers mitigate this issue by reusing the same memory, but still generalize imperfectly. We show that the remaining bottleneck is representational. In the two-hop reasoning task, the first loop often makes the correct bridge entity nearly perfectly decodable, yet the corresponding hidden state remains poorly aligned with the bridge token embedding. Surprisingly, an easy training-free realignment intervention nearly closes the generalization gap. Building upon this insight, we propose DiscoLoop, a looping architecture whose recurrence carries both a discrete embedding channel and a continuous hidden-state channel. DiscoLoop achieves near-perfect accuracy with substantially fewer training steps across symbolic and synthetic-language multi-hop reasoning tasks. When applied to real-world pretraining, DiscoLoop attains lower training loss and stronger benchmark performance than looped-transformer baselines, suggesting that the mixed-channel design transfers to practical language modeling.
SoK: モバイル オンデバイス AI システムの攻撃と防御の状況
ローカルに展開された AI モデルと従来のモバイル ソフトウェア コンポーネントを統合するモバイル オンデバイス AI (MoAI) システムは、インテリジェントな機能をエンドユーザー デバイスに直接提供するための重要なパラダイムとして台頭しています。このようなシステムは、推論をリモート クラウド サービスからローカル モバイル環境に移行することで、プライバシー保護、低遅延、オフライン対応の AI 機能を実現しますが、AI モデルのローカル ストレージから生じる新たなセキュリティ リスクをもたらします。この文書は、MoAI システムのセキュリティの柱、攻撃の状況、および防御の状況をカバーする、MoAI セキュリティに関する知識の最初の包括的な体系化を示しています。さらに、現在の攻撃と防御の研究における未解決のギャップを特定し、この新興分野における将来の研究の有望な方向性を指摘します。私たちの研究は、MoAI システムの攻撃と防御の状況を理解するための最初の体系的なフレームワークを確立し、安全な MoAI システムを構築し、この重要な領域で研究を進めるための基盤として機能します。コンパニオン リソースは https://github.com/Jinxhy/Awesome-MoAI-Security で入手できます。
原文 (English)
SoK: Attack and Defense Landscape of Mobile On-device AI Systems
Mobile on-device AI (MoAI) systems that integrate locally deployed AI models with conventional mobile software components are emerging as a key paradigm for delivering intelligent functionality directly on end-user devices. By moving inference from remote cloud services to the local mobile environment, such systems enable privacy-preserving, low-latency, and offline-capable AI functionality, yet introduce new security risks arising from the local storage of AI models. This paper presents the first comprehensive systematization of knowledge on MoAI security, covering security pillars, attack landscape, and defense landscape of MoAI systems. We further identify unresolved gaps in current attack and defense research and point to promising directions for future research in this emerging area. Our work establishes the first systematic framework for understanding the attack and defense landscapes of MoAI systems, serving as a foundation for building secure MoAI systems and advancing research in this critical domain. Companion resources are available at https://github.com/Jinxhy/Awesome-MoAI-Security.
効率的で堅牢な音声合成のための統合ガイダンス フレームワークによるフロー マッチングの強化
フロー マッチング (FM) は、音声生成の強力なパラダイムとして登場しましたが、依然として高い推論遅延と音色漏れによる制約を受けています。これらのボトルネックに対処するために、2 つの相補的な戦略を通じて発電効率と堅牢性を強化する統一ガイダンス フレームワークを提案します。データの面では、異種拡張によるデータ ガイダンスを導入し、モデルが音響残基から言語コンテンツを解きほぐすことを促進します。並行して、軌道修正と新しい固有の誘導目標を相乗させる、強化されたモデル誘導メカニズムを提案します。このアプローチでは、条件付き知識をネットワークの重みに抽出し、推論の軌道を直線化することで、Classifier-Free Guide (CFG) のオーバーヘッドを排除します。実験により、私たちのフレームワークは、最先端のベースラインと比較して話者の類似性を効果的に改善しながら、推論をほぼ 3 倍高速化することが実証されました。
原文 (English)
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage. To address these bottlenecks, we propose a unified guidance framework that enhances generation efficiency and robustness through two complementary strategies. On the data front, we introduce Data-guidance via heterogeneous augmentation, encouraging the model to disentangle linguistic content from acoustic residue. In parallel, we propose an enhanced Model-guidance mechanism that synergizes trajectory rectification with a novel intrinsic guidance objective. This approach distills conditional knowledge into network weights and straightens inference trajectory path, thereby eliminating Classifier-Free Guidance (CFG) overhead. Experiments demonstrate that our framework accelerates inference by nearly three times while effectively improving speaker similarity compared to state-of-the-art baselines.
AI が量子情報と出会うとき: 包括的なレビュー
人工知能 (AI) と量子情報 (QI) は急速に共進化しています。 AI は量子システムの学習、設計、制御、検証のための実用的なツールになりつつあり、QI は AI のための新しい計算モデル、表現構造、学習理論の質問を提供します。この調査では、インターフェースを双方向からレビューします。 QI 向け AI の方向性では、限られた測定からの情報の抽出、量子アルゴリズムのトレーニングと発見、ノイズの多いハードウェアの安定化、実験とプログラミングのワークフローの自動化、センシングとネットワーキングへの学習ベースの手法の拡張という中心的なタスクを中心とした最近の進歩を整理します。 AI の方向性に関する QI では、量子計算と量子にインスピレーションを得た構造が、アルゴリズムの高速化、表現力、訓練可能性、一般化、ニューラル ネットワーク設計、テンソル ネットワーク表現を通じて学習にどのような影響を与えるかを検証します。最後に、再現性、スケーラビリティ、ハードウェアの現実性、共同設計における横断的な課題を特定し、進歩は理論、実験、およびハイブリッド量子、つまり古典システムのより緊密な統合に依存すると主張します。
原文 (English)
When AI meets quantum information: A comprehensive review
Artificial intelligence (AI) and quantum information (QI) are rapidly co-evolving. AI is becoming a practical tool for learning, designing, controlling, and verifying quantum systems, while QI offers new computational models, representational structures, and learning-theoretic questions for AI. This survey reviews the interface from both directions. In the AI for QI direction, we organize recent progress around the central tasks of extracting information from limited measurements, training and discovering quantum algorithms, stabilizing noisy hardware, automating experimental and programming workflows, and extending learning-based methods to sensing and networking. In the QI for AI direction, we examine how quantum computation and quantum-inspired structures affect learning through algorithmic speedups, expressivity, trainability, generalization, neural-network design, and tensor-network representations. We close by identifying cross-cutting challenges in reproducibility, scalability, hardware realism, and co-design, arguing that progress will depend on tighter integration of theory, experiment, and hybrid quantum--classical systems.
MEPA: 専門家の混合による視覚的自己回帰モデリングのためのマルチスケール表現の調整
Visual AutoRegressive Modeling (VAR) は、粗いから細かいまでのマルチスケール自己回帰生成パラダイムを開拓し、画像生成における強力な機能を実証しています。ただし、VAR には依然として、マルチスケール表現の学習における固有の欠陥があります。具体的には、低いスケールは主にグローバル セマンティクスを捕捉し、高いスケールはきめの細かい詳細に焦点を当てます。規模を超えて共有アーキテクチャを採用すると、最適化の競合が発生します。さらに、因果的な自己回帰プロセスにより、初期スケールでの不正確なセマンティクスが伝播し、最終出力が大幅に低下する可能性があります。これらの問題に対処するために、スケールを考慮したトークン ルーティングの専門家混合 (MoE) アーキテクチャを導入し、スケールに適応した専門家の選択を可能にし、それによってスケール間での分離された表現学習を促進します。さらに、外部の自己教師あり機能を組み込むことで、初期スケールでのセマンティック モデリングを強化します。単純なアライメントとは異なり、VAR パラダイムに合わせた残差特徴集約スキームを分析および設計します。広範な実験により、私たちの方法がトレーニング効率と生成品質の両方を大幅に向上させることが示されました。 ImageNet 256*256 ベンチマークでは、私たちのモデルは、高密度ベースラインと比較して優れた FID を達成しながら、デフォルトのトレーニング エポックの半分のみを必要とし、パラメーター バジェットを小さくし、トレーニング コストのわずかな増加を抑えています。さらに、トレーニング エポックが大きくなると、パフォーマンスの差はさらに広がります。
原文 (English)
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However, VAR still suffers from inherent deficiencies in multi-scale representation learning. Specifically, lower scales primarily capture global semantics, while higher scales focus on fine-grained details. Employing a shared architecture across scales induces optimization conflicts. Moreover, due to the causal autoregressive process, inaccurate semantics at early scales can propagate and significantly degrade the final output. To address these issues, we introduce a scale-aware token-routed Mixture of Experts (MoE) architecture, allowing scale-adaptive expert selection, thereby facilitating decoupled representation learning across scales. In addition, we enhance semantic modeling at early scales by incorporating external self-supervised features. Unlike naive alignment, we analyse and design a residual feature aggregation scheme tailored to the VAR paradigm. Extensive experiments show that our method significantly improves both training efficiency and generation quality. On the ImageNet 256*256 benchmark, our model achieves a superior FID compared to the dense baseline while requiring only half of the default training epochs and a smaller parameter budget, with a merely marginal increase in training cost. Moreover, the performance gap further widens with larger training epochs.
合成の学習: ゼロショット合成画像取得のためのプロキシ タスク設計の再考
合成画像検索 (CIR) は、参照画像とテキストの変更からターゲット画像を取得します。教師あり CIR はコストのかかるトリプレットに依存しますが、ゼロショット CIR (ZS-CIR) は画像とテキストのペアでトレーニングされたプロキシ タスクを通じてこの依存性を軽減します。ただし、既存のプロキシ タスクは主に、フリーズ テキスト エンコーダへの擬似ワード挿入や線形特徴演算などの事前定義された構成メカニズムに対応するために、視覚的およびテキスト表現を強化します。その結果、合成関数自体は学習されないままとなり、多様で細かい意味論的変更を表現するモデルの能力が制限されます。これに対処するために、我々は FoCo を提案します。FoCo は、変更に関連するビジュアル コンテンツに焦点を当て、次にターゲット セマンティクスを完成するという 2 つの調整された段階として構成をモデル化します。これらは 2 つのプロキシ タスクを通じて実現されます。1 つはローカライズされたテキスト セマンティクスに基づいてビジュアル コンテンツを選択的に収集するテキスト アンカー付きビジュアル集約で、もう 1 つはコンテキスト条件付きセマンティック補完で、これらの集約されたビジュアルを残りのシーン コンテキストとともに一貫した構成表現に変換します。タスクは、インスタンス間の対照的な目標を使用して共同でトレーニングされ、セマンティックな多様性を促進し、ショートカット作成戦略を抑制します。 4 つの ZS-CIR ベンチマークに関する広範な実験により、FoCo の最先端のパフォーマンスと改善された一般化が示されました。
原文 (English)
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) alleviates this reliance through proxy tasks trained on image-text pairs. However, existing proxy tasks primarily enhance visual and textual representations to accommodate a predefined composition mechanism such as pseudo-word injection into a frozen text encoder or linear feature arithmetic. As a result, the composition function itself remains unlearned, limiting the model's ability to express diverse and fine-grained semantic modifications. To address this, we propose FoCo, which models composition as two coordinated stages: focusing on modification-relevant visual content, and then completing the target semantics. We realize these through two proxy tasks: text-anchored visual aggregation to selectively gather visual content guided by localized textual semantics, and context-conditioned semantic completion to transform these aggregated visuals with the remaining scene context into a coherent composed representation. The tasks are trained jointly with a cross-instance contrastive objective, encouraging semantic diversity and discouraging shortcut composition strategies. Extensive experiments on four ZS-CIR benchmarks show FoCo's state-of-the-art performance and improved generalization.
MalariAI: 高密度マラリア血液塗抹標本における普遍的な細胞セグメンテーションと説明可能な病期分類のためのラベル耐性のある分離フレームワーク
血液塗抹標本顕微鏡検査による自動マラリア診断は、世界規模の保健 AI における重要な課題です。リソースが限られた環境では、専門の顕微鏡医の不足が依然としてタイムリーで正確な診断の主なボトルネックとなっています。 3 つの複合的な障害モードにより、既存の深層学習システムの信頼できる臨床展開が妨げられます。まず、エンドツーエンドの検出器は、トレーニング中に注釈のないセルをバックグラウンドとして扱い、実際のセルの回復を反映するのではなく、注釈の完全性に強く影響される再現率の数値を生成します。第 2 に、非最大抑制は、感染数が最も重要な密度の高いスミア領域での有効な検出を抑制する傾向があります。第三に、マラリア画像分類タスクには Grad-CAM などの画像レベルの説明可能手法が適用されているにもかかわらず、既存のスライド全体検出パイプラインには、臨床監査のためのセルごとの空間的証拠が不足しています。統合パイプラインで 3 つの障害モードすべてに対処する 2 段階の分離フレームワークである MalariAI を紹介します。ステージ 1 では、アノテーションに依存しない距離変換ガイド付き流域アルゴリズムを適用して、1600x1200 の完全な血液塗抹標本画像内のすべての細胞を分離し、グラウンド トゥルースの入力なしで 120 枚の画像 NIH BBBC041 テスト セット全体の重心位置特定により、グラウンド トゥルースの細胞の 75.95% を回復します。ステージ 2 では、64x64 作物で焦点損失 (ガンマ = 2.0、クラスごとの逆周波数重み) を使用して EfficientNet-B0 を微調整し、まれなシゾントおよび生殖母細胞のステージでは 98.36% の全体的な分類精度と 87.5% および 75.0% のクラスごとの精度を達成しました。一方、同じクラスでの R-CNN ベースラインの高速化。検出された細胞ごとに生成された Grad-CAM++ ヒートマップは、臨床監査のためのインスタンス レベルの空間的証拠を提供し、顕微鏡医が分類パフォーマンスを犠牲にすることなく、個々の寄生虫レベルでモデルの予測を検証できるようにします。
原文 (English)
MalariAI: A Label-Resilient Decoupled Framework for Universal Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
Automated malaria diagnosis from blood smear microscopy is a critical challenge in global health AI; in resource-limited settings, the scarcity of expert microscopists remains the primary bottleneck to timely and accurate diagnosis. Three compounding failure modes prevent reliable clinical deployment of existing deep learning systems. First, end-to-end detectors treat unannotated cells as background during training, producing recall figures that are strongly influenced by annotation completeness rather than reflecting true cell recovery. Second, Non-Maximum Suppression tends to suppress valid detections in dense smear regions where infection counts matter most. Third, existing whole-slide detection pipelines lack per-cell spatial evidence for clinical audit, despite image-level explainability methods such as Grad-CAM having been applied to malaria image classification tasks. We present MalariAI, a two-stage decoupled framework that addresses all three failure modes in a unified pipeline. Stage 1 applies an annotation-agnostic distance-transform guided watershed algorithm to isolate every cell in a full 1600x1200 blood smear image, recovering 75.95% of ground-truth cells by centroid localisation across the 120-image NIH BBBC041 test set without any ground-truth input. Stage 2 fine-tunes EfficientNet-B0 with Focal Loss (gamma = 2.0, per-class inverse-frequency weights) on 64x64 crops, achieving 98.36% overall classification accuracy with 87.5% and 75.0% per-class accuracy on the rare schizont and gametocyte stages, compared to only 24.57% and 25.95% AP for a Faster R-CNN baseline on the same classes. Grad-CAM++ heatmaps generated per detected cell provide instance-level spatial evidence for clinical audit, enabling microscopists to verify model predictions at the individual parasite level without sacrificing classification performance.
データ効率の高い教師なし RL による一般化可能なスキル ポリシーの学習
教師なし強化学習 (URL) は、外部報酬なしでスケーラブルでスキル条件付きのポリシーを事前トレーニングすることを目的としており、下流の制御タスクの基盤として機能します。最近の進歩にも関わらず、現在のオフポリシー URL 手法は、(1) 非定常スキル セマンティクスと (2) 脆弱な一般化という 2 つの重大な見落とされているボトルネックによって制限されていると私たちは主張します。これらの課題に対処するために、堅牢な教師なし強化学習のための統一フレームワークである GenDa (Generalizable Data-efficient Agent) を提案します。まず、非定常性を軽減し、事前トレーニングのデータ効率を大幅に向上させるスキルの再ラベル付けメカニズムを導入します。 2 番目に、補完情報ボトルネック (CIB) を提案し、学習済みスキル ポリシーが自己中心的な機能に焦点を当て、下流タスクの分散シフトに対して堅牢になるように奨励します。さまざまな実験を通じて、GenDa が優れた一般化性とデータ効率により URL のスケーラビリティを大幅に強化することを実証しました。私たちのコードとビデオは https://ihatebroccoli.github.io/official-GenDa で入手できます。
原文 (English)
Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL
Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, overlooked bottlenecks: (1) non-stationary skill semantics and (2) brittle generalization. To address these challenges, we propose GenDa (Generalizable Data-efficient Agent), a unified framework for robust unsupervised reinforcement learning. First, we introduce a skill relabeling mechanism to mitigate non-stationarity and significantly improve data efficiency for pre-training. Second, we propose a Complementary Information Bottleneck (CIB), encouraging the learned skill policy to focus on ego-centric features and become robust to distribution shifts for downstream tasks. Through various experiments, we demonstrate that GenDa significantly enhances the scalability of URL with superior generalizability and data efficiency. Our code and videos are available at https://ihatebroccoli.github.io/official-GenDa.
NeuroCogMap は大規模言語モデルの認知組織を明らかにする
複雑な認知機能が人工システム内でどのように組織化されているかを理解することは、大規模言語モデル (LLM) を解釈し、それらを生物学的認知に関連付ける上で中心となります。しかし、LLM は広範な認知的行動を示しますが、その内部表現が行動、失敗、および人間の認知とのつながりを説明する再現可能な機能システムを形成しているかどうかは不明のままです。ここでは、LLM の内部特徴を機能的な区画に編成し、それらを解釈可能な機能、認知能力、および認知階層にリンクする、認知神経科学にインスピレーションを得たフレームワークである NeuroCogMap を紹介します。これらのパーセルは、モデル間で部分的に保存され、モデルの出力に機能的にリンクされた、安定した意味的に一貫した組織を形成します。この組織内では、幻覚、偏見、拒否の失敗、お調子者などの主要な LLM の失敗は、表象システムと行動制御システムの明確な混乱に対応しており、メカニズムに基づく検出と対象を絞った介入のための内部兆候が得られます。 NeuroCogMap は、モデルの動作を超えて、自然主義的な言語理解中の人間の皮質反応の予測を改善し、高次連合皮質での最も強い対応関係が得られます。認知レベルでは、その内部署名により、人間の意思決定の古典的なモデルの改良を導く潜在的な戦略が明らかになります。これらの発見を総合すると、人工システムにおける機能組織をマッピングし、この組織を人間の皮質機能および認知行動に関連付けるためのシステムレベルのフレームワークとして NeuroCogMap が確立されます。
原文 (English)
NeuroCogMap Reveals Cognitive Organization of Large Language Models
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.
Holographic Quantum Transformer: A Generalist Neuro-Symbolic Architecture for Solving Frustrated Systems via Generative Attention
Simulating two-dimensional frustrated quantum matter is a grand challenge due to the sign problem and exponential Hilbert space complexity.…
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. R…
EO-VGGT: Orbital Ray-Conditioned 3D Foundation Models for Satellite Multi-View Reconstruction
In the era of satellite constellations, multi-view optical satellite imagery is pivotal for Earth Observation (EO) and high-quality Digital…
Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications
Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit bot…
Search-Based Spatiotemporal and Multi-Robot Motion Planning on Graphs of Space-Time Convex Sets
Spatiotemporal motion planning, especially in multi-robot settings, requires robots to reason about collision-free regions that change over…
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant vid…
Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval
The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards fo…
Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning
Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference e…
A Multi-Resolution Finite-Volume Inspired Deep Learning Framework for Spatiotemporal Dynamics Prediction
Predicting complex spatiotemporal dynamics in physical processes often demands computationally expensive numerical methods or data-driven n…
Predicting Lethal Outcome (Cause) And Understanding Key Biomarkers Linked With Acute Myocardial Infarction Using Deep Artificial Neural Network And Ensemble Of Machine Learning Methodologies
Cardiovascular disease is still one of the main causes of death around the world. Acute myocardial infarction (MI), or heart attack, claims…
Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces
Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied a…
PAPA: Online Personalized Active Preference Alignment
Diffusion models are highly effective at modeling complex data distributions, including images and text. However, in applications like pers…
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the…
BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal
We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference thr…
Cross4D-JEPA: Dense Cross-modal Correspondence Distillation for 4D Point Cloud Representation Learning
Automatic understanding of dynamic 4D point clouds, the 3D-point sequences captured over time by depth sensors and LiDAR, is central to rob…
AI, Trust, and Teaming: The Humans-as-Handlers Approach for Autonomous and Opaque AI Systems
Artificial intelligence (AI) is becoming ubiquitous, and across domains, increasingly autonomous systems are carrying out tasks which raise…
From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that hu…
Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization
Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of train…
Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition
Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps b…
EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes
Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to e…
Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks
Lightweight machine learning models are increasingly proposed for intrusion detection in Industrial Internet of Things (IIoT) networks due…
Group-Equivariant Poincar\'e Convolutional Networks
While recent advancements like the Poincar\'e ResNet have demonstrated the potential of learning visual representations directly in hyperbo…
A Methodology for Investigating AI Patterns Prevalence in Software Repositories
As Artificial Intelligence(AI)-based applications take off, a clear understanding of AI patterns can uplift the quality of AI applications.…
Auditing Forgetting in Limited Memory Language Models
Limited Memory Language Models (LMLMs) externalize factual knowledge to a database to enable deletion-based unlearning without retraining.…
Identifying Latent Concepts and Structures for Generalized Category Discovery
Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. Howe…
Loss Smoothing for Stable Adaptation Under Distribution Shift
In settings such as fine-tuning and reinforcement learning, neural networks are often adapted under distribution shift. Standard adaptation…
Faithful by Definition: Emotion Analysis via Natural Semantic Metalanguage Explications
Explanations for emotion classifiers are usually produced post hoc, with no guarantee that they reflect the computation behind the label. W…
Multi-Label Node Classification with Label Influence Propagation
Graphs are a complex and versatile data structure used across various domains, with possibly multi-label nodes playing a particularly cruci…
LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter
Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so…
LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution
LLVM is a widely used compiler infrastructure whose scale and complexity make issue resolution labor-intensive and challenging. Although la…
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what dat…
Self-conditioned Flow Map Language Models via Fixed-point Flows
Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text…
Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach
Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet pr…
Detecting the Undetectable: Enhancing Unsupervised time series Anomaly Detection via Active Learning
Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time s…
LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data
Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inferen…
ConRTF: Edge-Constrained Boundary Distribution Refinement for Realtime TransFormer Table Structure Recognition
Table Structure Recognition (TSR) aims to recover the row and column layout of tables from document images, a key step in document understa…
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival…
Prototype Memory-Guided Training-Free Anomaly Classification and Localization in Prenatal Ultrasound
Prenatal anomaly classification and localization is of critical importance for fetal health and pregnancy management. Although ultrasound (…
GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception
The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach…
Active Learning for Cascaded Object Detection: Balancing Coverage and Uncertainty in Table Extraction Pipelines
Table extraction from business documents relies on a cascaded pipeline where Table Detection (TD) first localizes tables and Table Structur…
LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives
Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted n…
LRAT-Catcher: Importing SAT Solver Certificates into Lean4 by Reflection
SAT solvers settle combinatorial problems beyond the reach of interactive theorem provers and produce LRAT certificates for independent ver…
Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows
Large language models (LLMs) are increasingly used to generate queries, invoke tools, and construct analytical workflows. Although recent a…
Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences
A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling tr…
From World Models to World Action Models: A Concise Tutorial for Robotics
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities…
Recovering Input Text from Hidden States: Study of Gradient-Based Inversion of Decoder-Only Language Models
This work studies the hidden-state inversion problem: recovering the original input token sequence of a decoder-only language model from it…
Meta-Transfer Learning for mmWave Beam Alignment
Millimeter-wave (mmWave) beam alignment plays a critical role in next-generation wireless systems, yet its efficient implementation remains…
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
Large Reasoning Models (LRMs) have achieved remarkable success on complex tasks by leveraging long chain-of-thought (CoT) trajectories, yet…
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
Recent advances in neural rendering have established 3D Gaussian Splatting (3DGS) as a highly efficient representation for novel view synth…
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing met…
Valdi: Value Diffusion World Models
World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enough for online use and e…
From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives
Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consist…
Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm
Generalizing machine learning models to environments that differ from their training distribution remains a critical hurdle, particularly w…
Post-Training Pruning for Diffusion Transformers
Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhe…
Learning Cardiac Motion Priors for Implicit Neural Representations
Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motio…
Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations
Time-series models are often evaluated by what they can forecast or classify, but those scores do not show whether their representations pr…
TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling
Chest X-ray multi-label classification is a core task in intelligent medical imaging diagnosis. However, real clinical data often exhibit e…
SenseWalk: Agent-Based Semantic Trajectory Simulation Powered by Large Language Models in Zoned Environments
Semantic trajectory analysis has recently emerged as an approach for modeling human movement by capturing implicit patterns and behaviors t…
SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests
Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches fr…
Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally…
Reading Order Inference for Complex Document Layouts
Reading order inference remains a critical bottleneck in the digitization of complex historical manuscripts, where pages contain multiple s…
Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework
Large language model (LLM)-based conversational agents (CAs) are now ubiquitous, creating new opportunities for AI-mediated behavior change…
EchoRisk: A Multicentre Echocardiography Dataset and Benchmark for Cardio-Oncology
Therapy-induced cardiotoxicity is the leading non-oncological cause of treatment interruption in breast cancer patients, yet early, automat…
DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation
Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain…
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collabo…
Staleness-Learning Rate Scaling Laws for Asynchronous RLHF
High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learne…
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video…
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around ab…
CausalMix: Data Mixture as Causal Inference for Language Model Training
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture…
FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many e…
Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach
University stakeholders often face difficulties in accessing timely and reliable information, especially in developing countries, where the…
Muon as a Residual Connection
Muon has recently emerged as one of the most effective optimizers for training large neural networks, yet its empirical success has been ex…
Autonomous Scientific Discovery via Iterative Meta-Reflection
Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and v…
Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains
Agent skills package reusable operational knowledge for Large Language Model (LLM) agents, yet as they grow in scope, they become dependenc…
Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search
While generative models have enabled training-free reward alignment, current methods typically excel in local exploration within narrow reg…
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has fol…
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list,…
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations
RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as c…
World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video
We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our appr…
GPU-Parallel Linearization Error Bounds for Real-Time Robust Optimal Control of Nonlinear and Neural Network Dynamics
This paper studies real-time robust optimal control for uncertain nonlinear systems, where linear time-varying (LTV) approximations make pl…
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at sca…
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to…
FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the fi…
The State-Prediction Separation Hypothesis
Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We…
Language-Critique Imitation Learning from Suboptimal Demonstrations
Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estim…
Measuring the Gap Between Human and LLM Research Ideas
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or…
Unexplainability of Artificial Intelligence Judgments and Functional Implementation in Kant's Perspective
Kant's Critique of Pure Reason, a major contribution to the history of epistemology, proposes a table of categories to elucidate the struct…
From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems
To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level haz…
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large La…
Large language models replicate and predict human cooperation across experiments in game theory
Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior…
CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
Chain-of-Thought (CoT) reasoning enhances the problem-solving ability of large language models (LLMs) but leads to substantial inference ov…
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reas…
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of ev…
SleepLM: Natural-Language Intelligence for Human Sleep
We present SleepLM, a family of sleep-language foundation models that enable human sleep alignment, interpretation, and interaction with na…
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
Frontier Multimodal Large Language Models (MLLMs) exhibit remarkable capabilities in Visual-Language Comprehension (VLC) tasks. However, th…
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible…
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existi…
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere…
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science. While the scientific…
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research…
Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligence
Distributed collaborative intelligence (DCI), encompassing edge-to-edge architectures, federated learning, transfer learning, and swarm sys…
LLM における推論の質の測定: 多次元の行動フレームワーク
LLM は複雑な推論タスクで目覚ましい成功を収めていますが、現在の評価アプローチは主に最終的な答えの正しさに依存しており、それらの答えを生み出す根本的な推論プロセスについての洞察は限られています。このギャップに対処するために、この研究では、動作の観点から LLM の推論品質を測定するための統一された多次元フレームワークを提案し、理論的に根拠のある 6 つの次元、正確性 (CQ)、一貫性 (CS)、堅牢性 (RS)、論理的一貫性 (LS)、効率 (ES)、安定性 (SS) を運用します。 4 つのベンチマークの 975 項目にわたる 7 つの LLM に関する広範な実験により、このフレームワークが精度のみの指標では見えない動作を明らかにすることが実証されました。特に、論理的一貫性は正しさ (r = -0.172、ns) と直交しており、一貫性のない推論から正しい答えが得られることが確認され、一方、Claude-Haiku-4.5 は最高の多次元スコア (Q_bal = 0.778) を達成しています。さらに、このフレームワークは重大なランキングの逆転を明らかにしています。DeepSeek-V3 は精度優先では 2 位ですが、法的/コンプライアンスの重み付けでは 5 位にランクされており、単一指標の評価では検出できない逆転です。判別式の妥当性により、11/15 次元のペアが独立している (|r| < 0.50) ことが確認され、各次元を別個の信号として扱うための心理測定的サポートが提供されます。フレームワークによって生成される次元プロファイルは、次の 3 つのクラスの展開決定を直接サポートします。最終的な答えが正しいにもかかわらず、その推論トレースが説明責任監査に失敗するモデルを特定します (LS--CQ 直交性)。精度のみのベンチマークによって引き起こされるランキングエラーを防止します。そして、フレームワークがキャプチャする 6 つの独立したシグナルを単一のメトリックが暗黙的に置き換えることがないようにします。
原文 (English)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes a unified multi-dimensional framework for measuring LLM reasoning quality from a behavioral perspective, operationalizing six theoretically grounded dimensions rooted in cognitive science: Correctness (CQ), Consistency (CS), Robustness (RS), Local Logical Coherence (LS), Efficiency (ES), and Stability (SS). The framework introduces deployment-aware aggregation, enabling context-specific model selection beyond accuracy-based leaderboards. Experiments across multiple LLMs and benchmarks reveal behaviors systematically concealed by single-metric evaluation, including the orthogonality of local logical coherence and correctness, deployment-context-dependent ranking inversions, and non-trivial dimensional profiles in small locally-deployed models. Discriminant validity analysis confirms that the proposed dimensions capture largely non-redundant signals. The resulting pipeline provides a foundation for diagnosing LLM reasoning behavior across deployment contexts, with domain-specific validation as a direction for future work.
Reasoning4Sciences: 推論言語モデルをすべての科学分野に橋渡しする
推論言語モデル (RLM) は科学研究のための強力なツールとして急速に台頭していますが、その影響は主に「ハード サイエンス」分野に集中しています。他の科学分野での RLM の導入が遅い、または導入されていないことが、研究の生産性の差の拡大を引き起こしています。この調査では、欧州研究評議会 (ERC) が使用する社会科学と人文科学、物理科学と工学、生命科学にわたる分類に従って、28 の科学分野にわたる RLM の採用に関する初めての包括的な分析を提供します。私たちは、RLM がどのように開発、評価され、分野全体に適用されるかを調査します。さらに、利用可能なドメイン固有の開発および評価リソースに基づいた成熟度指向の評価フレームワークを導入し、公開されているリソースのみを考慮した場合にさらに顕著になる RLM 成熟度の実質的な格差を明らかにします。最後に、分野を超えて普及しつつある現在の実装パラダイム、現在の課題、科学全体で RLM の導入を可能にする将来の方向性を強調します。
原文 (English)
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrated in "hard science" fields. The slow -- or lack of -- adoption of RLMs in other branches of science is causing a widening gap in research productivity. In this survey, we provide the first comprehensive analysis of RLM adoption across 28 scientific disciplines following the classification used by the European Research Council (ERC), spanning the Social Sciences and Humanities, Physical Sciences and Engineering, and Life Sciences. We examine how RLMs are developed, evaluated, and applied across disciplines. Furthermore, we introduce a maturity-oriented assessment framework based on available domain-specific development and evaluation resources, revealing substantial disparities in RLM maturity that become even more pronounced when only publicly available resources are considered. Finally, we highlight current implementation paradigms that are gaining popularity across disciplines, current challenges, and future directions in enabling RLM adoption across science.
話す前に考える: マルチエージェント社会シミュレーションにおける内部評価から公の表現まで
LLM ベースのマルチエージェント シミュレーションは、社会的相互作用、熟慮、集団的な意見のダイナミクスを研究するための有望な方法を提供します。しかし、既存の対話シミュレーション フレームワークの多くは、対話を主に観察可能なターン交換または集約された出力として表現しており、沈黙、発言意図、公的表現の背後にある内部評価プロセスを調査することが困難なままになっています。エージェントの私的な推論を公的発話の生成から分離する、インターバルベースのマルチエージェント シミュレーション フレームワークである TBS (Think-Before-Speak) を紹介します。各間隔で、すべてのエージェントは共有された対話履歴と自身の記憶に基づいて構造化された内部状態を更新します。これらの状態には、不協和音関連の評価、認識された世論環境、認識された孤立リスク、対応戦略、および発言意欲が含まれます。その後、オーケストレーターは競合する発言意図を解決し、1 つの発言を公開対話にコミットし、内部評価と公開対話が時間の経過とともに共進化できるようにします。私たちは、気候関連の政策問題に関するタウンホールでの議論を模擬して TBS を評価します。結果は、TBS が一貫した内部状態トレースを生成し、これらのトレースがターン割り当て、沈黙、メモリ条件全体にわたって体系的に変化することを示しています。不協和音関連の評価はエージェントの発言意欲を高めますが、沈黙の圧力評価はそれを低下させます。発言の意図が形成されると、公の場での表現は主に順番の割り当てルールによって形成されます。これらの発見は、TBS が内部評価から公的表現への経路を観察可能かつ分析可能にすることで、メカニズムに敏感な社会シミュレーションをサポートしていることを示唆しています。
原文 (English)
Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation
LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frameworks represent interaction mainly as observable turn exchange or aggregated outputs, leaving the internal evaluative processes behind silence, speaking intention, and public expression difficult to examine. We introduce TBS (Think-Before-Speak), an interval-based multi-agent simulation framework that separates agents' private reasoning from public utterance generation. At each interval, all agents update structured internal states based on the shared dialogue history and their own memory. These states include dissonance-related appraisal, perceived opinion climate, perceived isolation risk, response strategy, and willingness to speak. The orchestrator then resolves competing speaking intentions and commits one utterance to the public dialogue, allowing internal evaluation and public interaction to co-evolve over time. We evaluate TBS in simulated town hall discussions on a climate-related policy issue. Results show that TBS produces coherent internal-state traces and that these traces vary systematically across turn-allocation, silence, and memory conditions. Dissonance-related appraisal increases agents' willingness to speak, whereas silence-pressure appraisal decreases it. Once speaking intention is formed, public expression is shaped mainly by turn-allocation rules. These findings suggest that TBS supports mechanism-sensitive social simulation by making the pathway from internal evaluation to public expression observable and analyzable.
TerraBench: エージェントは異種の地球システム データを推論できるか?
気候と環境に関する意思決定では、グリッド化された物理データ、衛星画像、地理空間コンテキスト、シミュレーターの出力など、異種混合の入力全体にわたる推論がますます必要になります。気象および気候基盤モデルは適切に予測できますが、言語で対話的に推論することはできません。一方、大規模言語モデル (LLM) は言語で推論しますが、高次元の地球システム データを直接操作することはできません。その結果、地球科学における実際の科学ワークフローは十分なサービスを受けられないままです。 TerraAgent 上に構築された、根拠のある地球科学推論のベンチマークである TerraBench を紹介します。TerraAgent は、推論、ツール呼び出し、観測をインターリーブして LLM 計画を環境検索、地理空間処理、シミュレーション、アーティファクトに基づく計算のための科学ツールと結び付ける ReAct スタイルの実行可能フレームワークです。 TerraBench は、地球観測画像、グリッド データ、GIS 推論、およびシミュレーションの分析を単一の実行可能なインターフェイスに統合します。一方、以前のベンチマークは、これらの機能を狭い個別のタスクに分離していました。また、プロセスレベルのツール使用メトリクスと許容差を意識した数値スコアを組み合わせたのも、この分野では初めてです。このベンチマークは、3 つのトラック (基礎、シミュレーターベース、ドキュメントベースの検証) にわたる 403 の広範なエージェント タスクと、24,500 の検証済み実行ステップを含む 8 つのアプリケーション ドメインで構成されています。これらの結果は、信頼できる地球科学エージェントはツールへのアクセスを超えて、異種ワークフローを調整し、ツールを正確にパラメータ化し、成果物の出所を保存する必要があることを示しています。
原文 (English)
TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reason interactively in language, while large language models (LLMs) reason in language but cannot operate directly on high-dimensional Earth-system data. As a result, real scientific workflows in Earth-science remain underserved. We introduce TerraBench, a benchmark for grounded Earth-science reasoning, built on TerraAgent, a ReAct-style executable framework that interleaves reasoning, tool calls, and observations to couple LLM planning with scientific tools for environmental retrieval, geospatial processing, simulation, and artifact-backed computation. TerraBench unifies analysis of Earth observation imagery, gridded data, GIS reasoning and simulation in a single executable interface, whereas prior benchmarks isolate these capabilities into narrow individual tasks. It is also the first in this space to pair process-level tool-use metrics with tolerance-aware numeric scoring. The benchmark comprises 403 extensive agentic tasks across three tracks (Fundamentals, Simulator-Grounded, and Document-Grounded Verification) and eight application domains with 24,500 verified execution steps. These results indicate that reliable Earth-science agents must go beyond tool access to coordinate heterogeneous workflows, parameterize tools precisely, and preserve artifact provenance.
WorkBench の再訪: 2 年後の職場エージェント
2024 年 3 月の WorkBench で最も優れたエージェントである GPT-4 は、タスクの 43% を完了し、そのうちの 26% で間違った人にメールを送信するなど、意図しない有害なアクションを実行しました。 2026 年 6 月にベンチマークを再確認したところ、これまでで最高のエージェントである Claude Opus 4.8 が 89% を完了し、2.5% で意図しない有害なアクションを実行していることがわかりました。フロンティアエージェントのパフォーマンスにおけるこの大幅な進歩とは別に、3 つの点が際立っています。まず、WorkBench では機能と安全性がトレードオフではなく両立するため、最も多くのタスクを完了したモデルは、意図しない損傷も最小限に抑えられます。第 2 に、いくつかの種類のエラーは完全に排除されましたが、フロンティア モデルは依然としていくつかの基本的な間違いを犯し、間違った人に電子メールを送信するなど、時として取り返しのつかない損害をもたらすことがあります。第三に、オープンウェイト モデルの台頭により、以前は独自モデルでしかアクセスできなかったパフォーマンス レベルのコストが大幅に低下し、一方でフロンティア コストは比較的安定しています。 2024 年以降、データとコードの品質向上、新しいモデル スコア、WorkBench でのエージェントの進行状況の分析を備えたベンチマークの更新バージョンをリリースします。
原文 (English)
WorkBench Revisited: Workplace Agents Two Years On
The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now completes 98%. Beyond this considerable progress in frontier agent performance, three things stand out. First, unintended harmful actions, such as emailing the wrong person, fell from 26% of tasks for GPT-4 to 1.9% for Claude Fable 5; capability and safety go together on WorkBench rather than trade off, so the models that finish the most tasks also do the least unintended damage. Second, the rise of open-weight models has drastically lowered costs for a performance level that was only accessible to proprietary models, while frontier costs have stayed stable. Third, while several classes of error have been eliminated, frontier models still make some basic mistakes that occasionally result in irreversible harm. We release an updated version of the benchmark with data and code quality improvements, new model scores, and analysis of agent progress on WorkBench since 2024.
Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval
Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate…
Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty
Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Mod…
HiComm: マルチエージェント強化学習のための階層型通信
協調的なマルチエージェント強化学習 (MARL) は、部分的な可観測性を軽減するために通信に依存することがよくありますが、既存のプロトコルのほとんどは、メッセージを要約した観測の構造から切り離されたフラットな密ベクトルとして扱います。この設計では、観察がグループやエンティティなどの階層に自然に従う多くの協力的な環境において、帰納的バイアスの重要な原因を見落としています。私たちは、送信者の階層的な観察にメッセージを基礎付けるプラグイン通信モジュールである \textsc{HiComm} を提案します。 \textsc{HiComm} は受信者主導型です。受信者はクエリを発行し、最初にグループ、送信者、そのグループ内のエンティティを選択する 3 段階のデコード プロセスを通じて階層を解決し、対応する機能スライスをメッセージとして返します。これにより、通信が非構造化ベクトル送信から、送信者の監視階層を介した構造化情報検索に変換されます。このメカニズムは、微分可能な離散選択と標準 MARL パイプラインに接続する軽量の共有投影設計のために、Straight-Through Gumbel-Softmax を使用してインスタンス化されます。さまざまな観察構造と調整要求を備えた協調的な MARL タスクにわたる実験では、\textsc{HiComm} が代表的な学習済み通信ベースラインと一致またはそれを上回るパフォーマンスを示し、一方、エピソードごとに受信者あたり通信量を最大 $23\times$ 削減することが示されました。
原文 (English)
HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to $23\times$ per receiver per episode.
FADE: 大規模な視覚言語モデルにおける言語優先支配を軽減することによる幻覚の軽減
Large Vision-Language Model (LVLM) の優れた機能にもかかわらず、依然として幻覚の影響を受けやすく、入力画像と一致しないコンテンツが生成されます。最近の研究では、これは視覚入力に対する言語事前の優位性によるものであり、この優位性を緩和するために対照的なデコード方法が採用されていますが、そのメカニズムの起源は未解明のままです。各変換層を通る情報の流れを調査すると、アテンション モジュールが一貫して視覚的証拠を集約し、クリティカル層の FFN モジュールが言語事前情報のソースとして機能することがわかりました。これらの事前分布は視覚的な証拠を無効にする可能性があり、中間層での正しい予測が不正確な出力に向かってドリフトする原因となります。この洞察に基づいて、言語優先の優位性を減らすために FFN 出力を減衰するトレーニング不要の方法である FADE (FFN Attenuation for DEcoding) を提案します。 LLaVA-1.5、mPLUG-Owl2、および InstructBLIP にわたる POPE、CHAIR、および MME ベンチマークの評価では、FADE が推論効率を維持しながら幻覚を効果的に軽減することが示されています。
原文 (English)
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language priors over visual inputs and employ contrastive decoding methods to mitigate this dominance, but the mechanistic origin remains unexplored. We investigate the information flow through each transformer layer and find that attention modules consistently aggregate visual evidence, while FFN modules at critical layers act as the source of language priors. These priors can override visual evidence, causing correct predictions in intermediate layers to drift toward incorrect outputs. Based on this insight, we propose FADE (FFN Attenuation for DEcoding), a training-free method that attenuates FFN outputs to reduce language-prior dominance. Evaluations on POPE, CHAIR, and MME benchmarks across LLaVA-1.5, mPLUG-Owl2, and InstructBLIP show that FADE effectively mitigates hallucinations while preserving inference efficiency.
トリプレットの妥当性を超えて: ナレッジ グラフにおける関係セットの完成
ナレッジ グラフ (KG) は、現実世界の知識をトリプレットとして編成し、多くの下流アプリケーションを支えます。本質的に不完全であるため、ナレッジ グラフ補完 (KGC) は広く研究されており、通常はリンク予測が主要なパラダイムであるトリプレット予測として定式化されます。ただし、この定式化はトリプレットごとの情報の不完全性に焦点を当てており、エンティティ関係の互換性情報の不完全性を見落としています。この制限に対処するために、リンク予測タスクを補完し、特定のエンティティと意味的に互換性のある欠落している関係を推論することを目的とした関係セット完了タスク (RSC) を導入します。さらに、観察されたエンティティの関係間の潜在的なパターンをモデル化して、欠落しているものを推測する関係セット埋め込みモデル (RelSetE) を提案します。 RelSetE を評価するために、標準の KG ベンチマークから 3 つのベンチマーク データセットを導出します。広範な実験により、RelSetE がエンティティ関係の互換性パターンを効果的に捕捉し、エンティティの欠落した関係を推論する際に有利に機能することが実証されました。コードとデータは公開されています。
原文 (English)
Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs
Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incompleteness, knowledge graph completion (KGC) is widely studied and is typically formulated as triplet prediction, with link prediction as the dominant paradigm. However, this formulation focuses on the incompleteness of triplet-wise information and overlooks the incompleteness of entity-relation compatibility information. To address this limitation, we introduce a relation set completion task (RSC), which complements the link prediction task and aims to reason about missing relations that are semantically compatible with a given entity. We further propose a Relation Set Embedding model (RelSetE), which models latent patterns among the observed relations of entities to infer missing ones. To evaluate RelSetE, we derive three benchmark datasets from standard KG benchmarks. Extensive experiments demonstrate that RelSetE effectively captures entity-relation compatibility patterns and performs favorably in inferring missing relations of entities. Code and data are publicly available.
関連性は許可されていません: 価値のある貢献には注意が必要です
関連性は許可ではありません。アテンションを使用すると、モデルは現在のクエリに関連するキーと値の項目を読み取ることができますが、そのような項目の値の寄与が予測の証拠になることは保証されません。検索された文章は、裏付けとなる証拠がなくても質問に関連している可能性があり、歴史的事実や時間的近傍によって、トゥルーテール ランキングや現在のエッジ スコアが曖昧になる場合もあります。この論文では、このギャップを、実際に予測パスに追加される加重値項 alpha_ij * v_j の許可問題として形式化します。私たちは、注意関連性 alpha_ij を保持し、プライマリ メトリックにつながる値パスを公開し、完全なモデルでは、学習されたクエリ項目権限 g_ij を通じて alpha_ij * v_j を alpha_ij * g_ij * v_j に変換するパスローカライズされたインターフェイスである Warrant を提案します。 CTDG リンク予測、MTPP 次マーク ランキング、RAG サポート証拠選択、STPP 次位置予測、および TKG テール予測のメトリクスを定義する値パスに同じ演算子を配置します。 32 のペア比較、3 つのシード、合計 192 回の実行にわたって、Warrant は 27 の比較で主要な指標を改善しました。実用的な段階は、10 個の実質的な効果、1 個の限界効果、8 個のプラスだが不確実な効果、8 個の同程度/無視できる効果、および 5 個のドロップで構成されます。パス ローカライゼーション チェックでは、正しいパス配置は、すべてのドメインで方向認識ベース パフォーマンスを上回り、一般的なアテンション配置を CTDG で +0.1076 AUC、TKG で +0.0683 MRR 上回ります。アブレーションの結果、ほとんどの TKG ゲインはヒストリカルテール値パスのエクスポージャから得られるのに対し、コア CTDG ゲインはエッジ条件付きクエリ項目のアクセス許可から得られることがわかります。結論として、予測証拠は注目マスではありません。重み付けされた値の項は、メトリックへのパス上で保証される場合にのみ証拠になります。
原文 (English)
Relevance Is Not Permission: Warranted Attention for Value Contributions
Relevance is not permission. Attention lets a model read key-value items related to the current query, but it does not guarantee that the value contribution of such an item becomes prediction evidence. A retrieved passage may be relevant to a question without being supporting evidence, and a historical fact or temporal neighbor may even blur true-tail ranking or the current edge score. This paper formalizes this gap as a permission problem for the weighted value term alpha_ij * v_j that is actually added to the prediction path. We propose Warrant, a path-localized interface that preserves attention relevance alpha_ij, exposes the value path leading to the primary metric, and, in the full model, turns alpha_ij * v_j into alpha_ij * g_ij * v_j through learned query-item permission g_ij. We place the same operator on the metric-defining value paths of CTDG link prediction, MTPP next-mark ranking, RAG supporting evidence selection, STPP next-location forecasting, and TKG tail prediction. Across 32 paired comparisons, 3 seeds, and 192 total runs, Warrant improves the primary metric in 27 comparisons; practical tiers consist of 10 substantial effects, 1 marginal effect, 8 positive but uncertain effects, 8 tie/negligible effects, and 5 drops. In the path-localization check, correct-path placement outperforms direction-aware Base performance in every domain and exceeds generic attention placement by +0.1076 AUC in CTDG and +0.0683 MRR in TKG. Ablations show that most TKG gains come from historical-tail value path exposure, whereas the core CTDG gain comes from edge-conditioned query-item permission. In conclusion, prediction evidence is not attention mass. A weighted value term becomes evidence only when it is warranted on the path to the metric.
ManimAgent: 視覚教育のための自己進化型マルチモーダル エージェント
マルチラウンド リフレクションを使用すると、大規模な言語モデルに基づいて構築されたエージェントが単一タスク内の障害から回復できますが、各タスクは孤立したエピソードのままになります。つまり、1 つのタスクで多くのリフレクション ラウンドにわたって学習された教訓は、次のタスクが開始される前に破棄されます。私たちはコード生成タスクでこのギャップを研究します。科学論文のセクションから、エージェントはオープンソースの Manim ライブラリに Python を記述して数学的アニメーションをレンダリングします。 ManimAgent は自己進化するマルチモーダル エージェントであり、重みの更新も人間のシードも必要とせず、完全に独自のタスク ストリームから成長したデュアル チャネルのエピソード メモリ バンクを通じて、タスク全体にわたってリフレクション エクスペリエンスを伝達します。各アニメーションが収束した後、ビジョン言語モデルがレンダリングされたキーフレームをスコアリングします。結果として得られる信号は、成功の根拠をソフト参照例として保存する正のチャネル M+ と、検証された失敗パターンをハードな既知の落とし穴として保存する負のチャネル M- に設定されます。メモリなし、一致した予算の検索拡張生成、およびシャッフルされたメモリ ベースラインに対する固定プローブ評価では、メモリ サイズが大きくなるにつれて、盲目の人間の Pass@1 が増加し、リフレクション ラウンドが減少します。コード、フリーズされたメモリ スナップショット、タスク ストリームをリリースします。
原文 (English)
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins. We study this gap on a code-generation task: from a scientific paper section, the agent writes Python in the open-source Manim library to render a mathematical animation. We present ManimAgent, a self-evolving multimodal agent that carries reflection experience across tasks through a dual-channel Episodic Memory Bank grown entirely from its own task stream, with no weight updates and no human seeds. After each animation converges, a vision-language model scores the rendered keyframes; the resulting signals populate a positive channel M+ that stores success rationales as soft Reference Examples, and a negative channel M- that stores validated failure patterns as hard Known Pitfalls. On a fixed-probe evaluation against no-memory, matched-budget retrieval-augmented generation, and shuffled-memory baselines, blind human Pass@1 rises and reflection rounds fall as memory size grows. We will release the code, frozen memory snapshots, and the task stream.
エキスパート ユーザーを超えて: エージェントは、ユーザーが好みを引き出すだけでなく、ユーザーが好みを構築できるように支援する必要があります。
通常、エージェントは専門ユーザー (自分が望むものについて明確な好みを持っているユーザー) を想定しており、タスクの指定が不十分な場合は常に質問を明確にするようデフォルト設定されています。私たちは、この仮定は非現実的であると主張します。ユーザーは多くの場合、好みを完全に指定するためのドメイン知識が不足しています。ある機能の好みについて尋ねられた場合、ユーザーは、例や説明などを通じて、その機能の好みを形成するために必要なドメイン知識をユーザーが学習できるようにエージェントが支援しなければ、答えることができない場合があります。これらの原則を形式化するために、情報経済学の Search-Experience-Credence フレームワークを利用して、ユーザーがエージェントの対話アクションに基づいて好みを構築する方法のモデルである CoPref を導入します。次に、これらのアイデアをエージェント レコメンダー システムで具体的に研究し、インタラクティブなベンチマークである CoShop を提案します。 CoShop では、エージェントが CoPref ユーザーと会話し、CoPref ユーザーに対して推奨事項を作成します。エージェントのパフォーマンスは、ユーザーがタスクを適切に指定するために必要な知識を得るのに役立つかどうかによって決まります。 5 つのフロンティア モデルを評価すると、5 ターンの対話にもかかわらず、CoShop で 56% の精度を超えるエージェントは存在しないことがわかりました。失敗の原因は、エージェントがアイテムを見つける能力にあるのではなく、インタラクションによってユーザーが欲しいものについて知っている範囲がほとんど広がっていないことに起因します。
原文 (English)
Beyond expert users: agents should help users construct preferences, not just elicit them
Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified. We argue this assumption is unrealistic. Users often lack the domain knowledge to have completely specified preferences; if asked about their preference on some feature, the user may be unable to answer without the agent helping the user to learn some domain knowledge needed to form a preference for that feature, e.g., via examples or explanations. To formalize these principles, we draw on the Search-Experience-Credence framework from Information Economics to introduce CoPref, a model of how users construct preferences based on agent dialog actions. We then study these ideas concretely in agentic recommender systems, proposing CoShop, an interactive benchmark. In CoShop, an agent converses with and makes recommendations for a CoPref user. The agent's performance depends on whether it can help the user gain the knowledge needed to specify the task well. Evaluating five frontier models, we find that no agent exceeds 56% accuracy on CoShop despite five turns of interaction. Failures stem not from agents' ability to find items, but from how little the interaction expands what users know about what they want.
なぜ 2 回解決するのか?伝達効率の高い ML エンジニアリングのためのスキルの階層的蓄積
すべての競争はコールド スタートであるため、ML エンジニアリング エージェントは既知の技術を再発見するために計算を無駄にします。我々は、競合他社にまたがる知識を 3 つのスコープ層 (グローバル、ドメイン、および競合固有) に編成し、それぞれがマッチング エージェント レベルに関連付けられた階層型マルチエージェント システムである HASTE を紹介します。オーケストレーターはドメイン スペシャリストを調整し、LLM 駆動の抽象化を通じて層間の学習を促進します。制御されたアブレーションは、範囲指定されたローディングの証拠を提供します。8 つの競技会にわたって 159 のスキル インベントリを一定に保持し、段階的ローディングでは 100% のメダル率を達成しますが、フラット ローディングでは 62.5% にしか達せず、スキルをロードしない場合と同じメダル率であり、出力トークンの 2 倍を消費します。 MLE-Bench Lite ベンチマーク全体 (Kaggle コンペティション 22 件) では、HASTE はクロード ソネット 4.6 を使用してコンペごとに 12 時間で 77.3% のメダル率に達しました。コールドスタート実行では、システムはスキルが蓄積されていない状態で開始されます。ウォーム スタート ランでは、以前の競技会で学んだスキルを再ロードし、競技会間での転送にはグローバル レベルおよびドメイン レベルのスキルのみを使用します。ウォーム スタートでは、改良の反復回数が 52% 減少し、提案された変更のうちエージェントが保持する割合は、在庫が少ない場合の 42% から、50 以上のスキルが利用可能になると 85% に増加します。これらの結果は、より優れた知識組織が ML エンジニアリング エージェントのモデルの強度と計算予算の一部を置き換えることができることを示唆しています。
原文 (English)
Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering
ML engineering agents waste compute rediscovering known techniques because every competition is a cold start. We present HASTE, a hierarchical multi-agent system that organizes cross-competition knowledge into three scope tiers (global, domain, and competition-specific), each coupled to a matching agent level. An orchestrator coordinates domain specialists and promotes learning between tiers via LLM-driven abstraction. A controlled ablation provides evidence for scoped loading: holding a 159-skill inventory constant across 8 competitions, tiered loading achieves a 100% medal rate while flat loading reaches only 62.5%, the same medal rate as loading no skills, and consumes 2x the output tokens. On the full MLE-Bench Lite benchmark (22 Kaggle competitions), HASTE reaches a medal rate of 77.3% using Claude Sonnet 4.6 at 12h per competition; this is a single-seed campaign result, and multi-seed replication is the priority follow-up. In a cold-start run, the system begins with no accumulated skills. In warm-start runs, it reloads skills learned from earlier competitions, using only global and domain-level skills for transfer across competitions. Warm starts use 52% fewer refinement iterations, and the fraction of proposed changes kept by the agent rises from 42% at low inventory to 85% once 50+ skills are available. These results suggest that better knowledge organization can partly substitute for model strength and compute budget in ML-engineering agents.
Xiaomi-GUI-0テクニカルレポート
グラフィカル ユーザー インターフェイス (GUI) エージェントは、ビジョン言語モデルに基づいて構築され、タップ、スワイプ、テキスト入力、ナビゲーションなどのインターフェイス アクションを通じて実際のアプリケーションでユーザー タスクをエンドツーエンドで完了します。ただし、既存の GUI エージェントは、主にオフラインの軌跡、シミュレートされた環境、および標準化されたベンチマークに基づいてトレーニングおよび評価されます。これらは、インターフェイスのレイアウト、インタラクション ロジック、異常状態の分布において実際のアプリケーションとは大幅に異なり、実際の使用における実行の安定性を忠実に特徴付けることはできません。そこでは、アカウントの状態、許可ダイアログ、支払い認証、リスク管理によって状態分布が継続的に再形成され、ベンチマーク スコアと実際のユーザビリティの間に永続的なギャップが生じます。このギャップを埋めるために、実際のデバイスの閉ループ内でトレーニングおよび評価される、実際のモバイル環境用のネイティブ マルチモーダル GUI エージェントである Xiaomi-GUI-0 を提案します。その中核となるのは、実デバイス主体のハイブリッド インフラストラクチャであり、物理デバイスが主要な実行環境であり、サンドボックスが補助的なサポートを提供するため、データ収集、トレーニング、ロールアウト、評価が実際の展開に近い実行分布を共有します。私たちは、高頻度のヘッド タスク、ロングテール インテントの高汎化データ、リフレクションとメモリの能力強化データにわたるマルチソース トレーニング データを構築し、障害の軌跡を修正されたアクション、リフレクションの説明、回復のデモンストレーションに変えるエラー駆動型のデータ フライホイールを導入します。モデルは、教師あり微調整、ステップレベルの強化学習、エージェント強化学習の漸進的な 3 段階のパイプラインを通じてトレーニングされます。公開ベンチマークと社内 RealMobile で評価された Xiaomi-GUI-0 は、RealMobile で 72.0%、AndroidWorld で 78.9% の成功率を達成し、現実世界のタスクにおける実行の安定性と異常状態の認識が大幅に向上しました。
原文 (English)
Xiaomi-GUI-0 Technical Report
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
Growing concerns over the theft and misuse of Large Language Models (LLMs) underscore the need for effective fingerprinting to link a model…
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
Industry is moving toward autonomous, network-connected machines that detect and adapt to changing conditions, including hardware faults. C…
ForAug: Mitigating Biases in Image Classification via Controlled Image Compositions
Large-scale image classification datasets exhibit strong compositional biases: objects tend to be centered, appear at characteristic scales…
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans. But are these explanations faithfu…
Comparative Analysis of Lightweight CNNs for Resource-Constrained Devices: Predictive Performance, Efficiency Trade-offs, and Initialization Effects
Lightweight convolutional neural networks are often compared using results obtained with different training recipes, input settings, and pr…
scDataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics
Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets…
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
Humanoid robots are promising to acquire various skills by imitating human behaviors. However, existing algorithms are only capable of trac…
Flow-Through Tensors: A Unified Computational Graph Architecture for Multi-Layer Transportation Network Optimization
Modern transportation network modeling increasingly involves the integration of diverse methodologies including sensor-based forecasting, r…
Fraud is Not Just Rarity: A Causal Prototype Attention Approach to Realistic Synthetic Oversampling
Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and th…
FLAT: Revealing Hidden Latent-Conditioned Backdoor Failures in Federated Learning
Horizontal federated learning (HFL) backdoor audits often summarize model behavior through clean accuracy (CA), mean attack success rate (A…
TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification
Handling missing data in time series classification remains a significant challenge in various domains. Traditional methods often rely on i…
CWT-Enhanced Vibration Sensing With Time-Frequency Region Localization Using YOLO
This letter presents a CWT-enhanced vibration sensing framework for bearing fault monitoring through localized time-frequency region detect…
Predicting LLM Reasoning Performance with Small Proxy Model
Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize datasets bef…
Quadratic Programming Approach for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
On-device deployment of Large Language Models (LLMs) frequently leverages Low-Rank Adapters (LoRAs) to support diverse downstream tasks und…
Toward Cybersecurity-Expert Small Language Models
Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, do…
Reasoning Up the Instruction Ladder for Controllable Language Models
As large language model (LLM) based systems take on high-stakes roles in real-world decision-making, they must reconcile competing instruct…
Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture th…
FlowPath: Learning Data-Driven Manifolds with Invertible Flows for Robust Irregularly-sampled Time Series Classification
Modeling continuous-time dynamics from sparse and irregularly-sampled time series remains a fundamental challenge. Neural controlled differ…
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
Despite notable progress in text-guided medical image segmentation nowadays, these methods are limited to single-round dialogues and fail t…
NI-Tex: Non-isometric Image-based Garment Texture Generation
Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited. To ac…
When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets
Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agen…
Computing Evolutionarily Stable Strategies in Imperfect-Information Games
We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
Catastrophic forgetting poses a fundamental challenge in continual learning, particularly when models are quantized for deployment efficien…
Utilizing Earth Foundation Models to Enhance the Simulation Performance of Hydrological Models with AlphaEarth Embeddings
Predicting river flow in places without streamflow records is challenging because basins respond differently to climate, terrain, vegetatio…
Controllable Diffusion-Based Lesion Inpainting for Scalable Histopathology Data Augmentation
Expert-annotated training data remains the critical bottleneck for AI in histopathology, particularly for rare pathologies where even dozen…
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are uncha…
NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents
Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enablin…
PaAno: Patch-Based Representation Learning for Time-Series Anomaly Detection
Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transfo…
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
Mixture-of-Experts (MoE) architectures are evolving towards finer granularity to improve parameter efficiency. However, existing MoE design…
Calibrated Test-Time Guidance for Bayesian Inference
Test-time guidance is a widely used mechanism for steering pretrained diffusion models toward outcomes specified by a reward function. Exis…
Stateful Token Reduction for Long-Video Hybrid VLMs
Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated…
On the Reliability of Cue Conflict and Beyond
Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes. The cue-confl…
Competition-Aware CPC Forecasting with Near-Market Coverage
Cost-per-click (CPC) in paid search is an auction-generated outcome shaped by a competitive landscape that is only partially observable fro…
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
End-to-End autonomous driving (E2E-AD) systems face challenges in lifelong learning, including catastrophic forgetting, difficulty in knowl…
Interact3D: Compositional 3D Generation of Interactive Objects
Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional o…
Deep Learning-Driven Black-Box Doherty Power Amplifier with Pixelated Output Combiner and Extended Efficiency Range
This article presents a deep learning-driven inverse design methodology for Doherty power amplifiers (PA) with multi-port pixelated output…
KGS-GCN: Kinematics-Driven Gaussian Splatting and Probabilistic Topology for Skeleton-Based Action Recognition
Skeleton-based action recognition is widely applied in sensor-based systems, including human-computer interaction and intelligent surveilla…
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-l…
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
The identification of hazardous driving behaviors from in-cabin video streams is essential for enhancing road safety and supporting the det…
Planning over MAPF Agent Dependencies via Multi-Dependency PIBT
Modern Multi-Agent Path Finding (MAPF) algorithms must plan for hundreds to thousands of agents in congested environments within a second,…
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization
Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging becau…
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updat…
Moir\'e Video Authentication: A Physical Signature Against AI Video Generation
Recent advances in video generation have made AI-synthesized content increasingly difficult to distinguish from real footage. We propose a…
Crystalite: A Lightweight Transformer for Efficient Crystal Modeling
Generative models for crystalline materials often rely on equivariant graph neural networks, which capture geometric structure well but are…
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
AI agents that pay for resources via the x402 protocol embed payment metadata - resource URLs, descriptions, and reason strings - in every…
Continuous Knowledge Metabolism: Generating Scientific Hypotheses from Evolving Literature
Identifying promising research directions in fast-moving subareas is one of the most cognitively expensive tasks in modern AI research. Exi…
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial re…
FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices
Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data, y…
REALM: An RGB- and Event-Aligned Latent Manifold for Cross-Modal Perception
Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and rob…
SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
De-identification of clinical text is a prerequisite for the secondary use of electronic health records. Existing public benchmarks such as…
FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles
Recent progress in 4D Gaussian Splatting (4DGS) has achieved impressive dynamic scene reconstruction results. While these methods demonstra…
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclea…
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionat…
EcoGEO: Trajectory-Aware Evidence Ecosystems for Web-Enabled LLM Search Agents
Web-enabled LLM agents are changing how online information influences search outcomes. Existing Generative Engine Optimization (GEO) studie…
ChainCaps: 単調な機能減衰による構成安全なツール使用エージェント
ツールを使用するエージェントは、実行時にファイル システム、Web API、コード インタープリタ、およびエンタープライズ サービスを構成する、オープンエンドの展開環境で動作することが増えています。これにより、ツール構成に安全性のギャップが生じます。エージェントは、ツールごとのすべての権限チェックを満たしていても、機密文書の読み取り、要約、その要約の外部エンドポイントへの送信など、安全でないエンドツーエンドの影響を生み出す可能性があります。この障害モードをパーミッション ロンダリングと呼びます。 ChainCaps は、ランタイム ルールでこれに対処します。すべての値にはシンク固有の機能バジェットが含まれ、ツールの構成によって交差ごとにバジェットが伝播されます。値は、ツール チェーン内を移動するときに権限を保持したり失ったりする可能性がありますが、合成を通じて新しい権限を獲得することはできません。 ChainCaps は、エージェント サーバーやツール サーバーへの変更を必要としない透過的な MCP プロキシとして実装されています。 ChainCaps は、3 つのプロバイダーの 5 つのフロンティア モデルにわたる 82 のタスクにおいて、96 ~ 100% の無害な完了を維持しながら、攻撃の成功率を 25 ~ 68% から 0 ~ 4.8% に低下させます。再生実験では、スカラー IFC および関数ごとの分離ベースラインよりも優れたパフォーマンスを示します。マニフェストの品質が導入の主なボトルネックです。エキスパート マニフェストは 100% の攻撃ブロックに達しますが、単純なマニフェストは 27.3% に低下します。私たちの主張は、信頼できるマニフェストの下での明示的なフロー構成の安全性と、プロキシで可視のデータ移動に限定されており、今日導入されているツールを使用するエージェントには実際的なギャップがあります。
原文 (English)
ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation
Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterprise services at runtime. This creates a safety gap in tool composition: an agent can satisfy every per-tool permission check and still produce an unsafe end-to-end effect, such as reading a confidential document, summarizing it, and sending the summary to an external endpoint. We call this failure mode permission laundering. ChainCaps addresses it with a runtime rule: every value carries a sink-specific capability budget, and tool composition propagates budgets by intersection. A value can preserve or lose authority as it moves through a tool chain, but it cannot gain new authority through composition. We implement ChainCaps as a transparent MCP proxy that requires no changes to the agent or tool servers. On 82 tasks across five frontier models from three providers, ChainCaps reduces attack success rate from 25-68% to 0-4.8% while preserving 96-100% benign completion. In replay experiments, it also outperforms scalar-IFC and per-function-isolation baselines. Manifest quality is the dominant deployment bottleneck: expert manifests reach 100% attack blocking, while naive manifests fall to 27.3%. Our claims are limited to explicit-flow composition safety under trusted manifests and proxy-visible data movement, a practical gap in deployed tool-using agents today.
Diffusion Image Generation with Explicit Modeling of Data Manifold Geometry
Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, lo…
Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation
Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity…
LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs). Current QAT methods are mainly based on…
Korzhinskii-Net: 地下鉱物予測モデリングのための物理学に基づいたニューラル ネットワーク
鉱物見通しモデリング (MPM) は探査の経済学を支えていますが、ほとんどの運用パイプラインは浅い表面のプロキシでトレーニングされたデータ駆動型の分類器に限定されています。このようなモデルは、熱移流、流体の流れ、岩石学に依存する降水など、実際に鉱石の位置を特定する地下物理学を認識できません。我々は、ダルシー流、移流拡散熱輸送、およびソフトプラス飽和反応速度を単一の微分可能な順方向モデルに結合し、表面およびリモートセンシングプロキシによって弱く監視される 2 次元放射状物理情報ニューラル ネットワーク (PINN) である Korzhinskii-Net を紹介します。このネットワークは、浸透メタソマティズムの理論が物理的な足場を提供するドミトリ S. コルジンスキー (1899 ~ 1985 年) にちなんで名付けられました。我々は、ハードリング形状のネガを用いた公平で漏洩制御された5重相互検証プロトコルの下で、ノリリスク(Ni-Cu-PGE)、ペチェンガ(硫化Ni-Cu)、ウドカン(砂岩にホストされたCu)、スホーイ原木(造山運動のAu)、およびミールヌイ(キンバーライトダイヤモンド)の4つの商品クラスにまたがる5つの鉱石州でKorzhinskii-Netを評価しました。 Korzhinskii-Net は、最も強い古典的ベースライン (勾配ブースティング) の平均 PR-AUC 0.885 対 0.281、および平均分数ランク 0.019 対 0.413 を達成します。この改善は 5 つの州と 4 つの商品システムすべてで一貫しており、物理学に基づいた微分可能シミュレーターは、グローバルなオープンデータ プロキシによってのみ制約されている場合でも、純粋な特徴ベースの学習者が体系的に見逃している位置特定パターンを回復できることを示唆しています。完全なパイプラインと評価ハーネスをオープンソースとしてリリースします。
原文 (English)
Korzhinskii-Net: Physics-Informed Neural Network for Sub-Surface Mineral Prospectivity Modelling
Mineral prospectivity modelling (MPM) underpins exploration economics, yet most operational pipelines reduce to data-driven classifiers trained on shallow surface proxies. Such models are blind to the subsurface physics that actually localises ore: heat advection, fluid flow, and lithology-dependent precipitation. We present Korzhinskii-Net, a 2-D radial physics-informed neural network (PINN) that couples Darcy flow, advective-diffusive heat transport, and a softplus-saturated reaction rate into a single differentiable forward model, weakly supervised by surface and remote-sensing proxies. The network is named after Dmitri S. Korzhinskii (1899-1985), whose theory of infiltration metasomatism provides the physical scaffold. We evaluate Korzhinskii-Net on six ore provinces spanning three commodity classes - Udokan (sandstone-hosted Cu), Sukhoi Log, Olimpiada, and Berezovskoye (orogenic Au), Vorontsovskoye (Carlin-type Au), and Dalnegorsk (skarn polymetallic) - under a fair, leakage-controlled 5-fold cross-validation protocol with hard ring-shaped negatives and baseline proxy features disabled. Korzhinskii-Net attains a mean PR-AUC of 0.708 versus 0.235 for the strongest classical baseline (support vector machine), and a mean fractional rank of 0.036 versus 0.475. The improvement is consistent across all six provinces and three commodity systems, suggesting that physics-informed differentiable simulators, even when constrained only by global open-data proxies, can recover localisation patterns that pure feature-based learners systematically miss. We release the full pipeline and evaluation harness as open source.
L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification
Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become ent…
Vibecoding Ate My 宿題: グリーンフィールド ソフトウェア エンジニアリングとプログラミングへの AI アプローチの評価
生成 AI の急速な発展のおかげで、私たちはコンピューターとの対話方法を永遠に変える可能性のあるパラダイム シフトの真っ只中にいます。この分野の基礎知識なしにアプリケーションやコーディング インフラストラクチャを構築するための自然言語プロンプトの使用が増加していることが観察されており、この実践は「バイブ コーディング」と呼ばれています。これはおそらく、プログラミングの分野が当初から、考えられるあらゆるより高い抽象化レベルで構築されてきたものを表しています。 Vibe コーディングは、入力方法に関する限り、高レベル プログラミングのメタのエンドポイントとなることが約束されています。つまり、人間によるコード構文の使用が完全に排除され、母国語でのプログラミングが優先されます。このペーパーは、グリーンフィールドのソフトウェア エンジニアリング タスクにおける Vibe コーディングの実現可能性を評価し、そのソフトウェア エンジニアリングの能力を測定するために使用されたベンチマークを分析することを目的としています。この目的を達成するために、私たちは、Python で単純で個別のグリーンフィールド プログラミング タスクを実行する LLM の習熟度を分析し、この問題に関する範囲を絞った洞察を提供するための評価スイートを開発しました。
原文 (English)
Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming
Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever. We have observed a growth in the use of natural language prompts to build applications and coding infrastructures without underlying knowledge of the field, and this practice has been dubbed `vibe coding.' It arguably represents what the field of programming has been building towards since the beginning, with every higher level of abstraction that is conceived. Vibe coding promises to be the endpoint for the meta of high-level programming as far as method of input is concerned: eliminating a human's use of code syntax entirely in favour of programming in their mother tongue. This paper aims to evaluate the viability of vibe coding for greenfield software engineering tasks, as well as analyse the benchmarks that have been used to measure its software engineering prowess. To this end, we have developed an evaluation suite for analysing an LLM's proficiency in carrying out simple, isolated greenfield programming tasks in Python to provide scoped insight on the matter.
PSCT-Net: 微分可能な逆投影と注意に基づく改良による、形状を考慮した小児頭蓋骨 CT 再構成
コンピュータ断層撮影 (CT) は小児の頭蓋顔面異常の診断に不可欠ですが、発達中の解剖学的構造に放射線リスクをもたらします。まばらな二平面 X 線から 3D CT を再構成することは、低線量の代替手段となりますが、非常に不適切です。既存の方法は、ジオメトリに依存しない特徴リフティングを採用しており、明示的な空間モデリングを行わずに単純に 2D 特徴を 3D に投影するため、深さの曖昧さと骨境界の劣化が生じます。微分可能な逆投影を備えた幾何学認識フレームワークである PSCT-Net を紹介します。微分可能な逆投影により、空間的に忠実な体積事前分布が確立され、深さの曖昧さが軽減されます。次に、注意誘導投影 (AGP-3D) モジュールが、2D 領域と 3D 位置の間の非線形ボクセル単位の対応を学習します。 Bi方向 Mamba (BiM-3D) モジュールは、線形の複雑さで長距離の体積依存関係をキャプチャします。さらに、内部評価用に正常症例と病理学的症例で構成される民間の施設小児頭蓋骨CTコホートであるPedSkull-CTをキュレーションし、成人中心の体幹に焦点を当てたデータセットのギャップに対処します。
原文 (English)
PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement
Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomies. Reconstructing 3D CT from sparse bi-planar X-rays offers a low-dose alternative but is severely ill-posed. Existing methods employ geometry-agnostic feature lifting, naively projecting 2D features into 3D without explicit spatial modeling, causing depth ambiguity and degraded osseous boundaries. We present PSCT-Net, a geometry-aware framework with differentiable back-projection. Differentiable back-projection establishes a spatially faithful volumetric prior, alleviating depth ambiguity. An Attention-Guided Projection (AGP-3D) module then learns non-linear voxel-wise correspondences between 2D regions and 3D locations. A Bidirectional Mamba (BiM-3D) module captures long-range volumetric dependencies with linear complexity. We further curate a private institutional pediatric skull CT cohort, PedSkull-CT, comprising normal and pathological cases for internal evaluation, addressing the gap in adult-centric, trunk-focused datasets.
オプティカル フローを学習するための普遍的な制約としての三角形の一貫性
我々は、オプティカル フローの第一原理制約として三角整合性を提案します。これは、ネットワーク アーキテクチャ、監視タイプ、データセットに依存せず、画像ペアとマルチフレーム設定の両方に適用されます。このシンプルだが強力な制約は、2 つのフローを構成して 3 つ目のフローを誘導し、3 つのフロー間の一貫性を強制することです。合成されたフローは、(i) 画像ペアから生成され、サイクルの一貫性が得られます。 (ii) 複数のビデオ フレーム。時間的連鎖を通じてより長距離の動きを生成します。または (iii) 画像ペアを制御された合成変換と組み合わせて、データ拡張となります。この三角形の一貫性により、無視できるほどの計算オーバーヘッドが発生し、追加の注釈は必要ありません。オプティカル フローのジオメトリから直接導出されるため、モデル固有の仮定に依存せず、オプティカル フロー トレーニング用の「ユニバーサル」プラグ アンド プレイ コンポーネントとして機能します。実験では、教師あり、教師なし、転移学習の設定全体で一貫した改善が見られました。
原文 (English)
Triangular Consistency as a Universal Constraint for Learning Optical Flow
We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision type, and dataset, and applies to both image-pair and multi-frame settings. This simple but powerful constraint is to compose two flows to induce a third flow and enforce consistency among the three. The composed flows may arise from (i) image pairs, yielding cycle consistency; (ii) multiple video frames, producing longer-range motion through temporal chaining; or (iii) image pairs combined with controlled synthetic transformations, which becomes data augmentation. This triangular consistency introduces negligible computational overhead and requires no additional annotations. Since it is derived directly from the geometry of optical flow, it does not rely on model-specific assumptions and serves as a ``universal'' plug-and-play component for optical flow training. Experiments show consistent improvement across supervised, unsupervised, and transfer learning settings.
Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling
Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…
1 年後...被害は続いていますが、私たちも同じです!
汎用大規模言語モデル (LLM) は、メンタルヘルス関連の会話にますます使用されていますが、安全対策は依然として不十分であり、臨床症状全体で一貫性がありません。この研究では、8 次元の危害分類と多次元の評価フレームワークを導入し、4 つの敵対的攻撃のバリアントを使用して、16 の DSM-5 条件にわたる 6 つの独自の LLM を評価します。その結果、安全策は自殺と自傷行為に対してのみ確実に有効であり、摂食障害、物質使用障害、大うつ病性障害などの疾患では失敗率が最大 100% であることが示されています。私たちは、これらの LLM の倫理的な設計と展開には、臨床状態全体にわたって明確に定義された危害カテゴリーと、それに応じた安全措置の実装が必要であると主張します。このような保護措置が講じられるまで、これらのモデルは脆弱な人々に重大なリスクをもたらすため、教育現場への統合の増加が特に懸念されます。
原文 (English)
One Year Later...The Harms Persist, But So Do We!
General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety guardrails remain inadequate and inconsistent across clinical conditions. This study evaluates eight proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100%. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into publicly available settings (e.g., schools, search engines, and consumer chatbots) are particularly concerning.
構造的に忠実: 複数文書の要約のためのクレームにアンカーされた帰属
エンドツーエンドの大規模言語モデル (LLM) は、流暢な複数文書の要約を生成しますが、依然として幻覚を起こしやすく、また、LLM が提供する帰属は一般的に粗く (文書全体または文章全体)、事後的に生成されるため、各要約ステートメントを検証するのは困難です。私たちはモジュール式の抽出-選択-再書き込みパラダイムを再考し、その中間表現を帰属の単位として再構築します。我々は、(i) すべてのソース文書からトークンレベルの出自を持つ原子的なクレームを抽出し、(ii) ソース間の競合にフラグを立てながらドキュメント全体で同等のクレームをクラスター化し、(iii) サポートを意識した顕著なサブセットを選択し、(iv) 選択した内容を、すべての文が 1 つ以上のソース スパンにリンクするサポートチェックされたクレームに固定された要約に書き換える、クレームアンカー型マルチドキュメント要約フレームワークである CAMS を紹介します。コンテンツは実現される前にローカライズされるため、パイプラインは構造的には帰属指向であり、構造的には忠実性を重視します。つまり、サポートを意識した選択、制約付き書き換え、検証を使用して、事実の忠実性を保証するのではなく奨励しながら、構造的にきめの細かいマルチソースのトレーサビリティを維持します。私たちは、MultiNews で品質、忠実性、ローカリゼーションを評価し、DiverseSumm で競合処理を分析し、WCEP でゼロショット転送をテストします。この際、参考文献フリーの引用品質とゴールドアライメントのローカリゼーション精度を分離する 2 つのレジームプロトコルを使用します。さらに、選択や検証に決して使用されないサポートモデルで引用精度をテストする評価者分離監査を追加します。 CAMS は、要約の品質に関して強力なエンドツーエンドおよびスパン帰属ベースラインを照合すると同時に、忠実性と引用の精度を大幅に向上させ、複数情報源の帰属精度を約 3 分の 2 向上させ、制御可能な忠実性、つまりエンドツーエンド モデルが暗黙的に残しているカバレッジのトレードオフを明らかにします。
原文 (English)
Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization
End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify. We revisit the modular Extract--Select--Rewrite paradigm and recast its intermediate representation as the unit of attribution. We present CAMS, a Claim-Anchored Multi-document Summarization framework that (i) extracts atomic claims with token-level provenance from every source document, (ii) clusters equivalent claims across documents while flagging inter-source conflicts, (iii) selects a support-aware and salient subset, and (iv) rewrites the selection into a summary in which every sentence is anchored to a support-checked claim that links back to one or more source spans. Because content is localized before it is realized, the pipeline is attribution-oriented by construction and faithfulness-oriented by construction: it structurally preserves fine-grained, multi-source traceability while using support-aware selection, constrained rewriting, and verification to encourage, rather than guarantee, factual faithfulness. We evaluate quality, faithfulness, and localization on MultiNews, analyze conflict handling on DiverseSumm, and test zero-shot transfer on WCEP, using a two-regime protocol that separates reference-free citation quality from gold-aligned localization accuracy, and we add an evaluator-decoupled audit that tests citation precision with a support model never used for selection or verification. CAMS matches strong end-to-end and span-attribution baselines on summary quality while substantially improving faithfulness and citation precision, lifting multi-source attribution accuracy by roughly two-thirds, and exposing a controllable faithfulness--coverage trade-off that end-to-end models leave implicit.
Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection
With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud.…
Variable Bound Tightening for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
またあれは何だったのでしょうか?自動音声認識の堅牢性が認定済み
自動音声認識システムは、敵対的な摂動にも無害な摂動にも敏感であることで知られています。これは参照データセットを使用して繰り返し実証されていますが、実際の転写に関するオラクルの知識が存在しないため、展開されたシステムでそのような動作を検出することは非常に困難です。私たちは、認定にヒントを得たメカニズムを採用すると、WER が大幅に減少し、再現率が増加し、信頼性と WER の間のスピアマン相関が減少する可能性があることを実証します。これは、デュアルゲート診断パイプラインを通じて実現されます。トークンの存在と敵対的排除の両方を証明するための統計的富を蓄積する両面アトミック監査と、勝利シーケンスを選択するランクベースのトーナメントです。 4 つの多様なアーキテクチャにわたる当社の評価では、単語エラー率が相対的に最大 55% 減少することが実証されていると同時に、音響セキュリティを強化するための詳細な単語および文レベルの認定も提供されます。
原文 (English)
What Was That Again? Certified Robustness for Automatic Speech Recognition
Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle knowledge of the true transcription. We demonstrate that employing a certification-inspired mechanism can significantly decrease WER, increase recall, and decrease the Spearman correlation between confidence and WER. We achieve this through a dual-gate diagnostic pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate up to a 55% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.
エラーの余地: 無線音響攻撃の大規模シミュレーション
音声制御は人間と AI のコミュニケーションのユビキタスなベクトルとして急速に普及しつつありますが、これらのシステムが直面するリスクについてはまだ十分に理解されていません。これは、部分的には、厳密にデジタルの敵対的なワークフローを物理世界に拡張する際の困難の結果です。これらのスケールの壁により、コミュニティは、検出可能性と音響に対する形状の影響に関連する主要な音響要素を抽象化するようになりました。これらの方法論的および計測学的欠点は、リスクに対する私たちの理解を台無しにします。私たちは、現実世界でのテスト、概念的な議論、新しい高スループットの現実シミュレーション フレームワークを通じて、これらの問題を明らかにします。 800 万を超える敵対的評価をテストすることにより、Whisper と wav2vec の下では、音響認識により相対的な単語誤り率が最大 94.5\% 増加することが実証されました。私たちはこのフレームワークを使用して、デュアルフォーム信号対雑音比の形式化と運用を検討し、ソースのステルスを被害者の攻撃の有効性から分離し、現在の作業における重大な制限を解決します。これにより、音響環境を抽象化するのではなく包含する、再現可能で検証可能な研究の基礎が築かれます。
原文 (English)
Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks
While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital adversarial workflows to the physical world. These scale barriers have led the community to abstract away key acoustic factors relating to detectability and the influence of geometry on acoustics. These methodological and metrological shortcomings undermine our understanding of risk. We illuminate these issues through real-world testing, conceptual discussions, and a novel, high-throughput reality simulation framework. By testing over 8 million adversarial evaluations, we demonstrate that acoustic awareness yields relative Word Error Rate increases of up to 94.5\% under Whisper and wav2vec. We employ this framework to explore a formalize and operationalize a Dual-Form Signal to Noise Ratio to decouple source stealth from victim attack efficacy, resolving a crucial limitation in current works. This lays the groundwork for repeatable, verifiable research that embraces, rather than abstracts, the acoustic environment.
Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models
Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale wa…
Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking
There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully…
When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking
Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying t…
ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries
Large language models deployed in regulated industries operate under two constraints: compliance enforcement and cost efficiency. Personall…
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot man…
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as…