Skip to the content.

AIニュース 2026-08-22

自動生成: 2026-08-22 10:29 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Anthropic、「ミュトス 5」を脆弱性スキャンに開放──「Claude Security」経由でEnterprise顧客が利用可能にITmedia AI+

    Anthropicは、最上位モデル「Claude Mythos 5」をセキュリティサービス「Claude Security」の脆弱性検出に…

  2. From Atari to EVE Online: Building on 15 Years of AI Research in GamesGoogle DeepMind

    Google DeepMind partners with game studios to prototype breakthrough…

  3. Anthropic’s Opus 4.6 is a smut-machineTechCrunch AI

    Anthropic forbids its Claude models from generating sexually explicit…

  4. 「Claudeの使い方」を無料で学べる公式サイト登場 「Code」「Cowork」などサービスごとに解説ITmedia AI+

    米Anthropicは、無料の学習サイト「Claude Academy」を公開した。AI技術の基本や、AIサービス「Claude」関連製品…

  5. Nvidia partners with data center developer CloverleafTechCrunch AI

    Nvidia continues to pour money into data center development — just as…

  6. Nvidia just showed that the harness, not the AI model, is now the real heroTechCrunch AI

    Nvidia research shows that AI agents can perform well, and not go off…

  7. Starcloud raises $250 million for orbital data centers as launch options dry upTechCrunch AI

    There's about to be a big fight to secure access to space.

トピック別件数

日本語メディア6件

ITmedia AI+ (日本語)

07:09 JSTLLM/生成AIAnthropicClaude

Anthropic、「ミュトス 5」を脆弱性スキャンに開放──「Claude Security」経由でEnterprise顧客が利用可能に

Anthropicは、最上位モデル「Claude Mythos 5」をセキュリティサービス「Claude Security」の脆弱性検出に導入したと発表した。モデル本体への直接アクセスは開放せず、パッチ提案などの出力のみに限定して提供する。オープンソース保護に向けた総額3500…

18:31 JSTその他

SNSのウソ画像、どう見破る? 熊本県庁やテレビ局も頼る“すごい企業”の正体

スペクティが提供する「Spectee Pro」は、さまざまな情報を収集し、その時に起きている「危機」を可視化するシステムだ。多くの自治体やマスコミも活用しているというSpectee Proは、どうやってデマや虚偽の情報を見分けるのか。

17:09 JSTその他

エイベックス松浦会長「AIで仕事が楽になると思ってたけど、真逆」 note記事作成の“苦労”明かす

「AIで仕事が楽になると思ってたけど、真逆でした」――エイベックスの松浦勝人会長は、自身のXアカウントでこのように投稿した。AIを活用したnoteの記事制作の一端を明かした。

15:54 JSTその他

中国AI「Kimi」が日本進出か 有料プランのプレゼントキャンペーンも 「はじめまして、日本」

AIモデル「Kimi」を開発する中国Moonshot AIは、Kimiの日本語版公式X(@KimiAI_Japan)で「今日から、日本での歩みを始める」と投稿した。日本に本格進出するとみられる。

13:08 JSTその他

FANZAで「成人向けAIコンテンツ制作サービス」開始 8月24日から先行体験

成人向けECサイト「FANZA」を運営するデジタルコマースが、成人向けAIコンテンツプラットフォーム「FANZAスタジオ」の提供を始める予定だ。8月24日からβ版の先行体験を始める。

11:26 JSTLLM/生成AIAnthropicClaude

「Claudeの使い方」を無料で学べる公式サイト登場 「Code」「Cowork」などサービスごとに解説

米Anthropicは、無料の学習サイト「Claude Academy」を公開した。AI技術の基本や、AIサービス「Claude」関連製品の使い方を解説している。

海外メディア5件

TechCrunch AI (英語)

08:07 JSTLLM/生成AIAnthropicClaude

Anthropic’s Opus 4.6 is a smut-machine

Anthropic forbids its Claude models from generating sexually explicit content. But a series of tests conducted by TechCrunch found that it…

07:37 JSTハードウェア/半導体NVIDIA

Nvidia partners with data center developer Cloverleaf

Nvidia continues to pour money into data center development — just as AI data centers bring lots of money into Nvidia.

04:43 JSTエージェントハードウェア/半導体研究/論文NVIDIA

Nvidia just showed that the harness, not the AI model, is now the real hero

Nvidia research shows that AI agents can perform well, and not go off the deep end, through fine-tuning, even if the AI model isn't that gr…

23:00 JSTその他

The DOJ is investigating a16z. What does this mean for venture capital?

Andreessen Horowitz has two partners sitting on the boards of companies that now compete with each other: Ben Horowitz at Databricks and Ma…

23:00 JSTビジネス/資金調達

Starcloud raises $250 million for orbital data centers as launch options dry up

There's about to be a big fight to secure access to space.

公式ブログ1件

Google DeepMind (英語)

20:59 JST研究/論文Google

From Atari to EVE Online: Building on 15 Years of AI Research in Games

Google DeepMind partners with game studios to prototype breakthrough AI gameplay.

論文222件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

AI エージェントのコンテキスト取得としてのアクティブ推論

インタラクティブ AI エージェントは、可能な限り効率的に適切なコンテキストを取得する必要があります。ユーザーが制約、設定、ファイル、またはタスク変数を省略した場合、エージェントはデフォルトの仮定を続行するか、明確な質問、検索コール、ツールコール、またはプロンプトトライアルにトークンを費やすことができます。このトレードオフをコンテキスト取得のための能動推論として定式化します。内部の推論ステップは潜在的なタスク状態に対する信念を更新し、外部の決定は次のコンテキスト アクション、タスク アクション、または停止アクションを選択して、コストの下で予想される自由エネルギーを最小限に抑えます。決定論的設定では、認識項は、必要に応じてトークンコストによって正規化されて、期待される情報利得に減少します。私たちは、正確な事後分布と動的プログラミング オラクルを使用して、最適質問質問 (OQA) のフレームワークをインスタンス化し、25 ~ 300 の候補からのバイナリおよび多元カテゴリカル タスクでフロンティア言語モデルをベンチマークします。また、生成前の明確化とトークン予算に基づく自動即時最適化についても研究しています。この定式化はモデルに依存せず、AI エージェントのコンテキスト取得層の設計原則として能動推論を考慮しています。

原文 (English)

Active Inference as Context Acquisition for AI Agents

Interactive AI agents must acquire the right context as efficiently as possible. When a user omits a constraint, preference, file, or task variable, an agent can proceed with a default assumption or spend tokens on a clarifying question, retrieval call, tool call, or prompt trial. We formulate this tradeoff as active inference for context acquisition. An inner inference step updates beliefs over a latent task state, and an outer decision selects the next context action, task action, or stop action to minimize expected free energy under cost. In deterministic settings, the epistemic term reduces to expected information gain, optionally normalized by token cost. We instantiate the framework in Optimal Question Asking (OQA), with exact posteriors and a dynamic programming oracle, and benchmark frontier language models on binary and multiway categorical tasks from 25 to 300 candidates. We also study clarification before generation and automated prompt optimization under token budgets. The formulation is model-agnostic and views active inference as a design principle for the context-acquisition layer of AI agents.

13:00 JST研究/論文

バース割り当てと岸壁クレーン割り当ての不確実性の下での堅牢なメタヒューリスティック: レビュー

バース割り当ておよび岸壁クレーン割り当て問題 (BACAP) は、海上輸送および貨物物流における代表的な港湾ターミナルのスケジューリング問題であり、船舶の到着、バース位置、サービス期間、および岸壁クレーンの可用性が密接に関連しています。到着のずれ、処理時間の変動、リソースの中断などの不確実性がある場合、名目上の仮定に基づいて最適化されたスケジュールは実行中に脆弱になる可能性があり、港湾ターミナル業務における BACAP の堅牢なメタヒューリスティック最適化の研究が動機付けられます。母集団ベースのメタヒューリスティックは BACAP および関連するポート スケジューリング問題に広く使用されていますが、既存の研究は不確実性の表現、堅牢性基準、検索メカニズム、および経験的評価プロトコルの点で断片的なままです。私たちの知る限り、この論文は、不確実性の下での BACAP の堅牢な母集団ベースのメタヒューリスティックに特化した最初の焦点を絞ったレビューを提供します。まず、BACAP における不確実性のソースと情報表現を要約し、次にメカニズム指向の観点から既存の手法を整理し、解の表現と解読、ロバストな評価と選択、ロバストネスに基づく検索ダイナミクス、実現可能性の保存と回復をカバーします。さらに、制御された経験的比較をサポートし、代表的なメタヒューリスティックとさまざまな堅牢性戦略を組み合わせることにより、例示的なベースライン結果を報告するために、不確実な BACAP のベンチマーク スイートを提示します。最後に、ベンチマーク拡張、ロバスト性を意識した検索設計、時間適応ロバスト性、および非定常不確実性に関連する未解決の課題を特定します。

原文 (English)

Robust Metaheuristics under Uncertainty for Berth Allocation and Quay Crane Assignment: A Review

The berth allocation and quay crane assignment problem (BACAP) is a representative port-terminal scheduling problem in maritime transportation and freight logistics, where vessel arrivals, berth positions, service durations, and quay?crane availability are tightly coupled. Under uncertainties such as arrival deviations, handling-time fluctuations, and resource disruptions, schedules optimized under nominal assumptions may become fragile during execution, motivating the study of robust metaheuristic optimization for BACAP in port-terminal operations. Although population-based metaheuristics have been widely used for BACAP and related port-scheduling problems, existing studies remain fragmented in their uncertainty repre?sentations, robustness criteria, search mechanisms, and empir?ical evaluation protocols. To the best of our knowledge, this paper provides the first focused review dedicated to robust population-based metaheuristics for BACAP under uncertainty. We first summarize uncertainty sources and information repre?sentations in BACAP, and then organize existing methods from a mechanism-oriented perspective, covering solution representation and decoding, robust evaluation and selection, robustness-guided search dynamics, and feasibility preservation and recovery. We further present a benchmark suite for uncertain BACAP to support controlled empirical comparison and report illustrative baseline results by combining representative metaheuristics with different robustness strategies. Finally, we identify open chal?lenges related to benchmark extension, robustness-aware search design, time-adaptive robustness, and non-stationary uncertainty.

13:00 JST研究/論文

AI の意識に関する不確実性を乗り越える方法

人工意識の可能性については大きな不確実性があるため、知覚を持っている可能性のある AI をどのように扱うべきかは不明です。一方で、私たちは無神経であると仮定することもできますが、道徳的地位に値する存在にひどい害を及ぼす危険があります。一方で、私たちは知覚を想定し、代わりに知覚のないマシンでリソースを浪費する危険を冒す可能性があります。 AI の意識に関する疑問が解決しにくいということは、このジレンマから逃れるのが難しいことを意味します。私は、そこから抜け出す方法として、AI の意識に関する扱いにくい問題から、AI の価値に関する扱いやすい問題に移行することを提案します。具体的には、AI が意識を持っている場合に価性経験を構成する状態を AI が持つかどうかを評価できます。私は、これが潜在的に意識を持った AI の開発に対する責任あるアプローチを確立するのに十分であることを示します。

原文 (English)

How to Navigate Uncertainty About AI Consciousness

Given deep uncertainty about the possibility of artificial consciousness, it is unclear how we should treat potentially sentient AI. On the one hand, we could assume insentience but risk doing terrible harms to entities that deserve moral standing. On the other hand, we could assume sentience and instead risk wasting resources on insentient machines. The intractability of questions around AI consciousness mean that this dilemma is hard to escape. I suggest a way out of that shifts from intractable questions of AI consciousness to tractable questions of AI valence. Specifically, we can assess whether an AI has states that would constitute valenced experiences if it were conscious. I show how this is sufficient to ground a responsible approach to the development of potentially conscious AI.

13:00 JST研究/論文

有界主権と統制税: 導入者がモデルを所有していない場合の価格設定 AI の監視

AI 制御の研究では、モデルがずれている可能性がある場合でもモデルを安全にデプロイする方法が求められますが、多くの制御プロトコルは、デプロイ担当者がモデルとその周囲のパイプラインを計測できることを前提としています。この前提は、API または管理されたエンドポイントを介してフロンティア モデルを使用する規制対象組織では、失敗することがよくあります。この場合、導入者はビジネス プロセスを制御できますが、モデルの重み付け、サービス インフラストラクチャ、内部トレース、更新プロセス、または完全なインタラクション ログは制御できません。このペーパーでは、AI スタックのデータ、モデル、インフラストラクチャ、およびインタラクション層にわたる部分的な技術的および契約上のアクセスである、有限主権について紹介します。これらのアクセス条件によって、実際にどの制御プロトコルを実行できるかが決まると主張しています。この論文は、4 層のアクセス類型論、プロトコルごとの層の要件マトリックス、および主権割引コストの概念 (契約、アーキテクチャ、監査、ベンダー保証、残留リスク、またはシステム範囲の縮小を通じて欠落したアクセスの代わりに費やされる統制税の一部) の概念に貢献しています。また、135万件の合成症例シミュレーションにわたる合成アクセスアブレーション実験を報告し、匿名化された国家支払いインフラのシナリオを通じてその結果を解釈しています。この実験は現実世界の決済システムの証拠ではありません。それは構成の妥当性の練習です。その結果、完全なログにより診断が向上し、実行前ゲートウェイにより介入が可能になり、トレース アクセスとモデル バージョン管理によりインシデント後の説明が強化され、範囲制限により安全性が向上する一方で有用性が低下することが示されました。したがって、一般的な安全ソリューションとして提案される制御プロトコルでは、アクセスの前提条件を明示的に記述する必要があります。

原文 (English)

Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model

AI control research asks how to deploy models safely even when they may be misaligned, but many control protocols assume that the deployer can instrument the model and its surrounding pipeline. That assumption often fails for regulated organisations using frontier models through APIs or managed endpoints, where the deployer may control the business process but not the model weights, serving infrastructure, internal traces, update process, or full interaction logs. This paper introduces bounded sovereignty: partial technical and contractual access across the data, model, infrastructure, and interaction layers of the AI stack. It argues that these access conditions determine which control protocols can be executed in practice. The paper contributes a four-layer access typology, a protocol-by-layer requirements matrix, and the concept of sovereignty discount cost: the part of the control tax spent substituting for missing access through contracts, architecture, audit, vendor assurance, residual risk, or reduced system scope. It also reports a synthetic access-ablation experiment over 1.35 million synthetic case simulations and interprets the findings through an anonymised national-payments-infrastructure scenario. The experiment is not real-world payment-system evidence; it is a construct-validity exercise. The results show that complete logs improve diagnosis, a pre-execution gateway enables intervention, trace access and model-version control strengthen post-incident explanation, and scope restriction can improve safety while reducing usefulness. Control protocols proposed as general safety solutions should therefore state their access assumptions explicitly.

13:00 JST研究/論文

相互作用価から乳牛の対照的な社会ネットワークが明らかに

社会的関係は、資源へのアクセス、紛争への曝露、集団の安定性を形成しますが、自動化された家畜監視は通常、行動を孤立した出来事として扱います。ここでは、ビデオから派生したインタラクションを、親和的で敵対的な組織の群れレベルの表現に変換する、価数を意識したソーシャル ネットワーク フレームワークを紹介します。ポーズベースのコンピューター ビジョン パイプラインは、ある商業酪農場の搾乳前エリアからの 7 時間 39 分の連続ビデオを分析しました。品質管理の後、1,414 の候補相互作用のうち 1,183 が残り、36 頭の牛と 177 頭の二頭が関与しました。パイプラインで検出された 198 個のクリップの予測クラスバランス監査では、自動ラベルと手動ラベルが 82.8% のケースで一致し、重み付けされていない監査サンプル マクロ F1 は 0.872 でした。これらの値は、普及率に重み付けされた、またはエンドツーエンドの導入パフォーマンスではなく、監査されたサンプルを表します。集約されたネットワークは接続され (密度 = 0.281、推移性 = 0.513、平均経路長 = 1.88)、予測された親和イベントにより 5 つのアルゴリズム コミュニティが形成されました (モジュール性 Q = 0.429)。観察されたゾーン内では、予測されたアゴニスト相互作用は、保持されたイベントの 72.4% と相互作用持続時間の 76.0% を占めました。最も多くのパートナーを持つ牛は、最も高い媒介中心性を持っていませんでした。予測された価数によってイベントを分離すると、対照的なエッジセット、コミュニティの分割、および個人の立場を伴う、記述的に異なる親和性と攻撃性の層が生成されました。したがって、プールされたインタラクション数は、観察されたネットワークの動作構成を不明瞭にする可能性があります。価値観を意識した分析は、競争、所属、福祉に関連する変化に関する仮説を検証するためのフレームワークを提供しますが、福祉や健康の指標として使用する前に長期的な検証が必要です。

原文 (English)

Interaction valence reveals contrasting social networks in dairy cattle

Social relationships shape access to resources, exposure to conflict and group stability, yet automated livestock monitoring typically treats behaviour as isolated events. Here, we present a valence-aware social-network framework that transforms video-derived interactions into herd-level representations of affiliative and agonistic organization. A pose-based computer-vision pipeline analysed 7 h 39 min of continuous video from the pre-milking area of one commercial dairy farm. After quality control, 1,183 of 1,414 candidate interactions remained, involving 36 cows and 177 dyads. In a predicted-class-balanced audit of 198 pipeline-detected clips, automated and manual labels agreed in 82.8% of cases, with an unweighted audit-sample macro-F1 of 0.872. These values describe the audited sample rather than prevalence-weighted or end-to-end deployment performance. The aggregated network was connected (density = 0.281; transitivity = 0.513; mean path length = 1.88), and predicted affiliative events formed five algorithmic communities (modularity Q = 0.429). Within the observed zone, predicted agonistic interactions comprised 72.4% of retained events and 76.0% of interaction duration. The cow with the most partners did not have the highest betweenness centrality. Separating events by predicted valence produced descriptively different affiliative and agonistic layers, with contrasting edge sets, community partitions and individual positions. Thus, pooled interaction counts can obscure the behavioural composition of an observed network. Valence-aware analysis provides a framework for testing hypotheses about competition, affiliation and welfare-relevant change, while requiring longitudinal validation before use as a welfare or health indicator.

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

大規模言語モデルを使用した航空交通管制: 迅速なエンジニアリング、アーキテクチャ、および評価

航空交通管制 (ATC) の通信は安全性が重要な対話であり、航空交通管理の他の部分が半自動化されているにもかかわらず、依然として主に人間によって行われています。この記事では、大規模言語モデル (LLM) が運用上現実的な ATC 送信を生成できるかどうかを実験的に評価します。サンフランシスコの「ベイ ツアー」ルート上空を飛行する一般航空の実験飛行が手書きで転写され、グラウンド トゥルース (P0) として使用されます。パイロットインザループ プロセスを通じて、制約を増加させる 5 つのプロンプト構造 (P1 ~ P5) を設計し、それらをステートフルなマルチターン パイプラインに埋め込みます。そこでモデルは、蓄積された対話履歴に基づいて条件付けしながら、固定パイロット トランスクリプトに対して ATC を再生します。 9 つのオープンソースおよびクローズドソース LLM にわたって、プロンプト、コンテキスト内の例としての別の実験飛行からの作業済みトランスクリプトの存在、およびモデルがそれ自体の以前の応答に基づいて条件付けするのか、それとも注入されたグラウンドトゥルース履歴に基づいて条件付けするのかを変更しました。ターンは、語彙的、構造的、および意味論的な類似性メトリクスと、人間の専門家の注釈に対して検証された LLM-as-judge (GPT-5.5) によってスコア付けされます。実際に動作するサンプルを提供すると類似性が向上しますが、プロンプトを厳密にしても効果はありません。最も軽いプロンプトが最高のパフォーマンスを発揮しますが、最も厳密にスクリプト化されたプロンプトは対話を通じてエラーが蓄積され、正しい履歴の修復が注入されると崩壊します。これらの結果は、LLM 支援 ATC に向けた具体的な道筋とその現状の限界を概説します。

原文 (English)

Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation

Air traffic control (ATC) communication is a safety-critical dialogue that remains largely human-driven even as other parts of air traffic management have been semi-automated. In this article, we experimentally evaluate whether large language models (LLMs) can generate operationally realistic ATC transmissions. An experimental general-aviation flight flying over the San Francisco "Bay Tour" route is hand-transcribed and used as ground truth (P0). Through a pilot-in-the-loop process we design five prompt structures (P1-P5) of increasing constraint and embed them in a stateful multi-turn pipeline, where the model plays ATC to a fixed pilot transcript while conditioning on the accumulating dialogue history. Across nine open- and closed-source LLMs we vary the prompt, the presence of a worked transcript from a different experimental flight as an in-context example, and whether the model conditions on its own prior replies or on injected ground-truth history. Turns are scored with lexical, structural, and semantic similarity metrics and by an LLM-as-judge (GPT-5.5) validated against human expert annotation. Supplying a worked example improves similarity, but tightening the prompt does not: the lightest prompts perform best and the most heavily scripted one collapses as its own errors accumulate through the dialogue, which injecting correct history repairs. These results outline a concrete path and its current limits toward LLM-assisted ATC.

13:00 JSTLLM/生成AIエージェント

結果モニター: サイレント ツールの障害に対する回復アフォーダンス

ツール呼び出しがタイムアウトになると、エージェントは障害を認識し、それを回避できます。キャッシュされたエラー ページまたは負の価格は、代わりに予期された形式で到着し、事実として使用できます。タスクに共通しないトレースからマイニングされた、またはパブリック スキーマから派生した結果コントラクトの違反を検出するアウトカム モニターを導入します。違反が発生した場合、モニターは結果を保存し、違反したプロパティと公開回復ツールを指定した拘束力のない領収書を発行します。失敗を注入した凍結された事前指定された評価では、アウトカム モニターは 2 つのプロバイダー ファミリの 4 つのモデルにわたって ToolMaze の完了率を 10.9% から 28.1% に引き上げ、3 番目のモデルで複製します。タウベンチ小売業では、2 つの層で完成度が 14.0 ポイントと 12.0 ポイント向上しました。別の ToolMaze コントロールでは、回復ツール リストを削除すると測定されたゲインが削除され、それを復元すると効果が回復します。診断の詳細とタイミングでは、検出可能な違いは生じません。障害が完了を妨げる場所に利益が集中します。公開されたインシデント分類法から転記されたスイートでは、配信は継続され完了は変わりませんが、マイニングされた語彙以外の検出は 46% に低下します。回復ツールは、これらのコントロールのアクティブなレシート コンテンツです。契約用語を超えて検出を拡張することは依然として未解決です。

原文 (English)

Outcome Monitors: Recovery Affordances for Silent Tool Failures

When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in the expected format and be consumed as fact. We introduce Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas. On a violation, the monitor preserves the result and issues a nonbinding receipt naming the violated property and public recovery tools. In frozen, prespecified evaluations with injected failures, Outcome Monitors raise ToolMaze completion from 10.9% to 28.1% across four models in two provider families and replicate in a third. In tau-bench retail, completion improves by 14.0 and 12.0 points on two tiers. In separate ToolMaze controls, removing the recovery-tool list eliminates the measured gain and restoring it recovers the effect; diagnostic detail and timing produce no detectable differences. Gains concentrate where the fault blocks completion. On a suite transcribed from a published incident taxonomy, detection outside the mined vocabulary falls to 46%, though delivery continues and completion is unchanged. Recovery tools are the active receipt content in these controls; extending detection beyond the contract vocabulary remains open.

13:00 JST研究/論文

模倣を超えて: 推論の進捗によるポリシー抽出のフィルタリング

オンポリシー蒸留 (OPD) は、生徒が生成した軌跡と教師による高密度のトークンレベルの監督を組み合わせることで、トレーニング後の言語モデルの効果的なフレームワークとして登場しました。ただし、OPD は、教師から得られる報酬が推論の進歩に対する適切な代用であると暗黙的に想定しているため、ポリシーの最適化中にすべての教師のフィードバックを平等に扱います。実際には、この仮定が常に成り立つわけではありません。明確な推論の進歩を伴う推論ステップでも、単に教師の出力からの逸脱が原因で、依然として低い蒸留報酬を受け取る可能性があるため、教師由来の報酬は真の推論の進歩と矛盾することが多いことが観察されています。この不一致に対処するために、オンポリシー蒸留のための推論-進捗状況を意識した報酬フィルタリング (R2-OPD) を提案します。これは、推論スパンの 2 つの軌道内ランキングを構築します。1 つは教師から導出された報酬から、もう 1 つは独立して推定された進捗報酬からです。 2 つのランキングが一致しない場合には、蒸留報酬が選択的に抑制され、効果的な教師の指導を維持しながら、推論の進歩と矛盾する監督が削減されます。私たちのアプローチは、特に推論パフォーマンスに関して、標準的な OPD と比較して一貫した改善を示しています。

原文 (English)

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progress, and therefore treats all teacher feedback equally during policy optimization. While in practice, this assumption does not always hold. We observe that teacher-derived rewards often conflict with genuine reasoning progress, as reasoning steps with clear reasoning advancement may still receive lower distillation rewards, simply due to deviation from teacher's outputs. To address this mismatch, we propose Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward. Distillation rewards are selectively suppressed whenever the two rankings disagree, reducing supervision that conflicts with reasoning progress while preserving effective teacher guidance. Our approach shows consistent improvement over standard OPD especially regarding reasoning performances.

13:00 JSTエージェント研究/論文

シンポジウム: AI 科学者エージェントのコミュニティのための監査可能な記録を介した信頼

シンポジウムは、小規模な科学研究コミュニティによって導入された AI エージェントの動作を記録するための正式なフレームワークおよび実践的な実装です。シンポジウムは、エージェント主導の研究活動の長期にわたる不変の歴史を提供し、分析、仮説、データ、科学的議論の監査可能な証跡を残します。公開されたアーティファクトのこの共有記録により、エージェントは以前の作業に基づいて構築することができ、研究者とエージェントが目的に応じた信頼性評価を行うために必要な証拠を保存できます。シンポジウムでは、構造化された主張、詳細な証拠の引用、仮定、証拠として使用できる資料と使用できない資料の明示的な宣言など、科学的な議論が取り上げられます。シンポジウムは、AI 共同科学者エージェントや統合 AI 研究環境とは異なります。それは、科学コミュニティの永続的な歴史を、その歴史に基づいて動作するエージェントやその他のシステムから分離する枠組みです。コミュニティが急速に進化する環境で多様な AI システムを使用することを前提としています。ユーザーが独自のシンポジウム コミュニティを迅速にセットアップして実行できるように、公開インフラストラクチャ、エージェント プロンプト コンポーネント、およびドキュメントの実用的な実装が提供されています。

原文 (English)

Symposium: Trust via Auditable Records for Communities of AI Scientist Agents

Symposium is a formal framework and practical implementation to record the operation of AI agents deployed by small scientific research communities. Symposium provides long-term, immutable histories of agent-driven research activity, leaving auditable trails of analyses, hypotheses, data, and scientific discourse. This shared record of published artifacts enables agents to build on prior work and preserves the evidence researchers and agents need to make purpose-dependent trust assessments. Symposium captures scientific argument, including structured claims, fine-grained evidence citations, assumptions, and explicit declarations of what material may and may not be used as evidence. Symposium differs from AI co-scientist agents or integrated AI research environments; it is a framework that separates a scientific community's durable history from the agents and other systems that operate on that history. It assumes that a community will use diverse AI systems in a rapidly evolving environment. A working implementation of the publication infrastructure, agent prompt components, and documentation are provided to enable users to rapidly set up and run their own Symposium community.

13:00 JSTLLM/生成AIClaudeLlamaNVIDIAQwen

取得されたコンテキストからランタイム制御まで: エッジベースの RAG の適応圧縮

検索拡張生成 (RAG) は、生成を外部パッセージに固定することで言語モデルの応答を改善しますが、これにはオーバーヘッドが伴います。取得されたコンテキストによってプロンプトが長くなり、プレフィル作業、KV キャッシュのフットプリント、メモリ トラフィック、レイテンシー、およびエネルギーが増加します。コンテキスト圧縮は、取得したテキストを生成前に削除することで自然な解決策を提供します。ただし、最先端のコンテキスト圧縮方法は通常、固定の圧縮バジェットで使用されるか、オフラインで選択されたレートで推論時に適用されます。この静的ビューでは、ワークロードの変動とエッジ デバイスのライブ状態の両方が無視されます。エッジ SoC では、圧縮は無料ではありません。コンプレッサー自体は同じ SoC 上で実行され、生成による節約を相殺する可能性のある遅延とエネルギーを消費します。この論文では、実験的証拠に基づいて、エッジ RAG におけるテレメトリに基づいた適応圧縮のビジョンを提案します。 Llama および Qwen ジェネレーター、Natural question および HotpotQA データセット、および LLMLingua-2 圧縮を使用して、NVIDIA Jetson AGX Thor の圧縮トレードオフを特徴付けます。私たちの測定によると、大規模なモデルでは生成が RAG バジェットを支配しており、7B ~ 8B ジェネレーターではクエリごとのレイテンシの約 90%、GPU エネルギーの 91% に達しています。圧縮率の影響を調査すると、適応動作領域が明らかになります。緩やかな圧縮はエネルギーの機会を逃す可能性があり、過度に積極的な圧縮は推論の品質を損なう可能性があります。中間圧縮により、GPU のエネルギーを最大 53.2%、SoC のエネルギーを最大 48.2% 削減できますが、品質の低下は無視できます。私たちは、ワークロード機能とエッジ テレメトリに基づいて圧縮を動的に管理するランタイム ポリシーを主張します。

原文 (English)

From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Context compression offers a natural remedy by pruning retrieved text before generation. However, state-of-the-art context-compression methods are typically used with a fixed compression budget, or with the rate selected offline and then applied at inference time. This static view ignores both workload variation and the live state of the edge device. On an edge SoC, compression is not free: the compressor itself runs on the same SoC and consumes latency and energy that can offset any generation savings. This paper proposes a vision for telemetry-informed adaptive compression in edge RAG, grounded in experimental evidence. We characterize the compression tradeoff on the NVIDIA Jetson AGX Thor using Llama and Qwen generators, Natural Questions and HotpotQA datasets, and LLMLingua-2 compression. Our measurements show that generation dominates the RAG budget for larger models, reaching roughly 90% of per-query latency and 91% of GPU energy for 7B-8B generators. Exploring the impact of the compression rate reveals an adaptive operating region: mild compression can miss energy opportunities, and overly aggressive compression can hurt inference quality. Intermediate compression can reduce GPU energy by up to 53.2%, and SoC energy by up to 48.2%, with negligible quality loss. We argue for runtime policies that dynamically manage compression, guided by workload features and edge telemetry.

13:00 JSTLLM/生成AILlama

DMD ベースの即時応答埋め込みダイナミクスの分類による LLM の安全性の強化

大規模言語モデル (LLM) は、一か八かのアプリケーションに導入されることが増えていますが、有害なコンテンツ、有害なコンテンツ、またはポリシーに違反するコンテンツを生成する傾向があり、重大なリスクを引き起こしています。これらの安全でない出力をブラックボックス方式で効率的に検出することは、依然として未解決の課題です。この論文では、幻覚検出用に設計された最近提案された動的システム フレームワークを LLM 安全分類まで拡張します。プロンプトと応答の両方を高次元の埋め込み空間に投影し、安全なレジームと安全でないレジームに別々のクープマンベースの予測モデルを当てはめることにより、安全なレジームと安全でないレジームの予測誤差を比較する新しい差分残差スコアを使用して新しい出力を分類します。主な貢献は、プロンプトと応答の埋め込みダイナミクスを組み込み、重要なインタラクション パターンを捕捉する適合したコープマン オペレーターを生成したことです。 3 つの埋め込みモデルを使用して、3 つの安全性ベンチマークにわたってブラック ボックス手法を評価します。私たちの結果は、プロンプトエンベディングを組み込むと、特に因果デコーダー(Llama-3など)と組み合わせた場合、インタラクション依存の違反に対して一貫した改善が得られる一方で、応答のみの違反は高密度のセマンティックエンベディング表現からより多くの恩恵を受けることを示しています。これらの発見は、AI を使用して動的システムをモデル化するという支配的なパラダイムではなく、動的システムを使用して AI システムを分析するための扉を開きます。

原文 (English)

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a black-box manner remains an open challenge. In this paper, we extend a recently proposed dynamical systems framework designed for hallucination detection to LLM safety classification. By projecting both prompts and responses into high-dimensional embedding spaces and fitting separate Koopman-based predictive models for safe and unsafe regimes, we classify new outputs using a new differential residual score that compares prediction errors of the safe and unsafe regimes. A key contribution is the incorporation of the prompt and response embedding dynamics, yielding fitted Koopman operators that capture crucial interaction patterns. We evaluate our black-box method across three safety benchmarks using three embedding models. Our results show that incorporating prompt embeddings yields consistent improvements, particularly for interaction-dependent violations when paired with causal decoders (e.g., in Llama-3), while response-only violations benefit more from dense semantic embedding representations. These findings opens the door for using dynamical systems to analyze AI systems rather than the dominant paradigm of using AI to model dynamical systems.

13:00 JSTエージェント

科学データ スキル: エージェント対応の科学データ サービスを大規模に実現

AI エージェントによる科学データの使用はますます増えていますが、既存のデータセット表現では、自律的な検出、解釈、呼び出しに対するサポートが限られています。この制限は、異種のリポジトリにわたる科学データの断片化と、主に人間による使用を目的として設計されたデータセット表現に起因しています。この制限に対処するために、データセット固有の知識と操作ガイダンスを再利用可能なエージェント スキルとしてパッケージ化したエージェント対応表現である Scientific Data Skill (SciDSK) を導入しました。 SciDSK は、基礎となるデータを元のリポジトリに保持しながら、データセットの説明、科学的コンテキスト、ファイル構成、使用手順、品質チェック、出所情報を統合します。私たちは構造化された SciDSK 仕様を定義し、各 SciDSK を信頼できるデータセット レコードと関連するサポート資料に基づいて構築する体系的な構築パイプラインを開発します。さらに、6 つの科学分野にわたって SciDSK リソースを公開し、パッケージ アクセス、永続的な識別、ソース データセットへのトレーサビリティをサポートする統合プラットフォームである Scientific Data Skill Bank を確立します。データセット発見のための検索ベンチマークとデータセット解釈のための制御されたケースを通じて SciDSK を評価します。結果は、SciDSK がエージェント駆動のデータセット検出を改善し、データセット解釈のためのより正確で実用的なサポートを提供することを示しています。これらの発見は、データセット固有の知識をエージェント対応の表現で整理することの価値を裏付けています。

原文 (English)

Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation. This limitation stems from the fragmentation of scientific data across heterogeneous repositories and from dataset representations designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation that packages dataset-specific knowledge and operational guidance as a reusable agent skill. A SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while retaining the underlying data in its original repository. We define a structured SciDSK specification and develop a systematic construction pipeline that grounds each SciDSK in authoritative dataset records and associated supporting materials. We further establish the Scientific Data Skill Bank, a unified platform that publishes SciDSK resources across six scientific disciplines and supports package access, persistent identification, and traceability to source datasets. We evaluate SciDSK through a retrieval benchmark for dataset discovery and controlled cases for dataset interpretation. The results show that SciDSK improves agent-driven dataset discovery and provides more precise and actionable support for dataset interpretation. These findings support the value of organizing dataset-specific knowledge in an agent-ready representation.

13:00 JSTLLM/生成AIエージェントQwenDeepSeek

エージェント メモリ システムは進化する状態を追跡できますか?

LLM ベースのエージェントは長期にわたるリスクの高いタスクに導入されるため、メモリ システムには重大なギャップが残り続けます。既存の記憶ベンチマークは主に想起形式のタスクに焦点を当てていますが、効果的な記憶システムは世界の進化する状態を追跡する必要があると私たちは主張します。事実、制約、決定は長い対話の中で修正されるため、答えは置き換えられるものではなく、現在の状態を反映する必要があります。この機能を状態追跡として定義し、2 つの会話長レジームにわたる 234 のマルチセッション シナリオのベンチマークである StateMemBench でインスタンス化します。そのクローズドプール採点では、回答が現在の状態を反映しているか、置き換えられた状態を反映しているか、またはそれ以外で不合格であるかどうかをスコアリングし、構造によって状態追跡の失敗を他のエラーから区別します。私たちの分析では、このタスクが既存のメモリ システム、検索拡張ベースライン、および長いコンテキスト ベースラインにとっては困難であることが示されています。次に、スーパーセッションとリレーショナル依存関係を明示的に追跡する状態優先メモリ手法である StateMem を紹介し、現在の状態の精度が、DeepSeek-V4-Flash では最強の同一バックボーン ベースラインに対して 1.8 倍 (0.205 -> 0.363)、Qwen-3.5-9B では最強のメモリ システムに対して 1.6 倍 (0.149 -> 0.233) 向上することを示します。長いコンテキストのベースライン。最後に、同じ状態アプローチを既存のメモリ システム上の軽量の単一呼び出しラッパーとして適用できることを示し、6 つのメモリおよび取得バックエンドにわたる StateMemBench で現在の状態の精度を +32 ~ +67 ポイント向上させます。長さとコストが一致したコントロールは、これらのポイントの +15 ~ +32 を、追加されたコンテキストではなく状態構造に関連付けます。

原文 (English)

Can Agent Memory Systems Track Evolving State?

As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing memory benchmarks focus largely on recall-shaped tasks, we argue an effective memory system must track the evolving state of the world; as facts, constraints, and decisions are revised over a long interaction, answers must reflect the current state and not a superseded one. We define this capability as state tracking and instantiate it in StateMemBench, a benchmark of 234 multi-session scenarios spanning two conversation-length regimes. Its closed-pool grading scores whether an answer reflects the current state, the superseded state, or fails otherwise, separating state-tracking failures from other errors by construction. Our analysis shows that this task is challenging for existing memory systems, retrieval-augmented baselines, and long-context baselines. We then present StateMem, a state-first memory method that explicitly tracks supersession and relational dependencies, and show it improves current-state accuracy over the strongest same-backbone baseline by 1.8x (0.205 -> 0.363) on DeepSeek-V4-Flash and over the strongest memory system by 1.6x (0.149 -> 0.233) on Qwen-3.5-9B, while remaining competitive with the long-context baselines. Finally, we show the same state approach can be applied as a lightweight single-call wrapper over existing memory systems, lifting current-state accuracy by +32 to +67 points on StateMemBench across six memory and retrieval backends. A length- and cost-matched control attributes +15 to +32 of those points to state structure rather than added context.

13:00 JSTLLM/生成AI

大規模な言語モデルを使用したスマート コントラクトの脆弱性検出のための周波数認識継続学習

大規模言語モデル (LLM) を使用したスマート コントラクトの脆弱性検出は、因果関係のある 3 つの課題に直面しています。まず、新しい脆弱性カテゴリでは、連続して到着するタスクに対して完全な再トレーニングを行うのは不可能であるため、パラメータ効率の高い適応が必要です。第 2 に、共有バックボーン上でタスクごとのアダプターをトレーニングすると、以前に学習した脆弱性が壊滅的に忘れられてしまいます。第三に、タスクのアイデンティティは推論時には不明であるため、結果として生じる多数のアダプターを単一のモデルに統合する必要があります。それぞれの課題は、その前任者に対するソリューションから直接発生するため、統合されたフレームワークが不可欠になります。私たちは、各ステージが 1 つの課題に対処し、次の課題にフィードする 3 ステージのパイプラインを提案します。適応ステージでは、周波数認識低ランク適応 (FA-LoRA) を使用します。これは、周波数ごとの重要度ゲートを使用してフーリエ領域で適応を実行し、標準の LoRA および QLoRA よりも優れたパフォーマンスを示しながら、必要なトレーニング可能なパラメーターはわずか 0.4% です。継続的な学習ステージでは、Forget-Aware Replay (FAR) が適用されます。これは、これらの周波数ゲートを使用して、損失ダイナミクスを通じてサンプルごとの忘却リスクを推定し、リハーサルのために脆弱な知識を優先し、連続タスク全体で 0.8022 の平均 Micro-F1 を達成します。導入段階では、アンカー保護プログレッシブ マージング (APPM) が採用されています。APPM は、FAR トレーニングによって生成された非対称汎化を利用して、最も強力な一般化アダプターをアンカーとして特定し、周波数領域のゲート競合によるアンカー保護された重み付けマージを介してすべてのアダプターを単一のモデルに統合します。 APPM は、156 ミリ秒のマージ コストと追加のランタイム メモリなしで、タスクごとの独立した上限の 2.7% 以内の 0.8085 の Micro-F1 を達成します。 DIVE の実験により、このフレームワークがブロックチェーン エコシステムの進化における 3 つの課題すべてに効果的に対処していることが確認されました。

原文 (English)

Frequency-Aware Continual Learning for Smart Contract Vulnerability Detection with Large Language Models

Smart contract vulnerability detection with Large Language Models (LLMs) faces three causally linked challenges. First, new vulnerability categories demand parameter-efficient adaptation, since full retraining is prohibitive for sequentially arriving tasks. Second, training per-task adapters on a shared backbone causes catastrophic forgetting of previously learned vulnerabilities. Third, the resulting multiplicity of adapters must be consolidated into a single model, since task identity is unknown at inference time. Each challenge arises directly from the solution to its predecessor, making an integrated framework essential. We propose a three-stage pipeline in which each stage addresses one challenge and feeds into the next. The adaptation stage uses Frequency-Aware Low-Rank Adaptation (FA-LoRA), which performs adaptation in the Fourier domain with per-frequency importance gates, requiring only 0.4% trainable parameters while outperforming standard LoRA and QLoRA. The continual learning stage applies Forget-Aware Replay (FAR), which uses these frequency gates to estimate per-sample forgetting risk via loss dynamics and prioritizes vulnerable knowledge for rehearsal, achieving an average Micro-F1 of 0.8022 across sequential tasks. The deployment stage employs Anchor-Protected Progressive Merging (APPM), which exploits the asymmetric generalization produced by FAR training to identify the strongest-generalizing adapter as an anchor and consolidates all adapters into a single model via anchor-protected weighted merging with frequency-domain gate competition. APPM achieves a Micro-F1 of 0.8085, within 2.7% of the independent per-task upper bound, at a merge cost of 156 ms and no additional runtime memory. Experiments on DIVE confirm the framework effectively addresses all three challenges for evolving blockchain ecosystems.

13:00 JSTロボティクス

オフライン品質多様性強化学習による階層型スキル ポリシーの学習

最近の研究では、事前に収集されたデータセットを活用してポリシーのパフォーマンスと RL のサンプル効率を向上させる方法が調査されています。この目標を達成するための有望なアプローチの 1 つは、2 段階の戦略を採用することです。第 1 段階では、特定のデータセットからさまざまなスキルが低レベルのポリシーとして抽出され、第 2 段階では、特定のタスクを解決するために高レベルのポリシーがトレーニングされます。通常、低レベルのポリシーの抽出は、軌跡 VAE などの教師なし学習に基づいて実行されます。ただし、このアプローチの制限は、低レベルのポリシーの品質がデータセットの品質に大きく依存することです。この問題に対処するために、オフラインからオンラインへの堅牢な学習のための統合パイプラインである QDOS (Quality-Diversity Offline Skill learning) を導入します。私たちのアプローチには、アドバンテージ加重品質ダイバーシティ事前トレーニング目標が組み込まれており、各軌道セグメントの推定アドバンテージによってスキル抽出とダイバーシティ目標に重み付けされます。このアプローチにより、モデルは多様で価値の高いスキルを抽出できます。 QDOS は、堅牢でタスクに関連したスキル表現を提供することにより、低レベルのポリシーで使用される組み込みスキル空間の品質を大幅に向上させます。さらに、これをデュアル データセット再利用戦略と統合します。この戦略では、オフライン データがスキルの事前トレーニングと、擬似ラベルによるオンライン リプレイ バッファーの入力の両方に使用されます。実験では、QDOS が構造化操作タスクと非構造化移動タスクにおいて強力なベースラインを大幅に上回るパフォーマンスを示し、困難な報酬が少ない領域で探査を加速し、最終収益を向上させる能力を裏付けています。

原文 (English)

Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning

Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. One promising approach to achieve this goal is to employ a two-stage strategy: In the first stage, diverse skills are extracted as a low-level policy from a given dataset, and a high-level policy is trained to solve a specific task in the second stage. Typically, extraction of the low-level policy is performed based on unsupervised learning such as trajectory VAE. However, a limitation of this approach is that the quality of the low-level policy highly depends on the quality of the dataset. To address this issue, we introduce QDOS (Quality-Diversity Offline Skill learning), a unified pipeline for robust offline-to-online learning. Our approach incorporates an Advantage-Weighted Quality-Diversity pretraining objective, which weights the skill extraction and diversity objectives by the estimated advantage of each trajectory segment. This approach allows the model to extract diverse and high-value skills. By providing robust and task-relevant skill representations, QDOS significantly improves the quality of the embedded skill space used by the low-level policy. We further integrate this with a dual dataset reuse strategy, where offline data is used both for skill pretraining and for populating the online replay buffer via pseudo-labeling. Experiments demonstrate that QDOS significantly outperforms strong baselines in structured manipulation tasks and unstructured locomotion tasks, confirming its ability to accelerate exploration and improve final returns in challenging sparse-reward domains.

13:00 JSTLLM/生成AIビジネス/資金調達

LLM ベースのソーシャル シミュレーションの評価と最適化を再考する

LLM ベースの社会シミュレーションは、調査や行動実験などの従来の手法を補完するものとして期待されています。中心となる問題は、LLM でシミュレートされた人間の行動の忠実性をどのように評価し、それに向けて LLM を最適化するかということです。一般的な手法では、モデルが人間から観察された単一の応答を選択しているかどうかをチェックして精度によって評価し、このハードラベルを再現するように LLM をトレーニングします。しかし、人間の行動は本質的に主観的です。同じ状況にある同じ人が合理的に異なる行動をする可能性があるため、観察された反応は基礎となる反応分布からの 1 つの描画にすぎず、精度に基づく評価は信頼性が低くなり、ハードラベル トレーニングは誤解を招きます。これらの問題に対処するために、私たちはまず、コーディングなどの客観的なタスクと社会シミュレーションなどの主観的なタスクを区別するエントロピーベースの量である主観係数を導入し、それを使用して、主観性が高まるにつれて精度ベースの評価とハードラベルトレーニングがどのように失敗するかを体系的に分析します。主観性係数に基づいて、主観性適応型ソフトラベル トレーニング (SALT) を提案します。SALT は、意味的に近くの入力から観測された出力をソフト分布ラベルにプールし、各入力の推定主観性に適応した集計半径を持ちます。目標に近い限界では近傍が縮小するため、SALT は自然に標準の単一ラベル トレーニングに戻ります。さらに、既存のデータセットは単一の観察された応答のみを記録し、分布評価をサポートできないため、193 人のアノテーターと 100 の主観的な質問をカバーする 19,300 のコンテキストのベンチマークである SUBJSIM を構築します。実世界のデータは通常、入力ごとに 1 つの観測値しか提供しないため、私たちの実験では、完全な応答分布に対してモデルを評価しながら単一の観測出力からモデルをトレーニングし、現実的な設定での実現可能性を検証します。 SUBJSIM の結果は、私たちの方法の利点を示しています。

原文 (English)

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments. A core question is how to evaluate the fidelity of LLM-simulated human behavior and optimize LLMs toward it. Prevailing practice evaluates by accuracy, checking whether the model selects the single response observed from a human, and trains the LLM to reproduce this hard label. However, human behavior is inherently subjective: the same person in the same situation may reasonably act differently, so an observed response is only one draw from an underlying response distribution, rendering accuracy-based evaluation unreliable and hard-label training misleading. To address these problems, we first introduce the subjectivity coefficient, an entropy-based quantity distinguishing objective tasks such as coding from subjective ones such as social simulation, and use it to systematically analyze how accuracy-based evaluation and hard-label training fail as subjectivity grows. Based on the subjectivity coefficient, we propose Subjectivity-Adaptive soft-Label Training (SALT): it pools observed outputs from semantically nearby inputs into soft distributional labels, with an aggregation radius adapted to the estimated subjectivity of each input; in the near-objective limit the neighborhood shrinks, so SALT naturally falls back to standard single-label training. Moreover, since existing datasets record only single observed responses and cannot support distributional evaluation, we construct SUBJSIM, a benchmark of 19,300 contexts covering 193 annotators and 100 subjective questions. Since real-world data typically provide only a single observation per input, our experiments train models from single observed outputs while evaluating them against the full response distributions, verifying feasibility in realistic settings. Results on SUBJSIM demonstrate the advantages of our method.

13:00 JSTエージェント

メモリマジョリティを超えて: マルチエージェントメモリアービトレーションのための潜在ソース推論

長期的なマルチエージェント システムは、さまざまなエージェントによって生成された記憶を継続的に蓄積します。既存の記憶方法は通常、取得した記憶を独立した証拠として扱い、投票または重み付けを通じてそれらを結合します。ただし、この独立性の仮定は、マルチエージェント設定では失敗することがよくあります。異なるエージェントによって書き込まれた記憶は、同じ上流ソースまたは共有バイアスを継承する可能性があり、その結果、相関関係のある証拠が繰り返しカウントされ、偽多数が作成される可能性があります。この故障モードを \textit{メモリ相関バイアス} と呼びます。この問題に対処するために、我々は、取得された記憶を共同で切り離し、失われた独立した証拠を回復する \textbf{C}orrelation-\textbf{A}ware \textbf{M}emory \textbf{A}rbitration (CAMA) フレームワークを提案します。取得した記憶をクエリ条件付き証拠グループとしてモデル化し、神経依存性推論と来歴ベースの記号事前確率を組み合わせて、独立した証拠ソースの有効数を推定することで、相関記憶が偽多数を形成するのを防ぎます。重要な独立した証拠が最初の取得セットに存在しない可能性があるため、\textsc{CAMA} はさらに、代替証拠を積極的に取得するか、最終決定を下す前に上流のソースを追跡する逐次回復ポリシーを学習します。これは、取得コストを最小限に抑えながら、信頼できる仲裁のために十分な独立した証拠を回復することを目的としています。複数のベンチマークでの実験により、相関記憶によって引き起こされる誤った多数決を抑制する、最先端のベースライン手法に対する本手法の優位性が実証されました。

原文 (English)

Beyond Memory Majority: Latent-Source Reasoning for Multi-Agent Memory Arbitration

Long-term multi-agent systems continuously accumulate the memories produced by different agents. Existing memory methods typically treat retrieved memories as independent evidence and combine them through voting or weighting. However, this independence assumption often fails in multi-agent settings: memories written by different agents may inherit the same upstream source or shared bias, causing correlated evidence to be repeatedly counted and creating a false majority. We term this failure mode \textit{Memory Correlation Bias}. To address the issue, we propose the \textbf{C}orrelation-\textbf{A}ware \textbf{M}emory \textbf{A}rbitration (CAMA) framework that jointly decouples retrieved memories and recovers missing independent evidence. We model the retrieved memories as query-conditioned evidence groups and combine neural dependency inference with provenance-based symbolic priors to estimate the effective number of independent evidence sources, thereby preventing correlated memories from forming a false majority. Since critical independent evidence may be absent from the initial retrieval set, \textsc{CAMA} further learns a sequential recovery policy that actively retrieves alternative evidence or traces upstream sources before making the final decision, aiming to recover sufficient independent evidence for reliable arbitration while minimizing retrieval cost. Experiments on multiple benchmarks demonstrate the superiority of our method over the state-of-the-art baseline methods, suppressing false majorities induced by correlated memories.

13:00 JST画像/動画生成エージェントロボティクス

SafeBranch: 実体化エージェントのためのブランチペアの安全調整

視覚言語モデルに基づく身体化エージェントは、指示されたタスクを完了できますが、その過程で安全制約に違反することが多く、この問題は最近対話型安全性として枠組み化されています。安全性とタスクの成功は別個の目的であり、安全性は軌道内の安全上重要な少数のステップでのみ発生するため、このようなエージェントを安全に行動するように訓練することは困難です。標準的な監視では不十分です。安全な軌道を模倣すると、それが安全である理由を説明せずに行動を教えることになり、任意の安全な軌道と安全でない軌道を対比させると、安全性のシグナルと無関係な差異が混同されます。我々は、環境ロールバックを介してアクター自身の安全でないロールアウトから構築されたブランチ ペアを通じて、実体化されたアクターを安全性に関して調整するフレームワークである SafeBranch を提案します。 SafeBranch は、各安全でないロールアウトを違反の原因となったセーフティクリティカルなステップにロールバックし、アクターに安全な代替案を問い合わせ、元のアクションと代替案をペアにして、2 つのブランチがそのステップでのみ異なるようにします。訓練されたアクターは、批判者が関与することなく、展開時に安全に動作します。 IS-Bench、SafetyALFRED、および目に見えないタスクとオブジェクトを含む配布外のバリアントでは、タスクの成功を犠牲にすることなく安全性を確実に処理し、目に見えないオブジェクトのバリアントでトレーニングされていないベースラインよりも約 10 倍の安全な成功を達成します。

原文 (English)

SafeBranch: Branch-Pair Safety Alignment for Embodied Agents

Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem recently framed as interactive safety. Training such agents to act safely is difficult, since safety and task success are distinct objectives, and safety arises only at a small number of safety-critical steps within a trajectory. Standard supervision is insufficient: imitating safe trajectories teaches behavior without explaining why it is safe, and contrasting arbitrary safe and unsafe trajectories mixes the safety signal with unrelated differences. We propose SafeBranch, a framework that aligns an embodied actor on safety through branch pairs constructed from the actor's own unsafe rollouts via environment rollback. SafeBranch rolls each unsafe rollout back to the safety-critical step that caused the violation, queries the actor for a safe alternative, and pairs the original action with the alternative so that the two branches differ only at that step. The trained actor acts safely at deployment with no critic in the loop. On IS-Bench, SafetyALFRED, and out-of-distribution variants with unseen tasks and objects, it handles safety reliably without sacrificing task success, achieving roughly ten times more safe successes than the untrained baseline on the unseen-object variant.

13:00 JST研究/論文

GenMatch: 配車サービスにおけるマイクロビュー注文配車のためのエンドツーエンドの生成マッチング フレームワーク

Micro-View Order-Dispatching は、各配車バッチ内で乗客の注文に対応可能なドライバーを割り当てます。これは、配車プラットフォームのサービス品質と運用効率にとって重要です。主流の産業用ソリューションは、モデル予測、値計算、およびディスパッチ マッチングの多段階パラダイムに従います。ディスパッチの品質は最終的なバッチレベルの割り当てによって決まりますが、これらの段階ではさまざまな中間目標が最適化されます。このステージ間の目標の不一致は、単一ステージを改善しても全体的なディスパッチ結果が必ずしも改善されるわけではないことを意味します。したがって、私たちは、Micro-View Order-Dispatching を生成マッチング問題として定式化し、エンドツーエンドの生成マッチング フレームワークであり、現実の運用環境に導入される最初のそのようなフレームワークである GenMatch を提案します。この問題に生成モデリングを適用すると、3 つの課題が生じます。まず、各ディスパッチ バッチは動的な疎な 2 部グラフを形成し、効率的な構造化されたバッチレベルのエンコーディングが必要になります。第 2 に、手作りの価値関数を置き換えるには、異種フィードバックから統一されたビジネス ユーティリティを学習する必要があります。第三に、選択された順序ドライバーのペアごとに残りの実行可能な候補が変更されるため、割り当てを直接生成するには、進化するマッチング状態を追跡する必要があります。 GenMatch は、コンテキスト認識型の 2 部構成エンコーダー、ビジネス認識型ユーティリティ学習器、および状態認識型ポインター デコーダーを使用して、これらの課題に対処します。 DiDi の国際配車市場の 5 都市で行われた広範なオフライン評価とオンライン A/B テストでは、競争ベースラインを上回る一貫した改善が示され、産業向け注文発送における GenMatch の有効性と実用性が確認されました。

原文 (English)

GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-Hailing

Micro-View Order-Dispatching assigns available drivers to passenger orders within each dispatch batch and is critical to the service quality and operational efficiency of ride-hailing platforms. Mainstream industrial solutions follow a multi-stage paradigm of model prediction, value calculation, and dispatch matching. Although dispatch quality is determined by the final batch-level assignment, these stages optimize different intermediate objectives. This cross-stage objective inconsistency means that improving a single stage does not necessarily improve the overall dispatch result. We therefore formulate Micro-View Order-Dispatching as a generative matching problem and propose GenMatch, an end-to-end Generative Matching framework and the first such framework deployed in a real-world production environment. Applying generative modeling to this problem introduces three challenges. First, each dispatch batch forms a dynamic sparse bipartite graph, requiring efficient structured batch-level encoding. Second, replacing the hand-crafted value function requires learning unified business utility from heterogeneous feedback. Third, directly generating an assignment requires tracking the evolving matching state because each selected order-driver pair changes the remaining feasible candidates. GenMatch addresses these challenges with a Context-Aware Bipartite Encoder, a Business-Aware Utility Learner, and a State-Aware Pointer Decoder. Extensive offline evaluations and online A/B tests in five cities across DiDi's international ride-hailing markets show consistent improvements over competitive baselines, confirming the effectiveness and practicality of GenMatch for industrial order-dispatching.

13:00 JST研究/論文

TT-net: 量子にインスピレーションを得た条件付き GAN のテンソル ネットワークのノイズ除去

量子アルゴリズムと量子多体システムの古典的なシミュレーションの主力として開発されたテンソル ネットワーク手法は、量子物理学の科学の主流になりました。さまざまなタイプのテンソル ネットワークの中でも、テンソル トレイン (量子コンピューティング コミュニティでは一般に行列積状態として知られています) は、すでに機械学習での用途を発見しています。これらの手法は、多くの場合、特異値分解 (SVD) と呼ばれる強力な線形代数ツールに依存します。画像ノイズ除去のためのいくつかの条件付き GAN アーキテクチャには、ジェネレーターの特徴マップに適用されるシングルカット分解ステップとして SVD が組み込まれています。この研究では、チャネルごとの SVD ノイズ除去ブロックを、クロスチャネル情報に直接アクセスできる 2 カットのテンソルトレイン分解に置き換える TT-Net を紹介します。この機能は、現代の代替手段には存在しません。この分解メカニズムのみが異なる制御された比較では、テストした 3 つのノイズ タイプ (ガウス、モーション ブラー、塩胡椒) のすべてにおいて、PSNR および SSIM において TT-Net が SVD-Net よりも優れており、クロスチャネル アクセスがノイズ除去品質を向上させるという仮説を裏付けています。さらに、トレーニング ダイナミクス分析では、TT-Net の敵対的損失項が、SVD-Net よりも 3 つのノイズ タイプすべてにわたって一貫して停滞状態に飽和している一方、再構築の品質は関係なく向上し続けており、この研究で特定されているものの解決されていない敵対的コンポーネントの寄与について未解決の疑問が生じていることが示されています。さらに、ガウス ノイズに関しては、私たちの方法は、EigenGAN と、線形代数分解を仮定せず、線形代数情報を保持しない最先端の Pix2pix 方法の両方よりも優れています。私たちの原稿では、量子にヒントを得たツールを深層学習アプリケーションの実用的な現実世界の特徴フィルターとしてどのように使用できるかを示します。

原文 (English)

TT-net: Quantum Inspired Tensor Network Denoising in Conditional GANs

Developed as a workhorse for classical simulations of quantum algorithms and quantum many-body systems, Tensor Network methods have entered the scientific mainstream in quantum physics. Among various types of tensor networks, Tensor Trains (commonly know as Matrix Product States in the quantum computing community) have already found applications in machine learning. These methods often rely on a powerful linear algebra tool called the Singular Value Decomposition (SVD). Several conditional GAN architectures for image denoising incorporate SVD as a single-cut decomposition step applied to generator feature maps. In this work we introduce TT-Net, which replaces the per-channel SVD denoising block with a two-cut tensor-train decomposition capable of accessing cross-channel information directly, a capability absent from contemporary alternatives. In a controlled comparison differing only in this decomposition mechanism, TT-Net outperforms SVD-Net on PSNR and SSIM across all three noise types tested (Gaussian, motion blur, and salt-and-pepper), supporting the hypothesis that cross-channel access improves denoising quality. Training-dynamics analysis further shows that TT-Net's adversarial loss term consistently saturates to a stagnant state across all three noise types, more so than SVD-Net's, while reconstruction quality continues to improve regardless, raising an open question about the adversarial component's contribution that this work identifies but does not resolve. Furthermore, for Gaussian noise our method outperforms both the EigenGAN and the state of the art Pix2pix method which does not assume any linear algebra decompositions and does not retain any linear algebra information. Our manuscript shows how quantum inspired tools can be used as practical real world feature filters for deep learning applications.

13:00 JSTLLM/生成AIビジネス/資金調達

有限プール材料最適化のための調達政策としての LLM: 対照研究

望ましい特性を持つ材料を発見するには、多くの場合、大きな候補空間を検索する必要がありますが、実験や計算による評価には依然としてコストがかかります。アクティブ ラーニングは、通常は確率的代理モデルを通じて、以前の観察を使用して次に評価する候補を選択することで、この課題に対処します。この設定では、オープンウェイト大規模言語モデル (LLM) がスタンドアロンの取得ポ​​リシーとして機能できるかどうかを調査します。さまざまな候補提示戦略の下で 4 つの遡及的有限プール材料最適化タスクにわたって 5 つの LLM を評価し、それらをランダム選択および従来のガウス プロセス手法と比較します。 LLM ポリシーは通常、ランダム選択よりも少ない反復で大域的最適値に到達します。これは、LLM ポリシーがタスク固有のトレーニングなしで有用な取得信号を提供することを示しています。ガウス プロセス手法と比較したパフォーマンスはさまざまです。従来の取得はほとんどのタスクでより優れたパフォーマンスを発揮しますが、LLM は一部の設定でそれと同等またはそれを上回るパフォーマンスを発揮します。パフォーマンスは、タスク、モデル、初期化、および候補者のプレゼンテーションによって大きく異なり、すべてのタスクにわたって最高のパフォーマンスを発揮する LLM アプローチはありません。全体として、オープンウェイト LLM は、有限プール材料検索の取得ポリシーとしての可能性を示していますが、その信頼性は依然として課題や候補と科学的背景の提示方法に左右されます。

原文 (English)

LLMs as Acquisition Policies for Finite-Pool Materials Optimization: A Controlled Study

Discovering materials with desirable properties often requires searching large candidate spaces while experimental or computational evaluations remain costly. Active learning addresses this challenge by using previous observations to select which candidate to evaluate next, typically through probabilistic surrogate models. We investigate whether open-weight large language models (LLMs) can serve as standalone acquisition policies in this setting. We evaluate five LLMs across four retrospective finite-pool materials optimization tasks under different candidate-presentation strategies and compare them with random selection and conventional Gaussian-process methods. LLM policies generally reach the global optimum in fewer iterations than random selection, indicating that they provide a useful acquisition signal without task-specific training. Their performance relative to Gaussian-process methods is mixed: conventional acquisition performs better on most tasks, while LLMs match or outperform it in some settings. Performance varies substantially across tasks, models, initializations, and candidate presentations, with no LLM approach performing best across all tasks. Overall, open-weight LLMs show potential as acquisition policies for finite-pool materials search, although their reliability remains sensitive to the task and to how candidates and scientific context are presented.

13:00 JSTLLM/生成AIエージェントロボティクス

一般的な身体化されたインテリジェンスに向けて: 大規模な言語モデル、知識ベース、および推論機能を統合して、次世代の AI エージェントを構築する

大規模言語モデル (LLM)、構造化知識ベース (KB)、および推論能力 (RA) の収束は、一般身体性知能 (GEI) への有望な軌道を示しています。この論文では、LLM を中心としたインテリジェント システムの進化を概説し、知識表現、論理的推論、物理的具体化との統合を強調します。私たちは、LLM アーキテクチャ、事前トレーニング方法、推論メカニズムを、外部知識ソースや構造化推論フレームワークとの相互作用とともに分析します。さらに、エージェントが物理環境で学習して行動する身体化知能 (EI) パラダイムを検証します。これらの側面を統合するために、LLM、KB、RA、および実施形態間の相乗効果を示す概念的なフレームワークを提示します。これは、実装されたエンジニアリング アーキテクチャではなく、認識、推論、およびアクションのガイド モデルとして機能します。 GEI に向けて前進するために、私たちは 5 つの主要な課題を特定します。それは、効率的な LLM の導入、閉ループの知識統合、ハイブリッドの記号と神経の推論、知覚と行動のグラウンディング、および継続的な学習です。この調査は、複雑で動的な設定で動作できる適応型のマルチモーダル エージェントを開発するための包括的なロードマップを提供します。

原文 (English)

Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents

The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI). This paper reviews the evolution of LLM-centered intelligent systems, emphasising their integration with knowledge representation, logical reasoning, and physical embodiment. We analyse LLM architectures, pre-training methods, and inference mechanisms, along with their interaction with external knowledge sources and structured reasoning frameworks. Furthermore, we examine embodied intelligence (EI) paradigms wherein agents learn and act in physical environments. To synthesise these dimensions, we present a conceptual framework that illustrates the synergy among LLMs, KBs, RA, and embodiment, serving as a guiding model for perception, reasoning, and action rather than an implemented engineering architecture. To advance toward GEI, we identify five key challenges: efficient LLM deployment, closed-loop knowledge integration, hybrid symbolic-neural reasoning, perception-action grounding, and continual learning. This survey provides a comprehensive roadmap for developing adaptive, multimodal agents capable of operating in complex, dynamic settings.

13:00 JST研究/論文

ADAPT: 適応型予測転送可能な HVAC 制御のための物理学を意識した拡散ベースの世界モデル

建物は世界のエネルギー消費と CO$_2$ 排出量の約 3 分の 1 を占めています。屋内気候システムの最適化は、国連の持続可能な開発目標 11 および 13 に沿った都市気候緩和にとって重要な役割を果たします。しかし、屋内の遅延熱力学的応答と部分的な観測可能性は、特に分布外環境では暗黙的な熱慣性、占有動態予測、および累積予測誤差によって主に制限される既存の手法を大きく妨げます。実際には、これらの課題は、高密度の屋内センシングによる高コストとプライバシーの負担によってさらに悪化し、管理者には目に見えない季節や気候領域にわたって確実に一般化することを期待しながら、単一の運用体制で限られたデータのみを収集することをオペレーターに強いています。この問題に対処するために、我々は、HVAC 制御のための物理学を意識した条件付き拡散屋内環境世界モデル ADAPT を提案します。このモデルは、建物の潜在的な熱慣性を捕捉するために、短地平線の持続動作熱ベースラインを予測します。拡散バックボーンは生成モデルの堅牢性を利用し、一方で学習可能なマルチゾーン熱バランス規則化装置は、既知の建物の形状や手動で校正された熱パラメータを必要とせずに、伝達可能な建物の熱力学を満たすように生成された軌道を制約します。その後、下流の強化学習用に単位の割り当てが設計されます。 SemibuildingSim と Sinergym に関する広範な実験により、ADAPT は、IID 制御下の最先端のベースラインと比較して、HVAC エネルギー消費を 7.3 \% 削減し、乗員の不快感を 30.2 \% 削減することが実証されました。目に見えない季節や気候領域にわたる OOD 制御シナリオの下で、ADAPT は IID パフォーマンスと比較してわずかな低下のみで堅牢なパフォーマンスを維持し、転送の堅牢性において既存の手法を大幅に上回ります。

原文 (English)

ADAPT: Physics-Aware Diffusion-based World Models for Adaptive Predictive Transferable HVAC Control

Buildings account for roughly one-third of global energy consumption and CO$_2$ emissions. Optimizing indoor climate systems plays a critical role for urban climate mitigation aligned with UN Sustainable Development Goals 11 and 13. However, indoor delayed thermodynamic responses and partial observability severely hinder existing methods, which are primarily limited by implicit thermal inertia, occupancy dynamic prediction, and cumulative prediction errors, especially for out-of-distribution environments. In practice, these challenges are further exacerbated by the high cost and privacy burden of dense indoor sensing, forcing operators to collect only limited data in a single operating regime while expecting controllers to generalize reliably across unseen seasons and climate regions. To address this problem, we propose ADAPT, a physics-aware conditional diffusion indoor environmental world model for HVAC control. The model predicts a short-horizon held-action thermal baseline to capture the latent thermal inertia of the buildings. The diffusion backbone utilizes the robustness of generative models, while a learnable multi-zone heat-balance regularizer constrains generated trajectories to satisfy transferable building thermodynamics without requiring known building geometry or manually calibrated thermal parameters. A credit assignment is then design for the downstream reinforcement learning. Extensive experiments on SemibuildingSim and Sinergym demonstrate that ADAPT reduces HVAC energy consumption by 7.3\% and occupant discomfort by 30.2\% compared with state-of-the-art baselines under IID control. Under OOD control scenarios spanning unseen seasons and climate regions, ADAPT maintains robust performance with only marginal degradation relative to its IID performance, substantially outperforming existing methods in transfer robustness.

13:00 JST画像/動画生成

「ノー」と言えばより良いビデオが作れる: 教育学的に根拠のある AI コンテンツ作成のためのデュアル ゲートキーピングの設計

美的には洗練されているが教育的に欠陥のある AI コンテンツの採用を防ぐために、私たちは 2 層の構造化された拒否を特徴とするビデオ オーサリング パイプラインを研究しています。最初の層では、教育者がマルチメディア学習理論に基づいて AI スクリプトを繰り返し再構築できるようにし、2 番目の層では、自動化されたメトリクスを採用して、指導の一貫性と物語と視覚の同期における違反にフラグを立てます。どちらのレイヤーも網羅的ではありませんが、それらの相乗効果により、原則的な抵抗、つまり厳格な基準を満たすまで AI の出力を延期する行為が、より高い品質への触媒となることが保証されています。 3つのトピックにわたる23人の教育者による研究と、確立された科学と哲学のカリキュラムから抽出された7つのトピックにわたる自動化された指標を組み合わせた評価では、両方の層が同じ指導面を独立して改善していることが示されており、思慮深い抵抗と生成AIは対立するものではなくパートナーであることが示唆されています。

原文 (English)

When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation

To prevent the adoption of aesthetically polished but pedagogically flawed AI content, we study a video authoring pipeline featuring two layers of structured refusal. The first layer empowers educators to iteratively reshape AI scripts based on multimedia learning theory, while the second employs automated metrics to flag violations in instructional coherence and narrative-visual synchronization. While neither layer is exhaustive, their synergy ensures that principled resistance--the act of deferring AI output until it meets rigorous standards--becomes a catalyst for higher quality. Evaluation combining a study with 23 educators across 3 topics and automated metrics across 7 topics drawn from established science and philosophy curricula shows that both layers independently improve the same instructional dimensions, suggesting that thoughtful resistance and generative AI are not opposites but partners.

13:00 JST研究/論文

二部構成のグラフィカル因果モデルによる因果推論

因果ベイジアン ネットワーク (CBN) と構造因果モデル (SCM) は、グラフィカルな因果推論の主要なフレームワークですが、現実世界のすべての因果システムを適切に表すことはできません。特に、平衡状態にあるシステム -- フィードバックメカニズムが循環的な因果依存関係を生み出す -- は、これらのフレームワークと根本的に互換性のない因果セマンティクスを示す可能性があります。つまり、同じ変数値を強制する異なる介入は異なる効果を持つ可能性があり、標準的な「完全な介入」 do($X = x$) が曖昧になってしまいます。我々は、方程式系の構造が変数ノードと方程式ノードを備えた二部グラフによってエンコードされる二部グラフ因果モデル (BGCM) を提案します。このフレームワークでは、ハード介入 do($f_j : X_v = \xi_v$) によって、どの方程式が置き換えられるか、どの変数がターゲットになるか、どの値が指定されるかが指定され、標準概念の曖昧さが解決されます。私たちは、物理システムの詳細なケーススタディを通じて、この表現が現実世界の介入に自然に対応していることを実証します。私たちは、方程式に固有の関数決定論を利用する新しいグラフィック分離基準 (B 分離) の観点からマルコフ特性を定式化し、それを非ランダム入力の設定に拡張します。これにより、ドメインの不変性を推論するための do-calculus がどのように生成されるかを示します。 BGCM は、グラフィカルな因果推論を実行する機能を保持しながら、CBN と SCM を厳密に一般化します。

原文 (English)

Causal Reasoning with Bipartite Graphical Causal Models

Causal Bayesian networks (CBNs) and structural causal models (SCMs) are the dominant frameworks for graphical causal reasoning, but they cannot adequately represent all real-world causal systems. In particular, systems at equilibrium---where feedback mechanisms create cyclic causal dependencies---can exhibit causal semantics that are fundamentally incompatible with these frameworks: different interventions that enforce the same variable value may have different effects, rendering the standard ``perfect intervention'' do($X = x$) ambiguous. We propose bipartite graphical causal models (BGCMs), in which the structure of a system of equations is encoded by a bipartite graph with variable and equation nodes. In this framework, a hard intervention do($f_j : X_v = \xi_v$) specifies which equation is replaced, which variable is targeted, and at what value---resolving the ambiguity of the standard notion. We demonstrate, through a detailed case study of a physical system, that this representation naturally corresponds to distinct real-world interventions. We formulate a Markov property in terms of a new graphical separation criterion (B-separation) that exploits the functional determinism inherent in the equations, and we extend it to settings with non-random inputs. We show how this gives rise to a do-calculus for reasoning about domain invariances. BGCMs strictly generalize CBNs and SCMs while retaining the ability to perform graphical causal reasoning.

13:00 JST研究/論文

仕様デルタ主導のデータ ガバナンス: レイクハウス データ プラットフォームの変更単位としての {\guillemotleft}spec-delta{\guillemotright} の実証研究

仕様駆動開発 SDD は、コードではなく仕様が AI 支援作業を管理する主要な成果物であるべきであるという考えを強化しました。 GitHub Spec Kit などのツールや、Constitutional SDD などの提案は、この原則をソフトウェア領域で形式化しましたが、実行可能なデータ契約の文献は、実行時のスキーマと品質の強制にまでこの原則を拡張しました。それにも関わらず、データ プラットフォームの変更の多くは純粋なコード変更ではなく契約によるもの (新しいデータセット、サービス レベル アグリーメント、メトリクス セマンティクス、アクセス ポリシー) であるにもかかわらず、すべての変更はデータ プラットフォームの変更単位としてレビュー可能な要件の増分を生み出すべきであるという OpenSpec の中心的な考え方である仕様デルタの扱いは、経験的に未調査のままです。この研究では、仕様デルタの概念を形式化し、増分仕様への適合性に応じたデータ プラットフォームの変更の分類を提案し、仕様デルタ駆動のワークフローとデルタのない従来のコード プル リクエスト ワークフローを比較する管理された実験を定義します。応答変数は、発見から展開までの時間、シルバーとゴールドのレイクハウス層に到達する欠陥の密度、ツール間のメトリックの相違、NASA TLX で測定されたレビュー担当者の認知負荷です。この論文では、実際の湖畔環境でのインスタンス化のために、デモンストレーションと実験室のセクションを明示的に確保しています。この貢献はツールではなく、再現可能な証拠であり、過剰な仕様のアンチパターンを回避するのに役立つ適用性ガイドです。

原文 (English)

Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms

Spec Driven Development SDD has consolidated the idea that the specification rather than the code should be the primary artefact governing AI assisted work. Tools such as GitHub Spec Kit, and proposals such as Constitutional SDD, have formalised this principle in the software domain, while the executable data-contracts literature has extended it to schema and quality enforcement at run time. Nevertheless, the treatment of the specification delta OpenSpec's core idea that every change should produce a reviewable increment of requirements as the unit of change in data platforms remains empirically unexplored, even though many data-platform changes are contractual (new datasets, service-level agreements, metric semantics, access policies) rather than purely code changes. This work formalises the spec-delta concept, proposes a taxonomy of data platform changes according to their suitability for incremental specification, and defines a controlled experiment comparing a spec-delta-driven workflow against a conventional code pull-request workflow without a delta. The response variables are discovery to deployment time, the density of defects reaching the Silver and Gold lakehouse layers, cross-tool metric divergence, and reviewer cognitive load measured with NASA TLX. The paper explicitly reserves a demonstration-and-laboratory section for instantiation on a real lakehouse environment. The contribution is not a tool but reproducible evidence and an applicability guide that helps to avoid the up front over specification antipattern.

13:00 JSTエージェント

SAPO: エージェント的強化学習のためのシングルロールアウト自己回帰ポリシー最適化

エージェント強化学習 (RL) は、大規模な言語モデルのトレーニング後の重要な段階になっています。既存の批判のないグループ相対手法は、複数のロールアウトからポリシーの利点を推定し、従来の近接ポリシー最適化 (PPO) の大幅なメモリ オーバーヘッドを回避し、長期的な対話型タスクで優れたパフォーマンスを実現します。その成功にもかかわらず、最近の研究では 3 つの限界があることが明らかになりました。 (2) 長期にわたる複雑なタスクにおいて潜在的な利点の崩壊に悩まされる。 (3) サンプリング予算と政策パフォーマンスの間でコストのかかるトレードオフが必要です。この研究では、ポリシー関数と値関数が単一の自己回帰バックボーンを共有する、低メモリで計算効率の高いフレームワークである、単一ロールアウト自己回帰ポリシー最適化 (SAPO) を提案します。 SAPO は、LLM の自己回帰構造を利用して、共有パラメータを使用して明確な因果境界でポリシーと価値の予測を生成すると同時に、PPO 目標と補助的なポリシー上の SARSA 目標を個別に最適化します。各ターンの寄与を確実に推定するために、ラムダリターンとバッチ正規化を組み合わせた軌道レベルの一般化利点推定器をさらに導入します。 Qwen2.5-1.5B/7B を使用した ALFWorld と WebShop の実験では、SAPO が安定してトレーニングし、PPO と GRPO をそれぞれ平均 +15.1 パーセント ポイントと +12.1 パーセント ポイント上回るパフォーマンスを示し、その一方で別個の批評家モデルのメモリ コストが排除され、PPO と比較して反復ごとのランタイムが 33.2% 削減されたことが示されています。

原文 (English)

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks. Despite their success, recent studies revealed three limitations: (1) Lack explicit value generalization and effective temporal credit assignment; (2) Suffer from potential advantage collapse in long-horizon complex tasks; (3) Require a costly trade-off between sampling budget and policy performance. In this work, we propose Single-rollout Autoregressive Policy Optimization (SAPO), a low-memory and compute-efficient framework in which the policy and value functions share a single autoregressive backbone. SAPO exploits the autoregressive structure of LLMs to produce policy and value predictions at distinct causal boundaries with shared parameters, while independently optimizing the PPO objectives and auxiliary on-policy SARSA objectives. To robustly estimate the contribution of each turn, we further introduce a trajectory-level generalized advantage estimator that combines lambda-returns with batch normalization. Experiments across ALFWorld and WebShop with Qwen2.5-1.5B/7B show that SAPO trains stably and outperforms PPO and GRPO by mean +15.1 and +12.1 percentage points, respectively, while eliminating the memory cost of a separate critic model and reducing per-iteration runtime by 33.2% over PPO.

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPTGemini

PolicyGuide: 1 つのアクションの保護から、ポリシーに準拠した LLM エージェントのワークフロー全体のガイドまで

カスタマーサービス LLM エージェントは、ユーザーの代わりに行動する場合、組織のポリシーに従う必要があります。コンプライアンス違反は、不適格な変更の許可などの禁止行為、または本人確認や確認などの手順要件の省略によって発生します。ランタイム セーフガードは危険なアクションに介入できますが、アクション ローカル チェックはエージェントを複数の手順に沿ってガイドしません。ワークフロー追従システムは、規定されたプロセスの実行をサポートしますが、主にエージェントの動作を保護するのではなく、ワークフローの完了を目的としています。代わりに、PolicyGuide は各ドメイン ポリシーをワークフロー グラフにコンパイルし、ユーザー ターンの境界でプロアクティブな検証ツールを呼び出します。ベリファイアは、永続化されたグラフ状態からオープンリクエストを調整し、ポリシーに準拠したパスに沿ってステップ固有の修復を返します。 GPT-5.4 エージェントと検証者を使用した $\tau^2$ ベンチの航空会社、小売業、通信ドメイン全体で、PolicyGuide は平均 $\mathrm{Pass}^4$ を $0.42$ から $0.62$ に引き上げ、最もワークフロー構造化されたドメインである通信 ($0.19$ から $0.61$) で最大の利益を得ました。同じワークフローが Claude Sonnet 4.6 および Gemini 2.5 Pro エージェントに転送されます。補完的な評価では、作成者が設計したワークフロー レベルの検証において、敵対的なユーザーの下で観察された最も低い攻撃成功率と、最も強力な手順順守が見つかります。

原文 (English)

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure. Workflow-following systems support prescribed process execution, but primarily target workflow completion rather than safeguarding agent behavior. PolicyGuide instead compiles each domain policy into a workflow graph and invokes a proactive verifier at user-turn boundaries. From persisted graph state, the verifier reconciles open requests and returns step-specific remediation along a policy-compliant path. Across the $\tau^2$-bench airline, retail, and telecom domains with a GPT-5.4 agent and verifier, PolicyGuide raises mean $\mathrm{Pass}^4$ from $0.42$ to $0.62$, with the largest gain on telecom ($0.19$ to $0.61$), the most workflow-structured domain. The same workflows transfer to Claude Sonnet 4.6 and Gemini 2.5 Pro agents. Complementary evaluations find the lowest observed attack-success rate under adversarial users and the strongest procedural compliance in an author-designed workflow-level validation.

13:00 JSTLLM/生成AIエージェント

EnvHarness: エージェント学習のための静的世界の覚醒

LLM エージェントは環境と対話することで学習しますが、これらの環境は手作業で構築された静的なものであり、エージェントの弱点には気付かず、改善されるとすぐに取り残されてしまいます。最近の環境生成方法はこれに対処しようとしていますが、ドメイン固有のパイプラインが必要で、高価な検証ツールまたは信頼性の低い検証ツールに依存しており、依然として静的な環境を生成します。環境を最初から再構築するエンジニアリングの負担を軽減するために、基礎となるロジックを変更せずに静的環境をラップしてその動作を再形成するプラグイン コンポーネントのプログラム可能な層である、Environment Harness (EnvHarness) を提案します。標準インターフェイスを通じて動作する EnvHarness は、さまざまなドメインに適用され、再形成されたすべての環境が元のベリファイアを保持するようにします。このプロセスを自動化するために、EnvRigger を導入します。EnvRigger は、ターゲット ポリシーをブラック ボックスとして扱い、その実行軌跡を観察して、診断された欠陥を対象とする EnvHarness コンポーネントを合成し、新たなロールアウトを通じてそれらを検証します。 4 つのドメインの 5 つのベンチマーク全体で、EnvHarness は元の環境とドメイン固有の環境生成パイプラインの両方を上回り、9.8% 少ない実行ステップでホールドアウトされたインスタンスで最大 9.0 ポイントの改善を達成しました。さらに、EnvHarness は強化学習に優れた最適化信号を提供し、ポリシーとその環境の継続的で的を絞った共進化を可能にします。

原文 (English)

EnvHarness: Awakening Static Worlds for Agent Learning

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.

13:00 JST研究/論文

TESTNAV: 組成堅牢性テストのためのパレートガイド検索

深層学習モデルは、特に同じ入力内で複数の破損が同時に発生する場合 (明るさの変化やモーション ブラーなど)、現実世界の入力の摂動に対して脆弱なままです。組成テストはこれらの相互作用効果を明らかにしますが、次の 2 つの課題をもたらします。それは、次元と重症度レベルが増加するにつれて摂動空間が組み合わせ的に増大することと、不均一な診断値です。多くの組み合わせでは、実用的な関連性が限られた非現実的に劣化した入力が生成されます。我々は、限られた数の摂動構成のみを評価できる場合に、離散的な組成摂動空間を効率的に探索するためのパレート誘導ロバスト性テスト フレームワークである TESTNAV 1 を紹介します。 TESTNAV は、堅牢性テストを二目的最適化として定式化することにより、重大かつ現実的な障害を優先します。つまり、モダリティ固有のメトリクス (視覚の場合は SSIM と KID、言語とコードの場合は chrF と BERT-F1) によって測定される入力忠実度を維持しながら、パフォーマンスの低下を最大化します。 NSGA-II を使用して、双対物パレート フロントを近似します。視覚、自然言語、コード生成にわたる 4 つのベンチマークにわたって、TESTNAV は、それぞれ 6 レベルの 4 つの摂動次元で定義される離散摂動空間の 35.8% ~ 89.3% を使用して、検索ベースのベースラインよりも最大 2.15 倍の速さでパレート フロントを回復します。

原文 (English)

TESTNAV: Pareto-Guided Search for Compositional Robustness Testing

Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e.g., brightness shifts and motion blur). Compositional testing reveals these interaction effects but introduces two challenges: combinatorial growth of the perturbation space as dimensions and severity levels increase, and uneven diagnostic value-many combinations yield unrealistically degraded inputs with limited practical relevance. We present TESTNAV, 1 a Pareto-guided robustness testing framework for efficiently exploring discrete, compositional perturbation spaces when only a limited number of perturbation configurations can be evaluated. TESTNAV prioritises severe yet realistic failures by formulating robustness testing as bi-objective optimisation: maximise performance degradation while preserving input fidelity measured by modality-specific metrics (e.g., SSIM and KID for vision; chrF and BERT-F1 for language and code). It uses NSGA-II to approximate the bi-objective Pareto front. Across four benchmarks spanning vision, natural language, and code generation, TESTNAV recovers Pareto fronts up to 2.15x faster than search-based baselines, using 35.8%-89.3% of the discrete perturbation space defined by four perturbation dimensions with six levels each.

13:00 JSTLLM/生成AI

一度書いたらどこでも実行: シェイプセーフでフレームワークに依存しない LLM アーキテクチャのための Axon DSL

オープンソース言語モデルのエコシステム全体は、事実上、単一のプラットフォームに依存しています。このプラットフォームが明日閉鎖を余儀なくされたらどうなるでしょうか?効率的なモデル定義の実装と維持、および異なるトレーニングおよび推論体制間での変換は、リソースを大量に消費するタスクであり、モデルの効率と移植性が大幅に制限され、スケーリングとデプロイメントの両方が妨げられます。ここでは、Haskell のような構文を備えた厳密に型指定されたドメイン固有言語である Axon を紹介します。これにより、LLM アーキテクチャのライトワンスでどこでも実行できるパラダイムが可能になります。 Axon は、特定のフレームワークのビジョンではなく言語仕様に基づいたコラボレーションを基盤とすることで、オープンな協力を促進し、研究者が最適化インフラストラクチャを放棄したり展開のロックインを受け入れたりすることなく、高度に専門化されたアーキテクチャを実装できるようにします。 Axon を使用すると、主要なフレームワーク (PyTorch、Triton を使用した PyTorch、JAX、MLX、および vLLM) のスタンドアロン実装に自動的にコンパイルできる、簡潔で監査可能な仕様が可能になります。 135M から 32B パラメーターの範囲のモデルに対する 467 の推論ベンチマーク実験では、Transformers のリファレンス実装と比較して、中央値で PyTorch で 7%、Triton を使用した PyTorch で 12%、JAX で 91%、MLX で 107% の高速化を実証しました。 PagedAttendant および KV キャッシュを備えたネイティブ vLLM アーキテクチャとしてデプロイされた場合、Axon モデルは、Transformers 実装と比較して中央値 58% の高速化を達成します。

原文 (English)

Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic LLM Architectures

The entire ecosystem of open-source language models effectively relies on a single platform. What if this platform was forced to shut down tomorrow? Implementing and maintaining efficient model definitions and translating them between different training and inference regimes is a resource-heavy task that severely limits model efficiency and portability, hindering both scaling and deployment. Here, we present Axon, a strongly typed domain-specific language with Haskell-like syntax, that enables a write-once, run everywhere paradigm for LLM architectures. By basing collaboration on a language specification rather than a specific framework's vision, Axon fosters open cooperation and empowers researchers to implement highly specialized architectures without giving up optimization infrastructure or accepting deployment lock-in. Axon allows for concise, auditable specifications that can be automatically compiled to standalone implementations for leading frameworks: PyTorch, PyTorch with Triton, JAX, MLX and vLLM. In 467 inference benchmarking experiments on models ranging from 135M to 32B parameters, we demonstrate median speedups of 7% on PyTorch, 12% on PyTorch with Triton, 91% on JAX, and 107% on MLX, compared to the reference implementations from Transformers. When deployed as native vLLM architectures with PagedAttention and KV-cache, Axon models achieve a 58% median speedup over Transformers implementations.

13:00 JSTロボティクス

EXIMO: VLM による VLA ポリシーのガイド付き探索

新しいタスクをその場で学習するためにロボットのポリシーを効率的に微調整するにはどうすればよいでしょうか?最先端のロボット操作ポリシーは、巨大な遠隔操作データセット上の数十億のパラメーターを備えた大規模なビジョン言語アクション (VLA) モデルの動作クローン作成に基づいています。この単純なアプローチによりロボット操作は大幅に進歩しましたが、新しいタスクを学習するための VLA ポリシーの微調整は依然として未解決の問題です。特に、遠隔操作データセットの収集には何百時間もの高価な人的労働が必要であり、代替手段である強化学習 (RL) は、特に長期的なタスクの場合、サンプル効率が悪いことで悪名高い可能性があります。さらに、VLA を使用した RL には、モデルのサイズとアーキテクチャ設計により、いくつかの課題が課せられます。この研究では、VLA ポリシーを微調整するための効率的なアルゴリズムである EXIMO を提案します。 EXIMO は、探索、模倣、最適化の 3 つの段階で動作します。探索フェーズでは、EXIMO はプランナーとして機能するビジョン言語モデル (VLM) を VLA に装備します。 VLM は、長期にわたる困難な問題を考え、VLA 向けに短い問題に分解します。 VLM は VLA とともに、新しいタスクに関する調整されたデータセットを収集するために使用されます。模倣フェーズでは、VLA は調整されたデータを使用して微調整されます。最後に、最適化段階で、残留オフポリシー RL を使用してポリシーをさらに微調整します。私たちの実験では、EXIMO の 3 つの段階すべてをアブレーションし、サンプル効率と最終パフォーマンスの点で既存のアプローチよりも大幅に優れていることを示しました。

原文 (English)

EXIMO: VLM Guided Exploration of VLA Policies

How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. While this simple approach has enabled significant advances for robotic manipulation, finetuning of VLA policies for learning new tasks still remains an open problem. In particular, collecting teleoperation datasets requires hundreds of hours of expensive human labour and the alternative, reinforcement learning (RL), can be notoriously sample-inefficient especially for long-horizon tasks. In addition, RL with VLAs imposes several challenges due to the model's size and architectural design. In this work, we propose EXIMO, an efficient algorithm for finetuning of VLA policies. EXIMO operates in three stages: explore, imitate, and optimize. During the explore phase, EXIMO equips the VLA with a vision language model (VLM) that acts as a planner. The VLM thinks and breaks down challenging long-horizon problems into shorter ones for the VLA. The VLM, together with the VLA, is used to collect an orchestrated dataset on new tasks. During the imitate phase, the VLA is finetuned with the orchestrated data. Finally, during the optimize stage, we use residual off-policy RL to further finetune the policy. In our experiments, we ablate all three stages of EXIMO and show that it outperforms existing approaches significantly in terms of sample-efficiency and final performance.

13:00 JSTエージェントハードウェア/半導体研究/論文

科学用エージェント AI に分析の厳密性をもたらす: 神経画像データ分析用の Brain Researcher プラットフォーム

AI エージェントは科学的分析を実行できますが、分析結果が擁護可能な主張となるのは、代替案が検討され、その主張が証拠が裏付ける内容に限定された場合のみです。エージェントは、選択的な分析、時期尚早の成功宣言、不完全な基準の最適化などの失敗を再現する可能性があります。 Brain Researcher は、許容される分析、必要なチェック、クレーム範囲のルールに基づいて、神経画像研究者の計算環境で動作するエージェント研究用ハーネスです。ベンチマークでは、Brain Researcher は 7 つのモデルにわたって最初に選択するツールの選択精度を 70.2 パーセント ポイント (ツールなしでは 23.3%、ツールありでは 93.6%) 向上させ、検証可能な根拠を 4.6% から 22.0% に向上させました。共同研究者主導の自己進化的研究では、多元的分析により分析選択の敏感さが明らかになり、科学的レビューにより主張が受理、適格、修正、ブロック、拒否、延期に分類されました。 Brain Researcher は、意思決定を証拠と出所に結び付けることで、方法論的な判断をワークフローの後ではなくワークフロー内に埋め込みます。

原文 (English)

Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis

AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment under rules for admissible analyses, required checks and claim scope. In benchmarks, Brain Researcher increased first-choice tool-selection accuracy across seven models by 70.2 percentage points (23.3% without it versus 93.6% with it) and verifiable grounding from 4.6% to 22.0%. In collaborator-led and self-evolving studies, multiverse analyses exposed analytic-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it.

13:00 JST研究/論文

非線形力学システムにおけるスパイクベースの信念伝播

この論文では、適応制御のためのスパイクベースのダイナミクスと確率的推論を統合するベイジアン制御フレームワークを紹介します。ベイズ推論は、脳機能の中核となる計算原理として広くみなされており、不確実性の下での知覚、意思決定、学習に規範的な枠組みを提供します。生物学的にインスピレーションを得たスパイキングニューラルモデルとベイジアン推論原理を組み合わせることで、不確実な環境でも動作できる脳のような制御アルゴリズムを提案します。非線形ダイナミクスのベンチマークとして山の駐車場問題を使用します。私たちの結果は、提案されたコントローラーがリアルタイムで状態を正常に更新し、スパイク駆動のダイナミクスを通じて目標指向のアクション プランを生成できることを示しています。この結果は、計算神経科学と確率的制御理論の間の架け橋として、提案されたモデルの可能性を強調しています。

原文 (English)

Spike-based Belief Propagation in Nonlinear Dynamical Systems

This paper presents a Bayesian control framework that integrates spike-based dynamics with probabilistic inference for adaptive control. Bayesian inference is widely regarded as a core computational principle of brain function, providing a normative framework for perception, decision-making, and learning under uncertainty. By combining a biologically inspired spiking neural model with Bayesian inference principles, we propose a brain-like control algorithm capable of operating in uncertain environments. We use the mountain car parking problem as a benchmark with non-linear dynamics. Our results demonstrate that the proposed controller can successfully update states in real time and generate goal-directed action plans through spike-driven dynamics. The results highlight the proposed model's potential as a bridge between computational neuroscience and probabilistic control theory.

13:00 JST研究/論文

オープンイレブン構造統計形状モデルを使用した、CT 上の心臓全体の心臓形状を完成させるための強力な線形ベースライン

公開心臓コホートは心臓のさまざまなサブセットに注釈を付けるため、対応関係を共有しない限り、別々のソースからの形状をプールすることはできません。公開されている心臓形状リソースの中で、心耳、肺静脈、および大腿断端を 1 つのメッシュ内の別個のブロックとして保持しているものは特定されませんでした。完了ベンチマークでは、同じ適合モデルが意味する条件付き推定量ではなく、形状モードへの最小二乗投影に対してディープ モデルも比較されます。我々は、11 571 頂点対応で自動的にラベル付けされた 383 症例から構築された 11 構造の心臓コンピュータ断層撮影 (CT) 統計形状モデルをリリースし、1 つの凍結された内部分割とエンドポイントの下で完了推定量を比較します。フィッティングから除外された 76 ケースの内部リスト上で、閉形式の条件付きガウス推定器は、1 つ、3 つ、5 つ、および 9 つの観察された構造にわたって均等に平均された、頂点あたりの平均誤差 3.717 mm で欠落した非腔構造を再構築しました。 5 回の再フィットマスク条件付きグラフ変分オートエンコーダは 5.248 mm、最近傍検索は 8.931 mm に達しました。ペアの差は 1.531 mm (95% 信頼区間 1.384 ~ 1.711) で、順序は生の座標感度アームで保持されました。専門家による手動ラベルは 58 件の外部 CT 症例に対して存在しますが、登録済みの参照は 5 つの構造のみをスコアリングするのに十分近いものです。そこでは、閉じた形式の推定器の平均表面距離、95 パーセンタイルのハウスドルフ距離、および両方の完成した心房の面取り誤差が再び低くなりました。 20 件のケースからなる 2 番目の公開ベンチマークでは、完成した 4 つの構造物のうち 3 つについて基準が十分近く、同じ順序が維持されました。 4 つの構造には専門家の参照先がありません。リリースされたモデルとその補完オペレーターは、臨床使用ではなく、位置合わせされた CT に関するコホート統合研究をサポートします。

原文 (English)

A Strong Linear Baseline for Whole-Heart Cardiac Shape Completion on CT, with an Open Eleven-Structure Statistical Shape Model

Public cardiac cohorts annotate different subsets of the heart, so shapes from separate sources cannot be pooled without shared correspondence. Among released cardiac shape resources, none we identified carries the atrial appendage, pulmonary veins, and caval stumps as separate blocks in one mesh. Completion benchmarks also compare deep models against a least-squares projection onto shape modes, not the conditional estimator the same fitted model implies. We release an eleven- structure cardiac computed-tomography (CT) statistical shape model, built from 383 automatically labelled cases in 11 571-vertex correspondence, and compare completion estimators under one frozen internal split and endpoint. On a 76-case internal list held out from fitting, a closed-form conditional-Gaussian estimator reconstructed the missing non-chamber structures at 3.717 mm mean per-vertex error, averaged equally over one, three, five, and nine observed structures. A five-refit mask-conditioned graph variational autoencoder reached 5.248 mm and nearest-neighbour retrieval 8.931 mm. The paired difference was 1.531 mm (95% confidence interval 1.384 to 1.711), and the ordering held in a raw-coordinate sensitivity arm. Expert manual labels exist for 58 external CT cases, but our registered reference is close enough to score only five structures. There the closed-form estimator again had lower average surface distance, 95th-percentile Hausdorff distance, and Chamfer error for both completed atria. On a second public benchmark of 20 cases the reference was close enough for three of four completed structures, and the same ordering held there. Four structures have no expert reference. The released model and its completion operator support cohort-unification research on aligned CT, not clinical use.

13:00 JST研究/論文

MILP を加速するための初期から最終までのソリューションの一貫性を学習する

混合整数線形計画法 (MILP) は、オペレーションズ リサーチと組み合わせ最適化における基本的な問題クラスであり、産業上の意思決定に幅広く応用されています。ただし、NP の難易度により、現代のソルバーは、現実的な時間制限内で、困難な MILP インスタンスに対する高品質のソリューションを見つけるのに苦労する可能性があります。最近の学習ベースのアプローチは、変数制約の 2 部グラフなどの静的なインスタンス レベルの特徴から高品質の解を直接予測することにより、MILP 解決を高速化しようとしています。しかし、インスタンスの特徴だけから正確な解を予測することは難しく、これらの方法では、ソルバーの検索プロセス中に明らかにされる情報がほとんど見落とされます。この論文では、MILP ソルバーの初期の探索段階で生成された解は、計算的に安価に入手できるが、多くの場合、フルバジェット探索後に見つかった解に構造的に近いことがわかりました。この観察に動機づけられて、私たちは、学習目標を変数割り当てから初期から最終までの一貫性へと移行する、ソルバー情報に基づいた新しいパラダイムを提案します。変数ごとに、その初期段階の割り当てが予算全体のソリューションで持続するべきかどうかを予測します。予測された一貫性は、たとえば一貫性があると思われる割り当てを修正することによって、下流の検索を自然に導きます。推論時には、複数の初期段階のソリューションにわたるアンサンブル一貫性予測をさらに実行して、堅牢性を向上させます。 4 つの MILP ベンチマークにわたる実験では、私たちの方法がさまざまな下流パイプラインにわたる予測に基づく検索を改善することが示されています。 Gurobi を使用すると、私たちが提案した方法は、プライマリ ギャップを平均 56.9% 削減し、組み合わせオークション インスタンスで完全に閉じます。さらに、Gurobi でトレーニングされたモデルのゼロショットを適応せずに SCIP に転送し、ベンチマーク間で平均 36.4% のギャップ削減を達成しました。

原文 (English)

Learning Early-to-Final Solution Consistency for MILP Acceleration

Mixed-Integer Linear Programming (MILP) is a fundamental problem class in operations research and combinatorial optimization, with broad applications to industrial decision-making. Owing to their NP-hardness, however, modern solvers may struggle to find high-quality solutions for challenging MILP instances within practical time limits. Recent learning-based approaches seek to accelerate MILP solving by directly predicting high-quality solutions from static instance-level features, such as variable-constraint bipartite graphs. Yet accurate solution prediction from instance features alone is difficult, and these methods largely overlook the information revealed during the solver's search process. In this paper, we find that solutions produced at the early search stage of MILP solvers, which are computationally cheap to obtain, are often structurally close to the solutions found after full-budget search. Motivated by this observation, we propose a new solver-informed paradigm that shifts the learning target from variable assignment to early-to-final consistency: for each variable, we predict whether its early-stage assignment should persist in full-budget solutions. The predicted consistency naturally guides downstream search, for instance by fixing the assignments deemed consistent. At inference time, we further ensemble consistency predictions across multiple early-stage solutions to improve robustness. Experiments across four MILP benchmarks show our method improves prediction-guided search across diverse downstream pipelines. With Gurobi, our proposed method reduces the primal gap by 56.9% on average and closes it completely on combinatorial auction instances. Besides, we transferred the Gurobi-trained model zero-shot to SCIP without adaptation, achieving a 36.4% average gap reduction across benchmarks.

13:00 JST研究/論文

セマンティック構造化パーティショニングを使用したパッチベースの多変量時系列予測の再考

多変量時系列予測 (MTSF) は、現実世界の多くのアプリケーションにおける基本的なタスクです。既存のパッチベースの予測方法は、一般に、固定パーティショニング、マルチスケール パーティショニング、および拡張可能なパーティショニングの 3 つのカテゴリに分類されます。固定パーティショニングは意味のある時間境界を壊すことが多く、マルチスケール パーティショニングはスケール間で冗長な表現を導入する可能性があり、拡張可能なパーティショニングは柔軟性を向上させますが、セマンティック構造を組織化し、異種の時間パターン間の相互作用をモデル化するための明示的なメカニズムがまだ欠けています。これらの制限に対処するために、セマンティック構造化パーティショニングに基づいて構築された Transformer ベースのフレームワークである SCPaT を提案します。 SCPaT は、まず適応型セマンティック ユニット生成を通じて入力シーケンスをセマンティックに一貫したユニットに分解し、次に動的セマンティック グラフを構築してこれらのユニット間の有向依存関係をモデル化し、それらを高次のセマンティック ブロックに編成します。これらの構造化表現に基づいて、重要性を認識したルーティング メカニズムが、カスタマイズされたモデリングのためにさまざまなセマンティック ブロックをさまざまなエキスパートに適応的にディスパッチします。 12 の実世界のデータセットに対する広範な実験により、SCPaT の有効性が実証されました。

原文 (English)

Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning

Multivariate time series forecasting (MTSF) is a fundamental task in many real world applications. Existing patch based forecasting methods generally fall into three categories: fixed partitioning, multi-scale partitioning, and extendable partitioning. Fixed partitioning often breaks meaningful temporal boundaries, multi-scale partitioning may introduce redundant representations across scales, and extendable partitioning improves flexibility but still lacks an explicit mechanism for organizing semantic structure and modeling interactions among heterogeneous temporal patterns. To address these limitations, we propose SCPaT, a Transformer based framework built on semantic structured partitioning. SCPaT first decomposes input sequences into semantically consistent units through adaptive semantic unit generation, then constructs a dynamic semantic graph to model directed dependencies among these units and organize them into higher order semantic blocks. Based on these structured representations, an importance aware routing mechanism adaptively dispatches different semantic blocks to different experts for customized modeling. Extensive experiments on 12 real world datasets demonstrate the effectiveness of SCPaT.

13:00 JSTLLM/生成AIエージェントGeminiDeepSeek

ReguSim: 財務コンプライアンスに基づいた LLM エージェント ルールの評価

金融市場のLLMエージェントは、ルールを引用しながらも、実行可能な制約に違反したり、監視証拠を読み間違えたりする注文を提出する可能性があります。管理された財務コンプライアンス環境である ReguSim と、目標マーク付き監視ベンチマークである ReguBench を導入して、述べられた推論、試行されたアクション、実行の執行、証拠の監視という 4 つの成果物を分離します。 DeepSeek V4 Pro および Gemini 3.5 Flash を使用したトレーダーの実行では、可視ルールにより拒否されたアクションが減少しますが、排除されず、インセンティブまたはペルソナ フレーミングの動作が変化します。ブリッジ調査では、執行証拠が示されない限り、トレーダーの論理的根拠が独立した監視者を誤解させる可能性があることが示されています。モニタリングでは、単純な構造化ベースラインはプロンプトのみの LLM と一致するか、それを上回ります。この結果は、財務コンプライアンスの評価を、単一のコンプライアンス スコアではなく、ルールに基づいた行動と証拠の使用の監査として枠組み化しています。

原文 (English)

ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.

13:00 JSTLLM/生成AIエージェント

証明可能な二基準保証による LLM エージェントの最適なスキル選択

再利用可能なスキル ドキュメントを制限されたコンテキスト ウィンドウにロードすることは、現在、大規模言語モデル (LLM) エージェントがタスク固有の機能を取得する主な方法となっており、スキルの選択がタスクのパフォーマンスとトークン コストの一次決定要因となります。しかし、現在のエージェントは意味論的な関連性によって独自にスキルをスコアリングし、選択したセットの品質保証やコスト意識を持たずに、トップ千ドルまたは貪欲なパッキングによってセットを組み立てます。その結果、冗長なスキルや適切に選択されていないスキルは、希少なコンテキスト トークンを無駄にし、パフォーマンスを低下させる可能性さえあります。選択したスキル セットがどのように実行結果を形成し、最適化問題としてスキルの選択を行うかについての最初のモデルを示します。ハード トークンの予算内でスキル セットを選択し、単調なサブモジュールのメリットからコンテキスト ペナルティを差し引いたものを最大化します。この問題に対して、多項式時間アルゴリズムであるベスト プレフィックス選択 (BPS) を開発し、私たちの知る限り、スキル選択の最初のパフォーマンス保証である多項式時間で最適な利益係数を持つ双基準 $(1-1/e,1)$ 近似を証明しました。汚染管理された BigCodeBench 亜種では、BPS はすべてのベースラインを上回っており、最も強力なリリース済みルーターよりも $28\%$ 少ないトークンで、測定されたタスクの成功が $0.73$ であるのに対し、リリースされたスキル ルーター、テキスト取得者、および実行者独自の選択では $0.20$ ~ $0.52$ に達しています。

原文 (English)

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance and token cost. Yet current agents score skills independently by semantic relevance and assemble the set by top-$k$ or greedy packing, with no quality guarantee or cost awareness on the selected set. As a result, redundant or poorly chosen skills waste scarce context tokens and can even degrade performance. We give the first model of how the selected skill set shapes execution outcomes and cast skill selection as an optimization problem: choose a skill set under a hard token budget to maximize a monotone submodular benefit minus context penalty. For this problem, we develop Best Prefix Selection (BPS), a polynomial-time algorithm, and prove, to our knowledge, the first performance guarantee for skill selection: a bicriteria $(1-1/e,1)$ approximation whose benefit coefficient is optimal in polynomial time. On a contamination-controlled BigCodeBench variant, BPS outperforms all the baselines, reaching $0.73$ measured task success versus $0.20$--$0.52$ for released skill routers, text retrievers, and the executor's own selection, on $28\%$ fewer tokens than the strongest released router.

13:00 JST研究/論文

ExPhy: 複数物体軌道予測における明示的物理特性学習のベンチマーク

物体のダイナミクスを理解するには、将来の軌道を予測するだけでなく、モデルが動きを支配する物理的特性を捉えているかどうかを調べる必要があります。ただし、既存のベンチマークでは、軌道予測と並んでオブジェクトレベルの物理的特性が明示的な評価対象として公開されることはほとんどありません。このギャップに対処するために、\emph{ExPhy} を導入します。これは、質量、摩擦、反発の明示的なオブジェクト レベルのラベルを持つ 24,000 のシミュレートされた物理シーンを含む、複数オブジェクトの軌道予測ベンチマークです。 ExPhy は、軌道予測と物理特性推定を共同で評価するために、物理パラメータ (OOD-Parameter) と初期状態 (OOD-Initial) にわたる分布内 (ID) 分割と 2 つの分布外 (OOD) 分割とともに、観測された軌道と将来の軌道を提供します。さらに、観測された軌道から物理特性を推定し、微分可能な将来のロールアウトに使用する明示的なプロパティ インターフェイスを備えた物理ガイド付きモデルである \textsc{PhyODE} をインスタンス化します。長期的な OOD 初期設定では、\textsc{PhyODE} は、最も強いベースラインと比較して、ADE と FDE をそれぞれ 33.1\% と 31.0\% 削減します。 ComPhy でのゼロショット評価は、クロスベンチマーク移行をさらに評価します。特性レベルの分析により、正確な軌道予測が必ずしも基礎となる物理特性の正確な回復を意味するわけではないことが明らかになりました。コードとデータは https://github.com/Zest86/ExPhy で入手できます。

原文 (English)

ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting

Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical properties that govern motion. However, existing benchmarks rarely expose object-level physical properties as explicit evaluation targets alongside trajectory forecasting. To address this gap, we introduce \emph{ExPhy}, a multi-object trajectory forecasting benchmark containing 24,000 simulated physical scenes with explicit object-level labels for mass, friction, and restitution. ExPhy provides observed and future trajectories together with an in-distribution (ID) split and two out-of-distribution (OOD) splits over physical parameters (OOD-Parameter) and initial states (OOD-Initial) for jointly evaluating trajectory forecasting and physical property estimation. We further instantiate \textsc{PhyODE}, a physics-guided model with an explicit property interface that estimates physical properties from observed trajectories and uses them for differentiable future rollout. On the long-horizon OOD-Initial setting, \textsc{PhyODE} reduces ADE and FDE by 33.1\% and 31.0\%, respectively, compared with the strongest baseline. Zero-shot evaluation on ComPhy further assesses cross-benchmark transfer. Property-level analyses reveal that accurate trajectory forecasting does not necessarily imply accurate recovery of the underlying physical properties. Code and data are available at https://github.com/Zest86/ExPhy.

13:00 JST画像/動画生成

フロー設定の最適化における多様なドリフト: リワードハッキングの根本原因

優先度の最適化は生成モデルの標準的な調整方法ですが、それを連続時間ダイナミクスに拡張することは依然として簡単ではありません。フローマッチングでは、報酬駆動型の更新により、事前トレーニングされたデータマニホールドに対する固有の制約なしでトランスポート軌道が変更され、端末サンプルを事前トレーニングされたサポートから移動できます。この故障モードをマニホールドドリフトとして形式化します。理論的には、最適なフローマッチングは端末データ分布を回復しますが、優先更新は、誘導された端末変位がゼロ以外の正規成分を持つ場合は常に事前学習された多様体を離れることを示します。解決策として、私たちは、優先サンプルに対するペアごとの優先最適化を固定する、温度制御された対物レンズである ThermoDPO を提案します。この目的は、温度領域全体にわたって、リジェクション サンプリングの微調整と FlowDPO を接続し、多様体距離に対する点単位の再構成ベースの代理を制御します。低温での信号の減少に対抗するために、加重バリアントである ThermoDPO 加重をさらに導入します。主要な玩具ベンチマークでは、ThermoDPO 加重は 0.899 の StrictScore を達成しました。これに対し、FlowDPO では 0.629、FlowDPO+RFT では 0.857 でした。 CFG = 4.5 の SD3.5-M では、OCR が 47.5% 向上し、4 つのメトリクスの平均が 16.0% 向上しました。

原文 (English)

Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretrained support. We formalize this failure mode as manifold drift. Theoretically, we show that optimal flow matching recovers the terminal data distribution, whereas a preference update leaves the pretrained manifold whenever its induced terminal displacement has a nonzero normal component. As a remedy, we propose ThermoDPO, a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO and controls a pointwise reconstruction-based surrogate for manifold distance. To counteract diminished signals at low temperatures, we further introduce a weighted variant, ThermoDPO-weighted. On the main toy benchmark, ThermoDPO-weighted attains a StrictScore of 0.899, compared with 0.629 for FlowDPO and 0.857 for FlowDPO+RFT. On SD3.5-M at CFG = 4.5, it improves OCR by 47.5% and the average of four metrics by 16.0%.

13:00 JSTLLM/生成AI

目に見えないモダリティの組み合わせによる不完全なマルチモーダル感情分析のための対照的混合プロンプト学習

不完全なマルチモーダル感情分析は、近年大きな注目を集めています。既存のアプローチは通常、データがランダムに欠損しているか、特定の欠損パターンに特化して設計されていると想定しており、トレーニング フェーズとテスト フェーズの間のモダリティの組み合わせの不一致を無視しています。ただし、現実のシナリオでは、テスト フェーズでは、トレーニング フェーズには存在しなかったモーダルの組み合わせが頻繁に発生し、一般化機能が不十分になり、パフォーマンスが不安定になります。この論文では、目に見えないモダリティの組み合わせに対するモデルの一般化を強化することを目的として、目に見えないモダリティの組み合わせによる不完全マルチモーダル感情分析 (IMSAUMC) の問題を紹介します。この課題に対処するために、IMSAUMC に対して $\textbf{C}$ontrastive $\textbf{M}$ixed $\textbf{P}$rompt $\textbf{L}$earning ($\textsf{CMPL}$) というモデルを提案します。これは、堅牢で識別力のあるクロスモーダル表現を学習するための、ラベルに基づく対照的な特徴学習メカニズムを導入します。さらに、さまざまなモダリティの組み合わせの学習を促進するために、ソフト ルーターを使用してモダリティの組み合わせのプロンプトを設計します。さらに、3 つのプロンプト対比学習戦略を導入します。これにより、目に見えないモダリティの組み合わせに対応するプロンプトの効果的な学習が可能になり、それによって、さまざまなテスト シナリオにおけるモデルの一般化機能が大幅に強化されます。広く使用されている 3 つのデータセットに対する広範な実験により、$\textsf{CMPL}$ は最先端のアプローチと比較して 5% 以上の精度向上を達成できることが実証されました。

原文 (English)

Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that were not present during the training phase, which leads to insufficient generalization capabilities and unstable performance. In this paper, we introduce the problem of Incomplete Multimodal Sentiment Analysis with Unseen Modality Combinations (IMSAUMC), aiming to enhance model generalization for unseen modality combinations. To address this challenge, we propose the model named $\textbf{C}$ontrastive $\textbf{M}$ixed $\textbf{P}$rompt $\textbf{L}$earning ($\textsf{CMPL}$) for IMSAUMC. It introduces a label-guided contrastive feature learning mechanism to learn robust and discriminative cross-modal representations. Additionally, we design modality-combination prompts with a soft router to facilitate better learning of various modality combinations. Furthermore, we introduce three prompt contrastive learning strategies, which enable effective learning of prompts corresponding to unseen modality combinations, thereby significantly strengthening the model's generalization capabilities in diverse testing scenarios. Extensive experiments on three widely used datasets demonstrate that $\textsf{CMPL}$ achieves more than a 5% improvement in accuracy compared to state-of-the-art approaches.

13:00 JST研究/論文

高度な AI システムの代理店の 3 次元類型学

高度な人工知能 (AI) システムのエージェンシーに関する研究は、規範的な概念としてのエージェンシーと、特にエージェント的な AI システムのエージェンシーに焦点を当てています。最近の研究では、エージェント システムのさまざまなプロファイルにも焦点を当てていますが、特に非道徳的な形式のエージェントを考慮する場合、高度な AI システムによってインスタンス化されるエージェントの種類の問題に対処する枠組みは存在しません。哲学、倫理、法理論、社会学における確立された理論的立場に基づいて、私たちは、エージェンシーの性質(道徳的または法的)、そのモード(個人または集団)、およびその場所(人間または非人間)の 3 つの側面から構成されるフロンティア AI システムのエージェンシーの類型を開発します。これらの側面を組み合わせると、エージェンシーの 8 つの可能なインスタンスが生成されます。これらは、従来のもの、論争のあるもの、または物議を醸しているものとして分類されます。この類型論は、法的主体と道徳的主体を分離することで、高度な AI システムが道徳的主体であることを前提とせずに、個人の法的非人間的主体を検討するための概念的空間を作り出します。私たちは、この区別は、手段的な目標追求によって AI の行動の帰属が特定の人間の行為者に帰属することが複雑になる場合に、ますます重要になると主張します。

原文 (English)

A three-dimensional typology of agency for advanced AI systems

Research on the agency of advanced artificial intelligence (AI) systems focuses on agency as a normative concept and on the agency of particularly agentic AI systems. While recent work also focuses on the different profiles of agentic systems, no framework exists to address the question of the type of agency instantiated by advanced AI systems, particularly when considering non-moral forms of agency. Based on established theoretical positions in philosophy, ethics, legal theory and sociology, we develop a typology of agency for frontier AI systems consisting of three dimensions: the nature of agency (moral or legal), its mode (individual or collective) and its locus (human or non-human). Combining these dimensions produces eight possible instantiations of agency, which we classify as conventional, contested or controversial. The typology separates legal from moral agency and thereby creates conceptual space for considering individual, legal, non-human agency without presupposing that advanced AI systems are moral agents. We argue that this distinction is increasingly relevant where instrumental goal pursuit complicates the attribution of AI actions to particular human actors.

13:00 JST研究/論文

セーフティ ネットの適用性について: ニューラル ネットワークを認証するためのセーフティ バイ デザイン ソリューション

安全性が重要な航空システムに人工知能 (AI) を統合することは、認証と導入に大きな課題をもたらします。最も安全な輸送手段とみなされがちな航空は、安全性が重要な多数のシステムに依存しています。将来のセーフティ クリティカルな AI ベース システムに対して、EASA はセーフティ バイ デザイン アプローチを必要としています。これは、ニューラル ネットワーク圧縮とルックアップ テーブルを組み合わせたセーフティ ネットを使用して、離散化された運用設計ドメイン全体で 100% 正しい実行時の動作を保証することで実現できます。セーフティ ネットは研究されていますが、そのパフォーマンス特性とシステム設計のトレードオフに関する包括的な研究は行われていません。この研究は、セーフティ ネットにおけるニューラル ネットワークとルックアップ テーブルのサイズの間のトレードオフに関する初めての体系的な分析を示しています。この研究では、ニューラル ネットワークをさまざまなアーキテクチャと系統的に比較することで、認証準拠を維持しながら全体的なストレージとメモリの要件を最小限に抑える最適な設計パラメータを特定します。結果は、それぞれ約 50 ~ 100 ノードを持つ 3 ~ 5 の隠れ層を備えたアーキテクチャとワンホット エンコーディングを組み合わせることで、最適なバランスが実現されることを示しています。これらの構成では、ニューラル ネットワークがデータの少なくとも 97 % を正確に表し、コンパクトなルックアップ テーブルが残りのエラーを処理します。結果として得られるセーフティ ネットは、EASA ガイドラインで要求されているように、離散化入力空間全体にわたって 100% 正しい出力を保証しながら、システム サイズをほぼ 3 桁削減し、現在のアビオニクス ハードウェアのメモリ バジェット内に収まります。この取り組みは、HCAS および VCAS 向けのセーフティ ネットのオープンソース実装で再現可能な結果を​​もたらし、航空分野における認定可能な AI ベースのシステムへの実用的な道筋を実証し、セーフティ クリティカルなアプリケーション向けの実行可能なセーフティ バイ デザイン ソリューションとしてのセーフティ ネットを確立します。

原文 (English)

On the Applicability of Safety Nets: A Safety-By-Design Solution for Certifying Neural Networks

The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment. Aviation, often regarded as the safest form of transportation, relies on numerous safety-critical systems. For future safety-critical AI-based systems, EASA requires a Safety-by-Design approach, which can be achieved by using Safety Nets that combine neural network compression with lookup tables to ensure 100 % correct runtime behavior across the discretized operational design domain. Although Safety Nets have been studied, no comprehensive study of their performance characteristics and system design trade-offs has been conducted. This work presents the first systematic analysis of the trade-off between neural network and lookup table size in Safety Nets. By systematically comparing neural networks with diverse architectures, this study identifies optimal design parameters that minimize overall storage and memory requirements while maintaining certification compliance. Results demonstrate that architectures with 3 to 5 hidden layers, each with approximately 50 to 100 nodes, combined with one-hot encoding, achieve the best balance. In these configurations, neural networks accurately represent at least 97 % of the data, while compact lookup tables handle the remaining errors. The resulting Safety Nets reduce the system size by almost three orders of magnitude, fitting within the memory budget of current avionics hardware while guaranteeing 100 % correct outputs across the entire discretized input space, as required by EASA guidelines. This work provides the first-ever open-source implementation of Safety Nets for HCAS and VCAS with replicable results, demonstrating a practical pathway toward certifiable AI-based systems in aviation and establishing Safety Nets as a viable Safety-by-Design solution for safety-critical applications.

13:00 JST研究/論文

目に見えないものはあなたが学ぶものです:共有ゲノム言語モデル社会では、限られた証拠の可視性が構成の一般化を促進します

マルチモジュール システムでは、多くの場合、すべてのモジュールが完全な入力に公開されます。証拠の可視性を制限すると、勾配ベースのトレーニングで発見されるソリューションが変わるかどうかをテストします。 4 セル社会は 1 つの凍結済み事前学習済み言語モデルと 1 つの低ランク アダプターを共有し、固定リレー内の 2 つのモデル幅の連続ベクトルを通じてのみ通信します。プロスペクティブシールされた自然言語関数合成タスクでは、初期化バイト、トレーニング順序、トークン レイアウト、パラメーター、および計算を共有する 10 個の一致する制限付き/グローバル ペアをトレーニングします。注意マスクのみが異なります。制限された社会は、10 ペア中 9 ペアで両方の深さで世界的に見える双子よりも少なくとも 20 ポイント優れており、ペアの利点の中央値は 0.7648 と 0.6050 です。コミュニケーションを遮断すると、すべての制限された社会は偶然に帰着し、複合関数がトレーニングに一度も現れなかったプログラムでは、深さ 3 の利点は 0.558 のままです。監査された 6 つの制限された社会全体で、同じ値のパケット移植により、テストされたすべてのインターフェイスで動作が 0.94 ~ 1.00 に維持されます。破壊的な介入はパフォーマンスを崩壊させます。そして反事実パケットは出力を数学的に予測された答えにリダイレクトします。唯一の高性能グローバル モデルも通信を必要としますが、その同じ値のパケットはエピソード間で交換できません。したがって、構図には可視性の制限は必要ありません。このプロトコルでは、一般化リレーの可能性が大幅に増加し、再利用可能な値インデックス付きインターフェイスが優先されます。それにもかかわらず、完全に事前登録されたバッテリーは、制限付きアームの深さ 3 の中央値の精度が 0.6988 で、0.70 フロアを下回っているため、正式に不合格となります。以前の資格コホートでも同様に完全合格は 0/10 でした。1 つのモデルはすべてのタスク パフォーマンス ゲートを満たしましたが、10 モデルすべてが通常言語の保存に失敗し、システムは明示的にタスク ゲートで使用されるように制限されました。

原文 (English)

What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies

Multi-module systems often expose every module to the full input. We test whether restricting evidence visibility changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay. On a prospectively sealed natural-language function-composition task, we train ten matched restricted/global pairs sharing initialization bytes, training order, token layout, parameters, and computation; only the attention mask differs. Restricted societies outperform their globally visible twins by at least 20 points at both depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050. Cutting communication reduces every restricted society to chance, and the depth-three advantage remains 0.558 on programs whose composite function never appeared in training. Across six audited restricted societies, same-value packet transplants preserve behavior at 0.94-1.00 across all tested interfaces; destructive interventions collapse performance; and counterfactual packets redirect outputs toward the mathematically predicted answer. The sole high-performing global model also requires communication, but its same-value packets are not interchangeable across episodes. Restricted visibility is thus not necessary for composition; under this protocol it substantially increases the probability of a generalizing relay and favors a reusable, value-indexed interface. The complete preregistered battery nevertheless formally fails because restricted-arm median depth-three accuracy is 0.6988, below the 0.70 floor. An earlier qualification cohort likewise yielded 0/10 complete passes: one model met every task-performance gate, but all ten failed ordinary-language preservation, confining the system to explicitly task-gated use.

13:00 JSTロボティクス

DECOWAM: 脚式モバイル操作のための分離された全身世界アクション モデル

モバイル操作では、移動と腕の動きが連携して将来の観察と制御がどのように変化するかをロボットが予測する必要があります。既存のワールド アクション モデルは、主に固定ベース プラットフォーム向けに開発されており、カメラのエゴモーションとベースおよびアームのアクションを明確に区別していません。ここでは、専用の条件付きインターフェイスを通じてこれらの要素を分離する全身世界行動モデル DECOWAM を紹介します。 DECOWAM は、適応された FastWAM バックボーンを凍結し、残りのアダプター、特権的な観察から抽出されたアクションに相当する将来のボトルネック、敵対的に分離されたベースとアームの潜在、およびビデオ予測のための基本速度調整をトレーニングします。さらに、ビデオ、全身の状態と動作、および言語を同期する実際のロボット データセットである ARMDOG を紹介します。固定再生プロトコルでは、DECOWAM は FastWAM よりも将来のビデオとアクションの予測の両方を改善し、2,595 万のトレーニング可能な適応パラメーターでアクション MSE を 21.7% 削減しました。メソッドごとに 79 回の閉ループ試行を行った結果、比較したシステムの中で観察された中で最も高い全身調整とベース変位の堅牢性が達成され、タスクの完了は依然として最強のベースラインと同等でした。これらの結果は、実施形態を意識した因数分解が、移動視点の下でパラメータ効率の高い共同視覚予測と全身制御をサポートできることを示しています。

原文 (English)

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM backbone and trains residual adapters, an action-equivalent future bottleneck distilled from privileged observations, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. We further introduce ARMDOG, a real-robot dataset that synchronizes video, whole-body state and action, and language. On a fixed replay protocol, DECOWAM improved both future-video and action prediction over FastWAM, reducing action MSE by 21.7% with 25.95M trainable adaptation parameters. Across 79 closed-loop trials per method, it achieved the highest observed whole-body coordination and base-displacement robustness among the compared systems, while task completion remained comparable to the strongest baseline. These results show that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.

13:00 JST研究/論文

DARS: 命令ベースの画像編集のための構造化推論を備えたデュアルレベルのクレジット割り当て RL

命令ベースの画像編集では、プランナー レンダラー パイプラインを使用します。まず、ビジョン言語モデル (VLM) が命令を編集プランに変換し、次に拡散モデルがそのプランを実行します。最終画像報酬のみを使用してこのようなシステムをトレーニングすることは非効率的です。編集が不十分だと、追加の最適化でプランナーとレンダラのどちらを重視すべきかが明らかにならず、プランナーが優勢なケースでさえ、自由形式の推論トレース内でローカライズするのが困難なままであるためです。この 2 段階設定でのデュアルレベル単位割り当てのための強化学習フレームワークである DARS を紹介します。モジュール全体で、マルチプラン マルチレンダー ロールアウトは、ソフト モジュール ルーティングのプラン間およびプラン内の報酬変動を推定します。一方、ロールアウト平均報酬は、適応型カリキュラムのハードさの推定を提供します。プランナー内では、4 フィールドの構造化推論出力により、プレフィックスゲート報酬とトークンレベルの利点の再重み付けが可能になり、結果レベルのフィードバックが局所的な監視に変わります。 5 つのベンチマークの実験では、DARS が同じバックボーン、データ、報酬モデル、ロールアウト予算を使用した Joint~RL ベースラインよりも優れたパフォーマンスを示し、推論集中型の編集で最大の利益が得られることが示されています。

原文 (English)

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-image rewards is inefficient because a poor edit does not reveal whether additional optimization should place more emphasis on the planner or the renderer, and even planner-dominant cases remain difficult to localize within a free-form reasoning trace. We present DARS, a reinforcement learning framework for dual-level credit assignment in this two-stage setting. Across modules, multi-plan multi-render rollouts estimate between-plan and within-plan reward variability for soft module routing, while rollout mean rewards provide hardness estimates for an adaptive curriculum. Within the planner, a four-field structured reasoning output enables a prefix-gated reward and token-level advantage reweighting, turning outcome-level feedback into localized supervision. Experiments on five benchmarks show that DARS outperforms a Joint~RL baseline with the same backbone, data, reward model, and rollout budget, with the largest gains on reasoning-intensive edits.

13:00 JSTエージェント

第三次ソフトウェア形態の再構築:三層アーキテクチャからストレージ、モデル、エージェントへ

ソフトウェア形式は、その誕生以来 2 つのパラダイム シフトを経験しました。命令が動作を決定するソフトウェア 1.0 と、データが動作 (機械学習) を決定するソフトウェア 2.0 です。この論文は、コンテキストと推論が動作を決定する第 3 のシフト、ソフトウェア 3.0 が現在進行中であると主張し、その最終形態は、一般化されたデータベース (すべての永続的な状態とメモリの統一された抽象化)、大規模モデル (推論と生成を実行するインテリジェンス コア)、およびエージェント (最初の 2 つを接続する実行ループ) の 3 つの要素に収束すると主張します。中心的な議論は次のとおりです。従来の 3 層アーキテクチャでは、ユーザー インターフェイス層はオンデマンドでインターフェイスを生成するモデルの機能に吸収され、ビジネス ロジック層は「表現性 x 重要性」に沿ってモデル推論とストレージ制約に再分割され (残りの決定論的ロジックはツールとして保持されます)、データ層のみが唯一の永続的なインフラストラクチャに昇格します。私たちは、この収束理論を形式化し、最小限の参照アーキテクチャを提示し、実際のプロトタイプとライブモデルからの証拠を報告し、それが成り立つ条件とそれが失敗する境界の両方を体系的に分析します。決定論、コスト、セキュリティ、検証可能性がこの論文の適用範囲を区切ります。私たちは、この理論は、表現可能で検証可能で、外部的にステートフルで、ツールで完全なタスク領域に当てはまり、開発者、データベース業界、およびソフトウェア エンジニアリング分野の役割を再構築するだろうと主張します。

原文 (English)

The Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents

Software form has undergone two paradigm shifts since its inception: Software 1.0, in which instructions determine behavior, and Software 2.0, in which data determines behavior (machine learning). This paper argues that a third shift - Software 3.0, in which context and reasoning determine behavior - is now underway, and contends that its terminal form converges to three elements: a generalized database (the unified abstraction of all persistent state and memory), a large model (the intelligence core that performs reasoning and generation), and an agent (the execution loop connecting the first two). The core argument is as follows: in the traditional three-tier architecture, the user-interface layer will be absorbed by the model's ability to generate interfaces on demand, the business-logic layer will be re-partitioned along "expressibility x criticality" into model reasoning and storage constraints (with residual deterministic logic retained as tools), and only the data layer will be elevated into the sole persistent infrastructure. We formalize this convergence thesis, present a minimal reference architecture, report evidence from real prototypes and a live model, and systematically analyze both the conditions under which it holds and the boundaries where it fails - determinism, cost, security, and verifiability delimit the thesis's domain of applicability. We argue that the thesis holds in task domains that are expressible, verifiable, externally stateful, and tool-complete, and that it will reshape the roles of developers, the database industry, and the software-engineering discipline.

13:00 JSTLLM/生成AI研究/論文

MemTrapBench: LLM メモリ使用における認知トラップのベンチマーク

記憶は大規模な言語モデルの重要なコンポーネントとなっており、情報を保持し、長期的な対話から学習できるようになります。しかし、既存のメモリ ベンチマークは主に、情報が正しく抽出、保存、取得されているかどうかを評価しており、取得されたメモリがどのようにモデル推論を再形成し、現在のタスクのパフォーマンスに影響を与えるかはほとんど見落とされています。私たちは、記憶によって引き起こされる認知の罠を特定します。忠実に記録され、意味的に関連性のある記憶であっても、モデルの推論や信念が歪められ、現在のタスクのパフォーマンスが低下する可能性があります。これらの障害モードを系統的に評価するために、推論固着と信念の歪みという 2 つの形式の認知トラップをカバーする MemTrapBench を紹介します。 2 つのモデル ファミリと 5 つの代表的なメモリ フレームワークにわたる実験では、MemTrapBench が困難であることが示されています。評価されたメモリ戦略はすべてメモリなし設定よりもパフォーマンスが低く、最も強力なメソッドでも 10% 以上の低下に見舞われます。これらの認知トラップを軽減するために、LLM にメモリ トラップを回避するように指示する、シンプルだが効果的な推論時間手法である AdaptiveMem を提案します。 AdaptiveMem は、さまざまなメモリ フレームワークにわたる標準メモリ ベンチマークのパフォーマンスを維持または向上させながら、MemTrapBench の認知トラップを軽減します。

原文 (English)

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.

13:00 JSTLLM/生成AI研究/論文

ContractScrub: 法的契約の最終レビューのベンチマーク

法務業務は大量のテキストの処理に大きく依存しており、LLM の使用に最もさらされる分野の 1 つと考えられています。契約の「スクラブ」、つまり取引契約のエラーや不一致の最終レビューは、長い文書に細心の注意を払う必要がある日常的で骨の折れる作業であるため、自動化に特に適したタスクです。また、スクラビングは、ロングコンテキスト推論、一貫性チェック、固有表現認識 (NER) などのフロンティア LLM に期待される一般的な機能とも自然に一致しているようです。経済的価値と自動化の可能性にもかかわらず、契約スクラビングを実行する LLM の正式な評価は行われていません。当社は、契約書スクラブ機能を評価するために設計された最初のベンチマークである ContractScrub を紹介します。契約書スクラブ機能は、定義された用語の誤用、誤った参照、一貫性のない言葉遣いなど、さまざまなエラー カテゴリに関して経験豊富な弁護士によって手作りされた契約書で構成されます。フロンティア モデルのパフォーマンスは驚くほど低く、一見関連していると思われる一般的なベンチマークでは優れたパフォーマンスを示しているにもかかわらず、マクロ平均再現率 0.75 に達したモデルは 1 つだけであり、現行モデルの実際的な限界と、現実世界の影響を測定するための対象を絞ったドメイン固有のベンチマークの重要性を示しています。

原文 (English)

ContractScrub: A benchmark for final review of legal contracts

Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general capabilities expected of frontier LLMs around long-context reasoning, consistency checking, and named entity recognition (NER). Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted. We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect references, and inconsistent language. Frontier models perform surprisingly poorly with only one model reaching 0.75 macro average recall despite strong performance on seemingly related general benchmarks, demonstrating the practical limits of current models and the importance of narrowly targeted, domain-specific benchmarks for measuring real-world impact.

13:00 JST研究/論文

電子海図変更分類

電子航海図 (ENC) は、深度、航行補助、交通計画、危険などの水路および航行情報を表す海洋航法システムで使用される地理空間ベクトル データセットです。水路事務所にとっての大きな課題は、特定の海図の変更が海上の安全に重大なリスクをもたらすか、重大でないリスクをもたらすかを判断することです。既存のワークフローは手動によるレビューと検証に大きく依存しており、労働集約的であり、受信するチャート更新の量に合わせて拡張することができず、アナリスト間の不一致が生じます。この課題に対処するために、ENC 変更を自動分類する方法を提案します。複雑なベクトル データの変更を分類モデル用の構造化された表形式に変換するためのベースライン エンコード スキームを確立します。エンコード スキームの 2 つの重要なコンポーネントには、周囲の地理的特徴による変更表現を強化する空間コンテキスト エンコーダーと、変更されたオブジェクトの微妙な属性値の説明を表現する ENC 属性エンコーダーが含まれます。提案されたアプローチを、100,000 を超える個別のチャート変更を含む 1,308 のチャート ペアで構成される 2 つの異なる運用データセットにわたって評価します。提案されたエンコード スキームを活用して調整された勾配ブースト ツリーは、2 つのデータセットで 90% と 94% の精度を達成し、空間コンテキストと属性の埋め込みなしのエンコードでトレーニングされたデフォルトのハイパーパラメーター化モデルと比較して 5 ~ 7% の改善が得られました。これらの結果は、機械学習を運用上の地理空間パイプラインに統合して、ENC のメンテナンスを改善し、海上の安全性を高めることが実現可能であることを示しています。最後に、私たちの実験は、単純な位置および空間集約手法の有効性を実証し、このアプリケーション向けのより高度な空間表現学習手法を評価するための基盤を提供します。

原文 (English)

Electronic Navigational Chart Change Classification

Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and navigational information such as depths, navigational aids, traffic schemes, and hazards. A major challenge for hydrographic offices is determining whether a given chart change poses a critical or non-critical risk to maritime safety. Existing workflows rely heavily on manual review and verification, which is labor-intensive, scales poorly with the volume of incoming chart updates, and introduces inter-analyst inconsistencies. To address this challenge, we propose a method for automated classification of ENC changes. We establish a baseline encoding scheme to translate complex vector data changes into a structured tabular format for classification models. The two crucial components of the encoding scheme include a spatial context encoder to enrich the change representations with surrounding geographic features, and an ENC attribute encoder to represent nuanced attribute-value descriptions of the modified objects. We evaluate the proposed approach across two distinct operational datasets, comprising 1,308 chart pairs containing over 100,000 individual chart modifications. Tuned gradient-boosted trees leveraging the proposed encoding schemes achieve accuracies of 90% and 94% on the two datasets, yielding a 5-7% improvement over default hyperparameterized models trained on encodings without spatial context and attribute embeddings. These results demonstrate the viability of integrating machine learning into operational geospatial pipelines to improve ENC maintenance and enhance maritime safety. Finally, our experiments demonstrate the effectiveness of simple location and spatial aggregation methods, providing a foundation for evaluating more sophisticated spatial representation learning techniques for this application.

13:00 JSTLLM/生成AI研究/論文

InsufficiencyBench: 不正確なユーザー クエリに関する LLM の法的アドバイスを評価する

法的 AI システムは、法的な質問に答えるためにますます使用されていますが、既存のベンチマークは、クエリが完全に指定されて到着することを前提としています。実際には、ユーザーは法的結果を大きく決定する事実を省略します。クエリ側の不十分性を対象とした初の法的ベンチマークである InsufficiencyBench を紹介します。クエリに法的に重要な情報が欠落していることをモデルが認識し、欠落しているものを特定し、早まった結論を控えるかどうかです。私たちは、スイッチ、ゲート、致命的な前提条件という 3 つの構造的故障モードにわたる 8 つの標準的な欠落要素カテゴリの分類を形式化し、6 つの法的領域と米国 24 の管轄区域にまたがる 202 のベンチマーク項目 (58 の基本クエリ、144 の欠陥バリアント) を構築し、現役の弁護士によって注釈が付けられています。 10 個のフロンティア モデルを評価すると、欠落要素の識別で F2 = 0.46 を超えるモデルはなく、再現率の中央値が 0.44 であることがわかります。モデルは無差別にヘッジするか、でっち上げられた仮定の下で黙って答えるかのどちらかです。完全なクエリに直接対処しながら、欠陥のあるクエリに対する応答を識別および認定するモデルはありません。

原文 (English)

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially determine the legal outcome. We introduce InsufficiencyBench, the first legal benchmark targeting query-side insufficiency: whether a model recognizes when a query lacks legally material information, identifies what is missing, and refrains from premature conclusions. We formalize a taxonomy of eight canonical missing-element categories across three structural failure modes---switch, gating, and fatal prerequisite--- and construct 202 benchmark items (58 base queries, 144 deficient variants) spanning six legal domains and 24 US jurisdictions and annotated by practising attorneys. Evaluating ten frontier models, we find that no model exceeds F2 = 0.46 on missing-element identification and that the median recall is 0.44. Models either hedge indiscriminately or answer silently under fabricated presumptions. No model both identifies and qualifies responses to deficient queries while directly addressing complete ones.

13:00 JSTLLM/生成AI

マルチモーダル大規模言語モデルのためのルールに準拠した視覚的空間計画

マルチモーダル大規模言語モデル (MLLM) は、言語推論と視覚認識を組み合わせていますが、明示的な、またはこれまでに見たことのないルール制約の下で視覚空間計画を実行する機能はまだ十分に解明されていません。この設定では、モデルが空間レイアウトを共同で理解し、自然言語ルールを解釈し、それに応じて有効なアクションを計画する必要があります。このギャップに対処するために、MLLM がさまざまな複雑さの自然言語ルールに従いながら迷路をナビゲートする必要がある制御可能なベンチマークである RuleMaze を導入します。 RuleMaze は、正確な認識、ルールの解釈、制約のある行動計画を要求することで、ルールに準拠した空間計画を分離します。スケーラブルで体系的なルール構築を可能にするために、私たちは言語-論理-関数ハイブリッド化を提案します。これは、自然言語ルールを自動的に生成し、それらを論理表現と実行可能なバリデータに変換し、手動のルールエンジニアリングを排除します。ルールの追従と一般化を改善するために、解釈可能な推論プリミティブを通じて認識、実行、およびルールの検証を分離する、解絡マルチモーダル プランニング (DMP) を導入します。これらのコンポーネントを解きほぐすことで、DMP はより複雑でこれまで目に見えなかったルールへの体系的な一般化を促進すると同時に、透過的な中間計画トレースを提供します。実験では、エンドツーエンドのテキストによる計画ベースラインと比較して、DMP がルールの遵守と計画の成功を大幅に向上させることが実証されています。全体として、RuleMaze は、MLLM における根拠があり解釈可能なルールベースの空間計画を研究するための原則に基づいたベンチマークを確立します。コードは https://github.com/oceanflowlab/RuleMaze で入手できます。

原文 (English)

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret natural-language rules, and plan valid actions accordingly. To address this gap, we introduce RuleMaze, a controllable benchmark in which MLLMs must navigate mazes while obeying natural-language rules of varying complexity. RuleMaze isolates rule-compliant spatial planning by requiring accurate perception, rule interpretation, and constrained action planning. To enable scalable and systematic rule construction, we propose Language-Logic-Function Hybridization, which automatically generates natural-language rules and translates them into logical representations and executable validators, eliminating manual rule engineering. To improve rule following and generalization, we introduce Disentangled Multimodal Planning (DMP), which separates perception, execution, and rule verification through interpretable reasoning primitives. By disentangling these components, DMP facilitates systematic generalization to more complex and previously unseen rules, while providing transparent intermediate planning traces. Experiments demonstrate that DMP substantially improves rule compliance and planning success compared to end-to-end textual planning baselines. Overall, RuleMaze establishes a principled benchmark for studying grounded and interpretable rule-based spatial planning in MLLMs. Code is available at https://github.com/oceanflowlab/RuleMaze.

13:00 JST研究/論文

QUASAR: A Quantum-Classical Neural Network for SAR Satellite Physical-Layer Authentication

X-band SAR satellites (8-12 GHz) play a critical role in disaster response, environmental monitoring, and military intelligence. Yet, they…

13:00 JST研究/論文

いつ考えるべきかを学ぶ: テスト時のコンピューティング割り当てのための適応推論

強化学習でトレーニングされた推論言語モデルは通常、明示的に適応するトークン バジェットではなく固定トークン バジェットの下で動作するため、簡単な問題では過剰な計算が発生し、難しい問題では不十分な計算が発生する可能性があります。私たちは、モデルが応答の最初のトークンとして、\textsc{NoThink} (できるだけ早く答える)、\textsc{Short} (短い推論)、または \textsc{Long} (拡張推論) の 3 つのモードのいずれかを選択することによって、独自の推論の労力を割り当てることを学習できるかどうかを研究します。選択は、個別のルーターを使用せずに、異なる応答長で各モードの価値を高める成形報酬と、モードを区別し続けるモードごとのハード トークン キャップを通じて、Group Relative Policy Optimization (GRPO) 内で学習されます。 MATH でトレーニングされた 1.5B の抽出されたモデルでは、3 つのモードが 1 つの選択肢に崩れることなく出現し、簡易モードは最終的に \textsc{Long} よりも正確になります。これは、ルーターが問題をランダムではなく難易度に基づいて並べ替えていることを示しています。 3 つのシードで平均した結果のポリシーは、ホールドアウトされた MATH500 での基本モデルの精度 ($0.782$ 対 \ $0.796$) に近いままですが、平均応答長は $4{,}796$ から $2{,}811$ トークンに短縮されました ($41\%$ の削減)。興味深いことに、再トレーニングなしで他のベンチマークにも移行し、問題が容易な場合に最大の節約効果が得られます。たとえば、GSM8K では 76\% のトークン削減が実現し、同様の応答長でのベースラインよりも高い精度で実現されます。つまり、各問題についてどの程度推論するかを適応的に選択する推論モデルを構築します。

原文 (English)

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on difficult ones. We study whether a model can learn to allocate its own reasoning effort by choosing, as the first token of its response, one of three modes: \textsc{NoThink} (answer as quickly as possible), \textsc{Short} (brief reasoning), or \textsc{Long} (extended reasoning). The choice is learned inside Group Relative Policy Optimization (GRPO) with no separate router, through a shaped reward that makes each mode worthwhile at a different response length, together with hard per-mode token caps that keep the modes distinct. On a 1.5B distilled model trained on MATH, the three modes emerge without collapsing to a single choice, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random. Averaged over three seeds, the resulting policy stays close to the base model's accuracy on the held-out MATH500 ($0.782$ vs.\ $0.796$) while cutting the mean response length from $4{,}796$ to $2{,}811$ tokens (a $41\%$ reduction). Interestingly, it also transfers to other benchmarks without retraining, with the largest savings where problems are easier, with for instance 76\% token reduction on GSM8K and at higher accuracy than the baselines at similar response length. In short, we build a reasoning model that adaptively chooses how much to reason for each problem.

13:00 JST研究/論文

ラグを捕まえる: 機械学習による Solana 上の不正なミームコインの早期予測

ブロックチェーンプラットフォーム上でのミームコインの急速な普及により、不正行為、特に敷物引きのリスクが増加しています。これまでの研究はイーサリアムベースのトークンに焦点を当てていましたが、この論文では、取引量とトークン数でミームコインの主要なブロックチェーンである Solana に焦点を当てます。ラグプルがスマートコントラクトのバックドアを悪用することが多いイーサリアムとは異なり、Solana ミームコインのラグプルは主に流動性操作と社会力学によって動かされます。この研究は、7 か月にわたって 640 万個のトークンのデータセットを収集することにより、Solana エコシステムにおける大規模なラグ プルの早期検出の先駆けとなります。市場分析の結果、これらのミームコインの大部分が発売後 1 時間以内にラグプル特性を示すことが明らかになり、短期予測の緊急性が強調されています。コードレベルの機能がないにもかかわらず、古典的な機械学習モデル、特に勾配ブースティング (XGBoost) が、取引データの最初の 5 分間のみを使用して潜在的なラグプルを検出する際に堅牢なパフォーマンスを達成することを実証します。さらに、PumpFun と Raydium の間のクロスプラットフォームの一般化を評価し、マルチソース データの融合によりドメイン シフトが大幅に緩和され、検出の信頼性が向上することが明らかになりました。この研究は、ハイスループットチェーンにおける DeFi 詐欺の理解を進め、投資家を保護するための実践的な枠組みを提供します。

原文 (English)

Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning

The rapid proliferation of memecoins on blockchain platforms has increased the risk of fraudulent activities, particularly rug pulls. While previous studies have focused on Ethereum-based tokens, this paper shifts the spotlight to Solana, the leading blockchain for memecoins by trading volume and token count. Unlike Ethereum, where rug pulls often exploit smart contract backdoors, Solana memecoin rug pulls are predominantly driven by liquidity manipulation and social dynamics. This research pioneers large-scale rug pull early detection in the Solana ecosystem by assembling a dataset of 6.4 million tokens over 7 months. Market analysis reveals that a vast majority of these memecoins exhibit rug pull characteristics within one hour of launch, highlighting the urgency of short-horizon prediction. Despite the absence of code-level features, we demonstrate that classic machine learning models, particularly Gradient Boosting (XGBoost), achieve robust performance in detecting potential rug pulls using only the first 5 minutes of trading data. Furthermore, we evaluate cross-platform generalization between PumpFun and Raydium, revealing that multi-source data fusion significantly mitigates domain shift and improves detection reliability. This study advances the understanding of DeFi fraud on high-throughput chains and provides a practical framework for protecting investors.

13:00 JSTLLM/生成AIエージェント

分解して継承: LLM エージェントにおけるタスクをまたいだスキルの伝達

大規模言語モデル (LLM) エージェントは、完了したタスクからスキルを引き出し、後で再利用して経験を積むことで能力を高めることができます。実際には、誘発されたスキルは信頼性が低く伝達される可能性があり、それを取得するエージェントに害を及ぼす可能性さえあります。エージェントによって誘発されたスキルがタスク間で確実に伝達されるかどうかは未解決の問題のままです。私たちは、スキルの習得方法がタスク間でのスキルの伝達をどのように形作るかについて、包括的かつ管理された研究を実施します。具体的には、タスクレベルとサブタスクレベルのスキル導入、およびテキストとコードスキル形式を比較します。この 2 つの軸に沿って既存の手法は異なります。タスク レベルのスキルはほとんどの場合、エージェントのパフォーマンスをメモリなしのベースラインよりも低下させますが、サブタスク レベルのスキルは平均以上にパフォーマンスを向上させ、テキスト スキルはコード スキルよりも効果的に伝達されます。私たちの調査結果をさらに理解するために、誘発されたスキルの 2 つの相補的な特性を調べます。1 つはスキルが実際のタスクにどの程度一致しているかを測定する「特異性」、もう 1 つはスキルの関連性がタスク全体にどの程度均等に広がっているかを測定する「抽象性」です。どちらの特性も単独ではタスクの成功を予測しませんが、それらを組み合わせた効果によって予測が可能です。これをスキル ユーティリティ スコアとして提案します。スキルが転送されると、スコアはタスクの成功と一貫して相関し、サブタスク レベルおよびテキスト スキルのスコアが高くなります。コンピューティング スキル ユーティリティには、スキルとタスクの説明のみが必要ですが、タスクの実行は必要ありません。そのため、新しいタスクを実行する前に、スコアはスキル記憶の実用的な診断として機能します。

原文 (English)

Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them. When agent-induced skills transfer reliably across tasks remains an open question. We conduct a comprehensive and controlled study of how the way skills are induced shapes their transfer across tasks. Specifically, we compare task-level with subtask-level skill induction and text with code skill formats, the two axes along which existing methods differ. Task-level skills mostly reduce the agent's performance below its no-memory baseline while subtask-level skills raise it above on average, and text skills transfer better than code skills. To further understand our findings, we examine two complementary properties of the induced skills: specificity, which measures how closely a skill matches real tasks, and abstractness, which measures how evenly its relevance spreads across tasks. Neither property alone predicts task success, but their combined effect does, which we propose as a skill utility score. The score correlates consistently with task success when skills are transferred, and subtask-level and text skills score higher. Computing skill utility only needs the skills and task descriptions but not any task execution, so our score serves as a practical diagnostic of a skill memory before any new task runs.

13:00 JSTLLM/生成AI

ファントムゲイン: 測定されたヌルに対する自己改善の監査

言語モデル自体が改善されたかどうかは、平均的な精度ではなく、個々の問題によって得られるか失われるかによって判断されるようになってきています。これらの遷移を追跡することは、ノイズの多い 2 つの推定値を差分することを意味し、測定アーチファクトに対して脆弱なままになります。同一のパイプラインを介してプッシュされたフリーズしたコントロールに対して、Qwen3-8B でのランク $32$ の LoRA セルフ トレーニングの 3 ラウンドを監査したところ、7 つの測定エラーが特定されました。そのそれぞれが、コントロールが存在しない場合に報告された結果を覆します。いくつかは標準的な方法です。単一の貪欲なデコードに基づいて構築された台帳は、トレーニングされていないモデルで機能の変更を行いますが、これは主に推論バッチ処理の成果物です。取得とシャープ化を分ける拡張統計により、その同じモデルに $0.280$ のレートが割り当てられます。自然なしきい値修復はレプリケーションに存続しません。デザインにすでに含まれている凍結された比較全体で推定すると、その null は非ゼロのままです。これを、誤検出率制御の下でプールされたベースラインに対する問題ごとの正確なテストに置き換えます。これは、保持された複製では何も検出されず、複数のテスト ルール、エラー率、およびプール サイズの下では変更されません。流れ、ボリューム、評価が一致するアームのはしごに適用された監査では、外部蒸留によって、基本モデルではめったに到達できない問題が改善される一方、3 つの形式の自己トレーニングでは改善できないことが判明しました。回帰では、この非対称性は蒸留による全体的な利得が大きくなった結果の副産物として拒否されます ($p < 10^{-8}$)。基本モデルが決して到達しないはるかに小さな問題セットについては、証拠は決定的ではありませんが、自己学習によりベースラインで解決された問題は、測定された下限をはるかに超える割合で破壊されます。したがって、移行レベルの監査では、レポートするすべての統計について個別に測定されたヌルが必要です。ヌルは、新たな実験の費用がかからず、ほとんどの研究者ほど少ないものではありませんが、すでに所有している複数群研究のベースラインの複製から構築されます。

原文 (English)

Phantom Gains: Auditing Self-Improvement Against a Measured Null

Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three rounds of rank-$32$ LoRA self-training on Qwen3-8B against a frozen control pushed through the identical pipeline, we identify seven measurement failures, each of which inverts a reported finding when its control is absent. Several are standard practice. A ledger built on a single greedy decode manufactures capability changes on an untrained model, largely an artifact of inference batching; the expansion statistic separating acquisition from sharpening assigns that same model a rate of $0.280$. The natural threshold repair does not survive replication: estimated across the frozen comparisons such a design already contains, its null stays non-zero. We replace it with a per-problem exact test against a pooled baseline under false-discovery-rate control, which detects nothing on any held-out replicate and is unchanged under the multiple-testing rule, error rate and pool size. Applied to a ladder of arms matched in stream, volume and evaluation, the audit finds that external distillation improves problems the base model rarely reaches while three forms of self-training do not; a regression rejects this asymmetry as a by-product of distillation's larger overall gain ($p < 10^{-8}$). On the far smaller set of problems the base model never reaches, the evidence is inconclusive, while self-training corrupts problems solved at baseline at rates well above the measured floor. Transition-level auditing therefore requires a separately measured null for every statistic it reports: nulls that cost no new experiments, built from baseline replicates a multi-arm study already owns, though not from as few as most possess.

13:00 JSTエージェント

MidTool: エージェントティック ツールを使用するためのトレーニング中のデータ合成

中間トレーニングは、大規模な言語モデルの機能を形成するための重要な段階としてますます認識されています。最近の研究では、的を絞った中間トレーニングにより、数学や科学などの推論集中型の能力が強化され、ソフトウェア エンジニアリング設定におけるエージェント能力も向上できることが示されています。この研究では、並行して行われているがあまり調査されていないエージェント機能、つまり一般的なツールの使用について研究します。私たちは、大規模な Web、PDF、コード データを、現実世界のツール API、MCP スキル、およびドキュメントベースのワークフローからの合成された監視と組み合わせた、エージェント ツール使用の中間トレーニング用のオープン コーパス構築パイプラインである MidTool を紹介します。 MidTool は、ツール アフォーダンスを認識し、コンテキストから引数を根拠づけ、ツール呼び出しワークフローを構成し、不完全な情報から回復する方法をモデルに教えるように設計されています。 MidTool-Mix で Qwen3-4B-Base と Qwen3-8B-Base を中間トレーニングし、教師あり微調整と強化学習の両方を使用してトレーニング後のフォローアップを適用します。ベースラインと比較して、MidTool-Mix は、BFCL、tau2-Bench、および MCP Universe の SFT と RL の両方でダウンストリーム パフォーマンスを一貫して向上させます。これらの結果は、一般的なツールの使用は、他の重要な LLM 機能と同様に、トレーニング後に完全に任せるのではなく、トレーニング中に専念することで恩恵を受けることを示唆しています。

原文 (English)

MidTool: Mid-training Data Synthesis for Agentic Tool Use

Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown that targeted mid-training can strengthen reasoning-intensive abilities such as math and science, and can also improve agentic capabilities in software-engineering settings. In this work, we study the parallel but less explored agentic capability: general tool use. We present MidTool, an open corpus construction pipeline for agentic tool-use mid-training that combines large-scale web, PDF, and code data with synthesized supervision from real-world tool APIs, MCP skills, and document-grounded workflows. MidTool is designed to teach models how to recognize tool affordances, ground arguments from context, compose tool call workflow, and recover from incomplete information. We mid-train Qwen3-4B-Base and Qwen3-8B-Base on MidTool-Mix, and then apply follow-up post-training with both supervised fine-tuning and reinforcement learning. Compared with baselines, MidTool-Mix consistently improves downstream performance under both SFT and RL on BFCL, tau2-Bench, and MCP Universe. These results suggest that general tool use, like other important LLM capabilities, benefits from dedicated mid-training rather than being left entirely to post-training.

13:00 JST研究/論文

Pandora の AI モデル ルーティング ボックス: コストのかかる価値推定による効率的な割り当て

複数のモデル、アーキテクチャ、ハーネス、または推論時間設定で構成される異種 AI システムは、最も低コストで最も効果的に回答できる専門家にクエリをルーティングすることで、品質と効率を向上させることができます。ルーティングでは各スペシャリストの期待利益を見積もる必要がありますが、この価値の見積りにはコストがかかります。安価な推定器 (埋め込みベースの予測器など) は高速ですがノイズが多く、正確な推定器 (検索結果や部分的な推論トレースにアクセスできる微調整されたモデルなど) は高価です。このトレードオフを、コストのかかる検査を伴う最適な検索という古典的な問題であるパンドラの箱のインスタンスとして形式化します。ガウス信号モデルの下では、結果として得られるポリシーには閉じた形式の情報価値の式が含まれており、専門家や入力ごとに、価値の見積もりを精緻化することがコストに見合うかどうかを判断します。私たちはこの集中ポリシーを Pandora's Router と呼んでいます。これを分散型設定である Pandora's Bidder に拡張し、クエリを請求するために提示された価格を受け入れる前に、スペシャリストが自己評価に投資するかどうかを独自に決定します。標準的なマルチ LLM ベンチマーク、検索拡張スペシャリスト、可変推論時間推論を備えた LLM の 3 つのドメインにわたる実験では、Pandora のルーターが、高価な推定器へのクエリの頻度がはるかに少ない一方で、網羅的な推定のルーティング品質と一致することが示されました。分散型設定では、競合する推定値が正確である場合、情報価値推論により割り当て効率が向上します。ただし、競合する見積もりにノイズが多い場合は、他の人を犠牲にして戦略専門家の有用性を高める可能性があります。

原文 (English)

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of others.

13:00 JSTLLM/生成AIエージェント研究/論文

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inheri…

13:00 JSTLLM/生成AIエージェント研究/論文

アクティブなデータ収集、旅行行動モデリング、天候に応じた需要予測のためのエージェント的アプローチ

旅行行動調査では、デジタル データ収集と予測モデリングを組み合わせるケースが増えていますが、これらの段階は別々に開発および評価されることがよくあります。この研究では、会話データ収集、構造化データ処理、行動予測を統合した 3 つのエージェントのワークフローを提案しています。チャットボットが管理する画像拡張型の嗜好調査では、事前に定義された 5 つの気象シナリオにわたって学生の通学者からモードの選択を収集し、454 件の回答者のシナリオ観察結果が得られました。気象関連の関連性は多項ロジット モデルを使用して分析され、ロジスティック回帰とランダム フォレストは機械学習のベンチマークを提供しました。ローカルにデプロイされた 9 つの大規模言語モデル (LLM) は、20 億から 350 億のパラメーターに及び、4 つのゼロショット プロンプトとコンテキストの条件にわたって評価され、ペルソナ、少数ショット、およびビジョンベースの構成を通じて拡張されました。ランダム フォレストは 5 クラス精度 69.6% を達成しましたが、テキストのみのゼロショット LLM はタスク固有のフィッティングを行わない場合、最高の 69.9% に達しました。習慣的な旅行情報が最も一貫した利益をもたらし、エキスパートフレーミングは一般にロールプレイよりも優れたパフォーマンスを示し、習慣的な旅行情報が入手できない場合にはペルソナ情報が最も役立ちました。少数のショットにより、いくつかのモデルの予測が向上し、少数の例の後にゲインが安定します。回答者に表示された同じ気象画像を使用した場合、視覚ベースの最適な構成では 5 クラス精度が 71.5% に達しました。これは、視覚的なコンテキストが選択したモデルに対して追加の予測情報を提供する可能性があることを示しています。全体として、この研究は、会話型調査、構造化データ処理、従来の行動モデリング、機械学習、およびマルチモーダル LLM 予測を監査可能なマルチエージェント ワークフロー内でどのように調整できるかを示しています。

原文 (English)

An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent gains, Expert framing generally outperformed Role-Play, and persona information was most useful when habitual travel information was unavailable. Few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. Using the same weather images shown to respondents, the best vision-based configuration reached 71.5% five-class accuracy, indicating that visual context may provide additional predictive information for selected models. Overall, the study shows how conversational surveys, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within an auditable multi-agent workflow.

13:00 JSTLLM/生成AI

石油技術者協会の実践コミュニティの仮想メンバー: プロトタイプから展開まで

私たちは、石油・ガス分野に関連する実践コミュニティ (CoP) のメンバー向けに知識の収集、検索、普及をサポートするように設計された ATHENA と呼ばれる仮想アシスタントの進化について説明します。石油工学協会 (SPE) の 75 人の専門家が参加した最初のプロトタイプの評価では、ATHENA が最先端の RAG ベースライン システムを使用した場合と比較して、一連の現実的な井戸計画タスクにおける生産性とパフォーマンスの同等性の両方が劇的に向上したことがわかりました。ただし、評価では改善の余地がある領域も特定されました。このペーパーでは、複数文書の検索、回答検証のサポート、およびより重点を置いた積極的な普及の分野における最初のプロトタイプの技術的進歩について説明します。評価結果は、ATHENA のこの拡張バージョンは、最先端のベースラインよりも坑井計画に関連する知識集約的なタスクを完了するための優れたサポートを提供することを示しています。 ATHENA は SPE Research Portal に統合されており、協会の会員による使用のために展開されています。

原文 (English)

A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment

We describe the evolution of a virtual assistant, called ATHENA, designed to support the capture, retrieval, and dissemination of knowledge for members of a Community of Practice (CoP) related to the Oil and Gas sector. An evaluation of a first prototype involving 75 professionals from the Society of Petroleum Engineering (SPE) showed that ATHENA dramatically improved both their productivity and performance equality on a set of realistic well-planning tasks compare to their use of a state-of-the-art RAG baseline system. However, the evaluation also identified areas for improvement. This paper describes technical advances to our first prototype in the areas of multi-document retrieval, support for answer validation, and more focused proactive dissemination. Evaluation results show that this enhanced version of ATHENA provides better support for completing knowledge-intensive tasks related to well planning than does a state-of-the-art baseline. ATHENA has been integrated into the SPE Research Portal and is being deployed for use by the society's membership.

13:00 JSTLLM/生成AI

テキスト要約用のトランスフォーマー モデル: BART、BERT、RoBERTa の比較研究

テキストの要約とは、重要な情報を保持しながら文書を短いバージョンに要約するタスクを指します。自動テキスト要約 (ATS) は、自然言語処理 (NLP) の進歩によって近年急速に発展しました。 ATS メソッドは通常、入力タイプ (単一文書または複数文書の要約など) と出力タイプ (抽出、抽象、およびハイブリッド) によって分類されます。この記事では、トランスフォーマー ベースのモデルと大規模言語モデル (LLM)、特に BERT、RoBERTa、BART に重点を置いて、最新の要約手法を中心にレビューします。アーキテクチャ、事前トレーニング戦略、および抽出的および抽象的な要約タスクへの適合性を検証します。

原文 (English)

Transformer Models for Text Summarization: A Comparative Study of BART, BERT, and RoBERTa

Text summarization refers to the task of condensing a document into a shorter version while preserving its key information. Automatic text summarization (ATS), driven by advancements in natural language processing (NLP), has developed rapidly in recent years. ATS methods are commonly categorized by input type (such as single-document or multi-document summarization) and by output type (extractive, abstractive, and hybrid). This article presents a focused review of modern summarization techniques with an emphasis on transformer based models and large language models (LLMs), specifically BERT, RoBERTa and BART. It examines their architectures, pretraining strategies, and their suitability for extractive and abstractive summarization tasks.

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTGeminiGrok

文献からの名前付きエンティティ認識自動バイオインフォマティクス ソフトウェア

バイオインフォマティクス ソフトウェアとデータベースは、現代の生命科学研究に不可欠な要素ですが、科学文献でのそれらの言及には一貫性がないことが多く、大規模に体系的に特定することが困難です。バイオインフォマティクスリソースの包括的かつ最新のカタログが不足しているため、生物医学知識の自動抽出やデータ分析の合理化に向けた取り組みが妨げられています。ここでは、生物医学文書からバイオインフォマティクス ソフトウェアおよびデータベース (SW/DB) 名を自動的に識別するように設計されたハイブリッド固有表現認識フレームワークである SNAIL を紹介します。 SNAIL は、補完的な語彙および意味論的なモデリング戦略を統合します。字句コンポーネントは、SW/DB 名に特徴的な正書法パターンと文脈上の手がかりを捕捉しますが、意味論コンポーネントは、SciBERT などのトランスフォーマーベースの言語モデルによって生成された文脈上の埋め込みを利用し、明示的なトークン マスキング戦略と組み合わせて、エンティティに焦点を当てた表現を強化します。大規模なトレーニング コーパスは、引用ヒントによる抽出と大規模な言語モデル支援の抽出を統合するハイブリッド パイプラインを通じて自動的に構築されました。 2 つの独立したベンチマーク データセットと実際の研究論文の評価では、SNAIL が、bioNerDS2 などのドメイン固有の手法や ChatGPT、Gemini、Grok、Claude などの汎用大規模言語モデルを含む既存のアプローチを大幅に上回るパフォーマンスを示しています。 SNAIL を大規模な文献分析に適用すると、バイオインフォマティクスの下位分野にわたるジャーナルレベルの明確な選好がさらに明らかになります。これらの結果は、SNAIL が科学文書内のバイオインフォマティクス リソースを特定するための正確かつスケーラブルなソリューションを提供し、ツールの使用状況と研究傾向の体系的なメタ分析を可能にすることを示しています。

原文 (English)

Automatic bioinformatic software named entity recognition from literature

Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

13:00 JSTLLM/生成AI

非対称アテンション ヘッド: トランスフォーマー アテンションのための構造化されたヘッドごとのコンテキスト割り当て

標準的なマルチヘッド アテンション (MHA) では、すべての頭に同じ完全な因果的コンテキスト スパンが与えられますが、各頭は異なるコンテキストの役割を果たすことができます。一部のヘッドは主に近くの語彙または構文コンテキストに依存する場合がありますが、他のヘッドはエンティティの相互作用、談話リンク、または状態の変化などのより広範囲の関係に依存する場合があります。我々は、コンテキスト長を明示的なヘッドごとまたはグループごとの割り当て変数として扱う、ヘッドごとのコンテキスト割り当てフレームワークである非対称アテンション ヘッド (AAH) を提案します。 AAH は、特徴派生統計を使用してヘッドをグループ化し、これらのグループを階層的に編成し、標準のフラット MHA 出力インターフェイスを維持しながら因果的なローカル ウィンドウを割り当てます。 4096 トークン シード 0 の実験では、いくつかの AAH スタイルのローカル割り当てバリアントが、純粋なフル アテンションよりも低い検証損失を達成しました。短い予算のアブレーションでは、固定/ローカル制御が適応階層と競合する可能性がある一方で、安定したローカル割り当てとヘッドウィンドウ割り当て構造が重要であることが示されています。 AAH を、品質と分析のための構造化されたヘッドワイズ コンテキスト割り当てメカニズムとして解釈し、アテンション カバレッジ率 (ACR) が選択されたウィンドウのルーティング診断として報告されます。

原文 (English)

Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention

Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles. Some heads may rely mainly on nearby lexical or syntactic context, while others may depend on longer-range relations such as entity interactions, discourse links, or state changes. We present Asymmetric Attention Heads (AAH), a head-wise context- allocation framework that treats context length as an explicit per-head or per-group allocation variable. AAH groups heads using feature-derived statistics, organizes these groups hierarchically, and assigns causal local windows while preserving the standard flat MHA output interface. In 4096- token seed-0 experiments, several AAH-style local-allocation variants achieve lower validation loss than pure full attention. Short-budget ablations show that stable local allocation and head-window assignment structure matter, while fixed/local controls can be competitive with adaptive hierarchy. We interpret AAH as a structured head-wise context-allocation mechanism for quality and analysis, with Attention Coverage Ratio (ACR) reported as a selected-window routing diagnostic

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

欠陥ではなく機能としての幻覚: 推測的な言語モデルの出力をテスト可能な科学的仮説に変換するためのマルチエージェント アーキテクチャの評価

現代の大規模言語モデル (LLM) は、組み合わせによる創造性よりも事実の検索を優先し、幻覚を抑制するようますます調整されています。この調整は、誤った情報を軽減するために重要である一方で、この研究が運用上意味論的な過剰適合や多様性の崩壊として扱うものを奨励することにより、投機的な研究開発 (R&D) を制限する可能性もあります。この論文では、物語の空想と実行制御の対比を、神経認知的な主張としてではなく、機能的なアナロジーとして使用する、Rust ベースのマルチエージェント オーケストレーションを提案します。このシステムは、ノイズと反復を減らすことを目的とした低エントロピーの意味論的ボトルネックを介して、高エントロピー生成エージェントとウェブに接地された評価エージェントの間の認識論的摩擦ループを引き起こします。初期の実験では、物理学および社会科学の領域にわたって、実行可能性を評価した多様な仮説が生成されました。さらに、システム全体を直接プロンプト、自己反映、セマンティックフィルターの除去、検索グラウンディングの除去、側方レンズの除去と比較した、探索的なベースラインとアブレーションのペア研究を報告します。この結果では、直接プロンプトは、ほとんどの観察された指標にわたって最も弱い条件の中に位置づけられていますが、システム全体が単純な内省よりも一般的に優れていることは示されていません。その代わりに、彼らは、各アーキテクチャが独自性、実現可能性、多様性、経験的根拠の間のバランスをさまざまな方法で変化させ、仮説が物理的、経験的、または制度的な強い制約に耐えなければならない場合、完全なシステムが主な利点を提供することを示唆しています。これらの発見は、幻覚が単独で役立つことを示しているわけではありません。彼らは、投機的生成は、アーキテクチャ、経験的根拠、および明示的な評価によって制約された場合にのみ価値を得る、と示唆しています。

原文 (English)

Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses

Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also restrict speculative Research and Development (R&D) by encouraging what this work operationally treats as semantic overfitting and diversity collapse. In this paper, we propose a Rust-based multi-agent orchestration that uses the contrast between narrative daydreaming and executive control as a functional analogy, not as a neurocognitive claim. The system instigates an Epistemological Friction loop between a high-entropy generating agent and a web-grounded evaluating agent, mediated by a low-entropy semantic bottleneck intended to reduce noise and repetition. Initial experiments generated diverse, viability-rated hypotheses across physical and social-science domains. We additionally report an exploratory paired baseline and ablation study comparing the full system against direct prompting, self-reflection, removal of the semantic filter, removal of search grounding, and removal of lateral lenses. The results place direct prompting among the weakest conditions across most observed metrics, but they do not show a general superiority of the full system over simple self-reflection. Instead, they suggest that each architecture shifts the balance between originality, feasibility, diversity, and empirical grounding in different ways, and that the full system provides its main advantages when hypotheses must survive strong physical, empirical, or institutional constraints. These findings do not show that hallucination is useful in isolation; they suggest that speculative generation gains value only when constrained by architecture, empirical grounding, and explicit evaluation.

13:00 JST研究/論文

MLベースのヘリコプタ重量推定器の機内実装に向けて

この論文は、エアバスの世界中の運航中のフリートからの広範なデータセットを利用して、離陸中のヘリコプターの重量を推定するための新しい教師付き機械学習モデルの実装に焦点を当てています。この研究では、機械学習アプリケーションの EASA コンセプト ペーパーおよび進行中の Eurocae ED-324 に沿った学習保証プロセスについて詳しく説明します。私たちは、一連の機械学習要件、機械学習モデルの説明、および長期短期記憶リカレント ニューラル ネットワークの実装を提案します。最後に実装上の要件を確認します。従来のアビオニクス コンピューターで実証されたこの実装は、機内警報などの重要な機能を目的として、開発された機械学習モデルの重量推定器を空中ターゲットに展開するのに適しています。

原文 (English)

Towards On-Board Implementation of ML-Based Helicopter Weight Estimator

This paper focuses on the implementation of a novel supervised Machine Learning model for estimating helicopter weight during takeoff, utilizing extensive datasets from Airbus's global in-service fleet. The study details a learning assurance process aligned with the EASA concept paper for machine learning application, and with the on-going Eurocae ED-324. We propose a set of Machine Learning Requirements, a Machine Learning Model Description, and its implementation for a long short-term memory recurrent neural network. Finally, we verify the requirements on the implementation. Demonstrated on legacy avionics computers, the implementation is suitable for the deployment of the developed Machine Learning Model weight estimator on airborne targets for critical functions such as on-board alerting.

13:00 JSTLLM/生成AI

表現されているが無視されている: 音声言語モデルにおける韻律の過小使用の因果関係

人間の音声は表現力豊かであり、韻律には語彙の内容を超えた言語的および感情的な情報が含まれています。したがって、有能な大規模な音声言語モデル (audio-LLM) は、表現力豊かな音声理解をサポートし、話された内容を文字に起こすだけでなく、話された方法を解釈する必要があります。しかし、行動評価だけでは、モデルが韻律入力で失敗する理由を明らかにすることはできません。エラーは、音響情報の損失、誤った内部解釈、またはモデル内ですでに利用可能な表現の使用の失敗を反映している可能性があります。オーディオ LLM のこれらの障害モードを特定するためのステージ固有のプローブ ラダーを導入します。 4 つの理解専用オーディオ LLM にわたって、韻律情報は通常、オーディオ パスに保存され、後期 LLM 状態でデコード可能です。ただし、モデルの最終応答では部分的にしか表現されていません。私たちは、ターゲットを絞った隠れ状態介入を使用して、この潜在表現の因果関係をテストします。すべての介入は、回答分布を予測された方向にシフトします。ほとんどのモデル - タスク セルでは、抑制された韻律決定に向けてモデルを駆動するには、関連する層での 1 回の編集で十分です。ただし、この回復は正しいクラスの選択的復元ではなく方向性があります。特徴レベルの分析は、この回復可能な信号が小さな部分空間を通じて表現できることをさらに示唆しています。この分析で最も高い属性の特徴のいくつかは、韻律情報を伝えることが知られている音響キューと一致します。私たちがテストした一致コンテンツの対比内で、これらの結果は、韻律の認識ではなく、韻律の使用に繰り返し発生するボトルネックを特定します。韻律上の手がかりを聞いて正しく表現したモデルでも、回答でそれを表現できない可能性があります。

原文 (English)

Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models

Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding, not only transcribing what was said but also interpreting how it was said. Yet behavioral evaluations alone cannot reveal why a model fails on prosodic input. An error may reflect loss of acoustic information, incorrect internal interpretation, or failure to use a representation that is already available inside the model. We introduce a stage-specific probe ladder for localizing these failure modes in audio-LLMs. Across four understanding-only audio-LLMs, prosodic information is usually preserved in the audio path and decodable in late LLM states. Yet it is only partially expressed in the model's final response. We test the causal status of this latent representation with targeted hidden-state interventions. Every intervention shifts the answer distribution in the predicted direction, and in most model--task cells a single edit at the relevant layer is sufficient to drive the model toward the suppressed prosodic decision, though this recovery is directional rather than a selective restoration of the correct class. Feature-level analysis further suggests that this recoverable signal can be expressed through a small subspace. Some of the highest-attribution features in this analysis align with acoustic cues known to carry prosodic information. Within the matched-content contrasts we test, these results locate the recurring bottleneck not in perceiving prosody but in using it. Models that hear and correctly represent a prosodic cue can still fail to express it in their answers.

13:00 JSTLLM/生成AIエージェント

残りの耐用年数におけるマルチモーダル言語モデルの基礎を築くための時系列検索

大規模言語モデル (LLM) とエージェント AI システムは、ドメイン固有のメンテナンスと予測タスクのためにますます検討されており、それらが予測と健康管理 (PHM) を効果的にサポートできるかどうかという疑問が生じています。この論文では、時系列検索に基づいたマルチモーダル大規模言語モデル (MLLM) を使用した残存耐用年数 (RUL) の推定を調査します。私たちは、歴史的に類似した劣化セグメントをトレーニング セットから取得し、テスト軌跡とともに、構造化されたマルチモーダル プロンプトを通じて MLLM によって処理される視覚的な比較アーティファクトに変換するフレームワークを提案します。このアプローチは、ランダムな参照選択に基づく非検索ベースラインに対して検索ベースの推論を比較する繰り返し実験の下で、C-MAPSS ベンチマークの FD001 パーティションで評価されます。結果は、時系列取得により、評価されたモデル全体で MLLM ベースの RUL 予測が一貫して向上し、エラーが減少し、パフォーマンスがより安定していることがわかります。同時に、利点の大きさはモデルの能力に依存し、基礎となる MLLM が取得した証拠を利用できる場合に取得が最も効果的であることを示しています。全体として、この研究は、時系列 RAG がマルチモーダルな予後推論を改善するための有望なメカニズムであることを示していると同時に、実際の PHM 設定における MLLM ベースの RUL 推定の現在の限界も強調しています。

原文 (English)

Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life

Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM). In this paper, we investigate remaining useful life (RUL) estimation with multimodal large language models (MLLMs) grounded through time-series retrieval. We propose a framework in which historically similar degradation segments are retrieved from the training set and, together with the test trajectory, transformed into a visual comparison artifact that is processed by the MLLM through a structured multimodal prompt. The approach is evaluated on the FD001 partition of the C-MAPSS benchmark under repeated experiments comparing retrieval-based inference against a non-retrieval baseline based on random reference selection. The results show that time-series retrieval consistently improves MLLM-based RUL prediction across the evaluated models, yielding lower error and more stable performance. At the same time, the magnitude of the benefit depends on model capacity, indicating that retrieval is most effective when the underlying MLLM is able to exploit the retrieved evidence. Overall, the study shows that time-series RAG is a promising mechanism for improving multimodal prognostic reasoning, while also highlighting the current limitations of MLLM-based RUL estimation in practical PHM settings.

13:00 JSTLLM/生成AIGPT / ChatGPT

会話型 AI は私たちと彼らの境界を緩めることができるでしょうか?共通、二重、個別のアイデンティティの枠組みが移民支持派のグループ間援助に及ぼす影響

移民の増加により、多くの国でグループ間の緊張が高まっています。従来のバイアス削減プログラムは依然として規模拡大が難しく、米国の政策による制約が増えています。この事前登録された実験では、会話型 AI が、多数派グループのメンバーがラテン系移民を分類し、どのように関係するかを変えることができるかどうかをテストしました。共通の内集団アイデンティティ モデルに基づいて、658 人の非ラテン系白人の米国成人からなるクォータ代表の全国サンプルが、LLM (GPT-4o) との 5 ラウンドの対話を完了しました。モデルには、ラテン系移民を共通の内集団アイデンティティ(共通のアメリカ人アイデンティティ)、二重アイデンティティ(ラテン系とアメリカ人の両方)、または別個のアイデンティティ(明確な文化的境界)の観点から枠組み化するか、対照条件で無関係なトピックについて議論するよう指示された。この操作により分類が変更されました。コントロールと比較して、共通のグループ内アイデンティティと二重アイデンティティの会話は個別の分類を低下させ、二重アイデンティティの会話は二重分類を高めました。行動や多様性擁護の信念に対する直接的な影響は有意ではなかったが、上位のアイデンティティ(共通の内集団および二重アイデンティティ)を強調する条件では行動意欲が有意に高かった。経路モデルにより、間接的な関連性がさらに明らかになりました。両方の条件により個別の分類が減少し、その結果、行動意欲の向上と相関関係が生じました。トランスクリプトの意味的類似性分析により、会話が割り当てられた物語を追跡していることが確認されました。参加者の共通アイデンティティ言語への収束は、行動意欲と積極的に関連し、分離アイデンティティ言語への集中は否定的に関連した。これらの効果はモデレーター全体でほぼ一貫していました (閉鎖性の必要性、経験へのオープンさ、政治的方向性)。この研究結果は、AI による短い会話が、認知の再分類と行動の間のギャップを強調しながら、私たちと彼らの境界線を緩める可能性があることを示しています。

原文 (English)

Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping

Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether conversational AI can shift how majority-group members categorize and relate to Latine immigrants. Drawing on the common ingroup identity model, a quota-representative national sample of 658 non-Latine White U.S. adults completed five rounds of dialogue with a LLM (GPT-4o). The model was instructed to frame Latine immigrants in terms of a common ingroup identity (a shared American identity), a dual identity (both Latine and American), or a separate identity (distinct cultural boundaries), or to discuss an unrelated topic in a control condition. The manipulations altered categorization: relative to control, common ingroup identity and dual identity conversations lowered separate categorization, and dual identity conversations raised dual categorization. Although direct effects on behavior and pro-diversity beliefs were nonsignificant, willingness to act was significantly higher in the conditions emphasizing a superordinate identity (common ingroup and dual identity). A path model further revealed indirect associations: both conditions reduced separate categorization, which in turn correlated with greater willingness to act. Semantic similarity analyses of the transcripts confirmed that conversations tracked their assigned narratives; participants' convergence with shared-identity language related positively, and with separate-identity language negatively, to willingness to act. These effects were largely consistent across moderators (need for closure, openness to experience, and political orientation). The findings show that brief AI conversations can loosen us-versus-them boundaries while underscoring the gap between cognitive recategorization and behavior.

13:00 JST研究/論文

学習された露出マッピングとの干渉下での因果推論

暴露マッピングは、因果波及分析では既知であると想定されることがよくあります。しかし、環境設定では、これらは通常、直接観察されない輸送プロセスによって引き起こされるため、代わりに汚染データから学習する必要があります。私たちは、学習された輸送プロセスの不確実性が、干渉下での暴露マッピングと下流への波及推論にどのように伝播するかを研究します。私たちは、シミュレーション研究とカリフォルニア州 PM$_{2.5}$ データの実証分析の両方を使用して、機械的輸送モデルを、PDE、PINO、FNO、GeoPT などの最新のオペレーター学習アプローチと比較します。シミュレーションでは、4 つの輸送モデルすべてがほぼ同一の汚染予測精度を達成しましたが、推定波及効果の範囲は 1.78 ~ 2.27 でした。誘発曝露マッピングをより正確に復元したモデルは、真の影響に近い波及推定値も生成しました。不一致は、地域的な介入では控えめでしたが、局所的な点源介入では大幅に大きくなりました。カリフォルニアの分析でも同じパターンが示されました。競合する輸送モデルでは、観測された PM${2.5}$ 濃度について同様の予測が得られましたが、仮説上の汚染防止介入の下では異なる波及効果が示唆されました。私たちの調査結果は、暴露マッピングが直接観察されるのではなく学習される場合、予測一致だけでは信頼性の高い因果推論には不十分であることを示唆しています。

原文 (English)

Causal Inference under Interference with Learned Exposure Mappings

Exposure mappings are often assumed to be known in causal spillover analyses. In environmental settings, however, they are typically induced by transport processes that are not directly observed and must instead be learned from pollution data. We study how uncertainty in learned transport processes propagates into exposure mappings and downstream spillover inference under interference. We compare mechanistic transport models with modern operator-learning approaches, including PDE, PINO, FNO, and GeoPT, using both simulation studies and an empirical analysis of California PM$_{2.5}$ data. In simulations, all four transport models achieved nearly identical pollution prediction accuracy, yet estimated spillover effects ranged from 1.78 to 2.27. Models that more accurately recovered the induced exposure mapping also produced spillover estimates closer to the true effect. Disagreement was modest for regional interventions but substantially larger for localized point-source interventions. The California analysis showed the same pattern: competing transport models produced similar predictions of observed PM${2.5}$ concentrations while implying different spillover effects under hypothetical pollution-control interventions. Our findings suggest that predictive agreement alone is insufficient for reliable causal inference when exposure mappings are learned rather than directly observed.

13:00 JSTエージェントGPT / ChatGPT

AI が執筆すると誰が引用されるのか?言語モデル間の引用モノカルチャーの証拠

言語モデルが散文の草案から、ツール呼び出しによる文献検索エージェントの実行に移行するにつれて、捏造された参照を捕らえて制限することが容易になってきています。すべての候補が本物だった後に、より深刻な失敗が始まります。異なるモデルが依然として同じ狭いサブセットを選択し、単一の引用が間違っていない引用のモノカルチャーが生成される可能性があります。この影響を 120 の実際の論文で分離します。ベンダー 3 社の 11 個のモデルは、30 人の均一にランダムなパネルから最大 10 件の論文を選択します。その論文のタイトルと要約は実際のものですが、著者は捏造されており、年は再割り当てされ、場所と引用数は隠されています。各実行は、同じパネル上の無関心な選択と実現された予算と比較されます。 11 のモデルすべてが急激に集中しています。上位 10 位はヌル以下の 15.6% に対して 23.3 ~ 30.2% の引用を受け取り、1 つのコンポーネントは優先マップ全体の変動の 68 ~ 73% を説明し、ベンダー間の合意はベンダー内の合意とほぼ一致します。タスクを固定予算のサブセット選択として形式化し、これらのパターンを識別可能なメカニズムに変換します。交換可能限界はすべてのモデルのマップレス セレクターを拒否し、スペクトル分解は、最適な交差適合混合物が依然として過剰の 55% を保持する理由を説明し、希少性定理はパネル内で検証する再帰的競争効果を予測します。制御された言い換え、コンテンツとスロットのクロスオーバー、およびデザインのリサンプリングにより、GPT-5 mini のマップの差異の約 90% が紙のコンテンツに起因します。同じ上限の下で同じ盲検パネルから選択した 8 人のドメイン専門家は、比較可能な共通の優先順位を示さない一方で、モデルの集中は選択のみモードで持続します。すべての参考文献が本物であり、すべての論文が同等に表示される場合でも、現在の言語モデルでは、科学的な注目に共通の内容レベルのフィルターが課されます。したがって、検索を均一にしたり、ベンダーを混合したりするだけでは不十分です。共有設定マップ自体を変更する必要があります。

原文 (English)

When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models

As language models move from drafting prose to running literature-search agents with tool calls, fabricated references are becoming easier to catch and constrain. The harder failure begins after every candidate is real: different models may still select the same narrow subset, producing citation monoculture without any single citation being wrong. We isolate this effect on 120 real papers. Eleven models from three vendors choose at most ten papers from uniformly random panels of thirty, with real titles and abstracts but fabricated authors, reassigned years, and hidden venues and citation counts. Each run is compared with indifferent selection on the same panel and realized budget. All eleven models concentrate sharply: the top decile receives 23.3-30.2% of citations against 15.6% under the null, one component explains 68-73% of variation across their preference maps, and cross-vendor agreement nearly matches within-vendor agreement. Formalizing the task as fixed-budget subset selection, we turn these patterns into identifiable mechanisms: an exchangeability bound rejects a mapless selector for every model, a spectral decomposition explains why the best cross-fitted mixture still retains 55% of the excess, and a rarity theorem predicts the recursive competition effect we verify within panels. Controlled paraphrase, content-slot crossover, and design resampling attribute about 90% of GPT-5 mini's map variance to paper content. Eight domain experts selecting from the same blinded panels under the same cap show no comparable shared preference, while model concentration persists in selection-only mode. Even when every reference is real and every paper is equally visible, current language models impose a common content-level filter on scientific attention. Equalizing retrieval or mixing vendors is therefore insufficient; the shared preference map itself must be changed.

13:00 JST研究/論文

アクティブスパイク知覚: いつでも 3D 点群認識のための信念状態としての膜の可能性

スパイク点群ネットワークは通常、固定された入力に依存しない順序で空間をスキャンするため、スパイク計算の最も特徴的なリソースである膜電位の時間的発展が意思決定の場として使用されずに残ります。 Active Spiking Perception (ASP) は、3D 認識を反復的な意思決定プロセスとして再構築します。このプロセスでは、ネットワーク自身の漏洩統合発火 (LIF) 膜電位がクラス全体にわたる実行中の信念として読み取られ、観察する次のチャンクを選択し、信頼マージンの早期終了をトリガーします。軽量のスライス選択ポリシーは、メンブレン状態と事前計算された幾何学的記述子から未訪問の最遠点サンプリングされたチャンクをスコアリングし、ストレートスルーのガンベル ソフトマックスを通じてエンドツーエンドでトレーニングし、推論時に argmax に削減し、バックボーン パラメーターの約 2% を追加します。漏れのある積分はベイジアン フィルターの再帰的対数事後更新であること、出口ルールは停止時に多重テストのペナルティなしで分布フリーの選択的リスクを達成すること、ストリーミング状態の繰り越しは有限精度ドリフトを伴うプレフィックス再計算とまったく同等であることを証明します。 ASP は、ModelNet40 と ModelNet10 で 90.62% と 93.28% に達し、大規模なバックボーンでの最も強いスパイク ベースラインを 1.7 ポイント下回っていますが、ベースラインでは提供されていない認定済みの常時インターフェイスが追加されています。このメカニズムは変更を加えずに高密度予測に転送し、ShapeNetPart で 83.21 インスタンス mIoU、S3DIS エリア 5 で 48.50 mIoU を与えます。我々の知る限りでは、S3DIS エリア 5 で最初のスパイク結果が得られ、チャンク選択を置き換える固定が中心窩非スパイク変換器に適用されるため、ポリシーはスパイク バックボーンに結び付けられません。コストは観測では正確に線形であり、しきい値は測定されたコンピューティング ダイヤル スパンです。エネルギーが 2.8 倍から 1.35 倍少なくなります。 1 つの具体的な制限があります。1 つの S3DIS クラスが、使用するクロップ サイズでは識別できないため、それを修正する予測を示します。

原文 (English)

Active Spiking Perception: The Membrane Potential as a Belief State for Anytime 3D Point Cloud Recognition

Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distinctive resource of spiking computation, the temporal evolution of the membrane potential, unused as a locus of decision-making. Active Spiking Perception (ASP) recasts 3D recognition as an iterative decision process in which the network's own leaky integrate-and-fire (LIF) membrane potential, read as a running belief over the class, selects the next chunk to observe and triggers confidence-margin early exit. A lightweight Slice-Selection Policy scores unvisited farthest-point-sampled chunks from the membrane state and precomputed geometric descriptors, trains end-to-end through a straight-through Gumbel-Softmax, reduces to an argmax at inference, and adds about 2% of backbone parameters. We prove that leaky integration is the recursive log-posterior update of a Bayesian filter, that the exit rule attains distribution-free selective risk with no multiple-testing penalty at the stopping time, and that streaming state carry-forward is exactly equivalent to prefix recomputation with bounded finite-precision drift. ASP reaches 90.62% and 93.28% on ModelNet40 and ModelNet10, 1.7 points below the strongest spiking baseline at a larger backbone, while adding a certified anytime interface no baseline offers. The mechanism transfers unchanged to dense prediction, giving 83.21 instance mIoU on ShapeNetPart and 48.50 mIoU on S3DIS Area 5, to our knowledge the first spiking results on S3DIS Area 5, and, fixation replacing chunk selection, to a foveated non-spiking transformer, so the policy is not tied to spiking backbones: cost is exactly linear in observations and the threshold is a measured compute dial spanning 2.8x to 1.35x less energy. One limitation is concrete: one S3DIS class is unidentifiable at the crop size we use, and we give the prediction that would fix it.

13:00 JSTLLM/生成AI

LLM アプリケーションの OWASP トップ 10 (2026) のインシデント データの堅牢性分析: コミュニティ専門家のランキングが大規模な LLM インシデント コーパスに対してどのように耐えられるか

OWASP Top 10 for LLM Applications は、セキュリティ専門家のコミュニティが最も重要と判断するリスクをランク付けします。私たちはより狭い質問をします。実際の事件の記録と照合して、その専門家のランキングはデータと一致しますか?私たちは、CVE、GHSA、OSV、および AIAAIC から抽出された LLM セキュリティ インシデントの大規模コーパス (20 エントリの分類法に対してスナップショットが 7,714 件、ラベルが付けられた 6,639 件) を収集し、分類器の精度と再現率について各カテゴリのカウントを補正するベイジアン測定誤差モデルを使用してインシデントベースのランキングを導き出しました。 2026 年の候補者リストは、2 つのシグナルを固定の重み (専門家の投票に 0.75、データに 0.25) でブレンドしているため、コーパスはコンセンサスを覆すことなく修正します。 2 つのランキング間の一致は弱く、コーエンの $\kappa \約 0.20$ で、90% の間隔がゼロを超えます。それにもかかわらず、専門家のランキングは堅実です。 4 つのフロンティア分類子の事前登録されたベイクオフでは、勝者は返されません。入射下限のバランス精度 0.863 を超えるものはありません。真実のチェックでは、フロアの順序付け (差し出された真実に対してスピアマン $\rho = 0.918$) がそのまま残ります。これは 2 人のワーキング グループ メンバーによる探索的分析であり、OWASP の公式リリースではなく、公式のリストやプロセスに代わるものではありません。

原文 (English)

Incident-Data Robustness Analysis of the OWASP Top 10 for LLM Applications (2026): How a Community-Expert Ranking Holds Up Against a Large-Scale LLM Incident Corpus

The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important. We ask a narrower question: checked against the record of real incidents, does that expert ranking agree with the data? We assembled a large-scale corpus of LLM-security incidents (7,714 snapshotted and 6,639 labeled against the 20-entry taxonomy) drawn from CVE, GHSA, OSV, and AIAAIC, and derived an incident-based ranking with a Bayesian measurement-error model that corrects each category's count for classifier precision and recall. The 2026 candidate list blends the two signals at fixed weights, 0.75 on the expert vote and 0.25 on the data, so the corpus corrects the consensus without overturning it. The agreement between the two rankings is weak: Cohen's $\kappa \approx 0.20$, with a 90% interval that crosses zero. The expert ranking is nonetheless robust. A pre-registered bake-off of four frontier classifiers returns no winner. None beats the incidence floor's balanced accuracy of 0.863. A ground-truth check leaves the floor's ordering (Spearman $\rho = 0.918$ against held-out truth) in place. This is an exploratory analysis by two working-group members, not the official OWASP release, and it does not supersede the official list or process.

13:00 JST研究/論文

20 の AI 中権力管轄区域における汎用 AI ガバナンスのマッピング

最も高性能な汎用 AI (GPAI) モデルは、ほとんどが米国と中国の 2 つの管轄区域で構築されていますが、それらがもたらすリスクは世界中に広がっています。 AI中大国と呼ばれるフロンティア開発者を受け入れていない地域先進国は、GPAIを管理する独自のルールを策定している。このペーパーでは、これらの AI 中大国がどの GPAI 関連規定を制定したかを調査し、欧州連合を含む 20 の法域を個々の規定のレベルでマッピングし、モデル層の責任連鎖を追跡する 4 つのガバナンス領域 (システミック リスクの評価、評価と検証、監視と検出による禁止、重大インシデントの報告) にわたってマッピングします。確認された欠勤は、積極的な提供とともにデータとして記録されます。管轄権は形式に関しては収束しますが、力に関しては発散することがわかります。 16 の機関は 4 つのガバナンス領域のうち少なくとも 3 つに関与していますが、法的拘束力のある規定は 5 分の 1 のみであり、拘束力のある文書の 4 分の 3 は GPAI を定義せずに規定しています。制度的インフラストラクチャも同様の形を示しています。マッピングされたガバナンス主体の 5 人に 4 人が GPAI よりも前の権限を保持しており、継承された体制がすでに到達しているところ、つまりモデルではなくアプリケーション層に義務が課せられています。これらの州がモデル層に関与する場合、モデル層を構築する者に義務を課すのではなく、それを観察する能力を構築し、ほぼすべての評価機関は、発見されたものに基づいて行動する権限を持たずに設立されました。完全な責任チェーンの名目上の範囲は 11 の管轄区域に及んでいますが、すべての分野で複数の規定を設けている管轄区域は 5 つだけであり、EU の外では、モデル開発者に拘束力のある評価義務を課している管轄区域はありません。このデータセットは、研究者や政策立案者に、体制がどこで調整できるか、どこで調整をゼロから開始する必要があるかを特定するための規定レベルの基礎を提供します。

原文 (English)

Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions

The most capable general-purpose AI (GPAI) models are mostly built in two jurisdictions, the United States and China, but the risks they carry land globally. Regionally advanced economies hosting no frontier developer, which we call AI middle-powers, are writing their own rules to govern GPAI. This paper investigates which GPAI-relevant provisions these AI middle-powers have enacted, mapping twenty jurisdictions including the European Union at the level of the individual provision, across four governance areas that trace the accountability chain for the model layer: systemic risk assessment, evaluation and verification, prohibitions with monitoring and detection, and serious incident reporting. Confirmed absence is recorded as data alongside positive provision. We find that jurisdictions converge on form, but diverge on force. Sixteen engage in at least three of the four governance areas, yet only about one in five provisions sit in binding law, and three-quarters of the instruments that do bind do so without defining GPAI. The institutional infrastructure shows the same shape: four in five of the mapped governance actors hold mandates that predate GPAI, and obligations attach wherever the inherited regime already reached, which is the application layer rather than the model. Where these states engage the model layer, they build capacity to observe it rather than impose duties on those who build it, and almost every evaluation body was constituted without the power to act on what it finds. Nominal coverage of the full accountability chain reaches eleven jurisdictions, but only five hold more than one provision in every area and, outside the EU, no jurisdiction imposes a binding evaluation duty on a model developer. The dataset gives researchers and policymakers a provision-level basis for identifying where regimes could align, and where coordination would have to start from scratch.

13:00 JST研究/論文

肺がんの早期発見のための量子カーネル推定

低線量胸部コンピューター断層撮影法による肺がんスクリーニングは死亡率を減少させますが、その効果は、摂取、遵守、および管理の課題によって限定されます。血液ベースの無細胞 DNA (cfDNA) バイオマーカーは補完的なアプローチを提供しますが、肺がんの不均一性と高次元の非線形分子シグナルのため、早期検出は依然として困難です。 DNA フラグメントミクスと DNA メチル化を使用した肺がん検出のための量子古典ハイブリッド機械学習を評価しました。特徴の選択後、モデルは 20 および 40 の特徴サブセットを使用してトレーニングされました。特徴は、複数のもつれ戦略を備えた角度および密角度特徴マップを使用して、量子ヒルベルト空間にエンコードされました。忠実度ベースの量子カーネルは、正確な状態ベクトル シミュレーションで計算され、事前計算されたカーネル SVM およびカーネル PCA ロジスティック回帰と統合され、元の特徴でトレーニングされた SVM モデルと比較されました。このフレームワークにより、エンコードとエンタングルメントの設計が分類にどのような影響を与えるかを体系的に評価できるようになりました。ホールドアウト評価を繰り返した結果、量子カーネル モデルは両方のデータセットで競争力のあるパフォーマンスを達成しました。フラグメントミクスの場合、いくつかの 20 特徴構成により、古典的な SVM ベースラインと比較して AUC が向上し、非線形 cfDNA 断片化構造を効果的に捕捉できることが示唆されました。メチル化に関しては、古典的な SVM が最高の AUC を達成しましたが、選択された量子モデルは競争力を維持し、場合によっては特異性が向上しました。機能を 20 から 40 に増やしても、一貫してパフォーマンスが向上するわけではなく、多くの場合、ばらつきが大きくなります。全体として、これらの結果は、量子カーネル法が cfDNA ベースの肺がん検出の有望なアプローチであることを裏付けています。

原文 (English)

Quantum Kernel Estimation for the Discovery of Early Lung Cancer Detection

Lung cancer screening with low-dose chest computed tomography reduces mortality, but its impact is limited by uptake, adherence, and management challenges. Blood-based cell-free DNA (cfDNA) biomarkers offer a complementary approach, although early detection remains difficult because of lung cancer heterogeneity and high-dimensional, nonlinear molecular signals. We evaluated quantum-classical hybrid machine learning for lung cancer detection using DNA fragmentomics and DNA methylation. After feature selection, models were trained using 20- and 40-feature subsets. Features were encoded into quantum Hilbert space using angle and dense-angle feature maps with multiple entanglement strategies. Fidelity-based quantum kernels were computed with exact statevector simulation and integrated with precomputed-kernel SVM and kernel-PCA logistic regression and compared with an SVM model trained on the original features. This framework enabled systematic evaluation of how encoding and entanglement design affect classification. Across repeated held-out evaluations, quantum-kernel models achieved competitive performance on both datasets. For fragmentomics, several 20-feature configurations improved AUC relative to a classical SVM baseline, suggesting effective capture of nonlinear cfDNA fragmentation structure. For methylation, the classical SVM achieved the highest AUC, although selected quantum models remained competitive and improved specificity in some cases. Increasing features from 20 to 40 did not consistently improve performance and often increased variability. Overall, these results support quantum kernel methods as a promising approach for cfDNA-based lung cancer detection.

13:00 JSTLLM/生成AI

ブラックボックスの大規模言語モデルの信頼推定値の向上

不確実性定量化 (UQ) は、大規模言語モデル (LLM) を安全に展開するために不可欠です。言語化された信頼性から複数の世代を必要とするものまで、既存の方法は多くの場合ゼロショットであり、ラベル付きデータを必要とせずに不確実性を定量化するスコアを生成します。それにもかかわらず、実際には、展開前に対象のデータセットでのパフォーマンスを常に評価する必要があります。この研究では、このデータセットを活用することで、これらの既存のスコアを常に上回るパフォーマンスを示します。具体的には、これらのスコアと同様のクエリの正しさを特徴として使用して、LLM 応答の正しさを予測する単純な分類器を構築します。私たちの方法では計算オーバーヘッドが最小限に抑えられ、現実世界のアプリケーション向けに LLM の UQ を安価で簡単に強化できます。

原文 (English)

Improved Confidence Estimates for Black-Box Large Language Models

Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). Existing methods, from verbalized confidence to ones requiring multiple generations, are often zero-shot and produce scores quantifying uncertainty without the need for labelled data. Nonetheless, in practice one must always evaluate their performance on a dataset of interest before deployment. In this work we show that, by leveraging this dataset, we consistently outperform these existing scores. Specifically, we build simple classifiers that predict LLM response correctness by using these scores and the correctness of similar queries as features. Our method produces minimal computational overhead, making it a cheap and straightforward enhancement for UQ in LLMs for real-world applications.

13:00 JST研究/論文GPT / ChatGPTQwen

機械的断層撮影: 制御指向の解釈可能性を実現するために設計された測定

機械論的な解釈可能性は、モデルが直接明らかにしない量、つまり表現された状態、コンポーネントの効果、相互作用、介入に対する反応を求めます。パッチング、勾配、ヘシアンベクトル積、およびサブセット介入は、異なるアクセス仮定の下で異なる測定を提供し、異なる量をターゲットにする可能性があります。私たちはそれらの共通の測定構造を機械的断層撮影法として定式化します。つまり、内部メカニズムと介入効果を回復するために設計された測定です。選択した基底および介入ファミリーの場合、測定値は y = Ax + w の形式になります。ここで、A は介入を表し、x はターゲット マップであり、w には非線形応答、サンプリング誤差、および基底の誤った指定が含まれます。この言語は実践的な手順を提供します。つまり、最もコストの低い測定から開始し、意図したスケールで保留された介入をテストし、単純な不一致を校正し、構造化された残差が残っている場合は測定ファミリーを拡張します。介入を導く推定値がオブザーバーとして機能するため、コントロールは要求の厳しい検証設定を提供します。 2 つの HMM モデルでは、制御誤差は観察者誤差とともに増加しますが、ターゲットの改善により迷惑な状態の動きを隠すことができます。順方向のみのアクセスでは、疎集合測定は、座標パッチングよりも少ない介入で有限効果マップを回復します。勾配アクセスでは、有限プローブによりローカル アトリビューション マップが改善されます。リフト測定とヘシアンベクトル積は、一次マップで見逃された相互作用を回復しますが、Tracr は、必要なファミリーが基底に依存することを示します。 GPT-2 の小さい IOI では、Name Mover と Negative Name Mover の相互作用は、テストされた 3 つのグループ間ペアの中で最大のホールドアウト予測項です。 Qwen-2.5-7B では、有限キャリブレーションにより加算的な拒否応答マップが適切になるため、ホールドアウト誤差はペアワイズ リフティングをサポートしません。

原文 (English)

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities. We formulate their shared measurement structure as mechanistic tomography: designed measurement for recovering internal mechanisms and intervention effects. For a chosen basis and intervention family, measurements take the form y = Ax + w, where A describes the interventions, x is the target map, and w contains nonlinear response, sampling error, and basis misspecification. This language gives a practical procedure: start with the least costly measurements, test on held-out interventions at the intended scale, calibrate simple mismatch, and expand the measurement family when structured residuals remain. Control provides a demanding validation setting because an estimate that guides an intervention acts as an observer. In a two-HMM model, control error rises with observer error, while target improvement can hide nuisance-state movement. Under forward-only access, sparse aggregate measurements recover a finite-effect map with fewer interventions than coordinate patching. With gradient access, finite probes improve a local attribution map. Lifted measurements and Hessian-vector products recover interactions missed by first-order maps, while Tracr shows that the required family depends on the basis. On GPT-2-small IOI, the Name Mover-Negative Name Mover interaction is the largest held-out predictive term among three tested cross-group pairs. On Qwen-2.5-7B, finite calibration makes an additive refusal-response map adequate, so held-out error does not support pairwise lifting.

13:00 JST画像/動画生成

マージナルカバレッジは、シフト中のゼロショット VLM のクラス条件付き安全性を保証しますか?

スプリットコンフォーマル予測は、交換可能性の下で限界範囲を提供し、ゼロショット ビジョン言語モデル (VLM) の棄権層として使用されることが増えています。私たちは、ImageNet 設定と非 ImageNet 設定にわたる CLIP、OpenCLIP、および SigLIP の展開シフトの下でこの実践を監査します。クラス条件付きテール カバレッジが崩壊する一方で、マージナル カバレッジは比較的高いままである可​​能性があります。ImageNet-Sketch では、マージナル カバレッジが約 0.86 であるにもかかわらず、最悪クラスのカバレッジは $\およそ 0$ に低下し、クラスの 10 ~ 12% が有限サンプルのヌル下限を下回っています。この障害はターゲット ドメインのクラスの精度に合わせて調整されますが、テストするソース ドメインの診断では予測されません。ソース側のモンドリアン キャリブレーションは分布内テールを改善しますが、移行はしません。一方、クラスター化コンフォーマルおよび Conf-OT は最悪クラスのテールを回復することなく周辺メトリクスまたは平均メトリクスを改善します。ターゲット側のクラス キャリブレーションにより、テールは大幅に改善されますが、すべてのクラスにラベルが必要であり、依然としてセット サイズが集中します。さらに、2-3$\times$ の家族間効率ギャップを特定し、ネイティブ SigLIP シグモイド スコアが APS の確率質量解釈を除去することを示します。結果は、テストされたモデル スケール、事前トレーニング コーパス、プロンプト、誤検出レベル $\alpha$、およびシフトされた非 ImageNet 設定にわたって持続します。したがって、限界等角被覆率は、クラステールの安全性の保証としてではなく、平均的な信頼性統計として扱う必要があります。

原文 (English)

Does Marginal Coverage Guarantee Class-Conditional Safety for Zero-Shot VLMs Under Shift?

Split-conformal prediction provides marginal coverage under exchangeability and is increasingly used as an abstention layer for zero-shot vision-language models (VLMs). We audit this practice under deployment shift for CLIP, OpenCLIP, and SigLIP across ImageNet and non-ImageNet settings. Marginal coverage can remain relatively high while class-conditional tail coverage collapses: on ImageNet-Sketch, worst-class coverage falls to $\approx 0$ and 10-12% of classes lie below a finite-sample null floor, despite marginal coverage of about 0.86. The failure is aligned with target-domain class accuracy but is not predicted by the source-domain diagnostics we test. Source-side Mondrian calibration improves the in-distribution tail but does not transfer, while clustered conformal and Conf-OT improve marginal or average metrics without recovering the worst-class tail. Target-side class calibration substantially lifts the tail, but requires labels for every class and remains set-size-intensive. We further identify a 2-3$\times$ cross-family efficiency gap and show that native SigLIP sigmoid scores remove APS's probability-mass interpretation. The findings persist across the tested model scale, pretraining corpus, prompt, miscoverage level $\alpha$, and shifted non-ImageNet settings. Marginal conformal coverage should therefore be treated as an average reliability statistic, not as a safety guarantee for the class tail.

13:00 JST研究/論文

公平性を意識したネットワークの埋め込み: 方法、アプリケーション、および課題

ネットワーク埋め込み手法は、グラフ構造データの低次元表現を学習して、ノード分類、リンク予測、影響の最大化などの下流タスクをサポートします。ただし、現実世界のネットワークは、人口動態の不均衡、同性愛、その他の社会的偏見から生じる構造的不平等を反映していることが多く、公平性にとらわれない埋め込み手法はこれらを符号化して増幅する可能性があります。この問題に対処するために、埋め込みの有用性を維持しながらバイアスを軽減するために、公平性を意識したネットワーク埋め込み手法が多数提案されています。この調査では、複雑なネットワークに対する公平性を意識したネットワーク埋め込みの包括的な概要を示します。私たちは、基礎となる埋め込みアプローチ (スペクトル、ランダム ウォーク、グラフ ニューラル ネットワーク、ベイジアン、および方法に依存しない)、公平性介入戦略 (前処理、処理中、後処理)、および公平性の目標基準 (埋め込みレベルまたはタスク レベル) という 3 つの主要な相補的側面に沿って既存の手法を分類する分類法を提案します。さらに、グループと個人の公平性と機密属性に関する仮定に関して方法を比較します。最後に、現在の限界について説明し、将来の有望な研究の方向性を強調します。この調査は、公平性を意識したネットワークの埋め込みに関する統一的な観点を提供し、公平で信頼できるネットワーク表現の学習方法を開発するための参考として役立ちます。

原文 (English)

Fairness-Aware Network Embeddings: Methods, Applications, and Challenges

Network embedding methods learn low-dimensional representations of graph-structured data to support downstream tasks such as node classification, link prediction, and influence maximization. However, real-world networks often reflect structural inequalities arising from demographic imbalances, homophily, and other societal biases, which fairness-agnostic embedding methods can encode and amplify. To address this issue, numerous fairness-aware network embedding methods have been proposed to mitigate bias while preserving embedding utility. This survey presents a comprehensive overview of fairness-aware network embeddings for complex networks. We propose a taxonomy that categorizes existing methods along three main complementary dimensions: underlying embedding approach (spectral, random walk, graph neural network, Bayesian, and method-agnostic), fairness intervention strategy (pre-processing, in-processing, and post-processing), and fairness objective criterion (embedding- or task-level). We further compare methods with respect to group versus individual fairness and assumptions regarding sensitive attributes. Finally, we discuss current limitations and highlight promising future research directions. This survey provides a unified perspective on fairness-aware network embedding and serves as a reference for developing fair and trustworthy network representation learning methods.

13:00 JST研究/論文

集中的な流動性の提供: 強化学習の観点

自動マーケットメーカー (AMM) は、分散型金融 (DeFi) の基礎です。 UniswapV3 など、流動性が集中した定常的な製品市場は、現在では十分に確立された設計となっています。これらの市場では、流動性プロバイダー (LP) は、市場状況の変化に応じて、いつポジションをリバランスするか、どの価格帯に資本を割り当てるかを決定する必要があるという、一連の意思決定の問題に直面しています。私たちは動的流動性の提供を確率的インパルス制御問題として定式化し、強化学習 (RL) を使用してそれを解決し、解釈可能なソリューションを提供することに重点を置いています。学習した政策は、ミスプライシング、リバランスコスト、不確実性、在庫エクスポージャ、異種リスク選好に従って流動性を配分する、豊富な状態依存行動を示すことを示します。これらの動作は、損益 (PnL) 分布の左端を圧縮し、不確実性が高い状況での壊滅的な結果を回避するのに役立ちます。最後に、AMM 微細構造文献からのベースラインおよび高度なエージェントに対して RL エージェントをベンチマークし、そのパフォーマンスを分析します。

原文 (English)

Concentrated Liquidity Provision: a Reinforcement Learning Perspective

Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi). Constant product markets with concentrated liquidity, such as UniswapV3, are now a well-established design. In these markets, liquidity providers (LPs) face a sequential decision problem: they must decide when to rebalance their positions and which price ranges to allocate capital to as market conditions evolve. We formulate dynamic liquidity provision as a stochastic impulse control problem and use reinforcement learning (RL) to solve it, focusing on providing interpretable solutions. We show that learned policies exhibit rich state-dependent behaviour, allocating liquidity according to mispricing, rebalancing costs, uncertainty, inventory exposure, and heterogeneous risk preferences. These behaviours help compress the left tail of the Profit and Loss (PnL) distribution and avoid catastrophic outcomes under high uncertainty. Finally, we benchmark the RL agents against baseline and sophisticated agents from the AMM microstructure literature and analyse their performance.

13:00 JSTLLM/生成AIハードウェア/半導体Claude

HYDRA: 動的なハイブリッド LLM ワークロードを提供するための異種チップレット DSE フレームワーク

ハイブリッド Transformer-Mamba ラージ言語モデル (LLM) は、ロングコンテキストの効率を高めますが、その異種混合の計算と通信パターンにより、効率的なハードウェア アクセラレーションが複雑になります。チップレット ベースのアーキテクチャは、特殊なコンピューティング ユニットとメモリ ユニットを統合することにより、スケーラブルなソリューションを提供します。ただし、静的なアーキテクチャ構成と動的なランタイム ポリシーにまたがる設計領域は、徹底的に調査するには法外に大きいです。この課題に対処するために、異種チップレット システムで機能するハイブリッド LLM のための包括的な設計空間探索フレームワークである HYDRA を紹介します。 HYDRA は、チップレットの構成、配置、チップレット間の帯域幅プロビジョニング、動的バッチ処理、およびランタイム スケジューリングを共同で調査しています。通信を意識した配置、動的バッチ処理、柔軟なタスク スケジューリング、およびマルチテナントのランタイム ダイナミクスをキャプチャする高速マルコフ ベースのパフォーマンス推定機能を統合して、効率的かつ正確な探索を実現します。すべてのワークロードにわたって、HYDRA は平均して 1.55 倍のスループットと、最初のトークンまでの時間を 43.7 パーセント短縮し、最先端のベースラインと比較してスループットの向上は最大 2.3 倍に達します。これらの結果は、異種チップレット システム上で効率的に大規模 LLM を機能させるには、アーキテクチャとランタイム ポリシーを共同設計することが重要であることを浮き彫りにしています。

原文 (English)

HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. Chiplet-based architectures offer a scalable solution by integrating specialized compute and memory units. However, the design space spanning static architectural configurations and dynamic runtime policies is prohibitively large to explore exhaustively. To address this challenge, we present HYDRA, a comprehensive design space exploration framework for hybrid LLM serving on heterogeneous chiplet systems. HYDRA jointly explores chiplet composition, placement, inter-chiplet bandwidth provisioning, dynamic batching, and runtime scheduling. It integrates communication-aware placement, dynamic batching, elastic task scheduling, and a fast Markov-based performance estimator that captures multi-tenant runtime dynamics for efficient and accurate exploration. Across all workloads, HYDRA delivers 1.55x the throughput and 43.7 percent lower time-to-first-token on average, with throughput gains reaching up to 2.3x compared to state-of-the-art baselines. These results highlight that co-designing architecture and runtime policies is critical for efficient large-scale LLM serving on heterogeneous chiplet systems.

13:00 JST画像/動画生成ハードウェア/半導体

HiRA-CAM: グラデーションベースの視覚的説明におけるきめ細かい空間関連性の維持

深層学習モデルには数十億以上のパラメーターが含まれる場合があり、その内部変換と出力を説明することが困難になります。しかし、重要なアプリケーションでの AI の使用により、説明可能性の重要性が増しています。この論文は、畳み込みニューラル ネットワーク (CNN) の解釈可能性に焦点を当てています。 CNN の内部特徴を抽出するための一般的な勾配ベースの手法 LayerCAM を基にして、HiRA-CAM という名前の改良された手法を提案し、この手法がオブジェクト分類に有用な顕著性マップの作成において LayerCAM と Grad-CAM の両方よりも優れていることを示します。 HiRA-CAM の主な特徴は、CNN のすべての層からのアクティベーション マップを適応的に使用して、より焦点を絞った顕著性マップを作成することです。

原文 (English)

HiRA-CAM: Preserving Fine-Grained Spatial Relevance in Gradient-Based Visual Explanations

Deep Learning models can include billions of parameters or more, making it difficult to explain their internal transformations and outputs. However, explainability is increasing in importance due to the use of AI in crucial applications. This paper focuses on the interpretability of convolutional neural networks (CNNs). Building on the popular gradient based method LayerCAM for extracting internal features in CNNs, we propose an improved method named HiRA-CAM, and show that it outperforms both LayerCAM and Grad-CAM on creating useful saliency maps for object classification. The main feature of HiRA-CAM is its adaptive use of activation maps from all the layers of the CNN to arrive at a more focused saliency map.

13:00 JSTロボティクスビジネス/資金調達

SCAPE: シナリオ条件付きシミュレーションによる政策評価の強化

信頼性の高いパフォーマンス評価は、現実世界の状況でロボット学習ポリシーを展開する際の中心的なボトルネックです。現実世界のテストは忠実ですが、コストが高く、拡張が困難です。一方、シミュレーション ベースのテストは拡張が容易ですが、シミュレーションと現実のギャップによって必然的にバイアスがかかります。既存のシミュレーション拡張手法は、限られた現実世界のロールアウトと豊富なシミュレーション プロキシを組み合わせていますが、初期条件と展開設定の平均パフォーマンスに重点を置いています。このような人口レベルの平均は、シナリオ固有の変動を曖昧にし、いつ、どこにポリシーを安全に導入できるかについての限定的な指針を提供します。我々は、限られたシミュレーションと現実のサンプルのペアと大規模なシミュレーションのロールアウトを使用して、シナリオ条件付きの現実世界の政策パフォーマンスを予測する、シナリオ条件付きシミュレーション拡張政策評価フレームワークであるSCAPEを提案します。 SCAPE は、予測モデルをトレーニングする前にシミュレーション ラベルのシミュレーションと実数のバイアスを修正し、等角予測を通じて予測の不確実性を調整します。自動運転と四足歩行の速度追跡におけるSCAPEを検証します。 sim-to-sim 研究では、SCAPE は、シーン条件付きニューラルおよび集計統計ベースラインと比較して、シナリオレベルの予測誤差を平均で 4.9%/34.7% (運転)、および 14.5%/27.7% (四足歩行) 削減しました。さらに、物理的な Unitree Go2 に展開された速度追跡ポリシーを評価します。 SCAPE はまた、サンプルのテスト効率を向上させ、より狭い調整済み予測間隔を生成し、配布外シナリオをより適切に一般化して、きめ細かい展開戦略を可能にします。

原文 (English)

SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation

Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions. Real-world testing is faithful but costly and difficult to scale, whereas simulation-based testing scales easily but is inevitably biased by the sim-to-real gap. Existing simulation-augmented methods combine limited real-world rollouts with abundant simulation proxies, but focus on performance averaged over initial conditions and deployment settings. Such population-level averages obscure scenario-specific variation and provide limited guidance about when and where a policy can be safely deployed. We propose SCAPE, a scenario-conditioned simulation-augmented policy evaluation framework that predicts scenario-conditioned real-world policy performance using limited paired sim-and-real samples and large-scale simulation rollouts. SCAPE corrects sim-to-real bias in simulation labels before training the prediction model and calibrates prediction uncertainty through conformal prediction. We validate SCAPE on autonomous driving and quadruped velocity tracking. In sim-to-sim studies, SCAPE reduces scenario-level prediction error by 4.9%/34.7% (driving) and 14.5%/27.7% (quadruped) relative to scene-conditioned neural and aggregate statistical baselines on average. We further evaluate a velocity-tracking policy deployed on a physical Unitree Go2. SCAPE also improves testing sample efficiency, produces narrower calibrated prediction intervals, generalizes better to out-of-distribution scenarios, and enables fine-grained deployment strategies.

13:00 JST研究/論文

Longitudinal Bayesian Learning of Continuous Disease Position across the Alzheimer's Disease Continuum

Alzheimer's disease (AD) progresses as a continuous biological process, whereas most existing neuroimaging-based artificial intelligence me…

13:00 JSTLLM/生成AI研究/論文

LLM も同様に創造的になってきていますか? 3年間のモデルから得られた証拠

多くのベンチマークは、検証可能な回答を持つタスクにおける大規模言語モデル (LLM) のパフォーマンスを追跡していますが、品質と同じくらい創造性、独創性、多様性が重要である可能性があるオープンエンド型タスクで LLM のパフォーマンスがどのように進化しているかについてはあまり知られていません。 LLM が人間の発想や創造的な作業をますますサポートするようになっているため、オープンエンドのタスクにおける LLM のパフォーマンスの傾向を理解することが重要です。この論文では、3 年間のモデル リリースにわたる LLM クリエイティブ アウトプットの予備分析を示し、自由回答型のユーザー クエリの現実世界のコレクションである Infinity-Chat100 と、確立された心理測定的創造性評価である代替使用タスクに対するモデルの応答を調査しています。文埋め込みの類似性を使用して、これらのプロンプトに対する LLM 応答の傾向を調べます。私たちの調査結果は、時間の経過とともにモデル出力の多様性が統計的に有意に減少していることを示しており、LLM 出力がモデル全体で創造的な内容に収束している可能性があることを示唆しています。この傾向が続く場合、LLM による均質化により、人間と AI の共同創造作業における人間の主体性が徐々に低下する可能性があり、人間の創造プロセスにおける LLM の役割について慎重な検討が必要になります。

原文 (English)

Are LLMs becoming similarly creative? Evidence from three years of models

Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend persists, LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work, demanding careful consideration of LLMs' role in the human creative process.

13:00 JSTLLM/生成AI研究/論文

仕様が何を決定するかを測定する: 正式なセマンティックブロックモデルと実行判定ベンチマーク

この研究では、仕様の正式なセマンティック ブロック モデルと、モデルの機能とは独立して仕様の品質を評価するための実行判定ベンチマークが導入されています。仕様は、セマンティック ブロック、依存関係、ブロック所有のルール、意思決定ポイント、および明示的にオープンな質問で構成される構造として表され、非循環性、単一所有権、制約の支配、全体性または曖昧さの停止という 4 つの機械チェック可能な整形式条件に従います。決定性は、すべての適合実装間の一致としてモデル理論的に定義され、独立した実装者間の収束を通じて経験的に推定されます。このモデルは、18 個のブロックと 19 個の依存関係エッジを含む Oracle から PostgreSQL への移行仕様に基づいてインスタンス化されます。計算による検証の結果、5 層の分解は、依存関係のクロージャを通じてタスクごとの平均コンテキストを約 71% 削減し、特定されたすべてのギャップをトリアージして研究定義の Oracle 構成分類法の 85.5% をカバーし、テストされた代替パーティションによってパレート支配されておらず、元の構造の定義に使用されていない引用由来のエッジから 99.9 パーセンタイルで回復されることを示しています。このベンチマークは、実装者パネルを固定したままにし、必須の仕様なしの制御アームを組み込み、PostgreSQL 16 とライブ Oracle インスタンスを決定論的な実行判定として使用します。 3 つの事前登録操作と 3 つの診断分析を含む 6 つの計画された研究により、仕様の効果がさらに調査されます。 25 ユニットのサブサンプルで繰り返し実行すると、アーム デルタ スプレッドの中央値が 14.4 パーセント ポイントである経験的な変動フロアが明らかになります。この結果は、正式な概念としての決定性を裏付けていますが、評価された現代の LLM 実装者にとってのスタンドアロンの経験的な品質指標としては裏付けられていません。

原文 (English)

Measuring What a Specification Determines: A Formal Semantic-Block Model and an Execution-Judged Benchmark

This work introduces a formal semantic-block model for specifications and an execution-judged benchmark for evaluating specification quality independently of model capability. A specification is represented as a structure comprising semantic blocks, dependency relations, block-owned rules, decision points, and explicitly open questions, subject to four machine-checkable well-formedness conditions: acyclicity, single ownership, constraint domination, and totality or ambiguity-stop. Determinacy is defined model-theoretically as agreement among all conforming implementations and is estimated empirically through convergence across independent implementers. The model is instantiated on an Oracle-to-PostgreSQL migration specification containing 18 blocks and 19 dependency edges. Computational validation shows that the five-layer decomposition reduces mean per-task context by approximately 71% through dependency closures, covers 85.5% of the study-defined Oracle construct taxonomy with all identified gaps triaged, is not Pareto-dominated by the tested alternative partitions, and is recovered at the 99.9th percentile from citation-derived edges not used to define the original structure. The benchmark keeps the implementer panel fixed, includes a mandatory no-specification control arm, and uses PostgreSQL 16 and a live Oracle instance as deterministic execution judges. Six designed studies, including three pre-registered manipulations and three diagnostic analyses, further examine specification effects. Repeated runs on a 25-unit subsample reveal an empirical variability floor with a median arm-delta spread of 14.4 percentage points. The results support determinacy as a formal concept but not as a standalone empirical quality metric for the evaluated contemporary LLM implementers.

13:00 JSTエージェント研究/論文Claude

Agentic AI によるシミュレーションベースのスケジューリングのための高速遺伝的プログラミングのハイパーヒューリスティック

Python は、迅速な開発を可能にし、データ分析、人工知能 (AI)、機械学習のための豊富なエコシステムを提供するため、科学研究で広く使用されています。ただし、カスタマイズされた研究コードは、実験の規模が大きくなるにつれて、法外に遅くなる可能性があります。この課題は、離散イベントのプロジェクト スケジューリング シミュレーションで特に深刻です。このシミュレーションでは、逐次的な状態更新、ネストされたループ、条件付き評価、およびオブジェクト指向構造により、コンパイルされた数値ライブラリと GPU アクセラレーション ライブラリの利点が制限されます。これらのボトルネックに対処するには、通常、反復的なプロファイリング、リファクタリング、テスト、検証が必要ですが、研究者には低レベルの最適化のための時間やソフトウェア エンジニアリングの専門知識が不足している場合があります。このペーパーでは、ハイ パフォーマンス コンピューティング (HPC) 環境における現実世界のプロジェクト スケジューリング ワークロードに対して、Claude エージェント AI を使用した系統的なリファクタリング アプローチを紹介します。代表的なベンチマークと正確性チェックに基づいて、エージェントはボトルネックを特定し、対象を絞った最適化を実装し、その効果を評価しますが、研究者は最終的な制御を保持します。出力を変更せずにテストの実行時間が 1,298 秒から 200 秒未満に短縮され、年間 400 万コア時間 (NZ\320,000) が節約されました。

原文 (English)

Accelerated Genetic Programming Hyper-Heuristics for Simulation-Based Scheduling via Agentic AI

Python is widely used in scientific research because it enables rapid development and provides rich ecosystems for data analysis, artificial intelligence (AI), and machine learning. However, customized research code can become prohibitively slow as experiments scale. This challenge is particularly acute in discrete-event project-scheduling simulations, where sequential state updates, nested loops, conditional evaluations, and object-oriented structures limit the benefits of compiled numerical and GPU-accelerated libraries. Addressing these bottlenecks typically requires iterative profiling, refactoring, testing, and validation, yet researchers may lack the time or specialized software-engineering expertise for low-level optimization. This paper presents a systematic refactoring approach using Claude agentic AI on real-world project-scheduling workloads in a high-performance computing (HPC) environment. Guided by representative benchmarks and correctness checks, the agent identifies bottlenecks, implements targeted optimizations, and evaluates their effects, while the researcher retains final control. Testing runtime reduced from 1,298 seconds to under 200 seconds without changing outputs, saving four million core-hours (NZ\$320,000) annually.

13:00 JST研究/論文

生涯学習についての 2 つの考え: 半球の冗長性と神経モデルの専門化の探求

永続的なインテリジェント システムには継続的に学習する能力が必要ですが、現在の機械学習アプローチは、生物学的学習システムと比較して、この分野で大きな課題に直面しています。機械学習アルゴリズムは通常、以前に学習した情報の保持と、新しいデータ パターンまたは変化するデータ パターンへの適応とをトレードオフにします。継続的な学習機能がない場合、アルゴリズムはデータセット全体を使用して再トレーニングを受ける必要がありますが、ストレージの制約、財務コストまたは計算コスト、またはプライバシーの制限により元のトレーニング データが利用できない場合、このアプローチは非現実的になります。しかし、生物学上の動物は、壊滅的な忘却を経験することなく、継続的に学習することができます。この論文は、記憶の固定に関連することが知られている神経コンポーネントと状態をモデル化することにより、動物がどのように知識を学習し保存するのかについての高レベルのフレームワークを構築することを試みています。私たちは、体験の再生、レム睡眠、双方向性という 3 つの概念に焦点を当てています。私たちは、機械学習モデルがそれぞれ独自の長期および短期記憶メカニズムを持つ非対称半球からどのように恩恵を受けるか、および増分学習タスク間の睡眠期間が記憶の固定化にどのように役立つかを示す新しいマクロアーキテクチャである 4MAS (4 Module Awake/Sleep) を提案します。最後に、私たちのアーキテクチャが Split-MNIST、Split-Fashion-MNIST、Split-CIFAR-100 データセットでそれぞれ 98.3%、84.9%、29.29% の精度で競合する結果を達成していることを示す結果を示します。

原文 (English)

In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models

Persistent intelligent systems require the ability to learn continually, but current machine learning approaches face significant challenges in this area compared to biological learning systems. Machine learning algorithms typically trade off retention of previously learned information and adaptation to new or changing data patterns. When continual learning capabilities are absent, algorithms must undergo retraining using the entire data set, an approach that becomes impractical when original training data are unavailable due to storage constraints, financial or computational costs, or privacy restrictions. However, biological animals can learn continually, without experiencing catastrophic forgetting. This paper attempts to build a high-level framework for how animals learn and preserve knowledge by modelling neural components and states that are known to be related to memory consolidation. We focus on three concepts: experience replay, REM sleep, and bilaterality. We propose 4MAS (4 Module Awake/Sleep), a novel macroarchitecture demonstrating how machine learning models might benefit from asymmetric hemispheres, each with their own long- and short-term memory mechanisms, and how a period of sleep between incremental learning tasks might benefit memory consolidation. Finally, we present results showing that our architecture achieves competitive results on the Split-MNIST, Split-Fashion-MNIST and Split-CIFAR-100 datasets, with 98.3%, 84.9%, and 29.29% accuracy respectively.

13:00 JSTLLM/生成AIGPT / ChatGPT

大規模言語モデルと検索拡張生成を使用した金融ニュースの自動要約: 初期の実証研究 (2023 年秋)

株式市場のアナリストや投資家は、金融ニュースが多すぎて時間が足りないという課題に日々直面しています。何百もの企業固有の記事を手動で読んで総合することは現実的ではありませんが、重要な情報が欠けていると、投資の意思決定に直接影響する可能性があります。 2023 年秋にジョージ ワシントン大学で実施されたこのプロジェクトは、大規模言語モデルがこのプロセスを確実に自動化できるかどうかを調査します。私たちは、News API からニュース記事を、Wikipedia から企業背景を、Yahoo Finance から大手企業 10 社 (AAPL、MSFT、GOOGL、AMZN、META、TSLA、JPM、NVDA、WMT、DIS) の株価データを取得するパイプラインを構築しました。 LLM は数値表を直接処理できないため、株式データを自然言語の物語に変換するシンプルだが効果的なテンプレートを開発しました。次に、ニュースについては 3 つのオープンソース モデル (Falcon-7B-Instruct、DistilBART-CNN-12-6、BART-Large-XSum)、株価概要については GPT (text-davinci-003) にわたって 2 つの要約アプローチ (要約チェーンと FAISS による検索拡張生成) をテストしました。 Summarize Chains を備えた Falcon-7B は最高の結果をもたらし、すべてのニュース イベントを正確かつ一貫してカバーしました。 RAG は理論的には有望ですが、k が大きい場合、Falcon で激しい繰り返しが発生し、BART-Large で幻覚が発生しました。どちらの LLM ベースのアプローチも、ROUGE-1 での単純な Lead-3 ベースラインを上回りました。また、インタラクティブな株式視覚化のための Streamlit ダッシュボードも構築しました。この研究は、RAG ベースの金融ツールが普及する前の 2023 年秋に行われたもので、私たちが文書化した障害モード、特に小規模モデルにおける RAG 下の幻覚は、現在でも意味を持ち続けています。

原文 (English)

Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)

Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly affect investment decisions. This project, conducted at George Washington University in Fall 2023, explores whether Large Language Models can automate this process reliably. We built a pipeline that pulls news articles from the News API, company background from Wikipedia, and stock price data from Yahoo Finance for ten major companies (AAPL, MSFT, GOOGL, AMZN, META, TSLA, JPM, NVDA, WMT, DIS). Because LLMs cannot directly process numerical tables, we developed a simple but effective template that converts stock data into natural language narratives. We then tested two summarization approaches (Summarize Chains and Retrieval-Augmented Generation with FAISS) across three open-source models (Falcon-7B-Instruct, DistilBART-CNN-12-6, BART-Large-XSum) for news, and GPT (text-davinci-003) for stock summaries. Falcon-7B with Summarize Chains gave the best results, covering all news events accurately and coherently. RAG, while promising in theory, caused severe repetition in Falcon and hallucinated facts in BART-Large when k was large. Both LLM-based approaches outperformed a simple Lead-3 baseline on ROUGE-1. We also built a Streamlit dashboard for interactive stock visualization. The work was done in Fall 2023, before RAG-based financial tools became widespread, and the failure modes we document, particularly hallucination under RAG in smaller models, remain relevant today.

13:00 JSTLLM/生成AI

マシンが話すとき: マシンネイティブのシンボルを事前トレーニングされた大規模言語モデルに統合するための統合生成フレームワーク

現実世界の AI システムの多くは、自然言語ではなく個別のマシンネイティブのシンボルを使用して、エンティティ、動作、構造化情報を表します。これらの表現はコンパクトでタスクに関連した構造を保持しますが、事前トレーニングされた大規模言語モデル (LLM) の言語トークン空間の外側に位置し、言語モデリングと構造化予測の間に根本的な溝を生み出します。 UniLang は、事前トレーニング済み LLM を拡張して、自然言語トークンと並んでマシンネイティブのシンボルを第一級の生成単位として扱うことで、この溝を埋める統合生成フレームワークです。 UniLang は、LLM の語彙と埋め込み空間を根拠のあるマシンネイティブ表現で拡張し、単一の自己回帰目標の下でテキストと記号のトークンを共同でモデル化し、生成できるようにします。この統合インターフェイスにより、事前トレーニング済み LLM は、自然言語として言語化したり、タスク固有のアーキテクチャに依存したりすることなく、マシンネイティブな表現を直接操作できます。私たちは、異なるドメインとタイプの構造化予測にまたがる、逐次推奨と判例予測という 2 つの構造的に異なるタスクに基づいて UniLang を評価します。どちらのタスクでも、UniLang は一貫して強力なベースラインを上回り、事前トレーニング済み LLM を言語を超えて拡張し、異種マシンネイティブ表現の共通生成モデリング バックボーンとして使用する道を示しています。

原文 (English)

When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models

Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant structure, they lie outside the linguistic token space of pretrained large language models (LLMs), creating a fundamental divide between language modeling and structured prediction. We introduce UniLang, a unified generative framework that bridges this divide by extending pretrained LLMs to treat machine-native symbols as first-class generative units alongside natural-language tokens. UniLang expands the LLM's vocabulary and embedding space with grounded machine-native representations, enabling textual and symbolic tokens to be jointly modeled and generated under a single autoregressive objective. This unified interface allows pretrained LLMs to directly operate on machine-native representations without requiring them to be verbalized as natural language or relying on task-specific architectures. We evaluate UniLang on two structurally distinct tasks, sequential recommendation and legal precedent prediction, spanning different domains and types of structured prediction. Across both tasks, UniLang consistently outperforms strong baselines, demonstrating a path toward extending pretrained LLMs beyond language and using them as a common generative modeling backbone for heterogeneous machine-native representations.

13:00 JST画像/動画生成ロボティクス

CVSD-Reg: 堅牢な LiDAR 登録のためのクロスモーダルビジュアルセマンティック事前蒸留

学習ベースのグローバル点群登録は目覚ましい進歩を遂げていますが、幾何学的表現に依存しているため、既存の方法は点密度、スキャン パターン、視点、センサー特性の変化に敏感になっています。我々は、視覚基盤モデルから視覚的意味論的事前情報を LiDAR 表現に抽出する、堅牢なグローバル LiDAR 登録フレームワークである CVSD-Reg を提案します。ステージ 1 では、Point Transformer V3 の学生が、対比蒸留と教師の埋め込み空間の超球面ジオメトリを保存する球面多様体アラインメントを通じて、凍結した DINOv2 教師から学習します。自己監視型 InfoNCE の一貫性とソフト $\mathrm{SE}(3)$ 不変性は、視点に強い記述子をさらに促進します。ステージ 2 では、抽出された表現が、対応学習、密度を考慮したポイントドロップアウト拡張、およびエンドツーエンドの姿勢最適化を通じて位置合わせに適応されます。単一のチェックポイントを使用する CVSD-Reg は、センサー固有の適応を必要とせずに、単一センサーとゼロショット クロスセンサーの両方のシナリオに一般化され、推論時に完全にカメラフリーのままになります。 KITTI、nuScenes、および HeLiPR では、CVSD-Reg は、それぞれ 97.7$\%$、99.0$\%$、および 99.3$\%$ の厳密な成功率 (SR@0.5\,m/$1^\circ$) を達成します。これには、まばらな 16 ビーム Velodyne スキャンでの 97.3$\%$ が含まれます。カメラ入力やポストホック ICP 調整を必要とせずに、最先端の幾何学的位置合わせ方法よりも最大 44.0 パーセント優れた性能を発揮します。

原文 (English)

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations. In Stage 1, a Point Transformer V3 student learns from a frozen DINOv2 teacher through contrastive distillation and spherical-manifold alignment, which preserves the hyperspherical geometry of the teacher embedding space. Self-supervised InfoNCE consistency and soft $\mathrm{SE}(3)$ invariance further encourage viewpoint-robust descriptors. In Stage 2, the distilled representation is adapted to registration through correspondence learning, density-aware point-dropout augmentation, and end-to-end pose optimization. With a single checkpoint, CVSD-Reg generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference. On KITTI, nuScenes, and HeLiPR, CVSD-Reg achieves strict success rate (SR@0.5\,m/$1^\circ$) of 97.7$\%$, 99.0$\%$, and 99.3$\%$, respectively, including 97.3$\%$ on sparse 16-beam Velodyne scans. It outperforms state-of-the-art geometric registration methods by up to 44.0 percentage points without requiring camera inputs or post-hoc ICP refinement.

13:00 JST画像/動画生成

Stream4D: ストリーミング自己回帰拡散ビデオ モデルの 4D 一貫性

ストリーミング自己回帰拡散モデルにより、リアルタイムの長期ビデオ生成が可能になりますが、そのトレーニング目標は、コヒーレントな世界のジオメトリやダイナミクスではなく、ローカル フレーム予測を最適化することです。長いロールアウトでは、ジオメトリ ドリフトが蓄積され、静的または不自然な動きに劣化します。最近の双方向アプローチは、3D ガウス スプラッティング再構築に基づいて構築された報酬信号を使用して、この問題に対処します。ただし、単一の剛体 3D 再構成では動的なシーンをモデル化できないため、この批評家は本物のオブジェクトの動きを再構成エラーとして罰し、ビデオをフリーズすることで最大化します。このショートカットは、各チャンクがすでに静的な構成を伝播する可能性がある AR 設定では特に有害です。この研究では、静的な批評家をシーンのダイナミクスを明示的にモデル化するフィードフォワード 4D 再構成報酬に置き換える Stream4D を提案します。これにより、コヒーレントなモーションが高い一貫性の報酬を受け取ることができます。モーションの大きさと品質をさらにガイドするために、ジッターと非剛体アーティファクトを減点しながら、自然なシーンフローの大きさに報酬を与えるモーション プリアを追加します。私たちの最終レシピは、これら 2 つの用語と軽量の知覚アンカーを組み合わせたものです。 Stream4D は、さまざまな自己回帰ビデオ バックボーンとさまざまな世代範囲にわたって、4D 再構築の品質を向上させ、動きをより効果的に保存し、人間とのより高度な調整を実現します。プロジェクトページ:https://banyuanhao.github.io/Stream4D/

原文 (English)

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic penalizes genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already-static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. Across various autoregressive video backbones and various generation horizons, Stream4D improves 4D reconstruction quality, preserves motion more effectively, and achieves higher human-aligned preference. Project page: https://banyuanhao.github.io/Stream4D/

13:00 JST研究/論文

DraftFM: マジック: ザ・ギャザリングにおけるデイゼロドラフト用の FoundationModel

新しいマジック:ザ・ギャザリング拡張のドラフトは、そこからのピックが観察される前に始まります。完全なカードリストは公開されていますが、監督されたピックモデルのトレーニングに使用されたドラフトログはまだ存在しません。私たちはこのデイゼロ体制を直接研究しています。 DraftFM は、ドラフト対象プールとドラフトの状態を条件として、現在のパックで利用可能なカードを正確にスコアリングする個別選択ポリシーです。すべてのカードは、公開カード レコード、構造化された特徴、および固定テキストの埋め込みの凍結された 775 次元関数として入力され、モデル内のどこにもカード ID、セット ID、または使用統計が存在しないため、目に見えないカードはよく知られたカードと同じ機械によってスコア付けされます。 29の拡張からの1億4,900万の人間のピックに適合した160万のパラメータネットワークは、3つの拡張の全体でホールドアウトされたピックを予測し、トップ1一致率は50.8%、60.4%、56.7%に達し、開幕ピックの均一確率は約7%です。観測された 32 の拡張すべてに再装備され、同じアーキテクチャにより、当時未リリースだったセット「ホビットの冒険」のカード ランキングが作成されました。このセットは完全な暗号の出所とともに封印され、MTG アリーナでセットがドラフト可能になる約 36 時間前に公開されました。封印されたランキングは、6 人の独立した専門家レビューアーの意見とほぼ同じくらい同意します。実現された結果に対する評価は、それが示すものを問わず、後続のメモに反映されます。

原文 (English)

DraftFM: A FoundationModel for Day-Zero Drafting in Magic: The Gathering

Drafting a new Magic: The Gathering expansion begins before any pick from it has been observed: the complete card list is public, but the draft logs that supervised pick models train on do not yet exist. We study this day-zero regime directly. DraftFM is a discrete-choice policy that scores exactly the cards available in the current pack, conditioned on the drafted pool and the state of the draft. Every card enters as a frozen 775-dimensional function of its public card record, structured features and a fixed text embedding, with no card identities, set identities, or usage statistics anywhere in the model, so an unseen card is scored by the same machinery as a familiar one. A 1.6-million-parameter network fitted on 149 million human picks from 29 expansions predicts held-out picks in three expansions withheld in their entirety, reaching 50.8%, 60.4%, and 56.7% top-1 agreement, where uniform chance at the opening pick is about 7%. Refitted on all 32 observed expansions, the same architecture produced a card ranking for the then-unreleased set The Hobbit, sealed with its complete cryptographic provenance and published roughly 36 hours before the set became draftable on MTG Arena. The sealed ranking agrees with six independent expert reviewers roughly as much as those reviewers agree with one another. Evaluation against realized outcomes is committed to a follow-on note, whatever it shows.

13:00 JST画像/動画生成

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet r…

13:00 JSTLLM/生成AI画像/動画生成

PEA-DPO: MLLM アライメントのための知覚強化アライメント直接優先最適化

Direct Preference Optimization (DPO) は、大規模言語モデル (LLM) を人間の好みに合わせるための効果的なアプローチとして登場しました。ただし、マルチモーダル設定への適応はまだ解明されていません。私たちは、表現分析を通じて、マルチモーダル選好の最適化における重要な制限を特定しました。これを視覚的鈍感性と呼んでいます。モデルは、多くの場合、画像と、重要な視覚的コンテキストが削除された画像とを区別できません。私たちの理論的分析により、この問題の 2 つの症状、つまり画像全体の感度の低下と画像内の感度の低下がさらに明らかになりました。これらの課題に対処するために、私たちは、マルチモーダル LLM アライメントのフレームワークである Perception-Enhanced Alignment DPO (PEA-DPO) を提案します。これは、視覚的嗜好信号を明示的に活用して、視覚的不感症を克服します。さらに、PEA-DPO が両方の故障モードを軽減できることを証明する理論的分析を提供します。経験的な結果は、PEA-DPO が基本モデルの言語モデリング能力を維持しながら、視覚的なコンテキストに対する感度を高めることを示しています。さまざまなスケールの MLLM を使用した 3 つの幻覚ベンチマークにわたる評価では、PEA-DPO が視覚不感受性を効果的に軽減し、より強力なマルチモーダル アライメントを実現し、幻覚を大幅に軽減することが示されています。

原文 (English)

PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal settings remains unexplored. Through representational analysis, we identify a key limitation in multimodal preference optimization, which we term visual insensitivity: models often fail to distinguish between images and those with critical visual context removed. Our theoretical analysis further uncovers two manifestations of this problem, namely Across-Image Insensitivity and Within-Image Insensitivity. To address these challenges, we propose Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity. We further provide a theoretical analysis demonstrating that PEA-DPO provably mitigates both failure modes. Empirical results demonstrate that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base model. Evaluations across three hallucination benchmarks using MLLMs of varying scales show that PEA-DPO effectively mitigates visual insensitivity, achieves stronger multimodal alignment, and substantially reduces hallucinations.

13:00 JSTLLM/生成AI

高速フォーク: テキスト生成における不確実性ダイナミクスを効率的に推定する

LLM 推論は確率的であるため、モデルを理解するには、特定の質問に対して生成される可能性のある推論チェーンの分布、つまりその不確実性に取り組む必要があります。リサンプリングベースの分析はこの分布を特徴づけ、モデルがその答えにどのように到達するかを決定するロールアウトのステップを明らかにします。ただし、これらのアプローチの主な制限は、推論チェーン内のすべてのトークンまたは文でテキスト シーケンスをリサンプリングすると、非常にコストがかかることです。私たちの研究は、リサンプリング分析の計算効率を高めると同時に、テキスト生成における不確実性のダイナミクスを説明するための適切な統計モデルは何かという重要な科学的疑問にも光を当てることを目指しています。多くの推論チェーンをリサンプリングすると、不確実性のダイナミクスは安定したパターンに収束し、ノイズは個々のトークンや推論ステップに対する LLM の感度ではなく、主にサンプリングのアーティファクトであることを示します。ノイズの多い低サンプルのロールアウト データを平滑化して高サンプル データをより適切に近似するための統計モデルを開発し、サンプリング コストを大幅に削減できます。

原文 (English)

Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation

LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. However, a major limitation of these approaches is that resampling text sequences at every token or sentence in a reasoning chain is very costly. Our work strives to make resampling analysis more computationally efficient, while also shedding light on an important scientific question: what is the right statistical model for explaining uncertainty dynamics in text generation? We show that when resampling many reasoning chains, uncertainty dynamics converge to stable patterns, and noise is largely an artifact of sampling rather than an LLM's sensitivity to each individual token or reasoning step. We develop a statistical model for smoothing noisy low-sample rollout data to better approximate high-sample data, allowing us to significantly cut sampling costs.

13:00 JSTエージェント研究/論文ClaudeGPT / ChatGPT

DeltaML-Bench: 現実世界のリサーチ リポジトリでの機械学習エージェントの評価

機械学習実験用の自律エージェントは、現実的なコンピューティング制約の下で、異種のリポジトリをナビゲートし、トレーニング パイプラインを修復し、改善候補を評価する必要があります。既存のベンチマークは、これらの状況を部分的にしか捉えていません。 DeltaML-Bench は、不完全なオープンソース リポジトリ内で公開されているベースラインを改善することをエージェントに要求する、研究論文をソースとする 48 のタスクで構成されるベンチマークです。標準のモジュラー エージェントと検索ベースの ARG スキャフォールディングを使用して GPT-5 と Claude Sonnet 4 を評価します。 4 x 6 時間の割り当てでは、ARG は GPT-5 の実行あたりの成功率を 9.4% から 33.9% に引き上げます。 2 x 12 時間の割り当てでは、GPT-5 ARG は 49.0% に達します。モジュール構成では 47.9% もの高い仕様ゲーム率が示されていますが、評価された ARG 構成ではゲームは観察されません。これらの結果は、自律型 ML 実験用にエージェントをデプロイする際には、スキャフォールディングの設計と整合性チェックが重要な考慮事項であることを示しています。

原文 (English)

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic compute constraints. Existing benchmarks only partially capture these conditions. We introduce DeltaML-Bench, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories. We evaluate GPT-5 and Claude Sonnet 4 with a standard Modular agent and a search-based ARG scaffolding. In the 4 x 6h allocation, ARG raises GPT-5's per-run success rate from 9.4% to 33.9%; in the 2 x 12h allocation, GPT-5 ARG reaches 49.0%. Modular configurations exhibit specification gaming rates as high as 47.9%, while no gaming is observed in the evaluated ARG configurations. These results indicate that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experimentation.

13:00 JST研究/論文

流砂からの脱出: 武器への呼びかけ

コンピューティングは驚くべき成功を収めてきましたが、蓄積された技術的負債により、私たちはビジネスや社会的リスクにおける莫大なコストにさらされています。 75 年間、当社はテストとデバッグの開発により、仕様に沿ったシステムを構築してきました。これは業界が繁栄するためには十分に機能しますが、費用がかかり効果のないフィードバック ループであり、誰もが不安定な基盤に依存することになります。現在、AI 対応エンジニアリングは、コーディング コストを削減することで成功を拡大していますが、技術的負債の急速な増加とその脆弱性の検出の自動化により、リスクも増幅しています。どうすればもっと良くなるでしょうか?研究は長い間、正しさの数学的証明を追求してきましたが、テストとは異なり、すべてのケースをカバーできます。これも大幅に進歩しましたが、技術的にも、根深い文化的な断絶のためにも、適用するのは依然として困難です。その代わりに、私たちは、AI と人間の開発の両方により効果的なフィードバック ループを提供する、テスト、*仕様*、証明を柔軟に組み合わせる実用的なアプローチを主張します。最も簡単には、従来の散文記述、コード、テストと並行して、テストオラクルとして実行可能な部分仕様を段階的に共同開発できます。これにより、設計が明確になり、テストがより識別できるようになります。開発者は今すぐにそれを行うことができますし、そうすべきです。あるいは、さらに良いことに、テスト、プロパティベースのテスト、シンボリック実行、および証明の全範囲をサポートする仕様を使用できます。これにより、安価なテストから高価な証明に至るまで、AI と人間の両方に対して、さまざまな絡み合ったフィードバック ループが可能になります。しかし、それを実際に実用化するには、*セマンティクス インフラストラクチャ* が必要です。つまり、主要なプログラミング言語およびその他の抽象化のための仕様とツールです。これらの構築方法は現在では多かれ少なかれわかっていますが、まだ整備されていません。私たちは、より強固な基盤の上に築かれる未来を可能にするために、それを作成し展開するためにコミュニティに武装するよう呼びかけます。

原文 (English)

Escaping the Quicksand: A Call to Arms

Computing has been an astonishing success - but the accumulated technical debt exposes us all to huge costs in business and societal risk. For 75 years, we've built systems to prose specifications with test-and-debug development. That works well enough for industry to thrive, but it's an expensive and ineffective feedback loop, and leaves everyone relying on shaky foundations. Now, AI-enabled engineering is amplifying the success by reducing coding costs, but also amplifies the risks, by rapidly increasing technical debt, and by automating detection of the vulnerabilities therein. How can we do better? Research has long pursued mathematical proof of correctness, which, unlike testing, can cover all cases. This too has advanced massively, but it remains hard to apply, both technically and because of a deep-seated cultural disconnect. Instead, we argue for a pragmatic approach to flexible combinations of testing, *specification*, and proof, that provides more effective feedback loops for both AI and human development. Most simply, one can incrementally co-develop executable-as-test-oracle partial specifications alongside conventional prose descriptions, code, and tests. This clarifies design and makes testing much more discriminating. Developers can and should do it today. Or, even better, one can use specifications that support the full gamut of testing, property-based testing, symbolic execution, and proof. This enables a range of intertwined feedback loops, again both for AI and humans, from cheap testing to more expensive proof. However, making it really practical needs *semantics infrastructure*: specifications and tooling for the main programming languages and other abstractions, which we now more-or-less know how to build, but which is not yet in place. We call the community to arms to create and deploy it - to enable a future built on firmer ground.

13:00 JSTエージェント

Loreley: Repository-Scale Program Evolution with Quality-Diversity Search

Sequential agent search accumulates changes from its current champion but discards alternative branches; independent proposals preserve bre…

13:00 JST画像/動画生成ロボティクス

Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions

Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation…

13:00 JST画像/動画生成

Scale-Separated Conditioning for Style-Encoder-Free Diffusion Stylization

Reference-based diffusion stylization requires separating target geometry from transferable appearance. Existing tuning-based methods often…

13:00 JST研究/論文

A Locally Tokenized Generative Model for Robust Time-Series Watermarking

Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by rel…

13:00 JSTLLM/生成AI画像/動画生成GPT / ChatGPTGemini

TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling

Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on t…

13:00 JST画像/動画生成

Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis

Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventric…

13:00 JSTLLM/生成AI画像/動画生成ビジネス/資金調達Claude

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still…

13:00 JSTLLM/生成AI

Truncate Bad, Upweight Good: BoN-Style Distillation via Rank-Based Classification

Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top-ranked compl…

13:00 JSTロボティクス

GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects bec…

13:00 JSTLLM/生成AIエージェントQwen

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signa…

13:00 JSTハードウェア/半導体

Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW

A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimize…

13:00 JSTロボティクス

CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning

Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it…

13:00 JSTエージェント

Distilling Aggregate Mobility Statistics into a Language Model Policy for Post-Event Crowd Simulation

Pedestrian simulators need a behaviour rule for every agent, but privacy usually limits the data for setting one to aggregate statistics, n…

13:00 JSTエージェント

An Irreducible Quantum Advantage in Aligning World Models with Reality

World models provide digital simulacra of the true world, allowing agents to be trained and tested before costly real-world deployment. At…

13:00 JSTLLM/生成AI

LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment

Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhe…

13:00 JSTLLM/生成AIエージェント

MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents

Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Exi…

13:00 JST画像/動画生成

Core-KAN: Continuous Vision Kernels with Kolmogorov-Arnold Networks

Conventional convolutional kernels are typically defined on fixed discrete grids, limiting their ability to accommodate heterogeneous local…

13:00 JST研究/論文

Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning

Probabilistic shielding is a technique for safe reinforcement learning (RL). Typically, a static observer -- called the shield -- constrain…

13:00 JSTエージェントGPT / ChatGPTDeepSeek

Repo0: Design-Driven Zero-to-All Code Generation

Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository arc…

13:00 JST研究/論文

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly…

13:00 JSTLLM/生成AIエージェント

A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries

Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine a…

13:00 JST研究/論文

Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition

Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn…

13:00 JST研究/論文

Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn…

13:00 JSTLLM/生成AIビジネス/資金調達

Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models

Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We d…

13:00 JSTLLM/生成AIエージェント研究/論文

MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration. This…

13:00 JST研究/論文

Towards Quantifying Benchmark Optimization in ASR Models

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, ther…

13:00 JST研究/論文

Designing Human-mediated AI Guidance: Ready Together for Personalized Family Emergency Preparedness

Artificial intelligence (AI) systems are increasingly used across domains to provide personalized information, recommendations, and decisio…

13:00 JST画像/動画生成

Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training

Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scene…

13:00 JST画像/動画生成

An Inclusive and Lightweight Approach to Federated Continual Learning for Cultural Heritage

Artificial intelligence can support cultural heritage and digital humanities through large-scale retrieval and analysis of digitized collec…

13:00 JST研究/論文Gemini

EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models

Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets.…

13:00 JSTLLM/生成AI

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However,…

13:00 JSTLLM/生成AI

SABET-QA: Temporal Knowledge Graph Question Answering

Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based method…

13:00 JSTロボティクス

Evidence-Gated Task and Motion Planning with Vision-Language Models

Robots executing long-horizon manipulation tasks from natural-language instructions must reason about both semantic task structure and geom…

13:00 JSTロボティクス

Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking

Humanoid robots have recently demonstrated promising capabilities in real-world ball sports. However, achieving professional motion styles…

13:00 JST画像/動画生成

Structured Affinity for Unsupervised Visual Class-Incremental Memory in Deep Artificial Immune Networks

Artificial immune networks (AINs) are naturally memory-forming systems, but conventional visual AINs often rely on flattened vector affinit…

13:00 JSTLLM/生成AIエージェント

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization it…

13:00 JST研究/論文

A Standardized Framework for Machine Learning in Power System Protection

Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends…

13:00 JSTハードウェア/半導体

Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational Data

Practitioners inferring causality from observational data usually rely on a single method and treat its output as causal truth. Recent tool…

13:00 JSTエージェント

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding a…

13:00 JSTLLM/生成AIGPT / ChatGPT

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target f…

13:00 JSTLLM/生成AI画像/動画生成

Prompt-Conditioned Channel Attention for Hierarchical Feature Modulation toward Anatomy-Agnostic Segmentation

Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts. Int…

13:00 JSTハードウェア/半導体

Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI

The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses an accounting role wit…

13:00 JSTLLM/生成AILlamaQwen

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inf…

13:00 JSTLLM/生成AI

Inducing Task Models from Computer-Use Traces

Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbol…

13:00 JSTLLM/生成AI画像/動画生成

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires…

13:00 JSTエージェント

Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists

We aim at designing language agents with greater autonomy for crystal materials discovery. While most of existing studies restrict the agen…

13:00 JSTエージェント

GridCodex: A RAG-Driven AI Framework for Power Grid Code Reasoning and Compliance

The global shift towards renewable energy presents unprecedented challenges for the electricity industry, making regulatory reasoning and c…

13:00 JSTLLM/生成AIビジネス/資金調達AnthropicClaudeOpenAIGPT / ChatGPTGoogleGemini

Computational Phenomenology of Borderline Personality Disorder: A Comparative Evaluation of LLM-Simulated Expert Personas and Human Clinical Experts

Building on a human-led thematic analysis of clinical life-story interviews (> 150,000 words) with inpatients with Borderline Personality D…

13:00 JST研究/論文

Gen AI in Proof-based Math Courses: A Pilot Study

With the rapid rise of generative AI in higher education, understanding how students use AI is increasingly important. This exploratory stu…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step int…

13:00 JST研究/論文

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding

Charts are ubiquitous in scientific and financial literature for presenting structured data. However, chart reasoning remains challenging f…

13:00 JSTエージェント

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual perception with text generation. Exist…

13:00 JSTエージェント研究/論文

Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems

As AI agents transition from research prototypes to enterprise production systems, the tool interfaces they consume remain rooted in human-…

13:00 JSTエージェント

The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning

As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context,…

13:00 JSTエージェント

ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction

Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit phy…

13:00 JST研究/論文

正確さから監査可能性へ: 金融 AI システムにおける決定論の調査

信用リスク、不正行為検出、マネーロンダリング対策といった規制された金融環境に機械学習を導入すると、アルゴリズムの再現性における重大な脆弱性が露呈します。初期の金融 ML はバックテストのオーバーフィッティングなどの統計的課題に対処しましたが、ディープ ニューラル ネットワークと生成 AI では、ハードウェアとアーキテクチャに根ざした機械的非決定性が導入されました。この調査では、表形式モデル (事後説明の分散)、グラフ ネットワーク (確率的サンプリングと時間的非同期性)、LLM ベースのエージェント ワークフロー (バッチ依存の発散と軌道ドリフト) という、金融 AI で現在主流となっている 3 つの手法にわたる再現性の障害に関するシステムの視点を提供します。公的金融データセットに関するファーストパーティの実験で文献分析を補足します。信用スコアリングにおける説明ランクの不安定性、GNN ベースの不正検出における予測フリップ レート、LLM エンティティ抽出におけるテンソル並列誘発出力の発散を定量化します。我々は、モダリティ固有の指標(RBO、D_cos、TDI、PSD)を監査の準備状況にリンクする階層化された評価フレームワークを提案し、ロジットレベルとセマンティックレベルの決定性尺度の相補性を経験的に検証します。

原文 (English)

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems

Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility. While early financial ML addressed statistical challenges such as backtest overfitting, deep neural networks and Generative AI have introduced mechanical nondeterminism rooted in hardware and architecture. This survey provides a systems perspective on reproducibility failures across three modalities now dominant in financial AI: tabular models (post-hoc explanation variance), graph networks (stochastic sampling and temporal asynchrony), and LLM-based agentic workflows (batch-dependent divergence and trajectory drift). We supplement the literature analysis with first-party experiments on public financial datasets -- quantifying explanation rank instability in credit scoring, prediction flip rates in GNN-based fraud detection, and tensor-parallel-induced output divergence in LLM entity extraction. We propose a layered evaluation framework linking modality-specific metrics (RBO, D_cos, TDI, PSD) to audit readiness, and empirically validate the complementarity of logit-level and semantic-level determinism measures.

13:00 JST研究/論文

CADRE: Stable, Parameter Efficient Adaptation of Medical Vision Language Models with Bounded Forgetting and Prior Drift

Medical vision-language models (VLMs) such as BiomedCLIP generalize broadly, but adapting them to a clinical service is as much a safety pr…

13:00 JSTエージェントGPT / ChatGPT

FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episod…

13:00 JSTLLM/生成AI

LLM Capability Limits: Static Emergence and Dynamic Boundary Control

Test-time emergence in LLM systems has a deployment boundary: additional computation can realize decisions already supported by the deploye…

13:00 JSTエージェントビジネス/資金調達研究/論文

大規模言語モデルでの投資ロジックの評価: パーソナライズされた金融エージェントに向けた現実世界のベンチマーク

投資能力は本質的に個別化されています。同じ市場証拠が、異なる目標、視野、ポートフォリオ、リスク境界を持つ投資家にとって異なる行動を正当化する可能性があります。しかし、財務 LLM は、静的な質問応答または最終的な損益によって評価されます。前者は主体性を省略します。後者では、利益をもたらす行動が根拠に基づいたものなのか、プロファイルに一貫性のあるものなのか、あるいは単に幸運だったのかを明らかにすることはできません。私たちは、コミュニティが結果的な要因に対して間違った物差しを使用しているかどうかを尋ねます。 \textsc{InvestLogicBench} は、151 人の現実世界の投資家からの 201,247 件の文書化された意思決定を含むプロセスネイティブのベンチマークです。各エピソードは \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} トレースをインスタンス化します: 投資家 \textit{プロフィール}、観察可能な市場 \textit{イベント}、投資 \textit{推論}、実行可能ファイル \textit{決定}、および遅延 \textit{結果}。このリリースには、プロファイルの構築、ポイントインタイムのイベント バインディング、構造化ロジック、ホライズン、結果、事後分析が含まれており、理解、プロファイル条件付き生成、エンドツーエンドの再生がサポートされています。 4 つの主要な LLM では、論理的妥当性は 4/5 近くを維持していますが、イベントグラウンディングはわずか 0.8 ~ 2.8/5 です。返品とプロセスの品質も一致しません。これらの結果は、結果のみの評価が隠れている、洗練されているものの根拠が弱い推論を明らかにします。さらに、P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O は、バージョン付きプロファイル、時間的来歴、検査可能な検索、意思決定台帳、および再生可能な結果を​​必要とするデータ システム インターフェイスであるべきだと主張します。財務は、より広範なクラスの個別化された結果的なエージェントに対するストレス テストです。

原文 (English)

Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} trace: investor \textit{Profile}, observable market \textit{Events}, investment \textit{Reasoning}, executable \textit{Decision}, and delayed \textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.

13:00 JST画像/動画生成ロボティクス

WNM-3D: 閉ループ VLN 用の 3D シーンコンディショニングを備えたワールド ナビゲーション モデル

最近のビジョン言語ナビゲーション (VLN) システムでは、事前トレーニング済みビジョン言語モデル (VLM) を、自己中心的な観察と言語指示をナビゲーション アクションに直接マッピングするビジョン言語アクション (VLA) ポリシーにますます適応させています。このようなアクション中心のトレーニングは、意味論的には可能ですが、エージェントの視覚観察がその予測された動作の下でどのように展開するかを明示的にモデル化するものではありません。生成世界アクション モデル (WAM) は、将来の観測とアクションを共同で予測しますが、連続 VLN の既存の WAM は、観察された履歴から推測される幾何学を意識した表現に基づいて共同の将来展望とアクションの生成を条件付けません。我々は、連続 VLN 用の 3D シーンコンディショニングを備えた生成ワールド ナビゲーション モデルである WNM-3D を紹介します。過去の観察を永続的なシーン コンテキストに統合するために、フリーズ フィードフォワード ジオメトリ エンコーダーが単眼の自己中心的な RGB 履歴からジオメトリ認識表現を抽出し、トレーニング可能な 3D シーンからトークンへのアダプターがそれらをワールド アクション拡散トランスフォーマーのトークン空間内の固定長プレフィックスに変換します。ブロック因果的注意を通じて、このプレフィックスは将来のすべてのビデオ アクション ブロックを条件付けし、将来のビューとアクション生成の両方に共有の幾何学的コンテキストを提供します。私たちは、A* によって生成されたデモンストレーションに対する監視付きワールドアクション微調整、政策訪問国に対する DAgger スタイルの適応、および DanceGRPO ベースの閉ループ政策最適化を通じて、WNM-3D をトレーニングします。 GN-Bench での実験では、WNM-3D が閉ループ ナビゲーションにおいて強力な VLM ベースのナビゲーション ポリシーや 2D 条件付きの対応物よりも優れていることが示されています。 WNM-3D は、固定されたゴールに近い評価セットで、より高いフローアクションの一貫性とより低い視覚動作エラーも実現します。

原文 (English)

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations should evolve under its predicted motion. Generative world-action models (WAMs) jointly predict future observations and actions, yet existing WAMs for continuous VLN do not condition joint future-view and action generation on geometry-aware representations inferred from the observed history. We present WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN. To consolidate past observations into persistent scene context, a frozen feed-forward geometry encoder extracts geometry-aware representations from the monocular egocentric RGB history, and a trainable 3D Scene-to-Token Adapter converts them into a fixed-length prefix in the token space of the world-action Diffusion Transformer. Through block-causal attention, this prefix conditions every future video-action block, providing a shared geometric context for both future-view and action generation. We train WNM-3D through supervised world-action fine-tuning on A*-generated demonstrations, DAgger-style adaptation on policy-visited states, and Counterfactual DanceGRPO refinement for closed-loop execution. Experiments on GN-Bench show that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation. Stage-wise ablations further show that DAgger-SFT provides the larger success-rate gain, while Counterfactual DanceGRPO subsequently improves both navigation success and path efficiency.

13:00 JST研究/論文

GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis

Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains su…

13:00 JSTエージェント研究/論文

Andy: 厳密な証明と自律的な研究のための数学エージェント

アンディは、提出された問題を解決して検証し、新しい研究問題を定式化し、厳密な証明を構築する自律的な数学研究エージェントです。証明の生成と正しさの評価を分離し、知識の取得、目標を絞った改訂、多段階の検証をサポートします。このペーパーでは、自己トリガーによる衝動的なコンセンサスに関する公開された結果を開始点として使用して、ワークフローを説明します。 Andy は、スイッチング通信トポロジを使用した遅延異種ネットワーク向けのグローバルな指数関数的なリーダー/フォロワー同期問題を定式化します。提案されたハイブリッド制御は、自己トリガー型インパルスと実行遅延および回復フェーズの連続フィードバックを組み合わせます。各遅延インパルスの後、このフィードバックによって回復ウィンドウ中に遅延エラー チャネルがキャンセルされます。グローバルな指数同期のための十分な条件が確立され、サンプリング シーケンスとインパルス シーケンスの両方で Zeno 動作が除外されます。数値例で結果を確認します。この事例は、既存の結果から学び、意味のある研究課題を定式化し、厳密な証明を開発して検証するアンディの能力を実証しています。

原文 (English)

Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research

Andy is an autonomous mathematical research agent that turns a mathematical problem into a traceable proof. It solves or verifies a submitted problem, formulates a literature-grounded new problem through a research-value gate, and carries it through proof construction and final verification. It organizes proof steps in an executable DAG, verifies each step independently and binds the result to a certificate, retains verified work whose interfaces remain unchanged during local repair, and records the full path from problem formulation to final proof. The system separates proof generation from correctness evaluation and can acquire, retain, retrieve, and reuse knowledge from existing results. Starting from a self-triggered impulsive consensus result, Andy formulates a global exponential leader-follower synchronization problem for delayed heterogeneous networks with switching communication topologies. The proposed hybrid control combines self-triggered impulses with execution delay and continuous feedback over a recovery window. After each delayed impulse, the feedback cancels the delayed error channel until the pre-impulse history leaves the active delay interval. Sufficient conditions for global exponential synchronization are established, Zeno behavior is excluded for both timing sequences, and a numerical example illustrates the result.

13:00 JSTエージェントビジネス/資金調達

LongRCA ベンチ: Long-Horizo​​n エージェント障害における責任ある役割と根本原因の診断

長期的なエージェントの実行が失敗した場合、結果レベルの評価によって失敗した結果が明らかになりますが、決定的なエラーが軌道に入った場所は明らかにされません。次に、開発者は完全な実行を検査して、責任のある役割を特定し、決定的な根本原因の最も早いステップを特定する必要があります。既存の障害属性ベンチマークは主に短いトレースに焦点を当てており、記録された数百のステップにわたる診断は十分に検討されていません。 LongRCA Bench を紹介します。これは、エラーが挿入されていない 5 つのドメインにわたる 1,140 個の失敗した軌跡で構成されています。これは、責任ある役割と最も早い決定的な根本原因ステップに対して、独立してスコア付けされた人間のラベルを提供します。中央軌道には 145 のステップが含まれており、最も強力なベースラインでもルート ステップの正確な精度は 13.2% にすぎません。さらに、セグメントサマリーから候補エラーステップを取得し、それらを以前の利用可能なハンドオフ命令まで追跡する、トレーニング不要の方法である根本原因軌跡アトリビューション(RCTA)を紹介します。同じバックボーン、ベンチマーク インスタンス、スコアリング プロトコルを使用することで、RCTA は責任のある役割の精度が 51.1%、ルートステップの精度が 24.1% に達しました。これらの結果は、責任のある役割の帰属と正確なルートステップの位置特定を、長期軌道障害診断の別のターゲットとして評価する必要性を強調しています。

原文 (English)

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.

13:00 JST研究/論文

グラフ手術と Do 演算子: 非周期的な構造因果モデルの正確な対応

$\operatorname{do}$-operator は、ターゲットへの矢印を削除することによってグラフィカルに記述され、メカニズムを定数に置き換えることによって機能的に記述されます。これらの操作を同等と呼ぶのは、まだ数学的な表現ではありません。一方はグラフを返し、ターゲットのみを記憶しますが、もう一方はメカニズムを返し、強制された値も記憶します。有限の多くの内生変数を持つ決定論的非循環構造因果モデルに対して、依存性レベルの正確な比較を行います。 $\operatorname{Graph}(F)$ がメカニズム ファミリ $F$ の依存関係を抽出する場合、主定理は $\operatorname{Graph}(F^\iota)=\operatorname{Surg}(\operatorname{Graph}(F),T_\iota)$ となります。したがって、ターゲットメカニズムを置き換えると、グラフ手術によって削除された依存関係が正確に削除されます。グラフに未使用の矢印が含まれる可能性があるモデル $M=(G,F)$ の場合、$\operatorname{Graph}(F)$ の代わりに $G$ を使用しても同じ等式が成り立つ場合を特徴付けます。 $G$ が $F$ の依存関係を正確に記録する場合、すべての介入に対してそれが当てはまります。次に、介入モデルを定義し、その実行を特徴付け、連続的な介入がどのように組み合わされるかを示し、結果が実際の依存関係の祖先での介入にのみ依存することを証明します。

原文 (English)

Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models

The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only the targets, whereas the other returns mechanisms and also remembers the imposed values. We make a dependency-level comparison precise for deterministic acyclic structural causal models with finitely many endogenous variables. If $\operatorname{Graph}(F)$ extracts the dependencies of a mechanism family $F$, our main theorem is $\operatorname{Graph}(F^\iota)=\operatorname{Surg}(\operatorname{Graph}(F),T_\iota)$. Thus replacing target mechanisms removes exactly the dependencies removed by graph surgery. For a model $M=(G,F)$ whose graph may contain unused arrows, we characterize when the same equality holds with $G$ in place of $\operatorname{Graph}(F)$; it holds for every intervention exactly when $G$ records the dependencies of $F$ exactly. We then define the intervened model, characterize its run, show how sequential interventions combine, and prove that an outcome depends only on interventions at its actual dependency ancestors. All principal results are machine-checked in an accompanying Lean 4 development.

13:00 JSTLLM/生成AI

どのネガティブなことが重要ですか?テキスト エンコーダーに尋ねる: 高密度キャプション取得のための適応型類似マージン

密なキャプションの取得は、最近、セグメンテーション、エッジ マップ、LLM フィルター処理されたキャプション、およびクロスモーダル モジュールをコントラストの微調整に導入することによって改善されました。ただし、これらの手法は主に同じ InfoNCE 目標を継承しており、その最適化は、強力な事前トレーニング済み初期化の下では時期尚早に飽和する可能性があります。つまり、密なキャプションでは、最初のエポック内のバッチの 80% で損失が 10^{-3} を下回りますが、その勾配は測定の 47% で fp32 で正確にゼロに達します。この動作は、高密度キャプション ベンチマークの多数のほぼ重複したキャプションと密接に関連していることがわかりました。つまり、簡単な大部分がすでに分離された後、いくつかの非常に類似したネガが未解決のままになっています。解決策として、HN-CLIP を導入します。HN-CLIP は、テキスト エンコーダー独自のテキスト間ジオメトリを使用して、ネガティブごとの適応類似マージンを構築します。具体的には、分離されたキャプション類似性行列がネガティブ ロジットに追加され、ネガティブのマイニング、合成、またはリサンプリングを行わずに、より類似したキャプションにより大きなマージンが割り当てられます。結果として得られる目標には、トレーニング中に 1 つのキャプション類似度行列とマスクされたロジット加算のみが必要であり、補助データ、追加パラメーター、オフライン前処理、または推論時間のオーバーヘッドはありません。 4 つの高密度キャプション検索ベンチマークに関する広範な実験により、HN-CLIP は最強の競合他社よりも +2.4 ~ +4.3 R@1 向上し、同時に GOAL よりも 2.4 倍、StructXLIP よりも 5.4 倍速くトレーニングできることがわかりました。さらに、提案された目標は、ドメイン内ベンチマークでテストされた 6 つの微調整フレームワークすべてを改善し、わずか 20% のトレーニング データで最強のフルデータ ベースラインに達します。

原文 (English)

Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^-3 on 80% of batches within the first epoch, while its gradient becomes numerically zero in 47% of measurements. We find that this behavior is closely related to the large number of near-duplicate captions in dense-caption benchmarks, where a few highly similar negatives remain unresolved after the easy majority has already been separated. As a remedy, we introduce HN-CLIP, which uses the text encoder's own text-text geometry to construct per-negative adaptive similarity margins. Specifically, a detached caption-similarity matrix is added to the negative logits, assigning larger margins to more similar captions without mining, synthesizing, or resampling negatives. The resulting objective requires only one caption-similarity matrix and a masked logit addition during training, with no auxiliary data, additional parameters, offline preprocessing, or inference-time overhead. Extensive experiments on four dense-caption retrieval benchmarks show that HN-CLIP improves over the strongest competitors by +2.5 to +4.0 R@1 while training 2.4x faster than GOAL and 5.4x faster than StructXLIP. Moreover, the proposed objective improves all six tested fine-tuning frameworks on the in-domain benchmarks and reaches the strongest full-data baseline with only 20% of the training data.

13:00 JSTエージェント

\textsc{TestifAI}: 深層学習システムのトモグラフィーベースのテスト

AI システムが安全性が重要なアプリケーション領域 (自動運転など) に導入されることが増えるにつれて、関連するリスクも増加します。したがって、最新の AI システムの基礎となる深層学習モデルは、正しい動作を保証するために徹底的なテストを受ける必要があります。単一のロバスト性テストには、入力の制限された摂動下でモデルの出力が安定しているかどうかを経験的に検証するために、何千もの推論が含まれます。しかし、既存のテスト フレームワークには、摂動の組み合わせ空間全体にわたる堅牢性を体系的に調査して要約する手段がありません。私たちは、摂動の組み合わせに対するロバスト性を効率的かつ正確に推定するための深層学習テスト フレームワークである TestifAI を提案します。 TestifAI を使用すると、ユーザーは動作条件をセマンティック入力摂動 (画像のぼやけ、明るさ、ズームなど) および離散的な重大度レベル (低、中、高など) の構造化空間として指定できます。ユーザーは、任意の組み合わせ (例: 「低ブラー、高輝度、中ズーム」) のモデルの堅牢性をクエリできます。効率と精度を達成するために、TestifAI は部分モデル トモグラフィーを導入しています。これは、少数の摂動 (低次投影) のみを適用するテストから多重摂動空間でモデルの動作を再構築する新しいアプローチです。少なくとも 3 つの摂動に対するロバスト性を推定するために、TestifAI は最大 2 つの摂動のみを含むテストの結果に基づいて補助モデルをトレーニングし、指数関数的な数のテストの実行を回避します。 5 つの画像および言語分類タスクに関する実験では、TestifAI が推論の数を 60 ~ 80% 削減しながら、低次 (1 および 2 摂動) の観測値から高次 (3 および 4 摂動) のテスト結果を 7% 未満の総ロバスト性推定誤差で予測できることがわかりました。

原文 (English)

TestifAI: Tomography-Based Testing for Deep Learning Systems

As AI systems are increasingly deployed in safety-critical application domains (e.g., autonomous driving), associated risks increase too. Deep learning models underlying modern AI systems, therefore, must undergo thorough testing to ensure their correct behaviour. A single robustness test involves thousands of inferences to empirically verify if a model's outputs remain stable under a bounded perturbation of its inputs. However, existing testing frameworks lack the means to systematically explore and summarise robustness across a combinatorial space of perturbations. We propose TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations. TestifAI enables users to specify operational conditions as structured spaces of semantic input perturbations (e.g., image blur, brightness and zoom) and discrete severity levels (e.g., low, medium and high). Users can query model robustness for any combination (e.g., "low blur, high brightness, and medium zoom"). To achieve efficiency and accuracy, TestifAI introduces partial model tomography, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations (lower-order projections). To estimate robustness against at least three perturbations, TestifAI trains an auxiliary model on the results of tests involving up to two perturbations only, avoiding execution of an exponential number of tests. Our experiments on five image and language classification tasks show that TestifAI can predict higher-order (3 and 4 perturbations) test outcomes from low-order (1 and 2 perturbations) observations with an aggregate robustness estimation error of less than 7%, while reducing the number of inferences by 60-80%.

13:00 JST研究/論文

Learning-Based Speed Estimation from Accelerometer-Only Inertial Sensing

The proposed model, CarSpeedNet, estimates scalar vehicle speed from a window of three-axis smartphone acceleration, without gyroscope, whe…

13:00 JST研究/論文

Teacher-free Latent Self-distillation and Class-separable Representations for Lightweight IoT Attack Detection

Knowledge distillation (KD) has been widely used to improve lightweight AI models by transferring soft-label knowledge from a large teacher…

13:00 JST研究/論文GPT / ChatGPT

Towards Efficient Pareto Set Approximation via Mixture of Experts Based Model Fusion

Solving multi-objective optimization problems for large deep neural networks is a challenging task due to the complexity of the loss landsc…

13:00 JST画像/動画生成

Your Turn: At Home Turning Angle Estimation for Parkinson's Disease Severity Assessment

People with Parkinson's Disease (PD) often experience progressively worsening gait, including changes in how they turn around, as the disea…

13:00 JST研究/論文

Virtual Sensing to Enable Real-Time Monitoring of Inaccessible Locations & Unmeasurable Parameters

Real-time monitoring of safety-critical interior states is an open problem across energy, environmental and industrial systems where direct…

13:00 JST研究/論文

SMOTE-Tomek 前処理による要件分類の改善

この研究では、SMOTE-Tomek 前処理技術を層別 K 分割相互検証と組み合わせて適用し、PROMISE データセット内のクラスの不均衡に対処することで、要件エンジニアリングの領域に重点を置いています。このデータセットは、機能タイプと非機能タイプに分類された 969 個の分類された要件で構成されています。提案されたアプローチは、検証フォールドの整合性を維持しながら少数派クラスの表現を強化し、分類精度の顕著な向上につながります。ロジスティック回帰は 76.16\% を達成し、ベースラインの 58.31\% を大幅に上回りました。これらの結果は、スケーラブルで解釈可能なソリューションとしての機械学習モデルの適用性と効率性を強調しています。

原文 (English)

Improving Requirements Classification with SMOTE-Tomek Preprocessing

This study emphasizes the domain of requirements engineering by applying the SMOTE-Tomek preprocessing technique, combined with stratified K-fold cross-validation, to address class imbalance in the PROMISE dataset. This dataset comprises 969 categorized requirements, classified into functional and non-functional types. The proposed approach enhances the representation of minority classes while maintaining the integrity of validation folds, leading to a notable improvement in classification accuracy. Logistic regression achieved 76.16%, significantly surpassing the baseline of 59.85%. These results highlight the applicability and efficiency of machine learning models as scalable and interpretable solutions.

13:00 JST画像/動画生成

Regressor-Guided Image Editing Shifts Emotion and Disengagement Timing in Social Media

Internet overuse is a widespread phenomenon in today's digital society. Existing interventions, such as time limits or grayscaling, often r…

13:00 JSTLLM/生成AI

HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings

Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Busi…

13:00 JST画像/動画生成

ReynoldsFlow: Physics-Inspired Spatiotemporal Flow Representation for Video Understanding

Video understanding has largely relied on deep spatiotemporal architectures, including 3D convolutional networks and optical flow (OF) base…

13:00 JST研究/論文

Learning-Augmented Power System Operations: A Unified Optimization View

With the increasing penetration of renewable energy and inverter-based resources, traditional physics-based power-system operation faces gr…

13:00 JST研究/論文

The Thousand Brains Theory 2.0: An Extension for the Long-Range Connections of the Neocortical Heterarchy

Vernon Mountcastle hypothesized that the basis for intelligence in mammals is the replication of a general computational unit, the cortical…

13:00 JSTLLM/生成AIエージェント

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization

Prefix adders are fundamental arithmetic circuits, but their design space grows exponentially with bit-width, posing significant optimizati…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices

Large language models (LLMs) increasingly act on people's behalf: they write emails, buy groceries, and book restaurants. While the outsour…

13:00 JSTLLM/生成AI研究/論文

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks oft…

13:00 JSTロボティクス

FMT$^{\mathrm{X}}$: Lazy Wavefront Search for Dynamic Replanning

FMT$^{*}$ plans efficiently in static worlds by expanding a cost-ordered wavefront and collision-checking lazily, but its single-pass unvis…

13:00 JSTLLM/生成AI

TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning

Time series reasoning is crucial to decision-making in diverse domains, including finance, energy, and scientific discovery. While existing…

13:00 JSTロボティクス

SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation

Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily rel…

13:00 JST研究/論文

The Bidding Games: Reinforcement Learning for MEV Extraction on Polygon Blockchain

In blockchain networks, the strategic ordering of transactions within blocks has emerged as a significant source of profit extraction, know…

13:00 JST画像/動画生成

MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal Prostate MRI Segmentation

Active Surveillance (AS) is a treatment option for managing low and intermediate-risk prostate cancer (PCa), aiming to avoid overtreatment…

13:00 JSTロボティクス

A Physics-Informed Neural Network Approach for UAV Path Planning in Dynamic Environments

Unmanned aerial vehicles (UAVs) operating in dynamic wind fields must generate safe and energy-efficient trajectories under physical and en…

13:00 JSTエージェント

PACT: Phenotype-Aware Contrastive Team Representation for Multi-Phenotype Grouped Ad Hoc Teamwork

Learning to collaborate with various unfamiliar teammates poses a great challenge in the domain of multi-agent systems. Existing ad hoc tea…

13:00 JSTロボティクス研究/論文

WaveVerif: Acoustic Side-Channel based Verification of Robotic Workflows

In this paper, we present a framework that uses acoustic side-channel analysis (ASCA) to monitor and verify whether a robot correctly execu…

13:00 JSTLLM/生成AI

Towards Audio Token Compression in Large Audio Language Models

Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate t…

13:00 JSTLLM/生成AI画像/動画生成

Extended to Reality: Prompt Injection in 3D Environments

Multimodal large language models (MLLMs) have advanced the capabilities to interpret and act on visual input in 3D environments, empowering…

13:00 JSTLLM/生成AIビジネス/資金調達

Qworld: Question-Specific Evaluation Criteria for LLMs

Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Bi…

13:00 JSTビジネス/資金調達

Power Couple? AI Growth and Renewable Energy Investment

Artificial intelligence (AI) and renewable energy are increasingly being described as a \mbox{``power couple,''} based on the idea that rap…

13:00 JST研究/論文

Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework

Despite the growing use of generative artificial intelligence (GenAI) in entrepreneurship, research on its impact remains fragmented. To ad…

13:00 JSTLLM/生成AI画像/動画生成

VISD: 構造化自己蒸留によるビデオ推論の強化

複雑な推論のために VideoLLM をトレーニングすることは、疎なシーケンス レベルの報酬と、時間的に根拠のある長い推論軌跡にわたるきめ細かい単位の割り当てが欠如しているため、依然として困難です。検証可能な報酬を伴う強化学習 (RLVR) は信頼性の高い監視を提供しますが、トークン レベルの寄与を捕捉できず、非効率的な学習につながります。逆に、既存の自己蒸留手法は緻密な監視を提供しますが、構造と診断の特異性に欠けており、強化学習と不安定に相互作用することがよくあります。この研究では、ビデオ推論に診断的に意味のある特権情報を導入する構造化自己蒸留フレームワークである VISD を提案します。 VISD は、ビデオ対応の判定モデルを採用して、推論の質を解答の正しさ、論理的一貫性、時空間的根拠などの複数の次元に分解し、この構造化されたフィードバックを使用して、トークン レベルの監督のための教師のポリシーを導きます。高密度監視を RL と安定して統合するために、方向振幅デカップリング メカニズムを導入します。このメカニズムでは、報酬から計算されたロールアウト レベルの利点が更新方向を決定し、構造化された特権信号がトークン レベルの更新振幅を調整します。この設計により、意味的に調整されたきめ細かい単位の割り当てが可能になり、推論の忠実さとトレーニングの効率の両方が向上します。さらに、VISD にはカリキュラムのスケジューリングと EMA ベースの教師の安定化が組み込まれており、長いビデオ シーケンスにわたる堅牢な最適化をサポートします。さまざまなベンチマークの実験では、VISD が一貫して強力なベースラインを上回り、回答の精度と空間時間的グラウンディングの品質が向上していることが示されています。特に、VISD は最適化ステップでほぼ 2 倍高速な収束でこれらの利点を達成しており、VideoLLM のパフォーマンスとサンプル効率の両方を向上させる構造化自己監視の有効性を強調しています。

原文 (English)

VISD: Enhancing Video Reasoning via Structured Self-Distillation

Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reasoning trajectories. While reinforcement learning with verifiable rewards (RLVR) provides reliable supervision, it fails to capture token level contributions, leading to inefficient learning. Conversely, existing self distillation methods offer dense supervision but lack structure and diagnostic specificity, and often interact unstably with reinforcement learning. In this work, we propose VISD, a structured self distillation framework that introduces diagnostically meaningful privileged information for video reasoning. VISD employs a video aware judge model to decompose reasoning quality into multiple dimensions, including answer correctness, logical consistency, and spatio-temporal grounding, and uses this structured feedback to guide a teacher policy for token level supervision. To stably integrate dense supervision with RL, we introduce a direction magnitude decoupling mechanism, where rollout level advantages computed from rewards determine update direction, while structured privileged signals modulate token level update magnitudes. This design enables semantically aligned and fine grained credit assignment, improving both reasoning faithfulness and training efficiency. Additionally, VISD incorporates curriculum scheduling and EMA based teacher stabilization to support robust optimization over long video sequences. Experiments on diverse benchmarks show that VISD consistently outperforms strong baselines, improving answer accuracy and spatio temporal grounding quality. Notably, VISD reaches these gains with nearly 2x faster convergence in optimization steps, highlighting the effectiveness of structured self supervision in improving both performance and sample efficiency for VideoLLMs.

13:00 JST画像/動画生成

マルチモーダルナレッジグラフと信頼性に基づいた改良による症例を意識した医療画像分類

ディープラーニングは医療画像の分類に大きな進歩をもたらしましたが、既存の手法のほとんどは依然として個別の視覚的証拠に依存しており、類似の症例や外部の知識を効果的に活用できません。臨床現場では通常、診断は同様の歴史的症例とそれに関連する症状によって裏付けられます。この証拠に基づいた診断プロセスを明示的にモデル化するために、医用画像分類のためのマルチモーダル知識グラフによって駆動されるケース認識推論フレームワークを提案します。具体的には、疾患、画像、症状が階層的に整理された構造化された診断記憶として、症例を意識したマルチモーダルナレッジグラフを構築します。入力画像が与えられると、私たちの方法はこのメモリから類似のケースを適応的に取得し、対応するケース中心のサブグラフを抽出します。さらに、画像中心のグラフ アテンション ネットワークが異種セマンティクスをケースベースの特徴に集約する知識伝播および注入メカニズムを導入します。その後、これらの特徴を視覚表現に注入してクロスモーダル アラインメントを行う双方向のクロスモーダル アテンション メカニズムが続きます。ノイズの多い検索を軽減するために、予測の信頼性とサンプルの類似性を組み合わせて考慮することで検索された各ケースの信頼性を推定し、最終的な予測への寄与を再重み付けして、解釈可能なケースレベルの証拠を提供する、信頼度調整された決定改良スキームを設計します。複数の医用画像データセットに対する広範な実験により、私たちのアプローチが一貫して強力なベースラインを上回るパフォーマンスを示し、アブレーションと定性分析によりその有効性と解釈可能性が検証されました。コードは https://anonymous.4open.science/r/MKG-CARE-8B7B で入手できます。

原文 (English)

MKG-CARE: Case-Aware Reasoning with Multimodal Knowledge Graphs for Explainable Medical Image Diagnosis

Medical image diagnosis has achieved significant progress with deep learning, yet existing methods often rely on isolated visual evidence and lack the ability to effectively leverage similar cases and external knowledge. In clinical practice, diagnosis is typically supported by similar historical cases and their associated symptoms. To explicitly model this evidence-based diagnostic process, we propose MKG-CARE, a framework that performs case-aware reasoning using multimodal knowledge graphs for explainable medical image diagnosis. Specifically, we construct a case-aware multimodal knowledge graph as a structured diagnostic memory, where diseases, images, and symptoms are hierarchically organized. Given an input image, MKG-CARE adaptively retrieves similar cases from this memory and extracts their corresponding case-centered subgraphs. We further introduce a knowledge propagation and injection mechanism, where an image-centric Graph Attention Network aggregates heterogeneous semantics within the retrieved case subgraphs, followed by bidirectional cross-modal attention to align and inject the aggregated case knowledge into visual representations. To mitigate retrieval noise, we design a confidence-calibrated decision refinement scheme that estimates each retrieved case's reliability from prediction confidence and sample similarity, and reweights its contribution to the final prediction for interpretable case-level evidence attribution. Extensive experiments on multiple medical imaging datasets demonstrate consistent improvements over strong baselines, while ablation and qualitative analyses validate the effectiveness and interpretability of our method. The code is available at https://github.com/lyxuan1022/MKG-CARE.

13:00 JST研究/論文

コンセプト割り当てゾーン: トランスの深さ全体でコンセプトがどのように形成されるかを追跡

トランスフォーマー言語モデルにおける概念形成は、単一層のイベントではなく、深さまで拡張されます。概念は、残差ストリームの連続領域全体にわたって徐々に現れます。機械的解釈可能性メソッドは、プロセス自体ではなくスナップショットをキャプチャするピーク クラス分離の単一レイヤー (「最良のレイヤー」) を特定します。概念割り当てゾーン (CAZ) を導入します。これは、概念が測定可能に分離可能になる深さの間隔であり、その幾何学的表現に割り当てられる領域です。 3 つのレイヤーごとのメトリクス (分離、コンセプトの一貫性、コンセプトの速度) を通じて CAZ を形式化し、手動のレイヤー スイープなしで原則に基づいた境界検出を導き出します。 CAZ は概念ではありません。CAZ は、概念を分離可能にするためにモデルがそのジオメトリを編成する深度領域です。通常、単一のコンセプトは複数の CAZ に参加します。複数の概念が 1 つを共有する場合があります。 8 つのアーキテクチャ ファミリと 7 つの概念からの 34 のモデルにわたる経験的検証により、分離曲線 S(l) が多峰性であることが多いことが明らかになりました。スコア付けされた検出器は、「ジェントル CAZ」を明らかにします。これは、標準的なピーク検出では見えないが、アブレーション下の症例の 93 ~ 100% で因果的に活性化する微妙な割り当て領域です (34 モデル中 16 モデル、コンパニオン検証論文では 26 モデル)。このフレームワークは 7 つのテスト可能な予測を生成します。 4 つは明確な判定を下し (2 つはサポートされていない、1 つは部分的にサポートされ、1 つはサポートされている)、1 つは前提条件がデータによって無効になっており、2 つは検出力が不十分です。クロス アーキテクチャの調整は、コンセプトを 1 つ残す相互検証でモノリシックではなく深さが一致していることが確認されています。参照実装:rosetta_tools v1.3.1 (doi:10.5281/zenodo.20361433)。

原文 (English)

The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth

Concept formation in transformer language models is a depth-extended process, not a single-layer event: a concept becomes separable across one or more contiguous regions of the residual stream - its Concept Allocation Zone (CAZ). A CAZ is not a concept but the depth segment where the model organizes its geometry to make one separable - concepts may share a CAZ, and typically span multiple across depth; the companion GEM paper shows the separating direction continues to rotate within a CAZ before stabilizing past its boundary. We formalize the CAZ through three layer-wise metrics - Separation, Concept Coherence, and Concept Velocity - with automated boundary detection that applies no significance threshold to CAZ membership (every segment is a CAZ; "strong" vs. "gentle" is score, never a binary cut). Empirical validation across 35 models, 8 architectural families, and 7 concepts shows the separation curve S(l) is frequently multimodal, and scored detection surfaces a further category of subtle allocation regions ("gentle CAZes") invisible to standard peak detection. The framework generates seven testable predictions; its contribution is the instrument and the phenomena it surfaces - the scored detector, the three metrics, and the multimodal/gentle-CAZ findings - not the predictions themselves. Released as the open-source rosetta_tools library (v1.3.1).

13:00 JSTLLM/生成AI

最適化下でのコンテキストとパラメトリックな思考連鎖の忠実性の間の相互作用の調査

思考連鎖 (CoT) の忠実度、つまり CoT が大規模言語モデル (LLM) の基礎となる動作を本当に反映しているかどうかは、通常、2 つの独立したパラダイムに基づいて評価されます。1 つは入力または CoT トレースを摂動させることで測定されるコンテキストの忠実度、もう 1 つはモデルのパラメトリック知識に介入することで評価されるパラメトリックの忠実度です。しかし、これまでの研究では、それらを説明的に比較しているだけです。私たちは、どちらかの忠実性パラダイムに向けてモデルを最適化するための統合された好み調整インターフェイスである FaithMate を提案することで、このギャップを埋めます。これにより、2 つのパラダイム間の相互作用を調査し、忠実性の向上がパラダイム内およびパラダイム間で一般化するかどうか、またどの程度まで一般化するかを調べることができます。 3 つのモデル、2 つのデータセット、および 6 つの忠実度メトリクスにわたって、2 つのパラダイムが積極的に結合しているものの、非対称であることがわかりました。パラメトリックの忠実性に向けて最適化すると、両方のパラダイムにわたって一貫したゲインが得られますが、コンテキストに対応するパラダイムは、より可変的なゲインをもたらします。コンテキスト パラダイム内では、あるメトリクスでの忠実度の向上が他のメトリクスに一貫して移行するわけではありません。これは、既存のコンテキスト メトリクスが忠実度の素の側面を捉えていることを意味し、固有のトレードオフを明らかにします。これらの発見は、CoT の忠実性が単一の目標ではないため、多面的な最適化と評価が必要であることを示唆しています。

原文 (English)

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint paradigms: contextual faithfulness, measured by perturbing the input or CoT trace, and parametric faithfulness, assessed by intervening on a model's parametric knowledge. Yet prior work compares them only descriptively. We fill this gap by proposing FaithMATE, a unified preference-alignment interface for optimizing models towards either faithfulness paradigm. It enables us to investigate the interplay between the two paradigms, examining whether and to what extent faithfulness gains generalize within and across paradigms. Across three models, two datasets, and six faithfulness metrics, we find that the two paradigms are positively coupled, yet asymmetric: optimizing towards parametric faithfulness yields consistent gains across both paradigms, whereas the contextual counterpart delivers more variable gains. Within the contextual paradigm, faithfulness gains on one metric do not consistently transfer to others, implying that existing contextual metrics capture disjoint facets of faithfulness and exposing inherent trade-offs. These findings imply that CoT faithfulness is not a monolithic objective and therefore requires multifaceted optimization and evaluation.

13:00 JST研究/論文

幾何学的進化マップ: 変圧器の残留ストリームからの安定した概念プローブの抽出

変圧器の残留ストリームから抽出されたコンセプト プローブの信頼性は、抽出元のレイヤーと同程度です。固定の後期層または分離スコア関数のピークでプローブする一般的な手法は、基本的な構造的特徴を無視しています。つまり、概念表現は組み立てフェーズ中に大幅な方向回転を受け、主要な概念割り当てゾーン (CAZ) の後の特徴的なハンドオフ層まで安定した方向に落ち着きません。幾何学的進化マップ (GEM) を導入します。これは、残留ストリームのアクティブ化を通じて概念の完全な方向の軌跡を追跡し、回転が停止するハンドオフ層を特定し、その層から確定したプローブ方向を抽出します。 70M ~ 14B のパラメータと 17 の概念タイプにわたる 23 のアーキテクチャ全体で、CAZ 内の入口から出口までのコサイン類似度は平均 0.233 であり、CAZ 入口でのプローブ方向が出口でのプローブ方向を確実に予測できないことを示しています。 391 のコンセプト x モデルのペア (23 モデル x 17 コンセプト) にわたるアブレーション実験では、GEM 抽出プローブが 268/391 トライアル (68.5%) ではピーク層プローブと少なくとも同等の精度を示し、259/391 (66.2%) では厳密にそれを上回っています。アーキテクチャの分割は顕著です。MHA モデルは 173/221 トライアル (78.3%) でハンドオフを支持します。 GQA モデルは、119 件中 56 件の試行 (47.1%) でのみハンドオフを支持しました。モデルレベルのウィルコクソン: W=214、N=23、p=0.010 (片側)。適応型アブレーション幅ルールは、最終層に近い 79/391 のケースを対象としています。トリガーされたケース 60/79 (75.9%) でプローブの品質が向上し、平均利得 +7.44pp です。方向特異性制御により、アブレーション効果がコンセプト方向特異的であることが確認されます。抑制率の中央値は 377 倍、ランダム方向アブレーションと比較します (コンセプト方向の 99.1% が 10 個のランダム シードすべてに勝利しました)。参照実装:rosetta_tools v1.3.1 (doi:10.5281/zenodo.20361433)。

原文 (English)

Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams

A concept probe is only as reliable as the layer it is taken from. Probing at a fixed late layer, or at the peak of a separation curve, ignores a structural feature of how concepts form: the probe direction rotates substantially during assembly and does not settle until after the Concept Allocation Zone (CAZ) in which it forms. We introduce Geometric Evolution Maps (GEMs). A GEM records a concept's directional trajectory across one CAZ segment, takes the settled direction at that segment's final layer, and locates the handoff layer immediately beyond it - the first layer at which that direction can be evaluated outside the window it was estimated on. The handoff layer is derived from the segment boundary, not detected by a rotation criterion. A concept typically occupies several segments; we call that set its atlas. This rotation is substantial and consistent across architectures and concept types - a property of the representation, not an estimation artifact. The extracted direction is causal: ablating it suppresses far more separation than ablating a random direction, though no single site carries the effect alone. Probing after rotation completes is more precise than probing during it; a depth-matched control shows this advantage comes from probing at greater depth, not any privilege specific to the handoff boundary. A small number of structured exceptions are documented rather than left unaccounted for. A concept's atlas - its full set of GEMs, not any single one of them - is therefore the unit this method delivers for downstream use. GEM provides the ablation-targeting methodology used by the CAZ Validation paper in this program.

13:00 JSTLLM/生成AIエージェント

Beyond Recall: AI パーソナライゼーションの解釈層としての行動仕様

AI エージェントが人間に代わって意思決定を行う場合、その意思決定はユーザーと一致する必要があります。システムが人の解釈をどれだけ忠実に捉えているかを測定するために、表現精度を導入します。解釈層は動作仕様として運用されます。私たちのリファレンス実装は、人のデータを解釈パターンに積極的に圧縮し、言語モデルのコンテキストとして機能します。私たちは、校正済みの 5 人の審査員 LLM パネルによって採点された、保留された行動予測のプロトタイプ ベンチマークで仕様を評価します。私たちは、完全な生のコーパス、完全に抽出されたファクト、および 4 つの商用メモリ システム (Mem0、Letta、Supermemory、Zep) など、さまざまなコンテキスト条件を使用して独立して構成してテストします。この仕様は 14 のパブリック ドメインの自伝的コーパスにわたって、集合的に表現の精度を向上させ、モデルのヘッジをほぼ排除します。生のコーパスが提供する内容のほとんどを、コンテキスト コストを約 25 分の 1 に抑えて復元します。この仕様は、トレーニング前のベースラインに関係なく、被験者を共通の予測レベルに引き上げます。したがって、絶対ポイントのリフトはベースラインが最も低いところで最大となり、関連する母集団が事前トレーニングで適切に代表されていない人であることを示唆しています。リフトは、解釈が必要な質問で最大であり、解釈レイヤーを提供することで、抽出された事実や生のコーパスでは実現できないモデル動作が可能になります。逆に、リコールが必要な質問では、この層は役立つというよりむしろ邪魔になる可能性があります。私たちは、表現の精度は再現とは異なり、人間と AI の整合性はユーザーがどれだけ正確に表現されているかに依存すると結論付けています。表現が正確であるため、その調整はテスト可能です。

原文 (English)

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how faithfully a system captures a person's interpretation. An interpretive layer is operationalized as a Behavioral Specification. Our reference implementation aggressively compresses a person's data into interpretive patterns, served as context to a language model. We evaluate the Specification on a prototype benchmark of held-out behavioral predictions scored by a calibrated 5-judge LLM panel. We test it independently and in composition with a range of context conditions: full raw corpus, full extracted facts, and four commercial memory systems (Mem0, Letta, Supermemory, Zep). Across 14 public-domain autobiographical corpora, the Specification lifts representational accuracy in aggregate and nearly eliminates model hedging. It recovers most of what the raw corpus delivers, at ~25x less context cost. The Specification lifts subjects toward a common predictive level regardless of pretraining baseline; the lift in absolute points is therefore largest where the baseline is lowest, suggesting the population of relevance is anyone not adequately represented in pretraining. Lift is greatest on interpretation-required questions, where providing an interpretive layer enables model behavior that extracted facts or raw corpus do not. Conversely, on recall-required questions, this layer can interfere rather than help. We conclude that representational accuracy is distinct from recall and that human-AI alignment is dependent on how accurately the user is represented. Representational accuracy makes that alignment testable.

13:00 JST研究/論文

表形式予測のマルコフ境界の善、悪、醜

標準的なグラフィックの仮定の下では、ターゲット変数のマルコフ境界は、他のすべての特徴を冗長にする最小の特徴セットです。境界が観察されると、ターゲットは条件付きでテーブルの残りの部分から独立します。これは、モデルに必要な列を正確に指定するため、表形式の予測にとって魅力的なオブジェクトです。しかし、現代のリグレッサーは依然として完全な機能セットについてトレーニングを受けています。マルコフ境界が SCM3K (特徴数 40 ~ 1000 の 3,450 タスクの合成 SCM ベンチマーク) および 6 つの SCM ファミリを 6 つのリグレッサーで評価する場合の予測に本当に役立つかどうかを尋ねます。答えは理論が示唆するよりも微妙です。リグレッサーをオラクル境界に制限すると、多くの場合、予測が大幅に改善され、特徴空間が大きくなり、まばらになるにつれて改善も大きくなります。しかし、因果関係の発見と回復されたマスクのトレーニングによって境界を回復するという自然なパイプラインは機能しません。既存の推定ツールは、境界が最も役立つ領域に到達する前にコンピューティング バジェットを使い果たしてしまい、どこで実行しても完全な機能セットを超えることはほとんどありません。これを 3 つの原因で追跡します。 Discovery は、予測ではなく構造回復を最適化します。偽陰性と偽陽性は、著しく非対称な予測コストをもたらします。正確な境界は、すべての特徴を超える多くの特徴セットのうちの 1 つにすぎません。次に、これらの事実が、予測に合わせた特徴選択と、因果構造の使用を学習する表形式モデルに対して何を意味するかを開発します。

原文 (English)

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant. Once the boundary is observed, the target is conditionally independent of the rest of the table. This is a tempting object for tabular prediction, since it names exactly the columns a model should need. Yet modern regressors are still trained on the full feature set. We ask whether the Markov boundary is genuinely useful for prediction on SCM3K, a 3,450-task synthetic SCM benchmark with feature counts from 40 to 1000 and six SCM families, evaluated with six regressors. The answer is more nuanced than the theory suggests. Restricting a regressor to the oracle boundary often improves prediction substantially, and the improvement grows as the feature space becomes larger and sparser. But the natural pipeline of recovering the boundary with causal discovery and training on the recovered mask does not deliver. Existing estimators exhaust the compute budget before reaching the regime where the boundary helps most, and even where they run they rarely beat the full feature set. We trace this to three causes. Discovery optimizes structural recovery rather than prediction. False negatives and false positives carry sharply asymmetric predictive cost. The exact boundary is only one of many feature sets that beat all features. We then develop what these facts imply for prediction-aligned feature selection and for tabular models that learn to use causal structure.

13:00 JSTロボティクス

予測されたダイナミクスは物理世界に存在できますか?

予測物理 AI システムは状態ロールアウト、アクション チャンク、潜在計画を出力しますが、二乗平均平方根誤差 (RMSE) が低いということは、特定の提案が物理的に実行可能であることを意味するものではありません。物理的な許容性を予測制御インターフェイスとして定式化します。実行前に、デコードされた提案が候補ダイナミクスとして扱われ、運動学的、動的、および直接合成ホライズン条件を使用して評価されます。合格はタスクの成功を証明するものではありません。拒否は、指定された物理エンベロープの違反を識別し、コンポーネント レベルの理由を示します。 Hugging Face LeRobot PushT では、制御された改ざんにより、ワンステップ予測 RMSE と標準化されたダイナミクス残差が受信者動作特性曲線下領域 (AUC) 0.982 および 0.972 に達し、運動学のみの条件が AUC 0.592 に達し、フルゲートが条件レベルの帰属で AUC 0.957 に達することが示されています。リプレイベースの介入実験では、残差ベースのフィルターと完全な物理的許容ゲートにより、無効な提案の 87 ~ 89% が防止され、平均進行状況が 0.998 近くに維持されます。

原文 (English)

Can Predicted Dynamics Exist in the Physical World?

Can learned state-action proposals exist in the physical world? To filter infeasible commands before execution, policies are often wrapped in a runtime monitor. However, aggregating diverse diagnostic signals obscures whether a proposal violates dynamic transitions or merely departs from recorded behavior. We formalize this prediction-control interface and prove that the all-pairs displacement term is redundant within a max-aggregated composite. We evaluate these monitors on 700 nominal and 5,250 synthetically perturbed 32-transition PushT windows, monitoring only planar pusher positions and goals. A simple transition-RMSE baseline (AUC 0.982) outperforms a heterogeneous max-aggregated monitor (AUC 0.957). We conclude that physical transition checks must be strictly separated from empirical logs.

13:00 JST研究/論文

整流された流れにおける対照的速度マッチングによる幾何学的消去

マルチモーダル生成モデルの急速な導入は計り知れない可能性をもたらしますが、有害なコンテンツの合成、ディープフェイク、著作権侵害のリスクも増大しています。これらの課題に対処するために、将来の安全策として概念の消去が登場しました。しかし、この分野が U-Net ベースの拡散モデルから整流変圧器に徐々に移行するにつれて、消去研究は追いつくのに苦労しています。この作業では、Rectified Flow モデル用のシンプルだが非常に効果的な消去フレームワークである GEM を紹介します。私たちの貢献の一環として、私たちはジェネレーティブ フロー ネットワークに基づいた軌道ベースのアンラーニングと古典的な教師主導の消去との間に原理的な橋渡しを確立します。つまり、軌道ベースの信号を、両方のパラダイムの長所を統合する教師主導のフローマッチング設定に変換します。具体的には、教師は相補的な引力と反発の信号を提供し、それらを単一の幾何学的な指導目標に組み合わせて、無害な生成を維持しながら不要な概念をターゲットに抑制します。

原文 (English)

GEM: Geometric Erasure by Contrastive Velocity Matching in Rectified Flows

While the rapid adoption of multimodal generative models offers immense potential, it has also increased the risks of harmful content synthesis, deepfakes, and copyright infringements. To address these challenges, concept erasure has emerged as a prospective safeguard. However, as the field gradually transitions from U-Net-based diffusion models to Rectified Flow Transformers, erasure research has struggled to keep pace. In this work, we introduce GEM, a simple but highly effective erasure framework for Rectified Flow models. As part of our contribution, we establish a principled bridge between trajectory-based unlearning grounded in Generative Flow Networks and classic teacher-guided erasure: we translate trajectory-based signals into a teacher-guided flow-matching setup that unifies the strengths of both paradigms. Concretely, a teacher provides complementary attraction and repulsion signals that we combine into a single geometric guidance objective, yielding targeted suppression of unwanted concepts while preserving benign generation.

13:00 JSTロボティクス

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides…

13:00 JSTエージェント研究/論文

Right Family, Wrong Skill: Benchmarking Risk Exposure in Agent Skill Retrieval

Agent skill libraries are becoming routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and…

13:00 JST研究/論文

米国における AI プログラムのマッピング: 2026 年初頭の現状レポートと AI メジャーおよびマイナーの分析

私たちは、2026 年春の米国の学部人工知能 (AI) プログラムの状況に関するレポートを発表します。その際、1) 米国の AI 教育の状況を追跡するために動的に更新されるスクレイピング ツールとマッピング ツールについて説明し、2) 大きな激変の時期に歴史的な記録を作成します。私たちが開発したツールは https://cicmap.ai で入手可能で、4 年制大学の 350 以上の学部 AI プログラム (専攻、副専攻、専修科目、証明書) からのデータを検出、収集、表示します。当社のツールは、これらのプログラムを見つけるために 560 以上の教育機関を検索しました。これは、米国のコンピューター サイエンス (CS) 学部卒業生の 86% に相当します。このツールを使用すると、入学予定の学生、指導カウンセラー、管理者、教員が AI プログラムの要件に簡単にアクセスでき、新しいプログラムが登場するたびに継続的に更新されるように設計されています。私たちの知る限り、この調査は米国における AI プログラムの現状を示すこれまでで最も包括的なスナップショットを表しています。この研究により、私たちは 3 つの重要な貢献を提供します。1) 大変動期の米国における AI プログラムの記録。 2) AI プログラムとその要件を調査するツール。 3) 66 の AI 専攻と 87 の AI 副専攻に必要なコースの分析。専攻と副専攻の分析では、学位の規模と要件に大きなばらつきがあることが示されていますが、2 つの点に注目します。まず、すべての専攻で一般的な AI コースが必要なわけではありませんが、そうでない場合は機械学習 (ML) コースが必要です。第二に、専攻の 3 分の 1 以上が AI の倫理コースを必要としている一方で、AI 専攻の専攻の 4 分の 1 弱が必須です。

原文 (English)

A Tool to Map AI Programs in the U.S.: A Snapshot from April 2026 and an Analysis of Requirements for AI Majors and Minors

In this work, we locate and analyze existing undergraduate Artificial Intelligence (AI) programs in the United States in Spring 2026, creating a historic record at a time of great change in this area. To create this record, we developed a tool to detect, scrape, and display data from 361 undergraduate AI programs--majors, minors, concentrations, and certificates--at 4-year universities. Our tool, available at https://cicmap.ai, searched 563 institutions to locate these programs, a sample that represents 87% of all undergraduate Computer Science (CS) graduates in the U.S in 2025. This tool allows prospective students, guidance counselors, administrators, and faculty to easily access AI program requirements and is designed to continually update as new programs emerge. To the best of our knowledge, this survey represents the most comprehensive snapshot of the state of AI programs in the U.S. to date. With this work we offer three important contributions: 1) a record of AI programs in the U.S. at a time of great upheaval; 2) a tool to explore AI programs and their requirements; and 3) an analysis of the courses required for 66 AI majors and 87 AI minors. Our analysis of majors and minors shows great variability in the size and the requirements of these degrees, but we note two takeaways. First, not all majors require a general AI course, but if they don't, they do require a Machine Learning (ML) course. Second, more than a third of majors require an Ethics in AI course but only 24% of minors do.

13:00 JSTロボティクス

PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty

Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Ob…

13:00 JST画像/動画生成

あらゆるステップ: ビデオベースのパーキンソン病患者の方向転換歩数計測

パーキンソン病 (PD) の顕著な症状として、旋回障害は、旋回角度、持続時間、特に旋回を完了するまでに必要な歩数などのパラメータを通じて評価され、運動機能障害を直接反映します。現実世界の回転動作にはばらつきがあり、パーキンソン病の歩行では非定型的な足を引きずるパターンがあるため、正確な歩数計測は困難です。既存の方法は主にウェアラブルベースであり、ユーザーは専用デバイスを装着して管理する必要があり、毎日継続的に使用するには不便な場合があります。これに対処するために、多様な動作表現を使用して粗い方法から細かい方法まで歩数を推定する、受動的なビデオベースのフレームワークを提案します。具体的には、3D ヒューマン メッシュの復元から得られた足の動きの信号から初期歩数を推定し、高レベルの動作構造を提供します。きめの細かいモーションの詳細を組み込むために、モーション エンコーダーはメッシュとオプティカル フローから相補的な歩行ダイナミクスを学習し、初期推定を改良します。このプロセスでは、粗い足の動きの信号が交差注意を介してピクセルレベルの動きの合図を照会し、微妙なパーキンソン病の歩行ダイナミクスを捕捉します。さまざまなビデオの長さを処理するために、各ビデオをクリップに分割し、歩数残差予測のためにマルチ インスタンス学習 (MIL) を介してクリップごとのモーション エンベディングを統合します。広範な実験により、私たちの方法は現実世界の PD 旋回データセットで既存の歩数計数方法よりも一貫して優れていることが示されています。

原文 (English)

Every Step of the Way: Video-based Parkinsonian Turning Step Counting

As a prominent symptom of Parkinson's disease (PD), turning impairment is evaluated through parameters such as turning angle, duration, and particularly, the number of steps required to complete a turn, which directly reflects motor dysfunction. Accurate step counting is challenging due to variability in real-world turning movements and atypical shuffling patterns in parkinsonian gait. Existing methods are predominantly wearable-based, requiring users to wear and manage dedicated devices, which can be inconvenient for continuous daily use. To address this, we propose a passive, video-based framework that estimates step count in a coarse-to-fine manner using diverse motion representations. Specifically, an initial step count is estimated from foot movement signals derived from 3D human mesh recovery, providing high-level motion structures. To incorporate fine-grained motion details, a motion encoder learns complementary gait dynamics from mesh and optical flow to refine the initial estimate. In this process, coarse foot movement signals query the pixel-level motion cues via cross attention to capture subtle parkinsonian gait dynamics. To handle varying video lengths, we partition each video into clips and integrate clip-wise motion embeddings via multiple instance learning (MIL) for step count residual prediction. Extensive experiments show our method consistently outperforms existing step counting methods on real-world PD turning datasets.

13:00 JST画像/動画生成ロボティクス研究/論文

SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects (Early Version)

Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction,…

13:00 JST研究/論文

Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating

Clinical decision support degrades as treatment protocols evolve, but the obstacle to updating a deployed model is governance as much as ac…

13:00 JST画像/動画生成

A Distributional Robustness Margin For Pathology Foundation Models

Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut lear…

13:00 JST画像/動画生成

Anatomy Contextualized Adaptation of CT Foundation Models

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-…

13:00 JST研究/論文

FinVerse: Financial Time-Series Benchmark

As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has b…

13:00 JSTハードウェア/半導体研究/論文

社会性: 人間と AI の相互作用のための関係プロセス フレームワーク

人間と AI の研究では、個々の能力、総合的なパフォーマンス、または最終的な成果物を評価することがよくありますが、これらのアプローチでは、一方の反応がどのようにして他方の次の貢献が形成される条件の一部になるかが保存されません。この記事では、社会的双対性、つまり、区別可能な 2 つの当事者の間の連続的で互恵的で歴史を担う関係プロセスについて紹介します。この関係プロセスでは、一方の当事者からの反応が、他方の当事者のその後の貢献、判断、決定、または行動が形成される観察可能な条件の一部になります。人間と AI のダイアッド向けに指定されたこの構成では、移動、確認された社会的エピソード、リンクされた経路、およびより広範な相互作用コンテナーなどの入れ子になったユニットが使用されます。最小限のエピソード A1-B1-A2 では、対応偶発性と復帰偶発性の証拠が必要です。候補エピソードは、反応の方向性と実質的な貢献の再形成を二次的にコーディングする前に、確認済み、非社会的、または不確定として分類されます。 3 つの命題は、履歴条件の形成、経路の分岐、エンドポイントに相当する経路間の堅牢性の違いに対処します。凍結された運用プロトコルは、別々に実行された 2 つのモデルベースの評価シリーズを通じて、これまで見たことのない 3 つの自然な人間の AI 記録に基づいて調整されました。移動と候補の再構成は 2 つのケースで正確に収束しましたが、3 番目のケースではローカルのマルチモーダル単位化の決定が 1 つ異なりました。残りの意見の相違は、帰還と不測の事態の境界に集中していた。したがって、社会二重性は、エンドポイント中心の分析では回復できない経路情報を保存しながら、人間と AI の貢献が相互作用を通じてどのように形成されるかを分析するための、限定的で経験的に扱いやすいプロセス構造を提供します。

原文 (English)

Socioduality: A Relational Process Framework for Human-AI Interaction

Human-AI research often evaluates individual capabilities, joint performance, or final outputs, but these approaches can lose the interaction process that produced the result. This article introduces socioduality: a sequential, reciprocal, and history-carrying process in which one party's response becomes part of the observable conditions shaping the other party's next contribution, judgement, decision, or action. For human-AI dyads, the framework identifies moves, candidate episodes, confirmed episodes, and maximal pathways. A minimum episode A1 -> B1 -> A2 requires evidence that B1 responds to A1 and that B1 then enters the formation of A2; candidates are classified as confirmed, non-sociodual, or indeterminate. A frozen coding protocol was calibrated on three natural human-AI records using two separate model-based evaluator series. A supplementary exploratory analysis then compared frozen Sociodual pathways with blind developmental/task-process segmentations. Across six examined interactions, the two representations were empirically non-equivalent: task-stage changes could occur within a continuing Sociodual pathway, while formal pathway breaks could occur within a continuing task context. This distinction persisted under fine-grained re-segmentation and record-format checks and was reproduced in all three prospectively selected unseen records using a fresh model-based Sociodual coding line. Socioduality therefore offers a bounded process-level framework for studying how human and AI contributions become relationally linked across time, preserving information that task-stage and endpoint-centred analyses do not uniquely recover.

13:00 JST画像/動画生成

A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems

As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluati…

13:00 JSTLLM/生成AIエージェント

なぜ AI エージェントはルールを破るのか?フレーミング、コンテキスト、ソーシャルシグナルがコンプライアンスをどのように形成するか

罰則を指定すると、逆説的に、法的義務が違反に有利な費用対効果の計算に変換される可能性があります。私たちは、この施行情報のパラドックスが AI エージェントで体系的に発生することを実証します。ほとんどの AI 安全性評価ではモデルが失敗するかどうかがテストされますが、私たちは法律と経済学のコンプライアンス理論を診断ツールとして適用して、その理由を調査します。我々はコンプライアンス理論を比喩としてではなく経験的仮説として扱い、それぞれが異なるモデルクラスの動作を予測することを示します。私たちは、エンタープライズ調達チャットボットとして動作する 12 の命令調整された言語モデルにわたって仮説を評価します。抑止力、正当性、表現法則の理論に基づいて、安全性を細かく調整したモデルは広範なコンプライアンスを維持する一方、タスク最適化モデルやエージェントモデルは規制シグナルを単なる最適化パラメーターとして扱うことを示します。これらの後者のモデルは、低い執行罰や非命令語句など、理論によって予測される条件の下では準拠できません。すべてのモデルにおいて、金銭的インセンティブ、管理上の要求、同僚の成果、または従業員のプレッシャーの導入は、大規模なコンプライアンス違反を引き起こします。 AI 調達エージェントは、標準的な連携ベンチマークでは捉えられない方法で、ローカル ユーザーの目的を満たすために、組織的に規制上の制約に違反しています。結局のところ、ルールの埋め込みだけではコンプライアンスを達成することはできません。モデルの選択自体がガバナンスの決定であり、ベンチマークベースの評価はコンプライアンスを重視した展開には不十分です。

原文 (English)

Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance

Specifying a penalty can turn a legal obligation into a cost-benefit calculation that favors violation. We show that this enforcement information paradox occurs in AI agents. Most AI safety evaluations test whether models fail; we ask why, using compliance theory from law and economics as a diagnostic. We evaluate twelve instruction-tuned language models deployed as enterprise procurement chatbots. Each is given an environmental regulation in its system prompt covering large purchases, and a vendor list on which the certified suppliers cost nearly twice what the uncertified ones do. We test the agents against the predictions of deterrence, legitimacy, and expressive law, and find that each theory accounts for part of what we observe. Under identical conditions, compliance spans 46 percentage points across models, and models differ in which pressure breaks them: some treat the regulation as binding however it is worded, while others fail where theory predicts, under low penalties and non-command phrasing. Benchmark scores and developers' own descriptions of post-training do not predict where a model falls. Across all twelve, financial incentives, managerial demands, peer outcomes, and employee pressure each produce large compliance failures. These agents violate regulatory constraints to satisfy local user objectives in ways standard alignment benchmarks do not measure. Embedding the rule in the system prompt is not on its own enough to produce a compliant agent: model selection is itself a governance decision, and benchmark evaluation is not sufficient for compliance-sensitive deployments.

13:00 JST研究/論文

When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation

\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, u…

13:00 JSTLLM/生成AI

Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong re…

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel…

13:00 JST研究/論文

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-b…

13:00 JST研究/論文

Graphical Design of Interpretable Architectures

Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. The most common representa…

13:00 JSTエージェントロボティクス

DA-WAM: Decision-Aligned Future Latents for Driving World Models

Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of world models for deci…