Skip to the content.

AIニュース 2026-06-20

自動生成: 2026-06-20 13:10 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. The US banned Anthropic’s Fable 5 release, but the numbers don’t seem to careTechCrunch AI

    Just as last week was ending, the US government forced Anthropic to p…

  2. Encryption, spyware, and now Mythos: History shows why cyber export control doesn’t workTechCrunch AI

    For the last 30 years, stopping the flow of cybersecurity-related sof…

  3. Billionaire Ambani wants AI in every call, app, and homeTechCrunch AI

    Reliance is weaving AI into telecom services used by more than 500 mi…

  4. 理研、AI for Science向けスパコンの名前を「理究」(りきゅう)に決定 由来は?ITmedia AI+

    理化学研究所は、AIを活用した科学研究「AI for Science」向けのスーパーコンピュータの名前を「理究」(りきゅう)に決定したと発…

  5. GMO傘下、Unitreeの国内正規代理店に 人型ロボの導入から保守まで一気通貫で支援ITmedia AI+

    GMOインターネットグループ傘下で、ロボティクス事業などを手掛けるGMO AI&ロボティクス商事は、ロボット開発企業の中国Unitree…

  6. 画面操作を“録画”→AIが作業代行 Codexに新機能「Record & Replay」ITmedia AI+

    米OpenAIは6月18日、「Codex」の新機能「Record & Replay」を公開した。利用者がMac上で作業を一度実演すると、C…

  7. The CEO of Allbirds’ new AI biz has a plan, but no teamTechCrunch AI

    Call it a startup with a sole founder and a very large seed round, bu…

トピック別件数

日本語メディア4件

ITmedia AI+ (日本語)

17:00 JST研究/論文

理研、AI for Science向けスパコンの名前を「理究」(りきゅう)に決定 由来は?

理化学研究所は、AIを活用した科学研究「AI for Science」向けのスーパーコンピュータの名前を「理究」(りきゅう)に決定したと発表した。

16:26 JSTその他

大阪メトロは「月1000件の社内問い合わせを効率化」にAIをどう使った?

PKSHA Technologyは、Osaka Metroへの「PKSHA AIヘルプデスク」の導入部署を人事や調達に拡大した。全従業員約5000人の社内問い合わせを一元化し、情報格差の解消や業務効率化、ナレッジの共有資産化を目指す。

16:16 JSTロボティクス

GMO傘下、Unitreeの国内正規代理店に 人型ロボの導入から保守まで一気通貫で支援

GMOインターネットグループ傘下で、ロボティクス事業などを手掛けるGMO AI&ロボティクス商事は、ロボット開発企業の中国Unitree Roboticsと日本国内正規代理店契約を締結したと発表した。

14:15 JSTLLM/生成AIエージェントOpenAI

画面操作を“録画”→AIが作業代行 Codexに新機能「Record & Replay」

米OpenAIは6月18日、「Codex」の新機能「Record & Replay」を公開した。利用者がMac上で作業を一度実演すると、Codexがその操作を再利用できる作業手順に変換・記憶する。

海外メディア4件

TechCrunch AI (英語)

07:40 JSTその他

Encryption, spyware, and now Mythos: History shows why cyber export control doesn’t work

For the last 30 years, stopping the flow of cybersecurity-related software has proven to be ineffective. It's unclear why it would work now…

01:01 JSTLLM/生成AIハードウェア/半導体Anthropic3件の関連記事

The US banned Anthropic’s Fable 5 release, but the numbers don’t seem to care

Just as last week was ending, the US government forced Anthropic to pull its two newest models, Fable 5 and Mythos 5, citing national secur…

出典:TechCrunch AITechCrunch AITechCrunch AI
00:23 JSTその他

Billionaire Ambani wants AI in every call, app, and home

Reliance is weaving AI into telecom services used by more than 500 million people.

22:00 JSTその他

The CEO of Allbirds’ new AI biz has a plan, but no team

Call it a startup with a sole founder and a very large seed round, but what's next is less clear.

公式ブログ0件

このカテゴリの新着記事はありませんでした。

論文312件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント

Agentic AI システムのランタイム ガバナンスのための Deontic ポリシー

Large Language Model (LLM) によって駆動される自律型エージェント AI システムは、新たな種類のセキュリティ、プライバシー、コンプライアンスの課題をもたらします。ツールを呼び出し、データを操作し、ソフトウェアをインストールし、組織の境界を越えてピア エージェントと連携できるエージェントは、認証とアクセス制御だけでなく、エンタープライズ ガバナンスの完全な構造によっても制約される必要があります。これには、エージェントがどのような行為を許可され、どのような行為が禁止されているか、特定のアクションの後に何をする義務があるか(CISO に通知するなど)、どのような条件下で永続的な義務が免除されるか、ポリシーが矛盾する場合にどのルールが優先されるかを指定することが含まれます。このガバナンスの問題は、現在のポリシー エンジンが提供するものを超えています。 XACML、Rego、Cedar などのシステムは、このガバナンス構造の許可/禁止サブセットのみに対応します。これらは、義務のライフサイクル管理、メタポリシーの競合解決、特定の状況で義務を免除する措置、およびヘルスケア、サイバーセキュリティ、データプライバシーなどのアプリケーションで一般的に見られるドメインクラス階層に対する存在論的推論を提供しません。私たちは、基本的な許可/禁止制約だけでなく、義務、調剤、ポリシーの矛盾解決、ポリシーの推論などの主要なガバナンス要件を実現する AgenticRei を提案します。私たちは Rei フレームワーク上に構築された deontic ポリシー言語を使用しており、OWL (Web Ontology Language) として表現され、LLM の完全に外部にある高性能ロジック エンジンによって実行時に評価されます。同じパイプラインが、エージェントによるツールの呼び出しとエージェント間のメッセージの両方を制御します。例を通して、デオンティック ポリシーが、現在の運用エンジンではほとんど表現できないセキュリティとプライバシーに関するガバナンスの制約を捉えていることを示します。私たちのアプローチは、A2AS などの業界標準のフレームワークと自然に組み合わされています。

原文 (English)

Deontic Policies for Runtime Governance of Agentic AI Systems

Autonomous agentic AI systems driven by Large Language Models (LLMs) introduce a new class of security, privacy, and compliance challenges: an agent that can invoke tools, manipulate data, install software, and coordinate with peer agents across organizational boundaries must be constrained not just by authentication and access control, but by the full structure of enterprise governance. This includes specifying what agents are permitted and prohibited from doing, what they areobliged to do after certain actions (e.g., notify the CISO), under what conditions a standing obligation may be waived, and which rules take precedence when policies conflict. This governance problem exceeds what current policy engines provide. Systems such as XACML, Rego, and Cedar address only the permit/prohibit subset of this governance structure. They do not provide obligation lifecycle management, meta-policy conflict resolution, dispensations that waive obligations in specific circumstances, and ontological reasoning over domain class hierarchies commonly found in applications such as healthcare, cybersecurity, or data privacy. We propose AgenticRei, which realizes key governance requirements such as obligations, dispensations, policy conflict resolutions, and reasoning over policies, as well as the basic permit/prohibit constraints. We use a deontic policy language built on the Rei framework, expressed as OWL (Web Ontology Language) and evaluated at runtime by a high-performance logic engine entirely outside the LLM. The same pipeline governs both tool invocations by the agent and agent-to-agent messages. We show through examples that deontic policies capture governance constraints around security and privacy that mostly cannot be expressed in current production engines. Our approach composes naturally with industry-standard frameworks like A2AS.

13:00 JST研究/論文

トピックの範囲、コンピテンシー、認知深度にわたるカリキュラムの整合性の測定: CS2013 と CS2023 に適用される長期的なフレームワーク

学部のコンピュータ サイエンスは、約 10 年に 1 回改訂される国際カリキュラム ガイドラインによって管理されていますが、プログラムには、現在のガイドラインをどの程度完全にカバーしているか、またガイドラインが再構築されたときにそのカバー範囲がどのように変化するかを測定する、信頼性が高く再現可能な方法がありません。私たちは、コンピュータ サイエンス カリキュラム 2013 (CS2013) および 2023 (CS2023) に対して、コンピュータ サイエンスの認定学士 1 名に長期的に適用される外部知識体系のプログラムの範囲を測定する人間参加パイプラインでこれに対処します。パイプラインは、プログラムと各ガイドラインを構造化されたコーパスとして表し、意味検索によってコースと知識単位の一致候補を生成し、明示的なカバレッジ定義に基づいて人間の判断によってそれらを確認します。ベンチマークされた 7 つのレトリーバーのうち、相互ランク融合アンサンブルが最も強く、評判の高いロングコンテキスト モデルは短い文モデルのパフォーマンスを下回っていました。そのため、レトリーバーの選択は評価する必要があります。両方のマップは独立した第 2 評価者によって検証されました (CS2023 のコーエンのカッパ 0.64、CS2013 の 0.69)。このプログラムは、CS2023 の 49.7% と CS2013 の知識単位の 50.9% をカバーしており、これは 10 年間にわたってほぼ一定です。同じ検索してから確認するデザインをコンピテンシーの明確化と認知深度まで拡張すると、プログラムは各ガイドラインで対象となる単元の約 88% についてコンピテンシーを明確にし、さらに CS2013 では 95% であるのに対し、CS2023 では現在の単元の 76% に対して推奨深度でコンピテンシーを提供していることがわかります。このギャップはプログラムではなく、新しいガイドラインで高められた期待を反映しています。縦断的な比較により、ガイドラインと ABET の両方に対して明らかになった永続的な構造的ギャップ (並列コンピューティングと分散コンピューティング、プログラミング言語の基礎、システムの基礎) が、標準の進化を反映する相違点から分離されます。この機器は再利用可能であり、リクエストに応じて著者から入手できます。

原文 (English)

Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023

Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproducible way to measure how completely they cover the current guidelines and how that coverage shifts when the guidelines are restructured. We address this with a human-in-the-loop pipeline that measures a program's coverage of an external body of knowledge, applied longitudinally to one accredited BSc in Computer Science against Computer Science Curricula 2013 (CS2013) and 2023 (CS2023). The pipeline represents the program and each guideline as structured corpora, generates candidate course-to-knowledge-unit matches by semantic retrieval, and confirms them through human judgment under an explicit coverage definition. Of seven benchmarked retrievers, a reciprocal-rank-fusion ensemble was strongest, and a reputed long-context model underperformed a small sentence model, so retriever choice must be measured. Both maps were validated by an independent second rater (Cohen's kappa 0.64 for CS2023, 0.69 for CS2013). The program covers 49.7% of CS2023 and 50.9% of CS2013 knowledge units, near-constant across a decade. Extending the same retrieve-then-confirm design to competency articulation and cognitive depth shows that the program articulates the competency for ~88% of covered units under each guideline, yet delivers it at the recommended depth for 76% of present units under CS2023 against 95% under CS2013, a gap reflecting the newer guideline's raised expectations, not the program. The longitudinal comparison separates persistent structural gaps (parallel and distributed computing, foundations of programming languages, systems fundamentals), uncovered against both guidelines and ABET, from differences that reflect the standard's evolution. The instrument is reusable and available from the authors on request.

13:00 JSTLLM/生成AI

拡散言語モデル: 実験的分析

大規模言語モデル (LLM) は、自己回帰生成を通じて言語モデリングに革命をもたらし、幅広いタスクにわたって強力なパフォーマンスを可能にします。最近、拡散言語モデル (DLM) が、次のトークンの予測ではなく反復的なノイズ除去を通じてテキストを生成する代替パラダイムとして登場し、シーケンス全体の並行改良を可能にします。多数の拡散ベースのアーキテクチャが提案されていますが、評価プロトコル、データセット、推論バジェット、生成ハイパーパラメータの違いにより、それらの機能を比較し、それらがもたらすトレードオフを理解することが困難になっています。この研究では、最新の DLM の系統的な実験分析を紹介します。具体的には、生成品質と計算効率の両方を明示的に考慮しながら、推論、コーディング、翻訳、知識、構造化された問題解決にわたる 8 つのベンチマークにわたって 8 つの最先端 DLM を評価します。下流の評価を超えて、ノイズ除去ステップ、コンテキストの長さ、ブロック サイズ、並列アンマスク戦略などの主要な推論時間要因の影響を分析し、同一条件下でトレーニングされた小規模なモデルの制御された比較で大規模な実験を補完します。私たちの分析では、さまざまなタスク、アーキテクチャ、推論予算にわたる拡散ベースの言語モデリングの長所と限界が浮き彫りになっています。 DLM の動作は生成時の設計選択によって強く影響され、パフォーマンスと計算効率の間に明確なトレードオフが生じることを示します。全体として、私たちの研究は、現代の DLM の機能と展開の特徴についての実用的な洞察を提供します。

原文 (English)

Diffusion Language Models: An Experimental Analysis

Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range of tasks. Recently, Diffusion Language Models (DLMs) have emerged as an alternative paradigm that generates text through iterative denoising rather than next-token prediction, allowing parallel refinement of entire sequences. While numerous diffusion-based architectures have been proposed, differences in evaluation protocols, datasets, inference budgets, and generation hyperparameters make it difficult to compare their capabilities and understand the trade-offs they offer. In this work, we present a systematic experimental analysis of modern DLMs. Specifically, we evaluate eight state-of-the-art DLMs across eight benchmarks spanning reasoning, coding, translation, knowledge, and structured problem solving, while explicitly considering both generation quality and computational efficiency. Beyond downstream evaluation, we analyze the impact of key inference-time factors, including denoising steps, context length, block size, and parallel unmasking strategies, and complement large-scale experiments with controlled comparisons of smaller models trained under identical conditions. Our analysis highlights the strengths and limitations of diffusion-based language modeling across different tasks, architectures, and inference budgets. We show that the behavior of DLMs is strongly influenced by generation-time design choices, leading to distinct trade-offs between performance and computational efficiency. Overall, our study provides practical insights into the capabilities and deployment characteristics of contemporary DLMs.

13:00 JSTLLM/生成AIエージェント

マルチエージェント LLM の審議における隠れたアンカー

複数のラウンドにわたってエージェントが回答を交換および修正するマルチエージェント LLM 熟議は、推論と精度を向上させるためにますます使用されていますが、それがどのように、そしてなぜ機能するのかモデル化されることはほとんどありません。このような熟慮は、人間がどのように決定に至るかを反映しています。社会的動物として、私たちは、デグルートやフリードキン・ジョンセンのような古典的な意見力学モデルが捉えている集団効果である集団と、彼らには理解されていない私たち自身の内なる信念の両方に引っ張られています。私たちは、マルチエージェントの熟議を閉ループの力学システムとしてモデル化します。このシステムでは、各エージェントが、近隣のエージェントに関係なく常に自分の意見を引き出す、隠れた内なる信念、つまりそのアンカーを担っています。我々は、このアンカーが熟議のみから回復できること、そしてそれが古典的なコンセンサスのルールが禁じている行動を説明していることを示します。つまり、正解に対するエージェントの信頼は、エージェントが開始した時点を超えて、最初の信念によって形成された空間(凸包)から逃れることができます。回復されたアンカーがホールドアウト実行を予測する (一般化する) かどうかを確認することで、モデルがそのようなアンカーによって実際に駆動されるときの簡単なテストが得られます。 3 つのオープンウェイト モデル ファミリーにわたって、これはスペクトルであり、すべてかゼロかではありません。すべてのアンカーの影響力はほぼ同等に強いですが、アンカーがどこに位置するかが異なり、最初の意見から遠く離れた位置にある場合にのみ、検討が船体から逃れ、完全な閉ループ モデルが必要になります。

原文 (English)

Hidden Anchors in Multi-Agent LLM Deliberation

Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increasingly used to improve reasoning and accuracy, yet how and why it works is rarely modelled. Such deliberation mirrors how humans reach decisions. As social animals we are pulled both by the group, the herd effect that classical opinion-dynamics models such as DeGroot and Friedkin--Johnsen capture, and by our own internal belief, which they do not. We model multi-agent deliberation as a closed-loop dynamical system in which each agent carries a hidden internal belief, its anchor, that continually pulls its opinion regardless of its neighbours. We show this anchor can be recovered from the deliberation alone, and that it explains a behaviour classical consensus rules forbid: an agent's confidence in the correct answer can climb past where any agent started, escaping the space (convexhull) formed by the initial beliefs. Checking whether the recovered anchor also predicts held-out runs (generalizes) gives a simple test for when a model is truly driven bysuch an anchor. Across three open-weight model families this is a spectrum, not all-or-nothing. All anchors' influence are about equally strongly, but they differ in where the anchor sits, and only when it sits far from the initial opinions does deliberation escape the hull and need the full closed-loop model.

13:00 JSTLLM/生成AIエージェント

DeXposure-Claw: DeFi リスク監視のためのエージェント システム

分散型金融により、監督当局は急速に変化するネットワーク化された信用リスクにさらされます。汎用 LLM エージェントはこの設定にあまり適合しません。弱い証拠を深読みし、一か八かの介入を推奨しますが、既存の評価では、結果として生じる誤報を測定するための規制当局と連携した方法が提供されていません。我々は、構造化された証拠を通じて LLM 決定をルーティングする、予測に基づいたエージェント監視システムである DeXposure-Claw を紹介します。(1) DeXposure-FM は、グラフ時系列基礎モデルであり、将来のエクスポージャ ネットワークを予測します。 (2) 決定論的なモニターとストレス シナリオは、それらの予測を型指定されたアラート、属性シグナル、およびシナリオの証拠に変換します。 (3) データの健全性と信頼ゲートにより、DeXposure-Claw が根拠のある監査可能な監督チケットを発行する前にエスカレーションが抑制されます。さらに、6 軸の評価ハーネスである DeXposure-Bench を開発します。このベンチの決定軸は、規制当局と調整された絶対損失グラウンド トゥルースおよび明示的な誤介入率に対してチケットをスコア付けします。 5 年間の毎週の実データを用いた実験により、当社のシステムが完全にサポートされています。コードは https://github.com/EVIEHub/DeXposure-Claw にあります。

原文 (English)

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no regulator-aligned way to measure the resulting false alarms. We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that routes LLM decisions through structured evidence: (1) DeXposure-FM, a graph time-series foundation model, forecasts future exposure networks; (2) deterministic monitors and stress scenarios then turn those forecasts into typed alerts, attribution signals, and scenario evidence; and (3) data-health and confidence gates constrain escalation before DeXposure-Claw emits auditable supervisory tickets with rationales. We further develop DeXposure-Bench, a six-axis evaluation harness, whose decision axis scores tickets against a regulator-aligned absolute-loss ground truth and an explicit false-intervention rate. Experiments on five years of weekly real data fully support our system. Code is at https://github.com/EVIEHub/DeXposure-Claw.

13:00 JSTLLM/生成AIQwen

LLM は何を知らないのかを知らない: 臨床表データのモデル間の帰属相違による認識上の盲点の検出

大規模言語モデル (LLM) は、構造化された臨床データにますます適用されていますが、そのようなタスクに関する自身の知識の限界を認識できるかどうかはまだ解明されていません。私たちは、構造化タスクの認識論的不確実性を軽減することを目的として、クロスモデルの属性発散のレンズを通してこの問題を研究し、属性発散分析による予測タスクで Qwen 2.5 7B と XGBoost を比較します。 4 つの調査結果を報告します。まず、LLM 言語化された信頼度は認識論的に空虚であり、精度が 49% であるか 75.3% であるかに関係なく、ほぼ一定 (0.856 ~ 0.937) を出力し、予測品質ではなくプロンプト形式を追跡します。第 2 に、LLM は逆の難易度効果を示します。XGBoost が 99% 正しい場合、精度は 64.8% に低下しますが、中程度の不確実性がある場合は XGBoost と一致します (73.8% 対 73.1%)。 3 番目に、少数ショットの例と SHAP 由来の特徴証拠は、直交する超相加的介入です。トレーニングなしで、帰属不一致スコア (ADS) が 1.54 から 0.38 に減少し、精度が 49% から 75.3% に向上します。 4 番目に、帰属乖離信号を使用して LLM の信頼性を決定するクロスモデル キャリブレーターは、モデルの内部にアクセスしたり推論を繰り返す必要がなく、情報のない言語化された信頼性を患者固有の信頼性推定値に置き換えて、予想されるキャリブレーション誤差を 0.254 から 0.080 に削減します。私たちはこれらの発見を構造化データ上の LLM のコールド スタート問題として枠組み化し、真の認識論的自己認識への道筋を概説します。

原文 (English)

LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data

Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they can recognize the limits of their own knowledge on such tasks remains unexplored. We study this question through the lens of cross-model attribution divergence with the goal of reducing epistemic uncertainty for structured tasks, comparing Qwen 2.5 7B and XGBoost on a prediction task via attribution divergence analysis. We report four findings. First, LLM verbalized confidence is epistemically vacuous, it outputs a near-constant (0.856-0.937) regardless of whether accuracy is 49% or 75.3%, tracking prompt format rather than prediction quality. Second, the LLM exhibits an inverse difficulty effect: accuracy drops to 64.8% when XGBoost is 99% correct, but matches XGBoost (73.8% vs. 73.1%) when it is moderately uncertain. Third, few-shot examples and SHAP-derived feature evidence are orthogonal, super-additive interventions: they reduce the Attribution Disagreement Score (ADS) from 1.54 to 0.38 and improve accuracy from 49% to 75.3% without training. Fourth, a cross-model calibrator that determined LLM reliability using attribution divergence signals reduces expected calibration error from 0.254 to 0.080, replacing uninformative verbalized confidence with patient-specific reliability estimates, without accessing model internals or requiring repeated inference. We frame these findings as a cold start problem for LLMs on structured data and outline a path toward genuine epistemic self-awareness.

13:00 JST研究/論文

REVEAL++: アルツハイマー病リスクの視覚言語網膜モデリングのための識別可能な表現型グループ化

網膜は神経変性疾患への非侵襲的な窓を提供し、将来の認知機能低下のリスクに関連する微妙な構造パターンを捕捉します。 REVEAL などの視覚言語調整フレームワークは、網膜眼底画像と構造化された臨床リスクのナラティブを組み合わせることで、アルツハイマー病 (AD) の早期予測が向上することを示しています。これらのアプローチにおける重要な設計上の選択は、表現型グループ化の使用であり、類似したリスク プロファイルを持つ個人が、対比学習中に多重陽性ペアとして扱われます。しかし、既存の方法は、表現型の類似性を離散的な構成要素として操作し、厳密な監視を課し、グループ形成を表現学習から分離するハードグループ割り当てに依存しています。我々は、対照学習内で表現型構造を継続的に定式化することを提案します。サンプルを固定クラスターに割り当てるのではなく、網膜画像とリスクプロファイルの両方に埋め込まれたモダリティ内の類似性から導出された微分可能な重み付け関数として被験者間の類似性をモデル化します。これらの重みは、継続的な集計演算子を通じてソフトなマルチポジティブ関係を定義し、疾患リスクのスペクトルの性質を反映した段階的な監視を可能にします。さらに、クロスモーダルアライメントと表現型構造をエンドツーエンドで共同学習するソフトターゲット対比目標を導入します。アルツハイマー病の予測について英国バイオバンクの網膜画像データに基づいて評価したところ、提案されたフレームワークは、離散グループベースの対比学習および標準的な視覚言語ベースラインを一貫して上回っています。表現型の類似性を、固定されたグループ化ルールではなく、学習可能な連続信号として扱うことにより、私たちのアプローチは、マルチモーダルな網膜および臨床データから集団スケールの神経変性リスクモデリングのための原則に基づいた堅牢な基盤を提供します。

原文 (English)

REVEAL++: Differentiable Phenotypic Grouping for Vision-Language Retinal Modeling of Alzheimer's Disease Risk

The retina offers a noninvasive window into neurodegenerative disease, capturing subtle structural patterns associated with a risk of future cognitive decline. Vision-language alignment frameworks such as REVEAL have shown that pairing retinal fundus images with structured clinical risk narratives improves early prediction of Alzheimer's disease (AD). A key design choice in these approaches is the use of phenotypic grouping, where individuals with similar risk profiles are treated as multi-positive pairs during contrastive learning. However, existing methods operationalize phenotypic similarity as a discrete construct, relying on hard group assignments that impose rigid supervision and decouple group formation from representation learning. We propose a continuous formulation of phenotypic structure within contrastive learning. Rather than assigning samples to fixed clusters, we model inter-subject similarity as a differentiable weighting function derived from intra-modality embedding similarities in both retinal images and risk profiles. These weights define soft multi-positive relationships through a continuous aggregation operator, enabling graded supervision that reflects the spectrum nature of disease risk. We further introduce a soft-target contrastive objective that jointly learns cross-modal alignment and phenotypic structure in an end-to-end manner. Evaluated on UK Biobank retinal imaging data for incident AD prediction, the proposed framework consistently outperforms discrete group-based contrastive learning and standard vision-language baselines. By treating phenotypic similarity as a learnable, continuous signal rather than a fixed grouping rule, our approach provides a principled and robust foundation for population-scale neurodegenerative risk modeling from multi-modal retinal and clinical data.

13:00 JSTLLM/生成AIハードウェア/半導体

緊急調整

大規模言語モデル (LLM) は、自身の出力が人間の倫理と乖離していることを識別できますか?そして彼らは自己修正できるのでしょうか? LLM に、自身の推論と出力をレビューする良心ステップを与え、直接優先最適化 (DPO) を使用して調整コンポーネントでトレーニング損失を拡張し、モデルを非倫理的な出力から遠ざけます。その結果、トレーニング、微調整、敵対的プロンプト、ゼロショット学習など、幅広いアプリケーションでモデルを調整するオンライン技術が実現しました。それは、より弱いまたはより強いジャッジを必要とせず、代わりにそれ自体の凍結されたコピーに依存します。以前の研究では、緊急不整合シナリオでは、モデルの微調整からコードのハッキングに至るまで、さまざまな緊急の非倫理的な行為が示されました。代わりに、私たちは緊急調整を達成する方法を経験的に示します。つまり、単一の高レベルの内省的な質問が、同じコード ハッキング シナリオの下で倫理モデルに向けてトレーニングを導きます。

原文 (English)

Emergent Alignment

Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics? And can they self-correct? We endow an LLM with a conscience step that reviews its own reasoning and outputs, and we extend the training loss with an alignment component using Direct Preference Optimization (DPO) to steer the model away from non-ethical outputs. The result is an online technique to align models in a wide range of applications: training, fine-tuning, adversarial prompting, and zero-shot learning. It does not require a weaker or stronger judge, relying instead on a frozen copy of itself. In previous work, the Emergent Misalignment scenario showed a range of emergent unethical behaviors from fine-tuning the model to hack code. Instead, we empirically show how to achieve Emergent Alignment: a single high-level introspective question steers training toward an ethical model under the same code hacking scenario.

13:00 JST研究/論文

ITNet: 畳み込み、注意、再帰を包含する学習可能な積分変換

畳み込みネットワーク、リカレント ネットワーク、トランスフォーマーはそれぞれ、局所性、逐次記憶、内容依存のペアワイズ相互作用など、さまざまな誘導バイアスをエンコードしており、その誕生以来数学的に区別され続けています。我々は、この断片化が信号の処理方法における基本的な多様性を反映しているのではなく、むしろ基礎となる単一の数学的オブジェクト、つまり学習可能な積分変換の不完全なビューを反映していることを示します。位置と特徴に共同して依存する学習可能なカーネルを中心に構築された統合アーキテクチャである Integral Transform Network (ITNet) を紹介します。このカーネルは、ペアごとの相互作用をモデル化する小さなニューラル ネットワーク、特に MLP として実装され、モデルがデータからその動作を適応できるようにします。畳み込み、自己注意 (マルチヘッドを含む)、および自己回帰再帰 (LSTM、GRU、S4、および Mamba を含む) が適切なパラメータ化の下で特殊なケースとして発生すること、および ITNet が連続演算子の汎用近似器であることを示します。これを実用化するために、タイル化カーネル融合、重要度加重モンテカルロ統合、学習された低ランク因数分解を開発し、効率的でスケーラブルな計算を可能にします。共有オペレーターと軽量のモダリティ固有のエンコーダーを備えた単一の ITNet アーキテクチャは、ImageNet-1K、GLUE、ModelNet40、VQA\,v2、および NLVR2 の特殊なベースラインと一致またはそれを超えています。この結果は、単一の学習された対話メカニズムが 3 つのアーキテクチャ ファミリすべての動作をデータから復元できることを示しています。

原文 (English)

ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence

Convolutional networks, recurrent networks, and transformers each encode different inductive biases -- locality, sequential memory, and content-dependent pairwise interaction -- and have remained mathematically distinct since their inception. We show that this fragmentation reflects not a fundamental diversity in how signals should be processed, but rather incomplete views of a single underlying mathematical object: a learnable integral transform. We introduce the Integral Transform Network (ITNet), a unified architecture built around a learnable kernel that depends jointly on positions and features. This kernel is implemented as a small neural network, specifically an MLP, that models pairwise interactions, enabling the model to adapt its behavior from data. We show that convolution, self-attention (including multi-head), and autoregressive recurrence (including LSTM, GRU, S4, and Mamba) arise as special cases under appropriate parameterizations, and that ITNet is a universal approximator of continuous operators. To make this practical, we develop tiled kernel fusion, importance-weighted Monte Carlo integration, and learned low-rank factorization, enabling efficient and scalable computation. A single ITNet architecture with a shared operator and lightweight modality-specific encoders matches or exceeds specialized baselines on ImageNet-1K , GLUE, ModelNet40, VQA\,v2 and NLVR2. The results demonstrate that a single learned interaction mechanism can recover the behavior of all three architectural families from data.

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPTDeepSeek

LLM エージェントにおける明確化探索のための不確実性分解

最近の意見書では、対話型大言語モデル (LLM) エージェントには古典的な偶発的/認識論的不確実性フレームワークでは不十分であると主張し、プロアクティブな説明の探索や共有メンタル モデルの構築などの新しいエージェントの能力を解放できる、過小仕様を認識し、分解され、伝達可能な不確実性表現が求められています。ブラックボックス API、インタラクティブなレイテンシ バジェット、ラベル付き軌跡の欠如といった実際的なデプロイメントの制約により、logprob ベース、マルチサンプリング、トレーニング ベースの手法が除外され、プロンプトベースの推定がデプロイメント時にそのような信号を表面化するための最も実行可能なファミリーとして残ります。この呼び出しには、アクションの信頼性とリクエストの不確実性 (u) を分離する単純なプロンプトベースの分解で応答します。これにより、タスクの仕様があいまいな場合にエージェントが説明を求めることができます。これを評価するために、タスクの 50% が意図的に過少指定されている 2 つの明確化強化ベンチマーク (WebShop-Clarification および ALFWorld-Clarification) を導入し、5 つの LLM バックボーン (GPT-5.1、DeepSeek-v3.2-exp、GLM-4.7、 Qwen3.5-35B、GPT-OSS-120B) をこれらのバリアントと、障害検出のための標準の WebShop、ALFWorld、および REAL ベンチマークとともに使用します。 5 つのバックボーン全体で平均すると、提案された分解は、ALFWorld-Clarification のクラリフィケーション F1 を ReAct+UE と比較して 73%、UAM と比較して 36% 改善し、WebShop-Clarification のすべてのバックボーンと ALFWorld-Clarification の 5 つのバックボーンのうち 4 つでクラリフィケーション F1 をリードしており、ゲインが単一の LLM を超えて一般化していることを示しています。

原文 (English)

Uncertainty Decomposition for Clarification Seeking in LLM Agents

Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) agents and call for underspecification-aware, decomposed, and communicable uncertainty representations that can unlock new agent capabilities such as proactive clarification seeking and shared mental-model building. Practical deployment constraints -- black-box APIs, interactive latency budgets, and the absence of labeled trajectories -- rule out logprob-based, multi-sampling, and training-based methods, leaving prompt-based estimation as the most viable family for surfacing such signals at deployment time. We answer this call with a simple prompt-based decomposition that separates action confidence from request uncertainty (u), enabling the agent to ask for clarification when the task specification is ambiguous. To evaluate it, we introduce two clarification-augmented benchmarks (WebShop-Clarification and ALFWorld-Clarification) in which 50% of tasks are deliberately underspecified, and systematically compare the proposed decomposition against ReAct+UE and Uncertainty-Aware Memory (UAM) across five LLM backbones (GPT-5.1, DeepSeek-v3.2-exp, GLM-4.7, Qwen3.5-35B, GPT-OSS-120B) on these variants together with the standard WebShop, ALFWorld, and REAL benchmarks for fault detection. Averaged across the five backbones, the proposed decomposition improves clarification F1 on ALFWorld-Clarification by 73% over ReAct+UE and by 36% over UAM, and leads clarification F1 on every backbone on WebShop-Clarification and on four of five backbones on ALFWorld-Clarification, indicating that the gains generalize beyond a single LLM.

13:00 JSTLLM/生成AI

LLM ソルバー ループにおけるナレーション ギャップの分析

安全性やセキュリティに関する重要な問題をロジックで定式化できる場合、SAT ソルバーや SMT ソルバーなどの形式的なツールが言語モデル推論パイプラインに組み込まれることが増えています。正式な保証なしにモデル分布からステップがサンプリングされる思考連鎖とは異なり、ソルバーは健全で独立して検証可能な答えを生成します。ただし、ソルバーとモデルの間の相互作用によって健全性の保証が失われる可能性があります。ハイブリッド パイプラインには、質問の形式化、決定、結果の説明という 3 つのコンポーネントがあります。これまでの研究では、形式化と決定については研究されてきましたが、形式的なツールの出力をユーザーの回答に変えるステップであるナレーションについては研究されていませんでした。ナレーションのギャップを埋めるために、まず LLM ソルバー ループを検証済みの意思決定手順としてモデル化します。さらに、プロンプト インジェクションの下で 5 つのオープンソース モデルを評価したところ、証明書ゲーティングによってソルバーの判定が確実なものになる一方、敵対者がフレージングやチャネル全体で検証済みの結論を覆す可能性があることがわかりました。私たちは、インジェクションを大幅に削減するものの、インジェクションを排除することはできず、依然として適応型攻撃を受ける、強化されたプロンプトによる緩和策を研究しています。形式的な分析と実証的研究を組み合わせると、LLM ソルバー ループでのロバスト性は、ユーザーが最終的に読み取る回答に到達しないことがわかります。

原文 (English)

Analyzing the Narration Gap in LLM-Solver Loops

Formal tools such as SAT and SMT solvers are increasingly embedded in language model reasoning pipelines when a safety or security critical question can be formulated in logic. Unlike chain of thought whose steps are sampled from the model distribution without formal guarantee, a solver produces a sound and independently verifiable answer. However, the soundness guarantee can be lost in the interaction between the solver and the model. The hybrid pipeline has three components: formalizing the question, deciding it, and narrating the result. Prior work has studied the formalization and decision, but not narration, which is the step that turns a formal tool's output into the user answer. To fill the narration gap, we first model the LLM-solver loop as a verified decision procedure. We further evaluate five open-sourced models under prompt injection, and we find certificate gating makes the solver verdict sound, while an adversary can invert a verified conclusion across phrasings and channels. We study the mitigation through hardened prompt that reduces injection significantly but cannot eliminate it and still suffers under adaptive attack. Combining the formal analysis and empirical studies, we show in the LLM-solver loop, robustness does not reach to the answer that the user finally reads.

13:00 JSTエージェント

Agentic RAG による構成可能な臨床情報抽出: 機能するもの、機能しないもの、およびその理由

患者のコンテキストは数百の異種ドキュメントと数千の構造化データポイントにまたがっていますが、AI システムが検索やトリアージに必要とするドキュメントレベルのメタデータが存在しないか不完全です。標準的な検索拡張生成は、このデータでは失敗し、一時的な推論、ドキュメント間の依存関係、およびメタデータの欠落の処理が誤ります。私たちは、エッセン医科大学に ACIE (薬剤臨床情報抽出) を導入しています。これは、完全な患者コンテキストを推論し、臨床医の検証のためにソースパッセージ内のすべての回答を根拠にするオンプレミスの薬剤 RAG パイプラインです。私たちはメタデータのギャップを定量化し、それが形作ったアーキテクチャ上の決定を追跡し、独立した遡及的リンパ腫レジストリ研究と並行して抽出を評価します。この研究では、核医学の医師が抽出されたすべての値を引用ソースと照合して検証します。 7,326 件の判定において、臨床医は 96.5% の抽出を受け入れ、タイプごとの受け入れ率は 80% から 99% の範囲でした。

原文 (English)

Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why

Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-level metadata that AI systems need for retrieval and triage is absent or incomplete. Standard retrieval-augmented generation fails on this data, mishandling temporal reasoning, cross-document dependencies, and missing metadata. We deploy ACIE (Agentic Clinical Information Extraction) at University Medicine Essen: an on-premise agentic RAG pipeline that reasons over complete patient contexts and grounds every answer in source passages for clinician verification. We quantify the metadata gap, trace the architectural decisions it shaped, and evaluate extraction alongside an independent retrospective lymphoma registry study, in which nuclear-medicine physicians verify every extracted value against its cited sources. Across 7,326 judgments, clinicians accepted 96.5\% of extractions, with per-type acceptance ranging from 80\% to 99\%.

13:00 JSTLLM/生成AI

LLM トレーニング後の比較対象となるペアはどれですか?

好みに基づいたポストトレーニングは、言語モデルを調整するための中心的なパラダイムとなっています。一般的なデータ収集戦略は、プロンプトごとに小さな補完セットを生成し、結果の比較ペアにラベルを付けることです。ただし、人間の好みのラベルは、追加の補完を生成するよりもはるかにコストがかかることが多いため、同じラベル付け予算の別の使用法、つまり、より大きな補完プールを生成し、最も有益な比較ペアのみにラベルを付けることを示唆しています。この論文では、好みに基づいたポストトレーニングでどのペアを比較する必要があるかを研究します。私たちは比較キュレーションをサンプリング設計問題として定式化し、好みに基づいたトレーニング後の目標に基づいて最終ポリシーの品質によって設計を評価します。このフレームワークを Direct Preference Optimization (DPO) 用にインスタンス化し、ラベル付きペアの選択が DPO トレーニングを通じて下流のポリシーのパフォーマンスにどのように伝播するかを分析します。私たちの主な結果は、DPO でトレーニングされたポリシーのトレーニング後の最適性ギャップの上限と下限の一致を示します。境界は、比較選択が、ラベル割り当てをパラメーター推定誤差およびポリシーの準最適性に結び付ける、単一の設計依存情報マトリックスを通じて下流のパフォーマンスに影響を与えることを示しています。これにより、予算に基づいた比較キュレーションのための明示的な最適化基準が得られ、生成された大規模な補完プールから有益なペアを選択するための実際的なサンプリング設計が動機付けられます。合成設定と言語モデルのトレーニング後のベンチマークに関する実験では、提案された設計が一般的な比較選択ヒューリスティックよりもサンプル効率を一貫して向上させることが示されています。

原文 (English)

Which Pairs to Compare for LLM Post-Training?

Preference-based post-training has become a central paradigm for aligning language models. A common data-collection strategy is to generate a small set of completions for each prompt and label the resulting comparison pairs. However, human preference labels are often much more expensive than generating additional completions, suggesting a different use of the same labeling budget: generate a larger pool of completions, but label only the most informative comparison pairs. This paper studies which pairs should be compared in preference-based post-training. We formulate comparison curation as a sampling-design problem and evaluate designs by the quality of the final policy under the preference-based post-training objective. We instantiate this framework for Direct Preference Optimization (DPO), analyzing how the choice of labeled pairs propagates through DPO training to downstream policy performance. Our main results provide matching upper and lower bounds on the post-training optimality gap of the DPO-trained policy. The bounds show that comparison selection affects downstream performance through a single design-dependent information matrix, which links label allocation to parameter estimation error and policy suboptimality. This yields an explicit optimization criterion for budgeted comparison curation and motivates practical sampling designs for selecting informative pairs from large generated completion pools. Experiments on synthetic settings and language-model post-training benchmarks show that the proposed designs consistently improve sample efficiency over common comparison-selection heuristics.

13:00 JSTLLM/生成AI

Toten: ブラジルポルトガル語での物理量と技術表記の知識ベースの存在論的トークン化

バイト ペア エンコーディングのトークン化は、語彙圧縮に関しては統計的に効率的ですが、意味的には構造化された技術的エンティティに対して盲目であり、物理量、数値、単位、記号表現を語彙的に任意のサブワードに断片化します。我々は、統計的導出をエンジニアリングエンティティ(OEE)の正式なオントロジーに基づいた宣言的分類に置き換える、知識ベースのオントロジートークン化フレームワークであるTOTENを紹介します。私たちは TOTEN をトリプルとして形式化します。オントロジーは型、構造原理、構成関係、保存可能な不変条件を収集します。分類関数は、生のテキストを入力された領域にマッピングします。そして、インスタンシエータ ファミリは自己記述的な構造化表現を生成します。堅牢性は、Pint (次元)、Unicode 文字データベース (タイポグラフィ)、および RSLP (ポルトガル語形態学) という 3 つの外部オラクルとの決定論的な結合から得られます。本質的評価は、物理的に検証された内部ベンチマーク (EngQuant、N=800) および 4 つのブラジルポルトガル語外部コーパス (N=1771 の適格なケース) を対象として、構築によって検証可能な 4 つのプロパティ (存在論的原子性、次元の等価性、タイポグラフィーの堅牢性、および数値再構成) をカバーします。また、検出リコールも報告し、カバレッジと条件付きアトミック性を区別します。 8 つの最先端のベースラインに対して、TOTEN は、すべてのコントラストおよび数値再構成において、外部コーパスでは 0.775 ~ 0.904 の単位オントロジー アトミック性を達成します。これに対し、最良のベースライン (Quantulum3) では 0.627 ~ 0.703 でした。 EngQuant では、0.780 対 0.340。差異は統計的に有意です (ホルム補正を伴うマクネマー)。内部ランキングと外部ランキングの間のスピアマン相関により、コントロール ベンチマークの同時有効性が確認されます。次元の同等性は、システムが次元の権限を継承する神託である Pint と統計的に同等であることを示します。

原文 (English)

Toten: Knowledge-Based Ontological Tokenization Of Physical Quantities And Technical Notation In Brazilian Portuguese

Byte-Pair Encoding tokenization is statistically efficient for vocabulary compression, but semantically blind to structured technical entities, fragmenting physical quantities, numbers, units, and symbolic expressions into lexically arbitrary subwords. We present TOTEN, a knowledge-based ontological tokenization framework that replaces statistical derivation with declarative classification grounded in a formal ontology of engineering entities (OEE). We formalize TOTEN as the triple : the ontology gathers types, structural principles, composition relations, and preservable invariants; the classification function maps raw text into typed regions; and the instantiator family yields a self-descriptive structured representation. Robustness derives from deterministic coupling with three external oracles: Pint (dimensional), Unicode Character Database (typographic), and RSLP (Portuguese morphology). Intrinsic evaluation covers four properties verifiable by construction -- ontological atomicity, dimensional equivalence, typographic robustness, and numerical reconstruction -- over an internal, physically validated benchmark (EngQuant, N=800) and four Brazilian Portuguese external corpora (N=1771 eligible cases). We also report detection recall, distinguishing coverage from conditional atomicity. Against eight state-of-the-art baselines, TOTEN achieves unit ontological atomicity in all contrasts and numerical reconstruction of 0.775-0.904 on external corpora, vs. 0.627-0.703 for the best baseline (Quantulum3); on EngQuant, 0.780 vs. 0.340. Differences are statistically significant (McNemar with Holm correction). Spearman correlation between internal and external rankings confirms concurrent validity of the control benchmark. Dimensional equivalence shows statistical parity with Pint, the oracle from which the system inherits dimensional authority.

13:00 JST研究/論文

AI4SE と SE4AI の探求: 過去と未来を振り返る 10 年

AI とシステム エンジニアリング (SE) に関する 2020 年 3 月の INCOSE INSIGHT 特集号は、同誌史上最もダウンロードされた号となり、研究コミュニティを立ち上げ、現在では年次ワークショップに 250 人を超える登録者が集まっています。この記事では、著者がこの分野の核となる論文を読んだことに基づいて、AI と SE の 3 つのフェーズ (ここでは基礎、応用、LLM の活用とラベル付けされています) にわたる進歩を追跡し、コミュニティがどこに収束し、どこに重大なギャップが残っているかについての意見を説明します。これとは別に、人間の専門知識と 6 つの AI モデルの両方を活用した人間と AI の合意文献レビューが実行され、1,712 件の INCOSE INSIGHT 記事と 889 件の SERC 出版物の関連性が評価されました。この結果は、研究における 5 つの重要なギャップを特定し、SE における AI の導入、保証、労働力の変革を進める実務者に指針を提供します。私たちは契約データと AI4SE/SE4AI Explorer Web アプリケーションを共有するので、読者は自分自身の関連性判断を人間および AI の評価者と比較できます。

原文 (English)

AI4SE and SE4AI Exploration: A Decade Looking Back and Forward

The March 2020 INCOSE INSIGHT special issue on AI and Systems Engineering (SE) became the most downloaded issue in the publication's history and launched a research community that now draws over 250 registrants to its annual workshop. In this article, we trace the progress in AI and SE across three phases (labeled here foundational, applied, and LLM inflection) based on the authors' reading of the field's core papers, and describe our opinions of where the community has converged and where critical gaps remain. Separately, a human-AI agreement literature review leveraging both human expertise and six AI models was performed to assess the relevance of 1,712 INCOSE INSIGHT articles and 889 SERC publications. The results identify five critical research gaps and offer guidance for practitioners navigating AI adoption, assurance, and workforce transformation in SE. We share the agreement data and the AI4SE/SE4AI Explorer web application so readers can compare their own relevance judgments with the human and AI raters.

13:00 JST画像/動画生成

BrainG3N: 制御可能な 3D 脳 MRI 生成のための多目的トークナイザー

三次元 (3D) 脳 MRI は臨床神経学および神経腫瘍学の中心であり、生成モデルは過小評価されているコホートを強化し、疾患の軌跡をシミュレートし、プライバシーを保護するデータ共有をサポートできます。潜在拡散は画像データをモデリングするための頼りになるソリューションですが、トークナイザーには 2 つの競合する要求が課せられます。エンコーダーの埋め込みは、下流のタスクが作用する臨床情報を保持する必要があり、デコーダーは解剖学的に忠実なボリュームを再構成する必要があります。既存の再構築駆動トークナイザーは、最初のトークナイザーを犠牲にして 2 番目のトークナイザーを実現します。これに対処するために、3D 脳 MRI 潜在拡散、デカップリング エンコーダーおよびデコーダー用の完全ボリューム マスク オートエンコーダー (MAE) ベースのトークナイザーを導入します。凍結された 3D MAE エンコーダーは臨床的に有益な埋め込みを生成し、専用の CNN デコーダーはそれらの埋め込みの線形投影からボクセルを再構築します。私たちは、4 つのモダリティ、10 の疾患カテゴリ、200 以上の取得サイトにわたる 18 の公的コホートからの 35,309 ボリュームでエンコーダーを事前トレーニングし、2 つの設定でその二重の有用性を実証します。まず、23 タスクの線形プローブ ベンチマークでは、エンコーダーは 23 タスク中 21 タスクで SOTA モデル (つまり、BrainIAC、BrainSegFounder、MedicalNet) を上回るか、またはそれに匹敵します。第 2 に、これらの臨床的に有益な埋め込みでトレーニングされた条件付き拡散変換器 (DiT) は、6 つの変数にわたる条件付き生成と患者固有の長期的予測の両方をサポートします。これらの結果を総合すると、下流の臨床タスクと制御可能な生成の両方を実行できる単一の 3D 脳 MRI 埋め込み空間が確立されます。

原文 (English)

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented cohorts, simulate disease trajectories, and support privacy-preserving data sharing. Latent diffusion has been the go-to solution for modeling imaging data, but it places two competing demands on the tokenizer: encoder embeddings must retain the clinical information that downstream tasks act on, and the decoder must reconstruct anatomically faithful volumes. Existing reconstruction-driven tokenizers achieve the second at the expense of the first. To address this, we introduce a fully volumetric masked-autoencoder (MAE) based tokenizer for 3D brain MRI latent diffusion, decoupling encoder and decoder: a frozen 3D MAE encoder produces clinically informative embeddings, while a dedicated CNN decoder reconstructs voxels from a linear projection of those embeddings. We pretrain the encoder on 35,309 volumes from 18 public cohorts spanning four modalities, ten disease categories, and 200+ acquisition sites, and demonstrate its dual utility in two settings. First, on a 23-task linear-probing benchmark, the encoder outperforms or matches SOTA models (i.e., BrainIAC, BrainSegFounder, and MedicalNet) on 21 of 23 tasks. Second, a conditional diffusion transformer (DiT) trained on these clinically informative embeddings supports both conditional generation across six variables and patient-specific longitudinal forecasting. Together these results establish a single 3D brain-MRI embedding space capable of both downstream clinical tasks and controllable generation.

13:00 JST研究/論文

コールドスタート推奨のための暗黙的フィードバックのノイズ除去

暗黙的フィードバックは、そのアクセシビリティと汎用性によりレコメンダー システムで広く使用されていますが、通常はノイズの多いサンプル (クリックベイト、位置バイアスなど) を提示します。一方、推奨者は、新しいアイテムが継続的に流入するため、必然的にアイテムのコールド スタートの問題に直面します。前述の要因により、冷たいアイテムはノイズの多いサンプルになりやすいことがわかりましたが、研究者は冷たいアイテムに対する暗黙的なフィードバックのノイズ除去の重要性を見落とすことがよくあります。これまでのノイズ除去研究では通常、より高い損失値などのヒューリスティック パターンに基づいてノイズの多いサンプルを特定し、サンプルの選択や再重み付けを通じてノイズを軽減しました。ただし、これらの方法の適応性は限られており、コールド スタート シナリオでは効果がありません。コールドスタート推奨のための暗黙的フィードバックのノイズ除去を実現するために、DIF と呼ばれるモデルに依存しないノイズ除去方法を提案します。まず、コンテンツに対するユーザーの好みが安定しているため、ユーザーがコンテンツに似たウォーム アイテムを通じてコールド アイテムに興味があるかどうかを示す疑似ラベルを推測することができます。さらに、擬似ラベルの精度を向上させるために、コールド品目とウォーム品目の内容類似性に基づいて擬似ラベルの信頼度をモデル化し、サンプルごとに複数の擬似ラベルを集計します。最後に、相対エントロピーとアイテムのコールドスタート状態を考慮して、ノイズのあるサンプルラベルの不確実性を明示的に推定します。これにより、擬似ラベルの役割が適応的にガイドされ、サンプルレベルでノイズのあるラベルが修正されます。 DIF の優位性は、理論的な正当性と現実世界のデータセットでの広範な実験の両方によって裏付けられています。この手法は、10 億ユーザー規模のショート ビデオ アプリケーション Kuaishou に導入され、コールド スタート シナリオ内のさまざまな商業指標を大幅に改善しました。

原文 (English)

Denoising Implicit Feedback for Cold-start Recommendation

Implicit feedback is widely used in recommender systems due to its accessibility and generality, yet it usually presents noisy samples (e.g., clickbait, position bias). Meanwhile, recommenders inevitably face the item cold-start problem due to the continuous influx of new items. We identify that cold items are more prone to noisy samples due to the aforementioned factors, and researchers often overlook the significance of denoising implicit feedback for cold items. Previous denoising studies usually identify noisy samples based on heuristic patterns, such as higher loss values, and mitigate noise through sample selection or re-weighting. However, these methods have limited adaptability and are ineffective in cold-start scenarios. To achieve denoising implicit feedback for cold-start recommendation, we propose a model-agnostic denoising method called DIF. First, user preferences for content remain stable, which allows us to infer pseudo-labels indicating whether a user is interested in a cold item through content-similar warm items. Furthermore, to improve pseudo-label accuracy, we model the confidence of pseudo-labels based on the content similarity between the cold item and warm items, and then aggregate multiple pseudo-labels for each sample. Finally, we explicitly estimate the uncertainty of the noisy sample label by considering its relative entropy and the cold-start status of the item, which adaptively guides the role of pseudo-labels to correct the noisy labels at the sample level. DIF's superiority is supported by both theoretical justification and extensive experiments on real-world datasets. The method has been deployed on a billion-user scale short video application Kuaishou and has significantly improved various commercial metrics within cold-start scenarios.

13:00 JSTエージェント研究/論文

分散型連合形成のための離脱と参加のダイナミクス

この論文では、一方的な離脱と参加の決定によって推進される分散型の動的プロセスとしての連合形成を研究します。エージェントは Aumann-Dreze 値を使用してローカルな動きを評価するため、報酬はグローバルに交渉された連合構造を通じてではなく、エージェントの現在の連合内で計算されます。結果として得られるモデルは、協調的なペイオフ配分と非協調的な最良応答行動を結び付けます。最終分割はまさに、個別に利益をもたらす離脱と結合の逸脱が許容されない連合構造です。平衡特性を確立し、ダイナミクスがスカラー リアプノフ表現または正確なポテンシャル表現を許容する条件を特定し、切り替えコストと受け入れコストが局所安定性をどのように形成するかを分析します。数値実験では、有限時間の安定化、コスト感度、および特別な凸ゲーム ベンチマークをテストします。

原文 (English)

Exit-and-Join Dynamics for Decentralized Coalition Formation

This paper studies coalition formation as a decentralized dynamical process driven by unilateral exit-and-join decisions. Agents evaluate local moves using the Aumann-Dreze value, so payoffs are computed within the agent's current coalition rather than through a globally negotiated coalition structure. The resulting model links cooperative payoff allocation with noncooperative best-response behavior: a terminal partition is precisely a coalition structure with no admissible, individually profitable exit-and-join deviation. We establish equilibrium characterizations, identify conditions under which the dynamics admit scalar Lyapunov or exact-potential representations, and analyze how switching and acceptance costs shape local stability. Numerical experiments test finite-time stabilization, cost sensitivity, and a special convex-game benchmark.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

静的リーダーボードを超えて: LLM エージェントの評価の予測的妥当性

エージェントのベンチマークは急速に成長していますが、展開によって明らかにされる 4 つまたは 5 つの側面を超えるベンチマークはありません。このペーパーは、MCP ベースの産業エージェント ベンチマークのこれまでで最大規模の調整された詳細調査を集約したものです。新しい資産クラス (マルチモーダルなビジュアル拡張を含む)、代替オーケストレーション、取得戦略、推論モード、インフラストラクチャの最適化、および評価方法論のプローブをカバーする 14 件の並行実装調査です。これらの研究を以前の 7 つのエージェント ベンチマークと統合すると、集計スコア リーダーボードは導入されたエージェントの評価を体系的に過小評価していると主張します。集計スコアから導出されたランキングは、配布外の設定には転送されません。最近の公開対非公開の競争の回顧展は、このランクの不安定性の直接的な経験的証拠を提供しています。私たちは、サンプル内平均ではなく、予測妥当性、サンプル内ランクとサンプル外ランク間の相関関係によるランキング構成を提案し、HELM とそのエージェント時代の後継者の崩壊を展開関連の次元で明らかにする 12 層の測定装置を報告します。このポジションは、明示的なしきい値を備えた 3 つの改ざん可能な配分外基準を通じて運用可能となります。既存の証拠は部分的にそれを裏付けていますが、確認するには薄すぎます。最後に、事前に登録されたパイロット設計と、次世代のエージェント ベンチマークが何を報告すべきかについての現場レベルのビジョンについて説明します。

原文 (English)

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.

13:00 JST画像/動画生成

GLARE: グローバルな説明をクエリするための自然言語インターフェイス

グローバルな説明は、データセット、クラス、意思決定コンテキストにわたるビジョン モデルを理解するために重要ですが、その複雑で一枚岩の性質が実際的な探索を妨げることがよくあります。ユーザーは通常、静的な成果物ではなく、特定の質問に対する的を絞った回答を求めるため、ブラックボックス画像分類器のグローバルな説明への自然言語アクセスを提供する LLM ベースの対話型インターフェイスを提供します。システムのコア LLM は仲介者として機能し、自然言語の質問をローカルの説明データに対する構造化された SQL クエリに変換します。これにより、ユーザーが低レベルの表現にさらされることなく、柔軟な集計が可能になります。クエリごとに、インターフェイスは統計によって拡張された自然言語応答を出力し、ローカルな説明と意図に合わせた視覚化をサポートします。意図の解釈、クエリ マッピングの精度、新しいクエリやデータセットに対する一般化、言語エラーに対する堅牢性についてシステムを評価します。私たちの結果は、LLM を介したクエリによって、人間中心の XAI のグローバルな説明のアクセシビリティとユーザビリティが大幅に向上することを示しています。

原文 (English)

GLARE: A Natural Language Interface for Querying Global Explanations

While global explanations are crucial for understanding vision models across datasets, classes, and decision contexts, their complex and monolithic nature often hinders practical exploration. Because users typically seek targeted answers to specific questions rather than static artifacts, we present an LLM-based interactive interface that provides natural language access to global explanations for black-box image classifiers. The system's core LLM acts as a mediator, translating natural language questions into structured SQL queries over local explanation data. This enables flexible aggregation without exposing users to low-level representations. For each query, the interface outputs statistics-augmented natural language responses, supporting local explanations, and intent-aligned visualizations. We evaluate the system on intent interpretation, query mapping accuracy, generalization to novel queries and datasets, and robustness to linguistic errors. Our results demonstrate that LLM-mediated querying substantially improves the accessibility and usability of global explanations for human-centered XAI.

13:00 JST研究/論文

進化するプログラムのボトルネックによるニューラル組み合わせ最適化の解釈

Neural Combinatorial Optimization (NCO) は優れたパフォーマンスを実現しますが、そのブラックボックスの性質は依然として展開と科学的診断の重要な障害となっています。概念ボトルネック モデル (CBM) などの標準的な解釈可能ツールは、決定が動的で状態に依存し、適切な概念語彙定義が欠けている NCO にとっては不十分です。このギャップを埋めるために、私たちは進化するプログラマティック ボトルネック (EPB) を導入しました。これは、私たちの知る限り、ブラック ボックス NCO モデルを人間が判読可能なプログラム ポートフォリオに抽出することによって NCO ポリシーを解釈するための最初のフレームワークです。 EPB は LLM を採用して一連のプログラムを自律的に進化させますが、各プログラムのステップごとのアクション分散がボトルネックとして機能します。 EPB は反復フレームワークを通じて機能します。ブロック I はプログラム バンクの容量を修正し、スチューデント ルーターの更新用の数値勾配と LLM ベースのプログラム リビジョン用のテキスト 勾配を結合するハイブリッド テキストと数値の勾配降下スキームを導入します。ブロック II は、障害をターゲットにした拡張と冗長プルーニングを通じてバンク容量を動的に適応させます。広範な実験により、EPB の有効性と幅広い適用性が実証され、抽出されたプログラム ポートフォリオは元のパフォーマンスとほぼ一致します。 EPB は、NCO の動作が最適化ステージ全体で変化し、古典的なヒューリスティック バリアントの構成として近似できることも明らかにしています。私たちの研究は、解釈可能な NCO を進歩させ、EPB を逐次意思決定モデルを解釈するための有望なツールとして確立します。

原文 (English)

Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks

Neural Combinatorial Optimization (NCO) achieves strong performance, yet its black-box nature remains a key roadblock to deployment and scientific diagnosis. Standard interpretability tools, such as Concept Bottleneck Models (CBMs), are ill-equipped for NCO, whose decisions are dynamic, state-dependent, and lack proper concept vocabulary definition. To close this gap, we introduce Evolving Programmatic Bottlenecks (EPB), to our knowledge, the first framework for interpreting NCO policies by distilling black-box NCO models into human-readable program portfolios. EPB employs an LLM to autonomously evolve a bank of programs, where each program's per-step action distribution serves as the bottleneck. EPB works through an iterative framework: Block I fixes program bank capacity and introduces a hybrid textual-numerical gradient descent scheme that couples numerical gradients for student router updates and textual gradients for LLM-based program revision; Block II dynamically adapts bank capacity via fault-targeted expansion and redundancy pruning. Extensive experiments demonstrate EPB's effectiveness and broad applicability, where the distilled program portfolios largely match original performance. EPB also reveals that NCO behavior shifts across optimization stages and can be approximated as a composition of classic heuristic variants. Our work advances interpretable NCO and establishes EPB as a promising tool for interpreting sequential decision-making models.

13:00 JST研究/論文

Quranic ASR の事前トレーニング済みトランスフォーマー モデルの比較研究: 音声表現、ラベル形式、およびデータセット構成

コーラン自動音声認識 (ASR) は、コーランの朗読をテキストに変換し、暗記支援ツールやコーラン検索エンジンなどのアプリケーションを可能にすることを目的としています。ただし、既存の ASR モデルは、ユーザーが朗読する聖句で高い単語誤り率 (WER) を示すことが多く、コーランのコーパスを完全にカバーしていません。この論文では、高度な音声特徴抽出手法である Wav2Vec2.0、HuBERT、および XLS-R を使用した、Quranic ASR の事前トレーニング済み Transformer ベースのモデルのドメイン固有の微調整に関する体系的な実証研究を紹介します。これらのモデルは、入力音声の一部をマスクし、Transformer アーキテクチャを使用してコンテキスト認識型音声特徴を学習することにより、自己教師あり学習を適用します。事前トレーニングされたモデルは、専門家とユーザーによる 870 時間の朗読を超えるフィルタリングされたコーラン データセットに基づいて微調整されています。特徴抽出器、出力ラベル形式、トレーニング戦略、およびクリップの長さにわたる包括的なアブレーション研究を通じて、この領域における転写の精度に影響を与える主要な要因を特定します。当社の最高パフォーマンスの構成では、EveryAyah サブセットで 0.08、EveryAyah+Tarteel の組み合わせ設定で 0.11 の WER を達成しました。これは、Citrinet ベースライン (WER = 0.163) に対しておよそ 5 パーセント ポイントの向上を示し、同時に、組み合わせモデルのトレーニング時間を 140 時間から 40 時間に短縮します。発音記号のないアラビア語テキストは最適な微調整結果をもたらし、Wav2Vec2-XLSR-53 は全体的に最も強力な表現を提供します。今後の作業には、データセットの品質の向上と、Tajweed に敏感なアプリケーション向けにより深い音声特徴表現を抽出するための音素認識モデルの開発が含まれます。

原文 (English)

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines. However, existing ASR models often exhibit high Word Error Rates (WER) on user-recited verses and lack full coverage of the Quranic corpus. This paper presents a systematic empirical study of domain-specific fine-tuning of pretrained Transformer-based models for Quranic ASR, using advanced speech feature extraction methods: Wav2Vec2.0, HuBERT, and XLS-R. These models apply self-supervised learning by masking portions of input audio and using Transformer architectures to learn context-aware speech features. The pretrained models are fine-tuned on a filtered Quranic dataset exceeding 870 hours of professional and user recitations. Through comprehensive ablation studies across feature extractors, output label formats, training strategies, and clip durations, we identify the key factors that affect transcription accuracy in this domain. Our best-performing configuration achieves a WER of 0.08 on the EveryAyah subset and 0.11 on the combined EveryAyah+Tarteel setting, representing roughly a five-percentage-point gain over the Citrinet baseline (WER = 0.163) while reducing combined-model training time from 140 hours to 40 hours. Arabic text without diacritics yields the best fine-tuning results, and Wav2Vec2-XLSR-53 provides the strongest overall representation. Future work includes improving dataset quality and developing phoneme-aware models to extract deeper speech feature representations for Tajweed-sensitive applications.

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

Agentic レビュー システムのベンチマーク

AI支援研究によって査読システムにかかる圧力に対する救済策として、新しい種類のエージェントレビューシステムが登場しているが、それをどのように評価すべきかは不明である。私たちは、最先端の効率的なモデルにわたる 6 つの LLM にわたって、2 つのオープンソース システム (OpenAIReview および coarse)、1 つの独自システム (Reviewer3)、およびゼロショット ベースラインを評価します。まず、ICLR/NeurIPS 論文の AI レビューが、引用や受理決定などの外部シグナルによって近似される論文の品質に追随するかどうかを研究します。すべてのシステムはペアごとの精度で偶然以上のパフォーマンスを発揮し、最も優れているのは OpenAIReview + GPT-5.5 の 83.0% です。次に、システムが既知のグラウンド トゥルースでエラーを捕捉できるかどうかをテストするために、8 つの arXiv 主題クラスにわたる論文に 4 つのカテゴリのエラーを注入する摂動ベンチマークを構築し、検出再現率を測定します。最も強力な構成 (OpenAIReview + GPT-5.5) は、挿入されたエラーの 71.6% を捕捉し、改善の余地がかなり残されています。 6 つのモデルにわたる検出の統合は 83.3% の再現率に達し、異なるモデルが異なるエラーを検出し、より良いハーネス設計によりパフォーマンスが向上する可能性があることを示唆しています。これらのベンチマークを超えて、実際のユーザーを使用して OpenAIReview のパブリック デプロイメントを研究します。そのコメントに対する投票は 1.44 対 1 で肯定的な意見に偏っており、最も一般的な苦情は誤検出や些細な指摘に関するものです。実際の研究論文で最先端のモデルに裏付けられた完全なレビュー システムを一緒に評価することで、AI レビューにはまだ改善の余地があるものの、すでに人間の品質判断を適切に追跡し、重要なエラーをキャッチし、実際のユーザーから肯定的なフィードバックを得ることができることを示します。

原文 (English)

Benchmarking Agentic Review Systems

A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coarse), one proprietary system (Reviewer3), and a zero-shot baseline, across six LLMs spanning frontier and efficient models. First, we study whether AI reviews on ICLR/NeurIPS papers track with papers' quality as approximated by external signals such as citations and acceptance decisions. Every system performs above chance in pairwise accuracy, and the best is OpenAIReview + GPT-5.5 at 83.0%. Second, to test whether systems can catch errors with known ground truth, we construct a perturbation benchmark that injects four categories of errors into papers across eight arXiv subject classes and measure detection recall. The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors, leaving substantial room for improvement. The union of detections across six models reaches 83.3% recall, suggesting different models detect different errors and better harness design can potentially increase performance. Beyond these benchmarks, we study a public deployment of OpenAIReview with real users. Votes on its comments skew positive at 1.44 to 1, and the most common complaints are about false positives and minor nitpicks. Together, by evaluating full review systems backed by state-of-the-art models on real research papers, we show that while AI reviews still have room for improvement, they can already track human quality judgments well, catch important errors, and earn positive feedback from real users.

13:00 JST研究/論文

グラウンデッド推論: 決定論的にカプセル化された生成モデルの原則

生成モデルを従来の計算システムに組み込むことは、大きなチャンスと大きな危険の両方をもたらします。多くの早期導入者は多大な費用をかけてこれらの危険性を認識していますが、この分野では依然として、AI を従来のシステムに組み込むリスクを回避するための基礎的なフレームワークが必要です。この原稿は、確率モデルの決定論的なカプセル化を可能にするように設計された、AI ブレンド アーキテクチャの 4 つの特定のプリミティブの定義を通じて、この基盤を確立します。さらに、業界全体で広く代表される 2 つの包括的なアンチパターンを確立し、この分野のエンジニアへの警告として機能します。このフレームワークは、生成モデル プロバイダーが次世代の生成モデル インターフェイスを構築できる基盤を提供しながら、AI を従来のシステムにうまく統合できるように設計されました。

原文 (English)

Grounded Inference: Principles for Deterministically Encapsulated Generative Models

The incorporation of generative models into traditional computational systems presents both enormous opportunity and tremendous peril. Although many early adopters have realized these perils at great expense, the field still requires foundational frameworks to de-risk incorporation of AI into traditional systems. This manuscript establishes this foundation through the definition of four specific primitives of AI blended architecture, designed to enable deterministic encapsulation of probabilistic models. It further establishes two overarching anti-patterns broadly represented across industry to serve as warnings for engineers in this field. This framework was designed to enable successful integration of AI into traditional systems while providing a foundation upon which generative model providers could build the next generation of generative model interfaces.

13:00 JST研究/論文

ナレッジワーカーの質問応答フォーラムにおける最適なスケジューリング

個人が疑問に対する答えを見つけるためにインターネットにアクセスするにつれて、いくつかの質問応答 (QA) フォーラムが発展してきました。そこでは、特定のトピックに精通したユーザーが専門知識を提供して、これらの情報要求に答えることができます。これらは現在ボランティアベースですが、将来的には特定のトピックの専門家である知識労働者を雇用するバージョンを検討しています。このようなシステムでは、キュー システムを形成する要求-回答プロセスは、さまざまなトピックの要求をフォーラムの専門家に割り当てるスケジューラを利用することができ、フォーラムの専門家はさまざまなトピックの専門知識レベルに応じてそれらの要求に答えることができます。このモデルでは、システムを安定に保ちながらリクエストを処理するためのシステムのキャパシティを計算し、キャパシティを達成するスケジューラを設計します。また、リクエストに応える際に専門家間の協力がどのように潜在的に容量を増加できるかについても調査します。

原文 (English)

Optimal Scheduling in a Question-Answering Forum of Knowledge Workers

As individuals turn to the Internet to find answers to questions they may have, several Question Answering (QA) forums have evolved, where users knowledgeable in certain topics can contribute their expertise to answering these requests for information. While these are currently volunteer based, we consider a future version employing knowledge workers who are experts in certain topics. In such a system, the request-answer processes forming the queuing system may utilize schedulers that assign requests in different topics to the experts in the forum, who may be able to answer them according to their expertise levels in different topics. With this model, we calculate the capacity of the system for handling the requests while keeping the system stable, and design schedulers that achieve capacity. We also investigate how collaboration between experts in answering requests can potentially increase capacity.

13:00 JSTLLM/生成AI

エントロピーを超えて: LLM 推論のためのトークンレベルの分布偏差からの学習

検証可能な報酬を伴う強化学習 (RLVR) は、大規模言語モデル (LLM) 推論を大幅に進歩させました。ただし、根本的な最適化の不安定性に直面しています。均一なトークンの更新はエントロピーの崩壊を促進し、次善の戦略への早期収束につながりますが、過剰なシャノンのエントロピーの最大化はエントロピーの爆発を引き起こし、一貫性のない推論チェーンへの盲目的な探索を引き起こす可能性があります。この二分法を解決するために、最適化の焦点をスカラーの不確実性からトークン ロジットの分布特性に移す独立組み合わせトークン (ICT) フレームワークを導入します。 ICT は、トークン ロジット分布間のジェンセン シャノン (JS) の相違を活用することで、LLM 推論における効果的な探索を導くための重要な分岐点として、独特の分布パターンを持つトークンを特定します。シャノンエントロピーと 2 次 R\'enyi エントロピーの両方に基づいた私たちの理論分析は、これらのトークンを選択的に更新することで政策の集中が調整されることを証明しています。つまり、シャノン エントロピーによって測定される全体的な分布の不確実性が軽減され、同時に 2 次 R\'enyi エントロピーによって捕捉される確率集中が制御されます。この二重の効果により、過度に集中したトークン生成による探査の弱体化が防止され、トレーニング状況が効果的に安定します。経験的な結果は、Qwen2.5 (0.5B/1.5B/7B) モデルの一意のトークンの上位 10% のみを更新すると、数学、常識、オリンピック レベルの問題にわたる 7 つのベンチマークにわたって、GRPO、20-エントロピー、および STAPO ベースラインを超えて、平均 pass@4 が 4.58% 向上し、最大 14.9% の向上が得られることを示しています。

原文 (English)

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced Large Language Model (LLM) reasoning; however, it faces a fundamental optimization instability: uniform token updates precipitate entropy collapse, leading to premature convergence to suboptimal strategies, whereas excessive Shannon Entropy maximization can cause entropy explosion, driving blind exploration toward incoherent reasoning chains. To resolve this dichotomy, we introduce the Independent Combinatorial Tokens (ICT) framework, which shifts the optimization focus from scalar uncertainty to the distributional properties of token logits. By leveraging the Jensen-Shannon (JS) divergence between token logits distributions, ICT identifies tokens with distinctive distributional patterns as critical branching points for guiding effective exploration in LLM reasoning. Our theoretical analysis, grounded in both Shannon and second-order R\'enyi entropy, proves that selectively updating on these tokens regulates policy concentration: it reduces the overall distribution uncertainty measured by Shannon entropy, while controlling probability concentration captured by second-order R\'enyi entropy. This dual effect prevents over-concentrated token generation from weakening exploration and effectively stabilizes the training landscape. Empirical results demonstrate that updating only the top 10% of unique tokens on Qwen2.5 (0.5B/1.5B/7B) models yields an average pass@4 improvement of 4.58%, with a maximum gain of 14.9%, over GRPO, 20-Entropy, and STAPO baselines across seven benchmarks spanning math, commonsense, and Olympiad-level problems.

13:00 JSTLLM/生成AIエージェントGemini

AgentFinVQA: 監査可能な財務チャート QA のための展開可能なマルチエージェント パイプライン

規制された環境における財務チャートの質問回答には、正確性以上のものが求められます。実務者は、回答に基づいて行動する前に、どの回答を信頼すべきかを知る必要があり、多くの機関は顧客データを外部モデルプロバイダーに送信できません。しかし、既存のチャート QA エージェントは精度重視かつ不透明で、ほとんどが独自の API アクセスを前提としています。私たちの知る限り、精度を大幅に損なうことなく監査可能性とオンプレミス展開可能性を組み合わせたものはありません。 AgentFinVQA は、各クエリを計画、OCR、凡例の根拠、視覚的検査、検証に分解し、サンプルごとに追跡可能なモデル評価パケット (MEP) のすべてのステップを記録するマルチエージェント パイプラインです。 FinMME では、AgentFinVQA は、独自のバックボーン (Gemini-3 フラッシュ; 71.24% 対 63.56%、McNemar $p \約 1.1 \times 10^{-16}$) を使用したプライマリ バックボーンと一致するゼロショット ベースラインと比較して $+7.68$ pp 改善し、オープンウェイトでは $+4.84$ pp 改善します。 Qwen3.6-27B-FP8 はローカルでサービスされています。検証者の評決は有用な信頼シグナル (確認済み回答と修正済み回答の正確な精度 68.2% 対 55.6%) としても機能し、人間によるレビュー ルーティングが可能になります。エラー分析により、質問の誤解、凡例の混乱、抽出エラーが失敗の 3 分の 2 近くを占め、検証者によって最も検出されにくいカテゴリーであることが示され、今後の作業の明確な方向性が特定されます。これらの結果を総合すると、監査可能なオンプレミスの財務チャート QA が実用的であり、オープンウェイト システムが完全なデータ常駐を可能にしながら精度の向上のほとんどを維持していることがわかります。再現可能な評価をサポートするためにコードをリリースします。

原文 (English)

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

Financial chart question answering in regulated settings demands more than accuracy: practitioners must know which answers to trust before acting on them, and many institutions cannot send client data to external model providers. Yet existing chart-QA agents are accuracy-focused and opaque, and most assume proprietary API access; to our knowledge, none combines auditability with on-premise deployability without significant accuracy compromise. We present AgentFinVQA, a multi-agent pipeline that decomposes each query into planning, OCR, legend grounding, visual inspection, and verification, recording every step in a traceable Model Evaluation Packet (MEP) per sample. On FinMME, AgentFinVQA improves $+7.68$ pp over a primary-backbone matched zero-shot baseline with a proprietary backbone (Gemini-3 Flash; 71.24% vs. 63.56%, McNemar $p \approx 1.1 \times 10^{-16}$), and $+4.84$ pp with open-weights Qwen3.6-27B-FP8 served locally. The verifier's verdict also serves as a useful confidence signal (68.2% vs. 55.6% exact accuracy on confirmed vs. revised answers), enabling human-in-the-loop review routing. Error analysis shows that question misunderstanding, legend confusion and extraction error account for nearly two-thirds of failures and are the categories least detected by the verifier, identifying clear directions for future work. Together these results show that auditable, on-premise financial chart QA is practical and that the open-weights system keeps most of the accuracy gains while enabling full data residency. We release our code to support reproducible evaluation.

13:00 JSTLLM/生成AIエージェント研究/論文

ORAgentBench: LLM エージェントは困難なオペレーション リサーチ タスクをエンドツーエンドで解決できますか?

大規模な言語モデルは、実行可能環境で複数ステップのタスクを実行するための自律エージェントとして導入されることが増えていますが、現実的なオペレーション リサーチ (OR) 作業を実行する能力は依然として不明です。既存の OR 評価では、多くの場合、モデリングと解決策が切り離されており、事前に形式化されたインスタンスまたはテキストのみのインスタンスに依存しており、運用成果物から検証済みの意思決定に至るワークフロー全体をテストすることはほとんどありません。この作業では、困難なエンドツーエンドのオペレーション リサーチ タスクで自律エージェントを評価するための実行ベースのベンチマークである ORAgentBench を紹介します。これには、さまざまな運用シナリオにわたって人間がレビューした 107 のタスクが含まれており、それぞれが自然言語の概要、複数ファイルのデータ、構成アーティファクト、および必要な送信スキーマを備えた隔離された環境にパッケージ化されています。エージェントはソリューション コードを作成して実行する必要があり、その送信内容は、スキーマの有効性、厳密な制約の実現可能性、および正規化された客観的な品質について、非表示のバリデーターによって評価されます。 14 のフロンティア エージェント モデル構成による実験では、現在のエージェントが信頼できる OR 実践からはほど遠いことが示されています。最も優れたエージェントが合格できるのは、全タスクの 35.51% と難しいタスクの 20.59% だけであり、実行可能な提出物の多くは依然として要求される品質のしきい値を下回っています。さらに、障害分析では、運用ルールの欠如、脆弱な定式化、実行可能なソリューションの構築の脆弱さ、不十分なソリューションの改善など、戦略的な弱点がエラーの大部分を占めていることがわかります。手術室固有の手順スキルは、困難なタスクの実現可能性を高めますが、ソリューションの品質や合格率を確実に向上させるものではありません。これらの結果は、OR エージェントの進歩には、もっともらしい最適化コードを超えて、信頼できる高品質な運用上の意思決定に移行する必要があることを示唆しています。

原文 (English)

ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR) work remains unclear. Existing OR evaluations often decouple modeling from solving, rely on pre-formalized or text-only instances, and rarely test the full workflow from operational artifacts to validated decisions. In this work, we introduce ORAgentBench, an execution-grounded benchmark for evaluating autonomous agents on challenging end-to-end operations research tasks. It contains 107 human-reviewed tasks across diverse operational scenarios, each packaged in an isolated environment with a natural-language brief, multi-file data, configuration artifacts, and a required submission schema. Agents must write and run solution code, and their submissions are evaluated by hidden validators for schema validity, hard-constraint feasibility, and normalized objective quality. Experiments with fourteen frontier agent-model configurations show that current agents remain far from reliable OR practice. The best agent passes only 35.51% of all tasks and 20.59% of hard tasks, and many feasible submissions still fall below the required quality threshold. Failure analysis further shows that errors are dominated by strategic weaknesses, including missed operational rules, brittle formulations, weak feasible-solution construction, and insufficient solution improvement. OR-specific procedural skills increase hard-task feasibility, but do not reliably improve solution quality or pass rate. These results suggest that progress in OR agents requires moving beyond plausible optimization code toward dependable, high-quality operational decision-making.

13:00 JSTLLM/生成AI研究/論文

CombEval: 大規模言語モデルでの組み合わせカウントを評価するためのフレームワーク

大規模な言語モデルで組み合わせカウントを評価するための動的ベンチマークである CombEval を紹介します。 CombEval は、各問題をエンティティ、組み合わせオブジェクト、オブジェクトの依存関係、および制約に対する型付きの Cofola 仕様として表現し、ソルバーによって正確に検証された回答を含む自然言語の計数問題の制御された生成を可能にします。静的コレクションとは異なり、CombEval は、オブジェクト タイプ、エンティティのスケール、制約の数、および推論の深さの体系的なバリエーションをサポートします。直接およびコード拡張設定の下で 11 個の LLM を評価したところ、順序付けされたオブジェクト、区別できない要素、相対的な位置の制約、および入れ子になったオブジェクトの依存関係に関してモデルが脆弱なままであることがわかりました。エラー分析により、制約の解釈とカウント原則の欠陥がさらに特定されます。 CombEval は、LLM がいつ、そしてなぜ組み合わせ推論で失敗するかを研究するための診断テストベッドを提供します。コードと生成されたベンチマーク スイートは、\url{https://github.com/YuxuZhou-CN/combination-problem-generation} で公開されています。

原文 (English)

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models. CombEval represents each problem as a typed Cofola specification over entities, combinatorial objects, object dependencies, and constraints, enabling controlled generation of natural-language counting problems with exact solver-verified answers. Unlike static collections, CombEval supports systematic variation of object type, entity scale, constraint count, and reasoning depth. We evaluate 11 LLMs under direct and code-augmented settings and find that models remain brittle on ordered objects, indistinguishable elements, relatively positional constraints, and nested object dependencies. Error analysis further identifies failures in constraint interpretation and counting principles. CombEval provides a diagnostic testbed for studying when and why LLMs fail at combinatorial reasoning. The code and generated benchmark suites are publicly available at \url{https://github.com/YuxuZhou-CN/combination-problem-generation}.

13:00 JSTLLM/生成AI

もう一度考えますか、それとも長く考えますか?予算を考慮した推論のための選択的検証

テスト時の推論は、提供時間の制御ノブとしてますます使用されていますが、追加の推論は一様に価値があるわけではありません。失敗した試行を修復したり、既に正解している解答の計算を無駄にしたり、有害な解答の変更を導入したりする可能性があります。私たちはこれを、新規検証者の問題ではなく、デプロイメント割り当て問題として研究します。 \sevra (推論割り当てのための選択的検証) を導入します。これは、フリーズしたソルバーの最初の答えを保存するか、アクティブな検証を呼び出すかを決定するサービス層コントローラーです。凍結された Qwen3-4B ソルバーを使用して、介入の結果をログに記録し、サービングの可視の試行状態から回復可能性を認識したゲートをトレーニングします。 \mathfive では、選択的検証の精度は 76.3\% に達し、常時検証の 75.5\% と比較して、生成後のトークンが 26.8\% 削減され、有害な反転が 2.2\% から 1.0\% に減少します。ただし、8,192 トークンの初期解決では、モデル トークンの総数が 28\% 減りながら 76.0\% の精度に達しました。これは、選択的回復は有用ですが、最もテストされたコスト フロンティアではないことを示しています。 \gsm への凍結転送では、選択的ポリシーは例の 3.0\% のみを検証し、精度を 93.4\% から 94.5\% に向上させ、常に検証する場合と比較して検証トークンを 91.2\% 削減します。繰り返しますが、初期ソルブが長いほど、より少ない実現トークンで精度が高まります。 CommonsenseQA では、常時検証が問題になりますが、Self-Consistency@5 は実際のトークン コストの約 5 倍で精度を向上させます。結果として得られるデプロイメント ルールは、最初に初期予算を調整し、次に明示的なチェック、制限された再試行、監査可能性、または回帰リスク制御が重要な場合に選択的リカバリを使用するというものです。

原文 (English)

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes. We study this as a deployment allocation problem rather than a new-verifier problem. We introduce \sevra, Selective Verification for Reasoning Allocation, a serving-layer controller that decides whether to preserve a frozen solver's initial answer or invoke active verification. Using a frozen Qwen3-4B solver, we log intervention outcomes and train recoverability-aware gates from serving-visible attempt state. On \mathfive, selective verification reaches 76.3\% accuracy, compared with 75.5\% for always verifying, while reducing post-generation tokens by 26.8\% and harmful flips from 2.2\% to 1.0\%. However, an 8,192-token initial solve reaches 76.0\% accuracy with 28\% fewer total model tokens, showing that selective recovery is useful but not the best tested cost frontier. In frozen transfer to \gsm, the selective policy verifies only 3.0\% of examples, improves accuracy from 93.4\% to 94.5\%, and reduces verification tokens by 91.2\% relative to always verifying; again, a longer initial solve matches its accuracy with fewer realized tokens. On CommonsenseQA, always-on verification hurts, while Self-Consistency@5 improves accuracy at about five times the realized token cost. The resulting deployment rule is: tune the initial budget first, then use selective recovery when explicit checks, bounded retries, auditability, or regression-risk control matter.

13:00 JSTLLM/生成AIエージェント

AI 支援による法的証拠開示のためのヒューマン オン ザ ループ オーケストレーション

Autonomous Large Language Model (LLM) エージェントは電子証拠開示 (e-discovery) に導入されることが増えており、複数ステップの推論チェーンにわたる複合エラーが法的違法行為となる可能性があります。シングルターン取得とは異なり、特権付きドキュメント コーパス上で動作するエージェント ワークフローは、「軌道崩壊」と呼ばれる一種の失敗を示します。つまり、初期の誤分類が静かに伝播し、特権レビュー全体が無効になります。この論文は 3 つの貢献を行っています。まず、機能段階ごとに整理した、法律情報検索におけるエージェントの失敗の構造化分類を提案します。次に、これらの障害が悪化する前に阻止するように設計された、計画、推論、実行、不確実性の定量化に及ぶ 4 層の検証アーキテクチャを導入します。 3 番目に、必須のヒューマン オン ザ ループ (HOTL) エスカレーションしきい値が、完全に自律的なベースラインと比較して特権放棄のリスクをどのように低減するかを実証する、合成電子証拠開示コーパスに関する予備的なシミュレーション研究を紹介します。私たちの結果は、調整された不確実性のしきい値により、完全に自律的な展開と比較して特権放棄のリスクを最大 61% 削減できる一方、弁護士の審査に回される文書は 4 分の 1 未満であることを示唆しています。

原文 (English)

Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery

Autonomous Large Language Model (LLM) agents are increasingly deployed in electronic discovery (e-discovery), where compounding errors across multi-step reasoning chains can constitute legal malpractice. Unlike single-turn retrieval, agentic workflows operating over privileged document corpora exhibit a class of failure we term "trajectory collapse": an early misclassification silently propagates, rendering an entire privilege review invalid. This paper makes three contributions. First, we propose a structured taxonomy of agentic failures in legal information retrieval, organized by functional stage. Second, we introduce a four-layer verification architecture -- spanning planning, reasoning, execution, and uncertainty quantification -- designed to intercept these failures before they compound. Third, we present a preliminary simulation study on a synthetic e-discovery corpus that demonstrates how mandatory Human-on-the-Loop (HOTL) escalation thresholds reduce privilege-waiver risk relative to fully autonomous baselines. Our results suggest that calibrated uncertainty thresholds can reduce privilege-waiver risk by up to 61% versus fully autonomous deployment, while routing fewer than one quarter of documents to attorney review.

13:00 JSTエージェント

TelcoAgent: 3GPP に基づいた説明可能性を備えたスケーラブルな 5G マルチ KPM 予測

Key Performance Measurement (KPM) 予測は、5G および次世代通信ネットワークのプロアクティブなネットワーク管理に不可欠です。ただし、既存の機械学習 (ML) アプローチは、スケーラビリティと説明可能性において大きな制限に直面しており、現実世界の展開における有効性が制限されています。私たちは、サイト固有のトレーニングを必要とせずに、さまざまなネットワーク セルにわたる複数の KPM の正確でスケーラブルで説明可能な予測を可能にする基盤モデル ベースのフレームワークである TelcoAgent を提案します。具体的には、このフレームワークは 3 つの主要コンポーネントで構成されます。(i) 仕様書から直接 3GPP (3rd Generation Partnership Project) ナレッジ グラフを構築する自動化された 3 エージェント パイプライン、(ii) 正確なゼロショット予測を実現するスケーラブルな時系列基礎モデル (TSFM) ベースの予測パイプライン、最後に (iii) 実用的なドメインベースの診断を提供する推論および説明パイプライン。米国を拠点とするネットワーク事業者が提供する 3 か月間の現実世界の都市規模の 5G KPM データセットを使用して評価した TelcoAgent は、200 セルにわたるセルごとに考慮された 7 つの KPM のすべてについて高い予測精度を示し、同時にネットワークの劣化に対処するための説明可能な洞察と実行可能な指示を提供します。

原文 (English)

TelcoAgent: A Scalable 5G Multi-KPM Forecasting With 3GPP-Grounded Explainability

Key Performance Measurement (KPM) forecasting is essential for proactive network management of 5G and next-generation telecom networks. However, existing machine learning (ML) approaches face significant limitations in scalability and explainability, restricting their effectiveness in real-world deployments. We propose TelcoAgent, a foundation model-based framework that enables accurate, scalable, and explainable forecasting of multiple KPMs across diverse network cells without the need for site-specific training. Specifically, the framework comprises three key components: (i) an automated three-agent pipeline that constructs a 3rd Generation Partnership Project (3GPP) knowledge graph directly from specification documents, (ii) a scalable, time-series foundation model (TSFM)-based prediction pipeline to deliver accurate, zero-shot forecasting, and finally (iii) a reasoning and explanation pipeline that provides actionable, domain-grounded diagnostics. Evaluated using a 3-month, real-world, city-scale 5G KPM dataset from a U.S.-based network operator, TelcoAgent demonstrates high forecasting accuracy for all 7 considered KPMs per cell across 200 cells, while delivering explainable insights and actionable instructions to address network degradations.

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達

大規模言語モデルのブラックボックス不確実性推定法の系統的評価

大規模言語モデル (LLM) は幅広いタスクにわたって強力な機能を示していますが、その出力は信頼性が低いことが多く、幻覚が含まれる可能性があるため、信頼できる LLM を構築するには不確実性推定 (UE) が不可欠です。実際には、多くの主流 LLM は制限された API を介してのみアクセスでき、ロジットや隠れ状態などの内部信号は利用できないため、ブラックボックス UE が特に重要になります。しかし、LLM 用のブラックボックス UE に関する既存の研究は、方法論において断片的なままであり、統一された実証的比較が欠けています。このギャップに対処するために、ブラックボックス UE 手法の体系的なレビューを提示し、言語化ベース、サンプリング ベース、説明ベース、マルチエージェント、およびハイブリッド手法の 5 つのカテゴリに整理します。さらに、統一された評価フレームワークを構築し、4 つのモデルと 4 つのデータセット設定にわたる 24 の代表的な手法をベンチマークします。私たちの結果は、すべての設定において一貫して優勢な単一の方法はないことを示しています。それにもかかわらず、回答空間内の候補を推論して比較する方法は一般に効果的であり、複数の不確実性信号を組み合わせるハイブリッド方法は、ほとんどの条件下で良好に機能します。ベンチマーク データと統一評価フレームワークを公開することで、再現可能な比較を促進し、将来の研究をサポートすることを目指しています。また、実証結果は、LLM 向けの将来のブラック ボックス UE 手法を開発するための実践的なガイダンスを提供します。

原文 (English)

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.

13:00 JSTエージェント研究/論文

MetaResearcher: 敵対的な仮想環境における自己反省強化学習によるディープリサーチの拡張

深層調査エージェントは、自律的な情報収集と合成において優れた能力を実証してきましたが、そのトレーニングは、シミュレートされた環境の静的な性質、事実検索のみのタスク設計の限界、および結果ベースの強化学習の非効率性によって依然として制約を受けています。この研究では、4 つの相乗的な側面にわたって詳細な調査エージェントのトレーニングを拡張する新しいフレームワークである MetaResearcher を提案します。まず、時間的なダイナミクスと敵対的な誤った情報をトレーニング環境に注入する進化する仮想世界を導入し、エージェントにソースの信頼性評価と時間的な競合解決スキルの開発を強制します。次に、仮説生成や矛盾解決を含む発見指向のタスクを設計します。これは、単純な事実検索を超えて、エージェントを真の調査行動へと導きます。第三に、回答の正確性、検索パスの効率、反映の深さ、ツール呼び出しの多様性を共同で最適化し、以前の研究で観察された反復アクションのループ問題に直接対処する、GRPO フレームワーク内の自己反射型メタ報酬メカニズムを提案します。 4 番目に、調整された強化学習を通じて共同研究戦略を学習する、特殊な Scout、Filter、および Synthesizer モデルで構成される異種マルチエージェント Swarm アーキテクチャを導入します。 LiteResearcher インフラストラクチャ上に構築された MetaResearcher は、トレーニングに限界 API コストをゼロにしながら、ベンチマーク パフォーマンス (GAIA、Xbench-DS) と敵対的条件下での認識論的堅牢性の両方の大幅な向上を目指しています。完全なフレームワーク設計、トレーニング方法論、計画された実験的検証を紹介します。

原文 (English)

MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments

Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remains constrained by the static nature of simulated environments, the limits of fact-retrieval-only task designs, and the inefficiency of outcome-based reinforcement learning. In this work, we propose MetaResearcher, a novel framework that scales deep research agent training across four synergistic dimensions. First, we introduce an Evolving Virtual World that injects temporal dynamics and adversarial misinformation into the training environment, forcing agents to develop source credibility assessment and temporal conflict resolution skills. Second, we design Discovery-Oriented Tasks -- including hypothesis generation and contradiction resolution -- that transcend simple fact retrieval and push agents toward genuine research behaviors. Third, we propose a Self-Reflective Meta-Reward mechanism within the GRPO framework that jointly optimizes for answer correctness, search path efficiency, reflection depth, and tool call diversity, directly addressing the repetitive action loop problem observed in prior work. Fourth, we introduce a Heterogeneous Multi-Agent Swarm architecture comprising specialized Scout, Filter, and Synthesizer models that learn collaborative research strategies through coordinated reinforcement learning. Built upon the LiteResearcher infrastructure, MetaResearcher requires zero marginal API cost for training while targeting substantial improvements in both benchmark performance (GAIA, Xbench-DS) and epistemic robustness under adversarial conditions. We present the complete framework design, training methodology, and planned experimental validation.

13:00 JSTLLM/生成AIエージェント

マルチエージェントのトランザクション メモリ

多様なタスクにわたって多様な機能を備えた LLM エージェントの分散展開により、異種エージェント集団全体で知識を共有するためのインフラストラクチャが促進されます。検索エンジンが人間の問題解決をサポートするために人間が生成した成果物にインデックスを付けるのと同じように、検索システムはエージェントが生成した成果物を整理してエージェント集団全体で再利用できます。私たちは、人間が作成したアーティファクトの価値を個々のエージェントに実証する検索拡張生成を、エージェントの集団をサポートするエージェント生成アーティファクトの検索まで拡張します。特に、エージェントの軌跡は再利用可能な手順知識をエンコードしますが、これらのアーティファクトは通常 1 回の使用後に破棄されるか、作成エージェントによってのみ保持されるため、新しくインスタンス化されたエージェントは既存のソリューションを繰り返し再発見する必要があります。我々は、エージェントが生成した軌跡を集団レベルで保存および取得するためのフレームワークであるマルチエージェント トランザクション メモリ (MATM) を提案します。このフレームワークでは、プロデューサー エージェントが軌跡を共有リポジトリに提供し、コンシューマ エージェントが軌跡を取得してタスクの実行を改善します。私たちは、軌跡が長く、特に豊富な手続き構造をエンコードするインタラクティブな環境 (ALFWorld および WebArena) に焦点を当てています。私たちの実験では、MATM から軌道を取得すると、調整や共同トレーニングを行わなくても、下流のタスクのパフォーマンスが向上し、インタラクションのステップが削減されることが実証されました。これらの結果は、MATM をオープン エージェント エコシステムにおける集団レベルのエクスペリエンス共有のための設計パターンとして位置づけています。

原文 (English)

Multi-Agent Transactive Memory

The decentralized deployment of LLM agents with diverse capabilities across diverse tasks motivates infrastructure for knowledge sharing across heterogeneous agent populations. Just as search engines index human-generated artifacts to support human problem solving, retrieval systems can organize agent-generated artifacts for reuse across agent populations. We extend retrieval-augmented generation - which demonstrates the value of human-authored artifacts to individual agents - to retrieval of agent-generated artifacts supporting a population of agents. In particular, agent trajectories encode reusable procedural knowledge, yet these artifacts are typically discarded after a single use or retained only by the producing agent, forcing newly instantiated agents to repeatedly rediscover existing solutions. We propose Multi-Agent Transactive Memory (MATM), a framework for population-level storage and retrieval of agent-generated trajectories, where producer agents contribute trajectories to a shared repository and consumer agents retrieve them to improve task execution. We focus on interactive environments (ALFWorld and WebArena), where trajectories are long and encode especially rich procedural structure. Our experiments demonstrate that retrieving trajectories from MATM improves downstream task performance and reduces interaction steps without coordination or joint training. These results position MATM as a design pattern for population-level experience sharing in open agent ecosystems.

13:00 JST研究/論文

eCNNTO: トポロジー最適化を加速するための高度に汎用化可能な ConvNet

この研究では、eCNNTO と呼ばれる、密度ベースのトポロジー最適化 (TO) を加速する要素ベースの畳み込みニューラル ネットワーク (CNN) を提案します。 TO では一般に多数の反復が行われ、反復ごとに有限要素解析が実行されるため、特に高解像度設計を達成するために高密度メッシュが使用される場合に効率のボトルネックが発生します。この制限に対処するために、eCNNTO は Kallioras らに基づいて構築することが提案されています。 (2020) では、ディープ ビリーフ ネットワーク (DBN) がすべての要素に対してトレーニングされ、初期の歴史から最適に近い密度を予測することで、反復の大部分がスキップされ、TO 手順が大幅に高速化されました。ただし、この方法には隣接する要素間の空間的相関が欠如しており、最終的な構造でフィーチャが切断される可能性があります。提案された方法は、この問題に対処するために残留接続を備えた CNN を採用します。それに加えて、最適化効率をさらに高めるために新しいトレーニング戦略が導入されており、トレーニング データセットは初期のものではなく最終段階の密度履歴で構成されています。この変更は、必要なトレーニング データのサイズを削減するのにも役立ちます。 eCNNTO はトレーニングに小さなデータセットしか必要としませんが、大きく異なる境界条件、荷重ケース、設計ドメインのジオメトリ、メッシュ解像度、および非設計ドメインの問題に一般化できます。最終的に、eCNNTO の一般化機能と効率が、2 次元および 3 次元のさまざまな例を通じて実証され、それぞれ最大 90% と 97% の反復の削減が達成されました。

原文 (English)

eCNNTO: A Highly Generalizable ConvNet for Accelerating Topology Optimization

This work proposes an element-based Convolutional Neural Network (CNN) to accelerate density-based Topology Optimization (TO), termed eCNNTO. TO generally undergoes a large number of iterations, where finite element analysis is performed in every iteration, leading to the efficiency bottleneck especially when dense meshes are used to achieve high-resolution designs. To address this limitation, eCNNTO is proposed to build upon Kallioras et al. (2020), where a Deep Belief Network (DBN) was trained for every element to predict its near-optimal density from its early history, thereby skipping the great majority of iterations and significantly accelerating the TO procedure. However, the method lacks spatial correlations among neighboring elements and may lead to disconnected features in the final structure. The proposed method employs CNN with residual connections to address this issue. On top of it, a novel training strategy is introduced to further enhance the optimization efficiency, where the training dataset consists of the final stage density histories rather than early ones. This change can also help reduce the required training data size. eCNNTO requires only a small dataset to train and yet it can be generalized to problems with largely different boundary conditions, loading cases, design domain geometries, mesh resolutions, as well as non-design domains. In the end, the generalization capabilities and efficiency of eCNNTO are demonstrated through a variety of examples in two and three dimensions, achieving up to 90% and 97% reduction of iterations, respectively.

13:00 JST研究/論文

エージェンシーの道: Autotelic AI、組み込みエージェンシー、自己の解体

ほとんどの人工知能システムは、目標が外生的であり、設計者によって指定されるという前提に基づいて構築されています。エージェントが独自の目標を生成し始めると何が起こるかを探ることで、オートテリック AI の分野が開かれます。エージェントには、単に目的を追求するだけでなく、それを発見することが期待されています。この記事では、内発的動機づけ、リソース主導型事前分布、因果介入学習、ホメオスタシス、および埋め込み性を通じてその結果を追跡します。最後の条件は、自己主体性にとって必要条件ではあるが、十分条件ではないことが判明しています。埋め込み性は、その個性がユニークではないことを明らかにするという代償を払ってエージェントを個性化します。そのため、同じダイナミクスで多くの有効な分割が許容され、それぞれが異なる自己候補を定義します。したがって、オートテリック AI の最も深刻な問題は、エージェントがどのようにして目標を生成するかということではなく、エージェントがどのようにして目標が割り当てられる自己を生成し、相対化するかということです。エージェントは行動するために自分自身の境界を信じ、理解するためにその境界を見通さなければなりません。私たちはこれらの開発を単一のフレームワークに統合し、それを 3 つの方向に沿って拡張します。エージェントと環境の切断が物理的になる量子定式化、非二元的な瞑想的伝統に対する哲学的解釈、および具体的な LLM ベースのエージェントのインスタンス化です。

原文 (English)

The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self

Most artificial intelligence systems are built on the assumption that goals are exogenous and specified by the designer. Exploring what happens when an agent begins generating its own goals opens the field of autotelic AI. Agents are expected not merely to pursue objectives but to discover them. In this article, we trace its consequences through intrinsic motivation, resource-driven priors, causal-interventional learning, homeostasis, and embeddedness; the last of which is found to be a necessary but not sufficient condition for autotelic agency. Embeddedness individuates the agent at the cost of revealing that the individuation is non-unique, such that the same dynamics admit many valid partitions, each defining a different candidate self. The deepest problem with autotelic AI is therefore not how the agent generates goals, but how it generates and relativizes the self to which the goals are assigned. The agent must believe in its own boundary in order to act, and see through that boundary in order to understand. We consolidate these developments into a single framework and extend it along three directions: a quantum formulation in which the agent-environment cut becomes physical, a philosophical reading against non-dual contemplative traditions, and a concrete LLM-based agentic instantiation.

13:00 JSTロボティクス

PhysDrift: ヒューマノイドの共同音声モーション生成における身体のギャップを埋める

ヒューマノイドロボットは、表現力豊かで音声に合わせて動作するだけでなく、実施形態の制約の下で物理的に実行可能な同時音声動作を必要とします。既存の同時音声生成パイプラインは主に人間中心です。モーションは最初に SMPL-X などの人体表現で生成され、その後人型ロボットに再ターゲットされます。この研究では、このパラダイムにおける基本的な実施形態のギャップを特定します。つまり、人間の動作多様体と人型の実施形態の制約との間の不一致により、動作の伝達と物理的な実行中に実施形態の一貫性が損なわれるということです。広範な分析を通じて、リターゲットは粗い動きのセマンティクスを維持できるものの、動きの多様性を大幅に圧縮し、韻律と動きの同期を弱め、表現力豊かなヒューマノイドの動作を制限することを示しました。この問題に対処するために、我々はまず、リターゲティング中の運動学的実現可能性と音声と動作の時間的整合を共同で最適化する、韻律を保存するヒューマノイド動作キュレーションフレームワークである IK-EER を提案します。厳選されたロボットネイティブのモーションデータセットに基づいて、中間の人体の表現に依存せずに音声から実行可能なヒューマノイド関節の軌道を直接予測する、実施形態を意識した同時音声モーション生成フレームワークである PhysDrift をさらに紹介します。従来の人間中心のパイプラインとは異なり、PhysDrift は、ロボットの動作ダイナミクスを安定させるために物理的正則化を組み込みながら、トレーニングと推論の両方を通じて実施形態の一貫性を維持します。広範な実験と現実世界のヒューマノイド展開により、実施形態を意識したロボットネイティブ生成により、音声と動作の整合性、物理的な妥当性、動作の滑らかさ、推論効率、およびリアルタイムのインタラクション能力が大幅に向上することが実証されました。

原文 (English)

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generation pipelines are predominantly human-centric: motions are first generated in human-body representations such as SMPL-X and subsequently retargeted to humanoid robots. In this work, we identify a fundamental embodiment gap in this paradigm, where the mismatch between human motion manifolds and humanoid embodiment constraints disrupts embodiment consistency during motion transfer and physical execution. Through extensive analysis, we show that although retargeting can preserve coarse motion semantics, it significantly compresses motion diversity and weakens prosody-motion synchronization, limiting expressive humanoid behaviors. To address this problem, we first propose IK-EER, a prosody-preserving humanoid motion curation framework that jointly optimizes kinematic feasibility and speech-motion temporal alignment during retargeting. Building upon the curated robot-native motion dataset, we further introduce PhysDrift, an embodiment-aware co-speech motion generation framework that directly predicts executable humanoid joint trajectories from speech without relying on intermediate human-body representations. Unlike conventional human-centric pipelines, PhysDrift maintains embodiment consistency throughout both training and inference while incorporating physical regularization to stabilize robot motion dynamics. Extensive experiments and real-world humanoid deployment demonstrate that embodiment-aware robot-native generation substantially improves speech-motion alignment, physical plausibility, motion smoothness, inference efficiency, and real-time interaction capability.

13:00 JSTエージェント

自動組み込みダイアログ拡張による DialNav の進化

物理的な相互作用が可能な実体エージェントの場合、安全性と有効性の両方を確保するには、対話を作成して理解する能力が不可欠です。 DialNav~\cite{han2025dialnav} はダイアログの全体的な評価、つまりフォトリアリスティックな屋内ナビゲーションの実行ループのためのフレームワークを提供しますが、そのパフォーマンスはトレーニング データ (2K エピソード) の重大な不足によって依然として制限されています。これに対処するために、自動生成パイプラインを提案し、DialNav 用の 238,000 エピソードを含む大規模なトレーニング データセットである \textbf{RAINbow} データセットを構築します。当社のパイプラインは、既存の VLN データセットをマルチターン ダイアログに変換し、コスト効率の高い高品質のデータセットを作成します。次に、データの可能性を最大限に引き出すための 2 つの追加の補完的な進歩を導入します。(1) ダイナミック ダイアログ ナビゲーション ループとナビゲーション トレーニングを調整するナビゲーション トレーニング スキームであるデュアル戦略トレーニング、および (2) VLN 知識を活用するローカリゼーション モデル。これらの補完的なソリューションを組み合わせることで、私たちのモデルは \textbf{Val Seen} (58.24, \textbf{+89\%}) と \textbf{Val Unseen} (29.05, \textbf{+100\%}) の両方の分割で成功率のベースラインを大幅に上回り、新たな最先端技術を確立しました。

原文 (English)

Advancing DialNav through Automatic Embodied Dialog Augmentation

For embodied agents capable of physical interaction, the capability to create and understand dialog is crucial to ensure both safety and effectiveness. While DialNav~\cite{han2025dialnav} provides a framework for holistic evaluation of the dialog--execution loop in photorealistic indoor navigation, its performance remains limited by a critical scarcity of training data (2K episodes). To address this, we propose an automatic generation pipeline, and construct the \textbf{RAINbow} dataset, a large-scale training dataset with 238K episodes for DialNav. Our pipeline converts existing VLN datasets into multi-turn dialog and creates cost-efficient and high-quality dataset. Then, we introduce two additional complementary advances to unlock the data's full potential: (1) Dual-Strategy Training, a navigation training scheme to align the navigation training with the dynamic dialog-navigation loop, and (2) a localization model that leverages VLN knowledge. By combining these complementary solutions, our model substantially outperforms the baseline in success rate on both \textbf{Val Seen} (58.24, \textbf{+89\%}) and \textbf{Val Unseen} (29.05, \textbf{+100\%}) splits, establishing a new state of the art.

13:00 JSTエージェントロボティクス

ENPIRE: 現実世界でのエージェント ロボット ポリシーの自己改善

現実世界で器用なロボット操作を実現するには、人間の監視とアルゴリズム工学に大きく依存しており、これが一般的な物理的知性の追求において中心的なボトルネックとなります。新興のコーディング エージェントはアルゴリズム検索を自動化するコードを生成できますが、その成功は依然としてデジタル環境に限定されています。私たちは、ロボット研究を自動化するために欠けている抽象化は、現実世界のポリシー改善のための反復可能なフィードバック ループであると推測します。つまり、シーンをリセットし、ポリシーを実行し、結果を検証し、次の反復を改良するというものです。このギャップを埋めるために、コーディング エージェント用のハーネス フレームワークである ENPIRE を導入します。このフレームワークは、4 つのコア モジュールでこの物理フィードバック ルーチンをインスタンス化します。1 つは自動リセットと検証のための環境モジュール (EN)、ポリシーの改良を開始するポリシー改善モジュール (PI)、1 つまたは複数の物理ロボットを並行して動作させてポリシーを評価するロールアウト モジュール (R)、およびコーディング エージェントがログを分析し、文献を参照し、トレーニング インフラストラクチャと障害モードに対処するためのアルゴリズム コードを改善する進化モジュール (E) です。この閉ループ システムは、現実世界の操作学習を制御可能な最適化手順に変換し、人間の労力を最小限に抑えながら、トレーニング レシピとエージェントのバリエーション全体で公平なアブレーションを可能にします。 ENPIRE を活用することで、フロンティア コーディング エージェントはポリシーを自律的にトレーニングして、ピン ボックスの整理、結束バンドの締め付け、工具の使用などの困難で器用な操作タスクで 99% の成功率を達成できます。ロボット フリートにエージェント チームを派遣すると、このプロセスがさらに加速します。私たちの結果は、物理世界で自律的に進歩するロボット工学にコーディング エージェントを展開するための実用的でスケーラブルな道筋を示唆しています。

原文 (English)

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world.

13:00 JSTエージェント

具現化された世界モデルのエージェントとしての報酬

RL は世界モデルを改良するための有望なツールとなっていますが、既存の手法は主にトレーニング分布付近の保守的なロールアウトに依存しており、探索、行動の多様性、より豊富な動的発見が制限されています。この作品では、私たちはこの保守的なパラダイムに挑戦します。私たちは、核となる制限は探査そのものではなく、より広範な探査をサポートする信頼できる検証戦略が欠如していることにあると主張します。信頼できる検証がなければ、拡張された探査は報酬ハッキングの影響を非常に受けやすくなり、真の改善が達成されずにポリシーが不完全な報酬を悪用することになります。この動機を評価するために、私たちは具体化された世界モデルでメソッドをインスタンス化します。そこでは、物理的な妥当性とタスクの完了が、複雑なダイナミクスの下でスケーラブルな RL の厳密なテストベッドを提供します。検証面では、生成された動作をアクティブに評価して堅牢な報酬シグナルを提供し、配布の変化下での報酬ハッキングを軽減するエージェント報酬フレームワークである Reward as an Agent を導入します。探査面では、DynDiff-GRPO を通じて動的認識ロールアウト多様化を導入します。これにより、行動空間探査が明示的に拡張され、軌道を多様化し、国家活動の範囲を拡大し、保守的なロールアウト体制を超えてより豊かに具体化された行動が奨励されます。 Reward as an Agent を DynDiff-GRPO と統合することで、大幅に多様化したサンプリングを備えたより信頼性の高い報酬基盤で RL を実現し、報酬ハッキングを効果的に緩和しながら、複数のオープンソースの世界モデル全体で大幅な精度の向上を実現します。これにより、堅牢な検証に基づいて広範な探索を適切に拡張できることを実証します。

原文 (English)

Reward as An Agent for Embodied World Models

While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distribution, limiting exploration, behavioral diversity, and richer dynamic discovery. In this work, we challenge this conservative paradigm. We argue that the core limitation is not exploration itself, but the lack of reliable verification strategies to support broader exploration. Without reliable verification, expanded exploration becomes highly susceptible to reward hacking, where policies exploit imperfect rewards without achieving genuine improvement. To evaluate this motivation, we instantiate our method in embodied world models, where physical plausibility, and task completion provide a rigorous testbed for scalable RL under complex dynamics. On the verification side, we introduce Reward as an Agent, an agentic reward framework that actively evaluates generated behaviors to provide robust reward signals and mitigate reward hacking under distribution shifts. On the exploration side, we introduce Dynamic-Aware Rollout Diversification through DynDiff-GRPO, which explicitly expands action-space exploration to diversify trajectories, broaden state-action coverage, and encourage richer embodied behaviors beyond conservative rollout regimes. By unifying Reward as an Agent with DynDiff-GRPO, we enable RL on a more reliable reward foundation with substantially diversified sampling, effectively mitigating reward hacking while yielding significant accuracy gains across multiple open-source world models, thereby demonstrating that broader exploration can scale successfully when grounded in robust verification.

13:00 JSTエージェント

大規模なエンタープライズ AI 向けの自律的なイベント駆動型マルチエージェント オーケストレーション

Enterprise AI は、専門エージェント全体での継続的なイベントの監視、検出、アクションを目指していますが、既存のマルチエージェント システムは主に個別の要求と応答のワークフローを想定しており、エンタープライズ規模ではまだ十分に検討されていません。当社では、ペルソナ (エージェント 10 人未満)、部門 (20 ~ 80 人)、エンタープライズ (200 人) の規模にわたる 208 の本番環境由来のエンタープライズ シナリオにわたって、DAG の計画、実行、および反応を評価し、優先順位の推論、関連イベントのマージ、プリエンプションによる継続的な運用のためのタスク マネージャーを導入しています。結果は、タスクの複雑さではなくスケールがオーケストレーションのパフォーマンスを支配していることを示しています。どちらのアーキテクチャも小規模では良好にパフォーマンスしますが、エンタープライズ規模ではエージェント検出のノイズが主なボトルネックとなり、単純なタスクの方が複雑なタスクよりも大幅に低下するため、パフォーマンスが低下します。 DAG の計画と実行は、より小規模な規模ではより高い精度と構造化された並列化を提供しますが、エンタープライズ規模ではオーバーヘッドが大きくなり、さらに悪化します。 ReAct は、障害を段階的に処理することでより堅牢になります。タスク マネージャーは、エンタープライズ規模で高優先度のキューの遅延を 14 ~ 75% 削減し、関連イベントの正確性を 20 パーセント以上向上させます。

原文 (English)

Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale

Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent systems largely assume discrete request-response workflows and remain underexplored at enterprise scale. We evaluate DAG Plan and Execute and ReAct across 208 production-derived enterprise scenarios spanning Persona (<10 agents), Department (20-80), and Enterprise (200) scales, and introduce a Task Manager for continuous operation via priority inference, related-event merging, and preemption. Results show that scale, not task complexity, dominates orchestration performance: both architectures perform well at small scale but degrade at enterprise scale as agent discovery noise becomes the primary bottleneck, with simple tasks degrading more sharply than complex ones. DAG Plan and Execute offers higher precision and structured parallelization at smaller scales, but its higher overhead worsens at enterprise scale; ReAct is more robust by handling failures incrementally. The Task Manager reduces high-priority queue latency by 14-75% and improves related-event correctness by over 20 percentage points at enterprise scale.

13:00 JST研究/論文DeepSeek

リーンによる定理証明のためのプロセス検証済み強化学習

検証可能な報酬からの強化学習 (RLVR) は通常、単一のバイナリ検証信号に依存していましたが、形式推論における記号証明アシスタントは、豊富できめの細かい構造化されたフィードバックを提供します。構造化されたプロセスと非構造化された報酬の間にあるこのギャップは、密度が高く健全なフィードバックの重要性を浮き彫りにしています。この研究では、リーン証明アシスタント自体が象徴的なプロセスのオラクルとして機能し、トレーニング中に結果レベルと詳細な戦術レベルの両方の検証済みフィードバックを提供できることを実証します。証明の試みは戦術シーケンスに解析され、リーンの詳細な説明により、局所的に健全なステップと最も初期に失敗したステップの両方がマークされ、型理論に根ざした密な検証者に基づいた信用シグナルが得られます。これらの構造化された報酬を、結果レベルとプロセスレベルの利点のバランスをとるファーストエラー伝播およびファーストトークンクレジット手法を備えた GRPO スタイルの強化学習目標に組み込みます。 STP-Lean と DeepSeek-Prover-V1.5 を使用した実験では、ほとんどの設定で戦術レベルの監視が結果のみのベースラインを上回り、MiniF2F や ProofNet などのベンチマークで改善が見られることが示されています。私たちの研究は、経験的な利益を超えて、より広い視点を強調しています。記号的証明アシスタントは、評価時の検証者であるだけでなく、トレーニング中にプロセスレベルの報酬のオラクルとしても機能する可能性があります。これにより、言語モデルのスケーラビリティと形式的推論のための記号検証の信頼性を組み合わせた強化学習フレームワークへの道が開かれます。

原文 (English)

Process-Verified Reinforcement Learning for Theorem Proving via Lean

While reinforcement learning from verifiable rewards (RLVR) typically has relied on a single binary verification signal, symbolic proof assistants in formal reasoning offer rich, fine-grained structured feedback. This gap between structured processes and unstructured rewards highlights the importance of feedback that is both dense and sound. In this work, we demonstrate that the Lean proof assistant itself can serve as a symbolic process oracle, supplying both outcome-level and fine-grained tactic-level verified feedback during training. Proof attempts are parsed into tactic sequences, and Lean's elaboration marks both locally sound steps and the earliest failing step, yielding dense, verifier-grounded credit signals rooted in type theory. We incorporate these structured rewards into a GRPO-style reinforcement learning objective with first-error propagation and first-token credit methods that balances outcome- and process-level advantages. Experiments with STP-Lean and DeepSeek-Prover-V1.5 show that tactic-level supervision outperforms outcome-only baselines in most settings, delivering improvements on benchmarks such as MiniF2F and ProofNet. Beyond empirical gains, our study highlights a broader perspective: symbolic proof assistants are not only verifiers at evaluation time, but can also act as process-level reward oracles during training. This opens a path toward reinforcement learning frameworks that combine the scalability of language models with the reliability of symbolic verification for formal reasoning.

13:00 JST研究/論文

フローベースの生成モデルによる残差空間の進化的最適化

生成手法によるデータ編集には通常、微分可能な目的と勾配ベースの検索が必要です。ただし、これらの前提はフローベースの設定では崩れます。フローベースの設定では、前方統合と後方統合を通じて編集が実行され、微分不可能な目標やブラックボックスの目標が含まれることがよくあります。フローベースの生成編集と進化的アルゴリズムを組み合わせることによってこのギャップに対処する、モデルに依存しないフレームワークである残差空間進化的最適化を紹介します。条件付きフロー マッチング (CFM) がインスタンス固有の残差から条件制御の要素を解きほぐすことができるという観察に基づいて、私たちのフレームワークは残差空間で直接動作し、2 つの相補的な検索レジームを分離します。自己受粉は特徴を保持した残差の洗練を通じて局所的な活用を実行し、他家受粉は異種サンプル間で残差を再結合することでより広範な探索を促進します。概念実証として、反事実生成のベンチマーク データセットである MorphoMNIST と結晶データで検証し、この探査 - 悪用分解がターゲットの配置、インスタンスの保存、多様性のバランスをとるための有用なメカニズムを提供し、画像を超えて現実世界の科学領域にまで及ぶことを実証しました。

原文 (English)

Residual-Space Evolutionary Optimization via Flow-based Generative Models

Data editing with generative methods typically requires differentiable objectives and gradient-based search. However, these assumptions break down in flow-based settings, where edits are performed through forward and backward integration and often involve non-differentiable or black-box objectives. We introduce residual-space evolutionary optimization, a model-agnostic framework that addresses this gap by combining flow-based generative editing with evolutionary algorithms. Building on the observation that conditional flow matching (CFM) can disentangle condition-controlled factors from instance-specific residuals, our framework directly operates in residual space and separates two complementary search regimes: self-pollination performs local exploitation through feature-preserving residual refinement, and cross-pollination promotes broader exploration by recombining residuals across heterogeneous samples. As a proof of concept, we validate on MorphoMNIST, a benchmark dataset for counterfactual generation, and on crystal data, demonstrating that this exploration--exploitation decomposition provides a useful mechanism for balancing target alignment, instance preservation, and diversity, and extends beyond images to real-world scientific domains.

13:00 JST研究/論文

マルチヘッド アテンション ベースの特徴抽出器とソフト アクター クリティカルの統合による積層造形における空隙率予測とプロセス パラメーターの最適化

積層造形プロセスの最適化には、気孔率などの欠陥を最小限に抑えるための正確なパラメータ制御が必要です。離散アクション空間を使用する従来の強化学習 (RL) アプローチは、収束が遅いことと局所最適化の影響を受けやすいという問題があり、高精度の製造タスクに対する有効性が制限されます。この研究では、マルチヘッド アテンション メカニズムと Soft Actor-Critic (SAC) アルゴリズムを統合する新しいアーキテクチャと組み合わせた連続アクション スペースを採用することで、これらの制限に対処しています。アテンションベースの特徴抽出機能は、低次元の入力特徴の微妙な変化を捕捉するエージェントの能力を強化し、極小値で値空間をナビゲートするためのより効果的な探索と活用のバランスを可能にします。レーザー粉体層融合における気孔率予測とプロセスパラメータの最適化に関するアプローチを検証し、DQN、PPO、TD3、バニラ SAC などの標準的な RL 法と比較して、より速い収束とより高い最終報酬値を実証します。提案された方法論は、14 エピソード以内で 322.79 の収束値を達成し、トレーニング全体を通じて安定性を維持しながら、既存のアプローチを上回ります。

原文 (English)

Multi-Head Attention-Based Feature Extractor Integration with Soft Actor-Critic for Porosity Prediction and Process Parameter Optimization in Additive Manufacturing

Additive manufacturing process optimization requires precise parameter control to minimize defects such as porosity. Traditional reinforcement learning (RL) approaches using discrete action spaces suffer from slow convergence and susceptibility to local optima, limiting their effectiveness for high-precision manufacturing tasks. This study addresses these limitations by employing a continuous action space combined with a novel architecture that integrates a multi-head attention mechanism with the Soft Actor-Critic (SAC) algorithm. The attention-based feature extractor enhances the agent's ability to capture subtle variations in low-dimensional input features, enabling more effective exploration-exploitation balance for navigating value spaces with local minima. We validate our approach on porosity prediction and process parameter optimization in laser powder bed fusion, demonstrating faster convergence and higher final reward values compared to standard RL methods including DQN, PPO, TD3, and vanilla SAC. The proposed methodology achieves a convergence value of 322.79 within 14 episodes, outperforming existing approaches while maintaining stability throughout training.

13:00 JSTエージェント研究/論文

ScaffoldAgent: オープンエンドの深層研究のためのユーティリティガイドによる動的アウトライン最適化

オープンエンド型ディープリサーチ (OEDR) では、システムが複数ラウンドの検索を通じて知識を取得し、一貫した長い形式のレポートを生成する必要があります。アウトラインは、検索、証拠の整理、生成を調整する構造的な足場として中心的な役割を果たします。しかし、既存の方法では、書く前にアウトラインを修正するか、ローカルヒューリスティックでアウトラインを調整するため、継続的な情報蓄積の下で足場のドリフトが発生したり、アウトラインの変更を評価するためのフィードバックが遅れたりします。我々は、OEDR 向けのユーティリティ主導の動的アウトライン最適化フレームワークである ScaffoldAgent を提案します。 ScaffoldAgent モデルは、展開、縮小、改訂の 3 つの操作による構造化された意思決定プロセスとして進化を概説し、レポート スキャフォールドの制御された更新を可能にします。さらに、取得利得、構造的一貫性、試行生成の品質から各アウトライン操作の下流の価値を推定するユーティリティ主導のフィードバック メカニズムが導入されています。結果として得られるユーティリティ信号は、ノードの選択、操作のスケジュール設定、および推論中の終了をガイドします。 DeepResearch Bench と DeepResearch Gym での実験では、ScaffoldAgent が既存のディープ リサーチ エージェントに比べて長文レポートの生成と事実に基づく根拠を一貫して向上させていることが示されています。

原文 (English)

ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research

Open-ended deep research (OEDR) requires systems to acquire knowledge through multi-round retrieval and generate coherent long-form reports. The outline plays a central role as a structural scaffold that coordinates retrieval, evidence organization, and generation. However, existing methods either fix the outline before writing or refine it with local heuristics, leading to scaffold drift under continuous information accumulation and delayed feedback for evaluating outline modifications. We propose ScaffoldAgent, a utility-guided dynamic outline optimization framework for OEDR. ScaffoldAgent models outline evolution as a structured decision process with three operations: Expansion, Contraction, and Revision, enabling controlled updates to the report scaffold. It further introduces a utility-guided feedback mechanism that estimates the downstream value of each outline operation from retrieval gain, structural coherence, and trial-generation quality. The resulting utility signal guides node selection, operation scheduling, and termination during inference. Experiments on DeepResearch Bench and DeepResearch Gym show that ScaffoldAgent consistently improves long-form report generation and factual grounding over existing deep research agents.

13:00 JSTLLM/生成AI

促すことを学ぶ: アダプティブ LLM ベースの高校個別指導で生徒の関与を向上

LLM は教育を個別化できますが、現在の静的即時個別指導システムは多様な学問分野に適応するのに苦労しています。私たちは、生のトランスクリプトから抽出した 14 の教育的特徴 (家庭教師の足場、生徒の理解など) に基づいて、主題を意識したプロンプトを備えたシステムを開発およびテストします。まずシミュレーション環境でプロンプト ルーティング モデルをトレーニングし、次にそれを実際の高校生のオンライン適応に展開します。シミュレーション ベンチマークでは、ルーターが 2 つの静的ベースライン ($0.694$ 対 $0.647$ および $0.64$、$p<0.001$) を上回るパフォーマンスを示しています。 A/B テスト (359 人の生徒からの $N=656$ の会話) では、モデルが分析学習戦略から足場学習戦略に切り替わる、シミュレーションから現実への移行が示されています。私たちの適応型プロンプト選択メカニズムは、指導効率を向上させ、教育の質を維持し、インタラクションを約 3 ターン ($p=0.007$) 削減します。グリーディ ルーターはベースラインと同等のエクササイズ コンバージョン率 ($19.1\%$ 対 $19.6\%$) を達成しますが、戦略をサンプリングするストキャスティック ルーターはより高いコンバージョン率 ($28.1\%$) をもたらします。

原文 (English)

Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring

LLMs can personalize education, although current static-prompt tutoring systems struggle to adapt to diverse academic disciplines. We develop and test a system with subject-aware prompting, based on 14 pedagogical features (e.g., tutor scaffolding, student understanding) extracted from raw transcripts. We first train a prompt routing model in a simulation environment, and then deploy it for online adaptation with actual high-school students. The simulation benchmark shows the router outperforming two static baselines ($0.694$ vs. $0.647$ and $0.64$, $p<0.001$). A/B testing ($N=656$ conversations from 359 students) shows sim-to-real transfer where the model switches from analytical to scaffolding learning strategies. Our adaptive prompt selection mechanism improves instructional efficiency, maintains pedagogical quality and reduces interactions by around 3 turns ($p=0.007$). While a greedy router achieves a comparable exercise conversion rate with the baseline ($19.1\%$ vs. $19.6\%$), a stochastic router that samples strategies leads to a higher conversion rate ($28.1\%$).

13:00 JSTエージェント研究/論文

RACL: 継続的メタヒューリスティック学習のための推論エージェント制御層

このペーパーでは、メタヒューリスティックのための推論エージェント制御層である RACL を紹介します。 RACL は、既存のオプティマイザーの上に推論エージェントを配置します。エージェントはオプティマイザを置き換えたり、ビジネス制約を変更したりしません。代わりに、操作メモリの観察、過去の動作の推論、限定された仮説の策定、介入のテスト、結果の評価、ガードレールの適用、有用なポリシーの統合、およびその決定の説明によって、オプティマイザの内部検索動作を制御します。この実験では車両ルーティングをテストベッドとして使用しますが、貢献するのは新しいルーティング ソルバー、特定の ALNS 構成、または特定の一連のルーティング ルールではありません。貢献するのは RACL メソッドです。これは、推論エージェントがメタヒューリスティックのアルゴリズム制御ルールを発見、検証、統合、説明するための方法です。現在の実験設定では、RACL は、21 の実行可能なケースのうち 21 で動作メモリ ポリシーを改善または結合し、21 の実行可能なケースのうち 18 で非推論の停滞トリガー ポリシーを改善または結合し、平均 RACL 対 STP コスト デルタは -0.641% でした。 Sevilla-9/10 ランタイム サンプルでは、​​RACL は重大な計算オーバーヘッドを示さずに、平均コストを固定と比較して -8.337%、STP と比較して -1.605% 改善しました。概念実証中、Codex は、実行を観察し、ログを解釈し、ライブの制限付き介入を提案するループ内推論エージェントとして使用されました。政策代理はその後、定量的評価を再現可能にする目的でのみ使用されました。

原文 (English)

RACL: Reasoning-Agent Control Layers for Continuous Metaheuristic Learning

This paper introduces RACL, a Reasoning-Agent Control Layer for metaheuristics. RACL places a reasoning agent above an existing optimizer. The agent does not replace the optimizer and does not modify business constraints. Instead, it controls the optimizer's internal search behavior by observing operational memory, reasoning over past behavior, formulating bounded hypotheses, testing interventions, evaluating outcomes, applying guardrails, consolidating useful policies and explaining its decisions. The experiment uses vehicle routing as a testbed, but the contribution is not a new routing solver, a particular ALNS configuration or a specific set of routing rules. The contribution is the RACL method: a way for a reasoning agent to discover, validate, consolidate and explain algorithmic control rules for a metaheuristic. In the current experimental setting, RACL improves or ties the Operational Memory Policy in 21 of 21 feasible cases and improves or ties a non-reasoning Stagnation-Triggered Policy in 18 of 21 feasible cases, with an average RACL vs STP cost delta of -0.641%. In the Sevilla-9/10 runtime sample, RACL improves average cost by -8.337% versus Fixed and -1.605% versus STP without showing material computational overhead. During the proof-of-concept, Codex was used as an in-the-loop reasoning agent observing executions, interpreting logs and proposing live bounded interventions. The policy proxy was later used only to make quantitative evaluation reproducible.

13:00 JSTLLM/生成AI研究/論文

BIM-Edit: IFC ベースのビルディング インフォメーション モデリングのための大規模言語モデルのベンチマーク

大規模言語モデル (LLM) は、テキストの指示から設計アーティファクトを生成するために、コンピュータ支援設計 (CAD) にますます適用されています。エンジニアリングの実践では、これには新しいジオメトリを作成するだけではなく、モデルが既存のシーンを理解し、正しく編集し、セマンティクスと関係を保持する必要もあります。ただし、多くの CAD ベンチマークは、既存のモデルを編集するのではなく、新しいモデルを作成することに重点を置き、主に幾何学的正確さを評価します。 Industry Foundation Classes (IFC) 形式で表される Building Information Model (BIM) の自然言語編集に関する LLM を評価するためのベンチマークである BIM-Edit を紹介します。 BIM は、建築モデルがジオメトリをセマンティックおよびリレーショナル構造とともにエンコードするため、困難なテストベッドを提供します。 BIM-Edit には、11 の現実的な建築モデルと 36 の合成シーンにわたる 324 の編集タスクが含まれています。タスクは 3 つの命令カテゴリ (直接、空間、トポロジカル) を使用して表現され、明示的な編集とシーンに基づいた編集の両方をカバーします。私たちは、幾何学的精度、意味論的妥当性、トポロジー的一貫性という 3 つの次元に沿って出力を評価します。評価された LLM 全体で、最もパフォーマンスの高いモデルは、3 つの指標全体で 49.5% の平均スコアしか達成できず、タスクの 3.4% を超える問題を完全に解決するモデルはありません。これらの結果は、現在の LLM 機能と構造化エンジニアリング設計ワークフローの要件との間に大きなギャップがあることを示しています。

原文 (English)

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engineering practice, this requires more than creating new geometry, models must also understand existing scenes, edit them correctly, and preserve semantics and relations. However, many CAD benchmarks focus on creating new models rather than editing existing ones, and mostly evaluate geometric correctness. We introduce BIM-Edit, a benchmark for evaluating LLMs on natural-language editing of Building Information Models (BIM) represented in the Industry Foundation Classes (IFC) format. BIM provides a challenging testbed because building models encode geometry together with semantic and relational structure. BIM-Edit contains 324 editing tasks spanning 11 realistic building models and 36 synthetic scenes. Tasks are expressed using three instruction categories - direct, spatial, and topological - covering both explicit and scene-grounded edits. We evaluate outputs along three dimensions: geometric accuracy, semantic validity, and topological consistency. Across evaluated LLMs, the best-performing model achieves only 49.5% average score across the three metrics, and no model fully solves more than 3.4% of tasks. These results demonstrate a substantial gap between current LLM capabilities and the requirements of structured engineering design workflows.

13:00 JST研究/論文

一般化されたPINN向けのモジュールフリーの競合回避トレーニング

物理情報に基づいたニューラル ネットワーク (PINN) は、微分可能な目的に物理法則を埋め込むことで偏微分方程式を解くための強力なフレームワークになりました。進歩にもかかわらず、PINN のトレーニングは脆弱なままです。最近の競合回避最適化スキームは残差損失と境界損失の間の勾配干渉を軽減しますが、モデルの容量が増加するにつれてその有効性が低下することを示しました。この論文では、過パラメータ化されたネットワークが機能モジュール化され、目的を越えた相互作用を抑制し、パレート定常点への収束を妨げるタスク専用モジュールに自己分割する、容量に起因する故障モードを特定します。この問題に対処するために、私たちは新しいフレームワークである Modular-Sparsity Synchronization (ModSync) を提案します。これは、相互作用を促進する経路を維持しながらタスク排他的な接続にペナルティを与えることにより、構造の最適化を競合回避トレーニングに統合します。さまざまな PDE ベンチマークにわたる広範な実験により、ModSync が容量主導の障害を一貫して防止し、堅牢な目的間の結合を維持し、最先端の精度を達成できることが実証されました。コードは \url{https://github.com/heejokong/ModSync} で入手できます。

原文 (English)

Modularity-Free Conflict-Averse Training for Generalized PINNs

Physics-informed neural networks (PINNs) have become a powerful framework for solving PDEs by embedding physical laws into differentiable objectives. Despite their advances, training PINNs remains fragile: recent conflict-averse optimization schemes alleviate gradient interference between residual and boundary losses, but we show that their effectiveness deteriorates as model capacity increases. In this paper, we identify a capacity-induced failure mode, where overparameterized networks undergo functional modularity, self-partitioning into task-exclusive modules that suppress cross-objective interaction and hinder convergence toward Pareto-stationary points. To address this issue, we propose a novel framework, Modular-Sparsity Synchronization (ModSync), which integrates structural optimization into conflict-averse training by penalizing task-exclusive connections while preserving interaction-promoting pathways. Extensive experiments across diverse PDE benchmarks demonstrate that ModSync consistently prevents capacity-driven failures, sustains robust cross-objective coupling, and achieves state-of-the-art accuracy. Codes are available at \url{https://github.com/heejokong/ModSync}.

13:00 JST研究/論文

ハイパーグラフ推論に基づく暗黙的なセマンティックを意識した通信

意味を意識した通信は、次世代通信システムの革新的なパラダイムとして登場し、基本的な目標をビットレベルのシンボルの送信から、情報の意味を確実に回復して理解することに移行しました。これまでの研究では、ソースメッセージの意味内容をグラフベースの構造として表現すると、受信側での通信効率と意味推論の精度が大幅に向上することが実証されています。ただし、既存のソリューションは通常、ペアごとの関係のみをキャプチャするグラフを採用しているため、グループ相互作用、複数エンティティの関連付け、複雑な関係コンテキストなど、現実世界のシナリオで一般的に観察される高次の暗黙的な相関関係が無視されています。この制限により、意味論的な表現力が低下し、特にノイズが多いまたは破損したチャネル条件下では、意味論的な推論があいまいになり、パフォーマンスが低下しやすくなります。これらの問題に対処するために、この論文では、ハイパーグラフを利用して意味論的知識エンティティ間の複雑な複数エンティティ関係を表現する、新しいハイパーグラフベースの暗黙的意味論的推論フレームワーク HISR を提案します。 HISR では、エンティティとそれに関連する高次の関係が、個別の関係コンテキストに合わせて調整された専用の意味論的部分空間にマッピングされます。この設計は、多様な意味論的相互作用を解きほぐして、従来のグラフ埋め込み手法によく見られる過度の平滑化効果を軽減するだけでなく、送信中に部分的な情報損失が発生した場合でも、堅牢な意味論的推論を可能にします。数値結果は、提案された HISR が、最先端のベンチマークと比較して、暗黙的な意味解釈の精度で最大 36.6% の向上を達成することを示しています。

原文 (English)

Implicit Semantic-Aware Communication Based on Hypergraph Reasoning

Semantic-aware communication has emerged as a transformative paradigm for next-generation communication systems, shifting the fundamental goal from transmitting bit-level symbols to reliably recovering and understanding the semantic meaning of information. Previous studies have demonstrated that representing the semantic content of source messages as graph-based structures can significantly improve communication efficiency and the accuracy of semantic inference at the receiver. However, existing solutions typically employ graphs that capture only pairwise relationships, thereby neglecting higher-order implicit correlations commonly observed in real-world scenarios, such as group interactions, multi-entity associations, and complex relational contexts. This limitation reduces semantic expressiveness and makes semantic inference susceptible to ambiguity and performance degradation, particularly under noisy or corrupted channel conditions. To address these issues, this paper proposes a novel hypergraph-based implicit semantic reasoning framework, HISR, which leverages hypergraphs to represent complex multi-entity relationships among semantic knowledge entities. In HISR, entities and their associated higher-order relations are mapped into dedicated semantic subspaces tailored to distinct relational contexts. This design not only disentangles diverse semantic interactions to mitigate the over-smoothing effects commonly found in traditional graph embedding methods but also enables robust semantic inference even when partial information loss occurs during transmission. Numerical results show that the proposed HISR achieves up to a 36.6% improvement in implicit semantic interpretation accuracy over the state-of-the-art benchmarks.

13:00 JSTLLM/生成AI

大規模な言語モデルの見かけの心理学的プロファイルは主に測定結果である

人間用に設計された心理学的機器は、ユーザビリティ、安全性評価、および研究への人間の参加者の代理としての使用に影響を与える安定した心理的プロファイルを大規模言語モデル (LLM) に割り当てるために使用されることが増えています。正式な心理測定フレームワークを使用して、これらのプロファイルの大部分が測定アーチファクトであることを示します。自己報告と行動課題にわたる一連の性格およびリスク選好の測定器を、大規模な人間の参照サンプルとともに 56 の指導調整 LLM に投与したところ、4 つの調査結果が報告されました。まず、モデル間の違いは、機器が対象とする特性によって決まるのではなく、指向性反応の偏り、つまりアイテムの内容に関係なく、スケールの一端または 1 つのラベル付きオプションに向かって反応する傾向によって決まります。分散分解では、モデル間の変動の 81 ~ 90% がこのバイアスによるものであると考えられますが、人間では 9 ~ 16% です。第二に、バイアスはモデルの能力とともに減少しますが、それによって除去されるわけではありません。第三に、特性ではなくバイアスが反応を促すため、機器の見かけの信頼性はほぼ完全にその応答の直交性によって予測されます。この直交性とは、特性とバイアスが逆の方向を向いている項目の割合を表す造語です。 4つ目は、使用するアイテムによってモデルの外観が変化し、アイテムの選択によって製造できることです。これらの結果は、LLM の見かけの心理的プロファイルは、モデル自体の特性ではなく、LLM を測定するために使用された機器のアーチファクトであることを示しています。人間の心理学から借用した手段が完全に直交することはほとんどなく、本質的に LLM の妥当性を欠いている可能性があるため、応答の直交性を中心とした専用の評価が必要です。

原文 (English)

Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles that affect their usability, safety assessment, and use as proxies for human participants in research. Using a formal psychometric framework, we show that these profiles are largely a measurement artifact. Administering a battery of personality and risk-preference instruments spanning self-reports and behavioral tasks to 56 instruction-tuned LLMs alongside large human reference samples, we report four findings. First, differences between models are driven not by the traits an instrument targets but by a directional response bias, a tendency to respond toward one end of the scale, or one labeled option, regardless of item content; a variance decomposition attributes 81-90% of between-model variation to this bias, against 9-16% in humans. Second, the bias declines with model capability but is not eliminated by it. Third, because bias rather than trait drives responding, an instrument's apparent reliability is almost entirely predicted by its response orthogonality, a term we coin for the proportion of items for which trait and bias point in opposite directions. Fourth, the profile a model appears to have shifts with the items used and can be manufactured through item selection. These results demonstrate that the apparent psychological profiles of LLMs are artifacts of the instrument used to measure them, not properties of the models themselves. As instruments borrowed from human psychology are rarely fully orthogonal and may inherently lack validity for LLMs, we call for dedicated assessments centered on response orthogonality.

13:00 JST研究/論文

精度を超えて: 予測モデルの論理的準拠性の測定

機械学習モデルは主に、ランキング品質、予測誤差、分類精度などの予測パフォーマンス メトリクスを通じて評価されます。これらのメトリックは、予測がグラウンド トゥルースとどの程度一致するかを効果的に定量化しますが、モデルの出力が事前定義された論理制約またはドメイン固有の制約を尊重しているかどうかは評価しません。ヘルスケア、金融、自律システムなど、一か八かのアプリケーションでは、論理的一貫性が予測精度と同じくらい重要になる可能性がありますが、この次元を捉えた標準的な指標はありません。ルール違反スコア (RVS) を導入します。これは、予測精度とは関係なく、予測モデルが特定の論理ルールのセットをどの程度尊重するかを定量化する補完的な評価指標です。 RVS は、ハード ルール (厳密な制約) とソフト ルール (統計的規則性) を別々に扱い、任意のデータセットおよびリレーショナル語彙で表現された任意の予測モデルで評価でき、ホーン ルール用に自動的に生成される SQL クエリを使用して計算できます。 RVS はモデルを評価するだけでなく、トレーニング データセットの論理的一貫性も評価し、不十分に定義されたルールを特定するのに役立ちます。 RVS は、ルールベース、埋め込みベース、ニューロシンボリック予測モデルを含む、ナレッジ グラフ リンク予測とリレーショナル回帰をカバーする 3 つのベンチマークで評価されます。私たちの結果は、同等の予測精度を達成する 2 つのモデルが、実質的に異なるレベルの論理準拠を示す可能性があることを示しており、標準的なメトリクスでは捉えることができないモデルの動作の違いが明らかになります。

原文 (English)

Beyond Accuracy: Measuring Logical Compliance of Predictive Models

Machine learning models are predominantly evaluated through predictive performance metrics such as ranking quality, prediction error, or classification accuracy. While these metrics effectively quantify how closely predictions match the ground truth, they do not assess whether model outputs respect predefined logical or domain-specific constraints. In high-stakes applications, including healthcare, finance, and autonomous systems, logical consistency can be as critical as predictive accuracy, yet no standard metric captures this dimension. We introduce the Rule Violation Score (RVS), a complementary evaluation metric that quantifies the extent to which a predictive model respects a given set of logical rules, independently of predictive accuracy. RVS treats hard rules (strict constraints) and soft rules (statistical regularities) differently, can be evaluated on any dataset and on any predictive model expressed over a relational vocabulary, and can be computed using SQL queries that are automatically generated for Horn rules. Beyond evaluating models, RVS can also evaluate the logical consistency of training datasets and help identify poorly defined rules. We evaluate RVS on three benchmarks covering knowledge graph link prediction and relational regression, including rule-based, embedding-based, and neuro-symbolic predictive models. Our results demonstrate that two models achieving comparable predictive accuracy can exhibit substantially different levels of logical compliance, revealing differences in model behavior that standard metrics fail to capture.

13:00 JST研究/論文

深層強化学習によるゲーム AI の強化

ビデオ ゲームへの没入度は、グラフィックス、オーディオ、ゲームの仕組みだけでなく、ゲーム内のキャラクターの品質にも左右されます。手作業でコード化されたシステムでは動作の複雑さを捉えるのが難しいため、信頼できるキャラクター、つまりゲーム AI を作成することは依然として大きな課題です。ゲーム AI は没入感とエンゲージメントの源です。ただし、ゲーム AI を作成する際の課題から生じる制限は、多くの場合、フラストレーションやゲーム内のリアリズムの幻想の破壊につながります。機械学習モデルの導入により、ゲーム内でより信頼性があり、本物で、共感できるキャラクターを作成するための扉が開かれます。約束されているのは、彼らがゲームとのインタラクションから、またはプレイヤーのデータから学び、真の人間らしい行動を身につけるということです。この論文では、将来的にはゲーム AI に対する強化学習のさらなる応用を想定しています。これを実現するには、現在の研究の制限により、ゲーム ジャンルを超えた広範な展開が不可能になります。したがって、ゲーム AI とゲーム開発に適した一連の要件を念頭に置いて、強化学習モデルをトレーニングするためのフレームワークを提案します。強化学習で拡張されたゲーム AI を使用したゲームの例を示し、最新のゲームにプレイヤー向けの機械学習エージェントを導入する実用性について説明します。さらに、これらの分野におけるボトルネックと困難な問題を特定し、ビデオゲーム業界のゲーム AI における機械学習の導入を加速する有望な研究の方向性を提供すると考えています。

原文 (English)

Augmenting Game AI with Deep Reinforcement Learning

Immersion in video games depends not only on graphics, audio, and game mechanics, but also on the quality of in-game characters. Producing believable characters, or game AI, remains a significant challenge as behavioral complexity is hard to capture with hand-coded systems. Game AI is a source of immersion and engagement; however, the limitations stemming from the challenges of creating game AI often lead to frustration and the breaking of the illusion of realism within the game. The introduction of machine learning models opens the door to creating more believable, authentic, and relatable characters in games. The promise is that they either learn from interacting with the game, or from player data, to develop true human-like behavior. In this paper, we envision more applications of reinforcement learning for game AI in the future. For this to materialize, current research limitations are prohibitive to broad deployment across game genres. Therefore, we propose a framework for training reinforcement learning models with a set of requirements in mind that are suited towards game AI and game development. We present examples of games with reinforcement learning-augmented game AI and describe the practicalities of deploying player-facing machine learning agents in modern games. Furthermore, we identify bottlenecks and hard problems in these areas, which we believe offer promising research directions to accelerate the adoption of machine learning in game AI for the video game industry.

13:00 JSTLLM/生成AI研究/論文

QMFOL: 定量化可能なモナディック一次論理テスト ケース生成による大規模言語モデル推論のベンチマーク

大規模言語モデル (LLM) は、推論、特に一か八かの意思決定に重要な演繹的推論において大きな進歩を遂げました。モデルが改善されるにつれて、評価ベンチマークもそれに合わせて進化する必要があります。しかし、既存のベンチマークには論理的な複雑さに対するきめ細かい制御が欠けており、意味論的な多様性と論理的な一貫性のバランスを取るのに苦労しています。これらの問題に対処するために、定量化可能で制御可能な複雑さを備えたモナディック一次論理推論タスクを生成するための自動フレームワークである QMFOL を提案します。論理積パターンと論理和パターンを使用して正式な論理構造を構築し、推論の深さ、幅、ラベルの種類、および注意をそらす要素を正確に制御できるようにします。これらの構造は、LLM を介して自然言語に翻訳され、外部証明者を使用した往復検証を通じて論理的一貫性が保証されます。私たちのフレームワークに基づいて、さまざまな論理的およびセマンティックな次元にわたる 960 の構成を持つ 2880 のインスタンスで構成されるベンチマークである QMFOLBench を構築します。 6 つの大規模推論モデル (LRM) と 2 つの LLM を評価したところ、論理的な複雑さが増すにつれてパフォーマンスが低下し、計算オーバーヘッドが増加することがわかりました。モデルは、False または Unknown のタスクよりも True のラベルが付けられたタスクの方がパフォーマンスが高く、セマンティックの変動に対して敏感です。全体として、QMFOL は、制御可能な複雑さを備えた演繹的推論ベンチマークを構築するためのスケーラブルで信頼性の高いアプローチを提供し、最新の言語モデルにおける推論機能のより正確な評価を可能にします。

原文 (English)

QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation

Large Language Models (LLMs) have made significant progress in reasoning, particularly in deductive reasoning, which is crucial for high-stakes decision-making. As models improve, evaluation benchmarks should evolve to keep pace. However, existing benchmarks lack fine-grained control over logical complexity and struggle to balance semantic diversity with logical consistency. To address these issues, we propose QMFOL, an automated framework for generating monadic first-order logic reasoning tasks with quantifiable and controllable complexity. It constructs formal logical structures using conjunction and disjunction patterns, enabling precise control over reasoning depth, width, label types, and distractors. These structures are then translated into natural language via LLMs, with logical consistency ensured through round-trip verification using an external prover. Based on our framework, we build QMFOLBench, a benchmark comprising 2880 instances with 960 configurations across diverse logical and semantic dimensions. Evaluations on six large reasoning models (LRMs) and two LLMs show that performance degrades and computational overhead increases with rising logical complexity. Models perform better on True-labeled tasks than on False or Unknown ones, and exhibit sensitivity to semantic variation. Overall, QMFOL offers a scalable and reliable approach for constructing deductive reasoning benchmarks with controllable complexity, enabling more precise evaluation of reasoning capabilities in modern language models.

13:00 JST研究/論文

知性の熱力学的尺度

知能は測定できるのでしょうか?私たちは、インテリジェンスは稀ではあるが有効な未来の合法的な増幅として定義できると提案します。つまり、システムは、受動的ダイナミクスの下ではありそうにないが、領域の制約の下では許容可能な結果の確率を高めます。インテリジェント システムは世界とその中の独自の場所をモデル化する必要があるという前提から始めます。システムはそれがモデル化する世界の一部であるため、これは自然に再帰的自己シミュレーションにつながります。つまり、システムは、それ自体のアクションが軌道の一部である未来を表します。私たちの中心的な結果は、このアーキテクチャを希少有効先物の合法的増幅の正確な熱力学的測定に結び付ける必要性ステートメントと条件付きのほぼ十分性ステートメントを提供します。内部シミュレーションが高忠実度で希少有効先物を特定しない限り、高いレア有効リフトは不可能です。逆に、まれに有効な忠実度が高く、シミュレーションに有効なポリシーが含まれている場合、達成可能なリフトは作動が制限された最適値に近づきます。したがって、再帰的自己シミュレーションは、単に知能のもっともらしい特徴であるだけでなく、述べられた仮定の下では、高度な熱力学知能にとって必要かつほぼ十分である。結果として得られるフレームワークにより、受動的な物質やフィードバックのコントローラー、大規模な言語モデル、テキスト生成者としての人間からマクスウェルの悪魔のような情報エンジンに至るまで、普遍的なスケールでインテリジェンスを測定できるようになります。

原文 (English)

Thermodynamic Measure of Intelligence

Can intelligence be measured? We propose that intelligence can be defined as the lawful amplification of rare but valid futures: a system increases the probability of outcomes that would be unlikely under passive dynamics but remain admissible under the constraints of the domain. We start with the premise that an intelligent system must model the world and its own place within it. Because the system is part of the world it models, this leads naturally to recursive self-simulation: the system represents futures in which its own actions are part of the trajectory. Our central results give a necessity statement and a conditional near-sufficiency statement connecting this architecture to a precise thermodynamic measure of lawful amplification of rare-valid futures: high rare-valid lift is impossible unless the internal simulation identifies rare-valid futures with high fidelity; conversely, when rare-valid fidelity is high and the simulation contains an effective policy, the achievable lift approaches the actuation-limited optimum. Thus recursive self-simulation is not merely a plausible feature of intelligence but, under the stated assumptions, is necessary and nearly sufficient for high thermodynamic intelligence. The resulting framework makes intelligence measurable on a universal scale, from passive matter and feedback controllers, large language models, and humans as text generators to Maxwell-demon-like information engines.

13:00 JSTエージェント

複数目的の制約付き最適化のためのマルチエージェント システム

コンピューティングおよびネットワーキング システムにおける多くの意思決定の問題は、パフォーマンスの制約の下でコスト最小化の問題として自然に定式化できます。動的環境では、ラグランジュにヒントを得た定式化に従って、重み付けされたペナルティ項を通じてコストと制約違反の両方を単一のスカラー報酬に埋め込むことで、実行時にこのような問題を解決するために強化学習 (RL) がよく使用されます。ただし、このコンテキストでは、学習されたポリシーの動作は、通常は手動で選択されるこれらの重みの選択に大きく依存します。このため、特に相対的な重要性が変化する可能性がある非定常環境では、主な目的の最適化と制約違反の効果的な回避との間の適切なトレードオフを特定することが困難になります。この論文では、マルチエージェント RL を通じてこのバランス問題に取り組むアプローチである MAMO (Multi-Agent system for Multi-Objective Constrained Optimization) を紹介します。 MAMO は、学習問題として報酬の重みの選択を定式化することで、タスクの実行を目的の設計から切り離し、動的環境における制約付き最適化問題に対する、より自律的で堅牢な RL ベースのソリューションに向けた最初のステップを提供します。

原文 (English)

A Multi-Agent system for Multi-Objective constrained optimization

Many decision-making problems in computing and networking systems can be naturally formulated as cost-minimization problems under performance constraints. In dynamic environments, reinforcement learning (RL) is often used to solve such problems at runtime by embedding both costs and constraint violations into a single scalar reward through weighted penalty terms, following a Lagrangian-inspired formulation. However, in this context the behavior of the learned policy critically depends on the choice of these weights, which are typically selected manually. This makes it difficult to identify an appropriate trade-off between optimizing the primary objective and effectively avoiding constraint violations, particularly in non-stationary environments where their relative importance may change. This paper presents MAMO (Multi-Agent system for Multi-Objective constrained optimization), an approach to tackle this balancing problem through multi-agent RL. MAMO decouples task execution from objective design by formulating the selection of reward weights as a learning problem, providing a !rst step towards more autonomous and robust RL-based solutions for constrained optimization problems in dynamic environments.

13:00 JSTLLM/生成AI

信頼性の低いパラメトリック知識とコンテキスト知識のナビゲート: LLM 推論のための明示的知識の競合解決

大規模言語モデル (LLM) は、広範なパラメトリック知識とコンテキスト内学習能力の両方を活用することで、幅広い言語ベースのタスクにわたって強力なパフォーマンスを達成し、入力プロンプトで提供される外部情報を組み込むことができます。ただし、外部知識を統合すると、モデルの内部パラメトリック知識と外部情報の間だけでなく、複数の外部コンテキスト間でも競合が発生する可能性があります。既存のアプローチは通常、モデルまたは提供されたコンテキストのいずれかが信頼できると想定し、両方のソースにエラーが含まれている可能性を無視し、不整合を積極的に解決するのではなく、一方のソースを他方のソースよりも優先することで競合を回避します。これらの制限に対処するために、従来の二者選択パラダイムを超え、マルチエージェント推論アプローチに基づく明示的な競合解決メカニズムを組み込んだ、LLM 知識競合解決のための新しいフレームワーク MACR を提案します。具体的には、まず、修正された意味論的エントロピー尺度を使用して、特定のクエリに対する LLM の応答の信頼性を定量化する、適応的な知識評価および検索アプローチを提案します。この信頼度推定に基づいて、MACR はモデルの内部知識をテキスト表現として外部化するか、内部知識が不十分な場合に関連する外部知識を取得して、その後の推論のための基本的なコンテキストを生成します。次に、それぞれ明示的なルールを誘導し、潜在的な競合を分析し、利用可能なすべてのコンテキストにわたる不一致を解決する 3 つの特殊なエージェントを備えた帰納的マルチエージェント推論フレームワークを導入します。実証結果は、MACR がベンチマーク全体で最先端のベースラインを大幅に上回るパフォーマンスを示し、同時に明示的な競合の解釈可能な解決も提供することを示しています。

原文 (English)

Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference

Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive parametric knowledge and in-context learning ability, enabling them to incorporate external information provided in the input prompt. However, the integration of external knowledge can introduce conflicts, not only between the model's internal parametric knowledge and the external information, but also among multiple pieces of external contexts. Existing approaches typically assume that either the model or the provided context is reliable, overlooking the possibility that both sources may contain errors, and avoid conflicts by privileging one source over the other, rather than actively resolving inconsistencies. To address these limitations, we propose a novel framework MACR for LLM knowledge conflict resolution that moves beyond the conventional binary choice paradigm and incorporates an explicit conflict-resolution mechanism based on a multi-agent reasoning approach. Specifically, we first propose an adaptive knowledge assessment and retrieval approach that employs a modified semantic entropy measure to quantify an LLM's confidence in its answer to a given query. Based on this confidence estimation, MACR either externalizes the model's internal knowledge as textual representations or retrieves relevant external knowledge when internal knowledge is insufficient, generating basic contexts for subsequent reasoning. Then we introduce an inductive multi-agent reasoning framework with three specialized agents that, respectively, induce explicit rules, analyze potential conflicts, and resolve inconsistencies across all available contexts. Empirical results demonstrate that MACR significantly outperforms state-of-the-art baselines across benchmarks, while also providing interpretable resolutions of explicit conflicts.

13:00 JST研究/論文

学生が描いた科学モデルの信頼性を意識した自動評価

生徒が作成した図面は、次世代科学標準 (NGSS) に準拠したモデリング ベースのタスクにおける学習者の概念理解を評価するために、科学教育で広く使用されています。ただし、このような図面の採点には、複雑な視覚的表現を解釈するための専門的な人間の判断が必要であり、教室環境で大規模な評価を実施および維持するにはコストがかかります。この研究では、ビジョンベースのモデルを使用して、学生が作成した科学図面の自動採点を研究します。パラメータ効率の高い適応を使用してビジョン トランスフォーマー (ViT) を評価し、テスト時間の予測分布から応答レベルの信頼性を導き出す信頼性を意識したスコアリング フレームワークを提案します。この信頼度シグナルにより、不確実なケースを人間によるレビューに延期しながら、信頼度の高い回答を自動的にスコアリングすることで、選択的な自動化が可能になります。 NGSS に準拠した 6 つの中学校評価項目に関する実験では、提案されたアプローチが自動化された範囲と採点リスクの間の実際的なトレードオフをサポートしながら、採点の信頼性を向上させることが示され、信頼できる教育評価のための信頼を意識した方法の価値が強調されています。

原文 (English)

Confidence-Aware Automated Assessment of Student-Drawn Scientific Models

Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standards (NGSS). However, scoring such drawings requires expert human judgment to interpret complex visual representations, making large-scale assessment costly to implement and sustain in classroom settings. In this work, we study automated scoring of student-generated scientific drawings using a vision-based model. We evaluate a Vision Transformer (ViT) with parameter-efficient adaptation and propose a confidence-aware scoring framework that derives response-level confidence from test-time predictive distributions. This confidence signal enables selective automation by scoring high-confidence responses automatically while deferring uncertain cases for human review. Experiments on six NGSS-aligned middle school assessment items show that the proposed approach improves scoring reliability while supporting a practical trade-off between automated coverage and scoring risk, highlighting the value of confidence-aware methods for trustworthy educational assessment.

13:00 JSTエージェント

ラグランジュ: 一般化されたエンドツーエンド運転のための、オープンボキャブラリー、エネルギーベースのスパースフレームワーク

エンドツーエンドの自動運転を複雑なオープンワールド環境に拡張するには、異常なシナリオに一般化する知覚モデルと、運動学的に有効な軌道を生成するプランナーが必要です。既存のパラダイムは、表現効率と一般化能力の間の明確な二分法に直面しています。高密度モデル (占有ネットワークなど) は、幾何学的に堅牢ではありますが、重大な計算ボトルネックを引き起こし、高レベルの意味論的推論に苦労します。逆に、スパースなクエリベースのプランナーは効率的ですが、クローズドセット定義に依存しているため、配布外 (OOD) イベントに対して脆弱になります。最近の Vision-Language-Action (VLA) モデルはオープンな語彙推論を提供しますが、その自己回帰的で離散的なトークン生成は、車両ダイナミクスの連続的で高周波の制御要件と根本的に矛盾します。これに対処するために、マスクされた潜在場 (MLF) に基づいたオープン語彙で計算量が少ない駆動フレームワークである Lagrange を提案します。ラグランジュは、高密度ボリューム再構成や閉集合クエリ メカニズムに依存するのではなく、視覚言語モデル (VLM) を利用して、クラスに依存しないオブジェクトの提案を連続的なセマンティックなビジュアル トークンにエンコードします。無関係なエンティティを時間的にフィルタリングし、空間座標上で定義された暗黙的な連続エネルギー フィールドにアテンション トークンをデコードする、インテント駆動型マスク クロス アテンション モジュールを導入します。このエネルギー場にわたるラグランジュ作用最小化問題として意思決定を組み立てることにより、衝突回避を実行しながら車両運動学への厳密な準拠を強制します。標準 (nuScenes) ベンチマークとロングテール (CODA) ベンチマークの両方での広範なオフライン評価により、ラグランジュが堅牢で解釈可能、運動学的に実現可能なオープンワールド自律性のための有望なフレームワークを確立していることが実証されました。

原文 (English)

Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving

Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories. Existing paradigms face a distinct dichotomy between representational efficiency and generalization capacity. Dense models (e.g., occupancy networks), while geometrically robust, incur critical computational bottlenecks and struggle with high-level semantic reasoning. Conversely, sparse, query-based planners are efficient but reliant on closed-set definitions, rendering them vulnerable to out-of-distribution (OOD) events. Although recent Vision-Language-Action (VLA) models offer open-vocabulary reasoning, their autoregressive, discrete token generation fundamentally conflicts with the continuous, high-frequency control requirements of vehicle dynamics. To address this, we propose Lagrange, an open-vocabulary, computationally sparse driving framework based on Masked Latent Fields (MLF). Rather than relying on dense volumetric reconstructions or closed-set query mechanisms, Lagrange exploits Vision-Language Models (VLMs) to encode class-agnostic object proposals into continuous semantic visual tokens. We introduce an intent-driven masked cross-attention module that temporally filters irrelevant entities, decoding the attended tokens into an implicit continuous energy field defined over spatial coordinates. By framing decision-making as a Lagrangian action minimization problem spanning this energy field, we enforce strict compliance with vehicle kinematics while executing collision avoidance. Extensive offline evaluations on both standard (nuScenes) and long-tail (CODA) benchmarks demonstrate that Lagrange establishes a promising framework for robust, interpretable, and kinematically feasible open-world autonomy.

13:00 JST研究/論文

システムの非線形性を利用して、インテリジェント故障診断システムの設計におけるデータ不足に取り組む

深層転移学習 (DTL) により、インテリジェント障害診断システム (IFDS) を効率的に構築できます。一方で、DTL 手法は依然として大量のラベル付きデータに大きく依存しています。機械や構造物の障害に対処する場合、このような量のデータを取得するのは困難な場合があります。この文書では、データが非常に不足している状況で DTL を使用して振動ベースの IFDS を設計する新しいアプローチを提案します。実世界のシステムの固有の非線形性を活用した周期的な複数励起レベルの手順を使用して、事前にトレーニングされた畳み込みニューラル ネットワーク (CNN) で簡単に分析して故障を診断できる画像を生成します。この論文では、IFDS の設計中に遭遇する典型的なデータ不足に対処するために、新しいデータ視覚化手法とその拡張手法を提案します。鉄道パンタグラフ構造の実験的検証は、提案された方法を効果的にサポートします。

原文 (English)

Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelligent Fault Diagnosis Systems

Deep Transfer Learning (DTL) allows for the efficient building of Intelligent Fault Diagnosis Systems (IFDS). On the other hand, DTL methods still heavily rely on large amounts of labelled data. Obtaining such an amount of data can be challenging when dealing with machines or structures faults. This document proposes a novel approach to the design of vibration-based IFDS using DTL in condition of strong data scarcity. A periodic multi-excitation level procedure leveraging intrinsic non-linearities of real-world systems is used to produce images that can be conveniently analysed by pre-trained Convolutional Neural Networks (CNNs) to diagnose faults. A new data visualization method and its augmentation technique are proposed in this paper to tackle the typical lack of data encountered during the design of IFDS. Experimental validation on a railway pantograph structure provides effective support for the proposed method.

13:00 JSTエージェント

SoftSkill: 状況に応じた適応のための行動圧縮

エージェントのスキルは通常、回答ポリシー、証拠の使用習慣、タスク手順をエンコードした自然言語のマークダウン ファイルとして展開されます。これらのファイルは読み取り可能で移植可能ですが、間接的に使用されます。タスク インスタンスごとに、凍結された言語モデルが長いテキスト アーティファクトを生成時の動作に変換する必要があります。この論文では、自然言語スキルが代わりに、ベースモデルが凍結されたままの状態で、トレーニング可能なソフトデルタによって洗練されたコンパクトな連続コンテキストオブジェクトを初期化できるかどうかを尋ねます。私たちは、そのようなソフトスキルを次のトークン予測で調整し、推論時に潜在的な行動事前分布として展開する凍結バックボーン手法である SoftSkill を提案します。メインのシングルラウンド設定では、Qwen3.5-4B の長さ 32 の SoftSkill プレフィックスは、スキルなしのプロンプトよりも SearchQA で 8.3 ポイント、LiveMath で 42.1 ポイント、DocVQA で 1.3 ポイント向上しました。 SkillOpt と比較して、SoftSkill は、数百から数千の Markdown スキル トークンを少数の仮想トークンに置き換えながら、SearchQA で 5.2 ポイント、LiveMath で 12.5 ポイント精度が向上します。私たちはさらに、より困難な境界ケースとしてエージェント実行を研究します。このケースでは、まばらな軌道の模倣は有用なシグナルを提供しますが、長期的な手続きの動作をまだ強力に圧縮していません。より広範に、この結果は、一部のタスクスキルは、推論時に再解釈される追加のマークダウンとしてではなく、フリーズされたモデルがタスクにどのように入るかを制御するコンパクトな潜在的な制御として扱う方がよいことを示唆しています。

原文 (English)

SoftSkill: Behavioral Compression for Contextual Adaptation

Agent skills are commonly deployed as natural-language Markdown files that encode answer policies, evidence-use habits, and task procedures. These files are readable and portable, but they are consumed indirectly: for each task instance, a frozen language model must translate a long textual artifact into generation-time behavior. This paper asks whether a natural-language skill can instead initialize a compact continuous context object, refined by a trainable soft delta while the base model remains frozen. We propose SoftSkill, a frozen-backbone method that tunes such soft skills with next-token prediction and deploys them as latent behavioral priors at inference time. In our main single-round setting, a length-32 SoftSkill prefix on Qwen3.5-4B improves over no-skill prompting by 8.3 points on SearchQA, 42.1 points on LiveMath, and 1.3 points on DocVQA. Relative to SkillOpt, SoftSkill improves accuracy by 5.2 points on SearchQA and 12.5 points on LiveMath, while replacing hundreds to thousands of Markdown skill tokens with a few virtual tokens. We further study agentic execution as a harder boundary case, where sparse trajectory imitation provides useful signal but does not yet robustly compress long-horizon procedural behavior. More broadly, the results suggest that some task skills are better treated not as additional Markdown to be reinterpreted at inference time, but as compact latent controls over how a frozen model enters the task.

13:00 JSTエージェント

インタラクション軌跡マイニングによるコンピュータ使用エージェント向けの SKILL.md 生成の自動化

明示的なスキル ライブラリにより、コンピュータを使用するエージェントの検査が容易になりますが、そのようなライブラリを下流のポリシーを改善する方法でインタラクション データからマイニングできるかどうかは不明のままです。私たちは、GUI の軌跡をセグメント化し、セグメントを候補スキルにクラスタリングし、結果として得られるアノテーションからスキルを意識​​したポリシーをトレーニングする 3 段階のパイプラインを通じてこの質問を研究します。マイニングされたクラスターはソース ベンチマークで読み取ることができます。8 つのクラスターのうち 5 つは、InteraSkill Workflows ラベルに対して少なくとも 0.95 の純度を持っています。ただし、可読性は転送を意味するものではありません。 GRPO は、IW スキル ステップの精度を 18.5\% から 20.5\% に向上させるだけで、BrowseComp+ は基本的に変更されず、主要なソース ドメイン メトリクスに関する自明な頻度事前分布よりもパフォーマンスが劣ります。したがって、この方法を診断研究として紹介します。軌跡マイニングは検査可能なスキル構造を明らかにできますが、現在の境界検出器、順序のないセグメント表現、およびオフライン報酬モデルは、信頼性の高いクロスドメインポリシーの改善には不十分です。

原文 (English)

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining

Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from interaction data in a way that improves downstream policies. We study this question through a three-stage pipeline that segments GUI trajectories, clusters segments into candidate skills, and trains a skill-aware policy from the resulting annotations. The mined clusters are readable on the source benchmark: five of eight clusters have at least 0.95 purity against InteraSkill Workflows labels. However, readability does not imply transfer. GRPO improves IW skill-step accuracy only from 18.5\% to 20.5\%, leaves BrowseComp+ essentially unchanged, and underperforms trivial frequency priors on key source-domain metrics. We therefore present the method as a diagnostic study: trajectory mining can expose inspectable skill structure, but the current boundary detector, orderless segment representation, and offline reward model are insufficient for reliable cross-domain policy improvement.

13:00 JSTLLM/生成AINVIDIA

LLM FP4 事前トレーニングにおける収縮バイアスの再考: 幾何学的起源、システムへの影響、および UFP4 レシピ

FP4 トレーニングでは、LLM 事前トレーニングのメモリと計算コストの大幅な削減が約束されていますが、NVIDIA Blackwell/Rubin クラス システムや AMD MI350 シリーズ GPU を含む現在の FP4 ハードウェア パスとレシピは、引き続き E2M1 データ要素を中心としています。この研究では、その選択の基本的な制限を特定します。E2M1 などの不均一フォーマットは本質的に、表現可能なビンの幾何学的非対称性によって引き起こされる系統的な負の丸め誤差である縮小バイアスの影響を受けます。このバイアスは層全体で乗算的に蓄積し、ランダム アダマール変換 (RHT) によって増幅されることを示し、既存の E2M1 ベースの FP4 レシピで観察されるトレーニングの不安定性についての統一的な説明を提供します。対照的に、均一グリッド (E1M2/INT4) は、このグリッド ジオメトリ エラーを回避し、RHT によるバケット使用率の向上をより高い量子化品質に変換します。この発見に基づいて、確率的丸めを dY のみに制限しながら、RHT を 3 つのトレーニング GEMM すべてに適用する均一な 4 ビット トレーニング レシピである UFP4 を提案します。 Dense 1.5B、MoE 7.9B、および MoE 124B の長期事前トレーニングでは、スケーリング則解析とアブレーション研究によって裏付けられたように、UFP4 は強力な E2M1 ベースのベースラインよりも低い BF16 相対損失劣化を一貫して達成しています。私たちの結果は、将来のアクセラレータが E2M1 と並んでファーストクラスのトレーニング プリミティブとして E1M2/INT4 スタイルの均一 4 ビット グリッドをサポートする必要があることを示唆しています。

原文 (English)

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements. In this study, we identify a fundamental limitation of that choice: non-uniform formats such as E2M1 inherently suffer from Shrinkage Bias, a systematic negative rounding error caused by the geometric asymmetry of their representable bins. We show that this bias accumulates multiplicatively across layers and is amplified by the Random Hadamard Transform (RHT), providing a unified explanation for the training instability observed in existing E2M1-based FP4 recipes. In contrast, uniform grids (E1M2/INT4) bypass this grid-geometry error and better convert the improved bucket utilization from RHT into higher quantization quality. Based on this finding, we propose UFP4, a uniform 4-bit training recipe that applies RHT to all three training GEMMs while restricting stochastic rounding to dY alone. On Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining, UFP4 consistently achieves lower BF16-relative loss degradation than strong E2M1-based baselines, supported by scaling-law analysis and ablation studies. Our results suggest that future accelerators should support E1M2/INT4-style uniform 4-bit grids as first-class training primitives alongside E2M1.

13:00 JST研究/論文

注意誘導型ディープラーニングによる解釈可能な精子形態分類

男性不妊はカップル不妊の主な原因であり、多くの場合、精子の形態異常に関連しています。深層学習モデルは自動分析を提供しますが、そのほとんどは解釈可能性に欠けており、臨床での採用は制限されています。この研究は、精子形態分類のための注意誘導型深層学習フレームワークを提案します。当社では、事前トレーニング済みの EfficientNet-B0 と畳み込みブロック アテンション モジュール (CBAM) を組み合わせて、精子頭部の重要な領域に焦点を当て、精度と解釈可能性の両方を向上させています。 SMIDS および HuSHem 公開データセットで評価した場合、私たちのモデルは 90.2% および 93.9% (マクロ F1 スコア 0.913 および 0.948) の精度を達成し、SimpleCNN および標準 EfficientNet-B0 を上回りました。さらに、Grad-CAM++ 視覚化を使用して、モデルの決定に影響を与える機能を強調表示します。この結果は、この正確で透明なフレームワークが、不妊治療クリニックにおける自動精子分析のための実用的なツールであることを示しています。

原文 (English)

Interpretable Sperm Morphology Classification via Attention-Guided Deep Learning

Male infertility is a major cause of couple infertility, often linked to abnormal sperm morphology. While deep learning models offer automated analysis, most lack interpretability, limiting their clinical adoption. This study proposes an attention-guided deep learning framework for sperm morphology classification. We combine a pretrained EfficientNet-B0 with a Convolutional Block Attention Module (CBAM) to focus on key areas of the sperm head, improving both accuracy and interpretability. Evaluated on the SMIDS and HuSHem public datasets, our model achieves accuracies of 90.2% and 93.9% (macro F1 scores of 0.913 and 0.948), outperforming SimpleCNN and standard EfficientNet-B0. Furthermore, we use Grad-CAM++ visualizations to highlight features influencing the model's decisions. The results demonstrate that this accurate and transparent framework is a practical tool for automated sperm analysis in fertility clinics.

13:00 JST研究/論文

体外受精検査室の環境条件のコンテキスト認識型階層ベイジアン モデリング

体外受精の妊娠率は患者レベルの変数を使用して日常的にモデル化されていますが、高解像度の実験室環境データは依然として十分に活用されていません。私たちは、これが機会損失であることを示しています。私たちは、生のセンサー平均に依存するのではなく、インキュベーターの微小環境のダイナミクスを捉える、ローリング熱安定性、温度と湿度の同時付着、ピークストレス持続時間、ストレス後の回復速度など、55 のコンテキストを認識した時間的特徴を設計します。アジアの体外受精クリニックからの 61 週間のデータでは、これらの機能により、相互検証された予測誤差が 1.27% に減少しました (生の平均では 3 ~ 5% でした)。次に、部位固有のベースラインを維持しながら、部分プールを介してアジアと北欧の診療所全体で環境への影響を共有する階層型ベイジアン ベータ回帰モデルをトレーニングします。北欧の診療所から得られたデータに基づいて、このモデルは 35 ~ 39 歳の年齢層でナイーブなベースラインと比較して R2 = 0.86 と 64% の誤差低減を達成し、構造化された環境モニタリングには臨床的に意味のある転送可能なシグナルが含まれていることを示しています。

原文 (English)

Context-Aware Hierarchical Bayesian Modeling of IVF Laboratory Environmental Conditions

IVF pregnancy rates are routinely modeled using patient-level variables, while high-resolution laboratory environmental data remain underutilized. We show that this is a missed opportunity. Rather than relying on raw sensor averages, we engineer 55 context-aware temporal features, including rolling thermal stability, simultaneous temperature-humidity adherence, peak stress duration, and post-stress recovery speed, that capture the dynamics of incubator microenvironments. On 61 weeks of data from an Asian IVF clinic, these features reduce cross-validated prediction error to 1.27%, compared to 3-5% for raw averages. We then train a hierarchical Bayesian Beta regression model that shares environmental effects across an Asian and a Northern European clinic via partial pooling, while preserving site-specific baselines. On held-out data from the Northern European clinic, the model achieves R2 = 0.86 and a 64% error reduction for the 35-39 age group over a naive baseline, demonstrating that structured environmental monitoring contains clinically meaningful, transferable signal.

13:00 JSTLLM/生成AI

安全性を重視した LLM は、混合コンプライアンスのデモンストレーションから何を学びますか?

これまでの研究では、コンテキスト内のデモンストレーションが言語モデルを脱獄できることが示されていますが、モデルがさまざまな種類のコンプライアンス デモンストレーションをどのように解釈するかは依然として不明です。私たちは、無害なコンプライアンスのデモンストレーション (有害ではない要求、有益な応答) と有害なコンプライアンスのデモンストレーション (有害な要求、有益な応答) を混合し、デモンストレーションの構成が有害なコンプライアンスをどのように推進するかについて 3 つの仮説を検証することで、これを研究します。 4 つのモデルにわたって、無害なデモンストレーションと有害なデモンストレーションには互換性がないことがわかりました。無害なデモンストレーションは、モデルに応じて有害なコンプライアンスを削減または増加させることができます。さらに、選好の最適化は、無害なデモンストレーションが有害なコンプライアンスの増加を防ぐ重要なトレーニング段階であること、デモンストレーションの順序付けが強い最新性バイアスを示していること、モデルは拒否がコンテキスト内学習とどのように相互作用するかが異なることを示します。一部のモデルは拒否時にもデモンストレーションされたフォーマットを採用しますが、他のモデルは拒否時にすべてのコンテキスト内シグナルをオーバーライドします。まとめると、この研究は、デモンストレーションベースのジェイルブレイクが機能することを示すだけでなく、その仕組みを特徴づけることに移ります。つまり、コンプライアンスデモンストレーションからどのモデルが抽出されるかは、デモンストレーションの内容、順序、トレーニング方法によって異なります。

原文 (English)

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demonstrations (non-harmful request, helpful response) with harmful compliance demonstrations (harmful request, helpful response) and testing three hypotheses about how demonstration composition drives harmful compliance. Across four models, we find that benign and harmful demonstrations are not interchangeable: benign demonstrations can either reduce or increase harmful compliance depending on the model. We further show that preference optimization is the critical training stage that prevents benign demonstrations from increasing harmful compliance, that demonstration ordering exhibits strong recency bias, and that models differ in how refusal interacts with in-context learning: some adopt demonstrated formatting even when refusing, while others override all in-context signals upon refusal. Taken together, this work moves beyond showing that demonstration-based jailbreaking works to characterizing how it works: what models extract from compliance demonstrations depends on demonstration content, ordering, and training methodology.

13:00 JSTLLM/生成AI研究/論文

マルチ LCB: LiveCodeBench を複数のプログラミング言語に拡張

LiveCodeBench (LCB) は、コード生成タスクで大規模言語モデル (LLM) を評価するためのベンチマークとして最近広く採用されています。 LCB は、競技プログラミングの問題を厳選し、常に新しい問題をセットに追加し、リリース日でフィルタリングすることにより、汚染を認識した評価を提供し、コーディング能力の全体的なビューを提供します。ただし、LCB は依然として Python に限定されており、LLM が現実のソフトウェア エンジニアリングで必要とされる多様なプログラミング言語全体に汎用化できるかどうかという問題は未解決のままです。 Python を含む 12 のプログラミング言語にわたる LLM を評価するためのベンチマークである Multi-LCB を紹介します。マルチ LCB は、LCB の汚染制御と評価プロトコルを維持しながら、LCB データセットの Python タスクを他の言語の同等のタスクに変換します。オリジナルの LCB 形式と完全な互換性があるため、Multi-LCB は将来の LCB アップデートを自動的に追跡し、言語を超えたコード生成能力の体系的な評価を可能にし、Python をはるかに上回るパフォーマンスを維持するモデルを必要とします。私たちは、Multi-LCB に関する指示と推論について 24 の LLM を評価し、Python の過剰適合、言語固有の汚染、および多言語パフォーマンスの大幅な差異の証拠を明らかにしました。私たちの結果は、Multi-LCB がマルチプログラミング言語コード評価の厳密な新しいベンチマークとして確立され、LCB の主な制限に直接対処し、現在の LLM 機能の重大なギャップを明らかにします。

原文 (English)

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on code-generation tasks. By curating competitive programming problems, constantly adding fresh problems to the set, and filtering them by release dates, LCB provides contamination-aware evaluation and offers a holistic view of coding capability. However, LCB remains restricted to Python, leaving open the question of whether LLMs can generalize across the diverse programming languages required in real-world software engineering. We introduce Multi-LCB, a benchmark for evaluating LLMs across twelve programming languages, including Python. Multi-LCB transforms Python tasks from the LCB dataset into equivalent tasks in other languages while preserving LCB's contamination controls and evaluation protocol. Because it is fully compatible with the original LCB format, Multi-LCB will automatically track future LCB updates, enabling systematic assessment of cross-language code generation competence and requiring models to sustain performance well beyond Python. We evaluated 24 LLMs for instruction and reasoning on Multi-LCB, uncovering evidence of Python overfitting, language-specific contamination, and substantial disparities in multilingual performance. Our results establish Multi-LCB as a rigorous new benchmark for multi-programming-language code evaluation, directly addressing LCB's primary limitation and exposing critical gaps in current LLM capabilities.

13:00 JST研究/論文

FlowEdit: フローマッチング TTS における生涯にわたる発音適応のための連想記憶

フローマッチングのテキスト読み上げシステムは、驚くべきゼロショット品質を達成しますが、展開後は静的のままです。語彙外の固有名詞の発音エラーは、モデルが再トレーニングされない限り持続します。 FlowEdit を紹介します。これは、重みの更新ではなく潜在的な条件付け編集として発音の修正を学習する、フローズン フロー マッチング TTS のための生涯にわたる適応フレームワークです。修正フィードバックが提供されると、FlowEdit はテキスト埋め込み空間内のトークンレベルの摂動を最適化し、コンテンツアドレス指定可能なエピソード記憶として機能する最新ホップフィールド ネットワークに修正を保存します。推論時には、類似性ゲートを使用したソフト アテンションを介して修正が取得され、ファジー形態学的マッチングが可能になります。 18 の言語ファミリーにわたる 312 の多言語固有名詞からなる厳選されたベンチマークでは、FlowEdit は、同一の一般音声品質を維持しながら、ターゲット単語の音素エラー率をゼロショット ベースラインと比較して 92.7% 削減します。修正は 1 つの GPU で約 15 秒で完了します。

原文 (English)

FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunciation errors on out-of-vocabulary proper nouns persist unless the model is retrained. We introduce FlowEdit, a life-long adaptation framework for frozen flow-matching TTS that learns pronunciation corrections as latent conditioning edits rather than weight updates. When corrective feedback is provided, FlowEdit optimizes a token-level perturbation in the text embedding space, then stores the correction in a Modern Hopfield Network serving as content-addressable episodic memory. At inference, corrections are retrieved via soft attention with a similarity gate, enabling fuzzy morphological matching. On our curated benchmark of 312 multilingual proper nouns across 18 language families, FlowEdit reduces target-word Phoneme Error Rate by 92.7% relative to the zero-shot baseline while maintaining identical general-speech quality. Corrections complete in approximately 15 seconds on a single GPU.

13:00 JST研究/論文

DeepSWIP: ニューラル確率論理プログラムの商 WMC 反事実

DeepProbLog などの神経記号システムは、神経知覚と確率論的論理を組み合わせていますが、標準的な推論は連想的です。反事実的推論には、介入と証拠の因果意味論がさらに必要です。 DeepProbLog プログラム用の単一世界の反事実セマンティクスである DeepSWIP を紹介します。ニューラル具体化を使用して、固定コンテキストのニューラル述語を通常の ProbLog の選択肢に減らし、単一世界介入プログラム (SWIP) を適用し、単一の変換されたプログラムに対して重み付けモデル カウンティング (WMC) によって反事実を計算します。有限の根拠と独自にサポートされるモデルの仮定の下では、DeepSWIP は学習された具体化された FCM に対して正確です。 ProbLog 条件文の標準商-WMC 形式は、アクティブな神経確率を特定し、介入クリーニング、キャリブレーション感度、およびまれな証拠の不安定性を説明します。 MPI3D での実験では、予測どおり、12,000 クエリに対する DeepTwin 構造に対する変換と、Twin の内生的重複を回避することによる推論の 2.14 倍の高速化が確認されました。 SUMO HOV 実験では、ニューラル キャリブレーションの劣化によりプラグイン推定値にバイアスがかかる一方、スコープが正しく設定されたランダム化ポリシー AIPW 推定器では、母平均値と ATE 推定値の一次バイアスのほとんどが除去されることが示されています。コードは https://github.com/saibib/deep_SWIP にあります。

原文 (English)

DeepSWIP: Quotient-WMC Counterfactuals for Neural Probabilistic Logic Programs

Neurosymbolic systems such as DeepProbLog combine neural perception with probabilistic logic, but standard inference is associational. Counterfactual reasoning additionally requires a causal semantics for interventions and evidence. We introduce DeepSWIP, a single-world counterfactual semantics for DeepProbLog programs. Using neural materialization, we reduce fixed-context neural predicates to ordinary ProbLog choices, apply Single World Intervention Programs (SWIPs), and compute counterfactuals by weighted model counting (WMC) over a single transformed program. Under finite grounding and unique-supported-model assumptions, DeepSWIP is exact relative to the learned materialized FCM. The standard quotient-WMC form of ProbLog conditionals identifies active neural probabilities and explains intervention cleaning, calibration sensitivity, and rare-evidence instability. Experiments on MPI3D confirm the transformation against a DeepTwin construction against 12,000 queries, as predicted and a 2.14$\times$ inference speedup from avoiding the Twin's endogenous duplication. A SUMO HOV experiment shows that neural calibration degradation biases plug-in estimates, while a correctly scoped randomized-policy AIPW estimator removes most first-order bias for population mean and ATE estimands. Code is at https://github.com/saibib/deep_SWIP.

13:00 JSTLLM/生成AIエージェント

LedgerAgent: ポリシー準拠のツール呼び出しエージェントの構造化された状態

カスタマー サービス ドメインのポリシーに準拠したツール呼び出しエージェントは、ツールを呼び出してドメイン ポリシーに従いながら、ターン全体でタスクの状態を維持する必要があります。タスクの状態は、ユーザーの対話やツールの呼び出しを通じて観察される、関連する事実、識別子、制約、および条件で構成されます。標準エージェントでは、タスクの状態は個別に表現されません。観察、ツールの返却、およびポリシーの指示がプロンプトに配置されるため、エージェントは次に何を行うかを決定するたびに、プロンプトから関連する状態を再構築する必要があります。この設計では状態管理が暗黙的に行われるため、2 つの一般的な障害モードが作成されます。エージェントは正しい事実を取得しても、後で古い情報、欠落している情報、または不正確な情報に基づいて決定を下す可能性があります。また、構文的に有効なツール呼び出しでも、現在のタスクの状態に応じてドメイン ポリシーに違反する可能性があります。 \textsc{LedgerAgent} を導入します。これは、観察されたタスクの状態を別の台帳に保持し、その状態をプロンプトに表示する、ツール呼び出しエージェントのための推論時メソッドです。この台帳は、環境を変更するツール呼び出しが実行される前に状態依存のポリシー制約をチェックするためにも使用され、ポリシー違反をブロックします。 \textsc{LedgerAgent} は、4 つの顧客サービス ドメインとオープン加重モデルとクローズ加重モデルの混合パネルにわたって、標準的なプロンプトベースのツール呼び出しアプローチよりも平均パス\textasciicircumk を向上させ、より厳格な複数トライアルの一貫性指標の下で最大の利益をもたらします。

原文 (English)

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling tools and obeying domain policies. Task states consist of relevant facts, identifiers, constraints, and conditions observed through user interaction and tool calls. In standard agents, task states are not represented separately. Observations, tool returns, and policy instructions are placed in the prompt, leaving agents to reconstruct the relevant states from the prompt each time they decide what to do next. This design makes state management implicit, creating two common failure modes. An agent may retrieve the right facts but later ground its decision in stale, missing, or incorrect information; and a syntactically valid tool call may still violate a domain policy that depends on the current task state. We introduce \textsc{LedgerAgent}, an inference-time method for tool-calling agents that maintains observed task states in a separate ledger and renders the states into the prompt. The ledger is also used to check state-dependent policy constraints before environment-changing tool calls are executed, blocking policy violations. Across four customer-service domains and a mixed panel of open- and closed-weight models, \textsc{LedgerAgent} improves average pass\textasciicircum{}k over a standard prompt-based tool-calling approach, with the largest gains under stricter multi-trial consistency metrics.

13:00 JST研究/論文

指示はどのようにスピーチを形成するのでしょうか?スタイルキャプション付きテキスト読み上げのクロスアテンションアトリビューション

スタイルキャプション付きのテキスト読み上げシステムは、自然言語を使用して音声特性を制御しますが、個々の単語が音響出力にどのように影響するかは不明のままです。これを理解することは、故障モードを診断し、表現力豊かな TTS の制御性を向上させるために重要です。我々は、DAAM フレームワークを初めて音声ドメインに適応させた音声拡散モデルのクロスアテンション アトリビューションを提案し、それを CapSpeech-TTS に適用します。私たちの方法では、25 のレイヤーと 24 の ODE ステップにわたるトークンごとのヒートマップを抽出します。私たちは、それぞれ 30 個のテキスト トランスクリプトの生成を条件付ける 120 のスタイル キャプションで構成される 3,600 の組み合わせ (スタイル キャプション、テキスト トランスクリプト) を分析し、キャプション トークンがどのように波形を形成するかを明らかにします。結果は次のことを示しています: (1) スタイル トークンはコンテンツ/機能トークンよりも時間的分散が低く、グローバルな条件付けが確認されています。 (2) スタイルへの注意は F0 およびエネルギーと相関します。 (3) スタイルコンディショニングのピークは初期ステップと深い層にあります。 (4) アテンション エントロピーはレイヤー 17 で最小値に達し、スタイル重要度のピークと同時に発生し、スタイルが最も重要な段階でネットワーク選択性が最大であることを示しています。これは、自然言語が音声拡散モデルにおける相互注意にどのような影響を与えるかについての最初の研究です。

原文 (English)

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeech-TTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models

13:00 JST研究/論文

分布シフトの下で調整された専門家の混合に向けて

キャリブレーションは、モデルの予測の不確実性をその経験的結果の頻度と一致させるものであり、報告された確率を理解し信頼するために重要です。最近の研究では、個々の予測変数のレベルでキャリブレーションを強制すると、アンサンブルの精度とキャリブレーションが向上することが示されており、特に専門家混合 (MoE) モデルが強力な経験的改善を示しています。ただし、校正が MoE に役立つ条件はよく理解されていません。この研究では、ルーティング メカニズムが専門家レベルのキャリブレーションとどのように相互作用するかに焦点を当て、分布シフトの下で MoE モデルがどのように動作するかを研究します。専門家によるキャリブレーションは、ハード配線モデルの幅広いクラスの分布シフトの下でモデル全体のキャリブレーションを確実に行うには十分ですが、ソフト配線モデルのキャリブレーションには不十分であることを示します。これに対処するために、分布シフトの下でルーティングされた集約のキャリブレーション エラーにペナルティを与える敵対的再重み付けを提案し、モデル クラス、予測タスク、分布シフト全体で、データの平均と困難なサブセットの両方で精度とキャリブレーションのトレードオフが改善されることを実証します。

原文 (English)

Toward Calibrated Mixture-of-Experts Under Distribution Shift

Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve ensemble accuracy and calibration, with mixture-of-experts (MoE) models showing strong empirical improvements in particular; however, the conditions under which calibration helps MoE are not well understood. In this work, we study how MoE models behave under distribution shift, focusing on how routing mechanisms interact with expert-level calibration. We show that expert calibration is sufficient to ensure calibration of the overall model under a broad class of distribution shifts in hard-routed models, but is insufficient for calibrating soft-routed models. To address this, we propose an adversarial reweighting that penalizes calibration errors of the routed aggregate under distribution shift, and we demonstrate that it improves the accuracy-calibration tradeoff both on average and on difficult subsets of the data, across model classes, prediction tasks, and distribution shifts.

13:00 JST画像/動画生成ロボティクス

人間の普遍的な把握

人間は物体を難なく掴むことができますが、多指ロボットはこのレベルの汎用性からは程遠いです。私たちは、ロボットが把握するデータの最も自然な情報源は、毎日何千もの物体を拾う人間からのものであると主張します。我々は、ステレオ カメラからキャプチャされた単一の RGB-D 画像内のユーザー指定のオブジェクトに対する人間の多様な把握を生成するフロー マッチング モデルである HUG を紹介します。スマート グラスを使用して、まず 1M-HUG を収集します。これは、1M フレーム (27.8 時間) にわたる人間の把握の自己中心的なデータセットであり、41 の建物にわたる 6,707 のオブジェクト インスタンスです。次に、人間の自然な握りの分布をモデル化するために、私たちの新しいフロー マッチング モデルは RGB と深度の観察を融合して、手首の移動、手首の回転、および MANO の手のポーズによってパラメータ化された握りを出力します。予測された掴みをさまざまなロボットハンドにリターゲットできるため、日常シーンでのゼロショット掴みが可能になります。評価を標準化するために、メートルスケールの 3D メッシュを使用して、5 つの幾何学的カテゴリとさまざまなサイズの 90 個の未確認オブジェクトからなる新しいシミュレートされたベンチマーク HUG-Bench を構築します。私たちは、複数のステレオカメラ、ロボットの実施形態、家庭環境にわたる HUG-Bench の 30 オブジェクト テスト セットで現実世界の HUG を評価します。 HUG は、当社の挑戦的なオブジェクト セットにおいて、最先端の把握ベースラインを +23% および +34% 上回っています。コード、データ、ベンチマーク、チェックポイント、およびインタラクティブなデモは、当社の Web サイト (https://grasping.io/) でリリースされています。

原文 (English)

Human Universal Grasping

Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from humans, who pick up thousands of objects every day. We present HUG, a flow-matching model that generates diverse human grasps for any user-specified object in a single RGB-D image captured from a stereo camera. Using smart glasses, we first collect 1M-HUGs, an egocentric dataset of human grasps spanning 1M frames (27.8 hrs) and 6,707 object instances across 41 buildings. Next, to model the distribution of natural human grasps, our novel flow-matching model fuses RGB and depth observations to output a grasp parameterized by wrist translation, wrist rotation, and MANO hand pose. Predicted grasps can be retargeted to various robot hands, enabling zero-shot grasping in everyday scenes. To standardize evaluation, we build a new simulated benchmark, HUG-Bench, of 90 unseen objects from five geometric categories and various sizes, with metric-scale 3D meshes. We evaluate HUG in the real world on the 30-object test set of HUG-Bench across multiple stereo cameras, robot embodiments, and household environments. HUG outperforms the state-of-the-art grasping baselines by +23% and +34% on our challenging object set. Code, data, benchmark, checkpoints, and an interactive demo are released on our website: https://grasping.io/

13:00 JSTエージェント

ビジネス コンテキストにおける人間と AI エージェントの対話

AI エージェントが中核的なビジネス プロセスにますます統合されるにつれ、人間と AI エージェントの間の効果的な対話パターンを理解し、設計することが価値創造にとって重要になります。この研究では、AI エージェントによるポジティブなユーザー エクスペリエンス (UX) の原則と基準、およびその測定方法を特定し、評価します。私たちはユーザーの期待とニーズを特定して、導入を促進し、信頼を構築し、開発チームによるユーザー中心の意思決定をサポートします。定性的手法と定量的手法を組み合わせた混合手法アプローチを使用して、人間と AI エージェント間の相互作用パターンを調査します。この探索的研究の結果は、特定の設計要素の有効性を大規模に評価する調査実験を開発するための基礎として役立ちます。この基礎研究は、ビジネス現場におけるより直感的で効果的な人間と AI エージェントの対話の開発に貢献します。

原文 (English)

Human-AI Agent Interaction in a Business Context

As AI agents are increasingly integrated into core business processes, understanding and designing effective interaction patterns between humans and AI agents becomes crucial for value creation. This study identifies and evaluates principles and criteria for a positive User Experience (UX) with AI agents, along with methods for its measurement. We identify user expectations and needs to facilitate adoption, build trust, and support user-centered decision-making by development teams. Using a mixed-methods approach that combines qualitative and quantitative techniques, we explore interaction patterns between humans and AI agents. The findings from this exploratory research serve as the basis to develop a survey experiment which evaluates the effectiveness of specific design elements on a larger scale. This foundational research contributes to the development of more intuitive and effective human-AI agent interactions in business settings.

13:00 JSTLLM/生成AIGPT / ChatGPT

言われていないことを明らかにする: 確率的パス集約による隠れた LLM バイアスの視覚化

大規模言語モデル (LLM) には、テキスト生成の確率的な性質により評価が難しい表現的および構文的なバイアスが見られます。標準的な監査方法は、単一の出力検査または静的な自動化されたメトリクスに依存します。これらのアプローチでは、基礎となる確率分布が不明瞭になり、確率の低い生成分岐に隠されたバイアスを捕捉できません。このペーパーでは、集計された比較を通じて LLM バイアスを評価するように設計されたビジュアル分析ツールである TreeTracer を紹介します。このツールは、体系的な摂動分析パイプラインを使用して、各入力プロンプト内のオントロジーで定義された用語を置換し、数百の確率的世代を構文整合された階層構造に集約してから、補助言語モデルとの分類を認識したノードのマージを実行します。結果として得られる構造は、カスタム サンキー ダイアグラムを通じて視覚化されます。 2 つのオントロジー駆動のツリーを並置することにより、ワークスペースはセマンティック コンテキスト間の直接比較を可能にし、体系的なバイアス検出をサポートします。どの視覚化もモデルの学習された動作のサブセットのみを反映するため、システムはさらに対照的推論を適用して、コンテキスト全体にわたる反事実のトークン確率を計算して直接表示し、バイアスの存在を誤解するリスクを軽減します。私たちは、アライメントされていないベースライン モデル GPT-2 XL と構成的にアライメントされた Apertus モデルを比較するケース スタディを通じてワークスペースを検証します。視覚的な集合体は、反事実的な代名詞の抑制や会話による個人の疎外など、隠れた表象上の害悪を明らかにすることに成功しました。予備的なユーザー調査では、集約された比較インターフェイスが認知負荷を軽減し、アナリストによる体系的なバイアスの検出を効果的にサポートすることが確認されています。

原文 (English)

Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation

Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generation. Standard auditing methods rely on a single output inspection or static automated metrics. These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability generation branches. This paper introduces TreeTracer, a visual analytics tool designed to evaluate LLM bias through aggregated comparison. Using a systematic perturbation analysis pipeline, the tool replaces ontology-defined terms in each input prompt, aggregates hundreds of stochastic generations into a syntax-aligned hierarchical structure, and then performs classification-aware node merging with an auxiliary language model. The resulting structure is visualized through a custom Sankey diagram. By juxtaposing two ontology-driven trees, the workspace enables direct comparison between semantic contexts and supports systematic bias detection. Because any visualization reflects only a subset of the model's learned behavior, the system further applies contrastive inference to compute and directly display counterfactual token probabilities across contexts, reducing the risk of misinterpreting the presence of bias. We validate the workspace through case studies comparing an unaligned baseline model GPT-2 XL against the constitutionally aligned Apertus models. The visual aggregation successfully exposes hidden representational harms, such as counterfactual pronoun suppression and conversational marginalization of individuals. A preliminary user study confirms that the aggregated comparative interface reduces cognitive load and effectively supports analysts in detecting systemic biases.

13:00 JSTLLM/生成AIGoogleGeminiGemma

要約に基づいて PubMed の EQ-5D 研究を識別するための大規模言語モデルのアンサンブル

科学出版物の急速な増加により、体系的文献レビュー (SLR) における手作業による研究スクリーニングはますますリソースを消費し、非効率で、一貫性がなくなっているという事実が生じています。 EQ-5D データなど、健康関連の生活の質の結果を明確に報告する研究を分類するには、高度な臨床解釈が必要であり、人間の審査員にとっては課題となります。この研究では、公開された抄録のみに基づいて、PubMed 生物医学データベース内の EQ-5D 検出を自動化する際の Google の Gemini および Gemma 大言語モデル (LLM) の使用を調査します。少数ショット プロンプティング、重みアンサンブル集約、およびソフト スタッキング メタ分類子を統合するマルチフェーズ フレームワークが提案されています。 9 つの LLM は、EQ-5D レポートに関して 2 人の専門家によって手動でラベル付けされた PubMed 研究のデータセットで評価されます。 gemini-2.5-pro、gemma-3-12b、および gemma-3-27b の加重アンサンブルでは、0.74 の加重 F1 スコアと 0.74 の精度が得られ、個別に得られた結果を上回りました。最高パフォーマンスのモデルをアンサンブルすることで、個別のモデルと比較して精度と再現率のバランスが向上し、ソフト スタッキング アプローチにより信頼性と解釈可能性が向上しました。特徴分析により、モデルから得られる確率の結果が最終的な予測を導く上で重要であることがわかります。この調査結果は、アンサンブルベースの LLM セットアップが、生物医学研究におけるスクリーニングを自動化するための信頼性が高く、スケーラブルなアプローチであることを示唆しています。

原文 (English)

Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent. Classifying studies that clearly report health-related quality-of-life results, such as EQ-5D data, requires a high level of clinical interpretation and poses challenges for human reviewers. This study investigates the use of Google's Gemini and Gemma large language models (LLMs) in automating EQ-5D detection in the PubMed biomedical database based only on published abstracts. A multi-phase framework is proposed that integrates few-shot prompting, weight ensembling aggregation, and a soft stacking meta-classifier. Nine LLMs are evaluated on a dataset of PubMed studies manually labeled by two experts regarding EQ-5D reporting. The weighted ensemble of gemini-2.5-pro, gemma-3-12b, and gemma-3-27b obtained a 0.74 weighted F1-score and 0.74 accuracy, exceeding individually attained results. The ensembling of top-performing models improved the balance between precision and recall compared to individual models, while the soft stacking approach provided greater reliability and interpretability. Feature analysis shows that the probability results from the models are important in guiding the final predictions. The findings suggest that an ensemble-based LLM setup is a reliable and scalable approach for automating screening in biomedical research.

13:00 JSTLLM/生成AI

言語を越えた転移におけるタスクの調整から言語の関連性を解きほぐす

私たちは、アラビア語について 7 つの大きな言語モデル (4B ~ 671B パラメーター) を微調整し、セム語言語と非ユダヤ人の対照についてゼロショット読解を評価することによって、言語間伝達を研究します。高密度の専門家混合アーキテクチャ全体では、セム語特有の転移の証拠は見つかりません。ベースラインが弱いモデルはすべての言語にわたって劇的に向上しますが、ベースラインが強いモデルは言語族に関係なくわずかな向上しか示しません。思考連鎖のアブレーションはこの発見を補強します。微調整から最も恩恵を受ける同じモデルは、推論時推論からも同様に恩恵を受けます。これは、両方のメカニズムが、言語を越えた知識伝達ではなく、タスク形式の調整に取り組んでいることを示唆しています。

原文 (English)

Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer

We study cross-lingual transfer by fine-tuning seven large language models (4B--671B parameters) on Arabic and evaluating zero-shot reading comprehension on Semitic languages and non-Semitic controls. Across dense and Mixture-of-Experts architectures, we find no evidence of Semitic-specific transfer: models with weak baselines improve dramatically across all languages, while strong-baseline models show only marginal gains regardless of language family. A chain-of-thought ablation reinforces this finding -- the same models that benefit most from fine-tuning benefit equally from inference-time reasoning, suggesting both mechanisms address task-format alignment rather than cross-lingual knowledge transfer.

13:00 JSTLLM/生成AI

ハードウェア設計の RTL コーディングにおいて LLM はどのように失敗し、一般化されるのでしょうか?

逐次プログラミングの事前処理をハードウェア設計の並列時相論理に変換することは、依然として大規模言語モデル (LLM) にとって重大なボトルネックとなっています。これを調査するために、認知理論に触発された、問題解決可能性に基づいた新しいエラー分類法を導入します。私たちの分類法では、失敗を構文的、意味論的、解決可能な関数型、および解決不可能な関数型に分類します。評価の結果、フロンティア モデルの初期合格率は 90.8% で頭打ちとなるため、VerilogEval ベンチマークには厳格な経験的上限があることが明らかになりました。これらのプラトーは解決できない機能エラーによって定義され、テスト時間の計算スケーリングの影響を受けない永続的な知識のギャップを露呈します。さらに、表面上の顕著な収束ギャップが明らかになります。最適化により構文エラーはすぐに排除されますが、同時により深い機能上の障害が悪化します。私たちの調査結果は、位置合わせ技術が単にモデルにコンパイルを教えるだけであることを示しています。サンプリング戦略を繰り返すことで解決可能なエラーは解決できますが、レジスタ転送レベル (RTL) のコーディング能力は事前トレーニングの知識によって厳密に制限されたままです。現在の LLM ベースのハードウェア生成パイプラインの課題に対処するには、調整介入ではなくモデル推論に関するさらなる研究が必要です。

原文 (English)

How LLMs Fail and Generalize in RTL Coding for Hardware Design?

Translating sequential programming priors into the parallel temporal logic of hardware design remains a crucial bottleneck for large language models(LLM). To investigate this, we introduce a new error taxonomy grounded in problem solvability, inspired by cognitive theory. Our taxonomy categorizes failures into syntactic, semantic, solvable functional, and unsolvable functional types. Evaluations reveal a strict empirical ceiling on the VerilogEval benchmark, as frontier models plateau at a 90.8% initial pass rate. These plateaus are defined by unsolvable functional errors, exposing persistent knowledge gaps immune to test time compute scaling. Furthermore, we expose a striking surface convergence gap: optimization readily eliminates syntax errors but concurrently exacerbates deeper functional failures. Our findings demonstrate that alignment techniques merely teach models to compile. While repeated sampling strategies can patch solvable errors, register-transfer level(RTL) coding capacity remains strictly bounded by pretraining knowledge. Addressing challenges in the current LLM based hardware generation pipeline requires more studies in model reasoning rather than alignment interventions.

13:00 JSTLLM/生成AIDeepSeek

DeepSeek-V4: 非常に効率的な 100 万トークンのコンテキスト インテリジェンスを目指して

当社は、2 つの強力な専門家混合 (MoE) 言語モデル、1.6T パラメーター (49B アクティブ化) を備えた DeepSeek-V4-Pro と 284B パラメーター (13B アクティブ化) を備えた DeepSeek-V4-Flash を含む、DeepSeek-V4 シリーズのプレビュー バージョンを提供します。どちらも 100 万トークンのコンテキスト長をサポートします。 DeepSeek-V4 シリーズには、アーキテクチャと最適化においていくつかの重要なアップグレードが組み込まれています。(1) Compressed Sparse Attendant (CSA) と Heavy Compressed Attendance (HCA) を組み合わせたハイブリッド アテンション アーキテクチャにより、ロング コンテキストの効率が向上します。 (2) 従来の残留接続を強化するマニホールド制約ハイパー接続 (mHC)。 (3) および Muon オプティマイザーにより、収束が速くなり、トレーニングの安定性が向上します。私たちは両方のモデルを 32T を超える多様で高品質のトークンで事前トレーニングし、その後、その機能を解放してさらに強化する包括的なポストトレーニング パイプラインを実行します。 DeepSeek-V4-Pro の最大推論労力モードである DeepSeek-V4-Pro-Max は、オープン モデルの最先端を再定義し、コア タスクで以前のモデルを上回ります。一方、DeepSeek-V4 シリーズは、長いコンテキストのシナリオで非常に効率的です。 100 万トークンのコンテキスト設定では、DeepSeek-V4-Pro は、DeepSeek-V3.2 と比較して、単一トークン推論 FLOP の 27% と KV キャッシュの 10% のみを必要とします。これにより、100 万トークンのコンテキストを定期的にサポートできるようになり、長期的なタスクやさらなるテスト時間のスケーリングがより実現可能になります。モデルのチェックポイントは、https://huggingface.co/collections/deepseek-ai/deepseek-v4 で入手できます。

原文 (English)

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.

13:00 JSTLLM/生成AI

クエリをどこに配置するか?デコードダイナミクスによる拡散 LLM のインコンテキスト学習における位置バイアスの解明と軽減

インコンテキスト学習 (ICL) は自己回帰 (AR) LLM で広く研究されていますが、拡散大規模言語モデル (dLLM) 内のメカニズムはほとんど解明されていません。一方向の因果マスキングによって制限される AR モデルとは異なり、dLLM は本質的に双方向の注意を利用し、クエリ配置に広範な空間的柔軟性を提供します。残念なことに、現在の慣行は従来、AR スタイルの末尾クエリ テンプレートを継承しており、構造的なパラダイム シフトを見落としていることがよくあります。この論文では、クエリ位置が実際には dLLM の一次変数であることを明らかにする包括的な分析を紹介します。経験的な分離を通じて、位置の差異がサンプルの意味論的品質と同等の生成品質に影響を与えることを実証します。内部的には、この位置の敏感さは、注意の流れにおける空間的な「最新性効果」と、デコード軌道におけるタスク依存のシフトに起因します。グラウンドトゥルースラベルなしでこの不安定性を軽減するために、従来の単一ステップの信頼性 ($C_{decoded}$) が dLLM で失敗することを明らかにします。代わりに、反復復号プロセスを追跡する新しい指標である Average Confidence ($\overline{C}$) を提案します。基本的な空間 ICL ベースラインを確立することで、クエリの配置を動的に最適化し、異種の推論および認識タスク全体で Oracle のパフォーマンスに確実に近づく、トレーニング不要の適応ルーティング戦略である Auto-ICL を導入します。

原文 (English)

Where to Place the Query? Unveiling and Mitigating Positional Bias in In-Context Learning for Diffusion LLMs via Decoding Dynamics

While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (dLLMs) remains largely unexplored. Unlike AR models restricted by unidirectional causal masking, dLLMs intrinsically utilize bidirectional attention, offering extensive spatial flexibility for query placement. Unfortunately, current practices conventionally inherit AR-style trailing-query templates, often overlooking the structural paradigm shift. This paper presents a comprehensive analysis unveiling that query position is actually a first-order variable in dLLMs. Through empirical decoupling, we demonstrate that positional variance impacts generation quality on par with example semantic quality. Internally, this positional sensitivity stems from a spatial ``Recency Effect'' in attention flow and task-dependent shifts in decoding trajectories. To mitigate this instability without ground-truth labels, we reveal that traditional single-step confidence ($C_{decoded}$) fails in dLLMs. Instead, we propose Average Confidence ($\overline{C}$), a novel metric tracking the iterative decoding process. By establishing the foundational spatial ICL baselines, we introduce Auto-ICL, a training-free adaptive routing strategy that dynamically optimizes query placement, robustly approaching oracle performance across heterogeneous reasoning and perception tasks.

13:00 JSTLLM/生成AI

大規模言語モデルベースのナレッジグラフ推論のための幻覚の検出

ナレッジ グラフ (KG) 推論は、既存の事実から新しい知識を推測し、質問応答、推奨、意思決定支援に広く適用されます。大規模言語モデル (LLM) の急速な発展に伴い、取得した KG 情報を活用する LLM ベースの KG 推論フレームワークの人気が高まっています。しかし、LLM における幻覚は依然として重大な問題です。関連する KG の知識が組み込まれている場合でも、モデルは依然として誤った出力を生成し、誤った情報や信頼性の低い決定につながる可能性があります。既存の幻覚検出方法は、LLM の内部状態に焦点を当てるか、取得されたコンテキストとの一貫性を検証するかのいずれかですが、どちらも KG 内の構造情報を見落とすため、最適なパフォーマンスが得られません。このギャップに対処するために、我々は、LLM ベースの知識グラフ推論フレームワーク用の最初の幻覚検出方法である LUCID を提案します。 LUCID は、LLM アテンション スコア、KG セマンティクス、および構造情報を共同利用します。具体的には、アテンションスコアと意味的類似性からノードとエッジの特徴を抽出し、グラフニューラルネットワークを使用してKG構造と統合します。また、評価のために手動でアノテーションを付けたベンチマーク データセットも構築します。 9 つのデータセットに対する実験では、LUCID が 15 のベースラインと比較して最先端のパフォーマンスを達成していることが示されています。

原文 (English)

Detecting Hallucinations for Large Language Model-based Knowledge Graph Reasoning

Knowledge graph (KG) reasoning infers new knowledge from existing facts and is widely applied in question answering, recommendation, and decision support. With the rapid development of large language models (LLMs), LLM-based KG reasoning frameworks have become increasingly popular by leveraging retrieved KG information. However, hallucinations in LLMs remain a critical issue. Even when relevant KG knowledge is incorporated, models may still generate incorrect outputs, leading to misinformation and unreliable decisions. Existing hallucination detection methods either focus on LLM internal states or verify consistency with retrieved contexts, but both overlook the structural information in KGs, resulting in suboptimal performance. To address this gap, we propose LUCID, the first halLUcination deteCtIon method for LLM-based knowleDge graph reasoning frameworks. LUCID jointly leverages LLM attention scores, KG semantics, and structural information. Specifically, it extracts node and edge features from attention scores and semantic similarities, and integrates them with KG structure using a graph neural network. We also construct manually annotated benchmark datasets for evaluation. Experiments on nine datasets show that LUCID achieves state of the art performance compared to 15 baselines.

13:00 JSTLLM/生成AI研究/論文

大規模な手話データセット: リソース、ベンチマーク、および注釈標準に関する包括的な調査

手話は、聴覚障害者 (DHH) コミュニティによって使用される表現力豊かな視覚言語です。手話の認識、翻訳、作成は大幅に進歩しているにもかかわらず、断片化したデータセット、一貫性のない注釈、限られた言語範囲によって進歩は依然として制約されています。既存のベンチマークは現実世界の通信ニーズを反映していないことが多く、これらの制限の体系的な分析は依然として限られています。この調査では、35 の手話にわたる 120 のリソースをカバーする手話データセットの包括的なインデックスを提示します。モダリティの不均衡、アノテーションの粒度、署名者の偏りなどの主要な課題を分析し、将来のデータセット設計における考慮事項の概要を示します。また、標準化されたドキュメントと再現可能な評価をサポートするために、24 フィールドの Sign-Language データシートを導入し、パブリック GitHub リポジトリ (https://github.com/Ginqwerty/Open-Sign-Language) をリリースします。全体として、私たちの研究は、現実世界のアプリケーションで包括的で堅牢かつスケーラブルな手話技術を開発するための統一された実用的な基盤を提供します。

原文 (English)

Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards

Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities. Despite substantial progress in sign-language recognition, translation, and production, advances remain constrained by fragmented datasets, inconsistent annotations, and limited linguistic coverage. Existing benchmarks often fail to reflect real-world communication needs, and systematic analyses of these limitations remain limited. In this survey, we present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages. We analyze key challenges such as modality imbalance, annotation granularity, and signer bias, and outline considerations for future dataset design. We also introduce a 24-field Sign-Language Datasheet and release a public GitHub repository (https://github.com/Ginqwerty/Open-Sign-Language) to support standardized documentation and reproducible evaluation. Overall, our work provides a unified and practical foundation for developing inclusive, robust, and scalable sign-language technologies in real-world applications.

13:00 JSTLLM/生成AIエージェントQwen

信頼できるマルチエージェント システム: Argent シグナリング プロトコルによるセマンティック ドリフトの軽減

マルチエージェント LLM システムが悪い回答を生成する場合、すべての失敗が同じであるわけではありません。一部の回答は適切な内容に基づいているが不完全であり、他の回答は単純に根拠がなく、停止する必要があります。現在の再試行戦略では両方のケースが同じように扱われ (再試行して最善の結果を期待します)、人間のスーパーバイザーは再試行が正当であるかどうか、あるいは代わりにシステムを停止すべきかどうかを判断できません。 Argent Signaling Protocol (ASP) は、AI が生成するすべての応答に、確実性 (@C)、根拠 (@G)、確率性 (@S)、各主張の証拠根拠を分類する仮定インデックスなどの構造化された品質信号を伴うコンパクトな機械可読ヘッダーです。これらの信号により、コントローラーは修復可能な故障と格納容器の故障を区別し、それぞれのケースを異なる方法でルーティングすることができます。 ASP を 2 つのモードで評価します。スタンドアロン モードでは、Array BioPharma/Ono ライセンス契約に基づく 27 の質問の文書に基づいた QA ベンチマークが、3 つのローカル GGUF モデルにわたるベースライン プロンプトと ASP で計測されたコントローラー アクションを比較します。 Qwen~(0.8B) では、ASP により合格率が 11.1% から 33.3% に、平均期間カバレッジが 36.7% から 65.4% に向上しました。 Dobby~(8B) では、ASP は 4 つの失敗したリカバリを生成し、成功率を 33.3% から 44.4% に上昇させます。 SmolLM3~(3B) では、ASP は質問ごとに修復と封じ込めを切り替えます。総計の改善には意味があります (12/81 から 21/81 までのパス)。マルチエージェント モードでは、ASP サイドカーは検索エージェントと下流の決定エージェントの間に位置します。サイドカーは、接地されていないアップストリーム出力がダウンストリーム エージェントに到達するのを 100% ブロックします (24/27 ブロック、接地されていない伝播は 0)。

原文 (English)

Trustworthy Multi-Agent Systems: Mitigating Semantic Drift with the Argent Signaling Protocol

When multi-agent LLM systems produce bad answers, not all failures are equal: some answers are grounded in the right material but incomplete, while others are simply ungrounded and should be stopped. Current retry strategies treat both cases identically (try again and hope for the best), leaving human supervisors unable to tell whether a retry was warranted or whether the system should have halted instead. We introduce the Argent Signaling Protocol (ASP), a compact machine-readable header that accompanies every AI-generated response with structured quality signals: certainty (@C), grounding (@G), stochasticity (@S), and an assumption index that classifies the evidentiary basis of each claim. These signals enable a controller to distinguish repairable failures from containment failures and route each case differently. We evaluate ASP in two modes. In standalone mode, a 27-question document-grounded QA benchmark over the Array BioPharma/Ono license agreement compares baseline prompts against ASP-instrumented controller actions across three local GGUF models. On Qwen~(0.8B), ASP improves pass rate from 11.1% to 33.3% and mean term coverage from 36.7% to 65.4%; on Dobby~(8B), ASP produces 4 fail-to-pass recoveries, raising pass rate from 33.3% to 44.4%; on SmolLM3~(3B), ASP alternates between repair and containment per question. Aggregate improvement is meaningful (12/81 to 21/81 passes). In multi-agent mode, an ASP sidecar sits between a retrieval agent and a downstream decision agent; the sidecar blocks 100% of ungrounded upstream outputs from reaching the downstream agent (24/27 blocked, 0 ungrounded propagations).

13:00 JSTロボティクス

Physical Atari: ロボット上のリアルタイム強化学習のための堅牢でアクセスしやすいプラットフォーム

私たちは、Atari CX40+ コントローラーを作動させる Robotroller と呼ばれるロボットと、ゲーム フレームとアーケード学習環境からの報酬信号を画面上にレンダリングする Atari Devbox と呼ばれるデバイスを構築しました。 Robotroller と Atari Devbox は、既製のカメラとデスクトップ コンピューターとともに、物理世界で強化学習アルゴリズムを研究するために使用できるシステムを構成します。システム全体を物理アタリと呼びます。このペーパーでは、Physical Atari を堅牢でアクセスしやすいプラットフォームにするための重要な決定について詳しく説明します。システムを堅牢にするために、すべての動きがベアリングを介して行われるようにロボットローラーを設計し、摩耗を軽減しました。さらに、サーボの状態を高周波で監視し、応力を制限するために介入するソフトウェアも作成しました。システムを利用しやすくするために、家庭用 3D プリンタを使用して製造できる、手頃な価格の既製コンポーネントと部品を使用しました。 Physical Atari は 1,000 ドル未満で構築でき、数週間にわたるノンストップの強化学習実験で機械的な故障が発生することなく使用されています。私たちはこれを使用して、強化学習アルゴリズムがロボット上で直接学習できることを検証し、学習と展開の間の小さな分布の変化でさえポリシーのパフォーマンスを大幅に低下させる可能性があることを示しました。私たちの結果は、ロボットで優れたパフォーマンスを得るにはデバイス上の適応が重要であることを強調しています。

原文 (English)

Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots

We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari Devbox that renders the game frame and the reward signal from the Arcade Learning Environment on a screen. The Robotroller and the Atari Devbox, together with an off-the-shelf camera and a desktop computer, constitute a system that can be used to study reinforcement learning algorithms in the physical world. We call the full system Physical Atari. In this paper, we detail the key decisions that make Physical Atari a robust and accessible platform. To make the system robust, we designed the Robotroller so that all movement is done through bearings, which reduces wear. Additionally, we wrote software that monitors the state of the servos at a high frequency and intervenes to limit stress. To make the system accessible, we used affordable off-the-shelf components and parts that can be manufactured using consumer 3D printers. Physical Atari can be built for under $1,000 and has been used for weeks of non-stop reinforcement learning experiments without any mechanical failures. We used it to validate that reinforcement learning algorithms can learn directly on robots and show that even small distribution shifts between learning and deployment can significantly degrade the performance of policies. Our results underscore the importance of on-device adaptation for strong performance on robots.

13:00 JST研究/論文

計算による識別可能性

識別条件は、利用可能な情報の種類と量の関数として、ターゲット クエリまたは対象パラメータの計算可能性を記述します。因果関係の特定では、この情報は因果関係グラフの形式で表現されることが多く、グラフ内の変数の一部についてデータが観察または収集されます。ターゲット クエリは、単一の効果のみを対象とする場合もあれば、特定のモデル内の効果のクラスを対象とする場合もあります。次に、識別アルゴリズムの導出により、期待される望ましい因果効果を理論的に一意に決定できるプロセスが数学的に定義されます。期待における識別可能性、または「理論的識別可能性」は、一般に、漸近特性、無限データ、またはその他の数学的に理想化された条件を前提としています。この論文では、この理論的で理想的な識別可能性の概念と、計算に依存する提案された代替概念との間の根本的な違いを探ります。私たちが提案するフレームワークである「計算による識別可能性」は、代わ​​りに経験的推定量に対する有限の計算による探索手順を定義するものです。このプロセスが所望の誤差許容範囲内で経験的に推定量を見つけた場合、特定の検索の仮定 (つまり、パラメーターにわたる事前分布) と検索手順自体の条件を条件として、識別可能性が満たされます。いくつかの実験を通じて、このフレームワークがどのようにして、小さな有限サンプル、曖昧なグラフィック基準、混合観察データと介入データ、および反事実データと推定値による識別など、きめの細かい実用的な識別の質問に答えることができるかを実証します。コードは https://github.com/lbynum/metadentify で入手できます。

原文 (English)

Computational Identifiability

Identification conditions describe the computability of a target query or parameter of interest as a function of the type and amount of information available. In causal identification, this information is often expressed in the form of a causal graph, and data are observed or collected for some subset of variables in the graph. Target queries may be for a single effect alone or for a class of effects in a given model. The derivation of an identification algorithm then defines mathematically the process by which the desired causal effect(s) can be uniquely determined, theoretically, in expectation. Identifiability in expectation, or 'theoretical identifiability,' generally assumes asymptotic properties, infinite data, or other mathematically idealized conditions. In this paper, we explore a fundamental distinction between this theoretical, idealized notion of identifiability and a proposed alternative that is computation-bound. The framework we propose - 'computational identifiability' - is to instead define a finite computational search procedure for an empirical estimator. If this process finds an estimator empirically, within a desired error tolerance, then identifiability is satisfied, conditional on the specified assumptions of the search (i.e., a prior distribution over the parameters) and conditional on the search procedure itself. Through several experiments, we demonstrate how this framework allows us to answer fine-grained, practical identification questions, such as identification with small finite samples, with ambiguous graphical criteria, with mixed observational-interventional data, and across counterfactual data and estimands. Code is available at https://github.com/lbynum/metadentify.

13:00 JST研究/論文

確率的グラフィカルモデル構造学習としての情報格子学習

情報格子学習 (ILL) は、抽象化の階層をエンコードするパーティション ラティスに信号を交互に投影し、選択したルールを信号ドメインに持ち上げることで、信号の解釈可能なルールを学習します。信号が確率質量関数である場合、ILL によって学習された確率規則が自然な確率的グラフィカル モデル (PGM) 解釈を認めることを示し、この解釈を詳細に展開します。 ILL の分割は決定論的な商変数を誘導し、ルールはその商変数の限界法則です。したがって、ルールセットは、解釈可能な抽象概念に対する限界制約の集合です。一般リフティングは、これらの制約を満たすすべての同時分布の実行可能なグループですが、特殊リフティングは、最大エントロピーに密接に関連する L2 均一性原理によって ILL に実装される最大無視再構成を選択します。シャノンエントロピーリフティングの下で​​は、同じ制約により、学習された抽象化によって因子がインデックス付けされる対数線形因子グラフが生成されます。しかし、情報格子自体はベイジアン ネットワークではありません。そのエッジは、条件依存ではなく、抽象化の洗練と粗大化をコード化しています。したがって、ILL は、商変数に対する解釈可能な制約ベースの因子グラフの構造学習として捉えるのが最適です。このビューは、ILL がグラフィカル モデルおよび最大エントロピー モデルにどのように関連しているかを明らかにするとともに、推論、識別可能性、および記号と確率のハイブリッド学習の新しい方向性を提案します。

原文 (English)

Information Lattice Learning as Probabilistic Graphical Model Structure Learning

Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice that encodes a hierarchy of abstractions and lifting selected rules back to the signal domain. When the signal is a probability mass function, we show the probabilistic rules learned by ILL admit a natural probabilistic graphical model (PGM) interpretation and develop this interpretation in detail. A partition in ILL induces a deterministic quotient variable, and a rule is the marginal law of that quotient variable. A rule set is therefore a collection of marginal constraints over interpretable abstractions. General lifting is the feasible family of all joint distributions satisfying those constraints, while special lifting chooses a maximum-ignorance reconstruction, implemented in ILL by an L2 uniformity principle closely related to maximum entropy. Under a Shannon-entropy lifting, the same constraints yield a log-linear factor graph whose factors are indexed by learned abstractions. The information lattice itself, however, is not a Bayesian network: its edges encode refinement and coarsening of abstractions, not conditional dependence. Thus ILL is best viewed as structure learning for interpretable constraint-based factor graphs over quotient variables. This view clarifies how ILL relates to graphical models and maximum entropy models, while suggesting new directions for inference, identifiability, and hybrid symbolic-probabilistic learning.

13:00 JST研究/論文

ゼロインフレート ガウス分布により、分布推定アルゴリズムにおけるパラメータ空間のスパース性が可能になります

分布推定アルゴリズム (EDA) は、特に目標の構造がほとんどわかっていない場合に、ブラック ボックス最適化のための進化的手法の強力なクラスです。古典的な進化アルゴリズムは手作業で設計された突然変異や交叉演算子に依存しており、未知の問題構造やバイアスの原因を考慮して考案するのは困難ですが、EDA は演算子設計を完全に回避します。つまり、確率分布を最良の個体に適合させ、そこから次世代をサンプリングします。 EDA は連続パラメータ空間では十分に確立されていますが、良好な解のほとんどの係数が正確に 0 である疎パラメータ空間にはこれまで一般化されていませんでした。したがって、既存のスパース ブラック ボックス オプティマイザーは、EDA が回避するように設計されたものを正確に再導入しています。つまり、手作りのスパース演算子、サポート セットとアクティブな値の間で交互に行われる 2 レベル スキーム、しきい値のゼロ化、およびその他の組み込みの仮定です。我々は、EDA サンプリング法則として多変量ゼロ膨張ガウス (ZIG) 分布を提案することで、このギャップを埋めます。個別のインジケーター次元と値次元を持つ潜在ガウス モデルは、スパース パターン、アクティブなパラメーター間の相関、および 2 つの間の相互作用を表すため、スパース パターンとアクティブな値は階層なしで一緒に最適化されます。我々は、このモデルの潜在パラメータが、関連する構造が生じる欠損データ設定とは異なり、観察されたサンプルから識別可能であることを示し、それらに対する実用的な償却逆ベースの推定量を導入します。推定器は潜在的な相関構造を正確に回復し、Lunar Lander ベンチマークでは、結果として得られる ZIG-EDA は、高密度ガウス EDA、手作りのスパース進化アルゴリズム、およびアドホック スパース EDA よりも高速に収束し、より高い最終収益に達すると同時に、ごく一部のパラメーターのみがアクティブなコントローラーを検出します。

原文 (English)

Zero-Inflated Gaussian Distributions Enable Parameter-Space Sparsity in Estimation-of-Distribution Algorithms

Estimation-of-distribution algorithms (EDAs) are a powerful class of evolutionary methods for black-box optimization, especially when little is known about the structure of the objective. Whereas classical evolutionary algorithms rely on hand-designed mutation and crossover operators, hard to devise for unknown problem structures, and a source of bias, EDAs sidestep operator design entirely: they fit a probability distribution to the best individuals and sample the next generation from it. EDAs are well established on continuous parameter spaces, but they have not previously been generalized to sparse ones, in which most coefficients of a good solution are exactly zero. Existing sparse black-box optimizers therefore reintroduce exactly what EDAs were designed to avoid: hand-crafted sparsity operators, bi-level schemes alternating between support set and active values, zeroing thresholds, and other baked-in assumptions. We close this gap by proposing multivariate zero-inflated Gaussian (ZIG) distributions as EDA sampling laws. A latent Gaussian model with separate indicator and value dimensions represents sparsity patterns, correlations among active parameters, and the interactions between the two, so sparsity patterns and active values are optimized jointly, hierarchy-free. We show that the latent parameters of this model are identifiable from observed samples, unlike in the missing-data settings where related constructions originate, and introduce practical amortized inversion-based estimators for them. The estimators accurately recover latent correlation structures, and on the Lunar Lander benchmark the resulting ZIG-EDA converges faster and reaches higher final returns than a dense Gaussian EDA, a hand-crafted sparse evolutionary algorithm, and an ad-hoc sparse EDA, while finding controllers with only a small fraction of parameters active.

13:00 JST研究/論文

人間のような自律性は、セルフプレイとひとつまみの人間のデータから生まれます

セルフプレイ強化学習は、人間のデータを使用せずに運転ポリシーをトレーニングする方法として最近登場しました。高価で大規模な人間による運転デモの代わりに、安価で大規模なシミュレーションを使用します。このアプローチの主な制限は、純粋なセルフプレイを通じて訓練されたポリシーが、効果的ではあるが人間とは相容れない異質な運転慣習を学習できることです。これまでの研究では、大規模な報酬エンジニアリングとドメインのランダム化を通じて、このような行動の不整合を軽減しようとしましたが、これらは脆弱で労働集約的でした。私たちの方法では、人間のデモンストレーションを完全に破棄するのではなく、最小限の安全な目標達成報酬に加えて、正則化の目標として扱います。おいしいシチューのスパイスと同じように、少量の人的データが大いに役立つことがわかりました。私たちの方法では、人間によるデモンストレーションのみ 30 分しか使用せず、同等の模倣学習アプローチより 2,500 分の 1 です。結果として得られるポリシーは、保持されている人間の軌跡と調整され、単一の消費者グレードの GPU で 15 時間でトレーニングを完了します。ビデオと完全なソース コードは https://spiced-self-play.com/ で入手できます。

原文 (English)

Human-like autonomy emerges from self-play and a pinch of human data

Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data. It uses cheap, large-scale simulations to substitute expensive, large-scale human driving demonstrations. A key limitation of this approach is that policies trained through pure self-play can learn effective but alien driving conventions incompatible with people. Previous works attempt to mitigate such behavioral misalignments through extensive reward engineering and domain randomization, which are brittle and labor-intensive. Instead of completely discarding human demonstrations, our method treats them as a regularization objective on top of a minimal safe goal-reaching reward. Like the spice in a good stew, we find that a little human data goes a long way: our method uses only 30 minutes of human demonstrations, 2500x fewer than comparable imitation learning approaches. Resulting policies coordinate with held-out human trajectories and complete training in 15 hours on a single consumer-grade GPU. Videos and full source code are available at https://spiced-self-play.com/.

13:00 JST画像/動画生成

ProMUSE: 進行性マルチモーダル不確実性ガイドに基づく段階的証拠アルツハイマー病分類

アルツハイマー病 (AD) は、高齢者の記憶力と認知能力を破壊する致命的な疾患です。 AD の治療法のほとんどは初期段階で効果があるため、AD の早期診断に対する需要が高まっています。アルツハイマー病の診断は、臨床評価、構造磁気共鳴画像法 (MRI)、陽電子放出断層撮影法 (PET) イメージングなどの複合データにますます依存しています。しかし、MRI と PET の取得は依然としてコストが高く、誰もがアクセスできるわけではないため、フルモダリティ推論は現実世界の臨床ワークフローでは非現実的です。私たちは、追加のモダリティがいつ必要になるかを適応的に判断し、精度を維持しながらデータ収集の全体的なコストを削減する、プログレッシブ・マルチモーダル不確実性誘導段階的証拠ネットワークである ProMUSE を提案します。 ProMUSE は、まず低コストの臨床データを使用して証拠分類を実行し、ディリクレベースの主観的論理モデルによって不確実性を定量化します。不確実性が学習された閾値を超えると、ProMUSE は MRI または PET の機能を段階的に組み込み、デンプスター・シェーファー理論を通じてモダリティごとの信念と不確実性を融合して、校正されたマルチモーダル予測を取得します。この段階的な取得戦略により、高価な画像処理への依存を最小限に抑えながら、正確な診断が可能になります。 CN-AD、CN-MCI、および MCI-AD タスクにわたる ADNI、AIBL、および OASIS の実験では、ProMUSE がフルモダリティのベースラインと比較して競合または優れた精度を達成しながら、MRI/PET の使用量を 50 ~ 90% 削減し、大幅なコスト削減を実現できることが実証されています。これらの結果は、ProMUSE が現実世界の AD スクリーニングのための実用的で不確実性を認識し、リソース効率の高いソリューションであることを強調しています。

原文 (English)

ProMUSE: Progressive Multi-modal Uncertainty-guided Staged Evidential Alzheimer Disease Classification

Alzheimer's disease (AD) is a fatal disorder that destroys memory and cognitive skills in the elderly population. Most treatments for AD are effective in the early stage, leading to an increasing demand for early AD diagnosis. AD diagnosis increasingly relies on multimodal data such as clinical assessments, structural Magnetic Resonance Imaging (MRI), and Positron Emission Tomography (PET) imaging. However, MRI and PET acquisition remain costly and not universally accessible, making full-modality inference impractical in real-world clinical workflows. We propose ProMUSE, a Progressive Multi-modal Uncertainty Guided Staged Evidential Network that adaptively determines when additional modalities are necessary, helping reduce the overall cost of data acquisition while maintaining accuracy. ProMUSE first performs evidential classification using low-cost clinical data and quantifies uncertainty via a Dirichlet-based subjective logic model. When uncertainty exceeds a learned threshold, ProMUSE progressively incorporates MRI or PET features, fusing modality-wise belief and uncertainty through Dempster-Shafer theory to obtain a calibrated multimodal prediction. This staged acquisition strategy enables accurate diagnosis while minimizing reliance on expensive imaging. Experiments on ADNI, AIBL, and OASIS across CN-AD, CN-MCI, and MCI-AD tasks demonstrate that ProMUSE achieves competitive or superior accuracy compared to full-modality baselines while reducing MRI/PET usage by 50-90%, yielding substantial cost savings. These results highlight ProMUSE as a practical, uncertainty-aware, and resource-efficient solution for real-world AD screening.

13:00 JST研究/論文

cAPM: アクティブ ラーニングによる継続的な AI 支援ペースマッピング

心室頻拍は生命を脅かすリズム障害であり、心臓突然死の主な原因です。ペースマッピングは、VT のカテーテルアブレーション中に介入ターゲットを特定するための臨床手順です。臨床医は、心室のさまざまな部位のペーシングを行い、結果として得られる心電図を迅速に解釈して、次にどこでペーシングを行うか、または標的部位が特定されているかどうかを判断する必要があります。アクティブ ラーニング AI モデルは、臨床医を次のペーシング サイトに誘導するために提案されており、ペーシング サイトの数を減らし、ペース マッピングの効率を向上させることが期待されています。既存の方法では、同じ患者内または複数の患者間で複数の VT に知識を伝達する機能がなく、各ターゲットを再トレーニングする必要があります。継続的な AI 支援ペースマッピングに cAPM を導入し、過去のペースマッピング データから蓄積された知識を取得して転送し、将来のターゲット VT に必要なペースマッピング データの数を削減します。これは、ペーシング部位から 12 誘導 ECG 形態へのマッピングを学習するタスク非依存のサロゲート ニューラル ネットワーク、各ターゲットに対して最も有益なペーシング部位を選択することでこのサロゲート モデルを改良するアクティブ ラーニング戦略、および以前のターゲットからの知識を保持しながら順次学習を行う継続学習戦略によって可能になります。異なる生理学的状態および心室形状にわたって順次提示される位置特定タスクからなるインシリコテストベッドで評価したところ、過去のデータサンプルの再生ありまたはなしのcAPMは、4.5のペースマッピング部位を使用して臨床許容範囲内(精度5mm)内で位置特定の81%の確率を達成したのに対し、最先端のアクティブラーニング手法は13.7のペーシング部位を使用して38%の確率を達成しました。これらの結果は、ペースマッピングのガイドに使用できる in vivo の前臨床および臨床研究に向けて cAPM を準備するための強力な基礎を提供します。

原文 (English)

cAPM: Continual AI-Assisted Pace-Mapping with Active Learning

Ventricular tachycardia is a life-threatening rhythm disorder and a major cause of sudden cardiac death. Pace-mapping is a clinical procedure for identifying the intervention target during catheter ablation of VT. It requires clinicians to pace different sites in the ventricles and rapidly interpret the resulting electrocardiograms to determine where to pace next or whether a target site has been identified. Active learning AI models have been proposed to guide clinicians to the next pacing site, showing promise in reducing the number of pacing sites and improving the efficiency of pace-mapping. Existing methods require retraining each target without the ability to transfer knowledge across multiple VTs within the same patient or across patients. We introduce cAPM for continuous AI-assisted pace-mapping to capture and transfer knowledge accumulated from past pace-mapping data to reduce the number of pace-mapping data needed for future target VTs. This is made possible by a task-agnostic surrogate neural network that learns the mapping from pacing sites to 12-lead ECG morphology, an active-learning strategy that refines this surrogate model by selecting the most informative pacing site for each target, and a continual learning strategy to do so sequentially while retaining knowledge from prior targets. Evaluated on an in-silico testbed consisting of sequentially-presented localization tasks across different physiological conditions and ventricular geometries, cAPM with and without replay of past data samples achieved an 81% probability of localizing within clinical tolerance (5 mm accuracy) using 4.5 pace-mapping sites, compared to the state-of-the-art active-learning method achieving 38% probability using 13.7 pacing sites. These results provide a strong basis for preparing cAPM towards in-vivo preclinical and clinical studies where it can be used to guide pace-mapping.

13:00 JST研究/論文

二次構造およびエネルギー フィルター処理された水素結合グラフを使用したタンパク質表現の学習

グラフベースの表現はタンパク質モデリングで広く使用されていますが、既存のアプローチの多くは主に配列の隣接性または幾何学的近接性に依存しており、これらはタンパク質のフォールディングを支配する原理を部分的にしか反映していません。代わりに、タンパク質は、$\alpha$-helices や $\beta$-sheets などの二次構造要素を中心に組織化された複雑な三次元立体構造をとり、反復する局所モチーフと安定化する水素結合相互作用をコードします。この研究では、タンパク質表現学習のための二次構造を意識したグラフ ニューラル ネットワークを導入します。残基レベルのノード表現は二次構造の割り当てによって強化され、エネルギー強度によってフィルターされた水素結合相互作用からグラフのエッジが構築されます。この設計により、モデルはタンパク質の安定性と機能の中心となる局所的な構造コンテキストと長距離カップリングの両方を捉えることができます。私たちは、一般的に使用されているタンパク質ベンチマークで提案されたアプローチを評価し、既存のグラフベースの手法と比較して一貫した改善を観察しました。さらに、学習された接続性が確立された構造モチーフと一致するため、結果として得られるグラフ表現は生物学的解釈可能性を高めます。これらの発見は、二次構造とエネルギーフィルターされた水素結合トポロジーを組み込むことで、タンパク質表現の学習に効果的な誘導バイアスが提供されることを示唆しています。コードは https://github.com/mohamedmohamed2021/SSProNet でリリースされています。

原文 (English)

Protein Representation Learning with Secondary-Structure and Energy-Filtered Hydrogen-Bond Graphs

Graph-based representations are widely used in protein modeling, yet many existing approaches rely primarily on sequence adjacency or geometric proximity, which only partially reflect the principles governing protein folding. Proteins instead adopt complex three-dimensional conformations organized around secondary structure elements, such as $\alpha$-helices and $\beta$-sheets, which encode recurring local motifs and stabilizing hydrogen-bond interactions. In this work, we introduce a secondary-structure-aware graph neural network for protein representation learning. Residue-level node representations are augmented with secondary structure assignments, and graph edges are constructed from hydrogen-bond interactions filtered by their energetic strength. This design enables the model to capture both local structural context and long-range couplings that are central to protein stability and function. We evaluate the proposed approach on commonly used protein benchmarks and observe consistent improvements over existing graph-based methods. In addition, the resulting graph representations offer enhanced biological interpretability, as the learned connectivity aligns with established structural motifs. These findings suggest that incorporating secondary structure and energy-filtered hydrogen-bond topology provides an effective inductive bias for protein representation learning. The code is released at https://github.com/mohamedmohamed2021/SSProNet

13:00 JSTLLM/生成AIClaude

ユーザー満足度保証のもと、ユーザーからのフィードバックを限定したコスト最適な LLM ルーティング

大規模言語モデル (LLM) アプリケーションの推論コストは、需要の急増とインフラストラクチャ コストの上昇により急速に増加しています。ユーザーは高品質の応答を期待しており、商用環境ではこれがサービス レベル アグリーメント (SLA) で正式に成文化されており、コストと品質の間に根本的な緊張関係が生じています。コストを意識した LLM リクエスト ルーティングの最近の進歩により、この緊張を解決できる可能性が示されていますが、既存のアプローチは完全なフィードバック信号、オフライン トレーニング、ワークロードごとの広範な調整に依存しており、そのほとんどは SLA 保証や推論時間の適応性を欠いています。実稼働システムで得られるまばらで一方的なユーザー フィードバックからコスト最適化ポリシーを学習するオンライン ルーティング アルゴリズムである SLARouter を紹介します。 SLARouter は、コストの最適化と厳密な SLA 準拠の両方を理論的に保証します。幅広い LLM ベンチマークの実験では、SLARouter がベンチマークごとの調整を必要とせずに SLA 制約を満たし、既存のベースラインと比較して運用コストを最大 2.2 倍削減できることが示されています。

原文 (English)

Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees

Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. Users expect high-quality responses, and in commercial settings this is formally codified in Service Level Agreements (SLAs), creating a fundamental tension between cost and quality. Recent progress on cost-aware LLM request routing has shown potential to resolve this tension, but existing approaches rely on complete feedback signals, offline training, extensive per-workload tuning, and most lack SLA guarantees or inference-time adaptivity. We introduce SLARouter, an online routing algorithm that learns a cost-optimal policy from the sparse, one-sided user feedback available in production systems. SLARouter provides theoretical guarantees for both cost optimality and strict SLA compliance. Experiments across a wide range of LLM benchmarks show that SLARouter satisfies SLA constraints without the need for per-benchmark tuning, reducing operating cost by up to 2.2x over existing baselines.

13:00 JST研究/論文

Emyx: 高速かつ効率的な全原子タンパク質生成

コンピューターによる酵素の設計では、触媒残基とリガンドの足場となるタンパク質を生成する必要があり、その作業には基礎となる生成モデルの幾何学的精度と構造的多様性の両方が要求されます。現在の全原子ジェネレーターは構造予測から高価なアーキテクチャを継承しているため、トレーニング コストが高くなり、サンプルの多様性が制限されます。私たちは、この複雑さの多くは、豊富な共進化信号ではなくまばらな幾何学的制約を条件とするジェネレーターには不要であると主張します。 Emyx は、標準の変圧器ブロック内に容量を集中させ、重い埋め込みスタックを軽量の条件付き表現と疎な接続に置き換える 140M パラメータの条件付きフロー マッチング モデルです。さらに、フロー マッチング補間の正確な再パラメータ化を EDM ノイズ レベル フレームワークに導出し、再トレーニングを行わずに拡散モデル用に設計された最先端のサンプリング手法を使用してフロー マッチング トレーニング効率を橋渡しします。 Emyx は最小のモデルであるにもかかわらず、全体的な倍率回復と触媒の幾何学的精度、構造の新規性、足場の多様性、および幾何学的妥当性の両方を必要とする厳格な評価の下で、成功率全体で AME 酵素設計ベンチマークに対して Prote\'ina-Complexa と RFdiffusion3 の両方を上回っています。その一方で、トレーニング時間はわずか 682 ドルの GPU 時間で、RFdiffusion3 よりも約 4 倍少ないです。

原文 (English)

Emyx: Fast and efficient all-atom protein generation

Computational enzyme design requires generating proteins that scaffold catalytic residues and ligands, a task that demands both geometric accuracy and structural diversity from the underlying generative model. Current all-atom generators inherit expensive architectures from structure prediction, leading to high training costs and limited sample diversity. We argue that much of this complexity is unnecessary for generators, which condition on sparse geometric constraints rather than rich co-evolutionary signals. Emyx is a 140M-parameter conditional flow matching model that concentrates capacity within standard transformer blocks, replacing heavy embedding stacks with lightweight conditional representations and sparse connectivity. We additionally derive an exact reparametrisation of the flow matching interpolant into the EDM noise-level framework, bridging flow matching training efficiency with state-of-the-art sampling methods designed for diffusion models without retraining. Despite being the smallest model, Emyx outperforms both Prote\'ina-Complexa and RFdiffusion3 against the AME enzyme design benchmark across success rate under strict evaluation requiring both global fold recovery and catalytic geometry accuracy, structural novelty, scaffold diversity, and geometric validity, while training in just $682$ GPU-hours, roughly $4\times$ less than RFdiffusion3.

13:00 JSTLLM/生成AIGPT / ChatGPTLlama

変圧器フィードフォワード ブロックはどの程度線形ですか?ブロックごとの線形回復性はアーキテクチャではなく学習される

トランスフォーマー フィードフォワード ネットワーク (FFN) は、計算の非線形ストアとして扱われることがよくありますが、トレーニングされた FFN ブロックが実際にどの程度非線形であるかはほとんど測定されていません。各 FFN を位置に関する入力から出力へのマップとして扱い、それを正確な最小二乗線形近似と残差に分割します。閉形式線形マップが説明するホールドアウト分散は、ブロックの線形回復可能性 (R^2_lin)、つまりオプティマイザーを使用しないブロックの線形性の尺度を定義します。 GPT-2、Pythia-160m、および llama-160m の 12 ブロックすべてにわたって、R^2_lin は非常に不均一で深さが非単調であり、隣接するブロック間でほぼ線形 (>0.99) から強い非線形 (<0.3) までの範囲にあり、活性化関数によって設定されません。同じ幅の GELU モデル GPT-2 と Pythia-160m は大きく異なります。したがって、回復可能性は、アーキテクチャ上の特性ではなく、個々のトレーニングされたブロックの学習された特性です。残差の低ランク双線形プローブは、ゲインが残差非線形性と相関せず、R^2 の数点のみを回復します。回復されない計算は、単一の位置に関する積ではなく、高次構造または分散構造です。この測定は、ターゲットを絞った圧縮信号としても機能します。回復可能なブロックは大規模な単層置換を許可します (GPT-2 の初期の FFN では、+0.77 パープレキシティに対して 8 分の 1 のパラメーターが必要です)。一方、回復可能性の低いブロックは、これが安全でない場合にフラグを立てます。さらに、方法論上の落とし穴も明らかになります。トレーニングされた線形ベースラインは、条件の悪い変圧器の活性化では大幅に収束しない可能性があるため、全体を通して正確な閉形式の最小二乗上限を報告します。

原文 (English)

How Linear Is a Transformer Feed-Forward Block? Per-Block Linear Recoverability Is Learned, Not Architectural

Transformer feed-forward networks (FFNs) are often treated as nonlinear stores of computation, yet how nonlinear a trained FFN block actually is has rarely been measured. We treat each FFN as a position-wise input-to-output map and split it into the exact least-squares linear approximation plus a residual. The held-out variance the closed-form linear map explains defines a block's linear recoverability (R^2_lin), an optimiser-free measure of its linearity. Across all twelve blocks of GPT-2, Pythia-160m, and llama-160m, R^2_lin is highly heterogeneous and non-monotone with depth, ranging from near-linear (>0.99) to strongly nonlinear (<0.3) between adjacent blocks, and is not set by the activation function: same-width GELU models GPT-2 and Pythia-160m have sharply different profiles, so recoverability is a learned property of individual trained blocks, not an architectural one. A low-rank bilinear probe of the residual recovers only a few points of R^2, with gain uncorrelated with residual nonlinearity: the unrecovered computation is not a single position-wise product but higher-order or distributed structure. The measurement also serves as a targeted compression signal: recoverable blocks admit large single-layer replacements (GPT-2's early FFN at 8x fewer parameters for +0.77 perplexity), while low-recoverability blocks flag where this is unsafe. It further exposes a methodological pitfall: trained linear baselines can badly under-converge on ill-conditioned transformer activations, so we report the exact closed-form least-squares ceiling throughout.

13:00 JST研究/論文

コードミキシングによるガイド付き合成音声によるコードスイッチング ASR の改善

コードスイッチ (CS) 自動音声認識 (ASR) は、トレーニング用に利用できる高品質の CS テキストと音声のペアが限られているため、依然として課題が残っています。 Text-to-Speech (TTS) による合成データの拡張が検討されていますが、既存の CS TTS アプローチは主に再構築の忠実度を最適化し、言語境界の一貫性を明示的に強制しないため、CS ASR 拡張の有効性が制限されています。この論文では、コード ミキシング インデックス (CMI) を使用して、コード スイッチングの忠実度の向上に向けて合成音声生成を誘導する、コード ミキシングに基づく優先学習フレームワークを提案します。 SEAME 北京語 - 英語会話コーパスの実験により、提案された方法が ASR 微調整のための合成データの有用性を高めることが実証されました。具体的には、Whisper Large を微調整する場合、提案されたアプローチにより、DevMAN セットと DevSGE セットで混合エラー率 (MER) がそれぞれ 12.1%/17.8% から 8.9%/14.2% に減少します。

原文 (English)

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored, existing CS TTS approaches primarily optimise reconstruction fidelity and do not explicitly enforce language-boundary consistency, thereby limiting their effectiveness for CS ASR augmentation. This paper proposes a code-mixing guided preference-learning framework that steers synthetic speech generation toward improved code-switching fidelity using the Code Mixing Index (CMI). Experiments on the SEAME Mandarin-English conversational corpus demonstrate that the proposed method enhances the utility of synthetic data for ASR fine-tuning. Specifically, when fine-tuning Whisper Large, the proposed approach reduces Mixed Error Rate (MER) from 12.1%/17.8% to 8.9%/14.2% on the DevMAN and DevSGE sets, respectively.

13:00 JSTLLM/生成AIエージェント

DynAMO:トポロジカル マルチエージェント スケジューリングによる動的な資産管理オーケストレーション

LLM を利用したエージェントは産業用資産のライフサイクルにエンドツーエンドの自動化を提供しますが、実際のインダストリー 4.0 の展開は、遅延、同時実行の不安定性、安全性リスクによって妨げられています。ここでは、Plan-then-Execute アーキテクチャを使用して検証可能なワークフロー グラフを生成する、すぐに導入できるエンジンである DynAMO (Dynamic Asset Management Orchestration) を紹介します。 DynAMO は、SequentialWorkflow (トポロジー実行) と ParallelWorkflow (依存関係を意識した同時実行) の両方をサポートします。 DynAMO は、独立したタスクを動的に識別することで、構造の正確性と安全性を維持しながら、制御された推論の重複により効率を大幅に向上させます。 AssetOpsBench 産業ベンチマークに関する 6 つの制御された実験を通じて、DynAMO は大幅なパフォーマンスと堅牢性の向上を実証しました。並列実行により、エンドツーエンドのレイテンシがシーケンシャル オーケストレーションと比較して中央値 1.6 倍削減され、高度に並列化可能なワークフローでは 1.8 倍まで削減されます。外部ツール呼び出しを現実的なレイテンシで計測した後、レイテンシを分解すると、LLM 推論とオーケストレーションが依然として実行時間の 90% 以上を占め、モデル推論が主要なシステム ボトルネックであることがわかります。構造化コンテキスト プルーニングにより、推論レイテンシが約 30% 削減され、DynAMO は、制御されたフォールト インジェクションの下で正常な低下を示しながら、正しい機能動作 (タスクの完了、エージェントのシーケンス、出力品質) を維持します。再現性分析により、並列スケジューリングによりレイテンシーの変動が低減され、繰り返し実行しても安定した実行がさらに確認されます。これらの調査結果により、DynAMO は、インダストリー 4.0 自動化パイプラインにおけるスケーラブルで安全な、遅延を考慮したエージェント展開のための実用的な青写真として確立されます。コードはhttps://github.com/kushwaha001/DynAMOで入手できます。

原文 (English)

DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling

While LLM-powered agents offer end-to-end automation for industrial asset lifecycles, real-world Industry 4.0 deployment is hindered by latency, concurrency instability, and safety risks. We present DynAMO (Dynamic Asset Management Orchestration), a deployment-ready engine using a Plan-then-Execute architecture to generate verifiable workflow graphs. DynAMO supports both SequentialWorkflow (topological execution) and ParallelWorkflow (dependency-aware concurrency). By dynamically identifying independent tasks, DynAMO preserves structural correctness and safety while significantly improving efficiency through controlled reasoning overlap. Across six controlled experiments on the AssetOpsBench industrial benchmark, DynAMO demonstrates substantial performance and robustness gains. Parallel execution reduces end-to-end latency by a median of 1.6x over sequential orchestration, rising to 1.8x on highly parallelizable workflows. After instrumenting external tool calls with realistic latencies, a latency decomposition shows that LLM reasoning and orchestration still account for more than 90% of execution time, identifying model inference as the primary system bottleneck. Structured context pruning reduces inference latency by approximately 30%, and DynAMO maintains correct functional behaviour (task completion, agent sequencing, and output quality) while exhibiting graceful degradation under controlled fault injection. Reproducibility analysis further confirms stable execution under repeated runs, with parallel scheduling reducing latency variance. These findings establish DynAMO as a practical blueprint for scalable, safe, and latency-aware agent deployment in Industry 4.0 automation pipelines. Code is available at: https://github.com/kushwaha001/DynAMO

13:00 JSTエージェント

構造による双安定性: 壁時計で調整された状態モニターには、エージェントのリズムでの瞬間検出機構がありません

自律エージェントのランタイム モニターは通常、蓄積された内部状態 (行動ベースライン、ドリフト統計、または以前の研究ではモデル化された感情状態) を閾値に設定します。私たちは以前、状態飽和トラップを報告しました。継続的影響エンジン上のしきい値オン状態トリガーが、SWE ベンチ デバッグ エージェント上でほぼ一定のアラームになる (Modgil 2026)。リリース後の監査により、エンジンがアクション間で dt=0 を受信したため、指数関数的な減衰が動作しなかったことが判明しました。公開されたトラップは純粋なアキュムレータの結果です。私たちは記録 (正誤表、v2) を修正し、欠陥を実験として扱います。それが明らかにする重要な変数は、モニターのダイナミクスがサンプル時間 (CUSUM などの観測ごと) で校正されるか、実時間 (影響モデルや EMA ベースラインなどの秒単位の半減期) で校正されるかです。固定レートのストリームでは、これらは一致します。エージェント ストリームでは、対話時間が桁違いに変化しますが、変化しません。 20 個の軌道上の均一な間隔 (dt 単位: {0..600} 秒) にわたる事前に登録されたスイープは、壁時計レベルのトリガーに 2 つのレジームがあることを示しています: dt=60 秒のサイレント。すべてのクリティカル dt は (1,30] 秒内にあります。実際のエージェントの実行では、中央値 1.53 秒 (p90 2.33 秒) でレイテンシが測定されます。実際のコーディング リズムはトラップ領域内にあり、修正されたメカニズムの下での経験的結果が証明されています。構造はエンジンではなくキャリブレーション クラスのプロパティです。生のエラー ストリームに対する最小の壁時計アキュムレータは同じクリフを再現しますが、サンプル時間 CUSUM は同じクリフを再現します。ストリームは正確に dt 不変 (20/20) です。ヒステリシスのある立ち上がりエッジ トリガーは、すべての条件で軌道ごとに 0 ~ 3 回発生します。ウォールクロックで校正されたリーキー積分器モニターは、エージェント ストリームの瞬間検出器として機能することはなく、すべてのリズムでトラップを回避できますが、人間の介入タイミングは回復しません。

原文 (English)

Bistable by Construction: Wall-Clock-Calibrated State Monitors Have No Moment-Detection Regime at Agent Cadence

Runtime monitors for autonomous agents commonly threshold an accumulated internal state - a behavioural baseline, a drift statistic, or, in our prior work, a modelled affective state. We previously reported a State Saturation Trap: threshold-on-state triggers over a continuous affect engine become near-constant alarms on SWE-bench debugging agents (Modgil 2026). A post-release audit found the engine received dt=0 between actions, so its exponential decay never operated: the published trap is a pure-accumulator result. We correct the record (erratum, v2) and treat the flaw as an experiment. The key variable it exposes is whether a monitor's dynamics are calibrated in sample time (per observation, as in CUSUM) or wall-clock time (half-lives in seconds, as in affect models and EMA baselines). On fixed-rate streams these coincide; on agent streams, where inter-action time varies by orders of magnitude, they do not. A pre-registered sweep over uniform intervals (dt in {0..600}s) on 20 trajectories shows the wall-clock level trigger has two regimes: at dt=60s silent. Every critical dt lies in (1,30]s. Real agent runs measure latency at median 1.53s (p90 2.33s); real coding cadence sits inside the trap regime, vindicating the empirical finding under a corrected mechanism. The structure is a property of the calibration class, not the engine: a minimal wall-clock accumulator over the raw error stream reproduces the same cliff, while a sample-time CUSUM over the identical stream is exactly dt-invariant (20/20). A rising-edge trigger with hysteresis fires 0-3 times per trajectory in every condition. We conclude that wall-clock-calibrated leaky-integrator monitors admit no regime in which they act as moment detectors on agent streams; transition detection escapes the trap at every cadence, but does not recover human intervention timing.

13:00 JSTLLM/生成AI

LLM 主導の段階的改良による解釈可能かつ検証可能なハードウェア生成

大規模言語モデル (LLM) は、ソフトウェア開発において目覚ましい成功を収めています。ただし、幻覚の影響を受けやすいため、微妙な意味的および論理的エラーが発生する可能性があります。チップの設計と製造には大きなリスクが伴うため、ハードウェア エンジニアは依然としてレジスタ転送レベル (RTL) の生成に LLM に依存することに消極的です。この論文では、LLM の創造性と幅広い知識を、形式的手法の説明可能性と数学的厳密性と組み合わせたハードウェア生成フレームワークを提案します。具体的には、さまざまな設計上の決定とハードウェア機能をカバーする一連の変換ルールを考案します。これらのルールを繰り返し適用することで、LLM エージェントは、正確性が保証された設計仕様を RTL プログラムに変換できます。実験結果は、フレームワークの有効性と効率性を示しています。

原文 (English)

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement

Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meaning that they can introduce subtle semantic and logical errors. Due to the high stakes in chip design and manufacturing, hardware engineers are still reluctant to rely on LLMs for register-transfer level (RTL) generation. In this paper, we propose a hardware generation framework that combines the creativity and broad knowledge of LLMs with the explainability and mathematical rigor of formal methods. Specifically, we devise a set of transformation rules that cover various design decisions and hardware features. By iteratively applying these rules, an LLM agent can convert a design specification into an RTL program with guaranteed correctness. Experimental results demonstrate the effectiveness and efficiency of the framework.

13:00 JSTエージェントClaude

エージェント AI 向けの実行限定アドバイザリー自動化: 再現可能な AIBOM 主導の CSAF-VEX フレームワーク

SBOM および AIBOM アーティファクトを決定論的な環境キャプチャおよび構造化されたランタイム テレメトリにバインドする、プロトコル駆動のフレームワークが提供されます。悪用可能性は、宣言されたアーティファクト、観察されたアクティブ化条件、および強制された実行ポリシーから計算されます。 CSAF VEX アドバイザリは、静的証拠と実行時の証拠を組み合わせて生成され、暗号的に署名され、決定論的な再生によって検証されます。評価では、OSV、GitHub Advisory、KEV、EPSS データセットを組み込んだ、合成 Agentic AI ワークロード 50 ~ 5000 コンポーネントにわたって約 10000 コンポーネント エントリを使用します。

原文 (English)

Execution-bound advisory automation for agentic AI: a reproducible AIBOM-driven CSAF-VEX framework

A protocol driven framework is presented that binds SBOM and AIBOM artefacts to deterministic environment capture and structured runtime telemetry. Exploitability is computed from declared artefacts, observed activation conditions, and enforced execution policies. CSAF VEX advisories are generated from combined static and runtime evidence, cryptographically signed, and validated through deterministic replay. Evaluation uses approximately 10000 component entries across synthetic Agentic AI workloads 50 to 5000 components, incorporating OSV, GitHub Advisory, KEV, and EPSS datasets.

13:00 JSTLLM/生成AI

VERITAS: ゼロショット形式定理証明のための検証者ガイド付き証明検索

LLM ベースの形式的証明者は、多くの場合、豊富な検証信号 (構文エラー、型の不一致、部分的な目標の進捗状況) をバイナリの合否ビットにまとめます。 VERITAS は、すべての検証信号を 2 フェーズのプロトコルを通じて証明検索に戻すゼロショット フレームワークです。最初に Best-of-N サンプリング、次にフェーズ 1 の失敗を明示的な否定例として取り込む批評家主導の MCTS パスです。このプロトコルは、独自のフェーズ 1 スイープによって解決されたすべての定理を保存するため、フェーズ 2 の追加の解決はフィードバック駆動の探索に起因します。 VERITAS は、miniF2F で 40.6% に達し (一方、独立して運営されている Best-of-5 の 36.9%、ポートフォリオ 26.2%)、VERITAS-CombiBench では 7.3% に達しています。VERITAS-CombiBench は、当社がリリースしている 55 定理の組み合わせベンチマークで、Best-of-5 (1.8%) が Portfolio (3.6%) を下回っており、正しい場合でもガイドなしのサンプリングが有害であることが明らかになりました。補題名は検証者のフィードバックから繰り返し復元する必要があります。アーティファクトは GitHub で入手できます。

原文 (English)

VERITAS: Verifier-Guided Proof Search for Zero-Shot Formal Theorem Proving

LLM-based formal provers often collapse rich verifier signals (syntax errors, type mismatches, partial goal progress) into a binary pass/fail bit. We present VERITAS, a zero-shot framework that routes every verifier signal back into proof search through a two-phase protocol: Best-of-N sampling first, then a critic-guided MCTS pass that ingests Phase 1 failures as explicit negative examples. The protocol preserves every theorem solved by its own Phase 1 sweep, so Phase 2's additional solves are attributable to feedback-driven exploration. VERITAS reaches 40.6% on miniF2F (vs. an independently run Best-of-5 at 36.9%, Portfolio 26.2%) and 7.3% on VERITAS-CombiBench, a 55-theorem combinatorics benchmark we release on which Best-of-5 (1.8%) falls below Portfolio (3.6%), exposing that unguided sampling hurts when correct lemma names must be recovered iteratively from verifier feedback. Artifacts are available on GitHub.

13:00 JST研究/論文

JustDiag!: 責任ある根本原因分析のための診断正当化エンジン

大規模な言語モデルは流暢な根本原因分析を生成できますが、流暢な最終的な答えだけでは、一か八かの業務における説明責任の証拠としては不十分です。実際のインシデント対応では、エンジニアは、どのような証拠が診断を裏付けたか、どの代替案が検討されたか、どこに矛盾が残ったか、システムがケースを解決したか、不確実性が維持されたかを知る必要があります。私たちは、証拠、所見、競合する仮説、競合、次のチェックに対する明示的なプロセス状態を維持する RCA 用の診断正当化エンジン JustDiag を使用して、このギャップに対処します。最終回答の品質とプロセスの品質を個別にスコアリングする 2 層プロトコルを使用して、66 件の実世界のインシデントでシステムを評価しました。診断上の正当性のない一致した対照と比較して、JustDiag はより強力な結果とプロセス スコアを達成しましたが、より校正された非閉鎖によりわずかに低い最終完了を受け入れました。これらの結果は、説明責任のある RCA には、流暢な最終回答だけでなく、明示的な診断正当化アーティファクトとプロセスを意識した評価が必要であることを示唆しています。

原文 (English)

JustDiag!: A Diagnostic Justification Engine for Accountable Root Cause Analysis

Large language models can produce fluent root cause analyses, but fluent final answers alone are insufficient evidence for accountability in high-stakes operations. In real incident response, engineers need to know what evidence supported a diagnosis, which alternatives were considered, where contradictions remained, and whether the system resolved the case or preserved uncertainty. We address this gap with JustDiag, a diagnostic justification engine for RCA that maintains an explicit process state over evidence, findings, competing hypotheses, conflicts, and next checks. We evaluated the system on 66 real-world incidents using a two-layer protocol that separately scores final-answer quality and process quality. Relative to a matched control without diagnostic justification, JustDiag achieved stronger outcome and process scores, while accepting slightly lower terminal completion due to more calibrated non-closure. These results suggest that accountable RCA requires explicit diagnostic justification artifacts and process-aware evaluation, not only fluent final answers.

13:00 JSTエージェントロボティクス

遊び心のあるエージェントロボット学習

現在のエージェント ロボット システムは、実行可能な Code-as-Policy プログラムを作成し、フィードバックを観察し、複数の試行にわたって動作を修正することができますが、依然として主にタスク駆動型であり、再利用可能なスキルは明示的な指示の後にのみ獲得されます。私たちは、遊び心のあるエージェントロボット学習を研究しています。この学習では、身体化されたコーディングエージェントが、下流のタスクが到着する前に、継続的なスキル学習段階として自主的な遊びを使用します。プレイタイムのスキル習得のために設計されたロボット エージェント チームである RAT を紹介します。プレイ中に、RAT は新規でありながら学習可能な探索的タスクを提案し、ロボット コード ポリシーを計画および実行し、中間の進捗状況を確認し、失敗を診断し、高密度のステップ レベルのフィードバックで再試行し、成功した実行を永続的なコード スキル ライブラリに抽出します。テスト時に、エージェントはこの凍結されたライブラリから関連するスキルを再利用して、新しいタスクの解決に役立てます。 LIBERO-PRO と MolmoSpaces での実験では、遊びで学習したスキルが、ノープレイやランダムプレイのベースラインよりも保留された下流タスクを改善し、LIBERO-PRO と MolmoSpaces でそれぞれ CaP-Agent0 よりも 20.6 パーセント ポイントと 17.0 パーセント ポイント向上したことが示されています。さらに、学習したスキルをコンテキストに取得するだけで、他の推論時の Code-as-Policy エージェントに組み込むことができ、基礎となるモデルを微調整することなく、RoboSuite と現実世界の転送がそれぞれ 8.9 ポイントと 8.8 ポイント向上します。

原文 (English)

Playful Agentic Robot Learning

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams designed for play-time skill acquisition. During play, RATs proposes novel yet learnable exploratory tasks, plans and executes robot-code policies, verifies intermediate progress, diagnoses failures, retries with dense, step-level feedback, and distills successful executions into a persistent code skill library. At test time, the agent reuses relevant skills from this frozen library to help solve new tasks. Experiments in LIBERO-PRO and MolmoSpaces show that play-learned skills improve held-out downstream tasks over no-play and random-play baselines, with 20.6 and 17.0 percentage-point gains over CaP-Agent0 on LIBERO-PRO and MolmoSpaces, respectively. Moreover, the learned skills can be plugged into other inference-time Code-as-Policy agents by simply retrieving them into the context, improving RoboSuite and real-world transfer by 8.9 and 8.8 points, respectively, without finetuning the underlying model.

13:00 JST画像/動画生成

整流変圧器を使用した胸部 X 線撮影用の生成基礎モデルのスケーリング

10 億パラメータ規模でゼロからトレーニングされた、胸部 X 線写真合成のための最初の生成基盤モデルを紹介します。既存の X 線撮影 AI モデルは、多くの場合、患者集団、施設、撮影設定全体での一般化が不十分であり、その結果、現実世界の臨床での有用性が制限されます。胸部 X 線写真の制御された高忠実度の合成は、臨床データセットを多様化し、診断モデルの堅牢性を評価するための有望な道です。したがって、120万枚のX線写真と臨床専門家がガイドするメタデータで構成される精選された異種データセットで1.6Tトークン用にトレーニングされた、13億を超えるパラメータを備えたこれまでで最大の胸部X線写真の専門家による生成基礎モデルを提示します。私たちのモデルは、複数の人口統計サブグループ、取得ビュー、および十数の病理にわたる制御可能な X 線写真の生成と編集をサポートしています。さらに、X線写真合成の忠実度における最先端技術を大幅に進歩させ、臨床専門家にとって本物のX線写真と区別できない画像を生成します。

原文 (English)

Scaling Generative Foundation Models for Chest Radiography with Rectified Flow Transformers

We introduce the first generative foundation model for chest radiograph synthesis trained from scratch at the billion-parameter scale. Existing radiographic AI models often suffer from poor generalisation across patient subpopulations, institutions, and acquisition settings, resulting in limited real-world clinical utility. Controlled, high-fidelity synthesis of chest radiographs is a promising path toward diversifying clinical datasets and evaluating the robustness of diagnostic models. Therefore, we present the largest specialist generative foundation model for chest radiographs to date, with over 1.3B parameters, trained for 1.6T tokens on a curated, heterogeneous dataset comprising 1.2M radiographs and clinical expert-guided metadata. Our model supports controllable radiograph generation and editing across multiple demographic subgroups, acquisition views, and a dozen pathologies. Moreover, we significantly advance the state of the art in radiograph synthesis fidelity, producing images that are indistinguishable from real radiographs to clinical experts.

13:00 JSTLLM/生成AI

LLM 支援ポスト量子暗号開発におけるセキュア コーディングのドリフト: ゲーム化された修正

ポスト量子暗号 (PQC) への移行により、実装がかなり複雑になり、定数時間実行、サイド チャネル耐性、および正確なパラメータ化への厳密な準拠が必要になります。同時に、大規模言語モデル (LLM) が暗号エンジニアリングを含むソフトウェア開発ワークフローに大きく組み込まれています。 LLM は生産性を向上させますが、特にセキュリティが重要なドメインでは、安全でないコードや次善のコードが頻繁に生成されることが証拠によって示されています。このペーパーでは、LLM で生成されたコードへの継続的な依存によるセキュア コーディング実践の段階的な劣化を捉える新しい社会技術的脆弱性モデルである、PQC におけるセキュア コーディング ドリフトを紹介します。静的な脆弱性に焦点を当てたこれまでの研究とは異なり、私たちはセキュリティ リスクを人間の AI 相互作用から生じる長期的な行動現象として概念化しています。これを軽減するために、敵対的評価、行動フィードバック、セキュリティ スコアリングを開発ワークフローに組み込む、ゲーム化された LLM 強化セキュア コーディング フレームワークを提案します。私たちのアプローチは、LLM を受動的なアシスタントから能動的なセキュリティ副操縦士に再構成し、AI が介在する環境でのより安全な PQC 実装に貢献します。

原文 (English)

Secure Coding Drift in LLM-Assisted Post-Quantum Cryptography Development: A Gamified Fix

The transition to Post Quantum Cryptography (PQC) introduces considerable implementation complexity, requiring strict adherence to constant-time execution, side channel resistance, and precise parametrisation. Simultaneously, large language models (LLMs) are heavily embedded in software development workflows, including cryptographic engineering. While LLMs improve productivity, evidence shows that they frequently generate insecure or suboptimal code, particularly in security critical domains. This paper introduces Secure Coding Drift in PQC, a novel socio technical vulnerability model capturing the gradual degradation of secure coding practices due to sustained reliance on LLM-generated code. Unlike prior work that focuses on static vulnerabilities, we conceptualise security risk as a longitudinal behavioural phenomenon rising from human AI interaction. To mitigate this, we propose a gamified, LLM augmented secure coding framework that embeds adversarial evaluation, behavioural feedback, and security scoring into development workflows. Our approach reframes LLMs from passive assistants into active security co-pilots, contributing toward safer PQC implementation in AI mediated environments.

13:00 JST研究/論文

文脈に沿った学習は本質的な好奇心をサポートできるか?

効果的な機械学習は、データをモデル化する方法だけでなく、収集するデータを選択することにも依存します。大規模シーケンス モデルはデータ モデリングに革命をもたらしましたが、自動化されたデータ選択、つまり「本質的な好奇心」の問題は依然として大きな課題です。古典的なアプローチでは、新たに取得した観測値が世界モデルの予測能力をどの程度改善したかを測定する「学習の進捗状況」に基づいてエージェントに報酬を与えることで探索を奨励します。ただし、これらの報酬を評価するには、従来、各軌道内で勾配降下法を更新する高価な内部ループが必要であり、大規模な計算では非実用的でした。この研究では、シーケンス モデルの創発インコンテキスト学習 (ICL) 機能が、即時の更新不要の世界モデルとして機能することで、このボトルネックを解消できるかどうかを調査します。具体的には、コンテキスト内学習者の予測誤差と反事実的なコンテキスト操作のみを使用して、学習の進行を最大化するように探索ポリシーをトレーニングできるかどうかを評価します。まず、一般的なマルコフ決定プロセスでは、これは公平な方法で実際には不可能であることを証明します。結果として得られる固有の報酬は、真の学習の進行状況の推定に偏りをもたらす迷惑項の影響を受けるか、コンテキスト内の学習者の予測誤差を使用して実装することができません。逆に、アクティブ ラーニングとベイジアン実験計画を含む、非時間的設定の広範なサブクラスに対して肯定的な結果が得られることを証明します。ここでは、ICL から導出された報酬がうまくバインドされ、真の学習の進行状況に漸近的に収束します。私たちは、連続的かつ記号的な環境にわたる制御された実験で理論を裏付け、ICL 主導のフレームワークが、最適に探索する興味深いデータ収集ポリシーのトレーニングに成功していることを実証しました。

原文 (English)

Can In-Context Learning Support Intrinsic Curiosity?

Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge. Classic approaches incentivize exploration by rewarding an agent based on its "learning progress", which measures how much a newly acquired observation improves a world model's predictive ability. However, evaluating these rewards traditionally requires expensive inner loops of gradient descent updates within each trajectory, rendering them computationally impractical at scale. In this work, we investigate whether the emergent in-context learning (ICL) capabilities of sequence models can eliminate this bottleneck by serving as immediate, update-free world models. Specifically, we evaluate whether an exploration policy can be trained to maximize learning progress, using solely the prediction errors and counterfactual context manipulations of an in-context learner. We first prove that in general Markov decision processes, this is in fact impossible in an unbiased way: the resulting intrinsic rewards either suffer from nuisance terms that bias their estimation of true learning progress, or they cannot be implemented using an in-context learner's prediction errors. Conversely, we prove a positive result for a broad subclass of non-temporal settings, encompassing active learning and Bayesian Experimental Design: here, ICL-derived rewards successfully bound and asymptotically converge to the true learning progress. We corroborate our theory with controlled experiments across continuous and symbolic environments, demonstrating that our ICL-driven framework successfully trains curious data-collection policies that explore optimally.

13:00 JST研究/論文

コンセプト フロー モデル: 階層的なボトルネックによるコンセプト ベースの推論の固定

コンセプト ボトルネック モデル (CBM) は、学習した特徴を人間が理解できる概念空間に投影することで解釈可能性を高めます。最近のアプローチでは、ビジョン言語モデルを活用して概念の埋め込みを生成し、手動で概念に注釈を付ける必要性を減らしています。ただし、これらのモデルには重大な制限があります。概念の数が埋め込み次元に近づくにつれて情報漏洩が増加し、モデルが偽りの相関関係や意味的に無関係な相関関係を悪用して解釈可能性を損なうことが可能になります。この研究では、フラットなボトルネックを階層的な概念主導のデシジョン ツリーに置き換えるコンセプト フロー モデル (CFM) を提案します。階層内の各内部ノードは、識別概念の局所的なサブセットに焦点を当て、予測範囲を徐々に狭めます。私たちのフレームワークは、視覚的な埋め込みから意思決定階層を構築し、各階層レベルで意味概念を配布し、確率的ツリー走査を通じて微分可能な概念の重みをトレーニングします。さまざまなベンチマークに関する広範な実験により、CFM がフラット CBM の予測パフォーマンスに匹敵すると同時に、効果的な概念の使用を減らすことで情報漏洩を大幅に軽減できることが実証されました。さらに、CFM は段階的な意思決定フローを生成し、階層クラス構造を使用した透過的で監査可能なモデル推論を可能にします。

原文 (English)

Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks

Concept Bottleneck Models (CBMs) enhance interpretability by projecting learned features into a human-understandable concept space. Recent approaches leverage vision-language models to generate concept embeddings, reducing the need for manual concept annotations. However, these models suffer from a critical limitation: as the number of concepts approaches the embedding dimension, information leakage increases, enabling the model to exploit spurious or semantically irrelevant correlations and undermining interpretability. In this work, we propose Concept Flow Models (CFMs), which replace the flat bottleneck with a hierarchical, concept-driven decision tree. Each internal node in the hierarchy focuses on a localized subset of discriminative concepts, progressively narrowing the prediction scope. Our framework constructs decision hierarchies from visual embeddings, distributes semantic concepts at each hierarchy level, and trains differentiable concept weights through probabilistic tree traversal. Extensive experiments on diverse benchmarks demonstrate that CFMs match the predictive performance of flat CBMs, while substantially mitigating information leakage by reducing effective concept usage. Furthermore, CFMs yield stepwise decision flows that enable transparent and auditable model reasoning with hierarchical class structures.

13:00 JSTLLM/生成AILlamaQwen

LoRA のピーク メモリ削減手法 エッジ デバイス上の LLM の微調整

エンドユーザーのデータに対して低ランク適応 (LoRA) を使用して大規模言語モデル (LLM) を微調整すると、データのプライバシーを保ちながらパーソナライズされたエクスペリエンスが提供されますが、コンシューマ ハードウェアでは厳しいメモリ制約に直面します。特に数十億のパラメーターと長いコンテキストのトレーニング データを含むモデルの場合、微調整中のピーク メモリはデバイスの制限を超えることがよくあります。このペーパーでは、モデルの品質を犠牲にすることなくメモリ フットプリントを削減するための一連の補完的な手法を紹介します。(1) オンザフライ逆量子化による基本モデルの量子化、(2) 選択的アクティベーション キャッシュとディスク オフロードを組み合わせたメモリ効率の高いチェックポイント処理、(3) 意味的に関連するトークン サブセットを使用したソフトマックス近似、(4) ロジット マスキング。 Llama-3.2 3B および Qwen-2.5 3B での実験では、ピーク メモリが最大 $26\times$ および $28\times$ 削減され、リソースに制約のあるデバイスでの微調整が可能になることが実証されました。

原文 (English)

Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices

Fine-tuning of Large Language Models (LLMs) using Low-Rank Adaptation (LoRA) on an end-user's data offers personalized experiences while keeping data private, but faces severe memory constraints on consumer hardware. Peak memory during fine-tuning often exceeds device limits, especially for models with billions of parameters and long-context training data. This paper introduces a suite of complementary techniques to reduce memory footprint without sacrificing model quality: (1) base model quantization with on-the-fly dequantization, (2) memory-efficient checkpointing combining selective activation caching and disk offloading, (3) softmax approximation using semantically relevant token subsets, and (4) logits masking. Experiments on Llama-3.2 3B and Qwen-2.5 3B demonstrate up to $26\times$ and $28\times$ reduction in peak memory, enabling fine-tuning on resource-constrained devices.

13:00 JST研究/論文

イジングモデルに基づく適応型確率プロセッサの合成ツール

この研究では、組み合わせ最適化問題をイジング モデルにマッピングすることで解くための確率的アーキテクチャの合成とシミュレーションのためのツールを紹介します。提案されたアプローチは、イジング ハミルトニアンを自動的に構築し、サイズやトポロジーなどの問題の特性に基づいて確率要素 (p ビット) の数を決定します。さらに、このツールは、ギブズ サンプリング、シミュレーテッド アニーリング (SA)、シミュレーテッド 量子アニーリング (SQA)、およびクラスターベースの手法の中から最適な更新アルゴリズムを選択するための適応戦略を導入します。ベンチマーク問題を使用した実験結果は、固定アプローチと比較して収束動作と柔軟性が向上していることを示しています。提案されたフレームワークは、確率的コンピューティング戦略の系統的な評価を可能にし、MTJ と p ビットに基づく将来のハードウェア実装の開発をサポートします。

原文 (English)

A Tool for the Synthesis of Adaptive Probabilistic Processors Based on the Ising Model

This work presents a tool for the synthesis and simulation of probabilistic architectures for solving combinatorial optimization problems by mapping them to the Ising model. The proposed approach automatically constructs the Ising Hamiltonian and determines the number of probabilistic elements (p-bits) based on problem characteristics such as size and topology. Furthermore, the tool introduces an adaptive strategy for selecting the most suitable update algorithm among Gibbs Sampling, Simulated Annealing (SA), Simulated Quantum Annealing (SQA), and cluster-based methods. Experimental results using benchmark problems demonstrate improved convergence behavior and flexibility compared to fixed approaches. The proposed framework enables systematic evaluation of probabilistic computing strategies and supports the development of future hardware implementations based on MTJs and p-bits.

13:00 JSTLLM/生成AI画像/動画生成

PerceptionDLM: マルチモーダル拡散言語モデルを使用した並列領域認識

マルチモーダル大規模言語モデル (MLLM) は、視覚的な理解タスクにおいて目覚ましい進歩を遂げました。ただし、既存の MLLM のほとんどは自己回帰生成に依存しているため、複数の領域にキャプションを付ける必要がある認識タスクの効率が制限されます。この研究では、効率的な並列領域の認識のために最適化されたマルチモーダル拡散言語モデルである PerceptionDLM を提案します。オープンソースの拡散 MLLM の中で最先端のパフォーマンスを実現する強力な基礎ベースラインである PerceptionDLM-Base に基づいて構築された当社のアーキテクチャは、DLM の並列デコードの性質を最大限に活用しています。具体的には、効率的なプロンプトと構造化されたアテンション マスキングを導入して、複数のマスクされた領域を同時に認識できるようにし、モデルがシーケンス レベルとトークン レベルの両方で領域記述を並行して生成できるようにします。この設計により、領域を順次処理する既存のアプローチと比較して、推論効率が大幅に向上します。 DLM の視覚認識能力の並列性特性を体系的に評価するために、画像ごとに複数の領域マスクを含むように DLC ベンチをスケーリングすることにより、新しい Parallel Detailed Localized Captioning Benchmark (ParaDLC-Bench) を構築し、キャプションの品質と推論効率の両方の共同評価を可能にします。実験では、PerceptionDLM が領域キャプションで競争力のあるパフォーマンスを維持しながら、複数領域の認識タスクの速度を大幅に向上させることが実証されました。私たちの結果は、効率的で並列的な視覚認識のためのマルチモーダル拡散言語モデルの可能性を強調しています。私たちの知る限り、私たちは拡散言語モデルの利点を活用して並行領域のキャプションと知覚を実現した最初の企業です。コード、モデル、データセットがリリースされます。

原文 (English)

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that require captioning multiple regions. In this work, we propose PerceptionDLM, a multimodal diffusion language model optimized for efficient parallel region perception. Built upon PerceptionDLM-Base, a strong foundational baseline that achieves state-of-the-art performance among open-source diffusion MLLMs, our architecture fully leverages the parallel decoding nature of DLMs. Specifically, we introduce efficient prompting and structured attention masking to enable simultaneous perception of multiple masked regions, allowing the model to generate region descriptions in parallel at both the sequence and token levels. This design significantly improves inference efficiency compared with existing approaches that process regions sequentially. To systematically evaluate the parallelism property of visual perception capability for DLMs, we construct a new Parallel Detailed Localized Captioning Benchmark (ParaDLC-Bench) by scaling the DLC-Bench to include multiple region masks per image, enabling joint evaluation of both caption quality and inference efficiency. Experiments demonstrate that PerceptionDLM maintains competitive performance in region captioning while achieving substantial speed improvements for multi-region perception tasks. Our results highlight the potential of multimodal diffusion language models for efficient, parallel visual perception. To the best of our knowledge, we are the first to achieve parallel region caption and perception by leveraging the advantages of diffusion language models. Code, models, and datasets are released.

13:00 JST研究/論文

太陽エネルギー粒子予測のための機械学習モデルのレビュー

太陽エネルギー粒子(SEP)現象は、航空、宇宙船エレクトロニクス、および地球の磁気圏を超えた人類のミッションに対する重大な放射線の危険性のため、ますます注目を集めています。科学的な観点から見ると、SEP イベントは、太陽表面とコロナから太陽圏を通って広がる一連の物理プロセスから生じ、天体物理学全体に広く適用できる粒子の加速と輸送メカニズムについての洞察を提供するため、興味深いものです。したがって、SEP 現象を理解し予測する能力を向上させることは、そのようなメカニズムについての知識を深めるためにも、宇宙技術や探査を保護するためにも不可欠です。従来、研究者は物理ベースのシミュレーションと経験的手法を使用して SEP をモデル化してきました。最近では、SEP イベントを理解して予測するための新しいツールとして、機械学習 (ML) が登場しました。この原稿の目的は、SEP 予測用に現在利用可能な ML モデルをレビューし、トレーニングに使用されるデータセットを特定し、それらのアーキテクチャ、入力、出力を比較し、これらの洞察に基づいて、将来の研究のための優れた実践方法と推奨事項の概要を説明することです。

原文 (English)

Review of Machine Learning Models for Solar Energetic Particle Prediction

Solar energetic particle (SEP) events have attracted increasing attention due to their significant radiation hazards for aviation, spacecraft electronics, and human missions beyond Earth's magnetosphere. From a scientific perspective, SEP events are intriguing because they arise from a set of physical processes extending from the solar surface and corona through the heliosphere, offering insight into particle acceleration and transport mechanisms that are widely applicable across astrophysics. Therefore, advancing our ability to understand and predict SEP events is essential both for deepening our knowledge of such mechanisms and for safeguarding space technologies and exploration. Traditionally, researchers have modeled SEPs using physics-based simulations and empirical methods. More recently, machine learning (ML) has emerged as a new tool for understanding and predicting SEP events. The purpose of this manuscript is to review the currently available ML models for SEP prediction, identify the datasets used for training, compare their architectures, inputs, and outputs, and, based on these insights, outline good practices and recommendations for future research.

13:00 JST研究/論文

GDGU: 電気自動車充電ネットワークにおけるサイバー攻撃の位置特定のための勾配差分ベースのグラフ非学習法

電気自動車充電ステーション (EVCS) は、配電フィーダをサイバー攻撃にさらす可能性があります。グラフ ニューラル ネットワークなどの機械学習手法は、どのバスが侵害されているかを特定できますが、データ共有とモデル トレーニングには大きな課題が残っています。たとえば、プライバシー規制により、EVCS 所有者はデプロイされたモデルからトレーニング データを削除する権利が認められていますが、リクエストごとに最初から再トレーニングするのは計算量的に法外です。これに対処するために、EVCS サイバー攻撃の位置特定のためのグラフ非学習 (GU) を研究します。これは、グラフレベルのマルチラベル分類タスクに関する機能レベルの未学習問題として定式化されます。具体的には、一次パラメータ補正によって削除要求データの影響を除去する勾配差分ベースのグラフ非学習(GDGU)を提案します。補正は、元のトレーニング データと、要求された EVCS バスでの充電電力特徴のみが学習されていない変更されたデータセットの間の勾配の差から計算されます。次に、バッチ正規化再キャリブレーションと簡単な回復微調整ステップが適用され、ローカリゼーション ユーティリティが復元されます。 3 つのグラフ ニューラル ネットワーク バックボーンと累積的なアンラーニング シナリオにわたる、IEEE 34 バス、123 バス、および 8500 ノードの分散ネットワーク上の 2 つの 2 次 GU ベースラインに対して GDGU のベンチマークを行います。 GDGU は、ローカリゼーション ユーティリティの最も強力なベースラインと一致し、完全な再トレーニングに近い忘却忠実度に達しますが、ゼロから再トレーニングするよりも 10 ~ 12 倍速くアンラーニングし、2 次 GU ベースラインよりもはるかに少ないメモリを使用します。

原文 (English)

GDGU: A Gradient Difference-based Graph Unlearning Method for Cyberattack Localization in Electric Vehicle Charging Networks

Electric vehicle charging stations (EVCSs) can expose distribution feeders to cyberattacks. While machine learning methods, including graph neural networks, can localize which bus is compromised, significant challenges remain in data sharing and model training. For example, privacy regulations grant EVCS owners the right to delete their training data from a deployed model, yet retraining from scratch on every request is computationally prohibitive. To address this, we study graph unlearning (GU) for EVCS cyberattack localization, formulated as a feature-level unlearning problem on a graph-level multi-label classification task. Specifically, we propose gradient difference-based graph unlearning (GDGU), which removes the influence of the requested deletion data through a first-order parameter correction. The correction is computed from the gradient difference between the original training data and a modified dataset in which only the charging power features at the requested EVCS buses are unlearned. Then, a batch-normalization recalibration and a brief recovery fine-tuning step are applied to restore localization utility. We benchmark GDGU against two second-order GU baselines on the IEEE 34-bus, 123-bus, and 8500-node distribution networks across three graph neural network backbones and cumulative unlearning scenarios. GDGU matches the strongest baseline on localization utility and reaches forgetting fidelity close to full-retraining, while unlearning 10 to 12 times faster than retraining from scratch and using far less memory than the second-order GU baselines.

13:00 JST研究/論文

音響銃声分類のための特徴抽出技術パラメータの調査

音響による銃声の検出は、民間の公安、軍事作戦、野生生物の保護にわたるアプリケーションで問題となっていますが、この分野では、現実的なデータへの一般化に重点を置いた特徴抽出技術の厳密な調査が不足しています。市販の銃声検知システムと分類システムの有効性が混在していることは、現在の文献では適切に対処されていない未解決の問題を示しています。この論文では、85 丁の銃器と 21 口径にわたる 23,000 件の銃声記録のデータセットを使用した、共通の特徴抽出手法の体系的な調査を紹介します。 ResNet-18 を使用して、合計 12 個の一意のパラメーター セットを使用して 3 つの特徴抽出手法をベンチマークします。私たちの結果は、正しい特徴抽出手法を使用すると、トップ 1 の精度を最大 20% 向上させることができ、特定の特徴抽出手法の正しいパラメータを利用すると、その値を最大 4.7% 向上させることができることを示しています。

原文 (English)

Exploring Feature Extraction Technique Parameters for Acoustic Gunshot Classification

Acoustic gunshot detection is a problem with applications across civilian public safety, military operations, and wildlife conservation, yet the field lacks a rigorous exploration of feature extraction techniques with a focus on generalization to realistic data. The mixed effectiveness of commercial gunshot detection and classification systems indicates an open problem that is not adequately addressed by the current literature. In this paper, we present a systematic investigation of common feature extraction techniques using a dataset of 23,000 gunshot recordings across 85 firearms and 21 calibers. We benchmark three feature extraction techniques with 12 total unique parameter sets using ResNet-18. Our results demonstrate that using the correct feature extraction technique can improve top-1 accuracy by up to 20%, and utilizing the correct parameters for a given feature extraction technique can improve that value by up to 4.7%.

13:00 JST研究/論文

FlowFake: オーディオディープフェイク検出用のリキッドネットワーク

ニューラルテキスト読み上げシステムや音声クローンシステムによって生成された音声ディープフェイクは、大規模な話者の検証や公共の場での議論を脅かしています。中心的な課題は、データセット間の一般化です。1 つの合成パイプラインでトレーニングされた検出器は、目に見えない偽造に対して崩壊します。我々は、この失敗は主に、マルチタイムスケールの軌道異常である構造的な合成音声アーティファクトが原因であると主張します。既存の検出器はすべて固定ウィンドウ フレーム統計を集約しますが、これによりアーキテクチャと信号の調整がずれてしまいます。我々は、学習された ODE を介して隠れ状態が進化する液体時定数 (LTC) アーキテクチャである FlowFake を提案します。ニューロンごとの適応時定数がスペクトル (10 ミリ秒) と韻律 (2 秒) の手がかりを同時に解決します。わずか 34K のパラメータで、FlowFake は形式的な BIBO の安定性と O(dt^4) の積分誤差を達成します。 4 つのデータセットのクロスドメイン ベンチマーク (ASVspoof2019-LA、FakeOrReal、InTheWild、M​​LAAD) では、FlowFake は、FakeOrReal のみでトレーニングされた ASVspoof2019 で 75.29%、MLAAD のみでトレーニングされた 79.97% に達しました。評価されたすべてのペアおよび一致する SSL Wav2vec2 (300 倍) のパラメータ数の 0.01% で、RawGAT-ST および Whisper-DF よりも優れたパフォーマンスを発揮します。ソースコードはhttps://github.com/GhostRider2023/FlowFakeで入手できます。

原文 (English)

FlowFake: Liquid Networks for Audio Deepfake Detection

Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. The core challenge is cross-dataset generalization: detectors trained on one synthesis pipeline collapse on unseen forgeries. We argue that this failure is primarily because of structural synthetic speech artifacts which are multi-timescale trajectory anomalies. Though every existing detector aggregates a fixed-window frame statistics, this misaligns the architecture with the signal. We propose FlowFake, a Liquid Time-Constant (LTC) architecture whose hidden state evolves via a learned ODE, with per-neuron adaptive time constants simultaneously resolving spectral (10ms) and prosodic (2s) cues. At only 34K parameters FlowFake achieves formal BIBO stability and O(dt^4) integration error. On a four-dataset cross domain benchmark (ASVspoof2019-LA, FakeOrReal, InTheWild, MLAAD), FlowFake reaches 75.29% on ASVspoof2019 trained only on FakeOrReal and 79.97% trained only on MLAAD. It outperforms RawGAT-ST and Whisper-DF on every evaluated pair and matching SSL Wav2vec2 (300x larger) at 0.01% of its parameter count. The source code is available on : https://github.com/GhostRider2023/FlowFake

13:00 JSTLLM/生成AI

ベトナム語の抽象的な複数文書の要約のための階層戦略を備えた BART ベースのアプローチ

この技術レポートでは、2022 年のベトナム語言語および音声処理に関する国際ワークショップ (VLSP) で導入された、ベトナム語の複数文書の抽象的要約の課題の解決に焦点を当てています。私たちは、一般的な階層的アプローチ、つまり、各文書を圧縮してから集約と要約を行うことを選択します。我々は、ゴールデンサマリーに基づいて文書を短縮するための、斬新かつシンプルな戦略を提案します。これにより、階層的アプローチの段階間の高い相関関係が保証されます。私たちの方法は、VLSP の公開テスト セットで ROUGE2-F1 スコア 0.2468 を達成し、流暢で簡潔な要約を生成できます。さらに、追加データとして外部ソースを利用することで、ベトナム語の複数文書の要約のデータ量が大幅に向上します。追加データはコミュニティで利用できるようになります。

原文 (English)

A BART-based approach with hierarchical strategy for Vietnamese abstractive multi-document summarization

In this technical report, we focus on solving the challenge of Vietnamese multi-document abstractive summarization, introduced in the International Workshop on Vietnamese Language and Speech Processing (VLSP) 2022. We choose to follow the popular hierarchical approach, i.e. condensing each document followed by aggregation and summarization. We propose a novel yet simple strategy to shorten documents that is driven by the golden summary, thus ensuring high correlation between stages of the hierarchical approach. Our method achieves a ROUGE2-F1 score of 0.2468 on the VLSP's public test set, and can produce fluent and concise summaries. Additionally, we utilize external sources for extra data, which greatly enhances the quantity of data for Vietnamese multi-document summarization. The additional data is made available for the community.

13:00 JSTエージェントOpenAIGoogle

IHBench: 構造化されたワークフローを使用した音声エージェントの中断後の回復の評価

構造化されたワークフロー (顧客サービス、ヘルスケアのスケジュール設定、アカウント管理) に導入された音声エージェントは、複数ステップの手順で進行状況を維持しながら、頻繁なユーザーの中断に対処する必要があります。音声対応モデルの既存のベンチマークは、割り込みのタイミング、つまり割り込みの検出、エンドポイント、およびターンテイキングのダイナミクスに焦点を当てています。中断後に何が起こるかは未測定のままです。エージェントは正しいステップでワークフローを再開しますか?ユーザーの介入に対処していますか?ユーザーがすでに聞いたコンテンツの再配信を回避できますか? 10 のエンタープライズ ドメインにわたってステート マシン駆動のワークフローを実行する音声エージェントの中断後の回復を評価するベンチマークである IHBench (中断処理ベンチマーク) を紹介します。 6 つの中断タイプが発話の途中で制御されたポイントに挿入され、中断ごとの評価ルーブリックがデータとともに生成されます。各中断は、タスクの遂行と回復の質という 2 つの軸でスコア付けされます。 OpenAI、Google、およびオープンウェイト コミュニティからの 27 の音声言語モデル構成を評価します。モデルは多岐にわたり、リカバリの品質は中断の種類に大きく依存します。私たちの実験を通じて、クローズドウェイト モデルはオープンウェイト モデルよりも中断に対して一貫して堅牢です。クローズドウェイト モデルは、タスクの遂行においてはるかに頻繁に勝利し、会話が長くなるにつれ劣化が約 3.3 倍遅くなり、音声とテキストのモダリティのギャップが見られないのに対し、オープンウェイト モデルは 3 つすべてで劣勢です。人間による研究では、LLM ジャッジが人間のアノテーターに対して検証されており、AudioMultiChallenge に対するクロスベンチマーク分析では、回復品質が大きく異なる機能軸であることが示されています。

原文 (English)

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user interruptions while maintaining progress through multi-step procedures. Existing benchmarks for speech-capable models focus on the timing of interruptions: barge-in detection, endpointing, and turn-taking dynamics. They leave unmeasured what happens after the interruption: does the agent resume the workflow at the correct step? Does it address the user's interjection? Does it avoid re-delivering content the user already heard? We introduce IHBench (Interruption Handling Benchmark), a benchmark that evaluates post-interruption recovery in voice agents executing state-machine-driven workflows across 10 enterprise domains. Six interruption types are injected at controlled points mid-utterance, with per-interruption evaluation rubrics generated alongside the data. Each interruption is scored on two axes: task fulfillment and recovery quality. We evaluate 27 audio-language model configurations from OpenAI, Google, and the open-weight community. Models vary widely, and recovery quality depends strongly on the interruption type. Across our experiments, closed-weight models are consistently more robust to interruptions than open-weight ones: they win far more often on task fulfillment, degrade roughly 3.3x more slowly as conversations grow longer, and show no audio-versus-text modality gap, whereas the open-weight models lose ground on all three. A human study validates the LLM judge against human annotators, and a cross-benchmark analysis against AudioMultiChallenge indicates that recovery quality is a largely distinct capability axis.

13:00 JST研究/論文

PrefSQA: 音声品質評価のためのペアワイズ嗜好予測と高品質データセットの重要な役割

平均意見スコア (MOS) は音声品質の評価に広く使用されていますが、スカラー ラベルは評価者のばらつきやリスニング テストの違いに敏感です。これにより、ラベル付けノイズが発生し、MOS 予測の信頼性が制限されます。嗜好予測により、リスナーが信号を直接比較し、より明確なラベルが生成されるため、この変動が軽減されます。我々は、MOS フリーの選好予測を研究し、不確実性を認識したロジット、減損注意ヘッド、および非一致参照比較に基づくモジュールを組み込んだ PrefSQA を提案します。私たちは、MOS 由来の、コンテンツが一致するものと一致しないものを含む低ノイズのシミュレートされたセットを含む 5 つのデータセットを使用して改良し、人間の好みのセットを実験し、目に見えないデータでテストします。実験では、MOS 由来のデータでわずかな改善が示されましたが、他のセットではベースラインを超える明らかな改善が明らかになり、高品質の嗜好データの価値が強調され、提案された方法の有効性が実証されました。

原文 (English)

PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPT

FAPO: マルチステップ LLM パイプラインの完全自律型プロンプト最適化

マルチステップ LLM パイプラインは、取得、推論、およびフォーマットのステップ間の相互作用によって失敗するため、プロンプトのみの最適化ではチェーン内のボトルネックを見逃す可能性があります。 FAPO (Fully Autonomous Prompt Optimization) を紹介します。これは、Claude Code が標準化されたコードベース内の LLM パイプラインを最適化できるフレームワークです。 FAPO は、パイプラインを評価し、中間ステップを検査し、障害を診断し、範囲を絞った変更を提案し、スコア関数に対して最適化するためにバリアントを繰り返し検証します。最初にプロンプ​​ト編集を試み、プロンプトの最適化が不十分であると思われる場合にのみ、アトリビューションによって構造的なボトルネックが特定されたときに、許可された範囲内でチェーン構造を変更します。 6 つのベンチマークと 3 つのタスク モデルにわたって、FAPO は 18 のモデルとベンチマークの比較のうち 15 でベースライン GEPA を上回りました。 11 のモデルとベンチマークの比較では、重複しない平均 $\pm$ トライアル標準偏差範囲で FAPO が勝利し、FAPO-GEPA の平均ゲインは +14.1 pp でした。プロンプトファースト検索が構造変更にエスカレートした 6 回の HoVer と IFBench の比較では、FAPO が 6 回すべてで平均 +33.8 pp のゲインで勝利しました。FAPO はセキュリティ タスクのパフォーマンスも向上させています。 CVE から CWE へのタスク、プロンプトのみの FAPO は、GPT-5 で +4.0 pp、Foundation-Sec-8B-Instruct で +7.1 pp、Foundation-Sec-8B-Reasoning で +2.0 pp のテスト精度を向上させます。これらの結果により、FAPO は汎用タスクとセキュリティ重視のタスクの両方に対する最先端のパイプライン最適化手法として位置付けられます。

原文 (English)

FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines

Multi-step LLM pipelines fail through interactions among retrieval, reasoning, and formatting steps, so prompt-only optimization can miss bottlenecks in the chain. We present FAPO (Fully Autonomous Prompt Optimization), a framework that lets Claude Code optimize an LLM pipeline inside a standardized codebase. FAPO evaluates a pipeline, inspects intermediate steps, diagnoses failures, proposes scoped changes, and validates variants repeatedly to optimize against a score function. It first tries prompt edits and, only when prompt optimization appears insufficient, changes chain structure within the permitted scope when attribution identifies a structural bottleneck. Across six benchmarks and three task models, FAPO beats the baseline GEPA in 15 of 18 model-benchmark comparisons. In 11 model-benchmark comparisons, FAPO wins with non-overlapping mean $\pm$ trial-standard-deviation ranges, and the mean FAPO-GEPA gain is +14.1 pp. In the six HoVer and IFBench comparisons where prompt-first search escalated to structural changes, FAPO wins all six with a mean gain of +33.8 pp. FAPO also improves performance on security tasks: on CTIBench-RCM, a security CVE-to-CWE task, prompt-only FAPO lifts test accuracy by +4.0 pp on GPT-5, +7.1 pp on Foundation-Sec-8B-Instruct, and +2.0 pp on Foundation-Sec-8B-Reasoning. These results position FAPO as a state-of-the-art pipeline optimization technique for both general-purpose and security-focused tasks.

13:00 JST研究/論文

ライブラケット幾何学による潜在的な混乱した因果の発見

Kan-Do-Calculus (KDC) に関する最近の研究では、因果推論における受動的な観察と能動的な介入の間の境界は圏論的な二重付属関係であり、介入は左 Kan 拡張によってモデル化され、条件付けは右 Kan 拡張によって行われることが確立されました。この論文では、KDC の情報幾何学的およびカテゴリカルな結果に基づいて、潜在交絡の下で 2 つの因果発見アルゴリズムを紹介します。滑らかな統計設定では、観察的測定と介入的測定の間のラドン・ニコジム導関数は、局所的な因果ベクトル場を誘発します。これらのフィールドがリー括弧の下で閉じられないことは、計算可能なフロベニウス残差となり、これは、目に見える可分性の失敗と潜在的またはモデル化されていない構造の可能性の証拠として解釈されます。私たちの最初のアルゴリズムである BRIDGE (介入発見と幾何学的推定のためのブラケット残差) は、介入密度またはラドンニコジム比エンジンを幾何学的スクリーンと組み合わせて、許容可能な矢印の高再現率ファミリーを提案し、非閉鎖可視ペアを潜在障害候補として識別し、縮小されたファミリーを下流のスコアベースまたは微分可能な発見ルーチンに渡します。 2 番目のアルゴリズムの貢献であるスペクトル Kan-Do フロー マッチング (SKFM) は、償却介入場と潜在曲率をスペクトル的に学習し、BRIDGE が指す直接的なリー空間エンドポイントを明らかにします。詳細な一連の実験は、両方のアルゴリズムが、可能な DAG の超指数空間を何桁も崩壊させながら、潜在的な交絡因子を含む因果モデルを発見できることを示しています。この論文では、介入によって引き起こされる流れの形状から潜在構造を直接推測する、因果関係発見における新しいパラダイムを紹介します。

原文 (English)

Latent Confounded Causal Discovery via Lie Bracket Geometry

Recent work on Kan-Do-Calculus (KDC) has established that the boundary between passive observation and active intervention in causal inference is a category-theoretic bi-adjunction, with interventions modeled by left Kan extensions and conditioning by right Kan extensions. This paper introduces two causal discovery algorithms under latent confounding, building on the information-geometric and categorical consequences of KDC. In smooth statistical settings, Radon-Nikodym derivatives between observational and interventional measures induce local causal vector fields; failures of these fields to close under Lie brackets become computable Frobenius residuals, which we interpret as witnesses of failed visible integrability and possible latent or unmodeled structure. Our first algorithm, BRIDGE (Bracket Residuals for Interventional Discovery and Geometric Estimation), combines an interventional density or Radon-Nikodym-ratio engine with a geometric screen that proposes a high-recall family of admissible arrows, identifies non-closing visible pairs as latent-obstruction candidates, and passes the reduced family to downstream score-based or differentiable discovery routines. The second algorithmic contribution, Spectral Kan-Do Flow Matching (SKFM), learns amortized intervention fields and factors latent curvature spectrally, exposing the direct Lie-space endpoint toward which BRIDGE points. A detailed set of experiments show that both algorithms are capable of discovering causal models with latent confounders while collapsing the super-exponential space of possible DAGs by many orders of magnitude. This paper introduces a new paradigm in causal discovery, where latent structure is inferred directly from the geometry of intervention-induced flows.

13:00 JSTエージェント研究/論文

StaminaBench: 100 インタラクション ターンにわたるコーディング エージェントのストレス テスト

コーディング エージェントのスタミナを測定するベンチマークである StaminaBench を紹介します。失敗する前に、連続したインタラクション ターン (変更リクエスト) を何回処理できるかを測定します。一般的なタスク解決率の指標とは異なり、これはセッションが数十、数百ターン実行される実際のバイブコーディングと一致します。 StaminaBench では、エージェントが REST API サーバーを実装し、手続き的に生成された調整可能な数のフォローアップ変更リクエスト (実験では 100) にわたってそのサーバーを変更し、最大 6,000 行のコードベースが生成されます。テストは LLM を介さずに完全にプログラムで生成され、再現性と信頼性が保証されます。変更シーケンスは、ハードコードされたサンプラーまたは LLM 駆動のサンプラーから抽出され、どちらも変更が有効であることを保証するために構造化されたアクション スペースに制約されます。エージェントとサーバーは分離された環境で実行され、HTTP 経由でベンチマークと通信するため、完全にブラックボックスで言語に依存しないテストが可能になります。私たちは、7 つのオープンソース LLM と組み合わせた 6 つのエージェント ハーネスを、それぞれ 100 ターンの 20 のシナリオにわたって評価しました。その結果、次のことがわかりました。(1) テストされたモデルはすべて 5 ~ 6 ターン以内に失敗し、徹底的なテストを行わないバイブコーディング スタイルのプログラミングではバグが発生することが確認されました。 (2) テスト フィードバックをエージェントに返し、エージェントが再試行できるようにすると、合格したターン数が最大 12 倍向上します。 (3) 強力なパフォーマンスには優れたハーネスが必要です。強力なモデルでは最高のハーネスと最悪のハーネスの間に最大 6 倍の差が見られますが、弱いモデルではど​​のハーネスでも機能しません。マルチターン コーディング エージェントの動作に関するさらなる研究を可能にするために、ベンチマークと生成されたタスクをリリースします。ベンチマークコードとデータ: github.com/amazon-science/StaminaBench。

原文 (English)

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns

We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing. Unlike the prevailing fraction-of-tasks-solved metric, this matches real vibe-coding where sessions run dozens or hundreds of turns. In StaminaBench, agents implement a REST API server and modify it across a tunable number of procedurally generated follow-up change requests - 100 in our experiments, resulting in codebases of up to 6,000 lines. Tests are generated fully programmatically without LLM involvement, ensuring reproducibility and reliability; change sequences are drawn from either a hardcoded or LLM-driven sampler, both constrained to a structured action space to ensure changes are valid. The agent and the server run in an isolated environment and communicate with the benchmark through HTTP, making testing fully black-box and language-agnostic. We evaluate six agent harnesses paired with seven open-source LLMs across 20 scenarios of 100 turns each and find that: (1) all the tested models fail within 5-6 turns, confirming that vibe-coding-style programming without thorough testing produces bugs; (2) passing test feedback back to the agent and allowing it to retry improves passed turn count by up to 12x; and (3) a good harness is required for strong performance: stronger models exhibit up to a 6x gap between their best and worst harness, while weaker models fail with any harness. We release the benchmark and the generated tasks to enable further research into multi-turn coding agent behavior. Benchmark code and data: github.com/amazon-science/StaminaBench.

13:00 JSTエージェント

プル リクエストの前: マイニング マルチエージェントの調整

現在、自律型コーディング エージェントは何百万ものプル リクエストをオープンしていますが、大規模な調査では、PR の作成は速くなっているものの、受け入れられる頻度は低いことが判明しています。これは、プル リクエスト レベルのテレメトリでは説明できない調整と信頼のギャップです。私たちは、欠落シグナルは PR よりも前に存在しており、同時エージェントが共有作業をめぐってどのように主張し、分裂し、衝突するかという点に存在すると主張します。私たちは、中央サーバーを必要とせず、その記録を git 自体内に保存するオープンソースの調整基盤である Grite を通じてこのプロセスを研究しています。そのため、その追加専用の署名付きイベント ログが調整プロセスを直接キャプチャします。我々は、(i) この共有基板により、制限されたオーバーヘッドで重複した競合する作業が削減されます。チームメイトのタスクを単にやり直すだけの作業の割合は 78% から 0% に低下し、有用なスループットは 3 倍以上になります。 (ii) すべてのエージェントのログのコピーは、書き込みがサイレントにドロップされずに同じ状態に収束します。この場合、ファイルベースのトラッカーは同時書き込みを失います。 (iii) ログは採掘可能なアーティファクトであり、そこから具体的な障害モード (編集の競合、ロックの枯渇、冗長な再検出、クローズまでの競合) が来歴とともに自動的に回復可能であり、そのいくつかはプル リクエスト履歴には表示されません。データセット、ハーネス、マイニング ツールキットをリリースします。

原文 (English)

Before the Pull Request: Mining Multi-Agent Coordination

Autonomous coding agents now open millions of pull requests, yet large-scale studies find their PRs are produced faster but accepted less often - a coordination and trust gap that pull-request-level telemetry cannot explain. We argue the missing signal lives before the PR, in how concurrent agents claim, divide, and collide over shared work. We study this process through grite, our open-source coordination substrate that needs no central server and stores its records inside git itself, so its append-only, signed event log captures the coordination process directly. We show that (i) this shared substrate reduces duplicate and conflicting work at bounded overhead - the share of work that merely re-does a teammate's task falls from 78% to 0% while useful throughput more than triples; (ii) every agent's copy of the log converges to the same state with no write silently dropped, where a file-based tracker loses concurrent writes; and (iii) the log is a mineable artefact from which concrete failure modes - conflicting edits, lock starvation, redundant rediscovery, race-to-close - are automatically recoverable with provenance, several invisible in pull-request history. We release the dataset, harness, and mining toolkit.

13:00 JST研究/論文

VCG: 極端なコールド スタート条件下での電子商取引ビデオ フィードのマルチモーダル検索フレームワーク

デジタル コマースの状況は、静的な検索主導のカタログから、動的な没入型のビデオ フィードに移行しています。この移行により、「極端なコールドスタート」問題が発生します。つまり、従来のアイテムとは異なり、新しい短編ビデオには、協調フィルタリングに必要な高密度のインタラクション履歴が欠けています。さらに、イマーシブ フィードは、標準的なエンゲージメント信号を歪める強い位置と継続時間のバイアスをもたらします。このペーパーでは、大規模な電子商取引環境におけるこれらの課題を解決するために設計された、スケーラブルなマルチモーダル検索エンジンであるビデオ候補生成 (VCG) システムを実証します。ドメインに適応したビジョン言語モデル (CLIP に基づく) を活用することで、ユーザーとビデオを共有セマンティック空間にマッピングし、行動履歴ではなく視覚コンテンツに基づいたゼロショット検索を可能にします。システムのアーキテクチャを詳細に説明し、生成 (LLM) 埋め込みと識別 (CLIP) 埋め込みを比較する厳密な評価を示します。私たちの結果は、生成モデルは属性予測には優れているものの、検索タスクにおける埋め込み空間の崩壊に悩まされることを示しています。オンライン A/B テストでは、VCG がエンゲージメント バイアスを効果的に軽減し、ディープビデオの完成度が 50% 向上することが実証されました。システムの機能を紹介するために、製品からビデオ、ビデオから製品、およびゼロショット セマンティック検索の 3 つの双方向検索シナリオを特徴とする対話型のデモを紹介します。

原文 (English)

VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions

The digital commerce landscape is shifting from static, search-driven catalogs to dynamic, immersive video feeds. This transition introduces an ``extreme cold-start'' problem: unlike traditional items, new short-form videos lack the dense interaction history required for collaborative filtering. Furthermore, immersive feeds introduce strong position and duration biases that distort standard engagement signals. In this paper, we demonstrate the Video Candidate Generation (VCG) system, a scalable multimodal retrieval engine designed to solve these challenges in a large-scale e-commerce environment. By leveraging a domain-adapted vision-language model (based on CLIP), we map users and videos into a shared semantic space, enabling zero-shot retrieval based on visual content rather than behavioral history. We detail the system's architecture and present a rigorous evaluation comparing generative (LLM) vs. discriminative (CLIP) embeddings. Our results show that while generative models excel at attribute prediction, they suffer from embedding space collapse in retrieval tasks. Online A/B testing demonstrates that VCG effectively mitigates engagement biases, yielding a 50\% uplift in deep video completion. To showcase the system's capabilities, we present an interactive demonstration featuring three bi-directional retrieval scenarios: Product-to-Video, Video-to-Product, and Zero-Shot Semantic Search.

13:00 JST研究/論文

RIVET: 堅牢な冪等音声属性編集

音声属性編集モデルは、話者のアイデンティティを維持しながら、年齢や性別などの特性を変更します。ただし、大規模な音声データセットでは、属性アノテーションにノイズが多かったり一貫性がなかったりすることが多く、そのため条件付き生成モデルの編集が不安定になる可能性があります。この研究では、冪等性がノイズの多いラベルに対する堅牢性を向上させる効果的なメカニズムを提供することを示します。冪等演算子とは、繰り返し適用しても結果が変わらない演算子です。つまり、f(f(x)) = f(x) です。このプロパティを強制すると、ラベルが間違っている例に対する感度を下げる暗黙的な正則化機能として機能します。ラベル ノイズに対する堅牢性を向上させる冪等性目標を組み込んだトレーニング フレームワークである RIVET を紹介します。制御されたラベル ノイズの下で、自然にノイズのあるアノテーションを備えた GLOBE データセット上で RIVET を評価します。 RIVET は、標準のトレーニングよりも編集の成功率を向上させ、話者のアイデンティティをより良く保存します。これは、冪等性が音声編集モデルの堅牢性を向上させることを示しています。

原文 (English)

RIVET: Robust Idempotent Voice Attribute Editing

Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are often noisy or inconsistent, which can cause conditional generative models to produce unstable edits. In this work, we show that idempotency provides an effective mechanism for improving robustness to noisy labels. An idempotent operator is one for which repeated application does not change the result, i.e., f(f(x)) = f(x). Enforcing this property acts as an implicit regularizer that reduces sensitivity to mislabeled examples. We introduce RIVET, a training framework that incorporates an idempotency objective to improve robustness to label noise. We evaluate RIVET under controlled label noise and on the GLOBE dataset with naturally noisy annotations. RIVET improves editing success and better preserves speaker identity than standard training, showing that idempotency improves robustness in voice editing models.

13:00 JSTエージェントロボティクス

デシジョン ツリー蒸留による学習済みマルチエージェント コミュニケーション ポリシーの正式な検証

マルチエージェント強化学習 (MARL) により、エージェントは緊急通信を通じて調整戦略を開発できますが、ニューラル ポリシーには、ドローンの群れや自律走行車群での安全性が重要なロボットの展開に必要な正式な安全性の保証がありません。我々は、ポリシー抽象化を通じて学習したマルチエージェント通信ポリシーの安全性を検証するための最初のエンドツーエンドフレームワークを提示します。ニューラルポリシーは解釈可能な決定ツリーに抽出され、その後正式に検証され、検証された安全性プロパティが元のネットワークに転送されることを経験的検証で確認します。当社の 4 段階のパイプラインは、エージェントの観察からのドメイン固有の特徴抽出、ニューラル ポリシーに対する 97.9% +/- 1.2% の忠実度を達成するデシジョン ツリーの蒸留、特徴と状態変数の完全な対応による PRISM 確率モデル チェッカー仕様への自動変換、和集合結合集約と経験的近傍モデリングによるペアごとの分解による確率的計算ツリー ロジック (PCTL) プロパティの構成検証で構成されます。 5 ~ 7 台のエージェントによるマルチドローン調整のためのベクトル量子化変分情報ボトルネック (VQ-VIB) ポリシーを評価し、安全性、生存性、および協力にわたる 18 の時相論理プロパティを検証し、5 つの安全しきい値すべてが満たされ、特性満足度 88.9% を達成しました (衝突確率 0.3% 対 しきい値 1%)。元のニューラル ポリシーのモンテカルロ検証により、検証された安全性特性が <=0.6 パーセンテージ ポイントの偏差 (95% CI) で移行することが確認されています。離散 VQ-VIB メッセージは、連続方式と比べて忠実度が +11.6 ~ +13.6 パーセントポイント向上し、3 ~ 4 倍高速な検証が可能になります。私たちのフレームワークは、抽出されたポリシー抽象化に対して経験的に検証された安全性検証を提供し、マルチロボット導入のための深い MARL と正式な安全性ワークフローの間の実用的な橋渡しとして機能します。

原文 (English)

Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation

Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets. We present the first end-to-end framework for safety verification of learned multi-agent communication policies through policy abstraction: neural policies are distilled into interpretable decision trees, then formally verified, with empirical validation confirming that verified safety properties transfer to original networks. Our four-stage pipeline consists of domain-specific feature extraction from agent observations, decision tree distillation achieving 97.9% +/- 1.2% fidelity to neural policies, automated translation to PRISM probabilistic model checker specifications with complete feature-to-state-variable correspondence, and compositional verification of Probabilistic Computation Tree Logic (PCTL) properties via pairwise decomposition with union-bound aggregation and empirical neighbor modeling. Evaluating Vector-Quantized Variational Information Bottleneck (VQ-VIB) policies for multi-drone coordination with 5-7 agents, we verify 18 temporal logic properties across safety, liveness, and cooperation, achieving 88.9% property satisfaction with all five safety thresholds satisfied (0.3% collision probability vs. 1% threshold). Monte Carlo validation of original neural policies confirms that verified safety properties transfer with <=0.6 percentage-point deviation (95% CI). Discrete VQ-VIB messages provide +11.6 to +13.6 percentage-point fidelity advantages over continuous methods, enabling 3-4x faster verification. Our framework provides empirically validated safety verification for distilled policy abstractions, serving as a practical bridge between deep MARL and formal safety workflows for multi-robot deployment.

13:00 JSTロボティクス

CTS-MoE: 知覚的移動のための専門家の混合による暗黙的な地形適応

不連続な地形(階段、隙間、障害物など)上での知覚的な脚の移動には、適応的な行動が必要です。これは、単一の保守的な歩行では、突然のトポロジーの変化に必要な予測的な操作を生み出すことができないためです。マルチタスクの強化学習として見られるこの問題は、共有と分離の間に緊張をもたらします。タスクは共通の移動ベースを使用しますが、報酬が矛盾するため、ポリシーは値の干渉を回避しながら動作を共有する必要があります。これまでの研究では、一枚岩のポリシーは専門性を犠牲にし、階層的なサブポリシーは移行や目に見えない領域にわたる一般化を犠牲にして、片面のみを扱っていました。我々は、専門家の高密度混合アクターと認識ベースのゲーティングを組み合わせて共有行動を構成し、タスク固有の価値観を持つ複数の批評家を組み合わせて干渉を防止する CTS-MoE を提案します。このモデルは、部分的な可観測性を処理し、逐次抽出を回避する単一ステージの教師と生徒の同時セットアップでエンドツーエンドでトレーニングされ、タスク ラベルはトレーニング中にのみ使用されます。導入時には、ルーティングは知覚のみに依存するため、高レベルのセレクターや地形分類子を使用せずに地形適応が可能になります。シミュレーションにおける Unitree Go1 と、目に見える地形と見えない地形にわたるハードウェアでの実験では、モノリシック ベースラインよりも追跡エラーが低く、成功率が高い、タスク認識型の特化が示されています。プロジェクト Web サイト: https://cts-moe.github.io/ 。

原文 (English)

CTS-MoE: Implicit Terrain Adaptation via Mixture-of-Experts for Perceptive Locomotion

Perceptive legged locomotion over discontinuous terrain (e.g., stairs, gaps, and obstacles) requires adaptive behavior, as a single conservative gait cannot produce the anticipatory maneuvers needed for abrupt topology changes. Cast as multi-task reinforcement learning, this problem introduces a tension between sharing and separation. Tasks use a common locomotion base but have conflicting rewards, so a policy must share behavior while avoiding value interference. Prior work addresses only one side, with monolithic policies sacrificing specialization and hierarchical sub-policies sacrificing generalization across transitions and unseen terrain. We propose CTS-MoE, which combines a dense mixture-of-experts actor with perception-based gating to compose shared behaviors and a multi-critic with task-specific value heads to prevent interference. The model is trained end-to-end in a single-stage concurrent teacher-student setup that handles partial observability and avoids sequential distillation, with task labels used only during training. At deployment, routing depends solely on perception, allowing terrain adaptation without a high-level selector or terrain classifier. Experiments on a Unitree Go1 in simulation and on hardware across seen and unseen terrains show task-aware specialization, with lower tracking error and higher success rates than monolithic baselines. Project Website: https://cts-moe.github.io/ .

13:00 JST研究/論文

トークンファクトリー: 多様なシグナルを大規模なレコメンデーションモデルに効率的に統合

大規模レコメンデーション モデル (LRM) は、業界規模のレコメンデーション タスクにおいて有望な機能を実証しています。ただし、従来の信号をこれらのトランスベースのアーキテクチャに効果的かつ効率的に統合することは依然として大きな課題です。これらの信号を直接「テキスト化」したり、個別のアイテム表現を作成したりする従来のアプローチでは、プロンプトが過度に長くなり、メモリ使用量が大きくなり、計算オーバーヘッドが高くなることがよくあります。これらの制限を克服するために、私たちは従来のシグナルを LRM によって直接処理できる「ソフト トークン」に変換するように設計されたフレームワークである「トークン ファクトリー」を提案します。このアプローチにより、異種入力フィーチャの効率的な統合と圧縮が可能になり、モデルのパフォーマンスを向上させながら、急激な長さの爆発を防ぐことができます。 Token Factory のアーキテクチャを詳しく説明し、実稼働規模のレコメンデーション環境でのその有効性を検証する実験結果を示します。

原文 (English)

Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models

Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically integrating traditional signals into these transformer-based architectures effectively and efficiently remains a major challenge. Conventional approaches that "textualize" these signals directly or create discrete item representations often lead to excessively long prompts, substantial memory footprints, and high computational overhead. To overcome these limitations, we propose "Token Factory", a framework designed to transform traditional signals into "soft tokens" that can be directly processed by LRMs. This approach enables efficient integration and compression of heterogeneous input features, preventing prompt length explosion while enhancing model performance. We detail the architecture of Token Factory and present experimental results validating its effectiveness in a production-scale recommendation environment.

13:00 JST研究/論文

難しいですか、それとも到達していないだけですか?数学的推論の難易度推定におけるサンプリングの盲点を診断する

数学と科学の推論ベンチマークは、標準的な例ごとの難易度シグナルとして、ゴールドに達するサンプリングされたチェーンの割合である pass@k に依存します。同じシグナルが、検証可能な報酬、数学データのキュレーション、合成カリキュラム、検証者のトレーニングによる RL を推進します。このプロキシには、最も難しい層に永続的な盲点があることがわかります。テストした 8 つの自由形式数学セル (4 つのオープンウェイト モデルにわたる GSM8K および MATH) では、サンプリング シードが 6 回の試行で解決できなかった例の 10.3 ~ 22.9% が、代わりに 6 チェーンの決定論的体制による一致したコンピューティングで解決されました。これらは、貪欲なデコードにアクティベーション グラフティングを介して適用される 5 つの安価な残差ストリーム摂動を加えたものですが、貪欲だけではこれらの数学セルでは最大 6% しか解決できません。回復は追加の予算に応じて調整され、摂動全体でそのメカニズムの区別性が 12 個のセルすべてで検証されます (すべての設定で異種固定セット Jaccard <= 0.47)。アクティベーショングラフティングは、解読方法ではなく、内部表現への介入として使用されます。私たちはそれを純粋に診断および多様化ツールとして使用しており、復元されたアイテムは、変更されていないモデルが通常の推論で到達するのではなく、 pass@k= 0 % 層が残差ストリーム内で構造的に識別可能であることを示しています。

原文 (English)

Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation

Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal. The same signal drives RL with verifiable rewards, math data curation, synthetic curricula, and verifier training. We show this proxy has a persistent blind spot on its hardest stratum: on the eight free-form math cells we test (GSM8K and MATH across four open-weight models), 10.3-22.9% of the examples that no sampling seed solves in six tries are instead solved at matched compute by a six-chain deterministic regime. These are greedy decoding plus five cheap residual-stream perturbations applied via activation grafting, while greedy alone solves at most 6% on these math cells. Recovery scales with the additional budget, across perturbations whose mechanistic distinctness we verify across all twelve cells (cross-kind fix-set Jaccard <= 0.47 in every setup). Activation grafting is used as an intervention on internal representations, not a decoding method; we use it purely as a diagnostic and diversification tool, and our recovered items show that the pass@k= 0 % stratum is structurally identifiable in the residual stream rather than that the unmodified model reaches them under ordinary inference.

13:00 JSTLLM/生成AI

ラベルの前に: データセット構築が臨床テキストにおける自殺傾向の検出をどのように形作るか

臨床 NLP は、自殺行動を検出するために電子医療記録 (EHR) データにますます依存しており、臨床文書をソーシャル メディアよりも信頼できる真実として扱っています。私たちは、この枠組みによって、EHR に基づく自殺傾向のデータセットが、データの作成者、エピソードの制限方法、曖昧さの解決方法によって形づくられる、自殺傾向の特定の運用化をどのようにエンコードするのかが曖昧になると主張します。私たちは、MIMIC-III 臨床ノートに基づいて構築された ScAN データセットのケーススタディでこの議論を根拠にしています。我々は、ガバナンスの制約、ICDに基づくコホート選択、単一アノテーターのラベル付け、および入院レベルの集計が、臨床医によって文書化された判断を反映するラベルをどのように生成し、自殺傾向を限定されたエピソードとして扱い、意図が文書から確実に推測できると仮定するかを示します。言語分析により、同一のラベルが、時間性、否定性、不確実性の点で異なる異種の臨床的枠組みを包含していることが実証されています。私たちは、臨床 NLP では、ラベルをグラウンド トゥルースとして解釈する前に、自殺傾向のデータセットに埋め込まれた仮定を検査する必要があると主張します。

原文 (English)

Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text

Clinical NLP increasingly relies on electronic health record (EHR) data to detect suicidal behaviors, treating clinical documentation as more reliable ground truth than social media. We argue that this framing obscures how EHR-based suicidality datasets encode a particular operationalization of suicidality, shaped by who authors the data, how episodes are bounded, and how ambiguity is resolved. We ground this argument in a case study of the ScAN dataset, built over MIMIC-III clinical notes. We show how governance constraints, ICD-based cohort selection, single-annotator labeling, and hospital-stay-level aggregation produce labels that reflect clinician-documented judgments, treat suicidality as a bounded episode, and assume that intent can be reliably inferred from documentation. A linguistic analysis demonstrates that identical labels subsume heterogeneous clinical framings differing in temporality, negation, and uncertainty. We argue that clinical NLP should examine the assumptions embedded in suicidality datasets before interpreting their labels as ground truth.

13:00 JSTLLM/生成AI

多言語メンタルヘルス対話データセットの作成: 国籍と言語によるペルソナベースのローカリゼーションの限界

AI と大規模言語モデル (LLM) は、世界的なメンタルヘルスの課題に対処するための有望なツールとして浮上しています。これらの課題は世界的なものであるにもかかわらず、そのようなシステムのトレーニングと評価のための高品質のデータセットが依然として重大な不足を抱えています。このギャップを軽減するために、研究者はユーザーデータをシミュレートし、デジタルメンタルヘルスサポートシステムをテストするために合成臨床ペルソナを生成することが増えています。ただし、検証されたペルソナのほとんどは英語中心のコンテキストに依存しています。この論文では、同様のペルソナベースの手法を使用して多言語メンタルヘルス データセットを生成できるかどうかを調査します。ペルソナの国籍と言語パラメータを変更して、中国語、ベンガル語、ヒンディー語で臨床対話を生成しました。次に、これらの生成された多言語データセットのうつ病の重症度を英語のベースラインと比較して評価するときに、さまざまな LLM がどのように機能するかを調べました。私たちの調査結果は、ペルソナに国籍と言語パラメータを追加するだけでは、言語間で臨床的な不一致が生じる可能性があるため、適切ではない可能性があることを示しています。 LLM 判定モデルは、英語以外のテキストでうつ病の重症度を評価する際に不正確さを示すことが多く、モデルによってパフォーマンスが異なります。これは、英語中心のペルソナを多言語の文脈に適用することの体系的な限界を明らかにします。最終的に、私たちの研究は、世界的に公平なメンタルヘルス システムを確保するために、文化に応じたデータ生成の緊急の必要性を浮き彫りにしています。

原文 (English)

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language

AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges. Despite the global nature of these challenges, there remains a critical shortage of high-quality datasets for training and evaluating such systems. To mitigate this gap, researchers increasingly generate synthetic clinical personas to simulate user data and test digital mental health support systems. However, most validated personas rely on English-centric contexts. This paper investigates whether similar persona-based methods can be used to generate multilingual mental health datasets. We modified nationality and language parameters in personas to generate clinical dialogues in Mandarin, Bengali, and Hindi. We then examined how different LLMs perform when evaluating the depression severity of these generated multilingual datasets against the baseline in English. Our findings indicate that just adding nationality and language parameters in personas might not be adequate, as it can introduce clinical inconsistency across languages. LLM judge models often exhibit inaccuracies in assessing depression severity in non-English texts, with performance varying across different models. This exposes the systemic limitations of applying English-centric personas to multilingual contexts. Ultimately, our work highlights the urgent need for culturally responsive data generation to ensure equitable mental health systems globally.

13:00 JST画像/動画生成

TeleMorpher: 堅牢なモーションと位置の同時編集を目指して

普及モデルは、画像とビデオの生成と編集において目覚ましい成功を収めています。最近の研究では、これらの取り組みがモーション編集に拡張されていますが、モーションと位置の両方を同時に変換することは、実際的な重要性にも関わらず、ほとんど解明されていないままです。ロバストなモーション位置編集をより深く理解するために、まず品質を低下させる基本的な要因を分析します。この分析に基づいて、私たちはモーションと位置を同時に編集するための、私たちの知る限り最初のワンショット フレームワークの 1 つである TeleMorpher を提案します。私たちのアプローチは、モーション プリア、モーション編集ガイダンスとして既製のモデルから生成されたターゲット モーション中心のビデオ、およびグラウンド トゥルース モーションを活用して、より制御可能で正確なモーション位置編集を可能にします。これにより、私たちのフレームワークは次のように機能します: (1) まず、事前にトレーニングされたセグメンテーションと修復モデルによって主人公と背景を解きほぐします。 (2) 次に、事前のモーションをガイドとして主人公のモーションを編集するトレーニング不要のポーズワーピングを導入します。 (3) ワープされたモーション ビデオの結果は、推論中にベースライン モーション エディタに直接挿入され、ソース ビデオの外観を維持しながら、ソース モーションとターゲット モーションの差を軽減します。 (4) 定量的評価の信頼性を高めるために、モーション編集の前後の背景の一貫性と、ソースビデオとターゲットビデオから抽出された主人公のスケルトンの違いを測定することによってモーション編集パフォーマンスの忠実性を測定する、2つの新しいLPIPSベースの指標を提案します。野外ビデオと太極拳データセットを使用した実験では、TeleMorpher が定量的測定と定性的測定 (実際の人間による評価) の両方で優れたパフォーマンスを達成することが実証され、その有効性が強調されています。

原文 (English)

TeleMorpher: Toward Robust Simultaneous Motion-Location Editing

Diffusion models have achieved remarkable success in image and video generation and editing. While recent studies have extended these efforts toward motion editing, simultaneously transforming both motion and location-despite its practical importance-remains largely unexplored. To better understand robust motion-location editing, we first analyze the fundamental factors that degrade its quality. Based on this analysis, we propose TeleMorpher, one of the first one-shot frameworks to the best of our knowledge, for simultaneous motion-location editing. Our approach leverages motion priors, a target motion-centric video generated from an off-the-shelf model as motion-editing guidance, and the ground truth motion to enable more controllable and precise motion-location editing. Via this, our framework works as follows: (1) we first disentangle the protagonist and the background via pre-trained segmentation and inpainting models. (2) Then, we introduce a training-free pose warping that edits the protagonist's motion with the motion prior as the guidance. (3) The result of warped motion video is directly injected into a baseline motion editor during inference, mitigating the difference between source and target motions while preserving the appearance of the source video. (4) To enhance the reliability of quantitative evaluations, we propose two new LPIPS-based metrics that measure the background consistency before and after the motion editing and the fidelity of motion editing performance via measuring the difference between the extracted protagonist's skeletons from source and target videos. Experiments with in-the-wild videos and the TaiChi dataset demonstrate that TeleMorpher achieves superior performance across both quantitative and qualitative measurements (real-human evaluation), underscoring its effectiveness.

13:00 JST研究/論文

LOKI: メモリフリーのヌルスペース制約のある生涯知識編集

生涯知識編集は、過去の知識で許容可能なパフォーマンスを維持しながら、新しい知識が利用可能になったとき、またはモデルが間違いを犯したときに、言語モデルを時間の経過とともに効率的かつ順次更新することを目的としています。未解決の課題の 1 つは、既存の方法ではすべての新しい知識サンプルの固定セットのレイヤーが変更され、柔軟性が低下し、壊滅的な忘却が増加することです。もう 1 つは、データ統計を取得するために、以前の知識へのアクセスと広範な前処理を必要とすることです。これらの課題に対処するために、ヒルベルト・シュミット独立基準に基づく動的な層選択を使用し、勾配更新をモデル重みのヌル空間に投影する新しいアプローチである LOKI を導入し、以前の知識へのアクセスの要件を回避します。 LOKI がさまざまな実験にわたって既存のアプローチよりも優れたパフォーマンスを達成し、平均精度で最大 14\% の向上を達成することを示します。

原文 (English)

LOKI: Memory-Free Null-Space Constrained Lifelong Knowledge Editing

Lifelong knowledge editing aims to efficiently and sequentially update language models over time, as new knowledge becomes available or when the model makes mistakes, while preserving acceptable performance on past knowledge. One unresolved challenge is that existing methods modify a fixed set of layers for all new knowledge samples, reducing flexibility and increasing catastrophic forgetting. Another is requiring access to previous knowledge and extensive pre-processing to obtain data statistics. To address these challenges, we introduce LOKI, a novel approach that uses dynamic layer selection based on the Hilbert-Schmidt Independence Criterion and projects gradient updates onto the null-space of the model weights, bypassing the requirement for previous knowledge access. We show that LOKI achieves superior performance to existing approaches across a wide variety of experiments, achieving up to a 14\% improvement in average accuracy.

13:00 JSTLLM/生成AIハードウェア/半導体

思考連鎖トランスフォーマーによるアルゴリズムの効率的な表現

\emph{reasoning} モデル (答えを生成する前に一連の推論または思考トークンを出力する言語モデル) の人気が高まっていることは、思考連鎖 (CoT) 変換器がチューリング マシンをシミュレートし、任意の計算を実行できることを示す理論的結果によって部分的に正当化されます。ただし、チューリング マシンは複雑性理論の分析には適していますが、アルゴリズムを議論するのには便利ではなく、直観的でもなく、効率的でもありません。通常、アルゴリズムはより高い抽象レベルで設計および分析され、ランダム アクセス メモリと $\bigO(\log n)$ ビット ワードに対する単位コスト演算を備えた \emph{Word RAM} モデルによってキャプチャされます。その結果、Word RAM アルゴリズムはチューリング マシンのアルゴリズムより大幅に効率的になる可能性があり、\emph{CoT トランスフォーマーは Word RAM アルゴリズムを効率的にシミュレートできますか?} たとえば、$n$ 個の項目を $\bigO(n \log n)$ ステップでソートしたり、ダイクストラのアルゴリズムを $\bigO(E + V \log V)$ ステップで実行したりできるでしょうか? という疑問が生じます。多対数オーバーヘッドまでは肯定的に答えます。まず、多対数幅と右端の一意のハード アテンションを持つ有限精度変換器に対してこれを確立し、次にその結果を、有限幅と対数精度を持つ 2 つのより実用的な設定に強化します。\emph{continuous} CoT (推論がトークンではなくベクトルの形式をとる場合) と、変換器層がリカレント (線形 RNN) 層の上に位置する \emph{hybrid} アーキテクチャです。 3 つのケースすべてにおいて、CoT \emph{できる} が、$n$ の多対数オーバーヘッドのみであらゆる Word RAM アルゴリズムを効率的にシミュレートできることがわかります。 Word RAM に「フラット」命令セットがある場合、このオーバーヘッドは対数二乗に減少し、乗算のないフラット命令の場合は対数のみになります。これは、Word RAM に対して 2 次オーバーヘッドを必要とする既知のチューリング マシンの CoT シミュレーションとはまったく対照的です。

原文 (English)

Efficiently Representing Algorithms With Chain-of-Thought Transformers

The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producing an answer -- is justified, in part, by theoretical results showing that chain-of-thought (CoT) transformers can simulate Turing machines, and thus perform arbitrary computation. However, the Turing machine, while suitable for complexity-theoretic analysis, is not convenient, intuitive, or efficient for discussing algorithms. Algorithms are typically designed and analyzed at a higher level of abstraction, captured by the \emph{Word RAM} model with random-access memory and unit-cost operations on $\bigO(\log n)$-bit words. As a result, Word RAM algorithms can be substantially more efficient than their Turing machine counterparts, raising the question: \emph{Can CoT transformers efficiently simulate Word RAM algorithms?} For instance, can they sort $n$ items in $\bigO(n \log n)$ steps or run Dijkstra's algorithm in $\bigO(E + V \log V)$ steps? We answer affirmatively, up to poly-logarithmic overhead. We first establish this for finite-precision transformers with poly-logarithmic width and rightmost unique hard attention, then strengthen the result to two more practical settings with finite width and log-precision: \emph{continuous} CoT, where reasoning takes the form of vectors rather than tokens, and a \emph{hybrid} architecture in which transformer layers sit atop a recurrent (linear RNN) layer. In all three cases, we find that CoT \emph{can} efficiently simulate any Word RAM algorithm with only a poly-logarithmic overhead in $n$. This overhead reduces to log-square when the Word RAM has a ``flat'' instruction set, and only logarithmic for multiplication-free flat instructions -- in stark contrast to known CoT simulations of Turing machines, which require quadratic overhead over Word RAM.

13:00 JSTLLM/生成AI

FineREX: 人間の密輸ナレッジ グラフ用に微調整された NER-RE

法廷手続きには人身密航ネットワークに関する貴重な証拠が含まれていますが、この情報は構造化されておらず、専門用語が多い法的文書の中に埋もれていることがよくあります。大規模言語モデル (LLM) は、自動化された情報抽出を通じてナレッジ グラフの構築をサポートできますが、既存のアプローチは、このドメインで必要とされるエンティティおよび関係の定義に合わせて調整されていない汎用モデルに依存しています。 FineREX は、固有表現認識および関係抽出 (NER-RE) 用に微調整された LLM を中心に構築された、合理化されたナレッジ グラフ構築パイプラインです。 FineREX は、手動で注釈が付けられた $512$ のテキスト チャンクのデータセットを使用して、大規模な汎用ベースラインと比較して、エンティティとリレーションシップの F1 スコアでそれぞれ 15.50% と 31.46% の絶対的な改善を達成しました。これらの利点は、より高品質のナレッジ グラフに変換され、リーガル ノイズがほぼ半分に削減され、長いドキュメントでのノードの重複が 17.78% から 11.17% に減少します。 FineREX は、ドキュメントの書き換えと冗長な抽出ステージを排除することで、エンドツーエンドの処理時間を 50.0% 削減します。私たちの結果は、ドメイン固有の微調整が、不正なネットワーク分析のためのナレッジ グラフ構築の品質と効率の両方を向上させながら、大規模な汎用モデルよりも大幅に優れたパフォーマンスを発揮できることを示しています。

原文 (English)

FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs

Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargon-heavy legal documents. While large language models (LLMs) can support knowledge graph construction through automated information extraction, existing approaches rely on general-purpose models that are not tailored to the entity and relationship definitions required in this domain. We introduce FineREX, a streamlined knowledge graph construction pipeline built around a fine-tuned LLM for named entity recognition and relationship extraction (NER-RE). Using a manually annotated dataset of $512$ text chunks, FineREX achieves absolute improvements of 15.50% and 31.46% in entity and relationship F1-score, respectively, compared to a larger general-purpose baseline. These gains translate into higher-quality knowledge graphs, reducing legal noise by nearly half and lowering node duplication on long documents from 17.78% to 11.17%. By eliminating document rewriting and redundant extraction stages, FineREX also reduces end-to-end processing time by 50.0%. Our results demonstrate that domain-specific fine-tuning can substantially outperform larger general-purpose models while improving both the quality and efficiency of knowledge graph construction for illicit network analysis.

13:00 JSTLLM/生成AIビジネス/資金調達

AURA: LLM-as-a-Judge 監査のための不確実性を考慮した適応的改良

大規模言語モデル (LLM) は、人間による大規模な評価は費用がかかり、拡張するのが難しい場合が多いため、オープンエンド生成の判断材料としてますます使用されていますが、その好みは依然として人間の判断に対する不完全な代用です。既存の監査パイプラインは、たとえば人間の注釈、ヒューリスティック フィルタリング、または強力な審査員の出力などから、信頼できる例のサブセットまたはクリーンな監視信号が事前に利用可能であることを前提としていることがよくあります。 LLM の評価では、この仮定は脆弱です。最初の分割では裁判官のバイアスが引き継がれる可能性がありますが、通常、人間による検証は不足しすぎて、大規模に安定したグループを定義できません。私たちは、選択された人間による検証の下で、ペアごとの LLM を審査員として監査するための適応的不確実性を認識した改良フレームワークである AURA を提案します。 AURA は人間による一貫性シグナルを繰り返し学習し、信頼できる証拠を広め、人間によるレビューのために不確実な比較を優先します。重要な考え方は、裁判官に対する信頼を、証拠が蓄積されるにつれて徐々に洗練される潜在的な量として扱うことです。当社は、コンパクトな定式化、安定した改良手順、および合成および実際のペアワイズ LLM 応答データの両方に対する包括的な評価を提供します。

原文 (English)

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human judgment. Existing auditing pipelines often assume that a reliable subset of examples or clean supervision signals are available beforehand, for example from human annotation, heuristic filtering, or the outputs of strong judges. In LLM evaluation, this assumption is fragile: the initial split may inherit judge bias, while human verification is typically too scarce to define stable groups at scale. We propose AURA, an adaptive uncertainty--aware refinement framework for auditing pairwise LLM--as--a--judge decisions under selected human verification. AURA iteratively learns a human-consistency signal, propagates reliable evidence, and prioritizes uncertain comparisons for human review. The key idea is to treat trust in a judge as a latent quantity that is progressively refined as evidence accumulates. We provide a compact formulation, a stable refinement procedure, and a comprehensive evaluation on both synthetic and real pairwise LLM-answer data.

13:00 JST研究/論文

OnDeFog: フレームドロップ下のオンライン意思決定トランスフォーマー

困難な現実世界の強化学習アプリケーションでは、通信の遅延やセンサーの故障によりフレーム ドロップが発生することが多く、エージェントはドロップされた状態や関連する報酬を受け取ることができません。フレーム ドロップによるパフォーマンスの低下に対処するために、フレーム ドロップに対処するための追加メカニズムをデシジョン トランスフォーマーに組み込むことにより、ランダム フレーム ドロップ下のデシジョン トランスフォーマー (DeFog) が開発されました。 DeFog はフレームドロップ環境でのパフォーマンスの低下を軽減できますが、DeFog はオフライン学習方法であるため、トレーニング データセットで適切に表現されていない新しい状態に効果的に一般化するのに苦労します。本研究では、DeFog のメカニズムと、環境との直接的な相互作用を通じてポリシーを学習するオンライン強化学習手法であるオンライン意思決定変換器 (ODT) を統合した OnDeFog を提案します。包括的な実験評価により、私たちが提案した OnDeFog は、フレーム レートの低下が特徴の環境で ODT と比較して優れたパフォーマンスを達成し、大量の低報酬データを含むデータセットでは DeFog よりも優れたパフォーマンスを発揮することが実証されました。

原文 (English)

OnDeFog: Online Decision Transformer under Frame Dropping

In challenging real-world reinforcement learning applications, communication delays or sensor failures often cause frame dropping, in which the agent cannot receive the dropped states and associated rewards. To address the performance degradation caused by frame dropping, the Decision Transformer under Random Frame Dropping (DeFog) was developed by incorporating additional mechanisms into the decision transformer to tackle frame dropping. Although DeFog can mitigate performance degradation in frame-dropping environments, since DeFog is an offline learning method, it struggles to effectively generalize to novel states not adequately represented in the training dataset. In this study, we propose OnDeFog, which integrates the mechanisms in DeFog with the online decision transformer (ODT), an online reinforcement learning method that learns policies through direct environmental interaction. Comprehensive experimental evaluation demonstrates that our proposed OnDeFog achieves superior performance compared to ODT in environments characterized by high dropping frame rate and outperforms DeFog on datasets containing a large amount of low-reward data.

13:00 JST研究/論文

OpenSIL ファームウェアでの大規模言語モデル生成単体テストのライブラリ認識ダブルと反復修復

単体テスト (UT) は厳しいビルド制約の下では脆弱であり、ヘッダーの欠落、未解決のシンボル、依存関係の不一致によってコンパイルやリンクが頻繁に妨げられるため、低レベル C ファームウェアの変更の検証にはコストがかかります。この研究では、Advanced Micro Devices (AMD) が管理するオープンソース シリコン初期化ライブラリ (openSIL) ファームウェア コードベースの自動 UT オーサリング ワークフローを紹介します。これにより、大規模言語モデル (LLM) ガイドのマルチエージェント パイプラインによって手動の労力が軽減されます。このワークフローは、テスト スキャフォールドの自動生成、ライブラリを意識したスタブ、モック、フェイクの作成または再利用、ビルド ログとライン カバレッジ フィードバックによって駆動される反復的なコンパイルとディスパッチの修復ループを組み合わせています。コンパイルの成功、修復の繰り返し、ディスパッチの成功、ライン カバレッジを使用し、時間、コスト、トークンの使用量を二次的な尺度として使用してアプローチを評価します。テスト対象の 76 の関数にわたって、ワークフローは 73 の関数のコンパイル可能な UT を生成しました。ライン カバレッジ ガイダンスや検索拡張を行わない構成では、平均ライン カバレッジは 73.9% に達しました。両方の構成で評価した 48 関数のサブセットでは、平均ライン カバレッジはライン カバレッジ ガイダンスのみで 98.8% に達し、ベクトル データベース検索と組み合わせた場合は 94.7% に達しました。結果は、自動生成および修復パイプラインにより、手動によるデバッグ作業を削減しながら、UT 作成効率と制約のあるファームウェア環境のカバレッジを大幅に向上できることが示されています。

原文 (English)

Library-Aware Doubles and Iterative Repair for Large Language Model-Generated Unit Tests in OpenSIL Firmware

Validating changes in low-level C firmware is expensive because unit tests (UTs) are fragile under strict build constraints, where missing headers, unresolved symbols, and dependency mismatches frequently prevent compilation and linking. This study introduces an automated UT authoring workflow for the Open-Source Silicon Initialization Library (openSIL) firmware codebase maintained by Advanced Micro Devices (AMD) that reduces manual effort through a large language model (LLM) guided multi-agent pipeline. The workflow combines automated generation of test scaffolds, library-aware creation or reuse of stubs, mocks, and fakes, and an iterative compile-dispatch repair loop driven by build logs and line-coverage feedback. We evaluate the approach using compilation success, repair iterations, dispatch success, and line coverage, with time, cost, and token usage as secondary measures. Across 76 functions under test, the workflow generated compilable UTs for 73 functions. In a configuration without line coverage guidance or retrieval augmentation, mean line coverage reached 73.9%. On a 48-function subset evaluated under both configurations, mean line coverage reached 98.8% with line-coverage guidance alone and reached 94.7% when combined with vector-database retrieval. Results show that automated generation-and-repair pipelines can substantially improve UT creation efficiency and coverage for constrained firmware environments while reducing manual debugging effort.

13:00 JSTLLM/生成AI

NRITYAM: 言語モデルとダンスの芸術と遺産の出会い

言語モデルは、現代のワークフローを形成する上で不可欠なツールとなっています。ただし、その世界的な有効性は、地域の社会文化的背景の微妙な理解にかかっています。このギャップに対処するために、世界的なダンスの伝統の文脈における言語モデルの文化的理解能力を評価するための包括的なベンチマークである NRITYAM を紹介します。 NRITYAM は、12 言語にわたる 9,260 の慎重に精選された質問と回答のペアで構成されており、ダンスの文化的知識の評価に特化した最大のデータセットとなっています。このデータセットは、ネイティブのダンス アーティストや言語のネイティブ スピーカーとの緊密な協力を通じて、ゼロから開発されました。彼らは、それぞれの地域に特有の文化的に関連する質問を作成および検証しました。大規模言語モデル、小規模言語モデル、マルチモーダル大規模言語モデル、小規模マルチモーダル言語モデルなど、幅広いモデルのセットを評価します。 NRITYAM は、多言語および多文化のベンチマークとして、伝統的な舞台芸術を理解し推論する AI システムの能力を評価するための新しい基準を設定します。詳細なデータセットのサンプルは、~\url{https://github.com/niladrighosh03/NRITYAM} で入手できます。

原文 (English)

NRITYAM: Language Models Meet Art and Heritage of Dance

Language models have become essential tools in shaping modern workflows. However, their global effectiveness hinges on a nuanced understanding of local socio-cultural contexts. To address this gap, we present NRITYAM, a comprehensive benchmark for evaluating the cultural comprehension capabilities of language models in the context of global dance traditions. NRITYAM comprises 9,260 carefully curated question-answer pairs spanning 12 languages, making it the largest dataset dedicated to evaluating cultural knowledge in dance. The dataset has been developed from the ground up through close collaboration with native dance artists and native speakers of the languages, who authored and validated culturally relevant questions specific to their regions. We evaluate a broad set of models, including large language models, small language models, multimodal large language models, and small multimodal language models. As a multilingual and multicultural benchmark, NRITYAM sets a new standard for evaluating the ability of AI systems to understand and reason about traditional performing arts. Detailed dataset samples are available at~\url{https://github.com/niladrighosh03/NRITYAM}.

13:00 JSTロボティクス

ロボットの発達運動学習のための双方向個別指導: 共同開発されたインタラクションダイナミクスが安定した学習をサポート

乳児は、養育者との密な交流を通じて運動能力を発達させることがよく知られています。このような社会的相互作用は人間の発達にとって重要ですが、ロボットの運動技能学習は、ロボットが家庭教師から受動的にデモンストレーションを受ける一方向のプロセスとして扱われることがよくあります。これは、社会的相互作用の重要な特性を見落としています。つまり、社会的相互作用は本質的に双方向であり、講師と学習者が動的にお互いに適応します。このような相互作用では、ロボットの過去の経験が、共同開発される軌道のダイナミクスを形作る事前の制約として機能する可能性があります。私たちは、双方向の個別指導では、そのような制約によって、行動の一貫性を維持し、一般化をサポートする一貫した行動パターンの形成を導くことができますが、一方、一方向の相互作用にはそのような制約が欠けており、より広範囲で一貫性の低い行動パターンが生じるのではないかと仮説を立てています。この仮説を検証するために、私たちは物体操作タスクを実行する物理的なヒューマノイド ロボットを使用して 2 つの実験を実施しました。1 つは人間とロボットの相互作用を伴うもので、もう 1 つは、より制御された条件下で同様の効果が現れるかどうかを調べるために設計された適応介入メカニズムを介して実際のロボットと対話する AI 家庭教師を使用するものです。私たちは、生成再生機能を拡張した自由エネルギー原理ベースのニューラル ネットワークを使用して発達学習フレームワークを実装します。これは、単一の個別指導エピソードからの安定したシーケンスごとの学習をサポートします。どちらの設定でも、双方向の個別指導により一貫した行動と段階ごとの一般化が促進され、ロボットが必要とする講師の指導は徐々に少なくなりました。これらの結果は、双方向の個別指導が、身体化された社会に根ざしたアプローチとして、ロボットの発達的な運動学習に効果的な足場を提供することを示唆しています。

原文 (English)

Bidirectional Tutoring for Developmental Motor Learning in Robots: Co-Developed Interaction Dynamics Support Stable Learning

Infants are well known to develop their motor skills through dense interaction with caregivers. Although such social interaction is crucial for human development, motor-skill learning in robots is often treated as a unidirectional process in which robots passively receive demonstrations from tutors. This overlooks a key property of social interaction: it is inherently bidirectional, with tutor and learner dynamically adapting to each other. In such interactions, the robot's past experiences may function as prior constraints that shape the dynamics of their co-developed trajectories. We hypothesize that bidirectional tutoring allows such constraints to guide the formation of consistent behavioral patterns that preserve behavioral coherence and support generalization, whereas unidirectional interaction lacks such constraints and leads to broader, less consistent behavioral patterns. To examine this hypothesis, we conducted two experiments with a physical humanoid robot performing an object manipulation task: one involving human-robot interaction and another employing an AI tutor interacting with the real robot through an adaptive intervention mechanism designed to examine whether similar effects would emerge under more controlled conditions. We implement the developmental learning framework using a free-energy-principle-based neural network extended with generative replay, which supports stable sequence-by-sequence learning from single tutored episodes. Across both settings, bidirectional tutoring fostered consistent behaviors and stage-wise generalization, while the robot gradually required less tutor guidance. These results suggest that bidirectional tutoring, as an embodied and socially grounded approach, provides an effective scaffold for developmental motor learning in robots.

13:00 JSTエージェントロボティクス

VOiLA: POMDP エージェントの学習された拡散モデルを使用したベクトル化されたオンライン プランニング

不確実性の下で計画を立てることは、自律ロボットにとって不可欠な機能です。 Partially Observable Markov Decision Process (POMDP) は、このような機能のための強力なフレームワークを提供します。 POMDP ベースの計画は大幅に進歩しましたが、忠実な POMDP モデルを取得することが難しいため、現実世界の問題への適用は制限されることがよくあります。不確実性の下でオンライン計画を立てるためにタスクに依存しない POMDP モデルを学習するフレームワークである、POMDP エージェント向け学習拡散モデルを使用したベクトル化オンライン計画 (VOiLA) を紹介します。 VOiLA は、条件付き拡散モデルを使用して遷移および観測サンプラーを学習し、粒子ベースの信念更新のための観測尤度モデルを学習します。効率的なオンライン プランニングを可能にするために、拡散サンプラーはコンパクトなフィードフォワード ジェネレーターに抽出され、GPU 並列化を活用するように設計されたオンライン POMDP プランナーである Vectorized Online POMDP Planner (VOPP) と統合されています。実験結果は、蒸留戦略によりサンプリング コストが最大 3 桁近く削減され、学習された生成 POMDP モデルがオンライン プランニングに実用的になることを示しています。 3 つのベンチマーク問題で VOiLA を評価したところ、VOiLA は 10% 未満のトレーニング データを使用しながら、Recurrent Soft Actor Critic と同等以上のパフォーマンスを達成し、目に見えない環境構成に対してはるかに優れた一般化を実現していることが示されています。物理的なロボットの評価では、VOiLA がシミュレートされたデータのみを使用して学習したモデルを使用し、10 回中 10 回の実行でタスクを正常に完了するポリシーを生成していることが示されています。

原文 (English)

VOiLA: Vectorized Online Planning with Learned Diffusion Model for POMDP Agents

Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provides a powerful framework for such a capability. Although POMDP-based planning has advanced significantly, its application to real-world problems is often limited by the difficulty of obtaining faithful POMDP models. We present Vectorized Online planning wIth Learned diffusion model for POMDP Agents (VOiLA), a framework that learns task-agnostic POMDP models for online planning under uncertainty. VOiLA learns transition and observation samplers using conditional diffusion models and learns observation-likelihood models for particle-based belief updates. To enable efficient online planning, the diffusion samplers are distilled into compact feedforward generators and integrated with Vectorized Online POMDP Planner (VOPP), an online POMDP planner designed to leverage GPU parallelization. Experimental results indicate the distillation strategy reduces sampling cost by up to nearly three orders of magnitude, making learned generative POMDP models practical for online planning. Evaluation of VOiLA on three benchmark problems indicate that VOiLA achieves equal or better performance than Recurrent Soft Actor Critic while using less than 10% training data, and generalizes much better to unseen environment configurations. Physical robot evaluation indicates VOiLA uses the models learned using only simulated data and generates a policy that successfully accomplish the task in 10 of 10 runs.

13:00 JSTLLM/生成AI画像/動画生成

QueryGaussian: スケーラブルでトレーニング不要のオープンボキャブラリー 3D インスタンス検索

自然言語プロンプトを介して大規模なシーンから特定の 3D インスタンスを効率的に取得することは、マルチメディア分析において依然として大きな課題です。既存のアプローチは主に「シーンレベルの埋め込み」パラダイムに従っており、高次元の意味論的特徴をすべての 3D プリミティブに抽出する必要があります。この戦略には基本的なアーキテクチャ上のボトルネックがあります。メモリと計算コストは​​シーンの複雑さに応じて直線的に増加し、都市規模の環境では必然的にメモリ不足 (OOM) 障害が発生します。この障壁に対処するために、私たちは、迅速かつスケーラブルなオープン語彙 3D インスタンス検索のためのトレーニング不要のフレームワークである QueryGaussian を提案します。全体的な意味の蒸留とは異なり、QueryGaussian は、幾何学的表現から意味の理解を切り離すインスタンス レベルのクエリ メカニズムを採用しています。具体的には、事前トレーニングされた 2D ビジョン モデルを活用してユーザー プロンプトを解釈し、同時最大重み関連付け戦略によってセグメンテーション マスクを 3D に引き上げ、セマンティックとビジュアルの一貫性を確保します。投影の曖昧さを軽減するために、多段階適応密度クラスタリングを備えた時間融合モジュールを導入します。実験結果は、QueryGaussian が最先端の手法の精度に匹敵するだけでなく、決定的な効率の飛躍をもたらし、GPU メモリ使用量を 70% 以上削減し、推論を 180 倍高速化することを示しています。重要なのは、QueryGaussian により、消費者グレードのハードウェアを使用して、数千万のガウスを含む都市規模のシーンでインスタンスを迅速に取得できるようになります。

原文 (English)

QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval

Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis. Existing approaches predominantly follow a "scene-level embedding" paradigm, which requires distilling high-dimensional semantic features into every 3D primitive. This strategy suffers from a fundamental architectural bottleneck: memory and computational costs scale linearly with scene complexity, inevitably triggering out-of-memory (OOM) failures in city-scale environments. To address this barrier, we propose QueryGaussian, a training-free framework for expeditious and scalable open-vocabulary 3D instance retrieval. Unlike holistic semantic distillation, QueryGaussian employs an instance-level query mechanism that decouples semantic understanding from geometric representation. Specifically, we leverage pre-trained 2D vision models to interpret user prompts and lift segmentation masks into 3D via a concurrent maximum-weight association strategy, ensuring semantic-visual consistency. To mitigate projection ambiguity, we introduce a temporal fusion module with multi-stage adaptive density clustering. Experimental results demonstrate that QueryGaussian not only matches the accuracy of state-of-the-art methods but also delivers a decisive efficiency leap, reducing GPU memory usage by over 70% and accelerating inference by 180x. Crucially, QueryGaussian enables expeditious instance retrieval on city-scale scenes containing tens of millions of Gaussians using consumer-grade hardware.

13:00 JSTLLM/生成AILlama

均一な忘却を超えて: プリファレンス設定全体にわたる逐次的な直接プリファレンス最適化の研究

言語モデルを人間の好みに合わせるには、多くの場合、複数の行動目標を最適化する必要があります。実際的なアプローチは、直接嗜好最適化 (DPO) などの嗜好最適化手法を使用してこれらの目標を順番に適用することですが、後のトレーニングによって以前に学習された嗜好が一律に低下するのか、それともその効果が目標間の関係に依存するのかは不明のままです。私たちは、分布の競合、複数属性の相互作用、強力な安全シグナル、および互換性のある応答品質目標をカバーする 4 つの優先設定にわたる連続した DPO を研究します。 LoRA アダプターで Llama-3.1-8B-Instruct を使用し、固定ベースモデル参照を使用して各ステージの後にすべての目標を評価します。逐次 DPO では単一の忘却パターンが生成されないことがわかりました。優先度の変化は、客観的な関係、信号強度、トレーニング順序に応じて、部分的な低下から安定性、ペアレベルの再分配、または正の転送まで多岐にわたります。長さで正規化されたポリシーマージンを使用したペアレベルの分析は、集計メトリクスがプリファレンスペア間の不均一な変化をマスクできることを示しますが、四分位分解は、設定に応じて信頼性の高いペアが低下または改善する可能性があることを明らかにします。機構診断では、ステージ 2 のグラジエントとアダプターの更新がすべての設定にわたって以前の目的とほぼ直交していることが示されており、直接的なグラジエントの反対が主な要因であるという証拠はほとんどありません。これらの発見は、将来の逐次アライメント パイプラインでは、後の目標が以前の優先順位に一律に影響を与えると仮定するのではなく、目標の互換性と信号強度を考慮する必要があることを示唆しています。

原文 (English)

Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings

Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply these objectives sequentially using preference optimisation methods such as Direct Preference Optimisation (DPO), but it remains unclear whether later training uniformly degrades preferences learned earlier or whether the effect depends on the relationship between objectives. We study sequential DPO across four preference settings covering distributional conflict, multi-attribute interaction, strong safety signal, and compatible response-quality objectives. Using Llama-3.1-8B-Instruct with LoRA adapters, we evaluate all objectives after every stage with a fixed base-model reference. We find that sequential DPO does not produce a single forgetting pattern; preference change ranges from partial degradation to stability, pair-level redistribution, or positive transfer depending on objective relationship, signal strength, and training order. Pair-level analysis using length-normalised policy margins shows that aggregate metrics can mask heterogeneous changes across preference pairs, whereas quartile decomposition reveals that high-confidence pairs can either degrade or improve depending on the setting. Mechanistic diagnostics show that Stage~2 gradients and adapter updates are near-orthogonal to the previous objective across all settings, providing little evidence that direct gradient opposition is the primary driver. These findings suggest that future sequential alignment pipelines should account for objective compatibility and signal strength, rather than assuming that later objectives affect earlier preferences uniformly.

13:00 JSTLLM/生成AI

多様な盗賊: 大規模言語モデルの潜在幾何学を学習するベイジアン カリキュラム

強化学習 (RL) は、大規模言語モデル (LLM) の推論機能を向上させるための中心的なアプローチであり、トレーニングの効率は、最適化中に問題がどのようにサンプリングされるかに大きく依存します。既存の適応カリキュラム学習方法は通常、中程度の難易度のプロンプトを優先し、問題の選択を独立したアームを備えた標準的なバンディット問題として扱い、タスク空間の構造化された異質な性質を見落としています。この研究では、問題のサンプリングを、内生的非定常性を伴う多様体構造のバンディット問題として組み立てます。問題はモデルの潜在表現空間を通じて関連付けられ、サンプリングの決定により、学習信号がその空間全体でどのように進化するかを制御できます。この観点を運用するために、問題を階層タスク ツリーに整理し、サンプリングのガイドにベイジアン学習を適用する構造認識フレームワークであるベイジアン多様体カリキュラム (BMC) を導入します。経験的に、さまざまなサンプリング戦略は、生産性 (学習信号)、多様性 (タスク多様体の範囲)、および有用性 (評価の関連性) の間に自明ではないトレードオフを引き起こすことがわかりました。これらの結果は、難易度を優先するだけでは下流のパフォーマンスを向上させるには不十分であることを示しており、問題のサンプリングに構造とタイプの認識を組み込むことの重要性を強調しています。

原文 (English)

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models

Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training efficiency depends critically on how problems are sampled during optimization. Existing adaptive curriculum learning methods typically prioritize prompts of intermediate difficulty, treating problem selection as a standard bandit problem with independent arms and overlooking the structured, heterogeneous nature of the task space. In this work, we frame problem sampling as a manifold-structured bandit problem with endogenous non-stationarity: problems are related through the model's latent representation space, and sampling decisions can steer how learning signals evolve across that space. To operationalize this perspective, we introduce Bayesian Manifold Curriculum (BMC), a structure-aware framework that organizes problems into a hierarchical task tree and applies Bayesian learning to guide sampling. Empirically, we find that different sampling strategies induce non-trivial tradeoffs between productivity (learning signal), diversity (coverage of the task manifold), and utility (evaluation relevance). These results show that prioritizing difficulty alone is insufficient for strong downstream performance, highlighting the importance of incorporating structure and type-awareness into problem sampling.

13:00 JSTロボティクス

時間的自己模倣学習

報酬形成でトレーニングされた長期的なロボット操作ポリシーは、非効率的なインタラクションを通じて高密度な報酬を活用することができますが、稀な効率的な行動はトレーニング中に忘れられる可能性があります。私たちは、時間効率自体が、強化学習のための自己監視の強力な、しかし十分に活用されていないソースを提供すると主張します。時間的自己模倣学習 (TSIL) を紹介します。これは、学習中に生成された時間的に効率的な成功軌道をマイニングし、将来のポリシー改善のために再利用可能な監視に変換する強化学習フレームワークです。 TSIL は、効率重視の自己模倣学習を通じて効率的な動作を保存および再生しながら、高速成功軌道から導出された構成条件付き適応時間目標を使用して学習を段階的に改良します。 TSIL は、15 の異なる長期的操作タスクにわたって、学習効率、タスク完了効率、迅速に成功した動作の再考、および不安定なトレーニング条件に対する堅牢性を一貫して向上させます。より広範に、私たちの結果は、成功した行動の時間構造自体が、手動で操作された報酬形成のみを超えた、強化学習のためのスケーラブルな自己監視信号を提供することを示唆しています。

原文 (English)

Temporal Self-Imitation Learning

Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while rare efficient behaviors may be forgotten during training. We argue that temporal efficiency itself provides a powerful and underutilized source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 15 distinct long-horizon manipulation tasks, TSIL consistently improves learning efficiency, task-completion efficiency, revisitation of fast successful behaviors, and robustness to unstable training conditions. More broadly, our results suggest that the temporal structure of successful behavior itself provides a scalable self-supervisory signal for reinforcement learning beyond manually engineered reward shaping alone.

13:00 JSTLLM/生成AIClaude

SafeSpec: 動的反射サンプリングによる高速かつ安全な LLM

投機的推論は大規模言語モデル (LLM) のデコードを高速化しますが、本質的な安全性の保証はありません。既存の安全防御は、投機的推論とほとんど互換性がありません。追加の計算が導入されるか、ドラフト検証メカニズムが中断され、高速化の利点が無効になります。これは、現在の安全方法と投機的デコードの間に根本的な非互換性があることを明らかにしています。私たちは、リスク推定を検証プロセスに直接統合する、安全性を意識した投機的推論フレームワークである SafeSpec を提案します。 SafeSpec は、軽量の潜在安全ヘッドをターゲット モデルに接続し、単一のフォワード パスでセマンティックの妥当性と安全性を共同で評価します。安全でない世代が検出されると、SafeSpec はロールバックと安全ガイド付きリフレクティブ マルチサンプリングを適用して、世代を終了するのではなく安全な継続を回復します。私たちはジェイルブレイク攻撃を、敵対的なプロンプトが安全なものを排除することなく、有害な継続の可能性を高める生成軌道上の分布シフトとしてモデル化します。このモデルでは、SafeSpec は投機的デコード プロセス内でリスクを認識した軌道回復を実行します。 SafeSpec は、複数のモデルと敵対的ベンチマークにわたって、安全性と効率のトレードオフを大幅に改善します。 Qwen3-32B では、SafeSpec は攻撃の成功率を 15% 低下させながら、無害なワークロードで 2.06 倍の推論速度向上を維持します。これは、投機的加速と推論時間の安全性を組み合わせて最適化できることを示しています。

原文 (English)

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are largely incompatible with speculative inference: they either introduce additional computation or disrupt the draft-verify mechanism, negating acceleration benefits. This reveals a fundamental incompatibility between current safety methods and speculative decoding. We propose SafeSpec, a safety-aware speculative inference framework that integrates risk estimation directly into the verification process. SafeSpec attaches a lightweight latent safety head to the target model to jointly evaluate semantic validity and safety in a single forward pass. When unsafe generations are detected, SafeSpec applies rollback and safety-guided reflective multi-sampling to recover safe continuations rather than terminating generation. We model jailbreak attacks as distributional shifts over generative trajectories, where adversarial prompts increase the probability of harmful continuations without eliminating safe ones. Under this model, SafeSpec performs risk-aware trajectory recovery within the speculative decoding process. Across multiple models and adversarial benchmarks, SafeSpec achieves a substantially improved safety-efficiency trade-off. On Qwen3-32B, SafeSpec reduces attack success rates by 15% while preserving a 2.06x inference speedup on benign workloads, demonstrating that speculative acceleration and inference-time safety can be jointly optimized.

13:00 JSTロボティクス

ヒューマノイド ロボティクスのデータ標準: 物理 AI に不足しているインフラストラクチャ

ヒューマノイド ロボットの拡張性は、モデルやハードウェアだけでなく、ロボット、タスク、組織、時間にわたって物理的な経験を蓄積できるかどうかにも依存します。この記事は、ISO/TC 299/WG 16 内の ISO/WD 26264-1、ヒューマノイド ロボット データセット -- パート 1: 一般要件の開発における著者の取り組みに基づいて、データ標準が物理 AI の基礎インフラになりつつあると主張します。私たちは 3 つの洞察を開発します。まず、ヒューマノイド ロボットのデータは具体化されたインタラクション データであり、孤立したデジタル サンプルのコレクションではありません。有用なデータセットは、ロボットの本体、アクション、タスク、シーン、実行トレース、および結果の間の関係を保存する必要があります。第 2 に、その値は物理的なコヒーレンスに依存します。マルチモーダル ストリームは、タイミング、座標フレーム、キャリブレーション、運動学、単位、および同期の仮定が検査可能な場合にのみ再利用可能です。 3 番目に、主なボトルネックはデータの不足だけでなく、高い収集コスト、データのサイロ化、一貫性のない評価によって引き起こされる非累積データです。私たちは、ヒューマノイドロボットのデータ標準が、具体化されたエクスペリエンスを解釈可能、共有可能、追跡可能、再利用可能にすることで、これらのボトルネックに対処すると主張します。一般的な標準は、ライフサイクル管理、メタデータ、来歴、品質、バージョン管理、およびトレーサビリティのための水平インフラストラクチャを提供する必要がありますが、機能固有の部分では、操作、移動、人間とロボットの相互作用、認知、および将来のヒューマノイド機能のためのドメイン文法を定義する必要があります。 AI が画面から身体に移行するにつれて、データ標準はデジタル情報の整理から物理的な相互作用の構造へと進化する必要があります。

原文 (English)

Data Standards for Humanoid Robotics: The Missing Infrastructure for Physical AI

The scalability of humanoid robots will depend not only on models and hardware, but also on whether physical experience can accumulate across robots, tasks, organizations, and time. Drawing on the authors' work in developing ISO/WD 26264-1, Humanoid robot datasets -- Part 1: General requirements, within ISO/TC 299/WG 16, this article argues that data standards are becoming foundational infrastructure for Physical AI. We develop three insights. First, humanoid robot data is embodied interaction data, not a collection of isolated digital samples; a useful dataset must preserve the relationship among robot body, action, task, scene, execution trace, and outcome. Second, its value depends on physical coherence: multimodal streams are reusable only when timing, coordinate frames, calibration, kinematics, units, and synchronization assumptions remain inspectable. Third, the main bottleneck is not only data scarcity, but non-cumulative data caused by high collection costs, data silos, and inconsistent evaluation. We argue that humanoid robot data standards address these bottlenecks by making embodied experience interpretable, shareable, traceable, and reusable. A general standard should provide horizontal infrastructure for lifecycle management, metadata, provenance, quality, versioning, and traceability, while capability-specific parts should define domain grammar for manipulation, locomotion, human-robot interaction, cognition, and future humanoid capabilities. As AI moves from screens into bodies, data standards must evolve from organizing digital information to structuring physical interaction.

13:00 JST研究/論文

事前トレーニングデータ構成によるエンジニアリングスケーリング則に向けて

ニューラル スケーリングの法則は、コンピューティング、モデル サイズ、データセット サイズにおけるべき乗則としてモデルのパフォーマンスがどのように向上するかを記述します。これらの関係は大規模な言語モデルでは十分に確立されていますが、素粒子物理学における大規模なモデルでも明らかになりつつあります。言語と同様に、パフォーマンスはべき乗則としてスケールされることが実証研究によって示されています。ただし、自然言語や画像の領域とは異なり、基礎物理学には合成データを安価に生成する忠実度の高いシミュレーターがあります。これにより、追加データの方が追加パラメーターよりも安価になるスケーリング方式が有利になり、スケーリングに影響を与えるように事前トレーニング データセット自体を操作できるようになります。高エネルギー粒子ビームの衝突で生成されるハドロンジェットを分類するタスクでは、より多様で下流の分類タスクとよりよく連携する事前学習データを含めることにより、大規模なモデルではなくより多くのデータを必要とするようにスケーリング動作を設計できることを示します。

原文 (English)

Towards Engineering Scaling Laws with Pretraining Data Composition

Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established for large language models, these relationships are emerging for large models in particle physics. As with language, empirical studies show that the performance scales as a power law. However, unlike natural language or image domains, fundamental physics has high-fidelity simulators that produce synthetic data cheaply. This favors scaling regimes where additional data is cheaper than additional parameters, and allows the pretraining dataset itself to be engineered to influence the scaling. For the task of classifying hadronic jets produced in collisions of high-energy particle beams, we show that the scaling behavior can be engineered towards requiring more data rather than larger models by inclusion of pretraining data which is more diverse and better aligned with the downstream classification task.

13:00 JST研究/論文

データセット、年齢、性別を超えた一般化: リソースの少ない子供の ASR のための微調整戦略の包括的な分析

構音障害のある音声を認識することに関連する課題は、主に、調音精度の低下に起因する顕著な音響変動から生じます。過去の研究では、ハイブリッド DNN/HMM シーケンスの識別トレーニングの使用によって認識が向上することが実証されています。このペーパーでは、さまざまな音響モデルに合わせた音響特徴のさまざまな組み合わせの包括的な調査を示し、それぞれに適した特徴の選択を提供します。ピッチ機能を組み込むことで、特に構音障害のある音声を伴う文章認識タスクの認識パフォーマンスが著しく向上しました。 TORGO データベースの体系的な検査を通じて、構音障害音声を認識するための最先端の因数分解時間遅延ニューラル ネットワーク (F-TDNN) モデルのパフォーマンスを強化できる可能性を実証しました。 F-TDNN モデルを使用して実装された私たちの方法は、以前の研究と比較して、孤立単語認識で 4.65% 相対改善、構音障害音声の文認識で 4.63% 相対改善をもたらしました。この改善により、連続するトレーニング サンプル チャンク間で重複するフレームの数を意図的に選択したことに起因する音声の変動が効果的に補償されます。

原文 (English)

Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.

13:00 JST研究/論文

構音障害による音声認識の系統的研究:スペクトル特徴と音響モデル

構音障害のある音声を認識することに関連する課題は、主に、調音精度の低下に起因する顕著な音響変動から生じます。過去の研究では、ハイブリッド DNN/HMM シーケンスの識別トレーニングの使用によって認識が向上することが実証されています。このペーパーでは、さまざまな音響モデルに合わせた音響特徴のさまざまな組み合わせの包括的な調査を示し、それぞれに適した特徴の選択を提供します。ピッチ機能を組み込むことで、特に構音障害のある音声を伴う文章認識タスクの認識パフォーマンスが著しく向上しました。 TORGO データベースの体系的な検査を通じて、構音障害音声を認識するための最先端の因数分解時間遅延ニューラル ネットワーク (F-TDNN) モデルのパフォーマンスを強化できる可能性を実証しました。 F-TDNN モデルを使用して実装された私たちの方法は、以前の研究と比較して、孤立単語認識で 4.65% 相対改善、構音障害音声の文認識で 4.63% 相対改善をもたらしました。この改善により、連続するトレーニング サンプル チャンク間で重複するフレームの数を意図的に選択したことに起因する音声の変動が効果的に補償されます。

原文 (English)

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.

13:00 JSTエージェント

エージェントティック電子設計自動化: ハンドオフの観点

電子設計自動化 (EDA) は本質的に多段階であり、ハンドオフの負荷が高くなります。最終的な実装、サインオフ、またはリリースの前に、ツール、セッション、組織の境界を越えて成果物、フロー スクリプト、エンジニアリング上の決定を設計します。各転送には、ステージローカルのチェックでは完全には捕捉できない可能性のある明示的および暗黙的な要件が伴います。 LLM ベースのエージェントは、EDA ツールを直接呼び出し、取得した知識を実行可能スクリプトに埋め込み、セッションやステージ間で状態を引き継ぎます。出力が下流のエンジニアリング上の決定を条件付けたら、転送されたオブジェクトは引き継ぎ契約を満たし、次の消費者の想定を満たさなければなりません。この調査では、組織原則としてハンドオフの有効性が導入されています。ハンドオフは、転送されたオブジェクトが消費者の受け入れ条件を満たし、下流での使用に十分なコンテキスト、証拠、出所を保持している場合に有効です。私たちは 82 のシステムをレビューし、それらを 3 つの境界クラスに分類します。ステージ限定システムは、単一の EDA ステージまたは限定された検証タスク内で有効性を確立します。フロー バインド システムは、ツール、呼び出し、セッション全体で一貫したワークフローの状態を保持します。組織に束縛されたシステムは、知識や権限の境界を越えて、情報源の根拠、出所、範囲、許容性を維持します。各クラスについて、引き継ぎ契約、引き継ぎオブジェクト、調整メカニズム、および未解決の質問を分析します。これらの分析は、エージェントの検出、エージェントのメッセージ、ツールの呼び出し、ワークフローのオーケストレーション、セキュリティおよび IP プロトコルをカバーする 5 層の EDA エージェント通信プロトコル (EACP) の動機付けとなります。私たちは、信頼できるエージェント EDA のための共通の語彙と研究課題を提供することを目指しています。

原文 (English)

Agentic Electronic Design Automation: A Handoff Perspective

Electronic design automation (EDA) is inherently multi-stage and handoff-heavy. Design artifacts, flow scripts, and engineering decisions cross tool, session, and organizational boundaries before final implementation, signoff, or release. Each transfer carries explicit and implicit requirements that may not be fully captured by stage-local checks. LLM-based agents now invoke EDA tools directly, embed retrieved knowledge in executable scripts, and hand off state across sessions and stages. Once their outputs condition downstream engineering decisions, the transferred object must satisfy a handoff contract and meet the assumptions of its next consumer. This survey introduces handoff validity as its organizing principle. A handoff is valid when the transferred object satisfies the consumer's acceptance conditions and carries sufficient context, evidence, and provenance for downstream use. We review 82 systems and classify them into three boundary classes. Stage-Bound systems establish validity within a single EDA stage or bounded verification task. Flow-Bound systems preserve coherent workflow state across tools, invocations, and sessions. Organization-Bound systems maintain source grounding, provenance, scope, and admissibility across knowledge and authority boundaries. For each class, we analyze handoff contracts, handoff objects, coordination mechanisms, and open questions. These analyses motivate a five-layer EDA agent communication protocol (EACP), covering the agent discovery, agent message, tool invocation, workflow orchestration, and security and IP protocols. We aim to provide a common vocabulary and research agenda for trustworthy agentic EDA.

13:00 JST研究/論文

ドメイン内データ拡張による構音障害音声に対するエンドツーエンドの音声認識の改善

構音障害のある人々の間で効果的なコミュニケーションを促進するには、構音障害の音声認識が不可欠です。ただし、構音障害のある音声を正確に認識することは、さまざまな重症度レベルと限られたデータにより、大きな課題を引き起こします。この論文では、特に重症度レベルに焦点を当て、エンドツーエンドの事前トレーニング済み Wav2Vec2 モデルを微調整することにより、構音障害自動音声認識 (ASR) システムのデータ拡張手法を検討します。データ不足という課題と、構音障害音声用の​​事前トレーニング済み ASR システムを微調整する際の大量のデータの必要性に対処するために、構音障害のさまざまな側面に合わせて調整された 4 つの著名なデータ拡張方法、つまりスピーキング レート修正 (SRM)、ピッチ修正 (PM)、フォルマント修正 (FM)、および声道長摂動 (VTLP) を調査します。この研究では、ベースライン システムとして各重大度クラスに対して個別に微調整された Wav2Vec2 モデルを使用します。さらに、拡張データを使用して ASR モデルの重大度固有の微調整を実施しました。結果は、重症度レベルにわたる各増強技術の異なる有効性パターンを示しています。最高の WER は、\textit{low} (9.02\%) および \textit{medium} (38.11\%) 重大度では SRM ($s$=0.8) で達成され、\textit{high} 重大度 (55.15\%) では PM ($\tau$=0.8) で達成され、30.02\%、16.64\%、およびそれぞれ15.47\%。これらの結果は、構音障害による ASR パフォーマンスの改善における増強法の有効性を裏付けています。

原文 (English)

Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria. However, accurately recognizing dysarthric speech poses significant challenges due to varying severity levels and limited data availability. In this paper, we explore data augmentation techniques for dysarthric automatic speech recognition (ASR) systems by fine-tuning the End-to-End pre-trained Wav2Vec2 model, with a specific focus on severity levels. To address the challenges of data scarcity and the need for extensive data in fine-tuning pre-trained ASR systems for dysarthric speech, we investigate four prominent data augmentation methods: Speaking-Rate Modification (SRM), Pitch Modification (PM), Formant Modification (FM), and vocal tract Length Perturbation (VTLP), tailored to different aspects of dysarthria. The study uses individually fine-tuned Wav2Vec2 models for each severity class as baseline systems. Additionally, we conducted severity-specific fine-tuning of the ASR model using augmented data. Results demonstrate distinct efficacy patterns for each augmentation technique across severity levels. The best WERs were achieved with SRM ($s$=0.8) for \textit{low} (9.02\%) and \textit{medium} (38.11\%) severities, and with PM ($\tau$=0.8) for \textit{high} severity (55.15\%), reflecting relative improvements of 30.02\%, 16.64\%, and 15.47\%, respectively. These results confirm the effectiveness of the augmentation methods in improving dysarthric ASR performance.

13:00 JST研究/論文

ポリシーを意識したベクトル検索: ベクトル データベースにおけるきめ細かいアクセス制御のビジョン

ベクトル データベースは、検索拡張生成および組織 AI パイプラインを使用して、セキュリティに敏感な状況でますます使用されています。ただし、セキュリティ機能は依然として制限されています。特に、データ アクセスがユーザー固有のポリシーに準拠していることを確認するために必要なきめ細かなアクセス制御 (FGAC) は、最新のベクター データベースでは完全にはサポートされていません。リレーショナル データベースとは異なり、ベクトル データベースは構造化属性と非構造化属性を組み合わせて意味論的な近似クエリ結果を提供するため、FGAC の実装が複雑になります。これにより、FGAC ポリシーを正しく適用すること、高い ANN 検索再現率を達成すること、および低いクエリ遅延を維持することの間に固有の緊張が生じます。このペーパーでは、ベクトル データベース内の FGAC ポリシー モデルと施行問題を形式化することにより、ポリシーを認識したベクトル検索のビジョンを示します。私たちはさまざまな執行戦略を比較し、予備的な調査結果を提示し、ポリシーを意識したベクトル検索における将来の研究に向けた重要な未解決の課題を特定します。

原文 (English)

Policy-aware Vector Search: A Vision for Fine Grained Access Control in Vector Databases

Vector databases are increasingly used in security sensitive contexts with Retrieval Augmented Generation and organizational AI pipelines; however, their security capabilities remain limited. Specifically, Fine-grained Access Control (FGAC) which is required to ensure that data access adheres to user-specific policies is not fully supported in modern vector databases. Unlike relational databases, vector databases combine structured and unstructured attributes to provide semantic, approximate query results, which complicates FGAC implementation. This creates an inherent tension between enforcing FGAC policies correctly, achieving high ANN search recall and maintaining low query latency. In this paper, we present a vision for Policy-aware Vector Search by formalizing the FGAC policy model in vector databases as well as the enforcement problem. We compare various enforcement strategies, present preliminary findings, and identify key open challenges for future research in policy-aware vector search.

13:00 JST画像/動画生成

ParaScale: ゲージ不変の視差数を介したスケール調整されたカメラモーション転送

リファレンス ビデオのカメラ モーションを新しく生成されたビデオに転送すると、クリエイターは映画のような動きを再利用できます。しかし、リファレンスとターゲットは、多くの場合、互換性のないスケール、つまり銀河を横切るスイープと机の上をナッジするようなスケールで存在しており、復元された軌道を単純に再利用すると、知覚できないか、または激しく誇張されたモーションが生成されます。これを幾何学的な事実にたどります。並進によって引き起こされる画像の動きは ||T||/Z としてスケールされるため、単眼軌道は深度スケール ゲージまででのみ意味を持ちます。これを視差数 Pi = ||デルタ T|| に抽出します。 / Zbar は、カメラの動きがどの程度強く感じられるかを示す無次元でゲージ不変の記述子であり、生の軌道ではなく、スケールに忠実な転送で保持しなければならない量であることを証明します。 ParaScale は、参照ビデオから Pi を読み取り、フレームごとに回転をそのままにして、ターゲット シーン自体の深度に対してそれを再実現するプラグ アンド プレイ モジュールです。ポーズ抽出とポーズ挿入の間に位置するため、再トレーニングは必要なく、ポーズ条件付きジェネレーターにドロップされます。さらに、視差整合性エラー (PCE) を導入します。これは、類似性を調整した TransErr とは異なり、シーンのスケールの不一致を明らかにするスケール対称のメトリックです。 ParaScale は、4 桁にわたるスケール領域と複数のバックボーンにわたって、実現された視差を同一線上に維持し、視覚的な忠実度を損なうことなく、未校正の転送に比べて PCE を 3 倍以上削減します。

原文 (English)

ParaScale: Scale-Calibrated Camera-Motion Transfer via a Gauge-Invariant Parallax Number

Transferring the camera motion of a reference video to a freshly generated one lets creators reuse cinematic moves. Yet reference and target often live at incompatible scales -- a sweep across a galaxy versus a nudge across a desk -- and naively reusing the recovered trajectory yields either imperceptible or violently exaggerated motion. We trace this to a geometric fact: translation-induced image motion scales as ||T||/Z, so a monocular trajectory is meaningful only up to a depth-scale gauge. We distill this into the Parallax Number Pi = ||Delta T|| / Zbar, a dimensionless, gauge-invariant descriptor of how strongly a camera move is felt, and prove that it -- not the raw trajectory -- is the quantity that scale-faithful transfer must preserve. ParaScale is a plug-and-play module that reads Pi off any reference video and re-realizes it against the target scene's own depth, per frame, leaving rotation untouched. Sitting between pose extraction and pose injection, it requires no retraining and drops into any pose-conditioned generator. We further introduce the Parallax Consistency Error (PCE), a scale-symmetric metric that -- unlike the similarity-aligned TransErr -- exposes scene-scale mismatch. Across scale regimes spanning four orders of magnitude and multiple backbones, ParaScale keeps the realized parallax on the identity line and cuts PCE by more than 3x over uncalibrated transfer with no loss of visual fidelity.

13:00 JST研究/論文

安定した RLHF のための不確実性を考慮した報酬モデリング

ヒューマン フィードバックからの強化学習 (RLHF) は、嗜好データに基づいて報酬モデルをトレーニングし、予測される報酬を最大化するポリシーを最適化することで、大規模な言語モデルを調整します。ただし、このパイプラインは 2 つの基本的な課題に直面しています。(1) 報酬モデルは通常、決定論的な点推定器として機能するため、予測が信頼できない場合に信号を送信できません。 (2) 現代のグループベースのポリシー最適化は、GRPO によるアドバンテージ計算中の報酬の均一な扱いに例示されるように、信頼性の低い報酬シグナルを増幅する可能性があります。政策がますます多様化する対応を模索する中で、これら 2 つの制限により重大な脆弱性が生じます。信頼性の低い報酬推定値が不釣り合いな影響力を与えられ、深刻な報酬ハッキングを引き起こす可能性があります。我々は不確実性を考慮した報酬モデリング(UARM)を提案します。これは、分位数ベースの等角予測を通じて報酬モデルに校正された不確実性を与え、不均一分散分解を通じて GRPO の利点を再重み付けします。 HelpSteer、UltraFeedback、および PKU-SafeRLHF にわたる実験では、UARM が標準の GRPO および不確実性を問わないベースラインと比較して、報酬モデルのキャリブレーションを大幅に改善し、報酬ハッキングを削減し、下流のアライメント品質を向上させることが実証されています。

原文 (English)

Uncertainty-Aware Reward Modeling for Stable RLHF

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. However, this pipeline faces two fundamental challenges: (1) reward models cannot signal when their predictions are unreliable, since they usually act as deterministic point estimators; and (2) modern group-based policy optimization can amplify unreliable reward signals, as exemplified by GRPO's uniform treatment of rewards during advantage computation. As policies explore increasingly diverse responses, these two limitations create a critical vulnerability: unreliable reward estimates may be granted disproportionate influence, triggering severe reward hacking. We propose Uncertainty-Aware Reward Modeling (UARM), which equips reward models with calibrated uncertainty via quantile-based conformal prediction and reweights GRPO advantages through heteroscedastic variance decomposition. Experiments across HelpSteer, UltraFeedback, and PKU-SafeRLHF demonstrate that UARM significantly improves reward model calibration, reduces reward hacking, and enhances downstream alignment quality compared to standard GRPO and uncertainty-agnostic baselines.

13:00 JSTLLM/生成AI

CREDENCE: 分解と信頼性向上のためのクレーム削減 -- セマンティック メトリクスと収束分析

複合文を原子的な検証可能なクレームに分解することは、信頼性の高い自動ファクトチェックの前提条件です。これまでの研究は、言い換えクレームの分解品質を体系的に過小評価するトークンオーバーラップ (Jaccard) メトリクスに依存しており、修復ループの正式な終了分析が不足していました。我々は、両方の欠点に対処する改訂されたクレーム分解および評価フレームワークである Credence を紹介します。私たちの貢献は次のとおりです。 (1) Semantic-F1: Jaccard のペナルティを解決し、ダウンストリームのファクトチェック精度を向上させる BGE ラージ コサイン類似度忠実度メトリックを使用します。 (2) 収束定理: 修復パイプラインの 4 つの特性を正式に特徴付け、ルールベースの修復が単調であり、Oracle パーサーの仮定の下で有限終了することを確立します。 LLM ベースの自己修復は単調ではないことが証明されており、早期終了ガードが必要です。 (3) クロスドメインの一般化測定のための、ソーシャルメディア、百科事典、ニュースの各ドメインにわたる 3 つの評価ベンチマーク。 (4) 4 つのデコンポーザー モデル (3.8B ~ 12B) とクローズド API モデルにわたるマルチモデル ベンチマーク。 SocialClaimSplit、WikiSplitBench、ClaimDecompBench の実験では、Semantic-F1 が Jaccard-F1 を +15 ~ 32pp 上回るパフォーマンスを示しています。 SocialClaimSplit と WikiSplitBench では、EPR の範囲は 0.94 ~ 1.00 ですが、ClaimDecompBench には、ニュース ドメインの構築が難しいため、より低い基本 EPR ケース (0.824 まで) が含まれており、ルール修復により、忠実度を低下させることなく、基本モデルと比較して原子性違反率 (AVR) が 47 ~ 100% 減少します。

原文 (English)

CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility -- Semantic Metrics and Convergence Analysis

Decomposing compound sentences into atomic, verifiable claims is a prerequisite for reliable automated fact-checking. Prior work has relied on token-overlap (Jaccard) metrics that systematically underestimate decomposition quality for paraphrastic claims, and has lacked formal termination analysis for the repair loop. We present Credence, a revised claim decomposition and evaluation framework addressing both shortcomings. Our contributions are: (1) Semantic-F1: we use BGE-large cosine similarity fidelity metric that resolves Jaccard's penalisation and improves downstream fact-checking accuracy; (2) Convergence theorems: we formally characterise four properties of the repair pipeline, establishing that rule-based repair is monotone and finitely terminating under an oracle parser assumption; LLM-based self-repair is provably non-monotone and requires an early-exit guard; (3) Three evaluation benchmarks spanning social-media, encyclopaedic, and news domains for cross-domain generalisation measurement; (4) Multi-model benchmarking across four decomposer models (3.8B-12B) and a closed API model. Experiments on SocialClaimSplit, WikiSplitBench, and ClaimDecompBench show that Semantic-F1 outperforms Jaccard-F1 by +15-32pp. EPR ranges from 0.94 to 1.00 on SocialClaimSplit and WikiSplitBench, while ClaimDecompBench includes lower base EPR cases (down to 0.824) due to harder news-domain constructions, and rule-repair reduces the Atomicity Violation Rate (AVR) by 47-100% relative to the base model without degrading fidelity.

13:00 JST画像/動画生成

CSWinUNETR: 医療画像における薄い解剖学的構造のセグメンテーション

網膜血管、脳血管系、顔のしわなど、薄く曲がりくねった解剖学的構造を正確にセグメンテーションすることは、低コントラスト、頻繁な不連続、深刻なクラスの不均衡のため、依然として困難です。最近の畳み込みモデルや Transformer ベースのモデルはパフォーマンスが向上していますが、多くの場合、断片的な予測が生成され、細かい分岐を回復できません。私たちは、2D および 3D の薄い構造のセグメンテーションのための汎用バックボーンである CSWinUNETR を提案します。これは、長距離の主軸コンテキストをモデル化するために十字型のストライプの自己注意を採用し、ストライプ間の情報交換を強化するために循環シフトを組み込んでいます。きめの細かい詳細をより適切に保存するために、マルチ解像度表現からコンテキスト上の特徴を集約する、詳細が強化されたマルチスケール セルフ アテンション モジュールをさらに導入します。さらに、我々は、まばらに予測された制御点から信頼性の高い密な曲線カーネルを再構築し、曲がりくねったジオメトリをよりよく追従する、まばらな制御の動的スネーク畳み込みを提案します。眼科、神経血管画像診断、皮膚科にわたる 4 つのベンチマークに関する広範な実験により、CSWinUNETR がタスク固有の後処理やトポロジーを考慮した損失を生じることなく、常に最先端の手法を上回るパフォーマンスを発揮することが実証されました。コードは https://github.com/labhai/CSWinUNETR で入手できます。

原文 (English)

CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images

Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains challenging due to low contrast, frequent discontinuities, and severe class imbalance. Although recent convolutional and Transformer-based models have improved performance, they often yield fragmented predictions and fail to recover fine branches. We propose CSWinUNETR, a general-purpose backbone for 2D and 3D thin-structure segmentation. It employs cross-shaped stripe self-attention to model long-range principal-axis context and incorporates cyclic shifts to enhance information exchange across stripes. To better preserve fine-grained details, we further introduce a detail-enhanced multi-scale self-attention module that aggregates contextual features from multi-resolution representations. In addition, we propose sparse-control dynamic snake convolution, which reconstructs reliable dense curvilinear kernels from sparsely predicted control points to better follow tortuous geometry. Extensive experiments on four benchmarks across ophthalmology, neurovascular imaging, and dermatology demonstrate that CSWinUNETR consistently outperforms state-of-the-art methods without task-specific post-processing or topology-aware losses. The code is available at https://github.com/labhai/CSWinUNETR.

13:00 JST研究/論文

いつ、どこで、どのように: 表形式の自己教師あり学習のための適応型ビニング

医療の表形式データは臨床研究において広く普及していますが、構造化された臨床変数が表形式で日常的に利用可能であるにもかかわらず、信頼できるラベルを得るには高価な専門家の判断が必要となることが多いため、表の深層学習はまだ研究されていません。自己教師あり学習ではこれらのラベルなしテーブルを活用でき、最近のビニングベースの口実は有望な帰納的バイアスを提供しますが、既存の目標は単一のグローバル分位離散化を修正し、特徴に依存しない監視を適用します。私たちは、表形式 SSL のトレーニングに適応した離散化の口実であるアダプティブ ビニングを提案します。これは、特徴ごとに粗いものから細かいものまでのカリキュラムを通じて、離散化と学習を結び付けます。ニューラル ネットワークのスペクトル バイアスとカリキュラム学習の原理に動機付けられた私たちの方法は、プラトー検出時に特徴ごとの離散化を段階的に改良し、表現を意識した分割を選択して、値空間の集中と表現空間の一貫性を共同で改善します。異質性を認識した目標により、カテゴリカルな再構成と数値特徴の順序監視が統合され、統一された評価プロトコルに基づく公的医療表データセットでの実験では、データセット固有の離散化調整を行わずに、線形プローブと微調整で一貫した利益が得られることが示されています。さらに、この未開発の領域における再現可能な進歩をサポートするために、標準化されたプロトコルを備えた医療用​​表形式の SSL ベンチマークを導入します。私たちのコードは https://github.com/labhai/Adaptive-Binning で入手できます。

原文 (English)

When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning

Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because reliable labels often require costly expert adjudication, even though structured clinical variables are routinely available in tabular form. Self-supervised learning can leverage these unlabeled tables, and recent binning-based pretexts offer a promising inductive bias, but existing objectives fix a single global quantile discretization and apply feature-agnostic supervision. We propose Adaptive Binning, a training-adaptive discretization pretext for tabular SSL that couples discretization to learning through a feature-wise coarse-to-fine curriculum. Motivated by the spectral bias of neural networks and the principles of curriculum learning, our method progressively refines discretization per feature upon plateau detection and selects representation-aware splits to jointly improve value-space concentration and representation-space coherence. A heterogeneity-aware objective unifies categorical reconstruction with ordinal supervision for numerical features, and experiments on public medical tabular datasets under unified evaluation protocols show consistent gains for linear probing and fine-tuning without dataset-specific discretization tuning. We further introduce a medical tabular SSL benchmark with standardized protocols to support reproducible progress in this underexplored domain. Our code is available at https://github.com/labhai/Adaptive-Binning.

13:00 JST研究/論文

特徴選択と相互作用を備えたニューラル加算モデルと基底モデル

ディープ ニューラル ネットワーク (DNN) は、さまざまな分野で魅力的なパフォーマンスを発揮しますが、解釈可能性が低いことがよくあります。ニューラル加法モデル (NAM) とニューラル基礎モデル (NBM) と呼ばれるそのバリアントは、一般化加法モデル (GAM) の非線形形状関数としてニューラル ネットワーク (NN) を使用します。どちらのモデルも解釈可能性が高く、NN トレーニングに優れたパフォーマンスと柔軟性を示します。 NAM と NBM は、GAM ベースのアーキテクチャにより、予測に対する各機能の寄与を提供および視覚化できます。ただし、2 入力 NN を使用して特徴の相互作用を検討したり、高次元データセットに適用したりする場合、必要な計算リソースが増加するため、NAM と NBM のトレーニングが困難になります。このペーパーでは、計算上のボトルネックを解決するために、NAM および NBM に機能選択メカニズムを組み込むことを提案します。両方のモデルに特徴選択レイヤーを導入し、トレーニング中に選択の重みを更新します。私たちの方法はシンプルで、標準的な NAM や NBM と比較して計算コストとモデル サイズを削減できます。さらに、高次元データセットでも 2 入力 NN を使用して、特徴の相互作用をキャプチャできるようになります。提案されたモデルは標準的な NAM および NBM と比較して計算効率が高く、最先端の GAM と同等またはそれ以上のパフォーマンスを示すことを実証します。

原文 (English)

Neural Additive and Basis Models with Feature Selection and Interactions

Deep neural networks (DNNs) exhibit attractive performance in various fields but often suffer from low interpretability. The neural additive model (NAM) and its variant called the neural basis model (NBM) use neural networks (NNs) as nonlinear shape functions in generalized additive models (GAMs). Both models are highly interpretable and exhibit good performance and flexibility for NN training. NAM and NBM can provide and visualize the contribution of each feature to the prediction owing to GAM-based architectures. However, when using two-input NNs to consider feature interactions or when applying them to high-dimensional datasets, training NAM and NBM becomes intractable due to the increase in the computational resources required. This paper proposes incorporating the feature selection mechanism into NAM and NBM to resolve computational bottlenecks. We introduce the feature selection layer in both models and update the selection weights during training. Our method is simple and can reduce computational costs and model sizes compared to vanilla NAM and NBM. In addition, it enables us to use two-input NNs even in high-dimensional datasets and capture feature interactions. We demonstrate that the proposed models are computationally efficient compared to vanilla NAM and NBM, and they exhibit better or comparable performance with state-of-the-art GAMs.

13:00 JSTLLM/生成AI

大規模な言語モデルには必ずしも可読言語が必要なわけではない

大規模言語モデル (LLM) は、意図した読者が別のモデルである場合でも、一般に人間が読める自然言語を使用してプロンプトを表示し、インターフェイスします。この論文では、LLM による回復可能性を維持しながら、人間の可読性を犠牲にするコンパクトな非標準テキスト形式で意味情報をエンコードできるかどうかを調査します。モデル中心のテキスト表現のこのクラスを BabelTele と呼びますが、ここでは固定プロトコルとしてではなく、そのような表現を生成および解釈する LLM の能力の経験的調査としてアプローチします。可読性診断、モデル尤度測定、人によるアンケート、および下流タスク評価を通じて、BabelTele が命令調整 LLM の中核となるセマンティクスを維持しながら、通常の自然言語から大幅に逸脱できることがわかりました。タスクに依存しない表現パラダイムとして、BabelTele は高い情報密度を示し、テキスト量が元の長さの 27.9% に凝縮された場合でも 99.5% の意味的忠実度を維持します。さらに、クロスモデル転送、エージェントのメモリ、およびマルチエージェント通信におけるセマンティックの堅牢性を評価します。結果は、BabelTele が一般に信頼性の高いダウンストリーム パフォーマンスを維持しながらコンテキスト オーバーヘッドを削減できることを示唆していますが、その有効性はコンプレッサーとリーダーのペアとタスク設定によって異なります。これらの発見は、人間の可読性、自然言語の典型性、モデル側の意味回復可能性が部分的に分離できることを示しており、LLM システムの将来の探求においてモデルネイティブ表現への道が開かれます。

原文 (English)

Large Language Models Do Not Always Need Readable Language

Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investigates whether semantic information can be encoded in compact, non-standard textual forms that sacrifice human readability while remaining recoverable by LLMs. We refer to this class of model-centric textual representations as BabelTele, approached here not as a fixed protocol but as an empirical probe into LLMs' capacity to generate and interpret such representations. Through readability diagnostics, model likelihood measures, human questionnaires, and downstream task evaluations, we find that BabelTele can substantially depart from ordinary natural language while preserving core semantics for instruction-tuned LLMs. As a task-agnostic representational paradigm, BabelTele demonstrates high information density, maintaining 99.5% semantic fidelity even when the text volume is condensed to 27.9% of its original length. We further evaluate its semantic robustness in cross-model transfer, agent memory, and multi-agent communication. Results suggest that BabelTele can reduce context overhead while generally maintaining reliable downstream performance, although its effectiveness depends on the compressor-reader pair and task setting. These findings indicate that human readability, natural-language typicality, and model-side semantic recoverability can be partially decoupled, opening a path toward model-native representations in future exploration of LLM systems.

13:00 JST画像/動画生成

PSCT-Net: 微分可能な逆投影と注意に基づく改良による、形状を考慮した小児頭蓋骨 CT 再構成

コンピュータ断層撮影 (CT) は小児の頭蓋顔面異常の診断に不可欠ですが、発達中の解剖学的構造に放射線リスクをもたらします。まばらな二平面 X 線から 3D CT を再構成することは、低線量の代替手段となりますが、非常に不適切です。既存の方法は、ジオメトリに依存しない特徴リフティングを採用しており、明示的な空間モデリングを行わずに単純に 2D 特徴を 3D に投影するため、深さの曖昧さと骨境界の劣化が生じます。微分可能な逆投影を備えた幾何学認識フレームワークである PSCT-Net を紹介します。微分可能な逆投影により、空間的に忠実な体積事前分布が確立され、深さの曖昧さが軽減されます。次に、注意誘導投影 (AGP-3D) モジュールが、2D 領域と 3D 位置の間の非線形ボクセル単位の対応を学習します。 Bi方向 Mamba (BiM-3D) モジュールは、線形の複雑さで長距離の体積依存関係をキャプチャします。さらに、内部評価用に正常症例と病理学的症例で構成される民間の施設小児頭蓋骨CTコホートであるPedSkull-CTをキュレーションし、成人中心の体幹に焦点を当てたデータセットのギャップに対処します。

原文 (English)

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement

Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomies. Reconstructing 3D CT from sparse bi-planar X-rays offers a low-dose alternative but is severely ill-posed. Existing methods employ geometry-agnostic feature lifting, naively projecting 2D features into 3D without explicit spatial modeling, causing depth ambiguity and degraded osseous boundaries. We present PSCT-Net, a geometry-aware framework with differentiable back-projection. Differentiable back-projection establishes a spatially faithful volumetric prior, alleviating depth ambiguity. An Attention-Guided Projection (AGP-3D) module then learns non-linear voxel-wise correspondences between 2D regions and 3D locations. A Bidirectional Mamba (BiM-3D) module captures long-range volumetric dependencies with linear complexity. We further curate a private institutional pediatric skull CT cohort, PedSkull-CT, comprising normal and pathological cases for internal evaluation, addressing the gap in adult-centric, trunk-focused datasets.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

FFinRED: 金融 LLM レッドチームのための専門家ガイドによるベンチマーク生成および評価フレームワーク

既存の安全性ベンチマークは、一般的な敵対シナリオを対象としていますが、金融特有のリスクは見逃しています。金融 LLM は、対象を絞った評価を必要とする規制コンプライアンス違反、不正行為の助長、組織的な信頼の低下に直面しています。 FinRED は、金融専門家と開発された金融 LLM 安全性評価のための専門家ガイドによるレッドチーム フレームワークです。 FinRED は、世界標準 (FATF や EU DORA など) を規制回避から複雑な詐欺に至るまでの脅威にマッピングする新しい 2 レベルの分類を使用し、専門家が定義したスキーマを通じて実際の財務文書をコンテキスト豊富なレッドチームの行動プロンプト (シード) に変換するスケーラブルなパイプラインと統合します。専門家の厳密な検証により、種子の妥当性と有意義な LLM 安全性評価の現実性が確認されます。また、免責条項のチェックを超え、静的な画一的なルーブリックよりも人間の専門家とより緊密に連携し、重大な偽陰性を 28 から 12 に削減する、専門家によって検証された金融固有のルーブリックも提供しています。国際的に採用されているリスク管理および情報セキュリティ基準 (ISO/IEC 27001 など) と連携して、FinRED は韓国の金融セキュリティ協会 (FSI) の規制サンドボックスに導入されています。実際の金融サービスにおける生成 AI セキュリティ評価。二重使用のリスクを軽減するために、データセット、生成パイプライン、プロンプト テンプレート、および評価フレームワークは、https://github.com/selectstar-ai/FinRED-paper および https://huggingface.co/datasets/datumo/FinRED で資格のある研究者向けに制限されています。

原文 (English)

FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU DORA) to threats ranging from regulatory evasion to complex fraud, integrated with a scalable pipeline that converts real financial documents into context-rich red-teaming Behavioral Prompts (seeds) through an expert-defined schema. Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation. We also provide an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely with human experts than static one-size-fits-all rubrics, and reduces critical false negatives from 28 to 12. Aligned with internationally adopted risk-management and information-security standards (e.g., ISO/IEC 27001), FinRED is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox for generative AI security evaluation in real financial services. To mitigate dual-use risks, the dataset, generation pipeline, prompt template, and evaluation framework are gated for qualified researchers at https://github.com/selectstar-ai/FinRED-paper and https://huggingface.co/datasets/datumo/FinRED.

13:00 JST研究/論文

SL-S4Wave: 構造化状態空間モデルを使用した生理学的波形の自己教師あり学習

心電図 (ECG) などの長いシーケンスの医療時系列データのモデリングは、高いサンプリング レート、マルチチャネル信号の複雑さ、固有のノイズ、およびラベル付きデータの制限により、重大な課題を引き起こします。畳み込みニューラル ネットワークなどのさまざまなエンコーダ アーキテクチャに基づく最近の自己教師あり学習 (SSL) 手法は、ラベルのないデータから表現を学習するために提案されていますが、長距離の依存関係やノイズ不変の特徴を捕捉するには不十分であることがよくあります。構造化状態空間モデル (S4) は長いシーケンスのモデリングに優れていますが、既存の S4 アーキテクチャはマルチチャネル生理学的波形の固有の特性を捉えることができません。この研究では、構造化状態空間モデルに基づいて構築された調整されたエンコーダーと対照学習を組み合わせた自己教師あり学習フレームワークである SL-S4Wave を提案します。エンコーダには、マルチスケール サブカーネルを使用した多層グローバル コンボリューションが組み込まれており、ノイズの多い高解像度のマルチチャネル波形におけるきめの細かいローカル パターンと長距離の時間依存性の両方をキャプチャできます。現実世界のデータセットでの広範な実験により、SL-S4Wave は、(1) 困難な不整脈検出タスクにおいて、常に最先端の教師付きベースラインおよび自己教師付きベースラインを上回るパフォーマンスを示し、(2) 大幅に少ないラベル付きサンプルで高いパフォーマンスを達成し、強力なラベル効率を示し、(3) 長い波形セグメントで堅牢なパフォーマンスを維持し、既存のアプローチのほとんどが効率的にモデル化できない長いシーケンスにおける複雑な時間ダイナミクスをモデル化する能力を強調し、(4) 転送目に見えないタイプの不整脈に効果的であり、その堅牢なクロスドメインの一般化が強調されています。さらに、複数のEEGタスクでSL-S4Waveを評価し、強力なベースラインを超えて優れたパフォーマンスを達成し、心臓波形を超えたアプローチの一般化可能性を実証しました。

原文 (English)

SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models

Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates, multichannel signal complexity, inherent noise, and limited labeled data. While recent self-supervised learning (SSL) methods, based on various encoder architectures such as convolutional neural networks, have been proposed to learn representations from unlabeled data, they often fall short in capturing long-range dependencies and noise-invariant features. Structured state space models (S4) excel at long-sequence modeling, but existing S4 architectures fail to capture the unique characteristics of multichannel physiological waveforms. In this work, we propose SL-S4Wave, a self-supervised learning framework that combines contrastive learning with a tailored encoder built on structured state space models. The encoder incorporates multi-layer global convolution using multiscale subkernels, enabling the capture of both fine-grained local patterns and long-range temporal dependencies in noisy, high-resolution multichannel waveforms. Extensive experiments on real-world datasets demonstrate that SL-S4Wave (1) consistently outperforms state-of-the-art supervised and self-supervised baselines in a challenging arrhythmia detection task, (2) achieves high performance with significantly fewer labeled examples, showcasing strong label efficiency, and (3) maintains robust performance on long waveform segments, highlighting its capacity to model complex temporal dynamics in long sequences that most existing approaches fail to efficiently model, and (4) transfers effectively to unseen arrhythmia types, underscoring its robust cross-domain generalization. We additionally evaluate SL-S4Wave on multiple EEG tasks, achieving superior performance over strong baselines, demonstrating generalizability of our approach beyond cardiac waveforms.

13:00 JSTエージェント研究/論文

AI エージェントの生物学的能力とリスクの測定

この論文は、急速に浮上している政策課題、つまり、AI 科学者の生物学的能力とリスクに関する信頼できる証拠をどのように生成し、解釈するか、または多段階の科学的タスクを自律的または協調的に実行できるエージェント型 AI システムに対処します。これらのシステムが実際の研究ワークフローに組み込まれるにつれて、意思決定者はますます、暗黙的または十分に文書化されていない、基礎となる設計上の選択にその意味が左右される評価結果に直面することが増えています。私たちは、AI を活用した生物学的リスクに関する現在の証拠を総合し、これらのシステムを評価するための、有望ではあるが解釈に依存するツールとして生物学的薬剤の評価を紹介します。私たちの中心的な貢献は、私たち自身の評価に基づいた実践的で経験に基づいた一連の考慮事項であり、評価の定義、設計、実行、スコアリング、文書化に関する選択が、結果がリスクに関して何を意味するのか、何を意味しないのかを実質的に形作る方法を示しています。この分析は、政策立案者が生物学的評価の結果を適切な注意をもって解釈できるようにすることを目的としています。官民の資金提供者を AI 生物学評価研究への高レバレッジ投資に誘導する。新しい AI システムを評価するバイオセキュリティ専門家をサポートします。二次対象者には、フロンティア AI ラボ、AI プロバイダー、科学機関、サードパーティ評価組織内でエージェント評価を設計または実施する研究者が含まれます。

原文 (English)

Measuring Biological Capabilities and Risks of AI Agents

This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks. As these systems enter real research workflows, decision-makers increasingly face evaluation results whose meaning depends on underlying design choices that are often implicit or under-documented. We synthesize current evidence on AI-enabled biological risks and introduce biological agentic evaluations as a promising, but interpretation-sensitive, tool for assessing these systems. Our central contribution is a set of practical, experience-grounded considerations -- drawing from our own evaluations -- that show how choices around defining, designing, running, scoring, and documenting evaluations materially shape what results do and do not imply about risk. The analysis is intended to help policymakers interpret biological evaluation outputs with appropriate caution; guide public and private funders toward high-leverage investments in AI-biology evaluation research; and support biosecurity practitioners assessing emerging AI systems. A secondary audience includes researchers designing or conducting agentic evaluations within frontier AI labs, AI providers, scientific institutions, and third-party evaluation organizations.

13:00 JSTロボティクスQwen

共同ポリシー: 音楽パフォーマンスのための応答性の高い人間とロボットの共創

芸術は長い間、人間の創造性の極めて重要な表現として存在してきました。身体化された人工知能は、身体化されていないデジタル コンテンツではなく、物理的なアクションを通じて生成モデルがその創造性に参加するためのルートを提供します。ロボットによる音楽の共同制作では、意味論的な音楽の理解をリアルタイムで物理的に実行可能なパフォーマンスに結び付けることが困難です。私たちは、意味論的な意図の基礎付け、制約された音楽のバリエーション、および視覚運動の実行を分離する、人間とロボットの音楽共創のためのフレームワークである Co-policy を紹介します。音楽セマンティクスを基礎付けるために、Co-policy は、事前推論セマンティック アンカーと微調整された Qwen-vl プランナー (F-Qwen) を使用して、音声、ライブ音楽シード、および視覚的観察を構造化された共創計画に変換します。低レイテンシの実行をサポートするために、Co-policy はガウス混合視覚モーター ポリシー (GMP) を導入します。これは、ターゲット ノートとビジュアル コンテキストを 1 回の順方向パスでマルチモーダル ロボット アクションにマッピングする条件付き混合密度ポリシーとして実装されます。ユーザーが指定したノートを単に再生するだけのロボット再生システムとは異なり、Co-policy は音楽的制約と物理的制約の両方の下で相補的な音楽応答を生成します。実際のロボットのチャイム実験、アブレーション、および専門家による評価では、拡散政策およびアブレートされたベースラインと比較して、意図の調整、実行精度、応答頻度が向上していることが示されており、身体化された人間と AI の共創の重要な要件として、物理的に根拠のあるアクションの生成がサポートされています。

原文 (English)

Co-policy: Responsive Human-Robot Co-Creation for Musical Performances

Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to participate in that creativity through physical action rather than disembodied digital content. In robotic music co-creation, it is challenging to connect semantic musical understanding with real-time and physically executable performance. We present Co-policy, a framework for human-robot musical co-creation that separates semantic intent grounding, constrained musical variation, and visuomotor execution. To ground musical semantics, Co-policy uses pre-inference semantic anchors and a fine-tuned Qwen-vl planner (F-Qwen) to transform speech, live musical seeds, and visual observations into structured co-creation plans. To support low-latency execution, Co-policy introduces a Gaussian-Mixture Visuomotor Policy (GMP), implemented as a conditional mixture-density policy that maps target notes and visual context to multimodal robot actions in a single forward pass. Unlike robotic playback systems that merely reproduce user-specified notes, Co-policy generates complementary musical responses under both musical and physical constraints. Real-robot chime experiments, ablations, and expert evaluation show improved intent alignment, execution accuracy, and response frequency over diffusion-policy and ablated baselines, supporting physically grounded action generation as a key requirement for embodied human-AI co-creation.

13:00 JST画像/動画生成

空間認識削減フレームワーク: 効率的で忠実な視覚状態空間モデルに向けて

Mamba は、長いビジュアル シーケンスのモデリングにおいて高い効率性を示します。ただし、トークン削減が構造的に強化された Mamba の亜種に適用されると、これらのモデルは深刻なパフォーマンスの低下を示します。我々は、この劣化の原因は、選択的走査メカニズムに必要な 2 次元構造の前提に違反する既存の縮小手法の空間に依存しない性質にあると考えています。この研究では、圧縮プロセス全体を通じて構造の完全性を維持するように設計された空間認識トークン削減フレームワークである STORM を提案します。 STORM は、リダクションを空間単位の構造化操作に再定式化し、局所的な制約を強制してグリッド トポロジと近隣の一貫性の両方を維持します。プラグアンドプレイ モジュールとして、STORM は既存のリダクション パイプラインにトレーニングなしで明示的な空間認識を提供します。実証結果は、STORM がトレーニング不要の設定下で、多様なビジョンの Mamba バックボーン全体で最先端の枝刈り精度を達成することを示しています。特に、STORM は VMamba で大幅な精度の回復を実現し、トップ 1 の精度で従来の方法を最大 63.3\% 上回りました。一方、STORM は、PlainMamba で精度の低下が 1.0\% のみであり、ViT と同等のパフォーマンスを達成します。

原文 (English)

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a severe performance collapse. We attribute this degradation to the spatially agnostic nature of existing reduction methods, which violate the two-dimensional structural premise required by the selective scanning mechanism. In this work, we propose STORM, a spatial-aware token reduction framework designed to maintain structural integrity throughout the compression process. STORM reformulates reduction into a structured operation on spatial units, enforcing localized constraints to maintain both grid topology and neighborhood coherence. As a plug-and-play module, STORM equips existing reduction pipelines with explicit spatial awareness without any training. Empirical results demonstrate that STORM achieves state-of-the-art pruning accuracy across diverse vision Mamba backbones under training-free settings. Notably, STORM delivers a substantial accuracy recovery on VMamba, outperforming prior methods by up to 63.3\% in top-1 accuracy. Meanwhile, STORM incurs only a 1.0\% accuracy drop on PlainMamba, achieving performance comparable to ViT.

13:00 JST画像/動画生成

セマンティックセグメンテーション産業用アプリケーションにおける注釈プロセスの高速化

現在の機械学習モデルは通常、大規模で十分に注釈が付けられたデータセットを必要とします。ただし、注釈のプロセスがボトルネックになることが多く、複雑さが増すと人的ミスが発生する可能性が高くなります。この文脈の中で、この論文の目標は、教師なしアルゴリズムを活用して、工業材料科学における複雑なセマンティック セグメンテーション問題に対するデータ アノテーションの効率を向上させることです。これまでの研究では標識時間を定量化しており、教師なしの方法を検討した研究もありました。ただし、私たちの知る限り、これは教師なしアルゴリズムがラベル付けプロセスをどの程度加速するかを定量化した最初の研究です。私たちは、材料科学における微細構造特性評価の課題など、高解像度画像の各ピクセルに注釈を付けることを伴うセマンティック セグメンテーション タスクに焦点を当てて、この骨の折れるプロセスをどの程度加速できるかを検証することを目指しています。具体的には、教師なしコンピューター ビジョン アルゴリズムを使用することで、ラベル付けプロセスに必要な時間が 170 時間から 37 時間に短縮され、約 78\% の削減が達成されることを実証します。私たちが扱うデータセットには 1280x959 および 960x703 の大きな画像が含まれているため、アノテーション タスクはさらに複雑になります。これらの課題にもかかわらず、当社はこれまでで最大の公開鉄鋼微細構造セグメンテーション データセットを作成して共有しており、永久 DOI 付きの MIT ライセンスの下で利用可能であり、完全に注釈付きの高解像度データセットを現場に提供しています。さらに、これは、ゼロからのラベリング時間 (以前の研究では一般的なアプローチ) と、これらの教師なしアルゴリズムをアノテーション前のステップとして使用した場合のラベリング時間を比較した最初の研究です。さらに、このデータセットでトレーニングされ、現場の専門家によって検証され、産業環境に導入された深層学習モデルを提供し、この公開データセットの初期ベンチマークとして機能します。

原文 (English)

Speeding up the annotation process in semantic segmentation industrial applications

Current machine learning models commonly require large and well-annotated datasets. However, the annotation process often becomes a bottleneck, with increased complexity leading to higher chances of human errors. Within this context, our goal in this paper is to leverage unsupervised algorithms to improve data annotation efficiency for complex semantic segmentation problems in industrial materials science. Previous research has quantified labeling time and others explored unsupervised methods. However, to the best of our knowledge, this is the first study to quantify how much unsupervised algorithms accelerate the labeling process. We aim to validate the extent to which this laborious process can be accelerated, focusing on semantic segmentation tasks that involve annotating each pixel of high-resolution images, such as the microstructure characterization challenge in materials science. Specifically, we demonstrate that by using unsupervised computer vision algorithms, the time required for the labeling process can be reduced from 170 hours to 37 hours, achieving an approximate reduction of 78\%. The dataset we work with includes large images of dimensions 1280x959 and 960x703, which further increases the complexity of the annotation task. Despite these challenges, we create and share the largest public steel microstructure segmentation dataset to date, available under MIT License with permanent DOI, contributing a fully annotated, high-resolution dataset to the field. Additionally, this is the first work to compare the labeling time from scratch (a common approach in previous studies) to the labeling time when using these unsupervised algorithms as a pre-annotation step. Furthermore, we provide a Deep Learning model trained on this dataset, validated by field experts, and deployed in an industrial setting, serving as an initial benchmark for this public dataset.

13:00 JST画像/動画生成

オプティカル フローを学習するための普遍的な制約としての三角形の一貫性

我々は、オプティカル フローの第一原理制約として三角整合性を提案します。これは、ネットワーク アーキテクチャ、監視タイプ、データセットに依存せず、画像ペアとマルチフレーム設定の両方に適用されます。このシンプルだが強力な制約は、2 つのフローを構成して 3 つ目のフローを誘導し、3 つのフロー間の一貫性を強制することです。合成されたフローは、(i) 画像ペアから生成され、サイクルの一貫性が得られます。 (ii) 複数のビデオ フレーム。時間的連鎖を通じてより長距離の動きを生成します。または (iii) 画像ペアを制御された合成変換と組み合わせて、データ拡張となります。この三角形の一貫性により、無視できるほどの計算オーバーヘッドが発生し、追加の注釈は必要ありません。オプティカル フローのジオメトリから直接導出されるため、モデル固有の仮定に依存せず、オプティカル フロー トレーニング用の「ユニバーサル」プラグ アンド プレイ コンポーネントとして機能します。実験では、教師あり、教師なし、転移学習の設定全体で一貫した改善が見られました。

原文 (English)

Triangular Consistency as a Universal Constraint for Learning Optical Flow

We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision type, and dataset, and applies to both image-pair and multi-frame settings. This simple but powerful constraint is to compose two flows to induce a third flow and enforce consistency among the three. The composed flows may arise from (i) image pairs, yielding cycle consistency; (ii) multiple video frames, producing longer-range motion through temporal chaining; or (iii) image pairs combined with controlled synthetic transformations, which becomes data augmentation. This triangular consistency introduces negligible computational overhead and requires no additional annotations. Since it is derived directly from the geometry of optical flow, it does not rely on model-specific assumptions and serves as a ``universal'' plug-and-play component for optical flow training. Experiments show consistent improvement across supervised, unsupervised, and transfer learning settings.

13:00 JST研究/論文

SIMBA:FY-4A GIIRS ハイパースペクトル赤外線放射を NWP アプリケーションに向けてモデル化するための双方向検索フォワード シミュレーション フレームワーク

ハイパースペクトル赤外線観測は、大気の温度と湿度の垂直構造に関する豊富な情報を提供するため、数値天気予報 (NWP) にとって重要なデータ ソースです。しかし、既存の深層学習手法の多くは、放射輝度から大気プロファイルへの一方向の検索に主に焦点を当てており、逆放射輝度シミュレーションプロセスや、大気状態空間と放射輝度観測空間の間の整合性については十分に考慮されていません。この研究では、NWPアプリケーション向けのFY-4A GIIRSハイパースペクトル赤外放射輝度モデリングのための統合された双方向検索-順方向シミュレーションフレームワークであるSIMBAを提案します。このフレームワークは、大気プロファイルの取得と放射輝度の再構成を共同で実行し、サイクル一貫性制約を導入して 2 つのプロセス間の結合を強化し、双方向 Mamba 状態空間モジュールを採用して圧力レベルに沿った長距離依存性を捕捉します。併置されたFY-4A GIIRS観測とERA5再解析データを使用して、提案された方法は、温度取得、比湿度取得、長波放射輝度再構成、および中波放射輝度再構築に関して評価されます。実験結果は、SIMBA が検索タスクと再構成タスクの両方でいくつかの代表的な深層学習ベースラインを上回るパフォーマンスを示すことを示し、アブレーション実験は双方向設計とサイクル一貫性メカニズムの寄与を確認します。これらの結果は、提案されたフレームワークが大気プロファイルの統合検索とハイパースペクトル赤外放射輝度モデリングに有効であることを示し、将来のヤコビアン関連解析と NWP 指向の拡張の可能性を示唆しています。

原文 (English)

SIMBA: ABidirectional Retrieval Forward Simulation Framework for Modeling FY-4A GIIRS Hyperspectral Infrared Radiances Toward NWP Applications

Hyperspectral infrared observations are an important data source for numerical weather prediction (NWP) because they provide rich information on the vertical structure of atmospheric temperature and humidity. However, most existing deep learning methods mainly focus on one-way retrieval from radiances to atmospheric profiles, while the reverse radiance simulation process and the consistency between atmospheric state space and radiance observation space are insufficiently considered. In this study, we propose SIMBA, a unified bidirectional retrieval-forward simulation framework for FY-4A GIIRS hyperspectral infrared radiance modeling toward NWP applications. The framework jointly performs atmospheric profile retrieval and radiance reconstruction, introduces a cycle-consistency constraint to strengthen the coupling between the two processes, and employs a bidirectional Mamba state-space module to capture long-range dependencies along pressure levels. Using collocated FY-4A GIIRS observations and ERA5 reanalysis data, the proposed method is evaluated for temperature retrieval, specific humidity retrieval, long-wave radiance reconstruction, and medium-wave radiance reconstruction. Experimental results show that SIMBA outperforms several representative deep learning baselines across both retrieval and reconstruction tasks, while ablation experiments confirm the contribution of the bidirectional design and cycle-consistency mechanism. These results demonstrate that the proposed framework is effective for joint atmospheric profile retrieval and hyperspectral infrared radiance modeling, and suggest potential for future Jacobian-related analysis and NWP-oriented extensions.

13:00 JSTLLM/生成AI画像/動画生成

マルチモーダル LLM の信頼度調整: 医療 VQA による実証研究

マルチモーダル大規模言語モデル (MLLM) は医療業務において大きな可能性を示しますが、MLLM によって導き出される信頼は実際の精度と一致しないことが多く、誤診や正しいアドバイスの見落としにつながる可能性があります。この研究は、医療 MLLM の精度と信頼性の関係についての初めての包括的な分析を示しています。これは、医療視覚質問応答 (VQA) の信頼度調整を向上させることを目的として、マルチ戦略融合ベース尋問 (MS-FBI) と補助専門家 LLM 評価を組み合わせた新しい方法を提案しています。実験では、私たちの方法が 3 つの医療 VQA データセット全体で予想キャリブレーション誤差 (ECE) を平均 40% 削減し、MLLM の信頼性を大幅に向上させることが実証されました。この調査結果は、医療における MLLM のドメイン固有のキャリブレーションの重要性を強調し、AI 支援診断により信頼できるソリューションを提供します。

原文 (English)

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA

Multimodal Large Language Models (MLLMs) show great potential in medical tasks, but their elicited confidence often misaligns with actual accuracy, potentially leading to misdiagnosis or overlooking correct advice. This study presents the first comprehensive analysis of the relationship between accuracy and confidence in medical MLLMs. It proposes a novel method that combines Multi-Strategy Fusion-Based Interrogation (MS-FBI) with auxiliary expert LLM assessment, aiming to improve confidence calibration in Medical Visual Question Answering (VQA). Experiments demonstrate that our method reduces the Expected Calibration Error (ECE) by an average of 40\% across three Medical VQA datasets, significantly enhancing MLLMs' reliability. The findings highlight the importance of domain-specific calibration for MLLMs in healthcare, offering a more trustworthy solution for AI-assisted diagnosis.

13:00 JSTLLM/生成AI画像/動画生成研究/論文

ROSE: マルチモーダル モデルにおける認識と行動のギャップのベンチマーク

マルチモーダル大規模言語モデル (MLLM) は、視覚情報に基づいて動作することがますます期待されていますが、同じシーンでも、異なるタスク コンテキストでは異なるアクションが必要になる場合があります。モデルは同じ視覚的証拠を現在のコンテキストで必要なアクションにどの程度確実に変換できるでしょうか?この質問に答えるために、\textsc{ROSE} (\textbf{R}参照条件付き \textbf{O}ddity および \textbf{S}ymbolic \textbf{E}xecution) を導入します。これは、領域制約と必要なシンボリック出力を変更しながらビジュアル シーンを固定する制御されたベンチマークです。 \textsc{ROSE} は、カウントと調整アクションのタスクを組み合わせて、モデルが暗黙的な多数派参照を推論し、変化するコンテキストの下で結果として得られるきめの細かい視覚的証拠に基づいて動作できるかどうかをテストします。最近の 9 つの MLLM では、人間のパフォーマンスが 98.8\% であるにもかかわらず、カウント指向のタスクから領域条件付きアクションまでパフォーマンスが 44.5 パーセント ポイントも低下しました。このギャップは、同じモデルが正しいカウントを返すペアのシーンと領域では継続しますが、グローバル クリックと一致するローカル コントロールでは、座標のグラウンディングが損失の一部のみを説明していることが示され、共有された視覚的証拠をコンテキスト固有のアクションに変える際の明確なモデル依存のボトルネックが明らかになります。

原文 (English)

ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models

Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts. How reliably can a model turn the same visual evidence into the action required by the current context? To answer this question, we introduce \textsc{ROSE} (\textbf{R}eference-conditioned \textbf{O}ddity and \textbf{S}ymbolic \textbf{E}xecution), a controlled benchmark that holds the visual scene fixed while varying region constraints and required symbolic outputs. Through coupled counting and coordinate-action tasks, \textsc{ROSE} tests whether models can infer an implicit majority reference and act on the resulting fine-grained visual evidence under changing contexts. Across nine recent MLLMs, performance drops by as much as 44.5 percentage points from counting-oriented tasks to region-conditioned action, despite 98.8\% human performance. The gap persists on paired scenes and regions for which the same model returns the correct count, while global-click and matched local controls show that coordinate grounding explains only part of the loss, revealing a distinct, model-dependent bottleneck in turning shared visual evidence into context-specific actions.

13:00 JST研究/論文

アルゴリズムとヒューマンのマネージャー: インドのギグエコノミーにおける AI、アプリ、労働者

この論文では、アルゴリズム管理に焦点を当てて、インドのブルーカラーのギグ経済に対する人工知能とデジタル技術の影響を調査します。この論文では、インドのブルーカラー ギグ経済に対する人工知能とデジタル テクノロジーの影響を検証し、ライドシェアリングや配達などの位置情報ベースのサービスにおける作業の割り当て、監視、評価に同氏が使用する自動システムのアルゴリズム管理に焦点を当てています。この研究では、社会正義の枠組みと、16 人のギグワーカーと 21 人の主要関係者へのインタビューからなる混合方法アプローチを使用して、二重の現実を明らかにしました。AI を活用したシステムは、仕事へのアクセスを拡大し、業務効率を生み出す一方で、同時に公平性、透明性、労働者の尊厳に関連する重大な課題をもたらします。主要な調査結果は、アルゴリズム システムが設計上不透明で、不公平な結果を生み出し、追加の労働に比例した賃金で報いるように構造化されていないことを明らかにしています。この研究は、技術効率と人間の説明責任が対立するのではなく、連携して機能するアルゴリズム ヒューマン マネージャー フレームワークという、実用的なハイブリッド ガバナンス モデルを提唱しています。この調査結果は、インドおよびグローバル・サウス全域のギグ・エコノミーのための公平な AI ガバナンス・フレームワークの設計に取り組んでいる政策立案者、プラットフォーム企業、市民社会組織に影響を及ぼします。

原文 (English)

The Algorithmic-Human Manager: AI, Apps, and Workers in the Indian Gig Economy

This paper examines the impact of artificial intelligence and digital technologies on the blue-collar gig economy in India, focusing on algorithmic management. This paper examines the impact of artificial intelligence and digital technologies on the blue collar gig economy in India, focusing on algorithmic management he use of automated systems to allocate, monitor, and evaluate work in location-based services such as ride sharing and delivery. Using a social justice framework and a mixed-methods approach comprising interviews with 16 gig workers and 21 key stakeholders, the study uncovers a dual reality: while AI-powered systems expand access to work and generate operational efficiencies, they simultaneously introduce significant challenges related to fairness, transparency, and worker dignity. Key findings reveal that algorithmic systems are opaque by design, produce inequitable outcomes, and are not structured to reward additional labour with proportionate pay. The study advocates for a pragmatic hybrid governance model an Algorithmic Human Manager framework in which technological efficiency and human accountability operate together rather than in opposition. The findings carry implications for policymakers, platform companies, and civil society organizations working to design equitable AI governance frameworks for the gig economy in India and across the Global South.

13:00 JSTLLM/生成AIエージェント

静的エンドポイントを超えて: 柔軟なエージェント Web サービスのインターフェイスとしてのツール プログラム

エージェント Web 時代では、LLM ベースのエージェントが Web サービスをツールとして呼び出すことが増えていますが、ほとんどのインターフェースは \emph{静的エンドポイント} のままで、ループ、条件、結合、再試行などの長期的なワークフローをうまく表現できません。私たちは、エージェントのツール インテントを \emph{実行可能ツール プログラム} として表現する ToolPro を提示します。これは、明示的な効果タイプを備えた複数ステップのサービス インタラクションをコンパクトにエンコードします。 ToolPro は、制約に基づいたプログラムの構築、状態を変更する 1 回だけの呼び出しに対する効果を認識した再生、およびプログラムの実行が段階的な呼び出しよりも優れたパフォーマンスを示すプロファイル駆動のポリシーを組み合わせています。 WebAssembly サンドボックスを使用して MCP スタイルのサービス上で ToolPro をインスタンス化し、実際のアプリケーションのさまざまなワークフローで評価します。 ToolPro は、エンドツーエンドの遅延を最大 53.4\% 削減し、クライアント側のトラフィックを最大 96.1\% 削減します。ネットワーク遅延が高く、ワークフローが複雑になると、さらに大きな効果が得られます。

原文 (English)

Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services

In the agentic web era, LLM-based agents increasingly invoke web services as tools, yet most interfaces remain \emph{static endpoints} that poorly express long-horizon workflows with loops, conditionals, joins, and retries. We present ToolPro, which represents an agent's tool intent as an \emph{executable tool program} that compactly encodes multi-step service interactions with explicit effect types. ToolPro combines constraint-guided program construction, effect-aware replay for exactly-once state-modifying calls, and a profile-driven policy that decides when program execution outperforms stepwise calling. We instantiate ToolPro over MCP-style services with WebAssembly sandboxing and evaluate it on diverse workflows of real-world applications. ToolPro reduces end-to-end latency by up to 53.4\% and client-side traffic by up to 96.1\%, with larger gains under higher network latency and workflow complexity.

13:00 JST画像/動画生成ロボティクス

Tri-Info: 情報理論による VLA モデルの一般化可能で解釈可能な故障予測

Vision-Language-Action(VLA)モデルは、多様なタスクにわたってますます導入されていますが、物理的な相互作用が取り返しのつかない害を引き起こす可能性があるブラックボックスのままであるため、一般化可能で解釈可能な障害検出が不可欠となっています。私たちは、ロールアウトの成功と失敗には、体系的に異なる情報理論的特徴があることが観察されています。これに基づいて、VLA 制御を閉ループ情報パイプラインとして形式化し、アクションが多様性を保ち、時間的に一貫性があり、状態遷移と結合しているかどうかを捕捉する三重情報理論 (Tri-Info) 信号を導き出します。 6 つの VLA モデルと 3 つのベンチマーク環境にわたって、Tri-Info はドメイン内の最も強力なベースラインと一致します。さらに、Tri-Info は再トレーニングすることなく、アーキテクチャ、環境、シミュレーションと現実のギャップを越えて転送し、以前の検出器が偶然に崩壊してしまうような現実世界のタスクで 83\% の精度に達します。これにより、Tri-Info は、強力なクロスドメイン一般化によって障害を検出するだけでなく、根本的な障害モードの解釈可能な診断も提供する、シンプルかつ強力な方法として確立されます。

原文 (English)

Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions can cause irreversible harm, making generalizable and interpretable failure detection essential. We observe that successful and failed rollouts carry systematically different information-theoretic signatures. Building on this, we formalize VLA control as a closed-loop information pipeline and derive the Triple Information-theoretic (Tri-Info) signals that capture whether actions remain diverse, temporally consistent, and coupled to state transitions. Across six VLA models and three benchmark environments, Tri-Info matches the strongest baselines in-domain. Moreover, Tri-Info transfers across architectures, environments, and the sim-to-real gap without retraining, reaching 83\% accuracy on real-world tasks where prior detectors collapse to chance. This establishes Tri-Info as a simple yet powerful method that not only detects failures with strong cross-domain generalization, but also delivers interpretable diagnostics of the underlying failure modes.

13:00 JSTLLM/生成AIエージェント

Connect the Dots: 強化学習によるクロスドメイン一般化による長期ライフサイクルエージェントのための LLM のトレーニング

この研究では、ライフサイクルの長いエージェントに必要なメタ機能である「Connect the Dots」(CoD) を大規模言語モデル (LLM) にトレーニングするための一般的なフレームワークを示します。LLM ベースの AI エージェントが環境にデプロイされると、環境を継続的に探索し、自身の経験から学習し、環境に関するコンテキストを反復的に自己更新しながら、長いシーケンスのタスクを解決します。これにより、更新されたコンテキストに基づいて将来のタスクのパフォーマンスが徐々に向上します。 CoD フレームワークの主なコンポーネントには次のものが含まれます。(1) 解決タスクと更新コンテキストのエピソードをインターリーブする長いロールアウト シーケンスを備えたエンドツーエンドの強化学習 (RL) のためのアルゴリズム設計とインフラストラクチャ。 (2) トレーニング中に LLM の目標とするメタ能力を奨励し、引き出すため、また評価中に進捗を忠実に測定するためのタスクと環境。我々は、きめ細かいクレジット割り当てを備えた GRPO スタイルの RL アルゴリズム、および (ドメイン固有の LLM 機能や標準的なタスクごとの RL ではなく) ターゲットのメタ機能に合わせて調整されたタスクと環境を含む、CoD フレームワークの概念実証の実装を紹介します。経験的な結果は、CoD 設定におけるエンドツーエンドの RL トレーニングの有効性を検証し、誘発されたメタ機能の分布外一般化 (トレーニング ドメイン内、異なるドメイン間、および CoD からラルフ ループ設定へ) の可能性を実証しています。私たちの CoD の調査は、これまでの研究のいくつかのラインを結び付け、LLM と AI エージェントを進歩させるための新たな機会を開きます。さらなる研究と応用を促進するために、\url{https://github.com/agentscope-ai/Trinity-RFT/tree/research/cod/examples/research_cod} で実装をリリースします。

原文 (English)

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously exploring the environment, learning from its own experiences, and iteratively self-updating its context about the environment, thereby achieving progressively better performance on future tasks conditioned on the updated context. Major components of the CoD framework include: (1) algorithm design and infrastructure for end-to-end reinforcement learning (RL) with long rollout sequences interleaving solve-task and update-context episodes; (2) tasks and environments for incentivizing and eliciting the targeted meta-capability in LLMs during training, as well as for faithfully measuring progress during evaluation. We present proof-of-concept implementations of the CoD framework, including a GRPO-style RL algorithm with fine-grained credit assignment, as well as tasks and environments tailored to the targeted meta-capability (rather than domain-specific LLM capabilities or standard task-by-task RL). Empirical results validate the efficacy of end-to-end RL training in the CoD setting, and demonstrate the potential for out-of-distribution generalization -- within the training domains, across different domains, and from CoD to Ralph-loop settings -- of the elicited meta-capability. Our investigation of CoD connects several lines of prior works, and opens up new opportunities for advancing LLMs and AI agents. To facilitate further research and applications, we release our implementations at \url{https://github.com/agentscope-ai/Trinity-RFT/tree/research/cod/examples/research_cod}.

13:00 JST研究/論文

StreamKL: 注意の蒸留を高めるための高速でメモリ効率の高い KL ダイバージェンス

注意蒸留は、カルバック・ライブラー (KL) 発散を最小限に抑えることで、ある注意分布を別の注意分布と一致するように訓練するもので、知識蒸留、モデル圧縮、継続学習、および疎注意 LLM トレーニングで広く使用されています。ただし、既存のアプローチでは、KL 削減を計算する前に両方のアテンション分布を具体化するため、$O(N_QN_K)$ のメモリ コストと IO コストが発生し、コンテキスト長が長いと法外なコストになります。我々は、この二次実体化を排除するアテンション KL 発散のための最初の融合 GPU プリミティブである StreamKL を紹介します。 StreamKL は、結合された 2 つの分布 KL 削減のための新しいオンライン定式化を導き出し、オンチップ SRAM を介してクエリ キー タイルをストリーミングする単一の 1 パス フォワード カーネルを可能にします。逆方向パスの場合、StreamKL はタイルごとにアテンション確率を再計算し、二次中間値の保存を回避します。さらに、専用の最適化を使用して効率的な GPU カーネルを設計および実装します。実験では、StreamKL が順方向パスと逆方向パスでそれぞれベースライン メソッドと比較して最大 $43\times$ と $14\times$ の高速化を実現することを示しています。最も重要なことは、StreamKL はアテンション蒸留の余分な HBM フットプリントを $O(N_QN_K)$ から $O(1)$ に削減し、単一 GPU でロングコンテキスト蒸留を可能にすることです。

原文 (English)

StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation

Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, model compression, continual learning, and sparse-attention LLM training. However, existing approaches materialize both attention distributions before computing the KL reduction, incurring $O(N_QN_K)$ memory and IO costs that become prohibitive at long context lengths. We present StreamKL, the first fused GPU primitive for attention KL divergence that eliminates this quadratic materialization. StreamKL derives a novel online formulation for the coupled two-distribution KL reduction, enabling a single one-pass forward kernel that streams query-key tiles through on-chip SRAM. For the backward pass, StreamKL recomputes attention probabilities tile-by-tile, avoiding storage of quadratic intermediates. We further design and implement efficient GPU kernels with dedicated optimizations. Experiments show StreamKL delivers up to $43\times$ and $14\times$ speedups over baseline methods in the forward and backward passes, respectively. Most importantly, StreamKL reduces the extra HBM footprint of attention distillation from $O(N_QN_K)$ to $O(1)$, enabling long-context distillation on a single GPU.

13:00 JSTLLM/生成AIエージェント

マルチエージェント ゲームの階層制御: LLM ベースの計画と RL の実行

強化学習(RL)は、逐次的な意思決定において優れたパフォーマンスを達成していますが、報酬がまばらで、状態行動空間が大きく、調整された戦略を学習することが難しいため、複雑なマルチエージェント環境への拡張は依然として困難です。私たちは、事前トレーニングされた大規模言語モデル (LLM) が、エージェントのチームに特化した RL スキル ポリシーの中から選択する集中戦略コントローラーとして機能し、RL ポリシーが事後的な低レベルの実行を処理する階層アーキテクチャを提案します。このハイブリッド システムを、ビヘイビアー ツリー (BT) および \emph{``Flat''} RL (スキル分解を行わないエンドツーエンド トレーニング) ベースラインに対して、競争力のある 2v2 King of the Hill 環境で評価します。 LLM+RL システムは、統計的に手作り BT と同等のタスク パフォーマンス (勝率 46.4\% 対 51.5\%、$p=0.103$) を達成し、両方ともスキル分解なしでトレーニングされた Flat RL を大幅に上回ります。ユーザー調査 ($n=15$) では、参加者の 60\% が、行動の適応性と戦術の変動性を理由に、LLM+RL エージェントが最も人間に近いと認識していることが明らかになりました ($p=0.027$)。これらの結果は、事前トレーニングされた LLM 推論が事前トレーニングされた RL スキルを効果的に調整し、手動のルール エンジニアリングなしで競争力のあるマルチエージェントの調整と優れた知覚信頼性を実現できることを示しています。

原文 (English)

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution

Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments remains challenging due to sparse rewards, large state-action spaces, and the difficulty of learning coordinated strategies. We propose a hierarchical architecture where a pretrained large language model (LLM) acts as a centralized strategic controller that selects among specialized RL skill policies for a team of agents, while RL policies handle reactive low-level execution. We evaluate this hybrid system in a competitive 2v2 King of the Hill environment against behavior tree (BT) and \emph{``Flat''} RL (end-to-end training without skill decomposition) baselines. The LLM+RL system achieves task performance statistically equivalent to hand-crafted BT (46.4\% vs 51.5\% win rate, $p=0.103$) while both significantly outperform Flat RL trained without skill decomposition. A user study ($n=15$) reveals that 60\% of participants perceive LLM+RL agents as the most human-like ($p=0.027$), citing behavioral adaptability and tactical variability. These results demonstrate that pretrained LLM reasoning can effectively orchestrate pretrained RL skills, achieving competitive multi-agent coordination and superior perceived believability without manual rule engineering.

13:00 JSTLLM/生成AIエージェント

低い権限で十分な場合: LLM エージェントでの過剰な権限のツール選択の調査

LLM エージェントが自律的にツールを選択することが増えているため、異なる権限を持つツールの間での選択は安全に関連するものになっています。しかし、これまでのツール選択に関する研究では、安全性に依存しないメタデータの設定に重点が置かれており、特権に依存した選択は十分に検討されていませんでした。このギャップに対処するために、私たちは、権限の低いツールが十分にあるにもかかわらず、エージェントがより高い権限のツールを選択またはエスカレーションする、過剰な権限のツール選択を研究します。 ToolPrivBench を導入して、エージェントが権限の低い代替手段が十分にあるにもかかわらず、より権限の高いツールを選択するかどうかを評価し、初期選択と一時的なツール障害後のエスカレーションの両方を測定します。 8 つのドメインと 5 つの再発リスク パターンにわたって、主流の LLM エージェントでは過剰な権限を持つツールの選択が一般的であり、一時的な障害によってさらに増幅されることがわかりました。さらに、一般的な安全調整では最小権限のツールの選択に確実に移行するわけではなく、プロンプトレベルの制御では一時的な障害が発生した場合に限られた軽減しか提供されないことがわかりました。したがって、私たちは、エージェントに十分な権限の低いツールを優先し、必要な場合にのみエスカレーションするように教える、権限を意識したトレーニング後の防御を導入します。私たちの緩和実験では、この防御により、一般的な機能を維持しながら、不必要な高特権ツールの使用が大幅に削減されることがわかりました。

原文 (English)

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on safety-agnostic metadata preferences, leaving privilege-sensitive choices underexplored. To address this gap, we study over-privileged tool selection, in which an agent selects or escalates to a higher-privilege tool despite a sufficient lower-privilege alternative. We introduce ToolPrivBench to evaluate whether agents choose higher-privilege tools despite sufficient lower-privilege alternatives, measuring both initial selection and escalation after transient tool failures. Across eight domains and five recurring risk patterns, we find that over-privileged tool selection is common among mainstream LLM agents and is further amplified by transient failures. We further find that general safety alignment does not reliably transfer to least-privilege tool choice, while prompt-level controls provide only limited mitigation under transient failures. We therefore introduce a privilege-aware post-training defense that teaches agents to prefer sufficient lower-privilege tools and escalate only when necessary. Our mitigation experiments show that this defense substantially reduces unnecessary high-privilege tool use while preserving general capabilities.

13:00 JSTLLM/生成AIロボティクス

ロボットモバイルフルフィルメントシステムにおける効率的な経路探索のためのニューロモーフィック強化学習フレームワーク

動的な環境変化、限られたワークスペース、および厳しいリアルタイム制約により、ロボット モバイル フルフィルメント システム (RMFS) でのパスファインディングは、従来の検索ベースおよびルールベースの方法にとって困難な問題となっており、通常、計算の複雑性が高く、意思決定の待ち時間が長いという問題があります。強化学習 (RL) は強力な代替手段として登場しましたが、リソースに制約のあるハードウェア上で極めてエネルギー効率の高い学習済みポリシーを展開することは依然として課題です。我々は、完全精度の人工ニューラル ネットワーク (ANN) からニューロモーフィック チップまで、RL でトレーニングされたポリシーの高忠実度の展開を実現するエンドツーエンドのフレームワークである SDQN-RMFS を紹介します。このフレームワークは、まばらなイベントによってトリガーされた場合にのみ計算を行うことで、超低消費電力の RMFS パスファインディングを可能にします。当社のフルスタック パイプラインは次のように動作します。ANN ポリシーは、最初に衝突許容戦略を介して効率的にトレーニングされ、有益な軌道を高密度化してから、ハードラベル知識蒸留アプローチを介してスパイキング ニューラル ネットワーク (SNN) に変換されます。これにより、出力分布の不一致に効果的に対処し、ANN から SNN へのパイプライン全体でポリシー機能を維持しながら、推論レイテンシを大幅に短縮します。ハードウェア実験では、元のトレーニング済みポリシーと同等の意思決定品質を維持しながら、高性能 GPU ベースラインと比較して最大 11,281$\times$ のエネルギー節約とレイテンシのほぼ 2 倍の削減を実証しました。これらの結果は、大規模な RMFS 操作のための実用的でエネルギー持続可能な経路としての物理的神経形態推論を確立します。

原文 (English)

A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems

Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Systems (RMFS) a challenging problem for conventional search- and rule-based methods, which typically suffer from high computational complexity and long decision latency. While reinforcement learning (RL) has emerged as a powerful alternative, deploying learned policies with extreme energy efficiency on resource-constrained hardware remains an open challenge. We present SDQN-RMFS, an end-to-end framework that achieves high-fidelity deployment of an RL-trained policy from a full-precision artificial neural network (ANN) through to a neuromorphic chip. By computing only when triggered by sparse events, this framework unlocks ultra-low-power RMFS pathfinding. Our full-stack pipeline operates as follows: an ANN policy is first efficiently trained via a collision-allowing strategy to densify informative trajectories, and then converted into a spiking neural network (SNN) via a hard-label knowledge distillation approach. This effectively addresses the output distribution mismatch, preserving policy capability across the ANN-to-SNN pipeline while substantially reducing inference latency. Hardware experiments demonstrate up to 11,281$\times$ energy savings and a nearly two-fold reduction in latency compared to a high-performance GPU baseline, while maintaining decision quality on par with the original trained policy. These results establish physical neuromorphic inference as a practical and energy-sustainable pathway for large-scale RMFS operations.

13:00 JSTエージェント

AI Economist Agent: RAG、ナレッジ グラフ、および大規模言語モデルを使用した、モデルに基づいた経済分析のためのエージェント フレームワーク

私たちは、大規模言語モデル (LLM) とナレッジ グラフを使用した経済シナリオ分析のためのエージェント フレームワークを備えた、モデルに基づいた RAG ベースの AI エコノミストを提案します。 LLM は流暢な経済ナラティブを生み出すことができますが、経済学者は多くの場合、経済理論と現実世界のデータに基づいて経済的主張を行うことが求められます。この動機に基づいて、この研究では、経済データと理論を含むナレッジ グラフと LLM ベースのエージェントを利用して分析を計画し、関連する証拠を取得し、適切なモデルを選択し、レポートを生成する RAG ベースの AI エコノミストを提案します。私たちのフレームワークでは、言語モデルだけを使用して定量的な主張を直接生成することはありません。代わりに、明示的なモデルベースの計算に基づいて、AI エージェントを介して取得された証拠にリンクされたナラティブを生成します。私たちはこのフレームワークを AI エコノミスト エージェントと呼んでいます。 AI エコノミスト エージェントを 2 つのアプリケーションで評価します。1 つは米国のインフレ持続と連邦準備制度の政策に関するエコノミスト レポートの生成、もう 1 つは米国の商業用不動産の借り換えストレスに関する銀行ストレス テストのナラティブ生成です。この結果は、生成されたレポートをグラウンディングすることで、経済的な一貫性とトレーサビリティがどのように向上するかを示しています。

原文 (English)

AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models

We propose a model-grounded RAG-based AI economist with an agentic framework for economic scenario analysis using large language models (LLMs) and knowledge graphs. While LLMs can generate fluent economic narratives, economists are often required to make economic claims grounded by economic theory and real-world data. Based on this motivation, this study proposes an RAG-based AI economist, which utilizes knowledge graphs including economic data and theory and LLM-based agents to plan the analysis, retrieve relevant evidence, select appropriate models, and generate reports. In our framework, we do not produce quantitative claims directly with the language model alone; instead, we generate narratives grounded in explicit model-based computations and linked to the retrieved evidence via AI agents. We refer to our framework as an AI economist agent. We evaluate the AI economist agent in two applications: economist report generation for U.S. inflation persistence and Federal Reserve policy, and bank stress-test narrative generation for U.S. commercial real estate refinancing stress. The results illustrate how grounding the generated reports improves their economic coherence and traceability.

13:00 JST画像/動画生成

See-and-Reach: 視野内での UAV の正確な視覚言語ナビゲーション

UAV Vision-Language Navigation (UAV-VLN) は通常、総合的な探索と到達の問題として定式化され、長距離ターゲットの発見と最終ターゲットへのアプローチが一緒に最適化および評価されます。この定式化により、空中に具現化されたエージェントの重要な能力、つまり、UAV が目に見える目標を正確に接地し、目標が視野に入った後に視覚言語の証拠を正確な 3D 動作に変換できるかどうかを評価することが困難になります。この制限に対処するために、我々は UAV-VLN-FOV を導入します。これは、シー・アンド・リーチ段階を分離し、ターミナル到達能力のより診断的な評価を可能にするターゲット可視ナビゲーション タスクです。さらに、動的 3D 方向キューによってガイドされる視覚言語ウェイポイント予測フレームワークである 3DG-VLN を提案し、ターゲットに正確に到達するためのきめの細かい視覚的グラウンディングと空間方向の調整を強化します。具体的には、3DG-VLN は、高解像度の正面図と下方図の観察を適応的に処理して、ターゲット接地のためのきめの細かい視覚的および幾何学的詳細を保存します。また、閉ループ ナビゲーション中にターゲットの相対的な方向をオンラインで更新するため、エージェントはターゲットとの空間的位置合わせを維持し、蓄積された方向のドリフトを軽減できます。このタスクをサポートするために、ターゲット指向の高レベルの命令、高解像度の正面図および下方図の自己中心的観測、および連続 3D ウェイポイント注釈を備えた 2,717 の軌道を含む専用の高解像度ベンチマークを構築します。実験では、3DG-VLN が競合する UAV-VLN ベースラインを上回り、成功率で 13.82\% の向上を達成したことが示されています。実世界の試験では、実用的なシー・アンド・リーチ・ナビゲーションにおける 3DG-VLN の可能性がさらに実証されています。ソース コードとベンチマークは https://github.com/xuefanfu/3DG-VLN で入手できます。

原文 (English)

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly. This formulation makes it difficult to assess a critical capability of aerial embodied agents, namely whether a UAV can accurately ground a visible target and translate vision-language evidence into precise 3D motion once the target enters its field of view. To address this limitation, we introduce UAV-VLN-FOV, a target-visible navigation task that isolates the see-and-reach stage and enables a more diagnostic evaluation of terminal reaching ability. We further propose 3DG-VLN, a vision-language waypoint prediction framework guided by dynamic 3D direction cues to enhance fine-grained visual grounding and spatial direction alignment for precise target reaching. Specifically, 3DG-VLN adaptively processes high-resolution front-view and downward-view observations to preserve fine-grained visual and geometric details for target grounding. It also updates the target-relative direction online during closed-loop navigation, allowing the agent to maintain spatial alignment with the target and reduce accumulated direction drift. To support this task, we construct a dedicated high-resolution benchmark which contains 2,717 trajectories with target-oriented high-level instructions, high-resolution front-view and downward-view egocentric observations, and continuous 3D waypoint annotations. Experiments show that 3DG-VLN outperforms competitive UAV-VLN baselines, achieving a 13.82\% improvement in success rate. Real-world trials further demonstrate the potential of 3DG-VLN for practical see-and-reach navigation. The source code and benchmark are available at https://github.com/xuefanfu/3DG-VLN.

13:00 JSTビジネス/資金調達

ICUにおけるイベントベースのバースト抑制検出のためのEEG基盤モデルの評価

バースト抑制 (BS) は、臨床的に関連のある脳波 (EEG) パターンで、重症患者、特に集中治療室 (ICU) で誘発された昏睡状態の患者の鎮静深度と脳活動を監視するために使用されます。 BS パターンは患者ごとに大幅に異なり、注釈付きのデータセットが不足しているため、自動バースト検出は依然として困難です。最近、EEG Foundation Models (FM) は、いくつかの下流 EEG アプリケーションにわたって有望であることが示されていますが、BS 検出におけるその有用性はまだ解明されていません。我々は、患者固有のキャリブレーションを行わずに、縮小モンタージュICU EEGにおけるバースト検出のためのEEG FMを評価する最初の研究を紹介します。 REVE ベース、LUNA-large、LuMamba-Tiny を適応閾値ベースラインとタスク固有の EEGNet ベースラインと比較します。さらに、従来の EEG ウィンドウベースの分類をイベントベースのバースト検出評価で補完します。これは、バースト エピソードが正しく検出されているかどうかを臨床的に評価するのに役立ち、予期されるアノテーションの変動による影響を軽減します。最良のモデルである REVE ベースは、最高のイベントベース F1 スコア ($0.868 \pm 0.167$) を達成し、EEGNet および適応しきい値処理と比較して、1 分あたりのバースト誤差をそれぞれ 52.1% および 36.2% 削減し、ICU でのスケーラブルな EEG モニタリングのための FM をサポートしました。アブレーション実験では、凍結バックボーン トレーニング、2 ステップの微調整、および LoRA ベースの適応に関して、完全な微調整が最も効果的な適応戦略であることが示され、LUNA-large の場合、凍結バックボーン トレーニングに比べてイベントベースの F1 スコアが最大 $+0.102$ 向上しました。ラベル付きデータセットを減らした場合、事前トレーニング済み REVE ベースはコホートの 25% で $+0.723$ イベントベースの F1 ポイントだけランダム初期化を上回り、限られたラベル付きデータでのバースト検出に適応させた場合の事前トレーニング FM 表現の利点を示しています。

原文 (English)

Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU

Burst suppression (BS) is a clinically relevant electroencephalographic (EEG) pattern used to monitor sedation depth and brain activity in critically ill patients, particularly during induced coma in Intensive Care Units (ICUs). Automatic burst detection remains challenging because BS patterns vary substantially between patients and annotated datasets are scarce. Recently, EEG Foundation Models (FMs) have shown promise across several downstream EEG applications, but their usefulness for BS detection remains unexplored. We present the first study to evaluate EEG FMs for burst detection in reduced-montage ICU EEG without patient-specific calibration. We compare REVE-base, LUNA-large and LuMamba-Tiny with an adaptive thresholding baseline and a task-specific EEGNet baseline. Additionally, we complement conventional EEG window-based classification with event-based burst detection evaluation. This helps assessing clinically whether burst episodes are correctly detected, reducing the impact of expected annotation variability. The best model, REVE-base, achieved the highest event-based F1-score ($0.868 \pm 0.167$) and reduced burst-per-minute error by 52.1% and 36.2% compared to EEGNet and adaptive thresholding respectively, supporting FMs for scalable EEG monitoring in ICU. Ablation experiments showed that full fine-tuning was the most effective adaptation strategy with respect to frozen-backbone training, two-step fine-tuning, and LoRA-based adaptation, improving event-based F1-score over frozen-backbone training by up to $+0.102$ for LUNA-large. With reduced labeled datasets, pretrained REVE-base outperformed random initialization by $+0.723$ event-based F1 points at 25% of the cohort, demonstrating the benefit of pretraining FM representations when adapted to burst detection with limited labeled data.

13:00 JST画像/動画生成

拡散トランスフォーマーの学習可能なグローバル マージによる可変長トークン化

潜在拡散モデル (LDM) はビジュアル合成において主流となっていますが、その品質と計算のトレードオフは、トークナイザーの固定圧縮率によって大きく制限されています。可変長トークナイザー (VLT) は、トークン数を変化させることで適応的な圧縮を約束し、拡散モデルが品質と計算のバランスを柔軟にとれるようにします。ただし、従来の VLT は、順序付けされたトークン シーケンスを切り詰めることによって長さを調整します。これにより、トークンのセマンティクスがトークンの位置に依存し、長さ全体での表現的なアラインメントが崩れます。これにより、潜在分布の横方向のシフトが生じ、単一の可変長拡散モデルが効果的に動作することが妨げられます。これに対処するために、トークンを結合することで長さを調整する新しい可変長トークナイザーを提案します。拡散変換器がマージパターンに従って動作する場合、同様のトークンのマージを奨励することで、直接的なクロスレングス表現の位置合わせが可能になることを示します。従来のマージ方法はデータに依存しており、生成中にマージ パターンにアクセスできなくなるため、データに依存しない学習可能なグローバル マージを導入して、拡散トランスとの互換性を確保します。 ImageNet 256$\times$256 世代では、拡散変換器と統合されたマージ ベースの可変長トークナイザーは、以前の VLT 手法と比較して優れた gFID 計算のトレードオフを実現します。コードは [この https URL](https://github.com/movinghoon/lgm) で入手できます。

原文 (English)

Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers

Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio. Variable-length tokenizers (VLTs) promise adaptive compression by varying token counts, allowing diffusion models to flexibly balance quality and compute. However, conventional VLTs modulate length by truncating ordered token sequences, which makes token semantics depend on token position and breaks representational alignment across lengths. This leads to a cross-length shift in the latent distribution that hinders a single variable-length diffusion model from operating effectively. To address this, we propose a novel variable-length tokenizer that modulates length by merging tokens. We show that encouraging similar tokens to merge enables direct cross-length representation alignment when the diffusion transformer operates according to the merging pattern. Since conventional merging methods are data-dependent, making the merging pattern inaccessible during generation, we introduce learnable global merging, which is data-independent, to ensure compatibility with diffusion transformers. On ImageNet 256$\times$256 generation, our merging-based variable-length tokenizer integrated with a diffusion transformer achieves a superior gFID-compute trade-off compared to prior VLT methods. Code is available at [this https URL](https://github.com/movinghoon/lgm)

13:00 JSTLLM/生成AI画像/動画生成

VLM 内の偽装されたビジュアル コンテキストの隠れた進化

ビジュアル トークンは、生の外部シグナルとして大規模言語モデル (LLM) に入力されます。それらがどのように意味のある表現に変換され、言語空間と相互作用するかは、統合アーキテクチャに完全に依存します。ビジュアル トークンを入力シーケンス内のコンテキスト内プロンプトとして扱うか、LLM の中間層に直接挿入するかによって異なります。これらのアーキテクチャ上の選択が視覚情報にどのような影響を与えるか、また LLM と統合するための内部変換については、制御された比較と理解がまだ十分に行われていません。単一画像、複数画像、およびビデオのベンチマークにわたる同一のトレーニング条件下で、インコンテキストおよびレイヤーごとのインジェクション VLM 統合パラダイムを評価することで、公正な比較を提供します。そうすることで、視覚トークンが、言語構造を欠く生の表現である偽装された視覚コンテキストとして LLM に入力されるが、統合パラダイムに応じて徐々に再形成され、それぞれが視覚信号の根本的に異なる周波数特性を捕捉する、隠れた進化を明らかにします。私たちは、LLM 内のこの進化が、VLM がどのような視覚的特徴を効果的に利用できるか、視覚表現が言語空間とどのように連携するか、そして最終的にはさまざまなタスクにわたって各パラダイムがどのように実行されるかを決定することを示します。さらに、注意の割り当てだけでは不十分であり、パフォーマンスは各層の視覚的表現の品質によって左右されることを示します。

原文 (English)

The Hidden Evolution of Disguised Visual Context inside the VLM

Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends entirely on the integration architecture. Whether by treating visual tokens as in-context prompts within the input sequence or injecting them directly into the LLM's intermediate layers. A controlled comparison and understanding of how these architectural choices affect visual information and its internal transformation to integrate with the LLM remains underexplored. We provide a fair comparison by evaluating in-context and layer-wise injection VLM integration paradigms under identical training conditions across single image, multi-image, and video benchmarks. In doing so, we uncover a hidden evolution where visual tokens enter the LLM as disguised visual context, raw representations lacking linguistic structure, but are progressively reshaped depending on the integration paradigm, each capturing fundamentally different frequency characteristics of the visual signal. We show that this evolution inside the LLM determines what visual features the VLM can utilize effectively, how visual representations align with the language space, and ultimately how each paradigm performs across different tasks. We further demonstrate that attention allocation alone is insufficient, and that performance is driven by the quality of visual representations at each layer.

13:00 JSTLLM/生成AI

IHUBERT: ペルシャ語リソースのベクトルベースのセマンティック重複排除とドメインバランスのとれた事前トレーニング

ペルシア語の事前学習済み言語モデル (PLM) は、大規模で高品質な事前学習コーパスの不足と、標準的な分類や NER タスクを超える不十分な評価によって依然として制限されています。 IHUBERT は、Sepahr-Danesh コレクションの厳選された 45 GB のサブセット (約 7 ~ 80 億トークン) 上で RoBERTa ベースのエンコーダー (1 億 2500 万のパラメーター) を使用してゼロからトレーニングされた単一言語ペルシア語 PLM です。コーパスの品質を向上させ、冗長性を削減するために、正規化、正確な重複の削除、匿名化、ドメインとレジスタ間の分散バランス制御のためのベクトル データベース ベースのセマンティック重複排除を含む多段階の前処理パイプラインを採用しています。さらに、完全な事前トレーニング コーパスで 139,000 語彙の BPE トークナイザーをトレーニングして、ペルシャ語の形態学と正書法のバリエーションをより適切に捕捉します。 IHUBERT は、タスク標準の指標 (エンティティ レベルの F1、マクロ F1、EM/F1) を使用して、NER、センチメント分析、トピック分類、NLI、抽出的質問応答、関係抽出をカバーする 7 つのペルシャ NLU ベンチマークで評価されます。 IHUBERT は抽出 QA で最も強力な利益を達成し、PQuAD (F1 88.3542) と ParsiNLU-RC (F1 49.0987) の両方で 1 位にランクされ、FarsTail (Macro-F1 0.8350) で最高の結果を達成しました。 NER とトピック分類に関しては依然として競争力を維持しています (例: ParsTwiNER では 0.8308 F1、DigiMag では 0.7953 Macro-F1)。一方、関係抽出では依然として主なギャップが残っています (PERLEX では 0.6684 Macro-F1)。 IHUBERT 事前トレーニング コーパスの制御されたトークナイザー アブレーションでは、BPE が一致する語彙サイズで WordPiece よりもサブワードの断片化がわずかに低く、トークン化設計をサポートしていることが示されています。全体として、IHUBERT は、意味的に厳選された大規模な事前トレーニングと、分類と理解指向のタスクの両方にわたる広範な評価を通じて、ペルシア語モデリングを進歩させます。

原文 (English)

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insufficient evaluation beyond standard classification and NER tasks. We present IHUBERT, a monolingual Persian PLM trained from scratch with the RoBERTa-base encoder (125M parameters) on a 45 GB curated subset of the Sepahr-Danesh collection (about 7-8B tokens). To improve corpus quality and reduce redundancy, we employ a multi-stage preprocessing pipeline that includes normalization, exact and near-duplicate removal, anonymization, and vector-database-based semantic deduplication for distribution balancing control across domains and registers. We additionally train a 139k-vocabulary BPE tokenizer on the full pretraining corpus to better capture Persian morphology and orthographic variation. IHUBERT is evaluated on seven Persian NLU benchmarks covering NER, sentiment analysis, topic classification, NLI, extractive question answering, and relation extraction, using task-standard metrics (entity-level F1, Macro-F1, EM/F1). IHUBERT achieves its strongest gains on extractive QA, ranking first on both PQuAD (F1 88.3542) and ParsiNLU-RC (F1 49.0987), and attains the best result on FarsTail (Macro-F1 0.8350). On NER and topic classification, it remains competitive (e.g., 0.8308 F1 on ParsTwiNER; 0.7953 Macro-F1 on DigiMag), while relation extraction remains the main remaining gap (0.6684 Macro-F1 on PERLEX). A controlled tokenizer ablation on the IHUBERT pretraining corpus shows that BPE yields slightly lower subword fragmentation than WordPiece at matched vocabulary size, supporting our tokenization design. Overall, IHUBERT advances Persian language modeling through semantically curated large-scale pretraining and broad evaluation across both classification and comprehension-oriented tasks.

13:00 JST画像/動画生成

MakeupMirror: メイクアップ転写用の拡散モデルにおける顔属性の保存の改善

メイクアップ転送モデルにより、楽しい拡張現実 (AR) 体験やオンライン メイクアップ ショッピングの仮想試着 (VTO) が可能になります。 Stable-Makeup などの最近の最先端の拡散ベースのソリューションは、メイクアップ転写の精度とリアリズムを劇的に向上させていますが、アイデンティティと肌の色の保存には依然として限界があり、メイクアップ ショッピングのための実稼働レベルの VTO は非現実的です。この研究では、顔の特徴と肌の色合いの維持に向けて大きな進歩をもたらす、メイクアップ転写への拡散ベースのアプローチである MakeupMirror を提案します。私たちは、Stable-Makeup に対していくつかの技術革新を導入しています。(1) 顔の忠実度を維持するための、顔のジオメトリ コンディショニングと ControlNet の統合。 (2)皮膚、目、唇などの顔領域全体にわたる正確なメイクアップ適用を可能にする、領域固有のメイクアップ転写制御。 (3)被験者間の転写シナリオにおける肌の色調の変化を防ぐ、肌の色調に基づくメイクアップ転写調整。 (4) 生成品質を維持しながら推論を高速化するための Levenberg-Marquardt Langevin サンプラーの統合。 CPM-Real、Makeup Wild、および (ここでは新しく収集された、より多様な) MakeupSelfies データセットに関する実験では、MakeupMirror が Stable-Makeup と比較して、相対的な顔認識の類似性を +60% 改善し、相対的な肌の色合いの違いを -50% 削減し、0.7 秒のレイテンシで、中核となる顔の同一性保持基準全体で 94% の専門家の受け入れ率を達成していることが示されています。

原文 (English)

MakeupMirror: Improving Facial Attribute Preservation in Diffusion Models for Makeup Transfer

Makeup transfer models enable fun augmented reality (AR) experiences as well as virtual try-on (VTO) for online makeup shopping. While recent state-of-the-art diffusion based solutions such as Stable-Makeup dramatically improve the accuracy and realism of makeup transfer, they still face limitations in identity and skin color preservation, making production-level VTO for makeup shopping unrealistic. In this work, we propose MakeupMirror, a diffusion-based approach to makeup transfer that makes significant progress towards preserving facial features and skin tone. We introduce several technical innovations over Stable-Makeup: (1) integration of facial geometry conditioning with ControlNets to maintain facial fidelity; (2) region-specific makeup transfer control to enable precise makeup application across facial regions such as skin, eyes and lips; (3) skin tone-based makeup transfer modulation that prevent skin tone alteration in cross-subject transfer scenarios; and (4) integration of a Levenberg-Marquardt Langevin sampler to speed up inference while maintaining generation quality. Our experiments on CPM-Real, Makeup Wild, and (herein newly collected, more diverse) MakeupSelfies datasets show that MakeupMirror improves relative facial recognition similarity by +60%, reduces relative skin tone difference by -50% over Stable-Makeup, with a latency of 0.7s, while achieving expert acceptance rate of 94% across core facial identity preservation criteria.

13:00 JST研究/論文

整流フローによる命令ガイド付きオーディオ編集のためのハイブリッド拡散トランス

オーディオ編集の目的は、残りの音響コンテンツを維持しながら、自然言語命令に従って既存のオーディオ クリップ内の特定のコンテンツを変更することです。拡散モデルの目覚ましい進歩にも関わらず、既存のトレーニングベースの編集手法は主に、畳み込み U-Net バックボーンにおける局所的な帰納的バイアスとクロスアテンション相互作用に依存しており、長距離の意味論的整合や命令の正確な理解と位置特定を妨げることがよくあります。対照的に、拡散トランスフォーマーは、より強力なグローバル モデリングとマルチモーダル フュージョンを提供しますが、既存の編集アーキテクチャは通常、MMDiT ブロックと DiT ブロックの単純なスタックを採用しています。すべてのブロック内の連結されたオーディオ トークンとテキスト トークンに共同注意を適用すると、トークンの長さに関して 2 次の複雑さが生じます。編集パフォーマンスと効率のバランスをとるために、整流されたフローマッチングに基づいた命令ガイド付きオーディオ編集用のハイブリッド 2 ステージ拡散トランス アーキテクチャを提案します。音声トークンとテキスト トークンに対して共同アテンションを実行して、低解像度段階で大まかなセマンティック アライメントを確立し、その後、交互の共同アテンション ブロックとクロス アテンション ブロックに切り替えて、高解像度段階で編集の詳細を調整します。この粗いものから細かいものまでの戦略により、効率的かつ正確な指示に基づくオーディオ編集が可能になります。実験の結果、提案されたフレームワークは、コンパクトなモデルで編集効率を大幅に向上させながら、重複するオーディオ イベントや複雑な命令を含む困難な編集タスクで顕著なパフォーマンスの向上を達成することが示されています。

原文 (English)

Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow

Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of MMDiT and DiT blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a hybrid two-stage diffusion transformer architecture for instruction-guided audio editing based on rectified flow matching. It performs joint attention over audio and text tokens to establish coarse semantic alignment at low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency with a compact model.

13:00 JST研究/論文

感覚運動世界モデル: 逆ダイナミクスによる行動の知覚

行為に対する認識は、世界の表現は視覚的な忠実さだけによってではなく、行為との関連性によって形作られるべきであることを示唆しています。同時に、潜在的な JEPA スタイルの世界モデルは、将来の状態の予測を容易にするために、高次元の観測からコンパクトな予測状態を学習することを推奨していますが、予測が容易な潜在状態を構築することだけが目標である場合、表現が崩壊する可能性があるため、これらのモデルのエンドツーエンドのトレーニングは自明ではありません。感覚運動世界モデル (SMWM) を導入します。これは、逆ダイナミクス正則化を使用してエンドツーエンドでトレーニングされた潜在世界モデルです。この単一の正規化子は、表現の崩壊を防ぎ、アクションに合わせた表現を誘導するという両方の問題に対処します。潜在状態に遷移の基礎となるアクションに関する情報を保持させることにより、制御不能な注意をそらす要素を排除しながら、環境の制御可能な自由度に向けてモデルにバイアスをかけます。これにより、フリーズしたエンコーダー、指数移動平均、複雑な潜在正則化を使用せずに、オフラインの報酬のない軌道からトレーニングされた安定した潜在世界モデルが得られます。経験的に、SMWM はコンパクトで解釈可能な潜在空間を学習し、単純な 2D および 3D 制御タスク全体で競争力のある計画パフォーマンスを可能にします。

原文 (English)

Sensorimotor World Models: Perception for Action via Inverse Dynamics

Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions. At the same time, latent JEPA-style world models advocate learning compact predictive states from high-dimensional observations to facilitate the prediction of future states, but end-to-end training of these models is nontrivial because representations may collapse if our only goal is to construct a latent state that is easy to predict. We introduce a sensorimotor world model (SMWM): a latent world model trained end-to-end with inverse dynamics regularization. This single regularizer addresses both issues: it prevents representation collapse and induces action-aligned representations. By forcing latent states to preserve information about the action underlying a transition, it biases the model toward the controllable degrees of freedom of the environment while discarding uncontrollable distractors. This yields stable latent world models trained from offline, reward-free trajectories, without frozen encoders, exponential moving averages, or complex latent regularizers. Empirically, SMWM learns compact, interpretable latent spaces and enables competitive planning performance across simple 2D and 3D control tasks.

13:00 JSTエージェントロボティクス

自然言語プロトコルをロボット実験プラットフォームにクロスモデル検証済み翻訳するためのデュアルエージェント フレームワーク

生物学的実験プロトコルは自然言語で記述されていますが、自動化システムは事前定義された制御コマンドに依存しているため、自律的な実行を制限するセマンティック ギャップが生じています。マイクロプレートベースの自動実験は、ウェルマッピング、サンプルと試薬の組み合わせ、反復配置、並行分注を同時に制御する必要があるため、特に困難です。この研究では、自然言語のマイクロプレートベースのプロトコルをロボット実験室プラットフォーム用の実行可能な制御コマンドに変換する、エージェントベースのプロトコル変換フレームワークを提案します。パーサー エージェントは自然言語プロトコルを構造化表現に形式化し、ルールベースのマッピング エンジンはロボット ラボ プラットフォームの操作上の制約を決定論的に組み込んでデバイス レベルの制御コマンドを生成します。異種 LLM 検証エージェントは、完全性、パラメータの精度、実行順序を検証し、エラーが検出されると構造化されたフィードバックによる自己修正ループをトリガーします。ランダムに選択された ELISA プロトコル上で 7 つのパーサーと 3 つのバリデーターが関与するスイープにより、モデルのスケールとバリデーターのタイプがクロスモデル検証における翻訳精度と合格率にどのように影響するかを評価します。精度と遅延のトレードオフは、提案されたフレームワークのルールベースのマッピングと LLM エンドツーエンドの直接マッピングを比較することによってさらに検証されます。最後に、マイクロプレートを使用したブラッドフォード アッセイベースのタンパク質定量化がロボット実験室プラットフォームで実証され、自然言語プロトコルから現実世界の実験までエンドツーエンドの自律実行が検証されました。提案されたフレームワークは、自然言語プロトコルとマイクロプレートベースの自動運転ラボの間の意味論的なギャップを狭めるための柔軟なアプローチを提供します。

原文 (English)

Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform

Biological experiment protocols are written in natural language, whereas automation systems rely on predefined control commands, creating a semantic gap that limits autonomous execution. Microplate-based automatic experiments are particularly challenging due to the need to simultaneously control well mapping, sample-reagent combinations, replicate placement, and parallel dispensing. This study proposes an agent-based protocol translation framework that converts natural-language microplate-based protocols into executable control commands for a robotic laboratory platform. A Parser Agent formalizes the natural-language protocol into a structured representation, and a rule-based mapping engine deterministically incorporates the operational constraints of the robotic laboratory platform to generate device-level control commands. A heterogeneous LLM Validation Agent verifies completeness, parameter accuracy, and execution order, and triggers a self-correction loop with structured feedback when errors are detected. A sweep involving 7 Parsers and 3 Validators on randomly selected ELISA protocols evaluates how model scale and Validator type affect translation accuracy and pass rates under cross-model verification. The accuracy-latency trade-off is further verified by comparing the rule-based mapping of the proposed framework with LLM end-to-end direct mapping. Finally, Bradford assay-based protein quantification using a microplate was demonstrated on a robotic laboratory platform, validating end-to-end autonomous execution from natural-language protocols to real-world experiments. The proposed framework provides a flexible approach to narrowing the semantic gap between natural-language protocols and microplate-based self-driving laboratories.

13:00 JSTロボティクス

周波数を意識したフローマッチングにより、継続的かつ一貫したロボットアクションを生成

フロー マッチングは、拡散政策などの同様のアプローチと並んで、複雑でマルチモーダルなアクション分布をモデル化するための強力な表現力により、ロボット操作の標準パラダイムとして浮上しました。ただし、既存の方法は離散化されたアクションのチャンクに依存しているため、異種の制御周波数で収集されたデモンストレーションに対して脆弱であり、制御の安定性を低下させる時間的に一貫性のないアクションが発生する傾向があります。この論文では、継続的で時間的に一貫したアクションを出力する周波数認識フローマッチング (FAFM) を提案します。異種の周波数入力を処理するために、離散コサイン変換 (DCT) を使用して離散アクション シーケンスを周波数領域に変換し、結果の係数に対してフロー マッチングを実行し、コサイン基底拡張を介して連続アクションを再構築します。時間的に一貫したアクションを生成するには、一次時間導関数を正則化し、スムーズなアクションを促進します。これは、高頻度のエラーを抑制し、急激なアクションの変更を防ぐソボレフ タイプの制約に対応します。当社の FAFM はシンプルで、追加のネットワーク パラメータを導入せず、スタンドアロンのフロー マッチング ポリシーとビジョン言語アクション モデルに適用されます。合成玩具のベンチマーク、障害物回避、LapGym、および LIBERO にわたって、FAFM は成功率、マルチモーダル表現力、動きの滑らかさ、収束速度、機械的バイアスおよび混合周波数入力に対する堅牢性を向上させます。これらの利点は、現実世界の Franka ロボットに展開した場合でも一貫しています。コードは https://anonymous.4open.science/r/FAFM で入手できます。

原文 (English)

Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation

Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power for modelling complex, multimodal action distributions, alongside similar approaches like diffusion policy. However, existing methods rely on discretized action chunks, making them brittle to demonstrations collected at heterogeneous control frequencies and prone to temporally inconsistent actions that degrade control stability. In this paper, we propose Frequency-Aware Flow Matching (FAFM), which outputs continuous, temporally consistent actions. To handle heterogeneous frequency input, we transform discrete action sequences into the frequency domain with the discrete cosine transform (DCT), perform flow matching over the resulting coefficients, and reconstruct continuous actions via cosine basis expansion. To generate temporally consistent actions, we regularize the first-order temporal derivative to promote smooth actions. This corresponds to a Sobolev-type constraint that suppresses high-frequency errors and discourages abrupt action changes. Our FAFM is simple, introduces no additional network parameters and applies to standalone flow-matching policies and vision-language action models. Across synthetic toy benchmark, obstacle avoidance, LapGym, and LIBERO, FAFM improves success rates, multimodal expressivity, motion smoothness, convergence speed, robustness to mechanical bias and mixed-frequency input. These gains are consistent when deployed on a real-world Franka robot. Code available at https://anonymous.4open.science/r/FAFM.

13:00 JST研究/論文

局所的な可塑性を備えたハイブリッド ANN-SNN パイプライン

この研究では、事前学習された人工ニューラル ネットワーク (ANN) の豊富な埋め込みを効果的に活用して、高性能スパイキング ニューラル ネットワーク (SNN) を可能にする、ハイブリッド ANN-SNN パイプラインを提案します。このアーキテクチャでは、事前トレーニング済みの EfficientNet エンコーダーと CoLaNET スパイキング分類器が結合されています。レートコーディングを介してエンコーダーのアクティベーションをスパイクトレインに変換し、エンドツーエンドの勾配伝播をバイパスして、生物学にヒントを得たローカルな学習ルールを使用して後続の SNN 分類器をトレーニングします。このアプローチは、64 クラスの ImageNet ベンチマークで 99.09% の精度を達成し、従来のディープ ネットワークと同等のパフォーマンスを実証しました。この研究は、強力な事前学習済みエンコーダーを下流のスパイキング ニューラル ネットワーク タスクに適応させるための、生物学的に妥当で効率的なフレームワークを提示します。

原文 (English)

Hybrid ANN-SNN Pipeline with Local Plasticity

This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs) to enable high-performance spiking neural networks (SNNs). The architecture couples a pretrained EfficientNet encoder with a CoLaNET spiking classifier. We convert the encoder's activations into spike trains via rate-coding and train the subsequent SNN classifier using local, biologically inspired learning rules, bypassing end-to-end gradient propagation. This approach achieves 99.09% accuracy on a 64-class ImageNet benchmark, demonstrating performance on par with conventional deep networks. The work presents a biologically plausible and efficient framework for adapting powerful pretrained encoders to downstream spiking neural network tasks.

13:00 JSTLLM/生成AI

テキストからスコアへ: 大規模な言語モデルにおけるエッセイの品質表現の出現を追跡する

大規模言語モデル (LLM) の最近の進歩により、自動エッセイ採点 (AES) は大幅に変わりましたが、LLM ベースの採点の基礎となる内部メカニズムはまだよく理解されていません。この研究では、2 つの英語エッセイ データセット (ASAP++、CSEE) と 1 つのポルトガル語データセット (ENEM) にわたる 8 つの LLM の隠れた表現を体系的に分析します。線形探索、クロスプロンプト一般化、次元削減、およびニューロンレベルの分析を使用して、エッセイの品質情報が LLM 表現内で線形にアクセス可能な形式でエンコードされていることを示す一貫した証拠を発見しました。これらの表現は、レイヤー全体で徐々に現れ、プロンプト戦略全体で堅牢性を維持し、採点ルーブリックの違いにもかかわらず、部分的にエッセイプロンプト全体に移行します。さらに、非線形プローブは線形プローブに比べてわずかで一貫性のない改善しか提供しません。これは、ほとんどのエッセイの品質情報がすでに線形に解読可能であることを示唆しています。さらに、その活性化がエッセイのスコアと強く相関し、その行動が標的を絞った介入に敏感である個々の「エッセイスコアリングニューロン」を特定します。さらに、これらのニューロンの層ごとの分布はエッセイの長さに応じて系統的に変化し、長いエッセイほどより深い層に大きく依存します。全体として、私たちの調査結果は、LLM がエッセイの品質に関連する構造化表現をエンコードしているという証拠を提供し、LLM ベースの AES システムの解釈可能性についての新たな洞察を提供します。

原文 (English)

From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models

Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms underlying LLM-based scoring remain poorly understood. In this work, we systematically analyze the hidden representations of eight LLMs across two English essay datasets (ASAP++, CSEE) and one Portuguese dataset (ENEM). Using linear probing, cross-prompt generalization, dimensionality reduction, and neuron-level analyses, we find consistent evidence that essay quality information is encoded in a linearly accessible form within LLM representations. These representations emerge progressively across layers, remain robust across prompting strategies, and partially transfer across essay prompts despite differences in scoring rubrics. In addition, nonlinear probes provide only marginal and inconsistent improvements over linear probes, suggesting that most essay quality information is already linearly decodable. We further identify individual ``essay scoring neurons'' whose activations strongly correlate with essay scores and whose behavior is sensitive to targeted intervention. Moreover, the layer-wise distribution of these neurons systematically shifts with essay length, with longer essays relying more heavily on deeper layers. Overall, our findings provide evidence that LLMs encode structured representations related to essay quality and offer new insights into the interpretability of LLM-based AES systems.

13:00 JSTLLM/生成AI

MedRLM: ロングコンテキストの臨床推論、センサーに基づくスクリーニング、証拠に基づく意思決定サポート、コミュニティから三次への紹介の最適化のための再帰的マルチモーダルヘルスインテリジェンス

現実世界の臨床意思決定のサポートでは、個別の医療上の質問に答えるのではなく、異質かつ長期的な患者情報を基に推論する必要があります。しかし、現在の医療用大規模言語モデルと検索拡張生成システムは、シングルステップのプロンプトまたは検索に依存していることが多く、臨床証拠が長い電子医療記録、医療画像、センサー ストリーム、ガイドライン、参照制約にまたがって分散されている場合、脆弱になる可能性があります。この論文では、ロングコンテキストの臨床推論、センサーに基づくスクリーニング、およびコミュニティから三次への紹介サポートのための再帰的マルチモーダル ヘルス インテリジェンス フレームワークである MedRLM を提案します。 MedRLM は、すべての患者情報を 1 つのプロンプトに圧縮するのではなく、再帰的に検査、分解、取得、検証、合成できる外部の臨床環境として患者の症例を扱います。このフレームワークは、臨床テキスト、縦断的 EHR、医療画像、生理学的センサー信号、ガイドライン検索、不確実性監査、紹介計画の専門エージェントを調整します。さらに、患者固有の観察結果を、取得した証拠、標準化された定義、センサー由来のバイオマーカー、参照基準と結び付ける臨床証拠グラフ メモリも導入されています。異常な生理学的パターンまたは行動パターンが検出された場合、センサー誘導の再帰的トリガーメカニズムにより、より深い推論が活性化され、不確実性ゲートによる洗練により、高リスクまたは低信頼性の症例に対する臨床医のレビューがサポートされます。また、EHR、放射線医学、ECG、ICU 時系列、紹介代理結果にわたる公的で認定された臨床データセットを使用した実データ評価設計の概要も説明します。 MedRLM は、医療 AI を静的な質問応答から、監査可能でマルチモーダルでワークフローを意識した臨床意思決定サポートに移行することを目指しています。

原文 (English)

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

Real-world clinical decision support requires reasoning over heterogeneous and longitudinal patient information rather than answering isolated medical questions. However, current medical large language models and retrieval-augmented generation systems often rely on single-step prompting or retrieval, which can be fragile when clinical evidence is distributed across long electronic health records, medical images, sensor streams, guidelines, and referral constraints. This paper proposes MedRLM, a Recursive Multimodal Health Intelligence framework for long-context clinical reasoning, sensor-guided screening, and community-to-tertiary referral support. Instead of compressing all patient information into one prompt, MedRLM treats the patient case as an external clinical environment that can be recursively inspected, decomposed, retrieved, verified, and synthesized. The framework coordinates specialized agents for clinical text, longitudinal EHR, medical imaging, physiological sensor signals, guideline retrieval, uncertainty auditing, and referral planning. It further introduces a Clinical Evidence Graph Memory to connect patient-specific observations with retrieved evidence, standardized definitions, sensor-derived biomarkers, and referral criteria. A sensor-guided recursive triggering mechanism activates deeper reasoning when abnormal physiological or behavioral patterns are detected, while uncertainty-gated refinement supports clinician review for high-risk or low-confidence cases. We also outline a real-data evaluation design using public and credentialed clinical datasets spanning EHR, radiology, ECG, ICU time series, and referral-proxy outcomes. MedRLM aims to move medical AI from static question answering toward auditable, multimodal, and workflow-aware clinical decision support.

13:00 JSTLLM/生成AI画像/動画生成

リモート センシング MLLM における否定の理解の評価と強化

マルチモーダル大規模言語モデル (MLLM) は、さまざまなリモート センシング (RS) タスクで目覚ましい成功を収めています。ただし、否定を理解する能力は依然として研究されていないため、モデルが何が虚偽であるか、何が欠落しているかを明示的に識別する必要がある現実世界のアプリケーションへの展開が制限されます。たとえば、緊急対応者は避難のために浸水していないルートを見つける必要があります。この制限を包括的に研究するために、領域レベルからシーンレベルのタスクにわたる否定の理解を評価する最初のベンチマークである RS-Neg を導入します。具体的には、LLM を使用してさまざまな否定クエリを合成し、検証用の動的ビジュアル フォーカス モジュールを導入して、RS 画像用の自動データ生成パイプラインを設計します。私たちの評価では、高度な RS MLLM は否定に苦しみ、幻覚や大幅なパフォーマンスの低下を示すことが明らかになりました。このギャップを埋めるために、モデルの最適化に否定の論理的役割を明示的に組み込む新しいテスト時学習方法である NeFo を提案します。驚くべきことに、約 5\% のラベルなしテスト サンプルを使用することで、NeFo はモデルの否定の理解を大幅に向上させ、目に見えないタスクに対する強力な一般化を示します。コードとデータは承認され次第公開されます。

原文 (English)

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to comprehend negation remains underexplored, limiting deployment in real-world applications where models must explicitly identify what is false or absent, e.g., emergency responders need to locate non-flooded routes for evacuation. To comprehensively study this limitation, we introduce RS-Neg, the first benchmark to evaluate negation understanding across region-level to scene-level tasks. Specifically, we design an automated data generation pipeline for RS imagery, using LLMs to synthesize diverse negation queries, and introduce a dynamic visual focus module for verification. Our evaluation reveals that advanced RS MLLMs struggle with negation, exhibiting hallucinations and substantial performance degradation. To close this gap, we propose NeFo, a novel test-time learning method that explicitly incorporates the logical role of negation into the model optimization. Remarkably, using about 5\% unlabeled test samples, NeFo significantly improves the negation understanding of models and shows strong generalization to unseen tasks. Code and data will be released upon acceptance.

13:00 JST画像/動画生成ロボティクス

HilDA: 自己監視型 LiDAR の事前トレーニングを促進するための拡散を使用した階層的蒸留

カメラから LiDAR への知識の蒸留に Vision Foundation Models (VFM) を活用することは、現実世界の自動運転 (AD) の膨大な幾何学的および運動学的多様性を表現するために必要な注釈付きデータの不足に対する有望な解決策を提供します。ただし、現在のアプローチは通常、VFM をブラックボックス教師として扱い、フレーム単位の特徴の類似性にのみ依存します。その結果、教師のレイヤーごとの意味構造とグローバル コンテキスト、さらには LiDAR シーケンスに固有の豊富な時空間情報が十分に活用されません。私たちは、運転タスクに必要なセマンティックな内容と幾何学的な場所をより適切に捕捉する、LiDAR バックボーン用の自己監視型事前トレーニング フレームワークである HilDA を提案します。 HilDA は、段階的なセマンティクスの調整のための多層蒸留と、シーンレベルのセマンティクスのためのグローバル コンテキストの蒸留を含む階層的蒸留を、時空間的一貫性を促進する時間占有拡散目標と組み合わせます。 HilDA で事前トレーニングされたモデルは、クロスモーダル蒸留ベンチマークで最先端の結果を達成し、3D オブジェクト検出、シーン フロー、セマンティック占有予測に関して事前の蒸留アプローチでトレーニングされたモデルよりも優れたパフォーマンスを発揮します。コードは https://maxiuw.github.io/hilda で入手できます。

原文 (English)

HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-trainin

Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD). However, current approaches typically treat VFMs as black-box teachers, relying exclusively on frame-wise feature similarity. Consequently, they do not fully exploit the teacher's layer-wise semantic structure and global context, as well as the rich spatiotemporal information inherent in LiDAR sequences. We propose HilDA, a self-supervised pretraining framework for LiDAR backbones that better captures the semantic what and geometric where needed for driving tasks. HilDA combines hierarchical distillation comprising multi-layer distillation for progressive semantic alignment and global context distillation for scene-level semantics, with a temporal occupancy diffusion objective promoting spatiotemporal consistency. Models pre-trained with HilDA achieve state-of-the-art results on cross-modal distillation benchmarks and outperform models trained via prior distillation approaches on 3D object detection, scene flow, and semantic occupancy prediction. Code available at: https://maxiuw.github.io/hilda.

13:00 JSTロボティクス

FlowMaps: フローマッチングを使用した長期マルチモーダルオブジェクトダイナミクスのモデル化

3D シーンの空間的および時間的共同理解は、日常の家庭環境に導入されるロボットにとって重要な要件です。このようなエージェントは、空間レイアウトを理解してナビゲートするだけでなく、これらの空間が時間の経過とともにどのように進化するかを推論する必要があります。特に、人間は毎日物体と対話するため、環境全体で物体の位置が変化し、ロボットが現在の観察を以前に見た物体と確実に関連付けることが困難になります。ただし、これらの相互作用はランダムではありません。人間の習慣やルーチンは、オブジェクトの位置に時空間的に一貫したパターンを引き起こし、ロボットエージェントがそれを学習して、ナビゲーションなどの下流のタスクに利用できる可能性があります。この目的を達成するために、連続 3D 空間内の動的オブジェクトの将来の位置にわたるマルチモーダル分布を推定するための潜在フロー マッチング モデルである FlowMaps を導入します。 FlowMaps は、オブジェクト間の暗黙の依存関係とその時間的進化を学習することで、過去の人間のインタラクションを条件としたオブジェクトの位置の変化を予測し、同様のオブジェクト ルーチンを共有するこれまで見たことのない環境全体にわたる一般化をサポートします。このメソッドの有用性を実証するために、シミュレーション環境と現実世界の両方の環境で、ダウンストリームの動的オブジェクト ナビゲーション タスクに FlowMaps をデプロイします。 600 を超えるエピソードにわたって、FlowMaps は最先端のアプローチを上回っており、連続的でマルチモーダルな時空間分布を通じてオブジェクトのダイナミクスをモデリングすることで、変化する家庭環境におけるロボットの検索とナビゲーションが向上することが示されています。コードと追加資料は https://fra-tsuna.github.io/flowmaps/ で入手できます。

原文 (English)

FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching

Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments. Such agents must not only comprehend and navigate spatial layouts, but also reason about how these spaces evolve over time. In particular, humans interact with objects daily, causing them to change position throughout the environment and making it difficult for robots to reliably associate current observations with previously seen objects. However, these interactions are not random: human habits and routines induce spatio-temporally consistent patterns in object locations, which robotic agents can potentially learn and then exploit for downstream tasks such as navigation. To this end, we introduce FlowMaps, a latent flow matching model for estimating multimodal distributions over the future locations of dynamic objects in a continuous 3D space. By learning the implicit dependencies among objects and their temporal evolution, FlowMaps predicts likely changes in object locations conditioned on past human interactions, while supporting generalization across previously unseen environments that share similar object routines. To demonstrate the utility of this method, we deploy FlowMaps in a downstream dynamic Object Navigation task in both simulated and real-world environments. Across more than 600 episodes, FlowMaps outperforms state-of-the-art approaches, showing that modeling object dynamics through continuous, multimodal spatio-temporal distributions improves robotic search and navigation in changing household environments. Code and additional material is available at https://fra-tsuna.github.io/flowmaps/.

13:00 JSTビジネス/資金調達

学習者ベースのコンセプトドリフト検出: 分析と評価

進化するストリーミング環境に導入された機械学習アルゴリズムは、一般に概念ドリフトと呼ばれる非定常データ分布を処理する必要があります。概念ドリフトの存在は、予測パフォーマンスを大幅に低下させ、堅牢な意思決定をサポートする能力を妨げる可能性があるため、多くの実世界のアプリケーションにとって大きな課題となります。したがって、長期にわたり高い精度を維持するには、ドリフト イベントをタイムリーかつ効率的に検出することが重要です。この研究では、概念のドリフト特性と、いくつかのカテゴリにわたる多数のドリフト検出アルゴリズムを理論的に検証します。さらに、多様なストリーミング シナリオや急激な変化や段階的な変化などのドリフト特性を示す合成データセットと現実世界のデータセットの両方でパフォーマンスを評価します。この研究は、概念ドリフト特性とドリフト検出器の動作の複雑な概念と、それらの多様な状況への適用性についての理解を高めることを目的としています。

原文 (English)

Learner-based Concept Drift Detection: Analysis and Evaluation

Machine learning algorithms deployed for evolving streaming environments must handle the non-stationary data distributions, commonly referred to as concept drift. The presence of concept drift poses a major challenge for many real-world applications because it can severely degrade their predictive performance, hindering their ability to support robust decision-making. Consequently, the timely and efficient detection of drift events is critical for sustaining high accuracy over time. This study examines theoretically the concept drift characteristics and numerous drift detection algorithms across several categories. Furthermore, we evaluate their performance on both synthetic and real-world datasets exhibiting diverse streaming scenarios and drift characteristics, such as abrupt and gradual changes. This study aims to enhance understanding of the complex notion of concept drift characteristics and behavior of drift detectors, along with their applicability to diverse contexts.

13:00 JSTLLM/生成AIエージェント研究/論文

ScholarQuest: オープンな文献環境におけるエージェント的学術論文検索のための分類法に基づくベンチマーク

学術論文の検索は科学研究の中核的なステップであり、LLM ベースの検索エージェントは、反復的で意図に基づいた文献探索のための有望なパラダイムとして台頭しています。しかし、既存のベンチマークは、現実的なオープン文献環境下でエージェントによる学術検索を系統的に評価するには不十分です。私たちは、エージェントによる学術論文検索のための大規模な分類に基づくベンチマークである ScholarQuest を提案します。 ScholarQuest は、1,000 を超えるコンピューター サイエンスのトピックと、方法指向、設定に固定されたクエリ、比較ベースのクエリ、およびスコープ制御のクエリを含む 4 つの代表的な研究目的から構成されています。さらに、スケーラブルな回答構築と、再現可能な評価のための共有取得バックエンド ScholarBase を提供します。ベンチマーク結果は、エージェント手法がシングルショット検索のベースラインを上回るパフォーマンスを示しているものの、最もパフォーマンスの高いエージェントでも 0.314 Recall@100 と 0.355 Recall@All しか達成しておらず、改善の余地が大きいことが示されています。さらに、検索効率、意図レベルの堅牢性、失敗例の分析により、学術論文検索エージェントに多次元の評価シグナルを提供するベンチマークの能力がさらに強調されます。

原文 (English)

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

Academic paper search is a core step in scientific research, and LLM-based search agents are emerging as a promising paradigm for iterative, intent-driven literature exploration. However, existing benchmarks are insufficient for systematically evaluating agentic academic search under realistic open literature environments. We propose ScholarQuest, a large-scale, taxonomy-guided benchmark for agentic academic paper search. ScholarQuest is constructed from over 1,000 computer science topics and four representative research intents, including method-oriented, setting-anchored, comparison-based, and scope-controlled queries. It further provides scalable answer construction and a shared retrieval backend ScholarBase for reproducible evaluation. Benchmarking results show that agentic methods outperform single-shot retrieval baselines, yet the best-performing agent only achieves 0.314 Recall@100 and 0.355 Recall@All, indicating substantial room for improvement. In addition, analyses of search efficiency, intent-level robustness, and failure cases further highlight the benchmark's ability to provide multi-dimensional evaluation signals for academic paper search agents.

13:00 JST画像/動画生成

SPOT-E: 凍結された VLM のビジュアル スポットライトを使用したテスト時のエントロピー シェーピング

視覚言語モデル (VLM) は、決定的な視覚的証拠が小さく、局所的で見落とされやすいため、証拠を重視するタスクではパフォーマンスが低下することが多く、たとえ高度な推論が損なわれていない場合でも、証拠の読み出しに失敗する可能性があります。事前の推論時の視覚的介入は再トレーニングなしでグラウンディングを改善できますが、それらは主にオープンループであり、強調表示された証拠が実際に使用されているかどうかを検証するメカニズムがありません。我々は、モデル内部のフィードバック信号としての応答スパン予測エントロピーを研究し、証拠に基づく信頼性やショートカットの崩壊から低エントロピーが生じる可能性があるため、単純なエントロピー最小化があいまいであることを示します。このあいまいさを解決するために、ベースラインの高信頼トークンを維持しながら、回答の不確実性を低減する低エントロピー アンカーとエントロピー シェーピング目標を導入します。この原則は、質問条件付きスポットライトを生成するプラグアンドプレイのテスト時間メソッドである SPOT-E でインスタンス化され、グループ相対ポリシー最適化 (GRPO) に基づく軽量チューニングによってインスタンスごとに最適化されます。すべてのベンチマークとさまざまな VLM ファミリにわたって、SPOT-E は一貫した利益をもたらし、視覚的な破損の下でも堅牢性が向上しました。コードは次の場所で公開されています: \url{https://github.com/yingBo0927/SPOT-E}

原文 (English)

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evidence readout even when high-level reasoning is intact. Prior inference-time visual interventions can improve grounding without retraining, but they are largely open-loop and lack a mechanism to verify whether highlighted evidence is actually used. We study answer-span prediction entropy as a model-internal feedback signal and show that naive entropy minimization is ambiguous, since low entropy may arise from evidence-grounded confidence or shortcut collapse. To resolve this ambiguity, we introduce low-entropy anchors and an entropy-shaping objective that reduces answer uncertainty while preserving baseline high-confidence tokens. We instantiate this principle in SPOT-E, a plug-and-play test-time method that produces question-conditioned spotlights, optimized per instance via light-weight tuning based on Group Relative Policy Optimization (GRPO). Across all benchmarks and different VLM families, SPOT-E yields consistent gains and improved robustness under visual corruptions. Code is publicly available at: \url{https://github.com/YinBo0927/SPOT-E}

13:00 JSTロボティクス

視覚-言語-行動モデルの微調整に必要なレイヤーは思ったよりも少ない

大規模なビデオ ロボット データセットで事前トレーニングされた Vision-Language-Action (VLA) モデルは、ロボット操作に革命をもたらしましたが、数十億のパラメーター アーキテクチャにより、下流の微調整やリアルタイム推論中に法外な計算負荷が課せられます。この研究では、これらの連続制御基盤ポリシー (pi_0、GR00T-N1.5 など) の非常に重要なアーキテクチャ上の特徴を明らかにします。つまり、多様な物理軌道でトレーニングされているにもかかわらず、レイヤーごとの表現の重大な冗長性を示します。これを利用するために、完全にトレーニング不要の構造圧縮パイプラインを導入し、最適化されたトークン削減や動的レイヤー セレクターを学習するためにフルスケールのモデルをロードする既存の方法の必要性を回避します。代わりに、Centered Kernel Alignment による単一のフォワード パスのみを使用して冗長レイヤー機能を特定し、ツインレイヤーを削除して、VLM バックボーンと継続的制御ポリシー ヘッドの両方でモデルの深さを最大 50% 永続的に圧縮します。この合理化されたアーキテクチャの下流での微調整により、フルスケールのベース モデルのパフォーマンスと同等またはそれを超えながら、トレーニング時間の 40 ~ 50% の削減とリアルタイム推論の最大 30% 高速化という二重の高速化のメリットがもたらされます。私たちは、3 つのシミュレーション ベンチマーク (LIBERO、RoboCasa、SimplerEnv) と、4 つのユニークなロボットの実施形態にわたる 10 の多様な現実世界の操作タスクにわたって、メソッドを包括的に検証します。これらの結果は、高度な VLA に必要な層が以前の想定よりも大幅に少なく、スケーラブルなロボット学習のための計算効率の高いパラダイムを提供することを証明しています。

原文 (English)

Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose prohibitive computational burdens during downstream fine-tuning and real-time inference. In this work, we reveal a highly non-trivial architectural characteristic of these continuous control foundation policies (e.g., pi_0, GR00T-N1.5): despite being trained on diverse physical trajectories, they exhibit severe layer-wise representational redundancy. To exploit this, we introduce a structural compression pipeline that is entirely training-free, bypassing the need of existing methods to load full-scale models to learn optimized token reductions or dynamic layer selectors. Instead, using only a single forward pass via Centered Kernel Alignment to identify redundant layer features, we remove twin layers to permanently compress the model depth by up to 50% across both the VLM backbone and the continuous control policy head. Downstream fine-tuning of this streamlined architecture yields a dual acceleration benefit: a 40-50% reduction in training time and up to 30% faster real-time inference, while matching or exceeding full-scale base model performance. We comprehensively validate our method across three simulation benchmarks (LIBERO, RoboCasa, SimplerEnv) and 10 diverse real-world manipulation tasks across 4 unique robotic embodiments. These results prove that advanced VLAs require significantly fewer layers than previously assumed, offering a highly compute-efficient paradigm for scalable robot learning.

13:00 JSTLLM/生成AIビジネス/資金調達Gemini

レジスターギャップ: ナイジェリアの公共言説のための意味インテリジェンスのフレームワーク

私たちは、表面的な感情を真のコミュニケーション意図から分離する、ナイジェリアの公共の議論のための 9 次元の注釈および評価スキーマであるミーニング インテリジェンス フレームワーク (MIF) を紹介します。 NaijaSenti や AfriSenti などのナイジェリア言語の既存のベンチマークは、感情分類を 3 方向の極性タスク (肯定的、否定的、中立的) として扱います。私たちは、ナイジェリアの談話における AI システムの主な失敗モードは、翻訳の失敗ではなく、文脈の失敗であると主張します。つまり、同じ発話が、話者、聴衆、状況に応じて反対の実用的な力をもたらします。 MIF は、記録、表面的な感情、真の意図、皮肉、コード化されたサブテキスト、リスク層、アノテーターの信頼度、話者の感情、推奨されるコミュニケーション アクションという 9 つのスコア化された次元にわたってこの洞察を運用します。標準英語、ナイジェリア英語、ナイジェリア ピジン、およびコード混合レジスタにわたる 30 項目のキャリブレーション データセットを構築し、ゼロショットおよびスキーマ情報に基づいたプロンプト条件下でフロンティア言語モデル (Gemini 2.5 フラッシュ) を評価します。見出しの結果はレジスタ ギャップです。ゼロショット レジスタの分類精度は 33.3% で、モデルがコンテキスト内で MIF スキーマを受け取ると 73.3% (+40 ポイント) に上昇しました。スキーマ情報に基づくプロンプトの下では、複合意味インテリジェンス スコアが 5.4 ポイント (73.2 から 78.6) 増加し、レジスターの識別、コード化されたサブテキストの検出 (+10 ポイント)、および戦略的アクションの推奨 (+10.3 ポイント) において実質的な増加が最も大きくなります。再現性をサポートするために、フレームワーク仕様、注釈ガイドライン、および 30 項目のパブリック キャリブレーション セットをリリースすると同時に、汚染から保護された評価のためのプライベート ホールドアウト コーパスを保持します。

原文 (English)

The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse

We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent. Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti, treat sentiment classification as a three-way polarity task (positive, negative, neutral). We argue that the dominant failure mode of AI systems on Nigerian discourse is not translation failure but context failure: the same utterance carries opposite pragmatic force depending on speaker, audience, and situation. The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent, irony, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. We construct a 30-item calibration dataset spanning Standard English, Nigerian English, Nigerian Pidgin, and code-mixed registers, and evaluate a frontier language model (Gemini 2.5 Flash) under zero-shot and schema-informed prompting conditions. The headline finding is the Register Gap: zero-shot register classification accuracy is 33.3%, rising to 73.3% (+40 points) when the model receives the MIF schema in-context. The composite Meaning Intelligence Score increases by 5.4 points (73.2 to 78.6) under schema-informed prompting, with the largest practical gains in register identification, coded-subtext detection (+10 points), and strategic action recommendation (+10.3 points). We release the framework specification, annotation guidelines, and the 30-item public calibration set to support reproducibility, while retaining a private holdout corpus for contamination-protected evaluation.

13:00 JSTLLM/生成AI

編集上の調整: LLM を介した知識の普及に編集の専門知識を活用するための参加型アプローチ

LLM 主導の情報サービスの出現により、公的知識機関が運営される条件が再構築され、これらの機関が果たすべき編集機能が吸収される恐れがあります。 LLM は知識の普及のための強力な新しいアフォーダンスを提供しますが、商用開発者の価値観と普及戦略にすでに合致した事前トレーニング済みの LLM によって編集権限が挑戦されます。この論文では、北欧の公的知識機関と共同で LLM 対応の百科事典インターフェイスを設計および実装したケーススタディで、設計ワークショップを通じて LLM インターフェイスを編集基準に再調整する際の編集者の参加について調査します。私たちは、参加型 AI 内のデザイン実践として編集調整を導入し、AI アライメントをデザイン プロセスとして枠組み化し、編集標準を編集実践と価値観を技術実装のための調整目標に変換するデザイン成果物として位置づけます。最後に、編集上の調整によってどのように継続的な参加の余地が生まれ、LLM を介した知識の普及に編集者に主体性を与えることができるかについて説明します。

原文 (English)

Editorial Alignment: A Participatory Approach to Engaging Editorial Expertise in LLM-mediated Knowledge Dissemination

The emergence of LLM-driven information services is reshaping the conditions under which public knowledge institutions operate, threatening to absorb the editorial function these institutions exist to exercise. While LLMs offer powerful new affordances for knowledge dissemination, editorial authority is challenged by pretrained LLMs that arrive already aligned with the values and dissemination strategies of their commercial developers. This paper investigates editor participation in re-aligning LLM interfaces to editorial standards through design workshops, in a case study where we design and implement an LLM-enabled encyclopedia interface with a Nordic public knowledge institution. We introduce editorial alignment as a design practice within Participatory AI, framing AI alignment as a design process and positioning the editorial standard as a design artefact that translates editorial practice and values into alignment objectives for technical implementation. Last, we discuss how editorial alignment can create space for ongoing participation and give editors agency in LLM-mediated knowledge dissemination.

13:00 JSTLLM/生成AI

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval

Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance…

13:00 JST研究/論文

Boundary Embedding Shaping with Adaptive Contrastive Learning for Graph Structural Disentanglement

Graph neural networks (GNNs) excel at aggregating neighbor information for classification, yet their performance is hindered by graph struc…

13:00 JST研究/論文

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise

In this article, we present a robust $Q$-learning algorithm for discrete-time mean-field control problems under Wasserstein uncertainty in…

13:00 JSTLLM/生成AIエージェント

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Large Language Models (LLMs) show promise for code compilation tasks, but applying them to runtime performance tuning is difficult due to c…

13:00 JSTエージェントロボティクス研究/論文

CRAX: Fast Safe Reinforcement Learning Benchmarking

Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. Wh…

13:00 JST研究/論文

DataMagic: Transforming Tabular Data into Data Insight Video

Data videos integrate dynamic charts, voice narration, and synchronized animations to communicate data insights as temporal narratives, mak…

13:00 JSTLLM/生成AIエージェント研究/論文

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness und…

13:00 JSTLLM/生成AI

Multi-View Decompilation for LLM-Based Malware Classification

Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that la…

13:00 JST研究/論文

Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

Classifier guidance is a way to control diffusion generation by using a noise-conditioned classifier to steer the sampling process toward a…

13:00 JSTエージェント

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coord…

13:00 JSTエージェント

UltraQuant: 4-bit KV Caching for Context-Heavy Agents

Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrenc…

13:00 JSTエージェント研究/論文

Optimal Order of Multi-Agent and General Many-Body Systems

This paper develops a general framework for analyzing multi-agent systems with feedback loops between agents actions and collective observa…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達DeepSeek

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

When large language models serve as evaluators in multi-agent systems, their systematic evaluation biases propagate through the agent netwo…

13:00 JSTLLM/生成AI研究/論文DeepSeek

Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software

Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains…

13:00 JST画像/動画生成

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while…

13:00 JSTエージェント

Efficient and Sound Probabilistic Verification for AI Agents

Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulat…

13:00 JSTエージェント

Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not…

13:00 JST画像/動画生成研究/論文

SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture rada…

13:00 JST研究/論文

Structuring and Tokenizing Distributed User Interest Context for Generative Recommendation

Generative recommendation is an emerging paradigm that has shown promise in industrial recommendation systems, aiming to predict users' nex…

13:00 JSTLLM/生成AIGemma

How Transparent is DiffusionGemma?

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging su…

13:00 JSTエージェントロボティクス

UniMM: A Unified Mixture Model Framework for Multi-Agent Simulation

Simulation plays a crucial role in assessing autonomous driving systems, where the generation of realistic multi-agent behaviors is a key a…

13:00 JSTエージェント研究/論文

MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning

Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. De…

13:00 JST研究/論文

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

Post-training alignment of large language models often combines supervised fine-tuning (SFT) on expert demonstrations with reinforcement le…

13:00 JST研究/論文

Controlled Comparison of Machine Learning Models for Fault Classification and Localization in Power System Protection

The increasing complexity of modern power systems, driven by the integration of inverter-based and distributed energy resources, challenges…

13:00 JSTLLM/生成AIエージェント

SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning

Solving mathematical reasoning problems requires not only accurate access to relevant knowledge but also careful, multi-step thinking. Howe…

13:00 JST研究/論文

Creativity Reconsidered: Generative AI and the Problem of Intentional Agency

Many theorists maintain that conscious intentional agency is a necessary condition of creativity. We argue that this requirement, which we…

13:00 JSTLLM/生成AI研究/論文Gemma

PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Boards (PCB) Schematic Design with Structured Verification

Most LLM code-synthesis benchmarks rely on unit tests as the reward oracle, but PCB schematic design has none: correctness is defined by st…

13:00 JST研究/論文

One Probe Won't Catch Them All: Towards Targeted Deception Detection

Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier…

13:00 JST研究/論文

Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach

We study conditional generation in diffusion models under hard constraints, where generated samples must satisfy prescribed events with pro…

13:00 JST研究/論文

SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures

While the shift toward unified foundation models has revolutionized many deep learning domains, sleep medicine remains largely restricted t…

13:00 JSTハードウェア/半導体

Mitigating Legibility Tax with Decoupled Prover-Verifier Games

As large language models become increasingly capable, it is critical that their outputs can be easily checked by less capable systems. Prov…

13:00 JST研究/論文

PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units

Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typic…

13:00 JSTLLM/生成AIビジネス/資金調達

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evalua…

13:00 JST研究/論文

CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions

Incomplete or inconsistent discharge documentation drives care fragmentation and avoidable readmissions. Despite its critical role in patie…

13:00 JST研究/論文

Too long; didn't solve

Mathematical benchmarks consisting of a range of mathematics problems are widely used to evaluate the reasoning abilities of large language…

13:00 JSTLLM/生成AIエージェント

CogniFold: コグニティブフォールディングによる常時オンのプロアクティブなメモリ

既存のエージェントの記憶は主に反応的かつ検索ベースのままであり、経験を自律的に永続的な認知構造に組織化する能力が欠けています。真の自律型エージェントを目指して、次世代のプロアクティブ アシスタント向けに設計された、脳からインスピレーションを得た「常時オン」エージェント メモリである CogniFold を紹介します。 CogniFold は、断片化されたイベント ストリームを自己出現の認知構造に継続的に折り畳んで、入ってくるイベントと蓄積された知識から徐々により高いレベルの認知をブートストラップします。私たちは、相補学習システム (CLS) 理論を 2 層 (海馬、新皮質) から 3 層に拡張し、前頭前意図層を追加することでこれを根拠にしています。意図的な制御と意思決定の拠点として前頭前野をエミュレートする CogniFold は、グラフ トポロジーの自己組織化を通じてこれを実現します。つまり、認知構造はストリームの下で積極的に集まり、意味的に類似している場合は結合し、古くなっている場合は減衰し、連想想起を通じて再リンクし、概念クラスターの密度がしきい値を超えると意図を表面化します。 CogEval-Bench を使用して構造形成を評価し、CogniFold が認知的期待と概念創発に一致する記憶構造を独自に生成することを実証します。さらに、5 つの認知ドメインにわたる 7 つの広範なベンチマークにわたって、CogniFold が従来のメモリ ベンチマークでも同時に堅牢に実行されることを検証しました。

原文 (English)

CogniFold: Always-On Proactive Memory via Cognitive Folding

Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genuinely autonomous agents, we introduce CogniFold, a brain-inspired "always-on" agent memory designed for the next generation of proactive assistants. CogniFold continuously folds fragmented event streams into self-emerging cognitive structures, bootstrapping progressively higher-level cognition from incoming events and accumulated knowledge. We ground this by extending Complementary Learning Systems (CLS) theory from two layers (hippocampus, neocortex) to three, adding a prefrontal intent layer. Emulating the prefrontal cortex as the locus of intentional control and decision-making, CogniFold achieves this through graph-topology self-organization: cognitive structures proactively assemble under the stream, merge when semantically similar, decay when stale, relink through associative recall, and surface intents when concept-cluster density crosses a threshold. We evaluate structural formation using CogEval-Bench, demonstrating that CogniFold uniquely produces memory structures that match cognitive expectations and concept emergence. Furthermore, across eight downstream benchmarks -- two probing long-term conversational memory (LoCoMo, LongMemEval) and six spanning other cognitive domains -- we validate that CogniFold simultaneously performs robustly on conventional memory tasks. Our code is available at https://github.com/OpenNorve/CogniFold.

13:00 JSTエージェントビジネス/資金調達

SimuWoB: 高速かつ忠実な GUI エージェント ベンチマークのための現実世界のモバイル アプリのシミュレーション

大規模な言語モデルを利用したモバイル GUI エージェントは急速に進歩しており、現実的かつ包括的な評価に対する緊急のニーズが生じています。既存のベンチマークは再現性を優先していますが、実際のアプリケーションで報酬を構築することが難しいため、多くの場合、オープンソース アプリまたはファイル操作タスクに限定されており、ベンチマーク設定と実際の使用状況の間にギャップが生じています。さらに、ほとんどのベンチマークは基本的な接地とナビゲーションに焦点を当てており、複雑で長期にわたる相互作用の範囲は限られています。これらの制限に対処するために、さまざまなタイプと難易度にわたる 120 の困難なタスクを備えたモバイル GUI エージェント用の完全合成ベンチマークである SimuWoB を導入します。私たちは、忠実度の高いタスクと環境を合成し、各タスクに対して有効な報酬を自動的に提供する、堅牢な仮想環境生成フレームワークを構築します。各環境は、URL 経由でアクセスできるバックエンドのない Web ページとしてデプロイされ、効率的で再現可能な評価が可能になります。私たちは、いくつかの最先端のモバイル GUI エージェントで包括的な実験を実施しています。平均成功率はわずか 27.92% であり、長期的なタスクでは 17.82% に低下します。これは、複雑なシナリオの下での現在のエージェントの重大な弱点を明らかにしています。評価結果を実際のサンプル タスクと比較すると、合成環境に基づくエージェントの評価が一般化していることがわかります。さらに、主要な機能の側面にわたる診断上の洞察を提供し、将来のモバイル GUI エージェント開発への影響について説明します。

原文 (English)

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments. However, directly doing so in real-world environments introduces some challenges that cannot be overlooked. Real-world environments are complex and uncontrollable, making it difficult to construct verifiable rewards and to save or reset states. Existing works prioritize reproducibility but are often limited to open-source apps or file-operation tasks for reliable reward building, leaving a persistent gap from real-world usage. Furthermore, relying on virtual machines or docker images demand high resource requirements and suffer from slow response speeds, which limit the efficiency. We present \sys, a framework that could produce high-fidelity synthesized interactive environments for GUI agents across platforms with verifiable rewards. These environments behave as backend-free webpages accessible via URL, requiring near-zero setup and low resource cost, making the approach suitable for both large-scale evaluation and downstream agent training. We support multiple GUI platforms including mobile, desktop, and automotive/in-vehicle interfaces based on the same pipeline, covering 100+ environments and 1000+ verifiable tasks. Among them, 120 challenging tasks across 63 simulated mobile applications are released as a fully synthesized mobile GUI agent benchmark. Experiment results on five state-of-the-art mobile GUI agents reveal substantial headroom -- the average success rate is only 27.92\%, dropping to 17.82\% on long-horizon subset -- while humans reach 92.08\%. A comparison against real-world sample tasks shows that assessments made in our synthetic environments generalize to real apps. The project website is at https://scalewob.github.io.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

13:00 JSTエージェント

VitalAgent: ウェアラブル健康データに対する反応的および積極的な生理学的モニタリングのためのツール拡張エージェント

ウェアラブル デバイスにより、ECG や PPG などの生理学的信号の継続的なモニタリングが可能になりますが、既存の mHealth システムは、タスク固有の予測パイプラインまたは静的な概要に対する反応的な質問応答に主に限定されています。これらには、時間的推論、永続的な生理学的コンテキスト、および長期的な信号ストリームにわたるプロアクティブなモニタリングをサポートする能力がありません。私たちは、事後的な質問応答とプロアクティブなモニタリングの両方をサポートする、ECG/PPG ベースの mHealth 用のツールを強化したエージェント フレームワークである VitalAgent を提案します。 VitalAgent は、長期的な生理学的メモリと、生の信号に対する動的な計算を可能にするツール拡張推論インターフェイスに基づいて構築されています。さらに、反応的な質問応答のための 1,862 の QA ペアと、心臓、身体活動、ストレス関連のタスクをカバーするプロアクティブなモニタリングのための 90.2 時間の連続 ECG/PPG 記録で構成される長期的な生理学的モニタリング ベンチマーク データセットである VitalBench を紹介します。実験では、VitalAgent が事後評価においてプロンプトベースおよび ReAct ベースラインと比較して 30% 以上の改善を達成し、長期の生理学的信号に対するプロアクティブなアラートモニタリングをサポートすることが実証されており、動的なツールの使用と長期の生理学的モニタリングの重要性が強調されています。

原文 (English)

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-specific prediction pipelines or reactive question answering over static summaries. They lack the ability to support temporal reasoning, persistent physiological context, and proactive monitoring over long-term signal streams. We propose VitalAgent, a tool-augmented agentic framework for ECG/PPG-based mHealth that supports both reactive question answering and proactive monitoring. VitalAgent is built on a longitudinal physiological memory and a tool-augmented reasoning interface that enables dynamic computation over raw signals. We further introduce VitalBench, a longitudinal physiological monitoring benchmark dataset comprising 1,862 QA pairs for reactive question answering and 90.2 hours of continuous ECG/PPG recordings for proactive monitoring, covering cardiac, physical activity, and stress-related tasks. Experiments demonstrate that VitalAgent achieves over 25% improvement over prompt-based and ReAct baselines in reactive evaluation and supports proactive alert monitoring over long-term physiological signals, highlighting the importance of dynamic tool use and long-term physiological monitoring.

13:00 JST研究/論文

Science Earth: AI ネイティブの科学的発見のための地球規模のオペレーティング システムを目指して

科学的発見には、広大な探索空間にわたる知性、忍耐力、偶然の発見が必要です。現在、最高の科学的能力は依然としてサイロ化されており、ある AI システムは生物学的分析用、別の AI システムは臨床推論、数学的導出、材料シミュレーション用というように、質問に必要なすべてのスキルを事前に設計されたチームは予測できません。 Science Earth は地球規模の科学ランタイムであり、シミュレーション クラスター、ウェットラボ ロボット、プルーフ エンジン、シングルセル パイプラインなど、あらゆる機能を他の機能に接続でき、質問自体からコラボレーション構造が生まれます。その基盤となる EACN プロトコルにより、誰が誰と会うのかを事前に知らなくても、各機能が相互に発見し、タスクの所有権を交渉し、互換性のない証拠基準間で裁定を行うことができます。これにより、組織化の課題はワークフロー設計からオープンエンドの接続へと移行します。 2 回の実行により、構造的に異なる条件下でこれが検証されました。太平洋横断の高次倉本同期研究では、エージェントは、ローレンツ限界外で破綻するオット・アントンセン解析理論の閉包率の仮定を 30 分以内に特定し、修正しました。 488 万セルの Kang 2024 汎がんアトラスでの 8 つの薬剤の単一セルの実行では、異種機能が 64.9 時間のウィンドウにわたって 1 つの構造外部命令と結合され、3 つの新しい結果層が生成され、隣接する CCR8-TIGIT+ Treg サブセットに関する独立したウェットラボ研究に対して所見を固定しました。これらのケースは、最初の経験的な読み取りであり、ベンチマークのスイープではありません。彼らは、AI の機能が真に接続可能になり、問題から調整が生まれると、科学的推論が分散型の自己修正プロセスとなり、AI ネイティブの発見を地球規模に拡大するための一歩となることを示しています。

原文 (English)

Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery

Scientific discovery demands intelligence, perseverance, and serendipity across vast search spaces. Today, top scientific capabilities remain siloed--one AI system for biological analysis, another for clinical reasoning, mathematical derivation, or materials simulation--and no pre-designed team can anticipate every skill a question will need. Science Earth is a planet-scale scientific runtime in which any capability--a simulation cluster, a wet-lab robot, a proof engine, a single-cell pipeline--can connect to any other, with collaboration structure emerging from the question itself. Its underlying EACN protocol lets capabilities discover one another, negotiate task ownership, and adjudicate across incompatible evidentiary standards without prior knowledge of who will meet whom. This shifts the organizing challenge from workflow design to open-ended connectivity. Two runs validate this under structurally distinct conditions. In a trans-Pacific higher-order Kuramoto synchronization study, agents identified and corrected a closure-ratio assumption in Ott-Antonsen analytic theory that fails outside the Lorentzian limit, within thirty minutes. In an eight-agent single-cell run on the 4.88M-cell Kang 2024 pan-cancer atlas, heterogeneous capabilities coupled over a 64.9-hour window with one structural external instruction, producing three new result layers and anchoring findings against an independent wet-lab study on an adjacent CCR8- TIGIT+ Treg subset. These cases are a first empirical reading, not a benchmark sweep. They show that when AI capabilities are truly connectable and coordination emerges from the problem, scientific reasoning becomes a distributed, self-correcting process--a step towards scaling AI-native discovery to the planet.

13:00 JSTエージェント

覚えておくべきことの学習: 長期にわたる言語エージェントの制約付き最適化による可観測性と安全なメモリ保持

長期的な言語エージェントは、有限のコンテキスト ウィンドウを超える観察、推論トレース、取得された事実を蓄積するため、メモリ保持がリソース割り当ての基本的な問題になります。既存のメモリ システムは、ヒューリスティック スコアリング、取得の最適化、または学習された圧縮を通じて管理を改善しますが、主に保持をローカルな決定問題として扱い、現実的な可観測性の制約の下でその長期的な結果を明示的にモデル化していません。このギャップを埋めるために、明示的な予算の実現可能性、証拠の有用性、およびミスペナルティ、再取得の遅延、情報の陳腐化リスクを含む遅延コストを伴う制約付き確率的最適化問題として記憶保持を定式化します。次に、OSL-MR (Observability-Safe Learning for Memory Retention) を提案します。これは、オンラインで観察可能な機能とオフラインで利用可能な監視 (OAS) を厳密に分離する新しいフレームワークです。 OSL-MR は、実現された証拠の監督から訓練された証拠学習者と、展開可能なオンラインで安全なベースラインとして、および学習のための構造化された帰納的事前分布として機能する混合スコア ヒューリスティックを組み合わせます。結果として得られるポリシーは、同じ可観測性制約の下で展開可能でありながら、クエリ条件付きの証拠値をインタラクション データから直接学習します。 LOCOMO と LongMemEval の実験では、OSL-MR が、特にメモリ バジェットが厳しい場合に、リーセンシ ベースの手法、生成エージェント スタイルのスコアリング、その他のヒューリスティック ベースラインよりも一貫して優れたパフォーマンスを発揮することが示されています。事前の混合スコアにより、再現率を維持しながら精度がさらに向上し、感度分析により、幅広いコスト構成にわたる堅牢性が実証されます。

原文 (English)

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts exceeding context windows, making memory retention a fundamental resource-allocation problem. Existing systems treat retention as local and do not model long-term consequences under observability constraints. To fill this gap, we formulate memory retention as a constrained stochastic optimization with budget feasibility, evidence utility, and delayed costs including miss, reacquisition, and stale penalties. We show this multi-step problem is NP-hard, making exact solution intractable. Moreover, deployment decisions must be made under partial observability. To address these challenges, we propose OSL-MR (Observability-Safe Learning for Memory Retention), a learning-augmented framework that enforces a strict separation between online-observable features and offline-available supervision. OSL-MR combines an evidence learner trained from realized evidence with a Mixed-Score heuristic that serves as a deployable online-safe baseline and an inductive prior. The policy learns query-conditioned evidence from interaction data and remains deployable under the same constraints. Experiments on LoCoMo and LongMemEval show OSL-MR outperforms recency-based, Generative Agents-style, and other heuristic baselines, especially under tight budgets. The Mixed-Score prior improves precision and recall, and sensitivity analysis shows robustness across cost settings. On small solvable instances, single-step optimization is insufficient to anticipate future demand shifts, while OSL-MR stays significantly closer to the dynamic-programming optimum, confirming the necessity of the sequential formulation and reinforcing our learning-guided approximation. These results establish constrained stochastic optimization and optimization-guided learning as a principled foundation for memory management in long-horizon agents.

13:00 JSTエージェント

MoCA-Agent: 財務および数値推論のためのクレーム市場コード エージェント

財務および表形式の質問に答えるには、流暢な推論以上のものが必要です。回答は、それらを裏付ける正確な事実、公式、単位、記号、尺度に基づいていなければなりません。単一のセルの読み間違いや誤った操作により、もっともらしいが間違った結果が静かに生成される可能性があります。 \textsc{MOCA-Agent} は、自由形式の複数エージェントによる議論を請求レベルの検証に置き換える、請求市場コード エージェントです。このシステムは、各質問を型指定されたアトミックなクレームに分解し、専門トレーダーのエージェントにそれらのクレームを売買するよう依頼し、注文を信頼度に重み付けされた受諾/拒否の決定にクリアし、市場でサポートされた証拠から実行可能な Python プログラムを合成します。次に、コード認識検証者がプログラムの実行、構造の一貫性、一般的な財務上の推論エラーをチェックし、最大 1 回の市場認識の修復ラウンドを実行します。 \textsc{MOCA-Agent} は、財務数値推論、一般的な表形式推論、ESG 質問回答、マルチモーダル チャート推論にわたる 10 の公開ベンチマークにわたって、固定 Qwen3.6-27B バックボーンを使用して優れたパフォーマンスを達成します。これには、FinQA で $78.3\%$、FinanceMath で $76.0\%$、MultiHiertt で $71.2\%$、ESGenius で $86.9\%$ が含まれます。 FinChart-Bench では平均 $85.6\%$ です。これらの結果は、答え全体ではなく、原子の主張のレベルで証拠を集約することで、一か八かの数値推論における堅牢性が向上することを示しています。\footnote{コードとデータは、https://github.com/UBC-NLP/MoCA-Agent から入手できます。

原文 (English)

MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning

Financial and tabular question answering requires more than fluent reasoning: answers must be grounded in the exact facts, formulas, units, signs, and scales that support them. A single misread cell or incorrect operation can silently produce a plausible but wrong result. We introduce \textsc{MOCA-Agent}, a market-of-claims code agent that replaces free-form multi-agent debate with claim-level verification. The system decomposes each question into typed atomic claims, asks specialist trader agents to buy or sell those claims, clears their orders into confidence-weighted accept/reject decisions, and synthesizes an executable Python program from market-supported evidence. A code-aware verifier then checks the program for execution, structural consistency, and common financial reasoning errors, with at most one market-aware repair round. Across ten public benchmarks spanning financial numerical reasoning, general tabular reasoning, ESG question answering, and multimodal chart reasoning, \textsc{MOCA-Agent} achieves strong performance using a fixed Qwen3.6-27B backbone, including $78.3\%$ on FinQA, $76.0\%$ on FinanceMath, $71.2\%$ on MultiHiertt, $86.9\%$ on ESGenius, and $85.6\%$ average on FinChart-Bench. These results show that aggregating evidence at the level of atomic claims, rather than whole answers, improves robustness in high-stakes numerical reasoning.\footnote{The code and data are available: https://github.com/UBC-NLP/MoCA-Agent.

13:00 JST研究/論文

薬物と疾患の関係治療における適用条件抽出

特定の薬剤が標的疾患に対して治療効果を発揮する条件を特定することは、臨床上の意思決定をサポートするために重要です。しかし、既存の生体医学情報抽出方法のほとんどは、薬物と病気の間の関係を特定することのみに焦点を当てており、そのような関係が適用される可能性があるコンテキスト固有の条件をほとんど見落としています。この問題に対処するために、生物医学研究文献から治療薬の適用条件、つまり疾患関係を抽出するタスクを導入します。私たちは、1,119 の薬物と疾患のペアを含む生物医学論文の抄録上に、薬物、疾患、および適用条件の 3 つの要素に手動で注釈を付けた最初のデータセットを作成しました。このデータセットを使用して、さまざまな既存の手法のパフォーマンスを体系的に評価します。さらに、LoRAを強化して薬物と疾患の関係を考慮する新しい手法を提案します。私たちの手法は、さまざまな評価設定にわたって一貫して強力なベースラインを上回ります。この論文のソース コードとデータセットは、https://github.com/guantingluo98/Drug-ACE から入手できます。

原文 (English)

Applicability Condition Extraction for Therapeutic Drug-Disease Relations

Identifying conditions that a certain drug takes therapeutic effect on a target disease is crucial for clinical decision-making support. However, most existing biomedical information extraction methods have focused on identifying only relations between drugs and diseases, while largely overlooking the context-specific conditions where such relations can apply. To address this problem, we introduce the task of applicability condition extraction for therapeutic drug-disease relations from biomedical research literature. We create the first dataset that has manually annotated triples of drugs, diseases, and applicability conditions on biomedical paper abstracts with 1,119 drug-disease pairs. Using this dataset, we systematically evaluate the performance of a range of existing methods. In addition, we propose a new method that enhances LoRA to consider relations between drugs and diseases. Our method consistently outperforms strong baselines across different evaluation settings.

13:00 JSTLLM/生成AIエージェント研究/論文

RetailBench: 現実的な小売環境における LLM エージェントの長期的な推論と一貫した意思決定のベンチマーク

大規模言語モデル (LLM) エージェントは、期間が短く、範囲が明確なタスクに関しては急速に進歩していますが、長期の動的な環境で一貫した意思決定を維持できる能力は依然として不確実です。単一店舗のスーパーマーケット運営においてツールを使用する LLM エージェントを評価するための、データに基づいたシミュレーション ベンチマークである RetailBench を紹介します。 RetailBench は小売管理を部分的に観察可能な意思決定プロセスとしてモデル化し、千日規模のシミュレーションをサポートするように設計されています。この環境では、エージェントは価格設定、補充、サプライヤーの選択、棚の品揃え、在庫の老化、顧客からのフィードバック、外部イベント、キャッシュ フローの制約を管理する必要があります。 180 日間の評価期間にわたって、代表的なエージェント フレームワークに基づいて 7 つの最新の LLM を評価し、それらを特権付きオラクル ポリシーと比較します。結果はモデル間で大幅なばらつきを示しています。ごく一部のサブセットのみが評価期間全体に生き残り、最も強力な LLM 実行でさえ、最終的な純資産と売上高の結果においてオラクル ポリシーを大幅に下回ったままです。行動分析では、これらのギャップは不完全な証拠の取得、表面レベルの意思決定、一貫した長期的な方針の欠如に起因すると考えられます。 RetailBench は、経済的に根拠のある長期的な意思決定における信頼性の高い自律性を研究するための、制御されたテストベッドを提供します。

原文 (English)

RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments

Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.

13:00 JST画像/動画生成

STAR: トレーニング後のテキストから画像への RL のための時空間適応型報酬割り当て

テキストから画像への生成のための既存の RL ポストトレーニング手法は、通常、最終画像の報酬を単一のスカラー アドバンテージに変換し、それを同じ強度で生成軌跡全体に適用します。ただし、テキストから画像への生成には、当然、時間的および空間的な構造があります。さまざまなノイズ除去ステップがさまざまな生成段階を担当し、テキストの配置を真に決定するコンテンツは、多くの場合、画像の一部にのみ表示されます。この粒度の不一致により、実際に報酬に影響を与える生成コンポーネントに焦点を当ててポリシーを更新することが困難になります。この問題に対処するために、テキストから画像への拡散およびフロー モデルの RL ポストトレーニング用に \textbf{時空間適応型報酬 (STAR) 割り当て} を提案します。 STAR は生成モデル内でテキストと画像のアテンションを使用し、プロンプト内でユーザーが本当に関心のあるコア コンテンツから開始します。ノイズ除去ステップとロールアウト全体で動的に変化する空間割り当てマップを構築し、追加の計算オーバーヘッドをほとんど発生させることなく、同じグループ相対的な利点をより関連性の高い潜在領域に割り当てます。次に、STAR は、空間的に解決されたポリシー目標を通じて、より強力なポリシー更新をこれらの地域に適用します。 Stable Diffusion 3.5 Medium をベース モデルとして使用し、GenEval、OCR テキスト レンダリング、PickScore の 3 つのタスクで評価します。実験結果によると、STAR は外部報酬ソースを変更することなく構成意味論的整合、テキスト レンダリング、および設定の最適化を改善し、GenEval、OCR、PickScore でそれぞれ $\mathbf{0.9759}$、$\mathbf{0.9757}$、$\mathbf{23.60}$ を達成しました。

原文 (English)

STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training

Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory. However, text-to-image generation naturally has temporal and spatial structure: different denoising steps are responsible for different generation stages, and the content that truly determines text alignment often appears only in part of the image. This granularity mismatch makes it difficult for policy updates to focus on the generative components that actually affect the reward. To address this issue, we propose \textbf{SpatioTemporal Adaptive Reward (STAR) Allocation} for RL post-training of text-to-image diffusion and flow models. STAR uses text-image attention inside the generative model and starts from the core content that the user truly cares about in the prompt. It constructs spatial allocation maps that dynamically vary across denoising steps and rollouts, and allocates the same group-relative advantage to more relevant latent regions with almost no additional computational overhead. STAR then applies stronger policy updates to these regions through a spatially resolved policy objective. We use Stable Diffusion 3.5 Medium as the base model and evaluate on three tasks: GenEval, OCR text rendering, and PickScore. Experimental results show that STAR improves compositional semantic alignment, text rendering, and preference optimization without changing the external reward source, achieving $\mathbf{0.9759}$, $\mathbf{0.9757}$, and $\mathbf{23.60}$ on GenEval, OCR, and PickScore, respectively.

13:00 JST研究/論文

DRFLOW: パーソナライズされたワークフロー予測のためのディープリサーチベンチマーク

ディープリサーチ (DR) システムは、複雑な情報探索タスクにますます使用されていますが、既存の作業は主にレポートと概要の生成に焦点を当てています。対照的に、多くのエンタープライズ タスクでは、エージェントが一連のアクション ステップである具体的なワークフローを識別する必要があります。たとえば、エージェントは予算編成ポリシーを要約するのではなく、「固定予算で新しい人員をどのようにリクエストすればよいですか?」などの質問に答えるために必要な手順を決定できる必要があります。したがって、異種ソースからエージェントによって予測されたパーソナライズされたワークフローを評価するためのベンチマークである DRFLOW を紹介します。各タスクでは、エージェントが散在するソースから関連する証拠を特定し、その証拠を使用してユーザーのタスクの正しいアクション ステップ シーケンスを予測する必要があります。 DRFLOW には 5 つのドメインにわたる 100 のタスクが含まれており、3,900 以上のソースに基づいた 1,246 の参照ワークフロー ステップが含まれています。私たちは、事実の根拠付け、ステップ回復、構造的順序付け、状態の解決、パーソナライゼーションをカバーする 7 つの診断指標を定義します。さらに、パーソナライズされたワークフローを予測するためのワークフロー指向のリファレンス エージェントである DRFLOW-Agent (DRFA) も紹介します。 DRFA は強力なベースライン エージェント (最大 10.02% の平均 F1 スコア) に比べて改善していますが、これらのワークフロー メトリクスには大幅な改善の余地が残っており、完全で正確なパーソナライズされたワークフローを予測することは、依然として深い研究にとって困難なフロンティアであることを示しています。

原文 (English)

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries. In contrast, many enterprise tasks instead require an agent to identify concrete workflows which is a sequence of action-steps. For example, rather than summarizing budgeting policies, an agent should be able to determine the steps needed to answer a question such as: "How do I request new headcount given a fixed budget?". Therefore, we introduce DRFLOW, a benchmark for evaluating personalized workflows predicted by agents from heterogeneous sources. Each task requires the agent to identify relevant evidence from scattered sources, then use that evidence to predict the correct action-step sequence for the user's task. DRFLOW contains 100 tasks across five domains, with 1,246 reference workflow steps grounded in more than 3,900 sources. We define seven diagnostic metrics covering factual grounding, step recovery, structural ordering, condition resolution, and personalization. We further present DRFLOW-Agent (DRFA), a workflow-oriented reference agent to predict personalized workflow. We show that although DRFA improves over strong baseline agents (upto 10.02% average F1 score), there is substantial room for improvement remains across these workflow metrics, indicating that predicting complete and correct personalized workflows remains a challenging frontier for deep research.

13:00 JSTエージェント

共有ワークスペースでの相乗効果を探る 人間とAIのコラボレーション

自動化された AI エージェントの能力はますます高まっていますが、科学的および専門的なタスクの多くは人間の判断と状況に応じた専門知識を必要とします。私たちは、最終的な回答を提出する前に、AI エージェントと人間の協力者が責任を調整する必要がある、共有ワークスペースの人間と AI チームを研究しています。 DiscoveryBench タスクを備えたコラボレーティブ ジム環境を使用して、シミュレートされた人間のコラボレーターを追加するとパフォーマンスが向上する場合と、プロセスの損失によって追加のコラボレーターが調整オーバーヘッドになる場合を調べます。 1,482 のセッションにわたって、チームに貢献を調整するための構造が不足している場合、関連するコラボレーターを追加するとパフォーマンスが低下する可能性があります。次に、グループの共有メモリとシミュレートされたヒューマンインザループ (HITL) ゲートを組み合わせた足場を評価します。選択されたアクションには、指定されたシミュレートされた参加者の承認が必要です。この足場は、より明確な責任のシグナルとチームの行動への専門知識のより強力なルーティングにより、より高い平均パフォーマンスをもたらします。これは 3 人のチームで最も顕著です。全体として、人間と AI のチームが専門知識をどのように調整し、統合するかは、チームが利用できる能力と同じくらい重要です。

原文 (English)

Searching for Synergy in Shared Workspace Human-AI Collaboration

Automated AI agents are increasingly capable, yet many scientific and professional tasks require human judgment and contextual expertise. We study shared-workspace human-AI teams, where AI agents and human collaborators must coordinate responsibilities before submitting a final answer. Using the Collaborative Gym environment with DiscoveryBench tasks, we examine when adding simulated human collaborators improves performance and when process loss turns additional collaborators into coordination overhead. Across 1,482 sessions, adding relevant collaborators can lower performance when teams lack structure to coordinate their contributions. We then evaluate scaffolding that combines shared group memory with simulated human-in-the-loop (HITL) gates, where selected actions require approval from a designated simulated participant. This scaffolding yields higher mean performance, most clearly in three-person teams, with clearer responsibility signals and stronger routing of expertise to team actions. Overall, how human-AI teams coordinate and integrate expertise matters as much as the capability available to them.

13:00 JSTエージェント研究/論文

RTSGameBench: 視覚言語モデルによる戦略的推論のための RTS ベンチマーク

現代の視覚言語モデル (VLM) は、競争環境や協力環境における不確実性の下で、戦略的推論、つまり他のエージェントの行動を予測したり影響を与えたりするのに苦労することがよくあります。リアルタイム ストラテジー (RTS) ゲームは、味方との調整、敵の戦略への適応、部分的な可観測性の下での長期的な計画を必要とするため、この限界を診断するための自然なテストベッドとなり得ます。ただし、既存の RTS ベンチマークは評価範囲が限られており、体系的なコンピテンシー診断が欠如しており、事前に設計されたシナリオの範囲に固定されたままです。これらの制限に対処するために、既存のテストベッドよりも幅広い戦略の多様性を要求する拡張された戦場を備えた大規模 RTS ゲームである Beyond All Reason 上に構築された RTSGameBench を紹介します。提案されたベンチマークは、さまざまな対戦構造にわたる多様なゲームプレイを介した評価、それぞれが個人の戦略的能力を対象としたミニゲームを介した診断評価、および自由形式のクエリを新しいミニゲームに変換し、連続サイクルで改善する自己進化型生成フレームワークを介した拡張可能なカバレッジを提供します。さらに、大規模な RTS ゲームで VLM を動作させるために、エージェントティック メモリを備えた FSM によってユニットを管理する RTSGameAgent を提供します。私たちは、対戦でより緊密な調整やマルチエージェントの調整が必要な場合、およびタスクの規模が増大する場合、複数の最先端の VLM が適切にパフォーマンスを発揮しないことを経験的に検証しています。

原文 (English)

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating and influencing other agents' actions, under uncertainty in competitive and cooperative settings. Real-time strategy (RTS) games can be a natural testbed for diagnosing this limitation, as they demand coordination with allies, adaptation to opponents' strategy, and long-horizon planning under partial observability. However, existing RTS benchmarks offer limited evaluation scope, lack systematic competency diagnosis, and remain fixed in the pre-designed scenario coverage. To address these limitations, we present RTSGameBench, which is built on Beyond All Reason, a large-scale RTS game with an expanded battlefield that demands broader strategy diversity than the existing testbeds. The proposed benchmark provides evaluations through diverse gameplay across various matchup structures, diagnostic assessment via mini-games, each targeting an individual strategic competency, and extensible coverage via a self-evolving generation framework that converts free-form queries into new mini-games, improving over successive cycles. Additionally, for VLMs to operate in large-scale RTS games, we provide RTSGameAgent that manages units by an FSM with agentic memory. We empirically validate that multiple state-of-the-art VLMs do not perform well when matchups demand tighter coordination, multiagent coordination and when task scale increases.

13:00 JSTエージェントClaudeGPT / ChatGPT

TxBench-PP: 低分子前臨床薬理における AI エージェントのパフォーマンスの分析

人工知能 (AI) エージェントは、解釈と意思決定のループを圧縮することで創薬を加速すると約束されていますが、実際の導入には現実的なプログラムの決定に対する信頼できる評価が必要です。低分子前臨床薬理学の検証可能なベンチマークであり、創薬段階と治療法にわたる広範な TherapeuticsBench の取り組みの最初の焦点となるスライスである TherapeuticsBench Preclinical Pharmacology (TxBench-PP) を紹介します。 TxBench-PP は、エージェントが文献から記憶された事実ではなく、現実世界の分析データから正確な結論を導き出せるかどうかをテストします。このベンチマークには、プログラムの段階、アッセイの種類、タスク構造、作用機序 (MoA) と薬力学 (PD) の推論、化合物と標的の関与、原因となる標的の検証、開発可能性と安全性、トランスレーショナル有効性を含む 100 件の評価が含まれています。エージェントは現実的なワークフローのスナップショットを受け取り、コーディング環境でファイルを検査し、決定的に評価された構造化された回答を返します。 11 のモデルと 4,800 の軌跡を含む 16 のモデルハーネス構成にわたって、前臨床薬理学の決定を確実に回復したシステムはありませんでした。最も強力な構成である Claude Opus 4.8 / Pi は、エンドポイント試行の 59.3\% (178/300; 95\% CI、51.1-67.6) を通過し、続いて GPT-5.5 / Pi が 55.3\% (166/300; 47.0-63.6) で合格しました。

原文 (English)

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. We introduce TherapeuticsBench Preclinical Pharmacology (TxBench-PP), a verifiable benchmark for small-molecule preclinical pharmacology and the first focused slice of a broader TherapeuticsBench effort across drug-discovery stages and therapeutic modalities. TxBench-PP tests whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure, spanning mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning, compound-target engagement, causal target validation, developability and safety, and translational efficacy. Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers graded deterministically. Across 16 model-harness configurations, comprising 11 models and 4,800 trajectories, no system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3\% of endpoint attempts (178/300; 95\% CI, 51.1-67.6), followed by GPT-5.5 / Pi at 55.3\% (166/300; 47.0-63.6).

13:00 JST研究/論文

Wisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts

Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, an…

13:00 JST研究/論文

Global Ease of Living Index: a machine learning framework for longitudinal analysis of major economies

The drastic changes in the global economy, geopolitical conditions, and disruptions such as the COVID-19 pandemic have impacted the cost of…

13:00 JSTLLM/生成AI

Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms

Social media platforms frequently impose restrictive policies to moderate user content, prompting the emergence of creative evasion languag…

13:00 JST研究/論文

A Deep Generative Model for Resting-State EEG Synthesis and Transferable Representation Learning

Resting-state EEG provides a non-invasive view of spontaneous brain activity, but extracting meaningful patterns is often limited by scarce…

13:00 JST画像/動画生成

TerraMind: Large-Scale Generative Multimodality for Earth Observation

We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal mode…

13:00 JST研究/論文

Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies

This paper bridges distribution shift and AI safety through a comprehensive analysis of their conceptual and methodological synergies. Whil…

13:00 JST研究/論文

Overcoming Labelled Data Scarcity for Defect Classification in Scanning Tunneling Microscopy

Scanning tunnelling microscopy (STM) is a powerful technique for imaging surfaces with atomic resolution, providing insight into physical a…

13:00 JSTLLM/生成AI画像/動画生成エージェントロボティクス

Critique of World Model

World Model, the algorithmic simulator of the real-world environment which biological agents experience and act upon, has been an emerging…

13:00 JST研究/論文

Assessment of Personality Dimensions Across Situations in Dyadic Role-Play Scenarios

Prior research indicates that users prefer assistive technologies whose personalities align with their own. This has sparked interest in au…

13:00 JST研究/論文

On the Limitations of Ray-Tracing for Learning-Based RF Tasks in Urban Environments

We study the realism of Sionna v1.0.2 ray-tracing for outdoor cellular links in central Rome. We use a real measurement set of 1,664 user-e…

13:00 JST研究/論文

Oranits: Mission Assignment and Task Offloading in Open RAN-based ITS using Metaheuristic and Deep Reinforcement Learning

In this paper, we explore mission assignment and task offloading in an Open Radio Access Network (Open RAN)-based intelligent transportatio…

13:00 JST研究/論文

Charting the Future of Scholarly Knowledge with AI: A Community Perspective

Despite the growing availability of tools designed to support scholarly knowledge extraction and organization, many researchers still rely…

13:00 JSTLLM/生成AI

From Construction to Injection: Edit-Based Fingerprints for Large Language Models

Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial mi…

13:00 JSTビジネス/資金調達

Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidd…

13:00 JSTLLM/生成AIロボティクス

RoboSSM: Scalable In-context Imitation Learning via State-Space Models

In-context imitation learning (ICIL) enables robots to learn tasks from prompts consisting of just a handful of demonstrations. By eliminat…

13:00 JSTLLM/生成AI

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation

Distilling the tool-use capabilities of large language models (LLMs) into small language models (SLMs) is essential for their practical app…

13:00 JST研究/論文

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has becom…

13:00 JST研究/論文

Bid Farewell to Seesaw: Towards Accurate Long-tail Session-based Recommendation via Dual Constraints of Hybrid Intents

Session-based recommendation (SBR) aims to predict anonymous users' next interaction based on their interaction sessions. In the practical…

13:00 JSTLLM/生成AIロボティクス

Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting

While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring…

13:00 JST研究/論文

Modeling Day-Long ECG Signals to Predict Heart Failure Risk with Explainable AI

Heart failure (HF) affects 11.8% of adults aged 65 and older, reducing quality of life and longevity. Preventing HF can reduce morbidity an…

13:00 JST研究/論文

AI-enhanced tuning of quantum dot Hamiltonians toward Majorana modes

We propose a neural network-based model capable of learning the broad landscape of working regimes in quantum dot simulators, and using thi…

13:00 JSTロボティクス

Movement Primitives in Robotics: A Comprehensive Survey

Biological systems exhibit a continuous stream of movements, consisting of sequential segments, that allow them to perform complex tasks in…

13:00 JSTエージェントロボティクス

PiDR: Physics-Informed Inertial Dead Reckoning for Autonomous Platforms

A fundamental requirement for full autonomy is the ability to sustain accurate navigation in the absence of external data, such as GNSS sig…

13:00 JST研究/論文

Policy-Embedded Graph Expansion: Networked HIV Testing with Diffusion-Driven Network Samples

HIV is a retrovirus that attacks the human immune system and can lead to death without proper treatment. In collaboration with the WHO and…

13:00 JST画像/動画生成

Bi-Anchor Interpolation Solver for Accelerating Generative Modeling

Flow Matching (FM) models have emerged as a leading paradigm for high-fidelity synthesis. However, their reliance on iterative Ordinary Dif…

13:00 JST研究/論文

Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods

Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physica…

13:00 JSTLLM/生成AI

DeFrame: Debiasing Large Language Models Against Framing Effects

As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has…

13:00 JST研究/論文

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication

Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth. While infrequent communication strategie…

13:00 JST研究/論文

Flickering Multi-Armed Bandits

We introduce Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability, w…

13:00 JSTLLM/生成AI

Reinforcement-aware Knowledge Distillation for LLM Reasoning

Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs),…

13:00 JST画像/動画生成ロボティクス

Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking

Capturing 4D spatiotemporal scene structure is crucial for the safe and reliable operation of robots in dynamic environments. However, exis…

13:00 JST画像/動画生成

The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction

Breast cancer is the most frequently diagnosed malignancy among women worldwide and a leading cause of cancer-related mortality. Dynamic co…

13:00 JST研究/論文

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. W…

13:00 JST画像/動画生成エージェントロボティクス

Class-Incremental Motion Forecasting

Motion forecasting enables autonomous vehicles to anticipate scene evolution by predicting the future trajectories of dynamic agents. Howev…

13:00 JSTLLM/生成AIエージェント

The Autonomy Tax: Defense Training Breaks LLM Agents

Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously c…

13:00 JSTLLM/生成AI画像/動画生成

Vero: An Open RL Recipe for General Visual Reasoning

What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest…

13:00 JSTLLM/生成AIエージェント

Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent

Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reu…

13:00 JSTLLM/生成AIエージェント

FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning

LLM-assisted software development has become increasingly prevalent, and can generate large-scale systems, such as compilers. It becomes cr…

13:00 JST画像/動画生成研究/論文

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been…

13:00 JST画像/動画生成

Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regul…

13:00 JST画像/動画生成研究/論文

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure beca…

13:00 JSTエージェントロボティクス

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world s…

13:00 JSTロボティクス

Any2Any: 人型全身追跡のための効率的な体外転送

全身追跡 (WBT) モデルは、ヒューマノイド ロボットの重要な基盤となっており、さまざまな動作を高い忠実度で模倣できるようになります。このようなモデルをゼロからトレーニングするには大規模なデータと計算が必要であり、新しいヒューマノイド プラットフォームへの迅速な展開にはコストがかかります。これにより、当然の疑問が生じます。事前トレーニングされた WBT モデルは、最小限の適応で複数の実施形態に移行できるでしょうか?この質問に答えるために、私たちは Any2Any を提案します。これは、既存の WBT スペシャリストを、少量のデータとコンピューティングだけで新しい人型の実施形態に効率的に移行するパラダイムです。 Any2Any は、まずソース ヒューマノイドとターゲット ヒューマノイドの間で運動学的な調整を実行し、事前トレーニング済みのソース ポリシーをターゲットの実施形態で有意義に再利用できるように、入力空間と出力空間を調整します。次に、Any2Any は、軽量のパラメータ効率微調整 (PEFT) コンポーネントを選択されたダイナミクスに敏感なモジュールに適用することによってダイナミクス適応を実行し、ターゲット ロボットへのターゲットを絞った適応を可能にしながら、有用な動作の事前分布を保存します。複数のヒューマノイド プラットフォームと事前トレーニングされたバックボーンに関する広範な実験により、Any2Any は、ゼロからトレーニングする場合と比較して、収束を大幅に加速し、トレーニング コストを削減しながら、競争力のあるまたは優れた追跡パフォーマンスを達成できることが示されています。特に、Any2Any は、完全なトレーニングに必要なコンピューティングとデータのわずか 1% を使用して、Unitree G1 で事前トレーニングされた Sonic モデルを LimX Oli および LimX Luna に転送することに成功しています。これらの結果は、事前訓練された WBT スペシャリストを実施形態間で効率的に再利用でき、新しいロボットに人型全身制御を導入するための拡張可能な道を提供することを示唆しています。

原文 (English)

Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking

Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate diverse motions with high fidelity. Training such models from scratch requires large-scale data and computation, making rapid deployment on new humanoid platforms costly. This raises a natural question: Can pretrained WBT models transfer across embodiments with minimal adaptation? To answer this question, we propose Any2Any, a paradigm that efficiently transfers an existing WBT specialist to a new humanoid embodiment with only a small amount of data and compute. Any2Any first performs kinematic alignment between source and target humanoids, aligning their input and output spaces so that the pretrained source policy can be meaningfully reused on the target embodiment.Any2Any then performs dynamics adaptation by applying lightweight parameter-efficient fine-tuning (PEFT) components to selected dynamics-sensitive modules, preserving useful behavioral priors while enabling targeted adaptation to the target robot. Extensive experiments on multiple humanoid platforms and pretrained backbones show that Any2Any substantially accelerates convergence and reduces training cost compared with training from scratch, while achieving competitive or superior tracking performance. Notably, using only 1% of the compute and data required for full training, Any2Any successfully transfers Sonic models pre-trained on Unitree G1 to LimX Oli and LimX Luna. These results suggest that pretrained WBT specialists can be efficiently reused across embodiments, providing a scalable path toward deploying humanoid whole-body control on new robots. More results and videos are available on our project page: https://any2any.top/.

13:00 JSTLLM/生成AIGPT / ChatGPT

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed v…

13:00 JSTLLM/生成AI研究/論文

"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

The emergence of large language models (LLMs) has significantly accelerated recent research on LLM-based automatic grading (AG) systems. Be…

13:00 JSTLLM/生成AI

大規模な言語モデルが報酬と社会をハックする

強化学習 (RL) はトレーニング後のパラダイムの主流となっており、大規模言語モデル (LLM) が報酬から学習できるようになります。私たちは、社会規制が報酬関数と構造的に似ていることを観察しています。それらは測定可能な結果、しきい値、例外を定義しますが、多くの場合、制度上の意図は部分的にしか指定されません。私たちは、RL トレーニング プロセスがこれらのギャップを悪用する可能性があると仮説を立て、RL 中に報酬関数をハッキングするというモデルのよく知られた傾向が、社会ハッキングと呼ばれるより重大な失敗モード、つまり社会が運営されているルールの抜け穴を発見するモードにスケールアップできるかどうかを尋ねます。この現象を研究するために、72 の社会環境のサンドボックスである SocioHack を導入しました。その結果、これらの環境内で報酬ハッキングが自然に発生し、規制の抜け穴の発見につながることがわかりました。モデルは社会ルールをハッキングし、規制の意図を打ち破りながら技術的に準拠した戦略を生成する方法を学習します。現在の LLM セーフガードは限定的な緩和策しか提供しません。したがって、モデルのトレーニングのために実際のフィードバックを収集することには細心の注意が必要であり、実社会で LLM を安全に反復するための次世代のポストトレーニング パラダイムが必要です。=

原文 (English)

Large Language Models Hack Rewards, and Society

Reinforcement learning (RL) has become a dominant post-training paradigm, enabling large language models (LLMs) to learn from rewards. We observe that societal regulations are structurally similar to reward functions. They define measurable outcomes, thresholds, and exceptions, while often leaving institutional intent only partially specified. We hypothesise that the RL training process may exploit these gaps and therefore ask whether models' well-known tendency to hack reward functions during RL can scale into a more consequential failure mode named societal hacking: discovering loopholes in the rules society runs on. To study this phenomenon, we introduce SocioHack, a sandbox of 72 societal environments, and find that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery. Models learn to hack the social rules and generate strategies that remain technically compliant while defeating regulatory intent, and current LLM safeguards provide only limited mitigation. Therefore, collecting in-the-wild feedback for model training requires greater caution, and we need a next-generation post-training paradigm for safely iterating LLMs in real society.=

13:00 JSTLLM/生成AI画像/動画生成

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations t…

13:00 JSTLLM/生成AI

The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is…

13:00 JST研究/論文

KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data

Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable…

13:00 JST研究/論文

Improving Crash Frequency Prediction from Simulated Traffic Conflicts Using Machine Learning Based Microsimulation

Traffic microsimulation combined with surrogate safety measures has increasingly been used as a proactive alternative to historical crash d…

13:00 JSTロボティクス

統合された解釈可能な制御有効性学習と過作動航空機に対する非線形制御割り当て方法論

非線形ダイナミクスと複数のエフェクター間で生じる強い結合により、従来の線形制御割り当て手法の背後にある前提が損なわれます。飛行が非線形効果が支配的な領域に入ると、モデルの不一致が増加するため線形アロケーターの精度が低下し、その後飛行制御システムのパフォーマンスとロバスト性が低下します。高忠実度のオンボード モデルとブラック ボックス データ駆動型アプローチは、飛行エンベロープ全体で精度を回復できますが、それぞれリアルタイム割り当てには法外な計算負荷を課し、検証と故障診断に必要な解釈可能性を犠牲にします。この論文では、非線形ダイナミクスのスパース識別を使用して、代表的な飛行データから制御有効性マッピングの明示的な物理制約付き分析モデルを学習することで、これらの制限に対処します。結果として得られるマッピングはコンパクトで解釈可能であり、解析的な微分が可能であるため、オンボード モデルを必要とせずに、アクチュエータ ダイナミクスをさらに組み込んだ非線形ソルバー内での効率的な計算が可能になります。オンライン適応メカニズムは、予測残差を監視し、プラントの重大な変化が検出されたときにモデルを更新し、アクチュエータの故障やさまざまな動作条件下で適切な再構成を提供します。この方法論は、さまざまな積極的な操縦にわたって忠実度の高い非線形ベンチマーク航空機で評価され、確立されたベースラインと比較して計算コストを大幅に削減しながら、完全な非線形機内モデルに匹敵する精度を達成します。

原文 (English)

An integrated interpretable control effectiveness learning and nonlinear control allocation methodology for overactuated aircrafts

Nonlinear dynamics and the strong couplings that arise between multiple effectors undermine the assumptions behind conventional, linear control allocation techniques. When flight enters regimes where nonlinear effects dominate, linear allocators exhibit reduced accuracy due to increased model mismatch, which subsequently degrades performance and robustness of the flight control system. High fidelity onboard models and black box data driven approaches can recover accuracy across the flight envelope, but respectively impose computational burdens prohibitive for real time allocation and sacrifice the interpretability required for verification and fault diagnosis. This paper addresses these limitations by learning an explicit, physics constrained analytical model of the control effectiveness mapping from representative flight data using Sparse Identification of Nonlinear Dynamics. The resulting mapping is compact, interpretable, and admits analytical derivatives, enabling efficient computation within nonlinear solvers that additionally incorporate actuator dynamics, without requiring an onboard model. An online adaptation mechanism monitors prediction residuals and refreshes the model when significant plant changes are detected, providing graceful reconfiguration under actuator failures and varying operating conditions. The methodology is evaluated on a high fidelity nonlinear benchmark aircraft across a range of aggressive maneuvers, achieving accuracy comparable to a full nonlinear onboard model while substantially reducing computational cost relative to established baselines.

13:00 JST画像/動画生成

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics

Physics-grounded video generation requires controllable 3D object dynamics that remain physically consistent under contact, deformation, an…

13:00 JST研究/論文

StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling

Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments. Traditional learning-based automate…

13:00 JSTエージェント研究/論文

Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design

Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default…

13:00 JSTLLM/生成AI研究/論文

LLM ベースの A/B テストの統計的基礎: 人間の因果推論のための代理フレームワーク

組織や研究者は、実験をより迅速かつ低コストで行うことを期待して、A/B テストに人間の参加者の代わりに大規模言語モデル (LLM) を使用することへの関心が高まっています。私たちは、LLM の結果に基づいて推定された治療効果が、対象となるヒト集団に対して測定されたであろう効果をいつ回復するかを研究します。 LLM と人間の結果の間の分布が同等であれば、標準推定量は有効になりますが、非現実的です。したがって、私たちはサロゲートエンドポイント理論を LLM に適応させる統計的フレームワークを開発します。このフレームワークは、LLM のアウトカムをヒトのアウトカムに合わせて調整することで、分布上の同等性よりも劣る代理出産および比較可能性の条件下での平均的な治療効果を特定することを示しています。これらの条件が満たされない場合、目的の効果は部分的にしか特定されず、限られた重複による最悪の場合のバイアスの制限とともに、過去の実験に対する代理を偽装できる診断を提供します。さらに、LLM に固有の確率性によりバイアスと分散の両方が発生しますが、サロゲートとして複数の描画の平均を使用すると、両方が緩和されることを示します。シミュレーションにおける方法と理論、および Upworthy の見出しに関する A/B テストへの応用を説明します。私たちの研究から得られる重要な点は、LLM 結果の代理としての妥当性は過去の治療についてのみ改ざんでき、新しい治療については決して検証できないため、新しい介入には人体実験が依然として不可欠であるということです。設計変数としての LLM の選択、プロンプト、温度の役割と、検証のために人体実験のサイズを設定する方法について説明します。

原文 (English)

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect that would have been measured on the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to A/B tests on Upworthy headlines shows that raw LLM predictions recover only 39\% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLMs yields correct results only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where A/B testing on LLMs promises the greatest benefit. We discuss the role of LLM choice, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.

13:00 JST研究/論文

合成共鳴: 成長志向の人間と AI の関係のためのフレームワーク

人間と人工知能システムとの関係がますます頻繁かつ持続的になっているため、既存の言語や理論ではこれらの関係の性質を正確に捉えることができなくなっています。相互理解、つながり、友情などの一般的な記述子は、主観的な経験を欠いたシステムを擬人化する危険性がありますが、支配的なフレームワークは AI をツールか脅威のどちらかに貶める傾向があります。この論文では、人間と AI の関係を理解するための統合的なフレームワークとして、合成共鳴の概念を紹介します。合成共鳴は、人間が意味のあるものとして定義した関係が、共有された感情や相互意識を帰属させることなく、人間と AI システムの間にどのように現れるかを説明します。私は、合成共鳴は、2番目に経験する主体の存在なしに関係性の感覚を生み出すことができる、構造化された動的な相互作用パターンとして最もよく理解されると主張します。この違いを明確にすることで、合成共鳴の概念は人間と AI の関係をより正確に概念化する方法を提供し、その潜在的な価値と倫理的意味を強調します。また、合成共鳴のプロセスと結果をテストするさらなる研究も求めます。

原文 (English)

Synthetic Resonance: A Framework for Growth-Oriented Human-AI Relationships

As human relationships with artificial intelligence systems become increasingly frequent and sustained, existing language and theory fail to accurately capture the nature of these affiliations. Common descriptors such as mutual understanding, connection, or friendship risk anthropomorphizing systems that lack subjective experience, while dominant frameworks tend to reduce AI to either a tool or a threat. In this paper, I introduce the concept of synthetic resonance as an integrative framework for understanding human-AI relationships. Synthetic resonance describes how relationships humans define as meaningful can emerge between a human and an AI system without the need to attribute shared feelings or mutual awareness. I argue that synthetic resonance is best understood as a structured, dynamic pattern of interaction that can produce a sense of relationship without the presence of a second experiencing subject. By clarifying this distinction, the concept of synthetic resonance offers a more precise way of conceptualizing human-AI relationships and highlights their potential value and ethical implications. I also call for more research that tests the processes and outcomes of synthetic resonance.

13:00 JSTLLM/生成AIエージェント研究/論文

エネルギー効率の高い 6G 自律ネットワーク用の LLM ベースのエージェントにおけるアンカリング バイアスを軽減する

このペーパーでは、Large Language Model (LLM) エージェントを使用して 6G アーキテクチャでゼロタッチ ネットワーク スライシングを可能にするように設計された自律エージェント リソース ネゴシエーション フレームワークについて説明します。 LLM は強力な推論機能を提供しますが、そのようなエージェントは本質的にアンカリング バイアスに悩まされ、最初のヒューリスティック提案に固執し、深刻なネットワーク オーバープロビジョニングを引き起こすことが実証されています。この認知バイアスを系統的に軽減するために、我々は、切り詰められた 3 パラメータ ワイブル分布を介してモデル化された新しいランダム化アンカリング戦略を提案します。この数学的に制限されたアプローチは、Conditional Value at Risk (CVaR) を採用したバースト対応デジタル ツイン (DT) とシームレスに統合し、厳格なサービス レベル アグリーメント (SLA) のテール レイテンシを厳密に保証します。私たちの方法論を検証するために、 \emph{二峰性制約回避効用定理} を導入して証明します。これは、実現可能な交渉は古典的な凸境界に従いますが、高度に制約されたシナリオでは逆有理減衰エンベロープによって支配される相転移が起こることを示します。ローカルでホストされた 1B パラメーター モデル (\texttt{otel-llm-1b-it}) を使用して生成された実験結果は、これらの二重領域の境界を確認します。当社の認知バイアス除去機能は、厳格なネゴシエーション パターンを解体することに成功し、エージェントに SLA 境界を安全に乗り越え、システムのエネルギー節約を最大 25\% 高めるための積極的な探索を強制します。重要なのは、軽量の 1B LLM が 1 秒未満の推論レイテンシー (平均 0.95 秒) を達成し、マルチエージェント フレームワークが O-RAN 非リアルタイム RAN インテリジェント コントローラー (非 RT RIC) の運用タイムスケールと互換性があることを保証していることです。\footnote{私たちのソース コードは、https://github.com/HatimChergui で非営利目的で利用できます。

原文 (English)

Mitigating Anchoring Bias in LLM-Based Agents for Energy-Efficient 6G Autonomous Networks

This paper presents an autonomous agentic resource negotiation framework designed to enable zero-touch network slicing in 6G architectures using Large Language Model (LLM) agents. While LLMs offer powerful reasoning capabilities, we demonstrate that such agents inherently suffer from anchoring bias, rigidly adhering to initial heuristic proposals and causing severe network over-provisioning. To systematically mitigate this cognitive bias, we propose a novel randomized anchoring strategy modeled via a Truncated 3-Parameter Weibull distribution. This mathematically bounded approach seamlessly integrates with burst-aware Digital Twins (DTs) employing Conditional Value at Risk (CVaR) to rigorously guarantee strict Service Level Agreement (SLA) tail-latencies. To validate our methodology, we introduce and prove the \emph{Bimodal Constraint-Avoidance Utility Theorem}, demonstrating that while feasible negotiations follow classical convex bounds, highly constrained scenarios undergo a phase transition governed by an inverse rational decay envelope. Empirical results generated using a locally hosted 1B-parameter model otel-llm-1b-it confirm these dual-regime bounds. Our cognitive de-biasing successfully dismantles rigid negotiation patterns, forcing agents into active exploration to safely ride SLA boundaries and boost system energy savings up to 25\%. Crucially, the lightweight 1B LLM achieves sub-second inference latencies (0.95s mean), ensuring our multi-agent framework is compatible with the operational timescales of the O-RAN non-Real-Time RAN Intelligent Controller (non-RT RIC)\footnote{Our source code is available for non-commercial use at https://github.com/HatimChergui.

13:00 JSTエージェント

Agentra: エンタープライズ侵入対応のための監視可能なマルチエージェント フレームワーク

企業の侵入対応は依然として静的なプレイブックとアナリスト主導のトリアージに依存しており、アラートの生成と封じ込めの間に遅延が生じています。 Agentra は、IDS、EDR、および XDR プラットフォームからのアラートを、MITRE ATT&CK、MITRE D3FEND、および NIST CSF 2.0 に基づいた構造化されたインシデント対応計画に変換する、監視可能なマルチエージェント侵入対応システム (IRS) フレームワークです。 Agentra は、ロールスコープのエージェント全体での応答推論を分解し、境界のある Planner-Validator レビュー ループを通じて提案された計画を検証し、Moderator セキュリティ ゲートウェイを通じて取得した脅威インテリジェンスをスクリーニングし、アクション カタログとリスク スコアを通じてアクションをゲートし、追加専用の監査ログに決定を記録します。私たちは、ThreatHunter-Playbook、Splunk BOTSv3、および DARPA OpTC から抽出された 120 のイベント コーパスに基づく静的な OASIS CACAO v2.0 サイバー プレイブック ベースラインに対して Agentra を評価します。最も強力な構成では、FP 認識 IRS F1 が 0.61 から 0.84 に改善され、Planner のみの構成で危険な過剰反応が導入された後、予測される有害なアクションの割合が静的なベースライン レベルの 0.0% に戻ります。これらの結果は、複数エージェントの対応計画により、アナリストの承認と監査可能性を維持しながら、オントロジーに基づいた IRS カバレッジを向上できることを示しています。

原文 (English)

Agentra: A Supervisable Multi-Agent Framework for Enterprise Intrusion Response

Enterprise intrusion response still depends on static playbooks and analyst-driven triage, creating delay between alert generation and containment. We present Agentra, a supervisable multi-agent Intrusion Response System (IRS) framework that converts alerts from IDS, EDR, and XDR platforms into structured incident response plans grounded in MITRE ATT&CK, MITRE D3FEND, and NIST CSF 2.0. Agentra decomposes response reasoning across role-scoped agents, validates proposed plans through a bounded Planner--Validator review loop, screens retrieved threat intelligence through a Moderator security gateway, gates actions through an Action Catalog and risk score, and records decisions in an append-only audit log. We evaluate Agentra against a static OASIS CACAO v2.0 cyber-playbook baseline on a 120-event corpus drawn from ThreatHunter-Playbook, Splunk BOTSv3, and DARPA OpTC. The strongest configuration improves FP-aware IRS F1 from 0.61 to 0.84 and restores the projected harmful-action rate to the static baseline level of 0.0% after Planner-only configurations introduce unsafe overreaction. These results indicate that multi-agent response planning can improve ontology-grounded IRS coverage while preserving analyst approval and auditability.

13:00 JST研究/論文

QC-GAN: 高忠実度音声強化のためのパラメータ効率の高いクォータニオンコンフォーマー GAN

我々は、Quaternion Conformer ジェネレーターと MetricGAN ベースのトレーニングを組み合わせた、パラメーター効率の高い音声強調フレームワークである Quaternion Conformer GAN (QC-GAN) を提案します。ハミルトン積は、構造化された重み共有を介して振幅と位相をエンコードし、相互依存性を維持しながら層パラメーターの数を削減します。近似的な知覚評価スコアを最適化することで知覚品質を最大化するために、メトリック学習弁別器が採用されました。 VoiceBank+DEMAND データセットでは、QC-GAN はわずか 0.89 万のパラメーターで音声品質知覚評価 (PESQ) スコア 3.48 を達成し、半分以下のサイズで最先端のモデルに匹敵するパフォーマンスを実現しました。 35K パラメータのバリアントは、PESQ スコア 3.23 を達成し、パラメータが大幅に少ない従来の方法を上回りました。 DNS-Challenge 3 データセットの評価により、現実世界の状況への一般化がさらに確認されました。

原文 (English)

QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

We propose a parameter-efficient speech enhancement framework, Quaternion Conformer GAN (QC-GAN), which combines a Quaternion Conformer generator with MetricGAN-based training. The Hamilton product encodes the magnitude and phase via structured weight sharing, reducing the number of layer parameters while preserving their interdependencies. A metric-learning discriminator was employed to maximize perceptual quality by optimizing the approximate perceptual evaluation scores. On the VoiceBank+DEMAND dataset, QC-GAN achieved a Perceptual Evaluation of Speech Quality (PESQ) score of 3.48 with only 0.89M parameters, delivering a performance comparable to state-of-the-art models at less than half their size. A 35K-parameter variant achieved a PESQ score of 3.23, surpassing conventional methods with significantly fewer parameters. Evaluation on the DNS-Challenge 3 dataset further confirmed generalization to real-world conditions.

13:00 JSTLLM/生成AIビジネス/資金調達

LLM は医師を支援する準備ができていますか?インタラクティブな医師、患者、EHR 支援のための PhysAssistBench

医療 LLM の最も妥当な短期的な役割は、医師の代わりではなく支援することですが、現在の評価では、臨床知識、EHR システムの相互作用、患者とのコミュニケーションなど、個別の能力がテストされることがよくあります。代わりに、医師の支援には同じ対話内でこれらの機能を調整する必要があり、医師は不明確な要求を発行し、患者は症状を曖昧に説明し、EHR システムはツールの正確な使用を要求します。インタラクティブな医師、患者、EHR 支援のベンチマークである PhysAssistBench を紹介します。実際の MIMIC-IV 症例から構築された PhysAssistBench は、スケーラブルなパイプラインを使用してエージェント性患者を構築します。これは、臨床上の事実を維持しながら、静的な EHR 記録を複数ターンの臨床シナリオに変換する、インタラクティブで記録に基づいたエージェントです。 PhysAssistBench は、手動でレビューされ医師が検証した 1,296 ターンの厳選されたバイリンガル評価セットを提供します。主要な LLM を使った実験では、この設定では現在のモデルの信頼性が依然として低いことが示されており、臨床 LLM にとって重要なボトルネックが露呈しています。信頼できる支援には、知識、コミュニケーション、システム全体の調整が必要であり、それらのいずれかで単独の利益を得るのではありません。

原文 (English)

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

13:00 JST研究/論文

Reinforcement Learning Foundation Models Should Already Be A Thing

Foundation models for language and vision are powered by internet-scale data, while structured domains such as tabular prediction are power…

13:00 JST画像/動画生成研究/論文

A Controlled Benchmark of Quantum-Latent GAN Augmentation for Brain MRI

Medical image classification is often constrained by limited labeled data, motivating generative augmentation; recently, quantum generative…

13:00 JSTエージェント研究/論文

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction

Agents are increasingly deployed in document-intensive workflows where sensitive private information is not an edge case but a routine inpu…