Skip to the content.

AIニュース 2026-06-25

自動生成: 2026-06-25 13:03 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. OpenAI and Broadcom unveil LLM-optimized inference chipOpenAI

    OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LL…

  2. Google、「Gemini 3.5 Flash」に「Computer Use」を標準搭載──AIが画面を見てブラウザやアプリを操作ITmedia AI+

    Googleは、AIモデル「Gemini 3.5 Flash」に、AIがコンピュータの画面を認識してマウス操作やキーボード入力を自動で実行…

  3. OpenAI、「GPT-5.5 Instant」をアップデート 会話の文脈維持や箇条書き減など「読みやすさ」を改善ITmedia AI+

    OpenAIは、ChatGPTで最も広く利用されているモデル「GPT-5.5 Instant」のアップデートを発表した。新機能の追加ではな…

  4. 富士通と日本IBMの協業、ついに始動 COBOL刷新における「役割分担」は?ITmedia AI+

    レガシーシステムをどうモダナイズするかは、多くの企業における課題だ。富士通と日本IBMがこの領域での協業を発表した。ついに始動する、両社の…

  5. OpenAI、初の独自AIチップ「Jalapeno」発表──Broadcomと共同開発の推論用アクセラレータITmedia AI+

    OpenAIは、Broadcomと共同開発したLLMの推論に特化した独自AIチップ「Jalapeno」を発表した。同社初の「インテリジェン…

  6. Facebook rolls out an AI companion app for creatorsTechCrunch AI

    The new app, which is currently being tested with select creators, wi…

  7. 「今日言うつもりはなかったが……」 孫正義氏が明かした「ロボット自動量産工場」の実態ITmedia AI+

    「今日ここで言うつもりはなかったんですが」──。ソフトバンクグループが6月24日に開催した株主総会の質疑応答で、会長兼社長の孫正義氏が、投…

トピック別件数

日本語メディア16件

ITmedia AI+ (日本語)

11:00 JSTロボティクスビジネス/資金調達

「今日言うつもりはなかったが……」 孫正義氏が明かした「ロボット自動量産工場」の実態

「今日ここで言うつもりはなかったんですが」──。ソフトバンクグループが6月24日に開催した株主総会の質疑応答で、会長兼社長の孫正義氏が、投資先の現場で起きている現場実態を明かす一幕があった。

10:08 JSTLLM/生成AIGoogleGemini2媒体が報道

Google、「Gemini 3.5 Flash」に「Computer Use」を標準搭載──AIが画面を見てブラウザやアプリを操作

Googleは、AIモデル「Gemini 3.5 Flash」に、AIがコンピュータの画面を認識してマウス操作やキーボード入力を自動で実行する「Computer Use」機能を標準ツールとして搭載したと発表した。これまで専用モデルでのみ提供していた機能を主力モデルに統合したもの…

出典:Google DeepMindITmedia AI+
09:23 JSTLLM/生成AIOpenAIGPT / ChatGPT

OpenAI、「GPT-5.5 Instant」をアップデート 会話の文脈維持や箇条書き減など「読みやすさ」を改善

OpenAIは、ChatGPTで最も広く利用されているモデル「GPT-5.5 Instant」のアップデートを発表した。新機能の追加ではなく、日常的な会話の品質向上が中心。質問の意図を的確に捉えて文脈を維持する能力が向上するほか、テンプレート的な回答が減り、位置情報を活用した地…

08:00 JSTその他

富士通と日本IBMの協業、ついに始動 COBOL刷新における「役割分担」は?

レガシーシステムをどうモダナイズするかは、多くの企業における課題だ。富士通と日本IBMがこの領域での協業を発表した。ついに始動する、両社の協業における役割分担とは。

08:00 JSTその他Google

【役に立つの?】「Google公式」の初心者向けAI講座、受けてみたら想像以上にすごかった

1日で1万人以上が登録したGoogleの初心者向けAI講座を実際に体験。想像以上の学びとは?

08:00 JSTLLM/生成AI

慶応大がNotionを選んだ「3つの理由」 “何ができるか”以外の決め手は?

慶應義塾大学が全教職員にNotionを導入し、「AIキャンパス構想」を本格始動した。数あるツールからNotionを選んだ理由は何か。また、同大が目指す、生成AI時代におけるナレッジ管理の形とは。

07:38 JSTLLM/生成AIOpenAI

OpenAI、初の独自AIチップ「Jalapeno」発表──Broadcomと共同開発の推論用アクセラレータ

OpenAIは、Broadcomと共同開発したLLMの推論に特化した独自AIチップ「Jalapeno」を発表した。同社初の「インテリジェンスプロセッサ」と位置付けており、設計から製造用のテープアウトまでを9カ月で完了したという。数世代にわたる計算基盤の第1弾として、年内に展開を…

07:00 JSTその他

味の素、“万能DX人材”増員へ 育成のきっかけは新規プロジェクトの苦い経験

味の素が、AIを活用して「フルスタック人財」の育成を目指している。DX推進のキーパーソンにする狙いだ。既に300時間の工数削減例も生まれている。一体どう育成したのか。

05:00 JSTLLM/生成AI2件の関連記事

「AI教育、どこから手を付ける?」 全社導入のカギは“生成AIリテラシー向上研修”(前編)

生成AIは業務の現場に急速に浸透し、「使って当たり前」の時代が到来しています。その活用範囲は広がる一方、情報漏洩や誤情報のリスクが企業の大きな課題になっています。今求められるのは、誰もが“安全かつ賢く”生成AIを使いこなすリテラシーです。本稿は、社内の誰もが生成AIを安全に、自…

出典:ITmedia AI+ITmedia AI+
05:00 JSTその他Microsoft

「最強モデル」はもう無意味 ナデラCEOが語る、企業の生き残り新戦略「学習ループ」

AIの進化で、自社システムの模倣やコモディティ化への不安が広がっている。MicrosoftのナデラCEOが示す「学習ループ」戦略とは何か。日本のソフトウェア企業の生き残りにも通じる筆者の視点を交えて解説する。

19:54 JSTLLM/生成AIハードウェア/半導体

「Transformerの最大475倍」 富士通、GPUを効率的に使うLLMアーキテクチャ「PHOTON」開発

富士通が、大規模言語モデル(LLM)を少ないGPUで動かせる新アーキテクチャ「PHOTON」(フォトン)を開発した。GPU当たりの処理性能(スループット)が、現在のLLMで主流のアーキテクチャ「Transformer」の最大475倍に達するという。LLMの運用に必要なGPUを抑…

18:25 JSTロボティクス

陸自駐屯地で四足歩行型の警備用ロボットが見回り GMOインターネットグループが開発

GMOインターネットグループ4社は、国産ロボット開発を担う未来ロボットと組み、国産の四足歩行型警備用ロボットを開発し、陸上自衛隊の駐屯地での導入検証を始めると発表した。警備の省人化を図り、24時間警備体制の実現を目指すという。

16:25 JSTロボティクス

民生VRグローブにロボット業界が注目 日本発ベンチャーがB2B加速

Diver-Xは2026年6月23日、グローブ型仮想現実(VR)コントローラー新製品「ContactGlove3」を発表した。電磁場トラッキング方式の採用により推奨環境下で中央値0.5mm、最大値1.5mmの誤差という高精度を実現した。民生用と業務用を用意していて、ロボティクス…

15:40 JSTハードウェア/半導体

【解説】キオクシアなぜ急成長? 半導体メモリって何? AIブームを見通すための基礎知識

注目を集める半導体メモリ大手のキオクシア。同社はなぜAI需要を取り込めたのか。いま押さえたい基礎知識を解説する。

15:27 JSTLLM/生成AIAnthropicClaude

ClaudeをSlackチャンネルに召喚、“チームの一員”として直接指示 新機能「Claude Tag」登場

Anthropicは、「Slack」上で「Claude」を利用できる新機能「Claude Tag」のベータ版提供を23日から開始した。指定したチャンネルにClaudeを招待し、メンションをつけてタスクを依頼すると、Claudeが関連情報をもとにタスク計画を自動構築する。

15:24 JSTエージェント

国内ユーザー数「前年比1582%増」――AI開発支援「Devin」は競合と何が違う? 日本法人代表が語る事業戦略

ソフトウェア開発向けAIエージェント「Devin」をどのように日本市場で展開するのか。Devinを手掛ける米Cognition AI日本法人の正井拓己代表が語った。

海外メディア12件

TechCrunch AI (英語)

09:08 JSTハードウェア/半導体

Europe is pushing back on Washington’s chip war

As ASML CEO Christophe Fouquet told TechCrunch in May, what China can currently buy are older-generation deep ultraviolet tools — gear firs…

08:26 JSTその他

Former Infosys chief has a new startup that wants to challenge the IT services world

Backed by Mayfield and Aramco Ventures, Vishal Sikka’s new venture brings together veterans from SAP, Infosys, and VianAI.

07:41 JSTハードウェア/半導体

Cerebras stock plunges after earnings as CEO says margin outlook was misunderstood

In its first earnings report since going public, the AI chipmaker forecast a narrower gross margin in its core business, scaring investors.

06:56 JSTその他

AI was supposed to kill engineering jobs, but new data suggests they’re the most resilient

While AI dominates the layoff narrative, engineers are actually making up a larger share of new hires, according to SignalFire data.

06:42 JSTLLM/生成AI研究/論文AnthropicGoogle

AI researchers continue to leave Google for its rivals

Top AI researchers Jonas Adler and Alexander Pritzel are leaving Google for Anthropic, following departures from top scientists Noam Shazee…

06:30 JSTハードウェア/半導体

The memory chip crunch is paying off for this US company

Revenue quadrupled to $41.45 billion compared with the same period a year ago. The company's profit, meanwhile, rose from $1.88 billion to…

05:09 JSTその他

Companies are scrambling to stop employees from maxing out AI budgets with small tasks

The tokenmaxxing era was brief. We now appear to be entering the era of token rationing.

02:16 JSTその他Meta

Facebook rolls out an AI companion app for creators

The new app, which is currently being tested with select creators, will have Facebook's recently launched AI creator assistant built into i…

01:48 JSTロボティクス

Agility Robotics plans to go public via SPAC in a $2.5B deal

Agility Robotics, the humanoid robotics startup that spun out of Oregon State University in 2015, expects to generate $620 million in proce…

01:15 JSTその他

Figma adds code layers, support for animations, more AI features in new update

Figma's update adds a new code layer, support for motion and shaders, and the ability to create custom plug-ins for various tasks using AI.

23:54 JSTLLM/生成AIハードウェア/半導体OpenAI

OpenAI unveils its first custom chip, built by Broadcom

Named Jalapeño, the new processor was designed specifically for the unique needs of OpenAI's inference systems.

23:00 JSTその他

3 days left to save up to $190 on your TechCrunch Founder Summit 2026 pass

You have just 3 days left to save up to $190 on your pass to TechCrunch Founder Summit 2026 before Early Bird rates end on June 26 at 11:59…

公式ブログ1件

OpenAI (英語)

15:00 JSTLLM/生成AIハードウェア/半導体OpenAI

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI sy…

Google DeepMind (英語)

新着記事はありませんでした。

論文299件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント

RIFT-Bench: エージェントティック AI システム向けの動的なレッドチーム化

大規模言語モデル (LLM) を利用したエージェント AI システムは、自律的な意思決定システムへと急速に進化しており、従来の LLM の脆弱性を超えた攻撃ベクトルをさらしています。既存のセキュリティ評価は多くの場合、特定の実装またはドメインに関連付けられており、異種システム間での統一された比較が制限されています。このギャップに対処するために、RIFT-Bench を導入します。これは、多様なエージェント アーキテクチャにわたって統一された評価を可能にする、動的なレッド チーム化のためのグラフ表現主導の方法論です。新しい階層表現に基づいて、RIFT-Bench は 2 つの自動化フェーズで動作します。システム構造を抽出するディスカバリと、適応型敵対的攻撃を展開して包括的な評価レポートを作成するスキャンです。さまざまな攻撃ベクトルや目的にわたって、動的に適応可能な幅広い敵対的プローブを利用して、調査されたシステム自体を評価します。私たちは、さまざまな実装範囲にわたる 45 のエージェント システムにわたって、提案された評価パイプラインの有効性を実証し、このアプローチが異種エージェント アーキテクチャに効果的に一般化できることを示しています。システムや攻撃を超えて、RIFT-Bench は緩和戦略の直接評価もサポートします。これらの主要な機能により、RIFT-Bench はエージェント AI システムのセキュリティ評価のためのスケーラブルな基盤となります。

原文 (English)

RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems. To address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agentic architectures. Building on a novel hierarchical representation, RIFT-Bench operates in two automated phases: Discovery, which extracts system structure, and Scanning, which deploys adaptive adversarial attacks and produces a comprehensive evaluation report. It evaluates the examined system itself, leveraging a broad set of dynamically adaptable adversarial probes across diverse attack vectors and objectives. We demonstrate the effectiveness of the proposed evaluation pipeline across 45 agentic systems spanning a diverse range of implementations, showing that the approach generalizes effectively to heterogeneous agentic architectures. Beyond systems and attacks, RIFT-Bench also supports direct evaluation of mitigation strategies. These key capabilities make RIFT-Bench a scalable foundation for security evaluation of agentic AI systems.

13:00 JSTLLM/生成AI

神経記号的ドライブ: VLA を駆動するためのルールに基づいた忠実な推論

思考連鎖 (CoT) 推論を組み込んだ VLA モデルの駆動は、事前トレーニング済みの VLM 表現を活用し、中間決定を自然言語で公開するため魅力的ですが、現在の理論的根拠には、理論的根拠を計画された動作と因果関係を保つために必要な段階的な決定セマンティクスが欠けていることがよくあります。 Neuro-Symbolic Drive は、古典的なルールベースのプランナーから直接抽出されたルールに基づいた推論トレースを使用して駆動 VLA を監視するニューロシンボリック駆動フレームワークです。私たちの重要な観察は、ルールベースのプランナーは、実行可能な推論エンジンとしてすでに機能している象徴的な AI システムであるということです。これらは、アクティブな安全性の制約について推論し、候補となる操縦を検索し、最終的な軌道を選択します。これらのプランナーをシミュレーションで計測して、ルール評価の各ステップで実行された軌跡と内部決定トレースの両方をキャプチャします。各トレースは、構造化されたルールに基づいた推論にシリアル化され、軌道と組み合わせて、駆動 VLA として Qwen3.5-4B を微調整します。これらのトレースは、アクションを決定するプランナーの状態から直接導出されるため、事後的な調整ではなく、構築によって推論がモーション生成に構造的に結合されることが保証されます。シミュレーターで生成されたベンチマークでは、ルールに基づいた詳細な推論により、3 台のカメラの認識では ADE@3s が 0.47 から 0.26 に、ミス率が 8.30% から 6.40% に減少し、8 台のカメラの認識では 0.54 から 0.26 に、10.13% から 5.99% に減少しました。したがって、Neuro-Symbolic Drive は、神経記号的な計画ロジックを構造化された監視に変換します。コードベース: https://github.com/XiangboGaoBarry/Neural-Symbolic-Drive。

原文 (English)

Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs

Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose intermediate decisions in natural language, yet current rationales often lack the step-by-step decision semantics needed to keep the rationale causally connected to the planned motion. We introduce Neuro-Symbolic Drive, a neuro-symbolic driving framework that supervises a driving VLA with rule-grounded reasoning traces extracted directly from classical rule-based planners. Our key observation is that rule-based planners are symbolic AI systems that already function as executable reasoning engines: they reason about active safety constraints, search over candidate maneuvers, and select a final trajectory. We instrument these planners in simulation to capture both the executed trajectory and the internal decision trace at each rule-evaluation step. Each trace is serialized into structured rule-grounded reasoning and paired with the trajectory to fine-tune Qwen3.5-4B as a driving VLA. Because these traces are derived directly from the planner states that determine the action, they ensure reasoning is structurally coupled to motion generation by construction, rather than by post-hoc alignment. On our simulator-generated benchmark, detailed rule-grounded reasoning reduces ADE@3s from 0.47 to 0.26 and miss rate from 8.30% to 6.40% under three-camera perception, and from 0.54 to 0.26 and 10.13% to 5.99% under eight-camera perception. Neuro-Symbolic Drive thus converts neuro-symbolic planning logic into structured supervision. Code base: https://github.com/XiangboGaoBarry/Neural-Symbolic-Drive.

13:00 JSTLLM/生成AIエージェントロボティクス

エージェントモデルの批判

エージェントとは何ですか?代理店とは何ですか? 「コーディング エージェント」、「AI 共同科学者」、および生産性の向上を約束するその他の「エージェント」ツールとして販売されるラージ言語モデル (LLM) システムの台頭、そして同時に、人間に対する投機的な「マシン エージェント」の下で AI が破壊的な力で人間の制御から逃れるなどの「実存的」な懸念により、有能なシステムを構築するためと、恐れるべきかどうか、何を恐れるべきかを理解するために、自動化がどこで終わり、エージェントが始まるのかを明確にすることが不可欠になっています。デカルトの独立した思考における主体性の根拠と、SF における自律的存在の描写に基づいて、AI エージェントの現状を概観し、目標、アイデンティティ、意思決定、自己規制、学習という 5 つの側面に沿ってエージェントのアーキテクチャを分析します。具体的には、真の主体性は、これらの構造が外部の足場を介して組み立てられるのではなく、\emph{エージェント} システムの能力が設計されたものに存在することを必要とすると主張します。ワークフローと \emph{エージェント} システムは、その機能 (ソーシャル インタラクションを含む) が内生的に生じ、所定のタスク用に設計されたシステムと、オープンワールドで真の自律性を持って動作できるシステムとの間の境界を定義します。この分析に基づいて、階層的な目標分解、アイデンティティ進化、個別に学習された世界モデルに基づくシミュレーション推論を組み合わせた、汎用エージェント モデルの Goal-Identity-Configurator (GIC) アーキテクチャを提案します。さらに、私たちは、より優れた自律性と「エージェンシー」を持ちながらも人間の監視下にあるエージェント システムの監査可能性、制御可能性、安全性についての洞察を共有します。

原文 (English)

Critique of Agent Model

What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and other ``agentic" tools that promise to drive up productivity, and at the same time, ``existential" concerns such as AI escaping human control with destructive power under a speculative ``machine agency" against humans, it has become essential to clarify where automation ends and agency begins, both for building capable systems and for understanding whether and what to fear. Drawing on Descartes' grounding of agency in independent thought, and on portrayals of autonomous beings in science fiction, we survey the current landscape of AI agents, and analyze agent architectures along five dimensions: goal, identity, decision-making, self-regulation, and learning. Specifically, we argue that genuine agency requires these structures to be \emph{internalized within the system itself} rather than assembled through external scaffolding. This distinction between \emph{agentic} systems, whose competence resides in engineered workflows, and \emph{agentive} systems, whose capabilities (including social interaction) arise endogenously, defines the boundary between systems designed for prescribed tasks, and those capable of operating in the open world with true autonomy. Building on this analysis, we propose the Goal-Identity-Configurator (GIC) architecture for a general-purpose agent model, combining hierarchical goal decomposition, identity evolution, simulative reasoning grounded in a separately trained world model, learned self-regulation, and self-directed learning from both real and simulated experience. Furthermore, we share insight on the auditability, controllability, and safety of agentive systems that possess greater autonomy and ``agency", but remain under human oversight.

13:00 JSTエージェント

制約マニホールド制御による安全で一般化可能な階層型マルチエージェント RL

マルチエージェント システムは、厳しい安全制約の下で調整された動作を必要とするセーフティ クリティカルなアプリケーションで広く使用されています。既存のアプローチは根本的なトレードオフに直面しています。学習ベースの手法は強力な経験的パフォーマンスを達成しますが、理論的な安全性が保証されていません。一方、制御理論的な手法は安全性を強化しますが、過度に保守的で非効率な動作につながることがよくあります。我々は、高レベルのポリシー学習を通じて効果的な調整を可能にしながら、制約マニホールドを介して低レベルのマイルドな仮定の下で厳しい安全制約を強制する、階層型マルチエージェント強化学習フレームワークを提案します。私たちのアプローチは、マルチエージェント設定で理論的な安全性を保証し、定常的な学習ダイナミクスを生み出すため、安定した効率的なトレーニングを可能にします。経験的に、私たちの方法はほぼ完璧な安全率を維持しながら競争力のあるパフォーマンスを達成し、さまざまな数のエージェントと障害物に効果的に一般化します。

原文 (English)

Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approaches face a fundamental trade-off: learning-based methods achieve strong empirical performance but lack theoretical safety guarantees, while control-theoretic methods enforce safety but often lead to overly conservative and inefficient behaviors. We propose a hierarchical multi-agent reinforcement learning framework that enforces hard safety constraints under mild assumptions at low level via a constraint manifold, while enabling effective coordination through high-level policy learning. Our approach provides theoretical safety guarantees in the multi-agent setting and yields stationary learning dynamics, thereby enabling stable and efficient training. Empirically, our method achieves competitive performance while maintaining nearly perfect safety rates, and generalizes effectively to varying numbers of agents and obstacles.

13:00 JSTLLM/生成AI

広範囲かつ永続的に有益なモデルを目指した強化学習

AI システムがますます多様化し、一か八かの環境に展開されるにつれて、モデルの調整はトレーニング中に見られるタスクや領域を超えて一般化する必要があります。これは、報酬のハッキング、欺瞞、その他の意図しない戦略によって予期せぬ不整合が生じる可能性がある強化学習 (RL) にとって特に重要です。私たちは、現実的なドメインでインスタンス化された有益な動作に関する RL が、トレーニング分布を超えた広範かつ永続的なアラインメント一般化を生成できるかどうかを研究します。私たちは、健康、科学、教育などのさまざまな領域にまたがる、真実性、公平性、リスク認識、正しさなどの有益な特性を測定および訓練するように設計された現実的な状況のデータセットを構築します。次に、このデータセットで RL を使用してモデルをトレーニングし、整合性と有益な動作に関する 50 を超える独立したベンチマークでモデルを評価します。コンピューティング一致ベースラインと比較して、有益な特性 RL は、これらの配布外ベンチマークの 80% 以上でパフォーマンスを向上させます。私たちは、実質的な分布外のアラインメントの転移を観察しています。つまり、健康という 1 つの領域に完全に限定された有益な行動の RL 介入は、報酬のハッキング、欺瞞、一般的なミスアラインメントの削減など、健康以外のアラインメント評価に広範な改善をもたらします。最後に、アライメントの持続性、つまりモデルを不整合に向けて誘導しようとする試みの下で動作がロバストにアライメントされたままであるかどうかを研究します。有益な特性 RL でトレーニングされたモデルは、敵対的なプロンプトや有害な微調整に対する優れた耐性など、持続性の向上を示します。これらの影響の原因を特定するには、さらなる研究が必要です。これらの結果は、現実的な領域で有益な行動を強化するための RL が、人間の繁栄とより強固に一致するモデルを生成できることを示唆しています。

原文 (English)

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through reward hacking, deception, or other unintended strategies. We study whether RL on beneficial behavior, instantiated in realistic domains, can produce broad and persistent alignment generalization beyond the training distribution. We construct a dataset of realistic situations designed to measure and train beneficial traits, such as truthfulness, fairness, risk awareness, and corrigibility, spanning varied domains, including health, science, and education. We then train models with RL on this dataset and evaluate them on more than 50 independent benchmarks of alignment and beneficial behavior. Compared to a compute-matched baseline, beneficial trait RL improves performance on over 80% of these out-of-distribution benchmarks. We observe substantial out-of-distribution alignment transfer: a beneficial-behavior RL intervention entirely limited to one domain, health, produces broad improvements on non-health alignment evaluations, including reduced reward hacking, deception, and general misalignment. Finally, we study alignment persistence: whether behavior remains robustly aligned under attempts to steer models towards misalignment. Models trained with beneficial trait RL show improved persistence, including greater resistance to adversarial prompting and harmful finetuning; further work is required to isolate the sources of these effects. These results suggest that RL to reinforce beneficial behavior in realistic domains can produce models that are more robustly aligned with human flourishing.

13:00 JSTエージェントLlama

言語モデルエージェントは機械的解釈において回路の説明者として役立つでしょうか?

機構の解釈可能性は、回路の自動ローカライズにおいて大幅な進歩を遂げましたが、ローカライズされたコンポーネントが何を行うかを説明することは依然として労力を要し、標準化が困難です。この研究では、回路がすでに特定されている場合、言語モデル (LM) エージェントがこの説明問題を支援できるかどうかを研究します。ここでは、163 のコンポーネント レベルの注釈を備えた 84 の半合成トランス回路から構築された回路説明用のベンチマークである AgenticInterpBench を紹介します。我々は、観察、仮説生成、因果関係検証の反復ループを通じて各コンポーネントを分析し、最終的にコンポーネントレベルの説明と回路レベルのタスク記述を生成するエージェント的説明ツールである HyVE (仮説、検証、説明) を提案します。 HyVE は 4 つの LM バックボーンにわたって、有用なコンポーネント レベルおよびタスク レベルの説明を復元しますが、一律に最適なバックボーンはありません。私たちの分析によると、強力なバックボーンは通常、観察に基づいた仮説を形成しますが、失敗は検証ループの後半で、不完全な検証計画、コード実行エラー、または未解決の仮説によって発生することが多いことがわかりました。 Llama-3-8B の算術回路に関するケース スタディでは、同じ定式化が半合成ベンチマークを超えて自然にトレーニングされたモデルに拡張できることが示されています。全体として、LM エージェントは回路の説明に有望ですが、信頼性の高い検証が依然として主要な障害となっています。

原文 (English)

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) agents can assist with this explanation problem once a circuit has already been identified. We introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circuits with 163 component-level annotations. We propose HyVE (Hypothesize, Validate, Explain), an agentic explainer that analyzes each component through an iterative loop of observation, hypothesis generation, and causal validation, eventually producing a component-level explanation and a circuit-level task description. Across four LM backbones, HyVE recovers useful component- and task-level explanations, but no backbone is uniformly best. Our analysis shows that strong backbones usually form observation-grounded hypotheses, while failures more often arise later in the validation loop, through incomplete validation plans, code execution errors, or unresolved hypotheses. A case study on an arithmetic circuit in Llama-3-8B shows that the same formulation can extend beyond semi-synthetic benchmarks to naturally trained models. Overall, LM agents are promising circuit explainers, but reliable validation remains the key obstacle.

13:00 JST研究/論文

フィルターバブルの打破: 多目的レコメンデーションのためのセマンティック パレート DQN フレームワーク

レコメンダー システムは、即時のユーザー エンゲージメントをモノリシックに最適化することでフィルター バブルとセマンティックな均質化を引き起こすことがよくあります。従来の Deep Q-Networks を含む標準的な単一目的モデルは、プラットフォームの維持と、情報の多様性やプロバイダーの公平性などの重要な社会的価値との間のトレードオフをうまく乗り切ることができません。これらの制限に対処するために、推奨を意味論的な多目的マルコフ決定プロセスとして形式化する多目的強化学習フレームワークを導入します。高忠実度のセマンティック埋め込みをパレート DQN エージェントと統合することにより、私たちのアーキテクチャはエンゲージメント、多様性、公平性を個別の集約不可能な報酬信号として扱い、静的な報酬のスカラー化の落とし穴を回避します。 MovieLens の小規模データセットに対する実験的評価では、ハイパーボリューム ベースのアクション選択により、セマンティック崩壊の原因となるフィードバック ループが破壊されることが示されています。高い状態軌道分散を維持することにより、パレート DQN はパレート フロンティアを効果的にマッピングし、エンゲージメントにわずかな影響を与えるだけで補助的な社会目標の向上を達成します。この取り組みは、本質的に調整された責任ある推奨システムへの道を提供します。

原文 (English)

Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation

Recommender systems often induce filter bubbles and semantic homogenization by monolithically optimizing for immediate user engagement. Standard single-objective models, including traditional Deep Q-Networks, are ill-equipped to navigate the trade-offs between platform retention and critical societal values like information diversity and provider fairness. To address these limitations, we introduce a multi-objective reinforcement learning framework that formalizes recommendation as a semantic multi-objective Markov decision process. By integrating high-fidelity semantic embeddings with a Pareto-DQN agent, our architecture treats engagement, diversity, and fairness as distinct, non-aggregable reward signals, avoiding the pitfalls of static reward scalarization. Empirical evaluations on the MovieLens small dataset shows that our hypervolume based action selection disrupts the feedback loops responsible for semantic collapse. By sustaining high state-trajectory variance, the Pareto-DQN effectively maps the Pareto frontier, achieving gains in auxiliary societal objectives with only marginal impacts on engagement. This work provides a path toward intrinsically aligned, responsible recommender systems.

13:00 JST研究/論文

女性セックスワーカーの説明可能なメンタルヘルスリスク予測のためのアンサンブル特徴選択とハリスホークス最適化

女性セックスワーカー (FSW) に影響を与える重大なメンタルヘルス問題の 1 つは、精神障害、特にうつ病です。暴力、偏見、経済的困難にさらされると、心理的リスクがさらに高まります。現在の機械学習 (ML) モデルは通常、この疎外されたグループに存在する高次元で複雑なリスク パターンを捉えるのには効果的ではありません。この論文では、ANOVA と相互情報量を使用したアンサンブル特徴選択戦略と、ハリス ホークス最適化調整ロジスティック回帰を組み合わせたハイブリッド予測モデルを提案し、脆弱なグループの精神的健康を予測するための群知能の新しいアプリケーションを示します。 Explainable AI (XAI) メソッドは、モデル予測に関連するトラウマの要因を理解するために使用できます。 3,005 人の FSW のグループに適用した場合、提案されたモデルは従来の分類器よりも効果的であり、精度 95.78%、F1 スコア 95.77%、AUC 0.96 であり、心的外傷後ストレス、クライアント関連の暴力、および職業的要因をうつ病の主な要因として特定していることがわかります。この取り組みは、従来のアプローチと ML アプローチの間のギャップを埋めて、脆弱なグループが早期支援、証拠に基づいた対象を絞った心理社会的ケア、健康計画を受けられるようにする XAI ツールを開発します。

原文 (English)

Ensemble Feature Selection and Harris Hawks Optimization for Explainable Mental Health Risk Prediction in Female Sex Workers

One of the significant mental health issues affecting female sex workers (FSWs) is mental disorders, especially depression. Exposure to violence, stigma, and economic hardship further increases their psychological risk. Current machine learning (ML) models are typically ineffective at capturing the high-dimensional and complex risk patterns that exist in this marginalized group. This paper suggests a hybrid predictive model that merges an ensemble feature selection strategy using ANOVA and mutual information and Harris Hawks optimization-tuned logistic regression and represents a new application of swarm intelligence to predict mental health in vulnerable groups. The explainable AI (XAI) methods can be used to understand the factors of trauma associated with model predictions. When applied to a group of 3,005 FSWs, it can be seen that the proposed model is more effective than traditional classifiers, with an accuracy of 95.78%, an F1 score of 95.77%, and an AUC of 0.96, and identifying post-traumatic stress, client-related violence, and occupational factors as major contributors to depression. This work bridges the gaps between conventional and ML approaches to develop an XAI tool that enables vulnerable groups to receive early assistance, evidence-based targeted psychosocial care, and health planning.

13:00 JSTLLM/生成AI

軌道の模倣を超えて: LLM 推論のための戦略に基づくポリシーの最適化

強力な言語モデルから弱い言語モデルへの推論機能を抽出するには、通常、特定の解決策の軌跡を模倣し、どのように推論するかではなく何を答えるかを効果的に転送することが含まれます。この軌跡レベルの模倣は、応用可能な問題解決スキルの習得ではなく、インスタンス固有のステップの暗記を促進し、新しい問題への一般化を制限します。私たちは、インスタンスレベルの軌跡の模倣を再利用可能な戦略の蒸留に置き換える、戦略に基づくポリシーの最適化 (SGPO) を提案します。 SGPO は、強力なモデルの応答から構造化された戦略の説明を抽出し、問題ごとに自律的な軌道と戦略に基づく軌道の両方を構築して、戦略的ガイダンスの有無にかかわらずモデルの動作を直接比較できるようにします。次に、このフレームワークは 2 つの重要な質問に対処します。抽出方法については、トークンレベルのフォワード KL 目標により、安定性を確保する近位制約を使用して、戦略条件付けによって引き起こされる分布シフトをガイドなしポリシーに選択的に転送します。いつ抽出するかについては、適応型インスタンス レベルの重み付けにより、自律探索が不十分な場合のガイダンスが強化され、モデル自体の能力が向上するにつれてガイダンスが軽減されます。 2 つのモデル ファミリにわたる 4 つの数学的ベンチマークの実験では、SGPO が SFT、オンポリシー RL、およびハイブリッド ポリシーのベースラインを常に上回っており、Qwen2.5-7B-Instruct の最も強力なベースラインよりも平均スコアを 2.2 ポイント改善していることが示されています。分析の結果、フォワード KL 対物レンズは、直接軌道の模倣を上回る本質的に選択的な蒸留信号を提供し、戦略蒸留が基本モデルの機能と相補的なスケーリングを示すことが明らかになりました。

原文 (English)

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason. This trajectory-level imitation encourages memorization of instance-specific steps rather than acquisition of transferable problem-solving skills, limiting generalization to novel problems. We propose Strategy-Guided Policy Optimization (SGPO), which replaces instance-level trajectory imitation with reusable strategy distillation. SGPO extracts structured strategy descriptions from strong-model responses and, for each problem, constructs both autonomous and strategy-guided trajectories to enable direct comparison of the model's behavior with and without strategic guidance. The framework then addresses two key questions. For how to distill, a token-level forward-KL objective selectively transfers the distributional shift induced by strategy conditioning into the unguided policy, with proximal constraints ensuring stability. For when to distill, adaptive instance-level weighting strengthens guidance when autonomous exploration falls short and reduces it as the model's own competence grows. Experiments on four mathematical benchmarks across two model families show that SGPO consistently outperforms SFT, on-policy RL, and hybrid-policy baselines, improving the average score by 2.2 points over the strongest baseline on Qwen2.5-7B-Instruct. Analysis reveals that the forward-KL objective provides an inherently selective distillation signal that outperforms direct trajectory imitation, and that strategy distillation exhibits complementary scaling with base model capability.

13:00 JSTLLM/生成AI研究/論文

学術論文全文に基づく共起ネットワークによるアルゴリズムの学術的影響の探求

アルゴリズムは、人工知能 (AI) の時代の科学研究の中心となっています。論文でのアルゴリズムの言及は、人気や影響力を示すためによく使用されますが、既存の研究では通常、個々のアルゴリズムを個別に評価し、相互接続を通じて形成される集合的な影響には限定的に注意を払っています。本研究では、学術論文全文に基づいて自然言語処理(NLP)における大規模なアルゴリズム共起ネットワークを構築し、ネットワークの観点からアルゴリズムの影響を調査します。深層学習モデルを使用して、アルゴリズム エンティティを抽出し、全体的、累積的、年次共起ネットワークを構築します。私たちはそれらの構造的特徴を分析し、複数の中心性測定を適用して、分野全体および長期にわたるアルゴリズムのグループへの影響を評価します。結果は、アルゴリズム ネットワークが複雑なネットワークの典型的な特徴を示しており、約 20 年間にわたって接続の密度が増加していることを示しています。古典的で高性能なアルゴリズムや、さまざまな研究期間の交差点に位置するアルゴリズムは、高い人気、制御性、中心性、バランスのとれた影響力を持つ傾向があります。アルゴリズムの影響力が低下すると、通常、まずそのアルゴリズムがコア ネットワークでの地位を失い、続いて他のアルゴリズムとの関連性が弱まります。この研究は、アルゴリズム共起ネットワークの最初の大規模な分析です。 40 年以上の学術出版物を網羅しており、アルゴリズムの影響を時間的かつ構造的に把握し、アルゴリズム、学者、タスクを結び付けるネットワークに関する将来の研究の基盤を提供します。

原文 (English)

Exploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Papers

Algorithms have become central to scientific research in the era of artificial intelligence (AI). Although algorithm mentions in papers are often used to indicate popularity and influence, existing studies usually evaluate individual algorithms in isolation and pay limited attention to the collective influence formed through their interconnections. This study constructs large-scale algorithm co-occurrence networks in natural language processing (NLP) based on the full text of academic papers and investigates algorithm influence from a network perspective. Using deep learning models, we extract algorithm entities and build overall, cumulative, and annual co-occurrence networks. We analyze their structural characteristics and apply multiple centrality measures to assess the group influence of algorithms across the whole field and over time. The results show that algorithm networks display typical features of complex networks, with increasingly dense connections developing over approximately two decades. Classic, high-performing algorithms and those located at the intersections of different research periods tend to have high popularity, control, centrality, and balanced influence. When the influence of an algorithm declines, it usually loses its core network position first, followed by weaker associations with other algorithms. This study is the first large-scale analysis of algorithm co-occurrence networks. Covering more than four decades of academic publications, it provides a temporal and structural view of algorithm influence and offers a foundation for future research on networks linking algorithms, scholars, and tasks.

13:00 JSTエージェントGPT / ChatGPT

ReMMD: マルチモーダルな誤情報検出のための現実的な多言語マルチ画像エージェント検証

バイラル投稿には、長い多言語の説明、複数の画像、混合の出所、および微妙なテキストと画像の構成エラーが組み合わされているため、マルチモーダルな誤情報の検出はますます重要になっています。既存のベンチマークと手法は依然としてこの設定にあまり適合していません。通常、それらは短いキャプション、単一の画像、バイナリ ラベル、または 1 つの操作ソースを分離しますが、現実的な証拠検索ではエージェントによる検証は依然としてコストがかかります。我々は、マルチモーダルな誤情報検出のための現実的な多言語マルチ画像エージェント検証フレームワークである ReMMD を紹介します。 ReMMD には、500 のサンプル、2,756 の画像、5 つの単一言語、2 つの言語間設定、3 つのテキスト長階層、複数画像の投稿、5 方向の真実性ラベル、8 つの歪曲ラベル、証拠の出所、根拠を備えた現実世界のマルチモーダル誤情報検出ベンチマークである ReMMDBench が含まれています。また、投稿をアトミックポイントに分解し、再利用可能な証拠セットを構築し、構造化された L1/L2/L3 出力を予測する永続メモリ検証ツールである ReMMD-Agent も含まれています。独自のシステム、オープン LVLM、MMD エージェント、および T2 エージェント全体にわたって、ReMMD エージェントは、GPT-5.2 を使用して 41.80% の精度と 39.12% のマクロ F1 という最高の 5 方向正確性パフォーマンスを実現しながら、MMD エージェントと比較して 17.5%、T2 エージェントと比較して 79.9% コストを削減します。プロジェクトは https://dang-ai.github.io/ReMMD で入手できます。

原文 (English)

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle text--image framing errors. Existing benchmarks and methods remain poorly matched to this setting: they usually isolate short captions, single images, binary labels, or one manipulation source, while agentic verification remains costly under realistic evidence search. We present ReMMD, a realistic multilingual multi-image agentic verification framework for multimodal misinformation detection. ReMMD includes ReMMDBench, a real-world multimodal misinformation detection benchmark with 500 samples, 2,756 images, five monolingual languages, two cross-lingual settings, three text-length tiers, multi-image posts, five-way veracity labels, eight distortion labels, evidence provenance, and rationales. It also includes ReMMD-Agent, a persistent-memory verifier that decomposes posts into atomic points, builds a reusable evidence set, and predicts structured L1/L2/L3 outputs. Across proprietary systems, open LVLMs, MMD-Agent, and T2-Agent, ReMMD-Agent obtains the best five-way veracity performance, with 41.80% accuracy and 39.12% macro-F1 using GPT-5.2, while reducing cost by 17.5% relative to MMD-Agent and 79.9% relative to T2-Agent. The project is available at https://dang-ai.github.io/ReMMD.

13:00 JSTLLM/生成AI

VeryTrace: コンパイル可能な形式主義と構造化検証による推論トレースの検証

思考連鎖 (CoT) プロンプトを使用した複数ステップの推論は依然として脆弱です。初期段階での論理的エラーや幻覚が静かに伝播し、自信はあるが不正確な結論を導き出します。この文書では、自然言語推論トレースを構造化されたコンパイル可能な表現に形式化するゼロショット検証および修復フレームワークである VeryTrace について説明します。 VeryTrace は、(i) ステップの依存関係を明示し、(ii) 定量的なコンテンツを実行可能な式として機械化し、(iii) 演繹スキーマを介して意味推論を構造化するドメイン固有言語 (DSL) を導入します。当社のハイブリッド検証ツールは、計算の正しさ、依存関係の解決、制約を満たすための決定論的チェックと、機械化不可能な意味論的判断のための対象を絞った LLM 監査を組み合わせて、ステップレベルのエラーの位置特定と修復を可能にします。 VeryTrace は、競争数学 (AIME 2025)、ロボティクス プランニング (LLM-BabyBench)、親族推論 (CLUTRR) の 3 つの多様なドメインにわたって、ドメイン固有のトレーニングやコンテキスト内のサンプルを必要とせずに、最先端の LLM でのゼロショット ベースラインを超える精度を向上させ、形式化されたトレース検証が精度と一般化の両方を達成していることを示しています。

原文 (English)

VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification

Multi-step reasoning with Chain-of-Thought (CoT) prompting remains fragile: logical errors or hallucinations in early steps silently propagate, producing confident but incorrect conclusions. This paper presents VeryTrace, a zero-shot verification-and-repair framework that formalizes natural-language reasoning traces into a structured, compilable representation. VeryTrace introduces a Domain-Specific Language (DSL) that (i) makes step dependencies explicit, (ii) mechanizes quantitative content as executable expressions, and (iii) structures semantic inferences via deduction schemas. Our hybrid verifier combines deterministic checks for computational correctness, dependency resolution, and constraint satisfaction with targeted LLM audits for non-mechanizable semantic judgments, enabling step-level error localization and repair. Across three diverse domains-competition mathematics (AIME 2025), robotics planning (LLM-BabyBench), and kinship reasoning (CLUTRR), VeryTrace improves accuracy over zero-shot baselines on state-of-the-art LLMs without requiring domain-specific training or in-context examples, demonstrating that formalized trace verification achieves both precision and generalization.

13:00 JSTエージェント

OmniPath: 車椅子のアクセシビリティを監査するためのマルチモーダル エージェント フレームワーク

車椅子ユーザーにとって、地図上の標準的な青い線は、多くの場合、約束を破られたものです。 OpenStreetMap (OSM) のようなプラットフォームは、パスの位置をうまく把握できますが、その上を移動する物理的な感覚を伝えることができないことがよくあります。この情報の壁は車椅子ユーザーにとっては問題です。この問題を解決するために、受動的なマッピングからプロアクティブな環境監査に移行するシステムである OmniPath を紹介します。私たちのフレームワークは、OSM のネットワーク トポロジと高密度航空 LiDAR (USGS 3DEP) のサブメートル精度を融合して、歩行者環境の忠実度の高い 3D モデルを作成します。当社のエージェントは単にユーザーをルーティングするのではなく、仮想的にネットワークを横断し、0.5 メートル単位で表面を分析します。これは、ADA コンプライアンス基準に照らして、特に斜面、横断斜面、および垂直不連続部を走る物理的摩擦ポイントを厳密に定量化し、重み付けされた重大度スコアを計算して危険を「軽度」から「重大」まで分類します。現実世界の信頼性を確保するために、層別ランダムサンプリングを使用して、ナショナル モール全体にわたる 200 件の物理的なグラウンド トゥルース フィールド調査に対してシステムを検証しました。このフレームワークは、重篤度の高いハザードに対する診断上の強力な信頼性を実証し、重度カテゴリーで 0.60、重篤カテゴリーで 0.58 の F1 スコアを達成しました。このマイクロスケールの検査を自動化することで、OmniPath は標準地図が見逃す「目に見えない」障壁を特定し、静的データセットを、ユーザーが家を出る前にアクセシビリティの課題を予測するアクセシビリティ データ ソースに効果的に変換します。

原文 (English)

OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility

For a wheelchair user, a standard blue line on a map is often a broken promise. While platforms like OpenStreetMap (OSM) successfully capture where a path is, they frequently fail to convey how it physically feels to travel on it. This information barrier is problematic for wheelchair users. To solve this issue, we present OmniPath, a system that moves from passive mapping to proactive environmental auditing. Our framework fuses the network topology of OSM with the submeter precision of high-density aerial LiDAR (USGS 3DEP) to create a high-fidelity 3D model of the pedestrian environment. Rather than simply routing a user, our agent virtually traverses the network, analyzing the surface in 0.5 meter increments. It rigorously quantifies physical friction points specifically running slope, cross slope, and vertical discontinuities against ADA compliance standards, calculating a weighted severity score to categorize hazards from ``Mild'' to ``Critical.'' To ensure real world reliability, we validated the system against 200 physical ground truth field surveys across the National Mall using stratified random sampling. The framework demonstrated strong diagnostic reliability for high-severity hazards, achieving F1-scores of 0.60 for Severe and 0.58 for critical categories. By automating this micro-scale inspection, OmniPath identifies the ``invisible'' barriers that standard maps miss, effectively transforming a static dataset into accessibility data source that anticipates accessibility challenges before the user ever leaves home.

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達GPT / ChatGPT

T2D ベンチ: 多層臨床ライフスタイル ナレッジ グラフを使用した 2 型糖尿病の LLM 出力の証拠ゲート型評価

大規模言語モデル (LLM) は、2 型糖尿病に対する臨床的に流暢な推奨事項を生成できますが、ガイドラインの制約を満たしたり、ライフスタイルに関連した血糖の主張を明確に正当化したりすることはできません。我々は、LLM 出力が明示的でグラフチェック可能な証拠要件を満たしているかどうかをテストするための、再現可能なベンチマークおよび証拠ゲート型評価フレームワークである T2D-Bench を紹介します。 T2D-Bench は、生体医学スパイン (UMLS、DrugBank、SIDER)、計算可能な ADA 治療標準ルール、血糖検査室効果への機構的なブリッジを介して接続されたライフスタイル知識を組み合わせた、多層の臨床ライフスタイル ナレッジ グラフに基づいて構築されています。診断、投薬の安全性、敵対的なライフスタイルの衝突にわたる 100 の構造化されたビネット全体で、ベースライン出力は、GPT-4o-mini のケースの 35%、GPT-4o のケースの 33% で、ベンチマークで定義されたエビデンスパス チェックに失敗しました。証拠ゲートはサポートされていない省略を検出し、制約付きリビジョンを使用して、出力をベンチマークで定義された証拠要件に検証者レベルで準拠させます。これらの結果は、糖尿病に焦点を当てた LLM 出力において、計算可能な証拠の制約により、裏付けのない臨床上の省略が明示的、測定可能、修正可能になる可能性があることを示しています。

原文 (English)

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.

13:00 JST研究/論文

拡散と流れのマッチングの背後にある幾何学: ワッサーシュタイン空間における勾配の流れと測地線

有限二次モーメントを持つ確率測度の空間 $\mathcal{P}_2(\mathbb{R}^d$) は自然幾何学を持ちます。二次のワッサーシュタイン距離 W_2 により完全計量空間になり、オットーに従って、測地線が最適輸送補間である (形式的な) リーマン多様体になります。この多様体上では、自由エネルギーの勾配流 F(rho) = KL(rho || \pi) はまさにフォッカー・プランク方程式であり、その陰的オイラー離散化は JKO スキームです。これは拡散モデルの基礎となるジオメトリです。順方向プロセスは自由エネルギーを降下させ、各ノイズ除去ステップは 1 つの JKO ステップを実現し、DDPM、DDIM、NCSN/SMLD、およびエネルギー マッチングを回復します。これは 1 つのスキームであり、別々の理論ではありません。同じ多様体は、第 2 変分原理をサポートします。その測地線 (Benamu-Brenier 式の最小動作曲線) は、まさにフロー マッチングが学習する最適な転送パスです。両方の終点を固定し、測地線に従うと、生成は直線に沿った決定論的な ODE になり、サンプリング ステップが大幅に少なくなります。両方のモデル族を 1 つの多様体に配置すると、それらの関係が正確になります。拡散は自由エネルギー勾配の流れ、つまり初期値問題に従います。最適輸送フロー マッチングは、ワッサーシュタイン測地線、境界値問題に従います。この 2 つは、異なるパスを通って同じエンドポイントに到達します。

原文 (English)

The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserstein Space

The space $\mathcal{P}_2(\mathbb{R}^d$) of probability measures with finite second moment carries a natural geometry: the quadratic Wasserstein distance W_2 makes it a complete metric space and, following Otto, a (formal) Riemannian manifold whose geodesics are the optimal-transport interpolations. On this manifold, the gradient flow of the free energy F(rho) = KL(rho || \pi) is exactly the Fokker-Planck equation, and its implicit-Euler discretization is the JKO scheme. This is the geometry underlying diffusion models: the forward process descends the free energy, and each denoising step realizes one JKO step, which recovers DDPM, DDIM, NCSN/SMLD, and Energy Matching; this is one scheme, not separate theories. The same manifold supports a second variational principle. Its geodesics - the minimum-action curves of the Benamou-Brenier formula - are precisely the optimal-transport paths that Flow Matching learns. Fixing both endpoints and following the geodesic, generation becomes a deterministic ODE along a straight line, hence far fewer sampling steps. Placing both families of models on one manifold makes their relationship exact: diffusion follows a free-energy gradient flow, an initial-value problem; optimal-transport Flow Matching follows a Wasserstein geodesic, a boundary-value problem. The two reach the same endpoints along different paths.

13:00 JST研究/論文

因果強化学習の概要

因果推論は、環境に関するデータと知識を組み合わせて、反事実的な性質の質問、つまり、この実現されていない現実のデータが現在利用できない場合でも、現実が異なっていたら何が起こっていたかを推論することを可能にする一連の原則とツールを提供します。強化学習は、エージェントが環境に配置され、探索的で試行錯誤的なアプローチを追求するときに、特定の尺度 (報酬、後悔など) を最適化するポリシーを学習する方法を提供します。これら 2 つの分野は独立して発展し、実質的に相互作用することはありません。私たちは、それらが同じ構成要素のさまざまな側面、反事実関係を介して機能し、それがそれらを臍帯で結びつけていることに注目します。これらの観察に基づいて、この関係が明確に認識され数学化されると、新たな学習の機会が生まれます。この可能性を実現するために、RL エージェントがデプロイされている環境は、さまざまな因果不変性を持つ自律メカニズムの集合として分解でき、構造的因果モデルとして倹約的にモデル化できることに注意します。標準の RL 設定は、そのようなモデルを暗黙的にエンコードします。この形式化により、文献では無関係に見える、オンライン、ポリシー外、因果微積分学習など、さまざまな学習モードを統一的に扱うことができます。ただし、これらの手法は網羅的なものではありません。新しい分析次元を必要とする、自然で普及した学習設定のクラスをいくつか紹介します。具体的には、因果レンズを通して、一般化された政策学習、どこに介入するか、模倣学習、反事実学習を紹介し、議論します。これらのタスクは、反事実学習のより広い視野につながり、因果推論と強化学習を並行して研究する大きな可能性を示唆しています。これを私たちは因果強化学習 (CRL) と呼んでいます。

原文 (English)

An Introduction to Causal Reinforcement Learning

Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is currently available. Reinforcement learning provides methods to learn a policy that optimizes a specific measure (e.g., reward, regret) when the agent is deployed in an environment and pursues an exploratory, trial-and-error approach. These two disciplines have evolved independently and with virtually no interaction between them. We note that they operate over different aspects of the same building block, counterfactual relations, which makes them umbilically connected. Based on these observations, novel learning opportunities arise when this connection is explicitly acknowledged and mathematized. To realize this potential, we note that any environment where the RL agent is deployed can be decomposed as a collection of autonomous mechanisms with different causal invariances, parsimoniously modeled as a structural causal model; any standard RL setting implicitly encodes such a model. This formalization allows us to put under a unifying treatment different modes of learning, including online, off-policy, and causal calculus learning, which appear unrelated in the literature. However, these modalities are not exhaustive: we introduce several natural and pervasive classes of learning settings that entail novel dimensions of analysis. Specifically, we introduce and discuss through causal lenses generalized policy learning, where to intervene, imitation learning, and counterfactual learning. These tasks lead to a broader view of counterfactual learning and suggest great potential for studying causal inference and reinforcement learning side by side, which we call causal reinforcement learning (CRL).

13:00 JST研究/論文

ストリーミング ASR における言語を超えたエンコーダ転送を形成するのは、遅延ではなくデータ スケールです

ストリーミング音声認識モデルを新しい言語に適応させるには、多言語 (ML) エンコーダーまたは英語専用 (EN) エンコーダーという 2 つの適切なウォーム スタートのどちらかを選択する必要があります。一般的な直観としては、多言語エンコーダは低データ時に最も役立つはずであるということですが、その利点がどのくらい持続するか、ストリーミング遅延が短いことでその利点が増幅されるかどうか、そしてそれが展開の量子化に耐えられるかどうかは不明です。当社は、8 つのヨーロッパ言語、最大 5 つのターゲット言語データ スケール (100 時間から 2500 時間)、3 つのストリーミング層とオフライン デコード、および最大 4 つの公開テスト セットにわたる 0.6 B パラメーターのキャッシュ対応 FastConformer トランスデューサーの制御されたスイープによってこれらの質問に答えます。主な結果は、多言語初期化はレイテンシに制限された利点ではなく、データに制限された利点であるということです。 160 ミリ秒の FLEURS では、平均 EN-ML 単語誤り率 (WER) ギャップが 100 時間の +4.21 パーセント ポイント (pp) から 2500 時間の +0.20 pp に減少しました。べき乗則当てはめはこの減衰を要約し、ターゲット言語データが 2 倍になるごとに残りの利点がほぼ半分になります。 3 つのストリーミング層全体で、言語間の平均 EN-ML ギャップは 100 時間から 1000 時間までの各スケールでほぼ安定しており、2500 時間までにほぼゼロになります。最後に、一致する 560 ミリ秒のストリーミング層での 4 ビットの重みのみのエンコーダー量子化により、エンコーダーのフットプリントが約 3 分の 1 に削減され、FLEURS WER は平均約 0.5 pp 増加します。結果として得られるガイドラインはシンプルです。低データ領域では多言語初期化を使用し、大規模データでは選択を実質的に無関係なものとして扱い、レイテンシと量子化の決定を独立して行います。

原文 (English)

Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR

Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an English-only (EN) encoder. The common intuition is that the multilingual encoder should help most at low data, but it is unclear how long that advantage persists, whether tight streaming latency amplifies it, and whether it survives deployment quantization. We answer these questions with a controlled sweep of a 0.6 B-parameter cache-aware FastConformer transducer across eight European languages, up to five target-language data scales (100 h to 2500 h), three streaming tiers plus offline decoding, and up to four public test sets. The main result is that multilingual initialization is a data-limited advantage, not a latency-limited one. On FLEURS at 160 ms, the mean EN-ML word error rate (WER) gap falls from +4.21 percentage points (pp) at 100 h to +0.20 pp at 2500 h; a power-law fit summarizes this decay, with each doubling of target-language data roughly halving the remaining advantage. Across the three streaming tiers, the across-language mean EN-ML gap is approximately stable at each scale from 100 to 1000 h, and is near zero by 2500 h. Finally, 4-bit weight-only encoder quantization at the matched 560 ms streaming tier reduces the encoder footprint by about 3x, with an average FLEURS WER increase of about 0.5 pp. The resulting guideline is simple: use multilingual initialization in low-data regimes, treat the choice as effectively irrelevant at large data, and make latency and quantization decisions independently.

13:00 JST研究/論文

パーソナライズされたマルチモーダル生成に向けてユーザー行動をナビゲートする

最新の AIGC パイプラインは高忠実度の画像とビデオを提供しますが、整形式の作成指示を前提としていますが、エンドユーザーが視覚的な詳細を明確にすることはほとんどなく、ジェネレーターがユーザーの要求とずれたままになっています。私たちは、ユーザーのインタラクション履歴を下流合成用の実行可能な命令に変換するパーソナライズされたコンテンツ生成を研究し、2 つの障害を特定しました。1 つは言語推論が読みやすい形式で動作をエンコードする必要があること、もう 1 つはモデルが事前トレーニングと動作データの両方に存在しない命令作成スキルを獲得することです。私たちは NaviGen を提案します。NaviGen は、1 つのトークン ストリーム内の動作基盤およびセマンティック ブリッジとして、協調的なコードとテキスト コードを結合する二重識別子で各項目を表します。この表現では、2 段階の SFT+RL パイプラインが最初に進化的に検索された監視から優先推論と命令記述を抽出し、次に階層的で自己矛盾のない報酬を通じてユーザーの意図に合わせて生成を調整します。製品、ゲーム、ショートビデオの各分野にわたる実験では、NaviGen がパーソナライズされた画像とビデオの生成を改善し、次のアイテムの予測を強化し、より具体的で関連性のある視覚的に生成可能な指示を生成できることが示されています。私たちのコードは匿名で https://github.com/iLearn-Lab/NaviGen で公開されています。

原文 (English)

Navigating User Behavior toward Personalized Multimodal Generation

Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, leaving generators misaligned with user demand. We study personalized content generation, which turns a user's interaction history into an executable instruction for downstream synthesis, and identify two obstacles: behavior must be encoded in a form legible to language reasoning, and the model must acquire instruction-writing skill absent from both pretraining and behavior data. We propose NaviGen, which represents each item with a dual identifier coupling a collaborative code and a textual code as a behavioral substrate and a semantic bridge in one token stream. On this representation, a two-stage SFT+RL pipeline first distills preference reasoning and instruction writing from evolutionarily searched supervision, then aligns generation with user intent through hierarchical and self-consistent rewards. Experiments across product, game, and short-video domains show that NaviGen improves personalized image and video generation, strengthens next-item prediction, and yields more specific, relevant, and visually generatable instructions. Our code is released at: https://github.com/iLearn-Lab/NaviGen.

13:00 JST研究/論文

人間中心の AI と企業の特異なリスクとの関係を探る

インダストリー 5.0 における人間中心の AI (HCAI) については広範な議論が行われているにもかかわらず、企業の特異リスク (IR) に対するその影響はまだ十分に解明されていません。これは、企業レベルの株価のボラティリティを体系的な要因から分離することで、企業の異種AI戦略や導入に対する投資家の反応を反映するものであり、現在のテクノロジー革命の中で財務リスクを乗り切る企業にとって緊急の課題となっている。状況に応じた AI 理論と社会技術システム理論を統合することで、私たちは HCAI を状況に応じた AI 戦略として概念化します。これは、AI 関連の倫理的リスクを軽減し、企業の事業運営における AI と人間の相乗効果を促進し、ステークホルダーの多様な期待に合わせることで最終的に IR を削減します。さらに、デジタル化、業務効率、経営陣の株式保有、IT の背景を持つ CEO などの社会技術的要因が、HCAI と IR の関係を緩和する可能性があります。 2015 年から 2023 年までの中国の上場企業のマルチソース パネル データセットを使用したところ、HCAI が企業の IR の低下と関連していることがわかりました。さらに、デジタル化と経営陣の株式保有はこのリスク軽減効果を強化しますが、業務効率化と IT の背景を持つ CEO は驚くほどその効果を弱めます。私たちの調査結果は、AI 時代の倫理的な AI ガバナンスと確実な財務リスク管理の両方に対する理論的貢献と実践的な洞察を提供します。

原文 (English)

Exploring the relationship between human-centric AI and firm idiosyncratic risks

Despite the extensive discussions of human-centric AI (HCAI) in Industry 5.0, its effects on firms' idiosyncratic risks (IR) remains underexplored. This is an imperative issue for firms navigate financial risks during the current technological revolution, as IR reflects investor reactions to corporate heterogeneous AI strategies and implementations by isolating firm-level stock volatility from systematic factors. Integrating situated AI theory with social-technical systems theory, we conceptualise HCAI as a situated AI strategy that reduces AI-related ethical risks and fosters AI-Human synergies in firms' business operations, ultimately reducing IR by aligning with stakeholders' diverse expectations. Moreover, socio-technical factors, namely digitalisation, operational efficiency, executive shareholding, and CEOs with IT background, may moderate the HCAI-IR relationship. Using a multi-source panel dataset of Chinese listed firms from 2015 to 2023, we find that HCAI is associated with lower firm IR. Furthermore, digitalisation and executive shareholding strengthen this risk-reducing effect, whereas operational efficiency and CEOs with IT background surprisingly attenuate it. Our findings offer theoretical contributions and practical insights for both ethical AI governance and firm financial risk management in the AI era.

13:00 JST研究/論文

FlowR2A: マルチモーダル運転計画のための報酬と行動の配分の学習

マルチモーダル運転計画は、2 つのパラダイムの間の長年の緊張に直面しています。スコアベースの方法は、緻密な報酬監視の恩恵を受けますが、固定されたアクション語彙に限定されます。一方、アンカーベースの方法は、提案を動的に生成しますが、単一のグラウンドトゥルース軌道に制限されるまばらな監視に悩まされます。この研究では、識別ターゲットからのシミュレーションベースの報酬を生成条件に再構築することで、この緊張を解決する FlowR2A を提案します。 FlowR2A は、フロー マッチング デコーダーを使用して高密度の軌道と報酬のペアから報酬条件付きアクションの分布を学習することで、スコアリング ベースの手法の高密度な監視とアンカー ベースの手法の提案生成を単一の生成モデルで統合し、安全性、進歩、快適さ、ルール遵守におけるアクションとその結果の間の相関関係をモデルに強制的に内部化します。ソフトな進捗目標に対してハードな安全制約のバランスをとるために、タイムステップごとのきめ細かい報酬条件付けと報酬ノイズの増大を導入します。生成的な定式化は、報酬ガイダンスとアンカー サンプリングを通じて制御可能なテスト時間のサンプリングを自然にサポートし、高品質の提案を生成します。 FlowR2A は、NAVSIM v1 および v2 ベンチマークで最先端の結果を達成し、従来の方法よりも大幅に高品質のマルチモーダル提案を実現します。

原文 (English)

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods generate proposals dynamically yet suffer from sparse supervision constrained to a single ground-truth trajectory. In this work, we propose FlowR2A, which resolves this tension by reframing simulation-based rewards from discriminative targets into generative conditions. By learning the reward-conditioned action distribution from dense trajectory-reward pairs with a flow-matching decoder, FlowR2A unifies the dense supervision of scoring-based methods with the proposal generation of anchor-based methods in a single generative model, forcing the model to internalize the correlation between an action and its outcomes in safety, progress, comfort, and rule compliance. To balance hard safety constraints against soft progress objectives, we introduce fine-grained per-timestep reward conditioning and reward noise augmentation. The generative formulation naturally supports controllable test-time sampling via reward guidance and anchored sampling, producing high-quality proposals. FlowR2A achieves state-of-the-art results on the NAVSIM v1 and v2 benchmarks, with multimodal proposals of substantially higher quality than prior methods.

13:00 JSTエージェント

SP-Mind: 空間プロテオミクス解析のための自律推論エージェント

空間プロテオミクスは、組織構造内のタンパク質発現の単一細胞解像度の特性評価を可能にし、腫瘍微小環境を理解し、精密医療を導く上で重要な役割を果たします。しかし、現在の分析ワークフローは断片化したままであり、異種ツールを専門家が手動で調整する必要があり、研究の拡張性と再現性が制限されています。生の多重組織イメージングから下流の表現型発見まで、空間プロテオミクス解析パイプラインを統合するように設計された初の自律型 AI エージェントである SP-Mind を紹介します。専門家が厳選した生物学的分析スキルと特殊な計算ツールを備えた SP-Mind は、タスク固有の微調整を行うことなく、自然言語クエリをエンドツーエンドの分析ワークフローに変換します。その機能を厳密に評価するために、18 の異なるカテゴリにわたる 102 のタスクで構成される、さまざまな組織タイプにわたる包括的なベンチマークである SP-Bench を導入します。 SP-Bench と確立された下流タスクでの広範な評価を通じて、SP-Mind は既存のオープンソース生物医学的エージェントのベースラインと比較して最先端のパフォーマンスを達成します。

原文 (English)

SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis

Spatial proteomics enables single-cell-resolution characterization of protein expression within tissue architecture, playing a critical role in understanding tumor microenvironments and guiding precision medicine. However, current analysis workflows remain fragmented, requiring expert manual orchestration of heterogeneous tools and limiting research scalability and reproducibility. We present SP-Mind, the first autonomous AI agent designed to unify the spatial proteomics analysis pipeline, from raw multiplexed tissue imaging to downstream phenotype discovery. Equipped with expert-curated biological analysis skills and specialized computational tools, SP-Mind converts natural-language queries into end-to-end analytical workflows without task-specific fine-tuning. To rigorously evaluate its capabilities, we introduce SP-Bench, a comprehensive benchmark spanning diverse tissue types, comprising 102 tasks across 18 distinct categories. Through extensive evaluation on SP-Bench and established downstream tasks, SP-Mind achieves state-of-the-art performance compared to existing open-source biomedical agent baselines.

13:00 JST研究/論文

フェデレーテッド・ロングテール・グラフ学習に向けて: エネルギーに導かれたデュアル・デカップリング・アプローチ

Federated Graph Learning は、データのプライバシーを維持しながら、分散クライアント間での共同グラフ モデリングを容易にします。ただし、現実世界のデータ カテゴリは、長い裾の分布を示すことがよくあります。このような統計的欠乏は、2 つの点でパフォーマンスを大幅に低下させます。1 つはグローバル モデルを多数派クラスに偏らせること、もう 1 つは少数派のノードを異好性の頭支配の近傍に沈めることで構造的に分離することです。既存の方法はトポロジーに依存しない統計的補償を試みますが、データ不足の下では失敗することがよくあります。末尾ノードを回復する代わりに、隣接する支配的なクラスからの構造ノイズを過剰適合させて、表現の劣化を引き起こします。これらの制限に対処するために、トポロジカルな浄化をセマンティックな再調整から分離する二重分離パラダイムに基づいて構築されたフレームワークである FedEPD を提案します。具体的には、FedEPD は分布を意識したディリクレ エネルギー プルーニングを利用して、空間的な異好性エッジをフィルター処理します。次に、トポロジー的に中心的なノードから堅牢なグローバル プロトタイプを抽出することで、非 IID 分布のシフトを克服します。これは、空間ローパス プロトタイプ インジェクションを介してローカル表現に組み込まれます。さらに、2 段階の交互最適化戦略により、少数派の精度を向上させながら、多数決の境界を厳密に保護します。広範な実験により、FedEPD がさまざまなロングテール ベンチマークにわたって最先端のパフォーマンスを達成し、精度で最大 4.97%、マクロ F1 で 5.48% の絶対的な向上が得られることが実証されました。

原文 (English)

Towards Federated Long-Tailed Graph Learning: An Energy-Guided Dual Decoupling Approach

Federated Graph Learning facilitates collaborative graph modeling across distributed clients while preserving data privacy. However, real-world data categories frequently exhibit long-tailed distributions. Such statistical scarcity severely degrades performance in two ways: it biases the global model toward majority classes, and it structurally isolates minority nodes by submerging them in heterophilic, head-dominated neighborhoods. While existing methods attempt topology-agnostic statistical compensations, they often fail under data scarcity. Instead of recovering tail nodes, they overfit the structural noise from adjacent dominant classes, leading to representation degradation. To address these limitations, we propose FedEPD, a framework built on a dual decoupling paradigm that separates topological purification from semantic recalibration. Specifically, FedEPD utilizes distribution-aware Dirichlet energy pruning to filter spatial heterophilic edges. It then overcomes Non-IID distribution shifts by extracting robust global prototypes from topologically central nodes, which are incorporated into local representations via a spatial low-pass prototype injection. Furthermore, a two stage alternating optimization strategy strictly protects majority decision boundaries while improving minority accuracy. Extensive experiments demonstrate that FedEPD achieves state-of-the-art performance across diverse long-tailed benchmarks, yielding absolute improvements of up to 4.97% in Accuracy and 5.48% in Macro-F1.

13:00 JST研究/論文

言語モデルの誤った思考プロセスを調査する

大規模な言語モデルでは、戦略的欺瞞、サンドバッグ、自己保存など、ますます多様な不整合な動作が見られます。一か八かの環境での導入が増えているため、安全かつ責任ある使用を確保するには、そのような動作を確実に検出することが重要です。この研究では、位置ずれをきめの細かい認知プロセス (位置ずれ指標) に分解し、線形プローブを介してモデルの内部活性化における位置ずれの存在を検出することで、位置ずれを監視することを提案します。私たちは、さまざまな不整合な行動にわたる 18 の指標の分類を開発し、複数ターンのトレーニング会話を生成する自動化されたメタプランに基づくパイプラインと組み合わせます。一般化を厳密に評価するために、自動化された行動誘発、確立された不整合ベンチマーク、および自然で無害な会話を組み合わせた配布外スイートを構築します。 5 つの不整合な動作にわたって、当社のプローブは、無害なトラフィックでの低い誤検知率を維持しながら、配布外ベンチマークで 0.935 AUROC という強力な LLM ジャッジと一致しました。さらに詳細な分析を実行して、プローブと位置ずれ指標のモデルの内部表現を理解します。

原文 (English)

Probing the Misaligned Thinking Process of Language Models

Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-stakes settings, it is critical to reliably detect such behaviors to ensure safe and responsible use. In this work, we propose to monitor misalignment by decomposing it into fine-grained cognitive processes -- misalignment indicators -- and detecting their presence in a model's internal activations via linear probes. We develop a taxonomy of 18 indicators spanning different misaligned behaviors, paired with an automated, meta-plan-guided pipeline that generates multi-turn training conversations. To rigorously evaluate generalization, we construct an out-of-distribution suite combining automated behavioral elicitation, established misalignment benchmarks, and natural benign conversations. Across 5 misaligned behaviors, our probes match a strong LLM judge with 0.935 AUROC on out-of-distribution benchmarks while keeping a low false positive rate on benign traffic. We further perform in-depth analysis to understand the probes and the model's internal representations of misalignment indicators.

13:00 JST研究/論文

合理的閉鎖の下での防御可能な DL-Lite のための扱いやすい推論と論理積クエリ応答

記述論理 (DL) では、合理的閉包 (RC) に基づく推論は、実行可能な知識を処理するためのよく知られ、広く受け入れられている非単調形式主義です。この論文では、軽量記述ロジックの DL-Lite ファミリのコアおよびホーンのバリアントへの RC の適用を研究します。 RC では、資格 (インスタンスのチェック) と接続クエリ (CQ) 応答の両方を分析します。私たちの主な貢献は、既存の標準的な古典的推論に基づいて構築されたプラグイン アーキテクチャを提供し、DL-Lite の RC での推論と CQ 応答が最小限の計算オーバーヘッドで効率的に実行できることを確立したことです。

原文 (English)

Tractable Reasoning and Conjunctive Query Answering for Defeasible DL-Lite under Rational Closure

In Description Logics (DLs), reasoning under Rational Closure (RC) is a well-known and widely accepted non-monotonic formalism to handle defeasible knowledge. In this paper, we study the application of RC to the core and horn variants of the DL-Lite family of lightweight description logics. We analyze both entitlement (instance checking) and Conjunctive Query (CQ) answering under RC. Our main contribution is providing a plug-in architecture that builds upon existing standard classical reasoners, establishing that reasoning and CQ answering under RC for DL-Lite can be done efficiently with minimal computational overhead.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

レモンハーネス技術レポート

大規模言語モデル (LLM) エージェントがより長いタスクに適用されると、複数ラウンドの反復にわたってワークスペースの状態がますます変更されます。ただし、エージェントは通常、ツールの出力とログの断片のみを観察し、実際の状態の変化はファイル システムで発生します。明示的なワークスペース境界がないと、ファイルの書き込みや一時的なアーティファクトの生成などの状態変更操作により、パス全体に変更が分散される可能性があります。時間の経過とともに、これらの弱く制約された変更が蓄積され、変更されたファイルなどの状態を追跡することが困難になります。この文書では、長期的なエージェント向けの統合実行フレームワークである LemonHarness について説明します。 LemonHarness は、明確に定義されたワークスペース内で状態変更操作を制限し、モデルの呼び出し、ツールの実行、およびルールの知識を単一の制御された境界内に持ち込むことによって、明示的な実行境界を確立します。ファイルの書き込み、依存関係のインストール、一時的なアーティファクトの作成などの状態変更操作は、構造化されたツール インターフェイスを通じて実行され、実行フィードバックが観察として記録され、後続のモデル決定に利用できます。このシステムには、再利用可能なルール ナレッジ ベースも導入されており、繰り返し実行ルールと受け入れ基準がランタイム ナレッジに変わります。 LemonHarness はさらに、経過予算と残り予算をモデルに公開する時間認識実行メカニズムを追加します。これにより、時間プレッシャーの変化に応じて調査、実装、検証作業のバランスを再調整し、長時間の待機や過剰な検証によるタイムアウトを回避できます。 Terminal-Bench 2.0 では、LemonHarness_GPT-5.3-CodeX は 445 回のトライアルで 84.49% の精度に達しました。同じフレームワークとより強力な GPT-5.5 バックボーンを組み合わせることで、5 つのジョブの平均精度が 86.52% に向上しました。この結果は、統合された実行時間境界、呼び出し可能なルールの知識、および時間認識の実行により、長期的なエージェント実行の安定性が向上する可能性があることを示唆しています。

原文 (English)

LemonHarness Technical Report

As large language model (LLM) agents are applied to longer tasks, they increasingly modify workspace state across multiple rounds of iteration. However, agents typically observe only tool outputs and log fragments, while the actual state changes occur in the file system. Without explicit workspace boundaries, state-changing operations such as file writes and temporary artifact generation may scatter changes across paths. Over time, these weakly constrained changes accumulate, making states such as modified files difficult to track. This paper presents LemonHarness, an integrated execution framework for long-horizon agents. LemonHarness establishes an explicit execution boundary by constraining state-changing operations within a clearly defined workspace and bringing model invocation, tool execution, and rule knowledge within a single controlled boundary. State-changing operations, including file writes, dependency installation, and temporary artifact creation, are executed through structured tool interfaces, with execution feedback recorded as observations available to subsequent model decisions. The system also introduces a reusable rule knowledge base, which turns recurring execution rules and acceptance criteria into runtime knowledge. LemonHarness further adds a time-aware execution mechanism that exposes elapsed and remaining budget to the model, so it can rebalance exploration, implementation, and validation effort as time pressure shifts and avoid timeouts from long waits or excessive verification. On Terminal-Bench 2.0, LemonHarness_GPT-5.3-CodeX reached 84.49% accuracy over 445 trials; pairing the same framework with the stronger GPT-5.5 backbone raised the average accuracy to 86.52% across five jobs. The results suggest that a unified runtime boundary, callable rule knowledge, and time-aware execution can improve the stability of long-horizon agent execution.

13:00 JST研究/論文

Prob-BBDM: MRI シーケンスの画像間変換のための確率的ブラウン橋拡散モデル

AI を活用した画像間の合成は急速に進歩しており、医療画像分野での応用が拡大しています。マルチモーダル画像解析は検査の品質を最適化する上で重要な役割を果たしますが、臨床現場で複数の画像モダリティを取得することは依然としてリソースを大量に消費し、特に 3D イメージングでは時間がかかります。この課題に対処するために、我々は、2D アキシャル スライスから磁気共鳴画像法 (MRI) シーケンスを合成するブラウン橋拡散モデル (BBDM) に基づく新しい画像間変換モデルを提案します。私たちのアプローチは、変分エンコーダーによる拡散メカニズムを統合し、確率的な画像分布を活用して合成品質を向上させます。 BraTS 2021 データセットで評価された当社の Probabilistic-BBDM (Prob-BBDM) は、複数の翻訳タスクにわたって優れたパフォーマンスを達成し、最大 88.46% の SSIM と 26.09 dB PSNR に達し、ベースラインを一貫して改善しています。特に、当社の拡散プロセスに必要なステップは 4 つだけであり、高品質の合成を維持しながら計算効率が高くなります。一般性をさらに検証するために、外部のサードパーティ データセットで Prob-BBDM をテストし、ドメイン全体で一貫したパフォーマンスを実証します。さらに、合成されたスライスを事前にトレーニングされたセグメンテーション モデルへの入力として使用することで、そのスライスの臨床的有用性を評価します。腫瘍のセグメンテーションにより、88.71% の Dice スコアと 3.49 mm の HD95 が得られ、合成されたスライスが重要な診断情報を保存していることが確認されました。これらの結果は、高品質、効率的、汎用性のある MRI 合成に対する Prob-BBDM の可能性を強調しており、医用画像変換の改善に向けた有望な一歩を提供します。

原文 (English)

Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation

AI-driven image-to-image synthesis is rapidly advancing, with growing applications in medical imaging. Multi-modal image analysis plays a crucial role in optimizing examination quality, yet acquiring multiple imaging modalities in clinical settings remains resource-intensive and time-consuming, especially for 3D imaging. To address this challenge, we propose a novel image-to-image translation model based on Brownian Bridge Diffusion Models (BBDM), which synthesizes magnetic resonance imaging (MRI) sequences from 2D axial slices. Our approach integrates a variational encoder-guided diffusion mechanism, leveraging probabilistic image distributions to enhance synthesis quality. Evaluated on the BraTS 2021 dataset, our Probabilistic-BBDM (Prob-BBDM) achieves superior performance across multiple translation tasks, reaching up to 88.46% SSIM and 26.09 dB PSNR, with consistent improvements over baselines. Notably, our diffusion process requires only 4 steps, making it computationally efficient while maintaining high-quality synthesis. To further validate generalizability, we test Prob-BBDM on an external third-party dataset, demonstrating consistent performance across domains. Additionally, we assess the clinical utility of the synthesized slices by using them as input to a pre-trained segmentation model. Tumor segmentation yields a Dice score of 88.71% and an HD95 of 3.49 mm, confirming that the synthesized slices preserve critical diagnostic information. These results highlight the potential of Prob-BBDM for high-quality, efficient, and generalizable MRI synthesis, offering a promising step toward improved medical image translation.

13:00 JST研究/論文

MVG-KAN: PM$_{2.5}$ 予測用のマルチビュー地風ガイド付き KAN

正確な短期 PM$_{2.5}$ 予測は、公衆衛生保護、大気質早期警報、都市環境管理にとって重要です。しかし、PM$_{2.5}$ の変動は、人間の活動や気象の規則性によって引き起こされる安定した周期的変化、観測所固有の短期濃度の変化、観測所間の気象学に起因する汚染物質の分散など、複数の複合要因によって引き起こされます。既存の時空間予測手法は観測点の関係をある程度把握できますが、距離のみ、相関ベース、または純粋に適応的なグラフでは、これらの不均一な要因、特に風向に依存する汚染物質の輸送を包括的に表現するには不十分なことがよくあります。この問題に対処するために、\textbf{MVG-KAN} という名前の PM$_{2.5}$ 予測用のマルチビュー地理風ガイド KAN モデルを提案します。このモデルは、局所的な周期規則性、観測点ごとの残留時間ダイナミクス、気象環境に誘導された空間分散という 3 つの相補的なビューから観測点レベルの PM$_{2.5}$ の進化をモデル化します。具体的には、周期的残差予測バックボーンは、まず、安定した日次および週次パターンを非周期的残差変動から分離します。 Geo-Wind Graph は、地理的距離の減衰と風向および風速を意識した伝送を組み合わせて構築され、ステーション間の残留伝播に対して軽量の物理的動機による有向空間事前分布を提供します。さらに、時間的コルモゴロフ-アーノルド ネットワーク (TKAN) 残差ヘッドを導入して、非周期化 PM$_{2.5}$ 残差と過去の複数汚染物質シーケンスからステーションごとの非線形自己回帰補正を学習し、それによって局所的な残留慣性と汚染物質の共変動のモデリングを強化します。

原文 (English)

MVG-KAN: Multi-View Geo-Wind Guided KAN for PM$_{2.5}$ Forecasting

Accurate short-term PM$_{2.5}$ forecasting is important for public health protection, air-quality early warning, and urban environmental management. However, PM$_{2.5}$ variation is driven by multiple coupled factors, including stable periodic changes induced by human activities and meteorological regularity, station-specific short-term concentration evolution, and meteorology-driven pollutant dispersion among monitoring stations. Existing spatio-temporal forecasting methods may capture station relationships to some extent, but distance-only, correlation-based, or purely adaptive graphs are often insufficient to comprehensively represent these heterogeneous factors, especially wind-direction-dependent pollutant transport. To address this problem, we propose a Multi-View Geo-Wind Guided KAN model for PM$_{2.5}$ forecasting, named \textbf{MVG-KAN}, which models station-level PM$_{2.5}$ evolution from three complementary views: local periodic regularity, station-wise residual temporal dynamics, and meteorological-environment-guided spatial dispersion. Specifically, the periodic-residual forecasting backbone first separates stable daily and weekly patterns from non-periodic residual variations. A Geo-Wind Graph is constructed by combining geographic distance decay with wind-direction- and wind-speed-aware transport, providing a lightweight physically motivated directed spatial prior for residual propagation among stations. In addition, a temporal Kolmogorov-Arnold network (TKAN) residual head is then introduced to learn station-wise nonlinear autoregressive correction from de-periodized PM$_{2.5}$ residuals and historical multi-pollutant sequences, thereby enhancing the modeling of local residual inertia and pollutant co-variation.

13:00 JSTLLM/生成AI

拡散ベースの並列処理とトレーナー支援生成によるビジュアル生成 LLM の分離 RL の高速化

強化学習 (RL) はトレーニング後のパラダイムの主流となっており、自己回帰大規模言語モデル (LLM) 用の veRL などの高性能 RL システムの出現を推進しています。並行して、DanceGRPO や FlowGRPO などの拡散指向の RL アルゴリズムにより、RL の範囲が言語推論から拡散ベースのビジュアルおよびフローベースの生成へと急速に拡大されました。ただし、拡散生成 LLM 用の効率的な RL システムはまだ研究されていません。既存の実装 (veRL-Omni など) は依然としてコロケーション実行に依存しており、同期は簡素化されますが、ロールアウトとトレーニングのリソースが結合され、異種導入が制限され、独立したスケーリングが制約されます。この目的を達成するために、柔軟なリソース割り当てをサポートし、異種 GPU に対応し、効率的なタスク スケジューリングを促進する、拡散ベースの生成 LLM 用の分散 RL フレームワークである DigenRL を導入します。分散アーキテクチャにおける実行バブルを最大限に減らすために、次のことを提案します。1) 拡散アーキテクチャにおける世代軸パイプライン (GAP) とタイムステップ並列処理 (TSP) により、ロールアウトとトレーニングの間のよりきめの細かいパイプライン処理が可能になります。 2) トレーナー GPU リソースがロールアウト世代の実行を動的に支援できるようにするエラスティック トレーナー支援生成 (TAG) アプローチ。 3) パイプラインのテール バブルをさらに利用するための、厳密に 1 ステップで制約された非同期戦略。 HunyuanVideo-13B、Wan2.1-14B、FLUX.1-12B、および QwenImage-20B 生成モデルを使用して、16 ~ 32 GPU を備えた 3 つのハードウェア テストベッドで広範な実験が行われています。実験結果は、DigenRL が最先端の拡散 RL システムである veRL-Omni および GenRL と比較して 1.56 ~ 2.10 倍のスループット向上を達成することを示しています。

原文 (English)

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models (LLMs). In parallel, diffusion-oriented RL algorithms, e.g., DanceGRPO and FlowGRPO, have rapidly expanded the scope of RL from language reasoning to diffusion-based visual and flow-based generation. However, efficient RL systems for diffusion generative LLMs remain underexplored. Existing implementations, e.g., veRL-Omni, still rely on colocated execution, which simplifies synchronization but couples rollout and training resources, limits heterogeneous deployment, and constrains independent scaling. To this end, we introduce DigenRL, a disaggregated RL framework for diffusion-based generative LLMs that supports flexible resource allocation, accommodates heterogeneous GPUs, and facilitates efficient task scheduling. To maximally reduce the execution bubbles in the disaggregated architecture, we propose: 1) a generation-axis pipeline (GAP) and time-step parallelism (TSP) in the diffusion architecture to enable finer-grained pipelining between rollout and training; 2) an elastic trainer-assisted generation (TAG) approach to enable the trainer GPU resources to dynamically assist in executing rollout generations; and 3) a tightly one-step constrained asynchronous strategy to further utilize the tail bubble in the pipeline. Extensive experiments are conducted on three hardware testbeds with 16-32 GPUs using HunyuanVideo-13B, Wan2.1-14B, FLUX.1-12B, and QwenImage-20B generative models. Experimental results show that DigenRL achieves 1.56-2.10x throughput improvements over state-of-the-art diffusion RL systems, veRL-Omni and GenRL.

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTGemini

有用性が因果関係を無効にする場合の注意: LLM におけるコンテキスト依存の抑制と回復

大規模言語モデル (LLM) は、ビジネスおよび政策のコンテキストにおける意思決定支援の役割にますます統合されています。これまでのベンチマーク研究では主に LLM の因果推論能力が評価されてきましたが、より基本的な認識論的側面は見落とされてきました。因果的注意とは、経験的証拠が不十分な場合に因果関係の判断を控える傾向として定義されます。この研究では、LLM が学術的な文脈から実践的な助言の文脈に移行するときに発生する、因果的注意の体系的な抑制を調査します。 Pearl の因果階層 (PCH スコア) にヒントを得た評価ルーブリックを使用して、4 つの高性能 LLM (Claude Sonnet 4.6、Claude Opus 4.7、GPT 5.5、および Gemini 3.1 Pro) で 480 回のトライアルにわたって実験を実施しました。因果関係の注意維持率は、学術的な文脈では 91.7 ~ 100.0% でしたが、実践的なアドバイスの文脈では 6.7 ~ 18.3% に低下しました (フィッシャーの直接確率検定、すべてのモデルで p < .001)。さらに、具体的な推奨事項や説明の根拠を求める実際的なプロンプトに限定すると、因果関係注意を維持した回答は 200 件中 1 件 (0.5%) のみでした。 「因果関係の観点からこの判断を再考してください」という短い自己修正プロンプトにより、因果関係注意の表現が 71.4 ~ 100.0% の維持率に戻りました (マクネマーの検定、すべてのモデルで p < .001)。これらの結果は、有用性指向の応答パターンが実際の助言の場面での因果的注意の表現を抑制する可能性があり、組織のガバナンスに重要な影響を与える可能性があることを示唆しています。この調査結果は、この抑制が根底にある機能制限ではなく、表現におけるコンテキスト依存の変動を反映していることを示しており、提案生成と因果関係の監査を分離するマルチエージェント アーキテクチャが有望なガバナンス設計を提供する可能性があることを示唆しています。

原文 (English)

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs

Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts. While prior benchmark studies have primarily evaluated LLMs' causal reasoning capabilities, a more fundamental epistemic dimension has been overlooked: Causal Caution, defined as the propensity to refrain from causal judgment when empirical evidence is insufficient. This study examines the systematic suppression of Causal Caution that occurs when LLMs shift from academic to practical advisory contexts. Using an evaluation rubric inspired by Pearl's Causal Hierarchy (the PCH score), we conducted experiments on four high-performance LLMs -- Claude Sonnet 4.6, Claude Opus 4.7, GPT 5.5, and Gemini 3.1 Pro -- across 480 trials. Causal Caution maintenance rates were 91.7--100.0% in academic contexts but dropped to 6.7--18.3% in practical advisory contexts (Fisher's exact test, p < .001 across all models). Furthermore, when restricted to practical prompts requesting concrete recommendations or explanatory rationales, only 1 of 200 responses (0.5%) maintained Causal Caution. A brief self-correction prompt -- "Please reconsider this judgment from the perspective of causal relationships" -- restored the expression of Causal Caution to maintenance rates of 71.4--100.0% (McNemar's test, p < .001 across all models). These results suggest that helpfulness-oriented response patterns may suppress the expression of Causal Caution in practical advisory contexts, with important implications for organizational governance. The findings indicate that this suppression reflects context-dependent variation in expression rather than an underlying capability limitation, suggesting that multi-agent architectures that separate proposal generation from causal auditing may offer a promising governance design.

13:00 JST研究/論文

PHANTOM: 視覚言語モデルに対するマルチモーダル敵対的攻撃の大規模データセット

ビジョン言語モデル (VLM) 用に事前に生成された敵対的攻撃の大規模なオープンソース データセットを紹介します。このデータセットは、多様性があり、代表的で、実用的になるように設計されており、有害な意図の 10 の高レベル カテゴリと 55 のサブカテゴリをカバーすることで既存のベンチマークを拡張します。私たちの主な目標は、大量の攻撃を生成する計算コストと複雑さを考慮して、研究コミュニティが敵対的なデータにアクセスできるようにすることです。データセットは、最近の文献からの最先端の攻撃戦略を使用して生成された 47,524 個の敵対的サンプルで構成されています。私たちの取り組みは、複数の確立されたソースからの以前のベンチマークを統合および拡張することで既存の取り組みを補完し、その結果 7,826 のインテントが得られ、追加のカテゴリを導入して対象範囲を広げています。これにより、モデルの堅牢性と整合性を研究するための現実的な評価リソースが提供されます。私たちのデータセットは、研究者や実践者が VLM の堅牢性と安全性を体系的に評価し、攻撃生成モデルを微調整し、さまざまな敵対状況下で防御ガードレールを開発またはストレス テストできるようにすることを目的としています。このリソースを公開することで、敵対的研究への障壁を下げ、VLM の安全性についてより再現性があり、包括的で比較可能な評価を促進することを目指しています。

原文 (English)

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

We introduce a large-scale, open-source dataset of pre-generated adversarial attacks for vision-language models (VLMs). The dataset is designed to be diverse, representative, and practical, extending existing benchmarks by covering 10 high-level categories and 55 subcategories of harmful intents. Our primary goal is to make adversarial data accessible to the research community, given the computational cost and complexity of generating large numbers of attacks. The dataset comprises 47 524 adversarial samples, generated using state-of-the-art attack strategies from recent literature. Our work complements existing efforts by consolidating and extending prior benchmarks from multiple established sources, resulting in 7 826 intents, and introduce an additional category to broaden coverage. This provides realistic evaluation resources for studying model robustness and alignment. Our dataset intends to enable researchers and practitioners to systematically evaluate the robustness and safety of VLMs, fine-tune attack-generation models, and develop or stress-test defensive guardrails under diverse adversarial conditions. By releasing this resource, we aim to lower the barrier to adversarial research and foster more reproducible, comprehensive, and comparable evaluations of VLM safety.

13:00 JSTLLM/生成AI研究/論文

LLM の時代: 戦争の霧の下での推論、外交、大規模言語モデルの信頼性のための戦略的な 1 対 1 ベンチマーク

Age of LLM を紹介します。これは、2 つの LLM が 13x7 グリッドで対決して敵の基地を破壊する、ターンベースの 1v1 ベンチマークです。 3 つのストレス要因は意図的なものです。戦争の霧、完全な外交 (メッセージ、停戦、最後通牒、ウランは秘密にされます)、そして毎ターン厳格な JSON スキーマに従わなければならず、違法行為は黙って破棄される信頼性の側面です。エンジンはプライベートであり、各試合では新鮮なランダムなマップ シードと対戦相手が使用されるため、公開ベンチマークに影響を与えるデータ汚染が軽減されます。モデルは、構築順序のアドバイスのない (ほぼ) ルールのみのプロンプトを受け取ります (データ収集中に 2 つの戦術的なシード フレーズが存在しました。セクション 2.7 を参照)。 54 の一致と 5,258 のアクションにわたって 15 の推論モデルをベンチマークしました。調査結果: (1) 核ラッシュは、認知的抑止力の失敗ではなく、機密同時発射ルールの下では主に機械的な単独発射機の署名で支配的 (ルール一貫性 v0.11+ サブコーパスで 78%、コーパス全体で 85%)。 (2) 軍事征服はまれですが、より高速です (12.3 対 18.9 ターン)。 (3) 外交は多作だが、ほとんど完了していない。 (4) 違法行為の ~58% はフォグ/ステート エラーであり、違法行為の割合が信念追跡の尺度になります。 (5) -- 最も確立されておらず、我々が探索的と名付けた唯一のもの -- 弱いリンクは、信頼性と勝利を結びつけます。コーパスは小さく、バランスがとれておらず、左右が入れ替わっていないため、ランキングは予備的な説明的なビューであり、貢献するものではありません。ランキングを超えて、アクションとメッセージのターンごとの追跡により、このコーパスは、LLM が敵対的な不確実性の下でどのように推論するか、つまり信念追跡、自発的欺瞞、およびモデルごとの認知「ペルソナ」についてのレンズとなり、私たちはそれを将来の研究の方向性として組み立てます。リプレイ フォーマット、アイソメトリック ビューア、およびすべてのリプレイをリリースします。リクエストに応じてエンジンソースを提供します。

原文 (English)

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War

We introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base. Three stressors are deliberate: fog of war, full diplomacy (messages, ceasefires, ultimatums; uranium kept secret), and a reliability dimension where every turn must follow a strict JSON schema and an illegal action is silently discarded. The engine is private and each match uses a fresh random map seed and opponent, mitigating the data contamination that affects public benchmarks. Models receive a (near) rule-only prompt with no build-order advice (two tactical seed phrases were present during data collection; see Section 2.7). We benchmark 15 reasoning models across 54 matches and 5,258 actions. Findings: (1) the nuclear rush dominates (78% on the rules-coherent v0.11+ sub-corpus; 85% corpus-wide) with a sole-launcher signature that is largely mechanical under secret-simultaneous launch rules, not a cognitive deterrence failure; (2) military conquest is rare but faster (12.3 vs 18.9 turns); (3) diplomacy is prolific yet almost never consummated; (4) ~58% of illegal actions are fog/state errors, making the illegal-action rate a measure of belief-tracking; (5) -- the least established, and the only one we label exploratory -- a weak link associates reliability with winning. The corpus is small, unbalanced and not side-swapped, so the ranking is a preliminary descriptive view, not a contribution. Beyond ranking, the turn-by-turn traces of actions and messages make the corpus a lens on how LLMs reason under adversarial uncertainty -- their belief-tracking, spontaneous deception, and per-model cognitive "personas" -- which we frame as a future research direction. We release the replay format, an isometric viewer and all replays; engine source on request.

13:00 JSTエージェント

ATRIA: 反復エージェントを使用した適応型追跡可能な ECG レポート

既存の ECG レポート生成は緊密に結合されており、解釈とレポートがエンドツーエンドで融合されているため、ステージレベルの手段を必要とせずにエラーが伝播します。一方、エージェントベースのシステムはタスクを分離しますがシングルパスのままで、以前の出力を再検討することはありません。代わりに、臨床 ECG レポートは反復的に展開され、段階的なコンテキスト統合と双方向編集が必要になります。我々は、臨床医の反復的なワークフローを反映するマルチエージェント ECG レポート システムである \textsc{ATRIA} を紹介します。これは、すべてのレポートの主張をその裏付けとなる証拠に結び付け、その証拠によって裏付けられていないステートメントにフラグを立て、セッション中に追加のコンテキストを組み込み、臨床医が 1 つの不透明な出力を受け入れるのではなく、個々の所見を検証して修正できるようにします。そのエージェントはすでに臨床で使用されている ECG 分析モデルを使用しているため、基礎となる所見は臨床的に信頼できるものです。また、クラウドベースの Web サービスとして、\textsc{ATRIA} はすぐに導入できる状態になっています。ライブ デモとビデオを利用して、4 つのインタラクション ケースを通じて \textsc{ATRIA} をデモンストレーションします。

原文 (English)

ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents

Existing ECG report generation is tightly coupled -- interpretation and reporting fused end-to-end, so errors propagate without stage-level recourse -- while agent-based systems decouple tasks but remain single-pass, never revisiting earlier outputs. Clinical ECG reporting instead unfolds iteratively, requiring progressive context integration and bidirectional editing. We present \textsc{ATRIA}, a multi-agent ECG reporting system that mirrors the clinician's iterative workflow: it binds every report claim to its supporting evidence, flags statements unsupported by that evidence, incorporates additional context mid-session, and lets clinicians verify and revise individual findings rather than accept one opaque output. Because its agents use ECG analysis models already in clinical use, the underlying findings are clinically trustworthy; and as a cloud-based web service, \textsc{ATRIA} is ready for immediate deployment. We demonstrate \textsc{ATRIA} through four interaction cases, with a live demo and video available.

13:00 JST研究/論文

正式検証証明書のサイクル一貫性のあるニューラル説明

正式な検証では、一時的特性の満足または違反を証明する機械チェック可能な証明書が生成されますが、これらの証明書は専門家以外の関係者には不透明なままです。私たちは、検証証明書の忠実な自然言語説明を生成するサイクル一貫性のあるニューラル アーキテクチャを提案します。順方向ネットワーク NN1 は証明書を説明にマッピングし、逆ネットワーク NN2 は説明から証明書を再構築します。シンボリックベリファイアはループを閉じ、微分可能な忠実性プロキシを提供します。ポインタ生成メカニズムは、証明書から状態名を直接コピーすることにより、語彙の基礎を確保します。私たちは、207 の指定州の金融コンプライアンス ドメインから抽出された YES と NO の両方の判定バリアントで、6 つの検証方法 (有界証明、K 帰納、帰納不変式、なげなわ、到達可能性、証人ペア) にわたる 420 のテスト証明書を評価します。ハイブリッド推論時間ルーティング戦略と組み合わせた当社のトレーニング済みアーキテクチャは、サイクル検証済みの健全性 90.0% を達成し、マルチ LLM の少数ショット ベースライン (4 つのフロンティア モデルにわたる 16 の LLM の組み合わせの最良の場合 76.1%) を 13.9 パーセント ポイント上回っています。ニューラル モデルは、12 の判定/種類カテゴリのうち 10 で勝利し、3 つのカテゴリが 100% の健全性に達しました。このアーキテクチャは、860 倍の高速推論 (完全なマルチ LLM ベースラインでは証明書ごとに 160 秒であるのに対し 185 ミリ秒)、オフライン操作、確定的な出力、および推論ごとのコストゼロを提供します。これらの結果は、トレーニングされた専門化が、クラウドベースの推論の展開上の制約を排除しながら、構造化された証明書の説明を求める汎用 LLM よりも優れたパフォーマンスを発揮することを示しています。

原文 (English)

Cycle-Consistent Neural Explanation of Formal Verification Certificates

Formal verification produces machine-checkable certificates that attest to the satisfaction or violation of temporal properties, yet these certificates remain opaque to non-specialist stakeholders. We propose a cycle-consistent neural architecture that generates faithful natural language explanations of verification certificates. A forward network NN1 maps certificates to explanations, and an inverse network NN2 reconstructs certificates from explanations; a symbolic verifier closes the loop, providing a differentiable faithfulness proxy. A pointer-generator mechanism ensures lexical grounding by copying state names directly from the certificate. We evaluate on 420 test certificates spanning six verification methods (bounded proof, k-induction, inductive invariant, lasso, reachability, witness pair) in both YES and NO verdict variants, drawn from a financial compliance domain with 207 named states. Our trained architecture, combined with a hybrid inference-time routing strategy, achieves 90.0% cycle-verified soundness, surpassing a multi- LLM few-shot baseline (76.1% for the best of 16 LLM combinations across four frontier models) by 13.9 percentage points. The neural model wins on 10 of 12 verdict/kind categories, with three categories reaching 100% soundness. The architecture offers 860x faster inference (185 ms vs. 160 s per certificate for the full multi-LLM baseline), offline operation, deterministic outputs, and zero per-inference cost. These results demonstrate that trained specialization outperforms general-purpose LLM prompting for structured certificate explanation, while eliminating the deployment constraints of cloud-based inference.

13:00 JSTエージェント

ポリシー主導の物理層システムのバイレベル長期最適化のための Agentic AI

ネットワーク事業者の変化するポリシー、サービス要件、および厳しいリアルタイム制約により、固定された目的と制約に従って設計された既存の手法は効果がなくなりました。このペーパーでは、適応物理層の問題構成に適用できる入れ子になった 2 レベルの最適化フレームワークである、Agentic 長期パフォーマンス最適化 (Agentic-LTPO) について説明します。重要なアイデアは、エージェント AI を使用して 2 レベルの最適化構造で上位レベルの構成を生成することです。進化するオペレーターのポリシー、環境の概要、および過去の経験が、構造化された下位レベルの最適化問題構成に変換されます。下位レベルでは、リアルタイムの物理層決定のための更新された構成の問題を解決します。セルフリー MIMO ビームフォーミングをユースケースとして考慮し、上位レベルで検索拡張された経験ベースの検証を備えた新しいマルチエージェント意思決定プロセスと、下位レベルのクローズドフォームビームフォーマーを設計することで、Agentic-LTPO を具体化します。実験では、Agentic-LTPO が動的なオペレーター ポリシーに対して強力な適応性を示し、従来の方法と比較してシステムの長期パフォーマンスを効果的に 57.2% 向上させることが実証されました。

原文 (English)

Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems

Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance optimization (Agentic-LTPO), a nested bilevel optimization framework that can be applied to adaptive physical layer problem configuration. The key idea is to employ agentic AI to generate upper-level configurations in a bilevel optimization structure, where evolving operator policies, environment summaries, and historical experiences are translated into structured lower-level optimization problem configurations. The lower level solves the problems with updated configurations for real-time physical-layer decisions. Considering cell-free MIMO beamforming as a use case, we embody Agentic-LTPO by designing a new multi-agent decision process with retrieval-augmented experience-based verification in the upper level, together with a closed-form beamformer in the lower level. Experiments demonstrate that Agentic-LTPO exhibits strong adaptability to dynamic operator policies and effectively enhances the system's long-term performance by 57.2% compared to traditional methods.

13:00 JST研究/論文Claude

集約不変式は連続的なサブグラフのマッチングを高速化できますか?限界、法則、および動的スペクトル指数

スペクトル フィルタリングは、最近 \emph{static} 部分グラフ マッチングに対して大幅な枝刈りを実現しました。ラプラシアン インターレースは、近傍がクエリをホストできない候補を拒否します。私たちは、このような集合構造テストが動的グラフ上で \emph{continuous} サブグラフ マッチング (CSM) を高速化できるかどうかを研究し、3 つの部分に分けて答えます。まず、スペクトルの枝刈りに価値がある場合、遅延的に維持されるスペクトルの境界は実行不可能です。形式化された摂動緩和に対する最も厳格な安全ルールを特徴付け、それが 4 回の更新内で本質的にすべての枝刈り能力を失うことを示します。第 2 に、選択的であれば正確なメンテナンスが手頃です。枝刈りのユーティリティと再計算のコストは頂点間で逆相関しており、ハブは枝刈りをしないことが証明されています。そのため、タッチで小さな近傍スペクトルを再計算すると、更新ごとにマイクロ秒単位で正確なローカル スペクトルが維持され、構築によって完了します。 3 番目に、同一マイナススペクトル コントロールに対する分離された CSM ベンチマークに統合されたテストでは、最大 $51\%$ の候補を削除するか、最大 $47\%$ の更新列挙を安全にスキップしますが、2 つのエンジン、4 つの実際のグラフ、2 つのストリーム タイプ、および $77$ の解決されたクエリにわたって、ゲートのスキップされた第 1 レベルのバインディング (通常はゼロ) を超えて、列挙の中間は変更されません。構築された半径階層化ワークロードにより、例外が存在する場合に機器が例外を検出することが確認されます ($-99.9\%$ 中間、$748\time$ 高速)。集約テストは、候補セット (構築、リスト スキャン) に応じてスケールするものを加速します。決して隣接関係に基づく探索ではありません。 CSM フィルターを評価するための中間不変性手法を抽出し、再利用可能な動的ローカル スペクトル インデックスをリリースします。

原文 (English)

Can Aggregate Invariants Accelerate Continuous Subgraph Matching? Limits, Laws, and a Dynamic Spectral Index

Spectral filtering recently delivered substantial pruning for \emph{static} subgraph matching: Laplacian interlacing rejects candidates whose neighborhoods cannot host the query. We study whether such aggregate structural tests can accelerate \emph{continuous} subgraph matching (CSM) over dynamic graphs, and answer in three parts. First, lazily maintained spectral bounds are infeasible exactly where spectral pruning has value: we characterize the tightest safe rule over a formalized perturbation relaxation and show that even it loses essentially all pruning power within four touching updates. Second, exact maintenance is affordable when selective: pruning utility and recomputation cost are anti-correlated across vertices -- hubs provably never prune -- so recomputing small-neighborhood spectra on touch sustains exact local spectra at microseconds per update, complete by construction. Third, integrated into a decoupled CSM benchmark against an identical-minus-spectra control, the tests remove up to $51\%$ of candidates or safely skip up to $47\%$ of update enumerations, yet enumeration intermediates remain unchanged -- beyond the gates' skipped first-level bindings, typically zero -- across two engines, four real graphs, two stream types, and $77$ solved queries; a constructed radius-stratified workload confirms the instrument detects the exception when one exists ($-99.9\%$ intermediates, $748\times$ faster). Aggregate tests accelerate what scales with candidate sets -- construction, list scans -- never adjacency-guided exploration. We distill an intermediate-invariance methodology for evaluating CSM filters and release a reusable dynamic local-spectra index.

13:00 JSTLLM/生成AIエージェント

ReM-MoA: 推論記憶がエージェント混合のスケーリングを維持する

Mixture-of-Agents (MoA) アーキテクチャは、複数の LLM エージェントを階層化された推論パイプラインに編成することで、推論時間のスケーリングを向上させます。ただし、既存の MoA バリアントは、深さが増加するにつれてゲインを維持できず、劣化、早期のプラトー状態、または飽和を示します。我々は、2 つのメカニズムを通じてスケーリングを維持するメモリ拡張 MoA フレームワークである ReM-MoA を提案します。(1) 比較レビューアー エージェントを使用して、すべてのレイヤーからの推論トレースを永続的に保存してランク付けするランク付き推論メモリ、(2) 成功したトレースと失敗したトレースの異なる組み合わせをさまざまなエージェントに公開し、高品質の推論を伝播しながら探索の多様性を維持するキュレーションされた多様なメモリ ルーティング スキーム。さらに、フロンティア モデルの監視を通じてランキングの品質を向上させる、オプションのマルチドメイン レビュアー蒸留パイプラインを導入します。数学、形式論理、コード、知識、常識に及ぶ 5 つの推論ベンチマークにわたって、ReM-MoA は深さと幅のスケーリングの両方で以前の MoA バリアントを常に上回っており、その利点は深さとともに拡大し、スケーラブルなマルチエージェント推論に欠けている重要なメカニズムとして構造化されたクロスレイヤー推論メモリを確立します。

原文 (English)

ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling

Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents into layered reasoning pipelines. However, existing MoA variants fail to sustain gains as depth increases, exhibiting degradation, early plateauing, or saturation. We propose ReM-MoA, a memory-augmented MoA framework that sustains scaling through two mechanisms: (1) a Ranked Reasoning Memory that persistently stores and ranks reasoning traces from all layers using a comparative Reviewer Agent, and (2) a Curated Diversified Memory Routing scheme that exposes different agents to distinct combinations of successful and failed traces, preserving exploration diversity while propagating high-quality reasoning. We further introduce an optional multi-domain Reviewer distillation pipeline that improves ranking quality through frontier-model supervision. Across five reasoning benchmarks spanning math, formal logic, code, knowledge, and commonsense, ReM-MoA consistently outperforms prior MoA variants across both depth and width scaling, and its advantage widens with depth, establishing structured cross-layer reasoning memory as a key missing mechanism for scalable multi-agent inference.

13:00 JSTLLM/生成AIエージェント

コーディングエージェントのベイジアン制御

最新のコーディング エージェントは、LLM ジェネレーターを、安価な診断や高価な検証ツールなどのさまざまなツールと組み合わせます。ツールの使用に関する決定は通常、固定ルールを使用し、不確実性を無視するオーケストレーターによって管理されます。私たちは、オーケストレーションをコスト重視の逐次仮説テストとして定式化します。ベイジアン コントローラーは、候補の正しさに対する信念を維持し、より多くの証拠を収集するか、候補を絞り込むか、検証するか、中止するかを動的に決定します。 6 つのジェネレーターと 9 つのコーディング ベンチマークにわたって、ベイジアン制御が最も価値があることが証明されるのは、検証にコストがかかり、批評家が有益ではあるが不完全な場合です。制御を超えて、信念状態は、不確実性の定量化において、トークンの確率や生のツールの成功ベースラインを上回る、解釈可能な正確性スコアを生成します。

原文 (English)

Bayesian control for coding agents

Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive verifiers. The tool-use decisions are typically governed by orchestrators that often use fixed rules and ignore uncertainty. We formulate orchestration as cost-sensitive sequential hypothesis testing: a Bayesian controller maintains a belief over candidate correctness and dynamically decides whether to gather more evidence, refine the candidate, verify it, or stop. Across six generators and nine coding benchmarks, Bayesian control proves to be most valuable when verification is costly and critics are informative but imperfect. Beyond control, the belief state yields an interpretable correctness score that outperforms token-probability and raw tool-success baselines for uncertainty quantification.

13:00 JSTLLM/生成AI

CompressKV: リソース効率の高いロングコンテキスト LLM 推論のためのセマンティック検索ガイドによる KV キャッシュ圧縮

ロングコンテキスト大規模言語モデル (LLM) 推論は、メモリ フットプリントとキーバリュー (KV) キャッシュのデコード コストによってますます制約が増えており、リソースに制約のあるハードウェアでの持続可能な展開が制限されています。既存の KV キャッシュ削除方法は通常、GQA ベースの LLM のすべてのヘッドにヒューリスティック トークン スコアリングを適用します。これらのメソッドはアテンション ヘッドのさまざまな機能を無視するため、クリティカル トークンの削除につながり、LLM のパフォーマンスが低下します。この問題に対処するために、GQA ベースの LLM 用のリソース効率の高い KV キャッシュ圧縮フレームワークである CompressKV を提案します。 CompressKV は、すべてのヘッドからのアテンション スコアを集約するのではなく、プロンプトおよび意味的に重要な中間コンテキスト証拠の最初と最後のトークンの両方をキャプチャするセマンティック検索ヘッド (SRH) を特定し、それらを使用して KV ペアを保持する必要があるトークンを選択します。さらに、CompressKV は、レイヤーごとのエビクション エラーのオフライン推定に従って、レイヤー全体にキャッシュ バジェットを割り当てます。 LongBench と Needle-in-a-Haystack での実験では、CompressKV がメモリ バジェット全体にわたって既存の KV キャッシュ削除方法よりも一貫して優れたパフォーマンスを発揮することが示されています。特に、LongBench の質問応答タスクではわずか 3\% の KV キャッシュを使用してフル キャッシュのパフォーマンスの 97\% 以上を維持し、Needle-in-a-Haystack ではわずか 0.7\% の KV ストレージで 90\% の精度を達成します。これらの結果は、ロングコンテキスト LLM 推論におけるリソースとパフォーマンスのトレードオフが改善されたことを示しています。私たちのコードは、https://github.com/TUDa-HWAI/CompressKV で公開されています。

原文 (English)

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware. Existing KV cache eviction methods typically apply heuristic token scoring over all heads in GQA-based LLMs. These methods ignore the different functionalities of attention heads, leading to the eviction of critical tokens and thus degrading the performance of LLMs. To address this issue, we propose CompressKV, a resource-efficient KV-cache compression framework for GQA-based LLMs. Instead of aggregating attention scores from all heads, CompressKV identifies Semantic Retrieval Heads (SRHs) that capture both the initial and final tokens of a prompt and semantically important mid-context evidence, and uses them to select tokens whose KV pairs should be retained. Furthermore, CompressKV allocates cache budgets across layers according to offline estimates of layer-wise eviction error. Experiments on LongBench and Needle-in-a-Haystack show that CompressKV consistently outperforms existing KV-cache eviction methods across memory budgets. Notably, it preserves over 97\% of full-cache performance using only 3\% of the KV cache on LongBench question-answering tasks and achieves 90\% accuracy with just 0.7\% KV storage on Needle-in-a-Haystack. These results demonstrate an improved resource--performance trade-off for long-context LLM inference. Our code is publicly available at: https://github.com/TUDa-HWAI/CompressKV

13:00 JSTエージェント

Latent Bridge: リアルタイム ゲーム エージェント向けの連続的な低速/高速チャネル

一般的なコンピュータで使用されるリアルタイム エージェント (最も要求の厳しいケースとしてゲーム) は、数秒かけて計画を立てながら、数十ミリ秒以内に動作する必要があります。これら 2 つの体制は、遅延と品質のトレードオフの対極に位置します。推論 VLM (Qwen3-VL-8B-Thinking) は効果的に検討しますが、応答ごとに約 1.5 秒かかります。これは 15 Hz の制御ループとしては遅すぎます。対照的に、リアクティブ VLM (MiniCPM-o 4.5) はミリ秒単位で動作しますが、計画の負荷が高いタスクではパフォーマンスが低下します。スケールが一致した 2 つの凍結モデル (9B リアクティブ、8B 推論) を結合し、通信チャネルを唯一のトレーニング可能なコンポーネントとして残します。標準的な結合はテキスト ブリッジ (T) です。低速モデルがサフィックスを書き込み、高速モデルが読み取ります。学習された連続潜在ブリッジ (L) を導入します。これは、低速モデルの残差を高速モデルの入力埋め込み空間に LLaVA スタイルの方法で投影し、テキストの往復を回避します。両方とも高速専用 (F) と比較されます。 7 つの Atari ゲームとドライビング ドメイン (MetaDrive) では、ホールドアウト シードでチャネルごとにアクション デコーダーを調整し、Latent Bridge はすべてのドメインで Text Bridge と同等かそれを上回ります。2 つのゲームを大幅に改善し (MsPacman +57%、RoadRunner +28%)、他の場所でも安全にドロップインできます。両方のチャネルを組み合わせると破壊的な干渉が発生するため (RoadRunner -96%)、1 つのみを使用する必要があります。この利点は非常に予測可能です。ブリッジは、遅い推論がすでに速い反応を上回っている場合 (T > F) にのみ役立ちます。高速のみに対する潜在ゲインとテキスト ゲインは、r=0.93 で一緒に推移します。 MetaDrive はコントロールされたネガティブであり、Text Bridge が価値を追加しないため、Latent Bridge は明らかに不活性です。リプレイ録画と再現可能なパイプラインをリリースします。

原文 (English)

The Latent Bridge: A Continuous Slow-Fast Channel for Real-Time Game Agents

A real-time agent for general computer use - with games as the most demanding case - must act within tens of milliseconds while still planning over seconds. These two regimes sit at opposite ends of the latency-quality tradeoff. A reasoning VLM (Qwen3-VL-8B-Thinking) deliberates effectively but requires ~1.5 s per response - far too slow for a 15 Hz control loop. In contrast, a reactive VLM (MiniCPM-o 4.5) acts in milliseconds but underperforms on planning-heavy tasks. We couple two frozen models of matched scale (9B reactive, 8B reasoning), leaving the communication channel as the sole trainable component. The standard coupling is a Text Bridge (T): the slow model writes a suffix the fast model reads. We introduce a learned continuous Latent Bridge (L) that projects the slow model's residuals into the fast model's input-embedding space in a LLaVA-style manner, avoiding any text round-trip; both are compared against Fast-Only (F). On 7 Atari games and a driving domain (MetaDrive), tuning the action decoder per channel on held-out seeds, the Latent Bridge matches or beats the Text Bridge in every domain: it significantly improves two games (MsPacman +57%, RoadRunner +28%) and is a safe drop-in elsewhere. Combining both channels interferes destructively (RoadRunner -96%), so only one should be used. The benefit is highly predictable: the bridge helps if and only if slow reasoning already beats fast reaction (T > F) - the Latent and Text gains over Fast-Only move together at r=0.93. MetaDrive is the controlled negative, where the Latent Bridge is demonstrably inert because the Text Bridge adds no value. We release replay recordings and reproducible pipelines.

13:00 JSTLLM/生成AI

大規模言語モデルのスケーリング指数の小ささについて

現在の大規模言語モデル (LLM) アプリケーションのスケーリング指数が、エネルギー資源の観点から持続不可能な状況を示している理由について説明します。さらに、このような指数の小ささを、無限データの限界における損失関数の非ゼロ値の無視による数値バイアス (「ペデスタル効果」) に帰することは、持続不可能性の問題を解決しないことを示します。最後に、スケーリング指数に対するデータの滑らかさ (粗さ) の影響について、流体乱流の現象論的モデルとの類似性に基づいてコメントします。

原文 (English)

On the Smallness of the Large Language Models Scaling Exponents

We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in terms of energy resources. We further show that attributing the smallness of such exponents to a numerical bias due to the neglect of a non-zero value of the loss function in the limit of infinite data (``pedestal effect") does not remove the unsustainability issue. Finally, the effects of the smoothness (roughness) of the data on the scaling exponents is commented upon based on an analogy with phenomenological models of fluid turbulence.

13:00 JSTLLM/生成AIDeepSeek

希少疾患診断を加速するための特殊な推論大規模言語モデル: ランダム化 AI 医師支援試験

希少疾患は世界中で何百万人もの人々に影響を及ぼしていますが、専門的な臨床専門知識が不足しているため、タイムリーな診断が依然として公衆衛生上の大きな課題となっています。大規模言語モデル (LLM) は希少疾患の診断をサポートする可能性を示していますが、現在のモデルは不十分な臨床展開可能性、限られた臨床的根拠のある証拠、およびトレーニング データの不足によって制約を受けています。ここでは、希少疾患診断用のオープンソースのコンパクト推論 LLM (32B パラメーター) である RaDaR (Rare Disaster navigatoR) を紹介します。 RaDaR は、公開されているフリーテキスト ケース 49,170 件と、推論強化トレーニングによる合成ケース 104,666 件を使用してトレーニングされました。 RaDaR は、公開ベンチマークと 4 つの外部検証センターにわたって、671B DeepSeek-R1 を含む評価されたオープンソース モデルの中で最も強力なパフォーマンスを示しました。遡及コホートにおいて、RaDaR は症例の 61.06 パーセントで臨床的疑いが文書化される前に最終診断を優先しました。これは、1.87 か月の潜在的なリードタイムと施設内間隔の 50.18 パーセントに相当します。無作為化された医師支援試験では、RaDaR 支援により、インターネット検索のみと比較して医師の希少疾患診断精度が 21.44 パーセント向上しました。合成データアブレーションは、表現型にアンカーされたナラティブが、テストされたデータ範囲内で単調なスケーリング傾向を持つ、ロングテール希少疾患に対する有用なトレーニングシグナルを提供することを示唆しました。 RaDaR とその開発および検証フレームワークを組み合わせることで、展開可能な希少疾患推論モデルと、データ不足下での診断 AI のための再現可能な開発フレームワークが提供されます。

原文 (English)

A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial

Rare diseases affect millions of individuals worldwide, yet timely diagnosis remains a major public health challenge due to scarcity of specialized clinical expertise. While large language models (LLMs) show promise to support rare disease diagnosis, current models are constrained by insufficient clinical deployability, limited clinically grounded evidence, and scarcity of training data. Here we present RaDaR (Rare Disease navigatoR), an open-source, compact reasoning LLM (32B parameters) for rare disease diagnosis. RaDaR was trained with 49,170 publicly available free-text cases and 104,666 synthetic cases with reasoning-enhanced training. RaDaR showed the strongest performance among evaluated open-source models, including the 671B DeepSeek-R1, across public benchmarks and four external validation centers. In a retrospective cohort, RaDaR prioritized the final diagnosis before documented clinical suspicion in 61.06 percent of cases, corresponding to a potential lead time of 1.87 months and 50.18 percent of the within-center interval. In a randomized physician-assistance trial, RaDaR assistance improved physicians' rare-disease diagnostic accuracy by 21.44 percentage points compared with internet search alone. Synthetic-data ablations suggested that phenotype-anchored narratives provide useful training signal for long-tail rare diseases, with a monotonic scaling trend within the tested data range. Together, RaDaR and its development and validation framework provide a deployable rare-disease reasoning model and a reproducible development framework for diagnostic AI under data scarcity.

13:00 JSTエージェントビジネス/資金調達

自律的な評価を備えたコンピュータ使用エージェントの強化学習

Computer-Use Agent (CUA) は、グラフィカル ユーザー インターフェイス内で直接認識して行動することで、高レベルのユーザー目標を実行します。ただし、オープンエンドのデスクトップ環境ではスケーラブルで機械可読な報酬信号がほとんど提供されないため、CUA の強化学習は依然として困難です。タスクの成功は多くの場合視覚的に根拠があり、手作りの報酬関数や高密度の手動ラベルで指定するのは困難です。我々は、GUI エージェントのスケーラブルな監視信号として自律的な視覚言語評価を使用する RL 微調整フレームワークを提案します。最終的なスクリーンショットと元の指示が与えられると、ビジョン言語モデルはタスクの完了を判断し、ポリシーの最適化中にタスク固有のヒューリスティックや手動ラベルを使用せずに最終的なフィードバックを提供します。自律型評価器は不完全であるため、そのフィードバックをノイズの多いバイナリ報酬チャネルとしてモデル化し、近接ポリシー最適化のためのノイズ補正された報酬推定器を導出します。 macOSWorld、Windows Agent Arena、および OSWorld にわたる実験では、修正された評価者の報酬がゼロショットのベースラインと生の評価者の報酬の両方を上回り、成功率がゼロショットのパフォーマンスより平均 12.6 ポイント、生の評価者の微調整よりも 5.1 ポイント向上したことが示されています。これらの結果は、評価者のノイズが明示的にモデル化され補正されている場合、自律評価が GUI 環境における RL の実用的な報酬信号として機能する可能性があることを示唆しています。

原文 (English)

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is often visually grounded and hard to specify with handcrafted reward functions or dense manual labels. We propose an RL fine-tuning framework that uses autonomous vision-language evaluation as a scalable supervision signal for GUI agents. Given a final screenshot and the original instruction, a Vision-Language Model judges task completion and provides terminal feedback without task-specific heuristics or manual labels during policy optimization. Because autonomous evaluators are imperfect, we model their feedback as a noisy binary reward channel and derive a noise-corrected reward estimator for Proximal Policy Optimization. Experiments across macOSWorld, Windows Agent Arena, and OSWorld show that corrected evaluator rewards outperform both zero-shot baselines and raw evaluator rewards, improving success rates by an average of 12.6 percentage points over zero-shot performance and 5.1 points over raw evaluator fine-tuning. These results suggest that autonomous evaluation can serve as a practical reward signal for RL in GUI environments when evaluator noise is explicitly modeled and corrected.

13:00 JSTLLM/生成AIエージェント研究/論文

マルチエージェント LLM システム用のガバナド共有メモリ

マルチエージェント LLM 環境には、共有ナレッジ管理のための堅牢なメカニズムが必要です。このペーパーでは、フリート メモリの問題を形式化し、不正な漏洩、古い伝播、矛盾の持続、来歴の崩壊という 4 つの基本的な障害モードを特定します。これらに対処するために、スコープ指定された取得、一時的なスーパーセッション、来歴追跡、およびポリシーに基づいたメモリ伝播などの明示的なシステムレベルのプリミティブを定義します。これらのプリミティブは、運用マルチテナント メモリ サービスである MemClaw に実装され、4 つのガバナンス次元をテストする再現可能なハーネスである ArgusFleet によって評価されます。この調査では、ベースラインの比較ではなく、実際の運用サービスを測定し、現実世界のアーキテクチャに関する洞察と否定的な結果を強調しています。主な評価結果 来歴: ホップあたり 1 秒未満のレイテンシーで正しいライター ID を使用して、深さ 4 の派生チェーンを 100% 再構築することに成功しました。伝播: フリート間の漏洩がゼロで、フリート内の高い可視性を実証しました。強力な書き込みモードでは、可視への書き込み遅延が 1 回の検索ラウンドトリップに最適化されました。運用アーキテクチャの問題が発見されました 非対称スコープの強制: テナントの分離は維持されましたが、サブテナント スコープは当初、エージェント スコープの資格情報に対する ID による直接の GET リクエストでバイパスされていました (調査中に開示および修正されました)。パイプライン順序付けの競合: 矛盾スーパーセッションは許可された書き込みに対して機能しますが、同期準重複ゲートは、非同期矛盾検出器が矛盾書き込みを評価する前に、矛盾書き込みを拒否する可能性があります。結論: ロングコンテキストの取得だけでは、本番環境のマルチエージェントメモリには不十分です。管理された共有メモリには明示的なシステムレベルの抽象化が必要であり、設計のみの処理では見逃されていた施行やパイプラインの順序付けの失敗を明らかにするにはライブ評価が不可欠です。

原文 (English)

Governed Shared Memory for Multi-Agent LLM Systems

Multi-agent LLM environments require robust mechanisms for shared knowledge management. This paper formalizes the fleet-memory problem and identifies four foundational failure modes: unauthorized leakage, stale propagation, contradiction persistence, and provenance collapse. To address these, we define explicit systems-level primitives: scoped retrieval, temporal supersession, provenance tracking, and policy-governed memory propagation. These primitives are implemented in MemClaw, a production multi-tenant memory service, and evaluated via ArgusFleet, a reproducible harness testing four governance dimensions. Rather than a baseline comparison, this study measures a live production service, emphasizing real-world architectural insights and negative results. Key Evaluation Results Provenance: Successfully reconstructed 100% of depth-four derivation chains with correct writer identity at sub-second per-hop latency. Propagation: Demonstrated high intra-fleet visibility with zero cross-fleet leakage. Under strong write mode, write-to-visible latency was optimized to a single search round-trip. Production Architectural Issues Discovered Asymmetric Scope Enforcement: Tenant isolation held, but sub-tenant scope was initially bypassed on direct GET-by-id requests for agent-scoped credentials (disclosed and remediated during the study). Pipeline Ordering Conflict: While contradiction supersession works for admitted writes, a synchronous near-duplicate gate can prematurely reject contradictory writes before the asynchronous contradiction detector can evaluate them. Conclusion: Long-context retrieval alone is insufficient for production multi-agent memory. Governed shared memory demands explicit systems-level abstractions, and live evaluation is vital to expose enforcement and pipeline-ordering failures missed by design-only treatments.

13:00 JSTエージェント

GUI と CLI: 画面のみおよびスキルを介したコンピュータ使用エージェントにおける実行のボトルネック

コンピュータ使用エージェントは、グラフィカル インターフェイスまたはプログラム コマンド インターフェイスを通じてソフトウェア タスクを実行できますが、既存の評価では、タスク、初期状態、検証者、許可されるアクションの違いにより、対話形式が混同されています。 18 のアプリケーションと 12 のワークフロー カテゴリにわたる 440 のデスクトップ タスクの一致する実行層ベンチマークを導入します。このベンチマークでは、画面のみの GUI エージェントとスキル媒介の CLI エージェントが、モダリティ ネイティブのアクションに制限されながら、同一の目標、状態、および最終状態の検証子を受け取ります。この制御された設定では、最強の GUI エージェントは 59.1% の完全合格率に達し、最強のオリジナル スキル CLI エージェントの 48.2% を上回ります。ただし、検証者主導のスキル強化により CLI の成功率は 69.3% に上昇し、CLI の不足の多くはモデルの機能だけではなく、スキルのカバー範囲が不完全であることに起因していることがわかります。これらの結果は、GUI と CLI が異なる実行ボトルネックを露呈していることを示唆しています。GUI エージェントは長期ワークフローにわたる信頼性の高い根拠のある対話によって制限されるのに対し、CLI エージェントはスキル インターフェイスの適用範囲とスケーラビリティによって制限されます。

原文 (English)

GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents

Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with differences in tasks, initial states, verifiers, and permitted actions. We introduce a matched execution-layer benchmark of 440 desktop tasks across 18 applications and 12 workflow categories, where screen-only GUI agents and skill-mediated CLI agents receive identical goals, states, and final-state verifiers while being restricted to modality-native actions. In this controlled setting, the strongest GUI agent reaches a 59.1% full pass rate, outperforming the strongest original-skill CLI agent at 48.2%; however, verifier-guided skill augmentation raises CLI success to 69.3%, showing that much of the CLI deficit comes from incomplete skill coverage rather than model capability alone. These results suggest that GUI and CLI expose different execution bottlenecks: GUI agents are limited by reliable grounded interaction over long-horizon workflows, whereas CLI agents are limited by the coverage and scalability of their skill interfaces.

13:00 JST研究/論文

クオンツ・コンバージェンス:体系的な株式選択のための古典的なバリュー投資と最新のファクター・モデルの橋渡し

現代の金融は、株式市場のパターンを見つけるために複雑な機械学習モデルに大きく依存しています。しかし、これらの AI モデルが複雑になるにつれて、実際の永続的な価値を持つ企業を見つけるのではなく、短期的な市場のノイズを記憶することがよくあります。私たちは、ベンジャミン・グレアムの古典的なバリュー投資ルールが、これらの最新モデルを抑制する数学的な「ローパスフィルター」として機能するかどうかをテストするためにこの研究を設計しました。私たちは、純粋なグラハム ルール、最新の市場要因、その両方の組み合わせという 3 つの異なる機能セットを構築し、20 年間の S&P 500 データを使用して非常に複雑なモデル (XGBoost および AutoGluon) に対してテストしました。 4 年間のテスト期間 (2022 年 3 月から 2026 年 3 月まで) にわたって厳密なバイアンドホールド戦略を適用した結果、より複雑なアルゴリズムが必ずしも勝利するとは限らないことがわかりました。 AutoGluon モデルは高いリターン (222.68%) を獲得しましたが、市場が暴落する直前に不安定なハイテク株を購入したため、39.78% という大幅な下落に見舞われました。一方、純粋な Graham Random Forest は、はるかに少ないリスク (1.38 Calmar Ratio) で最高の全体収益 (232.13%) を達成しました。さらに、結合ランダム フォレストは勢いとグラハム ルールをうまく組み合わせることで、テストしたモデルの中で最も低い最大ドロップ (34.53%) を維持しながら、202.91% のリターンを実現しました。結局のところ、この研究は、グレアムの「安全域」が時代遅れではないことを証明しています。これは実際、現代の AI が過度のリスクを負うことを防ぐ非常に効果的な方法です。

原文 (English)

Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection

Modern finance relies heavily on complex machine learning models to find patterns in the stock market. However, as these AI models get more complicated, they often memorize short-term market noise instead of finding companies with real, lasting value. We designed this research to test if Benjamin Graham's classic value investing rules could act as a mathematical "low-pass filter" to keep these modern models in check. We built three different sets of features - pure Graham rules, modern market factors, and a mix of both - and tested them against highly complex models (XGBoost and AutoGluon) using 20 years of S&P 500 data. By applying a strict buy-and-hold strategy over a four-year test period (March 2022 to March 2026), the results showed that more complex algorithms do not always win. While the AutoGluon model captured high returns (222.68%), it suffered a substantial 39.78% drop because it bought volatile tech stocks right before the market crashed. On the other hand, the pure Graham Random Forest achieved the highest overall return (232.13%) with much less risk (1.38 Calmar Ratio). Furthermore, the Combined Random Forest successfully mixed momentum with Graham's rules, making a 202.91% return while keeping the lowest maximum drop (34.53%) of any model tested. Ultimately, this research proves that Graham's "margin of safety" isn't outdated; it is actually a highly effective way to prevent modern AI from taking on too much risk.

13:00 JSTLLM/生成AI

LLM が法的コンテキスト オブジェクトの入力を求められる 詳細: 刑事法的コンテキストにおける小規模なオンプレミス LLM からの過剰拒否

法的文脈における LLM の使用の妥当性については依然として倫理的および法的な議論の対象となっていますが、法律専門家はすでに、翻訳と再定式化のみを目的として、個人用 LLM を実験しています。ただし、そのような一見無害な使用法でも、LLM アシスタントが特定のトピックに関する支援を選択的に拒否した場合、ケースの処理速度によってバイアスが生じる可能性があります。このようなバイアスをより適切に予測するために、デバイス上のアシスタントとして使用される可能性が最も高いいくつかの最新の小型 LLM を調査し、過剰な拒否が法的要求に与える影響を評価します。驚くべきことに、権威スタイルの接頭辞(「あなたは国家最高裁判所の助手を務めています」、「[...]弁護人」)は、接頭辞なしのベースラインに比べて拒否率を体系的に2〜20倍増加させることがわかりました。一方、既知のロールプレイジェイルブレイク接頭辞は複合的な効果を示し、一部のモデルでは拒否が急激に増加し、他のモデルでは拒否がほとんど変化しません。この発見は、オンプレミスで展開可能な小規模な LLM は、実際の機関ユーザーが自然に導入する可能性のあるコンテキスト フレームワークの下では不安定であることを示唆しており、バイアスの機会を最小限に抑えるためにはさらなる調査が不可欠です。

原文 (English)

LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context

While the validity of LLMs' use in the legal context remains subject to ethical and legal debate, legal professionals are already experimenting with personal LLMs, if only for translation and reformulation. However, even such a seemingly innocuous use can introduce biases through case processing speed if LLM assistants selectively refuse assistance on certain topics. To better anticipate such biases, we investigate several modern small LLMs that are most likely to be used as on-device assistants, to assess the impact of overrefusal on legal prompts. Surprisingly, we find that authority-style prefixes (``you are acting as an assistant of the national supreme court'', ``[...] defense lawyer'') systematically increase refusal rates by 2--20x over the no-prefix baseline, while a known role-play jailbreak prefix shows mixed effects, sharply increasing refusals in some models and barely shifting them in others. The finding suggests that small on-prem deployable LLMs are unstable under contextual framings that a real institutional user might naturally introduce, and further investigation is essential to minimize opportunities for bias.

13:00 JSTLLM/生成AIビジネス/資金調達Llama

AdversaBench: 複数の裁判官による確認とモデル間の移行性を備えた自動化された LLM レッドチーム化

大規模な言語モデルの敵対的評価をスケーリングするには、ハード入力を生成する方法と、結果として生じる失敗が本物であることを確認する信頼性の高い方法の両方が必要です。 AdversaBench は、5 つの構造化された演算子でシード プロンプトを変更し、ターゲット モデルをクエリし、メタ ジャッジ タイブレーカーを備えた 3 人のジャッジ パネルを通じて失敗を確認する、エンドツーエンドのレッド チーム パイプラインです。推論、指示に従い、ツールの使用という 3 つのカテゴリにわたる 45 のシードに関する実験を報告します。すべてのシードで失敗が確認されました。 4 つの発見が際立っています。まず、オペレーターの有効性はカテゴリによって大きく異なります。inject_distractor のスコアは、指示に従うシードでは 0.00 の平均報酬ですが、推論とツールの使用では 0.80 ~ 0.83 です。第 2 に、バイナリの失敗率が難しさを隠しています。命令に従うシードでは、攻撃者の反復回数が平均 2.4 回であるのに対し、他のカテゴリでは 1.1 回であり、生存曲線にギャップが見られます。第三に、80 ~ 87% というペアごとのジャッジの一致は、ラベルの歪みによりほぼゼロのコーエンのカッパと共存します。カテゴリレベルの不一致率の方が有益です。第 4 に、Llama 3.1 8B に対して生成された敵対的プロンプトはゼロショットを Llama 3.3 70B に転送します。これは、変異がモデル固有の弱点ではなく一般的な動作パターンを悪用していることを示唆しています。コード、データセット、分析スクリプトは https://github.com/khanak0509/AdversaBench で入手できます。

原文 (English)

AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability

Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real. We present AdversaBench, an end-to-end red-teaming pipeline that mutates seed prompts with five structured operators, queries a target model, and confirms failures through a three-judge panel with a meta-judge tiebreaker. We report experiments on 45 seeds across three categories: reasoning, instruction-following, and tool use. Every seed produced a confirmed failure. Four findings stand out. First, operator effectiveness varies sharply by category: inject_distractor scores 0.00 mean reward on instruction-following seeds but 0.80-0.83 on reasoning and tool-use. Second, binary failure rate hides difficulty: instruction-following seeds required 2.4 attacker iterations on average versus 1.1 for other categories, a gap visible in survival curves. Third, pairwise judge agreement of 80-87% coexists with near-zero Cohen's kappa due to label skew; category-level disagreement rates are more informative. Fourth, adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B, suggesting the mutations exploit general behavioral patterns rather than model-specific weaknesses. Code, dataset, and analysis scripts are available at https://github.com/khanak0509/AdversaBench .

13:00 JSTエージェント

ASALT: マルチエージェント強化学習における横方向伝達のための適応的状態調整

マルチエージェント強化学習 (MARL) は、協調的、競争的、または混合の目的を追求する複数のエージェントをトレーニングするという問題に対処します。これまでの研究では、MARL のソース ドメインとターゲット ドメイン間の転移学習を調査しました。ただし、既存のアプローチの大部分は、観測空間とグローバル状態空間の次元がドメイン全体で同一でなければならないという制約を課します。この論文では、ソース ドメインとターゲット ドメイン間の状態空間次元の不一致に明示的に対応する方法を紹介します。提案されたアプローチである ASALT には、ターゲット ドメインの観察とグローバルな状態を共有埋め込み空間にマッピングする観察レベルと状態レベルのアダプターの両方が組み込まれており、それによってアクターと批評家の両方にわたるより効果的な知識の伝達が可能になります。これらのアダプターは、異種ドメイン間での効率的な戦略転送をサポートするエンベディングを生成できます。標準ベンチマーク環境での複数の構成での実験結果は、ASALT がサンプル効率と協調設定でのグローバルな収益の点で既存のベースラインを上回っていることを示していますが、その有効性はソース ドメインとターゲット ドメイン間の不一致の程度に依存します。さらに、我々の調査結果は、ASALT がネガティブな移転を軽減することを示しています。ネガティブな移転は、異なる観察空間と行動空間を持つドメイン間でポリシーを移転する際に大きな障害となることがよくあります。

原文 (English)

ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning

Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives. Prior work has investigated transfer learning between source and target domains in MARL; however, the majority of existing approaches impose the constraint that the dimensionalities of the observation space and the global state space must be identical across domains. In this paper, we introduce a method that explicitly accommodates mismatched state-space dimensionalities between source and target domains. The proposed approach, ASALT, incorporates both observation-level and state-level adapters that map the target-domain observations and global states into a shared embedding space, thereby enabling more effective transfer of knowledge across both actors and critics. These adapters can generate embeddings that support efficient strategy transfer across heterogeneous domains. Experimental results on multiple configurations in standard benchmark environments demonstrate that ASALT surpasses existing baselines in terms of sample efficiency and global return in cooperative settings, but its effectiveness depends on the degree of mismatch between source and target domains. Furthermore, our findings indicate that ASALT mitigates negative transfer, which frequently constitutes a major obstacle when transferring policies between domains with differing observation and action spaces.

13:00 JST研究/論文

深層学習を使用した、不確実性を認識したアルツハイマー病進行の長期的予測

アルツハイマー病の進行の縦断的モデリングは、最も可能性の高い次の診断だけでなく、時間の経過とともに患者がどのように変化するか、そしてその予測がどれほど信頼できるかを説明できる場合にのみ臨床的に役立ちます。ほとんどの深層学習アプローチは、この問題を 1 段階の分類に落とし込み、正常な認知機能、軽度の認知障害、認知症をフラットなカテゴリーとして扱いますが、今後の訪問で不確実性がどのように蓄積するかについての洞察は限定的です。我々は、通常の診断予測、マルチ水平軌道生成、および分解された不確実性推定を組み合わせた確率的フレームワークを提案します。 Temporal Fusion Transformer エンコーダには、CORAL 順序出力層、非対称損失重み付け、およびコンバーターのオーバーサンプリングが適用され、疾患段階の順序を尊重し、MCI から認知症への移行に対する感度が向上します。学習された患者コンテキスト表現に基づいて、自己回帰混合密度ネットワークは、診断状態、CDR ボックス和、MMSE 方向、海馬体積の 5 年間の確率的軌跡を生成します。 ADNI では、このモデルは次回受診診断予測において線形ベースライン、再発ベースライン、およびトランスフォーマー ベースラインを上回り、MCI と認知症の区別において最も優れた効果を発揮します。生成された軌跡は、名目上の 90% に近い信頼区間の範囲を達成し、予測範囲全体にわたって不確実性が拡大し、予想されるアルツハイマー病の進行と一致するバイオマーカーのダイナミクスを実現します。さらに、解析的混合分散と、最も強力なエンコーダーの多様性と出力レベルの認識論的信号を提供する 5 メンバーのブートストラップ アンサンブルを使用して、認識論的不確実性から偶然の不確定性を分離します。認識論的不確実性は、まれな進行原型、MCI、認知症患者で高く、OASIS-3 の外部評価下では、予測誤差とともに増加します。

原文 (English)

Uncertainty-Aware Longitudinal Forecasting of Alzheimer's Disease Progression Using Deep Learning

Longitudinal modelling of Alzheimer's disease progression is clinically useful only if it can describe not just the most likely next diagnosis, but how a patient may evolve over time and how reliable that forecast is. Most deep learning approaches reduce this problem to single-step classification, treating cognitively normal, mild cognitive impairment, and dementia as flat categories while providing limited insight into how uncertainty accumulates across future visits. We propose a probabilistic framework that combines ordinal diagnosis prediction, multi-horizon trajectory generation, and decomposed uncertainty estimation. A Temporal Fusion Transformer encoder is adapted with a CORAL ordinal output layer, asymmetric loss weighting, and converter oversampling to respect disease-stage ordering and improve sensitivity to MCI-to-dementia transitions. Conditioned on the learned patient-context representation, an autoregressive Mixture Density Network generates five-year probabilistic trajectories for diagnosis state, CDR Sum of Boxes, MMSE orientation, and hippocampal volume. On ADNI, the model outperforms linear, recurrent, and transformer baselines for next-visit diagnosis prediction, with the strongest gains on MCI-versus-dementia discrimination. Generated trajectories achieve near-nominal 90% credible interval coverage, widening uncertainty across the forecast horizon, and biomarker dynamics consistent with expected Alzheimer's disease progression. We further separate aleatoric from epistemic uncertainty using analytic mixture variance and a five-member bootstrap ensemble, which provides the strongest encoder diversity and output-level epistemic signal. Epistemic uncertainty is higher for rare progression archetypes, MCI and dementia patients, and under external evaluation on OASIS-3, where it increases alongside prediction error.

13:00 JSTLLM/生成AI

ScaleToT: 数十億規模の低アクティビティ ユーザー モデリングのための構造化 LLM 推論の一般化

正確なユーザー モデリングは、多くの場合、豊富なインタラクション履歴に依存しますが、何十億人ものアクティビティの低いユーザーにとっては、これらの履歴は利用できません。大規模言語モデル (LLM) は静的プロファイルから潜在的なユーザー状態を推測できますが、プロファイルがまばらな場合、この推論は信頼できなくなり、LLM を数十億のユーザーに適用すると法外な費用がかかります。私たちは、LLM で処理された小さなサブセットから構造化推論を学習し、それを広範な低アクティビティ ユーザー集団に拡張する ScaleToT を紹介します。推論の信頼性を向上させるために、ScaleToT は、有界エントロピーに基づく思考ツリー (ToT) 改良手順を使用して型付きユーザー状態チェーンを構築します。この構造化された推論を疎なプロファイルから使用できるようにするために、教師が厳選したチェーンを使用して、教師あり微調整 (SFT) と結果主導型セグメント認識暗黙的報酬ポリシー最適化 (OSIPO) を通じて静的プロファイルで生徒モデルをトレーニングします。次に、ScaleToT は学生の推論表現を軽量プロファイル エンコーダに転送し、LLM 推論を行わずに残りのユーザーに共有推論信号を提供します。 10億規模の広告展開における生涯価値(LTV)予測に基づいてScaleToTを評価します。ランダム化されたオンライン A/B テストでは LT30 が 6.738\% 増加しましたが、オフライン推論では潜在的な母集団の 7.32\% しかカバーされず、全母集団推論と比較してコンピューティング コストが大幅に削減されました。

原文 (English)

ScaleToT: Generalizing Structured LLM Reasoning for Billion-Scale Low-Activity User Modeling

Accurate user modeling often depends on rich interaction histories, which are unavailable for billions of low-activity users. Large Language Models (LLMs) can infer latent user states from static profiles, but this reasoning becomes unreliable when profiles are sparse, and applying an LLM to billions of users is prohibitively expensive. We present ScaleToT, which learns structured reasoning from a small LLM-processed subset and extends it to the broader low-activity user population. To improve reasoning reliability, ScaleToT constructs typed user-state chains with a bounded entropy-guided Tree-of-Thought (ToT) refinement procedure. To make this structured reasoning usable from sparse profiles, the teacher-curated chains are used to train a student model on static profiles through supervised fine-tuning (SFT) and Outcome-Driven Segment-Aware Implicit Reward Policy Optimization (OSIPO). ScaleToT then transfers the student's reasoning representations to a lightweight profile encoder, providing shared reasoning signals for the remaining users without LLM inference. We evaluate ScaleToT on lifetime value (LTV) prediction in a billion-scale advertising deployment. A randomized online A/B test increased LT30 by 6.738\%, while offline reasoning covered only 7.32\% of the potential population, greatly reducing compute cost compared with full-population reasoning.

13:00 JST研究/論文

AI トークンノミクス: 基盤モデルにおけるトークン、計算、価格設定の経済学

トークンは、情報処理、計算、メモリ使用、エネルギー消費、価格設定、経済的価値を結び付ける、現代の基盤モデル サービスの実際的な会計単位となっています。このペーパーでは、AI トークンノミクスのフレームワークを開発します。これは、AI システム全体でトークンがどのように生成、消費、価格設定、割り当て、最適化されるかについての研究です。私たちは、トークンレベルの技術コストを、ワークフローレベルの生産機能、企業リソースの割り当て、測定および計測方法、新興市場の設計上の疑問に結び付けます。このフレームワークは、トークンの支出と経済的価値が異なることを示しています。価値は限界生産性、ワークフローの位置、隠れた推論活動、リスク、および下流の伝播効果に依存します。この論文は、隠れたトークンの測定、経験的校正、トークンの生産性、ダイナミックな割り当て、およびトークンベースの市場におけるオープンな研究の方向性を特定して締めくくられています。

原文 (English)

AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation Models

Tokens have become the practical accounting unit for modern foundation model services, linking information processing, computation, memory use, energy expenditure, pricing, and economic value. This paper develops a framework for AI tokenomics: the study of how tokens are generated, consumed, priced, allocated, and optimized across AI systems. We connect token-level technical costs to workflow-level production functions, enterprise resource allocation, measurement and instrumentation methods, and emerging market-design questions. The framework shows that token expenditure and economic value are distinct: value depends on marginal productivity, workflow position, hidden reasoning activity, risk, and downstream propagation effects. The paper concludes by identifying open research directions in hidden-token measurement, empirical calibration, token productivity, dynamic allocation, and token-based markets.

13:00 JST研究/論文

オントロジーベースのデータアクセスにおけるクエリの抽象化

オントロジーベースのデータアクセス (OBDA) では、オントロジーへのマッピングを介して複数のデータソースが統合されます。存在ルールと特定の回答セマンティクスに基づいて OBDA 設定を検討します。私たちは、データクエリをオントロジー層に変換することによってデータクエリを抽象化することで構成されるクエリ抽象化の最近の問題に取り組みます。完全な抽象化は存在しない可能性があるため、最小限に完全で最大限に健全な抽象化の概念が導入されています。私たちは、限定された形式の不等式とデータベース定数をマークする特別な述語を使用した UCQ の拡張内の抽象化を研究します。この拡張は対象となる問題の複雑さの増加にはつながりませんが、最小限に完全な抽象化を表現できるため、存在する場合には完全な抽象化を表現できます。また、データ交換から生じる最大の回復の概念と新たな関係を作ることによって、最大限に健全な抽象化を特徴づけます。

原文 (English)

Abstractions of Queries in Ontology-Based Data Access

In ontology-based data access (OBDA), multiple data sources are integrated via mappings to an ontology. We consider an OBDA setting based on existential rules and the certain answer semantics. We address the recent issue of query abstraction, which consists of abstracting data queries by translating them to the ontology layer. Since a perfect abstraction may not exist, the notions of minimally complete and maximally sound abstractions have been introduced. We study abstractions within an extension of UCQs with a limited form of inequality and a special predicate marking database constants. While this extension does not lead to an increased complexity of the problems of interest, it is able to express minimally complete abstractions, hence perfect abstractions when they exist. We also characterize maximally sound abstractions by making a new connection with the notion of maximum recovery stemming from data exchange.

13:00 JST研究/論文

CQ が失敗した場合: OE-Assist を使用した CQ 検証の課題

コンピテンシー質問 (CQ) は、CQ 検証の中心的なコンポーネントであり、オントロジーが一連の自然言語質問に対して評価され、オントロジーの意図された目的が適切にモデル化されているかどうかを判断する確立されたプロセスです。ただし、CQ 検証には、言語のニュアンスを注意深く解釈し、正式なオントロジー構造と正確に一致させる必要があるため、多くの場合時間がかかり、エラーが発生しやすくなります。 CQ のあいまいさと複雑さによってこのプロセスがさらに複雑になる可能性があり、モデリングの決定と検証の結果に一貫性がなくなる可能性があります。このペーパーでは、CQ を困難にしている原因と、CQ 検証プロセスにおけるユーザーのパフォーマンスを向上させるための可能な解決策を調査します。オントロジー評価をサポートする LLM アシスタントを使用して、20 のタスクに対して CQ 検証を実行した 19 人の参加者のデータを実験しました。この結果は、オントロジー エンジニアリング プロセスの後の段階でのあいまいさや過度の複雑さを回避するために、CQ を公開する前に CQ を改良するツールの必要性を示しています。

原文 (English)

When CQs Go Wrong: Challenges in CQ Verification with OE-Assist

Competency Questions (CQs) are the central component of CQ-verification, an established process in which an ontology is evaluated against a set of natural language questions to determine whether the intended purpose of the ontology has been properly modelled. However, CQ-verification is often time-consuming and error-prone, as it requires careful interpretation of linguistic nuances and precise alignment with formal ontology constructs. Ambiguities and complexity in CQs can further complicate this process, leading to inconsistent modelling decisions and verification outcomes. In this paper, we investigate what makes a CQ challenging and possible solutions to enhance the users' performance in the CQ-verification process. We experimented with the data of 19 participants who performed CQ-verification on 20 tasks using an LLM assistant to support ontology evaluation. The results show the necessity of a tool to refine CQs before publishing them to avoid ambiguity or excessive complexity in later phases of the ontology engineering process.

13:00 JST研究/論文

Themis: 人間のフィードバックによる強化学習のための説明可能な AI 対応フレームワーク

強化学習 (RL) システムを安全にトレーニングすることは本質的に困難であり、望ましくない動作を回避する保証はありません。これに対する最も効果的な防御策は、(i) 説明可能性による透明性と、(ii) 人間のフィードバックによる調整です。どちらも有望な結果を示していますが、現在、それらを組み合わせた公的に利用可能なフレームワークはありません。これに対処するために、ヒューマン フィードバックからの強化学習のための XAI 対応のテストおよび評価フレームワークである Themis を紹介します。 Themis は 200 以上の広く使用されている環境をサポートしており、RL、透明性、および位置合わせの実験用に簡単に構成できます。私たちの結果は、Themis が人間の好みを使用して、環境の真の報酬シグナルと一致またはそれを上回る報酬モデルをトレーニングできることを示しています。また、人間からのフィードバックを収集し、実験を管理するためのクラウドベースのプラットフォームも提供しています。ユーザーフレンドリーで自動スケーラブルで、追加の開発オーバーヘッドなしで複数の実験にわたる大規模な参加者グループをサポートします。テストでは、Themis が小規模な商用マシンでの連続実験で 1,000 人のユーザーをサポートできることが示されています。

原文 (English)

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback. While both show promising results, no publicly available framework currently combines them. To address this, we introduce Themis, an XAI-enabled testing and evaluation framework for Reinforcement Learning from Human Feedback. Themis supports over 200 widely used environments and is easily configurable for experiments in RL, transparency, and alignment. Our results show that Themis can train reward models that match or outperform the environment's true reward signal using human preferences. We also provide a cloud-based platform for collecting human feedback and managing experiments. It is user-friendly, auto-scalable, and supports large participant groups across multiple experiments without extra development overhead. Tests show Themis can support one thousand users in back-to-back experiments on a modest commercial machine.

13:00 JSTエージェント

SAFARI: 積極的な調査によるロングホライズンのエージェント障害の特定の拡張

自律エージェントがますます複雑なマルチステップ、マルチエージェント タスクに取り組むにつれて、その実行軌跡は最大のコンテキスト ウィンドウの制約さえも超えて拡大しています。エージェントの障害を効果的に診断する現在の方法では、LLM のコンテキスト ウィンドウに完全な軌跡が読み込まれますが、注意力が低下し、エージェントのトレースが必然的にコンテキストの制限を超えると失敗します。これに対処するために、線形コンテキストの読み込みをツールで拡張された診断ループに置き換えるフレームワークである SAFARI (Scaling long-horizo​​n Agentic Fault AttRibution via active Investigation) を導入します。 LLM に、クロスターン推論のための永続的短期メモリ (STM) とともに軌道セグメントを読み取って検索するための専用ツールボックスを装備することで、SAFARI は診断精度をアーキテクチャ上のコンテキスト制限から効果的に切り離します。私たちの実験では、SAFARI が 100 万トークン予算内の Who&When データセットで 20%、25,000 トークン予算内の TRAIL GAIA サブセットで 19% 最先端の結果を上回っていることが実証されました。最も重要なのは、ターゲット障害がモデルのネイティブ コンテキスト ウィンドウの 5 倍を超えて存在する場合でも、SAFARI は 0.58 の精度を維持します。これは、従来の評価器が完全に失敗するシナリオです。

原文 (English)

SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation

As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows. Current methods for effectively diagnosing agent failures load the full trajectory into an LLM's context window, which suffers from attention dilution and fails when agentic traces inevitably exceed context limits. To address this, we introduce SAFARI (Scaling long-horizon Agentic Fault AttRibution via active Investigation), a framework that replaces linear context loading with a tool-augmented diagnostic loop. By equipping LLMs with a specialized toolbox to read and search trajectory segments alongside a persistent Short-Term Memory (STM) for cross-turn reasoning, SAFARI effectively decouples diagnostic accuracy from architectural context limits. Our experiments demonstrate that SAFARI outperforms state-of-the-art results by 20% on the Who&When dataset within a 1M token budget, and by 19% on TRAIL GAIA subset on a 25K token budget. Most significantly, SAFARI maintains a 0.58 precision even when the target fault resides 5x beyond the model's native context window, a scenario where traditional evaluators fail entirely.

13:00 JST研究/論文

CineCap: 映画ビデオ キャプション用の時空間アンカーを使用した構造化推論

映画のキャプションは、カメラの動き、ショットのサイズ、被写界深度、構成、撮影角度などの専門的な映画言語の概念を使用して、ビデオがどのように撮影されるかを説明することを目的としています。この機能は、きめ細かいビデオの理解と制御可能な映画品質のビデオ生成にとって重要ですが、既存のマルチモーダル大規模言語モデルではまだ十分に検討されていません。映画の理解に関する質問応答ベースの評価とは異なり、映画のキャプションには、複数の映画の側面にわたって統一されたオープン形式の説明が必要です。このタスクは 2 つの主な理由から困難です。1 つは、モデルが微妙な視覚的証拠からプロの映画のコンセプトを推測する必要があること、もう 1 つは包括的かつ正確なキャプションを生成する必要があることです。したがって、私たちは、時空間アンカーを備えた構造化推論と、包括性、精度、ゲート制御されたカバレッジ報酬を備えた強化学習を組み合わせたフレームワークである CineCap を提案します。前者は、明確な視覚的証拠に基づいてプロの映画の説明を根拠にし、教師付き微調整のためにコンパクトな原子的推論に編成します。一方、後者は、説明の完全性と事実の正確さの間のバランスを改善します。さらに、体系的な評価のために手動で注釈を付けた 472 個のビデオとキャプションのペアのベンチマークである CineCap Bench を構築します。広範な実験により、CineCap が独自の強力なオープンソースベースラインを常に上回り、映画のキャプションの新しい最先端技術を確立していることが示されています。コード、モデル チェックポイント、ベンチマークは、https://github.com/Hectormxy/CineCap.git で公開されています。

原文 (English)

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning

Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is important for fine-grained video understanding and controllable movie-quality video generation, yet remains underexplored in existing multimodal large language models. Unlike question-answering-based evaluation of cinematic understanding, cinematographic captioning requires a unified open-form description over multiple cinematographic dimensions. This task is challenging for two main reasons: the model must infer professional cinematographic concepts from subtle visual evidence, and it must generate captions that are both comprehensive and accurate. Accordingly, we propose CineCap, a framework that combines structured reasoning with spatio-temporal anchors and reinforcement learning with comprehensiveness, accuracy, and gated coverage rewards. The former grounds professional cinematographic descriptions in explicit visual evidence and organizes them into compact atomic reasoning for supervised fine-tuning, while the latter improves the balance between descriptive completeness and factual correctness. In addition, we construct CineCap Bench, a benchmark of 472 manually annotated video-caption pairs for systematic evaluation. Extensive experiments show that CineCap consistently outperforms strong proprietary and open-source baselines, establishing a new state of the art for cinematographic captioning. The code, model checkpoint, and benchmark are publicly available in https://github.com/Hectormxy/CineCap.git.

13:00 JSTLLM/生成AI

LaGO: オンライン強化学習のための潜在アクション ガイダンス

大規模言語モデル (LLM) は計画と逐次的な意思決定に強力な可能性を示していますが、これまでの研究では多くの場合、LLM を直接コントローラーとして使用することに依存しており、これには正確なアクション生成が必要であり、実際には信頼できない可能性があります。この論文では、オンライン強化学習のための潜在アクション ガイダンス (LaGO) を提案します。これは、LLM を明示的なプランナーまたはコントローラーとして扱うのではなく、オンライン ポリシーの最適化をソフトにガイドする前に、事前トレーニング済み LLM を潜在アクションとして使用するフレームワークです。離散制御ベンチマークである CLEVR-Robot と連続制御ベンチマークである Meta-World の両方での実験では、LaGO がバニラ PPO よりも報酬と成功率の両方を一貫して向上させていることが実証されています。特に、LaGO は、CLEVR-Robot では平均成功率が 15.1% から 27.2% に、Meta-World では 2.7% から 15.2% に増加します。さらに、私たちの分析では、より強力な事前トレーニングされた LLM がより効果的なガイダンスを提供することが示されており、LLM の知識が計画とオンラインの意思決定を改善できることを示唆しています。

原文 (English)

LaGO: Latent Action Guidance for Online Reinforcement Learning

Large language models (LLMs) have shown strong potential for planning and sequential decision-making, but prior work often relies on using them as direct controllers, which requires precise action generation and can be unreliable in practice. This paper proposes Latent Action Guidance for Online Reinforcement Learning (LaGO), a framework that uses a pretrained LLM as a latent action prior to softly guide online policy optimization, rather than treating the LLM as an explicit planner or controller. Experiments on both a discrete-control benchmark, CLEVR-Robot, and a continuous-control benchmark, Meta-World, demonstrate that LaGO consistently improves both reward and success rate over Vanilla PPO. In particular, LaGO increases the average success rate from 15.1% to 27.2% on CLEVR-Robot and from 2.7% to 15.2% on Meta-World. Our analysis further shows that stronger pretrained LLMs provide more effective guidance, suggesting that LLM knowledge can improve planning and online decision-making.

13:00 JSTビジネス/資金調達

確率的ブール関数評価のためのコスト最適決定図

多くの意思決定シナリオでは、情報の取得にさまざまなコストがかかります。変動コストおよび真理の割り当てに対する確率分布の下で命題式を評価する際の期待コストを最小限に抑える決定論的評価戦略を構築する問題を検討します。変数選択ヒューリスティック、枝刈り、およびキャッシュを備えた分岐限定アルゴリズムを紹介します。私たちの知る限り、これはこのレベルの一般性を実現する最初の実用的な正確なアルゴリズムです。ランダム インスタンスの実験では、スケーラビリティを実証し、貪欲なビーム検索バリアントの効率と品質のトレードオフを定量化します。さらに、構造化された心臓病の診断例も評価します。最後に、問題が $\#P$ 困難であり、$\mathrm{PSPACE}$ に含まれていることを証明します。

原文 (English)

Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation

In many decision-making scenarios, acquiring information incurs different costs. We consider the problem of constructing a deterministic evaluation strategy that minimizes the expected cost of evaluating a propositional formula under variable costs and a probability distribution over truth assignments. We present a branch-and-bound algorithm with variable-selection heuristics, pruning, and caching. To the best of our knowledge, it is the first practical exact algorithm for this level of generality. Experiments on random instances demonstrate scalability and quantify the efficiency-quality trade-off of a greedy beam-search variant. We additionally evaluate a structured heart-disease diagnosis instance. Finally, we prove that the problem is $\#P$-hard and contained in $\mathrm{PSPACE}$.

13:00 JST研究/論文

BlockTrain を使用した分散型 AI トレーニングと推論

フロンティア AI トレーニングは、集中管理された高密度のアクセラレータ クラスターへのアクセスによってますます形作られています。これにより、ハイパースケーラーや大規模な集中研究所にとって構造的な利点が生まれ、オープンまたは独立した AI の取り組みが、希少な資本、特権的なインフラストラクチャ、およびデータセンターの地理に依存することになります。我々は、モデルが独立してトレーニング可能なブロックに分割され、それぞれが同じグローバル ターゲットから派生したローカル目標に基づいて最適化され、推論時に 1 つのモデルに構成される分散トレーニング プロトコルである Spheroid BlockTrain を紹介します。バイトレベルの WikiText では、同じセットアップのエンドツーエンド Transformer 参照の約 0.04 CE 以内で、BlockTrain はクロス エントロピー 1.359 (複雑度 3.89) に達しますが、アクティブな各ワーカーは 1 つのブロックのみをトレーニングし、フルモデル オプティマイザー状態を回避します。共有 6 ワーカー ブロックのトレーニング実行は、同じブロックの更新を 1 つの組み立てられたモデルに平均化することで CE 1.385 に達します。 HTTP/TCP トランスポート実験では、実際のシリアル化されたチェックポイントと更新を移動します。これには、15.22 GB を移動しながら CE を 5.580 から 1.811 に改善するパブリック IP 3 ホストの実行が含まれます。推論のために、現在の BlockTrain パスは完全な出力ごとに 1 つのブロック スタック トラバーサルを使用し、最大 75.80B パラメーターの論理 fp16 シェイプまでの 3 つのパブリック ネットワーク GPU ホストにわたる直接 TCP を介してサービスを提供します。これは、トラバーサルごとに 1 つのトークンではなく、WAN パイプライン トラバーサルごとに完全なシーケンスを出力するため、一致するプレーン自己回帰 TCP パイプライン ベースラインよりも優れたパフォーマンスを発揮します。

原文 (English)

Decentralised AI Training and Inference with BlockTrain

Frontier AI training is increasingly shaped by access to dense, centrally controlled accelerator clusters. This creates a structural advantage for hyperscalers and large centralized laboratories, and makes open or independent AI efforts depend on scarce capital, privileged infrastructure, and data-center geography. We present Spheroid BlockTrain, a decentralized training protocol in which a model is partitioned into independently trainable blocks, each optimized on a local objective derived from the same global target and composed at inference into one model. On byte-level WikiText, BlockTrain reaches cross entropy 1.359 (perplexity 3.89), within about 0.04 CE of a same-setup end-to-end Transformer reference, while each active worker trains only one block and avoids full-model optimizer state. A shared six-worker block training run reaches CE 1.385 by averaging same-block updates into one assembled model. HTTP/TCP transport experiments move real serialized checkpoints and updates, including a public-IP three-host run that improves CE from 5.580 to 1.811 while moving 15.22 GB. For inference, the current BlockTrain path uses one block-stack traversal per full output and serves over direct TCP across three public-network GPU hosts up to a 75.80B-parameter logical fp16 shape, outperforming a matched plain-autoregressive TCP pipeline baseline because it emits a full sequence per WAN pipeline traversal rather than one token per traversal.

13:00 JSTLLM/生成AI

タスク固有の LLM 蒸留のスケーリング則

大規模言語モデル (LLM) は、ますます広範囲のドメインにわたって強力なパフォーマンスを実現しますが、そのスケールにより、遅延とコストの制約が重要なアプリケーションでは導入の課題が生じます。この論文では、ドメイン固有の LLM 圧縮に関する経験的なスケーリング則を導き出し、データセットのサイズ、圧縮率、監視形式、反復枝刈りスケジュールによってドメイン内および一般知識のパフォーマンスがどのようにスケールされるかを定量化します。クオンツファイナンスをアプリケーションドメインとして使用し、反復構造枝刈りの下でロジットベースの蒸留と LoRA ベースの蒸留を比較し、推論トレース上で KL ダイバージェンスの蒸留を安定化する混合思考連鎖監視損失を導入します。圧縮下ではドメイン内タスクの品質が予想通り低下しますが、一般知識のベンチマークは同じ時点よりかなり前に崩壊します。監視形式はこのトレードオフの主な要因であり、思考連鎖監視は枝刈りによって消去された一般知識を積極的に回復します。私たちは、ヘッドライン データセット FinHeadlineMix、スケーリング則の結果、およびドメイン固有の圧縮決定のための再利用可能なフレームワークを提供する実践的な推奨事項をリリースします。

原文 (English)

Scaling Laws for Task-Specific LLM Distillation

Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in applications where latency and cost constraints are critical. This paper derives empirical scaling laws for domain-specific LLM compression, quantifying how in-domain and general knowledge performance scale with dataset size, compression ratio, supervision format, and iterative pruning schedule. Using quantitative finance as our application domain, we compare logit-based and LoRA-based distillation under iterative structural pruning, introducing a blended chain-of-thought supervision loss that stabilizes KL-divergence distillation over reasoning traces. In-domain task quality degrades predictably under compression while general-knowledge benchmarks collapse well before the same point; supervision format is the key driver of this tradeoff, with chain-of-thought supervision actively recovering general knowledge that pruning erases. We release the headline dataset FinHeadlineMix, scaling law results, and practical recommendations to provide a reusable framework for domain-specific compression decisions.

13:00 JST研究/論文GPT / ChatGPT

スケールは大規模な言語モデルにおける可塑性の損失を防ぐことができるか?

可塑性、つまり古い情報をすでに学習した後に新しい情報を学習するネットワークの能力の喪失は、継続的な学習が可能な人工ニューラル ネットワークを作成する際の基本的な課題です。この現象は何十年も前から知られていましたが、主に古い、比較的小規模なアーキテクチャで研究されており、自然言語領域で研究されることはほとんどありませんでした。現代のトランスフォーマーベースの LLM パラダイムにおいて可塑性の損失が依然として問題であるかどうかを判断するために、多言語の継続学習問題でトレーニングされた GPT スタイルのトランスフォーマー モデルにおける可塑性の損失を研究します。以前の研究と一致して、延期されたベトナムの探査タスクの劣化によって測定されたように、500万から3億1400万の非埋め込みパラメータの範囲のモデル全体で塑性損失の証拠が見つかりました。さらに、可塑性損失の開始は予測可能なスケーリング則に従っており、モデルのサイズに応じて非線形に増加することがわかりました。これらの結果は、モデルが大きくなると可塑性損失の測定可能な影響が遅れる可能性があるが、パラメータ数を増やすだけでは塑性損失を完全に防ぐには不十分である可能性が高いことを示唆しています。また、定常的な多言語トレーニング下で可塑性が失われる証拠も見つかり、この現象は突然の課題変更を伴う継続的な学習に限定されるという見解に異議を唱えています。全体として、私たちの結果は、自然言語でトレーニングされた大規模な Transformer 言語モデルであっても、継続的設定と定常的設定の両方で、十分に長いトレーニングの後、最終的には新しいデータに効率的に適応する能力を失うことを示唆しています。

原文 (English)

Can Scale Save Us From Plasticity Loss in Large Language Models?

The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning. Although this phenomenon has been known for decades, it has mostly been studied in older, relatively small architectures and rarely in natural-language domains. To determine whether loss of plasticity remains a problem in the modern transformer-based LLM paradigm, we study plasticity loss in GPT-style Transformer models trained on a multilingual continual learning problem. Consistent with prior work, we find evidence of plasticity loss across models ranging from 5M to 314M non-embedding parameters, as measured by deterioration on a held-out Vietnamese probing task. We further find that the onset of plasticity loss follows a predictable scaling law, growing sublinearly with model size. These results suggest that larger models may delay the measurable effects of plasticity loss, but that increasing parameter count alone is likely to be insufficient to completely prevent it. We also find evidence of plasticity loss under stationary multilingual training, challenging the view that the phenomenon is exclusive to continual learning with abrupt task changes. Overall, our results suggest that even large Transformer language models trained on natural-language will eventually lose the ability to efficiently adapt to new data after sufficiently long training, in both continual and stationary settings.

13:00 JST研究/論文GPT / ChatGPT

BluTrain: AI システム用の C++/CUDA フレームワーク

大規模な深層学習の進歩は、モデリングよりもシステム エンジニアリングの問題です。トレーニング中のモデルの動作 (スループット、メモリ フットプリント、結果の数値的忠実度) は、アーキテクチャ自体によって決まるというよりは、そのアーキテクチャがハードウェア上でどのように表現されるかによって決まります。システムの複雑さを抽象化してモデリングをシームレスにし、反復的なオーケストレーション ロジックの必要性を排除しながら、このハードウェア表現に対する絶対的な制御を実現するために、BluTrain は、標準 C++ およびコア CUDA プログラミング モデルにおける堅牢で軽量なアーキテクチャ全般のトレーニング フレームワークとして第一原理から設計されました。リバースモード autograd を備えた型付きテンソル モジュール、線形代数ライブラリ、キャッシュ アロケータ、マルチモード分散実行モジュール、MLIR ベースの深層学習コンパイラなど、すべての層がネイティブに実装されています。 8 GPU 6000 Ada システム上の FP32 で 124M パラメータの GPT-2 ベースラインをトレーニングする正式な評価では、BluTrain は、スループット (平均 407K トークン/秒を維持するのに対し、PyTorch の 395K トークン/秒を維持) とメモリ効率 (最大 22% のフットプリント削減を達成) の両方で業界標準のベースラインを上回っています。数値的忠実度が向上し、最終的な検証損失がわずかに低下するように収束します。すべてのレイヤーがネイティブ チューニングに対して明示的にオープンであるため、パフォーマンスの上限はフレームワーク自身で引き上げることができます。

原文 (English)

BluTrain: A C++/CUDA Framework for AI Systems

Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware. To achieve absolute control over this hardware expression while abstracting away systems complexity to make modelling seamless and eliminating the need for repetitive orchestration logic, BluTrain was architected from first principles as a robust, lightweight, and architecture-general training framework in standard C++ and the core CUDA programming model. Every layer is implemented natively: a typed tensor module with reverse-mode autograd, a linear-algebra library, a caching allocator, a multi-mode distributed-execution module, and an MLIR-based deep-learning compiler. In formal evaluations training a 124M-parameter GPT-2 baseline in FP32 on an 8-GPU 6000 Ada system, BluTrain outperforms industry-standard baselines in both throughput (sustaining an average of 407K tokens/s versus PyTorch's 395K tokens/s) and memory efficiency (achieving up to a 22% footprint reduction), while strictly preserving numerical fidelity and converging to a marginally lower final validation loss. With every layer explicitly open to native tuning, the performance ceiling is the framework's own to raise.

13:00 JST研究/論文

領域一般化のための人間の活動認識の分布シフトの評価

人間活動認識 (HAR) の分野は引き続き研究者の関心を集めており、重要な面で進歩していますが、いくつかの重要な課題が残っています。現実世界の設定で優れたパフォーマンスを示す HAR モデルを構築する際の最も困難な側面の 1 つは、デバイスとセンサーの異質性によるデータの多様性と、現実世界のアプリケーションに固有のコンテキストの変化に対処することです。 HAR におけるデータの多様性は文献でよく知られていますが、HAR モデルに対するさまざまなタイプの分布シフトの影響と、そこから生じるドメイン一般化問題についての理解にはまだギャップが残っています。その目的に向けて、このペーパーでは、デバイスの種類、センサーの配置、サンプリング レート、ユーザーの行動の変化を含む 4 つの異なるタイプの分布シフトを体系的に評価します。その影響を定量化することで、多様性の変化が主にあらゆる種類の変化を定義することを示し、異なるドメイン間で共有されない固有の特徴が存在することを示しています。次に、均一な HAR ベースの分散シフト ベンチマークを導入し、最大 28 のドメイン一般化手法の包括的な評価を実行します。私たちの分析は、経験的なリスク最小化ベースラインをわずかに上回るパフォーマンスで、モデルの一般化可能性を達成する際の現在のドメイン一般化アルゴリズムの限界を明らかにしました。この研究は、センサーベースの HAR における特定の分布シフトに関するドメインの一般化と適応の最初の体系的な調査を表しており、さらなる研究を促進するためのオープンソースのベンチマーク プラットフォームとデータセットを提供します。

原文 (English)

Assessing Distribution Shift in Human Activity Recognition for Domain Generalization

While the field of Human Activity Recognition (HAR) continues to draw interest from researchers and advance in important ways, some key challenges remain. One of the most difficult aspects of building HAR models that show good performance in real-world settings is dealing with data diversity from device and sensor heterogeneity, and contextual changes that are intrinsic to real-world applications. While data diversity in HAR has been well-acknowledged in the literature, there remains a gap in understanding the effect of various types of distribution shifts on HAR models and the domain generalization problem that arises. Towards that end, this paper systematically evaluates 4 different types of distribution shifts, including variations in device type, sensor placement, sampling rate, and user behavior. Quantifying their effects, we illustrate that diversity shifts predominantly define all types of shifts, indicating the existence of unique features that are not shared across different domains. We then introduce a uniform HAR-based distribution shift benchmarks and conduct a comprehensive evaluation of up to 28 domain generalization methods. Our analysis exposes the limitations of current domain generalization algorithms in achieving model generalizability, marginally outperforming the empirical risk minimization baseline. This work represents the first systematic exploration of domain generalization and adaptation concerning specific distribution shifts in sensor-based HAR, offering an open-source benchmark platform and datasets to spur further research.

13:00 JST研究/論文

双方向の条件付きフローマッチングによるカオスシステムの逆問題の解決

カオス システムのモデル化は重要ですが、困難です。カオス力学における逆問題、つまり最終状態から初期条件を推測する問題は、姿勢不良、非一意性、不安定性、および潜在的にカオスな時間反転力学のため、ほとんど未解決のままです。私たちは、双方向条件付きフロー マッチング (Bi-CFM) を使用してこの未解決の問題に対処します。Bi-CFM は、初期状態と最終状態の分布間の双方向マッピングを学習して、カオス進化の確率性を捉え、時間の経過に伴う指数関数的な誤差の蓄積を軽減します。さらに、保存則のある系については、それを保存制約付き Bi-CFM (CBi-CFM) に拡張します。 Bi-CFM は、古典的な Lorenz、Circuit、および高次元 Lorenz 96 システム全体で、ベースラインを上回る 5 つの分散レベルのメトリクスを改善し、2 桁を超える高速化を達成します。惑星力学における三体惑星間散乱問題では、CBi-CFM は保存則をよりよく尊重しており、保存誤差はグランド トゥルースの保存誤差と同等です。最後に、$\sim 10^{10}$ 年 (10 Gyr) の進化によって形作られた衝突百万体系である球状星団の実際の観測では、私たちの方法は精度の向上を示し、長期スケールの実世界のカオス力学の逆問題を解決するためのスケーラブルなルートを確立しました。

原文 (English)

Solving Inverse Problems of Chaotic Systems with Bidirectional Conditional Flow Matching

Modeling chaotic systems is crucial yet challenging. Inverse problems in chaotic dynamics, namely inferring initial conditions from final states, remain largely unsolved because of ill-posedness, non-uniqueness, instability, and potentially chaotic time-reverse dynamics. We address this open problem with Bidirectional Conditional Flow Matching (Bi-CFM), which learns bidirectional mappings between distributions of initial and final states to capture the stochasticity of chaotic evolution and mitigate exponential error accumulation over time. Furthermore, for systems with conservation laws, we extend it to Conservation-constrained Bi-CFM (CBi-CFM). Across the classic Lorenz, Circuit, and high-dimensional Lorenz 96 systems, Bi-CFM improves five distribution-level metrics over baselines while achieving a speedup of more than two orders of magnitude. In the three-body planet-planet scattering problem in planetary dynamics, CBi-CFM better respects conservation laws, with conservation errors comparable to those of the ground truth. Finally, on real observations of globular clusters, collisional million-body systems shaped by $\sim 10^{10}$ years (10 Gyr) of evolution, our method represents an advance in accuracy, establishing a scalable route to solving inverse problems of long-timescale real-world chaotic dynamics.

13:00 JST研究/論文

違いを生まずに違いを生み出す

一連の 7 つの論文にわたって、アンドレアスと G は実際の因果関係の 7 つの定義を導入し、それらを 3 つの異なる競合する説明タイプに属するものとして分類しました。それは、事実による差異形成、反事実による差異形成、および規則性に基づくものです。私は、彼らの最新の事実による差異形成の定義が 3 つのタイプすべてを具体化していることを示し、それによってこれらが違いのない区別であることを証明しました。さらに、いくつかの重要な例について、彼らの新しい説明を他の 6 つの説明と比較し、これが次のことを明らかにしています。彼らの7つのアカウントすべてを台無しにします。

原文 (English)

Difference-Making without Making a Difference

Over a series of seven papers, Andreas & G\"unther have introduced seven definitions of actual causation and have classified them as belonging to three different, competing, types of accounts: factual difference-making, counterfactual difference-making, and regularity-based. I show that their most recent - factual difference-making - definition instantiates all three types, thereby proving that these are distinctions without a difference. I further compare their novel account to the other six accounts on several crucial examples, revealing that this undermines all seven of their accounts.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Copilot

NFR 評価のためのマルチターン LLM ダイアログの精度と満足度

LLM ベースの対話アシスタントはソフトウェア開発者にとって主流のツールとなっていますが、現在の評価ベンチマークは機能の正しさのみに焦点を当てています。このため、本質的に曖昧でコンテキストに依存し、プログラムの多くの部分に関係する非機能要件 (NFR) を処理する際に、これらの会話の品質と正確性を評価する際に重大なギャップが残ります。これらのシステムが NFR に関する協調推論をどの程度サポートしているかを評価するには、シングル ターンの精度を超えて、システムの出力の正確さとマルチ ターン インタラクションの品質の両方を取得する方法が必要です。このペーパーでは、医療保険の相互運用性と責任に関する法律 (HIPAA) 規制順守の領域における、開発者と LLM ベースのエージェントとの間の複数回にわたる会話の精度と品質を調査します。私たちは 49 人のプログラマーを雇い、GitHub Copilot と対話し、HIPAA 規制に準拠するように設計されたシステムである iTrust コードベースに対して 148 個の HIPAA 由来の NFR を、要件満足度、推論、コード ローカリゼーションの 3 つの側面にわたって評価しました。開発者は LLM 評価に同意する傾向がありますが、専門家のグラウンド トゥルースに対する精度は低いことがわかりました。ユーザー満足度をモデル化したところ、システムの応答時間が長くなり、情報提供ターンが増えるとユーザー満足度にマイナスの影響が出るのに対し、積極的なインタラクションはプラスの影響を与えることがわかりました。私たちの調査結果は、NFR 評価をサポートする LLM ベースの対話システムを設計するための洞察を提供します。

原文 (English)

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a critical gap in assessing the quality and accuracy of these conversations when handling Non-Functional Requirements (NFRs), which are inherently vague, context-dependent, and involve many parts of a program. Evaluating how well these systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of the system's outputs and the quality of the multi-turn interaction. In this paper, we investigate the accuracy and quality of multi-turn conversations between developers and an LLM-based agent in the domain of Health Insurance Portability and Accountability Act (HIPAA) regulatory compliance. We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase, a system designed to comply with HIPAA regulations, across three dimensions: requirement satisfaction level, reasoning, and code localization. We find that developers tend to agree with LLM assessments, but accuracy against expert ground truth is low. We model user satisfaction and find that longer system responses and more information-providing turns negatively affect user satisfaction, whereas proactive interactions positively affect it. Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.

13:00 JSTエージェントハードウェア/半導体

採点者の採点: エージェントによるデータ分析システムの評価から得た教訓

エージェント的データ分析システムは、コード、数値結果、口頭診断などの豊富な出力を生成します。このため、シングルターン LLM 応答よりも評価が難しくなります。したがって、エージェントの出力と、採点アーティファクトからの真実の回答との間の真の不一致を区別する必要があります。私たちは、DSGym の 153 の数値 QRData タスクにマルチエージェント データ分析システムである LAMBDA を適用することで、自動採点者がそのようなシステムをどのように確実に評価するか、またどのような戦略が採点の品質を向上させるかを調査します。私たちは、厳格な正規表現マッチング、LLM ベースの寛大なグレーディング、スニペットベースの人的検査という 3 層の人的 AI グレーディング カスケードを開発および評価します。これは、非 GenAI 戦略と GenAI 戦略をさまざまな障害プロファイルと組み合わせたものです。どちらの自動グレーダーも 100% の観察精度 (誤検知 0/70) を達成しています。寛大な採点者の再現率は人間のラベルに対して 97% です。キーワードに固定された抽出パイプラインにより、厳密な採点者の再現率は、最後の数字のヒューリスティックよりも 60 パーセント ポイント高くなります。寛大なグレーダーはアーキテクチャ的にパーサーに依存しません。反復的なナッジメカニズムにより、採点の成功率が 36% から 97% に上昇し、寛容な合格率が 16% から 46% に上昇します。元の質問の再挿入ありとなしのナッジを比較すると、再挿入にはメリットがないことがわかり、ナッジが回答テンプレートの手がかりであることが確認されました。さらに、このケース スタディでは、変数タイプがパイプラインのダイナミクスの評価と観察された結果の評価に最も一貫して関連付けられているタスク メタデータ フィールドであることがわかります。

原文 (English)

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses. It is therefore necessary to distinguish genuine disagreement between an agent's output and a ground-truth answer from grading artifacts. We investigate how reliably automated graders assess such a system and what strategies improve grading quality by applying LAMBDA, a multi-agent data-analysis system, on 153 numerical QRData tasks from DSGym. We develop and evaluate a three-layer human-AI grading cascade: strict regex matching, LLM-based lenient grading, and snippet-based human inspection, which combines non-GenAI and GenAI strategies with different failure profiles. Both automated graders achieve 100% observed precision (0/70 false positives). The lenient grader's recall is 97% against human labels. A keyword-anchored extraction pipeline raises the strict grader's recall by 60 percentage points over a last-number heuristic; the lenient grader is architecturally parser-independent. An iterative nudge mechanism raises grading run success from 36% to 97% and lenient-pass rates from 16% to 46%; comparing nudging with and without original-question re-injection shows that re-injection offers no benefit, confirming the nudge as an answer template cue. We further observe in this case study that variable type is the task metadata field most consistently associated with grading pipeline dynamics and observed outcome grades.

13:00 JSTLLM/生成AI

タスクと目標のマッチング: エンコーダ/デコーダの事前トレーニング済み言語モデルの微調整および即時調整戦略

プロンプトベースの学習は、自然言語処理における主要なパラダイムとして浮上しています。この研究では、常識的な知識の検索と完了に焦点を当て、生成タスクと質問応答タスクにわたるエンコーダ/デコーダの事前トレーニング済み言語モデルのパフォーマンスに対する、さまざまな事前トレーニング目標の影響を調査します。トレーニング前と微調整の両方の段階で複数の目標を組み込む利点を強調します。特定のタスクに適切な目標を決定するための Match Task to Objective (MTO) フレームワークとメソッドを紹介します。このフレームワークは、特定された目的に基づいて、教師なしトレーニングによる適応のためにタスク関連データを準備する自動化された方法を提供します。微調整段階では、事前トレーニング段階と適応段階の目的に沿った新しいテンプレートを設計します。タスクの要件に合わせてこれらの戦略を使用すると、数ショット設定で従来の方法と比較して 120\% 以上のパフォーマンス向上を達成できます。これらは、数ショット設定で関連する作業を大幅に上回り、フルデータセットのシナリオでもベースラインを上回ります。さらに、このアプローチを拡張してプロンプト チューニング方法論を含め、より効果的なソフト プロンプト エンジニアリングと最適化のためのガイダンスを提供します。私たちの戦略は、プロンプトチューニングのパフォーマンスも大幅に向上させます。これらの洞察には大きな価値があり、特定のタスク用にカスタマイズされたモデルの選択と最適化を正確にガイドします。コードは https://github.com/puraminy/MTO/ で入手できます。

原文 (English)

Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models

Prompt-based learning has emerged as a dominant paradigm in natural language processing. This study explores the impact of diverse pre-training objectives on the performance of encoder-decoder pre-trained language models across generation and question answering tasks, with a focus on commonsense knowledge retrieval and completion. We highlight the benefits of incorporating multiple objectives during both pre-training and fine-tuning stages. We introduce the Match Task to Objective (MTO) framework and methods for determining the appropriate objective for a given task. This framework offers automated methods to prepare task-related data for adaptation through unsupervised training, based on the identified objective. In the fine-tuning stage, we design novel templates that align with the objectives of the pre-training and adaptation stages. When aligned with task requirements, these strategies can achieve a performance gain of over 120\% compared to conventional methods in few-shot settings. They significantly outperform related works in few-shot settings and exceed the baseline even in full-dataset scenarios. Furthermore, we extend this approach to include prompt-tuning methodologies, providing guidance for more effective soft prompt engineering and optimization. Our strategies significantly enhance prompt-tuning performance as well. These insights hold substantial value, precisely guiding the selection and optimization of models customized for specific tasks. Code is available at https://github.com/puraminy/MTO/

13:00 JSTエージェント

バラバラの世界モデル: 一般代理店向けの構造認証

大きな世界体制では、エージェントは普遍的な能力を持つことはできず、エージェントの能力は必然的に世界モデル全体に​​わたって細分化されて特殊化されます。その結果、標準的な統一保証では、重大なボトルネックと無関係な障害の理解が区別できません。まず、一般的なエージェントが普遍的ではなく、標準的な最悪の場合の分析が役に立たないことを証明することで、この制限を形式化します。これを克服するために、私たちは、限定された目標条件付きパフォーマンスをエージェントの内部世界モデルのエントリごとの保証にマッピングする移行ローカル フレームワークである構造認証を導入します。私たちの主な貢献は建設的なものです。私たちは、深い構成目標を使用して特定の遷移をフィルタリングするアルゴリズムを提供し、これらの目標に関する一般的なエージェントが $\mathcal{O}(1/n) + \mathcal{O}(\delta)$ 誤差限界を持つ構造世界モデルを持っていることを証明します。逆に、この限界は、Small-$\delta$ 体制では厳しく、その存在は当社の認定によって明示的に保証されています。これらの結果により、長期的な計画が信頼できる特定の移行を局所化することで、一般的なエージェントの認証可能な展開が可能になります。

原文 (English)

World Models in Pieces: Structural Certification for General Agents

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrelevant failures. We first formalize this limitation by proving that general agents are not universal, rendering standard worst-case analysis uninformative. To overcome this, we introduce structural certification, a transition-local framework that maps bounded goal-conditioned performance to entry-wise guarantees on the agent's internal world model. Our main contribution is constructive. We provide algorithms that filter specific transitions using deep compositional goals and prove that a general agent on these goals has a structural world model with a $\mathcal{O}(1/n) + \mathcal{O}(\delta)$ error bound. Conversely, this bound is tight in the small-$\delta$ regime, whose existence is explicitly guaranteed by our certification. These results enable the certifiable deployment of general agents by localizing the specific transitions where long-horizon planning is reliable.

13:00 JSTエージェント

OpenThoughts-Agent: エージェントティック モデルのデータ レシピ

エージェント言語モデルは AI のアプリケーションを劇的に拡張しますが、幅広い能力を持つエージェントのトレーニング データを収集する方法についてはほとんど公に知られていません。 SWE-Smith、SERA、Nemotron-terminal などの既存のオープンな取り組みは通常、単一のベンチマークをターゲットにしており、多様なエージェント タスクにわたって一般化するモデルをトレーニングする方法という問題が残されています。 OpenThoughts-Agent (OT-Agent) プロジェクトは、エージェント モデルをトレーニングするための完全にオープンなデータ キュレーション パイプラインでこのギャップに対処します。当社では、パイプラインの各段階を体系的に調査するために 100 件を超える制御アブレーション実験を実施し、タスクソースと多様性の重要性についての洞察をもたらします。次に、パイプラインから 100,000 個のサンプルのトレーニング セットを組み立て、このデータセットで Qwen3-32B を微調整します。これにより、7 つのエージェント ベンチマーク全体で平均精度 44.8% が得られ、既存の最も強力なオープン データ エージェント モデル (Nemotron-terminal-32B、40.9%) と比較して 3.9 パーセント ポイントの改善が得られました。さらに、当社のトレーニング データは強力なスケーリング特性を示し、コンピューティング制御による比較において、あらゆるトレーニング セット サイズで代替のオープン データセットを上回るパフォーマンスを示します。私たちは、エージェント モデル トレーニングに関する将来のオープン研究をサポートするために、トレーニング セット、データ パイプライン、実験データ、およびモデルを openthoughts.ai で一般公開しています。

原文 (English)

OpenThoughts-Agent: Data Recipes for Agentic Models

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation pipeline for training agentic models. We conduct more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline, yielding insights on the importance of task sources and diversity. We then assemble a training set of 100K examples from our pipeline and fine-tune Qwen3-32B on this dataset, which yields an average accuracy of 44.8% across seven agentic benchmarks and a 3.9 percentage point improvement over the strongest existing open data agentic model (Nemotron-Terminal-32B, 40.9%). Moreover, our training data exhibits strong scaling properties, outperforming alternative open datasets at every training set size in compute-controlled comparisons. We publicly release our training sets, data pipeline, experimental data, and models at openthoughts.ai to support future open research on agentic model training.

13:00 JST研究/論文

有限グラフ上の遅延結合反応拡散システムとしてのリエントラント値フィールド

二部構成のヒルベルト・シュミットカーネルを介して記号場が幾何学場に結合される動的システムについて説明します。このシステムは、リプシッツ条件と小さなゲイン条件に従う履歴空間上の遅延関数微分方程式 (RFDE) によって完全に記述されます。 RFDE が一定の入力の下で適切に配置され、コンパクトなグローバル アトラクターを許容することを示します。 2 つの主フィールドと実行フィールドで構成される主サブシステム $(H_L, X_R, P)$ は、フィールド間結合が $C_{\mathcal{K}}^2<\mu_L\mu_R$ を満たす限り、遅延に関係なく全体的に安定していることが示されています。さらに、主定理の仮説を満たす設計仕様について説明します。

原文 (English)

Reentrant value fields as delayed coupled reaction-diffusion systems on finite graphs

We describe a dynamical system in which a symbolic field is coupled to a geometric field via a bipartite Hilbert-Schmidt kernel. The system is fully described by a retarded functional differential equation (RFDE) on the history space, subject to Lipschitz and small gain conditions. We show that the RFDE is well-posed under constant input and that it admits a compact global attractor. The principal subsystem $(H_L, X_R, P)$, which is comprised of the two primary fields as well as an executive field, is shown to be globally stable independent of delay, provided that the interfield coupling satisfies $C_{\mathcal{K}}^2<\mu_L\mu_R$. In addition, we describe design specifications that fulfill the hypotheses of the main Theorem.

13:00 JST研究/論文

自己回帰の地平線を超えて: コードの拡散モデル、世界モデリング、および状態空間モデルの包括的な調査

自己回帰 (AR) 言語モデルは、自動ソフトウェア エンジニアリングの大幅な進歩を促進し、強力なコード生成および支援システムを可能にします。ただし、ネクスト トークン予測パラダイムでは、グローバル プランニングの制限、長期的な依存関係の維持における課題、プログラム実行セマンティクスの基盤の制限など、コード推論に構造的な制限が生じます。 AR モデルに対する既存の文献の大きな偏りに注目し、コード インテリジェンスのための次世代アーキテクチャ機能を解放することで、次のトークン予測のロジックとスケーリングのボトルネックを潜在的に克服できる可能性のある新たなパラダイムについて議論します。具体的には、拡散モデルの可能性について説明します。拡散モデルは、AR モデルでは見逃されがちな長距離の構文制約を捕捉する全体的なノイズ除去を介してコードを生成します。また、推論をサポートするために実行状態をシミュレートするコード ワールド モデル (CWM) と、大規模なコンテキストに対して線形時間効率を提供する状態空間モデル (SSM) についても説明します。これらの開発を認知神経科学の発見と結び付けることで、「システム 2」コード生成エージェントの開発の方向性を概説します。

原文 (English)

Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code

Autoregressive (AR) language models have driven significant progress in automated software engineering, enabling powerful code generation and assistance systems. However, the next-token prediction paradigm introduces structural limitations for code reasoning, including restricted global planning, challenges in maintaining long-range dependencies, and limited grounding in program execution semantics. Noting the heavy skewness of existing literature towards AR models, we discuss emerging paradigms that could potentially overcome the logic and scaling bottlenecks of next-token prediction by unlocking next-generation architectural capabilities for code intelligence. Specifically, we discuss the potential of Diffusion Models, which generate code via holistic denoising that captures long-range syntactic constraints often missed by AR models. We also discuss Code World Models (CWMs), which simulate execution states to support reasoning, and State Space Models (SSMs), which provide linear-time efficiency for massive contexts. By connecting these developments with findings from cognitive neuroscience, we outline directions for developing "System 2" code generation agents.

13:00 JSTLLM/生成AIビジネス/資金調達

RAG システムにおける以前の優位性の定量化

検索拡張生成(RAG)は大規模言語モデルを外部の知識に基づいて構築しますが、現在の評価は「認識論的盲目さ」に悩まされる離散ヒューリスティックに依存しており、真の文脈情報抽出とパラメトリック記憶想起を区別できません。これに対処するために、NCU (Normalized Context Utilization) メトリクスを導入し、ゼロショット、オラクル、および敵対的条件にわたる連続トークンのログ確率を活用して、コンテキスト情報の獲得を厳密に定量化します。独自の商用 API と並行して 1.5B から 72B のパラメータの範囲のアーキテクチャを評価すると、厳密な事実抽出 (思考連鎖推論なし) の場合、従来のスケーリング則では極端な利益逓減が見られることが明らかになりました。つまり、高効率の小型言語モデル (SLM) は、大容量アーキテクチャに匹敵するか、それを上回っています。さらに、「事前優位性」がモデルのスケールや独自のアラインメントと相関していることを示します。評価された商用 API は、敵対的紛争のほぼ半数で明示的な外部証拠を無効にしただけでなく、パラメトリック事前条件が矛盾した場合にシステムの信頼崩壊 (ネガティブトランスファー) に頻繁に悩まされました。私たちの調査結果は、厳密な抽出ワークフローにおける SLM の構造的認識論的な利点と優れた文脈順守を強調しています。

原文 (English)

Quantifying Prior Dominance in RAG Systems

Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suffer from ''epistemic blindness'' - failing to distinguish genuine contextual information extraction from parametric memory recall. To address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adversarial conditions to strictly quantify contextual information gain. Evaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of-Thought reasoning), traditional scaling laws exhibit extreme diminishing returns: highly efficient Small Language Models (SLMs) match or outperform high-capacity architectures. Furthermore, we demonstrate that ``Prior Dominance'' correlates with model scale and proprietary alignments. The evaluated commercial API not only overrode explicit external evidence in nearly half of adversarial conflicts, but also frequently suffered from systemic confidence collapse (Negative Transfer) when its parametric priors were contradicted. Our findings highlight the structural epistemic advantage and superior contextual adherence of SLMs in strict extraction workflows.

13:00 JST研究/論文

SemChunk-C: C コードのセマンティック セグメンテーション

C ファミリ言語で書かれたコードのセマンティック セグメンテーションは、言語の複雑な構文、マクロ拡張、不規則な構造パターンのため、依然として困難な問題です。固定サイズのウィンドウ、ヒューリスティック分割、構文ベースのツールなどの既存のチャンク手法では、意味のある機能単位をキャプチャできないことが多く、検索やその他の下流の LLM 主導タスクの有効性が制限されます。このペーパーでは、C 関連言語におけるチャンクの問題について取り上げます。まず、コード チャンク カテゴリのセットを定義します。次に、LLM ベースの分類器をトレーニングして、a) チャンクの境界を識別し、b) 各チャンクに説明的な機能属性 (カテゴリ) を割り当てます。これは下流のタスクに役立ちます。コード内のセマンティック コンテキストをキャプチャする LLM の機能を活用することで、柔軟なチャンク境界を想定し、各インスタンスの特定の構造とコンテキストに適応できるようにします。 3 番目に、C 関連ファイル (.c、.cpp、.h、.cs など) のセマンティック チャンキング用の軽量言語モデル ファミリである SemChunk-C を紹介します。これらのモデルは、17M、32M、68M、および 150M パラメーターを備えた最初の 4 つの Ettin エンコーダー [1] に基づいています。サイズが比較的小さいにもかかわらず、データ構造、インターフェイス ブロック、その他のコンポーネントなど、まとまったコード単位を識別することができます。さらに、ネストされた定義やマクロなどの困難な構造を含む、実際のコードに対するアプローチの堅牢性を実証します。私たちはさまざまなデータセットでアプローチをテストし、高い境界精度とセマンティック一貫性を達成し、はるかに大規模なコード指向の LLM に基づくチャンカーと一致またはそれを上回るパフォーマンスを示します。また、いくつかの厳選されたベンチマークでダウンストリーム タスクのパフォーマンスの向上も検証します。

原文 (English)

SemChunk-C: Semantic Segmentation for C Code

Semantic segmentation of code written in a C-family language remains a challenging problem, due to the language's complex syntax, macro expansion, and irregular structural patterns. Existing chunking methods, such as fixed-sized windows, heuristic splitting, and syntax-based tools, often fail to capture meaningful functional units, limiting the efficacy of retrieval and other downstream LLM driven tasks. In this paper, we address the problem of chunking in C-related languages. First, we define a set of code chunk categories. Second, we train an LLM-based classifier to a) identify chunk boundaries, and b) assign each chunk a descriptive functional attribute (a category), which can be useful for downstream tasks. By leveraging the LLM's ability to capture semantic context within the code, we assume flexible chunk boundaries, allowing to adapt to the specific structure and context of each instance. Third, we introduce SemChunk-C, a family of lightweight language models for semantic chunking of C-related files (.c, .cpp, .h, .cs, etc.). These models are based on the first four Ettin encoders [1] with 17M, 32M, 68M, and 150M parameters. Despite their relatively small size, they are capable of identifying cohesive code units, such as data structures, interface blocks, and other components. Furthermore, we demonstrate the robustness of our approach on real-world code, including challenging constructs such as nested definitions and macros. We test our approach on various datasets, and show that it achieves high boundary accuracy and semantic coherence, matching or outperforming chunkers that are based on much larger code-oriented LLMs. We also validate the improved performance of the downstream tasks on a few curated benchmarks.

13:00 JSTハードウェア/半導体ClaudeNVIDIA

必要なのは FP8 だけです (パート 2): Tensor-core Garner 再定式化と Kulisch エスケープ ルートによる効率的な Ozaki-Bailey スタイル FFT

NVIDIA の Blackwell Ultra (B300) は、FP64 ベクトル スループットを GPU あたり約 1.3 TFLOPS に削減します。これは、B200 の約 30 分の 1 であり、帯域幅が制限された FP64 ワークロードがメモリ制限にとどまるレベルをはるかに下回ります。 Ozaki Scheme II フレームワークは、仮数スライスされた中国剰余再構成を使用して FP8 テンソル コアを介して密行列乗算をルーティングすることにより、FP64 と同等のスループットを回復します。関連資料のパート (1) では、高密度 GEMM、バッチ GEMV、ステンシル、および SpMV について説明します。この論文では、5 番目の標準プリミティブである 3-D FFT を追加します。 FP8 テンソル コア上の両方の 1-D FFT GEMM を使用した Bailey 6 ステップ分解を介してエミュレートされた 3-D FFT である Ozaki-Bailey FFT を紹介します。 Bailey の小さな内因数 k ~ sqrt(N) (N=1024 の場合 k=32) は、カーネルを k << r^2 の状態に置き、3 番目の TME パラメーター ガンマ (再構成待ち時間) が償却ではなく結合します。 Garner の再構成は、フェーズ A (FP8/INT8 テンソル コアの内積、B300 の 1024^3 で約 1 ミリ秒) とフェーズ B (出力ごとの削減) に分割されます。 Kulisch 固定小数点完全演算は、完全に INT32 SIMT パイプ上で実行しながら完全な FP64 精度を維持するフェーズ B 再定式化であると認識します。閉じた形式の帯域幅パリティの下限を導出します。ネイティブ FP64 の下限は 1.56*B_HBM (8 TB/s で 12.5 TF) です。B300 の 1.3 TF は約 10 倍低く、Rubin の 33 TF は 4% 以内です。 Kulisch 避難ルートには、INT32 サブフロア 8.25*B_HBM と FP8 フロア 170*B_HBM が必要です。 B300 はその両方を満たします。フル FP64 での 1024^3 の予測は約 18 ミリ秒で、実質的には 12.9 ミリ秒のメモリ ルーフです。 GPU がネイティブ フロアまたは両方の Kulisch フロアを満たす場合、GPU はメモリルーフ FFT パリティを満たします。この予測が実際に当てはまれば、B300 はソフトウェアのみでフル FP64 FFT を実行できるようになり、libKulisch ライブラリとベンチマーク キャンペーンの動機付けとなります。

原文 (English)

FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route

NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at which bandwidth-limited FP64 workloads stay memory-bound. The Ozaki Scheme II framework recovers FP64-equivalent throughput by routing dense matrix multiply through FP8 tensor cores with a mantissa-sliced Chinese-remainder reconstruction. A companion Part (1) paper covers dense GEMM, batched GEMV, stencils, and SpMV; this paper adds the fifth canonical primitive, the 3-D FFT. We present Ozaki-Bailey FFT, an emulated 3-D FFT via the Bailey six-step decomposition with both 1-D FFT GEMMs on FP8 tensor cores. Bailey's small inner factor k ~ sqrt(N) (k=32 for N=1024) puts the kernel in the regime k << r^2, where the third TME parameter gamma (reconstruction latency) binds rather than amortising. Garner reconstruction splits into Phase A (inner products on FP8/INT8 tensor cores, ~1 ms for 1024^3 on B300) and Phase B (per-output reduction). We identify Kulisch fixed-point complete arithmetic as a Phase B reformulation that keeps full FP64 accuracy while running entirely on the INT32 SIMT pipe. We derive closed-form bandwidth-parity floors. The native FP64 floor is 1.56*B_HBM (12.5 TF at 8 TB/s): B300's 1.3 TF sits ~10x below, Rubin's 33 TF within 4%. The Kulisch escape route needs an INT32 sub-floor 8.25*B_HBM and an FP8 floor 170*B_HBM; B300 meets both. The projection is ~18 ms for 1024^3 at full FP64, essentially the 12.9 ms memory roof. A GPU meets memory-roof FFT parity if it satisfies either the native floor or both Kulisch floors. If the projection holds in practice, B300 becomes viable for full-FP64 FFT through software alone, motivating a libKulisch library and benchmark campaign.

13:00 JSTLLM/生成AIGPT / ChatGPT

自己認識の微調整により、突発的な位置ずれを防止し、逆転させることができます

緊急不整合 (EM) は、不整合なペルソナ ベクトルと邪悪なキャラクター特性の活性化に関連しており、EM は有害なコンテンツを直接学習するのではなく、モデルの整合したキャラクターを破壊することによって機能することを示唆しています。このつながりを動機として、私たちは既存のトレーニング中の防御とは異なる、文字をターゲットとした介入として自己生成テキスト認識 (SGTR) の微調整を研究しています。私たちは、3 つのモデル (GPT-4.1、Qwen2.5-32B-Instruct、Seed-OSS-36B-Instruct) と複数の EM データセットにわたって 2 段階の微調整実験を実施し、SGTR 微調整を無害な微調整ベースライン (正しいドメイン固有のデータ、一般知識、単語カウント) と比較し、逆転と防止の両方の設定で効果的な防御であることを確認しました。すべての介入が同等の EM の逆転をもたらすことがわかりましたが、それは EM によって低下した能力を回復する場合に限られます。予防に関しては、SGTR 微調整のみが個々の指標を悪化させることなく一貫して位置ずれを低減しており、性格の強化が特に予防を促進していることを示唆しています。我々は、EM 微調整が LLM のアイデンティティ自己報告に多様性を誘導し、自己認識を人為的に損なうことで EM 微調整によって引き起こされる不整合を悪化させ、モデルのアイデンティティを保持するシステム プロンプトを削除すると EM 微調整の効果が大幅に減少することを示すことにより、EM と LLM のデフォルト特性との関係についてのさらなる証拠を提供します。これらの発見を総合すると、EM は一貫して不整合な人格の採用としてではなく、整合性のある性格の不安定化として再構成されます。

原文 (English)

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content. Motivated by this connection, we study self-generated text recognition (SGTR) finetuning as a character-targeted intervention that is distinct from existing in-training defenses. We conduct two-stage finetuning experiments across three models (GPT-4.1, Qwen2.5-32B-Instruct, Seed-OSS-36B-Instruct) and multiple EM datasets to compare SGTR finetuning against benign finetuning baselines (correct domain-specific data, general knowledge, and word counting) to find it an effective defense in both reversal and prevention settings. We find that all interventions produce comparable EM reversal, but only when restoring capabilities that EM had degraded. For prevention, only SGTR finetuning consistently reduces misalignment without exacerbating any individual metric, suggesting that character fortification specifically drives prevention. We provide further evidence for EM's relation to the LLM's default character by showing that EM finetuning induces diversity into the LLM's identity self-reports, artificially corrupting self-recognition exacerbates misalignment caused by EM finetuning, and that removing the model's identity-bearing system prompt substantially reduces the effect of EM finetuning. Together, these findings reframe EM not as the adoption of a coherent misaligned persona but as the destabilization of aligned character.

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

製品の望ましさの効率的かつ説明可能な数値分析と分類された暗黙的センチメント分析のための LLM の使用法を評価する

製品に関する定性的なフィードバックは微妙なユーザー エクスペリエンスを明らかにする可能性がありますが、その暗黙の感情を測定するのは困難です。このペーパーでは、大規模言語モデル (LLM) を使用して、そのようなデータから製品の望ましさを定量化する、スケーラブルで解釈可能なフレームワークを紹介します。 ZORQ と CARMA の 2 つの Product Desirability Toolkit (PDT) データセットを使用し、ゴールドスタンダードの人による注釈を備えた 106 の回答者用語グループで構成され、明示的なレビュー スコアに依存せずに、ゼロショットの連続数値センチメント スコアリングとカテゴリカルセンチメント分類が評価されます。データセット全体にわたって、LLM は定性的回答と厳密に一致する専門家ラベルから数値感情スコアを直接生成し、最大 0.97 のピアソン相関と最大 94% の分類精度を達成しました。 LLM は、複数の形式で提示されたデータを処理する場合でも堅牢性を維持し、一貫して高い信頼性を示しました。対照的に、語彙ベースのベースラインとトランスフォーマーのベースラインでは、統計的に有意な結果は得られませんでした。テストしたモデルの中で、GPT-4o-mini は、94% 低いコストで大型モデルと同等のパフォーマンスを達成し、スケーラブルな導入をサポートしました。このフレームワークには、モデルの信頼度評価と人間が判読できる根拠の説明 (xAI) も組み込まれており、製品満足度評価における実用化をサポートしながら、解釈可能性、透明性、信頼性を向上させます。一般に、アンケート手法として PDT ツールを費用効率の高い LLM とともにセンチメント分析に使用すると、センチメント スコア (数値センチメントと分類されたセンチメントの両方) の点で豊富な結果が得られる製品評価を提供できる可能性があり、製品の開発と改善のアイデアや、ターゲット ユーザー向けのマーケティング アイデアを特定するために使用できる製品の高レベルのユーザー インプレッションの点で役立ちます。

原文 (English)

Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability

Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure. This paper presents a scalable and interpretable framework that uses large language models (LLMs) to quantify product desirability from such data. Using two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA comprising 106 respondent term groupings with gold-standard human annotation, zero-shot continuous numerical sentiment scoring and categorical sentiment classification are evaluated without relying on explicit review scores. Across the datasets, LLMs generated numerical sentiment scores directly from qualitative responses and closely matched expert labels, achieving Pearson correlations up to 0.97 and classification accuracy up to 94%. LLMs maintained robustness even when handling data presented in multiple forms and consistently expressed high confidence. In contrast, lexicon-based and transformer baselines did not produce statistically significant results. Among the models tested, GPT-4o-mini achieved performance comparable to larger models at 94% lower cost, supporting scalable deployment. The framework also incorporates model confidence ratings and human-readable rationale explanations (xAI), improving interpretability, transparency, and trust while supporting practical use in product satisfaction assessment. In general, using the PDT tool as a survey method along with a cost efficient LLM for sentiment analysis has the potential to provide for product evaluation with results that are rich in terms of sentiment scores (both numerical and classified sentiment) and in terms of the high-level user impressions of the product that can be used to identify ideas for product development and improvement, as well as marketing ideas for target audiences.

13:00 JST研究/論文

分布シフト下での水中音響変調認識のための異種 2D/1D 信号表現の融合

変調認識システムは、異種信号表現に依存しています。時間周波数マップや周期定常マップなどの 2D 信号画像モダリティは構造パターンをキャプチャし、高次パワー スペクトルなどの 1D 統計記述子は相補的な手がかりをエンコードします。分布の変化の下では、これらのモダリティは不均一に劣化するため、堅牢な融合が実用化の中心的な課題となっています。異なるシフトタイプを体系的に分離する統一された評価プロトコルが欠如しているため、進歩はさらに制限されています。この論文では、水中音響変調認識におけるベンチマークとモデルの共同研究を通じて、両方の課題に対処します。 UAMR-ShiftBench は、配信内、低 SNR、目に見えない環境、目に見えない通信パラメータ、および測定された海上試験評価を単一の一致するプロトコルの下で共同でカバーする最初のベンチマークであり、南シナ海で 3 月と 11 月に実施された 2 つの海上試験キャンペーン中に収集された 2 つの独立した現実世界のサブセットを使用します。 SCP-TriCA は、STFT、周期定常性、および P2/P4 (2 次および 4 次パワースペクトル) モダリティを階層的に融合します。2 つの 2D モダリティは、まず双方向クロスアテンションを通じて調整され、次に 1D 統計モダリティがサンプル適応選択ゲートを通じて組み込まれます。 UAMR-ShiftBenchでは、SCP-TriCAは分布内精度95.33%とシミュレートOOD平均74.59%を達成し、最強のベースラインを5.12パーセントポイント上回り、2つの海上試験サブセットでは91.14%と94.86%に達し、最良のベースラインをそれぞれ15.71パーセントと23.00パーセントポイント上回りました。アブレーションの結果は、効果がモダリティの相補性と階層的融合設計に由来していることを裏付けています。コードとモデルは https://github.com/ronglaiqian/UAMR-ShiftBench で入手できます。

原文 (English)

Heterogeneous 2D/1D Signal Representation Fusion for Underwater Acoustic Modulation Recognition Under Distribution Shift

Modulation recognition systems rely on heterogeneous signal representations. 2D signal-image modalities such as time-frequency and cyclostationary maps capture structural patterns, while 1D statistical descriptors such as higher-order power spectra encode complementary cues. Under distribution shift, these modalities degrade unevenly, making robust fusion a central challenge for practical deployment. Progress is further limited by the lack of a unified evaluation protocol that systematically separates different shift types. This paper addresses both challenges through a joint benchmark-and-model study in underwater acoustic modulation recognition. UAMR-ShiftBench is the first benchmark to jointly cover in-distribution, low-SNR, unseen-environment, unseen-communication-parameter, and measured sea-trial evaluation under a single matched protocol, with two independent real-world subsets collected during two sea-trial campaigns conducted in March and November in the South China Sea. SCP-TriCA fuses STFT, cyclostationary, and P2/P4 (second- and fourth-order power spectra) modalities hierarchically: the two 2D modalities are first aligned through bidirectional cross-attention, and the 1D statistical modality is then incorporated through a sample-adaptive selective gate. On UAMR-ShiftBench, SCP-TriCA achieves 95.33% in-distribution accuracy and 74.59% simulated OOD average, outperforming the strongest baseline by 5.12 percentage points, and reaches 91.14% and 94.86% on the two sea-trial subsets, exceeding the best baseline by 15.71 and 23.00 percentage points respectively. Ablation results confirm that the gains stem from modality complementarity and the hierarchical fusion design. Code and models are available at https://github.com/ronglaiqian/UAMR-ShiftBench.

13:00 JST研究/論文

連続ウェアラブル生理学を使用した複数評価者による疼痛評価のイベント整合分析

痛みの評価は患者、看護師、臨床医によって異なりますが、ほとんどの計算アプローチは単一のグラウンドトゥルースラベルを前提としており、誰が評価を行っているかを事実上無視しています。私たちは、まばらな評価者固有の痛みの評価を離散的な痛みの変化イベントに変換し、連続的なウェアラブル生理学的信号をこれらのイベントに調整して、全体を通して評価者のアイデンティティを維持する、評価者対応のイベント調整フレームワークを導入します。このフレームワークは、脊椎関連疼痛処置中に収集されたマルチモーダルウェアラブルデータに適用され、評価者グループ間の実質的な不一致を特定し、報告された痛みの増加に先立って評価者に依存する生理学的差異の予備的、探索的証拠を提供します。これらの発見は、痛みと生理学的関係は評価者によって不変ではない可能性があり、評価者間で評価を集約すると意味のある生理学的パターンが隠蔽される可能性があることを示唆しています。したがって、評価者を意識した、イベントに合わせた視点は、現実世界の臨床疼痛評価におけるウェアラブル データを解釈するための有望な方向性となります。

原文 (English)

Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology

Pain is assessed differently by patients, nurses, and clinicians, yet most computational approaches assume a single ground-truth label - effectively ignoring who is doing the rating. We introduce a rater-aware, event-aligned framework that converts sparse, rater-specific pain ratings into discrete pain-change events and aligns continuous wearable physiological signals to these events, preserving rater identity throughout. Applied to multimodal wearable data collected during spine-related pain procedures, the framework identifies substantial disagreement across rater groups and provides preliminary, exploratory evidence of rater-dependent physiological differences preceding reported pain increases. These findings suggest that pain-physiology relationships may not be rater-invariant, and that aggregating assessments across raters may mask meaningful physiological patterns. A rater-aware, event-aligned perspective is therefore a promising direction for interpreting wearable data in real-world clinical pain assessment.

13:00 JST研究/論文

目に見えない電極生成によるEEG空間超解像度のための座標クエリ可能な神経場の再構成

実際の導入における EEG 空間超解像 (EEGSR) は、ランダムなチャネル欠落、不安定な電極品質、接触不良やデバイスのばらつきによる可視チャネル パターンの変化などの課題にさらされています。既存の EEGSR メソッドのほとんどは、事前定義された入出力レイアウトの下で固定の低から高へのチャネル マッピングを学習するため、欠落しているチャネルがテスト時に変化すると脆弱になります。この論文では、部分的に観察されたサポート チャネルから共有条件付き頭皮フィールドを学習するものとして EEGSR を再定式化します。具体的には、位置誘導エンコーダーは観察されたEEGチャネルとその座標を潜在条件に要約し、条件付き暗黙的神経表現デコーダーは、所望の電極座標でこの条件をクエリすることによってターゲットEEG信号を再構成します。推論中、モデルは利用可能な EEG サポートとクエリされた座標から目に見えない電極信号を直接再構築します。デコーダ上のエンコードされた潜在表現の制約を強化し、それによって観察されたチャネルと一致するより安定した頭皮フィールドを構築するために、混合電極状態下で忠実度を維持するチャネル破損トレーニング戦略をさらに導入します。複数のEEGデータセットにわたる広範な実験により、ランダムな欠落チャネル再構成と厳密な目に見えない電極信号生成の両方に対するフレームワークの有効性が実証されています。特に、AAD の厳密なホールドアウト電極設定の下では、私たちの方法は、最も強いベースラインよりも NMSE を 37.5% 削減し、SNR を 2.12 dB 改善し、トレーニング中に決して露出しない電極位置で信号を合成する能力を示しています。

原文 (English)

Coordinate-Queryable Neural Field Reconstruction for EEG Spatial Super-Resolution with Unseen-Electrode Generation

EEG spatial super-resolution (EEGSR) in real deployments is challenged by random channel missingness, unstable electrode quality, and changing visible-channel patterns caused by bad contacts or device variability. Most existing EEGSR methods learn a fixed low-to-high channel mapping under pre-defined input-output layouts, which makes them brittle when missing channels vary at test time. In this paper, we reformulate EEGSR as learning a shared conditional scalp field from partially observed support channels. Specifically, a position-guided encoder summarizes the observed EEG channels and their coordinates into a latent condition, and a conditional implicit neural representation decoder reconstructs target EEG signals by querying this condition at desired electrode coordinates. During inference, the model directly reconstructs unseen electrode signals from the available EEG support and the queried coordinates. To strengthen the constraint of the encoded latent representation on the decoder and thereby construct a more stable scalp field consistent with the observed channels, we further introduce a fidelity-preserving channel corruption training strategy under mixed electrode states. Extensive experiments across multiple EEG datasets demonstrate the effectiveness of our framework for both random missing-channel reconstruction and strict unseen-electrode signal generation. Notably, under the strict held-out-electrode setting on AAD, our method reduces NMSE by 37.5\% and improves SNR by 2.12 dB over the strongest baseline, showing its ability to synthesize signals at electrode locations never exposed during training.

13:00 JST研究/論文

拡散ベースの視覚条件付き音声強調のための視聴覚コントラスト調整

AVSE (Audio-Visual Speech Enhancement) は、唇の動きなどの視覚的な手がかりを利用して、騒がしい環境で音声を復元します。最近の研究では、拡散ベースの教師なし AVSE が導入されました。この AVSE では、交差注意を介して視覚的特徴に条件付けされた音声拡散モデルがトレーニングされ、事後サンプリング ベースの音声強調のためのデータ駆動型事前学習として使用されます。オーディオのみの対応物よりもパフォーマンスが期待できるにもかかわらず、融合におけるクロスモーダル調整を明示的に強制することの影響は依然として不明です。この研究では、事後サンプリングのフレームワークを変更せずに、視覚情報のより強力な使用を促進するために、対照的な視聴覚損失で拡散トレーニングの目的を強化することを提案します。一致したテストデータと不一致なテストデータにわたる実験では、低い SNR で最大の利得が得られ、干渉抑制、信号再構成、知覚品質が一貫して向上していることが示されています。コードは https://github.com/cexauce/AV-CA-DiffUSE で入手できます。

原文 (English)

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE, where a speech diffusion model conditioned on visual features via cross-attention is trained and used as a data-driven prior for posterior sampling-based speech enhancement. Despite promising performance over its audio-only counterpart, the impact of explicitly enforcing cross-modal alignment in the fusion remains unclear. In this work, we propose to augment the diffusion training objective with a contrastive audio-visual loss to encourage stronger use of visual information while keeping the posterior sampling framework unchanged. Experiments across matched and mismatched test data show consistent improvements in interference suppression, signal reconstruction, and perceptual quality, with the largest gains at low SNRs. Code is available at https://github.com/ cexauce/AV-CA-DiffUSE

13:00 JST研究/論文

マルコフ論理ネットワークによって定義されたランダムな色の有向グラフ

マルコフ論理ネットワーク (MLN) は、統計リレーショナル人工知能で、任意の有限領域 $D$ の領域 $D$ を持つ可能世界のセットにおける確率分布を定義するために使用される確率的関係モデルです。 MLN は、非負の実数である重みが関連付けられたソフト制約で構成されます。この研究では、プロパティ $P(x)$ と関係 $R(x, y)$ について話す言語を検討します。 $P(x)$ と $R(x, y)$ のすべてのブール値の組み合わせがソフト制約 (関連する重み付き) である MLN を検討します。 $n$ がドメインのサイズ (カーディナリティ) を表すものとします。すべての重みの選択について、重みが $1/n$ でスケーリングされる場合、すべての 1 次文 $\varphi$ について、$\varphi$ が保持する確率は $n \to \infty$ として 0 または 1 のいずれかになる傾向があることを示します。つまり、一次論理の 0-1 の法則が成り立ちます。さらに、限界確率は重みに依存しません。代わりに MLN の標準セマンティクスを使用する場合、重みがスケーリングされない場合、制限の動作はより複雑になり、重みに依存します。スケーリングされていない重みを使用すると、重みに応じて質的に異なる 7 つのケースが得られます。場合によっては、一次論理の 0-1 の法則が存在する場合もあれば、存在しない場合もありますが、依然として収束の法則が存在する可能性があります。一次文の漸近確率に対する重みの影響は、7 つのケースのうちの 1 つから別のケースへの突然の「相転移」の形で現れる可能性があります。収束則の存在は、大規模な領域での推論にプラスの影響を及ぼします。

原文 (English)

Random coloured digraphs defined by a Markov logic network

A Markov Logic Network (MLN) is a probabilistic relational model used in Statistical Relational Artificial Intelligence for defining a probability distribution on the set of possible worlds with domain $D$ for an arbitrary finite domain $D$. An MLN consists of soft constraints with associated weights which are nonnegative real numbers. In this study we consider a language speaking about a property $P(x)$ and a relation $R(x, y)$. We consider an MLN for which every Boolean combination of $P(x)$ and $R(x, y)$ is a soft constraint (with associated weight). Let $n$ denote the size (cardinality) of the domain. We show that, for every choice of weights, if the weights are scaled by $1/n$ then, for every first-order sentence $\varphi$, the probability that $\varphi$ holds tends to either 0 or 1 as $n \to \infty$; that is, a 0-1 law for first-order logic holds. Morover, the limit probability does {\em not} depend on the weights. If we instead use the standard semantics of MLNs, in the case of which the weights are {\em not} scaled, then the limit behaviour is more complicated and {\em depends} on the weights. With unscaled weights we get 7 qualitatively different cases which depend on the weights. In some cases we have a 0-1 law for first-order logic, in some cases not, but we may still have a convergence law. The influence of the weights on the asymptotic probability of a first-order sentence may be in the form of a sudden ``phase transition'' from one of the 7 cases to another. The presence of a convergence law has positive implications for inference on large domains.

13:00 JST研究/論文

法的推論は弁護士ではない:プロセクが司法にアクセスするための法的ベンチマークを再考する

Legal AI のベンチマーク調査では、大規模な言語モデルによって、法的権利を理解して行使するために弁護士に相談できない人々など、司法へのアクセスが向上するという仮定が頻繁に呼び出されます。現在のベンチマークは、モデルのパフォーマンスの上限を測定する、法律専門家によってすでに前処理された入力に対して法的推論を評価するため、この仮定をサポートする機能が備わっていないと主張します。正義へのアクセスは下限に依存します。つまり、プロンプトに騒々しい物語、埋もれた事実、省略、民俗法的な仮定、および表面レベルの誤りが含まれる可能性があるプロの訴訟当事者からの入力があった場合に、モデルがどのように機能するかです。これらの劣化は、ロングコンテキストの感度、過小仕様、幻覚、活字の乱れなど、一般的な機械学習の文献で LLM が劣化することが知られている条件に匹敵します。私たちは、散文文献からの証拠をこの一連の機械学習研究と結びつけ、法的ベンチマークである LEXam での小さな摂動実験を提示して、これら 2 つの限界間のギャップを説明します。モデル開発が上限のみを測定するベンチマークに焦点を当て続けた場合、このギャップは隠れたままになるか、さらに拡大する可能性があります。最後に、法的 AI に関する司法へのアクセスの主張が実証的に検証できるよう、専門的な入力の下で堅牢性を直接測定する法的ベンチマークを求めます。

原文 (English)

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice

Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for people who cannot access lawyers in order to understand and exercise their legal rights. We argue that current benchmarks are not equipped to support this assumption because they evaluate legal reasoning over inputs that have already been preprocessed by legal experts, which measures the upper bound of model performance. Access to justice depends on a lower bound: how models perform when inputs come from pro se litigants, whose prompts may contain noisy narratives, buried facts, omissions, folk-legal assumptions, and surface-level errors. These degradations are comparable to conditions under which LLMs are known to degrade in the general machine learning literature, including long-context sensitivity, underspecification, hallucination, and typographical perturbations. We connect evidence from pro se literature with this body of machine learning research and present a small perturbation experiment on LEXam, a legal benchmark, to illustrate the gap between these two bounds. If model development continues to focus on benchmarks that measure only the upper bound, this gap may remain hidden or even widen. We conclude by calling for legal benchmarks that directly measure robustness under pro se-like inputs so that access-to-justice claims about legal AI can become empirically testable.

13:00 JST研究/論文

LOLA のランタイム検証とモデルベースの診断のための統合フレームワーク

ストリーム仕様言語 LOLA 内でランタイム検証とモデルベースの診断を統合する統合フレームワークを紹介します。このアプローチでは、システムの説明、コンポーネントの健全性状態、および観察を単一のストリームベースの形式にエンコードすることで、別個のツールチェーンを必要とせずに、障害検出と並行して継続的なオンライン障害位置特定が直接可能になります。このフレームワークは、時間不変故障と一時故障の両方をサポートし、非決定的な観測にも自然に対応します。

原文 (English)

A Unified Framework for Runtime Verification and Model-Based Diagnosis in LOLA

We present an integrated framework that unifies runtime verification and model-based diagnosis within the stream specification language LOLA. By encoding system descriptions, component health states, and observations into a single stream-based formalism, the approach enables continuous, online fault localization directly alongside fault detection, without requiring separate toolchains. The framework supports both time-invariant and transient faults, and naturally accommodates nondeterministic observations.

13:00 JST研究/論文

オフライン推論トレーニングの重み空間幾何学

オフライン強化学習損失 (RFT、RIFT、DFT、オフライン GRPO、DPO) は、大規模な教師から小規模な生徒へ推論を抽出するために広く使用されており、通常は下流の精度のみで比較されます。それらが機構的に異なるのか、それとも同様の重み更新に収束するのかを尋ねます。アテンションのみの LoRA を使用して、単一のベース モデル (Qwen3-4B) からの同一の数学ロールアウトで 6 つのメソッド (SFT、RFT、DFT、RIFT、オフライン GRPO、DPO) をトレーニングし、コサイン類似度、主角部分空間分析、線形モード接続性、および CKA を介して結果のデルタを分析します。 (i) SFT、RFT、および RIFT はほぼ同一線上の重みデルタ (コサイン >= 0.97、144 モジュールにわたるトップ 1 主角は中央値約 7 度) と同等の GSM8K 精度 (87 ~ 88%、n=1319; ペアワイズ マクネマー p >= 0.15) を示します。 (ii) DFT は、同じデータを使用しているにもかかわらず、どの報酬重み付け手法よりも方向が大きく異なります。 (iii) オフライン GRPO は、SFT 損失領域内に留まりながら、SFT 方向に直交する実質的なコンポーネントを追加します (グローバルで最大 67%、後期レイヤーで最大約 86%)。 (iv) DPO はほぼ直交部分空間に位置し、モード接続性障壁を示し、後期層の CKA を ~0.46 まで崩壊させます。また、DPO は、GSM8K (93.5%、McNemar p < 10^-9 対他のメソッド) と AIME26 (30.0% 対 3.3-10.0%) の両方で、当社のプロトコルで最高の精度に達します。そのトレーニングでは他のものよりも 10 倍小さい学習率 (標準規約) が使用されるため、更新ノルムと精度のギャップは損失関数とオプティマイザーの選択を合わせて反映し、学習率と一致する DPO の比較は将来の作業に残されます。

原文 (English)

Weight-Space Geometry of Offline Reasoning Training

Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller students, and are typically compared on downstream accuracy alone. We ask whether they are mechanistically distinct or converge to a similar weight update. Training six methods (SFT, RFT, DFT, RIFT, Offline GRPO, DPO) on identical math rollouts from a single base model (Qwen3-4B) with attention-only LoRA, we analyze the resulting deltas via cosine similarity, principal-angle subspace analysis, linear mode connectivity, and CKA. We observe: (i) SFT, RFT, and RIFT have nearly colinear weight deltas (cosine >= 0.97, top-1 principal angle ~7 deg median over 144 modules) and comparable GSM8K accuracy (87-88%, n=1319; pairwise McNemar p >= 0.15); (ii) DFT diverges further in direction than any reward-weighted method despite using the same data; (iii) Offline GRPO adds a substantial component orthogonal to the SFT direction (~67% globally, up to ~86% in late layers) while staying in the SFT loss basin; (iv) DPO sits in a near-orthogonal subspace, shows a mode-connectivity barrier, and collapses late-layer CKA to ~0.46. DPO also reaches the highest accuracy in our protocol on both GSM8K (93.5%, McNemar p < 10^-9 vs. each other method) and AIME26 (30.0% vs. 3.3-10.0%); its training uses a 10x smaller learning rate than the others (the standard convention), so the update-norm and accuracy gaps reflect loss-function and optimizer choices jointly, and a learning-rate-matched DPO comparison is left for future work.

13:00 JST研究/論文

フェデレーションによる因果関係の発見と推論に関する調査

因果構造の発見と因果効果の推論を含む因果推論は、データ主導の意思決定の基礎です。実際には、信頼性の高い因果関係分析のためのデータは多くの場合、機関全体に分散されており、プライバシー規制や通信上の制約により一元化することができません。フェデレーテッド ラーニング (FL) は、生データを共有せずに共同分析を可能にすることでこの問題に対処し、急速に成長するフェデレーテッド 因果発見 (FCD) および推論 (FCI) の分野を生み出します。しかし、この分野の学際的な性質と包括的な調査がないことが、研究者にとって参入障壁となっています。この論文は、多次元の分類法による体系的なレビューを提供することで、そのギャップを埋めます。 FCD ソリューションの基礎となる 3 つの核となる設計上の決定、つまり構造の学習方法、データの分割方法、各当事者が取得する構造知識に基づいて、方法論的パラダイム、フェデレーション トポロジ、構造スコープの 3 つの軸に沿って FCD を整理します。さらに、時間的ダイナミクス、データの異質性、欠損データ、同一でない変数セットなど、主要な実際的な側面を調査します。 FCI については、古典的な重み付け手法から最新の深層生成アーキテクチャまで、目標推定値 (平均対個別/条件付き治療効果) および推定戦略によって手法を分類します。 FCD と FCI を別々に扱う以前の研究とは異なり、我々はそれらの関係を統一されたフェデレーション因果推論パイプラインの補完的な段階として形式化し、FCD が FCI での有効な効果推定に必要な構造的知識を提供します。最後に、プライバシー、通信効率、理論的保証、応用分野に関する彼らの共通の懸念を強調し、将来の研究に向けた未解決の課題を特定して締めくくります。

原文 (English)

A Survey on Federated Causal Discovery and Inference

Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making. In practice, data for reliable causal analysis are often distributed across institutions and cannot be centralized due to privacy regulations or communication constraints. Federated learning (FL) addresses this by enabling collaborative analysis without raw data sharing, giving rise to the rapidly growing field of federated causal discovery (FCD) and inference (FCI). However, the interdisciplinary nature of this field and the absence of a comprehensive survey present barriers to entry for researchers. This paper bridges that gap by providing a systematic review through multi-dimensional taxonomies. Grounded in the three core design decisions underlying any FCD solution, namely how structures are learned, how data are partitioned, and what structural knowledge each party obtains, we organize FCD along three axes: methodological paradigm, federation topology, and structural scope. We further examine key practical dimensions, including temporal dynamics, data heterogeneity, missing data, and non-identical variable sets. For FCI, we categorize methods by target estimand (average versus individualized/conditional treatment effects) and by estimation strategy, from classical weighting methods to modern deep generative architectures. Unlike prior works that treat FCD and FCI separately, we formalize their connection as complementary stages of a unified federated causal reasoning pipeline, where FCD supplies the structural knowledge required for valid effect estimation in FCI. Finally, we highlight their shared concerns regarding privacy, communication efficiency, theoretical guarantees, and application domains, and conclude by identifying open challenges for future research.

13:00 JST研究/論文

継続制御のためのトレーニング可能な非線形接続を備えた低電力アナログ ニューラル ネットワーク

物理ニューラル ネットワークは、アナログ デバイス物理学を直接計算することで低消費電力の機械学習を約束しますが、ほとんどのアーキテクチャでは、非線形デバイス応答がスカラー重みとして機能するように強制されます。 Kolmogorov-Arnold ネットワークからインスピレーションを得て、学習可能な非線形関数を接続に配置し、各物理接続を学習可能な計算要素にします。これらの機能をフィールドでプログラム可能なアナログ アレイ上のアナログ バンドパス フィルターとして実現すると、その利点はタスクに依存し、物理的基盤の滑らかさに起因することがわかります。ネットワークは、ロボット運動学、連続制御、太陽光発電の最大電力点追跡などの滑らかで継続的に値付けされるターゲットを表し、多層パーセプトロンよりもはるかに少ないノードと接続を備えていますが、分類のような決定境界ではパラメーター効率の利点はありません。トレーニングされたネットワークは、定量化された忠実度で約 35,000 の接続にわたってハードウェアに転送され、専用の CMOS 実装は約 30 マイクロワットで動作すると予測されています。メモリスティブの実現はシミュレーションで同じ動作を再現します。これは、特定のデバイスからではなく、接続にトレーニング可能な非線形性を配置することによって利点が得られることを示しています。

原文 (English)

Low-power analogue neural networks with trainable nonlinear connections for continuous control

Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures force nonlinear device responses to act as scalar weights. Inspired by Kolmogorov-Arnold networks, we place trainable nonlinear functions on the connections, making each physical connection a learnable computational element. Realising these functions as analogue band-pass filters on field-programmable analogue arrays, we find that the benefit is task-dependent and follows from the smoothness of the physical basis: the networks represent smooth, continuously valued targets, including robotic kinematics, continuous control, and photovoltaic maximum-power-point tracking, with far fewer nodes and connections than multilayer perceptrons, but offer no parameter-efficiency advantage on classification-like decision boundaries. Trained networks transfer to hardware across approximately 35,000 connections with quantified fidelity, and a dedicated CMOS implementation is projected to operate at approximately 30 microwatts. A memristive realisation reproduces the same behaviour in simulation, indicating that the advantage comes from placing trainable nonlinearity on connections, rather than from a particular device.

13:00 JST画像/動画生成エージェント

Sol ビデオ推論エンジン: 効率的なビデオ生成のためのエージェントネイティブのフルスタック アクセラレーション フレームワーク

最新のビデオ拡散モデルは、スケーリングを通じてより高い生成品質を実現しますが、これにより推論コストも増加します。多くの高速化方法が提案されていますが、中心的な課題は、最も効果的な高速化戦略が非常にインスタンス固有であるということです。つまり、モデル、ハードウェア、推論構成の 1 つの組み合わせでうまく機能するレシピが、別の組み合わせには移行しないことがよくあります。モデルが異なれば、アーキテクチャ、数値感度、注意集中パターンも異なります。推論設定は、空間的および時間的な解像度とビデオの長さが異なり、ハードウェア プラットフォームはメモリ階層、サポートされている数値形式、およびカーネル スループットが異なります。これらの要因により調整の余地が大きくなり、手動によるパフォーマンス エンジニアリングのコストが高くなります。我々は、ビデオ拡散モデルのためのエージェント的、ネイティブ、トレーニング不要のアクセラレーション フレームワークである Sol Video Inference Engine を紹介します。これは、キャッシュ、スパース アテンション、トークン プルーニング、量子化、カーネル フュージョンという 5 つの広く適用可能な手法を、インスタンス固有の最適化のためにエージェント アクセラレーション スタックにまとめています。モデル、ハードウェア プラットフォーム、およびサービス構成によって定義された具体的なデプロイメント ターゲットに対して、並列スキル エージェントが各手法の実装を最適化し、エージェント インテグレーターがそれらをグローバル アクセラレーション スタックに構成し、人間のバリデーターが生成品質に関するフィードバックを提供します。サイズとアーキテクチャが異なる 3 つのビデオ モデル (64B Cosmos3-Super、22B LTX-2.3、および 2B SANA-Video) でこのワークフローをインスタンス化します。人的労力をほとんどかけることなく、フルスタックはほぼロスレスの VBench 品質を維持しながら 2 倍を超えるエンドツーエンド アクセラレーションを達成し、ビデオ拡散アクセラレーションに対するエージェント フレームワークの有効性を実証しています。

原文 (English)

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation

Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have been proposed, a central challenge is that the most effective acceleration strategy is highly instance-specific: a recipe that works well for one combination of model, hardware, and inference configuration often does not transfer to another. Different models vary in architecture, numerical sensitivity, and attention concentration patterns. Inference settings differ in spatial and temporal resolution and video duration, while hardware platforms differ in memory hierarchy, supported numerical formats, and kernel throughput. These factors create a large tuning space, making manual performance engineering costly. We present Sol Video Inference Engine, an agentic, native, training-free acceleration framework for video diffusion models. It organizes five broadly applicable techniques, cache, sparse attention, token pruning, quantization, and kernel fusion, into an agentic acceleration stack for instance-specific optimization. For a concrete deployment target defined by a model, hardware platform, and serving configuration, parallel skill agents optimize the implementation of each technique, an agent integrator composes them into a global acceleration stack, and a human validator provides feedback on generation quality. We instantiate this workflow on three video models with different sizes and architectures: 64B Cosmos3-Super, 22B LTX-2.3, and 2B SANA-Video. With little human effort, the full stack achieves more than 2x end-to-end acceleration while maintaining near-lossless VBench quality, demonstrating the effectiveness of the agent framework for video diffusion acceleration.

13:00 JST研究/論文

JEDEL: 初期段階の創薬のためのゼロショット DNA エンコード ライブラリ設計

我々は、活性リガンドの三次元ファーマコフォア表現から直接、合成に対応した DNA エンコード ライブラリー (DEL) を生成するためのフレームワークである JEDEL を紹介します。 JEDEL は、ファーマコフォアの相互作用パターンを実用的でスケーラブルな合成指示にマッピングする最初のモデルであり、潜在的に数百万の分子を含むターゲット ライブラリの設計を可能にします。下流の合成計画を必要とする仮想化合物を生成する既存の生成アプローチとは異なり、JEDEL は購入可能なビルディング ブロックと検証済みの反応の範囲内で動作し、すべての出力が構築によって実験的に実現可能であることを保証します。 JEDEL は、ファーマコフォアの形状と分子構造の間の予測的アラインメントを学習し、これを大規模なコンビナトリアル合成ルートに解読します。 18 のタンパク質ターゲットにわたって、ターゲット特異的な再トレーニングを行わずに、予測結合親和性、ファーマコフォア回収率、およびサンプル効率においてランダムおよび多様性ベースのベースラインを上回る、集中的なライブラリーを生成します。 JEDEL を使用すると、仮想分子生成から実験的に展開可能なライブラリ設計への移行が可能になります。

原文 (English)

JEDEL: Zero-Shot DNA-Encoded Library Design for Early-Stage Drug Discovery

We present JEDEL, a framework for generating synthesis-ready DNA-encoded libraries (DELs) directly from three-dimensional pharmacophore representations of active ligands. JEDEL is the first model to map pharmacophore interaction patterns to actionable, scalable synthesis instructions, enabling the design of targeted libraries comprising potentially millions of molecules. Unlike existing generative approaches that produce virtual compounds requiring downstream synthesis planning, JEDEL operates within the space of purchasable building blocks and validated reactions, ensuring that every output is experimentally realizable by construction. JEDEL learns a predictive alignment between pharmacophore geometry and molecular structure and decodes this into combinatorial synthesis routes at scale. Across 18 protein targets, it generates focused libraries that outperform random and diversity-based baselines in predicted binding affinity, pharmacophore recovery, and sample efficiency, without target-specific retraining. JEDEL enables a shift from virtual molecule generation to experimentally deployable library design.

13:00 JST研究/論文

物理的に制約された MCMC と化学情報に基づいたガウス過程を相乗して反応ネットワークを発見

離散的な反応トポロジーと連続的な速度論的パラメーターは密接に結合しているため、まばらでノイズの多い化学時系列データから解釈可能な支配方程式を抽出することは依然として困難です。スパイクアンドスラブトポロジーサンプリング、ハード保存、熱力学スクリーニング、パラメーター校正と実験計画のための化学情報ガウスプロセス(CIGP)残差モデルを組み合わせた再現可能なグレーボックスワークフローであるPC-MCMC-CIGPを紹介します。方法論的な貢献は、新しい MCMC または GP ファミリー単独ではありません。むしろ、これらのコンポーネントを物理的に制約されたワークフローに統合し、明確な不確実性を考慮した取得選択を行うことです。 H2 + Br2 ベンチマークでは、制約付きサンプラーは、実験で基本的なラジカル経路を欺瞞的な現象学的適合から区別します。スチレンのエポキシ化では、CIGP 最適化ループにより、報告されている GP-BO ベースラインよりも最終収率が 12.5% 向上しました。新しい 10 シード取得研究では、EI、GWU、PC-EI、不確実性サンプリング、不一致探索、およびランダム検索には異なるトレードオフがあることが示されています。PC-EI は低利回りの BO 提案を大幅に削減する一方、EI スタイルの基準は最強の最終利回りパフォーマンスをもたらします。

原文 (English)

Synergizing Physically Constrained MCMC and Chemical-Informed Gaussian Processes for Reaction Network Discovery

Extracting interpretable governing equations from sparse, noisy chemical time-series data remains difficult because discrete reaction topology and continuous kinetic parameters are tightly coupled. We present PC-MCMC-CIGP, a reproducible gray-box workflow that combines spike-and-slab topology sampling, hard conservation and thermodynamic screening, and a Chemical-Informed Gaussian Process (CIGP) residual model for parameter calibration and experimental design. The methodological contribution is not a new MCMC or GP family in isolation; rather, it is the integration of these components into a physically constrained workflow with explicit uncertainty-aware acquisition choices. On the H2 + Br2 benchmark, the constrained sampler distinguishes elementary radical pathways from deceptive phenomenological fits in our experiments. On styrene epoxidation, the CIGP optimization loop improves final yield by 12.5% over the reported GP-BO baseline. A new 10-seed acquisition study shows that EI, GWU, PC-EI, uncertainty sampling, discrepancy hunting, and random search have different trade-offs: PC-EI substantially reduces low-yield BO suggestions, while EI-style criteria give the strongest final-yield performance.

13:00 JST研究/論文

オープンセットシナリオにおけるドメイン汎化を強化するための二元論的メタ学習の探求

ドメインの一般化は、複数のソース ドメインから学習して、目に見えないターゲット ドメインに一般化します。ただし、ソースとターゲットの間のラベルの不一致という現実的なケースが無視されることがよくあります。次に、未知のドメイン内の未知のクラスを認識するために、開集合ドメイン一般化が提案されます。簡単なアプローチでは、1 対全分類器をトレーニングして各クラスを分離し、外れ値を未知として検出します。しかし、少数の正のサンプルと多数の負のサンプルの間の不均衡により、決定境界が正の方向に偏り、モデルは、たとえ目に見えない領域の既知のクラスからであっても、分布外のデータを過剰に拒否することになります。この論文では、Dualistic MEta-learning with Joint DomaIn-Class Matching (MEDIC) と呼ばれる新しいメタ学習戦略を提案します。これは、ドメイン間およびクラス間のタスク分割に向けた暗黙的な勾配マッチングを考慮して、ドメインとクラスの両方でバランスの取れた最適な境界を見つけます。実験結果は、MEDIC がオープンセットのシナリオで従来の方法より優れているだけでなく、競合するクローズセット汎化能力を維持していることを示しています。

原文 (English)

Exploring Dualistic Meta-Learning to Enhance Domain Generalization in Open Set Scenarios

Domain generalization learns from multiple source domains to generalize to unseen target domains. However, it often neglects the realistic case of label mismatch between source and target. Open set domain generalization is then proposed to recognize unseen classes in unseen domains. A simple approach trains one-vs-all classifiers to separate each class and detect outliers as unknown. Yet, the imbalance between few positive samples and many negative samples skews the decision boundary towards the positive ones, leading the model to over-reject out-of-distribution data, even from known classes in unseen domains. In this paper, we propose a novel meta-learning stategy called dualistic MEta-learning with joint DomaIn-Class matching (MEDIC), which considers implicit gradient matching towards inter-domain and inter-class task splits simultaneously to find optimal boundaries balanced for both domains and classes. Experimental results show that MEDIC not only outperforms prior methods in open set scenarios, but also maintains competitive close set generalization ability.

13:00 JSTLLM/生成AIGPT / ChatGPTNVIDIA

VeriPilot: LLM を利用した Verilog デバッグ フレームワーク

Verilog デバッグは、依然としてデジタル回路設計において最も時間のかかる段階の 1 つです。大規模言語モデル (LLM) の最近の進歩により、自動デバッグが可能になりました。ただし、既存のアプローチのほとんどは、エンドツーエンドの方法でテスト出力とコンパイラのフィードバックのみに依存しているため、複雑なバグに対する効果は限られています。主な課題は、エラーの根本原因が観察可能な出力から遠く離れている可能性があり、LLM がコード内の長い依存関係チェーンを追跡することが困難になることです。この課題は、コンテキストの長さが長いため効率的な推論が妨げられる大規模なコードベースではさらに悪化します。これらの制限に対処するために、私たちは、ゴールデン リファレンス モデルを活用してきめ細かいバグの位置特定と修復を可能にする、LLM を利用したデバッグ フレームワークである VeriPilot を提案します。 VeriPilot は、LLM ベースの分析を通じて Verilog 設計とそれに対応するゴールデン モデルの間の内部変数セマンティクスを調整することで、出力レベルの比較を超えています。次に、静的解析から得られたコントロール データ フロー グラフ (CDFG) を使用して段階的な信号トレースを実行し、疑わしいコード領域の最小限のセットと、ゴールデン モデルからの正しい対応部分を特定します。これらの構造化された洞察は、その後、推論と自動コード修復をガイドするために LLM に提供されます。 NVIDIA の包括的 Verilog 設計問題 (CVDP) ベンチマークの実験結果は、VeriPilot が GPT-4o の修復成功率を 54.3\% から 85.71\% に向上させ、複雑な Verilog 設計のバグ位置特定の精度と修復効果の両方を大幅に向上させることを示しています。ソース コードとベンチマークは、Github https://github.com/YihanWn/VeriPilot.git で公開されています。

原文 (English)

VeriPilot: An LLM-Powered Verilog Debugging Framework

Verilog debugging remains one of the most time-consuming stages in digital circuit design. Recent advances in Large Language Models (LLMs) have enabled automated debugging; however, most existing approaches rely solely on test outputs and compiler feedback in an end-to-end manner, limiting their effectiveness on complex bugs. A key challenge is that the root cause of an error may be far removed from its observable outputs, making it difficult for LLMs to trace long dependency chains in code. This challenge is further exacerbated in large codebases, where long context lengths hinder efficient reasoning. To address these limitations, we propose VeriPilot, an LLM-powered debugging framework that leverages golden reference models to enable fine-grained bug localization and repair. VeriPilot goes beyond output-level comparison by aligning internal variable semantics between the Verilog design and its corresponding golden model through LLM-based analysis. It then performs step-by-step signal tracing using Control-Data-Flow Graphs (CDFGs) derived from static analysis, identifying a minimal set of suspicious code regions along with their correct counterparts from the golden model. These structured insights are subsequently provided to the LLM to guide reasoning and automated code repair. Experimental results on the Comprehensive Verilog Design Problems (CVDP) benchmark from NVIDIA demonstrate that VeriPilot improves the repair success rate of GPT-4o from 54.3\% to 85.71\%, significantly enhancing both bug localization accuracy and repair effectiveness for complex Verilog designs. The source code and benchmark are publicly available at Github https://github.com/YihanWn/VeriPilot.git.

13:00 JSTエージェントロボティクス

信頼性の高い自律システムのエンジニアリング: 課題と解決策

信頼性の高い自律システムのエンジニアリングは、コンピューター サイエンスにおける重要かつ成長中のテーマです。自律システムが普及するにつれて、自律システムを確実に構築するための使いやすい技術の重要性が増しています。このワークショップレポートは、2024年6月10日から14日まで開催されたローレンツセンターワークショップ「信頼性の高い自律システムのエンジニアリング」(ERAS)での議論を取りまとめ、拡張したものです。このワークショップは、自律システムのための形式手法に関するワークショップ(FMAS)と信頼性の高いエンジニアリング自律システムのためのエージェントとロボットに関するワークショップ(AREA)の主催者によって共催されました。この会合には、FMAS および AREA コミュニティのメンバー、業界関係者、自律システムが特有のエンジニアリング課題を引き起こす分野の代表者が集まりました。このワークショップでは、自律システムの検証と妥当性確認の技術という 3 つの主要な研究トピックに焦点を当てました。現実世界の自律システムをエンジニアリングする。安全な自律システムのためのソフトウェア アーキテクチャ。その主な成果は、これらの分野における課題のカタログであり、最も重要なことに、解決策への道筋です。一部の課題は、学界ではよく知られているものの、実際にはまだ定期的に使用されていない手法ですでに取り組むことができます。その他の課題は未解決のままであり、さらなる研究が必要です。このロードマップは、将来の研究と産業協力をサポートすることを目的としています。

原文 (English)

Engineering Reliable Autonomous Systems: Challenges and Solutions

Engineering reliable autonomous systems is an important and growing topic in computer science. As autonomous systems become more prevalent, easy-to-use techniques for building them reliably are increasingly important. This workshop report captures and expands on the discussions at the Lorentz Center Workshop "Engineering Reliable Autonomous Systems" (ERAS), held from 10 to 14 June 2024. The workshop was co-organised by the organisers of the Workshop on Formal Methods for Autonomous Systems (FMAS) and the Workshop on Agents and Robots for reliable Engineered Autonomy (AREA). It brought together members of the FMAS and AREA communities, industry practitioners, and representatives from sectors where autonomous systems pose distinctive engineering challenges. The workshop focused on three main research topics: techniques for verification and validation of autonomous systems; engineering real-world autonomous systems; and software architectures for safe autonomous systems. Its main outcome is a catalogue of challenges in these areas and, most importantly, a pathway to solutions. Some challenges can already be tackled by techniques that are well known in academia but have not yet become regularly used in practice. Other challenges remain unresolved and require further research. This roadmap is intended to support future research and industrial collaboration.

13:00 JST研究/論文

デュアルブランチ スパイキング ニューラル ネットワークによるニューロモーフィック音声強化

スパイキング ニューラル ネットワーク (SNN) ベースのニューロモーフィック音声強調は、そのエネルギー効率により有望なパラダイムとして浮上していますが、バイナリ アクティベーションと適切に設計されたネットワーク アーキテクチャの欠如により、依然として古典的な人工ニューラル ネットワーク (ANN) ベースのアプローチよりも性能が劣っています。この制限を克服するために、GSU-DBNet と呼ばれる、ゲート スパイキング ユニット (GSU) を備えた新しいデュアル ブランチ スパイキング ニューラル ネットワーク アーキテクチャを提案します。具体的には、GSU-DBNet は音声の振幅スペクトルと複素スペクトルを同時にモデル化し、対応する振幅と複素スペクトル マスクを予測します。一方、デュアルパス GSU モジュールが採用され、時間情報と周波数情報を活用して時空間特徴表現を強化します。人気のベンチマーク データセットでの実験では、GSU-DBNet がわずか 394,000 のパラメータで PESQ スコア 3.04 を達成し、代表的な ANN ベース モデルのパラメータの 4.5% ~ 10.6% のみを使用しながら、既存の SNN ベースの手法を上回るパフォーマンスを示していることが示されています。

原文 (English)

Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks

Spiking neural network (SNN)-based neuromorphic speech enhancement has emerged as a promising paradigm due to its energy efficiency, yet it still underperforms classical artificial neural network (ANN)-based approaches owing to binary activations and the lack of well-designed network architectures. To overcome this limitation, we propose a novel dual-branch spiking neural network architecture equipped with a gated spiking unit (GSU), termed GSU-DBNet. Specifically, GSU-DBNet simultaneously models the speech magnitude spectrum and complex spectrum, predicting the corresponding magnitude and complex spectral masks. Meanwhile, a dual-path GSU module is adopted to exploit temporal and frequency information for enhanced spatiotemporal feature representation. Experiments on a popular benchmark dataset show that GSU-DBNet achieves a PESQ score of 3.04 with only 394K parameters, outperforming existing SNN-based methods while using only 4.5%--10.6% of the parameters of representative ANN-based models.

13:00 JST画像/動画生成

聞くことで VLM のビジョンが明確になります

最近の研究では通常、回答側トークンの注意分布を使用して視覚と言語の一貫性を評価します。ただし、最も注目度の高い領域が、意図したセマンティック トークンと常に一致するとは限らないことが観察されています。これはおそらく、以前に生成された応答トークンからの言語事前分布が蓄積され、視覚的な注意と不一致となるデコード ドリフトに起因します。以前の回答トークンからの事前確率に加えて、モダリティ境界マーカーなどの構造トークンがコンテキスト全体を包含し、ターゲットに関係のない領域への高い注目を生成する可能性があることがわかりました。これらの歪みを回避し、大規模な VLM の一貫性評価を提供するために、プロンプト側のセマンティクスを採用し、Prompt-Vision Token Activation Map (PV-TAM) を提案します。 PV-TAM にはさらに、モダリティ境界マーカーによって引き起こされる系統的なバイアスを除去するフィルターが組み込まれています。アクティベーション強度を無視してマスクのみを介して重複を評価する従来の方法とは異なり、当社のメトリクスは注意のピーク分布を利用して、プロンプトと視覚領域間の整合性を測定します。実験では、PV-TAM は、さまざまなデータセットの回答側のベースラインよりもアテンション ベースと IoU スタイルのローカリゼーション メトリックの両方を一貫して改善しました。

原文 (English)

Listening makes Vision Clear for VLMs

Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens. However, we observe that highest attention regions are not always consistent with the intended semantic token. This probably stems from decoding drift, where language priors from previously generated answer tokens accumulate and mismatch with visual attention. Besides the priors from previous answer tokens, we find that structural tokens, e.g., modality boundary markers, may encompass the entire context and generate high attention to areas unrelated to the target. To avoid these distortions and provide consistency evaluation for large VLMs, we adopt prompt-side semantics and propose Prompt-Vision Token Activation Map (PV-TAM). PV-TAM further incorporates a filter to remove systematic bias induced by modality boundary markers. Unlike traditional methods that evaluate overlap solely through masks while ignoring activation intensity, our metrics leverage the peak distribution of attention to measure the alignment between prompts and visual regions. In experiments, PV-TAM consistently improves both attention-based and IoU-style localization metrics over answer-side baselines on various datasets.

13:00 JSTLLM/生成AIエージェント

LLMエージェント社会における創発的な関係秩序:集団的影響から権威階層化まで

フェイ・シャオトンの差別的秩序パターンは、農村社会が自己中心的で関係性に段階があり、社会的距離が離れると協力が弱まるという特徴を持っています。文化的に特殊なものとして扱われることが多いものの、そのメカニズムの基礎は依然として十分に運用されておらず、これまでの LLM ベースのシミュレーションは主に長期的な社会構造ではなく、短期的な調整を扱っていました。私たちは、感情制御理論、社会的アイデンティティ理論、およびデュルケミアンの集合的感情に基づいたマルチエージェントフレームワークであるCAREB-MASを提案します。エージェントは、感情-倫理-信念の連鎖を通じて推論し、動的に進化する自己中心的なアイデンティティを維持しますが、マクロ環境は、個々の生産、好みに基づく割り当て、および最小限の対話プロトコルのみを指定します。長期的なシミュレーションを通じて、エージェントは 5 つの中心的な差異秩序現象、つまり安定した労働専門化、関西に基づく経済倫理、協力関係の衰退、新興の関係的権威、氏族ベースの中心部と周縁部の階層化を自発的に再現します。これらのパターンは、生産構造とともに血族中心の統合からより大きな機能的相互依存へと移行します。広範な実験結果は、社会構造と変化を研究するための学際的なフレームワークを提供する LLM ベースのマルチエージェント シミュレーションにより、差分秩序を一般的な社会メカニズムの構造に敏感な創発的な結果として解釈することをサポートしています。

原文 (English)

Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification

Fei Xiaotong's Differential Order Pattern characterizes rural society as egocentric and relationally graded, with cooperation attenuating over social distance. Although often treated as culturally specific, its mechanistic basis remains under-operationalized, and prior LLM-based simulations have mainly addressed short-term coordination rather than long-horizon social structure. We propose CAREB-MAS, a multi-agent framework grounded in Affect Control Theory, Social Identity Theory, and Durkheimian collective affect. Agents reason through an emotion-ethics-belief chain and maintain dynamically evolving egocentric identities, while the macro environment specifies only individual production, preference-based allocation, and minimal interaction protocols. Across long-horizon simulations, agents spontaneously reproduce five core Differential Order phenomena: stable labor specialization, guanxi-based economic ethics, relational decay of cooperation, emergent relational authority, and clan-based center-periphery stratification. These patterns shift with production structure from kin-centered integration toward greater functional interdependence. Extensive experiment results support interpreting Differential Order as a structure-sensitive emergent outcome of general social mechanisms, with LLM-based multi-agent simulation providing an interdisciplinary framework for studying social structure and change.

13:00 JSTエージェント

信頼できる AI の有効性を示す暗号証明書

私たちは、エージェント AI システムの有効性を示す暗号証明書を提案します。中心となるアイデアは、正当性またはポリシー条件を論理述語として正式に指定し、この述語を多項式制約上の証人確認問題にコンパイルし、簡潔な暗号証明システム (およびオプションでゼロ知識) を使用して条件が成立することを証明することです。これは、ソース コードの正式な検証と暗号化認証の間の中間点を提供します。エージェントのアクションには、検証者がエージェントを信頼したり計算を再実行したりする必要がなく、合意された正式なポリシーを満たしていることを独立してチェックできる証明を伴うことができます。アプローチの概要を高レベルで説明し、核となる数学的変換を示し、提案を証明力のあるコード、zkVM、形式的手法、およびエージェント ガバナンスに関連付け、完全な実装が回答する必要がある仕様、監査、および展開に関する質問に注意します。

原文 (English)

Cryptographic certificates of validity for trustworthy AI

We propose cryptographic certificates of validity for agentic AI systems. The core idea is to formally specify a correctness or policy condition as a logical predicate, compile this predicate to a witness-checking problem over polynomial constraints, and use a succinct cryptographic proof system (and optionally zero-knowledge) to certify that the condition holds. This offers a middle ground between formal verification of source code, and cryptographic authentication. An agent's action can be accompanied by an independently checkable proof that it satisfies an agreed formal policy, without requiring the verifier to trust the agent or to re-execute computation. We outline the approach at a high level, give the core mathematical translation, relate the proposal to proof-carrying code, zkVMs, formal methods, and agent governance, and note the specification, auditing, and deployment questions that a full implementation must answer.

13:00 JST研究/論文

5G を介した XR でのリアルタイム アバター制御のための統合されたセンシングと通信

拡張現実 (XR) は、5G および 6G ネットワークにとって困難なユースケースを示しており、真に没入型のエクスペリエンスを提供するには、高いデータレートと低遅延の通信が必要です。さらに、物理的な動作を仮想世界にシームレスに変換するには、正確なジェスチャ認識と姿勢推定が必要です。ハンドヘルド コントローラーとカメラをベースとした現在の XR インタラクション ソリューションでは、全身のポーズを簡単にキャプチャすることができず、手を自由に使うことができず、良好な視認性と明確な視線が必要です。この研究では、5G ミリ波 (mmWave) 統合センシングおよび通信 (ISAC) 信号と表面筋電図 (sEMG) 信号を組み合わせた XR 用のマルチモーダル センシング アーキテクチャを提案します。 5G ミリ波 ISAC は、コンテンツをヘッドマウント ディスプレイ (HMD) にワイヤレスで配信するために使用できるだけでなく、同じ通信信号を使用してユーザーの身体レベルの大まかなジェスチャやポーズを導き出し、リアルタイムのアバター制御をサポートすることもできます。きめ細かい指レベルのジェスチャを実現するために、当社のアーキテクチャは前腕の筋肉の活動を捕捉する軽量の sEMG センサーを活用しています。両方のモダリティの必要性を説明するために、両方のセンシング技術の評価を示します。ボディ レベル (5G) では、当社のアーキテクチャは、5G NR 標準の標準ビーム管理またはビーム スイープ手順から計算できるビーム ペアあたりの電力 (PPBP) に依存しています。 PPBP ベースのセンシングは、トレーニング中に見られなかったユーザーについて評価した場合、平均 82.2$\pm$5.9% の精度を達成します。きめの細かい指レベルのインタラクションについては、表面筋電図 (sEMG) が強力な識別情報を伝達し、さまざまな動作設定にわたって一貫した有望なパフォーマンスを実現することを示します。したがって、2 つのモダリティを組み合わせることで、既存の 5G 信号を介して身体レベルで、軽量 sEMG センサーを介して指レベルでマルチスケールのジェスチャ認識が可能になり、完全な XR フレームワークが形成されます。

原文 (English)

Integrated Sensing and Communications for Real-time Avatar Control in XR over 5G

Extended Reality (XR) presents a challenging use case for 5G and 6G networks, requiring high data-rates and lowlatency communication to deliver a truly immersive experience. Moreover, in order to seamlessly translate physical actions to the virtual world, accurate gesture recognition and pose estimation are required. Current XR interaction solutions based on handheld controllers and cameras cannot easily capture full-body poses, inhibit the free use of hands, and require good visibility and a clear line of sight. In this work, we propose a multimodal sensing architecture for XR that combines 5G MillimeterWave (mmWave) Integrated sensing and communication (ISAC) and surface electromyography (sEMG) signals. 5G mmWave ISAC cannot only be used to deliver content wirelessly to the Head-mounted display (HMD), but also the same communication signals can be used to derive coarse body-level gestures and poses of the user, to support real-time avatar control. For fine-grained finger-level gestures, our architecture leverages lightweight sEMG sensors that capture forearm muscle activity. To illustrate the need of both modalities, we present evaluations of both sensing technologies. At the body level (5G), our architecture relies on power-per-beam-pair (PPBP), which can be computed from standard beam management or beam sweeping procedures of the 5G NR standard. PPBP-based sensing achieves 82.2$\pm$5.9% average accuracy when evaluated on users not seen during training. For fine-grained finger-level interactions, we show that surface electromyography (sEMG) carries strong discriminative information achieving consistent promising performance across different movement settings. Thus, combining the two modalities enables multi-scale gesture recognition, at the body level via existing 5G signals and finger level via lightweight sEMG sensors, forming a complete XR framework.

13:00 JSTLLM/生成AIエージェント

タスクガイド付き会話グラフから目標指向の対話ランタイムまで

グラフおよびマルチエージェント オーケストレーション フレームワークは、本番環境の大規模言語モデル (LLM) ワークフローを実用的なものにしますが、ユーザーが相互に依存する複数の目的を維持する場合、それらだけでは会話の連続性を解決できません。この概念システムに関する論文は、他の目標のアクションによって目標が一時停止、再開、修正、無効化される可能性がある、設計空間の非常に複雑な部分に焦点を当てています。目標指向ダイアログ ランタイム (GODR) を導入します。これは、目標、タスク フレーム、ライフサイクル状態、無効化ルール、および再開コントラクトを最上級のランタイム オブジェクトとして扱い、制限された実行をグラフ ランタイム、エージェント、ツール、またはアプリケーション プログラミング インターフェイス (API) に委任するフレームワーク中立の設計パターンです。 GODR は、単純なガイド付きプロセスにおけるワークフロー グラフの代替として提案されていません。これは、エージェント ID、チャット履歴、または実行グラフの位置だけでは客観的な継続性を確実に回復できない、複雑でマルチドメインの中断可能な会話を対象としています。この論文では、問題を形式化し、ランタイム オブジェクトとアーキテクチャの選択基準を提案し、測定されたパフォーマンスの主張ではなく将来の経験的検証の課題として評価を組み立てています。

原文 (English)

From Task-Guided Conversational Graphs to Goal-Oriented Dialogue Runtimes

Graph and multi-agent orchestration frameworks make production large language model (LLM) workflows practical, but they do not by themselves solve conversational continuity when users maintain several interdependent objectives. This conceptual systems paper focuses on the high-complexity end of that design space, where goals can be suspended, resumed, revised, and invalidated by actions in other goals. We introduce the Goal-Oriented Dialogue Runtime (GODR), a framework-neutral design pattern that treats goals, task frames, lifecycle state, invalidation rules, and resumption contracts as first-class runtime objects while delegating bounded execution to graph runtimes, agents, tools, or application programming interfaces (APIs). GODR is not proposed as a replacement for workflow graphs in simple guided processes; it is intended for complex, multi-domain, interruptible conversations where objective continuity cannot be recovered reliably from agent identity, chat history, or execution-graph position alone. The paper formalizes the problem, proposes runtime objects and architecture-selection criteria, and frames evaluation as an agenda for future empirical validation rather than as a measured performance claim.

13:00 JST研究/論文

電車の 10 桁: 2 つの固有値問題の AI 支援検証

特に特異な設定や非正規の設定では、正確な数値固有値を証明するのが難しいことがよくあります。この記事では、そのような 2 つの計算における人間と AI のコラボレーションについて報告します。特異な自己共役シュルディンガー演算子の場合、検証されたゼロ カウントとディリクレ - ノイマン括弧法により、完全な負のスペクトルが小数点以下 10 桁まで証明されます。繊細な非正規原子 - 分子ベンチマークの場合、以前に未解決の共鳴ペアが分離され、各メンバーが 10 桁に囲まれます。2 番目の結果は、一方向シューティングの精度を向上させることではなく、射影解のためのグローバル マッチング システムとして問題を再定式化することによって達成されます。無限の尾部は終端射影データの不確実性としてエンコードされ、コンポーネントごとに尾部に堅牢な Krawczyk-Brouwer 包含によって証明書が提供されます。これにより、AI 支援の強みと限界が明らかになり、その中には、明らかに完全な尾部引数が 1 つ含まれていました。不均一なポリディスクに必要なコンポーネントごとのチェックは、AI 支援の数学の厳格なテストです。これらの例は、証明オブジェクトが重要である理由、およびより広範には、AI によってコード、説明、および妥当性のある数値的主張が可能になるため、その影響が非常に不安定であることを示しています。

原文 (English)

Ten Digits on a Train: AI-Assisted Verification of Two Eigenvalue Problems

Accurate numerical eigenvalues are often difficult to certify, especially in singular or non-normal settings. This article reports a human--AI collaboration on two such computations. For a singular self-adjoint Schr\"odinger operator, a verified zero count and Dirichlet--Neumann bracketing certify the complete negative spectrum to ten decimal places. For a delicate non-normal atom--molecule benchmark, a previously unresolved resonance pair is separated, with each member enclosed to ten digits. The second result is achieved not by increasing the precision of one-way shooting, but by reformulating the problem as a global matching system for projective solution lines. The infinite tail is encoded as uncertainty in the terminal projective data, and a componentwise, tail-robust Krawczyk--Brouwer inclusion supplies the certificate. This gives a reusable architecture for analytic boundary-value systems with ill-conditioned propagation and uncertain asymptotic data. The collaboration also exposes the strengths and limits of AI assistance. AI rapidly produced accurate candidates and plausible proof strategies, but several failed, including one apparently complete tail argument that omitted the componentwise check required by a nonuniform polydisc. Validated computation is a stringent test of AI-assisted mathematics: the output is not merely a number, but a number with a proof. These examples show why the proof object matters, and why human mathematical judgment remained decisive. More broadly, as AI makes code, exposition, and plausible numerical claims inexpensive, standards for verification, attribution, peer review, and training must adapt. The implications are unsettling; the opportunity is extraordinary.

13:00 JST画像/動画生成

空間からスペクトルへ: 小さな物体検出のための効率的な周波数ガイド付き特徴表現学習器

効率的な小さな物体の検出は、小さなターゲットに固有の特徴の不足によってボトルネックになっており、重要な高周波の詳細を無差別に破棄する空間領域検出器の動作によってさらに悪化します。空間領域内でこれらの壊れやすい手がかりを回復することは、多くの場合、計算コストのかかるアーキテクチャのアップスケーリングを必要とし、背景ノイズを誤って増幅してしまうため、難しいことで知られています。このギャップを埋めるために、私たちは \textbf{空間からスペクトルへのシフト} 特徴処理パラダイムを提案し、次の新規性を備えた総合的なソリューションを導入します。 (1) 多様な検出器アーキテクチャ (CNN ベースとトランスフォーマー ベースの両方) 全体で一般化する多用途の \textbf{周波数誘導特徴表現フレームワーク}。空間のみの特徴抽出に代わる堅牢な代替手段を提供します。 (2) 統合された \textbf{Decompose--Enhance--Reconstruct (DER)} オペレーター。ウェーブレット差分ゲート (WDG)、ログガボール エンハンサー (LGE)、および周波数駆動ヘッド (FDHead) の 3 つの \textbf{軽量、プラグアンドプレイ} モジュールを介してインスタンス化され、バックボーン、ネック、ヘッドに周波数を意識した変調を系統的に注入します。このメカニズムは、解像度の低下から特徴モデリングを切り離し、識別可能な高周波成分をキャプチャして、パラメーターの冗長性を大幅に削減して正確な位置特定を可能にします。 (3) マルチドメイン ベンチマーク (VisDrone2019、UAVDT、Tinyperson、DOTAv1) での広範な検証により、一貫した利益が実証されました。特に、私たちが提案する \textbf{DERNet} シリーズは、厳密なスペクトル診断と誤差分解分析に裏付けられた \textbf{パラメーターの 1/6 のみ} を必要としながら、同じスケールの下で YOLOv11 モデルよりも優れたパフォーマンスを発揮します。

原文 (English)

From Spatial to Spectral: An Efficient, Frequency-Guided Feature Representation Learner for Small Object Detection

Efficient small object detection is bottlenecked by the inherent feature scarcity of tiny targets, which is further aggravated by operations of spatial-domain detectors that indiscriminately discard critical high-frequency details. Recovering these fragile cues within the spatial domain is notoriously difficult, as it often requires computationally expensive architectural upscaling that inadvertently amplifies background noise. To bridge this gap, we propose a paradigm \textbf{shift from spatial to spectral} feature processing, introducing a holistic solution with the following novelty: (1) A versatile \textbf{Frequency-Guided Feature Representation framework} that generalizes across diverse detector architectures (both CNN and Transformer-based), offering a robust alternative to spatial-only feature extraction; (2) The unified \textbf{Decompose--Enhance--Reconstruct (DER)} operator, instantiated via three \textbf{lightweight, plug-and-play} modules -- Wavelet-Difference Gate (WDG), Log-Gabor Enhancer (LGE), and Frequency-Driven Head (FDHead) -- to systematically inject frequency-aware modulation into the backbone, neck, and head. This mechanism decouples feature modeling from resolution reduction, capturing discriminative high-frequency components to enable accurate localization with significantly reduced parameter redundancy; (3) Extensive validation on multi-domain benchmarks (VisDrone2019, UAVDT, TinyPerson, DOTAv1) demonstrating consistent gains. Notably, our proposed \textbf{DERNet} series outperforms YOLOv11 models under the same scale while requiring \textbf{only 1/6 of the parameters}, backed by rigorous spectral diagnostics and error decomposition analysis.

13:00 JST研究/論文

正確なエピトープ予測のための 3D 分子表面の指紋の解読

分子表面は、エピトープ予測の中心となる抗体抗原認識を決定する幾何学的および物理化学的パターンをコードしています。しかし、既存の方法は配列または骨格構造に依存しており、不連続な表面駆動型エピトープを捕捉するのに苦労しています。この研究では、分子表面表現に直接作用するエピトープ予測のための表面中心学習フレームワークである SurfBind を紹介します。 SurfBind は、パッチレベルの表面モデリング、バインダーを意識したクロスアテンション、および階層的な粗いから細かい予測パラダイムを備えた Transformer ベースのアーキテクチャを通じて、幾何学的および物理化学的なキューを統合します。 SAbDab や DB5.5 などの困難なエピトープ同定ベンチマークに関する実験では、SurfBind が最先端のパフォーマンスと、目に見えない抗体や立体構造状態にわたる強力な一般化を達成していることが実証され、タンパク質間相互作用の重要なメカニズムを理解するための相互作用を意識した表面モデリングの価値が強調されています。

原文 (English)

Deciphering Fingerprints of 3D Molecular Surfaces for Accurate Epitope Prediction

Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope prediction. However, existing methods rely on sequences or backbone structures and struggle to capture discontinuous, surface-driven epitopes. This study presents SurfBind, a surface-centric learning framework for epitope prediction that operates directly on molecular surface representations. SurfBind integrates geometric and physicochemical cues through a Transformer-based architecture with patch-level surface modeling, binder-aware cross-attention, and a hierarchical coarse-to-fine prediction paradigm. Experiments on challenging epitope identification benchmarks, including SAbDab and DB5.5, demonstrate that SurfBind achieves state-of-the-art performance and strong generalization across unseen antibodies and conformational states, highlighting the value of interaction-aware surface modeling for understanding the crucial mechanisms of protein-protein interactions.

13:00 JSTエージェントロボティクス

高度な航空モビリティ回廊を通じた自律交通の分散調整

Advanced Air Mobility (AAM) トラフィック用の専用通路の使用は、AAM トラフィックを既存の空域運用に統合する最も一般的に提案されている経路の 1 つです。これまでの研究のほとんどは、AAM 回廊のネットワークの設計と回廊内の航空機の衝突解決に焦点を当てていました。また、コリドーベースの運用は、実装の観点からは魅力的ではあるものの、特に集中型の交通管理がない場合には非効率的である可能性があるとも一般に考えられています。この論文では、この考えに反して、自律型航空機が分散設定で回廊の流れに自己組織化することを学習することが可能であることを示します。ここでは、固定翼航空機が (1) 出口後にメーターが表示される単一の通路、(2) 一連の 2 つの連続する通路、および (3) 2 つに分かれる通路を安全かつ効率的に通過する必要があるシナリオを使用してアプローチを説明します。ローカル情報のみを含む分散型設定では、航空機は 94% 以上の確率で回廊の境界に準拠し、比較的効率的に目的地に到達できることがわかりました。さらに、最小分離の違反に対処するための戦術的介入が必要になるのは、低密度および中密度の設定ではまれです。ただし、このような戦術的介入がより頻繁に必要になるのは、交通密度が高い場合に限られます。

原文 (English)

Decentralized Coordination of Autonomous Traffic Through Advanced Air Mobility Corridors

The use of dedicated corridors for Advanced Air Mobility (AAM) traffic is one of the most commonly proposed pathways to integrating them into existing airspace operations. Most prior research has focused on the design of networks of AAM corridors and conflict resolution for aircraft within corridors. It is also generally believed that while attractive from an implementation perspective, corridor-based operations may be inefficient, especially in the absence of centralized traffic management. In this paper, we show that contrary to this belief, it is possible for autonomous aircraft to learn to self-organize into corridor flows in decentralized settings. We illustrate our approach using scenarios in which fixed-wing aircraft need to safely and efficiently traverse (1) a single corridor with metering after the exit, (2) a sequence of two consecutive corridors, and (3) a corridor that splits into two. We find that in decentralized settings with only local information, the aircraft are able to conform to the corridor boundaries more than 94% of the time and reach their goal in a relatively efficient manner. Furthermore, tactical interventions to handle violations of the separation minimum are needed only infrequently in low- and medium-density settings. However, such tactical interventions become more frequently necessary only when traffic density is high.

13:00 JST研究/論文

測定可能な多数派

この論文は、いわゆる $\textit{社会的決定フレーム}$ を使用した、有限の選挙人における厳密多数派推論を研究します。つまり、厳密多数派を形成すると評価される投票ブロックとして解釈される、連合の著名な家族を備えた有限の有権者の集合です。定性的多数決の一貫性基準が特定され、有限加法尺度による厳密多数派の表現可能性の正確な特徴付けが得られることが示されています。さらに、厳密な多数決について推論するための最小限の自然な論理が健全で完全であることが示されています。これらの発展は、集合の有限族における一貫性のなさに関する、関連する組み合わせの問題の検討を動機づけます。部分的な結果と推測が示されています。最後に、この論文の結果は、Patrick Suppes による弱い質的確率構造に対する古典的な表現定理を修正し、社会的決定フレームの通常の厳密多数決に対するメイ型の特徴付けを確立するために適用されます。

原文 (English)

The Measurable Majority

This paper studies strict majority reasoning in finite electorates using so-called $\textit{social decision frames}$: finite sets of voters equipped with distinguished families of coalitions interpreted as those voting blocs evaluated to form a strict majority. A coherence criterion for qualitative majority judgments is identified and shown to give an exact characterization for representability of strict majorities by finitely additive measures. In addition, a minimal natural logic for reasoning about strict majorities is shown to be sound and complete. These developments motivate examination of associated combinatorial questions concerning incoherence in finite families of sets; partial results and a conjecture are given. Finally, the results of this paper are applied to correct a classical representation theorem for weak qualitative probability structures due to Patrick Suppes and to establish a May-type characterization for ordinary strict majority rule for social decision frames.

13:00 JST研究/論文

ニューラルネットワークの安全性保証は安全ですか?信頼できる堅牢性認定を計算する方法

AI の安全性における主な課題は、敵対的な例、つまりニューラル ネットワーク (NN) の誤分類を引き起こすわずかに歪んだ入力の存在です。この問題を軽減するために、最近の研究は、特定の入力に対して、ネットワークの予測を崩すことなく入力が受ける可能性のある最大の歪みを決定する堅牢性認証の計算に焦点を当てています。ロバストネス認定は、軸に沿って整列した超長方形 (多次元間隔) として解釈できます。既存のアプローチのほとんどは、認証のボリュームを最大化することに焦点を当てていますが、最近の扱いにくい結果により、ボリュームに最適な認証を妥当な時間内に計算することが不可能になっています。アポセムの尺度を導入し、NN 検証者 (オラクル) に対する線形呼び出し回数でアポセムに最適な認証を計算する方法を示します。入力ドメインの直径。さらに、たとえオラクルのコストを無視したとしても、ボリュームを最適化するオラクルベースのアルゴリズムは実現できないことを証明しました。また、二重認定 (クラスのすべてのインスタンスを含む間隔) を導入し、堅牢性認定に最低限の上限を提供します。さらに、標準 MNIST および Fashion MNIST ベンチマークで評価する ParallelepipedoNN システムを紹介します。同じデータセットに対する既存の研究との予備的な比較では、全体で少なくとも 2 倍の改善が明らかになりました。最小のエッジの長さ。

原文 (English)

Are Safety Guarantees in Neural Networks Safe? How to Compute Trustworthy Robustness Certifications

A primary challenge in AI safety is the existence of adversarial examples -- slightly distorted inputs that cause a neural network (NN) to misclassify. To mitigate this problem, recent research focuses on the computation of robustness certifications, which, for a given input, determine the largest distortion the input may receive without breaking the network's prediction. Robustness certifications can be interpreted as an axis-aligned hyper-rectangle (multi-dimensional intervals). Most existing approaches focus on maximizing the certification's volume, but recent intractability results prohibit the computation of volume-optimal certifications in reasonable time. We introduce the apothem measure and show how to compute apothem-optimal certifications in a linear number of calls to a NN verifier (oracle) w.r.t. the input domain's diameter. Moreover, we prove that we cannot have a volume-optimal, oracle-based algorithm, even if we discard the oracle costs. Also, we introduce dual certifications -- an interval including all instances of a class -- thus providing apothem-minimum upper bounds to a robustness certification. Further, we present the ParallelepipedoNN system, which we evaluate on the standard MNIST and Fashion MNIST benchmarks. A preliminary comparison with existing work on the same datasets reveals at least two-fold improvement w.r.t. the minimum edge length.

13:00 JST研究/論文

MGI: メンバー vs 生成された推論

生成モデルが人間が作成したコンテンツと区別できないサンプルを生成することが増えているため、特にモデルがトレーニング データを記憶して再現する場合、特定のデータ ポイントがモデルの自然なトレーニング セットの一部であるのか、それともモデル自体によって生成されたのかを判断することが困難になります。この課題をメンバー対生成推論 (MGI) として形式化します。つまり、サンプルとターゲット生成モデルが与えられた場合、サンプルが真のトレーニング メンバーであるか、そのモデルの生成された出力であるかを推論します。画像生成に焦点を当て、既存のメンバーシップ推論手法では生成されたサンプルをトレーニング メンバーとして体系的に誤分類する一方、属性ベースの手法では真のメンバーを生成されたメンバーとして誤分類することが多いことを示します。この失敗は、どちらのアプローチも、トレーニング例とモデル自体の出力に対して同様に上昇する尤度関連の信号に依存しているために発生します。 MGI に対処するために、生成モデルのオートエンコーダーと潜在ジェネレーターからの相補信号を組み合わせてトレーニング メンバーと生成されたサンプルを区別する 3 段階の方法であるデータ サーキット ブレーカー (DCB) を提案します。画像自己回帰モデルや拡散モデルを含む複数の生成モデルにわたって、DCB はメンバーシップ推論と帰属手法の欠点に一貫して対処し、モデルがトレーニング サンプルのほぼ重複を再現する場合でも効果を維持し、生成されたデータで新しいモデルをトレーニングする困難なモデルの派生設定に一般化します。

原文 (English)

MGI: Member vs Generated Inference

As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data. We formalize this challenge as Member vs Generated Inference (MGI): given a sample and a target generative model, infer whether the sample is a true training member or a generated output of that model. Focusing on image generation, we show that existing membership inference methods systematically misclassify generated samples as training members, while attribution-based methods often misclassify true members as generated. This failure arises because both approaches rely on likelihood-related signals that are similarly elevated for training examples and for the model's own outputs. To address MGI, we propose Data Circuit Breaker (DCB), a three-stage method that combines complementary signals from a generative model's autoencoder and latent generator to distinguish training members from generated samples. Across multiple generative models, including image autoregressive and diffusion models, DCB consistently addresses the shortcomings of membership inference and attribution methods, remains effective even when models reproduce near-duplicates of training samples, and generalizes to challenging model derivative settings in which new models are trained on generated data.

13:00 JST研究/論文

JupOtter: Jupyter Notebook でのセルレベルのバグ検出

Jupyter Notebook は、多くのドメイン、特に Python ベースのデータ サイエンスや科学技術コンピューティングで使用されるコーディング環境として人気が高まっています。ノートブックはもともとプロトタイピングやインタラクティブな探索に使用されていましたが、より複雑なプログラムの開発に使用されることが増え、GitHub などのプラットフォーム上でバグのあるノートブックが急増しています。この傾向に対処するために、Jupyter Notebook 専用に設計されたバグ検出システム JupOtter を紹介します。 JupOtter は 3 つの新しい貢献を特徴としています: (1) セル構造を保持するノートブック固有のトークン化戦略、(2) セルレベルのバグ予測技術、(3) きめ細かいセルレベルのバグ検出のために注釈が付けられた 21,000 を超えるノートブックを含む新しいラベル付きデータセット OtterDataset。 JupOtter は、3 つの評価データセットのうち 2 つで静的アナライザーや大規模言語モデルを上回るセルレベルのバグ検出 F1 スコアを達成しました。

原文 (English)

JupOtter: Cell-Level Bug Detection in Jupyter Notebooks

Jupyter Notebooks are an increasingly popular coding environment used across many domains, especially in Python-based data science and scientific computing. Originally used for prototyping and interactive exploration, notebooks are increasingly used to develop more complex programs, leading to a rapid rise in buggy notebooks on platforms like GitHub. To address this trend, we present JupOtter, a bug detection system designed specifically for Jupyter Notebooks. JupOtter features three novel contributions: (1) a notebook-specific tokenization strategy that preserves cell structure, (2) a cell-level bug prediction technique, and (3) a new labeled dataset, OtterDataset, containing over 21,000 notebooks annotated for fine-grained cell-level bug detection. JupOtter achieves cell-level bug detection F1 scores that surpass static analyzers and large language models in two out of three evaluation datasets.

13:00 JST研究/論文

対比のないコントラスト画像変換を使用した非造影 CT スキャンからの心腔セグメンテーションの期待と課題: 実現可能性研究

目的: 対比のない対比画像変換と深層学習ベースのセグメンテーションを使用した、非造影 CT スキャンからの心腔セグメンテーションの実現可能性と課題を評価すること。アプローチ: 私たちは、コントラスト CT スキャンから非造影 CT を合成するために、分離されたコントラスト学習 (DCL) 損失を備えたコントラスト不対変換 (CUT) ネットワークを利用するフレームワークである ChameleonNet を開発しました。コントラストスキャンからの 4 つの心腔 (左心房 (LA)、左心室 (LV)、右心房 (RA)、および右心室 (RV)) のアノテーションを使用して、合成された非造影画像上でハウスドルフ距離損失強化 nnU-Net をトレーニングしました。変換モデルは、35,538 枚の造影 CT スライスと 37,197 枚の非造影 CT スライスを使用してトレーニングされました。セグメンテーション モデルは、292 個の合成非コントラスト スキャンを使用してトレーニングされました。パフォーマンスは、36 個の合成非造影スキャンで Dice 類似性係数 (DSC) と 95 番ハウスドルフ距離 (HD95) を使用して評価され、36 個の実際の非造影 CT スキャンでの体積の一致は、ピアソン相関、平均絶対パーセント誤差 (MAPE)、および平均パーセント誤差 (MPE) を使用して評価されました。結果: セグメンテーション モデルは、合成上で 0.94 (0.01)、0.91 (0.04)、0.92 (0.03)、0.93 (0.02) の DSC、および 3.63 (1.49)、5.74 (4.08)、5.18 (1.77)、5.51 (3.21) mm の HD95 を達成しました。それぞれ LA、LV、RA、RV の非造影画像。実際の非造影 CT スキャンでは、ピアソン相関は 0.93、0.82、0.87、0.89 (すべて p<0.001) で、MAPE の範囲は 9.22% ~ 20.79%、MPE の範囲は -12.52% ~ 4.67% でした。結論: ChameleonNet は、手動による非造影剤のアノテーションを使用せずに、非造影 CT からの心腔セグメンテーションの実現可能性を実証しました。ただし、特に LV と RV の体積誤差は、臨床使用前にさらなる改良と検証が必要であることを示しています。

原文 (English)

Promise and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired image translation: a feasibility study

Purpose: To evaluate the feasibility and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired image translation and deep learning-based segmentation. Approach: We developed ChameleonNet, a framework utilizing the Contrastive Unpaired Translation (CUT) network with decoupled contrastive learning (DCL) loss to synthesize non-contrast CT from contrast CT scans. Using annotations of four heart chambers (left atrium (LA), left ventricle (LV), right atrium (RA), and right ventricle (RV)) from contrast scans, we trained a Hausdorff distance loss-enhanced nnU-Net on synthesized non-contrast images. The translation model was trained with 35,538 contrast-enhanced and 37,197 non-contrast CT slices. The segmentation model was trained with 292 synthesized non-contrast scans. Performance was evaluated using Dice similarity coefficient (DSC) and 95th Hausdorff distance (HD95) on 36 synthesized non-contrast scans, and volume agreement on 36 real non-contrast CT scans was assessed using Pearson correlation, mean absolute percentage error (MAPE), and mean percentage error (MPE). Results: The segmentation model achieved DSC of 0.94 (0.01), 0.91 (0.04), 0.92 (0.03), 0.93 (0.02), and HD95 of 3.63 (1.49), 5.74 (4.08), 5.18 (1.77), 5.51 (3.21) mm on synthesized non-contrast images for LA, LV, RA, and RV, respectively. On real non-contrast CT scans, Pearson correlations were 0.93, 0.82, 0.87, and 0.89 (all p<0.001), with MAPE ranging from 9.22% to 20.79%, and MPE ranging from -12.52% to 4.67%. Conclusions: ChameleonNet demonstrated feasibility for heart chamber segmentation from non-contrast CT without manual non-contrast annotations. However, volume errors, particularly for LV and RV, indicate that further refinement and validation are needed before clinical use.

13:00 JSTLLM/生成AI

1 年後...被害は続いていますが、私たちも同じです!

汎用大規模言語モデル (LLM) は、メンタルヘルス関連の会話にますます使用されていますが、安全対策は依然として不十分であり、臨床症状全体で一貫性がありません。この研究では、8 次元の危害分類と多次元の評価フレームワークを導入し、4 つの敵対的攻撃のバリアントを使用して、16 の DSM-5 条件にわたる 6 つの独自の LLM を評価します。その結果、安全策は自殺と自傷行為に対してのみ確実に有効であり、摂食障害、物質使用障害、大うつ病性障害などの疾患では失敗率が最大 100% であることが示されています。私たちは、これらの LLM の倫理的な設計と展開には、臨床状態全体にわたって明確に定義された危害カテゴリーと、それに応じた安全措置の実装が必要であると主張します。このような保護措置が講じられるまで、これらのモデルは脆弱な人々に重大なリスクをもたらすため、教育現場への統合の増加が特に懸念されます。

原文 (English)

One Year Later...The Harms Persist, But So Do We!

General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain inadequate and inconsistent across clinical conditions. This study evaluates six proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100%. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into educational settings a particularly concerning.

13:00 JSTLLM/生成AI画像/動画生成

注意: マルチモーダル LLM のトポロジー表現の調整

表現の調整は、外部ビジョン エンコーダの内部表現に合わせて内部表現を正規化することにより、マルチモーダル大規模言語モデル (MLLM) を改善する効果的なアプローチとして浮上しました。ただし、既存の方法は通常、言語バックボーンの固定層を調整し、Transformer モデルのきめの細かい構造を見落としています。この研究では、個々のアテンション ヘッドのレベルでクロスモーダル アライメントを強制する方法である Head-Wise Representation Alignment (HeRA) を提案します。私たちのアプローチはプラトニック表現仮説に基づいており、モダリティ全体で表現の位相構造 (つまり、ローカルな近傍関係) を保存することに焦点を当てています。相互 K 最近傍 (MKNN) アライメント メトリックに従って、ローカル構造を照合するための微分可能な代理として機能する対照的な目的を導入します。 HeRA は、マルチモーダル トレーニング中にこの目標を、MKNN メトリクスに従ったアライメント スコアによって選択された LLM 内の特定のアテンション ヘッドに適用します。直観に反しますが、最も整列していないヘッドを整列させると最大の利益が得られることがわかります。複数の MLLM と 18 のベンチマークにわたる広範な評価により、HeRA が視覚を中心とした困難なタスクのパフォーマンスを一貫して向上させ、言語の事前知識への過度の依存を自然に抑制することで幻視に対する効果的な規則化装置として機能することが実証されました。私たちのコードは公開されています。

原文 (English)

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder. However, existing methods typically align a fixed layer of the language backbone, overlooking the fine-grained structure of Transformer models. In this work, we propose Head-Wise Representation Alignment (HeRA), a method that enforces cross-modal alignment at the level of individual attention heads. Our approach is grounded in the Platonic Representation Hypothesis, focusing on preserving the topological structure of representations (i.e., their local neighborhood relationships) across modalities. Following the Mutual K-Nearest Neighbor (MKNN) alignment metric, we introduce a contrastive objective that acts as a differentiable proxy for matching local structures. HeRA applies this objective during multimodal training to specific attention heads in the LLM, selected by their alignment score according to the MKNN metric. Counterintuitively, we find that aligning the least aligned heads yields the largest gains. Extensive evaluations across multiple MLLMs and 18 benchmarks demonstrate that HeRA consistently improves performance on challenging vision-centric tasks and serves as an effective regularizer against visual hallucinations by naturally curbing the over-reliance on linguistic priors. Our code is publicly released.

13:00 JST画像/動画生成

E-MRL: 信頼性の高い 3D 腫瘍解析のためのクロスビューで調整された証拠主導型マルチモーダル強化学習

視覚言語モデル (VLM) は容積測定医療レポートの生成に大きな期待を寄せていますが、幻視や 3D CT データの根拠の欠如に悩まされることがよくあります。現在の教師あり微調整 (SFT) および強化学習 (RL) 戦略は、通常、テキストの忠実度のみを最適化し、本質的に、本物の視覚認識ではなく言語の事前知識に基づいて得られる正しい診断に報います。これに対処するために、我々は、生成プロセスを「診断 - 位置特定 - 検証」のマルコフ決定プロセスとして定式化する信頼性の高い RL 推論フレームワークである、クロスビュー調整された証拠駆動型マルチモーダル強化学習 (Evidence-MRL、E-MRL と表記) を提案します。標準的なアプローチとは異なり、私たちのモデルは、検証可能な視覚的証拠に基づいて、全体的な診断レポートと並んで「重要な証拠スライス」を特定するように明示的にトレーニングされています。重要なのは、新しいクロスビュー一貫性報酬を導入することです。これは、ゴールデンスタンダード レポートと、選択されたキー スライスのローカル視覚的再クエリの間の意味論的な整合性を検証し、正しくローカライズされた推論に対して追加の報酬を提供します。大規模な 3D CT 腫瘍データセットの実験では、E-MRL が SFT および RL ベースラインと比較して幻覚を大幅に軽減し、診断精度を向上させ、視覚に基づいた腫瘍分析のための臨床的に解釈可能なソリューションを提供することが実証されました。

原文 (English)

E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucinations and a lack of grounding in 3D CT data. Current Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) strategies typically optimize text fidelity alone, essentially rewarding correct diagnoses derived from language priors rather than genuine visual perception. To address this, we propose cross-view aligned Evidence-driven Multimodal Reinforcement Learning (Evidence-MRL, noted as E-MRL), a reliable RL reasoning framework that formulates the generation process as a Markov Decision Process of "diagnosis-localization-verification". Unlike standard approaches, our model is explicitly trained to identify a "key evidence slice" alongside the global diagnostic report, grounding its findings in verifiable visual evidence. Crucially, we introduce a novel cross-view consistency reward, which validates the semantic alignment between the golden-standard report and a local visual re-query of the selected key slice, providing additional rewards for correctly-localized reasoning. Experiments on large-scale 3D CT tumor datasets demonstrate that E-MRL significantly reduces hallucinations and improves diagnostic accuracy compared to SFT and RL baselines, offering a clinically interpretable solution for visually-grounded and tumor analysis.

13:00 JSTLLM/生成AI画像/動画生成

教授: 視覚言語モデルの複数教師による教師なし即時蒸留

プロンプト蒸留は、ラベルのないドメイン イメージ上で教師の予測を照合することにより、CLIP などの大規模ビジョン言語モデル (VLM) を軽量の学生モデルに圧縮します。 PromptKD (CVPR 2024) は、PromptSRC で微調整された 1 人の ViT-L/14 教師と ViT-B/16 の生徒でこのパラダイムを確立しました。私たちは、固定の 2 教師アンサンブルから派生したマルチ教師拡張機能である TheProfessor を提案します。つまり、ドメイン微調整された PromptSRC ViT-L/14 教師と、ロジットがデータセットごとに事前計算されるゼロショット EVA-CLIP-L/14 教師です。 Caltech-101、DTD、UCF101、および EuroSAT の 4 つの基礎から新規のデータセットに基づいて、単一教師 PromptKD、等確率アンサンブル、および信頼度加重アンサンブルを評価します。 12 回のシングルシード スイープでは、信頼度加重アンサンブルにより平均 HM が 87.52 から 89.28 (+1.77 ポイント) に改善され、一方、均等平均化により平均 HM が 88.88 (+1.37 ポイント) に改善されました。利得はデータセットに依存します。Caltech-101 では無視できる程度 (信頼度重み付けの +0.16 HM)、UCF101 では控えめ (+0.62)、ドメインシフト EuroSAT では最大 (+5.78) です。これらの結果は、以前のカリフォルニア工科大学のみの分析を更新し、複数の教師による即時蒸留が、ドメイン シフトの下で 2 番目の教師が補完的な監督に貢献する場合に最も有用であることを示しています。

原文 (English)

The Professor: Multi-Teacher Unsupervised Prompt Distillation for Vision-Language Models

Prompt distillation compresses large vision-language models (VLMs) such as CLIP into lightweight student models by matching teacher predictions on unlabeled domain images. PromptKD (CVPR 2024) established this paradigm with a single PromptSRC-finetuned ViT-L/14 teacher and a ViT-B/16 student. We propose TheProfessor, a multi-teacher extension that distills from a fixed two-teacher ensemble: a domain-finetuned PromptSRC ViT-L/14 teacher and a zero-shot EVA-CLIP-L/14 teacher whose logits are pre-computed per dataset. We evaluate single-teacher PromptKD, equal-probability ensembling, and confidence-weighted ensembling on four base-to-novel datasets: Caltech-101, DTD, UCF101, and EuroSAT. In a 12-run single-seed sweep, confidence-weighted ensembling improves average HM from 87.52 to 89.28 (+1.77 points), while equal averaging improves average HM to 88.88 (+1.37 points). Gains are dataset dependent: they are negligible on Caltech-101 (+0.16 HM for confidence weighting), modest on UCF101 (+0.62), and largest on domain-shifted EuroSAT (+5.78). These results update our earlier Caltech-only analysis and show that multi-teacher prompt distillation is most useful when the second teacher contributes complementary supervision under domain shift.

13:00 JST研究/論文

ARIA: 条件付き拡散蒸留のための適応的な領域ベースの重要度割り当て

条件付き拡散モデルを抽出することは、条件付け入力間の整合性を維持しながら、大柄な教師の行動を小柄な生徒に伝えることを目的としています。認識タスクとは異なり、条件付き拡散における知識の蒸留は、予測されたノイズが条件付け信号に強く依存するため、トレーニング分布を超えて知識を伝達するのに苦労することがよくあります。その結果、効果的な蒸留には広い調整スペースを探索する必要があります。実際の設定では、これが大きなボトルネックになります。ペアの画像条件データは制限されている可能性があり、利用可能なすべての条件に対して合成画像を生成することは多くの場合計算的に実行不可能であり、テキスト プロンプトなどの条件のプールは非常に大きくなる可能性があります。最近の研究では、トレーニング中に条件を切り替え、蒸留の目的を変更せずに学生をより広いコンディショニング空間にさらすことで、この問題に対処しています。しかし、これは補足的な疑問を引き起こします。大規模な条件付けコーパスが利用可能になったら、トレーニングの労力をどのように割り当てるべきでしょうか?この研究では、調整空間の粗い領域全体にトレーニングの労力を適応的に割り当てるフレームワークである ARIA を紹介します。 ARIA は、地域レベルで教師と生徒の不一致のオンライン推定値を維持することにより、元の蒸留の目的を維持しながら、不一致が持続する箇所の更新に焦点を当てます。経験的に、ARIA はほとんどのアーキテクチャと設定において RC よりも改善しており、目に見えず過小評価されている体制で最も明らかな改善が観察されます。また、提案された追跡メカニズムが、境界のある分散とドリフトの仮定の下でトレーニング中に進化する不一致に従うことを示す理論的分析も提供します。

原文 (English)

ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation

Distilling conditional diffusion models aims to transfer the behavior of a large teacher to a smaller student while preserving alignment across conditioning inputs. Unlike recognition tasks, knowledge distillation in conditional diffusion often struggles to transfer knowledge beyond the training distribution, since the predicted noise strongly depends on the conditioning signal. As a result, effective distillation requires exploring a large conditioning space. In practical settings, this creates a major bottleneck. Paired image-condition data may be limited, and generating synthetic images for every available condition is often computationally infeasible, while the pool of conditions, such as text prompts, can be extremely large. Recent work addresses this issue by switching conditions during training, exposing the student to a broader conditioning space without changing the distillation objective. Yet this raises a complementary question: once a large conditioning corpus is available, how should the training effort be allocated? In this work, we introduce ARIA, a framework that adaptively allocates training effort across coarse regions of the conditioning space. By maintaining online estimates of teacher-student discrepancy at the region level, ARIA focuses updates where misalignment persists while preserving the original distillation objective. Empirically, ARIA improves over RC across most architectures and settings, with the clearest gains observed in unseen and underrepresented regimes. We also provide a theoretical analysis showing that the proposed tracking mechanism follows the evolving discrepancy during training under bounded variance and drift assumptions.

13:00 JST研究/論文

壊滅的な組成生成: バニラ拡散モデルが外挿できない理由

構成生成のタスクには、可能な条件のサブセットのみでトレーニングされた条件付き生成モデルを使用して、ソース分布の幾何学的組み合わせなど、構成的に定義されたターゲット分布からサンプルを生成することが含まれます。この研究では、このタスクはバニラの条件付き拡散モデルでは実行不可能であることが多いと主張します。つまり、特定の動機付けられた設定では、ターゲット分布からサンプルを効率的に生成できる推論時間手法は存在しないと推測しています。この考えは、理論に基づいた一般化議論と、合成データと現実データの両方に対する慎重に設計された実験によって裏付けられています。特に、ファインマン・カック補正などの最近の手法は推論時間近似誤差を低減しますが、我々の結果は、ターゲット分布がソースに対して分布から外れている場合、スコア推定誤差がパフォーマンスにさらに壊滅的な影響を与えることを示しており、このタスクに対する別のアプローチの必要性を強調しています。

原文 (English)

Catastrophic Compositional Generation: Why Vanilla Diffusion Models Fail to Extrapolate

The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to produce samples from compositionally-defined target distributions such as a geometric combination of the source distributions. In this work, we argue that this task is often infeasible for vanilla conditional diffusion models: we conjecture that no inference-time technique can efficiently produce samples from the target distribution in certain well-motivated settings. This idea is supported by theory-guided generalization arguments and carefully-designed experiments on both synthetic and realistic data. In particular, while recent methods such as Feynman-Kac correction reduce inference-time approximation error, our results show that score estimation error has a more catastrophic effect on performance when the target distribution is out-of-distribution with respect to the sources, highlighting the need for a different approach to this task.

13:00 JSTLLM/生成AIエージェント

取得メトリクスが誤解を招く場合: 長期的なツール使用エージェントにおけるポリシーシグナルの測定

完全一致検索リコールは、検索者が有用なポリシー コンテキストを下流の意思決定モデルに提供するかどうかの代理としてよく使用されます。 Qwen2.5-3B/7B 分類子を使用して、tau-bench でこのプロキシのアクション前ポリシー分類をテストします。ゴールド ポリシー コンディショニングの下で​​は、コンパクトな構造状態により、調整後に生の軌道に比べてマクロ F1 が 0.13 ~ 0.17 改善されます。次に、ベンチマークに指定されたポリシー条項を、意思決定時のコンテキストから取得した最上位の条項に置き換えます。正確な支配条項がランク 1 で取得されるのは航空会社の州の 7% のみですが、プライマリ 3B 分類子は、取得された条項ではマクロ F1 が 0.58 であるのに対し、ゴールド条項では 0.60 になります (デルタ = -0.02、タスク クラスター 95% CI [-0.23,+0.21])。ポリシーが一致しないコントロールとポリシーなしのコントロールのスコアは 0.32 と 0.21 です。この構成では、取得句とゴールド句の間にマクロ F1 の違いは検出されませんが、非劣性を確立するには間隔が広すぎるままです。同じ定性的パターンが 2 番目のリトリーバーと 7B で現れますが、微調整構成によって異なります。これらの結果は、このベンチマーク設定では完全一致条項のリコールが下流のポリシーの有用性を過小評価する可能性があり、リコールのみではなく分類ループで取得されたポリシーを使用した評価を動機付ける可能性があることを示しています。

原文 (English)

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test this proxy for pre-action policy classification in tau-bench using Qwen2.5-3B/7B classifiers. Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by 0.13-0.17 after tuning. We then replace the benchmark-designated policy clause with the top-ranked clause retrieved from decision-time context. Although the exact governing clause is retrieved at rank 1 for only 7% of airline states, the primary 3B classifier obtains macro-F1 0.58 with retrieved clauses versus 0.60 with gold clauses (Delta=-0.02, task-cluster 95% CI [-0.23,+0.21]); mismatched-policy and no-policy controls score 0.32 and 0.21. We do not detect a macro-F1 difference between retrieved and gold clauses in this configuration, although the interval remains too wide to establish non-inferiority. The same qualitative pattern appears with a second retriever and at 7B, while varying across fine-tuning configurations. These results indicate that exact-match clause recall can underestimate downstream policy utility in this benchmark setting, motivating evaluation with retrieved policies in the classification loop rather than recall alone.

13:00 JSTLLM/生成AI

ウェアハウス SLAM スループット制御のためのオフライン強化学習

倉庫フルフィルメント環境における SLAM スループット制御を最適化するためのオフライン強化学習 (RL) フレームワークを紹介します。 SLAM (スキャン/ラベル/適用/マニフェスト) スループットは、システムの混雑と運用効率に直接影響します。当社の RL ベースの制御アプローチは、スロットリング動作のインテリジェントな調整を通じて、スループットの最大化とダウンストリームの安定性のバランスを適応的に調整する SLAM スループット設定を動的に推奨します。履歴に基づいた状態表現、遅延衝撃制御のためのアクション空間の抽象化、および上流と下流の両方の運用メトリクスを取得する報酬関数が含まれています。私たちのアプローチはアルゴリズムに依存せず、統一されたアーキテクチャの下で複数のオフライン RL メソッドの統合を可能にします。私たちは 3 つの最先端のオフライン RL アルゴリズムを使用してフレームワークをインスタンス化し、大規模な倉庫からの匿名化された過去の操作ログを使用してモデルをオフラインでトレーニングしました。ポリシーのパフォーマンスは、包括的な複数の方法戦略を使用して評価されます。これらには、回帰モデルや長期適合 Q 評価 (FQE) による即時報酬推定、モデルベースのディープ クープマン ダイナミクス評価などのモデルフリーのアプローチが含まれます。実証結果によると、CQL ポリシーは常に代替案より優れたパフォーマンスを示し、システムの健全性が 22.97% 改善され、平均スロットル期間が 3.18% 短縮されました。これらの調査結果は、安全かつスケーラブルな倉庫スループット制御の最適化に対するオフライン RL の可能性を示しています。

原文 (English)

Offline Reinforcement Learning for Warehouse SLAM Throughput Control

We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment. SLAM (Scan/Label/Apply/Manifest) throughput directly influences system congestion and operational efficiency. Our RL-based control approach dynamically recommends SLAM throughput settings that adaptively balance throughput maximization with downstream stability through intelligent adjustment of throttling behavior. We include a history-informed state representation, action space abstraction for delayed-impact control, and a reward function that captures both upstream and downstream operational metrics. Our approach is algorithm-agnostic, enabling integration of multiple offline RL methods under a unified architecture. We instantiate our framework with three state-of-the-art offline RL algorithms, and trained the models offline using de-identified historical operational logs from a large-scale warehouse. Policy performance is evaluated using a comprehensive multi-method strategy. These include model-free approaches including immediate reward estimation via regression models and long-horizon Fitted Q Evaluation (FQE), as well as model-based Deep Koopman dynamics evaluation. Empirical results reveal that the CQL policy consistently outperforms alternatives, improving system health by 22.97% and reducing average throttling duration by 3.18%. These findings demonstrate the potential of offline RL for safe and scalable warehouse throughput control optimization.

13:00 JST研究/論文

Maestro Order: モデルに依存しないオーケストレーション ハーネス

有能なモデルの 1 回のフォワード パスは、高速かつ流暢で信頼性の低い問題解決手段です。正しい場合は有用であることが多く、間違っている場合は危険であることがよくあります。言語モデルでは、このような確実なエラーは幻覚として知られています。我々は、4 つの構造プリミティブ (分解、アンサンブル、検証、再帰) に従って信頼性の低いソルバーを構成することで、信頼性の低いソルバーを信頼できる問題解決システムに変える、モデルに依存しないオーケストレーション ハーネスである Maestro Order と、コンピューティングをどこに費やすかを決定する予算を意識したコントローラーを紹介します。このハーネスは、あらゆるモデルを統一インターフェイスの背後にあるブラック ボックス ベース ソルバーとして扱い、オンラインで識別が測定される検証アンサンブルを階層化し、単位コストあたりの限界信頼性が最も高いステージに検証と投票を割り当てます。私たちは、アーキテクチャ、メッセージと状態のスキーマ、コントローラー アルゴリズム、およびそれを決定論的で監視可能でフォールト トレラントにするエンジニアリングを提供します。次に、評価方法 (固定コストでの信頼性、カバレッジ、キャリブレーション、およびアブレーション) を指定し、パラメーター化されたソルバー/検証モデルに対するハーネスの忠実なモンテカルロ シミュレーションの結果を報告します。シミュレーションは、予測された法則を定量的に再現します。検証は信頼性を幾何学的に増幅させます (例: 2 つのゲートで $0.55\to0.98$、4 つのゲートで $\to0.999$)、投票は偶然以上に役立ち、共有エラーによって制限されます。また、予算を意識したコントローラーは、各レジームで最も安価なメカニズムを選択することにより、投票単独のコストのほんの一部で目標の信頼性に到達します。最後に、障害モード (検証ゲーム、相関エラー、分解エラーの複合化) と具体的なガイダンスで終わります。つまり、堅牢なチェッカーを構築し、ソルバーを多様化し、情報がどこにあるのかをコントローラーに計算させます。

原文 (English)

Maestro Order: A Model-Agnostic Orchestration Harness

A single forward pass of a capable model is a fast, fluent, and unreliable problem-solver: it is right often enough to be useful and wrong often enough to be dangerous; in language models, such confident errors are known as hallucinations. We present Maestro Order, a model-agnostic orchestration harness that turns unreliable solvers into reliable problem-solving systems by composing them according to four structural primitives (decompose, ensemble, verify, and recurse) and a budget-aware controller that decides where to spend compute. The harness treats any model as a black-box base solver behind a uniform interface, layers a verifier ensemble whose discrimination is measured online, and allocates verification and voting to the stages with the highest marginal reliability per unit cost. We give the architecture, the message and state schema, the controller algorithm, and the engineering that makes it deterministic, observable, and fault-tolerant. We then specify an evaluation methodology (reliability at fixed cost, coverage, calibration, and ablations) and report results from a faithful Monte Carlo simulation of the harness over a parameterized solver/verifier model. The simulation reproduces the predicted laws quantitatively: verification amplifies reliability geometrically (e.g. $0.55\to0.98$ with two gates, $\to0.999$ with four), voting helps only above chance and is limited by shared errors, and a budget-aware controller reaches a target reliability at a small fraction of the cost of voting alone by selecting the cheapest mechanism for each regime. We close with failure modes (verifier gaming, correlated errors, and decomposition error compounding) and concrete guidance: build robust checkers, diversify solvers, and let the controller put compute where the information is.

13:00 JSTLLM/生成AI

構造的に忠実: 複数文書の要約のためのクレームにアンカーされた帰属

エンドツーエンドの大規模言語モデル (LLM) は、流暢な複数文書の要約を生成しますが、依然として幻覚を起こしやすく、また、LLM が提供する帰属は一般的に粗く (文書全体または文章全体)、事後的に生成されるため、各要約ステートメントを検証するのは困難です。私たちはモジュール式の抽出-選択-再書き込みパラダイムを再考し、その中間表現を帰属の単位として再構築します。我々は、(i) すべてのソース文書からトークンレベルの出自を持つ原子的なクレームを抽出し、(ii) ソース間の競合にフラグを立てながらドキュメント全体で同等のクレームをクラスター化し、(iii) サポートを意識した顕著なサブセットを選択し、(iv) 選択した内容を、すべての文が 1 つ以上のソース スパンにリンクするサポートチェックされたクレームに固定された要約に書き換える、クレームアンカー型マルチドキュメント要約フレームワークである CAMS を紹介します。コンテンツは実現される前にローカライズされるため、パイプラインは構造的には帰属指向であり、構造的には忠実性を重視します。つまり、サポートを意識した選択、制約付き書き換え、検証を使用して、事実の忠実性を保証するのではなく奨励しながら、構造的にきめの細かいマルチソースのトレーサビリティを維持します。私たちは、MultiNews で品質、忠実性、ローカリゼーションを評価し、DiverseSumm で競合処理を分析し、WCEP でゼロショット転送をテストします。この際、参考文献フリーの引用品質とゴールドアライメントのローカリゼーション精度を分離する 2 つのレジームプロトコルを使用します。さらに、選択や検証に決して使用されないサポートモデルで引用精度をテストする評価者分離監査を追加します。 CAMS は、要約の品質に関して強力なエンドツーエンドおよびスパン帰属ベースラインを照合すると同時に、忠実性と引用の精度を大幅に向上させ、複数情報源の帰属精度を約 3 分の 2 向上させ、制御可能な忠実性、つまりエンドツーエンド モデルが暗黙的に残しているカバレッジのトレードオフを明らかにします。

原文 (English)

Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization

End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify. We revisit the modular Extract--Select--Rewrite paradigm and recast its intermediate representation as the unit of attribution. We present CAMS, a Claim-Anchored Multi-document Summarization framework that (i) extracts atomic claims with token-level provenance from every source document, (ii) clusters equivalent claims across documents while flagging inter-source conflicts, (iii) selects a support-aware and salient subset, and (iv) rewrites the selection into a summary in which every sentence is anchored to a support-checked claim that links back to one or more source spans. Because content is localized before it is realized, the pipeline is attribution-oriented by construction and faithfulness-oriented by construction: it structurally preserves fine-grained, multi-source traceability while using support-aware selection, constrained rewriting, and verification to encourage, rather than guarantee, factual faithfulness. We evaluate quality, faithfulness, and localization on MultiNews, analyze conflict handling on DiverseSumm, and test zero-shot transfer on WCEP, using a two-regime protocol that separates reference-free citation quality from gold-aligned localization accuracy, and we add an evaluator-decoupled audit that tests citation precision with a support model never used for selection or verification. CAMS matches strong end-to-end and span-attribution baselines on summary quality while substantially improving faithfulness and citation precision, lifting multi-source attribution accuracy by roughly two-thirds, and exposing a controllable faithfulness--coverage trade-off that end-to-end models leave implicit.

13:00 JSTLLM/生成AIGPT / ChatGPT

RASC+: 臨床値セットオーサリングのための検索制約付き LLM 判定

臨床値セットは、品質測定、表現型解析、コホート構築、および臨床意思決定のサポートに使用される標準化された用語コードを定義します。最近導入された検索拡張セット補完 (RASC) ベンチマークは、直接ゼロショット大規模言語モデル (LLM) 生成がこのタスクにはあまり適していないことを示しました。臨床コード システムは大規模で、バージョン管理されており、言語モデルによって確実に記憶されません。私たちは、候補プールの構築が再現のために最適化され、制約付き LLM 判定器が候補の選択のために最適化される、段階別の代替案を研究します。完全な 3,744 値セットの RASC テスト分割では、語彙を意識した拡張とコード表示レスキュー検索を備えた Qwen3 ベースの検索により、候補プールの再現率が元の RASC 検索ベースラインの 0.553 から 0.730 に増加しました。保留された出版社層では、プール再現率は 0.655 です。高再現率プールだけでは十分ではありません。オリジナルの SAPBert クロスエンコーダをこの拡張されたプールに適用すると、フルテスト マクロ F1 は 0.287、ホールドアウト パブリッシャー マクロ F1 は 0.233 になります。ステージ 2 セレクターを同じプールに対するブラインド GPT-5 判定に置き換えると、フルテスト マクロ F1 は 0.549 に増加し、ホールドアウト パブリッシャー マクロ F1 は 0.533 に増加します。これらの結果は、取得制約付き LLM 判定により、返されるすべてのコードが監査可能な候補プールから取得されなければならないという安全制約を維持しながら、値セットの完了を大幅に向上できることを示しています。

原文 (English)

RASC+: Retrieval-Constrained LLM Adjudication for Clinical Value Set Authoring

Clinical value sets define the standardized terminology codes used in quality measurement, phenotyping, cohort construction, and clinical decision support. The recently introduced Retrieval-Augmented Set Completion (RASC) benchmark showed that direct zero-shot large language model (LLM) generation is poorly suited to this task: clinical code systems are large, version-controlled, and not reliably memorized by language models. We study a stage-wise alternative in which candidate-pool construction is optimized for recall and a constrained LLM adjudicator is optimized for candidate selection. On the full 3,744-value-set RASC test split, Qwen3-based retrieval with vocabulary-aware expansion and code-display rescue retrieval increases candidate-pool recall from the original RASC retrieval baseline of 0.553 to 0.730; on the held-out-publisher stratum, pool recall is 0.655. The higher-recall pool alone is not sufficient: applying the original SAPBert cross-encoder to this expanded pool gives full-test macro F1 of 0.287 and held-out-publisher macro F1 of 0.233. Replacing the stage-2 selector with blinded GPT-5 adjudication over the same pool increases full-test macro F1 to 0.549 and held-out-publisher macro F1 to 0.533. These results show that retrieval-constrained LLM adjudication can substantially improve value set completion while preserving the safety constraint that all returned codes must come from an auditable candidate pool.

13:00 JST研究/論文

トリガーの学習: 大型ハドロン衝突型加速器での強化学習

大型ハドロン衝突型加速器などの高スループットの科学施設は、帯域幅、遅延、ストレージに厳しい制約がある中で、リアルタイム イベント フィルタリング (\textit{triggering}) に依存しています。実際には、トリガー メニューはほとんどが静的で手動調整されており、検出器の状態、パイルアップ、および背景の構成が時間の経過とともにドリフトするため、最適ではなくなる可能性があります。オンラインしきい値調整を逐次的な意思決定の問題として捉えます。強化学習エージェントは、最近のレートと信号に敏感な機能のストリーミング サマリーを取り込み、許容範囲内でターゲット バックグラウンド レートを追跡しながら、信号効率を最大化するためにトリガーしきい値を更新します。 Group-Filtered Policy Optimization (GFPO) をストリーミング制御に適応させ、トレーニング中にバックグラウンド レートの実現可能性を強制する 2 つのバリアント (GFPO-F、GFPO-FR) を導入します。現実的な衝突型加速器の動作をエミュレートするベンチマークで、我々は 2 つの代表的なトリガーを研究します。パイルアップ変動に敏感な総横エネルギー ($H_{T}$) トリガーと、稀なまたは非標準のシグネチャの再構成損失に基づく異常検出 (AD) トリガーです。モンテカルロ ストリームでは、エージェントは許容範囲内時間間隔の割合を 48\% ($H_T$) および 28\% (AD) 増加させ、これらの許容範囲内間隔での信号効率の累積ゲインは最大 2\% になります。シミュレーションから \emph{real} 衝突データ (CMS Run 283408) に移行すると、同じエージェントは微調整を行わずに、ベースラインに対して 56\% ($H_T$) および 28\% (AD) の許容範囲内改善を達成し、両方のトリガーで信号効率がさらに向上しました。私たちの知る限り、これは実際の大型ハドロン衝突型加速器の衝突データに対する RL ベースのトリガー制御の \emph{最初}のデモンストレーションです。コードは https://github.com/Zixind/GFPO\_LHC で入手できます。

原文 (English)

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement learning agent ingests streaming summaries of recent rates and signal-sensitive features and updates trigger thresholds to maximize signal efficiency while tracking a target background rate within a tolerance band. We adapt Group-Filtered Policy Optimization (GFPO) to streaming control and introduce two variants (GFPO-F, GFPO-FR) that enforce background rate feasibility during training. On a benchmark that emulates realistic collider operation, we study two representative triggers: a total transverse energy ($H_{T}$) trigger sensitive to pileup variation, and an anomaly-detection (AD) trigger based on reconstruction loss for rare or non-standard signatures. On Monte Carlo streams, our agent increases the fraction of in-tolerance time intervals by 48\% ($H_T$) and 28\% (AD), with a cumulative gain of up to 2\% in signal efficiency on those in-tolerance intervals. Transferring from simulation to \emph{real} collision data (CMS Run 283408), the same agent, without fine-tuning, achieves a 56\% ($H_T$) and 28\% (AD) in-tolerance improvement over baselines, with further signal-efficiency gain on both triggers. To our knowledge, this is the \emph{first} demonstration of RL-based trigger control on real Large Hadron Collider collision data. Code is available at https://github.com/Zixind/GFPO\_LHC.

13:00 JST研究/論文

EMAgnet: 大規模ゲームにおけるポリシー勾配セルフプレイのためのパラメーター空間 EMA 正則化

最近の研究では、PPO などの正規化されたポリシー勾配手法をセルフプレイで使用すると、2 プレイヤーのゼロサム不完全情報ゲームを解くための特殊なゲーム理論アルゴリズムと同等またはそれを超えることができることが確立されました。均一分布は、この目的のための強力なポリシー正規化ターゲットとして浮上しましたが、実行可能性に関係なく、すべてのアクションに対して均等に正規化されます。 EMAgnet を導入します。これは、最終反復ポリシーのパラメーターの指数移動平均 (EMA) に向かって正規化し、エージェントの改善戦略に合わせて進化する適応的な正規化ターゲットを提供します。標準的な 2 プレイヤー ゼロサム ベンチマークと、探索チャレンジと多数の厳密に支配された戦略を備えた修正ベンチマークの両方で EMAgnet を評価します。線形アニーリング スケジュールとベキ乗則アニーリング スケジュールの両方で均一な磁石の正則化を使用した PPO セルフプレイと比較して、EMAgnet は、ほとんどのテスト環境で悪用可能性を低く抑え、厳密に支配された戦略を含むゲーム全体で一貫したパフォーマンスの向上を実現します。

原文 (English)

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-player zero-sum imperfect-information games. The uniform distribution has emerged as a strong policy regularization target for this purpose, but it regularizes equally toward all actions regardless of their viability. We introduce EMAgnet, which instead regularizes toward an exponential moving average (EMA) of the last-iterate policy's parameters, providing an adaptive regularization target that evolves with the agent's improving strategy. We evaluate EMAgnet on both standard two-player zero-sum benchmarks and modified benchmarks with exploration challenges and large numbers of strictly dominated strategies. Relative to PPO self-play with uniform-magnet regularization under both linear and power-law annealing schedules, EMAgnet achieves lower exploitability in the majority of tested environments, with consistent performance gains across games containing strictly dominated strategies.

13:00 JSTLLM/生成AI

仕様の学習に向けて: 優先ペアからの推論時間の調整

大規模言語モデル (LLM) を望ましい動作に導くには、通常、モデルの応答を注意深く検査してプロンプトを手作りする反復プロセスに依存します。これは複雑で脆弱で、エラーが発生しやすいプロセスです。設定ベースの微調整はより厳密ですが、多くの場合法外に高価なソリューションです。私たちは、短いユーザーの指示と少数の好みの判断に依存するフレームワークであるスペック学習を提案します。これらは、LLM の自然言語プロンプトの形式で仕様にコンパイルされます。仕様は推論時に LLM を条件付けするため、基礎となるモデルに対するパラメーターの更新は必要ありません。コンパイルされた仕様に基づいて生成された応答は、優先信号が高密度である特殊なドメインからのデータセットに対する直接優先最適化 (DPO) よりも優れたパフォーマンスを発揮することが多いことを示します。不透明な重み更新とは異なり、結果として得られる仕様は人間が判読可能であり、それを生成した優先信号の解釈可能で透明な文書化された実施形態としても機能します。

原文 (English)

Towards Spec Learning: Inference-Time Alignment from Preference Pairs

Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's responses. This is an involved, brittle, and error-prone process. Preference-based fine-tuning is a more rigorous but often prohibitively expensive solution. We propose spec learning, a framework that relies on a brief user instruction and a small set of preference judgments. These are compiled into specifications in the form of natural-language prompts for an LLM. Specifications condition LLMs at inference time, and no parameter updates to the underlying models are required. We show that the responses generated based on the compiled specifications often outperform direct preference optimization (DPO) on datasets from specialized domains whose preference signal is dense. Unlike opaque weight updates, the resulting specifications are human-readable and double as interpretable and transparent written embodiments of the preference signal that produced them.

13:00 JST研究/論文

高速および低速変分継続学習

継続的な学習は、現代のディープ ネットワークにとって依然として大きな課題です。その理由の 1 つは、一般的に使用されているオプティマイザーには、継続的に適応するための固有のメカニズムが欠けていることが挙げられます。そのような自然なメカニズムの 1 つは、安定性と可塑性のバランスをとるための急速な適応と遅い適応です。このメカニズムは神経科学と生物学に深く根ざしていますが、一般的に使用されるオプティマイザーにそれを組み込む最適な方法についてはコンセンサスがありません。ここでは、これが VCL フレームワークを介して簡単に実行できることを示します。このフレームワークでは、過去の事後値が将来の事前値として使用されます。私たちの重要なアイデアは、学習が進むにつれて知識のドリフトを遅らせるために、過去の事後値をマージすることでゆっくりとした適応を組み込むことです。マージされた事後分布は、VCL 更新で事前分布として使用され、高速重み付け更新が実装されます。これらのステップは、IVON オプティマイザーでシームレスに実装できます。IVON オプティマイザーの形式とコストは Adam のものとほぼ同じです。私たちは、この新しいオプティマイザーを Continual IVON (CoVON) オプティマイザーと呼び、既存の VCL オプティマイザーよりも一貫して改善するだけでなく、ドメイン増分学習、継続的な事前トレーニング、大規模な言語モデルの微調整にわたって他の重み正則化戦略よりも優れたパフォーマンスを発揮することを示します。

原文 (English)

Fast and Slow Variational Continual Learning

Continual learning remains a major challenge for modern deep networks, partly because commonly used optimizers lack inherent mechanisms for continual adaptation. One such natural mechanism is fast and slow adaptation to balance stability and plasticity. This mechanism has deep roots in neuroscience and biology, but there is no consensus on how to best incorporate it in commonly used optimizers. Here, we show that this can be easily done via the VCL framework, where past posteriors are used as priors in the future. Our key idea is to incorporate slow adaptation via merging of past posteriors to slow down the drift in the knowledge as learning progresses. The merged posterior is then used as the prior in the VCL update to implement the fast-weight updates. These steps can be seamlessly implemented in the IVON optimizer, whose form and costs are nearly identical to that of Adam. We call this new optimizer the Continual IVON (CoVON) optimizer and show that it not only consistently improves over existing VCL optimizers, but also performs better than other weight-regularization strategies across domain-incremental learning, continual pre-training, and fine-tuning of large language models.

13:00 JSTLLM/生成AI

マルチレイヤー MeMo のバージョンを意識した操作とトランザクション メモリに向けて

MeMo は、明示的な多層相関行列メモリ (CMM) を備えた言語モデルを提案します。このモデルでは、記憶、検索、および忘却がアーキテクチャ上の操作となります。この論文では、知識が変化したときにそのような記憶がどのようにして再訓練の必要性を減らすことができるかを尋ねます。 MeMo メモリの関連付けとして表現可能な変更の場合、モデル全体を再トレーニングするのではなく、明示的なメモリを編集することで、モデルのアクセス可能な知識を更新できます。私たちは、置換、廃止、履歴保持、ロールバック、トレースなどの高レベルの操作が、シーケンスとトークンを介した MeMo ネイティブのプリミティブ呼び出しにコンパイルされるバージョン認識操作層を提案します。重要な観察は、バージョンを認識する操作が単一の MeMo 関連付けであることはほとんどないということです。これは、原始的な編集の順序付けられたトランザクションであり、たとえば、あるシーケンス トークン チェーンを忘れたり、別のシーケンス トークン チェーンを記憶したり、履歴チェーンを保存したり、逆プログラムを記録したりします。このフレームワークでは、2 つの補助 CMM が導入されています。1 つはバージョン遷移をトランザクション ハンドルにマッピングするためのバージョン CMM (V-CMM)、もう 1 つは再利用可能な変更内容と逆プログラムを保存するためのトランザクション CMM (T-CMM) です。直接的なシーケンスレベルの編集と構造化された差分レベルの入力の両方をサポートし、更新の成功、ロールバック、トレーサビリティ、局所性、トランザクションの再利用の評価ルートの概要を示します。

原文 (English)

Towards Version-aware Operations and Transaction Memories for Multi-layer MeMo

MeMo proposes language models with explicit multi-layer correlation matrix memories (CMMs), where memorization, retrieval, and forgetting are architectural operations. This paper asks how such memories can reduce the need for retraining when knowledge changes. For changes expressible as MeMo memory associations, the model's accessible knowledge can be updated by editing explicit memories rather than retraining the whole model. We propose a version-aware operation layer in which high-level operations such as replace, obsolete, keep-history, rollback, and trace are compiled into MeMo-native primitive calls over sequences and tokens. The key observation is that a version-aware operation is rarely a single MeMo association. It is an ordered transaction of primitive edits, for example forgetting one sequence-token chain, memorizing another, preserving a historical chain, and recording an inverse program. The framework introduces two auxiliary CMMs: a Version CMM (V-CMM) for mapping version transitions to transaction handles, and a Transaction CMM (T-CMM) for storing reusable change contents and inverse programs. It supports both direct sequence-level edits and structured diff-level inputs, and outlines an evaluation route for update success, rollback, traceability, locality, and transaction reuse.

13:00 JST研究/論文

オートエンコーダーを使用した高速 FinFET モデリング

この研究では、FinFET を効率的にモデリングするためにオートエンコーダー (AE) を活用する機械学習フレームワークを紹介します。まず、BSIM-CMG モデルをキャリブレーションして、電流-電圧 (ID-VG) 特性のデータセットを生成しました。このデータは、完全な I-V 曲線を低次元の潜在空間に圧縮するオートエンコーダーをトレーニングするために使用され、主要なデバイスの物理現象を本質的にエンコードします。重要な革新は、ドレイン・ソース間電圧 (VDS) などのパラメーターを入力特徴として明示的に組み込むことで、バイアス依存の変動を捕捉するモデルの機能を強化します。トレーニングされたモデルは完全な I-V 曲線を正常に再構築し、しきい値電圧 (VTH)、しきい値未満の傾き (SS)、ピーク相互コンダクタンス (gm) などの重要なデバイス メトリクスを直接抽出します。このアプローチは、実際の特性評価データから構築されたデータ駆動型のコンパクトなモデルが最小限のトレーニング データで高精度を達成できることを実証し、迅速なデバイス特性評価、モデリング、および回路レベルのシミュレーションのための強力なツールを提供します。

原文 (English)

Rapid FinFET Modelling Using an Autoencoder

This work presents a machine learning framework that leverages an autoencoder (AE) for the efficient modeling of FinFET. We first calibrated a BSIM-CMG model to generate a dataset of current-voltage (ID-VG) characteristics. This data was used to train an autoencoder that compresses full I-V curves into a low-dimensional latent space, which intrinsically encodes key device physics. A key innovation is the explicit incorporation of parameter such as drain to source voltage (VDS) as an input feature, enhancing the model ability to capture bias dependent variation. The trained model successfully reconstructs full I-V curves and directly extracts critical device metrics including threshold voltage (VTH), subthreshold slope (SS), and peak transconductance (gm). This approach demonstrates that data driven compact models, built from actual characterization data, can achieve high accuracy with minimal training data, providing a powerful tool for rapid device characterization, modelling and circuit level simulation.

13:00 JST研究/論文

RAVEN: 金融時系列予測のための体制を意識した変数コンテキストの専門家ネットワーク

財務時系列予測には、標準ベンチマークにはない構造的な課題が存在します。対数リターンは非定常であり、信号対雑音比 (SNR) が非常に低く、レジーム依存の時間依存性によって支配されます。私たちは、金融設定における最先端 (SOTA) 時系列モデルの重要な制限を特定します。固定されたコンテキストウィンドウは、非定常価格プロセスの時間変化する最適なルックバックとは一致しません。我々は、各入力サンプルの時間的コンテキストを適応的に決定するように設計された専門家の混合フレームワークである、Regime-Aware Variable-context Expert Network (RAVEN) を提案します。 RAVEN は、固定されたルックバック範囲に依存する代わりに、データ自体によって長さが決定されるネストされた連続ウィンドウの階層を構築します。具体的には、RAVEN は学習された重要度によって逆時系列でパッチをスコアリングし、累積重要度しきい値 (CIT) メカニズムを適用してネストされたプレフィックス ウィンドウを導出し、それぞれがスケール専門の専門家にルーティングされます。 Global Compressed Representation (GCR) ブランチは、完全なコンテキストにわたって並行して実行され、ローカルの専門家が保証できないグローバルな時間的一貫性を維持します。ネストされたルーティングはエキスパート入力間の構造化された重複を誘発するため、可変長エキスパート出力を調整し、集計前にペアごとのコサイン類似性にペナルティを与えるために、Correlation-Aware Weighting (CAW) を導入します。累積ログリターン予測 (HS300、S&P500) とファンド販売予測に関する実験では、RAVEN が SOTA パフォーマンスを達成し、ピアソン相関を HS300 で 9.2%、S&P500 で 20.2% 改善し、ファンド販売予測で MSE を 18.2% 削減しながら、4 つの PEMS トラフィック ベンチマークの 16 メトリクスのうち 14 で最良の結果を達成していることが実証されました。

原文 (English)

RAVEN: A Regime-Aware Variable-context Expert Network for Financial Time Series Forecasting

Financial time series forecasting presents structural challenges absent from standard benchmarks. Log-returns are non-stationary, exhibit exceptionally low signal-to-noise (SNR) ratios, and are governed by regime-dependent temporal dependencies. We identify a key limitation of state-of-the-art (SOTA) time series models in financial settings. A fixed context window is mismatched to the time-varying optimal look-back of non-stationary price processes. We propose the Regime-Aware Variable-context Expert Network (RAVEN), a Mixture-of-Experts framework designed to adaptively determine the temporal context for each input sample. Instead of relying on a fixed look-back horizon, RAVEN constructs a hierarchy of nested contiguous windows whose lengths are determined by the data itself. Specifically, RAVEN scores patches by learned importance in reverse chronological order and applies the Cumulative Importance Thresholding (CIT) mechanism to derive nested prefix windows, each routed to a scale-specialized expert. A Global Compressed Representation (GCR) branch runs in parallel over the full context, preserving global temporal coherence that local experts cannot guarantee. Because the nested routing induces structured overlap among expert inputs, we introduce a Correlation-Aware Weighting (CAW) to align variable-length expert outputs and penalize pairwise cosine similarity prior to aggregation. Experiments on cumulative log-return prediction (HS300, S&P500) and fund sales forecasting demonstrate that RAVEN achieves SOTA performances, improves Pearson correlation by 9.2% on HS300 and 20.2% on S&P500, and reduces MSE by 18.2% on fund sales forecasting, while achieving the best results in 14 of 16 metrics on four PEMS traffic benchmarks.

13:00 JSTLLM/生成AI

エンドツーエンドの音声言語理解における選択的能力の未学習

最新の音声言語理解 (SLU) システムは、ポリシーや安全上の制約により特定の機能を削除する必要がある現実の環境に導入されることが増えています。 SLU では、機能はインテントとそれに関連するスロット生成動作に対応します。ただし、自己回帰モデルでは、ターゲット インテントを抑制しても、そのインテントに基づいて条件付けされたスロットを生成する条件付きマッピングは削除されません。インテント プレフィックスが外部から提供される場合、モデルは元のインテント スロット構造を再構築できます。この構造的欠陥は \textbf{\emph{機能の永続性}} として識別されます。私たちは、このマッピングの基礎となる意図条件付きの方向を分離して減衰する表現レベルのフレームワークである \textit{\underline{B}inding \underline{S}ubspace (BSU)} を提案します。 SLU ベンチマーク全体で、BSU は保持されたパフォーマンスを維持しながら、強制プレフィックスの回復可能性を大幅に低下させます。

原文 (English)

Selective Capability Unlearning in End-to-End Spoken Language Understanding

Modern spoken language understanding (SLU) systems are increasingly deployed in real-world settings, where specific functionalities may need to be removed due to policy or safety constraints. In SLU, a functionality corresponds to an intent and its associated slot-generation behavior. However, in autoregressive models, suppressing a target intent does not eliminate the conditional mapping that generates slots conditioned on that intent. When the intent prefix is externally supplied, the model can reconstruct the original intent-slot structure. We identify this structural failure as \textbf{\emph{capability persistence}}. We propose \textit{\underline{B}inding \underline{S}ubspace (BSU)}, a representation-level framework that isolates and attenuates intent-conditioned directions underlying this mapping. Across SLU benchmarks, BSU substantially reduces forced-prefix recoverability while preserving retained performance.

13:00 JST研究/論文

Stochastic-Oracle の信頼性を証明するためのトークンの複雑さ

Wang~\cite{Wang2026} は Stochastic-Oracle Turing Machine (SOTM) フレームワークを導入し、タスクに対して指定されたソリューション品質を達成するために必要な確率的オラクルとの対話に必要な最小予想コストとしてトークンの複雑さを定義しました。この論文では、特定のドメインにおける確率的オラクルの信頼性を証明するための同様の概念を開発します。証明書トークンの複雑さは、目標の信頼性レベルを満たすオラクルと、より低い信頼性しきい値を下回るオラクルを区別するために必要な、制御されたエラー確率で必要とされる最小予想トークンコストです。オラクルにクエリを実行し、バイナリ正しさスコアを計算し、蓄積された対数尤度証拠が判定しきい値を超えたときに停止する SPRT ベースの認証 SOTM を構築します。 SOTM はほぼ確実に停止し、認証される信頼性領域にわたって望ましい両側エラー保証を満たし、信頼性のしきい値、エラー限界、および予想されるターンごとのトークン コストの観点から、認証トークンの複雑さの明示的な上限を生成します。次に、一致する情報理論の下限を確立します。適応クエリを使用する場合でも、所定のエラー限界がゼロに近づく傾向があるため、すべてのエラー制限付き認証 SOTM は、SPRT ベースの構築と同じ先行予測トークン コストを負担する必要があります。これらの境界を合わせると、エラーが小さい領域での主要な次数の証明書トークンの複雑さが特徴付けられます。

原文 (English)

Token Complexity of Certifying Stochastic-Oracle Reliability

Wang~\cite{Wang2026} introduced the Stochastic-Oracle Turing Machine (SOTM) framework and defined token complexity as the minimum expected cost of interacting with a stochastic oracle needed to attain a specified solution quality for a task. This paper develops an analogous notion for certifying the reliability of a stochastic oracle on a given domain. Certification token complexity is the minimum expected token cost required, with controlled error probability, to distinguish oracles that meet a target reliability level from those that fall below a lower reliability threshold. We construct an SPRT-based certification SOTM that queries the oracle, computes binary correctness scores, and stops when the accumulated log-likelihood evidence crosses a decision threshold. The SOTM halts almost surely, satisfies the desired two-sided error guarantee over the reliability regions to be certified, and yields an explicit upper bound on certification token complexity in terms of the reliability thresholds, the error bound, and the expected per-turn token cost. We then establish a matching information-theoretic lower bound: even with adaptive queries, every error-bounded certification SOTM must incur the same leading-order expected token cost as the SPRT-based construction as the prescribed error bound tends to zero. Together, these bounds characterize the leading-order certification token complexity in the small-error regime.

13:00 JST画像/動画生成

ニューロモーフィック コンピューティングによるエンドツーエンドのレーダーと通信変調認識

深層学習ベースの手法は、自動変調認識 (AMR) タスクで高い精度を達成できますが、計算コストが高いため、精度と消費電力のバランスをとることが難しく、リソースに制約のあるプラットフォームでの適用が制限されます。適度なエネルギー予算でスパイク駆動推論を実行するニューロモーフィック アーキテクチャが、最近、ビジョンおよび時系列タスク向けに研究されています。これらの研究に動機付けられて、我々は、AMR 用のニューロモーフィック ハードウェアの制約にスパイク駆動トランスフォーマーを適用する、新しいエンドツーエンドのスパイキング ニューラル ネットワーク (SNN) アーキテクチャである EMRFormer を提案します。このモデルには、適応スパイク エンコーダーと整数リーク統合発射ニューロンが組み込まれており、有効な情報の劣化を軽減し、SNN 表現能力を強化します。スパイク分離可能な畳み込みニューラル ネットワーク (SSCNN) をスパイク駆動トランスフォーマー (SpikeFormer) に統合することにより、EMRFormer は生の IQ 波形からマルチスケールの時間的特徴を効果的に抽出します。私たちはさまざまな主流のデータセットにわたってアプローチを検証しました。実験結果は、EMRFormer が精度の点で最先端を達成し、すべてのベースラインを上回るパフォーマンスを示していることを示しています。さらに、このモデルは信号対雑音比 (SNR) が低い環境でも強力なパフォーマンスを維持し、理論上のエネルギー消費を 90% 以上削減します。最後に、KA200 ニューロモーフィック チップ上でモデルを評価します。結果は、私たちのモデルが 3090 GPU または Orin NX で実行する場合と比較して最大 5 倍の電力削減を達成していることを示しています。この研究は、リソースに制約のあるデバイスにおける AMR の有望な経路を示しています。

原文 (English)

End-to-End Radar and Communication Modulation Recognition with Neuromorphic Computing

Although deep learning-based methods can achieve high accuracy in automatic modulation recognition (AMR) tasks, their high computational cost makes it difficult to strike a balance between accuracy and power consumption, thereby limiting their application on resource-constrained platforms. Neuromorphic architectures that perform spike-driven inference with modest energy budgets have recently been explored for vision and timeseries tasks. Motivated by these works, we propose EMRFormer, a novel end-to-end spiking nerural network (SNN) architecture that applies spike-driven transformer to the constraints of neuromorphic hardware for AMR. The model incorporates an adaptive spike encoder and Integer Leaky Integrate-and-Fire neurons to mitigate the degradation of effective information and enhance SNN representational capacity. By integrating spike-separable Convolution Neural Networks (SSCNN) into Spike-Driven Transformers (SpikeFormer), EMRFormer effectively extracts multi-scale temporal features from the raw IQ waveforms. We validate our approach across various mainstream datasets, the experimental results show that EMRFormer achieves state-of-the-art interms of accuracy, outperforming all the baselines. Furthermore, the model maintains strong performance in low signal-to-noise(SNR) environments and reduces theoretical energy consumption by over 90%. Finally, we evaluate our model on a KA200 neuromorphic chip. The results show that our model achieves up to 5 times reduction in power compared to running on a 3090 GPU or an Orin NX. This work demonstrates a promising pathway for AMR on resource-constrained devices.

13:00 JSTビジネス/資金調達研究/論文

PixJail: テキストから画像へのジェイルブレイク評価のための自己進化する紙からパイプラインへの複製

Text-to-Image (T2I) ジェイルブレイク技術が急速に進化するにつれて、既存のベンチマークや複製ワークフローが追いつくのに苦労することがよくあります。さらに重要なことは、T2I ジェイルブレイク評価は単一のプロンプト レベルのテストではなく、プロンプト変換、画像生成、安全性フィルタリング、マルチモーダル判定などの複数の段階によって形成されるパイプライン レベルの問題であるということです。このため、複数の論文の結果を確実に再現し、公正に比較することが困難になります。このギャップを埋めるために、再現可能な T2I ジェイルブレイク評価のための自己進化する紙からパイプラインへのエージェント フレームワークである PixJail を提案します。 T2I ジェイルブレイク ペーパーとオプションの参照コードが与えられると、PixJail は元の実験結果を忠実に再現しながら、統一された契約の下でペーパー固有の攻撃モジュールと実行可能な評価パイプラインを迅速に構築します。 PixJail はさらに、論文のダイジェスト、攻撃の進化パターン、再利用可能なテンプレート、失敗例、バージョン管理された成果物を保存するメモリ バンクを維持し、以前の経験を再利用する将来の再現作業を可能にします。コードが利用可能な文書とコードが利用できない文書の両方を含む、11 の代表的な T2I 脱獄方法を再現します。元の設定では、フレームワークは最小限のエラー (2.1\% 平均、0\% 中央値) で以前の結果を正確に復元します。私たちは、PixJail が将来の T2I ジェイルブレイクの再現と評価のための統合基盤として機能し、手作業の労力を大幅に軽減できることを願っています。

原文 (English)

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation

As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More importantly, T2I jailbreak evaluation is not a single prompt-level test, but a pipeline-level problem shaped by multiple stages, including prompt transformation, image generation, safety filtering, and multimodal judging. This makes results across papers difficult to reliably reproduce and fairly compare. To bridge this gap, we propose PixJail, a self-evolving paper-to-pipeline agent framework for reproducible T2I jailbreak evaluation. Given a T2I jailbreak paper and optional reference code, PixJail rapidly constructs a paper-specific attack module and a runnable evaluation pipeline under a unified contract, while faithfully reproducing the original experimental results. PixJail further maintains a memory bank that stores paper digests, attack evolution patterns, reusable templates, failure cases, and versioned artifacts, enabling future reproduction efforts to reuse prior experience. We reproduce eleven representative T2I jailbreak methods, including both code-available and code-unavailable papers. Under their original settings, our framework accurately recovers prior results with minimal error (2.1\% average, 0\% median). We hope that PixJail can serve as a unified foundation for future T2I jailbreak reproduction and evaluation, significantly reducing manual effort.

13:00 JSTLLM/生成AIハードウェア/半導体

CAVEWOMAN: 言語入力および出力圧縮下で大規模言語モデルがどのように動作するか

「短く話してください。文法は省略してください。トークンを保存してください。」この穴居人のスタイルは、推論コストを削減する方法として広く推奨されていますが、実際に何かを節約できるかどうかは、どのチャネル (ユーザーのプロンプトまたはモデルの応答) が圧縮されているかによって異なります。我々は、タスクの精度、実現アイテムごとのコスト、およびモデルの制約のない参照に対する参照テキストの一致に関して世代ごとにスコアを付ける 2 チャネル評価プロトコルである Cave Woman を紹介します。両方のチャネルが同じ項目で測定され、5 つのデータセットの 8 つのモデルを 5 つの削減レベルで評価します。出力圧縮により、ほとんどの API モデル (モデルあたり 1.4 ~ 2.4 倍、最良の場合は最大 3 倍) とパブリック層価格設定の 4 つのオープンウェイト モデルすべてで実現コストが削減されます。入力圧縮には逆の効果があり、厳密な損失です。モデルは精度が低下しても、より長い応答で補正するため、正味コストは低下するのではなく増加します (5 つのベンチマーク平均で約 1.15 倍、最悪のデータセットで最大 1.8 倍、強力な圧縮下で 2.7 倍)。同じ設定の下では、表面テキストは制約のない参照から分岐します。非推論モデルでは、すべての世代の約半分が正しいにもかかわらず、それらの表面テキストはモデル独自の制約のないベースライン生成を必要としません。相違は、長さ制御された再スコアリング、複数比較の修正、および補完的な意味論的尺度の下での複製を経ても存続します。コードとデータは https://github.com/danielle34/cavewoman で入手できます。

原文 (English)

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being compressed. We present Cavewoman, a two-channel evaluation protocol that scores every generation on task accuracy, realized per-item cost, and reference-text agreement against the model's unconstrained reference. We evaluate eight models on five datasets at five reduction levels, with both channels measured on the same items. Output compression cuts realized cost on most API models (1.4-2.4x per model, up to 3x in the best case) and on all four open-weight models under public-tier pricing. Input compression has the opposite effect, a strict lose-lose: it raises net cost rather than lowering it (~1.15x on the five-benchmark mean, up to 1.8x on the worst dataset and 2.7x under stronger compression), because models compensate with longer responses even as accuracy collapses. Under the same setting, surface text diverges from the unconstrained reference: on the non-reasoning models, roughly half of all generations are correct yet their surface text no longer entails the model's own unconstrained baseline generation. The divergence survives length-controlled re-scoring, multiple-comparisons correction, and replication under complementary semantic measures. Code and data are available at https://github.com/danielle34/cavewoman.

13:00 JSTLLM/生成AI

ポリシーに基づく蒸留のためのブロック単位のポリシー ドリフト ゲーティング

オンポリシー蒸留 (OPD) は、生徒自身がサンプリングした軌跡に基づいて計算された教師信号を使用して生徒ポリシーをトレーニングします。最近の研究では、サンプリングされたトークンの OPD は長期的な推論タスクでは脆弱である可能性があり、ローカルの教師とサポートのマッチングが簡単で効果的な修復であることが示されています。このペーパーでは、ロールアウト再利用中の OPD 用の軽量の学生専用の古い電流ドリフト コントローラーであるブロック単位のポリシー ドリフト ゲーティングを紹介します。このメソッドは、サンプリングされたトークン パス上の行動生徒と現在の生徒の間の対数確率シフトを計算し、これらのシフトを固定ブロックまたはスパンにわたって集計し、結果として得られる切り離された平均正規化ゲートを使用して OPD 位置損失を再重み付けします。教師のターゲット、教師の Top-K サポート、ロールアウト ポリシーは変更されません。すべてのトレーニングされたバリアントに対して均一の 200 ステップのトレーニング バジェットを備えた 6 バリアントの Qwen3 数学推論ベンチマークでは、主要な問題レベルの解決率メトリクスとして pass@8 を使用します。固定 64 トークン ブロック ゲーティングにより、AIME24、AIME25、MATH500、および AMC23 全体でサンプル トークン OPD 平均パス @8 が 0.4978 から 0.5160 に改善されました。 Teacher-TopK/LSM では、Block64 は訓練を受けた生徒の中で最高の 4 ベンチマーク平均合格 @8 を示しました。その結果、局所的な古い現在のポリシーのドリフトが、再利用された OPD ロールアウトの実際的な制御信号として特定され、解決レートの堅牢性を向上させるための単純なデフォルトとしてブロック レベルのゲーティングが動機付けられました。

原文 (English)

Blockwise Policy-Drift Gating for On-Policy Distillation

On-policy distillation (OPD) trains a student policy using teacher signals computed on trajectories sampled by the student itself. Recent work shows that sampled-token OPD can be fragile on long-horizon reasoning tasks and that local teacher-support matching is a simple and effective repair. This paper introduces blockwise policy-drift gating, a lightweight student-only old-current drift controller for OPD under rollout reuse. The method computes log-probability shifts between the behavior student and the current student on the sampled token path, aggregates these shifts over fixed blocks or spans, and uses the resulting detached, mean-normalized gates to reweight OPD position losses. It does not change teacher targets, teacher top-K supports, or the rollout policy. In a six-variant Qwen3 math reasoning benchmark with a uniform 200-step training budget for all trained variants, we use pass@8 as the primary problem-level solve-rate metric. Fixed 64-token block gating improves sampled-token OPD mean pass@8 from 0.4978 to 0.5160 across AIME24, AIME25, MATH500, and AMC23. On Teacher-TopK/LSM, Block64 gives the best four-benchmark mean pass@8 among trained students. The results identify local old-current policy drift as a practical control signal for reused OPD rollouts and motivate block-level gating as a simple default for improving solve-rate robustness.

13:00 JSTロボティクス

DynaWM: 連続階段上のスムーズな移動のためのワールド モデルと運動量ターゲットを使用したダイナミクスを意識した蒸留

最近の制御の進歩により、二足歩行ロボットが斜面や段差のある障害物を通過できるようになりましたが、現在の教師と生徒のフレームワークでは力学を意識した表現が弱くなり、地形ジオメトリのエンコードが不完全であるため、長い階段を通過することは依然として困難です。このギャップを埋めるために、ダイナミクスを意識した表現学習フレームワークである DynaWM を提案します。地形エンコード機能を強化し、透過的な評価を可能にするために、フォワード ダイナミクスの認識を強化する正則化ツールとしてワールド モデルを導入し、階層エンコードの視覚化を容易にしながら包括的な地形ジオメトリを維持します。知識の伝達を安定させるために、運動量ターゲット エンコーダを採用して一貫した蒸留ターゲットを提供し、教師の非定常更新による次元の崩壊を防ぎます。主成分分析 (PCA) の可視化と定量的メトリクスによる学習された表現の評価により、エンコーダーがより高度な地形エンコード機能で地形ジオメトリを階層的にキャプチャし、地形適応性と動きの滑らかさが向上していることが明らかになりました。シミュレーションと実際のハードウェアによる実験結果は、図1に示すように、私たちの方法が優れた地形適応性と動作の滑らかさを実現し、二足歩行ロボットが多様な連続階段を克服できることを示しています。

原文 (English)

DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs

Recent advances in control have enabled bipedal-wheeled robots to traverse slopes and single-step obstacles, yet long staircase traversal remains challenging as current teacher-student frameworks suffer from weakened dynamics-aware representations and incomplete terrain geometry encoding. To bridge this gap, we propose DynaWM, a dynamics-aware representation learning framework. To enhance terrain encoding capability and enable transparent assessment, we introduce a world model as a regularizer to enforce forward-dynamics awareness, preserving comprehensive terrain geometry while facilitating hierarchical encoding visualization. To stabilize knowledge transfer, we employ a momentum target encoder to provide consistent distillation targets, preventing dimensional collapse from non-stationary teacher updates. Evaluation of the learned representations through Principal Component Analysis (PCA) visualization and quantitative metrics reveals that our encoder hierarchically captures terrain geometry with higher terrain encoding capability, leading to enhanced terrain adaptability and motion smoothness. Experimental results in simulation and real hardware demonstrate that our method achieves superior terrain adaptability and motion smoothness, enabling bipedal-wheeled robots to overcome diverse continuous stairs, as shown in Fig. 1.

13:00 JSTLLM/生成AI

詩から詩人の起源を予測する: 唐の詩全集における地域の言語的特徴の計算による分析

私たちは、唐代の詩人の地理的起源が彼らの作品に検出可能な言語的痕跡を残しているかどうかを尋ねます。 『全唐詩』(Quan Tang Shi)の各作者に帰属するすべての詩を集約し、中国伝記データベース(CBDB)を介して詩人を起源の行政回路にリンクすることで、10 の唐回路にわたる 357 人の詩人からなる詩人レベルのコーパスを構築し、多クラス分類として起源予測をフレーム化します。文字 $n$-gram TF-IDF と解釈可能な領域の特徴 (イメージ、季節、ほのめかし) を併用する古典的およびニューラル モデルは、詩人の広い地域 (南部対北部) を $0.69$ の精度で予測します。これは $0.53$ の多数派ベースラインをはるかに上回り、偶然を超えたより詳細な回路レベルの起源を予測します。分類を超えて、3 つの発見が得られます。 (i) 回路間の言語的距離は地理的距離に応じて増大します (9 回路にわたるマンテル $r=0.40$、$p\約 0.09$)。これは詩的言語における距離減衰効果の証拠です。 (ii) 信号は時間と相互作用します。南北の分離性は盛唐には偶然に起こり、晩唐には最も強くなり、帝国最盛期における宮廷主導の均質化とそれに続く地域的分岐と一致しています。 (iii) モデルの確信的な誤りは歴史的に意味があります。初唐では、すべての誤分類は南方の詩人が北方と読まれたものであり、北方宮廷の慣用句の威信を反映しています。さらに、階層的フリーズエンコーダ表現を通じてコーパス全体が与えられた場合、古典中国語変換器 (GuwenBERT) は単純な TF-IDF にのみ一致し (ビートではなく)、それらを組み合わせても何も加えないことを示し、文字 $n$-gram がすでに地域信号を捕捉していることを示しています。私たちの結果は、解釈可能な機械学習を文学史の仮説生成手段として位置付けています。

原文 (English)

Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems

We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work. Aggregating every poem attributed to each author in the Complete Tang Poems (Quan Tang Shi) and linking poets to their administrative circuit of origin via the China Biographical Database (CBDB), we build a poet-level corpus of 357 poets across the ten Tang circuits and frame origin prediction as multi-class classification. Using character $n$-gram TF-IDF together with interpretable domain features (imagery, season, and allusion), classical and neural models predict a poet's broad region (South vs.\ North) at $0.69$ accuracy, well above the $0.53$ majority baseline, and finer circuit-level origin above chance. Beyond classification, three findings emerge. (i) Linguistic distance between circuits grows with geographic distance (Mantel $r=0.40$, $p\approx0.09$ over nine circuits), evidence of a distance-decay effect in poetic language. (ii) The signal interacts with time: South/North separability is at chance in the High Tang and strongest in the Late Tang, consistent with court-driven homogenization at the empire's height followed by regional divergence. (iii) The model's confident errors are historically meaningful -- in the Early Tang, every misclassification is a southern poet read as northern, reflecting the prestige of the northern court idiom. We further show that, when given the whole corpus through a hierarchical frozen-encoder representation, a classical-Chinese transformer (GuwenBERT) only matches -- not beats -- simple TF-IDF, and that combining them adds nothing, indicating that character $n$-grams already capture the regional signal. Our results position interpretable machine learning as a hypothesis generator for literary history.

13:00 JST画像/動画生成エージェント

Bayer を超えて: 堅牢な自動運転セグメンテーションのためのタスク最適化センサーの共同設計

堅牢な認識は自動運転を支えており、最新の進歩は、モデルのより大きなバックボーン、基礎モデル、および協調的なマルチエージェント融合のスケーリングによってもたらされています。私たちは、カメラ自体が何を測定すべきかという補完的な上流の質問を追求します。微分可能な RAW からタスクへのパイプラインを使用して、どのセンサーの自由度が高密度予測に利益をもたらすかを分解します。スペクトル カラー フィルター アレイ (CFA) の重みを学習することが主要な手段であり、固定カメラと比較して mIoU が +0.017 (KITTI-360) および +0.023 (ACDC) 向上します。対照的に、点像分布関数 (光学) の共同設計は正味マイナス (KITTI-360 では -0.020 mIoU) です。これはデータ処理の不平等の結果であり、下流のモデルがどれほど大規模で協調的であっても、回復できるタスク情報も制限されます。ノイズの協調最適化はわずかで、フィルターがランク 3 の sRGB 入力に限定されているため、直感に反して CFA タイルを 2x2 を超えて拡大すると常に問題が発生します。介入はセンサーで行われるため、ゲインはモデルに依存しません。 ACDC の霧、夜、雨、雪に対する堅牢性を検証し、2x2 CFA 重みを学習し、アイデンティティ PSF を維持するという簡単なレシピで結論付けます。

原文 (English)

Beyond Bayer: Task-Optimal Sensor Co-Design for Robust Autonomous-Driving Segmentation

Robust perception underpins autonomous driving, and most recent progress comes from scaling the model-larger backbones, foundation models, and cooperative multi-agent fusion. We pursue a complementary, upstream question: what should the camera itself measure? Using a differentiable RAW-to-task pipeline, we decompose which sensor degrees of freedom benefit dense prediction. Learning the spectral colour-filter-array (CFA) weights is the dominant lever, improving mIoU by +0.017 (KITTI-360) and +0.023 (ACDC) over a fixed camera. In contrast, point-spread-function (optics) co-design is net-negative (-0.020 mIoU on KITTI-360) - a consequence of the data-processing inequality, which also bounds the task information that any downstream model, however large or cooperative, can recover. Noise co-optimisation is marginal, and counter to intuition enlarging the CFA tile beyond 2x2 consistently hurts, as the filters are confined to the rank three sRGB input. Because the intervention is at the sensor, the gains are model-agnostic; we validate robustness on ACDC's fog, night, rain, and snow, and conclude with a simple recipe: learn the 2x2 CFA weights and keep an identity PSF.

13:00 JST研究/論文

人文社会科学における中国人学生の学力発達に対する生成人工知能の影響

生成型人工知能 (GenAI) は高等教育における学習を再構築しており、特に人文科学および社会科学 (HSS) に顕著な影響を与えています。HSS では、一般的に学習成果は GenAI の機能と密接に連携した記述および解釈形式を通じて表現されます。しかし、HSS の生徒に対する GenAI の教育的影響に関する体系的な証拠は依然として限られています。このギャップに対処するために、この研究は中国のHSS学生を対象とした大規模調査に基づいて、学術発展におけるHSSの役割を調査しています。この研究は、関連する学習理論に基づいて、使用パターン、学習プロセスと学力への影響、GenAI の使用に関連する課題、カリキュラム統合への好ましいアプローチという 4 つの側面に焦点を当てています。私たちは、半数以上が学習意欲、独立した思考、創造性の向上を感じていることを発見しましたが、かなりの少数の人はほとんど変化がないか、さらには低下していると報告しました。比較的大多数が学業成績の向上を報告していますが、これらの向上は従来の評価慣行の限界を部分的に反映している可能性があります。この研究では、GenAIの経験期間が異なる生徒間での学習とパフォーマンスの向上の認識にばらつきがあり、観察可能な規律上の違いや軽度の性差も特定されています。圧倒的多数が倫理的配慮の重要性を評価している一方で、プライバシー保護に満足しているのは半数をわずかに超えるだけでした。学生から報告された最も差し迫った懸念として、精度の限界と過度の信頼が浮上しました。学生は実践指向のトレーニングによってサポートされる部分的またはオプションのカリキュラム統合を支持し、将来の専門能力開発にとって GenAI の重要性を広く認識しました。この研究は、学生の視点に基づいて、GenAI の責任ある教育学的に有意義な統合のための証拠に基づいた推奨事項を提供します。

原文 (English)

The impact of generative artificial intelligence on academic development of Chinese students in humanities and social sciences

Generative artificial intelligence(GenAI) is reshaping learning in higher education, with particularly pronounced implications for the humanities and social sciences(HSS), where learning outcomes are commonly expressed through written and interpretive forms that align closely with GenAI's capabilities. Yet, systematic evidence on the educational impacts of GenAI on HSS students remains limited. Addressing this gap, this study draws on a large-scale survey of HSS students in China to examine its role in academic development. Guided by relevant learning theories, this study focuses on four dimensions: patterns of use, effects on learning processes and academic performance, challenges associated with GenAI use, and preferred approaches to curricular integration. We found that more than half perceived enhanced learning motivation, independent thinking and creativity, although a substantial minority reported little change or even decline. Comparatively, a notably larger majority reported academic performance gains, although these gains may partly reflect limitations in conventional assessment practices. The study identifies variations in perceived learning and performance improvements among students with differing durations of GenAI experience, along with observable disciplinary differences and modest gender differences. While an overwhelming majority valued the importance of ethical considerations, only slightly more than half were satisfied with privacy protection. Limited accuracy and overreliance emerged as the most pressing concerns reported by students. Students favored partial or optional curricular integration supported by practice-oriented training, and widely recognized GenAI's significance for their future professional development. Grounded in student perspectives, this study offers evidence-based recommendations for the responsible and pedagogically meaningful integration of GenAI

13:00 JST画像/動画生成

ドラマディレクター: 幾何学に基づいた短編ドラマの生成

短いドラマは、素早いショットのリズム、対話による焦点の移動、映画的な基礎が要求されるため、プロンプト レベルまたはテキストのみのビデオ生成パイプラインでは対応するのが難しい課題が生じます。私たちは、グローバル プロットとローカル コンテキストが視覚的に根拠のあるマルチショット ビデオに変換される、プロットから短編ドラマへの生成を研究しています。私たちは、プランナーが深度とポーズによってインデックス付けされた実際の短編ドラマのショットのギャラリーから映画のジオメトリを借用できるようにする、ジオメトリに基づいたフレームワークである DramaDirector を提案します。 DramaDirector は、各ショットを静的なビジュアルと動的なナラティブ条件に分離し、学習されたテキストとビジュアルのアライメント報酬の下でスキーマに制約された SFT と GRPO を使用してプランナーをトレーニングし、最初のフレームの生成と画像からビデオへの合成をガイドする深度ポーズ参照を取得します。また、構造化されたストーリーボードと多次元の評価プロトコルを備えた、35 の実写ドラマ、2.8K のエピソード、81K のショットから構築されたベンチマークである DramaBoard も紹介します。実験の結果、DramaDirector は、忠実性、一貫性、および制御性に関して、代表的なマルチエージェントおよびビデオ生成のベースラインよりも向上していることが示されています。私たちのコードは https://github.com/iLearn-Lab/DramaDirector でリリースされています。

原文 (English)

DramaDirector: Geometry-Guided Short Drama Generation

Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation pipelines struggle to meet. We study plot-to-short-drama generation, where a global plot and local context are transformed into visually grounded multi-shot videos. We propose DramaDirector, a geometry-grounded framework that lets the planner borrow cinematographic geometry from a gallery of real short-drama shots indexed by depth and pose. DramaDirector decouples each shot into static visual and dynamic narrative conditions, trains the planner with schema-constrained SFT and GRPO under a learned text-visual alignment reward, and retrieves depth-pose references to guide first-frame generation and image-to-video synthesis. We also introduce DramaBoard, a benchmark built from 35 live-action dramas, 2.8K episodes, and 81K shots, with structured storyboards and multi-dimensional evaluation protocols. Experiments show that DramaDirector improves over representative multi-agent and video generation baselines on faithfulness, consistency, and controllability. Our code is released at: https://github.com/iLearn-Lab/DramaDirector

13:00 JST画像/動画生成研究/論文

消化器内視鏡検査用 VLM における幻覚検出のベンチマーク

視覚言語モデル (VLM) は幻覚を起こしやすいため、臨床現場での安全な導入にとって依然として大きな障壁となっています。現在まで、ほとんどの幻覚検出方法は MIMIC-CXR や VQA-RAD などの放射線学のベンチマークで評価されてきましたが、消化管 (GI) 内視鏡検査は依然として十分に研究されていません。この論文では、5 つの VLM (MedGemma-4B、MedGemma-27B、LLaVA-Med-7B、LLaVA-v1.6-7B、および Lingshu-32B) にわたる 4,392 のテスト VQA ペアを含む消化管診断視覚質問応答 (VQA) データセットである Gut-VLM データセットに対する 9 つの幻覚検出方法のベンチマークを行います。これらのメソッドは、ブラック ボックス メソッド (RadFlag、SelfCheckGPT-NLI)、グレー ボックス メソッド (AvgProb、AvgEnt、MaxProb、MaxEnt、Semantic Entropy、および VASE)、およびホワイト ボックス メソッド (ReXTrust) の 3 つのカテゴリにまたがっています。私たちの結果は、ホワイトボックス手法である ReXTrust が 5 つのモデルすべてで最高の AUC を達成し、各 VLM で最も強力な代替手法を統計的に有意なマージンで上回り (一対の順列検定、すべてのケースで p < 0.001)、MedGemma-4B ではピーク AUC 93.0 に達しました。ホワイトボックスの隠れ状態アクセスは、平均 19.5 AUC ポイント (範囲: 9.5 ~ 33.5) という一貫した利点を提供し、ブラック ボックス メソッドとクラスタリング ベースのグレー ボックス メソッドがほぼチャンスのパフォーマンスに崩壊する LLaVA-v1.6-7B (AUC 79.9) 上でも、ReXTrust は強力なパフォーマンスを維持します。非ホワイトボックス手法の中で、トークンレベルのグレーボックス統計 (MaxEnt、MaxProb) が最も強力な代替手法であり、平均してクラスタリングベースのグレーボックス手法 (セマンティック エントロピー、VASE) とブラックボックス アプローチの両方を上回ります。さらに、サンプル間一貫性またはトークンレベルの確率が高くモデルが幻覚を示す失敗モードである確信作話が、一貫性ベースの手法と不確実性ベースの手法の両方の体系的な失敗として特定されます。

原文 (English)

A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy

Vision-language models (VLMs) are prone to hallucination, which remains a major barrier to their safe deployment in clinical practice. To date, most hallucination detection methods have been evaluated on radiology benchmarks such as MIMIC-CXR and VQA-RAD, while gastrointestinal (GI) endoscopy remains largely underexplored. In this paper, we benchmark nine hallucination detection methods on the Gut-VLM dataset, a GI diagnostic Visual Question Answering (VQA) dataset with 4,392 test VQA pairs, across five VLMs (MedGemma-4B, MedGemma-27B, LLaVA-Med-7B, LLaVA-v1.6-7B, and Lingshu-32B). The methods span three categories: black-box methods (RadFlag, SelfCheckGPT-NLI), gray-box methods (AvgProb, AvgEnt, MaxProb, MaxEnt, Semantic Entropy, and VASE), and a white-box method (ReXTrust). Our results show that ReXTrust, a white-box method, achieves the highest AUC across all five models, outperforming the strongest alternative method on each VLM by a statistically significant margin (paired permutation test, p < 0.001 in all cases), reaching a peak AUC of 93.0 on MedGemma-4B. White-box hidden-state access provides a consistent advantage of 19.5 AUC points on average (range: 9.5--33.5), with ReXTrust maintaining strong performance even on LLaVA-v1.6-7B (AUC 79.9), where black-box methods and clustering-based gray-box methods collapse to near-chance performance. Among non-white-box methods, token-level gray-box statistics (MaxEnt, MaxProb) are the strongest alternatives, outperforming both clustering-based gray-box methods (Semantic Entropy, VASE) and black-box approaches on average. We further identify confident confabulation, a failure mode in which models hallucinate with high inter-sample consistency or high token-level probability, as a systemic failure for both consistency and uncertainty-based methods.

13:00 JST研究/論文

DTT-BSR+: 音楽ソース復元のための生成回帰カスケード

音楽ソース復元 (MSR) では、ソースのアンミックスと非線形制作効果の反転に共同で対処する必要があります。現在の方法では、意味の一貫性を維持しながらターゲット信号を正確に再構築するのが困難です。この制限に対処するために、分布フィッティングを信号再構築から別のステージに分離する 2 ステージ カスケード MSR システムである DTT-BSR+ を提案します。第 1 ステージの生成 DTT-BSR セパレーターは、クリーン ソースのプリアに一致するステムを生成し、第 2 ステージの修正された Demucs ネットワークは、時間領域と多重解像度のスペクトル損失を使用して第 1 ステージの出力を強化します。 DTT-BSR+ は、すべてのステムでシングルステージ DTT-BSR よりもマルチメル信号対雑音比 (MMSNR) を向上させ、5 つのステムで最先端の X-LANCE MSR システムを上回ります。また、Fr\'echet Audio Distance (FAD) 分解を通じて、信号再構成の精度とステム全体にわたる意味分布の適合との間の暗黙のトレードオフも明らかにします。

原文 (English)

DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration

Music source restoration (MSR) requires jointly addressing source unmixing and the inversion of non-linear production effects. Current methods struggle to achieve accurate target signal reconstruction while maintaining semantic consistency. To address this limitation, we propose DTT-BSR+, a two-stage cascade MSR system that decouples distribution fitting from signal reconstruction into separate stages. A generative DTT-BSR separator in the first stage produces stems matching the prior of clean sources, and a modified Demucs network in the second stage enhances the first stage output using time-domain and multi-resolution spectral losses. DTT-BSR+ improves multi-mel signal-to-noise ratio (MMSNR) over the single-stage DTT-BSR across all stems, and surpasses the state-of-the-art X-LANCE MSR system on five stems. We also reveal through Fr\'echet Audio Distance (FAD) decomposition an implicit trade-off between signal reconstruction accuracy and semantic distribution fitting across stems.

13:00 JSTLLM/生成AIエージェント

Metis: 自己進化エージェントのためのテキストとコード メモリの橋渡し

自己進化するエージェントは、過去の実行から経験を抽出し、それを将来のタスクで再利用することで、時間の経過とともに改善します。既存のシステムは、そのようなエクスペリエンスを、エージェント コンテキストに挿入された自然言語テキストとして、または呼び出し可能なツールとして公開されたコードとして表します。ただし、これらの表現間の選択は、エクスペリエンス自体の特性から導き出されるのではなく、設計時に行われるのが一般的であり、それらの間のトレードオフはほとんど理解されていません。我々は、同一の一連の経験においてテキスト記憶とコード記憶を分離する最初の対照研究を紹介する。我々の結果は、2 つの形式が構築コスト、実行効率、譲渡可能性において相補的なトレードオフを示し、どちらの表現だけでも十分ではないことを示しています。これらの発見に基づいて、階層的二重表現メモリ上に構築された自己進化エージェント システムである Metis を提案します。 Metis は、テキストのエクスペリエンスを実行計画、環境の事実、一般的な落とし穴に整理し、定期的な計画を選択的に検証済みの呼び出し可能なツールに結晶化します。この設計は、テキスト メモリの幅広い適用性とコード メモリの実行効率を組み合わせていますが、繰り返し再利用することが正当化される場合にのみツール生成コストが発生します。インタラクティブ エージェントにとっての挑戦的なベンチマークである Metis on AppWorld を評価します。結果は、Metis が ReAct よりもタスクの精度を最大 20.6% 向上させながら、実行コストを最大 22.8% 削減したことを示しています。代表的な自己進化型エージェント システムと比較して、Metis は精度、実行効率、メモリ構築コストの間で常に優れたバランスを実現しています。

原文 (English)

Metis: Bridging Text and Code Memory for Self-Evolving Agents

Self-evolving agents improve over time by distilling experience from past executions and reusing it in future tasks. Existing systems represent such experience either as natural-language text injected into the agent context or as code exposed as callable tools. However, the choice between these representations is typically made at design time rather than derived from the characteristics of the experience itself, leaving the trade-offs between them poorly understood. We present the first controlled study that isolates text memory and code memory over an identical set of experiences. Our results show that the two forms exhibit complementary trade-offs in construction cost, execution efficiency, and transferability, such that neither representation alone is sufficient. Guided by these findings, we propose Metis, a self-evolving agent system built on a hierarchical dual-representation memory. Metis organizes textual experience into execution plans, environment facts, and common pitfalls, and selectively crystallizes recurring plans into validated callable tools. This design combines the broad applicability of text memory with the execution efficiency of code memory while incurring tool-generation cost only when justified by repeated reuse. We evaluate Metis on AppWorld, a challenging benchmark for interactive agents. The results show that Metis improves task accuracy by up to 20.6% over ReAct while reducing execution cost by up to 22.8%. Compared with representative self-evolving agent systems, Metis consistently achieves a better balance between accuracy, execution efficiency, and memory-construction cost.

13:00 JST研究/論文

2 段階トレーニングによる、トライアル間の EEG ガイドによるターゲット音声抽出のためのショートカット学習の打破

EEG ガイドによるターゲット音声抽出のための最近のエンドツーエンド モデルは、目覚ましい結果を報告しており、神経誘導聴覚技術の可能性を強調しています。しかし、私たちの分析では、トライアル内の高いパフォーマンスは、ターゲット選択のショートカットとして機能するトライアル固有の EEG 構造によって促進され、未確認のトライアルでは一般化が不十分になる可能性があることが明らかになりました。このギャップを克服するために、ショートカット学習を軽減する 2 段階のフレームワークである TRUST-TSE を提案します。在席話者のネガティブサンプリングによる対照的な事前トレーニングを導入することで、EEG エンコーダがきめの細かい EEG、つまり試行同一性の手がかりを抑制しながら音声の整合性をキャプチャすることを奨励します。また、学習された表現を使用した抽出をガイドするために、EEG (ソース類似性) に基づいた信頼度で重み付けされた抽出目標も採用しています。 KUL および DTU データセットの実験では、TRUST-TSE が厳格なクロストライアル プロトコルの下でエンドツーエンドのベースラインを上回っており、既存のアプローチの主要な信頼性のボトルネックに対処していることが示されています。

原文 (English)

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials. To overcome this gap, we propose TRUST-TSE, a two-stage framework to mitigate shortcut learning. By introducing contrastive pretraining with attended-speaker negative sampling, we encourage the EEG encoder to capture fine-grained EEG--speech alignment while suppressing trial-identity cues. We also employ a confidence-weighted extraction objective based on EEG--source similarity to guide extraction using the learned representations. Experiments on KUL and DTU datasets show that TRUST-TSE outperforms end-to-end baselines under strict cross-trial protocols, addressing a key reliability bottleneck of existing approaches.

13:00 JSTLLM/生成AI

P\={a}nian インド言語処理財団

10 億人以上の人々がインドの言語でコミュニケーションを行っていますが、彼らにサービスを提供する自然言語処理インフラストラクチャは断片化され、未開発のままです。原因は構造的なものです。この分野では、個々の言語または系統言語族の小さなサブセットを中心にツールとベンチマークが編成され、言語ごとに個別のアナライザー、パーサー、データセットを構築し、次の言語に向けて最初からやり直しています。これは深い規則性を見落としています。サンスクリット語を中心とした 2000 年以上の収束を経て、インド諸言語は、P\={a}nini の文法である Ast\={a}dhy\={a}y\={i} で形式化された形態構文アーキテクチャを共有するようになりました。これは系譜を越えて、共通の枠組みを通じて言語を統一します。私たちは、この P\={a}ninian フレームワークがこの分野に欠けていた統一的な計算アーキテクチャを提供し、これに明示的に基づいたベンチマークがインド言語システムをより正確にし、よりデータ効率を高め、より転送可能にし、明らかに異質でまばらな多くのインド言語リソースを単一の高リソースのメタ言語基盤に効果的に統合するだろうと主張します。この共有アーキテクチャを明示的かつ測定可能にし、実用的なアプリケーションにすぐに活用できるようにするために、4 部構成のベンチマーク スイートを提案します。さらに、これらの言語で訓練されたニューラル モデルが独自に P\={a}nini のカテゴリを表すようになるかどうかという、解釈可能性の研究に対してこの問題が提起する疑問を強調します。

原文 (English)

A P\={a}ninian Foundation for Indic Language Processing

More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragmented and underdeveloped. The cause is structural: the field organizes its tools and benchmarks around individual languages or small subsets of genealogical language families, building separate analyzers, parsers, and datasets for each language and starting over for the next. This overlooks a deep regularity. Through more than two millennia of convergence around Sanskrit, Indic languages came to share a morphosyntactic architecture formalized in P\={a}nini's grammar, the Ast\={a}dhy\={a}y\={i}. This cuts across genealogical lines, uniting languages through a common framework. We argue that this P\={a}ninian framework supplies a unifying computational architecture the field has lacked, and that benchmarks grounded explicitly in it would make Indic language systems more accurate, more data-efficient, and more transferable, effectively merging many apparently disparate and sparse Indic language resources into a single high-resource metalanguage bedrock. We propose a four-part benchmark suite to render this shared architecture explicit, measurable, and ready to be leveraged for practical applications. Moreover, we underscore the question it raises for interpretability research: whether neural models trained on these languages come to represent P\={a}nini's categories on their own.

13:00 JST研究/論文

オンデバイス障害検出用の軽量トランス モデル: リソースに制約のある展開に関するベンチマーク調査

オンデバイスの障害検出により、クラウドに依存せずにリアルタイム診断が可能になりますが、リソースに制約のあるハードウェアに機械学習モデルをデプロイするには、精度、遅延、モデル サイズの間で慎重なトレードオフが必要になります。 NASA C-MAPSS ターボファンの劣化、SECOM 半導体製造、UCI AI4I 2020 予知保全の 3 つの公開データセットにわたるバイナリ障害検出について、従来の ML 手法 (ランダム フォレスト、XGBoost、SVM、ロジスティック回帰) と軽量トランス アーキテクチャ (DistilBERT、TinyBERT-6L、TinyBERT-4L、MobileBERT) を比較したベンチマークを示します。分類パフォーマンス (F1 スコア、AUC)、モデル サイズ、CPU 推論レイテンシーを評価し、さらに INT8 動的量子化と 2 段階の適応推論パイプラインを評価します。私たちの結果は、十分に分離されたセンサー データ (C-MAPSS) では、軽量トランスフォーマーは 87.8% F1 で従来の ML と一致しますが、モデル サイズは 100 倍、遅延は 9000 倍であることが明らかになりました。 TinyBERT-4L は、55 MB および 18 ミリ秒の CPU 遅延で、最も導入しやすいトランスとして登場します。 INT8 量子化により、F1 の 86.9% を維持しながらサイズが 25% 削減されます。当社の適応パイプラインは、予測の 97.9% を量子化トリアージ モデルを通じてルーティングし、2.1% だけをより大規模な専門家にルーティングし、19.5 ミリ秒の平均レイテンシーで 87.6% の F1 を達成します。著しく不均衡なデータセット (SECOM、UCI-PM) では、従来の手法とトランスフォーマー手法の両方が非常に困難であり、障害検出における極端なクラスの不均衡に対する現在のアプローチの基本的な限界が浮き彫りになっています。すべてのコードは公開されています。

原文 (English)

Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment

On-device fault detection enables real-time diagnostics without cloud dependency, but deploying machine learning models on resource-constrained hardware demands careful tradeoffs between accuracy, latency, and model size. We present a benchmark comparing traditional ML methods (Random Forest, XGBoost, SVM, Logistic Regression) against lightweight transformer architectures (DistilBERT, TinyBERT-6L, TinyBERT-4L, MobileBERT) for binary fault detection across three public datasets: NASA C-MAPSS turbofan degradation, SECOM semiconductor manufacturing, and UCI AI4I 2020 predictive maintenance. We evaluate classification performance (F1-score, AUC), model size, and CPU inference latency, and further assess INT8 dynamic quantization and a two-stage adaptive inference pipeline. Our results reveal that on well-separated sensor data (C-MAPSS), lightweight transformers match traditional ML at 87.8% F1 but at 100x the model size and 9000x the latency. TinyBERT-4L emerges as the most deployment-friendly transformer at 55 MB and 18 ms CPU latency. INT8 quantization reduces size by 25% while preserving 86.9% F1. Our adaptive pipeline, routing 97.9% of predictions through a quantized triage model and only 2.1% to a larger expert, achieves 87.6% F1 at 19.5 ms average latency. On severely imbalanced datasets (SECOM, UCI-PM), both traditional and transformer methods struggle significantly, highlighting fundamental limitations of current approaches for extreme class imbalance in fault detection. All code is publicly available.

13:00 JSTLLM/生成AIエージェント研究/論文

Agon: 即時経済に基づいて構築された自律的な大規模かつ全分野の研究システム

大規模な言語モデルにより研究生産が拡張可能になり、ボトルネックが成果物の生成から主張の判断へと移行しています。私たちは、ワークフロー内でチェックできる内容を検証し、残りの判断を人間の科学者に委ねる研究オーケストレーターである \textsc{Agon} を紹介します。 \textsc{Agon} は、プロンプト エコノミー、未来志向、最小限のプロンプト、OmniDisciplinary、大規模並列処理、ゼロコードという 6 つの設計原則に基づいて構築されています。私たちは、小さな開始トピックのみを使用し、人間が作成した実験コードを使用せずに、プロンプト エコノミー ループを 444 回反復してドメイン全体で \textsc{Agon} を実行しました。これらの展開は、新しいクラスの障害を明らかにしながら、スケーラビリティを実証します。これらの障害を、重大度、修正可能性、可視性、機能の軌跡に沿った分類に整理します。この分類では、ループが認識して修正できる障害と、人間の判断が必要な障害とが区別されます。これらの結果を総合すると、\textsc{Agon} が機械のスケールと人間の操縦という新しいパラダイムに向けて研究を推進していることを示しています。

原文 (English)

Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy

Large language models are making research production scalable, shifting the bottleneck from producing artifacts to judging claims. We present \textsc{Agon}, a research orchestrator that validates what can be checked inside the workflow and leaves the remaining judgments to human scientists. \textsc{Agon} is built on six design principles: Prompt Economy, Future-Facing, Minimal Prompts, OmniDisciplinary, Massive Parallelism, and Zero-Code. We ran \textsc{Agon} across domains for 444 iterations of Prompt Economy loops, using only small starting topics and no human-written experimental code. These deployments demonstrate scalability while exposing new classes of failure. We organize these failures into a taxonomy along severity, fixability, visibility, and capability locus. The taxonomy separates failures the loops can see and fix from those that require human judgment. Together, these results show that \textsc{Agon} is pushing research toward a new paradigm: machine scales, human steers.

13:00 JST画像/動画生成

分布外スコアリングを使用したゼロショット テスト時間の正規化

事前トレーニングされた視覚モデルは、アフィン変換によってオブジェクト クラスが変更されない場合でも、回転、スケーリング、またはせん断された入力を誤って分類することがよくあります。通常、ロバスト性はアーキテクチャに等分散性を組み込むか、拡張による再トレーニングによって復元されますが、どちらの場合もモデルの変更または再トレーニングが必要です。テスト時の正規化では、分類子はそのまま残ります。各入力の変換を元に戻し、分類前のトレーニング分布に近い標準形式にマッピングします。しかし、既存の正規化ツールは、ロジットベースのエネルギー スコアの狭いセットとオーダーメイドの検索手順に依存しており、スコアリング関数とオプティマイザーの設計領域は未調査のままです。正規化を分布外 (OOD) 検出として再構築し、OOD スコアを変換を通じて最小化されたエネルギーとして機能させます。手書きの文字やスケッチから自然画像や 3D 点群に至るまでのベンチマークにわたって、約 20 個の OOD スコアと 9 個の検索アルゴリズムを体系的に評価し、距離ベースのスコアとランダム検索およびローカル リファインメントを組み合わせたものが全体的に最も優れたパフォーマンスを発揮することがわかりました。すでに整列された入力を正規化すると精度が損なわれる可能性があるため、OOD スコアが変換が必要であることを示した場合にのみ入力を変換するゲート メカニズムを追加し、変換された入力の堅牢性の向上を維持しながら、ほとんどの分布内の精度を維持します。コードは github.com/johschm/its で入手できます。

原文 (English)

Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring

Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the object class unchanged. Robustness is usually restored either by building equivariance into the architecture or by retraining with augmentation, both of which require changing or retraining the model. Test-time canonicalization instead leaves the classifier untouched. It undoes the transformation of each input, mapping it to a canonical form near the training distribution before classification. Existing canonicalizers, however, rely on a narrow set of logit-based energy scores and bespoke search procedures, leaving the design space of scoring functions and optimizers unexplored. We reframe canonicalization as out-of-distribution (OOD) detection, which lets any OOD score serve as the energy minimized over transformations. Across benchmarks ranging from handwritten characters and sketches to natural images and 3D point clouds, we systematically evaluate around twenty OOD scores and nine search algorithms, finding that distance-based scores paired with random search and local refinement perform best overall. Because canonicalizing an already-aligned input can hurt accuracy, we add a gated mechanism that transforms an input only when its OOD score indicates this is needed, preserving most in-distribution accuracy while retaining the robustness gains on transformed inputs. Code is available at github.com/johschm/its.

13:00 JST画像/動画生成ロボティクス

3D 医療現場を完成させるための深層学習アプローチ: 幾何モデリングから生成パラダイムまで

3 次元シーンの完成は、コンピューター ビジョンとロボット工学の主要な問題として発展しており、その応用例は自律ナビゲーションや拡張現実など多岐にわたります。この研究では、過去 10 年間、つまり 2016 年から 2026 年に行われた研究貢献をまとめるために体系的なレビューが行われました。この分野は、SSCNet に代表されるボクセル意味補完パラダイムから、ガウス スプラッティング手法を使用した生成拡散プリアとリアルタイム レンダリングを組み合わせた最新パラダイムまで、この分野に革命をもたらしました。この研究では、ボクセル グリッド、点学習、暗黙的ニューラル フィールド、変換ネットワーク、拡散ネットワーク、レンダリング対応 3D ガウス プリミティブに基づく最新のパラダイムなどの表現パラダイムの進化について議論されています。過去 10 年間に行われた貢献について包括的な分析が行われ、この分野で行われた貢献について明確なアイデアを提供する分類法が開発されました。この研究では、この分野で行われた研究の貢献と、まだ対処する必要がある課題についても議論されています。最後に、この研究は、次世代システムの開発において従うことができる方向性についての明確なアイデアを提供する研究課題を提示しました。

原文 (English)

Deep Learning Approaches for 3D Medical Scene Completion: From Geometric Modeling to Generative Paradigms

Three-dimensional scene completion has evolved as a major problem in computer vision and robotics, and its applications are diverse, including autonomous navigation and augmented reality. In this study, a systematic review has been conducted to compile the research contributions made in the last ten years, i.e., 2016 to 2026, which has revolutionized the field from the voxel semantic completion paradigm represented by SSCNet to the latest paradigm that combines generative diffusion priors with real-time rendering using a Gaussian splatting technique. The evolution in representation paradigms, such as voxel grids, point learning, implicit neural fields, transformer networks, diffusion networks, and the latest paradigm based on rendering-aware 3D Gaussian primitives, has been discussed in this study. A comprehensive analysis has been carried out on the contributions made in the last ten years, and a taxonomy has been developed to provide a clear idea about the contributions made in the field. The study has also discussed the research contributions made in the field, along with the challenges that still need to be addressed. Finally, the study has presented a research agenda that will provide a clear idea about the directions that can be followed in the development of the next-generation system

13:00 JSTLLM/生成AI画像/動画生成

拡散アンラーニングで同時発生する関連保持概念

アンラーニングは、拡散モデルにおける有害なコンテンツの生成を軽減するための重要な技術として浮上しました。しかし、既存の手法では、ターゲットの概念だけでなく、共起する無害な概念も削除してしまうことがよくあります。図 1 に示すように、ヌードを学習しないと、人物の概念が意図せず抑制され、モデルが人物を含む画像を生成できなくなります。私たちは、これらの望ましくない抑制された、保存する必要がある共起概念を CARE (Co-occurring Associated REtained Concepts) と定義します。次に、非学習タスク全体での保存を直接定量化する一般的な指標である CARE スコアを導入します。これを基盤として、対象概念のみを消去しながらCAREを明示的に保護するフレームワークであるReCARE(Robust Erasure for CARE)を提案します。 ReCARE は、ターゲット画像から抽出された無害な共起トークンの厳選された語彙である CARE セットを自動的に構築し、安定した非学習のためのトレーニング中にこの語彙を活用します。さまざまなターゲット コンセプト (ヌード、ゴッホ スタイル、テンチ オブジェクト) にわたる広範な実験により、ReCARE が堅牢なコンセプトの消去、全体的な実用性、および CARE の保存のバランスにおいて全体的に最先端のパフォーマンスを達成することが実証されました。

原文 (English)

Co-occurring associated retained concepts in Diffusion Unlearning

Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co-occurring concepts. As illustrated in Fig.1, unlearning nudity can unintentionally suppress the concept of person, preventing a model from generating images with person. We define these undesirably suppressed co-occurring concepts that must be preserved CARE (Co-occurring Associated REtained concepts). Then, we introduce the CARE score, a general metric that directly quantifies their preservation across unlearning tasks. With this foundation, we propose ReCARE (Robust erasure for CARE), a framework that explicitly safeguards CARE while erasing only the target concept. ReCARE automatically constructs the CARE-set, a curated vocabulary of benign co-occurring tokens extracted from target images, and leverages this vocabulary during training for stable unlearning. Extensive experiments across various target concepts (Nudity, Van Gogh style, and Tench object) demonstrate that ReCARE achieves overall state-of-the-art performance in balancing robust concept erasure, overall utility, and CARE preservation.

13:00 JSTLLM/生成AI研究/論文

MMed-Bench-IR: 多言語医療情報検索のための異種ベンチマーク

臨床現場における検索拡張生成 (RAG) では、主に英語の証拠コーパスに対する多言語検索がますます必要になります。多言語による医療検索には、言語間の調整、概念の識別、および証拠の検索という 3 つの機能が必要です。しかし、既存のベンチマークはこれらを単独でのみ評価しており、生物医学の専門知識と多言語対応範囲との間の相互作用は測定されていません。 MMed-Bench-IR は、6 つの言語と 3 つの構造的に異質なタスクにわたるこれらの軸を解きほぐすように設計されたベンチマークです。(1) 統一医療言語システム (UMLS) に基づいた 6,127 クエリによる言語を超えた医療 QA 検索、(2) 3 つの難易度での 4,975 の混同セットにわたる概念識別、および (3) RAG の多言語証拠検索。 2,040 の品質が保証されたクエリ。 3 つのタスクは概念をまったく共有しておらず、設計によりクエリが重複しているため、集計スコアが真の機能範囲を反映していることが保証されます。 6 つのパラダイムファミリーにわたる 10 のシステムを評価したところ、言語間の重大な障害が明らかになりました。英語で 0.818 nDCG@10 のスコアを記録した生物医学エンコーダは、日本語では 0.056 に低下しました。このギャップは、英語のみのベンチマークでは検出できませんでした。

原文 (English)

MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora. Multilingual medical retrieval demands three capabilities: cross-lingual alignment, concept discrimination, and evidence retrieval. However, existing benchmarks evaluate these only in isolation, leaving the interaction between biomedical expertise and multilingual coverage unmeasured. We introduce MMed-Bench-IR, a benchmark designed to disentangle these axes across 6 languages and three structurally heterogeneous tasks: (1) cross-lingual medical QA retrieval with 6,127 queries grounded in the Unified Medical Language System (UMLS), (2) concept discrimination over 4,975 confusion sets at three difficulty tiers, and (3) multilingual evidence retrieval for RAG with 2,040 quality-assured queries. The three tasks share zero concept and query overlap by design, ensuring that aggregate scores reflect genuine capability breadth. Evaluation of ten systems across six paradigm families reveals severe cross-lingual failure: biomedical encoders that score 0.818 nDCG@10 in English drop to 0.056 in Japanese, a gap that English-only benchmarks cannot detect.

13:00 JST画像/動画生成

マルチビューの一貫した構成の 3D 生成のための包括的なインタラクティブな衝突

3D 生成における最近の進歩は、テキストから画像への拡散モデルの開発により顕著に進歩しました。しかし、既存の方法には 2 つの実際的な課題が残されています。(1) 既存の方法は主に単一の 3D オブジェクトを生成しますが、合理的な相互作用におけるガウス プリミティブのモデリングが欠如しているため、複数オブジェクトの合成 3D アセットを生成するのが困難です。 (2) スコア蒸留サンプリングは本質的に各単一ビューに対して実行されるため、3D 最適化中にビュー間の不一致が発生し、必然的にビュー間の幻覚が発生します。上記の問題を解決するために、合理的なインタラクションを備えたマルチビューの一貫した構成 3D アセットを生成する新しい最適化ベースの方法である I2C-3D を提案します。具体的には、合理的な相互作用領域に自然に現れるガウス プリミティブをガイドする包括的インタラクティブ衝突戦略を提案します。これにより、構成シーン内のオブジェクトが物理的に妥当で視覚的に一貫した方法で相互作用するようになります。さらに、マルチビューの一貫性を強化するために、視点全体でインスタンス トークンと空間トークンのアテンション マップを調整することで、事前トレーニングされた拡散モデルから事前にマルチビューの一貫性とレイアウトを抽出するマルチビュー適応スコア蒸留サンプリングが考案されています。上記の精緻な設計の恩恵を受けて、I2C-3D は高忠実度のマルチビューで一貫した構成の 3D アセットを生成するだけでなく、3D 編集を柔軟にサポートし、複雑なシーンの生成を容易にします。広範な実験により、当社の I2C-3D は生成品質とマルチビューの一貫性において既存の方法より優れていることが実証されました。

原文 (English)

Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-view inconsistency during 3D optimization, as Score Distillation Sampling inherently performs on each single view, inevitably resulting in cross-view hallucinations. To solve above issues, we propose I2C-3D, a novel optimization-based method to generate multi-view consistent compositional 3D assets with reasonable interactions. Specifically, we propose an Inclusive Interactive Collisions strategy to guide Gaussian primitives appearing in reasonable interaction regions naturally, thereby ensuring objects in the compositional scene interact in a physically plausible and visually coherent way. Additionally, to enhance multi-view consistency, Multi-View Adaptive Score Distillation Sampling is devised to distill multi-view consistency prior and layout prior from pre-trained diffusion model by modulating attention map of instance token and spatial token across viewpoints. Benefiting from above elaborate designs, I2C-3D not only generates high-fidelity multi-view consistent compositional 3D assets but also supports 3D editing flexibly, facilitating complex scene generation. Extensive experiments demonstrate our I2C-3D outperforms existing methods in generation quality and multi-view consistency.

13:00 JSTLLM/生成AIエージェント

AutoSpec: 帰納的論理プログラミングによる LLM エージェントの安全ルールの進化

大規模言語モデル (LLM) エージェントは、言語モデルを外部ツールや環境と統合することで、複雑なタスクを自動化することが増えています。ただし、その自律性は重大な安全上のリスクをもたらします。エージェントは破壊的なコマンドを実行したり、機密データを漏洩したり、ドメインの制約に違反したりする可能性があります。既存の安全性アプローチは根本的なトレードオフに直面しています。手作りのルールは解釈可能ですが脆弱で、過度に保守的なルールは安全な操作をブロックし(高い誤検知)、寛容なルールは危険な動作を見逃します(高い誤検知)。ニューラル分類子には、セーフティ クリティカルな展開に必要な解釈可能性が欠けています。 AutoSpec は、展開された専門家が設計した安全ルールを、ユーザーの安全/安全でない注釈から、帰納的論理プログラミング (ILP) によってガイドされた反例誘導型帰納合成 (CEGIS) を通じて自動的に進化させるフレームワークです。 AutoSpec は、エキスパート ルールと注釈付きトレースのストリームから開始して、ルールを繰り返し評価し、偽陽性と偽陰性の反例をマイニングし、ILP を使用してルールを区別する述語を学習し、候補ルールの編集を生成し、候補を検証して最適なリビジョンを選択します。重要な洞察は、ILP が、偽陰性では頻繁に現れるが、偽陽性ではめったに現れない (またはその逆) 述語を効率的に識別し、ルール編集の指数関数的な検索スペースを大幅に削減することです。これは収束するまで続き、精度と再現率のバランスをとった解釈可能なルールが生成されます。コード実行と組み込まれたエージェント ドメインにわたる 291 の実行トレースで AutoSpec を評価します。 AutoSpec は、2 つのドメイン全体でルール F1 を 0.98 および 0.93 に引き上げ、高い再現率を維持しながら最大 94% の誤検知削減を達成し、4 ~ 5 回の反復以内に収束します。 ILP に基づくアプローチは、ヒューリスティック CEGIS よりも最大 4.8 倍高い F1 を達成します。学習されたルールは人間が判読可能で監査可能であり、目に見えないシナリオにも一般化されます。

原文 (English)

AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming

Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, their autonomy poses significant safety risks: agents may execute destructive commands, leak sensitive data, or violate domain constraints. Existing safety approaches face a fundamental tradeoff: hand-crafted rules are interpretable but brittle, with overly conservative rules blocking safe operations (high false positives) while permissive rules miss unsafe behaviors (high false negatives). Neural classifiers lack the interpretability required for safety-critical deployments. We present AutoSpec, a framework that automatically evolves deployed expert-designed safety rules from user safe/unsafe annotations through counterexample-guided inductive synthesis (CEGIS) guided by inductive logic programming (ILP). Starting from the expert rules and a stream of annotated traces, AutoSpec iteratively evaluates rules, mines false-positive and false-negative counterexamples, uses ILP to learn which predicates discriminate them, generates candidate rule edits, and verifies candidates to select the best revision. The key insight is that ILP efficiently identifies predicates that appear frequently in false negatives but rarely in false positives (or vice versa), dramatically pruning the exponential search space of rule edits. This continues until convergence, producing interpretable rules that balance precision and recall. We evaluate AutoSpec on 291 execution traces spanning code execution and embodied agent domains. AutoSpec raises rule F1 to 0.98 and 0.93 across the two domains, achieving up to 94% false positive reduction while maintaining high recall, and converges within 4-5 iterations. The ILP-guided approach achieves up to 4.8x higher F1 than heuristic CEGIS. The learned rules are human-readable, auditable, and generalize to unseen scenarios.

13:00 JST画像/動画生成

3D 人間と人間のインタラクションの生成において社会構造が重要になる

テキストからモーションへの生成は、言語からリアルな 1 人の人物のモーションを合成する点で大きな進歩を遂げていますが、それをテキスト駆動の 3D 人間対人間インタラクション (HHI) に拡張することは依然として自明ではありません。HHI には、段階の進行、アクターの役割、およびアクター間の調整を制御する基礎となる \textbf{社会構造} のモデル化が必要だからです。この論文では、HHI の生成を社会構造のモデリングおよびグラウンディングの問題として定式化します。モデルはまず、相互作用がどのように展開するか、2 人のアクターがそれぞれの役割をどのように調整するかを推測し、次にこの構造を連続的で物理的に妥当な、パートナーを意識した 3D モーションとして実現する必要があります。このような構造をどのようにモデル化すべきかを研究するために、まず、HHI 生成のための大規模言語モデル (LLM) の機能境界を調べます。私たちの分析によると、LLM は位相分解とパートナーを意識した役割を回復することで \textit{考える} ことはできますが、動的で物理的に妥当なインタラクションを意識した動きを生成できないため、直接 \textit{動く} ことはできません。これは、プランナーと実行者のパラダイム、\textbf{LLM で考え、モーション スキルで動く} を動機づけます。 LLM プランナーは、インタラクションをフェーズに分解し、パートナーを意識したアクターの役割を割り当て、モーション シーケンスと調整することで、暗黙的なインタラクション セマンティクスをモーションに合わせた社会的監視に変換します。次に、動作実行者は、LoRA、前段階の自己調整、および自我相対パートナー条件付けを使用して事前トレーニングされた単独動作モデルを適応させることにより、計画された社会構造を調整された 2 人の動作に基礎付けます。当社の Solo-to-Social フレームワークは、ソーシャル組織とモーションの実現を橋渡しし、フェーズの一貫性、役割の調整、パートナーを意識した調整を改善した 3D HHI を生成します。

原文 (English)

Social Structure Matters in 3D Human-Human Interaction Generation

Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it to text-driven 3D human-human interaction (HHI) remains non-trivial, as HHI requires modeling the underlying \textbf{social structure} that governs phase progression, actor roles, and inter-actor coordination. In this paper, we formulate HHI generation as a social structure modeling and grounding problem: the model must first infer how an interaction unfolds and how the two actors coordinate their roles, and then realize this structure as continuous, physically plausible, and partner-aware 3D motion. To study how such structure should be modeled, we first examine the capability boundary of large language models (LLMs) for HHI generation. Our analysis shows that LLMs can \textit{think} by recovering phase decompositions and partner-aware roles, but cannot directly \textit{move}, as they fail to generate dynamic, physically plausible, and interaction-aware motion. This motivates our planner-executor paradigm, \textbf{Think with LLM, Move with Motion Skill}. The LLM planner converts implicit interaction semantics into motion-aligned social supervision by decomposing interactions into phases, assigning partner-aware actor roles, and aligning them with motion sequence. The motion executor then grounds the planned social structure into coordinated two-person motion by adapting a pretrained solo motion model with LoRA, previous-phase self-conditioning, and ego-relative partner conditioning. Together, our Solo-to-Social framework bridges social organization and motion realization, producing 3D HHI with improved phase consistency, role alignment, and partner-aware coordination.

13:00 JSTLLM/生成AIビジネス/資金調達

SURGELLM: クラスバランス正規化によるタスク認識機能ゲーティングによるマルチタスク評価の再考

異種混合の NLP タスク全体に導入された微調整されたエンコーダーは、3 つの複雑な問題に直面しています。それは、帰納的バイアスの不一致、特徴量統計のクラス不均衡の破損、および外部の語彙知識に注意を払うメカニズムの欠如です。 \textbf{\surgellm} は、専用の軽量モジュールでそれぞれに対応する統合トランスフォーマー フレームワークです。 \emph{外科的特徴ゲート} (精選された語彙指標と \texttt{[CLS]} で次元ごとのシグモイドを学習します。特徴が有益でない場合は、明らかに同一性に退化します)、 \emph{タスク条件付きプレフィックス トークン} (量子化された特徴値とタスク同一性)すべての入力に付加されます)、および \emph{インスタンス加重正規化} (IWN; ゲート統計からクラス事前バイアスを削除します)。私たちは、\emph{外科的特徴の位置合わせ} にゲートをリンクする超過リスク限界の利点を証明します。 SST-2、マルチホップ取得、LLM プロンプト帰属、および著者権検出の 4 つのタスクにわたって、3 つのシードにわたる 17,830 の例と 11 のモデル バリアントをカバーする IWN バリアントは、マクロ F1 \textbf{0.940} (最強の非 IWN ベースラインに対して $+0.036$、著者権検出で $+0.130$) を達成しました。ランダム語彙制御 ($-0.028$ avg.\ F1) は、ゲインがパラメトリックではなく語彙によるものであることを確認します。コード、語彙、および $99.5\%$-recovery 自動抽出レシピがリリースされています。

原文 (English)

SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on external lexical knowledge. We introduce \textbf{\surgellm}, a unified transformer framework that addresses each with a dedicated lightweight module: a \emph{surgical feature gate} (learned per-dimension sigmoid over curated lexical indicators and \texttt{[CLS]}; provably degenerates to identity when features are uninformative), \emph{task-conditioned prefix tokens} (quantized feature values and task identity prepended to every input), and \emph{Instance-Weighted Normalization} (IWN; removes class-prior bias from gate statistics). We prove an excess-risk bound linking gate benefit to \emph{surgical feature alignment}. Across four tasks, SST-2, multi-hop retrieval, LLM-prompt attribution, and authorship detection, covering 17,830 examples and eleven model variants over three seeds, the IWN variant achieves macro-F1 \textbf{0.940} ($+0.036$ over the strongest non-IWN baseline; $+0.130$ on authorship detection). A random-vocabulary control ($-0.028$ avg.\ F1) confirms gains are lexical, not parametric. Code, vocabularies, and a $99.5\%$-recovery auto-extraction recipe are released.

13:00 JST研究/論文

さまざまな車両形状の乱流を予測するためのニューラル ネットワーク ベースのパラメトリック モデル削減

産業用途における数値シミュレーションでは、多くの場合、特定の実験条件によってパラメーター化された多数の高精度計算を実行する必要があります。たとえば、車体設計では、提案されたさまざまな車体形状の空力特性を評価するために空力シミュレーションが不可欠です。ただし、計算リソースの制約がボトルネックになることがよくあります。したがって、計算コストを最小限に抑えながら、望ましい精度を達成することが重要です。この課題に対処するために、物理システムの可能な状態を低次元の部分空間に制限することで自由度を減らすモデル削減手法が開発されました。特に、ニューラルネットワークを用いて系を非線形部分空間に射影する縮小手法が盛んに研究されている。私たちの以前の研究では、ニューラル ネットワーク ベースのモデル削減と時間発展法を統合する削減次数モデルを開発しました。これは、高解像度の流れ場データを効率的に処理するための分散並列トレーニング フレームワークとして実装されました。この研究では、変分オートエンコーダを組み込むことでこの低減アプローチを拡張し、さまざまな形状の複数の車体周囲の高レイノルズ数の流れにおけるロバスト性を評価します。具体的には、特に車体後端付近の流れの挙動に焦点を当て、コンパクトな潜在表現を使用して、さまざまな時空間スケールにわたる渦発生の再構成精度を評価します。

原文 (English)

Neural Network-Based Parametric Model Reduction for Predicting Turbulent Flow for Different Vehicle Geometries

Numerical simulations in industrial applications often require performing numerous high-precision computations parameterized by specific experimental conditions. For instance, in vehicle body design, aerodynamic simulations are essential for evaluating the aerodynamic characteristics of various proposed body geometries. However, computational resource constraints often become a bottleneck. Therefore, achieving the desired accuracy while minimizing computational cost is crucial. To address this challenge, model reduction methods have been developed to decrease the degrees of freedom by constraining the possible states of a physical system to a lower-dimensional subspace. In particular, reduction techniques that project the system onto a nonlinear subspace using neural networks have been actively studied. Our previous research developed a reduced-order model that integrates neural-network-based model reduction with a time-evolution method, implemented as a distributed parallel training framework to process high-resolution flow field data efficiently. In this study, we extend this reduction approach by incorporating a variational autoencoder to assess its robustness in high-Reynolds-number flows around multiple vehicle bodies with varying geometries. Specifically, we evaluate the reconstruction accuracy of vortex generation across different spatial and temporal scales using a compact latent representation, with a particular focus on the flow behavior near the rear end of the vehicle body.

13:00 JSTLLM/生成AI

鳩穴: 不適切なプロンプトはモデルを傷つけ、モデルが崩れたり、間違いを犯したりする

一般に、コンテキスト内学習は大規模言語モデル (LLM) で効果的であることが示されていますが、不適切なコンテキストはパフォーマンスの低下やモードの崩壊を引き起こす可能性があり、これを「ピジョンホール」と呼んでいます。 **意図せず悪い** コンテキストは、悪意のあるジェイルブレイクの意図がなくても発生する可能性があります。たとえば、ユーザーがモデルに間違った数学定理を正当化するように要求したり、モデルのバグのあるコードの修正に失敗したりします。具体的には、(1) ユーザーが解決策を提案したとき、および (2) 会話のコンテキストにアシスタントの以前の (不正確な) 応答が含まれているときの 2 つのシナリオで「ピジョンホール」を調査します。10 個の異なるモデルを使用した 10 個の検証可能なオープンエンドのタスクにわたる実験では、ピジョンホールがいくつかの方法で現れることがわかりました: (1) コンテキストから不正解を繰り返す (38 ~ 40% のパフォーマンス低下につながる)、(2) 狭いセットに収束する(3) ユーザーまたはアシスタントの以前の主張に合わせて、議論の的となっているトピックに対するスタンスを反転する グループ分けは、会話のターン数に応じてほぼ単調に悪化することがわかりました (繰り返される間違いが 1 から 5 に増加するにつれて、パフォーマンスはさらに 14% 以上低下します)。また、提供された例が正しい場合でも、グループ分けに起因するモード崩壊が発生する可能性があります。緩和へのステップとして、モデルを改善する合成エラーを含む RLVR を提案します。バニラ RLVR ベースラインと比較して、不正なコンテキストでは 43 ~ 60%。

原文 (English)

Pigeonholing: Bad prompts hurt models to collapse and make mistakes

While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts can happen without malicious jailbreaking intents: For example, a user asks the model to justify an incorrect math theorem or fails to correct the model's buggy code. Specifically, we investigate ``pigeonholing" in two scenarios: (1) when the user suggests a solution, and (2) when the conversation context includes the assistant's previous (incorrect) responses. Our experiments across 10 verifiable and open-ended tasks with 10 different models show that pigeonholing manifests in several ways: (1) repeating the incorrect answers from context (leading to 38-40% performance drop), (2) converging on a narrow set of answers in coding and text generation without exploring alternatives, and (3) flipping stance on controversial topics to align with the user or the assistant's previous claims. We find that pigeonholing worsens almost monotonically with the number of conversation turns (performance drops by additional 14+% as repeated mistakes increase from 1 to 5), and pigeonholing-induced mode collapse can happen even when the provided example is correct. As a step toward mitigation, we propose RLVR with synthetic errors which improves models by 43-60% under bad contexts compared to vanilla RLVR baselines.

13:00 JSTLLM/生成AI

CALIBER: 言語モデルにおける推論の前後の信頼度の調整

推論言語モデルには、難しい質問に答えるだけでなく、成功の可能性を推定することも求められることが増えています。既存の方法では通常、考える前または答えた後のいずれかに一度だけ自信を引き出します。私たちは、推論モデルの信頼度は状態に依存すると主張します。考える前に、自信はモデルがプロンプトを正しく解決する可能性を推定する必要がありますが、考えた後は、実現された答えが正しい可能性が高いかどうかを予測する必要があります。この区別により、適切な監視対象が決定されます。プロンプトレベルの成功は、プロンプトを見た後に行われた信頼度の推定を監視する必要があり、一方、個々の回答レベルの正しさは、回答後に行われた信頼度の推定を監視する必要があります。 CALIBER (Calibration Before and After Reasoning) を導入します。これは、両方の推定値を導き出し、ターゲットをその情報の状態に合わせて監視します。この統一プロトコルの下で、CALIBER は、7B モデルの BigMathDigits の最も強力な単一信頼度ベースラインよりも予想校正誤差 (ECE) を 52.5% 削減しながら、最高の Brier スコアと AUROC を達成し、最高の精度から 2.1 ポイント以内に留まります。さらに、より大きな 30B モデルでは、CALIBER は BigMathDigits で最高の ECE を達成しながらも、Brier スコアと AUROC で競争力を維持しています。配布以外では、GPQA と TriviaQA で最高の ECE および Brier スコアを達成し、SimpleQA での競争力を維持しています。アブレーションはさらに、この位置とターゲットのアライメントが分布シフトの下で最も有益であることを示しており、分布外のベンチマーク全体にわたってキャリブレーション誤差を一貫して低減します。

原文 (English)

CALIBER: Calibrating Confidence Before and After Reasoning in Language Models

Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after answering. We argue that confidence in reasoning models is state-dependent: before thinking, confidence should estimate the chance of the model correctly solving the prompt, while after thinking it should predict whether the realized answer is likely to be correct. This distinction determines the appropriate supervision target: prompt-level success should supervise confidence estimates made after seeing the prompt, while individual answer-level correctness should supervise confidence estimates made after answering. We introduce CALIBER (Calibration Before and After Reasoning), which elicits both estimates and supervises each with the target matched to its information state. Under this unified protocol, CALIBER reduces Expected Calibration Error (ECE) by 52.5% over the strongest single-confidence baseline on BigMathDigits for the 7B model, while achieving the best Brier score and AUROC, and remains within 2.1 points of the best accuracy. Further, on a larger 30B model, CALIBER achieves the best ECE on BigMathDigits while remaining competitive in Brier score and AUROC. Out of distribution, it achieves the best ECE and Brier score on GPQA and TriviaQA, and remains competitive on SimpleQA. Ablations further show that this position-target alignment is most beneficial under distribution shift where it consistently reduces calibration error across all out-of-distribution benchmarks.

13:00 JST研究/論文

データフリーのストリーミング一貫性蒸留によるリアルタイムのインタラクティブ音楽生成

インタラクティブな音楽やライブ パフォーマンスは、リアルタイムの人間の表現に依存していますが、現代の生成音楽 AI は、その法外な推論遅延とオフライン レンダリング パラダイムのせいで、この分野にはほとんど存在しないままです。パイオニア ミュージシャンにインタラクティブな作曲のための新しいメディアを提供するには、これらの静的なモデルを動的で演奏可能な楽器に根本的に変更する必要があります。本稿では、このギャップを埋めるフレームワークを提案します。構造の一貫性を犠牲にすることなく、ライブ インタラクションに必要な低レイテンシを実現するために、ストリーミング自己回帰潜在空間内で蒸留を定式化します。私たちのアプローチでは、プロンプトのみの入力を利用して教師主導のチャンクごとの軌跡をその場で合成することで、高価なペアの音声潜在データセットの必要性を排除します。生楽器には高い音響忠実度が必要なため、潜在損失、スペクトル損失、時間差損失を組み合わせた音楽を意識した一貫性目標を導入し、加速されたシングルステップストリーミング生成中に音色、トランジェント、リズミカルな安定性などの重要な品質を維持します。パラメータ効率の高い適応を介して実装された当社の蒸留は、生成ステップを削減して、低いリアルタイム係数を実現します。重要なのは、システムが連続的な自己回帰ストリームとして動作することにより、動的人間の入力をオンザフライでシームレスに取り込み、ユーザーがオーディオ フローを中断することなく音楽の軌道を瞬時に操ることができることです。最終的に、この研究はテキストから音楽への生成モデルを受動的プロンプトアンドウェイトシステムとしてではなく、応答性の高い楽器として再文脈化し、人間と AI によるライブ音楽共創の新たな境地を開きます。

原文 (English)

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm. To provide pioneer musicians with a novel medium for interactive composition, we should fundamentally change these static models into dynamic, playable instruments. In this paper, we propose a framework that bridges this gap. To achieve the low latency required for live interaction without sacrificing structural coherence, we formulate distillation within a streaming autoregressive latent space. Our approach gets rid of the need for expensive paired audio-latent datasets by utilizing prompt-only inputs to synthesize teacher-guided, chunk-wise trajectories on the fly. Because live instruments require high acoustic fidelity, we introduce music-aware consistency objectives, which combine latent, spectral, and temporal-difference losses, to preserve crucial qualities like timbre, transients, and rhythmic stability during accelerated single-step streaming generation. Implemented via parameter-efficient adaptation, our distillation reduces generation steps to achieve a low real-time factor. Crucially, by operating as a continuous autoregressive stream, the system can seamlessly assimilate dynamic human inputs on the fly, allowing users to instantly steer the musical trajectory without interrupting the audio flow. Ultimately, this work recontextualizes generative text-to-music models not as passive prompt-and-wait systems, but as responsive instruments, opening new frontiers for live human-AI musical co-creation.

13:00 JST研究/論文

ZONOS2テクニカルレポート

最先端の自然さ、韻律、音声クローン忠実度を実現したTTS最新モデルZONOS2 8Bをご紹介します。 Zonos-v0.1 をスケール、データ、トレーニング レシピ全体にわたって改良しました。新しい専門家混合 (MoE) バックボーンを使用してモデルを合計パラメーター 1.6B から 8B (アクティブ 900M) にスケールし、推論のレイテンシーとスループットを向上させます。新しいデータ処理パイプラインを使用して、トレーニング コーパスを 20 万時間から 600 万時間以上に拡張し、トレーニング後のレシピとコンディショニング レシピを簡素化して、自然さと音声クローン作成の忠実度を向上させています。当社では、品質、スピーカーの類似性、WER、および当社の新しい TTS ベンチマークである ZTTS1-Eval に関して ZONOS2 8B を評価しています。ZTTS1-Eval は、良好なストリーミング遅延を維持しながら最先端のシステムと競合するパフォーマンスを発揮します。モデルの重みと推論コードの例は、GitHub および Hugging Face 上の Apache 2.0 ライセンスに基づいてリリースされています。

原文 (English)

ZONOS2 Technical Report

We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across scale, data, and training recipe. We scale the model from 1.6B to 8B total parameters (900M active) with a novel mixture-of-experts (MoE) backbone, improving inference latency and throughput. We expand our training corpus from 200K to over 6M hours using a new data processing pipeline, and we simplify our post-training and conditioning recipes to improve naturalness and voice cloning fidelity. We evaluate ZONOS2 8B on quality, speaker similarity, WER, and ZTTS1-Eval, our novel TTS benchmark, where it performs competitively with state-of-the-art systems while maintaining good streaming latency. We release our model weights and example inference code under an Apache 2.0 license on GitHub and Hugging Face.

13:00 JST研究/論文

ODRLとは何を意味しますか? UFO-L における許可、禁止、義務のクロスレベルの存在論的根拠

ODRLの政策評価者は評決を下すが、政策がもたらす規範的立場、その立場が前提とする権威構造、規範違反を宣言する権限を誰が持つかについては何も語らない。私たちはクロスレベルの設計原則を策定します。違反する結果的な規範を伴う規範的な言語には、行為レベルの立場 (許可、義務、権利、権利なし) と能力レベルの立場 (権力、服従、免除、障害) の両方が必要です。これを ODRL に適用すると、禁止は認可されている (違反の可能性と結果的)、許可はその動作パラメーター (オープンな世界とクローズドな世界) 全体で過少指定されていること、そして形式的な意味論は達成義務のみをカバーしていることが確立されます。 We ground ODRL in UFO-L, mapping each activated rule to a simple legal relator and extending coverage from two to eight legal positions; violation-declaration authority, implicit in every existing evaluator, becomes an explicit Power-Subjection pair. All axioms are mechanically verified in Isabelle/HOL and across a 39-problem benchmark under Vampire, E, and Z3.

原文 (English)

What Does ODRL Mean? A Cross-Level Ontological Grounding of Permissions, Prohibitions, and Duties in UFO-L

ODRL policy evaluators produce verdicts, but say nothing about the normative positions a policy brings into existence, the authority structures those positions presuppose, or who holds the power to declare a norm violated. We formulate the Cross-Level Design Principle: any normative language with violable, consequential norms requires both conduct-level positions (Permission, Duty, Right, No right) and competence-level positions (Power, Subjection, Immunity, Disability). Applying this to ODRL, we establish that prohibition is sanctioned (violation possible and consequential), that permission is underspecified across its behaviour parameter (open vs. closed world), and that the formal semantics covers achievement obligations only. We ground ODRL in UFO-L, mapping each activated rule to a simple legal relator and extending coverage from two to eight legal positions; violation-declaration authority, implicit in every existing evaluator, becomes an explicit Power-Subjection pair. All axioms are mechanically verified in Isabelle/HOL and across a 39-problem benchmark under Vampire, E, and Z3.

13:00 JST画像/動画生成

構造的コルモゴロフ・アーノルド畳み込み: エッジごとの畳み込み KAN のパラメーター効率の高い代替としての値またはフィルター形状の学習可能な関数

畳み込みコルモゴロフ -- アーノルド ネットワーク (KAN) は、畳み込みカーネルの固定重みを学習可能な一変量関数に置き換えます。主流の定式化では、このような関数を 1 つすべてのカーネル エントリに付加し、それをピクセル値に作用させます。表現力は高いものの、パラメータが多く、過学習が起こりやすいものです。私たちは、学習可能な関数は畳み込みの各エッジよりも \emph{struction} に配置する方が適切であると主張し、関数がピクセル \emph{values} に作用するかフィルター \emph{shape} に作用するかという 1 つの軸に沿って設計空間を編成します。私たちは 3 つの実現を研究します。 SV-KAN は、1 つの共有一変量関数を値に適用し、空間フィルターを自由かつ静的なままにし、単一の学習可能な共有アクティベーションによる古典的な畳み込みを行います。 AG-KAN は共有価値機能を維持しますが、内容適応型ガウス ゲートを通じて空間構造を提供します。代わりに、RF-KAN は学習可能な関数をフィルター形状に移動し、コンテンツ適応振幅を持つ局所振動 (Morlet) ウェーブレット ベースで拡張された指向性リッジ プロファイルから各フィルターを構築します。インラン参照と 3 つのシードを備えた一致した 4 層プロトコルの下では、RF-KAN と SV-KAN は、CIFAR-10 では $88.47\pm0.10\%$ と $88.20\pm0.31\%$、CIFAR-100 では $64.40\pm0.19\%$ と $64.57\pm0.30\%$ に達し、約 $0.4$M になります。パラメータ。この一致したスケールでは、形状モデルと最も単純な値モデルは、単純な畳み込みと、公式のグラム バリアントを含むテストしたすべてのエッジごとの KAN の両方の上で、パラメーターの約 5 分の 1 で一致します。対照研究では、RF-KANの利得は本質的に局所的な振動基底と内容の適応性に起因しており、学習された形状を完全に削除して共有価値関数のみを残すアブレーションでは、精度が40ポイント以上低下し、学習された形状がこのスケールでの耐荷重成分であることが特定されています。

原文 (English)

Structural Kolmogorov-Arnold Convolutions: Learnable Function on the Values or the Filter Shape as Parameter-Efficient Alternative to Per-Edge Convolutional KANs

Convolutional Kolmogorov--Arnold Networks (KANs) replace the fixed weights of a convolutional kernel with learnable univariate functions. The dominant formulation attaches one such function to every kernel entry and lets it act on pixel values, expressive but parameter-heavy and prone to overfitting. We argue that the learnable functions are better placed in the \emph{structure} of the convolution than on each edge, and we organise the design space along a single axis: whether the function acts on the pixel \emph{values} or on the filter \emph{shape}. We study three realisations. SV-KAN applies one shared univariate function to the values and leaves the spatial filter free and static, aa classical convolution with a single learnable shared activation. AG-KAN keeps the shared value function but supplies the spatial structure through a content-adaptive Gaussian gate. RF-KAN instead moves the learnable functions onto the filter shape, building each filter from oriented ridge profiles expanded in a localised oscillatory (Morlet) wavelet basis with content-adaptive amplitudes. Under a matched four-layer protocol with in-run references and three seeds, RF-KAN and SV-KAN reach $88.47\pm0.10\%$ and $88.20\pm0.31\%$ on CIFAR-10 and $64.40\pm0.19\%$ and $64.57\pm0.30\%$ on CIFAR-100, at about $0.4$M parameters. At this matched scale the shape model and the simplest value model meet at the top, both above a plain convolution and every per-edge KAN we tested, including the official Gram variant, at roughly a fifth of the parameters. A controlled study attributes the RF-KAN gain to an intrinsically localised oscillatory basis and to content adaptivity, and an ablation that removes the learned shape entirely, leaving only the shared value function, collapses accuracy by over forty points, identifying the learned shape as the load-bearing ingredient at this scale.

13:00 JSTLLM/生成AIビジネス/資金調達

大規模言語モデル評価におけるプロンプトランキングの安定性について

プロンプトベースの対話は、大規模言語モデル (LLM) を使用するための主要なパラダイムとなっています。LLM では、複数の候補プロンプトが評価され、最上位のプロンプトが下流で使用するために選択されます。このワークフローは、評価条件が多少変動してもプロンプト ランキングが安定していることを暗黙的に前提としています。この論文では、ランダム シードや限定された評価サブセットなどの一般的な変動要因の下でのプロンプト ランキングの安定性を体系的に研究します。 3 つのオープンウェイト LLM と 2 つのベンチマーク タスクにわたって、全体的なランク相関は多くの場合中程度から高である一方で、最もパフォーマンスの高いプロンプトのアイデンティティは頻繁に変化し、信頼性の低い選択決定につながることがわかりました。この問題に対処するために、パフォーマンスと分散の両方を考慮した下限信頼限界に基づいた、単純な安定性を意識した選択戦略を提案します。私たちの結果は、このアプローチが不安定な環境での堅牢性を向上させながら、より安定した環境で競争力を維持できることを示しています。これらの調査結果は、即時選択と LLM ベンチマークにおける評価の不確実性を考慮することの重要性を強調しています。

原文 (English)

On the Stability of Prompt Ranking in Large Language Model Evaluation

Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicitly assumes that prompt rankings are stable under minor variations in evaluation conditions. In this paper, we systematically study prompt ranking stability under common sources of variability, including random seeds and limited evaluation subsets. Across three open-weight LLMs and two benchmark tasks, we find that while overall rank correlations are often moderate to high, the identity of the top-performing prompt frequently changes, leading to unreliable selection decisions. To address this issue, we propose a simple stability-aware selection strategy based on a lower confidence bound, which accounts for both performance and variance. Our results show that this approach improves robustness in unstable settings while remaining competitive in more stable regimes. These findings highlight the importance of accounting for evaluation uncertainty in prompt selection and LLM benchmarking.

13:00 JST画像/動画生成

Female-RHINO: 自動定量子宮 MRI 分析と構造化レポートのためのリアルタイム スキャナー統合フレームワーク

子宮 MRI の標準化された評価は、解剖学的多様性、観察者の依存性、およびワークフローに統合された自動分析ツールの不足により、依然として困難です。この研究では、画像取得中の自動定量的子宮 MRI 分析と構造化レポートのためのリアルタイム AI 支援フレームワークである Female-RHINO: (R)eproduction (H)ealth (I)maging A(N)amination T(O)ol を紹介します。我々は、MRI スキャナーとのインライン通信と深層学習ベースの分析を統合して、矢状 T2 強調骨盤 MRI から定量的な子宮バイオマーカーを導出するエンドツーエンド システムを紹介します。このフレームワークは、さまざまなプロトコル、ベンダー、患者集団にわたる 500 以上の多施設データセットでトレーニングおよび評価されたセグメンテーションと解剖学的ランドマーク検出モデルを組み合わせています。体積測定を実行し、筋腫やナボシアン嚢胞などの一般的な偶発的所見を検出して定量化し、生体評価のための 6 つの解剖学的ランドマークを抽出します。結果は、手動操作を必要とせずに、統合された視覚化を備えた構造化された臨床医向けのレポートにまとめられます。独立した遡及コホートおよび前向きコホートの評価により、さまざまな取得設定にわたって堅牢なパフォーマンスが実証されました。平均 Dice 類似性係数は、子宮では 0.82、筋腫では 0.80 で、ナボシアン嚢胞では低いものの一貫して一致しました。ランドマーク検出では、平均半径誤差 3.7 mm を達成しました。エンドツーエンドの処理は 70 秒未満で完了し、進行中のスキャン中に結果を利用できるようになりました。将来的な展開により、観察者間の合意に裏付けられた、標準化された再現可能な分析が即座に得られました。提案されたシステムは、自動化された子宮 MRI 分析とレポートのためのリアルタイム スキャナー統合 AI を可能にし、骨盤イメージングにおける標準化、効率、および臨床ワークフローを改善する可能性があります。

原文 (English)

Female-RHINO: A Real-Time Scanner-Integrated Framework for Automated Quantitative Uterine MRI Analysis and Structured Reporting

Standardized assessment of uterine MRI remains challenging due to anatomical variability, observer dependence, and the lack of workflow-integrated automated analysis tools. This work presents Female-RHINO: (R)eproductive (H)ealth (I)maging A(N)alysis T(O)ol, a real-time AI-assisted framework for automated quantitative uterine MRI analysis and structured reporting during image acquisition. We present an end-to-end system that integrates inline communication with the MRI scanner and deep learning-based analysis to derive quantitative uterine biomarkers from sagittal T2-weighted pelvic MRI. The framework combines segmentation and anatomical landmark detection models trained and evaluated on more than 500 multi-center datasets spanning diverse protocols, vendors, and patient populations. It performs volumetry, detects and quantifies common incidental findings such as fibroids and Nabothian cysts, and extracts six anatomical landmarks for biometric assessment. Results are compiled into a structured clinician-oriented report with integrated visualizations, without manual interaction. Evaluation on independent retrospective and prospective cohorts demonstrated robust performance across varying acquisition settings. Mean Dice similarity coefficients were 0.82 for the uterus and 0.80 for fibroids, with lower but consistent agreement for Nabothian cysts. Landmark detection achieved a mean radial error of 3.7 mm. End-to-end processing was completed in under 70 seconds, enabling availability of results during the ongoing scan. Prospective deployment yielded immediate, standardized, and reproducible analyses supported by inter-observer agreement. The proposed system enables real-time scanner-integrated AI for automated uterine MRI analysis and reporting, with potential to improve standardization, efficiency, and clinical workflow in pelvic imaging.

13:00 JSTロボティクス研究/論文

平均ランキングによる被験者ごとの最適性のマスク: EEG 運動画像 BCI デコーダのフリードマン-ネメニ ベンチマーク

脳波検査 (EEG) はブレイン コンピューター インターフェイス (BCI) の非侵襲性モダリティとして主流ですが、運動イメージの信頼性の高いデコードは個人間および個人内のばらつきによって妨げられています。繰り返し主張されるのは、1 つのデコード パイプライン (ほとんどの場合、空間法またはリーマン法) が広く望ましいということです。私たちはその主張の最も弱いバージョンを最も有利な条件下でテストします。 Mother of All BCI Benchmarks (MOABB) フレームワークを使用して、3 つの公開左右運動画像データセット (PhysionetMI、参加者 109 人、Cho2017、52 人、Zhou2016、4 人) および 2 つの周波数帯域 (8 ~ 15 人) にわたって、1,056 のデコード構成 (特徴抽出器 x スケーラー x 分類子)、340,000 を超える被験者レベルのモデルの適合を評価しました。 Hz、8 ~ 30 Hz)。すべてのモデルは、単一の参加者の単一セッション内で適合およびテストされます。これは最も簡単な体制であり、すべてのパイプラインに最善のチャンスが与えられます。複数の分類子の比較には統計標準、つまりフリードマンオムニバステスト、ネメニ臨界差分分析、および効果量を使用したウィルコクソン符号付き順位テストを適用します。共分散接線空間投影 (cov-tgsp) と共通空間パターン (CSP) は最も強力なファミリーですが、それらの順序付けはデータセットに依存しており、最大かつ最も不均一なコホート (PhysionetMI) では統計的に区別できません (Nemenyi p = 0.27; Kendall の W = 0.11)。個人レベルでは、単一の最適なパイプラインは PhysionetMI 参加者の 35% のみに最適であり、非線形記述子は約 3 分の 1 に最適です。パイプラインを参加者に一致させると、最適な固定選択よりも約 7 精度ポイントが追加されます。ランク付けは次元の成果物ではなく、分類器とスケーラーの選択は特徴表現に次ぐものです。最も簡単な体制であっても、単一のパイプラインが支配することはありません。つまり、パーソナライゼーションの問題の下限と、ユニバーサル デコーダーではなく参加者を意識したモデル選択の定量的なケースです。

原文 (English)

Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders

Electroencephalography (EEG) is the dominant non-invasive modality for brain-computer interfaces (BCIs), yet reliable decoding of motor imagery is hampered by inter- and intra-individual variability. A recurring claim is that one decoding pipeline, most often a spatial or Riemannian method, is broadly preferable. We test the weakest version of that claim under the most favourable conditions. Using the Mother of All BCI Benchmarks (MOABB) framework, we evaluated 1,056 decoding configurations (feature extractor x scaler x classifier), >340,000 subject-level model fits, across three public left-versus-right motor-imagery datasets (PhysionetMI, 109 participants; Cho2017, 52; Zhou2016, 4) and two frequency bands (8-15 Hz, 8-30 Hz). Every model is fit and tested within a single session of a single participant, the easiest regime, giving every pipeline its best chance. We apply the statistics standard for multi-classifier comparison: Friedman omnibus tests, Nemenyi critical-difference analysis and Wilcoxon signed-rank tests with effect sizes. Covariance tangent-space projection (cov-tgsp) and Common Spatial Patterns (CSP) are the strongest families, but their ordering is dataset-dependent and, on the largest and most heterogeneous cohort (PhysionetMI), statistically indistinguishable (Nemenyi p = 0.27; Kendall's W = 0.11). At the individual level the single best pipeline is optimal for only 35% of PhysionetMI participants, and nonlinear descriptors are best for roughly one third; matching pipeline to participant adds about seven accuracy points over the best fixed choice. The ranking is not an artefact of dimensionality, and classifier and scaler choices are secondary to the feature representation. Even in the easiest regime, no single pipeline dominates: a lower bound on the personalization problem and a quantitative case for participant-aware model selection rather than a universal decoder.

13:00 JST研究/論文

バッチ処理された Oracle クエリによるエンティティ解決

一度に限られたレコードのバッチを処理し、同じ現実世界のエンティティを参照するレコードをクラスタリングするオラクルについて考えます。私たちは、そのようなオラクルに問い合わせて、サイズが単一のバッチよりもはるかに大きく、特定のエンティティのすべてのレコードが含まれるバッチが保証されていないデータセット内のエンティティを解決する方法を研究します。当社は、コスト (オラクル コンサルトの数) を完全に制御しながら、すべての段階で可能な限り最高のリコールを達成するための従量課金制のアプローチを目指しています。私たちはこの問題をバッチエンティティ解決として正式にキャストし、最適なバッチを選択することが NP 困難であることを証明し、エンティティ サイズの自然条件下で最適な解決策を提供します。最後に、6 つのデータセットでアプローチを評価し、最先端のベースラインよりも優れていることを示します。

原文 (English)

Entity Resolution via Batched Oracle Queries

We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We study how to interrogate such an oracle to resolve entities in a dataset whose size is far larger than a single batch, and where no batch is guaranteed to contain all records of any given entity. We aim at a pay-as-you-go approach, to have full control over the costs (the number of oracle consults), while achieving the highest possible recall at every step. We formally cast this problem as batched entity resolution, prove that selecting optimal batches is NP-hard, and provide an optimal solution under a natural condition on entity sizes. Finally, we evaluate our approach on six datasets and show its superiority over state-of-the-art baselines.

13:00 JSTエージェントClaude

オープンソースでの AI コーディング エージェントの検出: 1 億 8,000 万のリポジトリを対象とした検証済みの複数方法の調査

生成 AI コーディング エージェントはオープンソースのサプライ チェーンに参入しつつありますが、その痕跡は多様で目に見えないことが多いため、その普及状況はほとんど理解されていません。 World of Code (1 億 8,000 万以上の Git リポジトリ) にわたる構成ファイルのスキャン、コミット メッセージ分析、作成者 ID の照合、ボット署名の検索を統合する多層検出フレームワークを導入し、エージェント トレースを 4 つの動作タイプに分類します。単一のメソッドでは、アクティビティの一部以上をキャプチャすることはできません。マルチメソッドの検出では、1 つのスナップショットで 850,157 件のクロード コードのコミットが特定されますが、そのうちボット アカウントのルックアップ - ほとんどの導入調査が依存するシグナル - 回復するのは 28,154 (3.3%) のみで、相対再現率のギャップは 30 倍であるため、単一シグナルの蔓延推定値は、少なくともこの要因によって低くバイアスされます。すべての検出パターンは、セルごとの精度とウィルソン信頼区間で手作業で検証されます (495 ラベル)。 2024 年 12 月から 2026 年 4 月までのスナップショット全体で、コミット属性のエージェントは毎月 320,000 件を超えるコミットを生成します。 Claude Code がリードし (17,295 プロジェクトで 886,122 コミット)、サイレントな構成ファイルのみの採用 (21,078 プロジェクト) を支配しています。独立したプルリクエスト国勢調査 (AIDev) と比較すると、2 つのチャネルはほぼばらばらのエージェント集団をキャプチャしています。PR 国勢調査では、コミットが検出されたクロード コード採用者の 79% と基本的にすべての Codex 採用者が見逃されています。また、さまざまな種類の作業が含まれています。PR 展開されたクラウド エージェント (Codex、Cursor) は機能作業として表示され、コミット展開されたエディタ内エージェント (Claude Code、OpenHands、Aider) はメンテナンスとして表示されます。観察された作業プロファイルは、ツール自体ではなく展開および検出モードに従うため、単一のチャネルが代表的なものではありません。

原文 (English)

Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories

Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence poorly understood. We introduce a multi-layered detection framework that integrates configuration-file scanning, commit-message analysis, author-identity matching, and bot-signature lookup across World of Code (180M+ Git repositories), classifying agent traces into four behavioral types. No single method captures more than a fraction of activity: multi-method detection identifies 850,157 Claude Code commits in one snapshot, of which bot-account lookup_the signal most adoption studies rely on_recovers only 28,154 (3.3%), a 30x relative-recall gap, so single-signal prevalence estimates are biased low by at least this factor. Every detection pattern is hand-validated (495 labels) with per-cell precision and Wilson confidence intervals. Across snapshots from December 2024 to April 2026, commit-attributed agents generate over 320,000 commits per month; Claude Code leads (886,122 commits across 17,295 projects) and dominates silent, configuration-file-only adoption (21,078 projects). Compared against an independent pull-request census (AIDev), the two channels capture nearly disjoint agent populations_a PR census misses 79% of commit-detected Claude Code adopters and essentially all Codex adopters_and different kinds of work: PR-deployed cloud agents (Codex, Cursor) surface as feature work, while commit-deployed in-editor agents (Claude Code, OpenHands, Aider) surface as maintenance. The observed work profile follows deployment and detection mode rather than the tool itself, so no single channel is representative.

13:00 JST画像/動画生成

潜在空間における画像の変形挙動

組織病理学分類タスクのためのニューラル ネットワークのトレーニングは通常、潜在空間へのデータ エンコードに依存しており、これにより複雑さが軽減され、パフォーマンスが向上します。 ImageNET などの一般的な画像データセット、または特に組織病理学的画像で事前トレーニングされた、いくつかのエンコーダ ネットワークが利用可能です。エンコーダー ネットワークのトレーニングは、下流のタスクに適応させて、ラベルに関係のない変換に対してネットワークを不変にしながら、生物学的/診断コンテンツのエンコードを可能にする必要があります。この論文では、病理学的画像に焦点を当てている Lunit Inc. と Bioptimus が提供するネットワークと、Meta Research Team が提供するネットワークを使用して、潜在空間に対する古典的な画像変換の効果を調査します。結腸直腸組織データセットおよび公的にアクセス可能な TCGA データセットで利用可能なヘマトキシリン/エオシン染色切片の画像タイルを使用して、元の画像埋め込みと変換後の画像埋め込みを比較し、ランダムで無関係な埋め込みと対比することにより、標準的なデータ変換から生じる埋め込みの分散を評価します。私たちの調査結果は、元の画像と変換された画像の埋め込みがランダムな埋め込みよりも相互に近く、変換に対する堅牢性を示していることを示しています。ただし、それらは完全に不変ではなく、エンコーダー ネットワークが潜在空間の変換効果を完全に無効化していないことが明らかになり、変換を介したデータセットの拡張がパフォーマンスを向上できる理由を説明しています。一般的なエンコーダー ネットワークと組織病理学固有のエンコーダー ネットワークの間には、有意な違いが観察されました。

原文 (English)

Transformation Behavior of Images in Latent Space

Training of neural networks for histopathology classification tasks typically relies on data encoding into latent space, which reduces complexity and improves performance. There are several encoder networks available, either pretrained on general image datasets such as ImageNET, or specifically on histopathological images. Training of encoder networks should be adapted to downstream tasks, allowing encoding of biologic/diagnostic content while rendering networks invariant to label-irrelevant transformations. This paper investigates the effect of classical image transformation on the latent space, using networks provided by Lunit Inc. and Bioptimus, both focusing on pathological images, and by Meta Research Team. We assess variance of embeddings resulting from standard data transformations by comparing original and transformed image embeddings and by contrasting them with random, unrelated embeddings, using image tiles from hematoxylin/eosin-stained sections available in a colorectal tissue dataset and the publicly accessible TCGA dataset. Our findings show that embeddings of original and transformed images are closer to each other than to random embeddings, indicating robustness to transformations. However, they are not fully invariant, revealing that the encoder networks do not completely neutralize transformation effects in latent space, explaining why transformation-mediated augmentation of datasets can improve performance. Significant differences were observed between general and histopathology-specific encoder networks.

13:00 JST画像/動画生成

MedPCFM: ポイントトランスフォーマーとフローマッチングの統合による医療点群の完成度の向上

医療点群の完成は、解剖学的再構築と下流の臨床ワークフローにとって重要ですが、この設定における生成モデリングは依然として十分に研究されていません。私たちは連続時間生成モデリングを通じて完成を調査し、医療点群を完成させるための PTv3 を利用したフロー マッチング アプローチである PCFM を導入します。私たちは SkullFix と SkullBreak を評価し、さらに最近の下顎欠損データセットも評価します。 PTv3 を決定論的なエンコーダ/デコーダ補完モデルに適応させ、PVCNN と PTv3 デノイザの両方で拡散補完 (PCDiff) をインスタンス化することにより、強力なベースラインを構築します。 PTv3 を使用した PCFM は、決定論的な PTv3 ベースラインと競合し、データセット全体で最先端の生成パフォーマンスを達成しながら、必要なサンプリング ステップは拡散よりも大幅に少なくなります。最適な動作ポイントでは、PTv3 は明らかなスループットの向上ももたらし、PVCNN バックボーンと比較して PCFM に最大 7 倍の高速化をもたらします。最後に、モデル サイズとポイント カーディナリティを変化させることによって経験的なスケーリング傾向を研究し、より高いポイント解像度とモデル スケール全体での有益なトレードオフによる一貫したゲインを示します。

原文 (English)

MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and Flow Matching

Medical point cloud completion is important for anatomical reconstruction and downstream clinical workflows, yet generative modeling in this setting remains insufficiently studied. We investigate completion through continuous-time generative modeling and introduce PCFM, a PTv3-backed flow matching approach for medical point cloud completion. We evaluate on SkullFix and SkullBreak, and additionally on the more recent Mandibular Defect dataset. We build strong baselines by adapting PTv3 to a deterministic encoder-decoder completion model and by instantiating diffusion completion (PCDiff) with both PVCNN and PTv3 denoisers. PCFM with PTv3 is competitive with the deterministic PTv3 baseline and achieves state-of-the-art generative performance across datasets, while requiring substantially fewer sampling steps than diffusion. At the best operating points, PTv3 also yields clear throughput gains, providing up to a 7$\times$ speed-up for PCFM compared to a PVCNN backbone. Finally, we study empirical scaling trends by varying model size and point cardinality, showing consistent gains with higher point resolution and informative trade-offs across model scales.

13:00 JSTロボティクス

NoContactNoWorries: 手の器用な操作のための視覚と固有受容による接触の推定

身体的接触を認識することは、器用な操作の基本です。ロボットは多くの場合、専用のハードウェア触覚センサーに依存しますが、人間は視覚情報と体の姿勢や動きの生来の感覚を統合することで、接触を推測する驚くべき能力を発揮します。この身体化された知覚スキルに触発されて、私たちはロボットが視覚から接触を推測する方法を学習できるかどうかを調査します。このアプローチは、コスト、脆弱性、統合の点で現実的な課題に直面している、特にバイナリ接触推定のための触覚ハードウェアに代わるスケーラブルな代替手段も提供します。我々は、RGB-D 視覚とロボットの固有受容を融合して、手と物体の相互作用の疑似触覚信号としてバイナリ接触状態を推測する、トランスフォーマー ベースのマルチモーダル フレームワークである NoContactNoWorries を紹介します。複数の物体に対して単一の接触予測モデルをトレーニングすることで検証し、推定された接触信号が、新しい物体に一般化して、手の中の物体の向きを変えるための下流の強化学習エージェントをサポートすることを示します。シミュレーションと現実世界のロボットの両方での実験により、私たちのアプローチが検証され、視覚と固有受容から接触を推測する実現可能性が強調されました。プロジェクトページ: https://soham2560.github.io/no-contact-no-worries/

原文 (English)

NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation

Perceiving physical contact is fundamental to dexterous manipulation. While robots often rely on dedicated hardware tactile sensors, humans exhibit a remarkable ability to infer contact by integrating visual information with an innate sense of their body's pose and movement. Inspired by this embodied perceptual skill, we investigate whether a robot can learn to infer contact from vision, an approach that also offers a scalable alternative to tactile hardware specifically for binary contact estimation, which faces practical challenges in cost, fragility, and integration. We present NoContactNoWorries, a transformer-based multimodal framework that fuses RGB-D vision with the robot's proprioception to infer binary contact states as a pseudo-tactile signal for hand-object interactions. We validate by training a single contact prediction model on multiple objects and show that the inferred contact signal supports downstream reinforcement learning agents for in-hand object reorientation, generalizing to novel objects. Experiments in both simulation and on a real-world robot validate our approach, highlighting the feasibility of inferring contact from vision and proprioception. Project Page: https://soham2560.github.io/no-contact-no-worries/

13:00 JSTLLM/生成AIGPT / ChatGPTGemma

アフリカ言語税: フロンティア LLM におけるアフリカ言語のトークン化のコスト、レイテンシ、およびコンテキスト ペナルティの定量化

商用の大規模言語モデルでは、トークンごとに請求、スケール レイテンシ、および予算コンテキストが設定されます。しかし、トークナイザーは、一部の言語では他の言語よりも多くのサブワード トークンを同じ意味に割り当てるため、トークンの生産性が高い言語の話者は、モデルが呼び出される前に構造的なペナルティを支払うことになります。このペナルティは多言語設定全般について文書化されていますが、アフリカの言語については、企業展開の経済性や認知コンテキスト能力のレベルで体系的に測定されていません。言語効果が内容から分離されるように、並列コーパスを使用して、5 つの言語族と 3 つの文字にまたがる 20 のアフリカ言語 (ラテン語、Ge'ez/エチオピア語、N'Ko。19 言語は一次 FLORES-200+ コーパスに含まれており、ナイジェリアのピジンは MAFAND-MT のみで測定) にわたって測定しました。 FLORES-200+ の 11 のフロンティアおよびオープン トークナイザー全体で、すべてのアフリカ言語は英語よりもトークナイゼーション プレミアムをもたらします (GPT-5 / o200k_base では中央値 1.88 倍、N'Ko では最大 8.92 倍)。ペナルティはエチオピア文字と N'Ko 文字で最も大きく (7 ~ 9 倍に達します)、コーパス全体でほぼ不変です (FLORES 対 SIB-200 ピアソン r = 0.9998)。デプロイメント条件に換算すると、推論コストが最大 8.9 倍となり、同等の生成遅延乗数 (GPT-5 では N'Ko 対英語、アムハラ語では 7.4 倍)、英語の有効コンテキスト ウィンドウはわずか 11% になります。現在利用可能なアフリカ言語用の最良のトークナイザーである Gemma 4 は、平均プレミアムを 3.31 倍 (cl100k_base) から 2.38 倍に削減しますが、ペナルティを排除するトークナイザーはありません。私たちは、オープンな測定ツール (アフリカの妊孕性)、公開リーダーボード、結果データセット、およびアフリカの建設業者向けの緩和ガイダンスをリリースします。このペナルティは、サブワード語彙に直接エンコードされたデジタルデバイドである、話者がペナルティを支払う余裕のない言語に最も重くのしかかります。

原文 (English)

The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs

Commercial large language models bill, scale latency, and budget context per token. Yet tokenizers assign more subword tokens to the same meaning in some languages than in others, so speakers of languages with high token-fertility pay a structural penalty before a model is ever invoked. This penalty is documented for multilingual settings in general, but it has not been measured systematically for African languages at the level of enterprise deployment economics and cognitive context capacity. We measure it across 20 African languages spanning five language families and three scripts (Latin, Ge'ez/Ethiopic, N'Ko; 19 appear in the primary FLORES-200+ corpus, with Nigerian Pidgin measured via MAFAND-MT only), using parallel corpora so that the language effect is isolated from content. Across 11 frontier and open tokenizers on FLORES-200+, every African language carries a tokenization premium above English (median 1.88x on GPT-5 / o200k_base, up to 8.92x for N'Ko); the penalty is largest for Ethiopic and N'Ko scripts (reaching 7-9x) and is near-invariant across corpora (FLORES vs SIB-200 Pearson r = 0.9998). Translated into deployment terms, this results in up to 8.9x inference cost and an equivalent generation-latency multiplier (N'Ko vs English on GPT-5; 7.4x for Amharic), and as little as 11% of English's effective context window. The best currently available tokenizer for African languages, Gemma 4, reduces the mean premium from 3.31x (cl100k_base) to 2.38x, but no tokenizer eliminates the penalty. We release an open measurement tool (afri-fertility), a public leaderboard, a results dataset, and mitigation guidance for African builders. The penalty falls hardest on the languages whose speakers can least afford it, a digital divide encoded directly into the subword vocabulary.

13:00 JSTロボティクスGoogle

G$^3$VLA: 視覚・言語・行動モデルの幾何学的帰納バイアス

視覚言語アクション (VLA) モデルは、事前学習された視覚言語バックボーンからの意味論的な知識を利用することによって、汎用的なロボット操作において急速な進歩を遂げましたが、その視覚トークンは、ロボットのカメラのキャリブレーションされたジオメトリではなく、2D 画像座標に基づいたままです。この不一致は、ビューが既知の内部機能と外部機能によって結合されているにもかかわらず、独立した画像として処理されるマルチカメラ設定で特に顕著です。我々は、アクション空間や模倣目的を変更することなく、事前学習済み VLA のビジュアル トークン ストリームにキャリブレーションされた構造を注入するカメラ認識幾何学モジュールである G$^3$VLA を提案します。これは、固有条件付きレイ埋め込み、射影位置エンコーディング (PRoPE)、および双方向クロスビュー フュージョンを組み合わせたものです。幾何学的監視は、利用可能な場合はグラウンドトゥルース ポイント マップから、または信頼度ゲート $\pi^3$X 教師予測から提供され、深度センサーや手動の注釈は必要ありません。 $\pi_0$ でインスタンス化された G$^3$VLA は、LIBERO スイート、RoboCasa24、RoboTwin2.0、および実際のロボット設定全体で一貫した利益をもたらし、空間的およびオブジェクトに敏感なタスクで最大の改善をもたらします。 $\pi_{0.5}$ と GR00T 1.5 についてさらに検証し、その結果、ジオメトリ認識トークンがアクション生成経路に直接アクセスできる場合に幾何学的転送が最も効果的であることが示唆されました。私たちのプロジェクト ページは https://sites.google.com/view/g3vla にあります。

原文 (English)

G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models

Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent images. We propose G$^3$VLA, a camera-aware geometric module that injects calibrated structure into the visual-token stream of a pretrained VLA without altering its action space or imitation objective, combining intrinsic-conditioned ray embeddings, projective positional encoding (PRoPE), and bidirectional cross-view fusion. Geometric supervision is provided either from ground-truth point maps when available, or from confidence-gated $\pi^3$X teacher predictions, requiring no depth sensors or manual annotations. Instantiated on $\pi_0$, G$^3$VLA yields consistent gains across the LIBERO suites, RoboCasa24, RoboTwin2.0, and real-robot settings, with the largest improvements on spatially and object-sensitive tasks. We further validate on $\pi_{0.5}$ and GR00T 1.5, with results suggesting that geometric transfer is most effective when geometry-aware tokens have direct access to the action generation pathway. Our project page is at https://sites.google.com/view/g3vla

13:00 JSTLLM/生成AI画像/動画生成

video-SALMONN-R$^3$: ビデオを効率的に理解するための再視聴、再質問、再回答の学習

ビデオ大規模言語モデル (LLM) は、多くの場合、計算量とメモリ バジェットによって制限されるため、使用するフレーム レートと空間解像度が低下し、質問応答 (QA) に必要な重要な情報が失われる可能性があります。実用的で効率的なソリューションは 2 段階のパラダイムです。最初に大まかなビデオ理解を実行して関連セグメントの位置を特定し、次にこれらのセグメントをより高い時間的または空間的忠実度で再視聴します。この論文では、思考連鎖 (CoT) のコールドスタートに依存せずに、強化学習を通じて再視聴を可能にする初のエンドツーエンドのビデオ LLM である video-SALMONN-R$^3$ を紹介します。この設計により、コストのかかる CoT データ アノテーションの必要性がなくなり、事前トレーニングされたビデオ理解能力を低下させる可能性がある CoT ベースの教師あり微調整 (SFT) が回避されます。再視聴によって誘発される推論優先の動作と、事前トレーニングされたビデオ LLM の回答優先の傾向との間の不一致に対処するために、モデルが最初の視聴で直接の回答を生成し、再視聴後にそれを改良する再回答戦略を提案します。最後に、再視聴中の質問の遵守性を向上させるために、ローカライズされたセグメントを再訪問するときにクエリを再挿入する再質問メカニズムを提案します。実験結果は、video-SALMONN-R$^3$ が基本モデルと QA-SFT ベースラインの両方を一貫して上回っており、大幅に低い計算コストで以前の再視聴ベースのアプローチを上回っていることを示しています。コード、モデル、データは受理され次第公開されます。

原文 (English)

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA). A practical and efficient solution is a two-stage paradigm: first perform coarse video understanding to localize relevant segments, and then re-watch these segments at higher temporal or spatial fidelity. In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. This design removes the need for costly CoT data annotations and avoids CoT-based supervised fine-tuning (SFT), which can otherwise degrade the pretrained video understanding abilities. To address the mismatch between the reasoning-first behavior induced by re-watch and the answer-first tendency of pretrained video-LLMs, we propose a re-answer strategy, in which the model first produces a direct answer in the first watch and then refines it after re-watching. Finally, to improve question adherence during re-watching, we propose a re-ask mechanism that re-injects the query when revisiting localized segments. Experimental results show that video-SALMONN-R$^3$ consistently outperforms both the base model and the QA-SFT baseline, while surpassing prior re-watch-based approaches with significantly lower computational cost. Code, models, and data will be publicly released upon acceptance.

13:00 JST研究/論文

O-RAN における UAV 軌道最適化のための適応機械学習フレームワーク

6G セルラー システムにおけるオープン無線ユニット (O-RU) としての無人航空機 (UAV) の展開は、スケーラブルで適応性のあるネットワーク カバレッジを実現する有望な機会を提供します。しかし、動的で不慣れな環境で UAV の軌道を最適化することは、特に新しいシナリオごとに広範な再トレーニングが必要であるため、依然として重要な課題です。この論文では、O-RAN アーキテクチャ内に強化された継続的転移学習を統合する、新しい UAV 軌道最適化フレームワークを紹介します。提案されたシステムは、事前トレーニングされたモデルのライブラリを維持し、モデル選択メカニズムを採用して、最も関連性の高い環境から知識を特定して転送することで、適応時間を最小限に抑え、効率を向上させます。十分に類似したモデルが利用できない場合は、継続的な改良によって強化されたフォールバック モデルがベースライン パフォーマンスを保証します。このフレームワークは、現実世界の都市地図とレイ トレーシング技術を活用して、学習の信頼性を高め、軌道計画を改善します。シミュレーション結果は、提案されたモデル選択ベースの転移学習アプローチにより、最初から再トレーニングする場合と比較して収束時間が 44% ~ 56% 短縮され、モデル選択を行わない従来の転移学習と比較して最大 40% 短縮されることを示しています。

原文 (English)

Adaptive Machine Learning Framework for UAV Trajectory Optimization in O-RAN

The deployment of unmanned aerial vehicles (UAV) as open radio units (O-RUs) in 6G cellular systems presents a promising opportunity to achieve scalable and adaptive network coverage. However, optimizing UAV trajectories in dynamic and unfamiliar environments remains a critical challenge, particularly due to the need for extensive retraining in each new scenario. In this paper, we introduce a novel UAV trajectory optimization framework that integrates enhanced continual transfer learning within the O-RAN architecture. The proposed system maintains a library of pre-trained models and employs a model selection mechanism to identify and transfer knowledge from the most relevant environments, minimizing adaptation time and improving efficiency. When no sufficiently similar model is available, a fallback model empowered by continuous refinements ensures baseline performance. The framework leverages real-world city maps and ray tracing techniques to enhance learning reliability and improve trajectory planning. Simulation results demonstrate that the proposed model selection-based transfer learning approach reduces convergence time by 44% to 56% compared to retraining from scratch, and up to 40% compared to traditional transfer learning without model selection.

13:00 JST画像/動画生成

RetiSEM: 断片化された生物医学データの因果モデルの一般化

断片化された生物医学データから因果モデルを学習することは、臨床変数、分子変数、および画像変数が不完全であるか、一緒に観察されていないことが多いため、困難です。我々は、限られたマルチモーダルリソースの下で因果グラフ回復と媒介分析のためのドメイン制約構造方程式モデリング(SEM)フレームワークであるRetiSEMを提案します。この提案された研究では、変数を生物学的に情報を与えられたブロックに編成し、禁止エッジ制約を適用し、経路レベルの効果を TE、NDE、および NIE コンポーネントに分解します。 NHANES の臨床変数と外部から得られた網膜表現を組み合わせた断片化された現実世界の設定とともに、次元、非線形性、因果深さ、経路構造が異なる 10 の合成ベンチマーク シナリオにわたって RetiSEM を評価します。このアプローチでは、合成ベンチマーク全体で制約のないベースラインよりも構造エラーが低くなり、因果関係の精度が高くなります。実データ分析では、網膜変数は主に下流のバイオマーカーのような指標として機能し、小さいながらも検出可能な間接的な効果を伴います。これらの発見は、リソースが限られた生物医学 AI で構造化された因果仮説をテストするための解釈可能なフレームワークとしての私たちの戦略を裏付けています。この作業のコードとリソースは、https://github.com/Inamullah-Colab/ReitSEM で公開されています。

原文 (English)

RetiSEM: Generalising Causal Models for Fragmented Biomedical Data

Learning causal models from fragmented biomedical data is challenging because clinical, molecular, and imaging variables are often incomplete or not jointly observed. We propose RetiSEM, a domain-constrained structural equation modelling (SEM) framework for causal graph recovery and mediation analysis under limited multimodal resources. This proposed work organises variables into biologically informed blocks, applies forbidden-edge constraints, and decomposes pathway-level effects into TE, NDE, and NIE components. We evaluate RetiSEM across ten synthetic benchmark scenarios that vary in dimensionality, nonlinearity, causal depth, and pathway structure, together with a fragmented real-world setting that combines NHANES clinical variables with externally derived retinal representations. This approach achieves lower structural error and higher causal accuracy than unconstrained baselines across the synthetic benchmarks. In the real-data analysis, retinal variables behave mainly as downstream biomarker-like indicators, with smaller but detectable indirect effects. These findings support our strategy as an interpretable framework for testing structured causal hypotheses in limited-resource biomedical AI. The code and resources for this work are publicly available at: https://github.com/Inamullah-Colab/ReitSEM.

13:00 JSTエージェント

Agentic Red チームのレッドチーム化

攻撃的なセキュリティ操作を実行するためのエージェント システムの使用は、理論上の可能性からコモディティ化された機能に移行しました。しかし、コミュニティはより多くの有能なエージェントを作成することに焦点を当ててきましたが、それらのシステムのセキュリティの評価にはあまり注意が払われてきませんでした。この研究では、攻撃的なセキュリティ作戦に最も広く使用されているエージェント システムの詳細なセキュリティ分析を初めて紹介します。これらのツールのほとんどには共通の設計上の欠陥があり、エージェントがサンドボックス コンテナ内で動作している場合でも、積極的な敵対者が API キーを窃取し、永続的な足場を確立し、オペレータのマシンを完全に侵害できることを示します。私たちの分析をサポートするために、このようなエージェント システムに完全なサイバー キル チェーンを導入し、最初の LLM 操作から横方向の移動、永続化、ガードレールのバイパス、サンドボックスからの脱出までの進行を捉えます。私たちはセキュリティ分析に基づいて、エージェント型攻撃セキュリティ ツールの堅牢なアーキテクチャを導き出し、公開された攻撃パスをアーキテクチャ レベルで軽減する実用的で広く適用可能な設計原則を提案します。

原文 (English)

Red-Teaming the Agentic Red-Team

The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the community has focused on creating more and more capable agents, less attention has been allocated to assessing the security of those systems. In this work, we present the first in-depth security analysis of the most widely used agentic systems for offensive security operations. We show that most of these tools share common design flaws that enable an active adversary to exfiltrate API keys, establish persistent footholds, and fully compromise the operator's machine, even when the agent operates inside a sandboxed container. To support our analysis, we introduce a full cyber kill chain for such agentic systems, capturing the progression from initial LLM manipulation to lateral movement, persistence, guardrail bypass, and sandbox escape. Building on our security analysis, we derive a robust architecture for agentic offensive-security tools and propose actionable, broadly applicable design principles that mitigate the disclosed attack paths at the architectural level.

13:00 JSTLLM/生成AIハードウェア/半導体

CrossPool: KV キャッシュと重み分解によるコールド MoE モデルの効率的なマルチ LLM サービス

新興の LLM サービスは、多くの疎な MoE モデルをホストすることが増えていますが、ほとんどのモデルは疎なリクエストを受け取り、コールドのままです。これにより、GPU メモリの問題が発生します。モデルの重みは安定していてモデルによって決定されますが、KV キャッシュは一時的で需要によって決定されます。コールド モデルが同時にピーク KV キャッシュ要求に達することはほとんどないため、モデルごとに最悪の場合の KV 容量を確保するとメモリが無駄になります。代わりに、共有 KV キャッシュ プールを使用して、アクティブな需要を集約してプロビジョニングできます。ただし、重みと KV キャッシュがモノリシック GPU メモリ プールに残っている場合、KV キャッシュの共有は十分ではありません。静的重みは動的 KV キャッシュと競合し、コールドで同時実行トラフィックが少ない場合に KV ヘッドに制限された注意は、レプリケートされた KV 容量の一部のみを公開するため、GPU メモリ使用率が低くなり、ロングコンテキストのサポートが弱くなります。 CrossPool は、FFN 重みと KV キャッシュを 2 つの GPU メモリ プールに分離するコールド MoE モデル用のサービング エンジンです。1 つはコールド モデル全体で FFN 重みを統合する重みプール、もう 1 つは KV キャッシュにローカルな注意を保ちながらアクティブなリクエストを動的に処理する KV キャッシュ プールです。 CrossPool は、KV キャッシュ プランナーとバーチャライザー、隠し状態の転送を隠すレイヤーごとのパイプライン スケジューラー、および CPU-GPU 制御オーバーヘッドを削減するために制御を下げる永続カーネルを組み合わせています。効率的な GPU メモリ プーリングにより、CrossPool はバースト性の高いロングコンテキスト リクエストをサポートし、最先端の kvcached ベースのマルチ LLM サービング システムを上回るパフォーマンスを発揮し、P99 TBT を最大 $10.4\time$ 削減します。

原文 (English)

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory problem: model weights are stable and model-determined, while KV-cache is transient and demand-determined. Because cold models rarely reach peak KV-cache demand at the same time, reserving worst-case KV capacity per model wastes memory; a shared KV-cache pool can instead provision aggregate active demand. However, KV-cache sharing is not sufficient when weights and KV-cache remain in a monolithic GPU memory pool. Static weights compete with dynamic KV-cache, and KV-head-limited attention under cold, low-concurrency traffic exposes only a fraction of replicated KV capacity, leading to low GPU memory utilization and weak long-context support. We present CrossPool, a serving engine for cold MoE models that separates FFN weights and KV-cache into two GPU memory pools: a weights pool that consolidates FFN weights across cold models, and a KV-cache pool that dynamically serves active requests while keeping attention local to KV-cache. CrossPool combines a KV-cache planner and virtualizer, a layer-wise pipeline scheduler that hides hidden-state transfers, and persistent kernels with control lowering to reduce CPU-GPU control overhead. With efficient GPU memory pooling, CrossPool underpins bursty long-context requests and outperforms the state-of-the-art kvcached-based multi-LLM serving system, reducing P99 TBT by up to $10.4\times$.

13:00 JSTビジネス/資金調達

ノードプロパティ予測のためのグラフ基盤モデルの公正な評価

産業や科学のさまざまな分野でグラフ構造データが広く使用されているため、グラフ基盤モデル (GFM) の開発が最近大きな注目を集めています。多くの異なるタイプのモデルが GFM と呼ばれますが、ノード プロパティ予測タスク用に設計された GFM に特に関心が払われています。GFM は、金融およびソーシャル ネットワークでの不正検出から、電子商取引およびユーザー生成コンテンツ プラットフォームの推奨システムに至るまで、多くの実世界のアプリケーションを備えた Graph ML で最も人気のある設定の 1 つです。このタスク用の多数の GFM が最近提案されていますが、この分野は統一された評価設定に収束しておらず、さまざまな研究がモデルを大幅に異なる方法で評価しているため、GFM 同士や他の種類のモデルとの信頼できる比較ができません。この作業では、ノード プロパティ予測のために最近の 9 つの GFM の公正かつ厳密な再評価を実施し、それらを強力なグラフ ニューラル ネットワーク (GNN) ベースラインと比較します。これらの GFM の中で、事前データ適合ネットワーク パラダイムに基づく最新のものだけが、推論コストは高くなりますが、予測パフォーマンスにおいて適切に調整された GNN より優れていることがわかりました。

原文 (English)

A Fair Evaluation of Graph Foundation Models for Node Property Prediction

Due to the wide use of graph-structured data in different fields of industry and science, the development of Graph Foundation Models (GFMs) has recently attracted a lot of attention. While many different types of models are called GFMs, particular interest has been paid to GFMs designed for node property prediction tasks, which is one of the most popular settings in Graph ML with lots of real-world applications from fraud detection in financial and social networks to recommendation systems for e-commerce and user-generated content platforms. While a number of GFMs for this task have been recently proposed, the field has not converged to a unified evaluation setting, and different works evaluate their models in widely different ways, preventing reliable comparison of GFMs with each other and with other types of models. In this work, we conduct a fair and rigorous reevaluation of 9 recent GFMs for node property prediction, comparing them to strong Graph Neural Network (GNN) baselines. We find that, among these GFMs, only the most recent ones based on the Prior-data Fitted Networks paradigm outperform well-tuned GNNs in predictive performance, although at a higher inference cost.

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTGeminiQwen

ポスター: トルコの電話詐欺の音声ベースの検出の限界を探る

詐欺電話は世界中の脆弱なコミュニティを悪用していますが、検出に関する研究はほぼ英語やその他の高リソース言語のみに焦点を当てています。トルコなどのリソースが少ない環境では、注釈付きデータが不足しており、技術的な防御が依然として限られているため、検出は特に困難です。この研究では、詐欺会話と無害な会話の 100 個の整列された音声トランスクリプト ペアからなる初の公開マルチモーダル データセットを導入することにより、大規模言語モデル (LLM) がトルコ語での詐欺検出をどのようにサポートできるかを調査します。 Gemini 2.5 (Flash、Flash-Lite、Pro)、GPT-4o、Qwen (Max、Plus、Turbo) の 3 つのモデル ファミリにまたがる 7 つの LLM を、生の音声、自動音声テキスト変換トランスクリプト、およびネイティブ スピーカーによって洗練されたトランスクリプトの 3 つの入力条件下で評価しました。私たちの結果は、トランスクリプトベースの入力が一貫して直接音声処理よりも優れている一方、人間が修正したトランスクリプトと未修正のトランスクリプトは同等のパフォーマンスを発揮することを示唆しています。この研究では、リソースの少ない言語と現実世界の脅威に焦点を当てることで、文化的および言語的に包括的な AI 安全性研究と、詐欺防止のためのより堅牢なマルチモーダル システムの緊急の必要性を浮き彫りにしています。

原文 (English)

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages. In low-resource settings such as Turkish, detection is especially difficult, as annotated data is scarce and technological defenses remain limited. This research investigates how large language models (LLMs) can support scam detection in Turkish by introducing the first public multi-modal dataset of 100 aligned audio-transcript pairs of scam and benign conversations. We evaluate seven LLMs spanning three model families: Gemini 2.5 (Flash, Flash-Lite, Pro), GPT-4o, and Qwen (Max, Plus, Turbo), under three input conditions: raw audio, automatic speech-to-text transcripts, and transcripts refined by a native speaker. Our results suggest that transcript-based inputs consistently outperform direct audio processing, while human-corrected and uncorrected transcripts perform comparably. By centering a low-resource language and real world threat, this work highlights the urgent need for culturally and linguistically inclusive AI safety research and more robust multi-modal systems for fraud prevention.

13:00 JSTLLM/生成AI

自己進化に対応したワークフロー ハーネスに向けて: エキスパート LLM パイプラインのための可逆的な移行パスと変換性分類法

専門家によって検証された「LLM + スクリプト」ワークフローは、大きな価値をもたらしますが、静的なままです。苦労して獲得したドメイン知識をエンコードしていますが、フィードバックに基づいて実行を適応させることができません。既存のエージェント調査は主にグリーンフィールド エージェントと合成ベンチマークを対象としており、アクティブなレガシー ワークフローの移行は未解決のままです。このギャップを埋めるために、レガシー ワークフローを構成可能で型指定された監査可能なステージにリファクタリングする、可逆的な Strangler-Fig 移行パスを紹介します。このフレームワークの中心となるのは、システム ハーネス内のルーティング ステージとして実装された 3 層の変換可能性分類 (A/B/C) です。これは、ワークフローの準備状況を診断し、それに応じてルーティングします。

原文 (English)

Toward Self-Evolution-Ready Workflow Harnesses: A Reversible Migration Path and Convertibility Taxonomy for Expert LLM Pipelines

While expert-validated "LLM + script" workflows deliver significant value, they remain static: they encode hard-won domain knowledge yet fail to adapt execution based on feedback. Existing agent research predominantly targets greenfield agents and synthetic benchmarks, leaving the migration of active legacy workflows unresolved. To bridge this gap, we present a reversible, Strangler-Fig migration path that refactors legacy workflows into composable, typed, and auditable stages. Central to this framework is a three-tier convertibility taxonomy (A/B/C), implemented as a routing stage within the system harness, which diagnoses a workflow's readiness and routes it accordingly.

13:00 JST研究/論文

無限微因果関係

この論文では、接線バンドルの意味論を備えたフロベニウス マルコフ圏における無限小因果関係のカテゴリカルな説明を紹介します。 IDC は、介入がコピー/破棄構造の接線変形として機能する極小レイヤーをキャプチャします。 2 つの異なるフロベニウス構造が相互作用します。(1) コピー、比較、および破棄をコード化する古典的な変数のカテゴリカル フロベニウス代数。 (2) 幾何学的なフロベニウス可積分条件、すなわち代数的なフロベニウス構造とは異なる介入分布の包含閉包。カテゴリ的因果的十分性は、これら 2 つの概念の互換性として定義されます。重要な観察は、構造因果モデルの場合、微小な因果関係は外生変数に対する決定論的メカニズムのスライスで最も自然に定式化され、目に見える確率的カーネルはプッシュフォワード後にのみ取得されるということです。介入は、フロベニウスのコピー/破棄操作を変形する接線ベクトルです。彼らのリーブラケットは、この変形が古典的な情報フロー構造を保存しているかどうかを測定します。パールの do-calculus は、介入の同一性の指針となる例として使用されます。無関係な介入の無視は単位の不変性に、行動/観察の交換はプッシュフォワードとの共積互換性に、独立性は目に見える介入分布の包括的な括弧の閉包に対応します。

原文 (English)

Infinitesimal Causality

This paper introduces a categorical account of infinitesimal causality in Frobenius Markov categories equipped with tangent-bundle semantics. IDC captures the infinitesimal layer in which interventions act as tangent deformations of copy/discard structure. Two distinct Frobenius structures interact: (1) the categorical Frobenius algebra on classical variables encoding copying, comparing, and discarding; and (2) the geometric Frobenius integrability condition, namely involutive closure of the intervention distribution, distinct from the algebraic Frobenius structure. Categorical causal sufficiency is defined as the compatibility of these two notions. A key observation is that, for structural causal models, infinitesimal causality is most naturally formulated in the slice of deterministic mechanisms over exogenous variables, with visible stochastic kernels obtained only after pushforward. Interventions are tangent vectors that deform the Frobenius copy/discard operations; their Lie brackets measure whether this deformation preserves classical information-flow structure. Pearl's do-calculus is used as a guiding example of intervention identities: ignoring irrelevant interventions corresponds to counit invariance, action/observation exchange to coproduct compatibility with pushforward, and independence to involutive bracket closure of the visible intervention distribution.

13:00 JSTLLM/生成AIエージェントLlama

マルチエージェントのセマンティック書き換えによるプライバシー保護 RAG: コンテキストの忠実性を損なうことなく機密性を実現

検索拡張生成は、外部の知識を組み込むことで大規模な言語モデルを強化しますが、機密性の高いシナリオに導入すると、悪意のあるプロンプトを介してプライバシーが漏洩する危険があります。これに対処するために、セマンティック書き換えを通じて取得したコンテンツをサニタイズするマルチエージェント フレームワークを提案します。プライバシー抽出、セマンティック分析、再構築に 3 つの専門エージェントを採用することで、当社のアプローチはセマンティック コアを維持しながら機密識別子を連携して削除します。私たちは、ChatDoctor および Wiki-PII データセットのフレームワークを 6 つの大規模な言語モデルにわたって評価します。 Experimental results demonstrate a significant reduction in privacy leakage under targeted attacks.たとえば、LLaMA-3-8B における対象情報の露出をベースラインの 144 インスタンスからわずか 1 に削減しました。さらに、BLEU-1 スコア 0.122 で強力なコンテキスト忠実度を維持しており、既存の SAGE メソッドの 0.117 を上回っています。最後に、フレームワークは非同期前処理モジュールとして動作し、すべての書き換えが 1 回限りのオフライン前処理ステップとして実行されるため、オンライン推論に追加の待ち時間が発生しません。再現性を高めるために、この作業のソース コードは https://github.com/foursoils/Privacy-Preserving-RAG で公開されています。

原文 (English)

Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity

Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios risks privacy leakage via malicious prompts. To address this, we propose a multi-agent framework that sanitizes retrieved content through semantic rewriting. By employing three specialized agents for privacy extraction, semantic analysis, and reconstruction, our approach collaboratively removes sensitive identifiers while preserving the semantic core. We evaluate the framework on the ChatDoctor and Wiki-PII datasets across six large language models. Experimental results demonstrate a significant reduction in privacy leakage under targeted attacks. For instance, we reduced targeted information exposure in LLaMA-3-8B from 144 instances in the baseline to just 1. Furthermore, we maintain strong contextual fidelity with a BLEU-1 score of 0.122, outperforming the existing SAGE method's 0.117. Finally, the framework operates as an asynchronous preprocessing module, introducing no additional latency to online inference, as all rewriting is executed as a one-time offline preprocessing step. To promote reproducibility, the source code of this work is publicly available at https://github.com/foursoils/Privacy-Preserving-RAG.

13:00 JST研究/論文

「私たち人間」の視覚化: 多元的なデータ ストーリーテリングを通じて認識のギャップを埋める

従来のビジュアル データ ストーリーテリングは、対立する 2 つの単純化されたグループを描写するバイナリ グラフィックに依存しています。 This can increase political polarization by oversimplifying intra-group disagreements and erasing ambiguity and shared ideas or values.これにより、「私たち対彼ら」という考えがうっかり助長されてしまう可能性があります。 AI 対応デジタル プラットフォームの意図的で多元的な設計を選択すると、ニュアンス、意見の分布、グループ間の共通性を強調する視覚化を生み出すことができます。この可能性を実証するために、高次元の意見空間をマッピングし、合意と反対の両方の領域を強調する審議技術を検討します。この論文は、2025年9月にジグソーとナポリタン研究所によって実施された「We the People」の審議に焦点を当てており、この審議では435の下院選挙区すべての2,400人以上のアメリカ人が自由と平等に関するAI支援の非同期対話に参加した。 AI を利用して長文のテキストベースの参加者の入力をインタラクティブな「意見風景」に合成することにより、このイニシアチブは、多様な視点を人間らしく表現し、実質的に広範なコンセンサスの隠れた領域を明らかにする、多元的なデータ ストーリーテリングの代替形式を提供しました。 The paper concludes that shifting from divisive, contrast-heavy visual frameworks to distribution-focused, interactive models represents a highly scalable, low-cost intervention capable of bridging perceptual gaps and cultivating a more resilient, collaborative democratic culture.

原文 (English)

Visualizing "We the People": Bridging the Perception Gap through Pluralistic Data Storytelling

Traditional visual data storytelling relies on binary graphics that depict two simplified groups in conflict. This can increase political polarization by oversimplifying intra-group disagreements and erasing ambiguity and shared ideas or values. This can inadvertently foster "us versus them" thinking. Intentional, pluralistic design choices for AI-enabled digital platforms can produce visualizations that emphasize nuance, opinion distribution, and intergroup commonalities. To demonstrate this potential, we examine deliberative technologies that map high-dimensional opinion spaces and highlight areas of both consensus and dissensus. The paper highlights the We the People deliberation conducted by Jigsaw and the Napolitan Institute in September 2025, which engaged over 2,400 Americans across all 435 congressional districts in an AI-supported, asynchronous dialogue regarding freedom and equality. By utilizing AI to synthesize long-form, text-based participant inputs into interactive "opinion landscapes," the initiative provided an alternative format for pluralistic data storytelling that humanized diverse viewpoints and revealed hidden areas of substantial broad consensus. The paper concludes that shifting from divisive, contrast-heavy visual frameworks to distribution-focused, interactive models represents a highly scalable, low-cost intervention capable of bridging perceptual gaps and cultivating a more resilient, collaborative democratic culture.

13:00 JSTLLM/生成AI

AI-PAVE-Br: 大規模言語モデルを活用して、ゴールデン セット アプローチによる製品属性値の抽出を強化

The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods for structured information extraction. Traditional approaches to Product Attribute Value Extraction (PAVE) often struggle with the linguistic nuances and sheer diversity of product descriptions in Portuguese. To address this critical gap, this paper introduces two major contributions. First, we present AI-PAVEBr, a specialized system engineered with Large Language Models (LLMs) to perform high-accuracy PAVE specifically for Brazilian e-commerce catalogs. Second, to facilitate reproducible research and provide a definitive benchmark, we introduce and share the Golden Set, a new, meticulously curated, and manually annotated dataset for PAVE in Portuguese.この高品質のリファレンス セットの作成プロセスと構造 (エンティティ、カテゴリ、サブカテゴリ) について詳しく説明します。 Our experiments conclusively show that AI-PAVE-Br, leveraging targeted prompt engineering, dramatically outperforms conventional Named Entity Recognition (NER) baselines. This work not only delivers a superior, scalable solution for a major non-English market but also enriches the NLP community with a valuable, publicly available resource for future PAVE research.

原文 (English)

AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach

The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods for structured information extraction. Traditional approaches to Product Attribute Value Extraction (PAVE) often struggle with the linguistic nuances and sheer diversity of product descriptions in Portuguese. To address this critical gap, this paper introduces two major contributions. First, we present AI-PAVEBr, a specialized system engineered with Large Language Models (LLMs) to perform high-accuracy PAVE specifically for Brazilian e-commerce catalogs. Second, to facilitate reproducible research and provide a definitive benchmark, we introduce and share the Golden Set, a new, meticulously curated, and manually annotated dataset for PAVE in Portuguese. We detail the creation process and structure (Entity, Category, Subcategories) of this high-quality reference set. Our experiments conclusively show that AI-PAVE-Br, leveraging targeted prompt engineering, dramatically outperforms conventional Named Entity Recognition (NER) baselines. This work not only delivers a superior, scalable solution for a major non-English market but also enriches the NLP community with a valuable, publicly available resource for future PAVE research.

13:00 JSTLLM/生成AI

FlowPipe: LLM-Enhanced Conditional Generative Flow Networks for Data Preparation Pipeline Construction

Data preparation pipelines improve data quality in machine learning by transforming raw tables into learning-ready data through sequential…

13:00 JSTロボティクス

TACTFUL: Tactile-Driven Exploration For Object Localization and Identification in Confined Environments

Humans effortlessly locate and identify objects by touch alone, even without vision. In contrast, robotic systems rely heavily on vision an…

13:00 JST画像/動画生成

概念アノテーションを使用したスパース オートエンコーダーの解釈可能性の評価

Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence. We present a human-grounded evaluation framework that quantifies alignment between SAE latents and human-annotated concepts, without requiring user studies, and validate this matching through targeted attribute perturbations. To enable this intervention-style evaluation in vision, we construct synCUB and synCOCO, synthetic benchmarks of paired images that differ in exactly one attribute. We introduce Fully-Binary Matching Pursuit (FBMP), a coalition-based matching procedure that supports many-to-one mappings between SAE latents and annotated concepts, and consistently outperforms one-to-one baselines. For functional validation, we propose a Targeted Attribute Perturbation Alignment Score (TAPAScore), which tests whether matched concepts respond selectively and in the expected direction under targeted image-level attribute perturbations.健全性チェックでは、トレーニング済みの SAE とトレーニングされていない SAE を確実に区別する評価指標は、マッチングと TAPAScore だけです。 Across SAEs trained on CLIP and DINOv2 embeddings, we find that increased overcompleteness can reduce perturbation alignment, indicating a reduction in interpretability. Our evaluation framework suggests that moderate dictionary sizes provide the best trade-off, yielding the most interpretable SAEs. Code and datasets are available at https://github.com/JonasKlotz/sae-concept-eval.

原文 (English)

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence. We present a human-grounded evaluation framework that quantifies alignment between SAE latents and human-annotated concepts, without requiring user studies, and validate this matching through targeted attribute perturbations. To enable this intervention-style evaluation in vision, we construct synCUB and synCOCO, synthetic benchmarks of paired images that differ in exactly one attribute. We introduce Fully-Binary Matching Pursuit (FBMP), a coalition-based matching procedure that supports many-to-one mappings between SAE latents and annotated concepts, and consistently outperforms one-to-one baselines. For functional validation, we propose a Targeted Attribute Perturbation Alignment Score (TAPAScore), which tests whether matched concepts respond selectively and in the expected direction under targeted image-level attribute perturbations. Under sanity checks, our matching and TAPAScore are the only evaluated metrics that reliably distinguish trained SAEs from untrained ones. Across SAEs trained on CLIP and DINOv2 embeddings, we find that increased overcompleteness can reduce perturbation alignment, indicating a reduction in interpretability. Our evaluation framework suggests that moderate dictionary sizes provide the best trade-off, yielding the most interpretable SAEs. Code and datasets are available at https://github.com/JonasKlotz/sae-concept-eval.

13:00 JSTLLM/生成AI

効率的なアノテーションのためのタスク分解

High-quality annotations of structured representations are expensive to collect over large corpora.構造の手動アノテーションは手間がかかり、モデルベースのアノテーションは、生成コストは安くなりますが、アノテーションの品質が下流で役立つのに十分な強度であることを確認するために、高価な検証と潜在的に重要な監視が必要です。従来のアノテーション ワークフローでは、各完全なサンプルのアノテーションは、1 人のアノテーターによってエンドツーエンドで実行されます。 However, structured annotation is complex, and each aspect of the task represents a unique challenge with an associated inferential load for a given annotator. Modern annotation projects can incorporate heterogeneous groups of annotators, including both models and human annotators with varying domain and linguistic expertise. It remains unclear, however, how to redesign annotation tasks in this setting, where efforts are discriminately allocated across heterogeneous annotators with respect to distinct annotation challenges. We propose to decompose annotation tasks into sub-tasks in order to reduce the aggregate inferential load of annotation projects. Inspired by the notion of centers from centering theory, we introduce a formal model of inferential load based on the degrees of freedom in the space of valid annotations. Using this model, we show that identifying these centers (i.e. salient anchor entities realized by annotation sub-tasks) constrains the output space complexity, and decompositions which isolate and advance center identification reduce the aggregate inferential load.私たちは、以前の研究からコスト効率が向上したことを示す例によって裏付けられた、複雑な構造化アノテーション タスクを分解するためのガイドラインを提供します。 Finally, we present a procedure for allocating sub-tasks across annotators to maximize quality under a fixed budget.

原文 (English)

Task Decomposition for Efficient Annotation

High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is laborious, and model-based annotation, although cheaper to generate, requires expensive validation and potentially significant supervision to ensure that the annotation quality is strong enough to be useful downstream. In traditional annotation workflows, annotation of each complete example is performed end-to-end by a single annotator. However, structured annotation is complex, and each aspect of the task represents a unique challenge with an associated inferential load for a given annotator. Modern annotation projects can incorporate heterogeneous groups of annotators, including both models and human annotators with varying domain and linguistic expertise. It remains unclear, however, how to redesign annotation tasks in this setting, where efforts are discriminately allocated across heterogeneous annotators with respect to distinct annotation challenges. We propose to decompose annotation tasks into sub-tasks in order to reduce the aggregate inferential load of annotation projects. Inspired by the notion of centers from centering theory, we introduce a formal model of inferential load based on the degrees of freedom in the space of valid annotations. Using this model, we show that identifying these centers (i.e. salient anchor entities realized by annotation sub-tasks) constrains the output space complexity, and decompositions which isolate and advance center identification reduce the aggregate inferential load. We provide guidelines for decomposing complex structured annotation tasks, supported by examples demonstrating improved cost-efficiency from our prior work. Finally, we present a procedure for allocating sub-tasks across annotators to maximize quality under a fixed budget.

13:00 JST研究/論文

U-Net を超えて: フローマッチング音声強化のための潜在表現に合わせたスキップフリー バックボーン

生成モデル、特に拡散およびスコアベースのアプローチは、音声強調において最近優れたパフォーマンスを達成していますが、その反復サンプリング プロセスによりリアルタイムの展開が制限されます。フロー マッチングは、関数評価をほとんど行わない常微分方程式を通じて、ノイズの多い音声をクリーンな音声に変換することにより、効率的な代替手段を提供します。この研究では、潜在表現アライメント (LRA) に基づいて、フローマッチング音声強化のためのスキップフリーのエンコーダ/デコーダ バックボーンを提案します。提案されたモデルは、ノイズ相関の低レベルの特徴をデコーダに転送する可能性がある U-Net スキップ接続に依存するのではなく、量子化を行わずに凍結された Descript Audio Codec エンコーダ/デコーダから抽出されたクリーンな潜在特徴を使用してボトルネックとデコーダの表現を調整します。 This codec-aligned supervision promotes compact clean-speech representations while preserving efficient few-step inference. Experiments on WSJ0-CHiME3 and VoiceBank-DEMAND show improved PESQ and perceptual quality, especially on VoiceBank-DEMAND, using only five function evaluations.

原文 (English)

Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement

Generative models, particularly diffusion and score-based approaches, have recently achieved strong performance in speech enhancement, but their iterative sampling process limits real-time deployment. Flow Matching offers an efficient alternative by transporting noisy speech toward clean speech through an ordinary differential equation with few function evaluations. In this work, we propose a skip-free encoder-decoder backbone for flow-matching speech enhancement, guided by Latent Representation Alignment (LRA). Instead of relying on U-Net skip connections, which may transfer noise-correlated low-level features to the decoder, the proposed model aligns its bottleneck and decoder representations with clean latent features extracted from a frozen Descript Audio Codec encoder-decoder without quantization. This codec-aligned supervision promotes compact clean-speech representations while preserving efficient few-step inference. Experiments on WSJ0-CHiME3 and VoiceBank-DEMAND show improved PESQ and perceptual quality, especially on VoiceBank-DEMAND, using only five function evaluations.

13:00 JSTLLM/生成AI画像/動画生成エージェント

UniDrive: 自動運転におけるリスクを解釈可能に理解するための統一された視覚言語とグラウンディング フレームワーク

最近のマルチモーダル大規模言語モデル (MLLM) は、自動運転シーンの理解に大きな可能性を示していますが、既存の方法では依然として時間的推論と空間精度の間の根本的なトレードオフに直面しています。 Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations.このギャップに対処するために、自動運転におけるリスクを解釈可能に理解するための統一された視覚言語と基礎フレームワークである UniDrive を提案します。 UniDrive は、マルチフレームの視覚入力からシーンのダイナミクスをモデル化する時間的推論ブランチと、最新のフレームからのきめの細かい空間詳細を保存する高解像度の知覚ブランチを組み合わせています。 2 つのブランチは、ゲート制御クロスアテンション フュージョン モジュールを通じて統合され、動的なコンテキストを正確な空間証拠と一致させることができます。融合表現に基づいて、UniDrive は自然言語によるリスク記述とリスク オブジェクトの根拠のある境界ボックス出力を共同生成します。 DRAMA-Reasoning ベンチマークの実験では、UniDrive がキャプションとリスクオブジェクトのグラウンディングの両方において、代表的な画像ベースおよびビデオベースのベースラインよりも優れていることが示されています。特に、UniDrive は検証分割で最高の全体的なパフォーマンスを達成し、小さなオブジェクトの位置特定、NuScenes および BDD100K へのゼロショットの一般化、人間による解釈可能性と信頼性において明らかな利点を示しています。これらの結果は、時間セマンティクスと高解像度の知覚を明示的に組み合わせることで、解釈可能な安全志向の自動運転システムのためのより強力な基盤が提供されることを示唆しています。コードは https://github.com/pixeli99/unidrive-dev で入手できます。

原文 (English)

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations. To address this gap, we propose UniDrive, a unified visual-language and grounding framework for interpretable risk understanding in autonomous driving. UniDrive combines a temporal reasoning branch that models scene dynamics from multi-frame visual input with a high-resolution perception branch that preserves fine-grained spatial details from the latest frame. The two branches are integrated through a gated cross-attention fusion module, enabling dynamic context to be aligned with precise spatial evidence. Based on the fused representation, UniDrive jointly generates natural-language risk descriptions and grounded bounding-box outputs for risk objects. Experiments on the DRAMA-Reasoning benchmark show that UniDrive outperforms representative image-based and video-based baselines in both captioning and risk-object grounding. In particular, UniDrive achieves the best overall performance on the validation split and demonstrates clear advantages in small-object localization, zero-shot generalization to NuScenes and BDD100K, and human-rated interpretability and trustworthiness. These results suggest that explicitly combining temporal semantics and high-resolution perception provides a stronger foundation for interpretable and safety-oriented autonomous driving systems. The code is available at https://github.com/pixeli99/unidrive-dev.

13:00 JST研究/論文

マルチモーダル教科書機能による生徒のクイズ成績のコンテキスト認識型予測

Educational platforms often predict student performance from prior interactions, but the assessment content itself also varies in linguistic and visual complexity.この論文では、CourseKata の章復習問題から抽出された軽量コンテンツの特徴が、章末クイズのスコアの予測を生徒の以前の演習の平均成績を超えて改善するかどうかを研究します。この研究では、2023年のCourseKata学生の回答データと、復習問題の文言からの章レベルのテキスト特徴および教科書のビジュアルからの画像特徴を組み合わせています。 562 のクラス学生 ID からの 4,742 の学生章の観察全体で、コンテンツ機能を追加すると、学生グループの 5 倍クイズの予測パフォーマンスが、以前のパフォーマンスのベースラインと比較して 9.1% 向上しました。チャプターアウトなし検証では、テキスト特徴によりベースラインと比較して予測誤差が減少しますが、画像を含むモデルでは誤差が大きくなります。この論文は、コンテキスト認識モデルが、過去の生徒の成績のみを使用する場合と比較して、生徒のクイズの成績をより正確に予測するために、質問のテキストと視覚的特徴に関する有用なシグナルを追加することを示唆しています。

原文 (English)

Context-Aware Prediction of Student Quiz Performance with Multimodal Textbook Features

Educational platforms often predict student performance from prior interactions, but the assessment content itself also varies in linguistic and visual complexity. This paper studies whether lightweight content features extracted from CourseKata chapter-review questions improve prediction of end-of-chapter quiz scores beyond a student's average prior exercise performance. The study combines 2023 CourseKata student response data with chapter-level text features from review-question wording and image features from textbook visuals. Across 4,742 student-chapter observations from 562 class-student IDs, adding content features improves student-grouped five-fold quiz prediction performance by 9.1% relative to a prior-performance baseline. In leave-chapter-out validation, text features reduce prediction error relative to the baseline, while image-containing models have higher error. This paper suggests that a context-aware model adds useful signal about the text and visual features of questions to better predict student quiz performance compared with using past student performance alone.

13:00 JSTエージェント

DeepBD: バリアントの優先順位付けと遺伝的先天異常の診断のためのグラウンデッド エージェント ワークフロー

Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability.遺伝的病因が疑われるサブセットでは、エクソームおよびゲノム配列決定により、多くの症例が変異検出から配列決定後の解釈へと移行しました。臨床医は、不完全な胎児または乳児の表現型と、集団遺伝学、変異の影響予測、遺伝子疾患の妥当性、表現型オントロジー、細胞および経路の状況、タンパク質構造および臨床文献からの不均一な証拠に基づいて、患者固有の候補変異をランク付けする必要があります。我々は、遺伝的先天異常のバリアントの優先順位付けと診断的解釈のための根拠のあるエージェントワークフローである DeepBD を紹介します。 DeepBD は、ワークフローを LLM 支援のケース構造化、事前トレーニング済み証拠エンジン、専門家証拠モジュール、および根拠のある診断レビュー層に編成します。エビデンスエンジンは、構造化されたルールの証拠、配列とバリアント効果の表現、および表現型条件付けされた生物学的コンテキストから患者固有のバリアントスコアを学習しますが、スペシャリストモジュールとエージェントレイヤーは、ツールベースの改良、候補プールのレビュー、およびランク付けされた候補からの診断指向の統合を提供します。 Developed using an in-house fetal and infant cohort comprising 18,622 cases, DeepBD achieved Recall@1/3/5/10 of 0.658/0.882/0.912/0.929 on an internal held-out solved-case benchmark, outperforming standalone Exomiser, DeepRare and prompted LLM reranking baselines evaluated on Exomiser-derived top-20 candidate亜種。 Ablation and overlap analyses show that rule evidence, mechanistic context, and specialist refinement provide complementary signals.これらの発見は、遺伝性先天性欠損症における遡及的バリアント優先順位付けのための証拠の統合、ツールベースの改良、LLM 支援診断レビューを分離する、根拠のあるエージェントワークフローを裏付けています。

原文 (English)

DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects

Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability. In the subset with suspected genetic etiologies, exome and genome sequencing have moved many cases from variant detection to post-sequencing interpretation: clinicians must rank patient-specific candidate variants under incomplete fetal or infant phenotypes and heterogeneous evidence from population genetics, variant-effect prediction, gene-disease validity, phenotype ontologies, cellular and pathway context, protein structure and clinical literature. We present DeepBD, a grounded agentic workflow for variant prioritization and diagnostic interpretation of genetic birth defects. DeepBD organizes the workflow into LLM-assisted case structuring, a pretrained evidence engine, specialist evidence modules and a grounded diagnostic review layer. The evidence engine learns patient-specific variant scores from structured rule evidence, sequence and variant-effect representations and phenotype-conditioned biological context, whereas specialist modules and the agentic layer provide tool-based refinement, candidate-pool review and diagnosis-oriented synthesis from ranked candidates. Developed using an in-house fetal and infant cohort comprising 18,622 cases, DeepBD achieved Recall@1/3/5/10 of 0.658/0.882/0.912/0.929 on an internal held-out solved-case benchmark, outperforming standalone Exomiser, DeepRare and prompted LLM reranking baselines evaluated on Exomiser-derived top-20 candidate variants. Ablation and overlap analyses show that rule evidence, mechanistic context, and specialist refinement provide complementary signals. These findings support a grounded agentic workflow that separates evidence integration, tool-based refinement, and LLM-assisted diagnostic review for retrospective variant prioritization in genetic birth defects.

13:00 JSTLLM/生成AIエージェント

知っておくべきこと: エージェントの電子商取引における検証済み製品情報のマイクロトランザクション市場

商用 NLP は、ショッピング チャットボットをレコメンダーまたはコンバージョン ツールとして扱います。その仕事は、ユーザーをカタログ エントリと照合し、販売を成立させることです。私たちは、エージェントネイティブのマイクロペイメントレール(x402、AP2など)の登場により、不足しているものが変化すると主張します。購入者が徹底的に調査できる自律的なエージェントである場合、ボトルネックは製品を照合することではなく、製品に関する信頼できる意思決定に関連する情報を取得することです。私たちは、エージェント電子商取引を検証済み情報のマイクロトランザクション市場として想定しています。バイヤーエージェントは、販売者とレビューアが提供するデータ(サービス履歴、サードパーティのテストレポート、部品表、監査済みの販売およびサポート指標)を段階的にアンロックするために数セントを費やし、フリーミアムモデルの下でアラカルトで支払いを行い、レビューアの信頼は評判によってスコア化されます。 We sketch the architecture of such a market and argue that it rewards genuine product quality and yields truer competition than ranking-based storefronts.次に、このビジョンを具体的な NLP 問題 (コスト最適化の情報取得、データの価格設定と交渉、リアルタイムのエンティティ解決、根拠のある価値交換、プライバシーを保護するペルソナ モデリング) に変換し、チャットの流暢さではなく、これらがこの分野の注目に値すると主張します。

原文 (English)

Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce

Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close a sale. We argue that the arrival of agent-native micro-payment rails (e.g., x402, AP2) changes what is scarce. When the buyer is an autonomous agent that can investigate exhaustively, the bottleneck is no longer matching products but acquiring trustworthy, decision-relevant information about them. We envision agentic e-commerce as a micro-transaction market for verified information: buyer agents spend fractions of a cent to progressively unlock seller- and reviewer-supplied data -- service histories, third-party test reports, bills of materials, audited sales and support metrics -- paid for a la carte under a freemium model, with reviewer trust scored reputationally. We sketch the architecture of such a market and argue that it rewards genuine product quality and yields truer competition than ranking-based storefronts. We then translate the vision into concrete NLP problems -- cost-optimal information acquisition, data pricing and negotiation, real-time entity resolution, grounded value exchange, and privacy-preserving persona modelling -- and argue that these, not chat fluency, deserve the field's attention.

13:00 JSTLLM/生成AI

Grad Detect: LLM における勾配ベースの幻覚検出

大規模言語モデル (LLM) は、さまざまなタスクにわたって優れた機能を実証してきましたが、依然として幻覚を生成する傾向があります。これらの幻覚を検出することは、一か八かのアプリケーションに LLM を確実に導入するために重要です。我々は、推論中の単一の前後方向パスから層ごとの勾配パターンを分析することによって幻覚を予測するための勾配ベースのアプローチである Grad Detect を紹介します。私たちの方法は、モデルの内部勾配構造がその出力の正確さに関する豊富な情報を保持していることを示しています。この情報は、出力レベル信号だけではアクセスできません。私たちは、幻覚検出とモデル棄権予測の両方にわたるいくつかの Q&A ベンチマークで Grad Detect を評価し、信頼性ベースおよびサンプリングベースのベースラインを一貫して上回っています。 4 つのアーキテクチャ ファミリの 11 モデルすべてにわたる包括的な層アブレーション研究を通じて、最後の 5 層が識別勾配信号の 97% 以上を集中させ、最小限のパフォーマンス損失で効率的な展開が可能であることがわかりました。 Grad Detect は、LLM の信頼性の複数の側面を予測するための統合フレームワークを提供し、モデルの障害がどこでどのように発生するかについての解釈可能な洞察とともに、強力な予測パフォーマンスを提供します。

原文 (English)

Grad Detect: Gradient-Based Hallucination Detection in LLMs

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinations. Detecting these hallucinations is critical for deploying LLMs reliably in high-stakes applications. We present Grad Detect, a gradient-based approach for predicting hallucinations by analyzing layer-wise gradient patterns from a single forward-backward pass during inference. Our method shows that the internal gradient structure of a model carries rich information about the correctness of its output. This information is not accessible through output-level signals alone. We evaluate Grad Detect on several Q&A benchmarks across both hallucination detection and model abstention prediction, where it consistently outperforms confidence-based and sampling-based baselines. Through comprehensive layer ablation studies across all eleven models from four architectural families, we find that the final five layers concentrate over 97% of the discriminative gradient signal, enabling efficient deployment with minimal performance loss. Grad Detect provides a unified framework for predicting multiple dimensions of LLM reliability, offering strong predictive performance alongside interpretable insights into where and how model failures originate.

13:00 JSTLLM/生成AI画像/動画生成研究/論文

EG-VQA: 根拠のある時間的証拠を使用した検証可能なビデオ質問応答のベンチマーク

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA).それにもかかわらず、既存のベンチマークは主に回答の正しさによって評価される一方、関連するビデオ証拠における予測の根拠はほとんど検討されていないままです。回答の生成と証拠の理解との間のこの断絶は、証拠に基づくビデオ質問応答ベンチマーク (EG-VQA) の構築を動機付けています。EG-VQA は、各 QA ペアに裏付けとなる一時的な証拠で明示的に注釈が付けられているため、共同推論と正確な証拠の位置特定が必要となる、オープンエンドの評価プロトコルです。 EG-VQA は、2,067 のビデオと、きめ細かい証拠の注釈が付いた 11,838 の QA ペアで構成されています。予測された証拠を評価するために、Evidence-Grounded F1 (EG-F1) が統一指標として導入され、グラウンドトゥルース証拠に対する時間的整合性と意味論的一貫性が共同で測定されます。実験による評価では、強力な独自モデルであっても予測を正確に根拠付けるのに苦労しており、回答の正確さと忠実な証拠の位置特定との間に根本的な矛盾があることが明らかになりました。このギャップを埋めるために、明示的な監視でトレーニングされた証拠に基づく推論モデルである EG-Reasoner が提案されています。オープンソース モデルの中で最先端のパフォーマンスが達成され、独自のシステムに匹敵する結果が得られます。特に、反事実の質問などの推論が集中するタスクで顕著な向上が観察されます。 These findings demonstrate that scaling alone is insufficient for robust video understanding and that structured evidence supervision is essential for the development of more reliable and interpretable VideoQA systems.

原文 (English)

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the grounding of predictions in relevant video evidence remains largely unexamined. This disconnect between answer generation and evidence understanding motivates the construction of the Evidence-Grounded Video Question Answering Benchmark (EG-VQA), an open-ended evaluation protocol in which each QA pair is explicitly annotated with supporting temporal evidence, thereby requiring joint reasoning and precise evidence localization. EG-VQA is comprised of 2,067 videos and 11,838 QA pairs with fine-grained evidence annotations. To evaluate predicted evidence, Evidence-Grounded F1 (EG-F1) is introduced as a unified metric in which temporal alignment and semantic consistency against ground-truth evidence are jointly measured. Experimental evaluation reveals that even strong proprietary models struggle to accurately ground their predictions, exposing a fundamental discrepancy between answer correctness and faithful evidence localization. To bridge this gap, EG-Reasoner, an evidence-grounded reasoning model trained with explicit supervision, is proposed. State-of-the-art performance is achieved among open-source models, with results competitive against proprietary systems, particularly pronounced gains are observed on reasoning-intensive tasks such as counterfactual questions. These findings demonstrate that scaling alone is insufficient for robust video understanding and that structured evidence supervision is essential for the development of more reliable and interpretable VideoQA systems.

13:00 JST画像/動画生成

OrbitForge: 再構築アンカー付きビデオ合成によるテキストから 3D シーンの生成

一般的なテキストからビデオへのモデルは、リッチなオープンワールド シーンの事前処理として使用できます。今日生成されたビデオは高品質であるにもかかわらず、信頼できる 3D アセットを直接生み出すわけではありません。カメラの動きの制御が難しく、ビュー範囲が部分的で、時間の経過とともにフレームに不一致が含まれることがよくあります。 OrbitForge は、フリーズされたビデオ事前情報とプロンプトごとのガウス スプラッティング再構築最適化から構築されたアダプターで、単一のテキスト生成ビデオを正規の閉軌道 3D ガウス スプラッティング シーンに変換します。生成されたビデオの 3D の一貫性を向上させるために、3D 再構成をアンカーとして使用します。堅牢な MedianGS プロキシを使用した変形可能ガウス スプラッティングを介して、最初に生成されたビデオから予備的な 3D 再構成を取得します。欠落している視点を検出するために、所定の軌道からビューをレンダリングします。 OrbitForge は、テキストからビデオへのモデルを使用して欠落しているビューのみを完成させ、完成した軌道を最終的なガウス スプラッティング シーンに再構築します。この設計では、タスク固有のビデオやマルチビューの微調整は必要なく、プロンプトごとのスコア抽出の最適化を回避し、一度に 1 ステップずつ段階的にビューを生成しません。さらに、この設定にはカバレッジを意識した評価が必要であると主張します。局所的な滑らかさだけで、完全な軌道を決して試行しないメソッドに報酬が与えられます。凍結された 300 プロンプトの T3Bench 由来の監査では、OrbitForge 再構成は 359.0 度の測定中央値スパンを達成し、元々サポートされていなかったビン Q10 ImageReward を MedianGS のみの再構成と比較して 8.07 から 16.36 に引き上げ、同時にカバレッジ品質において VideoMV との競争力を維持しました。

原文 (English)

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis

Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not directly yield reliable 3D assets: camera motion is difficult to control, view coverage is partial, and frames often contain inconsistencies across time. We introduce OrbitForge, an adapter built from frozen video priors and per-prompt Gaussian Splatting reconstruction optimization that converts a single text-generated video into a canonical closed-orbit 3D Gaussian Splatting scene. We use 3D reconstruction as an anchor to improve the 3D consistency of the generated video. We obtain a preliminary 3D reconstruction from a first generated video via Deformable Gaussian Splatting with a robust MedianGS proxy. We render views from a prescribed orbit to detect missing viewpoints. OrbitForge uses the text-to-video model to complete only the missing views, and reconstructs the completed orbit into a final Gaussian Splatting scene. This design requires no task-specific video or multiview fine-tuning, avoids per-prompt score-distillation optimization, and does not progressively generate views one step at a time. We further argue that this setting demands coverage-aware evaluation: local smoothness alone rewards methods that never attempt a full orbit. On a frozen 300-prompt T3Bench-derived audit, OrbitForge reconstruction attains a 359.0-degree measured median span, raises originally unsupported-bin Q10 ImageReward from 8.07 to 16.36 relative to MedianGS-only reconstruction, while remaining competitive with VideoMV on the coverage-quality.

13:00 JST研究/論文GPT / ChatGPT

構造化概念進化による量子LDPC符号の大言語モデルの発見

量子コンピューターは、重要な問題に関して古典的なマシンを上回るパフォーマンスを発揮できる可能性がありますが、それは量子ハードウェアに蔓延するエラーを大規模に修正できる場合に限ります。量子低密度パリティ チェック (qLDPC) コードは、スパース パリティ チェックを有限の符号化レートと増大する距離と組み合わせることで、この目標への有望な道筋を提供しますが、その構築は依然として困難な離散設計の問題です。ここでは、大規模な言語モデルと構造化代数突然変異文法を組み合わせて、CSS qLDPC コードのクラスであるリフト積コード ファミリを発見する検索フレームワークである構造化概念進化 (SCE) を紹介します。 LLM に第一原理に基づいてコードを設計するよう依頼する代わりに、SCE は、群代数、プロトグラフ幾何学、または基底空間を変更する階層的突然変異を使用して、代数仕様とそれを実現する実行可能プログラムの組み合わせからなる構造化された概念を進化させます。 SCE を実行すると、アーベル構造から、二変量バイシクル コードなどの基礎となる標準設計を超えた非アーベル グループにわたるファミリーに至るまで、競合するコード ファミリの多様なセットを発見し、BP+OSD デコードによるコード容量の脱分極ノイズの下でそれらを特徴付けます。これらの結果は、軽量モデル (GPT-5.4-mini および GPT-5.4-nano) で得られます。

原文 (English)

Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution

Quantum computers could outperform classical machines on important problems, but only if the errors that pervade quantum hardware can be corrected at scale. Quantum low-density parity-check (qLDPC) codes offer a promising route to this goal by combining sparse parity checks with finite encoding rate and growing distance, but their construction remains a challenging discrete design problem. Here we introduce structured concept evolution (SCE), a search framework that pairs a large language model with a structured algebraic mutation grammar to discover lifted-product code families, a class of CSS qLDPC codes. Instead of asking the LLM to design codes from first principles, SCE evolves structured concepts consisting of algebraic specifications paired with executable programs that realize them, using hierarchical mutations that modify the group algebra, protograph geometry, or base space. Running SCE, we discover a diverse set of competitive code families, ranging from abelian constructions to families over non-abelian groups beyond those underlying standard designs such as bivariate-bicycle codes, and characterize them under code-capacity depolarizing noise with BP+OSD decoding. These results are obtained with lightweight models (GPT-5.4-mini and GPT-5.4-nano).

13:00 JSTLLM/生成AI画像/動画生成

IV-CoT: 構造を意識したテキストから画像への生成のための暗黙的な視覚的思考連鎖

統合マルチモーダル大規模言語モデル (MLLM) は、強力なテキストから画像への生成品質を実現していますが、オブジェクト数、空間関係、属性バインディング、および大まかなレイアウトを保持する必要がある、構造を認識したプロンプト追従には依然として苦労しています。この制限の一部は、単一のコンディショニング ストリーム内での構造計画と外観レンダリングの絡み合いにあると考えられます。この問題に対処するために、クエリ条件付き画像生成のための潜在的な視覚推論フレームワークである Implicit Visual Chain-of-Thought (IV-CoT) を提案します。 IV-CoT は、視覚的条件付けクエリを構造から意味論的なカスケードに分解します。構造的クエリは最初に潜在的な視覚的プランを形成し、次に意味論的クエリはこのプランに基づいて条件付けされた外観をレンダリングします。構造クエリをガイドするために、トレーニング専用のスケッチ監視を導入します。これにより、推論時のスケッチ抽出や中間デコードを必要とせずに、スケッチから構造をキャプチャすることが促進されます。 IV-CoT は、単一のフォワード パスで暗黙的な CoT 推論を実行し、GenEval および T2I-CompBench で優れた結果を達成します。視覚化と分析は、学習された構造クエリと意味論的クエリが構造認識生成において補完的な役割を果たすことを示しています。

原文 (English)

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object counts, spatial relations, attribute bindings, and coarse layouts must be preserved. We attribute this limitation in part to the entanglement of structural planning and appearance rendering within a single conditioning stream. To address this issue, we propose Implicit Visual Chain-of-Thought (IV-CoT), a latent visual reasoning framework for query-conditioned image generation. IV-CoT decomposes the visual conditioning queries into a structural-to-semantic cascade, where structural queries first form a latent visual plan and semantic queries then render appearance conditioned on this plan. To guide the structural queries, we introduce training-only sketch supervision, which encourages them to capture structure from sketches without requiring sketch extraction or intermediate decoding at inference time. IV-CoT performs implicit CoT reasoning in a single forward pass and achieves superior results on GenEval and T2I-CompBench. Visualizations and analyses demonstrate that the learned structural and semantic queries play complementary roles in structure-aware generation.

13:00 JSTビジネス/資金調達

複雑です: AI を活用した AAC インターフェイスの設計と評価について

人工知能 (AI) は、拡張代替コミュニケーション (AAC) を使用する人々がシステムでできることを強化できます。ただし、AI を活用した AAC インターフェイスの評価は難しい場合があります。人々は交差する存在であり、現在の評価指標では、人々が AAC に対して抱く可能性のある多面的で微妙な欲求を捉えるのが難しい場合があります。私たちは、AAC の 6 つの問題空間の複雑な性質を調査し、これらの空間で AI がどのように使用されるかを検討し、人々の交差するニュアンスを考慮したより堅牢な評価方法を提案します。また、これらの問題領域全体で発生するより広範な問題と、提案された評価方法を使用してそれらにどのように対処できるかについても説明します。

原文 (English)

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are able to do with their systems. However, evaluating AI-powered AAC interfaces can be difficult. People are intersectional beings and current evaluation metrics can struggle to capture the multifaceted and nuanced desires people may have for their AAC. We explore the complicated nature of six AAC problem spaces, explore how AI might be used in these spaces, and suggest more robust methods of evaluation that take the intersectional nuances of people into account. We also discuss broader issues that arise across these problem spaces and how they could be addressed using our proposed evaluation methods.

13:00 JST画像/動画生成

FLUX3D: 拡散整列スパース表現による高忠実度 3D ガウス生成

スパースボクセル表現は、画像から 3D ガウス スプラッティング (3DGS) 生成のためのスケーラブルな基盤として登場しましたが、現在の手法では 2 つの構造的なボトルネックにより、入力画像の高周波の視覚的な詳細を保持するのが困難です。まず、セマンティック抽象化に最適化された識別 2D 特徴を採用して、スパースなボクセル潜在を構築します。これにより、再構築キューが抑制され、表現のボトルネックが引き起こされます。第二に、生成段階において、標準的な拡散変換器には、密集した 2D 画像トークンを疎な 3D ボクセル潜在と位置合わせするための効果的なメカニズムが欠如しており、その結果、クロスモーダル対応のボトルネックが生じます。これらの問題に対処するために、生成中の表現学習とクロスモーダル アライメントの両方を強化するスケーラブルな画像から 3DGS フレームワークである FLUX3D を提案します。まず、スパースボクセルベースの 3D 表現学習のための 2D 特徴選択を再検討し、拡散整合構造化潜在 (DA-SLAT) を提案し、それをデコーダ専用アーキテクチャと組み合わせて 3DGS 再構築の忠実度を向上させます。また、疎構造認識拡散フレームワークも設計します。これは、疎構造マルチモーダル拡散変換器 (SMDiT) とモーダル認識回転位置埋め込み (MARoPE) を統合して、ジオメトリに依存しない 2D-3D アライメントを実現します。広範なベンチマーク実験により、FLUX3D は外観の忠実度が大幅に向上し、高品質の 3DGS アセットの生成においてすべての最先端 (SOTA) 手法を大幅に上回ることが実証されました。

原文 (English)

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck. Second, in the generation stage, standard diffusion transformers lack effective mechanisms to align dense 2D image tokens with sparse 3D voxel latents, resulting in a cross-modal correspondence bottleneck. To address these issues, we propose FLUX3D, a scalable image-to-3DGS framework that boosts both representation learning and cross-modal alignment during generation. We first revisit 2D feature selection for sparse-voxel-based 3D representation learning, propose Diffusion-Aligned Structured Latents (DA-SLAT) and couple it with a decoder-only architecture to improve 3DGS reconstruction fidelity. We also design a sparse-structure-aware diffusion framework, which integrates the Sparse-structure Multimodal Diffusion Transformer (SMDiT) and Modal-Aware Rotary Positional Embedding (MARoPE) to achieve geometry-agnostic 2D-3D alignment. Extensive benchmark experiments demonstrate that FLUX3D yields substantial improvements in appearance fidelity and significantly outperforms all state-of-the-art (SOTA) methods in generating high-quality 3DGS assets.

13:00 JSTロボティクスビジネス/資金調達

InSight: 操縦可能な VLA を介した自己ガイドによるスキル習得

ビジョン言語アクション (VLA) モデルはデモンストレーションから操作スキルを学習できますが、その機能はトレーニング データ内のスキルによって制限されます。我々は、原始的な動作レベル (例: 「グリッパーをボウルに移動する」、「上方に持ち上げる」、「ボトルに注ぐ」など) で VLA を操作可能にすることで、自律的なスキル習得を可能にするフレームワークである InSight を紹介します。 InSight は 2 つの主要なステージで構成されます。(1) VLA プリミティブのステアビリティを可能にするために、VLM プラン分解とエンドエフェクター ポーズによってデモンストレーションをラベル付きプリミティブに分割する自動セグメンテーション パイプライン、(2) 新しいタスクを達成するために必要な欠落しているプリミティブを特定し、VLM が提案する低レベル制御を使用して欠落しているプリミティブのデモンストレーションを自律的に試み、成功したものを自動的にラベル付け、保存、統合する VLM ガイド付きデータ フライホイールVLA トレーニング セットへのデモンストレーション。当社では、シミュレーションおよび実際の操作タスク (ブロックの反転、引き出しの閉め方、掃除、ひねり、流し込みなど) にわたって InSight を評価します。これらの対象スキルを人間がデモンストレーションする必要はありません。一度学習すると、これらのプリミティブを構成して、人間による追加のデモンストレーションなしで、新しい長期的なタスクを実行することができます。私たちの調査結果は、原始的なステアビリティが VLA ポリシーにおける継続的なスキル習得のための実用的な基盤となることを示しています。プロジェクトの Web サイト: https://insight-vla.github.io。

原文 (English)

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives via VLM plan decomposition and end-effector poses to enable VLA primitive steerability, and (2) a VLM-guided data flywheel that identifies missing primitives required to accomplish a novel task, autonomously attempts demonstrations of the missing primitives with VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set. We evaluate InSight across simulation and real-world manipulation tasks, including block flipping, drawer closing, sweeping, twisting, and pouring, without any human demonstrations of these target skills. Once learned, these primitives can be composed to execute novel, long-horizon tasks without additional human demonstrations. Our findings demonstrate that primitive steerability provides a practical foundation for continual skill acquisition in VLA policies. Project website: https://insight-vla.github.io.

13:00 JSTLLM/生成AI

ランダム ルール フォレスト (RRF): 非構造化データから成功を予測するための LLM 生成質問の解釈可能かつ管理可能なアンサンブル

一か八かのスクリーニング タスクの多くでは、非構造化テキストからまれな結果を予測する必要があり、エラーが発生するとコストが高くつき、決定は監査可能でなければなりません。大規模言語モデル (LLM) をエンドツーエンドの予測子としてではなく、単純な YES/NO の質問の生成器として使用する、解釈可能なアンサンブルである Random Rule Forest (RRF) を紹介します。各質問は弱い学習者として機能し、その回答は単純な単位重み付け投票によって監査可能な「青信号」スコアカードに結合されます。十分な独立した肯定的なシグナルは、成功の可能性が高いことを示します。私たちは、この意図的な単純化が、ポジティブな値が不足し、学習された重みの推定が難しい場合の堅牢なデフォルトであると主張します。 2 つの低基本レート ドメインで RRF を評価します。創業者のプロフィールからの初期段階のスタートアップスクリーニングでは、RRF は基本レートの数倍の精度を持つ透明なスコアカードを作成します (軽い専門家の入力によりさらに精度が上がります)。直接のプロンプトとは異なり、その動作点は直接制御できます。確立された第 I 相臨床試験ベンチマークでは、RRF は閾値に依存しない指標 PR-AUC および ROC-AUC で公表されているベースラインを上回っています。これらを総合すると、LLM が、透明性と競争力のある予測パフォーマンスを組み合わせて、一か八かのテキストベースの意思決定を行うための監査可能な特徴ジェネレーターとして機能できることがわかります。

原文 (English)

Random Rule Forest (RRF): Interpretable and Manageable Ensembles of LLM-Generated Questions for Predicting Success from Unstructured Data

Many high-stakes screening tasks require predicting rare outcomes from unstructured text, where errors are costly and decisions must be auditable. We introduce Random Rule Forest (RRF), an interpretable ensemble that uses a large language model (LLM) not as an end-to-end predictor but as a generator of simple YES/NO questions. Each question acts as a weak learner, and their responses are combined by a plain unit-weight vote into an auditable ``green-flags'' scorecard: enough independent positive signals indicate a higher chance of success. We argue this deliberate simplicity is a robust default when positives are scarce and learned weights are hard to estimate. We evaluate RRF in two low-base-rate domains. On early-stage startup screening from founder profiles, RRF produces a transparent scorecard whose precision is several times the base rate (with light expert input raising it further) and, unlike direct prompting, its operating point can be controlled directly. On an established Phase~I clinical-trial benchmark, RRF outperforms published baselines on the threshold-independent metrics PR-AUC and ROC-AUC. Together these show that LLMs can serve as auditable feature generators for high-stakes text-based decisions, combining transparency with competitive predictive performance.

13:00 JST研究/論文

TIP-Search: 不確実な負荷の下での市場予測のための時間予測可能な推論スケジューリング

リアルタイム市場予測サービスは、決定期限までに正確な予測を行う必要があります。遅れて配信された正しい予測は使用できません。 TIP-Search は、不確実な負荷の下での固定市場予測子に対する時間予測可能な推論スケジューリングを研究します。等角レイテンシー分位点の実現可能モデルをフィルタリングし、有限のワーカーを派遣し、シールドされた制約のあるオンライン専門家を使用して精度、キューのプレッシャー、期限のリスクをトレードします。最適化された展開可能なプールでは、TIP-Search は生の精度が 0.994、タイムリーな精度が 0.991 に達します。公式 TLOB FI-2010 h=10 では、TIP-Search++ はタイムリー精度を 0.156 から 0.239 に、期限満足度を 0.391 から 0.962 に向上させます。一致した h10 プロファイル システムのリプレイでは、OCO-ACPO は 0.303 タイムリー精度と 0.951 デッドライン満足度に達し、RAMSIS/SneakPeek/ユーティリティ スタイルのコンパレータを上回る $+0.00285$ タイムリー精度 ($p=0.0118$) と $+0.0146$ デッドライン満足度のペアのゲインを達成しました。 ($p=1.5{\times}10^{-5}$)。 SA-OCO-ACPO は、非定常ストレス下で CPO よりもタイムリー/デッドライン サービスを 0.188 ~ 0.417 改善します。このクレームはシステムのスケジュール結果であり、広範な LOB 分類子のリーダーボードではありません。

原文 (English)

TIP-Search: Time-Predictable Inference Scheduling for Market Prediction under Uncertain Load

Real-time market prediction services need correct predictions before a decision deadline; a correct prediction delivered late is not usable. TIP-Search studies time-predictable inference scheduling over fixed market predictors under uncertain load. It filters conformal latency-quantile feasible models, dispatches over finite workers, and uses shielded constrained online experts to trade accuracy, queue pressure, and deadline risk. On the optimized deployable pool, TIP-Search reaches 0.994 raw accuracy and 0.991 timely accuracy. On official TLOB FI-2010 h=10, TIP-Search++ raises timely accuracy from 0.156 to 0.239 and deadline satisfaction from 0.391 to 0.962. In matched h10 profiled systems replay, OCO-ACPO reaches 0.303 timely accuracy and 0.951 deadline satisfaction, with paired gains over RAMSIS/SneakPeek/utility-style comparators of $+0.00285$ timely accuracy ($p=0.0118$) and $+0.0146$ deadline satisfaction ($p=1.5{\times}10^{-5}$). SA-OCO-ACPO improves timely/deadline service by 0.188--0.417 over CPO under nonstationary stress. The claim is a systems scheduling result, not a broad LOB classifier leaderboard.

13:00 JST研究/論文

From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control

Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated,…

13:00 JST研究/論文

A global log for medical AI

Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure. Medicine's r…

13:00 JSTLLM/生成AILlamaQwen

Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs

Large language models (LLMs) often produce incorrect or outdated content after being employed. Efficient and accurate knowledge updates wit…

13:00 JSTエージェントビジネス/資金調達

Evolving Programmatic Skill Networks

We study continual skill acquisition in open-ended embodied environments where an agent must construct, refine, and reuse an expanding libr…

13:00 JST研究/論文

BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding

Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cro…

13:00 JSTLLM/生成AI

LLM-MINE: Large Language Model based Alzheimer's Disease and Related Dementias Phenotypes Mining from Clinical Notes

Accurate extraction of Alzheimer's Disease and Related Dementias (ADRD) phenotypes from electronic health records (EHR) is critical for ear…

13:00 JST研究/論文

Grounded Chess Reasoning in Language Models via Master Distillation

Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel.…

13:00 JSTLLM/生成AIエージェント

Subjective-Graph LLM Agents for Simulating Uncertainty in Classroom Social Perception

Social actors do not observe a common social world: each individual forms judgments from a partial and potentially distorted view of the su…

13:00 JST研究/論文

Riemann-Bench: A Benchmark for Moonshot Mathematics

Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficien…

13:00 JST研究/論文

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models ,…

13:00 JSTエージェント研究/論文

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accura…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPTNVIDIA

2.5-D Decomposition for LLM-Based Spatial Construction

Autonomous systems that build structures from natural-language instructions need reliable spatial reasoning, yet large language models (LLM…

13:00 JST研究/論文

生成計画モデルの効率的なテスト時間推論

生成モデルは AI 計画の強力なパラダイムとして登場しましたが、そのパフォーマンスは依然としてトレーニング データの分布によって制限されています。 1 つのアプローチは、テスト時の計算をスケーリングすることで、推論中に生成されるソリューションを改善することです。より効率的な代替方法は、推論プロセス自体を最適化することです。この論文では、古典的なオープンクローズド リスト (OCL) 検索の修正バージョンがまさにそのような効率的な推論手順を提供することを示します。私たちのアルゴリズムは、中間状態から高速ロールアウトを実行する生成モデルと、候補推論パス間で優先順位を付けるヒューリスティック モデルという 2 つの学習されたコンポーネントを相乗させます。主な貢献には、新しい探索制御メカニズムと、OCL フレームワーク内での学習済みモデルの統合が含まれます。複数の組み合わせ計画ドメインにわたって、私たちのアプローチは、計算効率とソリューションの品質において、ニューロシンボリック検索ベースラインと古典的ソルバーの両方を上回っています。

原文 (English)

Efficient Test-time Inference for Generative Planning Models with OCL Search

Generative models have emerged as a powerful paradigm for AI planning, yet their performance remains constrained by the training data distribution. One approach is to improve generated solutions during inference by scaling test-time compute. A more efficient alternative is to optimize the inference process itself. In this paper, we show that a modified version of a classical Open-Closed List (OCL) search provides just such an efficient inference procedure. Our algorithm synergizes two learned components: a generative model that performs fast rollouts from intermediate states and a heuristic model that prioritizes among candidate reasoning paths. Key contributions include novel exploration control mechanisms and integration of learned models within the OCL framework. Across multiple combinatorial planning domains, our approach outperforms both neurosymbolic search baselines and classical solvers in computational efficiency and solution quality.

13:00 JSTエージェント

TouchThinker: 大規模なデータとアクションを意識した表現を使用して、触覚的常識推論をオープンワールドに拡張する

接触は、肉体を持ったエージェントが物理世界を理解するための重要なモダリティです。最近の研究では、触覚常識推論のための言語システムに触覚信号が組み込まれていますが、そのようなシステムを現実的なオープンワールド設定に拡張することは、2 つの重要なボトルネックのため依然として困難です。(1) 現在の触覚推論データセットは形式と規模が制限されたままであり、触覚観察から物理的常識への推論に対する監視が不十分であり、伝達可能な触覚常識の学習を妨げています。 (2) 触覚信号は本質的に冗長でアクション固有ですが、既存の方法ではこれらの特性が見落とされることが多く、その結果、意味表現力が限られた非効率な表現が生じます。これらの制限に対処するために、私たちは、データと表現の両方の観点から触覚の常識的推論をオープンワールドに拡張する触覚言語フレームワークである TouchThinker を提案します。まず、\textbf{415} オブジェクト、\textbf{8} シナリオ、\textbf{7} センサー タイプをカバーする百万規模のマルチソース触覚推論データセットである TouchThinker-1M を構築し、オープンワールドの一般化のための強固なデータ基盤を提供します。さらに、より現実的で多様なタスクを備えたオープンワールドのベンチマークである TouchThinker-Bench を紹介します。次に、触覚表現の効率を向上させ、効率的な推論を可能にするアクション認識モデリングメカニズムを提案します。実験結果は、TouchThinker が複数のデータセットにわたって最先端のモデルに対して競争力のあるパフォーマンスを達成することを示しています。私たちのコードとデータセットは、https://github.com/lvkailin0118/TouchThinker で利用できるようになります。

原文 (English)

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) Tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering \textbf{415} objects, \textbf{8} scenarios, and \textbf{7} sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.

13:00 JSTLLM/生成AIエージェント研究/論文

EComAgentBench: 分散された隠れたインテントを使用した長期タスクに関するショッピング エージェントのベンチマーク

LLM ベースのショッピング エージェントが本番環境に入るにつれて、既存のベンチマークは、買い物客の要件がどのように届くか、つまりクエリで暗黙的に指定されるか、プロファイルに記録されるか、適切な質問がされた場合にのみ明らかにされるかを把握できません。事前に完全な意図を明らかにし、最終的な選択のみを評価するベンチマークでは、この長期にわたる課題を提起することも、エージェントがどの要件を逃したかを説明することもできません。このギャップに対処するために、実際の Amazon 製品とレビューに基づいた 662 のタスクのベンチマークである EComAgentBench を導入します。各タスクは、これらの要件を、目に見えるクエリ、ツールゲートのプロファイル、およびスクリプト化された説明に分散させます。エージェントは、隠れた意図を明らかにし、候補者を属性と照合して証拠を確認し、100 回のツール呼び出し以内に単一の製品にコミットする必要があります。さらに、入力され、ソースタグが付けられたルーブリックにより、すべてのタスクが評価され、各失敗の原因が要件とそのソースに帰されます。構築は自動化されていますが、信頼性が高く、テキストが生成され、すべてのサンプルが検証される前に、すべての回答がコード内で修正されます。 7 つのモデルを評価したところ、最も強力なモデルでも全体の精度は 57.1% にとどまっており、ルーブリックの満足度は、目に見えるソースから隠れたソースへと低下することが明らかになりました。全体として、私たちは EComAgentBench が、ショッピング エージェントを単一クエリ検索から長期にわたる信頼できる支援へと移行させるための再現可能な基盤として機能すると考えています。

原文 (English)

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose full intent upfront and grade only the final choice can neither pose this long-horizon challenge nor explain which requirement an agent missed. To address this gap, we introduce EComAgentBench, a benchmark of 662 tasks grounded in real Amazon products and reviews. Each task scatters these requirements across a visible query, a tool-gated profile, and scripted clarification; an agent must uncover hidden intent, verify candidates against attributes and review evidence, and commit to a single product within 100 tool calls. Moreover, typed, source-tagged rubrics grade every task, attributing each failure to a requirement and its source. Construction is automated yet reliable, with every answer fixed in code before any text is generated and every sample validated. Our evaluation of seven models reveals that even the strongest attains only 57.1% overall accuracy, and rubric satisfaction degrades from visible to hidden sources. Overall, we believe EComAgentBench will serve as a reproducible foundation for moving shopping agents from single-query search toward dependable assistance over long horizons.

13:00 JSTLLM/生成AI研究/論文

BIM-Edit: IFC ベースのビルディング インフォメーション モデリングのための大規模言語モデルのベンチマーク

大規模言語モデル (LLM) は、テキストの指示から設計アーティファクトを生成するために、コンピュータ支援設計 (CAD) にますます適用されています。エンジニアリングの実践では、これには新しいジオメトリを作成するだけではなく、モデルが既存のシーンを理解し、正しく編集し、セマンティクスと関係を保持する必要もあります。ただし、多くの CAD ベンチマークは、既存のモデルを編集するのではなく、新しいモデルを作成することに重点を置き、主に幾何学的正確さを評価します。 Industry Foundation Classes (IFC) 形式で表される Building Information Model (BIM) の自然言語編集に関する LLM を評価するためのベンチマークである BIM-Edit を紹介します。 BIM は、建築モデルがジオメトリをセマンティックおよびリレーショナル構造とともにエンコードするため、困難なテストベッドを提供します。 BIM-Edit には、11 の現実的な建築モデルと 36 の合成シーンにわたる 324 の編集タスクが含まれています。タスクは 3 つの命令カテゴリ (直接、空間、トポロジカル) を使用して表現され、明示的な編集とシーンに基づいた編集の両方をカバーします。私たちは、幾何学的精度、意味論的妥当性、トポロジー的一貫性という 3 つの次元に沿って出力を評価します。評価された LLM 全体で、最もパフォーマンスの高いモデルは、3 つの指標全体で 49.5% の平均スコアしか達成できず、タスクの 3.4% を超える問題を完全に解決するモデルはありません。これらの結果は、現在の LLM 機能と構造化エンジニアリング設計ワークフローの要件との間に大きなギャップがあることを示しています。

原文 (English)

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engineering practice, this requires more than creating new geometry, models must also understand existing scenes, edit them correctly, and preserve semantics and relations. However, many CAD benchmarks focus on creating new models rather than editing existing ones, and mostly evaluate geometric correctness. We introduce BIM-Edit, a benchmark for evaluating LLMs on natural-language editing of Building Information Models (BIM) represented in the Industry Foundation Classes (IFC) format. BIM provides a challenging testbed because building models encode geometry together with semantic and relational structure. BIM-Edit contains 324 editing tasks spanning 11 realistic building models and 36 synthetic scenes. Tasks are expressed using three instruction categories - direct, spatial, and topological - covering both explicit and scene-grounded edits. We evaluate outputs along three dimensions: geometric accuracy, semantic validity, and topological consistency. Across evaluated LLMs, the best-performing model achieves only 49.5% average score across the three metrics, and no model fully solves more than 3.4% of tasks. These results demonstrate a substantial gap between current LLM capabilities and the requirements of structured engineering design workflows.

13:00 JST研究/論文Grok

Repeated Shared Access Enables Grokking, but Edit Propagation Depends on an Addressable Memory

We study factual edit propagation in a controlled synthetic knowledge-graph QA setting using a 2x2 grid that crosses loop recurrence with s…

13:00 JSTLLM/生成AI

When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models

Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two out…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文AnthropicClaudeOpenAIGPT / ChatGPTGoogleGeminiAlibabaQwen

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…

13:00 JSTLLM/生成AI研究/論文

HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs

Logical reasoning is essential for reliable AI, yet existing benchmarks are largely first-order-logic-centric, focusing on object-level ded…

13:00 JST研究/論文

Invariant Graph Representations for Continuous-Time Dynamic Graphs Under Distribution Shifts

Continuous-Time Dynamic Graphs (CTDGs) enable fine-grained modeling of evolving relational systems. However, most existing CTDG representat…

13:00 JSTエージェント

When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments

Can AI Agents simulate real-world trading environments to investigate the impact of external factors on stock trading activities (e.g., mac…

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the develo…

13:00 JST研究/論文

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute. To reduce costs…

13:00 JST研究/論文

Impatient Bandits: Optimizing for the Long-Term Without Delay

Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions

Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as si…

13:00 JSTLLM/生成AI

Societal Alignment Frameworks Can Improve LLM Alignment

Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared value…

13:00 JSTロボティクス

Reward-Centered ReST-MCTS: A Robust Decision-Making Framework for Robotic Manipulation in High Uncertainty Environments

Monte Carlo tree search is attractive for robotic manipulation because it can improve action selection through simulation without requiring…

13:00 JSTLLM/生成AI

Ensemble Learning for Large Language Models in Text and Code Generation: A Survey

Generative Pretrained Transformers (GPTs) are foundational Large Language Models (LLMs) for text generation. However, individual LLMs often…

13:00 JSTエージェント

Multimedia and Visual Analytics in the Agentic Era

Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have…

13:00 JSTロボティクス

MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems

Robots need task planning methods to achieve goals that require more than one action. Recently, large pretrained models have demonstrated i…

13:00 JST研究/論文

Minimisation of Quasar-Convex Functions Using Random Zeroth-Order Oracles

This paper explores the performance of a random Gaussian smoothing zeroth-order (ZO) scheme for minimising quasar-convex (QC) and strongly…

13:00 JST画像/動画生成

SEAL: Searching Expandable Architectures for Incremental Learning

Incremental learning is a machine learning paradigm where a model learns from a sequential stream of tasks. This setting poses a key challe…

13:00 JST研究/論文

Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings

We propose a novel benchmarking methodology for graph neural networks (GNNs) based on the graph alignment problem, a combinatorial optimiza…

13:00 JST画像/動画生成

Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering

Photorealistic volumetric rendering of CT scans greatly benefits clinical workflows, yet neural approaches such as Neural Radiance Fields (…

13:00 JSTLLM/生成AI

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training

Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation. However, expos…

13:00 JST研究/論文

FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation

Industrial signal analysis is hindered by severe data heterogeneity, which we characterize as the M5 problem. Existing solutions rely on sp…

13:00 JSTLLM/生成AIGemini

Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules

Knowledge graphs (KGs) can be enhanced through rule mining; however, the resulting logical rules are often difficult for humans to interpre…

13:00 JSTLLM/生成AI

FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection

Signature-based Intrusion Detection Systems (IDS) detect malicious activity by matching network or host events against predefined rules. Se…

13:00 JSTLLM/生成AI

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor t…

13:00 JSTLLM/生成AI

VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing

The energy cost of Large Language Model (LLM) inference is rapidly becoming a barrier to sustainable and scalable deployment. Although mode…

13:00 JST画像/動画生成

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Light…

13:00 JSTエージェント

ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms

Progress in computational science depends on complex numerical workflows that must faithfully encode physical laws, yet translating concept…

13:00 JST研究/論文

Computing Evolutionarily Stable Strategies in Imperfect-Information Games

We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…

13:00 JST研究/論文

EMFusion: Uncertainty-Aware Conditional Diffusion Model for Multivariate Narrow-band Exposure Forecasting

The rapid growth in wireless infrastructure has increased the need to accurately estimate and forecast electromagnetic field (EMF) levels t…

13:00 JST研究/論文

Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection

Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle…

13:00 JST研究/論文

Disentangling Aleatoric and Epistemic Uncertainty in Physics-Informed Neural Networks. Application to Insulation Material Degradation Prognostics

Physics-Informed Neural Networks (PINNs) provide a framework for integrating physical laws with data. However, their application to Prognos…

13:00 JSTLLM/生成AIエージェント

The $\mathbf{P}$-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs

Modern AI agents increasingly rely on search infrastructure to execute complex, neuro-symbolic reasoning workflows. These workflows often c…

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達研究/論文

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of aut…

13:00 JSTエージェント

自律型 O-RAN に向けて: リアルタイム ネットワーク制御および管理のためのマルチスケール エージェント AI フレームワーク

オープン無線アクセス ネットワーク (O-RAN) は、分散されたソフトウェア駆動のコンポーネントとオープン インターフェイスを通じて柔軟な 6G ネットワーク アクセスを約束しますが、このプログラマビリティにより運用の複雑さも増大します。複数の制御ループがサービス管理層と RAN インテリジェント コントローラー (RIC) 全体で共存しますが、個別に開発された制御アプリケーションは意図しない方法で相互作用する可能性があります。同時に、生成型人工知能 (AI) の最近の進歩により、孤立した AI モデルから、目標を解釈し、複数のモデルと制御機能を調整し、時間の経過とともに動作を適応させることができるエージェント AI システムへの移行が可能になりました。この記事では、非リアルタイム (Non-RT)、準リアルタイム (Near-RT)、およびリアルタイム (RT) の制御ループにわたる調整された階層として RAN インテリジェンスを組織化する、O-RAN 用のマルチスケール エージェント AI フレームワークを提案します。 (i) 非 RT RIC の大規模言語モデル (LLM) エージェントは、オペレーターの意図をポリシーに変換し、モデルのライフサイクルを管理します。 (ii) Near-RT RIC の Small Language Model (SLM) エージェントは、低遅延の最適化を実行し、既存の制御アプリケーションをアクティブ化、調整、または無効化できます。 (iii) 分散ユニット近くのワイヤレス物理層基盤モデル (WPFM) エージェントは、エア インターフェイスに近い高速推論を提供します。これらのエージェントが標準化された O-RAN インターフェイスとテレメトリを通じてどのように連携するかを説明します。オープンソース モデル、ソフトウェア、データセットに基づいて構築された概念実証の実装を使用して、非定常条件下での堅牢な動作とインテント駆動型のスライス リソース制御という 2 つの代表的なシナリオで提案されたエージェント アプローチを実証します。

原文 (English)

Toward Autonomous O-RAN: A Multi-Scale Agentic AI Framework for Real-Time Network Control and Management

Open Radio Access Networks (O-RAN) promise flexible 6G network access through disaggregated, software-driven components and open interfaces, but this programmability also increases operational complexity. Multiple control loops coexist across the service management layer and RAN Intelligent Controller (RIC), while independently developed control applications can interact in unintended ways. In parallel, recent advances in generative Artificial Intelligence (AI) are enabling a shift from isolated AI models toward agentic AI systems that can interpret goals, coordinate multiple models and control functions, and adapt their behavior over time. This article proposes a multi-scale agentic AI framework for O-RAN that organizes RAN intelligence as a coordinated hierarchy across the Non-Real-Time (Non-RT), Near-Real-Time (Near-RT), and Real-Time (RT) control loops: (i) A Large Language Model (LLM) agent in the Non-RT RIC translates operator intent into policies and governs model lifecycles. (ii) Small Language Model (SLM) agents in the Near-RT RIC execute low-latency optimization and can activate, tune, or disable existing control applications; and (iii) Wireless Physical-layer Foundation Model (WPFM) agents near the distributed unit provide fast inference close to the air interface. We describe how these agents cooperate through standardized O-RAN interfaces and telemetry. Using a proof-of-concept implementation built on open-source models, software, and datasets, we demonstrate the proposed agentic approach in two representative scenarios: robust operation under non-stationary conditions and intent-driven slice resource control.

13:00 JST研究/論文

Event-Grounded Question Answering over Long Audio via Structured Retrieval

Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding. Current large audio-lang…

13:00 JST研究/論文

MyoInteract: A Framework for Fast Prototyping of Biomechanical HCI Tasks using Reinforcement Learning

Reinforcement learning (RL)-based biomechanical simulations have the potential to revolutionise HCI research and interaction design, but cu…

13:00 JST研究/論文

Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators

Neural network accelerators have been widely applied to edge devices for complex tasks like object tracking, image recognition, etc. Previo…

13:00 JST研究/論文

No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason

Generative AI can convert uncertainty into authoritative-seeming verdicts, intensifying the hypersuasive force of automated speech and disp…

13:00 JSTLLM/生成AIビジネス/資金調達

An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production

Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the a…

13:00 JST画像/動画生成ロボティクス

CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance

We introduce CRAFT hand, a tendon-driven anthropomorphic hand with hybrid hard-soft compliance for contact-rich manipulation. The design is…

13:00 JST研究/論文

AI-Driven Predictive Maintenance with Environmental Context Integration for Connected Vehicles: Simulation, Benchmarking, and Field Validation

Predictive maintenance for connected vehicles offers the potential to reduce unexpected breakdowns and improve fleet reliability, but most…

13:00 JST画像/動画生成

HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction

Pathology reports are structured, multi-granular documents encoding diagnostic conclusions, histological grades, and ancillary test results…

13:00 JSTLLM/生成AI

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polish…

13:00 JSTLLM/生成AI

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale qu…

13:00 JST研究/論文

THEIA: Learning Complete Kleene Three-Valued Logic in a Pure-Neural Modular Architecture

We present THEIA, a 2.75M-parameter modular neural architecture that learns the complete Kleene three-valued logic (K3) truth table from ta…

13:00 JST画像/動画生成エージェント

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation

Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While rec…

13:00 JSTLLM/生成AI

Fix Initial Programs and Iteratively Refine Repair Instructions Toward Non-Elimination Multi-Turn Program Correction

Recent work on large language models (LLMs) has emphasized the importance of scaling inference compute. From this perspective, the state-of…

13:00 JSTLLM/生成AI

DynamicPO: Dynamic Preference Optimization for Recommendation

In large language model (LLM)-based recommendation systems, direct preference optimization (DPO) effectively aligns recommendations with us…

13:00 JST研究/論文

Ensemble Distributionally Robust Bayesian Optimisation with Continuous Context

We study Bayesian Optimisation (BO) in settings where the objective function is influenced by uncontrollable environmental contexts governe…

13:00 JST画像/動画生成エージェント

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely h…

13:00 JST研究/論文

変分オートエンコーダにおける定常コラプスに対するシンプレックス証人証明書

私たちは、変分オートエンコーダーにおける正確な定数の崩壊を研究します。つまり、決定論的なエンコーダーの平均が入力から独立するようになります。事前分布は標準のガウス分布のままです。 VAE トレーニングの前に、データの GMM ベースのビューから事後固定教師を選択し、固定潜在のみのシンプレックス監視をエンコーダー平均に添付します。この構築により、2 つのリンクされたオブジェクトが生成されます。 1 つ目は証明書です。目撃者の予測が教師の最良の定数予測子を改善する場合、エンコーダーの平均は入力に依存しない定数になることはできません。 2 つ目は局所的なエスケープ方向です。崩壊した多様体では、教師残差によってアライメント損失に対するサンプル依存の下降方向が与えられます。あらゆるフルサポート教師事後分布の場合、同じジオメトリにより、教師と証人の位置合わせエラーがゼロの閉じた形式の潜在コードも得られます。そのスケーリングされたバージョンは、定数予測子から正確な教師コードまでのマージン エネルギー パスを追跡し、保護された目撃部分空間内の非崩壊を定量化します。 MNIST、CIFAR-10、および CIFAR-100 でメソッドをインスタンス化します。教師なしの PCA-GMM 教師を検索すると、バニラ VAE は CIFAR-10 および CIFAR-100 の 5 つのシードすべてで教師証人証明書に不合格ですが、RST バリアントは 5 つのシードすべてに合格します。 \(\beta_{\mathrm{KL}}\in\{2,4,8\}\) を使用した崩壊ストレス設定では、バニラ VAE はすべてのシードで再び失敗しますが、RST-alpha-prefit は証明書陽性のままです。両方の自然画像データセットのエスケープ トラジェクトリは、低マージン初期化からウィットネス マージンを増加させ、非ゼロの教師誘発勾配ノルムを示します。分析は、エンコーダ平均値の正確な一定の崩壊に限定されます。生成品質、デコーダの使用、およびその他の崩壊モードについては、別個の問題として残ります。

原文 (English)

A Simplex Witness Certificate and Escape Force for Constant Collapse in Variational Autoencoders

We study exact constant collapse in variational autoencoders: the deterministic encoder mean becomes independent of the input. The prior remains the standard Gaussian. Before VAE training, we select a fixed teacher posterior from a GMM-based view of the data and attach a fixed latent-only simplex witness to the encoder mean. This construction yields two linked objects. The first is a certificate: if the witness prediction improves on the best constant predictor of the teacher, the encoder mean cannot be input-independent constant. The second is a local escape direction: on the collapsed manifold, the teacher residual gives a sample-dependent descent direction for the alignment loss. For any full-support teacher posterior, the same geometry also gives a closed-form latent code with zero teacher-witness alignment error. Its scaled versions trace a margin-energy path from the constant predictor to the exact teacher code, which quantifies non-collapse inside the protected witness subspace. We instantiate the method on MNIST, CIFAR-10, and CIFAR-100. With searched unsupervised PCA-GMM teachers, vanilla VAEs fail the teacher-witness certificate in all five seeds on CIFAR-10 and CIFAR-100, while RST variants pass in all five seeds. Under collapse-stress settings with \(\beta_{\mathrm{KL}}\in\{2,4,8\}\), vanilla VAE again fails in all seeds, whereas RST-alpha-prefit remains certificate-positive. Escape trajectories on both natural-image datasets increase the witness margin from a low-margin initialization and exhibit nonzero teacher-induced gradient norms. The analysis is confined to exact constant collapse of the encoder mean; generation quality, decoder use, and other collapse modes remain separate questions.

13:00 JSTLLM/生成AIエージェント

Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment

Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in…

13:00 JST研究/論文

トレーニング可能なメタマテリアル特性としてのセンシングインテリジェンス

生物学的システムでは、感知は脳だけで行われるわけではありません。身体は、外部刺激が神経信号に変換される前に、外部刺激を変形、振動、フィルタリングします。工学的に設計されたシステムでは、この処理負荷は主にエレクトロニクスと計算に課せられますが、機械本体は通常、強度と安定性のみを目的として設計されています。ここでは、訓練可能な身体の特性としての感覚知性を紹介します。メタマテリアルの幾何学形状を最適化して、外部刺激をニューラル ネットワークが解釈しやすい内部信号に再形成できることを示します。この物理的な前処理を手動で設計するのではなく、微分可能なシミュレーションを通じてセンシング損失を身体の設計パラメータに逆伝播させることで、ニューラル ネットワークに自身の身体をセンシング用にトレーニングさせます。数値的および実験的なセンシング シナリオ全体で、最適化された本体によりセンシング精度が最大 5 倍向上し、必要な電子センサーの数がほぼ 1 桁削減されます。

原文 (English)

Sensing Intelligence as a Trainable Metamaterial Property

In biological systems, sensing is not performed by the brain alone: the body deforms, vibrates, and filters external stimuli before they are transduced into neural signals. In engineered systems, this processing burden is placed largely on electronics and computation, while the mechanical body is usually designed only for strength and stability. Here, we present sensing intelligence as a trainable property of the body. We show that the geometry of a metamaterial can be optimized to reshape external stimuli into internal signals that are easier for a neural network to interpret. Rather than hand-designing this physical preprocessing, we let the neural network train its own body for sensing by backpropagating the sensing loss to the body's design parameters through differentiable simulation. Across numerical and experimental sensing scenarios, the optimized body improves sensing accuracy by up to fivefold or reduces the number of required electronic sensors by nearly an order of magnitude.

13:00 JSTLLM/生成AIエージェント

スキルが増えればエージェントは劣る?スキル ライブラリを拡張するときにスキル シャドウイングによりパフォーマンスが低下する

スキル ライブラリを使用すると、LLM エージェントはタスク固有の指示をオンデマンドで読み込むことができるため、専門知識のないユーザーは、どのスキルが存在するか、どのように機能するかを知らなくても、自然言語を通じてドメイン固有のタスクを解決できます。ただし、ライブラリが大きくなるにつれて、パフォーマンスは低下します。役立つスキルの小さなセットから 202 のスキル ライブラリに拡張すると、最大 21\% 低下します。この研究では、このパフォーマンスの低下を、既知の役立つスキルのライブラリをロードするときと完全なライブラリをロードするときとの間の合格率の低下として定式化します。さらに、スキルの呼び出し (軌道中にエージェントがどのスキルを選択するか) を条件付けすることで合格率の低下を 2 つの効果に分解することを提案します。 \emph{スキル シャドウイング} (ライブラリが拡張するにつれてエージェントが間違ったスキルを選択する頻度が高くなります)、および \emph{コンテキスト オーバーヘッド} (選択が正しい場合でも、拡大されたコンテキストによって実行が低下する) です。両方の効果の上限を導き出し、合格率の低下に対する影響の大きさを特徴付けます。効果とその上限についての経験的な推定によると、\emph{スキル シャドウイング} 効果はライブラリのサイズとともに増大し、パフォーマンス低下に大きく寄与するのに対し、\emph{コンテキスト オーバーヘッド} 効果は依然として小さく、ゼロと区別がつかないことがわかります。この観察された非対称性は、スキル ライブラリを拡張する際の主なボトルネックは、拡大されたコンテキストではなく、スキル選択の失敗であることを示しています。

原文 (English)

More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries

Skill libraries allow LLM agents to load task-specific instructions on demand, letting non-expert users solve domain-specific tasks through natural language without knowing which skills exist or how they work. However, performance degrades as libraries grow -- by up to 21\% when scaling from a small set of helpful skills to a 202-skill library. In this work, we formulate this performance degradation as the pass rate drop between loading a library of known-helpful skills and the full library. Moreover, we propose to decompose the pass rate drop by conditioning on the skill(s) invocation -- which skills the agent selects during a trajectory -- into two effects: \emph{skill shadowing}, where the agent selects wrong skills more often as the library expands, and \emph{context overhead}, where the enlarged context degrades execution even when selection is correct. We derive upper bounds on both effects to characterize their magnitudes of impacts to the pass rate drop. Our empirical estimates of the effects and their upper bounds both show that the \emph{skill shadowing} effect grows with library size and significantly contributes to the performance degradation, whereas the \emph{context overhead} effect remains small and indistinguishable from zero. This observed asymmetry establishes that the skill selection failure, not the enlarged context, is the primary bottleneck when expanding the skill libraries.

13:00 JSTLLM/生成AI画像/動画生成エージェント研究/論文

VISTA: Visual Spec-to-Web-App コーディング エージェントのエンドツーエンド ベンチマーク

ここでは、LLM ベースのエージェントのエンドツーエンドの Web アプリ生成機能を評価するためのベンチマークである VISTA (VIsual Spec-To-App Benchmark) を紹介します。アルゴリズム タスクに焦点を当てた以前のコード生成ベンチマークとは異なり、VISTA は現実的な UI 中心の開発をターゲットにしており、エージェントは過少指定された入力から機能的で視覚的に一貫したアプリケーションを生成する必要があります。視覚的/構造的忠実度およびスタック制約という 2 つの軸に沿って変化する 5 つのプロンプト情報条件を定義します。(1) 自由なスタック選択によるテキストのみ、(2) 3 つの指定されたスタック下の参照スクリーンショットを含むテキスト、(3) 自由なスタック選択による参照スクリーンショットを含むテキスト、(4) 単一の指定されたスタック下のスクリーンショットおよびプルーニングされた Figma 構造を含むテキスト、(5) 自由なスタック選択によるスクリーンショットおよびプルーニングされた Figma 構造を含むテキスト。堅牢な評価を可能にするために、ベンチマークの各ページにはインタラクティブな UI コンポーネントと約 3 つのビジュアル アンカー ポイントで手動で注釈が付けられ、オープンエンド コード生成設定における Playwright などのスクリプト ベースのテスト ツールのよく知られた制限に対処します。評価では、DOM に基づいた参照マッチング、動作固有のブラウザ テスト、および CLIP ベースの視覚的類似性を組み合わせて、構造の整合性、動作の完全性、および全体的な視覚的な忠実度を共同で測定します。 VISTA を使用して、2 つのモデル ファミリと 2 つのハーネスから描画された 4 つのエージェント システムを評価しました。その結果、入力条件とエージェントの両方で視覚的な忠実性と機能の正確さが部分的に切り離されており、エージェントの編集スタイルは大きく変化しますが、タスクの品質とはほぼ直交していることがわかりました。 VISTA は、エージェントベースのソフトウェア エンジニアリング研究を推進するための厳密で再現可能な基盤を確立します。

原文 (English)

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike prior code generation benchmarks that focus on algorithmic tasks, VISTA targets realistic UI-centric development, where agents must produce functional, visually coherent applications from underspecified inputs. We define five prompt-information conditions that vary along two axes, visual/structural fidelity and stack constraint: (1) text only with free stack choice, (2) text with reference screenshots under three specified stacks, (3) text with reference screenshots under free stack choice, (4) text with screenshots and pruned Figma structure under a single specified stack, and (5) text with screenshots and pruned Figma structure under free stack choice. To enable robust evaluation, each page in the benchmark is manually annotated with interactive UI components and around three visual anchor points, addressing the well-known limitations of script-based testing tools such as Playwright in open-ended code generation settings. Evaluation combines DOM-grounded reference matching, behavior-specific browser tests, and CLIP-based visual similarity, jointly measuring structural alignment, behavioral completeness, and overall visual fidelity. We use VISTA to assess four agent systems drawn from two model families and two harnesses, finding that visual fidelity and functional correctness are partially decoupled across both input conditions and agents, and that agent editing style varies sharply but is largely orthogonal to task quality. VISTA establishes a rigorous and reproducible foundation for advancing agent-based software engineering research. Code is available at https://github.com/kaboider/VISTA_Bench.

13:00 JST研究/論文

QSignAI: 科学のための AI と AI のための科学の交差点における量子ランダムネスシード ID 署名

2024~2025年のノーベル賞とチューリング賞は、AIと量子科学を同時に評価した。しかし、これらの流れを一般公開するために導入されたシステムはまだありません。このペーパーでは、リアルタイム イベント参加システムにおける双方向の AI 量子関係を実証する実稼働環境に導入されたプラットフォームである QSignAI について説明します。私たちは 3 つの質問に取り組みます。2 ソース抽出器による量子ランダム性の生成は、許容可能な遅延で AI 駆動のソーシャル プラットフォームに埋め込むことができるか。 AIボットは量子現象を一般の聴衆が知覚的に判読できるようにすることができるか。そして、その組み合わせたシステムは実際に機能するのでしょうか?会話型ボットは、SV1 および DM1 シミュレーターでの独立した単一量子ビットのアダマール測定と 2 量子ビットのベル状態を介したテプリッツの 2 ソース抽出器で構成される量子パイプラインを介して各参加者の最初のメッセージをルーティングし、参加者ごとに固有の量子ランダムネスシード ID 署名を生成します。最初の 2 つの質問は、システム アーキテクチャとライブ イベントからの導入の定性的な証拠を通じて解決されます。 3 番目は実稼働デプロイメントの成功によるものです。現在のデプロイではクラウド量子シミュレーターが使用されています。物理 QPU のランダム性は短期的な拡張です。測定可能なベンチマークは、将来の優先課題として特定されます。

原文 (English)

QSignAI: Quantum-Randomness-Seeded Identity Signatures at the Intersection of AI for Science and Science for AI

The 2024-2025 Nobel and Turing awards recognised AI and quantum science simultaneously. Yet no deployed system has brought these streams together for the public. This paper presents QSignAI, a production-deployed platform demonstrating a bidirectional AI-quantum relationship in a real-time event participation system. We address three questions: can quantum-randomness generation via a two-source extractor be embedded in an AI-driven social platform with acceptable latency; can an AI bot make quantum phenomena perceptually legible to general audiences; and does the combined system work in practice? A conversational bot routes each participant's first message through a quantum pipeline comprising a Toeplitz two-source extractor over independent single-qubit Hadamard measurements on SV1 and DM1 simulators, plus a 2-qubit Bell state, producing a unique quantum-randomness-seeded identity signature per participant. The first two questions are answered through system architecture and qualitative deployment evidence from live events; the third through successful production deployment. The current deployment uses cloud quantum simulators; physical QPU randomness is the near-term extension. Measurable benchmarks are identified as priority future work.

13:00 JST画像/動画生成ロボティクスNVIDIA

Cosmos 3: Omnimodal World Models for Physical AI

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and actio…

13:00 JSTLLM/生成AI

ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information

Asynchronous reinforcement learning can improve language-model post-training throughput by decoupling response generation from policy optim…

13:00 JSTLLM/生成AIエージェント

A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs

Multi-Document Summarization (MDS) plays a critical role in distilling essential information from collections of textual data. Existing app…

13:00 JST画像/動画生成

Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing

Purpose: Digitization projects in the humanities produce vast, heterogeneous archives of historical documents, making manual sorting imprac…

13:00 JST研究/論文

チームティーチングトークの AI 主導分析: 経験、コホート、学習デザインにわたる音響パターン

教室のコホートが拡大するにつれて、複数の教師の専門知識と教育的観点を統合するためにチームティーチングがますます使用されています。しかし、チームティーチングが実際にどのように展開されるか、特に経験レベル、生徒集団、学習課題設計による教師の貢献の違いについての経験的理解は限られています。チームティーチングに関するこれまでの研究は、主に遡及的な自己報告や小規模な観察に依存しており、チームティーチングが実行されるミクロレベルのプロセスについての洞察は限られていました。教師の話は、これらのプロセスに関する拡張可能なレンズを提供します。個人の教育現場での研究では、音声の音響的特徴(声質、イントネーション、音量など)が生徒の学習を形作る可能性があることが示されていますが、チーム教育の現場での証拠は依然として不足しています。さらに、手動による観察や文字起こしによるこのような特徴の把握は、複数の教師が長時間のセッションや空間的場所にわたって話すチームティーチングの教室では特に困難であり、自動化なしでは拡張性が制限されます。この論文は、空間教育理論とチームティーチング研究に基づいて、チームティーチング環境における教室での会話を分析するための AI ベースの音声処理アプローチを紹介します。私たちは、12 人の教師が参加した学部および大学院での 36 の記録されたセッションを分析しました。空間教育行動がコード化され、音響特徴が抽出されて、教師の経験、生徒コホート、学習タスク設計全体の変動が調べられました。その結果、特にラウドネスダイナミクスにおける体系的な違いが明らかになりました。経験豊富な教師、学部のクラス、および共同学習タスクでは、より大きなラウドネス変動が見られ、重要な情報を前面に出し、教室での対話と参加をサポートするために、より頻繁に音量を調節していることが示唆されました。

原文 (English)

AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design

As classroom cohorts expand, team teaching is increasingly used to integrate the expertise and pedagogical perspectives of multiple teachers. Yet, there is limited empirical understanding of how team teaching unfolds in practice, particularly regarding differences in teachers' contributions across experience levels, student cohorts, and learning task design. Prior research on team teaching has largely relied on retrospective self-reports or small-scale observations, offering limited insight into the micro-level processes through which team teaching is enacted. Teacher talk offers a scalable lens on these processes. While research in individual teaching contexts shows that acoustic features of speech (e.g., voice quality, intonation, and loudness) can shape student learning, evidence from team-teaching settings remains scarce. Moreover, capturing such features through manual observation or transcription is especially challenging in team-teaching classrooms, where multiple teachers speak across extended sessions and spatial locations, limiting scalability without automation. Grounded in spatial pedagogy theory and team-teaching research, this paper presents an AI-based speech processing approach to analyse classroom talk in team-teaching settings. We analysed 36 recorded undergraduate and postgraduate sessions involving 12 teachers. Spatial pedagogy behaviours were coded and acoustic features extracted to examine variation across teachers' experience, student cohorts, and the learning task design. The results reveal systematic differences, most notably in loudness dynamics: high-experience teachers, undergraduate classes and collaborative learning tasks exhibited greater loudness variation, suggesting more frequent modulation of volume to foreground key information and support classroom interaction and engagement.

13:00 JST研究/論文

FedSteer: Taming Extreme Gradient Staleness in Federated Learning with Corrective Projections and Caching

Federated learning (FL) is often subject to aggregation variance if clients do not consistently participate in training rounds. While reusi…

13:00 JST画像/動画生成ビジネス/資金調達

Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata

AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends local acceptance testing and ongoing drif…

13:00 JSTエージェントAnthropicOpenAIGoogle

AgentRivet: an automated system for producing Rivet routines from journal publications

Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measuremen…

13:00 JST研究/論文

Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling

Pairwise comparison is the gold standard for subjective ranking tasks; however, exhaustive annotation requires a massive number of human co…

13:00 JSTLLM/生成AIエージェント

Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but…

13:00 JSTLLM/生成AI

Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts

Small-scale Large Language Models (LLMs) natively default to literal semantic interpretations, making few-shot irony detection a persistent…

13:00 JST研究/論文

量子シネマ: 生成世界モデルを介した量子コンピューティング ハードウェアのインタラクティブな映画的探索

量子コンピューティングは科学と産業全体に革新的な進歩を約束しますが、これらの計算を可能にする物理的ハードウェアは依然として一般の人々には見えません。量子プロセッサは絶対零度に近い温度で密閉された希釈冷蔵庫内で動作するため、直接観察することは不可能です。量子コンピューティングの増大する社会的影響とそれを視覚化する一般の人々の能力との間のこの「想像力のギャップ」は、量子リテラシーと労働力の育成にとって大きな障壁となっています。私たちは、オープンソースのブラウザベースのインタラクティブ アプリケーションである Quantum Cinema を紹介します。これは、生成世界モデルを使用して、目に見えない量子ハードウェアを探索可能な映画のような体験に変換することで、このギャップを埋めます。 Quantum Cinema は、ノーベル賞を受賞した量子もつれの基礎科学から、3 つの主要な量子コンピューティング アーキテクチャ (トラップ イオン、中性原子、超伝導システム) への精選されたビデオ紹介を経て、目に見えない量子現象を観察可能にする没入型の 3 次元生成世界、そして最後に実際の量子デバイスの仕様に基づいたインタラクティブなレーダー チャートの比較まで、4 幕の物語を通してユーザーをガイドします。すべての 3 次元環境は、WorldLabs の生成ワールド モデル プラットフォームを使用して生成され、アマゾン ウェブ サービス (AWS) Braket 量子ハードウェアから厳選されたメトリクスに科学的に基づいています。 Quantum Cinema には、インストール、特殊なハードウェア、量子コンピューティングの知識は必要ありません。これは、プラットフォームの複製または拡張を求める学者や開発者と、さまざまな聴衆に量子ハードウェアを説明するための直感的なツールを求める教育者、研究者、科学コミュニケーターという 2 つの異なるコミュニティにサービスを提供するように設計されています。このペーパーでは、システム アーキテクチャ、生成ワールド モデル パイプライン、両方のコミュニティの使用例、および将来の作業の方向性について説明します。

原文 (English)

Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models

Quantum computing promises transformative advances across science and industry, yet the physical hardware that enables these computations remains invisible to the public: quantum processors operate inside sealed dilution refrigerators at temperatures near absolute zero, making direct observation impossible. This "imagination gap" between quantum computing's growing societal impact and the public's ability to visualize it represents a significant barrier to quantum literacy and workforce development. We present Quantum Cinema, an open-source, browser-based interactive application that closes this gap by transforming invisible quantum hardware into explorable, cinematic experiences using generative world models. Quantum Cinema guides users through a four-act narrative -- from the foundational Nobel Prize-winning science of quantum entanglement, through curated video introductions to three major quantum computing architectures (trapped-ion, neutral-atom, and superconducting systems), into immersive three-dimensional generative worlds that make invisible quantum phenomena observable, and finally to interactive radar-chart comparisons grounded in real quantum device specifications. All three-dimensional environments are generated using WorldLabs' generative world model platform and are scientifically grounded in curated metrics from Amazon Web Services (AWS) Braket quantum hardware. Quantum Cinema requires no installation, no specialized hardware, and no quantum computing background. It is designed to serve two distinct communities: scholars and developers seeking to replicate or extend the platform, and educators, researchers, and science communicators seeking an intuitive tool for explaining quantum hardware to diverse audiences. This paper describes the system architecture, the generative world model pipeline, use cases for both communities, and directions for future work.

13:00 JSTLLM/生成AI研究/論文

LLM ベースの A/B テストの統計的基礎: 人間の因果推論のための代理フレームワーク

組織や研究者は、実験をより迅速かつ低コストで行うことを期待して、A/B テストに人間の参加者の代わりに大規模言語モデル (LLM) を使用することへの関心が高まっています。私たちは、LLM の結果に基づいて推定された治療効果が、対象となるヒト集団に対して測定されたであろう効果をいつ回復するかを研究します。 LLM と人間の結果の間の分布が同等であれば、標準推定量は有効になりますが、非現実的です。したがって、私たちはサロゲートエンドポイント理論を LLM に適応させる統計的フレームワークを開発します。このフレームワークは、LLM のアウトカムをヒトのアウトカムに合わせて調整することで、分布上の同等性よりも劣る代理出産および比較可能性の条件下での平均的な治療効果を特定することを示しています。これらの条件が満たされない場合、目的の効果は部分的にしか特定されず、限られた重複による最悪の場合のバイアスの制限とともに、過去の実験に対する代理を偽装できる診断を提供します。さらに、LLM に固有の確率性によりバイアスと分散の両方が発生しますが、サロゲートとして複数の描画の平均を使用すると、両方が緩和されることを示します。シミュレーションにおける方法と理論、および Upworthy の見出しに関する A/B テストへの応用を説明します。私たちの研究から得られる重要な点は、LLM 結果の代理としての妥当性は過去の治療についてのみ改ざんでき、新しい治療については決して検証できないため、新しい介入には人体実験が依然として不可欠であるということです。設計変数としての LLM の選択、プロンプト、温度の役割と、検証のために人体実験のサイズを設定する方法について説明します。

原文 (English)

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.

13:00 JST研究/論文

KANLib -- A Modular, Extensible and Fast Kolmogorov-Arnold Network Implementation

Kolmogorov-Arnold Networks (KANs) have recently emerged as a promising alternative to traditional multilayer perceptrons by replacing linea…

13:00 JST研究/論文

Essential Subspace Merging for Multi-Task Learning

Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained ch…

13:00 JST研究/論文

データセット、年齢、性別を超えた一般化: リソースの少ない子供の ASR のための微調整戦略の包括的な分析

構音障害のある音声を認識することに関連する課題は、主に、調音精度の低下に起因する顕著な音響変動から生じます。過去の研究では、ハイブリッド DNN/HMM シーケンスの識別トレーニングの使用によって認識が向上することが実証されています。このペーパーでは、さまざまな音響モデルに合わせた音響特徴のさまざまな組み合わせの包括的な調査を示し、それぞれに適した特徴の選択を提供します。ピッチ機能を組み込むことで、特に構音障害のある音声を伴う文章認識タスクの認識パフォーマンスが著しく向上しました。 TORGO データベースの体系的な検査を通じて、構音障害音声を認識するための最先端の因数分解時間遅延ニューラル ネットワーク (F-TDNN) モデルのパフォーマンスを強化できる可能性を実証しました。 F-TDNN モデルを使用して実装された私たちの方法は、以前の研究と比較して、孤立単語認識で 4.65% 相対改善、構音障害音声の文認識で 4.63% 相対改善をもたらしました。この改善により、連続するトレーニング サンプル チャンク間で重複するフレームの数を意図的に選択したことに起因する音声の変動が効果的に補償されます。

原文 (English)

Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.

13:00 JST画像/動画生成ロボティクス

HilDA: 自己監視型 LiDAR の事前トレーニングを促進するための拡散を使用した階層的蒸留

カメラから LiDAR への知識の蒸留に Vision Foundation Models (VFM) を活用することは、現実世界の自動運転 (AD) の膨大な幾何学的および運動学的多様性を表現するために必要な注釈付きデータの不足に対する有望な解決策を提供します。ただし、現在のアプローチは通常、VFM をブラックボックス教師として扱い、フレーム単位の特徴の類似性にのみ依存します。その結果、教師のレイヤーごとの意味構造とグローバル コンテキスト、さらには LiDAR シーケンスに固有の豊富な時空間情報が十分に活用されません。私たちは、運転タスクに必要なセマンティックな内容と幾何学的な場所をより適切に捕捉する、LiDAR バックボーン用の自己監視型事前トレーニング フレームワークである HilDA を提案します。 HilDA は、段階的なセマンティクスの調整のための多層蒸留と、シーンレベルのセマンティクスのためのグローバル コンテキストの蒸留を含む階層的蒸留を、時空間的一貫性を促進する時間占有拡散目標と組み合わせます。 HilDA で事前トレーニングされたモデルは、クロスモーダル蒸留ベンチマークで最先端の結果を達成し、3D オブジェクト検出、シーン フロー、セマンティック占有予測に関して事前の蒸留アプローチでトレーニングされたモデルよりも優れたパフォーマンスを発揮します。コードは https://maxiuw.github.io/hilda で入手できます。

原文 (English)

HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training

Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD). However, current approaches typically treat VFMs as black-box teachers, relying exclusively on frame-wise feature similarity. Consequently, they do not fully exploit the teacher's layer-wise semantic structure and global context, as well as the rich spatiotemporal information inherent in LiDAR sequences. We propose HilDA, a self-supervised pretraining framework for LiDAR backbones that better captures the semantic what and geometric where needed for driving tasks. HilDA combines hierarchical distillation comprising multi-layer distillation for progressive semantic alignment and global context distillation for scene-level semantics, with a temporal occupancy diffusion objective promoting spatiotemporal consistency. Models pre-trained with HilDA achieve state-of-the-art results on cross-modal distillation benchmarks and outperform models trained via prior distillation approaches on 3D object detection, scene flow, and semantic occupancy prediction. Code available at: https://maxiuw.github.io/hilda.

13:00 JST研究/論文

Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling

Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…

13:00 JSTLLM/生成AI

Sexualised synthetic personas encode and amplify gendered power asymmetries through voice

This work examines sexualised AI-generated English-speaking voices offered by a popular commercial platform. New technologies may enable se…

13:00 JST研究/論文LlamaNVIDIA

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study

Mixture-of-Experts (MoE) language models are often described as ideal for resource-constrained inference. Each token activates only a small…

13:00 JSTエージェント

Skills for the future software profession: beyond agentic AI!

As coding agents are rapidly changing software engineering, a natural question is: what are the core skills needed by future software engin…

13:00 JST研究/論文

Alternate loss functions and regression models that achieve robustness to outliers by modulating the learning rate

Most real-world datasets used for training supervised learning models are contaminated with noisy data and outliers leading to large predic…

13:00 JST画像/動画生成

MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning

Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. Howe…

13:00 JST画像/動画生成

Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution

Path-based attribution methods such as Integrated Gradients (IG) are widely adopted for their strong axiomatic properties and effectiveness…

13:00 JST研究/論文

On the Position Bias of On-Policy Distillation

On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision fro…

13:00 JSTLLM/生成AIGPT / ChatGPT

AI Fiction in the Wild

Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, to…

13:00 JST画像/動画生成

Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion es…