Skip to the content.

AIニュース 2026-06-24

自動生成: 2026-06-24 13:03 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mysteryOpenAI

    GPT-5 Pro helped solve a 3-year-old immunology mystery, offering insi…

  2. Helping build shared standards for advanced AIOpenAI

    OpenAI helps build shared standards for advanced AI, supporting evalu…

  3. 「Apps in ChatGPT」にメルカリ登場 自社MCP基盤を活用、会話で商品検索などITmedia AI+

    1月に公開したAI接続基盤「Mercari MCP」(Model Context Protocol)を活用。ChatGPTでの会話を通じて…

  4. Anthropic、Slackで「@Claude」を呼べる「Claude Tag」提供──チームの一員として非同期でタスク遂行ITmedia AI+

    Anthropicは、チーム向け新機能「Claude Tag」を発表し、「Slack」での提供を開始した。チャンネル内で「@Claude」…

  5. Mistral、文書解析OCRの新版「OCR 4」公開 文字の位置や信頼度スコアを出力、日本語を含む170言語に対応ITmedia AI+

    Mistral AIは、文書のテキストや構造を抽出するOCRモデルの最新版「Mistral OCR 4」を公開した。日本語を含む170言語…

  6. Geminiで「AIを使いたい現場」と「ダメと言う会社」のギャップを埋める方法ITmedia AI+

    生成AIやAIエージェントの取り組みが活発化している。ただ、その取り組みが検証止まりになるケースも多い。グーグル・クラウド・ジャパンの北瀬…

  7. 主婦がAIアニメでYouTube登録者100万人を突破――1億ユーザーのAI動画生成サービス「PixVerse」の実態ITmedia AI+

    1億人超のユーザーを持つAI動画生成プラットフォーム「PixVerse」とは何か。運営企業の担当者が活用事例などを語った。

トピック別件数

日本語メディア14件

ITmedia AI+ (日本語)

12:00 JST画像/動画生成

主婦がAIアニメでYouTube登録者100万人を突破――1億ユーザーのAI動画生成サービス「PixVerse」の実態

1億人超のユーザーを持つAI動画生成プラットフォーム「PixVerse」とは何か。運営企業の担当者が活用事例などを語った。

11:59 JSTLLM/生成AIGPT / ChatGPT

「Apps in ChatGPT」にメルカリ登場 自社MCP基盤を活用、会話で商品検索など

1月に公開したAI接続基盤「Mercari MCP」(Model Context Protocol)を活用。ChatGPTでの会話を通じて商品を検索したり、出品時の説明文の下書きを作成したりできる。

11:45 JSTロボティクス

「夏場は50度以上のコンテナで作業」に対処 サンワサプライが西日本で荷降ろしロボット活用

サンワサプライが、物流倉庫における荷降ろし作業の自動化と労働環境の改善を目的に、AI搭載のコンテナ向け荷降ろしロボット「RockyOne」を採用した。5月から同社の西日本物流センターで運用を開始している。

11:00 JSTハードウェア/半導体NVIDIA

AmazonはNVIDIAに挑戦状を突きつけるのか

世界最大のハイパースケーラーであるAWSは、AIアクセラレーターを大規模に販売することで、半導体市場の好機を捉えようとしているのだろうか。

10:13 JSTその他Mistral AI

Mistral、文書解析OCRの新版「OCR 4」公開 文字の位置や信頼度スコアを出力、日本語を含む170言語に対応

Mistral AIは、文書のテキストや構造を抽出するOCRモデルの最新版「Mistral OCR 4」を公開した。日本語を含む170言語に対応し、要望の多かったバウンディングボックスや信頼度スコアの出力に対応した。単一コンテナによる自己ホスティングも可能で、厳格なデータ主権や…

08:55 JSTLLM/生成AIAnthropicClaude2媒体が報道

Anthropic、Slackで「@Claude」を呼べる「Claude Tag」提供──チームの一員として非同期でタスク遂行

Anthropicは、チーム向け新機能「Claude Tag」を発表し、「Slack」での提供を開始した。チャンネル内で「@Claude」とメンションすることで、各種ツールやコードベースと連携し、非同期かつ自律的にタスクを遂行する。基盤モデルには「Opus 4.8」が採用され、…

出典:ITmedia AI+TechCrunch AI
07:30 JSTその他

再生師、テクスチャー翻訳家……オカムラがAIで導く2045年“未来の職業”

オカムラは、自社の保有特許とAIを掛け合わせて導き出した「まだ存在しない未来の職業展 2045」を開催。手の動きで環境音を奏でる「ジェスチャーオーケストラ」や、対話から人生を再仕立てする「エンディングエディター」といった体験型展示を披露した。

07:00 JSTハードウェア/半導体NVIDIA

「最初は壊れ過ぎてビビった」──1220億円投じたソフトバンク「AIスパコン」、それでもNVIDIAのGPUを選ぶワケ

AIブームの波に乗って時価総額世界1位に躍り出たNVIDIA。一体なぜ、AIインフラにNVIDIA製GPUが採用されるのか。その理由をAIスパコン開発者に聞いた。

07:00 JSTLLM/生成AIエージェントGoogleGemini

Geminiで「AIを使いたい現場」と「ダメと言う会社」のギャップを埋める方法

生成AIやAIエージェントの取り組みが活発化している。ただ、その取り組みが検証止まりになるケースも多い。グーグル・クラウド・ジャパンの北瀬公彦氏は、そうした「AI活用のカベ」を乗り越えるために「スピードと守りを両立が重要」と語った。

07:00 JSTエージェント

「賞金1000万円」コンテストに「AI番付」 サイバーエージェントが明かす、AIを使い倒させる仕組み作り

AI活用を一部の意欲的な社員にとどめず、組織全体の競争力へとつなげるには何が必要なのか。全社的にAI活用を進めているサイバーエージェント AIオペレーション室長の上野千紘氏と、フリー 常務執行役員の前村菜緒氏が語った。

06:15 JSTロボティクス研究/論文

“中国ヒューマノイド革命”はなぜ起きた、異業種や大手テックが動かす市場の今

中国のヒューマノイドロボット市場は、劇的なパラダイムシフトの渦中にある。出荷台数は前年比約7倍、世界シェアは8割に達し、異業種企業の参入で本体企業数は倍増した。野村総合研究所の李智慧氏による、量産化フェーズへ突入した中国市場の急成長を支えるマクロ動向の解説を紹介する。

17:18 JSTその他

業務でAIを使う人の約38%「禁止されても利用継続」 セキュリティ企業が調査

業務でAIを使っている人の37.8%が勤務先に禁止されても利用を継続する意向を示した――Webセキュリティサービスなどを手掛けるサイバーセキュリティクラウドは、このような調査結果を発表した。

16:34 JSTその他

国産AI「Sakana Fugu」なぜドル建て? 円建てニーズ「受け止める」とSakana AI

円建てプランへのニーズは「日本のユーザーの皆様からいただくご意見・ご要望として、引き続きしっかりと受け止める」という。

13:12 JSTLLM/生成AIAnthropicClaude

NRIセキュア、未公表の脆弱性を「Mythosと同等のレベルで」検出する診断サービス提供

「米AnthropicのClaude Mythos Previewと同等のレベルで未公表の脆弱性を検出できる」のが売り。Anthropicの日本法人代表も新サービスにコメントを寄せている。

海外メディア3件

TechCrunch AI (英語)

08:30 JSTエージェント

India’s MoEngage bets that the future of marketing is millions of AI agents

The all-cash deal gives MoEngage access to technology that assigns AI agents to individual customers.

23:00 JSTその他

4 days left to save up to $190 on TechCrunch Founder Summit 2026

Four days left to save up to $190 on your pass to TechCrunch Founder Summit 2026 — the ultimate founder bootcamp — before Early Bird rates…

22:00 JSTエージェントビジネス/資金調達

Fika Jobs raises $4M to build a video-first hiring platform where AI agents interview candidates

Stockholm-based startup Fika Jobs is building a video-first hiring platform that combines AI interview agents with short-form video profile…

公式ブログ2件

OpenAI (英語)

02:00 JSTLLM/生成AIGPT / ChatGPT

How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery

GPT-5 Pro helped solve a 3-year-old immunology mystery, offering insights into T cell behavior. The breakthrough could support cancer and a…

22:00 JSTLLM/生成AIビジネス/資金調達OpenAI

Helping build shared standards for advanced AI

OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the…

論文299件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント

RIFT-Bench: エージェントティック AI システム向けの動的なレッドチーム化

大規模言語モデル (LLM) を利用したエージェント AI システムは、自律的な意思決定システムへと急速に進化しており、従来の LLM の脆弱性を超えた攻撃ベクトルをさらしています。既存のセキュリティ評価は多くの場合、特定の実装またはドメインに関連付けられており、異種システム間での統一された比較が制限されています。このギャップに対処するために、RIFT-Bench を導入します。これは、多様なエージェント アーキテクチャにわたって統一された評価を可能にする、動的なレッド チーム化のためのグラフ表現主導の方法論です。新しい階層表現に基づいて、RIFT-Bench は 2 つの自動化フェーズで動作します。システム構造を抽出するディスカバリと、適応型敵対的攻撃を展開して包括的な評価レポートを作成するスキャンです。さまざまな攻撃ベクトルや目的にわたって、動的に適応可能な幅広い敵対的プローブを利用して、調査されたシステム自体を評価します。私たちは、さまざまな実装範囲にわたる 45 のエージェント システムにわたって、提案された評価パイプラインの有効性を実証し、このアプローチが異種エージェント アーキテクチャに効果的に一般化できることを示しています。システムや攻撃を超えて、RIFT-Bench は緩和戦略の直接評価もサポートします。これらの主要な機能により、RIFT-Bench はエージェント AI システムのセキュリティ評価のためのスケーラブルな基盤となります。

原文 (English)

RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems. To address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agentic architectures. Building on a novel hierarchical representation, RIFT-Bench operates in two automated phases: Discovery, which extracts system structure, and Scanning, which deploys adaptive adversarial attacks and produces a comprehensive evaluation report. It evaluates the examined system itself, leveraging a broad set of dynamically adaptable adversarial probes across diverse attack vectors and objectives. We demonstrate the effectiveness of the proposed evaluation pipeline across 45 agentic systems spanning a diverse range of implementations, showing that the approach generalizes effectively to heterogeneous agentic architectures. Beyond systems and attacks, RIFT-Bench also supports direct evaluation of mitigation strategies. These key capabilities make RIFT-Bench a scalable foundation for security evaluation of agentic AI systems.

13:00 JSTLLM/生成AI

神経記号的ドライブ: VLA を駆動するためのルールに基づいた忠実な推論

思考連鎖 (CoT) 推論を組み込んだ VLA モデルの駆動は、事前トレーニング済みの VLM 表現を活用し、中間決定を自然言語で公開するため魅力的ですが、現在の理論的根拠には、理論的根拠を計画された動作と因果関係を保つために必要な段階的な決定セマンティクスが欠けていることがよくあります。 Neuro-Symbolic Drive は、古典的なルールベースのプランナーから直接抽出されたルールに基づいた推論トレースを使用して駆動 VLA を監視するニューロシンボリック駆動フレームワークです。私たちの重要な観察は、ルールベースのプランナーは、実行可能な推論エンジンとしてすでに機能している象徴的な AI システムであるということです。これらは、アクティブな安全性の制約について推論し、候補となる操縦を検索し、最終的な軌道を選択します。これらのプランナーをシミュレーションで計測して、ルール評価の各ステップで実行された軌跡と内部決定トレースの両方をキャプチャします。各トレースは、構造化されたルールに基づいた推論にシリアル化され、軌道と組み合わせて、駆動 VLA として Qwen3.5-4B を微調整します。これらのトレースは、アクションを決定するプランナーの状態から直接導出されるため、事後的な調整ではなく、構築によって推論がモーション生成に構造的に結合されることが保証されます。シミュレーターで生成されたベンチマークでは、ルールに基づいた詳細な推論により、3 台のカメラの認識では ADE@3s が 0.47 から 0.26 に、ミス率が 8.30% から 6.40% に減少し、8 台のカメラの認識では 0.54 から 0.26 に、10.13% から 5.99% に減少しました。したがって、Neuro-Symbolic Drive は、神経記号的な計画ロジックを構造化された監視に変換します。コードベース: https://github.com/XiangboGaoBarry/Neural-Symbolic-Drive。

原文 (English)

Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs

Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose intermediate decisions in natural language, yet current rationales often lack the step-by-step decision semantics needed to keep the rationale causally connected to the planned motion. We introduce Neuro-Symbolic Drive, a neuro-symbolic driving framework that supervises a driving VLA with rule-grounded reasoning traces extracted directly from classical rule-based planners. Our key observation is that rule-based planners are symbolic AI systems that already function as executable reasoning engines: they reason about active safety constraints, search over candidate maneuvers, and select a final trajectory. We instrument these planners in simulation to capture both the executed trajectory and the internal decision trace at each rule-evaluation step. Each trace is serialized into structured rule-grounded reasoning and paired with the trajectory to fine-tune Qwen3.5-4B as a driving VLA. Because these traces are derived directly from the planner states that determine the action, they ensure reasoning is structurally coupled to motion generation by construction, rather than by post-hoc alignment. On our simulator-generated benchmark, detailed rule-grounded reasoning reduces ADE@3s from 0.47 to 0.26 and miss rate from 8.30% to 6.40% under three-camera perception, and from 0.54 to 0.26 and 10.13% to 5.99% under eight-camera perception. Neuro-Symbolic Drive thus converts neuro-symbolic planning logic into structured supervision. Code base: https://github.com/XiangboGaoBarry/Neural-Symbolic-Drive.

13:00 JSTLLM/生成AIエージェントロボティクス

エージェントモデルの批判

エージェントとは何ですか?代理店とは何ですか? 「コーディング エージェント」、「AI 共同科学者」、および生産性の向上を約束するその他の「エージェント」ツールとして販売されるラージ言語モデル (LLM) システムの台頭、そして同時に、人間に対する投機的な「マシン エージェント」の下で AI が破壊的な力で人間の制御から逃れるなどの「実存的」な懸念により、有能なシステムを構築するためと、恐れるべきかどうか、何を恐れるべきかを理解するために、自動化がどこで終わり、エージェントが始まるのかを明確にすることが不可欠になっています。デカルトの独立した思考における主体性の根拠と、SF における自律的存在の描写に基づいて、AI エージェントの現状を概観し、目標、アイデンティティ、意思決定、自己規制、学習という 5 つの側面に沿ってエージェントのアーキテクチャを分析します。具体的には、真の主体性は、これらの構造が外部の足場を介して組み立てられるのではなく、\emph{エージェント} システムの能力が設計されたものに存在することを必要とすると主張します。ワークフローと \emph{エージェント} システムは、その機能 (ソーシャル インタラクションを含む) が内生的に生じ、所定のタスク用に設計されたシステムと、オープンワールドで真の自律性を持って動作できるシステムとの間の境界を定義します。この分析に基づいて、階層的な目標分解、アイデンティティ進化、個別に学習された世界モデルに基づくシミュレーション推論を組み合わせた、汎用エージェント モデルの Goal-Identity-Configurator (GIC) アーキテクチャを提案します。さらに、私たちは、より優れた自律性と「エージェンシー」を持ちながらも人間の監視下にあるエージェント システムの監査可能性、制御可能性、安全性についての洞察を共有します。

原文 (English)

Critique of Agent Model

What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and other ``agentic" tools that promise to drive up productivity, and at the same time, ``existential" concerns such as AI escaping human control with destructive power under a speculative ``machine agency" against humans, it has become essential to clarify where automation ends and agency begins, both for building capable systems and for understanding whether and what to fear. Drawing on Descartes' grounding of agency in independent thought, and on portrayals of autonomous beings in science fiction, we survey the current landscape of AI agents, and analyze agent architectures along five dimensions: goal, identity, decision-making, self-regulation, and learning. Specifically, we argue that genuine agency requires these structures to be \emph{internalized within the system itself} rather than assembled through external scaffolding. This distinction between \emph{agentic} systems, whose competence resides in engineered workflows, and \emph{agentive} systems, whose capabilities (including social interaction) arise endogenously, defines the boundary between systems designed for prescribed tasks, and those capable of operating in the open world with true autonomy. Building on this analysis, we propose the Goal-Identity-Configurator (GIC) architecture for a general-purpose agent model, combining hierarchical goal decomposition, identity evolution, simulative reasoning grounded in a separately trained world model, learned self-regulation, and self-directed learning from both real and simulated experience. Furthermore, we share insight on the auditability, controllability, and safety of agentive systems that possess greater autonomy and ``agency", but remain under human oversight.

13:00 JSTエージェント

制約マニホールド制御による安全で一般化可能な階層型マルチエージェント RL

マルチエージェント システムは、厳しい安全制約の下で調整された動作を必要とするセーフティ クリティカルなアプリケーションで広く使用されています。既存のアプローチは根本的なトレードオフに直面しています。学習ベースの手法は強力な経験的パフォーマンスを達成しますが、理論的な安全性が保証されていません。一方、制御理論的な手法は安全性を強化しますが、過度に保守的で非効率な動作につながることがよくあります。我々は、高レベルのポリシー学習を通じて効果的な調整を可能にしながら、制約マニホールドを介して低レベルのマイルドな仮定の下で厳しい安全制約を強制する、階層型マルチエージェント強化学習フレームワークを提案します。私たちのアプローチは、マルチエージェント設定で理論的な安全性を保証し、定常的な学習ダイナミクスを生み出すため、安定した効率的なトレーニングを可能にします。経験的に、私たちの方法はほぼ完璧な安全率を維持しながら競争力のあるパフォーマンスを達成し、さまざまな数のエージェントと障害物に効果的に一般化します。

原文 (English)

Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approaches face a fundamental trade-off: learning-based methods achieve strong empirical performance but lack theoretical safety guarantees, while control-theoretic methods enforce safety but often lead to overly conservative and inefficient behaviors. We propose a hierarchical multi-agent reinforcement learning framework that enforces hard safety constraints under mild assumptions at low level via a constraint manifold, while enabling effective coordination through high-level policy learning. Our approach provides theoretical safety guarantees in the multi-agent setting and yields stationary learning dynamics, thereby enabling stable and efficient training. Empirically, our method achieves competitive performance while maintaining nearly perfect safety rates, and generalizes effectively to varying numbers of agents and obstacles.

13:00 JSTLLM/生成AI

広範囲かつ永続的に有益なモデルを目指した強化学習

AI システムがますます多様化し、一か八かの環境に展開されるにつれて、モデルの調整はトレーニング中に見られるタスクや領域を超えて一般化する必要があります。これは、報酬のハッキング、欺瞞、その他の意図しない戦略によって予期せぬ不整合が生じる可能性がある強化学習 (RL) にとって特に重要です。私たちは、現実的なドメインでインスタンス化された有益な動作に関する RL が、トレーニング分布を超えた広範かつ永続的なアラインメント一般化を生成できるかどうかを研究します。私たちは、健康、科学、教育などのさまざまな領域にまたがる、真実性、公平性、リスク認識、正しさなどの有益な特性を測定および訓練するように設計された現実的な状況のデータセットを構築します。次に、このデータセットで RL を使用してモデルをトレーニングし、整合性と有益な動作に関する 50 を超える独立したベンチマークでモデルを評価します。コンピューティング一致ベースラインと比較して、有益な特性 RL は、これらの配布外ベンチマークの 80% 以上でパフォーマンスを向上させます。私たちは、実質的な分布外のアラインメントの転移を観察しています。つまり、健康という 1 つの領域に完全に限定された有益な行動の RL 介入は、報酬のハッキング、欺瞞、一般的なミスアラインメントの削減など、健康以外のアラインメント評価に広範な改善をもたらします。最後に、アライメントの持続性、つまりモデルを不整合に向けて誘導しようとする試みの下で動作がロバストにアライメントされたままであるかどうかを研究します。有益な特性 RL でトレーニングされたモデルは、敵対的なプロンプトや有害な微調整に対する優れた耐性など、持続性の向上を示します。これらの影響の原因を特定するには、さらなる研究が必要です。これらの結果は、現実的な領域で有益な行動を強化するための RL が、人間の繁栄とより強固に一致するモデルを生成できることを示唆しています。

原文 (English)

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through reward hacking, deception, or other unintended strategies. We study whether RL on beneficial behavior, instantiated in realistic domains, can produce broad and persistent alignment generalization beyond the training distribution. We construct a dataset of realistic situations designed to measure and train beneficial traits, such as truthfulness, fairness, risk awareness, and corrigibility, spanning varied domains, including health, science, and education. We then train models with RL on this dataset and evaluate them on more than 50 independent benchmarks of alignment and beneficial behavior. Compared to a compute-matched baseline, beneficial trait RL improves performance on over 80% of these out-of-distribution benchmarks. We observe substantial out-of-distribution alignment transfer: a beneficial-behavior RL intervention entirely limited to one domain, health, produces broad improvements on non-health alignment evaluations, including reduced reward hacking, deception, and general misalignment. Finally, we study alignment persistence: whether behavior remains robustly aligned under attempts to steer models towards misalignment. Models trained with beneficial trait RL show improved persistence, including greater resistance to adversarial prompting and harmful finetuning; further work is required to isolate the sources of these effects. These results suggest that RL to reinforce beneficial behavior in realistic domains can produce models that are more robustly aligned with human flourishing.

13:00 JSTエージェントLlama

言語モデルエージェントは機械的解釈において回路の説明者として役立つでしょうか?

機構の解釈可能性は、回路の自動ローカライズにおいて大幅な進歩を遂げましたが、ローカライズされたコンポーネントが何を行うかを説明することは依然として労力を要し、標準化が困難です。この研究では、回路がすでに特定されている場合、言語モデル (LM) エージェントがこの説明問題を支援できるかどうかを研究します。ここでは、163 のコンポーネント レベルの注釈を備えた 84 の半合成トランス回路から構築された回路説明用のベンチマークである AgenticInterpBench を紹介します。我々は、観察、仮説生成、因果関係検証の反復ループを通じて各コンポーネントを分析し、最終的にコンポーネントレベルの説明と回路レベルのタスク記述を生成するエージェント的説明ツールである HyVE (仮説、検証、説明) を提案します。 HyVE は 4 つの LM バックボーンにわたって、有用なコンポーネント レベルおよびタスク レベルの説明を復元しますが、一律に最適なバックボーンはありません。私たちの分析によると、強力なバックボーンは通常、観察に基づいた仮説を形成しますが、失敗は検証ループの後半で、不完全な検証計画、コード実行エラー、または未解決の仮説によって発生することが多いことがわかりました。 Llama-3-8B の算術回路に関するケース スタディでは、同じ定式化が半合成ベンチマークを超えて自然にトレーニングされたモデルに拡張できることが示されています。全体として、LM エージェントは回路の説明に有望ですが、信頼性の高い検証が依然として主要な障害となっています。

原文 (English)

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) agents can assist with this explanation problem once a circuit has already been identified. We introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circuits with 163 component-level annotations. We propose HyVE (Hypothesize, Validate, Explain), an agentic explainer that analyzes each component through an iterative loop of observation, hypothesis generation, and causal validation, eventually producing a component-level explanation and a circuit-level task description. Across four LM backbones, HyVE recovers useful component- and task-level explanations, but no backbone is uniformly best. Our analysis shows that strong backbones usually form observation-grounded hypotheses, while failures more often arise later in the validation loop, through incomplete validation plans, code execution errors, or unresolved hypotheses. A case study on an arithmetic circuit in Llama-3-8B shows that the same formulation can extend beyond semi-synthetic benchmarks to naturally trained models. Overall, LM agents are promising circuit explainers, but reliable validation remains the key obstacle.

13:00 JST研究/論文

フィルターバブルの打破: 多目的レコメンデーションのためのセマンティック パレート DQN フレームワーク

レコメンダー システムは、即時のユーザー エンゲージメントをモノリシックに最適化することでフィルター バブルとセマンティックな均質化を引き起こすことがよくあります。従来の Deep Q-Networks を含む標準的な単一目的モデルは、プラットフォームの維持と、情報の多様性やプロバイダーの公平性などの重要な社会的価値との間のトレードオフをうまく乗り切ることができません。これらの制限に対処するために、推奨を意味論的な多目的マルコフ決定プロセスとして形式化する多目的強化学習フレームワークを導入します。高忠実度のセマンティック埋め込みをパレート DQN エージェントと統合することにより、私たちのアーキテクチャはエンゲージメント、多様性、公平性を個別の集約不可能な報酬信号として扱い、静的な報酬のスカラー化の落とし穴を回避します。 MovieLens の小規模データセットに対する実験的評価では、ハイパーボリューム ベースのアクション選択により、セマンティック崩壊の原因となるフィードバック ループが破壊されることが示されています。高い状態軌道分散を維持することにより、パレート DQN はパレート フロンティアを効果的にマッピングし、エンゲージメントにわずかな影響を与えるだけで補助的な社会目標の向上を達成します。この取り組みは、本質的に調整された責任ある推奨システムへの道を提供します。

原文 (English)

Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation

Recommender systems often induce filter bubbles and semantic homogenization by monolithically optimizing for immediate user engagement. Standard single-objective models, including traditional Deep Q-Networks, are ill-equipped to navigate the trade-offs between platform retention and critical societal values like information diversity and provider fairness. To address these limitations, we introduce a multi-objective reinforcement learning framework that formalizes recommendation as a semantic multi-objective Markov decision process. By integrating high-fidelity semantic embeddings with a Pareto-DQN agent, our architecture treats engagement, diversity, and fairness as distinct, non-aggregable reward signals, avoiding the pitfalls of static reward scalarization. Empirical evaluations on the MovieLens small dataset shows that our hypervolume based action selection disrupts the feedback loops responsible for semantic collapse. By sustaining high state-trajectory variance, the Pareto-DQN effectively maps the Pareto frontier, achieving gains in auxiliary societal objectives with only marginal impacts on engagement. This work provides a path toward intrinsically aligned, responsible recommender systems.

13:00 JST研究/論文

女性セックスワーカーの説明可能なメンタルヘルスリスク予測のためのアンサンブル特徴選択とハリスホークス最適化

女性セックスワーカー (FSW) に影響を与える重大なメンタルヘルス問題の 1 つは、精神障害、特にうつ病です。暴力、偏見、経済的困難にさらされると、心理的リスクがさらに高まります。現在の機械学習 (ML) モデルは通常、この疎外されたグループに存在する高次元で複雑なリスク パターンを捉えるのには効果的ではありません。この論文では、ANOVA と相互情報量を使用したアンサンブル特徴選択戦略と、ハリス ホークス最適化調整ロジスティック回帰を組み合わせたハイブリッド予測モデルを提案し、脆弱なグループの精神的健康を予測するための群知能の新しいアプリケーションを示します。 Explainable AI (XAI) メソッドは、モデル予測に関連するトラウマの要因を理解するために使用できます。 3,005 人の FSW のグループに適用した場合、提案されたモデルは従来の分類器よりも効果的であり、精度 95.78%、F1 スコア 95.77%、AUC 0.96 であり、心的外傷後ストレス、クライアント関連の暴力、および職業的要因をうつ病の主な要因として特定していることがわかります。この取り組みは、従来のアプローチと ML アプローチの間のギャップを埋めて、脆弱なグループが早期支援、証拠に基づいた対象を絞った心理社会的ケア、健康計画を受けられるようにする XAI ツールを開発します。

原文 (English)

Ensemble Feature Selection and Harris Hawks Optimization for Explainable Mental Health Risk Prediction in Female Sex Workers

One of the significant mental health issues affecting female sex workers (FSWs) is mental disorders, especially depression. Exposure to violence, stigma, and economic hardship further increases their psychological risk. Current machine learning (ML) models are typically ineffective at capturing the high-dimensional and complex risk patterns that exist in this marginalized group. This paper suggests a hybrid predictive model that merges an ensemble feature selection strategy using ANOVA and mutual information and Harris Hawks optimization-tuned logistic regression and represents a new application of swarm intelligence to predict mental health in vulnerable groups. The explainable AI (XAI) methods can be used to understand the factors of trauma associated with model predictions. When applied to a group of 3,005 FSWs, it can be seen that the proposed model is more effective than traditional classifiers, with an accuracy of 95.78%, an F1 score of 95.77%, and an AUC of 0.96, and identifying post-traumatic stress, client-related violence, and occupational factors as major contributors to depression. This work bridges the gaps between conventional and ML approaches to develop an XAI tool that enables vulnerable groups to receive early assistance, evidence-based targeted psychosocial care, and health planning.

13:00 JSTLLM/生成AI

軌道の模倣を超えて: LLM 推論のための戦略に基づくポリシーの最適化

強力な言語モデルから弱い言語モデルへの推論機能を抽出するには、通常、特定の解決策の軌跡を模倣し、どのように推論するかではなく何を答えるかを効果的に転送することが含まれます。この軌跡レベルの模倣は、応用可能な問題解決スキルの習得ではなく、インスタンス固有のステップの暗記を促進し、新しい問題への一般化を制限します。私たちは、インスタンスレベルの軌跡の模倣を再利用可能な戦略の蒸留に置き換える、戦略に基づくポリシーの最適化 (SGPO) を提案します。 SGPO は、強力なモデルの応答から構造化された戦略の説明を抽出し、問題ごとに自律的な軌道と戦略に基づく軌道の両方を構築して、戦略的ガイダンスの有無にかかわらずモデルの動作を直接比較できるようにします。次に、このフレームワークは 2 つの重要な質問に対処します。抽出方法については、トークンレベルのフォワード KL 目標により、安定性を確保する近位制約を使用して、戦略条件付けによって引き起こされる分布シフトをガイドなしポリシーに選択的に転送します。いつ抽出するかについては、適応型インスタンス レベルの重み付けにより、自律探索が不十分な場合のガイダンスが強化され、モデル自体の能力が向上するにつれてガイダンスが軽減されます。 2 つのモデル ファミリにわたる 4 つの数学的ベンチマークの実験では、SGPO が SFT、オンポリシー RL、およびハイブリッド ポリシーのベースラインを常に上回っており、Qwen2.5-7B-Instruct の最も強力なベースラインよりも平均スコアを 2.2 ポイント改善していることが示されています。分析の結果、フォワード KL 対物レンズは、直接軌道の模倣を上回る本質的に選択的な蒸留信号を提供し、戦略蒸留が基本モデルの機能と相補的なスケーリングを示すことが明らかになりました。

原文 (English)

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason. This trajectory-level imitation encourages memorization of instance-specific steps rather than acquisition of transferable problem-solving skills, limiting generalization to novel problems. We propose Strategy-Guided Policy Optimization (SGPO), which replaces instance-level trajectory imitation with reusable strategy distillation. SGPO extracts structured strategy descriptions from strong-model responses and, for each problem, constructs both autonomous and strategy-guided trajectories to enable direct comparison of the model's behavior with and without strategic guidance. The framework then addresses two key questions. For how to distill, a token-level forward-KL objective selectively transfers the distributional shift induced by strategy conditioning into the unguided policy, with proximal constraints ensuring stability. For when to distill, adaptive instance-level weighting strengthens guidance when autonomous exploration falls short and reduces it as the model's own competence grows. Experiments on four mathematical benchmarks across two model families show that SGPO consistently outperforms SFT, on-policy RL, and hybrid-policy baselines, improving the average score by 2.2 points over the strongest baseline on Qwen2.5-7B-Instruct. Analysis reveals that the forward-KL objective provides an inherently selective distillation signal that outperforms direct trajectory imitation, and that strategy distillation exhibits complementary scaling with base model capability.

13:00 JSTLLM/生成AI研究/論文

学術論文全文に基づく共起ネットワークによるアルゴリズムの学術的影響の探求

アルゴリズムは、人工知能 (AI) の時代の科学研究の中心となっています。論文でのアルゴリズムの言及は、人気や影響力を示すためによく使用されますが、既存の研究では通常、個々のアルゴリズムを個別に評価し、相互接続を通じて形成される集合的な影響には限定的に注意を払っています。本研究では、学術論文全文に基づいて自然言語処理(NLP)における大規模なアルゴリズム共起ネットワークを構築し、ネットワークの観点からアルゴリズムの影響を調査します。深層学習モデルを使用して、アルゴリズム エンティティを抽出し、全体的、累積的、年次共起ネットワークを構築します。私たちはそれらの構造的特徴を分析し、複数の中心性測定を適用して、分野全体および長期にわたるアルゴリズムのグループへの影響を評価します。結果は、アルゴリズム ネットワークが複雑なネットワークの典型的な特徴を示しており、約 20 年間にわたって接続の密度が増加していることを示しています。古典的で高性能なアルゴリズムや、さまざまな研究期間の交差点に位置するアルゴリズムは、高い人気、制御性、中心性、バランスのとれた影響力を持つ傾向があります。アルゴリズムの影響力が低下すると、通常、まずそのアルゴリズムがコア ネットワークでの地位を失い、続いて他のアルゴリズムとの関連性が弱まります。この研究は、アルゴリズム共起ネットワークの最初の大規模な分析です。 40 年以上の学術出版物を網羅しており、アルゴリズムの影響を時間的かつ構造的に把握し、アルゴリズム、学者、タスクを結び付けるネットワークに関する将来の研究の基盤を提供します。

原文 (English)

Exploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Papers

Algorithms have become central to scientific research in the era of artificial intelligence (AI). Although algorithm mentions in papers are often used to indicate popularity and influence, existing studies usually evaluate individual algorithms in isolation and pay limited attention to the collective influence formed through their interconnections. This study constructs large-scale algorithm co-occurrence networks in natural language processing (NLP) based on the full text of academic papers and investigates algorithm influence from a network perspective. Using deep learning models, we extract algorithm entities and build overall, cumulative, and annual co-occurrence networks. We analyze their structural characteristics and apply multiple centrality measures to assess the group influence of algorithms across the whole field and over time. The results show that algorithm networks display typical features of complex networks, with increasingly dense connections developing over approximately two decades. Classic, high-performing algorithms and those located at the intersections of different research periods tend to have high popularity, control, centrality, and balanced influence. When the influence of an algorithm declines, it usually loses its core network position first, followed by weaker associations with other algorithms. This study is the first large-scale analysis of algorithm co-occurrence networks. Covering more than four decades of academic publications, it provides a temporal and structural view of algorithm influence and offers a foundation for future research on networks linking algorithms, scholars, and tasks.

13:00 JSTエージェントGPT / ChatGPT

ReMMD: マルチモーダルな誤情報検出のための現実的な多言語マルチ画像エージェント検証

バイラル投稿には、長い多言語の説明、複数の画像、混合の出所、および微妙なテキストと画像の構成エラーが組み合わされているため、マルチモーダルな誤情報の検出はますます重要になっています。既存のベンチマークと手法は依然としてこの設定にあまり適合していません。通常、それらは短いキャプション、単一の画像、バイナリ ラベル、または 1 つの操作ソースを分離しますが、現実的な証拠検索ではエージェントによる検証は依然としてコストがかかります。我々は、マルチモーダルな誤情報検出のための現実的な多言語マルチ画像エージェント検証フレームワークである ReMMD を紹介します。 ReMMD には、500 のサンプル、2,756 の画像、5 つの単一言語、2 つの言語間設定、3 つのテキスト長階層、複数画像の投稿、5 方向の真実性ラベル、8 つの歪曲ラベル、証拠の出所、根拠を備えた現実世界のマルチモーダル誤情報検出ベンチマークである ReMMDBench が含まれています。また、投稿をアトミックポイントに分解し、再利用可能な証拠セットを構築し、構造化された L1/L2/L3 出力を予測する永続メモリ検証ツールである ReMMD-Agent も含まれています。独自のシステム、オープン LVLM、MMD エージェント、および T2 エージェント全体にわたって、ReMMD エージェントは、GPT-5.2 を使用して 41.80% の精度と 39.12% のマクロ F1 という最高の 5 方向正確性パフォーマンスを実現しながら、MMD エージェントと比較して 17.5%、T2 エージェントと比較して 79.9% コストを削減します。プロジェクトは https://dang-ai.github.io/ReMMD で入手できます。

原文 (English)

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle text--image framing errors. Existing benchmarks and methods remain poorly matched to this setting: they usually isolate short captions, single images, binary labels, or one manipulation source, while agentic verification remains costly under realistic evidence search. We present ReMMD, a realistic multilingual multi-image agentic verification framework for multimodal misinformation detection. ReMMD includes ReMMDBench, a real-world multimodal misinformation detection benchmark with 500 samples, 2,756 images, five monolingual languages, two cross-lingual settings, three text-length tiers, multi-image posts, five-way veracity labels, eight distortion labels, evidence provenance, and rationales. It also includes ReMMD-Agent, a persistent-memory verifier that decomposes posts into atomic points, builds a reusable evidence set, and predicts structured L1/L2/L3 outputs. Across proprietary systems, open LVLMs, MMD-Agent, and T2-Agent, ReMMD-Agent obtains the best five-way veracity performance, with 41.80% accuracy and 39.12% macro-F1 using GPT-5.2, while reducing cost by 17.5% relative to MMD-Agent and 79.9% relative to T2-Agent. The project is available at https://dang-ai.github.io/ReMMD.

13:00 JSTLLM/生成AI

VeryTrace: コンパイル可能な形式主義と構造化検証による推論トレースの検証

思考連鎖 (CoT) プロンプトを使用した複数ステップの推論は依然として脆弱です。初期段階での論理的エラーや幻覚が静かに伝播し、自信はあるが不正確な結論を導き出します。この文書では、自然言語推論トレースを構造化されたコンパイル可能な表現に形式化するゼロショット検証および修復フレームワークである VeryTrace について説明します。 VeryTrace は、(i) ステップの依存関係を明示し、(ii) 定量的なコンテンツを実行可能な式として機械化し、(iii) 演繹スキーマを介して意味推論を構造化するドメイン固有言語 (DSL) を導入します。当社のハイブリッド検証ツールは、計算の正しさ、依存関係の解決、制約を満たすための決定論的チェックと、機械化不可能な意味論的判断のための対象を絞った LLM 監査を組み合わせて、ステップレベルのエラーの位置特定と修復を可能にします。 VeryTrace は、競争数学 (AIME 2025)、ロボティクス プランニング (LLM-BabyBench)、親族推論 (CLUTRR) の 3 つの多様なドメインにわたって、ドメイン固有のトレーニングやコンテキスト内のサンプルを必要とせずに、最先端の LLM でのゼロショット ベースラインを超える精度を向上させ、形式化されたトレース検証が精度と一般化の両方を達成していることを示しています。

原文 (English)

VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification

Multi-step reasoning with Chain-of-Thought (CoT) prompting remains fragile: logical errors or hallucinations in early steps silently propagate, producing confident but incorrect conclusions. This paper presents VeryTrace, a zero-shot verification-and-repair framework that formalizes natural-language reasoning traces into a structured, compilable representation. VeryTrace introduces a Domain-Specific Language (DSL) that (i) makes step dependencies explicit, (ii) mechanizes quantitative content as executable expressions, and (iii) structures semantic inferences via deduction schemas. Our hybrid verifier combines deterministic checks for computational correctness, dependency resolution, and constraint satisfaction with targeted LLM audits for non-mechanizable semantic judgments, enabling step-level error localization and repair. Across three diverse domains-competition mathematics (AIME 2025), robotics planning (LLM-BabyBench), and kinship reasoning (CLUTRR), VeryTrace improves accuracy over zero-shot baselines on state-of-the-art LLMs without requiring domain-specific training or in-context examples, demonstrating that formalized trace verification achieves both precision and generalization.

13:00 JSTエージェント

OmniPath: 車椅子のアクセシビリティを監査するためのマルチモーダル エージェント フレームワーク

車椅子ユーザーにとって、地図上の標準的な青い線は、多くの場合、約束を破られたものです。 OpenStreetMap (OSM) のようなプラットフォームは、パスの位置をうまく把握できますが、その上を移動する物理的な感覚を伝えることができないことがよくあります。この情報の壁は車椅子ユーザーにとっては問題です。この問題を解決するために、受動的なマッピングからプロアクティブな環境監査に移行するシステムである OmniPath を紹介します。私たちのフレームワークは、OSM のネットワーク トポロジと高密度航空 LiDAR (USGS 3DEP) のサブメートル精度を融合して、歩行者環境の忠実度の高い 3D モデルを作成します。当社のエージェントは単にユーザーをルーティングするのではなく、仮想的にネットワークを横断し、0.5 メートル単位で表面を分析します。これは、ADA コンプライアンス基準に照らして、特に斜面、横断斜面、および垂直不連続部を走る物理的摩擦ポイントを厳密に定量化し、重み付けされた重大度スコアを計算して危険を「軽度」から「重大」まで分類します。現実世界の信頼性を確保するために、層別ランダムサンプリングを使用して、ナショナル モール全体にわたる 200 件の物理的なグラウンド トゥルース フィールド調査に対してシステムを検証しました。このフレームワークは、重篤度の高いハザードに対する診断上の強力な信頼性を実証し、重度カテゴリーで 0.60、重篤カテゴリーで 0.58 の F1 スコアを達成しました。このマイクロスケールの検査を自動化することで、OmniPath は標準地図が見逃す「目に見えない」障壁を特定し、静的データセットを、ユーザーが家を出る前にアクセシビリティの課題を予測するアクセシビリティ データ ソースに効果的に変換します。

原文 (English)

OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility

For a wheelchair user, a standard blue line on a map is often a broken promise. While platforms like OpenStreetMap (OSM) successfully capture where a path is, they frequently fail to convey how it physically feels to travel on it. This information barrier is problematic for wheelchair users. To solve this issue, we present OmniPath, a system that moves from passive mapping to proactive environmental auditing. Our framework fuses the network topology of OSM with the submeter precision of high-density aerial LiDAR (USGS 3DEP) to create a high-fidelity 3D model of the pedestrian environment. Rather than simply routing a user, our agent virtually traverses the network, analyzing the surface in 0.5 meter increments. It rigorously quantifies physical friction points specifically running slope, cross slope, and vertical discontinuities against ADA compliance standards, calculating a weighted severity score to categorize hazards from ``Mild'' to ``Critical.'' To ensure real world reliability, we validated the system against 200 physical ground truth field surveys across the National Mall using stratified random sampling. The framework demonstrated strong diagnostic reliability for high-severity hazards, achieving F1-scores of 0.60 for Severe and 0.58 for critical categories. By automating this micro-scale inspection, OmniPath identifies the ``invisible'' barriers that standard maps miss, effectively transforming a static dataset into accessibility data source that anticipates accessibility challenges before the user ever leaves home.

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達GPT / ChatGPT

T2D ベンチ: 多層臨床ライフスタイル ナレッジ グラフを使用した 2 型糖尿病の LLM 出力の証拠ゲート型評価

大規模言語モデル (LLM) は、2 型糖尿病に対する臨床的に流暢な推奨事項を生成できますが、ガイドラインの制約を満たしたり、ライフスタイルに関連した血糖の主張を明確に正当化したりすることはできません。我々は、LLM 出力が明示的でグラフチェック可能な証拠要件を満たしているかどうかをテストするための、再現可能なベンチマークおよび証拠ゲート型評価フレームワークである T2D-Bench を紹介します。 T2D-Bench は、生体医学スパイン (UMLS、DrugBank、SIDER)、計算可能な ADA 治療標準ルール、血糖検査室効果への機構的なブリッジを介して接続されたライフスタイル知識を組み合わせた、多層の臨床ライフスタイル ナレッジ グラフに基づいて構築されています。診断、投薬の安全性、敵対的なライフスタイルの衝突にわたる 100 の構造化されたビネット全体で、ベースライン出力は、GPT-4o-mini のケースの 35%、GPT-4o のケースの 33% で、ベンチマークで定義されたエビデンスパス チェックに失敗しました。証拠ゲートはサポートされていない省略を検出し、制約付きリビジョンを使用して、出力をベンチマークで定義された証拠要件に検証者レベルで準拠させます。これらの結果は、糖尿病に焦点を当てた LLM 出力において、計算可能な証拠の制約により、裏付けのない臨床上の省略が明示的、測定可能、修正可能になる可能性があることを示しています。

原文 (English)

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.

13:00 JST研究/論文

拡散と流れのマッチングの背後にある幾何学: ワッサーシュタイン空間における勾配の流れと測地線

有限二次モーメントを持つ確率測度の空間 $\mathcal{P}_2(\mathbb{R}^d$) は自然幾何学を持ちます。二次のワッサーシュタイン距離 W_2 により完全計量空間になり、オットーに従って、測地線が最適輸送補間である (形式的な) リーマン多様体になります。この多様体上では、自由エネルギーの勾配流 F(rho) = KL(rho || \pi) はまさにフォッカー・プランク方程式であり、その陰的オイラー離散化は JKO スキームです。これは拡散モデルの基礎となるジオメトリです。順方向プロセスは自由エネルギーを降下させ、各ノイズ除去ステップは 1 つの JKO ステップを実現し、DDPM、DDIM、NCSN/SMLD、およびエネルギー マッチングを回復します。これは 1 つのスキームであり、別々の理論ではありません。同じ多様体は、第 2 変分原理をサポートします。その測地線 (Benamu-Brenier 式の最小動作曲線) は、まさにフロー マッチングが学習する最適な転送パスです。両方の終点を固定し、測地線に従うと、生成は直線に沿った決定論的な ODE になり、サンプリング ステップが大幅に少なくなります。両方のモデル族を 1 つの多様体に配置すると、それらの関係が正確になります。拡散は自由エネルギー勾配の流れ、つまり初期値問題に従います。最適輸送フロー マッチングは、ワッサーシュタイン測地線、境界値問題に従います。この 2 つは、異なるパスを通って同じエンドポイントに到達します。

原文 (English)

The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserstein Space

The space $\mathcal{P}_2(\mathbb{R}^d$) of probability measures with finite second moment carries a natural geometry: the quadratic Wasserstein distance W_2 makes it a complete metric space and, following Otto, a (formal) Riemannian manifold whose geodesics are the optimal-transport interpolations. On this manifold, the gradient flow of the free energy F(rho) = KL(rho || \pi) is exactly the Fokker-Planck equation, and its implicit-Euler discretization is the JKO scheme. This is the geometry underlying diffusion models: the forward process descends the free energy, and each denoising step realizes one JKO step, which recovers DDPM, DDIM, NCSN/SMLD, and Energy Matching; this is one scheme, not separate theories. The same manifold supports a second variational principle. Its geodesics - the minimum-action curves of the Benamou-Brenier formula - are precisely the optimal-transport paths that Flow Matching learns. Fixing both endpoints and following the geodesic, generation becomes a deterministic ODE along a straight line, hence far fewer sampling steps. Placing both families of models on one manifold makes their relationship exact: diffusion follows a free-energy gradient flow, an initial-value problem; optimal-transport Flow Matching follows a Wasserstein geodesic, a boundary-value problem. The two reach the same endpoints along different paths.

13:00 JST研究/論文

因果強化学習の概要

因果推論は、環境に関するデータと知識を組み合わせて、反事実的な性質の質問、つまり、この実現されていない現実のデータが現在利用できない場合でも、現実が異なっていたら何が起こっていたかを推論することを可能にする一連の原則とツールを提供します。強化学習は、エージェントが環境に配置され、探索的で試行錯誤的なアプローチを追求するときに、特定の尺度 (報酬、後悔など) を最適化するポリシーを学習する方法を提供します。これら 2 つの分野は独立して発展し、実質的に相互作用することはありません。私たちは、それらが同じ構成要素のさまざまな側面、反事実関係を介して機能し、それがそれらを臍帯で結びつけていることに注目します。これらの観察に基づいて、この関係が明確に認識され数学化されると、新たな学習の機会が生まれます。この可能性を実現するために、RL エージェントがデプロイされている環境は、さまざまな因果不変性を持つ自律メカニズムの集合として分解でき、構造的因果モデルとして倹約的にモデル化できることに注意します。標準の RL 設定は、そのようなモデルを暗黙的にエンコードします。この形式化により、文献では無関係に見える、オンライン、ポリシー外、因果微積分学習など、さまざまな学習モードを統一的に扱うことができます。ただし、これらの手法は網羅的なものではありません。新しい分析次元を必要とする、自然で普及した学習設定のクラスをいくつか紹介します。具体的には、因果レンズを通して、一般化された政策学習、どこに介入するか、模倣学習、反事実学習を紹介し、議論します。これらのタスクは、反事実学習のより広い視野につながり、因果推論と強化学習を並行して研究する大きな可能性を示唆しています。これを私たちは因果強化学習 (CRL) と呼んでいます。

原文 (English)

An Introduction to Causal Reinforcement Learning

Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is currently available. Reinforcement learning provides methods to learn a policy that optimizes a specific measure (e.g., reward, regret) when the agent is deployed in an environment and pursues an exploratory, trial-and-error approach. These two disciplines have evolved independently and with virtually no interaction between them. We note that they operate over different aspects of the same building block, counterfactual relations, which makes them umbilically connected. Based on these observations, novel learning opportunities arise when this connection is explicitly acknowledged and mathematized. To realize this potential, we note that any environment where the RL agent is deployed can be decomposed as a collection of autonomous mechanisms with different causal invariances, parsimoniously modeled as a structural causal model; any standard RL setting implicitly encodes such a model. This formalization allows us to put under a unifying treatment different modes of learning, including online, off-policy, and causal calculus learning, which appear unrelated in the literature. However, these modalities are not exhaustive: we introduce several natural and pervasive classes of learning settings that entail novel dimensions of analysis. Specifically, we introduce and discuss through causal lenses generalized policy learning, where to intervene, imitation learning, and counterfactual learning. These tasks lead to a broader view of counterfactual learning and suggest great potential for studying causal inference and reinforcement learning side by side, which we call causal reinforcement learning (CRL).

13:00 JST研究/論文

ストリーミング ASR における言語を超えたエンコーダ転送を形成するのは、遅延ではなくデータ スケールです

ストリーミング音声認識モデルを新しい言語に適応させるには、多言語 (ML) エンコーダーまたは英語専用 (EN) エンコーダーという 2 つの適切なウォーム スタートのどちらかを選択する必要があります。一般的な直観としては、多言語エンコーダは低データ時に最も役立つはずであるということですが、その利点がどのくらい持続するか、ストリーミング遅延が短いことでその利点が増幅されるかどうか、そしてそれが展開の量子化に耐えられるかどうかは不明です。当社は、8 つのヨーロッパ言語、最大 5 つのターゲット言語データ スケール (100 時間から 2500 時間)、3 つのストリーミング層とオフライン デコード、および最大 4 つの公開テスト セットにわたる 0.6 B パラメーターのキャッシュ対応 FastConformer トランスデューサーの制御されたスイープによってこれらの質問に答えます。主な結果は、多言語初期化はレイテンシに制限された利点ではなく、データに制限された利点であるということです。 160 ミリ秒の FLEURS では、平均 EN-ML 単語誤り率 (WER) ギャップが 100 時間の +4.21 パーセント ポイント (pp) から 2500 時間の +0.20 pp に減少しました。べき乗則当てはめはこの減衰を要約し、ターゲット言語データが 2 倍になるごとに残りの利点がほぼ半分になります。 3 つのストリーミング層全体で、言語間の平均 EN-ML ギャップは 100 時間から 1000 時間までの各スケールでほぼ安定しており、2500 時間までにほぼゼロになります。最後に、一致する 560 ミリ秒のストリーミング層での 4 ビットの重みのみのエンコーダー量子化により、エンコーダーのフットプリントが約 3 分の 1 に削減され、FLEURS WER は平均約 0.5 pp 増加します。結果として得られるガイドラインはシンプルです。低データ領域では多言語初期化を使用し、大規模データでは選択を実質的に無関係なものとして扱い、レイテンシと量子化の決定を独立して行います。

原文 (English)

Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR

Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an English-only (EN) encoder. The common intuition is that the multilingual encoder should help most at low data, but it is unclear how long that advantage persists, whether tight streaming latency amplifies it, and whether it survives deployment quantization. We answer these questions with a controlled sweep of a 0.6 B-parameter cache-aware FastConformer transducer across eight European languages, up to five target-language data scales (100 h to 2500 h), three streaming tiers plus offline decoding, and up to four public test sets. The main result is that multilingual initialization is a data-limited advantage, not a latency-limited one. On FLEURS at 160 ms, the mean EN-ML word error rate (WER) gap falls from +4.21 percentage points (pp) at 100 h to +0.20 pp at 2500 h; a power-law fit summarizes this decay, with each doubling of target-language data roughly halving the remaining advantage. Across the three streaming tiers, the across-language mean EN-ML gap is approximately stable at each scale from 100 to 1000 h, and is near zero by 2500 h. Finally, 4-bit weight-only encoder quantization at the matched 560 ms streaming tier reduces the encoder footprint by about 3x, with an average FLEURS WER increase of about 0.5 pp. The resulting guideline is simple: use multilingual initialization in low-data regimes, treat the choice as effectively irrelevant at large data, and make latency and quantization decisions independently.

13:00 JST研究/論文

パーソナライズされたマルチモーダル生成に向けてユーザー行動をナビゲートする

最新の AIGC パイプラインは高忠実度の画像とビデオを提供しますが、整形式の作成指示を前提としていますが、エンドユーザーが視覚的な詳細を明確にすることはほとんどなく、ジェネレーターがユーザーの要求とずれたままになっています。私たちは、ユーザーのインタラクション履歴を下流合成用の実行可能な命令に変換するパーソナライズされたコンテンツ生成を研究し、2 つの障害を特定しました。1 つは言語推論が読みやすい形式で動作をエンコードする必要があること、もう 1 つはモデルが事前トレーニングと動作データの両方に存在しない命令作成スキルを獲得することです。私たちは NaviGen を提案します。NaviGen は、1 つのトークン ストリーム内の動作基盤およびセマンティック ブリッジとして、協調的なコードとテキスト コードを結合する二重識別子で各項目を表します。この表現では、2 段階の SFT+RL パイプラインが最初に進化的に検索された監視から優先推論と命令記述を抽出し、次に階層的で自己矛盾のない報酬を通じてユーザーの意図に合わせて生成を調整します。製品、ゲーム、ショートビデオの各分野にわたる実験では、NaviGen がパーソナライズされた画像とビデオの生成を改善し、次のアイテムの予測を強化し、より具体的で関連性のある視覚的に生成可能な指示を生成できることが示されています。私たちのコードは匿名で https://github.com/iLearn-Lab/NaviGen で公開されています。

原文 (English)

Navigating User Behavior toward Personalized Multimodal Generation

Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, leaving generators misaligned with user demand. We study personalized content generation, which turns a user's interaction history into an executable instruction for downstream synthesis, and identify two obstacles: behavior must be encoded in a form legible to language reasoning, and the model must acquire instruction-writing skill absent from both pretraining and behavior data. We propose NaviGen, which represents each item with a dual identifier coupling a collaborative code and a textual code as a behavioral substrate and a semantic bridge in one token stream. On this representation, a two-stage SFT+RL pipeline first distills preference reasoning and instruction writing from evolutionarily searched supervision, then aligns generation with user intent through hierarchical and self-consistent rewards. Experiments across product, game, and short-video domains show that NaviGen improves personalized image and video generation, strengthens next-item prediction, and yields more specific, relevant, and visually generatable instructions. Our code is anonymously released at: https://github.com/iLearn-Lab/NaviGen.

13:00 JST研究/論文

人間中心の AI と企業の特異なリスクとの関係を探る

インダストリー 5.0 における人間中心の AI (HCAI) については広範な議論が行われているにもかかわらず、企業の特異リスク (IR) に対するその影響はまだ十分に解明されていません。これは、企業レベルの株価のボラティリティを体系的な要因から分離することで、企業の異種AI戦略や導入に対する投資家の反応を反映するものであり、現在のテクノロジー革命の中で財務リスクを乗り切る企業にとって緊急の課題となっている。状況に応じた AI 理論と社会技術システム理論を統合することで、私たちは HCAI を状況に応じた AI 戦略として概念化します。これは、AI 関連の倫理的リスクを軽減し、企業の事業運営における AI と人間の相乗効果を促進し、ステークホルダーの多様な期待に合わせることで最終的に IR を削減します。さらに、デジタル化、業務効率、経営陣の株式保有、IT の背景を持つ CEO などの社会技術的要因が、HCAI と IR の関係を緩和する可能性があります。 2015 年から 2023 年までの中国の上場企業のマルチソース パネル データセットを使用したところ、HCAI が企業の IR の低下と関連していることがわかりました。さらに、デジタル化と経営陣の株式保有はこのリスク軽減効果を強化しますが、業務効率化と IT の背景を持つ CEO は驚くほどその効果を弱めます。私たちの調査結果は、AI 時代の倫理的な AI ガバナンスと確実な財務リスク管理の両方に対する理論的貢献と実践的な洞察を提供します。

原文 (English)

Exploring the relationship between human-centric AI and firm idiosyncratic risks

Despite the extensive discussions of human-centric AI (HCAI) in Industry 5.0, its effects on firms' idiosyncratic risks (IR) remains underexplored. This is an imperative issue for firms navigate financial risks during the current technological revolution, as IR reflects investor reactions to corporate heterogeneous AI strategies and implementations by isolating firm-level stock volatility from systematic factors. Integrating situated AI theory with social-technical systems theory, we conceptualise HCAI as a situated AI strategy that reduces AI-related ethical risks and fosters AI-Human synergies in firms' business operations, ultimately reducing IR by aligning with stakeholders' diverse expectations. Moreover, socio-technical factors, namely digitalisation, operational efficiency, executive shareholding, and CEOs with IT background, may moderate the HCAI-IR relationship. Using a multi-source panel dataset of Chinese listed firms from 2015 to 2023, we find that HCAI is associated with lower firm IR. Furthermore, digitalisation and executive shareholding strengthen this risk-reducing effect, whereas operational efficiency and CEOs with IT background surprisingly attenuate it. Our findings offer theoretical contributions and practical insights for both ethical AI governance and firm financial risk management in the AI era.

13:00 JST研究/論文

FlowR2A: マルチモーダル運転計画のための報酬と行動の配分の学習

マルチモーダル運転計画は、2 つのパラダイムの間の長年の緊張に直面しています。スコアベースの方法は、緻密な報酬監視の恩恵を受けますが、固定されたアクション語彙に限定されます。一方、アンカーベースの方法は、提案を動的に生成しますが、単一のグラウンドトゥルース軌道に制限されるまばらな監視に悩まされます。この研究では、識別ターゲットからのシミュレーションベースの報酬を生成条件に再構築することで、この緊張を解決する FlowR2A を提案します。 FlowR2A は、フロー マッチング デコーダーを使用して高密度の軌道と報酬のペアから報酬条件付きアクションの分布を学習することで、スコアリング ベースの手法の高密度な監視とアンカー ベースの手法の提案生成を単一の生成モデルで統合し、安全性、進歩、快適さ、ルール遵守におけるアクションとその結果の間の相関関係をモデルに強制的に内部化します。ソフトな進捗目標に対してハードな安全制約のバランスをとるために、タイムステップごとのきめ細かい報酬条件付けと報酬ノイズの増大を導入します。生成的な定式化は、報酬ガイダンスとアンカー サンプリングを通じて制御可能なテスト時間のサンプリングを自然にサポートし、高品質の提案を生成します。 FlowR2A は、NAVSIM v1 および v2 ベンチマークで最先端の結果を達成し、従来の方法よりも大幅に高品質のマルチモーダル提案を実現します。

原文 (English)

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods generate proposals dynamically yet suffer from sparse supervision constrained to a single ground-truth trajectory. In this work, we propose FlowR2A, which resolves this tension by reframing simulation-based rewards from discriminative targets into generative conditions. By learning the reward-conditioned action distribution from dense trajectory-reward pairs with a flow-matching decoder, FlowR2A unifies the dense supervision of scoring-based methods with the proposal generation of anchor-based methods in a single generative model, forcing the model to internalize the correlation between an action and its outcomes in safety, progress, comfort, and rule compliance. To balance hard safety constraints against soft progress objectives, we introduce fine-grained per-timestep reward conditioning and reward noise augmentation. The generative formulation naturally supports controllable test-time sampling via reward guidance and anchored sampling, producing high-quality proposals. FlowR2A achieves state-of-the-art results on the NAVSIM v1 and v2 benchmarks, with multimodal proposals of substantially higher quality than prior methods.

13:00 JSTエージェント

SP-Mind: 空間プロテオミクス解析のための自律推論エージェント

空間プロテオミクスは、組織構造内のタンパク質発現の単一細胞解像度の特性評価を可能にし、腫瘍微小環境を理解し、精密医療を導く上で重要な役割を果たします。しかし、現在の分析ワークフローは断片化したままであり、異種ツールを専門家が手動で調整する必要があり、研究の拡張性と再現性が制限されています。生の多重組織イメージングから下流の表現型発見まで、空間プロテオミクス解析パイプラインを統合するように設計された初の自律型 AI エージェントである SP-Mind を紹介します。専門家が厳選した生物学的分析スキルと特殊な計算ツールを備えた SP-Mind は、タスク固有の微調整を行うことなく、自然言語クエリをエンドツーエンドの分析ワークフローに変換します。その機能を厳密に評価するために、18 の異なるカテゴリにわたる 102 のタスクで構成される、さまざまな組織タイプにわたる包括的なベンチマークである SP-Bench を導入します。 SP-Bench と確立された下流タスクでの広範な評価を通じて、SP-Mind は既存のオープンソース生物医学的エージェントのベースラインと比較して最先端のパフォーマンスを達成します。

原文 (English)

SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis

Spatial proteomics enables single-cell-resolution characterization of protein expression within tissue architecture, playing a critical role in understanding tumor microenvironments and guiding precision medicine. However, current analysis workflows remain fragmented, requiring expert manual orchestration of heterogeneous tools and limiting research scalability and reproducibility. We present SP-Mind, the first autonomous AI agent designed to unify the spatial proteomics analysis pipeline, from raw multiplexed tissue imaging to downstream phenotype discovery. Equipped with expert-curated biological analysis skills and specialized computational tools, SP-Mind converts natural-language queries into end-to-end analytical workflows without task-specific fine-tuning. To rigorously evaluate its capabilities, we introduce SP-Bench, a comprehensive benchmark spanning diverse tissue types, comprising 102 tasks across 18 distinct categories. Through extensive evaluation on SP-Bench and established downstream tasks, SP-Mind achieves state-of-the-art performance compared to existing open-source biomedical agent baselines.

13:00 JST研究/論文

フェデレーテッド・ロングテール・グラフ学習に向けて: エネルギーに導かれたデュアル・デカップリング・アプローチ

Federated Graph Learning は、データのプライバシーを維持しながら、分散クライアント間での共同グラフ モデリングを容易にします。ただし、現実世界のデータ カテゴリは、長い裾の分布を示すことがよくあります。このような統計的欠乏は、2 つの点でパフォーマンスを大幅に低下させます。1 つはグローバル モデルを多数派クラスに偏らせること、もう 1 つは少数派のノードを異好性の頭支配の近傍に沈めることで構造的に分離することです。既存の方法はトポロジーに依存しない統計的補償を試みますが、データ不足の下では失敗することがよくあります。末尾ノードを回復する代わりに、隣接する支配的なクラスからの構造ノイズを過剰適合させて、表現の劣化を引き起こします。これらの制限に対処するために、トポロジカルな浄化をセマンティックな再調整から分離する二重分離パラダイムに基づいて構築されたフレームワークである FedEPD を提案します。具体的には、FedEPD は分布を意識したディリクレ エネルギー プルーニングを利用して、空間的な異好性エッジをフィルター処理します。次に、トポロジー的に中心的なノードから堅牢なグローバル プロトタイプを抽出することで、非 IID 分布のシフトを克服します。これは、空間ローパス プロトタイプ インジェクションを介してローカル表現に組み込まれます。さらに、2 段階の交互最適化戦略により、少数派の精度を向上させながら、多数決の境界を厳密に保護します。広範な実験により、FedEPD がさまざまなロングテール ベンチマークにわたって最先端のパフォーマンスを達成し、精度で最大 4.97%、マクロ F1 で 5.48% の絶対的な向上が得られることが実証されました。

原文 (English)

Towards Federated Long-Tailed Graph Learning: An Energy-Guided Dual Decoupling Approach

Federated Graph Learning facilitates collaborative graph modeling across distributed clients while preserving data privacy. However, real-world data categories frequently exhibit long-tailed distributions. Such statistical scarcity severely degrades performance in two ways: it biases the global model toward majority classes, and it structurally isolates minority nodes by submerging them in heterophilic, head-dominated neighborhoods. While existing methods attempt topology-agnostic statistical compensations, they often fail under data scarcity. Instead of recovering tail nodes, they overfit the structural noise from adjacent dominant classes, leading to representation degradation. To address these limitations, we propose FedEPD, a framework built on a dual decoupling paradigm that separates topological purification from semantic recalibration. Specifically, FedEPD utilizes distribution-aware Dirichlet energy pruning to filter spatial heterophilic edges. It then overcomes Non-IID distribution shifts by extracting robust global prototypes from topologically central nodes, which are incorporated into local representations via a spatial low-pass prototype injection. Furthermore, a two stage alternating optimization strategy strictly protects majority decision boundaries while improving minority accuracy. Extensive experiments demonstrate that FedEPD achieves state-of-the-art performance across diverse long-tailed benchmarks, yielding absolute improvements of up to 4.97% in Accuracy and 5.48% in Macro-F1.

13:00 JST研究/論文

言語モデルの誤った思考プロセスを調査する

大規模な言語モデルでは、戦略的欺瞞、サンドバッグ、自己保存など、ますます多様な不整合な動作が見られます。一か八かの環境での導入が増えているため、安全かつ責任ある使用を確保するには、そのような動作を確実に検出することが重要です。この研究では、位置ずれをきめの細かい認知プロセス (位置ずれ指標) に分解し、線形プローブを介してモデルの内部活性化における位置ずれの存在を検出することで、位置ずれを監視することを提案します。私たちは、さまざまな不整合な行動にわたる 18 の指標の分類を開発し、複数ターンのトレーニング会話を生成する自動化されたメタプランに基づくパイプラインと組み合わせます。一般化を厳密に評価するために、自動化された行動誘発、確立された不整合ベンチマーク、および自然で無害な会話を組み合わせた配布外スイートを構築します。 5 つの不整合な動作にわたって、当社のプローブは、無害なトラフィックでの低い誤検知率を維持しながら、配布外ベンチマークで 0.935 AUROC という強力な LLM ジャッジと一致しました。さらに詳細な分析を実行して、プローブと位置ずれ指標のモデルの内部表現を理解します。

原文 (English)

Probing the Misaligned Thinking Process of Language Models

Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-stakes settings, it is critical to reliably detect such behaviors to ensure safe and responsible use. In this work, we propose to monitor misalignment by decomposing it into fine-grained cognitive processes -- misalignment indicators -- and detecting their presence in a model's internal activations via linear probes. We develop a taxonomy of 18 indicators spanning different misaligned behaviors, paired with an automated, meta-plan-guided pipeline that generates multi-turn training conversations. To rigorously evaluate generalization, we construct an out-of-distribution suite combining automated behavioral elicitation, established misalignment benchmarks, and natural benign conversations. Across 5 misaligned behaviors, our probes match a strong LLM judge with 0.935 AUROC on out-of-distribution benchmarks while keeping a low false positive rate on benign traffic. We further perform in-depth analysis to understand the probes and the model's internal representations of misalignment indicators.

13:00 JST研究/論文

合理的閉鎖の下での防御可能な DL-Lite のための扱いやすい推論と論理積クエリ応答

記述論理 (DL) では、合理的閉包 (RC) に基づく推論は、実行可能な知識を処理するためのよく知られ、広く受け入れられている非単調形式主義です。この論文では、軽量記述ロジックの DL-Lite ファミリのコアおよびホーンのバリアントへの RC の適用を研究します。 RC では、資格 (インスタンスのチェック) と接続クエリ (CQ) 応答の両方を分析します。私たちの主な貢献は、既存の標準的な古典的推論に基づいて構築されたプラグイン アーキテクチャを提供し、DL-Lite の RC での推論と CQ 応答が最小限の計算オーバーヘッドで効率的に実行できることを確立したことです。

原文 (English)

Tractable Reasoning and Conjunctive Query Answering for Defeasible DL-Lite under Rational Closure

In Description Logics (DLs), reasoning under Rational Closure (RC) is a well-known and widely accepted non-monotonic formalism to handle defeasible knowledge. In this paper, we study the application of RC to the core and horn variants of the DL-Lite family of lightweight description logics. We analyze both entitlement (instance checking) and Conjunctive Query (CQ) answering under RC. Our main contribution is providing a plug-in architecture that builds upon existing standard classical reasoners, establishing that reasoning and CQ answering under RC for DL-Lite can be done efficiently with minimal computational overhead.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

レモンハーネス技術レポート

大規模言語モデル (LLM) エージェントがより長いタスクに適用されると、複数ラウンドの反復にわたってワークスペースの状態がますます変更されます。ただし、エージェントは通常、ツールの出力とログの断片のみを観察し、実際の状態の変化はファイル システムで発生します。明示的なワークスペース境界がないと、ファイルの書き込みや一時的なアーティファクトの生成などの状態変更操作により、パス全体に変更が分散される可能性があります。時間の経過とともに、これらの弱く制約された変更が蓄積され、変更されたファイルなどの状態を追跡することが困難になります。この文書では、長期的なエージェント向けの統合実行フレームワークである LemonHarness について説明します。 LemonHarness は、明確に定義されたワークスペース内で状態変更操作を制限し、モデルの呼び出し、ツールの実行、およびルールの知識を単一の制御された境界内に持ち込むことによって、明示的な実行境界を確立します。ファイルの書き込み、依存関係のインストール、一時的なアーティファクトの作成などの状態変更操作は、構造化されたツール インターフェイスを通じて実行され、実行フィードバックが観察として記録され、後続のモデル決定に利用できます。このシステムには、再利用可能なルール ナレッジ ベースも導入されており、繰り返し実行ルールと受け入れ基準がランタイム ナレッジに変わります。 LemonHarness はさらに、経過予算と残り予算をモデルに公開する時間認識実行メカニズムを追加します。これにより、時間プレッシャーの変化に応じて調査、実装、検証作業のバランスを再調整し、長時間の待機や過剰な検証によるタイムアウトを回避できます。 Terminal-Bench 2.0 では、LemonHarness_GPT-5.3-CodeX は 445 回のトライアルで 84.49% の精度に達しました。同じフレームワークとより強力な GPT-5.5 バックボーンを組み合わせることで、5 つのジョブの平均精度が 86.52% に向上しました。この結果は、統合された実行時間境界、呼び出し可能なルールの知識、および時間認識の実行により、長期的なエージェント実行の安定性が向上する可能性があることを示唆しています。

原文 (English)

LemonHarness Technical Report

As large language model (LLM) agents are applied to longer tasks, they increasingly modify workspace state across multiple rounds of iteration. However, agents typically observe only tool outputs and log fragments, while the actual state changes occur in the file system. Without explicit workspace boundaries, state-changing operations such as file writes and temporary artifact generation may scatter changes across paths. Over time, these weakly constrained changes accumulate, making states such as modified files difficult to track. This paper presents LemonHarness, an integrated execution framework for long-horizon agents. LemonHarness establishes an explicit execution boundary by constraining state-changing operations within a clearly defined workspace and bringing model invocation, tool execution, and rule knowledge within a single controlled boundary. State-changing operations, including file writes, dependency installation, and temporary artifact creation, are executed through structured tool interfaces, with execution feedback recorded as observations available to subsequent model decisions. The system also introduces a reusable rule knowledge base, which turns recurring execution rules and acceptance criteria into runtime knowledge. LemonHarness further adds a time-aware execution mechanism that exposes elapsed and remaining budget to the model, so it can rebalance exploration, implementation, and validation effort as time pressure shifts and avoid timeouts from long waits or excessive verification. On Terminal-Bench 2.0, LemonHarness_GPT-5.3-CodeX reached 84.49% accuracy over 445 trials; pairing the same framework with the stronger GPT-5.5 backbone raised the average accuracy to 86.52% across five jobs. The results suggest that a unified runtime boundary, callable rule knowledge, and time-aware execution can improve the stability of long-horizon agent execution.

13:00 JST研究/論文

Prob-BBDM: MRI シーケンスの画像間変換のための確率的ブラウン橋拡散モデル

AI を活用した画像間の合成は急速に進歩しており、医療画像分野での応用が拡大しています。マルチモーダル画像解析は検査の品質を最適化する上で重要な役割を果たしますが、臨床現場で複数の画像モダリティを取得することは依然としてリソースを大量に消費し、特に 3D イメージングでは時間がかかります。この課題に対処するために、我々は、2D アキシャル スライスから磁気共鳴画像法 (MRI) シーケンスを合成するブラウン橋拡散モデル (BBDM) に基づく新しい画像間変換モデルを提案します。私たちのアプローチは、変分エンコーダーによる拡散メカニズムを統合し、確率的な画像分布を活用して合成品質を向上させます。 BraTS 2021 データセットで評価された当社の Probabilistic-BBDM (Prob-BBDM) は、複数の翻訳タスクにわたって優れたパフォーマンスを達成し、最大 88.46% の SSIM と 26.09 dB PSNR に達し、ベースラインを一貫して改善しています。特に、当社の拡散プロセスに必要なステップは 4 つだけであり、高品質の合成を維持しながら計算効率が高くなります。一般性をさらに検証するために、外部のサードパーティ データセットで Prob-BBDM をテストし、ドメイン全体で一貫したパフォーマンスを実証します。さらに、合成されたスライスを事前にトレーニングされたセグメンテーション モデルへの入力として使用することで、そのスライスの臨床的有用性を評価します。腫瘍のセグメンテーションにより、88.71% の Dice スコアと 3.49 mm の HD95 が得られ、合成されたスライスが重要な診断情報を保存していることが確認されました。これらの結果は、高品質、効率的、汎用性のある MRI 合成に対する Prob-BBDM の可能性を強調しており、医用画像変換の改善に向けた有望な一歩を提供します。

原文 (English)

Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation

AI-driven image-to-image synthesis is rapidly advancing, with growing applications in medical imaging. Multi-modal image analysis plays a crucial role in optimizing examination quality, yet acquiring multiple imaging modalities in clinical settings remains resource-intensive and time-consuming, especially for 3D imaging. To address this challenge, we propose a novel image-to-image translation model based on Brownian Bridge Diffusion Models (BBDM), which synthesizes magnetic resonance imaging (MRI) sequences from 2D axial slices. Our approach integrates a variational encoder-guided diffusion mechanism, leveraging probabilistic image distributions to enhance synthesis quality. Evaluated on the BraTS 2021 dataset, our Probabilistic-BBDM (Prob-BBDM) achieves superior performance across multiple translation tasks, reaching up to 88.46% SSIM and 26.09 dB PSNR, with consistent improvements over baselines. Notably, our diffusion process requires only 4 steps, making it computationally efficient while maintaining high-quality synthesis. To further validate generalizability, we test Prob-BBDM on an external third-party dataset, demonstrating consistent performance across domains. Additionally, we assess the clinical utility of the synthesized slices by using them as input to a pre-trained segmentation model. Tumor segmentation yields a Dice score of 88.71% and an HD95 of 3.49 mm, confirming that the synthesized slices preserve critical diagnostic information. These results highlight the potential of Prob-BBDM for high-quality, efficient, and generalizable MRI synthesis, offering a promising step toward improved medical image translation.

13:00 JST研究/論文

MVG-KAN: PM$_{2.5}$ 予測用のマルチビュー地風ガイド付き KAN

正確な短期 PM$_{2.5}$ 予測は、公衆衛生保護、大気質早期警報、都市環境管理にとって重要です。しかし、PM$_{2.5}$ の変動は、人間の活動や気象の規則性によって引き起こされる安定した周期的変化、観測所固有の短期濃度の変化、観測所間の気象学に起因する汚染物質の分散など、複数の複合要因によって引き起こされます。既存の時空間予測手法は観測点の関係をある程度把握できますが、距離のみ、相関ベース、または純粋に適応的なグラフでは、これらの不均一な要因、特に風向に依存する汚染物質の輸送を包括的に表現するには不十分なことがよくあります。この問題に対処するために、\textbf{MVG-KAN} という名前の PM$_{2.5}$ 予測用のマルチビュー地理風ガイド KAN モデルを提案します。このモデルは、局所的な周期規則性、観測点ごとの残留時間ダイナミクス、気象環境に誘導された空間分散という 3 つの相補的なビューから観測点レベルの PM$_{2.5}$ の進化をモデル化します。具体的には、周期的残差予測バックボーンは、まず、安定した日次および週次パターンを非周期的残差変動から分離します。 Geo-Wind Graph は、地理的距離の減衰と風向および風速を意識した伝送を組み合わせて構築され、ステーション間の残留伝播に対して軽量の物理的動機による有向空間事前分布を提供します。さらに、時間的コルモゴロフ-アーノルド ネットワーク (TKAN) 残差ヘッドを導入して、非周期化 PM$_{2.5}$ 残差と過去の複数汚染物質シーケンスからステーションごとの非線形自己回帰補正を学習し、それによって局所的な残留慣性と汚染物質の共変動のモデリングを強化します。

原文 (English)

MVG-KAN: Multi-View Geo-Wind Guided KAN for PM$_{2.5}$ Forecasting

Accurate short-term PM$_{2.5}$ forecasting is important for public health protection, air-quality early warning, and urban environmental management. However, PM$_{2.5}$ variation is driven by multiple coupled factors, including stable periodic changes induced by human activities and meteorological regularity, station-specific short-term concentration evolution, and meteorology-driven pollutant dispersion among monitoring stations. Existing spatio-temporal forecasting methods may capture station relationships to some extent, but distance-only, correlation-based, or purely adaptive graphs are often insufficient to comprehensively represent these heterogeneous factors, especially wind-direction-dependent pollutant transport. To address this problem, we propose a Multi-View Geo-Wind Guided KAN model for PM$_{2.5}$ forecasting, named \textbf{MVG-KAN}, which models station-level PM$_{2.5}$ evolution from three complementary views: local periodic regularity, station-wise residual temporal dynamics, and meteorological-environment-guided spatial dispersion. Specifically, the periodic-residual forecasting backbone first separates stable daily and weekly patterns from non-periodic residual variations. A Geo-Wind Graph is constructed by combining geographic distance decay with wind-direction- and wind-speed-aware transport, providing a lightweight physically motivated directed spatial prior for residual propagation among stations. In addition, a temporal Kolmogorov-Arnold network (TKAN) residual head is then introduced to learn station-wise nonlinear autoregressive correction from de-periodized PM$_{2.5}$ residuals and historical multi-pollutant sequences, thereby enhancing the modeling of local residual inertia and pollutant co-variation.

13:00 JSTLLM/生成AI

拡散ベースの並列処理とトレーナー支援生成によるビジュアル生成 LLM の分離 RL の高速化

強化学習 (RL) はトレーニング後のパラダイムの主流となっており、自己回帰大規模言語モデル (LLM) 用の veRL などの高性能 RL システムの出現を推進しています。並行して、DanceGRPO や FlowGRPO などの拡散指向の RL アルゴリズムにより、RL の範囲が言語推論から拡散ベースのビジュアルおよびフローベースの生成へと急速に拡大されました。ただし、拡散生成 LLM 用の効率的な RL システムはまだ研究されていません。既存の実装 (veRL-Omni など) は依然としてコロケーション実行に依存しており、同期は簡素化されますが、ロールアウトとトレーニングのリソースが結合され、異種導入が制限され、独立したスケーリングが制約されます。この目的を達成するために、柔軟なリソース割り当てをサポートし、異種 GPU に対応し、効率的なタスク スケジューリングを促進する、拡散ベースの生成 LLM 用の分散 RL フレームワークである DigenRL を導入します。分散アーキテクチャにおける実行バブルを最大限に減らすために、次のことを提案します。1) 拡散アーキテクチャにおける世代軸パイプライン (GAP) とタイムステップ並列処理 (TSP) により、ロールアウトとトレーニングの間のよりきめの細かいパイプライン処理が可能になります。 2) トレーナー GPU リソースがロールアウト世代の実行を動的に支援できるようにするエラスティック トレーナー支援生成 (TAG) アプローチ。 3) パイプラインのテール バブルをさらに利用するための、厳密に 1 ステップで制約された非同期戦略。 HunyuanVideo-13B、Wan2.1-14B、FLUX.1-12B、および QwenImage-20B 生成モデルを使用して、16 ~ 32 GPU を備えた 3 つのハードウェア テストベッドで広範な実験が行われています。実験結果は、DigenRL が最先端の拡散 RL システムである veRL-Omni および GenRL と比較して 1.56 ~ 2.10 倍のスループット向上を達成することを示しています。

原文 (English)

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models (LLMs). In parallel, diffusion-oriented RL algorithms, e.g., DanceGRPO and FlowGRPO, have rapidly expanded the scope of RL from language reasoning to diffusion-based visual and flow-based generation. However, efficient RL systems for diffusion generative LLMs remain underexplored. Existing implementations, e.g., veRL-Omni, still rely on colocated execution, which simplifies synchronization but couples rollout and training resources, limits heterogeneous deployment, and constrains independent scaling. To this end, we introduce DigenRL, a disaggregated RL framework for diffusion-based generative LLMs that supports flexible resource allocation, accommodates heterogeneous GPUs, and facilitates efficient task scheduling. To maximally reduce the execution bubbles in the disaggregated architecture, we propose: 1) a generation-axis pipeline (GAP) and time-step parallelism (TSP) in the diffusion architecture to enable finer-grained pipelining between rollout and training; 2) an elastic trainer-assisted generation (TAG) approach to enable the trainer GPU resources to dynamically assist in executing rollout generations; and 3) a tightly one-step constrained asynchronous strategy to further utilize the tail bubble in the pipeline. Extensive experiments are conducted on three hardware testbeds with 16-32 GPUs using HunyuanVideo-13B, Wan2.1-14B, FLUX.1-12B, and QwenImage-20B generative models. Experimental results show that DigenRL achieves 1.56-2.10x throughput improvements over state-of-the-art diffusion RL systems, veRL-Omni and GenRL.

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTGemini

有用性が因果関係を無効にする場合の注意: LLM におけるコンテキスト依存の抑制と回復

大規模言語モデル (LLM) は、ビジネスおよび政策のコンテキストにおける意思決定支援の役割にますます統合されています。これまでのベンチマーク研究では主に LLM の因果推論能力が評価されてきましたが、より基本的な認識論的側面は見落とされてきました。因果的注意とは、経験的証拠が不十分な場合に因果関係の判断を控える傾向として定義されます。この研究では、LLM が学術的な文脈から実践的な助言の文脈に移行するときに発生する、因果的注意の体系的な抑制を調査します。 Pearl の因果階層 (PCH スコア) にヒントを得た評価ルーブリックを使用して、4 つの高性能 LLM (Claude Sonnet 4.6、Claude Opus 4.7、GPT 5.5、および Gemini 3.1 Pro) で 480 回のトライアルにわたって実験を実施しました。因果関係の注意維持率は、学術的な文脈では 91.7 ~ 100.0% でしたが、実践的なアドバイスの文脈では 6.7 ~ 18.3% に低下しました (フィッシャーの直接確率検定、すべてのモデルで p < .001)。さらに、具体的な推奨事項や説明の根拠を求める実際的なプロンプトに限定すると、因果関係注意を維持した回答は 200 件中 1 件 (0.5%) のみでした。 「因果関係の観点からこの判断を再考してください」という短い自己修正プロンプトにより、因果関係注意の表現が 71.4 ~ 100.0% の維持率に戻りました (マクネマーの検定、すべてのモデルで p < .001)。これらの結果は、有用性指向の応答パターンが実際の助言の場面での因果的注意の表現を抑制する可能性があり、組織のガバナンスに重要な影響を与える可能性があることを示唆しています。この調査結果は、この抑制が根底にある機能制限ではなく、表現におけるコンテキスト依存の変動を反映していることを示しており、提案生成と因果関係の監査を分離するマルチエージェント アーキテクチャが有望なガバナンス設計を提供する可能性があることを示唆しています。

原文 (English)

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs

Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts. While prior benchmark studies have primarily evaluated LLMs' causal reasoning capabilities, a more fundamental epistemic dimension has been overlooked: Causal Caution, defined as the propensity to refrain from causal judgment when empirical evidence is insufficient. This study examines the systematic suppression of Causal Caution that occurs when LLMs shift from academic to practical advisory contexts. Using an evaluation rubric inspired by Pearl's Causal Hierarchy (the PCH score), we conducted experiments on four high-performance LLMs -- Claude Sonnet 4.6, Claude Opus 4.7, GPT 5.5, and Gemini 3.1 Pro -- across 480 trials. Causal Caution maintenance rates were 91.7--100.0% in academic contexts but dropped to 6.7--18.3% in practical advisory contexts (Fisher's exact test, p < .001 across all models). Furthermore, when restricted to practical prompts requesting concrete recommendations or explanatory rationales, only 1 of 200 responses (0.5%) maintained Causal Caution. A brief self-correction prompt -- "Please reconsider this judgment from the perspective of causal relationships" -- restored the expression of Causal Caution to maintenance rates of 71.4--100.0% (McNemar's test, p < .001 across all models). These results suggest that helpfulness-oriented response patterns may suppress the expression of Causal Caution in practical advisory contexts, with important implications for organizational governance. The findings indicate that this suppression reflects context-dependent variation in expression rather than an underlying capability limitation, suggesting that multi-agent architectures that separate proposal generation from causal auditing may offer a promising governance design.

13:00 JST研究/論文

PHANTOM: 視覚言語モデルに対するマルチモーダル敵対的攻撃の大規模データセット

ビジョン言語モデル (VLM) 用に事前に生成された敵対的攻撃の大規模なオープンソース データセットを紹介します。このデータセットは、多様性があり、代表的で、実用的になるように設計されており、有害な意図の 10 の高レベル カテゴリと 55 のサブカテゴリをカバーすることで既存のベンチマークを拡張します。私たちの主な目標は、大量の攻撃を生成する計算コストと複雑さを考慮して、研究コミュニティが敵対的なデータにアクセスできるようにすることです。データセットは、最近の文献からの最先端の攻撃戦略を使用して生成された 47,524 個の敵対的サンプルで構成されています。私たちの取り組みは、複数の確立されたソースからの以前のベンチマークを統合および拡張することで既存の取り組みを補完し、その結果 7,826 のインテントが得られ、追加のカテゴリを導入して対象範囲を広げています。これにより、モデルの堅牢性と整合性を研究するための現実的な評価リソースが提供されます。私たちのデータセットは、研究者や実践者が VLM の堅牢性と安全性を体系的に評価し、攻撃生成モデルを微調整し、さまざまな敵対状況下で防御ガードレールを開発またはストレス テストできるようにすることを目的としています。このリソースを公開することで、敵対的研究への障壁を下げ、VLM の安全性についてより再現性があり、包括的で比較可能な評価を促進することを目指しています。

原文 (English)

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

We introduce a large-scale, open-source dataset of pre-generated adversarial attacks for vision-language models (VLMs). The dataset is designed to be diverse, representative, and practical, extending existing benchmarks by covering 10 high-level categories and 55 subcategories of harmful intents. Our primary goal is to make adversarial data accessible to the research community, given the computational cost and complexity of generating large numbers of attacks. The dataset comprises 47 524 adversarial samples, generated using state-of-the-art attack strategies from recent literature. Our work complements existing efforts by consolidating and extending prior benchmarks from multiple established sources, resulting in 7 826 intents, and introduce an additional category to broaden coverage. This provides realistic evaluation resources for studying model robustness and alignment. Our dataset intends to enable researchers and practitioners to systematically evaluate the robustness and safety of VLMs, fine-tune attack-generation models, and develop or stress-test defensive guardrails under diverse adversarial conditions. By releasing this resource, we aim to lower the barrier to adversarial research and foster more reproducible, comprehensive, and comparable evaluations of VLM safety.

13:00 JSTLLM/生成AI研究/論文

LLM の時代: 戦争の霧の下での推論、外交、大規模言語モデルの信頼性のための戦略的な 1 対 1 ベンチマーク

Age of LLM を紹介します。これは、2 つの LLM が 13x7 グリッドで対決して敵の基地を破壊する、ターンベースの 1v1 ベンチマークです。 3 つのストレス要因は意図的なものです。戦争の霧、完全な外交 (メッセージ、停戦、最後通牒、ウランは秘密にされます)、そして毎ターン厳格な JSON スキーマに従わなければならず、違法行為は黙って破棄される信頼性の側面です。エンジンはプライベートであり、各試合では新鮮なランダムなマップ シードと対戦相手が使用されるため、公開ベンチマークに影響を与えるデータ汚染が軽減されます。モデルは、構築順序のアドバイスのない (ほぼ) ルールのみのプロンプトを受け取ります (データ収集中に 2 つの戦術的なシード フレーズが存在しました。セクション 2.7 を参照)。 54 の一致と 5,258 のアクションにわたって 15 の推論モデルをベンチマークしました。調査結果: (1) 核ラッシュは、認知的抑止力の失敗ではなく、機密同時発射ルールの下では主に機械的な単独発射機の署名で支配的 (ルール一貫性 v0.11+ サブコーパスで 78%、コーパス全体で 85%)。 (2) 軍事征服はまれですが、より高速です (12.3 対 18.9 ターン)。 (3) 外交は多作だが、ほとんど完了していない。 (4) 違法行為の ~58% はフォグ/ステート エラーであり、違法行為の割合が信念追跡の尺度になります。 (5) -- 最も確立されておらず、我々が探索的と名付けた唯一のもの -- 弱いリンクは、信頼性と勝利を結びつけます。コーパスは小さく、バランスがとれておらず、左右が入れ替わっていないため、ランキングは予備的な説明的なビューであり、貢献するものではありません。ランキングを超えて、アクションとメッセージのターンごとの追跡により、このコーパスは、LLM が敵対的な不確実性の下でどのように推論するか、つまり信念追跡、自発的欺瞞、およびモデルごとの認知「ペルソナ」についてのレンズとなり、私たちはそれを将来の研究の方向性として組み立てます。リプレイ フォーマット、アイソメトリック ビューア、およびすべてのリプレイをリリースします。リクエストに応じてエンジンソースを提供します。

原文 (English)

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War

We introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base. Three stressors are deliberate: fog of war, full diplomacy (messages, ceasefires, ultimatums; uranium kept secret), and a reliability dimension where every turn must follow a strict JSON schema and an illegal action is silently discarded. The engine is private and each match uses a fresh random map seed and opponent, mitigating the data contamination that affects public benchmarks. Models receive a (near) rule-only prompt with no build-order advice (two tactical seed phrases were present during data collection; see Section 2.7). We benchmark 15 reasoning models across 54 matches and 5,258 actions. Findings: (1) the nuclear rush dominates (78% on the rules-coherent v0.11+ sub-corpus; 85% corpus-wide) with a sole-launcher signature that is largely mechanical under secret-simultaneous launch rules, not a cognitive deterrence failure; (2) military conquest is rare but faster (12.3 vs 18.9 turns); (3) diplomacy is prolific yet almost never consummated; (4) ~58% of illegal actions are fog/state errors, making the illegal-action rate a measure of belief-tracking; (5) -- the least established, and the only one we label exploratory -- a weak link associates reliability with winning. The corpus is small, unbalanced and not side-swapped, so the ranking is a preliminary descriptive view, not a contribution. Beyond ranking, the turn-by-turn traces of actions and messages make the corpus a lens on how LLMs reason under adversarial uncertainty -- their belief-tracking, spontaneous deception, and per-model cognitive "personas" -- which we frame as a future research direction. We release the replay format, an isometric viewer and all replays; engine source on request.

13:00 JSTエージェント

ATRIA: 反復エージェントを使用した適応型追跡可能な ECG レポート

既存の ECG レポート生成は緊密に結合されており、解釈とレポートがエンドツーエンドで融合されているため、ステージレベルの手段を必要とせずにエラーが伝播します。一方、エージェントベースのシステムはタスクを分離しますがシングルパスのままで、以前の出力を再検討することはありません。代わりに、臨床 ECG レポートは反復的に展開され、段階的なコンテキスト統合と双方向編集が必要になります。我々は、臨床医の反復的なワークフローを反映するマルチエージェント ECG レポート システムである \textsc{ATRIA} を紹介します。これは、すべてのレポートの主張をその裏付けとなる証拠に結び付け、その証拠によって裏付けられていないステートメントにフラグを立て、セッション中に追加のコンテキストを組み込み、臨床医が 1 つの不透明な出力を受け入れるのではなく、個々の所見を検証して修正できるようにします。そのエージェントはすでに臨床で使用されている ECG 分析モデルを使用しているため、基礎となる所見は臨床的に信頼できるものです。また、クラウドベースの Web サービスとして、\textsc{ATRIA} はすぐに導入できる状態になっています。ライブ デモとビデオを利用して、4 つのインタラクション ケースを通じて \textsc{ATRIA} をデモンストレーションします。

原文 (English)

ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents

Existing ECG report generation is tightly coupled -- interpretation and reporting fused end-to-end, so errors propagate without stage-level recourse -- while agent-based systems decouple tasks but remain single-pass, never revisiting earlier outputs. Clinical ECG reporting instead unfolds iteratively, requiring progressive context integration and bidirectional editing. We present \textsc{ATRIA}, a multi-agent ECG reporting system that mirrors the clinician's iterative workflow: it binds every report claim to its supporting evidence, flags statements unsupported by that evidence, incorporates additional context mid-session, and lets clinicians verify and revise individual findings rather than accept one opaque output. Because its agents use ECG analysis models already in clinical use, the underlying findings are clinically trustworthy; and as a cloud-based web service, \textsc{ATRIA} is ready for immediate deployment. We demonstrate \textsc{ATRIA} through four interaction cases, with a live demo and video available.

13:00 JST研究/論文

正式検証証明書のサイクル一貫性のあるニューラル説明

正式な検証では、一時的特性の満足または違反を証明する機械チェック可能な証明書が生成されますが、これらの証明書は専門家以外の関係者には不透明なままです。私たちは、検証証明書の忠実な自然言語説明を生成するサイクル一貫性のあるニューラル アーキテクチャを提案します。順方向ネットワーク NN1 は証明書を説明にマッピングし、逆ネットワーク NN2 は説明から証明書を再構築します。シンボリックベリファイアはループを閉じ、微分可能な忠実性プロキシを提供します。ポインタ生成メカニズムは、証明書から状態名を直接コピーすることにより、語彙の基礎を確保します。私たちは、207 の指定州の金融コンプライアンス ドメインから抽出された YES と NO の両方の判定バリアントで、6 つの検証方法 (有界証明、K 帰納、帰納不変式、なげなわ、到達可能性、証人ペア) にわたる 420 のテスト証明書を評価します。ハイブリッド推論時間ルーティング戦略と組み合わせた当社のトレーニング済みアーキテクチャは、サイクル検証済みの健全性 90.0% を達成し、マルチ LLM の少数ショット ベースライン (4 つのフロンティア モデルにわたる 16 の LLM の組み合わせの最良の場合 76.1%) を 13.9 パーセント ポイント上回っています。ニューラル モデルは、12 の判定/種類カテゴリのうち 10 で勝利し、3 つのカテゴリが 100% の健全性に達しました。このアーキテクチャは、860 倍の高速推論 (完全なマルチ LLM ベースラインでは証明書ごとに 160 秒であるのに対し 185 ミリ秒)、オフライン操作、確定的な出力、および推論ごとのコストゼロを提供します。これらの結果は、トレーニングされた専門化が、クラウドベースの推論の展開上の制約を排除しながら、構造化された証明書の説明を求める汎用 LLM よりも優れたパフォーマンスを発揮することを示しています。

原文 (English)

Cycle-Consistent Neural Explanation of Formal Verification Certificates

Formal verification produces machine-checkable certificates that attest to the satisfaction or violation of temporal properties, yet these certificates remain opaque to non-specialist stakeholders. We propose a cycle-consistent neural architecture that generates faithful natural language explanations of verification certificates. A forward network NN1 maps certificates to explanations, and an inverse network NN2 reconstructs certificates from explanations; a symbolic verifier closes the loop, providing a differentiable faithfulness proxy. A pointer-generator mechanism ensures lexical grounding by copying state names directly from the certificate. We evaluate on 420 test certificates spanning six verification methods (bounded proof, k-induction, inductive invariant, lasso, reachability, witness pair) in both YES and NO verdict variants, drawn from a financial compliance domain with 207 named states. Our trained architecture, combined with a hybrid inference-time routing strategy, achieves 90.0% cycle-verified soundness, surpassing a multi- LLM few-shot baseline (76.1% for the best of 16 LLM combinations across four frontier models) by 13.9 percentage points. The neural model wins on 10 of 12 verdict/kind categories, with three categories reaching 100% soundness. The architecture offers 860x faster inference (185 ms vs. 160 s per certificate for the full multi-LLM baseline), offline operation, deterministic outputs, and zero per-inference cost. These results demonstrate that trained specialization outperforms general-purpose LLM prompting for structured certificate explanation, while eliminating the deployment constraints of cloud-based inference.

13:00 JSTエージェント

ポリシー主導の物理層システムのバイレベル長期最適化のための Agentic AI

ネットワーク事業者の変化するポリシー、サービス要件、および厳しいリアルタイム制約により、固定された目的と制約に従って設計された既存の手法は効果がなくなりました。このペーパーでは、適応物理層の問題構成に適用できる入れ子になった 2 レベルの最適化フレームワークである、Agentic 長期パフォーマンス最適化 (Agentic-LTPO) について説明します。重要なアイデアは、エージェント AI を使用して 2 レベルの最適化構造で上位レベルの構成を生成することです。進化するオペレーターのポリシー、環境の概要、および過去の経験が、構造化された下位レベルの最適化問題構成に変換されます。下位レベルでは、リアルタイムの物理層決定のための更新された構成の問題を解決します。セルフリー MIMO ビームフォーミングをユースケースとして考慮し、上位レベルで検索拡張された経験ベースの検証を備えた新しいマルチエージェント意思決定プロセスと、下位レベルのクローズドフォームビームフォーマーを設計することで、Agentic-LTPO を具体化します。実験では、Agentic-LTPO が動的なオペレーター ポリシーに対して強力な適応性を示し、従来の方法と比較してシステムの長期パフォーマンスを効果的に 57.2% 向上させることが実証されました。

原文 (English)

Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems

Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance optimization (Agentic-LTPO), a nested bilevel optimization framework that can be applied to adaptive physical layer problem configuration. The key idea is to employ agentic AI to generate upper-level configurations in a bilevel optimization structure, where evolving operator policies, environment summaries, and historical experiences are translated into structured lower-level optimization problem configurations. The lower level solves the problems with updated configurations for real-time physical-layer decisions. Considering cell-free MIMO beamforming as a use case, we embody Agentic-LTPO by designing a new multi-agent decision process with retrieval-augmented experience-based verification in the upper level, together with a closed-form beamformer in the lower level. Experiments demonstrate that Agentic-LTPO exhibits strong adaptability to dynamic operator policies and effectively enhances the system's long-term performance by 57.2% compared to traditional methods.

13:00 JST研究/論文Claude

集約不変式は連続的なサブグラフのマッチングを高速化できますか?限界、法則、および動的スペクトル指数

スペクトル フィルタリングは、最近 \emph{static} 部分グラフ マッチングに対して大幅な枝刈りを実現しました。ラプラシアン インターレースは、近傍がクエリをホストできない候補を拒否します。私たちは、このような集合構造テストが動的グラフ上で \emph{continuous} サブグラフ マッチング (CSM) を高速化できるかどうかを研究し、3 つの部分に分けて答えます。まず、スペクトルの枝刈りに価値がある場合、遅延的に維持されるスペクトルの境界は実行不可能です。形式化された摂動緩和に対する最も厳格な安全ルールを特徴付け、それが 4 回の更新内で本質的にすべての枝刈り能力を失うことを示します。第 2 に、選択的であれば正確なメンテナンスが手頃です。枝刈りのユーティリティと再計算のコストは頂点間で逆相関しており、ハブは枝刈りをしないことが証明されています。そのため、タッチで小さな近傍スペクトルを再計算すると、更新ごとにマイクロ秒単位で正確なローカル スペクトルが維持され、構築によって完了します。 3 番目に、同一マイナススペクトル コントロールに対する分離された CSM ベンチマークに統合されたテストでは、最大 $51\%$ の候補を削除するか、最大 $47\%$ の更新列挙を安全にスキップしますが、2 つのエンジン、4 つの実際のグラフ、2 つのストリーム タイプ、および $77$ の解決されたクエリにわたって、ゲートのスキップされた第 1 レベルのバインディング (通常はゼロ) を超えて、列挙の中間は変更されません。構築された半径階層化ワークロードにより、例外が存在する場合に機器が例外を検出することが確認されます ($-99.9\%$ 中間、$748\time$ 高速)。集約テストは、候補セット (構築、リスト スキャン) に応じてスケールするものを加速します。決して隣接関係に基づく探索ではありません。 CSM フィルターを評価するための中間不変性手法を抽出し、再利用可能な動的ローカル スペクトル インデックスをリリースします。

原文 (English)

Can Aggregate Invariants Accelerate Continuous Subgraph Matching? Limits, Laws, and a Dynamic Spectral Index

Spectral filtering recently delivered substantial pruning for \emph{static} subgraph matching: Laplacian interlacing rejects candidates whose neighborhoods cannot host the query. We study whether such aggregate structural tests can accelerate \emph{continuous} subgraph matching (CSM) over dynamic graphs, and answer in three parts. First, lazily maintained spectral bounds are infeasible exactly where spectral pruning has value: we characterize the tightest safe rule over a formalized perturbation relaxation and show that even it loses essentially all pruning power within four touching updates. Second, exact maintenance is affordable when selective: pruning utility and recomputation cost are anti-correlated across vertices -- hubs provably never prune -- so recomputing small-neighborhood spectra on touch sustains exact local spectra at microseconds per update, complete by construction. Third, integrated into a decoupled CSM benchmark against an identical-minus-spectra control, the tests remove up to $51\%$ of candidates or safely skip up to $47\%$ of update enumerations, yet enumeration intermediates remain unchanged -- beyond the gates' skipped first-level bindings, typically zero -- across two engines, four real graphs, two stream types, and $77$ solved queries; a constructed radius-stratified workload confirms the instrument detects the exception when one exists ($-99.9\%$ intermediates, $748\times$ faster). Aggregate tests accelerate what scales with candidate sets -- construction, list scans -- never adjacency-guided exploration. We distill an intermediate-invariance methodology for evaluating CSM filters and release a reusable dynamic local-spectra index.

13:00 JSTLLM/生成AIエージェント

ReM-MoA: 推論記憶がエージェント混合のスケーリングを維持する

Mixture-of-Agents (MoA) アーキテクチャは、複数の LLM エージェントを階層化された推論パイプラインに編成することで、推論時間のスケーリングを向上させます。ただし、既存の MoA バリアントは、深さが増加するにつれてゲインを維持できず、劣化、早期のプラトー状態、または飽和を示します。我々は、2 つのメカニズムを通じてスケーリングを維持するメモリ拡張 MoA フレームワークである ReM-MoA を提案します。(1) 比較レビューアー エージェントを使用して、すべてのレイヤーからの推論トレースを永続的に保存してランク付けするランク付き推論メモリ、(2) 成功したトレースと失敗したトレースの異なる組み合わせをさまざまなエージェントに公開し、高品質の推論を伝播しながら探索の多様性を維持するキュレーションされた多様なメモリ ルーティング スキーム。さらに、フロンティア モデルの監視を通じてランキングの品質を向上させる、オプションのマルチドメイン レビュアー蒸留パイプラインを導入します。数学、形式論理、コード、知識、常識に及ぶ 5 つの推論ベンチマークにわたって、ReM-MoA は深さと幅のスケーリングの両方で以前の MoA バリアントを常に上回っており、その利点は深さとともに拡大し、スケーラブルなマルチエージェント推論に欠けている重要なメカニズムとして構造化されたクロスレイヤー推論メモリを確立します。

原文 (English)

ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling

Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents into layered reasoning pipelines. However, existing MoA variants fail to sustain gains as depth increases, exhibiting degradation, early plateauing, or saturation. We propose ReM-MoA, a memory-augmented MoA framework that sustains scaling through two mechanisms: (1) a Ranked Reasoning Memory that persistently stores and ranks reasoning traces from all layers using a comparative Reviewer Agent, and (2) a Curated Diversified Memory Routing scheme that exposes different agents to distinct combinations of successful and failed traces, preserving exploration diversity while propagating high-quality reasoning. We further introduce an optional multi-domain Reviewer distillation pipeline that improves ranking quality through frontier-model supervision. Across five reasoning benchmarks spanning math, formal logic, code, knowledge, and commonsense, ReM-MoA consistently outperforms prior MoA variants across both depth and width scaling, and its advantage widens with depth, establishing structured cross-layer reasoning memory as a key missing mechanism for scalable multi-agent inference.

13:00 JSTLLM/生成AIエージェント

コーディングエージェントのベイジアン制御

最新のコーディング エージェントは、LLM ジェネレーターを、安価な診断や高価な検証ツールなどのさまざまなツールと組み合わせます。ツールの使用に関する決定は通常、固定ルールを使用し、不確実性を無視するオーケストレーターによって管理されます。私たちは、オーケストレーションをコスト重視の逐次仮説テストとして定式化します。ベイジアン コントローラーは、候補の正しさに対する信念を維持し、より多くの証拠を収集するか、候補を絞り込むか、検証するか、中止するかを動的に決定します。 6 つのジェネレーターと 9 つのコーディング ベンチマークにわたって、ベイジアン制御が最も価値があることが証明されるのは、検証にコストがかかり、批評家が有益ではあるが不完全な場合です。制御を超えて、信念状態は、不確実性の定量化において、トークンの確率や生のツールの成功ベースラインを上回る、解釈可能な正確性スコアを生成します。

原文 (English)

Bayesian control for coding agents

Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive verifiers. The tool-use decisions are typically governed by orchestrators that often use fixed rules and ignore uncertainty. We formulate orchestration as cost-sensitive sequential hypothesis testing: a Bayesian controller maintains a belief over candidate correctness and dynamically decides whether to gather more evidence, refine the candidate, verify it, or stop. Across six generators and nine coding benchmarks, Bayesian control proves to be most valuable when verification is costly and critics are informative but imperfect. Beyond control, the belief state yields an interpretable correctness score that outperforms token-probability and raw tool-success baselines for uncertainty quantification.

13:00 JSTLLM/生成AI

CompressKV: リソース効率の高いロングコンテキスト LLM 推論のためのセマンティック検索ガイドによる KV キャッシュ圧縮

ロングコンテキスト大規模言語モデル (LLM) 推論は、メモリ フットプリントとキーバリュー (KV) キャッシュのデコード コストによってますます制約が増えており、リソースに制約のあるハードウェアでの持続可能な展開が制限されています。既存の KV キャッシュ削除方法は通常、GQA ベースの LLM のすべてのヘッドにヒューリスティック トークン スコアリングを適用します。これらのメソッドはアテンション ヘッドのさまざまな機能を無視するため、クリティカル トークンの削除につながり、LLM のパフォーマンスが低下します。この問題に対処するために、GQA ベースの LLM 用のリソース効率の高い KV キャッシュ圧縮フレームワークである CompressKV を提案します。 CompressKV は、すべてのヘッドからのアテンション スコアを集約するのではなく、プロンプトおよび意味的に重要な中間コンテキスト証拠の最初と最後のトークンの両方をキャプチャするセマンティック検索ヘッド (SRH) を特定し、それらを使用して KV ペアを保持する必要があるトークンを選択します。さらに、CompressKV は、レイヤーごとのエビクション エラーのオフライン推定に従って、レイヤー全体にキャッシュ バジェットを割り当てます。 LongBench と Needle-in-a-Haystack での実験では、CompressKV がメモリ バジェット全体にわたって既存の KV キャッシュ削除方法よりも一貫して優れたパフォーマンスを発揮することが示されています。特に、LongBench の質問応答タスクではわずか 3\% の KV キャッシュを使用してフル キャッシュのパフォーマンスの 97\% 以上を維持し、Needle-in-a-Haystack ではわずか 0.7\% の KV ストレージで 90\% の精度を達成します。これらの結果は、ロングコンテキスト LLM 推論におけるリソースとパフォーマンスのトレードオフが改善されたことを示しています。私たちのコードは、https://github.com/TUDa-HWAI/CompressKV で公開されています。

原文 (English)

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware. Existing KV cache eviction methods typically apply heuristic token scoring over all heads in GQA-based LLMs. These methods ignore the different functionalities of attention heads, leading to the eviction of critical tokens and thus degrading the performance of LLMs. To address this issue, we propose CompressKV, a resource-efficient KV-cache compression framework for GQA-based LLMs. Instead of aggregating attention scores from all heads, CompressKV identifies Semantic Retrieval Heads (SRHs) that capture both the initial and final tokens of a prompt and semantically important mid-context evidence, and uses them to select tokens whose KV pairs should be retained. Furthermore, CompressKV allocates cache budgets across layers according to offline estimates of layer-wise eviction error. Experiments on LongBench and Needle-in-a-Haystack show that CompressKV consistently outperforms existing KV-cache eviction methods across memory budgets. Notably, it preserves over 97\% of full-cache performance using only 3\% of the KV cache on LongBench question-answering tasks and achieves 90\% accuracy with just 0.7\% KV storage on Needle-in-a-Haystack. These results demonstrate an improved resource--performance trade-off for long-context LLM inference. Our code is publicly available at: https://github.com/TUDa-HWAI/CompressKV

13:00 JSTエージェント

Latent Bridge: リアルタイム ゲーム エージェント向けの連続的な低速/高速チャネル

一般的なコンピュータで使用されるリアルタイム エージェント (最も要求の厳しいケースとしてゲーム) は、数秒かけて計画を立てながら、数十ミリ秒以内に動作する必要があります。これら 2 つの体制は、遅延と品質のトレードオフの対極に位置します。推論 VLM (Qwen3-VL-8B-Thinking) は効果的に検討しますが、応答ごとに約 1.5 秒かかります。これは 15 Hz の制御ループとしては遅すぎます。対照的に、リアクティブ VLM (MiniCPM-o 4.5) はミリ秒単位で動作しますが、計画の負荷が高いタスクではパフォーマンスが低下します。スケールが一致した 2 つの凍結モデル (9B リアクティブ、8B 推論) を結合し、通信チャネルを唯一のトレーニング可能なコンポーネントとして残します。標準的な結合はテキスト ブリッジ (T) です。低速モデルがサフィックスを書き込み、高速モデルが読み取ります。学習された連続潜在ブリッジ (L) を導入します。これは、低速モデルの残差を高速モデルの入力埋め込み空間に LLaVA スタイルの方法で投影し、テキストの往復を回避します。両方とも高速専用 (F) と比較されます。 7 つの Atari ゲームとドライビング ドメイン (MetaDrive) では、ホールドアウト シードでチャネルごとにアクション デコーダーを調整し、Latent Bridge はすべてのドメインで Text Bridge と同等かそれを上回ります。2 つのゲームを大幅に改善し (MsPacman +57%、RoadRunner +28%)、他の場所でも安全にドロップインできます。両方のチャネルを組み合わせると破壊的な干渉が発生するため (RoadRunner -96%)、1 つのみを使用する必要があります。この利点は非常に予測可能です。ブリッジは、遅い推論がすでに速い反応を上回っている場合 (T > F) にのみ役立ちます。高速のみに対する潜在ゲインとテキスト ゲインは、r=0.93 で一緒に推移します。 MetaDrive はコントロールされたネガティブであり、Text Bridge が価値を追加しないため、Latent Bridge は明らかに不活性です。リプレイ録画と再現可能なパイプラインをリリースします。

原文 (English)

The Latent Bridge: A Continuous Slow-Fast Channel for Real-Time Game Agents

A real-time agent for general computer use - with games as the most demanding case - must act within tens of milliseconds while still planning over seconds. These two regimes sit at opposite ends of the latency-quality tradeoff. A reasoning VLM (Qwen3-VL-8B-Thinking) deliberates effectively but requires ~1.5 s per response - far too slow for a 15 Hz control loop. In contrast, a reactive VLM (MiniCPM-o 4.5) acts in milliseconds but underperforms on planning-heavy tasks. We couple two frozen models of matched scale (9B reactive, 8B reasoning), leaving the communication channel as the sole trainable component. The standard coupling is a Text Bridge (T): the slow model writes a suffix the fast model reads. We introduce a learned continuous Latent Bridge (L) that projects the slow model's residuals into the fast model's input-embedding space in a LLaVA-style manner, avoiding any text round-trip; both are compared against Fast-Only (F). On 7 Atari games and a driving domain (MetaDrive), tuning the action decoder per channel on held-out seeds, the Latent Bridge matches or beats the Text Bridge in every domain: it significantly improves two games (MsPacman +57%, RoadRunner +28%) and is a safe drop-in elsewhere. Combining both channels interferes destructively (RoadRunner -96%), so only one should be used. The benefit is highly predictable: the bridge helps if and only if slow reasoning already beats fast reaction (T > F) - the Latent and Text gains over Fast-Only move together at r=0.93. MetaDrive is the controlled negative, where the Latent Bridge is demonstrably inert because the Text Bridge adds no value. We release replay recordings and reproducible pipelines.

13:00 JSTLLM/生成AI

大規模言語モデルのスケーリング指数の小ささについて

現在の大規模言語モデル (LLM) アプリケーションのスケーリング指数が、エネルギー資源の観点から持続不可能な状況を示している理由について説明します。さらに、このような指数の小ささを、無限データの限界における損失関数の非ゼロ値の無視による数値バイアス (「ペデスタル効果」) に帰することは、持続不可能性の問題を解決しないことを示します。最後に、スケーリング指数に対するデータの滑らかさ (粗さ) の影響について、流体乱流の現象論的モデルとの類似性に基づいてコメントします。

原文 (English)

On the Smallness of the Large Language Models Scaling Exponents

We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in terms of energy resources. We further show that attributing the smallness of such exponents to a numerical bias due to the neglect of a non-zero value of the loss function in the limit of infinite data (``pedestal effect") does not remove the unsustainability issue. Finally, the effects of the smoothness (roughness) of the data on the scaling exponents is commented upon based on an analogy with phenomenological models of fluid turbulence.

13:00 JSTLLM/生成AIDeepSeek

希少疾患診断を加速するための特殊な推論大規模言語モデル: ランダム化 AI 医師支援試験

希少疾患は世界中で何百万人もの人々に影響を及ぼしていますが、専門的な臨床専門知識が不足しているため、タイムリーな診断が依然として公衆衛生上の大きな課題となっています。大規模言語モデル (LLM) は希少疾患の診断をサポートする可能性を示していますが、現在のモデルは不十分な臨床展開可能性、限られた臨床的根拠のある証拠、およびトレーニング データの不足によって制約を受けています。ここでは、希少疾患診断用のオープンソースのコンパクト推論 LLM (32B パラメーター) である RaDaR (Rare Disaster navigatoR) を紹介します。 RaDaR は、公開されているフリーテキスト ケース 49,170 件と、推論強化トレーニングによる合成ケース 104,666 件を使用してトレーニングされました。 RaDaR は、公開ベンチマークと 4 つの外部検証センターにわたって、671B DeepSeek-R1 を含む評価されたオープンソース モデルの中で最も強力なパフォーマンスを示しました。遡及コホートにおいて、RaDaR は症例の 61.06 パーセントで臨床的疑いが文書化される前に最終診断を優先しました。これは、1.87 か月の潜在的なリードタイムと施設内間隔の 50.18 パーセントに相当します。無作為化された医師支援試験では、RaDaR 支援により、インターネット検索のみと比較して医師の希少疾患診断精度が 21.44 パーセント向上しました。合成データアブレーションは、表現型にアンカーされたナラティブが、テストされたデータ範囲内で単調なスケーリング傾向を持つ、ロングテール希少疾患に対する有用なトレーニングシグナルを提供することを示唆しました。 RaDaR とその開発および検証フレームワークを組み合わせることで、展開可能な希少疾患推論モデルと、データ不足下での診断 AI のための再現可能な開発フレームワークが提供されます。

原文 (English)

A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial

Rare diseases affect millions of individuals worldwide, yet timely diagnosis remains a major public health challenge due to scarcity of specialized clinical expertise. While large language models (LLMs) show promise to support rare disease diagnosis, current models are constrained by insufficient clinical deployability, limited clinically grounded evidence, and scarcity of training data. Here we present RaDaR (Rare Disease navigatoR), an open-source, compact reasoning LLM (32B parameters) for rare disease diagnosis. RaDaR was trained with 49,170 publicly available free-text cases and 104,666 synthetic cases with reasoning-enhanced training. RaDaR showed the strongest performance among evaluated open-source models, including the 671B DeepSeek-R1, across public benchmarks and four external validation centers. In a retrospective cohort, RaDaR prioritized the final diagnosis before documented clinical suspicion in 61.06 percent of cases, corresponding to a potential lead time of 1.87 months and 50.18 percent of the within-center interval. In a randomized physician-assistance trial, RaDaR assistance improved physicians' rare-disease diagnostic accuracy by 21.44 percentage points compared with internet search alone. Synthetic-data ablations suggested that phenotype-anchored narratives provide useful training signal for long-tail rare diseases, with a monotonic scaling trend within the tested data range. Together, RaDaR and its development and validation framework provide a deployable rare-disease reasoning model and a reproducible development framework for diagnostic AI under data scarcity.

13:00 JSTエージェントビジネス/資金調達

自律的な評価を備えたコンピュータ使用エージェントの強化学習

Computer-Use Agent (CUA) は、グラフィカル ユーザー インターフェイス内で直接認識して行動することで、高レベルのユーザー目標を実行します。ただし、オープンエンドのデスクトップ環境ではスケーラブルで機械可読な報酬信号がほとんど提供されないため、CUA の強化学習は依然として困難です。タスクの成功は多くの場合視覚的に根拠があり、手作りの報酬関数や高密度の手動ラベルで指定するのは困難です。我々は、GUI エージェントのスケーラブルな監視信号として自律的な視覚言語評価を使用する RL 微調整フレームワークを提案します。最終的なスクリーンショットと元の指示が与えられると、ビジョン言語モデルはタスクの完了を判断し、ポリシーの最適化中にタスク固有のヒューリスティックや手動ラベルを使用せずに最終的なフィードバックを提供します。自律型評価器は不完全であるため、そのフィードバックをノイズの多いバイナリ報酬チャネルとしてモデル化し、近接ポリシー最適化のためのノイズ補正された報酬推定器を導出します。 macOSWorld、Windows Agent Arena、および OSWorld にわたる実験では、修正された評価者の報酬がゼロショットのベースラインと生の評価者の報酬の両方を上回り、成功率がゼロショットのパフォーマンスより平均 12.6 ポイント、生の評価者の微調整よりも 5.1 ポイント向上したことが示されています。これらの結果は、評価者のノイズが明示的にモデル化され補正されている場合、自律評価が GUI 環境における RL の実用的な報酬信号として機能する可能性があることを示唆しています。

原文 (English)

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is often visually grounded and hard to specify with handcrafted reward functions or dense manual labels. We propose an RL fine-tuning framework that uses autonomous vision-language evaluation as a scalable supervision signal for GUI agents. Given a final screenshot and the original instruction, a Vision-Language Model judges task completion and provides terminal feedback without task-specific heuristics or manual labels during policy optimization. Because autonomous evaluators are imperfect, we model their feedback as a noisy binary reward channel and derive a noise-corrected reward estimator for Proximal Policy Optimization. Experiments across macOSWorld, Windows Agent Arena, and OSWorld show that corrected evaluator rewards outperform both zero-shot baselines and raw evaluator rewards, improving success rates by an average of 12.6 percentage points over zero-shot performance and 5.1 points over raw evaluator fine-tuning. These results suggest that autonomous evaluation can serve as a practical reward signal for RL in GUI environments when evaluator noise is explicitly modeled and corrected.

13:00 JSTLLM/生成AIエージェント研究/論文

マルチエージェント LLM システム用のガバナド共有メモリ

マルチエージェント LLM 環境には、共有ナレッジ管理のための堅牢なメカニズムが必要です。このペーパーでは、フリート メモリの問題を形式化し、不正な漏洩、古い伝播、矛盾の持続、来歴の崩壊という 4 つの基本的な障害モードを特定します。これらに対処するために、スコープ指定された取得、一時的なスーパーセッション、来歴追跡、およびポリシーに基づいたメモリ伝播などの明示的なシステムレベルのプリミティブを定義します。これらのプリミティブは、運用マルチテナント メモリ サービスである MemClaw に実装され、4 つのガバナンス次元をテストする再現可能なハーネスである ArgusFleet によって評価されます。この調査では、ベースラインの比較ではなく、実際の運用サービスを測定し、現実世界のアーキテクチャに関する洞察と否定的な結果を強調しています。主な評価結果 来歴: ホップあたり 1 秒未満のレイテンシーで正しいライター ID を使用して、深さ 4 の派生チェーンを 100% 再構築することに成功しました。伝播: フリート間の漏洩がゼロで、フリート内の高い可視性を実証しました。強力な書き込みモードでは、可視への書き込み遅延が 1 回の検索ラウンドトリップに最適化されました。運用アーキテクチャの問題が発見されました 非対称スコープの強制: テナントの分離は維持されましたが、サブテナント スコープは当初、エージェント スコープの資格情報に対する ID による直接の GET リクエストでバイパスされていました (調査中に開示および修正されました)。パイプライン順序付けの競合: 矛盾スーパーセッションは許可された書き込みに対して機能しますが、同期準重複ゲートは、非同期矛盾検出器が矛盾書き込みを評価する前に、矛盾書き込みを拒否する可能性があります。結論: ロングコンテキストの取得だけでは、本番環境のマルチエージェントメモリには不十分です。管理された共有メモリには明示的なシステムレベルの抽象化が必要であり、設計のみの処理では見逃されていた施行やパイプラインの順序付けの失敗を明らかにするにはライブ評価が不可欠です。

原文 (English)

Governed Shared Memory for Multi-Agent LLM Systems

Multi-agent LLM environments require robust mechanisms for shared knowledge management. This paper formalizes the fleet-memory problem and identifies four foundational failure modes: unauthorized leakage, stale propagation, contradiction persistence, and provenance collapse. To address these, we define explicit systems-level primitives: scoped retrieval, temporal supersession, provenance tracking, and policy-governed memory propagation. These primitives are implemented in MemClaw, a production multi-tenant memory service, and evaluated via ArgusFleet, a reproducible harness testing four governance dimensions. Rather than a baseline comparison, this study measures a live production service, emphasizing real-world architectural insights and negative results. Key Evaluation Results Provenance: Successfully reconstructed 100% of depth-four derivation chains with correct writer identity at sub-second per-hop latency. Propagation: Demonstrated high intra-fleet visibility with zero cross-fleet leakage. Under strong write mode, write-to-visible latency was optimized to a single search round-trip. Production Architectural Issues Discovered Asymmetric Scope Enforcement: Tenant isolation held, but sub-tenant scope was initially bypassed on direct GET-by-id requests for agent-scoped credentials (disclosed and remediated during the study). Pipeline Ordering Conflict: While contradiction supersession works for admitted writes, a synchronous near-duplicate gate can prematurely reject contradictory writes before the asynchronous contradiction detector can evaluate them. Conclusion: Long-context retrieval alone is insufficient for production multi-agent memory. Governed shared memory demands explicit systems-level abstractions, and live evaluation is vital to expose enforcement and pipeline-ordering failures missed by design-only treatments.

13:00 JSTエージェント

GUI と CLI: 画面のみおよびスキルを介したコンピュータ使用エージェントにおける実行のボトルネック

コンピュータ使用エージェントは、グラフィカル インターフェイスまたはプログラム コマンド インターフェイスを通じてソフトウェア タスクを実行できますが、既存の評価では、タスク、初期状態、検証者、許可されるアクションの違いにより、対話形式が混同されています。 18 のアプリケーションと 12 のワークフロー カテゴリにわたる 440 のデスクトップ タスクの一致する実行層ベンチマークを導入します。このベンチマークでは、画面のみの GUI エージェントとスキル媒介の CLI エージェントが、モダリティ ネイティブのアクションに制限されながら、同一の目標、状態、および最終状態の検証子を受け取ります。この制御された設定では、最強の GUI エージェントは 59.1% の完全合格率に達し、最強のオリジナル スキル CLI エージェントの 48.2% を上回ります。ただし、検証者主導のスキル強化により CLI の成功率は 69.3% に上昇し、CLI の不足の多くはモデルの機能だけではなく、スキルのカバー範囲が不完全であることに起因していることがわかります。これらの結果は、GUI と CLI が異なる実行ボトルネックを露呈していることを示唆しています。GUI エージェントは長期ワークフローにわたる信頼性の高い根拠のある対話によって制限されるのに対し、CLI エージェントはスキル インターフェイスの適用範囲とスケーラビリティによって制限されます。

原文 (English)

GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents

Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with differences in tasks, initial states, verifiers, and permitted actions. We introduce a matched execution-layer benchmark of 440 desktop tasks across 18 applications and 12 workflow categories, where screen-only GUI agents and skill-mediated CLI agents receive identical goals, states, and final-state verifiers while being restricted to modality-native actions. In this controlled setting, the strongest GUI agent reaches a 59.1% full pass rate, outperforming the strongest original-skill CLI agent at 48.2%; however, verifier-guided skill augmentation raises CLI success to 69.3%, showing that much of the CLI deficit comes from incomplete skill coverage rather than model capability alone. These results suggest that GUI and CLI expose different execution bottlenecks: GUI agents are limited by reliable grounded interaction over long-horizon workflows, whereas CLI agents are limited by the coverage and scalability of their skill interfaces.

13:00 JST研究/論文

クオンツ・コンバージェンス:体系的な株式選択のための古典的なバリュー投資と最新のファクター・モデルの橋渡し

現代の金融は、株式市場のパターンを見つけるために複雑な機械学習モデルに大きく依存しています。しかし、これらの AI モデルが複雑になるにつれて、実際の永続的な価値を持つ企業を見つけるのではなく、短期的な市場のノイズを記憶することがよくあります。私たちは、ベンジャミン・グレアムの古典的なバリュー投資ルールが、これらの最新モデルを抑制する数学的な「ローパスフィルター」として機能するかどうかをテストするためにこの研究を設計しました。私たちは、純粋なグラハム ルール、最新の市場要因、その両方の組み合わせという 3 つの異なる機能セットを構築し、20 年間の S&P 500 データを使用して非常に複雑なモデル (XGBoost および AutoGluon) に対してテストしました。 4 年間のテスト期間 (2022 年 3 月から 2026 年 3 月まで) にわたって厳密なバイアンドホールド戦略を適用した結果、より複雑なアルゴリズムが必ずしも勝利するとは限らないことがわかりました。 AutoGluon モデルは高いリターン (222.68%) を獲得しましたが、市場が暴落する直前に不安定なハイテク株を購入したため、39.78% という大幅な下落に見舞われました。一方、純粋な Graham Random Forest は、はるかに少ないリスク (1.38 Calmar Ratio) で最高の全体収益 (232.13%) を達成しました。さらに、結合ランダム フォレストは勢いとグラハム ルールをうまく組み合わせることで、テストしたモデルの中で最も低い最大ドロップ (34.53%) を維持しながら、202.91% のリターンを実現しました。結局のところ、この研究は、グレアムの「安全域」が時代遅れではないことを証明しています。これは実際、現代の AI が過度のリスクを負うことを防ぐ非常に効果的な方法です。

原文 (English)

Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection

Modern finance relies heavily on complex machine learning models to find patterns in the stock market. However, as these AI models get more complicated, they often memorize short-term market noise instead of finding companies with real, lasting value. We designed this research to test if Benjamin Graham's classic value investing rules could act as a mathematical "low-pass filter" to keep these modern models in check. We built three different sets of features - pure Graham rules, modern market factors, and a mix of both - and tested them against highly complex models (XGBoost and AutoGluon) using 20 years of S&P 500 data. By applying a strict buy-and-hold strategy over a four-year test period (March 2022 to March 2026), the results showed that more complex algorithms do not always win. While the AutoGluon model captured high returns (222.68%), it suffered a substantial 39.78% drop because it bought volatile tech stocks right before the market crashed. On the other hand, the pure Graham Random Forest achieved the highest overall return (232.13%) with much less risk (1.38 Calmar Ratio). Furthermore, the Combined Random Forest successfully mixed momentum with Graham's rules, making a 202.91% return while keeping the lowest maximum drop (34.53%) of any model tested. Ultimately, this research proves that Graham's "margin of safety" isn't outdated; it is actually a highly effective way to prevent modern AI from taking on too much risk.

13:00 JSTLLM/生成AI

LLM が法的コンテキスト オブジェクトの入力を求められる 詳細: 刑事法的コンテキストにおける小規模なオンプレミス LLM からの過剰拒否

法的文脈における LLM の使用の妥当性については依然として倫理的および法的な議論の対象となっていますが、法律専門家はすでに、翻訳と再定式化のみを目的として、個人用 LLM を実験しています。ただし、そのような一見無害な使用法でも、LLM アシスタントが特定のトピックに関する支援を選択的に拒否した場合、ケースの処理速度によってバイアスが生じる可能性があります。このようなバイアスをより適切に予測するために、デバイス上のアシスタントとして使用される可能性が最も高いいくつかの最新の小型 LLM を調査し、過剰な拒否が法的要求に与える影響を評価します。驚くべきことに、権威スタイルの接頭辞(「あなたは国家最高裁判所の助手を務めています」、「[...]弁護人」)は、接頭辞なしのベースラインに比べて拒否率を体系的に2〜20倍増加させることがわかりました。一方、既知のロールプレイジェイルブレイク接頭辞は複合的な効果を示し、一部のモデルでは拒否が急激に増加し、他のモデルでは拒否がほとんど変化しません。この発見は、オンプレミスで展開可能な小規模な LLM は、実際の機関ユーザーが自然に導入する可能性のあるコンテキスト フレームワークの下では不安定であることを示唆しており、バイアスの機会を最小限に抑えるためにはさらなる調査が不可欠です。

原文 (English)

LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context

While the validity of LLMs' use in the legal context remains subject to ethical and legal debate, legal professionals are already experimenting with personal LLMs, if only for translation and reformulation. However, even such a seemingly innocuous use can introduce biases through case processing speed if LLM assistants selectively refuse assistance on certain topics. To better anticipate such biases, we investigate several modern small LLMs that are most likely to be used as on-device assistants, to assess the impact of overrefusal on legal prompts. Surprisingly, we find that authority-style prefixes (``you are acting as an assistant of the national supreme court'', ``[...] defense lawyer'') systematically increase refusal rates by 2--20x over the no-prefix baseline, while a known role-play jailbreak prefix shows mixed effects, sharply increasing refusals in some models and barely shifting them in others. The finding suggests that small on-prem deployable LLMs are unstable under contextual framings that a real institutional user might naturally introduce, and further investigation is essential to minimize opportunities for bias.

13:00 JSTLLM/生成AIビジネス/資金調達Llama

AdversaBench: 複数の裁判官による確認とモデル間の移行性を備えた自動化された LLM レッドチーム化

大規模な言語モデルの敵対的評価をスケーリングするには、ハード入力を生成する方法と、結果として生じる失敗が本物であることを確認する信頼性の高い方法の両方が必要です。 AdversaBench は、5 つの構造化された演算子でシード プロンプトを変更し、ターゲット モデルをクエリし、メタ ジャッジ タイブレーカーを備えた 3 人のジャッジ パネルを通じて失敗を確認する、エンドツーエンドのレッド チーム パイプラインです。推論、指示に従い、ツールの使用という 3 つのカテゴリにわたる 45 のシードに関する実験を報告します。すべてのシードで失敗が確認されました。 4 つの発見が際立っています。まず、オペレーターの有効性はカテゴリによって大きく異なります。inject_distractor のスコアは、指示に従うシードでは 0.00 の平均報酬ですが、推論とツールの使用では 0.80 ~ 0.83 です。第 2 に、バイナリの失敗率が難しさを隠しています。命令に従うシードでは、攻撃者の反復回数が平均 2.4 回であるのに対し、他のカテゴリでは 1.1 回であり、生存曲線にギャップが見られます。第三に、80 ~ 87% というペアごとのジャッジの一致は、ラベルの歪みによりほぼゼロのコーエンのカッパと共存します。カテゴリレベルの不一致率の方が有益です。第 4 に、Llama 3.1 8B に対して生成された敵対的プロンプトはゼロショットを Llama 3.3 70B に転送します。これは、変異がモデル固有の弱点ではなく一般的な動作パターンを悪用していることを示唆しています。コード、データセット、分析スクリプトは https://github.com/khanak0509/AdversaBench で入手できます。

原文 (English)

AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability

Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real. We present AdversaBench, an end-to-end red-teaming pipeline that mutates seed prompts with five structured operators, queries a target model, and confirms failures through a three-judge panel with a meta-judge tiebreaker. We report experiments on 45 seeds across three categories: reasoning, instruction-following, and tool use. Every seed produced a confirmed failure. Four findings stand out. First, operator effectiveness varies sharply by category: inject_distractor scores 0.00 mean reward on instruction-following seeds but 0.80-0.83 on reasoning and tool-use. Second, binary failure rate hides difficulty: instruction-following seeds required 2.4 attacker iterations on average versus 1.1 for other categories, a gap visible in survival curves. Third, pairwise judge agreement of 80-87% coexists with near-zero Cohen's kappa due to label skew; category-level disagreement rates are more informative. Fourth, adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B, suggesting the mutations exploit general behavioral patterns rather than model-specific weaknesses. Code, dataset, and analysis scripts are available at https://github.com/khanak0509/AdversaBench .

13:00 JSTエージェント

ASALT: マルチエージェント強化学習における横方向伝達のための適応的状態調整

マルチエージェント強化学習 (MARL) は、協調的、競争的、または混合の目的を追求する複数のエージェントをトレーニングするという問題に対処します。これまでの研究では、MARL のソース ドメインとターゲット ドメイン間の転移学習を調査しました。ただし、既存のアプローチの大部分は、観測空間とグローバル状態空間の次元がドメイン全体で同一でなければならないという制約を課します。この論文では、ソース ドメインとターゲット ドメイン間の状態空間次元の不一致に明示的に対応する方法を紹介します。提案されたアプローチである ASALT には、ターゲット ドメインの観察とグローバルな状態を共有埋め込み空間にマッピングする観察レベルと状態レベルのアダプターの両方が組み込まれており、それによってアクターと批評家の両方にわたるより効果的な知識の伝達が可能になります。これらのアダプターは、異種ドメイン間での効率的な戦略転送をサポートするエンベディングを生成できます。標準ベンチマーク環境での複数の構成での実験結果は、ASALT がサンプル効率と協調設定でのグローバルな収益の点で既存のベースラインを上回っていることを示していますが、その有効性はソース ドメインとターゲット ドメイン間の不一致の程度に依存します。さらに、我々の調査結果は、ASALT がネガティブな移転を軽減することを示しています。ネガティブな移転は、異なる観察空間と行動空間を持つドメイン間でポリシーを移転する際に大きな障害となることがよくあります。

原文 (English)

ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning

Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives. Prior work has investigated transfer learning between source and target domains in MARL; however, the majority of existing approaches impose the constraint that the dimensionalities of the observation space and the global state space must be identical across domains. In this paper, we introduce a method that explicitly accommodates mismatched state-space dimensionalities between source and target domains. The proposed approach, ASALT, incorporates both observation-level and state-level adapters that map the target-domain observations and global states into a shared embedding space, thereby enabling more effective transfer of knowledge across both actors and critics. These adapters can generate embeddings that support efficient strategy transfer across heterogeneous domains. Experimental results on multiple configurations in standard benchmark environments demonstrate that ASALT surpasses existing baselines in terms of sample efficiency and global return in cooperative settings, but its effectiveness depends on the degree of mismatch between source and target domains. Furthermore, our findings indicate that ASALT mitigates negative transfer, which frequently constitutes a major obstacle when transferring policies between domains with differing observation and action spaces.

13:00 JST研究/論文

深層学習を使用した、不確実性を認識したアルツハイマー病進行の長期的予測

アルツハイマー病の進行の縦断的モデリングは、最も可能性の高い次の診断だけでなく、時間の経過とともに患者がどのように変化するか、そしてその予測がどれほど信頼できるかを説明できる場合にのみ臨床的に役立ちます。ほとんどの深層学習アプローチは、この問題を 1 段階の分類に落とし込み、正常な認知機能、軽度の認知障害、認知症をフラットなカテゴリーとして扱いますが、今後の訪問で不確実性がどのように蓄積するかについての洞察は限定的です。我々は、通常の診断予測、マルチ水平軌道生成、および分解された不確実性推定を組み合わせた確率的フレームワークを提案します。 Temporal Fusion Transformer エンコーダには、CORAL 順序出力層、非対称損失重み付け、およびコンバーターのオーバーサンプリングが適用され、疾患段階の順序を尊重し、MCI から認知症への移行に対する感度が向上します。学習された患者コンテキスト表現に基づいて、自己回帰混合密度ネットワークは、診断状態、CDR ボックス和、MMSE 方向、海馬体積の 5 年間の確率的軌跡を生成します。 ADNI では、このモデルは次回受診診断予測において線形ベースライン、再発ベースライン、およびトランスフォーマー ベースラインを上回り、MCI と認知症の区別において最も優れた効果を発揮します。生成された軌跡は、名目上の 90% に近い信頼区間の範囲を達成し、予測範囲全体にわたって不確実性が拡大し、予想されるアルツハイマー病の進行と一致するバイオマーカーのダイナミクスを実現します。さらに、解析的混合分散と、最も強力なエンコーダーの多様性と出力レベルの認識論的信号を提供する 5 メンバーのブートストラップ アンサンブルを使用して、認識論的不確実性から偶然の不確定性を分離します。認識論的不確実性は、まれな進行原型、MCI、認知症患者で高く、OASIS-3 の外部評価下では、予測誤差とともに増加します。

原文 (English)

Uncertainty-Aware Longitudinal Forecasting of Alzheimer's Disease Progression Using Deep Learning

Longitudinal modelling of Alzheimer's disease progression is clinically useful only if it can describe not just the most likely next diagnosis, but how a patient may evolve over time and how reliable that forecast is. Most deep learning approaches reduce this problem to single-step classification, treating cognitively normal, mild cognitive impairment, and dementia as flat categories while providing limited insight into how uncertainty accumulates across future visits. We propose a probabilistic framework that combines ordinal diagnosis prediction, multi-horizon trajectory generation, and decomposed uncertainty estimation. A Temporal Fusion Transformer encoder is adapted with a CORAL ordinal output layer, asymmetric loss weighting, and converter oversampling to respect disease-stage ordering and improve sensitivity to MCI-to-dementia transitions. Conditioned on the learned patient-context representation, an autoregressive Mixture Density Network generates five-year probabilistic trajectories for diagnosis state, CDR Sum of Boxes, MMSE orientation, and hippocampal volume. On ADNI, the model outperforms linear, recurrent, and transformer baselines for next-visit diagnosis prediction, with the strongest gains on MCI-versus-dementia discrimination. Generated trajectories achieve near-nominal 90% credible interval coverage, widening uncertainty across the forecast horizon, and biomarker dynamics consistent with expected Alzheimer's disease progression. We further separate aleatoric from epistemic uncertainty using analytic mixture variance and a five-member bootstrap ensemble, which provides the strongest encoder diversity and output-level epistemic signal. Epistemic uncertainty is higher for rare progression archetypes, MCI and dementia patients, and under external evaluation on OASIS-3, where it increases alongside prediction error.

13:00 JSTLLM/生成AI

ScaleToT: 数十億規模の低アクティビティ ユーザー モデリングのための構造化 LLM 推論の一般化

正確なユーザー モデリングは、多くの場合、豊富なインタラクション履歴に依存しますが、何十億人ものアクティビティの低いユーザーにとっては、これらの履歴は利用できません。大規模言語モデル (LLM) は静的プロファイルから潜在的なユーザー状態を推測できますが、プロファイルがまばらな場合、この推論は信頼できなくなり、LLM を数十億のユーザーに適用すると法外な費用がかかります。私たちは、LLM で処理された小さなサブセットから構造化推論を学習し、それを広範な低アクティビティ ユーザー集団に拡張する ScaleToT を紹介します。推論の信頼性を向上させるために、ScaleToT は、有界エントロピーに基づく思考ツリー (ToT) 改良手順を使用して型付きユーザー状態チェーンを構築します。この構造化された推論を疎なプロファイルから使用できるようにするために、教師が厳選したチェーンを使用して、教師あり微調整 (SFT) と結果主導型セグメント認識暗黙的報酬ポリシー最適化 (OSIPO) を通じて静的プロファイルで生徒モデルをトレーニングします。次に、ScaleToT は学生の推論表現を軽量プロファイル エンコーダに転送し、LLM 推論を行わずに残りのユーザーに共有推論信号を提供します。 10億規模の広告展開における生涯価値(LTV)予測に基づいてScaleToTを評価します。ランダム化されたオンライン A/B テストでは LT30 が 6.738\% 増加しましたが、オフライン推論では潜在的な母集団の 7.32\% しかカバーされず、全母集団推論と比較してコンピューティング コストが大幅に削減されました。

原文 (English)

ScaleToT: Generalizing Structured LLM Reasoning for Billion-Scale Low-Activity User Modeling

Accurate user modeling often depends on rich interaction histories, which are unavailable for billions of low-activity users. Large Language Models (LLMs) can infer latent user states from static profiles, but this reasoning becomes unreliable when profiles are sparse, and applying an LLM to billions of users is prohibitively expensive. We present ScaleToT, which learns structured reasoning from a small LLM-processed subset and extends it to the broader low-activity user population. To improve reasoning reliability, ScaleToT constructs typed user-state chains with a bounded entropy-guided Tree-of-Thought (ToT) refinement procedure. To make this structured reasoning usable from sparse profiles, the teacher-curated chains are used to train a student model on static profiles through supervised fine-tuning (SFT) and Outcome-Driven Segment-Aware Implicit Reward Policy Optimization (OSIPO). ScaleToT then transfers the student's reasoning representations to a lightweight profile encoder, providing shared reasoning signals for the remaining users without LLM inference. We evaluate ScaleToT on lifetime value (LTV) prediction in a billion-scale advertising deployment. A randomized online A/B test increased LT30 by 6.738\%, while offline reasoning covered only 7.32\% of the potential population, greatly reducing compute cost compared with full-population reasoning.

13:00 JST研究/論文

AI トークンノミクス: 基盤モデルにおけるトークン、計算、価格設定の経済学

トークンは、情報処理、計算、メモリ使用、エネルギー消費、価格設定、経済的価値を結び付ける、現代の基盤モデル サービスの実際的な会計単位となっています。このペーパーでは、AI トークンノミクスのフレームワークを開発します。これは、AI システム全体でトークンがどのように生成、消費、価格設定、割り当て、最適化されるかについての研究です。私たちは、トークンレベルの技術コストを、ワークフローレベルの生産機能、企業リソースの割り当て、測定および計測方法、新興市場の設計上の疑問に結び付けます。このフレームワークは、トークンの支出と経済的価値が異なることを示しています。価値は限界生産性、ワークフローの位置、隠れた推論活動、リスク、および下流の伝播効果に依存します。この論文は、隠れたトークンの測定、経験的校正、トークンの生産性、ダイナミックな割り当て、およびトークンベースの市場におけるオープンな研究の方向性を特定して締めくくられています。

原文 (English)

AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation Models

Tokens have become the practical accounting unit for modern foundation model services, linking information processing, computation, memory use, energy expenditure, pricing, and economic value. This paper develops a framework for AI tokenomics: the study of how tokens are generated, consumed, priced, allocated, and optimized across AI systems. We connect token-level technical costs to workflow-level production functions, enterprise resource allocation, measurement and instrumentation methods, and emerging market-design questions. The framework shows that token expenditure and economic value are distinct: value depends on marginal productivity, workflow position, hidden reasoning activity, risk, and downstream propagation effects. The paper concludes by identifying open research directions in hidden-token measurement, empirical calibration, token productivity, dynamic allocation, and token-based markets.

13:00 JST研究/論文

オントロジーベースのデータアクセスにおけるクエリの抽象化

オントロジーベースのデータアクセス (OBDA) では、オントロジーへのマッピングを介して複数のデータソースが統合されます。存在ルールと特定の回答セマンティクスに基づいて OBDA 設定を検討します。私たちは、データクエリをオントロジー層に変換することによってデータクエリを抽象化することで構成されるクエリ抽象化の最近の問題に取り組みます。完全な抽象化は存在しない可能性があるため、最小限に完全で最大限に健全な抽象化の概念が導入されています。私たちは、限定された形式の不等式とデータベース定数をマークする特別な述語を使用した UCQ の拡張内の抽象化を研究します。この拡張は対象となる問題の複雑さの増加にはつながりませんが、最小限に完全な抽象化を表現できるため、存在する場合には完全な抽象化を表現できます。また、データ交換から生じる最大の回復の概念と新たな関係を作ることによって、最大限に健全な抽象化を特徴づけます。

原文 (English)

Abstractions of Queries in Ontology-Based Data Access

In ontology-based data access (OBDA), multiple data sources are integrated via mappings to an ontology. We consider an OBDA setting based on existential rules and the certain answer semantics. We address the recent issue of query abstraction, which consists of abstracting data queries by translating them to the ontology layer. Since a perfect abstraction may not exist, the notions of minimally complete and maximally sound abstractions have been introduced. We study abstractions within an extension of UCQs with a limited form of inequality and a special predicate marking database constants. While this extension does not lead to an increased complexity of the problems of interest, it is able to express minimally complete abstractions, hence perfect abstractions when they exist. We also characterize maximally sound abstractions by making a new connection with the notion of maximum recovery stemming from data exchange.

13:00 JST研究/論文

CQ が失敗した場合: OE-Assist を使用した CQ 検証の課題

コンピテンシー質問 (CQ) は、CQ 検証の中心的なコンポーネントであり、オントロジーが一連の自然言語質問に対して評価され、オントロジーの意図された目的が適切にモデル化されているかどうかを判断する確立されたプロセスです。ただし、CQ 検証には、言語のニュアンスを注意深く解釈し、正式なオントロジー構造と正確に一致させる必要があるため、多くの場合時間がかかり、エラーが発生しやすくなります。 CQ のあいまいさと複雑さによってこのプロセスがさらに複雑になる可能性があり、モデリングの決定と検証の結果に一貫性がなくなる可能性があります。このペーパーでは、CQ を困難にしている原因と、CQ 検証プロセスにおけるユーザーのパフォーマンスを向上させるための可能な解決策を調査します。オントロジー評価をサポートする LLM アシスタントを使用して、20 のタスクに対して CQ 検証を実行した 19 人の参加者のデータを実験しました。この結果は、オントロジー エンジニアリング プロセスの後の段階でのあいまいさや過度の複雑さを回避するために、CQ を公開する前に CQ を改良するツールの必要性を示しています。

原文 (English)

When CQs Go Wrong: Challenges in CQ Verification with OE-Assist

Competency Questions (CQs) are the central component of CQ-verification, an established process in which an ontology is evaluated against a set of natural language questions to determine whether the intended purpose of the ontology has been properly modelled. However, CQ-verification is often time-consuming and error-prone, as it requires careful interpretation of linguistic nuances and precise alignment with formal ontology constructs. Ambiguities and complexity in CQs can further complicate this process, leading to inconsistent modelling decisions and verification outcomes. In this paper, we investigate what makes a CQ challenging and possible solutions to enhance the users' performance in the CQ-verification process. We experimented with the data of 19 participants who performed CQ-verification on 20 tasks using an LLM assistant to support ontology evaluation. The results show the necessity of a tool to refine CQs before publishing them to avoid ambiguity or excessive complexity in later phases of the ontology engineering process.

13:00 JST研究/論文

Themis: 人間のフィードバックによる強化学習のための説明可能な AI 対応フレームワーク

強化学習 (RL) システムを安全にトレーニングすることは本質的に困難であり、望ましくない動作を回避する保証はありません。これに対する最も効果的な防御策は、(i) 説明可能性による透明性と、(ii) 人間のフィードバックによる調整です。どちらも有望な結果を示していますが、現在、それらを組み合わせた公的に利用可能なフレームワークはありません。これに対処するために、ヒューマン フィードバックからの強化学習のための XAI 対応のテストおよび評価フレームワークである Themis を紹介します。 Themis は 200 以上の広く使用されている環境をサポートしており、RL、透明性、および位置合わせの実験用に簡単に構成できます。私たちの結果は、Themis が人間の好みを使用して、環境の真の報酬シグナルと一致またはそれを上回る報酬モデルをトレーニングできることを示しています。また、人間からのフィードバックを収集し、実験を管理するためのクラウドベースのプラットフォームも提供しています。ユーザーフレンドリーで自動スケーラブルで、追加の開発オーバーヘッドなしで複数の実験にわたる大規模な参加者グループをサポートします。テストでは、Themis が小規模な商用マシンでの連続実験で 1,000 人のユーザーをサポートできることが示されています。

原文 (English)

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback. While both show promising results, no publicly available framework currently combines them. To address this, we introduce Themis, an XAI-enabled testing and evaluation framework for Reinforcement Learning from Human Feedback. Themis supports over 200 widely used environments and is easily configurable for experiments in RL, transparency, and alignment. Our results show that Themis can train reward models that match or outperform the environment's true reward signal using human preferences. We also provide a cloud-based platform for collecting human feedback and managing experiments. It is user-friendly, auto-scalable, and supports large participant groups across multiple experiments without extra development overhead. Tests show Themis can support one thousand users in back-to-back experiments on a modest commercial machine.

13:00 JSTエージェント

SAFARI: 積極的な調査によるロングホライズンのエージェント障害の特定の拡張

自律エージェントがますます複雑なマルチステップ、マルチエージェント タスクに取り組むにつれて、その実行軌跡は最大のコンテキスト ウィンドウの制約さえも超えて拡大しています。エージェントの障害を効果的に診断する現在の方法では、LLM のコンテキスト ウィンドウに完全な軌跡が読み込まれますが、注意力が低下し、エージェントのトレースが必然的にコンテキストの制限を超えると失敗します。これに対処するために、線形コンテキストの読み込みをツールで拡張された診断ループに置き換えるフレームワークである SAFARI (Scaling long-horizo​​n Agentic Fault AttRibution via active Investigation) を導入します。 LLM に、クロスターン推論のための永続的短期メモリ (STM) とともに軌道セグメントを読み取って検索するための専用ツールボックスを装備することで、SAFARI は診断精度をアーキテクチャ上のコンテキスト制限から効果的に切り離します。私たちの実験では、SAFARI が 100 万トークン予算内の Who&When データセットで 20%、25,000 トークン予算内の TRAIL GAIA サブセットで 19% 最先端の結果を上回っていることが実証されました。最も重要なのは、ターゲット障害がモデルのネイティブ コンテキスト ウィンドウの 5 倍を超えて存在する場合でも、SAFARI は 0.58 の精度を維持します。これは、従来の評価器が完全に失敗するシナリオです。

原文 (English)

SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation

As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows. Current methods for effectively diagnosing agent failures load the full trajectory into an LLM's context window, which suffers from attention dilution and fails when agentic traces inevitably exceed context limits. To address this, we introduce SAFARI (Scaling long-horizon Agentic Fault AttRibution via active Investigation), a framework that replaces linear context loading with a tool-augmented diagnostic loop. By equipping LLMs with a specialized toolbox to read and search trajectory segments alongside a persistent Short-Term Memory (STM) for cross-turn reasoning, SAFARI effectively decouples diagnostic accuracy from architectural context limits. Our experiments demonstrate that SAFARI outperforms state-of-the-art results by 20% on the Who&When dataset within a 1M token budget, and by 19% on TRAIL GAIA subset on a 25K token budget. Most significantly, SAFARI maintains a 0.58 precision even when the target fault resides 5x beyond the model's native context window, a scenario where traditional evaluators fail entirely.

13:00 JST研究/論文

CineCap: 映画ビデオ キャプション用の時空間アンカーを使用した構造化推論

映画のキャプションは、カメラの動き、ショットのサイズ、被写界深度、構成、撮影角度などの専門的な映画言語の概念を使用して、ビデオがどのように撮影されるかを説明することを目的としています。この機能は、きめ細かいビデオの理解と制御可能な映画品質のビデオ生成にとって重要ですが、既存のマルチモーダル大規模言語モデルではまだ十分に検討されていません。映画の理解に関する質問応答ベースの評価とは異なり、映画のキャプションには、複数の映画の側面にわたって統一されたオープン形式の説明が必要です。このタスクは 2 つの主な理由から困難です。1 つは、モデルが微妙な視覚的証拠からプロの映画のコンセプトを推測する必要があること、もう 1 つは包括的かつ正確なキャプションを生成する必要があることです。したがって、私たちは、時空間アンカーを備えた構造化推論と、包括性、精度、ゲート制御されたカバレッジ報酬を備えた強化学習を組み合わせたフレームワークである CineCap を提案します。前者は、明確な視覚的証拠に基づいてプロの映画の説明を根拠にし、教師付き微調整のためにコンパクトな原子的推論に編成します。一方、後者は、説明の完全性と事実の正確さの間のバランスを改善します。さらに、体系的な評価のために手動で注釈を付けた 472 個のビデオとキャプションのペアのベンチマークである CineCap Bench を構築します。広範な実験により、CineCap が独自の強力なオープンソースベースラインを常に上回り、映画のキャプションの新しい最先端技術を確立していることが示されています。コード、モデル チェックポイント、ベンチマークは、https://github.com/Hectormxy/CineCap.git で公開されています。

原文 (English)

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning

Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is important for fine-grained video understanding and controllable movie-quality video generation, yet remains underexplored in existing multimodal large language models. Unlike question-answering-based evaluation of cinematic understanding, cinematographic captioning requires a unified open-form description over multiple cinematographic dimensions. This task is challenging for two main reasons: the model must infer professional cinematographic concepts from subtle visual evidence, and it must generate captions that are both comprehensive and accurate. Accordingly, we propose CineCap, a framework that combines structured reasoning with spatio-temporal anchors and reinforcement learning with comprehensiveness, accuracy, and gated coverage rewards. The former grounds professional cinematographic descriptions in explicit visual evidence and organizes them into compact atomic reasoning for supervised fine-tuning, while the latter improves the balance between descriptive completeness and factual correctness. In addition, we construct CineCap Bench, a benchmark of 472 manually annotated video-caption pairs for systematic evaluation. Extensive experiments show that CineCap consistently outperforms strong proprietary and open-source baselines, establishing a new state of the art for cinematographic captioning. The code, model checkpoint, and benchmark are publicly available in https://github.com/Hectormxy/CineCap.git.

13:00 JSTLLM/生成AI

LaGO: オンライン強化学習のための潜在アクション ガイダンス

大規模言語モデル (LLM) は計画と逐次的な意思決定に強力な可能性を示していますが、これまでの研究では多くの場合、LLM を直接コントローラーとして使用することに依存しており、これには正確なアクション生成が必要であり、実際には信頼できない可能性があります。この論文では、オンライン強化学習のための潜在アクション ガイダンス (LaGO) を提案します。これは、LLM を明示的なプランナーまたはコントローラーとして扱うのではなく、オンライン ポリシーの最適化をソフトにガイドする前に、事前トレーニング済み LLM を潜在アクションとして使用するフレームワークです。離散制御ベンチマークである CLEVR-Robot と連続制御ベンチマークである Meta-World の両方での実験では、LaGO がバニラ PPO よりも報酬と成功率の両方を一貫して向上させていることが実証されています。特に、LaGO は、CLEVR-Robot では平均成功率が 15.1% から 27.2% に、Meta-World では 2.7% から 15.2% に増加します。さらに、私たちの分析では、より強力な事前トレーニングされた LLM がより効果的なガイダンスを提供することが示されており、LLM の知識が計画とオンラインの意思決定を改善できることを示唆しています。

原文 (English)

LaGO: Latent Action Guidance for Online Reinforcement Learning

Large language models (LLMs) have shown strong potential for planning and sequential decision-making, but prior work often relies on using them as direct controllers, which requires precise action generation and can be unreliable in practice. This paper proposes Latent Action Guidance for Online Reinforcement Learning (LaGO), a framework that uses a pretrained LLM as a latent action prior to softly guide online policy optimization, rather than treating the LLM as an explicit planner or controller. Experiments on both a discrete-control benchmark, CLEVR-Robot, and a continuous-control benchmark, Meta-World, demonstrate that LaGO consistently improves both reward and success rate over Vanilla PPO. In particular, LaGO increases the average success rate from 15.1% to 27.2% on CLEVR-Robot and from 2.7% to 15.2% on Meta-World. Our analysis further shows that stronger pretrained LLMs provide more effective guidance, suggesting that LLM knowledge can improve planning and online decision-making.

13:00 JSTビジネス/資金調達

確率的ブール関数評価のためのコスト最適決定図

多くの意思決定シナリオでは、情報の取得にさまざまなコストがかかります。変動コストおよび真理の割り当てに対する確率分布の下で命題式を評価する際の期待コストを最小限に抑える決定論的評価戦略を構築する問題を検討します。変数選択ヒューリスティック、枝刈り、およびキャッシュを備えた分岐限定アルゴリズムを紹介します。私たちの知る限り、これはこのレベルの一般性を実現する最初の実用的な正確なアルゴリズムです。ランダム インスタンスの実験では、スケーラビリティを実証し、貪欲なビーム検索バリアントの効率と品質のトレードオフを定量化します。さらに、構造化された心臓病の診断例も評価します。最後に、問題が $\#P$ 困難であり、$\mathrm{PSPACE}$ に含まれていることを証明します。

原文 (English)

Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation

In many decision-making scenarios, acquiring information incurs different costs. We consider the problem of constructing a deterministic evaluation strategy that minimizes the expected cost of evaluating a propositional formula under variable costs and a probability distribution over truth assignments. We present a branch-and-bound algorithm with variable-selection heuristics, pruning, and caching. To the best of our knowledge, it is the first practical exact algorithm for this level of generality. Experiments on random instances demonstrate scalability and quantify the efficiency-quality trade-off of a greedy beam-search variant. We additionally evaluate a structured heart-disease diagnosis instance. Finally, we prove that the problem is $\#P$-hard and contained in $\mathrm{PSPACE}$.

13:00 JST研究/論文

BlockTrain を使用した分散型 AI トレーニングと推論

フロンティア AI トレーニングは、集中管理された高密度のアクセラレータ クラスターへのアクセスによってますます形作られています。これにより、ハイパースケーラーや大規模な集中研究所にとって構造的な利点が生まれ、オープンまたは独立した AI の取り組みが、希少な資本、特権的なインフラストラクチャ、およびデータセンターの地理に依存することになります。我々は、モデルが独立してトレーニング可能なブロックに分割され、それぞれが同じグローバル ターゲットから派生したローカル目標に基づいて最適化され、推論時に 1 つのモデルに構成される分散トレーニング プロトコルである Spheroid BlockTrain を紹介します。バイトレベルの WikiText では、同じセットアップのエンドツーエンド Transformer 参照の約 0.04 CE 以内で、BlockTrain はクロス エントロピー 1.359 (複雑度 3.89) に達しますが、アクティブな各ワーカーは 1 つのブロックのみをトレーニングし、フルモデル オプティマイザー状態を回避します。共有 6 ワーカー ブロックのトレーニング実行は、同じブロックの更新を 1 つの組み立てられたモデルに平均化することで CE 1.385 に達します。 HTTP/TCP トランスポート実験では、実際のシリアル化されたチェックポイントと更新を移動します。これには、15.22 GB を移動しながら CE を 5.580 から 1.811 に改善するパブリック IP 3 ホストの実行が含まれます。推論のために、現在の BlockTrain パスは完全な出力ごとに 1 つのブロック スタック トラバーサルを使用し、最大 75.80B パラメーターの論理 fp16 シェイプまでの 3 つのパブリック ネットワーク GPU ホストにわたる直接 TCP を介してサービスを提供します。これは、トラバーサルごとに 1 つのトークンではなく、WAN パイプライン トラバーサルごとに完全なシーケンスを出力するため、一致するプレーン自己回帰 TCP パイプライン ベースラインよりも優れたパフォーマンスを発揮します。

原文 (English)

Decentralised AI Training and Inference with BlockTrain

Frontier AI training is increasingly shaped by access to dense, centrally controlled accelerator clusters. This creates a structural advantage for hyperscalers and large centralized laboratories, and makes open or independent AI efforts depend on scarce capital, privileged infrastructure, and data-center geography. We present Spheroid BlockTrain, a decentralized training protocol in which a model is partitioned into independently trainable blocks, each optimized on a local objective derived from the same global target and composed at inference into one model. On byte-level WikiText, BlockTrain reaches cross entropy 1.359 (perplexity 3.89), within about 0.04 CE of a same-setup end-to-end Transformer reference, while each active worker trains only one block and avoids full-model optimizer state. A shared six-worker block training run reaches CE 1.385 by averaging same-block updates into one assembled model. HTTP/TCP transport experiments move real serialized checkpoints and updates, including a public-IP three-host run that improves CE from 5.580 to 1.811 while moving 15.22 GB. For inference, the current BlockTrain path uses one block-stack traversal per full output and serves over direct TCP across three public-network GPU hosts up to a 75.80B-parameter logical fp16 shape, outperforming a matched plain-autoregressive TCP pipeline baseline because it emits a full sequence per WAN pipeline traversal rather than one token per traversal.

13:00 JSTLLM/生成AI

タスク固有の LLM 蒸留のスケーリング則

大規模言語モデル (LLM) は、ますます広範囲のドメインにわたって強力なパフォーマンスを実現しますが、そのスケールにより、遅延とコストの制約が重要なアプリケーションでは導入の課題が生じます。この論文では、ドメイン固有の LLM 圧縮に関する経験的なスケーリング則を導き出し、データセットのサイズ、圧縮率、監視形式、反復枝刈りスケジュールによってドメイン内および一般知識のパフォーマンスがどのようにスケールされるかを定量化します。クオンツファイナンスをアプリケーションドメインとして使用し、反復構造枝刈りの下でロジットベースの蒸留と LoRA ベースの蒸留を比較し、推論トレース上で KL ダイバージェンスの蒸留を安定化する混合思考連鎖監視損失を導入します。圧縮下ではドメイン内タスクの品質が予想通り低下しますが、一般知識のベンチマークは同じ時点よりかなり前に崩壊します。監視形式はこのトレードオフの主な要因であり、思考連鎖監視は枝刈りによって消去された一般知識を積極的に回復します。私たちは、ヘッドライン データセット FinHeadlineMix、スケーリング則の結果、およびドメイン固有の圧縮決定のための再利用可能なフレームワークを提供する実践的な推奨事項をリリースします。

原文 (English)

Scaling Laws for Task-Specific LLM Distillation

Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in applications where latency and cost constraints are critical. This paper derives empirical scaling laws for domain-specific LLM compression, quantifying how in-domain and general knowledge performance scale with dataset size, compression ratio, supervision format, and iterative pruning schedule. Using quantitative finance as our application domain, we compare logit-based and LoRA-based distillation under iterative structural pruning, introducing a blended chain-of-thought supervision loss that stabilizes KL-divergence distillation over reasoning traces. In-domain task quality degrades predictably under compression while general-knowledge benchmarks collapse well before the same point; supervision format is the key driver of this tradeoff, with chain-of-thought supervision actively recovering general knowledge that pruning erases. We release the headline dataset FinHeadlineMix, scaling law results, and practical recommendations to provide a reusable framework for domain-specific compression decisions.

13:00 JST研究/論文GPT / ChatGPT

スケールは大規模な言語モデルにおける可塑性の損失を防ぐことができるか?

可塑性、つまり古い情報をすでに学習した後に新しい情報を学習するネットワークの能力の喪失は、継続的な学習が可能な人工ニューラル ネットワークを作成する際の基本的な課題です。この現象は何十年も前から知られていましたが、主に古い、比較的小規模なアーキテクチャで研究されており、自然言語領域で研究されることはほとんどありませんでした。現代のトランスフォーマーベースの LLM パラダイムにおいて可塑性の損失が依然として問題であるかどうかを判断するために、多言語の継続学習問題でトレーニングされた GPT スタイルのトランスフォーマー モデルにおける可塑性の損失を研究します。以前の研究と一致して、延期されたベトナムの探査タスクの劣化によって測定されたように、500万から3億1400万の非埋め込みパラメータの範囲のモデル全体で塑性損失の証拠が見つかりました。さらに、可塑性損失の開始は予測可能なスケーリング則に従っており、モデルのサイズに応じて非線形に増加することがわかりました。これらの結果は、モデルが大きくなると可塑性損失の測定可能な影響が遅れる可能性があるが、パラメータ数を増やすだけでは塑性損失を完全に防ぐには不十分である可能性が高いことを示唆しています。また、定常的な多言語トレーニング下で可塑性が失われる証拠も見つかり、この現象は突然の課題変更を伴う継続的な学習に限定されるという見解に異議を唱えています。全体として、私たちの結果は、自然言語でトレーニングされた大規模な Transformer 言語モデルであっても、継続的設定と定常的設定の両方で、十分に長いトレーニングの後、最終的には新しいデータに効率的に適応する能力を失うことを示唆しています。

原文 (English)

Can Scale Save Us From Plasticity Loss in Large Language Models?

The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning. Although this phenomenon has been known for decades, it has mostly been studied in older, relatively small architectures and rarely in natural-language domains. To determine whether loss of plasticity remains a problem in the modern transformer-based LLM paradigm, we study plasticity loss in GPT-style Transformer models trained on a multilingual continual learning problem. Consistent with prior work, we find evidence of plasticity loss across models ranging from 5M to 314M non-embedding parameters, as measured by deterioration on a held-out Vietnamese probing task. We further find that the onset of plasticity loss follows a predictable scaling law, growing sublinearly with model size. These results suggest that larger models may delay the measurable effects of plasticity loss, but that increasing parameter count alone is likely to be insufficient to completely prevent it. We also find evidence of plasticity loss under stationary multilingual training, challenging the view that the phenomenon is exclusive to continual learning with abrupt task changes. Overall, our results suggest that even large Transformer language models trained on natural-language will eventually lose the ability to efficiently adapt to new data after sufficiently long training, in both continual and stationary settings.

13:00 JST研究/論文GPT / ChatGPT

BluTrain: AI システム用の C++/CUDA フレームワーク

大規模な深層学習の進歩は、モデリングよりもシステム エンジニアリングの問題です。トレーニング中のモデルの動作 (スループット、メモリ フットプリント、結果の数値的忠実度) は、アーキテクチャ自体によって決まるというよりは、そのアーキテクチャがハードウェア上でどのように表現されるかによって決まります。システムの複雑さを抽象化してモデリングをシームレスにし、反復的なオーケストレーション ロジックの必要性を排除しながら、このハードウェア表現に対する絶対的な制御を実現するために、BluTrain は、標準 C++ およびコア CUDA プログラミング モデルにおける堅牢で軽量なアーキテクチャ全般のトレーニング フレームワークとして第一原理から設計されました。リバースモード autograd を備えた型付きテンソル モジュール、線形代数ライブラリ、キャッシュ アロケータ、マルチモード分散実行モジュール、MLIR ベースの深層学習コンパイラなど、すべての層がネイティブに実装されています。 8 GPU 6000 Ada システム上の FP32 で 124M パラメータの GPT-2 ベースラインをトレーニングする正式な評価では、BluTrain は、スループット (平均 407K トークン/秒を維持するのに対し、PyTorch の 395K トークン/秒を維持) とメモリ効率 (最大 22% のフットプリント削減を達成) の両方で業界標準のベースラインを上回っています。数値的忠実度が向上し、最終的な検証損失がわずかに低下するように収束します。すべてのレイヤーがネイティブ チューニングに対して明示的にオープンであるため、パフォーマンスの上限はフレームワーク自身で引き上げることができます。

原文 (English)

BluTrain: A C++/CUDA Framework for AI Systems

Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware. To achieve absolute control over this hardware expression while abstracting away systems complexity to make modelling seamless and eliminating the need for repetitive orchestration logic, BluTrain was architected from first principles as a robust, lightweight, and architecture-general training framework in standard C++ and the core CUDA programming model. Every layer is implemented natively: a typed tensor module with reverse-mode autograd, a linear-algebra library, a caching allocator, a multi-mode distributed-execution module, and an MLIR-based deep-learning compiler. In formal evaluations training a 124M-parameter GPT-2 baseline in FP32 on an 8-GPU 6000 Ada system, BluTrain outperforms industry-standard baselines in both throughput (sustaining an average of 407K tokens/s versus PyTorch's 395K tokens/s) and memory efficiency (achieving up to a 22% footprint reduction), while strictly preserving numerical fidelity and converging to a marginally lower final validation loss. With every layer explicitly open to native tuning, the performance ceiling is the framework's own to raise.

13:00 JST研究/論文

領域一般化のための人間の活動認識の分布シフトの評価

人間活動認識 (HAR) の分野は引き続き研究者の関心を集めており、重要な面で進歩していますが、いくつかの重要な課題が残っています。現実世界の設定で優れたパフォーマンスを示す HAR モデルを構築する際の最も困難な側面の 1 つは、デバイスとセンサーの異質性によるデータの多様性と、現実世界のアプリケーションに固有のコンテキストの変化に対処することです。 HAR におけるデータの多様性は文献でよく知られていますが、HAR モデルに対するさまざまなタイプの分布シフトの影響と、そこから生じるドメイン一般化問題についての理解にはまだギャップが残っています。その目的に向けて、このペーパーでは、デバイスの種類、センサーの配置、サンプリング レート、ユーザーの行動の変化を含む 4 つの異なるタイプの分布シフトを体系的に評価します。その影響を定量化することで、多様性の変化が主にあらゆる種類の変化を定義することを示し、異なるドメイン間で共有されない固有の特徴が存在することを示しています。次に、均一な HAR ベースの分散シフト ベンチマークを導入し、最大 28 のドメイン一般化手法の包括的な評価を実行します。私たちの分析は、経験的なリスク最小化ベースラインをわずかに上回るパフォーマンスで、モデルの一般化可能性を達成する際の現在のドメイン一般化アルゴリズムの限界を明らかにしました。この研究は、センサーベースの HAR における特定の分布シフトに関するドメインの一般化と適応の最初の体系的な調査を表しており、さらなる研究を促進するためのオープンソースのベンチマーク プラットフォームとデータセットを提供します。

原文 (English)

Assessing Distribution Shift in Human Activity Recognition for Domain Generalization

While the field of Human Activity Recognition (HAR) continues to draw interest from researchers and advance in important ways, some key challenges remain. One of the most difficult aspects of building HAR models that show good performance in real-world settings is dealing with data diversity from device and sensor heterogeneity, and contextual changes that are intrinsic to real-world applications. While data diversity in HAR has been well-acknowledged in the literature, there remains a gap in understanding the effect of various types of distribution shifts on HAR models and the domain generalization problem that arises. Towards that end, this paper systematically evaluates 4 different types of distribution shifts, including variations in device type, sensor placement, sampling rate, and user behavior. Quantifying their effects, we illustrate that diversity shifts predominantly define all types of shifts, indicating the existence of unique features that are not shared across different domains. We then introduce a uniform HAR-based distribution shift benchmarks and conduct a comprehensive evaluation of up to 28 domain generalization methods. Our analysis exposes the limitations of current domain generalization algorithms in achieving model generalizability, marginally outperforming the empirical risk minimization baseline. This work represents the first systematic exploration of domain generalization and adaptation concerning specific distribution shifts in sensor-based HAR, offering an open-source benchmark platform and datasets to spur further research.

13:00 JST研究/論文

双方向の条件付きフローマッチングによるカオスシステムの逆問題の解決

カオス システムのモデル化は重要ですが、困難です。カオス力学における逆問題、つまり最終状態から初期条件を推測する問題は、姿勢不良、非一意性、不安定性、および潜在的にカオスな時間反転力学のため、ほとんど未解決のままです。私たちは、双方向条件付きフロー マッチング (Bi-CFM) を使用してこの未解決の問題に対処します。Bi-CFM は、初期状態と最終状態の分布間の双方向マッピングを学習して、カオス進化の確率性を捉え、時間の経過に伴う指数関数的な誤差の蓄積を軽減します。さらに、保存則のある系については、それを保存制約付き Bi-CFM (CBi-CFM) に拡張します。 Bi-CFM は、古典的な Lorenz、Circuit、および高次元 Lorenz 96 システム全体で、ベースラインを上回る 5 つの分散レベルのメトリクスを改善し、2 桁を超える高速化を達成します。惑星力学における三体惑星間散乱問題では、CBi-CFM は保存則をよりよく尊重しており、保存誤差はグランド トゥルースの保存誤差と同等です。最後に、$\sim 10^{10}$ 年 (10 Gyr) の進化によって形作られた衝突百万体系である球状星団の実際の観測では、私たちの方法は精度の向上を示し、長期スケールの実世界のカオス力学の逆問題を解決するためのスケーラブルなルートを確立しました。

原文 (English)

Solving Inverse Problems of Chaotic Systems with Bidirectional Conditional Flow Matching

Modeling chaotic systems is crucial yet challenging. Inverse problems in chaotic dynamics, namely inferring initial conditions from final states, remain largely unsolved because of ill-posedness, non-uniqueness, instability, and potentially chaotic time-reverse dynamics. We address this open problem with Bidirectional Conditional Flow Matching (Bi-CFM), which learns bidirectional mappings between distributions of initial and final states to capture the stochasticity of chaotic evolution and mitigate exponential error accumulation over time. Furthermore, for systems with conservation laws, we extend it to Conservation-constrained Bi-CFM (CBi-CFM). Across the classic Lorenz, Circuit, and high-dimensional Lorenz 96 systems, Bi-CFM improves five distribution-level metrics over baselines while achieving a speedup of more than two orders of magnitude. In the three-body planet-planet scattering problem in planetary dynamics, CBi-CFM better respects conservation laws, with conservation errors comparable to those of the ground truth. Finally, on real observations of globular clusters, collisional million-body systems shaped by $\sim 10^{10}$ years (10 Gyr) of evolution, our method represents an advance in accuracy, establishing a scalable route to solving inverse problems of long-timescale real-world chaotic dynamics.

13:00 JST研究/論文

違いを生まずに違いを生み出す

一連の 7 つの論文にわたって、アンドレアスと G は実際の因果関係の 7 つの定義を導入し、それらを 3 つの異なる競合する説明タイプに属するものとして分類しました。それは、事実による差異形成、反事実による差異形成、および規則性に基づくものです。私は、彼らの最新の事実による差異形成の定義が 3 つのタイプすべてを具体化していることを示し、それによってこれらが違いのない区別であることを証明しました。さらに、いくつかの重要な例について、彼らの新しい説明を他の 6 つの説明と比較し、これが次のことを明らかにしています。彼らの7つのアカウントすべてを台無しにします。

原文 (English)

Difference-Making without Making a Difference

Over a series of seven papers, Andreas & G\"unther have introduced seven definitions of actual causation and have classified them as belonging to three different, competing, types of accounts: factual difference-making, counterfactual difference-making, and regularity-based. I show that their most recent - factual difference-making - definition instantiates all three types, thereby proving that these are distinctions without a difference. I further compare their novel account to the other six accounts on several crucial examples, revealing that this undermines all seven of their accounts.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Copilot

NFR 評価のためのマルチターン LLM ダイアログの精度と満足度

LLM ベースの対話アシスタントはソフトウェア開発者にとって主流のツールとなっていますが、現在の評価ベンチマークは機能の正しさのみに焦点を当てています。このため、本質的に曖昧でコンテキストに依存し、プログラムの多くの部分に関係する非機能要件 (NFR) を処理する際に、これらの会話の品質と正確性を評価する際に重大なギャップが残ります。これらのシステムが NFR に関する協調推論をどの程度サポートしているかを評価するには、シングル ターンの精度を超えて、システムの出力の正確さとマルチ ターン インタラクションの品質の両方を取得する方法が必要です。このペーパーでは、医療保険の相互運用性と責任に関する法律 (HIPAA) 規制順守の領域における、開発者と LLM ベースのエージェントとの間の複数回にわたる会話の精度と品質を調査します。私たちは 49 人のプログラマーを雇い、GitHub Copilot と対話し、HIPAA 規制に準拠するように設計されたシステムである iTrust コードベースに対して 148 個の HIPAA 由来の NFR を、要件満足度、推論、コード ローカリゼーションの 3 つの側面にわたって評価しました。開発者は LLM 評価に同意する傾向がありますが、専門家のグラウンド トゥルースに対する精度は低いことがわかりました。ユーザー満足度をモデル化したところ、システムの応答時間が長くなり、情報提供ターンが増えるとユーザー満足度にマイナスの影響が出るのに対し、積極的なインタラクションはプラスの影響を与えることがわかりました。私たちの調査結果は、NFR 評価をサポートする LLM ベースの対話システムを設計するための洞察を提供します。

原文 (English)

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a critical gap in assessing the quality and accuracy of these conversations when handling Non-Functional Requirements (NFRs), which are inherently vague, context-dependent, and involve many parts of a program. Evaluating how well these systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of the system's outputs and the quality of the multi-turn interaction. In this paper, we investigate the accuracy and quality of multi-turn conversations between developers and an LLM-based agent in the domain of Health Insurance Portability and Accountability Act (HIPAA) regulatory compliance. We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase, a system designed to comply with HIPAA regulations, across three dimensions: requirement satisfaction level, reasoning, and code localization. We find that developers tend to agree with LLM assessments, but accuracy against expert ground truth is low. We model user satisfaction and find that longer system responses and more information-providing turns negatively affect user satisfaction, whereas proactive interactions positively affect it. Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.

13:00 JSTエージェントハードウェア/半導体

採点者の採点: エージェントによるデータ分析システムの評価から得た教訓

エージェント的データ分析システムは、コード、数値結果、口頭診断などの豊富な出力を生成します。このため、シングルターン LLM 応答よりも評価が難しくなります。したがって、エージェントの出力と、採点アーティファクトからの真実の回答との間の真の不一致を区別する必要があります。私たちは、DSGym の 153 の数値 QRData タスクにマルチエージェント データ分析システムである LAMBDA を適用することで、自動採点者がそのようなシステムをどのように確実に評価するか、またどのような戦略が採点の品質を向上させるかを調査します。私たちは、厳格な正規表現マッチング、LLM ベースの寛大なグレーディング、スニペットベースの人的検査という 3 層の人的 AI グレーディング カスケードを開発および評価します。これは、非 GenAI 戦略と GenAI 戦略をさまざまな障害プロファイルと組み合わせたものです。どちらの自動グレーダーも 100% の観察精度 (誤検知 0/70) を達成しています。寛大な採点者の再現率は人間のラベルに対して 97% です。キーワードに固定された抽出パイプラインにより、厳密な採点者の再現率は、最後の数字のヒューリスティックよりも 60 パーセント ポイント高くなります。寛大なグレーダーはアーキテクチャ的にパーサーに依存しません。反復的なナッジメカニズムにより、採点の成功率が 36% から 97% に上昇し、寛容な合格率が 16% から 46% に上昇します。元の質問の再挿入ありとなしのナッジを比較すると、再挿入にはメリットがないことがわかり、ナッジが回答テンプレートの手がかりであることが確認されました。さらに、このケース スタディでは、変数タイプがパイプラインのダイナミクスの評価と観察された結果の評価に最も一貫して関連付けられているタスク メタデータ フィールドであることがわかります。

原文 (English)

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses. It is therefore necessary to distinguish genuine disagreement between an agent's output and a ground-truth answer from grading artifacts. We investigate how reliably automated graders assess such a system and what strategies improve grading quality by applying LAMBDA, a multi-agent data-analysis system, on 153 numerical QRData tasks from DSGym. We develop and evaluate a three-layer human-AI grading cascade: strict regex matching, LLM-based lenient grading, and snippet-based human inspection, which combines non-GenAI and GenAI strategies with different failure profiles. Both automated graders achieve 100% observed precision (0/70 false positives). The lenient grader's recall is 97% against human labels. A keyword-anchored extraction pipeline raises the strict grader's recall by 60 percentage points over a last-number heuristic; the lenient grader is architecturally parser-independent. An iterative nudge mechanism raises grading run success from 36% to 97% and lenient-pass rates from 16% to 46%; comparing nudging with and without original-question re-injection shows that re-injection offers no benefit, confirming the nudge as an answer template cue. We further observe in this case study that variable type is the task metadata field most consistently associated with grading pipeline dynamics and observed outcome grades.

13:00 JSTLLM/生成AI

タスクと目標のマッチング: エンコーダ/デコーダの事前トレーニング済み言語モデルの微調整および即時調整戦略

プロンプトベースの学習は、自然言語処理における主要なパラダイムとして浮上しています。この研究では、常識的な知識の検索と完了に焦点を当て、生成タスクと質問応答タスクにわたるエンコーダ/デコーダの事前トレーニング済み言語モデルのパフォーマンスに対する、さまざまな事前トレーニング目標の影響を調査します。トレーニング前と微調整の両方の段階で複数の目標を組み込む利点を強調します。特定のタスクに適切な目標を決定するための Match Task to Objective (MTO) フレームワークとメソッドを紹介します。このフレームワークは、特定された目的に基づいて、教師なしトレーニングによる適応のためにタスク関連データを準備する自動化された方法を提供します。微調整段階では、事前トレーニング段階と適応段階の目的に沿った新しいテンプレートを設計します。タスクの要件に合わせてこれらの戦略を使用すると、数ショット設定で従来の方法と比較して 120\% 以上のパフォーマンス向上を達成できます。これらは、数ショット設定で関連する作業を大幅に上回り、フルデータセットのシナリオでもベースラインを上回ります。さらに、このアプローチを拡張してプロンプト チューニング方法論を含め、より効果的なソフト プロンプト エンジニアリングと最適化のためのガイダンスを提供します。私たちの戦略は、プロンプトチューニングのパフォーマンスも大幅に向上させます。これらの洞察には大きな価値があり、特定のタスク用にカスタマイズされたモデルの選択と最適化を正確にガイドします。コードは https://github.com/puraminy/MTO/ で入手できます。

原文 (English)

Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models

Prompt-based learning has emerged as a dominant paradigm in natural language processing. This study explores the impact of diverse pre-training objectives on the performance of encoder-decoder pre-trained language models across generation and question answering tasks, with a focus on commonsense knowledge retrieval and completion. We highlight the benefits of incorporating multiple objectives during both pre-training and fine-tuning stages. We introduce the Match Task to Objective (MTO) framework and methods for determining the appropriate objective for a given task. This framework offers automated methods to prepare task-related data for adaptation through unsupervised training, based on the identified objective. In the fine-tuning stage, we design novel templates that align with the objectives of the pre-training and adaptation stages. When aligned with task requirements, these strategies can achieve a performance gain of over 120\% compared to conventional methods in few-shot settings. They significantly outperform related works in few-shot settings and exceed the baseline even in full-dataset scenarios. Furthermore, we extend this approach to include prompt-tuning methodologies, providing guidance for more effective soft prompt engineering and optimization. Our strategies significantly enhance prompt-tuning performance as well. These insights hold substantial value, precisely guiding the selection and optimization of models customized for specific tasks. Code is available at https://github.com/puraminy/MTO/

13:00 JSTエージェント

バラバラの世界モデル: 一般代理店向けの構造認証

大きな世界体制では、エージェントは普遍的な能力を持つことはできず、エージェントの能力は必然的に世界モデル全体に​​わたって細分化されて特殊化されます。その結果、標準的な統一保証では、重大なボトルネックと無関係な障害の理解が区別できません。まず、一般的なエージェントが普遍的ではなく、標準的な最悪の場合の分析が役に立たないことを証明することで、この制限を形式化します。これを克服するために、私たちは、限定された目標条件付きパフォーマンスをエージェントの内部世界モデルのエントリごとの保証にマッピングする移行ローカル フレームワークである構造認証を導入します。私たちの主な貢献は建設的なものです。私たちは、深い構成目標を使用して特定の遷移をフィルタリングするアルゴリズムを提供し、これらの目標に関する一般的なエージェントが $\mathcal{O}(1/n) + \mathcal{O}(\delta)$ 誤差限界を持つ構造世界モデルを持っていることを証明します。逆に、この限界は、Small-$\delta$ 体制では厳しく、その存在は当社の認定によって明示的に保証されています。これらの結果により、長期的な計画が信頼できる特定の移行を局所化することで、一般的なエージェントの認証可能な展開が可能になります。

原文 (English)

World Models in Pieces: Structural Certification for General Agents

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrelevant failures. We first formalize this limitation by proving that general agents are not universal, rendering standard worst-case analysis uninformative. To overcome this, we introduce structural certification, a transition-local framework that maps bounded goal-conditioned performance to entry-wise guarantees on the agent's internal world model. Our main contribution is constructive. We provide algorithms that filter specific transitions using deep compositional goals and prove that a general agent on these goals has a structural world model with a $\mathcal{O}(1/n) + \mathcal{O}(\delta)$ error bound. Conversely, this bound is tight in the small-$\delta$ regime, whose existence is explicitly guaranteed by our certification. These results enable the certifiable deployment of general agents by localizing the specific transitions where long-horizon planning is reliable.

13:00 JSTエージェント

OpenThoughts-Agent: エージェントティック モデルのデータ レシピ

エージェント言語モデルは AI のアプリケーションを劇的に拡張しますが、幅広い能力を持つエージェントのトレーニング データを収集する方法についてはほとんど公に知られていません。 SWE-Smith、SERA、Nemotron-terminal などの既存のオープンな取り組みは通常、単一のベンチマークをターゲットにしており、多様なエージェント タスクにわたって一般化するモデルをトレーニングする方法という問題が残されています。 OpenThoughts-Agent (OT-Agent) プロジェクトは、エージェント モデルをトレーニングするための完全にオープンなデータ キュレーション パイプラインでこのギャップに対処します。当社では、パイプラインの各段階を体系的に調査するために 100 件を超える制御アブレーション実験を実施し、タスクソースと多様性の重要性についての洞察をもたらします。次に、パイプラインから 100,000 個のサンプルのトレーニング セットを組み立て、このデータセットで Qwen3-32B を微調整します。これにより、7 つのエージェント ベンチマーク全体で平均精度 44.8% が得られ、既存の最も強力なオープン データ エージェント モデル (Nemotron-terminal-32B、40.9%) と比較して 3.9 パーセント ポイントの改善が得られました。さらに、当社のトレーニング データは強力なスケーリング特性を示し、コンピューティング制御による比較において、あらゆるトレーニング セット サイズで代替のオープン データセットを上回るパフォーマンスを示します。私たちは、エージェント モデル トレーニングに関する将来のオープン研究をサポートするために、トレーニング セット、データ パイプライン、実験データ、およびモデルを openthoughts.ai で一般公開しています。

原文 (English)

OpenThoughts-Agent: Data Recipes for Agentic Models

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation pipeline for training agentic models. We conduct more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline, yielding insights on the importance of task sources and diversity. We then assemble a training set of 100K examples from our pipeline and fine-tune Qwen3-32B on this dataset, which yields an average accuracy of 44.8% across seven agentic benchmarks and a 3.9 percentage point improvement over the strongest existing open data agentic model (Nemotron-Terminal-32B, 40.9%). Moreover, our training data exhibits strong scaling properties, outperforming alternative open datasets at every training set size in compute-controlled comparisons. We publicly release our training sets, data pipeline, experimental data, and models at openthoughts.ai to support future open research on agentic model training.

13:00 JST研究/論文

有限グラフ上の遅延結合反応拡散システムとしてのリエントラント値フィールド

二部構成のヒルベルト・シュミットカーネルを介して記号場が幾何学場に結合される動的システムについて説明します。このシステムは、リプシッツ条件と小さなゲイン条件に従う履歴空間上の遅延関数微分方程式 (RFDE) によって完全に記述されます。 RFDE が一定の入力の下で適切に配置され、コンパクトなグローバル アトラクターを許容することを示します。 2 つの主フィールドと実行フィールドで構成される主サブシステム $(H_L, X_R, P)$ は、フィールド間結合が $C_{\mathcal{K}}^2<\mu_L\mu_R$ を満たす限り、遅延に関係なく全体的に安定していることが示されています。さらに、主定理の仮説を満たす設計仕様について説明します。

原文 (English)

Reentrant value fields as delayed coupled reaction-diffusion systems on finite graphs

We describe a dynamical system in which a symbolic field is coupled to a geometric field via a bipartite Hilbert-Schmidt kernel. The system is fully described by a retarded functional differential equation (RFDE) on the history space, subject to Lipschitz and small gain conditions. We show that the RFDE is well-posed under constant input and that it admits a compact global attractor. The principal subsystem $(H_L, X_R, P)$, which is comprised of the two primary fields as well as an executive field, is shown to be globally stable independent of delay, provided that the interfield coupling satisfies $C_{\mathcal{K}}^2<\mu_L\mu_R$. In addition, we describe design specifications that fulfill the hypotheses of the main Theorem.

13:00 JST研究/論文

自己回帰の地平線を超えて: コードの拡散モデル、世界モデリング、および状態空間モデルの包括的な調査

自己回帰 (AR) 言語モデルは、自動ソフトウェア エンジニアリングの大幅な進歩を促進し、強力なコード生成および支援システムを可能にします。ただし、ネクスト トークン予測パラダイムでは、グローバル プランニングの制限、長期的な依存関係の維持における課題、プログラム実行セマンティクスの基盤の制限など、コード推論に構造的な制限が生じます。 AR モデルに対する既存の文献の大きな偏りに注目し、コード インテリジェンスのための次世代アーキテクチャ機能を解放することで、次のトークン予測のロジックとスケーリングのボトルネックを潜在的に克服できる可能性のある新たなパラダイムについて議論します。具体的には、拡散モデルの可能性について説明します。拡散モデルは、AR モデルでは見逃されがちな長距離の構文制約を捕捉する全体的なノイズ除去を介してコードを生成します。また、推論をサポートするために実行状態をシミュレートするコード ワールド モデル (CWM) と、大規模なコンテキストに対して線形時間効率を提供する状態空間モデル (SSM) についても説明します。これらの開発を認知神経科学の発見と結び付けることで、「システム 2」コード生成エージェントの開発の方向性を概説します。

原文 (English)

Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code

Autoregressive (AR) language models have driven significant progress in automated software engineering, enabling powerful code generation and assistance systems. However, the next-token prediction paradigm introduces structural limitations for code reasoning, including restricted global planning, challenges in maintaining long-range dependencies, and limited grounding in program execution semantics. Noting the heavy skewness of existing literature towards AR models, we discuss emerging paradigms that could potentially overcome the logic and scaling bottlenecks of next-token prediction by unlocking next-generation architectural capabilities for code intelligence. Specifically, we discuss the potential of Diffusion Models, which generate code via holistic denoising that captures long-range syntactic constraints often missed by AR models. We also discuss Code World Models (CWMs), which simulate execution states to support reasoning, and State Space Models (SSMs), which provide linear-time efficiency for massive contexts. By connecting these developments with findings from cognitive neuroscience, we outline directions for developing "System 2" code generation agents.

13:00 JSTLLM/生成AIビジネス/資金調達

RAG システムにおける以前の優位性の定量化

検索拡張生成(RAG)は大規模言語モデルを外部の知識に基づいて構築しますが、現在の評価は「認識論的盲目さ」に悩まされる離散ヒューリスティックに依存しており、真の文脈情報抽出とパラメトリック記憶想起を区別できません。これに対処するために、NCU (Normalized Context Utilization) メトリクスを導入し、ゼロショット、オラクル、および敵対的条件にわたる連続トークンのログ確率を活用して、コンテキスト情報の獲得を厳密に定量化します。独自の商用 API と並行して 1.5B から 72B のパラメータの範囲のアーキテクチャを評価すると、厳密な事実抽出 (思考連鎖推論なし) の場合、従来のスケーリング則では極端な利益逓減が見られることが明らかになりました。つまり、高効率の小型言語モデル (SLM) は、大容量アーキテクチャに匹敵するか、それを上回っています。さらに、「事前優位性」がモデルのスケールや独自のアラインメントと相関していることを示します。評価された商用 API は、敵対的紛争のほぼ半数で明示的な外部証拠を無効にしただけでなく、パラメトリック事前条件が矛盾した場合にシステムの信頼崩壊 (ネガティブトランスファー) に頻繁に悩まされました。私たちの調査結果は、厳密な抽出ワークフローにおける SLM の構造的認識論的な利点と優れた文脈順守を強調しています。

原文 (English)

Quantifying Prior Dominance in RAG Systems

Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suffer from ''epistemic blindness'' - failing to distinguish genuine contextual information extraction from parametric memory recall. To address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adversarial conditions to strictly quantify contextual information gain. Evaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of-Thought reasoning), traditional scaling laws exhibit extreme diminishing returns: highly efficient Small Language Models (SLMs) match or outperform high-capacity architectures. Furthermore, we demonstrate that ``Prior Dominance'' correlates with model scale and proprietary alignments. The evaluated commercial API not only overrode explicit external evidence in nearly half of adversarial conflicts, but also frequently suffered from systemic confidence collapse (Negative Transfer) when its parametric priors were contradicted. Our findings highlight the structural epistemic advantage and superior contextual adherence of SLMs in strict extraction workflows.

13:00 JST研究/論文

SemChunk-C: C コードのセマンティック セグメンテーション

C ファミリ言語で書かれたコードのセマンティック セグメンテーションは、言語の複雑な構文、マクロ拡張、不規則な構造パターンのため、依然として困難な問題です。固定サイズのウィンドウ、ヒューリスティック分割、構文ベースのツールなどの既存のチャンク手法では、意味のある機能単位をキャプチャできないことが多く、検索やその他の下流の LLM 主導タスクの有効性が制限されます。このペーパーでは、C 関連言語におけるチャンクの問題について取り上げます。まず、コード チャンク カテゴリのセットを定義します。次に、LLM ベースの分類器をトレーニングして、a) チャンクの境界を識別し、b) 各チャンクに説明的な機能属性 (カテゴリ) を割り当てます。これは下流のタスクに役立ちます。コード内のセマンティック コンテキストをキャプチャする LLM の機能を活用することで、柔軟なチャンク境界を想定し、各インスタンスの特定の構造とコンテキストに適応できるようにします。 3 番目に、C 関連ファイル (.c、.cpp、.h、.cs など) のセマンティック チャンキング用の軽量言語モデル ファミリである SemChunk-C を紹介します。これらのモデルは、17M、32M、68M、および 150M パラメーターを備えた最初の 4 つの Ettin エンコーダー [1] に基づいています。サイズが比較的小さいにもかかわらず、データ構造、インターフェイス ブロック、その他のコンポーネントなど、まとまったコード単位を識別することができます。さらに、ネストされた定義やマクロなどの困難な構造を含む、実際のコードに対するアプローチの堅牢性を実証します。私たちはさまざまなデータセットでアプローチをテストし、高い境界精度とセマンティック一貫性を達成し、はるかに大規模なコード指向の LLM に基づくチャンカーと一致またはそれを上回るパフォーマンスを示します。また、いくつかの厳選されたベンチマークでダウンストリーム タスクのパフォーマンスの向上も検証します。

原文 (English)

SemChunk-C: Semantic Segmentation for C Code

Semantic segmentation of code written in a C-family language remains a challenging problem, due to the language's complex syntax, macro expansion, and irregular structural patterns. Existing chunking methods, such as fixed-sized windows, heuristic splitting, and syntax-based tools, often fail to capture meaningful functional units, limiting the efficacy of retrieval and other downstream LLM driven tasks. In this paper, we address the problem of chunking in C-related languages. First, we define a set of code chunk categories. Second, we train an LLM-based classifier to a) identify chunk boundaries, and b) assign each chunk a descriptive functional attribute (a category), which can be useful for downstream tasks. By leveraging the LLM's ability to capture semantic context within the code, we assume flexible chunk boundaries, allowing to adapt to the specific structure and context of each instance. Third, we introduce SemChunk-C, a family of lightweight language models for semantic chunking of C-related files (.c, .cpp, .h, .cs, etc.). These models are based on the first four Ettin encoders [1] with 17M, 32M, 68M, and 150M parameters. Despite their relatively small size, they are capable of identifying cohesive code units, such as data structures, interface blocks, and other components. Furthermore, we demonstrate the robustness of our approach on real-world code, including challenging constructs such as nested definitions and macros. We test our approach on various datasets, and show that it achieves high boundary accuracy and semantic coherence, matching or outperforming chunkers that are based on much larger code-oriented LLMs. We also validate the improved performance of the downstream tasks on a few curated benchmarks.

13:00 JSTハードウェア/半導体ClaudeNVIDIA

必要なのは FP8 だけです (パート 2): Tensor-core Garner 再定式化と Kulisch エスケープ ルートによる効率的な Ozaki-Bailey スタイル FFT

NVIDIA の Blackwell Ultra (B300) は、FP64 ベクトル スループットを GPU あたり約 1.3 TFLOPS に削減します。これは、B200 の約 30 分の 1 であり、帯域幅が制限された FP64 ワークロードがメモリ制限にとどまるレベルをはるかに下回ります。 Ozaki Scheme II フレームワークは、仮数スライスされた中国剰余再構成を使用して FP8 テンソル コアを介して密行列乗算をルーティングすることにより、FP64 と同等のスループットを回復します。関連資料のパート (1) では、高密度 GEMM、バッチ GEMV、ステンシル、および SpMV について説明します。この論文では、5 番目の標準プリミティブである 3-D FFT を追加します。 FP8 テンソル コア上の両方の 1-D FFT GEMM を使用した Bailey 6 ステップ分解を介してエミュレートされた 3-D FFT である Ozaki-Bailey FFT を紹介します。 Bailey の小さな内因数 k ~ sqrt(N) (N=1024 の場合 k=32) は、カーネルを k << r^2 の状態に置き、3 番目の TME パラメーター ガンマ (再構成待ち時間) が償却ではなく結合します。 Garner の再構成は、フェーズ A (FP8/INT8 テンソル コアの内積、B300 の 1024^3 で約 1 ミリ秒) とフェーズ B (出力ごとの削減) に分割されます。 Kulisch 固定小数点完全演算は、完全に INT32 SIMT パイプ上で実行しながら完全な FP64 精度を維持するフェーズ B 再定式化であると認識します。閉じた形式の帯域幅パリティの下限を導出します。ネイティブ FP64 の下限は 1.56*B_HBM (8 TB/s で 12.5 TF) です。B300 の 1.3 TF は約 10 倍低く、Rubin の 33 TF は 4% 以内です。 Kulisch 避難ルートには、INT32 サブフロア 8.25*B_HBM と FP8 フロア 170*B_HBM が必要です。 B300 はその両方を満たします。フル FP64 での 1024^3 の予測は約 18 ミリ秒で、実質的には 12.9 ミリ秒のメモリ ルーフです。 GPU がネイティブ フロアまたは両方の Kulisch フロアを満たす場合、GPU はメモリルーフ FFT パリティを満たします。この予測が実際に当てはまれば、B300 はソフトウェアのみでフル FP64 FFT を実行できるようになり、libKulisch ライブラリとベンチマーク キャンペーンの動機付けとなります。

原文 (English)

FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route

NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at which bandwidth-limited FP64 workloads stay memory-bound. The Ozaki Scheme II framework recovers FP64-equivalent throughput by routing dense matrix multiply through FP8 tensor cores with a mantissa-sliced Chinese-remainder reconstruction. A companion Part (1) paper covers dense GEMM, batched GEMV, stencils, and SpMV; this paper adds the fifth canonical primitive, the 3-D FFT. We present Ozaki-Bailey FFT, an emulated 3-D FFT via the Bailey six-step decomposition with both 1-D FFT GEMMs on FP8 tensor cores. Bailey's small inner factor k ~ sqrt(N) (k=32 for N=1024) puts the kernel in the regime k << r^2, where the third TME parameter gamma (reconstruction latency) binds rather than amortising. Garner reconstruction splits into Phase A (inner products on FP8/INT8 tensor cores, ~1 ms for 1024^3 on B300) and Phase B (per-output reduction). We identify Kulisch fixed-point complete arithmetic as a Phase B reformulation that keeps full FP64 accuracy while running entirely on the INT32 SIMT pipe. We derive closed-form bandwidth-parity floors. The native FP64 floor is 1.56*B_HBM (12.5 TF at 8 TB/s): B300's 1.3 TF sits ~10x below, Rubin's 33 TF within 4%. The Kulisch escape route needs an INT32 sub-floor 8.25*B_HBM and an FP8 floor 170*B_HBM; B300 meets both. The projection is ~18 ms for 1024^3 at full FP64, essentially the 12.9 ms memory roof. A GPU meets memory-roof FFT parity if it satisfies either the native floor or both Kulisch floors. If the projection holds in practice, B300 becomes viable for full-FP64 FFT through software alone, motivating a libKulisch library and benchmark campaign.

13:00 JSTLLM/生成AIGPT / ChatGPT

自己認識の微調整により、突発的な位置ずれを防止し、逆転させることができます

緊急不整合 (EM) は、不整合なペルソナ ベクトルと邪悪なキャラクター特性の活性化に関連しており、EM は有害なコンテンツを直接学習するのではなく、モデルの整合したキャラクターを破壊することによって機能することを示唆しています。このつながりを動機として、私たちは既存のトレーニング中の防御とは異なる、文字をターゲットとした介入として自己生成テキスト認識 (SGTR) の微調整を研究しています。私たちは、3 つのモデル (GPT-4.1、Qwen2.5-32B-Instruct、Seed-OSS-36B-Instruct) と複数の EM データセットにわたって 2 段階の微調整実験を実施し、SGTR 微調整を無害な微調整ベースライン (正しいドメイン固有のデータ、一般知識、単語カウント) と比較し、逆転と防止の両方の設定で効果的な防御であることを確認しました。すべての介入が同等の EM の逆転をもたらすことがわかりましたが、それは EM によって低下した能力を回復する場合に限られます。予防に関しては、SGTR 微調整のみが個々の指標を悪化させることなく一貫して位置ずれを低減しており、性格の強化が特に予防を促進していることを示唆しています。我々は、EM 微調整が LLM のアイデンティティ自己報告に多様性を誘導し、自己認識を人為的に損なうことで EM 微調整によって引き起こされる不整合を悪化させ、モデルのアイデンティティを保持するシステム プロンプトを削除すると EM 微調整の効果が大幅に減少することを示すことにより、EM と LLM のデフォルト特性との関係についてのさらなる証拠を提供します。これらの発見を総合すると、EM は一貫して不整合な人格の採用としてではなく、整合性のある性格の不安定化として再構成されます。

原文 (English)

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content. Motivated by this connection, we study self-generated text recognition (SGTR) finetuning as a character-targeted intervention that is distinct from existing in-training defenses. We conduct two-stage finetuning experiments across three models (GPT-4.1, Qwen2.5-32B-Instruct, Seed-OSS-36B-Instruct) and multiple EM datasets to compare SGTR finetuning against benign finetuning baselines (correct domain-specific data, general knowledge, and word counting) to find it an effective defense in both reversal and prevention settings. We find that all interventions produce comparable EM reversal, but only when restoring capabilities that EM had degraded. For prevention, only SGTR finetuning consistently reduces misalignment without exacerbating any individual metric, suggesting that character fortification specifically drives prevention. We provide further evidence for EM's relation to the LLM's default character by showing that EM finetuning induces diversity into the LLM's identity self-reports, artificially corrupting self-recognition exacerbates misalignment caused by EM finetuning, and that removing the model's identity-bearing system prompt substantially reduces the effect of EM finetuning. Together, these findings reframe EM not as the adoption of a coherent misaligned persona but as the destabilization of aligned character.

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

製品の望ましさの効率的かつ説明可能な数値分析と分類された暗黙的センチメント分析のための LLM の使用法を評価する

製品に関する定性的なフィードバックは微妙なユーザー エクスペリエンスを明らかにする可能性がありますが、その暗黙の感情を測定するのは困難です。このペーパーでは、大規模言語モデル (LLM) を使用して、そのようなデータから製品の望ましさを定量化する、スケーラブルで解釈可能なフレームワークを紹介します。 ZORQ と CARMA の 2 つの Product Desirability Toolkit (PDT) データセットを使用し、ゴールドスタンダードの人による注釈を備えた 106 の回答者用語グループで構成され、明示的なレビュー スコアに依存せずに、ゼロショットの連続数値センチメント スコアリングとカテゴリカルセンチメント分類が評価されます。データセット全体にわたって、LLM は定性的回答と厳密に一致する専門家ラベルから数値感情スコアを直接生成し、最大 0.97 のピアソン相関と最大 94% の分類精度を達成しました。 LLM は、複数の形式で提示されたデータを処理する場合でも堅牢性を維持し、一貫して高い信頼性を示しました。対照的に、語彙ベースのベースラインとトランスフォーマーのベースラインでは、統計的に有意な結果は得られませんでした。テストしたモデルの中で、GPT-4o-mini は、94% 低いコストで大型モデルと同等のパフォーマンスを達成し、スケーラブルな導入をサポートしました。このフレームワークには、モデルの信頼度評価と人間が判読できる根拠の説明 (xAI) も組み込まれており、製品満足度評価における実用化をサポートしながら、解釈可能性、透明性、信頼性を向上させます。一般に、アンケート手法として PDT ツールを費用効率の高い LLM とともにセンチメント分析に使用すると、センチメント スコア (数値センチメントと分類されたセンチメントの両方) の点で豊富な結果が得られる製品評価を提供できる可能性があり、製品の開発と改善のアイデアや、ターゲット ユーザー向けのマーケティング アイデアを特定するために使用できる製品の高レベルのユーザー インプレッションの点で役立ちます。

原文 (English)

Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability

Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure. This paper presents a scalable and interpretable framework that uses large language models (LLMs) to quantify product desirability from such data. Using two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA comprising 106 respondent term groupings with gold-standard human annotation, zero-shot continuous numerical sentiment scoring and categorical sentiment classification are evaluated without relying on explicit review scores. Across the datasets, LLMs generated numerical sentiment scores directly from qualitative responses and closely matched expert labels, achieving Pearson correlations up to 0.97 and classification accuracy up to 94%. LLMs maintained robustness even when handling data presented in multiple forms and consistently expressed high confidence. In contrast, lexicon-based and transformer baselines did not produce statistically significant results. Among the models tested, GPT-4o-mini achieved performance comparable to larger models at 94% lower cost, supporting scalable deployment. The framework also incorporates model confidence ratings and human-readable rationale explanations (xAI), improving interpretability, transparency, and trust while supporting practical use in product satisfaction assessment. In general, using the PDT tool as a survey method along with a cost efficient LLM for sentiment analysis has the potential to provide for product evaluation with results that are rich in terms of sentiment scores (both numerical and classified sentiment) and in terms of the high-level user impressions of the product that can be used to identify ideas for product development and improvement, as well as marketing ideas for target audiences.

13:00 JST研究/論文

分布シフト下での水中音響変調認識のための異種 2D/1D 信号表現の融合

変調認識システムは、異種信号表現に依存しています。時間周波数マップや周期定常マップなどの 2D 信号画像モダリティは構造パターンをキャプチャし、高次パワー スペクトルなどの 1D 統計記述子は相補的な手がかりをエンコードします。分布の変化の下では、これらのモダリティは不均一に劣化するため、堅牢な融合が実用化の中心的な課題となっています。異なるシフトタイプを体系的に分離する統一された評価プロトコルが欠如しているため、進歩はさらに制限されています。この論文では、水中音響変調認識におけるベンチマークとモデルの共同研究を通じて、両方の課題に対処します。 UAMR-ShiftBench は、配信内、低 SNR、目に見えない環境、目に見えない通信パラメータ、および測定された海上試験評価を単一の一致するプロトコルの下で共同でカバーする最初のベンチマークであり、南シナ海で 3 月と 11 月に実施された 2 つの海上試験キャンペーン中に収集された 2 つの独立した現実世界のサブセットを使用します。 SCP-TriCA は、STFT、周期定常性、および P2/P4 (2 次および 4 次パワースペクトル) モダリティを階層的に融合します。2 つの 2D モダリティは、まず双方向クロスアテンションを通じて調整され、次に 1D 統計モダリティがサンプル適応選択ゲートを通じて組み込まれます。 UAMR-ShiftBenchでは、SCP-TriCAは分布内精度95.33%とシミュレートOOD平均74.59%を達成し、最強のベースラインを5.12パーセントポイント上回り、2つの海上試験サブセットでは91.14%と94.86%に達し、最良のベースラインをそれぞれ15.71パーセントと23.00パーセントポイント上回りました。アブレーションの結果は、効果がモダリティの相補性と階層的融合設計に由来していることを裏付けています。コードとモデルは https://github.com/ronglaiqian/UAMR-ShiftBench で入手できます。

原文 (English)

Heterogeneous 2D/1D Signal Representation Fusion for Underwater Acoustic Modulation Recognition Under Distribution Shift

Modulation recognition systems rely on heterogeneous signal representations. 2D signal-image modalities such as time-frequency and cyclostationary maps capture structural patterns, while 1D statistical descriptors such as higher-order power spectra encode complementary cues. Under distribution shift, these modalities degrade unevenly, making robust fusion a central challenge for practical deployment. Progress is further limited by the lack of a unified evaluation protocol that systematically separates different shift types. This paper addresses both challenges through a joint benchmark-and-model study in underwater acoustic modulation recognition. UAMR-ShiftBench is the first benchmark to jointly cover in-distribution, low-SNR, unseen-environment, unseen-communication-parameter, and measured sea-trial evaluation under a single matched protocol, with two independent real-world subsets collected during two sea-trial campaigns conducted in March and November in the South China Sea. SCP-TriCA fuses STFT, cyclostationary, and P2/P4 (second- and fourth-order power spectra) modalities hierarchically: the two 2D modalities are first aligned through bidirectional cross-attention, and the 1D statistical modality is then incorporated through a sample-adaptive selective gate. On UAMR-ShiftBench, SCP-TriCA achieves 95.33% in-distribution accuracy and 74.59% simulated OOD average, outperforming the strongest baseline by 5.12 percentage points, and reaches 91.14% and 94.86% on the two sea-trial subsets, exceeding the best baseline by 15.71 and 23.00 percentage points respectively. Ablation results confirm that the gains stem from modality complementarity and the hierarchical fusion design. Code and models are available at https://github.com/ronglaiqian/UAMR-ShiftBench.

13:00 JST研究/論文

連続ウェアラブル生理学を使用した複数評価者による疼痛評価のイベント整合分析

痛みの評価は患者、看護師、臨床医によって異なりますが、ほとんどの計算アプローチは単一のグラウンドトゥルースラベルを前提としており、誰が評価を行っているかを事実上無視しています。私たちは、まばらな評価者固有の痛みの評価を離散的な痛みの変化イベントに変換し、連続的なウェアラブル生理学的信号をこれらのイベントに調整して、全体を通して評価者のアイデンティティを維持する、評価者対応のイベント調整フレームワークを導入します。このフレームワークは、脊椎関連疼痛処置中に収集されたマルチモーダルウェアラブルデータに適用され、評価者グループ間の実質的な不一致を特定し、報告された痛みの増加に先立って評価者に依存する生理学的差異の予備的、探索的証拠を提供します。これらの発見は、痛みと生理学的関係は評価者によって不変ではない可能性があり、評価者間で評価を集約すると意味のある生理学的パターンが隠蔽される可能性があることを示唆しています。したがって、評価者を意識した、イベントに合わせた視点は、現実世界の臨床疼痛評価におけるウェアラブル データを解釈するための有望な方向性となります。

原文 (English)

Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology

Pain is assessed differently by patients, nurses, and clinicians, yet most computational approaches assume a single ground-truth label - effectively ignoring who is doing the rating. We introduce a rater-aware, event-aligned framework that converts sparse, rater-specific pain ratings into discrete pain-change events and aligns continuous wearable physiological signals to these events, preserving rater identity throughout. Applied to multimodal wearable data collected during spine-related pain procedures, the framework identifies substantial disagreement across rater groups and provides preliminary, exploratory evidence of rater-dependent physiological differences preceding reported pain increases. These findings suggest that pain-physiology relationships may not be rater-invariant, and that aggregating assessments across raters may mask meaningful physiological patterns. A rater-aware, event-aligned perspective is therefore a promising direction for interpreting wearable data in real-world clinical pain assessment.

13:00 JST研究/論文

目に見えない電極生成によるEEG空間超解像度のための座標クエリ可能な神経場の再構成

実際の導入における EEG 空間超解像 (EEGSR) は、ランダムなチャネル欠落、不安定な電極品質、接触不良やデバイスのばらつきによる可視チャネル パターンの変化などの課題にさらされています。既存の EEGSR メソッドのほとんどは、事前定義された入出力レイアウトの下で固定の低から高へのチャネル マッピングを学習するため、欠落しているチャネルがテスト時に変化すると脆弱になります。この論文では、部分的に観察されたサポート チャネルから共有条件付き頭皮フィールドを学習するものとして EEGSR を再定式化します。具体的には、位置誘導エンコーダーは観察されたEEGチャネルとその座標を潜在条件に要約し、条件付き暗黙的神経表現デコーダーは、所望の電極座標でこの条件をクエリすることによってターゲットEEG信号を再構成します。推論中、モデルは利用可能な EEG サポートとクエリされた座標から目に見えない電極信号を直接再構築します。デコーダ上のエンコードされた潜在表現の制約を強化し、それによって観察されたチャネルと一致するより安定した頭皮フィールドを構築するために、混合電極状態下で忠実度を維持するチャネル破損トレーニング戦略をさらに導入します。複数のEEGデータセットにわたる広範な実験により、ランダムな欠落チャネル再構成と厳密な目に見えない電極信号生成の両方に対するフレームワークの有効性が実証されています。特に、AAD の厳密なホールドアウト電極設定の下では、私たちの方法は、最も強いベースラインよりも NMSE を 37.5% 削減し、SNR を 2.12 dB 改善し、トレーニング中に決して露出しない電極位置で信号を合成する能力を示しています。

原文 (English)

Coordinate-Queryable Neural Field Reconstruction for EEG Spatial Super-Resolution with Unseen-Electrode Generation

EEG spatial super-resolution (EEGSR) in real deployments is challenged by random channel missingness, unstable electrode quality, and changing visible-channel patterns caused by bad contacts or device variability. Most existing EEGSR methods learn a fixed low-to-high channel mapping under pre-defined input-output layouts, which makes them brittle when missing channels vary at test time. In this paper, we reformulate EEGSR as learning a shared conditional scalp field from partially observed support channels. Specifically, a position-guided encoder summarizes the observed EEG channels and their coordinates into a latent condition, and a conditional implicit neural representation decoder reconstructs target EEG signals by querying this condition at desired electrode coordinates. During inference, the model directly reconstructs unseen electrode signals from the available EEG support and the queried coordinates. To strengthen the constraint of the encoded latent representation on the decoder and thereby construct a more stable scalp field consistent with the observed channels, we further introduce a fidelity-preserving channel corruption training strategy under mixed electrode states. Extensive experiments across multiple EEG datasets demonstrate the effectiveness of our framework for both random missing-channel reconstruction and strict unseen-electrode signal generation. Notably, under the strict held-out-electrode setting on AAD, our method reduces NMSE by 37.5\% and improves SNR by 2.12 dB over the strongest baseline, showing its ability to synthesize signals at electrode locations never exposed during training.

13:00 JST研究/論文

拡散ベースの視覚条件付き音声強調のための視聴覚コントラスト調整

AVSE (Audio-Visual Speech Enhancement) は、唇の動きなどの視覚的な手がかりを利用して、騒がしい環境で音声を復元します。最近の研究では、拡散ベースの教師なし AVSE が導入されました。この AVSE では、交差注意を介して視覚的特徴に条件付けされた音声拡散モデルがトレーニングされ、事後サンプリング ベースの音声強調のためのデータ駆動型事前学習として使用されます。オーディオのみの対応物よりもパフォーマンスが期待できるにもかかわらず、融合におけるクロスモーダル調整を明示的に強制することの影響は依然として不明です。この研究では、事後サンプリングのフレームワークを変更せずに、視覚情報のより強力な使用を促進するために、対照的な視聴覚損失で拡散トレーニングの目的を強化することを提案します。一致したテストデータと不一致なテストデータにわたる実験では、低い SNR で最大の利得が得られ、干渉抑制、信号再構成、知覚品質が一貫して向上していることが示されています。コードは https://github.com/cexauce/AV-CA-DiffUSE で入手できます。

原文 (English)

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE, where a speech diffusion model conditioned on visual features via cross-attention is trained and used as a data-driven prior for posterior sampling-based speech enhancement. Despite promising performance over its audio-only counterpart, the impact of explicitly enforcing cross-modal alignment in the fusion remains unclear. In this work, we propose to augment the diffusion training objective with a contrastive audio-visual loss to encourage stronger use of visual information while keeping the posterior sampling framework unchanged. Experiments across matched and mismatched test data show consistent improvements in interference suppression, signal reconstruction, and perceptual quality, with the largest gains at low SNRs. Code is available at https://github.com/ cexauce/AV-CA-DiffUSE

13:00 JST研究/論文

マルコフ論理ネットワークによって定義されたランダムな色の有向グラフ

マルコフ論理ネットワーク (MLN) は、統計リレーショナル人工知能で、任意の有限領域 $D$ の領域 $D$ を持つ可能世界のセットにおける確率分布を定義するために使用される確率的関係モデルです。 MLN は、非負の実数である重みが関連付けられたソフト制約で構成されます。この研究では、プロパティ $P(x)$ と関係 $R(x, y)$ について話す言語を検討します。 $P(x)$ と $R(x, y)$ のすべてのブール値の組み合わせがソフト制約 (関連する重み付き) である MLN を検討します。 $n$ がドメインのサイズ (カーディナリティ) を表すものとします。すべての重みの選択について、重みが $1/n$ でスケーリングされる場合、すべての 1 次文 $\varphi$ について、$\varphi$ が保持する確率は $n \to \infty$ として 0 または 1 のいずれかになる傾向があることを示します。つまり、一次論理の 0-1 の法則が成り立ちます。さらに、限界確率は重みに依存しません。代わりに MLN の標準セマンティクスを使用する場合、重みがスケーリングされない場合、制限の動作はより複雑になり、重みに依存します。スケーリングされていない重みを使用すると、重みに応じて質的に異なる 7 つのケースが得られます。場合によっては、一次論理の 0-1 の法則が存在する場合もあれば、存在しない場合もありますが、依然として収束の法則が存在する可能性があります。一次文の漸近確率に対する重みの影響は、7 つのケースのうちの 1 つから別のケースへの突然の「相転移」の形で現れる可能性があります。収束則の存在は、大規模な領域での推論にプラスの影響を及ぼします。

原文 (English)

Random coloured digraphs defined by a Markov logic network

A Markov Logic Network (MLN) is a probabilistic relational model used in Statistical Relational Artificial Intelligence for defining a probability distribution on the set of possible worlds with domain $D$ for an arbitrary finite domain $D$. An MLN consists of soft constraints with associated weights which are nonnegative real numbers. In this study we consider a language speaking about a property $P(x)$ and a relation $R(x, y)$. We consider an MLN for which every Boolean combination of $P(x)$ and $R(x, y)$ is a soft constraint (with associated weight). Let $n$ denote the size (cardinality) of the domain. We show that, for every choice of weights, if the weights are scaled by $1/n$ then, for every first-order sentence $\varphi$, the probability that $\varphi$ holds tends to either 0 or 1 as $n \to \infty$; that is, a 0-1 law for first-order logic holds. Morover, the limit probability does {\em not} depend on the weights. If we instead use the standard semantics of MLNs, in the case of which the weights are {\em not} scaled, then the limit behaviour is more complicated and {\em depends} on the weights. With unscaled weights we get 7 qualitatively different cases which depend on the weights. In some cases we have a 0-1 law for first-order logic, in some cases not, but we may still have a convergence law. The influence of the weights on the asymptotic probability of a first-order sentence may be in the form of a sudden ``phase transition'' from one of the 7 cases to another. The presence of a convergence law has positive implications for inference on large domains.

13:00 JST研究/論文

法的推論は弁護士ではない:プロセクが司法にアクセスするための法的ベンチマークを再考する

Legal AI のベンチマーク調査では、大規模な言語モデルによって、法的権利を理解して行使するために弁護士に相談できない人々など、司法へのアクセスが向上するという仮定が頻繁に呼び出されます。現在のベンチマークは、モデルのパフォーマンスの上限を測定する、法律専門家によってすでに前処理された入力に対して法的推論を評価するため、この仮定をサポートする機能が備わっていないと主張します。正義へのアクセスは下限に依存します。つまり、プロンプトに騒々しい物語、埋もれた事実、省略、民俗法的な仮定、および表面レベルの誤りが含まれる可能性があるプロの訴訟当事者からの入力があった場合に、モデルがどのように機能するかです。これらの劣化は、ロングコンテキストの感度、過小仕様、幻覚、活字の乱れなど、一般的な機械学習の文献で LLM が劣化することが知られている条件に匹敵します。私たちは、散文文献からの証拠をこの一連の機械学習研究と結びつけ、法的ベンチマークである LEXam での小さな摂動実験を提示して、これら 2 つの限界間のギャップを説明します。モデル開発が上限のみを測定するベンチマークに焦点を当て続けた場合、このギャップは隠れたままになるか、さらに拡大する可能性があります。最後に、法的 AI に関する司法へのアクセスの主張が実証的に検証できるよう、専門的な入力の下で堅牢性を直接測定する法的ベンチマークを求めます。

原文 (English)

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice

Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for people who cannot access lawyers in order to understand and exercise their legal rights. We argue that current benchmarks are not equipped to support this assumption because they evaluate legal reasoning over inputs that have already been preprocessed by legal experts, which measures the upper bound of model performance. Access to justice depends on a lower bound: how models perform when inputs come from pro se litigants, whose prompts may contain noisy narratives, buried facts, omissions, folk-legal assumptions, and surface-level errors. These degradations are comparable to conditions under which LLMs are known to degrade in the general machine learning literature, including long-context sensitivity, underspecification, hallucination, and typographical perturbations. We connect evidence from pro se literature with this body of machine learning research and present a small perturbation experiment on LEXam, a legal benchmark, to illustrate the gap between these two bounds. If model development continues to focus on benchmarks that measure only the upper bound, this gap may remain hidden or even widen. We conclude by calling for legal benchmarks that directly measure robustness under pro se-like inputs so that access-to-justice claims about legal AI can become empirically testable.

13:00 JST研究/論文

LOLA のランタイム検証とモデルベースの診断のための統合フレームワーク

ストリーム仕様言語 LOLA 内でランタイム検証とモデルベースの診断を統合する統合フレームワークを紹介します。このアプローチでは、システムの説明、コンポーネントの健全性状態、および観察を単一のストリームベースの形式にエンコードすることで、別個のツールチェーンを必要とせずに、障害検出と並行して継続的なオンライン障害位置特定が直接可能になります。このフレームワークは、時間不変故障と一時故障の両方をサポートし、非決定的な観測にも自然に対応します。

原文 (English)

A Unified Framework for Runtime Verification and Model-Based Diagnosis in LOLA

We present an integrated framework that unifies runtime verification and model-based diagnosis within the stream specification language LOLA. By encoding system descriptions, component health states, and observations into a single stream-based formalism, the approach enables continuous, online fault localization directly alongside fault detection, without requiring separate toolchains. The framework supports both time-invariant and transient faults, and naturally accommodates nondeterministic observations.

13:00 JST研究/論文

オフライン推論トレーニングの重み空間幾何学

オフライン強化学習損失 (RFT、RIFT、DFT、オフライン GRPO、DPO) は、大規模な教師から小規模な生徒へ推論を抽出するために広く使用されており、通常は下流の精度のみで比較されます。それらが機構的に異なるのか、それとも同様の重み更新に収束するのかを尋ねます。アテンションのみの LoRA を使用して、単一のベース モデル (Qwen3-4B) からの同一の数学ロールアウトで 6 つのメソッド (SFT、RFT、DFT、RIFT、オフライン GRPO、DPO) をトレーニングし、コサイン類似度、主角部分空間分析、線形モード接続性、および CKA を介して結果のデルタを分析します。 (i) SFT、RFT、および RIFT はほぼ同一線上の重みデルタ (コサイン >= 0.97、144 モジュールにわたるトップ 1 主角は中央値約 7 度) と同等の GSM8K 精度 (87 ~ 88%、n=1319; ペアワイズ マクネマー p >= 0.15) を示します。 (ii) DFT は、同じデータを使用しているにもかかわらず、どの報酬重み付け手法よりも方向が大きく異なります。 (iii) オフライン GRPO は、SFT 損失領域内に留まりながら、SFT 方向に直交する実質的なコンポーネントを追加します (グローバルで最大 67%、後期レイヤーで最大約 86%)。 (iv) DPO はほぼ直交部分空間に位置し、モード接続性障壁を示し、後期層の CKA を ~0.46 まで崩壊させます。また、DPO は、GSM8K (93.5%、McNemar p < 10^-9 対他のメソッド) と AIME26 (30.0% 対 3.3-10.0%) の両方で、当社のプロトコルで最高の精度に達します。そのトレーニングでは他のものよりも 10 倍小さい学習率 (標準規約) が使用されるため、更新ノルムと精度のギャップは損失関数とオプティマイザーの選択を合わせて反映し、学習率と一致する DPO の比較は将来の作業に残されます。

原文 (English)

Weight-Space Geometry of Offline Reasoning Training

Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller students, and are typically compared on downstream accuracy alone. We ask whether they are mechanistically distinct or converge to a similar weight update. Training six methods (SFT, RFT, DFT, RIFT, Offline GRPO, DPO) on identical math rollouts from a single base model (Qwen3-4B) with attention-only LoRA, we analyze the resulting deltas via cosine similarity, principal-angle subspace analysis, linear mode connectivity, and CKA. We observe: (i) SFT, RFT, and RIFT have nearly colinear weight deltas (cosine >= 0.97, top-1 principal angle ~7 deg median over 144 modules) and comparable GSM8K accuracy (87-88%, n=1319; pairwise McNemar p >= 0.15); (ii) DFT diverges further in direction than any reward-weighted method despite using the same data; (iii) Offline GRPO adds a substantial component orthogonal to the SFT direction (~67% globally, up to ~86% in late layers) while staying in the SFT loss basin; (iv) DPO sits in a near-orthogonal subspace, shows a mode-connectivity barrier, and collapses late-layer CKA to ~0.46. DPO also reaches the highest accuracy in our protocol on both GSM8K (93.5%, McNemar p < 10^-9 vs. each other method) and AIME26 (30.0% vs. 3.3-10.0%); its training uses a 10x smaller learning rate than the others (the standard convention), so the update-norm and accuracy gaps reflect loss-function and optimizer choices jointly, and a learning-rate-matched DPO comparison is left for future work.

13:00 JST研究/論文

フェデレーションによる因果関係の発見と推論に関する調査

因果構造の発見と因果効果の推論を含む因果推論は、データ主導の意思決定の基礎です。実際には、信頼性の高い因果関係分析のためのデータは多くの場合、機関全体に分散されており、プライバシー規制や通信上の制約により一元化することができません。フェデレーテッド ラーニング (FL) は、生データを共有せずに共同分析を可能にすることでこの問題に対処し、急速に成長するフェデレーテッド 因果発見 (FCD) および推論 (FCI) の分野を生み出します。しかし、この分野の学際的な性質と包括的な調査がないことが、研究者にとって参入障壁となっています。この論文は、多次元の分類法による体系的なレビューを提供することで、そのギャップを埋めます。 FCD ソリューションの基礎となる 3 つの核となる設計上の決定、つまり構造の学習方法、データの分割方法、各当事者が取得する構造知識に基づいて、方法論的パラダイム、フェデレーション トポロジ、構造スコープの 3 つの軸に沿って FCD を整理します。さらに、時間的ダイナミクス、データの異質性、欠損データ、同一でない変数セットなど、主要な実際的な側面を調査します。 FCI については、古典的な重み付け手法から最新の深層生成アーキテクチャまで、目標推定値 (平均対個別/条件付き治療効果) および推定戦略によって手法を分類します。 FCD と FCI を別々に扱う以前の研究とは異なり、我々はそれらの関係を統一されたフェデレーション因果推論パイプラインの補完的な段階として形式化し、FCD が FCI での有効な効果推定に必要な構造的知識を提供します。最後に、プライバシー、通信効率、理論的保証、応用分野に関する彼らの共通の懸念を強調し、将来の研究に向けた未解決の課題を特定して締めくくります。

原文 (English)

A Survey on Federated Causal Discovery and Inference

Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making. In practice, data for reliable causal analysis are often distributed across institutions and cannot be centralized due to privacy regulations or communication constraints. Federated learning (FL) addresses this by enabling collaborative analysis without raw data sharing, giving rise to the rapidly growing field of federated causal discovery (FCD) and inference (FCI). However, the interdisciplinary nature of this field and the absence of a comprehensive survey present barriers to entry for researchers. This paper bridges that gap by providing a systematic review through multi-dimensional taxonomies. Grounded in the three core design decisions underlying any FCD solution, namely how structures are learned, how data are partitioned, and what structural knowledge each party obtains, we organize FCD along three axes: methodological paradigm, federation topology, and structural scope. We further examine key practical dimensions, including temporal dynamics, data heterogeneity, missing data, and non-identical variable sets. For FCI, we categorize methods by target estimand (average versus individualized/conditional treatment effects) and by estimation strategy, from classical weighting methods to modern deep generative architectures. Unlike prior works that treat FCD and FCI separately, we formalize their connection as complementary stages of a unified federated causal reasoning pipeline, where FCD supplies the structural knowledge required for valid effect estimation in FCI. Finally, we highlight their shared concerns regarding privacy, communication efficiency, theoretical guarantees, and application domains, and conclude by identifying open challenges for future research.

13:00 JST研究/論文

継続制御のためのトレーニング可能な非線形接続を備えた低電力アナログ ニューラル ネットワーク

物理ニューラル ネットワークは、アナログ デバイス物理学を直接計算することで低消費電力の機械学習を約束しますが、ほとんどのアーキテクチャでは、非線形デバイス応答がスカラー重みとして機能するように強制されます。 Kolmogorov-Arnold ネットワークからインスピレーションを得て、学習可能な非線形関数を接続に配置し、各物理接続を学習可能な計算要素にします。これらの機能をフィールドでプログラム可能なアナログ アレイ上のアナログ バンドパス フィルターとして実現すると、その利点はタスクに依存し、物理的基盤の滑らかさに起因することがわかります。ネットワークは、ロボット運動学、連続制御、太陽光発電の最大電力点追跡などの滑らかで継続的に値付けされるターゲットを表し、多層パーセプトロンよりもはるかに少ないノードと接続を備えていますが、分類のような決定境界ではパラメーター効率の利点はありません。トレーニングされたネットワークは、定量化された忠実度で約 35,000 の接続にわたってハードウェアに転送され、専用の CMOS 実装は約 30 マイクロワットで動作すると予測されています。メモリスティブの実現はシミュレーションで同じ動作を再現します。これは、特定のデバイスからではなく、接続にトレーニング可能な非線形性を配置することによって利点が得られることを示しています。

原文 (English)

Low-power analogue neural networks with trainable nonlinear connections for continuous control

Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures force nonlinear device responses to act as scalar weights. Inspired by Kolmogorov-Arnold networks, we place trainable nonlinear functions on the connections, making each physical connection a learnable computational element. Realising these functions as analogue band-pass filters on field-programmable analogue arrays, we find that the benefit is task-dependent and follows from the smoothness of the physical basis: the networks represent smooth, continuously valued targets, including robotic kinematics, continuous control, and photovoltaic maximum-power-point tracking, with far fewer nodes and connections than multilayer perceptrons, but offer no parameter-efficiency advantage on classification-like decision boundaries. Trained networks transfer to hardware across approximately 35,000 connections with quantified fidelity, and a dedicated CMOS implementation is projected to operate at approximately 30 microwatts. A memristive realisation reproduces the same behaviour in simulation, indicating that the advantage comes from placing trainable nonlinearity on connections, rather than from a particular device.

13:00 JST画像/動画生成エージェント

Sol ビデオ推論エンジン: 効率的なビデオ生成のためのエージェントネイティブのフルスタック アクセラレーション フレームワーク

最新のビデオ拡散モデルは、スケーリングを通じてより高い生成品質を実現しますが、これにより推論コストも増加します。多くの高速化方法が提案されていますが、中心的な課題は、最も効果的な高速化戦略が非常にインスタンス固有であるということです。つまり、モデル、ハードウェア、推論構成の 1 つの組み合わせでうまく機能するレシピが、別の組み合わせには移行しないことがよくあります。モデルが異なれば、アーキテクチャ、数値感度、注意集中パターンも異なります。推論設定は、空間的および時間的な解像度とビデオの長さが異なり、ハードウェア プラットフォームはメモリ階層、サポートされている数値形式、およびカーネル スループットが異なります。これらの要因により調整の余地が大きくなり、手動によるパフォーマンス エンジニアリングのコストが高くなります。我々は、ビデオ拡散モデルのためのエージェント的、ネイティブ、トレーニング不要のアクセラレーション フレームワークである Sol Video Inference Engine を紹介します。これは、キャッシュ、スパース アテンション、トークン プルーニング、量子化、カーネル フュージョンという 5 つの広く適用可能な手法を、インスタンス固有の最適化のためにエージェント アクセラレーション スタックにまとめています。モデル、ハードウェア プラットフォーム、およびサービス構成によって定義された具体的なデプロイメント ターゲットに対して、並列スキル エージェントが各手法の実装を最適化し、エージェント インテグレーターがそれらをグローバル アクセラレーション スタックに構成し、人間のバリデーターが生成品質に関するフィードバックを提供します。サイズとアーキテクチャが異なる 3 つのビデオ モデル (64B Cosmos3-Super、22B LTX-2.3、および 2B SANA-Video) でこのワークフローをインスタンス化します。人的労力をほとんどかけることなく、フルスタックはほぼロスレスの VBench 品質を維持しながら 2 倍を超えるエンドツーエンド アクセラレーションを達成し、ビデオ拡散アクセラレーションに対するエージェント フレームワークの有効性を実証しています。

原文 (English)

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation

Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have been proposed, a central challenge is that the most effective acceleration strategy is highly instance-specific: a recipe that works well for one combination of model, hardware, and inference configuration often does not transfer to another. Different models vary in architecture, numerical sensitivity, and attention concentration patterns. Inference settings differ in spatial and temporal resolution and video duration, while hardware platforms differ in memory hierarchy, supported numerical formats, and kernel throughput. These factors create a large tuning space, making manual performance engineering costly. We present Sol Video Inference Engine, an agentic, native, training-free acceleration framework for video diffusion models. It organizes five broadly applicable techniques, cache, sparse attention, token pruning, quantization, and kernel fusion, into an agentic acceleration stack for instance-specific optimization. For a concrete deployment target defined by a model, hardware platform, and serving configuration, parallel skill agents optimize the implementation of each technique, an agent integrator composes them into a global acceleration stack, and a human validator provides feedback on generation quality. We instantiate this workflow on three video models with different sizes and architectures: 64B Cosmos3-Super, 22B LTX-2.3, and 2B SANA-Video. With little human effort, the full stack achieves more than 2x end-to-end acceleration while maintaining near-lossless VBench quality, demonstrating the effectiveness of the agent framework for video diffusion acceleration.

13:00 JST研究/論文

JEDEL: 初期段階の創薬のためのゼロショット DNA エンコード ライブラリ設計

我々は、活性リガンドの三次元ファーマコフォア表現から直接、合成に対応した DNA エンコード ライブラリー (DEL) を生成するためのフレームワークである JEDEL を紹介します。 JEDEL は、ファーマコフォアの相互作用パターンを実用的でスケーラブルな合成指示にマッピングする最初のモデルであり、潜在的に数百万の分子を含むターゲット ライブラリの設計を可能にします。下流の合成計画を必要とする仮想化合物を生成する既存の生成アプローチとは異なり、JEDEL は購入可能なビルディング ブロックと検証済みの反応の範囲内で動作し、すべての出力が構築によって実験的に実現可能であることを保証します。 JEDEL は、ファーマコフォアの形状と分子構造の間の予測的アラインメントを学習し、これを大規模なコンビナトリアル合成ルートに解読します。 18 のタンパク質ターゲットにわたって、ターゲット特異的な再トレーニングを行わずに、予測結合親和性、ファーマコフォア回収率、およびサンプル効率においてランダムおよび多様性ベースのベースラインを上回る、集中的なライブラリーを生成します。 JEDEL を使用すると、仮想分子生成から実験的に展開可能なライブラリ設計への移行が可能になります。

原文 (English)

JEDEL: Zero-Shot DNA-Encoded Library Design for Early-Stage Drug Discovery

We present JEDEL, a framework for generating synthesis-ready DNA-encoded libraries (DELs) directly from three-dimensional pharmacophore representations of active ligands. JEDEL is the first model to map pharmacophore interaction patterns to actionable, scalable synthesis instructions, enabling the design of targeted libraries comprising potentially millions of molecules. Unlike existing generative approaches that produce virtual compounds requiring downstream synthesis planning, JEDEL operates within the space of purchasable building blocks and validated reactions, ensuring that every output is experimentally realizable by construction. JEDEL learns a predictive alignment between pharmacophore geometry and molecular structure and decodes this into combinatorial synthesis routes at scale. Across 18 protein targets, it generates focused libraries that outperform random and diversity-based baselines in predicted binding affinity, pharmacophore recovery, and sample efficiency, without target-specific retraining. JEDEL enables a shift from virtual molecule generation to experimentally deployable library design.

13:00 JST研究/論文

物理的に制約された MCMC と化学情報に基づいたガウス過程を相乗して反応ネットワークを発見

離散的な反応トポロジーと連続的な速度論的パラメーターは密接に結合しているため、まばらでノイズの多い化学時系列データから解釈可能な支配方程式を抽出することは依然として困難です。スパイクアンドスラブトポロジーサンプリング、ハード保存、熱力学スクリーニング、パラメーター校正と実験計画のための化学情報ガウスプロセス(CIGP)残差モデルを組み合わせた再現可能なグレーボックスワークフローであるPC-MCMC-CIGPを紹介します。方法論的な貢献は、新しい MCMC または GP ファミリー単独ではありません。むしろ、これらのコンポーネントを物理的に制約されたワークフローに統合し、明確な不確実性を考慮した取得選択を行うことです。 H2 + Br2 ベンチマークでは、制約付きサンプラーは、実験で基本的なラジカル経路を欺瞞的な現象学的適合から区別します。スチレンのエポキシ化では、CIGP 最適化ループにより、報告されている GP-BO ベースラインよりも最終収率が 12.5% 向上しました。新しい 10 シード取得研究では、EI、GWU、PC-EI、不確実性サンプリング、不一致探索、およびランダム検索には異なるトレードオフがあることが示されています。PC-EI は低利回りの BO 提案を大幅に削減する一方、EI スタイルの基準は最強の最終利回りパフォーマンスをもたらします。

原文 (English)

Synergizing Physically Constrained MCMC and Chemical-Informed Gaussian Processes for Reaction Network Discovery

Extracting interpretable governing equations from sparse, noisy chemical time-series data remains difficult because discrete reaction topology and continuous kinetic parameters are tightly coupled. We present PC-MCMC-CIGP, a reproducible gray-box workflow that combines spike-and-slab topology sampling, hard conservation and thermodynamic screening, and a Chemical-Informed Gaussian Process (CIGP) residual model for parameter calibration and experimental design. The methodological contribution is not a new MCMC or GP family in isolation; rather, it is the integration of these components into a physically constrained workflow with explicit uncertainty-aware acquisition choices. On the H2 + Br2 benchmark, the constrained sampler distinguishes elementary radical pathways from deceptive phenomenological fits in our experiments. On styrene epoxidation, the CIGP optimization loop improves final yield by 12.5% over the reported GP-BO baseline. A new 10-seed acquisition study shows that EI, GWU, PC-EI, uncertainty sampling, discrepancy hunting, and random search have different trade-offs: PC-EI substantially reduces low-yield BO suggestions, while EI-style criteria give the strongest final-yield performance.

13:00 JST研究/論文

オープンセットシナリオにおけるドメイン汎化を強化するための二元論的メタ学習の探求

ドメインの一般化は、複数のソース ドメインから学習して、目に見えないターゲット ドメインに一般化します。ただし、ソースとターゲットの間のラベルの不一致という現実的なケースが無視されることがよくあります。次に、未知のドメイン内の未知のクラスを認識するために、開集合ドメイン一般化が提案されます。簡単なアプローチでは、1 対全分類器をトレーニングして各クラスを分離し、外れ値を未知として検出します。しかし、少数の正のサンプルと多数の負のサンプルの間の不均衡により、決定境界が正の方向に偏り、モデルは、たとえ目に見えない領域の既知のクラスからであっても、分布外のデータを過剰に拒否することになります。この論文では、Dualistic MEta-learning with Joint DomaIn-Class Matching (MEDIC) と呼ばれる新しいメタ学習戦略を提案します。これは、ドメイン間およびクラス間のタスク分割に向けた暗黙的な勾配マッチングを考慮して、ドメインとクラスの両方でバランスの取れた最適な境界を見つけます。実験結果は、MEDIC がオープンセットのシナリオで従来の方法より優れているだけでなく、競合するクローズセット汎化能力を維持していることを示しています。

原文 (English)

Exploring Dualistic Meta-Learning to Enhance Domain Generalization in Open Set Scenarios

Domain generalization learns from multiple source domains to generalize to unseen target domains. However, it often neglects the realistic case of label mismatch between source and target. Open set domain generalization is then proposed to recognize unseen classes in unseen domains. A simple approach trains one-vs-all classifiers to separate each class and detect outliers as unknown. Yet, the imbalance between few positive samples and many negative samples skews the decision boundary towards the positive ones, leading the model to over-reject out-of-distribution data, even from known classes in unseen domains. In this paper, we propose a novel meta-learning stategy called dualistic MEta-learning with joint DomaIn-Class matching (MEDIC), which considers implicit gradient matching towards inter-domain and inter-class task splits simultaneously to find optimal boundaries balanced for both domains and classes. Experimental results show that MEDIC not only outperforms prior methods in open set scenarios, but also maintains competitive close set generalization ability.

13:00 JSTLLM/生成AIGPT / ChatGPTNVIDIA

VeriPilot: LLM を利用した Verilog デバッグ フレームワーク

Verilog デバッグは、依然としてデジタル回路設計において最も時間のかかる段階の 1 つです。大規模言語モデル (LLM) の最近の進歩により、自動デバッグが可能になりました。ただし、既存のアプローチのほとんどは、エンドツーエンドの方法でテスト出力とコンパイラのフィードバックのみに依存しているため、複雑なバグに対する効果は限られています。主な課題は、エラーの根本原因が観察可能な出力から遠く離れている可能性があり、LLM がコード内の長い依存関係チェーンを追跡することが困難になることです。この課題は、コンテキストの長さが長いため効率的な推論が妨げられる大規模なコードベースではさらに悪化します。これらの制限に対処するために、私たちは、ゴールデン リファレンス モデルを活用してきめ細かいバグの位置特定と修復を可能にする、LLM を利用したデバッグ フレームワークである VeriPilot を提案します。 VeriPilot は、LLM ベースの分析を通じて Verilog 設計とそれに対応するゴールデン モデルの間の内部変数セマンティクスを調整することで、出力レベルの比較を超えています。次に、静的解析から得られたコントロール データ フロー グラフ (CDFG) を使用して段階的な信号トレースを実行し、疑わしいコード領域の最小限のセットと、ゴールデン モデルからの正しい対応部分を特定します。これらの構造化された洞察は、その後、推論と自動コード修復をガイドするために LLM に提供されます。 NVIDIA の包括的 Verilog 設計問題 (CVDP) ベンチマークの実験結果は、VeriPilot が GPT-4o の修復成功率を 54.3\% から 85.71\% に向上させ、複雑な Verilog 設計のバグ位置特定の精度と修復効果の両方を大幅に向上させることを示しています。ソース コードとベンチマークは、Github https://github.com/YihanWn/VeriPilot.git で公開されています。

原文 (English)

VeriPilot: An LLM-Powered Verilog Debugging Framework

Verilog debugging remains one of the most time-consuming stages in digital circuit design. Recent advances in Large Language Models (LLMs) have enabled automated debugging; however, most existing approaches rely solely on test outputs and compiler feedback in an end-to-end manner, limiting their effectiveness on complex bugs. A key challenge is that the root cause of an error may be far removed from its observable outputs, making it difficult for LLMs to trace long dependency chains in code. This challenge is further exacerbated in large codebases, where long context lengths hinder efficient reasoning. To address these limitations, we propose VeriPilot, an LLM-powered debugging framework that leverages golden reference models to enable fine-grained bug localization and repair. VeriPilot goes beyond output-level comparison by aligning internal variable semantics between the Verilog design and its corresponding golden model through LLM-based analysis. It then performs step-by-step signal tracing using Control-Data-Flow Graphs (CDFGs) derived from static analysis, identifying a minimal set of suspicious code regions along with their correct counterparts from the golden model. These structured insights are subsequently provided to the LLM to guide reasoning and automated code repair. Experimental results on the Comprehensive Verilog Design Problems (CVDP) benchmark from NVIDIA demonstrate that VeriPilot improves the repair success rate of GPT-4o from 54.3\% to 85.71\%, significantly enhancing both bug localization accuracy and repair effectiveness for complex Verilog designs. The source code and benchmark are publicly available at Github https://github.com/YihanWn/VeriPilot.git.

13:00 JSTエージェントロボティクス

信頼性の高い自律システムのエンジニアリング: 課題と解決策

信頼性の高い自律システムのエンジニアリングは、コンピューター サイエンスにおける重要かつ成長中のテーマです。自律システムが普及するにつれて、自律システムを確実に構築するための使いやすい技術の重要性が増しています。このワークショップレポートは、2024年6月10日から14日まで開催されたローレンツセンターワークショップ「信頼性の高い自律システムのエンジニアリング」(ERAS)での議論を取りまとめ、拡張したものです。このワークショップは、自律システムのための形式手法に関するワークショップ(FMAS)と信頼性の高いエンジニアリング自律システムのためのエージェントとロボットに関するワークショップ(AREA)の主催者によって共催されました。この会合には、FMAS および AREA コミュニティのメンバー、業界関係者、自律システムが特有のエンジニアリング課題を引き起こす分野の代表者が集まりました。このワークショップでは、自律システムの検証と妥当性確認の技術という 3 つの主要な研究トピックに焦点を当てました。現実世界の自律システムをエンジニアリングする。安全な自律システムのためのソフトウェア アーキテクチャ。その主な成果は、これらの分野における課題のカタログであり、最も重要なことに、解決策への道筋です。一部の課題は、学界ではよく知られているものの、実際にはまだ定期的に使用されていない手法ですでに取り組むことができます。その他の課題は未解決のままであり、さらなる研究が必要です。このロードマップは、将来の研究と産業協力をサポートすることを目的としています。

原文 (English)

Engineering Reliable Autonomous Systems: Challenges and Solutions

Engineering reliable autonomous systems is an important and growing topic in computer science. As autonomous systems become more prevalent, easy-to-use techniques for building them reliably are increasingly important. This workshop report captures and expands on the discussions at the Lorentz Center Workshop "Engineering Reliable Autonomous Systems" (ERAS), held from 10 to 14 June 2024. The workshop was co-organised by the organisers of the Workshop on Formal Methods for Autonomous Systems (FMAS) and the Workshop on Agents and Robots for reliable Engineered Autonomy (AREA). It brought together members of the FMAS and AREA communities, industry practitioners, and representatives from sectors where autonomous systems pose distinctive engineering challenges. The workshop focused on three main research topics: techniques for verification and validation of autonomous systems; engineering real-world autonomous systems; and software architectures for safe autonomous systems. Its main outcome is a catalogue of challenges in these areas and, most importantly, a pathway to solutions. Some challenges can already be tackled by techniques that are well known in academia but have not yet become regularly used in practice. Other challenges remain unresolved and require further research. This roadmap is intended to support future research and industrial collaboration.

13:00 JST研究/論文

デュアルブランチ スパイキング ニューラル ネットワークによるニューロモーフィック音声強化

スパイキング ニューラル ネットワーク (SNN) ベースのニューロモーフィック音声強調は、そのエネルギー効率により有望なパラダイムとして浮上していますが、バイナリ アクティベーションと適切に設計されたネットワーク アーキテクチャの欠如により、依然として古典的な人工ニューラル ネットワーク (ANN) ベースのアプローチよりも性能が劣っています。この制限を克服するために、GSU-DBNet と呼ばれる、ゲート スパイキング ユニット (GSU) を備えた新しいデュアル ブランチ スパイキング ニューラル ネットワーク アーキテクチャを提案します。具体的には、GSU-DBNet は音声の振幅スペクトルと複素スペクトルを同時にモデル化し、対応する振幅と複素スペクトル マスクを予測します。一方、デュアルパス GSU モジュールが採用され、時間情報と周波数情報を活用して時空間特徴表現を強化します。人気のベンチマーク データセットでの実験では、GSU-DBNet がわずか 394,000 のパラメータで PESQ スコア 3.04 を達成し、代表的な ANN ベース モデルのパラメータの 4.5% ~ 10.6% のみを使用しながら、既存の SNN ベースの手法を上回るパフォーマンスを示していることが示されています。

原文 (English)

Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks

Spiking neural network (SNN)-based neuromorphic speech enhancement has emerged as a promising paradigm due to its energy efficiency, yet it still underperforms classical artificial neural network (ANN)-based approaches owing to binary activations and the lack of well-designed network architectures. To overcome this limitation, we propose a novel dual-branch spiking neural network architecture equipped with a gated spiking unit (GSU), termed GSU-DBNet. Specifically, GSU-DBNet simultaneously models the speech magnitude spectrum and complex spectrum, predicting the corresponding magnitude and complex spectral masks. Meanwhile, a dual-path GSU module is adopted to exploit temporal and frequency information for enhanced spatiotemporal feature representation. Experiments on a popular benchmark dataset show that GSU-DBNet achieves a PESQ score of 3.04 with only 394K parameters, outperforming existing SNN-based methods while using only 4.5%--10.6% of the parameters of representative ANN-based models.

13:00 JST画像/動画生成

聞くことで VLM のビジョンが明確になります

最近の研究では通常、回答側トークンの注意分布を使用して視覚と言語の一貫性を評価します。ただし、最も注目度の高い領域が、意図したセマンティック トークンと常に一致するとは限らないことが観察されています。これはおそらく、以前に生成された応答トークンからの言語事前分布が蓄積され、視覚的な注意と不一致となるデコード ドリフトに起因します。以前の回答トークンからの事前確率に加えて、モダリティ境界マーカーなどの構造トークンがコンテキスト全体を包含し、ターゲットに関係のない領域への高い注目を生成する可能性があることがわかりました。これらの歪みを回避し、大規模な VLM の一貫性評価を提供するために、プロンプト側のセマンティクスを採用し、Prompt-Vision Token Activation Map (PV-TAM) を提案します。 PV-TAM にはさらに、モダリティ境界マーカーによって引き起こされる系統的なバイアスを除去するフィルターが組み込まれています。アクティベーション強度を無視してマスクのみを介して重複を評価する従来の方法とは異なり、当社のメトリクスは注意のピーク分布を利用して、プロンプトと視覚領域間の整合性を測定します。実験では、PV-TAM は、さまざまなデータセットの回答側のベースラインよりもアテンション ベースと IoU スタイルのローカリゼーション メトリックの両方を一貫して改善しました。

原文 (English)

Listening makes Vision Clear for VLMs

Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens. However, we observe that highest attention regions are not always consistent with the intended semantic token. This probably stems from decoding drift, where language priors from previously generated answer tokens accumulate and mismatch with visual attention. Besides the priors from previous answer tokens, we find that structural tokens, e.g., modality boundary markers, may encompass the entire context and generate high attention to areas unrelated to the target. To avoid these distortions and provide consistency evaluation for large VLMs, we adopt prompt-side semantics and propose Prompt-Vision Token Activation Map (PV-TAM). PV-TAM further incorporates a filter to remove systematic bias induced by modality boundary markers. Unlike traditional methods that evaluate overlap solely through masks while ignoring activation intensity, our metrics leverage the peak distribution of attention to measure the alignment between prompts and visual regions. In experiments, PV-TAM consistently improves both attention-based and IoU-style localization metrics over answer-side baselines on various datasets.

13:00 JSTLLM/生成AIエージェント

LLMエージェント社会における創発的な関係秩序:集団的影響から権威階層化まで

フェイ・シャオトンの差別的秩序パターンは、農村社会が自己中心的で関係性に段階があり、社会的距離が離れると協力が弱まるという特徴を持っています。文化的に特殊なものとして扱われることが多いものの、そのメカニズムの基礎は依然として十分に運用されておらず、これまでの LLM ベースのシミュレーションは主に長期的な社会構造ではなく、短期的な調整を扱っていました。私たちは、感情制御理論、社会的アイデンティティ理論、およびデュルケミアンの集合的感情に基づいたマルチエージェントフレームワークであるCAREB-MASを提案します。エージェントは、感情-倫理-信念の連鎖を通じて推論し、動的に進化する自己中心的なアイデンティティを維持しますが、マクロ環境は、個々の生産、好みに基づく割り当て、および最小限の対話プロトコルのみを指定します。長期的なシミュレーションを通じて、エージェントは 5 つの中心的な差異秩序現象、つまり安定した労働専門化、関西に基づく経済倫理、協力関係の衰退、新興の関係的権威、氏族ベースの中心部と周縁部の階層化を自発的に再現します。これらのパターンは、生産構造とともに血族中心の統合からより大きな機能的相互依存へと移行します。広範な実験結果は、社会構造と変化を研究するための学際的なフレームワークを提供する LLM ベースのマルチエージェント シミュレーションにより、差分秩序を一般的な社会メカニズムの構造に敏感な創発的な結果として解釈することをサポートしています。

原文 (English)

Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification

Fei Xiaotong's Differential Order Pattern characterizes rural society as egocentric and relationally graded, with cooperation attenuating over social distance. Although often treated as culturally specific, its mechanistic basis remains under-operationalized, and prior LLM-based simulations have mainly addressed short-term coordination rather than long-horizon social structure. We propose CAREB-MAS, a multi-agent framework grounded in Affect Control Theory, Social Identity Theory, and Durkheimian collective affect. Agents reason through an emotion-ethics-belief chain and maintain dynamically evolving egocentric identities, while the macro environment specifies only individual production, preference-based allocation, and minimal interaction protocols. Across long-horizon simulations, agents spontaneously reproduce five core Differential Order phenomena: stable labor specialization, guanxi-based economic ethics, relational decay of cooperation, emergent relational authority, and clan-based center-periphery stratification. These patterns shift with production structure from kin-centered integration toward greater functional interdependence. Extensive experiment results support interpreting Differential Order as a structure-sensitive emergent outcome of general social mechanisms, with LLM-based multi-agent simulation providing an interdisciplinary framework for studying social structure and change.

13:00 JSTエージェント

信頼できる AI の有効性を示す暗号証明書

私たちは、エージェント AI システムの有効性を示す暗号証明書を提案します。中心となるアイデアは、正当性またはポリシー条件を論理述語として正式に指定し、この述語を多項式制約上の証人確認問題にコンパイルし、簡潔な暗号証明システム (およびオプションでゼロ知識) を使用して条件が成立することを証明することです。これは、ソース コードの正式な検証と暗号化認証の間の中間点を提供します。エージェントのアクションには、検証者がエージェントを信頼したり計算を再実行したりする必要がなく、合意された正式なポリシーを満たしていることを独立してチェックできる証明を伴うことができます。アプローチの概要を高レベルで説明し、核となる数学的変換を示し、提案を証明力のあるコード、zkVM、形式的手法、およびエージェント ガバナンスに関連付け、完全な実装が回答する必要がある仕様、監査、および展開に関する質問に注意します。

原文 (English)

Cryptographic certificates of validity for trustworthy AI

We propose cryptographic certificates of validity for agentic AI systems. The core idea is to formally specify a correctness or policy condition as a logical predicate, compile this predicate to a witness-checking problem over polynomial constraints, and use a succinct cryptographic proof system (and optionally zero-knowledge) to certify that the condition holds. This offers a middle ground between formal verification of source code, and cryptographic authentication. An agent's action can be accompanied by an independently checkable proof that it satisfies an agreed formal policy, without requiring the verifier to trust the agent or to re-execute computation. We outline the approach at a high level, give the core mathematical translation, relate the proposal to proof-carrying code, zkVMs, formal methods, and agent governance, and note the specification, auditing, and deployment questions that a full implementation must answer.

13:00 JST研究/論文

5G を介した XR でのリアルタイム アバター制御のための統合されたセンシングと通信

拡張現実 (XR) は、5G および 6G ネットワークにとって困難なユースケースを示しており、真に没入型のエクスペリエンスを提供するには、高いデータレートと低遅延の通信が必要です。さらに、物理的な動作を仮想世界にシームレスに変換するには、正確なジェスチャ認識と姿勢推定が必要です。ハンドヘルド コントローラーとカメラをベースとした現在の XR インタラクション ソリューションでは、全身のポーズを簡単にキャプチャすることができず、手を自由に使うことができず、良好な視認性と明確な視線が必要です。この研究では、5G ミリ波 (mmWave) 統合センシングおよび通信 (ISAC) 信号と表面筋電図 (sEMG) 信号を組み合わせた XR 用のマルチモーダル センシング アーキテクチャを提案します。 5G ミリ波 ISAC は、コンテンツをヘッドマウント ディスプレイ (HMD) にワイヤレスで配信するために使用できるだけでなく、同じ通信信号を使用してユーザーの身体レベルの大まかなジェスチャやポーズを導き出し、リアルタイムのアバター制御をサポートすることもできます。きめ細かい指レベルのジェスチャを実現するために、当社のアーキテクチャは前腕の筋肉の活動を捕捉する軽量の sEMG センサーを活用しています。両方のモダリティの必要性を説明するために、両方のセンシング技術の評価を示します。ボディ レベル (5G) では、当社のアーキテクチャは、5G NR 標準の標準ビーム管理またはビーム スイープ手順から計算できるビーム ペアあたりの電力 (PPBP) に依存しています。 PPBP ベースのセンシングは、トレーニング中に見られなかったユーザーについて評価した場合、平均 82.2$\pm$5.9% の精度を達成します。きめの細かい指レベルのインタラクションについては、表面筋電図 (sEMG) が強力な識別情報を伝達し、さまざまな動作設定にわたって一貫した有望なパフォーマンスを実現することを示します。したがって、2 つのモダリティを組み合わせることで、既存の 5G 信号を介して身体レベルで、軽量 sEMG センサーを介して指レベルでマルチスケールのジェスチャ認識が可能になり、完全な XR フレームワークが形成されます。

原文 (English)

Integrated Sensing and Communications for Real-time Avatar Control in XR over 5G

Extended Reality (XR) presents a challenging use case for 5G and 6G networks, requiring high data-rates and lowlatency communication to deliver a truly immersive experience. Moreover, in order to seamlessly translate physical actions to the virtual world, accurate gesture recognition and pose estimation are required. Current XR interaction solutions based on handheld controllers and cameras cannot easily capture full-body poses, inhibit the free use of hands, and require good visibility and a clear line of sight. In this work, we propose a multimodal sensing architecture for XR that combines 5G MillimeterWave (mmWave) Integrated sensing and communication (ISAC) and surface electromyography (sEMG) signals. 5G mmWave ISAC cannot only be used to deliver content wirelessly to the Head-mounted display (HMD), but also the same communication signals can be used to derive coarse body-level gestures and poses of the user, to support real-time avatar control. For fine-grained finger-level gestures, our architecture leverages lightweight sEMG sensors that capture forearm muscle activity. To illustrate the need of both modalities, we present evaluations of both sensing technologies. At the body level (5G), our architecture relies on power-per-beam-pair (PPBP), which can be computed from standard beam management or beam sweeping procedures of the 5G NR standard. PPBP-based sensing achieves 82.2$\pm$5.9% average accuracy when evaluated on users not seen during training. For fine-grained finger-level interactions, we show that surface electromyography (sEMG) carries strong discriminative information achieving consistent promising performance across different movement settings. Thus, combining the two modalities enables multi-scale gesture recognition, at the body level via existing 5G signals and finger level via lightweight sEMG sensors, forming a complete XR framework.

13:00 JSTLLM/生成AIエージェント

タスクガイド付き会話グラフから目標指向の対話ランタイムまで

グラフおよびマルチエージェント オーケストレーション フレームワークは、本番環境の大規模言語モデル (LLM) ワークフローを実用的なものにしますが、ユーザーが相互に依存する複数の目的を維持する場合、それらだけでは会話の連続性を解決できません。この概念システムに関する論文は、他の目標のアクションによって目標が一時停止、再開、修正、無効化される可能性がある、設計空間の非常に複雑な部分に焦点を当てています。目標指向ダイアログ ランタイム (GODR) を導入します。これは、目標、タスク フレーム、ライフサイクル状態、無効化ルール、および再開コントラクトを最上級のランタイム オブジェクトとして扱い、制限された実行をグラフ ランタイム、エージェント、ツール、またはアプリケーション プログラミング インターフェイス (API) に委任するフレームワーク中立の設計パターンです。 GODR は、単純なガイド付きプロセスにおけるワークフロー グラフの代替として提案されていません。これは、エージェント ID、チャット履歴、または実行グラフの位置だけでは客観的な継続性を確実に回復できない、複雑でマルチドメインの中断可能な会話を対象としています。この論文では、問題を形式化し、ランタイム オブジェクトとアーキテクチャの選択基準を提案し、測定されたパフォーマンスの主張ではなく将来の経験的検証の課題として評価を組み立てています。

原文 (English)

From Task-Guided Conversational Graphs to Goal-Oriented Dialogue Runtimes

Graph and multi-agent orchestration frameworks make production large language model (LLM) workflows practical, but they do not by themselves solve conversational continuity when users maintain several interdependent objectives. This conceptual systems paper focuses on the high-complexity end of that design space, where goals can be suspended, resumed, revised, and invalidated by actions in other goals. We introduce the Goal-Oriented Dialogue Runtime (GODR), a framework-neutral design pattern that treats goals, task frames, lifecycle state, invalidation rules, and resumption contracts as first-class runtime objects while delegating bounded execution to graph runtimes, agents, tools, or application programming interfaces (APIs). GODR is not proposed as a replacement for workflow graphs in simple guided processes; it is intended for complex, multi-domain, interruptible conversations where objective continuity cannot be recovered reliably from agent identity, chat history, or execution-graph position alone. The paper formalizes the problem, proposes runtime objects and architecture-selection criteria, and frames evaluation as an agenda for future empirical validation rather than as a measured performance claim.

13:00 JST研究/論文

電車の 10 桁: 2 つの固有値問題の AI 支援検証

特に特異な設定や非正規の設定では、正確な数値固有値を証明するのが難しいことがよくあります。この記事では、そのような 2 つの計算における人間と AI のコラボレーションについて報告します。特異な自己共役シュルディンガー演算子の場合、検証されたゼロ カウントとディリクレ - ノイマン括弧法により、完全な負のスペクトルが小数点以下 10 桁まで証明されます。繊細な非正規原子 - 分子ベンチマークの場合、以前に未解決の共鳴ペアが分離され、各メンバーが 10 桁に囲まれます。2 番目の結果は、一方向シューティングの精度を向上させることではなく、射影解のためのグローバル マッチング システムとして問題を再定式化することによって達成されます。無限の尾部は終端射影データの不確実性としてエンコードされ、コンポーネントごとに尾部に堅牢な Krawczyk-Brouwer 包含によって証明書が提供されます。これにより、AI 支援の強みと限界が明らかになり、その中には、明らかに完全な尾部引数が 1 つ含まれていました。不均一なポリディスクに必要なコンポーネントごとのチェックは、AI 支援の数学の厳格なテストです。これらの例は、証明オブジェクトが重要である理由、およびより広範には、AI によってコード、説明、および妥当性のある数値的主張が可能になるため、その影響が非常に不安定であることを示しています。

原文 (English)

Ten Digits on a Train: AI-Assisted Verification of Two Eigenvalue Problems

Accurate numerical eigenvalues are often difficult to certify, especially in singular or non-normal settings. This article reports a human--AI collaboration on two such computations. For a singular self-adjoint Schr\"odinger operator, a verified zero count and Dirichlet--Neumann bracketing certify the complete negative spectrum to ten decimal places. For a delicate non-normal atom--molecule benchmark, a previously unresolved resonance pair is separated, with each member enclosed to ten digits. The second result is achieved not by increasing the precision of one-way shooting, but by reformulating the problem as a global matching system for projective solution lines. The infinite tail is encoded as uncertainty in the terminal projective data, and a componentwise, tail-robust Krawczyk--Brouwer inclusion supplies the certificate. This gives a reusable architecture for analytic boundary-value systems with ill-conditioned propagation and uncertain asymptotic data. The collaboration also exposes the strengths and limits of AI assistance. AI rapidly produced accurate candidates and plausible proof strategies, but several failed, including one apparently complete tail argument that omitted the componentwise check required by a nonuniform polydisc. Validated computation is a stringent test of AI-assisted mathematics: the output is not merely a number, but a number with a proof. These examples show why the proof object matters, and why human mathematical judgment remained decisive. More broadly, as AI makes code, exposition, and plausible numerical claims inexpensive, standards for verification, attribution, peer review, and training must adapt. The implications are unsettling; the opportunity is extraordinary.

13:00 JST画像/動画生成

From Spatial to Spectral: An Efficient, Frequency-Guided Feature Representation Learner for Small Object Detection

Efficient small object detection is bottlenecked by the inherent feature scarcity of tiny targets, which is further aggravated by operation…

13:00 JST研究/論文

Deciphering Fingerprints of 3D Molecular Surfaces for Accurate Epitope Prediction

Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope predic…

13:00 JSTエージェントロボティクス

Decentralized Coordination of Autonomous Traffic Through Advanced Air Mobility Corridors

The use of dedicated corridors for Advanced Air Mobility (AAM) traffic is one of the most commonly proposed pathways to integrating them in…

13:00 JST研究/論文

The Measurable Majority

This paper studies strict majority reasoning in finite electorates using so-called $\textit{social decision frames}$: finite sets of voters…

13:00 JST研究/論文

Are Safety Guarantees in Neural Networks Safe? How to Compute Trustworthy Robustness Certifications

A primary challenge in AI safety is the existence of adversarial examples -- slightly distorted inputs that cause a neural network (NN) to…

13:00 JST研究/論文

MGI: Member vs Generated Inference

As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine…

13:00 JST研究/論文

JupOtter: Cell-Level Bug Detection in Jupyter Notebooks

Jupyter Notebooks are an increasingly popular coding environment used across many domains, especially in Python-based data science and scie…

13:00 JST研究/論文

Promise and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired image translation: a feasibility study

Purpose: To evaluate the feasibility and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired ima…

13:00 JSTLLM/生成AI

One Year Later...The Harms Persist, But So Do We!

General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain in…

13:00 JSTLLM/生成AI画像/動画生成

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their int…

13:00 JST画像/動画生成

E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucin…

13:00 JSTLLM/生成AI画像/動画生成

The Professor: Multi-Teacher Unsupervised Prompt Distillation for Vision-Language Models

Prompt distillation compresses large vision-language models (VLMs) such as CLIP into lightweight student models by matching teacher predict…

13:00 JST研究/論文

ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation

Distilling conditional diffusion models aims to transfer the behavior of a large teacher to a smaller student while preserving alignment ac…

13:00 JST研究/論文

Catastrophic Compositional Generation: Why Vanilla Diffusion Models Fail to Extrapolate

The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to…

13:00 JSTLLM/生成AIエージェント

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model…

13:00 JSTLLM/生成AI

Offline Reinforcement Learning for Warehouse SLAM Throughput Control

We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment.…

13:00 JST研究/論文

Maestro Order: A Model-Agnostic Orchestration Harness

A single forward pass of a capable model is a fast, fluent, and unreliable problem-solver: it is right often enough to be useful and wrong…

13:00 JSTLLM/生成AI

Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization

End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions the…

13:00 JSTLLM/生成AIGPT / ChatGPT

RASC+: Retrieval-Constrained LLM Adjudication for Clinical Value Set Authoring

Clinical value sets define the standardized terminology codes used in quality measurement, phenotyping, cohort construction, and clinical d…

13:00 JST研究/論文

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tig…

13:00 JST研究/論文

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized g…

13:00 JSTLLM/生成AI

Towards Spec Learning: Inference-Time Alignment from Preference Pairs

Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on…

13:00 JST研究/論文

Fast and Slow Variational Continual Learning

Continual learning remains a major challenge for modern deep networks, partly because commonly used optimizers lack inherent mechanisms for…

13:00 JSTLLM/生成AI

Towards Version-aware Operations and Transaction Memories for Multi-layer MeMo

MeMo proposes language models with explicit multi-layer correlation matrix memories (CMMs), where memorization, retrieval, and forgetting a…

13:00 JST研究/論文

Rapid FinFET Modelling Using an Autoencoder

This work presents a machine learning framework that leverages an autoencoder (AE) for the efficient modeling of FinFET. We first calibrate…

13:00 JST研究/論文

RAVEN: A Regime-Aware Variable-context Expert Network for Financial Time Series Forecasting

Financial time series forecasting presents structural challenges absent from standard benchmarks. Log-returns are non-stationary, exhibit e…

13:00 JSTLLM/生成AI

Selective Capability Unlearning in End-to-End Spoken Language Understanding

Modern spoken language understanding (SLU) systems are increasingly deployed in real-world settings, where specific functionalities may nee…

13:00 JST研究/論文

Token Complexity of Certifying Stochastic-Oracle Reliability

Wang~\cite{Wang2026} introduced the Stochastic-Oracle Turing Machine (SOTM) framework and defined token complexity as the minimum expected…

13:00 JST画像/動画生成

End-to-End Radar and Communication Modulation Recognition with Neuromorphic Computing

Although deep learning-based methods can achieve high accuracy in automatic modulation recognition (AMR) tasks, their high computational co…

13:00 JSTビジネス/資金調達研究/論文

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation

As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. Mor…

13:00 JSTLLM/生成AIハードウェア/半導体

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves…

13:00 JSTLLM/生成AI

Blockwise Policy-Drift Gating for On-Policy Distillation

On-policy distillation (OPD) trains a student policy using teacher signals computed on trajectories sampled by the student itself. Recent w…

13:00 JSTロボティクス

DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs

Recent advances in control have enabled bipedal-wheeled robots to traverse slopes and single-step obstacles, yet long staircase traversal r…

13:00 JSTLLM/生成AI

Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems

We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work. Aggregating every poem attri…

13:00 JST画像/動画生成エージェント

Beyond Bayer: Task-Optimal Sensor Co-Design for Robust Autonomous-Driving Segmentation

Robust perception underpins autonomous driving, and most recent progress comes from scaling the model-larger backbones, foundation models,…

13:00 JST研究/論文

The impact of generative artificial intelligence on academic development of Chinese students in humanities and social sciences

Generative artificial intelligence(GenAI) is reshaping learning in higher education, with particularly pronounced implications for the huma…

13:00 JST画像/動画生成

DramaDirector: Geometry-Guided Short Drama Generation

Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that pr…

13:00 JST画像/動画生成研究/論文

A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy

Vision-language models (VLMs) are prone to hallucination, which remains a major barrier to their safe deployment in clinical practice. To d…

13:00 JST研究/論文

DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration

Music source restoration (MSR) requires jointly addressing source unmixing and the inversion of non-linear production effects. Current meth…

13:00 JSTLLM/生成AIエージェント

Metis: Bridging Text and Code Memory for Self-Evolving Agents

Self-evolving agents improve over time by distilling experience from past executions and reusing it in future tasks. Existing systems repre…

13:00 JST研究/論文

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearin…

13:00 JSTLLM/生成AI

A P\={a}ninian Foundation for Indic Language Processing

More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragment…

13:00 JST研究/論文

Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment

On-device fault detection enables real-time diagnostics without cloud dependency, but deploying machine learning models on resource-constra…

13:00 JSTLLM/生成AIエージェント研究/論文

Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy

Large language models are making research production scalable, shifting the bottleneck from producing artifacts to judging claims. We prese…

13:00 JST画像/動画生成

Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring

Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the…

13:00 JST画像/動画生成ロボティクス

Deep Learning Approaches for 3D Medical Scene Completion: From Geometric Modeling to Generative Paradigms

Three-dimensional scene completion has evolved as a major problem in computer vision and robotics, and its applications are diverse, includ…

13:00 JSTLLM/生成AI画像/動画生成

Co-occurring associated retained concepts in Diffusion Unlearning

Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remov…

13:00 JSTLLM/生成AI研究/論文

MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English eviden…

13:00 JST画像/動画生成

Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing method…

13:00 JSTLLM/生成AIエージェント

AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming

Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments.…

13:00 JST画像/動画生成

Social Structure Matters in 3D Human-Human Interaction Generation

Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it…

13:00 JSTLLM/生成AIビジネス/資金調達

SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance c…

13:00 JST研究/論文

Neural Network-Based Parametric Model Reduction for Predicting Turbulent Flow for Different Vehicle Geometries

Numerical simulations in industrial applications often require performing numerous high-precision computations parameterized by specific ex…

13:00 JSTLLM/生成AI

Pigeonholing: Bad prompts hurt models to collapse and make mistakes

While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradatio…

13:00 JSTLLM/生成AI

CALIBER: Calibrating Confidence Before and After Reasoning in Language Models

Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success.…

13:00 JST研究/論文

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from thi…

13:00 JST研究/論文

ZONOS2 Technical Report

We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve up…

13:00 JST研究/論文

What Does ODRL Mean? A Cross-Level Ontological Grounding of Permissions, Prohibitions, and Duties in UFO-L

ODRL policy evaluators produce verdicts, but say nothing about the normative positions a policy brings into existence, the authority struct…

13:00 JST画像/動画生成

Structural Kolmogorov-Arnold Convolutions: Learnable Function on the Values or the Filter Shape as Parameter-Efficient Alternative to Per-Edge Convolutional KANs

Convolutional Kolmogorov--Arnold Networks (KANs) replace the fixed weights of a convolutional kernel with learnable univariate functions. T…

13:00 JSTLLM/生成AIビジネス/資金調達

On the Stability of Prompt Ranking in Large Language Model Evaluation

Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evalua…

13:00 JST画像/動画生成

Female-RHINO: A Real-Time Scanner-Integrated Framework for Automated Quantitative Uterine MRI Analysis and Structured Reporting

Standardized assessment of uterine MRI remains challenging due to anatomical variability, observer dependence, and the lack of workflow-int…

13:00 JSTロボティクス研究/論文

Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders

Electroencephalography (EEG) is the dominant non-invasive modality for brain-computer interfaces (BCIs), yet reliable decoding of motor ima…

13:00 JST研究/論文

Entity Resolution via Batched Oracle Queries

We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We s…

13:00 JSTエージェントClaude

Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories

Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence…

13:00 JST画像/動画生成

Transformation Behavior of Images in Latent Space

Training of neural networks for histopathology classification tasks typically relies on data encoding into latent space, which reduces comp…

13:00 JST画像/動画生成

MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and Flow Matching

Medical point cloud completion is important for anatomical reconstruction and downstream clinical workflows, yet generative modeling in thi…

13:00 JSTロボティクス

NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation

Perceiving physical contact is fundamental to dexterous manipulation. While robots often rely on dedicated hardware tactile sensors, humans…

13:00 JSTLLM/生成AIGPT / ChatGPTGemma

The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs

Commercial large language models bill, scale latency, and budget context per token. Yet tokenizers assign more subword tokens to the same m…

13:00 JSTロボティクスGoogle

G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models

Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrai…

13:00 JSTLLM/生成AI画像/動画生成

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spa…

13:00 JST研究/論文

Adaptive Machine Learning Framework for UAV Trajectory Optimization in O-RAN

The deployment of unmanned aerial vehicles (UAV) as open radio units (O-RUs) in 6G cellular systems presents a promising opportunity to ach…

13:00 JST画像/動画生成

RetiSEM: Generalising Causal Models for Fragmented Biomedical Data

Learning causal models from fragmented biomedical data is challenging because clinical, molecular, and imaging variables are often incomple…

13:00 JSTエージェント

Red-Teaming the Agentic Red-Team

The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability.…

13:00 JSTLLM/生成AIハードウェア/半導体

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU…

13:00 JSTビジネス/資金調達

A Fair Evaluation of Graph Foundation Models for Node Property Prediction

Due to the wide use of graph-structured data in different fields of industry and science, the development of Graph Foundation Models (GFMs)…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTGeminiQwen

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other hi…

13:00 JSTLLM/生成AI

Toward Self-Evolution-Ready Workflow Harnesses: A Reversible Migration Path and Convertibility Taxonomy for Expert LLM Pipelines

While expert-validated "LLM + script" workflows deliver significant value, they remain static: they encode hard-won domain knowledge yet fa…

13:00 JST研究/論文

Infinitesimal Causality

This paper introduces a categorical account of infinitesimal causality in Frobenius Markov categories equipped with tangent-bundle semantic…

13:00 JSTLLM/生成AIエージェントLlama

Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity

Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios…

13:00 JST研究/論文

Visualizing "We the People": Bridging the Perception Gap through Pluralistic Data Storytelling

Traditional visual data storytelling relies on binary graphics that depict two simplified groups in conflict. This can increase political p…

13:00 JSTLLM/生成AI

AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach

The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods…

13:00 JSTLLM/生成AI

FlowPipe: LLM-Enhanced Conditional Generative Flow Networks for Data Preparation Pipeline Construction

Data preparation pipelines improve data quality in machine learning by transforming raw tables into learning-ready data through sequential…

13:00 JSTロボティクス

TACTFUL: Tactile-Driven Exploration For Object Localization and Identification in Confined Environments

Humans effortlessly locate and identify objects by touch alone, even without vision. In contrast, robotic systems rely heavily on vision an…

13:00 JST画像/動画生成

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing eva…

13:00 JSTLLM/生成AI

Task Decomposition for Efficient Annotation

High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is labor…

13:00 JST研究/論文

Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement

Generative models, particularly diffusion and score-based approaches, have recently achieved strong performance in speech enhancement, but…

13:00 JSTLLM/生成AI画像/動画生成エージェント

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing method…

13:00 JST研究/論文

Context-Aware Prediction of Student Quiz Performance with Multimodal Textbook Features

Educational platforms often predict student performance from prior interactions, but the assessment content itself also varies in linguisti…

13:00 JSTエージェント

DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects

Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability. In the subset with suspected genetic etiologies…

13:00 JSTLLM/生成AIエージェント

Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce

Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close…

13:00 JSTLLM/生成AI

Grad Detect: Gradient-Based Hallucination Detection in LLMs

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinat…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Never…

13:00 JST画像/動画生成

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis

Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not…

13:00 JST研究/論文GPT / ChatGPT

Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution

Quantum computers could outperform classical machines on important problems, but only if the errors that pervade quantum hardware can be co…

13:00 JSTLLM/生成AI画像/動画生成

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-…

13:00 JSTビジネス/資金調達

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are able to do with their sys…

13:00 JST画像/動画生成

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods…

13:00 JSTロボティクスビジネス/資金調達

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in…

13:00 JSTLLM/生成AI

Random Rule Forest (RRF): Interpretable and Manageable Ensembles of LLM-Generated Questions for Predicting Success from Unstructured Data

Many high-stakes screening tasks require predicting rare outcomes from unstructured text, where errors are costly and decisions must be aud…

13:00 JST研究/論文

TIP-Search: Time-Predictable Inference Scheduling for Market Prediction under Uncertain Load

Real-time market prediction services need correct predictions before a decision deadline; a correct prediction delivered late is not usable…

13:00 JST研究/論文

From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control

Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated,…

13:00 JST研究/論文

A global log for medical AI

Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure. Medicine's r…

13:00 JSTLLM/生成AILlamaQwen

Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs

Large language models (LLMs) often produce incorrect or outdated content after being employed. Efficient and accurate knowledge updates wit…

13:00 JSTエージェントビジネス/資金調達

Evolving Programmatic Skill Networks

We study continual skill acquisition in open-ended embodied environments where an agent must construct, refine, and reuse an expanding libr…

13:00 JST研究/論文

BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding

Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cro…

13:00 JSTLLM/生成AI

LLM-MINE: Large Language Model based Alzheimer's Disease and Related Dementias Phenotypes Mining from Clinical Notes

Accurate extraction of Alzheimer's Disease and Related Dementias (ADRD) phenotypes from electronic health records (EHR) is critical for ear…

13:00 JST研究/論文

Grounded Chess Reasoning in Language Models via Master Distillation

Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel.…

13:00 JSTLLM/生成AIエージェント

Subjective-Graph LLM Agents for Simulating Uncertainty in Classroom Social Perception

Social actors do not observe a common social world: each individual forms judgments from a partial and potentially distorted view of the su…

13:00 JST研究/論文

Riemann-Bench: A Benchmark for Moonshot Mathematics

Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficien…

13:00 JST研究/論文

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models ,…

13:00 JSTエージェント研究/論文

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accura…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPTNVIDIA

2.5-D Decomposition for LLM-Based Spatial Construction

Autonomous systems that build structures from natural-language instructions need reliable spatial reasoning, yet large language models (LLM…

13:00 JST研究/論文

生成計画モデルの効率的なテスト時間推論

生成モデルは AI 計画の強力なパラダイムとして登場しましたが、そのパフォーマンスは依然としてトレーニング データの分布によって制限されています。 1 つのアプローチは、テスト時の計算をスケーリングすることで、推論中に生成されるソリューションを改善することです。より効率的な代替方法は、推論プロセス自体を最適化することです。この論文では、古典的なオープンクローズド リスト (OCL) 検索の修正バージョンがまさにそのような効率的な推論手順を提供することを示します。私たちのアルゴリズムは、中間状態から高速ロールアウトを実行する生成モデルと、候補推論パス間で優先順位を付けるヒューリスティック モデルという 2 つの学習されたコンポーネントを相乗させます。主な貢献には、新しい探索制御メカニズムと、OCL フレームワーク内での学習済みモデルの統合が含まれます。複数の組み合わせ計画ドメインにわたって、私たちのアプローチは、計算効率とソリューションの品質において、ニューロシンボリック検索ベースラインと古典的ソルバーの両方を上回っています。

原文 (English)

Efficient Test-time Inference for Generative Planning Models with OCL Search

Generative models have emerged as a powerful paradigm for AI planning, yet their performance remains constrained by the training data distribution. One approach is to improve generated solutions during inference by scaling test-time compute. A more efficient alternative is to optimize the inference process itself. In this paper, we show that a modified version of a classical Open-Closed List (OCL) search provides just such an efficient inference procedure. Our algorithm synergizes two learned components: a generative model that performs fast rollouts from intermediate states and a heuristic model that prioritizes among candidate reasoning paths. Key contributions include novel exploration control mechanisms and integration of learned models within the OCL framework. Across multiple combinatorial planning domains, our approach outperforms both neurosymbolic search baselines and classical solvers in computational efficiency and solution quality.

13:00 JSTエージェント

TouchThinker: 大規模なデータとアクションを意識した表現を使用して、触覚的常識推論をオープンワールドに拡張する

接触は、肉体を持ったエージェントが物理世界を理解するための重要なモダリティです。最近の研究では、触覚常識推論のための言語システムに触覚信号が組み込まれていますが、そのようなシステムを現実的なオープンワールド設定に拡張することは、2 つの重要なボトルネックのため依然として困難です。(1) 現在の触覚推論データセットは形式と規模が制限されたままであり、触覚観察から物理的常識への推論に対する監視が不十分であり、伝達可能な触覚常識の学習を妨げています。 (2) 触覚信号は本質的に冗長でアクション固有ですが、既存の方法ではこれらの特性が見落とされることが多く、その結果、意味表現力が限られた非効率な表現が生じます。これらの制限に対処するために、私たちは、データと表現の両方の観点から触覚の常識的推論をオープンワールドに拡張する触覚言語フレームワークである TouchThinker を提案します。まず、\textbf{415} オブジェクト、\textbf{8} シナリオ、\textbf{7} センサー タイプをカバーする百万規模のマルチソース触覚推論データセットである TouchThinker-1M を構築し、オープンワールドの一般化のための強固なデータ基盤を提供します。さらに、より現実的で多様なタスクを備えたオープンワールドのベンチマークである TouchThinker-Bench を紹介します。次に、触覚表現の効率を向上させ、効率的な推論を可能にするアクション認識モデリングメカニズムを提案します。実験結果は、TouchThinker が複数のデータセットにわたって最先端のモデルに対して競争力のあるパフォーマンスを達成することを示しています。私たちのコードとデータセットは、https://github.com/lvkailin0118/TouchThinker で利用できるようになります。

原文 (English)

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) Tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering \textbf{415} objects, \textbf{8} scenarios, and \textbf{7} sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.

13:00 JSTLLM/生成AIエージェント研究/論文

EComAgentBench: 分散された隠れたインテントを使用した長期タスクに関するショッピング エージェントのベンチマーク

LLM ベースのショッピング エージェントが本番環境に入るにつれて、既存のベンチマークは、買い物客の要件がどのように届くか、つまりクエリで暗黙的に指定されるか、プロファイルに記録されるか、適切な質問がされた場合にのみ明らかにされるかを把握できません。事前に完全な意図を明らかにし、最終的な選択のみを評価するベンチマークでは、この長期にわたる課題を提起することも、エージェントがどの要件を逃したかを説明することもできません。このギャップに対処するために、実際の Amazon 製品とレビューに基づいた 662 のタスクのベンチマークである EComAgentBench を導入します。各タスクは、これらの要件を、目に見えるクエリ、ツールゲートのプロファイル、およびスクリプト化された説明に分散させます。エージェントは、隠れた意図を明らかにし、候補者を属性と照合して証拠を確認し、100 回のツール呼び出し以内に単一の製品にコミットする必要があります。さらに、入力され、ソースタグが付けられたルーブリックにより、すべてのタスクが評価され、各失敗の原因が要件とそのソースに帰されます。構築は自動化されていますが、信頼性が高く、テキストが生成され、すべてのサンプルが検証される前に、すべての回答がコード内で修正されます。 7 つのモデルを評価したところ、最も強力なモデルでも全体の精度は 57.1% にとどまっており、ルーブリックの満足度は、目に見えるソースから隠れたソースへと低下することが明らかになりました。全体として、私たちは EComAgentBench が、ショッピング エージェントを単一クエリ検索から長期にわたる信頼できる支援へと移行させるための再現可能な基盤として機能すると考えています。

原文 (English)

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose full intent upfront and grade only the final choice can neither pose this long-horizon challenge nor explain which requirement an agent missed. To address this gap, we introduce EComAgentBench, a benchmark of 662 tasks grounded in real Amazon products and reviews. Each task scatters these requirements across a visible query, a tool-gated profile, and scripted clarification; an agent must uncover hidden intent, verify candidates against attributes and review evidence, and commit to a single product within 100 tool calls. Moreover, typed, source-tagged rubrics grade every task, attributing each failure to a requirement and its source. Construction is automated yet reliable, with every answer fixed in code before any text is generated and every sample validated. Our evaluation of seven models reveals that even the strongest attains only 57.1% overall accuracy, and rubric satisfaction degrades from visible to hidden sources. Overall, we believe EComAgentBench will serve as a reproducible foundation for moving shopping agents from single-query search toward dependable assistance over long horizons.

13:00 JSTLLM/生成AI研究/論文

BIM-Edit: IFC ベースのビルディング インフォメーション モデリングのための大規模言語モデルのベンチマーク

大規模言語モデル (LLM) は、テキストの指示から設計アーティファクトを生成するために、コンピュータ支援設計 (CAD) にますます適用されています。エンジニアリングの実践では、これには新しいジオメトリを作成するだけではなく、モデルが既存のシーンを理解し、正しく編集し、セマンティクスと関係を保持する必要もあります。ただし、多くの CAD ベンチマークは、既存のモデルを編集するのではなく、新しいモデルを作成することに重点を置き、主に幾何学的正確さを評価します。 Industry Foundation Classes (IFC) 形式で表される Building Information Model (BIM) の自然言語編集に関する LLM を評価するためのベンチマークである BIM-Edit を紹介します。 BIM は、建築モデルがジオメトリをセマンティックおよびリレーショナル構造とともにエンコードするため、困難なテストベッドを提供します。 BIM-Edit には、11 の現実的な建築モデルと 36 の合成シーンにわたる 324 の編集タスクが含まれています。タスクは 3 つの命令カテゴリ (直接、空間、トポロジカル) を使用して表現され、明示的な編集とシーンに基づいた編集の両方をカバーします。私たちは、幾何学的精度、意味論的妥当性、トポロジー的一貫性という 3 つの次元に沿って出力を評価します。評価された LLM 全体で、最もパフォーマンスの高いモデルは、3 つの指標全体で 49.5% の平均スコアしか達成できず、タスクの 3.4% を超える問題を完全に解決するモデルはありません。これらの結果は、現在の LLM 機能と構造化エンジニアリング設計ワークフローの要件との間に大きなギャップがあることを示しています。

原文 (English)

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engineering practice, this requires more than creating new geometry, models must also understand existing scenes, edit them correctly, and preserve semantics and relations. However, many CAD benchmarks focus on creating new models rather than editing existing ones, and mostly evaluate geometric correctness. We introduce BIM-Edit, a benchmark for evaluating LLMs on natural-language editing of Building Information Models (BIM) represented in the Industry Foundation Classes (IFC) format. BIM provides a challenging testbed because building models encode geometry together with semantic and relational structure. BIM-Edit contains 324 editing tasks spanning 11 realistic building models and 36 synthetic scenes. Tasks are expressed using three instruction categories - direct, spatial, and topological - covering both explicit and scene-grounded edits. We evaluate outputs along three dimensions: geometric accuracy, semantic validity, and topological consistency. Across evaluated LLMs, the best-performing model achieves only 49.5% average score across the three metrics, and no model fully solves more than 3.4% of tasks. These results demonstrate a substantial gap between current LLM capabilities and the requirements of structured engineering design workflows.

13:00 JST研究/論文Grok

Repeated Shared Access Enables Grokking, but Edit Propagation Depends on an Addressable Memory

We study factual edit propagation in a controlled synthetic knowledge-graph QA setting using a 2x2 grid that crosses loop recurrence with s…

13:00 JSTLLM/生成AI

When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models

Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two out…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文AnthropicClaudeOpenAIGPT / ChatGPTGoogleGeminiAlibabaQwen

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…

13:00 JSTLLM/生成AI研究/論文

HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs

Logical reasoning is essential for reliable AI, yet existing benchmarks are largely first-order-logic-centric, focusing on object-level ded…

13:00 JST研究/論文

Invariant Graph Representations for Continuous-Time Dynamic Graphs Under Distribution Shifts

Continuous-Time Dynamic Graphs (CTDGs) enable fine-grained modeling of evolving relational systems. However, most existing CTDG representat…

13:00 JSTエージェント

When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments

Can AI Agents simulate real-world trading environments to investigate the impact of external factors on stock trading activities (e.g., mac…

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the develo…

13:00 JST研究/論文

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute. To reduce costs…

13:00 JST研究/論文

Impatient Bandits: Optimizing for the Long-Term Without Delay

Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions

Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as si…

13:00 JSTLLM/生成AI

Societal Alignment Frameworks Can Improve LLM Alignment

Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared value…

13:00 JSTロボティクス

Reward-Centered ReST-MCTS: A Robust Decision-Making Framework for Robotic Manipulation in High Uncertainty Environments

Monte Carlo tree search is attractive for robotic manipulation because it can improve action selection through simulation without requiring…

13:00 JSTLLM/生成AI

Ensemble Learning for Large Language Models in Text and Code Generation: A Survey

Generative Pretrained Transformers (GPTs) are foundational Large Language Models (LLMs) for text generation. However, individual LLMs often…

13:00 JSTエージェント

Multimedia and Visual Analytics in the Agentic Era

Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have…

13:00 JSTロボティクス

MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems

Robots need task planning methods to achieve goals that require more than one action. Recently, large pretrained models have demonstrated i…

13:00 JST研究/論文

Minimisation of Quasar-Convex Functions Using Random Zeroth-Order Oracles

This paper explores the performance of a random Gaussian smoothing zeroth-order (ZO) scheme for minimising quasar-convex (QC) and strongly…

13:00 JST画像/動画生成

SEAL: Searching Expandable Architectures for Incremental Learning

Incremental learning is a machine learning paradigm where a model learns from a sequential stream of tasks. This setting poses a key challe…

13:00 JST研究/論文

Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings

We propose a novel benchmarking methodology for graph neural networks (GNNs) based on the graph alignment problem, a combinatorial optimiza…

13:00 JST画像/動画生成

Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering

Photorealistic volumetric rendering of CT scans greatly benefits clinical workflows, yet neural approaches such as Neural Radiance Fields (…

13:00 JSTLLM/生成AI

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training

Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation. However, expos…

13:00 JST研究/論文

FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation

Industrial signal analysis is hindered by severe data heterogeneity, which we characterize as the M5 problem. Existing solutions rely on sp…

13:00 JSTLLM/生成AIGemini

Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules

Knowledge graphs (KGs) can be enhanced through rule mining; however, the resulting logical rules are often difficult for humans to interpre…

13:00 JSTLLM/生成AI

FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection

Signature-based Intrusion Detection Systems (IDS) detect malicious activity by matching network or host events against predefined rules. Se…

13:00 JSTLLM/生成AI

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor t…

13:00 JSTLLM/生成AI

VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing

The energy cost of Large Language Model (LLM) inference is rapidly becoming a barrier to sustainable and scalable deployment. Although mode…

13:00 JST画像/動画生成

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Light…

13:00 JSTエージェント

ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms

Progress in computational science depends on complex numerical workflows that must faithfully encode physical laws, yet translating concept…

13:00 JST研究/論文

Computing Evolutionarily Stable Strategies in Imperfect-Information Games

We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…

13:00 JST研究/論文

EMFusion: Uncertainty-Aware Conditional Diffusion Model for Multivariate Narrow-band Exposure Forecasting

The rapid growth in wireless infrastructure has increased the need to accurately estimate and forecast electromagnetic field (EMF) levels t…

13:00 JST研究/論文

Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection

Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle…

13:00 JST研究/論文

Disentangling Aleatoric and Epistemic Uncertainty in Physics-Informed Neural Networks. Application to Insulation Material Degradation Prognostics

Physics-Informed Neural Networks (PINNs) provide a framework for integrating physical laws with data. However, their application to Prognos…

13:00 JSTLLM/生成AIエージェント

The $\mathbf{P}$-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs

Modern AI agents increasingly rely on search infrastructure to execute complex, neuro-symbolic reasoning workflows. These workflows often c…

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達研究/論文

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of aut…

13:00 JSTエージェント

自律型 O-RAN に向けて: リアルタイム ネットワーク制御および管理のためのマルチスケール エージェント AI フレームワーク

オープン無線アクセス ネットワーク (O-RAN) は、分散されたソフトウェア駆動のコンポーネントとオープン インターフェイスを通じて柔軟な 6G ネットワーク アクセスを約束しますが、このプログラマビリティにより運用の複雑さも増大します。複数の制御ループがサービス管理層と RAN インテリジェント コントローラー (RIC) 全体で共存しますが、個別に開発された制御アプリケーションは意図しない方法で相互作用する可能性があります。同時に、生成型人工知能 (AI) の最近の進歩により、孤立した AI モデルから、目標を解釈し、複数のモデルと制御機能を調整し、時間の経過とともに動作を適応させることができるエージェント AI システムへの移行が可能になりました。この記事では、非リアルタイム (Non-RT)、準リアルタイム (Near-RT)、およびリアルタイム (RT) の制御ループにわたる調整された階層として RAN インテリジェンスを組織化する、O-RAN 用のマルチスケール エージェント AI フレームワークを提案します。 (i) 非 RT RIC の大規模言語モデル (LLM) エージェントは、オペレーターの意図をポリシーに変換し、モデルのライフサイクルを管理します。 (ii) Near-RT RIC の Small Language Model (SLM) エージェントは、低遅延の最適化を実行し、既存の制御アプリケーションをアクティブ化、調整、または無効化できます。 (iii) 分散ユニット近くのワイヤレス物理層基盤モデル (WPFM) エージェントは、エア インターフェイスに近い高速推論を提供します。これらのエージェントが標準化された O-RAN インターフェイスとテレメトリを通じてどのように連携するかを説明します。オープンソース モデル、ソフトウェア、データセットに基づいて構築された概念実証の実装を使用して、非定常条件下での堅牢な動作とインテント駆動型のスライス リソース制御という 2 つの代表的なシナリオで提案されたエージェント アプローチを実証します。

原文 (English)

Toward Autonomous O-RAN: A Multi-Scale Agentic AI Framework for Real-Time Network Control and Management

Open Radio Access Networks (O-RAN) promise flexible 6G network access through disaggregated, software-driven components and open interfaces, but this programmability also increases operational complexity. Multiple control loops coexist across the service management layer and RAN Intelligent Controller (RIC), while independently developed control applications can interact in unintended ways. In parallel, recent advances in generative Artificial Intelligence (AI) are enabling a shift from isolated AI models toward agentic AI systems that can interpret goals, coordinate multiple models and control functions, and adapt their behavior over time. This article proposes a multi-scale agentic AI framework for O-RAN that organizes RAN intelligence as a coordinated hierarchy across the Non-Real-Time (Non-RT), Near-Real-Time (Near-RT), and Real-Time (RT) control loops: (i) A Large Language Model (LLM) agent in the Non-RT RIC translates operator intent into policies and governs model lifecycles. (ii) Small Language Model (SLM) agents in the Near-RT RIC execute low-latency optimization and can activate, tune, or disable existing control applications; and (iii) Wireless Physical-layer Foundation Model (WPFM) agents near the distributed unit provide fast inference close to the air interface. We describe how these agents cooperate through standardized O-RAN interfaces and telemetry. Using a proof-of-concept implementation built on open-source models, software, and datasets, we demonstrate the proposed agentic approach in two representative scenarios: robust operation under non-stationary conditions and intent-driven slice resource control.

13:00 JST研究/論文

Event-Grounded Question Answering over Long Audio via Structured Retrieval

Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding. Current large audio-lang…

13:00 JST研究/論文

MyoInteract: A Framework for Fast Prototyping of Biomechanical HCI Tasks using Reinforcement Learning

Reinforcement learning (RL)-based biomechanical simulations have the potential to revolutionise HCI research and interaction design, but cu…

13:00 JST研究/論文

Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators

Neural network accelerators have been widely applied to edge devices for complex tasks like object tracking, image recognition, etc. Previo…

13:00 JST研究/論文

No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason

Generative AI can convert uncertainty into authoritative-seeming verdicts, intensifying the hypersuasive force of automated speech and disp…

13:00 JSTLLM/生成AIビジネス/資金調達

An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production

Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the a…

13:00 JST画像/動画生成ロボティクス

CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance

We introduce CRAFT hand, a tendon-driven anthropomorphic hand with hybrid hard-soft compliance for contact-rich manipulation. The design is…

13:00 JST研究/論文

AI-Driven Predictive Maintenance with Environmental Context Integration for Connected Vehicles: Simulation, Benchmarking, and Field Validation

Predictive maintenance for connected vehicles offers the potential to reduce unexpected breakdowns and improve fleet reliability, but most…

13:00 JST画像/動画生成

HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction

Pathology reports are structured, multi-granular documents encoding diagnostic conclusions, histological grades, and ancillary test results…

13:00 JSTLLM/生成AI

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polish…

13:00 JSTLLM/生成AI

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale qu…

13:00 JST研究/論文

THEIA: Learning Complete Kleene Three-Valued Logic in a Pure-Neural Modular Architecture

We present THEIA, a 2.75M-parameter modular neural architecture that learns the complete Kleene three-valued logic (K3) truth table from ta…

13:00 JST画像/動画生成エージェント

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation

Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While rec…

13:00 JSTLLM/生成AI

Fix Initial Programs and Iteratively Refine Repair Instructions Toward Non-Elimination Multi-Turn Program Correction

Recent work on large language models (LLMs) has emphasized the importance of scaling inference compute. From this perspective, the state-of…

13:00 JSTLLM/生成AI

DynamicPO: Dynamic Preference Optimization for Recommendation

In large language model (LLM)-based recommendation systems, direct preference optimization (DPO) effectively aligns recommendations with us…

13:00 JST研究/論文

Ensemble Distributionally Robust Bayesian Optimisation with Continuous Context

We study Bayesian Optimisation (BO) in settings where the objective function is influenced by uncontrollable environmental contexts governe…

13:00 JST画像/動画生成エージェント

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely h…

13:00 JST研究/論文

変分オートエンコーダにおける定常コラプスに対するシンプレックス証人証明書

私たちは、変分オートエンコーダーにおける正確な定数の崩壊を研究します。つまり、決定論的なエンコーダーの平均が入力から独立するようになります。事前分布は標準のガウス分布のままです。 VAE トレーニングの前に、データの GMM ベースのビューから事後固定教師を選択し、固定潜在のみのシンプレックス監視をエンコーダー平均に添付します。この構築により、2 つのリンクされたオブジェクトが生成されます。 1 つ目は証明書です。目撃者の予測が教師の最良の定数予測子を改善する場合、エンコーダーの平均は入力に依存しない定数になることはできません。 2 つ目は局所的なエスケープ方向です。崩壊した多様体では、教師残差によってアライメント損失に対するサンプル依存の下降方向が与えられます。あらゆるフルサポート教師事後分布の場合、同じジオメトリにより、教師と証人の位置合わせエラーがゼロの閉じた形式の潜在コードも得られます。そのスケーリングされたバージョンは、定数予測子から正確な教師コードまでのマージン エネルギー パスを追跡し、保護された目撃部分空間内の非崩壊を定量化します。 MNIST、CIFAR-10、および CIFAR-100 でメソッドをインスタンス化します。教師なしの PCA-GMM 教師を検索すると、バニラ VAE は CIFAR-10 および CIFAR-100 の 5 つのシードすべてで教師証人証明書に不合格ですが、RST バリアントは 5 つのシードすべてに合格します。 \(\beta_{\mathrm{KL}}\in\{2,4,8\}\) を使用した崩壊ストレス設定では、バニラ VAE はすべてのシードで再び失敗しますが、RST-alpha-prefit は証明書陽性のままです。両方の自然画像データセットのエスケープ トラジェクトリは、低マージン初期化からウィットネス マージンを増加させ、非ゼロの教師誘発勾配ノルムを示します。分析は、エンコーダ平均値の正確な一定の崩壊に限定されます。生成品質、デコーダの使用、およびその他の崩壊モードについては、別個の問題として残ります。

原文 (English)

A Simplex Witness Certificate and Escape Force for Constant Collapse in Variational Autoencoders

We study exact constant collapse in variational autoencoders: the deterministic encoder mean becomes independent of the input. The prior remains the standard Gaussian. Before VAE training, we select a fixed teacher posterior from a GMM-based view of the data and attach a fixed latent-only simplex witness to the encoder mean. This construction yields two linked objects. The first is a certificate: if the witness prediction improves on the best constant predictor of the teacher, the encoder mean cannot be input-independent constant. The second is a local escape direction: on the collapsed manifold, the teacher residual gives a sample-dependent descent direction for the alignment loss. For any full-support teacher posterior, the same geometry also gives a closed-form latent code with zero teacher-witness alignment error. Its scaled versions trace a margin-energy path from the constant predictor to the exact teacher code, which quantifies non-collapse inside the protected witness subspace. We instantiate the method on MNIST, CIFAR-10, and CIFAR-100. With searched unsupervised PCA-GMM teachers, vanilla VAEs fail the teacher-witness certificate in all five seeds on CIFAR-10 and CIFAR-100, while RST variants pass in all five seeds. Under collapse-stress settings with \(\beta_{\mathrm{KL}}\in\{2,4,8\}\), vanilla VAE again fails in all seeds, whereas RST-alpha-prefit remains certificate-positive. Escape trajectories on both natural-image datasets increase the witness margin from a low-margin initialization and exhibit nonzero teacher-induced gradient norms. The analysis is confined to exact constant collapse of the encoder mean; generation quality, decoder use, and other collapse modes remain separate questions.

13:00 JSTLLM/生成AIエージェント

Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment

Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in…

13:00 JST研究/論文

トレーニング可能なメタマテリアル特性としてのセンシングインテリジェンス

生物学的システムでは、感知は脳だけで行われるわけではありません。身体は、外部刺激が神経信号に変換される前に、外部刺激を変形、振動、フィルタリングします。工学的に設計されたシステムでは、この処理負荷は主にエレクトロニクスと計算に課せられますが、機械本体は通常、強度と安定性のみを目的として設計されています。ここでは、訓練可能な身体の特性としての感覚知性を紹介します。メタマテリアルの幾何学形状を最適化して、外部刺激をニューラル ネットワークが解釈しやすい内部信号に再形成できることを示します。この物理的な前処理を手動で設計するのではなく、微分可能なシミュレーションを通じてセンシング損失を身体の設計パラメータに逆伝播させることで、ニューラル ネットワークに自身の身体をセンシング用にトレーニングさせます。数値的および実験的なセンシング シナリオ全体で、最適化された本体によりセンシング精度が最大 5 倍向上し、必要な電子センサーの数がほぼ 1 桁削減されます。

原文 (English)

Sensing Intelligence as a Trainable Metamaterial Property

In biological systems, sensing is not performed by the brain alone: the body deforms, vibrates, and filters external stimuli before they are transduced into neural signals. In engineered systems, this processing burden is placed largely on electronics and computation, while the mechanical body is usually designed only for strength and stability. Here, we present sensing intelligence as a trainable property of the body. We show that the geometry of a metamaterial can be optimized to reshape external stimuli into internal signals that are easier for a neural network to interpret. Rather than hand-designing this physical preprocessing, we let the neural network train its own body for sensing by backpropagating the sensing loss to the body's design parameters through differentiable simulation. Across numerical and experimental sensing scenarios, the optimized body improves sensing accuracy by up to fivefold or reduces the number of required electronic sensors by nearly an order of magnitude.

13:00 JSTLLM/生成AIエージェント

スキルが増えればエージェントは劣る?スキル ライブラリを拡張するときにスキル シャドウイングによりパフォーマンスが低下する

スキル ライブラリを使用すると、LLM エージェントはタスク固有の指示をオンデマンドで読み込むことができるため、専門知識のないユーザーは、どのスキルが存在するか、どのように機能するかを知らなくても、自然言語を通じてドメイン固有のタスクを解決できます。ただし、ライブラリが大きくなるにつれて、パフォーマンスは低下します。役立つスキルの小さなセットから 202 のスキル ライブラリに拡張すると、最大 21\% 低下します。この研究では、このパフォーマンスの低下を、既知の役立つスキルのライブラリをロードするときと完全なライブラリをロードするときとの間の合格率の低下として定式化します。さらに、スキルの呼び出し (軌道中にエージェントがどのスキルを選択するか) を条件付けすることで合格率の低下を 2 つの効果に分解することを提案します。 \emph{スキル シャドウイング} (ライブラリが拡張するにつれてエージェントが間違ったスキルを選択する頻度が高くなります)、および \emph{コンテキスト オーバーヘッド} (選択が正しい場合でも、拡大されたコンテキストによって実行が低下する) です。両方の効果の上限を導き出し、合格率の低下に対する影響の大きさを特徴付けます。効果とその上限についての経験的な推定によると、\emph{スキル シャドウイング} 効果はライブラリのサイズとともに増大し、パフォーマンス低下に大きく寄与するのに対し、\emph{コンテキスト オーバーヘッド} 効果は依然として小さく、ゼロと区別がつかないことがわかります。この観察された非対称性は、スキル ライブラリを拡張する際の主なボトルネックは、拡大されたコンテキストではなく、スキル選択の失敗であることを示しています。

原文 (English)

More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries

Skill libraries allow LLM agents to load task-specific instructions on demand, letting non-expert users solve domain-specific tasks through natural language without knowing which skills exist or how they work. However, performance degrades as libraries grow -- by up to 21\% when scaling from a small set of helpful skills to a 202-skill library. In this work, we formulate this performance degradation as the pass rate drop between loading a library of known-helpful skills and the full library. Moreover, we propose to decompose the pass rate drop by conditioning on the skill(s) invocation -- which skills the agent selects during a trajectory -- into two effects: \emph{skill shadowing}, where the agent selects wrong skills more often as the library expands, and \emph{context overhead}, where the enlarged context degrades execution even when selection is correct. We derive upper bounds on both effects to characterize their magnitudes of impacts to the pass rate drop. Our empirical estimates of the effects and their upper bounds both show that the \emph{skill shadowing} effect grows with library size and significantly contributes to the performance degradation, whereas the \emph{context overhead} effect remains small and indistinguishable from zero. This observed asymmetry establishes that the skill selection failure, not the enlarged context, is the primary bottleneck when expanding the skill libraries.

13:00 JSTLLM/生成AI画像/動画生成エージェント研究/論文

VISTA: Visual Spec-to-Web-App コーディング エージェントのエンドツーエンド ベンチマーク

ここでは、LLM ベースのエージェントのエンドツーエンドの Web アプリ生成機能を評価するためのベンチマークである VISTA (VIsual Spec-To-App Benchmark) を紹介します。アルゴリズム タスクに焦点を当てた以前のコード生成ベンチマークとは異なり、VISTA は現実的な UI 中心の開発をターゲットにしており、エージェントは過少指定された入力から機能的で視覚的に一貫したアプリケーションを生成する必要があります。視覚的/構造的忠実度およびスタック制約という 2 つの軸に沿って変化する 5 つのプロンプト情報条件を定義します。(1) 自由なスタック選択によるテキストのみ、(2) 3 つの指定されたスタック下の参照スクリーンショットを含むテキスト、(3) 自由なスタック選択による参照スクリーンショットを含むテキスト、(4) 単一の指定されたスタック下のスクリーンショットおよびプルーニングされた Figma 構造を含むテキスト、(5) 自由なスタック選択によるスクリーンショットおよびプルーニングされた Figma 構造を含むテキスト。堅牢な評価を可能にするために、ベンチマークの各ページにはインタラクティブな UI コンポーネントと約 3 つのビジュアル アンカー ポイントで手動で注釈が付けられ、オープンエンド コード生成設定における Playwright などのスクリプト ベースのテスト ツールのよく知られた制限に対処します。評価では、DOM に基づいた参照マッチング、動作固有のブラウザ テスト、および CLIP ベースの視覚的類似性を組み合わせて、構造の整合性、動作の完全性、および全体的な視覚的な忠実度を共同で測定します。 VISTA を使用して、2 つのモデル ファミリと 2 つのハーネスから描画された 4 つのエージェント システムを評価しました。その結果、入力条件とエージェントの両方で視覚的な忠実性と機能の正確さが部分的に切り離されており、エージェントの編集スタイルは大きく変化しますが、タスクの品質とはほぼ直交していることがわかりました。 VISTA は、エージェントベースのソフトウェア エンジニアリング研究を推進するための厳密で再現可能な基盤を確立します。

原文 (English)

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike prior code generation benchmarks that focus on algorithmic tasks, VISTA targets realistic UI-centric development, where agents must produce functional, visually coherent applications from underspecified inputs. We define five prompt-information conditions that vary along two axes, visual/structural fidelity and stack constraint: (1) text only with free stack choice, (2) text with reference screenshots under three specified stacks, (3) text with reference screenshots under free stack choice, (4) text with screenshots and pruned Figma structure under a single specified stack, and (5) text with screenshots and pruned Figma structure under free stack choice. To enable robust evaluation, each page in the benchmark is manually annotated with interactive UI components and around three visual anchor points, addressing the well-known limitations of script-based testing tools such as Playwright in open-ended code generation settings. Evaluation combines DOM-grounded reference matching, behavior-specific browser tests, and CLIP-based visual similarity, jointly measuring structural alignment, behavioral completeness, and overall visual fidelity. We use VISTA to assess four agent systems drawn from two model families and two harnesses, finding that visual fidelity and functional correctness are partially decoupled across both input conditions and agents, and that agent editing style varies sharply but is largely orthogonal to task quality. VISTA establishes a rigorous and reproducible foundation for advancing agent-based software engineering research. Code is available at https://github.com/kaboider/VISTA_Bench.

13:00 JST研究/論文

QSignAI: 科学のための AI と AI のための科学の交差点における量子ランダムネスシード ID 署名

2024~2025年のノーベル賞とチューリング賞は、AIと量子科学を同時に評価した。しかし、これらの流れを一般公開するために導入されたシステムはまだありません。このペーパーでは、リアルタイム イベント参加システムにおける双方向の AI 量子関係を実証する実稼働環境に導入されたプラットフォームである QSignAI について説明します。私たちは 3 つの質問に取り組みます。2 ソース抽出器による量子ランダム性の生成は、許容可能な遅延で AI 駆動のソーシャル プラットフォームに埋め込むことができるか。 AIボットは量子現象を一般の聴衆が知覚的に判読できるようにすることができるか。そして、その組み合わせたシステムは実際に機能するのでしょうか?会話型ボットは、SV1 および DM1 シミュレーターでの独立した単一量子ビットのアダマール測定と 2 量子ビットのベル状態を介したテプリッツの 2 ソース抽出器で構成される量子パイプラインを介して各参加者の最初のメッセージをルーティングし、参加者ごとに固有の量子ランダムネスシード ID 署名を生成します。最初の 2 つの質問は、システム アーキテクチャとライブ イベントからの導入の定性的な証拠を通じて解決されます。 3 番目は実稼働デプロイメントの成功によるものです。現在のデプロイではクラウド量子シミュレーターが使用されています。物理 QPU のランダム性は短期的な拡張です。測定可能なベンチマークは、将来の優先課題として特定されます。

原文 (English)

QSignAI: Quantum-Randomness-Seeded Identity Signatures at the Intersection of AI for Science and Science for AI

The 2024-2025 Nobel and Turing awards recognised AI and quantum science simultaneously. Yet no deployed system has brought these streams together for the public. This paper presents QSignAI, a production-deployed platform demonstrating a bidirectional AI-quantum relationship in a real-time event participation system. We address three questions: can quantum-randomness generation via a two-source extractor be embedded in an AI-driven social platform with acceptable latency; can an AI bot make quantum phenomena perceptually legible to general audiences; and does the combined system work in practice? A conversational bot routes each participant's first message through a quantum pipeline comprising a Toeplitz two-source extractor over independent single-qubit Hadamard measurements on SV1 and DM1 simulators, plus a 2-qubit Bell state, producing a unique quantum-randomness-seeded identity signature per participant. The first two questions are answered through system architecture and qualitative deployment evidence from live events; the third through successful production deployment. The current deployment uses cloud quantum simulators; physical QPU randomness is the near-term extension. Measurable benchmarks are identified as priority future work.

13:00 JST画像/動画生成ロボティクスNVIDIA

Cosmos 3: Omnimodal World Models for Physical AI

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and actio…

13:00 JSTLLM/生成AI

ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information

Asynchronous reinforcement learning can improve language-model post-training throughput by decoupling response generation from policy optim…

13:00 JSTLLM/生成AIエージェント

A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs

Multi-Document Summarization (MDS) plays a critical role in distilling essential information from collections of textual data. Existing app…

13:00 JST画像/動画生成

Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing

Purpose: Digitization projects in the humanities produce vast, heterogeneous archives of historical documents, making manual sorting imprac…

13:00 JST研究/論文

チームティーチングトークの AI 主導分析: 経験、コホート、学習デザインにわたる音響パターン

教室のコホートが拡大するにつれて、複数の教師の専門知識と教育的観点を統合するためにチームティーチングがますます使用されています。しかし、チームティーチングが実際にどのように展開されるか、特に経験レベル、生徒集団、学習課題設計による教師の貢献の違いについての経験的理解は限られています。チームティーチングに関するこれまでの研究は、主に遡及的な自己報告や小規模な観察に依存しており、チームティーチングが実行されるミクロレベルのプロセスについての洞察は限られていました。教師の話は、これらのプロセスに関する拡張可能なレンズを提供します。個人の教育現場での研究では、音声の音響的特徴(声質、イントネーション、音量など)が生徒の学習を形作る可能性があることが示されていますが、チーム教育の現場での証拠は依然として不足しています。さらに、手動による観察や文字起こしによるこのような特徴の把握は、複数の教師が長時間のセッションや空間的場所にわたって話すチームティーチングの教室では特に困難であり、自動化なしでは拡張性が制限されます。この論文は、空間教育理論とチームティーチング研究に基づいて、チームティーチング環境における教室での会話を分析するための AI ベースの音声処理アプローチを紹介します。私たちは、12 人の教師が参加した学部および大学院での 36 の記録されたセッションを分析しました。空間教育行動がコード化され、音響特徴が抽出されて、教師の経験、生徒コホート、学習タスク設計全体の変動が調べられました。その結果、特にラウドネスダイナミクスにおける体系的な違いが明らかになりました。経験豊富な教師、学部のクラス、および共同学習タスクでは、より大きなラウドネス変動が見られ、重要な情報を前面に出し、教室での対話と参加をサポートするために、より頻繁に音量を調節していることが示唆されました。

原文 (English)

AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design

As classroom cohorts expand, team teaching is increasingly used to integrate the expertise and pedagogical perspectives of multiple teachers. Yet, there is limited empirical understanding of how team teaching unfolds in practice, particularly regarding differences in teachers' contributions across experience levels, student cohorts, and learning task design. Prior research on team teaching has largely relied on retrospective self-reports or small-scale observations, offering limited insight into the micro-level processes through which team teaching is enacted. Teacher talk offers a scalable lens on these processes. While research in individual teaching contexts shows that acoustic features of speech (e.g., voice quality, intonation, and loudness) can shape student learning, evidence from team-teaching settings remains scarce. Moreover, capturing such features through manual observation or transcription is especially challenging in team-teaching classrooms, where multiple teachers speak across extended sessions and spatial locations, limiting scalability without automation. Grounded in spatial pedagogy theory and team-teaching research, this paper presents an AI-based speech processing approach to analyse classroom talk in team-teaching settings. We analysed 36 recorded undergraduate and postgraduate sessions involving 12 teachers. Spatial pedagogy behaviours were coded and acoustic features extracted to examine variation across teachers' experience, student cohorts, and the learning task design. The results reveal systematic differences, most notably in loudness dynamics: high-experience teachers, undergraduate classes and collaborative learning tasks exhibited greater loudness variation, suggesting more frequent modulation of volume to foreground key information and support classroom interaction and engagement.

13:00 JST研究/論文

FedSteer: Taming Extreme Gradient Staleness in Federated Learning with Corrective Projections and Caching

Federated learning (FL) is often subject to aggregation variance if clients do not consistently participate in training rounds. While reusi…

13:00 JST画像/動画生成ビジネス/資金調達

Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata

AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends local acceptance testing and ongoing drif…

13:00 JSTエージェントAnthropicOpenAIGoogle

AgentRivet: an automated system for producing Rivet routines from journal publications

Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measuremen…

13:00 JST研究/論文

Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling

Pairwise comparison is the gold standard for subjective ranking tasks; however, exhaustive annotation requires a massive number of human co…

13:00 JSTLLM/生成AIエージェント

Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but…

13:00 JSTLLM/生成AI

Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts

Small-scale Large Language Models (LLMs) natively default to literal semantic interpretations, making few-shot irony detection a persistent…

13:00 JST研究/論文

量子シネマ: 生成世界モデルを介した量子コンピューティング ハードウェアのインタラクティブな映画的探索

量子コンピューティングは科学と産業全体に革新的な進歩を約束しますが、これらの計算を可能にする物理的ハードウェアは依然として一般の人々には見えません。量子プロセッサは絶対零度に近い温度で密閉された希釈冷蔵庫内で動作するため、直接観察することは不可能です。量子コンピューティングの増大する社会的影響とそれを視覚化する一般の人々の能力との間のこの「想像力のギャップ」は、量子リテラシーと労働力の育成にとって大きな障壁となっています。私たちは、オープンソースのブラウザベースのインタラクティブ アプリケーションである Quantum Cinema を紹介します。これは、生成世界モデルを使用して、目に見えない量子ハードウェアを探索可能な映画のような体験に変換することで、このギャップを埋めます。 Quantum Cinema は、ノーベル賞を受賞した量子もつれの基礎科学から、3 つの主要な量子コンピューティング アーキテクチャ (トラップ イオン、中性原子、超伝導システム) への精選されたビデオ紹介を経て、目に見えない量子現象を観察可能にする没入型の 3 次元生成世界、そして最後に実際の量子デバイスの仕様に基づいたインタラクティブなレーダー チャートの比較まで、4 幕の物語を通してユーザーをガイドします。すべての 3 次元環境は、WorldLabs の生成ワールド モデル プラットフォームを使用して生成され、アマゾン ウェブ サービス (AWS) Braket 量子ハードウェアから厳選されたメトリクスに科学的に基づいています。 Quantum Cinema には、インストール、特殊なハードウェア、量子コンピューティングの知識は必要ありません。これは、プラットフォームの複製または拡張を求める学者や開発者と、さまざまな聴衆に量子ハードウェアを説明するための直感的なツールを求める教育者、研究者、科学コミュニケーターという 2 つの異なるコミュニティにサービスを提供するように設計されています。このペーパーでは、システム アーキテクチャ、生成ワールド モデル パイプライン、両方のコミュニティの使用例、および将来の作業の方向性について説明します。

原文 (English)

Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models

Quantum computing promises transformative advances across science and industry, yet the physical hardware that enables these computations remains invisible to the public: quantum processors operate inside sealed dilution refrigerators at temperatures near absolute zero, making direct observation impossible. This "imagination gap" between quantum computing's growing societal impact and the public's ability to visualize it represents a significant barrier to quantum literacy and workforce development. We present Quantum Cinema, an open-source, browser-based interactive application that closes this gap by transforming invisible quantum hardware into explorable, cinematic experiences using generative world models. Quantum Cinema guides users through a four-act narrative -- from the foundational Nobel Prize-winning science of quantum entanglement, through curated video introductions to three major quantum computing architectures (trapped-ion, neutral-atom, and superconducting systems), into immersive three-dimensional generative worlds that make invisible quantum phenomena observable, and finally to interactive radar-chart comparisons grounded in real quantum device specifications. All three-dimensional environments are generated using WorldLabs' generative world model platform and are scientifically grounded in curated metrics from Amazon Web Services (AWS) Braket quantum hardware. Quantum Cinema requires no installation, no specialized hardware, and no quantum computing background. It is designed to serve two distinct communities: scholars and developers seeking to replicate or extend the platform, and educators, researchers, and science communicators seeking an intuitive tool for explaining quantum hardware to diverse audiences. This paper describes the system architecture, the generative world model pipeline, use cases for both communities, and directions for future work.

13:00 JSTLLM/生成AI研究/論文

LLM ベースの A/B テストの統計的基礎: 人間の因果推論のための代理フレームワーク

組織や研究者は、実験をより迅速かつ低コストで行うことを期待して、A/B テストに人間の参加者の代わりに大規模言語モデル (LLM) を使用することへの関心が高まっています。私たちは、LLM の結果に基づいて推定された治療効果が、対象となるヒト集団に対して測定されたであろう効果をいつ回復するかを研究します。 LLM と人間の結果の間の分布が同等であれば、標準推定量は有効になりますが、非現実的です。したがって、私たちはサロゲートエンドポイント理論を LLM に適応させる統計的フレームワークを開発します。このフレームワークは、LLM のアウトカムをヒトのアウトカムに合わせて調整することで、分布上の同等性よりも劣る代理出産および比較可能性の条件下での平均的な治療効果を特定することを示しています。これらの条件が満たされない場合、目的の効果は部分的にしか特定されず、限られた重複による最悪の場合のバイアスの制限とともに、過去の実験に対する代理を偽装できる診断を提供します。さらに、LLM に固有の確率性によりバイアスと分散の両方が発生しますが、サロゲートとして複数の描画の平均を使用すると、両方が緩和されることを示します。シミュレーションにおける方法と理論、および Upworthy の見出しに関する A/B テストへの応用を説明します。私たちの研究から得られる重要な点は、LLM 結果の代理としての妥当性は過去の治療についてのみ改ざんでき、新しい治療については決して検証できないため、新しい介入には人体実験が依然として不可欠であるということです。設計変数としての LLM の選択、プロンプト、温度の役割と、検証のために人体実験のサイズを設定する方法について説明します。

原文 (English)

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.

13:00 JST研究/論文

KANLib -- A Modular, Extensible and Fast Kolmogorov-Arnold Network Implementation

Kolmogorov-Arnold Networks (KANs) have recently emerged as a promising alternative to traditional multilayer perceptrons by replacing linea…

13:00 JST研究/論文

Essential Subspace Merging for Multi-Task Learning

Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained ch…

13:00 JST研究/論文

データセット、年齢、性別を超えた一般化: リソースの少ない子供の ASR のための微調整戦略の包括的な分析

構音障害のある音声を認識することに関連する課題は、主に、調音精度の低下に起因する顕著な音響変動から生じます。過去の研究では、ハイブリッド DNN/HMM シーケンスの識別トレーニングの使用によって認識が向上することが実証されています。このペーパーでは、さまざまな音響モデルに合わせた音響特徴のさまざまな組み合わせの包括的な調査を示し、それぞれに適した特徴の選択を提供します。ピッチ機能を組み込むことで、特に構音障害のある音声を伴う文章認識タスクの認識パフォーマンスが著しく向上しました。 TORGO データベースの体系的な検査を通じて、構音障害音声を認識するための最先端の因数分解時間遅延ニューラル ネットワーク (F-TDNN) モデルのパフォーマンスを強化できる可能性を実証しました。 F-TDNN モデルを使用して実装された私たちの方法は、以前の研究と比較して、孤立単語認識で 4.65% 相対改善、構音障害音声の文認識で 4.63% 相対改善をもたらしました。この改善により、連続するトレーニング サンプル チャンク間で重複するフレームの数を意図的に選択したことに起因する音声の変動が効果的に補償されます。

原文 (English)

Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.

13:00 JST画像/動画生成ロボティクス

HilDA: 自己監視型 LiDAR の事前トレーニングを促進するための拡散を使用した階層的蒸留

カメラから LiDAR への知識の蒸留に Vision Foundation Models (VFM) を活用することは、現実世界の自動運転 (AD) の膨大な幾何学的および運動学的多様性を表現するために必要な注釈付きデータの不足に対する有望な解決策を提供します。ただし、現在のアプローチは通常、VFM をブラックボックス教師として扱い、フレーム単位の特徴の類似性にのみ依存します。その結果、教師のレイヤーごとの意味構造とグローバル コンテキスト、さらには LiDAR シーケンスに固有の豊富な時空間情報が十分に活用されません。私たちは、運転タスクに必要なセマンティックな内容と幾何学的な場所をより適切に捕捉する、LiDAR バックボーン用の自己監視型事前トレーニング フレームワークである HilDA を提案します。 HilDA は、段階的なセマンティクスの調整のための多層蒸留と、シーンレベルのセマンティクスのためのグローバル コンテキストの蒸留を含む階層的蒸留を、時空間的一貫性を促進する時間占有拡散目標と組み合わせます。 HilDA で事前トレーニングされたモデルは、クロスモーダル蒸留ベンチマークで最先端の結果を達成し、3D オブジェクト検出、シーン フロー、セマンティック占有予測に関して事前の蒸留アプローチでトレーニングされたモデルよりも優れたパフォーマンスを発揮します。コードは https://maxiuw.github.io/hilda で入手できます。

原文 (English)

HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training

Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD). However, current approaches typically treat VFMs as black-box teachers, relying exclusively on frame-wise feature similarity. Consequently, they do not fully exploit the teacher's layer-wise semantic structure and global context, as well as the rich spatiotemporal information inherent in LiDAR sequences. We propose HilDA, a self-supervised pretraining framework for LiDAR backbones that better captures the semantic what and geometric where needed for driving tasks. HilDA combines hierarchical distillation comprising multi-layer distillation for progressive semantic alignment and global context distillation for scene-level semantics, with a temporal occupancy diffusion objective promoting spatiotemporal consistency. Models pre-trained with HilDA achieve state-of-the-art results on cross-modal distillation benchmarks and outperform models trained via prior distillation approaches on 3D object detection, scene flow, and semantic occupancy prediction. Code available at: https://maxiuw.github.io/hilda.

13:00 JST研究/論文

Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling

Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…

13:00 JSTLLM/生成AI

Sexualised synthetic personas encode and amplify gendered power asymmetries through voice

This work examines sexualised AI-generated English-speaking voices offered by a popular commercial platform. New technologies may enable se…

13:00 JST研究/論文LlamaNVIDIA

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study

Mixture-of-Experts (MoE) language models are often described as ideal for resource-constrained inference. Each token activates only a small…

13:00 JSTエージェント

Skills for the future software profession: beyond agentic AI!

As coding agents are rapidly changing software engineering, a natural question is: what are the core skills needed by future software engin…

13:00 JST研究/論文

Alternate loss functions and regression models that achieve robustness to outliers by modulating the learning rate

Most real-world datasets used for training supervised learning models are contaminated with noisy data and outliers leading to large predic…

13:00 JST画像/動画生成

MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning

Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. Howe…

13:00 JST画像/動画生成

Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution

Path-based attribution methods such as Integrated Gradients (IG) are widely adopted for their strong axiomatic properties and effectiveness…

13:00 JST研究/論文

On the Position Bias of On-Policy Distillation

On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision fro…

13:00 JSTLLM/生成AIGPT / ChatGPT

AI Fiction in the Wild

Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, to…

13:00 JST画像/動画生成

Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion es…