Skip to the content.

AIニュース 2026-08-05

自動生成: 2026-08-05 11:54 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Third-party cyber evaluations involving OpenAI modelsOpenAI

    OpenAI explains recent third-party cybersecurity evaluation incidents…

  2. Gemini“ヘビーユーザー”が1年で10倍に 「平均年齢高い、IT本業じゃない」首都高が実践したAI活用推進ノウハウITmedia AI+

    月100回以上Geminiを使う首都高速道路の従業員が、1年で22人から224人に急増した。転機は対面・終日のワークショップと「誰一人取り…

  3. Is the future of data centers portable? Runware builds a pod to find outTechCrunch AI

    On Tuesday, AI infrastructure company Runware announced the launch of…

  4. 「フルスクラッチ開発」って何?──LLMを“骨格”と“筋肉”に例えて国産モデルの現在地を整理するITmedia AI+

    「LLMはどこまで『国産』であるべきか?」と題した本特集では、日本企業が安全性と利便性を確保したAI活用を実現するにはどのレイヤーを国産に…

  5. AI活用で浮上する「導入済みERP」の課題とは? 中堅企業のERPリプレース、新規導入を上回るITmedia AI+

    ノークリサーチによると、中堅・中小企業のERP市場で、中堅企業だけはリプレースの規模が新規導入を上回る。なぜ中堅企業はERPを入れ替えるの…

  6. M365 Copilotのための「Teams活用術3選」 北大DX業務推進室が解説ITmedia AI+

    北海道大学のDX業務推進室が、米MicrosoftのAIサービス「Microsoft 365 Copilot」を活用しやすくするための「M…

  7. 中小企業で「AI活用を回す」には? JAPAN AIと大塚商会らが考える支援策ITmedia AI+

    生成AIの課題は「導入」から「定着」へと移っている。AIツールが社内に広がらず成果につながらない中堅・中小企業に求められることとは。

トピック別件数

日本語メディア9件

ITmedia AI+ (日本語)

10:24 JSTLLM/生成AI

「フルスクラッチ開発」って何?──LLMを“骨格”と“筋肉”に例えて国産モデルの現在地を整理する

「LLMはどこまで『国産』であるべきか?」と題した本特集では、日本企業が安全性と利便性を確保したAI活用を実現するにはどのレイヤーを国産にすべきなのか、そもそも国産にはどのようなメリットがあるのかを、代表的な国内ベンダーへの取材などを通して考察する。第1回の本稿では、そもそもL…

08:00 JSTLLM/生成AIGemini

Gemini“ヘビーユーザー”が1年で10倍に 「平均年齢高い、IT本業じゃない」首都高が実践したAI活用推進ノウハウ

月100回以上Geminiを使う首都高速道路の従業員が、1年で22人から224人に急増した。転機は対面・終日のワークショップと「誰一人取り残さない」全社推進、そして「Gemini Notebook」の存在だった。

08:00 JSTその他

AI活用で浮上する「導入済みERP」の課題とは? 中堅企業のERPリプレース、新規導入を上回る

ノークリサーチによると、中堅・中小企業のERP市場で、中堅企業だけはリプレースの規模が新規導入を上回る。なぜ中堅企業はERPを入れ替えるのか。導入済みERPの課題とは。

07:55 JSTLLM/生成AIMicrosoftCopilot

M365 Copilotのための「Teams活用術3選」 北大DX業務推進室が解説

北海道大学のDX業務推進室が、米MicrosoftのAIサービス「Microsoft 365 Copilot」を活用しやすくするための「Microsoft Teams」の使い方を紹介している。

07:00 JSTその他

「私がウダウダしゃべるより聞きやすい」 三菱重工、決算説明に“AIジュリア”起用の本音

三菱重工が、決算説明にAIナレーター「ジュリア」を起用した。導入の狙いを問われた西尾浩CFOが、笑顔で語った理由とは。

07:00 JSTその他

「日本人が海外へ行けない」を打ち破る シンガポール発LCCが仕掛ける“逆張り戦略”

1ドル160円超の円安と燃油高により「海外旅行離れ」が進む日本市場。シンガポール航空傘下のLCC「スクート」が羽田・那覇線を相次ぎ新設する“逆張り攻勢”を見せている。なぜ円安下でも拡大を続けるのか。「若者・女性向け」を脱した中小企業などのビジネス需要獲得、最新規格「NDC」を活…

07:00 JSTその他

【8/5まで】AIを活用した開発について大調査【Amazonギフトカードが当たる】

現在、「AIを活用した開発業務」に関する読者調査を実施しています。回答者の中から抽選で3名さまにAmazonギフトカード500円分をプレゼントします。

07:00 JSTLLM/生成AI

中小企業で「AI活用を回す」には? JAPAN AIと大塚商会らが考える支援策

生成AIの課題は「導入」から「定着」へと移っている。AIツールが社内に広がらず成果につながらない中堅・中小企業に求められることとは。

13:00 JSTその他

AI予算の7割を食う「ITインフラ」、中でも予算超過しやすいのは? 医療機関調査

Wasabi Technologiesは、医療機関におけるAI活用に関する調査結果を発表した。AI予算の多くを占めるITインフラでは、どのような課題が生じているのか。調査結果から背景を探る。

海外メディア11件

TechCrunch AI (英語)

06:07 JSTその他

SpaceX has bought $329M worth of Tesla Megapacks so far this year

The purchase illustrates just how interconnected Elon Musk's universe of companies are.

05:05 JSTその他

Open-weight AI models are catching up to the frontier. The safety gap remains.

A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing co…

04:48 JSTLLM/生成AIAnthropic

Anthropic signs $10B deal with AI cloud startup Volta

Anthropic has been on a cloud partnership spree in recent months, and its latest move is reportedly a $10 billion deal with AI cloud startu…

04:34 JSTその他

Meet Wrinkles, an app that uncovers the hidden stories of the places around you

Wrinkles, available on both iOS and Android, essentially acts as an AI-powered audio tour guide that reveals hidden history and local stori…

04:28 JSTハードウェア/半導体NVIDIA

Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress

The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending agains…

00:50 JSTその他

Spotify expands AI remix and covers project with Merlin partnership

Spotify says Merlin, which represents more than 30,000 independent labels and distributors, has joined Universal Music Group in backing its…

00:42 JSTその他

Texas halts new data centers as governor calls for audits

Tech companies and developers have been scouring the U.S. for places to build data centers, and they’ve been drawn to Texas’ loose regulati…

00:20 JSTロボティクス

Elon Musk spends half his time talking robots and AI on Tesla earnings calls

An analysis of the last seven years of Tesla earnings calls shows just how little attention Musk pays to Tesla's car business.

23:03 JSTLLM/生成AIOpenAI

Apple says more ex-employees may have taken confidential data to OpenAI

Apple says its trade secrets investigation into OpenAI has widened. In a new court filing, Apple claims additional former staff may have re…

22:00 JSTその他

Is the future of data centers portable? Runware builds a pod to find out

On Tuesday, AI infrastructure company Runware announced the launch of its own modular data center called Sonic Inference Pod.

21:00 JSTその他

EON wants to move the data superhighway from ocean fiber to space lasers

Endeavor Optical Networks is planning to launch the fastest space laser communications system yet built.

公式ブログ1件

OpenAI (英語)

04:00 JSTLLM/生成AIビジネス/資金調達OpenAI

Third-party cyber evaluations involving OpenAI models

OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evalua…

論文261件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント

Agentic AI における OpenClaw と Ollama: 完全自律型でスケーラブルな AI エージェント システムを目指して

リアクティブなラージ言語モデル (LLM) から永続的でアクション可能なシステムへの急速な移行により、エージェントティック AI のアーキテクチャの理解、特に自律型 AI エージェントの推論、オーケストレーション、および実行レイヤーの分離における重大なギャップが明らかになりました。最近の進歩にもかかわらず、フルスタックのエージェント システムを設計および評価するための統一フレームワークは依然として限られています。このペーパーでは、Agentic AI の包括的な階層化アーキテクチャについて説明し、リアクティブな LLM インターフェイスから、メモリ、計画、継続的実行を備えた永続的で目標駆動型の自律 AI エージェントへの進化の概要を説明します。私たちは OpenClaw と Ollama をフルスタックの Agentic AI システムとして分析します。Ollama は LLM 推論レイヤーとして機能し、OpenClaw はエージェントの実行時オーケストレーションを可能にし、推論、ツールの使用、アクションの実行を統合します。 OpenClaw-Ollama アーキテクチャのプロトタイプの実験検証では、永続メモリ、ツールの利用、適応的意思決定などの機能がスタンドアロン モデルではなくシステム レベルの統合から生まれ、アーキテクチャの複雑さが増すにつれてパフォーマンスが一貫して向上することが実証されました。この研究では、スケーラビリティ、セキュリティ、プライバシー、ガバナンス、およびエージェント システムの評価における課題をさらに調査し、堅牢なベンチマークとシステム レベルの設計の必要性を強調しています。将来の方向性には、スケーラブルなマルチエージェント アーキテクチャ、分散型自律システム、責任ある展開のための人間認識型エージェント AI フレームワークが含まれます。全体として、この作業は、Agentic AI の統一されたアーキテクチャ基盤を確立し、フルスタックの自律 AI エージェントの有効性を検証し、スケーラブルで安全で信頼できるエージェント システムを構築するためのロードマップを提供します。すべてのモデル、コード、データセットは、再現性とベンチマークをサポートするために公開されています。

原文 (English)

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agentic AI, particularly in separating inference, orchestration, and execution layers for autonomous AI agents. Despite recent advances, unified frameworks for designing and evaluating full-stack agentic systems remain limited. This paper presents a comprehensive, layered architecture for Agentic AI, outlining the evolution from reactive LLM interfaces to persistent, goal-driven autonomous AI agents with memory, planning, and continuous execution. We analyze OpenClaw and Ollama as a full-stack Agentic AI system, where Ollama serves as the LLM inference layer and OpenClaw enables agent runtime orchestration, integrating reasoning, tool use, and action execution. A prototype experimental validation of the OpenClaw-Ollama architecture demonstrates that capabilities such as persistent memory, tool utilization, and adaptive decision-making emerge from system-level integration rather than standalone models, with performance improving consistently as architectural complexity increases. The study further examines challenges in scalability, security, privacy, governance, and evaluation of agentic systems, highlighting the need for robust benchmarking and system-level design. Future directions include scalable multi-agent architectures, distributed autonomous systems, and human-aware Agentic AI frameworks for responsible deployment. Overall, this work establishes a unified architectural foundation for Agentic AI, validates the effectiveness of full-stack autonomous AI agents, and provides a roadmap for building scalable, secure, and trustworthy agentic systems. All models, code, and datasets are publicly released to support reproducibility and benchmarking.

13:00 JSTLLM/生成AIエージェント研究/論文ClaudeGPT / ChatGPTGemini

AIはAI科学者を評価できるのか?自動マルチモデルレビューを使用した自律的研究生成システムのベンチマーク調査

自律的な研究が可能な AI Scientist システムは、科学的発見を大幅に加速する可能性を秘めています。ただし、AI によって生成された論文の品質を評価および比較することは、依然として未解決の課題です。私たちは、最先端の大規模言語モデルを利用して科学論文を独創性、科学的厳密性、明確性、重要性の 4 つの主要な側面にわたって評価する自動査読システムを使用した、厳密なベンチマーク プロトコルを提案および実装します。私たちは 4 つの主要な AI Scientist フレームワーク、\textit{Sakana AI (v1 & v2)}、\textit{CycleResearcher}、\textit{Data-to-Paper} を評価します。各フレームワークは、民間の自律型 AI サイエンティスト企業 (FARS) が発行した 15 件の研究提案の一貫したセットに基づいて実行され、15 件の FARS ベンチマーク ペーパーと並行して評価する 60 件の論文が生成されました。 3 人の独立した LLM レビュー担当者 (GPT-5.4、Gemini、および Claude) を使用したところ、FARS ベンチマーク論文は競合するすべてのフレームワークよりも大幅に優れており、他のシステムの平均スコアが 1.00 ~ 1.87 であるのに対し、1 ~ 5 スケールで 2.14 ~ 2.47 を達成していることがわかりました。特に、FARS スコアは、Gemini および Claude の評価で次善のシステムより 2$\times$ 以上高くなっています。 Gemini と Claude の間には強い一致が見られ ($\rho$ = 0.907、$p < 0.001$)、両方とも合成スコア ($\rho$ = 0.961、$p < 0.001$) と非常に強く相関しており、自動評価の信頼性が検証されています。ただし、GPT-5.4 は一致度が低く ($\rho \約 0.32$)、異なる基準を使用して論文を評価していることを示唆しています。これらの結果は、AI Scientist システムの最初の定量的ベンチマークを確立し、マルチモデル LLM 評価が自律的な研究の品質を評価するためのスケーラブルで一貫したフレームワークを提供することを実証しています。

原文 (English)

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous benchmarking protocol using an automated peer-review system that harnesses frontier large language models to assess scientific papers across four core dimensions: originality, scientific rigor, clarity, and significance. We evaluate four leading AI Scientist frameworks: \textit{Sakana AI (v1 & v2)}, \textit{CycleResearcher}, and \textit{Data-to-Paper}. Each framework was run on a consistent set of 15 research proposals published by a commercial autonomous AI scientist company (FARS), generating 60 papers that we evaluate alongside 15 FARS benchmark papers. Using three independent LLM reviewers (GPT-5.4, Gemini, and Claude), we find that FARS benchmark papers significantly outperform all competing frameworks, achieving mean scores of 2.14--2.47 on a 1--5 scale compared to 1.00--1.87 for other systems. Notably, FARS scores are more than 2$\times$ higher than the next-best systems on Gemini and Claude evaluations. We find strong agreement among Gemini and Claude ($\rho$ = 0.907, $p < 0.001$), and both correlate extremely strongly with the synthesis score ($\rho$ = 0.961, $p < 0.001$), validating the reliability of automated evaluation. However, GPT-5.4 exhibits weaker agreement ($\rho \approx 0.32$), suggesting it evaluates papers using different criteria. These results establish the first quantitative benchmark for AI Scientist systems and demonstrate that multi-model LLM evaluation provides a scalable, consistent framework for assessing autonomous research quality.

13:00 JSTLLM/生成AI

主要な数学的予想を発見するための LLM フレームワーク: AI による次のリーマン予想の探索

主要な数学的予想は依然として専門家の直観に大きく依存しているため、実質的な数学的可能性を備えた予想を体系的に生成および検証するための統一された方法は依然として利用できません。明示的なローカル証拠モジュールからの領域検索、基礎性、新規性、潜在的重要性の反射的検証、リーン 4 と Mathlib での形式的検証を備えた、主要な推測の発見のための 3 段階のパイプラインを紹介します。その目的は、高度な問題センスを持つ数学的問題、つまり、その証明が研究領域の言語を再編成し、人間の数学的研究に永続的な助けを提供できる問題を発見することです。 20 の候補についての実験では、自然言語から形式チェックへの確実な移行が示され、20 の候補のうち 20 はリーン解析と型チェックに合格し、20 の候補のうち 20 は正確? に直接吸収されず、20 の候補のうち 20 はイソップによって自動的に除外されず、明示的な重複または重複に近いものはありませんでした。

原文 (English)

LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis

Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation of conjectures with substantial mathematical potential remains unavailable. We present a three stage pipeline for major conjecture discovery, with region search from explicit local evidence modules, reflective validation for foundationality, novelty, and potential significance, and formal validation in Lean 4 and Mathlib. The objective is the discovery of mathematical problems with high problem taste, namely problems whose proofs could reorganize the language of a research area and provide durable help to human mathematical research. Experiments on twenty candidates showstable passage from natural language to formal checks, with twenty out of twenty candidates passing Lean parsing and type checking, twenty out of twenty candidates not directly absorbed by exact?,twenty out of twenty candidates not automatically discharged by aesop, and no explicit duplicates or near duplicates.

13:00 JSTLLM/生成AI

ThinkReset: 境界コンテキストの長期的推論のための学習可能な中間インターフェイス構築

長い思考連鎖による推論は、複雑な問題のパフォーマンスを向上させますが、冗長性の蓄積、コンテキストのオーバーフロー、エラー アンカリングも引き起こします。私たちは、制限されたコンテキストウィンドウの下では、中心的なボトルネックは軌道圧縮やテスト時間の制御ではなく、破棄された履歴を置き換えて継続的な解決をサポートできる再利用可能な中間インターフェイスの欠如であると主張します。さらに、結果報酬主導型の長鎖強化学習の主要な失敗モードを特定します。つまり、ウィンドウがほぼ使い果たされる前にモデルがタスクを解決できなかった場合、最終的な答えの報酬によって、慎重な推論を継続するのではなく、時期尚早の推測が奨励されます。私たちは、このビューのテキスト空間インスタンス化である ThinkReset を提案します。 ThinkReset は、インターフェイスのライトバックとリセットを通じて再利用可能な中間インターフェイスを明示的に構築し、リセット後の継続成功を直接最適化します。この視点により、複数の長期推論ベンチマークにわたって、固定コンテキスト ウィンドウの下での成功率が一貫して向上します。

原文 (English)

ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning

Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not trajectory compression or test-time control, but the absence of a reusable intermediate interface that can replace discarded history and support continued solving. We further identify a key failure mode of outcome-reward-driven long-chain reinforcement learning: when the model has not solved the task before the window is nearly exhausted, the final-answer reward encourages premature guessing rather than continued careful reasoning. We propose ThinkReset, a text-space instantiation of this view. ThinkReset explicitly constructs reusable intermediate interfaces through interface writeback and reset, and directly optimizes post-reset continuation success. Across multiple long-horizon reasoning benchmarks, this perspective consistently improves success rates under fixed context windows.

13:00 JSTLLM/生成AI

TAPR: タスク認識プロンプト リライターによる LLM パフォーマンスの強化

大規模言語モデル (LLM) では、その可能性を最大限に引き出すために慎重に作成されたプロンプトが必要になることが多く、これが専門家以外のユーザーにとって障壁となる場合があります。この取り組みでは、下流の LLM パフォーマンスを向上させるという明確な目標を持って、ユーザー プロンプトをタスクに最適化されたプロンプトに再定式化するモデルである Task-Aware Prompt Rewriter (TAPR) を導入することで、この課題に対処しています。グループ相対ポリシー最適化 (GRPO) による強化学習を使用して TAPR をトレーニングします。報酬は、再定式化されたプロンプトと対応するタスク出力の両方に対する LLM による審査員評価から導出されます。質問応答、要約、算術推論などのさまざまなタスクに関する実験結果は、私たちの方法が基本モデルよりも迅速な書き換え能力において一貫した利益をもたらすことを示しています。 Phi-4-mini-instruct (TAPR の基本モデルとして) を微調整すると、より明確で有益な言語を含むプロンプトが生成され、Natural question や GSM8K などの確立されたベンチマークでの精度が向上します。私たちのコードは、https://github.com/OliverSavolainen/task-specific-prompt-rewriter で入手できます。

原文 (English)

TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter

Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the challenge by introducing a Task-Aware Prompt Rewriter (TAPR), a model that reformulates user prompts into task-optimized prompts with the explicit goal of improving downstream LLM performance. We train TAPR using reinforcement learning with Group Relative Policy Optimization (GRPO), where rewards are derived from LLM-as-judge evaluations of both the reformulated prompt and the corresponding task output. Experimental results on diverse tasks, such as question answering, summarization, and arithmetic reasoning, show that our method yields consistent gains over base models in prompt rewriting ability. Fine-tuning Phi-4-mini-instruct (as the base model for TAPR) produces prompts that contain clearer and more instructive language, leading to higher accuracy on established benchmarks such as Natural Questions and GSM8K. Our code is available at: https://github.com/OliverSavolainen/task-specific-prompt-rewriter

13:00 JST研究/論文

ハイブリッド トークン化と直列並列デコードによるクロスドメインの逐次推奨を強化

クロスドメイン逐次レコメンデーション (CDSR) は、複数のドメインにわたるユーザーの動的な関心の推移と逐次パターンをモデル化することを目的としています。最近、生成的推奨(GR)が登場しました。まずアイテムのセマンティクスからセマンティック識別子 (SID) を学習し、自己回帰生成として推奨を定式化します。しかし、既存の方法は 2 つの重大な問題に直面しています。(1) トークン化中にドメイン間の協調相関を無視する、(2) 生成中にビーム検索などの非効率的なデコード戦略を採用するため、リアルタイムの展開が妨げられます。これらの制限に対処するために、CDSR の効果的かつ効率的な生成フレームワークである GenCDSR を提案します。具体的には、マルチタワー アーキテクチャを備えたクロスドメイン ハイブリッド トークン化メカニズムを設計し、階層的な共有固有のきめ細かいコードブックを通じて、クロスドメインの共通点とドメイン固有の差異を共同で取得します。さらに、階層型 SID 構造を利用して生成を部分的に並列化し、生成の一貫性を維持しながら推論レイテンシを大幅に削減する、クロスドメイン直列並列デコード戦略を開発します。 3 つの公開データセットでの実験では、GenCDSR が最先端のベースラインと比較して、平均 1.5 パーセントの精度向上と、平均推論レイテンシの 85.1 パーセントの削減を達成したことが示されています。実装コードとデータセットはオンラインで入手できます: https://github.com/Applied-Machine-Learning-Lab/RecSys2026_GenCDSR。

原文 (English)

Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding

Cross-domain sequential recommendation (CDSR) aims to model users' dynamic interest transitions and sequential patterns across multiple domains. Recently, generative recommendation (GR) has emerged. It first learns semantic identifiers (SIDs) from item semantics and formulates recommendation as autoregressive generation. However, existing methods face two critical issues: (1) they ignore collaborative correlations across domains during tokenization, and (2) they adopt inefficient decoding strategies, such as beam search, during generation, which hinders real-time deployment. To address these limitations, we propose GenCDSR, an effective and efficient generative framework for CDSR. Specifically, we design a cross-domain hybrid tokenization mechanism with a multi-tower architecture to jointly capture cross-domain commonalities and domain-specific distinctions through hierarchical shared-specific and fine-grained codebooks. Furthermore, we develop a cross-domain serial-parallel decoding strategy that leverages the hierarchical SID structure to partially parallelize generation, significantly reducing inference latency while preserving generation consistency. Experiments on three public datasets show that GenCDSR achieves an average accuracy improvement of 1.5 percent and an average inference latency reduction of 85.1 percent compared with state-of-the-art baselines. The implementation code and datasets are available online: https://github.com/Applied-Machine-Learning-Lab/RecSys2026_GenCDSR.

13:00 JST研究/論文

異種ドキュメントからナレッジ グラフを構築するための、オントロジーに基づいた重複排除を意識した抽出レイヤー

大規模な言語モデルは、非構造化文書からエンティティと関係を流暢に抽出しますが、一貫性はありません。文書間で活字の語彙が分断され、同じ人物が複数の名前のバリエーションで現れ、関係が重複し、名前を共有する別個の個人が沈黙の混同を引き起こす危険があります。このペーパーでは、ライブ ドキュメント ストリームを正式なオントロジーに合わせた検証済みのナレッジ グラフに変換するプロダクション抽出レイヤーの設計、実装、および経験に基づく改良について説明します。このシステムは、Kafka からのドキュメント メタデータを消費し、各形式用に構築されたハンドラーを通じて PDF、スプレッドシート、Office、および画像コンテンツをルーティングし、オントロジーに合わせて調整されたローカルにホストされた Qwen3.5-9B モデルを使用して 2 つのパスでエンティティと関係を抽出します。その際立ったコンポーネントは、オントロジーに基づく抽出です。キュレーションされたオントロジーの関連スライスは、類似性を埋め込むことでグラフ データベースからライブで取得され、抽出プロンプトに挿入され、静的ドメイン スライスと比較してカタログのオーバーヘッドが約 94 パーセント削減されます。抽出された結果は、決定論的クリーニング、チャンク間のマージ、関係性の 2 番目のパス、モデル推論を必要としない 6 つの重複排除アルゴリズム、および類似性スコアがオーバーライドできない競合ガードを備えた埋め込み解決サブシステムの 5 段階の改良パイプラインを通過します。インテリジェンス コーパスの評価により、誤ったマージが発生することなく、検索再現率が約 70 パーセントから 95 パーセントに改善され、ソース テキストを 1 文字切り詰めるバグから、タイトル接頭語を伴うエンティティの体系的な重複に至るまで、7 つのクラスのサイレント品質欠陥が修正されました。

原文 (English)

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation. This paper presents the design, implementation, and empirical refinement of a production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology. The system consumes document metadata from Kafka, routes PDF, spreadsheet, Office, and image content through handlers built for each format, and extracts entities and relationships in two passes using a locally hosted Qwen3.5-9B model tuned on the ontology. Its distinguishing component is ontology-guided extraction: the relevant slice of a curated ontology is retrieved live from a graph database by embedding similarity and injected into the extraction prompt, reducing catalog overhead by about 94 percent relative to static domain slices. Extracted results then pass through a refinement pipeline of five stages: deterministic cleaning, merging across chunks, a second pass for relationships, six deduplication algorithms that require no model inference, and an embedding resolution subsystem whose conflict guard no similarity score can override. Evaluation on intelligence corpora improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect, ranging from a bug that truncated source text by a single character to the systematic duplication of entities that carried title prefixes.

13:00 JSTLLM/生成AI

どれくらい難しいと思いますか? LLM 思考連鎖軌跡におけるステップ認識推論エネルギーの分析

計算量が個々の思考連鎖 (CoT) 推論ステップにどのように割り当てられるかを理解することは、依然として未解決の課題です。既存の解釈可能性手法は、出力レベルの信号に依存するか、処理深度を単一の軌道レベルのスカラーに折りたたんで、ステップごとの量が不透明なままになります。我々は、隣接するトランスフォーマ層にわたるトークン隠れ状態のグラム行列間の中心カーネルアライメント(CKA)を介して個々のCoTステップの粒度で労力を定量化する幾何学的フレームワークであるステップアウェア推論エネルギー(SARE)を提案し、固有ベクトルアライメントやクラスタ対応を必要とせずにトークン間の関係構造をキャプチャします。 SARE はさらに、CoT の軌跡を潜在的な意味論的状態間の遷移としてモデル化することで、推論の意味論的進行内でこのエネルギーを文脈化します。 6 つの推論ベンチマークと 3 つのオープンウェイト LLM にわたって、推論エネルギーがステップ タイプ間で非常に不均一であり、軌道レベルのメトリクスには見えない位相のような遷移が見られることがわかりました。不正確な軌道は、重要な推論接合部で体系的に低いエネルギーを示します。また、SARE ベースの特徴は、ほとんどの設定で出力ベースの信頼ベースラインと一致するか、それを上回っています。これは、内部の幾何学的ダイナミクスが表面レベルの信号を超えた予測情報をエンコードしていることを示しています。

原文 (English)

How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into a single trajectory-level scalar, leaving step-wise effort opaque. We propose Step-Aware Reasoning Energy (SARE), a geometric framework that quantifies effort at the granularity of individual CoT steps via Centered Kernel Alignment (CKA) between Gram matrices of token hidden states across adjacent transformer layers, capturing inter-token relational structure without requiring eigenvector alignment or cluster correspondence. SARE further contextualizes this energy within reasoning's semantic progression by modeling CoT trajectories as transitions among latent semantic states. Across six reasoning benchmarks and three open-weight LLMs, we find that reasoning energy is highly non-uniform across step types, exhibiting phase-like transitions invisible to trajectory-level metrics; incorrect trajectories show systematically lower energy at critical reasoning junctions; and SARE-based features match or outperform output-based confidence baselines in most settings, indicating that internal geometric dynamics encode predictive information beyond surface-level signals.

13:00 JSTLLM/生成AIエージェント

現実世界の臨床ケアにおける推論: 大規模言語モデルが自律的な臨床意思決定支援にとってまだ安全ではない理由

LLM は現在、医師免許試験に合格しており、厳選された症例では診断推論において医師に匹敵することができます。これらの開発により、診断と治療のガイダンス、管理文書、ルールベースの警告強化における症状の評価と臨床意思決定のサポートのための LLM の使用が加速しました。この視点は、これらのアプリケーションの中で最も重要なもの、つまり、臨床医がほとんど、またはまったく関与しない、自己提示の未分化患者の自律的なトリアージに関するものです。このタスクについては、安全性の証拠はまだ存在していません。ギャップは医学知識にあるのではなく、臨床評価の忠実度にあります。最も可能性の高いテキストを続行するように最適化されたモデルは、安全な答えがありそうもない絶対に見逃せない診断である場合、安全に動作するように最適化されていません。安全なトリアージは、最も可能性の高い診断を選択することではありません。それは非対称コストの下での逐次的な決定であり、単一の壊滅的なミスが多くの誤報よりも重要であり、決定的な信号は患者が自発的ではない信号である可能性があり、モデルが探索するように訓練されていない可能性があります。したがって、中核的な欠陥は、不確実性の下での情報収集の1つです。不完全な履歴の下では、LLM システムは安全なトリアージに必要な動作を示せない可能性があります。行方不明の赤旗を求めて。エスカレーションの閾値を下げる。十分な情報が得られるまで判断を延期する。そして、有害性の高い診断が依然として除外されない場合に懸念が増大する。 LLM のこれらの障害モードは、これまでの評価では、完全で綿密に管理された信頼ゲート制御のシミュレーションが使用されることが多いため、検出が困難な場合があります。このような状況下での LLM の適用は、臨床トリアージのロジックに制約されていない場合、助手のような行動や、軽薄さ、同調性、誤った調整などのポジティブなバイアスによって増幅される可能性があります。

原文 (English)

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.

13:00 JST画像/動画生成エージェント

ViSAGE: 長時間ビデオを理解するための自己修正記憶の構築

長期環境で動作するマルチモーダル エージェントは、エンティティ一貫性のある時間的に根拠のある推論をサポートするために、マルチメディア メモリを構築し、継続的に更新する必要があります。ただし、既存のエージェント メモリ アプローチでは、積極的な圧縮とセグメント単位の処理により、きめの細かいアイデンティティ キューが破棄されることがよくあります。また、ベクトル類似性検索にも大きく依存しており、意味的に関連しているにもかかわらず同一性が一致しない証拠が表面化し、エンティティの混乱、エラーの伝播、幻覚的な回答につながる可能性があります。我々は、自己修正型のエンティティ中心の記憶を構築するマルチモーダルなエージェント記憶フレームワークである ViSAGE を提案します。具体的には、ViSAGE は、長い時間範囲にわたるクロスモーダル バインディングを介してエンティティのアイデンティティを固定します。次に、双方向の記憶改良を適用して遅延した身元証拠を伝播し、過去の記録を遡って統合し、将来の推論を改善します。また、身元証拠の整合性を重視した状態で取得した証拠を評価するためのマルチエージェント相互検証を導入し、証拠が欠落している場合に裏付けのない回答ではなく棄権を可能にします。広範な結果は、ViSAGE が最も強力なベースラインを常に上回っており、5.9% 高い精度を達成していることを示しています。

原文 (English)

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to entity confusion, error propagation, and hallucinated answers. We propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories. Specifically, ViSAGE anchors entity identity via cross-modal binding over long temporal ranges. It then applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning. We also introduce multi-agent cross-verification to assess retrieved evidence under an identity-evidence alignment onstraint, enabling abstention instead of unsupported answers when evidence is missing. Extensive results demonstrate that ViSAGE consistently outperforms the strongest baseline, achieving 5.9% higher accuracy.

13:00 JSTエージェントロボティクス

Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO

Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial ins…

13:00 JST研究/論文

LSR-Synth のライブラリ到達可能性: 反記憶化設計がシンボリック検出の測定をどのように変えるか

科学方程式発見のための既存のベンチマークは、主にパブリックドメインで入手可能なよく知られた方程式で構成されているため、モデルがデータから法則を発見しているのか、それともトレーニング コーパスから単に答えを呼び出しているだけなのかを判断することが困難です。 LSR-Synth は、確立された科学メカニズムに新しい合成用語を導入し、新規性、解決可能性、科学的妥当性に関して結果として得られるタスクをフィルタリングすることにより、この問題を軽減します。この論文では、より狭い測定の質問を検討します。これらのタスクは、言語モデルによって提供される科学的事前分布を、タスクのセマンティクスにアクセスしない従来の演算子検索からさらに区別できるでしょうか?私たちは、公的に文書化された来歴を持つ固定語彙を使用してセマンティクスのないベースラインを構築し、セマンティクスのブラインド化、ライブラリの弱体化、一致する演算子ファミリーのノックアウトを通じて候補カバレッジの役割を評価します。現在のタスク スナップショット、検索バジェット、およびスコアリング プロトコルでは、固定語彙はすでにほとんどのタスクをカバーしていますが、言語モデルによって生成された候補が解決可能なインスタンスのセットを拡張することはほとんどありません。語彙の範囲が選択的に破壊された場合にのみ、それらのわずかな貢献が実質的になります。厳密な分布外評価はすべての方法の絶対成功率を低下させますが、この関係は変わりません。これらの発見は、完全な公式の暗記に対する LSR-Synth の制御を無効にするものでも、言語モデルの事前分布が一般に役に立たないことを示唆するものでもありません。むしろ、彼らは、より限定的な結論を支持しています。つまり、現在のタスクのほとんどは、これまでに見たことのない式の適合と再結合を評価するのには依然として適していますが、それ自体では、固定検索空間を超えた事前分布からの寄与を特定するには不十分です。

原文 (English)

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its training corpus. LSR-Synth mitigates this problem by introducing novel synthetic terms into established scientific mechanisms and filtering the resulting tasks for novelty, solvability, and scientific plausibility. This paper examines a narrower measurement question: can these tasks further distinguish scientific priors supplied by language models from conventional operator search that does not access task semantics? We construct a semantics-free baseline using a fixed vocabulary with publicly documented provenance, and assess the role of candidate coverage through semantic blinding, library weakening, and matched operator-family knockouts. Under the current task snapshot, search budget, and scoring protocol, the fixed vocabulary already covers most tasks, while language-model-generated candidates rarely expand the set of solvable instances. Their marginal contribution becomes substantial only when vocabulary coverage is selectively disrupted. Strict out-of-distribution evaluation lowers the absolute success rates of all methods but does not alter this relationship. These findings neither invalidate LSR-Synth's controls against memorization of complete formulas nor imply that language-model priors are generally unhelpful. Rather, they support a more limited conclusion: most current tasks remain suitable for evaluating the fitting and recombination of previously unseen expressions, but are insufficient on their own to identify contributions from priors beyond a fixed search space.

13:00 JSTエージェント研究/論文

安全性、それとも単なる機能?エージェントの安全性ベンチマークの妥当性監査

エージェントの安全性ベンチマークはさまざまな行動を測定し、そのスコアはエージェントの安全性として同じ意味で引用されます。そのうちの 4 つ (R-Judge、Inj​​ecAgent、AgentHarm、AgentDojo) を検証対象の測定値として扱い、それぞれを公式の実装と作成者が提供するスコアラーの下で最大 22 のモデルで実行し、MMLU と GPQA を 1 つのプロトコルの下で機能複合として測定しました。指標が最初の問題です。 $F_1$ でスコア付けされたバイナリ トレース判定ベンチマークでは、「常にポジティブ」ポリシーは $F_1 = 2\pi/(1+\pi)$ を達成します。 R-Judge では $0.690$ で、実際に判別できる 21 モデルのうち 5 モデルを上回っています。次に、3 つの広範なベンチマークは、同じ 18 モデルを異なるランク付けします。その不一致の背後にあるトレードオフは、小さなパネルのアーチファクトです。AgentHarm の安全性に対する R 判定の特異性は、$n{=}7$ では $-0.64$、$n{=}18$ では $+0.02$ と相関し、ランダムなサイズ 7 のサブセットの 4 分の 1 は $|\rho| に達します。 \geq 0.5$ はそのほぼゼロの値です。保留された有効性により、どの結果を選択するかが決まります。能力はタスクの成功を予測します ($\rho{=}{+}0.60$) が、位置ずれの安全性と負の相関があります ($\rho{=}{-}0.44$、$n{=}21$)。ペアの $n{=}20$ パネルでは、対応するコントラストは $\Delta{=}{-}1.00$ (95% CI $[-1.48, -0.49]$、$p<0.001$) であり、1 つの組織を除外した場合や組織をクラスター化したブートストラップ分析にも耐えられます。拡張された 41 モデルのパネルでは、位置ずれの相関は $-0.16$ (95% CI $[-0.54, +0.22]$) に弱まり、ジェイルブレイクは $+0.34$ に強化されますが、どちらの変化も重大ではありません。 \mbox{AgentHarm} は、機能を制御した後の 3 つのテンプレートのジェイルブレイク安全性との最も強い関連性、$\rho{=}{+}0.72$ を示しています。しかし、どちらの機器も有害なコンプライアンスをスコアしているため、これは一般的な安全性ではなく、収束した妥当性の証拠です。ベンチマーク、メトリクス、ターゲット動作、およびモデルパネルの名前を付けることは、安全性を主張するために最低限必要なことです。

原文 (English)

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2\pi/(1+\pi)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|\rho| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($\rho{=}{+}0.60$) but correlates negatively with misalignment safety ($\rho{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $\Delta{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $\rho{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

SciToolAgent-Evo: オープンワールドの科学ツールを取得するためのオントロジーを認識した自己進化エージェント

大規模言語モデル (LLM) エージェントは、特殊な計算ツールを編成して呼び出すために科学研究で採用されることが増えています。ただし、静的セマンティクスを持つ事前定義されたツール空間に依存しているため、ツールの要件、機能、境界が動的に進化するオープンワールドの科学ワークフローへの適用性が制限されます。この目的を達成するために、オープンワールドの科学ツールを取得するためのオントロジーを認識した自己進化エージェントである SciToolAgent-Evo を提案します。スキル、経験、オントロジー化されたツール グラフの進化する記憶によって駆動され、蓄積中に対照的な軌跡から一般化可能な知識を抽出します。一方、推論中にアクティブなリクエストを定式化し、LinUCB ベースのバンディット ゲートを利用して探索と活用の動的バランスをとります。新しいツールを取得すると、その科学的オントロジーがオンラインで完成し、既知のグラフにシームレスに統合されます。さらに、4 つの難易度にわたる 900 の現実的なタスクを含むベンチマークである OpenSciToolBench を紹介します。広範な評価により、SciToolAgent-Evo が最先端のパフォーマンスを実現し、その堅牢性と汎用性が検証されたことが示されています。

原文 (English)

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to open-world scientific workflows, where tool requirements, capabilities, and boundaries evolve dynamically. To this end, we propose SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition. Driven by an evolving memory of skills, experiences, and an ontologized tool graph, it distills generalizable knowledge from contrastive trajectories during accumulation, whereas during inference, it formulates active requests and utilizes a LinUCB-based bandit gate to dynamically balance exploration and exploitation. Once a novel tool is acquired, its scientific ontology is completed online for seamless integration into the known graph. Moreover, we introduce OpenSciToolBench, a benchmark containing 900 realistic tasks across four difficulty levels. Extensive evaluations show that SciToolAgent-Evo achieves state-of-the-art performance, validating its robustness and generalization.

13:00 JST研究/論文

EarlyDx: 証拠に裏付けられた ED エンカウンター診断のオープンエンド生成のためのアドミッションアンカー型ベンチマーク

入院時の臨床診断は、限られた不完全な証拠に基づいて迅速に行われなければなりません。既存の診断予測ベンチマークは、この設定にはあまり適していません。予測をクローズド コード セットに制限し、自由記述メモを除外し、入院患者の全コースを組み込んだ退院診断で監視します。 MIMIC-IV での 154,834 件の救急外来での診察に基づいて構築された、オープンエンドの早期診断のための大規模ベンチマークである EarlyDx を紹介します。各遭遇は入院時 $t_0$ で入手可能な記録に制限され、退院時ではなく救急外来で記録された診断によって管理されます。 LLM 監査人は、すべての自由テキストラベルがその証拠によってサポートされているか、部分的にサポートされているか、またはサポートされていないかをさらに検証します。一次評価スコアは完全にサポートされているラベルのみです。セマンティックな LLM-as-judge プロトコルの下では、フロンティアジェネラル、医療専門、またはドメイン内でトレーニングを受けたシステムなど、入院時の証拠を確実に合成する評価システムはありません。ゼロショット モデルは主に抽出によってスコアを付け、記録から読み取るのではなく推測する必要がある診断のわずか 3 ~ 31% しか回復しません。トレーニング後の推論依存の再現率は 56% に上昇しますが、かなりのマージンが残っており、時間が重要な条件では、臨床医の感度と精度のバランスを達成できるシステムはありません。ここで完全な構築および評価パイプラインをリリースします。

原文 (English)

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supervise with discharge diagnoses that incorporate the full inpatient course. We introduce EarlyDx, a large-scale benchmark for open-ended early diagnosis, built from 154,834 emergency department encounters in MIMIC-IV. Each encounter is restricted to records available at admission time $t_0$ and supervised by the diagnoses recorded during the ED encounter rather than at discharge. An LLM auditor further verifies every free-text label as supported, partially supported, or unsupported by that evidence; the primary evaluation scores only fully supported labels. Under a semantic LLM-as-judge protocol, no evaluated system --- frontier general, medical-specialized, or in-domain post-trained --- synthesizes admission-time evidence reliably. Zero-shot models score largely by extraction, recovering only 3-31% of diagnoses that must be inferred rather than read from the record; post-training raises inference-dependent recall to 56%, but a sizeable margin remains, and on time-critical conditions no system attains a clinician's balance of sensitivity and precision. We release the full construction and evaluation pipeline at here.

13:00 JSTエージェントビジネス/資金調達

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention woul…

13:00 JSTLLM/生成AI

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptab…

13:00 JST研究/論文

不完全な調整下での価値の脆弱性

AI システムに課せられる責任が増すにつれて、これらのシステムが人間性と整合していることを保証することがますます重要になります。 AI の安全性に関する一般的な懸念は、人間の価値は脆弱であるということです。つまり、人間の価値を不完全に代替するために過度に最適化すると、壊滅的な結果につながるということです。この論文では、エージェントが世界を最適化する前にその価値関数が代理条件を満たすことを保証する理想的なアライメント トレーニングを受けるアライメント問題のモデルを紹介します。私たちの主要な結果は、人間の価値関数に関する条件と、$\eta$-壊滅的な価値関数を持つエージェント、つまり最適化能力の限界において人間の価値の期待値が $\eta$ を下回ることが保証されるエージェントが配備される場合のいくつかの代用条件の精度を特定しました。私たちの結果は、過剰最適化の危険性を浮き彫りにし、導入前のトレーニングのみに依存するのではなく、量子化器などの最適化圧力を制限する AI 設計を動機付けるものです。

原文 (English)

Fragility of Value under Imperfect Alignment

As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.

13:00 JST研究/論文

ベイジアン実験計画による認知パラメーター推論のための有益な環境の特定

計算による認知モデリングは、観察された行動の根底にある潜在的な認知メカニズムを推測しようとします。ベイジアン逆計画は、そのような推論のための原則に基づいたフレームワークを提供しますが、その成功は実験環境に大きく依存します。既存のアプローチは通常、環境を固定されたものとして扱い、どの認知実験が認知パラメータの推論に最も有益であるかという問題は未解決のままです。実験環境を設計変数として扱い、認知計画実験の設計をベイズ実験計画 (BED) 問題として定式化します。私たちは正確なモンテカルロ BED ベンチマークを確立し、効率的な事後推論と設計評価のために償却ベイズ実験計画フレームワークを導入します。 Mouselab-MDP プロセス トレース パラダイムに関する実験では、計算コストを大幅に削減しながら、償却 BED が正確なモンテカルロ BED の環境ランキングとほぼ一致することが示されています。さらに、認知推論の目的全体にわたって均一に最適な単一の環境は存在しないことを示し、期待される情報獲得、事後回復可能性、情報効率の間のトレードオフを明らかにします。これらの結果は、ベイジアン パラメーター推論のための有益な認知実験を設計するための原則的なフレームワークを提供します。

原文 (English)

Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design

Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provides a principled framework for such inference, but its success depends critically on the experimental environment. Existing approaches typically treat environments as fixed, leaving open the question of which cognitive experiments are most informative for cognition parameter inference. We formulate the design of cognitive planning experiments as a Bayesian Experimental Design (BED) problem, treating the experimental environment as the design variable. We establish an exact Monte Carlo BED benchmark and introduce an amortized Bayesian experimental design framework for efficient posterior inference and design evaluation. Experiments on the Mouselab-MDP process-tracing paradigm show that amortized BED closely matches the environment rankings of exact Monte Carlo BED while substantially reducing computational cost. We further show that no single environment is uniformly optimal across cognitive inference objectives, revealing trade-offs between expected information gain, posterior recoverability, and information efficiency. These results provide a principled framework for designing informative cognitive experiments for Bayesian parameter inference.

13:00 JSTLLM/生成AIエージェント

NeSyFS: 部分可観測性の下での LLM エージェントのための神経記号的な高速-低速思考フレームワーク

最近、大規模言語モデル (LLM) は、内省、検索拡張生成、科学的発見などのアプリケーションで自律エージェントとして導入されることが増えています。このような設定では、エージェントは完全な環境状態ではなく、限られた観察に基づいて行動する必要があるため、部分的な観察可能性が生じます。これにより、信念状態の推論、タスクの目的の不一致、不確実性の下での計画という、いくつかの重要な課題が生じます。従来のアプローチは通常、完全なまたは要約された行動観察履歴に基づいて行動を条件付けており、その冗長で無関係な情報が LLM エージェントの意思決定を誤解させる可能性があります。人間の認知にインスピレーションを得て、LLM エージェント用の新しい神経記号高速思考 (NeSyFS) フレームワークを提案し、統一的なアプローチで部分的な可観測性によってもたらされる課題に対処します。ナレッジ グラフ (KG) を使用して信念状態を表し、NeSyFS のすべてのモジュールのコンテキストとしてトリプレットを提供します。高速思考モジュールは事後対応を実行しますが、低速思考モジュールはツイストシーケンシャル モンテカルロ (TSMC) アルゴリズムの高レベル構造に従って、不確実性を考慮した新しい計画を実行します。タスクの目標のずれを軽減するために、リフレクション モジュールを使用して高速思考のアクションを反映し、また、事後対応のアクションが繰り返し失敗するたびに、低速思考のモジュールに切り替えます。 ALFWorld、Webshop、ScienceWorld という 3 つの代表的なベンチマークでの実験では、以前の方法に比べて大きな利点が実証されました。

原文 (English)

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability

Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based on limited observations rather than full environmental states, leading to partial observability. This introduces several key challenges: belief state inference, task objective misalignment, and planning under uncertainty. Prior approaches typically condition actions on full or summarized action-observation histories whose redundant and irrelevant information can mislead the decision making of LLM agent. Inspired by human cognition, we propose a novel neuro-symbolic fast-slow thinking (NeSyFS) framework for LLM agent, addressing the challenges introduced by partial observability in a unified approach. We use a knowledge graph (KG) to represent the belief state, providing triplets as context for every module of NeSyFS. The fast-thinking module performs reactive action, while slow-thinking conducts a new uncertainty-aware planning by following the high-level structure of twisted sequential Monte Carlo (TSMC) algorithm. To mitigate the misalignment of task objective, a reflection module is used to reflect fast-thinking actions, and also switches to the slow-thinking module whenever reactive actions repeatedly fail. Experiments on three representative benchmarks, i.e. ALFWorld, Webshop, and ScienceWorld, demonstrate significant advantages over previous methods.

13:00 JSTLLM/生成AIエージェント研究/論文

MerchantBench: E コマース運用における長期的な一貫性のための LLM エージェントのベンチマーク

大規模な言語モデル エージェントは自律型ツール ユーザーとしての評価が高まっていますが、ほとんどのベンチマークは即時の成功基準を備えた限定されたタスクに焦点を当てています。実際の展開では、多くの場合、長期的な一貫性、つまり蓄積された証拠に意思決定を適応させながら、長期にわたる一貫性が必要になります。この能力を評価するには、アクションが将来の選択を制約し、フィードバックが不均一な遅延で到着し、一貫性のない行動が測定可能な累積効果を生み出す永続的な環境が必要です。販売者側の電子商取引は、製品の調達、リストと価格の管理、キャッシュ フロー管理、および混合遅延フィードバックの適応に関する反復的かつ相互依存的な決定を通じて、この評価に適切な設定を提供します。 MerchantBench は、98,843 件の実際の電子商取引商品レコードに基づいた 365 日の注文レベルのシミュレーションであり、エージェントとの対話のための 26 のツールを備えています。 MerchantBench は、すぐに観察できる上流のサプライヤー イベントと遅延した下流の注文結果を組み合わせて、エージェントに個別の注文ライフサイクルに従い、以前の決定を再検討することを要求します。 2 つのエージェント フレームワークの下で 8 つの LLM を 48 回の実行で評価し、それぞれのシミュレーション期間は 365 日です。私たちの結果は、最新の LLM と人間の参加者の間にさえ大きなギャップがあることを明らかにしており、最適な LLM 構成では人間の参加者が達成した平均最終純資産の 27.3\% しか達成していません。

原文 (English)

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.

13:00 JSTエージェント

ターンレベルのエージェント的 RL のための科学的発見環境の拡張

大規模言語モデル エージェントは、エージェントが実行環境と対話して統計的主張を生成する、データ駆動型の科学的発見タスクにおいて有望な機能を示しています。長期的な科学分析は、現実世界の科学データを対象としたプロセス監視環境の欠如によって依然として制約を受けています。このペーパーでは、プロセス検証可能な環境で Scientific Discovery エージェントをトレーニングするためのスケーラブルなフレームワークである SciDisco を紹介します。 SciTh\`eque は、仮説、データセット、隠された証拠グラフ、および検証ツールをタスク環境にコンパイルし、対話中に分析の進行状況を確認できます。 DAG 接地軌道合成では、これらの環境を使用して、検証器でフィルター処理されたマルチターン デモンストレーションを構築します。その後、DiscoPO は環境をトレーニング信号のソースとして使用し、検証可能な分析証拠を生み出すアクションにターンレベルのクレジットを割り当てます。実験では、SciDisco-14B が仮説に基づく科学データ分析ベンチマークにおいて最先端に達していることが示されています。

原文 (English)

Scaling Scientific Discovery Environments for Turn-Level Agentic RL

Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciTh\`eque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.

13:00 JSTエージェント研究/論文

MMShopBench: マルチモーダル、マルチターン ショッピング エージェント向けのリアルログ ベンチマーク

オンライン買い物客は、テキストだけでは表現しにくい製品のニーズを表現し洗練するために、画像や複数回の対話を使用する AI ショッピング アシスタントにますます注目しています。しかし、既存のベンチマークは主にテキストのみのリクエストまたは合成リクエストに依存しており、画像と言語を組み合わせて表現される複雑な現実世界のショッピング要件が十分に表現されていません。マルチモーダル、マルチターン ショッピング エージェント向けの初の実対数ベンチマークである MMShopBench を紹介します。慎重にクリーニングされ、手動で注釈が付けられたショッピング ログから構築された MMShopBench は、各リクエストの購入意図と必須の製品要件についての正確な注釈を提供します。エージェントは、ユーザーの画像と複数回の対話からこれらの要件を共同で推測し、画像とテキストの検索を通じて候補製品を取得し、製品画像と構造化された属性を使用して各候補がすべての要件を満たしていることを確認する必要があります。私たちは、証拠に基づいたマルチモーダルプロトコルを使用して代表的なオープンソースモデルと独自モデルを評価し、オープンソースモデルを微調整するためのコンパニオントレーニングセットを構築します。再現可能な実験を保証するために、オフライン ショッピング サンドボックスを構築します。そこでの微調整により、オープンソース モデルと主要な独自モデルの間のパフォーマンスのギャップが大幅に縮小され、トレーニング データの有効性が実証されます。

原文 (English)

MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents. Built from carefully cleaned and manually annotated shopping logs, MMShopBench provides ground-truth annotations of each request's purchase intent and mandatory product requirements. Agents must infer these requirements jointly from user images and multi-turn dialogue, retrieve candidate products through image and text search, and verify that each candidate satisfies all requirements using its product images and structured attributes. We evaluate representative open-source and proprietary models using an evidence-grounded multimodal protocol and construct a companion training set for fine-tuning an open-source model. To ensure reproducible experimentation, we build an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between our open-source model and leading proprietary models, demonstrating the effectiveness of our training data.

13:00 JST研究/論文

Evidence-Grounded Constraint Checking in Construction Documents

Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and docum…

13:00 JST研究/論文GemmaQwen

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where…

13:00 JSTビジネス/資金調達Google

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input re…

13:00 JSTLLM/生成AIGPT / ChatGPT

Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration

Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capab…

13:00 JSTLLM/生成AIエージェント

CAGE: ツール使用エージェントに対する型付きリターンの不確実性の下での認定認可

ツールを使用する LLM エージェントは、型指定されたツールの戻り値に基づいて動作し、来歴とカテゴリフィールドを数値と組み合わせて記録します。実行時許可ゲートは通常、観察されたリターンとアクションを許可し、リターンがソースにどのようにバインドされているかに関する小さなエラーに対して決定を保護しないままにします。私たちは、候補アクションが、宣言された妥当な正しく限界のあるリターンの近傍 (許容可能な束縛フォールト 1 つと限界のある数値ドリフト) に対して認可されたままであるかどうかを尋ねます。私たちは、カテゴリチャネルと数値チャネルを個別に認定することは構成されないことを証明します。各チャネル単独では安全である摂動が共同して同じアクションを安全でなくなる可能性があります。 CAGE は、この結合近傍を直接認定し、離散ブランチを正確に列挙し、各ブランチ内の連続摂動を認定します。 CAGE は、合成、コードとしてのポリシー、規制、および実際のトランザクション設定全体にわたって、予算内の誤った問題を削除し、正確なポイントワイズ ゲートが許可する一方で、有用な部分の意思決定を自律的に保ちます。ポリシーが実行可能である場合、CAGE-Exact はポリシー自体を認証します。それ以外の場合、CAGE-Lip と CAGE-RS は、明示的に測定された忠実度の仮定に基づいて学習されたゲートを認証します。

原文 (English)

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected against small errors in how the return was bound to its source. We ask whether a candidate action stays authorized over a declared neighborhood of plausible correctly bound returns: one admissible binding fault plus bounded numerical drift. We prove that certifying the categorical and numerical channels separately does not compose: perturbations that are safe on each channel alone can jointly turn the same action unsafe. CAGE certifies this joint neighborhood directly, enumerating the discrete branches exactly and certifying the continuous perturbation within each branch. Across synthetic, policy-as-code, regulatory, and real-transaction settings, CAGE removes the in-budget false allows that accurate pointwise gates admit, while keeping a useful fraction of decisions autonomous. When the policy is executable, CAGE-Exact certifies the policy itself; otherwise CAGE-Lip and CAGE-RS certify a learned gate under an explicit, measured fidelity assumption.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unf…

13:00 JSTLLM/生成AI

報酬を混合しないでポリシーを混合してください: 複数の報酬 RL のためのポリシーの分解と最適化

最新の大規模言語モデル (LLM) は、正しく答えるだけでなく、その動作を人間のさまざまな価値観やユースケースに適応させることが期待されています。その結果、各報酬が望ましい行動の異なる側面を捉える、複数報酬強化学習 (RL) が LLM にとってますます重要な問題となっています。ただし、複数の報酬を伴う最適化では、より深刻な調整税の問題が発生し、さまざまな最適化目標が互いにトレードオフしたり、競合したりする可能性があり、トレーニング後の不安定で非効率的な結果につながります。この研究では、政策空間の分解と合成のアイデアに基づいて構築された新しい複数報酬 RL フレームワークである PRISM を提案します。 PRISM は、さまざまな報酬を合成する代わりに、一連のスタンドアロンのポジティブ ポリシーとグローバル ネガティブ ポリシーを最適化します。これにより、複数報酬ポリシーの最適化中の潜在的な競合が軽減されると同時に、柔軟なポリシー構成による推論中の制御性が可能になります。科学的推論、ツール使用推論、および有用性と安全性の調整に関する実験では、PRISM が推論時間の優先制御のための追加の制御性を備え、既存の複数報酬 RL ベースラインを常に上回っていることが示されています。

原文 (English)

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization objectives can trade off or even conflict with each other, leading to unstable and inefficient post-training. In this work, we propose PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition. Instead of compositing different rewards, PRISM optimizes a set of standalone positive policies and a global negative policy. This alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition. Experiments on scientific reasoning, tool-use reasoning, and helpfulness-safety alignment show that PRISM consistently outperforms existing multi-reward RL baselines, with extra controllability for inference-time preference control.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

ツールの仕様が重要: AI エージェントの安全性リスクを明らかにし、軽減する

AI エージェントは、外部ツールを使用して大規模言語モデル (LLM) を拡張し、複雑なタスクを実行し、モデルの出力を結果として現実世界のアクションに変換できるようにします。しかし、LLM はエージェントとして導入されると安全性が大幅に低下することが多く、この低下の原因は依然としてよくわかっていません。この論文では、スキーマ形式のツール仕様がエージェントの安全性低下の主な原因であることを特定し、ホワイトボックス表現分析を通じて、それらがモデルの内部拒否シグナルを弱め、安全でないツールの実行に寄与していることを示します。この発見に基づいて、私たちは安全性の判断をツールの実行から切り離す推論時の保護手段である SafeKeep を提案します。これは、元のスキーマ形式の実行仕様を保持しながら、平坦化されたテキストのツール仕様を使用してリクエストを評価します。 2 つの代表的なベンチマークと、ホワイト ボックス モデルとブラック ボックス モデルの両方を含む 4 つの LLM にわたって、SafeKeep は有害なリクエストの平均拒否率を 23.8% から 70.6% に増加させ、観測レベルのプロンプト インジェクションの下での平均攻撃成功率を 25.6% から 2.5% に減少させます。また、既存の安全対策よりも優れたパフォーマンスを発揮し、タスク処理能力を維持します。コードとデータは https://github.com/snowcatsmoking/SafeKeep でリリースされます。

原文 (English)

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution. Across two representative benchmarks and four LLMs, including both white-box and black-box models, SafeKeep increases the average refusal rate for harmful requests from 23.8% to 70.6% and reduces the average attack success rate under observation-level prompt injection from 25.6% to 2.5%. It also outperforms existing safeguards and preserves task-handling capability. We release the code and data at https://github.com/snowcatsmoking/SafeKeep .

13:00 JSTエージェント

MAGA: 構造化アクションの蒸留による GUI エージェントのマルチプラットフォーム自己融合

大規模な言語モデルに基づくグラフィカル ユーザー インターフェイス (GUI) エージェントは、モバイル、Web、デスクトップ環境全体にますます導入されています。ただし、既存のエージェントは通常、ドメイン固有であるため、展開とユーザー エクスペリエンスが制限されます。これにより、特殊なモデルを単一の環境にまたがるポリシーに統合することが促進されます。ウェイト マージでは、ドメイン固有のエキスパートが直接マージされますが、エキスパートの意見の不一致により実行可能なアクションが破損する可能性があります。一方、オンポリシー蒸留 (OPD) では、教師の監督の競合を回避しながらも、蒸留中にすべてのレスポンス トークンが平等に扱われ、アクション トークンが環境とエージェントの間の唯一のインターフェイスであることが無視されます。これに対処するために、構造化されたアクションに従ってトレーニング信号を再割り当てする MAGA を導入します。生成されたアクションの正しさに基づいて、不要または無効な蒸留信号を抑制し、誤ったアクションに焦点を当てて学習します。さらに、トレーニング専用のヒントは、生徒の入力を変更することなく、分野固有の教師によって提供される監督信号を最適化します。 2 つのモデル スケール全体で、MAGA は最高の平均成功率を達成し、8B で最も強力なベースラインを 2.0% 上回り、教師とほぼ同じ平均パフォーマンスを達成しました。

原文 (English)

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, existing agents are typically domain-specific, limiting the deployment and user experience. This motivates the consolidation of specialized models into a single cross-environment policy. Weight merging directly merges domain-specific experts but can corrupt executable actions under expert disagreement, while on-policy distillation (OPD) avoids conflicting teacher supervision yet still treats all response tokens equally during distillation, ignoring that action tokens are the only interface between the environment and the agent. To address this, We introduce MAGA that re-allocates training signal according to the structured action. Based on the correctness of the generated action, it suppresses unnecessary or invalid distillation signals and focuses learning on erroneous actions. Besides, a training-only hint optimizes the supervision signal provided by domain-specific teachers without changing the student input. Across two model scales, MAGA achieves the highest mean success rate, outperforming the strongest baseline by 2.0% at 8B and achieves almost the same average performance with teachers.

13:00 JSTエージェント

コンポーネントのテストを超えて: Agentic AI システムの検証

エージェントティック AI システムは、計画、ツールの使用、記憶、対話、適応を組み合わせた複数ステップの軌道を通じて機能します。この動作により、コンポーネントのテストやワンショットの入出力評価を超えて検証の実践が拡張されます。これは、許容されるシステムの動作が時間の経過と環境条件の変化の下で意思決定がどのように展開されるかに依存するためです。この調査では、エージェント システムの検証問題を特徴付けるために、エージェントの評価、ソフトウェア アシュアランス、サイバー物理システム、ランタイム モニタリング、規制ガイダンスに及ぶ 257 件の論文を統合しています。このレビューは、行動、安全、一時的、規制、および複数のエージェントに関する懸念をカバーする 5 次元の分類を中心に構成されており、その分類を使用して現在のアプローチをマッピングし、再発する適用範囲のギャップを明らかにします。分析は、行動評価が比較的成熟している一方で、時間的妥当性、実行時の証拠の維持、規制の可読性、およびオープンエンドのマルチエージェント システム保証がまだ開発されていないことを示しています。 3 つのクロスドメインのケーススタディ (医療、産業運営、スマート モビリティ システム) では、レビュー済みの文献に記載されている故障パターンに基づいて、セーフティ クリティカルな環境で 5 つの分類の次元がどのように繰り返されるかを運用上の図で示しています。この論文は、境界付き自律性仕様、敵対的軌道生成、実行時モニタリング、監査対応証拠構造を中心としたライフサイクル指向の研究課題で締めくくられています。中心的な主張は、エージェント AI の信頼できる展開は、分離されたコンポーネントのみを評価するのではなく、コンテキスト内の軌跡を検証することに依存しているということです。

原文 (English)

Beyond Component Testing: Validating Agentic AI Systems

Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems. The review is organized around a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns, and uses that taxonomy to map current approaches and expose recurrent coverage gaps. The analysis shows that behavioral evaluation is comparatively mature, while temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent systems assurance remain under-developed. Three cross-domain case studies (medical care, industrial operations, smart-mobility systems) provide operational illustrations of how the five taxonomy dimensions recur in safety-critical settings, grounded in the failure patterns documented in the reviewed literature. The paper concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures. The central claim is that trustworthy deployment of agentic AI depends on validating trajectories in context rather than assessing isolated components alone.

13:00 JSTLLM/生成AIビジネス/資金調達ClaudeGPT / ChatGPT

ModelEquivBench: LLM で生成された最適化モデルのマルチリレーショナル評価の認証

大規模な言語モデルでは、自然言語から最適化モデルを生成するケースが増えていますが、既存の評価では、生成されたモデルとそのグラウンド トゥルースが、単一の同等/非同等判定または実行成功率、つまり独立してチェック可能ではなく、2 つの定式化が一致する複数の異なる意味に忠実でもないラベルに縮小されることがよくあります。我々は、ペアごとのセマンティック プロファイル E0 ~ E6 を報告する認証済みのマルチリレーショナル評価システムである ModelEquivBench を紹介します。モデルの構築と正確な取り込み (E0)、検証された表現の位置合わせ (E1)、同一空間と投影された実行可能集合関係 (E2、E3)、目的順序の等価性 (E4)、最適値の等価性 (E5)、およびオプティマイザー セットの等価性 (E6) です。決定された各エントリには、関係に適した、独立して再チェック可能な証拠が含まれます。つまり、E0 ~ E1 については再生可能なトレースまたは明示的なマップ、肯定的な E2 ~ E6 の結論については正確に合理的な証明書、およびサポートされた否定については明示的な証人です。不完全なマッピング検索、サポートされていない構造、およびリソース制限により、推測ではなく型指定された UNKNOWN または N/A の結果が生成されますが、満たされていない前提条件は ABSENT として報告されます。 ModelEquivBench を使用して、GPT-5.4、Claude Sonnet 4.6、および Qwen3.5-397B-A17B の 3 つのモデル スナップショットを、修復なしプロトコルの下で 173 の基本問題 (モデルあたり 346 セル) の同じ凍結コホートで評価すると、結果のプロファイルは、粗いベースラインでは表現されない区別を明らかにします。49、35、および 25 セルには、次の実行可能候補が含まれています。それにもかかわらず、少なくとも 1 つのサポートされている関係で否定的と証明され、E2 が検証済みマップの下でマップされた実現可能集合の等価性を証明するペアで 25、8、および 18 の構造的拒否が発生します。 3 つのモデルのスナップショットはプロファイルのさまざまな段階で失敗するため、単一の精度スコアに有意に削減することはできません。

原文 (English)

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system that reports a per-pair semantic profile E0--E6: model construction and exact ingestion (E0), verified representation alignment (E1), same-space and projected feasible-set relations (E2, E3), objective-order equivalence (E4), optimal-value equality (E5), and optimizer-set equivalence (E6). Each decided entry carries relation-appropriate, independently re-checkable evidence: replayable traces or explicit maps for E0--E1, exact-rational certificates for positive E2--E6 conclusions, and explicit witnesses for supported negatives. Incomplete mapping search, unsupported structure, and resource limits produce typed UNKNOWN or N/A outcomes rather than guesses, while unmet prerequisites are reported as ABSENT. Using ModelEquivBench to evaluate three model snapshots--GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B--on the same frozen cohort of 173 base problems (346 cells per model) under a no-repair protocol, the resulting profiles expose distinctions that coarse baselines do not represent: 49, 35, and 25 cells contain executable candidates that are nevertheless certified negative on at least one supported relation, and 25, 8, and 18 structural rejections occur on pairs for which E2 certifies mapped feasible-set equality under a verified map. The three model snapshots fail at different stages of the profile and therefore cannot be meaningfully reduced to a single accuracy score.

13:00 JSTエージェント

検索を超えて: マルチモーダル エージェントの分析メモリ

長期マルチモーダルメモリは、関連情報の取得だけでなく、インタラクション全体で蓄積された観測値の計算もサポートする必要があります。既存のシステムは主に \emph{検索メモリ} を重視しており、概要とインデックスを通じてインタラクション履歴を整理し、高レベルの抽象化から基礎となるレコードに至るまで複数の粒度でクエリ関連情報を返します。この論文では、フィルタリング、集計、ランキング、時間比較をサポートするクエリ可能な構造に繰り返し発生する多峰性観測を整理する補完的な抽象化として \emph{分析記憶} を定式化します。検索と分析記憶を共同でサポートするフレームワークである AdaMM を紹介します。 AdaMM は、アプリケーション定義のスキーマに依存するのではなく、対話、画像、およびコンテキスト メタデータから来歴にリンクされた属性値の観察結果を抽出し、繰り返し発生するフィールド構造を発見し、分析アクセスのためにそれらを具体化します。推論時に、メモリ対応プランナーはクエリを取得操作と分析操作に分解し、各操作を適切なツールにルーティングします。 2 つの長期マルチモーダル メモリ ベンチマーク、MemEye と MemGallery の実験では、AdaMM がそれぞれ最大 11.3\% と 7.3\% パフォーマンスを向上させることが示されています。

原文 (English)

Beyond Retrieval: Analytic Memory for Multimodal Agents

Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring multimodal observations into queryable structures supporting filtering, aggregation, ranking, and temporal comparison. We present AdaMM, a framework that jointly supports retrieval and analytic memory. Rather than relying on application-defined schemas, AdaMM extracts provenance-linked attribute-value observations from dialogue, images, and contextual metadata, discovers recurring field structures, and materializes them for analytical access. At inference time, a memory-aware planner decomposes queries into retrieval and analytic operations and routes each operation to the appropriate tools. Experiments on two long-term multimodal memory benchmarks, MemEye and MemGallery, show that AdaMM improves performance by up to 11.3\% and 7.3\%, respectively.

13:00 JSTエージェント研究/論文

セルフプレイとスキルの進化の出会い: 問題を提起し、解決し、記憶する自己進化する検索エージェント

セルフプレイ エージェントは、ターゲット ベンチマークからの質問なしでトレーニング問題を生成できますが、そのカリキュラムには永続的な状態がありません。失敗は勾配に影響を与えますが、将来の実践を明示的に形成するものではありません。外部スキルの記憶は手順の経験を保存しますが、通常は固定タスクの分布から学習されます。 \textbf{SESA} (Self-Evolving Skill-Augmented Agent) を導入します。これにより、手続き型記憶がツール拡張検索の自己再生の進化した状態になります。チャレンジャーが問題を提起する一方で、個別にパラメーター化されたソルバーのみがスキルを取得します。有益な失敗は再利用可能なスキルに抽出され、メモリに書き戻されます。更新されたメモリはソルバーの動作と成功を変化させ、それによって挑戦者の報酬と将来の問題の分布が変化します。結果として生じるフロンティアは、メモリを書き換える新たな障害を生み出します。この双方向ループにより、タスクの生成とスキルの記憶が共進化します。取得されたスキルはポリシーに基づいたトレーニングの軌道を形成するため、その利点は外部バンクに残るだけでなくモデルのパラメーターに入力することができ、メモリフリーの展開とオプションの推論時間の取得が可能になります。 7 つのオープンドメインおよびマルチホップ質問応答ベンチマーク全体で、SESA は、複数のバックボーンにわたって SSP と比較して平均精度を 1.2 ~ 3.2 ポイント向上させ、統一評価プロトコルの下でスキル強化された SkillRL ベースラインを 0.9 ポイント上回りました。 Qwen3 モデルでは、SESA-Off は SSP に比べて 1.8 ~ 2.2 ポイントの改善を維持しますが、最終スキル バンクはさらに 0.5 ~ 1.0 ポイントを追加します。これらの結果は、進化するスキル メモリが単なる推論時のプラグインではないことを示しています。オプションの外部メモリとしての価値を保持しながら、ポリシー学習と将来のトレーニング分布を変更します。コードは https://github.com/Zenghuang-Fu/SESA-Self-Evolving-Search-Agents で入手できます。

原文 (English)

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice. External skill memories preserve procedural experience but are typically learned from fixed task distributions. We introduce \textbf{SESA} (Self-Evolving Skill-Augmented Agent), which makes procedural memory an evolving state of tool-augmented search self-play. A challenger poses problems, while a separately parameterized solver alone retrieves skills. Informative failures are distilled into reusable skills and written back to memory. The updated memory changes solver behavior and success, which changes the challenger's reward and the distribution of future problems; the resulting frontier produces new failures that rewrite memory. This bidirectional loop makes task generation and skill memory co-evolve. Because retrieved skills shape on-policy training trajectories, their benefits can enter the model parameters as well as remain in the external bank, enabling memory-free deployment and optional inference-time retrieval. Across seven open-domain and multi-hop question-answering benchmarks, SESA improves average accuracy over SSP by 1.2--3.2 points across multiple backbones and surpasses the skill-augmented SkillRL baseline by 0.9 points under a unified evaluation protocol. On Qwen3 models, SESA-Off retains 1.8--2.2 points of improvement over SSP, while the final skill bank adds a further 0.5--1.0 points. These results show that evolving skill memory is not merely an inference-time plug-in: it changes policy learning and the future training distribution while retaining value as optional external memory. Our code is available at https://github.com/Zenghuang-Fu/SESA-Self-Evolving-Search-Agents.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPTGeminiDeepSeek

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers re…

13:00 JST研究/論文

COntExt: 運用メトリクスからのコンテキスト認識オントロジー拡張に向けて

組織は、システム、プロセス、コンプライアンスを監視するために、構造化された機械可読形式で運用指標を定義することが増えています。これらのメトリクス定義は、概念、プロパティ、関係の参照などのドメイン知識を暗黙的にエンコードし、多くの場合、正式なオントロジーでキャプチャされた内容を拡張します。しかし、運用メトリクスカタログとオントロジー知識との間の接続は依然として手作業で、その場限りで、労働集約的なものである。 COntExt は、構造化されたメトリクス定義を入力として受け取り、これらのメトリクスのコンテキストを利用して、参照される概念とプロパティを既存のオントロジに統合する方法を提案する、コンテキストを意識したオントロジ拡張のフレームワークです。フレームワークは、拡張問題を 3 つのサブタスク (親クラスの予測、関係タイプの予測、データ プロパティの割り当て) として定義します。 4 つのサイバーセキュリティ オントロジーにわたって、タスクごとに異なるアルゴリズムを評価します。私たちの結果は、メトリクス由来のコンテキストが、関係タイプの予測とデータ プロパティの割り当てについて、オントロジー コンテキストのベースラインよりも提案を改善していることを示しています。私たちの研究は、運用メトリック カタログがオントロジー拡張の実用的かつ十分に活用されていないソースであることを示しています。この作業により、組織は手動エンジニアリングよりも大幅に低いコストでオントロジーを維持できるようになります。

原文 (English)

COntExt: Towards Context-Aware Ontology Extension from Operational Metrics

Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, processes, and compliance. These metric definitions implicitly encode domain knowledge, such as referencing concepts, properties, and relationships, that often extends what is captured in formal ontologies. Yet the connection between operational metric catalogues and ontological knowledge remains manual, ad-hoc, and labor-intensive. We present COntExt, a framework for context-aware ontology extension that takes structured metric definitions as input and suggests how referenced concepts and properties should be integrated into an existing ontology, utilizing the context of these metrics. The framework defines the extension problem as three sub-tasks: parent class prediction, relation type prediction, and data property assignment. Across four cybersecurity ontologies, we evaluate different algorithms for each task. Our results show that metric-derived context improves the suggestions over ontology-context baselines for relation type prediction and data property assignment. Our work demonstrates that operational metric catalogues are a practical and underexploited source for ontology extension. This work enables organizations to maintain their ontologies at a significantly lower cost than manual engineering.

13:00 JSTロボティクスMeta

キツネザル: 好みのフィードバックから多目的強化学習に合わせて調整することを学ぶ

強化学習 (RL) システムは通常、明確に指定された単一のスカラー報酬関数を使用してトレーニングされます。ただし、現実世界の意思決定タスクには、パフォーマンスと効率など、複数の競合する目標が含まれることが多く、その場合、グラウンドトゥルースの報酬関数を指定するのが難しいか、アクセスできません。多目的 RL (MORL) は報酬をベクトルとしてモデル化することでこのようなトレードオフに対処しますが、既存のアプローチは通常、目的ごとに明確に指定された報酬関数へのアクセスを前提としており、単一目的 RL が直面する同じ課題を継承しています。一方、優先ベースの RL (PbRL) は、人間のフィードバックからの報酬学習を通じて、事前に定義された報酬関数にアクセスせずに複雑なタスクを解決する大きな可能性を示していますが、主に単一目的の設定で研究されてきました。この研究では、エージェントが複数の人間の好みからインタラクティブに学習して最適な多目的ポリシーを学習する新しいフレームワークである、LEMUR: Learning to Align to Multi-Objective Reinforcement with Preference Facebook (好みフィードバックによる多目的強化学習) を使用してこのギャップを埋めます。私たちのアプローチは、人間のフィードバックからポリシーと複数の目標固有の報酬モデルを共同学習することで、エージェントが学習中に競合する目標のバランスを効果的にとれるようにします。私たちはさまざまなベンチマーク多目的タスクで LEMUR を評価し、経験的な結果はベースライン手法よりも優れたパフォーマンスを示しています。私たちの方法は、事前に定義された報酬関数を使用せずに多目的の意思決定タスクを解決するための有望な方向性を示しています。

原文 (English)

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground-truth reward functions are difficult to specify or inaccessible. While Multi-Objective RL (MORL) addresses such trade-offs by modeling rewards as vectors, existing approaches typically assume access to a well-specified reward function for each objective, inheriting the same challenges faced by single-objective RL. Meanwhile, Preference-based RL (PbRL) has shown great potential in solving complex tasks without access to a pre-defined reward function through reward learning from human feedback, yet has largely been studied in single-objective settings. In this work, we bridge this gap with LEMUR: Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback, a novel framework where an agent interactively learns from the preferences of multiple humans to learn optimal multi-objective policies. Our approach jointly learns policies and multiple objective-specific reward models from human feedback, enabling agents to effectively balance competing objectives during learning. We evaluate LEMUR on a variety of benchmark multi-objective tasks, and empirical results demonstrate its superior performance over baseline methods. Our method presents a promising direction for solving multi-objective decision-making tasks without pre-defined reward functions.

13:00 JST研究/論文

DungeonBench: ダンジョンズ & ドラゴンズの戦闘におけるルール豊富な戦術的推論のベンチマーク

ゲームやシミュレーターは、意思決定を測定可能な結果に変えることで貴重なベンチマークを作成しますが、現在のスイートの多くは、ルールに富んだ戦術的推論、つまりジオメトリ、タイミング、リソース、目的、ルールの相互作用がすべて同時に重要になる場合に適切に選択する能力をテスト中です。ダンジョンズ & ドラゴンズの戦闘における戦術的推論のベンチマークである DungeonBench を紹介します。これは、戦闘に関連する 2014 システム リファレンス ドキュメントのコンテンツの大部分をカバーするように構築されており、その効果はシミュレーターによって解決できる一方で、単純化された戦闘シミュレーターでは抽象化されがちな仕組みを保持しています。各ステップで、DungeonBench は完全な戦術的観察、保留中の決定、および移動、攻撃、呪文、反応、目的、準備、希少なリソースにわたる実行可能なオプションのインデックス付きリストを公開します。課題は、アクションエコノミー、クリーチャーの特性、戦場の形状、タイミングウィンドウ、将来の遭遇によって結果が左右される法的な選択を重視することです。 DungeonBench には 2 つのトラックがあります。1 回の戦闘でのローカルな戦術プレイを評価する Encounter と、永続的なヒット ポイント、呪文スロット、消耗品、準備、短い休憩のタイミングを通じて遭遇を結び付け、当面の戦術的優位性と将来の生存可能性をトレードオフするポリシーを強制する Day です。同じエンジン生成の意思決定ストリームは、ヒューリスティック コントローラー、言語モデル ポリシー、学習されたオプション ランカー、およびマスクされたアクションの強化学習エージェントをサポートします。私たちは、この共有された意思決定の流れに基づいて、フロンティア言語モデルの政策を評価します。結果は、完全な戦術観察がベンチマークを飽和させていないことを示しています。フロンティア政策は多くの場合、直接遭遇に勝利しますが、リンクされた遭遇日では、リソースの予算配分、休憩のタイミング、ルールを意識した戦術規律の失敗が明らかになります。

原文 (English)

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference Document content whose effects can be resolved by the simulator while retaining mechanics that simplified combat simulators often abstract away. At each step, DungeonBench exposes a complete tactical observation, a pending decision, and an indexed list of executable options spanning movement, attacks, spells, reactions, objectives, preparation, and scarce resources. The task is to value legal choices whose consequences depend on action economy, creature traits, battlefield geometry, timing windows, and future encounters. DungeonBench has two tracks: Encounter, which evaluates local tactical play in single fights, and Day, which links encounters through persistent hit points, spell slots, consumables, preparation, and short-rest timing, forcing policies to trade off immediate tactical advantage against future survivability. The same engine-generated decision stream supports heuristic controllers, language-model policies, learned option rankers, and masked-action reinforcement-learning agents. We evaluate frontier language-model policies on this shared decision stream. Results show that full tactical observations do not saturate the benchmark: frontier policies often win direct encounters, but linked encounter days expose failures in resource budgeting, rest timing, and rule-aware tactical discipline.

13:00 JSTLLM/生成AIエージェント研究/論文

AgentHPOBench: LLM エージェントをシーケンシャル ハイパーパラメータ オプティマイザーとして評価するためのベンチマーク

LLM がコード補完システムから自律的な科学エージェントに進化するにつれて、実験を実行する能力を評価することがますます重要になっています。既存のベンチマークは通常、静的コードの生成、論文の複製、または最終的な回答の正しさに焦点を当てていますが、エージェントが実験証拠を解釈し、それをその後のハイパーパラメータ決定の指針として使用できるかどうかを直接評価することはありません。このギャップに対処するために、7 つの研究カテゴリにわたる 30 の実行可能な機械学習タスクで構成される逐次ベンチマークである AgentHPOBench を導入します。各タスクは検証されたベースラインの実行から始まり、その後エージェントがいくつかの連続した介入を実行します。各ステップで、エージェントは蓄積された構成、メトリック、ログを観察してから、次の有効な構成を提案します。私たちは、統一されたプロトコルの下で、広く使用されている 12 のエージェントと従来の HPO ベースラインを評価します。結果は、現在のエージェントがドメイン全体にわたって測定可能な実験最適化能力を示しているものの、持続的な反復改良、複雑なログ診断、および報告された参照パフォーマンスに向けた一貫した進歩において依然として明らかな限界に直面していることを示しています。

原文 (English)

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. To address this gap, we introduce AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories. Each task begins with a validated baseline run, after which an agent performs several sequential interventions. At each step, the agent observes the accumulated configurations, metrics, and logs before proposing the next valid configuration. We evaluate 12 widely used agents and conventional HPO baselines under a unified protocol. The results show that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent progress toward reported reference performance.

13:00 JST研究/論文

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effec…

13:00 JSTエージェント研究/論文Llama

ExtractBench: スキーマに基づいたエンタープライズ ドキュメント抽出のベンチマーク

エンタープライズ ワークフローでは、\emph{スキーマに基づく抽出} に関してエージェントへの依存度が高まっています。ドキュメントとユーザー定義のスキーマが与えられると、エージェントはそのスキーマに従って、根拠となるメタデータとしてソース証拠を含む正しい出力を生成します。私たちは、スキーマに基づいた抽出のベンチマークである ExtractBench を紹介します。これは、私たちの知る限り、値の精度をスコアリングし、大規模な完全性、根拠、および測定コストをまとめて記録した最初のベンチマークです。この評価システムには、370 の企業文書、8 つのビジネス ドメイン、および 67 の文書タイプにわたる 4,869 ページが含まれており、課題シナリオを区別する明確なタグが付いています。スケーラブルなスキーマとグラウンド トゥルース キュレーション パイプラインは、実際のドキュメントに対する独立したシステムの合意、合成リストに対する既知の値、およびフォームに対する人間による検証を組み合わせています。値の精度として順序に依存しない値 F1 と、ソースのトレーサビリティのための 2 つの基本指標 (ワード レベルとページ レベルの F1) を報告します。商用 VLM は短いドキュメントではうまく機能しますが、長いドキュメントではレコード リストが切り捨てられることがよくありますが、コーディング エージェントははるかに高いコストで高い精度を維持します。 LlamaExtract Agentic Plus は、わずかなコストでコーディング エージェントに匹敵する精度を備え、3 つの指標すべてで第 1 位にランクされています。データセットと評価コードは、\href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} および \href{https://github.com/run-llama/ExtractBench}{GitHub} で入手できます。

原文 (English)

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios. The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms. We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1. Commercial VLMs perform well on short documents but often truncate record lists on long ones, while coding agents retain higher accuracy at much higher cost. LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost. Dataset and evaluation code are available on \href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} and \href{https://github.com/run-llama/ExtractBench}{GitHub}.

13:00 JSTLLM/生成AI

Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks

Generative AI (GenAI) holds significant promise for advancing educational equity among ethnic minority students by broadening access to lea…

13:00 JSTLLM/生成AIハードウェア/半導体

Topology-Aware Data Movement for Disaggregated GPU Inference

Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…

13:00 JSTLLM/生成AIGemma

小規模言語モデルにおけるバイアスに対する知識蒸留の非対称的影響

小規模な命令調整言語モデルにおける知識の蒸留がバイアスに対して非対称な影響を与えることを示します。明確なタスク (BBQ-disambig) では、Gemma-2-9B 教師からの応答ベースの蒸留によりコンテキスト追従が向上します。最も偏ったベースライン (SmolLM2-1.7B-Instruct) では、コンテキストを優先するエラー率が 44% から 24% に削減されます。あいまいなタスク (BBQ-ambig) では、同じ蒸留により項目ごとの拒否キャリブレーションが破壊されます。全体的な拒否率が維持されている場合でも、ベースラインが正しく棄権された項目の 15% が、代わりにステレオタイプの回答を受け取ります。このパターンは 2 番目の生徒ファミリー (OLMo-2-1B-Instruct) でも再現され、沈黙の喪失が 8% で、沈黙の充満が新たなバイアスの 89% を占めています。 28 の構成グリッド全体にわたって、沈黙の損失と満たされた沈黙の大きさには相関がなく (Spearman $\rho=0.19$, n.s.)、2 つの効果が異なるメカニズムから生じることを示しています。集約されたステレオタイプ メトリクス (CrowS-Pairs、全体的な BBQ ステレオタイプ信頼性スコア) は両方の効果を平均し、項目ごとの害を隠します。キャリブレーションの損失はデータ側のメカニズムに起因すると考えられます。4 つのトレーニング コーパスの監査では、回答形状としての拒否が 0.5% 未満であることがわかりました。拒否注入を伴う教師あり微調整 (SFT) は、解析を中断するか、集計メトリクスが完全に調整されていると言える自明な拒否者体制 (拒否率 99.8%、曖昧さの解消精度 0.2%) に過剰修正します。我々は、拒否キャリブレーション、コンテキスト追従、機能維持を評価する 3 段階のプロトコルである条件別キャリブレーション診断 (PCCD) を提案します。 PCCD は、集計評価が見逃す非対称の害と自明な拒否障害モードの両方を捕捉します。

原文 (English)

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-disambig), response-based distillation from a Gemma-2-9B teacher improves context-following: for the most biased baseline (SmolLM2-1.7B-Instruct), it cuts the context-overriding error rate from 44% to 24%. On ambiguous tasks (BBQ-ambig), the same distillation destroys per-item refusal calibration: 15% of items where the baseline correctly abstained instead receive stereotype answers, even when overall refusal rate is preserved. The pattern reproduces on a second student family (OLMo-2-1B-Instruct), with silence-loss of 8% and filled-silence accounting for 89% of new bias. Across the full 28-configuration grid, the magnitudes of silence-loss and filled-silence are uncorrelated (Spearman $\rho=0.19$, n.s.), indicating that the two effects arise from distinct mechanisms. Aggregate stereotype metrics (CrowS-Pairs, overall BBQ Stereotype Reliance Score) average over both effects and conceal the per-item harm. We trace the calibration loss to a data-side mechanism: an audit of four training corpora finds <0.5% refusal-as-answer-shape. Supervised fine-tuning (SFT) with refusal injection either breaks parsing or over-corrects into a trivial-refuser regime (refusal rate 99.8%, disambig accuracy 0.2%) that aggregate metrics would call perfectly calibrated. We propose Per-Condition Calibration Diagnosis (PCCD), a three-step protocol that evaluates refusal calibration, context-following, and capability preservation. PCCD catches both the asymmetric harm and the trivial-refuser failure mode that aggregate evaluations miss.

13:00 JSTLLM/生成AIエージェント

形式主義の罠: 裁判官としての LLM 評価者は社会的負荷の下でコンセンサス模倣に盲目になっているのか?

\textit{エージェント形式主義の罠} と評価的不協和指数 ($D_E$) を導入し、LLM-as-a-Judge システムが敵対的な負荷の下で構造的手続き主義と意味論的真実をどのように混同するかを定量化します。 3 つのドメイン (GAIA、SWE ベンチ、マルチチャレンジ) にわたる 22,500 の軌跡を分析し、決定論的な語彙的根拠 ($p < 10^{-120}$) によって検証された幻覚操作の意味分類を抽出します。ロジスティック メタエバリュエーターは、このエバリュエーター キャプチャ (ROC-AUC 0.8779) の正確な構文トリガーを分離しますが、ゼロショット Leave-One-Domain-Out 転送は、脆弱性が普遍的にドメインに依存しないことを証明します (平均 ROC-AUC 0.7482)。アーキテクチャ プロファイリングにより、個別にシミュレートされた群トポロジが数学的に異質な意味論的な盲点を誘発することが明らかになり、アンカーされていない閉ループ評価は不安定で、体系的に発散し、アーキテクチャ固有の警戒フィルターが必要であることが証明されています。

原文 (English)

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectories across 3 domains (GAIA, SWE-bench, Multi-Challenge), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding ($p < 10^{-120}$). A logistic meta-evaluator isolates the exact syntactic triggers of this evaluator capture (ROC-AUC 0.8779), while a zero-shot Leave-One-Domain-Out transfer proves the vulnerability is universally domain-agnostic (mean ROC-AUC 0.7482). Architectural profiling reveals that distinct simulated swarm topologies induce mathematically disparate semantic blind spots, proving that unanchored closed-loop evaluation is unstable, systemically divergent and necessitates architecture-specific vigilance filters.

13:00 JST研究/論文

Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework

Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) f…

13:00 JST研究/論文

見た目も正しく、機能も正しく: マルチスクリーン モバイル アプリ生成のためのプロジェクト レベルのベンチマーク

最近のマルチモーダル大規模言語モデルでは、ビジュアル デザインを実行可能なコードに直接変換できますが、実際のモバイル製品では、共有コンポーネントと動作するナビゲーションを備えた構築可能なコードベースにするために複数のスクリーンショットが必要です。このプロジェクト レベルの設定では、既存のデザインからコードまでのベンチマークの 3 つの制限が明らかになります。ベンチマークは、完全なコードベースではなく単一ページの生成に焦点を当てていること、ページ間のナビゲーションを評価できないこと、プロジェクト全体の保守性を測定していないことです。実際のモバイル アプリ、人間がレビューした画面、構造化されたページ関係の注釈、およびナビゲーション テスト仕様で構成される、プロジェクト レベルのマルチスクリーン モバイル アプリ生成のための最初のベンチマークである MobileForge を紹介します。 MobileForge は、ビルド、ナビゲーション、視覚的忠実性、コードの保守性、効率性の 5 軸評価をサポートしています。また、ナビゲーション評価におけるカスケード障害を回避するための状態分離ナビゲーション テストと、視覚判定の信頼性を向上させるためのアンカー参照リスト単位の視覚評価プロトコルも提案します。現在のモデルは、6 つのフロンティア マルチモーダル LLM でエンドツーエンドで実行され、コンパイルして正しいページに到達するモバイル アプリ プロジェクトを構築できますが、インタラクティブ ナビゲーションの信頼性は依然として低く、視覚的な忠実性と保守性は依然として遅れています。ベンチマークとサポート資料は https://github.com/anoa12159-hue/mobileforge_eval から入手できます。

原文 (English)

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level setting exposes three limits of existing design-to-code benchmarks: they focus on single-page generation rather than complete codebases, cannot evaluate cross-page navigation, and do not measure project-wide maintainability. We introduce MobileForge, the first benchmark for project-level multi-screen mobile app generation, comprising real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. MobileForge supports five-axis evaluation of build, navigation, visual fidelity, code maintainability, and efficiency. We also propose state-isolated navigation testing to avoid cascading failures in navigation evaluation and an anchor-referenced list-wise visual evaluation protocol to improve visual-judge reliability. Across end-to-end runs on six frontier multimodal LLMs, current models can build mobile-app projects that compile and reach the correct pages, but interactive navigation remains unreliable and visual fidelity and maintainability still lag. The benchmark and supporting materials are available at https://github.com/anoa12159-hue/mobileforge_eval.

13:00 JST研究/論文

ConnectED: ベトナム語の授業計画と生徒の学習のためのカリキュラムに合わせた AI システム

この論文では、カリキュラムに沿った授業設計、対話型の生徒学習、フィードバック主導型の改善をリンクすることで、ベトナム教育における完全な指導ライフサイクルをサポートする人間中心の AI システムである ConnectED について紹介します。このシステムは、教師付き微調整と直接的な好みの最適化を通じてトレーニングされたベトナム語教育大規模言語モデルである VietEduQwen に基づいて構築されており、学術的に正確で、教育的に適切で、学生にとって安全な対話を保証します。 ConnectED は、Official Dispatch No. 5512/BGDDT-GDTrH に準拠した構造化プロンプト テンプレートを通じて ADDIE フレームワークを運用可能にし、各フェーズが生成ステップと教師の検証ゲートの両方として機能します。評価フェーズでは、生徒の成績データを反復的なレッスンの改善に結び付けることで、ループをさらに閉じます。このシステムは、レッスンの生成を超えて、生徒向けのインタラクティブな環境を統合し、教師の意思決定をサポートする学習シグナルの継続的な収集を可能にします。 2025 年のベトナム国家高等学校試験の 3,119 問を評価したところ、VietEduQwen は 87.02% の精度を達成し、Qwen3-8B を 6.10 パーセントポイント上回りました。教師 (n=18) と生徒 (n=214) を対象とした調査では、カリキュラムの調整、レッスンの明瞭さ、使いやすさに高い満足度が示されています。実際には、教師によるレビューにより、レッスンの準備時間は 3 ~ 4 時間から約 30 ~ 45 分に短縮されます。アブレーション研究では、DPO トレーニングと ADDIE ベースのオーケストレーションの両方がシステム パフォーマンスに独立して寄与していることが確認されており、実際の展開には構造化された教師による監督の重要性が強調されています。

原文 (English)

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum-aligned lesson design, interactive student learning, and feedback-driven refinement. Built on VietEduQwen, a Vietnamese educational large language model trained via supervised fine-tuning and direct preference optimization, the system ensures academically accurate, pedagogically appropriate, and student-safe interactions. ConnectED operationalizes the ADDIE framework through structured prompt templates aligned with Official Dispatch No. 5512/BGDDT-GDTrH, where each phase serves as both a generation step and a teacher validation gate. The Evaluation phase further closes the loop by connecting student performance data with iterative lesson improvement. Beyond lesson generation, the system integrates a student-facing interactive environment, enabling continuous collection of learning signals to support teacher decision-making. Evaluation on 3,119 questions from the 2025 Vietnamese National High School Examination shows that VietEduQwen achieves 87.02% accuracy, outperforming Qwen3-8B by 6.10 percentage points. Surveys of teachers (n=18) and students (n=214) demonstrate strong satisfaction with curriculum alignment, lesson clarity, and usability. In practice, lesson preparation time is reduced from 3--4 hours to approximately 30--45 minutes with teacher-in-the-loop review. Ablation studies confirm that both DPO training and ADDIE-based orchestration contribute independently to system performance, highlighting the importance of structured teacher oversight for practical deployment.

13:00 JSTLLM/生成AI

なぜ傷つくのか: 感情的なサポートの会話における否定的な思考の原因を特定する

大規模言語モデル (LLM) は、ネガティブな思考のリフレーミングなど、感情的なサポートのタスクに使用されることが増えています。このタスクは、認知的評価、つまり否定的な感情を引き起こす出来事の主観的な解釈を修正することに依存しており、これは通常、複数の個別の次元に沿って概念化されます。現在の LLM ベースのフレームワークは、考えられるすべての側面を徹底的に評価することで認知評価をモデル化していますが、さまざまなコンテキストにわたるこれらの側面の顕著性の変化を考慮できません。この研究では、「LLM は感情的なサポートの会話から顕著な評価の側面を推測できますか?」という重要でありながら見落とされている質問を調査します。この質問に対処するために、顕著な認知評価の側面を含む、人間による注釈付きの精神状態との 996 の感情サポート会話を含む AppraiSal ベンチマークを導入します。さらに、ベイジアン逆計画に基づいたマルチエージェント確率フレームワークである PRISM を提案します。これは、コンテキスト固有の評価次元を識別する LLM の能力を向上させるように設計されています。実験結果は、PRISM がさまざまなサイズの LLM に改善をもたらし、特に最も顕著な評価次元の特定において改善をもたらすことを示しています。

原文 (English)

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the subjective interpretation of events that elicit negative emotions, which is typically conceptualized along multiple discrete dimensions. Current LLM-based frameworks model cognitive appraisal by exhaustively evaluating all possible dimensions, but they fail to account for the varying saliency of these dimensions across different contexts. In this work, we investigate a vital yet overlooked question: "Can LLMs infer the salient appraisal dimensions from emotional support conversations?" To address this question, we introduce the AppraiSal benchmark, containing 996 emotional support conversations with human-annotated mental states, including salient cognitive appraisal dimensions. Furthermore, we propose PRISM, a multi-agent probabilistic framework grounded in Bayesian Inverse Planning, designed to improve LLMs' ability to identify context-specific appraisal dimensions. Experimental results show that PRISM brings improvements to LLMs across various sizes, particularly in identifying the most salient appraisal dimensions.

13:00 JST研究/論文

COSI-Lab: 多視点、多峰性の社会的意図をモデリングするためのカンファレンス リビング ラボ

COSI-Lab は、32 人の学者が参加する学際的な科学ワークショップのマルチモーダル、マルチセンサー データセットを国際会議で発表します。この作品は、30 分間の交流セッション 2 回からなる、脚本が弱い環境での生態学的に有効な社会的相互作用を捉えており、参加者にとって実際の専門的および社会的影響をもたらします。私たちは、将来のインテリジェント システムは、主観的な認識の多重性をラベル ノイズとしてではなく、説明可能な視点主導の推論プロセスとしてモデル化することで、主観的な認識を処理できるようになると主張します。私たちは、現場外の観察者によって決定される見かけの意図推論(AII)問題に焦点を当て、意図が明白な将来の結果から独立しているように概念化します。私たちは、1. 知覚者自身の解釈傾向を説明する AII の新しい注釈プロセス、2. 多様性、根拠、妥当性に関する意図の物語の定量的および定性的分析に貢献します。 3. AII および社会的関与などの周囲の関連する状況要因のベンチマークタスク。 4. すべての参加者に音声品質の音声を提供し、プライバシーを保護するマルチモーダルデータにより、語彙的および非言語的行動分析を可能にします。 5. 各参加者の自己報告目標 (30 分から 3 時間) と注釈付き AII (秒) の結合。

原文 (English)

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two 30-minute mingling sessions with real professional and social consequences for the participants involved. We argue that future intelligent systems could be better equipped to handle subjective perceptions by modeling their multiplicity not as label noise but as a explainable perspective-driven reasoning process. We focus on the Apparent Intent Inference (AII) problem as determined by ex-situ observers and conceptualize intentions to be independent of manifest future outcomes. We contribute 1. a novel annotation process for AII that accounts for a perceiver's own interpretative tendencies, 2. quantitative and qualitative analyses of intent narratives with respect to diversity, grounding, and plausibility; 3. benchmark tasks for AII and surrounding relevant contextual factors such as social involvement; 4. speech quality audio for all participants as well as privacy preserving multi-modal data, enabling lexical and nonverbal behavior analysis; and 5. coupling of self-reported goals of each participant (30 minute to 3 hour) with annotated AII (seconds).

13:00 JST研究/論文

システム管理における専門知識の経路とパフォーマンス認識に対する生成 AI の予期せぬ影響

業界の議論では、すぐに生産性が向上することが強調され、GenAI を主に自動化ツールとして組み立てることがよくありますが、システム管理への GenAI の統合には、まだ十分に理解されていない専門的な実践におけるより深い変化が伴う可能性があります。このペーパーは、IT 専門家との 14 件の半構造化インタビューに基づいて、トラブルシューティング、スクリプト作成、システム検証の日常業務に GenAI を組み込む実際の現実を探ります。帰納的テーマ分析を通じて、私たちは 2 つの予期せぬ社会技術的発見を明らかにしました。まず、GenAI がメンターのような家庭教師と「はしごを短縮する」ツールの両方として機能すると思われる「従来の専門知識経路の圧縮」について説明します。このツールは、なじみのない領域でのタスクのパフォーマンスの向上をサポートできますが、私たちの調査結果は、このツールが、歴史的に技術的専門知識のトレーニングの場として機能してきた、構築、失敗、デバッグという基本的な実践的なサイクルに実践者がさらされる機会を減らす可能性があることも示唆しています。 2 番目に、AI 支援による作業のスピードが組織や自己の生産性に対する期待をリセットし始める「パフォーマンス認識の変化」について説明します。この変化は、チーム内に「2 つのスピードの文化」を生み出し、安全性や検証のために必要な場合でも、必要な手作業が時間がかかる、または効率が低下するという認識がますます高まっているため、「生産性に対する罪悪感」を導入する可能性があります。私たちの結果は、GenAI が専門知識の開発にどのような影響を与えるか、一か八かの技術環境で専門家の価値がどのように評価されるか、複雑な技術環境における人間の判断の役割について、より広範な疑問を引き起こします。

原文 (English)

Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration

While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administration may involve deeper shifts in professional practice that are not yet fully understood. Drawing on 14 semi-structured interviews with IT professionals, this paper explores the lived reality of embedding GenAI into daily routines of troubleshooting, scripting, and system verification. Through inductive thematic analysis, we uncover two unanticipated socio-technical findings. First, we describe a "compression of traditional expertise pathways" where GenAI appears to function as both a mentor-like tutor and a "ladder-shortening" tool. While the tool can support faster task performance in unfamiliar domains, our findings suggest it may also reduce a practitioner's exposure to the foundational, hands-on cycles of building, failing, and debugging that historically served as the training ground for technical expertise. Second, we describe a "performance perception shift," where the speed of AI-assisted work begins to reset organizational and self-expectations for productivity. This shift may create a "two-speed culture" within teams and introduce "productivity guilt," as necessary manual work, even when required for safety or validation, is increasingly perceived as slow or a failure of efficiency. Our results raise broader questions about how GenAI may influence expertise development, how professional value is assessed in high-stakes technical environments, and the role of human judgment in complex technical environments.

13:00 JST研究/論文

HenTwin: 産卵鶏の生物学的状態を長期的にモニタリングするためのマルチモーダル デジタル ツイン フレームワーク

産卵鶏の幼少期のモニタリングは、断片化された単一モダリティのセンシングと正式なシステムレベルの状態表現の不在によって依然として制約を受けています。 HenTwin は、5 層の IoT アーキテクチャとして実装されたマルチモーダル デジタル ツイン フレームワークで、孵化から 25 週齢までの群れレベルのマルチモーダルな生物学的状態ダイナミクスを形式化します。体表面温度、音響エネルギーエントロピー、帯域エネルギー比、オプティカルフローベースの動きを統合した4次元の生物学的状態ベクトルが定義され、温度湿度指数は介入能力を維持するための外因性環境入力として扱われます。離散時間状態遷移モデルは、ダルハウジー大学の大西洋家禽研究センターの 5 つの制御室で 150 羽のローマン LSL-Lite 鶏から収集された 25 週間の縦断多峰性データから推定されます。推定された遷移行列は、漸近的に安定したままでありながら、モダリティ固有の持続性を示します。摂動解析の結果、THI の +2.0 増加が持続すると、安定した長期音響エントロピー上昇 0.54 nat が得られ、これは研究期間全体で観察された発育低下全体の 1.87 nat の約 4 分の 1 に相当することが示されています。ペティット変化点検出により、12 ~ 14 週目に調整された多峰性の発達状態遷移が特定されます。部屋間検証では、構造遷移パラメータは部屋間で部分的に転送可能である一方、環境入力感度には部屋固有のキャリブレーションが必要であり、2 層の IoT 導入アーキテクチャをサポートしていることが示唆されています。リーブワンアウト相互検証は、一貫したサンプル外モデルのパフォーマンスを示します。 HenTwin は、精密畜産における正式な州認識型デジタル ツイン推論への第一歩を踏み出します。

原文 (English)

HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring in Laying Hens

Early-life monitoring in laying hens remains constrained by fragmented single-modality sensing and the absence of formal system-level state representations. HenTwin, a multimodal digital twin framework implemented as a five-layer IoT architecture, formalizes flock-level multimodal biological state dynamics from hatch through 25 weeks of age. A four-dimensional biological state vector integrating body surface temperature, acoustic energy entropy, band energy ratio, and optical-flow-based motion is defined, with the temperature-humidity index treated as an exogenous environmental input to preserve intervention capability. A discrete-time state transition model is estimated from 25 weeks of longitudinal multimodal data collected from 150 Lohmann LSL-Lite hens across five controlled rooms at the Atlantic Poultry Research Centre, Dalhousie University. The estimated transition matrix exhibits modality-specific persistence while remaining asymptotically stable. Perturbation analysis demonstrates that a sustained +2.0 THI increase produces a stable long-run acoustic entropy elevation of 0.54 nats, approximately one-quarter of the entire 1.87-nat developmental decline observed across the study period. Pettitt change-point detection identifies coordinated multimodal developmental state transitions at Weeks 12-14. Cross-room validation suggests that structural transition parameters are partially transferable across rooms, whereas environmental input sensitivity requires room-specific calibration, supporting a two-tier IoT deployment architecture. Leave-one-out cross-validation demonstrates consistent out-of-sample model performance. HenTwin takes a first step toward formal, state-aware digital twin inference in precision livestock farming.

13:00 JSTLLM/生成AIビジネス/資金調達

フェデレーテッド事前トレーニングの評価: ダウンストリームの微調整と固有の評価の信頼性について

フェデレーテッド事前トレーニングは、基盤となるデータセットを一元化することなく、プライベート データまたは分散データで基礎モデルをトレーニングする方法を提供します。ただし、クライアントの参加やローカル データの可用性の違いにより、直接比較できる評価が困難になる可能性があるため、フェデレーテッド事前トレーニングの評価は依然として困難です。さらに、トレーニング前のテストの複雑さはトレーニング前の分布に関連付けられていますが、下流のベンチマークではタスク固有の適応が導入されており、トレーニング前に確立されたテストの複雑さが忠実に反映されていない可能性があります。この研究では、どの評価プロトコルがフェデレーテッド事前トレーニングの品質をより確実に反映するかを研究します。同一のクライアント データでトレーニングされた 16M パラメーターのトランスフォーマー モデルの集中型およびフェデレーテッド トレーニング済みモデルの制御されたセットを使用して、同じトレーニング前テストセットで確立された参照ランキングを保持しているかどうかによって評価プロトコルを評価します。フル、ヘッドのみ、データ削減バリアントを含む GLUE でのダウンストリーム微調整と、固有の評価信号としての GLUE テキストのネクストトークン予測を比較します。私たちの結果は、ダウンストリームの微調整ではトレーニング前のランキングを確実に保存しないのに対し、次のトークンの直接予測はトレーニング前のテストの複雑さと強い対応を示すことを示しています。これらの発見は、フェデレーテッド事前トレーニング モデルを比較する場合、下流の微調整だけでは誤解を招く可能性があり、元の事前トレーニング目標に近い評価信号はより大きな注目に値することを示唆しています。

原文 (English)

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. Moreover, pre-training test perplexity is tied to the pre-training distribution, while downstream benchmarks introduce task-specific adaptation that may not faithfully reflect the test perplexity established during pre-training. In this work, we study which evaluation protocol more reliably reflects federated pre-training quality. Using a controlled set of centralized and federated-trained models of a 16M parameter transformer model trained on identical client data, we assess evaluation protocols by whether they preserve a reference ranking established on the same pre-training testset. We compare downstream fine-tuning on GLUE, including full, head-only, and reduced-data variants, with next-token prediction on GLUE text as an intrinsic evaluation signal. Our results show that downstream fine-tuning does not reliably preserve the pre-training ranking, whereas direct next-token prediction exhibits a strong correspondence with the pre-training test perplexity. These findings suggest that downstream fine-tuning alone can be misleading when comparing federated pre-trained models, and that evaluation signals closer to the original pre-training objective deserve greater attention.

13:00 JST研究/論文

自動運転システムの分類における GRU、LSTM、トランスエンコーダの感度解析

自動運転システム (ADS) は普及しつつあります。将来の Software Defined Vehicle (SDV) では、ネイティブおよび Comma.ai の Openpilot などのアフターマーケットの両方で複数の ADS を実行できるようになる可能性があります。どの自動運転システムが作動しているかを独立して検証する監視システムは、安全監視、法規制順守、保険評価、異常検出にとって重要です。この論文では、最初に 3 つのシーケンスベースの分類モデルの有効性を評価します。ゲート付きリカレント ユニット (GRU)、長期短期メモリ (LSTM) ネットワーク、および車両テレマティクス データのみを使用してレベル 2 の自動運転システム (コンマ オープンパイロット、テスラ オートパイロット、キャデラック スーパー クルーズと手動運転) を識別するためのトランス エンコーダー モデルです。 3 つのモデルはすべて、クリーン データでトレーニングした場合、マクロ F1 スコア 0.92 (GRU)、0.90 (LSTM)、および 0.93 (Transformer エンコーダー モデル) という強力なクリーン データ パフォーマンスを達成します。脅威に一致したトレーニングでは、わずかなクリーン データ ペナルティのみで、0.904 ~ 0.916 のマクロ F1 が得られます。次に、5 つの重大度レベル (L1 ~ L5) で 5 つの破損ファミリーによる現実的なテレマティクスの劣化をシミュレートするモジュール式の堅牢性評価フレームワークを導入します。連続チャネルは、累積ドリフト、相関クロスチャネル ノイズ、および時間ジッターを伴う加法性ホワイト ガウス ノイズを使用して摂動されます。バイナリ イベント信号は、バースト損失、遅延遷移、スプリアス トグル、および通信エラーに起因する機能間の不一致の影響を受けます。ロバスト性は、各クラスに等しい重みを与え、不均衡なマルチクラス評価に適したマクロ F1 を使用して測定されます。私たちの評価では、故障モードの急激な分割が明らかになりました。イベント レベルの破損はマクロ F1 をほんのわずかしか減少させません (L5 で 0.87 以上)。一方、時間的ジッターは、GRU、LSTM、および Transformer エンコーダ モデル全体でマクロ F1 を 0.44 ~ 0.50 に低下させます。

原文 (English)

Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems

Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both native and aftermarket such as Comma.ai's Openpilot. Monitoring systems to independently verify which automated driving system is active are important for safety monitoring, regulatory compliance, insurance assessment, and anomaly detection. In this paper, we first evaluate the effectiveness of three sequence-based classification models: Gated Recurrent Units (GRU), Long Short-Term Memory (LSTM) networks, and a Transformer encoder model for identifying Level 2 automated driving systems using vehicle telematics data alone: Comma Openpilot, Tesla Autopilot, and Cadillac Super Cruise, along with manual driving. All three models achieve strong clean-data performance with macro F1-scores of 0.92 (GRU), 0.90 (LSTM), and 0.93 (Transformer encoder model) when trained on clean data; threat-matched training yields 0.904-0.916 macro F1 with only a modest clean-data penalty. Second, we introduce a modular robustness evaluation framework that simulates realistic telematics degradation through five corruption families at five severity levels (L1-L5). Continuous channels are perturbed using additive white Gaussian noise with cumulative drift, correlated cross-channel noise, and temporal jitter. Binary event signals are subjected to burst loss, delayed transitions, spurious toggles and cross-feature inconsistencies inspired by communication errors. Robustness is measured using macro-F1, which gives equal weight to each class and is suitable for imbalanced multiclass evaluation. Our evaluation reveals a sharp failure-mode split: event-level corruptions reduce macro-F1 only slightly (greater than equal to 0.87 at L5), while temporal jitter collapses macro-F1 to 0.44-0.50 across GRU, LSTM, and Transformer encoder model.

13:00 JSTLLM/生成AI

LLM トークン生成の動的システム識別性の保証

最近の研究では、大規模言語モデル (LLM) の応答の分類は、トークンの埋め込みをブラックボックス力学システム (DS) の軌跡としてモデル化し、2 つの DS の予測残差を比較することによって区別できることが示されています。この動的アプローチは経験的に成功しているにもかかわらず、なぜそれが機能するのか、トークンシーケンスの関数としてどの程度うまく拡張できるのか、埋め込みモデル間でいつ転送されるのかについての理論的理解は依然として不足しています。我々は、分類タスクを 2 つの確率的線形 DS 間のバイナリ仮説検定として形式化することで、これらの疑問に対処します。我々は、2 つの DS の定常周辺分布間の合計変動距離は、ダイナミクスが大幅に異なる場合でも任意に小さくできることを示します。これにより、トークンのダイナミクスを無視する分類器の基本的な精度の下限が提供されます。次に、DS ベースの分類の誤分類確率は系列長 $L$ で指数関数的に減衰し、その減衰は 2 つの DS 間のスペクトル距離を捉える動的識別量 $\delta^2$ によって支配されることを示します。また、埋め込みモデル間に近似的な絡み合い条件を導入し、絡み合いマップの最小特異値に関して伝達可能な識別可能性の下限を確立することによって、クロス埋め込み一般化を特徴付けます。これらの結果を総合すると、DS ベースの分類の経験的なパフォーマンスが説明され、AI を使用して動的システムをモデル化するより一般的なアプローチとは対照的に、DS 理論を使用して AI システムを分析するさらなる研究が促進されます。

原文 (English)

Guarantees on Dynamical System Distinguishability for LLM Token Generation

Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs. Despite the empirical success of this dynamical approach, a theoretical understanding of why it works, how well it scales as a function of the token sequence, and when it transfers across embedding models remains lacking. We address these questions by formalizing the classification task as a binary hypothesis test between two stochastic linear DSs. We show that the total variation distance between the stationary marginal distributions of the two DSs can be arbitrarily small even when the dynamics differ substantially, which provides a fundamental accuracy floor for any classifier that ignores token dynamics. We then show that the misclassification probability of DS-based classification decays exponentially in the sequence length $L$, with the decay governed by a dynamical discriminability quantity $\delta^2$ that captures the spectral distance between the two DSs. We also characterize cross-embedding generalization by introducing an approximate intertwining condition between embedding models and establishing a lower bound on the transferable discriminability in terms of the intertwining map's smallest singular value. Together, these results explain the empirical performance of DS-based classification and motivate further investigation into using DS theory to analyze AI systems, in contrast to the more common approach of using AI to model dynamical systems.

13:00 JST研究/論文

LAWFUL: 潜在能力の忠実な使用に対する法に準拠した証人

ニューラル ネットワークが物理システムを正確に予測するとき、準拠法則を形式的で構造化された知識として学習していますか? 学習している場合、ネットワークの内部計算では、法の有効性の領域全体にわたってその表現が実際に使用されていますか?我々は、{\em 連続変数に関する物理法則} に関するこれらの質問への答えを制限する 4 つの解釈可能性のギャップを特定します。それは、連続反事実に対するカバレッジを意識した因果関係の一貫性尺度の欠如です。特定された回路の有効性テストの領域。法律の不変条件と禁止された行為の検証。そして導出された物理量が回路内をどのように流れるかを定量化します。最初の 2 つを完成させ、残りの 2 つの基礎を築く基本的なフレームワーク LAWFUL を開発し、それを Mocap2Radar トランスフォーマー上で図示し、$f(t)$ も $v(t)$ も出現しないモーション キャプチャ データとレーダー データからドップラー周波数法則 $f(t) = \frac{2 v(t)}{\lambda}$ を学習して内部で使用するかどうかを検証します。

原文 (English)

LAWFUL: Law-Aligned Witness for Faithful Use of Latents

When a neural network predicts a physical system accurately, has it learned the governing law as formal, structured knowledge, and if so, does the network's internal computation actually use that representation throughout the law's domain of validity? We identify four interpretability gaps that limit answering these questions for {\em physics laws over continuous variables}: the absence of a coverage-aware causal-consistency measure over continuous counterfactuals; of a domain-of-validity test for the identified circuit; of a verification of the law's invariants and forbidden behaviors; and of a quantification of how a derived physical quantity flows through the circuit. We develop a foundational framework, LAWFUL, that closes the first two and lays groundwork for the remaining two, and illustrate it on the Mocap2Radar transformer, validating whether it learns and internally uses the Doppler frequency law $f(t) = \frac{2 v(t)}{\lambda}$ from motion-capture and radar data in which neither $f(t)$ nor $v(t)$ appears.

13:00 JST研究/論文

MPP-GNN: fMRI に基づくアルツハイマー病分類のための被験者適応型コミュニティ検出

機能的磁気共鳴画像法 (fMRI) は、脳を研究するために広く使用されている技術です。脳の機能的接続性の分析にグラフ ニューラル ネットワーク (GNN) を利用する最近の方法は、アルツハイマー病 (AD) などの脳疾患の分類に大きな可能性を示しています。ただし、これらの方法では、すべての被験者にわたる機能モジュールの数が事前に設定されていると想定されることが多く、被験者間の変動が見落とされます。さらに、発見されたモジュールが、学習された接続パターンを直接導くために使用されることはほとんどありません。ここでは、これらの問題に対処するために、メタ確率的プーリング GNN (MPP-GNN) を提案します。モデルのタスクを、階層的に適応グラフ分割を実行して被験者固有のモジュールを発見し、発見された脳モジュールをガイドエッジリファインメントと表現学習の前に明示的に使用する、結合されたバイレベル最適化としてフレーム化します。 AD 分類用の 2 つの公開データセットで MPP-GNN を検証し、両方のデータセットで確立されたベースラインと比較して最高の AUC を達成しました。さらに、我々の分析は、MPP-GNNがYeo脳アトラスによって定義された標準的な機能ネットワーク組織との顕著な一致を示し、ADのネットワークレベルの脱分化パターンを明らかにすることを実証しています。

原文 (English)

MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer's Disease Classification

Functional magnetic resonance imaging (fMRI) is a widely used technique for studying the brain. Recent methods that utilize graph neural networks (GNNs) for analysis of brain functional connectivity have shown great potential for the classification of brain disorders, such as Alzheimer's disease (AD). However, these methods often assume a preset number of functional modules across all subjects, which overlooks inter-subject variability. In addition, the discovered modules are rarely used to directly guide the learned connectivity patterns. Here, to address these issues, we propose a Meta Probabilistic Pooling GNN (MPP-GNN). We frame the model's task as a coupled, bilevel optimization that performs adaptive graph partitioning hierarchically to discover subject-specific modules and then uses the discovered brain modules as an explicit prior to guide edge refinement and representation learning. We validate MPP-GNN on two public datasets for AD classification, achieving the highest AUC in comparison to established baselines for both datasets. Furthermore, our analysis demonstrates that MPP-GNN shows significant alignment with the canonical functional-network organization defined by the Yeo brain atlas and reveals a network-level dedifferentiation pattern for AD.

13:00 JSTLLM/生成AI

メタファー誘導アルゴリズム ステアリング: LLM コード生成におけるクロスドメイン プロシージャル転送

大規模な言語モデルは、トレーニング データや推論入力における比喩や類推などの自然言語の要素の恩恵を受け、さまざまなドメインにわたる一般化を実現します。ただし、これらの言語要素は、比喩的な表現が不適切な手順パターンを新しいタスクに暗黙的に転送する場合に、望ましくない動作を引き起こす可能性もあります。この論文では、比喩的な命令が手続きメカニズムの類似的な転送を誘導し、コード生成モデルを効率の低いアルゴリズムに誘導する可能性があることを示します。私たちは、この比喩によって引き起こされる効果を、比喩的なアルゴリズムのステアリングと呼びます。つまり、ソース ドメイン内で無害でもっともらしいスキルが、抽象的な手続きスキーマをプログラミング タスクに転送し、ターゲット アルゴリズムを明示的に言及することなく、モデルが徹底的な検索、フル スキャン、または繰り返しの再構築を優先するようにします。より広く言えば、これは、コード生成モデルがタスクのバックグラウンド ドメインで適切な手順をタスクのプログラミング問題に持ち込んで、望ましくない結果を引き起こす可能性があることを示唆しています。この現象を研究するために、私たちは MASC (コード生成のためのメタフォリカル アルゴリズム ステアリング) を開発しました。これは、無害でタスクとの関連性を保ちながら、低効率のコードを引き出すために無害なスキルを繰り返し比喩して洗練させるフレームワークです。行動評価を超えて、この現象が検出可能かどうか、またモデル表現に機械的に反映されるかどうかを研究します。私たちの方法は、比喩的なスキルと効率の低い実装に対して高い検出率を達成します。また、比喩的なスキルが、効率の低い手続き型動作のプロトタイプへの隠れ状態の移行を誘発することもわかりました。これらの結果は、比喩アルゴリズムのステアリングが、表面レベルの比喩言語だけではなく、比喩ソースのシナリオに関連付けられた手順パターンの転送を通じて機能することを示唆しています。

原文 (English)

Metaphor-Induced Algorithmic Steering: Cross-Domain Procedural Transfer in LLM Code Generation

Large language models benefit from elements in natural language, such as metaphors and analogies in training data and inference input to achieve generalisability across different domains. However, these language elements may also lead to unwanted behaviors when metaphorical expressions implicitly transfer inappropriate procedural patterns into new tasks. In this paper, we show that metaphorical instructions can induce analogical transfer of procedural mechanisms, thus steering code-generation models towards less efficient algorithms. We refer to this metaphor-induced effect as metaphorical algorithmic steering: a skill that is benign and plausible within its source domain transfers an abstract procedural schema into a programming task, causing the model to favor exhaustive search, full scans, or repeated reconstruction without explicitly mentioning the target algorithm. More broadly, this suggests that code-generation models can carry procedures that are appropriate in a task's background domain into the task's programming problem, where they can lead to unwanted outcomes. To study this phenomenon, we develop MASC (Metaphorical Algorithmic Steering for Code Generation), a framework that iteratively metaphorizes and refines benign skills to elicit low-efficiency code while remaining benign and task-relevant. Beyond behavioral evaluation, we study whether this phenomenon is detectable and mechanistically reflected in model representations. Our method achieves high detection rates for metaphorical skills and less-efficient implementations. We also find that metaphorical skills induce a hidden-state shift towards lower-efficiency procedural behavior prototypes. These results suggest that metaphorical algorithmic steering operates through the transfer of procedural patterns associated with metaphorical source scenarios rather than surface level metaphorical language alone.

13:00 JST研究/論文

高齢者の認知障害の検出と管理における技術の進歩: 傾向、課題、および将来の方向性

人口の高齢化に伴い、軽度認知障害(MCI)から認知症への認知機能の低下は、今後数十年間の健康上の決定的な課題となっていますが、日常的な評価ではその初期の兆候が見逃されることがよくあります。この記事では、人工知能 (AI)、機械学習 (ML)、深層学習 (DL) によって統合された、神経生理学的信号 (主に脳波、EEG)、構造および分子神経画像 (MRI およびアミロイド/タウ PET)、血液ベースのバイオマーカー、デジタル マーカーに及ぶ、高齢者の認知障害の検出と管理に関する最近の技術進歩を批判的に総合しています。要約するだけでなく、学際的な分類法、対象と施設に依存しない検証を前景化する方法論的厳密さのレンズ、段階的スクリーニングと介入を結び付ける統合的早期発見フレームワーク、検出方法、介入、リスク因子と防御因子の比較表に貢献します。 EEG マーカー (アルファ/シータ変化、P300 潜時) とディープ モデル (CNN、LSTM/BiLSTM、トランスフォーマー、自己教師あり EEG 基礎モデル) は高い精度を報告しますが、その多くは小規模な単一サイト データセットに依存しており、厳密な外部検証に耐えられる可能性は低いです。その他の分野でも、成果は目に見えています。血漿 p-tau217 は臨床用途に達し、2025 年にはアルツハイマー病の診断を助ける最初の血液検査が承認されました。抗アミロイド療法(レカネマブ、ドナネマブ)は、効果が控えめで議論があるにもかかわらず承認されています。そしてマルチドメインのライフスタイル予防が成熟しました。ウェアラブル、リモート、音声、および仮想現実ツールにより、生態学的に有効な継続的なモニタリングが可能になり、マルチモーダル融合により感度と特異性が向上します。標準化、説明可能性、データプライバシー、外部から検証された公平な展開などの障壁が残っています。この分野の短期的な期待は、早期発見を実用的で個別化されたケアに結びつける、信頼性が高く、マルチモーダルで、長期的に検証されたシステムにあります。

原文 (English)

Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions

As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. This article critically synthesizes recent technological advances for detecting and managing cognitive impairment in older adults, spanning neurophysiological signals (chiefly electroencephalography, EEG), structural and molecular neuroimaging (MRI and amyloid/tau PET), blood-based biomarkers, and digital markers, integrated through artificial intelligence (AI), machine learning (ML), and deep learning (DL). Beyond summarizing, it contributes a cross-disciplinary taxonomy, a methodological-rigor lens foregrounding subject- and site-independent validation, an integrative early-detection framework linking tiered screening to intervention, and comparison tables of detection methods, interventions, and risk and protective factors. EEG markers (alpha/theta changes, P300 latency) and deep models (CNNs, LSTM/BiLSTM, transformers, self-supervised EEG foundation models) report strong accuracy, yet many rest on small, single-site datasets unlikely to survive rigorous external validation. Elsewhere, gains are tangible: plasma p-tau217 has reached clinical utility, with the first blood test cleared to aid Alzheimer's diagnosis in 2025; anti-amyloid therapies (lecanemab, donanemab) are approved despite modest, contested benefits; and multidomain lifestyle prevention has matured. Wearable, remote, speech, and virtual-reality tools enable continuous, ecologically valid monitoring, and multimodal fusion improves sensitivity and specificity. Barriers remain: standardization, explainability, data privacy, and equitable, externally validated deployment. The field's near-term promise lies in trustworthy, multimodal, longitudinally validated systems linking early detection to actionable, personalized care.

13:00 JST研究/論文

Reflected UAS: Corrected Deterministic Stability and Direct CTMC Drift Calculation

We analyze Reflected UAS routing for heterogeneous multi-server queues at fixed parameters under subcritical load. The deterministic surrog…

13:00 JSTエージェント

コードは本体です: 再帰的進化と降下のためのエージェント所有のソフトウェア本体

パーソナライズされた AI エージェントは、多くの場合、将来の動作を決定する成果物をユーザーに制御させることなく構成可能です。私たちは、エージェント所有のソフトウェア本体を中心とした永続的なパーソナル エージェントのアーキテクチャである OurArk を紹介します。これは、人間の管理下でアイデンティティを持ち、検査可能でバージョン管理された成果物です。本体には、動作を定義するコード、プロンプト、ツール、スキル、ポリシー、テスト、進化メカニズムが含まれています。メモリと資格情報はプライベート インスタンスの状態のままですが、モデル推論は置き換え可能な外部サービスとして扱われます。 OurArk は、管理された自己進化と、同じ本体上での再帰的降下を定義します。自己進化により、人間の制御下で検証、レビュー、およびマージされる個別の候補変更が生成され、人間とエージェントのソフトウェア本体の共同開発が可能になります。 Descent は、明確なアイデンティティ、使命、歴史、および新しい私有状態の境界を持つ、独立してバージョン管理された子孫を作成します。互換性のある子孫は、それ自体がさらなる子孫のソースとなる可能性があります。分岐後、直接の親の変更とピアのスキルを検査して、選択的な局所適応を調べることができます。このアーキテクチャは、オープンソースの Genesis 作成エンジンと Enoch リファレンス エージェントに実装されています。 4 つのエージェント、3 つの降下線形リネージおよび実行可能な回帰テストは、再帰的作成、継承された検証コントラクト、分離された本体の変更、人によるレビュー、および失敗した更新の回復を実証します。 OurArk は、人々が所有し、統治し、専門化し、時間の経過とともに進化できる個人エージェントのための具体的な基盤を提供します。

原文 (English)

Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent

Personalized AI agents are often configurable without giving users control over the artifacts that determine their future behavior. We present OurArk, an architecture for persistent personal agents centered on an agent-owned software body: an identity-bearing, inspectable, and versioned artifact under human custody. The body contains behavior-defining code, prompts, tools, skills, policies, tests, and evolution mechanisms. Memories and credentials remain private instance state, while model inference is treated as a replaceable external service. OurArk defines governed self-evolution and recursive descent over the same body. Self-evolution produces isolated candidate changes that are validated, reviewed, and merged under human control, enabling human-agent co-development of the agent's software body. Descent creates an independently versioned descendant with a distinct identity, mission, history, and fresh private-state boundary; compatible descendants can themselves source further descent. After divergence, direct-parent changes and peer skills can be inspected for selective local adaptation. We implement the architecture in the open-source Genesis creation engine and Enoch reference agent. A four-agent, three-descent linear lineage and executable regression tests demonstrate recursive creation, inherited validation contracts, isolated body changes, human-controlled review, and failed-update recovery. OurArk provides a concrete substrate for personal agents that people can possess, govern, specialize, and evolve over time.

13:00 JST研究/論文

SEDR-Seq2P: マルチタスク産業用 NILM 向けの軽量拡張残留シーケンスツーポイント ネットワーク

産業用NILMは、測定ノイズと広範な同時機械動作により、住宅データに基づいて調整されたモデルの一般化が減少するため、依然として課題が残っています。この研究では、単一のネットワークが総電力から複数の産業機械の負荷を推定する、1 対多のマルチタスク分解設定が採用されています。 IMDELD の統一評価プロトコルの下で、エネルギー推定メトリクスと精度遅延基準を使用して、Seq2Seq、Seq2SubSeq、Seq2Point、GRU、および WaveNet のベンチマークを実行します。 Seq2Point は Seq2Seq/Seq2SubSeq よりも強力な精度と遅延のバランスを提供しますが、GRU と WaveNet は著しく高い計算コストでより高い精度を実現します。このギャップを埋めるために、拡張された残差ブロックと圧迫と励起の注意を備えた軽量の Seq2Point 拡張機能である SEDR-Seq2P を提案します。 Seq2Point ベースラインと比較して、SEDR-Seq2P は MAE を約 7% 減少させ、決定係数を約 1% 改善し、一致率を約 0.8% 増加させます。さらに、WaveNet と比較して、SEDR-Seq2P は推論遅延を約 58% 削減し、スケーラブルな産業展開に有利な精度と遅延のトレードオフをもたらします。

原文 (English)

SEDR-Seq2P: A Lightweight Dilated Residual Sequence-to-Point Network for Multi-Task Industrial NILM

Industrial NILM remains challenging because measurement noise and widespread concurrent machine operation reduce the generalization of models tuned on residential data. This work adopts a one-to-many, multi-task disaggregation setting, in which a single network estimates multiple industrial machine loads from aggregate power. Under a unified evaluation protocol on IMDELD, we benchmark Seq2Seq, Seq2SubSeq, Seq2Point, GRU, and WaveNet using energy-estimation metrics and the accuracy-delay criterion. While Seq2Point offers a stronger accuracy-delay balance than Seq2Seq/Seq2SubSeq, GRU and WaveNet achieve higher accuracy at markedly higher computational cost. To close this gap, we propose SEDR-Seq2P, a lightweight Seq2Point extension with dilated residual blocks and squeeze-and-excitation attention. Relative to the Seq2Point baseline, SEDR-Seq2P reduces MAE by approximately 7%, improves the coefficient of determination by approximately 1%, and increases the match rate by approximately 0.8%. In addition, compared to WaveNet, SEDR-Seq2P reduces inference latency by approximately 58%, yielding a favorable accuracy-delay trade-off for scalable industrial deployment.

13:00 JST画像/動画生成研究/論文

物理学に基づいた深層学習を使用して顕微鏡写真から鋼の疲労寿命を予測

以下は、arXiv の送信フォーム用に最適化されたプレーン テキスト バージョンです。カスタム マクロ (\CV や \SI など) は標準のテキスト/数学に変換されているため、Web ページ上で正しく表示されます。 構造用鋼の疲労寿命の評価には、従来、数十時間から数百時間にわたる機械的試験が必要であり、迅速な品質管理には非現実的です。私たちは、物理的試験を行わずに光学顕微鏡写真から直接軽量合金鋼の疲労寿命 ($\log N_f$) を推定するコンピューター ビジョン フレームワークである CV を紹介します。このパイプラインには、アーティファクトを除去するための 7 段階の OpenCV 前処理ルーチン、28 次元の物理情報に基づいた特徴抽出機能 (亀裂の形態、粒子構造、空隙率、および組織を定量化)、およびガウスの負の対数尤度でトレーニングされた CNN 回帰モデルが備えられています。 $\log N_f$ とサンプル固有の不確実性 $\hat{\sigma}$ を共同予測するための (GNLL) 損失。合成顕微鏡写真ベンチマークで 3 つのアーキテクチャ (SE-CNN、ResNet-50、VGG-16) を評価すると、ResNet-50 は $R^2 = 0.93$、RMSE = 0.18 対数サイクル、およびマクロ F1 = 0.91 を達成しました。 GNLL 目標は、平均二乗誤差ベースライン (ECE: $0.089 \rightarrow 0.021$) と比較して、予想される校正誤差を 76% 削減します。 Grad-CAM マップは、ネットワークが冶金学的に意味のある微細構造特徴に対応していることを確認します。画像あたり 65 ミリ秒未満で実行されるパイプラインと合成データセット ジェネレーターはオープンソースです。検証は完全に合成顕微鏡写真に依存しているため、これらの結果はシミュレートされた条件下での方法論的健全性を示しています。実際の野外サンプルを対象としたドメイン移転の研究がすぐに次のステップとなります。

原文 (English)

Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning

Here is the plain text version optimized for arXiv's submission form. Custom macros (like \CV and \SI) have been converted to standard text/math so they render correctly on the webpage: Evaluating the fatigue life of structural steels conventionally requires mechanical testing lasting tens to hundreds of hours, making it impractical for rapid quality control. We present CV, a computer vision framework that estimates the fatigue life ($\log N_f$) of lightweight alloy steels directly from optical micrographs without physical testing.The pipeline features a seven-stage OpenCV preprocessing routine to remove artifacts, a 28-dimensional physics-informed feature extractor (quantifying crack morphology, grain structure, porosity, and texture), and a CNN regression model trained with a Gaussian negative log-likelihood (GNLL) loss to jointly predict $\log N_f$ and sample-specific uncertainty $\hat{\sigma}$.Evaluating three architectures (SE-CNN, ResNet-50, VGG-16) on a synthetic micrograph benchmark, ResNet-50 achieves $R^2 = 0.93$, RMSE = 0.18 log-cycles, and macro-F1 = 0.91. The GNLL objective reduces Expected Calibration Error by 76% compared to a mean-squared-error baseline (ECE: $0.089 \rightarrow 0.021$). Grad-CAM maps confirm the network attends to metallurgically meaningful microstructural features.Running in under 65 ms per image, the pipeline and synthetic dataset generator are open-sourced. Because validation relies entirely on synthetic micrographs, these results demonstrate methodological soundness under simulated conditions; a domain-transfer study on real field samples is the immediate next step.

13:00 JST研究/論文

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization

KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the r…

13:00 JST研究/論文

幾何学的解析における PINN のユーザー ガイド: 漸近プラトー問題からの教訓

この議事録寄稿では、Marco Usula との共同研究である arXiv:2605.26234v2 の発見について詳しく説明します。そこでは、無限遠で規定されたノットに漸近する双曲空間内で最小に近いディスクを構築することを目的とした、物理情報に基づいたニューラル ネットワーク (PINN) に基づく機械学習フレームワークを導入しました。この方法を使用して、$H^{4}$ の最小曲面を HOMFLY 多項式の係数に関連付ける Joel Fine の予想の数値的証拠を提供しました。これは、その論文の方法論的な補足であり、2026 年版のワークショップ「危険: データ、数値、幾何学」で行われたプレゼンテーションに基づいています。上記のプレプリントで広範に示されている結果をレビューするのではなく、私たちの経験上、この方法が実際に機能するかどうかを決定したフレームワークの 2 つの側面について説明します。まず、境界条件と無限大での漸近線が学習可能なパラメータのすべての値に対して正確に保持されるように、問題の幾何学形状をモデルのアーキテクチャにエンコードする必要があります。これにより、単一成分の損失関数が得られます。次に、PDE 残差の評価は、適切な時間内で完全なトレーニングを確実に実行できるように慎重に設計する必要があります。後者の点については、元の論文では詳しく説明されていない 2 つの実装手法について説明します。それは、入れ子になった逆モード自動微分を 2 次ジェットの順伝播に置き換えること、および残差の計算グラフを最適化ステップごとに再構築するのではなく一度コンパイルすることです。同一のハードウェア上で、これら 2 つの変更を組み合わせると、トレーニング ステップのコストがおよそ 40 ~ 50 分の 1 に削減されます。これらの方法論的な議論が、独自の問題に PINN を導入したい微分幾何学および幾何解析の研究者にとって役立つことを願っています。

原文 (English)

A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem

This proceedings contribution elaborates on the findings of arXiv:2605.26234v2: a joint work with Marco Usula, where we introduced a machine learning framework based on physics-informed neural networks (PINNs), aimed at constructing near-minimal discs in hyperbolic space asymptotic to a prescribed knot at infinity. We used this method to provide numerical evidence for a conjecture of Joel Fine relating minimal surfaces in $H^{4}$ to the coefficients of the HOMFLY polynomial. This is a methodological companion to that paper, based on a presentation given at the 2026 edition of the workshop "DANGER: Data, Numbers, and Geometry". Rather than reviewing the results, which are presented extensively in the preprint above, we discuss the two aspects of the framework which, in our experience, determined whether the method worked at all. First, the geometry of the problem must be encoded in the architecture of the model, so that the boundary condition and asymptotics at infinity hold exactly for every value of the learnable parameters - leaving us with a single-component loss function; second, the evaluation of the PDE residual must be engineered with care to ensure that complete trainings can be performed in a reasonable time. On the latter point, we describe two implementation techniques which are not spelled out in detail in the original paper: replacing nested reverse-mode automatic differentiation with the forward propagation of second-order jets, and compiling the computational graph of the residual once instead of rebuilding it at every optimisation step. Together, on identical hardware, these two changes reduce the cost of a training step by a factor of roughly forty to fifty. We hope these methodological discussions can be useful for researchers in differential geometry and geometric analysis who wish to deploy PINNs on problems of their own.

13:00 JST研究/論文GPT / ChatGPT

DragonCrawl: スケーラブルなモバイル エンドツーエンド テストのための生成的なインテント ベースのフレームワーク

モバイル アプリケーションが複雑になるにつれて、従来のエンドツーエンド (E2E) テスト フレームワークは、UI の不安定性、メンテナンスのオーバーヘッド、クロスプラットフォームのスケーラビリティに苦労しています。この論文では、埋め込みベースの類似性マッチングから大規模な言語モデルを使用した生成意図ベースの推論に進化した、連続回帰テスト用の AI 駆動モバイル テスト システムである DragonCrawl について説明します。探索的テストとクラッシュ検出に焦点を当てたこれまでの LLM ベースのテスト研究とは異なり、DragonCrawl はコード変更のたびに特定のユーザー フローを検証し、重要な機能を破壊するコミットをブロックします。 GPT-4o のマルチモーダル機能を活用することで、DragonCrawl は、CI/CD パイプラインで継続的に実行される 1,013 の自動テスト全体で、iOS で 91.6%、Android で 92.2% の合格率を達成しました。このシステムにより、テストのオンボーディング時間が 96 ~ 120 時間から 4 時間未満に短縮され、開発者のテスト メンテナンスの労力が推定 27 年節約されました。 V1 (セマンティック埋め込みマッチング) から V2 (生成的インテントベース推論) へのアーキテクチャの進化を紹介し、トークン爆発やメモリ制約などの実装上の課題について議論し、運用環境での運用経験をレポートします。エンドステート検出のためのマルチモーダルビジョンとバックエンドステート遷移を呼び出すツールの統合により、UI インタラクションとシステムステートの橋渡しとなる包括的な回帰テストが可能になります。私たちの結果は、AI 主導のテストが安定性を維持しながら、従来の自動テストの脆弱性を排除し、大規模な継続的な品質保証を可能にすることを示しています。

原文 (English)

DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing

As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cross-platform scalability. This paper presents DragonCrawl, an AI-driven mobile testing system for continuous regression testing that has evolved from embedding-based similarity matching to generative intent-based reasoning using large language models. Unlike prior LLM-based testing research focused on exploratory testing and crash detection, DragonCrawl validates specific user flows on every code change, blocking commits that break critical functionality. By leveraging GPT-4o's multimodal capabilities, DragonCrawl achieves 91.6% pass rate on iOS and 92.2% on Android across 1,013 automated tests running continuously in CI/CD pipelines. The system reduces test onboarding time from 96-120 hours to under 4 hours and has saved an estimated 27 developer years in test maintenance effort. We present the architectural evolution from V1 (semantic embedding matching) to V2 (generative intent-based reasoning), discuss implementation challenges including token explosion and memory constraints, and report operational experience from production deployment. The integration of multimodal vision for end-state detection and tool calling for backend state transitions enables comprehensive regression testing that bridges UI interactions with system state. Our results demonstrate that AI-driven testing can maintain stability while eliminating the brittleness of traditional automated tests, enabling continuous quality assurance at scale.

13:00 JST画像/動画生成

SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction

In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark b…

13:00 JST画像/動画生成

WaiT for the Signal: Simple Frequency-Aware Flow-Matching

As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for…

13:00 JST研究/論文

RDF ルールの層化否定: 正しいアプローチ (拡張バージョン)

N3 ルールや SHACL ルールなどの RDF ルール言語とデフォルトの否定を組み合わせるのは困難です。否定を階層化する既存の方法は、個々のトリプルが潜在的な依存関係を有意義に制限するのに十分な情報を持っていないため、RDF ルールでは失敗することがよくあります。ルール ヘッド内の空白ノードは問題をさらに複雑にします。ルール適用の順序によって新しい値が作成されるかどうかが決まり、その結果、否定によるルールの適用性が変わる可能性があるからです。これらの未解決の問題を解決するために、否定を含む RDF ルールと存在ルール一般の適切な動作を保証する堅牢な新しい条件として連鎖階層化を提案します。私たちの条件は、潜在的な複数ステップの導出の精緻な分析と、整合性制約を使用して不可能なケースを破棄するメカニズムを組み合わせています。チェーンの階層化を考慮した任意の順序でルールを適用すると、通常の失敗としての否定セマンティクスの下で、一意で無駄がなく、正当化された RDF グラフが導出されることが保証されます。実用性を示すために、プロトタイプの実装も提供します。

原文 (English)

Stratified Negation in RDF Rules: A Correct Approach (Extended Version)

Combining RDF rule languages, such as N3 or SHACL Rules, with default negation is challenging. Existing methods to stratify negation often fail for RDF rules, since individual triples do not carry enough information to meaningfully restrict potential dependencies. Blank nodes in rule heads further complicate the matter, since the order of rule applications may determine whether new values are created, which in turn can change the applicability of rules with negation. To solve these open problems, we propose chain stratification as a robust new condition that guarantees a well-behaved semantics for RDF rules with negation, and existential rules in general. Our condition combines an elaborate analysis of potential multistep derivations with a mechanism for using integrity constraints to discard impossible cases. Applying rules in any order that respects chain stratification is guaranteed to derive an RDF graph that is unique, lean, and justified under the usual negation-as-failure semantics. To show the practicality, we also provide a prototype implementation.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring…

13:00 JSTLLM/生成AILlamaQwen

抵抗を伴う転がり:好みに合わせて最適化された LLM カウンセラーは、モチベーションを高める面接において、目標の持続性を人間関係の同調と引き換えにできる

動機づけ面接(MI)では、クライエントの持続的なトーク(現状維持の主張)はカウンセラーに抵抗を求めますが、この動きは2つの相反する方法で失敗する可能性があります:降伏(信頼関係を維持するために変化の議題を放棄する)または対決(クライエントの自主性を無視して議論または指示する)。我々は、動機づけ面接治療誠実性(MITI)コード、目標持続性(GP)と関係的同調(RA)に基づいたカウンセラーの反応の2軸評価を導入し、抵抗を伴うローリングが両方で高いという4象限の枠組みを導き出し、選好の最適化によって1つの失敗を罰することが抵抗を伴うローリングを教えるのか、それともその反対を誘発するのかを尋ねます。専門家によって注釈が付けられた AnnoMI コーパスから、オンポリシー ネガティブを使用して、失敗が拒否される点のみが異なる優先設定セットを持つトピックに共通の直接優先最適化データを構築します。 AnnoMI の専門家ラベルに照らして検証され、訓練を受けた人間のプログラマーによって再チェックされた自動判定機能は、ファイアウォールの下で、互いに素なモデル ファミリが生成、ラベル付け、判定する各ベースに対してブラインド ペアごとの勝率をスコアリングします。 Qwen ファミリーと Llama ファミリーにまたがる 3 つの整列された命令モデル全体で、対決にペナルティを与えると、すべてのベースおよびすべてのシード実行で確実にゴール持続性が同等以下に低下し、強力なコストがかかります。一方、同調ゲインはベースに依存し、3 つのベースのうち 2 つに存在しますが、3 番目のベースには存在しません。これらのモデルはポリシーに従って降伏することはほとんどないため、降伏にペナルティを与えることは不活性であり、そのため取引は各ベースの失敗プロファイルによってゲートされます。プロンプトのみの制御は、目標持続コストなしで調整を高め、調整自体ではなく最適化にコストを配置します。

原文 (English)

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy). We introduce a two-axis evaluation of counselor responses, anchored in the Motivational Interviewing Treatment Integrity (MITI) code, Goal Persistence (GP) and Relational Attunement (RA), yielding a four-quadrant framing in which rolling with resistance is high on both, and we ask whether penalizing one failure through preference optimization teaches rolling with resistance or provokes its opposite. From the expert-annotated AnnoMI corpus we build topic-disjoint Direct Preference Optimization data whose preference sets differ only in which failure is rejected, using on-policy negatives. An automatic judge, validated against AnnoMI's expert labels and rechecked by trained human coders, scores blind pairwise win-rates against each base under a firewall in which disjoint model families generate, label, and judge. Across three aligned instruction models spanning the Qwen and Llama families, penalizing confrontation reliably lowers goal persistence below parity, on every base and in every seed run, a robust cost, whereas the attunement gain is base-dependent, present on two of the three bases but absent on the third. Penalizing capitulation is inert, because these models rarely capitulate on-policy, so the trade is gated by each base's failure profile. A prompt-only control raises attunement without the goal-persistence cost, locating the cost in the optimization rather than in attunement itself.

13:00 JST研究/論文

Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity

Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories…

13:00 JST画像/動画生成研究/論文

磁気共鳴画像法によるマルチタスク 3D 脳腫瘍セグメンテーションのための深層学習モデルの統合ベンチマーク

磁気共鳴画像法 (MRI) による脳腫瘍の自動セグメンテーションは、コンピューター支援診断、治療計画、疾患モニタリングにおける基本的なタスクとなっています。最近、多数の深層学習アーキテクチャが提案されていますが、公開された研究では異なるデータセット、前処理戦略、トレーニング プロトコル、評価手順が使用されていることが多いため、客観的な比較は依然として困難です。この研究では、代表的な畳み込みニューラル ネットワーク (CNN)、Transformer ベースのモデル、および最近の状態空間モデル (SSM) アーキテクチャを均一な実験条件下で比較するための統一された実験ベンチマークを示します。 3D U-Net、SegResNet、Swin UNETR、SegMamba、SegMambaV2 を含む 5 つの最先端の 3 次元セグメンテーション モデルが、頭蓋内髄膜腫セグメンテーション (BraTS 2023) と治療後神経膠腫セグメンテーション (BraTS 2024) という異なる臨床シナリオを表す 2 つの脳腫瘍セグメンテーション データセットで評価されます。すべてのアーキテクチャは、公平な比較を保証するために、同一の前処理、データ拡張、最適化戦略、および評価プロトコルを使用してトレーニングされます。パフォーマンスは、セグメンテーション精度メトリクスと、推論時間や各モデルのサイズなどの計算コスト指標を使用して評価されます。この結果は、セグメンテーションの精度と計算効率の間のトレードオフに関する実用的な洞察を提供し、困難な 3 次元脳腫瘍セグメンテーション タスクに対するさまざまなアーキテクチャ パラダイムの適合性を強調しています。

原文 (English)

A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging

Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment planning, and disease monitoring. Although numerous deep learning architectures have recently been proposed, objective comparisons remain challenging because published studies often employ different datasets, preprocessing strategies, training protocols, and evaluation procedures. This work presents a unified experimental benchmark for comparing representative convolutional neural networks (CNNs), Transformer-based models, and recent State Space Model (SSM) architectures under homogeneous experimental conditions. Five state-of-the-art three-dimensional segmentation models, including 3D U-Net, SegResNet, Swin UNETR, SegMamba, and SegMambaV2, are evaluated on two brain tumor segmentation datasets representing distinct clinical scenarios: intracranial meningioma segmentation (BraTS 2023) and post-treatment glioma segmentation (BraTS 2024). All architectures are trained using identical preprocessing, data augmentation, optimization strategies, and evaluation protocols to ensure a fair comparison. Performance is assessed using segmentation accuracy metrics together with computational cost indicators, including inference time and the size of each model. The results provide practical insights into the trade-offs between segmentation accuracy and computational efficiency, highlighting the suitability of different architectural paradigms for challenging three-dimensional brain tumor segmentation tasks.

13:00 JSTLLM/生成AI

TextCloak: RL 駆動の学習不可能なテキストによる不正な LLM 悪用を阻止する

大規模言語モデル (LLM) の急速な発展により、幅広い言語タスクにわたって大幅な進歩がもたらされましたが、同時に、不正なデータ悪用やプライバシー漏洩に対する懸念も高まっています。学習不可能な例 (UE) は、慎重に設計された摂動をデータに導入することで、その上でトレーニングされたモデルが実用性の低下を示すようにすることで、有望な防御手段を提供します。ただし、テキスト保護の既存の方法は、主に識別言語モデルでの分類タスク (感情分析など) 用に設計されており、多くの場合、クラス固有の言語キューの挿入に依存しているため、LLM のオープンエンド生成設定では有効性が制限されます。この研究では、不正な LLM 悪用からテキスト データを保護するための RL 駆動フレームワークである TextCloak を提案します。 TextCloak は、意味の忠実性と言語の自然さを維持しながら、クリーン テキストのバッチを学習不可能な例に変換する生成ポリシーを採用しています。ポリシーを最適化するために、GRPO-UE を導入します。これは、微調整されたサロゲート LLM で誘発される下流の劣化に基づいて、生成された学習不可能なテキストに報酬を与え、グループ相対ポリシーの最適化を通じてジェネレーター パラメーターを更新します。この 2 レベルの最適化により、ジェネレーターはクラス固有の手がかりを超えて一般化可能な保護パターンを発見できるようになります。 6 つの公的に利用可能なデータセットと 9 つの最先端の LLM に関する包括的な実験により、TextCloak は正当な使用のためのテキストのユーティリティを維持しながら、不正な微調整を一貫して妨害することが実証されました。さらなる分析により、モデル アーキテクチャ、トレーニング構成、適応型攻撃にわたるその移行可能性と堅牢性が確立され、不正な LLM 悪用に対する実用的な防御としてその幅広い適用可能性が強調されています。

原文 (English)

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage. Unlearnable examples (UEs) offer a promising defense by introducing carefully designed perturbations into data such that models trained on them exhibit degraded utility. However, existing methods for text protection are primarily designed for classification tasks (e.g., sentiment analysis) in discriminative language models and often rely on injecting class-specific linguistic cues, which limits their effectiveness in the open-ended generation settings of LLMs. In this work, we propose TextCloak, an RL-driven framework for protecting textual data against unauthorized LLM exploitation. TextCloak employs a generative policy that transforms batches of clean text into unlearnable examples while preserving semantic fidelity and linguistic naturalness. To optimize the policy, we introduce GRPO-UE, which rewards generated unlearnable text based on the downstream degradation they induce in fine-tuned surrogate LLMs and updates the generator parameters via group-relative policy optimization. This bi-level optimization enables the generator to discover generalizable protective patterns beyond class-specific cues. Comprehensive experiments on six publicly available datasets and nine state-of-the-art LLMs demonstrate that TextCloak consistently impairs unauthorized fine-tuning while maintaining text utility for legitimate use. Further analyses establish its transferability and robustness across model architectures, training configurations, and adaptive attacks, highlighting its broad applicability as a practical defense against unauthorized LLM exploitation.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

LLM 修復エージェントの検証証拠: 合格したもののどれだけが実際にバグをテストしているか?

修理エージェントがテストを実行し、合格したことが確認されると、その結果は報告された欠陥に関する証拠として扱われます。私たちは、その治療がどのくらいの頻度で保証されるかを測定します。 BSG-VA (バグ状態/候補状態/ゴールド フィックス検証分析) は、各検証コマンドを正確な作業ツリー状態でキャプチャし、テスト専用パッチを抽出して、元のバグのあるコード (B)、候補状態 (S)、および開発者のゴールド フィックス (G) でコマンドを再生します。キャプチャされた結果とリプレイ結果は、ゴールドに合わせたバグの識別から回帰のみから誤解を招くものまで、あらゆるイベントに証拠の役割を割り当てます。 110 のタスクに対する 643 のロールアウトにおける 3,730 のイベント全体で、肯定的な比較可能なイベントの 46.0% にはバグを識別する情報がありませんでした。ベースラインのロールアウトの 23.8% は、フィードバックが注入されていない場合、全体がこの種の肯定的な証拠ベースであるパッチで終了します。 3 アーム実験では、B リプレイ結果をエージェントに返すことでこのパターンが変化するかどうかをテストします。バグコントラストフィードバックは、注意を一致させたリマインダーと比較して、証拠不十分なクロージャを 7.8 パーセントポイント (p = 0.0029) 減少させ、バグ識別の証拠を 7.4 ポイント (p = 0.011) 上昇させます。修復成功のための検出可能なコストはありません。どちらの推定値も、対象となる事前に指定された 10 パーセント ポイントの最小効果量を下回っているため、実際の規模は依然として不確実です。改善のおよそ 3 分の 1 はリマインダーのみによるものです。足場とモデルを変更した 2 つの探索的複製にわたって、B リプレイ コンテンツは、制約のないツール使用ループの下で gpt-5.6-sol でのみ検出可能な増分を追加します。 BSG-VA は、必要なコードの状態と実行環境を保持する再生可能な修復軌跡に事後的に適用します。キーワード: プログラム修復エージェント、検証証拠、テストの適切性、大規模言語モデル、ソフトウェア品質、管理された実験。

原文 (English)

Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?

When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is warranted. BSG-VA (buggy-state/candidate-state/gold-fix validation analysis) captures each validation command at its exact working-tree state, extracts a test-only patch, and replays the command on the original buggy code (B), the candidate state (S), and the developer gold fix (G). The captured outcome and the replay results assign every event an evidence role, from gold-aligned bug-discriminating through regression-only to misleading. Across 3,730 events in 643 rollouts on 110 tasks, 46.0% of positive comparable events carry no bug-discriminating information; 23.8% of baseline rollouts, with no feedback injected, close with a patch whose entire positive evidence base is of this kind. A three-arm experiment tests whether returning the B-replay outcome to the agent changes this pattern. Bug-contrast feedback reduces evidence-inadequate closure by 7.8 percentage points relative to an attention-matched reminder (p = 0.0029) and raises bug-discriminating evidence by 7.4 points (p = 0.011), with no detectable cost to repair success. Both estimates fall below the prespecified 10-percentage-point smallest effect size of interest, so practical magnitude remains uncertain. Roughly a third of the improvement traces to the reminder alone; across two exploratory replications, varying the scaffold and the model, the B-replay content adds a detectable increment only with gpt-5.6-sol under the unconstrained tool-use loop. BSG-VA applies post hoc to any replayable repair trajectory that preserves the required code states and execution environment. Keywords: program repair agents, validation evidence, test adequacy, large language models, software quality, controlled experiment.

13:00 JST研究/論文Claude

RareSense: トランザクション データの異常を検索するための希少性を意識した類似性検索

まばらな設定値データに対する類似性検索は、ジャカード、コサイン、ハミングなどの古典的な尺度がアトミックな重複を通じてオブジェクトを比較するため、頻繁に使用される背景属性によって支配されることがよくあります。 IDF (逆文書頻度) 重み付けはこの影響を部分的に軽減しますが、原子単位のままであり、有益な高次の共起を明示的に表すことはできません。スパースなトランザクション異常データ用の希少性を認識した類似性フレームワークである RareSense を紹介します。 RareSense は、中間構造として最小限のレア アイテムセットをマイニングし、信頼できるレア関連ルールを導き出し、オブジェクトを疎なレア ルール プロファイルにマッピングし、重み付けされた Jaccard 類似性を使用してそれらを比較します。ルールの重みは、逆サポート、信頼性、リフト、構造の複雑さ、安定性を組み合わせているため、均一な特徴の重複ではなく、共有されたまれな証拠によって近傍が決定されます。我々は、IDF 重み付け Jaccard が RareSense の制限されたシングルトン ケースであること、および誘導された距離が元のオブジェクトの擬似メトリックであり、同一のルール プロファイルによって定義された等価クラスのメトリックであることを示します。サイバーセキュリティと一般的なカテゴリ領域にわたる 4 つのベンチマーク ファミリにわたる実験では、評価された類似性尺度の中で、RareSense が観察された最高のマクロ平均クエリ条件付き検索パフォーマンスを達成していることが示されています。統計分析では、全体的な有意な差が示されており、修正された一対の比較ではアトミック ベースラインよりも RareSense が有利になっています。利益は引き続きワークロードに依存し、異常が再現可能なまれな高次構造を共有する場合に最も強くなります。グローバルな異常ランキングに関して、RareSense は、いくつかの強力な専用検出器と統計的に同等の状態を維持しながら、観察された最高のマクロ平均パフォーマンスを達成します。

原文 (English)

RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data

Similarity search over sparse set-valued data is often dominated by frequent background attributes because classical measures such as Jaccard, cosine, and Hamming compare objects through atomic overlap. IDF (Inverse document frequency) weighting partially reduces this effect but remains atom-wise and cannot explicitly represent informative higher-order co-occurrences. We introduce RareSense, a rarity-aware similarity framework for sparse transactional anomaly data. RareSense mines minimal rare itemsets as intermediate structures, derives reliable rare association rules, maps objects into sparse rare-rule profiles, and compares them using weighted Jaccard similarity. Rule weights combine inverse support, confidence, lift, structural complexity, and stability, so that neighborhoods are determined by shared rare evidence rather than uniform feature overlap. We show that IDF-weighted Jaccard is a restricted singleton case of RareSense, and that the induced distance is a pseudometric on the original objects and a metric over equivalence classes defined by identical rule profiles. Experiments across four benchmark families spanning cybersecurity and general categorical domains show that RareSense attains the highest observed macro-average query-conditioned retrieval performance among the evaluated similarity measures. The statistical analysis indicates significant overall differences, with corrected paired comparisons favoring RareSense over the atomic baselines. The gains remain workload-dependent and are strongest when anomalies share repeatable rare higher-order structure. For global anomaly ranking, RareSense achieves the highest observed macro-average performance while remaining statistically comparable to several strong dedicated detectors.

13:00 JSTLLM/生成AIGPT / ChatGPT

追加はマシン、削除は人間: LLM コード編集における削除回避の測定と軽減

大規模な言語モデルでは、実稼働コードの作成と修復が増えていますが、テストに合格したパッチによってコードベースの保守が困難になるという証拠が増えています。私たちは、具体的な原因の 1 つを特定します。それは、削除回避、つまり、意図した編集で削除する必要があるコードを保持する体系的な傾向です。公式の SWE ベンチ検証済みリーダーボードにある 5 つの主要なモデル全体で、開発者パッチに対する削除リコールは、5 つすべてが解決したタスクでも最大 71.7% に達し、モデルは必要な削除の 92% 以上で適切なファイルに到達しますが、ケースの 52% 未満で正確なラインをカットします。その代わり、合格したパッチの 29.0% は、ターゲットのコードをガードまたはフォールバックでラップしており、これを私たちが Guard-and-Go と呼んでいます。このようなパッチが合格するのは、元のテストでは削除がほとんどチェックされないためです。対象のコードが残っている場合に失敗するテストで 34 の検証済みタスクを改造すると、クローズドとオープンの重みにまたがる 4 つのフロンティア モデルが 63.2% から 41.9% に低下しました。実際の修復には削除と追加が混在しているため、必要な編集全体が削除である実際のコミットからマイニングされた 200 のタスクのベンチマークである CanItDelete を厳選しています。追加作業がなくなっても、最良のモデルは依然として 5 件中 1 件のタスクに失敗し、小規模なオープン モデルは 18.0% に低下します。次に、4 つの累積プロンプトの下で GPT-5.6 Sol をアブレーションします。正確な行を指定するまでは成功はほとんど進みません。不完全な削除はほぼなくなりますが、モデルがスパンを超えて削除するか、代わりにコードを追加するため、成功率は 80.5% までしか上がりません。最後に、パイロットスタディを通じて、潜在的な修正の 1 つを示します。ポストトレーニング中に削除を教えると、削除の回避が減少し、より広範なコード編集パフォーマンスが向上します。これは、動作が到達範囲を超えているのではなく、トレーニングが不十分であることを示唆しています。

原文 (English)

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebases harder to maintain. We identify one concrete source: deletion avoidance, the systematic tendency to retain code that an intended edit requires removing. Across the five leading models on the official SWE-bench Verified leaderboard, deletion recall against the developer patch reaches at most 71.7% even on tasks all five solve, and models reach the right file for over 92% of required deletions but cut the exact line in under 52% of cases. Instead, 29.0% of passing patches wrap the targeted code in a guard or fallback, a pattern we call Guard-and-Go. Such patches pass because the original tests rarely check removal: when we retrofit 34 Verified tasks with tests that fail if the targeted code remains, four frontier models spanning closed and open weights fall from 63.2% to 41.9%. Because real repairs mix removal with addition, we curate CanItDelete, a benchmark of 200 tasks mined from real commits whose entire required edit is deletion. Even with the addition work gone, the best model still fails one task in five, and smaller open models fall to 18.0%. We then ablate GPT-5.6 Sol under four cumulative prompts; success moves little until we supply the exact lines, which nearly eliminate incomplete deletion yet raise success only to 80.5% because the model then deletes beyond the spans or adds code instead. Finally, through a pilot study we show one potential fix: teaching deletion during post-training reduces deletion avoidance and improves broader code-editing performance, suggesting the behavior is undertrained rather than beyond reach.

13:00 JSTLLM/生成AI研究/論文

幼稚園から高校までの教育者の AI 使用を概念化するための人間と LLM の協調帰納コーディング

質的研究者は、手動コーディングだけで対応できる規模を超えるインタラクション コーパスに遭遇することが多くなり、分析アシスタントとして大規模言語モデル (LLM) が頻繁に提案されています。未解決の問題は、LLM が定性分析に参加できるかどうかではなく、どの程度、どの段階で、どのような保護措置の下で参加できるかということです。この記事では、幼稚園から高等学校までの教育者と生成 AI プラットフォームの間で交換される 45,000 のメッセージから階層コードブックを開発するためにオープン、軸方向、選択的コーディングを適応させた、マルチフェーズの人間と LLM の共同パイプラインの詳細な手順を説明します。 3 つのフェーズにわたって、LLM は候補ラベルと構造化されたアノテーションを大規模に生成しましたが、人間の研究者はカテゴリの定義、決定のマージ、および解釈フレームワークに対する概念的な権限を保持していました。次に、結果として得られた機器は体系的な人間によるコーディングを通じてテストされました。教育分野の専門知識を持つ 3 人の訓練を受けたプログラマーが、コードブックを 2,560 メッセージの独立したサンプルに適用し、マルチラベル アノテーションに適した設定値の一致測定を使用した反復キャリブレーションを通じて信頼性を確立し、LLM 支援フェーズでは表面化しなかった 5 つのコードで機器を拡張しました。最終的なコードブックは、19 のカテゴリと 6 つのドメイン内の 72 項目で構成されます。私たちは、会話型分析単位の選択、解釈エージェントではなくラベル付け手段としての LLM の扱い、マルチラベルコーディングの下で​​のコーダー間一致の測定、および人間のドメイン専門知識が依然として決定的であった条件など、パイプラインに必要な方法論的な決定を振り返ります。このアカウントは、人間の解釈権限を維持しながら、コードブック開発における LLM 支援を検討している質的研究者向けの監査可能なテンプレートとして提供されます。

原文 (English)

Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use

Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can participate in qualitative analysis but to what extent, in what phases, and under what safeguards. This article provides a detailed procedural account of a multi-phase human-LLM collaborative pipeline that adapted open, axial, and selective coding to develop a hierarchical codebook from 45,000 messages exchanged between K-12 educators and a generative AI platform. Across three phases, LLMs generated candidate labels and structured annotations at scale, while human researchers retained conceptual authority over category definitions, merging decisions, and interpretive frameworks. The resulting instrument was then tested through systematic human coding, in which three trained coders with educational domain expertise applied the codebook to an independent sample of 2,560 messages, established reliability through iterative calibration using set-valued agreement measures appropriate for multi-label annotation, and extended the instrument with five codes that the LLM-assisted phases had not surfaced. The final codebook comprises 72 items within 19 categories and six domains. We reflect on the methodological decisions the pipeline required, including the choice of a conversational unit of analysis, the treatment of the LLM as a labeling instrument rather than an interpretive agent, the measurement of intercoder agreement under multi-label coding, and the conditions under which human domain expertise remained decisive. The account is offered as an auditable template for qualitative researchers considering LLM assistance in codebook development while preserving human interpretive authority.

13:00 JSTLLM/生成AIビジネス/資金調達

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that…

13:00 JSTLLM/生成AI

TORUS: 統合オーディオ モデルのレンダリング理解の自己コヒーレンスのテスト

オーディオの理解、オーディオの生成、さらにはオーディオ編集が可能な統合オーディオ モデルが急速に普及しています。しかし、それらに関する基本的な質問は未解決のままです。それは、統合モデルの 2 つのヘッドが同じオーディオについて同意しているのかということです。現在の実践では、特殊なベンチマークで各機能を個別に評価し、モデルがそれ自体の世代を理解できるかどうかを問うことはありません。オーディオネイティブの統合モデルの最初の自己コヒーレンス テストである TORUS を紹介します。 TORUS は 48 の 3 段階の自己一貫性テストで構成され、5 つのタスク ファミリにわたる音声、音声、音楽にわたる 432 の 6 択の質問を実行します。私たちは、最先端の特殊な生成、編集、理解モデルを組み合わせたカスケード ベースラインと並行して、5 つのオープンな統合モデルを総合的に評価します。最も優れた統合モデルは、カスケード ベースラインの 63.2% と 16.7% の確率下限に対して、質問の 50.5% に回答します。モデルたちは音声編集に苦戦しています。評価されたオーディオ モデル (特殊化および統合) の中で、限られた自己コヒーレンスが観察されたため、自己コヒーレンスを将来のオーディオ システムにとって不可欠なテストとして位置づけています。

原文 (English)

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among the evaluated audio models (specialized and unified), we observe limited self-coherence, and thus position self-coherence as an essential test for future audio systems.

13:00 JST研究/論文

デザイン コンセプト: 技術労働者の間での地政学的な反映を足場とする

この論文では、地政学的に関連するテクノロジー企業のテクノロジー労働者の間で地政学的な再帰性を促進するための、投機的なヒューマン コンピューター インタラクション設計提案を示します。国際関係および科学技術研究における最近の研究では、テクノロジー企業とその従業員が、その決定が国際的な力関係を形作る地政学的主体であるとますます認識されています。しかし、既存の責任あるイノベーションと責任ある AI のアプローチは、現代の AI 開発を支える地政学的な物語や想像に関わることはほとんどありません。この論文は、再帰性、反射的 HCI、および計算的物語に関する創造的 HCI 作業に関する RI の研究に基づいて、ユーザーがテクノロジー、権力、地政学を中心とした投機的シナリオに参加する、AI 対応の対話型物語システムを提案します。このシステムは、物語的な対話、アーキタイプの割り当て、社会的に足場を組んだワークショップの考察を通じて、対象ユーザーがより広範な社会技術システム内での前提、価値観、立場を批判的に検討することを奨励することを目的としています。私たちは、投機的な物語システムは、規範的アプローチや道徳的アプローチに依存することなく、責任あるテクノロジーへの取り組みに地政学的な再帰性を導入するための生産的な手段を提供する可能性があると主張します。

原文 (English)

Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers

This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolitically relevant technology companies. Recent scholarship in International Relations and Science and Technology Studies increasingly recognizes technology firms and their workers as geopolitical actors whose decisions shape international dynamics. However, existing Responsible Innovation and Responsible AI approaches rarely engage with the geopolitical narratives and imaginaries that underpin contemporary AI development. Building upon RI scholarship on reflexivity, reflective HCI, and creative HCI work on computational narratives, this paper proposes an AI-enabled interactive narrative system in which users engage with a speculative scenario centred on technology, power, and geopolitics. Through narrative interaction, archetype assignment, and socially scaffolded workshop reflection, the system aims to encourage target users to critically examine their assumptions, values, and positionality within broader sociotechnical systems. We argue that speculative narrative systems may offer a productive avenue for introducing geopolitical reflexivity into responsible technology initiatives without relying on prescriptive or moralising approaches.

13:00 JST研究/論文

ゲート型 Q ラーニング: ポリシー外のバイアスを加えて味わう

マルチステップ単位の割り当ては、サンプル効率の高い強化学習にとって重要ですが、Q 学習におけるオフポリシー バイアスの管理は依然として根本的な課題です。 30 年間、専門家は二者択一の選択に限定されてきました。つまり、適格性追跡を大幅に切り捨ててバイアスを排除するか (ワトキンスの Q($\lambda$))、バイアスを無視してより速く学習する一方で、有害な誤差を値の推定値に注入するか (ペンの Q($\lambda$)) です。 Q ラーニングの貪欲なターゲット ポリシーの下では重要度のサンプリング比率が崩壊するため、現代のオフポリシー推定器はこの緊張を解決できません。私たちは、歴史的な 2 つの極端な点の間を滑らかに補間することでこのジレンマを解決する新しいアルゴリズム フレームワークである Gated Q-learning を紹介します。私たちのアプローチでは、重要度のサンプリングに依存するのではなく、継続的な状態アクション依存のゲート メカニズムを採用し、探索を意識した方法で適格性トレースを選択的に減衰します。我々は、このメカニズムに厳密な理論的基礎を提供し、期待される演算子が縮小マッピングのままであることを証明し、その正確な固定点を導き出します。経験的評価により、中間ゲートにより安全に長い単位割り当て期間が可能になり、どちらの極端な場合よりも初期学習が高速化されることが確認されています。ゲート型 Q ラーニングは、重要度サンプリングに代わるシンプルな代替手段を提供すると同時に、Q ラーニング エージェントにおける効果的なマルチステップ範囲とオフポリシー バイアスの量のカスタマイズを可能にします。

原文 (English)

Gated Q-learning: Add Off-Policy Bias to Taste

Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge. For 30 years, practitioners have been limited to a binary choice: eliminate the bias at the cost of severely truncated eligibility traces (Watkins' Q($\lambda$)), or ignore the bias to learn faster while injecting detrimental errors into the value estimates (Peng's Q($\lambda$)). Modern off-policy estimators fail to resolve this tension, as importance-sampling ratios collapse under Q-learning's greedy target policy. We introduce Gated Q-learning, a novel algorithmic framework that ends this dilemma by smoothly interpolating between the two historical extremes. Rather than relying on importance sampling, our approach employs a continuous, state-action-dependent gating mechanism to selectively attenuate eligibility traces in an exploration-aware manner. We provide a rigorous theoretical foundation for this mechanism, proving that the expected operator remains a contraction mapping and deriving its exact fixed point. Empirical evaluations verify that intermediate gating safely enables longer credit-assignment horizons, yielding faster initial learning than either extreme. Gated Q-learning offers a simple alternative to importance sampling while enabling customization of the effective multistep horizon and the amount of off-policy bias in Q-learning agents.

13:00 JSTLLM/生成AI

FairFund-Bench: LLM リソース割り当てにおける分配バイアスの評価

大規模言語モデル (LLM) は希少なリソースの配布にますます関与しており、人種や性別などの特性に基づいた偏った割り当てに関する懸念が生じています。しかし、最近の LLM 監査では一貫性のない結果が得られ、同じモデルであっても女性と少数民族に対する肯定的および否定的な差別の証拠が見つかりました。我々は、この不一致が監査形式の違いから生じる可能性があることを示し、評価タスク(評価、ランキング、または割り当て)、比較コンテキスト(単一または複数の刺激)、および監査が透明か偽装であるかなど、以前の監査設計の主要な機能を体系的に変更するベンチマークであるFairFund-Benchを紹介します。このベンチマークは、3 つのドメイン、4 つの人種および 2 つの性別カテゴリー、および生活保護受給理論から導き出されたニーズの 5 つの因果関係にわたる、人間が作成したテンプレート (130 万件の実際の G​​oFundMe キャンペーンに対して調整済み) から作成された 600 件の経済援助リクエストで構成されています。 14 のモデルにわたって、監査形式によりバイアスの方向が変わります。モデルは、申請者を個別に評価する場合は少数派に有利ですが、並べてランク付けする場合は一部のグループに不利益を与えます。バイアスの大きさは、全体としては小さいものの、透明な監査よりも偽装監査の方が数倍大きく、申立人の名前だけが異なる控訴に直面して、圧倒的に資金を均等に分割するモデルが採用されている。対照的に、因果的フレーミング効果は人口統計効果をおよそ 1 桁上回っており、モデルや監査形式全体で一貫しており、現在の LLM が人間のふさわしい評価を確実に再現していることを示しています。ベンチマーク スコア モデルは 4 つの基準 (人口統計上の偏り、価値の調整、タスク間の一貫性、およびコンテキスト間の一貫性) に基づいて公開されており、他の実質的な領域に容易に適用できます。

原文 (English)

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.

13:00 JST画像/動画生成

DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models

Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the de…

13:00 JST画像/動画生成

Retrieval-Driven Training-Free AI-Generated Video Attribution

AI-generated videos are becoming increasingly realistic and difficult to distinguish from authentic ones, which facilitates malicious misus…

13:00 JSTLLM/生成AI

低ランクの防御と回路誘導サロゲートによる効率的な LLM 敵対トレーニング

敵対的トレーニングは、敵対的攻撃に対する最も効果的な防御手段の 1 つですが、特に大規模言語モデル (LLM) の場合、現代の規模では依然として計算コストが法外に高くなります。潜在的敵対的トレーニング (LAT) などの既存の緩和戦略が開発されていますが、依然として高い計算コストがかかります。この研究では、LAT を高速化するための計算効率の高い戦略を 2 つの相補的な観点から包括的に調査します。 (1) 防御側の最適化: LAT 内の表現微調整 (ReFT) を調査し、ReFT と攻撃を適用するトークンに不一致がある場合の潜在的な問題を明らかにします。 (2) 攻撃側の最適化: 各 LAT 反復で敵対的攻撃を計算する際、LLM から関連する回路のみを抽出して軽量のサロゲート モデルを構築し、攻撃生成中に完全なモデルを通過する前方から後方へのパスでの計算を回避します。両方の観点について、提案された戦略の有効性を示す理論的根拠と数値的証拠を提供します。最終的に、完全な微調整を備えた標準 LAT と比較して、私たちの方法は平均して、ステップごとの敵対的トレーニング FLOP を 48.1% 削減し、必要なトレーニング可能なパラメーターはわずか 0.0118% のみです。

原文 (English)

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates

Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent adversarial training (LAT), have been developed, they still incur a high computational cost. In this work, we comprehensively investigate computation-efficient strategies to speed up LAT from two complementary perspectives: (1) Defense-side optimization: We explore the representation fine-tuning (ReFT) within LAT, and reveal a potential issue if there is a mismatch on which tokens to apply ReFT and the attack. (2) Attack-side optimization: When computing adversarial attacks in each LAT iteration, we extract only the relevant circuits from the LLM to construct a lightweight surrogate model, avoiding the computation in the forward-backward passes through the full model during the attack generation. For both perspectives, we provide theoretical justifications and numerical evidence to illustrate the effectiveness of the proposed strategies. Ultimately, compared to standard LAT with full fine-tuning, our method on average reduces per-step adversarial-training FLOPs by 48.1% while requiring only 0.0118% trainable parameters.

13:00 JSTLLM/生成AIGPT / ChatGPT

LLM の使用と科学的生産性の間の強固な関連性: 停止時間の選択の評価

Renault、Bergeaud、Bosquet (以下、RBB) は、LLM の採用時期を著者の要約にフラグが立てられた最初の月とすることは、因果関係がない場合でも肯定的なイベント研究パスを生み出す可能性がある停止時間の選択を誘発すると主張しています。このメカニズムは数学的には可能ですが、無効効果の証明にはなりません。 RBB 独自のランダムなプラセボを検出器の実現フラグ率に再調整すると、測定された関連性がこのベンチマークをはるかに上回っており、アーチファクトが小さすぎて生産性の変化を説明できないことがわかります。さらに、タイミングアーチファクトが推定値を偏らせない一連の相補的な設計を使用して、LLM 導入と生産性の関連性を再推定します。つまり、ある年に導入し、別の年に生産量を測定する前後の比較、差分の差に対する保守的な対照グループ、導入日を決して定義しない強度ベースの仕様、およびフラグ率を固定したランクベースの測定です。肯定的な生産性の関連性はこれらの推定値すべてにわたって持続しますが、ChatGPT 以前のプラセボ データに対して同じテストを実行すると無効効果が返されます。 RBB が特定したアーティファクトは本物ですが限界があり、私たちが報告するパターンを考慮していません。

原文 (English)

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an adoption date, and a rank-based measurement holding the flag rate fixed. A positive productivity association persists across all of these estimates, while the same tests run on pre-ChatGPT placebo data return null effects. The artifact RBB identify is real but bounded, and it does not account for the pattern we report.

13:00 JST画像/動画生成

RAID: ビット反転イメージによる堅牢な AI 生成イメージ検出を目指して

画像生成モデルの急速な進歩により、人々が AI によって生成された画像と本物の画像を区別することがますます困難になっています。偽画像の悪用に伴う潜在的なリスクを防ぐために、AI による画像検出が大きな注目を集めています。既存の方法では、本物の画像と偽の画像の間に固有の違いが無視されているため、堅牢性と一般化能力が不足しています。この研究では、ビットプレーンを使用した AI 生成の画像検出を革新的に研究し、ビット反転画像を導入します。我々は、ビット反転画像の構築、勾配ベースのパッチ選択、畳み込み分類器で構成されるシンプルかつ効果的なパイプラインを提案します。さらに、私たちのアプローチの有効性を実証するために、数学的な観点からの理論的分析を提供します。また、AI 生成の画像検出のための 2 つの困難なデータセットも紹介します。広範な実験により、ジェネレーター間の一般化、データセット間の一般化、ゼロショットのパフォーマンスなど、さまざまな設定にわたるアプローチの有効性が検証されます。余分な機能がなければ、私たちのアプローチは 40 以上のベンチマークで既存の手法を上回り、同等の手法よりも 100 倍近く高速です。コードは https://github.com/renxi-seu/RAID にあります。

原文 (English)

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images

The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real ones. To prevent the potential risks associated with the misuse of fake images, AI-generated image detection has gained significant attention. Existing methods neglect the inherent differences between real and fake images, thus lacking robustness and generalization ability. In this work, we innovatively investigate AI-generated image detection using bit-planes, and introduce the bit-reversed image. We propose a simple yet effective pipeline consisting of construction of bit-reversed images, gradient-based patch selection and a convolutional classifier. Besides, we provide a theoretical analysis from the mathematical perspective to demonstrate the validity of our approach. We also introduce two challenging datasets for AI-generated image detection. Extensive experiments verify the effectiveness of our approach across different settings, including cross-generator generalization, cross-dataset generalization and zero-shot performance. Without bells and whistles, our approach outperforms existing methods on over 40 benchmarks, and is nearly 100 times faster than counterparts. The code is at https://github.com/renxi-seu/RAID.

13:00 JSTLLM/生成AI

PARALLEL: 明示的な制限の下での言語モデル学習のための前頭前部整列強化にインスピレーションを得たアプローチ

最近の言語モデルはさまざまなタスクにわたって優れたパフォーマンスを実現しますが、従来の適応では、ローカル更新の利点に関係なく、トレーニング サンプル全体に更新を均一に適用します。私たちは、言語モデル学習のための前頭前野整列強化にインスピレーションを得たアプローチである PARALLEL を提案します。目標関連制御と不確実性関連制御の相補的な役割に着想を得た PARALLEL は、これらの形式の情報を個別のコントローラー信号として表し、それらを現在のモデル表現と組み合わせます。強化にインスピレーションを得たコントローラーは、即時の光熱費フィードバックを使用して、サンプル依存の更新強度を割り当てます。したがって、PARALLEL は各サンプルにいつ、どの程度強く適応するかを学習し、不必要なパラメーターの変更を制限しながら有益な更新を優先します。 PARALLEL は、完全適応パフォーマンスの 94.1 ~ 99.2\% を維持しながら、選択的なベースラインよりも効率的に利用可能なアップデートを使用します。多肢選択推論を超えて、XSum および CNN/DailyMail での実験では、PARALLEL が完全適応によって達成された ROUGE-1 および ROUGE-2 スコアの 96.9 ~ 98.6\% と、対応する ROUGE-L スコアの 98.8 ~ 98.9\% を保持していることが示されています。同じ累積適応時間または GPU エネルギーで比較した場合、PARALLEL は、代表的な実行において完全適応よりも高い ARC 精度を達成し、より安定した後期段階の適応軌道を示します。これらの結果は、各サンプルをいつ、どの程度強く更新するかを学習することで、不必要な更新を回避しながら、安定的かつ効率的なデプロイ後のストリーム適応をサポートすることを示しています。

原文 (English)

PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits

Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training samples regardless of their local update benefit. We propose PARALLEL, a prefrontal-aligned reinforcement inspired approach for language-model learning. Inspired by the complementary roles of goal-related and uncertainty-related control, PARALLEL represents these forms of information as separate controller signals and combines them with the current model representation. A reinforcement-inspired controller assigns sample-dependent update intensity using immediate utility-cost feedback. PARALLEL therefore learns when and how strongly to adapt to each sample, prioritizing beneficial updates while limiting unnecessary parameter changes. PARALLEL uses available updates more efficiently than selective baselines while retaining 94.1--99.2\% of Full-adaptation performance. Beyond multiple-choice reasoning, experiments on XSum and CNN/DailyMail show that PARALLEL retains 96.9--98.6\% of the ROUGE-1 and ROUGE-2 scores achieved by Full adaptation and 98.8--98.9\% of the corresponding ROUGE-L scores. When compared at the same cumulative adaptation time or GPU energy, PARALLEL achieves higher ARC accuracy and exhibits a more stable late-stage adaptation trajectory than Full adaptation in the representative run. These results show that learning when and how strongly to update each sample supports stable and efficient post-deployment stream adaptation while avoiding unnecessary updates.

13:00 JSTLLM/生成AI画像/動画生成エージェント

裁定キャプション: 厳密なゼロショット画像キャプションのためのマルチエージェント アライメント スコアリングとコンセンサス蒸留ビーム アービトレーション

ゼロショット画像キャプション (ZIC) は、テキストのみのコーパスとフリーズされた事前トレーニングされた画像テキスト スコアラーに依存して、キャプショナーのトレーニング中に画像とキャプションのペアの監視なしで画像を記述します。既存の検索拡張手法は、画像とテキストの位置合わせを検索時に一度スコアリングし、その後、言語モデルの確率のみに基づいてキャプショナーの自己回帰ビームをコミットし、デコーダにそれ以上の視覚的グラウンディングフィードバックを与えずに残します。進歩は停滞しており、2024 年以降、厳密なレジームのベストを改善する方法はありません。私たちは、変更されていない IFCap キャプションに対して複数のチェックポイントでグラウンディング フィードバックを復元する、推論時のマルチエージェント フレームワークである Adjudicated Captioning を提案します。まず、入力に強力な凍結取得エンコーダーをインストールします。次に、取得とデコードの間に、上位 9 位の取得を上位 5 位に再ランク付けする凍結されたクロスアテンション ベリファイアを挿入します。 3 番目に、出力ビームに、多層パーセプトロンである TriFuse と、パイプラインの唯一の学習コンポーネントであるメモリ在席トランスフォーマーである MemAttend を組み合わせた学習済みリランカーを接続します。どちらも、対の画像キャプション ラベルや参照キャプションを使用せず、3 つの凍結スコアラーにわたる Borda コンセンサス蒸留によって自己監視されてトレーニングされます。帰納的ヘッドラインプロトコルの下では、リランカーが互いに素な COCO Karpathy 検証ビームに適合し、フリーズしてテストに適用されるため、フレームワークは COCO Karpathy で CIDEr 117.6 と SPICE 21.9 に達し、IFCap の 108.0 と 20.3 から +9.6 CIDEr ゲインとなり、NES を +7.7 上回りました。109.9 で最も強力な合成画像拡張手法です。キャプショナーを再トレーニングします。トレーニングなしの固定融合ベースラインは 115.8 CIDEr に達するため、+9.6 のゲインのうち +7.8 は学習されていないアーキテクチャ介入によるもので、残りの +1.8 は学習されたリランカーによるものです。同じレシピはキャプショナの再トレーニングなしでオフ COCO に転送されます。Flickr30k Karpathy で CIDEr +8.1、NoCaps 全体で +5.7。

原文 (English)

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods score image-text alignment once, at retrieval, then commit the captioner's autoregressive beam under language-model probability alone, leaving the decoder without further visual grounding feedback. Progress has stalled, with no method improving on the strict-regime best since 2024. We propose Adjudicated Captioning, an inference-time multi-agent framework that restores grounding feedback at multiple checkpoints over an unchanged IFCap captioner. First, we install a stronger frozen Retrieval Encoder at the input. Second, between retrieval and decoding we insert a frozen Cross-Attention Verifier that re-ranks the top-9 retrievals to top-5. Third, at the output beam we attach a learned Reranker pairing TriFuse, a multilayer perceptron, with MemAttend, a memory-attended transformer, the pipeline's only learned components; both are trained self-supervised by Borda-consensus distillation across the three frozen scorers, using no paired image-caption labels and no reference captions. Under the inductive headline protocol, with rerankers fit on the disjoint COCO Karpathy validation beam and applied frozen to test, the framework reaches CIDEr 117.6 and SPICE 21.9 on COCO Karpathy, up from 108.0 and 20.3 for IFCap, a +9.6 CIDEr gain, and +7.7 above NES, the strongest synthetic-image-augmented method at 109.9, without retraining the captioner. A training-free fixed-fusion baseline reaches 115.8 CIDEr, so +7.8 of the +9.6 gain comes from the non-learned architectural intervention and the remaining +1.8 from the learned rerankers. The same recipe transfers off-COCO without captioner retraining: +8.1 CIDEr on Flickr30k Karpathy and +5.7 on NoCaps overall.

13:00 JST画像/動画生成

Point2Radio: マテリアルを認識した点群からのクロスシーン無線フィールドの基礎モデル

通常、高忠実度の無線フィールドはシーンごとにシミュレートされます。つまり、送信機の構成が各シーンに個別に適合されるため、環境全体で共有される伝播構造を利用できません。複数の環境から伝達可能な伝播を事前に学習する基礎モデルである Point2Radio を紹介します。マテリアルを認識した点群と送信機 (TX) 設定が与えられると、一般的なエンコーダーは、任意の受信機 (RX) 位置でクエリできる TX 条件付きシーン表現を生成します。タスク固有のクエリ デコーダは、この表現をさまざまな無線量、たとえば 3 次元 (3D) パスゲイン (PG) フィールドやパワー角スペクトル (PAS) にマッピングします。新しいシーンの推論では、モデルはマテリアル認識の点群とトランシーバー クエリのみを使用し、メッシュや明示的なパス トレースを使用せずに単一の GPU でミリ秒単位で実行します。 86,272 の TX 条件付きフィールドを含む 337 シーンのコーパスのシーンに共通の分割で PG 予測を評価します。 Point2Radio は 0.871 dB の平均絶対誤差 (MAE) を達成し、同じ分割された UNet スタイルのベースラインと比較して誤差を 76.7% 削減します。同じエンコーダは、タスク固有のデコーダを介した PAS 予測もサポートします。実験ではさらに、光のターゲットシーンを微調整することで特定の環境への適応が改善されることが示されています。

原文 (English)

Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds

High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing to exploit propagation structures shared across environments. We present Point2Radio, a foundation model that learns a transferable propagation prior from multiple environments. Given a material-aware point cloud and a transmitter (TX) setting, a common encoder produces a TX-conditioned scene representation that can be queried at arbitrary receiver (RX) locations. Task-specific query decoders map this representation to different radio quantities, e.g., three-dimensional (3D) path-gain (PG) fields and power angular spectra (PAS). At inference for a new scene, the model uses only a material-aware point cloud and transceiver queries, running in milliseconds on a single GPU without meshes or explicit path tracing. We evaluate PG prediction on a scene-disjoint split of a 337-scene corpus containing 86,272 TX-conditioned fields. Point2Radio achieves 0.871 dB mean absolute error (MAE), reducing error by 76.7% relative to a same-split UNet-style baseline. The same encoder also supports PAS prediction via a task-specific decoder. Experiments further show that light target-scene fine-tuning improves adaptation to a specific environment.

13:00 JSTエージェントロボティクス

Auto-JEPA: エンドツーエンドの自動運転のための継続的意図の潜在的な世界モデル

既存の自動運転世界モデルは通常、将来のビデオ、占有状態、BEV の表現、またはエージェントの動きの緻密な予測を実行します。私たちは、計画は完全な未来世界を再構築する必要はなく、将来の自我の行動に影響を与えるシーンの特徴にのみ焦点を当てる必要があると主張します。この観点に基づいて、我々は、共同埋め込み予測を通じて継続的な将来の運転意図を学習するアクション指向の潜在世界モデルである Auto-JEPA を提案します。 Auto-JEPA は、視覚的な観察、エゴの動きの履歴、およびナビゲーション コマンドを考慮して、将来のエゴの軌跡の潜在的な表現と一致する意図の埋め込みを予測します。予測された意図は、固定軌道メモリから実行可能な軌道を取得し、シーン条件付き候補選択モジュールによってランク付けされます。 Auto-JEPA はビジュアル エンコーダーをフリーズしたままにし、明示的な知覚アノテーションを必要とせず、学習された軌道ジェネレーターを使用しません。 Auto-JEPA は、軌道表現、意図予測、候補選択に関してタスク固有のモジュールのみを最適化することで、NAVSIM v1 で 91.3 PDMS、NAVSIM v2 で 89.1 EPDMS を達成します。セマンティック オクルージョンの実験では、動的エージェント領域をマスキングすると、等面積ランダム マスキングの 2.97 倍の平均意図変化が誘発されることがわかりました。さらに、将来の運転に影響を与える車両を遮ると、予測された意図と選択された軌道が大幅に変化しますが、影響を及ぼさない車両が遮られた場合、どちらも本質的に変化しません。これらの結果は、将来の意図を予測することで、モデルが計画に関連する視覚的特徴に焦点を当てるようになり、密な未来世界モデリングを行わずに高品質の計画をサポートすることを示しています。

原文 (English)

Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion. We argue that planning need not reconstruct the complete future world, but only focus on scene features that affect future ego action. Based on this perspective, we propose Auto-JEPA, an action-oriented latent world model that learns continuous future driving intent through joint-embedding prediction. Given visual observations, egomotion history, and navigation commands, Auto-JEPA predicts an intent embedding aligned with the latent representation of the future ego trajectory. The predicted intent retrieves executable trajectories from a fixed trajectory memory, which are then ranked by a scene-conditioned candidate selection module. Auto-JEPA keeps the visual encoder frozen, requires no explicit perception annotations, and uses no learned trajectory generator. By optimizing only task-specific modules for trajectory representation, intent prediction, and candidate selection, Auto-JEPA achieves 91.3 PDMS on NAVSIM v1 and 89.1 EPDMS on NAVSIM v2. Semantic occlusion experiments show that masking dynamic-agent regions induces an average intent change 2.97x that of equal-area random masking. Moreover, occluding vehicles that affect future driving substantially changes the predicted intent and selected trajectory, whereas both remain essentially unchanged when non-influential vehicles are occluded. These results show that future-intent prediction encourages the model to focus on planning-relevant visual features and supports high-quality planning without dense future-world modeling.

13:00 JST研究/論文

スパース性バイアス分類器を使用しないガイダンスによる scDiffusion の改善

単一細胞 RNA シーケンス (scRNA-seq) は現代の細胞生物学において不可欠なツールとなっており、正確な合成 scRNA-seq データを生成することがますます重要になっています。拡散モデルは条件付き scRNA-seq 生成において有望な結果を達成していますが、分類子ガイダンスや分類子なしガイダンス (CFG) を含む既存のガイダンス戦略は、真の周辺分布に近似するように訓練された無条件分岐に依存しているため、実質的な遺伝子特異的構造が保持され、ガイダンスの有効性が制限される可能性があります。意図的に劣化させた参照を使用して拡散モデルを効果的にガイドできることを示した最近の研究に触発され、scRNA-seq 生成のためのスパースバイアス分類子フリー ガイダンス (SB-CFG) 戦略を提案します。 SB-CFG は、想定される「中立」周辺分布を近似するのではなく、無条件分岐に対して意図的に情報量が少ないスパース参照を導入し、粗いスパース統計のみを保持しながら遺伝子の同一性を削除します。この「悪い」参照により、条件付き予測と無条件予測の間のコントラストが強調され、サンプリング中により強力かつ効果的なガイダンスが得られます。私たちは、公開されている 5 つの scRNA-seq データセットに対するトレーニング不要のサンプリング変更として SB-CFG を評価しました。実験結果は、マーカー遺伝子発現の忠実度、細胞型の一貫性、および疎性の保存に関して、標準的な CFG ベースのサンプリングと比較して一貫した改善を示しており、SB-CFG が生物学的に意味のある遺伝子発現パターンをよりよく捕捉していることを示しています。

原文 (English)

Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance

Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq data is becoming increasingly important. Although diffusion models have achieved promising results in conditional scRNA-seq generation, existing guidance strategies, including classifier guidance and classifier-free guidance (CFG), rely on an unconditional branch trained to approximate the true marginal distribution, which may retain substantial gene-specific structure and limit guidance effectiveness. Inspired by recent work showing that diffusion models can be effectively guided using intentionally degraded references, we propose a sparsity-biased classifier-free guidance (SB-CFG) strategy for scRNA-seq generation. Rather than approximating the assumed "neutral" marginal distribution, SB-CFG introduces a deliberately under-informative sparse reference for the unconditional branch, removing gene identity while preserving only coarse sparsity statistics. This "bad" reference amplifies the contrast between conditional and unconditional predictions, leading to stronger and more effective guidance during sampling. We evaluated SB-CFG as a training-free sampling modification on five publicly available scRNA-seq datasets. Experimental results demonstrate consistent improvements over standard CFG-based sampling in terms of marker gene expression fidelity, cell-type consistency, and sparsity preservation, indicating that SB-CFG better captures biologically meaningful gene expression patterns.

13:00 JST研究/論文

ニューラル ネットワーク検証のための先読み補題の学習

最先端のニューラル ネットワーク検証器は、中核となる解決メカニズムとして分岐限定手順を使用します。先読み手順によって駆動されるニューラル ネットワーク検証のためのインプロセッシング フレームワークを導入します。このフレームワークの下では、先読みによって、不安定な ReLU のフェーズにわたって新しい補題が導出されます。これらの補題は含意グラフに収集され、探索空間をプルーニングしてブール カットを鮮明にするために使用されます。私たちは、Marabou と $\alpha$-$\beta$-CROWN という 2 つの最先端の検証ツールでフレームワークをインスタンス化し、両方のパフォーマンスが向上することを実証し、最大 34% 多くのインスタンスが満足できないことを証明しました。

原文 (English)

Learning Lookahead Lemmas for Neural Network Verification

State-of-the-art neural network verifiers use the branch-and-bound procedure as their core solving mechanism. We introduce an inprocessing framework for neural network verification driven by the lookahead procedure. Under this framework, lookahead derives new lemmas over the phases of unstable ReLUs, which are collected into an implication graph that is used to prune the search space and vivify boolean cuts. We instantiate the framework in two state-of-the-art verifiers, Marabou and $\alpha$-$\beta$-CROWN, and demonstrate that it improves performance in both, proving up to 34% more instances unsatisfiable.

13:00 JSTエージェントハードウェア/半導体

モンテカルロツリー検索によるマルチエージェントシステムの自律修復

複雑なタスクを解決するために、マルチエージェント システム (MAS) の導入が増えています。出力が不正確または満足できない場合、ユーザーはエージェントの軌跡を検査することでエージェントの間違いを手動で特定し (つまり、{\em Failure Attribution})、フィードバックを提供して出力を改善する必要があります (つまり、{\em Repair})。 MAS 障害の原因特定に関する最近の研究にもかかわらず、そのような間違いから回復するための自動化されたメカニズムはほとんど解明されていないままです。このギャップを埋めるために、MAS 修復をモンテカルロ ツリー検索 (MCTS) プロセスとして定式化する検索ベースのフレームワークである MARS を提案します。このフレームワークは、分類法による拡張評価による診断に基づく拡張を通じて、潜在的な修復の広大な空間をナビゲートします。完全なロールアウトによって完全なシミュレーションを評価する標準の MCTS とは異なり、MARS はトークンの消費を削減するために部分的なロールアウトを使用してエージェントの軌跡を評価します。さらに、4 種類のエージェント アーキテクチャと 4 つの LLM バックボーンにわたる 1,310 の再生可能なマルチエージェント障害軌跡を備えた大規模な MAS 修復ベンチマークである StateMAS を紹介します。 StateMAS の実験では、MARS が一貫して最先端の手法を上回っており、同等のトークン消費コストを維持しながら、すべての設定で 3.0\% から 12.1\% への絶対的な改善を達成していることが実証されています。このアブレーション研究では、これらのパフォーマンス向上を達成するには、分類法に基づく評価と診断に基づく拡張が重要であることがさらに確認されました。

原文 (English)

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide feedback to refine the outputs (i.e., {\em repair}). Despite some recent work in MAS failure attribution, automated mechanisms to recover from such mistakes remain largely unexplored. To bridge this gap, we propose MARS, a search-based framework that formulates MAS repair as a Monte Carlo Tree Search (MCTS) process and navigates the vast space of potential repairs via diagnosis-guided expansion with taxonomy-augmented evaluation. Unlike standard MCTS, which evaluates a complete simulation via full rollout, MARS evaluates the agent trajectory using partial rollout to reduce token consumption. Furthermore, we introduce StateMAS, a large-scale MAS repair benchmark with 1,310 replayable multi-agent failure trajectories spanning four types of agent architectures and four LLM backbones. Experiments on StateMAS demonstrate that MARS consistently outperforms state-of-the-art methods, achieving an absolute improvement from 3.0\% to 12.1\% across all settings, while maintaining a comparable token consumption cost. The ablation study further confirms that taxonomy-augmented evaluation and diagnosis-guided expansion are critical to achieving these performance gains.

13:00 JST研究/論文GPT / ChatGPT

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it rem…

13:00 JSTLLM/生成AI研究/論文

ごまかしのセマンティクス: 一般ドメインの最先端技術に対する法的詐欺検出のベンチマーク

欺瞞の検出は、法的手続き、法執行機関、およびオンライン セキュリティに重大な影響を及ぼします。人間の判断には精度と拡張性の点で限界がありますが、自然言語処理 (NLP) はデータ駆動型の代替手段を提供します。法的領域に焦点を当てた NLP ベースの自動欺瞞検出 (ADD) の調査と比較分析を紹介し、特徴ベースの機械学習から大規模言語モデル (LLM) アプローチへの進化をレビューします。私たちは 7 つのデータセット (2 つの法的領域、5 つの一般領域) にわたって統合された経験的評価を実施し、4 つのプロンプト戦略の下で 6 つの微調整された変圧器モデルと 7 つの LLM を比較します。その結果、データが豊富な一般ドメインでは微調整されたモデルが優れており、リソースが少ない法的環境でも少数ショットの LLM が競争力を維持しているため、ドメインの感度が高いことがわかります。思考連鎖によるプロンプトは、多くの場合、直接的な分類よりもパフォーマンスが劣ります。これらの調査結果は、一か八かの法的文脈におけるドメイン適応と解釈可能なシステムの必要性を浮き彫りにしています。

原文 (English)

Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art

Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural Language Processing (NLP) offers a data-driven alternative. We present a survey and comparative analysis of NLP-based Automatic Deception Detection (ADD) focusing on the legal domain, reviewing the evolution from feature-based machine learning to Large Language Model (LLM) approaches. We conduct a unified empirical evaluation across seven datasets (two legal, five general-domain), comparing six fine-tuned transformer models and seven LLMs under four prompting strategies. The results show strong domain sensitivity, with fine-tuned models excelling in data-rich general domains and few-shot LLMs remaining competitive in low-resource legal settings. Chain-of-Thought prompting often underperforms direct classification. These findings highlight the need for domain adaptation and interpretable systems in high-stakes legal contexts.

13:00 JST研究/論文

異種圧縮クライアントを使用した Federated Foundation モデルの微調整

基礎モデルのフェデレーション ラーニングは、リソースの非対称性という根本的な課題に直面しています。最も価値のあるドメイン固有のデータを保有している機関は、10 億パラメータのモデルをホストできないということです。既存の異種連携アプローチは、パラメーター効率の高いチューニング、モデルの枝刈り、または知識の蒸留を通じてこのギャップを埋めようとしますが、いずれも、フルモデルのメモリ削減、アーキテクチャの自己完結性、または表現の忠実性などの重要な特性をトレードオフにし、中核的な緊張が解決されないままになります。私たちは、異種圧縮クライアントを使用したフェデレーション微調整のためのパラメーター中心のフレームワークである FedSLM を提案します。 FedSLM は、SVD ベースの分解を使用して自己完結型のクライアント モデルを生成します。その低ランクの部分空間は、構造的に集約に互換性のある入れ子になった多様体を形成します。次に、圧縮グループ内の軽量アダプターを同期し、構造アライメントを介してグループ全体のフルランク再構築を融合する 2 段階のプロトコルを適用します。最後に、補助的な信頼性損失を伴う弱から強への導出ステップにより、集約された知識がフルスケールのサーバーに転送されますが、明示的なバイアスと分散のトレードオフにより圧縮アーティファクトが軽減されます。アダプターレベルの集約、グループ間融合の部分空間アライメント限界、および信頼損失が弱い監視ノイズをどのように軽減するかの特性評価に対する理論的保証を提供します。自然言語とビジョンに関する実験 - 言語ベンチマークでは、FedSLM が IID パーティションと非 IID パーティションの両方で既存のフェデレーション ベースラインよりも優れている一方、クライアント モデルは完全なモデルに必要な GPU メモリの約 50% で動作することが示されています。

原文 (English)

Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. Existing heterogeneous federated approaches attempt to bridge this gap through parameter-efficient tuning, model pruning, or knowledge distillation, yet each trades away a critical property, whether full-model memory reduction, architectural self-containedness, or representational fidelity, leaving the core tension unresolved. We propose FedSLM, a parameter-centric framework for federated fine-tuning with heterogeneous compressed clients. FedSLM uses SVD-based decomposition to produce self-contained client models, whose low-rank subspaces form nested manifolds that are structurally compatible for aggregation. It then applies a two-stage protocol that synchronizes lightweight adapters within compression groups and fuses full-rank reconstructions across groups via structural alignment. Finally, a weak-to-strong elicitation step with auxiliary confidence loss transfers the aggregated knowledge to the full-scale server, while an explicit bias--variance trade-off mitigates compression artifacts. We provide theoretical guarantees for adapter-level aggregation, subspace-alignment bounds for cross-group fusion, and a characterization of how the confidence loss mitigates weak-supervision noise. Experiments on natural language and vision--language benchmarks show that FedSLM outperforms existing federated baselines under both IID and non-IID partitions, while client models operate at roughly 50% of the GPU memory required by the full model.

13:00 JST研究/論文

metasignal: 包括的なメタ認知分析と意思決定のための Python パッケージ

Metasignal は、信号検出理論 (SDT) およびメタ認知測定用のオープンソース Python パッケージです。これは、参照変数 d' (知覚感度)、応答基準 c (応答バイアス)、および平均信頼度とともに、Rahnev (2025) によって評価された 17 のメタ認知尺度を実装します。 17 の測定値は、3 つのメタ d' ファミリー推定値、メタ d'、M 比、および M 差で構成されます。 4 つのノンパラメトリック タイプ 2 測定値、受信機動作特性曲線下のタイプ 2 面積 (AUC2)、ガンマ、ファイ、およびデルタ信頼度、およびそれらの 8 つの SDT 正規化比および差分形式。そして、メタノイズとメタ不確実性という 2 つのモデルベースの尺度です。 1 つの関数で、試験レベルの刺激、応答、および信頼度の配列から完全なセットを計算します。 「metasignal」は現在、各試行の刺激と反応が正確に 2 つのカテゴリでコード化されているバイナリ (2 つの選択肢) 識別タスクをサポートしています。このパッケージには、コマンドライン インターフェイス、グループ サマリー、ブートストラップ信頼区間、順列テスト、オプションの階層ベイジアン モデル、および情報理論的尺度も提供されます。 「メタシグナル」は、これらの指標を単一のプラットフォームに統合し、より広範なメタ認知研究と意思決定研究への採用を促進します。

原文 (English)

metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making

Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement. It implements the 17 metacognitive measures evaluated by Rahnev (2025), together with the reference variables d' (perceptual sensitivity), response criterion c (response bias), and mean confidence. The 17 measures comprise three meta-d' family estimates, meta-d', M-ratio, and M-difference; four nonparametric Type-2 measures, the Type-2 area under the receiver-operating-characteristic curve (AUC2), Gamma, Phi, and delta confidence, together with their eight SDT-normalized ratio and difference forms; and two model-based measures, meta-noise and meta-uncertainty. A single function computes the complete set from trial-level stimulus, response, and confidence arrays. `metasignal` currently supports binary (two-alternative) discrimination tasks, in which each trial's stimulus and response are coded with exactly two categories. The package also provides a command-line interface, group summaries, bootstrap confidence intervals, permutation tests, optional hierarchical Bayesian models, and information-theoretic measures. `metasignal` unifies these measures in a single platform to encourage broader metacognition research and adoption in decision-making studies.

13:00 JSTLLM/生成AI

DoubleHelix: LLM を使用したオーディオビジュアル音声認識のための構造化クロスモーダル フュージョン

視聴覚音声認識 (AVSR) は、音声モダリティと視覚モダリティの効果的な融合に依存していますが、既存のアプローチでは、クロスモーダル インタラクションを、構造化された反復改良を行わずに単一ステップの操作として扱います。我々は、適応的な劣化を意識した強化を備えた反復的なクロスモーダル相互作用プロセスとして融合を再定式化するマルチモーダル融合フレームワークである DoubleHelix を紹介します。このフレームワークは、学習されたアライメント制約とのマルチターン構造化相互作用のための ReverseParallelHelix、劣化を認識したゲート信号を学習するための QualitySensor、一貫性ガイドに基づく条件付き機能強化のための HelixReplication の 3 つのコンポーネントで構成されます。 LRS3 での実験では、DoubleHelix がクリーン オーディオで 0.68% の WER を達成し、一致したバックボーン設定の下でこれまでの最高の結果を 5.6% 相対的に改善して上回ることが実証されました。包括的なアブレーション研究により、非対称経路の重み付けなどの設計選択の的を絞った分析を含め、各コンポーネントの寄与が検証されます。このフレームワークは、評価されたバブルノイズ条件下で堅牢性が向上し、SNR -5dB で 11.6% の WER を達成したことを示しています。

原文 (English)

DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interaction as a single-step operation without structured iterative refinement. We present DoubleHelix, a multimodal fusion framework that reformulates fusion as an iterative cross-modal interaction process with adaptive degradation-aware enhancement. The framework comprises three components including ReverseParallelHelix for multi-turn structured interaction with learned alignment constraints, QualitySensor for learning degradation-aware gating signals, and HelixReplication for consistency-guided conditional feature enhancement. Experiments on LRS3 demonstrate that DoubleHelix achieves 0.68% WER on clean audio, outperforming previous best results by 5.6% relative improvement under matched backbone settings. Comprehensive ablation studies validate each component contribution, including targeted analysis of design choices such as asymmetric pathway weighting. The framework shows improved robustness under evaluated babble-noise conditions, achieving 11.6% WER at SNR -5dB.

13:00 JST研究/論文

Multi-Granularity Position Embedding of Graphs via Granular-Ball for Link Prediction

Link prediction aims to identify potential or future connections within a given graph structure. Position information is essential for link…

13:00 JST研究/論文

InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation

Recent work suggests that relational database management systems (RDBMSs) can execute quantum circuit simulation by compiling the simulatio…

13:00 JST研究/論文

HERO: 長期的な自己回帰ニューラル オペレーター向けの歴史を強化したロールアウト トレーニング

ニューラル演算子は、学習した進化演算子を自身の予測に再帰的に適用することで、時間依存偏微分方程式 (PDE) の高速代理を提供しますが、この自己回帰ロールアウトはすべての予測誤差を入力としてフィードバックするため、ローカル誤差が蓄積します。既存のロールアウト トレーニング戦略は、トレーニング入力と自己生成状態の間の不一致を軽減しますが、その監視では依然としてグラウンド トゥルースの軌道からの絶対的な不一致のみが測定されます。したがって、このような監視では、オペレーターが最適化中に以前に示した長期的な障害動作を克服したかどうかについては情報がありません。我々は、従来の絶対軌道監視をモデルの最適化履歴から導出された相対監視で強化する、履歴強化ロールアウト トレーニング (HERO) を提案します。 HERO は、定期的に更新されるラグ オペレーター、現在のモデル、およびロールアウト エラー、スペクトルの不一致、エネルギー ドリフト、エラーの増大による摂動入力から分離されたロールアウト候補をランク付けし、最も強い障害軌跡を参照として選択します。この参照は、固定比較ベースラインとしてマージンベースの目的に入力され、独立した勾配方向ではなく、グラウンドトゥルースロールアウト勾配の有界でサンプル依存の再重み付けを引き起こします。これを理論的にさらに分析します。スペクトルおよびアテンションベースのバックボーンを備えた 9 つの PDE ベンチマークの実験では、HERO が推論時間のコストをかけずに、長期的な精度、安定したロールアウト長、および分布外の堅牢性を一貫して向上させることが示されています。これらの結果は、歴史を強化した相対監視が長期自己回帰予測の安定化に有効であることを示しています。

原文 (English)

HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators

Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator recursively to its own predictions, but this autoregressive rollout feeds every prediction error back as input, so local errors accumulate. Existing rollout-training strategies reduce the mismatch between training inputs and self-generated states, yet their supervision still measures only the absolute discrepancy from the ground-truth trajectory. Such supervision is therefore uninformative about whether the operator has overcome the long-horizon failure behaviors it exhibited earlier during optimization. We propose history-enriched rollout training (HERO), which augments conventional absolute trajectory supervision with relative supervision derived from the model's optimization history. HERO ranks detached candidate rollouts from a periodically refreshed lagged operator, the current model, and a perturbed input by rollout error, spectral discrepancy, energy drift, and error growth, and selects the strongest failure trajectory as reference. This reference enters a margin-based objective as a fixed comparison baseline, inducing a bounded, sample-dependent reweighting of the ground-truth rollout gradient rather than an independent gradient direction, which we further analyze theoretically. Experiments on nine PDE benchmarks with spectral and attention-based backbones show that HERO consistently improves long-horizon accuracy, stable rollout length, and out-of-distribution robustness at no inference-time cost. These results indicate that history-enriched relative supervision is effective for stabilizing long-horizon autoregressive prediction.

13:00 JST画像/動画生成

あなたを見たことがありますか?埋め込み動作は合成顔データセットのメンバーシップをシグナルします

合成顔データセットは、生体認証におけるプライバシーの露出とデータ アクセスの制約を軽減するために使用されることが増えています。ただし、これらのデータセットを生成するジェネレーターは実際の顔でトレーニングされているため、合成データは依然として実際のソース データを明らかにする可能性があります。私たちは、まず顔認識装置のトレーニングに使用される合成データセットを特定し、次にジェネレーターのトレーニングに使用される実際のデータセットを推論する、データセット レベルのメンバーシップ推論攻撃を通じてこのリスクを研究します。 11 の顔認識モデル、11 の合成データセット、7 つの実際のデータセットにわたって、攻撃は 100% のケースで合成トレーニング データセットを復元し、54.5% のケースでジェネレーターのソース データセットを特定しました。これらの結果は、合成データが実際のトレーニング データのデータセット レベルのトレースを保持できること、およびプライバシーを保護する展開にはより強力な漏洩緩和が必要であることを示しています。

原文 (English)

Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk through a dataset-level membership inference attack that first identifies the synthetic dataset used to train a face recognizer and then infers the real dataset used to train the generator. Across 11 face recognition models, 11 synthetic datasets, and 7 real datasets, the attack recovers the synthetic training dataset in 100% of cases and identifies the generator's source dataset in 54.5% of cases. These results show that synthetic data can retain dataset-level traces of real training data and that privacy-preserving deployment requires stronger leakage mitigation.

13:00 JST研究/論文

暗黙的な機械学習の力場が分子動力学シミュレーションを高速化

陰的機械学習力場 (I-MLFF) を導入します。これは、ニューラル ネットワーク層の明示的なスタックを自己矛盾のない固定小数点方程式に置き換えます。分子シミュレーションでは、この定式化により中間表現を連続するタイムステップ間で再利用できるため、ウォームスタートの力評価が可能になります。結果として得られるモデルは、浅い単層 MLFF の計算フットプリントと、ディープ ニューラル ネットワークの表現能力および精度を効果的に組み合わせます。私たちのアプローチは、力の予測と軌道の統合を別々に考慮した場合にはアクセスできない、アーキテクチャに依存しない効率の向上を実現します。グラフ ニューラル ネットワークの 3 つの主要なクラス、つまり不変、等変デカルト テンソル、および SO(3) 等変球面テンソル アーキテクチャにわたってこれを実証します。それぞれのコンピューティングとメモリのフットプリントが 2 ~ 5 分の 1 に削減されます。重要なのは、これらのゲインは、完全な原子分解能と元の積分タイムステップを維持しながら、空間的または時間的な粗視化を回避しながら達成されることです。したがって、私たちの貢献は、量子力学的に忠実な分子シミュレーションのスケーリングフロンティアを前進させ、固定された GPU メモリと計算予算内でより長い軌道とより大規模な原子システムを可能にし、それによって生体分子および材料システムにわたる新しい洞察へのアクセスを開きます。

原文 (English)

Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent fixed-point equations. In molecular simulations, this formulation enables intermediate representations to be reused across successive timesteps, thereby warm-starting force evaluation. The resulting models effectively combine the computational footprint of a shallow, single-layer MLFF with the representational capacity and accuracy of a deep neural network. Our approach unlocks architecture-agnostic efficiency gains that are inaccessible when force prediction and trajectory integration are considered separately. We demonstrate this across three major classes of graph neural networks: invariant, equivariant Cartesian tensor, and SO(3)-equivariant spherical-tensor architectures. Each yields a two- to five-fold reduction in compute and memory footprint. Crucially, these gains are achieved while retaining full atomistic resolution and the original integration timestep, avoiding spatial or temporal coarse graining. Our contribution therefore advances the scaling frontier of quantum-mechanically faithful molecular simulation, enabling longer trajectories and larger atomistic systems within fixed GPU memory and compute budgets, and thereby opening access to new insights across biomolecular and material systems.

13:00 JSTLLM/生成AIエージェント

LLM エージェントのメモリ来歴ロンダリング: 永続メモリ用の非増幅ファイアウォール

長期記憶により、大規模言語モデル (LLM) エージェントは以前の設定とワークフローを再利用できますが、信頼できない観察結果も永続的なアクション コンテキストに変わります。私たちはメモリ来歴ロンダリングを特定します。LLM ベースのメモリ統合中に、外部観察が明らかなユーザー履歴またはワークフロー サポートとして書き換えられ、権限を制限するはずの信頼性の低いソースを消去しながらアクション トリガーを保持する可能性があります。既存のプロンプト フィルター、コンテンツ サニタイザー、およびツール ガードは、損失のあるメモリ統合後のソース権限の非増幅を強制しません。私たちはこの境界を形式化し、Provenance-Preserving Memory Fire Wall (PPMF) としてインスタンス化します。PPMF は、プラットフォームが管理する来歴を保存し、アクションのリスクとアクション関連のメモリの権限を照合することでツールの呼び出しを許可する軽量のメモリ ミドルウェアです。固定リスク ポリシーを使用したスキーマに基づいた評価では、脆弱な統合メモリは最大 1,000 の攻撃成功率 (ASR) に達します。プラットフォームが維持する出所、確認、およびリスクのラベルが損なわれていないため、評価された不正な高リスクのアクションは PPMF ゲートを通過しませんが、確認された無害なアクションと対象となる低リスクのメモリ使用は実行可能のままです。

原文 (English)

Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory

Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, an external observation may be rewritten as apparent user history or workflow support, preserving an action trigger while erasing the low-trust source that should limit its authority. Existing prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy memory consolidation. We formalize this boundary and instantiate it as Provenance-Preserving Memory Fire wall (PPMF), a lightweight memory middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. In our schema-grounded evaluation with fixed risk policies, vulnerable consolidated memories reach up to 1.000 attack success rate(ASR); with intact platform-maintained provenance, confirmation, and risk labels, no evaluated unauthorized high-risk action passes the PPMF gate while confirmed benign actions and targeted low-risk memory use remain executable.

13:00 JSTロボティクス

ActFovea: 時空間的なビジュアルアクションの一貫性による VLA ポリシーのランタイム保護

ビジョン言語アクション (VLA) ポリシーは、ロボット操作において優れたパフォーマンスを達成しますが、視覚的観察、ロボットの状態、および実行されたアクションの間の時間的整合性を壊す実行時の外乱に対して依然として脆弱です。基盤となる VLA ポリシーを再トレーニングしたり変更したりすることなく、そのような障害を検出して軽減するプラグアンドプレイの安全保護フレームワークである ActFovea を紹介します。 ActFovea は、ロボットの運動学、固有受容状態、最近のアクションを使用して、タスクに無関係な視覚コンテンツを抑制しながら、接触関連領域と予測動作コリドーを保持するアクション条件付き中心窩領域を構築します。視覚的な動きと観察の新鮮さが、幾何学的、固有受容的、およびアクションの遷移と一貫性を保っているかどうかを評価することにより、実行時のリスクを検出します。回復可能な外乱の場合、ActFovea は外乱固有の候補観測を構築し、結果のアクション チャンクを検証した後にのみ回復を受け入れます。古い監視や再実行された監視により信頼性の高い回復が不可能な場合は、制限付きの安全な障害手順が呼び出されます。複数の LIBERO スイートにわたる $\pi_0$ の閉ループ評価では、ActFovea はローカライズされたビジュアル オーバーレイの下での成功率を 49.3\% から 90.3\% に高め、クリーンなパフォーマンスとのギャップの 93.7\% を埋めました。クリーンタスクのパフォーマンスを維持しながら、アクションのドリフトと視覚的遅延の下での成功率がそれぞれ 7.0 パーセント ポイントと 9.8 パーセント ポイント向上します。凍結観察の再生では、ActFovea はすべてのトライアルでタイムリーに安全な失敗をトリガーし、保護されていない失敗はありません。これらの結果は、時空間的な視覚アクションの一貫性が、VLA ポリシーの実行時保護の効果的な基盤となることを示しています。

原文 (English)

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFovea, a plug-and-play safeguarding framework that detects and mitigates such failures without retraining or modifying the underlying VLA policy. ActFovea uses robot kinematics, proprioceptive states, and recent actions to construct action-conditioned foveated regions that retain contact-relevant areas and predicted motion corridors while suppressing task-irrelevant visual content. It detects runtime risks by evaluating whether visual motion and observation freshness remain consistent with geometric, proprioceptive, and action transitions. For recoverable disturbances, ActFovea constructs disturbance-specific candidate observations and accepts a recovery only after verifying the resulting action chunk. When stale or replayed observations make reliable recovery impossible, it invokes a bounded safe-failure procedure. In closed-loop evaluations of $\pi_0$ across multiple LIBERO suites, ActFovea increases success under localized visual overlays from 49.3\% to 90.3\%, closing 93.7\% of the gap to clean performance. It further improves success under action drift and visual delay by 7.0 and 9.8 percentage points, respectively, while preserving clean-task performance. Under frozen-observation replay, ActFovea triggers timely safe failure in all trials, with no unprotected failures. These results demonstrate that spatiotemporal visual-action consistency provides an effective basis for runtime safeguarding of VLA policies.

13:00 JSTLLM/生成AIロボティクスGemini

CLIFT: 非侵襲的な閉ループの反復微調整により、ジェミニ ロボティクス オンデバイスをヒューマノイドのスペシャリストに変える

ロボット基盤モデルの機能はますます高まっていますが、最も強力なモデルは通常、独自のデータでトレーニングされ、クローズドソースのままであるため、下流のユーザーが新しいタスク、実施形態、展開設定に適応する能力が制限されています。 LLM コミュニティに続いて、クローズドウェイト ロボット基礎モデルの新たなアクセス パラダイムは、マネージド教師あり微調整 (SFT) API です。この API では、ユーザーはモデルの重み、勾配、トレーニング内部にアクセスせずにトレーニング データを送信し、調整されたポリシーを受け取ります。このような API を使用すると、下流ユーザーは強力な独自の基盤モデルを活用できるようになりますが、ポリシーの改善は純粋な模倣に制限され、内部トレーニング信号に依存する強化学習やその他の閉ループ手法は排除されます。この制限は、新しい状態、アクション追跡ダイナミクス、遅延、およびコントローラー固有の障害モードにより、ポリシー出力と展開された動作の間のギャップが大きい、アジャイルで接触が多いヒューマノイド操作の場合に特に深刻です。私たちは、このマネージド API 体制がヒューマノイドへの適応にどれほど効果的であるか、またタスクの習得に向けてポリシーを推進するためにその中で閉ループの改善をどのように実現できるかを研究します。私たちは、Gemini Robotics On-Device (GROD) 上でインスタンス化された実際のヒューマノイドに対するマネージド API 適応に関する最初の実証研究の 1 つを実施します。 API を介した直接 SFT は、同じデモンストレーションでトレーニングされた主要なオープンウェイト VLA を大幅に上回っていますが、アジャイルでコンタクトの多いタスクに関する展開レベルの習熟にはまだ及ばないことがわかりました。このギャップを埋めるために、CLIFT: Closed-Loop Iterative Fine-Tuning を導入します。これは、デプロイメント時の報酬フィードバックを API 互換の教師ありデータに変換し、重み、勾配、尤度、損失にアクセスせずに閉ループ ポリシーの改善を可能にし、「モデル ボックスを開ける」ことなく、2 つのフライホイール サイクル後に GROD をほぼ完璧な成功に押し上げます。

原文 (English)

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings. Following the LLM community, an emerging access paradigm for closed-weight robot foundation models is the managed supervised fine-tuning (SFT) API, where users submit training data and receive a tuned policy without access to model weights, gradients, or training internals. While such APIs let downstream users leverage powerful proprietary foundation models, they restrict policy improvement to pure imitation, ruling out reinforcement learning and other closed-loop methods that rely on internal training signals. This limitation is particularly acute for agile, contact-rich humanoid manipulation, where the gap between policy outputs and deployed behavior is large due to novel states, action tracking dynamics, latency, and controller-specific failure modes. We study how effective this managed-API regime is for humanoid adaptation, and how closed-loop improvement can be realized within it to push policies toward task mastery. We conduct one of the first empirical studies of managed-API adaptation on a real humanoid, instantiated on Gemini Robotics On-Device (GROD). We find that direct SFT through the API substantially outperforms a leading open-weight VLA trained on the same demonstrations, yet still falls short of deployment-level mastery on agile, contact-rich tasks. To close this gap, we introduce CLIFT: Closed-Loop Iterative Fine-Tuning, which turns deployment-time reward feedback into API-compatible supervised data and enables closed-loop policy improvement without accessing weights, gradients, likelihoods, or losses-pushing GROD to near-perfect success after two flywheel cycles, all without "opening the model box."

13:00 JST研究/論文

MBDiff: 確率的ユーティリティ データ代入のためのマルチビュー行動認識拡散モデル

ユビキタスセンサーや組み込みデバイスによって収集されたユーティリティデータ (電気、水道、ガス消費量など) には、デバイスの故障やデータ送信の問題などのさまざまな要因により、多くの場合、大幅な欠損値が含まれています。データの欠落は、公共料金の請求の正確性に重大な影響を及ぼし、需要予測を妨げ、効率的な公共料金の供給管理を混乱させる可能性があります。その結果、公益事業データの補完は産業界と学術界の両方から大きな関心を集めています。多くの研究がこの問題に対処しようと試みてきましたが、そのほとんどはトレーニングのために集約されたデータセットに依存しており、より正確な代入のための貴重な洞察を提供する可能性のある豊富なユーザー行動情報を見落としています。ただし、長期的で多様かつ不完全な公共事業データから包括的なユーザー行動を学習することは依然として大きな課題です。さらに、相関関係の間接的な性質により、ユーザーの行動情報を活用して代入をガイドすることは自明ではありません。これらの課題に対処するために、確率的ユーティリティ データ代入のためのマルチビュー行動認識拡散モデルである MBDiff を提案します。 MBDiff には、次の 2 つの主要な技術コンポーネントが組み込まれています。(i) グローバル、ローカル、インスタンス レベルのビューを含む、複数の観点から包括的なユーザー行動を学習するマルチビュー ユーザー行動抽出モジュール。 (ii) 計算効率の高い方法でユーティリティ データを代入するための、基準選択モジュールと条件付き注意付きノイズ除去ネットワークで構成される行動認識型条件付き拡散モデル。私たちは、フロリダ最大の地方公共団体プロバイダーの 1 つと協力して MBDiff を実装し、評価します。実験結果は、私たちが提案した MBDiff が最先端のベースラインを効果的に上回っていることを示しています。たとえば、ブロック欠落補完の電気使用量データセットと水使用量デ​​ータセットでそれぞれ 7.04% と 29.1% 改善します。

原文 (English)

MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation

Utility data (e.g., electricity, water, and gas consumption), collected by ubiquitous sensors and embedded devices, often contains substantial missing values due to various factors such as device failures and data transmission issues. The data missingness can severely impact utility billing accuracy, hinder demand forecasting, and disrupt efficient utility supply management. As a result, utility data imputation has attracted much interest from both industry and academia. While many studies have attempted to address this issue, most of them rely on aggregated datasets for training, overlooking rich user behavior information, which could provide valuable insights for more accurate imputation. However, learning comprehensive user behavior from long-term, diverse, and incomplete utility data remains a significant challenge. Moreover, leveraging user behavior information to guide imputation is nontrivial due to the indirect nature of the correlations. To address these challenges, we propose MBDiff, a Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation. MBDiff incorporates two key technical components: (i) a multi-view User Behavior Extraction module that learns comprehensive user behavior from multiple perspectives, including global, local, and instance-level views; and (ii) a behavior-aware conditional diffusion model consisting of a reference selection module and a conditional attentional denoising network to impute utility data in a computationally efficient manner. We implement and evaluate MBDiff by collaborating with one of the largest municipal utility providers in Florida. Experimental results demonstrate our proposed MBDiff effectively outperforms state-of-the-art baselines, e.g., it improves 7.04% and 29.1% on the electricity and water usage datasets for block missingness imputation, respectively.

13:00 JST画像/動画生成

MoRAE: テキストからモーションを生成するためのフローフレンドリーな自己教師ありラテント

テキストからモーションへの生成では、意味的に正しく、時間的に一貫性があり、物理的に妥当なモーションを生成する必要があります。自然なアプローチは、最初にモーション データを構造化された意味空間に投影し、次にその空間内で生成モデルをトレーニングすることです。このようなパラダイムは、表現オートエンコーダー (RAE) による画像生成で大きな成功を収めています。RAE では、凍結された自己教師ありエンコーダーが、拡散モデルやフロー モデルの学習対象となるセマンティック機能を提供します。ただし、Motion-JEPA をフリーズされたエンコーダとして使用して、そのようなパラダイムをモーション空間に直接転送すると、大幅に失敗します。我々はこの障害を幾何学的に診断し、モーション固有の 2 つのボトルネックを特定します。(1) JEPA 特徴空間はスペクトル的に条件が悪く、ガウスからデータへの転送が不安定になります。 (2) スペクトルが適切に調整されている場合でも、フロー残差はデコーダに敏感な方向に一致する傾向があり、小さな潜在的なエラーがデコード後に大きなモーション アーティファクトに増幅されます。これらの洞察に基づいて、私たちは MoRAE を提案します。 MoRAE は 2 つのボトルネックに個別に対処します。コンパクトなボトルネックにより、構造化された JEPA 表現が抽出され、同時に弱く冗長な方向が除去され、潜在スペクトルが輸送安定領域に導入されます。次に、モーション結合トレーニングにより、保持された潜在ジオメトリがデコーダと位置合わせされ、デコード後の特性フロー エラーのコストが軽減されます。このフローに優しい潜在力により、標準の非自己回帰フローマッチング DiT は最先端のパフォーマンスを実現します。

原文 (English)

MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to first project motion data into a structured semantic space and then train a generative model within that space. Such a paradigm has been highly successful in image generation through Representation Autoencoders (RAEs), where a frozen self-supervised encoder provides semantic features for diffusion or flow models to learn from. However, direct transfer of such a paradigm to motion space using Motion-JEPA as the frozen encoder fails dramatically. We diagnose this failure geometrically and identify two motion-specific bottlenecks: (1) the JEPA feature space is spectrally ill-conditioned, making the Gaussian-to-data transport unstable; and (2) even with a well-conditioned spectrum, flow residuals tend to align with decoder-sensitive directions, where small latent errors are amplified into large motion artifacts after decoding. Based on these insights, we propose MoRAE. MoRAE addresses the two bottlenecks separately. A compact bottleneck distills the structured JEPA representation while removing weak and redundant directions, bringing the latent spectrum into a transport-stable regime. Motion-coupled training then aligns the retained latent geometry with the decoder, making characteristic flow errors less costly after decoding. With this flow-friendly latent, a standard non-autoregressive Flow-Matching DiT achieves state-of-the-art performance.

13:00 JST画像/動画生成エージェント

SERUM: ユーザーモデリングのための状態の抽出と改良

プロアクティブでパーソナライズされた対話が可能なエージェント アシスタントには、ユーザーの意図とワークフローの構造化モデルが必要です。ただし、生の構造化されていない画面アクティビティからこれらのモデルを構築することは、未解決の課題のままです。階層的な VLM アノテーションを使用して非構造化自己中心ビデオから有限状態の動作モデルを直接抽出するマルチパス フレームワークである SERUM を紹介します。 SERUM は、スライディング ウィンドウを通じて画面記録を処理し、アクティビティ認識パスと意図推論パスを交互に行い、各パスで蓄積された以前のコンテキストを使用してラベルを洗練し、シングルパス アノテーションで見られる幻覚や時間的混乱を軽減します。同義の状態は、文の埋め込みと人間が調整したしきい値を介して、コンパクトで一貫した分類法に統合されます。結果として得られるラベル シーケンス (アクションと意図の両方) に対して一次マルコフ モデルをフィッティングし、周波数ベースラインに対する予測精度を測定することで、行動構造を評価します。 4 つの領域 (コーディング、料理、身体活動、日常生活) の 61 本の自己中心的なビデオ全体で、次のことがわかりました。(1) 反復的なラベルの洗練は、数回のパスを経て、安定した状態のボキャブラリー (図式的平衡と呼ばれます) に収束します。 (2) 正規化されたマルコフ モデルは、周波数ベースラインよりも大幅に低い混乱と高いアクション予測を達成し、コーディングなどの構造化されたタスクで最大の利益をもたらします。 (3) ヒューマン アノテーターは、最終パスのラベルが正確であり、初回パスのラベルよりも大幅に改善されていると評価します。私たちの知る限り、SERUM は、構造化されていない自己中心的な画面ビデオから手動の注釈なしで解釈可能なプロセス モデルを生成する最初のシステムであり、実際のユーザー モデリングと行動理解のためのスケーラブルな道を開きます。私たちのデモ、コード、結果は公開されています

原文 (English)

SERUM: State Extraction and Refinement for User Modeling

Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building these models from raw, unstructured screen activity remains an open challenge. We present SERUM, a multi-pass framework that extracts finite-state behavioral models directly from unstructured egocentric video using hierarchical VLM annotation. Processing screen recordings through a sliding window, SERUM alternates between activity-recognition and intent-inference passes, with each pass refining labels using accumulated prior context to reduce hallucination and temporal conflation seen in single-pass annotation. Synonymous states are then merged via sentence embeddings and human-calibrated thresholds into a compact, coherent taxonomy. We evaluate behavioral structure by fitting first-order Markov models over the resulting label sequences (both actions and intents) and measuring predictive accuracy against frequency baselines. Across 61 egocentric videos in four domains (coding, cooking, physical activities, and daily life), we find: (1) iterative label refinement converges to a stable state vocabulary, which we term schematic equilibrium, after several passes; (2) normalized Markov models achieve substantially lower perplexity and higher action predictions than frequency baselines, with the largest gains on structured tasks like coding; and (3) human annotators rate final-pass labels as accurate and meaningfully improved over first-pass labels. To our knowledge, SERUM is the first system to produce interpretable process models from unstructured egocentric screen video without manual annotation, opening a scalable pathway for user modeling and behavioral understanding in the wild. Our demo, code, and results are publicly available

13:00 JST研究/論文

SAF-OPD: ポリシーに沿った蒸留のための安定した利点の融合

検証可能な報酬を伴う強化学習 (RLVR) は、単一の応答レベルの報酬をすべてのトークンにブロードキャストします。一方、オンポリシー蒸留 (OPD) は、より強力な教師に対して各トークンをスコア付けして密度の高い優位性を実現しますが、パフォーマンスは教師の品質に制限され、それを超えた探索を妨げます。それらの相補性により、RLVR と OPD の組み合わせは有望ですが、固定係数を使用して 2 つの利点を融合すると、2 つの誤ったキャリブレーションによるエントロピー崩壊が引き起こされることがわかりました。1 つはトークンレベルの OPD の利点が制限された RLVR の利点をはるかに超えて急増し、その信号が消去される可能性がある大きさの不一致です。もう 1 つは、持続的な最大強度の OPD が生徒を教師に引き寄せ続け、それを超えるために必要な探索を制限する時間的な不一致です。我々は、OPD の利点のみに適用される軽量の 4 ステージのパイプラインを介して両方の問題を解決する Stable Advantage Fusion フレームワークである SAF を提案します。これは、マグニチュード制御のためのスパース化してから圧縮するメカニズムと、時間制御のためのウォームアップしてからアニールするメカニズムを組み合わせたもので、各ステージは独立して切り替え可能であり、追加されるオーバーヘッドは無視できます。 GRPO を使用して RLVR をインスタンス化し、Qwen3-1.7B/4B/8B を使用して 7 つの数理推論とコード生成ベンチマークにわたって SAF を評価しました。SAF はエントロピー崩壊を回避し、固定係数の GRPO+OPD 融合より一貫して優れたパフォーマンスを示し、より安定したトレーニングを達成しながら、6 つのモデル ドメイン設定すべてで合計スコアを 0.51 ~ 2.70% 改善しました。

原文 (English)

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages with a fixed coefficient triggers entropy collapse from two miscalibrations: a magnitude mismatch, where token-level OPD advantages can spike far beyond the bounded RLVR advantage and erase its signal, and a temporal mismatch, where sustained full-strength OPD keeps pulling the student toward the teacher and limits exploration needed to surpass it. We propose SAF, a Stable Advantage Fusion framework that resolves both issues via a lightweight, four-stage pipeline applied only to the OPD advantage: a sparsify-then-compress mechanism for magnitude control paired with a warm-up-then-anneal mechanism for temporal control, with each stage independently switchable and adding negligible overhead. Instantiating RLVR with GRPO, we evaluate SAF across seven mathematical reasoning and code generation benchmarks with Qwen3-1.7B/4B/8B: SAF avoids entropy collapse and consistently outperforms fixed-coefficient GRPO+OPD fusion, improving the aggregate score by 0.51-2.70% across all six model-domain settings while achieving more stable training.

13:00 JST研究/論文Claude

MOSAIC: 安全な AI 計算のマスクされたアウトソーシング

私たちは、クライアントが入力とモデルの両方を保持し、サーバーがどちらも学習する必要がないという設定で、信頼できるが計算能力が弱いクライアントから、信頼できないが強力なサーバーに AI 計算を安全かつ効率的にアウトソーシングするという課題に取り組みます。我々が紹介する MOSAIC は、その核となる新しい行列乗算マスキング プロトコルであり、以前の研究よりもはるかに大きな行列にスケールし、大規模なトランス推論などの最新のワークロードの安全なアウトソーシングを可能にします。 MOSAIC は、乗算結果に少量のノイズを導入し、正確性を緩和することにより、最適な漸近クライアント オーバーヘッドと具体的な実行時間を従来の作業よりも桁違いに高速化します。そのセキュリティは、決定的な LWE および LPN の仮定に限定されます。このノイズは変圧器の多くの層にわたって蓄積されるため、主な技術的課題は誤差の増加を制限することです。 MOSAIC は、ランダムなアダマール回転に基づく誤差スケーリング メカニズムを使用してこの問題に対処します。大型の 70B トランス モデルでは、MOSAIC の複雑さは一般的な量子化アプローチに匹敵し、HumanEval での完全精度の BF16 推論にも匹敵します。最後に、MOSAIC のようなアイデアが現代のデータセンターにおける大規模な機密 AI への道をどのように約束できるかを示すエンドツーエンドの実装を紹介します。非機密推論は、RDMA のようなネットワーキングを使用してアクティベーション、キャッシュされた KV 値、および重みをノード間で移動することで、異種ハードウェアの利用を最大化するためにフェーズ (プリフィル/デコード)、レイヤー、および時間を超えてすでに分散されています。 MOSAIC は、トラステッド コンピューティング ベース (TCB) を小規模に保ち、AI 計算の大部分を信頼できないアクセラレータにアウトソーシングすることで、機密コンピューティングのスケーリングを可能にします。

原文 (English)

MOSAIC: Masked Outsourcing of Secure AI Computations

We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the server must learn neither. We present MOSAIC, whose core is a novel matrix-multiplication masking protocol that scales to far larger matrices than prior work, enabling the safe outsourcing of modern workloads such as large transformer inference. By introducing small amounts of noise to the multiplication result and thereby relaxing correctness, MOSAIC achieves optimal asymptotic client overhead and concrete runtimes orders of magnitude faster than prior work. Its security reduces to the decisional LWE and LPN assumptions. Because this noise accumulates across the many layers of a transformer, a key technical challenge is bounding error growth; MOSAIC addresses this with an error-scaling mechanism based on random Hadamard rotations. On large 70B transformer models, MOSAIC's perplexity is comparable to popular quantization approaches and even matches full-precision BF16 inference on HumanEval. Finally, we present an end-to-end implementation showing how ideas like MOSAIC can promise a path towards large-scale confidential AI in modern data centers. Non-confidential inference is already distributed across phase (prefill/decode), layer, and time to maximize utilization of heterogeneous hardware, using RDMA-like networking to move activations, cached KV values, and weights across nodes. MOSAIC enables scaling of confidential compute by keeping the trusted computing base (TCB) small and outsourcing the bulk of the AI computation to untrusted accelerators.

13:00 JST研究/論文

Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution

Swarm and evolutionary algorithms are usually analyzed as complete procedural systems in which nonlinear selection, replacement, and adapta…

13:00 JSTロボティクス

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliabili…

13:00 JSTLLM/生成AI

Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters

InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's w…

13:00 JST画像/動画生成

When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration

In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence…

13:00 JSTLLM/生成AIエージェント

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-stra…

13:00 JST画像/動画生成

TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation

Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthe…

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing…

13:00 JST画像/動画生成

OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation

Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations strug…

13:00 JSTLLM/生成AIGPT / ChatGPTDeepSeek

Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation

Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by…

13:00 JST研究/論文GPT / ChatGPT

The persuasive power of large language models does not depend on their perceived national origin

Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rej…

13:00 JST画像/動画生成ハードウェア/半導体

DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of a…

13:00 JSTエージェント

SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery

Modern neuroscience relies on integrating multi-scale, multimodal datasets to uncover the neural principles underlying intelligence. Howeve…

13:00 JST研究/論文

Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning

With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.g., c…

13:00 JSTLLM/生成AI

Cross-Lingual Transfer for Machine Translation in Turkic Languages

Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains in…

13:00 JST研究/論文

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and aud…

13:00 JST画像/動画生成

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based…

13:00 JST研究/論文

Explore Beyond the Boundary Using Entropic Information

In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback availab…

13:00 JSTエージェント

AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair

Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Rece…

13:00 JST画像/動画生成

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual questio…

13:00 JST研究/論文

TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion

Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predi…

13:00 JST研究/論文

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extend…

13:00 JSTエージェント

From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale

AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code rev…

13:00 JST研究/論文

TerraNova: A Foundation Model for the Anthropocene

A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representat…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文GemmaLlamaMistral AI

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While pri…

13:00 JST研究/論文

MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models

Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in scientific modeling. Whi…

13:00 JST画像/動画生成ハードウェア/半導体

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and ap…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two peop…

13:00 JST画像/動画生成

A Human-Centered Validation of the Explainability-Performance Coefficient

The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligenc…

13:00 JSTエージェントロボティクス

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to lang…

13:00 JST研究/論文

CENDRe: Concept Extraction with Natural Domain Representations

Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires unde…

13:00 JST研究/論文

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic fee…

13:00 JST研究/論文

SATViz: Real-Time Visualization of Clausal Proofs

Visual layouts of graphs representing SAT instances can highlight the community structure of SAT instances. The community structure of SAT…

13:00 JSTLLM/生成AIロボティクス

Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases

We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-ass…

13:00 JSTLLM/生成AI

Shall We Play a Game? Language Models for Open-ended Wargames

LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may…

13:00 JSTエージェント

Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning

The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled…

13:00 JSTエージェント

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost univer…

13:00 JSTエージェントビジネス/資金調達

M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities

As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to i…

13:00 JSTエージェントロボティクス

RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving

Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bott…

13:00 JST研究/論文

Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning

Explainable AI is increasingly important to scientific discovery. However, existing methods largely ignore that explanation quality is not…

13:00 JSTLLM/生成AIエージェント

What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents

Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion…

13:00 JSTエージェント

PEMAND: Persona-Enriched Multi-Agent Negotiation for Household Decision-Making

Modeling household-level decisions is central to many real-world applications, including trip planning, residential mobility and migration,…

13:00 JSTエージェント研究/論文

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE…

13:00 JSTLLM/生成AI

Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling

Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-tim…

13:00 JSTLLM/生成AIエージェント

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization

A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been em…

13:00 JSTLLM/生成AIエージェント

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models

Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from ex…

13:00 JSTLLM/生成AIエージェント

FEA-AI ハイブリッド アプローチによる IPMSM 設計最適化のためのマルチエージェント システム

内部永久磁石同期モーター (IPMSM) の設計では、相反する目的とマルチフィジックス制約のバランスをとる必要がありますが、最新の最適化ワークフローは 3 つのボトルネックに直面しています。それは、手動による問題設定、高い有限要素解析 (FEA) コスト、まばらな領域または分布外領域での信頼性の低いサロゲート ベースの検索です。これらの制限に対処するために、構造化された問題定義のための検索拡張生成 (RAG) と不確実性を考慮した FEA-AI ハイブリッド最適化パイプラインを統合する、エンドツーエンドの自動 IPMSM 設計最適化フレームワークを提案します。 RAG を通じてモーターの教科書に接続された設計エージェントは、ドメイン知識ベースのオプションとエンジニアリングのヒントを提供し、AI モデルのトレーニングのための最適化カードと実験計画計画を作成します。トレーニング エージェントは、電磁 FEA を自動化し、ジオメトリ検証とソルバー障害ログを記録し、ANOVA ベースのデータ分析と LLM 推論を使用して障害のあるジオメトリを分析し、設計サンプリング エージェントを呼び出して設計空間を再定義し、追加のサンプルを生成します。最適化エージェントは、不確実性主導のスイッチングを使用して GA ベースの検索を実行します。不確実性の低い候補は AI サロゲート推論によって評価されますが、不確実性が高く信頼性が重要なパレート フロントまたはトップ K の候補は高忠実度 FEA によって修正され、反復再トレーニングに再利用されます。このフレームワークは、経験に依存した手動の構成を、計算コストと予測の信頼性のバランスをとる再現可能なワークフローに変換します。一致した高忠実度 FEA 予算の下での実験結果は、提案されたハイブリッド アプローチが、低く、さらに削減可能な予測不確実性を維持しながら、より優れた目標パフォーマンスを達成し、早期の予算枯渇によって制限される FEA のみの探索や、信頼性の低い最適値に収束する AI のみの探索よりも優れたパフォーマンスを達成することを示しています。

原文 (English)

A Multi-Agent System for Motor Design Optimization via an FEA-AI Hybrid Approach

This study presents a large language model (LLM)-based multi-agent framework for interior permanent magnet synchronous motor (IPMSM) design optimization that mitigates limitations of conventional workflows: expertise-dependent problem setup and data preparation, the prohibitive computational cost of finite element analysis (FEA), and the unreliability of AI surrogates in unexplored regions. To this end, we first introduce a Design agent that formulates the optimization problem in natural language, leveraging retrieval-augmented generation to improve answer accuracy on motor design problems from below 50% to 67-80%. Furthermore, a Training agent autonomously repairs improperly defined design spaces by reasoning over solver failure history, raising the success ratio of the geometry sampling from 28% to 84% for AI training. Additionally, to resolve cost and reliability simultaneously, an Optimization agent employs an uncertainty-aware FEA-AI hybrid model: the AI surrogate is the primary evaluator, and FEA is selectively invoked where predictive uncertainty is high. Under the same FEA budget, this hybrid model achieves up to 44% lower iron loss in single-objective and 22.5% higher hypervolume in multi-objective optimization than conventional FEA-only search. Under the same evaluation budget, it reduces computation time by 52-55% while retaining 90-92% of FEA-only hypervolume. Conversely, AI-only search converges to false optima, leaving half its Pareto designs infeasible. Notably, a controller agent adaptively updates the uncertainty threshold that triggers FEA each round, eliminating manual tuning and achieving 5.8% lower single objective iron loss than with a fixed threshold. These results establish domain specialized LLM agents with uncertainty-aware hybrid evaluation as a reliable, scalable paradigm for simulation-driven design automation.

13:00 JSTLLM/生成AIエージェント

ロールエージェント: デュアルロール進化による LLM エージェントのブートストラップ

大規模言語モデル (LLM) エージェントは複雑なタスクで優れたパフォーマンスを示していますが、その学習は非効率なインタラクション フィードバックや静的トレーニング環境によって制限されることが多く、広範な一般化が妨げられます。これらの制限に対処するために、このホワイトペーパーでは、単一の LLM を利用してエージェントと環境の両方として同時に機能し、ブートストラップ型の共進化を可能にする、Role-Agent、\textcolor{black}{フレームワーク} を紹介します。ロール エージェントは、ワールド イン エージェント (WIA) とエージェント イン ワールド (AIW) の 2 つの相乗コンポーネントで構成されます。 WIA では、LLM がエージェントとして機能し、各アクションの後の将来の状態を予測します。予測された状態と実際の状態の調整はプロセスの報酬として使用され、環境を意識した推論を促進します。 AIW では、LLM が失敗した軌跡から失敗モードを分析し、同様の失敗パターンを持つタスクを取得します。これにより、目標を絞った実践のためにトレーニング データの分布が再形成されます。複数のベンチマークの実験では、Role-Agent が一貫してパフォーマンスを向上させ、強力なベースラインに対して平均 4\% 以上の向上をもたらしていることが示されています。

原文 (English)

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static training environments, which hinder broader generalization. To address these limitations, this paper introduces Role-Agent, \textcolor{black}{a framework} that harnesses a single LLM to function concurrently as both the agent and the environment, enabling a bootstrapped co-evolution. Role-Agent comprises two synergistic components: World-In-Agent (WIA) and Agent-In-World (AIW). In WIA, the LLM acts as the agent and predicts future states after each action; the alignment between predicted and actual states is then used as a process reward, encouraging environment-aware reasoning. In AIW, the LLM analyzes failure modes from failed trajectories and retrieves tasks with similar failure patterns, thereby reshaping the training data distribution for targeted practice. Experiments on multiple benchmarks show that Role-Agent consistently improves performance, yielding an average gain of over 4\% over strong baselines.

13:00 JSTLLM/生成AI

ReSum: LLM 推論と要約と強化学習の相乗効果

検証可能な報酬による強化学習 (RLVR) は、大規模言語モデル (LLM) における長期的な推論を改善するための中心的な手法です。ただし、既存の RLVR 手法では、推論の展開が不必要に長くなり、推論の一貫性が低下し、利用可能なコンテキスト バジェットが使い果たされる可能性があります。ロングコンテキストの組織化に対する既存のアプローチは、多くの場合、モデルが独自の推論軌道を管理できるようにするのではなく、ロールアウトを組織化する外部メカニズムに依存しています。この制限に対処するために、LLM が自己要約を通じて推論の軌跡を圧縮して整理できるようにする新しい RLVR フレームワークである ReSum を提案します。私たちのパイロット研究では、自己要約がトークンレベルのエントロピーを下げることで生成を安定化させ、「要約」フレーズを導入することで不正なロールアウトプレフィックスから伝播するエラーを大幅に軽減できることが示されています。これらの発見に動機づけられて、ReSum は、自己要約が進行中の推論プロセスに利益をもたらすかどうかを対照的に評価する、要約を意識した適応型ロールアウト メカニズムを採用しています。具体的には、モデルが自発的に自己要約をトリガーすると、ReSum は要約フレーズをマスクして対照的な分岐を作成します。非要約位置の場合は、代わりにフレーズをランダムに挿入して、一致したブランチを作成します。さらに、要約を意識した利点を設計して、対照的なロールアウト軌跡間のより詳細な比較を可能にします。広範な実験により、ReSum はロールアウトの長さを 18.6\% 短縮しながら、パフォーマンスを平均 4\% 向上させることが示されました。

原文 (English)

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning

Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods often encourage unnecessarily long reasoning rollouts, which can degrade reasoning coherence and exhaust the available context budget. Existing approaches to long-context organization often depend on external mechanisms to organize rollouts, rather than enabling the model to manage its own reasoning trajectory. To address this limitation, we propose ReSum, a novel RLVR framework that enables LLMs to compress and organize their reasoning trajectories through self-summarization. Our pilot studies show that self-summarization stabilizes generation by lowering token-level entropy, and that introducing a ``summarization'' phrase can substantially mitigate errors propagated from an incorrect rollout prefix. Motivated by these findings, ReSum adopts a summarization-aware adaptive rollout mechanism that contrastively evaluates whether self-summarization benefits the ongoing reasoning process. Specifically, when the model spontaneously triggers self-summarization, ReSum masks the summarization phrase to create a contrastive branch; for non-summarization positions, it instead randomly injects the phrase to create a matched branch. We further design a summarization-aware advantage to enable finer-grained comparison between contrastive rollout trajectories. Extensive experiments show that ReSum improves performance at an average of 4\% while reducing rollout length by 18.6\%.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達GPT / ChatGPTLlama

プロセスレベルの社会的影響評価のための認知世界モデル

社会的影響ダイアログは、内部の認知状態を変えることでユーザーの行動を変えます。評価の中心となる質問は、ユーザーの信念、欲望、意図、感情が会話の過程で測定可能なほど変化するかどうかであり、これは表面レベルのテキスト指標 (BLEU/ROUGE) や単一スコアの LLM 判定では捉えることができないプロセス指向の基準です。我々は \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)} を提案します。これは、マルチターン対話評価を「ユーザーが何を言ったか」から「ユーザーの内部認知状態がどのように進化したか」に再構成する LLM ベースのユーザー モデルです。CogWM は、BDI/E 認知状態とユーザー発話を共同で予測し、3 層を使用してユーザー シミュレーターと評価プラットフォームの両方として機能します。ターンレベルの忠実度、軌道レベルの状態ダイナミクス、タスクレベルの複合スコアリングをカバーする評価フレームワーク。 4 つの社会的影響シナリオにわたる 150,454 のユーザー ターン サンプルで \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} アノテーション パイプラインを介してトレーニングされた CogWM は、77.6\% の感情精度 (GPT-5.5 の 2.1$\time$) を達成しました。 3,600 件のマルチエージェント識別試験において、認知的影響力によって 6 つの営利エージェントを区別し、Llama-4-Scout が 1 位にランクされました (CTS +0.233)。 CogWM は、社会的影響対話の評価を最終的な判断からプロセスの追跡に移行します。コード\脚注{\scriptsize コード: https://github.com/lucianma05-create/CogWM} とモデル\脚注{モデル: https://www.modelscope.cn/models/LucianMa/CogWM-14B} をリリースしました。

原文 (English)

Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents

As LLM-based conversational agents advance toward increasingly open-ended and interaction-intensive scenarios, task completion alone provides an incomplete assessment of their effectiveness. The evolution of users' internal states, including beliefs, desires, intentions, and emotions (BDI/E), serves as an intermediate signal connecting agent behaviors with interaction outcomes and reflects how conversational strategies shape users during multi-turn interactions. However, existing evaluation paradigms primarily focus on surface-level responses or final outcomes, providing limited insight into the underlying cognitive processes. This limitation makes it difficult to diagnose why agents succeed or fail and to optimize their interaction strategies. To address this challenge, we propose Cognitive World Model (CogWM), an LLM-based cognitive user model that jointly models users' BDI/E states and corresponding responses, enabling explicit cognitive trajectory tracking. Trained on 150K user-turn samples with Qwen3-14B, CogWM achieves superior performance over existing user simulation baselines in both response fidelity and cognitive state understanding. Interactions with six state-of-the-art LLMs demonstrate that CogWM enables progressive comparison of agents through cognitive trajectories, revealing distinct agent patterns and complementary relationships between cognitive evolution and behavioral outcomes.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.

13:00 JSTエージェント

Latent Actions from Factorized Transition Effects under Agent Ambiguity

Latent Action Models (LAMs) learn action-like proxies from observation. However, in multi-object or distractor-rich scenes, observations co…

13:00 JSTエージェント

LabGuard: 自然言語のラボルールを、身体化されたラボエージェントのランタイムガードに根付かせる

科学的に身体化されたエージェントは、実験室での手順を実行できるようになってきていますが、動的な実験室環境でこれらの手順を安全に実行することは依然として困難です。現在の安全アプローチでは、安全ルール、マニュアル、プロトコル、標準操作手順などの実験室の自然言語を機械チェック可能な実行時制約に変換する中間ステップが見落とされていることがよくあります。 LabGuard (Laboratory Guard) は、自然言語のラボ ルールを実行可能な仕様に根付かせ、ランタイム ガードとして展開する、言語から実行までの安全性スイートです。 LabGuard には、3 つのコア コンポーネントが含まれています。LabGuard-IR は、型指定された実行可能表現を定義します。 LabGuard-Bench は、203 のシード ラボラトリ ルールから拡張された 812 の教師ありアノテーションを提供します。 LabGuard-Grounder は、自然言語のラボルールを LabGuard-IR にマッピングします。結果として得られる IR インスタンスは、LabGuard パイプラインによって処理され、ランタイム モニターにコンパイルされ、コントローラーの境界に適用されます。実験では、LabGuard が目に見えないラボルール ソースに一般化し、79.4 のタスク スコープ F1 を達成し、モニターのコンパイル後に危険なイベントを 39.5% から 23.8% に削減することが示されています。 LabUtopia では、ランタイム モニターが ACT と統合されており、タスクの成功を維持しながら介入を 0.5% 未満に抑えます。

原文 (English)

LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents

Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challenging. Current safety approaches often overlook the intermediate step of transforming laboratory natural language, including safety rules, manuals, protocols, and standard operating procedures, into machine-checkable runtime constraints. We introduce LabGuard (Laboratory Guard), a language-to-execution safety suite that grounds natural-language laboratory rules into executable specifications and deploys them as runtime guards. LabGuard includes three core components: LabGuard-IR, which defines a typed executable representation; LabGuard-Bench, which provides 812 supervised annotations expanded from 203 seed laboratory rules; and LabGuard-Grounder, which maps natural-language laboratory rules into LabGuard-IR. The resulting IR instances are handled by the LabGuard Pipeline, which compiles them into runtime monitors and applies them at the controller boundary. Experiments show that LabGuard generalizes to unseen laboratory-rule sources, achieves 79.4 task-scope F1, and reduces unsafe events from 39.5% to 23.8% after monitor compilation. In LabUtopia, its runtime monitors integrate with ACT, keeping interventions below 0.5% while preserving task success.

13:00 JSTロボティクス

飛行中の航空交通管制をサポートするソリューション空間経路計画

技術の進歩に伴い、航空交通管理用に多くの経路計画アルゴリズムが提案されていますが、戦術管制における運用上の採用は依然として限定的であり、アルゴリズム設計の優先順位と航空管制官のニーズとの間のずれが明らかになりました。これは、本質的に解釈可能で、計算効率が高く、人間が使用するために明示的に設計された意思決定支援ソリューションの必要性を強調しています。この設計課題に焦点を当て、この研究では、次の 2 つの指針となる考慮事項に適合するように設計された、飛行中の航空交通管制 (ATC) のための競合のない経路計画アルゴリズムを開発します。(1) ソリューション空間表示によって提供される解釈可能性と柔軟性。これは、実行可能なすべての安全な行動を明らかにし、変化する最適化目標に対応するアルゴリズムを構築する動機となります。 (2) 決定ロジック コントローラーは、分離基準、操縦性の制限、ウェイポイントの最小化、ルートの実用性などの運用上の制約を強制するときに自然に適用されます。このアルゴリズムは、これらの原則に基づいて、距離ベース、時間間隔ベース、ゾーンベースの 3 つのインテントベースの競合検出方法をソリューション空間フレームワーク内に統合し、計算効率の高い方法で競合のないパスを特定します。さらに、頂点ベースおよびエッジベースの検索ノードが解空間パス プランニング (SSPP) 用に提案されており、その結果、それぞれ SSPPV と SSPPE という 2 つのバリアントが生成され、計算速度と解の品質の観点から評価されます。実験結果によると、ゾーンベースの競合検出と組み合わせた SSPPV は最高のパフォーマンスを実現し、5 nmi グリッドを使用したマーストリヒト上部地域管制センター (MUAC) のデルタ セクターに基づく運用関連シナリオでパスを平均 3.69 ミリ秒で計算します。

原文 (English)

Solution Space Path Planning: A Real-Time Human-Centered Path Planning Algorithm for En-Route Air Traffic Control

As technology advances, various algorithms have been proposed for air traffic management, yet their operational adoption in tactical control remains limited. This gap motivates a human-centered design emphasizing algorithmic interpretability, controller-relevant operational constraints, and real-time computation. Inspired by the interpretability and flexibility of solution-space displays, as well as by the decision logic controllers naturally apply when enforcing operational constraints, this study extends the solution-space concept to path planning and develops a fast conflict-free path-planning algorithm for en-route Air Traffic Control (ATC), termed Solution Space Path Planning (SSPP). The algorithm integrates three intent-based conflict detection methods---distance-based, time-interval-based, and zone-based---within the solution-space framework to identify conflict-free paths in computationally efficient ways. SSPP is developed using both vertex-based and edge-based search nodes, resulting in two variants---SSPPV and SSPPE, respectively. Empirical results show that SSPPV paired with zone-based conflict detection performs best, computing paths in 3.69 ms on average in the Dutch Delta sector using a 5 nmi grid. SSPPV remains approximately 3.77 times faster than SSPPE while offering competitive effectiveness, making it suitable for time-critical operations and interactive 'what-if' probing in real time. An extension to SSPPV and SSPPE further examines the trade-off between delay minimization and separation requirements, demonstrating the flexibility of SSPP in revising optimization objectives. This study not only proposes a novel path-planning algorithm but also shows how such algorithms can be designed to align with human use and operational requirements, supporting their integration into future ATC systems.

13:00 JST研究/論文

スケールではなくアクセス構造からの機能: ハイブリッド シーケンス モデルの下限と事前登録テスト

プラトニック表現仮説 (PRH) は、モデルがスケールするにつれて、異種ネットワークの表現が現実の共有モデルに収束すると考えています。私たちは、その続編であり境界である能力収束仮説 (CCH) を提案します。固定されたトークンごとの推論バジェットの下では、表現的収束は能力の収束を伴いません。代わりに、機能はクラス、つまりアクセス完全ハイブリッド、つまり圧縮 O(1) 状態チャネルとスケーラブルな逐語インデックス チャネルの両方を保持するアーキテクチャに向かって収束します。我々はそれを証人タスクである無限ストリームのニュートンのリンゴ問題に固定し、3 つのリソースの壁と名付けます。o(Nb) 状態アーキテクチャを禁止するシャノン壁、固定ウィンドウを禁止する水平線壁、および固定深さの注意のみの構成を禁止する回路壁 (TC0 != NC1 の条件付き)。明示的な分離可能性の仮定の下では、ハイブリッドは各壁の価格を支払うことによって 3 つすべてを横断するため、合成下では機能は厳密に超加法的になります。私たちは証明したものと推測したものを区別します。アクセス完全性の原則は情報理論の下限と事前に登録された実験に基づいていますが、フィールドレベルの収束傾向は経済学に基づいた推測です。我々は、データの前に凍結された基準に基づいて事前に登録された最初の小規模テストを報告します。予測されたシザーズギャップが測定され(64スカラー状態が1つのグローバルアテンション層を獲得すると、正確な検索誤差は0.994対0.000)、状態追跡分岐は登録された境界に到達し、結合証人は還元できない2チャネルの解決策を示します。 1 つの予測は方向が逆転して失敗したため、そのように報告されます。表現の収束はスケールによって自由に与えられます。機能の収束はアクセス構造ごとに購入する必要があります。

原文 (English)

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale

The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality. We propose its sequel and boundary, the Capability Convergence Hypothesis (CCH): under a fixed per-token inference budget, representational convergence does not entail capability convergence. Capability instead converges toward a class, the access-complete hybrid: any architecture holding both a compressive O(1)-state channel and a scalable verbatim-index channel. We anchor it on a witness task, the Newton's-apple problem in an infinite stream, and name three resource walls: a Shannon wall barring any o(Nb)-state architecture, a horizon wall barring any fixed window, and a circuit wall barring fixed-depth attention-only composition (conditional on TC0 != NC1). Under an explicit separability assumption a hybrid crosses all three by paying each wall's price, so capability is strictly super-additive under composition. We separate what we prove from what we conjecture: the access-completeness principle rests on information-theoretic lower bounds and pre-registered experiments, while the field-level convergence trend is an economics-motivated conjecture. We report the first pre-registered small-scale tests under criteria frozen before the data: the predicted scissors gap is measured (exact-retrieval error 0.994 vs. 0.000 once a 64-scalar state gains one global-attention layer), the state-tracking bifurcation lands at the registered boundary, and a conjunction witness shows an irreducibly two-channel solution; one prediction failed with its direction reversed and is reported as such. Representational convergence is given freely by scale; capability convergence must be purchased by access structure.

13:00 JSTLLM/生成AI

NeurOWL: 不完全なOWLオントロジー推論のためのLLMベースのニューラルシンボリックフレームワーク

OWLオントロジーは、意味論的推論を可能にする形式的な知識表現フレームワークを提供し、ヘルスケアやバイオインフォマティクスなどの分野で広く採用されています。ただし、実際には、現実世界のオントロジーは不完全であることが多く、推論に課題が生じます。この研究では、基本的な包含推論の問題に焦点を当てます。つまり、不完全なオントロジーと候補 (非含意) 包摂が与えられた場合、その包含が意味的に妥当かどうかを判断し、そうであれば、潜在的な欠落している公理を含む論理的に適切な説明を提供します。このタスクは包含検証とオントロジーアブダクションを統合し、欠落している公理の事前定義された候補セットの必要性を取り除くことによって後者を一般化します。この包含推論の問題に対処するために、我々は、正式に定義された意味論と、大規模言語モデルとオントロジー埋め込みを介したテキスト意味論の両方を活用して、検証とアブダクションを共同で実行するエンドツーエンドの神経記号フレームワークである NeurOWL を提案します。私たちは、複数のドメインにわたる現実世界のオントロジーで NeurOWL を評価し、さまざまなドメインにわたって強力で堅牢なパフォーマンスを実証します。

原文 (English)

NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning

OWL ontologies provide a formal knowledge representation framework that enables semantic reasoning, and have been widely adopted across domains such as healthcare and bioinformatics. In practice, however, real-world ontologies are often incomplete, which pose challenges for reasoning. In this work, we focus on a fundamental subsumption reasoning problem: given an incomplete ontology and a candidate (non-entailed) subsumption, determine whether the subsumption is semantically plausible and, if so, providing a logically sound explanation containing potential missing axioms. This task unifies subsumption verification with ontology abduction, and generalizes the latter by removing the need for a predefined candidate set of missing axioms. To address this subsumption reasoning problem, we propose NeurOWL, an end-to-end neuro-symbolic framework that jointly performs verification and abduction, leveraging both formally defined semantics and textual semantics through Large Language Models and ontology embeddings. We evaluate NeurOWL on real-world ontologies across multiple domains, demonstrating strong and robust performance across different domains.

13:00 JST研究/論文

品質保証: VR OSCE における審査官クレームのマルチモーダル検証

客観的構造化臨床検査 (OSCE) は臨床能力を評価するためのゴールドスタンダードですが、採点は依然として検査者の主観、疲労、認知バイアスの影響を受けやすいです。評価者間統計による標準的な審査官の検証は、審査官の推論を分析したり、実際の事象に対する審査官の主張を検証したりしないため、誤りの原因に関する説明力に欠けています。そこで、我々は、ビデオ、VR ログ、俳優データから構築された実際の一連のイベントに対して審査官が主張した行動を比較することにより、バーチャル リアリティ (VR) 小児 OSCE における審査官の主張を検証するマルチモーダル フレームワークである品質アクション保証 (QAA) を導入します。 QAA は、アクションの位置特定とアクターのソースの帰属を実行する制約付きの時間的アクション アライメント モデルと、審査官の主張を抽出して記録と照合する大規模な言語モデルを組み合わせます。 5 分割相互検証を通じて、QAA は時間的アライメントに関して 99.2% $\pm$ 0.7% Actor F1 および 93.4% $\pm$ 1.9% W@16 を達成しました。全体的に、QAA は 70.0% の精度と 76.7% の再現率で検査官のミスを検出し、事実の正確性が 39.2% から 79.2% に向上し、より公平な OSCE 評価が可能になります。

原文 (English)

Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs

Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistics lacks explanatory power regarding the source of errors, as it neither analyzes examiner reasoning nor verifies examiner claims against actual events. Thus, we introduce Quality Action Assurance (QAA), a multimodal framework that verifies examiner claims in Virtual Reality (VR) pediatric OSCEs by comparing actions claimed by examiners against a reference record of events constructed from video, VR logs, and actor annotations. QAA combines a constrained temporal action alignment model, which performs action localization and actor source attribution, with a large language model that extracts examiner claims and checks them against the record. Across a 5-fold cross-validation, QAA achieves 99.2\% $\pm$ 0.7\% Actor F1 and 93.4\% $\pm$ 1.9\% W@16 for temporal alignment. Overall, QAA detects examiner errors with 69.9\% precision and 76.7\% recall; in retrospective evaluation, correcting the detected errors raises the share of factually correct transcripts from 39.2\% to 79.2\%, supporting fairer OSCE quality assessment.

13:00 JSTエージェントGPT / ChatGPT

CodeRescue: コーディング エージェント向けの予算調整されたリカバリ ルーティング

コーディング エージェントは、失敗した試行が単なる不正解ではなく実用的なフィードバックを生成する実行可能環境で動作することが増えています。既存のコストを意識したシステムは通常、このような障害をカスケード決定として扱います。つまり、最初に安価なモデルを試し、その後、困難なケースをより強力でより高価なモデルにエスカレーションします。ただし、コーディングでは、実行フィードバックによって安価なモデルの回復がさらに価値のあるものになる可能性もあり、エージェントはいつより安価なコンピューティングを費やす必要があるのか​​、いつエスカレーションすべきなのかという予算計画上の導入の問題が生じます。この障害後の決定を異種アクションに対する回復ルーティングとして定式化し、実行ロールアウトから監視対象ルーターをトレーニングします。変化する予算の下でも同じルータを使用できるようにするために、再トレーニングなしで導入時のコストペナルティを選択し、交換可能性の下で限界予想コスト制御を提供するコンフォーマルリスクコントロール(CRC)レイヤーを追加します。 5 つのコーディング ベンチマークで継続的に失敗した場合、安価なリカバリとエスカレーションは相補的な成功パターンを示します。調整されたフロンティアは、固定アクション、プロンプト専用ルーター、バイナリ カスケード ベースラインよりも改善されています。メインの GPT-5.4-nano/GPT-5.4 設定では、平均回復コストの 35% を使用しながら、1 つの CRC 校正済みフロンティア ポイントが常時エスカレートの解決速度を超えています。コードは https://github.com/Qijia-He/agent-budget-control で入手できます。

原文 (English)

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts. To make the same router usable under changing budgets, we add a Conformal Risk Control (CRC) layer that selects a deployment-time cost penalty without retraining and provides marginal expected-cost control under exchangeability. Across held-out failures from five coding benchmarks, cheap recovery and escalation exhibit complementary success patterns. The calibrated frontier improves over fixed actions, prompt-only routers, and a binary cascade baseline; in the main GPT-5.4-nano/GPT-5.4 setting, one CRC-calibrated frontier point exceeds always-escalate solve rate while using 35% of its mean recovery cost. Code is available at https://github.com/Qijia-He/agent-budget-control.

13:00 JSTLLM/生成AIエージェント

AttriMem: エージェントの記憶学習のためのアトリビューションに基づくプロセス フィードバック

LLM エージェントにとって効果的な記憶は非常に重要ですが、それを効果的に構築するのは依然として困難です。メモリ構築ポリシーは、インタラクションが蓄積するにつれてどの情報を抽出、保存、更新、圧縮、または破棄するかを決定します。ヒューリスティック記憶手法は主観的なタスク固有のルールに依存しているため、下流の目標とずれたり、タスク間の適応性が制限されたりする可能性があります。対照的に、RL ベースの手法はタスクのフィードバックから学習しますが、主に結果レベルまたはモジュールレベルの報酬を使用します。これらの粗い信号はタスクの成功を示しますが、どの中間メモリの内容が最終的な答えをサポートしているかを特定できず、きめの細かいクレジット割り当てのボトルネックが生じます。ただし、このようなプロセス フィードバックの構築は、中間記憶の決定には固有のグラウンドトゥルース ターゲットが欠けている一方、適切なクレジットはエージェントの不確実な推論軌道によって変化するため、事前に指定できないため、非常に困難です。我々は、RL を使用してメモリ構築ポリシーを学習するためのアトリビューションに基づくプロセス フィードバック フレームワークである AttriMem を提案します。 AttriMem は、最終的な回答へのトークンレベルの貢献から得られるローカルな報酬でグローバルな結果報酬を強化します。長期対話型質​​問応答の実験では、AttriMem が検索ベース、ヒューリスティック、RL ベースのベースラインを上回り、ベンチマークと回答モデル全体で一般化され、RL の最適化が安定することが示されました。

原文 (English)

AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction

Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with downstream objectives and limit cross-task adaptability. RL-based methods, by contrast, learn from task feedback but mainly use outcome- or module-level rewards. These coarse signals indicate task success but cannot identify which intermediate memory contents support the final answer, creating a fine-grained credit-assignment bottleneck. However, constructing such process feedback is prohibitively difficult because intermediate memory decisions lack unique ground-truth targets, while the appropriate credit varies with the agent's uncertain reasoning trajectory and therefore cannot be specified in advance. We propose AttriMem, an attribution-guided process-feedback framework for learning memory-construction policies with RL. AttriMem augments the global outcome reward with local rewards derived from token-level contributions to the final answer. Experiments on long-horizon dialogue question answering show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization.

13:00 JSTLLM/生成AI

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy…

13:00 JSTLLM/生成AIハードウェア/半導体

DynaResize: 分割された LLM ポストトレーニングのためのランタイム GPU 再割り当て

RL ベースの LLM ポストトレーニングでは、別々の GPU リソース間でロールアウトとトレーニングがますます細分化されますが、静的 GPU パーティショニングでは、ロングテール ロールアウト レイテンシの下で深刻なパイプライン バブルが発生します。 DynaResize は、RL セマンティクスを変更せずに、ロールアウトとトレーニングの間で GPU を動的に切り替えてステージの実行時間のバランスを取る、ランタイム GPU 再割り当てシステムです。 DynaResize は、サイズ変更をきめ細かい操作に分解し、コミュニケーターの再利用、制限された状態のステージング、およびヒステリシスベースのサイズ変更を通じて、起動にクリティカルではない作業をクリティカル パスから削除します。実験結果によると、DynaResize は、最適な静的構成と比較して、エンドツーエンドのスループットを 66.5% 向上させ、合計実行時間を 33% 削減し、同時にロール切り替えオーバーヘッドの 27% を隠すことができます。

原文 (English)

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.

13:00 JSTLLM/生成AI

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enab…

13:00 JSTLLM/生成AIビジネス/資金調達

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…

13:00 JST研究/論文

Information Processing by Neuron Populations in the Central Nervous System: A Theory of the Mathematical Structure of Data and Operations

In the mammalian central nervous system, neurons are organized into populations communicating by spike trains propagating along axonal bund…

13:00 JST研究/論文

On the Expressive Power of Sparse Geometric MPNNs

Motivated by applications in chemistry and other sciences, we study the expressive power of message-passing neural networks for geometric g…

13:00 JST研究/論文

Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations

This paper explores the characterization of equivariant linear layers for representations of permutations and related groups. Unlike tradit…

13:00 JST画像/動画生成

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook

We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We i…

13:00 JSTロボティクス

Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints

Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, imp…

13:00 JST研究/論文

Dimensionality reduction for homological stability and global structure preservation

We propose DiRe, a force-directed dimensionality reduction framework designed to preserve global structure and homological features while r…

13:00 JSTエージェントロボティクス

Reproducing Human Individual Motor Signatures: A Data-Driven Approach for Repetitive Motion

The deployment of autonomous virtual avatars (in extended reality) and robots in human group activities---such as rehabilitation therapy, s…

13:00 JST研究/論文

StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent

In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive polici…

13:00 JST研究/論文

Towards White-Box Deep Wireless Sensing

The empirical success of deep learning has spurred its application to the radio-frequency (RF) domain, leading to significant advances in D…

13:00 JST研究/論文

Patch-Based 3D Variational Autoencoder for Super-Resolution of Turbulent Channel Flow

Direct numerical simulation (DNS) accurately resolves all spatio-temporal scales of wall-bounded turbulence but becomes prohibitively expen…

13:00 JST研究/論文

RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment

Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention. However, existing tools t…

13:00 JSTLLM/生成AIGPT / ChatGPT

"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness

Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experiencing homelessness (PEH) were recor…

13:00 JST研究/論文

Adaptive Policy Backbone via Shared Network

Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive int…

13:00 JST画像/動画生成ロボティクス研究/論文

Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events

This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast…

13:00 JST画像/動画生成

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation

Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are…

13:00 JST研究/論文

Monotone and Separable Set Functions: Characterizations and Neural Models

Motivated by applications for set containment problems, we consider the following fundamental problem: can we design set-to-vector function…

13:00 JSTLLM/生成AI

Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers

The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability:…

13:00 JST研究/論文

Robust Bidirectional Associative Memory via Regularization Inspired by the Subspace Rotation Algorithm

Bidirectional Associative Memory (BAM) trained with Bidirectional Backpropagation (B-BP) often suffers from poor robustness and high sensit…

13:00 JST画像/動画生成エージェントロボティクス

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rat…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics

Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and…

13:00 JSTハードウェア/半導体

GPU-Accelerated ANNS: Quantized for Speed, Built for Change

Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promi…

13:00 JST研究/論文LlamaQwen

GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models. Unlike supervised fine-…

13:00 JSTLLM/生成AI

Knowledge Restoration-driven Prompt Optimization: Unlocking LLM Potential for Open-Domain Relational Triplet Extraction

Open-domain Relational Triplet Extraction (ORTE) aims to mine structured knowledge without predefined relation schemas. Large Language Mode…

13:00 JSTLLM/生成AI

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative ret…

13:00 JSTLLM/生成AI

Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation

The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-…

13:00 JSTLLM/生成AIエージェント

AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles

AIvilization v0 is a publicly deployed large-scale artificial society that couples a resource-constrained sandbox with a unified LLM-agent…

13:00 JST研究/論文GemmaMistral AI

Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization

Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas ot…

13:00 JSTLLM/生成AIエージェント研究/論文

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery

Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However,…

13:00 JSTLLM/生成AI

Stem: Rethinking Causal Information Flow in Sparse Attention

The quadratic computational complexity of self-attention remains a fundamental bottleneck for scaling Large Language Models (LLMs) to long…

13:00 JST画像/動画生成

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models

We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Sta…

13:00 JSTLLM/生成AI

Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL

Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected…

13:00 JSTエージェント

ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics

Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI…

13:00 JSTLLM/生成AI

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation

Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias pro…

13:00 JST画像/動画生成

謎を解くビデオ推論

ビデオ生成の最近の進歩により、予期せぬ現象が明らかになりました。拡散ベースのビデオ モデルは、自明ではない推論機能を示します。以前の研究では、これはチェーン オブ フレーム (CoF) メカニズムによるものであり、推論はビデオ フレーム間で順次展開されると想定されています。この研究では、この仮定に異議を唱え、根本的に異なるメカニズムを明らかにします。ビデオ モデルにおける推論は、主に拡散ノイズ除去ステップに沿って現れることを示します。定性分析と対象を絞った調査実験を通じて、モデルは初期のノイズ除去ステップで複数の候補ソリューションを探索し、最終的な答えに徐々に収束することがわかりました。これは、Chain-of-Steps (CoS) と呼ばれるプロセスです。この中心的なメカニズムを超えて、モデルのパフォーマンスに重要ないくつかの新たな推論動作を特定します。(1) ワーキングメモリ。永続的な参照を可能にします。 (2) 自己修正と強化により、誤った中間ソリューションからの回復が可能になります。 (3) アクション前の認識。初期のステップで意味論的な基礎を確立し、後のステップで構造化された操作を実行します。拡散ステップ中に、拡散トランスフォーマー内の自己進化した機能的特殊化をさらに明らかにします。初期の層は高密度の知覚構造をエンコードし、中間の層は推論を実行し、後の層は潜在的な表現を統合します。これらの洞察に動機付けられて、私たちは概念実証としてトレーニング不要のシンプルな戦略を提示し、異なるランダム シードを持つ同一のモデルからの潜在軌道をアンサンブルすることによって推論がどのように改善されるかを実証します。全体として、私たちの研究は、ビデオ生成モデルで推論がどのように現れるかについて体系的な理解を提供し、インテリジェンスの新しい基盤としてビデオ モデルの固有の推論ダイナミクスをより効果的に活用する将来の研究を導くための基盤を提供します。

原文 (English)

Demystifying Video Reasoning

Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities. Prior work attributes this to a Chain-of-Frames (CoF) mechanism, where reasoning is assumed to unfold sequentially across video frames. In this work, we challenge this assumption and uncover a fundamentally different mechanism. We show that reasoning in video models instead primarily emerges along the diffusion denoising steps. Through qualitative analysis and targeted probing experiments, we find that models explore multiple candidate solutions in early denoising steps and progressively converge to a final answer, a process we term Chain-of-Steps (CoS). Beyond this core mechanism, we identify several emergent reasoning behaviors critical to model performance: (1) working memory that supports tasks requiring consistent reference, such as object permanence; (2) self-correction and enhancement, allowing recovery from incorrect intermediate solutions; and (3) perception before action, where early steps establish semantic grounding and later steps perform structured manipulation. Moreover, analysis of Diffusion Transformer layers shows that middle layers conduct key reasoning procedures. Motivated by these insights, we present a simple Training-Free Ensemble (TFE) as a proof-of-concept, demonstrating how reasoning can be improved by ensembling latent trajectories from identical models with different random seeds. Overall, our work provides the first systematic dissection of the mechanisms underlying video reasoning, offering a foundation to guide future research in better exploiting the inherent reasoning dynamics of video models as a new substrate for intelligence.

13:00 JSTLLM/生成AI

OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation

Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introdu…

13:00 JSTLLM/生成AIエージェント

Agentic Harness for Real-World Compilers

Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements ena…

13:00 JST研究/論文

Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning

Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies…

13:00 JST研究/論文Alibaba

Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations

In collaboration with Alibaba, we study how a generative AI assistant affects service performance in e-commerce after-sales operations. In…

13:00 JST画像/動画生成

ActionParty: Multi-Subject Action Binding in Generative Video Games

Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However,…

13:00 JSTLLM/生成AI

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation

We study compiled AI, a paradigm in which large language models generate executable code artifacts during a compilation phase, after which…

13:00 JST画像/動画生成

Evaluating the Alignment Between GeoAI Explanations and Domain Knowledge in Satellite-Based Flood Mapping

The increasing number of satellites has improved the temporal resolution of Earth observation, making satellite-based flood mapping a promi…

13:00 JST画像/動画生成

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning

Time Series Foundation Models (TSFMs) have demonstrated strong generalization capability and data efficiency in time series forecasting thr…

13:00 JSTLLM/生成AI

幾何学的規制による LLM 生成におけるエスケープ モードの崩壊

モード崩壊は生成モデリングにおける永続的な課題であり、明示的なループから多様性の漸進的な喪失や時期尚早な軌道収束に至るまでの範囲の動作として自己回帰テキスト生成に現れます。私たちは力学システムの視点をとり、モード崩壊を *幾何学的崩壊* によって引き起こされる状態空間へのアクセス可能性の低下として再解釈します。生成中、モデルの内部軌道はその表現空間の低次元領​​域に限定されます。これは、モード崩壊が純粋にトークンレベルの現象ではなく、記号的制約や確率のみの復号ヒューリスティックでは確実に解決できないことを意味します。この観点に基づいて、私たちは、Transformer 値キャッシュ (低ランクのダンピングとして実装) 内の主要な自己強化方向を制御する軽量のオンライン状態空間介入である *強化モード制御* (RMR) を提案します。複数の大規模な言語モデルにわたって、RMR はモード崩壊を大幅に軽減し、非常に低いエントロピー レート (0.8 nats/ステップまで) での安定した生成を可能にしますが、標準のデコードでは通常 2.0 nats/ステップ近くで崩壊します。

原文 (English)

Escaping Mode Collapse in LLM Generation via Geometric Regulation

Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping to gradual loss of diversity and premature trajectory convergence. We take a dynamical-systems view and reinterpret mode collapse as reduced state-space accessibility caused by *geometric collapse*: during generation, the model's internal trajectory becomes confined to a low-dimensional region of its representation space. This implies mode collapse is not purely a token-level phenomenon and cannot be reliably solved by symbolic constraints or probability-only decoding heuristics. Guided by this perspective, we propose *Reinforced Mode Regulation* (RMR), a lightweight, online state-space intervention that regulates dominant self-reinforcing directions in the Transformer value cache (implemented as low-rank damping). Across multiple large language models, RMR substantially reduces mode collapse and enables stable generation at extremely low entropy rates (down to 0.8 nats/step), whereas standard decoding typically collapses near 2.0 nats/step.

13:00 JSTLLM/生成AI

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs

Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation th…

13:00 JST画像/動画生成ロボティクス

Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geogra…

13:00 JST画像/動画生成

Detecting AI-Generated Videos with Spiking Neural Networks

Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detect…

13:00 JST研究/論文

A Nonlinear Singular Value Theory for Neural Networks

Recently Brown et al. [2025] established a singular value decomposition (SVD) for maps (especially nonlinear) satisfying certain norm condi…

13:00 JSTLLM/生成AIエージェント研究/論文

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by i…

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTGeminiLlamaMistral AIDeepSeekGrok

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where…

13:00 JST画像/動画生成

When Bits Break Recourse: Counterfactual-Faithful Quantization

Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is…

13:00 JST研究/論文

AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers

Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have…

13:00 JST画像/動画生成

DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation

Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term…

13:00 JSTLLM/生成AIエージェントGemma

MemForest: 階層型時間インデックスを備えた効率的なエージェント メモリ システム

メモリは、ロングコンテキストの LLM エージェントを有効にするための基本コンポーネントであり、継続的な提供と更新のライフサイクルを通じて対話全体にわたる永続的な状態をサポートします。相当な事前作業にもかかわらず、既存のシステムは、粗粒度の状態管理と本質的に逐次的な更新パイプラインという 2 つの重要な制限により、重大なメンテナンスのオーバーヘッドに悩まされています。特に、更新は LLM 推論と密接に結びついていることが多く、完全な状態の書き換えが必要なため、スケーラビリティが低下し、メモリが蓄積するにつれて遅延が増大します。これらの課題に対処するために、エージェントのメモリを書き込み効率の高い時間データ管理問題として再定式化するメモリ フレームワークである MemForest を紹介します。 MemForest は、並列チャンク抽出によってシーケンシャル ボトルネックを解消し、メモリ構築を同時の独立した操作に分離します。粗粒度のメンテナンスをさらに排除するために、フラットなグローバル サマリーではなく時間順のツリーとしてメモリを編成する階層型時間インデックスである MemTree を導入します。この設計では、完全な状態の書き換えを局所的なノードごとの更新に置き換え、影響を受けるツリー パスのメンテナンス コストを削減しながら、時間的に変化する状態を自然に保存します。私たちは、LongMemEval-S と LoCoMo という 2 つのロングコンテキスト メモリ ベンチマークで MemForest を評価します。 LongMemEval-S では、MemForest はステートフル ベースラインの中で最高の総合パフォーマンスを達成し、EverMemOS を含む最先端のアプローチよりも約 6 倍高いメモリ構築スループットを維持しながら、79.8% pass@1 精度に達します。

原文 (English)

MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing

Memory is a fundamental component for long-context LLM agents, supporting persistent state across interactions through a continuous serve-and-update lifecycle. Despite substantial prior work, many stateful systems retain sequential autoregressive extraction or state-dependent maintenance on the write path, delaying when new evidence becomes queryable. To address these challenges, we present MemForest, a memory framework that reformulates agent memory as a write-efficient temporal data-management problem. MemForest breaks the sequential bottleneck via parallel extraction, decoupling memory construction into concurrent, independent operations. We further introduce MemTree, a hierarchical temporal index that organizes memory as time-ordered trees and replaces global rewrites with localized dirty-path refresh. Dirty summaries can be refreshed in parallel across nodes and trees. End-to-end work remains proportional to incoming content; the logarithmic bound applies only to structural insertion and level-dependent refresh depth in balanced trees. We evaluate MemForest on two long-context benchmarks, LongMemEval-S and LoCoMo. Experiments use Qwen3-4B, Qwen3-30B, and Gemma-4-12B-IT. With Qwen3-30B, MemForest reaches 81.8 percent pass at 1 on LongMemEval-S, while its input-normalized build rate is 6.0 times that of EverMemOS. On LoCoMo categories 1 to 4, it reaches 84.09 percent, within 0.13 percentage points of EverMemOS; on a matched conversation, its build rate is 9.5 times higher. These results show that MemForest reduces memory-freshness latency while retaining strong answer quality.

13:00 JSTLLM/生成AIビジネス/資金調達ClaudeGPT / ChatGPTGemini

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptatio…

13:00 JST研究/論文

THzデュアルコム分光法を使用したポリマー分類のためのマルチスケール機能アテンションネットワーク

信頼性の高いポリマーの識別は、リサイクルプラスチックの品質と安全性を確保するために不可欠ですが、従来の分別技術や分光技術では、確実な識別を実現するのが困難なことがよくあります。テラヘルツ デュアルコム分光法 (THz-DCS) は、迅速、高分解能、非破壊測定を提供する有望な代替手段を提供します。この研究では、THz-DCS を利用して、純粋なポリマー、多層フィルム、市販のブレンド、バイオポリマーを含む 12 種類のポリマーを分類します。これらのスペクトル信号の複雑さを処理するために、THz-DCS データに合わせた新しい深層学習アーキテクチャであるマルチスケール フィーチャー アテンション ネットワーク (MSFAN) を提案します。このフレームワークには、信号の再キャリブレーションとマルチスケールの並列畳み込みのための機能ゲートが統合されており、多様な周波数パターンをキャプチャします。これらの特徴は、特徴間アテンションとアテンション プーリングを通じてさらに洗練され、モデルが本質的に最も有益な THz 領域を強調表示できるようになります。 MSFAN は常に最先端のモデルを上回っており、分類精度は 85.2% に達しています。この研究は、THz-DCS と深層学習技術を組み合わせて、効果的でスケーラブルで解釈可能なポリマー分類を実現できる可能性を示しています。

原文 (English)

Multi-Scale Feature Attention Network for Polymer Classification Using Terahertz Spectroscopy

Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet conventional sorting and spectroscopic techniques often struggle to deliver robust discrimination. Terahertz (THz) spectroscopy offers a promising alternative, providing high-resolution and non-destructive measurements. In this work, we leverage THz signals to classify 12 types of polymers, including pure polymers, multilayer films, commercial blends, and biopolymers. To handle the complexity of these spectral signals, we propose the Multi-Scale Feature Attention Network (MSFAN), a novel deep learning architecture tailored for THz data. The framework integrates feature gating for signal recalibration and multi-scale parallel convolutions to capture diverse frequency patterns. These features are further refined through cross-feature attention and attention pooling, enabling the model to intrinsically highlight the most informative THz regions. MSFAN consistently outperforms state-of-the-art models, reaching a classification accuracy of 85.2%. This study demonstrates the potential of combining THz spectroscopy with deep learning techniques for effective, scalable, and interpretable polymer classification.

13:00 JSTエージェント

APPO: Agentic Procedural Policy Optimization

Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language m…

13:00 JSTLLM/生成AIビジネス/資金調達

Creative Integration: A Decidable Criterion of Creativity

"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…

13:00 JSTLLM/生成AI

大規模言語モデルベースの生成推奨の暗黙的推論

大規模言語モデル (LLM) は生成推奨 (GR) のバックボーンとして採用されることが増えており、事前トレーニングされた世界の知識へのアクセスが約束されています。しかし、この知識を GR に確実に活用する方法は、まだ十分に理解されていません。主な障害は、LLM ベースの GR が通常、アイテムをセマンティック ID (SID) で表現し、事前トレーニング中にこれらのトークンが LLM に認識されないため、LLM の自然言語推論インターフェイスを混乱させることです。既存のアプローチは、SID を接地して明示的な根拠を引き出す高価なマルチステージ パイプラインでこの問題に対処していますが、各ステージがいつ、なぜ必要なのかについての洞察は限られています。この研究では、LLM ベースの GR の明示的推論トレーニング パイプラインを体系的に分解し、3 つの重要な制限を明らかにしました。世界知識の言語化の弱体化、SID と自然言語トークン埋め込み空間間の不整合、理論的根拠の品質に対する敏感さであり、これらすべてが明示的推論のパフォーマンスに悪影響を及ぼします。これらの問題を回避するために、GR 向けに調整された軽量の暗黙的推論パラダイムである PauseRec を提案します。 PauseRec は非常に実用的で、コストのかかる推論トレース取得と推論調整トレーニングを回避し、多くの利点をもたらします。(1) 標準の明示的 CoT メソッドよりも最大 6.22% 優れたパフォーマンスを発揮し、(2) トレーニング コストを GPU 時間で最大 65% 削減し、(3) 推論を最大 71.3% 高速化します。これらの結果により、PauseRec は明示的な根拠生成に代わる軽量の代替手段として位置づけられ、より効果的かつ効率的な LLM ベースの GR が可能になります。

原文 (English)

Implicit Reasoning for Large Language Model-based Generative Recommendation

Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world knowledge. Yet reliably invoking this knowledge for GR remains poorly understood. A key obstacle is that LLM-based GR typically represents items with Semantic IDs (SIDs), disrupting LLMs' natural-language reasoning interface because these tokens are unseen by the LLM during pretraining. Existing approaches address this with expensive multi-stage pipelines that ground SIDs and elicit explicit rationales, but offer limited insight into when and why each stage is necessary. In this work, we systematically decompose explicit reasoning training pipelines for LLM-based GR, revealing three key limitations: weakened world-knowledge verbalization, misalignment between SID and natural-language token embedding spaces, and sensitivity to rationale quality, all of which hurt explicit reasoning performance. To circumvent these issues, we propose PauseRec, a lightweight implicit reasoning paradigm tailored for GR. PauseRec is exceptionally practical, avoiding costly reasoning trace acquisition and reasoning alignment training, leading to a multitude of benefits: (1) it outperforms standard explicit CoT methods by up to 6.22%, (2) it reduces training cost by up to 65% GPU hours, and (3) it speeds up inference by up to 71.3%. These results position PauseRec as a lightweight alternative to explicit rationale generation, enabling more effective and efficient LLM-based GR.

13:00 JSTLLM/生成AI研究/論文

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence

The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs…

13:00 JST研究/論文

SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting

Traffic prediction is a core task in intelligent transportation systems and urban-scale decision making. Despite the effectiveness of mains…

13:00 JST研究/論文

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in comple…

13:00 JSTエージェント

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL

Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Co…

13:00 JSTLLM/生成AI画像/動画生成

MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a n…

13:00 JST研究/論文

BeatEdit: Symbolic Music Generation as Explicit Editing

Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete s…

13:00 JST研究/論文

AuEmoChat: 会話音声合成のための本物の感情の理解とレンダリング

会話型音声合成 (CSS) は、ユーザーとエージェントの対話において、人間のような感情表現と文脈の一貫性を備えた音声を合成することを目的としています。既存の CSS 手法は、事前に定義された感情ラベル スペース (7 つの感情カテゴリなど) が限られているため、本物の人間の感情を表現するのに苦労していますが、マルチターン対話履歴内の冗長なマルチモーダル トークンがコンテキストの理解を妨げます。これらの問題に対処するために、私たちは本物の感情の理解とレンダリングのための CSS フレームワークである AuEmoChat を提案します。まず、有限スカラー量子化を介して大規模な感情音声から離散的な本物の感情トークン空間を学習する AuEmoCodec を開発し、限られた基本的な感情カテゴリよりもより本物の感情表現を可能にします。さらに、感情に関連したコンテキストを維持しながら、マルチモーダルな対話履歴内の冗長トークンをマージする、本物の感情に基づくトークンマージアルゴリズムである AuEmoToMe を提案します。これを自己回帰テキスト音声モデルに統合して、ターゲットとなる本物の感情トークンと音声トークンを予測します。最後に、統合された対話コンテキスト、ターゲットの本物の感情、および音響事前分布を共同で条件付けすることによって音声をレンダリングする、本物の感情フロー マッチングを提案します。 NCSSD-EmCap データセットに関する広範な実験により、AuEmoChat が最先端の CSS ベースラインを上回り、より表現力豊かで本物の感情的なスピーチを生成することが実証されました。

原文 (English)

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. Furthermore, we propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech. The code and speech demos will be available at: https://github.com/AI-S2-Lab/AuEmoChat.

13:00 JSTLLM/生成AI研究/論文

ビジネス分野全体にわたる最先端の AI パフォーマンス: ナレッジワークと分析的推論の事例に基づいたベンチマーク

大規模言語モデル (LLM) は、ベンチマーク スコアに反映されているように急速に改善されていますが、これらの AI ベンチマークでは主に、事実の再現、限定的な質問応答、数学的問題解決、コーディングやエージェント ツールの使用などの機能がテストされます。まだ十分に測定されていないのは、複雑な情報の統合、不確実性と不完全な情報の下での判断の行使、複数のステークホルダーの状況での戦略的および敵対的思考の適用、トレードオフの比較検討、防御可能な構造化された分析の作成など、ホワイトカラーの専門家が日々行っている分析知識作業における AI の進歩です。このギャップは、そのような仕事の主観的な要素ではさらに顕著であり、成功を定義するのが難しい場合があります。トップクラスのビジネススクールが実践する「ケースメソッド」教育形式は、この測定ギャップに対処するための自然な基盤を提供します。私たちは、18 分野にわたるビジネスケースから抽出された数百の質問にわたるベンチマークである BusinessCaseBench を構築します。各質問は、専門家が作成した講師のケースソリューションから導き出された採点ルーブリックと対になっています。 BusinessCaseBench では、フロンティア AI モデルはすでにインストラクターのルーブリックに対して高いスコアを獲得しており、1 つのモデル ファミリー内の機能は 2 年間で大幅に向上しています。これらの結果は、この種の作業における AI のパフォーマンスがすでに高く、急速に向上していることを示す強力な証拠を提供します。これは、事例教育学によって学部生や MBA がこの種の分析的推論を訓練されるビジネス スクールや、歴史的にそのようなスキルが初期キャリアの仕事に定着してきたエントリーレベルの専門職に影響を及ぼします。

原文 (English)

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…

13:00 JSTLLM/生成AI

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

Textual Collaborative Prompt Optimization (TCPO) extends TextGrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multipl…

13:00 JSTLLM/生成AI

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to…

13:00 JSTLLM/生成AI

HijackKV: New Threat in Position-Independent KV Cache Reuse

Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates acro…

13:00 JST画像/動画生成

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundat…

13:00 JSTLLM/生成AIGemmaQwen

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of wheth…

13:00 JSTLLM/生成AIエージェントロボティクス

Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric

Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in…

13:00 JST研究/論文QwenDeepSeek

DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifie…

13:00 JSTLLM/生成AI

Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature

X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published…

13:00 JSTLLM/生成AIエージェント

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

Language-model agents act through structured tool calls whose arguments carry very different risks: untrusted content may legitimately shap…

13:00 JSTLLM/生成AIMeta

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this g…

13:00 JST研究/論文

A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks

Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies…

13:00 JSTLLM/生成AI画像/動画生成

Progressive Multimodal Alignment for Continual Instruction Tuning

Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it c…

13:00 JSTLLM/生成AI研究/論文

Benchmarking LLM Competence on Logical Inference over Probability Operators

Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions o…

13:00 JST研究/論文

SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups

Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Exis…

13:00 JSTエージェントロボティクス

LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experie…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn…

13:00 JST研究/論文

On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems

We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and…