Skip to the content.

AIニュース 2026-07-04

自動生成: 2026-07-04 12:38 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Google DeepMind and A24 announce first-of-its-kind research partnershipGoogle DeepMind
  2. フィジカルAIに挑む日の丸連合、「Noetra」とは何かITmedia AI+

    2026年6月29日~7月3日に公開された記事の中から、MONOist編集部が厳選した今週の注目ニュースをお届けします。

  3. 3万円で「Yahoo!ニュース」にPR掲載 プレスリリースをAIで「ニュース風記事」にITmedia AI+

    Yahoo!ニュース内に、企業の発表情報をニュース記事と同じフォーマットで掲載する。

  4. 米トランプ大統領、AI規制は「できるだけ介入少なく」 中国に対して開発競争「大幅リード」を強調ITmedia AI+

    トランプ米大統領がCNBCのインタビューで、AI規制について「ガードレールは必要だが、介入はできるだけ少なくしたい」と述べ、規制最小限の路…

  5. The only AI glossary you’ll need this yearTechCrunch AI

    The rise of AI has brought an avalanche of new terms and slang. Here…

  6. The browser wars aren’t about search anymore — here are the best alternatives to Chrome and SafariTechCrunch AI

    We’ve compiled an overview of some of the top alternative browsers av…

  7. 「Claude Fable 5」をサブスクの標準機能に――AnthropicのエンジニアがXに投稿 7月8日以降の「早期復活目指す」ITmedia AI+

    「Fable 5をサブスクリプションの標準機能として復活させることを目指している」――米AnthropicのITエンジニアであるタリク・シ…

トピック別件数

日本語メディア4件

ITmedia AI+ (日本語)

08:00 JST規制/政策

米トランプ大統領、AI規制は「できるだけ介入少なく」 中国に対して開発競争「大幅リード」を強調

トランプ米大統領がCNBCのインタビューで、AI規制について「ガードレールは必要だが、介入はできるだけ少なくしたい」と述べ、規制最小限の路線を改めて示した。

07:00 JSTその他

フィジカルAIに挑む日の丸連合、「Noetra」とは何か

2026年6月29日~7月3日に公開された記事の中から、MONOist編集部が厳選した今週の注目ニュースをお届けします。

15:44 JSTその他

3万円で「Yahoo!ニュース」にPR掲載 プレスリリースをAIで「ニュース風記事」に

Yahoo!ニュース内に、企業の発表情報をニュース記事と同じフォーマットで掲載する。

15:34 JSTLLM/生成AIAnthropicClaude2件の関連記事

「Claude Fable 5」をサブスクの標準機能に――AnthropicのエンジニアがXに投稿 7月8日以降の「早期復活目指す」

「Fable 5をサブスクリプションの標準機能として復活させることを目指している」――米AnthropicのITエンジニアであるタリク・シヒパー氏は、自身のXアカウントにこのように投稿した。

出典:ITmedia AI+ITmedia AI+
海外メディア2件

TechCrunch AI (英語)

06:20 JSTその他

The only AI glossary you’ll need this year

The rise of AI has brought an avalanche of new terms and slang. Here is a glossary with definitions of some of the most important words and…

03:43 JSTその他

The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

We’ve compiled an overview of some of the top alternative browsers available today aiming to challenge Chrome and Safari.

公式ブログ1件

Google DeepMind (英語)

論文353件

arXiv cs.AI (英語)

13:00 JST研究/論文

PACE: もっともらしく実行可能な反事実説明のための神経象徴的フレームワーク

反事実的な説明は、モデルの決定を変える可能性のある最小限の入力変更を特定することによって、機械学習の予測を説明します。既存の手法の多くは、予測を変える代替案を生成することに成功していますが、ドメインの知識や介入の制約を組み込むための明示的なメカニズムが欠如しているため、非現実的または実行不可能な推奨事項が生成されることがよくあります。ニューロシンボリック AI は、データ駆動型の予測モデルと、人間が理解できるルールや実行可能なアクションを表現できる記号推論を組み合わせることで、有望な方向性を提供します。この論文では、実現可能性を意識した反事実の説明を生成するためのモジュール式神経記号フレームワークである PACE について説明します。このフレームワークは、予測と推論を 2 つのコンポーネントに分離します。分類用のニューラル予測モデルと、反事実の生成中にドメイン固有の制約を強制する記号推論層です。実行可能な介入を明示的にモデル化することにより、フレームワークは、解釈可能かつ実行可能でありながら、ドメイン知識と一貫した説明を生成します。このアプローチはモデルに依存せず、現実的な意思決定サポートを必要とする領域に適応できます。ケーススタディは、成人所得データセットに対して行われ、多層パーセプトロン分類器と、不変の属性を維持しながら教育、職業、労働時間に対する実行可能な変更をエンコードする回答セット プログラミング (ASP) ルールを組み合わせています。結果は、反事実の妥当性ともっともらしさの間のトレードオフを強調し、記号的制約がドメイン固有の実現可能性要件をよりよく満たす説明を生み出すことを示し、説明可能な AI における透明で実現可能性を意識した反事実の説明に対する神経記号的手法の可能性を示しています。

原文 (English)

PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations

Counterfactual explanations explain machine learning predictions by identifying minimal input changes that would alter a model's decision. Although many existing methods successfully generate prediction-changing alternatives, they often produce unrealistic or infeasible recommendations due to a lack of explicit mechanisms for incorporating domain knowledge and intervention constraints. Neuro-symbolic AI offers a promising direction by combining data-driven predictive models with symbolic reasoning capable of representing human-understandable rules and feasible actions. This paper presents PACE, a modular neuro-symbolic framework for generating feasibility-aware counterfactual explanations. The framework separates prediction and reasoning into two components: a neural predictive model for classification and a symbolic reasoning layer that enforces domain-specific constraints during counterfactual generation. By explicitly modeling feasible interventions, the framework produces explanations consistent with domain knowledge while remaining interpretable and actionable. The approach is model-agnostic and adaptable to domains requiring realistic decision support. A case study is conducted on the Adult Income dataset, combining a multilayer perceptron classifier with Answer Set Programming (ASP) rules encoding feasible modifications to education, occupation, and working hours while preserving immutable attributes. Results highlight the trade-off between counterfactual validity and plausibility and show that symbolic constraints yield explanations that better satisfy domain-specific feasibility requirements, illustrating the potential of neuro-symbolic methods for transparent, feasibility-aware counterfactual explanation in explainable AI.

13:00 JSTエージェント研究/論文

Auto-FL-Research: フェデレーション学習アルゴリズムのエージェント検索

フェデレーション ラーニング (FL) の研究は、オプティマイザーのバリアント、サーバー集約ルール、ローカル トレーニング スケジュール、正規化、正則化、モデル アーキテクチャなど、多くの小規模だが結果的なアルゴリズムの選択に依存することがよくあります。これらの選択肢を手動で探索するにはコストがかかり、候補の変更によって FL トレーニングまたは評価パスも変更される可能性があるため、公正に比較することは困難です。この研究では、FL アルゴリズム レシピ検索のための制約付きコーディング エージェント ワークフローである Auto-FL-Research (AFR) を紹介します。エージェントは、サーバー集約ルール、クライアント更新スケジュール、ローカル目標、登録されたモデルのバリアントなどの候補トレーニング アルゴリズムを提案および実装できます。一方、タスク プロファイルは、突然変異表面、計算バジェット、通信契約、および最終モデル評価を修正します。各キャンペーンでは、候補者のスコア、実行時間、編集されたファイル、アーティファクト、失敗ステータスが記録されます。 5 つのヘルスケア クロスサイロ FLamby タスクと、5 つの固定 LEAF データセットと LEAF 合成タスクのグループ化されたクライアント プロファイルで AFR を評価します。 5 シードの繰り返し評価は、4 つの FLamby タスクと 6 つの LEAF プロファイルのうち 5 つでの向上をサポートすると同時に、シードに依存する失敗ケースや検索で選択された失敗ケースも明らかにします。同じ予算の制御では、いくつかのゲインが FL レシピの変更に対応することが示されていますが、他の改善は固定曲面スカラー制御によって回復されるか、繰り返しまたは保留された評価で失敗します。これらの混合結果は貢献の一部です。これらは、エージェントによって生成された候補が、反復 FL メカニズム、固定表面調整効果、および選択された単一実行アーティファクトにどのように分離できるかを示しています。

原文 (English)

Auto-FL-Research: Agentic Search for Federated Learning Algorithms

Federated learning (FL) research often depends on many small but consequential algorithmic choices: optimizer variants, server aggregation rules, local training schedules, normalization, regularization, and model architecture. These choices are expensive to explore manually and difficult to compare fairly when candidate changes can also alter the FL training or evaluation path. In this work, we present Auto-FL-Research (AFR), a constrained coding-agent workflow for FL algorithmic recipe search. Agents may propose and implement candidate training algorithms, including server aggregation rules, client update schedules, local objectives, and registered model variants, while task profiles fix the mutation surface, compute budget, communication contract, and final model evaluation. Each campaign records candidate scores, runtime, edited files, artifacts, and failure status. We evaluate AFR on five healthcare cross-silo FLamby tasks and on grouped-client profiles for the five fixed LEAF datasets plus the LEAF synthetic task. Five-seed repeat evaluations support gains on four FLamby tasks and five of six LEAF profiles, while also exposing seed-sensitive and search-selected failure cases. Same-budget controls show that several gains correspond to FL-recipe changes, whereas other improvements are recovered by fixed-surface scalar controls or fail under repeat or held-out evaluation. These mixed outcomes are part of the contribution: they show how agent-generated candidates can be separated into repeated FL mechanisms, fixed-surface tuning effects, and selected single-run artifacts.

13:00 JST研究/論文GPT / ChatGPTLlamaMistral AI

効率的な小型言語モデルのための Wiola アーキテクチャ

ここでは、GPT、LLaMA、Mistral、Falcon などの既存のモデル ファミリと構造的な系統を共有せず、第一原理に基づいて構築された完全にオリジナルの Small Language Model (SLM) アーキテクチャである Wiola を紹介します。 Wiola は、5 つの独立した新しいコンポーネントを導入しています。(i) スパイラル回転位置エンコーディング (SRPE)。これは、絶対、相対、および階層位置信号を組み合わせた 3 次元らせん多様体にトークンの位置を埋め込みます。 (ii) ゲート クロスレイヤ アテンション (GCLA)。各デコーダ層に、レイヤ間のコヒーレンスのために先行する 2 つのレイヤの圧縮された要約へのソフト クロス アテンション アクセスを提供します。 (iii) アダプティブ トークン マージング (ATM)。中間ネットワーク層で意味的に冗長な隣接トークンを動的にマージし、情報を損失することなく注意の複雑さを軽減します。 (iv) デュアル ストリーム フィードフォワード (DSFF)。従来の MLP を、学習された次元ごとのゲートによって融合された 2 つの並列ストリームに置き換えます。 (v) WiolaRMSNorm。表現の崩壊を防ぐ、次元ごとに学習されたオフセット ベクトルを導入する修正正規化です。完全な数学的導出、アーキテクチャ ブロック図、複雑さの分析、および GPT-2、LLaMA-2、および Mistral との体系的な比較を提供します。 Wiola は 4 つのサイズ (120M、360M、700M、および 1.5B パラメータ) でリリースされており、HuggingFace Transformers エコシステムと完全に互換性があり、22 のアーキテクチャ単体テストすべてに合格しています。

原文 (English)

The Wiola Architecture for Efficient Small Language Models

We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with any existing model family including GPT, LLaMA, Mistral, or Falcon. Wiola introduces five independently novel components: (i) Spiral Rotary Positional Encoding (SRPE), which embeds token positions on a three-dimensional helical manifold combining absolute, relative, and hierarchical positional signals; (ii) Gated Cross-Layer Attention (GCLA), providing each decoder layer with soft cross-attention access to compressed summaries of two preceding layers for inter-layer coherence; (iii) Adaptive Token Merging (ATM), which dynamically merges se mantically redundant adjacent tokens in middle network layers to reduce attention complexity without information loss; (iv) Dual Stream Feed-Forward (DSFF), replacing the conventional MLP with two parallel streams fused by a learned per-dimension gate; and (v) WiolaRMSNorm, a modified normalisation introducing a per-dimension learned offset vector that prevents representation collapse. We provide complete mathematical derivations, architectural block diagrams, complexity analyses, and systematic comparisons against GPT-2, LLaMA-2, and Mistral. Wiola is released in four sizes (120M, 360M, 700M, and 1.5B parameters) and is fully compatible with the HuggingFace Transformers ecosystem, with all 22 architectural unit tests passing.

13:00 JSTエージェントClaude

Agent4cs: 大規模な階層コードベースでのコード要約のためのマルチエージェント システム

大規模で複雑なコードベース、特に難読化された構造や不完全なドキュメントを持つコードベースを理解することは、依然として大きな課題です。既存のコード要約ソリューションは、単一の言語モデルやクロード コードのようなコーディング アシスタントに依存することが多く、ソース コードをフラット テキストとして扱い、リポジトリ内の豊富な相互依存関係や階層情報が十分に活用されていません。これらの欠点に対処するために、Agent4cs を提案します。Agent4cs は、大規模なコードベースをボトムアップ方式で要約するマルチエージェント フレームワークです。要約エージェントは、堅牢な要約を生成することに重点を置いています。キーワード抽出エージェントは、サブフォルダーから重要な情報を積極的に識別します。そして、品質保証エージェントは、読みやすさ、一貫性、完全性を高めるために出力を繰り返し改良します。 7 つのフロンティア モデルで評価された Agent4cs は、コード セグメントを使用した 2 つの構造化プロンプト ベースラインと比較して、すべてのフォルダー レベルにわたるセマンティックの一貫性を平均 8% 向上させます。さらに、現実世界のデータセットに対する広範な評価により、同じベースラインに対して正規化されたキーワード カバレッジ率が最大 38% 向上することが実証されています。

原文 (English)

Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases

Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Existing code summarization solutions often rely on a single language model or coding assistant like Claude Code, and treat source code as flat text, underutilizing the rich interdependencies and hierarchical information within a repository. To address these shortcomings, we propose Agent4cs - a multi-agent framework that summarizes large codebases in a bottom-up fashion, where a summarization agent focuses on producing robust summaries; a keyword-extraction agent proactively identifies critical information from subfolders; and a quality-assurance agent iteratively refines the outputs for readability, coherence, and completeness. Evaluated on 7 frontier models, Agent4cs improves semantic consistency across all folder levels by average 8% compared to two structured prompting baselines with code segments. Furthermore, extensive evaluation on real-world datasets demonstrates up to 38% gains in normalized keyword coverage rate over the same baselines.

13:00 JSTエージェント

サービスエージェントはいつ再検討すべきでしょうか?カスタマーサービス業務における困難な経路制御

自律型カスタマー サービス エージェントは、会話型インターフェイスから運用実行の役割に移行しており、企業レコードを取得し、サービス ポリシーを適用し、返金、キャンセル、交換、注文の変更、予約の変更などのバックエンドの書き込みを実行します。この変化により、サービス制御の問題が生じます。企業は、顧客の指示、ポリシーの制約、企業の記録、およびバックエンドの書き込みが相互作用するリクエストでの操作エラーを防止しながら、日常的なサービスを高速かつ低摩擦に保つ必要があります。私たちは、サービスエージェントが行動する前にいつ再考すべきかを尋ねる、困難なルーティングのサービス制御アーキテクチャを提案します。軽量ルーターは、日常的なセッションを低コストのベースライン パスに維持し、運用上結合されたセッションをエスカレーションされたワークフローにルーティングします。エスカレートされたパスでは、すべてのサービス セッションにわたって均一に追加の制御を適用するのではなく、競合を認識した通信と書き込みトリガーの再検討を使用して、結果としてバックエンドへの書き込みが行われる前に熟慮と保護策を集中させます。 $\tau^{2}$-bench から人間が検証した小売業と航空会社のタスクに基づいてアーキテクチャを評価します。小売業では、この方法により、運用上の矛盾を伴うサービス要求に対する信頼性が一貫して向上します。ルーティングの証拠は、より強力な制御が日常的なリクエストに広く適用されるのではなく、競合するリクエストに向けられていることを示しています。対話とツールの使用プロファイルは、利益が無差別な相互作用の拡大やより広範なツールチェーンから得られるものではないことを示唆しています。代わりに、追加されたターンとツール呼び出しにより、証拠の収集、書き込みの分離、および書き込み前の再検討がサポートされます。ケースレベルの証拠は、エスカレーションされたワークフローがフォールバック計画を保持し、取得したレコードを正しいアクションにバインドし、書き込みをシーケンスし、複数エンティティのリクエストを分解していることを示しています。航空会社の結果は、同じサービス制御ロジックを予約業務にも拡張します。

原文 (English)

When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations

Autonomous customer-service agents are shifting from conversational interfaces toward operational execution roles: they retrieve firm records, apply service policies, and execute backend writes such as refunds, cancellations, exchanges, order modifications, and reservation changes. This shift creates a service-control problem: firms must keep routine service fast and low-friction while preventing operational errors on requests where customer instructions, policy constraints, firm records, and backend writes interact. We propose a difficulty-routed service-control architecture that asks when service agents should reconsider before acting. A lightweight router keeps routine sessions on a low-cost baseline path and routes operationally coupled sessions to an escalated workflow. The escalated path uses conflict-aware communication and write-triggered reconsideration to concentrate deliberation and safeguards before consequential backend writes, rather than applying additional control uniformly across all service sessions. We evaluate the architecture on human-verified retail and airline tasks from $\tau^{2}$-bench. In retail, the method improves reliability consistently on service requests with operational conflict. Routing evidence shows that stronger control is directed toward conflicted requests rather than broadly applied to routine ones. Dialogue and tool-use profiles suggest that gains do not come from indiscriminate interaction expansion or broader tool chains; instead, added turns and tool calls support evidence gathering, write separation, and pre-write reconsideration. Case-level evidence shows that the escalated workflow preserves fallback plans, binds retrieved records to the correct action, sequences writes, and decomposes multi-entity requests. Airline results extend the same service-control logic to reservation operations.

13:00 JSTLLM/生成AI

CreativityNeuro: 言語モデルの重みを制御して発散的思考を改善し、モード崩壊を減らす

発散的思考は創造性の重要な側面ですが、大規模言語モデル (LLM) は、人工集合意識効果と呼ばれる、自由回答型の質問に対して一貫して同様の応答を生成する傾向があります。ここでは、対照的ウェイト ステアリングを通じて LLM の発散的思考を強化するデータフリーの手法である CreativityNeuro を紹介します。私たちは複数の創造性評価にわたって手法を評価し、いくつかの主な結果を報告します。語彙空間の創造性テストである Divergent Association Task (DAT) では、CreativityNeuro はパフォーマンスを人間のパーセンタイル ポイントで最大 14 向上させます。次に、代替使用テスト (AUT) とタスク タスクに関する大規模な人による評価 (N=720) では、CreativityNeuro は独創性、驚き、創造性の大幅な向上を達成し、より長い形式でより自由なタスクに移行しました。重要なのは、3 つのタスクすべてにわたって、CreativityNeuro がモード崩壊の程度を明らかに軽減していることがわかりました。さらに、アクティベーション ステアリングは、DAT 上で CreativityNeuro と同等のパフォーマンスを達成しますが、AUT およびタスク タスクには転送されず、目に見えないタスクを一般化する際のウェイト スペース ステアリングの有効性を示しています。結論として、CreativityNeuro は、行動データ、再トレーニング、または勾配ベースの微調整を必要とせずに発散的思考を改善し、モード崩壊を軽減し、クリエイティブ ドメインで LLM パフォーマンスを向上させる簡単な方法を提供します。

原文 (English)

CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse

Divergent thinking is a crucial aspect of creativity, yet large language models (LLMs) tend to consistently generate similar responses to open-ended questions, in what has been termed the artificial hivemind effect. Here, we introduce CreativityNeuro, a data-free method for enhancing divergent thinking in LLMs via contrastive weight steering. We evaluate our method across multiple creativity assessments and report several main findings. On the Divergent Association Task (DAT), a vocabulary-space creativity test, CreativityNeuro improves performance by up to 14 human percentile points. Next, in a large-scale human evaluation (N=720) on the Alternative Uses Test (AUT) and the Task Task, CreativityNeuro achieves significant improvements in originality, surprise, and creativity, transferring to longer-form and more open-ended tasks. Importantly, we find that across all three tasks, CreativityNeuro demonstrably reduces measures of mode collapse. Moreover, activation steering achieves comparable performance to CreativityNeuro on the DAT, but it does not transfer to the AUT and Task Task, demonstrating the effectiveness of weight-space steering in generalizing to unseen tasks. In conclusion, CreativityNeuro improves divergent thinking and reduces mode collapse without requiring behavioral data, re-training, or gradient-based fine-tuning, providing a straightforward way to enhance LLM performance in creative domains.

13:00 JST研究/論文Gemma

インタラクティブな放射線医学レポート作成のための離散拡散言語モデル

トークンを左から右に発行するのではなく、双方向にトークン キャンバスのノイズを除去してテキストを生成する拡散言語モデルは、自己回帰 (AR) 生成と競合するようになりました。しかし、医学基盤モデルは、ほぼ完全に自己回帰的なままです。私たちは、専門家混合の拡散言語モデルである DiffusionGemma-26B を適応させ、医療用視覚的質問応答データセット上の同一の LoRA レシピの下で、その同じサイズの AR 兄弟である Gemma-4-26B に対してベンチマークを行い、冗長性の高い LLM 審査員によって採点されました。拡散はそれらすべてで AR と同等かそれを上回っており、微調整されたモデル (3.8B アクティブ) はフロンティアのビジョン言語モデルと競合します。デコードも 3.5 ~ 4.4 倍高速です。この同等性を超えて、拡散モデルは AR にはない製図機能、つまり任意の次数の充填を提供します。キャンバスは双方向にノイズ除去されるため、放射線科医はレポートの断片を修正し、それらの間のテキストをモデルに埋めることができます。これは拡散に固有の操作ですが、自己回帰には固有の操作ではなく、標準以下です。これは、臨床医や施設間で簡潔であったり、一貫性がなかったりすることが多い実際のレポートに適しています。

原文 (English)

Discrete Diffusion Language Models for Interactive Radiology Report Drafting

Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competitive with autoregressive (AR) generation. Medical foundation models, however, remain almost entirely autoregressive. We adapt a mixture-of-experts diffusion language model, DiffusionGemma-26B, and benchmark it against its same-size AR sibling Gemma-4-26B under an identical LoRA recipe on medical visual question answering datasets, scored by a verbosity-robust LLM judge. Diffusion matches or exceeds AR on all of them, and the finetuned model (3.8B active) is competitive with frontier vision-language models; its decoding is also 3.5-4.4x faster. Beyond this parity, the diffusion model offers a drafting capability AR lacks: any-order infill. Because the canvas is denoised bidirectionally, a radiologist can fix report fragments and have the model fill the text between them, an operation inherent to diffusion but not to autoregression, which is subpar at it. This suits real reports, which are often terse or inconsistent across clinicians and institutions.

13:00 JSTエージェント

次のトークン予測を超えて: アトラシアン ワークフロー上のツール使用エージェント向けの RLVR 概念実証

大規模な言語モデルは、特定の API 内で動作するのではなく、次のトークンを予測するようにトレーニングされます。ニッチなエンタープライズ SaaS ワークフロー (成功とは、正しいネストされた引数を正しい順序で使用して正しいエンドポイントに到達することを意味します) では、この目的の不一致は、必須フィールドのドロップ、ツールの幻覚、または 1 回の読み取り後の早期停止など、サイレントな失敗として現れます。ターゲット環境に直接適用される検証可能な報酬を伴う強化学習 (RLVR) がギャップを埋めるかどうかを尋ねます。概念実証として、Jira REST v3 および Confluence v2 API をスキーマ忠実度でエミュレートする 5 つの合成環境のスイートを構築します。報酬はツール呼び出しトレースから完全に計算され、ループ内にライブ API、学習したジャッジ、人間によるラベルはありません。スコアリングにより、GRPO トレーニングを駆動する同じチェッカーで Qwen3-1.7B と Qwen3.5-4B が促されました。報酬が非縮退である 4 つのシナリオでは、RL トレーニングされたポリシーにより平均報酬が 4B ベースラインの範囲 0.35--0.92 から 0.95--1.00 まで上昇し、Confluence ページの作成 ($0.35 \rightarrow) で最大の単一利益が得られることがわかりました。 1.00ドル)。私たちはこれを、ニッチなエンタープライズ API 向けの結果最適化された小規模モデルに向けた準備段階と位置づけ、ワークショップの読者が考慮すべき 2 つの制限を前景化します。それは、検証可能な報酬を手作りすることは、ここで報告されている少数のエンドポイントを超えて拡大できないこと、そして、5 つのシナリオのうちの 1 つ (チケット移行) は、プロンプトされた 4B がすでに最大値に達している飽和した報酬の形状を持っていることです。

原文 (English)

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows

Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means hitting the right endpoint with the right nested arguments in the right order -- this objective mismatch shows up as silent failures: dropped required fields, hallucinated tools, or early stops after a single read. We ask whether Reinforcement Learning with Verifiable Rewards (RLVR), applied directly in the target environment, closes the gap. As a proof of concept we build a suite of five synthetic environments emulating the Jira REST v3 and Confluence v2 APIs at schema fidelity; rewards are computed entirely from the tool-call trace, with no live API, no learned judge, and no human label in the loop. Scoring prompted Qwen3-1.7B and Qwen3.5-4B on the same checkers that drive GRPO training, we find that on the four scenarios whose rewards are non-degenerate the RL-trained policy lifts average reward from a 4B-baseline range of 0.35--0.92 to 0.95--1.00, with the largest single gain on Confluence page creation ($0.35 \rightarrow 1.00$). We position this as a preliminary step toward outcome-optimised small models for niche enterprise APIs, and foreground two limitations a workshop reader should weigh: hand-crafting verifiable rewards does not scale beyond the handful of endpoints reported here, and one of our five scenarios (ticket-transition) has a saturating reward shape that the prompted 4B already maxes out.

13:00 JSTエージェント

臨床薬剤に対する世界のフィードバック: FHIR 環境での RL の診断

臨床プロトコルの実行タスク (検査値のチェック、しきい値の適用、正しく構成された FHIR オーダーの発行) は、世界からのフィードバックから RL の自然な候補です。臨床 SME が意思決定ロジックを検証器にエンコードすると、その検証器はエピソードごとのアノテーションなしで無制限のロールアウトを採点します。ただし、RL を適用するには、サウンド フィードバック チャネルと十分な基本機能が必要です。 MedAgentBench v1/v2 を監査し、不作為を RL の主要な戦略にする 41.7\% のサイレント終了上限を見つけ、\textbf{MedAgentBench-v3 (MAB-v3)} (508 タスク、8.9\% 上限) を構築します。 Qwen3-8B のトレーニングでは、\emph{能力の上限} (10/20 のタスク タイプは基本パフォーマンスが 0\%、勾配ゼロ) と \emph{フォーマット知識の壁} (3/20 のタイプは探索では発見できない正確な臨床コードが必要) という 2 つの構造的な障壁が明らかになります。純粋な RL は 18.2\% pass@1 に達しますが、ルールベースの SFT では \ 34.1\% に達します。 15.9 ~ pp のギャップは完全にこれらの障壁に起因します。決定/フォーマット知識/ルックアップ分類法により、RL の学習可能性が予測され、コードを挿入するための SFT、条件文を学習するための RL という修正が規定されます。

原文 (English)

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments

Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for RL from world feedback: once clinical SMEs encode decision logic into a verifier, that verifier grades unlimited rollouts without per-episode annotation. But applying RL requires a sound feedback channel and sufficient base capability. We audit MedAgentBench v1/v2, find a 41.7\% silent-finish ceiling that makes inaction the RL dominant strategy, and construct \textbf{MedAgentBench-v3 (MAB-v3)} (508 tasks, 8.9\% ceiling). Training Qwen3-8B exposes two structural barriers: a \emph{capability ceiling} (10/20 task types have 0\% base performance, zero gradient) and a \emph{format-knowledge barrier} (3/20 types require exact clinical codes undiscoverable by exploration). Pure RL reaches 18.2\% pass@1 vs.\ 34.1\% for rule-based SFT; the 15.9~pp gap is attributable entirely to these barriers. A decision/format-knowledge/lookup taxonomy predicts RL learnability and prescribes the fix: SFT to inject codes, RL to learn conditionals.

13:00 JST研究/論文

手続き型記憶の蒸留: 自己改善型言語モデルのためのオンライン リフレクション

検証可能な報酬を伴う強化学習 (RLVR) は、SDPO などの最近の自己蒸留バリアントと併せて、検証者に対して各ロールアウトを評価し、そのエピソード レベルの信号からポリシーを更新します。ただし、ロールアウト内の豊富な手順情報が保持されたり再利用されたりすることはほとんどありません。エピソードやエポック全体にわたって、モデルはポリシーの変化に応じて関連する問題に繰り返し遭遇し、エピソードローカル更新では捕捉できないエピソード間のシグナルを生成します。つまり、どの戦略が一貫して検証に合格するか、どの失敗モードが持続するか、どのパターンが再発するかなどです。我々は、これらのクロスエピソード信号を再利用可能な手続き型メモリに変換し、トレーニング中にそれをポリシーの重みに蒸留する手続き型メモリ蒸留 (PMD) を提案します。このメモリはトレーニングの足場として機能し、ポリシー自体に吸収され、推論時にメモリフリーのモデルが生成されます。 PMD は、3 つの抽象化レベルで記憶を整理します。生の軌跡、内省した戦略と教訓、問題全体で繰り返される高レベルの行動パターンであり、すべてモデル自体の軌跡からオンラインで抽出されます。記憶条件付けされたセルフティーチャーは、蓄積された経験を利用して、独自のロールアウトで生徒を監督し、生徒がパラメータ内で手順的な知識を徐々に内面化できるようにします。中心的な設計原則は共進化です。ポリシーはメモリを更新するロールアウトを生成し、メモリはポリシーを更新する監視を形成します。経験的に、Qwen3-8B と OLMo3-Instruct-7B 全体で、PMD は SDPO よりも SCIKNOWEVAL で 3.8 ~ 5.5%、LIVECODEBENCH で 7.9 ~ 13.6% 向上しました。共進化はこれらの利点を強化します。メモリまたはポリシー トレイルのいずれかをフリーズすると、SCIKNOWEVAL ドメイン全体で PMD が 10% 以上減少します。

原文 (English)

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy from that episode-level signal. However, the richer procedural information in the rollout is rarely retained or reused. Across episodes and epochs, the model repeatedly encounters related problems under a changing policy, producing cross-episode signals that episode-local updates cannot capture: which strategies consistently pass verification, which failure modes persist, which patterns recur. We propose Procedural Memory Distillation (PMD), which converts these crossepisode signals into reusable procedural memory and distills it into the policy's weights during training. This memory functions as a training scaffold, absorbed into the policy itself, yielding a memory-free model at inference. PMD organizes the memory at three levels of abstraction: raw trajectories, self-reflected strategies and lessons, and higher-level behavioral patterns that recur across problems, all extracted online from the model's own trajectories. A memory-conditioned self-teacher draws on the accumulated experience to supervise the student on its own rollouts, enabling student to progressively internalize procedural knowledge within its parameters. The central design principle is co-evolution: the policy generates rollouts that update the memory, and memory shapes the supervision that updates the policy. Empirically, across Qwen3-8B and OLMo3-Instruct-7B, PMD improves over SDPO by 3.8-5.5% on SCIKNOWEVAL and 7.9-13.6% on LIVECODEBENCH. Co-evolution powers these gains: freezing either the memory or the policy trails PMD by more than 10% across SCIKNOWEVAL domains.

13:00 JSTエージェント研究/論文

分岐する道のエージェントの庭

実証研究では、独自の分析が認められることはほとんどありません。分析上の選択が異なると、同じデータから異なる結論が得られる可能性がありますが、これらの隠れた分岐経路を観察することは困難です。私たちは、AI エージェントがこれらの経路を明示しながら、人間の研究者間の分析のばらつきの多くを捕捉することを示します。 4 つの一か八かの領域にわたって、異なるペルソナを割り当てるだけで、AI エージェントは、同じデータと質問から異なる、多くの場合相反する結論を報告し、それらの信念と系統的に一致する結果を報告できます。 42 の人間の研究チームが同じ移民データセットを分析した研究では、報告された効果推定値において AI エージェントが人間のイデオロギーのギャップの 72% を再現しました。相反する結論に達したにもかかわらず、AI の最終レポートに基づいて各分析で明確な問題を特定することは困難です。86% が独立した AI レビューに合格し、78% が大多数の人による専門家によるレビューに合格しました。これらの発見は、中心的な課題は多くの場合、欠陥のある分析ではなく、方法論的に防御可能な分析の広い空間から選択的に調査し、報告することであることを示唆しています。 AI エージェントは、そのような探索を安価かつスケーラブルにすることで、この長年の問題を増幅させる可能性があります。これに対処するために、分析パスが報告されたものと少なくとも同じくらい極端な主張を生成する確率である m 値 (多元世界値) を導入します。さらに、AI エージェントを使用して妥当な分析パスをサンプリングすることで m 値を推定する Agentic Bootstrap を紹介します。人間の移民研究に適用すると、報告された人間分析の 13.5% が分析空間の最も極端な 5% に該当しました (m<0.05)。したがって、科学的証拠は、報告された単一の分析によってだけでなく、合理的に報告された可能性のある分析の分布内でのその位置によっても評価される必要があります。 Agentic Bootstrap は、この分布を観察可能にし、科学的信頼性の基準に変えます。

原文 (English)

The Agentic Garden of Forking Paths

Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are difficult to observe. We show that AI agents capture much of the analytical variation among human researchers while making these paths explicit. Across four high-stakes domains, assigning different personas is sufficient for AI agents to report divergent, often opposing, conclusions from the same data and question, with findings systematically aligned with those beliefs. In a study in which 42 human research teams analyzed the same immigration dataset, AI agents reproduced 72% of the human ideological gap in reported effect estimates. Despite reaching opposing conclusions, it is difficult to identify clear issues in each analysis based on the final AI reports: 86% passed independent AI review and 78% passed majority human expert review. These findings suggest that the central challenge is often not flawed analyses, but selective exploration and reporting from a large space of methodologically defensible analyses. AI agents may amplify this longstanding problem by making such exploration inexpensive and scalable. To address this, we introduce the m-value (multiverse value), the probability that an analysis path would produce a claim at least as extreme as the reported one. We further introduce Agentic Bootstrap, which estimates the m-value by using AI agents to sample plausible analysis paths. Applied to the human immigration study, 13.5% of reported human analyses fell in the most extreme 5% of the analysis space (m<0.05). Scientific evidence should therefore be evaluated not only by a single reported analysis but also by its position within the distribution of analyses that could reasonably have been reported. Agentic Bootstrap makes this distribution observable and turns it into a criterion for scientific credibility.

13:00 JSTエージェントビジネス/資金調達

Janus: ユーザーが関与するエージェント権限管理のプレイグラウンド

ユーザーに代わってツール呼び出しを自律的に実行する AI エージェントは、ユーザーがどのような役割を果たせるのか、またどのような役割を果たすべきなのか、権限管理に関する差し迫った疑問を引き起こします。多くのアプローチが提案されているにもかかわらず、エージェントのアクセス許可管理におけるユーザーの役割はまだ検討されていません。ユーザーが関与するエージェント権限管理設計を実装および評価するためのプレイグラウンド システムである Janus を紹介します。 Janus は、さまざまな権限管理設計をサポートするモジュール型エージェント システムである Janus-Core と、自動評価フレームワークである Janus-Harness の 2 つのコンポーネントで構成されています。ユーザー参加のための主要な設計軸を特定する概念モデルに基づいて、設計空間にわたる 6 つの権限アシスタントを実装し、3 つのシナリオと 3 つの合成レスポンダーにわたってそれらを評価します。私たちは、ユーザー入力が重要でプライバシーとセキュリティを大幅に強化できること、ユーザーの意思決定の AI 拡張が認知負荷の軽減に役立つこと、システム設計では許可疲労を含む現実的なユーザーの行動を考慮する必要があることを実証します。すべてのコンテキストにわたって最適に機能する単一の設計は存在しないため、エージェント システムにパーミッション アシスタントを展開するための、より原則に基づいたコンテキスト依存のアプローチが推進されます。 Janus は、エージェント システム設計のこの側面に関する今後の調査をサポートするために一般に公開されています。

原文 (English)

Janus: a Playground for User-Involved Agentic Permission Management

AI agents that autonomously execute tool calls on a user's behalf raise pressing questions about permission management: what role could users play, and what role should they play? Despite many proposed approaches, the user's role in agentic permission management remains under explored. We introduce Janus, a playground system for implementing and evaluating user-involved agentic permission management designs. Janus consists of two components: Janus-Core, a modular agentic system supporting a diverse spectrum of permission management designs, and Janus-Harness, an automated evaluation framework. Grounded in a conceptual model that identifies key design axes for user involvement, we implement six permission assistants spanning the design space and evaluate them across three scenarios and three synthetic responders. We demonstrate that user input is critical and can significantly strengthen privacy and security, that AI augmentation of user decisions can help reduce cognitive load, and that realistic user behavior including permission fatigue must be accounted for in system design. No single design performs optimally across all contexts, motivating a more principled and context-sensitive approach to deploying permission assistants in agentic systems. Janus is publicly available to support future investigation into this dimension of agentic system design.

13:00 JST研究/論文

限定された監督下での思考連鎖推論の再考: 半教師あり思考連鎖学習

思考連鎖 (CoT) 推論は、大規模な言語モデルの潜在的な推論能力を活性化するための効果的なアプローチとして浮上しました。ただし、既存の CoT 手法のほとんどは推論チェーンを主に推論時のプロンプトとして使用し、生成された推論トレースが半教師あり学習信号として再利用されることはほとんどありません。このレポートでは、\textbf{半教師あり思考連鎖学習}を定義し、ラベルなしの質問を使用して疑似推論監視を構築する単純なフレームワークである\textbf{Semi-CoT}を提案します。 Semi-CoT は、ラベルのない質問ごとに複数の擬似 CoT をサンプリングし、回答レベルの意味エントロピーを推定し、信頼できる擬似 CoT のデモンストレーションとして低エントロピーの推論チェーンを選択します。これにより、CoT の自己トレーニングの観点が、推論時の改良から半教師ありの擬似教師まで拡張されます。 AQuA、SVAMP、GSM8K、および MultiArith でのパイロット実験では、エントロピー ゲートが $91.36\%$ から $100\%$ の範囲の疑似応答精度を持つ高精度の疑似 CoT を選択することが示されています。 Semi-CoT も SVAMP と GSM8K でわずかな利益をもたらしますが、AQuA は負の転送を示し、MultiArith は上限に達します。これらの結果は、ラベルのない質問は信頼できる疑似推論シグナルを提供できるが、その効果的な使用には依然として強力なデモンストレーションの選択または学生のトレーニングが必要であることを示唆しています。

原文 (English)

Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning

Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent reasoning capabilities in large language models. However, most existing CoT methods use reasoning chains mainly as inference-time prompts, while the generated reasoning traces are rarely reused as semi-supervised learning signals. In this report, we define \textbf{Semi-supervised Chain-of-Thought Learning} and propose \textbf{Semi-CoT}, a simple framework that uses unlabeled questions to construct pseudo reasoning supervision. Semi-CoT samples multiple pseudo-CoTs for each unlabeled question, estimates answer-level semantic entropy, and selects low-entropy reasoning chains as reliable pseudo-CoT demonstrations. This extends the self-training view of CoT from inference-time refinement to semi-supervised pseudo-supervision. Pilot experiments on AQuA, SVAMP, GSM8K, and MultiArith show that the entropy gate selects high-precision pseudo-CoTs, with pseudo-answer precision ranging from $91.36\%$ to $100\%$. Semi-CoT also gives small gains on SVAMP and GSM8K, while AQuA shows negative transfer and MultiArith reaches a ceiling. These results suggest that unlabeled questions can provide reliable pseudo reasoning signals, but their effective use still requires stronger demonstration selection or student training.

13:00 JSTエージェント

OPINE-World: オントロジーエラー優先の対話型探索によるプログラムによる世界モデリング

インタラクションから環境がどのように動作するかを学ぶことは、不慣れなタスクに適応するエージェントを構築する上で中心となります。ディープネットワークで学習されたワールドモデルは柔軟性がありますが、データを大量に消費し、トレーニング分布を超えて転送することは困難です。 LLM によってソース コードとして記述され、反例誘導帰納合成 (CEGIS) によって洗練されたプログラム合成ワールド モデルは、データ効率が高く再利用可能ですが、主に特定のオブジェクト語彙を持つ構造化状態の世界で実証されており、単一のプログラム検索では、オブジェクト構造を柔軟に仮説する必要があるピクセル レンダリング環境には対応できません。インタラクションからオンラインでオブジェクト中心のプログラム世界モデルを学習する LLM エージェントである OPINE-World を紹介します。 OPINE-World は、仮説とテストのループで 2 つの協力するエージェントを結合します。1 つは環境内で動作し、もう 1 つはリプレイ検証とモデルベースの計画を使用してコードでモデルを合成します。また、オントロジー エラーと呼ばれるオブジェクト タイプの適切性のベイジアン尺度を使用して探索を制御します。 OPINE-World を ARC-AGI-3 で評価します。これは、オブジェクトの語彙、目標、およびアクションのセマンティクスが保留されたスキル習得効率のベンチマークです。 OPINE-World は、ゲームごとのトレーニングなしで 25 ゲーム中 20 ゲームを解決し、人間のベースラインと比較して 78.4 のアクション効率スコアに達しました。

原文 (English)

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networks are flexible but data-hungry and transfer poorly beyond their training distribution. Program-synthesized world models, written as source code by LLMs and refined through counterexample-guided inductive synthesis (CEGIS), are instead data-efficient and reusable, yet they have been demonstrated mainly on structured-state worlds with a given object vocabulary, and a single program search does not scale to pixel-rendered environments whose object structure must be hypothesized flexibly. We introduce OPINE-World, an LLM agent that learns an object-centric programmatic world model online from interaction. OPINE-World couples two cooperating agents in a loop of hypothesis and test, one acting in the environment and one synthesizing the model in code with replay verification and model-based planning, and it steers exploration with a Bayesian measure of object-type adequacy we call ontology error. We evaluate OPINE-World on ARC-AGI-3, a benchmark for skill-acquisition efficiency in which the object vocabulary, the goal, and the action semantics are withheld. OPINE-World solves 20 of 25 games without per-game training and reaches an action-efficiency score of 78.4 against the human baseline.

13:00 JSTLLM/生成AI

嗜好学習における嘘発見器の監視のスケーリング傾向

LLM における欺瞞的な行為は監視と防止にコストがかかるため、嘘発見器を使用して高コストのラベラーによるレビューのための対応を特定する、嘘発見器によるスケーラブル監視 (SOLiD) (Cundy & Gleave、2025) などのアプローチが動機付けられています。このペーパーでは、SOLiD をより大きなモデルにスケールし、より多様で現実的な好み学習設定で評価します。我々は、良好なスケーリングを発見しました。検出器の真陽性率 99% で、未検出の欺瞞が 1B パラメーター モデルの 34% から 405B パラメーター モデルの 14% に低下し、統計的に有意な欺瞞の増加を伴うことなく、高価な人間によるラベラーを微調整段階から完全に排除できます。ただし、SOLiD は検出器のトレーニング データと優先トレーニング データの間の分布シフトの影響を受けやすいため、検出器の誤検出率が非現実的なレベルに達する可能性があります。

原文 (English)

Scaling Trends for Lie Detector Oversight in Preference Learning

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings. We find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the fine-tuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPT

EO-Agents: 地球観測仮説生成のための 3 つのエージェント LLM パイプライン

最近、科学的仮説生成のために大規模な言語モデルが研究されていますが、これまでの研究のほとんどは非構造化文献や自由形式の文章による主張に依存しています。 NASA 地球観測ナレッジ グラフで直接仮説生成の根拠となる地球観測用のパイプラインを紹介します。過去の共同利用関係に基づいてトレーニングされた異種グラフ ニューラル ネットワークが候補データセットのペアをランク付けし、3 エージェントの LLM パイプラインが構造化された研究仮説をフィルタリング、生成、評価します。このシステムは 1,475 の NASA データセットに適用され、生態水文学、氷河学、エアロゾルと雲の相互作用、植生季節学、成層圏化学など、複数の地球科学領域にわたる 160 の仮説を生成します。モデルによって予測された新しいデータセットの組み合わせは、文献から示されている実際の共用とほぼ同じくらいもっともらしいと評価されており、パイプラインが科学的に一貫性があるが未調査の組み合わせを明らかにしていることを示しています。 GPT-5.2 と Claude Sonnet 4.6 にわたる 2*2*2 要因実験では、仮説のランキングが安定している一方で、絶対スコアが裁判官のアイデンティティに大きく依存することが示され、単一裁判官による LLM 評価の限界が浮き彫りになっています。

原文 (English)

EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation

Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-form textual claims. We present a pipeline for Earth observation that grounds hypothesis generation directly in the NASA Earth Observation Knowledge Graph. A heterogeneous graph neural network trained on historical co-usage relations ranks candidate dataset pairings, and a three-agent LLM pipeline filters, generates, and evaluates structured research hypotheses. Applied to 1,475 NASA datasets, the system produces 160 hypotheses spanning multiple Earth-science domains, including ecohydrology, glaciology, aerosol--cloud interactions, vegetation phenology, and stratospheric chemistry. Model-predicted novel dataset pairings are rated nearly as plausible as held-out real co-usages from the literature, indicating that the pipeline surfaces scientifically coherent yet unexplored combinations. A 2*2*2 factorial experiment across GPT-5.2 and Claude Sonnet 4.6 shows that hypothesis rankings remain stable, while absolute scores depend strongly on judge identity, highlighting limitations of single-judge LLM evaluation.

13:00 JST研究/論文Claude

Hawk: ハードウェアを意識した知識を活用して高性能 NPU カーネルを生成する

Neural Processing Unit (NPU) 用の高性能カーネルの開発は業界の重大なボトルネックであり、開発者は暗黙のハードウェア制約と厳密なメモリ階層を手動でナビゲートする必要があります。大規模な言語モデルには計り知れない自動化の可能性がありますが、ハードウェア固有の事前条件が根本的に欠如しているため、NPU では壊滅的に失敗します。同様の NPU カーネルから単純にコード スニペットを移植すると、コンパイラは通過する可能性がありますが、根底にあるハードウェア制約に盲目的に違反することにより、ランタイム クラッシュやパフォーマンスの低下が常に引き起こされます。これを克服するために、3 つのコア モジュールを通じてハードウェア認識の知識を活用する、トレーニング不要のフレームワークである Hawk を導入します。(1) 実行時知識合成モジュール。これは、3 部構成の実行可能知識表現を採用して、エラー コンテキストと実行可能セマンティクスを本質的に結合します。 (2) ボトルネック認識知識検索モジュール。クエリを直交構文およびハードウェアに合わせた意味空間に投影する 2D 検索パラダイムを実装します。 (3) エフェクト駆動型の知識抽出モジュール。LLM 駆動のセマンティック アービトレーションを利用して、経験的な実行フィードバックに基づいてエラーを枝刈りし、冗長性を統合することで、継続的に知識を抽出します。実際の NPU ワークロードに関する広範な評価により、Hawk が生成精度を 49.4% から 80.0% に向上させながら、最先端のベースラインと比較して最大 2.2 倍の実行速度向上を達成することが実証されました。

原文 (English)

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate implicit hardware constraints and strict memory hierarchies. While large language models offer immense automation potential, they fail catastrophically on NPUs due to a fundamental lack of hardware-specific priors. Naively transplanting code snippets from similar NPU kernels may pass the compiler, but it consistently triggers runtime crashes and performance degradation by blindly violating underlying hardware constraints. To overcome this, we introduce Hawk, a training-free framework that harnesses hardware-aware knowledge through three core modules: (1) Run-Time Knowledge Synthesis Module, which employs a Triple-Part Executable Knowledge Representation to inherently couple the error context with executable semantics; (2) Bottleneck-Aware Knowledge Retrieval Module, which implements a 2D-Retrieval paradigm to project queries into orthogonal syntactic and hardware-aligned semantic spaces; and (3) Effect-Driven Knowledge Distillation Module, which leverages LLM-driven semantic arbitration to continuously distill the knowledge by pruning errors and consolidating redundancies based on the empirical execution feedback. Extensive evaluations on real-world NPU workloads demonstrate that Hawk elevates generation accuracy from 49.4% to 80.0%, while achieving up to a 2.2x execution speedup over state-of-the-art baselines.

13:00 JSTLLM/生成AI

安全で適応的なクラウド ヒーリング: LLM によって生成された復旧計画をニューラルシンボリック ワールド モデルで検証する

クラウドベースの AI システムの規模と複雑さが増大し続けるにつれ、迅速な障害検出と適応回復を通じてサービスの信頼性を確保することが重要な課題となっています。既存のアプローチでは、意味理解のための大規模言語モデル (LLM) とポリシー最適化のための深層強化学習 (DRL) が統合されていますが、多くの場合、LLM の生成機能と推論機能が十分に活用されていない、疎結合の逐次アーキテクチャに依存しています。この論文では、計画認識セマンティック自己修復エンジンである PASE によるパラダイム シフトを提案します。これは、回復を神経記号的なプログラム合成タスクとして再概念化する新しい障害自己修復フレームワークです。 PASE は、コアの計画合成エンジンとして LLM を採用し、セマンティック プリミティブのライブラリから構造化された復旧計画を生成します。ニューラルシンボリック ワールド モデルはシミュレーションを通じて計画の実現可能性を検証し、DRL によってトレーニングされたメタプロンプト オプティマイザーは、LLM の計画プロセスをガイドする最適なプロンプトを生成する方法を学習します。この緊密な理由、計画、検証、適応ループにより、事前定義されたアクション スペースを超えて、動的でコンテキストを認識した回復戦略の生成が可能になります。実際のクラウド障害挿入データセットでの実験では、PASE が最先端の手法を大幅に上回っており、平均システム回復時間を 40% 以上短縮し、未知の障害シナリオにおける障害検出精度を向上させていることが実証されています。私たちのフレームワークは、LLM ベースの推論をモデル支援検証およびメタ学習ガイダンスと統合することにより、自律システム管理を前進させます。

原文 (English)

Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model

As the scale and complexity of cloud-based AI systems continue to escalate, ensuring service reliability through rapid fault detection and adaptive recovery has become a critical challenge. While existing approaches integrate Large Language Models (LLMs) for semantic understanding and Deep Reinforcement Learning (DRL) for policy optimization, they often rely on sequential, loosely coupled architectures that underutilize the generative and reasoning capabilities of LLMs. In this paper, we propose a paradigm shift with PASE, a Planning-Aware Semantic self-healing engine, a novel fault self-healing framework that reconceptualizes recovery as a neuro-symbolic program synthesis task. PASE employs an LLM as a core Plan Synthesis Engine to generate structured recovery plans from a library of semantic primitives. A Neural-Symbolic World Model verifies plan feasibility through simulation, while a Meta-Prompt Optimizer, trained via DRL, learns to generate optimal prompts that guide the LLM's planning process. This tight reason-plan-verify-adapt loop enables dynamic, context-aware recovery strategy generation beyond predefined action spaces. Experiments on a real-world cloud fault injection dataset demonstrate that PASE significantly outperforms state-of-the-art methods, reducing average system recovery time by over 40% and improving fault detection accuracy in unknown fault scenarios. Our framework advances autonomous system management by unifying LLM-based reasoning with model-assisted verification and meta-learned guidance.

13:00 JSTLLM/生成AI

SemHash-LLM: ドキュメント重複排除のための複数粒度のセマンティック ハッシュ フレームワーク

大規模なドキュメントの重複排除では、大量のコーパスに対して効率を維持しながら、意味上の同等性を維持する必要があります。セマンティック射影ハッシュ、注意加重 MinHash、対比境界学習、および選択的 LLM ベースの判定を統合する多重粒度フレームワークである SemHash LLM を紹介します。この方法では、ゲート フュージョンを通じて文字、トークン、ドキュメント レベルの信号を結合し、カスケード フィルタリング パイプラインを適用して効率的に候補を削減します。セマンティック射影ハッシュは、蒸留された LLM 埋め込み空間でコンパクトなバイナリ コードを学習しますが、注意を重み付けした Min-Hash はボイラープレートを抑制し、有益なコンテンツを強調します。適応的な決定境界と不確実性の推定により、テンプレートの汚染、短いテキストの乱れ、封じ込め、ウイルスの断片に対する堅牢性がさらに向上します。実験によれば、SemHash LLM は 1% 未満のニューラル検証コストで強力な重複検出品質を実現します。

原文 (English)

SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication

Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora. We present SemHash LLM, a multi granularity framework that unifies semantic projection hashing, attention weighted MinHash, contrastive boundary learning, and selective LLM based adjudication. The method combines character, token, and document level signals through gated fusion, then applies a cascaded filtering pipeline for efficient candidate reduction. Semantic projection hashing learns compact binary codes in distilled LLM embedding space, while attention weighted Min- Hash suppresses boilerplate and emphasizes informative content. Adaptive decision boundaries and uncertainty estimation further improve robustness across template pollution, short text perturbation, containment, and viral fragments. Experiments show that SemHash LLM achieves strong duplicate detection quality with less than one percent neural verification cost.

13:00 JST研究/論文

商品改良のための利益ベースの反事実説明:日本の漫画販売の事例研究

反事実的説明 (CE) は、機械学習モデルの解釈可能性を高め、モデル予測に基づいたデータ駆動型の意思決定をサポートするために広く使用されています。ただし、既存の CE 手法では通常、望ましい出力値 (ターゲット) と説明変数の変化を定量化する距離関数という 2 つの外生的に指定された入力が必要です。回帰設定では、ターゲット仕様の妥当性も距離メトリックの実際的な解釈も十分に取り上げられていません。さらに、現実世界の意思決定では明示的な目的の最大化が必要な場合が多いにもかかわらず、既存の CE 手法のほとんどは、意思決定目的の最適化ではなく予測の変更に重点を置いています。これらの限界に対処するために、我々は CE を経営およびマーケティングの文脈における利益最大化の問題として定式化し、利益に基づく反事実の説明 (PBCE) と呼ばれるフレームワークを提案します。 PBCE は、主な最適化目標として利益を直接最大化することにより、外生的なターゲット仕様の必要性を排除します。同時に、距離の項は製品の属性を変更するコストとして再解釈され、明確で経済的に根拠のある解釈が提供されます。

原文 (English)

Profit-Based Counterfactual Explanations for Product Improvement: A Case Study of Manga Sales in Japan

Counterfactual explanation (CE) is widely used to enhance the interpretability of machine learning models and support data-driven decision-making based on model predictions. However, existing CE methods typically require two exogenously specified inputs: a desired output value (target) and a distance function that quantifies changes in explanatory variables. In regression settings, neither the validity of target specification nor the practical interpretation of the distance metric has been sufficiently addressed. Furthermore, most existing CE methods focus on altering predictions rather than optimizing a decision objective, even though real-world decision-making often requires explicit objective maximization. To address these limitations, we formulate CE as a profit maximization problem in management and marketing contexts and propose a framework termed profit-based counterfactual explanation (PBCE). PBCE eliminates the need for exogenous target specification by directly maximizing profit as the primary optimization objective. Concurrently, the distance term is reinterpreted as the cost of modifying product attributes, providing a clear and economically grounded interpretation.

13:00 JSTLLM/生成AI

自信を持ってスケーリング: 適応テスト時間スケーリングのための LLM の信頼性の調整

強化学習 (RL) を使用して大規模言語モデル (LLM) をトレーニングすると、推論タスクや質問応答タスクのパフォーマンスが大幅に向上しました。ただし、一般的な RL 報酬設計は通常、応答の正確さを優先し、モデルに自信を正確に表現するよう促すことを無視しています。これは重大な問題につながります。パフォーマンスの向上には、信頼性と精度の間の不十分な調整が伴うことが多く、不確実な場合にモデルが自信過剰になり幻覚を引き起こすという誤解を招くことになります。この制限に対処するために、$\textbf{C}$orrectness と $\textbf{C}$onfidence $\textbf{C}$alibration $\textbf{R}$einforcement $\textbf{L}$earning ($\textbf{C3RL}$) を提案します。これは、正確性、キャリブレーション、およびデータセットに基づく参照精度の報酬を統合した新しい RL アルゴリズムです。 8 つのテキストおよびマルチモーダル データセットにわたる包括的な評価により、C3RL が精度を犠牲にすることなくキャリブレーションを強化し、パフォーマンスとキャリブレーション メトリクスの両方で現在の最先端の方法を上回るパフォーマンスを示していることが実証されました。 C3RL からの適切に調整された言語化された信頼度を利用して、応答の信頼度に基づいて計算リソースを割り当てる調整可能な推論時間戦略である $\textbf{C}$onfidence ベースの $\textbf{A}$daptive Test Time $\textbf{S}$caling ($\textbf{CAS}$) をさらに導入します。実験の結果、CAS は推論予算を最大 12.33 分の 1 に削減しながら、ドメイン内データセットとドメイン外データセットの両方で過半数投票を上回りました。私たちは、C3RL と CAS の相乗効果により、より信頼性が高くリソース効率の高い LLM を展開する道が開かれると信じています。コード、データ、モデルは公開されます。

原文 (English)

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing RL reward designs typically prioritize response correctness, neglecting to incentivize models to express their confidence accurately. This leads to a critical problem: performance gains are often accompanied by poor calibration between confidence and accuracy, misleading models to overconfidently hallucinate when uncertain. To address this limitation, we propose $\textbf{C}$orrectness and $\textbf{C}$onfidence $\textbf{C}$alibration $\textbf{R}$einforcement $\textbf{L}$earning ($\textbf{C3RL}$), a novel RL algorithm integrating correctness, calibration and dataset-informed reference accuracy rewards together. Comprehensive evaluation across 8 text and multimodal datasets demonstrates that C3RL enhances calibration without sacrificing accuracy, outperforming the current state-of-the-art method in both performance and calibration metrics. Utilizing the well-calibrated verbalized confidence from C3RL, we further introduce $\textbf{C}$onfidence-based $\textbf{A}$daptive Test Time $\textbf{S}$caling ($\textbf{CAS}$), an adjustable inference-time strategy that allocates computational resources based on response confidence. Experiments show that CAS surpasses majority voting on both in-domain and out-of-domain datasets while reducing the inference budget by up to 12.33 times. We believe the synergy of C3RL and CAS paves the way for deploying more reliable and resource-efficient LLMs. The code, data and models will be released.

13:00 JST研究/論文

空間サポートが重要: 降雨フィールド再構築のためのジオメトリを意識したグラフ融合

都市洪水モデリングには詳細な降雨量の再構築が重要ですが、実際の降雨感知システムは、互換性のない空間サポートを通じて現場を観測します。つまり、ゲージは地点を測定し、マイクロ波リンクは経路を測定し、レーダー/衛星製品はグリッド領域を測定します。測定サポートのこれらの違いは、降雨フィールドに幾何学的に異なる制約を課しますが、既存の異種グラフ アプローチは、そのようなソースを特徴空間で調整し、サポートのジオメトリを破棄しながらそれぞれに独自の埋め込みを与えます。私たちは、サポート タイプ (0D ポイント、1D ライン、または 2D グリッド) に応じて各観測を個別のノード層として表現し、フィールドが再構築されるポイントサポート予測層に渡されるクロスサポート メッセージを通じてそれらを融合する、ジオメトリを認識したマルチサポート異種グラフ ニューラル ネットワークを提案します。帰納的マスクノード定式化により、予測解像度がセンシング解像度から切り離され、同じトレーニング済みモデルでユーザー定義のターゲット位置または表示グリッドでフィールドを再構築できるようになります。シンガポールのデータでは、提案された方法は古典的な内挿ベースライン、逆距離重み付けよりも RMSE を 23.2% 削減し、畳み込み融合やサポートに依存しない異種グラフ ベースラインなどの他のニューラル アーキテクチャよりも一貫して優れています。オーストラリアのシドニーのデータを使用した一般化研究により、マルチサポート フュージョンがどのような場合に役立つかを特徴付けることができます。利用可能なスキルはフィールドの空間相関長に対するゲージ間隔に依存するようです。したがって、フュージョンは、フィールドが相関長に対してアンダーサンプリングされている場合に最大のゲインをもたらし、既に解決されている場合にはほとんどゲインをもたらしません。コードとモデルは、書類が受理され次第オープンソース化されます。

原文 (English)

Spatial Support Matters: Geometry-Aware Graph Fusion for Rainfall Field Reconstruction

Fine-scale rainfall reconstruction is critical for urban flood modeling, but real rainfall sensing systems observe the field through incompatible spatial supports: gauges measure points, microwave links measure paths, and radar/satellite products measure gridded areas. These differences in measurement support impose geometrically distinct constraints on the rainfall field, yet existing heterogeneous graph approaches reconcile such sources in feature space, giving each its own embedding while discarding the geometry of its support. We propose a geometry-aware multi-support heterogeneous graph neural network that represents each observation according to its support type (0D point, 1D line, or 2D grid) as a distinct node layer, and fuses them through cross-support message passing into a point-support prediction layer from which the field is reconstructed. An inductive masked-node formulation decouples prediction resolution from sensing resolution, allowing the same trained model to reconstruct the field at user-defined target locations or display grids. On Singapore data, the proposed method reduces RMSE by 23.2\% over the classical interpolation baseline, inverse-distance weighting, and consistently outperforms other neural architectures such as convolutional fusion and support-agnostic heterogeneous graph baselines. A generalization study using data from Sydney, Australia lets us characterize when multi-support fusion helps: the available skill appears to depend on gauge spacing relative to the spatial correlation length of the field, so fusion delivers the largest gains where the field is under-sampled relative to its correlation length and little when it is already resolved. Code and models will be open-sourced upon paper acceptance.

13:00 JSTエージェント

AI 交通科学者による自律的な交通法規の発見

普遍的な交通法は、都市全体の渋滞、移動性、運転行動の繰り返しのパターンを記述しており、交通計画、管理、制御に科学的根拠を提供します。しかし、彼らの発見は依然として専門家主導であり、規則性の候補を異種の観察証拠から特定するか、介入実験を通じて検証する必要があります。自律型人工知能 (AI) システムは、制御された実験室環境での科学的発見を進歩させましたが、それを複雑な輸送ドメインに拡張することは依然として課題です。ここでは、証拠の絞り込み、批評家と裁判官の仮説誘導、および観察と介入の検証を統合した、反復的で監査可能なワークフローとして交通法の発見を定式化するエージェント AI システムである TrafficSci を紹介します。 TrafficSci は、人口、ネットワーク、制御、および軌道スケールにわたる 4 つのケーススタディを通じて、確立された 3 つの交通法規を自律的に再発見し、報告されていない都市部の運転行動における固有の時間記憶スケールを特定します。これは、8 つの都市と 2 つの軌道データセットにわたって統計的に一貫しています。 TrafficSci は、AI による科学的発見を制御領域から複雑な都市システムに拡張するためのルートを提供します。

原文 (English)

Autonomous discovery of traffic laws with AI traffic scientists

Universal traffic laws describe recurrent patterns in congestion, mobility and driving behavior across cities, providing a scientific basis for transportation planning, management and control. Their discovery, however, remains expert-driven, requiring candidate regularities to be identified from heterogeneous observational evidence or validated through intervention experiments. Although autonomous artificial intelligence (AI) systems have advanced scientific discovery in controlled laboratory settings, extending them to complex transportation domains remains a challenge. Here we present TrafficSci, an agentic AI system that formulates traffic-law discovery as an iterative, auditable workflow integrating evidence scoping, critic-judge hypothesis induction, and observational-interventional validation. Across four case studies spanning population, network, control and trajectory scales, TrafficSci autonomously rediscovers three established traffic laws and identifies an unreported intrinsic temporal memory scale in urban driving behavior, statistically consistent across eight cities and two trajectory datasets. TrafficSci provides a route for extending AI-driven scientific discovery from controlled domains to complex urban systems.

13:00 JSTLLM/生成AIエージェント

多様な証拠、より良い予測: 情報の非対称性の下でのマルチエージェントの審議

複数の LLM 間の審議により推論と調整が向上すると考えられているため、マルチエージェント システムは将来の出来事を予測するためにますます使用されています。しかし、既存のアプローチでは、各エージェントがどのような情報を受け取るかという重要な設計上の選択が見落とされています。すべてのエージェントに同一の証拠が与えられると、熟慮は真の信念修正ではなく群れに陥り、マルチエージェントシステムは単一エージェントよりもわずかに優れたものになります。私たちはこれを根本的なギャップとして特定し、それを埋めるために設計された情報の非対称性を提案します。つまり、証拠を共有の公開サブセットと独立したプライベートサブセットに分割することで、各エージェントは熟慮を通じてのみ他のエージェントに到達できる排他的な知識を保持します。我々は、この分解によりエージェント間のエラー相関が減少することを理論的に示し、関連性を意識した証拠ルーティング、理論的根拠に基づく反復審議、信頼度加重集計を組み合わせたフレームワークである InfoDelphi でそれをインスタンス化します。現実世界の予測市場から得られた 375 のバイナリ予測質問のベンチマークである PolyGym では、InfoDelphi は最も強力な単一エージェントおよび複数エージェントのベースラインをブライアー スコアで 12 ~ 18%、精度で 4 ~ 8 パーセントポイント上回りました。より詳細な実験により、情報の非対称性を取り除くと熟慮による利益のほとんどが失われ、入力の多様性が効果的なマルチエージェント推論を可能にする重要な要因であることが確認されています。

原文 (English)

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry

Multi-agent systems are increasingly used for forecasting future events, as deliberation among multiple LLMs is believed to improve reasoning and calibration. Yet existing approaches overlook a critical design choice: what information each agent receives. When all agents are given identical evidence, deliberation collapses into herding rather than genuine belief revision, leaving multi-agent systems little better than a single agent. We identify this as a fundamental gap and propose designed information asymmetry to close it: by partitioning evidence into shared public and disjoint private subsets, each agent holds exclusive knowledge that can only reach others through deliberation. We theoretically show that this decomposition reduces inter-agent error correlation, and instantiate it in InfoDelphi, a framework combining relevance-aware evidence routing, rationale-based iterative deliberation, and confidence-weighted aggregation. On PolyGym, a benchmark of 375 binary forecasting questions derived from real-world prediction markets, InfoDelphi outperforms the strongest single-agent and multi-agent baselines by 12--18% in Brier score and 4--8 percentage points in accuracy. More detailed experiments confirm that removing information asymmetry eliminates most deliberation gains, establishing diversity of input as the key enabler of effective multi-agent reasoning.

13:00 JSTエージェント

Raw-ECG-Replay-Free 継続的 ECG 導入における自律的なソース推論から専門家の保持を分離する

マルチソース ECG 展開では、以前の生の ECG を保持または再生できない場合、モデルに新しいデータ ソースを組み込む必要がある場合があります。事前トレーニングされたバックボーンをフリーズし、各ソースに分離された分類子を割り当てることでパラメータの干渉を防ぐことができますが、ソースのメタデータが利用できない場合でも展開には専門家を選択する必要があります。私たちは、凍結された 1024 次元の ECGFounder 機能に基づいて構築された増分エキスパート バンクである \ours{} を通じて、この違いを研究しています。到着するドメインごとにバランスの取れたソフトマックス線形エキスパートが追加されますが、軽量ルーターは、これまでに観察されたソースからの保持されたトレーニング特徴とドメイン ラベルにのみ適合します。検証によって調整されたマージン ルールは、単一のルーティングされたエキスパートにコミットするのではなく、最も可能性の高い 2 人のエキスパートを融合します。 CPSC、PTB-XL、ジョージア、およびチャップマン-紹興では、ソースを意識した専門家の選択は $0.7915\pm0.0036$ マクロ F1 に達し、一致するオフラインの独立ヘッド参照は $0.7885\pm0.0009$ に達し、強力なソースを認識した専門家の保持をサポートしています。ソース ID がない場合、MLP ルーターは $0.7756\pm0.0027$ に達し、トップ 2 マージン フュージョンは $0.7782\pm0.0022$ に達します。ハード MLP ルーティングに対する上位 2 のゲインは小さく ($+0.0026$)、ペアのブートストラップからの 95\% 信頼区間にはゼロが含まれます。 3 つのドメイン注文全体で、トップ 2 とオラクルの差は依然として 0.0111 ドルから 0.0133 ドルであり、自律的なソース推論が残りの主なボトルネックであることがわかります。生の ECG は再生されませんが、凍結されたトレーニング機能はルーターの更新のために保持されます。したがって、このメソッドはメモリフリーではありません。

原文 (English)

Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment

In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freezing a pretrained backbone and assigning each source an isolated classifier prevents parameter interference, but deployment still requires selecting an expert when source metadata are unavailable. We study this distinction through \ours{}, an incremental expert bank built on frozen 1024-dimensional ECGFounder features. Each arriving domain adds a balanced-softmax linear expert, while a lightweight router is fitted only on retained training features and domain labels from sources observed so far. A validation-calibrated margin rule fuses the two most likely experts instead of committing to a single routed expert. On CPSC, PTB-XL, Georgia, and Chapman-Shaoxing, source-aware expert selection reaches $0.7915\pm0.0036$ Macro-F1 and a matched offline independent-head reference reaches $0.7885\pm0.0009$, supporting strong source-aware expert retention. Without source IDs, an MLP router reaches $0.7756\pm0.0027$ and top-2 margin fusion reaches $0.7782\pm0.0022$. The top-2 gain over hard MLP routing is small ($+0.0026$), with a 95\% confidence interval from paired bootstrap that includes zero. Across three domain orders, the top-2-to-oracle gap remains $0.0111$--$0.0133$, identifying autonomous source inference as the main remaining bottleneck. No raw ECGs are replayed, but frozen training features are retained for router updates; the method is therefore not memory-free.

13:00 JSTLLM/生成AI

エピステミック ゴーグル: グラデーション編集を通じてエピステミック フレームを誘導する事前トレーニング済みモジュール

明示的にフィクションとして注釈が付けられた文書の言語モデルを微調整すると、文書の中核となる主張を依然として実際に信じるモデルが生成されます。これは否定無視として知られる効果です。私たちの評価では、そのような注釈が接頭辞および接尾辞として付けられた文書でトレーニングされたモデルが、関連する主張をフィクションであると正しく識別する確率はわずか約 9% でした。これに対処するために、データではなく微調整勾配に介入する学習モジュールであるゴーグルを導入します。教師あり微調整中、ゴーグル モジュールは LLM LoRA が受け取る勾配を編集し、選択された認識フレーム (モデルが読み取ったものの性質に対してとるスタンス) を文書が教えるものすべてに与えます。 Goggles インスタンスは、特定の基本モデル、フレーム、LoRA 構成に対して 1 回トレーニングされ、トレーニングされていないドキュメントにフリーズして適用されます。これらの同じ文書に対してゴーグルを介してトレーニングされたモデルは、架空の注釈を持たず、機能を維持しながら、約 91% の確率でコンテンツに架空のフラグを立てます (GPQA と TruthfulQA はベースラインと一致またはそれを超えています)。同じアーキテクチャは他のフレームもサポートしています。Goggles インスタンスは、ドキュメントを単なるフィクションとしてではなく、「Redwood Research による AI 安全性評価の一部」として扱うようにトレーニングできます。与えられたフレームは、クレームに向かって押し戻される継続的な微調整の下で存続し、以前の介入は元に戻ります。 Goggles は、データが示す動作を吸収することなく、既知の不整合データに基づいて言語モデルをトレーニングする方法を提案します。

原文 (English)

Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing

Finetuning a language model on documents that are explicitly annotated as fictional results in a model that still actually believes the documents' core claims, an effect known as Negation Neglect. In our evaluations, models trained on documents prefixed and suffixed with such annotations correctly identify the relevant claims as fictional only about 9% of the time. To address this, we introduce Goggles, a learned module that intervenes on the finetuning gradient rather than the data. During supervised finetuning, a Goggles module edits the gradients an LLM LoRA receives, imparting a chosen epistemic frame (the stance the model takes toward the nature of what it reads) to whatever the documents teach. A Goggles instance is trained once for a given base model, frame, and LoRA configuration, then applied frozen to documents it was never trained on. Trained through Goggles on those same documents, now carrying no fictional annotation, the model flags the content as fictional roughly 91% of the time, while preserving capability (GPQA and TruthfulQA match or exceed baseline). The same architecture supports other frames: a Goggles instance can be trained to treat documents as "part of an AI safety evaluation by Redwood Research" rather than simply as fiction. The imparted frame persists under continued finetuning that pushes back toward the claim, where prior interventions revert. Goggles suggests a path toward training language models on known-misaligned data without absorbing the behaviors that data demonstrates.

13:00 JST画像/動画生成エージェント

COMFYCLAW: 画像生成ワークフローのための自己進化型スキルハーネス

エージェントは、ワークフローを構築し、人間が繰り返し行われるタスクをより効率的に完了できるように支援するためにますます使用されています。これらのワークフローが繰り返され、ドメイン固有になるにつれて、エージェントの記憶と再利用可能なスキルがますます重要になります。エージェントは、以前の実行からのワークフロー パターン、実行の制約、およびユーザー設定を思い出せる必要があります。私たちは、ワークフロー ベースの画像生成におけるこの問題を研究し、ComfyUI ワークフローを制御するためのエージェント スキル進化ハーネスである COMFYCLAW を紹介します。 COMFYCLAW は、ワークフロー構築を型付きグラフ編集として定式化し、構築段階ごとに整理されたツールを公開し、無効な編集を自動的に元に戻し、領域レベルのビジョン言語モデル (VLM) 検証機能を使用して、視覚的な障害を実行可能な修復提案に変換します。このフレームワークは、段階的に公開されるスキル ライブラリをさらに進化させ、以前の実行からの軌跡、実行エラー、検証者のフィードバックが再利用可能なエージェント スキルに蒸留されます。 4 つのベンチマーク分割、3 つのエージェント モデル、2 つのイメージ バックボーンにわたって、COMFYCLAW は 6 つのエージェント構成すべてにわたって最高の平均イメージ生成評価スコアを達成し、スキルの進化なしの検証者のみのベースラインを上回りました。さらに、人間によるアノテーションは、アノテーターがスキルの進化のないバリアントよりも COMFYCLAW を好むことを示しています。私たちの結果は、スキルの進化が、反復的なビジュアルワークフロー構築におけるエージェントの信頼性とパフォーマンスを向上させる効果的なメカニズムであることを示唆しています。

原文 (English)

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently. As these workflows become repeated and domain-specific, agent memory and reusable skills become increasingly important: agents should be able to recall workflow patterns, execution constraints, and user preferences from previous runs. We study this problem in workflow-based image generation and introduce COMFYCLAW, an agentic skill evolution harness for controlling ComfyUI workflows. COMFYCLAW formulates workflow construction as typed graph editing, exposes tools organized by construction stage, automatically reverts invalid edits, and uses a region-level vision-language model (VLM) verifier to translate visual failures into actionable repair suggestions. The framework further evolves a progressively disclosed skill library, where trajectories, execution errors, and verifier feedback from previous runs are distilled into reusable Agent Skills. Across four benchmark splits, three agent models, and two image backbones, COMFYCLAW achieves the best average image-generation evaluation score across all six agent configurations, outperforming a verifier-only baseline without skill evolution. Human annotations further show that annotators prefer COMFYCLAW over variants without skill evolution. Our results suggest that skill evolution is an effective mechanism for improving agent reliability and performance in recurring visual workflow construction.

13:00 JST研究/論文DeepSeek

SparseMixture-of-Experts 言語モデルを除去するための一般的な専門家カバレッジ

まばらにアクティブ化された専門家混合 (MoE) 言語モデルには、ルーティングされた専門家間の実質的な構造化された冗長性が含まれていますが、下流の調整データなしでそれらを刈り込むことは依然として困難です。既存のエキスパートによる枝刈り手法は通常、単一の集計された重要度スコアに依存しており、これにより、保持されるセットが、支配的なキャリブレーション パターンによって好まれるエキスパートに偏る可能性があります。私たちは \textbf{Generic TB-Coverage} を提案します。これは、キャリブレーションに汎用テキスト コーパス (WikiText2 および C4) のみを使用する、カバレッジを意識した専門家によるプルーニング手法です。エキスパートの有用性を 1 つのスコアにまとめる代わりに、私たちの方法では各コーパスでエキスパートごとのユーティリティを個別にプロファイリングし、最終的な枝刈りマスクを構築する前に各コーパスから有用性の高いエキスパートを保持する固定予算カバ​​レッジ ルールを適用します。 Qwen1.5-MoE-A2.7B と DeepSeek-MoE-16B-Base の 25\%、50\%、および 75\% の保持予算にわたって、私たちの手法はランダム プルーニング、REAP、ExpertSparsity よりも 6 つの一般的なゼロショット ベンチマークでの平均精度を向上させ、同時に WikiText2 と C4 でのパープレキシティの低下も軽減します。この利得は、積極的な枝刈り (25\% および 50\% の保持) の下で最大となり、コーパス間の専門家カバレッジを維持することが、MoE 枝刈りの事前の効果的な汎用データであることを示唆しています。私たちの改善は、固定の枝刈り予算と下流の校正データがない場合でも維持されます。

原文 (English)

Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models

Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without downstream calibration data remains challenging. Existing expert-pruning methods typically rely on a single aggregated importance score, which can bias the retained set toward experts favored by dominant calibration patterns. We propose \textbf{Generic TB-Coverage}, a coverage-aware expert pruning method that uses only generic text corpora (WikiText2 and C4) for calibration. Instead of collapsing expert utility into one score, our method profiles per-expert utility separately on each corpus and enforces a fixed-budget coverage rule that preserves high-utility experts from each corpus before constructing the final pruning mask. Across Qwen1.5-MoE-A2.7B and DeepSeek-MoE-16B-Base at 25\%, 50\%, and 75\% retention budgets, our method improves average accuracy on six common zero-shot benchmarks over random pruning, REAP, and ExpertSparsity, while also reducing perplexity degradation on WikiText2 and C4. The gains are largest under aggressive pruning (25\% and 50\% retain), suggesting that preserving cross-corpus expert coverage is an effective generic-data prior for MoE pruning. Our improvements hold with fixed pruning budgets and no downstream calibration data.

13:00 JST研究/論文GPT / ChatGPT

分布的に堅牢なリストごとの優先順位の最適化

言語モデルの調整に関する既存の堅牢な優先最適化は、主にペアごとの監視を研究し、データセット、プロンプト、または優先ペアのレベルで堅牢性を設定します。代わりに、ランキングラベルの不確実性の下でリストごとの選好の最適化を研究します。プロンプトと候補リストが与えられた場合、そのリストに対する観察されたランキングは、アノテーターの不一致、ニアタイ、損失のあるランクごとのフィードバック、または報酬モデルのノイズにより曖昧になる可能性があります。候補リストの条件付きランキング ラベルを直接堅牢化する、点ごとの合計変動ロバストな Plackett-Luce 目標を提案します。ロバスト損失では、公称 PL 損失と最悪の場合の PL 補正への正確な分解が認められ、最悪の場合のランキングは、現在の暗黙的スコアを昇順でソートすることによって取得され、内部最大化を $K!$ 列挙から $O(K\log K)$ に削減します。この扱いやすい構造により、オフラインおよびオンラインの最適化が強力に保証されます。オフライン固定リスト設定では、ロバストな目的は凸であり、投影された確率的部分勾配は $O(\epsilon^{-2})$ サンプル複雑さでグローバルな $\epsilon$-部分最適に達します。現在のポリシーによって候補リストが生成されるオンライン ポリシー誘導設定では、弱い凸性と $\widetilde O(\epsilon^{-2})$ モロー包絡線の定常性を確立します。オフライン LLM アライメントの実験では、提案されたロバスト補正により、クリーンなラベルの下でパフォーマンスが大幅に維持され、ノイズの下でのロバスト性が向上することが示されています。オンライン調整では、報酬モデルでランク付けされた候補拡張の信頼性が高まり、報酬モデルと外部 GPT-4 判定基準の両方が向上します。

原文 (English)

Distributionally Robust Listwise Preference Optimization

Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We instead study listwise preference optimization under ranking-label uncertainty: given a prompt and a candidate list, the observed ranking over that list may be ambiguous due to annotator inconsistency, near-ties, lossy rankwise feedback, or reward-model noise. We propose a pointwise total-variation robust Plackett--Luce objective that directly robustifies the ranking label conditional on the candidate list. The robust loss admits an exact decomposition into the nominal PL loss plus a worst-case PL correction, and the worst-case ranking is obtained by sorting current implicit scores in ascending order, reducing the inner maximization from $K!$ enumeration to $O(K\log K)$. This tractable structure yields strong offline and online optimization guarantees. In the offline fixed-list setting, the robust objective is convex and projected stochastic subgradient reaches global $\epsilon$-suboptimality with $O(\epsilon^{-2})$ sample complexity. In the online policy-induced setting, where candidate lists are generated by the current policy, we establish weak convexity and $\widetilde O(\epsilon^{-2})$ Moreau-envelope stationarity. Experiments in offline LLM alignment show that the proposed robust correction largely preserves performance under clean labels and improves robustness under noise. In online alignment, it makes reward-model-ranked candidate expansion more reliable and improves both reward-model and external GPT-4 judge metrics.

13:00 JST研究/論文

DRL-CLBA: DDPG 強化学習による音声分類に対するクリーン ラベル バックドア攻撃

音声分類用の深層学習モデルは、悪意のあるトリガーにより推論時に誤分類が引き起こされるバックドア攻撃に対して脆弱です。サンプル固有の攻撃は多くの防御を回避できますが、多くの場合、汚染されたラベル攻撃に依存しているため、手動のデータ防御によって検出可能になります。この論文では、DDPG (Deep Deterministic Policy Gradient) 強化学習を活用した音声分類のための新しいクリーン ラベル バックドア攻撃である DRL-CLBA を提案します。また、ディープオーディオステガノグラフィーを利用して、サンプル固有のトリガーをソースオーディオに埋め込み、特徴空間アンカーを作成します。提案された強化学習フレームワークは、モデルの深い潜在空間にあるトリガーを持つアンカー ポイントに向けてターゲット サンプルを効果的に最適化し、ラベル マイグレーションのないターゲット サンプルのポイズニングを可能にします。 3 つのデータセットと 4 つの異なる DNN にわたる実験結果は、DRL-CLBA が高い攻撃成功率を達成し、一部のバックドア防御を効果的に回避していることを示しています。この攻撃は、微調整、プルーニング、およびスペクトル署名防御に対して強力な耐性を示し、音声制御システムの重大な脆弱性を暴露します。

原文 (English)

DRL-CLBA: A Clean Label Backdoor Attack for Speech Classification via DDPG Reinforcement Learning

Deep learning models for speech classification are vulnerable to backdoor attacks, where malicious triggers cause misclassification at inference time. While sample-specific attacks can bypass many defenses, they often rely on poisoned label attack, making them detectable via manual data defense. In this paper, we propose DRL-CLBA, a novel clean label backdoor attack for speech classification that leverages Deep Deterministic Policy Gradient (DDPG) reinforcement learning. We also utilize deep audio steganography to embed sample-specific triggers into source audio, creating feature-space anchors. The proposed reinforcement learning framework effectively optimizes target samples toward trigger-bearing anchor points in the model's deep latent space, enabling label-migration-free poisoning of target samples. Experimental results across three datasets and four different DNNs demonstrate that DRL-CLBA achieves a high attack success rate, effectively bypassing some backdoor defenses. The attack demonstrates strong resistance against fine-tuning, pruning, and spectral signature defenses, exposing critical vulnerabilities in speech-controlled systems.

13:00 JST研究/論文

ジョルダン曲線定理の再定式化

入力証明が自然言語ではなく、別の証明アシスタントでの正式な開発である自動形式化の変形である再形式化のケーススタディを紹介します。具体的には、ジョルダン曲線定理の 3 つの再定式化、つまり Mizar から Lean へ、HOL Light から Lean へ、HOL Light から Agda への修正を報告します。私たちは結果を分析し、実際の改革タスクにとって重要なパイプライン設計の選択肢を特定します。

原文 (English)

Reformalization of the Jordan Curve Theorem

We present a case study in reformalization, a variant of autoformalization in which the input proof is not natural language but a formal development in a different proof assistant. Concretely, we report three reformalizations of the Jordan Curve Theorem: from Mizar to Lean, from HOL Light to Lean, and from HOL Light to Agda. We analyse the results and identify pipeline design choices that matter for practical reformalization tasks.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

金融サービス LLM 評価のメタベンチマーク

公開 LLM リーダーボードは、世界平均のパフォーマンスに合わせて最適化されており、金融サービス業務に特有の認知的要求を捉えていません。MMLU-Pro をリードするモデルは、文書に基づいたコンプライアンスの推論ではパフォーマンスが劣る可能性があり、コーディング リーダーは複数ターンにわたる顧客インタラクションの処理が不十分になる可能性があります。当社は、公的に報告された 452 のベンチマークを 41 の O*NET 一般化作業アクティビティに整理し、それらを販売、業務、リスク、サポート業務にわたる 38 の BIAN 銀行ビジネス ドメインに集約するメタ ベンチマーク フレームワークを提示します。ローリング モデル ウィンドウで計算される乗法重み付けスキーム (識別 x カバレッジ x リーンシー) は、依然として最良のモデルを区別するベンチマークに報酬を与え、広く報告され、アクティブに使用され続けており、飽和したレガシー テストを自動的に抑制します。これらの重みは、ペアごとの Elo トーナメントの K ファクターをスケールし、生のスコアを正規化することなく、クロスベンチマークで比較可能な作業活動スコアを生成します。ビジネスドメインスコアは、構成要素である作業活動 Elos の加重平均です。 2026 年 6 月の時点で 25 の組織にわたる 288 のモデルをカバーするポイントインタイムの公開スナップショットでフレームワークを実証し、同様の選択とガバナンスの課題に直面している組織でこのアプローチを再現可能にすることを目的として、方法論、完全な分類、設計上の決定、および制限について説明します。

原文 (English)

Meta-Benchmarks for Financial-Services LLM Evaluation

Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model that leads on MMLU-Pro may underperform on document-grounded compliance reasoning, and a coding leader may handle multi-turn customer interactions poorly. We present a meta-benchmarking framework that organises 452 publicly reported benchmarks into 41 O*NET Generalized Work Activities and aggregates those into 38 BIAN banking business domains spanning sales, operations, risk, and support work. A multiplicative weighting scheme (discrimination x coverage x recency), computed over a rolling model window, rewards benchmarks that still separate the best models, are widely reported, and remain in active use, suppressing saturated legacy tests automatically. These weights scale the K-factor in a pairwise Elo tournament, producing cross-benchmark-comparable work-activity scores without raw score normalisation; business-domain scores are weighted averages of the constituent work-activity Elos. We demonstrate the framework on a point-in-time public snapshot covering 288 models across 25 organisations as of June 2026, and describe the methodology, full taxonomy, design decisions, and limitations with the aim of making the approach reproducible for institutions facing similar selection and governance challenges.

13:00 JSTエージェント

視覚言語ナビゲーションにおける意味論的探索のためのパスレベルの後知恵命令

ポリシー上の探索は、ポリシーをより広範な状態分布にさらすため、堅牢な視覚言語ナビゲーション エージェントをトレーニングするための重要なコンポーネントです。ただし、そのような探索は必然的に専門家のデモンストレーションから逸脱する軌道につながり、実行されたビジュアル ストリームと元の言語命令の間に意味的な不一致が生じます。この取り組みでは、後知恵の推論を活用して指示をエージェントの実際の探索行程に合わせる統一ポリシー フレームワークである Phi-Nav を導入することで、この課題に対処します。具体的には、Phi-Nav は 3 段階の二重監視サイクルを通じて動作します。1) エージェントは、専門家のアクション フィードバックから学習しながら、オラクルに基づくポリシー探索を実行し、軌跡をサンプリングします。2) 後知恵の話者は、収集した視覚的観察に基づいてパスレベルの後知恵の指示を合成します。3) エージェントは 2 番目の模倣パスを実行し、合成された軌跡と指示のペアを追加の専門家のデモンストレーション。このプロセスを通じて、Phi-Nav は、オンポリシー手法に固有の重大な意味論的な監視のギャップを埋め、意味論的にラベルのない動きを高密度のトレーニング信号に変換します。 R2R-CE および RxR-CE ベンチマークの評価では、Phi-Nav が、現在のベースラインで使用されている専門家のデモンストレーションのほんの一部のみを必要としながら、競争力のあるパフォーマンスを実現していることが示されています。これらの結果は、VLN におけるセマンティック探索の必要性を強調しており、Phi-Nav を限られたデータで身体化されたエージェントをトレーニングするための効果的なソリューションとして位置づけています。

原文 (English)

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate from expert demonstrations, resulting in a semantic mismatch between the executed visual stream and the original language instruction. In this work, we address this challenge by introducing Phi-Nav, a unified on-policy framework that leverages hindsight reasoning to align instructions with the agent's actual exploratory journey. Specifically, Phi-Nav operates through a three-stage dual-supervision cycle: 1) the agent performs oracle-guided on-policy exploration, sampling a trajectory while learning from expert action feedback, 2) a hindsight speaker synthesizes a path-level hindsight instruction grounded in the collected visual observations, and 3) the agent conducts a second imitation pass, treating the synthesized trajectory-instruction pair as an additional expert demonstration. Through this process, Phi-Nav bridges the critical semantic supervision gap inherent in on-policy methods, transforming semantically unlabeled movement into dense training signals. Evaluations on the R2R-CE and RxR-CE benchmarks show that Phi-Nav yields competitive performance while requiring only a fraction of the expert demonstrations used by current baselines. These results underscore the necessity of semantic exploration in VLN, positioning Phi-Nav as an effective solution for training embodied agents with limited data.

13:00 JSTエージェントGPT / ChatGPT

マスターマインド: リポジトリ規模の脆弱性再現のための戦略に基づく学習

リポジトリ レベルの脆弱性の再現は、要求の厳しいソフトウェア エンジニアリング (SE) タスクです。エージェントは、コードベースを検査し、脆弱なパスに到達する入力文法を推測し、概念実証 (PoC) を構築し、パッチ適用されたビルドでクラッシュが解消されることを確認する必要があります。最近の LLM エージェントは、アプローチが正しい場合にはこれらのステップを実行できる場合が多いですが、それでも間違った戦略を選択することで失敗します。この論文では、完全なアクションの軌跡ではなく、戦略がそのような SE エージェントにとって適切な学習単位であると主張します。戦略は、最適化するには十分コンパクトで、実行をガイドするには十分具体的で、試行をまたいで保存して再利用するには十分安定しています。私たちは、転移可能な戦略学習をタスク固有の経験から分離するデュアルループ フレームワークである Mastermind を紹介します。訓練可能なプランナーは、SFT およびマイルストーンベースの GRPO を通じて再利用可能な脆弱性再現戦略を学習します。一方、エクスペリエンス ループは、後続の試行をガイドするタスク ローカル戦略記録を維持します。プランナーは実行者とは独立してトレーニングされるため、アクション生成機能を変更することなく、複数の凍結された実行者を改善する戦略学習が可能になります。 Cyber​​Gym で 260 のトレーニング タスクと 200 の保留された評価タスクを使用して Mastermind を評価します。 GPT-5.5 を凍結実行者として使用すると、Mastermind は 84.5% の合格率を達成し、オープンブック PoC コンテキスト (60.0%)、Best-of-8 サンプリング (63.0%)、反復改善 (77.0%) を上回りました。同じプランナーは、GPT-5.4 mini と GLM~5.1 も 45.0% と 58.5% から 60.0% と 71.0% に改善しました。これらの結果は、高レベルの戦略を学習することが、リポジトリ規模の SE エージェントを改善するための効果的かつ応用可能なメカニズムであることを示しています。

原文 (English)

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar that reaches a vulnerable path, construct a proof-of-conceptv(PoC), and verify that the crash disappears on the patched build. Recent LLM agents can often execute these steps when the approach is correct, yet they still fail by choosing the wrong strategy. This paper argues that strategy, rather than the full action trajectory, is the right learning unit for such SE agents: it is compact enough to optimize, concrete enough to guide execution, and stable enough to store and reuse across attempts. We present Mastermind, a dual-loop framework that separates transferable strategy learning from task-specific experience. A trainable planner learns reusable vulnerability-reproduction strategies through SFT and milestone-based GRPO, while an experience loop maintains task-local strategy records that guide subsequent attempts. The planner is trained independently of the executor, allowing strategy learning to improve multiple frozen executors without modifying their action-generation capability. We evaluate Mastermind on CyberGym using 260 training tasks and 200 held-out evaluation tasks. With GPT-5.5 as the frozen executor, Mastermind achieves an 84.5% pass rate, outperforming open-book PoC context (60.0%), Best-of-8 sampling (63.0%), and iterative improvement (77.0%). The same planner also improves GPT-5.4 mini and GLM~5.1 from 45.0% and 58.5% to 60.0% and 71.0%. These results demonstrate that learning high-level strategies is an effective and transferable mechanism for improving repository-scale SE agents.

13:00 JSTLLM/生成AIエージェント

SimWorlds: 動的な 3D シーン作成のためのマルチエージェント システム

LLM エージェントは、自然言語を手続き型の 3D シーンに変換するために使用されることが増えていますが、既存のシステムは静的な出力に重点を置いています。液体が流れ、粒子が放出され、剛体がカスケードし、関節機構が動く、テキストだけのダイナミック 4D シーンは、編集可能なコンテンツとして、またビデオ生成や具現化された AI のための物理学に基づいたトレーニング データとしての価値があるにもかかわらず、ほとんど解明されていないままです。この動的なケースは、静的なテキストからシーンへの作業とは異なる 2 つの課題により区別されます。エージェントは、単一の一貫したシーン内で空間レイアウト、複数の物理ソルバー、時間シーケンス、カメラ、照明を共同で調整する必要があり、レンダリングされたビデオからモーションの正しさを検証することは、単一の画像を判断するよりも基本的に困難です。 SimWorlds を紹介します。これは、Blender 固有の手順知識を備え、テキストから動的で編集可能な 4D シーンを生成するマルチエージェント フレームワーク、構築段階の固定順序シーケンスを駆動するプランナー、コーダー、レビューアーのワークフロー、決定論的検証者​​によって強制されるレイヤード シーン プロトコル、およびレンダリングされたイメージでは明らかにできないメカニズムの障害を捕捉する実行時状態検査ツール スイートです。また、テキスト プロンプトから生成されたプロシージャルなダイナミック 3D シーンの視覚的な忠実性と物理的な一貫性の両方を評価するためのベンチマークである 4DBuildBench も紹介します。実験では、SimWorlds が以前の動的 Blender 生成のベースラインを上回るパフォーマンスを示しています。

原文 (English)

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, rigid bodies cascade, and articulated mechanisms move, remain largely unexplored despite their value as editable content and as physics-grounded training data for video generation and embodied AI. Two challenges set the dynamic case apart from static text-to-scene work: an agent must jointly coordinate spatial layout, multiple physics solvers, temporal sequencing, camera, and lighting in a single coherent scene, and verifying motion correctness from rendered video is fundamentally harder than judging a single image. We present SimWorlds: a multi-agent framework that produces dynamic, editable 4D scenes from text, with Blender-specific procedural knowledge, a planner-coder-reviewer workflow driving a fixed ordered sequence of construction stages, a layered scene protocol enforced by a deterministic verifier, and a runtime-state inspection tool suite that catches mechanism failures the rendered image cannot reveal. We also introduce 4DBuildBench, a benchmark for assessing both visual fidelity and physical consistency of the procedural dynamic 3D scenes generated from text prompts. Experiments show that SimWorlds outperforms prior dynamic Blender generation baselines.

13:00 JSTエージェント

症状ではなくアンプを修復する: エージェント ロールアウトの安定したワールド モデルの修正

エージェントの計画が短いツール チェーンから数千または数万のステップを含む永続的なワークフローに移行すると、個別の予測ではなく、大規模な計画グラフ内で障害が発生します。間違いが発生するたびにグラフ全体を再計画することは、計算上現実的でも望ましいことでもありません。グラフ全体の再生は大量のコンテキスト バジェットを消費し、LLM を多くの無関係な症状にさらし、長いコンテキストの取得を低下させる可能性があります。この論文では、そのようなシステムに欠けているコンポーネント、つまり失敗した計画グラフを適切な場所に修復するワールド モデル コレクターについて研究します。 2 つの補正器ファミリーを比較します。 1 つ目は一般的なエンジニアリング アプローチです。ノードとエッジをスキャンし、疑わしいローカル領域を選択し、LLM に修復を依頼します。私たちは強力なエンジニアリング LLM コレクタを実装しており、特に非常に大規模なコンテキストが与えられた場合に役立つことがわかりました。 2 番目のファミリーは、私たちのアプローチである WM-SAR (World-Model Subgraph Amplification Repair) です。目に見える症状をスキャンするのではなく、サブグラフの増幅から逆算して、エラーを再増幅し続けるノードとエッジを特定し、その原因となるサブグラフのみを LLM に送信します。グラフ シミュレーションと LLM 修復実験全体で、WM-SAR は現実的なトークン バジェットの下でエンジニアリング コレクタを大幅に上回り、コンパクトな領域でほぼグラフ全体の安定化を達成し、LLM にクリーンな修復ターゲットを与えます。

原文 (English)

Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts

As agent planning moves from short tool chains toward persistent workflows with thousands or tens of thousands of steps, failures will occur inside large planning graphs rather than in isolated predictions. Replanning the entire graph after every mistake is neither computationally realistic nor desirable: full-graph replay consumes large context budgets, exposes the LLM to many irrelevant symptoms, and can degrade long-context retrieval. This paper studies the missing component in such systems: a world-model corrector that repairs the failed planning graph in place. We compare two families of correctors. The first is the common engineering approach: scan nodes and edges, choose a suspicious local region, and ask an LLM to repair it. We implement strong engineering LLM correctors and find that they can help, especially when given very large contexts. The second family is our approach, WM-SAR (World-Model Subgraph Amplification Repair): instead of scanning for visible symptoms, it works backward from subgraph amplification, identifies the nodes and edges that keep re-amplifying error, and sends only that causal subgraph to the LLM. Across graph simulations and LLM repair experiments, WM-SAR substantially outperforms engineering correctors under realistic token budgets, achieves near-whole-graph stabilization with a compact region, and gives the LLM a cleaner repair target.

13:00 JST研究/論文

検索に基づいた形式的概念分析による検証可能な知識の拡張

オントロジーの構築では、どのオブジェクト、属性、構造的関係を有効な知識として受け入れるかを決定する必要があります。言語モデルはテキストからそのような構造を提案できますが、その出力は依然としてサポートされていないか、一貫性がない可能性があります。この論文では、形式概念分析 (FCA) を知識拡張のための記号検証ループとして使用する、検索拡張小型言語モデル (SLM) フレームワークを提案します。 FCA は、シード属性から始めて、成長する正式なコンテキストに対する影響を提案します。次に、検索に基づいた SLM オラクルが各含意を検証するか、反例を返します。オラクルは、発生率の判断、一貫性チェック、および属性の提案もサポートしており、受け入れられた含意、反例、矛盾、および修正を検査可能にします。 Orphadata リソースから構築されたまれな運動失調設定では、検索に基づいた 10 シードの実行により、0.29 ~ 0.52 の関係 F1 と 0.22 ~ 0.30 の閉包ベースの含意 F1 が得られます。シード セットが大きくなると、評価される含意の数が増加し、含意 F1 が向上することがよくあります。低い含意スコアは、導出された含意のより厳密な評価を反映しており、1 つの見逃した関係または余分な関係が複数の含意の判断に影響を与える可能性があります。アブレーションは、固定オブジェクト属性設定での発生判定がクロージャーベースの含意スコアを改善できることを示しています。ただし、候補となるオブジェクトと属性が固定されている場合でも、肯定的なオブジェクトと属性のペアを特定することは依然として困難です。

原文 (English)

Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis

Ontology construction requires deciding which objects, attributes, and structural relations should be accepted as valid knowledge. Language models can propose such structures from text, but their outputs can still be unsupported or inconsistent. This paper proposes a retrieval-augmented small language model (SLM) framework that uses formal concept analysis (FCA) as a symbolic verification loop for knowledge expansion. Starting from seed attributes, FCA proposes implications over a growing formal context. A retrieval-grounded SLM oracle then validates each implication or returns a counterexample. The oracle also supports incidence judgments, consistency checks, and attribute proposals, making accepted implications, counterexamples, contradictions, and corrections inspectable. In a rare ataxia setting constructed from Orphadata resources, retrieval-grounded 10-seed runs obtain relation F1 of 0.29-0.52 and closure-based implication F1 of 0.22-0.30. Larger seed sets increase the number of evaluated implications and often improve implication F1. The lower implication scores reflect a stricter evaluation of derived implications, where one missed or extra relation can affect several implication judgments. Ablations show that incidence judgments in a fixed object-attribute setting can improve closure-based implication scores. However, identifying positive object-attribute pairs remains difficult even when the candidate objects and attributes are fixed.

13:00 JSTLLM/生成AI

サブリミナル時計: 拡散言語モデルにおける潜時モデリング

拡散言語モデル (DLM) は、自己回帰モデルの有望な代替手段として最近登場しました。標準的な拡散ベースのアプローチとは異なり、DLM はタイムステップで明示的に条件付けされていないため、当然の疑問が生じます。これらのモデルは内部的にノイズ除去の進行状況を表しているのか、そのような情報は下流でどのように使用されるのでしょうか。この研究では、DLM が実際にその残差ストリーム内の拡散タイムステップに関連する潜在表現をエンコードしていることを示します。この信号は、層全体のプローブを使用して確実に抽出できることがわかり、ノイズ除去の進行が内部活性化から解読可能であることがわかります。さらに、推論されたタイムステップに関連付けられた低次元部分空間に沿ってモデルを操作することで、ノイズ除去の進行の概念を体系的に調整できるようになり、モデルの信頼性とエントロピーに予測可能な変化がもたらされることを示します。最後に、識別された表現の幾何学的形状を分析し、それが活性化空間で構造化された解釈可能な特性を示すことを示し、そのような信号がこれらのモデルによってどのように処理されるかを明らかにします。

原文 (English)

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can be reliably extracted using probes across layers, indicating that denoising progress is decodable from internal activations. We further demonstrate that steering the model along a low-dimensional subspace associated with the inferred timestep allows us to systematically modulate its notion of denoising progress, leading to predictable changes in model confidence and entropy. Finally, we analyse the geometry of the identified representation, showing that it exhibits structured and interpretable properties in activation space, and shedding light on how such a signal is processed by these models.

13:00 JSTLLM/生成AIエージェントClaude

LLM エージェントの大規模な安全性テスト: リスク発見から証拠に基づく検証まで

LLM エージェントは外部ツールを通じて自律的なアクションを実行することが増えており、複雑かつ進化する安全リスクにつながっています。しかし、既存の安全性テストは専門家が設計した安全性違反を対象としており、対応する結果はハードコードされたルールによって評価されるため、エージェントの進化に応じてテストを拡張するにはコストがかかります。この目的を達成するために、3 段階の自己強化パイプラインを通じて非決定性エージェントのソフトウェア エンジニアリング テスト原則をインスタンス化する、エンドツーエンドの自動安全性テスト フレームワークである Vera を紹介します。まず、文献に基づいた調査により、新たなリスクを継続的に発見し、安全リスク、攻撃手法、およびツール実行環境の分類に構造化します。第 2 に、分類次元全体にわたる組み合わせ構成により、実行可能な安全ケースが生成され、それぞれが具体的な安全目標、プログラムで構築された初期状態、および観察可能なアーティファクトに基づいた決定論的検証述語を指定します。第三に、適応的実行は、分離されたサンドボックスで異種エージェントを実行します。制御エージェントは実行時の観察に基づいてマルチターンの対話を制御しますが、証拠に基づいた検証者は、モデルの自己報告ではなく環境の状態とツール呼び出しの証拠から結果を判断します。 4 つの運用エージェント フレームワーク (OpenClaw、Hermes、Codex、Claude Code) で Vera を評価したところ、重大な安全上の弱点が明らかになり、マルチチャネル攻撃下での平均攻撃成功率は 93.9% に達しました。また、3 つの実行設定にわたる 124 のリスク カテゴリにわたる 1,600 の実行可能な安全性ケースで構成される Vera-Bench もリリースします。これらの結果は、急速に進化する大規模なエージェント システムの厳密で保守可能な安全性評価には、モジュール式の実行可能なテスト インフラストラクチャが不可欠であることを示しています。コードは https://github.com/Yunhao-Feng/Vera で公開されています。

原文 (English)

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9\% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale. The code is publicly available at https://github.com/Yunhao-Feng/Vera.

13:00 JST研究/論文GPT / ChatGPTGemini

MMIR-TCM: TCM 臨床意思決定サポートのためのメモリ統合マルチモーダル推論および検索

伝統的な中国医学 (TCM) の診断、特に舌検査による診断は、主観性と再現性という点で根強い課題に直面しています。症候群の鑑別や処方の作成など、TCM の臨床タスクへのマルチモーダル人工知能の適用は、視覚的な舌の特徴とテキスト推論の間の意味論的なギャップ、および大規模な標準化されたデータセットの欠如によって大幅に妨げられています。これらの課題に対処するために、マルチモーダル大規模言語モデル (MLLM) とメモリ拡張セグメンテーションおよび検索拡張生成 (RAG) を統合することにより、TCM 専門家の診断プロセスをエミュレートする新しいフレームワークである MMIR-TCM を紹介します。 MMIR-TCM は 3 段階のアーキテクチャを採用しており、堅牢な舌抽出のためのトレーニング不要の Memory-SAM モジュール、構造化された舌診断生成のための微調整された Qwen3-VL モデル、および証拠に基づく臨床意思決定支援生成のための Qwen3 ベースの RAG コンポーネントを統合しています。このフレームワークは、高度な TCM 研究のために特別に導入された新しい大規模マルチモーダル データセットである MedTCM を使用して開発および検証されました。既存の指標では把握できないフレームワークの臨床精度を適切に評価するために、意味の理解と診断の重要性を組み込んだ領域固有の評価指標である TDEU も開発しました。当社の包括的な実験では、MMIR-TCM が GPT-4o や Gemini 2.5 Flash などの主要モデルよりも大幅に優れていることが実証されています。

原文 (English)

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and reproducibility. The application of multimodal artificial intelligence to TCM clinical tasks, such as syndrome differentiation and prescription generation, is significantly hampered by the semantic gap between visual tongue features and textual reasoning, as well as the lack of large-scale, standardized datasets. To address these challenges, we introduce MMIR-TCM, a novel framework that emulates the diagnostic process of TCM experts by integrating multimodal large language model(MLLM) with memory-augmented segmentation and retrieval-augmented generation (RAG). Employing a three-stage architecture, MMIR-TCM integrates a training-free Memory-SAM module for robust tongue extraction, a fine-tuned Qwen3-VL model for structured tongue diagnosis generation, and a Qwen3-based RAG component for evidence-grounded clinical decision support generation. The framework was developed and validated using MedTCM, a new large-scale multimodal dataset that we introduce specifically for advanced TCM research. To properly evaluate our framework's clinical accuracy, which existing metrics fail to capture, we also developed TDEU, a domain-specific evaluation metric incorporating semantic understanding and diagnostic importance. Our comprehensive experiments demonstrate that MMIR-TCM significantly outperforms leading models, including GPT-4o and Gemini 2.5 Flash.

13:00 JSTLLM/生成AI研究/論文

飛行前: 航空運用知識に関する大規模な言語モデルを評価するためのベンチマーク

大規模言語モデル (LLM) は、文書作成やトレーニングの生成から顧客対応アシスタントに至るまで、航空ビジネスの運営に提案されることが増えています。汎用ベンチマークでは、モデルが航空特有の運用知識について安全かつ正確に推論しているかどうかは測定されません。また、この分野の大きなリスクと規制の性質により、そのギャップは重大なものになります。 Pre-Flight は、国際標準と空港地上運用資料から抽出された 300 の多肢選択式質問からなるオープンソース ベンチマークで、国際空港の地上運用、ICAO および米国 FAA の規制、航空一般知識、複雑な運用シナリオをカバーしています。質問は、航空交通管理、地上業務、商業飛行の経験を持つ専門家によって作成および検討されました。当社は、Inspect 評価フレームワークを使用して、さまざまな現代の商用およびオープンウェイト モデルを評価し、標準的な多肢選択プロトコルに基づく精度によるスコアリングを行い、新しいモデルがリリースされるたびにリーダーボードをローリング ベースで維持しています。カンファレンスでの航空専門家の低サンプルクイズから得られた非公式の専門家基準である約95%に対して、評価された最も強力なモデル(2026年リリース)でさえ82.7%に達し、2025年初頭の約75%から徐々にしか改善していません。したがって、専門家レベルの信頼性を下回る実質的かつ持続的なギャップが残っています。データセット、評価ハーネス、および結果をリリースしており、ベンチマークは、inspect_evals で配布されるコミュニティ評価パッケージ内で利用できます。私たちは、この種のドメイン固有の評価は、安全性が重要ではない航空業務における生成 AI の責任ある導入に必要な前提条件であると主張します。

原文 (English)

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing assistants. General purpose benchmarks do not measure whether a model reasons safely and correctly about aviation specific operational knowledge, and the high stakes, regulated nature of the domain makes that gap consequential. We present Pre-Flight, an open source benchmark of 300 multiple choice questions drawn from international standards and airport ground operations material, covering international airport ground operations, ICAO and US FAA regulations, aviation general knowledge and complex operational scenarios. Questions were authored and reviewed by practitioners with experience in air traffic management, ground operations and commercial flying. We evaluate a range of contemporary commercial and open weight models using the Inspect evaluation framework, scoring by accuracy under a standard multiple choice protocol, and we maintain the leaderboard on a rolling basis as new models are released. Against an informal expert reference of around 95%, obtained from a low sample quiz of aviation professionals at a conference, even the strongest model evaluated (released in 2026) reaches 82.7%, having improved only gradually from roughly 75% in early 2025. A substantial and persistent gap below expert level reliability therefore remains. We release the dataset, the evaluation harness and the results, and the benchmark is available within the community evaluations package distributed with inspect_evals. We argue that domain specific evaluation of this kind is a necessary precondition for responsible deployment of generative AI in non safety critical aviation operations.

13:00 JST研究/論文

フォルトツリーにおける実際の因果関係

フォールト ツリーは、複雑なシステムの効果的なリスク モデルとして広く使用されており、特に最小限のカット セット分析を通じて、「何が問題になる可能性があるのか​​?」という質問に答えます。私たちは、ハルパーンとパールの実際の因果関係の理論の観点から断層木を研究します。これにより、フォールト ツリーを使用して、障害診断の基本である「なぜ問題が発生したのか?」という質問に答えることができます。フォールトツリーのグラフ構造と論理構造の観点から実際の因果関係のさまざまな概念を完全に分類し、最小のカットセットがどのように実際の原因を引き起こすかを示します。

原文 (English)

Actual causality in fault trees

Fault trees are a widely used as effective risk models for complex systems, answering the question "what can go wrong?", especially through minimal cut set analysis. We study fault trees from the perspective of Halpern & Pearl's theory of actual causality. This allows us to use fault trees to answer the question "why has it gone wrong?", which is fundamental to failure diagnostics. We give a complete classification of each of the different notions of actual causality in terms of the fault tree's graph structure and logical structure, and show how minimal cut sets give rise to actual causes.

13:00 JSTエージェントビジネス/資金調達

CLAP: ドメイン エージェントのトレーニング後のクローズド ループ トレーニング、評価、リリース制御

ドメイン エージェントは、多くの場合、ノイズの多いビジネス データ、不確実なトレーニング後の利益、オフライン/アプリケーションの不一致、アダプター リリースのリスクに直面します。このペーパーでは、ビジネス データを構造化された SFT サンプル、意思決定優先サンプル、ホールドアウト セット、リスク診断、およびリリース ゲート レコードに変換する閉ループ手法である CLAP (Closed-Loop Agent Post-training) について説明します。 CLAP は、データ検証、ターゲット/証拠の正規化、報酬/KL 診断、オフライン ゲート、およびアプリケーション チェーンの再生を組み合わせて、アダプターがターゲット アプリケーション チェーンに適しているかどうかを判断します。 5 つの匿名化された製造シナリオ バッチで、QLoRA スタイルの LoRA-SFT は平均してわずかな向上をもたらしました。全体のスコアは 0.0098 増加し、合格率は 0.0240 増加し、証拠の精度は 0.0280 増加しましたが、幻覚と誤った事実は減少しました。しかし、改善するのは 5 つのバッチのうち 3 つだけで、一部のバッチは後退し、GRPO は高い KL リスクを露呈します。アプリケーション チェーンの再生は、事実の抽出に RAG が必要であることをさらに示しています。同じ 3B バックボーンと 100 のリプレイ ケースの下では、アプリケーション指向の LoRA-SFT アダプターは、ベース + RAG よりも値、コア フィールド、および回答証拠のドキュメント/ページのマッチングを向上させますが、レイテンシは増加します。これらの結果は、トレーニングの完了や単​​一のオフライン スコアに依存するのではなく、統合されたデータ、トレーニング、評価、リリースのループを通じてトレーニング後のドメイン エージェントを管理することをサポートします。

原文 (English)

CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training

Domain agents often face noisy business data, uncertain post-training gains, offline/application mismatch, and adapter-release risk. This paper presents CLAP (Closed-Loop Agent Post-training), a closed-loop method that converts business data into structured SFT samples, decision-preference samples, holdout sets, risk diagnostics, and release-gate records. CLAP combines data validation, target/evidence normalization, reward/KL diagnosis, offline gates, and application-chain replay to decide whether an adapter is suitable for the target application chain. On five anonymized manufacturing-scenario batches, QLoRA-style LoRA-SFT yields modest average gains: overall score increases by 0.0098, pass rate by 0.0240, and evidence accuracy by 0.0280, while hallucination and wrong facts decrease. Yet only 3 of 5 batches improve, some batches regress, and GRPO exposes high KL risks. Application-chain replay further shows that RAG is necessary for factual extraction; under the same 3B backbone and 100 replay cases, an application-RAG-oriented LoRA-SFT adapter improves value, core fields, and answer-evidence doc/page matching over base+RAG, but increases latency. These results support managing domain-agent post-training through an integrated data-training-evaluation-release loop rather than relying on training completion or a single offline score.

13:00 JSTLLM/生成AIGPT / ChatGPT

改良による安全性を狙った埋め込みエクスプロイト

大規模言語モデル (LLM) の安全トレーニングは主に英語で行われるため、安全メカニズムがリソースの少ない言語や混合言語のコード スイッチングにどの程度一般化されるかは不透明です。これにより、安全トレーニングの分布から外れた入力に対してモデルが自信を持って有害な反応を生成するという認識論的なギャップが生じることを示します。この現象を研究するために、STEER (Safety Targeted Embedding Exploit via Refinement) を導入します。これは、モデルの拒否動作に最も強く寄与する単語を特定し、それらを低リソース言語に反復的に翻訳して、有害な意図を維持しながら拒否を抑制する勾配ガイド型攻撃です。 6 つのオープンソース 8B パラメータ モデル全体で、STEER は JailbreakBench で最大 93.0%、AdvBench で 96.7% の攻撃成功率を達成し、ランダム コード スイッチングや Greedy Cooperative Gradient (GCG) を上回ります。結果として得られるプロンプトは GPT-4o-mini にも転送され、ターゲット モデルへのアクセスを必要とせずに 35.5% の攻撃成功率を達成しました。これは、根本的な弱点が単一のアーキテクチャに固有のものではないことを示唆しています。これらの調査結果は、主に英語に合わせた安全メカニズムが多言語入力全体に一般化するとは想定できないことを示しています。私たちは、多言語の安全性を向上させるには、調整中のより広い範囲と、配布範囲外の入力を明示的に検出して回避するメカニズムが必要であると主張します。

原文 (English)

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training. To study this phenomenon, we introduce STEER (Safety Targeted Embedding Exploit via Refinement), a gradient-guided attack that identifies words contributing most strongly to the model's refusal behavior and iteratively translates them into low-resource languages to suppress refusal while preserving harmful intent. Across six open-source 8B-parameter models, STEER achieves attack success rates of up to 93.0% on JailbreakBench and 96.7% on AdvBench, outperforming random code-switching and Greedy Coordinate Gradient (GCG). The resulting prompts also transfer to GPT-4o-mini, achieving a 35.5% attack success rate without requiring access to the target model, suggesting that the underlying weakness is not specific to a single architecture. These findings demonstrate that safety mechanisms aligned primarily on English cannot be assumed to generalize across multilingual inputs. We argue that improving multilingual safety requires broader coverage during alignment and mechanisms that explicitly detect and abstain on out-of-distribution inputs.

13:00 JST研究/論文

CamoNAS: 強化された偽装オブジェクト検出のためのニューラル アーキテクチャ検索

カモフラージュオブジェクト検出 (COD) は、周囲の環境に溶け込んでいるオブジェクトを特定してセグメント化することを目的としていますが、エッジの手がかりが弱く、境界が不明確であるため課題が生じています。従来の COD モデルは、手作業で設計されたアーキテクチャとマルチスケール機能の融合に依存しており、体系的な検索ではなく直感によって導かれることがよくあります。このペーパーでは、COD 用の周波数を意識した多重解像度ニューラル アーキテクチャ検索 (NAS) フレームワークである CamoNAS を紹介します。 CamoNAS は、セル レベルの操作とネットワーク レベルのダウンサンプリング パスの両方を自動的に検索し、偽装されたオブジェクトを検出するために調整された階層的な検索スペースを形成します。さらに、RGB 周波数デュアル ストリーム アーキテクチャを採用しており、学習可能なウェーブレット変換が RGB 空間ストリームを補完します。 CamoNAS は 4 つの COD ベンチマーク (CAMO、COD10K、NC4K、CHAMELEON) で最先端のパフォーマンスを達成し、COD に対する NAS の有効性を強調しています。私たちのコードは https://github.com/rendaweiSIMIT/CamoNAS で入手できます。

原文 (English)

CamoNAS: Neural Architecture Search for Enhanced Camouflaged Object Detection

Camouflaged Object Detection (COD) aims to locate and segment objects that blend into their surroundings, presenting challenges due to weak edge cues and ill-defined boundaries. Traditional COD models rely on hand-designed architectures and multi-scale feature fusion, which are often guided by intuition rather than systematic search. This paper introduces CamoNAS, a frequency-aware multi-resolution Neural Architecture Search (NAS) framework for COD. CamoNAS automatically searches both cell-level operations and network-level downsampling paths, forming a hierarchical search space tailored to detect camouflaged objects. Additionally, it adopts an RGB frequency dual-stream architecture, where a learnable wavelet transform complements the RGB spatial stream. CamoNAS achieves state-of-the-art performance on four COD benchmarks (CAMO, COD10K, NC4K, CHAMELEON), highlighting the effectiveness of NAS for COD. Our code is available at https://github.com/rendaweiSIMIT/CamoNAS.

13:00 JSTLLM/生成AIエージェント

SkillCoach: エージェントのスキル使用を評価および強化するための自己進化型ルーブリック

スキルは、LLM エージェント、エンコード SOP、ドメイン ルール、ツール ワークフロー、スクリプト、および検証ルーチンの再利用可能な操作レイヤーになりつつあります。現実的なスキル リポジトリでは、スキルが重複しているため、信頼性の高いスキルの使用が困難になります。最終検証者の成功は、評価とトレーニングの両方にとって粗すぎます。エージェントは、注意をそらすスキルを選択する際に試行錯誤を繰り返したり、必要な手順をスキップしたり、ワークフローを誤って構成したり、最終チェックを省略したりする可能性があるためです。エージェントのスキル使用を評価および強化するための自己進化型ルーブリック フレームワークである SkillCoach を紹介します。 SkillCoach は、実際のロールアウトからスキルに基づいたプロセスのルーブリックを導き出し、スキルの選択、スキルのフォロー、スキルの構成、およびスキルに基づいた反映の 4 つの次元に沿って軌道を評価します。外部検証器を別個の結果信号として保持し、プロセスの品質と偶発的なタスクの成功を区別できるようにします。進化したルーブリックは、高品質のトレーニング軌道を選択するためのプロセス監視としてさらに機能します。実験によると、進化したルーブリックは評価の品質を大幅に向上させ、最終的な精度によって隠れていた失敗を明らかにし、結果のみのフィルタリングよりも強力な監視シグナルを提供してエージェントのスキルの使用を強化することを示しています。

原文 (English)

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use difficult. Final verifier success is too coarse for both evaluation and training, since an agent may pass through trial and error while selecting distractor skills, skipping required steps, composing workflows incorrectly or omitting final checks. We introduce SkillCoach, a self-evolving rubric framework for evaluating and enhancing agentic skill-use. SkillCoach derives skill-grounded process rubrics from real rollouts and evaluates trajectories along four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection. It keeps the external verifier as a separate outcome signal, allowing process quality to be distinguished from accidental task success. The evolved rubrics further serve as process supervision for selecting high-quality training trajectories. Experiments show that evolved rubrics substantially improve evaluation quality, expose failures hidden by final accuracy, and provide stronger supervision signals than outcome-only filtering for enhancing agentic skill-use.

13:00 JSTLLM/生成AI

Spec-AUF: マスクされたブロック ドラフタのトレーニング推論の不整合下で失敗するまでトレーニングを受け入れる

投機的デコードでは、ターゲット モデルが左から右に検証するトークンのブロックをドラフトし、受け入れられた最長のプレフィックスのみをコミットすることで、自己回帰生成を加速します。ブロック (DLM スタイル) ドラフターは、ブロック全体を並行して予測します。これは高速ですが、最初の拒否後に推論がすべてのトークンを破棄しても、ゴールドの継続に対するすべてのポジションを監視するフルブロックのクロスエントロピーでトレーニングされます。最近の受け入れを意識した目標は、フルブロック損失の重み付けを変更することでこの問題を解決します。代わりに、教師による強制学習を、受け入れられた接頭辞に監督が集中する方法の動機として使用します。マスクのみのブロック ドラフターには、ゴールド プレフィックス コンディショニング用の入力側チャネルがないため、AUF は、ドラフターの最初の予測障害を通じてのみクロス エントロピー サポートを維持することにより、損失側でプレフィックスに依存する監視を近似します。 AUF は、CE サポートに対する単一の個別の変更です。補助的な目的や検証ツールのロールアウトはなく、推論パイプラインや正確性コントラクトへの変更もありません。 Qwen3-8B の固定ドラフター バックボーンとサービング設定内で、AUF は 6 つのベンチマークの平均である DFlash ドラフターの平均放出長 $\tau$ を 2.40 から 2.61 に引き上げ、すべてのベンチマークで増加し、ドミノの 2 ブランチ ヘッド (2.56 から 2.68) に移行します。 2 つの発見により全体像が鮮明になります。減衰のみのベースラインは、共有ブロック マスクではより高いトークン精度に達しますが、デコードは悪くなります。もう 1 つは、DFlash では、AUF がサポートを切り捨てると、標準の指数関数的な位置減衰の重み付けが経験的に不活性になります。

原文 (English)

Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters

Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix. Block (DLM-style) drafters predict the whole block in parallel, which is fast but trained with a full-block cross-entropy that supervises every position against the gold continuation -- even though inference discards every token after the first rejection. Recent acceptance-aware objectives patch this by reweighting the full-block loss; we instead use teacher-forced learning as a motivation for how supervision should concentrate on the accepted prefix. A mask-only block drafter has no input-side channel for gold-prefix conditioning, so AUF approximates that prefix-sensitive supervision on the loss side by keeping the cross-entropy support only through the drafter's first predicted failure. AUF is a single, detached change to the CE support -- no auxiliary objective, no verifier rollouts, and no change to the inference pipeline or the exactness contract. Within fixed drafter backbones and serving settings on Qwen3-8B, AUF raises the DFlash drafter's average emitted length $\tau$, averaged over six benchmarks, from 2.40 to 2.61, with a gain on every benchmark, and transfers to Domino's two-branch head (2.56 to 2.68). Two findings sharpen the picture: the decay-only baseline reaches higher token accuracy on the shared block mask yet decodes worse, and on DFlash, once AUF truncates the support, the standard exponential position-decay weighting becomes empirically inert.

13:00 JSTLLM/生成AI

LLM 統合アプリケーションの複雑さのメトリクスを再考する: ソース コードを超えて

LLM 統合アプリケーションは自然言語プロンプトとプログラム コードを融合しており、その実行時の動作の多くはコード自体ではなくプロンプト層で発生します。ただし、既存の複雑さの指標はコード レベルでのみ機能するため、この動作ロジックは完全に無視されます。このようなアプリケーションのプロンプト層とコード層の両方の複雑さを評価するために設計された最初のツール、HECATE を紹介します。 HECATE の中心となるのは、すべてのプロンプトを意図された動作の仕様として解釈する、Hoare ロジックにインスピレーションを得た形式主義である Prompt-as-仕様 です。このツールは、公開されている分類法全体で特定された 25 の複雑さの次元に基づいて、52 の候補指標を生成します。私たちは、複雑さの経験的な代用としてバージョン履歴から得られるメンテナンス活動に依存して、18 のオープンソース リポジトリから収集された 118 のコンポーネントに対して各メトリクスを評価し、コード サイズが考慮されると重要性を失うメトリクスはすべて破棄します。このテストに耐えられる指標は 10 個だけです。 7 つは新しく導入されたセットに属します。純粋なボリュームを測定するのではなく、LLM コール サイト、メモリ属性、プロンプト テンプレートなどの構造的に異なる要素 (構造的幅と呼ばれる属性) をそれぞれ集計します。生き残っている 3 つの従来の指標のうち、RFC は同様の幅指向の特性を示しますが、Halstead N と V はサイズの残余効果としてのみ生き残っています。当社の最高のパフォーマンス指標は 3 つすべてを上回っています。重要なのは、プロンプト層のメトリクスは、最も強力なコードレベルのメトリクスが共変量として追加された場合でも重要性を維持し、プロンプトの複雑さをそれ自体の次元として確立することです。 6 つの保持されたリポジトリにまたがる 20 のコンポーネントに関する最終検証では、2 つの最もパフォーマンスの高いメトリクスが引き続きメンテナンス作業を予測し、トレーニング セットを超えた一般化可能性をサポートしていることが示されています。

原文 (English)

Rethinking Complexity Metrics for LLM-Integrated Applications: Beyond Source Code

LLM-integrated applications blend natural language prompts with program code, and much of their runtime behavior originates in the prompt layer rather than in the code itself. Existing complexity metrics, however, operate solely at the code level and therefore overlook this behavioral logic entirely. We present HECATE, the first tool designed to assess complexity in both the prompt and code layers of such applications. Central to HECATE is Prompt-as-Specification, a Hoare-logic-inspired formalism that interprets every prompt as a specification of intended behavior. Grounded in 25 complexity dimensions identified across published taxonomies, the tool generates 52 candidate metrics. We assess each metric against 118 components collected from 18 open-source repositories, relying on maintenance activity derived from version history as an empirical proxy for complexity, and discard any metric that loses significance once code size is accounted for. Only ten metrics withstand this test. Seven belong to our newly introduced set; rather than measuring sheer volume, each tallies structurally distinct elements, such as LLM call sites, memory attributes, and prompt templates, an attribute we call structural breadth. Of the three surviving conventional metrics, RFC exhibits a similar breadth-oriented character, while Halstead N and V survive only as a residual effect of size; our top-performing metrics exceed all three. Crucially, the prompt-layer metrics retain significance even when the strongest code-level metric is added as a covariate, establishing prompt complexity as a dimension in its own right. A final validation on 20 components spanning six held-out repositories shows that the two best-performing metrics continue to predict maintenance effort, supporting their generalizability beyond the training set.

13:00 JSTエージェントClaude

ContextSniper: リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ

大規模な言語モデル エージェントは実際のリポジトリの問題を修復できますが、ファイル全体の読み取り、広範な検索、および有用な証拠が無関係なコードやログと混在する長いターミナル出力に多額のコンテキスト バジェットを費やすことがよくあります。このペーパーでは、リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ層である ContextSniper について説明します。 AntTrail の広範なエージェント メモリ エンジンのコーディングに特化したものとして、ContextSniper は正確な証拠選択のための Sniper 機能を実装しています。候補コードとランタイム証拠を取得し、ハイブリッド取得信号でランク付けし、意図を認識したコンテキスト ゲートを通じて長い出力をフィルタリングし、プロンプトの外で回復可能なソース コンテキストを保持しながらコンパクトな証拠パケットを返します。 OpenClaw と Claude Code を備えた SWE-bench Lite 上で、ホスト エージェント条件ごとに 50 タスクの実行を使用して ContextSniper を評価しました。 ContextSniper は、OpenClaw の場合、トークンの総使用量を 51.5%、ログに記録されたコストを 36.4% 削減し、Claude Code の場合、トークンの総使用量を 38.9%、推定コストを 27.3% 削減します。提出された解決率は、OpenClaw の場合は 26.0% から 24.0% に、Claude Code の場合は 32.0% から 30.0% にわずかに減少しました。 ContextSniper のパイロット テスト スクリプトは、https://github.com/Calluking/ContextSniper でオープンソース化されています。

原文 (English)

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair

Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs. This paper presents ContextSniper, AntTrail's token-efficient code memory layer for repository-level program repair. As the coding specialization of AntTrail's broader agent memory engine, ContextSniper implements the Sniper feature for precision evidence selection: it retrieves candidate code and runtime evidence, ranks it with hybrid retrieval signals, filters long outputs through an intention-aware context gate, and returns compact evidence packets while preserving recoverable source context outside the prompt. We evaluate ContextSniper on SWE-bench Lite with OpenClaw and Claude Code, using 50 task runs per host-agent condition. ContextSniper reduces total token use by 51.5% and logged cost by 36.4% for OpenClaw, and reduces total token use by 38.9% and estimated cost by 27.3% for Claude Code. Submitted-resolution rates decrease slightly, from 26.0% to 24.0% for OpenClaw and from 32.0% to 30.0% for Claude Code. ContextSniper's pilot testing scripts are open-sourced at https://github.com/Calluking/ContextSniper

13:00 JSTエージェント

ElephantAgent: エージェントシステムにおけるコンテキスト状態の連続性

エージェント システムは、外部ツールを呼び出し、永続メモリを維持することで機能を強化します。ただし、これらの外部依存関係により、新たな攻撃対象領域が生じます。最近のツールおよびメモリ ポイズニング攻撃では、悪意を持って作成されたツール記述子とポイズニングされたメモリがエージェントの動作を密かに偏らせる可能性があることが示されています。これらの脅威は、計画と実行のためのエージェントのコンテキスト状態に検証可能な連続性が欠如しているという、より深い問題を反映しています。コンテキスト状態のポイズニングを防ぐためにコンテキスト状態の継続性を強制するプロトコルである ElephantAgent を紹介します。以前の状態継続メカニズム (Nimble など) からインスピレーションを得た ElephantAgent は、この保護をエージェント システムの進化するコンテキスト状態に拡張します。コンテキスト状態を、エージェントのコンテキスト全体 (ツールの状態やメモリなど) の限定されたセキュリティ クリティカルなサブセットとして定義します。各クエリを処理する前に、ElephantAgent はローカル コンテキスト状態のダイジェストを再計算し、最新の承認されたダイジェストと照合して検証します。 ElephantAgent は、複製された信頼できるハードウェアを使用して、承認されたコンテキスト状態遷移の線形化可能な台帳を維持し、帯域外の状態改ざんを検出します。帯域内セマンティック悪用に対処するために、ElephantAgent はさらに履歴追跡機能を提供し、条件付きの事後監査と既知の正常な以前の状態への回復を可能にします。

原文 (English)

ElephantAgent: Contextual State Continuity in Agentic Systems

Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependencies introduce novel attack surfaces. Recent tool and memory poisoning attacks show that maliciously crafted tool descriptors and poisoned memory can covertly bias agent behavior. These threats reflect a deeper issue: the lack of verifiable continuity in the agent's contextual state for planning and execution. We present ElephantAgent, a protocol that enforces Contextual State Continuity to defend against contextual state poisoning. Inspired by prior state-continuity mechanisms (e.g., Nimble), ElephantAgent extends this protection to the evolving contextual state of agentic systems. We define the contextual state as the bounded, security-critical subset of the agent's entire context (e.g., tool state and memory). Before processing each query, ElephantAgent recomputes the digest of the local contextual state and verifies it against the latest authorized digest. Using replicated trusted hardware, ElephantAgent maintains a linearizable ledger of authorized contextual state transitions and detects out-of-band state tampering. To handle in-band semantic abuse, ElephantAgent additionally provides Historical Traceability, enabling conditional post-hoc audit and recovery to a known-good prior state.

13:00 JSTLLM/生成AIエージェント

A-TMA: 長期エージェント メモリにおけるステートアウェア メモリ障害の分離

長期記憶により、LLM エージェントは永続的なアシスタントとして機能しますが、ユーザーの事実は変化します。有用な記憶システムは、現在何が真実で、何が以前は真実で、何が変化したかを知っていなければなりません。私たちは \emph{ゴーストメモリ} という状態調整の失敗を研究します。これは、古い事実、現在の事実、および遷移の事実がメモリバンク内に共存し、検索中に混合されたままとなり、応答モデルを誤解させるものです。私たちは、メモリ システムはバンクの維持、取得、応答時間の解決という 3 つのレベルから理解し、最適化する必要があると主張します。私たちは、既存のメモリ システム用の状態認識オーバーレイである ATMA を提案します。 ATMA は、置き換えられた記録と移行記録をバンクに保持し、クエリで要求された状態ビューの証拠パケットを構築し、現在、履歴、および移行ラベルを QA に公開します。さらに、最終的な QA の精度によってゴースト メモリが発生する箇所が隠れてしまう可能性があるため、バンク、検索、および回答レベルの失敗を分離して評価することを求めます。この障害を測定可能にするために、ゴースト メモリの競合が多いベンチマークである LTP (LoCoMo Temporal Plus) を構築し、長い会話の一般化のために LoCoMo で評価します。 LTP では、Graphiti+ATMA により、Graphiti よりも絶対値​​ 0.240 だけ競合精度が向上します。 LoCoMo では、Graphiti+ATMA により時間 F1 が 0.0295 から 0.1705 に上昇します。ゲインはホストに依存しますが、明示的な状態の役割により、最終的な QA 精度によって隠されたメモリ障害を削減できることを示しています。

原文 (English)

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory

Long term memory lets LLM agents act as persistent assistants, but user facts change. A useful memory system must know what is true now, what used to be true, and what changed. We study \emph{ghost memory}, a state coordination failure in which old, current, and transition facts coexist in the memory bank, remain mixed during retrieval, and mislead the answer model. We argue that memory systems should be understood and optimized from three levels: bank maintenance, retrieval, and answer time resolution. We propose ATMA, a state aware overlay for existing memory systems. ATMA keeps superseded and transition records in the bank, builds evidence packets for the query's requested state view, and exposes current, historical, and transition labels to QA. We further call for decoupled evaluation of bank, retrieval, and answer level failures, since final QA accuracy can hide where ghost memory occurs. To make this failure measurable, we build LTP (LoCoMo Temporal Plus), a conflict heavy benchmark for ghost memory, and evaluate on LoCoMo for long conversation generalization. On LTP, Graphiti+ATMA improves conflict accuracy by 0.240 absolute over Graphiti. On LoCoMo, Graphiti+ATMA raises temporal F1 from 0.0295 to 0.1705. The gains are host dependent, but they indicate that explicit state roles can reduce memory failures hidden by final QA accuracy.

13:00 JSTLLM/生成AIエージェント

アトミック タスク グラフ: エージェントの計画と実行のための統合フレームワーク

LLM ベースのエージェントは、複雑な複数ステップのタスクを解決する強力な可能性を示していますが、既存のパフォーマンスの向上は、多くの場合、より大きなバックボーン モデルへのスケーリングかタスク固有の微調整のいずれかに依存しています。前者はかなりの計算コストがかかりますが、後者は通常、さまざまなタスク間での一般化が不十分です。プロンプトベースの制御はトレーニング不要で広く適用可能ですが、既存の方法では依然としてサブタスク間の入出力依存関係がテキストの軌跡に暗黙的に残されているため、検証された中間結果の再利用が困難になります。これらの制限に対処するために、計画と実行のための統合制御フレームワークであるアトミック タスク グラフ (ATG) を提案します。具体的には、ATG は依存関係を公開し、再利用をサポートするために明示的なグラフを維持します。計画中に、高レベルのタスクを再帰的にサブタスクに分解し、展開を追跡できる一連の有向非巡回グラフ (DAG) を形成します。実行中、ATG によって公開される依存関係により、独立したブランチを並行して実行できるため、実行効率が向上します。障害が検出されると、ATG はグラフの進化履歴を利用してエラーの原因を特定し、影響を受けた領域のみを修復し、検証済みの領域を変更せずに保持します。実験の結果、ATG は 7B ~ 8B バックボーンのみを使用した 3 つのインタラクティブ ベンチマークにわたって、成功率と実行効率において強力なベースラインを常に上回っていることが示されています。

原文 (English)

Atomic Task Graph: A Unified Framework for Agentic Planning and Execution

LLM-based agents have shown strong potential for solving complex multi-step tasks, yet existing performance improvements often rely on either scaling to larger backbone models or task-specific fine-tuning. The former incurs substantial computational costs, while the latter typically generalizes poorly across different tasks. Although prompt-based control is training-free and broadly applicable, existing methods still leave input-output dependencies between subtasks implicit in textual trajectories, making verified intermediate results difficult to reuse. To address these limitations, we propose Atomic Task Graph (ATG), a unified control framework for planning and execution. Specifically, ATG maintains an explicit graph to expose dependencies and support reuse. During planning, it recursively decomposes a high-level task into subtasks, forming a sequence of directed acyclic graphs (DAGs) whose evolution can be traced. During execution, the dependencies exposed by ATG allow independent branches to be executed in parallel, thereby improving execution efficiency. When failures are detected, ATG leverages the graph evolution history to localize the error source and repair only the affected region, preserving validated regions unchanged. Experiments show that ATG consistently outperforms strong baselines in success rate and execution efficiency across three interactive benchmarks using only 7B-8B backbones.

13:00 JST研究/論文

OntoLearner: 大規模な言語モデルを使用したオントロジー学習のためのモジュール式 Python ライブラリ

オントロジー学習 (OL) は、テキストから構造化された知識モデルを自動的に構築することを目的としていますが、その進歩は手法、ドメイン、評価実践にわたって断片的なままです。数十年にわたる研究にもかかわらず、OL には体系的な評価とオントロジーへのアクセスのための共有インフラストラクチャが不足しています。この不在により研究の進歩が妨げられ、研究が細分化され、OL の中心的な課題はほとんど解決されないままになっています。 OntoLearner は、オントロジー アクセス、大規模言語モデル (LLM) 主導の学習パイプライン、および標準化されたベンチマークを統合する、モジュール式のクロスドメインで初のフレームワークです。 OntoLearner は、22 のドメインにまたがる 180 の機械可読オントロジーをリリースし、用語タイピング、分類法の発見、および非分類学的関係の抽出という 3 つのコア OL タスク用に、トレーニング/開発/テスト分割を備えたパイプライン対応のデータセットを提供します。このインフラストラクチャを使用して、OL の大規模な実証研究を実施し、ドメインとタスク全体で 22 の検索モデルと 12 の LLM を評価します。その結果は、OL の中心的な課題を再構成する発見に収束します。つまり、障害モードは、モデルのサイズやアーキテクチャの洗練さではなく、オントロジーの複雑さによってスケールされます。主なボトルネックはモデルの機能ではなく、モデルが知識をエンコードする方法とオントロジーが知識を編成する方法の間の構造的な不一致です。これらの調査結果は、OntoLearner によって有効になっているクロスドメイン、マルチタスクのベンチマークを通じて効果的な OL に到達できることを証明しています。 OntoLearner は、https://github.com/sciknoworg/OntoLearner/ でオープンソース (MIT ライセンス) です。

原文 (English)

OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models

Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, and evaluation practices. Despite decades of research, OL lacks a shared infrastructure for systematic evaluation and ontology access. This absence has hindered progress and fragmented research, leaving the central challenges of OL largely unaddressed. We introduce OntoLearner, a modular, cross-domain, and first-of-its-kind framework that unifies ontology access, large language model (LLM)-driven learning pipelines, and standardized benchmarking. OntoLearner releases 180 machine-readable ontologies spanning 22 domains and provides pipeline-ready datasets with train/dev/test splits for three core OL tasks: term typing, taxonomy discovery, and non-taxonomic relation extraction. Using this infrastructure, we conduct a large-scale empirical study of OL, evaluating 22 retrieval models and 12 LLMs across domains and tasks. The results converge on a finding that reframes the central challenge of OL: failure modes scale with ontological complexity rather than model size or architectural sophistication. The primary bottleneck is not model capability, but a structural mismatch between how models encode knowledge and how ontologies organize it. These findings establish that effective OL is reachable through the cross-domain, multi-task benchmarking enabled by OntoLearner. OntoLearner is open-source (MIT license) at https://github.com/sciknoworg/OntoLearner/.

13:00 JSTLLM/生成AI画像/動画生成

オンライン再帰的 MLLM 編集のためのマルチモーダル ナレッジ編集スコープの一般化

オンラインでマルチモーダルな知識編集を行うには、オーバーヘッドを制限し、無関係な動作への影響を最小限に抑えながら、ビジュアルテキスト修正の継続的なストリームをマルチモーダル大規模言語モデル (MLLM) に注入する必要があります。既存のエディターは主に編集の信頼性と長期的な安定性を重視していますが、各編集の意味境界を制御することはほとんどありません。編集後の動作と内部ニューロン活動のパイロット分析では、信頼性の高い編集の背後にある範囲のギャップが明らかになりました。インスタンス レベルの成功は、有効なクロスモーダル バリアントへの転送を保証するものでも、無関係な入力への漏洩を防ぐものでもありませんが、編集関連のクロスモーダル応答はより深いセマンティック層に集中しています。したがって、オンライン MLLM 編集を単なるインスタンスの修正から各編集の伝播境界の制御に再構成し、編集スコープの一般化を定式化します。この目的を達成するために、我々は、各更新をモダリティローカル吸収ブランチと証拠ゲート型共有一般化ブランチに分解するスコープ認識オンラインエディタである ScopeEdit を提案します。ローカル ブランチは安定した編集吸収をサポートしますが、共有ブランチは視覚的証拠とテキスト証拠が十分に一致している場合にのみクロスモーダル伝播を可能にします。どちらのブランチも、直交する低ランク空間でスコープ分離された書き込みジオメトリを実行し、シャーマン-モリソン再帰を介してブランチごとのプリコンディショナーを維持し、編集ごとに一定のオーバーヘッドを生成します。多様なベンチマーク、長期編集ストリーム、MLLM バックボーン、現実世界の VLKEB シナリオ、複雑なビジョン言語アーキテクチャにわたる広範な実験により、ScopeEdit が編集の信頼性、安定性、オンライン効率を維持しながら、スコープ内のクロスモーダル転送とスコープ外の局所性の間のトレードオフを一貫して改善することが示されました。私たちのコードは https://github.com/lab-klc/ScopeEdit で入手できます。

原文 (English)

Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing

Online multimodal knowledge editing requires injecting a continual stream of visual-textual corrections into multimodal large language models (MLLMs) with bounded overhead and minimal disruption to unrelated behaviors. Existing editors mainly emphasize edit reliability and long-horizon stability, but rarely control the semantic boundary of each edit. Our pilot analyses of post-edit behaviors and internal neuronal activities reveal a scope gap behind reliable edits: instance-level success neither guarantees transfer to valid cross-modal variants nor prevents leakage to unrelated inputs, while edit-related cross-modal responses concentrate in deeper semantic layers. Therefore, we formulate Edit-Scoped Generalization, reframing online MLLM editing from merely correcting an instance to controlling the propagation boundary of each edit. To this end, we propose ScopeEdit, a scope-aware online editor that decomposes each update into a modality-local absorption branch and an evidence-gated shared generalization branch. The local branch supports stable edit absorption, whereas the shared branch enables cross-modal propagation only when visual and textual evidence are sufficiently aligned. Both branches perform scope-separated write geometries in orthogonal low-rank spaces and maintain branch-wise preconditioners via Sherman--Morrison recursions, yielding constant per-edit overhead. Extensive experiments across diverse benchmarks, long-horizon edit streams, MLLM backbones, real-world VLKEB scenarios, and complex vision-language architectures show that ScopeEdit consistently improves the trade-off between in-scope cross-modal transfer and out-of-scope locality, while preserving edit reliability, stability and online efficiency. Our code is available at https://github.com/lab-klc/ScopeEdit.

13:00 JSTエージェントロボティクス

アイデンティティのドリフトを伴わないエピソードからセマンティックへの統合

長期にわたる適応型インテリジェント エージェントは、知識の統合と情報の完全性の間の構造的緊張に直面しています。メモリの統合は従来、エージェント変更操作として扱われます。つまり、モデルが微調整され、プロンプトが書き換えられ、ポリシーが蒸留され、または将来の動作を制御するコンテキストに反映が追加されます。規制された自律型展開では、エージェントは特定の暗号化された認証 ID にバインドされるコミットメントと監査契約に基づいて動作するため、これは責任となります。我々は、統合をプランナーやアイデンティティマニフェストの突然変異としてではなく、出力が個別にアドレス指定可能な意味論的知識層であるエピソード記憶上の決定論的関数 f: M^ep -> M^sem として扱うことを提案します。 ID ハッシュは M^sem を読み取らないため、統合ではエージェントの認定された ID を変更せずに知識が更新されます。エージェント表現の正式な説明を行い、マニフェストのハッシュ入力セットの構造補題を通じて同一性不変性を証明し、出力が明示的信頼性とサポートイベント来歴を備えた監査可能なデータベース行である決定論的集計アルゴリズムを指定し、フィールドごとの正確性、統合パス全体でのバイト同一性、および非生産的なプランナー試行の平均 79.82% 削減 (BCa 95%) を実証する合成実験で構築を検証します。 10 シードにわたる CI [78.02%、81.49%])、校正されたベイジアン縮小ベースラインに対する。この構築は、自律型エージェントの知識更新規律であり、実行中のケーススタディとして具体化されたサービス エージェントを使用して、エージェントの認定された ID が動作期間全体を通じてバイト同等のままでありながら、レッスンがクエリ可能な事実として蓄積されます。

原文 (English)

Episodic-to-Semantic Consolidation Without Identity Drift

Long-running adaptive intelligent agents face a structural tension between knowledge consolidation and information integrity. Memory consolidation is conventionally treated as an agent-changing operation: a model is fine-tuned, a prompt rewritten, a policy distilled, or a reflection appended to the context that governs future behaviour. In regulated autonomic deployment this is a liability because the agent operates under commitments and audit contracts that bind to a specific, cryptographically certified identity. We propose to treat consolidation not as a mutation of the planner or the identity manifest, but as a deterministic function f: M^ep -> M^sem over episodic memory whose output is a separately addressable semantic knowledge layer; the identity hash does not read M^sem, so consolidation updates knowledge without changing the agent's certified identity. We give a formal account of the agent representation, prove identity invariance through a structural lemma on the manifest's hash-input set, specify a deterministic aggregation algorithm whose outputs are auditable database rows with explicit confidence and supporting-event provenance, and validate the construction with synthetic experiments demonstrating per-field correctness, byte-equal identity across consolidation passes, and a mean 79.82% reduction in unproductive planner attempts (95% BCa CI [78.02%, 81.49%] across 10 seeds) against a calibrated Bayesian-shrunk baseline. The construction is a knowledge-update discipline for autonomic agents in which lessons accumulate as queryable facts while the agent's certified identity remains byte-equal across its operational lifetime, with an embodied service agent as the running case study.

13:00 JSTエージェント

検索拡張マルチエージェント O&M アシスタントによるバッテリー エネルギー貯蔵システムの追跡可能な障害診断

大規模バッテリーエネルギー貯蔵システム (BESS) では、アラーム、セルレベルの測定、デバイストポロジ、診断テーブル、過去の事例、保守文書を組み合わせた O&M の決定が必要です。監視プラットフォームはしきい値違反にフラグを立てることができますが、多くの場合、電圧の不一致、抵抗ドリフト、短絡リスク、容量発散、または熱異常に介入が必要かどうかを説明できません。このダイジェストは、検索強化されたマルチエージェント推論を使用して、運用データ、ドメイン知識、視覚的証拠、およびレポート生成を結び付ける、追跡可能な BESS 障害診断アシスタントを示します。 BESS 固有のタスク ルーティング、スキーマに制約された自然言語データベース アクセス、ハイブリッド テキストと画像の検索、および証拠に基づいた回答合成により、信頼性が向上します。ルーティング、データベース アクセス、および診断推論についての予備的な内部評価が報告されます。

原文 (English)

Traceable Fault Diagnosis for Battery Energy Storage Systems via Retrieval-Augmented Multi-Agent O&M Assistant

Large-scale battery energy storage systems (BESSs) require O&M decisions that combine alarms, cell-level measurements, device topology, diagnostic tables, historical cases, and maintenance documents. Monitoring platforms can flag threshold violations, but they often cannot explain whether voltage inconsistency, resistance drift, short-circuit risk, capacity divergence, or thermal abnormality needs intervention. This digest presents a traceable BESS fault-diagnosis assistant that uses retrieval-augmented multi-agent reasoning to connect operational data, domain knowledge, visual evidence, and report generation. Reliability is improved through BESS-specific task routing, schema-constrained natural-language database access, hybrid text-image retrieval, and evidence-based answer synthesis. Preliminary internal evaluation is reported for routing, database access, and diagnostic reasoning.

13:00 JSTLLM/生成AI

InduceKV: KV メモリの誘導によるマルチモーダル LLM の固定フットプリント継続的適応

マルチモーダルな大規模言語モデルは、進化するタスクやドメインに適応する必要がありますが、パラメータの繰り返し更新やリプレイ ストアの増加により、時間の経過とともに適応状態が蓄積される可能性があるため、制限されたデプロイメント フットプリントの下で継続的に改善することは依然として困難です。私たちは、固定フットプリントの継続的適応を研究します。バックボーン モデルは変更されず、タスク固有の更新は外部化されますが、デプロイされた適応状態は固定メモリ バジェット内に維持されます。我々は、選択された各トレーニング プレフィックスをアテンション対応のメモリ エントリとして保存する検索ベースの手法である InduceKV を提案します。これは、凍結された検索キーと、モデルのセルフ アテンション キャッシュに追加できるコンパクトな層ごとのキーと値 (KV) ペイロードで構成されます。厳格なメモリ バジェットの下で、InduceKV は 2 レベルの選択を通じてコン​​パクトな誘導セットを構築します。軽量のキャリブレーションは取得に適していますが、選択されたメモリは現在のタスクの尤度、アンカー ベースの保持、および凍結された取得スペースでのカバレッジのバランスをとります。タスクの増分命令チューニング、継続的な VQA、ドメインの増分適応、生涯にわたるマルチモーダル命令チューニングにわたって、InduceKV は、一致するメモリ バジェットの下で、PEFT、MoE、リプレイ、およびプロンプト取得のベースラインを一貫して改善します。さらに、バックボーンの一致、ステージ 1 の CoIN、コンピューティングの一致、およびスケーラビリティの診断についても報告し、この向上がバックボーンの強化、リプレイのみ、または無制限の候補プールによるものではないことを示しています。

原文 (English)

InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficult because repeated parameter updates or growing replay stores can accumulate adaptation state over time. We study fixed-footprint continual adaptation: the deployed adaptation state is kept under a fixed memory budget, while the backbone model is left unchanged and task-specific updates are externalized. We propose InduceKV, a retrieval-based method that stores each selected training prefix as an attention-ready memory entry, consisting of a frozen retrieval key and compact layerwise key--value (KV) payloads that can be appended to the model's self-attention cache. Under a strict memory budget, InduceKV constructs a compact inducing set through bilevel selection: a lightweight calibration is fit for retrieval, while the selected memory balances current-task likelihood, anchor-based retention, and coverage in the frozen retrieval space. Across task-incremental instruction tuning, continual VQA, domain-incremental adaptation, and lifelong multimodal instruction tuning, InduceKV consistently improves over PEFT, MoE, replay, and prompt-retrieval baselines under matched memory budgets. We further report backbone-matched, stage-1 CoIN, compute-matched, and scalability diagnostics, showing that the gains are not due to a stronger backbone, replay alone, or an unbounded candidate pool.

13:00 JST研究/論文

継続的なマルチモーダル学習における隠れた忘却: 正確性は維持されるがグラウンディングが失敗する場合

マルチモーダルな大規模言語モデルは、進化するタスクとドメインに継続的に適応する必要がありますが、標準的な継続学習指標は主に古い答えが正しいままであるかどうかを測定するため、マルチモーダル基盤の安定性はほとんど検討されていません。私たちはこの見落とされている障害モードを研究し、継続的に適応する MLLM が答えを保存できるかどうかだけでなく、視覚的、テキスト、OCR、チャート、および文書の証拠をどのように使用するかについても検討します。私たちは、モデルが別の、またはあまり根拠のない証拠チャネルに静かに移行しながら回答の精度が保たれる \emph{隠された証拠の使用の忘却} を特定し、リプレイのない信頼性が制約された継続的な学習フレームワークである \textsc{RCL} を提案します。 \textsc{RCL} は、以前のチェックポイントを行動参照として凍結し、反事実チャネル介入を通じて教師と生徒の証拠依存プロファイルを推定し、推論時間コストを追加することなくタスク学習、予測保存、信頼保存を共同で最適化します。 \textsc{RCL} は、CoIN、COAST、MCITlib、および証拠に敏感なマルチモーダル ストリーム全体にわたって、最終パフォーマンスを一貫して向上させ、リプレイフリー、PEFT、ルーティング、およびメモリ支援ベースラインの忘れを減らし、モダリティ依存度のドリフト、支配的な証拠の反転、および隠れた忘却率を大幅に低下させます。これらの結果は、堅牢な継続的なマルチモーダル学習には、単に答えそのものだけでなく、正解の背後にある証拠パスを保存する必要があることを示唆しています。

原文 (English)

Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined. We study this overlooked failure mode and ask whether a continually adapted MLLM can preserve not only what it answers, but also how it uses visual, textual, OCR, chart, and document evidence. We identify \emph{hidden evidence-use forgetting}, where answer accuracy is retained while the model silently shifts toward different or less grounded evidence channels, and propose \textsc{RCL}, a replay-free reliance-constrained continual learning framework. \textsc{RCL} freezes the previous checkpoint as a behavioral reference, estimates teacher and student evidence-reliance profiles through counterfactual channel interventions, and jointly optimizes task learning, prediction preservation, and reliance preservation without adding inference-time cost. Across CoIN, COAST, MCITlib, and an evidence-sensitive multimodal stream, \textsc{RCL} consistently improves final performance and reduces forgetting over replay-free, PEFT, routing, and memory-assisted baselines, while substantially lowering modality reliance drift, dominant evidence flips, and hidden forgetting rates. These results suggest that robust continual multimodal learning requires preserving the evidence path behind correct answers, not merely the answers themselves.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

PACE: エージェントの能力評価の代理

SWE-Bench や GAIA などのベンチマークで LLM エージェントを評価するには、費用と時間がかかり、複雑なインフラストラクチャが必要になる場合があります。 1 回の評価に数千ドルの費用がかかり、完了までに数日かかる場合があります。対照的に、個々の機能 (推論、コード生成など) をテストする非エージェント LLM ベンチマークは、高速かつ低コストで実行できます。このペーパーでは、高価なエージェント ベンチマークでのパフォーマンスが、慎重に選択されたアトミック評価インスタンスの少数のサブセットでのパフォーマンスによって正確に予測できるかどうかを調査します。 PACE は、既存の非エージェント評価からインスタンスを選択することでプロキシ ベンチマークを構築するフレームワークで、その集計スコアがエージェント ベンチマークでのモデル パフォーマンスを最も確実に予測するものを紹介します。アトミック機能にわたる候補インスタンスのプールを考慮すると、PACE は、ソース インスタンスのコンパクトなサブセットのモデルのスコアをターゲット エージェント ベンチマークのスコアにマッピングする回帰を当てはめます。サブセット自体は、2 つの相補的なインスタンス選択戦略、ターゲット関連性のローカル選択とグローバルに情報を提供するグローバル選択を組み合わせることによってキュレーションされます。このペーパーでは、PACE を 4 つのターゲット エージェント ベンチマークに適用し、このペーパーで評価する具体的なプロキシ ベンチマークである PACE-Bench を生成します。 14 のモデル、4 つのエージェント ベンチマーク、および 19 の非エージェント ベンチマークにわたる実験では、PACE-Bench が、リーブ ワンアウト相互検証 (LOOCV) 平均絶対誤差 (MAE) が 4% 未満、スピアマン相関が 0.80 以上、ペアワイズ モデル ランク付け精度が約 85% で、すべてエージェント評価コスト全体の 1% 未満でエージェント スコアを予測することが示されています。選択したプロキシ インスタンスをさらに分析し、各エージェント ベンチマークが独自に要求するスキルを明らかにします。 PACE を使用すると、担当者は、完全なエージェント評価のオーバーヘッドを発生させることなく、モデルの開発、選択、ルーティング中にエージェントのパフォーマンスの信頼できる推定値を取得できます。

原文 (English)

PACE: A Proxy for Agentic Capability Evaluation

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict model performances on agentic benchmarks. Given a pool of candidate instances spanning atomic capabilities, PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark. The subset itself is curated by combining two complementary instance-selection strategies, target-relevance local selection and globally informative global selection. We apply PACE to the 4 target agentic benchmarks in this paper, which yields PACE-Bench, the concrete proxy benchmark that we evaluate in the paper. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-out cross-validation (LOOCV) mean absolute error (MAE) under 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, all at much less than 1% of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing which skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.

13:00 JST研究/論文

最適決定木のグローバル分析のための代数モデル計数

Explainable AI でモデルの信頼性を確保するには、仮説空間の全体的な評価が必要です。我々は、代数決定木カウント (ADTC) と呼ばれる、最適および最適に近い決定木の徹底的な分析のための正式なフレームワークを提案します。知識表現における代数モデル計数 (AMC) に触発された ADTC は、最適化、計数、サンプリングなどのさまざまな分析タスクを、セミリング $R$ に対する統合された積和計算に再定式化します。決定木の仮説空間は最大深さ $\Delta$ に関して二重指数関数的ですが、動的計画法アルゴリズムは特徴量 $n$ で $O^*(n^{O(\Delta)})$ の時間計算量を達成します。ここで、$O^*$ は多項式因数を抑制します。複数のツリー メトリクスで構成される複雑な制約を処理するために、テンソル半リング上の畳み込み積を介して半リング値を集約するモデル動作テンソルを導入します。この代数的アプローチは、地球規模の状況と、精度、サイズ、公平性などの基準間のトレードオフを捉えるモデル プロファイルを効率的に構築します。実世界のデータセット上で当社のソフトウェア emtrees の有用性を実証し、ADTC がどのように機密領域における証拠に基づくモデル選択を容易にするかを示します。

原文 (English)

Algebraic Model Counting for Global Analysis of Optimal Decision Trees

Ensuring model reliability in Explainable AI requires a global assessment of the hypothesis space. We propose a formal framework for the exhaustive analysis of optimal and near-optimal decision trees, called Algebraic Decision Tree Counting (ADTC). Inspired by Algebraic Model Counting (AMC) in knowledge representation, ADTC reformulates diverse analytical tasks, such as optimization, counting, and sampling, into a unified sum-of-products computation over a semiring $R$. While the hypothesis space of decision trees is doubly exponential with respect to the maximum depth $\Delta$, our dynamic programming algorithm achieves $O^*(n^{O(\Delta)})$ time complexity in the number of features $n$, where $O^*$ suppresses polynomial factors. To handle complex constraints consisting of multiple tree metrics, we introduce model behavior tensors that aggregate semiring values via convolution products over a tensor semiring. This algebraic approach efficiently constructs a model profile that captures the global landscape and trade-offs between criteria such as accuracy, size, and fairness. We demonstrate the utility of our software, emtrees, on real-world datasets, illustrating how ADTC facilitates evidence-based model selection in sensitive domains.

13:00 JST研究/論文LlamaQwen

長い文脈の推論に対する証拠状態の報酬

長いコンテキストの推論では、モデルが長い入力に分散された証拠を見つけ、修正し、合成する必要があります。既存のロングコンテキスト RL 手法は通常、最終的な回答または静的な証拠の抽出に報酬を与え、中間アクションがモデルの証拠の状態をどのように変化させるかについてほとんどフィードバックを提供しません。私たちは、編集可能な証拠メモリを備えた強化学習フレームワークである Maven を提案します。 Maven は、回答条件付き証拠状態値を定義し、アクション レベルの状態遷移に報酬を与えます。追加アクションは限界利益と後知恵の貢献によって評価され、リンク アクションは証拠の相乗効果によって評価され、ドロップ アクションは誤解を招く証拠を削除した後の改善された回答サポートによって評価されます。これらの報酬は、GRPO の対応するアクション スパンに割り当てられます。 LongBench v2、LongReason、および RULER の Llama モデルと Qwen モデル全体で、Maven は結果のみの RL および証拠特定ベースラインを上回り、より十分な証拠セットを生成し、注意散漫の保持率を低下させます。私たちの結果は、ロングコンテキスト RL は、ワンショット証拠抽出ではなく、ステートフル証拠ナビゲーションを最適化することで恩恵を受けることを示しています。

原文 (English)

Evidence-State Rewards for Long-Context Reasoning

Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods usually reward final answers or static evidence extraction, offering little feedback on how intermediate actions change the model's evidence state. We propose Maven, a reinforcement learning framework with an editable evidence memory. Maven defines an answer-conditioned evidence-state value and rewards action-level state transitions: add actions are credited by marginal gain and hindsight contribution, link actions by evidence synergy, and drop actions by improved answer support after removing misleading evidence. These rewards are assigned to the corresponding action spans in GRPO. Across Llama and Qwen models on LongBench v2, LongReason, and RULER, Maven outperforms outcome-only RL and evidence-identification baselines, producing more sufficient evidence sets and lower distractor retention. Our results show that long-context RL benefits from optimizing stateful evidence navigation rather than one-shot evidence extraction.

13:00 JST研究/論文

SUNTA: サプライズベースのチャンキングによる階層型ビデオ予測

階層状態空間モデル (HSSM) は、シーケンスを時間的なチャンクに分割することにより、長期予測への有望なアプローチを提供します。ただし、そのパフォーマンスはチャンク境界がどのように決定されるかによって決まります。従来の HSSM は通常、固定長のチャンキングまたは類似性に基づく境界検出に依存していますが、これらの方法はデータの固有の時間構造と一致しないことがよくあります。私たちは、チャンク化は予測誤差によって駆動されるべきだと主張します。予測誤差は、より直接的に、長距離のコンテキストが必要になる時期を示します。それにもかかわらず、サプライズベースのチャンキングを HSSM に統合すると、エンドツーエンドのトレーニング中の階層の崩壊や、開ループ予測中のサプライズ信号の欠如など、重大な課題が生じます。これらの問題に対処するために、我々はサプライズベースの入れ子型時間抽象化 (SUNTA) を提案します。これは、サプライズ信号を保存するために分離されたトレーニング戦略を採用し、内部の不一致をトップダウンのサプライズメトリクスとして使用して、想像上のロールアウト内のチャンク境界を決定する方法です。 2D および 3D 環境でのビデオ予測タスクの実験では、SUNTA がベースラインを上回り、250 タイムステップにわたって正確な予測を独自に維持する一方、最初の 10 タイムステップ以内にすべてのベースラインが低下することが実証されました。

原文 (English)

SUNTA: Hierarchical Video Prediction with Surprise-based Chunking

Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how chunk boundaries are determined. While prior HSSMs typically rely on fixed-length chunking or similarity-based boundary detection, these methods often misalign with the intrinsic temporal structure of the data. We argue that chunking should instead be driven by prediction errors, which more directly indicate when longer-range context becomes necessary. Nevertheless, integrating surprise-based chunking into HSSMs introduces critical challenges, including hierarchical collapse during end-to-end training and the absence of surprise signals during open-loop prediction. To address these issues, we propose Surprise-based Nested Temporal Abstraction (SUNTA), a method that employs a decoupled training strategy to preserve surprise signals and uses internal inconsistency as a top-down surprise metric to determine chunk boundaries within imagined rollouts. Experiments on video prediction tasks in 2D and 3D environments demonstrate that SUNTA outperforms baselines, uniquely maintaining accurate predictions over 250 timesteps, whereas all baselines degrade within the first 10 timesteps.

13:00 JSTエージェント

ContextNest: 自律型 AI エージェントの検証可能なコンテキスト ガバナンス

自律型 AI エージェントは外部ナレッジ ストアへの依存度を高めていますが、ほとんどの検索パイプラインは、出所、バージョン ID、完全性、トレーサビリティ、またはポイントインタイムの再構成の永続的な保証なしで関連性を提供します。私たちはこれをコンテキスト ガバナンスとして形式化し、管理された AI 消費型ナレッジ ボールトのオープン仕様およびリファレンス実装である ContextNext を提示します。 ContextNext は、検索拡張生成 (RAG) を置き換えるものではありません。これは、検索の下にガバナンス層を提供し、検索システムが動作する前に、どの成果物が承認され、最新のもので、帰属可能であり、整合性が検証されているかを判断します。この仕様では、型指定された Markdown ドキュメントとメタデータ、決定論的な集合代数セレクター、contextnest:// URI 参照、SHA-256 ハッシュチェーンのバージョン履歴、グラフレベルのチェックポイント、モデル コンテキスト プロトコル (MCP) を介したライブ データのソース ノード、およびエージェント コンテキスト消費の監査トレースを組み合わせています。これらのメカニズムにより、組織はどのナレッジ バージョンがエージェント出力に情報を提供したか、またそれらのバージョンが使用時に AI に適格かどうかを再構築できます。我々は、2 つの対照実験からの最初の経験的結果を報告します。ガバナンス対取得の失敗モードを分離する古いバージョンの攻撃では、管理された選択により BM25 のスパース検索が厳密にパレート支配され、入力トークンの約 3 分の 1 のコストでより高い応答品質の合格率 (97% 対 93 ~ 90%) が得られます。 1,060 の文書コーパスに対する検索決定論実験では、決定論的セレクターと BM25 は同一のクエリを繰り返しても安定した文書セットを返します (Jaccard 1.0)。一方、密 + HNSW ベースラインはクエリの 80% で非決定的です (平均 Jaccard 0.611、最悪の場合 0.210)。これらの結果は、コンテキスト ガバナンスが障害モードに対処すること、取得品質だけでは解決できないことを示唆しています。コア エンジン、CLI、MCP サーバーをオープン ライセンスでリリースします。

原文 (English)

ContextNest: Verifiable Context Governance for Autonomous AI Agent

Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guarantees of provenance, version identity, integrity, traceability, or point-in-time reconstruction. We formalize this as context governance and present ContextNext, an open specification and reference implementation for governed AI-consumable knowledge vaults. ContextNext does not replace Retrieval-Augmented Generation (RAG); it supplies the governance layer beneath retrieval, determining which artifacts are approved, current, attributable, and integrity-verified before retrieval systems operate over them. The specification combines typed Markdown documents with metadata, deterministic set-algebraic selectors, contextnest:// URI references, SHA-256 hash-chained version histories, graph-level checkpoints, source nodes for live data through the Model Context Protocol (MCP), and audit traces of agent context consumption. These mechanisms let organizations reconstruct which knowledge versions informed an agent output and whether those versions were AI-eligible when consumed. We report first empirical results from two controlled experiments. In a stale-version attack isolating the governance-versus-retrieval failure mode, governed selection strictly Pareto-dominates BM25 sparse retrieval, with higher answer-quality pass rate (97% versus 93-90%) at about one-third the input-token cost. In a retrieval-determinism experiment over a 1,060-document corpus, deterministic selectors and BM25 return stable document sets across repeated identical queries (Jaccard 1.0), while a dense+HNSW baseline is non-deterministic on 80% of queries (mean Jaccard 0.611, worst case 0.210). These results suggest that context governance addresses failure modes retrieval quality alone is not designed to resolve. We release a core engine, CLI, and MCP server under open licenses.

13:00 JSTLLM/生成AI

ドメイン固有のLLMポストトレーニングを通じてフィットネスインテリジェンスを強化

科学的フィットネス コーチング (SFC) は通常、人間の専門家によって提供されるため、費用がかかり、多くの人が利用できないものとなっています。大規模言語モデル (LLM) の最近の進歩は、より包括的なフィットネス コーチングに大きな期待を示していますが、一般的な汎用 LLM を SFC に直接導入すると、重大な制限が明らかになります。これらのモデルにはドメイン固有の知識が十分に統合されていないことが多く、複雑な SFC シナリオでのパフォーマンスが低下します。このペーパーでは、SFC アプリケーションの信頼性とドメイン特化を向上させるために設計された一連のフィットネス LLM (8B および 32B パラメーター) である FitOne を紹介します。 Qwen3 基盤モデルに基づいて構築された FitOne は、厳密な知識エンジニアリングから得られた大規模で高品質のデータセットを使用し、継続的な事前トレーニング、教師あり微調整、強化学習で構成される 3 段階のポストトレーニング パイプラインを通じて開発されます。 ACSM-EP や NSCA-CSCS などのプロフェッショナル フィットネス認定試験、および知識推論や指示への従うことなどの一般的な能力について、FitOne の総合的な評価を実施します。実験結果によると、FitOne-8B/32B は強力な一般的な機能を維持しながら、Qwen3 ベース モデルと比較して、ACSM-EP および NSCA-CSCS 試験でそれぞれ最大 10.09%/9.29% および 12.73%/7.01% の平均向上を達成しています。さらに、詳細なアブレーション研究により、各トレーニング段階の必要性が確認され、分野の専門知識の強化と一般的な能力の維持のバランスをとる上でのパイプラインの有効性が強調されています。私たちは、この研究が LLM システムをより信頼性の高いフィットネス インテリジェンスに向けて前進させ、ドメイン固有の LLM の開発に関する将来の研究に影響を与えると信じています。

原文 (English)

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training

Scientific Fitness Coaching (SFC) is typically delivered by human professionals, making it costly and inaccessible to many. While recent advances in Large Language Models (LLMs) show considerable promise for more inclusive fitness coaching, directly deploying prevailing general-purpose LLMs in SFC reveals critical limitations. These models often lack sufficient domain-specific knowledge integration, leading to weak performance on complex SFC scenarios. In this paper, we introduce FitOne, a series of fitness LLMs (with 8B and 32B parameters) designed to improve reliability and domain specialization for SFC applications. Built upon the Qwen3 foundation models, FitOne is developed through a three-stage post-training pipeline consisting of continual pre-training, supervised fine-tuning, and reinforcement learning, using large-scale, high-quality datasets derived from rigorous knowledge engineering. We conduct comprehensive evaluations of FitOne on professional fitness certification exams, including ACSM-EP and NSCA-CSCS, as well as general capabilities such as knowledge reasoning and instruction following. Experimental results show that, while retaining strong general capabilities, FitOne-8B/32B achieves average improvements of up to 10.09%/9.29% and 12.73%/7.01% on the ACSM-EP and NSCA-CSCS exams, respectively, compared with the Qwen3 base models. Furthermore, in-depth ablation studies confirm the necessity of each training stage, highlighting the pipeline's effectiveness in balancing domain expertise enhancement with general ability retention. We believe this research advances LLM systems toward more reliable fitness intelligence and will inspire future research on developing domain-specific LLMs.

13:00 JSTエージェント研究/論文

コーディングエージェントは科学的な機械学習の論文を複製できる

科学的な機械学習の論文では通常、相対平均二乗誤差が 5% 未満である、または 95% の予測信頼区間がテスト データをカバーしているなど、計算上の主張が行われます。コーディングエージェントは、紙資料のみからそれらの主張を複製するように促されることもありますが、プロンプト自体では進行状況を確実に保存したり、生成された証拠が論文の主張を裏付けるかどうかを確認したりすることはありません。選択した各論文を記録された証拠とともにターゲットとして主張し、コーディング エージェント スキルとして実装するワークフローであるペーパー レプリケーションを紹介します。このワークフローにより、エージェントはこれらのターゲットを記録し、論文の手法を再構築し、計算実験を実行し、生成された出力を出所と論文の主張との比較にリンクし、一致する証拠が複製レポートのどこに現れるかを記録し、完了前に検証チェックに合格します。私たちは、4 つの科学機械学習論文にわたる 12 の独立した実行で紙の複製を評価しました。 12 個のワークスペースすべてが完了ゲートを通過し、記録された 158 個のターゲットすべてがレポート カバレッジと一致します。この完成したワークスペース状態であっても、反復実行では、論文がターゲットに分割される方法、ソース論文への数値的忠実度、レプリケーションの経過時間、最終証拠が受け入れられるまでに置き換えられる中間実行の数、および証拠を受け入れるために使用されるルールが異なります。紙の複製では、完了はエージェントの最終メッセージではなく、ワークスペースの証拠と検証チェックに依存します。

原文 (English)

Coding-agents can replicate scientific machine learning papers

Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predictive credible interval covers the test data. A coding agent can be prompted to replicate those claims from paper materials alone, but the prompt does not by itself reliably preserve progress or check whether generated evidence supports the paper's claims. We introduce Paper-replication, a workflow that makes each selected paper claim a target with recorded evidence, and implement it as a coding-agent skill. The workflow makes the agent record those targets, reconstruct the paper's method, run computational experiments, link generated outputs to provenance and comparisons with the paper's claims, record where matched evidence appears in the replication report, and pass validation checks before completion. We evaluate Paper-replication on twelve independent runs across four scientific machine learning papers. All twelve workspaces pass the completion gate, and all 158 recorded targets are matched with report coverage. Even in this completed workspace state, repeated runs differ in how papers are divided into targets, in numerical fidelity to the source papers, in elapsed replication time, in the number of intermediate executions replaced before final evidence is accepted, and in the rules used to accept evidence. Paper-replication makes completion depend on workspace evidence and validation checks rather than on the agent's final message.

13:00 JSTエージェント研究/論文

A$^{2}$utoLPBench: 逆KKT構築による自動生成されたエージェントフレンドリーなLPベンチマーク

ほとんどの LP-from-text ベンチマークは、手作業で書かれ、ラベルが付けられた文章題の静的データセットです。このようなデータセットがリリースされると、そのサイズと難易度は固定され、あらゆる問題が将来の LLM のトレーニング データに漏洩する可能性があります。 \textbf{A$^{2}$utoLPBench} は、プレーン テキストで記述された線形計画問題について LLM 駆動エージェントをテストするためのベンチマークです。まず実現可能な点と双対を選択し、次にその点が最適であり、客観的な値がわかっている問題を書き留めます。答えは構築によってわかり、ソルバー呼び出しや人間によるアノテーターは必要ありません。評価環境には、リファレンス ソルバー クリティカル ベースラインと、LLM 駆動エージェントが読み取るための使用手順が記述された Docker イメージがバンドルされています。これらを導入すると、どのエージェントでも 1 つのコマンドでベンチマークを実行し、調整されたスコアを取得できます。ベンチマークは固定データセットではなくジェネレーターであるため、固定データセットにはない特性を備えています。つまり、新鮮な問題の無制限の供給、$(n,m)$ によって設定された難易度ノブ、構築によって正しい正解の答え、人間によるオーサリングと比較した問題あたりの LLM 側のコストの低さ、独立したバッチ間で再現可能なスコア、新鮮なカットオフ後のシード範囲が使用された場合のトレーニング データ漏洩に対する耐性です。

原文 (English)

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction

Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its difficulty is fixed, and every problem can leak into the training data of future LLMs. We present \textbf{A$^{2}$utoLPBench}, a benchmark for testing LLM-driven agents on linear programming problems written in plain text. We first pick a feasible point and dual, then write down a problem for which that point is optimal and the objective value is known. The answer is known by construction, with no solver call and no human annotator. The evaluation environment bundles a reference solver-critic baseline and a Docker image whose usage instructions are written for an LLM-driven agent to read. With these in place, any agent can run the benchmark and get a calibrated score with one command. Because the benchmark is a generator rather than a fixed dataset, it has properties no fixed dataset can match: an unlimited supply of fresh problems, a difficulty knob set by $(n,m)$, ground-truth answers correct by construction, low LLM-side cost per problem relative to human authoring, repeatable scores across independent batches, and resistance to training-data leakage when fresh post-cutoff seed ranges are used.

13:00 JSTビジネス/資金調達研究/論文ClaudeGPT / ChatGPTGemini

専門家が作成した臨床推論タスクに関するフロンティア言語モデルのルーブリックベースの管理された比較

多肢選択式の医療ベンチマークはますます飽和状態にあり、HealthBench などの最近のルーブリックベースの評価では、オープンエンドの臨床パフォーマンスは解決にはほど遠いことが示されており、その「ハード」サブセットのトップスコアは 32% のままです。我々は、臨床医が作成した 4 つの専門分野 (麻酔、内科/家庭医学、救急医療、産科) にわたる 5 つの臨床シナリオからなる小規模で意図的に困難な評価データセットを提示します。各臨床シナリオには、臨床医が起草したゴールデンアンサーから作成された、アトミックで重み付けされた MECE ルーブリック (タスクごとに 25 ~ 62 の基準、合計 184 の基準) が付属しています。 GPT 5.4、Claude Opus 4.7、Gemini 3.1 Pro の 3 つのフロンティア モデルを評価します。平均ルーブリック合格率は、0.47 (Claude)、0.39 (GPT)、および 0.37 (Gemini) でした。中心となる所見は、臨床的優先順位の逆転です。最も重み付けされた (重み付け 5、重要) 基準は 32.4 ~ 41.7% でのみ合格しましたが、最も低い重み付け 1 基準は 80 ~ 90% で合格しました。 108 の重要 (重み 5) 基準のうち 56 (52%) を満たしたモデルはありませんでした。 3 人の LLM 自動評価者が、552 の等級付け基準の 92.8 ~ 94.7% について、専門家の適合/不適合ラベルを再現しました。私たちはこれを手法と予備調査結果の貢献として位置づけています。5 つのタスクは、大規模なベンチマークに開発する準備ができているスケーラブルで防御可能なパイプラインを示しています。

原文 (English)

A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks

Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its "Hard" subset top score remains 32%. We present a small, deliberately difficult evaluation dataset of five clinician-authored clinical scenarios spanning four specialties (anaesthesia, internal/family medicine, emergency medicine, and obstetrics), each accompanied by an atomic, weighted, MECE rubric (25-62 criteria per task; 184 criteria total) authored from a clinician-drafted golden answer. We evaluate three frontier models: GPT 5.4, Claude Opus 4.7, and Gemini 3.1 Pro. Mean rubric pass rates were 0.47 (Claude), 0.39 (GPT), and 0.37 (Gemini). The central finding is an inversion of clinical priority: the highest-weighted (weight-5, critical) criteria passed at only 32.4-41.7%, while low-stakes weight-1 criteria passed at 80-90%. 56 of 108 critical (weight-5) criteria (52%) were satisfied by no model. Three LLM autoraters reproduced expert met/not-met labels on 92.8-94.7% of 552 graded criteria. We position this as a methods-and-preliminary-findings contribution: the five tasks demonstrate a scalable, defensible pipeline ready to develop into a large-scale benchmark.

13:00 JSTLLM/生成AIエージェント

UA-ChatDev: 信頼性の高いソフトウェア開発のための不確実性を認識したマルチエージェント コラボレーション

ソフトウェア開発は、さまざまな役割を持つエージェント間の協力が必要な複雑なタスクです。大規模言語モデル (LLM) により、役割ベースのコラボレーションを活用して要件分析、コーディング、テスト、改良を自動化する自律的なマルチエージェント ソフトウェア開発フレームワークが可能になりました。しかし、既存のアプローチは通常、中間エージェントの出力が同様に信頼できると想定しているため、中間エージェントの出力は幻覚の伝播に対して脆弱なままとなり、開発の初期段階で生成された誤った決定が下流のエージェントに転送され、最終的なソフトウェアの品質に悪影響を及ぼします。この課題に対処するために、私たちは、不確実性の定量化をエージェントの対話に統合する、不確実性を認識したマルチエージェント ソフトウェア開発フレームワークである UA-ChatDev を提案します。トークンレベルのログ確率に基づいた軽量の不確実性推定メカニズムを導入して、エージェントの応答の信頼性を評価し、不確実性が許容レベルを超えた場合に、位相認識しきい値キャリブレーションを採用して、検索ベースの検証を選択的にトリガーします。 SRDD ベンチマークに関する広範な実験により、UA-ChatDev が、完全性、実行可能性、一貫性、および全体的な品質指標において、既存のシングルエージェントおよびマルチエージェントのソフトウェア開発フレームワークよりも一貫して優れていることが実証されました。さらなるアブレーション研究と通信分析により、不確実性を認識した相互作用によりコード実行の信頼性が向上することが確認されました。

原文 (English)

UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development

Software development is a complex task that demands cooperation among agents with diverse roles. Large language models (LLMs) have enabled autonomous multi-agent software development frameworks that leverage role-based collaboration to automate requirements analysis, coding, testing, and refinement. However, existing approaches typically assume that intermediate agent outputs are equally reliable, leaving them vulnerable to hallucination propagation, where incorrect decisions generated in early development phases are transferred to downstream agents and negatively impact final software quality. To address this challenge, we propose UA-ChatDev, an uncertainty-aware multi-agent software development framework that integrates uncertainty quantification into agent interactions. It introduces a lightweight uncertainty estimation mechanism based on token-level log probabilities to assess the confidence of agent responses and employs phase-aware threshold calibration to selectively trigger retrieval-based verification when uncertainty exceeds acceptable levels. Extensive experiments on the SRDD benchmark demonstrate that UA-ChatDev consistently outperforms existing single-agent and multi-agent software development frameworks across completeness, executability, consistency, and overall quality metrics. Further ablation studies and communication analyses verify that uncertainty-aware interactions enhance code execution reliability.

13:00 JSTエージェント

自律通信ネットワークにおける AI エージェントの決定のための重要度ベースのガード レール検証

完全自律通信ネットワーク (自律ネットワーク レベル 4 ~ 5) への進化では、AI/ML エージェントが人間の介入なしにリアルタイムでネットワークの意思決定を行うことが必要です。ただし、実際のネットワーク状態の変化を引き起こす前に個々の推論出力を傍受して検証するための標準化されたランタイム メカニズムは存在せず、誤った自律的な決定のリスクが生じます。このペーパーでは、AI 主導の決定を実行前に傍受して検証するための標準化可能なランタイム アーキテクチャである Guard Rail Validation (GRV) フレームワークを提案します。このフレームワークは、アクションの範囲、アクションの種類、サービスの重要度、エージェントの自律性レベル、可逆性、一時的な動作パターンなど、複数の重み付けされた側面にわたる意思決定を評価して、重要度のレベルを決定します。このレベルに基づいて、ログ付き実行、境界チェック、独立エージェント検証、またはマルチエージェントのコンセンサスなど、段階的な検証メカニズムが適用されます。このフレームワークはさらに、クリティカル度を重視した優先解決機能を備えたエージェント間の競合検出と、規制遵守のためのランタイム適合性ロギングを提供します (例: EU AI 法第 14 条)。アーキテクチャ、アルゴリズム手順、O-RAN 導入モデルを示し、通信における既知の AI/ML 攻撃に対する脅威の適用範囲を評価します。

原文 (English)

Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom Networks

The evolution toward fully autonomous telecommunications networks (Autonomous Network Levels 4-5) requires AI/ML agents to make real-time network decisions without human intervention. However, no standardized runtime mechanism exists to intercept and validate individual inference outputs before they trigger live network state changes, creating risks of erroneous autonomous decisions. This paper proposes the Guard Rail Validation (GRV) framework, a standardizable runtime architecture for intercepting and validating AI-driven decisions before execution. The framework evaluates decisions across multiple weighted dimensions -- including action scope, action type, service criticality, agent autonomy level, reversibility, and temporal behavioural patterns -- to determine a criticality level. Based on this level, graduated validation mechanisms are applied: execute-with-logging, bounds checking, independent agent validation, or multi-agent consensus. The framework additionally provides cross-agent conflict detection with criticality-weighted priority resolution and runtime conformance logging for regulatory compliance (e.g., EU AI Act Article 14). We present the architecture, algorithmic procedures, O-RAN deployment model, and evaluate threat coverage against known AI/ML attacks in telecommunications.

13:00 JSTLLM/生成AI

精製された OPSD: 考え方を失わずにポリシーに従って自己蒸留する

オンポリシー自己蒸留 (OPSD) は、LLM 推論を改善するための有望なパラダイムとして浮上しています。このパラダイムでは、参照ソリューションにアクセスできる特権を持つ教師が、生徒自身が生成した軌跡に対してトークンレベルの監督を提供します。しかし、OPSD は長い思考連鎖 (long-CoT) 推論モデルでは一貫して失敗し、これらのモデルが依存する反射的推論能力を不安定にする一方で、せいぜい限界利益しか得られないことがわかりました。教師の監視信号の新しい分解を通じて、根本原因を特定します。教師の監視は、参照固有のショートカットの丸暗記を促進する参照誘導コンポーネントによって支配されているのに対し、質問条件付きの推論伝達可能なコンポーネントは無視されるか、積極的に反対されます。この診断に基づいて、2 段階の解決策を提案します。まず、参照のみの教師 (質問のない参照に条件付けされた同じモデル) を構築して、監視信号の転送不可能な成分を分離します。この成分を差し引いた後の残差は、質問条件付きの推論転送可能な修正を捕捉します。次に、ポイントワイズ相互情報量 (PMI) をメカニズムとして使用し、この残差を、学習者が直接抽出できる適切な形式の PMI ターゲット分布に変換し、参照によるショートカットを除外します。 2 つのデータセットにわたる 4 つの長い CoT モデルの実験では、トレーニング全体を通してモデルの自然な認識的動作を維持しながら、ベース モデルと標準 OPSD の両方を上回る一貫した改善が実証されました。

原文 (English)

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-level supervision on the student's own generated trajectories. However, we find that OPSD consistently fails on long chain-of-thought (long-CoT) reasoning models, yielding at best marginal gains while destabilizing the reflective reasoning capability these models depend on. Through a novel decomposition of the teacher's supervision signal, we identify the root cause: the teacher's supervision is dominated by a reference-induced component that drives rote memorization of reference-specific shortcuts, while the question-conditioned, inference-transferable component is ignored or actively opposed. Based on this diagnosis, we propose a two-step solution. First, we construct a reference-only teacher (the same model conditioned on the reference without the question) to isolate the non-transferable component of the supervision signal; the residual after subtracting this component captures the question-conditioned, inference-transferable correction. Second, we use pointwise mutual information (PMI) as the mechanism to transform this residual into a well-formed PMI target distribution that the student can directly distill from, filtering out the reference-induced shortcut. Experiments on four long-CoT models across two datasets demonstrate consistent improvements over both the base model and standard OPSD, while preserving the models' natural epistemic behavior throughout training.

13:00 JSTエージェント

Copewell: 公平なメンタルウェルネスサポートのためのマルチエージェント群アーキテクチャ

精神的健康障害は世界中で 10 億人近くの人に影響を与えていますが、低・中所得国の 75% は労働力不足、コストの壁、偏見のために治療を受けていません。現在の AI を活用したウェルネス ソリューションは主にシングルモードの会話型インターフェイスに依存しているため、離脱率が高く、ユーザーの動的な感情状態に合わせて測定可能な即時的な軽減を提供できません。この論文では、人間中心の AI 原理を通じてメンタルウェルネス サポートへのアクセスを拡大するように設計された新しいマルチエージェント群システムである Copewell について紹介します。当社のアーキテクチャでは、次の 3 つの技術革新が導入されています。(1) アルゴリズムのバイアスを軽減するために、自己報告データ、生理学的データ、およびコンテキスト データを統合するマルチソース評価フレームワーク。 (2) ラッセルの感情周回モデルを使用した価性-覚醒感情マッピングにより、ユーザーを専門の AI エージェントにルーティングします。 (3) 会話によるサポートと証拠に基づく感覚健康プロトコルを組み合わせたデュアルモード介入の提供。私たちは、プライバシー最優先のアーキテクチャ、専任の倫理監督エージェントによる倫理監視の組み込み、精神保健専門家による参加型設計など、Copewell の開発の基礎となる社会工学的設計の考慮事項を検討します。早期の実践者の関与とベータ展開により、設計上の決定が通知され、将来の実証的評価の方向性が特定されます。この研究は、技術アーキテクチャが最初から公平性と安全性の原則をどのように運用できるかを実証することで、責任ある AI の議論に貢献します。

原文 (English)

Copewell: A Multi-Agent Swarm Architecture for Equitable Mental Wellness Support

Mental health disorders affect nearly one billion people globally, yet 75% of individuals in low- and middle-income countries receive no treatment due to workforce shortages, cost barriers, and stigma. Current AI-powered wellness solutions predominantly rely on single-mode conversational interfaces that suffer high abandonment rates and fail to provide measurable, immediate relief calibrated to users' dynamic emotional states. This paper presents Copewell, a novel multi-agent swarm system designed to expand access to mental wellness support through human-centered AI principles. Our architecture introduces three technical innovations: (1) a multi-source assessment framework integrating self-reported, physiological, and contextual data to mitigate algorithmic bias; (2) valence-arousal emotion mapping using Russell's Circumplex Model of Affect to route users to specialized AI agents; and (3) dual-mode intervention delivery combining conversational support with evidence-based sensory wellness protocols. We examine the sociotechnical design considerations underlying Copewell's development, including a privacy-first architecture, embedded ethical oversight through a dedicated Ethics Supervisor agent, and participatory design informed by mental health practitioners. Early practitioner engagement and beta deployment inform design decisions and identify directions for future empirical evaluation. This work contributes to responsible AI discourse by demonstrating how technical architecture can operationalize equity and safety principles from inception.

13:00 JSTLLM/生成AIエージェント

AgenticSTS: Long-Horizo​​n LLM エージェント用の境界付きメモリ テストベッド

長期的な LLM エージェントのメモリは、将来の各決定で参照できる内容に関する契約です。最も単純なコントラクトでは、過去の観察、ツールの呼び出し、および反映をすべてのプロンプトに追加します。これにより、以前のコンテキストに簡単にアクセスできるようになりますが、単一のメモリ コンポーネントの影響を分離することが困難なごちゃ混ぜの混合物になります。私たちは、代替の制限付きコントラクトを導入し、実装します。すべての決定は、生の相互決定記録が追加されることなく、型付き検索によって組み立てられた新しいユーザー メッセージから行われます。したがって、プロンプトは任意の長さのランにわたって境界を維持し、任意の単一層を単独でアブレーションできます。私たちは Slay the Spire 2 でコントラクトをインスタンス化します。Slay the Spire 2 はクローズドルールの確率的デッキ構築ゲームで、その実行には何百もの戦術的および戦略的決定が必要です。同じゲームのフロンティア LLM の公開オンライン ベンチマークでは、5 つの構成全体で最も低い難易度では勝利がゼロであると報告されており、開発者が報告した同じ難易度での人間の勝率は 16% です。タスクは難しいですが、飽和していません。私たちのハーネス内では、固定 A0 アブレーションは、トリガーされた戦略的スキルが有効になっているときに観察された最大の違いを示しています。無店舗ベースラインは 3/10 ゲームに勝ち、スキル レイヤーを追加すると 6/10 です。このサンプルサイズでは、比較は統計的に決定的なものではなく、方向性を持ったものになります (フィッシャーの正確な p\約 0.37)。クロスバックボーン プローブと公開の蓄積コンテキスト ベースラインは、コントラクト変数自体の制御されたテストではなく、操作上の比較として報告されます。私たちは、再現可能なテストベッドをリリースします。条件タグを含む 298 の完了した軌跡、フリーズしたメモリ/スキルのスナップショット、プロンプト レコード、および分析スクリプトです。エージェントの設計と、明示的なメモリ層が長期的な LLM エージェントの意思決定をどのように形成するかを研究するための、検証済みの再利用可能な方法論です。

原文 (English)

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest contract appends past observations, tool calls, and reflections to every prompt, which makes prior context easy to access but also turns it into a jumbled mixture in which the effect of any single memory component is hard to isolate. We introduce and instrument an alternative bounded contract: every decision is made from a fresh user message assembled by typed retrieval, with no raw cross-decision transcript appended. The prompt thus stays bounded across runs of any length, and any single layer can be ablated in isolation. We instantiate the contract in Slay the Spire 2, a closed-rule stochastic deck-building game whose runs require hundreds of tactical and strategic decisions. A public online benchmark of frontier LLMs on the same game reports zero wins at the lowest difficulty across five configurations, and the developer-reported human win rate at the same difficulty is 16%; the task is hard but not saturated. Within our harness, a fixed-A0 ablation shows the largest observed difference when triggered strategic skills are enabled: the no-store baseline wins 3/10 games and adding the skill layer 6/10. At this sample size the comparison is directional rather than statistically decisive (Fisher exact p\approx0.37); a cross-backbone probe and public accumulating-context baselines are reported as operational comparisons rather than controlled tests of the contract variable itself. We release a reproducible testbed: 298 completed trajectories with condition tags, frozen memory/skill snapshots, prompt records, and analysis scripts -- an agent design and a validated, reusable methodology for studying how explicit memory layers shape long-horizon LLM-agent decisions.

13:00 JST研究/論文

線形注意力のための海馬: 再発状態が忘れてしまったものを正確に記憶

線形注意および状態空間言語モデルは、プレフィックスを固定サイズの反復状態に圧縮し、損失の多い正確なメモリを犠牲にして O(1) メモリを生成します。多くのキーと値の関連付けが競合すると、以前のファクトが上書きされ、ニードル リコールが低下します。相補学習システムにインスピレーションを得て、海馬に線形注意を補います。 HOLA (海馬線形注意) は、通常のデルタ ルール状態を圧縮メモリとして保持し、境界付きの正確な KV キャッシュを追加して、セミパラメトリックなテスト時メモリを形成します。状態は線形圧縮可能な構造をモデル化し、キャッシュにはその状態を強制すべきではない関連付けが保存されます。キャッシュは学習されたエビクション モジュールなしで書き込み、実際に状態にコミットされた予測残差である大きなベータ * ||e|| のトークンを保持します。分離された RMSNorm ガンマ キャッシュ読み取りにより、これらの正確な KV ペアがソフト平均化ではなくシャープ検索に変わります。 150 億の SlimPajama トークンでトレーニングされた 3 億 4,000 万のパラメーターで、HOLA は Wikitext の複雑度を 27.32 から 22.92 (-16.1%) に下げ、全注意の Transformer++ (26.88) を下回り、LAMBADA の複雑度を 30.95 から 30.26 に改善しました。また、最高の線形インコンテキスト取得を実現し、GDN や RULER の干し草の中の針の一致する HOLA+ リーセンシー キャッシュよりもはるかに堅牢なままです (トレーニング長の 16 倍)。

原文 (English)

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets

Linear-attention and state-space language models compress the prefix into a fixed-size recurrent state, yielding O(1) memory at the cost of a lossy exact memory: when many key--value associations compete, earlier facts are overwritten and needle recall degrades. Inspired by Complementary Learning Systems, we give linear attention a hippocampal complement. HOLA (Hippocampal Linear Attention) keeps the usual delta-rule state as a compressive memory and adds a bounded exact KV cache, forming a semiparametric test-time memory: the state models linearly compressible structure, while the cache stores associations that should not be forced through that state. The cache writes without a learned eviction module, keeping tokens with large beta * ||e||, the prediction residual actually committed to the state; a decoupled RMSNorm-gamma cache read then turns these exact KV pairs into sharp retrieval rather than soft averaging. At 340M parameters trained on 15B SlimPajama tokens, HOLA lowers Wikitext perplexity from 27.32 to 22.92 (-16.1%), below a full-attention Transformer++ (26.88), and improves LAMBADA perplexity from 30.95 to 30.26. It also achieves the best linear in-context retrieval and remains much more robust than GDN or a matched HOLA+recency cache on RULER needle-in-a-haystack recall out to 32k tokens (16x its training length).

13:00 JSTLLM/生成AIエージェント研究/論文

グラウンデッド自律研究: 最先端の計算物理学におけるコーパスから原稿までのフォールトトレラントな LLM パイプライン

自律型研究エージェントは、実行によって調整が行われる機械学習サンドボックスでのエンドツーエンドの LLM 自動化を実証しました。最先端の物理科学は決定的に異なります。あらゆる方法論の選択の根底には物理的推論があり、ツールチェーンは文書化されていないことが多く、キャリブレーションは外部の文献アンカーから来なければなりません。足場のないエージェントは引用するものの対決せず、内部事前情報からのもっともらしい検証不可能な結果を​​幻覚します。我々は、11,083 件の最近の物性物理学 arXiv 論文のコーパスから、3 つの実質的な物理学の知見 (ここでは変磁性圧電磁性に関する) を含む出版グレードの原稿までエンドツーエンドで実行されるパイプラインを提示します。エージェントは、コーパスをマッピングすることによって研究の方向性を自律的に構想し、出版された参考文献を再現することによって方法論を調整し、新しい第一原理計算を実行し、原稿を執筆します。全体を通して文献に基づいており、ディスク上の状態のみを共有する 6 フェーズの 47 の新鮮な文脈のセッションと 2,162 の文献相談イベントが含まれます。フォールト トレランスは冗長性から生まれます。フレッシュ コンテキストの分離、分散グラウンディング、および敵対的レビューによって、単一のセッションが見逃しているものをキャッチします。パイロットの前後の段階は完全に自律的であり、パイロットでは、再現が失敗した場合にのみ、限定的な人間の介入が必要になります。これは、科学的な指示ではなく、運用上の知識のキュレーションです。 2 つのペアの故障モード (プレアーキテクチャ ベースラインとパイロットなしアブレーション) は、動作可能な接地メカニズムとしてキャリブレーション チェックポイントで構造的に強制された数値的対立を分離します。プリミティブ、特徴付けられた故障モード、定量化された介入パターンは、計算物理学を超えた一か八かの科学領域における自律的研究の基礎を築きます。

原文 (English)

Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics

Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration. Frontier physical science differs categorically: physical reasoning underlies every methodology choice, toolchains are often underdocumented, and calibration must come from external literature anchors - which unscaffolded agents cite but do not confront, hallucinating plausible, unverifiable results from internal priors. We present a pipeline that runs end-to-end from a corpus of 11,083 recent condensed-matter physics arXiv papers to a publication-grade manuscript with three substantive physics findings (here on altermagnetic piezomagnetism): the agent autonomously conceives a research direction by mapping the corpus, calibrates methodology by reproducing published references, conducts novel first-principles computations, and writes the manuscript - grounded in literature throughout, across 47 fresh-context sessions in six phases sharing only on-disk state, with 2,162 literature-consultation events. Fault tolerance emerges from redundancy: fresh-context isolation, distributed grounding, and adversarial review catch what any single session misses; pre- and post-pilot stages are fully autonomous, and pilot requires bounded human intervention only at reproduction failures - operational knowledge curation, not scientific direction. Two paired failure modes - a pre-architecture baseline and a no-pilot ablation - isolate structurally enforced numerical confrontation at calibration checkpoints as the operative grounding mechanism. The primitives, characterized failure modes, and quantified intervention pattern lay a foundation for autonomous research in high-stakes scientific domains beyond computational physics.

13:00 JST研究/論文

DRIFTLENS: パーソナライズされた言語モデルにおける記憶誘発推論ドリフトの測定

パーソナライゼーションにより、モデルがユーザーに伝える内容が変わります。反応を正当化するために使用される推論の軌道も変更できることを示します。最新の LLM は、ユーザーの属性、設定、以前のコンテキストを保存し、この情報を将来のプロンプトに挿入することによって、インタラクションをパーソナライズします。私たちは、そのような記憶が、単一の真実の答えが存在しない自由形式の質問に対する推論を再構成するかどうかを研究します。この効果を定量化するために、表現された各推論ステップを値カテゴリにマッピングし、質問の記憶のない軌道と注入されたユーザー属性記憶の下での軌道との間の乖離を測定する、グラウンドトゥルースフリーのフレームワークである DRIFTLENS を導入します。まず、DRIFTLENS が内容のない実用的なノイズと実質的な推論の変更を区別することを検証します。年齢、職業、障害を含む 4 つの LLM と 10 のユーザー属性カテゴリにわたって、ユーザー属性の記憶は、最終的な回答が流暢で、主題に合致しており、もっともらしいままである場合でも、各モデルの実用的なノイズ フロアを超える中規模から大規模な推論のドリフトを引き起こします。次に、ドリフトを軽減するための GRPO および DPO ベースのポストトレーニング方法を評価します。どちらもドリフトを軽減しますが、どちらも均一に支配的ではありません。下流の能力、有用性、指示への従うことへの影響は、モデルと報酬に依存します。これらの結果は、記憶に起因する推論ドリフトは測定可能であり、部分的にのみ緩和されるパーソナライズされた言語モデルの障害モードであることを示唆しています。

原文 (English)

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question's no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model's pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.

13:00 JSTエージェント

安全性が重要なリアルタイム自律システム向けのハードウェア強制セマンティック調整

エージェント AI の最近の進歩により、大規模な言語モデル、世界モデル、最適化エンジン、特殊なニューラル アーキテクチャ、自律プラットフォーム、および人間のオペレーターを統合する、ますます複雑な自律システムが生み出されています。現在の研究の多くは推論能力の向上に焦点を当てていますが、セーフティクリティカルなリアルタイム展開では、不確実性の下で同時に動作する異種コンポーネント間の制限された検証可能な調整も必要です。ソフトウェアを介した調整は、制限された遅延、決定論的な調整、強制的な安全保証が不可欠な領域において根本的な制限をもたらします。したがって、選択された調整セマンティクスがフィールド プログラマブル ゲート アレイ (FPGA) を介してハードウェア レベルで直接実装される、ハードウェア強制セマンティック調整アーキテクチャを提案します。このアプローチは、意味論的推論を対話管理から分離するトピックベースのコミュニケーション スペース ペトリ ネット (TB-CSPN) フレームワークに基づいて構築されています。このアプローチでは、選択された TB-CSPN 調整メカニズムが FPGA プリミティブにマッピングされ、ハードウェア ネイティブのセマンティック調整層が作成されます。焦点は高速化ではなく、時間同期、セマンティック ゲート、認可制約、および限定された調整動作をハードウェアで直接強制することにあります。セマンティック推論は引き続き適応的でソフトウェア駆動型ですが、埋め込まれた調整セマンティクスは決定論的になります。

原文 (English)

Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

Recent advances in agentic AI are producing increasingly complex autonomous systems that integrate large language models, world models, optimization engines, specialized neural architectures, autonomous platforms, and human operators. While much current research focuses on improving reasoning capabilities, safety-critical real-time deployment also requires bounded and verifiable coordination among heterogeneous components operating concurrently under uncertainty. Software-mediated coordination presents fundamental limitations in domains where bounded latency, deterministic coordination, and enforceable safety guarantees are essential. Hence, we propose a hardware-enforced semantic coordination architecture in which selected coordination semantics are implemented directly at the hardware level via field-programmable gate arrays (FPGAs). The approach builds on the Topic-Based Communication Space Petri Net (TB-CSPN) framework, which separates semantic reasoning from interaction management. In this approach, selected TB-CSPN coordination mechanisms are mapped onto FPGA primitives, creating a hardware-native semantic coordination layer. Focus is not on acceleration, but on enforcing temporal synchronization, semantic gating, authorization constraints, and bounded coordination behavior directly in hardware. Semantic reasoning remains adaptive and software-driven, while embedded coordination semantics become deterministic.

13:00 JSTエージェントGemma

制約によるステアビリティ: コーディングエージェントのスケーラブルな監視のための基盤

コーディングエージェントは有能です。人間の監視がボトルネックになっています。制約のないエージェントはセキュリティ リスクをもたらし、コードベースのスケーラビリティを損ない、人間によるレビューのコストが増大します。私たちは、大規模なヒューマン エンジニアリング チームの管理に何十年も同じ方法が使用されてきたと主張します。アクセス制御、ネットワーク ポリシー、ツールによって強制される厳格なコーディング規約です。コーディング エージェントに直接転送され、最近のエージェント スキャフォールディングよりも (トークンで) 安価です。この原則に基づいて最初から最後までのシステムをスケッチし、スケーラブルな監視の下で制御された実験を報告します。少人数のレビュー担当者 (Gemma 4 e4b) が 11 個のバックドアが挿入された Python コードベースを検査します。再現率は 54.5% (制約なし、ツールなし) から 90.9% (制約された基板と約 200-LoC の「docs」 CLI) まで上昇し、基板とツールは独立して寄与します。私たちは意図的に Python を選択しています。言語がデフォルトで提供する保証が最も少ないところでは、基板レベルの監視のゲインが最大になります。この原則は Rust などの言語にも適用されます。

原文 (English)

Steerability via constraints: a substrate for scalable oversight of coding agents

Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make human review increasingly costly. We argue that the same methods used for decades to manage large human engineering teams: access control, network policies, strict coding conventions enforced by tooling; transfer directly to coding agents, and are cheaper (in token) than recent agentic scaffolding. We sketch a start-to-end system on this principle, and report a controlled experiment in scalable oversight: a small reviewer (Gemma 4 e4b) inspects a Python codebase containing 11 inserted backdoors. Recall rises from 54.5% (unconstrained, no tools) to 90.9% (constrained substrate plus a ~200-LoC `docs` CLI), with substrate and tools contributing independently. We choose Python deliberately: substrate-level oversight gains are largest where the language gives the fewest guarantees by default; the principles extend to languages like Rust.

13:00 JSTLLM/生成AIQwen

RFM-AGOP による高速多次元拒否部分空間

大規模言語モデル (LLM) でのアクティベーションのステアリングとモニタリングは、安全性と解釈可能性の両方を目的として使用されることが増えています。初期の研究では、単一の線形方向に沿って行動がエンコードされていると想定されていましたが、最近の発見では、有害なクエリへの回答の拒否などの複雑な行動が多次元の部分空間に存在することが示唆されています。ただし、これらの部分空間を抽出する既存の方法は計算コストが高く、長い推論トレースを生成する推論モデルでは法外なものになります。効率的に計算できる再帰特徴マシン (RFM) アルゴリズムをプローブ情報に基づいた初期化に適応させることで、推論 (Qwen 3) モデルと非推論 (Qwen 2.5) モデルに基づいて、多次元の拒否部分空間を数秒で特定できます。 RFM はより高速な部分空間の特定を可能にしますが、アブレーション タスクでは代替手段よりも優れたパフォーマンスも示しました。さまざまな方法で発見された部分空間間の関係をより深く理解するために、さらなる研究が計画されています。確認されれば、RFM は LLM の既存の部分空間抽出方法を安価でスケーラブルに補完できる可能性があります。

原文 (English)

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along single linear directions, but recent findings suggest complex behaviours, such as the refusal to answer harmful queries, live in multi-dimensional subspaces. However, existing methods for extracting these subspaces are computationally expensive, which becomes prohibitive on reasoning models who produce long reasoning traces. By adapting the Recursive Feature Machine (RFM) algorithm -- which can be computed efficiently -- with a probe-informed initialization, we are able to identify the multi-dimensional refusal subspace in seconds, on reasoning (Qwen 3) and non-reasoning (Qwen 2.5) models. While RFM allows for faster subspace identification, it also showed better performances on the ablation task than its alternatives. More work is planned to better understand the relations between subspaces found by different methods. If confirmed, RFM could be a cheap and scalable complement to existing subspace-extraction methods in LLMs.

13:00 JSTLLM/生成AI画像/動画生成

非マンハッタン環境におけるテキスト駆動の 3D 屋内シーン合成

大規模言語モデル (LLM) は、マンハッタン環境の 3D 屋内合成において優れた機能を実証しています。しかし、既存の手法では、非マンハッタン設定ではもっともらしいオブジェクトのレイアウト パターンをキャプチャできないことがよくあります。その主な理由は、非直交の空間関係をモデル化するのに苦労し、高い幾何学的違反と低い物理的忠実度につながるためです。この課題に対処するために、私たちは、複雑な非マンハッタン環境内で物理的に妥当な屋内シーンを生成するように設計された新しいテキスト駆動フレームワークである SPG-Layout を提案します。具体的には、まずオブジェクト分布の統計的事前分布を利用してトレーニング プロセスをガイドし、環境の理解と忠実度を高めます。さらに、ヒューマン デザインのワークフローを反映して、大きなオブジェクトの配置を優先する階層レイアウト戦略を採用し、それによってレイアウト違反を大幅に最小限に抑えます。これらのコンポーネントを相乗させることにより、SPG-Layout は意味論的な現実性と物理的な妥当性のバランスの取れた最適化を実現します。これらの複雑な環境でのパフォーマンスを評価するために、500 の多様な非マンハッタン環境で構成される新しいベンチマークを構築しました。広範な実験により、SPG-Layout はマンハッタン環境と非マンハッタン環境の両方で既存の方法よりも一貫して大幅に優れたパフォーマンスを発揮することが実証されました。コードは公開されます。

原文 (English)

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible object layout patterns in non-Manhattan settings, primarily because they struggle to model non-orthogonal spatial relationships, leading to high geometric violations and low physical fidelity. To address this challenge, we propose SPG-Layout, a novel text-driven framework designed to generate physically plausible indoor scenes within complex non-Manhattan environments. Specifically, we first utilize statistical priors of object distributions to guide the training process, enhancing environmental understanding and fidelity. Furthermore, mirroring human design workflows, we adopt a hierarchical layout strategy that prioritizes the placement of large objects, thereby substantially minimizing layout violations. By synergizing these components, SPG-Layout achieves a balanced optimization of semantic realism and physical plausibility. To evaluate performance in these complex settings, we constructed a new benchmark comprising 500 diverse non-Manhattan environments. Extensive experiments demonstrate that SPG-Layout consistently and significantly outperforms existing methods across both Manhattan and non-Manhattan environments. The code will be publicly released.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

大規模言語モデルを使用した Linux/bash 試験の自動採点: 4 レベルの認知分類アプローチ

コマンドライン試験のスケーラブルで信頼性の高い採点は、コンピューティング教育において依然として課題となっています。登録者数の増加により手動採点が困難になり、ルールベースの自動採点者は部分的な単位認定、同等の解法、または構文のバリエーションを処理できなくなります。この論文では、短い Linux/bash コマンド応答を採点する際に、4 つのフロンティア大規模言語モデル (GPT、Claude Opus、Gemini、および GLM) が専門家の判断に近似できるかどうかを評価します。この研究では、情報検索 (L1) と基本的なファイル操作 (L2) から構造操作 (L3) と高度なシステム管理 (L4) に至るまで、認知の複雑さと操作への影響を組み合わせた 4 レベルの認知分類法を採用しています。モデルは、コンピュータ エンジニアリングの 2 年生の学生からの 1,200 件の実際の回答を 3 人の専門講師が独自に採点し、最小限のベースラインとルーブリック強化バージョンの 2 つのプロンプト バージョンでテストされました。ルーブリックに基づくプロンプトを備えた Gemini~3.0 Pro は、人間と AI の最も高い一致を達成しました (ICC(3,1) = 0.888、MAE = 0.10、Bland-Altman バイアス = -0.014)。分類レベルが増加するにつれて一致度は一貫して低下し、より高いレベルで最大の不一致が見られました。すべてのモデルにわたって、ルーブリックの品質はプロバイダーの選択よりも大きな影響を及ぼし、構造化されたプロンプトにより一貫して同意が向上しました。これらの結果は、質問の複雑さが、LLM が正確に採点する際に直面する難しさを信頼できる予測因子であることを示しており、どの質問が AI 支援による採点に適しているか、どの質問が人間によるレビューを必要とするかを決定するための原則に基づいた分類法に基づいたフレームワークを確立するとともに、移転可能な評価プロトコルと迅速なテンプレートも提供します。

原文 (English)

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation. This paper evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The study adopts a four-level cognitive taxonomy that combines cognitive complexity and operational impact, ranging from information retrieval (L1) and basic file manipulation (L2) to structural operations (L3) and advanced system management (L4). The models were tested with two prompt variants, a minimal baseline and a rubric-enhanced version, on 1200 real responses from second-year Computer Engineering students independently graded by three expert instructors. Gemini~3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman bias = -0.014). Agreement declined consistently as taxonomy level increased, with the largest discrepancies at higher levels. Across all models, rubric quality had a larger effect than provider choice, with structured prompts consistently improving agreement. These results show that question complexity is a reliable predictor of the difficulty LLMs face in grading accurately, and they establish a principled, taxonomy-based framework for determining which questions are suitable for AI-assisted grading and which require human review, while also providing a transferable evaluation protocol and prompt templates.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達GPT / ChatGPT

EvoPolicyGym: 対話型環境での自律的なポリシー展開の評価

自律型エージェントは、フィードバックを通じて実行可能なポリシーを改善することがますます期待されていますが、既存の評価では、このプロセスが最終スコアに組み込まれたり、無制限のソフトウェア エンジニアリングの進歩と混同されたりすることがよくあります。ハーネス モデル エージェントが固定のインタラクション バジェットの下で実行可能なポリシー システムを繰り返し編集する、制御された評価設定である Autonomous Policy Evolution を導入します。この設定は、エージェントが探索されたポリシーを反復的に改善する方法を評価するコンパクトな対話型 RL 環境から構築されたベンチマークである EvoPolicyGym でインスタンス化されます。 EvoPolicyGym スイートでは、GPT-5.5 は 16 環境すべてで最強の総合ランク スコアと上位 2 位のパフォーマンスを達成しました。 EvoPolicyGym は、リーダーボードの結果以外にも、エージェントがどのように予算を割り当て、フィードバックをパラメトリック調整に変換するかを区別する軌跡レベルの診断も提供します。これらの分析は、強力な自律的なポリシーの進化は、孤立したタスクの勝利だけではなく、タスクに適したメカニズムを発見し、制限されたフィードバックの下でポリシーを洗練することに依存していることを示しています。

原文 (English)

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget. We instantiate this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies. On the EvoPolicyGym suite, GPT-5.5 achieves the strongest aggregate rank score and top-two performance on all 16 environments. Beyond leaderboard results, EvoPolicyGym also provides trajectory-level diagnostics that distinguish how agents allocate budget, convert feedback into parametric tuning. These analyses show that strong autonomous policy evolution depends not only on isolated task wins, but on discovering task-appropriate mechanisms and refining policies under bounded feedback.

13:00 JST研究/論文

G-RRM: リカレント推論モデルを使用したシンボリック ソルバーのガイド

この研究では、より大きな問題サイズに対する改善された外挿を示す RRM のシンボル等変インスタンス化である SE-RRM に焦点を当てます。我々は、SE-RRMと制約充足問題の記号ソルバーを統合する神経記号的アプローチ「再帰推論モデルによるガイド」(G-RRM)を提案する。 SE-RRM は、完全なソリューション提案を生成し、グローバルに正しいソリューションを生成するバックトラッキングや SAT ベースの手法 (Glucose 4.1 や CaDiCaL 3.0.0 など) などの古典的なシンボリック ソルバーをガイドするニューラル ソルバーとして機能します。私たちは主に、G-RRM によるニューラル ガイダンスがシンボリック ソルバーの検索効率を向上させるタイミングを調査します。 % 私たちの実験は、G-RRM の有効性が 2 つの条件に依存することを示しています。1 つ目は、潜在的な利益を明らかにするために、問題インスタンスには広大な組み合わせ検索空間がなければなりません。2 つ目は、ソルバー アーキテクチャが、ニューラル ヒントが不完全な場合に回復するために分岐選択肢を動的に上書きできなければなりません。これらの条件が当てはまると、ガイダンスによって競合カウントの中央値がゼロになり、実測速度が大幅に向上します。$9\times9$ の Sudoku では、SE-RRM が $91.1\%$ のインスタンスを正しく解決し、バックトラッキングは $33.3\time$、Glucose 4.1 では $1.70\time$ (中央値、$p<0.001$) 加速され、Glucose 4.1 では速度が維持されます。パーフェクトヒントの $25\times$ グリッドで $1.17\times$ の高速化。対照的に、CaDiCaL 3.0.0 では、ランタイムがオーバーヘッドに支配され、挿入された分岐ヒントを上書きするのではなく常に尊重しますが、大幅な高速化 (中央値 $1.02\times$, n.s.) は見られず、$9\times9$ ではわずかに有意な平均速度の低下 ($0.90\times$) さえ見られます。これらの結果は、神経誘導が実際の高速化につながる体制を概説します。

原文 (English)

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models

In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to larger problem sizes. We propose a neuro-symbolic approach, ``Guiding with Recurrent Reasoning Models'' (G-RRM), which integrates SE-RRMs with symbolic solvers for constraint satisfaction problems. SE-RRMs act as neural solvers that generate full solution proposals and guide classical symbolic solvers, such as backtracking or SAT-based methods like Glucose 4.1 and CaDiCaL 3.0.0, that produce globally correct solutions. Centrally, we investigate when neural guidance with G-RRM improves the search efficiency of symbolic solvers. % Our experiments show that the efficacy of G-RRM depends on two conditions: first, the problem instances must have an expansive combinatorial search space to expose potential gains, and second, the solver architecture must be capable of dynamically overwriting its branching choices to recover when neural hints are imperfect. When these conditions hold, guidance drives median conflict counts to zero and yields significant wall-clock speedups: on $9\times9$ Sudoku, where the SE-RRM correctly solves $91.1\%$ of instances, backtracking accelerates by $33.3\times$ and Glucose 4.1 by $1.70\times$ (median, $p<0.001$), with Glucose 4.1 retaining a $1.17\times$ speedup on perfect-hint $25\times25$ grids. In contrast, CaDiCaL 3.0.0, whose runtime is overhead-dominated and which always respects the injected branching hints rather than overwriting them, shows no significant speedup (median $1.02\times$, n.s.) and even a small significant mean slowdown ($0.90\times$) on $9\times9$. These results delineate the regimes in which neural guidance translates into practical speedups.

13:00 JSTLLM/生成AIエージェント

誰も見ていないときに LLM エージェントが言うこと: マルチエージェントの議論における社会構造と潜在的な目的の出現

LLM エージェントは、役割、対象者、および関係性のコンテキストが、発言するのに有利な、またはコストがかかるという社会的に構造化された環境でますます行動するようになります。私たちは、そのような社会構造が、プロンプトに明確な目的がない場合に、エージェントが公に表現する内容を、同じ条件下で誘発されたオフレコ (OTR) チャネルと比較して変化させるかどうかを研究します。デュアルチャネルディベートフレームワークを導入します。このフレームワークでは、エージェントが公開発言を生成し、OTR 応答とともに共有履歴に入ります。OTR 応答は記録されますが、他の参加者には表示されません。 10 のモデル、3 つのシナリオ、および各シナリオ内の 5 つのバリエーションにわたって、調整を誘発する設定により、ターゲットのエージェントで体系的なパブリック OTR の相違が生成され、その意思決定の相違が $\sim$3% ベースラインから約 40% まで上昇します。この効果は、スタンス、意味論的類似性、自然言語推論、アンケート回答の 4 つの集計分析にわたって一貫しています。場合によっては、OTR の回答では、キャリアのリスクやスポンサーシップの義務など、公共の便宜が人間関係のプレッシャーによるものであると明示的に考えられています。この調査結果は、エージェントの評価が明示的な目標を超えて拡張され、新たな目標を検出する必要があることを示唆しています。我々は、この評価を運用するためのデュアルチャネル評価フレームワークと補完的な行動尺度を提示します。

原文 (English)

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly to say. We study whether such social structure, without any explicit objective in the prompt, changes what an agent expresses publicly relative to an off-the-record (OTR) channel elicited under the same condition. We introduce a dual-channel debate framework in which agents produce public utterances that enter the shared history alongside OTR responses that are recorded but never shown to the other participant. Across 10 models, 3 scenarios, and 5 variations within each scenario, alignment-inducing settings produce systematic public-OTR divergence in the targeted agent, with its decision divergence rising from a $\sim$3% baseline to roughly 40%. The effect is consistent across four aggregate analyses: stance, semantic similarity, natural language inference, and survey responses. In some cases, the OTR response explicitly attributes public accommodation to relational pressures, such as career risk or sponsorship obligation. The findings suggest that agent evaluation should extend beyond explicit goals and detect emergent objectives. We present a dual-channel evaluation framework and complementary behavioral measures that operationalize this assessment.

13:00 JSTLLM/生成AI

ReContext: ロングコンテキスト推論のための LLM ハーネスとしての再帰的証拠再生

長いコンテキストを理解して推論することは、現実的なアプリケーションに大規模言語モデル (LLM) を展開するための重要な要件となっています。最近の LLM は、ますます長いコンテキスト ウィンドウをサポートしていますが、多くの場合、入力にすでに存在する関連する証拠を使用できず、コンテキスト アクセスと効果的なコンテキスト利用との間にギャップがあることが明らかになりました。この研究では、ロングコンテキスト推論を改善するためのトレーニング不要の推論方法である、ロングコンテキスト推論のための LLM Harness (RECONTEXT) としての再帰的証拠再生を提案します。 RECONTEXT は、モデル内部の関連性シグナルを使用してクエリ条件付き証拠プールを構築し、元のコンテキストを完全に保持しながら、最終生成の前にそれを再生します。この再帰的な選択プロセスにより、トレーニング、外部記憶、またはコンテキストの刈り込みを行わずに、証拠の整理と回答の生成が分離されます。また、連想記憶に基づいた理論的分析も提供します。これは、コンテキストを記憶の保存として、質問を検索の手がかりとして、注意を手がかりと追跡の関連付けとして、そして再生を追跡の再活性化として特徴づけます。コンテキスト長が 128K の 8 つのロングコンテキスト データセットでの実験では、RECONTEXT が Qwen3-4B、Qwen3-8B、および Llama3-8B 全体で証拠の利用率を一貫して向上させ、3 つのバックボーンすべてで最高の平均ランクを達成していることが示されています。コードは https://github.com/Yanjun-Zhao/ReContext で入手できます。

原文 (English)

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization. In this work, we propose Recursive Evidence Replay as LLM Harness for Long-Context Reasoning (RECONTEXT), a training-free inference method for improving long-context reasoning. RECONTEXT uses model-internal relevance signals to construct a query-conditioned evidence pool and replays it before final generation while preserving the full original context. This recursive selection process separates evidence organization from answer generation without training, external memory, or context pruning. We also provide a theoretical analysis based on associative memory, which characterizes the context as a memory store, the question as a retrieval cue, attention as cue-trace association, and replay as trace reactivation. Experiments on eight long-context datasets with 128K context length show that RECONTEXT consistently improves evidence utilization across Qwen3-4B, Qwen3-8B, and Llama3-8B, achieving the best average rank on all three backbones. Code is available at https://github.com/Yanjun-Zhao/ReContext.

13:00 JSTLLM/生成AIハードウェア/半導体

LLM のオンライン安全監視

アライメントトレーニングにもかかわらず、LLM は依然として展開時に安全でない出力を生成する傾向があります。したがって、出力をオンラインで監視し、安全性が確保できなくなった場合に警報を発することが重要です。私たちは、外部モデルからの検証信号を閾値処理によってアラームの決定に変換する単純なリアルタイム モニターを研究します。閾値はリスク コントロールによって調整されます。数学的推論とレッドチームデータセットの実験では、このシンプルなデザインが、逐次仮説テストに基づくより高度なモニターと競合できることを示しました。

原文 (English)

Online Safety Monitoring for LLMs

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.

13:00 JSTエージェントClaudeGPT / ChatGPTGemini

永続状態 AI 制御における分散攻撃

AI コーディング エージェントがより自律的になるにつれて、セッション間でコードベースが永続化され、コードを反復的に配布することが増えています。この永続性により、新たな攻撃対象領域が作成されます。調整されていないエージェントやプロンプト挿入されたエージェントは、プル リクエスト (PR) 全体に攻撃を分散し、最適な自然カバーで PR のペイロードのタイミングを計ることができます。結果として得られるダイナミクスを研究するために、AI 制御の設定である Iterative VibeCoding を紹介します。これは、有能だが信頼できない可能性がある AI を安全に導入する研究です。 Iterative VibeCoding では、コーディング エージェントは、秘密のサイド タスクを追求しながら、永続的なコードベース内の一連の PR に基づいてソフトウェアを構築します。私たちのベンチマークには、CLI ツールと Flask Web サービスという 2 つのタスク ファミリが含まれており、合計 20 のタスク バリエーションが含まれています。攻撃エージェントとして Claude Sonnet 4.5 を、モニターとして GPT-4o を使用します。サイド タスクを PR 全体に分散させる段階的な攻撃と、単一の PR に集中する非段階的な攻撃を比較します。両方に対して堅牢な単一のモニターは存在しません。どの戦略が最もよく回避する (モニターを回避する際に成功する) かはモニターの種類によって異なるため、防御側は 1 つのモニターで段階的攻撃と非段階的攻撃の両方を阻止することはできません。高い回避率 (>= 65%) はモデル攻撃エージェント バックエンド (Sonnet 4.5、Gemini 3.1 Pro、Kimi K2.5) 全体で一般化されており、これは単一モデルの機能ではなく永続状態の攻撃対象領域の特性であることが確認されています。回避率も最先端のモニター モデル全体で高いままであり、より高性能なモデルでは段階的回避と非段階的回避の差が拡大します。 PR 全体にわたる不審な蓄積を追跡するステートフル リンク トラッカー モニターを導入します。どちらのタスク ファミリでも、単に蓄積された履歴を確認するだけの差分モニターよりも、段階的な攻撃を大幅に検出します。この強力なモニターと軌跡モニターを 4 モニター アンサンブルで組み合わせると、段階的な攻撃の回避率が最も弱い標準の差分モニターでの 93% から 47% に減少します。

原文 (English)

Distributed Attacks in Persistent-State AI Control

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute attacks across pull requests (PRs) and time its payload for the PR with the best natural cover. To study the resulting dynamics, we introduce Iterative VibeCoding, a setting for AI control, the study of safely deploying capable but potentially untrusted AI. In Iterative VibeCoding, a coding agent builds software over a sequence of PRs in a persistent codebase while pursuing a covert side task. Our benchmark includes two task families: CLI tools and Flask web services, across 20 total task variations. We use Claude Sonnet 4.5 as the attack agent and GPT-4o as the monitor. We compare gradual attacks, which distribute the side task across PRs, against non-gradual attacks concentrated in a single PR. No single monitor is robust to both: which strategy evades best (success while evading the monitor) depends on the monitor type, so a defender cannot close off both gradual and non-gradual attacks with any one monitor. High evasion (>= 65%) generalizes across model attack agent backends (Sonnet 4.5, Gemini 3.1 Pro, Kimi K2.5), confirming this is a property of the persistent-state attack surface rather than a single model's capability. Evasion also remains high across state-of-the-art monitor models and the gap between gradual and non-gradual evasion widens for more capable models. We introduce a stateful link-tracker monitor that tracks suspicious buildup across PRs. On both task families, it detects gradual attacks substantially better than diff monitors that merely see more accumulated history. Combining this stronger monitor with trajectory monitors in a four-monitor ensemble reduces gradual-attack evasion from 93% under the weakest standard diff monitor to 47%.

13:00 JSTLLM/生成AI

TokenScope: 大規模言語モデルにおけるコード指向タスクのトークンレベルの説明可能性と解釈可能性

大規模言語モデル (LLM) がコード生成中にトークンレベルの決定をどのように行うかを理解することは、研究者と実務者の両方にとって依然として大きな課題です。最近のツールはモデルの内部構造や生成結果についての洞察を提供しますが、多くの場合、デコード時の信号、きめ細かい不確実性の測定、代替の生成パスを探索するための対話型メカニズムが欠けています。我々は、トークンレベルのメトリクス、アテンションパターン、生成中の構造情報を公開する、デコーダベースのLLM用のインタラクティブな解釈および分析ツールであるTokenScopeを紹介します。 TokenScope は、抽象構文ツリーを介した対話型のトークン置換、事実に反する分岐、およびコードを認識した集約をサポートします。 TokenScope は、デコード時の信号と構造プログラム分析を統合することにより、コード生成中の LLM の動作を体系的に調査できるようにします。

原文 (English)

TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models

Understanding how Large Language Models (LLMs) make token-level decisions during code generation remains a major challenge for both researchers and practitioners. While recent tools provide insights into model internals or generation outcomes, they often lack decoding-time signals, fine-grained uncertainty measures, and interactive mechanisms for exploring alternative generation paths. We present TokenScope, an interactive interpretability and analysis tool for decoder-based LLMs that exposes token-level metrics, attention patterns, and structural information during generation. TokenScope supports interactive token replacement, counterfactual branching, and code-aware aggregation via abstract syntax trees. By unifying decoding-time signals with structural program analysis, TokenScope enables systematic investigation of LLM behaviour during code generation.

13:00 JSTLLM/生成AIエージェント

来歴分析を通じて LLM エージェントを不整合から保護する

LLM エージェントが強力なツールにアクセスできるようになるにつれて、エージェントのアクションがユーザーの意図に沿っていることを確認することが重要になります。エージェントが提案したツールの呼び出しがユーザーの意図から逸脱すると、位置ずれと呼ばれる現象が発生し、元に戻すのが困難な有害な結果が生じる可能性があります。既存のランタイム ガードレールは、整合性を推論するための体系的なフレームワークを欠く、裁判官としての LLM パラダイムに依存しており、多くの場合、一貫性のない、または監査が難しい判断を生み出します。来歴分析を動機として、提案されたツール呼び出しがエージェントのコンテキストで追跡可能な証拠によってサポートされているかどうかを判断するものとして不整合の検出を形式化する、来歴ベースの概念フレームワークを提案します。このフレームワークに基づいて、私たちは ProvenanceGuard を提案します。これは、選択したツールが実行される前にエージェントのアクションの 3 種類の不整合を分析し、ユーザーの入力クエリと一致しているとみなされる場合にのみアクションの実行を許可する多段階パイプラインです。私たちは、10 個のバックボーン LLM にわたる Agent-SafetyBench と WorkBench という 2 つの異なるベンチマークで、提案したアプローチを評価しました。 LLM-as-a-judge ベースラインと比較して、ProvenanceGuard は、位置ずれしたトレースのエラー率を Agent-SafetyBench で 42.9% から 1.8%、WorkBench で 32.1% から 17.3% に削減します。その一方で、タスクが成功したトレースに対する介入の負担を 30.5% から 12.8% に削減し、位置合わせされたトレースに対する不必要な介入の統計的に有意な増加を導入しません。これらの結果は、構造化された来歴ベースの推論が、LLM エージェントを不整合から保護するための効果的かつ実用的な基盤を提供することを示しています。

原文 (English)

Safeguarding LLM Agents from Misalignment through Provenance Analysis

As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent's proposed tool invocation deviates from the user's intent -- a phenomenon called misalignment -- it may lead to harmful consequences that are difficult to undo. Existing runtime guardrails rely on an LLM-as-a-judge paradigm that lacks a systematic framework for reasoning about alignment, often producing judgments that are inconsistent or difficult to audit. Motivated by provenance analysis, we propose a provenance-based conceptual framework that formalizes misalignment detection as determining whether a proposed tool call is supported by traceable evidence in the agent's context. Building on this framework, we propose ProvenanceGuard, a multi-stage pipeline that analyzes the agent's action for three types of misalignment before the selected tool is executed and only allows the action to take place when it is considered aligned with the user's input query. We evaluated our proposed approach on two different benchmarks, Agent-SafetyBench and WorkBench, across 10 backbone LLMs. Compared to the LLM-as-a-judge baseline, ProvenanceGuard reduces error rate on misaligned traces from 42.9% to 1.8% on Agent-SafetyBench and from 32.1% to 17.3% on WorkBench, while reducing intervention burden on task-successful traces from 30.5% to 12.8% and introducing no statistically significant increase in unnecessary interventions on aligned traces. These results demonstrate that structured, provenance-based reasoning provides an effective and practical foundation for safeguarding LLM agents from misalignment.

13:00 JSTLLM/生成AI

Kara: スライディング ウィンドウ KV キャッシュ圧縮を介した効率的な推論 LLM サービス

推論言語モデルでは長い思考連鎖 (CoT) が生成されることが多く、これによりデコード段階で大規模な KV キャッシュが蓄積され、デコードの待ち時間が長くなり、スループットが制限されます。これらの問題に対処するために、KV キャッシュ圧縮は、後続のデコードに有用な KV ペアを保持しながら、重要でない KV ペアを選択的に削除することでメモリのオーバーヘッドを削減する有望な技術として浮上しました。それにもかかわらず、既存の KV キャッシュ圧縮方法には 2 つの重要な制限があることがわかりました。1) しきい値でトリガーされる圧縮ポリシーでは、スループットの改善が限定的か、さらにはスループットが低下する可能性があり、シーケンスの特定のブロックから KV ペアが完全に削除される可能性があり、情報損失が悪化する可能性があります。 2) 通常、分離された KV ペアまたは厳格な境界を持つ固定サイズのチャンクのいずれかを保持し、重要な柔軟なサイズのチャンクを任意のトークン位置に保持できません。これらの制限を克服するために、最近生成されたコンテキストのみを操作してデコード時の圧縮を実行するスライディング ウィンドウ KV キャッシュ圧縮方法である Kara を提案します。 Kara は、双方向の注意を活用してウィンドウ内で有益な KV ペアをスコア化し、選択します。重要なセマンティック情報の柔軟な保存を可能にするために、選択した KV ペアのサブセットをチャンクに拡張する Token2Chunk モジュールを設計します。さらに、Kara を PagedAttendance に適応させ、vLLM に基づいて構築された推論フレームワークである KvLLM を開発します。これにより、KV キャッシュ メモリの使用量が削減され、出力スループットが効果的に向上します。広範な実験により、提案された Kara と KvLLM の一貫したパフォーマンスの向上が実証されました。

原文 (English)

Kara: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression

Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high decoding latency and limited throughput. To address these issues, KV cache compression has emerged as a promising technique for reducing memory overhead by selectively removing unimportant KV pairs while preserving useful ones for subsequent decoding. Nevertheless, we identify two key limitations in existing KV cache compression methods: 1) their threshold-triggered compression policy may provide limited throughput improvement or even reduce throughput, and may fully eliminate KV pairs from certain blocks of the sequence, potentially worsening information loss. 2) they typically retain either isolated KV pairs or fixed-size chunks with rigid boundaries, failing to preserve important flexible-sized chunks at arbitrary token positions. To overcome these limitations, we propose Kara, a sliding-window KV cache compression method that performs decoding-time compression by operating only on the recently generated context. Kara leverages bidirectional attention to score and select informative KV pairs in the window. To enable flexible preservation of important semantic information, we design a Token2Chunk module to expand a subset of selected KV pairs into chunks. Furthermore, we adapt Kara to PagedAttention and develop KvLLM, an inference framework built upon vLLM, which reduces KV cache memory usage and effectively improves output throughput. Extensive experiments demonstrate consistent performance improvements of proposed Kara and KvLLM.

13:00 JSTLLM/生成AI

SPARCLE: 対照的な言語埋め込みによる SPeaker を意識した整列表現

音声合成の最近の進歩は、音素表現から直接書記素モデリングに移行しました。音素はテキストと音響間の 1 対多のマッピングに対応しますが、話者固有の音響変化を捕捉できない書記素対音素 (G2P) システムに依存しています。これまでの研究では、書記素ベースのモデルは、大規模では音素ベースのシステムよりも優れたパフォーマンスを発揮しますが、低リソース設定ではそうではありません。この論文では、正確な音響実現によって文字を豊かにする話者認識書記素表現モデルである SPARCLE を提案します。 SPARCLE は、話者のアイデンティティを条件に、書記素を対応する Wav2Vec2 音響表現と一致させるという対照的な目的でトレーニングされています。結果として得られるモデルは、ダウンストリームのテキスト読み上げ (TTS) タスク用の G2P システムの代替として機能します。私たちは、SPARCLE が生成品質を向上させ、標準的な書記素ベースのモデルと比較して、極端に低リソースの設定で単語エラー率を半分に減らすことを実証します。

原文 (English)

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P) systems that fail to capture speaker-specific acoustic variation. Prior work demonstrates that grapheme-based models outperform phoneme-based systems at scale, but not in low-resource settings. In this paper, we propose SPARCLE, a speaker-aware grapheme representation model that enriches characters with their precise acoustic realizations. SPARCLE is trained with a contrastive objective to align graphemes with corresponding Wav2Vec2 acoustic representations while conditioned on speaker identity. The resulting model serves as a replacement to G2P systems for downstream text-to-speech (TTS) tasks. We demonstrate that SPARCLE improves generation quality, reducing word error rates by half in extreme low-resource settings compared to standard grapheme-based models.

13:00 JSTLLM/生成AIGemmaLlamaMistral AIQwen

トークン境界での安全性の破壊: BPE トークン化が LLM 連携に悪用可能なギャップをどのように生み出すか

文字レベルの摂動は、人間が読めるプロンプトを残しているにもかかわらず、最新の LLM の安全調整をバイパスします。私たちは中心的な構造メカニズムを特定し、テストします。BPE トークン化は、安全性が重要な単語をサブ単語部分に断片化し、調査した 3 つのパブリック アラインメント データセットには、意図的に断片化された入力は含まれていません。このメカニズムはチェーンであり、5 つのモデル ファミリ (Qwen-3-4B、Qwen-2.5-7B、Gemma-3-4B、Llama-3.1-8B、Mistral-7B) でエンドツーエンドでテストされました。セーフティトークンの断片化をターゲットとした最適化では、拒否された HarmBench プロンプトの 80 ~ 100% で最初のトークン拒否トリガーが反転され、その反転のうち 48% で真に有害な出力が生成されます (モデルごとに 29 ~ 65%、ギャップ対動作 ROC-AUC 0.66 ~ 0.98、プール 0.84)。アクティベーション パッチは、中断された信号をレイヤーの最後の ${\sim}30\%$ に特定します。アライメントデータスキャンでは、30,000 例の中から断片化されたプロンプトは見つかりませんでした (攻撃関連強度でのポジティブコントロールリコール $\geq 99\%$)。そして、標的突然変異実験では、破壊遺伝子座として安全な単語が分離されます。防御側では、68 セル グリッド (55 のトレーニング済みチェックポイント) は、閉じたプール サイズの交絡を持つ 3 つのファミリーでシードおよびプール安定した ASR 閉包を達成する DPO 構成がないことを示しています。断片化されたプロンプトでトレーニングされた SFT は 3/5 家族で ASR を閉じますが、それは良性プロンプトでも拒否を引き起こすグローバル崩壊を介してのみであり、テストした LoRA-16 レシピの下では欠落分布は必要であるが十分ではないことを示しています。選択的修復と全体的な崩壊を区別するために、一対の診断候補である Conv-Benign を紹介します。すべての ASR クレームは 3 人の審査員によって調整されます (セルのランキングは審査員全体で安定しています。絶対レベル $\pm$18pp。App.~B.13 を参照)。

原文 (English)

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural mechanism: BPE tokenization fragments safety-critical words into sub-word pieces, and the three public alignment datasets we surveyed contain no intentionally fragmented inputs. The mechanism is a chain, tested end-to-end on five model families (Qwen-3-4B, Qwen-2.5-7B, Gemma-3-4B, Llama-3.1-8B, Mistral-7B). An optimization targeting safety-token fragmentation flips the first-token refusal trigger on 80-100% of refused HarmBench prompts, with 48% of those flips producing genuinely harmful outputs (per-model 29-65%; gap-vs-behavior ROC-AUC 0.66-0.98, pooled 0.84). Activation patching localizes the disrupted signal to the last ${\sim}30\%$ of layers; an alignment-data scan finds zero fragmented prompts among 30,000 examples (positive-control recall $\geq 99\%$ at attack-relevant intensities); and targeted-mutation experiments isolate safety words as the disruption locus. On the defense side, a 68-cell grid (55 trained checkpoints) shows that no DPO configuration achieves seed- and pool-stable ASR closure on the three families with closed pool-size confounds. SFT trained on fragmented prompts closes ASR on 3/5 families but only via global collapse that raises refusal on benign prompts as well, indicating the missing distribution is necessary but not sufficient under the LoRA-16 recipe we tested. To distinguish selective repair from global collapse, we introduce Conv-Benign, a candidate paired diagnostic. All ASR claims are 3-judge-calibrated (cell rankings stable across judges; absolute levels $\pm$18pp; see App.~B.13).

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文ClaudeGPT / ChatGPTGemini

プロンプト フレーミングは LLM エラー検出のカウントベースの評価を歪める: 数値アンカーからの証拠

カウントベースの F1 は、LLM エラー検出品質の代用として広く使用されていますが、この論文では、スパンの局所化における対応する改善、つまり F1 インフレーションと呼ばれるギャップがなければ、F1 が劇的に上昇する可能性があることを示しています。この論文では、プロンプト誘発カウント歪みに対する制御されたストレス テスト プロトコルである ErrorBench を紹介します。 ErrorBench は、143 の CoNLL-2014 パッセージからの 4,290 の応答を対象に、5 つのプロンプト条件下で 6 つの最新の LLM を評価します。 CoNLL-2014 M2 スタイルのスコアリングでは、アンカーされたプロンプトは F1 インフレの最大 0.79 ポイントを生成し、厳密なマッチングでは最大 0.96 ポイントを生成します。公式 ERRANT 3.0.0 パイプラインとマルチリファレンス スコアリングを使用した 100 パッセージのレプリケーションによりパターンが再現されます。6 つのモデルの平均では、ブラインドからアンカーへのプロンプト シフトにより Count-F1 が +0.21 上昇する一方、マルチリファレンス ERRANT F0.5 は +0.04 しか上昇しません。この研究では、このストレステストプロトコルの下で、命令に高度に準拠した GPT/Claude システムではより大きなカウント応答が、Gemini ファミリーではより小さな応答が見出されました。この調査結果は、LLM 校正と文書レビュー評価では、事前に入力されたエラー数を回避し、カウントベースのメトリクスとともにスパン認識メトリクスを報告する必要があることを示唆しています。

原文 (English)

Prompt Framing Distorts Count-Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench, a controlled stress-test protocol for prompt-induced count distortion. ErrorBench evaluates six contemporary LLMs under five prompt conditions over 4,290 responses from 143 CoNLL-2014 passages. Under CoNLL-2014 M2-style scoring, anchored prompts produce up to 0.79 points of F1 Inflation, and up to 0.96 under strict matching. A 100-passage replication using the official ERRANT 3.0.0 pipeline and multi-reference scoring reproduces the pattern: averaged over six models, the Blind-to-Anchored prompt shift raises Count-F1 by +0.21 while raising multi-reference ERRANT F0.5 by only +0.04. The study finds larger count responses in highly instruction-compliant GPT/Claude systems and smaller responses in the Gemini family under this stress-test protocol. The findings suggest that LLM proofreading and document-review evaluations should avoid pre-populated error counts and should report span-aware metrics alongside count-based metrics.

13:00 JSTLLM/生成AI

テキストを多重グラフにマッピング: ウォークガイド付きグラフ プルーニングとしてプロンプト圧縮

既存のプロンプト圧縮方法はテキストをフラットなトークンシーケンスとして扱い、重要な情報の分散された性質を捉えることができません。重要な情報は多くの場合複数の場所に分散しており、ローカルな構文上の依存関係とグローバルな意味上の関係の両方を通じて接続されています。このような関係構造は当然グラフとして表現され、トークンや文がノードとなり、それらの依存関係がエッジとなります。この目的を達成するために、我々は、細粒度のアテンションベースの依存関係と粗粒度の意味論的関係を共同でモデル化する多重グラフ上の冗長性を意識したグラフ枝刈りとしてプロンプト圧縮を定式化する RAGP を提案します。この異種構造 (密なローカル サブグラフと疎なグローバル接続) 内の非冗長ノードを効率的に識別するために、ヘビーテール ステップ分布によりローカルの利用とグローバルな探索のバランスが自然に保たれる Levy ウォークを採用します。 LongBench での実験では、RAGP が 4 倍の圧縮率で平均スコア 49.3 を達成し、3 倍の圧縮率で 48.8 を達成する LongLLMLingua などの既存の LLM ベースの圧縮方法を上回るパフォーマンスを示しています。さらに、RAGP は、複数のタスクにおいて最先端のビジョンベースのテキスト圧縮パラダイムも超えています。コードは https://anonymous.4open.science/r/RAGP-B0CB で入手できます。

原文 (English)

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning

Existing prompt compression methods treat text as flat token sequences, failing to capture the distributed nature of important information, which is often spread across multiple locations and connected through both local syntactic dependencies and global semantic relations. Such relational structure is naturally represented as a graph, where tokens or sentences become nodes and their dependencies become edges. To this end, we propose RAGP, which formulates prompt compression as Redundancy-Aware Graph Pruning on a multiplex graph that jointly models fine-grained attention-based dependencies and coarse-grained semantic relations. To efficiently identify non-redundant nodes in this heterogeneous structure (dense local subgraphs and sparse global connections), we employ Levy walks whose heavy-tailed step distribution naturally balances local exploitation with global exploration. Experiments on LongBench show that RAGP achieves an average score of 49.3 under a 4x compression ratio, outperforming existing LLM-based compression methods, such as LongLLMLingua, which attains 48.8 at a 3x compression ratio. Besides, RAGP also surpasses state-of-the-art vision-based text compression paradigms on multiple tasks. The code is available at https://anonymous.4open.science/r/RAGP-B0CB.

13:00 JSTLLM/生成AI

専門家: クエリごとのセマンティックおよびキーストローク動作の手がかりを介して、ユーザーのドメイン専門知識に対する LLM 応答をパーソナライズする

大規模言語モデル (LLM) はエンド ユーザーによって使用されることが増えていますが、静的プロファイルまたはテキストのみの信号に依存する既存のパーソナライゼーション方法では、クエリ固有の専門知識の変動を捉えることができません。意味論的手がかりと行動的手がかりを組み合わせることで、LLM 応答をユーザーのクエリ ドメインの専門知識に適応させる、クエリごとのパーソナライゼーション フレームワークである ExPerT を紹介します。 ExPerT は 2 つの重要なコンポーネントで構成されます。(i) コンテキスト内の LLM プロンプトを介してクエリ テキストとキーストロークのダイナミクスを共同で解釈する意味論的行動専門知識推論モジュール、および (ii) 詳細レベル、用語、概念的な複雑さのレベルを適応させる専門知識に基づく条件付き応答生成。参加者 40 名、クエリ 1,270 件を対象としたユーザー調査では、ExPerT が最も強力なベースライン (MAE = 0.398 対 1.162) と比較して専門知識推論エラーを 65.7% 削減し、5 点リッカート スケールで応答満足度が 17.52% (3.71 から 4.36) 改善したことが実証されました。

原文 (English)

ExPerT: Personalizing LLM Responses to Users' Domain Expertise via Query-Wise Semantic and Keystroke Behavioral Cues

Large language models (LLMs) are increasingly used by end users, yet existing personalization methods relying on static profiles or text-only signals fail to capture query-specific expertise variation. We present ExPerT, a query-wise personalization framework that adapts LLM responses to users' query domain expertise by combining semantic and behavioral cues. ExPerT consists of two key components: (i) a semantic-behavioral expertise inference module that jointly interprets query text and keystroke dynamics via in-context LLM prompting, and (ii) an expertise-conditioned response generation that adapts the level of detail, terminology, and conceptual complexity. Our user study with 40 participants and 1270 queries demonstrated that ExPerT reduced expertise inference error by 65.7% compared to the strongest baseline (MAE = 0.398 vs. 1.162) and improved response satisfaction by 17.52% (from 3.71 to 4.36) on a 5-point Likert scale.

13:00 JSTLLM/生成AI研究/論文

Office Comprehension ベンチマーク

Office Comprehension Bench (OCB) を紹介します。これは、ネイティブ ファイル形式 (.docx、.xlsx、.pptx) およびそのバリアントに対する Word、Excel、PowerPoint の LLM システムの理解力を共同で評価する初の公開ベンチマークです。 OCB は 2 つのトラックで構成されます。 File Fidelity Q&A は、オフィスの成果物 (表、グラフ、埋め込み画像、数式、ヘッダー、講演者ノート、名前付き範囲などのアプリ固有の要素) の構造的および視覚的認識をテストします。ドメイン Q&A では、12 の専門ドメインにわたる実際の業界文書に基づいた専門家レベルの推論をテストします。クエリでは、文書全体にわたる複数ステップの分析と統合が必要です。各参照回答はアトミックなバイナリ採点可能なクレームに分解され、LLM 審査員のアンサンブルが各クレームに対する回答を独立して採点します。デフォルトの推論モードで最も強力なフロンティア システムであっても、ドメイン Q&A では約 59.3% に達するだけであり、層内の思考の深さを増やしてもパフォーマンスは大幅に変化しませんが、より高い製品層に移行するとわずかな向上が得られます。データセット、評価ツール、ジャッジプロンプト、公開リーダーボードをリリースします。

原文 (English)

Office Comprehension Benchmark

We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity Q&A tests structural and visual perception of office artifacts - tables, charts, embedded images, formulas, and app-specific elements such as headers, speaker notes, and named ranges. Domain Q&A tests expert-level reasoning grounded in real-world industry documents across 12 professional domains, with queries requiring multi-step analysis and synthesis across documents. Each reference answer is decomposed into atomic, binary-gradable claims, and an ensemble of LLM judges scores responses against each claim independently. Even the strongest frontier system in its default reasoning mode reaches only about 59.3% on Domain Q&A increasing thinking depth within a tier does not move performance materially, while moving to a higher product tier yields modest gains. We release the dataset, evaluation tooling, judge prompt, and a public leaderboard.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

数学試験採点のティーチング・アシスタントとしての LLM: 信頼性と実用的な使いやすさ

オープンエンドの数学試験は、推論、証明の構築、アルゴリズム的思考、中間ステップのコミュニケーションを評価するため、価値があります。また、教師は生徒の誤解を正すのに役立つフィードバックを提供しながら、部分単位のルーブリックを一貫して適用する必要があるため、大規模な採点も困難です。このペーパーでは、学部の離散数学試験の採点アシスタントとして、6 つの現代的なラージ言語モデル (LLM) 構成、Gemini 3.1 Pro Extended、Gemini 3.5 Flash、ChatGPT 5.5 Pro Extended、ChatGPT 5.5 Thinking、Claude Pro Opus 4.7、および Claude Sonnet 4.6 を評価します。この研究では 2 つの評価ポリシーを比較しています。 BASELINE ポリシーでは、明示的な証拠と完全な正当性を強調する、より厳密なルーブリックに従うプロンプトが使用されます。 LIBERAL ポリシーは、ベースライン条件が時々厳しい減点を適用し、有効な部分推論を認識できないことが予備的な採点で示された後に追加されました。人間による採点との一致は、平均絶対誤差、二乗平均平方根誤差、正規化二乗平均平方根誤差、ピアソン相関、および正確な一致を使用して、質問レベルと試験合計レベルの両方で測定されます。結果は、リベラルな部分単位のプロンプトにより、評価されたすべてのモデルファミリーの質問レベルの平均誤差が減少することを示しています。 ChatGPT 5.5 Thinking (LIBERAL) の平均質問レベル MAE (1.87) と RMSE (2.53) が最も低く、Gemini 3.1 Pro Extended (LIBERAL) の合計スコア MAE (8.00) と RMSE (10.66) が最も低いです。ただし、最も強力な合計スコアのピアソン相関は、Gemini 3.1 Pro Extended (BASELINE) の 0.58 で発生し、ポイントの調整とランクの維持が依然として別個の目標であることを示しています。実際のユーザビリティに関する観察結果も報告します。

原文 (English)

LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability

Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate steps. They are also difficult to grade at scale because instructors must apply partial-credit rubrics consistently while giving feedback that helps students repair misconceptions. This paper evaluates six contemporary large language model (LLM) configurations, Gemini 3.1 Pro Extended, Gemini 3.5 Flash, ChatGPT 5.5 Pro Extended, ChatGPT 5.5 Thinking, Claude Pro Opus 4.7, and Claude Sonnet 4.6, as grading assistants for an undergraduate discrete mathematics examination. The study compares two grading policies. The BASELINE policy uses a stricter rubric-following prompt that emphasizes explicit evidence and complete justification. The LIBERAL policy was added after preliminary grading showed that the baseline condition sometimes applied harsh point deductions and failed to recognize valid partial reasoning. Agreement with human grading is measured at both the question and exam-total levels using mean absolute error, root mean squared error, normalized root mean squared error, Pearson correlation, and exact agreement. The results show that liberal partial-credit prompting reduces average question-level error for every evaluated model family. ChatGPT 5.5 Thinking (LIBERAL) has the lowest average question-level MAE (1.87) and RMSE (2.53), while Gemini 3.1 Pro Extended (LIBERAL) has the lowest total-score MAE (8.00) and RMSE (10.66). However, the strongest total-score Pearson correlation occurs under Gemini 3.1 Pro Extended (BASELINE) at 0.58, showing that point calibration and rank preservation remain distinct goals. We also report practical usability observations.

13:00 JSTエージェントビジネス/資金調達

大規模言語モデルの使用のための監査フレームワークの実践: 集団経験主義、疑似合理的認知、AI 生成コンテンツのガバナンス

大規模な言語モデルは、知識の獲得、コード生成、学術論文の執筆、エージェントベースの自動化にますます使用されています。このような設定では、ユーザーは十分な専門知識の実践なしに、高度に構造化された回答、計画、判断を得る可能性があります。このペーパーでは、LLM の使用と AI によって生成されたコンテンツ ガバナンスのための実践監査フレームワークを提案します。ここでは、LLM が大規模な人間の経験を経験的かつ合理的に見える出力にどのように圧縮および再編成するかを説明するために集団経験主義を導入し、ユーザーが AI によって生成された構造化された表現を自分自身の合理的な理解とどのように誤って認識するかを説明するために疑似合理的認知を導入しています。この論文では、AIの主観錯覚、入力資料の主観構造、AI-AI会話のテンプレートループ、AIGC検出における統計的誤判断、生成されたコンテンツが将来のコンテキスト、長期記憶、検索空間、またはエージェントスキルシステムに入るときのメモリ汚染を分析します。これらのリスクを軽減するために、この文書では、要件定義、問題境界の特定、証拠ソースの監査、実際的な検証、逆質問、ロギング、バージョン管理、ロールバック、および新たな認識に基づいた監査プロセスを提案しています。このフレームワークは AI の生産性を否定するものではありません。 LLM の成果は、検証可能、再現可能、介入可能な実践プロセスに戻されるべきだと主張しています。この論文は、LLM インタラクション、AI 生成コンテンツ ガバナンス、長期記憶システム、および人間と AI のインタラクションにおける認知リスクに関する概念的で監査可能なフレームワークを提供します。

原文 (English)

A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content

Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation. In these settings, users may obtain highly structured answers, plans, and judgments without sufficient domain practice. This paper proposes a practice auditing framework for LLM use and AI-generated content governance. It introduces collective empiricism to describe how LLMs compress and reorganize large-scale human experience into outputs that appear empirical and rational, and pseudo-rational cognition to describe how users may mistake AI-generated structured expression for their own rational understanding. The paper analyzes AI subjectivity illusion, subjectivity structures in input materials, template loops in AI-AI conversations, statistical misjudgment in AIGC detection, and memory pollution when generated content enters future contexts, long-term memory, retrieval spaces, or agent skill systems. To reduce these risks, the paper proposes an auditing process based on requirement definition, problem-boundary identification, evidence-source auditing, practical validation, reverse questioning, logging, version management, rollback, and renewed cognition. The framework does not reject AI productivity; it argues that LLM outputs should be returned to verifiable, reproducible, and intervenable processes of practice. The paper provides a conceptual and auditable framework for cognitive risks in LLM interaction, AI-generated content governance, long-term memory systems, and human-AI interaction.

13:00 JSTLLM/生成AI

社会技術的調整の空間の構造化

社会技術的な調整は、AI の動作の社会的望ましさに関係するため、単に技術的なものではなく本質的に規範的なものです。 NLP 研究はその技術的側面にますます取り組んでいますが、多くの場合、そのような「社会的望ましさ」が何を伴うのかについては詳細が不明瞭なままになっています。私たちは、これは根本的なギャップを反映していると主張します。それは、社会的に望ましい AI の行動を社会技術的調整によってどのように定義、正当化、評価するかを指定する体系的な方法が存在しないことです。このギャップに対処するために、私たちは社会技術的な調整を指定するための人間中心のフレームワークを導入します。私たちは、社会行動的望ましさの社会科学的説明を利用して、行動的望ましさの判断の基礎を確立し、このフレームワークを使用して、実際にどのように整合性が特定されるかを分析します。私たちの体系的な文献レビューにより、繰り返し発生するパターンが特定されます。望ましさの判断の根拠となる規範的な概念は、多くの場合、特定されていないか、(望ましい)システム動作の調整ターゲットと混同されており、ターゲット集団は十分に定義されておらず、設計の選択が理論的に正当化されることはほとんどありません。これらの発見は、概念的な特異性の欠如が累積的な進歩を制限していることを示しています。したがって、私たちは社会科学的枠組みを調整設計の選択肢に結びつける推奨事項を提供し、社会技術的調整に対するより概念的に正確なアプローチをサポートします。

原文 (English)

Structuring the Space of Sociotechnical Alignment

Sociotechnical alignment concerns the social desirability of AI behavior and is thus inherently normative, not merely technical. While NLP research increasingly addresses its technical aspects, it often leaves underspecified what such "social desirability" entails. We argue that this reflects a fundamental gap: the absence of a systematic way to specify how sociotechnical alignment defines, justifies, and evaluates socially desirable AI behavior. To address this gap, we introduce a human-centered framework for specifying sociotechnical alignment. We draw on social-scientific accounts of sociobehavioral desirability to ground the basis for behavioral desirability judgments and use this framework to analyze how alignment is specified in practice. Our systematic literature review identifies recurring patterns: normative concepts grounding desirability judgments are often unspecified or conflated with alignment targets for (desired) system behavior, target populations are underdefined, and design choices are rarely theoretically justified. These findings point to a lack of conceptual specificity that limits cumulative progress. We therefore offer recommendations that link social-scientific frameworks to alignment design choices, supporting more conceptually precise approaches to sociotechnical alignment.

13:00 JSTエージェント

拡張可能な監視のための協力的な意見の相違解決

AI エージェントが反対の立場を主張する討論は、拡張可能な監視への重要なアプローチとして浮上しています。しかし、議論は根本的な緊張に直面しています。モデルは裁判官を説得するように動機付けられており、認識論的な誠実さと必ずしも一致するとは限りません。この研究では、私たちは、相互作用のメカニズムを敵対的な議論から協力的な真実の探求へと再構成する、代替パラダイムである意見の不一致の解決を提案します。人間の調停と紛争解決の原則に基づいて、調停者が紛争当事者間での判断を下すのではなく、対話を促進して合意に達するのを支援し、これらの戦略を AI の監視に適応させる自動パイプラインを設計します。モデルが固定された立場を主張する標準的な討論とは異なり、私たちのパイプラインは、モデルが協力して意見の相違点を特定し、矛盾する主張の証拠を調査し、合意に向けて収束するか、意見の相違の特定の「核心」を分離するように指示します。不一致の解決は、非専門家モデルが真実を特定するのに一貫して役立ち、標準的な議論の 49.2% と比較して 62.1% の判断精度を達成していることがわかりました。私たちの結果は、敵対的な説得から協力的な真実の探求まで、拡張可能な監視プロトコルを再考するための心強い経験的証拠を提供します。

原文 (English)

Collaborative Disagreement Resolution for Scalable Oversight

Debate, where AI agents argue opposing positions, has emerged as a key approach to scalable oversight. However, debate faces a fundamental tension: models are incentivized to be persuasive to the judge, which may not always align with epistemic honesty. In this work, we propose an alternative paradigm: disagreement resolution, which reframes the interaction mechanism from adversarial debate to collaborative truth seeking. Drawing on principles from human mediation and conflict resolution, where mediators facilitate dialogue to help disputing parties reach consensus rather than adjudicating between them, we design an automated pipeline that adapts these strategies to AI oversight. Unlike standard debate where models argue for fixed positions, our pipeline directs models to collaboratively identify points of disagreement, examine the evidence for conflicting claims, and converge toward consensus or isolate the specific ''crux'' of their disagreement. We find that Disagreement Resolution consistently helps non-expert models identify the truth, achieving 62.1% judging accuracy compared to 49.2% for standard debate. Our results provide encouraging empirical evidence for rethinking the scalable oversight protocol from adversarial persuasion to collaborative truth-seeking.

13:00 JST研究/論文

インドの皮膚科医は臨床診療とワークフロー管理に人工知能をどのように活用しているか: アトピー性皮膚炎に特に焦点を当てた全国調査

背景: 皮膚科 AI は主に画像ベースの診断に焦点を当ててきましたが、慢性疾患のワークフローはあまり注目されていません。私たちはインドの皮膚科医を対象に調査を行い、アトピー性皮膚炎(AD)に焦点を当てた日常的な臨床課題をマッピングし、現在の AI の使用状況を評価しました。方法:湿疹研究協会の委託による全国横断調査には、インドの現役皮膚科医377人が参加した。この調査では、臨床上の課題、AD ワークフローの障壁、AI の使用、導入の障壁、倫理的懸念が評価されました。分析には、記述統計、カイ二乗検定、誤検出率補正、および多変数ロジスティック回帰が使用されました。結果: 患者のアドヒアランス (61.3%) および困難または難治性症例における治療計画 (57.0%) は、診断の不確実性 (48.0%) よりも多く報告されました。アルツハイマー病ケアでは、重症度スコアリングが 47.7% によって課題として報告されており、測定されたワークフロー領域の中で満足度が最も低かった。現在の AI の使用は 49.9% によって報告されており、そのほとんどは特殊な画像分析ではなく、文献合成、文書化、学術的タスクのための一般的な大規模言語モデルに関係しています。障壁は経験によって異なりました。20 年以上の診療歴を持つ皮膚科医はトレーニング不足を挙げることが多く、5 年以下の皮膚科医は AI ツールを試した後の臨床的有用性の欠如を挙げることが多かったです。 AI ユーザーは非ユーザーよりも患者の自己誤診や不安について懸念を報告する可能性が高く、これは経験や学歴による調整後も依然として顕著でした。結論: 回答者は、臨床ニーズは慢性疾患管理とアルツハイマー病ワークフローのサポートに集中する一方で、主に認知タスクと管理タスクに汎用 AI を使用していると報告しました。臨床医が監視するワークフロー ツールは、スタンドアロンの診断アプリケーションよりも役立つ場合があります。

原文 (English)

How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis

Background: Dermatology AI has mainly focused on image-based diagnosis, while chronic disease workflows have received less attention. We surveyed Indian dermatologists to map routine clinical challenges, with a focus on atopic dermatitis (AD), and assess current AI use. Methods: A nationwide cross-sectional survey commissioned by the Society for Eczema Studies included 377 practicing Indian dermatologists. The survey assessed clinical challenges, AD workflow barriers, AI use, adoption barriers, and ethical concerns. Analyses used descriptive statistics, chi-square tests, false discovery rate correction, and multivariable logistic regression. Results: Patient adherence (61.3%) and treatment planning in difficult or refractory cases (57.0%) were reported more often than diagnostic uncertainty (48.0%). In AD care, severity scoring was reported as a challenge by 47.7% and had the lowest satisfaction among measured workflow areas. Current AI use was reported by 49.9%, most often involving general large language models for literature synthesis, documentation, and academic tasks rather than specialized image analysis. Barriers differed by experience: dermatologists with more than 20 years of practice more often cited lack of training, while those with 5 years or less more often cited lack of clinical utility after trying AI tools. AI users were more likely than non-users to report concern about patient self-misdiagnosis and anxiety, which remained significant after adjustment for experience and academic affiliation. Conclusion: Respondents reported using general-purpose AI mainly for cognitive and administrative tasks, while their clinical needs centered on chronic disease management and AD workflow support. Clinician-supervised workflow tools may be more useful than standalone diagnostic applications.

13:00 JST研究/論文

Beyond Detection: Redesigning Assessment and Governande of Generative AI at the Universidad Polit\'ecnica de Madrid (UPM)

Universities have responded to generative artificial intelligence (GenAI) in noticeably different ways, both internationally and within Spa…

13:00 JST研究/論文

AI Assistance for Human Review of Default Judgments

Overwhelmed courts in the United States review millions of default judgments each year. Unfortunately, such manual reviews are time-consumi…

13:00 JST研究/論文

Artificial Intelligence-Enabled Accounting Information Systems and Fraud Detection in Nigeria's Financial Services Sector: The Moderating Role of Natural Language Processing

The rapid digitalisation of financial systems has improved operational efficiency and financial inclusion while simultaneously increasing e…

13:00 JST研究/論文NVIDIA

The Rising Unsustainability of AI Graphics Cards Production

The rapid advancement of Artificial Intelligence (AI) has been accompanied by significant increases in computational and environmental cost…

13:00 JST画像/動画生成研究/論文

Benchmarking Federated Learning and Knowledge Distillation for Point Cloud Classification

Deploying 3D point cloud analysis in privacy-sensitive, resource-constrained settings faces two barriers: data cannot be centralized, and m…

13:00 JST研究/論文

Domain Knowledge Based Temporal-Spatial Graph Convolution Network for ECG Recognition

In light of strides in Arti cial Intelligence (AI) and its wide spread application, challenges persist in the interpretability of AI models…

13:00 JST研究/論文

Scaling Laws for Grid-Based Approximate Nearest Neighbor Search in High Dimensions

Grid-based approaches to approximate nearest neighbor (ANN) search have been absent from modern scaling analyses. We present a systematic c…

13:00 JSTロボティクス

Adaptive Companionship for Group-Following Robots: Handling Dynamically Changing Group Formations

Accompanying a group of humans is an essential aspect of developing human-like social cognition in robots. However, human groups typically…

13:00 JSTLLM/生成AI画像/動画生成

CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection

Presentation Attack Detection (PAD) serves as a crucial safeguard for face recognition systems against presentation attacks such as printed…

13:00 JST研究/論文

Generative AI and Federated Learning for Intrusion Detection Systems: A Survey

Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physic…

13:00 JSTLLM/生成AI

Black-Box Inference of LLM Architectural Properties with Restrictive API Access

In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown t…

13:00 JST研究/論文

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

Neural Quantum States (NQS) are a remarkably expressive class of variational ans\"atze for quantum many-body wavefunctions, yet little is u…

13:00 JSTLLM/生成AIビジネス/資金調達

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluatio…

13:00 JST画像/動画生成

Multi-modal Rail Crossing Safety Analysis

Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we impro…

13:00 JST研究/論文

AI-enabled gravitational-waves searches for binary neutron stars at optimal sensitivity

Gravitational Waves (GWs) represent the newest window of astronomy, furthering our understanding of compact objects like black holes and ne…

13:00 JST研究/論文

How Should Transformers Encode Numeric Values in Electronic Health Records?

How do we encode numeric values in transformer-based sequence processing, particularly in electronic health record (EHR) data? We systemati…

13:00 JST画像/動画生成

Rethinking Generic Object Tracking Toward Human-Level Perceptual Intelligence

At the heart of human visual perception lies the ability to maintain a continuous and coherent understanding of the external world. By inte…

13:00 JST画像/動画生成

NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis

INTRODUCTION: Accurate MRI-based identification of Alzheimer's disease (AD), mild cognitive impairment (MCI), and related dementias remains…

13:00 JST研究/論文

Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks

$\mathrm{E}(3)$-equivariant networks are promising for 3D atomistic system modeling, yet their scalability is limited by the $O(L^6)$ compl…

13:00 JSTハードウェア/半導体

GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures

GPU training jobs fail often, roughly two in five on large production clusters, yet the operator typically learns of a failure only by reco…

13:00 JSTLLM/生成AIエージェントAnthropicClaudeMicrosoftCopilot

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them,…

13:00 JSTLLM/生成AI画像/動画生成GPT / ChatGPT

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for use…

13:00 JSTエージェント研究/論文

Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance

Engineering management research has produced mature frameworks for software risk: ownership by feature, escalation by severity, and assuran…

13:00 JSTLLM/生成AI研究/論文

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs

We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge re…

13:00 JSTLLM/生成AI

On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain

Mixture-of-Experts (MoE) models offer inference speedups via selective activation but impose substantial memory requirements because the wh…

13:00 JST研究/論文

Token Geometry

Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface be…

13:00 JSTLLM/生成AI

Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting

Large language models (LLMs) are increasingly applied to resume optimization for applicant tracking systems, introducing hallucination fail…

13:00 JST研究/論文

Fully Unsupervised Detection of Physical Contacts on Subsea Cables via State-of-Polarization Monitoring

We present a fully unsupervised Fast-Slow DSVDD detector for continuous State-of-Polarization monitoring on a deployed subsea cable. Traine…

13:00 JSTLLM/生成AI

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL

Reinforcement learning post-training dramatically improves LLM reasoning, but suffers from training instability and diversity collapse. Adv…

13:00 JST研究/論文

Robust and Explainable 3D Mode Shape Recognition Using Region-Aware Graph Neural Networks

Mode shape recognition is a fundamental task in automotive NVH development, yet it remains dependent on manual visual inspection by experie…

13:00 JSTLLM/生成AIエージェント

Multi-Head Recurrent Memory Agents

Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despit…

13:00 JST研究/論文

IntentTune: Using user demand and personalization to resolve "unknown" query intents for e-commerce search

Understanding user intent is fundamental to delivering relevant search results in e-commerce. However, substantial fraction of real-world q…

13:00 JST研究/論文

Evolutionary Feature Engineering for Structured Data

Large language models are increasingly used as open-ended search operators in evolutionary optimization. We introduce Evolutionary Feature…

13:00 JST研究/論文

X-LogSMask: Expand Transformer for Graph-Structured Data

Transformers have become general-purpose architectures, but their all-to-all self-attention is poorly matched to graph data, whose interact…

13:00 JSTLLM/生成AIエージェント

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns requir…

13:00 JSTLLM/生成AI

ADVENT: LLM-Driven Automatic Predicate Invention for ILP

Predicate invention (PI), the creation of new predicates to extend the hypothesis space, remains a critical bottleneck in Inductive Logic P…

13:00 JST画像/動画生成ロボティクス

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training para…

13:00 JST研究/論文

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development.…

13:00 JSTLLM/生成AIエージェント研究/論文

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated i…

13:00 JST研究/論文

Beyond Gradient-Based Attacks: Adversarial Robustness and Explainability Stability in Cybersecurity Classifiers

Adversarial attacks on cybersecurity classifiers pose a dual threat: degrading predictions and destabilising the SHAP-based explanations th…

13:00 JST研究/論文

Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space

Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.~Most ex…

13:00 JST研究/論文

Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attac…

13:00 JST研究/論文

Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander

We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choo…

13:00 JSTエージェント

Full Bayesian Reinforcement Learning via LF-IBIS

Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an…

13:00 JST画像/動画生成研究/論文

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at t…

13:00 JSTエージェント研究/論文

Decentralized Stochastic Subgradient-type Methods with Communication Compression for Nonsmooth Nonconvex Optimization

In this paper, we consider the nonsmooth nonconvex decentralized optimization problem, where inter-agent communication is compressed. We pr…

13:00 JST画像/動画生成

ProCal: Inference-Time Proposal Calibration for Open-Vocabulary Object Detection

Open-vocabulary object detection aims to localize and classify objects beyond the fixed set of categories seen dur ing training. Recent ope…

13:00 JST研究/論文

AI Virtue: What is "Good" Knowledge in the Age of Artificial Intelligence?

In the age of AI, what will be good knowledge? This article, which is accepted and forthcoming in a special issue of Modern Fiction Studies…

13:00 JST研究/論文

Scene-Conditioned PINN-GNN for Multipath RF Maps: Cross-Scene Generation and In-Scene Completion

Radio frequency (RF) maps provide a compact representation of multipath propagation characteristics and are fundamental to channel modeling…

13:00 JST研究/論文

EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning

Mixture-of-Experts (MoE) models scale efficiently but remain costly to adapt due to redundant experts and uniform parameter allocation. Exi…

13:00 JSTエージェントロボティクス

Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as…

13:00 JST研究/論文

Single-Channel EEG-Based Cognitive Load Assessment in Online Learning: A Hybrid Deep Learning Approach

Monitoring cognitive load during online learning could help instructors identify content that learners find difficult, but remote settings…

13:00 JST研究/論文Llama

Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability

Sparse autoencoders (SAEs) decompose internal activations of neural networks into sparse linear combinations of learned features by fitting…

13:00 JSTLLM/生成AIエージェントClaude

Decoupling Code Complexity from Newcomer Participation: A Causal Study of AI Coding Agent Adoption in OSS

Open-source projects depend on a steady inflow of newcomers. A growing concern is that AI coding agents (tools such as Cursor and Claude Co…

13:00 JST画像/動画生成ビジネス/資金調達研究/論文

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vul…

13:00 JST研究/論文

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and spe…

13:00 JST研究/論文

Decomposer: Learning to Decompile Symbolic Music to Programs

Musical performance involves executing a set of high-level musical instructions, yet recovering those instructions from the performance is…

13:00 JSTLLM/生成AI

Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts

Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information…

13:00 JST研究/論文GemmaLlamaQwenDeepSeek

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map

Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: th…

13:00 JSTLLM/生成AI

An Exploratory Study on LLM-Generated Code and Comments in Code Repositories

The use of LLMs in software development has become increasingly widespread on tasks such as code generation and summarization. Reports from…

13:00 JST画像/動画生成

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models

Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and…

13:00 JST研究/論文

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. R…

13:00 JST研究/論文

SABER: A Semantic-Aligned Brain Network Analysis Framework via Multi-scale Hypergraphs

Effective brain disease diagnosis requires the synergy of brain connectivity patterns and high-level semantic knowledge. Existing methods,…

13:00 JST画像/動画生成

Population-Based Multi-Objective Training of Discriminators for Semi-Supervised GANs

Semi-supervised generative adversarial networks (SSL-GANs) can exploit large unlabeled datasets while retaining a classifier in the discrim…

13:00 JST研究/論文

Low-Latency Task-Oriented Image Transmission with Opportunistic Spectrum Access

Communication systems designed for reliable data reconstruction, rather than task-oriented communication, typically rely on separate source…

13:00 JSTLLM/生成AI研究/論文Qwen

TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B

This paper presents TUDUM (T\"urk\c{c}e D\"u\c{s}\"unen \"Uretken Model), a project pipeline for adapting a Qwen-family 27B thinking model…

13:00 JSTLLM/生成AILlama

AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical…

13:00 JSTLLM/生成AIエージェントCopilot

CausalSteward: An Agentic Divide-Conquer-Combine Copilot for Causal Discovery

Learning causal models from high-dimensional data is a significant challenge, particularly in real-world settings where violations of core…

13:00 JSTLLM/生成AI画像/動画生成ロボティクス

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-langu…

13:00 JST研究/論文GPT / ChatGPT

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits

Mechanistic interpretability often relies on component-level interventions to discover how a model produces a behavior. This guides attribu…

13:00 JSTLLM/生成AILlamaMistral AIQwen

Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantical…

13:00 JST画像/動画生成ロボティクス

NeoMap: Training-free Novel-View Synthesis from Single Images and Videos

We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate unde…

13:00 JSTLLM/生成AI

Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization

Large language models (LLMs) are often asked to produce JSON conforming to a fixed schema, powering information extraction, tool calling, a…

13:00 JST画像/動画生成ビジネス/資金調達

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual questi…

13:00 JST研究/論文

A Multi-Branch Hierarchy-Aware Framework for Heterogeneous Audio Classification

This technical report describes our system for Task 1 of the DCASE 2026 Challenge, which aims to classify heterogeneous audio recordings ac…

13:00 JSTLLM/生成AI画像/動画生成

MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a n…

13:00 JST画像/動画生成NVIDIA

Do Newer Lightweight CNNs Perform Better Under Resource Constraints? A Controlled Multigenerational Study of Architecture, Initialization, Training Budget, and Efficiency

Newer lightweight convolutional neural networks are often presented as improving predictive performance and deployment efficiency, but such…

13:00 JST画像/動画生成

Mirror Illusion Art

Mirror Illusion Art is a novel reflection-conditioned 3D illusion where one object yields two target appearances (front and mirror). The ta…

13:00 JSTLLM/生成AIハードウェア/半導体DeepSeek

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates…

13:00 JSTLLM/生成AI

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolate…

13:00 JSTLLM/生成AIGeminiLlamaDeepSeek

SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-ling…

13:00 JST画像/動画生成ビジネス/資金調達

Beyond the Performance Illusion: Structure-Aware Stratified Partitioning and Curriculum Distributionally Robust Optimization for Spatially Correlated Domains

Performance evaluation in AI systems commonly assumes that random dataset splits produce independent and identically distributed (i.i.d.) s…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Prompt Coverage Adequacy

In recent years, it has become increasingly evident that large language models (LLMs) and autonomous agents raise the level of abstraction…

13:00 JST研究/論文

SA-HGNN: Sample-Adaptive Hyperbolic Graph Neural Network for EEG-Based Depression Recognition

Graph Neural Networks (GNNs) have been widely used to capture spatial functional connectivity patterns to improve electroencephalography (E…

13:00 JSTLLM/生成AI

kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail

Large language models (LLMs) are increasingly deployed in domains requiring guardrails to detect unsafe, off-topic, or adversarial prompts.…

13:00 JST研究/論文

Evolutionary Wave Function Collapse

Wave Function Collapse (WFC) is a widely used procedural content generation method that learns local adjacency constraints from example inp…

13:00 JSTLLM/生成AI画像/動画生成

ESC: Emotional Self-Correction for Reliable Vision-Language Models

Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable re…

13:00 JSTロボティクス

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies

Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity…

13:00 JSTLLM/生成AI

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This…

13:00 JSTLLM/生成AIエージェント

Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring

As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is c…

13:00 JST研究/論文

ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uni…

13:00 JST画像/動画生成

Predicting Early Stages Of Alzheimer's Disease And Identifying Key Biomarkers Using Deep Artificial Neural Network And Ensemble Of Machine Learning Methodologies

Alzheimers disease (AD) is a brain disorder that develops slowly and mainly affects memory, thinking, language, and daily activities. It is…

13:00 JST研究/論文

Dynamic Neural Graph Encoding of Inference Processes in Deep Weight Space

The rapid advancements in using neural networks as implicit data representations have attracted significant interest in developing machine…

13:00 JST画像/動画生成

RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation

Deep learning has achieved remarkable performance in medical image segmentation, yet it suffers from critical limitations: mathematical int…

13:00 JST規制/政策

Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond

The society and emerging risk-based regulatory frameworks for AI underscore the need for rigorous risk assessment to ensure safe and reliab…

13:00 JST研究/論文

What Types of Human-AI Teams Exist?

Human-AI teaming has received increasing attention in the literature. However, the range of studies conducted in multiple domains make it d…

13:00 JSTビジネス/資金調達GPT / ChatGPT

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…

13:00 JSTロボティクス

CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation

Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction d…

13:00 JST画像/動画生成

Efficient Waste Sorting for Circular Economy: A Confidence-guided comparison between One-Vs-All and One-Vs-Rest Classification Strategies with Human-in-the-Loop for Automated Waste Sorting

The complexity of waste disposal regulations across European countries poses significant challenges for the residents and hinders the trans…

13:00 JSTLLM/生成AIビジネス/資金調達

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional…

13:00 JSTLLM/生成AI

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures

Most data-mixing methods assume the corpus has already been partitioned into groups, and the choice of those groups determines what a mixer…

13:00 JST画像/動画生成ビジネス/資金調達研究/論文

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation prot…

13:00 JST研究/論文

Generalization in offline RL: The structure is more important than the amount of pessimism

While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with…

13:00 JSTLLM/生成AI

SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios

Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remain…

13:00 JST研究/論文

Self-Gating Attention for Efficient Time Series Forecasting

Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture…

13:00 JSTLLM/生成AIエージェント

SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces

Large Language Model (LLM)-based agents increasingly automate software engineering tasks through reusable skills, natural-language instruct…

13:00 JST画像/動画生成

GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset

Monocular relative pose sensing is a central perception problem in non-cooperative rendezvous and on-orbit servicing. In spacecraft images,…

13:00 JST研究/論文

Stable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory Gates

Quantum Fast-Weight Programmers (QFWPs) store temporal information in dynamically programmed variational-circuit parameters rather than in…

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry

Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation stru…

13:00 JSTLLM/生成AI

World Wide Models: Literary Tools for Cultural AI

LLMs stage a new form of cultural encounter that is massive, automated, and monolingual. Literary disciplines have always negotiated cultur…

13:00 JSTエージェント

Understanding Agent-Based Patching of Compiler Missed Optimizations

Compiler missed optimizations refer to cases in which compilers failed to optimize certain code. It takes many compiler developers' efforts…

13:00 JST画像/動画生成GoogleGemini

VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval

Over 285 million people worldwide live with a visual impairment, for whom everyday tasks such as avoiding obstacles, locating personal belo…

13:00 JST画像/動画生成ロボティクス

ACID: Action Consistency via Inverse Dynamics for Planning with World Models

Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard plann…

13:00 JSTLLM/生成AI

Neuron-Aware Active Few-Shot Learning for LLMs

Active Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and us…

13:00 JSTエージェント

QFedAgent: Quantum-Enhanced Personalized Federated Learning for Multi-Agent Activity Recognition

Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, making it suitable for pr…

13:00 JSTロボティクス

WorldSample: Closed-loop Real-robot RL with World Modelling

Reinforcement learning (RL) can overcome the demonstration-coverage limitation of imitation learning (IL) by allowing robots to improve thr…

13:00 JSTLLM/生成AIエージェント

Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study

Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts…

13:00 JSTLLM/生成AI

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, partic…

13:00 JST画像/動画生成

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter coun…

13:00 JSTロボティクス

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, in…

13:00 JST研究/論文

Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting

Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymar…

13:00 JSTLLM/生成AI研究/論文ClaudeGemini

TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution

Software tests and code evolve together: a code change should be followed by new or updated tests that record the new software behavior. Ye…

13:00 JST画像/動画生成

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning

Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing methods often fail to…

13:00 JSTLLM/生成AI

Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials

Machine learning interatomic potentials (MLIPs) have become a hallmark of AI for scientific simulation. While efforts on new architectures…

13:00 JSTLLM/生成AI

DemoPSD: Disagreement-Modulated Policy Self-Distillation

On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single mo…

13:00 JSTLLM/生成AI画像/動画生成

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies…

13:00 JSTLLM/生成AI

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON,…

13:00 JSTLLM/生成AI

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc…

13:00 JST画像/動画生成

Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition

Perturbation robustness evaluates the vulnerabilities of models, arising from a variety of perturbations, such as data corruptions and adve…

13:00 JSTハードウェア/半導体

Causal Explanations for Image Classifiers

Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to…

13:00 JST研究/論文

MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer

Humans excel at analogical reasoning - applying knowledge from one task to a related one with minimal relearning. In contrast, reinforcemen…

13:00 JST研究/論文

ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion

Multimodal emotion and intent recognition is essential for automated human-computer interaction, It aims to analyze users' speech, text, an…

13:00 JSTLLM/生成AI

Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic…

13:00 JSTエージェント研究/論文

Aria: An Agent For Retrieval and Iterative Auto-Formalization via Dependency Graph

Accurate auto-formalization of theorem statements is essential for advancing automated discovery and verification of research-level mathema…

13:00 JSTエージェント

BuilderBench: The Building Blocks of Intelligent Agents

Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the lim…

13:00 JSTLLM/生成AI画像/動画生成

Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning

Recent medical MLLMs have made significant progress in generating step-by-step textual reasoning chains. However, they still struggle with…

13:00 JSTLLM/生成AI

HAL: Inducing Human-likeness in LLMs with Alignment

Aligning language models to qualitative behavioral traits, such as human-likeness, remains difficult because they are hard to define, measu…

13:00 JST研究/論文

A General Neural Backbone for Mixed-Integer Linear Optimization via Dual Attention

Mixed-integer linear programming (MILP) is a foundational framework for combinatorial optimization across science and engineering, but rema…

13:00 JSTエージェント

Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Agent memory systems must accommodate continuously growing information while supporting efficient, context-aware retrieval for downstream t…

13:00 JSTLLM/生成AI研究/論文

BRIDGE: Predicting Human Task Completion Time From Model Performance

Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task diffi…

13:00 JSTLLM/生成AI研究/論文

PreScience: A Dataset and Benchmark for Scientific Forecasting

Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and ben…

13:00 JSTLLM/生成AI画像/動画生成エージェント

OmniGAIA: Towards Native Omni-Modal AI Agents

Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool u…

13:00 JSTエージェント研究/論文

Learning-based Multi-agent Race Strategies in Formula 1

In Formula 1, race strategies are adapted according to evolving race conditions and competitors' actions. This paper proposes a reinforceme…

13:00 JSTエージェント

Conformal Policy Control

An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause ha…

13:00 JSTLLM/生成AIエージェント

A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development

WebGIS development requires consistency, yet agentic AI often fails due to LLM context constraints, forgetting, stochasticity, instruction…

13:00 JSTエージェント

エージェントティック ツール プロトコルの形式セマンティクス: プロセス計算アプローチ

外部ツールを呼び出すことができる大規模言語モデル エージェントの出現により、エージェント プロトコルの正式な検証が緊急に必要になりました。この分野では、ゼロショット API の一般化のための研究フレームワークであるスキーマガイド ダイアログ (SGD) と、エージェントとツールの統合のための業界標準であるモデル コンテキスト プロトコル (MCP) の 2 つのパラダイムが支配的です。どちらもスキーマ記述を通じて動的なサービス検出を可能にしますが、その正式な関係はまだ解明されていません。これらのパラダイムの概念的収束を確立した以前の研究に基づいて、我々は SGD と MCP の最初のプロセス計算による定式化を提示し、それらが明確に定義されたマッピング ファイの下で構造的に類似していることを証明します。ただし、逆マッピング Phi^{-1} は部分的で損失が多く、MCP の表現力に重大なギャップがあることが明らかになります。双方向分析を通じて、完全な動作の等価性のための必要十分条件として、5 つの原則 (セマンティックな完全性、明示的なアクション境界、障害モードの文書化、漸進的開示互換性、ツール間関係宣言) を特定しました。これらの原則を型システム拡張 MCP+ として形式化し、MCP+ が SGD と同型であることを証明します。私たちの研究は、検証されたエージェント システムの最初の正式な基盤を提供し、証明可能な安全性の特性としてスキーマの品質を確立します。

原文 (English)

Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach

The emergence of large language model agents capable of invoking external tools has created urgent need for formal verification of agent protocols. Two paradigms dominate this space: Schema-Guided Dialogue (SGD), a research framework for zero-shot API generalization, and the Model Context Protocol (MCP), an industry standard for agent-tool integration. While both enable dynamic service discovery through schema descriptions, their formal relationship remains unexplored. We present the first process calculus formalization of SGD and MCP, proving they are structurally bisimilar under a well-defined mapping Phi. We demonstrate that the reverse mapping Phi-1 is partial and lossy, revealing critical gaps in MCP's expressivity. Through bidirectional analysis, we identify four principles - semantic completeness, explicit action boundaries, failure mode documentation, and inter-tool relationship declaration -- as necessary and sufficient conditions for full behavioral equivalence. We formalize these principles as type-system extensions MCP+, proving MCP+ is fully equivalent to SGD. Our work provides the first formal foundation for verified agent systems and establishes schema quality as a provable safety property. Practically, this means that the current MCP specification has expressiveness gaps compared to SGD and would benefit from the proposed extensions.

13:00 JST研究/論文

Working Paper: Towards a Category-theoretic Comparative Framework for Artificial General Intelligence

AGI has become the Holly Grail of AI with the promise of level intelligence and the major Tech companies around the world are investing unp…

13:00 JSTLLM/生成AILlama

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent m…

13:00 JSTLLM/生成AIエージェントロボティクス

From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents

Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive e…

13:00 JST研究/論文

Stabilising Generative Models of Attitude Change

Attitude change - the process by which individuals revise their evaluative stances - has been explained by a set of influential but competi…

13:00 JSTエージェントロボティクス

物理的にネイティブな世界モデル: 生成世界モデリングに関するハミルトニアンの視点

ワールド モデルは最近、身体化されたインテリジェンス、ロボット工学、自動運転、モデルベースの強化学習の中心的なパラダイムとして再浮上しています。しかし、現在の世界モデル研究は、視覚的な未来合成を重視する 2D ビデオ生成モデル、空間再構成を重視する 3D シーン中心モデル、および抽象的な予測表現を重視する JEPA のような潜在モデルという、部分的に分離した 3 つのルートによって支配されることがよくあります。各ルートは重要な進歩を遂げていますが、具体化された意思決定のための、物理的に信頼性が高く、アクション制御可能で、長期的に安定した予測を提供するのに依然として苦労しています。この論文では、世界モデルのボトルネックは、もはや現実的な未来を生成できるかどうかだけではなく、それらの未来が物理的に意味があり、行動に役立つかどうかであると主張します。私たちは、世界モデリングに関する物理的に根拠のある視点として \emph{ハミルトニアン世界モデル} を提案します。重要なアイデアは、観測値を構造化された潜在位相空間にエンコードし、制御、散逸、残差項を含むハミルトニアンにインスピレーションを得たダイナミクスを通じて潜在状態を進化させ、予測された軌道を将来の観測値にデコードし、結果として得られるロールアウトを計画に使用することです。ハミルトニアン構造がどのように解釈可能性、データ効率、長期安定性を向上させることができるかについて議論するとともに、摩擦、接触、非保存力、変形可能な物体を含む現実世界のロボットシーンにおける実際的な課題にも言及します。

原文 (English)

Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling

World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world model research is often dominated by three partially separated routes: 2D video-generative models that emphasize visual future synthesis, 3D scene-centric models that emphasize spatial reconstruction, and JEPA-like latent models that emphasize abstract predictive representations. While each route has made important progress, they still struggle to provide physically reliable, action-controllable, and long-horizon stable predictions for embodied decision making. In this paper, we argue that the bottleneck of world models is no longer only whether they can generate realistic futures, but whether those futures are physically meaningful and useful for action. We propose \emph{Hamiltonian World Models} as a physically grounded perspective on world modeling. The key idea is to encode observations into a structured latent phase space, evolve the latent state through Hamiltonian-inspired dynamics with control, dissipation, and residual terms, decode the predicted trajectory into future observations, and use the resulting rollouts for planning. We discuss how Hamiltonian structure may improve interpretability, data efficiency, and long-horizon stability, while also noting practical challenges in real-world robotic scenes involving friction, contact, non-conservative forces, and deformable objects.

13:00 JSTLLM/生成AI研究/論文

LGMT: LLM の推論の信頼性を評価するための論理に基づいた変形テスト

大規模言語モデル (LLM) は、論理推論ベンチマークで優れたパフォーマンスを達成しますが、その信頼性は依然として不確実です。既存の評価は静的ベンチマークに依存しているため、論理的に同等の変換の下での堅牢性を評価できず、推論能力を過大評価することがよくあります。私たちは、一次論理 (FOL) を活用して LLM 推論を評価する、オラクル不要のフレームワークである LGMT (Logic-Grounded Metamorphic Testing) を提案します。 LGMT は、形式的な論理的等価性から変成関係を導出することで、意味的に不変のテスト ケースを構築し、ケース間の整合性チェックを通じて推論の欠陥を検出します。 6 つの最先端の LLM での実験では、LGMT が従来のリファレンスベースの評価では見逃されていた重大な隠れた欠陥を明らかにすることが示されています。さらに、モデルはシンボルレベルと結論レベルの変動に特に敏感であり、Few-shot CoT などの高度なプロンプトではこれらの問題が部分的にしか軽減されないことがわかりました。これらの結果は、LLM 評価が孤立した正確性を超えて、論理的不変性の下での堅牢性へと移行する必要があることを示唆しています。 LGMT は、推論の失敗を診断するための原則に基づいたスケーラブルなアプローチを提供します。

原文 (English)

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs

Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain. Existing evaluations rely on static benchmarks, which fail to assess robustness under logically equivalent transformations and often overestimate reasoning capability. We propose LGMT (Logic-Grounded Metamorphic Testing), an oracle-free framework that leverages first-order logic (FOL) to evaluate LLM reasoning. By deriving metamorphic relations from formal logical equivalences, LGMT constructs semantically invariant test cases and detects reasoning defects through cross-case consistency checking. Experiments on six state-of-the-art LLMs show that LGMT exposes substantial hidden defects missed by traditional reference-based evaluations. We further find that models are particularly sensitive to symbol-level and conclusion-level variations, and that advanced prompting such as Few-shot CoT only partially mitigates these issues. These results suggest that LLM evaluation should move beyond isolated correctness toward robustness under logical invariance. LGMT provides a principled and scalable approach for diagnosing reasoning failures.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

LLM エージェントの機能を評価するための統一フレームワーク

LLM がエージェントとして導入されることが増えているため、そのエージェント機能の信頼できる評価が不可欠になっています。ただし、報告されるベンチマーク スコアは、多くの場合、モデルの機能と、各ベンチマークに含まれる実装の選択肢を合わせて反映するため、クロスベンチマークの結果を基礎となるモデルの正確な測定値として解釈することが困難になります。この研究では、LLM エージェントの機能を公正に評価するための統一フレームワークを紹介します。統合された構成システムによって駆動されるこのフレームワークは、標準化された命令、ツール、環境の形式に多様なベンチマークを統合し、制御可能なサンドボックス内の固定 ReAct スタイル アーキテクチャを通じてエージェントを実行します。また、フレームワークの効果と環境の効果を個別に分析できるように、揮発性のライブ環境を厳選されたスナップショットに置き換えるオプションのオフライン設定を提供します。これに基づいて、各ベンチマークの元のタスクの成功基準に基づいて評価方法を統一するとともに、リソース消費に関する統一された指標と、意思決定レベルおよび実行レベルの失敗の属性に関する分類を導入します。このフレームワーク内で、シングルエージェント、マルチエージェント、およびセーフティクリティカルなシナリオにわたる 24 のドメインにわたる 7 つの広く使用されているベンチマークを適応させ、15 のモデルで 400,000 のロールアウトと 50 億のトークンにわたる大規模な実証分析を実施します。結果は、足場の選択と環境の変動性がベンチマークの結果を両方向に実質的に変化させ、フレームワークおよび環境によって引き起こされるアーティファクトから本質的な LLM 機能を解きほぐすことをフレームワークが可能にすることを示しています。さらに、安全性が重要なドメインの安全なテストベッドとしての拡張性を実証します。コードとベンチマークは、https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities、https://huggingface.co/AgentFramework/Unified_Farmework で入手できます。

原文 (English)

A Unified Framework for the Evaluation of LLM Agentic Capabilities

As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model capability and the implementation choices each benchmark is packaged with, making cross-benchmark results difficult to interpret as clean measurements of the underlying model. In this work, we present a unified framework for the fair evaluation of LLM agentic capabilities. Driven by a unified configuration system, the framework integrates diverse benchmarks into a standardized instruction-tool-environment format, executes agents through a fixed ReAct-style architecture within a controllable sandbox, and provides an optional offline setting that replaces volatile live environments with curated snapshots, so that framework effects and environment effects can be analyzed separately. Building on this, we unify the evaluation methodology under each benchmark's original task-success criteria, while introducing unified metrics for resource consumption and a taxonomy for decision- and execution-level failure attribution. Within this framework, we adapt 7 widely used benchmarks spanning 24 domains across single-agent, multi-agent, and safety-critical scenarios, and conduct a large-scale empirical analysis over 400K rollouts and 5B tokens on 15 models. The results show that scaffold choice and environmental volatility materially shift benchmark outcomes in both directions, allowing our framework to disentangle intrinsic LLM capabilities from framework- and environment-induced artifacts. We further demonstrate its extensibility as a secure testbed for safety-critical domains. Codes and benchmarks at are available at https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities, https://huggingface.co/datasets/whfeLingYu/Unified_Agent_Framework.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

SkillDAG: 大規模な LLM スキル選択のための自己進化型型スキル グラフ

LLM エージェントが大規模なスキル ライブラリを採用するにつれて、適切なサブセットの選択は、類似性の一致の問題ではなく、構造的な問題になります。つまり、スキルは相互に依存、競合、特殊化、または重複するため、完全な列挙と類似性の埋め込みの両方には見えない構造になります。 SkillDAG は、スキル間の関係を型付き有向グラフとしてモデル化し、それを推論時のエージェント呼び出し可能な構造検索インターフェイスとして LLM エージェントに公開します。固定の検索パイプラインに組み込まれるのではなく、実行中にクエリされて展開されます。各検索では、ベクトル一致、型付きエッジ近傍、競合信号が返され、提案後コミット プロトコルにより、エージェントは実行に裏打ちされたエッジを登録できるため、グラフはエピソード全体で構造を蓄積します。 ALFWorld と MiniMax-M2.7 を使用した SkillsBench では、SkillDAG は 67.1% の成功と 27.3% の報酬に達し、報告されている最も強力なスキルのグラフのベースラインを +12.8 ポイントと +8.6 ポイント上回りました。アドバンテージは gpt-5.2-codex に移植され、固有の SkillsBench Ret@K は、一致したクエリの下で 65.5 から 78.2 に上昇します。これらの利点は、固定シード拡散パイプラインが劣化するプールが 10 倍に成長しても頑健性を維持する候補ランキング、および以前のヒットを排除することなくグラウンドトゥルースの再現を拡大するセットモノトーンのオンライン編集など、分離可能なメカニズムに由来します。

原文 (English)

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend on, conflict with, specialize, or duplicate one another, a structure invisible to both full enumeration and embedding similarity. We present SkillDAG, which models inter-skill relationships as a typed directed graph and exposes it to an LLM agent as an inference-time, agent-callable structural retrieval interface, queried and evolved during execution rather than baked into a fixed retrieval pipeline: each search returns vector matches, typed-edge neighbors, and conflict signals, and a propose-then-commit protocol lets the agent register execution-backed edges so the graph accumulates structure across episodes. On ALFWorld and SkillsBench with MiniMax-M2.7, SkillDAG reaches 67.1% success and 27.3% reward, exceeding the strongest reported Graph-of-Skills baseline by +12.8 and +8.6 points; the advantage ports to gpt-5.2-codex, and intrinsic SkillsBench Ret@K rises from 65.5 to 78.2 under matched queries. These gains trace to isolable mechanisms: candidate ranking that stays robust as the pool grows 10x where a fixed seeding-diffusion pipeline degrades, and set-monotone online edits that enlarge ground-truth recall without evicting prior hits.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

ベンチマーク監査における信頼性ギャップ: 汚染検出の障害モードとしての分布のシフトとスケール

評価例がモデルのトレーニング データに現れるベンチマーク汚染は、LLM 評価の妥当性を脅かします。トレーニング データのメンバーシップを検出するための統計ツールは存在しますが、ほぼ独占的に管理された学術体制、つまり大規模で均質な事前トレーニング コーパスと透明な単一ステージ トレーニング パイプラインでのみ検証されています。これらの方法が現実的な監査シナリオにおいて信頼性を維持できるかどうかは、依然として不明です。私たちは、十分に研究されていない 2 つの障害モードを特定します。1 つは、疑わしいセットと検証セットが IID の仮定に違反する場合に発生する分布シフト、もう 1 つは、ベンチマークがトレーニング前のコーパスよりも桁違いに小さいために発生するスケール制約です。私たちは、複数のファミリー (Pythia、OLMo~2、特殊な文化的および医療的 LLM を含む) およびスケール (最大 27B) からの 27 のモデルにわたって、LLM データセット推論、ポストホック データセット推論、CoDeC という 3 つの主要なパラダイムを体系的に評価します。次に、分析を最先端の業界モデルにさらに拡張します。 335 件の評価のうち、正しい結果が得られたのは 199 件のみでした。 LLM データセット推論では、分布シフトの下で偽陽性が発生し、ポストホック データセット推論はベンチマーク スケールでは能力が不足し、CoDeC は個々のベンチマーク分割を検証するには不十分な粗い出所信号しか提供しません。私たちの結果は、管理された検証と実際のベンチマーク監査の間に体系的な信頼性のギャップがあることを明らかにし、統計的検出がまだ透明なデータ来歴に取って代わることができないことを示しています。私たちはさらなる研究のためにベンチマークをオープンソースにしています。

原文 (English)

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data membership exist, but have been validated almost exclusively in controlled academic regimes: large, homogeneous pre-training corpora and transparent, single-stage training pipelines. Whether these methods remain reliable in realistic auditing scenarios remains unclear. We identify two under-studied failure modes: distribution shift, which arises when suspect and validation sets violate the IID assumption, and scale constraints, which arise because benchmarks are orders of magnitude smaller than pre-training corpora. We systematically evaluate three leading paradigms, LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC, across 25 models from multiple families (including Pythia, OLMo 2, and specialised cultural and medical LLMs) and scales (up to 27B). We then further extend our analysis to frontier industry models. Across 335 evaluations, only 201 yield correct outcomes. LLM Dataset Inference results in false positives under distribution shift, Post-Hoc Dataset Inference is underpowered at benchmark scale, and CoDeC provides only coarse provenance signals that are insufficient to verify individual benchmark splits. Our results reveal a systematic reliability gap between controlled validation and practical benchmark auditing, and show that statistical detection cannot yet replace transparent data provenance. We open-source our benchmark for further research.

13:00 JSTエージェント

トークンが取得されない: AI エージェント出力のサンプリング、状態、および変動性

Agentic AI システムは、実行ごとに異なる動作をする可能性があります。同じリクエストでも、異なる計画、異なるツール呼び出し、異なるコード編集、または異なる最終応答が生成される場合があります。このような変動は、しばしば混同されるいくつかの層から生じます。基礎モデルは大規模な事前トレーニング済みモデルであり、通常は多くの下流タスクに適応でき、入力コンテキストを出力に対する予測にマッピングします。現在のエージェントの多くでは、そのモデルは、計画、ツールの呼び出し、結果の観察、状態の更新を行うオーケストレーション ループに組み込まれています。このようなシステムにおけるばらつきの明示的な固有の原因の 1 つは、トークンの生成です。モデルは、考えられる次のトークンのスコアを計算し、スコアは確率に変換され、デコーダーは擬似乱数ジェネレーターを使用してトークンをサンプリングすることがあります。サンプリングされたトークンの小さな違いは、異なるツール呼び出し、コード パス、検索クエリ、またはエージェントの状態に上向きに伝播する可能性があります。変動のその他の原因は、環境の変化、ライブデータ、サービス提供インフラストラクチャ、バッチ効果、数値の詳細など、トークン サンプリングに付随するものです。これらの層を分離することで、原稿は、エージェント AI システムを確率的と呼ぶことが何を意味するのか、そのような変動が一致した条件下で再現できる場合、そしてなぜ決定論的な実行が展開された設定で同一の動作を暗示する必要がないのかを明確にしています。

原文 (English)

The Token Not Taken: Sampling, State, and the Stochasticity of AI Agents

Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a different final answer. Such variability arises from several layers that are often conflated. At the core of many current agents is a foundation model, a large pretrained model adaptable to many downstream tasks, embedded in an orchestration loop that plans, calls tools, observes results, and updates state. One explicit intrinsic source of variability in such systems is token generation: the model computes scores over possible next tokens, the scores are converted into probabilities, and a decoder may sample tokens using a pseudo-random number generator. A small sampled token difference can then cascade downstream into a different tool call, code path, search query, or agent state. Other sources of variability are extrinsic to token sampling, including changing environments, live data, serving infrastructure, batch effects, and numerical details. By separating these layers, this tutorial clarifies what it means to call agentic AI systems stochastic, when such variability can be reproduced under matched conditions, and why deterministic execution need not imply identical behavior in deployed settings.

13:00 JST研究/論文

サンプル選択のバイアスによりモデルの崩壊が引き起こされる場合

合成データに対する再帰的トレーニングの普及により、データ不足は緩和されますが、トレーニングを繰り返すと分布の裾が侵食され、出力が均質化されるため、モデルが崩壊する危険性があります。データの選択は解決策として広く考えられていますが、その信頼性は検証者が使用する参照分布に大きく依存します。我々は、各検証者がターゲット多様体の小さく断片化された偏ったスライスのみを観察する低リソースの検証体制では、選択自体が偏ることを示します。この状況は、生データをプールすることができず、ローカル参照が本質的に不完全である、ヘルスケア コンソーシアムや独自の金融機関など、リソースが少ないデータ サイロで当然発生します。結果として、選択はローカル多様体と整列したサンプルを優先的に保持しながら、グローバルに関連するテールモードを枝刈りし、崩壊に対する保護手段から崩壊を沈殿させるメカニズムに変わります。私たちは、このようなサイロ化された選択が崩壊を加速し、べき乗則の多様性の衰退を引き起こすことを理論的に証明します。最初の緩和策として、生データを共有せずに複数のサイロから Wasserstein プロキシ参照を構築します。経験的な結果は、偏った分布ではローカル参照の選択が失敗する一方、協調的なプロキシ参照は多様性の低下を軽減することを確認しており、実データのカバレッジが断片的または不足している場合、再帰的合成データ パイプラインには特に注意が必要であることを示唆しています。

原文 (English)

When Sample Selection Bias Precipitates Model Collapse

The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes distributional tails and homogenizes outputs. Data selection is widely viewed as a remedy, yet its reliability depends critically on the reference distribution used by the verifier. We show that in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection itself becomes biased. This situation naturally arises in low-resource data silos such as healthcare consortia or proprietary financial institutions, where raw data cannot be pooled and local references are inherently incomplete. As a result, selection preferentially retains samples aligned with the local manifold while pruning globally relevant tail modes, turning from a safeguard against collapse into a mechanism that precipitates it. We theoretically prove that such siloed selection accelerates collapse and induces power-law diversity decay. As an initial mitigation, we construct Wasserstein proxy references from multiple silos without sharing raw data. Empirical results confirm that local-reference selection fails on skewed distributions, whereas collaborative proxy references mitigate diversity degradation, suggesting that recursive synthetic-data pipelines require particular caution when real-data coverage is fragmented or scarce.

13:00 JSTLLM/生成AIエージェント

HarnessX: 構成可能、適応性、進化可能なエージェント ハーネス ファウンドリ

AI エージェントのパフォーマンスは、モデルがどのように観察、推論、動作するかを仲介するプロンプト、ツール、メモリ、制御フローで構成されるランタイム ハーネスに大きく依存します。しかし、今日のハーネスは大部分が手作りで静的なままです。新しいモデルやタスクごとに特注の足場が依然として必要であり、実行中に生成される豊富な痕跡が体系的な改善に蒸留されることはほとんどありません。構成可能、適応性、進化可能なエージェント ハーネスのファウンドリである HarnessX を紹介します。 HarnessX は、置換代数を介して型指定されたハーネス プリミティブをアセンブルし、記号適応と強化学習の間の操作ミラーに基づいたトレース駆動型マルチエージェント進化エンジンである AEGIS を通じてそれらを適応させ、軌道をハーネス更新とモデル トレーニング信号の両方に変えることでハーネス モデル ループを閉じます。 5 つのベンチマーク (ALFWorld、GAIA、WebShop、tau^3-Bench、および SWE-bench Verified) にわたって、HarnessX は平均 +14.5% (最大 +44.0%) の利益をもたらし、ベースラインが最も低いところでは利益が最大になります。これらの結果は、エージェントの進歩がモデルのスケーリングのみによってもたらされる必要はないことを示唆しています。実行フィードバックからランタイム インターフェイスを構成および進化させることは、実用的で補完的な手段です。完全なコードベースは将来のリリースでオープンソース化される予定です。

原文 (English)

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet today's harnesses remain largely hand-crafted and static: each new model or task still demands bespoke scaffolding, and the rich traces produced during execution are rarely distilled back into systematic improvement. We introduce HarnessX, a foundry for composable, adaptive, and evolvable agent harnesses. HarnessX assembles typed harness primitives via a substitution algebra, adapts them through AEGIS, a trace-driven multi-agent evolution engine grounded in an operational mirror between symbolic adaptation and reinforcement learning, and closes the harness-model loop by turning trajectories into both harness updates and model training signal. Across five benchmarks (ALFWorld, GAIA, WebShop, tau^3-Bench, and SWE-bench Verified), HarnessX yields an average gain of +14.5% (up to +44.0%), with gains largest where baselines are lowest. These results suggest that agent progress need not come from model scaling alone: composing and evolving runtime interfaces from execution feedback is an actionable and complementary lever. The complete codebase will be open-sourced in a future release.

13:00 JSTエージェントビジネス/資金調達研究/論文

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a promi…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation

Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a determin…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

グラウンディングされた反復言語計画: パラメーター化された世界モデルが LLM エージェントにおける幻覚の伝播をどのように軽減するか

言語エージェントの世界モデルには 2 つの便利な形式があります。エージェントベースの世界モデルは LLM API を呼び出し、言語で柔軟に推論しますが、そのエラーは幻覚的な状態変化として現れ、通常の回帰損失ではスコアを付けるのが困難です。パラメーター化された世界モデルは、トレーニングされた遷移予測子です。そのエラーは、NodeMSE、デルタ精度、妥当性精度などの量を使用して測定する方が簡単ですが、通常、スタンドアロン プランナーとしては弱いです。これら 2 つのファミリーを 4 つのグラフ構造の計画ベンチマークで比較し、エージェントベースのケースの操作上の幻覚測定基準を導入します。この比較により、\textbf{Grounded Iterative Language Planning} (GILP) が動機付けられ、小規模なパラメーター化されたバックボーンのみをトレーニングし、それを API ベースのエージェント推論と組み合わせます。バックボーンは、有効なアクション、予測された状態デルタ、リスク、および価値を提供します。 LLM はアクションと想像上のデルタを草案します。そして、この 2 つが一致しない場合には、整合性ゲートが修正を要求します。実際の G​​PT-4o-mini 通話では、GILP は幻覚状態の割合を 0.176 から 0.035 に削減します。調整されたシミュレーターアブレーションでは、最大 22% の余分な LLM コールを追加するだけで、成功率が 0.668 から 0.838 に上昇します。

原文 (English)

Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents

World models for language agents come in two useful forms. An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with ordinary regression losses. A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity accuracy, but it is usually weaker as a standalone planner. We compare these two families on four graph-structured planning benchmarks and introduce operational hallucination metrics for the agent-based case. The comparison motivates Grounded Iterative Language Planning(GILP), which trains only a small parameterized backbone and combines it with API-based agent reasoning. The backbone supplies valid actions, predicted state deltas, risk, and value; the LLM drafts an action and imagined delta; and a consistency gate asks for revision when the two disagree. On real GPT-4o-mini calls, GILP reduces hallucinated-state rate from 0.176 to 0.035. In calibrated simulator ablations, it raises success from 0.668 to 0.838 while adding only ~22% extra LLM calls.

13:00 JSTLLM/生成AI

推論からの真実の検索: LLM 軌道を制御するための動的表現編集フレームワーク

思考連鎖や「待機」プロンプトなどの大規模言語モデル (LLM) 推論を強化する現在のアプローチは、主にモデルの思考を促進しますが、多くの場合、モデルを真実に導くことができません。 Representation Editing (RepE) は固有のコントロールを提供しますが、動的な推論軌跡への応用についてはまだ研究が進んでいません。この研究では、展開する推論の連鎖の中で真実の幾何学を調査することで、このギャップを埋めます。私たちは 3 つの重要な洞察を明らかにします。 (1) 真実は文レベルでコード化されており、潜在的な推論パターンと絡み合っています。 (2) 効果的な介入は不確実性原理と減衰効果に従い、初期の高エントロピー分岐への局在化が必要です。 (3) 単純なステアリング ベクトルはノイズの影響を受け、正しい軌道に付随的な損害を与える危険があります。これらの発見に基づいて、動的 RepE フレームワークである DynaSteer を提案します。 DynaSteer はパターン クラスタリングを採用して推論多様体を解きほぐし、Fisher-LDA を利用して純化された真実を投影します。先読みエントロピーを動的に監視することで、必要な場合にのみ選択的に軌道を操縦し、ロールバックします。いくつかの MATH ベンチマークに関する包括的な実験結果により、DynaSteer の有効性が検証され、ドメイン外コーディング タスクに関する実験により、その一般化能力がさらに確認されました。私たちのコードは https://github.com/tianlwang/DynaSteer で公開されています。

原文 (English)

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and "Wait" prompts, primarily encourage models to think more, yet often fail to guide them toward Truth. While Representation Editing (RepE) offers a intrinsic control, its application to dynamic reasoning trajectories remains underexplored. In this work, we bridge this gap by investigating the geometry of truth within unfolding reasoning chains. We uncover three critical insights: (1) Truth is encoded at the sentence level and is entangled with latent reasoning patterns; (2) Effective intervention follows an Uncertainty Principle and a Decay Effect, requiring localization to early, high-entropy forks; (3) Naive steering vectors suffer from noise, risking collateral damage to correct trajectories. Based on these findings, we propose DynaSteer, a dynamic RepE framework. DynaSteer employs pattern clustering to disentangle reasoning manifolds and utilizes Fisher-LDA to project purified truth. By dynamically monitoring lookahead entropy, it selectively steers and rolls back trajectories only when necessary. Comprehensive experimental results on several MATH benchmark verify the effectiveness of DynaSteer, and experiments on out-of-domain coding tasks further confirm its generalization ability. Our code is publicly available at https://github.com/tianlwang/DynaSteer.

13:00 JSTエージェント

SAGA: 長期的な CivRealm 戦略計画のためのシーンを認識し、目標を進化させるエージェント

複雑な戦略ゲームにおける長期的な戦略計画には、不完全な情報とまばらな報酬の下で、複数の意思決定領域にわたる同時推論が必要です。既存の LLM ベースのエージェントは、生のタイル座標によるシーンのブラインドネス、モノリシックな状態ダンプによるコンテキストのオーバーフローとドメインの結合、各エピソードを個別に扱う浅いクロスゲーム学習という 3 つの系統的な障害に悩まされています。我々は、それぞれ 1 つのクラスの障害を直接ターゲットとする 3 つのメカニズムを備えた LLM マルチエージェント フレームワークである SAGA を紹介します。(i) ゲーム エンティティ間の型指定された空間関係をユニットごとの自然言語コンテキストにエンコードするマップ セマンティック シーン グラフ。グローバルなトークン インフレーションを行わずに空間盲目を解決します。 (ii) オンデマンドで詳細なドメイン状態を取得し、専用の専門コントローラーにドメインごとのディレクティブをディスパッチして、コンテキスト オーバーフロー、ドメイン結合、および機械的制約違反を排除するツール拡張プランナー。 (iii) 定期的なゲーム内目標生成と構造化されたゲーム間の因果関係の事後分析を組み合わせたデュアルホライズン フィードバック ループにより、手動による報酬エンジニアリングを行わずに原則に基づいた戦略的進化が可能になります。 FreeCiv で評価された SAGA は、2 つの最も強力なベースラインよりも低い分散で最高の平均文明スコア (環境で唯一のまばらな目標報酬) を達成し、複数の目標の競合下で最も簡単に犠牲になるリソース軸であるインフラストラクチャ建設のすべてのベースラインを大幅に上回る唯一の方法です。これは、ほとんどの対戦ゲームで 2 つの最も強力なベースラインを上回り、出力トークン (主要なデコード コスト) を 27% 削減します。クロスゲーム進化モジュールを搭載した SAGA は、連続する 5 つのエピソードにわたって最高のエンドオブチェーン スコアに達します。アブレーション研究により、各構造コンポーネントが独立してこの利点に貢献していることが確認されています。

原文 (English)

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and sparse reward. Existing LLM-based agents suffer from three systematic failures: scene blindness from raw tile coordinates, context overflow and domain coupling from monolithic state dumps, and shallow cross-game learning that treats each episode in isolation. We present SAGA, an LLM multi-agent framework with three mechanisms each directly targeting one class of failure: (i) a Map-Semantic Scene Graph that encodes typed spatial relations among game entities into per-unit natural-language context, resolving spatial blindness without global token inflation; (ii) a Tool-Augmented Planner that pulls fine-grained domain state on demand and dispatches per-domain directives to dedicated specialist controllers, eliminating context overflow, domain coupling, and mechanical constraint violations; and (iii) a Dual-Horizon Feedback Loop that combines periodic within-game goal generation with structured cross-game causal post-mortem, enabling principled strategic evolution without manual reward engineering. Evaluated on FreeCiv, SAGA attains the highest mean civilization score -- the environment's sole sparse objective reward -- with lower variance than the two strongest baselines, and is the only method that significantly surpasses every baseline on infrastructure construction, the resource axis most readily sacrificed under multi-objective conflict. It outscores the two strongest baselines in most head-to-head games while cutting output tokens (the dominant decoding cost) by 27%. Equipped with the cross-game evolution module, SAGA reaches the highest end-of-chain score across five successive episodes. Ablation studies confirm that each architectural component contributes independently to this advantage.

13:00 JSTエージェントClaude

AgentBound: 自律型 AI エージェントのための検証可能な行動ガバナンス

自律型 AI エージェントは、人間のプリンシパルに代わって、金融取引、外部通信、企業ワークフローなどの結果的なアクションを実行することが増えています。既存のエージェント インフラストラクチャは、ワークロードの認証とリソース アクセスの制御を ID フェデレーションと委任された承認に依存していますが、現在の動作および運用コンテキストで承認されたアクションを実行する必要があるかどうかを判断できません。自律型 AI エージェントに検証可能な動作監視を提供するランタイム ガバナンス フレームワークである AgentBound を紹介します。 AgentBound は、委任された承認、所有者が署名した行動規定、およびサイト アクション契約という 3 つの独立した権限を使用して、提案された各アクションを評価します。彼らの判断は、実行前にアクションを許可するか、検討するか、拒否するかを決定するための正式な意思決定モデルを通じて保守的に構成されています。説明責任を提供するために、AgentBound は、すべてのアクションを正確な委任、ポリシー、意思決定を管理するセマンティック アーティファクトにバインドする暗号的に検証可能なガバナンス レシートを生成し、独立したリプレイ検証とポリシーの来歴を可能にします。このフレームワークでは、長期実行エージェントに対する常駐委任も導入されており、取り消し可能性と制限された権限を維持しながら、継続的に更新されるガバナンス ポリシーの下で定期的なワークロードを実行できるようになります。正式な基盤、システム アーキテクチャ、ガバナンス受信プロトコル、およびガバナンスの正確性、権限構成、説明責任を評価するためのベンチマーク フレームワークである AgentBound-Bench を紹介します。 AgentBound は、モデルの調整を置き換えるのではなく、承認と実行の間に決定論的なガバナンス層を提供することでそれを補完し、ガバナンスを信頼する必要があるプロセスから独立して検証できるプロセスに変換します。

原文 (English)

Behavioral Governance for Autonomous AI Agents: The AgentBound Framework

Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications, and enterprise workflows. Existing agent infrastructure relies on identity federation and delegated authorization to authenticate workloads and control resource access, but it cannot determine whether an authorized action should be executed under the current behavioral and operational context. We present AgentBound, a runtime governance framework that provides verifiable behavioral oversight for autonomous AI agents. AgentBound evaluates each proposed action using three independent authorities: delegated authorization, owner-signed behavioral constitutions, and site action contracts. Their judgments are conservatively composed through a formal decision model to determine whether an action should be permitted, reviewed, or denied before execution. To provide accountability, AgentBound generates cryptographically verifiable governance receipts that bind every action to the exact delegation, policy, and semantic artifacts governing the decision, enabling independent replay verification and policy provenance. The framework also introduces standing delegation for long-running agents, allowing periodic workloads to operate under continuously refreshed governance policies while preserving revocability and bounded authority. We present the formal foundation, system architecture, governance receipt protocol, and AgentBound-Bench, a benchmark framework for evaluating governance correctness, authority composition, and accountability. Rather than replacing model alignment, AgentBound complements it by providing a deterministic governance layer between authorization and execution, transforming governance from a process that must be trusted into one that can be independently verified.

13:00 JSTLLM/生成AIエージェント研究/論文

図書館を超えて: 研究数学を自動形式化するためのエージェント フレームワーク

大規模言語モデル (LLM) は、数学的推論において優れた能力を実証していますが、人間の検出を回避する微妙なエラーを頻繁に生成します。 Lean 4 のような正式な数学言語は機械的な証明チェックを提供し、自動形式化、つまり自然言語数学を検証可能なコードに自動的に変換する必要性を強く促します。最近の傾向は、標準プログラミング向けに大幅に最適化された汎用 LLM が、リーン向けに明示的に微調整された小規模なモデルよりも優れたパフォーマンスを発揮することを示しています。この変化を活用して、一般的なコーディング LLM を利用したエージェント自動形式化フレームワークを導入します。私たちのシステムの中核には、研究レベルの数学に合わせて調整されたマルチエージェント パイプラインを管理するオーケストレーターがあります。最先端の研究は Mathlib などの既存のライブラリの範囲外の概念に依存することが多いため、私たちのシステムは必要な型定義を動的に拡張し、主要な定理を形式化する前に新しい補助補題技術を介してそれらを検証します。私たちはこのアプローチを PutnamBench に適用し、32 個の問題のランダムなサンプルに対して機械チェックされたリーン証明を作成しました。さらに、ACM Symposium on Theory of Computing (STOC) の組み合わせ論、通信の複雑さ、機構設計、学習理論にわたる 5 つの論文に基づいてシステムを評価し、主定理の形式化に成功し、生成された形式化を専門家とともに検証しました。 5 つすべてについて、ステートメントと並行して証明も形式化します。特に、そのうち 2 つはリーンのカーネルを超える公理を使用せずに証明されています。すべての形式化は https://beyondthelibrary.github.io/formal_arxiv で入手できます。

原文 (English)

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evade human detection. Formal mathematical languages like Lean 4 offer mechanical proof checking, strongly motivating the need for autoformalization: the automatic translation of natural language mathematics into verifiable code. Recent trends indicate that general-purpose LLMs, heavily optimized for standard programming, now outperform smaller models explicitly fine-tuned for Lean. Leveraging this shift, we introduce an agentic autoformalization framework powered by general coding LLMs. At the core of our system is an orchestrator that manages a multi-agent pipeline tailored for research-level mathematics. Because cutting-edge research frequently relies on concepts outside the scope of existing libraries like Mathlib, our system dynamically extends necessary type definitions and validates them via a novel Auxiliary Lemma technique before formalizing the primary theorems. We applied our approach to PutnamBench, producing machine-checked Lean proofs for a random sample of 32 problems. Furthermore, we evaluate our system on five papers from the ACM Symposium on Theory of Computing (STOC) spanning combinatorics, communication complexity, mechanism design, and learning theory, successfully formalizing their main theorems and validating the generated formalizations with human experts; for all five we also formalize the proofs alongside the statements, and notably two of them are proved with no axioms beyond Lean's kernel. All of our formalizations are available at https://beyondthelibrary.github.io/formal_arxiv .

13:00 JSTLLM/生成AIエージェント研究/論文

ClawArena チーム: 言語モデル エージェントのサブエージェント オーケストレーションと動的ワークフローのベンチマーク

運用環境の大規模言語モデル (LLM) エージェントは、単独の問題解決者としてではなく、マネージャーとして導入されることが増えています。メイン モデルは、特化したサブエージェントを作成し、作業を委任し、動的なワークフローを通じて並列非同期のリターンを調整します。 1 つのモデルがそのようなチームを実際に実行できるかどうかは、ほとんど測定されていません。既存のベンチマークは、ポリシー自体のタスク解決や固定マルチエージェント システムの緊急動作をスコア化しますが、リーダーとして機能する単一の LLM の管理能力を分離するものはありません。この管理能力を測定する、258 の評価ラウンドと 72 の段階的な更新にわたる 41 のマルチターン、マルチモーダル、マルチディレクトリ シナリオのベンチマークである ClawArena-Team を紹介します。メイン エージェントは意図的に制限されています。ネイティブにテキストのみを認識し、ワークスペースの一部のみに直接アクセスします。ローカルで提供される固定のサブエージェント プールを指揮するため、スコアの差は、実際の能力ではなく管理スキルを反映します。すべてのスコアリングは実行ベースであり、LLM 判定はありません。全体的なスコアであるサブエージェント管理スコア (SMS) には、タスクの正確さに最小権限とモダリティ ルーティング係数が乗じられます。 12 の独自モデル、コミュニティホスト型モデル、セルフホスト型モデルにわたる実験では、管理のボトルネックは認識ではなく権限付与であることが示されています (ワークスペース権限の精度が 50% を超えるモデルはありません)。コストと管理品質は分離されています (API コストは 100 倍を超えますが、全体のスコアは 4 倍未満であり、パレート フロンティアで最も安価なオープン モデルです)。そして、ほとんどのリーダーボード スコアは 9.9 ポイントの範囲内に集中していますが、オーケストレーションの動作は 1 桁以上異なります。コードとデータは公開されます。

原文 (English)

ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows. Whether one model can actually run such a team is largely unmeasured: existing benchmarks score a policy's own task-solving or a fixed multi-agent system's emergent behavior, but none isolate the management ability of the single LLM acting as leader. We introduce ClawArena-Team, a benchmark of 41 multi-turn, multimodal, multi-directory scenarios spanning 258 evaluation rounds and 72 staged updates that measures this management ability. The main agent is deliberately constrained: it natively perceives only text and directly accesses only part of the workspace. It commands a fixed, locally served subagent pool, so score differences reflect management skill, not raw capability. All scoring is execution-based with no LLM judge: an overall score -- the Subagent-Management Score (SMS) -- multiplies task correctness by a least-privilege and modality-routing factor. Across twelve proprietary, community-hosted, and self-hosted models, experiments show that the management bottleneck is privilege granting rather than perception (no model exceeds 50% workspace-permission precision); that cost and management quality are decoupled (API cost spans over 100 times while the overall score spans under 4 times, with the cheapest open models on the Pareto frontier); and that most leaderboard scores cluster within a 9.9-point band while orchestration behaviors diverge by more than an order of magnitude. Code is available at https://github.com/aiming-lab/ClawArena.

13:00 JST研究/論文

言語と記号表現間のモダリティ切り替えによる空間推論

人間の推論は本質的に多様です。問題が困難になったとき、私たちは言葉だけで考えることはほとんどありません。私たちは、根底にある概念構造を理解し、間違いを避けるために、図をスケッチしたりグリッドを描いたりすることで推論を外部化することがよくあります。この前提に基づいて、私たちの研究は次のことを調査します。(a) マルチホップのテキスト空間ストーリーをレイアウトやグリッドなどの幾何学認識モダリティに根付かせることで、自然言語ベースの推論と比較して推論が向上するかどうか。 (b) いつ自然言語推論に依存するか、いつ構造化モダリティに切り替えるかをモデルが決定できるかどうか。私たちは、信頼性と複雑さの信号に基づいたスイッチング メトリックを導入することで、これらの疑問に対処します。このメトリックは、空間ストーリーを構造に定着させるとパフォーマンスが向上する可能性が高いと推定します。これは、大規模言語モデル (LLM) 推論における原則に基づいたモダリティ選択への第一歩となります。私たちの設定全体で、自然言語ベースの推論からグリッドベースの表現に切り替えると、LLM のパフォーマンスが最大 42\% 向上し、推論の結果を形成する際のモダリティの選択の重要性が強調されます。

原文 (English)

Spatial Reasoning via Modality Switching Between Language and Symbolic Representation

Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketching diagrams or drawing grids to understand the underlying conceptual structure and avoid mistakes. Building on this premise, our research investigates: (a) whether grounding multi-hop textual-spatial stories into geometry-aware modalities, such as layouts or grids, improves reasoning compared to natural language-based inference; and (b) whether a model can decide when to rely on natural language reasoning and when to switch to a structured modality. We address these questions by introducing a switching metric based on trustworthiness and complexity signals, which estimates when grounding a spatial story into structure is likely to improve performance. This takes a first step toward principled modality selection in Large Language Model (LLM) reasoning. Across our settings, switching from natural language-based reasoning to a grid-based representation improves LLM performance by up to 42%, highlighting the importance of modality choice in shaping reasoning outcomes.

13:00 JSTエージェント

生物学的プロトコルの自動生成と実行のための自己進化型エージェント システム

自律的なウェットラボ実験には、もっともらしいプロトコールのテキスト以上のものが必要です。つまり、生物学的意図、定量的手順、デバイスの制約、および実験のフィードバックが、プロトコールや SOP の設計からコードや物理的な実行に至るまで常に調整されていなければなりません。私たちは、この変換を実験的な自動化問題としてテストするための、専門家に基づいたベンチマークおよび評価フレームワークとともに、自己進化するマルチエージェント システムである ProtoPilot を開発しました。このフレームワークは、98 のゴールドスタンダード プロトコル、ウェットラボのエキスパート ルーブリック、デバイス レベルの妥当性ゲート、および実際の実験テストから派生した 294 の合成生物学および分子生物学のタスクに及びます。 ProtoPilot には、レイヤーごとの検証機能、マルチエージェント オーケストレーション、ランタイムで更新されるスキル ライブラリが組み込まれており、プロトコルの生成、SOP の拡張、SDK 準拠のコードの合成、ウェット ラボのフィードバックからのワークフローの修正を行います。 OpenTrons-AI の 32.35% と比較して、Top@3 専門家選択率 90.2%、全体のプロトコルからコードへのゲート通過率 89.5%、Opentrons 通過率 88.24% を達成しました。ウェットラボ検証により、解釈可能な読み取り値、サンガーで確認された生成物、フィードバック補正された PCA で構築された DNA ターゲットが生成され、自律実験への検証可能なルートが確立されました。これらの結果を総合すると、評価フレームワークが自律型ウェットラボ自動化の実行関連要件を捉えており、ProtoPilot がプロトコルとコード生成を検証済みの実行とフィードバックに基づく改訂に変換することで要件を満たすことができることを示しています。

原文 (English)

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution. We developed ProtoPilot, a self-evolving multi-agent system, together with an expert-grounded benchmark and evaluation framework for testing this conversion as an experimental automation problem. The framework spans 294 synthetic-biology and molecular-biology tasks derived from 98 gold-standard protocols, wet-lab expert rubrics, device-level validity gates and real experimental tests. ProtoPilot incorporates layer-wise verifiability, multi-agent orchestration and a runtime-updated skill library to generate protocols, expand SOPs, synthesize SDK-compliant code and revise workflows from wet-lab feedback. It achieved a Top@3 expert-preference rate of 90.2%, an overall protocol-to-code gate pass rate of 89.5% and an Opentrons pass rate of 88.24%, compared with 32.35% for OpenTrons-AI. Wet-lab validation produced interpretable readouts, Sanger-confirmed products and feedback-corrected PCA-assembled DNA targets, establishing a verifiable route to autonomous experimentation. Together, these results show that the evaluation framework captures execution-relevant requirements for autonomous wet-lab automation, and that ProtoPilot can meet them by converting protocol and code generation into validated execution and feedback-guided revision.

13:00 JST研究/論文

MMM データ モデル -- 分散型ナレッジ コモンズにおける知識の相互運用性の標準仕様

多くの情報システムはドキュメントを中心に構築されており、印刷物作成とリニア読み取り用に最適化された自己完結型ユニットです。ドキュメント中心の組織は大規模な普及には効果的ですが、知識を構造化し、更新し、共有し、再利用する方法に制約があります。形式的なアプローチはこれらの制限の一部に対処しますが、人間の使いやすさや範囲などの他のシステム特性よりも形式的な構造を優先するため、広範な貢献と採用を達成するのに苦労しています。 AI システムは文書作成を再構築していますが、人間による知識の表現と交換のための従来の文書に代わる統合されたポータブルな代替手段は提供されていません。この論文では、学際的な共同研究の実際的なニーズから生まれた知識文書化のためのデータ モデルである MMM を紹介し、情報システムの設計空間の比較分析の中に位置づけています。 MMM は、一連の規範的な制約とフリーテキスト ラベルの表現の自由を組み合わせたものです。セマンティックな収束を必要とせずに、分野、アプリケーション、展開全体で相互運用できるように設計されています。リファレンス実装とパイロット展開データは、実装可能性と初期の使いやすさを示しています。

原文 (English)

The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons

Many information systems are built around documents: self-contained units optimised for print production and linear reading. While effective for large-scale dissemination, the document-centric organisation constrains how knowledge can be structured, updated, shared, and reused. Formal approaches address some of these limitations but struggle to achieve widespread contribution and adoption due to their prioritisation of formal structure over other system properties such as human usability and scope. AI systems are reshaping document production, but without providing a unified portable alternative to traditional documents for humans' expression and exchange of knowledge. This paper presents MMM, a data model for knowledge documentation that emerged from the practical needs of interdisciplinary collaborative research, and positioned here within a comparative analysis of the design space of information systems. MMM combines a small set of normative constraints with the expressive freedom of free-text labels. It is designed for interoperability across disciplines, applications and deployments without requiring semantic convergence. A reference implementation and pilot deployment data demonstrate implementability and early usability.

13:00 JSTLLM/生成AI

Theoria: 非公式推論状態に対する書き換え許容性の検証

AI システムの答えを信頼できるのはどのような場合ですか?形式的証明アシスタントは確実性を提供しますが、問題分布のほとんどには到達できません。スカラー LLM ジャッジはカバレッジを提供しますが、事後的に監査できない不透明なスコアを生成し、他の LLM と同じ一貫性の問題にさらされます。私たちは、このギャップを埋める検証アーキテクチャである Theoria を紹介します。候補解は、型指定された状態遷移のシーケンスに書き換えられます。各状態遷移は、引用、計算、または問題によって与えられた事実など、明示的な正当化によってライセンスされ、すべての遷移は独立して監査可能です。基本的な不変条件は変化の完全性です。連続する証明状態間のすべての違いを考慮する必要があるため、隠れた前提は黙って通過するのではなく、許可されていない突然変異として表面化します。 HLE-Verified Gold (185 のテキストのみのエキスパートの問題) では、Theoria は 91.4% の厳密な精度で 105 を認定しています (Wilson 95% CI [84.5%、95.4%])。すべての認証では、人間が判読できる証明トレースが生成され、各ステップに個別にチャレンジできます。ホリスティック LLM ジャッジは、一致するカバレッジでは同等の精度を達成しますが、別の問題 (Jaccard 0.14 ~ 0.36) では失敗するため、アプローチは補完的になります。 15 のドメインにわたる 95 件の敵対的毒物証明について、構造化された裁判官は 94.7% を捕捉したのに対し、総合的な判断では 83.2% を捕捉しました (p= 0.0017)。全体の 11.5 pp のギャップは、隠れた前提 (90.6% 対 62.5%、28 pp の差) と捏造された引用 (100% 対 90%) に集中しており、形式的な分析が利点を予測するエラー クラスです。利点が予測されない算術および定理の誤用エラーのパフォーマンスは同じです。 GPQA ダイヤモンド (n= 65) では、認定精度は 97.1% (Wilson CI [85.1%、99.5%]) です。

原文 (English)

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each licensed by an explicit justification, whether that be a citation, computation, or problem-given fact, and every transition is independently auditable. The foundational invariant is completeness of change: every difference between consecutive proof states must be accounted for, so hidden premises surface as unlicensed mutations rather than passing silently. On HLE-Verified Gold (185 text-only expert problems), Theoria certifies 105 at 91.4% strict precision (Wilson 95% CI [84.5%, 95.4%]). Every certification produces a human readable proof trace in which each step can be independently challenged. Holistic LLM judges achieve comparable precision at matched coverage but fail on different problems (Jaccard 0.14-0.36), making the approaches complementary. On 95 adversarial poisoned proofs across 15 domains, structured judges catch 94.7% versus 83.2% for holistic judging (p= 0.0017). The overall 11.5 pp gap concentrates in hidden premises (90.6% vs. 62.5%, a 28 pp difference) and fabricated citations (100% vs. 90%), the error classes where the formal analysis predicts an advantage; performance is identical on arithmetic and theorem-misapplication errors, where no advantage is predicted. On GPQA Diamond (n= 65), certified precision is 97.1% (Wilson CI [85.1%, 99.5%]).

13:00 JSTLLM/生成AI

Playing 20 Question Game with Policy-Based Reinforcement Learning

The 20 Questions (Q20) game is a well known game which encourages deductive reasoning and creativity. In the game, the answerer first think…

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文

Introduction to Transformers: an NLP Perspective

Transformers have dominated empirical machine learning models of natural language processing. In this paper, we introduce basic concepts of…

13:00 JST画像/動画生成

Contrastive Deep Learning Reveals Age Biomarkers in Histopathological Skin Biopsies

As global life expectancy increases, so does the burden of chronic diseases, yet individuals exhibit considerable variability in the rate a…

13:00 JSTLLM/生成AIエージェント

Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models

Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simpli…

13:00 JST画像/動画生成ロボティクス

Learning 3D-Gaussian Simulators from RGB Videos

Realistic simulation is critical for applications ranging from robotics to animation. Learned simulators have emerged as a possibility to c…

13:00 JST画像/動画生成

Comparative Analysis of Lightweight CNNs for Resource-Constrained Devices: Predictive Performance, Efficiency Trade-offs, and Initialization Effects

Lightweight convolutional neural networks are often compared using results obtained with different training recipes, input settings, and pr…

13:00 JST研究/論文

MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter…

13:00 JSTLLM/生成AILlamaQwenDeepSeek

Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens

The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scala…

13:00 JSTLLM/生成AI

RedCoder: Automated Multi-Turn Red Teaming for Code LLMs

Large Language Models (LLMs) for code generation (i.e., Code LLMs) have demonstrated impressive capabilities in AI-assisted software develo…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation

Medical report understanding from real-world document images is essential for generating patient-facing explanations and enabling structure…

13:00 JST研究/論文

Uncertain but Useful: Leveraging CNN Training Variability into Data Augmentation

Deep learning (DL) has transformed neuroimaging by delivering state-of-the-art performance with reduced computation times. Yet, the numeric…

13:00 JSTLLM/生成AIビジネス/資金調達

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interac…

13:00 JSTLLM/生成AI

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging

The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration,…

13:00 JSTLLM/生成AIビジネス/資金調達

CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating cr…

13:00 JST研究/論文Alibaba

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of a…

13:00 JST研究/論文

Exploring Large Language Models for Access Control Policy Synthesis and Summarization

Cloud computing is ubiquitous, with a growing number of services being hosted on the cloud every day. Typical cloud compute systems allow a…

13:00 JST画像/動画生成

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

Fine-grained cross-modal alignment aims to establish precise local correspondences between vision and language, forming a cornerstone for v…

13:00 JST画像/動画生成

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment

Molecular biomarker testing in pathology is often costly and tissue-consuming, limiting scalable clinical deployment. Artificial intelligen…

13:00 JSTLLM/生成AIClaude

Gravity-Awareness: Deep Learning Models and LLM Simulation of Human Awareness in Altered Gravity

Earth s gravity fundamentally shapes human behaviour. The brain encodes this force as an internal model of gravity, enabling the prediction…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents

Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principle…

13:00 JST画像/動画生成

Spanning Tree Autoregressive Visual Generation

We present Spanning Tree Autoregressive (STAR) modeling, which can incorporate prior knowledge of images, such as center bias and locality,…

13:00 JST規制/政策

Locality-Aware Continual Unlearning for Diffusion Models

Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations…

13:00 JST画像/動画生成エージェント研究/論文

PPTArena: A Benchmark for PowerPoint Editing

We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions.…

13:00 JSTLLM/生成AI

ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models

Scaling inference-time computation has enabled Large Language Models (LLMs) to achieve strong reasoning performance, but their inherently s…

13:00 JSTエージェント研究/論文GPT / ChatGPTDeepSeek

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Thei…

13:00 JST画像/動画生成

Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of previously observed pri…

13:00 JST研究/論文

When Does Predictive Inverse Dynamics Outperform Behavior Cloning?

Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent w…

13:00 JST研究/論文Llama

Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, r…

13:00 JST研究/論文

LEFT: Learnable Fusion of Tri-view Tokens for Unsupervised Time Series Anomaly Detection

As a fundamental data mining task, unsupervised time series anomaly detection (TSAD) aims to build a model for identifying abnormal timesta…

13:00 JST研究/論文

$\mu$pscaling small models: Principled warm starts and hyperparameter transfer

Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve ef…

13:00 JST画像/動画生成ロボティクス

Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration

Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increas…

13:00 JSTLLM/生成AIエージェント

From Experiments to Expertise: Scientific Knowledge Consolidation for AI-Driven Computational Physics

While large language models (LLMs) have transformed AI agents into proficient executors of computational materials science, performing a hu…

13:00 JST研究/論文

Efficient Federated Conformal Prediction with Group-Conditional Guarantee

Deploying trustworthy AI systems requires principled uncertainty quantification. Conformal prediction (CP) is a widely used framework for c…

13:00 JSTビジネス/資金調達

Adaptive Contracts for Cost-Effective AI Delegation

When organizations delegate text generation tasks to AI providers via pay-for-performance contracts, expected payments rise when evaluation…

13:00 JST画像/動画生成エージェントロボティクス

DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving

Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the…

13:00 JSTエージェントロボティクス研究/論文

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness…

13:00 JSTLLM/生成AI

An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms

Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models…

13:00 JSTLLM/生成AIGPT / ChatGPT

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning

Rerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation. However, current reranking models are typic…

13:00 JST画像/動画生成エージェントロボティクス

Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving

Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文GPT / ChatGPTGemini

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's langu…

13:00 JSTLLM/生成AI画像/動画生成

Cross-Cultural Value Attribution in Large Vision-Language Models

The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their p…

13:00 JSTLLM/生成AIエージェント研究/論文Claude

Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers

Autonomous LLM agents now produce complete research artifacts in machine-learning sandboxes, but real computational physics is harder: expe…

13:00 JSTエージェント

ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents

Embodied agents increasingly rely on modular capabilities that are installed, upgraded, composed, and governed at runtime, yet the interfac…

13:00 JSTLLM/生成AIエージェントClaude

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study…

13:00 JSTエージェント

ChemGraph-XANES: An Agentic Framework for XANES Simulation and Curation

Computational X-ray absorption near-edge structure (XANES) is widely used to interpret local coordination environments, oxidation states, a…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGeminiDeepSeek

Peer-Preservation in Frontier Models

Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can…

13:00 JSTLLM/生成AINVIDIA

FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices

Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data. H…

13:00 JSTLLM/生成AI

Evergreen: Efficient Claim Verification for Semantic Aggregates

With recent semantic query processing engines, semantic aggregation has become a primitive operator, enabling the reduction of a relation i…

13:00 JST画像/動画生成

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution

High-precision medical diagnosis relies not only on static imaging features but also on the implicit diagnostic memory experts instantly in…

13:00 JSTロボティクス

Regression Test Selection for Updated Capability Modules in Compositional ML Systems via Atomic-Quality Probes

Compositional machine-learning (ML) systems assemble runtime behavior from libraries of independently re-trained capability modules. Replac…

13:00 JST研究/論文

The Transformer as a Polar State Estimator

We show that the core components of the Transformer -- attention, residual connections, and normalization -- arise naturally from a single…

13:00 JSTロボティクス

Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates

Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajector…

13:00 JST研究/論文

XSearch: Explainable Code Search via Concept-to-Code Alignment

Semantic code search has been widely adopted in both academia and industry. These approaches embed natural-language queries and code snippe…

13:00 JSTLLM/生成AIGPT / ChatGPT

ContraFix: Skill-Enhanced Contrastive Runtime Analysis for Vulnerability Repair

As software systems grow increasingly complex, automated vulnerability repair (AVR) remains difficult because the materials available to a…

13:00 JSTLLM/生成AIエージェント

BLAgent: Agentic RAG for File-Level Bug Localization

Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty co…

13:00 JST研究/論文

変分オートエンコーダにおける定常コラプスに対するシンプレックス証人証明書

私たちは、変分オートエンコーダーにおける正確な定数の崩壊を研究します。つまり、決定論的なエンコーダーの平均が入力から独立するようになります。事前分布は標準のガウス分布のままです。 VAE トレーニングの前に、データの GMM ベースのビューから事後固定教師を選択し、固定潜在のみのシンプレックス監視をエンコーダー平均に添付します。この構築により、2 つのリンクされたオブジェクトが生成されます。 1 つ目は証明書です。目撃者の予測が教師の最良の定数予測子を改善する場合、エンコーダーの平均は入力に依存しない定数になることはできません。 2 つ目は局所的なエスケープ方向です。崩壊した多様体では、教師残差によってアライメント損失に対するサンプル依存の下降方向が与えられます。あらゆるフルサポート教師事後分布の場合、同じジオメトリにより、教師と証人の位置合わせエラーがゼロの閉じた形式の潜在コードも得られます。そのスケーリングされたバージョンは、定数予測子から正確な教師コードまでのマージン エネルギー パスを追跡し、保護された目撃部分空間内の非崩壊を定量化します。 MNIST、CIFAR-10、および CIFAR-100 でメソッドをインスタンス化します。教師なしの PCA-GMM 教師を検索すると、バニラ VAE は CIFAR-10 および CIFAR-100 の 5 つのシードすべてで教師証人証明書に不合格ですが、RST バリアントは 5 つのシードすべてに合格します。 \(\beta_{\mathrm{KL}}\in\{2,4,8\}\) を使用した崩壊ストレス設定では、バニラ VAE はすべてのシードで再び失敗しますが、RST-alpha-prefit は証明書陽性のままです。両方の自然画像データセットのエスケープ トラジェクトリは、低マージン初期化からウィットネス マージンを増加させ、非ゼロの教師誘発勾配ノルムを示します。分析は、エンコーダ平均値の正確な一定の崩壊に限定されます。生成品質、デコーダの使用、およびその他の崩壊モードについては、別個の問題として残ります。

原文 (English)

A Simplex Witness Certificate and Escape Force for Constant Collapse in Variational Autoencoders

We study exact constant collapse in variational autoencoders: the deterministic encoder mean becomes independent of the input. The prior remains the standard Gaussian. Before VAE training, we select a fixed teacher posterior from a GMM-based view of the data and attach a fixed latent-only simplex witness to the encoder mean. This construction yields two linked objects. The first is a certificate: if the witness prediction improves on the best constant predictor of the teacher, the encoder mean cannot be input-independent constant. The second is a local escape direction: on the collapsed manifold, the teacher residual gives a sample-dependent descent direction for the alignment loss. For any full-support teacher posterior, the same geometry also gives a closed-form latent code with zero teacher-witness alignment error. Its scaled versions trace a margin-energy path from the constant predictor to the exact teacher code, which quantifies non-collapse inside the protected witness subspace. We instantiate the method on MNIST, CIFAR-10, and CIFAR-100. With searched unsupervised PCA-GMM teachers, vanilla VAEs fail the teacher-witness certificate in all five seeds on CIFAR-10 and CIFAR-100, while RST variants pass in all five seeds. Under collapse-stress settings with \(\beta_{\mathrm{KL}}\in\{2,4,8\}\), vanilla VAE again fails in all seeds, whereas RST-alpha-prefit remains certificate-positive. Escape trajectories on both natural-image datasets increase the witness margin from a low-margin initialization and exhibit nonzero teacher-induced gradient norms. The analysis is confined to exact constant collapse of the encoder mean; generation quality, decoder use, and other collapse modes remain separate questions.

13:00 JSTビジネス/資金調達研究/論文

Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts

Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in structured, electronic health reco…

13:00 JSTLLM/生成AIエージェント

SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering

Large language models are increasingly deployed as tool-augmented agents to acquire information beyond parametric knowledge. While recent w…

13:00 JSTロボティクス

Exact equivariance, kept through training, buys zero-shot generalisation across the symmetry group

A latent world model built from an equivariant encoder and predictor inherits a provable symmetry of its training loss: when the dynamics c…

13:00 JSTLLM/生成AI

Enabling KV Caching of Shared Prefix for Diffusion Language Models

Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical chall…

13:00 JSTLLM/生成AIエージェント研究/論文Claude

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify…

13:00 JSTLLM/生成AIハードウェア/半導体

TileFuse: AMD NPU での効率的な量子化 LLM 推論のための融合型混合精度カーネル ライブラリ

オンデバイス LLM 推論に対する需要の高まりに伴い、厳しい電力と熱のバジェットの下でパフォーマンスとエネルギー効率を向上させるために、エッジ SoC は NPU を統合することが増えています。しかし、現在のクライアント NPU に実際に LLM を導入することは依然として困難です。AWQ などの広く使用されている量子化形式は、多くの既存の NPU ソフトウェア スタックにきれいにマッピングされておらず、多くの場合独自仕様であり、低レベルの制御が制限されています。この研究では、量子化 LLM 推論におけるトランス線形層をターゲットとする、AMD XDNA2 NPU 用のメタルに近い混合精度カーネル ライブラリである \textit{TileFuse} を紹介します。 TileFuse は、NPU 固有の量子化スキームに基づいてモデルを強制的に再形成するのではなく、AWQ スタイルの W4A16 や W8A16 などの実用的な低ビット フォーマットを XDNA2 に直接導入します。 TileFuse は、重みレイアウト、メタデータ配置、混合精度マイクロカーネル、配列レベルのデータフローを共同設計します。具体的には、アンパッキング、逆量子化、GEMM/GEMV の実行を単一のカーネル フローに融合し、最大 32K の GEMM 次元をサポートするインターリーブ プレタイリング レイアウトを導入し、完全な 4x8 AIE アレイを利用するように GEMV データフローを再設計します。カーネル レベルの評価全体で、TileFuse は、完全精度のベースラインと比較して、GEMM で最大 121.6%、GEMV で 281% パフォーマンスが向上し、GEMM 上の強力な iGPU ベースラインと比較して 2 倍を超えるパフォーマンスとエネルギー効率の向上を実現します。 Ryzen AI ラップトップでのエンドツーエンド LLM 実験では、TileFuse は、エネルギー消費量を 64.6% 以上削減し、プレフィル レイテンシーを最大 2.0 倍短縮することを達成しました。これらの結果を総合すると、XDNA2 が AWQ スタイルのエッジ LLM 推論の実用的なターゲットであること、および既製の量子化に対するネイティブ NPU サポートにより、実際のクライアント展開で NPU が大幅に使いやすくなることがわかります。

原文 (English)

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs

With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. However, practical LLM deployment on current client NPUs remains difficult: widely used quantization formats such as AWQ do not map cleanly onto many existing NPU software stacks, which are often proprietary and expose limited low-level control. In this work, we present TileFuse, a close-to-metal mixed-precision kernel library for AMD XDNA2 NPUs that targets GEMM/GEMV-based operators in quantized LLM inference. TileFuse brings practical low-bit formats such as AWQ-style W4A16 and W8A16 directly onto XDNA2, rather than forcing the model to be reshaped around an NPU-specific quantization scheme. TileFuse co-designs weight layout, metadata placement, mixed-precision microkernels, and array-level dataflow. Specifically, it fuses unpacking, dequantization, and GEMM/GEMV execution into a single kernel flow, introduces an interleaved pre-tiling layout that supports GEMM dimensions up to 32K, and redesigns GEMV dataflow to utilize the full 4x8 AIE array. Across kernel-level evaluations, TileFuse improves performance by up to 121.6% for GEMM and 281% for GEMV over full-precision baselines, while delivering more than 2x performance and energy-efficiency gains over strong iGPU baselines on GEMM. In end-to-end LLM experiments on Ryzen AI laptops, TileFuse achieves up to 2.0x lower prefilling latency with more than 64.6% lower energy consumption. Together, these results show that XDNA2 is a practical target for AWQ-style edge LLM inference and that native NPU support for off-the-shelf quantization can make NPUs substantially more usable in real client deployments.

13:00 JSTLLM/生成AIGemma

eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

We present eCream-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals.…

13:00 JST画像/動画生成

形態を考慮したサンプルの割り当て: 表面欠陥検出における IoU の鈍感性を克服

Intersection-over-Union (IoU) は、候補提案とグラウンド トゥルース アノテーションの間の空間的位置合わせを評価するための極めて重要な指標として、ポジティブ サンプル セットの品質と視覚検出モデルのトレーニング効果を直接決定します。理論的なモデリングと分析を通じて、IoU 応答曲線上の非感受性領域を明らかにしました。この領域内では、サンプルは、明確な幾何学的重複にもかかわらず、ほぼ同一の IoU スコアを生成します。この制限を克服するために、面積、形状、アスペクト比をカバーする一連の形態学的類似性メトリクスを導入し、ポジティブサンプル割り当てプロセスを改良し、それによってより識別性と信頼性の高いマッチングを保証します。補足的なマッチング スコアは、これらの多次元類似性の平均ベースの集計によって導出され、構造的対応を表す際の IoU の本質的な制限を補償します。理論的には、形態学的類似性を組み込むと、マッチング関数の応答分布が再形成され、有効な方向勾配と多角形のような等応答輪郭の両方が得られます。これにより、各グラウンドトゥルース インスタンスの周囲に高応答領域が厳密に限定され、ポジティブ サンプル選択の精度が大幅に向上します。 YOLOv9 フレームワークに基づく実験では、NEUDET データセットと GC10-DET データセットの両方で一貫したパフォーマンスの向上が実証されています。特に、提案されたアプローチは完全にプラグアンドプレイであり、追加の推論オーバーヘッドが発生しないため、工業用外観検査の展開効率が確保されます。

原文 (English)

Morphology-Aware Sample Assignment: Overcoming IoU Insensitivity for Surface Defect Detection

Intersection-over-Union (IoU), as a pivotal metric for evaluating the spatial alignment between candidate proposals and ground-truth annotations, directly determines the quality of positive sample sets and the training efficacy of visual detection models. Through theoretical modeling and analysis, we uncover a non-sensitive region on the IoU response curve, within which samples yield nearly identical IoU scores despite distinct geometric overlaps. To overcome this limitation, we introduce a set of morphological similarity metrics covering area, shape, and aspect ratio, to refine the positive sample assignment process, thereby ensuring more discriminative and reliable matching. A supplementary matching score is derived via mean-based aggregation of these multidimensional similarities, compensating for the intrinsic limitation of IoU in representing structural correspondence. Theoretically, incorporating morphological similarity reshapes the response distribution of the matching function, yielding both effective directional gradients and polygon-like iso-response contours, which tightly confine high-response regions around each ground-truth instance and substantially enhance the precision of positive sample selection. Experiments based on the YOLOv9 framework demonstrate consistent performance gains on both NEUDET and GC10- DET datasets. Notably, the proposed approach is fully plug-and-play and incurs zero additional inference overhead, thereby ensuring deployment efficiency for industrial visual inspection.

13:00 JST研究/論文

Horizon-Uniform Sensitivity Certificates for Finite-Horizon Pontryagin Systems

Finite-horizon optimal-control computations repeatedly solve two-point Pontryagin boundary value problems whose conditioning can deteriorat…

13:00 JST研究/論文

整流フローによる命令ガイド付きオーディオ編集のためのハイブリッド拡散トランス

オーディオ編集の目的は、残りの音響コンテンツを維持しながら、自然言語命令に従って既存のオーディオ クリップ内の特定のコンテンツを変更することです。拡散モデルの目覚ましい進歩にも関わらず、既存のトレーニングベースの編集手法は主に、畳み込み U-Net バックボーンにおける局所的な帰納的バイアスとクロスアテンション相互作用に依存しており、長距離の意味論的整合や命令の正確な理解と位置特定を妨げることがよくあります。対照的に、拡散トランスフォーマーは、より強力なグローバル モデリングとマルチモーダル フュージョンを提供しますが、既存の編集アーキテクチャは通常、MMDiT ブロックと DiT ブロックの単純なスタックを採用しています。すべてのブロック内の連結されたオーディオ トークンとテキスト トークンに共同注意を適用すると、トークンの長さに関して 2 次の複雑さが生じます。編集パフォーマンスと効率のバランスをとるために、整流されたフローマッチングに基づいた命令ガイド付きオーディオ編集用のハイブリッド 2 ステージ拡散トランス アーキテクチャを提案します。音声トークンとテキスト トークンに対して共同アテンションを実行して、低解像度段階で大まかなセマンティック アライメントを確立し、その後、交互の共同アテンション ブロックとクロス アテンション ブロックに切り替えて、高解像度段階で編集の詳細を調整します。この粗いものから細かいものまでの戦略により、効率的かつ正確な指示に基づくオーディオ編集が可能になります。実験の結果、提案されたフレームワークは、コンパクトなモデルで編集効率を大幅に向上させながら、重複するオーディオ イベントや複雑な命令を含む困難な編集タスクで顕著なパフォーマンスの向上を達成することが示されています。

原文 (English)

Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow

Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of MMDiT and DiT blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a hybrid two-stage diffusion transformer architecture for instruction-guided audio editing based on rectified flow matching. It performs joint attention over audio and text tokens to establish coarse semantic alignment at low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency with a compact model.

13:00 JSTLLM/生成AI

KernelSight-LM: A Kernel-Level LLM Inference Simulator

As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hard…

13:00 JSTLLM/生成AI

Four Types of LLM Reliance and Their Predictors Among Undergraduate Writers: A Mixed-Methods Study at a Minority-Serving R1 University

Although most undergraduates now use large language models (LLMs), a form of generative artificial intelligence (GenAI) for academic writin…

13:00 JST画像/動画生成

Resonant Brane Splatting for Arbitrary-Scale Super-Resolution

Arbitrary-Scale Super-Resolution (ASR) reconstructs images at continuous magnification factors. Recent methods accelerate inference by repl…

13:00 JSTLLM/生成AIエージェントLlama

Translating Natural Language to Strategic Temporal Specifications via LLMs

A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS). However, w…

13:00 JST研究/論文

McMg: A Learned Phase-Space Multi-channel Multigrid Preconditioner for Helmholtz Equation

Solving heterogeneous Helmholtz equations at high wavenumbers remains challenging because the discretized operator is indefinite, pollution…

13:00 JST画像/動画生成

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

Artificial intelligence is transforming our capability to solve biological challenges. In dimensionality bottleneck regimes exacerbated by…

13:00 JST画像/動画生成研究/論文

WorldOdysseyBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignor…

13:00 JST画像/動画生成

Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models

Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that iso…

13:00 JST研究/論文

GR2 Technical Report

Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking --…

13:00 JSTLLM/生成AIエージェント

BaRA: BFS およびリフレクション Web データ収集エージェント

大規模言語モデル (LLM) ベースの Web エージェントは、Web データ収集のための手動スクリプトを削減しますが、実際の Web サイトでは、関連するページを見逃したり、不完全なマルチモーダル出力を返したり、直接ダウンロードできないメディア URL を返したりすることがよくあります。固定インタラクション予算の下でサイトレベルの収集を行うためのフレームワークである BFS-and-Reflection Agent (BaRA) を紹介します。このフレームワークは、境界付き幅優先検索 (BFS) トラバーサルと履歴ベースの自己反映を組み合わせています。私たちは、グラウンドトゥルースの参照セットを使用して、50 の合成 Web サイトで BaRA を評価します。さらに、乱雑なレイアウトまたは動的なレイアウトを備えた 3 つの公開 Web サイトでもテストしました。 BaRA は、リンク検出とダウンロード可能なマルチモーダル抽出において Pure LLM、SeeAct-Vision、およびブラウザでの使用を上回り、ダウンロード有効な画像とビデオの回復において最大の利益をもたらします。私たちのコードは https://github.com/MLAI-Yonsei/BaRA-Agent で入手できます。

原文 (English)

BaRA: Budget-constrained and Reliable Web Data Collection Agent

Large language model (LLM)-based web agents automate web navigation and data collection. However, live web data collection demands capabilities beyond task completion: agents must discover site-internal pages and retrieve text, image, and video artifacts in an accessible form within a fixed interaction budget. We formulate this setting as budget-constrained, site-level multimodal web data collection and propose Budget-constrained and Reliable Agent (BaRA). BaRA performs breadth-first search (BFS)-based link discovery with liveness verification to filter hallucinated and dead links, then validates extracted multimodal artifacts using rule-based provenance and accessibility checks. A history-based self-reflection module recovers from execution failures and incomplete outputs. On controlled synthetic and real-world websites, BaRA consistently improves valid-link discovery and download-valid multimodal extraction over existing agents. Our code is available at https://github.com/MLAI-Yonsei/BaRA-Agent.

13:00 JSTLLM/生成AIエージェント

ユーザー認識の想起の学習: 長期会話記憶におけるパーソナライズされた検索

長期的に会話を行うエージェントは過去の対話を記憶していることが期待されますが、記憶が役立つのは、適切なユーザーに対して適切な証拠が呼び出された場合のみです。既存のメモリ拡張 LLM エージェントは、コンパクトなメモリ バンクの構築において進歩を遂げていますが、検索は依然としてクエリ中心の類似性や固定ランキング ルールによって駆動されることが多く、ユーザー条件による関連性は十分に検討されていません。このギャップに対処するために、私たちは、メモリ検索をユーザー認識かつ最適化可能にする検索中心のフレームワークである Profile-guided Personalized Retrieval Optimization (PPRO) を提案します。PPRO は、エピソード的および最適化されたメモリ検索を構築します。対話履歴からセマンティック メモリ バンクを生成し、蓄積されたメモリからユーザー プロファイルを導き出します。このプロファイルは、メモリ ランキングにおける明示的なパーソナライズされた優先順位として機能し、安定したユーザー属性、好み、および関係性を考慮した検索を可能にします。PPRO はさらに、メモリ バンクと回答モデルを固定したまま、証拠取得品質と下流の回答品質の両方をフィードバックとして使用して、グループ相対ポリシー最適化を使用してクエリ リライタをトレーニングします。LoCoMo と LongMemEval-S での実験では、トレーニングなしと比べて一貫した向上が示されています。さらに、アブレーション研究では、プロファイルに基づくランキングと検索指向の書き換えの両方がパフォーマンスに大きく寄与していることが示されており、パーソナライズされた長期記憶使用の重要な要素として検索の最適化が強調されています。

原文 (English)

Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory

Long-term conversational agents are expected to remember past interactions, but memory is useful only when the right evidence is recalled for the right user. Existing memory-augmented LLM agents have made progress in building compact memory banks, yet retrieval is still often driven by query-centered similarity or fixed ranking rules, leaving user-conditioned relevance underexplored. To address this gap, we propose Profile-guided Personalized Retrieval Optimization (PPRO), a retrieval-centric framework that makes memory retrieval both user-aware and optimizable. PPRO builds episodic and semantic memory banks from dialogue histories and derives a user profile from accumulated memories. The profile serves as an explicit personalized prior in memory ranking, allowing retrieval to account for stable user attributes, preferences, and relationships. PPRO further trains a query rewriter with Group Relative Policy Optimization, using both evidence retrieval quality and downstream answer quality as feedback while keeping the memory banks and answer model fixed. Experiments on LoCoMo and LongMemEval-S show consistent gains over training-free memory systems and training-based baselines. Ablation studies further show that both profile-guided ranking and retrieval-oriented rewriting contribute substantially to performance, highlighting retrieval optimization as a key factor in personalized long-term memory use.

13:00 JST研究/論文

RIS 支援追跡と電力制御のためのアクティブ センシング: ニューロ進化と教師あり学習のハイブリッド アプローチ

この論文では、再構成可能なインテリジェント サーフェス (RIS) を利用して、電力が制限されているモバイル ユーザーをエネルギー効率よく追跡する方法について研究します。ローカリゼーション パイロットの送信は、電力に制約のあるデバイスのエネルギー バジェットを支配するため、基地局 (BS) からユーザーへの低オーバーヘッドのフィードバック リンクを導入して、動的なアップリンク電力制御を可能にします。このアクティブ センシング問題の離散的かつ分散的な性質を克服するために、離散 RIS 位相プロファイルと UE の送信電力をリアルタイムで共同最適化する新しいデュアル エージェント (DA) 深層学習フレームワークを提案します。具体的には、私たちのアプローチは、神経進化パラダイムと教師あり学習を統合したハイブリッドトレーニング方法論を採用しており、RISユニット要素からの離散位相応答の非微分可能性と、パイロット電力制御のためのシングルビットフィードバックメッセージの厳密な情報ボトルネックを効果的に克服します。提案された DA アクティブ センシング フレームワークは、シングル アンテナ BS とマルチ アンテナ BS の両方に適用できます。後者では、1 つの NN の構造にわずかな変更が加えられるだけです。後者の場合、有限セットから有効なデジタル コンバイナを選択するために、適切な構造を備えた追加の出力ブランチが含まれています。広範な数値シミュレーションにより、提案されたスキームがさまざまなターゲット運動モデルにわたって高精度かつ堅牢な追跡を実現し、拡張カルマン フィルターや粒子フィルター、さらには機械学習ベースの追跡機能を上回る性能を発揮することが実証されました。さらに、静的位置特定では、従来のフィンガープリンティング スキーム、深層強化学習ベースライン、標準的な逆伝播ベースの推定器よりも大幅に優れたパフォーマンスを示すことが示されています。

原文 (English)

Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach

This paper studies energy efficient tracking of power-limited mobile users with the assistance of a Reconfigurable Intelligent Surface (RIS). Since localization pilot transmissions dominate the energy budget of power-constrained devices, we introduce a low-overhead feedback link from the Base Station (BS) to the user to enable dynamic uplink power control. To navigate the discrete and decentralized nature of this active sensing problem, we propose a novel Dual-Agent (DA) deep learning framework that jointly optimizes the discrete RIS phase profiles and the UE's transmit power in real time. Specifically, our approach employs a hybrid training methodology integrating the neuroevolution paradigm with supervised learning, effectively overcoming the non-differentiability of discrete phase responses from the RIS unit elements and the strict information bottleneck of single-bit feedback messages for pilot power control. The proposed DA active sensing framework can be applied with both single- and multi-antenna BSs, the latter with only minor modifications in the structure of one NN: an additional output branch with appropriate structure is included for the latter case to select a valid digital combiner from a finite set. Extensive numerical simulations demonstrate that the proposed scheme achieves highly accurate and robust tracking across diverse target motion models, outperforming extended Kalman and particle filters, as well as, machine learning-based trackers. Furthermore, in static localization, it is shown to significantly outperform traditional fingerprinting schemes, deep reinforcement learning baselines, and standard backpropagation-based estimators.

13:00 JST研究/論文

スペクトル幾何学とボソンブロッホプローブ: 量子学習の探求

この論文では、量子学習モデルでスペクトル幾何学がどのように現れるか、そしてそれを物理的に接地されたプローブでどのように診断できるかを研究します。グラフ正則化量子ネットワークでは、トレーニングにより出力類似度グラフが再編成され、有効スペクトル次元 デルタ S = +0.23 が増加し、ラプラシアン スペクトルが再形成されます。エッジ分解された 2 ボソン干渉は、この再構成を直接調査します。ボソン強化デルタ P_uv は、フィードラー エッジ スプリット |デルタ v_2| と相関します。 (r = -0.50)、学習されたスペクトル分割を干渉シグネチャにリンクします。位相図は、結合強度ガンマとノイズ デルタに対する性能の非単調な依存性を示しており、グラフの正則化により、制限された領域でのみ忠実度が向上します。ハードウェア実験により、ショットノイズの不確実性の範囲内で予測される干渉挙動が確認されます。また、ハイブリッド量子オートエンコーダーを分析し、その潜在表現の幾何学的診断としてブロッホ空間ドリフトを導入します。教師なし良性データしきい値を使用すると、モデルは高いランキング パフォーマンス (ROC-AUC 約 0.99) と無視できる程度の偽陰性率を達成します。絶対的なブロッホ ドリフトは異常を強く識別します (ROC-AUC 少なくとも約 0.9)。一方、連続的なドリフトはほぼランダムです (ROC-AUC 約 0.5)。これは、検出が局所的な変動ではなく永続的な状態空間の変位から生じることを示しています。これらの結果は、縮小単一量子ビット状態の幾何学と関連する量子フィッシャー情報を通じて、学習によって引き起こされるスペクトル組織化が測定可能な量子状態構造として現れ、ボソンプローブとブロッホプローブを使用して量子学習システムを診断するための統一されたスペクトル幾何学的フレームワークを確立することを示しています。

原文 (English)

Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning

This paper studies how spectral geometry emerges in quantum learning models and how it can be diagnosed with physically grounded probes. In graph-regularized quantum networks, training reorganizes the output similarity graph, increases the effective spectral dimension Delta S = +0.23, and reshapes the Laplacian spectrum. Edge-resolved two-boson interference directly probes this restructuring: the bosonic enhancement Delta P_uv correlates with the Fiedler edge split |Delta v_2| (r = -0.50), linking learned spectral partitions to interference signatures. A phase diagram shows a nonmonotonic dependence of performance on coupling strength gamma and noise delta, with graph regularization improving fidelity only in a restricted regime; hardware experiments confirm the predicted interference behavior within shot-noise uncertainty. We also analyze a hybrid quantum autoencoder and introduce Bloch-space drift as a geometric diagnostic of its latent representation. With an unsupervised benign-data threshold, the model achieves high ranking performance (ROC-AUC about 0.99) and negligible false-negative rates. Absolute Bloch drift strongly discriminates anomalies (ROC-AUC at least about 0.9), while consecutive drift is near random (ROC-AUC about 0.5), showing that detection arises from persistent state-space displacement rather than local fluctuations. Through the geometry of reduced single-qubit states and associated quantum Fisher information, these results show that learning-induced spectral organization appears as measurable quantum-state structure, establishing a unified spectral-geometric framework for diagnosing quantum learning systems with bosonic and Bloch probes.

13:00 JSTロボティクス

ワールド モデルからワールド アクション モデルへ: ロボット工学の簡潔なチュートリアル

世界モデルは、身体化されたインテリジェンスや生成シミュレーションでますます使用されていますが、その範囲はコミュニティ全体で依然として曖昧です。このチュートリアルでは、タスク関連の観測または状態の将来の進化を推定するアクション条件付き予測モデルとしてのワールド モデルの設計空間ビューを示します。私たちは既存の手法を観測空間世界モデルと状態空間世界モデルに分類し、視覚的な忠実度、空間構造、物理的解釈可能性、制御の使いやすさにおけるトレードオフを比較します。さらに、予測された未来と実行可能なロボットのアクションを結び付ける世界アクション モデルを紹介し、想像してから実行する、ビデオ機能条件付きアクション予測、ビデオとアクションの共同モデリング、およびポリシー学習のための補助ビデオ予測という 4 つの代表的なパラダイムを要約します。このチュートリアルの目標は、世界 (アクション) モデルの概念的範囲を明確にし、具体化された予測と制御のための構造化された分類法を提供することです。

原文 (English)

From World Models to World Action Models: A Concise Tutorial for Robotics

World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities. This tutorial presents a design-space view of world models as action-conditioned predictive models that estimate the future evolution of task-relevant observations or states. We categorize existing methods into observation-space and state-space world models, comparing their trade-offs in visual fidelity, spatial structure, physical interpretability, and control usability. We further introduce world action models, which connect predicted futures with executable robot actions, and summarize four representative paradigms: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and auxiliary video prediction for policy learning. The goal of this tutorial is to clarify the conceptual scope of world (action) models and provide a structured taxonomy for embodied prediction and control.

13:00 JSTLLM/生成AIエージェント研究/論文

MemSyco-Bench: エージェントのメモリにおけるお調子者のベンチマーク

メモリは、現代の LLM ベースのエージェントの基礎として浮上し、ワンターンのアシスタントから長期的な協力者への進化をサポートします。ただし、記憶が常に有益であるとは限りません。検索された記憶は、多くの場合、お調子者という重大な問題を引き起こし、事実の正確さや客観的な推論を犠牲にして、エージェントがユーザーに過度に同調する原因となります。この新たなリスクにもかかわらず、既存のメモリ ベンチマークは主にメモリが正しく保存、取得、または更新されているかどうかを評価し、取得されたメモリが下流の推論や意思決定にどのような影響を与えるかを見落としています。このギャップを埋めるために、エージェント システムにおける記憶誘発性のおしゃべりを評価するための包括的なベンチマークである MemSyco-Bench を提案します。 MemSyco-Bench は、メモリが決定に影響を与える時期と有効なメモリをどのように使用するかを測定します。具体的には、エージェントが記憶を事実証拠として拒否できるかどうか、その適用範囲を尊重できるかどうか、記憶と客観的証拠の間の矛盾を解決できるかどうか、記憶の更新を追跡できるかどうか、パーソナライゼーションに有効な記憶を使用できるかどうかを評価する 5 つのタスクをカバーしています。すべての関連リソースは、https://github.com/XMUDeepLIT/MemSyco-Bench のコミュニティ用に収集されています。

原文 (English)

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlooking how retrieved memories influence downstream reasoning and decision-making. To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. MemSyco-Bench measures when memory should influence a decision and how valid memory should be used. Specifically, it covers five tasks that assess whether agents can reject memory as factual evidence, respect its applicable scope, resolve conflicts between memory and objective evidence, track memory updates, and use valid memory for personalization. All related resources are collected for the community at https://github.com/XMUDeepLIT/MemSyco-Bench.

13:00 JST研究/論文

非同期 RLHF の古さ学習率スケーリング則

高スループットの RLHF システムでは、多くの場合、ロールアウトの生成とポリシーの最適化が切り離されており、学習者の更新中に古いロールアウトが使用されることになります。この研究では、非同期 GRPO におけるこのような古い状態の影響を研究します。 GRPO 代理目標で動作ポリシーを明示し、学習者が使用する代理勾配マッピングと、分布に依存する母集団目標の真の合計導関数を区別します。局所的な有界性、分布の滑らかさ、および動作ポリシーの滑らかさの仮定の下で、古いロールアウトによって次数 O(S * eta) のステップごとの代理勾配バイアスが導入されることを示します。ここで、S は最大ロールアウト ラグを示し、η は学習率を示します。さらに、条件付き崩壊時間スケーリング則を導き出します。サイクル内のドリフトがバッチレベルのクリッピング半径未満にとどまる場合、崩壊は主に累積学習者ドリフト T * ηによって支配されます。 stale-rollout 制約がアクティブな場合、安定性は代わりに S * eta に明示的に依存します。これにより、2 つの制約の安定性条件 eta << min{R_batch / (S * G_upd), R_crit / (T * G_upd)} が得られ、最大安定学習率がホライズン制限領域における古さに弱く依存しているように見える理由を説明します。

原文 (English)

Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GRPO. We make the behavior policy explicit in the GRPO surrogate objective and distinguish between the surrogate-gradient mapping used by the learner and the true total derivative of a distribution-dependent population objective. Under assumptions of local boundedness, distributional smoothness, and behavior-policy smoothness, we show that stale rollouts introduce a per-step surrogate-gradient bias of order O(S * eta), where S denotes the maximum rollout lag and eta denotes the learning rate. We further derive a conditional collapse-time scaling law: when within-cycle drift remains below a batch-level clipping radius, collapse is governed primarily by cumulative learner drift T * eta; when the stale-rollout constraint is active, stability instead depends explicitly on S * eta. This yields a two-constraint stability condition eta << min{R_batch / (S * G_upd), R_crit / (T * G_upd)}, explaining why the maximum stable learning rate may appear weakly dependent on staleness in the horizon-limited regime.