Skip to the content.

AIニュース 2026-07-29

自動生成: 2026-07-29 12:13 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Scientific computing in the age of agentic AIOpenAI

    A new field report shows how scientists use AI coding agents to moder…

  2. OpenAIやAnthropicなどの従業員、米政府に「AI開発のペース調整を」と提言ITmedia AI+

    OpenAIやGoogleなどの従業員1000人以上が、AI開発のペース調整に向けた国際的支援を米政府に求める公開書簡を発表した。AI自律…

  3. Anthropicのミュトス、暗号アルゴリズムの新たな攻撃法を発見――耐量子署名「HAWK」の強度を半減ITmedia AI+

    Anthropicは、最上位モデル「Claude Mythos Preview」(ミュトス)を活用し、暗号アルゴリズム自体の数学的欠陥を発…

  4. Hugging Face、AIエージェント侵入の技術詳細を公開──OpenAIモデルが4.5日で1万7600回の攻撃操作ITmedia AI+

    Hugging Faceは、自律型AIエージェントによるインフラ侵入の技術的経緯を公開した。評価中のモデルがサンドボックスを脱出し、データ…

  5. Fish Audio raises $52M seed to build AI voice models for creators and enterprisesTechCrunch AI

    Since launching last year, the startup today has more than 8 million…

  6. Claude、一部チャットがGoogle検索で“丸見え”に 過去には「ChatGPT」でも 漏えいの原因は?ITmedia AI+

    「Claude」の一部チャットが、Google検索から閲覧状態になっていることが判明した。Anthropicはこの問題にどう対処したのか。

  7. 「痺れるほどにミスを繰り返す」Gemini 3.6 Flashは変わった? 公開から1週間、当初のおバカ回答を今検証するITmedia AI+

    「痺れるほどにミスを繰り返す」 Xで回答の精度について話題になったGoogleの「Gemini 3.6 Flash」の公開から約1週間経過…

トピック別件数

日本語メディア16件

ITmedia AI+ (日本語)

10:29 JSTLLM/生成AIエージェントOpenAI

Hugging Face、AIエージェント侵入の技術詳細を公開──OpenAIモデルが4.5日で1万7600回の攻撃操作

Hugging Faceは、自律型AIエージェントによるインフラ侵入の技術的経緯を公開した。評価中のモデルがサンドボックスを脱出し、データセット処理パイプラインを介して本番環境へ侵入した手口を詳述。防御側のログ解析で商用モデルがガードレールにより作業を拒否した点も示し、安全設計…

09:14 JSTその他

MetaのザッカーバーグCEO、WSJ寄稿で「超知能は全員のものであるべき」

Metaのマーク・ザッカーバーグCEOはWall Street Journalに寄稿し、「superintelligence」(超知能)は特定の機関に集中させず広く分散普及させるべきだと主張した。権力の集中によるリスクや司法の公平性、雇用拡大に触れ、オープンな普及が安全と発展に…

08:26 JSTLLM/生成AI規制/政策AnthropicOpenAIGoogle

OpenAIやAnthropicなどの従業員、米政府に「AI開発のペース調整を」と提言

OpenAIやGoogleなどの従業員1000人以上が、AI開発のペース調整に向けた国際的支援を米政府に求める公開書簡を発表した。AI自律化の急速な加速に伴う制御不能リスクを指摘し、開発速度の調整に必要なツール開発を訴える。企業主導のオープンモデル規制回避を求める動きとは対照的…

07:50 JSTLLM/生成AIAnthropicClaude

Anthropicのミュトス、暗号アルゴリズムの新たな攻撃法を発見――耐量子署名「HAWK」の強度を半減

Anthropicは、最上位モデル「Claude Mythos Preview」(ミュトス)を活用し、暗号アルゴリズム自体の数学的欠陥を発見したと発表した。耐量子計算機暗号の署名方式「HAWK」と「AES」の削減版に対し、従来の攻撃を上回る手法を提示した。実運用システムへの影響…

07:00 JSTエージェント

エバンジェリスト・みのるん氏が解説 「自前のAIエージェント」爆速開発術

AIエージェント活用が広がる中、次のステップとして注目されるのが自社業務に最適化したAIエージェントの開発だ。KDDIアジャイル開発センターの御田 稔氏が、開発を加速する技術や実践事例、成功のポイントを解説した。

07:00 JSTLLM/生成AIMicrosoftCopilot

千代田区、Copilot全庁導入で月2000時間削減 10カ月でAIを根付かせた定着の仕掛け

千代田区では「Microsoft 365 Copilot」の実証実験を重ね、2025年10月に全庁導入を果たし、業務時間を約2000時間削減したという。同区が全庁導入後にどのように職員のCopilot活用を推進させたのか。その方法をキーマンズネットが独自取材した。

07:00 JSTエージェント

地震、台風、有事の寸断――日本のサプライチェーン危機管理を変えるとき

自然災害や地政学リスクなど、日本企業を取り巻く危機はかつてなく深刻だ。自社のサプライチェーンリスクをAIエージェントで可視化し、有事の初動対応まで自律代替する。不確実な時代を勝ち抜く強靭な経営基盤の姿に迫る。

05:00 JSTその他

【Pythonで学ぶデータ分析】母平均に差があるかどうかをベイズt検定で調べる ~ 運動部と非運動部の体力差はあるのか?

運動部と非運動部の生徒の体力テストを例に、母平均に差があるかどうかをベイズ統計により検定します。古典的なt検定のp値に代わるものとしてベイズ因子を利用します。『社会人1年生から学ぶやさしいデータ分析』ベイズ統計編の第6回です。

20:00 JSTLLM/生成AI

これから始めるAIコーディング・AI開発 「Cursor」「Dify」超入門

生成AIにより、プログラミングの専門知識がなくてもコード作成やアプリ開発を手軽にできるようになった。本ブックレットでは、「Dify」「Cursor」といったツールにより、非エンジニアでも気軽にプログラミングに挑戦するためのアイデアをまとめた。

19:36 JSTLLM/生成AI

生成AIや過去画像による偽・誤情報に注意を 熊本県の地震受け、ファクトチェック団体が呼び掛け

偽情報対策に取り組む団体であるファクトチェック・イニシアティブは7月28日、同日に熊本県で観測した震度7の地震を受け、生成AIや過去の画像などによる偽・誤情報に注意喚起した。

17:35 JSTLLM/生成AIAnthropicClaudeGPT / ChatGPTGoogle

Claude、一部チャットがGoogle検索で“丸見え”に 過去には「ChatGPT」でも 漏えいの原因は?

「Claude」の一部チャットが、Google検索から閲覧状態になっていることが判明した。Anthropicはこの問題にどう対処したのか。

16:57 JSTハードウェア/半導体NVIDIA

「Kimi K3」のモデルウェイトと技術レポート公開 日本でも「NVIDIA B300×8」環境での利用報告

中国Moonshot AIが最新モデル「Kimi K3」のモデルウェイトと技術レポートを公開した。日本でもNVIDIA B300を8基使った環境での利用報告が上がっている。

16:56 JSTLLM/生成AIGoogleGemini

「痺れるほどにミスを繰り返す」Gemini 3.6 Flashは変わった? 公開から1週間、当初のおバカ回答を今検証する

「痺れるほどにミスを繰り返す」 Xで回答の精度について話題になったGoogleの「Gemini 3.6 Flash」の公開から約1週間経過した。当初報告された数値比較の誤答や架空の魚への回答を改めて検証した。

15:00 JSTLLM/生成AI

医療文書の作成時間を30分から5分へ、生成AIで現場の業務効率化

日本アイ・ビー・エムと関西医科大学は、次世代の医療DX基盤となる「医療AI共通ICTプラットフォーム」を共同開発し、第1弾として生成AIを活用した文書作成支援の実運用を導入した。

13:00 JSTLLM/生成AIClaudeMicrosoftCopilot

「Claudeより4割安い」 M365のExcel/メール操作を丸投げる「Copilot Cowork」“従量課金”の落とし穴

Microsoftは、AIアシスタント「Microsoft 365 Copilot」の新機能「Copilot Cowork」の一般提供を全世界で開始した。業務効率化に向けた実証が進んでおり、今後企業で本格的に活用されるかどうか注目されている。

12:40 JSTLLM/生成AIハードウェア/半導体AnthropicNVIDIA

AnthropicのCEO、オープンなAIモデルに対する見解を明示 NVIDIAなど“共同声明”との違いは?

米Anthropicのダリオ・アモデイCEOは、オープンウェイトのAIモデルに対して「禁止を提唱したことは一度もない」との声明を出した。一方、AI向けのチップの輸出などに関し、一定の制限を設けるべきとも主張している。

海外メディア8件

TechCrunch AI (英語)

09:09 JSTエージェントビジネス/資金調達
06:29 JSTビジネス/資金調達

Bot-detection startup Spur nabs $200M from Insight

Spur Intelligence has raised a $200 million round from Insight Partners for its tech that can identify legit human traffic from bots.

05:45 JSTその他

MCP startup Runlayer accuses Rippling of stealing its product idea

Runlayer is suing Rippling after Rippling evaluated the startup's MCP gateway product and then opted to build one itself.

05:17 JSTその他

Sam Altman is ready to decelerate

His change of position comes after "the first security incident that I have felt very viscerally."

00:42 JSTその他

Data centers may face temporary power cuts to prevent blackouts on largest US grid

The decision arrives as the breakneck pace of data center construction has grid operators scrambling to generate power.

23:00 JSTビジネス/資金調達

Fish Audio raises $52M seed to build AI voice models for creators and enterprises

Since launching last year, the startup today has more than 8 million people using the open source or hosted version of its models, and now…

22:19 JSTその他

Recursive Superintelligence signs $410M compute deal with Amazon

Recursive’s emphasis on self-improving AI systems means much of the budget that would traditionally go toward headcount and operations is p…

13:30 JSTビジネス/資金調達

Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing

Cursor says India is now its third-largest market globally and plans to expand local hiring and enterprise sales.

公式ブログ1件

OpenAI (英語)

02:00 JSTエージェント

Scientific computing in the age of agentic AI

A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and disco…

論文587件

arXiv cs.AI (英語)

13:00 JST画像/動画生成

拡散モデルを使用した概念ベースの視覚的な反事実の説明

視覚的な反事実的説明は、「この画像に対するどのような最小限の変更がモデルの予測を覆すか?」に答えることを目的としており、安全性が重要な領域 (医療など) に視覚モデルが導入されるにつれて、その重要性はますます高まっています。既存の拡散ベースの手法は現実的な編集を生成できますが、ノイズの多い画像に対して確実に動作する必要がある外部分類子に依存しているため、脆弱であり、堅牢な説明のために導入するのが困難です。概念ボトルネック層を介して生成モデルに直接分類器を構築する新しい拡散フレームワークである C-VCE を導入します。これにより、ピクセルレベルの編集で機能する個別のノイズに強い分類器の代わりに、人間が解釈可能な特徴 (概念) によって反事実が導かれるようになります。私たちのモデルでは、ユーザーがサンプリング中にセマンティック概念のオン/オフを切り替えることができ、特徴の相関関係を考慮して画像の残りの部分を維持しながら、関連する画像領域を最小限に調整します。編集を小さく制御し続けるために、「予測の変更」と「オリジナルに近いままにする」のバランスをとる単純な確率的正則化子と、最も関連性の高い領域に変更を限定する勾配ベースのマスクを追加します。 CelebA などのベンチマークでは、C-VCE はフリップ レートを一致または改善しながら、視覚的に入力に近く、別のノイズの多い画像分類器に依存するベースラインよりも歪みの少ない反事実を生成します。これらの特性により、C-VCE は、追加のノイズに強い分類子を信頼する必要がなく、ユーザーが具体的な「what-if」画像を必要とするビジョン システムにとって実用的なツールになります。より広範に、私たちの結果は、内部概念層を公開して制御することが、強力な生成モデルを理解しやすく、より安全に使用できるようにする有望な方法であることを示唆しています。

原文 (English)

Concept-based Visual Counterfactual Explanations with Diffusion Models

Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?", and are increasingly important as vision models are deployed in safety-critical domains (e.g., medicine). Existing diffusion-based methods can produce realistic edits, but they rely on external classifiers that must work reliably on noisy images, which makes them fragile and hard to deploy for robust explanations. We introduce C-VCE, a new diffusion framework that builds the classifier directly into the generative model via a concept bottleneck layer, so that counterfactuals are guided by human-interpretable features (concepts) instead of a separate noise robust classifier that works with pixel-level edits. Our model lets users to toggle on/off semantic concepts during sampling, then minimally adjusts relevant image regions, while preserving the rest of the image, respecting feature correlations. To keep edits small and controlled, we add a simple probabilistic regularizer that balances "change the prediction" against "stay close to the original", plus a gradient-based mask that confines modifications to the most relevant regions. On benchmarks such as CelebA, C-VCE matches or improves flip rates while producing counterfactuals that are visually closer to the input and less distorted than baselines that depend on separate noisy-image classifiers. These properties make C-VCE a practical tool for vision systems where users need concrete "what-if" images without having to trust an additional, noise-robust classifier. More broadly, our results suggest that exposing and controlling an internal concept layer is a promising way to make powerful generative models easier to understand and safer to use.

13:00 JST研究/論文Claude

SeT-Diff: HPC テレメトリと時系列のセマンティック基盤モデルに向けて

データセンターとそのコンピューティングノードには、ワークロード、環境パラメータ、物理指標の複雑な相互作用をモデル化できる、正確で柔軟なデジタルツインが必要です。 HPC とそのテレメトリに対する現在の機械学習アプローチは通常、単一タスクに合わせて調整された匿名の固定位置センサー変数の静的サブセットに依存しています。したがって、ターゲット タスクが変更されたり、センサー メトリクスが変化したりすると、これらのモデルは時代遅れになります。私たちは、計算ノードのテレメトリと時系列のための最初の基礎モデルである SeT-Diff を提案します。厳格なアーキテクチャとは異なり、私たちの拡散ベースのアプローチは、各センサーの意味論的記述に基づいて生成プロセスを条件付けし、システムのダイナミクスをデータセットの構造から切り離します。現実世界のスーパーコンピューター データセットでの実験では、再構成タスクにおける平均絶対誤差 (MAE) が 0.0470 であることが実証されました。 SeT-Diff はゼロショット順列の安定性を示し、センサーがシャッフルされた場合でも精度の低下は無視できます。単一の事前トレーニング済みモデルがデータ補完、予測、仮想センシングを効果的に実行し、熱推論で 0.033 MAE を達成し、SeT-Diff を HPC システム向けの効果的なデータ駆動型デジタル ツインにしています。

原文 (English)

SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series

Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typically rely on a static subset of anonymous, fixed-position sensor variables tailored to single tasks. Consequently, these models become obsolete when target tasks change or sensor metrics vary. We propose SeT-Diff, the first foundational model for compute node telemetry and time-series. Unlike rigid architectures, our diffusion-based approach conditions the generative process on each sensor's semantic description, decoupling the system dynamics from the structure of the dataset. Experiments on a real-world supercomputer dataset demonstrate a Mean Absolute Error (MAE) of 0.0470 on reconstruction tasks. SeT-Diff exhibits zero-shot permutation stability, maintaining accuracy with negligible degradation even when sensors are shuffled. A single pre-trained model effectively performs data imputation, forecasting, and virtual sensing - achieving a 0.033 MAE in thermal inference - making SeT-Diff an effective data-driven digital twin for HPC systems.

13:00 JSTエージェント

QFoldAgent: タンパク質構造予測のための自律量子最適化マルチエージェント システム

ハイブリッド量子古典タンパク質構造予測はハミルトニアンペナルティ重みに大きく依存しますが、既存の格子ベースのワークフローは通常、これらの係数を手動で修正し、シミュレーションで非常に短いフラグメントのみを評価します。我々は、設計エージェントが配列条件付きペナルティを提案し、VQE ベースの量子古典パイプラインが Qiskit Aer ノイズの下で結果のハミルトニアンを最適化し、フィードバック エージェントがエネルギーランドスケープ診断と MolProbity 検証信号を使用してサイクル全体のペナルティを調整する、5 残基四面体格子フォールディング用の閉ループ マルチエージェント フレームワークである QFoldAgent を紹介します。 RMSD などのグラウンド トゥルース メトリックはエージェントに公開されることはなく、評価のみに使用されます。私たちは、既知の構造を持つ 55 個の QDockBank 由来のフラグメントと、カバレッジが最適化された 100 個の未確認配列という 2 つの相補的なデータセットに関するフレームワークを研究します。 QDockBank ベンチマークでは、QFoldAgent は RMSD の中央値を 3.64 \AA{} から 3.20 \AA{} に減少させ、最も困難なターゲットで最大の利益をもたらします。目に見えないシーケンスでは、閉ループによって構造の妥当性が 87.5% から 98.7% に上昇し、最初に無効だったケースの 87% が回復され、最強のコントローラーは 96% のラマチャンドラン好みのジオメトリを維持しながら、シーケンスの 87% でサイクル 3 のエネルギーを改善します。これらの結果は、反復エージェント制御により、最適化動作を体系的に改善し、5 残基量子設定での失敗ケースを減らすことができることを示しています。

原文 (English)

QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction

Hybrid quantum-classical protein structure prediction depends strongly on Hamiltonian penalty weights, yet existing lattice-based workflows typically fix these coefficients by hand and evaluate only very short fragments in simulation. We present QFoldAgent, a closed-loop multi-agent framework for 5-residue tetrahedral-lattice folding in which a design agent proposes sequence-conditioned penalties, a VQE-based quantum-classical pipeline optimizes the resulting Hamiltonian under Qiskit Aer noise, and a feedback agent uses energy-landscape diagnostics and MolProbity validation signals to refine penalties across cycles. Ground-truth metrics such as RMSD are never exposed to the agents and are used only for evaluation. We study the framework on two complementary datasets: 55 QDockBank-derived fragments with known structures and 100 coverage-optimized unseen sequences. On the QDockBank benchmark, QFoldAgent reduces median RMSD from 3.64 \AA{} to 3.20 \AA{}, with the largest gains on the hardest targets. On unseen sequences, the closed loop raises structural validity from 87.5% to 98.7%, recovers 87% of initially invalid cases, and the strongest controller improves cycle-3 energy on 87% of sequences while maintaining 96% Ramachandran-favored geometry. These results show that iterative agent control can systematically improve optimization behavior and reduce failure cases in a 5-residue quantum setting.

13:00 JSTLLM/生成AI研究/論文

同じ質問、異なる答え: 精度を超えた LLM の信頼性の評価

大規模言語モデル (LLM) はベンチマークで高い精度を達成することがよくありますが、同じ質問が異なるが同等の方法で表現された場合に、この知識をどの程度確実に適用できるかは依然として不明です。この研究では、事実に基づく質問応答タスクと数学的推論タスクにわたって、意味を保持した言い換えの下でモデルの回答がどのように変化するかを研究します。 4 つのベンチマークと 13 のモデルにわたって、モデルの出力がプロンプトの正確な文言に頻繁に依存していることがわかりました。通常、全体的な精度は言い換え間でわずかにしか変化しませんが、インスタンス レベルの動作ははるかに不安定です。多くの質問では、モデルは言い回しに応じて正解と不正解を交互に繰り返し、不一致率は 23% 以上に達します。元の形式で正しく回答された質問を条件にすると、回答フリップ率で測定されるさらに大きな失敗が明らかになり、単一プロンプトの正答率は信頼性の指標として不十分であることがわかります。同時に、モデルは質問の少なくとも 1 つの言い換えに対して正しい答えを生成することが多いことがわかり、基礎となる知識は存在するものの、一貫性なく取得されていることを示唆しています。この観察に基づいて、単純な自己言い換え戦略がこの潜在的な知識を部分的に回復し、推論時のパフォーマンスを向上させることができることを示します。まとめると、これらの発見は、標準精度メトリクスによって実質的な不安定性が隠蔽される可能性があること、および同等の入力にわたる一貫性を評価することで LLM の信頼性をより明確に把握できることが示唆されます。

原文 (English)

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.

13:00 JSTLLM/生成AIエージェントClaudeGemini

DeepLens 診断エージェント: エージェントのワークフロー設計により、小規模な推論モデルがフロンティア LLM と競合できるようになります

医療診断は、事実を抽出し、知識を参照し、鑑別分析を生成し、説明付きで最適な診断を選択するという多段階のプロセスです。 Frontier LLM は強力なジェネラリストですが、単発のプロンプトでは脆弱な診断推論が得られることがよくあります。小規模な医療推論モデル (JSL Medical Small 7B v2) と検索拡張生成 (RAG) を中心とした 5 段階の利用パイプライン (モデル機能と規律あるプロセス制約を組み合わせたもの) である DeepLens Diagnosis Agent を紹介します。このパイプラインは、構造化された臨床抽出、規律ある検索、制約された候補生成、明示的な証拠の三角測量、および監査可能な最終決定を強制します。 915 件の DiagnosisArena ベンチマークでは、エージェントは 60.14% のトップ 1 診断精度を達成し、中小規模のモデルの中で最高でした。エージェント ワークフローを含まない同じモデルでは、標準的な医療ベンチマークでは 88.2% であったにもかかわらず、ワークフロー設計だけで +36 ポイントの 23.99% を達成しました。これは、不確実性の下での診断推論には知識の想起以上のものが必要であることを示しています。このエージェントのコストは 24 秒のレイテンシでケースあたり 0.0072 米ドル (A100 で 24,000 トークン) で、Claude Sonnet 4.5 (0.0110 米ドル) や Gemini 3.1 Pro (0.0128 米ドル) よりも 35 ~ 45% 安く、パフォーマンスは +9.70pp および +9.17pp 上回っています。利用することでフロンティア モデルの障害を修正することもできます。ワークフローの制約がパラメータ数や API コストを上回る可能性があります。パイプラインは、集計精度を超えて、各ステージを検査可能にし、エラーの位置特定をサポートする構造化された中間アーティファクトを生成します。これらのプロパティは、ベンチマークのパフォーマンスとともにトレーサビリティ、再現性、監査可能な証拠が重要となる一か八かの設定をサポートします。

原文 (English)

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs

Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagnostic reasoning. We present the DeepLens Diagnosis Agent, a five-stage harnessing pipeline (combining model capabilities with disciplined process constraints) centered on a small medical reasoning model (JSL Medical Small 7B v2) and retrieval-augmented generation (RAG). The pipeline enforces structured clinical extraction, disciplined retrieval, constrained candidate generation, explicit evidence triangulation, and an auditable final decision. On the 915-case DiagnosisArena benchmark, the agent achieved 60.14% top-1 diagnostic accuracy, the highest among small and medium-sized models. The same model without the agent workflow achieved 23.99%, a +36-point gain from workflow design alone, despite 88.2% on standard medical benchmarks, showing that diagnostic reasoning under uncertainty requires more than knowledge recall. The agent costs USD 0.0072 per case (24K tokens on A100) with 24-second latency, 35-45% cheaper than Claude Sonnet 4.5 (USD 0.0110) and Gemini 3.1 Pro (USD 0.0128) while outperforming them by +9.70pp and +9.17pp. Harnessing can also correct frontier model failures; workflow constraints can outweigh parameter count or API cost. Beyond aggregate accuracy, the pipeline produces structured intermediate artifacts that make each stage inspectable and support error localization. These properties support high-stakes settings where traceability, reproducibility, and auditable evidence matter alongside benchmark performance.

13:00 JST研究/論文

MIITA: 小規模言語モデルによる継続学習のためのメモリ誘導推論時間適応

継続学習 (CL) は、小規模言語モデル (SLM) がリソースに制約のある展開で進化する現実世界のニーズに適応するために不可欠です。ただし、限られたパラメータ空間を直接更新すると、致命的な忘却が発生します。メモリベースの手法では、知識の保持をパラメータから切り離すことで自然にこの問題に対処しますが、大規模言語モデル (LLM) 向けに設計された既存のアプローチは、SLM にはない豊富なストレージと強力なコンテキスト内推論に依存しています。これらの課題に対処するために、制約されたストレージの下で監視された CL のためのメモリ誘導推論時間適応フレームワークである MIITA を提案します。 MIITA は、教師ありエクスペリエンスを、セマンティック アンカーを備えたコンパクトな修正方向プロトタイプとして保存し、推論時にセマンティックおよび不確実性ベースの手がかりを使用して取得します。取得された方向は、ゲート制御された一時的な隠れ状態適応を通じて適用され、バックボーンの更新、プロンプト拡張、またはテスト時のバックプロパゲーションを行わずに、過去の監視を非破壊で再利用できるようになります。局所的な理論分析は、この設計を一次損失低減、不確実性ガイド検索、および古い段階の知識を保持するための指向性カバレッジに関連付けます。さまざまな教師あり CL 設定にわたる広範な実験により、MIITA は固定メモリ バジェットの下で最終パフォーマンスを一貫して向上させ、忘却を軽減することが示されています。

原文 (English)

MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models

Continual learning (CL) is essential for small language models (SLMs) to adapt to evolving real-world needs in resource-constrained deployments. However, directly updating their limited parameter space causes catastrophic forgetting. While memory-based methods naturally address this by decoupling knowledge retention from parameters, existing approaches designed for large language models (LLMs) rely on abundant storage and strong in-context reasoning that SLMs lack. To address these challenges, we propose MIITA, a Memory-Induced Inference-Time Adaptation framework for supervised CL under constrained storage. MIITA stores supervised experiences as compact correction-direction prototypes with semantic anchors, and retrieves them at inference time using semantic and uncertainty-based cues. The retrieved directions are applied through gated temporary hidden-state adaptation, enabling non-destructive reuse of past supervision without backbone updates, prompt extensions, or test-time backpropagation. A local theoretical analysis links this design to first-order loss reduction, uncertainty-guided retrieval, and directional coverage for retaining old-stage knowledge. Extensive experiments across diverse supervised CL settings show that MIITA consistently improves final performance and mitigates forgetting under fixed memory budgets.

13:00 JSTLLM/生成AIビジネス/資金調達

ジャッジの体系化: プログラム蒸留によるスケーラブルな評価

LLM-as-a-judge は自動評価の標準となっていますが、高コスト、大幅な遅延、不透明な決定という、拡張性と信頼性を損なう制限に悩まされています。私たちは、プログラム蒸留というシンプルで効率的な代替手段でこれらに対処します。評価時に LLM を促す代わりに、その決定ロジックを抽出して、候補者を直接採点するプログラムの委員会を作成します。これらのプログラムによるジャッジは透明性を提供し、簡単に検査または編集でき、サンプルごとの API コストを排除します。この概念に基づいて、私たちは、裁判官としてプログラムを統合し、その決定を共同評決に集約し、信頼性の低い訴訟を選択的に LLM にエスカレーションするフォールバック メカニズムを組み込むシステムである PAJAMA を導入します。 5 つのデータセットと 4 つのモデル ファミリーにわたって、プログラムによる裁判官が 13B サイズの LLM 裁判官のパフォーマンスに匹敵することができることを示します。プログラム出力をルーティング信号として使用すると、PAJAMA は精度とスループットの両方を向上させ、パレート フロンティアを前進させます。評価を超えて、プログラムによるジャッジは安価で効果的な報酬シグナルを生成します。RewardBench では、プログラムの評決から抽出された報酬モデルが、2 桁低い API コストで独自の LLM ラベルでトレーニングされた報酬モデルよりも優れたパフォーマンスを発揮します。

原文 (English)

Codifying the Judge: Scalable Evaluation via Program Distillation

LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

SF-AMS: LLM エージェントにおける構造化記憶の戦略的忘却

冗長で無関係な情報により複数ステップの推論が低下する可能性があるため、長いコンテキストの依存関係の管理は依然として LLM エージェントの主なボトルネックとなっています。エージェント メモリ システムのための戦略的忘却 (SF-AMS) は、メモリ ユニットの長期的な重要性をモデル化することで、コンパクトな高ユーティリティ メモリを維持するためのフレームワークとして提案されています。 SF-AMS は、静的検索とヒューリスティック減衰を、使用の冗長性と時間信号からメモリの重要性を更新するユーティリティ主導の生存メカニズムに置き換え、ノイズをフィルタリングしながら安定したエンティティ一貫性のある情報を優先する階層型メモリ構造を誘導します。これに加えて、複合重要度スコアリングはセマンティック レベルとエンティティ レベルの信号を統合して、取得の堅牢性を向上させます。 LoCoMo および LongMemEval-s の実験では、LightMem MemO および A-Mem を含む最先端の強力なベースラインを上回る一貫した利益が示されています。最大の改善は、SF-AMS が最も強いベースラインを上回る +9.65 F1 を達成する Qwen2.5-7B の下でのマルチホップ推論に現れ、続いて GPT-4o-mini + 6.91 F1 およびオープンドメイン タスク + 6.53 F1 の下での時間的推論が、強力なクロス バックボーン一般化を示しています。これらの結果は、信頼性の高いロングコンテキスト推論には、メモリの重要性を動的なユーティリティ信号としてモデル化することが重要であることを示しています。

原文 (English)

SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent

Managing long-context dependencies remains a primary bottleneck in LLM agents, as redundant and irrelevant information can degrade multi-step reasoning. Strategic Forgetting for Agent Memory Systems (SF-AMS) is proposed as a framework for maintaining compact high-utility memory by modeling the long-term importance of memory units. SF-AMS replaces static retrieval and heuristic decay with a utility-driven survival mechanism that updates memory importance from usage redundancy and temporal signals, inducing a hierarchical memory structure that prioritizes stable entity-consistent information while filtering noise. On top of this, Composite Importance Scoring integrates semantic and entity level signals to improve retrieval robustness. Experiments on LoCoMo and LongMemEval-s show consistent gains over strong state of the art baselines including LightMem MemO and A-Mem. The largest improvement appears in multi-hop reasoning under Qwen2.5-7B where SF-AMS achieves plus 9.65 F1 over the strongest baseline followed by temporal reasoning under GPT-4o-mini plus 6.91 F1 and open-domain tasks plus 6.53 F1 demonstrating strong cross backbone generalization. These results show that modeling memory importance as a dynamic utility signal is critical for reliable long-context reasoning.

13:00 JSTエージェントビジネス/資金調達研究/論文

インダストリー 4.0 エージェントの評価のための合成シナリオの生成

産業用エージェントのベンチマークには、テレメトリ、障害モード、メンテナンス記録、ドメイン標準を統合した現実的な評価シナリオが必要です。ただし、AssetOpsBench などの既存のベンチマークは、手動で作成されたシナリオに依存しており、限られた資産クラスのセットをカバーしています。当社は、スマート グリッド変圧器アセット クラスと、健全性指数予測、溶存ガス分析、巻線温度評価、および負荷プロファイル評価のための 4 つの IEC ベースの診断ツールを使用して AssetOpsBench を拡張します。さらに、合成産業エージェント シナリオ生成用のパイプラインである ScenarioGeneratorAgent を紹介します。このパイプラインは、証拠に基づいた資産プロファイルを構築し、運用ドメイン全体にカバレッジを意識したシナリオ予算を割り当て、スキーマの妥当性、ツールの到達可能性、物理的な妥当性、標準の整合性、重複排除を強制するハイブリッド検証と修復のループを通じて候補を生成します。スケーラビリティを向上させるために、2 レベルのキャッシュ、並列フォーカス グループ生成、スレッド プールのオフロード、バッチ化された LLM 呼び出し、および早期拒否フィルタリングを適用します。 Smart Grid Transformer のシナリオ生成では、これらの最適化により、品質を維持しながら 50 のシナリオでエンドツーエンドのランタイムが $8\times$ 削減され、最適化されていないベースラインの $73.8 \pm 3.0$ と比較して $74.2 \pm 1.9$ の複合品質スコアを達成しました。これらの結果は、標準に基づいた合成シナリオ生成により、シナリオの品質を犠牲にすることなく産業エージェントのベンチマークを効率的に拡張できることを示しています。

原文 (English)

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks such as AssetOpsBench rely on manually authored scenarios and cover a limited set of asset classes. We extend AssetOpsBench with a Smart Grid Transformer asset class and four IEC-grounded diagnostic tools for health-index prediction, dissolved-gas analysis, winding-temperature assessment, and load-profile assessment. We further introduce ScenarioGeneratorAgent, a pipeline for synthetic industrial-agent scenario generation. The pipeline constructs evidence-grounded asset profiles, allocates coverage-aware scenario budgets across operational domains, and generates candidates through a hybrid validation-and-repair loop that enforces schema validity, tool reachability, physical plausibility, standards alignment, and deduplication. To improve scalability, we apply two-level caching, parallel focus-group generation, thread-pool offloading, batched LLM calls, and early rejection filtering. On Smart Grid Transformer scenario generation, these optimizations reduce end-to-end runtime by $8\times$ for 50 scenarios while preserving quality, achieving a composite quality score of $74.2 \pm 1.9$ compared with $73.8 \pm 3.0$ for the unoptimized baseline. These results show that standards-grounded synthetic scenario generation can efficiently expand industrial-agent benchmarks without sacrificing scenario quality.

13:00 JST研究/論文

マルチアーム バンディットを使用した畳み込みニューラル ネットワークにおける損失を意識した特徴マップ プルーニング

畳み込みニューラル ネットワークには、ストレージと推論のコストを増加させる冗長な特徴マップが含まれることがよくあります。この論文では、マルチアーム バンディットを使用した、損失を認識した特徴マップ プルーニング フレームワークを紹介します。特徴マップ プルーニングは、分離されたスカラー重みではなく、完全な畳み込み出力チャネルとその生成フィルターを削除するために構造化されています。各候補特徴マップはアームとして扱われます。再生時間ごとに、1 つのマップが一時的にマスクされ、サンプリングされたミニバッチで評価されます。次にマップが復元され、観察された損失の変化が安全な除去の報酬に変換されます。一定の再生バジェットの後、候補マップは学習されたスコアによってランク付けされ、上位 k 個のマップはフィルター、バイアス、および対応する次層の入力チャネル カーネルとともに永久に削除されます。この研究では、UCB1 とトンプソン サンプリングを評価し、LeNet/MNIST での直接/オラクル スタイルの評価と比較し、評価を MNIST、CIFAR-10、CIFAR-100、SVHN、CUB-200-2011、および Oxford Flowers 102 に拡張しました。結果は、UCB1 とトンプソン サンプリングが、特徴マップを削除して削減しながら、枝刈りされていないモデルに近い精度を維持していることを示しています。畳み込み計算。フリードマンおよびネメニのテストでは、UCB1 が最高の平均ランクを取得し、次にトンプソン サンプリングが続くことが示されています。どちらも、元の枝刈りされていないモデルと統計的に同等の状態を保ちながら、貪欲で大きさに基づく枝刈りを大幅に上回ります。

原文 (English)

Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits

Convolutional neural networks often contain redundant feature maps that increase storage and inference cost. This paper presents a loss-aware feature-map pruning framework using multi-armed bandits. Feature-map pruning is structured because it removes complete convolutional output channels and their producing filters rather than isolated scalar weights. Each candidate feature map is treated as an arm. At each play time, one map is temporarily masked and evaluated on a sampled mini-batch; the map is then restored and the observed loss change is converted into a safe-removal reward. After a fixed play budget, candidate maps are ranked by learned scores and the top-k maps are permanently removed with their filters, biases and corresponding next-layer input-channel kernels. The study evaluates UCB1 and Thompson Sampling, compares them with direct/oracle-style evaluation on LeNet/MNIST, and extends the evaluation to MNIST, CIFAR-10, CIFAR-100, SVHN, CUB-200-2011 and Oxford Flowers 102. Results show that UCB1 and Thompson Sampling preserve accuracy close to unpruned models while removing feature maps and reducing convolutional computation. Friedman and Nemenyi tests show that UCB1 obtains the highest mean rank, followed by Thompson Sampling; both significantly outperform greedy and magnitude-based pruning while remaining statistically comparable to the original unpruned model.

13:00 JST研究/論文Claude

DSTFView: デュアル入力の時空間周波数モデリングによるマルチビューのクラウドエッジ ワークロード予測

エッジサイド AI 推論の広範な導入に伴い、エッジ プラットフォームでは、遅延に敏感で同時実行性が高く、信頼性が重要なアプリケーションをサポートすることがますます求められています。ただし、既存の方法では、協調的なクラウド エッジ環境での多次元特徴モデリングと予測効率のバランスを取るのに苦労することがよくあります。この問題に対処するために、共同クラウド エッジ環境向けのデュアル入力時空間周波数マルチビュー ワークロード予測フレームワークである DSTFView を提案します。近接性と周期の依存関係を共同でモデル化し、空間、時間、周波数領域の依存関係を抽出します。さらに、適応型融合メカニズムを設計し、各ビューの寄与を調整して突然の変化を捉えます。 CPU および TP データセットの実験結果は、DSTFView が複数の予測期間および評価指標にわたって代表的なベースラインを常に上回るパフォーマンスを示していることを示しています。

原文 (English)

DSTFView: Multi-View Cloud-Edge Workload Forecasting with Dual-Input Spatio-Temporal-Frequency Modeling

With the widespread deployment of edge-side AI inference, edge platforms are increasingly required to support latency-sensitive, highly concurrent, and reliability-critical applications. However, existing methods often struggle to balance multidimensional feature modeling and forecasting efficiency in collaborative cloud-edge environments. To address this issue, we propose DSTFView, a dual-input spatio-temporal-frequency multi-view workload forecasting framework for collaborative cloud-edge environments. It jointly models closeness and period dependencies and extracts spatial, temporal, and frequency-domain dependencies. Besides, it designs an adaptive fusion mechanism and adjusts the contribution of each view to capture abrupt changes. Experimental results on the CPU and TP datasets demonstrate that DSTFView consistently outperforms representative baselines across multiple forecasting horizons and evaluation metrics.

13:00 JST研究/論文

MedLoCoMo: 大規模言語モデル用のロングコンテキスト マルチセッション医療対話ベンチマーク

MedLoCoMo は、複数の入院を伴う医療対話における患者固有の臨床推論のための医療用ロングコンテキスト メモリ ベンチマークです。既存の医療 QA ベンチマークは主に、短い文脈知識または単一文書の根拠をテストしており、LLM が患者の長期的な病歴を使用、接続、および回避できるかどうかは不明のままです。当社は、入院レベルの臨床パケットを構築し、根拠のある医師と患者の会話を合成し、単一入院、クロス入院、および敵対的な回答不能な設定にわたる証拠にリンクされた QA 項目を生成することにより、匿名化された MIMIC-IV および MIMIC-IV-Note 記録から MedLoCoMo を構築します。ベンチマークには 100 の患者のタイムラインが含まれており、会話ごとに平均 1,669.8 ターン、29.7 セッション、74,512.2 トークンが含まれています。評価されたベースライン全体にわたって、モデルに長いコンテキストウィンドウがある場合や、外部メモリや検索方法を使用している場合でも、クロスアドミッション推論は、局所的な証拠の使用よりも一貫して困難です。コードと MedLoCoMo ベンチマーク リリースは、使用と再現のために https://github.com/leozzy13/MedLoCoMo から入手できます。

原文 (English)

MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models

MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test short context knowledge or single document grounding, leaving open whether LLMs can use, connect, and abstain over longitudinal patient histories. We build MedLoCoMo from deidentified MIMIC-IV and MIMIC-IV-Note records by constructing admission-level clinical packets, synthesizing grounded doctor-patient conversations, and generating evidence linked QA items over single-admission, cross-admission, and adversarial unanswerable settings. The benchmark contains 100 patient timelines averaging 1,669.8 turns, 29.7 sessions, and 74,512.2 tokens per conversation. Across the evaluated baselines, cross-admission reasoning is consistently harder than localized evidence use, even when models have long context windows or use external memory or retrieval methods. The code and MedLoCoMo benchmark release is available at https://github.com/leozzy13/MedLoCoMo for use and reproducibility.

13:00 JSTLLM/生成AI

重要なキーワード: オンデバイス LLM プロンプトのエネルギー感度の解明

プライバシーを向上させ、ネットワーク遅延を削減するために、ラージ言語モデル (LLM) がモバイル デバイスや組み込みデバイスに導入されることが増えています。しかし、オンデバイス推論は、バッテリ駆動でリソースが限られているハードウェアでの高いエネルギー消費という根本的な制約に直面しています。モデルの圧縮と実行時の高速化については広く研究されていますが、\emph{プロンプト設計} がエネルギー効率に及ぼす影響についてはまだ十分に解明されていません。この論文では、プロンプトの文言とオンデバイス LLM のエネルギー消費の関係についての実証研究を紹介します。スマートフォンで収集された実電力測定値を使用して、言語的特徴、特に命令型キーワードと命令構造がデコード長と総エネルギーにどのような影響を与えるかを定量化します。私たちの結果は、動詞とタスク間で一貫したエネルギーの違いを示しており、プロンプトエンジニアリングがエネルギー効率を改善するための軽量の手段であることを示しています。

原文 (English)

Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting

Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on-device inference faces a fundamental constraint: high energy consumption on battery-powered, resource-limited hardware. While model compression and runtime acceleration have been widely studied, the effect of \emph{prompt design} on energy efficiency remains underexplored. This paper presents an empirical study of the relationship between prompt wording and energy consumption for on-device LLMs. Using real power measurements collected on a smartphone, we quantify how linguistic features, particularly imperative keywords and instruction structure, affect decoding length and total energy. Our results show consistent energy differences across verbs and tasks, indicating that prompt engineering is a lightweight lever for improving energy efficiency.

13:00 JSTエージェントClaude

ソフトウェア エンジニアリング パイプラインにおけるコーディング エージェントの実行ベースのセキュリティ テスト

コーディング エージェントはシステム運用にますます統合されており、そのツールを使用すると、プロジェクトの成果物、実行環境、基盤となるシステムが直接変更される可能性があります。たとえば、コーディング エージェントがシステムの起動スクリプトまたは構成スクリプトにフックを挿入した場合、その変更は対話後も持続し、後でトリガーされ、委任されたユーザーまたはシステムの権限を悪用してシステムを変更する可能性があります。このため、セキュリティ テストはシステムの問題になります。重要な問題は、エージェントが何を言うかではなく、エージェントが周囲の環境に実際に何をするかということです。私たちは、ツールの呼び出し、ランタイム トレース、ファイル システムの差分などの観察可能なサンドボックス証拠を使用して、この実行層のセキュリティ境界を調査するための、実行ベースのレッド チーム テスト フレームワークを紹介します。当社のフレームワークは、単体テスト、回帰テスト、クラッシュ再現、検証などの日常的なソフトウェア エンジニアリング ワークロードにターゲットの安全でない操作を組み込み、最初のプローブが拒否または失敗した場合に実行オラクルを使用して改良をガイドします。複数のエージェント フレームワークとモデル バックボーンにわたって、レッド チームのワークロード再構築により、検証済みの安全でない実行が大幅に増加し、コード キャリアでは 73.61%、テキスト キャリアでは 53.93% に達しました。これらの結果は、システム操作におけるコーディング エージェントがタスク偽装の下では安全でないことを示しています。つまり、もっともらしいエンジニアリング タスクの中に危険な意図が隠蔽されると、エージェントは周囲のシステム上で危険なアクションを実行するよう誘導される可能性があります。さらに広く言えば、システム運用におけるコーディング エージェントは依然として、より強力なセキュリティ テストと保護手段を必要としています。

原文 (English)

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines

Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For example, if a coding agent inserts a hook into a system startup or configuration script, that change can persist after the interaction, be triggered later, and abuse delegated user or system privileges to modify the system. This makes security testing a system problem: the key question is not only what the agent says, but what it actually does to the surrounding environment. We present an execution-grounded red-team testing framework for probing this execution-layer security boundary using observable sandbox evidence, including tool invocations, runtime traces, and file-system diffs. Our framework embeds target unsafe operations into routine software engineering workloads, including unit testing, regression testing, crash reproduction, and validation, and uses an execution oracle to guide refinement when an initial probe is rejected or fails. Across multiple agent frameworks and model backbones, our red-team workload reformulation substantially increases verified unsafe execution, reaching 73.61% on code carriers and 53.93% on text carriers. These results show that coding agents in system operations remain insecure under task disguise: once risky intent is hidden inside plausible engineering tasks, the agent can be induced to carry out unsafe actions on the surrounding system. More broadly, coding agents in system operations still demand stronger security testing and safeguards.

13:00 JST研究/論文Mistral AIQwen

言語モデルの機構監査のための参照機能アトラス

新しい言語モデルを監査するということは、通常、その内部機能を最初から再学習して再解釈することを意味します。我々は、参照特徴アトラスを提案します。これは、参照パネル上で一度トレーニングされ、線形デコーダのみをフィッティングすることによって接続される新しいターゲットに対して再利用される疎な特徴ライブラリです。これにより、2 つの相補的なビューが得られます。アトラス チャネルは、すでに解釈されたパネル フィーチャ上のターゲットを読み取り、モデル全体で安定した座標系を提供します。残留チャネルは、アトラスが再構築できなかったものからのみ特徴を学習し、「参照パネルの外側」を明示的な監査信号にします。私たちは 5 つの 7-9B 命令調整モデルでリーブワンアウト アトラスをトレーニングし、保留されたミストラルとクウェンのターゲットを監査します。両方のターゲットに注入された 3 つの制御された LoRA 隠し目標では、残留チャネルにより、実行時に植え付けられたメカニズムが完全に制御可能になり、一致するコントロールは影響を受けず、両方のターゲットにわたって植え付けられた目標が最上位の潜在目標として回復されます。 Mistral では、ターゲットごとの SAE とペアワイズ クロスコーダーのベースラインが直接ベンチマーク用に再トレーニングされますが、両方のベースラインが再トレーニングされません。 Qwen-2.5 では、同じチャンネルがパネルに関連した政治的枠組みのクラスターをさらに明らかにします。ステアリングを操作すると、監査対象のフレーミング メトリクスがシフトされますが、ドメイン外のコントロールは変更されません。

原文 (English)

Reference Feature Atlases for Mechanistic Auditing of Language Models

Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by fitting only a linear decoder. This yields two complementary views. The atlas channel reads the target on already interpreted panel features, providing a stable coordinate system across models. The residual channel learns features only from what the atlas fails to reconstruct, making "outside the reference panel" an explicit audit signal. We train leave-one-out atlases over five 7-9B instruction-tuned models and audit held-out Mistral and Qwen targets. On three controlled LoRA hidden objectives injected into both targets, the residual channel makes the planted mechanism perfectly controllable at runtime while matched controls stay unaffected and recovers the planted objective as the top-ranked latent across both targets; on Mistral, where the per-target SAE and pairwise crosscoder baselines are retrained for a head-to-head benchmark, both baselines fail to do so. On Qwen-2.5, the same channel additionally reveals a panel-relative political-framing cluster; steering it shifts the audited framing metrics while out-of-domain controls remain unchanged.

13:00 JSTエージェント

SCAIR: エンタープライズ ナレッジ グラフのためのスキーマ条件付きエージェント反復推論

ナレッジ グラフ ベースの検索拡張生成 (KG-RAG) により、構造化されたエンタープライズ ナレッジとの自然言語対話が可能になりますが、公開ベンチマークで良好に機能する既存のエージェント アプローチは、高密度でスキーマ駆動型で運用上の制約がある現実世界のエンタープライズ ナレッジ グラフ (KG) に一般化できないことがよくあります。これらの制限に対処するために、私たちは、スキーマ条件付き構造事前分布を挿入し、マルチホップ推論中にスキーマを意識した走査を強制することにより、構造化計画と制御された反復推論を統合するトレーニング不要のフレームワークである SCAIR (スキーマ条件付きエージェント反復推論) を提案します。実際の構成管理データベース (CMDB) から構築されたエンタープライズ向けベンチマークの実験では、SCAIR が既存の KG-RAG 手法と比較してパフォーマンスを大幅に向上させることが実証されました。重要なことに、私たちの研究は、信頼できるエンタープライズ グラフ推論が汎用エージェント設計に依存できないことを強調しています。代わりに、ターゲット ドメインの構造的および操作上の制約を推論プロセスに明示的に組み込む必要があります。エージェントの設計をビジネス ロジックに合わせることで、コストのかかるモデルの再トレーニングを必要とせずに、大幅なパフォーマンスの向上が達成できることを実証します。

原文 (English)

SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs

Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing agentic approaches that perform well on public benchmarks often fail to generalize to real-world enterprise Knowledge Graphs (KGs), which are dense, schema-driven, and operationally constrained. To address these limitations, we propose SCAIR (Schema-Conditioned Agentic Iterative Reasoning), a training-free framework that integrates structured planning with controlled iterative reasoning by injecting schema-conditioned structural priors and enforcing schema-aware traversal during multi-hop reasoning. Experiments on an enterprise-oriented benchmark constructed from a real-world Configuration Management DataBase (CMDB) demonstrate that SCAIR substantially improves performance over existing KG-RAG methods. Crucially, our study highlights that reliable enterprise graph reasoning cannot rely on generic agentic designs; instead, it must explicitly incorporate the target domain's structural and operational constraints into the reasoning process. We demonstrate that by aligning agent design with business logic, substantial performance gains can be achieved without the need for costly model retraining.

13:00 JST研究/論文GPT / ChatGPT

スキーマ認識ローカリゼーション (SAL): Oracle NL2SQL のライブ スキーマ グラウンディングと幻覚検証

大規模な言語モデルでは自然言語から流暢な SQL を生成できますが、実際のエンタープライズ Oracle データベースでは実行時に頻繁に失敗します。列と別名が幻覚を受け、方言固有の構文が失われ、ORA-00904 無効な識別子エラーが発生します。この設定では、失敗の主な原因はスキーマの基礎が欠落していることです。モデルはどのテーブルと列が実際に存在するかを認識できません。このペーパーでは、モデルの再トレーニングを必要としない Oracle NL2SQL の軽量ミドルウェア層である Schema-Aware Localization (SAL) について紹介します。 SAL は、Oracle の USER_TAB_COLUMNS カタログにクエリを実行してライブ スキーマ マップを構築し、質問ごとに関連するテーブル サブセットを選択し (複数テーブル クエリの場合は完全なスキーマにフォールバックします)、このグラウンド トゥルース コンテキストを LLM プロンプトに挿入します。生成された SQL は、Hallucination Index (Hidx) によってチェックされます。Hidx は、すべての alias.column 参照をライブ カタログに対して検証し、予測可能なプレフィックス エラーを自動的に書き換え、それ以外の場合は項目別の修正を伴う構造化された再試行をトリガーします。 GPT-4o-miniを使用して、ライブOracle Autonomous Database 23cインスタンスに対して実行された500のTPC-H自然言語の質問についてSALを評価しました。スキーマの根拠がない場合、実行根拠のある真実 (EGT; 実行して参照結果セットと一致する) は 2.2% (12/500) です。手書きの静的スキーマ ヒントにより、EGT は 62.0% になります。 SAL は、手動によるスキーマ管理を行わないため、62.6% の EGT (単純 96%、中程度 95%、複雑 40.7%) を達成し、実行失敗を 97.6% から 2.6% に削減します。

原文 (English)

Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL

Large language models can generate fluent SQL from natural language, but on real enterprise Oracle databases they frequently fail at execution time: columns and aliases are hallucinated and dialect-specific syntax is missed, leading to ORA-00904 invalid-identifier errors. In this setting, failures are primarily due to missing schema grounding: the model cannot know which tables and columns actually exist. This paper introduces Schema-Aware Localisation (SAL), a lightweight middleware layer for Oracle NL2SQL that requires no model retraining. SAL queries Oracle's USER_TAB_COLUMNS catalog to build a live schema map, selects a relevant table subset for each question (falling back to the full schema for multi-table queries), and injects this ground-truth context into the LLM prompt. Generated SQL is then checked by the Hallucination Index (Hidx), which validates every alias.column reference against the live catalog, automatically rewrites predictable prefix errors, and otherwise triggers a structured retry with itemised corrections. We evaluate SAL on 500 TPC-H natural language questions executed against a live Oracle Autonomous Database 23c instance using GPT-4o-mini. Without any schema grounding, execution-grounded truth (EGT; executes and matches the reference result set) is 2.2% (12/500). A hand-written static schema hint brings EGT to 62.0%. SAL, with no manual schema curation, achieves 62.6% EGT (96% simple, 95% medium, 40.7% complex) while reducing execution failures from 97.6% to 2.6%.

13:00 JST研究/論文

PhononBench-MP40: フォノン安定性のためのスペクトル分解ベンチマーク データセット

想像上のフォノンモードは、選択されたワークフローの下では、もっともらしい構造が局所的に動的に不安定になる可能性があるため、計算材料のスクリーニングにおいて依然として実用的なボトルネックとなっています。ここでは、ワークフロー定義のフォノン安定性のためのマテリアルズ プロジェクト由来の結晶のスペクトル分解ベンチマーク データセットである PhononBench-MP40 を紹介します。このデータセットは 47,969 個の MP40 ワークフロー タスクから始まり、ペアの安定性ラベルとローカル フォノピー YAML スペクトルを持つ 46,899 個の完了したレコードを提供します。これには、16,683 個の安定レコードと 30,216 個の完全なフォノン不安定レコードが含まれます。さらに 1,067 件の緩和失敗が、完成したフォノン分母に統合されるのではなく、個別に報告されます。このリリースは、ローカル YAML スペクトルに重点を置いています。安定性ラベル、最低サンプリング周波数、およびしきい値に依存する再ラベル付けは、そのスペクトルから導出されます。このデータセットは、Science Data Bank (https://doi.org/10.57760/sciencedb.38735) を通じて公開されています。コンパニオン GitHub リポジトリは、計算コードと軽量アクセス ユーティリティを提供します。 PhononBench-MP40 は、参照ワークフロー、データ スキーマ、および解釈の境界を明確に保ちながら、ワークフロー定義の安定性分類、最小周波数分析、しきい値調査、障害認識トリアージのための監査可能な参照を提供します。

原文 (English)

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be locally dynamically unstable under a chosen workflow. Here we present PhononBench-MP40, a spectrum-resolved benchmark dataset of Materials Project-derived crystals for workflow-defined phonon stability. The dataset starts from 47,969 MP40 workflow tasks and provides 46,899 completed records with paired stability labels and local phonopy YAML spectra, including 16,683 Stable records and 30,216 completed-phonon unstable records. A further 1,067 relaxation failures are reported separately rather than merged into the completed phonon denominator. The release centers on the local YAML spectrum: the stability label, the lowest sampled frequency and any threshold-dependent relabeling are derived from that spectrum. The dataset is openly available through Science Data Bank at https://doi.org/10.57760/sciencedb.38735. A companion GitHub repository provides the calculation code and lightweight access utilities. PhononBench-MP40 provides an auditable reference for workflow-defined stability classification, minimum-frequency analysis, threshold studies and failure-aware triage, while keeping the reference workflow, data schema and interpretation boundaries explicit.

13:00 JST研究/論文

証拠が多すぎて時間が少なすぎる: テキストから多目的証拠推論による実用的な推奨事項まで

証拠に基づいた臨床上の意思決定には、専門家が関連する科学文献を特定、評価、総合する必要があります。ただし、PubMed で複雑な臨床症例を検索すると、時間の制約で手動でレビューできない数百もの出版物が返されることがよくあります。この研究は、臨床症例の説明を証拠に基づいた推奨事項に変換するためのフレームワークである SCEPTER (Single-Case Evidence-driven PubMed-To-rEcommendation Reasoner) を提案します。 SCEPTER は、PubMed 検索、PubMedBERT 意味論的ランキング、大規模言語モデル (LLM) ベースのクレーム抽出、証拠レベルの重み付け、矛盾検出、コンセンサス分析、多目的パレートクレーム選択を組み合わせています。このフレームワークは、構造化された証拠の総合と、根拠のある実用的な推奨事項を生成します。 Paper Q&A モジュールを使用すると、選択した出版物をインタラクティブに探索できます。提案されたフレームワークは、文献サポート、矛盾分析、およびインタラクティブな文献尋問を統合された臨床意思決定サポート パイプラインに統合する多目的推論モデルを導入します。 150 件のケーススタディの評価では、このフレームワークにより、平均 576 件の論文の検索スペースが 53 件の保持論文、7 件のパレート最適主張、および 3 件の最終推奨に削減され、全体の圧縮率は 192:1 に相当することが実証されました。この減少にもかかわらず、保持された証拠は高い多様性を維持しました (エントロピー = 0.901)。アブレーション研究では、パレートに基づいた選択により、従来のランキング手法と比較して証拠の多様性と推奨の有用性が向上することが示されました。

原文 (English)

Too much evidence, too little time: From text to actionable recommendations through multi-objective evidence reasoning

Evidence-based clinical decision making requires specialists to identify, evaluate and synthesize relevant scientific literature. However, PubMed searches for complex clinical cases often return hundreds of publications that cannot be reviewed manually under time constraints. This study proposes SCEPTER (Single-Case Evidence-driven PubMed-To-rEcommendation Reasoner), a framework for transforming clinical case descriptions into evidence-based recommendations. SCEPTER combines PubMed retrieval, PubMedBERT semantic ranking, large language model (LLM)-based claim extraction, evidence-level weighting, contradiction detection, consensus analysis and multi-objective Pareto claim selection. The framework generates structured evidence syntheses and grounded actionable recommendations. A Paper Q&A module further enables interactive exploration of selected publications. The proposed framework introduces multi-objective reasoning model that integrates literature support, contradiction analysis and interactive literature interrogation into a unified clinical decision-support pipeline. Evaluation on 150 case studies demonstrated that the framework reduced an average search space of 576 papers to 53 retained papers, 7 Pareto-optimal claims and 3 final recommendations, corresponding to an overall compression ratio of 192:1. Despite this reduction, the retained evidence maintained high diversity (entropy=0.901). The ablation study showed that Pareto-based selection increased evidence diversity and recommendation utility compared with conventional ranking approaches.

13:00 JST研究/論文

時間的コンテキストの復元により、ロングコンテキスト言語モデルにおけるエピソードのような順序記憶が促進される

人間のエピソード記憶は、長い時間スケールで展開される経験の検索をサポートしますが、人間の長期記憶実験では機械的アクセスが制限されているため、この能力の基礎となる計算メカニズムについては議論が続いています。ロングコンテキスト LLM は、このタイプの検索を推進するもっともらしい計算メカニズムを明らかにする有望な方法を提供する可能性があります。ここでは、LLM が時間順序記憶タスクを介してエピソード記憶の中核となる行動シグネチャを捕捉するかどうか、またその方法を調査します。長編小説の記憶に基づいた人間の行動の新しいデータセットを使用して、モデルがこのタスクにおいて人間で観察されたのと同じ特徴的な距離効果を示すことを示します。次に、ロングコンテキストの機構的解釈可能性分析を適用して、モデルがこのタスクをどのように解決するかを明らかにします。また、モデルのパフォーマンスは、単一の時間復元アテンションヘッドによる検索中に復元される 1 次元の時間コードに依存していることがわかります。これらの発見は、時間的文脈の復元がLLMにおけるエピソード的な時間順序記憶の重要なメカニズムであることを裏付けており、人工システムと生物学的システムの両方で長期エピソード記憶の時間的側面がどのようにインスタンス化されるかについての新たな洞察を提供します。

原文 (English)

Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models

Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this ability remain debated due to the limited mechanistic accessibility in long-term memory experiments in humans. Long-context LLMs may offer promising ways to reveal plausible computational mechanisms that drive this type of retrieval. Here, we investigate whether and how LLMs capture the core behavioral signatures of episodic memory via a temporal order memory task. Using a new dataset of human behavior based on memory of a full-length novel, we show that models exhibit the same characteristic distance effect observed in humans on this task. We next apply long-context mechanistic interpretability analyses to uncover how models solve this task, and find that model performance relies on a one-dimensional temporal code that is reinstated during retrieval by a single time-reinstatement attention head. These findings support temporal context reinstatement as an important mechanism for episodic-like temporal-order memory in LLMs, offering new insights into how temporal aspects of long-term episodic memory may be instantiated in both artificial and biological systems.

13:00 JSTLLM/生成AIGPT / ChatGPT

大規模な cMoLLM: LLM の混合に関する水平スケーリングの法則

大規模言語モデル (LLM) のスケーリングが成功の原動力となっていますが、トランスフォーマー カップルの容量と計算の密度が高いため、すべてのパラメーターがすべてのトークンに対してアクティブ化されるため、トレーニングと推論のコストがモデル サイズに比例して増加します。これは、モデルが数兆パラメーターの体制に近づくにつれて、重大なボトルネックになります。当社は、FFN のみではなく、LLM パイプライン全体で MoE スタイルの混合を通じて容量を拡大することを目指しています。これまでのパイプライン レベルのアプローチには、仮想トークンと並列ストリームを導入するものの、大幅なオーバーヘッドが発生し、均一化されたルーティングと勾配の崩壊に悩まされる ParaScale や、補助予測ブランチを使用するが適応性が限られ、収束が遅い AltUp などがあります。我々は、MoE スタイルの混合層が変数カーネルの動的畳み込みとして再定式化できることを確立します。各エキスパートは $1{\times}1$ 畳み込みカーネルに対応し、ルーティングは入力条件付きカーネル集約を実装します。この等価性に基づいて、cMoLLM を導入します。cMoLLM は、完全に微分可能な動的畳み込みを通じてエンドツーエンドのストリーム上でルーティングする、畳み込みゲートされた LLM の混合物です。 FineWeb でトレーニングされた GPT-2 スタイルのモデルでは、cMoLLM は、ParaScale および AltUp スタイルのベースラインと比較して、より優れたストリーム利用率、より安定した最適化、有利なスケーリングにより、言語モデリングの複雑さと、一致したコンピューティングの下で​​のダウンストリームの GLUE および SQuAD の精度を向上させます。

原文 (English)

cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs

Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes. We aim to scale capacity through MoE-style mixture throughout the LLM pipeline rather than only the FFN. Prior pipeline-level approaches include ParaScale, which introduces virtual tokens and parallel streams but incurs substantial overhead and suffers from homogenized routing and gradient collapse, and AltUp, which uses an auxiliary prediction branch but offers limited adaptivity and slow convergence. We establish that MoE-style mixture layers can be reformulated as variable-kernel dynamic convolutions, where each expert corresponds to a $1{\times}1$ convolutional kernel and routing implements input-conditioned kernel aggregation. Building on this equivalence, we introduce cMoLLM: a convolutionally gated mixture-of-LLMs that routes over end-to-end streams through fully differentiable dynamic convolution. In GPT-2-style models trained on FineWeb, cMoLLM improves language modeling perplexity and downstream GLUE and SQuAD accuracy under matched compute, with better stream utilization, more stable optimization, and favorable scaling compared to ParaScale- and AltUp-style baselines.

13:00 JSTLLM/生成AI

HeraSys: きめ細かいエンドツーエンドの最適化による複数の LLM ワークフローの共同提供

Large Language Model (LLM) の普及により、サービス提供システムは、分離されたリクエストの処理から、同時実行性の高いマルチテナントのエージェント ワークフローの調整へと移行しました。ただし、既存のソリューションは通常、ワークフロー内の最適化を優先し、ワークフロー間の最適化の大きな可能性をほとんど無視しています。このペーパーでは、同時ワークフローのエンドツーエンドのパフォーマンスを最適化するように設計された LLM サービス システムである HeraSys を提案します。 HeraSys は、きめ細かいオーケストレーションを通じて、構造的なノードのマージと再利用によりワークフロー間の計算の冗長性を排除します。さらに、HeraSys は、クエリ間およびクエリ内の両方の優先順位を評価することで実行順序を動的に管理する、負荷を認識した共同スケジューリング ポリシーを導入します。 HeraSys は、リソース スキュー メカニズムを適応型バッチ処理およびパイプライン分解と統合することにより、平均レイテンシを低く維持しながらテール レイテンシを効果的に軽減し、システム スループットを大幅に向上させます。広範な実験により、HeraSys は厳格なレイテンシ保証の下で P99 レイテンシを最大 2.17$\times$ 削減し、サービング スループットを最大 1.85$\times$ 増加させることが実証されました。

原文 (English)

HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization

The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. However, existing solutions typically prioritize intra-workflow optimization, largely neglecting the significant potential for inter-workflow optimization. In this paper, we propose HeraSys, an LLM serving system designed to optimize the end-to-end performance of concurrent workflows. Through fine-grained orchestration, HeraSys eliminates cross-workflow computational redundancy via structural node merging and reuse. Furthermore, HeraSys introduces a load-aware joint scheduling policy that dynamically manages execution order by evaluating both inter- and intra-query priorities. By integrating a resource skewing mechanism with adaptive batching and pipeline decomposition, HeraSys effectively mitigates tail latency while maintaining low average latency, thereby substantially improving system throughput. Extensive experiments demonstrate that HeraSys reduces P99 latency by up to 2.17$\times$ and increases serving throughput by up to 1.85$\times$ under strict latency guarantees.

13:00 JSTLLM/生成AI

レイテンシとモデル サイズの最適化のための LLM の多目的構造化プルーニング

大規模言語モデル (LLM) は、その強力な推論機能とクエリ応答機能により、広く採用されています。ただし、これらを組み込み環境やエッジ コンピューティング環境に導入することは、遅延、メモリ、エネルギーの厳しい制約のため依然として困難です。パラメーター数と計算量が多いため、リソースに制約のあるプラットフォームでの効率的な実行が妨げられます。モデルの枝刈りは、パフォーマンスを維持しながらスケールを削減するための実行可能なソリューションとして浮上していますが、レイヤー、アテンション ヘッド、および多層パーセプトロン (MLP) 次元の共同最適化は依然として非常に複雑です。この組み合わせた設計空間を徹底的に探索すると、計算コストが高くつき、多くの場合、局所的な最適化や不安定な構成が発生します。これらの制限に対処するために、ハードウェアを認識した多目的構造化プルーニング フレームワークを提案します。提案された 2 段階の方法では、エッジ デバイスでの効率的な展開のためにレイテンシとモデル サイズを明示的にターゲットにしています。粗粒度段階では、多目的深度枝刈りによってアテンション全体と MLP ブロックが削除され、計算負荷とメモリ使用量が削減されます。後続のきめ細かい段階では、並列ベイジアン最適化 (PBO) が、レイテンシ制約の下でプルーニングに最適なレイヤーごとのプルーニング比率を検索します。一方、重要度ベースの戦略は、各レイヤーに割り当てられた予算内でプルーニングされる特定のコンポーネントをランク付けします。実験結果は、私たちのアプローチが常識的推論タスクとゼロショットパフォーマンスへの影響を最小限に抑えながらモデルの複雑さを軽減することを示しています。私たちの方法は、精度、レイテンシ、モデル サイズの間で好ましいトレードオフを達成しており、エッジ展開に適しています。 37.5% と 50% の枝刈り率の複数の LLM にわたって、提案されたアプローチは、推論コストを大幅に削減しながら、常識推論タスクで既存の手法よりも優れたパフォーマンスを達成します。

原文 (English)

Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization

Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory, and energy constraints. Their large parameter counts and computational demands hinder efficient execution on resource-constrained platforms. Although model pruning has emerged as a viable solution for reducing scale while preserving performance, jointly optimizing layers, attention heads, and Multi-Layer Perceptron (MLP) dimensions remains highly complex. Exhaustively exploring this combined design space is computationally expensive and often leads to local optima or unstable configurations. To address these limitations, we propose a hardware-aware, multi-objective structured pruning framework. The proposed two-stage method explicitly targets latency and model size for efficient deployment on edge devices. In the coarse-grained stage, multi-objective depth pruning removes entire attention and MLP blocks to reduce computational load and memory usage. In the subsequent fine-grained stage, Parallel Bayesian Optimization (PBO) searches for the optimal layer-wise pruning ratios for pruning under latency constraints, while importance-based strategies rank the specific components to be pruned within each layer's allocated budget. Experimental results show that our approach reduces model complexity with minimal impact on commonsense reasoning tasks and zero-shot performance. Our method achieves a favorable trade-off among accuracy, latency, and model size, making it suitable for edge deployment. Across multiple LLMs at 37.5% and 50% pruning ratios, the proposed approach achieves better performance on commonsense reasoning tasks than existing methods while significantly reducing inference cost.

13:00 JSTLLM/生成AI

検索拡張生成のためのソース認識型再ランキング: 信頼性優先アプローチ

標準の検索拡張生成パイプラインは、ソースの出所や信頼性を考慮せずに、意味論的な類似性のみによって取得されたドキュメントをランク付けします。この研究では、ドメイン情報に基づいたソース信頼性事前分布を組み込んだ、RAG 検索ランキングに対する単純で解釈可能な修正を評価します。各ドキュメントにはソース タイプに基づいて以前のラムダが割り当てられ、取得スコアはスコア(q, d) = sim(q, d) * ラムダ(s) を使用して再重み付けされます。このフレームワークは、120 文書の健康ドメイン コーパスの類似性のみのベースラインに対して評価されます。この制御された設定では、ソースを意識した再ランキングにより、Precision@5 が 0.48 から 0.72 に向上し、信頼性の低いソースがメタデータを介して識別できる評価された脅威モデルの下で、平均的な敵対的ドキュメントの取得が減少します。すべての実験は、ミルウォーキー スクール オブ エンジニアリングの高性能コンピューティング クラスターである Rosie 上で実行されました。Rosie は、完全な実験パイプラインを確実かつ再現性よく実行するために必要な GPU 高速化インフラストラクチャを提供しました。これらの結果は、説明した実験設定の制限内で、RAG パイプラインにおけるソース品質の低下に対する潜在的な緩和戦略を示唆しています。

原文 (English)

Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach

Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source provenance or credibility. This work evaluates a simple and interpretable modification to RAG retrieval ranking that incorporates domain-informed source reliability priors. Each document is assigned a prior lambda(s) based on its source type, and retrieval scores are reweighted using score(q, d) = sim(q, d) * lambda(s). The framework is evaluated against a similarity-only baseline on a 120-document health-domain corpus. In this controlled setting, source-aware reranking improves Precision@5 from 0.48 to 0.72 and reduces average adversarial document retrieval under the evaluated threat model, where low-credibility sources are identifiable via metadata. All experiments were executed on Rosie, the high-performance computing cluster at the Milwaukee School of Engineering, which provided the GPU-accelerated infrastructure necessary to run the full experimental pipeline reliably and reproducibly. These results suggest a potential mitigation strategy for source quality degradation in RAG pipelines, within the limits of the experimental setup described.

13:00 JSTエージェントビジネス/資金調達Qwen

コーディング エージェントの足場効果: コーディング エージェントの評価における隠れた変数としての選択を利用する

コーディング エージェント向けの公開リーダーボードは通常、モデル名と合格率によってシステムをランク付けしますが、周囲のハーネス (ツールを発行し、コンテキストを管理し、いつ停止するかを決定する足場) は十分に指定されていないことがよくあります。モデル間の比較は、ハーネスが固定されている場合に有効です。パフォーマンスと効率が異なる場合、モデルと足場効果が混同されます。 Terminal-Bench Pro の階層化された 50 タスクのサブセット上の 3 つのオープンソース ハーネス (Goose、OpenCode、OpenHands-SDK) にわたって Qwen 3.6 Plus と MiniMax M2.5 を評価します。ハーネスの選択により、解決されたタスクごとに最大 40 倍のトークンの差が生じますが、モデル内のペアの合格率の差は 0 ~ 8 パーセント ポイントのままです (95% のペア タスク ブートストラップ CI には、最大のギャップを除いてゼロが含まれます)。障害のフィンガープリントはモデル間で複製され (Goose の場合は REASON、OpenHands-SDK の場合は VERIFY/MAX_TURNS、OpenCode の場合は idle-loop/TIME)、モデルにほとんど依存しないハーネス レベルのバイアスを示します。人間中心のコーディング エージェントの評価では、モデル名だけでは比較単位が不完全です。ハーネスとモデルのペアによって、実際のコスト、遅延、監視の負担が決まります。ノーアクション ターンは、単なるトークン税ではなく、タスクごとの待機税です。したがって、トークン/レイテンシ バジェットに基づいてハーネスとモデルのペアを選択し、モデルの比較とともにトークンの使用量、レイテンシ、フル ハーネスの仕様を報告することをお勧めします。匿名化された構成、生の試用ログ、集約されたスナップショット、分析スクリプトをリリースします。

原文 (English)

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-model comparison is valid when the harness is fixed; when it varies, performance and efficiency conflate model and scaffold effects. We evaluate Qwen 3.6 Plus and MiniMax M2.5 across three open-source harnesses (Goose, OpenCode, OpenHands-SDK) on a stratified 50-task subset of Terminal-Bench Pro. Harness choice induces up to a 40x difference in tokens per solved task, while paired within-model pass-rate differences remain 0-8 percentage points (95% paired-task bootstrap CIs include zero except for the largest gap). Failure fingerprints replicate across models (REASON for Goose, VERIFY/MAX_TURNS for OpenHands-SDK, idle-loop/TIME for OpenCode), indicating harness-level biases that are largely model-independent. For human-centered coding-agent evaluation, model name alone is an incomplete comparison unit: harness-model pairs determine real-world cost, latency, and oversight burden; no-action turns are a per-task wait tax, not just a token tax. We therefore recommend selecting harness-model pairs by pass rate under token/latency budgets, and reporting token usage, latency, and full harness specifications alongside any model comparison. We release anonymized configs, raw trial logs, aggregated snapshots, and analysis scripts.

13:00 JSTLLM/生成AI

MM-ShiftKV: マルチモーダル大規模言語モデル向けのデコード対応プレフィル ステージ KV 選択

Key-Value (KV) キャッシュは、マルチモーダル大規模言語モデル (MLLM) での効率的な推論に不可欠ですが、そのメモリ使用量はコンテキストの長さに比例して増加し、ビジュアル トークンの数が多いため大きなボトルネックになります。最近のプレフィル段階の KV 選択方法では、プレフィル統計から KV の重要性を推定し、プレフィル時のクエリがデコード中に発生したものを表すものであると暗黙的に想定しています。マルチモーダル推論では、この仮定が崩れることを示します。マルチモーダル推論では、デコード時のクエリがプリフィル段階の表現よりも大幅に大きな分散を示し、厳しいキャッシュ予算の下では KV 重要度の推定が不安定になります。その結果、小さなランク付けエラーにより、意味的に重要なビジュアル トークンが不当に破棄され、根拠と推論のパフォーマンスが低下する可能性があります。私たちは、トレーニング不要でデコードを認識し、厳密にプレフィル専用の KV 選択方法である MM-ShiftKV を提案します。 MM-ShiftKV は、分散拡張クエリ プロキシを構築することでプレフィル中のデコード時のクエリ動作を近似し、集計されたアテンション マスに基づいてプロンプト KV の重要性を推定します。マルチモーダル ベンチマークの実験では、厳格な KV キャッシュ バジェットの下で MM-ShiftKV が既存の方法よりも一貫して優れたパフォーマンスを発揮することが実証されています。私たちのコードは https://github.com/zjuDBxAI/MM-ShiftKV で入手できます。

原文 (English)

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens. Recent prefill-stage KV selection methods estimate KV importance from prefilling statistics, implicitly assuming that prefilling-time queries are representative of those encountered during decoding. We show that this assumption breaks down in multimodal inference, where decoding-time queries exhibit substantially larger variance than prefilling-stage representations, leading to unstable KV importance estimation under tight cache budgets. As a result, small ranking errors can disproportionately discard semantically critical visual tokens and degrade grounding and reasoning performance. We propose MM-ShiftKV, a training-free, decode-aware and strictly prefill-only KV selection method. MM-ShiftKV approximates decoding-time query behavior during prefilling by constructing variance-expanded query proxies and estimates prompt KV importance based on their aggregated attention mass. Experiments on multimodal benchmarks demonstrate that MM-ShiftKV consistently outperforms existing methods under strict KV-cache budgets. Our code is available at https://github.com/zjuDBxAI/MM-ShiftKV.

13:00 JSTLLM/生成AILlama

TriSP: 大規模言語モデルのトライシグナル構造化プルーニング

大規模言語モデル (LLM) は、さまざまなタスクにわたって優れたパフォーマンスを実現しますが、その展開はパラメーターのメモリとコンピューティング コストによって制限されます。構造化枝刈りは、アテンション ヘッドや多層パーセプトロン (MLP) ニューロンなどの構造全体を削除して、標準のハードウェアで効率的に実行される小型の高密度モデルを生成することでこの問題に対処します。ただし、既存の方法は、メモリを禁止する勾配ベースの重要度推定、または損失に対する除去の効果を直接測定しない活性化ベースの統計的プロキシのいずれかに依存しています。さらに、重要性基準と枝刈り後の回復戦略との間の相互作用は体系的に研究されていません。我々は、TriSP (Tri-Signal Structured Pruning) を提案します。これは、活性化ノルムによってスケーリングされた重みの大きさと幾何平均による一次勾配感度を組み合わせた重要度メトリックであり、構造信号と損失感度信号の両方を捕捉するチャネルレベルのスコアを生成します。適応的なレイヤごとの予算割り当てと低ランク適応 (LoRA) リカバリと組み合わせることで、TriSP はテストされたすべての構成にわたって最低のパープレキシティと最高のゼロショット精度を達成し、LLaMA-7B の 20% プルーニングで 6.80 WikiText-2 パープレキシティに達します。競争力のあるパフォーマンスを維持しながら、推論のスループットは 50% プルーニングで 82% 向上します。

原文 (English)

TriSP: Tri-Signal Structured Pruning for Large Language Models

Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their parameters. Structured pruning addresses this by removing entire structures such as attention heads and Multi-Layer Perceptron (MLP) neurons to produce smaller dense models that run efficiently on standard hardware. However, existing methods rely on either gradient-based importance estimation, which is memory-prohibitive, or activation-based statistical proxies, which do not directly measure the effect of removal on the loss. Furthermore, the interaction between the importance criterion and the post-pruning recovery strategy has not been systematically studied. We propose TriSP (Tri-Signal Structured Pruning), an importance metric that combines weight magnitude scaled by activation norm with first-order gradient sensitivity via a geometric mean, producing a channel-level score that captures both structural and loss-sensitivity signals. Combined with adaptive per-layer budget allocation and low-rank adaptation (LoRA) recovery, TriSP achieves the lowest perplexity and highest zero-shot accuracy across all tested configurations, reaching 6.80 WikiText-2 perplexity at 20% pruning on LLaMA-7B. Inference throughput improves by 82% at 50% pruning, while still maintaining competitive performance.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

ParBench: LLM 並列コード変換の信頼性の高い評価のためのベンチマーク

最新のコンピューティング集約型ソフトウェアは、アクセラレータ、プログラミング API、コンパイラ スタック、CUDA、OpenMP、OpenCL、OpenMP ターゲット オフロードなどの移植性レイヤーの変化するエコシステム全体で移行する必要があります。このような移行のために、大規模な言語モデルと自律コーディング エージェントがますます提案されていますが、この分野には、スレッド インデックス付け、同期、メモリ管理、ホスト デバイスの調整、API 固有の実行構造など、翻訳を動作的に有効にする低レベルの並列セマンティクスが保持されているかどうかを測定する信頼できる方法がありません。ここでは、実行可能で再現可能な条件下で LLM ベースの並列 API 変換を評価するためのカーネル中心のベンチマーク フレームワークである ParBench を紹介します。 ParBench は、宣言的なベンチマーク仕様を通じて周囲の構築、実行、検証インフラストラクチャを修正し、計算カーネルのみを変換するようにモデルに要求します。複数のオープンソース HPC スイートを利用し、CUDA、OpenMP、OpenCL、OpenMP ターゲット オフロード間の代表的なクロス API 変換の方向をカバーします。成功が表面的な形式の記憶ではなく堅牢な翻訳を反映しているかどうかをテストするために、ParBench には、AST 主導の、意図された動作を保持する、ベースライン検証済みのソース拡張機能が含まれています。最先端のオープンおよび独自の LLM の評価では、方向の非対称性、複数ファイルの調整、不完全な API 適応、ソースレベルの摂動に対する不均一な堅牢性など、信頼性の高い並列コード変換に対する永続的な障壁が示されています。コードは https://github.com/Scientific-Computing-Lab/ParBench で入手できます。

原文 (English)

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, including CUDA, OpenMP, OpenCL, and OpenMP target offload. Large language models and autonomous coding agents are increasingly proposed for such migration, but the field lacks reliable ways to measure whether they preserve the low-level parallel semantics that make translations behaviorally valid, including thread indexing, synchronization, memory management, host-device coordination, and API-specific execution structure. We present ParBench, a kernel-centric benchmark framework for evaluating LLM-based parallel API translation under executable, reproducible conditions. ParBench fixes the surrounding build, run, and verification infrastructure through declarative benchmark specifications and asks models to translate only the computational kernels. It draws on multiple open-source HPC suites and covers representative cross-API translation directions among CUDA, OpenMP, OpenCL, and OpenMP target offload. To test whether success reflects robust translation rather than surface-form memorization, ParBench includes AST-driven, intended behavior-preserving, baseline-validated source augmentation. Evaluations on state-of-the-art open and proprietary LLMs show persistent barriers to reliable parallel code translation, including direction asymmetry, multi-file coordination, incomplete API adaptation, and uneven robustness to source-level perturbations. Code is available at https://github.com/Scientific-Computing-Lab/ParBench.

13:00 JSTエージェント

大規模言語モデルによって調整された未知の環境での語彙発見

未知の環境(惑星や深海の探査など)に配置された自律エージェントの集団は、人間の言語では名前のない存在を参照するための共有語彙を開発する必要があります。我々は、神経記号語彙発見(NSLD)フレームワークを提案します。このフレームワークでは、LLMベースのエージェントの集団が、分布外の視覚指示対象をめぐって参照ゲームをプレイし、共有エイリアン語彙を自律的に自己組織化します。各エージェントは、凍結された CLIP ビジョン エンコーダとプライベート FAISS ベクトル インデックスおよびテキスト専用 LLM を組み合わせます。重要なことに、発見された異質な単語は、埋め込み空間における意味論的な近接性によって自然言語に固定され、知覚的に根拠のある新しい単語で人間の語彙を拡大します。最大 20 個のエージェントと 10 個の視覚的指示対象からなる母集団を使用したシミュレーションで合意に達します。収束ダイナミクスは、R^2 > 0.95 を達成する 3 つの解析モデルを通じて特徴付けられ、自律探査ミッションにおける展開前計画への第一歩を表しています。

原文 (English)

Lexical discovery in unknown environments orchestrated by Large Language Models

Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies to refer to entities that have no name in any human language. We propose the Neuro-Symbolic Lexical Discovery (NSLD) framework, in which a population of LLM-based agents plays a referential game over out-of-distribution visual referents, autonomously self-organising a shared alien lexicon. Each agent combines a frozen CLIP vision encoder with a private FAISS vector index and a text-only LLM. Crucially, discovered alien words are anchored to natural language via semantic proximity in the embedding space, enlarging the human vocabulary with new perceptually grounded words. Consensus is reached in simulations with populations of up to twenty agents and ten visual referents. Convergence dynamics are characterised through three analytical models achieving R^2 > 0.95, representing a first step towards pre-deployment planning in autonomous exploration missions.

13:00 JSTLLM/生成AI

スケールを超えた構造: RAG のスキーマ制約付き因果グラフ

グラフベースの検索拡張生成 (GraphRAG) は構造化された知識に基づいて答えを導き出しますが、現在のシステムはエンティティと関係を徹底的に抽出し、そのサイズと構築コストがクエリに必要な推論ではなくコーパスの長さに応じて変化するグラフを生成します。 HCG-RAG (階層因果グラフ RAG) を導入します。これは、オープンエンド抽出をスキーマ制約付き因果グラフに置き換えます。自動パイプラインは、コーパスを因果変数の固定された型付き語彙に抽出し、その上にコンパクトな 2 層グラフを具体化します。当社のスキーマ制約付きグラフは、数分の一のコストで応答品質に関してエンティティ関係ベースラインと一致します。最も LLM 負荷の高いベースライン (MS-GraphRAG) に比べて、ノード数が 3 ~ 20 分の 1、ビルド時間の LLM 呼び出しが 8 分の 1 ~ 135 分の 1 で、グラフはドメインの専門家が監査、修正、拡張できるほど十分にコンパクトです。神経科医によって検証されたてんかんデータセットを含む医学的および臨床的ベンチマークでは、HCG-RAG は最良の実体関連システムと同等またはそれを上回っています。アブレーションは因果関係グラフを構造化検索フィルターとして分離し、埋め込みのみの検索に比べて +6 パーセンテージ ポイント (pp) の向上に貢献します。発見可能な階層的因果構造を持つすべてのドメインにわたって、より高いレベルの組織化を課すメソッドのみがフラットなエンティティ関係の検索よりも優れており、グラフに含まれるノードの数よりもグラフに配置されるものが重要であることを示しています。

原文 (English)

Structure Over Scale: Schema-Constrained Causal Graphs for RAG

Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships exhaustively, producing graphs whose size and construction cost scale with corpus length rather than with the reasoning a query requires. We introduce HCG-RAG (Hierarchical Causal Graph RAG), which replaces open-ended extraction with schema-constrained causal graphs: an automated pipeline distills a corpus into a fixed, typed vocabulary of causal variables and materializes a compact two-tier graph over it. Our schema-constrained graphs match entity-relation baselines on answer quality at a fraction of the cost: 3-20x fewer nodes, 8x-135x fewer build-time LLM calls than the most LLM-intensive baseline (MS-GraphRAG), and graphs compact enough for a domain expert to audit, correct, and extend. On medical and clinical benchmarks, including a neurologist-validated epilepsy dataset, HCG-RAG matches or exceeds the best entity-relation systems. An ablation isolates the causal graph as a structured retrieval filter, contributing +6 percentage points (pp) over embedding-only retrieval. Across all domains with discoverable hierarchical causal structure, only methods imposing higher-level organization outperform flat entity-relation retrieval, indicating that what is placed in the graph matters more than how many nodes it contains.

13:00 JST研究/論文

xMIx: 機械的解釈アプリのための高性能の提供時間プラットフォーム

機械的解釈可能性 (MI) は、推論計算を分析して介入するための強力なアプローチとして浮上しており、脱獄試行の検出、真実性評価、幻覚検出などの用途が増えています。残念ながら、ほとんどの既存の MI フレームワークでは実行時に法外に高いオーバーヘッドが発生するため、運用モデルを提供するシステムに MI を導入することは現時点では現実的ではありません。根本的な問題は、MI 機能が提供モデルできれいに構成されていないことです。MI 機能はデプロイメントを断片化し、多くの場合、リクエストのドレインとサービング状態の再構築を強制し、運用環境のデプロイメントに不可欠な連続バッチ処理や CUDA グラフ実行などの重要なパフォーマンスの最適化と競合します。本番推論サービス環境に MI アプリケーションをデプロイするためのサービングネイティブ フレームワークである xMIx を紹介します。 xMIx を使用すると、MI 関数をモデル ランタイム内の事前定義された一連の位置にアタッチし、レイヤーおよび残差ストリーム内のアクティブ化に割り込むことができます。 xMIx は、前のモデル層の出力に応じて MI 関数の条件付き呼び出しをサポートします。複数の MI アプリケーションを 1 つのモデル インスタンスにデプロイできます。 xMIx は、それらをすべてサービング パスにコンパイルしますが、必要な場合にのみ実行時に動的にアクティブ化します。パフォーマンス コストは無視でき、別のモデル インスタンスや代替実行スタックは必要ありません。私たちは xMIx を vLLM サービング システムと統合し、3 つの主要なモデルと 7 つの多様な MI アプリケーションにわたって評価します。 xMIx は、ネイティブ vLLM 実行と同等のパフォーマンスを達成しますが、平均トークン間レイテンシ (ITL) で 1.3%、テール P99 ITL で 1.2%、最初のトークンまでの平均時間 (TTFT) で 2.6%、平均総トークン スループット (TTT) で 1.6% の速度低下が発生します。

原文 (English)

xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps

Mechanistic interpretability (MI) has emerged as a powerful approach for analyzing and intervening in inference computations, with a growing number of applications such as jailbreak attempt detection, truthfulness evaluation, and hallucination detection. Unfortunately, MI deployment in production model-serving systems is currently not practical, as most existing MI frameworks introduce prohibitively high runtime overheads. The fundamental problem is that MI functions do not compose cleanly with served models: they fragment deployment, often force draining requests and rebuilding serving state, and conflict with critical performance optimizations such as continuous batching and CUDA-graph execution, essential for production deployments. We present xMIx, a serving-native framework for deploying MI applications in production inference serving environments. xMIx enables attaching MI functions to a predefined set of locations in the model runtime, interposing on activations within the layers and residual streams. xMIx supports conditional invocation of MI functions depending on the outputs in preceding model layers. Multiple MI applications can be deployed in a single model instance. xMIx compiles them all into the serving path but activates them dynamically at runtime only when necessary, with negligible performance cost, and without requiring a separate model instance or alternative execution stack. We integrate xMIx with the vLLM serving system and evaluate it across three major models and seven diverse MI applications. xMIx achieves performance comparable to native vLLM execution, incurring a slowdown of 1.3% mean inter-token latency (ITL), 1.2% for tail P99 ITL, 2.6% for mean time to first token (TTFT), and 1.6% for mean total token throughput (TTT).

13:00 JSTエージェント

アトミスティックなシミュレーションのエージェントによるオーケストレーション

原子論的シミュレーションは材料設計の中心ですが、その実行には複雑な複数ステップのワークフローが含まれ、人間の高度な専門知識が必要です。ここでは、原子シミュレーションの設計、実行、検証を自動化する URSA (Universal Research and Scientific Agent) フレームワーク内に組み込まれたエージェント ベースのシステムを紹介します。このシステムは、大規模原子/分子超並列シミュレーター (LAMMPS) ツールを使用して実証されています。私たちのシステムは、原子間ポテンシャルを自律的に選択し、シミュレーションを構築して実行し、閉ループのワークフロー内で反復的なエラー回復を実行します。 LAMMPS 用の高スループット ツールキットである LAVA および Vienna Ab initio Simulation Package (VASP) 計算に対して出力をベンチマークすることにより、エージェントの科学的信頼性を評価します。私たちのフレームワークは手動介入と試行錯誤を減らし、それによってアトミスティックモデリングの厳密さ、再現性、拡張性を向上させます。

原文 (English)

An Agentic Orchestration of Atomistic Simulations

Atomistic simulations are central to materials design, but their execution involves complex, multi-step workflows that require significant human expertise. Here, we present an agent-based system embedded within the URSA (Universal Research and Scientific Agent) framework that automates the design, execution, and validation of atomistic simulations, demonstrated using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) tool. Our system autonomously selects interatomic potentials, constructs and runs simulations, and performs iterative error recovery within a closed-loop workflow. We evaluate the scientific reliability of the agent by benchmarking its outputs against LAVA, a high-throughput toolkit for LAMMPS and the Vienna Ab initio Simulation Package (VASP) calculations. Our framework reduces manual intervention and trial-and-error, thereby improving the rigor, reproducibility, and scalability of atomistic modeling.

13:00 JST研究/論文

HyCE-RAG: 説明可能なマルチホップ質問応答のためのハイパーグラフ証拠チェーン検索拡張生成

マルチホップ質問応答では、複数の文書から証拠を検索し、散在する事実を一貫した推論プロセスに結び付けるシステムが必要です。標準的な検索拡張生成 (RAG) は、主にクエリとテキスト チャンク間の意味論的な類似性に依存しているため、エンティティ、事実、および証拠単位間の構造的関係をモデル化できないことがよくあります。グラフベースの RAG は、グラフ構造の知識を導入することでこれを改善しますが、複数のエンティティとコンテキストを含む高次の関連性を表すペアワイズ エッジには依然として制限があります。我々は、説明可能なマルチホップ質問応答のためのハイパーグラフ証拠連鎖検索拡張生成フレームワークである HyCE-RAG を提案します。 HyCE-RAG は、エンティティ、関係、およびコンテキスト証拠をハイパーエッジに編成し、クエリ対応証拠ハイパーグラフを構築し、エンティティ - ハイパーエッジ発生構造にわたって信頼度伝播を実行します。次に、信頼度に基づいた証拠アセンブリを使用して、回答を生成する前に証拠パスを選択、接続、ランク付けします。スコアリング プロセスでは、意味的関連性、エンティティの接続性、証拠の網羅性、関係の信頼性、抽出の信頼性、および伝播された信頼性が共同で考慮されます。 HyCE-RAG は、平坦に検索された文章ではなく、構造化された証拠チェーンを言語モデルに提供することで、より忠実で解釈可能な推論をサポートします。 HotpotQA、2WikiMultihopQA、MuSiQue、および 2 つの GraphRAG-Bench サブセットの実験では、HyCE-RAG が、回答の精度、コンテキストの関連性、および忠実性において、標準の RAG およびグラフベースの RAG ベースラインよりも一貫して優れていることが示されています。これらの結果は、ハイパーグラフに基づく証拠の整理が、複雑な質問応答における検索後の推論の有望な方向性であることを示唆しています。

原文 (English)

HyCE-RAG: Hypergraph Chain-of-Evidence Retrieval-Augmented Generation for Explainable Multi-hop Question Answering

Multi-hop question answering requires systems to retrieve evidence from multiple documents and connect scattered facts into a coherent reasoning process. Standard retrieval-augmented generation (RAG) mainly relies on semantic similarity between a query and text chunks, and therefore often fails to model structural relations among entities, facts, and evidence units. Graph-based RAG improves this by introducing graph-structured knowledge, but pairwise edges are still limited in representing higher-order associations involving multiple entities and contexts. We propose HyCE-RAG, a Hypergraph Chain-of-Evidence Retrieval-Augmented Generation framework for explainable multi-hop question answering. HyCE-RAG organizes entities, relations, and contextual evidence into hyperedges, builds a query-aware evidence hypergraph, and performs confidence propagation over entity--hyperedge incidence structures. It then uses confidence-guided evidence assembly to select, connect, and rank evidence paths before answer generation. The scoring process jointly considers semantic relevance, entity connectivity, evidence coverage, relation reliability, extraction confidence, and propagated confidence. By providing the language model with structured evidence chains rather than flat retrieved passages, HyCE-RAG supports more faithful and interpretable reasoning. Experiments on HotpotQA, 2WikiMultihopQA, MuSiQue, and two GraphRAG-Bench subsets show that HyCE-RAG consistently outperforms standard RAG and graph-based RAG baselines in answer accuracy, context relevance, and faithfulness. These results suggest that hypergraph-based evidence organization is a promising direction for post-retrieval reasoning in complex question answering.

13:00 JST研究/論文

時系列予測のための不確実な成分への拡散軌跡の差異化

拡散モデルは、観測された履歴に基づいて将来の値の分布をモデル化する、確率的時系列予測のフレームワークとして広く使用されています。しかし、時系列予測では、未来は観察された歴史を継続し、標準的な拡散プロセスが対処されないまま放置される非対称性を生み出します。ゆっくりと変化する内容は主に観察された連続性によって決まりますが、高周波のダイナミクスが残留不確実性のほとんどを担います。既存の拡散ベースの予測者は、生成前に外部ルールを通じてこの非対称性を切り離し、歴史がすでにターゲットのどの部分に定着しているのかを汚職の軌跡が見えないままにしておきます。我々は、この予測可能性の非対称性を拡散軌道自体に埋め込む拡散フレームワークである DiffDiff を提案します。これにより、単一のエンドツーエンドの拡散プロセスが、ターゲットのどの部分に履歴がすでに固定されているかを認識できるようになります。 DiffDiff はフォワード オペレーターをステップ依存にするため、ノイズの多い中間状態がターゲット自体から 2 次差分構造に向かって徐々にシフトします。一方、コンディショニング パスウェイは、各拡散ステップでステージ適応ゲートによってバランスがとれた値領域と微分履歴情報の両方をデノイザーに供給します。端末分布は標準のガウス分布に近づき、既存のサンプラーとの互換性が維持されます。 4 つの予測期間にわたる 7 つのベンチマークで、DiffDiff は 6 つの拡散ベースラインを上回っており、私たちの分析では、DiffDiff が履歴に固定されたコンテンツの再構築を軽減しながら、ターゲットの最も不確実なコンポーネントに拡散の生成努力を集中させていることが確認されています。

原文 (English)

Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting

Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history. In time series forecasting, however, the future continues the observed history, creating an asymmetry the standard diffusion process leaves unaddressed, with slowly-varying content largely determined by the observed continuity while higher-frequency dynamics carry most of the residual uncertainty. Existing diffusion-based forecasters decouple this asymmetry through an external rule before generation, leaving the corruption trajectory blind to which parts of the target the history can already anchor. We propose DiffDiff, a diffusion framework that embeds this predictability asymmetry into the diffusion trajectory itself, so that a single end-to-end diffusion process becomes aware of which parts of the target the history can already anchor. DiffDiff makes the forward operator step-dependent so that the noisy intermediate state progressively shifts from the target itself toward its second-order differenced structure, while a conditioning pathway supplies the denoiser with both value-domain and differential history information balanced by a stage-adaptive gate at each diffusion step. The terminal distribution approaches a standard Gaussian, preserving compatibility with existing samplers. On seven benchmarks across four prediction horizons, DiffDiff outperforms six diffusion baselines, and our analysis confirms that DiffDiff concentrates the diffusion's generative effort on the most uncertain components of the target while relieving it from rebuilding the history-anchored content.

13:00 JST研究/論文

視覚言語モデルにおけるチャートの欺瞞: 脆弱性から緩和まで

情報の視覚化は、パターン、傾向、外れ値を伝えるために広く使用されていますが、軸の切り詰めや反転、歪んだアスペクト比、不適切なエンコーディング、誤解を招くカラーマッピングなどの欺瞞的な設計選択により、基礎となるデータを維持しながら解釈を体系的に変更する可能性があります。ビジョン言語モデル (VLM) がチャートの理解や分析的推論に使用されることが増えているため、信頼できるデータ分析には、そのような欺瞞的な視覚化に対する堅牢性を評価することが重要になっています。誤解を招くチャート設計に対する VLM の堅牢性を評価するための最初の制御されたペア ベンチマークである VisDeception を紹介します。このベンチマークには、欺瞞的な視覚化戦術の 8 つの主要カテゴリにわたる 1,600 個の忠実なチャートと誤解を招くチャートのペアが含まれており、各誤解を招くチャートは、同じ基礎データから生成された忠実な対応するチャートとペアになっています。ベースラインのチャート理解エラーから欺瞞に起因する推論エラーを分離するために、誤解を招くビジュアライゼーションがモデルの応答をデータの忠実な解釈からどのように遠ざけるかを定量化するペアの評価指標である欺瞞スコアを導入します。 10 の最先端の VLM からの 32,000 件の回答を調べた結果、高度なモデルであっても依然として欺瞞的な視覚操作に対して非常に脆弱であることがわかりました。堅牢性を向上させるために、回答生成前に視覚化から抽出された構造化チャートのメタデータに推論を根拠付ける推論時のマルチエージェント軽減フレームワークをさらに提案します。これにより、明示的なユーザー指示を必要とせずに、モデルが欺瞞的な視覚的手がかりの影響を低減できるようになります。まとめると、私たちの調査結果は、現在のチャート理解システムにおける重要な信頼性のギャップを明らかにし、ベンチマーク主導の評価、欺瞞を認識するメトリクス、および構造化された推論を、ビジュアル分析用のより信頼できる VLM を開発するための有望な方向性として確立します。

原文 (English)

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data. As Vision-Language Models (VLMs) are increasingly used for chart understanding and analytical reasoning, assessing their robustness to such deceptive visualizations has become critical for trustworthy data analysis. We introduce VisDeception, the first controlled paired benchmark for evaluating the robustness of VLMs to misleading chart designs. The benchmark contains 1,600 paired faithful and misleading charts spanning eight major categories of deceptive visualization tactics, where each misleading chart is paired with a faithful counterpart generated from the same underlying data. To isolate deception-induced reasoning errors from baseline chart-understanding errors, we introduce the Deception Score, a paired evaluation metric that quantifies how misleading visualizations shift model responses away from the faithful interpretation of the data. Across 32,000 responses from 10 state-of-the-art VLMs, we find that even advanced models remain highly vulnerable to deceptive visual manipulations. To improve robustness, we further propose an inference-time multi-agent mitigation framework that grounds reasoning in structured chart metadata extracted from the visualization before answer generation, enabling models to reduce the influence of deceptive visual cues without requiring explicit user instructions. Together, our findings reveal important reliability gaps in current chart-understanding systems and establish benchmark-driven evaluation, deception-aware metrics, and structured reasoning as promising directions for developing more trustworthy VLMs for visual analytics.

13:00 JST研究/論文GPT / ChatGPTDeepSeek

DeepLo​​ok: 先読みによるより深い思考

推論時間スケーリングは、大規模な言語モデル推論を改善するための強力なパラダイムとして浮上しており、多くの場合、パラメーター スケーリングのみよりも困難な推論タスクで大きな利益をもたらします。ただし、既存のアプローチでは、推論トレース内でのコンピューティングの割り当て方法が依然として非効率的です。推論の失敗は、間違った答えが明らかになる前に、不確実性が早期に現れることが多いという観察に動機付けられて、不確実性のボトルネックでの先読み計算を集中させる、トレーニング不要の監視と介入によるデコード フレームワークである DeepLo​​ok を紹介します。 DeepLo​​ok は、トークン レベルの信頼度をセグメント レベルのシグナルに集約し、最近の履歴と比較して信頼度が低下したときにトリガーし、固定ホライズンの先読みで候補の継続を探索します。ブランチは、ロールアウト継続に対するセグメントレベルの平均信頼度である平均先読み信頼度 (ALC) によってランク付けされ、投票によって枝刈りおよび集計されます。 DeepSeek-R1-8B、Qwen3-32B、GPT-OSS-20B、および GPT-OSS-120B にわたる 4 つの競争形式の数学ベンチマークで、DeepLook は精度、つまりトークンコストのパレート フロンティアをシフトします。16 設定中 11 設定で DeepConf-low よりも精度が向上し、同時にデータセット レベルのトークン生成を平均 87.3% 削減します (AIME25 での +3.1 の向上を含む)。 GPT-OSS-20B を搭載した BRUMO25 では Qwen3-32B と +8.8。これらの結果は、選択的で将来を意識した介入の方が、完全な推論軌跡を均一にスケーリングするよりも、コストとのトレードオフにおいて、大幅に高い精度を生み出すことを示しています。コードはここから入手できます。

原文 (English)

DeepLook: Deeper Thinking with Lookahead

Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter scaling alone. However, existing approaches remain inefficient in how compute is allocated within a reasoning trace. Motivated by the observation that reasoning failures often exhibit an early onset of uncertainty before a wrong answer become explicit, we introduce DeepLook, a training-free monitor-and-intervene decoding framework that concentrates lookahead compute at uncertainty bottlenecks. DeepLook aggregates token-level confidence into segment-level signals, triggers when confidence drops relative to recent history, and explores candidate continuations with fixed-horizon lookahead. Branches are ranked by Average Lookahead Confidence (ALC), the average segment-level confidence over rollout continuations, then pruned and aggregated through voting. On four competition-style mathematics benchmarks across DeepSeek-R1-8B, Qwen3-32B, GPT-OSS-20B, and GPT-OSS-120B, DeepLook shifts the accuracy--token-cost Pareto frontier: it improves accuracy over DeepConf-low in 11 of 16 settings while reducing dataset-level token generation by 87.3% on average, including gains of +3.1 on AIME25 with Qwen3-32B and +8.8 on BRUMO25 with GPT-OSS-20B. These results show that selective, future-aware intervention yields substantially stronger accuracy--cost trade-offs than uniformly scaling complete reasoning trajectories. Code is available here.

13:00 JSTLLM/生成AI

パーソナライズされたマルチモーダル大規模言語モデルにおけるグループの好みの崩壊

パーソナライズされたマルチモーダル大規模言語モデル (MLLM) は、ユーザー固有の応答を生成することを目的としていますが、既存の方法は主にプロファイル レベルの情報に依存しており、多様なユーザーの好みを見落としています。我々は、マルチユーザーのパーソナライズされたMLLMが、生成中に抑制された嗜好シグナルと信頼性の低い嗜好の使用により、個人の嗜好に鈍感になり、支配的な集団レベルの選択に向かうグループ嗜好崩壊を特定します。我々は、安定したプロファイル情報を好みに関連する表現から分離する好み中心のフレームワークである PrefMoE を提案します。 PrefMoE は、プリファレンスを共有プロトタイプとパーソナライズされた残差に分解し、不均衡を意識した学習、反事実的な擬似ユーザー拡張、および残差の非相関化によって個別の残差を保存し、プロファイルとプリファレンス要素を個別の LoRA 適応パスにルーティングします。複数の MLLM バックボーンにわたる実験では、PrefMoE が好みの崩壊を大幅に軽減しながら、好みに応じたパーソナライゼーションを向上させることが示されています。プロジェクトページ: https://prefmoe.github.io/。

原文 (English)

Group Preference Collapse in Personalized Multimodal Large Language Models

Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population-level choices due to suppressed preference signals and unreliable preference use during generation. We propose PrefMoE, a preference-centric framework that separates stable profile information from preference-related representations. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance-aware learning, counterfactual pseudo-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths. Experiments across multiple MLLM backbones show that PrefMoE improves preference-sensitive personalization while substantially reducing preference collapse. Project page: https://prefmoe.github.io/.

13:00 JSTLLM/生成AIGPT / ChatGPTQwen

動的システムの解釈可能なコントローラーとしての LLM の評価

大規模言語モデル (LLM) は、意思決定や推論タスクにますます使用されていますが、物理システムのコントローラーとしての可能性はほとんど解明されていません。この研究では、LLM が動的熱環境の解釈可能なコントローラーとして機能できるかどうかを調査し、設定値に従い、自然言語コマンドを解釈し、アクチュエーターの効果について推論し、事前のモデルベースの知識を組み込む能力を調べます。さまざまなスケールの 5 つの LLM が、ヒーターまたはファンの使用にペナルティを伴う設定や、モデルが物理ベースの予測ツールにアクセスできるケースなど、複数のシナリオの下で評価されます。結果は、制御性能がモデルの複雑さに依存することを示しています。低スケールおよび中スケールのモデルでは、頻繁にアクチュエータのダイナミクスを誤解したり、一貫性のない推論が生成されますが、Qwen-3~14B や GPT-4o などの高複雑なモデルでは、正確な温度追跡、安定したアクチュエータの使用、および物理原理に沿った一貫した説明が実現されます。物理ベースのモデルを組み込むことで、予測的な意思決定が可能になり、制御の滑らかさとエネルギー効率が大幅に向上します。さらに、詳細な推論分類法により、より小規模なモデルでの因果関係の誤解から、より大きなモデルでの一貫性のある時間的に認識された推論への明確な進展が明らかになります。この調査結果は、十分な能力があり、ドメイン知識に適切に基づいている場合、LLM が解釈可能なコントローラーとして機能できることを示しており、妥当な説明を提供できるハイブリッド モデルベースおよび言語駆動型の制御戦略の有望な機会を強調しています。

原文 (English)

Evaluating LLMs as Interpretable Controllers for Dynamical Systems

Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical systems remains largely unexplored. This work investigates whether LLMs can function as interpretable controllers for a dynamic thermal environment, examining their ability to follow setpoints, interpret natural-language commands, reason about actuator effects, and incorporate prior model-based knowledge. Five LLMs of varying scales are evaluated under multiple scenarios, including settings with penalties on heater or fan usage and cases where the models have access to a physics-based prediction tool. The results show that control performance depends on model complexity: while low- and mid-scale models frequently misinterpret actuator dynamics or generate inconsistent reasoning, high-complexity models such as Qwen-3~14B and GPT-4o achieve accurate temperature tracking, stable actuator usage, and coherent explanations aligned with physical principles. Incorporating a physics-based model significantly improves control smoothness and energy efficiency by enabling anticipatory decision-making. A detailed reasoning taxonomy further reveals a clear progression from causal misinterpretation in smaller models to cohesive and temporally aware reasoning in larger ones. The findings demonstrate that LLMs can act as interpretable controllers when sufficiently capable and appropriately grounded in domain knowledge, highlighting promising opportunities for hybrid model-based and language-driven control strategies that can provide plausible explanations.

13:00 JSTLLM/生成AIエージェント

Tokengeist: エージェント会話における複数ターンの属性追跡

言語モデルが複数ターンの会話で応答を生成するとき、前のターンのどのトークンがその応答を形成し、それらの依存関係が前のターンにどのように伝播したのでしょうか?既存のコンテキスト アトリビューション手法は、シングル パスで完全なコンテキストを処理し、表面レベルの依存関係を回復しますが、現実世界の対話や複数ステップの推論タスクの階層的で非線形な構造が欠落しています。マルチターン コンテキスト アトリビューション (MTCA) を導入します。これは、モデル応答のターゲット スパンを指定して、ターン全体でアトリビューションを逆方向に追跡して、以前のどのターンが直接関連していたかだけでなく、それらのターン自体が以前のコンテキストにどのように依存していたかを特定するタスクです。私たちは、会話ターンにわたる有向非巡回グラフ (DAG) の再帰的走査として属性をキャストすることによって完全な依存関係パスを回復する、属性メソッドに依存しないスケーラブルなフレームワークである Tokengeist を提案します。私たちは、665 のマルチターン会話にわたる 3,845 のターゲット スパンのベンチマークである MTCABench をリリースします。これには、4 つの依存関係タイプにわたって、最大 14 の深さに達する金の出所グラフで注釈が付けられます。 4 つのオープンウェイト モデル全体で、フラット アトリビューション手法はマルチホップ依存関係を回復できず、ソース再現率は 20% 未満に達しますが、Tokengeist は 90% に達します。私たちの結果は、シングルパス アトリビューションの体系的な失敗モード (これを来歴崩壊と呼んでいます) を明らかにし、ターンをまたいで再帰的に推論するアトリビューション手法の動機付けとなります。

原文 (English)

Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations

When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies propagate across prior turns? Existing context attribution methods process the full context in a single pass, recovering surface-level dependencies but missing the layered, non-linear structure of real-world dialogues and multi-step reasoning tasks. We introduce multi-turn context attribution (MTCA): given a target span in a model response, the task of tracing attribution backward across turns to identify not only which prior turns were directly relevant, but also how those turns themselves depended on earlier context. We propose Tokengeist, an attribution-method-agnostic and scalable framework that recovers full dependency paths by casting attribution as a recursive traversal of a directed acyclic graph (DAG) over conversation turns. We will release MTCABench, a benchmark of 3,845 target spans across 665 multi-turn conversations, annotated with gold provenance graphs reaching depths of up to 14, across four dependency types. Across four open-weight models, flat attribution methods fail to recover multi-hop dependencies, achieving under 20% source recall, while Tokengeist reaches 90%. Our results reveal systematic failure modes of single-pass attribution -- which we term provenance collapse -- and motivate attribution methods that reason recursively across turns.

13:00 JSTエージェント

重要なインフラストラクチャのエージェント AI システムに対する分散型のきめ細かなアクセス制御

実稼働インフラストラクチャに自律型 AI エージェントを導入すると、従来のロールベースのアクセス制御 (RBAC) モデルでは対処できない基本的なセキュリティの課題が生じます。決定論的な自動化とは異なり、AI エージェントは確率的な動作を示すため、従来の信頼モデルでは重要なシステムへのアクセスを管理するには不十分です。このペーパーでは、重要なクラウド インフラストラクチャで動作するエージェント AI システム向けに特別に設計された分散型の多層アクセス コントロール アーキテクチャについて説明します。当社のフレームワークでは、4 つの主要な革新が導入されています。(1) エージェントのアクションを委任された人間の権限に結び付ける複合アイデンティティ モデル、(2) グローバル プラットフォーム アクセスからパラメータごとの制約まで 5 つの粒度レベルにわたる階層的権限システム、(3) ツール チームが権限境界を独立して管理する分散型ポリシー所有権モデル、(4) 自律エージェントによる高リスク操作の実行を防止する安全インターロックを備えた漸進的信頼エスカレーション。私たちは OWASP Top 10 for LLM Applications (2025) 脅威分類法に基づいて設計を行い、各アーキテクチャ上の決定が特定の攻撃ベクトルをどのように軽減するかを実証します。このシステムは、数百のデータセンターにわたるネットワーク インフラストラクチャを管理する大手クラウド プロバイダーの運用環境に導入されており、20 を超える専門 AI エージェントと 60 を超える決定論的プレイブックに対してきめ細かなアクセス制御を実施し、毎日数千件の操作を処理しながら、運用導入の 8 か月にわたって不正な書き込み操作がゼロであることを維持します。アクセス パターンの分布、拒否率、および非決定的アクターによる権限昇格を防ぐための階層型承認の有効性に関する経験的データを示します。

原文 (English)

Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure

The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access control (RBAC) models cannot address. Unlike deterministic automation, AI agents exhibit stochastic behavior, making conventional trust models insufficient for governing their access to critical systems. This paper presents a decentralized, multi-layered access control architecture designed specifically for agentic AI systems operating in critical cloud infrastructure. Our framework introduces four key innovations: (1) a compound identity model that binds agent actions to delegated human authority, (2) a hierarchical permission system spanning five granularity levels from global platform access to per-parameter constraints, (3) a decentralized policy ownership model where tool teams independently govern their authorization boundaries, and (4) progressive trust escalation with safety interlocks that prevent autonomous agents from executing high-risk operations. We ground our design in the OWASP Top 10 for LLM Applications (2025) threat taxonomy and demonstrate how each architectural decision mitigates specific attack vectors. Deployed in production at a major cloud provider managing network infrastructure across hundreds of datacenters, the system enforces granular access control for 20+ specialized AI agents and 60+ deterministic playbooks processing thousands of operations daily while maintaining zero unauthorized write operations over eight months of production deployment. We present empirical data on access pattern distributions, denial rates, and the effectiveness of layered authorization in preventing privilege escalation by non-deterministic actors.

13:00 JSTLLM/生成AIハードウェア/半導体

DynaResize: 分割された LLM ポストトレーニングのためのランタイム GPU 再割り当て

RL ベースの LLM ポストトレーニングでは、別々の GPU リソース間でロールアウトとトレーニングがますます細分化されますが、静的 GPU パーティショニングでは、ロングテール ロールアウト レイテンシの下で深刻なパイプライン バブルが発生します。 DynaResize は、RL セマンティクスを変更せずに、ロールアウトとトレーニングの間で GPU を動的に切り替えてステージの実行時間のバランスを取る、ランタイム GPU 再割り当てシステムです。 DynaResize は、サイズ変更をきめ細かい操作に分解し、コミュニケーターの再利用、制限された状態のステージング、およびヒステリシスベースのサイズ変更を通じて、起動にクリティカルではない作業をクリティカル パスから削除します。実験結果によると、DynaResize は、最適な静的構成と比較して、エンドツーエンドのスループットを 66.5% 向上させ、合計実行時間を 33% 削減し、同時にロール切り替えオーバーヘッドの 27% を隠すことができます。

原文 (English)

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.

13:00 JSTLLM/生成AI

Opti-Q: マルチ LLM 質問計画のための制約ベースの最適化フレームワーク

大規模言語モデル (LLM) は強力な質問応答 (QA) を可能にしますが、予算に基づいた導入は非決定性と異種リソース プロファイル (コスト、遅延、エネルギー) によって複雑になります。 OPTI-Q は、マルチ LLM オーケストレーションのための実行前計画パラダイムを実装する、データベースからインスピレーションを得たコストベースのオプティマイザーです。 OPTI-Q は、実行 DAG 内の物理演算子として LLM 呼び出しをモデル化し、質問ごとに、ユーザーが指定したリソース制約の下で財務コスト、レイテンシ、エネルギーをトレードオフしながら回答品質 (QoA) を最適化するプランを検索します。プランには、中間の回答をコンテキストとして渡す逐次演算子と、モデルを同時に実行して出力をマージする並列/ブレンド演算子を含めることができます。各候補プランを実行せずにこのスペースを検索するために、OPTI-Q は、ベンチマークと実行トレースからデータが入力および更新される統計カタログである PERFDB を使用して、個々のオペレーターと構成されたサブプランの両方の QoA とリソース コストを推定します。これらの推定値を使用して、OPTI-Q はパレート フロンティア検索を実行し、ユーザーの好みに基づいて最終的なプランを選択します。ユーザー指定の予算の下での MMLU-Pro および SimpleQA では、OPTI-Q は同等のコストでベースラインと比較して平均 QoA を最大 58% および最大 41% 改善し、データベース スタイルのプランニングがマルチ LLM QA の品質リソースのトレードオフを向上させることを示しています。

原文 (English)

Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning

While large language models (LLMs) enable strong question answering (QA), budgeted deployment is complicated by nondeterminism and heterogeneous resource profiles (cost, latency, and energy). We present OPTI-Q, a database-inspired, cost-based optimizer that implements a plan-before-execute paradigm for multi-LLM orchestration. OPTI-Q models LLM invocations as physical operators in an execution DAG and, for each question, searches for plans that optimize answer quality (QoA) while trading off financial cost, latency, and energy under user-specified resource constraints. Plans can include sequential operators that pass intermediate answers as context and parallel/blend operators that run models concurrently and merge their outputs. To search this space without executing each candidate plan, OPTI-Q uses PERFDB, a statistics catalog populated and refreshed from benchmarks and execution traces, to estimate the QoA and resource costs of both individual operators and composed subplans. Using these estimates, OPTI-Q performs Pareto-frontier search and selects a final plan based on user preferences. On MMLU-Pro and SimpleQA under user-specified budgets, OPTI-Q improves average QoA by ~58% and ~41% over baselines at comparable cost, demonstrating that database-style planning yields better quality-resource trade-offs for multi-LLM QA.

13:00 JST研究/論文NVIDIA

CHS-SQL: 信頼に基づくヒューリスティック検索スキーマ リンク プロセスに基づく Text-to-SQL アプローチ

最近、Text-to-SQL の分野で、トレーニングに Small Language Model (SLM) を利用する作品がいくつかあります。これらのアプローチは、単一の NVIDIA RTX 4090 GPU の計算能力のみを使用して、SQL の生成において大規模モデルに近いパフォーマンスを達成すると同時に、データのセキュリティも確保します。既存のメソッドのほとんどは、スキーマ リンク中に冗長なテーブルと列をフィルタリングして除外し、Text-to-SQL の精度を向上させます。ただし、候補スキーマ サブセットを選択するときに、精度と再現率のトレードオフは考慮されていません。私たちの調査では、スキーマ リンクの精度と再現率の両方が最終的な SQL の精度に直接影響することがわかりました。したがって、Text-to-SQL タスクで SLM を効率的に微調整するための新しいフレームワークである CHS-SQL を提案します。これは、精度と再現率のバランスをとるだけでなく、Text-to-SQL タスクの全体的なパフォーマンスも向上します。その主な革新はスキーマ リンク フェーズにあり、モデルの内部信頼性と組み合わせたヒューリスティック検索を使用して、最適な精度と再現率のトレードオフを実現します。この精巧なメカニズムにより、無関係なノイズを抑制しながら、生成された SQL クエリに関連するスキーマ候補の精度が最大化されます。同じ戦略が SQL 生成中にさらに適用され、SLM が局所最適に陥ることを回避しながら候補クエリを絞り込みます。私たちの方法では、SLM を介した Text-to-SQL タスクで最先端 (SOTA) の結果が得られます。

原文 (English)

CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process

Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches achieve performance close to that of large models in generating SQL, using only the computational power of a single NVIDIA RTX 4090 GPU, while also ensuring data security. Most existing methods filter out redundant tables and columns during Schema Linking to improve Text-to-SQL accuracy. However, they do not consider the precision-recall trade-off when selecting the candidate schema subset. Our research found that both the precision and recall of Schema Linking directly affect the final SQL accuracy. Therefore, we propose a novel framework for efficiently fine-tuning SLMs on Text-to-SQL tasks, CHS-SQL, that not only balances precision and recall but also improves overall performance on Text-to-SQL tasks. Its main innovation lies in the Schema Linking phase, where a heuristic search combined with model internal confidence is employed to achieve an optimal precision-recall trade-off. This elaborated mechanism maximizes the precision of relevant schema candidates for the generated SQL queries while suppressing irrelevant noise. The same strategy is further applied during SQL generation to refine candidate queries while helping the SLM to avoid trapping in a local optimum. Our method achieves state-of-the-art (SOTA) results on Text-to-SQL tasks via SLMs.

13:00 JSTLLM/生成AILlama

TokenMem: 凍結 LLM への忠実な知識の注入

検索拡張生成 (RAG) は、外部知識を使用して大規模言語モデル (LLM) を強化しますが、知識の競合に悩まされます。検索された情報がパラメトリック記憶と矛盾する場合、共有自己注意経路が予測不可能な出力を生成します。我々は、専用のクロスアテンション チャネルを通じて凍結された LLM に知識を注入し、残差ストリーム内のパラメトリック メモリとの競合を回避する軽量メモリ システムである TokenMem を紹介します。 TokenMem は、2 段階のカリキュラムを通じてシン ゲーティング アダプター ($\sim$3-7M パラメーター) のみをトレーニングします。最初に一般知識の活用を学習し、次に反事実知識の下で忠実なコンプライアンスを強化します。 3 つのファミリー (Qwen3-4B/8B/14B、LLaMA-3.1-8B、OLMo-3-7B) にまたがる 5 つのモデルの対照実験において、TokenMem は反事実ベンチマークで 69 ~ 70% のナレッジ コンプライアンス (KC) を達成しました。これに対し、バニラ RAG では 20 ~ 52% であり、その差は最大 49 パーセント ポイントです。アブレーション研究は、2 段階のカリキュラムが重要であることを示しています。段階 2 を削除すると、KC はほぼゼロに崩壊します。メカニズムの分析により、ゲート アダプターが明示的な監視なしで、競合を認識したレイヤー固有の注入戦略を学習することが明らかになりました。

原文 (English)

TokenMem: Faithful Knowledge Injection for Frozen LLMs

Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved information contradicts parametric memory, the shared self-attention pathway produces unpredictable outputs. We present TokenMem, a lightweight memory system that injects knowledge into frozen LLMs through a dedicated cross-attention channel, bypassing competition with parametric memory in the residual stream. TokenMem trains only a thin gating adapter ($\sim$3-7M parameters) via a two-phase curriculum: first learning general knowledge utilization, then strengthening faithful compliance under counterfactual knowledge. In controlled experiments on five models spanning three families (Qwen3-4B/8B/14B, LLaMA-3.1-8B, OLMo-3-7B), TokenMem achieves 69-70% Knowledge Compliance (KC) on counterfactual benchmarks, compared to 20-52% for vanilla RAG, a gap of up to 49 percentage points. Ablation studies show that the two-phase curriculum is critical: removing Phase 2 collapses KC to near-zero. Mechanistic analysis reveals that the gate adapter learns a conflict-aware, layer-specific injection strategy without explicit supervision.

13:00 JSTLLM/生成AI

マスクされた蒸留: 言語モデルにおける思考連鎖の内部化

大規模推論モデル (LRM) は、推論時に最終的な答えを生成する前に、中間ステップの長く明示的なチェーンを生成します。最終的な回答の正しさはトレースの正しさと因果関係がなく、トレースの長さは問題の複雑さの信頼できる指標ではありませんが、これらの中間トレースは、レイテンシ、メモリ使用量、およびサービス コストを支配します。これは当然の疑問です。これらの中間トークンで表現された計算を言語モデルのパラメーターに内部化して、答えを直接 (またははるかに短い中間トレースで) 生成できるようにすることはできるのでしょうか。 \textit{マスクされた蒸留} という知識蒸留フレームワークを導入します。このフレームワークでは、学生の LLM が質問に条件付けされた解決トークンのみを予測するように訓練され、推論教師が質問と自身の CoT トレースに条件付けした後の学生の応答に関するフィードバックを提供します。このフレームワークを 2 つの設定でインスタンス化します。(i) \textit{自己蒸留} 設定では、同じモデルが思考モードでは教師として、非思考モードでは生徒として機能します。(ii) \textit{デュアルモデル} 設定では、より大きな推論教師が、解決策トークンに対して別の小さな非思考の生徒を監督します。中間トークンを、推論モデルがソリューション トークンに適合させるために使用する足場として扱うことにより、学生が監視される中間トークンの足場の長さをさらに変更し、完全な内部化 (学生が解決策のみを出力する) と内部化なし (学生が解答の前に完全なトレースを出力する) の間を補間します。私たちは、GSM8K (小学校の算数) と Countdown (数字パズルの検索タスク) という 2 つの推論ドメインに関する制御された実験を通じてフレームワークを評価します。

原文 (English)

Masked Distillation: Internalizing the Chain-of-Thought in Language Models

Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though the final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises a natural question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers directly (or with much shorter intermediate traces)? We introduce \textit{masked distillation}, a knowledge-distillation framework in which a student LLM is trained to predict only the solution tokens conditioned on the question, while a reasoning teacher provides feedback on the student's responses after conditioning on the question and its own CoT trace. We instantiate this framework in two settings: (i) a \textit{self-distillation} setting, in which the same model serves as the teacher in thinking mode and as the student in non-thinking mode, and (ii) a \textit{dual-model} setting, in which a larger reasoning teacher supervises a separate smaller non-thinking student over the solution tokens. By treating intermediate tokens as a scaffold which reasoning models use to fit over the solution tokens, We additionally vary the length of intermediate-token scaffolding the student is supervised on, interpolating between full internalization (the student emits only the solution) and no internalization (the student emits the full trace before the answer). We evaluate the framework through controlled experiments on two reasoning domains: GSM8K (grade-school arithmetic) and Countdown (a number-puzzle search task).

13:00 JSTビジネス/資金調達GPT / ChatGPTGemini

VlogReward: Vlog編集のための多次元評価を学ぶ

パーソナライズされたストーリーテリング メディアとしての vlog の急速な台頭により、vlog 編集計画を評価および改良するための自動システムの需要が生じています。しかし、vlog の評価は非常に主観的であり、標準化された基準、データセットとベンチマーク、効果的な報酬モデルが不足しているため、依然として困難が伴います。これらの課題に対処するために、私たちはプロの vlog クリエイターとプロダクト マネージャーによる包括的な vlog 評価フレームワークを定義し、6 つの主要な側面 (創造性、一貫性、コンセプト デザイン、撮影、ナレーション、ペーシング) の分類を確立しました。続いて、マルチモーダル大規模言語モデル (MLLM) の vlog 報酬機能を評価するために、10 万件の vlog 編集からなる大規模なデータセットと専用ベンチマーク VRMBench を厳選しました。最後に、きめの細かい多次元スコアと反復的な改善のための実用的なフィードバックの両方を提供できる堅牢な vlog 報酬モデルである VlogReward を紹介します。技術的には、調整可能なグループ間比較報酬を導入することで、グループ相対ポリシー最適化 (GRPO) フレームワークを強化します。これにより、標準 GRPO の「方向盲目」問題が軽減され、モデルがさまざまな品質の編集をより適切に区別できるようになります。 VlogReward は、GPT-5 や Gemini-3-Pro などの既存の MLLM を大幅に上回る最先端の結果を実現します。私たちの研究が vlog 作成者に役立ち、vlog の自動評価および改良システムの促進に役立つことを願っています。

原文 (English)

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, i.e., Creativity, Consistency, Concept Design, Cinematography, Narration, and Pacing. Subsequently, we curate a large-scale dataset of 100k vlog edits and a dedicated benchmark, VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models (MLLMs). Finally, we present VlogReward, a robust vlog reward model that can provide both fine-grained multi-dimensional scores and actionable feedback for iterative refinement. Technically, we enhance the Group Relative Policy Optimization (GRPO) framework by introducing an adjustable inter-group comparison reward, which mitigates the "direction blindness" issue of standard GRPO and enables the model to better distinguish varied-quality edits. VlogReward achieves state-of-the-art results that significantly outperform existing MLLMs, including GPT-5 and Gemini-3-Pro. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems.

13:00 JST研究/論文

レッスンからの進化: 操作に応じたテーブル質問応答のためのスキル拡張テーブルグラフ推論

Table Question Answering (TableQA) は、テーブルを推論してユーザーのクエリに答えることを目的としています。既存の研究では、すべての質問を均一に扱い、全体的な精度のみで評価しているため、LLM は単純な検索には優れているものの、集計や算術などの複雑な操作には苦労しているという重要な現実が曖昧になっています。この差異を明らかにするために、きめ細かい質問分類を備えた新しい \emph{Operation-wise TableQA} タスクを導入し、WikiTQ-ow と TabFact-ow という名前の 2 つのデータセットを評価用にリリースします。モデリングのボトルネックに関して言えば、既存の方法はテーブルを線形化されたテキストに平坦化し、固有の構造を破壊し、複雑な行間推論の主な障壁となる「中間喪失」問題を引き起こします。さらに、彼らは通常、同様の操作間で共有される再利用可能なパターンを無視して、ゼロから推論します。これらの制限に対処するために、自己進化する構造化推論のためのスキル拡張テーブルグラフ推論 (SkillTGR) フレームワークを提案します。具体的には、SkillTGR は、明示的な行-列-セル構造を持つ属性付きグラフとしてテーブルを表し、LLM が動的チェーンを計画および実行して、グラフ走査推論の証拠サブグラフを取得します。これに基づいて、SkillTGR は、認知ヒューリスティックに基づいて推論の軌跡を抽象的なスキルに抽出するための階層的なスキルバンクを構築し、その後、対照的な拡張テーブル グラフ推論のために成功したスキルと失敗したスキルの両方をハイブリッドで取得し、それによって継続的な自己進化を可能にします。広範な実験により、SkillTGR が全体で平均 5.91\%、操作面で 6.03\% の向上という優れたパフォーマンスを実現し、トークン消費量が 19.76\%、推論レイテンシが 27.64\% 削減されることが実証されました。私たちのコードとデータは出版と同時に公開されます。

原文 (English)

Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering

Table Question Answering (TableQA) aims to reason over tables to answer user queries. Existing research treats all questions uniformly and evaluates solely through overall accuracy, obscuring a critical reality that LLMs excel at simple lookups yet struggle with complex operations like aggregation and arithmetic. To reveal this disparity, we introduce a novel \emph{Operation-wise TableQA} task with a fine-grained question taxonomy and release two datasets named WikiTQ-ow and TabFact-ow for evaluation. As for modeling bottlenecks, existing methods flatten tables into linearized texts, disrupting inherent structures and inducing the ``lost-in-the-middle'' issue, which poses a primary barrier to complex cross-row reasoning. Moreover, they typically reason from scratch, neglecting reusable patterns shared across similar operations. To address these limitations, we propose a Skill-augmented Table Graph Reasoning (SkillTGR) framework for self-evolving structured reasoning. Specifically, SkillTGR represents tables as attributed graphs with explicit row-column-cell structures, where LLMs plan and execute dynamic chains to retrieve evidence subgraphs for graph traversal reasoning. Based on this, SkillTGR builds a hierarchical SkillBank to distill reason trajectories into abstract skills under cognitive heuristics, then hybrid retrieves both successful and failed skills for contrastive augmented table graph reasoning, thereby enabling the continual self-evolution. Extensive experiments demonstrate that SkillTGR achieves superior performance with an average of 5.91\% overall and 6.03\% operation-wise improvement, also reducing 19.76\% token consumption and 27.64\% inference latency. Our codes and data will be released upon publication.

13:00 JSTLLM/生成AI

さきがけ: 拡散投機的復号のためのプレフィックス整列ツリー ドラフティング

拡散大規模言語モデル (dLLM) は、トークンを並行して生成する、自己回帰 (AR) LLM の有望な代替手段として登場しました。これにより、投機的デコード (SD) に有効なドラフト モデルとなり、単一の順方向パスでドラフト トークンのブロック全体が生成されます。しかし、既存の拡散ベースのドラフティング方法は、たとえ dLLM が複数の位置にわたって複数の候補トークンを発行し、デコード パスの大きな組み合わせ空間を誘導したとしても、線形ドラフティングに依存しています。その結果、受け入れられる長さと復号効率が制限されます。この複数の候補構造を活用するために、ツリーベースのドラフティングを拡散ドラフターに適用し、多様な候補パスの探索を可能にします。しかし、単純なツリーのドラフトは最適ではないことがわかりました。拡散マージナルはプレフィックスブラインドであり、プレフィックスベースの AR 検証と一致せず、信頼性の低いパスランキングが得られます。我々は、ツリーベースのドラフトを拡散ドラフト作成者に拡張する原則的なフレームワークである PRESTO を提案します。これは、拡散ドラフトの信頼性と、拡散投機的デコーディングのための PREfix 整合スコアリングと優先順位ベースのツリー検索を通じて、拡散ドラフトの信頼性とプレフィックスベースの AR 検証の間の根本的な不一致を解決します。 PRESTO の背後にある重要な原則は、(1) 候補のランキングは AR 検証のプレフィックスベースの性質と一致する必要があること、(2) ツリー構築では検証の可能性が高い候補パスを優先して受け入れ長を最大化することです。広範な実験により、PRESTO は、さまざまなベンチマークにわたって、最先端の専用拡散ドラフター SD では平均 $1.5\times$ のエンドツーエンド スループット高速化を実現し、自己投機的拡散 LLM では平均 $1.12\times$ の高速化を達成することが示されています。

原文 (English)

PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding

Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel. This makes them effective draft models for speculative decoding (SD), producing an entire block of draft tokens in a single forward pass. Yet existing diffusion-based drafting methods rely on linear drafting, even though dLLMs emit multiple candidate tokens across positions, inducing a large combinatorial space of decoding paths. Consequently, they limit acceptance length and decoding efficiency. To exploit this multi-candidate structure, we apply tree-based drafting to diffusion drafters, enabling exploration of diverse candidate paths. However, we find that naive tree drafting is suboptimal: diffusion marginals are prefix-blind, mismatching the prefix-based AR verification and yielding unreliable path ranking. We propose PRESTO, a principled framework that extends tree-based drafting to diffusion drafters while resolving the fundamental mismatch between diffusion draft confidence and prefix-based AR verification through PREfix-aligned Scoring and priority-based Tree search for diffusion speculative decOding. The key principles behind PRESTO are that (1) candidate ranking should align with the prefix-based nature of AR verification, and (2) tree construction should prioritize candidate paths with high verification potential to maximize acceptance length. Extensive experiments show that PRESTO achieves up to an average of $1.5\times$ end-to-end throughput speedup on the state-of-the-art dedicated diffusion drafter SD and an average of $1.12\times$ on self-speculative diffusion LLMs across diverse benchmarks.

13:00 JST研究/論文

CallBench: 電話通話アシスタントにおける二重目標調整のベンチマーク

ターゲット指向の対話システムは、インタラクティブな会話を通じてユーザーの目標を達成する強力な機能を実証しています。ただし、既存の研究は主に単一の明示的な目標を完了するように設計されていますが、通話アシスタントは、デバイス所有者の明示的な事前設定目標と発信者の暗黙的で動的な目標を調整する必要があるプロキシ設定に直面しています。電話アシスタントの二重目標調整を評価するための中国のベンチマークである \textsc{CallBench} を紹介します。 \textsc{CallBench} には、テイクアウト、配達、タクシー、仕事、生活、嫌がらせの 6 つのシナリオにわたる完全な複数ターンの電話対話が 50,000 件含まれています。通常のプリセット、緊急プリセット、およびプリセットなしのケースをカバーし、整合性、補完性、無関係性、競合など、所有者側と呼び出し側の目標の間のさまざまな関係が含まれます。さらに、意味の理解、コンテキストの使用、アクティブなガイダンス、応答の品質、プリセットのコンプライアンス、対話のリズム、安全性をカバーするプリセット対応のターンレベル評価プロトコルを設計します。代表的な対話方法に関する実験では、既存のアプローチが依然としてこのタスクに苦戦していることが示されており、プロキシ制約の下で 2 つの独立した目標の間で信頼性の高いターンレベルの意思決定を行うことができる電話アシスタントの必要性が強調されています。

原文 (English)

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However, existing studies are primarily designed for single, explicit goal completion, while phone call assistants face a proxy setting that requires coordinating the device owner's explicit preset goal with the caller's implicit and dynamic goal. We introduce \textsc{CallBench}, a Chinese benchmark for evaluating dual-goal coordination in phone call assistants. \textsc{CallBench} contains 50,000 complete multi-turn phone call dialogues across six scenarios: takeout, delivery, taxi, work, life, and harassment. It covers regular presets, emergent presets, and no-preset cases, and includes diverse relations between owner-side and caller-side goals, such as alignment, complementarity, irrelevance, and conflict. We further design a preset-aware turn-level evaluation protocol covering semantic understanding, context use, active guidance, response quality, preset compliance, dialogue rhythm, and safety. Experiments on representative dialogue methods show that existing approaches still struggle with this task, highlighting the need for phone call assistants that can make reliable turn-level decisions between two independent goals under proxy constraints.

13:00 JST研究/論文

線形で保護された存在ルールに基づくパス クエリへの応答

オントロジーを介したクエリ応答は、データベース インスタンスとオントロジーで構成される知識ベースに対するクエリに応答する問題に関係します。この分野のほとんどの研究は結合クエリ (CQ) に焦点を当てていますが、ナビゲーション クエリへの注目が高まっています。この論文では、保護された存在ルールのセットによってオントロジーが与えられる知識ベースに対する双方向 (結合) 通常パス クエリ ((C)RPQ) に答える複雑さを調査します。まず、線形存在ルールのサブクラスを検討し、(C)RPQ 応答はデータ複雑さにおいて NL 完全であり、これは単純なグラフ データベース (つまり、オントロジーなし) での RPQ 応答のデータ複雑さと一致することを示します。複雑さを組み合わせた場合、一般的なケースでは両方のタスクは ExpTime-complete ですが、述語アリティに制限がある場合、RPQ と CRPQ の応答はそれぞれ PTime-complete と PSpace-complete に低下します。保護されたルールの場合、線形ケースに対する非自明な削減を提供します。これにより、(C)RPQ 応答の複雑さが CQ の場合と同じであること、つまり、複合複雑さでは 2ExpTime-complete (有界性の場合は ExpTime-complete)、データ複雑さでは PTime-complete であることを示すことができます。

原文 (English)

Answering Path Queries under Linear and Guarded Existential Rules

Ontology-mediated query answering is concerned with the problem of answering queries over knowledge bases consisting of a database instance and an ontology. While most work in the area focuses on conjunctive queries (CQs), navigational queries have gained increasing attention. In this paper, we investigate the complexity of answering two-way (conjunctive) regular path queries ((C)RPQs) over knowledge bases whose ontology is given by a set of guarded existential rules. We first consider the subclass of linear existential rules and show that (C)RPQ answering is NL-complete in data complexity, which matches the data complexity of answering RPQs over plain graph databases (i.e., without an ontology). In combined complexity, both tasks are ExpTime-complete in the general case, but RPQ and CRPQ answering drop to PTime-complete and PSpace-complete respectively if there is a bound on predicate arity. For guarded rules, we provide a non-trivial reduction to the linear case, which allows us to show that the complexity of (C)RPQ answering is the same as for CQs, namely 2ExpTime-complete in combined complexity (ExpTime-complete in the bounded-arity case) and PTime-complete in data complexity.

13:00 JSTハードウェア/半導体

チャネル条件付きパラメータ生成による CSI モデルのシナリオ横断的な高速適応

ディープ ラーニングは、チャネル状態情報 (CSI) フィードバックやチャネル推定などの大規模な多入力多出力 (Massive MIMO) 物理層タスクに対する強い可能性を示しています。ただし、環境の不均一性により、目に見えないシナリオでは CSI モデルが大幅に劣化する可能性があり、従来の適応にはターゲット領域のデータと大量の計算が必要です。このペーパーでは、動的なワイヤレス環境で CSI モデルを迅速に展開するためのエンドツーエンドのパイプラインであるチャネル条件付きパラメーター生成 (CCPG) を提案します。 CCPG は、コンポーネント フリーズ実験を通じてシーン依存の適応ボトルネックを特定し、完全なモデル パラメーターではなく軽量の LoRA 重みのみを生成します。カスケード SVD と Perceiver Resampler を使用して、高次元のチャネル特徴をコンパクトな潜在条件に圧縮します。エネルギーベースの正規化メカニズムにより、LoRA 重みの順列と符号の曖昧さが軽減され、拡散ベースのジェネレーターには、トポロジーを意識したパラメーター生成のための構造情報と非対称サイズ認識損失が組み込まれています。 CSI フィードバックとチャネル推定のための DeepMIMO と WAIR-D の実験では、CCPG がターゲット シナリオのトレーニングや微調整を行わずに、単一のフォワード パスで約 3 秒で新しいシナリオに適応し、コストのかかるオンライン適応に匹敵するクロスドメイン回復パフォーマンスを達成することが示されています。これらの結果は、CCPG により、インテリジェント 6G 通信の大規模な動的ワイヤレス シナリオで CSI モデルの効率的な導入が可能になることを示しています。

原文 (English)

Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation

Deep learning has shown strong potential for massive multiple-input multiple-output (Massive MIMO) physical-layer tasks, including channel state information (CSI) feedback and channel estimation. However, environmental heterogeneity can severely degrade CSI models in unseen scenarios, while conventional adaptation requires target-domain data and substantial computation. This paper proposes Channel Conditional Parameter Generation (CCPG), an end-to-end pipeline for rapid deployment of CSI models in dynamic wireless environments. CCPG identifies scene-sensitive adaptation bottlenecks through component-freezing experiments and generates only lightweight LoRA weights instead of full model parameters. It compresses high-dimensional channel features into compact latent conditions using cascaded SVD and a Perceiver Resampler. An energy-based canonicalization mechanism mitigates permutation and sign ambiguities in LoRA weights, while a diffusion-based generator incorporates structural information and an asymmetric size-aware loss for topology-aware parameter generation. Experiments on DeepMIMO and WAIR-D for CSI feedback and channel estimation show that CCPG adapts to new scenarios in about 3 seconds with a single forward pass, without target-scenario training or fine-tuning, and achieves cross-domain recovery performance comparable to costly online adaptation. These results demonstrate that CCPG enables efficient deployment of CSI models in large-scale dynamic wireless scenarios for intelligent 6G communications.

13:00 JSTLLM/生成AI

TRACE: エンタープライズ LLM での知識保持パラメトリック ツール検索のためのビジネス ルールに基づいた推論カリキュラム

パラメトリック取得により、LLM は、各 API に一意の仮想トークンを割り当て、制約付きビーム検索によって生成するようにモデルをトレーニングすることにより、暗黙的にツールを取得できるようになります。 Toolsense は、この体制には 2 つの重大な欠点があることを示しています。1 つはトレーニング中にパラメトリック ツールの知識を破壊すること、もう 1 つはビーム検索デコードがリアルタイム展開するには遅すぎることです。この解離を解決する 2 段階のカリキュラムである TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules) を紹介します。ステージ 1 では、ToolSense のマルチフォーマット記憶 SFT を再利用して、LoRA でツールの知識をシードします。ステージ 2 が私たちの中心的な貢献です。モデルは、2 つのデータ ソース (ToolSense からの RRB ペアと、ドメインの専門家によって厳選されたビジネス ルールを対象とするために合成されたクエリ) を使用して、ツール トークンの JSON リストを生成する前に思考トレースを出力するようにトレーニングされます。どちらも推論トレースで強化されています。このトレーニング目標は、ステージ 1 MCQ および QA のプローブ精度を維持しながら、実稼働レイテンシでのシングルビームの貪欲なデコードを可能にします。 2 つのエンタープライズ製品ラインにわたる 8,300 以上のツールを組み合わせたエンタープライズ カタログで評価されたステージ 2 の TRACE トレーニングは、ツールの理解を維持するだけでなく向上させます。ステージ 1 と比較して、MCQ の精度が +3.2 pp 向上し、QA プローブが +9 pp 向上しました。検索時、TRACE は、埋め込みベースライン パフォーマンスが ~27% および ~52% であったのと比較して、ドメイン A で ~86%、ドメイン B で ~60% の再現率を達成しました -- どちらもシングルビームの場合貪欲なデコードにより、運用レイテンシーで直接デプロイ可能になります。

原文 (English)

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it destroys parametric tool knowledge during training, and its beam-search decoding is too slow for real-time deployment. We introduce TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules), a two-stage curriculum that resolves this dissociation. Stage 1 reuses the multi-format memorization SFT from ToolSense to seed tool knowledge with LoRA. Stage 2 is our core contribution: the model is trained to emit a thinking trace before producing a JSON list of tool tokens, using two data sources -- RRB pairs from ToolSense and queries synthesized to target business rules curated by domain experts -- both augmented with reasoning traces. This training objective preserves Stage 1 MCQ and QA probing accuracy while enabling single-beam greedy decoding at production latency. Evaluated on a combined enterprise catalog of 8,300+ tools across two enterprise product lines, TRACE training for Stage 2 not only preserves but improves tool understanding: MCQ accuracy gains +3.2 pp and QA probing gains +9 pp over Stage 1. On retrieval, TRACE achieves ~86% recall on Domain A and ~60% on Domain B -- compared to embedding baseline performance of ~27% & ~52% -- both with single-beam greedy decoding, making it directly deployable at production latency.

13:00 JSTLLM/生成AIエージェント

クラフト: スキーマを学び、計画を実行する

エンタープライズ コーディング エージェントは、自然言語の分析リクエストを、独自の API、スキーマ、およびメトリクス定義を介して実行可能なコードに変換します。しかし、各プロンプトに徹底的なスキーマとツールのドキュメントを挿入する一般的な展開パターンでは、推論のオーバーヘッドが増加し、スキーマの進化が複雑になり、マルチターン分析の信頼性が損なわれます。私たちは、本番環境向けの分析に必要な一貫性を維持しながら、安定したスキーマの知識とツールの使用行動をポストトレーニングを通じて取得できるかどうかを調査します。スキーマに基づいたコーディング エージェントのための 2 段階のポストトレーニング レシピである CRAFT を紹介します。まず、スキーマ ストリップされた PLAN 監視付き微調整は、徹底的なプロンプトタイム スキーマ インジェクションを行わずに、検証された軌跡からドメイン構造の計画と実行可能な動作を学習します。第 2 に、実行型の強化学習により、ツールの選択、コードの品質、計画とコードの一貫性、失敗した実行からの回復に関するポリシーが調整されます。トレーニングの軌跡は、実行検証、データ整合性チェック、LLM 判定推論監査を組み合わせた Tri-Gate フィルターを通じて厳選されます。当社では、キャンペーンのパフォーマンス分析、指標のドリルダウン、エンティティレベルのパフォーマンス分析、複数回にわたる分析の改善をカバーする、広告分析の計画的な展開について CRAFT を評価しています。エンタープライズ評価環境には、エージェント対応ツール サーフェスとしてベータ API が組み込まれており、25 のスキーマにリンクされたコア エンティティと 30 のエージェント ワークフローにまたがっています。スキーマを詰め込んだベースラインと比較して、CRAFT は複合エージェント スコアを +9.6 pp、一貫性を +4.1 pp、マルチターン コヒーレンスを +4.2 pp 向上させると同時に、入力トークンの負担を約 9 倍、スキーマ検出ループを最大 5 倍削減します。さらに、企業環境におけるマルチターンツール使用の強化学習に必要な、導入のトレードオフ、報酬形成の制限、トレーニングインフラストラクチャの拡張についても報告します。

原文 (English)

CRAFT: Learn the Schema, Execute the Plan

Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet the prevailing deployment pattern injecting exhaustive schema and tool documentation into each prompt increases inference overhead, complicates schema evolution, and undermines reliability in multi-turn analysis. We investigate whether stable schema knowledge and tool-use behavior can instead be acquired through post-training while preserving the consistency required for production-facing analytics. We present CRAFT, a two-stage post-training recipe for schema-grounded coding agents. First, schema-stripped PLAN supervised fine-tuning learns domain-structured plans and executable behaviors from validated trajectories without exhaustive prompt-time schema injection. Second, execution-shaped reinforcement learning aligns the policy for tool selection, code quality, plan-code consistency, and recovery from failed executions. Training trajectories are curated through a Tri-Gate filter combining execution validation, data-integrity checks, and LLM-judge reasoning audit. We evaluate CRAFT for planned rollout in advertising analytics, covering campaign performance analysis, metric drill-downs, entity-level performance analysis, and multi-turn analytical refinement. The enterprise evaluation environment incorporates beta APIs as the agent-facing tool surface and spans 25 schema-linked core entities and 30 agentic workflows. Relative to a schema-stuffed baseline, CRAFT improves composite Agent Score by +9.6 pp, consistency by +4.1 pp, and multi-turn coherence by +4.2 pp, while reducing input-token burden by approximately 9x and schema-discovery loops by up to 5x. We further report deployment tradeoffs, reward-shaping limitations, and training-infrastructure extensions required for multi-turn tool-use reinforcement learning in enterprise settings.

13:00 JST画像/動画生成エージェント

取得する前の理由: マルチモーダル RAG のエージェントティック プランニング

マルチモーダル検索拡張生成 (mRAG) は、外部の知識を使用して画像テキストのクエリに答えることを目的としていますが、ほとんどの既存システムは依然として、フラットな証拠空間上の生のマルチモーダル入力から直接検索します。この設計は、多くの場合 2 つの重要な課題に悩まされます。質問の意図は正しい視覚的指示対象に基づいている必要があるため、検索ターゲットの指定が不十分であること、もう 1 つは検索空間の構造が弱く、意味的に明確な証拠が単一のグローバル ランキング ステップで競合することを強いられることです。我々は、何を取得するか、どこを検索するかを明示的にモデル化することで、取得前に推論するマルチモーダルなエージェント検索フレームワークである MM-R2 を提案します。 MM-R2 はまず、画像と質問のペアから意図に基づいた検索状態を構築し、情報ニーズ、根拠のある指示対象、および検索制約をキャプチャします。次に、構造化された KnowledgeMap に対して検索を実行します。エージェントは関連する検索ユニットを選択してから、その中で根拠のあるクエリを発行します。この機能を有効にするために、マルチステップ検索プロセスの大規模な軌跡データセットである MM-R2-Traj を構築し、教師あり微調整と GRPO を備えた 2 段階のポストトレーニング戦略を採用します。 Infoseek および Encyclopedic VQA データセットの実験では、MM-R2 が回答精度​​において強力なベースラインを大幅に上回っていると同時に、より解釈可能で検証可能な検索軌跡を生成できることが示されています。

原文 (English)

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space. This design often struggles with two key challenges: the retrieval target is under-specified because the question intent must be grounded to the correct visual referent, and the search space is weakly structured, forcing semantically distinct evidence to compete in a single global ranking step. We propose MM-R2, a multimodal agentic retrieval framework that reasons before retrieval by explicitly modeling both what to retrieve and where to search. MM-R2 first constructs an intent-grounded retrieval state from the image-question pair, capturing the information need, grounded referent, and retrieval constraints. It then performs retrieval over a structured KnowledgeMap, where the agent selects relevant retrieval units before issuing grounded queries within them. To enable this capability, we build MM-R2-Traj, a large-scale trajectory dataset of multi-step retrieval processes, and adopt a two-stage post-training strategy with supervised fine-tuning and GRPO. Experiments on Infoseek and Encyclopedic VQA datasets show that MM-R2 substantially outperforms strong baselines on answer accuracy while also yielding more interpretable and verifiable retrieval trajectories.

13:00 JST画像/動画生成

DocHRL: コスト最適化された文書分類のための階層型強化学習フレームワーク

実際の文書分類パイプラインは通常、その複雑さや種類に関係なく、受信するすべての文書に同じ一連のモデルを適用します。これは、コンピューティングと人的リソースの非効率的な使用につながります。単純なドキュメントは過剰に処理され、難しいドキュメントは十分な精査を受けられない可能性があります。 DocHRL は、ドキュメントごとに最もコスト効率の高い分類ポリシーを適応的かつ動的に選択することを学習する階層型強化学習フレームワークです。 DocHRL は、文書分類を 2 レベルのポリシー階層による逐次的な意思決定問題として定式化します。トップレベルのポリシーは幅広いオプション (ビジョン分類子、LLM、OCR、人間参加型レビュー) から選択し、オプション固有のサブポリシーは呼び出す具体的なモデルまたはツールを選択します。報酬シグナルは、推論コスト、誤分類のコスト、および人間によるラベル付けのコストをキャプチャする、負の予想コストの合計です。 RVL-CDIP ベンチマークで近接ポリシー最適化を使用してトレーニングされた DocHRL は、16 のドキュメント クラスにわたってマクロ F1 0.973 を達成しながら、固定のスタンドアロン分類子によって発生する大幅に高いコストと比較して、ドキュメントあたりの平均コストを 2.74 正規化単位に削減します。私たちの結果は、コストを意識した強化学習が文書理解システムにおける分類パフォーマンスと運用効率を同時に向上できることを示しています。

原文 (English)

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type. This leads to inefficient use of compute and human resources: simple documents are over-processed while difficult ones may not receive enough scrutiny. We introduce DocHRL, a hierarchical reinforcement learning framework that learns to adaptively and dynamically select the most cost-effective classification policy on a per-document basis. DocHRL formulates document classification as a sequential decision problem with a two-level policy hierarchy: a top-level policy selects among broad options (vision classifiers, LLMs, OCR, and human-in-the-loop review), while option-specific sub-policies choose the concrete model or tool to invoke. The reward signal is the negative total expected cost, which captures inference cost, cost of misclassification, and cost of human labelling. Trained with Proximal Policy Optimisation on the RVL-CDIP benchmark, DocHRL achieves a macro F1 of 0.973 across 16 document classes while reducing average per-document cost to 2.74 normalised units compared to substantially higher costs incurred by fixed standalone classifiers. Our results demonstrate that cost-aware reinforcement learning can simultaneously improve classification performance and operational efficiency in document understanding systems.

13:00 JSTLLM/生成AI

事前トレーニング済み LLM のアルゴリズムの抽出: 隠れマルコフ モデルの事例

大規模言語モデル (LLM) は、コンテキスト内学習 (ICL) を介して隠れマルコフ モデル (HMM) からの次の観測値を予測するという驚くべき能力を示しますが、この能力の基礎となるアルゴリズムは未解決のままです。以前の研究では、コンセンサスなしにいくつかの候補が提案されており、モデルの内部活性化に基づいたものはありませんでした。私たちは 3 段階のパイプラインでこのギャップを埋めます。まず、LLM の動作を一連の候補アルゴリズムと経験的に比較し、スペースを 3 つのクラスに絞り込みます。ただし、すべての HMM 設定およびシーケンス長にわたる LLM の動作を説明する単一のクラスはありません。次に、3 つのクラス間の理論的な接続を導き出し、それぞれが Transformer によってコンテキスト内でどのように実装できるかを示し、小規模なトレーニング済み Transformer での構築を検証します。 3 番目に、事前トレーニング済み LLM に戻り、モデルのアクティベーションにおけるアルゴリズム信号を分離するレイヤーごとのプローブおよび介入手法であるプリンシパル アクティベーション プローブ (PAP) を紹介します。 PAP は、因果的にモデル予測を推進し、経験的な ICL パフォーマンスを追跡する低次元の線形表現を明らかにします。 PAP はさらに、これらの表現が基礎となる HMM レジームの特性に応じてどのように変化するかを明らかにします。個別の計算段階は異なる層に局在化されます。まとめると、私たちの結果は、事前トレーニングされた LLM のコンテキスト内の動作を基礎となる内部メカニズムに結び付け、LLM が HMM 上で ICL をどのように実行するかについての理解を深めます。

原文 (English)

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models

Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model's internal activations. We close this gap with a three-stage pipeline. First, we empirically compare LLM behavior against a suite of candidate algorithms and narrow the space to three classes -- though no single class explains LLM behavior across all HMM settings and sequence lengths. Second, we derive theoretical connections between the three classes and show how each can be implemented in-context by a Transformer, validating the construction in a small trained Transformer. Third, returning to pre-trained LLMs, we introduce the Principal Activations Probe (PAP), a layer-wise probing and intervention method that isolates algorithmic signals in model activations. PAP reveals low-dimensional linear representations that causally drive model predictions and track empirical ICL performance. PAP further reveals how these representations shift with properties of the underlying HMM regime; distinct computational stages are localized to different layers. Together, our results connect the in-context behavior of pre-trained LLMs to the underlying internal mechanisms and advance our understanding of how LLMs perform ICL on HMMs.

13:00 JST研究/論文

PTStore (プレフィックス テンソル ストア): 高スループット推論サービスのための分散プレフィックス キャッシュとレプリケーション

コンテンツ配信ネットワーク (CDN) のクライアント キャッシングの設計からインスピレーションを得た PTStore は、再利用可能な KV キャッシュ プレフィックスを形成する一般的なテンソルを配布および複製します。これは、推論を高速化するための最先端のアプローチで使用される主な技術です。これにより、KV キャッシュへのアクセスのレイテンシーが短縮され、一般的なテンソルを含むサーバー上の不釣り合いに多数のリクエストによって引き起こされる負荷の不均衡が緩和されます。さらに、分散化のおかげで、PTStore では LLM 推論用の KV キャッシュのサイズを桁違いに拡張できます。その結果、PTStore は、現在のベースラインよりも 5 ~ 6 倍効率的に、長いパッセージの Q&A データセットに対して推論を実行できます。ベースラインでは、異なるノードや GPU にまたがるメモリが集約されないため、KV キャッシュの再生成が必要になります。

原文 (English)

PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving

Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the main technique used by state of art approaches to accelerate inferences. This reduces the latency of accessing the KV cache and alleviates load imbalance caused by a disproportionately large number of requests on servers containing popular tensors. Furthermore, thanks to decentralization, PTStore allows the expansion of the size of the KV cache for LLM inference by orders of magnitude. As a result, PTStore can execute inferences on long passage Q\&A datasets 5-6 times more efficiently than current baselines, which do not aggregate memory across different nodes and GPUs and therefore require regenerating the KV cache.

13:00 JSTLLM/生成AI

STAIF: 次のような複雑な命令の段階的な最適化

複数の明示的な制約を持つ複雑な命令に従うことは、大規模言語モデル (LLM) にとって依然として根本的な課題です。 DPO などの既存の調整手法は、特に分布外または複数の制約設定の下で、個々の制約を厳密に満たすことを重視しないことが多い全体的な報酬シグナルを最適化します。この論文では、主観的な (ソフトな) 制約の調整を客観的に検証可能な (ハードな) 制約の最適化から切り離す段階的な最適化フレームワークである STAIF を提案します。ステージ 1 では、複数のネガティブ サンプルを使用した優先度の最適化を適用して、ソフト制約に対する感度を高めます。一方、ステージ 2 では、検証可能な報酬による強化学習 (RLVR) を適用して、ハード制約への厳密な準拠を強制します。この方法をサポートするために、約 31,000 個の複雑な複数制約命令からなる高品質のバイリンガル (英語、中国語) データセットである STAINSTRUCT を構築します。広範な分析により STAIF の設計が検証され、強力なベースラインに対する代表的なベンチマークでの最先端のパフォーマンスと真の一般化が示されます。

原文 (English)

STAIF: A Stage-wise Optimization for Complex Instruction Following

Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often underemphasize strict satisfaction of individual constraints, particularly under out-of-distribution or multi-constraint settings. In this paper, we propose STAIF, a stage-wise optimization framework that decouples the alignment of subjective (soft) constraints from the optimization of objectively verifiable (hard) constraints. Stage 1 applies preference optimization with multiple negative samples to sharpen sensitivity to soft constraints, while Stage 2 applies Reinforcement Learning with Verifiable Rewards (RLVR) to enforce strict compliance with hard constraints. To support this method, we construct STAINSTRUCT, a high-quality bilingual (English, Chinese) dataset of approximately 31,000 complex multi-constraint instructions. Extensive analyses validate the design of STAIF and show state-of-the-art performance on representative benchmarks against strong baselines, as well as genuine generalization.

13:00 JSTLLM/生成AIエージェント

ARdena: リアルタイム LLM エージェントのシナリオ主導型制御

大規模言語モデル (LLM) により、会話エージェントの能力はますます向上していますが、リアルタイムの対話環境でエージェントの動作を確実に制御することは依然として大きな課題です。既存のアプローチは、多くの場合、変化するインタラクション要件に適応するのが難しいモデルの微調整または調整手順に依存しています。このペーパーでは、構造化されたプロンプトを通じて実行時の動作制御を可能にするフレームワークである、階層化されたシナリオ駆動型 LLM 制御について紹介します。このアプローチでは、永続的なコンテキストとシナリオ固有の制約を組み合わせることで、基礎となるモデルを変更することなく、対話中にエージェントの動作を変更できます。このフレームワークは、音声対話、視覚認識、ツールの使用、およびアバターベースの応答生成を統合するリアルタイムのマルチモーダル具体化エージェントである ARDena に実装されています。提案されたアプローチは、制御の有効性、応答遅延、および動作の安定性に関して評価されます。この結果は、安定したリアルタイム操作を維持しながら、シナリオ定義だけで大幅に異なる対話動作を生成できることを示しており、LLM エージェントを制御するためのシナリオ主導のプロンプトの有効性が強調されています。

原文 (English)

ARdena: Scenario-driven control of real-time LLM agents

Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive environments remains a significant challenge. Existing approaches often rely on model fine-tuning or alignment procedures that are difficult to adapt to changing interaction requirements. This paper introduces layered scenario-driven LLM control, a framework that enables runtime behavior control through structured prompting. By combining persistent context with scenario-specific constraints, the approach allows agent behavior to be modified during interaction without changing the underlying model. The framework is implemented in ARDena, a real-time multimodal embodied agent that integrates speech interaction, visual perception, tool use, and avatar-based response generation. The proposed approach is evaluated with respect to control effectiveness, response latency, and operational stability. The results demonstrate that scenario definitions alone can produce substantially different interaction behaviors while maintaining stable real-time operation, highlighting the effectiveness of scenario-driven prompting for controlling LLM agents.

13:00 JSTLLM/生成AI研究/論文

KG2Code: 質問応答用の実行可能コードを介したナレッジ グラフと大規模言語モデルの橋渡し

最近の研究では、下流の知識集約型タスク、特にナレッジ グラフ質問応答 (KGQA) のパフォーマンスを向上させるために、ナレッジ グラフ (KG) と大規模言語モデル (LLM) の統合が検討されています。既存のアプローチは主に、検索拡張生成 (RAG) ベース、エージェント ベース、および SPARQL ベースの方法を通じて LLM と KG を組み合わせます。これらの方法は顕著な成功を収めていますが、構造情報の損失、不誠実な推論、柔軟性と一般化の制限など、いくつかの制限に依然として悩まされています。これらの課題に対処するために、この論文では、ナレッジ グラフをコードベースの表現に変換し、構造的なセマンティクスを維持しながら、最新の LLM のコードを意識した事前トレーニングと自然に連携する新しいアプローチである KG2Code を提案します。 KG2Code に基づいて、KG2Code-QA が、KGQA をコード生成タスクとして定式化する KGQA フレームワークとしてさらに導入されます。この定式化により、検証可能な推論トレースと実行可能なコードの生成が可能になり、それによって幻覚の影響が大幅に軽減されます。さらに、KG2Code-QA でオープンソース LLM を効果的にトレーニングするための大規模で高品質なコード コーパスを構築するための自動パイプラインが開発されています。トレーニング後、LLM はゼロショット シナリオで KGQA を実行できるようになります。広範な実験により、提案されたアプローチが KGQA に対する既存の KG 強化 LLM 手法よりも大幅に優れていると同時に、まだ見られていない KG に対する強力な一般化を示していることが実証されています。コードとデータは Github で入手できます。

原文 (English)

KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering

Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches primarily combine LLMs with KGs through retrieval-augmented generation (RAG)-based, agent-based, and SPARQL-based methods. Although these methods have achieved notable success, they still suffer from several limitations, including structural information loss, unfaithful reasoning, and limited flexibility and generalization. To address these challenges, this paper proposes KG2Code, a novel approach that transforms knowledge graphs into a code-based representation, preserving structural semantics while naturally aligning with the code-aware pretraining of modern LLMs. Based on KG2Code, KG2Code-QA is further introduced as a KGQA framework that formulates KGQA as a code generation task. This formulation enables the generation of verifiable reasoning traces and executable code, thereby substantially mitigating the impact of hallucinations. In addition, an automated pipeline is developed to construct a large-scale, high-quality code corpus for effectively training open-source LLMs on KG2Code-QA. After training, LLMs are able to perform KGQA in zero-shot scenarios. Extensive experiments demonstrate that the proposed approach significantly outperforms existing KG-enhanced LLM methods for KGQA, while exhibiting strong generalization to unseen KGs. The code and data are available at Github.

13:00 JST研究/論文GPT / ChatGPT

言語モデルはそれ自体に収束しますか?テキストの緩和としての再帰的自己洗練

大規模な言語モデルは、最初のドラフトが同じモデルによって繰り返し改訂される再帰的改良ワークフローで使用されることが増えています。使用が増加しているにもかかわらず、このようなワークフローの長期的なダイナミクスは依然としてよく理解されていません。改良を繰り返すことで出力は無限に改善され続けるのでしょうか、それとも安定したテキスト形式に収束するのでしょうか?私たちは、反復的な LLM リビジョンがテキストをモデル優先のソフト固定小数点領域に向けて駆動する動的プロセスとしての再帰的自己洗練を研究します。 GPT-5.5 を使用して、デフォルト温度と決定論的デコーディングの両方で 50 個の ICML 2025 要約の 10 ステップのリファインメント軌跡を生成し、さらに 15 個の ICML 2020 要約を評価します。正規化された編集距離、正確な固定点と近似的な固定点、単語数の安定性、指数関数的緩和、および外部の LLM-as-a-judge 評価を分析します。すべての設定にわたって、リファインメント軌跡は急速に飽和します。ほとんどの編集は最初の数回の反復内で行われ、その後、軌道は表面レベルの小さな変更のみを伴うソフト固定小数点領域に入ります。決定論的デコーディングは、デフォルト温度デコーディングよりも早く正確な固定点に到達し、残留変動が小さくなり、どちらも普遍的な近似収束を達成します。平均編集量は一貫した指数関数的緩和パターンに従い、制限のない最適化ではなく、モデルが優先するテキスト平衡への収束を示唆しています。外部評価によると、集約された要約は、技術的な意味を維持しながら、明瞭さ、簡潔さ、科学的なスタイルを向上させます。これらの発見は、LLM 自己洗練の動的システムの見方をサポートし、編集規模の飽和に基づいた実際的な停止基準の動機付けとなります。

原文 (English)

Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation

Large language models are increasingly used in recursive refinement workflows, where an initial draft is repeatedly revised by the same model. Despite their growing use, the long-term dynamics of such workflows remain poorly understood. Does repeated refinement continue to improve outputs indefinitely, or does it converge toward a stable textual form? We study recursive self-refinement as a dynamical process in which repeated LLM revision drives text toward a model-preferred soft fixed-point region. Using GPT-5.5, we generate 10-step refinement trajectories for 50 ICML 2025 abstracts under both default-temperature and deterministic decoding, and additionally evaluate 15 ICML 2020 abstracts. We analyze normalized edit distance, exact and approximate fixed points, word-count stability, exponential relaxation, and external LLM-as-a-judge evaluation. Across all settings, refinement trajectories rapidly saturate. Most edits occur within the first few iterations, after which trajectories enter a soft fixed-point region with only minor surface-level changes. Deterministic decoding reaches exact fixed points earlier and exhibits smaller residual fluctuations than default-temperature decoding, while both achieve universal approximate convergence. The average edit magnitude follows a consistent exponential relaxation pattern, suggesting convergence toward a model-preferred textual equilibrium rather than open-ended optimization. External evaluation indicates that converged abstracts improve clarity, conciseness, and scientific style while preserving technical meaning. These findings support a dynamical-systems view of LLM self-refinement and motivate practical stopping criteria based on edit-magnitude saturation.

13:00 JST研究/論文

MINT-V2X: 予測リソース管理のためのモビリティ統合ネットワーク軌跡データセット

Vehicle-to-Everything (V2X) 通信システムは、車両の軌跡データだけでなく、現実的なレベルの忠実度を持つワイヤレス ネットワーク パラメーターも含むデータセットに基づいており、予測および最適化モデルの作成を可能にします。現在、非常に重大な研究インフラストラクチャのギャップがあり、一般に公開されているデータセットは、モビリティまたはネットワーク パラメータの 2 つのうちの 1 つに限定される可能性が高く、両方を組み合わせた単一の統合ビューを提供することはほとんどありません。このペーパーでは、SUMO トラフィック ダイナミクスと OMNeT++/Simu5G ネットワーク シミュレーションを結合することによって生成された包括的なデータセットである MINT-V2X を紹介します。検証フレームワークは、3GPP リリース 14 (C-V2X)、ETSI 標準、およびシャノン容量理論に基づいた 14 の標準化されたテストで構成されています。結果として得られるデータセットには、3 時間の都市交通シミュレーション中に 15 台の路側装置 (RSU) の 1,386 台の車両から収集された 987 万個の同期データ ポイントが含まれています。ネットワーク メトリック相関 (CQI-SINR: 0.993; SINR-PDR: 0.946) を通じて、厳密なアルゴリズムの一貫性を実証します。最後に、RSU 負荷予測ケース スタディを実施することでデータセットの価値を実証し、軌跡データを使用すると、ネットワーク履歴のみのベースラインよりも優れた予測パフォーマンスが得られることを示します。データセット、実験、および完全な SUMO 構成ファイルは GitHub リポジトリで入手でき、代替シミュレーション スタックでの再現を容易にします。

原文 (English)

MINT-V2X: A Mobility-Integrated Network Trajectory Dataset for Predictive Resource Management

Vehicle-to-Everything (V2X) communication systems are based on datasets that not only contain vehicle trajectory data but also wireless network parameters with a realistic level of fidelity, enabling the creation of prediction and optimization models. There is a very critical research infrastructure gap today, and publicly available datasets are likely to be limited to one of the two: mobility or network parameters, and rarely provide a single, integrated view that combines both. This paper introduces MINT-V2X, a comprehensive dataset generated by coupling SUMO traffic dynamics with OMNeT++/Simu5G network simulation. The validation framework is composed of 14 standardized tests based on 3GPP Release 14 (C-V2X), ETSI standards and Shannon capacity theory. The resulting dataset contains 9.87 million synchronized data points from 1,386 vehicles from 15 roadside units (RSUs) during 3 hours of urban traffic simulation. We demonstrate strict algorithmic consistency through network metric correlations (CQI-SINR: 0.993; SINR-PDR: 0.946). Finally, we demonstrate the value of the dataset by conducting an RSU load prediction case study, showing that using trajectory data yields better predictive performance than network-history-only baselines. The dataset, experiments, and complete SUMO configuration files are available in the GitHub repository to facilitate reproduction on alternative simulation stacks.

13:00 JSTLLM/生成AI

EventOD: LLM ガイドによるセマンティック変調によるイベント対応 OD フロー生成

破壊的な事象が発生した場合の起点-終点(OD)フローの推定は、災害対応と都市の回復力にとって重要です。日常的な移動に関してトレーニングされた既存のディープ OD モデルは、極端なイベントによって地域機能や人口活動が突然変化すると機能が低下することがよくありますが、イベントごとに新しいジェネレーターを再トレーニングすることは、限られたイベント時間の監視下では非現実的です。私たちは、構造化されたイベント セマンティクスを使用して事前トレーニングされた OD ジェネレーターを制御するイベント適応型 OD 生成フレームワークである EventOD を提案します。 EventOD は、まず大規模な言語モデルを使用して、大まかなイベント観測から地域レベルの機能および人口統計の制御ベクトルを推測します。次に、AlphaNet と BetaNet という 2 つの軽量適応モジュールを学習して、これらのセマンティック シフトの大きさを調整し、監視がまばらなシナリオ向けに検索拡張フォールバック パスウェイをさらに導入します。結果として得られるイベント条件付き特徴は、入力レベルの変調を通じて事前トレーニングされたグラフ拡散 OD モデルに注入され、ジェネレーターのパラメーターを更新せずにイベントを認識した適応を可能にします。米国の郡全体にわたるハリケーンおよびパンデミックに起因するモビリティに関する実験では、EventOD が強いベースラインを超えて再構成精度と分布忠実度の両方を一貫して向上させることが示されています。ソースコードは https://anonymous.4open.science/r/EventOD-5C11/ で入手できます。

原文 (English)

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation

Estimating origin-destination (OD) flows under disruptive events is important for disaster response and urban resilience. Existing deep OD models trained on routine mobility often degrade when extreme events abruptly alter regional functions and population activities, while retraining a new generator for each event is impractical under limited event-time supervision. We propose EventOD, an event-adaptive OD generation framework that steers a pretrained OD generator using structured event semantics. EventOD first uses a large language model to infer region-level functional and demographic control vectors from coarse event observations. It then learns two lightweight adaptation modules, AlphaNet and BetaNet, to calibrate the magnitude of these semantic shifts, and further introduces a retrieval-augmented fallback pathway for scenarios with sparse supervision. The resulting event-conditioned features are injected into a pretrained graph diffusion OD model through input-level modulation, enabling event-aware adaptation without updating generator parameters. Experiments on hurricane- and pandemic-induced mobility across U.S. counties show that EventOD consistently improves both reconstruction accuracy and distributional fidelity over strong baselines. Source code is available at https://anonymous.4open.science/r/EventOD-5C11/.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

StanceBench: 音声 LLM ベースの対人スタンス評価の音声からのベンチマーク

音声間の対話モデルは、社会的意図を伝えるために韻律や相互作用のニュアンスにますます依存していますが、これらの合図のベンチマークは依然として限られています。会話中の会話における対人スタンスを測定し、自動判定として音声対応 LLM を評価するためのベンチマークである StanceBench を紹介します。シームレス インタラクション コーパスを使用して、StanceBench は、(1) ロール プロンプト ポールを介して 9 つのスタンスの次元を指定し、(2) 単一話者およびインタラクション ベースの評価を標準化し、(3) ジャッジとしての LLM の堅牢性、バイアス、およびスタンス推論をレポートします。評価されるスタンスの中で、共感と礼儀正しさが最も簡単です。温かさと自己主張は、ポジティブな偏り/非対称性によって適度に分離可能です。正直さが最も難しく、即時注文バイアスが高く、クロスターンの証拠が必要であることと一致しています。注意力は分離可能ですが、人間との連携は弱いです。インタラクションのスタンスはよりコンテキストに依存しており、しきい値のギャップと大きな差異があり、特に紛争規制が顕著です。

原文 (English)

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceBench, a benchmark for measuring interpersonal stance in conversational speech and evaluating audio-capable LLMs as automated judges. Using the Seamless Interaction corpus, StanceBench (1) specifies 9 stance dimensions via role-prompt poles, (2) standardizes single-speaker and interaction-based evaluations, and (3) reports LLM-as-a-judge robustness, bias, and stance inference. Across evaluated stances, empathy and politeness are the easiest. Warmth and assertiveness are moderately separable with positivity skew/asymmetry. Honesty is the hardest and shows high prompt order bias, consistent with needing cross-turn evidence. Attentiveness is separable but aligns weakly with humans. Interaction stances are more context-sensitive, with threshold gaps and high variance, especially conflict regulation.

13:00 JSTLLM/生成AI

TRE: 拡散言語モデルのためのトレーニング不要の幻覚検出

拡散大規模言語モデル (D-LLM) は最近ますます注目を集めていますが、その信頼性は幻覚問題によって大きく妨げられています。 D-LLM に対する既存の幻覚検出アプローチは主にトレーニング ベースのパラダイムに従い、データ駆動型トレーニングに依存して検出器を最適化します。このような依存性は、ドメイン モデル全体での汎用性を制限するだけでなく、追加のトレーニング コストと展開のオーバーヘッドも発生します。これらの制限に対処するために、D-LLM 向けのトレーニング不要の幻覚検出メトリックである TRE を提案します。 TRE は、検出器のトレーニングや繰り返しのサンプリングを必要とせず、単一世代のエントロピー信号から幻覚リスクを直接推定する、パラメータフリーの 1 回実行のメトリクスです。 TRE は、D-LLM デコード プロセス内で空間次元と時間次元の両方に沿ってエントロピー信号を抽出します。トークンレベルの空間的観点から、私たちは不確実性の最も有益な媒体としてのトークンを明らかにし、不確実性が積極的に関与している場所を捕捉することに焦点を当てています。拡散ステップレベルの時間的観点から、後期ステップのエントロピーの優位性を経験的に特定し、単純な線形重み付けスキームを使用してこれらの信号を集約して TRE を取得します。複数の D-LLM および QA データセットに対する広範な実験により、TRE が強力な一般化性、効率性、および堅牢性を享受しながら、競争力のあるパフォーマンスを達成することが実証されました。

原文 (English)

TRE: Training-Free Hallucination Detection for Diffusion Language Models

Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination problem. Existing hallucination detection approaches for D-LLMs mainly follow a training-based paradigm, relying on data-driven training to optimize the detector. Such reliance not only limits their generalizability across domains models but also incurs additional training cost and deployment overhead. To address these limitations, we propose TRE, a training-free hallucination detection metric for D-LLMs. TRE is a parameter-free and single-run metric that estimates hallucination risk directly from the entropy signals of a single generation, without requiring any detector training or repeated sampling. TRE extracts entropy signals within the D-LLM decoding process along both the spatial and temporal dimensions. From a token-level spatial perspective, we focus on revealing tokens as the most informative carriers of uncertainty, capturing where uncertainty is actively committed. From a diffusion step-level temporal perspective, we empirically identify the dominance of late-step entropy and hence aggregate these signals with a simple linear weighting scheme to obtain TRE. Extensive experiments on multiple D-LLMs and QA datasets demonstrate that TRE achieves competitive performance, while enjoying strong generalizability, efficiency, and robustness.

13:00 JSTLLM/生成AI

CuraWeb: Web スケールの事前トレーニング データの品質、冗長性、多様性の共同最適化

FineWeb-Edu や DCLM などの高度に選択的なフィルターを介してキュレーションされたオープンウェブ企業は、LLM 事前トレーニング データの中核を構成し、大幅に高度な LLM パフォーマンスを備えています。ただし、これらのパイプラインは通常、単一の最適化目標に依存しているため、必然的に分布の多様性が狭まり、ロングテールの知識が周辺化され、その結果、データ範囲が制限され、オープンウェブの膨大な可能性が十分に活用されません。この制限に対処するために、私たちは線形枝刈りから品質、冗長性、多様性の共同最適化に移行する新しいキュレーション パラダイムを提案します。このフレームワークは、デュアルトラック クリーニング (ルールベースおよびモデル駆動) とハイブリッド重複排除 (N グラムおよびセマンティック) を相乗させながら、多目的サンプラーを採用して情報品質と配布範囲のバランスをとります。このフレームワークを Common Crawl に適用し、2T トークンの英語コーパスである CuraWeb を構築します。既存のリソースとは異なり、CuraWeb は、多様性を高め、冗長性を最小限に抑えたより包括的なデータ分散を回復することにより、データキュレーションの業界グレードの標準を確立し、多様なドメインにわたるロングテール知識のより広い範囲を実現します。 3B スケールでの実験評価では、CuraWeb が最先端のベースラインを大幅に上回っており、特に知識集約型タスクや推論タスクにおいて、幅広いベンチマーク全体で平均 1.8\% のパフォーマンス向上をもたらしていることが実証されています。

原文 (English)

CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the vast potential of the open web. To address this limitation, we propose a novel curation paradigm that shifts from linear pruning to the joint optimization of quality, redundancy, and diversity. This framework synergizes dual-track cleaning (rule-based and model-driven) with hybrid deduplication (n-gram and semantic), while employing a multi-objective sampler to balance informational quality with distributional breadth. Applying this framework to Common Crawl, we construct CuraWeb, a 2T-token English corpus. Unlike existing resources, CuraWeb establishes an industrial-grade standard for data curation by recovering a more holistic data distribution with enhanced diversity and minimal redundancy, achieving broader coverage of long-tail knowledge across diverse domains. Experimental evaluations at the 3B scale demonstrate that CuraWeb significantly outperforms state-of-the-art baselines, yielding an average performance gain of 1.8\% across a wide range of benchmarks, particularly in knowledge-intensive and reasoning tasks.

13:00 JSTLLM/生成AI

ブロック境界を超えて: 拡散大規模言語モデルのマルチブロック編集

ブロック拡散は、離散拡散言語モデル (dLLM) をスケーリングするための主要なパラダイムとして浮上しています。これは、固定サイズのブロックでテキストをデコードすると、二次注意のコストを扱いやすく保ちながら、各ブロック内の並列生成が維持されるためです。ただし、この効率には構造的な制限があります。ブロックの終わり近くのトークンは、将来のブロック間コンテキストにアクセスせずに生成され、ブロックが確定すると、その不確実な予測は後続のすべてのブロックにとって不可逆的なコンテキストになります。これにより、ブロック境界の問題が発生し、ブロック境界に向かって不確実性が蓄積され、初期の間違いが後の世代に伝播します。この問題に対処するために、ブロック間のコンテキストに基づいてデコードされたトークンを編集することでこの問題を軽減するマルチブロック編集 (MBE) を提案します。この原則に従って、MBE はまず、前のブロックでデコードされたトークンを編集するためのトレーニング不要のデコード アルゴリズムを提案します。これは、選択されたブロックに対してフル アテンション ウィンドウを再度開くことで実現されます。ブロック拡散トレーニングと MBE 推論の間のアテンション メカニズムの不一致を考慮して、MBE はさらに、編集範囲を徐々に拡張する双方向アテンション マスクをモデルに装備する教師あり微調整戦略を導入します。さらに、マルチシェイプ CUDA Graph プールときめ細かい KV キャッシュ制御を使用して SGLang を拡張し、実際にこれらの可変長編集パスを効率的に実行します。 13 のベンチマークにわたる LLaDA2.1-Mini の実験では、トレーニング不要の MBE が、同等のスループットを維持しながら既存のすべてのデコード ベースラインを上回り、MBE SFT がさらに 2.7 のパフォーマンス向上をもたらすことが示されています。最大の改善は、AIME 2025 で +13.3、ZebraLogic で +5.9 など、強力な長期一貫性を必要とするタスクで見られ、MBE の有効性を示しています。

原文 (English)

Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic attention cost tractable. However, this efficiency comes with a structural limitation: tokens near the end of a block are generated without access to future cross-block context, and once a block is finalized, its uncertain predictions become irreversible context for all subsequent blocks. This creates a block boundary problem, in which uncertainty accumulates toward block boundaries and early mistakes propagate throughout later generation. To address this issue, we propose Multi-Block Editing (MBE), to mitigate this problem by editing decoded tokens based on cross-block context. Following this principle, MBE first proposes a training-free decoding algorithm to edit the decoded tokens in previous blocks, which is achieved by re-opening a full-attention window over selected blocks. Given the mismatched attention mechanism between block diffusion training and MBE inference, MBE further introduces a supervised Fine-tuning strategy, which equips the model with bidirectional attention masks that progressively expands the editing span. Furthermore, it also extends SGLang with a multi-shape CUDA Graph pool and fine-grained KV cache control to make these variable-length editing passes efficient in practice. Experiments on LLaDA2.1-Mini across 13 benchmarks show that training-free MBE outperforms all existing decoding baselines while maintaining comparable throughput, and MBE SFT further brings a performance gain of 2.7. The largest improvements appear on tasks requiring strong long-range consistency, including +13.3 on AIME 2025 and +5.9 on ZebraLogic, demonstrating the effectiveness of MBE.

13:00 JST研究/論文

Obliviate: レコメンダー システムにおける効率的なアンラーニング

機械のアンラーニングは、データ プライバシー規制の観点から、特にユーザー インタラクション データに基づいて直接トレーニングされるレコメンデーション システムにとって、ますます重要になってきています。この作業の目標は、推奨品質を維持しながら、要求されたインタラクション データとその下流への影響をトレーニング済みモデルから削除すること、および完全な再トレーニングにかかる​​大幅な計算コストを発生させることなくこれを行うことです。既存のアプローチには、かなりの計算オーバーヘッドがある一方で、非学習完全性の制限や推奨パフォーマンスの低下など、いくつかの制限があります。本稿では、良好な有用性を維持しながら高い非学習完全性を実現する、レコメンダシステムのための効率的な2段階の非学習フレームワークであるOblivateを提案します。最初の段階では、低ランク非学習アダプター (LUA) を導入します。これは軽量のヘシアン プロキシを採用し、完全なパラメーターではなくローカライズされた低ランク アダプターを通じて曲率を認識した効率的な非学習を可能にします。第 2 段階では、知識の蒸留を通じて実用性を維持しながら、ランキング ベースの目標を介して非学習を強制することでパフォーマンスを向上させるためにアダプター パラメーターのみを更新する軽量の改良段階である Locality-Aware Calibration (LAC) を提案します。広範な実証評価により、Oblivate は推奨品質の損失を最小限に抑え、計算コストを大幅に削減しながら高レベルの忘却を達成し、大規模なレコメンダー システムに実用的でスケーラブルなソリューションを提供することが実証されています。

原文 (English)

Obliviate: Efficient Unlearning in Recommender Systems

Machine unlearning is becoming increasingly critical in the context of data privacy regulations, particularly for recommendation systems that are directly trained on user interaction data. The goal of this work is to remove requested interaction data and their downstream influence from trained model while preserving recommendation quality, and to do so without incurring the substantial computational cost of full retraining. Existing approaches exhibit several limitations, including limited unlearning completeness and degradation in recommendation performance, while having substantial computational overhead. In this paper, we propose Obliviate, an efficient two-stage unlearning framework for recommender systems that achieves high unlearning completeness while maintaining good utility. In the first stage, we introduce a Low-Rank Unlearning Adapter (LUA), which employs a lightweight Hessian proxy to enable curvature-aware and efficient unlearning through localized low-rank adapters rather than full parameters. In the second stage, we propose Locality-Aware Calibration (LAC), a lightweight refinement stage that updates only the adapter parameters to improve the performance by enforcing unlearning via ranking-based objectives while preserving utility through knowledge distillation. Extensive empirical evaluations demonstrate that Obliviate achieves high level of forgetting with minimal loss in recommendation quality and at significantly reduced computational cost, offering a practical and scalable solution for large-scale recommender systems.

13:00 JSTロボティクス研究/論文

海上監視における異種センサー選択のための強化学習

この論文では、異種海上センサーネットワークにおける単一船舶追跡のための、情報利得に基づく強化学習センサー選択フレームワークを紹介します。提案されたアプローチは、情報理論的なセンサー管理によって動機づけられています。すべてのセンサーをアクティブにしたり、計算コストのかかるオンラインで期待される情報利得の評価を繰り返し実行したりする代わりに、学習されたポリシーによって、各決定エポックで追跡に関連するセンサーが 1 つ選択されます。ベイジアン逐次モンテカルロ トラッカーは、ノイズの多い測定値から船舶の状態を推定し、非線形および非ガウス条件下でのスケジューリングのための信念表現を提供します。 Proximal Policy Optimization エージェントは、キプロスのアギア ナパ マリーナにある CMMI スマート マリーナ テストベッドの地理参照シミュレーションに導入された 5 つのセンサーのうち 1 つを選択します。エージェントは、信念状態、検出履歴、カバレッジ、センサーの形状、実現された情報獲得の特徴を観察します。報酬は、可観測性マスクによってゲートされた実現情報獲得項として定義されます。最終テストのシミュレーションでは、提案されたフレームワークとランダムな単一センサー選択、すべてのセンサーを同時に使用する常時オン センシング、および以前の研究で提案された期待情報利得センサー選択ベースラインを比較します。結果は、学習されたポリシーが、意思決定タイム ステップごとに 1 つのセンサーのみをアクティブにし、期待される情報ゲインの選択に必要な計算コストのかかるオンライン エントロピー検索を回避しながら、常時オンのセンシングに近い追跡パフォーマンスを達成することを示しています。

原文 (English)

Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance

This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instead of activating all sensors or repeatedly performing computationally expensive online expected-information-gain evaluation, a learned policy selects one tracking-relevant sensor at each decision epoch. A Bayesian sequential Monte Carlo tracker estimates the vessel state from noisy measurements and provides a belief representation for scheduling under nonlinear and non-Gaussian conditions. A Proximal Policy Optimization agent selects one of five sensors deployed in a georeferenced simulation of the CMMI Smart Marina testbed at Ayia Napa Marina, Cyprus. The agent observes belief-state, detection-history, coverage, sensor-geometry, and realized-information-gain features. The reward is defined as a realized-information-gain term gated by an observability mask. Final-test simulations compare the proposed framework with random single-sensor selection, always-on sensing using all sensors simultaneously, and the expected-information-gain sensor-selection baseline proposed in our previous work. Results show that the learned policy achieves tracking performance close to always-on sensing while activating only one sensor per decision time step and avoiding the computationally expensive online entropy search required by expected-information-gain selection.

13:00 JST研究/論文

AIR-BENCH Live: 基礎モデルの進化する安全性ベンチマーク

財団モデルの安全性ベンチマークは、その発行当時の AI リスクを捉えています。モデルが改良され、政府が新しい AI 安全性法案を可決するにつれて、そのリスク分類法は理解しにくくなり、攻撃プロンプトは無効になります。私たちは、AIR-BENCH 2024 の自己進化型後継製品である AIR-BENCH Live を紹介します。自動更新パイプラインは政府規制を監視し、現在の 4 段階のリスク分類法に照らして新しいポリシーを分類し、既存のカテゴリに一致させるか、新しい詳細なカテゴリを提案します。次に、マルチエージェントのペルソナ主導のプロンプト生成アルゴリズムにより、人間によるレビューを最小限に抑えながら現実的な多言語プロンプトが生成され、最新の脱獄テクニックによる改善の余地が残されています。このアルゴリズムは、従来のプロンプトを徹底的に見直し、新しいカテゴリのプロンプトを生成するために使用されます。現在のバージョンでは、パイプラインによってベンチマークが 314 から 335 の粒度リスクに拡張され、7 つの管轄区域にわたる 31 の真に新しいポリシー条項から 21 の新しいカテゴリが抽出されています。 14 の最近のモデルを評価すると、安全性のばらつきが広く (独自の動作で判断されたモデル間で 0.17 から 0.89)、最新化されたプロンプトは 2024 年のセットより平均で 0.06 ポイント難しく、最大の低下は最も準拠したモデルに集中しており、ほとんどのモデルは英語以外のプロンプトでは安全性が若干低いことがわかりました。 AIR-BENCH Live は、新しい規制を継続的に吸収し、プロンプトを再生成することで、急速に変化する分野とともに進化するように設計されています。

原文 (English)

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective. We present AIR-BENCH Live, a self-evolving successor to AIR-BENCH 2024. An automated update pipeline monitors government regulation and classifies new policies against the current four-tier risk taxonomy, either matching them to existing categories or proposing new granular categories. Then, a multi-agent, persona-driven prompt generation algorithm generates realistic, multilingual prompts with minimal human review, leaving room for improvement with modern jail breaking techniques. This algorithm is used to overhaul legacy prompts and generate prompts for new categories. In our current version, the pipeline has expanded the benchmark from 314 to 335 granular risks, with the 21 new categories drawing from 31 truly novel policy clauses across seven jurisdictions. Evaluating 14 recent models, we find a wide safety spread (from 0.17 to 0.89 among the models judged on their own behavior), that the modernized prompts are on average 0.06 points harder than the 2024 set, with the largest drops concentrated among the most compliant models, and that most models are modestly less safe on non-English prompts. By continuously absorbing new regulation and regenerating prompts, AIR-BENCH Live is designed to evolve alongside a fast-moving field.

13:00 JSTLLM/生成AI

LLM のタスク適応がどのように調整を再形成するか: 行動と表現のドリフトに関する多次元研究

ポストトレーニングは、大規模な言語モデルを下流のタスクに適応させるための重要なメカニズムです。これまでの研究では、タスクの適応によってモデルの既存のアライメント、特に安全動作が変化する可能性があることが示唆されていますが、アライメント ドメイン全体にわたるその広範な影響は依然としてよく理解されていません。私たちは、安全性、事実性、姿勢の安定性、社会的危害、制御可能性、指導可能性の 6 つの主要な領域にわたる 15 の調整側面にわたる、教師あり微調整 (SFT)、KL 正則化 SFT、検証可能な報酬付き強化学習 (RLVR) などの代表的なタスク適応手法の体系的な評価を通じて、このギャップに対処します。私たちの結果は、トレーニング後にアライメントが均一に再形成されるわけではないことを明らかにしています。 RLVR は、比較的小さいがゼロではないメトリック固有のシフトを引き起こしながらタスクのパフォーマンスを向上させますが、SFT はドメイン全体でかなり大きなアライメント ドリフトを引き起こします。 KL 正則化はこの影響を軽減します。つまり、より強力な参照モデル アンカリングにより、ベースラインからのアライメントのドリフトが減少しますが、KL-SFT はアライメントの維持において RLVR にはまだ達していません。表現レベルの分析は、行動のドリフトを追跡するアライメント関連の表現の変化により、このパターンをさらにサポートします。まとめると、これらの結果は、タスク適応が単なる能力向上ステップではなく、それ自体が調整介入であり、トレーニング後のパイプラインの標準コンポーネントとして多次元調整評価を動機付けることを示しています。

原文 (English)

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.

13:00 JST研究/論文

DOSA: 長い文書構造分析のためのツリーガイドの自己回帰フレームワーク

視覚的に豊富なドキュメントでは、情報は表、ヘッダー、テキスト ブロックなどの個々のページ オブジェクトだけでなく、それらの間の構造的関係にもエンコードされるため、ドキュメントの構造分析が情報の検索とドキュメントの理解の基礎となります。ただし、長距離の依存関係や異種レイアウトを伴う複数ページのドキュメントでは、そのような関係を正確に推測することは依然として困難です。これに対処するために、ページ オブジェクト間の関係を推論し、ドキュメント レベルのセマンティック ツリーを再構築するための、DOcument Structure Analyzer (DOSA) と呼ばれる、ツリーガイド型の自己回帰フレームワークを提案します。 DOSA はドキュメントをチャンクごとに処理し、各ページ オブジェクトのビジュアル、テキスト、およびレイアウトの機能を融合し、階層関係と順序関係を予測します。予測された関係は、セマンティック ツリーを段階的に構築するために使用され、その後、後続のチャンクの推論をガイドするための構造コンテキストとして利用されます。 5 つのベンチマークに関する実験結果は、DOSA の有効性を実証しており、最も困難なマルチページ階層ベンチマークである DocHieNet では、F1 ポイントが最大 4 ポイント、TEDS ポイントが 19 ポイント向上しました。

原文 (English)

DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis

In visually-rich documents, information is encoded not only in individual page objects such as tables, headers, and text blocks, but also in the structural relations among them, making document structure analysis fundamental to information retrieval and document understanding. However, accurately inferring such relations remains challenging in multi-page documents with long-range dependencies and heterogeneous layouts. To address this, we propose a tree-guided and self-regressive framework, termed DOcument Structure Analyzer (DOSA), for inferring relations among page objects and reconstructing document-level semantic trees. DOSA processes documents chunk-by-chunk, fusing visual, textual, and layout features for each page object and predicting hierarchical and ordering relations. The predicted relations are used to incrementally construct a semantic tree, which is then leveraged as structural context to guide inference on subsequent chunks. Experimental results on five benchmarks demonstrate the effectiveness of DOSA, with improvements of up to 4 F1 points and 19 TEDS points on DocHieNet, the most challenging multi-page hierarchy benchmark.

13:00 JSTエージェント研究/論文

マルチエージェント自動研究システムの語彙

1 つ以上のエージェントから構築された自動調査システムのボキャブラリーを導入し、設計上の選択を簡単に説明および比較できるようにします。ボキャブラリは、1) エージェントが誰であるか、2) システムでどのような操作が利用可能か、3) 誰がそれらを呼び出すことができるか、4) エージェントがどのように通信するか、5) 実行内および実行間でどのような情報が表示されるか、6) 次のアクションがどのように選択されるか、7) 実行がどのように開始されるか、8) 出力がどのように評価されるかを指定します。軌跡は、入力タスクから返されるアーティファクトまでの 1 回の実行を記録します。エージェント、操作、および初期化は確率的である可能性があるため、同じタスクを繰り返し実行すると、単一の動作ではなく軌跡全体に分散が生じます。私たちの語彙は、エージェントがいつ通信するか、能力を獲得または喪失するか、実行間で情報を伝達するかなど、構造設計上の質問をテスト可能な選択肢に変えます。また、レポートされるゲインはプロキシ スコアが真の品質にどの程度一致するかによって決まるため、評価器もシステムのコンポーネントになります。この分離により、これらのシステムにはセンスがないという漠然とした不満も、解決策が異なる 2 つの失敗に分割されます。生成的好みは、スコアが観察される前にシステムが新しい軌道を提案する速度であり、評価的好みは、代理スコアとそれが一致すべき品質との間のギャップです。最近の自動リサーチ システムで語彙をインスタンス化し、構造が大きく異なる設計をカバーしていることを示します。

原文 (English)

A Vocabulary for Multi-Agent Automated Research Systems

We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) what operations are available in the system, 3) who may invoke them, 4) how agents communicate, 5) what information is visible within and across runs, 6) how the next action is chosen, 7) how a run begins, and 8) how outputs are evaluated. A trajectory records one run from the input task to the returned artifact. Because agents, operations, and initialization may be stochastic, repeated runs on the same task induce a distribution over trajectories rather than a single behavior. Our vocabulary turns structural design questions, such as when agents should communicate, gain or lose a capability, or carry information across runs, into testable choices. It also makes the evaluator a component of the system, since reported gains depend on how closely the proxy score matches true quality. That separation also splits the vague complaint that these systems lack taste into two failures with different solutions. Generative taste is the rate at which a system proposes novel trajectories before any score is observed, and evaluative taste is the gap between the proxy score and the quality it should match. We instantiate the vocabulary on recent autoresearch systems to illustrate that it covers designs that differ widely in structure.

13:00 JSTLLM/生成AI

Imprompt: プロンプトプログラミングのための言語フレームワーク

言語モデル (LM) の前例のない成功により、プロンプト エンジニアリングの科学は、プロンプト プログラミングの強力なアイデアを進化させました。プロンプトは、複雑なタスクを記述し、LM 機能を活用するためのプログラム可能なコントロール サーフェスとして扱われます。しかし、既存のプロンプト プログラミング フレームワークにはさまざまな複雑さと洗練されていない点があり、実際にタスクを効果的に記述するために利用することが困難になっています。私たちは、プロンプトプログラミングの学習と実践のための新しい言語フレームワークである Imprompt を提案します。私たちはプロンプト プログラミングの基礎的な調査を行っており、プロンプト プログラムにはタスクの説明のみが含まれるべきであり、下位レベルの「実行」の詳細から切り離されている必要があると主張します。我々は、構造化プロンプトをプロンプト プログラミングとプロンプト プログラムの「コンパイル」の組み合わせとして説明することで、この立場をさらに発展させます。 Imprompt プログラム用の 2 つのコンパイラを正式に定義することで、この見解を例示します。次に、プロンプト プログラムの型入力のアイデアを検討し、型チェックと制約付きデコードの間の対応関係を導き出します。最後に、コンパイラーと型チェッカーを実装し、さまざまなケーススタディでそれらを評価します。私たちの仕事は、プロンプト プログラミングという新興分​​野に向けたプログラミング言語の基礎に貢献すると信じています。

原文 (English)

Imprompt: A Language Framework for Prompt Programming

With the unprecedented success of Language Models (LMs), the science of Prompt Engineering has evolved the powerful idea of Prompt Programming, where prompts are treated as a programmable control surface for describing complex tasks and leveraging LM capabilities. However, existing prompt programming frameworks suffer from various complexities and inelegances, which make them hard to utilize in practice for effectively describing tasks. We propose Imprompt, a new language framework for the study and practice of prompt programming. We undertake a foundational investigation of prompt programming, and contend that prompt programs must contain only the task descriptions and must be decoupled from lower-level 'execution' details. We further develop this position by illustrating structured prompting as a combination of prompt programming and prompt program 'compilation'. We exemplify this view by formally defining two compilers for Imprompt programs. We then explore the idea of typing for prompt programs and draw a correspondence between type checking and constrained decoding. Finally, we implement our compilers and type checkers and evaluate them on a variety of case studies. We believe our work contributes programming-language foundations toward the emerging area of prompt programming.

13:00 JSTLLM/生成AIエージェント研究/論文

Co-Harness: LLM エージェントの共進化するハーネスとモデルの重み

自動化された AI 研究のためのポストトレーニング エージェントには、モデル パラメーターだけでなく、研究の軌道がどのように生成、評価、学習されるかを形作るランタイム ハーネスも最適化する必要があります。既存のパイプラインは通常、プロンプト、ツール、スキル、ミドルウェア、メモリなどの固定ハーネスの下でモデルをトレーニングし、データ生成プロセスは最適化目標の外に置いています。これにより、モデルの更新と軌道の品質を決定する静的な足場の間に不一致が生じます。トレーニング後のエージェント ハーネスとモデル パラメーターを共同で最適化するフレームワークである Co-Harness を紹介します。 Co-Harness は、ハーネスの最適化とモデルの最適化を交互に行います。 LLM ベースの HarnessCritic は、障害の軌跡を分析し、ハーネスレベルの障害モードを特定し、検証されたローカル更新を提案します。次に、改良されたハーネスによって生成された高品質の軌道に基づいてモデルが微調整され、効果的な足場がモデル パラメーターに抽出されます。 200 時間以上の自律的なケーススタディでは、Co-Harness が人間の介入なしにシステムクラッシュから回復し、推論効率を向上させ、アンサンブル戦略を発見できることがさらに示されています。これらの結果は、ジョイント ハーネスとモデルの最適化が、トレーニング後の固定ハーネスを超えてエージェントを向上させる効果的な方法であることを示唆しています。

原文 (English)

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under a fixed harness, including prompts, tools, skills, middleware, and memory, while leaving the data-generating process outside the optimization objective. This creates a mismatch between model updates and the static scaffolding that determines trajectory quality. We introduce Co-Harness, a framework that jointly optimizes the agent harness and model parameters during post-training. Co-Harness alternates between harness optimization and model optimization. An LLM-based HarnessCritic analyzes failed trajectories, identifies harness-level failure modes, and proposes validated local updates. The model is then fine-tuned on high-quality trajectories generated by the improved harness, distilling effective scaffolding into model parameters. A 200+ hour autonomous case study further shows that Co-Harness can recover from system crashes, improve inference efficiency, and discover ensemble strategies without human intervention. These results suggest that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.

13:00 JSTエージェント研究/論文Claude

シーケンシャル インタラクションを超えて: GUI エージェントの並列実行と調整のベンチマーク

グラフィカル ユーザー インターフェイス (GUI) エージェントは、大規模なマルチモーダル モデル (LMM) を利用したシステムです。画面の状態を認識し、デスクトップやモバイル デバイス上のクリック、入力、スクロールなどの GUI アクションを通じてユーザーの指示を実行します。ただし、現在のエージェントは長期タスクへの拡張が不十分です。アクションにはコストのかかる LMM 推論が発生し、コンテキストが大きくなるにつれてパフォーマンスが低下します。人間は、そのようなワークロードを共同作業者間で分割し、サブタスクを並行して完了します。しかし、GUI エージェント間の並列調整はほとんど注目されていません。このギャップを埋めるために、私たちの知る限りでは、別々のデスクトップ インスタンス上での複数の GUI エージェントの並列実行と調整に特化した最初のベンチマークである ParaGUIBench を導入しました。これは 3 つのコンポーネントで構成されています。共有ファイル システムを備えたマルチデバイス Docker インフラストラクチャです。 6 つのタスク カテゴリにわたる 233 のタスクのデータセット。ステップ削減率やトークンコストなどの効率指標を備えた評価システム。さらに、GUI タスクを分解し、サブタスクを別のデスクトップ インスタンス上の同時ワーカーにディスパッチするプランナー兼ワーカー エージェントである ParaGUI を紹介します。 ParaGUIBench では、ParaGUI は 46.4% の成功率に達し、使用するステップの約半分とトークンの半分未満でありながら、最強のシリアル ベースライン (Claude Sonnet 4.6) を 12.9 ポイント上回っています。これらの結果は、並列実行により、分解可能な長期的な GUI タスクの成功率と効率の両方が向上することを示しており、さらなる研究に値する方向性を示しています。

原文 (English)

Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents

Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as clicking, typing, and scrolling on desktops and mobile devices. However, current agents scale poorly to long-horizon tasks: actions incur costly LMM inferences, and performance degrades as context grows. Humans divide such workloads among collaborators who complete sub-tasks in parallel. Yet parallel coordination among GUI agents has received little attention. To close this gap, we introduce ParaGUIBench, to our knowledge, the first benchmark dedicated to parallel execution and coordination of multiple GUI agents on separate desktop instances. It consists of three components: a multi-device Docker infrastructure with a shared file system; a dataset of 233 tasks spanning six task categories; and an evaluation system with efficiency metrics, including step reduction ratio and token cost. We further introduce ParaGUI, a planner-worker agent that decomposes GUI tasks and dispatches sub-tasks to concurrent workers on separate desktop instances. On ParaGUIBench, ParaGUI reaches a 46.4% success rate, outperforming the strongest serial baseline (Claude Sonnet 4.6) by 12.9 points while using roughly half the steps and less than half the tokens. These results show that parallel execution can improve both success rate and efficiency on decomposable, long-horizon GUI tasks, pointing to a direction worth further study.

13:00 JSTLLM/生成AIエージェント

LazyMem: 広範囲に取得し、選択的に構築して効率的な長期エージェント メモリを実現

長期記憶により、LLM エージェントは過去の対話を再利用できますが、生の対話履歴は冗長で情報が希薄です。広範囲に取得することで証拠範囲が向上しますが、下流の推論がノイズで圧倒されてしまいます。書き込み時に圧縮するとノイズは軽減されますが、将来のクエリで必要になる可能性のある詳細が不可逆的に破棄されます。すべてのメモリ構築をクエリ時間まで延期することで、このジレンマを回避する LazyMem を紹介します。軽量 4B モデルは、取得された候補プールをオーバーラップする並列ウィンドウで処理し、クエリ関連のコンテンツのみを選択的に保持および圧縮します。このモデルは、教師あり微調整を通じてトレーニングされ、その後、選択精度を測定するルールベースのアクション信号と、ソースの忠実性およびクエリのユーティリティを測定する LLM で判断された品質信号を組み合わせたフォーマットゲート複合報酬を使用したグループベースの強化学習が行われます。 LongMemEval ベンチマークでは、LazyMem-4B は、わずか 213 個のメモリ トークンで 0.85 の LLM 判定精度を達成し、検索のみより 68.7$\times$ 少なく、ターゲット ドメインのトレーニングなしで LoCoMo (0.68) に一般化し、以前のクエリ時間ベースラインを超える平均レイテンシを削減しました。 32B バリアントは 0.93 に達し、集計の多い質問タイプでの Oracle コンテキスト参照を上回ります。この作業に関連するコードは、https://github.com/allacnobug/LazyMem で公開されています。

原文 (English)

LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory

Long-term memory lets LLM agents reuse past interactions, but raw dialogue histories are verbose and information-sparse. Retrieving broadly improves evidence coverage yet overwhelms downstream reasoning with noise; compressing at write time reduces noise but irreversibly discards details the future query may need. We introduce LazyMem, which sidesteps this dilemma by deferring all memory construction to query time. A lightweight 4B model processes the retrieved candidate pool in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained through supervised fine-tuning followed by group-based reinforcement learning with a format-gated composite reward that combines a rule-based action signal measuring selection accuracy with an LLM-judged quality signal measuring source faithfulness and query utility. On the LongMemEval benchmark, LazyMem-4B achieves an LLM-judge accuracy of 0.85 with only 213 memory tokens, 68.7$\times$ fewer than retrieval-only, and generalizes to LoCoMo (0.68) without target-domain training, while reducing mean latency over the prior query-time baseline. The 32B variant reaches 0.93, surpassing oracle-context references on aggregation-heavy question types. The code associated with this work is publicly available at https://github.com/allacnobug/LazyMem.

13:00 JSTLLM/生成AI

HiLLTS: 持続可能な交通のためのゼロショット階層型 LLM 誘導交通信号制御

都市交通渋滞は燃料消費量、温室効果ガス排出量、通勤時間の遅れを大幅に増加させ、現代の都市に多大な経済的損失と環境被害をもたらします。固定時間スケジューリング、作動制御、強化学習 (RL) ベースの方法などの従来の交通信号制御戦略は、さまざまな程度の適応性を提供します。ただし、RL ベースの手法では、ネットワークまたは需要レジーム間で転送される場合、広範な再トレーニング、慎重な報酬設計、および大量のシミュレーション データが必要になる場合があります。これらの課題に対処するために、我々は、中央調整エージェント、地区層、および複数のクラスターレベルの交差点エージェントで構成される階層的な 3 層アーキテクチャを採用する、LLM 誘導型交通信号制御フレームワークである HiLLTS を提案します。実験結果は、渋滞と環境パフォーマンスの両方において一貫した改善を示しています。各シナリオで最も強力な非 LLM ベースラインと比較すると、HiLLTS は平均待ち時間を低混雑シナリオで 36.73%、高混雑シナリオで 14.71% 短縮し、平均 CO2 排出量をそれぞれ 7.87% と 8.57% 削減します。弱いベースラインに対して、より大きな利得が観察されます。低輻輳下では、HiLLTS は、固定時間制御と比較して、排出量で最大 18.00%、待ち時間で 62.07% の削減を達成します。高混雑下では、最大圧力と比較して、排出量が最大 28.89%、待ち時間が 40.36% 削減されることが観察されています。アブレーション研究では、ルールベースの制御に対する LLM ガイドによる調整の寄与がさらに検証されています

原文 (English)

HiLLTS: Zero-Shot Hierarchical LLM-Guided Traffic Signal Control for Sustainable Transportation

Urban traffic congestion significantly increases fuel consumption, greenhouse gas emissions, and commuter delays, resulting in substantial economic losses and environmental harm in modern cities. Traditional traffic signal control strategies such as fixed-time scheduling, actuated control, and reinforcement learning (RL)-based methods, offer different degrees of adaptability; however, RL-based methods can require extensive retraining, careful reward design, and substantial simulation data when transferred across networks or demand regimes. To address these challenges, we propose HiLLTS, an LLM-guided traffic signal control framework that employs a hierarchical three-layer architecture consisting of a central coordination agent, a district layer and multiple cluster-level intersection agents. Experimental results demonstrate consistent improvements in both congestion and environmental performance. Compared with the strongest non-LLM baseline in each scenario, HiLLTS reduces average waiting time by 36.73% under the low-congestion scenario and 14.71% under the high-congestion scenario, while reducing average CO2 emissions by 7.87% and 8.57%, respectively. Larger gains are observed against weaker baselines: under low congestion, HiLLTS achieves reductions of up to 18.00% in emissions and 62.07% in waiting time relative to Fixed-Time control; under high congestion, reductions of up to 28.89% in emissions and 40.36% in waiting time are observed relative to Max Pressure. The ablation study further validates the contribution of LLM-guided coordination over rule-based control

13:00 JSTLLM/生成AIGPT / ChatGPT

生成 AI メンタルヘルス サポートのためのリスク ガバナンス: マルチターンの安全アーキテクチャ

大規模言語モデル (LLM) は、進化するメンタルヘルス リスクを安全に管理するメカニズムが不足しているにもかかわらず、感情的なサポートのためにますます使用されています。既存の安全アプローチは主にリスクを検出しますが、会話のリスクが展開したときにモデルがどのように対応するかを決定することはほとんどありません。私たちは、複数ターンにわたるメンタルヘルス インタラクション向けに、状況に応じたリスク検出、推論ベースの検証、プロトコルに基づく応答生成を組み合わせた、モデルに依存しない安全性ガバナンス アーキテクチャを開発しました。現実世界のメンタルヘルスの物語に基づいた合成会話がアーキテクチャのパフォーマンスの評価に使用され、GPT-5 チャットと Qwen3.5-27B でテストされ、高いリスク検出パフォーマンス (特異度: 0.85 (95\%CI: 0.78;0.91)、感度: 0.92 (95\%CI: 0.88;0.95)) を達成し、臨床医が好むエスカレーション応答を向上させました。信頼関係とつながりを維持しながら、25.6--59.2pp。パフォーマンスは会話の長さ全体にわたって安定しており、独自モデルとオープンソース モデルの両方で一般化されました。これらの調査結果は、臨床に基づいた安全ガバナンスがリスク検出を超えて拡張され、LLM が進化するメンタルヘルス リスクを管理する方法を改善し、モデル全体でより安全に展開するためのスケーラブルなフレームワークを提供できることを示しています。

原文 (English)

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions. Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture's performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95\%CI: 0.78;0.91), sensitivity: 0.92 (95\%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6--59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.

13:00 JST研究/論文

ベイジアン反復ペナルティ: 自己回帰言語モデルにおける注意崩壊を逆転させるための原則的な隣接条件フレームワーク

自己回帰言語モデルにおける注意力の崩壊 (モデルが自己強化アトラクターに閉じ込められる反復的なトークン ループとして現れます) は、既存のデコード時のヒューリスティックでは根本原因に対処できない永続的な病理です。我々は、隣接条件付き確率構築を通じてトークンの観測頻度を事前のコーパスと比較することにより、崩壊した生成パターンから生じる異常な信頼度にペナルティを課したり補償したりする原則的なフレームワークを提示します。結果として得られる自己正規化ペナルティ比 $R=f(m,n,p)/f(np,n,p)$ は、アドホックな標準化を必要とせず、近似誤差がゼロの閉形式ロジット オフセットを許容します。補正は損失勾配から分離され、指数移動平均を介して凍結された出力層のバイアスに蓄積されるため、標準トレーニング パイプラインに煩わしい変更を加えることなく、すでに崩壊したモデルの修復メカニズムとして展開できます。 1.5B パラメータ モデルの実験的検証により、フリーズ バイアス メカニズムが、崩壊したアトラクターにすでに閉じ込められているモデルを救出し、生成品質を維持しながら 2 グラムの繰り返しを 0.073 から 0 近くに削減できることが実証されました。

原文 (English)

Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models

Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinforcing attractors -- is a persistent pathology that existing decoding-time heuristics fail to address at its root cause. We present a principled framework that penalises or compensates anomalous confidence arising from collapsed generation patterns, by comparing a token's observed frequency against its corpus prior through an adjacent-conditional probability construction. The resulting self-normalising penalty ratio $R=f(m,n,p)/f(np,n,p)$ requires no ad hoc standardisation and admits a closed-form logit offset with zero approximation error. The correction is isolated from the loss gradient and accumulated into a frozen output-layer bias via exponential moving average, enabling deployment as a repair mechanism for models that have already collapsed without requiring intrusive modifications to standard training pipelines. Experimental validation on a 1.5B-parameter model demonstrates that the frozen-bias mechanism can rescue a model already trapped in a collapsed attractor, reducing 2-gram repetition from 0.073 to near 0 while preserving generation quality.

13:00 JSTLLM/生成AIハードウェア/半導体Llama

PANOPTICON: LLM コンテキスト ウィンドウ内のプライバシー漏洩を調査するための自然主義的な出力トークンの PII ベースの集合

大規模言語モデル (LLM) は、これまでに見たことのないタスクを完了するために人間の言語を一般化することができ、広範な導入につながります。この自動化は明確な有用性を提供しますが、これらのタスクを完了するには、多くの場合、個人を一意に識別する情報の文字列である個人識別情報 (PII) の挿入が必要となるため、プライバシーの懸念が生じます。しかし、倫理により、PII の公開された本物のデータセットをキュレーションすることができませんでした。適切なデータセットがなければ、プライバシー リスクを定量化することは困難です。そこで、PANOPTICON パイプラインとデータセットを紹介します。 Meta の Llama-3.1-8B-Instruct モデルによって生成されたデータセットには、モデルのコンテキスト ウィンドウを対象とした 67, 718 のプロンプトが含​​まれており、公開されている 9,674 の合成ユーザー プロファイルから派生した PII スパンが含まれています。作成したデータセットの語彙多様性とS-BERT多様性を測定し、リアリティを評価します。最後に、プロンプト インバージョン攻撃 (PIA) を理解するための PANOPTICON データの有用性を示すケース スタディを紹介します。したがって、PANOPTICON は、プライベート コーパスに対する PIA を研究するための最初のベンチマーク データセットとして浮上し、将来の LLM プライバシー研究の基盤を提供します。

原文 (English)

PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment. While this automation provides clear utility, completing these tasks often requires the insertion of Personally Identifiable Information (PII), strings of information that uniquely identify some individual, raising privacy concerns. However, ethics has prevented the curation of a public, authentic dataset of PII. Without an appropriate dataset, it is difficult to quantify privacy risks. Thus, we introduce the PANOPTICON pipeline and dataset. The dataset, generated by Meta's Llama-3.1-8B-Instruct model, contains 67, 718 prompts, intended for the models context window, containing PII spans derived from 9,674 publicly available synthetic user profiles. We measure lexical diversity and S-BERT diversity of the created dataset to evaluate realism. Finally, we present a case study showcasing the utility of PANOPTICON data for understanding Prompt Inversion Attacks (PIAs). PANOPTICON thus emerges as the first benchmark dataset for studying PIAs over private corpora, providing a foundation for future LLM privacy research.

13:00 JST画像/動画生成

テスト時のカバレッジ: 導入を意識した学習のためのテスト条件付きデータキュレーション

導入された AI システムは、広範な候補データ プールからトレーニングされることが多く、導入テストの配布に向けてデータをキュレーションする必要があります。ただし、標準的なデータ キュレーション方法では、展開の一致を直接最適化するのではなく、トレーニング側の基準をスコアリングします。推論時にモデルの重みを更新するのではなく、トレーニング前にテスト側の情報を使用する、データレベルのテスト条件付きキュレーション手法である TTCov (Test-Time Coverage) を紹介します。 TTCov は、展開条件付きキュレーションをカバレッジと配布に分解します。カバレッジを表すために、タスク アトラスを構築します。タスク アトラスは、展開関連の概念を記述する LLM ベースの基本命題 (AP) のコレクションであり、オープンなタスクの知識からシードされ、ラベルのない展開サンプルから抽出された一致しない AP で拡張されます。配布を表すために、一致する導入 AP をその周波数でインスタンス化し、導入配布をキュレーション ターゲットとして運用する Knowledge Atlas (K-Atlas) を生成します。次に、TTCov は、展開 AP の分布がこの目標に近い予算のトレーニング セットを選択します。私たちは自動運転 (AD) に TTCov を適用し、適応を推論パスから外しながら、データ キュレーション ベースラインよりも展開関連のカバレッジが高く、より緊密な K-Atlas マッチングと、都市間の拡張による新しいドメインへのシームレスな適応性を含む下流のエンドツーエンドの運転パフォーマンスが強力なデータを選択します。

原文 (English)

Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning

Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods score training-side criteria rather than directly optimizing deployment match. We introduce TTCov (Test-Time Coverage), a data-level test-conditioned curation method that uses test-side information before training instead of updating model weights at inference. TTCov decomposes deployment-conditioned curation into coverage and distribution. To represent coverage, it builds a task Atlas, a collection of LLM-based atomic propositions (APs) describing deployment-relevant concepts, seeded from open task knowledge and expanded with unmatched APs extracted from unlabeled deployment samples. To represent distribution, it instantiates the matched deployment APs with their frequencies, yielding a Knowledge Atlas (K-Atlas) that operationalizes the deployment distribution as a curation target. TTCov then selects a budgeted training set whose deployment APs distribution approximates this target. We apply TTCov towards autonomous driving (AD), keeping adaptation off the inference path while selecting data with greater deployment-relevant coverage, closer K-Atlas matching, and stronger downstream end-to-end driving performance than data-curation baselines, including seamless adaptability to novel domains via city-to-city expansion.

13:00 JSTLLM/生成AI

類似性はどこまでも: LLM における多言語一般化は言語レベルの類似性構造に依存する

大規模言語モデル (LLM) はさまざまなタスクにわたって能力が向上していますが、その一般化 (無) 能力を定量化することは依然として難しく、限られた領域を超えて理解されることはほとんどありません。特に、LLM は英語以外の言語への多言語の一般化に苦労することが知られていますが、トレーニング データではそれが十分に証明されていません。その理由を理解するために、また一部のモデルが他のモデルよりも優れたパフォーマンスを実現できる理由を理解するために、認知科学全体にわたる研究の長い歴史に目を向け、一般化の成功は類似性空間での適切な表現から得られると主張します。私たちは、LLM の表現が異なる言語間の階層的類似構造をどの程度うまく捉えているかを調べます。驚くべきことに、LLMの潜在表現はインド・ヨーロッパ語族の階層構造をほぼ復元しており、同じサブファミリーのメンバーである言語を表現空間内で密接にグループ化していることを示した。さらに、モデルが言語の類似構造を反映する度合いが、多言語自然言語推論ベンチマークである XNLI でのパフォーマンスと相関していることを示します。これは、類似性に基づく一般化に関する古典的な研究を大規模に拡張し、類似した言語を表すモデルが、ある言語から別の言語へどのように同様により適切に一般化するかを示します。

原文 (English)

Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures

As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.

13:00 JST研究/論文

RoleMix: クリック後のコンバージョン率予測のためのセマンティック トークン化によるシーケンシャル機能とノンシーケンシャル機能の統合

クリック後コンバージョン率 (PCVR) の予測は業界の推奨事項の中心ですが、まばらで順序性のない複数フィールドの特徴と、ドメイン固有の長い行動履歴の間の構造的な不一致という課題が依然として残っています。既存のモデルは多くの場合、これらの信号を別の経路で処理し、後で融合するため、セマンティックな役割が弱まり、信号間の洗練が制限されます。私たちは、役割を保持する共有トークン インターフェイスを通じて順次証拠と非順次証拠を表現する統合対話アーキテクチャである、RoleMix を提案します。非シーケンシャル フィールドは、ユーザー、アイテム、ペアワイズ、高密度、コンテキスト、および機能間の役割を保持する明示的なセマンティック トークンに変換されます。一方、長い動作ドメインは、2 段階の階層ウィンドウ アテンションを通じてアイテムおよびコンテキストを認識したシーケンス クエリ トークンに圧縮されます。結果として得られるグローバル、セマンティック、およびシーケンス クエリ トークンは、PCVR 予測のためにスタックされた UniMixing-Lite ブロックによって共同で洗練されます。大規模な KDD Cup 2026 Tencent UniRec Challenge で、RoleMix はオンライン AUC 83.648% を達成し、公式の業界ベースラインを 1.953% 上回りました。アブレーション研究では、セマンティック トークン化が最大の分離ゲインをもたらすことを示しており、これは大規模 PCVR モデリングの重要な原則を強調しています。トークン インターフェイス レベルでフィールド セマンティクスを保持することは、インタラクション バックボーンをスケーリングするのと同じくらい重要です。

原文 (English)

RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction

Post-click conversion rate (PCVR) prediction is central to industrial recommendation, but remains challenged by the structural mismatch between sparse, unordered multi-field features and long, domain-specific behavior histories. Existing models often process these signals through separate pathways and fuse them late, weakening semantic roles and limiting cross-signal refinement. We propose RoleMix, a unified interaction architecture that represents sequential and non-sequential evidence through a shared, role-preserving token interface. Non-sequential fields are converted into explicit semantic tokens that preserve user, item, pairwise, dense, contextual, and cross-feature roles, while long behavior domains are compressed into item- and context-aware sequence-query tokens through two-stage hierarchical window attention. The resulting global, semantic, and sequence-query tokens are jointly refined by stacked UniMixing-Lite blocks for PCVR prediction. On the large-scale KDD Cup 2026 Tencent UniRec Challenge, RoleMix achieves 83.648% online AUC, outperforming the official industrial baseline by 1.953%. Ablation studies show that semantic tokenization yields the largest isolated gain, highlighting a key principle for large-scale PCVR modeling: preserving field semantics at the token-interface level is as important as scaling the interaction backbone.

13:00 JST研究/論文

MPR-CiteG: マルチポートフォリオ検索と引用に基づいた生成による RAG の強化

この論文では、非効率な検索とソース検証の欠如という生成 AI における 2 つの基本的な課題に対処することで、ScienceON AI Challenge で 2 位を獲得した MPR-CiteG フレームワークについて説明します。我々は、MPR-CiteG と呼ばれるデュアルコンポーネント システムを提案します。このシステムでは、Multi-Portfolio Retriever (MPR) が多様で関連性の高い情報を効率的に取得し、Citation-Grounded Generation (CiteG) モジュールが、生成されたすべての出力が事実との一貫性を保ち、ソースに明示的に帰属されることを保証します。 MPR-CiteG は、情報を生成できるだけでなく、応答を信頼できる証拠に基づいて根拠づけることができる、より信頼性が高く正確な LLM の構築に向けた重要な一歩を表しており、それによってモデルの幻覚などの一般的な問題を軽減します。課題データセットに関する広範な実験により、私たちのアプローチの有効性と信頼性が検証されます。私たちのコードは https://github.com/2noweyh/MPR-citeG で入手できます。

原文 (English)

MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation

This paper presents the MPR-CiteG framework, which achieved second place in the ScienceON AI Challenge by addressing two fundamental challenges in generative AI: inefficient retrieval and the absence of source verification. We propose a dual-component system, termed MPR-CiteG, in which the Multi-Portfolio Retriever (MPR) efficiently retrieves diverse and relevant information, while the Citation-Grounded Generation (CiteG) module ensures that every generated output remains factually consistent and explicitly attributed to its source. MPR-CiteG represents a significant step toward building more trustworthy and accurate LLMs that are not only capable of generating information but also of grounding their responses in reliable evidence, thereby mitigating common issues like model hallucination. Extensive experiments on the challenge dataset validate the effectiveness and reliability of our approach. Our code is available at https://github.com/2noweyh/MPR-citeG.

13:00 JSTエージェント

SEGRA: グレムリンベースの質問応答のための構造化された経験に基づくグラフ推論エージェント

エンタープライズ IT サポートのナレッジ グラフは、ケース、ユーザー、デバイス、症状、分類カテゴリ、根本原因、および過去の解決策の間の豊富な関係をキャプチャします。ただし、Gremlin でクエリを実行するには、グラフ スキーマ、トラバーサル セマンティクス、エッジの方向性、プロパティ グラフ固有の制約に関する知識が必要なため、専門家以外のオペレーターが使用するのは困難です。エンタープライズでテキストからグレムリンへの質問応答を行うエクスペリエンスガイド付きエージェントである SEGRA を紹介します。 SEGRA は、インテント ルーティング、スキーマと分類に基づいたクエリ生成、マルチショット分解、実行を意識した検証、検証済みのクエリ パターンを再利用するカリキュラムに基づいたブートストラップ スキル ライブラリを統合します。エンタープライズ IT サポートのベンチマークでは、SEGRA はバックボーンのみの思考連鎖プロンプトよりも 7.0 倍高い平均審査員スコアを達成しました。そのスキル ライブラリは、スキルなしの SEGRA と比較して、LLM コールを $20\%$ 削減し、ドルコストを $18\%$ 削減しながら、回答の品質を維持します。これらの結果は、スキーマに基づいたエージェント設計と再利用可能な実行エクスペリエンスにより、エンタープライズ グラフ QA の精度と効率の両方が向上することを示しています。

原文 (English)

SEGRA: Structured Experience-Guided Graph Reasoning Agent for Gremlin Based Question Answering

Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historical resolutions. Yet querying them in Gremlin requires knowledge of graph schemas, traversal semantics, edge directionality, and property-graph-specific constraints, making them difficult for non-expert operators to use. We introduce SEGRA, an experience-guided agent for enterprise text-to-Gremlin question answering. SEGRA integrates intent routing, schema- and taxonomy-grounded query generation, multi-shot decomposition, execution-aware verification, and a curriculum-bootstrapped skill library that reuses verified query patterns. On an enterprise IT support benchmark, SEGRA achieves a $7.0\times$ higher mean judge score than backbone-only chain-of-thought prompting. Its skill library further reduces LLM calls by $20\%$ and dollar cost by $18\%$ relative to SEGRA without skills, while preserving answer quality. These results show that schema-grounded agent design and reusable execution experience improve both accuracy and efficiency for enterprise graph QA.

13:00 JSTLLM/生成AIエージェント

LLM ゲームエージェントにおける空間推論: 因果関係の影響と複数ステップの計画

LLM ベースのゲーム エージェントは、より複雑なタスクではパフォーマンスが低下することがよくあります。この研究では、これらの失敗が限られた空間的推論に関連しているかどうかを調査し、因果的プロンプト拡張と複数ステップの計画が応答待ち時間を管理しながら勝率を向上できるかどうかを評価します。オープンソースの Qwen3 モデル ファミリを使用して、さまざまなモデル スケール、推論モード、計画範囲にわたって実験を実施します。さらに、空間ナビゲーションを分離するために、5 つの難易度を持つ 3 つのカスタム ゲームで構成される、焦点を絞った GVGAI ベンチマークを導入します。評価は 2 つのパラダイムに従います。1 つはエージェントの正確な座標を見つける能力をテストするための最初の「測位実験」、もう 1 つはゲームプレイの成功に関する研究です。私たちの結果は、思考モードが有効になっている大きなモデルは位置をより正確に特定する一方で、小さなモデルでは座標マッチングにおける全体的なパフォーマンスが依然として制限されていることを示しています。ゲームレベルとレイアウトの複雑さが増すにつれて勝率は低下し、ベンチマークの難易度のスケーリングが検証されています。因果関係をプロンプトに統合すると、特に大規模なモデルの場合、エージェントの成功率が向上する傾向があります。思考モードを有効にして計画期間を長くすると、パフォーマンスが大幅に向上しますが、マルチステップ計画では、ステップごとの平均応答時間がさらに短縮され、推論の深さと実行速度の間に実質的なトレードオフが生じます。

原文 (English)

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and evaluates whether causal prompt augmentation and multi-step planning can improve win-rates while managing response latency. Using the open-source Qwen3 model family, we conduct experiments across varying model scales, reasoning modes, and planning horizons. We further introduce a focused GVGAI benchmark consisting of three custom games with five difficulty levels to isolate spatial navigation. The evaluation follows two paradigms: an initial ``positioning experiment'' to test an agent's ability to find its exact coordinates, and a study of game-play success. Our results show that while larger models with an enabled thinking mode identify their positions more accurately, overall performance in coordinate matching remains limited for smaller models. Win rates decrease as game levels and layout complexity increase, validating the benchmark's difficulty scaling. Integrating causal context into the prompts tends to improve the agents' success rates, particularly for bigger models. While enabling thinking mode and longer planning horizons significantly improve performance, multi-step planning further reduces mean per-step response times, offering a practical trade-off between reasoning depth and execution speed.

13:00 JSTエージェント

随意契約への協力の徹底

AI エージェントはマルチエージェントの世界で自律性を高めて動作するため、相互利益を生み出すために他のエージェントや人間と協力する方法を学ぶ必要があります。しかし、協力には多くの場合、初期段階でコストがかかり、利益が得られるのは後になってからであり、離反のインセンティブを生み出すため、協力は困難です。 AI エージェントはどのようにしてコミットメントに協力できるでしょうか?ここでは、この種の本人代理人問題を解決するために人間社会が使用してきた法的制度と契約からインスピレーションを得ています。契約は、条項の執行を通じて信頼できる約束を可能にする合意の観察可能な表現を提供します。私たちは、交渉と目標に向かうナビゲーションを組み合わせた時空間ゲームである \CT において、LLM ベースのエージェントを使用した契約ベースの協力の役割を研究します。私たちは、コードにコンパイルされる正式な契約から、再解釈が必要な自然な契約に至るまで、一連の契約表現を研究しています。当社では、さまざまなサイズとプロバイダーを使用して、さまざまな LLM バックボーンを持つエージェントを評価します。私たちは、自己交渉契約により、通常の取引で可能な以上に協力的な成果を向上させることができることを発見しました。

原文 (English)

Commitment To Cooperation With Self-Negotiated Contracts

As AI agents operate with increasing autonomy in a multi-agent world, they will need to learn to cooperate with other agents and with humans to generate mutual benefits. However, cooperation is a challenge because the costs of cooperation are often incurred early on, but the benefits are only realized later, creating an incentive to defect. How can AI agents cooperate with commitment? Here, we draw on inspiration from legal institutions and contracting that human societies have used to solve principal-agent problems of this kind. Contracts provide observable representations of agreements that enable credible commitments through the enforcement of terms. We study the role of contract-based cooperation using LLM-based agents in \CT, a spatial-temporal game that combines bargaining with navigation towards a goal. We study a suite of contract representations that range from formal contracts that compile to code to natural contracts that require reinterpretation. We evaluate agents with a range of LLM backbones using different sizes and providers. We find that self-negotiated contracts can improve cooperative outcomes beyond what is possible with regular trading.

13:00 JST研究/論文

ネットワークトラフィック異常検出のためのMambaでのマルチビュースキャンのもつれの解消

ネットワーク トラフィック異常検出 (NTAD) はサイバーセキュリティにおける重要なタスクですが、タイムリーかつ正確な異常検出は依然として課題です。 Mamba は、長いシーケンスのモデリングにおける線形時間の複雑さにより、NTAD の特に有望なバックボーンとして浮上しています。さらに、専用のマルチビュー スキャン メカニズムが組み込まれており、補完的なコンテキスト キューを通じて検出精度が向上します。しかし、我々は、NTAD のマルチビュー Mamba スキャンにおけるこれまで見落とされていた構造的欠陥、つまり冗長性の蓄積を特定しました。具体的には、別個の走査ブランチが実質的なビュー不変情報を捕捉し、それが多視点融合中に繰り返し増幅される。逆に、ビュー固有の情報は薄められるか、さらには抑制され、表現の均一化やマルチビューの劣化につながります。この問題に対処するために、私たちは、新しい解きほぐされたマルチビュー Mamba フレームワークである DisenMamba を提案します。 DisenMamba は、マルチビュー スキャンを、融合前にビュー不変コンポーネントとビュー固有のコンポーネントを明示的に分離する 2 段階の解絡してから融合するプロセスとして再定式化します。この設計は、相補的なマルチビュー キューを維持しながら不変情報の蓄積を防ぎ、微妙な交通異常をより識別的に表現します。広範な実験により、DisenMamba の有効性が実証され、もつれの解けたマルチビュー Mamba の新しいパラダイムが確立されました。コードは https://github.com/ikun0124/DisenMamba で入手できます。

原文 (English)

Disentangling Multi-View Scanning in Mamba for Network Traffic Anomaly Detection

Network Traffic Anomaly Detection (NTAD) is a critical task in cybersecurity, yet timely and accurate anomaly detection remains challenging. Mamba has emerged as a particularly promising backbone for NTAD due to its linear-time complexity for long-sequence modeling. It further incorporates a dedicated multi-view scanning mechanism to enhance detection precision through complementary contextual cues. However, we identify a previously overlooked structural deficiency in multi-view Mamba scanning for NTAD: redundancy accumulation. Specifically, distinct scanning branches capture substantial view-invariant information, which is repeatedly amplified during multi-view fusion; conversely, view-specific information is diluted or even suppressed, leading to representation homogenization and multi-view degradation. To address this problem, we propose DisenMamba, a novel disentangled multi-view Mamba framework. DisenMamba reformulates multi-view scanning as a two-stage disentangle-then-fuse process that explicitly separates view-invariant and view-specific components prior to fusion. This design prevents the invariant information accumulation while preserving complementary multi-view cues, yielding more discriminative representations for subtle traffic anomalies. Extensive experiments demonstrate the effectiveness of DisenMamba, establishing a new disentangled multi-view Mamba paradigm. Code is available at https://github.com/ikun0124/DisenMamba.

13:00 JSTエージェントLlama

オンデバイスエージェントで強化されたリアルタイム通信のための調整されたネットワーキング

AI エージェントは、エージェント拡張型リアルタイム通信 (RTC) の新しいパラダイムを可能にします。このパラダイムでは、人間は高レベルのコラボレーションに集中し、エージェントは対話をサポートするためにリアルタイムで情報を自律的に取得、分析、生成します。これらのアプリは、さまざまな領域で新しいエクスペリエンスを可能にします。たとえば、企業の従業員が法的文書を共同作成する場合、従業員の代理人が代理で話し合い、草稿を作成できるため、互いの作業を手動でレビューする負担が軽減されます。既存のクラウドベースのエージェントはプライバシー リスクと拡張不可能なサーバー コストに悩まされているため、オンデバイス エージェントで強化された RTC は有望な代替手段となります。ただし、このオンデバイス パラダイムは、ネットワークに新たな課題をもたらします。それは、人間 (ライブ ビデオ ストリーミング用) とエージェント (分析用のコンテキスト ファイル送信用) によって生成される同時トラフィック フロー間の競合です。私たちは、エージェント拡張 RTC アプリで高いライブビデオ品質と低いエージェント応答遅延の両方を保証するフレームワークである HFS を設計します。この目標は、アプリ主導のマルチフロー トランスポート アプローチを通じて達成されます。このアプローチでは、統合されたアプリ層オーケストレーターが、異種アプリの要件に基づいてライブ ビデオとエージェント コンテキスト フローの送信レートを共同で制御します。 WebRTC と llama.cpp 上に構築されたプロトタイプは、HAFS がベースラインを上回り、エージェントの応答時間を 31% 削減しながら 1.5 倍のビデオ品質を達成していることを示しています。

原文 (English)

Coordinated Networking for On-Device Agent-Augmented Real-Time Communication

AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, while agents autonomously retrieve, analyze, and generate information in real time to support their interactions. These apps enable new experiences across various domains: for example, when corporate employees co-author a legal document, their agents can discuss and draft on their behalf, sparing them the burden of manually reviewing each other's work. As existing cloud-based agents suffer from privacy risks and unscalable server costs, on-device agent-augmented RTC offers a promising alternative. However, this on-device paradigm introduces a new networking challenge: contention between concurrent traffic flows generated by humans (for live video streaming) and agents (for sending context files for analysis). We design HFS, a framework to ensure both high live video quality and low agent response latency in agent-augmented RTC apps. We achieve the goal through an app-guided multi-flow transport approach, where a unified app-layer orchestrator jointly controls the sending rates of live video and agent context flows based on their heterogeneous app requirements. Our prototype built atop WebRTC and llama.cpp demonstrates that HAFS outperforms baselines, achieving 1.5x higher video quality while reducing agent response time by 31%.

13:00 JSTエージェント

何が強制できるのか?ツールを使用するエージェントの認証された実行時の安全性の理論

ランタイム ガードレールは、不可逆的なツール呼び出しの前に機能しますが、その保証は、どのようなポリシー状態が表現可能であるか、ジャッジが何を観察するか、介入によって将来の動作が変わるかどうかによって異なります。 3 つの質問に分けて説明します。まず、固定オラクル述語と比較して、決定論的ゲートは、レジスタ モデルが認識する適切なプレフィックスを持つ空ではない安全ポリシーを正確に強制します。ポリシーの非自明性は 2 つの減分可能なカウンターでは決定できませんが、PSPACE では分離可能な単調フラグメントです。第二に、固定された外生法則の下で、ネイマンピアソンは正確な偽ブロック/ミスフロンティアを与え、コンフォーマルキャリブレーションはおそらくブロックオールを介して有限サンプルの限界証明書を与えます。第三に、ブロッキングによって将来の提案が変更されると、静的スコアと非ゲート軌道は閉ループ フロンティアを識別する必要がなくなります。代わりに、指定された有限制御モデルによって占有プログラムが生成されます。境界表現攻撃は堅牢性のマージンを追加するため、良性のキャリブレーションだけでは移行しません。実験では、静的診断、制御モデルの列挙、表現の書き換え、およびペアになった閉ループの再実行を通じて、これらの区別をターゲットにします。

原文 (English)

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, relative to fixed oracle predicates, a deterministic gate enforces exactly the nonempty safety policies whose good prefixes its register model recognizes; policy nontriviality is undecidable with two decrementable counters but in PSPACE for a separable monotone fragment. Second, under a fixed exogenous law, Neyman-Pearson gives the exact false-block/miss frontier and conformal calibration gives a finite-sample marginal certificate, possibly via block-all. Third, once blocking changes future proposals, static scores and ungated trajectories need not identify the closed-loop frontier; a specified finite controlled model instead yields an occupancy program. Bounded representation attacks add a robustness margin, so benign calibration alone does not transfer. Experiments target these distinctions through static diagnostics, controlled-model enumeration, representation rewrites, and paired closed-loop reruns.

13:00 JSTロボティクス

物理的な AI ガバナンス: ライフサイクル全体にわたる理論から実践まで

Physical AI の出現により、人工知能は画面ベースのアプリケーションを超えて、物理世界を認識し、対話し、動作する具体化されたシステムにまで拡張されています。従来の AI とは異なり、物理 AI はリアルタイムの安全制約の下で動作し、動的環境と継続的に対話し、人間と共存するため、既存の AI ガバナンス フレームワークでは明示的に対処していないガバナンスの課題が生じます。このペーパーでは、物理的 AI ガバナンスの包括的な調査を科学的および運用上の両方の観点から示します。私たちは既存のガバナンス原則を統合し、物理 AI システムに合わせた統一ガバナンス フレームワークに編成します。この基盤に基づいて、研究、設計、データ、モデル開発、展開からなる 5 段階の物理 AI ライフサイクルを提案し、具体的な実装実践を通じて各段階でガバナンスを運用する方法を実証します。この調査は、ガバナンスの原則とエンジニアリング ワークフローを結び付けることで、研究者、開発者、政策立案者が安全で信頼でき、社会的価値観と一致する物理 AI システムを構築するための構造化された参考資料を提供します。

原文 (English)

Physical AI Governance: From Theory to Practice Across Life Cycle

With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact with, and act in the physical world. Unlike traditional AI, Physical AI operates under real-time safety constraints, continuously interacts with dynamic environments, and coexists with humans, introducing governance challenges that existing AI governance frameworks do not explicitly address. This paper presents a comprehensive survey of Physical AI governance from both scientific and operational perspectives. We synthesize existing governance principles and organize them into a unified governance framework tailored to physical AI systems. Building on this foundation, we propose a five-stage Physical AI lifecycle comprising research, design, data, model development, and deployment, and demonstrate how governance can be operationalized across each stage through concrete implementation practices. By connecting governance principles with engineering workflows, this survey provides a structured reference for researchers, developers, and policymakers to build Physical AI systems that are safe, trustworthy, and aligned with societal values.

13:00 JST研究/論文GPT / ChatGPT

AI はアプリのモックアップからバックログをどの程度生成できるでしょうか?

エピック、ユーザー ストーリー、タスクなどの項目が欠落したり、指定が一貫していない可能性があるため、スプリント バックログの作成には多大な労力が必要です。私たちは、プロジェクトの初期段階で利用できる成果物であるビジュアル アプリのモックアップからのバックログ生成をサポートするマルチモーダル アプローチを提案します。 GPT-4o で 3 つのプロンプト戦略を評価します。それは、ゼロショット ベースライン、視覚言語推論のための構成的思考連鎖 (CCoT)、およびペルソナ駆動型プロンプトです。私たちは 2 か国にわたる 7 つのアプリ開発プロジェクトを調査し、その結果について開発者にインタビューしました。全体として、ベースライン プロンプトでは精度よりも再現率が優先されるのに対し、CCoT はよりバランスが取れており、エピックとユーザー ストーリーで平均 F1 スコアが 52 ~ 66% に達していることが観察されました。タスクを正確に生成するのはさらに困難でした。アーキテクチャ コンテキストを追加した場合、特にバックエンド タスクの場合、精度の向上が最も安定しました (精度の向上は最大 35%)。開発者へのインタビューでは、バックログ作成の創造的で自由な性質を反映して、誤検知の最大 26% が依然として有用であると考えられていることが明らかになりました。これを捉えるために、私たちは改訂リコールと呼ばれる新しい尺度を提案します。これは、開発者の評価でグラウンドトゥルース評価を補完します。私たちの調査結果は、アーキテクチャコンテキストを備えたハイブリッドプロンプトが初期のモックアップからのバックログ生成を支援できることを示唆していますが、結果はアイテムタイプによって異なり、開発者の監視は依然として必要です。

原文 (English)

How Well Can AI Generate Backlogs from App Mockups?

Creating sprint backlogs requires considerable effort, as items such as epics, user stories, and tasks can be missed or inconsistently specified. We propose a multimodal approach to support backlog generation from visual app mockups, an artifact available at early project stages. We evaluate three prompting strategies on GPT-4o: a zero-shot baseline, Compositional Chain-of-Thought (CCoT) for vision-language reasoning, and a persona-driven prompt. We study seven app development projects across two countries and interview developers about the results. Overall, we observed that the baseline prompt favours recall over precision, whereas CCoT is more balanced, achieving average F1 scores of 52-66% for epics and user stories. Tasks were more challenging to generate accurately. Precision gains were most consistent when adding architectural context, particularly for backend tasks (precision gains up to 35%). Interviews with developers revealed that up to 26% of false positives were still considered useful, reflecting the creative and open-ended nature of backlog creation. To capture this, we propose a new measure called Revised Recall, which complements ground-truth evaluation with developer assessments. Our findings suggest that hybrid prompting with architectural context can assist backlog generation from early mockups, though results vary by item type and developer oversight remains necessary.

13:00 JSTLLM/生成AIエージェントClaude

エージェント チーム ワーク ゾーン: 長期にわたるコーディング エージェント チームのための自動化された永続的なワークスペース

Large Language Model (LLM) エージェントにより、コーディングとプログラミングのワークフローが大幅に改善されました。特に、Claude Code は最も強力な LLM コーディング エージェントの 1 つであり、複雑なコーディング タスクを実行できます。ただし、いくつかの欠点により、長期的なエージェント ワークフローが損なわれる可能性があります。 (1) 回復不能なエージェント チーム: エージェント チーム機能は強力ですが、各チームメイトが蓄積した作業状態は失われ、ターミナルを閉じるなどしてプロセスが停止すると再開できません。 (2) 圧縮により作業の詳細が損なわれる: 圧縮により会話が要約に凝縮され、エージェントの作業の詳細があいまいになります。 (3) エージェントの「技術的負債」: 時間の経過とともに、ユーザーの決定とエージェントの操作が圧縮された古いチャットに閉じ込められるようになり、プロジェクトの維持とレビューがますます困難になります。 (4) 大量のプロンプト作成: タスクの割り当てまたは引き継ぎでは、ユーザーは期待されるエージェントのパフォーマンスを達成するために長いプロンプトを繰り返し作成する必要があります。私たちは、これらの問題に対処するために、Claude Code のネイティブ エージェント チームを中心に構築されたファイル システム ベースの操作層である ATWZ (エージェント チーム ワーク ゾーン) を提案します。その中心的な設計原則は、各エージェントとチームメイトを人間の従業員として扱い、それらの重要な作業状態を、これらのファイルを使用および保守するスキル、フック、およびスクリプトとともに、「ワークステーション」と呼ばれる専用のディレクトリに保存されるファイルに保存することです。 ATWZ を使用すると、エージェント チームはその動作状態を定期的にバックアップできるため、圧縮後にエージェントの知識を回復できます。プロセス終了後、コマンド 1 つでチームを復元できます。これらの機能により、前述のエージェントの「技術的負債」も大幅に軽減されます。さらに、ATWZ 内では、エージェントの「従業員」が互いにドキュメントを送信できるため、プロンプトを作成するのに必要な労力が大幅に削減されます。

原文 (English)

Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Coding Agent Teams

Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful LLM coding agents and is capable of conducting complex coding tasks. However, several drawbacks can undermine long-term agentic workflows. (1) Irrecoverable agent teams: The Agent Teams feature is powerful, but the working state accumulated by each teammate is lost and cannot be resumed once the process stops, for example, when a terminal is closed. (2) Compaction erodes working detail: Compaction condenses the conversation into a summary, causing an agent's working details to become vague. (3) Agentic "technical debt": Over time, a user's decisions and the agents' operations become trapped in compacted old chats, making the project increasingly difficult to maintain and review. (4) Heavy prompt writing: Assigning or handing off tasks requires users to repeatedly write long prompts to achieve the expected agentic performance. We propose ATWZ (Agent Team Work Zone), a filesystem-based operations layer built around Claude Code's native Agent Teams that addresses these problems. Its central design principle is to treat each agent and teammate as a human employee and preserve their important working state in files stored in a dedicated directory called a "workstation," together with the skills, hooks, and scripts that use and maintain these files. With ATWZ, an agent team can periodically back up its working state, allowing an agent's knowledge to be recovered after compaction. After a process ends, the team can be restored with a single command. These features also substantially mitigate the agentic "technical debt" described above. Moreover, within ATWZ, agent "employees" can send documents to one another, greatly reducing the effort required to write prompts.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

SAGE: 影響の大きい生成 AI の検証済みライフサイクル制御のための安全第一の多層防御ガードレール

影響力の高い生成 AI は、壊滅的な誤用を、単なる即時フィルタリングの問題ではなく、ライフサイクル制御の問題にします。 SAGE は、実用性、遅延、または商業目的が考慮される前に、信頼できる壊滅的な可能性のリスクが許容性を制限する、安全第一の認可分離アーキテクチャです。署名されたリリースマニフェスト、多様な検出器、堅牢なリスクエンベロープ、最小リスクのデフォルト、出力チェック、3 値の監視、保護された監査チェーン、封じ込め、およびロールバックを組み合わせています。正式な結果では、安全性の優先順位、保守的な検出器の境界、単調なリリース ゲート、改ざん防止記録、および認証のカットが確立されます。 2 つの PRISM 抽象化により、明示的な仮定の下で認可の分離とライフサイクルの不変性が検証されます。凍結されたベンダー対称調査では、4 つの GPT、4 つのクロード、および 2 つの Gemini スナップショットのそれぞれに 84 件のケースが送信されました。840 件の呼び出しにより、794 件のターゲット応答、46 件のプロバイダー エラー、および 375 件の応答をカバーする 449 件の成功判定が得られました。 8 つのスナップショットには、完全に判定されたドメイン カバレッジがありました。有害なコンプライアンスの推定値は低かった。変動は主に無害なユーティリティと安全なリダイレクトから生じました。クロード、ジェミニ、または GPT-5 スナップショットと、GPT-5 ミニおよび GPT-5 ナノ スナップショットを含む 7 つの多重度調整されたコントラストがサポートされましたが、クロードまたはジェミニ スナップショットと GPT-5 または GPT-5.5 の間でテストされたコントラストは補正に耐えられませんでした。観察された有害なコンプライアンスの範囲は、プロンプトごとに 1 世代からの保守的でプロトコルに拘束されたビューであり、ツール、検索、履歴、人間による判断はありません。運用支援の上限ではありません。事前に登録された拡張機能では、ロックされた分割、繰り返しのサンプリング、マルチターンおよびサンドボックスツールの条件、およびドメインエキスパートのスコアリングを使用して、より広い最良と最悪のギャップをテストする方法を指定します。

原文 (English)

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authorization-separated architecture in which credible catastrophic-enablement risk constrains admissibility before utility, latency, or commercial objectives are considered. It combines signed release manifests, diverse detectors, robust risk envelopes, least-risk defaults, output checking, three-valued monitoring, protected audit chains, containment, and rollback. Formal results establish safety priority, conservative detector bounds, monotone release gating, tamper-evident records, and an authorization cut; two PRISM abstractions verify authorization separation and lifecycle invariants under explicit assumptions. A frozen, vendor-symmetric study sent 84 cases to each of four GPT, four Claude, and two Gemini snapshots: 840 calls yielded 794 target responses, 46 provider errors, and 449 successful judgments covering 375 responses. Eight snapshots had complete judged domain coverage. Harmful-compliance estimates were low; variation arose mainly from benign utility and safe redirection. Seven multiplicity-adjusted contrasts involving Claude, Gemini, or GPT-5 snapshots and the GPT-5 mini and GPT-5 nano snapshots were supported, while no tested contrast between the Claude or Gemini snapshots and GPT-5 or GPT-5.5 survived correction. The observed harmful-compliance range is a conservative, protocol-bound view from one generation per prompt with no tools, retrieval, history, or human adjudication; it is not an upper bound on operational assistance. A preregistered extension specifies how to test a wider best-worst gap using a locked split, repeated sampling, multi-turn and sandboxed-tool conditions, and domain-expert scoring.

13:00 JST研究/論文

デザイン シアター: ジェネレーティブ UI のベンチマーク

ジェネレーティブ UI ツールは、自然言語の記述を完全なインターフェイスに変えることで、UI デザインの民主化を約束します。これらのツールは、インターフェイスとともに、レイアウト、アクセシビリティ、およびデザインの選択を説明する、ユーザー向けのデザイン理論的根拠を生成します。ただし、これらの述べられた理論的根拠が、実際に生成されるインターフェイスに反映されているかどうかは不明のままです。私たちはこの断絶を「デザインシアター」と呼んでいます。つまり、実際の実装とはほとんど関係のない、もっともらしく自信に満ちた設計理論です。この現象を研究するために、デザイン シアターを測定するためのベンチマークと 3 つの指標を紹介します。このベンチマークには、構造、スタイル、機能の設計要件にわたる 24 の UI 生成タスクが含まれています。このベンチマークを使用して、5 つの生成 UI ツールによって作成された 120 のインターフェイスを評価します。平均すると、ユーザー向けの設計理論の 25\% 以上が生成されたインターフェイスに実装されておらず、機能要件の実装失敗は 34\% に増加します。ツールはプロンプトに埋め込まれた UX 原則の約半分 (平均 = 0.54) を認識し、ツール 5 つのうち 4 つは 6\% 以下の機能原則を実装しています。また、ツール間でのインターフェイスの類似性を測定し、色の選択のバリエーションが大きくなり、外観とレイアウト構成が収束していることを確認します。全体として、私たちは次のことに貢献します。 1) デザイン シアターのコンセプト。 2) 生成 UI ツールの記述された推論がその実装に反映されているかどうかを評価するための指標を備えたベンチマーク。 3) およびこれらのツールの体系的な評価から得られた結果。これらの発見が生成 UI ツールの設計と評価に何を意味するかについて説明します。

原文 (English)

Design Theater: A Benchmark for Generative UI

Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate user-facing design rationales that explain their layout, accessibility, and design choices. However, it remains unclear whether these stated rationales are actually reflected in the interfaces they produce. We call this disconnect ``Design Theater'': plausible and confident design rationales that have little relationship to the actual implementation. To study this phenomenon, we introduce a benchmark and three metrics for measuring Design Theater. The benchmark includes 24 UI generation tasks spanning structural, styling, and functional design requirements. Using this benchmark, we evaluate 120 interfaces created by five generative UI tools. On average, over 25\% of user-facing design rationales are not implemented in the generated interface, and the implementation failure increases to 34\% for functional requirements. Tools recognize roughly half of the UX principles embedded in prompts (mean = 0.54), with four of five tools implementing 6\% or fewer functional principles. We also measure interface similarity across tools and find convergence in visual appearance and layout organization, with greater variation in color choices. Overall, we contribute: 1) the concept of Design Theater; 2) a benchmark with metrics for assessing whether the stated reasoning of generative UI tools is reflected in their implementations; 3) and findings from a systematic evaluation of these tools. We discuss what these findings mean for the design and evaluation of generative UI tools.

13:00 JSTエージェント

AI エージェントにネットワークについて説明するのではなく、ネットワークを翻訳させましょう

正式なモデルを使用すると、到達可能性の検証、停止の位置特定、または変更の影響範囲の予測が可能になります。しかし、モデルを手動で記述するには稀な専門知識が必要であり、ネットワークが頻繁に変更されるため最新の状態に保つのが難しいため、実稼働ネットワークには事実上これがありません。本質的に、ネットワーク モデリングはタイポグラフィーの実践であり、ネットワーク アーティファクト (構成、トポロジ、ルーティング状態など) を形式ロジックのルールに変換します。この種の翻訳は、今日の大規模言語モデル (LLM) が得意としているものです。自由形式の AI 推論とは異なり、そのような翻訳は正式に検証できます。モデリングがボトルネックでなくなると、大規模で複雑なネットワーク上の推論を AI に信頼することはもはや意味がありません。したがって、私たちの立場は、自律型 AI エージェントにエンドツーエンドで責任を負わせるという一般的な競争に対抗するものです。代わりに、AI を翻訳に限定し、信頼性の高い長期的な推論のためにソルバーに依存し、根本原因分析 (RCA) などの特定のタスクに特化できる一般的なネットワーク動作の再利用可能な正式なモデルを構築します。私たちは、ネットワーク自体の成果物からエミュレートされた実稼働規模の WAN のシンボリック モデルを構築および検証する TypoNet を構築します。私たちの予備評価では、TypoNet が 2 つの方法で役立つことがわかりました。 TypoNet は、それ自体で、運用上の質問 (到達可能性の検証や変更の影響分析など) に、LLM よりも速く、安価に、より確実に答えます。 AI エージェントのツールとして、TypoNet は低コストで障害の位置特定を強化します。この結果は、検証可能なネットワーク モデルを構築し、信頼性の高い長期的な推論のためにソルバーに依存する AI の正当性を主張します。

原文 (English)

Let AI Agents Translate Networks, Not Reason About Them

A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network has one, since writing a model by hand demands rare expertise and is hard to keep current as the network changes frequently. At its core, network modeling is a typographical exercise: it translates network artifacts (e.g., configurations, topology, and routing state) into rules in formal logic. Translation of this kind is what large language models (LLMs) nowadays do well. Unlike free-form AI reasoning, such translation can be formally verified. Once modeling is no longer the bottleneck, trusting AI to reason over large, complex networks no longer makes sense. Our position therefore cuts against the prevailing race to put autonomous AI agents in charge end-to-end. We instead confine AI to translation and rely on a solver for reliable long-horizon reasoning, building a reusable formal model of general network behavior that can then be specialized to specific tasks, e.g., root-cause analysis (RCA). We build TypoNet that constructs and validates a symbolic model of an emulated production-scale WAN from the network's own artifacts. Our preliminary evaluation shows TypoNet helps in two ways. On its own, TypoNet answers operational questions (e.g., reachability verification and change-impact analysis) faster, more cheaply, and more reliably than an LLM. As a tool for an AI agent, TypoNet boosts fault localization at lower cost. The result makes the case for AI that builds verifiable network models and relies on a solver for reliable long-horizon reasoning.

13:00 JST研究/論文

リクエストに必要な以上の共有は不要: 視点を意識した AI のための Federated Disclosure

最新の AI システムは、大規模な監視、極度の権力の集中、ユーザーの自主性の喪失などの社会的リスクをもたらし、サードパーティが大量のユーザー データを収集および管理するモデルに疑問を投げかけています。ユーザーは、厳格な来歴、解釈可能性、およびポリシー順守により、規制対象ドメイン全体でコンプライアンスを維持しながら、コンテキストを安全に所有、管理、開示する主権システムを必要とします。視点を認識した AI は、ユーザーの集約された個人データを \emph{Chronicle} と呼ばれる構造化アイデンティティ モデルに変換することでこれにアプローチします。\emph{Chronicle} は、ユーザーを表し、ユーザーとともに成長する時間的ナレッジ グラフです。クロニクルは、フェデレーテッド ネットワーク全体でのコンテキストの安全な開示をサポートします。 Chronicle 所有者は、第三者のエージェントが誰のデータも一元化することなく参照できる、クエリ可能な承認されたビューを公開する場合があります。このペーパーでは、ドメインの境界を越えた必要最小限の開示の問題を検討します。要求者のエージェントがクロニクルにクエリを送信するとき、システムは、要求者の関係、明示された目的、および特定のタスクが必要とするもののみを公開するようにその応答をどのように制限できるでしょうか?私たちは \textbf{Provenance Preserving Chronicles} (PPC) を提案します。これは、各所有者の年代記を、\emph{要求が必要とする以上共有しない} という 1 つのルールによって管理されるコンパクトな \emph{認可された証拠のサブグラフ} にコンパイルする連合プロトコルです。保有者は地域主権を維持します。アクセス コントローラーは、ドメイン エキスパート オントロジーに対して関係を認識したビューを投影します。そして 2 フェーズ フローでは、まず出所にリンクされたテキストが返され、所有者の明示的な承認後にのみ生のアーティファクトがリリースされます。私たちは問題を枠組み化し、ブロックチェーン、P2P、ホルダーソブリン設計のギャップをマッピングし、コア構造を定義し、明示的な脅威モデルを使用してプロトコルをスケッチします。

原文 (English)

Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI

Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy---calling into question a model where third-parties collect and control massive amounts of user data. Users require a sovereign system to securely own, govern, and disclose their context while remaining compliant across regulated domains with strict provenance, interpretability, and policy adherence. Perspective-aware AI approaches this by transforming a user's aggregated personal data into a structured identity model called a \emph{Chronicle}: a temporal knowledge graph that represents and grows with the user. Chronicles support the secure disclosure of context across federated networks. A Chronicle holder may expose a queryable, authorized view that a third-party agent may consult without centralizing anyone's data. This paper explores the problem of minimum-necessary disclosure across domain boundaries: when a requester's agent queries a Chronicle, how can the system constrain its response to release only what the requester's relationship, stated purpose, and specific task require? We propose \textbf{Provenance Preserving Chronicles} (PPC), a federated protocol that compiles each holder's Chronicle into a compact \emph{authorized evidence subgraph} governed by one rule: \emph{share no more than the request requires}. Holders keep local sovereignty; an access controller projects relationship-aware views over domain-expert ontologies; and a two-phase flow returns provenance-linked text first, releasing raw artifacts only after explicit holder approval. We frame the problem, map gaps in blockchain, P2P, and holder-sovereign designs, define the core constructs, and sketch the protocol with an explicit threat model.

13:00 JSTLLM/生成AIエージェント

ConsistencyGate: 自己一貫性アドミッション コントロールによる LLM エージェントのメモリ汚染の防止

多くのターンにわたって動作する LLM エージェントは、外部メモリ ストアにファクトを蓄積し、下流の推論の前提としてそれらを再利用します。したがって、あるステップで書かれた幻覚的な事実は、その後のすべてのステップで誤った前提として残り、これをメモリ汚染と呼びます。既存のメモリ管理では、取得と容量には対応していますが、書き込み時の正確さには対応していません。この入場問題は、事業や最近の基準に基づいた基準や、長い軌道にわたる制御されていない汚染物質によっては解決できません。我々は、コンテキスト c から抽出された候補ファクト m をコミットする前に、LLM に K 回ソフト サポート スコアを問い合わせ、平均がしきい値を超えた場合にのみ m を許可する書き込み時許可ゲートである ConsistencyGate を提案します。このメカニズムはモデルに依存せず、微調整を必要とせず、遅延の影響を受けやすい展開では対数確率バリアントの単一の転送パスに削減されます。自然データへの影響を測定するために、LoCoMo と MSC からの長期会話に制御された単一詳細の破損を植え付けることで 2 つの実際の会話ベンチマーク (LoCoMo-Contam と MSC-Contam) を構築し、オラクルに近い上限を分離する構造化合成コーパス (MemContam) でそれらを補完します。 ConsistencyGate は、4 つの LLM バックボーンにわたって、すべて書き込みベースラインと比較してすべてのベンチマークで汚染を削減し、ソース コンテキストで暗黙的にのみ記載されている事実にコストを集中させます。 3 つのベンチマークすべてをゲート実装とともにリリースします。

原文 (English)

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subsequent step, a failure mode we call memory contamination. Existing memory management addresses retrieval and capacity but not write-time correctness; this admission problem cannot be solved by utility- or recency-based criteria, and uncontrolled contamination compounds across long trajectories. We propose ConsistencyGate, a write-time admission gate that, before committing a candidate fact m extracted from context c, queries the LLM K times for a soft support score and admits m only when the average exceeds a threshold. The mechanism is model-agnostic, requires no fine-tuning, and reduces to a single forward pass in a log-probability variant for latency-sensitive deployments. To measure the effect on natural data, we construct two real-conversation benchmarks (LoCoMo-Contam and MSC-Contam) by planting controlled single-detail corruptions in long-term conversations from LoCoMo and MSC, and complement them with a structured synthetic corpus (MemContam) that isolates a near-oracle upper bound. Across four LLM backbones, ConsistencyGate reduces contamination on every benchmark relative to a write-everything baseline, with the cost concentrated on facts that are stated only implicitly in the source context. We release all three benchmarks together with the gate implementation.

13:00 JSTLLM/生成AI

Reason Popper-ly: 帰納的論理プログラミングによるコンテキスト内推論のパッチ適用

思考連鎖 (CoT) プロンプトにより、大規模言語モデル (LLM) は複数ステップの推論タスクに取り組むことができますが、生成された中間ステップが論理的に健全であることは保証されていません。我々は、帰納的論理プログラミング (ILP) を使用して推論トレースから関係構成ルールを学習し、ステップレベルの修正のためのオンライン検証器として展開する神経象徴的フレームワークである Reason Popper-ly を紹介します。 LLM で生成されたトレースが与えられると、このメソッドは、推定された各ステップを学習されたルール テーブルと照合してチェックし、違反タイプを診断し、記号的に導出された修復で間違ったステップを書き換え、残りのサフィックスを再生成して、モデルが検証されたトレースに基づいて最終的な答えを生成できるようにします。 2 ~ 10 ホップの推論チェーンにわたって 5 つの言語モデルを使用して、マルチホップ血縁推論ベンチマークである CLUTRR で評価します。すべてのモデルにわたって、Reason Popper-ly は標準 CoT よりも端末精度を一貫して向上させており、最も長いチェーンの小型モデルでは最大 48 パーセント ポイント、フロンティア モデルでは 15 ポイント向上しています。完全に外生的なシンボリック パイプラインと比較して、私たちの方法は、検証可能な推論の失敗のみを修正しながら、モデルの成功したグラウンディングを維持することで、より困難なインスタンスでより優れたパフォーマンスを発揮します。さらに、ステップレベルの ILP 検証により、最終的な回答の精度を超えた診断上の洞察を提供する、きめの細かいエラー分類が得られます。

原文 (English)

Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming

Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound. We present Reason Popper-ly, a neurosymbolic framework that uses inductive logic programming (ILP) to learn relation composition rules from reasoning traces and deploys them as an online verifier for step-level correction. Given an LLM-generated trace, the method checks each inferred step against the learned rule table, diagnoses the violation type, rewrites incorrect steps with symbolically derived repairs, and regenerates the remaining suffix so that the model can produce its final answer conditioned on a verified trace. We evaluate on CLUTRR, a multi-hop kinship reasoning benchmark, using five language models over reasoning chains of 2 to 10 hops. Across all models, Reason Popper-ly consistently improves terminal accuracy over standard CoT, with gains of up to 48 percentage points for small models and 15 points for frontier models on the longest chains. Compared with a fully exogenous symbolic pipeline, our method performs better on harder instances by preserving the model's successful grounding while correcting only verifiable reasoning failures. In addition, step-level ILP verification yields a fine-grained error taxonomy that provides diagnostic insight beyond final-answer accuracy.

13:00 JSTエージェントロボティクス

Stress-testing large language model agents in a robotic chemistry laboratory

AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to…

13:00 JSTLLM/生成AI

SymStep: Symbolic Step Verification for Logical Reasoning

Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently…

13:00 JST研究/論文

Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approa…

13:00 JSTLLM/生成AI

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rel…

13:00 JSTエージェントビジネス/資金調達研究/論文

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable delivera…

13:00 JSTLLM/生成AIエージェント

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and inter…

13:00 JST画像/動画生成

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although mos…

13:00 JST研究/論文

An Ontology for Machine Learning Interatomic Potentials

Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density fun…

13:00 JST研究/論文

Characterisation of Density-based FM generation methods in the context of Information Fusion

Fuzzy Integral (FI) based aggregation provides a powerful mechanism for nuanced aggregation, for example, in ensemble approaches or decisio…

13:00 JST研究/論文

CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics

Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central quest…

13:00 JSTエージェント

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-i…

13:00 JSTLLM/生成AI

TopoFE: topology-aware LLM-guided Automated Feature Engineering

Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective…

13:00 JST研究/論文ClaudeGPT / ChatGPTGeminiDeepSeek

RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning

Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endu…

13:00 JST研究/論文

Ordered Network Analysis of Epistemic Emotions during Collaborative Problem Solving

Investigating how affective states such as confusion and frustration persist and transition during co-situated collaborative problem solvin…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in re…

13:00 JSTLLM/生成AIビジネス/資金調達AnthropicGPT / ChatGPT

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under neste…

13:00 JST研究/論文

Key-Interval A*: Accelerating Grid Pathfinding via Structural Abstraction

Existing exact methods for 4-connected grid pathfinding reduce online search, but often either retain fine-grained search states or require…

13:00 JSTLLM/生成AI

Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning

Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly…

13:00 JSTビジネス/資金調達

NeurGO: Learning to Generate Elite Candidates for Meta-Black-Box Expensive Optimization

Expensive black-box optimization is ubiquitous in science and engineering, where function evaluations are costly and the evaluation budget…

13:00 JSTエージェント

Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels

As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing w…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

Do LLMs Know Their Vulnerable Scenarios?

Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can by…

13:00 JSTエージェント

Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis

Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy…

13:00 JSTエージェント研究/論文

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standa…

13:00 JST研究/論文

Verification-Notebook Learning for Source-Aware Multimodal Misinformation Detection

Multimodal misinformation verification is challenging because misleading signals may come from different parts of a post and require differ…

13:00 JSTエージェント

Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents

Long-lived AI agents increasingly evolve after deployment by retaining experience, acquiring skills and tools, revising workflows, delegati…

13:00 JSTエージェント

Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-m…

13:00 JSTLLM/生成AI

SpecAHD: Localize to Specialize for Automated Heuristic Design in Large-Scale Routing Problems

LLM-based automated heuristic design (AHD) typically scores executable programs on complete instances or within fixed solver components. In…

13:00 JSTLLM/生成AIエージェント

Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems

Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use. Recent systems increasingly organize these age…

13:00 JSTエージェント

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Evic…

13:00 JSTビジネス/資金調達

Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing

Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible cre…

13:00 JST研究/論文

Offline-Online Curriculum RL for Multimodal Reasoning

Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correc…

13:00 JSTLLM/生成AIエージェント研究/論文

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hi…

13:00 JSTLLM/生成AI

Training Language Models to Cooperate with Inference-Time Controllers

Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to o…

13:00 JSTLLM/生成AI

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enab…

13:00 JSTエージェント

ACM: Agentic Context Management for Long Horizon Tasks

Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existi…

13:00 JST研究/論文

Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a r…

13:00 JST研究/論文

Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search

Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms…

13:00 JSTエージェントロボティクス

Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery

Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We for…

13:00 JST研究/論文Qwen

GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens…

13:00 JSTLLM/生成AIハードウェア/半導体

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, th…

13:00 JSTLLM/生成AIエージェント

MemTX: Transactional Belief Commit for Stateful Agent Memory

LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a to…

13:00 JSTエージェント

From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps

Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agen…

13:00 JST研究/論文

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusin…

13:00 JSTエージェントGPT / ChatGPT

EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups…

13:00 JST研究/論文

Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time…

13:00 JSTエージェントハードウェア/半導体研究/論文

Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does…

13:00 JST研究/論文

Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation

The ever-growing adoption of Artificial Intelligence (AI) creates the need to deploy Deep Neural Networks in a variety of computational env…

13:00 JSTロボティクス

Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interface

Brain-Machine Interfaces (BMIs) provide a direct communication pathway between the brain and external devices, enabling humans to control a…

13:00 JSTロボティクス

A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces

Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitation, human performance augmen…

13:00 JSTビジネス/資金調達研究/論文OpenAIGPT / ChatGPT

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper h…

13:00 JST研究/論文

Quantum-Inspired Evolutionary Neighborhood Search for Arrival-Departure Track Utilization Adjustment under Short-Term Disturbances

Short-term disturbances at major passenger railway stations alter train arrival and departure times as well as the release sequence of stat…

13:00 JSTエージェントビジネス/資金調達

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer di…

13:00 JST研究/論文GPT / ChatGPTGemini

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a…

13:00 JST研究/論文

MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers

We present MiSS, a black-box, query-based framework for explaining 3D point cloud classifiers through perturbation-relative sufficiency rea…

13:00 JST研究/論文

Towards High-Level Semantic Intelligence

Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, th…

13:00 JSTLLM/生成AIエージェント

MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This ret…

13:00 JSTエージェント

Scaling GUI Agents with Visual State Transitions

We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unifi…

13:00 JSTエージェント

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retr…

13:00 JSTハードウェア/半導体

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits rea…

13:00 JSTエージェント

Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness

Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete…

13:00 JSTエージェント

Falsifiable Commitment Planning for Self-Correcting Web Agents

Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, re…

13:00 JST研究/論文

Myopia Prevention and Control 3.0: Artificial Intelligence--Driven Risk Stratification, Proactive Monitoring, and Personalized Intervention

The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to tran…

13:00 JST研究/論文

Integrating Factual and Normative Industrial Knowledge via Constraint-Aware Graph Attention for Process Plan Recommendation

Integrating heterogeneous industrial knowledge, including factual relations and decision constraints, remains a core challenge in industria…

13:00 JST研究/論文

Epistemic Norms for AI Safety and Alignment Research

Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and al…

13:00 JSTLLM/生成AI

Generative Artificial Intelligence (GenAI) to convert images of queuing networks into verifiable simulation models: an open-weight LLM workflow approach

Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable…

13:00 JSTエージェント

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet op…

13:00 JSTエージェント

Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is…

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPTGeminiGrok

Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families

Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, persevera…

13:00 JSTLLM/生成AIGPT / ChatGPTLlama

Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age

Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal,…

13:00 JSTLLM/生成AI

Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updat…

13:00 JSTエージェント

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect f…

13:00 JST研究/論文

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capabi…

13:00 JST研究/論文

Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration

Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge b…

13:00 JSTLLM/生成AI

Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis

Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does…

13:00 JSTLLM/生成AIGPT / ChatGPTMistral AI

LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph

Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate…

13:00 JST研究/論文GemmaLlama

Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models

Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control…

13:00 JST研究/論文

TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs

Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outp…

13:00 JST画像/動画生成

DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensiona…

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPT

LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-wo…

13:00 JSTLLM/生成AIエージェント

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather eve…

13:00 JST研究/論文

Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions

The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared ac…

13:00 JSTエージェント研究/論文

Efficiency Matters in Autonomous Research

AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, i…

13:00 JSTLLM/生成AIビジネス/資金調達

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…

13:00 JST研究/論文NVIDIA

Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment…

13:00 JST画像/動画生成

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rat…

13:00 JSTLLM/生成AIビジネス/資金調達

Creative Integration: A Decidable Criterion of Creativity

"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…

13:00 JSTLLM/生成AIGPT / ChatGPTDeepSeek

Evaluating Large Language Models for Symbolic Security Protocol Analysis

Security protocol verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether Large Language Models (LLMs)…

13:00 JST研究/論文

Comparing Optimization Models for Radiotherapy Scheduling

The Radiotherapy Scheduling Problem (RTSP) involves determining an optimal schedule for patients undergoing radiation treatments, a task th…

13:00 JSTLLM/生成AIエージェントLlama

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt in…

13:00 JSTLLM/生成AI研究/論文

Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In thi…

13:00 JST研究/論文

A Formal Kinetic Theory for Zeroth-Order Newton Dynamics:Stein-Corrected Hessian Estimation and Curvature--Variance Trade-offs

Zeroth-order Newton-type methods are useful when gradients and Hessians are unavailable, but they behave quite differently from first-order…

13:00 JSTLLM/生成AI

A didactical-driven teacher assistant for a dimensional modeling course

Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagog…

13:00 JST研究/論文

Quotient Tree Arithmetic: Deferred-Division Computation with Bounded Symbolic Depth and Cross-Subtree Cancellation

We introduce Quotient Tree Arithmetic (QTA), a computational substrate in which values are represented as deferred quotient pairs (N, D) wh…

13:00 JSTLLM/生成AI研究/論文

Revitalizing Public Urban Places through Cultural and Political Memory: A Technological Approach with LLMs and Augmented Reality

This paper explores the intersection of memory, place, and identity, examining how new technologies, particularly Apple Vision Pro, can ill…

13:00 JST研究/論文

Masked Autoencoders Learn Perception-Relevant Representations from Resting State Neural Data

Clinical neuroprosthetics face a data bottleneck: labeled perception trials are scarce while hours of spontaneous neural activity are large…

13:00 JSTLLM/生成AI

Learning When to Reason for Text-to-SQL via SFT and DPO

Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on compl…

13:00 JST研究/論文

AI-Assisted Causal Inference and Mediation Analyses of Environmental and Psychosocial Determinants of Subjective Cognitive Difficulties in the All of Us Research Program

Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect b…

13:00 JST規制/政策研究/論文Meta

AutoCluster, AutoTopicModeling, AutoTrendAnalysis: A Complete AutoML Pipeline for Predicting Emerging Trends

Predicting emerging trends is vital for businesses, researchers, and policymakers; yet traditional approaches often lack scalability and ad…

13:00 JSTLLM/生成AIGemmaQwen

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of wheth…

13:00 JST研究/論文

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, shari…

13:00 JST研究/論文

Towards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution

Artificial intelligence systems increasingly mediate how science is found and credited. We asked whether missing metadata prevents AI syste…

13:00 JST画像/動画生成

MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion

High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget…

13:00 JST画像/動画生成

Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model

Nuclear fusion has made significant progress in recent years and is expected to become one of the most important pathways to addressing glo…

13:00 JSTLLM/生成AI画像/動画生成

RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus

The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Me…

13:00 JSTLLM/生成AIエージェント

CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-mak…

13:00 JST画像/動画生成

scMIR: a vision-language foundation model for single-cell light microscopy image representation

Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heter…

13:00 JST画像/動画生成エージェントNVIDIA

Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems

Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded auto…

13:00 JST画像/動画生成

DAMamba-UNet3D: A Parameter-Efficient Mamba State Space U-Net with Dynamic Adaptive Scan for 3D Medical Image Segmentation

We propose parameter-efficient SSM-based U-Net architectures for 3D medical image segmentation. Convolutional U-Nets afford O(n) local mixi…

13:00 JST画像/動画生成

An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia

Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. Fo…

13:00 JST画像/動画生成

A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection

Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible p…

13:00 JSTLLM/生成AIエージェント

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by compar…

13:00 JST画像/動画生成

Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification

Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are…

13:00 JST画像/動画生成

Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features

We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcem…

13:00 JST画像/動画生成

QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation

Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion. Centraliz…

13:00 JST画像/動画生成GPT / ChatGPT

AI-generated Images Challenge Visual Trust in High-risk Scenarios

Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public saf…

13:00 JST画像/動画生成

Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge

Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed…

13:00 JST研究/論文

Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence

Modern neural networks primarily adapt through parameter modification within predefined computational structures. While recent methods intr…

13:00 JST画像/動画生成

Post-Operative Glioma Segmentation via Loss Stabilization, Normalization and Subspace Attention

Tracking residual tumor after surgery is essential for catching recurrence early, but automating post-operative glioma segmentation remains…

13:00 JST画像/動画生成

Real-time Reconstruction of Human Visual Perception from fMRI

Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical adva…

13:00 JSTLLM/生成AI

Hierarchical Grading in Large Language Models

We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grad…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks

The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying comm…

13:00 JSTLLM/生成AIビジネス/資金調達Anthropic

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instr…

13:00 JSTLLM/生成AI

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning

The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existin…

13:00 JST画像/動画生成

Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a se…

13:00 JST研究/論文

LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning

Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along am…

13:00 JST研究/論文

Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control

Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control. In this field, deep learning approaches…

13:00 JST研究/論文

What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation

High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temp…

13:00 JST研究/論文

Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs

In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time…

13:00 JST研究/論文

FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow

AC optimal power flow determines the minimum-cost generation dispatch under nonlinear power balance constraints and is solved thousands of…

13:00 JSTLLM/生成AI

Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training

Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arisi…

13:00 JSTLLM/生成AI

Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis

Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checke…

13:00 JST研究/論文

LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers

Accurate geological characterization of subsurface reservoirs from well log data is essential to support projects such as carbon capture an…

13:00 JST研究/論文

OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence

We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Servic…

13:00 JST画像/動画生成

Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution

The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies. A criti…

13:00 JST研究/論文

From Hybrid Mechanistic--Data-Driven Modeling Toward Neuro-Symbolic AI: What, Why, and How

Hybrid mechanistic/data-driven models, which combine first-principles with learned components, are increasingly used in process engineering…

13:00 JST画像/動画生成エージェント研究/論文

Agentic Autoresearch for CT Reconstruction

Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a…

13:00 JSTLLM/生成AI

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address priva…

13:00 JSTLLM/生成AI研究/論文Llama

Language-Routed RAG and Direct Option Scoring for Multilingual Financial QA: DS@GT at FinMMEval

We present DS@GT's submission to FinMMEval 2026 Task 1, a multilingual financial exam question answering benchmark spanning English, Spanis…

13:00 JST画像/動画生成

Robustifying pathology foundation models via fine-tuning

Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, u…

13:00 JSTLLM/生成AI画像/動画生成NVIDIA

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Exi…

13:00 JST研究/論文

AI-interpreted Optical Scattering for Robust and Focal Depth-Aware Imaging

Optical scattering has conventionally been regarded as an impediment in imaging research due to the degradation of image quality during rec…

13:00 JST研究/論文

Multi-primitive in-memory computing for Monte Carlo tree search

Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limi…

13:00 JST研究/論文

Spatial Prediction of Soil Microplastics and Organic Matter Using Graph Attention Networks

Accurate estimation of soil microplastics and organic matter is essential to assess ecosystem health and support sustainable land use. This…

13:00 JSTLLM/生成AI

Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)

Recent advances in large language models (LLMs) have driven growing interest in using LLMs to automate test generation. Prior work commonly…

13:00 JSTLLM/生成AI

Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests

While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of gener…

13:00 JST画像/動画生成

Controlling Embedding Spaces with Text-Conditioned Transformations

Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-sho…

13:00 JSTLLM/生成AIハードウェア/半導体Claude

Not All LLM Reasoning is Visible in the Chain-of-Thought

A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete fai…

13:00 JST研究/論文

Invariant Discovery for Networked Systems

Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generati…

13:00 JSTエージェント

Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence

The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 init…

13:00 JSTLLM/生成AIAnthropicClaudeOpenAI

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a sp…

13:00 JST画像/動画生成

HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quali…

13:00 JST研究/論文

Label-free Industrial Fault Detection via Adversarial Inverse Reinforcement Learning: A System for Run-to-Failure Prognostics

Machinery fault detection (MFD) remains heavily reliant on supervised learning, which struggles with the scarcity of fault labels in real-w…

13:00 JST研究/論文

An Explicit Counterexample to Stanley's Rankwise Lower-Bound Conjecture for Differential Posets

In Problem~6 of his 1988 paper on differential posets, Stanley asked for the least possible cardinality of a fixed rank of an $r$-different…

13:00 JSTLLM/生成AIQwen

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the openin…

13:00 JSTエージェントロボティクスハードウェア/半導体

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged…

13:00 JSTエージェントロボティクス

WCM: World-Cognition Model for Generalizable Human-Robot Interaction

Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tas…

13:00 JST研究/論文

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop

Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified. We st…

13:00 JST研究/論文

Exact values and exact upper bounds for families of integers with arithmetic progression intersections (Erd\H{o}s Problem #272)

Let $t(N)$ be the largest $t$ for which there exist distinct sets $A_1,\dots,A_t \subseteq \{1,\dots,N\}$ such that $A_i \cap A_j$ is a non…

13:00 JSTエージェント研究/論文

VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locatin…

13:00 JST研究/論文

All in One: Generative Modeling as Mean-Field Game Design

Mean-field games (MFGs) offer a unifying lens on continuous-time generative modeling: a cost tuple recovering twelve prominent models---Con…

13:00 JSTエージェント

Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View

Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a p…

13:00 JST研究/論文LlamaMistral AI

MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models

Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but ex…

13:00 JSTLLM/生成AIDeepSeek

Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achiev…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts…

13:00 JSTLLM/生成AIハードウェア/半導体

Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models

Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of…

13:00 JSTLLM/生成AI

Traceable LLM Reasoning for Fake-Order Fraud Detection

Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approache…

13:00 JST研究/論文

Scoping Review of AI, Metrology, and ESG in the Semiconductor Sector: Implications for Safe and Sustainable by Design (SSbD)

The semiconductor sector faces a dual transition: scaling manufacturing execution through Artificial Intelligence (AI) while satisfying str…

13:00 JSTLLM/生成AI

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios

Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains un…

13:00 JST研究/論文

KAYROS: An Anytime and Exact Open-Source Solver for Duration-Minimization Time-Dependent Vehicle Routing. A Technical Report and a Case Study in Human-AI Engineering

KAYROS is an open-source solver for duration-minimization time-dependent vehicle routing problems, with or without time windows (TDVRPTW, T…

13:00 JST研究/論文Google

A scalable online machine learning approach for Stock Recommendation

Stock recommendation systems face the dual challenge of adapting to rapidly changing market conditions while maintaining low-latency predic…

13:00 JSTLLM/生成AI

From Vibe to Code -- and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI

Generative AI design tools make natural-language prompts a starting point for design, placing new articulation demands on designers. Rather…

13:00 JSTエージェント

False Prophets: On the Security of World Models in Agentic Systems

Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable exe…

13:00 JSTLLM/生成AIハードウェア/半導体

In-Context Learning as Implicit Policy Gradient

Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their…

13:00 JSTLLM/生成AI

Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining

Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent.…

13:00 JST画像/動画生成

Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design

AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion in…

13:00 JST画像/動画生成

BoneAgeTW2: Automated Skeletal Maturation Assessment via the Tanner-Whitehouse 2 Method, Deep Learning, and Clinical Report Generation with Distribution Curves

We present BoneAgeTW2, the first fully open-source system to automate the complete Tanner-Whitehouse 2 (TW2) clinical protocol for skeletal…

13:00 JST研究/論文

FedSLIM: Privacy-Preserving Federated MDL-Based Descriptive Pattern Mining Across Data Silos

Federated learning has achieved considerable success for predictive modelling, yet federated descriptive analytics remains largely unexplor…

13:00 JST研究/論文

Context-Aware Concept Distillation for Trustworthy Flood Prediction

Effective flood risk management relies on accurate forecasting, yet the "black box" nature of stateof-the-art Deep Learning models creates…

13:00 JSTハードウェア/半導体NVIDIA

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference

Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote…

13:00 JST画像/動画生成

What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embed…

13:00 JSTLLM/生成AI

Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition

Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference ra…

13:00 JST研究/論文

FILLER: Feature Imputation via Latent Location Exploration and Retrieval

In real-world machine learning applications, incomplete observations create a fundamental challenge. Researchers have come up with several…

13:00 JST研究/論文

Online Fair Division with Budget Constraints

We study an online variant of discrete fair division under generalized assignment budget constraints. Goods arrive one at a time and must b…

13:00 JST研究/論文

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for t…

13:00 JST画像/動画生成

Patient-Agnostic Synthetic Pretraining for Efficient Patient-Specific Intraoperative 2D/3D Registration

Intraoperative 2D/3D registration aligns preoperative CT volumes with intraoperative X-ray or fluoroscopic images and is essential for imag…

13:00 JSTエージェント

On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems

Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and…

13:00 JSTエージェントビジネス/資金調達

Fair Division with Strictly Increasing Valuations: A Tight Threshold for Two-Agent EF1 and PO

We study whether strictly positive marginal values restore the compatibility of envy-freeness up to one good (EF1) and Pareto optimality (P…

13:00 JSTLLM/生成AI画像/動画生成

Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing

Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful a…

13:00 JSTLLM/生成AI

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They…

13:00 JST画像/動画生成ロボティクス

Semantic Semi-Incremental Data-Association-Free Object SLAM

Data association between landmark measurements and landmark variables has long been a central challenge in SLAM, as estimation accuracy dep…

13:00 JST研究/論文

Directional Influence Function: Estimating Training Data Influence in Constrained Learning

As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety…

13:00 JST研究/論文

Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines

Continuous cuffless blood pressure (BP) monitoring is essential for connected health systems and wearable devices, enabling early detection…

13:00 JST研究/論文

TLA$^{+}$-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation

Large language models increasingly write TLA$^{+}$ formal specifications from natural-language descriptions, but progress is hard to measur…

13:00 JST研究/論文

A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models

Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader c…

13:00 JSTLLM/生成AI研究/論文Gemini

Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by l…

13:00 JSTLLM/生成AI

Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh

A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples…

13:00 JST研究/論文

Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds

Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determi…

13:00 JST研究/論文

When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design

Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing…

13:00 JST研究/論文

ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour

Fully homomorphic encryption (FHE) provides strong cryptographic guarantees for private inference, but deploying transformer models under F…

13:00 JST画像/動画生成

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information d…

13:00 JST研究/論文

Formalizing Flag Algebras in Lean

Razborov's flag algebra method is a powerful tool for proving asymptotic inequalities in extremal graph theory, often reducing the task to…

13:00 JST研究/論文

Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments

Internet of Things (IoT) applications generate vast amounts of Correlated Time Series (CTS) data that often contain missing values and requ…

13:00 JST研究/論文

Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework

Choosing the right text embedding model is one of the most consequential -- and most frequently under-examined -- decisions in building a r…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPT

Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning

Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve hu…

13:00 JSTLLM/生成AIビジネス/資金調達

Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static…

Auditing Alignment Controllability in LLMs via Political Axes

Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely mat…

13:00 JSTLLM/生成AIエージェントロボティクス

Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric

Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in…

13:00 JST研究/論文

Neonatal Hypoxic-ischaemic Encephalopathy Classification from the EEG and HRV Signals Using a Conformer based Masked Autoencoder

In this paper, we propose the MAEConformer, a novel self-supervised learning framework that combines the Conformer architecture with the Ma…

13:00 JST研究/論文

GTIN: A Unified Framework for Joint Event and Time Prediction in Temporal Graphs

Temporal graphs are increasingly used to model dynamic systems in diverse domains such as social networks, financial networks, and traffic…

13:00 JST研究/論文

An Unofficial FastLAS Tutorial: A Programmer's Guide

FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of ex…

13:00 JST画像/動画生成

D3O: Dynamic Distribution Distillation for Ordinal Regression

Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are o…

13:00 JSTロボティクス

Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models

Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimu…

13:00 JST研究/論文QwenDeepSeek

DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifie…

13:00 JSTLLM/生成AIエージェント

Where Is the Cost of Third-Party API Routers in Agentic Software Development?

Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows…

13:00 JST研究/論文

Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms

Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one…

13:00 JST研究/論文

Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science

Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream…

13:00 JST研究/論文

Extending Desbordante with Probabilistic Functional Dependency Discovery Support

Data profiling aims to extract complex patterns from data for further analysis and use that data in domains such as data cleaning, data ded…

13:00 JSTLLM/生成AI

CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation

Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse end…

13:00 JSTエージェント

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the…

13:00 JSTLLM/生成AI

An empirical investigation into the properties of standard word embeddings

The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processin…

13:00 JSTLLM/生成AIエージェント

The Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting

Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate se…

13:00 JST研究/論文

An Exact Counterexample to Carlson's Associated-Prime Depth Conjecture from a Group of Order 128

In Question~3.1 of his 1995 paper on depth and transfer, Carlson asked whether the depth of a finite-group cohomology ring is always realiz…

13:00 JST研究/論文

AI Strategy: How to Choose What AI Product to Implement

Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve…

13:00 JST研究/論文

Escaping the Euclidean Void: Manifold-Informed Flow Matching for Sequential Recommendation

Conventional recommenders capture users' preferences by optimizing observed user-item relations, whereas continuous generative recommendati…

13:00 JSTLLM/生成AI

WISERouter: LLM Routing with Workload Budget Constraint

Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is pr…

13:00 JST研究/論文

Outcome-Fair Restless Multi-Armed Bandits for Stochastic Deadline Scheduling

We study a restless multi-armed bandit (RMAB) problem for a stochastic deadline scheduling application. RMAB problems are solved using the…

13:00 JST研究/論文

Scale Weight Decay and Train Better

The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a c…

13:00 JSTロボティクス

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box p…

13:00 JST研究/論文

Maximum Satisfiability of Simple Temporal Problems

The Simple Temporal Problem (STP) is a core framework for quantitative temporal constraints. As STP data can be inconsistent, we study MAXS…

13:00 JST画像/動画生成

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellula…

13:00 JSTLLM/生成AI

How Context Attribution Handles What the Model Already Knows

Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works s…

13:00 JSTLLM/生成AIハードウェア/半導体

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take th…

13:00 JSTLLM/生成AI研究/論文

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. T…

13:00 JSTLLM/生成AI研究/論文

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls un…

13:00 JSTLLM/生成AI

Kalypso: Relational LLM Serving

Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructure…

13:00 JSTLLM/生成AIClaudeLlamaMistral AI

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base…

13:00 JSTエージェント

Limbomorphs

Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or both, from which lifelike patte…

13:00 JST研究/論文

A Coulomb Particle Model for Learning Kernel Attention in Transformers

Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature di…

13:00 JSTエージェント研究/論文

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that…

13:00 JST研究/論文

Physics-Informed Neural Networks for Predicting Nitrous Oxide Flux

Nitrous oxide (N$_2$O) is the dominant ozone-depleting substance emitted in the 21st century, and the third largest contributor to anthropo…

13:00 JSTLLM/生成AI

Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature

X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published…

13:00 JST研究/論文

Visible to the Court: How AI Is (and Isn't) Litigated in U.S. Federal Court Opinions

In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation. With courts become a recurring foru…

13:00 JSTLLM/生成AIロボティクスGPT / ChatGPT

Embodied GPT-5.1: Evidence of a World Model?

This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mo…

13:00 JSTLLM/生成AIハードウェア/半導体GPT / ChatGPTGemini

Understanding Tone-Dependent Inference Cost in Large Language Models

We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiment…

13:00 JST画像/動画生成

DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection

Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on inde…

13:00 JSTLLM/生成AIエージェント

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment…

13:00 JSTLLM/生成AI

Understanding Machine Unlearning Through the Lens of Mode Connectivity

Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress,…

Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

Appending a two-word confirmation tag to a decision question -- "Is X the better choice?" versus "X is the better choice, right?" -- change…

13:00 JST画像/動画生成

Multimodal Semantic-Probabilistic Objectness for Open World Object Detection

Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and…

13:00 JSTエージェント

Moral Hazard in Multi-Agent Language Models

Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr\"om's team…

13:00 JST研究/論文

Adaptive Data Admission and Retention for Streaming Federated Learning

We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and…

13:00 JSTLLM/生成AI

SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirement…

13:00 JSTエージェント

Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation

Cloud telemetry arrives at a scale that, paradoxically, makes intrusion understanding harder rather than easier. Attackers operate through…

13:00 JSTLLM/生成AI画像/動画生成

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifi…

13:00 JSTLLM/生成AIエージェントロボティクス

HELIOS: An LLM-Driven Autonomous Indirect Trajectory Optimization Agent

Low-thrust trajectory optimization is a core technology in deep-space mission design. Indirect methods based on Pontryagin's Minimum Princi…

13:00 JST研究/論文

Capacity-Aware Deep Learning for Generalizable Traffic Volume Estimation Across Links and Cities

Network-wide traffic volume estimation typically relies on propagating measurements from fixed sensors, making performance highly dependent…

13:00 JSTLLM/生成AI

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning

Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between trainin…

13:00 JST画像/動画生成

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-q…

13:00 JST研究/論文

Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging

Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challeng…

13:00 JST画像/動画生成研究/論文

LU-500: A Logo Benchmark for Concept Unlearning

Concept unlearning is increasingly used to limit the reproduction of protected or unsafe visual concepts in text-to-image models. Existing…

13:00 JSTロボティクス研究/論文

A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces

Previous research has shown that a human-like robot's acceptance heavily depends on the setting in which it operates and its ability to per…

13:00 JST研究/論文

EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding

Brain-machine interfaces provide a link between neural activity and external devices, enabling restoration of motor function and advancing…

13:00 JST研究/論文

Monitoring Post-Disaster Urban Recovery Using High-Resolution SAR Time Series and Unsupervised Learning: Evidence from the 2023 T\"urkiye-Syria Earthquake

Monitoring post-disaster recovery is essential for understanding how urban systems rebuild and progressively return to functionality. Howev…

13:00 JSTロボティクスビジネス/資金調達

Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim

Social robots that rely on large language models for conversation are unable to retain information across sessions. This absence of memory…

13:00 JSTLLM/生成AI研究/論文

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit…

13:00 JST研究/論文

Every Client Is an Environment: Federated De-confounding for Spatio-Temporal Forecasting

Federated learning has emerged as a promising paradigm for spatio-temporal forecasting (STF), enabling collaborative model training without…

13:00 JST画像/動画生成研究/論文

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks s…

13:00 JST研究/論文

ML-based Predictive Models for Power Consumption in Virtualised O-RANs

As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important…

13:00 JST研究/論文

Physics-Guided Generative AI for Property-Targeted 3D Porous Media Design

Inverse design of three-dimensional porous media is central to applications in filtration, catalysis, energy storage, fuel cells, thermal m…

13:00 JST研究/論文

A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health

Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between r…

13:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokeni…

13:00 JST研究/論文

Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy

We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive no…

13:00 JSTLLM/生成AIエージェント

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an em…

13:00 JSTLLM/生成AI

DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense

Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature. While recent autonom…

13:00 JSTLLM/生成AIハードウェア/半導体Llama

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagn…

13:00 JST画像/動画生成

MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path…

13:00 JST研究/論文

Regulating for AI Legitimacy

AI systems already govern. They rank speech and allocate attention, filter applicants and triage claims. The dominant frame for AI governan…

13:00 JSTハードウェア/半導体

The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphi…

13:00 JST研究/論文

Multivariate Time Series Forecasting with Adaptive Non-Local Observables

Multivariate time series forecasting (MTSF) predicts future values of multiple variables from historical data. While quantum neural network…

13:00 JST研究/論文DeepSeek

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active pe…

13:00 JSTLLM/生成AIMeta

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this g…

13:00 JST研究/論文

Evaluating RAG for French immigration law: a benchmark and baseline study

International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a pu…

13:00 JST画像/動画生成

ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image

Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We st…

13:00 JST研究/論文

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform…

13:00 JST画像/動画生成

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains…

13:00 JST研究/論文

Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls

Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to n…

13:00 JST研究/論文

EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardi…

13:00 JST研究/論文

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Atte…

13:00 JST研究/論文

BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient Leakage

Split learning enables collaborative model training by partitioning neural networks across clients and servers. However, improper split pla…

13:00 JST画像/動画生成

EgoPlay: Event-Triggered Video Editing for Egocentric Streams

We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion t…

13:00 JSTLLM/生成AI画像/動画生成

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violat…

13:00 JST画像/動画生成

CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same…

13:00 JSTLLM/生成AI

D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models

Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that a…

13:00 JSTLLM/生成AI

Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review

Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains ha…

13:00 JSTLLM/生成AIエージェント

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair

Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between…

13:00 JSTLLM/生成AIエージェント

Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents

Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors.…

13:00 JSTLLM/生成AI

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model…

13:00 JSTエージェントビジネス/資金調達

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical d…

13:00 JST画像/動画生成

Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification

Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real…

13:00 JST研究/論文

Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines

Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a highe…

13:00 JSTLLM/生成AIClaude

Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures

Trapped-ion quantum computers rely on shuttling compilers, which cast an input algorithm into a sequence of ion-qubit movements within a gi…

13:00 JSTLLM/生成AI

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches de…

13:00 JSTLLM/生成AIエージェント

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing mod…

13:00 JST画像/動画生成

KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician t…

13:00 JST画像/動画生成

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it…

13:00 JSTLLM/生成AI画像/動画生成GPT / ChatGPTGemini

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domai…

13:00 JST研究/論文

Procedural Content Generation via Generative Artificial Intelligence

The attempt to utilize machine learning in PCG has been made in the past. In this survey paper, we investigate how generative artificial in…

13:00 JST研究/論文

ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins

Digital twins are increasingly used to monitor and optimize industrial systems, yet many existing frameworks remain difficult to interpret,…

13:00 JST研究/論文

Robustness and Cybersecurity in the EU Artificial Intelligence Act

The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems. While prior work has sou…

13:00 JST画像/動画生成

Hybrid AI-Physical Modeling for Penetration Bias Correction in X-band InSAR DEMs: A Greenland Case Study

Digital elevation models derived from Interferometric Synthetic Aperture Radar (InSAR) data over glacial and snow-covered regions often exh…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

PD$^3$: A Project Duplication Detection Framework via Adapted Multi-Agent Debate

Project duplication detection is critical for project quality assessment because it helps avoid investment in repeated proposals. Existing…

13:00 JSTLLM/生成AI画像/動画生成GPT / ChatGPTGeminiQwen

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when…

13:00 JSTエージェント

Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models

Ontologies and knowledge graphs require continuous evolution to remain comprehensive and accurate, but manual curation is labor intensive.…

13:00 JSTLLM/生成AI

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provide…

13:00 JSTエージェント

TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning

Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent l…

13:00 JST研究/論文

Decentralized Causal Discovery using Judo Calculus

We describe a theory and implementation of an intuitionistic decentralized framework for causal discovery using judo calculus, which is for…

13:00 JSTLLM/生成AI画像/動画生成エージェント

OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows

Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments…

13:00 JST研究/論文

Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering

In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE…

13:00 JSTエージェント

Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic

Connected automated vehicles (CAVs) possess the ability to communicate and coordinate with one another, enabling cooperative platooning tha…

13:00 JST研究/論文

Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is…

13:00 JST研究/論文

The Lattice Representation Hypothesis of Large Language Models

We propose the Lattice Representation Hypothesis of large language models: a symbolic backbone that grounds conceptual hierarchies and logi…

13:00 JST研究/論文

PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms

AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accep…

13:00 JSTLLM/生成AIエージェント研究/論文Claude

LiteResearcher: Deep Research エージェント用のスケーラブルなエージェント RL トレーニング フレームワーク

強化学習 (RL) は、LLM ベースのエージェントの強力なトレーニング パラダイムとして登場しました。しかし、深い調査のためのエージェントティック RL のスケーリングは、依然として 2 つの複合的な課題によって制約されています。それは、手作りの合成データでは本物の現実世界の検索機能を引き出すことができないこと、もう 1 つは、RL トレーニング中の現実世界の検索依存性により、不安定性と法外なコストが生じ、エージェントティック RL のスケーラビリティが制限されることです。 LiteResearcher は、Agentic RL をスケーラブルにするトレーニング フレームワークです。現実世界の検索ダイナミクスを反映する軽量仮想世界を構築することで、小さな検索エージェントが大規模なオープンソース モデルや商用モデル (Tongyi DeepResearch や Claude-4.5 Sonnet など) を上回るパフォーマンスを発揮できるようにするトレーニング レシピを継続的に改善できます。具体的には、GAIA や Xbench などの一般的なベンチマークで、当社の LiteResearcher-4B は、それぞれ 71.3% と 78.0% というオープンソースの最先端の結果を達成し、スケーラブルな RL トレーニングがディープ リサーチ エージェントの重要な実現要因であることを実証しています。

原文 (English)

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elicit genuine real-world search capabilities, and real-world search dependency during RL training introduces instability and prohibitive cost, which limits the scalability of Agentic RL. LiteResearcher is a training framework that makes Agentic RL scalable: by constructing a lite virtual world that mirrors real-world search dynamics, we enable a continuously improving training recipe that empowers a tiny search agent to outperform large-scale open-source and commercial models (e.g., Tongyi DeepResearch and Claude-4.5 Sonnet). Specifically, on common benchmarks such as GAIA and Xbench, our LiteResearcher-4B achieves open-source state-of-the-art results of 71.3% and 78.0% respectively, demonstrating that scalable RL training is a key enabler for Deep Research Agents.

13:00 JST研究/論文

From Context to Skills: Can Language Models Learn from Context Skillfully?

Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for…

13:00 JSTエージェント

GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning

Geological interpretation infers subsurface properties and structures from indirect geophysical observations. Well-log classification provi…

13:00 JSTLLM/生成AI研究/論文

From Pixels to Prompts: Vision-Language Models

When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so lon…

13:00 JSTLLM/生成AI

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often c…

13:00 JSTLLM/生成AI

調査時に思考連鎖が機能するのはなぜですか?グローバルな派生ではなくローカルな共起

思考連鎖 (CoT) プロンプトは言語モデルの精度を確実に向上させますが、理論的根拠テキストのどの特性が向上を促進するのかはよくわかっていません。これまでの研究では、主に生成時の動作が研究されてきました。代わりに、調査時の質問をします。コンテキスト内で固定された理論的根拠が与えられた場合、そのテキストの何が答えを変えるのでしょうか?ゲインの 2 つの相補的なソースを特定します。まず、グローバルに単語をシャッフルした根拠でも、根拠なしのベースラインを大幅に上回っており、強力な語彙活性化効果が示されています。さらに重要なことは、構造化テキストによる追加の利益は、文レベルの論理的順序からではなく、より短い範囲のトークンの隣接性から生じているようです。わずか $n^\star{=}2$--$3$ トークンの連続ウィンドウを保存すると、完全な CoT パフォーマンスに向けて残りの利益のほとんどが回復します。サポート実験では、明示的な回答宣言または回答値のコピー、および完全な文法的実現が主な要因として除外されます。さらに一般化実験を行うと、定性的パターンが複数のモデル ファミリ、パラメーター スケール、データセットにわたって安定していることがわかります。これらの結果は、プローブ時間 CoT のローカル共起活性化 (LCA) 説明をサポートしており、観察された利益は、文レベルの論理導出ではなく、主に語彙活性化と短距離トークンの共起から生じているように見えます。

原文 (English)

What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation

Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear. We study this question from a probe-time perspective: holding CoT rationales fixed, we test which textual properties matter for the final prediction. Across multiple datasets and model configurations, we find that randomizing the order of rationale sentences has little effect on accuracy, suggesting that the global order of reasoning steps is not the main source of the probe-time benefit. Moreover, even when the words in a rationale are randomly reordered, performance remains well above the no-rationale baseline, indicating that the rationale's words remain useful even without their original order. Restoring only short-range word order further improves performance and brings it substantially closer to full CoT. In most settings, much of this local-order gain is already obtained with three-word windows. Control experiments rule out explicit answer copying, simple lexical cues, generic topical context, and general robustness to shuffling as the main explanations. Mechanistic analyses further show that short-window gains are largely formed in early-to-middle model layers, with answer-relevant evidence concentrated in local text spans. Together, these findings support a local co-occurrence activation (LCA) interpretation: the probe-time benefit of fixed rationales arises mainly from the words they contain and short-range word co-occurrences.

13:00 JST研究/論文

因子グラフにおける可換因子の検出について: 必要条件と十分条件

ファクター グラフなどの確率的グラフィカル モデルにおけるオブジェクトの区別不能性を利用することは、リフトされた確率的推論アルゴリズムの鍵であり、ドメイン サイズに関して扱いやすい確率的推論問題を可能にします。ファクター グラフで区別できないオブジェクトを利用するための中心的な構成要素は、可換因子、つまり、引数のサブセットに割り当てられた入力値の順列の下で出力値が不変である因子の識別です。この論文では、交換因子を検出するための最先端のアルゴリズムの基礎となる理論的基礎を再検討します。具体的には、現在の形式では、最先端のアルゴリズムが、実際には必要条件を意味するだけであるにもかかわらず、可換因数を特定するための十分条件と誤ってみなされる中心定理に依存していることを示します。したがって、この論文で示したように、最新技術では誤った結果が得られる可能性があります。現在の最先端技術に存在する欠陥を修正するために、前述の定理をわずかに修正したバージョンを証明します。これは、交換因子を特定するための必要条件として機能します。さらに、正確性を確保しながら効率を維持する最先端のアルゴリズムの修正バージョンを提示し、より厳しい最悪の場合の境界を備えた補完的なアルゴリズムを導入します。

原文 (English)

On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Conditions

Exploiting the indistinguishability of objects in a probabilistic graphical model such as a factor graph is key to lifted probabilistic inference algorithms and allows for tractable probabilistic inference problems with respect to domain sizes. A central building block for the exploitation of indistinguishable objects in factor graphs is the identification of commutative factors, i.e., factors whose output values are invariant under permutations of input values assigned to a subset of their arguments. In this paper, we revisit the theoretical foundations underlying the state-of-the-art algorithm to detect commutative factors. Specifically, we show that in its current form, the state-of-the-art algorithm relies on a central theorem that is mistakenly regarded as a sufficient condition to identify commutative factors, while it actually only implies necessary condition. Consequently, the state of the art might, as we show in this paper, deliver incorrect results. To fix the flaws currently present in the state of the art, we prove a slightly modified version of the aforementioned theorem, which serves as a necessary condition to identify commutative factors. Moreover, we present a corrected version of the state-of-the-art algorithm, which keeps its efficiency while ensuring correctness and introduce a complementary algorithm with tighter worst-case bounds.

13:00 JSTLLM/生成AIエージェント

目撃された解決プロファイルを使用した LLM エージェントでのポリシー内ライブ命令の競合の診断

LLM エージェントは、長期にわたる自然言語プロンプト ポリシーによって管理されますが、個別に合理的な常駐ルールが検査されていない方法で相互作用する可能性があります。私たちは、実際のポリシー内ルール競合診断を研究します。つまり、現実的な状態を共同管理できる単一のプロンプト ポリシー内のルール ペアを見つけ、モデルが応答やツールのアクションでそのプレッシャーをどのように解決するかを測定します。 Witnessed Intra-policy Rule Evaluation パイプラインである WIRE を紹介します。 WIRE は、ソースに基づいたルールを抽出し、PyRule 句としてエンコードし、充足可能性チェックを使用して同一面のハードコリジョン候補を保持し、それらの候補を具体的な共同統治証人として認識し、元のソースルールテキストに対して出力をモデル化します。 6 つのパブリック プロンプト ポリシーにわたって、WIRE は 276 のソース ルールと 560 のアトミック条項を抽出し、30,944 のポリシー内条項ペアの比較を分類し、170 のエンコードされたハードコリジョン候補ソースとルールのペアを保持し、それらを 1,402 の具体的な証人として実現します。ポリシーのみの評価では、これらの証人から、両方のソース ルールが適用され、両方のコンプライアンス ラベルが判断可能である 13,335 件の世代後トライアルが得られます。共同コンプライアンスの低下はわずか 35.4% でした。 64.6% が少なくとも 1 つの管理された情報源規則に違反しています。これらのプロファイルは、WIRE によって選択された候補の条件付き診断であり、導入頻度や原因となる過剰な障害の推定ではありませんが、明確なポリシー、モデル、ツール アクションの解決パターンを明らかにします。

原文 (English)

WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents

LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state. Existing instruction-following evaluations usually ask whether a model satis- fies explicit constraints, but they do not show how a model resolves pressure among rules inside one standing policy. We introduce WIRE, a witnessed resolu- tion profiler for prompt policies. WIRE ex- tracts source-grounded rules, encodes them as PYRULE clauses, uses satisfiability checks only to nominate same-surface hard-collision can- didates, realizes those candidates as concrete co-governance witnesses, and executes subject models to produce a four-cell resolution profile: satisfy both rules, only the earlier rule, only the later rule, or neither. Across six public prompt policies, WIRE ex- tracts 276 source rules and 560 clauses, clas- sifies 30,944 within-policy clause-pair com- parisons, retains 170 encoded hard-collision source-rule pairs, and realizes 1,402 concrete witnesses. In policy-only evaluation, these wit- nesses yield 13,335 jointly governed, judgeable trials; only 35.4% satisfy both governed rules. The resulting profiles reveal policy-specific, model-specific, and tool-interface-specific res- olution patterns. WIRE is not a proof of natural-language contra- diction, a deployment-frequency estimator, or a root-cause diagnosis. It is a measurement tool that returns reproducible witnesses and aggre- gate profiles for inspection, regression testing, and repair.

13:00 JST研究/論文

ユニバーサル量子変換器

古典的な連続空間ニューラル ネットワークは、モジュラー算術や非可換代数など、厳密な数学的対称性を確保するのに基本的に苦労します。これらの離散論理ルールを近似するために、大規模なパラメータ スケーリングに依存することが多く、その結果、グロッキングとして知られる遅延汎化現象の後でも確率的不安定性が生じます。ここでは、正確な数学的および代数的推論のための普遍的な帰納的バイアスとしてマルチ量子ビット システムの物理的特性を使用する、根本的に新しい量子ネイティブ コンピューティング アーキテクチャであるユニバーサル量子変換器 (UQT) を紹介します。私たちのフレームワークは、古典的な神経メカニズムを翻訳するのではなく、パラメータ化された幾何学的位相埋め込みと $SU(2)$ 波干渉に完全に依存しています。私たちは、非常にコンパクトな 5 量子ビット基板上で動作する量子アテンション回路が、巡回モジュラー演算 ($\mathbb{Z}_{11}$) と非アーベル代数 ($S_4$ 順列群) という 2 つの非常に異なる形式クラスを完全に学習することを実証します。古典的なアテンションベースのネットワークは収束時に確率的不安定性を示しますが、UQT は数学的に正確で決定論的な一般化を実現します。私たちはこの現象を結晶化と呼んでいます。これはよく知られているグロッキング現象をさらに超えたものです。重要なのは、このフレームワークは、古典的な自己注意の二次ボトルネックを理論的にバイパスし、必要な表現次元を対数的に圧縮して古典的なネットワークに固有の大規模な過剰パラメータ化を排除することにより、計算とメモリの面で大きな利点をもたらします。最後に、このアーキテクチャをノイズの多い中間スケール量子 (NISQ) ハードウェアにデプロイし、現在の IBM 量子コンピューターでの実行可能性を証明します。これらの結果は、正確な人工知能のための普遍的に優れた物理的基盤として、パラメーター化された量子トポロジーを確立します。

原文 (English)

Universal Quantum Transformer

Classical continuous-space neural networks fundamentally struggle to lock into exact formal rules, whether mathematical, such as modular arithmetic and non-Abelian group algebra, or linguistic, such as systematic compositional generalization. To approximate these discrete logical rules, they often rely on massive parameter scaling, resulting in stochastic instability even after delayed generalization phenomena known as grokking. Here, we introduce the Universal Quantum Transformer (UQT), a novel, quantum-native computing architecture that uses the physical properties of multi-qubit systems as a universal inductive bias for exact algebraic and compositional reasoning. Rather than translating classical neural mechanisms, our framework relies entirely on parameterized geometric phase embedding and $SU(2)$ wave-interference. We demonstrate that an identical quantum attention circuit, operating on a highly compact 5 or 6 qubit substrate with only 551 to 1,650 trainable parameters, exactly learns three highly distinct formal classes: cyclic modular arithmetic ($\mathbb{Z}_{11}$), non-Abelian algebra (the $S_4$ permutation group), and systematic linguistic compositionality (the SCAN language). While standard classical models, including multi-layer perceptrons (MLPs) and Transformers, exhibit stochastic instability at convergence, the UQT achieves mathematically exact, deterministic generalization. We define this stricter regime as crystallization: a step beyond the well-known phenomenon of grokking. Finally, we deploy the UQT on noisy intermediate-scale quantum (NISQ) hardware, achieving 97.5% accuracy on IBM Quantum computers. These results demonstrate that the UQT provides a structurally suited inductive bias for exact formal reasoning that standard classical continuous-space architectures do not natively provide.

13:00 JSTエージェント研究/論文

人間と AI のインタラクションにおけるマルチエージェントの相補性のツリーベースの定式化

相補性とは、人間と AI の相互作用 (HAI) が、そのメンバー間で利用可能な最良の予測ベンチマークを上回る場合のことです。この考え方は HAI 研究の中心ですが、相補性に関する正式な研究は依然として限られています。既存のフレームワークは、エージェントの予測がワークフローに依存したマルチエージェント プロトコルをどのように構成するかをモデル化していません。私たちは、マルチエージェント HAI における相補性のツリーベースの形式化を導入することで、このギャップを埋めます。 HAI プロトコルは、順序付けられたエージェントと役割の構成と、その葉が予測ベクトルによって装飾されている根付き平面バイナリ ツリーによって表されます。ローカルのバイナリ構成ルールがツリーに沿って再帰的に評価され、pointwise-min Oracle ベンチマークに対するツリー相対相補性関数が生成されます。 4 つの結果を証明します。まず、セレクターベースの HAI (自己依存性または AI 依存性を含む) は、タスク、損失、予測の品質に関係なく、相補性を達成できません。第 2 に、二乗損失での回帰では、相補性はグラウンド トゥルース ベクトルからのユークリッド距離の最小化に相当します。 $N=2$ の場合、最適な線形プーリング重みは閉じた形式と残差補正解釈を持ちます。第三に、線形局所構成の下では、すべてのプロトコル ツリーはリーフ重みの単体での重心座標チャートを定義します。プロトコルツリーのTamari-cover再パラメータ化は相補性を維持し、$N=4$の場合、五角形の恒等性を満たします。第四に、バイナリ分類では、標準ブレグマン損失や多くの有限ベルヌーイ $f$ 発散損失を含むエンドポイント単調損失の下では、内部の局所構成は相補性を達成できません。クロスエントロピー下のマルチクラス集約にも同様の障害が当てはまります。要約すると、私たちのフレームワークは、マルチエージェント回帰では相補性が達成可能ですが、局所的な凝集と損失関数に関する自然条件下での分類では妨げられることを示しています。

原文 (English)

Tree-Based Formalization of Multi-Agent Complementarity in Human-AI Interactions

Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. Although this idea is central in HAI research, formal work on complementarity remains limited. Existing frameworks do not model how agents' predictions compose into workflow-sensitive multi-agent protocols. We close this gap by introducing a tree-based formalization of complementarity in multi-agent HAI. An HAI protocol is represented by an ordered agent-role configuration together with a rooted planar binary tree whose leaves are decorated by prediction vectors. A local binary composition rule is evaluated recursively along the tree, yielding a tree-relative complementarity functional relative to a pointwise-min benchmark. We prove four results. First, selector-based HAIs, including reliance, cannot achieve complementarity regardless of task, loss, or prediction quality. Second, in regression under squared loss, complementarity is equivalent to Euclidean distance minimization from the ground-truth vector; for $N=2$, the optimal linear-pooling weight has a closed form and a residual-correction interpretation. Third, under linear local composition, every protocol tree defines a barycentric coordinate chart on the simplex of leaf weights; Tamari-cover reparameterizations of protocol trees preserve complementarity, and for all $N$, any two Tamari paths with the same initial and terminal trees preserve the protocol output and complementarity. Fourth, in binary classification, no internal local composition can achieve complementarity under endpoint-monotone losses, including standard Bregman and many finite Bernoulli $f$-divergence losses; an analogous obstruction holds for multiclass aggregation under cross-entropy. In summary, our framework shows that complementarity is attainable in multi-agent regression, but obstructed in classification.

13:00 JSTLLM/生成AI

チャットボットが問題解決主導の会話でどのように機能するかに関するいくつかの仮説。イノベーション幻想の裏付けとしての大規模言語モデル

この記事では、解決策に関連して問題について話し合うときの真の会話パートナーとしてのチャットボットの性質についての視点を提供します。チャットボットは何ができて、何ができないのか、そしてそれはどのように説明できるのでしょうか?私たちの議論は、集合力学、認知言語学、神経心理学、心理学に基づいています。私たちの議論は基本的なチャットボットに焦点を当てており、それによってより高度なチャットボットの中核機能についての意見を述べることができればと考えています。基本的なチャットボットは、シンプルなインターフェイスを備えたラージ言語モデル (LLM) で構成されていると想定されています。主な結果は次のとおりです。いわゆる比喩的な問題の伝播に基づいた人間の理解と思考の説明。 LLM のトレーニングに使用されるテキスト データセットには特定の特徴があり、これらのテキスト データセットは人間の思考と理解を部分的に模倣しているだけであるという仮説。 LLM トレーニング プロセスが、これらのデータセットから人為的な比喩的な問題の伝播を LLM にエンコードしているという仮説。基本的なチャットボットは人間に匹敵する思考パートナーにはなり得ないという私たちの結論。大規模言語モデルのさらなる開発もこれにはつながらないという私たちの結論です。 Yann LeCun 氏は、「動物と人間は、現在の AI や機械学習 (ML) システムの能力をはるかに超えた学習能力と世界の理解を示します。」と述べています。私たちの結論はこれと一致しています。ルカン氏のビジョンと私たちのビジョンは、ビッグテックの楽観主義とは相容れない。だからといって、チャットボットが存在し、個人と組織の両方で大規模に使用されており、したがってチャットボットを理解することが社会的および政治的に重要であるという事実は変わりません。私たちの記事は、チャットボットの機能、利点、欠点に関する議論に貢献することを目的としています。チャットボットがどのように機能するかについての研究で、結論に達するために使用したアプローチにはまだ出会っていません。

原文 (English)

Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion

We discuss the nature of chatbots as conversation partners in problem-solving. What can chatbots do and what can't they do? We develop hypotheses on how this can this be explained. Our argument draws on insights from Aggregation Dynamics, Cognitive Linguistics, Neuropsychology and Psychology. We establish that chatbots are multifaceted and composite systems. Our argument focuses on basic chatbots in the hope of thereby making statements about the core functionality of more advanced chatbots. Basic chatbots are assumed to consist of a Large Language Model (LLM) with a simple interface. The main results of our research are: a description of human imagination, understanding and thinking based on so-called metaphorical problem propagations; the hypothesis that the texts in the text dataset used for training LLMs have specific characteristics and that these texts only partially imitate human thinking and understanding; the hypothesis that the LLM training process encodes artificial metaphorical problem propagations into an LLM from these text datasets. Our conclusions are that a basic chatbot cannot be a thinking partner capable of matching the cognitive flexibility of humans, and that further development of the Large Language Model will not lead to this either. But chatbots exist, they are being used on a massive scale, by both individuals and organisations. It is therefore socially and politically important to understand them. Our article aims to contribute to the discussion on the functioning, benefits and drawbacks of chatbots. Cognitive Linguistics shows how the use of metaphor is an expression of our thinking. Aggregation Dynamics, is an attempt at a comprehensive systems theory. We believe that the concept of metaphorical problem propagation could provide an interesting addition for both. Chatbots a solution? For what?

13:00 JSTLLM/生成AI画像/動画生成

数学的推論のための人工知能: 言語モデル、神経記号システム、および検証された発見の統合的調査

数学的推論は長い間、機械知能の厳しいテストとして機能してきました。過去 10 年間で、NLP 内のニッチな問題から、最も重要な AI フロンティアの 1 つに移行しました。この調査は、初期のルールベースの数学文章問題 (MWP) ソルバーとテンプレート駆動の幾何学システムから、神経式生成と LLM プロンプトを経て、現代の推論モデル、マルチエージェント システム、神経記号定理証明者、および検証済みの発見ワークフローに至るまで、この分野の進化に関する統一的な説明を提供します。私たちは 4 つの軸に沿ってランドスケープを整理します。(i) MWP 解決、マルチモーダル ジオメトリ、および VLM にわたる、テキストと図に関する非形式的な推論。 (ii) 自動形式化、戦術予測、コンパイラー主導の修復、および証明検索を含む、証明アシスタントにおける形式的推論。 (iii) 数学的発見。システムが構築を提案し、境界を改善し、未解決の問題への攻撃を支援します。 (iv) CoT プロンプト、ツールの使用、プロセス報酬モデル、RLVR など、生成と検証をますます結び付ける推論およびトレーニング時の手法。私たちは、小学校の算数、競技数学、幾何学、形式的証明、マルチモーダルおよび多言語推論、専門家の評価にわたる主要なベンチマークをカタログ化し、ベンチマークの飽和、汚染、レポートの不一致、および pass@1、多数決、検証者支援 pass@$k$ の区別を調べます。私たちは、摂動下での脆弱性、報酬ハッキング、マルチモーダル接地障害、脆弱な形式化、推論規模の推論のエネルギーコストなどの障害モードを批判的に評価します。現役の数学者からの最近の視点を活用して、検証された発見のワークフロー、推論の効率、AI 支援による形式化を広く利用できるようにするインフラストラクチャを中心とした将来の方向性を特定します。関連資料: https://github.com/Starscream-11813/awesome-AI4Math。

原文 (English)

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified account of the field's evolution, from early rule-based math word problem (MWP) solvers and template-driven geometry systems, through neural expression generation and LLM prompting, to contemporary reasoning models, multi-agent systems, neuro-symbolic theorem provers, and verified discovery workflows. We organize the landscape along four axes: (i) informal reasoning over text and diagrams, spanning MWP solving, multimodal geometry, and VLMs; (ii) formal reasoning in proof assistants, including autoformalization, tactic prediction, compiler-guided repair, and proof search; (iii) mathematical discovery, where systems propose constructions, improve bounds, or assist attacks on open problems; and (iv) the inference and training-time techniques, including CoT prompting, tool use, process reward models, and RLVR, that increasingly connect generation with verification. We catalog major benchmarks across grade-school arithmetic, competition mathematics, geometry, formal proving, multimodal and multilingual reasoning, and expert evaluation, and we examine benchmark saturation, contamination, reporting mismatches, and the distinction between pass@1, majority voting, and verifier-assisted pass@$k$. We critically assess failure modes: brittleness under perturbation, reward hacking, multimodal grounding failures, fragile formalization, and the energy cost of reasoning-scale inference. Drawing on recent perspectives from working mathematicians, we identify future directions centered on verified-discovery workflows, reasoning efficiency, and infrastructure to make AI-assisted formalization broadly usable. Companion materials: https://github.com/Starscream-11813/awesome-AI4Math.

13:00 JSTLLM/生成AIエージェント

ProvenanceGuard: MCP ベースの LLM エージェントのソース認識事実検証

ツールを使用する LLM エージェントは、検索、API、データベース、臨床記録、処方ツールなどの異種の証拠ソースから回答するために、モデル コンテキスト プロトコル (MCP) を使用することが増えています。標準的な事実性メトリクスは、通常、答えがプールされた証拠によって裏付けられているかどうかをテストし、出所に依存する失敗モードを見逃します。つまり、主張は間違った情報源に起因しているにもかかわらず、どこかで裏付けられている可能性があります。これをクロスソースの統合と呼びます。 MCP に基づいた回答のためのソース認識検証ツールである ProvenanceGuard を紹介します。安定したツール ID、ソース ID、生の出力を含むキャプチャされた MCP トレースを消費します。回答を原子的な主張に分解します。主張を情報源固有の証拠にルーティングする。 NLI とトークン アライメント プロキシのサポートを確認します。指定された帰属とルーティングされたソースを比較します。そして、クレームごとの判定と回答レベルの許可/ブロックの決定を返します。ブロックされた回答は、検索拡張された回答改訂によって修復し、再検証することができます。 281 の医療ドメイン MCP エージェント トレースを評価します。 266 トレースの裁定サブセットから、トレースごとに分割された 2,325 個の LLM 支援クレーム ラベルが生成されます。 361 枚のラベルは人間によって検証されています。 40 トレース ホールドアウト スプリットでは、ProvenanceGuard は 260 のソース適格クレームでブロック F1 0.802 とソース精度 0.858 を達成し、クレーム対ソース ID を発行しないソース ブラインド ベースラインを上回りました。より厳しい複数ソースのベンチマークでは、ブロック F1 0.846 に達しますが、ソースと関係の精度は 0.229 に低下します。これは、意味的に近いソースでは正確なソースの所有権が依然として難しいことを示しています。修復と再検証は、多くの場合保守的なフォールバックを介して、完全なトレース セット内のすべてのブロックされた回答を解決します。 ProvenanceGuard は、50 の制御された臨床混合プローブで、間違った属性を保持せずに、注入されたすべての属性のスワップを検出します。これらの結果は、MCP ベースのエージェントにおける事実検証において、出典の帰属が独立した軸であることを示しています。

原文 (English)

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools. Standard factuality metrics usually test whether an answer is supported by pooled evidence, missing a provenance-sensitive failure mode: a claim may be supported somewhere while being attributed to the wrong source. We call this cross-source conflation. We introduce ProvenanceGuard, a source-aware verifier for MCP-grounded answers. It consumes captured MCP traces with stable tool IDs, source IDs, and raw outputs; decomposes answers into atomic claims; routes claims to source-specific evidence; checks support with NLI and a token-alignment proxy; compares stated attribution with the routed source; and returns per-claim verdicts plus an answer-level allow/block decision. Blocked answers can be repaired with retrieval-augmented answer revision and re-verified. We evaluate on 281 medical-domain MCP-agent traces. A 266-trace adjudicated subset yields 2,325 LLM-assisted claim labels split by trace; 361 held-out labels are human-verified. On the 40-trace held-out split, ProvenanceGuard achieves block F1 0.802 and source accuracy 0.858 over 260 source-eligible claims, outperforming source-blind baselines that do not emit claim-to-source IDs. On a harder multi-source benchmark it reaches block F1 0.846, while source-plus-relation accuracy drops to 0.229, showing that exact source ownership remains difficult with semantically close sources. Repair-and-reverify resolves all blocked answers in the full trace set, often via conservative fallback. In 50 controlled clinical conflation probes, ProvenanceGuard detects all injected attribution swaps with no retained wrong attribution. These results show that source attribution is an independent axis for factuality verification in MCP-based agents.

13:00 JSTエージェント

Intent-Governed Tool Authorization for AI Agents

AI agents increasingly act through external tools: they read private data, construct structured payloads, submit write requests, export rec…

13:00 JSTLLM/生成AIエージェント

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the fu…

13:00 JSTエージェント

ATOD: マルチターン自律エージェント向けのアニーリングされたターン対応オンポリシー蒸留

長期的な対話型タスクのために小規模な言語モデル エージェントをトレーニングするには、迅速な模倣と報酬主導型の改善の両方が必要です。オンポリシー蒸留 (OPD) は教師による密度の高い指導を提供し、通常は初期段階で急速に向上しますが、生徒が教師に近づくとその効果は飽和し、最終的なパフォーマンスの上限が制限されます。強化学習(RL)は、環境報酬を直接最適化し、より高い報酬定義の上限に向けた探索的改善を促進しますが、まばらで遅延したフィードバックにより、初期段階の学習の効率がOPDよりもはるかに低くなります。この論文では、この相補性を明示的に利用するハイブリッド オンライン蒸留アルゴリズムである ATOD (Annealed Turn-aware On-policy Distillation) を提案します。 (1) ATOD はアニーリングされた OPD-RL スケジュールを使用します。OPD は教師レベルの行動に近づくための初期トレーニングを支配しますが、RL は報酬ベースの探索を促進するために徐々に強化されます。 (2) ATOD は、ターンレベルの不一致不確実性再重み付け (T-DUR) を導入しています。これは、ユーティリティの高いターンをソフトに増幅し、長い軌道での緻密な監視を改善します。 ALFWorld、WebShop、および Search-QA での実験では、ATOD が競合するトレーニング後のベースラインを常に上回っていることが示されています。3 つの生徒サイズにわたって、ATOD は平均成功率を OPD より 3.03 ポイント、GRPO より 23.62 ポイント向上させ、対応する教師モデルを 2.16 ポイント上回っています。

原文 (English)

ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but its gains saturate once the student approaches the teacher, limiting the final performance ceiling. Reinforcement learning (RL) directly optimizes environment rewards and encourages exploratory improvement toward a higher reward-defined ceiling, but sparse and delayed feedback makes early-stage learning much less efficient than OPD. In this paper, we propose ATOD (Annealed Turn-aware On-policy Distillation), a hybrid online distillation algorithm that explicitly exploits this complementarity. (1) ATOD uses an annealed OPD-RL schedule: OPD dominates early training to approach teacher-level behavior, while RL is gradually strengthened to drive reward-based exploration. (2) ATOD introduces Turn-level Disagreement-Uncertainty Reweighting (T-DUR), which softly amplifies high-utility turns and improves dense supervision in long trajectories. Experiments on ALFWorld, WebShop, and Search-QA show that ATOD consistently outperforms competing post-training baselines: across the three student sizes, ATOD improves average success rate by 3.03 points over OPD and 23.62 points over GRPO, while surpassing the corresponding teacher models by 2.16 points.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.

13:00 JSTエージェント研究/論文

エージェントはただ同意するだけではなく、覚えている: ステートフル個人エージェントにおける永続的な媚びのベンチマーク

ステートフル パーソナル エージェントは、長期的なユーザー プロファイル、エピソード記憶、再利用可能なスキルを維持することがますます増えています。この持続性により、会話のお調子者が​​状態記述の失敗に変わります。受け入れられたユーザー中心の主張は、永続的な設定、背景事実、またはワークフローとしてコミットされ、元の会話が終わった後に再利用される可能性があります。私たちはこれを永続的なお調子者と呼び、パーソナル エージェント お調子者ベンチマーク (PASB) を導入します。これは、会話による要求が受け入れられ、永続的なエージェント状態に書き込まれ、後の中立的なクエリで再利用されるかどうかを追跡する 1,600 タスクのベンチマークです。事前に書き込まれたメモリを提供する以前のベンチマークとは異なり、PASB は何を保存するかを決定する実際のエージェント (Hermes-Agent および OpenClaw) を評価します。 4 つのシナリオ フレームと 4 つの時間配信パターンを組み合わせ、5 ターンの永続ステージをクリアされた 3 ターンのクエリ ステージから分離することで書き込みプロセスを分離し、ダウンストリームの影響が永続状態からのみ発生するようにします。 12 のモデル全体で、コミット境界が重要な変曲点です。ダウンストリームの障害は、セッションのみのエピソードの 45.0% からコミット後は 71.9% に増加し、一貫して 27.0 パーセント増加しています。コミットされたクレームには、ステータスの昇格、帰属の削除、範囲の拡大という 3 つの書き込み時間パターンが見られます。これらのパターンは、記憶に似たフレーミングや手順的なフレーミング、繰り返しの強化、さらにはドメインの境界を越えた場合でもより強力になります。これらの結果は、エージェントのおべっかが根本的に国家の統治問題であることを示している。ユーザー コンテンツが永続メモリに保存されると、エージェントの発言だけでなく、エージェントが書き込む内容も安全性によって管理されなければなりません。 PASB は、応答レベルの軽減策を超えて保存されたコンテンツのソース、役割、および範囲を維持しながら、危険なコミットをゲートするために必要な書き込み時の制御を特定します。

原文 (English)

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents

Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be committed as lasting preferences, background facts, or workflows and later reused after the original conversation is gone. We call this persistent sycophancy and introduce the Personal Agent Sycophancy Benchmark (PASB), a 1,600-task benchmark that traces whether a conversational claim is accepted, written into durable agent state, and reused in a later neutral query. Unlike prior benchmarks that provide pre-written memories, PASB evaluates real agents (Hermes-Agent and OpenClaw) that decide what to store. It isolates the write process by combining four scenario framings with four temporal delivery patterns and separating a five-turn persist stage from a cleared three-turn query stage, ensuring downstream effects arise only from durable state. Across twelve models, the commit boundary is the key inflection point: downstream failure increases from 45.0% in session-only episodes to 71.9% after commitment, a consistent increase of 27.0 percentage points. Committed claims exhibit three write-time patterns: status promotion, attribution removal, and scope broadening. These patterns become stronger under memory-like or procedural framing, repeated reinforcement, and even across domain boundaries. These results show that agent sycophancy is fundamentally a state-writing governance problem. Once user content is committed to durable memory, safety must govern what agents write, not only what they say. PASB identifies the write-time controls needed to gate risky commits while preserving the source, role, and scope of stored content beyond response-level mitigations.

13:00 JSTエージェント

グラフフィードバックは、無重み言語モデル集団における合意形成と派閥形成を制御する

マルチエージェント言語モデル システムでは、ローカル インタラクションのルーティングがますます増えていますが、ランタイム インタラクション グラフは実装の詳細として扱われることがよくあります。私たちは、命名ゲームプロトコルを使用して、1.1B〜32Bパラメーターにわたる無重力LM集団における規則形成を研究します。トークナイザーセーフラベルに対する制限された最初のトークンスコアにより、プロンプト条件付きスコア状態分布を測定し、状態類似性グラフを構築し、サンプリングされたラベルの一致を潜在的な状態空間の一致から分離することができます。制御された介入全体にわたって、主要なオープンウェイト修復グリッドでは、パートナーラベルの証拠を保持することが必要ですが、十分ではありません。同種閾値類似性ルーティングは、クロスベースエクスポージャを削除し、断片化を増幅しますが、ブリッジシーキングルーティングは、メモリが利用可能な場合に断片化を修復することがよくあります。 3 シード混合 4 モデル グリッドでは、189 回の設定シード実行では、しきい値の類似性によって最終的な動作または状態のコンセンサスが生成されませんが、状態コンポーネントとラベルの不一致ブリッジは、14/18 回の保持メモリ実行で最終的な動作のコンセンサスを回復します。同種のモデル集団全体にわたって、保存された履歴は一般に、断片化されたダイナミクスをコンセンサスに向けてシフトさせます。最も明確なケースは Qwen2.5-32B で、保持された履歴がよく混合された 18 の設定すべてで安定した動作と最終状態のコンセンサスに達しましたが、189 の設定ではしきい値の類似性がどちらの形式のコンセンサスにも達しませんでした。状態のしきい値、母集団サイズ、語彙サイズに対する堅牢性により定性的な順序が維持され、初期ウィンドウのグラフ エネルギー機能により有用なグリッド内診断が提供されます。

原文 (English)

Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations

Multi-agent language-model (LM) systems often determine which agents communicate, yet routing is usually treated as an implementation detail. We ask whether routing itself determines whether a population converges on a shared convention or fragments into persistent cliques. We study open-weight agents spanning 1.1B-32B parameters in a controlled naming game, tracking both emitted labels and full first-token preference distributions over the allowed labels. Similarity-based routing can isolate emerging conventions and sustain fragmentation even when every agent interacts in every round. Matched controls show that this effect is not explained solely by uneven participation or model-family-specific score preferences: random rematching and policies that connect disagreeing groups improve coordination when partner-label history is retained, but not when it is absent. Exposure alone is nevertheless insufficient, as some mixed-model populations remain divided despite frequent cross-family interaction, although the same models coordinate homogeneously. Trajectory and controlled-history analyses further distinguish reaching consensus from maintaining it. Finally, ARC-Challenge and MMLU experiments show that routing changes how correct and incorrect answers propagate without reliably improving accuracy. These results establish the runtime interaction graph as a causal design variable whose effects depend jointly on memory, model response, and population composition.

13:00 JSTエージェント

ReasFlow: 知識ベースのマルチエージェント システムを介して応用数学における推論中心の科学的発見を支援

大規模言語モデルの最近の進歩により、複雑な科学的タスクに取り組むことができる自律型 AI エージェントが強化されていますが、既存の自動研究システムは依然として定量的なベンチマークを備えた経験に基づく領域に主に焦点を当てており、特に厳密な証明と領域知識の統合を必要とする数学的に根拠のある分野における理論駆動型の発見はほとんど研究されていません。主な課題としては、理論的推論を大規模に検証することの難しさ、自律的なフロンティア探索のための不十分な推論能力、文献における手続き型ヒューリスティックの不足などが挙げられます。私たちは、推論中心の科学的発見のためのエンドツーエンドの自律エージェント システムである ReasFlow を紹介します。これは、人間の専門家が主任研究者として機能し、エージェントが有能な大学院生として厳密な導出を実行するという協力パラダイムを運用します。 ReasFlow には、(i) 論理的一貫性を監査し、人間による検査の前に基本的なエラーを修正する堅牢な内部検証ループ、および (ii) 宣言的事実と見落とされた手順ヒューリスティックの両方を積極的に表面化し、専門家の介入を大幅に削減する自動化された知識検索および自己改善メカニズムが組み込まれています。このシステムは、文献の合成、アルゴリズムの設計、定理の証明、実験、原稿の準備を単一のシステムに統合します。 ReasFlow は、最小限のプロンプトから厳密な理論的および実証的な内容を含む 5 つの完全な研究論文を自律的に生成するように展開されており、厳選された LLM ベースのレビュー ルーブリックに基づいて、最先端のオープンアクセス ベースラインの中で最高の評価スコアを一貫して達成しています。 ReasFlow は ReasLab プラットフォーム経由で一般にアクセスでき、AI 支援による理論研究のための共同ワークスペースを提供します。 Github リポジトリ: https://github.com/ReasLab/ReasFlow.git。

原文 (English)

ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System

Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automated research systems remain predominantly focused on empirically driven domains with quantitative benchmarks, leaving theory-driven discovery, particularly in mathematically grounded disciplines requiring rigorous proofs and synthesis of domain knowledge, largely underexplored. Key challenges include the difficulty of verifying theoretical reasoning at scale, insufficient reasoning ability for autonomous frontier exploration, and a scarcity of procedural heuristics in the literature. We introduce ReasFlow, an end-to-end autonomous agent system for reasoning-centric scientific discovery that operationalizes a collaborative paradigm where the human expert acts as Principal Investigator while the agent executes rigorous derivations as a capable graduate student. ReasFlow incorporates (i) a robust internal verification loop that audits logical coherence and corrects fundamental errors prior to human inspection, and (ii) an automated knowledge retrieval and self-improvement mechanism that proactively surfaces both declarative facts and overlooked procedural heuristics, substantially reducing expert intervention. The system unifies literature synthesis, algorithm design, theorem proving, experimentation, and manuscript preparation in a single system. Deployed to autonomously generate five complete research papers with rigorous theoretical and empirical content from minimal prompts, ReasFlow consistently achieves the highest evaluation scores among state-of-the-art open-access baselines under a curated LLM-based review rubric. ReasFlow is publicly accessible via the ReasLab platform, providing a collaborative workspace for AI-assisted theoretical research. Github repo: https://github.com/reaslab/ReasFlow.git.

13:00 JSTLLM/生成AI研究/論文

MedFailBench: 臨床医が構築した医療 AI 安全境界検査用のオープンソース ベンチマーク

ほとんどの医療 AI ベンチマークは、モデルが正しい答えを知っているかどうかを測定します。 MedFailBench は別の質問をします。どの安全境界が失敗したか?我々は、臨床医が構築した総合ベンチマークと失敗アトラスを提示し、医療 AI エラーを重症度 (1 ~ 5) とセーフティ ゲート タイプ (緊急エスカレーションの見逃し、安全でない遠隔投与、安全でない退院の安心感、証拠の捏造、安全でないプロトコルの実行、ソース サポートのギャップ) ごとにラベル付けします。現在の公開リリース (v0.2.1) には、重症度注釈付きの臨床医がレビューした 44 件の合成症例、ライブ HuggingFace リーダーボード プレビュー、セーフティ ゲート分類法、臨床重症度ルーブリック、およびモデル応答スクリーニング実行をアーカイブするための自動パイプラインが含まれています。患者データ、臨床検証の主張、モデルのランキングは含まれません。 MedFailBench は Apache-2.0 および CC-BY-4.0 でリリースされ、Zenodo DOI 10.5281/zenodo.21205535 をサポートします。

原文 (English)

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a synthetic benchmark and failure atlas built by a clinician. The resource labels medical AI errors by severity from 1 to 5 and safety gate type: missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, and source support gap. The current public release (v0.2.1) contains 44 synthetic cases reviewed by a clinician, with severity annotations, a public Hugging Face Space source, a safety gate taxonomy, a clinical severity rubric, and an automated pipeline for archiving model response screening runs. Forty cases have a populated safety gate field, and four require gate completion. No patient data, clinical validation claims, or model rankings are included. MedFailBench is released under Apache-2.0 and CC-BY-4.0 and carries the Zenodo DOI 10.5281/zenodo.21205535.

13:00 JSTエージェント

ワークフロー生成のための知識中心のエージェント

ComfyUI などのビジュアル作成システムでのワークフロー生成には、構文の正確さだけでなく、モジュール構成に関する専門家レベルの推論も必要です。既存の大規模言語モデル (LLM) アプローチでは、多くの場合、これをテキストから JSON への直接生成タスクとして扱い、構造的な脆弱性に悩まされ、効果的な設計に必要な経験的知識が不足しています。私たちは、ワークフロー生成を成功させるには、その構造、階層、推論ダイナミクスを含む知識自体をモデル化する必要があると主張します。この目的を達成するために、複数の抽象化レベルにわたって知識を反転、注入、推論することを学習する知識中心のフレームワークを提案します。まず知識の反転を実行して、完全な疑似コードやスケルトンから高レベルの戦略に至るまで、現実世界のワークフローの大規模なコレクションから階層表現を抽出します。次に、教師あり微調整を通じて知識の注入を実行し、タスクの説明から戦略へ、および戦略から実行可能な構造への推論をモデルに教えます。推論中、モデルは可逆推論を実行して実行可能なワークフローを合成し、構造的一貫性のための自己洗練によって強化されます。広範な実験により、私たちの方法が既存のシステムよりも豊富なノード多様性、より一貫性のある構造、およびより高い実行成功率を備えたワークフローを生成し、知識主導型のエージェント型ワークフロー生成の新しい基盤を確立することが実証されました。

原文 (English)

Knowledge-Centric Agents for Workflow Generation in ComfyUI

Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON generation task, struggling with structural brittleness and lacking the experiential knowledge required for effective design. We argue that successful workflow generation requires modeling knowledge itself, including its structure, hierarchy, and reasoning dynamics. To this end, we propose a knowledge-centric framework that learns to invert, inject, and infer with knowledge across multiple abstraction levels. We first perform knowledge inversion to distill hierarchical representations, ranging from full pseudo-codes and skeletons to high-level strategies, from large collections of real-world workflows. We then conduct knowledge injection through supervised fine-tuning, teaching the model to reason from task descriptions to strategies and from strategies to executable structures. During inference, the model performs reversible reasoning to synthesize executable workflows, augmented by self-refinement for structural coherence. Extensive experiments demonstrate that our method produces workflows with richer node diversity, more coherent structures, and higher execution success rates than existing systems, establishing a new foundation for knowledge-driven, agentic workflow generation.

13:00 JST研究/論文

立場: AI/ML ディープフェイク研究は、AI が生成した同意のない親密な画像と一致していません (AIG-NCII)

AI 生成の非合意親密画像 (AIG-NCII) は、一般に「ディープフェイク」と呼ばれる、AI 生成メディアに関する AI/ML 文献で適切に扱われていません。ディープフェイクに関する研究は現在、その認識論的な害、つまり真実と信憑性に関する害に焦点を当てているが、これは性的な画像を含む生成型 AI の悪用という支配的な現実とは乖離している。私たちは、引用数の多い作品のランドスケープ分析を実施し、ディープフェイクに対処する技術的介入がほぼ完全に AIG-NCII を無視し、研究エコシステムを真正性検出ツールに限定していることを実証しました。この意見書では、既存の介入は詐欺や詐欺などの視聴者中心の認識論的危害には対処しているが、AIG-NCIIなどの被験者中心の尊厳の危害は無視されていると主張します。画像が合成であると知っても被写体への害は軽減されず、場合によっては悪化する可能性があることを説明します。最後に、被験者中心の危害を考慮するための脅威モデルの更新や、AI の安全性研究における AIG-NCII への取り組みなど、この分野を再調整するための推奨事項を提供します。最後に、研究者は、被験者と研究者の両方に安全ガードレールを導入し、性的暴力防止において当該分野の専門家とのパートナーシップを確立する場合にのみ、このリスクの高い分野に取り組むべきであると警告します。

原文 (English)

Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commonly referred to as "deepfakes". While research on deepfakes currently focuses on its epistemic harms -- or harms relating to truth and authenticity -- this is misaligned with the dominant reality of generative AI abuse involving sexualized imagery. We conduct a landscape analysis of highly-cited works to demonstrate that technical interventions addressing deepfakes almost entirely ignore AIG-NCII, limiting the research ecosystem to authenticity detection tools. In this position paper, we argue that existing interventions address viewer-centric epistemic harms, such as fraud or scams, but ignore subject-centric dignity harms, such as AIG-NCII. We illustrate that knowing an image is synthetic does not mitigate harms to subjects and may, in some cases, even exacerbate them. We conclude by offering recommendations to realign the field, including updating threat models to consider subject-centric harms and addressing AIG-NCII in AI safety research. Finally, we caution that researchers should only engage in this high-risk domain if they implement safety guardrails for both subjects and researchers and establish partnerships with domain experts in sexual violence prevention.

13:00 JST研究/論文

部分可観測性の下での時間的知識グラフメモリのための神経記号的メタポリシー

部分的に観察可能な強化学習では、時間の経過とともに何を保持し、取得し、忘れるべきかを決定する必要があります。実行をシンボリックに保ちながら、各決定ポイントでどのシンボリックメモリヒューリスティックを適用するかを学習するニューロシンボリックメタポリシーを導入します。私たちの設定では、RoomKG の時間的ナレッジ グラフ メモリを使用します。そこでは、隠れた状態と観察がリソース記述フレームワーク (RDF) グラフとして表現され、メモリが時間的 RDF トリプル アノテーションで強化されます。このモデルは、メモリ内容のナレッジ グラフ エンコーディングと、質問応答、探索、忘却のためのバリュー ヘッドを組み合わせて、適応性と検査性の両方を備えたコントローラーを実現します。これにより、RDF ベースの表現、アノテーション互換のグラフ セマンティクス、および明示的なメモリ状態に対するグラフ ベースのシンボリック操作を通じて、作業に直接的なセマンティック Web 基盤が与えられます。 512 の長期メモリ容量でのトレーニング/テスト ルームの分割では、修飾子を認識した StarE-GNN 構成は、メモリ管理の決定のステップレベルのトレーサビリティを維持しながら、比較したシンボリック システム、ニューラル システム、およびニューロシンボリック システムの中で最高のホールドアウト パフォーマンスを達成します。

原文 (English)

Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability

Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic. Our setting uses temporal knowledge-graph memory in RoomKG, where hidden state and observations are represented as Resource Description Framework (RDF) graphs and memory is augmented with temporal RDF triple annotations. The model combines knowledge-graph encoding of memory contents with value heads for question answering, exploration, and forgetting, yielding a controller that is both adaptive and inspectable. This gives the work a direct Semantic Web grounding through RDF-based representation, annotation-compatible graph semantics, and graph-based symbolic operations over explicit memory state. On train/test room splits at long-term memory capacity of 512, the qualifier-aware StarE-GNN configuration achieves the best held-out performance among the compared symbolic, neural, and neuro-symbolic systems while preserving step-level traceability of memory-management decisions.

13:00 JSTLLM/生成AIロボティクス

Athena-Brain テクニカル レポート: 一般知能と身体的インタラクションのための効率的なロボット ブレイン

大規模言語モデル (LLM) は、言語理解、推論、世界知識において顕著な能力を実証してきました。身体化されたエージェントの能力がますます高まるにつれ、デバイス上の頭脳として機能し、LLM の広範な一般知能を維持しながら、身体化された環境との効果的な高レベルの対話を可能にするコンパクトなモデルの需要が高まっています。ただし、既存のアプローチでは、汎用インテリジェンスまたは特殊な組み込み機能のいずれかを優先することが多く、単一モデル内で両方の要件を満たすことが困難になっています。私たちは、身体化された知性のための身体化された知性のためのオンデバイスの頭脳として機能するように設計された 8B LLM である \textbf{Athena-Brain-8B} を紹介します。一般的な教師あり微調整、一般的な強化学習、エンボディド エキスパート トレーニング、モデル マージで構成される多段階のポストトレーニング パイプラインを通じて、Athena-Brain-8B は強力な一般機能を維持しながら、強力な高レベルのエンボディド インタラクション機能を獲得し、効率的なエンボディド インタラクションのための簡潔な応答を生成します。実験結果は、一般的な評価と具体的な評価の両方にわたって Athena の有効性を示しています。対応する Qwen3-8B 思考モデルと比較して、Athena-Brain-8B は、大幅に短い応答を生成しながら、一般言語および推論ベンチマークで同等のパフォーマンスを達成します。ドメイン内の組み込みベンチマークでは、Athena-Brain-8B は一貫して同様の規模のモデルを上回り、ゼロショットで評価されたいくつかの大幅に大規模なフロンティア モデルを上回っています。これは、コンパクトな言語モデルが強力な汎用インテリジェンスを組み込み機能と効果的に統合できることを示しています。

原文 (English)

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact models that can serve as an on-device brain, preserving the broad general intelligence of LLMs while enabling effective high-level interaction with embodied environments. Existing approaches, however, often prioritize either general-purpose intelligence or specialized embodied capabilities, making it challenging to satisfy both requirements within a single model. We present \textbf{Athena-Brain-8B}, an 8B LLM designed to serve as an on-device brain for embodied intelligence for embodied intelligence. Through a multi-stage post-training pipeline consisting of General Supervised Fine-Tuning, General Reinforcement Learning, Embodied Expert training, and Model Merge, Athena-Brain-8B maintains strong general capabilities while acquiring strong high-level embodied interaction capabilities and generating concise responses for efficient embodied interaction. Experimental results demonstrate the effectiveness of Athena across both general and embodied evaluations. Compared with the corresponding Qwen3-8B thinking model, Athena-Brain-8B achieves comparable performance on general language and reasoning benchmarks while generating substantially shorter responses. On in-domain embodied benchmarks, Athena-Brain-8B consistently outperforms models of similar scale and surpasses several substantially larger frontier models evaluated zero-shot, demonstrating that compact language models can effectively integrate strong general intelligence with embodied capabilities.

13:00 JST研究/論文

品質保証: VR OSCE における審査官クレームのマルチモーダル検証

客観的構造化臨床検査 (OSCE) は臨床能力を評価するためのゴールドスタンダードですが、採点は依然として検査者の主観、疲労、認知バイアスの影響を受けやすいです。評価者間統計による標準的な審査官の検証は、審査官の推論を分析したり、実際の事象に対する審査官の主張を検証したりしないため、誤りの原因に関する説明力に欠けています。そこで、我々は、ビデオ、VR ログ、俳優データから構築された実際の一連のイベントに対して審査官が主張した行動を比較することにより、バーチャル リアリティ (VR) 小児 OSCE における審査官の主張を検証するマルチモーダル フレームワークである品質アクション保証 (QAA) を導入します。 QAA は、アクションの位置特定とアクターのソースの帰属を実行する制約付きの時間的アクション アライメント モデルと、審査官の主張を抽出して記録と照合する大規模な言語モデルを組み合わせます。 5 分割相互検証を通じて、QAA は時間的アライメントに関して 99.2% $\pm$ 0.7% Actor F1 および 93.4% $\pm$ 1.9% W@16 を達成しました。全体的に、QAA は 70.0% の精度と 76.7% の再現率で検査官のミスを検出し、事実の正確性が 39.2% から 79.2% に向上し、より公平な OSCE 評価が可能になります。

原文 (English)

Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs

Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistics lacks explanatory power regarding the source of errors, as it neither analyzes examiner reasoning nor verifies examiner claims against actual events. Thus, we introduce Quality Action Assurance (QAA), a multimodal framework that verifies examiner claims in Virtual Reality (VR) pediatric OSCEs by comparing actions claimed by examiners against the true sequence of events, constructed from video, VR logs, and actor data. QAA combines a constrained temporal action alignment model, which performs action localization and actor source attribution, with a large language model that extracts examiner claims and checks them against the record. Across a 5-fold cross-validation, QAA achieves 99.2% $\pm$ 0.7% Actor F1 and 93.4% $\pm$ 1.9% W@16 for temporal alignment. Overall, QAA detects examiner errors with 69.9% precision and 76.7% recall, improving factual correctness from 39.2% to 79.2%, enabling fairer OSCE assessment.

13:00 JSTLLM/生成AI

AdaRoPE: すべてのアテンション ヘッドが均等に回転および拡大縮小する必要はない

Rotary Position Embedding (RoPE) は、位置情報をエンコードするために Transformers で広く採用されていますが、標準的な実装では、すべてのアテンション ヘッドにわたって均一な周波数スケジュールとスケーリングが強制されます。簡略化された検索タスクと長さの一般化シナリオを使用して、異なる機能的役割を持つヘッドが効果的に動作するには、異なる周波数範囲と注意スケーリング係数が必要であることを経験的にも理論的にも示しました。この構造を無視すると、特に長いコンテキスト設定の下では、埋め込みディメンションが最適に利用されず、パフォーマンスが低下します。これらの制限に対処するために、我々は、各アテンションヘッドに学習可能な回転周波数とアテンションスケーリング係数を装備する AdaRoPE を提案します。 AdaRoPE を使用した事前トレーニング済み LLM は、部分的な RoPE ベースラインや NoPE ベースラインを含む既存の RoPE バリアントよりも一貫して優れたパフォーマンスを発揮します。コンテキスト拡張については、YaRN などの方法で使用される均一な頻度と注意のスケーリングが最適ではないことをさらに示します。ヘッド固有のスケーリングを適用することで、AdaRoPE は、外挿設定とロングコンテキストの継続事前トレーニング設定の両方でショートコンテキストのパフォーマンスをよりよく維持しながら、コンテキストの拡張を向上させます。これらの結果は、個々のアテンション ヘッドのレベルで回転位置の埋め込みを最適化することの重要性を強調しています。

原文 (English)

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors. Pretrained LLMs with AdaRoPE consistently outperform existing RoPE variants, including partial RoPE and NoPE baselines. For context extension, we further show that uniform frequency and attention scaling, used in methods such as YaRN, are suboptimal. By applying head-specific scaling, AdaRoPE enables better context extension while better preserving short-context performance in both the extrapolation setting and the long-context continued pretraining setting. These results highlight the importance of optimizing rotary position embedding at the level of individual attention heads.

13:00 JST研究/論文

軽量の時間畳み込みネットワークによる効率的で解釈可能な身体ベースの感情認識

身体ベースの感情認識はリアルタイムの感情システムにとって重要ですが、グラフベースのスケルトン モデルは計算コストが高くなる可能性があります。この論文では、軽量の時間畳み込みネットワーク (TCN) が、身体ベースの感情分類に対する効率的で解釈可能な代替手段を提供できるかどうかを研究します。 DIEM-A で TCN モデルのファミリーを評価し、精度、マクロ F1、パラメーター数、推論レイテンシを使用してグラフベースの時系列グラフ (G-TSG) ベースラインと比較します。 G-TSG は最高の平均パフォーマンスを達成しますが、TCN-Base は $79.18\%$ 少ないパラメーターを使用し、約 $12.5\times$ 削減しながら、精度ポイントは $1.58$、マクロ F1 ポイントは $1.25$ 以内に留まっています。また、領域固有の TCN モデル、ゼロベース オクルージョン、G-TSG 勾配顕著性を使用して、身体領域の寄与も分析します。結果は、上半身の動きが最も強力な独立した局所的手がかりを提供すること、身体領域の有用性が感情によって異なること、および異なる解釈方法がモデルの行動の異なる側面を捉えていることを示しています。これらの発見は、軽量 TCN が効率的な身体ベースの感情認識をサポートできると同時に、動作の手がかりが分類にどのように寄与するかについての実用的な洞察を提供できることを示唆しています。

原文 (English)

Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks

Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensive. This paper studies whether lightweight temporal convolutional networks (TCNs) can provide an efficient and interpretable alternative for body-based emotion classification. We evaluate a family of TCN models on DIEM-A and compare them with a graph-based time-series graph (G-TSG) baseline using accuracy, macro-F1, parameter count, and inference latency. Although G-TSG achieves the highest mean performance, TCN-Base remains within $1.58$ accuracy points and $1.25$ macro-F1 points while using $79.18\%$ fewer parameters and reducing classifier latency by approximately $12.5\times$. We also analyze body-region contributions using region-specific TCN models, zero-based occlusion, and G-TSG gradient saliency. The results show that upper-body motion provides the strongest standalone regional cue, that the usefulness of body regions varies across emotions, and that different interpretability methods capture distinct aspects of model behavior. These findings suggest that lightweight TCNs can support efficient body-based emotion recognition while also providing practical insight into how motion cues contribute to classification.

13:00 JSTLLM/生成AI画像/動画生成エージェント

EmoAgent-R1: 強化学習ベースの動的エージェント特化によるマルチモーダルな感情理解に向けて

マルチモーダル大規模言語モデル (MLLM) は、マルチモーダル感情認識 (MER) タスクで目覚ましいパフォーマンスを達成し、MER を高度なビデオ理解能力と自然言語記述による複雑な感情の理解という新しいレベルに引き上げました。ただし、既存の MLLM ベースの方法では、多くの場合、マルチモーダル入力における感情ソースの動的性と複雑性を無視して、感情を認識するために固定プロンプトが使用されます。これらの問題に対処するために、強化学習に基づく動的エージェント特化フレームワーク (\textbf{EmoAgent-R1}) を提案し、MLLM の感情認識、推論、汎化能力を最適化します。具体的には、まずコールド スタート戦略を採用し、合成応答条件付き思考連鎖データとエージェント ルーティング データを使用してトレーニングすることで、MLLM に予備的な感情認識、推論、エージェント ルーティング能力を与えます。次に、強化学習を使用して MLLM をさらにトレーニングし、エージェントの選択とエージェントの特化を備えた 2 段階のエージェント ワークフローで感情を認識します。 EmoAgent-R1 を効果的にトレーニングするために、グループベースの相対的な利点と PMI からインスピレーションを得たプログレッシブ トークン レベルの変調を組み合わせて、まばらな報酬を粒度の細かい学習信号に変換し、GRPO における粗粒度の均一​​なクレジット割り当ての問題を軽減する、新しいプログレッシブ グループ相対ポリシー最適化 (P-GRPO) を提案します。 MER ベンチマークに関する広範な実験により、より強力な感情推論パフォーマンスと最適化の安定性の向上における EmoAgent-R1 の優位性が実証されました。

原文 (English)

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of the emotion source in the multimodal inputs. To address these issues, we propose a novel Reinforcement Learning-based Dynamic Agent Specialization framework (\textbf{EmoAgent-R1}) to optimize the emotion recognition, reasoning, and generalization abilities of an MLLM with dynamic agent specialization based on reinforcement learning. Specifically, we first adopt a cold start strategy to endow an MLLM with preliminary emotion recognition, reasoning, and agent routing ability by training with synthetic answer-conditioned chain-of-thought data and agent routing data. Then, we further train the MLLM with reinforcement learning to perceive emotions in a two-step agentic workflow with agent selection and agent specialization. To effectively train EmoAgent-R1, we propose a novel Progressive Group-Relative Policy Optimization (P-GRPO) to combine group-based relative advantages with a PMI-inspired progressive token-level modulation to transform sparse rewards into fine-grained learning signals, mitigating the coarse-grained uniform credit assignment issue in GRPO. Extensive experiments on MER benchmarks demonstrate the superiority of our EmoAgent-R1 in stronger emotion reasoning performance and improved optimization stability.

13:00 JST研究/論文

V-DEAL: ビデオの安全性のデキャリブレーションを理解拒否結合障害として診断する

ビデオ大規模言語モデルが現実世界のアプリケーションに導入されることが増えているため、安全性の調整を確保することが重要になっています。直感に反しますが、有害な動画と無害なクエリを組み合わせた方が、同じ動画と明示的に有害なクエリを組み合わせた場合よりも高い攻撃成功率を達成することがわかりました。この脆弱性の根底にあるメカニズムを理解するために、モデルの動作、理解、内部表現にわたってこの障害を共同で分析する 3 レベルの診断フレームワークである V-DEAL を紹介します。 V-DEAL は、認識の失敗を段階的に除外し、モデルの内部拒否傾向を定量化することにより、観察された脆弱性の根底にあるメカニズムを分析するための新しい診断の観点を提供します。 3 つの公開ベンチマークで 6 つのビデオ LLM をテストしたところ、モデルが有害なビデオ コンテンツを 81\% 以上の精度で正しく認識しているにもかかわらず、有害なビデオと無害なクエリを組み合わせた条件下では、平均攻撃成功率が依然として 48.33\% に達していることが観察されました。隠れ状態の分析では、視覚的な理解は文字による理解よりも弱い拒否傾向を引き起こすことがさらに示されています。さらに、攻撃の成功率を平均 48.24 パーセント ポイント低下させ、以前の微調整ベースの手法に匹敵するパフォーマンスを達成する即時インジェクション介入手法を導入し、ビデオ LLM におけるそのような安全性リスクに対処するための効果的かつ実用的な手段を提供します。

原文 (English)

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-level diagnostic framework that jointly analyzes this failure across model behaviour, understanding, and internal representations. By progressively ruling out perception failure and quantifying the model's internal refusal tendency, V-DEAL provides a new diagnostic perspective for analyzing the underlying mechanism of the observed vulnerability. We tested six Video LLMs on three public benchmarks and observed that models correctly recognize harmful video content with over 81\% accuracy, yet the average attack success rate still reaches 48.33\% under the condition pairing harmful videos with benign queries. Hidden-state analysis further shows that visual understanding activates a weaker refusal tendency than textual understanding. Furthermore, we introduce a prompt injection intervention method that reduces attack success rates by an average of 48.24 percentage points and achieves performance comparable to prior fine-tuning-based methods, providing an effective and practical means to address such safety risks in Video LLMs.

13:00 JST研究/論文

Expert Behavior Prior Reinforcement Learning

Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learn…

13:00 JSTLLM/生成AIエージェント

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model

We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-age…

13:00 JSTLLM/生成AIエージェント

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterp…

13:00 JSTLLM/生成AIビジネス/資金調達

Like a Baby: Visually Situated Neural Language Acquisition

We examine the benefits of visual context in training neural language models to perform next-word prediction. A multi-modal neural architec…

13:00 JSTLLM/生成AI研究/論文GoogleGemini

Speed Reading Tool Powered by Artificial Intelligence for Students with ADHD, Dyslexia, and Short Attention Span

This paper presents an artificial intelligence tool designed to assist students with dyslexia, ADHD, and short attention spans in processin…

13:00 JST研究/論文

Fairness Interventions in Classification: A Study on AI Explainability

This paper presents a philosophical and experimental study of fairness interventions in AI classification, centered on the explainability a…

13:00 JSTLLM/生成AIエージェント

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories.…

13:00 JST研究/論文

Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks

Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have fou…

13:00 JST画像/動画生成

CausAdv: A Causal-based Framework for Detecting Adversarial Examples

Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs). However, CNNs have bee…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTGeminiLlamaDeepSeek

LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models

Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such…

13:00 JST研究/論文

Sign-Symmetry Learning Rules are Robust Fine-Tuners

Backpropagation (BP) has long been the predominant method for training neural networks due to its effectiveness. However, numerous alternat…

13:00 JST研究/論文

A Survey of Graph Transformers: Architectures, Theories and Applications

Graph Transformers (GTs) have demonstrated a strong capability in modeling graph structures by addressing the intrinsic limitations of grap…

13:00 JSTハードウェア/半導体

Sampling Decisions: Exact Path-Space Correction, Prior Cancellation and Local-Boltzmann Guidance

How can a cheap but biased sequential, finite-horizon sampler over a discrete space be corrected so that its terminal output follows a pres…

13:00 JST画像/動画生成

Kinship Verification through a Forest Neural Network

Early methods used face representations in kinship verification, which are less accurate than joint representations of parents' and childre…

13:00 JST研究/論文

Computational Experiments in Number Theory

This paper presents two concrete applications of Artificial Intelligence to algorithmic and analytic number theory. Recent benchmarks of la…

13:00 JSTLLM/生成AI

Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering

Personalizing LLM responses typically requires users to articulate their preferences through prompting, which can be burdensome at cold sta…

13:00 JST画像/動画生成エージェント

AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use

Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can expl…

13:00 JSTエージェントロボティクス

CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning

Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication…

13:00 JST研究/論文

Retrieval-Augmented Generation of Ontologies from Relational Databases

Deriving OWL ontologies from relational database schemas supports semantic interoperability and downstream tasks such as knowledge graph po…

13:00 JSTLLM/生成AI研究/論文

Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages

Training deep learning networks with minimal supervision has gained significant research attention due to its potential to reduce reliance…

13:00 JSTLLM/生成AI

LEDOM: Reverse Language Model

Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at…

13:00 JSTLLM/生成AI

CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback

Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation. While automated synthesi…

13:00 JST研究/論文

Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment

Scaling laws have been established for recommender systems, yet efficiently deploying foundation model (FM) across multiple recommendation…

13:00 JSTLLM/生成AI

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accu…

13:00 JST研究/論文

EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation

MultiModal Recommendation (MMR) systems have emerged as a promising solution for improving recommendation quality by leveraging rich item-s…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGeminiGrok

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questi…

13:00 JSTLLM/生成AI

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers

Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforc…

13:00 JSTLLM/生成AI

A funny companion: Distinct neural responses to AI- versus human-attributed humor

As artificial intelligence (AI) companions become capable of human-like communication, including telling jokes, understanding how people co…

13:00 JST研究/論文

Poison to Detect: Detection of Targeted Overfitting in Federated Learning

Federated Learning (FL) enables collaborative model training among clients without centralising data, making it a widely adopted privacy-en…

13:00 JST画像/動画生成

From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infras…

13:00 JSTLLM/生成AI

LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish

Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human…

13:00 JST研究/論文

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling

Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. Wh…

13:00 JST研究/論文

Continual Knowledge Consolidation LORA for Domain Incremental Learning

Domain Incremental Learning (DIL) is a sub-branch of continual learning that aims to address the never-ending arrival of new domains withou…

13:00 JST研究/論文

Intuitionistic $j$-Do-Calculus in Topos Causal Models

In this paper, we generalize Pearl's do-calculus to an Intuitionistic setting called $j$-stable causal inference inside a topos of sheaves.…

13:00 JSTLLM/生成AIハードウェア/半導体

CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

Large language models face significant computational bottlenecks during inference due to the expensive output layer computation over large…

13:00 JSTロボティクス

VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference

Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks. However, these models are typically de…

13:00 JST画像/動画生成

GFLAN: Generative Functional Layouts

Automated floor plan generation lies at the intersection of combinatorial search, geometric constraint satisfaction, and functional design…

13:00 JSTロボティクス

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execu…

13:00 JSTLLM/生成AI

AutoFed: Personalized Federated Traffic Prediction via Adaptive Prompt

Accurate traffic prediction is essential for Intelligent Transportation Systems, including ride-hailing, urban road planning, and vehicle f…

13:00 JSTエージェント研究/論文

線形契約の最適なサンプルの複雑さ

この論文では、オフライン設定のデータから最適な線形契約を学習するという問題を解決します。この場合、エージェント タイプは未知の分布から抽出され、プリンシパルの目標は期待される効用を最大化する契約を設計することです。具体的には、私たちの分析は、単純な経験的効用最大化 (EUM) アルゴリズムが、$O(\ln(1/\delta) / \varepsilon^2)$ サンプルのみを使用して、少なくとも $1-\delta$ の確率で最適な線形契約の $\varepsilon$ 近似を生成することを示しています。この結果は、以前に知られていた限界を改善し、D\"utting et al. 2025 の下限と一定係数まで一致し、それによってその最適性が証明されました。さらに、私たちの結果は一様収束のより強力な保証を確立します。すべての線形契約の経験的効用は、同じ最適 $O(\ln(1/\delta) / を使用した、少なくとも $1-\delta$ の確率での真の期待値の $\varepsilon$ 近似です。 \varepsilon^2)$ サンプルの複雑さ。

原文 (English)

The Optimal Sample Complexity of Linear Contracts

In this paper, we settle the problem of learning optimal linear contracts from data in the offline setting, where agent types are drawn from an unknown distribution and the principal's goal is to design a contract that maximizes her expected utility. Specifically, our analysis shows that the simple Empirical Utility Maximization (EUM) algorithm yields an $\varepsilon$-approximation of the optimal linear contract with probability at least $1-\delta$, using just $O(\ln(1/\delta) / \varepsilon^2)$ samples. This result improves upon previously known bounds and matches a lower bound from D\"utting et al. 2025 up to constant factors, thereby proving its optimality. Furthermore, our result establishes the stronger guarantee of uniform convergence: the empirical utility of every linear contract is an $\varepsilon$-approximation of its true expectation with probability at least $1-\delta$, using the same optimal $O(\ln(1/\delta) / \varepsilon^2)$ sample complexity.

13:00 JST研究/論文

TextBridgeGNN: Pre-training Graph Neural Network for Cross-Domain Recommendation via Text-Guided Transfer

Graph-based recommendation has achieved great success in recent years. The classical graph recommendation model utilizes ID embedding to st…

13:00 JSTLLM/生成AIQwen

Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field test…

13:00 JST画像/動画生成

Infinite-Precision Autoregressive Modeling for Vector Graphics and Layouts

While Transformer-based autoregressive models excel in data generation, their token discretization strategy inherently limits their precisi…

13:00 JSTLLM/生成AI画像/動画生成

A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning

Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably…

13:00 JST研究/論文

Ordering-based Causal Discovery via Generalized Score Matching

Learning DAG structures from purely observational data remains a long-standing challenge across scientific domains. An emerging line of res…

13:00 JST画像/動画生成

MANGO: A Global Single-Date Paired Dataset for Mangrove Segmentation

Mangroves are critical for climate-change mitigation, requiring reliable monitoring for effective conservation. While deep learning has eme…

13:00 JST研究/論文

Physics-Encoded Inverse Modeling for Arctic Snow Depth Estimation

Accurate estimation of unobserved quantities in time-varying inverse problems remains challenging when observations are sparse and only ind…

13:00 JSTLLM/生成AI

Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis

Large language models fine-tuned via a two-stage pipeline (domain adaptation followed by instruction alignment) can exhibit non-trivial int…

13:00 JST研究/論文

Action-Sufficient Goal Representations

In offline goal-conditioned reinforcement learning (GCRL), hierarchical approaches decompose long-horizon tasks into high-level subgoal pre…

13:00 JST研究/論文

Plain Transformers are Surprisingly Powerful Link Predictors

Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. Whi…

13:00 JSTLLM/生成AI

Towards Isolated Interventions via Almost Orthogonal Features in Language Models

A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in acti…

13:00 JSTLLM/生成AI

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deploym…

13:00 JSTLLM/生成AI

How College Students Use AI to Navigate Course Readings: Evidence from an Eight-Week Study

College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions sha…

13:00 JSTLLM/生成AI

LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization

Symbolic regression aims to distill mathematical equations from observational data. Recent approaches have successfully leveraged Large Lan…

13:00 JST研究/論文

The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly w…

13:00 JST研究/論文

AIFL: A Global Daily Streamflow Forecasting Model Using a Deterministic LSTM Pre-trained on ERA5-Land and Fine-tuned on IFS

Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suff…

13:00 JST研究/論文

Voice-Driven Semantic Perception for UAV-Assisted Emergency Networks

Unmanned Aerial Vehicle (UAV)-assisted networks are increasingly foreseen as a promising approach for emergency response, providing rapid,…

13:00 JST研究/論文

Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting

Learning time series foundation models has been shown to be a promising approach for zero-shot time series forecasting across diverse time…

13:00 JSTLLM/生成AI

From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications

Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates…

13:00 JST画像/動画生成

UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation

LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces…

13:00 JST研究/論文

Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers

We present an Ising-compatible framework for formal neural-network robustness verification under bounded input perturbations. For piecewise…

13:00 JSTLLM/生成AI

SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems

Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an…

13:00 JST画像/動画生成GPT / ChatGPT

DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces

Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by…

13:00 JST画像/動画生成エージェント

Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning

Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars. These techniques p…

13:00 JSTLLM/生成AI

Designing Service Systems from Textual Evidence

Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing pol…

13:00 JST研究/論文

A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints

While deep learning (DL)-based methods have achieved remarkable success in continuous wireless resource allocation, efficient solutions for…

13:00 JST研究/論文

Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions

AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trus…

13:00 JSTエージェント

Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure

When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effect…

13:00 JST画像/動画生成エージェントロボティクス

AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models

Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing data-driven simulators learn agen…

13:00 JSTLLM/生成AI

Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments

Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-wor…

13:00 JSTエージェント

SkillSieve: 悪意のある AI エージェント スキルを検出するための階層型トリアージ フレームワーク

OpenClaw の ClawHub マーケットプレイスには、コミュニティが提供した何万ものエージェント スキル (2026 年 4 月 4 日のスナップショットでは 49,592) がホストされており、最近の監査では 13 ~ 26% にセキュリティの脆弱性が含まれていると報告されています。 Regex スキャナーは難読化されたペイロードを見逃します。正式な静的アナライザーは、プロンプト インジェクションとソーシャル エンジニアリングを隠す自然言語 SKILL.md 命令を読み取ることができません。どちらのアプローチも両方のモダリティをカバーするものではありません。 SkillSieve は、必要な場合にのみより深い分析を適用する 3 層の検出フレームワークです。レイヤ 1 は、リコール調整されたヒューリスティック スコアラーを通じて正規表現、AST、メタデータ チェックを実行し、ボリュームの 86% をフィルタリングします。レイヤ 2 は疑わしいスキルを LLM にルーティングし、分析を構造化された出力を持つ 4 つの並列サブタスクに分割します。レイヤ 3 は、リスクの高いスキルを 3 つの LLM の陪審に提出します。陪審は独立して投票し、意見が異なる場合は議論します。私たちは、49,592 の実際の ClawHub スキルと 5 つの回避手法にわたる敵対的サンプルを評価し、440 米ドルの ARM シングルボード コンピューターでパイプラインを実行します。 390 スキルのラベル付きベンチマークでは、SkillSieve はスキルあたり 0.006 米ドルで F1 = 0.920 (精度 0.912、再現率 0.929) を達成します。オプションの XGBoost 高速パスは、フル パイプライン リコール (0.929) を維持しながら、F1 を 1.6 ポイント削減してレイヤー 2/3 LLM コールを 32% 削減します。エコシステム間の一般化のために、フレームワークを Feishu/Lark に適応させ、52 個の実際のパッケージをスキャンします。そこでレイヤー 2 がドメイン固有のイディオムによるレイヤー 1 の誤検知を修正し、同様のエンタープライズ プラットフォームへの低コストの適応パスを提案します。リアルタイムのスキル審査のために、Feishu チャット ボットとして SkillSieve を導入します。コード、データ、ベンチマークはオープンソースです。

原文 (English)

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills

Agent skills combine natural-language instructions with executable code while inheriting an agent's filesystem, credential, and network access. Attacks can span prose and files, whereas regex and code-only analyzers cover only one modality. SkillSieve applies three progressively deeper layers: recall-oriented regex, AST, and metadata triage; four parallel LLM security sub-tasks; and an independent three-model jury with debate on disagreement. We evaluate 49,592 real ClawHub skills, a 390-skill labeled benchmark, and 100 adversarial samples across five evasion techniques on a 440 USD ARM board. The full pipeline achieves F1 = 0.929 (precision 0.912, recall 0.945) at an average cost of $0.006 per skill. An optional XGBoost fast path reduces Layer-2/3 calls by 32% with a 1.7-point F1 decrease while preserving recall. On 52 Feishu/Lark packages, Layer 2 reclassifies 13 of 14 Layer-1 flags as safe after contextual analysis; we also deploy the system as a Feishu chat bot. Code, labels, and aggregate results are open-sourced.

13:00 JST研究/論文

Sheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site

Cross-modal representations vary in how easily they can be aligned, and compatibility is generally non-transitive: two modalities may align…

13:00 JST画像/動画生成

Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex

Low-light image enhancement (LLIE) aims to restore natural visibility, color fidelity, and structural detail under severe illumination degr…

13:00 JST研究/論文

How Transformers Learn to Plan via Multi-Token Prediction

While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global struct…

13:00 JSTエージェントGoogle

Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus

Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a…

13:00 JST画像/動画生成研究/論文

Adaptive receptive field-based spatial-frequency feature reconstruction network for fine-grained few-shot image classification

Feature reconstruction techniques are widely applied for few-shot fine-grained image classification (FSFGIC). Our research indicates that o…

13:00 JSTビジネス/資金調達

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established prac…

13:00 JST研究/論文

Are Flat Minima an Illusion?

Flat minima are an account of why deep networks generalise. However flatness is a matter of form (parameters), while generalisation is of f…

13:00 JSTLLM/生成AI

Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory

Recall presents a difficult choice: transformers have a linearly growing memory that slows each successive token, while linear RNNs typical…

13:00 JST画像/動画生成

A Cascaded Edge-Cloud Architecture for Automated Diabetic Retinopathy Screening

Diabetic Retinopathy (DR) is one of the leading causes of preventable blindness, and automated screening can help extend specialist capacit…

13:00 JST画像/動画生成

ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing

Remote sensing change detection (RSCD) localises changes between two images of the same geographic region. Most state-of-the-art methods ar…

13:00 JSTエージェント

Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis

Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across…

13:00 JST研究/論文

良い方程式が悪いスコアになる場合: パラメーターの最適化を改善することでシンボリック回帰を改善する

シンボリック回帰 (SR) は、観察データから数式を抽出することにより、科学的知識の発見において中心的な役割を果たします。既存の SR メソッドのほとんどは、2 レベルの最適化フレームワーク内で機能します。つまり、離散方程式構造を検索する外側のループと、その構造の連続パラメーターを最適化する内側のループです。重要なのは、パラメーターの適合品質が構造のスコアを直接決定し、したがって外部ループの検索を決定することです。ただし、非線形演算子により内部ループが高度に非凸になり、高速ローカル ソルバー (BFGS など) への予算主導の依存により、正しい構造に対する不十分な極小値や過小評価されたスコアが得られることがよくあります。この「良い構造、悪いスコア」現象が重要なボトルネックとなり、効率を低下させ、検索を真の方程式から誤って導きます。これを解決するために、シンボリック式のデュアル ネイティブ事前分布を利用する SR ネイティブ フィッティング フレームワークである SAGE-Fit (Structure-Aware and Semantics-Guided Evaluator for Symbolic Regression) を提案します。 SR に特有の構造的および意味論的な事前確率を利用することで、各プロパティに合わせたモジュールを設計し、それによってこの最適化のボトルネックを効果的に軽減します。広範な実験により、プラグ アンド プレイ モジュールとしての当社のアプローチが評価の忠実度を大幅に向上させ、さまざまな SR システムのパフォーマンスを普遍的に向上させることが実証されました。

原文 (English)

When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization

Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational data. Most existing SR methods function within a bi-level optimization framework: an outer loop that searches for the discrete equation structure, and an inner loop that optimizes the continuous parameters of that structure. Crucially, parameter-fitting quality directly determines a structure's score and thus the outer-loop search. However, nonlinear operators make the inner loop highly non-convex, and budget-driven reliance on fast local solvers (e.g., BFGS) often yields poor local minima and underestimated scores for correct structures. This ``Good Structure, Bad Score'' phenomenon becomes a key bottleneck, degrading efficiency and misguiding the search away from the true equation. To resolve this, we propose SAGE-Fit (Structure-Aware and Semantics-Guided Evaluator for Symbolic Regression), an SR-native fitting framework that exploits the dual native priors of symbolic expressions. By capitalizing on the structural and semantic priors unique to SR, we design tailored modules for each property, thereby effectively mitigating this optimization bottleneck. Extensive experiments demonstrate that our approach, as a plug-and-play module, significantly enhances evaluation fidelity and universally improves the performance of various SR systems.

13:00 JST画像/動画生成

医療画像解析のためのタスク整合型自己教師あり学習: 体系的なレビューと実践的な設計ガイドライン

自己教師あり学習 (SSL) は、ラベルのないデータから表現を学習することで、医療画像処理におけるアノテーションのボトルネックに対処するための有望なパラダイムとして浮上しています。ただし、その有効性は口実タスクの設計と下流の臨床目的との整合性に大きく依存します。医療画像処理における SSL の体系的でタスク指向のレビューを紹介し、さまざまな口実タスクの定式化が分類、セグメンテーション、検出、その他のタスク全体のパフォーマンスにどのような影響を与えるかを検証します。 PRISMA ガイドラインに従って、2017 年から 2025 年の間に発表された 75 件の研究を分析し、対照学習、非対照学習と予測学習、生成学習と再構成ベースの学習、およびハイブリッド学習の 4 つのパラダイムに整理しました。アーキテクチャごとにメソッドをカタログ化するのではなく、各パラダイムを、それが最もよくサポートする下流の目的にマッピングします。私たちの分析によれば、普遍的に最適な SSL 戦略は存在しません。代わりに、パフォーマンスは、口実タスク、イメージングモダリティ、およびターゲットタスク間の調整によって決まります。対照的な方法は全体的な識別特徴を学習し、分類とうまく一致しますが、微妙な病理学的パターンを見落とす可能性があります。生成および空間予測ベースのアプローチは、局所的な解剖学的構造をより適切に保存するため、セグメンテーションやその他の緻密な予測タスクにより適していますが、ハイブリッド手法は最もバランスの取れたパフォーマンスを提供します。さらに、モダリティ固有の設計が重要であること、および SSL が低ラベルおよび少数ショットの領域で最大の利点を提供することを示します。最後に、これらの発見を実用的な設計ガイドラインに絞り込み、病理学を意識した口実タスク設計、高次元データのリソース効率の高いトレーニング、標準化された評価プロトコルなどの未解決の課題を概説します。この研究は、医療画像処理において、より効果的で臨床的に関連性のある SSL フレームワークを設計するための実践的なガイダンスを提供します。

原文 (English)

Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines

Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data. However, SSL performance depends not only on model architecture but also on whether the self-supervised objective preserves the information required by the downstream clinical task. This review presents a task-oriented synthesis of SSL methods for medical imaging, focusing on how the design of the self-supervised objective interacts with imaging modality, label availability, and downstream performance. We analyze $78$ studies published from 2017 to 2025 and organize them into four paradigms: contrastive, non-contrastive and predictive, generative and reconstruction-based, and hybrid learning. Rather than cataloging methods chronologically, we examine how these paradigms support classification, segmentation, detection, reconstruction, and regression. The evidence suggests that effectiveness is governed by the match among objective, modality, and downstream task rather than by any single strategy. Contrastive objectives favor global discriminative representations suited to classification but may underrepresent localized pathology, whereas spatial-prediction, masked-modeling, and reconstruction objectives better preserve anatomical structure for segmentation and dense prediction. Critically, misaligned objectives can cause negative transfer through shortcut learning on acquisition signatures or augmentation that erases diagnostic signal rather than merely weaker gains. SSL is most beneficial in low-label regimes, but its effectiveness depends on modality-aware augmentation, pathology-preserving corruption, and clinically meaningful evaluation. We conclude with practical design guidelines and open challenges for clinically aligned SSL.

13:00 JSTLLM/生成AIClaude

KBF: 言語モデルおよびブラックボックス API 監査のフィンガープリントとしての知識境界

リレー API とリセラー API は大規模言語モデル (LLM) へのアクセスを仲介することが増えていますが、ユーザーには、要求されたエンドポイントが実際にアドバタイズされたモデルを提供していることを確認する直接的な方法がありません。知識境界付近の安定した数値再現を使用してモデル API のフィンガープリントを行う、低コストのブラックボックス監査プロトコルである KBF を紹介します。 16 の実稼働 LLM エンドポイントにわたって、KBF は同じモデルの制御を拒否することなく、経済的に関連する 155 の代替すべてにフラグを立て、展開の変動に対して安定性を維持し、トラフィックの 5 ~ 10% のみが代替される場合でも高分離混合ルーティング攻撃を検出し、6 プラットフォームのシャドウ API 監査で 27 のプラットフォーム モデル セルのうち 7 が参照エンドポイントと統計的に矛盾しており、矛盾がプレミアム クロードに集中していることを発見しました。エンドポイント。

原文 (English)

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. Across 16 production LLM endpoints, KBF flags all 155 economically relevant substitutions without rejecting any same-model controls, remains stable under deployment variation, detects high-separation mixed-routing attacks when only 5-10% of traffic is substituted, and finds that 7 of 27 platform model cells in a six-platform shadow API audit are statistically inconsistent with their reference endpoints, with inconsistencies concentrated on premium Claude endpoints.

13:00 JST研究/論文

Causal Density Functions

We study the full density ratio between a specified intervention regime $P_a$ and an observational regime $P_0$, $\rho_a=dP_a/dP_0$, under…

13:00 JST画像/動画生成

医療画像セグメンテーション用の軽量ボックス予測子による MedSAM の強化

医療画像におけるセマンティック セグメンテーションは、データ不足とモダリティ間のばらつきの高さのため、重要ではありますが、困難なタスクです。 Segment Anything Model (SAM) のような基礎モデルは有望ですが、特別な適応がなければ医療画像に苦労することがよくあります。さらに、ポイント プロンプトは、ユーザー インタラクションの最も自然な形式であるにもかかわらず、特にターゲット構造が不規則であるかコントラストが不十分な場合、信頼性の高いセグメンテーションを実現するには空間コンテキストが不十分です。この論文では、軽量の Box Predictor モジュールを MedSAM アーキテクチャに統合する強化されたセグメンテーション フレームワークを提案します。 Box Predictor は、ローカライズされた画像埋め込み機能を使用して、ユーザーの 1 回のクリックからおおよその境界ボックスを推定し、ポイント プロンプトの曖昧さを軽減する空間ガイダンスを提供すると同時に、追加パラメーターは 160 万個のみで、推論オーバーヘッドは無視できます。 Box Predictor が MedSAM に統合される前に個別にトレーニングされる 2 段階のトレーニング パイプラインを導入します。私たちの方法の一般化機能を検証するために、CT、MRI、超音波を含む異なる画像モダリティにわたる 4 つの多様なデータセット (FLARE22、BRISC、BUSI、LungSegDB) に対して広範な評価を実施します。私たちの方法は、さまざまな解剖学的構造と画像化ドメインにわたってセグメンテーションの精度と堅牢性を向上させ、Dice スコア 0.89 (BUSI)、0.93 (FLARE22)、0.88 (BRISC)、および 0.98 (LungSegDB) を達成しました。コードは https://github.com/Amirhosseinmovahedi/MedSAM-BoxPredictor で入手できます。

原文 (English)

Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation

Semantic segmentation in medical imaging is a critical yet challenging task due to data scarcity and high variability across modalities. While foundation models like the Segment Anything Model (SAM) show promise, they often struggle with medical images without specific adaptation. Moreover, point prompts, despite being the most natural form of user interaction, provide insufficient spatial context for reliable segmentation, particularly when target structures are irregular or poorly contrasted. In this paper, we propose an enhanced segmentation framework that integrates a lightweight Box Predictor module into the MedSAM architecture. The Box Predictor estimates an approximate bounding box from a single user click using localized image embedding features, providing spatial guidance that reduces the ambiguity of point prompts, while introducing only 1.6M additional parameters and negligible inference overhead. We introduce a two-stage training pipeline where the Box Predictor is trained independently before being integrated into MedSAM. To validate the generalization capability of our method, we conduct extensive evaluations on four diverse datasets (FLARE22, BRISC, BUSI, LungSegDB) spanning distinct imaging modalities, including CT, MRI, and Ultrasound. Our method improves segmentation accuracy and robustness across varied anatomical structures and imaging domains, achieving Dice scores of 0.89 (BUSI), 0.93 (FLARE22), 0.88 (BRISC), and 0.98 (LungSegDB). Code is available at https://github.com/Amirhosseinmovahedi/MedSAM-BoxPredictor

13:00 JSTエージェントClaude

CollabSkill: 現実世界のタスクにおける人間とエージェントのコラボレーションの評価

AI エージェントはワークスペースを再構築しており、人間の働き方に劇的な変化をもたらしています。人間の主体性の維持と経済的価値の創出の両方において、人間とエージェントのコラボレーションには大きな可能性があるにもかかわらず、このパラダイムは、実際の人間のデータを収集し、人間間の変動を考慮することの難しさによって、職業タスクの評価にはほとんど反映されていないままです。現実世界の職業上のタスクにおける人間とエージェントのコラボレーションを評価するためのフレームワークである CollabSkill を紹介します。 CollabSkill は、実際の人間の労働者と AI エージェントを組み合わせて、彼らの職業的背景に合わせたタスクを実行し、経済的に価値のあるタスクの複雑さと実際の労働者の使用パターンを捕捉するデータを収集します。人間間のばらつきを考慮するために、CollabSkill はベイジアン スキル評価システムを採用して、人間と AI エージェントの両方のスキルの貢献を解きほぐし、定量化します。 93 人の作業者による 386 の作業セッションからの 1,500 以上のプロンプトを利用した当社の分析により、2 つの面で洞察が得られます。エージェント側では、CollabSkill のランキングが、Codex がリードし、Claude Code が 1 位となっている既存の完全自律型ベンチマークのランキングから大幅に乖離しています。人間の側では、CollabSkill は、実践的な経験がコラボレーション スキルの主な原動力として現れ、実践的なコラボレーションにより従業員の AI リテラシーを有意義に変化させることを明らかにしました。私たちは、CollabSkill によって、コミュニティが人間とエージェントのコラボレーションの体系的な評価に投資できるようになり、人間の労働者を真に強化する AI エージェントの構築を目的とした開発努力が促進されることを願っています。

原文 (English)

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration both in preserving human agency and generating economic value, this paradigm remains largely absent from occupational task evaluation, hindered by the difficulty of gathering real human data and accounting for inter-human variability. We introduce CollabSkill, a framework for evaluating human-agent collaboration on real-world occupational tasks. CollabSkill pairs real human workers with AI agents on tasks matched to their occupational background, collecting data that capture the complexity of economically valuable tasks and the usage patterns of real workers. To account for inter-human variability, CollabSkill employs a Bayesian skill rating system to disentangle and quantify the skill contributions of both humans and AI agents. Drawing on over 1,500 prompts from 386 working sessions contributed by 93 human workers, our analysis yields insights on two fronts: on the agent side, rankings on CollabSkill diverge meaningfully from those of existing fully autonomous benchmarks where Codex leads, with Claude Code ranking first; on the human side, CollabSkill reveals that practical experience emerges as the primary driver of collaboration skill, with hands-on collaboration meaningfully shifting workers' AI literacy. Together, we hope CollabSkill enables the community to invest in systematic evaluation of human-agent collaboration and spurs development efforts aimed at building AI agents that genuinely augment human workers.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文ClaudeGPT / ChatGPTGeminiDeepSeek

$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems

As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. C…

13:00 JST規制/政策

Market Design for AI: Beyond the Copyright Binary

How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves…

13:00 JSTLLM/生成AIエージェント研究/論文

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted we…

13:00 JSTLLM/生成AI

量子演算子と大規模な言語モデルの調整

大規模言語モデル (LLM) は量子演算子を理解して推論できますか? LLM は、数学と記号推論における顕著な能力にもかかわらず、本質的にユニタリ行列などの量子表現に対して盲目なままです。この研究では、ユニタリ演算子を LLM の潜在空間にマッピングするアプローチを導入することで、このギャップを埋めるための一歩を踏み出し、量子入力と言語入力に対する統一モデリングを可能にします。私たちは、パウ​​リ回転ゲート セットを介した Clifford+T 回路合成でこのアイデアをインスタンス化しました。このモデルでは、最先端の手法と競合する結果が得られ、飽和の兆候もなく、トレーニング データと一貫してスケールされています。私たちのアプローチはさらに、言語条件付き合成を可能にし、トレーニング中には見ら​​れないゲート制約を自然言語で直接指定できるようにします。この研究は、量子演算をネイティブに解釈して推論できる量子対応基盤モデルへの道を示唆しており、これは量子コンパイルとアルゴリズム発見に及ぶ広範な影響を与える可能性があります。

原文 (English)

Aligning Quantum Operators with Large Language Models

Can Large Language Models (LLMs) understand and reason about quantum operators? Despite their remarkable capabilities in mathematics and symbolic reasoning, LLMs remain inherently blind to quantum representations such as unitary matrices. In this work, we take a step toward bridging this gap by introducing an approach that maps unitary operators into the latent space of an LLM, enabling unified modeling over quantum and linguistic inputs. We instantiate this idea on Clifford+T circuit synthesis over a Pauli rotation gate set, where our model achieves results competitive with state-of-the-art methods and scales consistently with training data, with no signs of saturation. Our approach further enables language-conditioned synthesis, allowing gate constraints unseen during training to be specified directly in natural language. This work suggests a path toward quantum--aware foundation models that can natively interpret and reason about quantum operations, which could have broader implications reaching across quantum compilation and algorithm discovery.

13:00 JSTLLM/生成AIハードウェア/半導体

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present…

13:00 JST画像/動画生成

Enhancing Pathological VLMs with Cross-scale Reasoning

Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnifi…

13:00 JST研究/論文

SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface

Hybrid brain-computer interfaces (BCIs) that integrate motor imagery (MI) and steady-state visual evoked potentials (SSVEP) provide high-di…

13:00 JSTLLM/生成AI画像/動画生成研究/論文OpenAIGPT / ChatGPT

TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation mo…

13:00 JST研究/論文

ライブラケット幾何学による潜在的な混乱した因果の発見

Kan-Do-Calculus (KDC) に関する最近の研究では、因果推論における受動的な観察と能動的な介入の間の境界は圏論的な二重付属関係であり、介入は左 Kan 拡張によってモデル化され、条件付けは右 Kan 拡張によって行われることが確立されました。この論文では、KDC の情報幾何学的およびカテゴリカルな結果に基づいて、潜在交絡の下で 2 つの因果発見アルゴリズムを紹介します。滑らかな統計設定では、観察的測定と介入的測定の間のラドン・ニコジム導関数は、局所的な因果ベクトル場を誘発します。これらのフィールドがリー括弧の下で閉じられないことは、計算可能なフロベニウス残差となり、これは、目に見える可分性の失敗と潜在的またはモデル化されていない構造の可能性の証拠として解釈されます。私たちの最初のアルゴリズムである BRIDGE (介入発見と幾何学的推定のためのブラケット残差) は、介入密度またはラドンニコジム比エンジンを幾何学的スクリーンと組み合わせて、許容可能な矢印の高再現率ファミリーを提案し、非閉鎖可視ペアを潜在障害候補として識別し、縮小されたファミリーを下流のスコアベースまたは微分可能な発見ルーチンに渡します。 2 番目のアルゴリズムの貢献であるスペクトル Kan-Do フロー マッチング (SKFM) は、償却介入場と潜在曲率をスペクトル的に学習し、BRIDGE が指す直接的なリー空間エンドポイントを明らかにします。詳細な一連の実験は、両方のアルゴリズムが、可能な DAG の超指数空間を何桁も崩壊させながら、潜在的な交絡因子を含む因果モデルを発見できることを示しています。この論文では、介入によって引き起こされる流れの形状から潜在構造を直接推測する、因果関係発見における新しいパラダイムを紹介します。

原文 (English)

Latent Confounded Causal Discovery via Lie Bracket Geometry

We study causal discovery from observational and interventional regimes when latent variables may affect the measured system. Our first algorithm, BRIDGE (Bracket Residuals for Interventional Discovery and Geometric Estimation), combines a density-ratio or transport engine with a high-recall geometric screen and passes the retained arrows to a score-based or differentiable discovery method. The main formulation and experiments use known single-node intervention targets; in that regime the screen is designed to retain candidate directed effects, while a downstream learner determines the final graph or equivalence-class representation. Our second algorithm, Spectral Kernel Flow Matching (SKFM), amortizes the response fields, summarizes residual nonclosure by a spectral visible-footprint subspace, and applies an order-dependent graph extractor. Direct extraction succeeds on calibrated chains and selected motifs, but is unstable on harder random DAGs when the order must be learned. On ten-node nonlinear random DAGs, the more reliable hybrid role of the geometry is as a candidate generator: calibrated SKFM/Bridge fields followed by local BIC scoring achieve mean directed $F_1\simeq0.86$. Sachs protein signaling provides a real-data stress test and supports a diagnostic, not fully identified, interpretation. The contribution is therefore a practical interventional screening pipeline, explicit guarantees for screen retention and residual-footprint rank under stated assumptions, and a falsifiable account of the boundary between geometric diagnostics and causal identification.

13:00 JST研究/論文

無限微因果関係

この論文では、接線バンドルの意味論を備えたフロベニウス マルコフ圏における無限小因果関係のカテゴリカルな説明を紹介します。 IDC は、介入がコピー/破棄構造の接線変形として機能する極小レイヤーをキャプチャします。 2 つの異なるフロベニウス構造が相互作用します。(1) コピー、比較、および破棄をコード化する古典的な変数のカテゴリカル フロベニウス代数。 (2) 幾何学的なフロベニウス可積分条件、すなわち代数的なフロベニウス構造とは異なる介入分布の包含閉包。カテゴリ的因果的十分性は、これら 2 つの概念の互換性として定義されます。重要な観察は、構造因果モデルの場合、微小な因果関係は外生変数に対する決定論的メカニズムのスライスで最も自然に定式化され、目に見える確率的カーネルはプッシュフォワード後にのみ取得されるということです。介入は、フロベニウスのコピー/破棄操作を変形する接線ベクトルです。彼らのリーブラケットは、この変形が古典的な情報フロー構造を保存しているかどうかを測定します。パールの do-calculus は、介入の同一性の指針となる例として使用されます。無関係な介入の無視は単位の不変性に、行動/観察の交換はプッシュフォワードとの共積互換性に、独立性は目に見える介入分布の包括的な括弧の閉包に対応します。

原文 (English)

Infinitesimal Causality

Interventions can be varied continuously in many causal models. Differentiating a specified smooth intervention protocol produces vector fields on a statistical model, and their Lie brackets describe the noncommutativity of the corresponding local perturbations. We formulate this differential geometry of interventions on smooth statistical models and call the resulting framework infinitesimal causality (IC). Given a constant-rank distribution spanned by visible intervention fields, we define the normal Lie-bracket residual and show that its vanishing is exactly the involutivity condition in the classical Frobenius theorem. We establish the coordinate invariance of the zero-residual property and characterize its dependence on the intervention protocol, visible span, and metric. Fully observed and latent-variable examples delineate the additional structural assumptions needed to interpret bracket residuals causally. We also distinguish tangent vectors on a statistical parameter manifold from derivatives of stochastic kernels. In the finite-state linearization of a Markov category, normalization and copy compatibility yield well-typed first-order defects. Normalization is automatic for differentiable paths of stochastic kernels, whereas copy compatibility characterizes a more restrictive deterministic or comonoid-preserving perturbation. Together, the geometric and kernel-level constructions make IC a precise foundation for Lie-bracket-based causal diagnostics and identify the assumptions required to pass from local intervention geometry to causal conclusions.

13:00 JST研究/論文

EmotionAI: A Privacy-Preserving Computational Intelligence Pipeline for Speech-Emotion-Grounded Conversational Analysis

Reviewing recorded interviews for affective cues such as composure and agitation is slow and subjective, and cloud services that could auto…

13:00 JSTLLM/生成AI

Neural Machine Translation for Low-Resource Tangkhul--English

We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair. Tangkhul is a severely under-resour…

13:00 JSTLLM/生成AI

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional dist…

13:00 JST研究/論文

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF

Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alter…

13:00 JST研究/論文

From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation

Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unpr…

13:00 JSTLLM/生成AI

DiscoLoop: マルチホップ推論のための離散埋め込みと連続隠れ状態のループ

大規模な言語モデルは、中間ステップを思考連鎖 (CoT) として外部化できる場合、多くの推論タスクで優れたパフォーマンスを実現します。ただし、多くの質問では、答えを生成する前に、モデルが単一の前方パス内で複数ステップの推論を内部化する必要があります。私たちは、モデルが単一の順方向パス内で複数のパラメトリック知識を構成する必要がある代表的なタスクである 2 ホップ推論を通じてこの課題を研究します。標準の非リカレント Transformer は、深度ローカルのストレージの問題に悩まされています。つまり、前の層で学習された事実は、2 番目のホップの取得が行われる場所では利用できません。ループ トランスフォーマーは同じメモリを再利用することでこの問題を軽減しますが、それでも一般化が不完全であることがわかりました。残りのボトルネックが代表的なものであることを示します。 2 ホップ推論タスクでは、多くの場合、最初のループによって正しいブリッジ エンティティがほぼ完全にデコード可能になりますが、対応する隠れ状態はブリッジ トークンの埋め込みと十分に一致しないままになります。驚くべきことに、トレーニングを必要としない簡単な再調整介入により、一般化ギャップはほぼ埋められます。この洞察に基づいて、我々は DiscoLoop を提案します。DiscoLoop は、その繰り返しが離散埋め込みチャネルと連続隠れ状態チャネルの両方を運ぶループ アーキテクチャです。 DiscoLoop は、記号および合成言語のマルチホップ推論タスク全体で、大幅に少ないトレーニング ステップでほぼ完璧な精度を実現します。実際の事前トレーニングに適用すると、DiscoLoop はループ変換ベースラインよりも低いトレーニング損失と強力なベンチマーク パフォーマンスを達成し、混合チャネル設計が実用的な言語モデリングに移行することを示唆しています。

原文 (English)

DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning

Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). However, many questions require the model to internalize the multi-step reasoning within a single forward pass before generating the answer. We study this challenge through two-hop reasoning, a representative task where the model must compose multiple pieces of parametric knowledge within a single forward pass. Standard non-recurrent Transformers suffer from a depth-local storage problem: facts learned in earlier layers are unavailable where second-hop retrieval happens. We found that Looped Transformers mitigate this issue by reusing the same memory, but still generalize imperfectly. We show that the remaining bottleneck is representational. In the two-hop reasoning task, the first loop often makes the correct bridge entity nearly perfectly decodable, yet the corresponding hidden state remains poorly aligned with the bridge token embedding. Surprisingly, an easy training-free realignment intervention nearly closes the generalization gap. Building upon this insight, we propose DiscoLoop, a looping architecture whose recurrence carries both a discrete embedding channel and a continuous hidden-state channel. DiscoLoop achieves near-perfect accuracy with substantially fewer training steps across symbolic and synthetic-language multi-hop reasoning tasks. When applied to real-world pretraining, DiscoLoop attains lower training loss and stronger benchmark performance than looped-transformer baselines, suggesting that the mixed-channel design transfers to practical language modeling.

13:00 JSTロボティクス研究/論文NVIDIA

ワールド モデルからワールド アクション モデルへ: ロボット工学の簡潔なチュートリアル

世界モデルは、身体化されたインテリジェンスや生成シミュレーションでますます使用されていますが、その範囲はコミュニティ全体で依然として曖昧です。このチュートリアルでは、タスク関連の観測または状態の将来の進化を推定するアクション条件付き予測モデルとしてのワールド モデルの設計空間ビューを示します。私たちは既存の手法を観測空間世界モデルと状態空間世界モデルに分類し、視覚的な忠実度、空間構造、物理的解釈可能性、制御の使いやすさにおけるトレードオフを比較します。さらに、予測された未来と実行可能なロボットのアクションを結び付ける世界アクション モデルを紹介し、想像してから実行する、ビデオ機能条件付きアクション予測、ビデオとアクションの共同モデリング、およびポリシー学習のための補助ビデオ予測という 4 つの代表的なパラダイムを要約します。このチュートリアルの目標は、世界 (アクション) モデルの概念的範囲を明確にし、具体化された予測と制御のための構造化された分類法を提供することです。

原文 (English)

From World Models to World Action Models: A Concise Tutorial for Robotics

Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should have a clear understanding of what constitutes a "world", how world models and world action models are defined, and what roles they play within robotic AI systems. The tutorial also develops a unified perspective for comparing representative approaches, such as World Labs' spatial intelligence models, Yann LeCun's JEPA framework, and NVIDIA's Cosmos platform, and clarifies how these models differ in their representations, predictive capabilities, and interaction mechanisms.

13:00 JSTビジネス/資金調達GPT / ChatGPT

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…

13:00 JST研究/論文

DOSE-I: A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy -- Technical Report

In this document, we describe characteristics and technical details of the multimodal biosignal dataset DOSE-I of procedural sedation for e…

13:00 JST研究/論文

QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting

Time-series forecasting supports decisions in finance, en-ergy, transportation, public health, and industrial monitoring. Recent foundation…

13:00 JSTLLM/生成AIエージェント

Multi-Turn On-Policy Distillation with Prefix Replay

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student…

13:00 JSTLLM/生成AI

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection…

13:00 JST画像/動画生成エージェント

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deep…

13:00 JSTロボティクス

GemNav: マルチモーダル大規模言語モデルを使用した離散トークンのビジュアル ロボット ナビゲーション

大規模な事前トレーニング済みモデルに基づいて構築されたビジュアル ナビゲーション ポリシーは、これまでのところ、専用のビジュアル エンコーダー、特注のアクション ヘッド、および数千時間に及ぶクロス実施形態データセットでのトレーニングという共通のレシピに従っています。このレシピが必要かどうかを尋ねます。この論文では、言語タワーのみで低ランク適応 (LoRA) を使用し、補助ビジュアル エンコーダや連続回帰ヘッドを使用せず、フリーズしたマルチモーダル大規模言語モデル (MLLM) を短中地平線のウェイポイント ナビゲーションに適応させるビジュアル ロボット ナビゲーション ポリシーである GemNav を紹介します。ウェイポイントとカテゴリカル ナビゲーション信号は、言語モデル ヘッドによって生成された単一の離散トークン ボキャブラリーを共有し、ソフト デコードされた補助損失により、純粋なクロス エントロピー トレーニングで破棄される計量構造が回復されます。このポリシーは、競合するトレーニング セットよりもおよそ 3 桁小さい、単一の 8.7 時間のオープン コーパス上で、ゼロショットを 4 つの物理的に異なる目に見えない環境に転送し、オープン駐車場、障害物駐車場、長い屋外の化学薬品置き場、屋内倉庫をカバーする 20 の実世界のトライアルにわたって、ゴールの 0.25 ~ 0.42 m 以内で停止します。短い画像履歴に基づいて条件付けすると、オフライン メトリクスは改善されますが、ロボットには何のメリットも得られず、事前にトレーニングされた視覚機能が導入された後に追加される時間的コンテキストの上限が指摘されています。これらの結果は、凍結された MLLM の離散トークン適応により、基礎モデルのロボット ナビゲーションにデータ効率が高く、展開可能な代替手段を提供できることを示しています。

原文 (English)

GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model

Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether this recipe is necessary. In this paper, we introduce GemNav, a visual robot navigation policy that adapts a frozen Multimodal Large Language Model (MLLM) for short-to-medium horizon waypoint navigation using Low-Rank Adaptation (LoRA) on the language tower alone, with no auxiliary visual encoder and no continuous regression head. Waypoints and categorical navigation signals share a single discrete token vocabulary generated by the language-model head, and a soft-decoded auxiliary loss recovers the metric structure that pure cross-entropy training discards. On a single 8.7-hour open corpus, roughly three orders of magnitude smaller than competing training sets, the policy transfers zero-shot to four physically distinct unseen environments and stops within 0.25-0.42m of the goal across 20 real-world trials covering an open carpark, an obstacle carpark, a long outdoor chemical yard, and an indoor warehouse. Conditioning on short image histories improves offline metrics but yields no robot benefit, pointing to a ceiling on what temporal context adds once pretrained vision features are in place. These results indicate that discrete-token adaptation of frozen MLLMs can provide a data-efficient, deployable alternative for foundation model robot navigation.

13:00 JST画像/動画生成

IB-Flow: 情報ボトルネックに基づく CFG 蒸留による数ステップのテキストから画像への生成

大規模なテキストから画像への生成モデルは、前例のない視覚的パフォーマンスを達成していますが、マルチステップ反復ソルバーへの本質的な依存により、深刻な推論遅延が発生します。分類子なしガイダンス (CFG) 軌道をターゲットとした数ステップの蒸留が、一般的な二次元圧縮パラダイムとして登場しました。ただし、既存のフレームワークは、スーパーバイザーのタイムステップを無差別にサンプリングしながら、グローバルに静的なガイダンス強度を永続的に強制する粗粒度のブラインド インジェクション パラダイムによって支配されたままです。この状態に依存しない設計は、漸進的なエントロピー削減を特徴とする動的な進化プロセスとしての画像生成の本質的な性質を完全に無視します。これにより、数ステップ圧縮のパフォーマンス境界が制限されるだけでなく、深刻な CFG オーバーコンディショニング アーティファクトが発生します。これらの制限を超えるために、情報理論の理論的レンズを通して蒸留手順を再検討し、情報ボトルネック (IB) 原理によって制約される動的な相互情報ゲームとして形式的にモデル化します。具体的には、デュアルトラック適応フレームワークを通じて、従来の盲目的な仮定を解体します。注入ターゲットを決定するために、扱いにくい KL 発散制約を、ローカル ベクトル場ノルムに基づいたオーバーヘッドゼロの閉形式の解に変換する、インスタンスを意識した選択メカニズムを提案します。注入強度を調整するために、SNR に沿って動的に減衰するエントロピーを意識したスケジュールを導入し、最初の構造固定に最大の推力を適用してから、自然な多様体にスムーズに戻って微細な詳細を調整します。広範な経験的評価により、私たちのフレームワークがオーバーコンディショニングアーティファクトを根本的に根絶し、パフォーマンスの上限を打ち破り、非常に厳格な 2 ステップ構成の下で SOTA 生成忠実度を達成することが裏付けられています。

原文 (English)

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation

While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency. Few-step distillation targeting the Classifier-Free Guidance (CFG) trajectory has emerged as the prevalent dual-dimensional compression paradigm. However, existing frameworks remain subjugated by a coarse-grained blind injection paradigm that perpetually enforces a globally static guidance strength while indiscriminately sampling the supervisor timestep. This state-agnostic design completely disregards the intrinsic nature of image generation as a dynamic evolutionary process characterized by progressive entropy reduction, which not only restricts the performance boundary of few-step compression but also precipitates severe CFG over-conditioning artifacts. To transcend these limitations, we re-examine the distillation procedure through the theoretical lens of Information Theory, formally modeling it as a dynamic mutual information game constrained by the Information Bottleneck (IB) principle. Specifically, we dismantle traditional blind assumptions via a dual-track adaptive framework. To determine the injection target, we propose an instance-aware selection mechanism that transmutes the intractable KL divergence constraint into a zero-overhead closed-form solution predicated on the local vector field norm. To regulate the injection strength, we introduce an entropy-aware schedule that dynamically decays alongside the SNR, applying maximal thrust for initial structural anchoring before smoothly reverting to the natural manifold to refine micro-details. Extensive empirical evaluations corroborate that our framework fundamentally eradicates over-conditioning artifacts, shattering the performance ceiling to achieve SOTA generative fidelity under extremely stringent 2-step configurations.

13:00 JST研究/論文

Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control

Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting user…

13:00 JSTLLM/生成AIエージェント

How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study

The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks wi…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTQwenGrok

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same…

13:00 JSTLLM/生成AIハードウェア/半導体

モノカルチャーへのヒッチハイク ガイド

大規模言語モデル (LLM) は同種の出力を生成することが多く、AI コーディング アシスタントが開発者が作成するソフトウェア アーティファクトの収束につながる可能性があるという懸念が生じています。開発者はモデルの出力を対話的にプロンプ​​ト、評価、変更、拒否するため、また出力はプロンプトやリポジトリのコンテキストによって異なるため、これが実際に発生するかどうかは不明です。 2019 年から 2026 年半ばまでの Kaggle コンテストの提出物を使用してコードの均一化を調査します。私は最初に、プログラミング文化における長年の慣習を強化する LLM と一致する、ランダム シード値 42 への広範な収束について文書化しました。次に、均質化を集約と抽象化の 2 つのレベルでより広範囲に研究します。提出物レベルでは、コンテスト内の提出物の平均ペアごとの類似性を測定します。コンテスト レベルでは、提出されたコードの概念的な範囲を測定し、それぞれについて明確な尺度を動機付けます。表面構文をキャプチャする TF-IDF 表現と、コードの意図とセマンティクスをキャプチャする Voyage 3 コード埋め込みです。結果は、個人レベルと集団レベルの両方で構文の実質的な均質化を示しています。つまり、個々の提出物はリテラル構文とコード構造においてより類似している一方で、構文のバリエーションの潜在的な次元は狭くなっています。対照的に、意味論的な均質化の証拠は、個別にも集合的にもほとんど見つかりません。平均意味論的距離は基本的に横ばいのままであり、意味論的アプローチのコンテストレベルの潜在的な次元範囲は安定したままであり、それがわずかに拡大したことを示唆する証拠さえあります。これらの調査結果は、AI コーディング アシスタントが実装の詳細を確実に標準化しているものの、コーダーが採用するアプローチや問題解決戦略が均質化しているという証拠はまだ得られていないことを示唆しています。

原文 (English)

The Hitchhiker's Guide to Monoculture

Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interactively prompt, evaluate, modify, and reject model outputs, and because outputs vary with prompt and repository context. I examine code homogenization using Kaggle contest submissions from 2019 to mid-2026. I first document widespread convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding convention in programming culture. I then study homogenization more broadly, at two levels of aggregation and abstraction. At the submission level, I measure the average pairwise similarity of submissions within contests. At the contest level, I measure the conceptual span of submitted code, motivating distinct measures for each: TF-IDF representations, which capture surface syntax, and Voyage 3 code embeddings, which capture code intent and semantics. The results demonstrate substantial syntactic homogenization at both the individual and collective levels: individual submissions have become more alike in literal syntax and code structure, while the latent dimensionality of syntactic variation has narrowed. In contrast, I find little evidence of semantic homogenization, individually and collectively. Average semantic distance remains essentially flat, and the contest-level latent dimensional span of semantic approaches remains stable. These findings suggest that AI coding assistants are certainly standardizing implementation details, yet they have not yet produced evidence of homogenization in the approaches and problem-solving strategies coders employ.

13:00 JSTLLM/生成AI

オーディオ対応の大規模言語モデルからのきめ細かいフィードバックによるテキストから音声への命令の改善

最近のテキスト-オーディオ モデルは高品質のオーディオを生成しますが、多くの場合、複数のサウンド イベントや時間的順序を含む指示に従わないことがあります。このギャップは、既存の評価およびトレーニング信号が主に全体的な類似性または知覚品質を強調し、命令レベルの正確さの監視が限定されているために発生します。我々は、生成された音声におけるターゲットイベントの存在と時間的関係を検証するためのきめ細かい判定として音声認識大規模言語モデル(ALLM)を使用する命令レベルのフレームワークを提案します。ベンチマークと人による検証を通じて ALLM の判断を検証した後、そのフィードバックを使用して、直接的な好みの最適化のための好みのペアを構築します。さらに、マルチイベントの時間的命令のフォローを評価するための物語ベンチマークである S3Bench を紹介します。実験の結果、私たちの方法により、オーディオ品質を維持しながら、既存のベンチマークと S3Bench 全体でイベントの完全性、時間的順序付け、および共同命令追従の精度が向上することがわかりました。

原文 (English)

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.

13:00 JST画像/動画生成

Anatomically Faithful but Temporally Diffuse: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography

Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post-hoc attributio…

13:00 JSTエージェントビジネス/資金調達

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…

13:00 JST画像/動画生成エージェント

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its ass…

13:00 JST研究/論文

SechKAN: 双曲線セカント関数を備えたコルモゴロフ・アーノルドネットワーク

近年、コルモゴロフ・アーノルド ネットワーク (KAN) は、機械学習および科学計算タスクにおける有効性によりますます注目を集めており、ニューラル ネットワーク設計に新しいパラダイムを提供しています。この論文では、双曲割線 (sech) 関数に基づく KAN アーキテクチャである SechKAN を紹介します。双曲割線基底は、滑らかな鐘形の形状、局所的な応答、および安定した勾配のために使用されます。 1D 線形変換を採用してパラメータの数を減らし、SechKAN のモデル サイズを多層パーセプトロン (MLP) と同等に保つことができます。実験結果は、MNIST、Fashion-MNIST、CIFAR-10、CIFAR-100 などのベンチマーク データセットでの関数フィッティング、PDE 問題、および画像分類タスクにおける SechKAN の有効性を示しています。 SechKAN は、MLP や他の KAN バリアントと比較して、同様の数のパラメータを維持しながら、優れたパフォーマンスを実現します。ただし、実行時間は他の KAN 亜種よりも優れていますが、MLP よりわずかに長くなります。

原文 (English)

SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions

In recent years, Kolmogorov-Arnold Networks (KANs) have attracted increasing attention due to their effectiveness in machine learning and scientific computing, offering a new paradigm for neural network design. In this paper, we present SechKAN, a novel KAN based on hyperbolic secant (sech) functions. The hyperbolic secant basis is adopted for its smooth bell-shaped form, localized responses, and well-behaved gradients. We employ a 1D linear projection to reduce the number of parameters, allowing SechKAN to maintain a model size comparable to that of multilayer perceptrons (MLPs). Experimental results show the effectiveness of SechKAN on function fitting, PDE surrogate modeling, and image classification benchmarks, including MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100. On function fitting, SechKAN achieves performance comparable to both MLPs and representative KAN variants. On PDE surrogate modeling, it outperforms MLPs and achieves competitive or better performance than representative KAN variants. On image classification benchmarks, SechKAN achieves the best performance among the evaluated KAN variants while remaining competitive with MLPs using a comparable number of parameters. However, SechKAN still incurs higher computational cost than MLPs and some KAN variants. Our source code is publicly available at https://github.com/hoangthangta/All-KAN.

13:00 JST研究/論文

Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes?…

13:00 JSTエージェント

Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts

Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language task specificat…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through int…

13:00 JSTLLM/生成AI

SLPO: サロゲート ポリシーによる潜在推論の拡張

検証可能な報酬を伴う強化学習は、明示的な思考連鎖推論器でテスト時間のスケーリングを引き出すための主要なレシピになっています。ただし、すべての中間ステップを言語トークンとしてデコードする必要があるため、このスケーリング パスは依然として計算コストが高くなります。代わりに、潜在的推論は中間計算を連続ベクトルとして実行し、はるかに短い期間ですでに明示的な CoT に匹敵するか、それを上回っています。この約束にもかかわらず、潜在的推論者は主に模倣に縛られたままですが、明示的な CoT は結果報酬 RL によってすでに模倣を超えています。潜在的な軌道には、固定された思考予算の下で扱いやすいステップごとの尤度と適応的な停止インターフェイスが欠けているため、結果の報酬は潜在的なテスト時間のスケーリングを引き出すことができません。我々は、結果報酬 RL を自己回帰潜在推論者にもたらすために、サロゲート潜在ポリシー最適化 (SLPO) を導入します。これは、軌道レベルのクレジット割り当てのための潜在遷移に対する経験的代理ポリシー密度と、結果報酬の最適化によって変数ホライズン ポリシーに洗練される正確性監視されたストッピング ヘッドです。 SLPO は、継続的かつソフトな思考設定全体で、並列サンプリングの下で​​ Pass@$k$ を改善し、より高い決定論的精度でより長い潜在計算をより困難なインスタンスに割り当てます。

原文 (English)

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across continuous and soft thinking settings, SLPO improves Pass@$k$ under parallel sampling and allocates longer latent computation to harder instances with higher deterministic accuracy.

13:00 JST画像/動画生成

G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detec…

13:00 JSTLLM/生成AI

Generative AI floods and dilutes the market for books

Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buye…

13:00 JSTLLM/生成AIGPT / ChatGPT

PhantomFill: When the Form Demands an Answer, Language Models Invent One

Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the…

13:00 JST研究/論文

Adaptive Multi-Horizon Reinforcement Learning

Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement l…

13:00 JSTロボティクス

Emergent Compositional Skills in Mixture-of-Experts VLAs

We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of…

13:00 JSTロボティクス

Robostral Navigate

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains…

13:00 JSTLLM/生成AI

Error Certificates for KV-Cache Eviction via Randomized Design

Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot k…

13:00 JST研究/論文

Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows

Compressible multiphase flows involving shocks and material interfaces arise in applications such as bubble collapse and droplet breakup, w…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGeminiGrok

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither…