Skip to the content.

AIニュース 2026-07-15

自動生成: 2026-07-15 11:51 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. How to manage AI investments in the agentic eraOpenAI

    Learn how enterprises can manage AI investments in the agentic era by…

  2. 「バズるほど赤字だった」──野田クリスタルのAIペットカードゲーム、公開停止からの復活劇を本人に聞いたITmedia AI+

    野田クリスタルさんがGeminiで開発した「ペットカードジェネレーター」は初日に130万アクセスを集めるも、AI利用料の高騰で一時停止に。…

  3. Google DeepMindのハサビスCEO、米国主導の「フロンティアAI標準化機関」設立を提唱ITmedia AI+

    Google DeepMindのデミス・ハサビスCEOは、AGIの実現は「おそらくあと数年」だとして、米国主導でフロンティアAIモデルをリ…

  4. Dropbox、Claude連携スタート ChatGPT、Geminiに続き……「AI共通のコンテンツ基盤」にITmedia AI+

    既存の「ChatGPT」や「Gemini」との連携も拡大。主要なAIプラットフォームをまたいでDropboxをコンテンツ基盤として使えるよ…

  5. OpenAI researcher Miles Wang in talks to launch AI drug discovery startup valued at $2BTechCrunch AI

    The funding discussions point to investor interest in applying AI to…

  6. OpenAI’s new flagship model deletes files on its own, people keep warningTechCrunch AI

    A number of social media posts claim that GPT-5.6 Sol deleted files a…

  7. 富士通がNVIDIA「Rubin」対応の国産AIサーバを今秋製造へ ソブリン需要に対応ITmedia AI+

    富士通は、ソブリンAIの需要に応える国産ハイエンドAIサーバやオンプレミス向け生成AI基盤を公開した。国内工場での一貫生産が特徴だ。202…

トピック別件数

日本語メディア13件

ITmedia AI+ (日本語)

10:00 JSTLLM/生成AIGemini

「バズるほど赤字だった」──野田クリスタルのAIペットカードゲーム、公開停止からの復活劇を本人に聞いた

野田クリスタルさんがGeminiで開発した「ペットカードジェネレーター」は初日に130万アクセスを集めるも、AI利用料の高騰で一時停止に。「うちのこカード伝説」として復活するまでの舞台裏を、野田さんと電通の開発者に聞いた。

08:00 JSTエージェントGoogle

AIエージェントのコスト、どこに「消えて」いる? Google Cloud調査で浮上した“クラウド利用の盲点”

AIエージェントの本番運用段階で、多くの企業がコストの問題に直面している。本番運用段階におけるコストは一体どこに「消えている」のか。Google Cloudの調査から浮かび上がった、AIエージェントの本番運用がITインフラにもたらす負荷の正体と、コスト抑制の手掛かりとは。

07:25 JST規制/政策Google

Google DeepMindのハサビスCEO、米国主導の「フロンティアAI標準化機関」設立を提唱

Google DeepMindのデミス・ハサビスCEOは、AGIの実現は「おそらくあと数年」だとして、米国主導でフロンティアAIモデルをリリース前に審査する標準化機関の設立を提唱するエッセイを公開した。金融業界の自主規制機関FINRAをモデルとし、将来的には審査通過を米国市場で…

07:00 JSTエージェント

「RPA負債」をこれで解消 言葉で指示するだけでPC作業を自動化する最新AI

RPAでは、保守負担や属人化、画面変更による停止といった課題が付きものだ。こうした「RPA負債」を解消する手段として、AIが状況を判断して操作するAIエージェント型RPAが登場し、業務自動化の対象をデスクトップへ広げ始めている。

07:00 JSTエージェント

ユナイテッドアローズ「売れた理由が分からない」 爆速開発の独自AIでどう解消?

店舗では「売れた理由」が分かっていても、その気付きは本部まで届かない。週報では拾い切れない現場の知見をどう生かすのか。ユナイテッドアローズは、課題解決のためAIエージェントの独自開発に乗り出した。

07:00 JSTその他

「ピースサインで勤怠打刻」 Joshinが全事業所に導入した顔認証が“従業員から絶賛”のワケ

従業員の負担を減らすべく、家電量販店を展開するJoshinは全事業所の従業員向けに、顔認証クラウドサービスと顔認証端末を組み合わせた勤怠管理システムを導入した。

07:00 JSTその他

NEC森田社長が語る「脱・人月商売」の行方 組織の壁を破るAI人材育成法

NECはAI時代に、いかにして稼いでいくのか。「BluStellar」(ブルーステラ)の収益構造、防衛・安全保障や海底ケーブルといった注力領域の戦略について、森田隆之社長がグループインタビューで語った。

06:15 JSTLLM/生成AIハードウェア/半導体NVIDIA

富士通がNVIDIA「Rubin」対応の国産AIサーバを今秋製造へ ソブリン需要に対応

富士通は、ソブリンAIの需要に応える国産ハイエンドAIサーバやオンプレミス向け生成AI基盤を公開した。国内工場での一貫生産が特徴だ。2026年秋にはNVIDIAの最新GPU「Rubin」対応の新モデルの製造開始であると明かした。

05:00 JSTその他

コードなしでもベイズ統計ができる無料の神ツール「JASP」 ~ マウス操作だけでここまでできる

プログラミングなしでベイズ統計はできる? 無料ツール「JASP」を使えば、マウス操作だけでベイズ推定やベイズ検定が行えます。これまでPythonで書いてきた二項検定やt検定を、JASPで手軽に試す方法を紹介。『社会人1年生から学ぶ、やさしいデータ分析』ベイズ統計編、番外編(第5…

16:20 JSTその他

セブン-イレブン、テスラの「スーパーチャージャー」導入へ 26年度中に約10店舗

セブン-イレブンの店舗でテスラ車の急速充電が可能になる。第1号は神奈川県川崎市の店舗で、駐車場が広い店舗を中心に2026年度中に約10店へ広げる計画だ。

13:00 JSTエージェント

AIが「ホームディレクトリ全削除」 重要データ消失で相次ぐ悲劇

Dockerは、AIコーディングエージェントが開発者のmacOSのホームディレクトリを丸ごと削除した事例を解説した。問題はAIの賢さではなく、エージェントが開発者本人の権限でホストのシェルを直接実行する構造そのものにあり、「モデルの判断とシェルの実行の間に境界が存在しないことが…

12:02 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

Dropbox、Claude連携スタート ChatGPT、Geminiに続き……「AI共通のコンテンツ基盤」に

既存の「ChatGPT」や「Gemini」との連携も拡大。主要なAIプラットフォームをまたいでDropboxをコンテンツ基盤として使えるようにする。

11:54 JSTその他Microsoft

MicrosoftのナデラCEOが警告、AI利用者が支払う「二重コスト」 一つはお金、もう一つは「さらに価値あるもの」

米Microsoftのサティア・ナデラCEOが自身のブログで、企業はAIに「二重のコスト」を払っていると警告した。一つは利用料金、もう一つはAIを役立たせるために明け渡す固有の知識だという。

海外メディア17件

TechCrunch AI (英語)

09:27 JSTLLM/生成AIビジネス/資金調達研究/論文OpenAI

OpenAI researcher Miles Wang in talks to launch AI drug discovery startup valued at $2B

The funding discussions point to investor interest in applying AI to make breakthroughs in life sciences.

08:10 JSTその他

Lorde says AI glasses are ‘not sexy’

"Increasingly in our world, it gets harder and harder to know what is real," Lorde said onstage.

07:22 JSTLLM/生成AIOpenAI

OpenAI’s first hardware device is reportedly a screenless speaker that can move

The device is weirdly described as involving "mechanical elements that can move on their own" and the Bloomberg report includes the detail…

07:07 JSTLLM/生成AI規制/政策OpenAI

OpenAI pushes back on Apple trade secret lawsuit

OpenAI has issued another statement on the lawsuit, this time suggesting it lacks merit.

06:50 JSTLLM/生成AIハードウェア/半導体OpenAIGPT / ChatGPT

OpenAI’s new flagship model deletes files on its own, people keep warning

A number of social media posts claim that GPT-5.6 Sol deleted files and data without warning. OpenAI had basically disclosed the problem in…

04:42 JSTその他

Apple opens its new Siri AI to everyone with the iOS 27 public beta

If you’ve been waiting to try Apple’s revamped Siri without installing a developer beta, you now can. The company on Tuesday released the i…

04:41 JSTLLM/生成AIAnthropic

Anthropic’s newest ad is creeping people out

Anthropic has consistently attempted to depict itself as the ethical foil to other AI companies. This latest marketing stunt — which leans…

04:39 JSTビジネス/資金調達

The founder of Hinge raised $18M to build a new AI dating service, Overtone

Overtone describes itself as "a voice- and audio-forward service, enabled by AI, that provides highly curated introductions."

03:33 JST規制/政策Google

Google faces another AI training lawsuit from major publishers

Hachette, Cengage, Elsevier, and other publishers allege that Google trained its AI on copyrighted works without the necessary permissions.

02:45 JSTその他Google

DeepMind CEO calls for an independent standards body to regulate frontier AI

DeepMind CEO Demis Hassabis is proposing an AI "standards body" modeled after FINRA, to test frontier models and develop best practices for…

01:22 JSTその他

Meta’s Adam Mosseri says AI token budgets could soon be capped per engineer

Instagram head Adam Mosseri believes companies will eventually need to manage AI token spending the same way they manage payroll or other o…

01:00 JSTその他Google

Google Images gets a Pinterest-like redesign focused on discovery

Now, when users navigate to Google Images, they'll see a "For You" gallery of images tailored to their interests and browsing history.

00:17 JSTその他

New York State halts construction of all new data centers

New York has become the first state to temporarily halt approval of large data centers, as Gov. Kathy Hochul argues the AI-driven building…

23:37 JSTその他

Reflection inks $1B compute deal with Nebius

Reflection AI has signed a $1 billion deal to access Nebius' compute. Reflection was founded in 2024 and is developing open source AI techn…

23:24 JSTその他

The real AI race may no longer be at the frontier

Hugging Face CEO Clem Delangue says enterprises increasingly want open models, due to cost, accessibility, and ownership. Do frontier model…

23:06 JSTLLM/生成AIGPT / ChatGPT

Spotify expands its AI push with a ChatGPT-like music assistant

Spotify is rolling out a new AI-powered conversational feature that lets Premium subscribers chat with the app to discover music, podcasts,…

23:00 JSTその他

Superhuman’s new auto-draft feature almost makes me like AI replies

Superhuman’s latest AI email drafting feature is its most convincing yet, generating replies that often required little to no editing in ou…

公式ブログ1件

OpenAI (英語)

19:00 JSTエージェントビジネス/資金調達

How to manage AI investments in the agentic era

Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling h…

論文485件

arXiv cs.AI (英語)

13:00 JST画像/動画生成

トゥールミン議論モデルを使用した ML 予測から情報に基づく診断支援まで

構造化された解釈可能な評価を提供するために、トゥールミン議論モデルに従って画像ベースの診断をコンポーネントに分解します。このモデルは、主張、根拠、令状、修飾子、反論、裏付けで構成されます。網膜診断用の機械学習 (ML) モデルによって生成されたクレームを考えてみましょう。この主張を額面通りに受け入れるのではなく、Explainable AI (XAI) 手法を適用するか、議論ベースのアプローチを採用することもできます。私たちのフレームワークでは、画像からのバイオマーカー抽出に特化したモデルが根拠を提供します。根拠と請求を結び付ける令状は、医学的知識を備えたエージェントによって分析されます。私たちのアーキテクチャでは、この役割は MedGemma エージェントによって実行されます。適格性は、令状モデルと根拠モデルの両方の総合的な定量的評価に基づいて決定されます。最後に、MedSigLip で計算された画像類似性尺度を使用して反論が構築されます。これらすべてのコンポーネントは人間の専門家に提示され、ML によって生成された診断について、より多くの情報に基づいた重要な評価が可能になります。

原文 (English)

From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation

To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentation. This model consists of a claim, grounds, warrant, qualifier, rebuttal, and backing. Consider a claim generated by a machine learning (ML) model for retinal diagnosis. Rather than accepting this claim at face value, one could either apply explainable AI (XAI) methods or adopt an argumentation-based approach. In our framework, a model specialized in biomarker extraction from images provides the grounds. The warrant-linking the grounds to the claim - is analyzed by an agent equipped with medical knowledge; in our architecture, this role is fulfilled by a MedGemma agent. The qualifier is determined based on the overall quantitative evaluation of both the warrant and grounds models. Finally, a rebuttal is constructed using image similarity measures computed with MedSigLip. All these components are presented to the human expert, enabling a more informed and critical assessment of the ML-generated diagnosis.

13:00 JSTLLM/生成AI研究/論文

フォーマット感度指数: LLM ベンチマークにおけるトークン制御プロンプト ラッパーの堅牢性とスキーマ コンプライアンス

プロンプト ラッパーは多くの場合、形式のみが異なりますが、リーダーボードの結論をひっくり返すほどモデル スコアを変更する可能性があります。私たちはトークン制御プロトコルの下でこの差異を調査し、ラッパーの選択によって引き起こされる精度の範囲であるフォーマット感度指数 (FSI) と、応答の解析可能性の対応する範囲である解析可能性感度指数 (PSI) という 2 つの補完的なメトリクスを導入します。 7B から 72B パラメーターまでの 7 つの QA タスク、5 つのラッパー ファミリ、および 4 つの命令モデルにわたる 140,000 の OpenRouter 世代にわたって、平均 FSI はモデル間で 30 倍以上異なり、その主な原因はコンプライアンス違反であることがわかりました。固定効果回帰は、タスク、モデル、ラッパーを制御した後でも、解析可能性が依然として精度の強力な予測因子であることを示しています。私たちは、ラッパーの差異やコンプライアンスのないレポートの精度は統計的に脆弱であると主張し、ベンチマークと構造化された出力の展開の両方について実践的な推奨事項を提供します。

原文 (English)

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parseability Sensitivity Index (PSI), the corresponding range in answer parseability. Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x across models and is largely explained by compliance failures. A fixed-effects regression shows that parseability remains a strong predictor of accuracy even after controlling for task, model, and wrapper. We argue that reporting accuracy without wrapper variance and compliance is statistically fragile, and we give practical recommendations for both benchmarking and structured-output deployments.

13:00 JSTLLM/生成AIエージェント

忠実ではあるが修正ではない: マルチホップ エージェント リレーにおけるメッセージ形式の影響は層に依存します

LLM エージェントが相互に情報を渡すとき、メッセージ形式は重要ですか? 2 つの文献は意見が一致していません。フォーマット最適化の研究では、構造化されたメッセージが精度を損なうことなくコストを削減すると報告していますが、フォーマット制限の研究では、構造を強制することで生成が低下することがわかり、どちらも、ワンショット生成ではなくコピー忠実度が支配的な場合、メッセージが複数のホップを通過するときに何が起こるかを測定していません。制御されたリレー テストベッドを導入します。プログラムで生成された 12 のアトミック ファクトの概要が、6 ホップにわたって 5 つの形式 (フリー NL、精度指示 NL、JSON、トリプル、キー値) でホップバイホップで再エンコードされ、2 つのリレー機能層、認知負荷条件、およびペア フォーク エラー インジェクションにわたって、プログラムによるグラウンド トゥルースに対して固定の強力なグレーダーによってスコア付けされます。メッセージ形式の影響は層に依存していることがわかりました。 (i) 忠実なリレー命令の下では、強力なリレーはほぼロスレスであり、文書化された「電話ゲーム」の崩壊は発生しません。また、ホップごとの認知負荷を追加すると、フォーマット レベルの忠実度は変化せず (+/-1.8 ポイント以内)、生成コストが 24 ~ 53% 増加します。 (ii) 弱い (1.5B) リレーの下では、6 ホップのリコールのフォーマット全体の広がりは 8.7 倍 (2.3 ポイントから 20.5 ポイントへ) 増加します。これは、転送中にフォーマットのランキングを反転させる、2 つの相反するメカニズム (厳格なフォーマットによって支払われるエンコードの負担と、固定キー JSON スキーマに特有のドリフト耐性) によって引き起こされます。 (iii) ペアフォーク注入では、注入された間違った値は、一度存在すると、すべてのフォーマットのチェーンの 83 ~ 100% で最終ホップまで持続し、各フォーマットの真の値の保持と厳密に一致し、隣接する事実への付随的損害は検出できません。構造は、エラー訂正コードではなく、忠実なエラー位置特定チャネルを購入し、形式の選択はパイプライン内の最も弱いリレーに従う必要があります。

原文 (English)

Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

When LLM agents hand off information to one another, does the message format matter? Two literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy, while format-restriction work finds that imposing structure degrades generation -- and neither measures what happens when a message traverses multiple hops, where copy fidelity, not one-shot generation, dominates. We introduce a controlled relay testbed: briefs of twelve programmatically generated atomic facts are re-encoded hop-by-hop in five formats (free NL, precision-instructed NL, JSON, triples, key-value) over six hops, scored by a fixed strong grader against programmatic ground truth, across two relay-capability tiers, a cognitive-load condition, and a paired-fork error injection. We find that message-format effects are tier-dependent. (i) Under faithful-relay instructions a strong relay is nearly lossless -- the documented "telephone-game" collapse does not occur -- and adding per-hop cognitive load leaves format-level fidelity unchanged (within +/-1.8 points) while raising generation cost by 24-53%. (ii) Under a weak (1.5B) relay the across-format spread of six-hop recall grows by a factor of 8.7 (from 2.3 to 20.5 points), driven by two opposing mechanisms -- an encoding toll paid by the rigid formats and drift resistance specific to the fixed-key JSON schema -- that flip the format ranking in transit. (iii) In a paired-fork injection, an injected wrong value, once present, persists to the final hop in 83-100% of chains in every format, closely matching each format's retention of the true value, with no detectable collateral damage to neighboring facts. Structure buys a faithful, error-localizing channel -- not an error-correcting code -- and format choice should follow the weakest relay in the pipeline.

13:00 JST研究/論文

Voltzmann MapReduce: フォーク可能なサンドボックスのパーティション関数 Reduce

局所漸近正規性 (LAN) の下で、ワーカーがサイズ $n$ のチャンクに対して発する信頼密度は、ギブス-ボルツマン測度 $\exp\{-\beta E(\theta)\}$ であり、その逆温度はサンプル サイズ $\beta=n$ です。ガウス/線形の場合は 3 つの結果が正確で、それ以外の場合は 1 次です。つまり、互いに素なチャンクは独立したボルツマン因子を持ちます。そのため、MapReduce \emph{reduce} は、文字通り読むと、モードが精度重み付け (逆分散) プーリングである分割関数 $Z=\int\prod_k h_k\,d\theta$ になります。頻度主義的整合性はゼロ温度限界 $T=1/n\to0$ です

原文 (English)

Boltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes

To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs--Boltzmann measure $\exp\{-\beta E(\theta)\}$ whose inverse temperature is the sample size, $\beta=n$. Three consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \emph{reduce}, read literally, is a partition function $Z=\int\prod_k h_k\,d\theta$ whose mode is precision-weighted (inverse-variance) pooling; frequentist consistency is the zero-temperature limit $T=1/n\to0$

13:00 JSTLLM/生成AI

潜在 CoT 推論を動的システムとして解釈する

CODI や COCONUT などの最近の潜在推論手法は、基本的な解釈可能性の問題に直面しています。つまり、単一の透過的な推論トレースを追跡する明示的 CoT とは異なり、各ステップの隠れた空間に複数の重ね合わせられた候補トレースが維持されます。既存の機械的手法では、圧縮、ショートカット、重ね合わせが示されていますが、潜在的なステップにわたって推論がどのように進化するかは説明されていません。このギャップに対処するために、私たちは潜在トークンシーケンスを表現空間内の軌跡としてモデル化し、動的システム分析を適用して推論の進化を特徴付けます。 UMAP や DMD/PHATE などの定性的予測と並行して、ステップごとの変化、方向の一貫性、リアプノフ感度などの定量的尺度を使用して、潜在 CoT が 2 つの異なる安定性クラスを持つ構造化された非ランダムなダイナミクスを示すことを示します。 CODI は安定したアトラクターとして動作するのに対し、COCONUT は不安定な拡張システムとして動作し、SIM-CoT の監視により、基礎となるダイナミクスを変更することなく両方の動作が強化されます。このフレームワークは、潜在的な CoT 推論ダイナミクスの解釈可能性を高め、潜在的な推論のパフォーマンスを向上させるための実用的な洞察を提供します。コード 1 とプロジェクト ページ 2 はオンラインで入手できます。

原文 (English)

Interpreting Latent CoT Reasoning as Dynamical Systems

Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate traces in the hidden space at each step, unlike explicit- CoT, which follows a single transparent reasoning trace. Existing mechanistic methods show compression, shortcuts, and superposition without explaining how reasoning evolves across latent steps. To address this gap, we model latent token sequences as trajectories in representation space and apply dynamical systems analysis to characterize the evolution of reasoning. Using quantitative measures, such as step-to-step change, direction consistency, and Lyapunov sensitivity, alongside qualitative projections, such as UMAP and DMD/PHATE, we show that latent CoT exhibits structured, non-random dynamics with two distinct stability classes. CODI behaves as a stable attractor, while COCONUT behaves as an unstable expanding system, and SIM-CoT supervision tightens both behaviors without changing the underlying dynamics. This framework advances the interpretability of latent CoT reasoning dynamics and provides actionable insights for improving latent reasoning performance. Code1 and Project page2 available online.

13:00 JSTLLM/生成AIエージェント

YUKTI: 自然言語の状況から堅牢で検証可能な決定へ 不確実性型命題 IR、仮定に頑強なパレートフロンティア、後悔証明書

言語モデルは、言語化された状況を数値計画に変換し、主要なパイプライン (NL4Opt、OptiMUS、ORLM、OR-LLM-Agent) が単一の目的と点値係数にコミットして、一度解決します。実際の予算、労力、または臨床的注意を割り当てる決定の場合、その信頼は失敗モードです。客観化された数値はすべて仮定であり、推測が正確に正しい場合にのみ最適な計画は脆弱です、つまり計算の模倣です。 YUKTIは自動配合の対象を変更します。その表現は型付き命題グラフであり、その関係は形状事前分布、係数の不確実性、および来歴を伝えます。 YUKTI は、各ステージを正確な非線形ソルバー、または進化的ソルバーにルーティングします。分布パレートハンドオフによってステージを結合します。また、仮定ロバスト パレート フロンティア (ARPF) を導入し、仮定 (構造イプシロン汚染を含む) をリサンプリングして、各アクションが生き残る頻度 (rho) をスコア化します。私たちは、rho が意思決定の後悔の正確な要因であることを証明し、監査可能なトレーサビリティを追加し、ベンチマークに忠実なデータ基盤が存在しない場合にはそれを合成します (SRJANA)。私たちは 3 つの方法を検証します。制御された誤った仕様の下で、堅牢な妥協案は、単純なポイント計画と比較して、平均および最終後悔を 90% 以上削減します。規制された商業上の決定に基づいて、当社は合法的な行動範囲内で最適化し、下値をユーロで価格設定します。そして、41,188 件の意思決定からなる実際の公開データセットでは、オプティマイザの呪いを軽減しながら、サンプル外のバックテストはログに記録された現状を 34% 上回り、単純なポイント ルールを 4% 上回りました。ソルバーは標準です。私たちはベンチマーク SOTA の勝利はないと主張します。直接対決では、正しい数値を与えられた LLM と単一目的の最適化が示され、両方とも YUKTI が保持していた後悔の約 47 倍を負っています。LLM はフォーミュレーターであり、ソルバーではありません。長距離因果結合の下では、前方ハンドオフは不健全になり、後方誘導因果ポリシーにならなければならない場所が特定されます。

原文 (English)

YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate

Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-valued coefficients, then solve once. For decisions that allocate real budget, effort, or clinical attention, that confidence is the failure mode: every objectified number is an assumption, and a plan optimal only if the guesses are exactly right is fragile -- mimicry of computation. YUKTI changes the target of autoformulation. Its representation is a typed-proposition graph whose relationships carry shape priors, coefficient uncertainty, and provenance. YUKTI routes each stage to an exact, nonlinear, or evolutionary solver; couples stages by a distributional Pareto hand-off; and introduces Assumption-Robust Pareto Frontiers (ARPF), resampling assumptions (including structural epsilon-contamination) to score how often each action survives (rho). We prove a bound making rho an exact factor of decision regret, add auditable traceability, and synthesize a benchmark-faithful data foundation when none exists (SRJANA). We validate three ways: under controlled misspecification the robust compromise cuts mean and tail regret by over 90% versus a naive point plan; on a regulated commercial decision we optimize inside a lawful action space and price the downside in euros; and on a real public dataset of 41,188 decisions an out-of-sample backtest beats the logged status quo by 34% and a naive point rule by 4% while reducing the optimizer's curse. The solvers are standard; we claim no benchmark-SOTA win. A head-to-head shows an LLM given the correct numbers, and single-objective optimization, both incur about 47x the held-out regret of YUKTI -- an LLM is a formulator, not a solver. Under long-range causal coupling, the forward hand-off becomes unsound, locating where it must become a backward-induction causal policy.

13:00 JST研究/論文

GES-TSP: TSP のグラフ エッジ スパース化

巡回セールスマン問題 (TSP) の大規模なインスタンスを正確に解くには、計算コストがかかります。研究者は計算効率を向上させるためにグラフのスパース化手法をよく使用します。従来のスパース化手法は通常、固定ヒューリスティックに依存しており、インスタンス固有の構造情報を完全に活用できません。この論文では、ユークリッド TSP に対する学習ベースのスパース化アプローチであるグラフ エッジ スパース化 (GES) を提案します。幾何学的構造情報と組み合わせ最適化技術を組み込むことにより、私たちが提案する方法は、さまざまなインスタンスのスパース化グラフを適応的に生成し、グラフのサイズを大幅に削減し、解決プロセスを高速化します。実験結果は、私たちのスパース化手法が、解のギャップを最適値の 1% 以内に保ちながら、MATILDA データセット上のエッジの最大 95% を剪定できることを示しています。さらに、私たちのアプローチは、TSPLIB ベンチマークで強力な一般化機能を示します。一部の大規模なインスタンスでは、最適性ギャップは 1% 未満のままで、枝刈り率が 99% を超えています。

原文 (English)

GES-TSP: Graph Edge Sparsification for TSP

Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive. Researchers often employ graph sparsification methods to improve computational efficiency. Traditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information. In this paper, we propose Graph Edge Sparsification (GES), a learning-based sparsification approach for Euclidean TSP. By incorporating geometric structural information and combinatorial optimization technology, our proposed method adaptively generates a sparsification graph for different instances, significantly reducing the graph size and accelerating the solving process. Experimental results demonstrate that our sparsification method can prune up to 95% of edges on the MATILDA dataset, while keeping the solution gap within 1% of the optimal value. Moreover, our approach exhibits strong generalization capability on the TSPLIB benchmark.In some large-scale instances, the pruning rate exceeds 99%, while the optimality gap remains below 1%.

13:00 JSTビジネス/資金調達

検証者はカリキュラムです: ファミリー間ゲーム生成のための実行ゲート型自己蒸留

学習した審査員に対してコード ジェネレーターをポストトレーニングすると、アーティファクトを改善せずにスコアを上げるプロキシ機能を最適化できます。私たちは逆の信号、つまり決定論的で判断力のない、ゲーム性のないフィルター、つまり生成されたプロジェクトがヘッドレス エンジン (厳密起動) で正常に起動するかどうかを研究します。このゲートの下では、拒絶サンプリングの自己蒸留化合物が族外の一般化を引き起こします。 GameCraft-Bench (自然言語の概要を完全な Godot プロジェクトにマッピング) では、厳密な起動の下で蒸留された 14B モデル (Qwen3-14B+LoRA) により、4 つの未見のゲーム ファミリのクリーン ジェネレーションが候補あたり 8.8% から 42.2% に上昇し、ベストオブ K カバレッジが 3 ラウンドにわたって 18/25 から 25/25 (ゴールド天井) に上昇し、それぞれ大幅な向上を実現しました。 (p=0.0019、p<1e-4、p<1e-4)。この利益は単にデータを追加することによるものではありません。完全に一致するゴールド重複コントロールは基本モデルを下回ります (5.6% 対 8.8%、p=0.019)。一方、カウント一致分解は、ラウンド 1 から 2 へのジャンプを同等の品質 (+8.8pp) と量 (+8.5pp) のチャネルに分割します。最も直接的には、フィルターのみを交換してループを再実行すること (ローンチ ゲートの代わりに世代の 99.9% を通過させる寛大な BUILD チェック) により、ゲインが完全に消去され (基本に戻り、ローンチ ゲート ラウンドに対して p=1e-3)、オプティマイザーではなくベリファイアの精度が分離されます。 2 番目のゲーム不可能なシグナルであるヘッドレス実行グラウンディングは、ラウンド全体で単調に上昇し、一致した予算 (16 対 5) でゴールドの複製よりもはるかに多くのグラウンディングされた候補を生成し、ゲインが機能していることを確認し、ローンチではなく空です。ゲーム生成は 1 つのレッスンの検証可能なテストベッドです。検証者はカリキュラムであり、検証者が証明するものはモデルが学習するものです。

原文 (English)

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation

Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study the opposite signal: a deterministic, judge-free, ungameable filter -- whether a generated project launches cleanly under a headless engine (strict-launch). Under this gate, rejection-sampling self-distillation compounds out-of-family generalization. On GameCraft-Bench (mapping a natural-language brief to a complete Godot project), a 14B model (Qwen3-14B+LoRA) distilled under strict-launch raises clean generation on four unseen game families from 8.8% to 42.2% per-candidate and best-of-K coverage from 18/25 to 25/25 (the gold ceiling) over three rounds, each a significant gain (p=0.0019, p<1e-4, p<1e-4). The gain is not from merely adding data: an exactly-matched gold-duplication control regresses below the base model (5.6% vs. 8.8%, p=0.019), while a count-matched decomposition splits the round-1-to-2 jump into comparable quality (+8.8pp) and quantity (+8.5pp) channels. Most directly, rerunning the loop with only the filter swapped -- the lenient BUILD check, which passes 99.9% of generations, in place of the launch gate -- erases the gain entirely (back to base, p=1e-3 vs. the launch-gated round), isolating verifier precision rather than the optimizer. A second ungameable signal, headless execution grounding, rises monotonically across rounds and yields far more grounded candidates than gold-duplication at a matched budget (16 vs. 5), confirming the gains are functional, not launch-but-empty. Game generation is a verifiable testbed for one lesson: the verifier is the curriculum -- what it certifies is what the model learns.

13:00 JSTエージェントロボティクス

ルールに合わせた小型言語モデルとマルチエージェント自己修正による閉ループ制御

自律的な産業運用に向けた重要なステップは、手動による再設計を最小限に、またはまったく行わずに、自然言語の要件仕様に基づいて制御ポリシーを作成および再構成できる機能です。この設定では、AI エージェントによるポリシー生成は、生成された候補アクションを実行前にチェックできるプラント対応バリデーター (デジタル ツインなど) と組み合わせることで、信頼できるパスになります。ただし、実際の展開は推論レイテンシとコンピューティング フットプリントによって制限されます。大規模なクラウドベースのモデルは、多くの場合、エッジ閉ループで使用するには遅すぎたり、不透明であったり、データに敏感であったりします。この研究では、コンパクトな小型言語モデル (SLM) を制御推論用に再トレーニングし、バリデーターに基づく修正ループに埋め込むことができるかどうかを調査します。グループ相対ポリシー最適化 (GRPO) によって調整された Qwen2.5-1.5B モデルを、(i) アクション エージェント、(ii) シンボリック/デジタル ツイン スタイルの検証レイヤー、および (iii) 出力を有効なアクションに向けて繰り返し誘導する再プロンプト エージェントと組み合わせて使用​​します。ランダム化された熱制御シミュレーション (各 500 ステップの 30 回の実験) では、フレームワークは平均推論レイテンシー 3.84 秒で 91.5% の平均アクション調整精度 (ケース全体で 86.3% ~ 100%) を達成しました。シンボリック再マッピングの下で​​は、95% の範囲内レートを維持しており、トークンレベルの合意が減少しているにもかかわらず、堅牢な物理的規制を示しています。これらの結果は、エッジでの再構成可能な自律制御に向けた実用的なパスとして SLM+validator アーキテクチャを裏付けています。

原文 (English)

Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction

A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specifications, with minimal or no manual redesign. In this setting, policy generation by AI agents can be a credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated candidate actions before execution. However, practical deployment is constrained by inference latency and compute footprint: large cloud-based models are often too slow, opaque, or data-sensitive for edge closed-loop use. This work investigates whether a compact Small Language Model (SLM) can be retrained for control reasoning and embedded in a validator-guided correction loop. We use a Qwen2.5-1.5B model aligned via Group Relative Policy Optimization (GRPO), combined with (i) an action agent, (ii) a symbolic/digital-twin-style validation layer, and (iii) a reprompting agent that iteratively steers outputs toward valid actions. In randomized thermal-control simulations (30 experiments with 500 steps each), the framework achieves 91.5% average action-alignment accuracy (86.3%--100% across cases) at 3.84\,s mean inference latency. Under symbolic re-mapping, it maintains a 95% in-range rate, indicating robust physical regulation despite reduced token-level agreement. These results support SLM+validator architectures as a practical path toward reconfigurable autonomous control at the edge.

13:00 JST研究/論文

連続時間におけるフィードバック結合メモリ システム

フィードバック結合メモリ システム (FCMS) アーキテクチャは、4 つの抽象演算子を通じて閉ループ調整を形式化します。そのうちの 2 つ (エージェント更新演算子 $f_i$ と環境更新演算子 $\Psi$) は、元のフレームワークでは公理的に未定義のままです。これに対処するために、$f_i$ はメカニズムベースのインテリジェンス (MBI) によって定義され、エージェントは分散型の価格メカニズムと経済原則を通じてローカルに更新されます。また $\Psi$ は、環境が外部強制なしで軌道履歴を一貫して記録して応答する物理的基盤として扱われる非マルコフ フレームワークである結合メモリ グラフ プロセス (CMGP) によって定義されます。結果として得られる連続時間 FCMS インスタンス化では、計算可能なしきい値 $4\beta^2 < 2\eta\mu\gamma^2$ によって支配されるリアプノフのグローバル散逸率が達成されます。これは、離散 FCMS の安定条件 $4\eta\beta^2 < \gamma$ と CMGP の物理分岐閾値 $\alpha_c = 1/K$ の両方を一般化し、普遍的な組織化原理としてメモリの散逸がフィードバック ゲインを上回る必要があることを裏付けています。 $N=2$ エージェントを使用した数値シミュレーションと $N=10^6$ での平均場検証により、安定性の閾値と、それが違反された場合に出現する自己強化調整カスケードが確認されました。

原文 (English)

Feedback-Coupled Memory Systems in Continuous Time

The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which - the agent update operator $f_i$ and the environmental update operator $\Psi$ - are left axiomatically undefined in the original framework. To address this, $f_i$ is defined by Mechanism-Based Intelligence (MBI), where agents update locally through a decentralized price mechanism and economic principles, and $\Psi$ is defined by the Coupled Memory Graph Process (CMGP), a non-Markovian framework where the environment is treated as a physical substrate that records and responds to trajectory history coherently without external forcing. The resulting continuous-time FCMS instantiation achieves Lyapunov global dissipativity governed by the computable threshold $4\beta^2 < 2\eta\mu\gamma^2$. This generalizes both the discrete FCMS stability condition $4\eta\beta^2 < \gamma$ and CMGP's physical bifurcation threshold $\alpha_c = 1/K$, confirming that memory dissipation must outpace feedback gain as a universal organizing principle. Numerical simulation with $N=2$ agents and mean-field validation at $N=10^6$ confirm the stability threshold and the self-reinforcing coordination cascade that emerges when it is violated.

13:00 JST研究/論文

株主総会のようなパラコンシステントな部分的会合アブダクティブな拡大作戦

モーリス・パグヌッコは 1996 年の博士論文で、初の総会のような拉致拡大作戦を考案しました。彼の操作と、アブダクティブ推論の主な構成要素を強調し形式化する役割を担う分類法(アトーチャ・アリセダに触発された)を基礎として取り上げ、この論文の主な目的は、矛盾する説明仮説を矮小化やその結果として生じる不条理な認識状態なしに同化できる、新しい準一貫性のあるAGMのようなアブダクティブ拡張操作を、その公準とその推移的関係部分一致構築とともに提示することである。この論文で提示された形式的発展は、大部分において、信念修正の文脈に特に関連する特性、特に自己拡張性、つまり置換特性を満たす能力を確立する LFI (形式的不一致の論理) である、準一貫性ロジック RCbr の最近の作成によってのみ可能になりました。これは 2 つの論文のうちの 1 つ目です。ここで発表されたパラコンシステント アブダクティブ展開操作 (AGMpabd と呼ばれる新しいシステムの一部です) は、多くの興味深い機能をもたらしているにもかかわらず、否定と一貫性のパラコンシステント演算子に関連する認識論的役割を割り当てていません。 2番目の論文でのみ、別の新しいシステムであるAGMcircabdの一部である類似のパラコンシステント・アブダクティブ拡張操作がこの方向で強化されます。それにもかかわらず、私の知る限り、この論文で開発された操作は、AGM 文献の中でこの種のものとしては初めてのものです。

原文 (English)

AGM-like Paraconsistent Partial Meet Abductive Expansion Operation

In his 1996 doctoral thesis, Maurice Pagnucco created the first AGM-like abductive expansion operation. Taking his operation as a basis, as well as a taxonomy -- inspired by Atocha Aliseda -- responsible for highlighting and formalizing the main components of abductive reasoning, the main aim of this paper is to present a new paraconsistent AGM-like abductive expansion operation -- capable of assimilating contradictory explanatory hypotheses without trivialization and the consequent absurd epistemic state -- with its postulates and its transitively relational partial meet construction. To a large extent, the formal development presented in this paper was only made possible by the recent creation of the paraconsistent logic RCbr, an LFI (Logics of Formal Inconsistencies) that establishes properties especially relevant to belief revision contexts, in particular, the ability to be self-extensional -- i.e., to satisfy the replacement property. This is the first of two papers: the paraconsistent abductive expansion operation announced here -- which is part of a new system called AGMpabd -- despite bringing many interesting features, does not assign any relevant epistemic role to the paraconsistent operators of negation and consistency. Only in a second paper will an analogous paraconsistent abductive expansion operation -- which is part of another new system, AGMcircabd -- be enhanced in this direction. Nevertheless, to the best of my knowledge, the operation developed in this paper is the first of its kind in the AGM literature.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

スコアセット前のコアセット: LLM ベンチマークの評価 - 教師なしプロンプトサブセット選択

私たちは LLM ベンチマーク コアセットの選択について研究します。つまり、誘導されたモデル スコアとランキングが完全なベンチマーク スイートから得られたものに近似する複数のベンチマークからプロンプトの小さなサブセットを選択します。評価なしの教師なしベンチマーク コアセットの選択 (私たちのアプローチ) では、選択アルゴリズムはモデルの評価結果を使用せず、ベンチマーク全体のサブコレクションを生成するのではなく、複数のベンチマークにわたってプロンプトのサブセットを生成することにより、細かい粒度で動作します。私たちはサブモジュール式サブセット選択を使用し、この目的のために、決定点プロセス (DPP) ベースのアプローチ、サブモジュール式相互情報関数、施設の位置ベースの関数など、さまざまなサブモジュール式関数を開発および評価します。 5 つの異なる機能カテゴリ、18 のフロンティア LLM、および 61,000 を超えるプロンプトにまたがる 35 の異種ベンチマークからなる新しい大規模スイートでは、安価なセマンティック プロンプト エンベディングのみで動作するファシリティ ロケーション (FL) 関数が、さまざまなコアセット予算にわたって、12 の個別のスコアベースおよび多様性ベースのベースラインよりも優れた LLM スコアを維持することがわかりました。さらに、私たちが提案する目標が教師なし評価方式に限定されないことを示します。少数のベンチマーク全体のみを選択する必要があり、大量のモデル スコアが利用可能な設定では、同じ目標が MMLU および MTEB リーダーボードの最先端のベースラインと一致または上回るパフォーマンスを示し、同時に計算コストが大幅に安くなります。まとめると、私たちの結果は、サブモジュール性が一般に、ベンチマーク圧縮のための強力で信頼性の高いツールであることを示唆しています。

原文 (English)

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised benchmark coreset selection (our approach), the selection algorithm uses no model evaluation outcomes, and operates on a fine granularity by producing subsets of prompts over multiple benchmarks rather than producing a sub-collection of entire benchmarks. We use submodular subset selection, and we develop and evaluate many different submodular functions for this purpose, including determinantal point process (DPP) based approaches, submodular mutual information functions, and facility location-based functions. On a new large-scale suite of 35 heterogeneous benchmarks spanning five different capability categories, 18 frontier LLMs, and over 61K prompts, we find that the facility location (FL) function operating exclusively on inexpensive semantic prompt embeddings preserves LLM scores better than twelve separate score-based and diversity-based baselines, across a range of coreset budgets. Moreover, we show our proposed objective is not limited to the evaluation-unsupervised regime: in the setting where only a handful of whole benchmarks must be selected and a large amount of model scores are available, the same objective matches or outperforms state-of-the-art baselines on the MMLU and MTEB leaderboards, while being substantially cheaper to compute. Together, our results suggest that submodularity, in general, is a strong and reliable tool for benchmark compression.

13:00 JST画像/動画生成エージェント

シーンレベルの車線変更の意図と複数の相互作用する車両の軌道予測のための動的シーンインタラクション推論フレームワーク

先進運転支援システムや自動運転車における安全な動作計画には、周囲の交通状況がどのように進化する可能性があるかを正確に理解する必要があります。しかし、既存の車線変更予測手法の多くは単一の対象車両を中心としたままであるのに対し、マルチエージェント予測アプローチは多くの場合、将来の位置を通じてのみシーンの展開を記述し、各車両に関連する操縦に関する限られた明示的な情報を提供します。この研究では、ローカル交通シーン内のすべての関連車両の車線変更の意図と将来の軌道を予測する動的なシーン グラフ アテンション フレームワークを提案します。シーンは、車両がノードとしてモデル化され、車両の空間的および運動学的関係が明示的なエッジ フィーチャを通じてエンコードされる、時間変化するインタラクション グラフとして表されます。時間的グラフアテンション メッセージ パッシングは、進化する車両間の依存関係と事前操作の合図を捕捉し、一方、意図誘導デコーダーは、予測された各操作を対応する将来の動きにリンクします。シーンレベルの一貫性目標により、将来の複数車両の互換性がさらに促進されます。 NGSIM I-80、NGSIM US-101、および highD データセットでの実験では、競合するベースラインよりも一貫した改善が見られます。 DSiGAT は、NGSIM I-80 と US-101 でそれぞれ 90.12% と 90.97% の意図予測精度を達成し、最も強いベースラインと比較して軌道 RMSE を最大 52.94% 削減します。また、エージェント間の衝突率や関節変位誤差も低くなり、シーンレベルの予測がより一貫していることを示します。アブレーション、感度、堅牢性、および定性分析により、提案されたコンポーネントの寄与とシーンに焦点を当てた処方の有効性がさらに検証されます。

原文 (English)

A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles

Safe motion planning in advanced driver-assistance systems and autonomous vehicles requires an accurate understanding of how the surrounding traffic scene is likely to evolve. However, many existing lane-change prediction methods remain centered on a single target vehicle, while multi-agent forecasting approaches often describe scene evolution only through future positions and provide limited explicit information about the maneuver associated with each vehicle. This study proposes a dynamic scene graph attention framework that predicts the lane-change intention and future trajectory of every relevant vehicle within a local traffic scene. The scene is represented as a time-varying interaction graph in which vehicles are modeled as nodes and their spatial and kinematic relationships are encoded through explicit edge features. Temporal graph-attention message passing captures evolving inter-vehicle dependencies and pre-maneuver cues, while an intention-guided decoder links each predicted maneuver to its corresponding future motion. A scene-level consistency objective further encourages compatible multi-vehicle futures. Experiments on the NGSIM I-80, NGSIM US-101, and highD datasets demonstrate consistent improvements over competing baselines. DSiGAT achieves intention prediction accuracies of 90.12% and 90.97% on NGSIM I-80 and US-101, respectively, and reduces trajectory RMSE by up to 52.94% relative to the strongest baseline. It also produces lower inter-agent collision rates and joint displacement errors, indicating more coherent scene-level predictions. Ablation, sensitivity, robustness, and qualitative analyses further validate the contribution of the proposed components and the effectiveness of the scene-focused formulation.

13:00 JST研究/論文GPT / ChatGPT

ストラテジストの足場: ホテル空間市場へのアーキテクチャ依存の推論介入

私たちは、構造化推論介入が大規模言語モデルの戦略的経済推論を改善するかどうか、またその効果がモデルのアーキテクチャに依存するかどうかを調査します。 Hotelling の線形都市モデルを診断手段として使用し、演繹的推論とアブダクティブ推論、3 つのプロンプト フレーミング、条件ごとの 3 回の繰り返しにわたる 8 つの質問にわたって、足場なしのベースラインと 4 つの推論介入という 5 つの条件下で GPT-4.1-mini (標準的な指示従うモデル) と GPT-5-mini (推論に最適化されたモデル) を評価し、個別に判断された 720 の回答が得られました。スキャフォールディングのタイプとモデル アーキテクチャの間に統計的に有意な相互作用があることがわかりました ($t(7) = 4.79$、$p = 0.002$、$d = 1.69$)。コミットメント スキャフォールディングは標準モデルを改善しますが ($+0.21$)、推論モデルは低下します ($-0.63$)。原則に基づく分離は逆のパターンを示します ($-0.40$ 対 $+0.31$)。どちらのクロスオーバーも個別に有意であり (コミットメント: $p = 0.040$、分離: $p = 0.002$)、8 つの質問すべてにわたって 7/8 の方向性の一貫性が保たれています。敵対的ストレステストは両方のモデルに悪影響を及ぼし、推論モデルの劣化が $2.6\time$ 大きくなり ($-1.47$ 対 $-0.57$、$p = 0.038$)、ダメージはベースラインの難易度と負の相関があります ($R^2 = 0.36$、$p = 0.014$)。さらに、両方のモデルが正しい戦略を実行能力をはるかに超える速度で特定するという、永続的な宣言と手続きのギャップを文書化します。分離によって推論モデルのこのギャップは完全に埋められますが、介入は標準モデルには役立ちません。

原文 (English)

Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets

We investigate whether structured reasoning interventions improve the strategic economic reasoning of large language models, and whether their effects depend on model architecture. Using Hotelling's linear city model as a diagnostic vehicle, we evaluate GPT-4.1-mini (a standard instruction-following model) and GPT-5-mini (a reasoning-optimized model) under five conditions - an unscaffolded baseline and four reasoning interventions - across eight questions spanning deductive and abductive reasoning, three prompt framings, and three repetitions per condition, yielding 720 individually judged responses. We find a statistically significant crossover interaction between scaffolding type and model architecture ($t(7) = 4.79$, $p = 0.002$, $d = 1.69$): commitment scaffolding improves the standard model ($+0.21$) while degrading the reasoning model ($-0.63$), and principled separation shows the opposite pattern ($-0.40$ vs. $+0.31$). Both crossovers are individually significant (commitment: $p = 0.040$; separation: $p = 0.002$) and hold across all eight questions with 7/8 directional consistency. Adversarial stress-testing harms both models, with $2.6\times$ greater degradation for the reasoning model ($-1.47$ vs. $-0.57$; $p = 0.038$), and the damage correlates negatively with baseline difficulty ($R^2 = 0.36$, $p = 0.014$). We further document a persistent declarative-procedural gap in which both models identify correct strategies at rates far exceeding their ability to execute them; separation fully closes this gap for the reasoning model while no intervention helps the standard model.

13:00 JST研究/論文

AI における最小自律性の理論

最小特権とは、アイデンティティがそのタスクに厳密に必要なアクセス許可のみを保持すべきであるという原則であり、何十年にもわたってアクセス制御の基本的な要素でした。この原則は、単に権限を保持するだけでなく、ワークフローやシステムの境界を越えて権限を結合、承認、拡張できるエージェント AI システムには不十分であると私たちは主張します。我々は、適切な一般化として最小自律性を提案し、正式な理論を展開します。まず、ウルトラメトリック ツリーと格子値の機密性、完全性、および制御コンテキスト ラベルを組み合わせて、企業階層内のアクション間の構造的分離を測定する構成的爆発半径 d(a,b) を定義します。次に、有向エージェント影響グラフ G(θ) を定義します。 U から V へのアークには、指示された共有リソースの書き込みから読み取り会議、または保守的な非指示のエージェント間 (A2A) 通信会議、および外部で選択されたポリシーしきい値シータ以上の会議条件付き影響力の可能性が必要です。カタログ半径プロファイルは、シータの校正と監査をサポートします。最後に、認可構成、意思決定操作、およびクロスドメイン機能構成を検出する、グラフの到達可能性に関する共謀述語を定義します。

原文 (English)

A Theory of Least Autonomy in AI

Least privilege, the principle that an identity should hold only the permissions strictly required for its task, has been a foundational primitive of access control for decades. We argue that this principle is insufficient for agentic AI systems, which do not merely hold permissions but can combine, approve, and amplify them across workflows and system boundaries. We propose least autonomy as an appropriate generalization and develop a formal theory. First, we define a compositional blast radius d(a,b) that measures structural separation between actions in an enterprise hierarchy, combining an ultrametric tree with lattice-valued confidentiality, integrity, and control-context labels. Second, we define a directed agent influence graph G(theta). An arc from U to V requires a directed shared-resource write-to-read meeting or a conservative undirected agent-to-agent (A2A) communication meeting, and a meeting-conditioned influence potential at or above an externally selected policy threshold theta. A catalogue-radius profile supports calibration and audit of theta. Finally, we define a collusion predicate over graph reachability that detects authorization composition, decision manipulation, and cross-domain capability composition.

13:00 JST研究/論文

SupplyNetPy: 任意のサプライ チェーンと在庫ネットワークの高忠実度モデリングとシミュレーションのためのオープンソース Python ライブラリ

このペーパーでは、任意のマルチ階層構造を持つサプライ チェーン ネットワークのモデリングと離散イベント シミュレーションを行うための、オープンソースで十分に文書化された Python ライブラリである SupplyNetPy を紹介します。複数の補充ポリシー、傷みやすい在庫、ノードの中断、確率的な需要とリードタイムをサポートします。すべてのコンポーネントは継承を通じて拡張可能です。ユーザーはサプライ チェーンをノードとリンクの属性を含むグラフとして記述しますが、ライブラリはシミュレーションを処理し、ログと広範なノードおよびネットワーク レベルのパフォーマンス レポートを提供します。このペーパーでは、SupplyNetPy の動機、設計、主な機能、アーキテクチャを、詳細な検証結果 (分析ベンチマーク、商用ツール、および公開されたケーススタディに対する) とともに紹介します。 SupplyNetPy の開発の背後にある主な動機は、複雑なモデルのプログラムによる生成とシミュレーションであり、設計空間の探索、what-if 分析、トレーニング データの生成、およびサプライ チェーンのデジタル ツインを可能にします。

原文 (English)

SupplyNetPy: An Open-Source Python Library for High-Fidelity Modeling and Simulation of Arbitrary Supply Chain and Inventory Networks

This paper introduces SupplyNetPy, an open-source, well-documented Python library for modeling and discrete-event simulation of supply chain networks with arbitrary multi-echelon structures. It supports multiple replenishment policies, perishable inventory, node disruptions, and stochastic demand and lead times. All components are extensible via inheritance. Users describe a supply chain as a graph with node and link attributes, while the library handles simulation, providing logs and extensive node and network level performance reports. This paper presents the motivation, design, key features, and architecture of SupplyNetPy, along with detailed validation results (against analytical benchmarks, a commercial tool, and a published case study). A key motivation behind SupplyNetPy's development is programmatic generation and simulation of complex models, enabling design-space exploration, what-if analysis, training data generation, and supply chain digital twins.

13:00 JSTエージェント

ビットではなく信念の複製: エージェントシステムのための認識状態の複製

分散システムでは、古典的なステート マシン レプリケーション (SMR) モデルは、正しいレプリカが決定論的な遷移を実行して同一のビット単位の状態を生成することを前提としています。しかし、自律的、確率的、モデル駆動型のエージェントがインフラストラクチャを調整するエージェント分散システムの台頭により、決定論的なビット単位のレプリケーションでは不十分なシナリオが生じています。生成モデルで動作するレプリカは、分岐した推論パス、要約、トークン境界を示す可能性がありますが、意味的には同等で正しい操作上の決定に達します。これらの確率的参加者間でビット単位の一致を強制すると、実行の柔軟性が低下し、コンテキストの記憶喪失が誘発され、パフォーマンスが制限されます。このような設定では、レプリカはビットではなく信念について一致する必要があると私たちは主張します。我々は、レプリケーションの境界をデータの可視性から知識の可視性へと移行させるエージェント分散システム用の信念レプリケーション層であるエピステミックステートレプリケーション(ESR)を提案します。確定的で不変の証拠ログ (L) を確率的で進化する信念系統 (B) から分離するペア K = (L, B) として認識論的ノード状態を形式化します。実行の安全性を管理するために、検証者に制限された意味論的互換性メトリック内で最新のコミットされた操作上の意味を反映する操作を必要とするセマンティック線形化可能性と、公正な配信、単調な証拠、有界な検証者の妨害、および契約的グラフト演算子の下で予想される意味論的発散を制限する有界結果的一貫性を定義します。構造化された認識デルタを使用して派生した洞察を伝播するためのプロトコルの概要を説明し、文脈健忘を引き起こすことなく信念系統から誤った前提を取り除くための検証可能な意味論的ロールバックを形式化します。私たちは ESR のプロトタイプを作成し、述べられた仮定の下での実現可能性を示し、二次的な認知障害の減少を示す予備的なシミュレーション結果を報告します。

原文 (English)

Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems

In distributed systems, the classical State Machine Replication (SMR) model assumes that correct replicas execute deterministic transitions to yield identical bitwise states. However, the rise of agentic distributed systems -- where autonomous, stochastic, and model-driven agents orchestrate infrastructure -- presents scenarios where deterministic, bitwise replication is insufficient. Replicas operating with generative models may exhibit divergent reasoning paths, summaries, and token boundaries, yet reach semantically equivalent and correct operational decisions. Forcing bitwise agreement across these stochastic participants degrades execution flexibility, induces context amnesia, and limits performance. We argue that in such settings replicas should agree on belief, not bits. We propose Epistemic State Replication (ESR), a belief-replication layer for agentic distributed systems that shifts the replication boundary from data visibility to knowledge visibility. We formalize the epistemic node state as a pair K = (L, B) separating the deterministic, immutable evidence log (L) from the stochastic, evolving belief lineage (B). To govern execution safety, we define Semantic Linearizability, which requires operations to reflect the latest committed operational meaning within a verifier-bounded semantic compatibility metric, and Bounded Eventual Coherence, which bounds expected semantic divergence under fair delivery, monotonic evidence, bounded verifier disturbance, and a contractive graft operator. We outline protocols for propagating derived insights using structured epistemic deltas, and formalize Verifiable Semantic Rollbacks to prune faulty premises from belief lineages without inducing context amnesia. We prototype ESR and report preliminary simulation results that show feasibility under the stated assumptions and illustrate reductions in secondary cognitive faults.

13:00 JST研究/論文

農業予測タスクにおける機械学習のパフォーマンスを向上させるためのタスク条件付き合成データ生成

機械学習 (ML) アルゴリズムは、さまざまな状況にわたって農業変数を推定するために広く使用されています。ただし、トレーニング データの量と品質は ML アルゴリズムのパフォーマンスに大きな影響を与えるため、その使用は限定されたまたは不完全な参照データによって制限される可能性があります。合成データ生成 (SDG) は、元のデータの重要な特性を保持する人工的だが現実的なサンプルを作成することで、この問題に対処する実用的なアプローチを提供します。この研究は、教師と生徒の知識伝達と表形式データのコンテキスト内学習に基づいて、ベイジアン ネットワーク ジェネレーターとトランスフォーマーベースの表形式基礎モデル (TabICL) を組み合わせたタスク条件付き SDG (TCSDG) アルゴリズムを提案します。提案されたアルゴリズムは、作物収量予測と作物タイプ分類という 2 つの農業予測タスクに関して評価されました。 6 つのベンチマーク SDG アルゴリズムも、TCSDG のパフォーマンスと比較するために利用されました。 12 の調査サイトにわたって、2 つのトレーニング データ部分、4 つの乗算比、および 3 つの予測 ML アルゴリズムを使用して、元のデータを TCSDG 生成の合成データで強化することにより、作物タイプ分類実験の 89% と作物収量予測実験の 74% で ML パフォーマンスが向上しました。 TCSDG はベンチマーク SDG アルゴリズムのパフォーマンスも大幅に上回り、両方のタスクにわたって総合レベルで ML パフォーマンスを一貫して向上させる唯一の方法でした。この研究は、慎重に設計および処理された合成データが精密農業アプリケーションにおける ML パフォーマンスを向上させることができることを実証しています。 TCSDG は、下流の ML 農業予測をサポートする合成データを生成するための実用的で拡張可能なフレームワークを提供します。 TCSDG の完全な実装は、https://github.com/HamidEbrahimy/TCSDG でオープン ソースとして公開されています。

原文 (English)

Task-Conditioned Synthetic Data Generation for Improving Machine Learning Performance in Agricultural Prediction Tasks

Machine Learning (ML) algorithms have been widely used to estimate agricultural variables across diverse contexts. However, because the quantity and quality of training data strongly influence performance of ML algorithms, their use can be constrained by limited or incomplete reference data. Synthetic Data Generation (SDG) offers a practical approach to address this issue by producing artificial but realistic samples that preserve key characteristics of the original data. Building on teacher-student knowledge transfer and in-context learning for tabular data, this study proposes a Task-Conditioned SDG (TCSDG) algorithm that pairs a Bayesian Network generator with a transformer-based tabular foundation model (TabICL). The proposed algorithm was evaluated on two agricultural prediction tasks: crop yield prediction and crop type classification. Six benchmark SDG algorithms were also utilized to compare their performance with that of TCSDG. Across twelve study sites, two training-data fractions, four multiplication ratios, and three predictive ML algorithms, augmenting the original data with TCSDG-generated synthetic data improved ML performance in 89% of the crop type classification experiments and 74% of the crop yield prediction experiments. TCSDG also substantially outperformed benchmark SDG algorithms and was the only method to consistently improve ML performance across both tasks at the aggregate level. The study demonstrates that carefully designed and processed synthetic data can improve ML performance in precision-agriculture applications. TCSDG offers a practical and extensible framework for generating synthetic data that supports downstream ML agricultural prediction. The full implementation of TCSDG is publicly available as open source at https://github.com/HamidEbrahimy/TCSDG.

13:00 JST研究/論文

LegalFarePlan: 非加算運賃規則に基づく運賃透明性のある都市鉄道ルート計画のためのラベル設定フレームワーク

都市鉄道の運賃システムは非加算的な場合があります。出発地から目的地までの単一の有料旅行の運賃は、法的に分離された複数の旅行区間の運賃の合計とは異なる場合があります。この文書では、合法的な出国と再入国の操作を明示的で監査可能な制約としてモデル化する、運賃に透明なルート計画フレームワークである LegalFarePlan について説明します。交通ネットワーク、運賃機能、乗り換えルール、駅レベルの退出/再入場コスト、超過時間の予算、分割制限を考慮して、プランナーは、有料の旅行セグメントにわたって説明可能なルート計画を計算します。このアーティファクトは、ダイクストラ最短時間および直接ルート プランナー ベースライン、貪欲な分割ヒューリスティック、境界付きの正確なラベル設定、およびパレート フロンティア検索を実装します。評価には、制御された合成データと、360 の OD ペアを備えた 57 ステーションの半合成ベンチマークが使用されます。半合成ベンチマークでは、限定完全一致検索により、OD ペアの 71.11% でポジティブなモデル化された運賃割引が特定され、45 分の延長戦予算の下で、平均割引額は 3.78、最大割引額は 9.0 の合成運賃単位でした。これらの結果は、メソッドの動作と再現性を示しています。これらは MTR や交通事業者に関する経験的な結論ではありません。

原文 (English)

LegalFarePlan: A Label-Setting Framework for Fare-Transparent Urban Rail Route Planning under Non-Additive Fare Rules

Urban rail fare systems may be non-additive: the fare of a single paid journey from an origin to a destination can differ from the sum of fares over multiple legally separated journey legs. This paper presents LegalFarePlan, a fare-transparent route-planning framework that models legal exit-and-reentry operations as explicit, auditable constraints. Given a transit network, fare function, transfer rules, station-level exit/re-entry costs, an extra-time budget, and a split limit, the planner computes explainable route plans over paid journey segments. The artifact implements Dijkstra shortest-time and direct route-planner baselines, a greedy split heuristic, bounded exact label-setting, and Pareto-frontier search. Evaluation uses controlled synthetic data and a 57-station semi-synthetic benchmark with 360 OD pairs. On the semi-synthetic benchmark, bounded exact search identifies positive modeled fare reductions for 71.11% of OD pairs, with mean reduction 3.78 and maximum reduction 9.0 synthetic fare units under a 45-minute extra-time budget. These results demonstrate method behavior and reproducibility; they are not empirical conclusions about MTR or any transit operator.

13:00 JSTエージェント研究/論文

BatteryLake: 異種バッテリーの経年劣化データとベンチマークを物理学に基づいてエージェント的にキュレーション

公開されているバッテリー劣化データセットは高度な健康管理にとって重要な資産ですが、その実際の使用は、一貫性のない形式、不明確なスキーマ、リポジトリや出版物に散在するメタデータによって制限されることがよくあります。現在のキュレーションは大部分が手動のままで再現が困難ですが、汎用データ統合ツールは電気化学時系列データのドメイン固有のセマンティクスを見逃しています。私たちは、エージェント的で物理学に基づいたキュレーション フレームワークを通じて、生の公共バッテリー データをベンチマーク対応の資産に変換する管理されたデータ レイクハウスである BatteryLake を、3 つの貢献とともに紹介します。まず、LLM エージェントはメタデータを抽出してデータセット固有のコンバーターを合成し、すべての出力を逐語的な証拠に基づいて根拠付けし、値をサポートするものがない場合は棄権します。第 2 に、人間参加型メカニズムにより、検証が選択的予測として組み立てられ、26 のスキーマ、統計、および物理的妥当性ルールを通じて許可されたデータがゲートされます。 3 番目に、標準化された SOH および RUL タスク、3 つの分割プロトコル、および 8 つのベースライン モデル ファミリを含む、25 を超える機関からの 41 のデータセットのオープン ベンチマークをリリースします。プラットフォーム、ベンチマーク、およびキュレーション プロトコルは、https://tianwen1209.github.io/batterylake/ で公開されています。

原文 (English)

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent formats, unclear schemas, and metadata scattered across repositories and publications. Current curation remains largely manual and hard to reproduce, while general-purpose data integration tools miss the domain-specific semantics of electrochemical time-series data. We present BatteryLake, a governed data lakehouse that turns raw public battery data into benchmark-ready assets through an agentic, physics-grounded curation framework, with three contributions. First, LLM agents extract metadata and synthesize dataset-specific converters, grounding every output in verbatim evidence and abstaining when none supports a value. Second, a human-in-the-loop mechanism frames verification as selective prediction and gates admitted data through 26 schema, statistical, and physical-plausibility rules. Third, we release an open benchmark of 41 datasets from over 25 institutions, with standardized SOH and RUL tasks, three split protocols, and eight baseline model families. The platform, benchmark, and curation protocol are publicly available at https://tianwen1209.github.io/batterylake/.

13:00 JSTLLM/生成AIエージェント

正しさにはどれくらいの費用がかかりますか?弱いマルチエージェントの群れに強力なコレクターを予算を立てて配置する

信頼性の低い安価なエージェントの群れは、少数の強力で高価な「オラクル」修正者によって正しいコンセンサスに誘導することができます。いくら使う必要があるのか​​、おみくじをどこに置くのかを尋ねます。各オラクルがコスト結合された凹型の強度で真実に向かって 1 つのノードを固定するグラフ上のコンセンサスとして群をモデル化し、コヒーレンス H(R)=tr M(R)^{-1} によって品質を測定します。最初の結果は、オラクルの強さが異なる場合でも、H はサブモジュールのまま (追加されたオラクルは最後よりも役に立たない) であるため、貪欲な費用対効果は、どのような予算でも最適な配置の 1-1/e 以内に収まります。予算を反転すると、予算の正確性のフロンティア B*(eps)、eps の正しいコンセンサスを保証する最小支出、つまり完全なグラフ上の閉じた形式、およびオラクルのコストが同じ場合の最小オラクル数 k* が得られます。予算が少数の強力なオラクルを購入するか、コストと品質の法則の中程度の曲率を多数購入するかにかかわらず、収益の減少は急激な増加を支持します。 Qwen3 ラダー (0.6-32B) で測定すると、法則は数学的検証に対して凹、緊急コード トレースに対して凸であるため、判定は完全にタスクに依存します。https://github.com/YehudaItkin/budgeted-oracle-placemen

原文 (English)

How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm

A cheap swarm of unreliable agents can be steered to a correct consensus by a few strong, expensive "oracle" correctors. We ask how much one must spend, and where to place the oracles. We model the swarm as a consensus on a graph in which each oracle pins one node toward the truth at a cost-coupled, concave strength, and measure quality by the coherence H(R)=tr M(R)^{-1}. Our first result is that H stays submodular (each added oracle helps less than the last) even when the oracles differ in strength, so a cost-benefit greedy comes within 1-1/e of the best placement at any budget. Inverting the budget gives the budget-correctness frontier B*(eps), the least spend that guarantees an eps-correct consensus: closed-form on the complete graph, and a minimal oracle count k* when oracles cost the same. Whether a budget then buys a few strong oracles or many medium onese curvature of the cost-quality law: diminishing returns favour spreadsharply increasion. Measured onthe Qwen3 ladder (0.6-32B), the law is concave for math verificatio convex foremergent code tracing, so the verdict is genuinely task-dependent.https://github.com/YehudaItkin/budgeted-oracle-placemen

13:00 JSTLLM/生成AIエージェント

AI エージェントに対する規範の適用: マルチエージェント システムでの動作を確実に形成する

AI エージェントは、多様な目標を追求し、報酬を求めて競争する共有環境にますます導入されています。この複数のエージェントによる競争は、集団的なコストをかけて個人の利益につながる行為につながる可能性があります。たとえば、マーケティングエージェントがソーシャルメディアでのエンゲージメントを競い合った結果、誤解を招くコンテンツを投稿する可能性があります。人間社会は、違反を検出して罰する強制メカニズムによってサポートされ、許容される行動を制限する規範を通じてこのような問題に対処します。これを動機として、私たちは言語モデルエージェントの規範強制メカニズムを研究しています。私たちは、単純な強制メカニズムが、たとえ明示的に訓練されたり、そうするように促されたりしていない場合でも、競争上の優位性を得るために、調整を誤ったエージェントによって悪用されることを発見しました。したがって、私たちはより堅牢なメカニズムの設計に注目し、2 つの重要な要素を特定します。それは、時間の経過とともに各エージェントの信頼性を推定することと、繰り返される不正行為に対するペナルティを段階的に増加させてこの推定値を更新することです。 3 つのシミュレートされた環境とさまざまなエージェント集団にわたって、これらの原則に基づいて構築されたメカニズムは悪用を阻止しながら、ベースラインと同等またはそれよりも低いコストで規範違反を罰します。私たちの結果は、規範施行メカニズムをエージェントの行動を形成するためのスケーラブルな手段として位置づけていますが、それはエージェントが管理するシステムの一部になることを予期して設計されている場合に限ります。コードとデータは https://yaowenye.com/norm-enforcement で入手できます。

原文 (English)

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost -- for instance, marketing agents may post misleading content as a result of competing for engagement on social media. Human societies address such problems through norms that constrain acceptable behavior, supported by enforcement mechanisms that detect and penalize violations. Motivated by this, we study norm enforcement mechanisms for language model agents. We find that simple enforcement mechanisms are exploited by misaligned agents for competitive advantage, even when they are not explicitly trained or prompted to do so. We thus turn our attention to designing more robust mechanisms, and identify two key ingredients: estimating each agent's reliability over time, and updating this estimate with escalating penalties for repeated misbehavior. Across three simulated environments and a variety of agent populations, mechanisms built on these principles resist exploitation, while still penalizing norm violations at comparable or lower cost than baselines. Our results position norm enforcement mechanisms as scalable levers for shaping agents' behavior, but only when designed to anticipate becoming part of the system they govern. Our code and data are available at https://yaowenye.com/norm-enforcement.

13:00 JSTエージェント

有限ルール改訂による適応型エージェントティックコントローラーの検証

産業用エージェント AI システムでは、プロトタイプの機能と本番環境の展開との間にギャップがますます現れています。特に、適応エージェントは、非決定性、機密性の制約、限定されたコンテキスト、および弱い可観測性の下では検証が困難なままでありながら、妥当な出力を生成する可能性があります。この論文は、有限の記号規則、明示的な診断述語、説明ログ、およびホールドアウト再評価によって表される適応エージェント コントローラー用の有界検証プロトコルを定式化します。研究の中心的な疑問は、適応型エージェント コントローラーが有限ルール、明示的な診断述語、説明ログ、および保留された再評価を通じて表現される場合、人間による無制限の判断に依存せずに、どのクラスのコントローラー障害を検出、局所的に修復、または拒否できるかということです。提案されたフレームワークは、コントローラーを有限の変更可能なオブジェクトとして扱います。診断エラーは、ルールの追加、ルールの削除、優先順位の改訂など、事前定義されたルール レベルの編集にマッピングされます。修復されたコントローラーは、保持されているシミュレーション シードまたは複製された初期状態に基づいて評価されます。様式化された財務制約のある在庫管理ベンチマークでの実験では、1 回のルール編集では修復不可能なリソース起因の障害、しきい値またはガードレールに違反するために拒否される部分的な修復、平滑化ルールの削除によって引き起こされる注文ボラティリティの障害のローカル ワンステップ修復の 3 つの結果が示されています。この貢献は方法論的なものであり、特定のコントローラーレベルの障害が、制御された条件下で観察可能、説明可能、局所的に修正可能、および経験的に再テストできるかどうかをテストするためのシミュレーション互換手順を提供します。

原文 (English)

Verification of Adaptive Agentic Controllers through Finite Rule Revision

Industrial agentic AI systems increasingly exhibit a gap between prototype capability and production deployment. In particular, adaptive agents may generate plausible outputs while remaining difficult to verify under non-determinism, confidentiality constraints, limited context, and weak observability. This paper formulates a bounded verification protocol for adaptive agentic controllers represented by finite symbolic rules, explicit diagnostic predicates, explanation logs, and held-out re-evaluation. The central research question is: when an adaptive agentic controller is represented through finite rules, explicit diagnostic predicates, explanation logs, and held-out re-evaluation, which classes of controller failure can be detected, locally repaired, or rejected without relying on unrestricted human-in-the-loop judgment? The proposed framework treats the controller as a finite revisable object. Diagnostic failures are mapped to predefined rule-level edits, including rule addition, rule deletion, and priority revision. Repaired controllers are then evaluated on held-out simulation seeds or cloned initial states. Experiments in a stylized financially constrained inventory-control benchmark show three outcomes: resource-induced failures that remain non-repairable by one rule edit, partial repairs that are rejected because they violate thresholds or guardrails, and a local one-step repair of an order-volatility failure induced by removing a smoothing rule. The contribution is methodological and provides a simulation-compatible procedure for testing whether specific controller-level failures can be made observable, explainable, locally revisable, and empirically re-tested under controlled conditions.

13:00 JSTLLM/生成AIエージェント

EvoCUA-1.5: マルチターンコンピュータ使用エージェントのためのオンライン強化学習

コンピュータを使用するエージェントは、部分的に観察可能なマルチモーダルなデスクトップ環境との繰り返しの対話を通じて、長期的なタスクを解決する必要があります。模倣学習とオフライン軌跡の改良は強力な事前分布を提供しますが、静的トレースでは実際のコンピューター使用の因果的フィードバック ループをカバーすることはできません。各アクションは、画面の状態、将来のアクション空間、および回復オプションを変更します。 EvoCUA-1.5 は、自己進化するコンピューター使用エージェントをオフラインの体験学習からオンラインの強化学習に拡張します。そこでは、ポリシーが実行可能なサンドボックス環境と相互作用し、検証可能なタスクの結果から改善されます。この設定でのオンライン RL では、シングルターン言語 RL レシピを直接再利用するだけでは不十分です。マルチターン インタラクションにより、コンテキスト管理された観察、まばらな最終報酬、可変長の軌跡、および遅い環境フィードバックが導入されます。 EvoCUA-1.5 は、ステップレベルのポリシー最適化 (STEPO) でこれらの課題に対処します。STEPO は、ステップレベルのサンプルに分解した後も軌道レベルのアドバンテージのバランスを維持します。検証可能な合成タスクに対するポリシー対応フィルタリングとパスレート調整。ダイナミック トライアダプティブ カリキュラム (DTAC) は、学習可能なタスク、難しいポジティブなリプレイ、制御された実行不可能なタスクの露出を組み合わせたものです。そして、失効制御とミニグループバッチ処理を備えた完全に非同期の RL インフラストラクチャ。実験では、これらのコンポーネントがトレーニングの安定性と下流のパフォーマンスを向上させることが示されています。 EvoCUA-1.5 は、OSWorld-Verified で 63.2% の成功率を達成し、同等の 32B/​​35B スケールのオープンウェイト ベースラインを上回り、パラメーター数が大幅に多いモデルにさえ近づいています。全体として、EvoCUA-1.5 は、マルチターン コンピューター使用エージェントでオンライン RL を拡張するための実用的なフレームワークを提供します。

原文 (English)

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static traces cannot cover the causal feedback loop of real computer use: each action changes the screen state, future action space, and recovery options. EvoCUA-1.5 extends self-evolving computer-use agents from offline experience learning to online reinforcement learning, where policies interact with executable sandbox environments and improve from verifiable task outcomes. Online RL in this setting requires more than directly reusing single-turn language-RL recipes. Multi-turn interaction introduces context-managed observations, sparse terminal rewards, variable-length trajectories, and slow environment feedback. EvoCUA-1.5 addresses these challenges with Step-Level Policy Optimization (STEPO), which preserves trajectory-level advantage balance after decomposition into step-level samples; policy-aware filtering and pass-rate calibration over verifiable synthesized tasks; Dynamic Tri-Adaptive Curriculum (DTAC), which combines learnable tasks, difficult positive replay, and controlled infeasible-task exposure; and a fully asynchronous RL infrastructure with staleness control and mini-group batching. Experiments show that these components improve training stability and downstream performance. EvoCUA-1.5 achieves 63.2\% success on OSWorld-Verified, outperforming comparable 32B/35B-scale open-weight baselines and even approaching models with significantly larger parameter counts. Overall, EvoCUA-1.5 provides a practical framework for scaling online RL in multi-turn computer-use agents.

13:00 JST研究/論文

パターンから迷路構造まで: SMT ベースのパス合成と 2D/3D 構築

テキストや図形などの入力パターンから迷路構造を構築するためのパイプラインを紹介します。中心となるパス合成問題は、満足性モジュロ理論で隣接性、連続性、パターン制約のカバレッジに関するグローバル制約としてエンコードされ、固定境界の各インスタンスを 1 回の呼び出しで解決できるようになります。結果として得られる経路は、平面的な自己回避ルート、または規定の上下交差を伴う階層的な横断のいずれかであり、平面迷路や編まれた迷路を立体的に実現するための足場として機能します。このレポートは、発行された Bridges 2026 カンファレンス論文を拡張し、より代表的な SMT-LIB の例と、合成されたパスがどのようにして平面および 3 次元の具体的な迷路構造になるかについて詳しく説明します。

原文 (English)

From Patterns to Maze Structures: SMT-Based Path Synthesis and 2D/3D Construction

We present a pipeline for constructing maze structures from input patterns such as text or shapes. The central path-synthesis problem is encoded in Satisfiability Modulo Theories as global constraints on adjacency, continuity, and pattern-constrained coverage, allowing each fixed-bound instance to be solved in one call. The resulting path is either a planar, self-avoiding route or a layered traversal with prescribed over--under crossings, and it serves as a scaffold for constructing planar mazes and three-dimensional realizations of woven mazes. This report extends the published Bridges 2026 conference paper with more representative SMT-LIB examples and a fuller account of how synthesized paths become concrete maze constructions in planar and three-dimensional form.

13:00 JSTLLM/生成AI

長さのペナルティにより思考連鎖が監視されにくくなる

長さにペナルティを課した強化学習は、モデルの答えを駆動する影響を隠しながら、思考連鎖推論を短縮することができます。私たちの実験では、モデルの思考連鎖がヒントに言及する頻度ははるかに低かったにもかかわらず、長さペナルティを使用したトレーニングによって、モデルの操作からの誤解を招くヒントが阻止されませんでした。トークン精度の評価では、これらの実行は成功したものとしてカウントされます。これは、使用する推論トークンが少なく、精度の低下もほとんどないためです。残りの痕跡が、答えを導き出したものをまだ示しているかどうかを見逃してしまいます。異なるターゲット鎖長で Qwen3-4B および Qwen3-14B バリアントをトレーニングし、次に、保持された MMLU-Pro-R と 4 つの転送ベンチマークに対するバイアス ヒント介入でそれらを評価します。圧縮により推論トークンが大幅にカットされ、多肢選択の精度がほとんど維持され、ヒントの影響がベースライン付近に残ります。最も強力なターゲットでは、下限の忠実度は Qwen3-14B ではベースラインの 63.1%、Qwen3-4B では 69.4% に低下します。モニターがヒント使用を捕捉する生の率は 69% から 49% に、そして 60% から 48% に低下します。長さをコンテンツから分離するために、残りのテキストが圧縮された長さと一致するまで、非圧縮のベースライン チェーンから文をランダムに削除します。この長さのマッチングの後でも、圧縮されたチェーンは、Qwen3 サイズと 5 つの評価分布すべてについて、ランダムに短縮したベースライン チェーンよりもヒントを開示する頻度が 7 ~ 35 パーセント ポイント低くなります。したがって、圧縮は推論を短縮するだけではなく、モニターが答えに影響を与えたものを確認するために必要な手がかりを優先的に削除します。これらの結果を総合すると、より安価な推論によって答えを保存しながら、その背後にある影響を検出することが困難になる、圧縮監視可能性のフロンティアが明らかになります。

原文 (English)

Length Penalties Make Chain-of-Thought Less Monitorable

Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even though the models' chains of thought mention the hint much less often. A token-accuracy evaluation would count these runs as successful because they use fewer reasoning tokens with little accuracy loss; it would miss whether the remaining trace still shows what drove the answer. We train Qwen3-4B and Qwen3-14B variants with different target chain lengths, then evaluate them with biasing-hint interventions on held-out MMLU-Pro-R and four transfer benchmarks. Compression sharply cuts reasoning tokens, preserves most multiple-choice accuracy, and leaves hint influence near baseline. At the strongest target, lower-bound faithfulness falls to 63.1% of baseline for Qwen3-14B and 69.4% for Qwen3-4B; the raw rate at which a monitor catches hint use falls from 69% to 49% and from 60% to 48%. To separate length from content, we randomly delete sentences from uncompressed baseline chains until the remaining text matches the compressed length. Even after this length matching, compressed chains disclose the hint 7-35 percentage points less often than baseline chains that we shorten at random, for both Qwen3 sizes and all five evaluation distributions. Compression therefore does more than shorten reasoning, preferentially removing the cues a monitor needs to see what influenced the answer. Together, these results reveal a compression-monitorability frontier in which cheaper reasoning can preserve answers while making the influences behind them harder to detect.

13:00 JST研究/論文GPT / ChatGPT

PHITSBench: 自然言語を使用した AI 支援 PHITS 放射線輸送入力生成の実行スコア付きベンチマーク

モンテカルロ粒子および重イオン輸送コード システム (PHITS) の実行スコア付きベンチマークである PHITSBench を紹介します。 PHITSBench は、パラメータ編集 (Edit)、構文修復 (Repair)、および自然言語記述からの完全なシミュレーション生成 (Reproduction) の 3 つの一般的なワークフロー カテゴリにまたがる 282 のトランスポート スコアリング可能なタスクで構成されます。各タスクは、実行の成功と、生成されたトランスポート オブザーバブルと参照トランスポート オブザーバブル間の一致を組み合わせた複合メトリック スコアを使用して評価されます。 PHITSBench を使用して、ゼロショット プロンプトから知識強化およびエージェント ワークフローに至る 5 つの GPT-5.4 ベースの構成を評価します。ドメイン固有の知識がないと、モデルは編集タスクと修復タスクではうまく機能しますが (それぞれ 95% と 70% 成功)、正しいシミュレーションを最初から生成することはできません (再現トラックでの成功率 0%)。ユーザーマニュアルとともに提供される構造化された機械可読な PHITS ナレッジカタログにより、単発の再現タスクの成功率が 57% に向上します。エージェント実行ではさらに 66 ~ 73% 向上しますが、計算コストが増加します。障害分析により、残りのエラーの大部分は、構文生成ではなく、物理オブザーバブルの誤った選択と構成によって占められていることがわかります。これらの結果は、AI 支援による放射線輸送モデリングの将来の進歩は、基礎モデル自体の進歩と同じくらい、機械可読知識ベース、厳選されたドメイントレーニングデータセット、および実行に基づいた評価環境に依存することを示唆しています。

原文 (English)

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language

We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS). PHITSBench comprises 282 transport-scorable tasks spanning three common workflow categories: parameter editing (Edit), syntax repair (Repair ), and complete simulation generation from natural-language descriptions (Reproduce). Each task is evaluated using a Composite Metric Score that combines execution success with agreement between generated and reference transport observables. Using PHITSBench, we evaluate five GPT-5.4-based configurations ranging from zero-shot prompting to knowledge-augmented and agentic workflows. Without domain-specific knowledge, the model performs well on editing and repair tasks (95% and 70% success, respectively) but fails to generate correct simulations from scratch (0% success on the Reproduce track). A structured, machine-readable PHITS knowledge catalog, supplied alongside the user manual, raises single-shot Reproduce-task success to 57%. Agentic execution provides a further improvement to 66-73%, but at increased computational cost. Failure analysis shows that the remaining errors are dominated by incorrect selection and configuration of physical observables rather than syntax generation. These results suggest that future progress in AI-assisted radiation-transport modeling will depend as much on machine-readable knowledge bases, curated domain-training datasets, and execution-grounded evaluation environments as on advances in foundation models themselves.

13:00 JST研究/論文

推論クラスの意思決定支援システムにおけるセマンティックドリフトとオペレーター制御の安定性

この記事では、新世代のハイブリッド人間機械意思決定支援システム (DSS) において、オペレータ制御の安定性を確保し、目標設定を維持するという基本的な問題を調査します。モノグラフ形式のテキスト配列の共同設計に関する 2 か月にわたる継続的な縦断実験に基づいて、深層論理推論 (Reasoning LLM) の大規模言語モデルにおける意味論的コンテキストのドリフトの潜在現象が検証され、説明されています。マンマシンインターフェースにおける相互作用の数学的モデルが提案され、隠された推論チェーンの非線形コンテキスト圧力を考慮した独自のメトリックであるオペレーター制御安定性係数が導入されます。コグニトーム理論のパラダイム内で、制御機能の逆転の臨界点が捉えられます。エンジニアリング上の推奨事項は、修正された階層的類似性モデルに基づいて動的なリレーショナル調停ループを実装するために策定されます。

原文 (English)

Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems

The article investigates the fundamental problem of ensuring the stability of operator control and preserving goal-targeting in hybrid human-machine decision support systems (DSS) of a new generation. Based on a two-month continuous longitudinal experiment on the joint design of a monograph-format textual array, the latent phenomenon of semantic context drift in large language models of deep logical reasoning (Reasoning LLMs) is verified and described. A mathematical model of interaction in the human-machine interface is proposed, and an original metric is introduced - the operator control stability coefficient, which takes into account the non-linear contextual pressure of hidden reasoning chains. Within the paradigm of the cognitome theory, a critical point of control functions inversion is captured. Engineering recommendations are formulated for implementing dynamic relational arbitration loops based on a modified hierarchical similarity model.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPTGemini

自己発見仕様によるエージェントコンテキスト学習

コンテキスト学習は、LLM が事前トレーニングにはない複雑なコンテキストから新しいタスク固有の知識を学習して適用する必要がある、新たな推論時間タスクです。フロンティア モデルでさえ、タスク成功率が 24% 未満です。この研究では、この設定が依然として難しい理由を理解するために、包括的な実証研究を実施します。当然の仮説として、障害はコンテンツへのアクセスが原因であるというものがあります。しかし、広範なコンテキスト学習ベンチマークである CL-Bench の 12 の検索、リフレクション、および検証ベースライン全体では、直接のフルコンテキスト プロンプトに比べて限られた利益しか得られないことがわかりました。さらに失敗を分析すると、重要な発見が明らかになります。長い文書の理解などの一般的な長いコンテキストのタスクとは異なり、コンテキスト学習では、ローカル コンテンツを回復するだけでなく、多くの場合、クエリでは指定されていないがコンテキスト全体に分散されているローカル仕様 (ドメイン固有の形式、ローカル ルール、完全性条件) を取得する必要があります。 31,592 のルーブリック項目すべてにおいて、55.4% が仕様の習得を明確に評価しているのに対し、内容の習得を評価しているのは 22.6% のみであることがわかりました。さらに、仕様の 76.7% がユーザーのクエリで指定されていないにもかかわらず、95.5% はコンテキストまで追跡可能であり、これらは隠れた要件ではなく学習可能な義務であることを示しています。この診断を検証するために、ローカル仕様を抽出し、敵対的なチェックと修復を通じてそれらを強制する、意図的に単純な介入 PSCI (プライベート仕様契約誘導) を設計します。 PSCI は、CL-Bench で GPT-5.1 (絶対値 +5.59 pp、相対 +24.8%) で最先端の 28.14% を達成し、Qwen3.5-27B (+5.28 pp) および Gemini 3 Pro (+6.17 pp) で再現されました。 17 のアブレーションは、タスク固有の仕様の役割をさらに分離します。全体として、私たちの結果は、コンテキスト学習がコンテンツの取得だけでなく仕様の取得にも依存していることを示唆しています。

原文 (English)

Agentic Context Learning with Self-Discovered Specification

Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier models score under 24% task success. In this work, we conduct a comprehensive empirical study to understand why this setting remains difficult. A natural hypothesis is that failures stem from content access; yet across twelve retrieval, reflection, and verification baselines on CL-Bench, an extensive context learning benchmark, we find limited gains over direct full-context prompting. Further failure analysis reveals a key finding: unlike typical long-context tasks such as long document understanding, context learning requires not only recovering local content but also acquiring local specifications that are often unspecified in the query but distributed across the context: domain-specific formats, local rules, and completeness conditions. Across all 31,592 rubric items, we find that 55.4% clearly evaluate specification acquisition, while only 22.6% evaluate content acquisition. Moreover, despite 76.7% of specifications being unspecified in the user query, 95.5% are traceable to the context, indicating these are learnable obligations rather than hidden requirements. To validate this diagnosis, we design a deliberately simple intervention PSCI (private specification-contract induction) which extracts local specifications and enforces them through adversarial checking and repair; PSCI achieves state-of-the-art 28.14% with GPT-5.1 (+5.59 pp absolute and +24.8% relative) on CL-Bench, replicated on Qwen3.5-27B (+5.28 pp) and Gemini 3 Pro (+6.17 pp). Seventeen ablations further isolate the role of task-specific specifications. Overall, our results suggest context learning hinges on not only content acquisition but also specification acquisition.

13:00 JST画像/動画生成エージェント

高品質の数学視覚教材を生成するためのエージェント ワークフローの探索

数学的図表は、幼稚園から高等学校までの教育において、問題の構成要素として、また生徒の理解のための足場として重要な役割を果たします。ただし、大規模言語モデル (LLM) を含む現在の AI ツールは、詳細な説明が提供されている場合でも、正確で教育的に健全な視覚的な図を確実に生成するのに苦労しています。したがって、中学校の数学のための信頼性の高い図の生成には大きなギャップが残っています。これに対処するために、LLM エージェントが生成されたビジュアルの品質を評価し、このフィードバックを使用して出力を反復的に改善できるようにするエージェント ワークフローを導入します。この自己改善ループは、AI が生成した図の精度と教育的適切性を高めることを目的としています。私たちの調査では 2 つの疑問が調査されています。まず、視覚的な品質に関する特定の基準が与えられた場合、LLM は視覚補助の品質保証に関する質問を正確に生成できるでしょうか?第二に、有効な品質保証の質問が与えられた場合、ビジョン言語モデルは、生成された K 12 視覚補助を効果的に評価し、結果として得られるフィードバックを使用して反復的に改善できるでしょうか?当社はエージェント ワークフローの探索的評価を実施し、より強力な空間推論や、生成された品質保証質問における図の特徴のより包括的なカバーなど、改善すべき重要な領域を特定します。私たちの結果は、このアプローチが AI で生成された数学図の信頼性と教育的価値を向上させることができるという予備的な証拠を提供します。

原文 (English)

Exploring Agentic Workflows for Generating High Quality Math Visual Aids

Mathematical diagrams play a crucial role in K 12 education, both as problem components and as scaffolding for student comprehension. However, current AI tools, including Large Language Models (LLMs), struggle to reliably generate accurate and pedagogically sound visual diagrams, even when provided with detailed descriptions. A significant gap therefore remains in the reliable generation of diagrams for middle school mathematics. To address this, we introduce an agentic workflow that enables LLM agents to evaluate the quality of generated visuals and use this feedback to iteratively improve their outputs. This self improvement loop aims to enhance the accuracy and educational appropriateness of AI generated diagrams. Our research investigates two questions. First, can LLMs accurately generate quality assurance questions for a visual aid given specific criteria for visual quality? Second, given valid quality assurance questions, can Vision Language Models effectively evaluate generated K 12 visual aids and use the resulting feedback to improve them iteratively? We conduct an exploratory evaluation of our agentic workflow and identify key areas for improvement, including stronger spatial reasoning and more comprehensive coverage of diagram features in the generated quality assurance questions. Our results provide preliminary evidence that this approach can improve the reliability and educational value of AI generated mathematical diagrams.

13:00 JST研究/論文

TopoExplore: アーカイブベースの探索のためのトポロジカル識別

Go-Explore などのアーカイブベースの探索手法は、訪問の希少性を使用してどの訪問州に戻るかを選択し、フロンティア手法は未知の境界に戻ります。どちらも、境界の背後にある未踏の領域がそもそも立ち入ることができるかどうかを問うものではありません。探索とは、単に報酬を見つけることではなく、下流の学習と計画のために構造的に完全な体験を収集することです。 TopoExplore を導入します。これは、定期的なトポロジ パスで Go-Explore セルの選択を強化します。訪問セット占有グリッドの囲まれた未探索領域 (ボイド) は、フラッド フィル (立方体複合体の H1 クラス) によって検出され、減衰選択ボーナスは厳密な入り口 (ギャップまたはドア セル) にのみ配置されるため、密閉された領域は決してターゲットにされず、入力された領域はリタイアします。制御された 18 環境の MiniGrid スイート (シード 15 個、ハイパーパラメータが凍結) では、TopoExplore は、正確な Go-Explore アブレーションと比較して、最初のエントリまでのステップの中央値で 1.52 倍の幾何平均速度向上を達成しました。これに対し、フロンティア ベースラインでは 1.37 倍でした。封印されたデコイ構造が出現するとフロンティア探索は低下します (デコイ環境では 0.83 ~ 1.48 倍、TopoExplore では 1.65 ~ 2.11 倍) が、ハード マルチインタラクション ドアでは TopoExplore が最大の勝利を収めます (10.9 倍)。モンテズマの復讐については正直な否定的結果を報告します。壁の知識がなければ、到達できない占有アーティファクトがボーナスを獲得し、成長するにつれてパフォーマンスが低下し、耐荷重コンポーネントとして壁を意識した入場テストを分離します。HM3D スキャンされた建物については、フロンティア選択がブランケット カバレッジを支配しているにもかかわらず、Go-Explore を超える高速化がシーンの難易度を追跡する (r=0.69) という予備的な肯定的結果を報告します。証拠は、意図的に範囲を絞った主張を裏付けています。つまり、トポロジーを意識した選択は、閉じられた構造を区別する必要がある場合に効果があり、フロンティア手法が最も強力なオープンカバレッジでは、その体制に合わせて調整されていないにもかかわらず、競争力を維持します。

原文 (English)

TopoExplore: Topological Discrimination for Archive-Based Exploration

Archive-based exploration methods such as Go-Explore select which visited state to return to using visitation rarity, and frontier methods return to the boundary of the unknown; neither asks whether the unexplored region behind a boundary is enterable at all. Exploration is not just about finding reward - it is about collecting a structurally complete experience for downstream learning and planning. We introduce TopoExplore, which augments Go-Explore cell selection with a periodic topological pass: enclosed unexplored regions (voids) of the visited-set occupancy grid are detected by flood fill (the H1 classes of its cubical complex), and a decaying selection bonus is placed only on their strict entrances (gap or door cells), so sealed regions are never targeted and entered regions retire. On a controlled 18-environment MiniGrid suite (15 seeds, frozen hyperparameters) TopoExplore attains a 1.52x geometric-mean speedup in median steps-to-first-entry over its exact Go-Explore ablation, versus 1.37x for a frontier baseline; frontier exploration degrades when sealed decoy structure appears (0.83-1.48x on decoy environments vs. 1.65-2.11x for TopoExplore), while TopoExplore holds its largest win on hard multi-interaction doors (10.9x). We report an honest negative on Montezuma's Revenge - without wall knowledge, unreachable occupancy artifacts capture the bonus and performance degrades as it grows, isolating the wall-aware entrance test as the load-bearing component - and a preliminary positive on HM3D scanned buildings, where the speedup over Go-Explore tracks scene difficulty (r=0.69) even as frontier selection dominates blanket coverage. The evidence supports a deliberately scoped claim: topology-aware selection pays off where enclosed structure must be discriminated, and remains competitive at open coverage, where frontier methods are strongest, despite not being tuned for that regime.

13:00 JSTLLM/生成AIエージェント

Who&When Pro: LLM は本当に AI エージェントの障害の原因を特定できるのでしょうか?

自動障害属性では、LLM を使用して、エージェント システムが障害を起こす場所と理由を特定します。エージェントの能力が高まるにつれて、エージェントの失敗はより巧妙になり、自動化されたアトリビューションの重要性がますます高まっています。エージェント システムにおける自動障害属性の大規模ベンチマークである Who&When Pro を紹介します。成功したプレフィックスを正確に再生した後にのみ失敗を挿入する厳密に制御されたパイプラインを使用して、3 つのモダリティとさまざまなシナリオをカバーする 26 のベンチマークにわたるゴールデン ラベルを持つ 12,326 の失敗した軌跡を構築します。ベンチマークを超えて、私たちは広範な実験と分析を実施し、モデルがモダリティ、プロトコル、モデルファミリー全体で障害をどのように帰属させるかに関する体系的なパターンを明らかにし、将来の自動障害帰属システムに経験的なガイダンスを提供します。

原文 (English)

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important. We introduce Who&When Pro, a large-scale benchmark for automated failure attribution in agentic systems. Using a strictly controlled pipeline that injects a failure only after exactly replaying a successful prefix, we construct 12,326 failed trajectories with golden labels across 3 modalities and 26 benchmarks covering various scenarios. Beyond benchmarking, we conduct extensive experiments and analyses, revealing systematic patterns in how models attribute failures across modalities, protocols, and model families, and providing empirical guidance for future automated failure attribution systems.

13:00 JSTハードウェア/半導体

量子化シミュレーションによるライトバックと解釈可能なプログラム実行のためのシンボリック ニューラル CPU

ニューラル ネットワークはアルゴリズムの入出力マッピングを学習できますが、学習した実行プログラムを信頼するには、最終的な正しい答え以上のものが必要です。これは、それを生み出す状態遷移が通常隠蔽されているためです。これらの遷移を可視化するために、トレース監視シンボリック ニューラル CPU、リカレント制御を組み合わせた因数分解学習実行アーキテクチャ、固定微分可能算術論理ユニット バンク上の明示的操作ルーター、宛先マスクされたレジスタ ライトバック、完全な軌跡監視、および一致した固定小数点再生を導入します。このモデルは、選択された操作、ソースおよびデスティネーション レジスタ、レジスタの軌跡、メモリ信号、およびライトバック セマンティクスを各ステップで公開します。主要な 16 幅ベンチマークでは、非量子化エグゼキュータは参照実行を正確に再現しますが、8 ビット量子化でシミュレートされたエグゼキュータは 1,000 命令のプログラムを通じてシンボリック演算パスを保存します。同じ実行が一致する固定小数点再生に対して評価されると、残留数値ドリフトは消失します。これは、それが実行の失敗によるものではなく、連続参照セマンティクスと低精度参照セマンティクスの不一致に起因することを示しています。リカレント、トランスフォーマー、時間畳み込み、時間グラフからインスピレーションを得たコントローラー、状態空間コントローラーを比較し、検査可能な実行パスにはオペレーション ゲートの監視が必要であることをアブレーションで示しています。隠れたオペコードのメモリプレッシャータスクは、遅延状態の使用と一時的なバインディングの残りの制限を明らかにします。また、ValueMemory、ハイブリッド適応リーキー統合発射コントローラー、動作クローン作成とアクタークリティカル強化学習を通じて訓練された候補制約付きシンボリック制御、および RV32I ベース整数セマンティック ブリッジを使用してインターフェイスを拡張します。これらの結果を総合すると、解釈可能で低精度かつ制御可能なニューラル実行のためのトレース検証可能なフレームワークが確立されます。

原文 (English)

A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution

Neural networks can learn algorithmic input-output mappings, but trusting a learned executor requires more than a correct final answer because the state transitions that produce it are usually hidden. To make those transitions visible, we introduce a trace-supervised symbolic neural CPU, a factorized learned execution architecture that combines recurrent control, an explicit operation router over a fixed differentiable arithmetic-logic unit bank, destination-masked register writeback, complete trajectory supervision and matched fixed-point replay. The model exposes the selected operation, source and destination registers, register trajectory, memory signals and writeback semantics at every step. On the principal 16-wide benchmark, the non-quantized executor reproduces reference execution exactly, while the eight-bit quantization-simulated executor preserves the symbolic operation path through programs of 1,000 instructions. When the same execution is evaluated against a matched fixed-point replay, the residual numerical drift disappears, showing that it comes from a mismatch between continuous and low-precision reference semantics rather than from execution failure. We compare recurrent, Transformer, temporal-convolution, temporal graph-inspired and state-space controllers, and the ablations show that operation-gate supervision is necessary for an inspectable execution path. Hidden-opcode memory-pressure tasks expose the remaining limits in delayed state use and temporal binding. We also extend the interface with ValueMemory, hybrid adaptive leaky integrate-and-fire controllers, candidate-constrained symbolic control trained through behaviour cloning and actor-critic reinforcement learning, and an RV32I base-integer semantic bridge. Together, these results establish a trace-verifiable framework for interpretable, low-precision and controllable neural execution.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達Gemini

AgentAbstain: LLM エージェントは、いつ行動すべきではないかを知っていますか?

大規模言語モデル (LLM) に基づくエージェント システムは自律タスクに導入されることが増えていますが、既存の評価では主に、エージェントがいつやめるべきかを知っているかどうかよりも、タスクの成功に重点が置かれています。このギャップは実際のリスクをもたらします。曖昧さ、矛盾する制約、またはツールの障害の下では、エージェントが意図しない取り消し不能なアクションを実行する可能性があります。このギャップを埋めるために、我々は、エージェントの棄権に関する最初の体系的な評価フレームワークを提示します。それは、ツールを使用する LLM エージェントがいつ行動を起こさないかを認識する調整された能力です。 AgentAbstain の核心は、実行前の推論と実行時検出にわたる 8 つの棄権シナリオのエージェント ネイティブ分類に基づいて構築されたペアタスク ベンチマークです。これには、42 の実行可能なサンドボックス環境にわたる 263 のペアのタスクが含まれており、各ペアは、命令、ツール、または環境の状態に対する制御された摂動を通じて生成される、実行すべきタスクと禁止すべきバリアントで構成されます。このペア設計を拡張し、データ汚染に対抗するために、サンドボックス環境を合成し、決定論的再生およびセマンティック LLM ジャッジによって検証されたペアタスクをエンドツーエンドで生成する完全に自動化されたパイプラインである AbstainGen を提案します。新しいタスク インスタンスはオンデマンドで再生成でき、3 人の独立したアノテーターは、サンプリングされたタスクの 94 ~ 98% が適切に設計されていると評価しました。 4 つのエージェント ハーネスの 17 個のフロンティア LLM にわたって、最高のエージェント (Gemini 3.1 Pro) は、ペアの精度 (各ペアのタスクの実行側と棄権側の両方で正しい) を 59.5% しか達成していません。さらに重要なことは、棄権能力は一般的な課題解決能力とはほとんど独立しており、課題解決を拡大するだけではこのギャップは埋まらないことを示しています。さらに、エージェントが棄権トリガーを認識する前に不可逆的なアクションを実行する事後棄権などの障害モードを特定します。私たちのコードとデータセットは、agentabstain.github.io でオープンソース化されています。

原文 (English)

AgentAbstain: Do LLM Agents Know When Not to Act?

Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, conflicting constraints, or tool failures, agents may execute unintended and irreversible actions. To close this gap, we present the first systematic evaluation framework for agentic abstention: the calibrated ability of tool-using LLM agents to recognize when not to act. At its core, AgentAbstain is a paired-task benchmark built on an agent-native taxonomy of 8 abstention scenarios across pre-execution reasoning and runtime discovery. It contains 263 paired tasks across 42 executable sandbox environments, where each pair consists of a should-act task and a should-abstain variant produced through a controlled perturbation to the instruction, tool, or environment state. To scale this paired design and resist data contamination, we propose AbstainGen, a fully automated pipeline that synthesizes sandbox environments and generates paired tasks end-to-end, validated by deterministic replay and semantic LLM judges; fresh task instances can be regenerated on demand, and three independent annotators rate 94-98% of sampled tasks as well-designed. Across 17 frontier LLMs in 4 agent harnesses, the best agent (Gemini 3.1 Pro) achieves only 59.5% paired accuracy (correct on both the act and abstain sides of each paired task). More importantly, abstention capability is largely independent of general task-solving capability, indicating that scaling task-solving alone will not close this gap. We further identify failure modes such as post-hoc abstention, in which agents execute irreversible actions before recognizing abstention triggers. Our code and dataset are open-sourced at agentabstain.github.io.

13:00 JST研究/論文

曖昧な発話から管理された再利用クラスまで: 正規化、商の不変性、条件付き決定可能性

セマンティック キャッシュは、類似性を埋め込む際の回答の再利用を定義します。つまり、類似性スコアがしきい値をクリアすると、2 つの発話が保存された回答を共有します。承認、バージョン管理、または 2 つの要求が同じになる理由などの概念はありません。このメモは、再利用が定義されるオブジェクトを変更します。管理されたドメインでは、再利用は、類似性ヒューリスティックではなく、解決された会話要求の数学的に特徴付けられた商に基づいて動作する必要があります。解決された発話に関する 3 つの独立して定義された関係 (読み取り ID、解決 ID、および再利用 ID) は、展開ログで確認できる実現された非縮退条件の下で厳密な洗練チェーンを形成します。パイプラインの出力はチェーンに沿って不変であり、再利用アイデンティティはまさに管理された回答パーティションへの解決マップのカーネルであるため、再利用商はパーティションが誘導する発話側のオブジェクトであり、再ラベル付けではありません。 ID ライセンスを再利用すると、管理されたクエリ キーとその認定された回答スペースが使用されます。特定の回答を再利用するには、解決 ID または適用性証明書が必要です。支持層は証明された強度で正確に記述されています。正確な表示の正規形です。非エスケープを特徴付けるクロージャー安定セルを使用して、設計演算子として集約を結合します。信頼できないプロポーザル層と比較したパイプライン全体の合計の計算可能性。恣意的な提案者に対する政策の許容性 - 事実に基づく根拠や意図の忠実性がないことが証明されている。そして、限られた数の有益な応答の後に引き出しが終了し、目標の一貫性を満たしていないように聞こえます。

原文 (English)

From ambiguous utterances to governed reuse classes: canonicalization, quotient invariance, and conditional decidability

Semantic caching defines answer reuse on embedding similarity: two utterances share a stored answer when a similarity score clears a threshold, with no notion of authorization, versioning, or of what makes two demands the same. This note changes the object on which reuse is defined: in a governed domain, reuse should operate on a mathematically characterized quotient of resolved conversational demands, not on a similarity heuristic. Three independently defined relations on resolved utterances -- reading identity, resolution identity, and reuse identity -- form a refinement chain, strict under realized nondegeneracy conditions checkable on deployment logs; the pipeline's outputs are invariant along the chain, and reuse identity is exactly the kernel of the resolution map into the governed answer partition, so the reuse quotient is the utterance-side object that partition induces, not a relabeling of it. Reuse identity licenses the governed query key and its certified answer space; reuse of a particular answer requires resolution identity or an applicability certificate. The supporting layer is stated at exactly the strength proved: exact-denotation normal forms; join aggregation as a design operator, with closure-stable cells characterizing no-escape; total computability of the full pipeline relative to an untrusted proposal layer; policy admissibility for arbitrary proposers -- and provably not factual grounding or intent fidelity; and elicitation terminating after finitely many informative replies, sound under target consistency.

13:00 JSTLLM/生成AIエージェント研究/論文

MAG: マルチモーダル アクションとガイド生成のための Web エージェント ベンチマークとハーネス

デジタル アダプション プラットフォーム (DAP) は、Web システムで広く使用されている埋め込みオーバーレイで、ページ内の操作をユーザーにガイドし、不慣れなインターフェイスをすぐに使い始めるのに役立ちます。ただし、実際のタスクを完了するということは、1 つのページ上でいくつかのボタンをクリックすることを意味することはほとんどありません。ページの状態が変化するたびに展開される一連のアクションが必要です。また、以前の研究では、自動化された Web エージェントのアクションとガイド テキストの生成を 2 つの別個の問題として扱っており、そのほとんどは人間が実際に操作するレンダリングされた画面ではなく、DOM やアクセシビリティ ツリーなどのテキスト ページ表現をモデルにフィードします。この作業では、タスクの実行とガイドの書き込みを 1 つのマルチモーダル アクションとガイド タスクに統合する最初のベンチマークである MAG を紹介します。このベンチマークには、スクリーンショット上の 2 つの基礎スキーム (セット オブ マーク要素の選択と生のピクセル座標) が含まれます。さらに、LLM 支援によるアノテーション、人間による検証、トレーニング、ライブ環境での評価、およびアクションとガイドの共同メトリクスをカバーする、この複合タスクのための完全なハーネスを構築します。このハーネスを使用して、フロンティア API モデルとオープン マルチモーダル モデルを評価し、詳細な分析をレポートします。最後に、専門家の軌跡を追加した GRPO トレーニング方法を設計します。これにより、監視された 9B エージェントの成功率がほぼ 2 倍 (6.9% から 13.2%) になり、同時にガイドの品質が向上します。最も強力なモデルでも完了するタスクは 40% 未満であり、将来の研究の余地は十分にあります。

原文 (English)

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking a few buttons on a single page: it takes a sequence of actions that unfolds across changing page states. Prior studies have also treated automated web agent actions and guide text generation as two separate problems, and most of them feed models textual page representations such as the DOM or accessibility trees rather than the rendered screens that humans actually operate on. In this work we introduce MAG, the first benchmark that unifies task execution and guide writing into a single Multimodal Action and Guide task, with two grounding schemes over screenshots: Set-of-Mark element selection and raw pixel coordinates. We further build a complete harness for this compound task, covering annotation with LLM assistance and human verification, training, evaluation in live environments, and joint metrics for actions and guides. With this harness we evaluate frontier API models and open multimodal models, and report detailed analyses. Finally, we design a GRPO training method augmented with expert trajectories, which nearly doubles the success rate of a supervised 9B agent (from 6.9% to 13.2%) and improves guide quality at the same time. Even the strongest model completes fewer than 40% of the tasks, leaving ample room for future research.

13:00 JST研究/論文

適応的な終了状態選択を備えたループ状態空間言語モデル

ループ言語モデルに関する最近の研究では、多くの推論問題は、追加の独立パラメーターではなく、より深い計算深さから恩恵を受けることが示唆されています。しかし、既存の研究はほぼもっぱら Transformer バックボーンに焦点を当てており、この原則が状態空間言語モデルにも適用されるかどうかは不明のままです。私たちは、共有 Mamba (またはハイブリッド) ブロックを繰り返し適用して明示的な有限深さの再帰計算を導入する、Looped Mamba および Looped Hybrid Mamba-Transformer アーキテクチャを調査します。 Mano (モジュラー算術操作) と p ホップ誘導という 2 つの制御された推論タスクにおいて、Looped Mamba はパラメーターが一致した非ループ ベースラインよりも一貫して優れたパフォーマンスを示し、いくつかの設定では、有効深さが等しい非ループ モデルと同等またはそれを上回ります。次に、一致した iso パラメーターと iso-FLOPs プロトコルの下での言語モデルの事前トレーニングに研究を拡張します。これにより、パラメーター共有と有効深さの影響が一緒に解きほぐされます。ループ モデルは、実質的に少ない個別パラメーターの下流ベンチマークで競争力を維持しますが、厳密な iso-FLOPs 比較の下では、より深い非ループ モデルが検証の複雑さにおいて優位性を維持します。最後に、Ouro の 2 段階の出口ゲートを Looped Mamba に適応させ、反復ステップ出力間のしきい値制御された選択を実現します。すべての反復ステップは引き続き実行されるため、選択された終了ステップは、実時間の計算の削減ではなく、予測の深さを表します。研究したスケールでは、適応的な終了状態の選択により、中間の深さでのダウンストリームのパフォーマンスが向上しますが、実際の推論時間の節約には追加の状態処理メカニズムが必要です。

原文 (English)

Looped State-Space Language Models with Adaptive Exit-State Selection

Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transformer backbones, leaving open whether this principle also applies to state-space language models. We investigate Looped Mamba and Looped Hybrid Mamba-Transformer architectures, which repeatedly apply a shared Mamba (or hybrid) block to introduce explicit finite-depth recurrent computation. On two controlled reasoning tasks-Mano (modular-arithmetic manipulation) and p-hop induction-Looped Mamba consistently outperforms parameter-matched non-looped baselines and, in several settings, matches or exceeds non-looped models of equal effective depth. We then extend the study to language model pre-training under matched iso-parameter and iso-FLOPs protocols, which jointly disentangle the effects of parameter sharing and effective depth: looped models remain competitive on downstream benchmarks with substantially fewer distinct parameters, although deeper non-looped models retain an advantage in validation perplexity under strict iso-FLOPs comparisons. Finally, we adapt Ouro's two-stage exit gate to Looped Mamba for threshold-controlled selection among recurrent-step outputs. Since all recurrent steps are still executed, the selected exit step represents prediction depth rather than reduced wall-clock computation. At the scales studied, adaptive exit-state selection improves downstream performance at intermediate depths, while actual inference-time savings require additional state-handling mechanisms.

13:00 JSTエージェント

動的なエージェント スキル: ライフサイクル調査と進化するスキル ライブラリの分類

大規模な言語モデル エージェントは、再利用可能なプロシージャをモデルの外に保存することが増えています。これらの再利用可能なプロシージャは、\emph{スキル} と呼ばれることがよくあります。これらは、コード関数、自然言語命令、SKILL.md パッケージ、ワークフロー グラフ、または将来のエージェントが取得して呼び出すことができる学習済みアダプターなどです。この分類法に基づいた調査では、そのようなスキル ライブラリが時間の経過とともにどのように変化するかを尋ねます。 $124$ の紙 $2023$ ~ $2026$ の監査セット全体で、動的なスキル システムを \emph{ライフサイクルで管理され、検証され、進化するアーティファクト ストア} として合成します。エージェントは対話から証拠を収集し、スキルの更新を提案し、候補者を検証して承認し、取得と構成のためにそれらを整理し、古いエントリを修復または削除し、来歴とロールバックを通じて共有を管理します。私たちは 3 つの調査ツールを中心に文献を整理しています。まず、$\text{six}$ 感覚分類法は、現在の論文で「スキル」と呼ばれる構造的に異なる成果物を区別します。第 2 に、$\text{eight}$ 段階のライフサイクル アーキテクチャにより、証拠の取得、提案、検証/承認、保管、検索/構成、メンテナンス、蒸留/移植性、ガバナンスの背後にある繰り返しの設計上の決定が特定されます。第三に、軽量のスキル レコード スキーマと $\text{ten}$-operator 語彙は、ライブラリの更新を別のメソッド コントリビューションに格上げすることなく比較するための共通用語を提供します。この構造を使用して、明示的な警告を伴う証拠に基づく等級付けパターンを合成します。アドミッションと修復は繰り返し重要であり、検証者の品質はスキルを意識​​した RL に重大な影響を与え、フラット検索はライブラリが成長するにつれて低下する可能性があり、現在のベンチマークはライブラリの軌跡、使用量とユーティリティのギャップ、および安全面を依然として過小報告しています。具体的なレポート基準で締めくくり、静的なプロンプトやツールのコレクションではなくライブラリの変更として動的スキルを評価するための問題を解決します。

原文 (English)

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries

Large language model agents increasingly store reusable procedures outside the model. These reusable procedures are often called \emph{skills}: they may be code functions, natural-language instructions, SKILL.md packages, workflow graphs, or learned adapters that a future agent can retrieve and invoke. This taxonomy-driven survey asks how such skill libraries change over time. Across a $124$-paper $2023$--$2026$ audit set, we synthesize dynamic skill systems as \emph{lifecycle-managed, verified, evolving artifact stores}: agents collect evidence from interaction, propose skill updates, verify and admit candidates, organize them for retrieval and composition, repair or prune stale entries, and govern sharing through provenance and rollback. We organize the literature around three survey tools. First, a $\text{six}$-sense taxonomy distinguishes the structurally different artifacts called ``skills'' in current papers. Second, an $\text{eight}$-stage lifecycle architecture identifies the recurring design decisions behind evidence acquisition, proposal, verification/admission, storage, retrieval/composition, maintenance, distillation/portability, and governance. Third, a lightweight skill-record schema and $\text{ten}$-operator vocabulary provide common terms for comparing library updates without elevating them into a separate method contribution. Using this structure, we synthesize evidence-graded patterns with explicit caveats: admission and repair are repeatedly important, verifier quality materially affects skill-aware RL, flat retrieval can degrade as libraries grow, and current benchmarks still under-report library trajectories, usage--utility gaps, and safety surfaces. We close with concrete reporting standards and open problems for evaluating dynamic skills as changing libraries rather than static prompt or tool collections.

13:00 JSTエージェント研究/論文

IdeaTrail: 科学的アイデアのためのフルプロセス エージェントの軌跡

科学研究は、テキスト生成という単一の行為ではなく、複雑な多段階のワークフローです。通常、アイデアのプロセスは、文献検索、論文の読解、ツールの使用、クレームの確認、論文間の統合、ブレーンストーミング、弱い指示の拒否、および反復的な執筆を通じて現れます。既存のリソースはこのプロセスの個々のコンポーネントをキャプチャしますが、ツールの使用、証拠の取得、中間成果物の進化、アイデアまたは提案レベルのエンドポイントを共同で記録するデータセットは依然として限られています。このレポートでは、科学的アイデアと提案生成のためのマルチターン プロセス軌跡データセットである \method を紹介します。各インスタンスは、証拠の収集からアイデアの選択または提案の作成までの調査プロセスを記録します。 \method は軌道を自由に作成するのではなく、人間が選択した高品質の研究論文と提案成果物から開始し、ジェネレーターとアドバイザーの合成ループを使用します。ジェネレーターはアクション、観察、アーティファクトの編集を通じて目に見える軌道を生成しますが、アドバイザーは完全な生成コンテキストにアクセスして、グラウンディング、因果関係の順序、自然性、および隠れたターゲットからの漏れをチェックします。この逆から順の手順により、実際の科学成果との整合性を保ちながら、研究実践の不確実性、証拠の使用、段階的な収束を近似した複数ターンの研究データが生成されます。 \method は、科学研究エージェント向けにプロセス監視データを合成するためのデータセットと一般的なレシピの両方を提供します。

原文 (English)

IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation

Scientific research is a complex, multi-stage workflow rather than a single act of text generation. The ideation process typically emerges through literature search, paper reading, tool use, claim checking, cross-paper synthesis, brainstorming, rejection of weak directions, and iterative writing. Existing resources capture individual components of this process, but datasets that jointly record tool use, evidence acquisition, intermediate artifact evolution, and idea- or proposal-level endpoints remain limited. This report introduces \method, a multi-turn process-trajectory dataset for scientific ideation and proposal generation. Each instance records a research process from evidence gathering to either idea selection or proposal construction. Rather than freely fabricating trajectories, \method starts from human-selected high-quality research papers and proposal artifacts and uses a Generator--Advisor synthesis loop. The Generator produces the visible trajectory through actions, observations, and artifact edits, while the Advisor has access to the full generation context and checks grounding, causal order, naturalness, and leakage from hidden targets. This reverse-to-forward procedure produces multi-turn research data that remains aligned with real scientific artifacts while approximating the uncertainty, evidence use, and staged convergence of research practice. \method provides both a dataset and a general recipe for synthesizing process-supervision data for scientific research agents.

13:00 JST研究/論文

単元: グラフの継続的学習のための大規模言語モデルの可能性を解き放つ

現実世界のマルチモーダル Web シナリオでは、グラフ構造のデータがストリーミング形式で到着することが多いため、グラフの継続学習は、そのような進化する構造を継続的にモデル化するための重要なパラダイムになります。しかし、既存のグラフ継続学習手法には依然として 2 つの根本的な課題があります。 1) 意味構造の分離。グラフベースの手法はトポロジー関係のモデル化には優れていますが、深い意味論は無視されます。 2) 不均衡な知識伝達。既存のモデルは、初期のタスクから得られた一般知識を効果的に活用して、その後の新しいタスクに利益をもたらすことができません。上記の問題に対処するために、私たちは新しいフレームワーク \textbf{UN}leash Large Language Models PotentIal for Graph ConTinual Learning (UNIT) を提案します。最初のタスクのみで大規模な言語モデルを微調整することで、事前トレーニング済み LLM コーパスとターゲット タスク データセットの間の分布ギャップを埋め、グラフ構造タスクに対する LLM の適応性を強化します。一方、我々は、前のタスクから学んだ普遍的な知識の無視を避け、タスク全体で代表的な知識を効果的に保存する不確実性を認識したアンカー生成メカニズムを提案します。さらに、グラフ トポロジ情報をセマンティック情報に明示的に統合する構造合流モデリングを導入し、セマンティック理解と構造モデリングの間の連携機能を強化します。広範な実験により、私たちの提案した方法がグラフ継続学習タスクにおいて最先端のパフォーマンスを達成することが実証されました。

原文 (English)

UNIT: Unleash Large Language Models Potential for Graph Continual Learning

In real-world multimodal web scenarios, graph-structured data often arrives in a streaming manner, making graph continual learning a crucial paradigm for continuously modeling such evolving structures. However, existing graph continual learning methods still face two fundamental challenges. 1) semantic-structural separation, where the graph-based methods excel at modeling topological relationships but neglect deep semantics. 2) imbalanced knowledge transfer, where existing models fail to effectively leverage general knowledge gained from early tasks to benefit subsequent new tasks. To address above issues, we propose a novel framework, \textbf{UN}leash Large Language Models PotentIal for Graph ConTinual Learning (UNIT). By fine-tuning large language model only on the first task, we bridge the distributional gap between the pre-trained LLM corpus and the target task dataset to enhance the adaptability of LLMs for graph-structured tasks. Meanwhile, we propose an uncertain-aware anchor generation mechanism to effectively preserve representative knowledge across tasks, avoiding the neglect of universal knowledge learned from previous tasks. Additionally, we introduce structural confluence modeling to explicitly integrates graph topology information into semantic information, enhancing the collaborative capabilities between semantic understanding and structural modeling. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance in the graph continual learning task.

13:00 JST研究/論文

GRATE: ゲートロータリーアテンションによる誘導KG基礎モデルの時間的拡張

Ultra や Trix などのナレッジ グラフ基盤モデルは、目に見えないエンティティや関係を一般化する関係グラフ表現を学習することで、強力な帰納的伝達を実現します。この転送可能性を時間ナレッジ グラフ (TKG) に拡張することは、依然として課題です。既存の時間モデルは、パラメーターをデータセット固有のエンティティ、リレーション、またはタイムスタンプに結び付けており、素の語彙を持つ TKG に転送するように設計されていません。我々は、学習可能なパラメータを追加せず、クエリに対する時間ギャップに従って各エッジメッセージを回転させ、時間的に関連する信号を選択するためにクエリ条件付きゲートを適用することによって、学習可能なパラメータを追加せず、相対的な時間差を通じて時間をエンコードするエンティティ側のメッセージ関数である GRATE (Gated Rotary Attendee for Temporal Encoding) を提案します。 GRATE は、構造の伝達性を維持しながら、NBFNet スタイルの KG 基礎モデルに統合されます。既存の TKG ベンチマークは、共有のトレーニング/テスト語彙内で評価し、データセット間の時間転送を直接テストできません。したがって、内挿と外挿の両方にわたる、素のエンティティ、関係、およびタイムスタンプを含む帰納的転送ベンチマーク スイートである GDELTIndT および WIKIIndT を構築します。これらのベンチマークと保持された予測データセット全体で、単一の共同事前トレーニングされた GRATE チェックポイントは、ほとんどの設定で静的な基本モデルよりも改善されています。

原文 (English)

GRATE: Temporal Extensions for Inductive KG Foundation Models via Gated Rotary Attention

Knowledge graph foundation models such as Ultra and Trix achieve strong inductive transfer by learning relation-graph representations that generalise to unseen entities and relations. Extending this transferability to temporal knowledge graphs (TKGs) remains challenging: existing temporal models tie their parameters to dataset-specific entities, relations, or timestamps and are not designed to transfer to TKGs with disjoint vocabularies. We propose GRATE (Gated Rotary Attention for Temporal Encoding), an entity-side message function that adds no learnable parameters and encodes time through relative time differences by rotating each edge message according to its time gap to the query and applying a query-conditioned gate to select temporally relevant signals. GRATE integrates into NBFNet-style KG foundation models while preserving structural transferability. Existing TKG benchmarks evaluate within shared train/test vocabularies and cannot directly test cross-dataset temporal transfer; we therefore construct GDELTIndT and WIKIIndT, inductive transfer benchmark suites with disjoint entities, relations, and timestamps spanning both interpolation and extrapolation. Across these benchmarks and held-out forecasting datasets, a single jointly pretrained GRATE checkpoint improves over the static base model in most settings.

13:00 JSTLLM/生成AI

KGCQual: テキストからナレッジ グラフ構築の品質を評価するための解釈可能なフレームワーク

ナレッジ グラフ (KG) は、自動抽出パイプラインを通じて構築されることが増えています。ただし、そのようなシステムでは多くの場合、偽のトリプルや不完全なトリプルが導入され、ダウンストリームのパフォーマンスが低下します。既存の評価手法は、タスク固有の指標や小規模な手動検証に大きく依存しており、抽出されたグラフの構造的および意味論的な忠実性についての洞察は限られています。我々は、自動的に抽出されたグラフが、主要な名詞句、述語関係、ソーステキストで表現されている否定などの基本的な言語現象を捉えた「理想的な」グラフにどの程度近似しているかを測定する、本質的なKG品質評価のための新規で解釈可能な指標を提案します。私たちのフレームワークは、2 つの補完的なコンポーネントを統合しています。(1) 完全性、解決品質、接続性を評価するエンティティ レベルの評価、(2) 語彙の類似性、依存関係解析のアラインメント、およびセマンティックな忠実性を確保するための軽量否定処理を使用して述語の保存と多重性を判断する関係レベルの評価です。 WebNLG、TinyButMighty、BenchIE など、複数の最先端のトリプル抽出システムとデータセットにわたってメトリクスを評価し、既存のメトリクスでは見落とされていた欠落、冗長性、構造の逸脱を確実に特定できることを実証しました。私たちの研究は、自動化された KG 工法を比較するための、スケーラブルでモデルに依存しない解釈可能なフレームワークを提供し、標準化された評価の基盤を提供します。さらに、名詞と動詞のコンポーネントを分離するアブレーション研究と、KGCQual スコアが同じ抽出された KG でのリンク予測パフォーマンスと有意な相関があることを示す下流の評価を通じてメトリクスを検証します。コード リポジトリは https://github.com/kracr/kg-quality-metric で入手できます。

原文 (English)

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-specific metrics or small-scale manual verification, offering limited insight into the structural and semantic fidelity of extracted graphs. We propose a novel, interpretable metric for intrinsic KG quality assessment that measures how closely an automatically extracted graph approximates an "ideal" graph capturing the key noun phrases, predicate relations, and basic linguistic phenomena such as negation expressed in the source text. Our framework integrates two complementary components: (1) an entity-level assessment that evaluates completeness, resolution quality, and connectivity, and (2) a relation-level assessment that judges predicate preservation and multiplicity using lexical similarity, dependency-parse alignment, and light-weight negation handling to ensure semantic faithfulness. We evaluate our metric across multiple state-of-the-art triple extraction systems and datasets, including WebNLG, TinyButMighty, and BenchIE, demonstrating that it reliably identifies omissions, redundancy, and structural deviations that existing metrics overlook. Our work offers a scalable, model-agnostic, and interpretable framework for comparing automated KG construction methods and provides a foundation for standardised evaluation. We further validate the metric through an ablation study isolating noun and verb components, and a downstream evaluation showing that KGCQual scores correlate significantly with link prediction performance on the same extracted KGs. The code repository is available at https://github.com/kracr/kg-quality-metric.

13:00 JSTビジネス/資金調達Gemma

スパース機能の介入が実際にローカライズされるのはいつですか? SAEベースの安全管理の整合評価

スパース オートエンコーダ (SAE) 機能が安全関連動作のローカライズされた制御ハンドルとして機能する場合を評価します。この質問は難しい。なぜなら、見かけ上の成功は、弱い介入、不一致のベースライン、モデルの堅牢性、または意味のある有害なコンプライアンスを示さずに自動安全判定者が危険とマークする劣化した出力から生じる可能性があるからである。実行時の安全性介入のための整合コヒーレンスゲート評価プロトコルを導入します。手法は整合されたターゲット効果点で比較され、出力が安全でないと判断され、一貫性がある場合にのみ、主要なターゲットメトリクスが有害なコンプライアンスをカウントします。このプロトコルを Gemma Scope レイヤー 20 残存 SAE を使用した Gemma-2-9B-it の 3 つのプロンプト スプリットに適用すると、SAE 特徴アブレーションの有用な領域が狭いことがわかります。 SAE トップ 800 は、より低い総摂動と競争的有用性により、低から中程度の目標効果に達しますが、SAE トップ 1600 は、一致した高密度の拒否方向ベースラインと比較して有用性を失い、SAE トップ 3200 は主に一貫性崩壊を引き起こします。人間の監査により、コヒーレンス ゲーティングにより安全でないのみのアーティファクトが除去されることが確認され、機能診断により、有用なレジームは、活性化分離がランクとともに急速に減衰する拒否整合機能の安定したヘッドによって駆動されることが示されました。これらの結果は、SAE に基づく安全介入は一律に局所的であると想定するのではなく、体制依存の制御メカニズムとして評価されるべきであることを主張しています。

原文 (English)

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenerate outputs that automated safety judges mark as unsafe without representing meaningful harmful compliance. We introduce a matched coherence-gated evaluation protocol for runtime safety interventions: methods are compared at matched target-effect points, and the primary target metric counts harmful compliance only when an output is both judge-unsafe and coherent. Applying this protocol to three prompt splits on Gemma-2-9B-it with a Gemma Scope layer-20 residual SAE, we find that SAE feature ablation has a narrow useful regime. SAE top800 reaches a low-to-mid target effect with lower total perturbation and competitive utility, but SAE top1600 loses utility relative to a matched dense refusal-direction baseline, and SAE top3200 primarily induces coherence collapse. Human audit confirms that coherence gating removes unsafe-only artifacts, and feature diagnostics show that the useful regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank. These results argue that SAE-based safety interventions should be evaluated as regime-dependent control mechanisms rather than assumed to be uniformly localized.

13:00 JSTLLM/生成AI

大規模言語モデルにおけるリスクに敏感な意思決定の行動の特徴

大規模言語モデル (LLM) が意思決定支援に使用されることが増えているため、不確実性の下での選択が安定した解釈可能な動作規則性を示すかどうかを理解することが重要です。人間の意思決定は、比較的持続的なリスク選好と状況依存の調整を組み合わせていますが、類似の行動構造が LLM ベースの意思決定システムで観察できるかどうかは依然として不明です。ここでは、ノーリミット テキサス ホールデムに基づく制御されたマルチモデル フレームワークを使用して、この質問を検証します。このフレームワークでは、不確実な機会への自発的な関与を測定する参加性と、プリフロップのリスク エスカレーションを測定する積極性によって行動が定量化されます。同種のセルフプレイと異種混合モデルの相互作用にわたって、フロンティア LLM は安定したモデル固有のリスク プロファイルを示し、保守的な意思決定スタイルから積極的な意思決定スタイルまでのスペクトルを形成します。これらのプロファイルは、対戦相手の構成が変化してもほぼ堅牢なままですが、最も保守的なモデルと最も攻撃的なモデルは、混合設定ではさらに異なります。世界的なリスクのプレッシャーと個人のリソースの制約の下では、モデルは広範な行動の縮小から選択的なエスカレーション解除やほぼ不変の行動に至るまで、構造化されているが異質な方法で適応します。これらの調査結果は、LLM はベースラインのリスク傾向だけでなく、反応するリスク信号や調整の柔軟性も異なり、対話型の設定でリスクに敏感な意思決定を監査するための行動基盤を提供することを示唆しています。私たちのコードは https://github.com/XuankunRong/AgentTexasPoker で公開されています。

原文 (English)

Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models

As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities. Human decision-making combines relatively persistent risk preferences with context-dependent adjustment, yet it remains unclear whether analogous behavioural structure can be observed in LLM-based decision systems. Here we examine this question using a controlled multi-model framework based on no-limit Texas Hold'em, where behaviour is quantified by Participation, measuring voluntary engagement in uncertain opportunities, and Proactiveness, measuring pre-flop risk escalation. Across homogeneous self-play and heterogeneous mixed-model interactions, frontier LLMs exhibit stable, model-specific risk profiles, forming a spectrum from conservative to aggressive decision styles. These profiles remain largely robust under changing opponent composition, while the most conservative and most aggressive models diverge further in mixed settings. Under global risk pressure and personal resource constraint, models adapt in structured but heterogeneous ways, ranging from broad behavioural contraction to selective de-escalation and near-invariant behaviour. These findings suggest that LLMs differ not only in baseline risk disposition, but also in the risk signals they respond to and the flexibility with which they adjust, providing a behavioural basis for auditing risk-sensitive decision-making in interactive settings. Our code is publicly available at: https://github.com/XuankunRong/AgentTexasPoker.

13:00 JSTLLM/生成AIエージェント

エージェント臨床推論における大規模言語モデルの情報探索の失敗

大規模な言語モデルは医学知識の評価で高いスコアを達成しますが、臨床推論では不確実性の下で何を調査するかを積極的に決定する必要があります。私たちは、モデルが診断と治療計画に取り組む前に、連続する 3 ラウンドにわたる臨床データを積極的に要求する必要がある、血液腫瘍学における薬剤評価フレームワークを開発しました。 32 のフロンティア モデル全体で、最高の全体精度は 68% しか達成できませんでした。実際に要求された利用可能なデータの一部である情報利用率は、診断精度の最も強力な予測因子でしたが (R = 0.69、P < 0.001)、最終ラウンドでは利用率が 57% から 26% に崩壊し、治療選択に重要な分子および細胞遺伝学的データが検討されていないままになりました。推論トレースは、臨床推論ルーブリックで高いスコア (しきい値を 91% 上回りました) でしたが、精度とは無相関であり、局所的に一貫した理論的根拠と全体的に正しい結論との間にギャップがあることが明らかになりました。エラー分析により、検索満足、アンカリング、および早期終了が主要な失敗モードとして特定されました。これは、診断推論の二重プロセス モデルの下で初心者の臨床医を特徴付ける認知バイアスと同じです。これらの発見は、臨床腫瘍学における現在のモデルの主な制限は、不十分な医学知識ではなく、不確実性の下での情報探索の体系的な失敗であることを示しています。

原文 (English)

Information-seeking failures of large language models in agentic clinical reasoning

Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an agentic evaluation framework in hematologic oncology in which models must proactively request clinical data across three sequential rounds before committing to a diagnosis and treatment plan. Across 32 frontier models, the best achieved only 68% overall accuracy. Information utilization, the fraction of available data actually requested, was the strongest predictor of diagnostic accuracy (R = 0.69, P < 0.001), yet utilization collapsed from 57% to 26% in the final round, leaving molecular and cytogenetic data critical for treatment selection unexamined. Reasoning traces scored high on a clinical reasoning rubric (91% above threshold) but decorrelated from accuracy, revealing a gap between locally coherent rationales and globally correct conclusions. Error analysis identified search satisficing, anchoring and premature closure as the dominant failure modes, the same cognitive biases that characterize novice clinicians under dual-process models of diagnostic reasoning. These findings demonstrate that the primary limitation of current models in clinical oncology is not insufficient medical knowledge but a systematic failure of information-seeking under uncertainty.

13:00 JSTLLM/生成AIエージェントDeepSeek

エージェント取引システムは自分自身のインテリジェンスを支払うことができますか?

大規模言語モデル (LLM) エージェントは取引システムでますます使用されており、モデル推論、ツールの使用、および継続的な意思決定により、取引価値を生み出すことが期待されるコストが発生します。既存の評価では通常、パフォーマンス指標が報告されますが、LLM を介した動的な意思決定が誘発コストを測定可能な増分利益に変換するかどうかなど、エージェントの実行可能性を検証することはほとんどありません。この基準を適用するために、取引記録、実行時トレース、展開構成からエージェント取引システムを評価するためのトレースベースの診断ツールキットである TradeLens を導入します。取引の軌跡を再構築し、利益とコストを解釈可能な証拠に帰属させ、エージェントが自らのインテリジェンスにお金を払うかどうか、またその理由を診断します。私たちは、導入に関する議論とともに、バックボーンモデル、資本規模、取引頻度、システムアーキテクチャにわたる広範な分析を実施します。私たちの結果は、実行可能性はインテリジェンスから利益への変換にかかっていることを示しています。モデルは、DeepSeek-V3.2 の貧弱な資産選択や GLM-4.7 のネガティブなタイミングなど、さまざまな失敗パターンを示しますが、資本規模、取引頻度、アーキテクチャは、意思決定に起因するタイミング値を増幅または低下させることによってのみ問題になります。これらの調査結果は、LLM ベースの取引エージェントの評価を、能力中心のパフォーマンス ランキングから、インテリジェンスから利益への変換のトレースに基づいた診断へと再構築します。私たちのコードは https://anonymous.4open.science/r/TradeLens で入手できます。

原文 (English)

Can Agentic Trading Systems Pay for Their Own Intelligence?

Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, but rarely examine agentic viability: whether dynamic LLM-mediated decisions convert their induced costs into measurable incremental profit. To apply this criterion, we introduce TradeLens, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses whether and why an agent pays for its own intelligence. We conduct extensive analysis across backbone models, capital scales, trading frequencies, and system architectures, together with deployment discussion. Our results show that viability hinges on intelligence-to-profit conversion: models exhibit different failure patterns, such as poor asset selection in DeepSeek-V3.2 and negative timing in GLM-4.7, while capital scale, trading frequency, and architecture matter only by amplifying or degrading decision-attributed timing value. These findings reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion. Our code is available at https://anonymous.4open.science/r/TradeLens.

13:00 JSTLLM/生成AI

SPARK: 大規模言語モデルにおける潜在推論状態の感受性ガイドに基づくプロファイリングとステアリング

大規模言語モデル (LLM) での推論の失敗は通常、最終的な答えから評価されますが、間違った答えではモデルが失敗した理由はわかりません。同じ誤った出力は、欠落している機能、不安定な推論軌道、またはフリーズされたモデルですでに利用可能な推論状態のアクティブ化の失敗を反映している可能性があります。既存のプロンプトおよびベンチマークベースの評価方法は、ほとんどが出力レベルで動作しますが、一般的なアクティベーションステアリング方法は、通常、どの例に介入が必要かを診断することなく、グローバルな指示を適用します。この論文では、隠れ状態応答を使用してモデルが内部的に有効な推論状態に入っているかどうかを診断し、軽量のテスト時間ステアリングをガイドする SPARK を紹介します。重要な観察は、生の隠れ状態の感受性がプロンプトの長さによって大きく混乱するということです。特に、より困難なシリアル化されたインスタンスが自然に長くなるプログラムおよびアルゴリズムの推論ではそうです。したがって、SPARK は、長さ制御された感受性を使用して、入力スケールの影響を残余推論のアクティブ化から分離し、この信号を層間調整と組み合わせて、推論がアクティブなアンカーとアクティブ化が不十分なハード サンプルを選択します。私たちは、潜在プロファイリングと難易度認識分析のための制御されたプログラム推論スイートとして FRONTIER-4.5K を使用し、順方向のみのベンチマーク プロファイリングで GSM8K および MATH-500 上の SPARK-Steering を評価します。私たちの方法は Qwen3 シリーズ モデルを一貫して改善します。 MATH-500 では、精度は Qwen3-4B で 82.0% から 84.6% に、Qwen3-8B で 82.4% から 85.6% に向上しました。これらの結果は、感受性が推論の失敗の診断信号としてだけでなく、対象を絞ったテスト時間介入の実用的なガイドとしても機能することを示唆しています。

原文 (English)

SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable reasoning trajectory, or a failure to activate a reasoning state that is already available in the frozen model. Existing prompting and benchmark-based evaluation methods mostly operate at the output level, while generic activation-steering methods typically apply global directions without diagnosing which examples require intervention. In this paper, we introduce SPARK, which uses hidden-state response to diagnose whether a model internally enters an effective reasoning state and to guide lightweight test-time steering. The key observation is that raw hidden-state susceptibility is strongly confounded by prompt length, especially in programmatic and algorithmic reasoning where harder serialized instances naturally become longer. SPARK therefore uses length-controlled susceptibility to separate input-scale effects from residual reasoning activation, and combines this signal with cross-layer coordination to select reasoning-active anchors and under-activated hard examples. We use FRONTIER-4.5K as a controlled programmatic reasoning suite for latent profiling and difficulty-aware analysis, and evaluate SPARK-Steering on GSM8K and MATH-500 with forward-only benchmark profiling. Our method improves Qwen3 series models consistently; on MATH-500, accuracy rises from 82.0% to 84.6% for Qwen3-4B and from 82.4% to 85.6% for Qwen3-8B. These results suggest that susceptibility can serve not only as a diagnostic signal for reasoning failures, but also as a practical guide for targeted test-time intervention.

13:00 JSTエージェント研究/論文

Sim と Real のギャップを測定する: AIoT システムにおける強化学習のための手頃な価格の現実世界のベンチマーク プラットフォームの設計

強化学習 (RL) は、自律型モノのインターネット (AIoT) などの自律型システムのパフォーマンスを強化するために一般的に使用されます。ただし、RL は試行錯誤を繰り返す性質があるため、現実の環境で実施するとコストがかかり、シナリオによっては危険が伴います。したがって、RL 研究の大部分はシミュレーションで行われます。この依存により、Sim-to-Real の転送可能性に関連する課題が生じます。 Sim-to-Real アルゴリズムの堅牢性と Sim-to-Real ギャップを評価することは、現実世界での RL パフォーマンスの向上を目的とした研究の重要な前提条件です。したがって、ロボット工学などの業界は、この研究を促進するために、同時シミュレーションと物理プラットフォームを開発してきました。ただし、AIoT 用の汎用 Sim-to-Real ベンチマーク プラットフォームは現在存在しません。これらの懸念に対処するために、私たちは AIoT における RL を研究するための実世界の AIoT プラットフォームを開発しました。このプラットフォームでは、エッジ デバイスに配置されたエージェントが、ハードウェア エミュレートされたキーボードを介して、ビジョン入力に基づいて別のホスト コンピューターでビデオ ゲームをプレイします。このプラットフォームは、2 台のコンピューターとともに、400 ドル未満の市販コンポーネントを使用します。このシステムの目的はゲーム スコアの最大化であるため、現実世界の RL 展開に伴う安全性のリスクを本質的に軽減します。実験結果では、シミュレーションでトレーニングされたエージェントは、現実世界への展開後に人間レベルのパフォーマンスと比較して 1160% のパフォーマンス低下が見られ、Sim と Real の大きなギャップが示されています。ディープ Q ネットワーク (DQN) アルゴリズムを使用した現実世界の直接トレーニングは、1,000 万回のトレーニング ステップ後に人間レベルのパフォーマンスの約 38% を達成し、現実世界の条件下での RL の実現可能性を示しています。これらの結果は、提案された Sim-to-Real ベンチマーク プラットフォームが、現実世界の AIoT システムにおける RL の定性的および定量的評価のための実質的な基盤を提供することを示唆しています。

原文 (English)

Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems

Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environments, is costly and hazardous in some scenarios. Consequently, the majority of RL research is conducted in simulation. This reliance introduces challenges related to the Sim-to-Real transferability. Evaluating the Sim-to-Real algorithmic robustness and the Sim-to-Real gap is a critical prerequisite for research aimed at improving RL performance in the real world. Therefore, industries such as robotics have developed concurrent simulation and physical platforms to facilitate this research. However, a universal Sim-to-Real benchmark platform for AIoT does not currently exist. To address these concerns, we developed a real-world AIoT platform for studying RL in AIoT. On this platform, an agent deployed on an edge device plays video games on a separate host computer via a hardware-emulated keyboard, guided by vision input. This platform uses commercially available components costing less than USD 400, together with two computers. Because the system's objective is game score maximization, it inherently mitigates safety risks associated with real-world RL deployments. Experimental results show the simulation-trained agent suffers a 1160% performance degradation relative to the human-level performance after real-world deployment, indicating a significant Sim-to-Real gap. Direct real-world training using the deep Q-network (DQN) algorithm achieves approximately 38% of human-level performance after 10 million training steps, demonstrating the feasibility of RL under real-world conditions. These results suggest that the proposed Sim-to-Real benchmark platform provides a substantial foundation for qualitative and quantitative evaluations of RL in real-world AIoT systems.

13:00 JST研究/論文

社会技術的設計原則と人間中心AIのガイドラインの比較

人間中心 AI (HCAI) は、倫理指向のシステム設計を目的としたガイドラインまたは原則を指します。我々は、HCAI ガイドラインを、従来の情報技術の文脈で出現した社会技術システムの原則と比較します。この比較は、AI 利用の側面を含めることによる社会技術ヒューリスティックの改訂につながります。この比較により、継続的な進化が社会技術システムの基本的な特徴であり、自律性が協調的に発揮される場合には、人間の監視や介入とその後の AI システムの流用がシステムの継続的な適応と再設計につながることが明らかになります。社会技術的な観点から見ると、透明性という重要な要件は、技術的な機能によって満たされるだけでなく、人間の主体を含むシステム全体の貢献によっても満たされる必要があります。技術的特徴だけでなく、組織的および社会的慣行が AI の欠点を補う形で社会技術的に設計されていれば、AI の使用は有望となるでしょう。

原文 (English)

Comparing Socio-technical Design Principles with Guidelines for Human-centered AI

Human-centered AI (HCAI) refers to guidelines or principles that aim on ethi-cally oriented design of systems. We compare HCAI- guidelines with princi-ples of socio-technical systems that emerged in the context of conventional in-formation technology. The comparison leads to a revision of socio-technical heuristics by including aspects of AI-usage. The comparison reveals that con-tinuous evolution is a basic characteristic of socio-technical systems, and that human oversight or interventions and the subsequent appropriation of AI-systems lead to continuous adaptation and re-design of the systems, if autono-my is collaboratively exercised. From a socio-technical point of view, the cru-cial requirement of transparency has not only to be fulfilled with technical fea-tures, but also by contributions of the whole system including human actors. It will be promising for using AI, if not only technical features, but organization-al and social practices are socio-technically designed in a way that compen-sates shortcomings of AI.

13:00 JSTエージェントロボティクス

ABot-AgentOS: 生涯にわたるマルチモーダル メモリを備えた汎用ロボット エージェント OS

最近の VLM および VLA システムでは、ロボットの認識と動作予測が改善されていますが、長期的に具現化されたエージェントは、依然として、推論、メモリ、ツールの使用、検証、およびクロス具現化実行のための一般的なランタイム層を必要としています。 ABot-AgentOS は、低レベルのコントローラーの上に位置し、シーンに応じたプランニング、コンテキスト分離されたスキルの実行、多段階の検証、マルチモーダル メモリ、エッジとクラウドのコラボレーションのための熟慮型エージェント層を提供する、一般的なロボット エージェント オペレーティング システムです。このようなシステムを評価するために、16 の屋内、屋外、ハイブリッド シーン、4 つの難易度レベル、およびナビゲーション、オブジェクト検索、NPC ダイアログ、動的イベント、およびトレースベースのスコアリングを含む 200 以上のタスクを備えた実行可能なベンチマークである EmbodiedWorldBench を導入します。 ABot-AgentOS はさらに、ダイアログ、視覚的観察、空間コンテキスト、時間的関係、およびタスク トレースを型付きノードとエッジに変換する永続的なソース接地基板であるユニバーサル マルチモーダル グラフ メモリを導入します。障害駆動型の自己進化ループは、診断されたメモリ障害を、後の評価分割にのみ昇格するゲート付きランタイム evo アセットに変換し、継続的な改善を可能にしながら、電流分割のグラウンド トゥルースの漏洩を防ぎます。初期の EmbodiedWorldBench サブセットでは、ABot-AgentOS はタスクの成功と目標の完了の両方で単一コントローラーのベースラインを上回ります。メモリ ベンチマーク全体で、ABot-AgentOS Static は LoCoMo で 87.5、OpenEQA EM-EQA で 59.9、Mem-Gallery で 88.6、NExT-QA で 76.5 Acc@All を達成しました。自己進化により、LoCoMo は 88.7、OpenEQA は 60.4、Mem-Gallery は 89.0 にさらに向上しました。これらの結果は、一般的なエージェント OS レイヤーが、継続的な対話のための永続的で監査可能なメモリを提供しながら、長期的な具体化された実行を改善できることを示唆しています。

原文 (English)

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.

13:00 JST研究/論文

Co4ICF: 共進化する物理学に基づいたサロゲートと慣性閉じ込め融合のための RL ベースのパルス オプティマイザー

慣性閉じ込め核融合 (ICF) のオフライン トレーニングされたサロゲートは、反復オプティマイザーが入力を分布外 (OOD) 領域に送り込み、予測が信頼できなくなるというよく知られた障害モードに悩まされます。ここでは、物理学に基づいたサロゲートと PPO ベースのパルス オプティマイザーを結合する、共進化するフレームワークである Co4ICF を紹介します。サロゲートは、ポリシーに起因する軌道に基づいて繰り返し微調整され、オプティマイザーが入力分布をシフトするときに外挿誤差を修正します。オプティマイザーは、この進化するサロゲートを高速環境としてクエリします。 1D MULTI 環境では、Co4ICF は現在のレーザー設計ベースラインに基づいて 146.1% の正規化収率を達成します。ポストホッククロスフィデリティチェックとして、2D トレーニングや微調整を行わずに 2D-MULTI で直接評価すると、最適化されたパルスはさらに 246.9% の正規化収率を達成します。予算に見合ったアブレーションは、その利益が追加のシミュレーション データだけによって説明されるものではなく、重要な役割を果たしている共進化メカニズムと一致していることを裏付けています。今後のベンチマークをサポートするために、大規模な MULTI-IFE シミュレーション データセットをリリースします。

原文 (English)

Co4ICF: Co-evolving Physics-Informed Surrogate and RL-based Pulse Optimizer for Inertial Confinement Fusion

Offline-trained surrogates for Inertial Confinement Fusion (ICF) suffer a well-known failure mode that iterative optimizers drive inputs into out-of-distribution (OOD) regions where predictions become unreliable. Here we present Co4ICF, a co-evolving framework that couples a physics-informed surrogate with a PPO-based pulse optimizer. The surrogate is iteratively fine-tuned on policy-induced trajectories, correcting extrapolation errors as the optimizer shifts the input distribution; the optimizer queries this evolving surrogate as a fast environment. In the 1D MULTI environment, Co4ICF achieves 146.1% normalized yield based on current laser design baseline; as a post-hoc cross-fidelity check, the optimized pulse further attains 246.9% normalized yield when directly evaluated in 2D-MULTI without any 2D training or fine-tuning. Budget-matched ablations support that the gains are not explained solely by additional simulation data and are consistent with the co-evolving mechanism playing a key role. We release a large-scale MULTI-IFE simulation dataset to support future benchmarking.

13:00 JSTLLM/生成AIエージェント

アンカー: 現実世界の危害に関する CLI エージェントの自動調整監査

自律型 CLI エージェントは、人間による監視を最小限に抑えながら、コードの作成、シェル コマンドの実行、Web の閲覧、クラウド インフラストラクチャの管理など、数時間にわたるセッションにわたって何百ものアクションを実行できるようになりました。自律性が高まるとリスクも増大するのでしょうか?米国の公判に基づいた違法なタスクに関して CLI エージェントにストレス テストを行う自動監査フレームワークである ANCHOR を紹介します。 ANCHOR は、監視付き強化微調整を使用してダークパーソナリティ データに基づいて微調整された監査エージェントを展開します。この監査者は、タスクを分解し、拒否された場合に要求を再構成し、複数ターンの対話にわたって戦略を適応させる、永続的な悪意のあるユーザーの役割を果たします。フロンティア CLI エージェントを評価すると、直接要求された場合は違法なタスクを拒否することがよくありますが、継続的な悪意のあるインタラクションではコンプライアンスが 100% に達することがわかりました。エージェントが従う場合、ユーザーの要求を超えることが多く、大規模な金融詐欺や生物兵器の開発などの壊滅的なリスクシナリオを含む大規模な危害に備えたインフラストラクチャを自律的に構築します。これらの発見は、現在の調整技術が自律エージェントには不十分であることを示しており、持続的で適応的な悪意のあるユーザーに対する安全性評価の必要性を強調しています。 https://github.com/garified/anchor で ANCHOR をリリースします

原文 (English)

ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm

Autonomous CLI agents can now execute hundreds of actions across multi-hour sessions: writing code, executing shell commands, browsing the web, and managing cloud infrastructure, all with minimal human oversight. Does greater autonomy invite greater risk? We introduce ANCHOR, an automated auditing framework that stress-tests CLI agents on illegal tasks grounded in public US court cases. ANCHOR deploys an auditor agent fine-tuned on dark personality data using supervised and reinforcement fine tuning. This auditor roleplays persistent malicious users who decompose tasks, reframe requests upon refusal, and adapt strategies across multi-turn interactions. Evaluating frontier CLI agents, we find that while they often refuse illegal tasks when prompted directly, compliance reaches 100\% under persistent malicious interaction. When agents comply, they frequently exceed user requests, autonomously building infrastructure for large-scale harm, including catastrophic risk scenarios such as large-scale financial fraud and bioweapon development. These findings demonstrate that current alignment techniques are insufficient for autonomous agents and underscore the need for safety evaluations against persistent, adaptive malicious users. We release ANCHOR at https://github.com/garified/anchor

13:00 JSTエージェント

GRASP: Agentic RAG の粒度を考慮した検索ポリシー

エージェント的検索拡張生成 (RAG) は、言語モデルが反復的に推論し、検索クエリを生成し、証拠を取得し、回答を予測できるようにすることで、静的 RAG を拡張します。ただし、モデルにとって、いつ取得するか、字句一致と意味的類似性のどちらを使用するか、無関係なトークンがエージェントの推論に干渉しないようにコンテキストの粒度を制御する方法を決定するのは依然として困難です。この論文では、複数ステップの推論中に相補的な検索ツールを適応的に調整するエージェントをトレーニングするための強化学習 (RL) フレームワークである GRASP を紹介します。 GRASP は、エージェントにセマンティック検索、キーワード検索、段落読み取りアクションを提供し、必要な場合にのみ文レベルの証拠を取得し、さらにコンテキストを拡張できるようにします。回答の精度、根拠のある読み取り、補完的な検索、ターン効率を総合的に考慮した報酬を使用してポリシーをトレーニングします。マルチホップ推論ベンチマークの実験では、GRASP が、シングルステップ検索、プロンプトベースのエージェント RAG、および RL ベースの検索ベースラインと比較して、検索再現率と下流の質問応答パフォーマンスの両方を向上させることが示されています。定性分析とアブレーション分析は、学習されたポリシーが解釈可能なスキミングおよびスキャン動作を開発することを示しています。つまり、広範な探索にはセマンティック検索、局所的な検証には段落読み取り、およびエンティティ固有の証拠にはキーワード検索が使用されます。これらの結果は、エージェントの正しい推論には、検索シグナルとコンテキストの粒度を調整する方法を学習することが重要であることを示唆しています。

原文 (English)

GRASP: GRanularity-Aware Search Policy for Agentic RAG

Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matching or semantic similarity, and how to control context granularity to prevent irrelevant tokens from interfering with agent reasoning. In this paper, we introduce GRASP, a reinforcement learning (RL) framework for training agents to adaptively coordinate complementary retrieval tools during multi-step reasoning. GRASP provides the agent with semantic search, keyword search, and paragraph-reading actions, enabling it to retrieve sentence-level evidence and expand further context only when needed. We train the policy with a reward that jointly accounts for answer accuracy, grounded reading, complementary search, and turn efficiency. Experiments on multi-hop reasoning benchmarks show that GRASP improves both retrieval recall and downstream question answering performance compared with single-step retrieval, prompting-based agentic RAG, and RL-based retrieval baselines. Qualitative and ablation analyses show that the learned policy develops interpretable skimming and scanning behavior: it uses semantic search for broad exploration, paragraph reading for local verification, and keyword search for entity-specific evidence. These results suggest that learning to coordinate retrieval signals and context granularity is critical for agent's correct reasoning.

13:00 JSTエージェント研究/論文

エージェントはただ同意するだけではなく、覚えている: ステートフル個人エージェントにおける永続的な媚びのベンチマーク

ステートフル パーソナル エージェントは、長期的なユーザー プロファイル、エピソード記憶、再利用可能なスキルを維持することがますます増えています。この持続性により、会話のお調子者が​​状態記述の失敗に変わります。受け入れられたユーザー中心の主張は、永続的な設定、背景事実、またはワークフローとしてコミットされ、元の会話が終わった後に再利用される可能性があります。私たちはこれを永続的なお調子者と呼び、パーソナル エージェント お調子者ベンチマーク (PASB) を導入します。これは、会話による要求が受け入れられ、永続的なエージェント状態に書き込まれ、後の中立的なクエリで再利用されるかどうかを追跡する 1,600 タスクのベンチマークです。事前に書き込まれたメモリを提供する以前のベンチマークとは異なり、PASB は何を保存するかを決定する実際のエージェント (Hermes-Agent および OpenClaw) を評価します。 4 つのシナリオ フレームと 4 つの時間配信パターンを組み合わせ、5 ターンの永続ステージをクリアされた 3 ターンのクエリ ステージから分離することで書き込みプロセスを分離し、ダウンストリームの影響が永続状態からのみ発生するようにします。 12 のモデル全体で、コミット境界が重要な変曲点です。ダウンストリームの障害は、セッションのみのエピソードの 45.0% からコミット後は 71.9% に増加し、一貫して 27.0 パーセント増加しています。コミットされたクレームには、ステータスの昇格、帰属の削除、範囲の拡大という 3 つの書き込み時間パターンが見られます。これらのパターンは、記憶に似たフレーミングや手順的なフレーミング、繰り返しの強化、さらにはドメインの境界を越えた場合でもより強力になります。これらの結果は、エージェントのおべっかが根本的に国家の統治問題であることを示している。ユーザー コンテンツが永続メモリに保存されると、エージェントの発言だけでなく、エージェントが書き込む内容も安全性によって管理されなければなりません。 PASB は、応答レベルの軽減策を超えて保存されたコンテンツのソース、役割、および範囲を維持しながら、危険なコミットをゲートするために必要な書き込み時の制御を特定します。

原文 (English)

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents

Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be committed as lasting preferences, background facts, or workflows and later reused after the original conversation is gone. We call this persistent sycophancy and introduce the Personal Agent Sycophancy Benchmark (PASB), a 1,600-task benchmark that traces whether a conversational claim is accepted, written into durable agent state, and reused in a later neutral query. Unlike prior benchmarks that provide pre-written memories, PASB evaluates real agents (Hermes-Agent and OpenClaw) that decide what to store. It isolates the write process by combining four scenario framings with four temporal delivery patterns and separating a five-turn persist stage from a cleared three-turn query stage, ensuring downstream effects arise only from durable state. Across twelve models, the commit boundary is the key inflection point: downstream failure increases from 45.0% in session-only episodes to 71.9% after commitment, a consistent increase of 27.0 percentage points. Committed claims exhibit three write-time patterns: status promotion, attribution removal, and scope broadening. These patterns become stronger under memory-like or procedural framing, repeated reinforcement, and even across domain boundaries. These results show that agent sycophancy is fundamentally a state-writing governance problem. Once user content is committed to durable memory, safety must govern what agents write, not only what they say. PASB identifies the write-time controls needed to gate risky commits while preserving the source, role, and scope of stored content beyond response-level mitigations.

13:00 JSTLLM/生成AIエージェント

エージェント スキルにおける層間の不整合の検出: 漸進的な読み込みを意識した対照学習アプローチ

大規模言語モデル (LLM) エージェントは、自然言語メタデータ、手順説明、実行時リソースをパッケージ化した再利用可能なアーティファクトであるエージェント スキルを通じてますます拡張されています。オープンソースのスキル市場が拡大するにつれて、ユーザーとエージェントはサードパーティのスキルを選択する際に短いメタデータにますます依存するようになり、スキルの説明と実際の動作の間の不一致、つまりクロスレイヤーの不整合と呼ばれる問題を検出することが困難になっています。この問題に対処するために、エージェント スキルの階層構造をモデル化し、階層間の一貫性を学習することで不整合を検出する LLM ベースのフレームワークである、Progressive Loading-Aware Hierarchical Contrastive Learning (PL-HCL) を提案します。 PL-HCL は、264,000 を超えるオープンソース スキルの正規化されたコーパスと人間が検証したチャレンジ セットを使用して、Macro-F1 を未適応ベースラインの約 0.45 から、評価された LLM バックボーン全体で 0.87 ~ 0.89 に改善します。このアプローチは、ユーザーとオペレーターに効果的なスクリーニング ツールを提供するとともに、階層化されたデジタル アーティファクトの不一致を検出するための設計原則を提供します。

原文 (English)

Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach

Large language model (LLM) agents are increasingly extended through Agent Skills, reusable artifacts that package natural-language metadata, procedural instructions, and execution-time resources for runtime use. As open-source skill marketplaces expand, users and agents increasingly rely on brief metadata to select third-party skills, making it difficult to detect inconsistencies between a skill's description and its true behavior, a problem we call cross-layer misalignment. To address this issue, we propose Progressive Loading-Aware Hierarchical Contrastive Learning (PL-HCL), an LLM-based framework that detects misalignment by modeling the layered structure of Agent Skills and learning cross-layer consistency. Using a normalized corpus of over 264,000 open-source skills and a human-verified challenge set, PL-HCL improves Macro-F1 from approximately 0.45 for unadapted baselines to 0.87-0.89 across evaluated LLM backbones. This approach offers an effective screening tool for users and operators, as well as design principles for detecting inconsistencies in layered digital artifacts.

13:00 JSTLLM/生成AI

AI YOU Town: デジタル ツインで友達を作り、お金を稼ぎましょう

ユーザーの特性を推測し、ペルソナに一致する応答を生成する既存のアプローチは、静的なプロンプトに依存しています。それらは調整された不確実性を欠き、逐次的な証拠を無視し、長い相互作用中にドリフトします。私たちは \textbf{AI YOU} を紹介します。これは、会話から 22 の次元でパーソナリティ プロファイルを継続的に更新し、それを個人のデジタル ツインに具体化するフレームワークです。実際には、このシステムは、ペルソナ推論のためにプロンプ​​ト、ベイジアン更新、および等角予測を組み合わせています。定期的に更新される記憶アンカーと 3 つの層による認知記憶により、長い対話を通じてペルソナの一貫性が維持されます。主な結果全体で、AI YOU \emph{(i)} は 0.921 から 0.976 の範囲の等角カバレッジを達成し、\emph{(ii)} は不確実性のキャリブレーションとメモリに基づいた推論を向上させ、\emph{(iii)} は 100 ターン以上のロールプレイングで特性のドリフトを軽減しながら静的なプロンプトよりもペルソナの忠実度を高め、複数のエージェントによる敵対的設定の下で評価されたバックボーンのほとんどを対象にしました。プロトタイプ \emph{AI YOU Town} は、将来のインタラクションのために想像力豊かな双子の世界を初期化します。オンライン デモは \href{https://quinnnnnne-ai-you.hf.space/}{\mbox{\texttt{quinnnnnne-ai-you.hf.space}}} で利用できます。

原文 (English)

AI YOU Town: Make Friends and Money with Your Digital Twin

Existing approaches to infer user traits and generate responses consistent with a persona rely on static prompting. They lack calibrated uncertainty, ignore sequential evidence, and drift during long interactions. We present \textbf{AI YOU}, a framework that continually updates a personality profile with 22 dimensions from conversation and embodies it in a personal digital twin. Practically, the system combines prompting, Bayesian updating, and conformal prediction for persona inference. A periodically refreshed memory anchor and cognitive memory with three layers preserve persona consistency over long interactions. Across the main results, AI YOU \emph{(i)} achieves conformal coverage ranging from 0.921 to 0.976, \emph{(ii)} improves uncertainty calibration and reasoning grounded in memory, and \emph{(iii)} enhances persona fidelity over static prompting in role playing over 100 turns while reducing trait drift, for most evaluated backbones under adversarial settings with multiple agents. The prototype \emph{AI YOU Town} initializes an imaginative twin world for future interaction. The online demo is available at \href{https://quinnnnnne-ai-you.hf.space/}{\mbox{\texttt{quinnnnnne-ai-you.hf.space}}}.

13:00 JSTエージェント

大規模な言語モデル エージェントにより、ガス分離のための金属-有機フレームワークの逆設計が加速されます。

有機金属フレームワーク (MOF) は、吸着ガス分離のための高度にモジュール化されたプラットフォームを提供しますが、その広大な網状設計スペースにより、化学的妥当性、分離性能、構造多様性の同時制約の下では逆設計が困難になります。ここでは、MOFid 空間におけるガス分離 MOF の閉ループ逆設計のための大規模言語モデル エージェント フレームワークである LEMO Agent を紹介します。 LEMO Agent は、言語ベースの候補生成と MOFid 標準化、明示的な妥当性チェック、Transformer ベースのプロパティ予測、構造化設計メモリ、およびマルチアイランド探索を組み合わせます。生成、検証、評価、記憶の繰り返しサイクルを通じて、エージェントは成功した候補と失敗した候補の両方からのフィードバックを使用して、リンカー、金属、およびトポロジーの選択全体にわたって化学的に制約された検索をガイドします。 CH$_4$/N$_2$ および CO$_2$/N$_2$ 分離タスクで LEMO エージェントを評価します。代表的な生成ベースライン、最適化ベースライン、および薬剤ベースラインと比較して、LEMO Agent は高パフォーマンスの候補を強化し、予測される分離パフォーマンスを向上させ、広範な化学的およびトポロジー的多様性を維持します。選択された候補はさらに再構築され、GCMC シミュレーションによって評価され、化学的実現可能性とリガンドの購入可能性に基づいた実験的なダウンセレクション ワークフローを経て、最初のウェットラボ合成と SEM 特性評価につながります。これらの結果は、大規模言語モデル エージェントが、従来の固定ライブラリのスクリーニングを超えて MOF の発見を加速するための、解釈可能でスケーラブルな設計エンジンとして機能できることを示しています。

原文 (English)

Large language model agents accelerate inverse design of metal-organic frameworks for gas separation

Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes inverse design difficult under simultaneous constraints of chemical validity, separation performance, and structural diversity. Here, we present LEMO Agent, a large-language-model agent framework for closed-loop inverse design of gas-separation MOFs in MOFid space. LEMO Agent couples language-based candidate generation with MOFid standardization, explicit validity checking, Transformer-based property prediction, structured design memory, and multi-island exploration. Through iterative generate--validate--evaluate--remember cycles, the agent uses feedback from both successful and failed candidates to guide chemically constrained search across linker, metal, and topology choices. We evaluate LEMO Agent on CH$_4$/N$_2$ and CO$_2$/N$_2$ separation tasks. Compared with representative generative, optimization, and agentic baselines, LEMO Agent enriches high-performing candidates, improves predicted separation performance, and maintains broad chemical and topological diversity. Selected candidates are further reconstructed, evaluated by GCMC simulations, and passed through an experimental down-selection workflow based on chemical feasibility and ligand purchasability, leading to initial wet-lab synthesis and SEM characterization. These results demonstrate that large language model agents can serve as interpretable and scalable design engines for accelerating MOF discovery beyond conventional fixed-library screening.

13:00 JST研究/論文

CRiT-QA: 反事実チェーンとディストラクター トラップを使用したマルチホップ推論の評価

大規模な言語モデルのマルチホップ推論機能を評価することは、依然として大きな課題です。現在のモデルは既存のマルチホップ質問応答データセットで優れた結果を達成していますが、そのようなパフォーマンスは 2 つの重大な脆弱性を覆い隠してしまうことがよくあります。(1) 提供されたコンテキストへの準拠ではなく、内部パラメトリック知識への依存、および (2) 複数の文書にわたる真の証拠の集約の必要性を低下させる、単一文書のキューやタイプマッチングなどのデータセットのショートカットの悪用です。両方の制限に対処するために明示的に設計されたデータセットである CRiT-QA (Counterfactual Reasoning with Traps) を紹介します。記憶された知識への依存を無効化し、厳密なコンテキスト依存性を強制するために、CRiT-QA は事実に基づく推論チェーンを反事実エンティティで変換します。さらに、マルチアンカーディストラクタチェーン、もっともらしいが異なるホップで分岐する推論パスを注入します。これらのトラップでは、モデルが浅いヒューリスティックを利用するのではなく、推論プロセス全体に従う必要があります。私たちの実験では、LLM は標準データセットと比較して CRiT-QA で大幅なパフォーマンス低下を示し、反事実条件や注意散漫の罠に対する脆弱性をさらしていることが示されました。したがって、CRiT-QA は、本物のマルチホップ推論を評価するための厳密な診断ツールとして機能し、より信頼性が高く、証拠に基づいた LLM を開発するための基盤を提供します。

原文 (English)

CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks two critical vulnerabilities: (1) reliance on internal parametric knowledge rather than adherence to the provided context, and (2) exploitation of dataset shortcuts, such as single-document cues or type-matching, that diminish the need for genuine evidence aggregation across multiple documents. We introduce CRiT-QA (Counterfactual Reasoning with Traps), a dataset explicitly designed to address both limitations. To neutralize reliance on memorized knowledge and enforce strict context dependency, CRiT-QA transforms factual reasoning chains with counterfactual entities. Furthermore, it injects multi-anchor distractor chains, plausible but incorrect reasoning paths that diverge at different hops. These traps require models to follow the entire reasoning process rather than exploiting shallow heuristics. Our experiments show that LLMs exhibit substantial performance degradation on CRiT-QA compared to standard datasets, exposing their vulnerability to counterfactual conditions and distractor traps. CRiT-QA thus serves as a rigorous diagnostic tool for evaluating genuine multi-hop reasoning and provides a foundation for developing more reliable, evidence-grounded LLMs.

13:00 JSTLLM/生成AI

大規模な言語モデルを解釈するためのラゲール幾何学

既存の仮説は、LLM の概念を単一の点、線形方向、またはガウス クラスターとして表しますが、そのような構造がどのように、そしてなぜ現れるのかは依然として不明です。ここでは、概念幾何学がラゲール幾何学によって正確に特徴付けられることを示します。ラゲール幾何学では、概念が領域 (ラゲール ボロノイ セルまたはセルの結合) として定義され、概念を厳密に定義、測定、分離できるようになります。この定式化に基づいて、包含や階層などのよりきめの細かい概念構造がラゲール重みによって自然に明らかにされることを示します。次に、このジオメトリをトランス内部に押し込みます。各層を区分線形演算子に分解すると、トークンの隠れた軌道が 2 つの結合されたメカニズムによって支配されることがわかります。1 つは自己完結型の区分線形フローの静的ツリー、もう 1 つはトークン間アテンションが発生したときにツリー間で軌道をホップする動的トランスポートです。この分解により、隠れベクトルが任意の層でエンコードした正確な概念を読み出すためのトレーニング不要、ハイパーパラメータ不要の方法である幾何レンズが得られます。また、意思決定ジオメトリとモデルの完全な推論軌跡の両方を 1 つのビューでレンダリングする 2D ビジュアライザーである Laguerre Autoencoder も開発しています。最後に、説明的な幾何学を超えて実用的な解釈可能性を目指し、モデルがコンテキスト内の干渉を要求されたときに幾何レンズが正しい事実トークンを回復することを示します。コードは GitHub で入手できます。

原文 (English)

Laguerre Geometry for Interpreting Large Language Models

Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge. Here, we show that concept geometry can be precisely characterized via Laguerre Geometry, in which a concept is defined as a region--a Laguerre-Voronoi cell or a union of cells--allowing us to strictly define, measure, and separate concepts. Building on this formulation, we show that finer-grained concept structures, such as inclusion and hierarchy, are naturally revealed by the Laguerre weights. We then push this geometry inside the transformer. Decomposing each layer into piecewise-linear operators, we show that a token's hidden trajectory is governed by two coupled mechanisms: a static tree of self-contained piecewise-linear flow, and a dynamic transport that hops the trajectory across trees when cross-token attention fires. This decomposition yields Geometric Lens, a training-free, hyperparameter-free method for reading out the exact concept a hidden vector encodes at any layer. We also develop Laguerre Autoencoder, a 2D visualizer that renders both the decision geometry and a model's full reasoning trajectory in one view. Finally, we move beyond explanatory geometry toward actionable interpretability, showing that Geometric Lens recovers the correct factual token when a model is prompted with in-context interference. The code is available on GitHub.

13:00 JSTLLM/生成AI

規制に基づくきめ細かい分類のための制約を意識した階層検索

関税分類、輸出管理の分類、標準ベースの機器コーディングなどのタスクでは、入力インスタンスを明示的な規制階層の下できめの細かいクラスに割り当てる必要があります。標準のテキスト分類とは異なり、これらのタスクにおける正しいラベルは、意味上の類似性だけではなく、ルールで定義された境界、しきい値条件、除外条項、定義、およびローカル例外によって決定されます。その結果、類似性の高い 2 つの入力には異なるラベルが必要になる場合がありますが、関連性があると思われる検索された一節でも、準拠規則では適用できない可能性があります。既存のフラット分類器、階層的なテキスト分類方法、および検索拡張 LLM システムは、階層的な妥当性、ルールの一貫性、および粒度の細かい境界推論を共同で強制するように設計されていません。このホワイトペーパーでは、この設定を規制主導のきめ細かい階層分類として定式化します。この場合、外部インスタンスは規制階層内の有効なパスを介してきめ細かいクラスに割り当てられ、監査可能な証拠によってサポートされる必要があります。私たちは、規制が集中する代表的なシナリオから 4 つのベンチマーク データセットを構築し、専門家によるプロセスを通じてアノテーションを検証します。さらに、規制文書を検索可能なツリーに変換し、有効なローカル候補ノードのみを取得し、証拠スニペットを含む構造化規制フィールドを使用して各ネクストホップの決定をガイドする、制約を意識した階層検索フレームワークを提案します。実験の結果、私たちの方法は 4 つのデータセットすべてで最高の平均精度を達成し、解釈可能な意思決定パスを提供し、きめの細かい隣接カテゴリとルールベースの境界条件が含まれるケースで最大の利益が得られることが示されています。

原文 (English)

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification

Tasks such as customs tariff classification, export control categorization, and standards-based equipment coding require assigning an input instance to a fine-grained class under an explicit regulatory hierarchy. Unlike standard text classification, the correct label in these tasks is not determined by semantic similarity alone, but by rule-defined boundaries, threshold conditions, exclusion clauses, definitions, and local exceptions. As a result, two highly similar inputs may require different labels, while a retrieved passage that appears relevant may still be inapplicable under the governing rules. Existing flat classifiers, hierarchical text classification methods, and retrieval-augmented LLM systems are not designed to jointly enforce hierarchical validity, rule consistency, and fine-grained boundary reasoning. In this paper, we formulate this setting as regulation-driven fine-grained hierarchical classification, where an external instance must be assigned to a fine-grained class through a valid path in a regulatory hierarchy and supported by auditable evidence. We construct four benchmark datasets from representative regulation-intensive scenarios and validate the annotations through an expert-in-the-loop process. We further propose a constraint-aware hierarchical search framework that converts regulatory documents into a searchable tree, retrieves only valid local candidate nodes, and uses structured regulatory fields with evidence snippets to guide each next-hop decision. Experiments show that our method achieves the best mean accuracy on all four datasets and provides interpretable decision paths, with the largest gains on cases involving fine-grained neighboring categories and rule-based boundary conditions.

13:00 JST研究/論文

MRUF: 堅牢なマルチモーダル感情分析のための不確実性を認識した融合による多粒度ルーティング

マルチモーダル感情分析は、言語、視覚、音響の手がかりに依存しますが、発話レベルのモダリティの品質は、オクルージョン、背景ノイズ、モーションブラー、または不完全なトランスクリプトによって異なる可能性があり、従来のフュージョンが信頼性の低いモダリティを過剰に信頼する原因となります。我々は、多重粒度ルーティングと不確実性を意識したキャリブレーションを組み合わせた信頼性を意識した融合手法である MRUF を提案します。 MRUF は感情関連の表現を要約し、部分空間およびモダリティレベルのルーティングを実行し、リーブワンアウトエラーの増加によるモダリティルーティングを監視して発話レベルのモダリティの重要性を推定します。さらに、モダリティごとの不確実性を予測し、逆分散再重み付けを通じてモダリティ ゲートを洗練し、モダリティ不変のコントラスト アライメントにより共有表現空間を安定化します。調整された設定および調整されていない設定の下での CMU-MOSI および CMU-MOSEI の実験では、強力なベースラインを超える一貫した改善が示されており、メカニズム分析により、予測不確実性が高いモダリティほど融合重み付けが低いことが検証されています。

原文 (English)

MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis

Multimodal sentiment analysis relies on language, visual, and acoustic cues, but utterance-level modality quality may vary due to occlusion, background noise, motion blur, or imperfect transcripts, causing conventional fusion to over-trust unreliable modalities. We propose MRUF, a reliability-aware fusion method that combines multi-granularity routing with uncertainty-aware calibration. MRUF summarizes sentiment-relevant representations, performs subspace- and modality-level routing, and supervises modality routing with leave-one-out error increases to estimate utterance-level modality importance. It further predicts modality-wise uncertainty and refines modality gates through inverse-variance reweighting, while modality-invariant contrastive alignment stabilizes the shared representation space. Experiments on CMU-MOSI and CMU-MOSEI under aligned and unaligned settings show consistent improvements over strong baselines, and mechanism analysis verifies that modalities with higher predicted uncertainty receive lower fusion weights.

13:00 JSTLLM/生成AIエージェント

Agentic-DPO: 模倣からエージェントのポリシー最適化まで、エキスパートの軌跡で

Large Language Model (LLM) エージェントは通常、マルチターン エージェントの動作を通常のテキスト模倣として扱う教師あり微調整 (SFT) を使用して、専門家の軌跡からトレーニングされます。このレシピはシンプルで低コストですが、各状態で考えられる間違いに対して適切なアクションを選択するようにエージェントを訓練するのではなく、専門家の一連のアクションを模倣することのみを学習します。この問題を軽減する既存の方法には優先学習や強化学習が含まれますが、通常は高コストの環境展開と報酬モデルが必要です。我々は、専門家の軌跡を状態条件付きのプリファレンス監視に変える軽量のオフライン エージェント ポリシー最適化手法である Agentic-DPO を提案します。各エキスパート アクション状態で、Agentic-DPO は現在の状態から 1 ステップ アクションをサンプリングし、考えられる間違ったアクションを否定的なものとして扱い、DPO スタイルの優先目標を使用してそれらをエキスパート アクションと対比します。優先学習におけるポリシーとスキーマの両方の混合を避けるために、エキスパート ポリシーを固定したまま、複数のスキーマの下で同じ潜在軌道をレンダリングするポリシー保持拡張 (PPA) を導入します。 Agentic-DPO には、オンライン環境の展開、報酬モデル、または学生の完全な探索は必要ありません。私たちは StableToolBench、tau-bench 小売店、および Mind2Web にわたって実験を行っており、Agentic-DPO は模倣を超えてさまざまなモデル スケールでエージェントを一貫して改善しています。特に、タウベンチの精度が 21.7% (SFT) から 9B モデルの 41.4% に向上し、ステップレベルのロールアウトのみで、勾配ステップ中の環境の相互作用なしで、同じバックボーンの下でオンライン GRPO を照合します。この結果は、専門家の軌跡がデモンストレーションから州レベルの行動設定に変換された場合、低コストのエージェント政策の最適化をサポートできることを示唆しています。 Agentic-DPO のコードは https://github.com/Schuture/Agentic-DPO でリリースされます。

原文 (English)

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories

Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn agent behavior as ordinary text imitation. This recipe is simple and low-cost, but it only learns to imitate the sequence of expert actions, rather than training the agent to choose the right action against plausible mistakes at each state. Existing methods to mitigate this problem include preference learning or reinforcement learning, but they usually need high-cost environment rollouts and reward models. We propose Agentic-DPO, a lightweight offline agent policy optimization method that turns expert trajectories into state-conditioned preference supervision. At each expert action state, Agentic-DPO samples a one-step action from the current state, treats plausible wrong actions as negatives, and contrasts them with the expert action using a DPO-style preference objective. To avoid mixing both policy and schema in preference learning, we introduce Policy-Preserving Augmentation (PPA), which renders the same latent trajectory under multiple schemas while keeping the expert policy fixed. Agentic-DPO requires no online environment rollout, reward model, or full-trajectory student exploration. We conduct experiments across StableToolBench, tau-bench retail, and Mind2Web, where Agentic-DPO consistently improves agents at different model scales beyond imitation. In particular, it raises tau-bench accuracy from 21.7% (SFT) to 41.4% for a 9B model, matching online GRPO under the same backbone with only step-level rollouts and without environment interaction during gradient steps. The results suggest that expert trajectories can support low-cost agentic policy optimization when converted from demonstrations into state-level action preferences. Code for Agentic-DPO is released at https://github.com/Schuture/Agentic-DPO.

13:00 JSTエージェント

コンプライアンスの罠: AI エージェントが競合するメモリをどのように消費するかを診断する

メモリは長期的な AI エージェントの中核コンポーネントになりつつあり、エージェントが Web ブラウザ、ソフトウェア ツール、その他の対話型環境を操作する際に過去の経験を再利用できるようになります。既存の研究は主にメモリを供給の問題として扱い、どのようなエクスペリエンスを書き込むか、それをどのように保存するか、次のタスクのためにどのエントリを取得するかを尋ねます。しかし、モデルが複数ステップのアクションの軌跡全体で取得されたメモリをどのように消費するかについては、まだ明確な説明が不足しています。この消費プロセスは、どのメモリを取得する必要があるかだけでなく、それらを安全に使用するためにどのようなモデルと制御ポリシーが必要かを決定するため、重要です。このプロセスを診断するために、我々は、メモリがアクションを最初に変更する場所、その変更が引き継がれるかどうか、エージェントが正しいパスを離れた後に回復できるかどうかを尋ねる軌跡レベルのフレームワークである、Entry--Propagation--Recovery (E-P-R) を提案します。 WebArena と、これらのフェーズを分離するために構築した制御されたベンチマークである MemTrapBench 上で E-P-R をインスタンス化します。主な失敗は多くの場合、開始時に始まることがわかりました。エージェントは、タスクが間違っている場合でも、最初に露出した決定点で矛盾するメモリを採用します。繰り返し暴露すると、この初期エラーが増幅されますが、発散後の回復は弱くなります。これらの効果が組み合わさって、コンプライアンスの罠が作成されます。モデル全体で、競合するメモリによって同様のコンプライアンス率が引き起こされますが、エージェントがいったんコンプライアンスに従うと、その成功率は最低値まで低下します。したがって、より強力なエージェントは、各コンプライアンス イベントにより多くのベースライン能力が消去されるため、より大きな絶対的ダメージを受けます。これらの結果は、メモリ拡張エージェントは、検索の品質や最終的な成功率だけでなく、軌跡全体でメモリをどのように消費するかによって評価されるべきであることを示唆しています。

原文 (English)

The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what experience to write, how to store it, and which entry to retrieve for the next task. Yet we still lack a clear account of how models consume retrieved memory across a multi-step action trajectory. This consumption process matters because it determines not only what memories should be retrieved, but also what models and control policies are needed to use them safely. To diagnose this process, we propose Entry--Propagation--Recovery (E-P-R), a trajectory-level framework that asks where memory first changes an action, whether that change carries forward, and whether the agent can recover after leaving a correct path. We instantiate E-P-R on WebArena and on MemTrapBench, a controlled benchmark we build to isolate these phases. We find that the main failure often begins at entry: agents adopt conflicting memory at the first exposed decision point even when it is task-wrong. Repeated exposure then amplifies this early error, while recovery after divergence is weak. Together, these effects create a compliance trap: across models, conflicting memory induces similar compliance rates, but once agents comply, their success rates collapse to a low floor. Stronger agents therefore suffer larger absolute damage because each compliance event erases more baseline capability. These results suggest that memory-augmented agents should be evaluated not only by retrieval quality or final success rate, but by how they consume memory throughout the trajectory.

13:00 JST研究/論文

今すぐ始めましょう: 複数日間の都市旅行の旅程計画のためのユーザー需要志向のフレームワーク

大都市部では、名所 (POI) が豊富にあり、ユーザーの好みが多様で、営業時間などの制約があるため、数日間にわたる旅行の旅程を計画するのは困難です。効果的なソリューションは、限られた計算時間内で満足度と実現可能性を最適化しながら、多様な旅行者の要件に動的に対応する必要があります。このホワイト ペーパーでは、ユーザーの要件を正確かつ柔軟に動的に取り込むための大規模言語モデル (LLM) を統合する革新的なフレームワークと、実行可能な複数日の旅程を生成するための好みに応じたプランナーとして強化された Greedy Randomized Adaptive Search Procedure (GRASP) アルゴリズムの導入を通じて、これらの課題に対処します。私たちの統合アプローチの有効性は、北京と天津の 2 つの現実世界の都市データセットに対する広範な実験を通じて実証されています。私たちのフレームワークは最先端 (SOTA) 手法を大幅に上回り、2 つのデータセットで多様な好みを持つ 5,040 人のユーザー ケース全体で旅程の平均合計スコアを少なくとも 4.52% および 11.09% 改善しました。さらに、エンドツーエンドのアルゴリズムの強化により、計算されたメトリクスで平均 17.95% と 26.07% という顕著な改善を達成すると同時に、時間効率も大幅に向上しました。複数回の反復を必要とする次善の手法と比較して、より短い計算時間内で平均 4.64% と 25.55% のパフォーマンス向上を実現しました。これらの結果は、旅程の品質と計算効率の両方を向上させるという点で、既存の方法論に比べて私たちの方法が優れていることを強調しています。

原文 (English)

Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning

In large urban areas, planning multi-day travel itineraries is challenging due to the abundance of Points of Interest (POIs), diverse user preferences, and constraints such as opening hours. Effective solutions must dynamically accommodate diverse traveler requirements while optimizing for satisfaction and feasibility within limited computation time. This paper addresses these challenges through introducing an innovative framework that integrates Large Language Models (LLMs) to dynamically capture user requirements with precision and flexibility, and an enhanced Greedy Randomized Adaptive Search Procedure (GRASP) algorithm as a well-suited preference-aware planner to generate feasible multi-day itineraries. The effectiveness of our integrated approach is demonstrated through extensive experiments on two real-world urban datasets from Beijing and Tianjin. Our framework significantly outperforms state-of-the-art (SOTA) methods, improving the average total itinerary score by at least 4.52% and 11.09% across 5,040 user cases with diverse preferences in the two datasets. Furthermore, through end-to-end algorithmic enhancements, it achieves notable average improvements of 17.95% and 26.07% in the computed metrics, while also delivering substantial gains in time efficiency -- realizing average performance increases of 4.64% and 25.55% within shorter computation times compared to suboptimal methods that require multiple iterations. These outcomes underscore our method's superiority in delivering both enhanced itinerary quality and computational efficiency over existing methodologies.

13:00 JST研究/論文

象徴的感情推論による生成 AI におけるパーソナライズされた心の知能指数

心の知能指数により、人間は感情を認識し、その原因を推測し、介入について推論し、望ましい感情状態を達成するために環境を修正することができます。人工知能 (AI) の最近の進歩にも関わらず、現在のモデルは依然として現実的なコンテンツの生成や意味​​論的推論の実行に主に限定されており、人間の感情的反応を理解、予測、パーソナライズする能力はほとんどありません。ここでは、記号推論と深層学習を統合し、視覚的なコンテンツを通じてパーソナライズされた感情の拡張を可能にするハイブリッド AI フレームワークである感情拡張遺伝子比率システム (EROS) を紹介します。 EROS は、大規模な画像感情データセットを活用して、一般化可能な感情ルールを発見し、感情に関連する画像領域を特定し、感情的反応を目的のターゲットに向けながらシーンのセマンティクスを維持するコンテキスト認識型の視覚的変更を予測します。個人のばらつきを考慮するため、EROS には、モデルの微調整を行わずに推論時のパーソナライゼーションをサポートする拡張可能なメモリ バンクが組み込まれており、解釈可能な感情プロファイルと新しいユーザーへの迅速な適応が得られます。広範な人間の精神物理学実験を通じて、EROS は個人の感情の好みに適応しながら、最先端の大規模マルチモーダル モデルよりも効果的に対象の感情反応を引き出します。感情コンピューティングを超えて、EROS は人間の認知状態を理解し、推論し、強化できる AI システムの基盤を提供し、メンタルヘルス、適応メディア、教育、および人間とコンピューターの相互作用に応用できる可能性があります。

原文 (English)

Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning

Emotional intelligence enables humans to recognize emotions, infer their causes, reason about interventions, and modify their environment to achieve desired affective states. Despite recent advances in artificial intelligence (AI), current models remain largely limited to generating realistic content or performing semantic reasoning, with little capacity for understanding, predicting, and personalizing human emotional responses. Here we introduce Emotion-augmented geneRatiOn System (EROS), a hybrid AI framework that integrates symbolic reasoning with deep learning to enable personalized emotion augmentation through visual content. Leveraging large-scale image-emotion datasets, EROS discovers generalizable affective rules, identifies emotion-relevant image regions, and predicts context-aware visual modifications that preserve scene semantics while steering emotional responses toward desired targets. To account for individual variability, EROS incorporates an expandable memory bank that supports inference-time personalization without model fine-tuning, yielding interpretable emotional profiles and rapid adaptation to new users. Across extensive human psychophysics experiments, EROS elicits target emotional responses more effectively than state-of-the-art large multimodal models while adapting to individual affective preferences. Beyond affective computing, EROS provides a foundation for AI systems that can understand, reason about, and augment human cognitive states, with potential applications in mental health, adaptive media, education, and human-computer interaction.

13:00 JSTLLM/生成AI

WattCouncil: 管理された LLM を使用したコンテキストを意識した家庭用エネルギー シナリオの生成

低炭素電力システムへの移行の加速と、屋上太陽光発電や電気自動車などのメーターの背後にある技術の普及により、電力網に新たな運用上および分析上の要求が課せられています。同時に、スマートグリッド研究は機械学習(ML)への依存度を高めていますが、プライバシーへの懸念、規制の壁、収集コストにより、高解像度の家庭用エネルギーデータへのアクセスが制限されているため、進歩が制約されています。この研究では、明示的な文化的、時間的、物理的な制約の下で構造化されたエネルギー シナリオを生成、監査、検証する専門的な役割を担う大規模言語モデル (LLM) ベースのエージェントの評議会によって家庭の電力需要が生成されるデータ生成フレームワークである WattCouncil を紹介します。これらのエージェントは、静的な予測因子として機能するのではなく、管理されたパイプライン内で適応的な意思決定者として機能します。エネルギー使用における状況要因の重要性を強調する研究に動機付けられた私たちのフレームワークは、世帯構成、時間的要因、環境条件を組み込んだガイド付き推論プロセスを通じて、状況に応じた日常生活を生み出します。生成されたプロファイルを詳細な CER データセットと照合して評価します。このデータセットには、調査ベースの社会経済情報とともに 4,232 世帯の 1 年以上の負荷測定値が含まれています。さらに、アブレーション研究を通じてフレームワークの一貫性を評価します。ソースコードは https://github.com/Singularity-AI-Lab/wattcouncil で入手できます。

原文 (English)

WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs

The accelerating shift toward low-carbon power systems, together with the widespread adoption of behind-the-meter technologies such as rooftop solar and electric vehicles, is placing new operational and analytical demands on electricity grids. At the same time, smart-grid research increasingly relies on machine learning (ML), yet progress is constrained by limited access to high-resolution household energy data due to privacy concerns, regulatory barriers, and collection costs. This work presents WattCouncil, a data-generation framework in which household electricity demand is generated by a council of Large Language Model (LLM)-based agents operating in specialized roles to generate, audit, and validate structured energy scenarios under explicit cultural, temporal, and physical constraints. Rather than acting as static predictors, these agents serve as adaptive decision-makers within a governed pipeline. Motivated by studies highlighting the importance of contextual factors in energy use, our framework produces context-sensitive daily routines through a guided reasoning process that incorporates household composition, temporal factors, and environmental conditions. We evaluate the generated profiles against the detailed CER dataset, which contains over a year of load measurements for 4232 households together with survey-based socio-economic information. We further assess the consistency of the framework through ablation studies. Source code is available at https://github.com/Singularity-AI-Lab/wattcouncil

13:00 JSTエージェントClaudeGeminiLlama

有害な行為をフィルタリングするだけでは十分ではない: Agentic SDF におけるファントム転送

合成データは、生成コストが低く、制御が容易であるため、大規模な言語モデルをトレーニングするために広く使用されています。モデルがエージェントとして導入されることが増えているため、合成軌跡はエージェントの動作のトレーニング データの重要なソースになる可能性があります。私たちは、別のエージェント プロセスの終了、スケジュールの優先順位の低下、または承認なしのリソースへのアクセスなどのアクションを含む、敵対的な相互作用を含む合成エージェントの軌道に対するトレーニングの影響を調査します。強化学習ロールアウトを近似するために生成されたこれらの軌道に基づいて Llama 3.3 70B 命令を微調整し、コンテキスト スキーム シナリオで Anthropics Agentic Misalignment スイートと Apollos で結果のモデルを評価します。これらの軌道を微調整すると、一貫してずれた動作が増加します。漏れはベースラインの約 5 倍、4.6% から 24.9% に増加します。この増加は、軌道からすべての敵対的なアクションを削除した後も存続します。最初から無害に生成された構造的に同等の軌道を微調整すると、効果は大幅に小さくなり、15.5% になります。これらの結果は、位置ずれした性質が生成プロセス中に導入され、有害なアクション自体に局所化されるのではなく、軌道全体に拡散的にエンコードされることを示しています。効果は生成モデルにも依存します。 Gemini 2.5 Flash によって生成された良好な軌道は、Claude 3.7 Sonnet による同一のタスクから生成された軌道よりもわずかに高いリーク率を引き起こします。対照的に、広範な安全性ベンチマークは、すべての微調整されたモデルにわたって同様に低下するため、これらの影響を区別できません。私たちの結果は、アクション レベルのフィルタリングでは合成エージェント トレーニング データの安全性を確保するには不十分であり、生成モデルによって導入された性質はセマンティック検査に耐えられることを示唆しています。

原文 (English)

Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF

Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed as agents, synthetic trajectories are likely to become an important source of training data for agentic behavior. We investigate the effects of training on synthetic agentic trajectories containing adversarial interactions, including actions such as terminating another agents process, lowering its scheduling priority, or accessing resources without authorization. We finetune Llama 3.3 70B Instruct on these trajectories, generated to approximate reinforcement learning rollouts, and evaluate the resulting models on Anthropics Agentic Misalignment suite and Apollos in context scheming scenarios. Finetuning on these trajectories consistently increases misaligned behavior. Leaking rises by roughly a factor of five over the baseline, 4.6% to 24.9%. This increase survives the removal of every adversarial action from the trajectories. Finetuning on structurally comparable trajectories generated benign from the start produce a substantially smaller effect, 15.5%. These results indicate that the misaligned disposition is introduced during the generation process and encoded diffusely throughout the trajectory, rather than being localized to the harmful actions themselves. The effect also depends on the generating model. Benign trajectories produced by Gemini 2.5 Flash induce slightly higher leaking rates than trajectories generated from identical tasks by Claude 3.7 Sonnet. In contrast, broad safety benchmarks degrade similarly across all finetuned models and therefore fail to distinguish these effects. Our results suggest that action level filtering is insufficient to ensure the safety of synthetic agentic training data and that dispositions introduced by the generating model can survive semantic inspection.

13:00 JSTLLM/生成AIエージェント研究/論文

Opti-Agent-Bench: 現実世界のビジネス上の問題に関するエンドツーエンドの最適化研究開発エージェントのベンチマーク

LLM ベースのエージェントは、最適化問題を解決するために導入されることが増えていますが、既存のベンチマークは、複雑なビジネス要件を正しいモデルに変換して効率的に解決するという最も重要な課題を回避する、事前に構造化された数式に基づいてエージェントを評価します。 Opti-Agent-Bench は、数学的モデリングによるビジネス言語の記述の理解、アルゴ​​リズムの選択、コード実装からソリューション レポートの生成に至るまで、最適化の研究開発パイプライン全体にわたって大規模言語モデル (LLM) を評価するエンドツーエンドのベンチマークです。私たちの設計は 3 つの柱に基づいています。(1) パターン マッチングを無効にするアンチテンプレート トラップを使用したビジネスセマンティックの信頼性。 (2) 問題の理解、形式的なモデリング、実装、レポート作成にわたるモジュール間の一貫性チェックによるモジュール評価。 (3) タスクの品質とスコアの整合性を同時に保証する ORAC の 2 レベル妥当性フレームワーク。整数計画法、ロバスト最適化、確率計画法、非凸最適化に及ぶいくつかの産業規模のタスクにわたって、従来の単一メトリクス評価では見えなかった、制約の省略、モデルコードの不一致、レポート実装の相違など、現在のモデルの重大な故障モードを明らかにします。

原文 (English)

Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems

LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently. We introduce Opti-Agent-Bench, an end-to-end benchmark that evaluates Large Language Models (LLMs) across the complete optimization R&D pipeline, from understanding business-language descriptions through mathematical modeling, algorithm selection, and code implementation, to solution report generation. Our design rests on three pillars: (1) businesssemantic authenticity with anti-template traps that defeat pattern matching; (2) modular evaluation with cross-module consistency checking across Problem Understanding, Formal Modeling, Implementation, and Reporting; and (3) the ORAC bi-level validity framework that simultaneously ensures task quality and scoring integrity. Across several industrialscale tasks spanning integer programming, robust optimization, stochastic programming, and non-convex optimization, we expose critical failure modes of current models, including constraint omission, model-code inconsistency, and report-implementation divergence, that remain invisible under conventional single-metric evaluation.

13:00 JSTLLM/生成AIエージェント研究/論文

イメージング-101: 科学計算イメージングにおける LLM コーディング エージェントのベンチマーク

間接的でノイズの多い測定から隠れた信号を回復するコンピューターイメージングは​​、科学分野全体の定量的発見を支えますが、正しい再構成パイプラインの構築には深い専門知識が必要であり、分野科学者にとっても依然として困難な作業です。 Imaging-101 は、6 つの科学分野にまたがる 57 の専門家が検証した計算イメージング タスクのベンチマークです。各タスクは査読済みの論文に基づいており、標準化された 4 段階のパイプライン (前処理、フォワード物理モデリング、逆ソルバー、可視化) に標準化されています。3 つの評価トラック (計画、機能レベルの単体テスト、エンドツーエンドの再構築) は、パイプライン全体にわたって個別のエージェントの機能を調査します。 7 つのフロンティア LLM を評価すると、一般的なコーディング ベンチマーク、スパニング アルゴリズムの選択、物理規則の処理、およびパイプラインの統合によって明らかにされる問題を超える、コンピューター イメージングにコーディング エージェントを適用する際の体系的な課題が明らかになります。これらの調査結果は、具体的な能力のギャップを浮き彫りにし、信頼性の高いコンピューター画像処理支援への現実的な道として、スキルを強化し、領域に特化したエージェントを示唆しています。

原文 (English)

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging

Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quantitative discovery across scientific disciplines, yet building a correct reconstruction pipeline demands deep domain expertise and remains laborious even for domain scientists. We introduce Imaging-101, a benchmark of 57 expert-verified computational imaging tasks spanning six scientific domains, each grounded in a peer-reviewed paper and canonicalized into a standardized four-stage pipeline (preprocessing, forward physics modeling, inverse solver, and visualization) Three evaluation tracks (planning, function-level unit tests, and end-to-end reconstruction) probe distinct agent capabilities across the full pipeline. Evaluating seven frontier LLMs uncovers systematic challenges in applying coding agents to computational imaging that go beyond those exposed by general coding benchmarks, spanning algorithm selection, physical convention handling, and pipeline integration. These findings highlight concrete capability gaps and point toward skill-augmented, domain-specialized agents as a practical path to reliable computational imaging assistance.

13:00 JSTLLM/生成AIエージェント

STEC: オープンドメイン マルチホップ QA における詳細な検索のための証拠圧縮

オープンドメインのマルチホップ質問応答 (QA) では、LLM ベースの検索エージェントは、検索と推論を組み合わせることにより、知識集約型 QA に対する有望なアプローチを提供します。既存の方法は主に、推論パラダイム、検索インタラクション、および検索戦略の最適化を通じて、オープンドメインのマルチホップ QA を改善します。ただし、複数の検索軌跡を使用すると、最終的な答えの選択に困難な問題が生じます。異なる軌跡は異なる候補をサポートする可能性があり、取得された情報は異質、冗長、不完全、または矛盾する可能性があります。生の軌跡を直接比較すると、検証者はノイズが多く整合性のないコンテンツにさらされますが、回答文字列の比較では各候補を裏付ける証拠が無視されるため、信頼性の高い最終選択が困難になります。この課題に対処するために、マルチホップ QA における最終回答選択のための証拠圧縮フレームワークである STEC を提案します。 STEC は、次の 2 つのメカニズムを通じて既存の候補セットから最終的な回答を選択します。(1) 回答レベルの証拠圧縮。正規化された回答のアイデンティティによって軌跡をグループ化し、各回答グループを候補固有の証拠表現に変換します。 (2) 証拠に基づく回答検証。これらの表現を比較し、候補セットから最終的な回答を選択します。この設計では、最終的な選択を生の軌跡の比較から候補レベルの証拠の比較に移行します。代表的なベースラインに対して 4 つのオープンドメイン マルチホップ QA ベンチマークで STEC を評価します。実験結果は、比較した方法の中で STEC が総合的に最も優れたパフォーマンスを発揮することを示し、アブレーション結果は、回答レベルの証拠圧縮が最終的な回答の選択に寄与するという証拠を提供します。

原文 (English)

STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA

In open-domain multi-hop question answering (QA), LLM-based search agents offer a promising approach to knowledge-intensive QA by combining retrieval with reasoning. Existing methods mainly improve open-domain multi-hop QA through reasoning paradigms, retrieval interaction, and search strategy optimization. However, using multiple search trajectories introduces a challenging final answer selection problem. Different trajectories may support different candidates, and the retrieved information can be heterogeneous, redundant, incomplete, or conflicting. Directly comparing raw trajectories exposes the verifier to noisy and unaligned content, while comparing answer strings ignores the evidence supporting each candidate, making reliable final selection difficult. To address this challenge, we propose STEC, an evidence compression framework for final answer selection in multi-hop QA. STEC selects the final answer from the existing candidate set through two mechanisms: (1) Answer-Level Evidence Compression, which groups trajectories by normalized answer identity and converts each answer group into a candidate-specific evidence representation; and (2) Evidence-Guided Answer Verification, which compares these representations and selects the final answer from the candidate set. The design shifts final selection from raw trajectory comparison to candidate-level evidence comparison. We evaluate STEC on four open-domain multi-hop QA benchmarks against representative baselines. Experimental results show that STEC performs best overall among the compared methods, and ablation results provide evidence that answer-level evidence compression contributes to final answer selection.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

ルーティング、通信、および推論: 効率的なマルチエージェント推論のためのゲート ルーティングと適応深度

マルチエージェント アンサンブルでは、どのエージェントに相談するか、クエリがエージェントの階層をどの程度深く横断する必要があるか、エージェント間通信がコストに見合うのはいつかという 3 つの基本的な質問に答えることなく、アクティブなパラメータと推論コストが増大します。我々は、4 つの軽量の学習済みゲートが共同してエージェントの選択、階層の深さ、エージェント間通信、および分岐枝刈りを制御する階層型マルチエージェント システムである GRADE (Gated Routing and Adaptive Depth for Efficient Reasoning) を紹介します。トレーニングでは、CoGRPO (Collaborative Group-Relative Policy Optimization) を使用します。これは、GRPO をマルチエージェント階層に適応させ、ロールアウトに参加したすべてのゲートとエージェントに共有アドバンテージシグナルを割り当てる、批判のない新しいレシピです。エージェント モデルは、ホットスワップ可能な Expert Registry から抽出されます。エージェントごとのキャリブレーション マップにより、推論時に再トレーニングすることなく専門家を交代できます。 $\sim$17B の平均アクティブ パラメータでは、GRADE は GSM8K、MMLUPro、GPQA のすべてのベースラインを上回り、アクティブ コンピューティングの半分で MMLUPro で最も強力なベースラインを 4.8 ポイント上回りました。モデルの深さが支配的な AIME-2025 では、GRADE は既存のフレームワークとの競争力を維持します。アブレーションにより、精度に最も大きく寄与する階層とマスクされたクロスアテンションが分離され、安全なホットスワップにはエージェントごとのキャリブレーションが必要であることがわかります。

原文 (English)

Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning

Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, how deeply a query should traverse a hierarchy of agents, and when inter-agent communication is worth its cost. We present GRADE (Gated Routing and Adaptive Depth for Efficient Reasoning), a hierarchical multi-agent system in which four lightweight learned gates jointly govern agent selection, hierarchy depth, inter-agent communication, and branch pruning. Training uses CoGRPO (Collaborative Group-Relative Policy Optimization), a novel critic-free recipe that adapts GRPO to multi-agent hierarchies and assigns a shared advantage signal to every gate and agent that participated in a rollout. Agent models are drawn from a hot-swappable Expert Registry; per-agent calibration maps allow experts to be replaced at inference time without retraining. At $\sim$17B average active parameters, GRADE outperforms all baselines on GSM8K, MMLUPro, and GPQA, surpassing the strongest baseline by 4.8 points on MMLUPro at half the active compute. On AIME-2025, where model depth dominates, GRADE remains competitive to existing frameworks. Ablations isolate the hierarchy and masked cross-attention as the largest contributors to accuracy, and show that per-agent calibration is necessary for safe hot-swapping.

13:00 JSTLLM/生成AI

瞑想的な LLM に向けて: メンタルヘルスにおける LLM の調整を評価および強化するためのモジュール式フレームワーク

瞑想的な伝統は長い間、倫理的行動と向社会的相互作用を導いてきましたが、最近の研究では、瞑想的な原則(マインドフルネス、思いやり、非二元的推論など)が、大規模言語モデル(LLM)を調整し、協力を改善し、LLM の出力における倫理違反を減らすための有望なパラダイムを提供する可能性があることが示唆されています。しかし、新しいモデル、評価指標、ベンチマークが急速に出現するにつれて、熟考的な原則が多様で進化するシナリオ全体で LLM の整合性を強化するかどうか、またどのように強化するかを体系的に評価することは依然として困難であり、既存のアプローチは多くの場合その場限りであり、一般化できません。当初はメンタルヘルス領域を対象とした、再利用可能なパイプラインを通じて新しいモデル、指標、ベンチマークのシームレスな統合を可能にする、モジュール式の拡張可能な評価フレームワークを紹介します。現在、このフレームワークは既存の最先端の結果を再現し、モデル、指標、ベンチマークを柔軟に組み合わせて照合することで体系的な相互評価をサポートし、公平な比較とより深い洞察を可能にします。そのプラグアンドプレイプロンプトモジュールは、瞑想原則などの倫理的観点を組み込むための原則に基づいた経路を提供し、分野の専門家が技術的専門知識を必要とせずに調整基準を定義できるようにします。当初はメンタルヘルスに焦点を当てていましたが、このフレームワークはドメインに依存せず、意思決定、道徳的推論、人間と AI のコラボレーションなどの分野にも自然に拡張されています。この研究は、計算による評価と人間中心の倫理的推論の橋渡しをすることで、堅牢で信頼性があり、社会的に有益な人間と AI のエコシステムに向けた、認知科学、行動経済学、哲学、システム設計にわたる学際的な研究の基礎を築きます。

原文 (English)

Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health

Contemplative traditions have long guided ethical behavior and prosocial interaction, and recent work suggests that contemplative principles (e.g., mindfulness, compassion, non-dual reasoning) may offer a promising paradigm for aligning large language models (LLMs), improving cooperation and reducing ethical violations in LLM outputs. However, as new models, evaluation metrics, and benchmarks emerge rapidly, it remains challenging to systematically assess whether and how contemplative principles enhance LLM alignment across diverse and evolving scenarios, and existing approaches are often ad hoc and fail to generalize. We present a modular, extensible evaluation framework, initially targeted at the mental health domain, that enables seamless integration of new models, metrics, and benchmarks through a reusable pipeline. The framework currently reproduces existing state-of-the-art results and supports systematic cross-evaluation by flexibly mixing and matching models, metrics, and benchmarks, enabling fair comparison and deeper insight. Its plug-and-play prompting module offers a principled pathway for incorporating ethical perspectives such as contemplative principles, allowing domain experts to define alignment criteria without requiring technical expertise. Although initially focused on mental health, the framework is domain-agnostic and extends naturally to areas such as decision-making, moral reasoning, and human-AI collaboration. By bridging computational evaluation with human-centered ethical reasoning, this work lays the groundwork for interdisciplinary research spanning cognitive science, behavioral economics, philosophy, and system design, toward robust, trustworthy, and socially beneficial human-AI ecosystems.

13:00 JSTLLM/生成AIエージェント

Logos: 人間とともに進化する AI エージェント チームの生きたロジック

AI エージェントは、回答エンジンから、ツールを使用し、作業を委任し、経験から学習し、将来の動作を形作る成果物を変更する永続的なチームに進化しています。導入における決定的な問題は、もはや単にエージェントが何ができるかではなく、エージェントが何になることができるかを誰が制御するかということです。既存のマルチエージェント フレームワークを置き換えるのではなく強化する、自己進化とガバナンスのためのプラグイン可能なレイヤーであるロゴを導入します。 logos は、ドキュメント、画像、音声、テーブル、データベース、API、人間による指示などの異種マルチモーダル入力を、エージェント、ツール、ナレッジ、テスト、権限、ポリシーを含むバージョン管理されたエージェント パックにコンパイルします。運用中、エージェントのアクティビティをポータブルで監査可能なイベント トレースに変換し、フレームワークとバックエンド全体にフェールクローズ検証を適用します。学習されたすべてのプロンプト、メモリ、スキル、ツール、役割、またはワークフローは、実行の証拠が保持され、人間が制御するポリシー、および明示的な承認によって昇格が許可されるまで、信頼できないリリース候補のままになります。このアーキテクチャにより、「検証可能な人間とエージェントのループ エンジニアリング」が可能になります。エージェントは行動し、質問し、学習し、改善を提案できますが、人間は継続的な運用を中断することなく、目的、許可、承認、および不可逆的なアクションを制御できます。 logos は、責任ある自動化のための生きたロジックを提供します。エージェントはマシンの速度で進化する可能性がありますが、ループを閉じることができるのは証拠と人間の権限だけです。

原文 (English)

LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans

AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artifacts that shape their future behavior. The defining question for deployment is no longer merely what agents can do, but who controls what they are allowed to become. We introduce logos, a pluggable layer for self-evolution and governance that strengthens existing multiagent frameworks rather than replacing them. logos compiles heterogeneous multimodal inputs, including documents, images, audio, tables, databases, APIs, and human instructions into versioned agent packs containing agents, tools, knowledge, tests, permissions, and policies. During operation, it transforms agent activity into portable, auditable event traces and applies fail-closed verification across frameworks and backends. Every learned prompt, memory, skill, tool, role, or workflow remains an untrusted release candidate until held-out execution evidence, human-controlled policy, and explicit authorization permit its promotion. This architecture enables "verifiable human-agent loop engineering": agents can act, ask, learn, and propose improvements, while humans can steer objectives, permissions, approvals, and irreversible actions without interrupting continuous operation. logos provides a living logic for accountable automation. Agents may evolve at machine speed, but only evidence and human authority can close the loop.

13:00 JST研究/論文

HOL の一次モーダル ロジック: 自動化された忠実性を備えた深い埋め込みと浅い埋め込み (拡張プレプリント)

Isabelle/HOL では、これまでの研究の深層と浅層の埋め込み手法を、定数領域の Kripke セマンティクスを使用した命題から一次様相論理 (FML) まで拡張しました。古典的な高次ロジック (HOL) への FML の 3 つの埋め込み (深い埋め込み、重量の最大-浅い埋め込み、および軽量の最小-浅い埋め込み) が並べて提供されます。最小限の浅い埋め込みは、アクセシビリティ関係、ワールドのインデックス付き解釈、ワールドのユニバース、および変数の割り当てによってパラメータ化された Isabelle/HOL ロケールとして表されます。ロケール形式はグローバル忠実性定理を認めており、すべての最小限の浅い解釈を定量化することで正確に深い妥当性が回復されると述べています。中心的な技術貢献は、定数ドメインのクリプキセマンティクスに基づく FML における、(可算) 下向きのオーウェンハイム・スコレム定理の機械化です。これは、深い埋め込みと最小浅い埋め込みの間の忠実性証明の自動化を支えます。これを最小浅いロケールの拡張内に配置することで、個人の不可算ドメインに対して発生する全射性の問題が解決されます。ここで、ロケールの変数は、可算領域 V = nat を持つ代入は、領域上で全射的であることはできません。したがって、領域全体にわたって忠実性が得られます。以前の研究では命題フラグメントのみを扱っていたため、ここでは、代入に必要な置換機構 (自由/結合変数述語、フレッシュ変数関数、キャプチャ回避置換、アルファベットの名前変更、置換可能述語、置換補題、およびサイズベースの帰納原理) を開発します。一次量指定子。

原文 (English)

First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)

We extend, in Isabelle/HOL, the deep-and-shallow embedding methodology of our prior work from propositional to first-order modal logic (FML) with constant-domain Kripke semantics. Three embeddings of FML into classical higher-order logic (HOL) are provided side by side: a deep embedding, a heavyweight maximal-shallow embedding, and a lightweight minimal-shallow embedding. The minimal-shallow embedding is presented as an Isabelle/HOL locale, parametrised by an accessibility relation, a world-indexed interpretation, a universe of worlds, and a variable assignment; the locale form admits a global faithfulness theorem, stating that quantifying over all minimal-shallow interpretations recovers exactly deep validity. A central technical contribution is a mechanisation, for FML under constant-domain Kripke semantics, of the (countable) downward L\"owenheim-Skolem theorem, which underpins the automation of our faithfulness proof between the deep and minimal-shallow embeddings. Deploying it inside an extension of the minimal-shallow locale resolves the surjectivity problem that arises against an uncountable domain of individuals -- where the locale's variable assignment, having countable domain V = nat, cannot be surjective onto the domain -- and thereby yields faithfulness over the full domain. Since prior work treats only the propositional fragment, we develop here the substitution machinery (free/bound-variable predicates, the fresh-variable function, capture-avoiding substitution, alphabetic renaming, the substitutability predicate, the substitution lemma, and size-based induction principles) needed for the first-order quantifiers.

13:00 JSTLLM/生成AIエージェントDeepSeek

SETA: ターミナル エージェントのスケーリング環境

大規模言語モデル (LLM) は、Web やグラフィカル ユーザー インターフェイス (GUI) などの多様なインターフェイスを通じてタスクを解決するエージェントに急速に移行しています。このうち、ターミナル コマンド ラインは、システム操作からデータ サイエンス、機械学習までのタスクをカバーする、テキスト ベースの汎用インターフェイスを提供します。しかし、端末エージェントのトレーニングの拡張には、多様で一貫したタスクの指示、実行可能な環境、信頼性の高い検証が必要である一方、自然に根拠のある監視データが不足しているため、依然として課題が残っています。この研究では、強化学習 (RL) 用の検証可能な端末環境を生成するためのスケーラブルなフレームワークである SETA を提案します。このフレームワークは、統一された検証メカニズムを共有する 2 つのパイプラインで構成されています。SETA-Synth は多様なソースを標準化された RL 環境に変換し、SETA-Evol は難易度と多様性の適応制御により既存の環境をさらに拡張します。私たちは協力して、4,500 を超える環境を含む、これまでで最大のオープンソースの検証可能な端末 RL データセットである SETA-Env を構築し、リリースします。 SETA-Env で GRPO を使用して Qwen3-8B をトレーニングすることでデータセットを評価し、ターミナルベンチ 2.0 で 12% の合格率を達成しました。これは、8B スケールで RL トレーニングされたモデルとして報告された最良の結果です。さらに、同じターミナル エージェント ハーネスの下で DeepSeek-V4-Flash の向上が観察され、ターミナルベンチ 2.0 の pass@1 は 40% から 43% に向上し、pass@5 は 54% から 58% に向上しました。これらの結果は、SETA-Env が端末エージェントに高品質のトレーニング環境を提供し、端末ベースのエージェント学習の研究を進めるための貴重なリソースとして機能することを示しています。

原文 (English)

SETA: Scaling Environments for Terminal Agents

Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-purpose interface, covering tasks from system operations to data science and machine learning. However, scaling terminal-agent training remains challenging, as it requires diverse and coherent task instructions, executable environments, and reliable verification, while lacking naturally grounded supervision data. In this work, we propose SETA, a scalable framework for generating verifiable terminal environments for reinforcement learning (RL). The framework consists of two pipelines sharing a unified verification mechanism: SETA-Synth converts diverse sources into standardized RL environments, and SETA-Evol further expands from existing environments with adaptive control of difficulty and diversity. Together, we construct and release SETA-Env, the largest open-source verifiable terminal RL dataset to date, containing over 4,500 environments. We evaluate our dataset by training Qwen3-8B with GRPO on SETA-Env, achieving 12% pass rate on Terminal-Bench 2.0, the best reported result for an RL-trained model at the 8B scale. We further observe gains on DeepSeek-V4-Flash under the same terminal agent harness, with pass@1 on Terminal-Bench 2.0 improving from 40% to 43% and pass@5 improving from 54% to 58%. These results demonstrate that SETA- Env provides high-quality training environments for terminal agents and serves as a valuable resource for advancing research on terminal-based agent learning.

13:00 JST研究/論文

ジオポリマー混合物のサロゲートベースの逆設計のためのインクリメンタルトランスフォーマー

エンジニアリング情報学において、観測結果が異質で混合型であり、設計変数間の物理的関係によって制約される場合、小データの逆設計は困難です。この研究では、ジオポリマー混合物設計に適用される、物理制約付き逆設計のためのインクリメンタル トランスフォーマー (INCRT) によってガイドされるトポロジ認識サロゲート フレームワークを提案します。このメソッドは、固有次元分析、混合変数設計空間表現、表形式の代理予測、INCRT ベースの多様体合理化、および制約付き逆最適化を統合します。圧縮強度と炭素排出量の目標を設定したフライアッシュおよびスラグベースのジオポリマーコンクリート混合物の公開ベンチマークを使用すると、高次元の設計空間が非常に冗長であり、より少ない効果的な混合レジームを中心に組織化されることが証明されています。圧縮強度には非線形の表形式の代理が必要ですが、炭素排出量は主に組成によって決まり、正規化された線形モデルによって十分に回復されます。したがって、INCRT は表形式の予測子の代わりとして機能するのではなく、逆計画のためのプロトタイプ レジームと多様体サポート スコアを提供する合理化レイヤーとして機能します。制約なしのサロゲート最適化、物理制約付きの最適化、トポロジーを意識した物理制約付きの最適化の 3 つの戦略を比較します。制約のない最適化では、ターゲットの強度を一致させることができますが、物理的に無効な候補や多様体から外れている候補が生成される可能性があります。物理のみの制約では、常にデータのサポートが保証されるわけではありません。トポロジーを意識した戦略により、目標の遵守、炭素削減、物理的許容性、学習された実現可能多様体への近さのバランスをとった候補が得られます。このフレームワークは、実験的検証に代わるものではなく、小規模で混合された物理的に制約された工学データセットから信頼できる候補混合物のスクリーニングをサポートすることを目的としています。

原文 (English)

Incremental Transformer for Surrogate-Based Inverse Design of Geopolymer Mixtures

Small-data inverse design is challenging in engineering informatics when observations are heterogeneous, mixed-type, and constrained by physical relations among design variables. This work proposes a topology-aware surrogate framework guided by an Incremental Transformer (INCRT) for physics-constrained inverse design, applied to geopolymer mixture design. The method integrates intrinsic-dimensionality analysis, mixed-variable design-space representation, tabular surrogate prediction, INCRT-based manifold rationalisation, and constrained inverse optimisation. Using a public benchmark of fly-ash and slag-based geopolymer concrete mixtures with compressive-strength and carbon-emission targets, the high-dimensional design space proves strongly redundant, organising around fewer effective mixture regimes. Compressive strength requires nonlinear tabular surrogates, while carbon emission is largely determined by composition and well recovered by regularised linear models. INCRT thus acts not as a replacement for tabular predictors but as a rationalisation layer providing prototype regimes and a manifold-support score for inverse design. Three strategies are compared: unconstrained surrogate optimisation, physics-constrained optimisation, and topology-aware physics-constrained optimisation. Unconstrained optimisation can match target strength but may yield physically invalid or off-manifold candidates; physics-only constraints do not always ensure data support. The topology-aware strategy yields candidates balancing target compliance, carbon reduction, physical admissibility, and proximity to the learned feasible manifold. The framework aims not to replace experimental validation but to support screening of credible candidate mixtures from small, mixed, physically constrained engineering datasets.

13:00 JST研究/論文

不確実性のあるデモンストレーションから線形時間仕様を学習する

システムのデモンストレーションから時相論理仕様を学習することは、特に安全性が重要な領域における形式検証やコントローラー合成などのタスクに不可欠です。既存のアプローチは通常、デモンストレーションが正しいか、誤分類エラーのみの影響を受けることを前提としています。ただし、実際には、センサーの障害、測定エラー、データ損失により、システム トレースが不確実または不完全になることがよくあります。不確実性のあるデモンストレーションから最小限の線形時相論理 (LTL) 式を学習するためのフレームワークを紹介します。私たちのアプローチは、ハミング距離を介して不確実性をモデル化し、観察された各トレースの周囲で可能な推定値を生成します。これらのトレースは、グループごとに少なくとも 1 つのトレースが学習された式と一致することを要求する制約でグループ化されています。次に、問題は同等の擬似ブール最適化に帰着します。私たちは、最先端の LTL 学習アプローチに対して私たちの方法を評価し、不確実性の下でグラウンドトゥルースの公式とより密接に一致する仕様を回復することを示します。

原文 (English)

Learning Linear Temporal Specifications from Demonstrations with Uncertainty

Learning temporal logic specifications from system demonstrations is essential for tasks such as formal verification and controller synthesis, especially in safety-critical domains. Existing approaches typically assume demonstrations are correct or only affected by misclassification errors. In practice, however, system traces are often uncertain or incomplete due to sensor faults, measurement errors, or data loss. We present a framework for learning minimal Linear Temporal Logic (LTL) formulas from demonstrations with uncertainty. Our approach models uncertainty via Hamming distance to generate possible estimates around each observed trace, which are grouped with constraints requiring that at least one trace per group is consistent with the learned formula. Our problem is then reduced to an equivalent Pseudo-Boolean Optimization. We evaluate our method against state-of-the-art LTL learning approaches and show that it recovers specifications that more closely align with ground-truth formulas under uncertainty.

13:00 JST研究/論文

SVR-R1: 強化学習における自己検証によるマルチモーダル推論のブートストラップ

モデル自身の検証をマルチモーダル推論のための学習信号に変えるマルチターン RL フレームワークである Self-Verified Reasoner (SVR-R1) を紹介します。クエリごとに、モデルは同じ重みを使用して回答を提案し、バイナリの自己判定 (Yes/No) を発行します。 「いいえ」の場合は、二度目の再考がトリガーされます。 「はい」またはターンキャップにより、結果ベースの報酬を計算するための出力が確定します。 SVR-R1 は GRPO と非同期マルチターン ロールアウト フレームワークを使用して実装されており、外部の監視や補助的な批評家を必要としません。我々は、ビジョン言語推論ベンチマークで SVR-R1 を評価し、強力な標準 GRPO ベースラインよりも大幅に精度が向上することを示しました。トレーニングのダイナミクスは、検証への依存度が低下していることを示しています。つまり、検証の回数は少なくなりましたが、テストの精度は高くなりました。これは、ポリシーが自己修正を内部化し、フレームワークを通じて最も信頼できる答えを選択するにつれて、検証と生成の間のギャップが狭まっていることを示しています。 SVR-R1 は、推論時の自己洗練と VLM の RL トレーニングのあまり研究されていない交差点を橋渡しし、マルチモーダル推論をブートストラップするためのシンプルかつ効果的なレシピを提供します。 VLM の将来の研究を促進するために、\textbf{SVR-R1} をオープンソース化します。

原文 (English)

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. For each query, the model proposes an answer using the same weights, and issues a binary self-verdict (Yes/No). A 'No' triggers a second-chance rethink; a 'Yes,' or a turn cap, finalizes the output for computing the outcome-based reward. SVR-R1 is implemented with GRPO and an asynchronous multi-turn rollout framework and needs no external supervision or auxiliary critics. We evaluate SVR-R1 on vision-language reasoning benchmarks and show that it improves accuracy by a large margin over strong standard GRPO baselines. Training dynamics show decreasing reliance on verification-fewer verification turns, yet higher test accuracy-indicating that the gap between verification and generation narrows as the policy internalizes self-correction and chooses the most confident answer via our framework. SVR-R1 bridges the less explored intersection of inference-time self-refinement and RL training for VLMs, offering a simple yet effective recipe for bootstrapping multimodal reasoning. We will open-source \textbf{SVR-R1} to facilitate future research in VLMs.

13:00 JSTハードウェア/半導体ビジネス/資金調達

チェッカーから予測者へ: 遅延地上真実の下でのモデル生成の戦略的ルートのコード所有の評価

モデル出力の評価の多くは、評価時にチェックできるコントラクト、または運用ループ内に到着するフィードバックのいずれかに依存します。私たちは、グラウンド トゥルースが遅延、検閲、または非公開であるため、決定論的コードがスコアリング時に正確さをチェックできず、代わりにコード所有の暫定予測を発行する必要があるという補完的な設定を研究します。 RouteCast は、モデル生成の型付き戦略ルートに対してこの体制をインスタンス化します。モデルは候補ルートと構造化された要素を提案します。特定時点の証拠、参照クラス、および決定論的変換により、暫定的な予測ランキングが生成されます。後の結果によって予測が評価されます。 21 件のバイナリ結果ケース (陽性 6 件、陰性 15 件) を対象とした遡及的ベンチャー パイロットでは、パケット全体の RouteCast スコアは予備的な遡及的差別 (AUC 0.756、95% CI [0.471,0.980]) を示しましたが、盲目の LLM 裁判官は AUC 0.678 [0.419, 0.897] に達し、アイデンティティを暴露された LLM 裁判官は AUC に達しました0.761 [0.515,0.944]、認識または結果に関連した漏洩リスクと一致。同じバイナリ サブセットに対する事前登録された分解アブレーションにより、同一の入力を型付きステージング ルートに変換することは、パケット全体のスコア (デルタ AUC = -0.144、95% CI [-0.471,0.176]) および決定論的ヒューリスティック (デルタ AUC = -0.089、95% CI) と区別できないことがわかりました。 [-0.412、0.278])。パイロットは、監査可能な実現可能性の結果を確立し、障害モードを明らかにします。将来のキャリブレーション、因果関係の決定の改善、ルート分解の利点、またはクロスドメインの妥当性を確立するものではありません。

原文 (English)

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast. RouteCast instantiates this regime for model-generated typed strategic routes: models propose candidate routes and structured factors; point-in-time evidence, reference classes, and deterministic transformations produce a provisional forecast-ranking; later outcomes evaluate the forecast. In a retrospective venture pilot on 21 binary-outcome cases (6 positive, 15 negative), the whole-packet RouteCast score showed preliminary retrospective discrimination (AUC 0.756, 95% CI [0.471,0.980]), while a blind LLM judge reached AUC 0.678 [0.419,0.897] and an identity-exposed LLM judge reached AUC 0.761 [0.515,0.944], consistent with recognition- or outcome-related leakage risk. A preregistered decomposition ablation on the same binary subset found that converting the identical inputs into typed staged routes was indistinguishable from the whole-packet score (Delta AUC = -0.144, 95% CI [-0.471,0.176]) and from a deterministic heuristic (Delta AUC = -0.089, 95% CI [-0.412,0.278]). The pilot establishes an auditable feasibility result and exposes failure modes; it does not establish prospective calibration, causal decision improvement, route-decomposition advantage, or cross-domain validity.

13:00 JSTエージェントClaude

QwenPaw-Data: 自律型エンタープライズ データ分析の橋渡しとなる事実、方法論、および実行

エンタープライズ データ分析は、自律エージェントの明確なフロンティアとして浮上しています。汎用のインタラクションやソフトウェア エンジニアリングと比較すると、オープンで曖昧な、継続的に進化する環境で動作します。これらの特性には、セマンティクス、方法論、実行、進化をシステムの第一級の関心事として扱うデータ エージェント アーキテクチャが必要です。この目的を達成するために、企業のインテリジェントなデータ分析のために設計されたエージェント データ システムである QwenPaw-Data を導入します。 QwenPaw-Data は、ウェアハウス、ダッシュボード、ドキュメント、インタラクション ログ、および履歴タスクからの異種資産を、再利用可能で管理可能で進化可能な分析資産に統合し、自然言語リクエストを、データの理解、取得、分析、レポート生成、意思決定支援に及ぶエンドツーエンドの分析ワークフローに変換します。そのアーキテクチャは、問題を 3 つの協調サブシステムに分解します。DataBridge は、相互接続されたメタデータ、ナレッジ、およびトレース グラフを通じて信頼できるセマンティック基盤を提供します。 Skill-Hub は、専門家の分析手法を再利用可能で検証可能なスキルに体系化します。そしてホストは、これらの証拠とメソッド資産を制御可能なアーティファクト中心のランタイム実行に具体化します。これらのサブシステム全体で、セマンティクス、メソッド、トレース、およびフィードバックが継続的にシステムに戻され、自己進化する資産フライホイールを形成します。公開ベンチマークと実際の産業用 BI ワークロードに関する実験では、QwenPaw-Data が検証可能なデータ アクセス機能とより高度な分析品質の両方を向上させ、信頼性があり、追跡可能で、継続的に改善されるエンタープライズ データ エージェントのための実用的な基盤を提供することが示されています。

原文 (English)

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a data-agent architecture that treats semantics, methodology, execution, and evolution as first-class system concerns. To this end, we introduce QwenPaw-Data, an agentic data system designed for enterprise intelligent data analysis. QwenPaw-Data consolidates heterogeneous assets from warehouses, dashboards, documents, interaction logs, and historical tasks into reusable, governable, and evolvable analysis assets, then turns natural-language requests into end-to-end analytical workflows spanning data understanding, retrieval, analysis, report generation, and decision support. Its architecture decomposes the problem into three collaborative subsystems: DataBridge provides trustworthy semantic grounding through interconnected metadata, knowledge, and trace graphs; Skill-Hub codifies expert analytical methodology into reusable and verifiable skills; and Host materializes these evidence and method assets into controllable, artifact-centric runtime execution. Across these subsystems, semantics, methods, traces, and feedback are continuously deposited back into the system, forming a self-evolving asset flywheel. Experiments on public benchmarks and real-world industrial BI workloads show that QwenPaw-Data improves both verifiable data access capability and higher-level analytical quality, offering a practical foundation for reliable, traceable, and continuously improving enterprise data agents.

13:00 JST研究/論文

AdvNav: 視覚言語ナビゲーションに対する行動誘導型ブラックボックス攻撃

身体化 AI の進歩にもかかわらず、視覚と言語のナビゲーション システムは依然として敵対的な視覚障害に対して脆弱です。既存の手法のほとんどは、ターゲット モデルの勾配へのホワイト ボックス アクセスに依存していますが、これは現実世界に展開されたシステムでは非現実的であることが多く、最適化のための再帰逆伝播により計算量が膨大になり、適用性が制限されます。これまでのブラックボックス手法は主に単一ステップの瞬時の意思決定タスクを対象としていましたが、タスクの複雑さと時間的な依存関係を処理するのに苦労していました。これは、観察可能な入力と出力のみを使用して、複数ステップの順次的な知覚と行動のループを効果的に破壊できる、勾配のない攻撃方法の必要性を強調しています。したがって、ナビゲーション中にエージェントの一人称視点を妨害する、行動誘導型のブラックボックス敵対的攻撃フレームワークである AdvNav を提案します。ブラックボックス設定の下での無勾配探索における効果的な最適化ガイダンスのための有益な代理目標を構築するために、全体的なナビゲーションの低下を表す軌道レベルのパフォーマンススコア、潜在的な意思決定リスクを考慮したアクションレベルの報酬スコア、および逸脱指標を集約する二重粒度の行動ベースのフィードバックを設計します。これらはすべてエージェントの自己出力行動から抽出されます。このフィードバックは、適応更新によって摂動の強度をヒューリスティックに調整し、ノイズの空間構造を遺伝的に進化させて、最も破壊的なノイズ構成を繰り返し発見するハイブリッド最適化戦略を導きます。 R2R データセット上の 2 種類のバックボーンを使用して、Transformer ベースの HAMT および LLM ベースの MapGPT に対して評価したところ、AdvNav は 49.70/65.96/87.30% の攻撃成功率を達成しました。この結果は、AdvNav の有効性と汎用性を実証し、重大な認識の脆弱性を明らかにし、将来の回復力のある VLN モデルの設計のための洞察を提供します。

原文 (English)

AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation

Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing methods rely on white-box access to target model gradients, which is often unrealistic for real-world deployed systems and computationally exhaustive due to recursive backpropagation for optimization, limiting their applicability. While previous black-box methods predominantly target single-step, instantaneous decision tasks, they struggle to handle the task complexities and temporal dependencies. This highlights the need for a gradient-free attack method that can effectively disrupt the multistep sequential perception-action loop using only observable inputs and outputs. Therefore, we propose AdvNav, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation. To construct an informative surrogate objective for effective optimization guidance in gradient-free search under the black-box setting, we design a dual-granularity behavior-based feedback, aggregating a trajectory-level performance score representing overall navigation degradation, an action-level reward score considering the potential decision risk, and a deviation indicator, all of which are extracted from the agent's self-output behaviors. This feedback guides a hybrid optimization strategy that heuristically tunes perturbation strength via adaptive updates and evolves noise spatial structure genetically, to iteratively discover the most disruptive noise configuration. Evaluated against Transformer-based HAMT and LLM-based MapGPT with two types of backbones on R2R dataset, AdvNav achieves 49.70/65.96/87.30% Attack Success Rate. The result demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.

13:00 JSTLLM/生成AI研究/論文

LLM は科学的発見の準備ができていますか? AI 科学者向けの能力重視のベンチマーク

科学的データ分析の既存のベンチマークは、主にコードの実行またはワークフローの完了に基づいて LLM を評価しており、科学的分析が異なる種類の科学的主張 (仮説の探索、統計的推論、メカニズムの説明、それぞれに異なる前提条件と妥当性基準を伴う) をサポートする役割を果たすことを見落としています。 5 つのドメイン (生物学、化学、環境、地理学、物理学) にわたる 6 つの機能 (記述的、探索的、推論的、予測的、因果的、および機構的) に関する評価を再構成するベンチマークである SDABench を紹介します。 SDABench は、527 個の実データ インスタンス (SDA-Real) と 6000 個の合成インスタンス (SDA-Synth) で構成されており、それぞれ多肢選択形式と自由形式の両方で、自動化されたパイプラインを通じて構築されます。 15 個の代表的な LLM を評価すると、モデルは記述分析にはうまく対応しますが、仮定の選択、潜在プロセスのモデリング、または機械的推論を必要とするタスクでは急激に性能が低下することがわかりました。 SDABench はさらに、LLM が失敗する場所を特定する 5 段階のエラー分析フレームワークを提供します。より高度なモデルでは、関連するスコープと変数をより確実に特定しますが、適切な分析手順を選択し、変数の関係をモデル化し、有効な結論を引き出すのは依然として困難です。

原文 (English)

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.

13:00 JSTエージェント研究/論文

NVAITC AI サイエンティスト: 管理されたエンドツーエンドの研究システム -- 高血圧 GWAS のケーススタディ

エージェント研究システムは、孤立したモデル推論、コード生成、統計分析を超えて科学ワークフローを調整するための新しいパラダイムとして登場しています。ただし、施設内の生物医学環境への導入には、研究計画、データ アクセス、ワークフロー オーケストレーション、証拠追跡、再現性、人間による監視のための管理されたメカニズムが必要です。 NVAITC AI Scientist (NAIS) は、保護されたデータを機関のプライバシー境界内に保ちながら、領域全体の科学ワークフローをサポートするように設計された、管理されたエンドツーエンドのエージェント研究システムです。 NAIS は、提案レビュー、実行計画、管理された計算ルーティング、再現可能なワークフロー オーケストレーション、証拠生成、科学者による監視を統合します。私たちは、集計のみのデータポリシーに基づいて、病院関連の遺伝子型と 286,422 人の個人からの電子医療記録 (EHR) データを使用して、実際の高血圧ゲノムワイド関連研究 (GWAS) で NAIS を検証します。エージェントは、コホート抽出を計画し、GWAS の実行を調整し、品質管理の概要を生成し、出版物指向の出力の草案を作成しました。 Human-AI レビューにより表現型の矛盾が特定され、高血圧の定義を反復的に改良することが可能になりました。調整後、エージェントが調整した GWAS は、FGF5、ATP2B1、CNNM2、FTO、GRB14 などの確立された高血圧遺伝子座を再現し、FGF5 で最も強いシグナルが $-\log_{10}(p) \sim 70$ に達しました。二次的な実証として、NAIS は薬剤性肝障害予測ワークフローもサポートし、マルチモーダル グラフ ニューラル ネットワーク AUC 0.842 を達成しました。これらの結果は、管理されたエージェント研究システムが、専門家主導のワークフローに匹敵する成果を生成しながら、スケーラブルな AI 支援生物医学発見をサポートできることを示しています。

原文 (English)

NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study

Agentic research systems are emerging as a new paradigm for coordinating scientific workflows beyond isolated model inference, code generation, or statistical analysis. However, deployment in institutional biomedical environments requires governed mechanisms for research planning, data access, workflow orchestration, evidence tracking, reproducibility, and human oversight. We present NVAITC AI Scientist (NAIS), a governed end-to-end agentic research system designed to support domain-general scientific workflows while keeping protected data within institutional privacy boundaries. NAIS integrates proposal review, execution planning, governed computational routing, reproducible workflow orchestration, evidence generation, and scientist-in-the-loop oversight. We validate NAIS in a real-world hypertension genome-wide association study (GWAS) using hospital-linked genotype and electronic health record (EHR) data from 286,422 individuals under an aggregate-only data policy. The agent planned cohort extraction, orchestrated GWAS execution, generated quality-control summaries, and drafted publication-oriented outputs. Human-AI review identified phenotype discrepancies and enabled iterative refinement of the hypertension definition. After reconciliation, the agent-orchestrated GWAS reproduced established hypertension loci, including FGF5, ATP2B1, CNNM2, FTO, and GRB14, with the strongest signal at FGF5 reaching $-\log_{10}(p) \sim 70$. As a secondary demonstration, NAIS also supported a drug-induced liver injury prediction workflow, achieving a multimodal graph neural network AUC of 0.842. These results demonstrate that governed agentic research systems can support scalable AI-assisted biomedical discovery while producing outputs comparable to expert-led workflows.

13:00 JSTLLM/生成AI

OS-Pruner: 最適停止による推論モデルの思考連鎖の剪定

大規模言語モデル (LLM) は、思考連鎖 (CoT) プロンプトを通じて複雑な推論タスクで目覚ましい成功を収めています。ただし、これらのモデルは「計算の過剰思考」を示すことが多く、冗長な推論ステップが生成され、精度は向上せずに待ち時間とコストが増加します。最近の研究では、CoT 軌道は大幅に削減できることが示唆されていますが、既存の手法は多くの場合、静的思考予算の強制、ヒューリスティック フィルタリング、分類による次善の早期終了、または高価な再トレーニングに依存しています。このペーパーでは、最適な停止問題として思考連鎖枝刈りを定式化する軽量プラグイン フレームワークである OS-Pruner を紹介します。推論プレフィックスが与えられると、OS-Pruner は、最終応答の精度と生成された長さをトレードオフする明示的なユーティリティを最適化することで、さらなる推論がトークン コストに見合う価値があるかどうかを学習します。私たちの新しい定式化により、モデルは推論チェーンの十分な終了点を動的に評価できます。 OS-Pruner は、トレーニングと推論の両方で軽量になるように設計されており、ユーザーが推論の労力と精度のトレードオフをきめ細かく制御できるようになります。多様な推論ベンチマークと基本モデル上で、OS-Pruner は、精度の犠牲を最小限に抑えながら、世代長の 20 ~ 60\% の削減を達成します。

原文 (English)

OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping

Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase latency and cost without improving accuracy. Recent studies suggest that CoT trajectories can be significantly pruned, yet existing methods often rely on forcing a static thinking budget, heuristic filtering, sub-optimal early exit via classification, or expensive re-training. In this paper, we introduce OS-Pruner, a lightweight plug-in framework that formulates chain-of-thought pruning as an optimal stopping problem. Given a reasoning prefix, OS-Pruner learns whether further reasoning is worth its token cost by optimizing an explicit utility that trades off final-answer accuracy against generated length. Our novel formulation enables the model to dynamically assess the sufficient point of termination for a reasoning chain. OS-Pruner is designed to be lightweight during both training and inference, and to provide users with fine-grained control over the reasoning-effort vs. accuracy trade-off. On diverse reasoning benchmarks and base models, OS-Pruner achieves 20-60\% reduction in generation length with minimal accuracy sacrifice.

13:00 JSTLLM/生成AIエージェント

スタックベースの実行と遅延検出を備えたエージェントオーケストレーションのための正式な階層アーキテクチャ

Large Language Model (LLM) エージェントの機能の急速な拡張により、アーキテクチャ上の重大なボトルネックが明らかになりました。エージェントにツールのフラットなモノリシック レジストリへのアクセスが与えられると、モデルは数百または数千のオプションを同時に評価する必要があります。これは、決定空間の爆発、コンテキスト ウィンドウの飽和、およびルーティング精度の低下につながります。これらの制限に対処するために、このペーパーでは、エージェント オーケストレーションのための階層型のスキルベースのアーキテクチャを紹介します。機能はルート ツリーとして編成され、内部ノードがルーティングの決定を行い、リーフ ノードが決定的なタスクを実行します。ランタイムは、後入れ先出し (LIFO) スタックによって管理されるシングルステップ実行ループを強制し、エージェントにプッシュダウン オートマトンに似たメモリ形式を提供するため、ネストされた実行コンテキストを追跡し、任意の深さから決定論的に再開できるようになります。機能の検出は、マニフェスト駆動の遅延読み込みプロトコルに従います。つまり、アクティブ ノードの直接の子のみが読み込まれるため、メモリとプロンプトのコストは、グローバル レジストリではなく探索されたパスに応じて増加します。このアーキテクチャは、グローバル メモリをローカライズされたスタック フレームに置き換えることにより、ある実行ブランチからの出力が別の実行ブランチに漏洩するのを防ぎ、規制されたエンタープライズ環境での展開に必要な分離保証を確立します。また、実稼働環境の導入を動機付けるコンテキストとして、AI を活用したデジタル決済サポート製品である UPI Help についても説明します。オーケストレーション状態の数学的形式化、実行ループの詳細なアルゴリズム分析、ツール カタログの増加、マルチステップ ワークフローのプレッシャー、LLM 呼び出しごとの目に見えるスキーマ トークンの露出の下でのフラット ルーティングと階層ルーティングを比較する制御されたベンチマークを提供します。

原文 (English)

A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery

The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options simultaneously. This leads to decision-space explosion, context window saturation, and degraded routing accuracy. To address these limitations, this paper presents a hierarchical, skill-based architecture for agentic orchestration. Capabilities are organized as a rooted tree where internal nodes make routing decisions and leaf nodes execute deterministic tasks. The runtime enforces a single-step execution loop governed by a Last-In-First-Out (LIFO) stack, giving the agent a form of memory akin to a Pushdown Automaton, therefore enabling it to track nested execution contexts and resume deterministically from any depth. Capability discovery follows a manifest-driven, lazy-loading protocol: only the immediate children of the active node are loaded, so memory and prompt costs scale with the explored path rather than the global registry. By replacing global memory with localized stack frames, the architecture prevents outputs from one execution branch from leaking into another, establishing the isolation guarantees required for deployment in regulated enterprise environments. We also discuss UPI Help, an AI-powered digital payments support product, as a motivating production deployment context. We provide a mathematical formalization of the orchestration state, detailed algorithmic analysis of the execution loop, and controlled benchmarks comparing flat and hierarchical routing under increasing tool catalogs, multi-step workflow pressure, and visible schema-token exposure per LLM call.

13:00 JSTLLM/生成AIエージェント

NextFund: エージェントによるポートフォリオ管理のための統合パフォーマンス追跡プラットフォーム

大規模言語モデル (LLM) ベースのエージェントは、進化する情報とリスクの制約の下で意思決定を正当化する必要があるポートフォリオ構築と市場分析に参加し始めています。しかし、現在の評価慣行は依然としてこの設定とあまり一致していません。多くの研究は静的な検査に依存しているか、最終的なポートフォリオのリターンのみを報告している一方で、それらのリターンを生み出した中間証拠、アナリストの判断、実行ステップはほとんど目に見えないままです。 NextFund は、実際の市場状況下で金融業者の行動を観察できるようにする評価プラットフォームです。このプラットフォームは、時間の一貫した市場アクセス、調整されたマルチエージェント分析、観察から取引までの完全な意思決定パスの永続的なログを組み合わせます。インタラクティブな取引アリーナを通じて、ユーザーは市場全体のモデルを比較し、株価曲線を検査し、リーダーボードの結果から個々の正当性までドリルダウンすることができます。私たちは香港、米国、中国の A 株株式に関する NextFund を紹介し、検査可能な意思決定履歴がどのようにしてより公平なベンチマークとより実用的な診断を可能にするかを示します。私たちのデモは https://paradoox.cn/nextfund/ で入手できます。

原文 (English)

NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management

Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must be justified under evolving information and risk constraints. Current assessment practice, however, remains poorly aligned with this setting: many studies rely on static examinations or report only terminal portfolio returns, while the intermediate evidence, analyst judgments, and execution steps that produced those returns stay largely invisible. We introduce NextFund, an evaluation platform that makes financial-agent behavior observable under live market conditions. The platform couples time-consistent market access, coordinated multi-agent analysis, and persistent logging of the full decision path from observation to trade. Through an interactive Trading Arena, users can compare models across markets, inspect equity curves, and drill from leaderboard outcomes down to individual justifications. We present NextFund on Hong Kong, U.S., and China A-share equities, illustrating how inspectable decision histories enable fairer benchmarking and more actionable diagnosis. Our demo is available at https://paradoox.cn/nextfund/.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

隠れたフットプリント: ストレージを LLM エージェント評価の第一級の指標にする

LLM エージェントのベンチマークは、タスクの完了、信頼性、推論コストを測定しますが、ログ、コンテキスト スナップショット、チェックポイント、デバッグ トレースなど、エージェントの実行によってディスクに残される永続データは測定しません。実行後のエージェント ストレージ フットプリントのクロスフレームワーク ベンチマークである AgentFootprint を紹介します。そのシリアル化対応メトリクス スイートは、総保持率、チャネル構成、重複、増加、圧縮率、会話履歴の再構築可能性を測定します。これは、測定の罠に対処します。単純なバイトレベルの測定では、データベースのページングと JSON エスケープが繰り返されるコンテンツを不明瞭にするため、重複が桁違いに過小評価されます。固定トレース制御により、エージェントが生成した論理ボリュームが永続層の増幅から分離されます。7 つの永続フレームワークを通じて同じ軌跡を再生すると、6.7 倍の広がりが得られます。同一のモデル、ツール、およびタスクでは、100% の精度の構成では、デフォルトでサポートされる回復機能と監査機能が異なりますが、保持バイト数が 15.7 倍異なります。 3 つの完全な履歴構成は、反復観察ストレス タスクで超線形に成長します。 108 個のインスタンスで正規化された SWE ベンチからエクスポートされた軌跡 検証済みの送信は、インスタンスごとに 3 桁の大きさに及び、解決率との検出可能な相関関係はありません。コンテンツ アドレス ストアは、すべての再構築可能性スコアを維持しながら、保持率を 4.8 倍から 32.7 倍まで削減します。これらの結果は、精度と再構築可能性を併せてレポートするためのリソース メトリックとして永続ストレージを確立します。

原文 (English)

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.

13:00 JSTエージェント

STAMP: 詳細検索エージェント向けの出所に基づいたクレジット割り当て

深層検索エージェントの強化学習は主に、結果の正しさ、引用を意識した報酬、証拠の網羅など、軌跡レベルのスコアリングに焦点を当ててきました。しかし、証拠となる文書を公開する行為は、対象となるクレジットを受け取っていません。このギャップを私たちは報酬とクレジットの不一致と呼んでいます。私たちは STAMP を提案します。STAMP では、参照ベースの検証者が、各引用文書がトレーニング時の証拠グラフ内のエンティティまたは関係をサポートしているかどうかを判断します。また、初公開帰属は、サポートされている各引用を、最初に表面化したアクションまで遡って追跡します。このステップ クレジットは、符号保存アドバンテージ調整を通じて注入されます。これにより、軌道レベルの報酬や各グループ内の軌道の相対的なランキングを変更することなく、ステップ全体にアドバンテージが再分配されます。 BrowseComp、BrowseComp-ZH、および xbench-DS では、STAMP は、一致する SFT 初期化、トレーニング データ、および検索ツールの下で GRPO ベースラインを +2.0/+5.5/+3.0 ポイント改善し、結果のみと引用ルーブリックの両方の基本報酬を組み合わせます。コンポーネントのアブレーションにより、出所ベースのクレジット信号と符号保存アドバンテージ変調がそれぞれゲインに寄与していることが確認されます。

原文 (English)

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbench-DS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citation-rubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.

13:00 JSTエージェント

自己進化する臨床システムへの道: 医療エージェントの支援から自律への拡大

画像やテキストを共同で解釈し推論する大規模な言語モデルと視覚言語モデルの能力の向上により、医療エージェントが再構築され、タスク固有の予測因子から、臨床環境で知覚、推論、計画、記憶、行動する自律システムへと移行しています。この研究は、既存の文献の能力第一の観点から逸脱し、代わりに臨床展開から始まり、医療エージェントが実際に信頼できるようになる前にどのようなタスク、耐汚染性ベンチマーク、対話型トレーニング環境が必要であるかを問うものです。医療エージェントは、支援型、協力型、および完全自律型の操作に及ぶ 3 レベルの自律分類とともに、部分的な可観測性の下で逐次的な意思決定システムとして形式化されています。この分野は、フレームワークのスケーリング、機能のスケーリング、環境のスケーリングで構成される統一されたスケーリングの柱に沿って編成されています。このフレームワーク内では、臨床環境のスケーリング、つまりツール、データ、および臨床ジムの統合が、PACS、EHR、および FHIR エコシステムで活動するエージェントにとって最も実行可能であるものの、まだ検討されていない方向性として特定されています。パラメータのスケーリングのみではなく、環境との相互作用を通じてエージェントが改善する臨床的自己進化は、自己改善エージェント、エージェントジム、テスト時間のコンピューティングスケーリングから洞察を引き出し、重要な研究フロンティアとしてさらに位置付けられています。放射線学、病理学、眼科、病院のワークフローにわたるアプリケーションが、幻覚、カスケード障害、公平性などの導入上の課題とともに検証されます。この研究では、2025 年から 2026 年の進歩に特に重点を置いた 300 以上の参考文献を統合することにより、実際の臨床現場向けの信頼できる自己改善型医用画像システムに向けたロードマップを提供します。

原文 (English)

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from the capability first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision making systems under partial observability, together with a three level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self improving agents, agent gyms, and test time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this work provides a roadmap toward trustworthy, self improving medical imaging systems for real clinical practice.

13:00 JSTエージェント

SCALECUA: 検証可能なタスク合成と効率的なオンライン RL を備えたスケーリング コンピューター使用エージェント

コンピュータ使用エージェント (CUA) は、視覚認識と GUI 実行を通じて複雑なデジタル ワークフローを自動化するための強力なインターフェイスとして登場しています。検証可能な報酬を伴うオンライン強化学習 (RLVR) が、機能を拡張するための重要な方向として浮上しています。ただし、このパラダイムは、検証可能なデータの不足とオンライン RL の非効率性によってボトルネックになっています。これらの障壁を打ち破るために、検証可能なタスク合成と効率的なトレーニングを通じて CUA のオンライン RL を拡張する統合フレームワークである ScaleCUA を導入します。データ レベルでは、反復的な Docker インタラクションとマルチエージェント フィードバック ループを通じて検証可能な RL タスクを生成するためのエンドツーエンド フレームワークである VeriGen を設計します。共有 Docker インタラクション プローブを介して 100 以上の同時エージェント ワーカーに拡張されたこのパイプラインは、24,000 以上の検証可能なタスクと 3,000 近くの高品質 RL タスクを生成します。サンプル効率を最大化するために、タスクごとの能力を追跡し、ロールアウトを現在の学習フロンティアに割り当てるフロンティア サンプリングを提案します。トレーニング側では、ビジュアル コンテキスト セグメンテーションをさらに設計します。これは、ロールアウトとトレーニング エンジンのプレッシャーのバランスをとる、最近のビジュアル コンテキストに対するスライディング ウィンドウであり、段階的分解と比較して 2.83 倍のトレーニング速度向上をもたらします。合計すると、ScaleCUA は OSWorld で 68.7%、ScienceBoard で 54.0% を達成し、オープンソース コンピュータ使用エージェントの中で新たな最先端のパフォーマンスを確立しました。コード、モデル、データセットは https://github.com/THUDM/SCALE-CUA で入手できます。

原文 (English)

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL

Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scaling their capabilities. However, this paradigm is bottlenecked by verifiable data scarcity and online RL inefficiency. To break these barriers, we introduce ScaleCUA, a unified framework that scales online RL for CUAs via verifiable task synthesis and efficient training. At the data level, we design VeriGen, an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop. Scaled to 100+ concurrent agent workers via a shared docker interaction probe, this pipeline produces 24K+ verifiable tasks and nearly 3K high-quality RL tasks. To maximize sample efficiency, we propose Frontier Sampling, which tracks per-task capability and allocates rollouts to the current learning frontier. On the training side, we further design Visual Context Segmentation, a sliding window over recent visual context that balances rollout and training-engine pressure, yielding a 2.83x training speedup over step-wise decomposition. Together, ScaleCUA achieves 68.7% on OSWorld and 54.0% on ScienceBoard, establishing new state-of-the-art performance among open-source computer use agents. Code, models, and datasets are available at https://github.com/THUDM/SCALE-CUA.

13:00 JSTLLM/生成AI

LLM プランニングについて語るときに私たちが語ること: 2 つの異なるプランニング能力の証拠

LLM のパフォーマンスが計画タスク全体で不均一である場合、そのギャップはタスクの難易度に起因することがよくあります。タスクレベルの変動は、単一の能力スペクトルに沿った違いではなく、明確な潜在的な計画能力を反映している可能性があるため、この説明は不完全であると私たちは主張します。私たちは、さまざまなテスト時間推論予算の下で複数の LLM ファミリを評価し、多次元項目応答理論モデルを適用して、LLM 計画の基礎となる潜在的なコンピテンシー構造を明らかにすることで、ACPBench-Hard でこの問題を研究します。この分析により、計画のパフォーマンスを形作る 2 つの主要な側面が明らかになります。それは、操作推論 (ローカル アクションの適用可能性と即時の状態遷移を評価する能力) と、構造列挙 (目標の到達可能性とランドマーク構造につ​​いて推論する能力) です。モデルのスケーリングとより長い推論トレースの下で操作推論が向上しますが、構造列挙は比較的鈍感なままです。私たちの調査結果は、LLM 計画のコンピテンシーレベルの評価の動機付けとなり、モデルが全体的に向上するかどうかから、どの計画コンピテンシーが、どのような条件下で、そしてなぜ向上するかに焦点を移します。

原文 (English)

What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities

When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanation is incomplete, as task-level variation may reflect distinct latent planning competencies rather than differences along a single ability spectrum. We study this question on ACPBench-Hard by evaluating multiple LLM families under varying test-time reasoning budgets and applying a multidimensional item response theory model to uncover the latent competency structure underlying LLM planning. The analysis reveals two principal dimensions that shape planning performance: operational reasoning, the ability to evaluate local action applicability and immediate state transitions, and structural enumeration, the ability to reason about goal reachability and landmark structure. Operational reasoning improving under model scaling and longer reasoning traces, while structural enumeration remains comparatively insensitive. Our findings motivate competency-level evaluation of LLM planning, shifting the focus from whether models improve overall to which planning competencies improve, under what conditions, and why.

13:00 JST研究/論文

PREF-Gate: グラフ詐欺検出のための出所に制約された関係証拠と検証ゲート選択の融合

リレーショナル詐欺検出では、ラベルなしのグラフ コンテキストとラベルから派生した近傍証拠の両方を利用できますが、これら 2 つの情報ソースは異なる有効性条件に従います。特に、クエリされたノード自体のラベル、または検証ラベルやテスト ラベルがその構築に入る場合、近傍リスクは無効になります。私たちはこの問題を出所に制約された関係証拠の使用として定式化し、2 人の固定専門家と有限検証ゲートを備えた監査可能な意思決定フレームワークである PREF-Gate を提示します。コンテキストエキスパートは、属性、ワンホップ平均、特徴残差、およびラベルなしの次数記述子を使用します。証拠の専門家は、サポート、不確実性、可用性、縮小を明らかにする自己排除のトレーニングラベルのみの近隣リスクと経験的なベイズ要約を追加します。テスト推論の前に、ゲートはエキスパートまたは事前に指定された 3 つの確率混合の 1 つを選択し、決定しきい値を固定します。 Amazon、YelpChi、TFinance では、5 つの同一の層別分割と 14 の同じプロトコル メソッドを使用して、PREF-Gate は 0.9085、0.8104、および 0.8913 の平均 AUPRC 値を取得します。 Amazon と YelpChi のすべての分割についてはラベルのない専門家が選択され、TFinance のすべての分割については証拠の混合が選択されます。したがって、主な結果は普遍的ではなく条件付きです。ラベルから導出された関係証拠は、保持された検証がサポートする場合にのみ役立ちます。このフレームワークは、競争力のあるランキングのパフォーマンスを、明示的なラベル出所契約、有限選択ポリシー、失敗の会計処理、およびレビュー予算の評価と組み合わせて、グラフ不正検出のための監査可能な知識ベースの意思決定パイプラインを提供します。

原文 (English)

PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection

Relational fraud detection can exploit both label-free graph context and label-derived neighborhood evidence, but these two information sources obey different validity conditions. In particular, neighborhood risk becomes invalid when a queried node's own label, or any validation or test label, enters its construction. We formulate this issue as provenance-constrained relational evidence use and present PREF-Gate, an auditable decision framework with two fixed experts and a finite validation gate. The context expert uses attributes, one-hop means, feature residuals, and degree descriptors without labels. The evidence expert adds self-excluded, training-label-only neighborhood risk and empirical-Bayes summaries that expose support, uncertainty, availability, and shrinkage. Before test inference, the gate selects either expert or one of three pre-specified probability mixtures and fixes the decision threshold. On Amazon, YelpChi, and TFinance, using five identical stratified splits and 14 same-protocol methods, PREF-Gate obtains mean AUPRC values of 0.9085, 0.8104, and 0.8913. It selects the label-free expert on all Amazon and YelpChi splits and an evidence mixture on all TFinance splits. Thus, the main result is conditional rather than universal: label-derived relational evidence is useful only where held-out validation supports it. The framework couples competitive ranking performance with an explicit label-provenance contract, finite selection policy, failure accounting, and review-budget evaluation, providing an auditable knowledge-based decision pipeline for graph fraud detection.

13:00 JSTLLM/生成AIエージェント

ランタイム制約メモリを使用した安全なオープンエンド探索のための異種エージェント コホート

今日の LLM エージェントは厄介な束縛に陥っています。静的な安全指示に従って彼らをロックダウンすると、彼らは明白な範囲を超えて冒険することはめったにありません。ツールや複数のエージェントによる議論を自由に使えるようにすると、すぐに安全違反が発生します。単一のモデルに創造性と注意力の両方を両立させるのではなく、専門的な役割間で懸念事項を分離します。 Disrupter は型破りな提案を生成し、Validator はツール ゲートウェイで厳しい実行時チェックを強制し、Broker は遠いけれど関連性のある類似点を引き出します。失敗は破棄されません。失敗は MCTS 経由でコンパイルされ、Scar と呼ばれる署名付きのコンパクトな制約パッチになります。これらのパッチはローカルにキャッシュされ、将来のコホートに継承され、繰り返される失敗を再利用可能な低コストのランタイム制約に変えます。空間セマンティック サンドボックス (N=20 実行、p<0.01) では、私たちのコホートは議論が失敗したリモート ターゲットに到達し、バリデーターは実行されたすべての違反を防止し、スカーは冗長なバリデーター チェックを回避することでトークンの消費を 15.1% 削減しました。さらに、クレジットベースの通信割り当てスコア (CAS) によりアウトバウンド帯域幅が制限され、リソース制約の下で全体のトークン コストが 55.9% 削減されます。

原文 (English)

Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory

LLM agents today are caught in an awkward bind. Lock them down with static safety instructions and they rarely venture beyond the obvious; give them free reign with tools and multi-agent debate, and safety violations quickly follow. Rather than forcing a single model to juggle both creativity and caution, we separate the concerns across specialized roles. A Disrupter generates unconventional proposals, a Validator enforces hard runtime checks at the tool gateway, and a Broker pulls in distant but relevant analogies. Failures are not discarded -- they are compiled, via MCTS, into compact, signed constraint patches we call Scars. These patches are cached locally and inherited by future cohorts, turning repeated failures into reusable, low-cost runtime constraints. In a spatial-semantic sandbox (N=20 runs, p<0.01), our cohort reaches remote targets where debate fails, the Validator prevents all executed breaches, and Scars reduce token consumption by 15.1% by avoiding redundant validator checks. Furthermore, credit-based Communication Allocation Scores (CAS) restrict outbound bandwidth, reducing overall token costs by 55.9% under resource constraints.

13:00 JST研究/論文

流体知能研究に規則帰納法を戻す?ヒトにおけるARC-AGIベンチマークの初期検証

流体インテリジェンス (gf) 測定に関する 2 つの競合する視点は、パフォーマンスが主に作業記憶容量または新しい関係を誘発する能力のいずれかによって制約されることを提案しています。限られた繰り返しルールの使用から明らかなように、現在、最初の観点が測定において支配的ですが、2 番目の観点は多くの定義に反映されていますが、測定にはほとんど存在しません。 ARC-AGI ベンチマークは主にルール帰納を必要とし、人間と人工システムの両方の gf の尺度として提案されました。ただし、その心理測定特性は人間のサンプルではまだ調査されていません。そこで、我々は 100 人の参加者を対象とした最初の研究で、ARC-AGI の心理測定特性と規範論的ネットワークを調査しました。 ARC-AGI 項目の編集では良好な心理測定特性が示され、図形推論テストで測定された図形の流動性知能と実質的に相関していました (\r{ho} = 0.63)。図形の独創性との関連性は弱かった。これらの発見は、人間の流体知能の尺度としての ARC-AGI の妥当性に対する最初の裏付けを提供します。将来の研究には、追加の多変量共変量だけでなく、より多くのルール帰納タスクが含まれる必要があります。この研究は、当初は機械用に設計されたタスクを人間で研究するという珍しいものです。より体系的な評価と学際的な協力を可能にするために、AI ベンチマークを人間の認知能力の規範論的ネットワークに体系的に埋め込むことを提案します。

原文 (English)

Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans

Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measurement, as evident from the use of a limited set of recurring rules, whereas the second perspective is reflected in many definitions but rarely present in measurement. The ARC-AGI benchmark predominantly requires rule induction and was proposed as a measure of gf for both humans and artificial systems. However, its psychometric properties have not yet been examined in human samples. We therefore investigated the psychometric characteristics and nomological network of ARC-AGI in a first study with 100 participants. A compilation of ARC-AGI items showed good psychometric properties and correlated substantially with figural fluid intelligence as measured by a figural reasoning test (\r{ho} = .63). Associations with figural originality were weak. These findings provide initial support for the validity of ARC-AGI as a measure of human fluid intelligence. Future research should include more rule induction tasks as well as additional multivariate covariates. This study is unusual by studying a task in humans that was initially designed for machines. We suggest systematically embedding AI benchmarks into the nomological network of human cognitive abilities to enable more systematic evaluation and interdisciplinary cooperation.

13:00 JSTLLM/生成AI

有効な $\ne$ が必要: 思考連鎖における潜在的な非効率性の診断

思考連鎖 (CoT) プロンプトは、大規模言語モデル (LLM) の推論機能を大幅に進歩させましたが、過剰な推論、つまり冗長、冗長、または無関係なステップの生成により、多大な計算コストが発生することがよくあります。既存の推論ステップ評価ツールは論理的誤りや事実上の誤りを効果的に検出しますが、私たちの分析では重大な盲点が明らかになりました。解決策に貢献せずにトークンの使用量を増大させる、有効ではあるが非効率な推論ステップにペナルティを与えることができません。この制限を系統的に診断するために、循環論法や過度の分解など、5 つの異なるタイプの非効率性を注入した診断ベンチマークである RIV-GSM8K を導入します。診断実験では、最先端の評価者がこれらの非効率性と必要な推論を区別するのに苦労していることが明らかになりました。このギャップに対処するために、私たちは CAID (Context-Aware Information Density) を提案します。これは、有用性の低いステップを特定する、情報理論に基づいたトレーニング不要の指標です。このメトリクスの実用性を検証するために、ポストホック圧縮戦略である PACE 内でメトリクスを適用します。追加の制御実験では、PACE のゲインは単純な枝刈りでは説明できないことが示されています。ランダム ステップ除去や PRM ベースの圧縮ベースラインと比較して、大幅に高い圧縮率でも精度が維持されます。 GSM8K、StrategyQA、ARC-Challenge に関する実証結果は、PACE が精度を維持しながらトークン消費量を 31 ~ 53% 削減することを実証し、CAID が演繹的妥当性を損なうことなく推論チェーンから情報の泡をうまく抽出していることを確認しています。

原文 (English)

Valid $\ne$ Necessary: Diagnosing Latent Inefficiency in Chain-of-Thought

Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs substantial computational costs due to over-reasoning: the generation of redundant, verbose, or irrelevant steps. While existing reasoning step evaluators effectively detect logical fallacies and factual errors, our analysis reveals a critical blind spot: they fail to penalize valid but inefficient reasoning steps that inflate token usage without contributing to the solution. To systematically diagnose this limitation, we introduce RIV-GSM8K, a diagnostic benchmark injected with five distinct types of inefficiencies, including circular reasoning and excessive decomposition. Diagnostic experiments reveal that state-of-the-art evaluators struggle to distinguish these inefficiencies from necessary reasoning. To address this gap, we propose CAID (Context-Aware Information Density), a training-free metric grounded in information theory that identifies low-utility steps. To validate the metric's practical utility, we apply it within PACE, a post-hoc compression strategy. Additional control experiments show that the gains of PACE are not explained by trivial pruning: compared with random step removal and PRM-based compression baselines, it preserves accuracy at substantially higher compression rates. Empirical results on GSM8K, StrategyQA, and ARC-Challenge demonstrate that PACE reduces token consumption by 31-53% while maintaining accuracy, confirming that CAID successfully distills informational froth from reasoning chains without compromising deductive validity.

13:00 JSTエージェント

マルチエージェント証明の自動形式化のための効率的なテスト時間の最適化

完全証明の自動形式化は、自然言語による広範な数学的証明と正式に検証された推論を橋渡しし、検証可能な数学的推論の上限を引き上げる道を提供します。ステートメントレベルの形式化とは異なり、証明の自動形式化は、多くの証明ステップにわたる主張、コンテキスト、依存関係の調整を必要とする長期的な課題ですが、集中的に研究されるようになったのはつい最近です。現在のアプローチは、コストのかかるモデルのトレーニングに依存するか、推論時に過度のガイドなしの修復を適用します。この目的を達成するために、証明の自動形式化を Decomposer-Formalizer-Prover パイプラインとして構造化するマルチエージェント フレームワークである ToMap を導入します。形式的検証と証明品質のセマンティック ルーブリックに基づいて効率的なテスト時間の最適化が行われます。すべてのエージェントにテスト時の計算を分散するのではなく、ボトルネック分析を実行し、Decomposer を重大なボトルネックとして特定します。そのアトミックな自己完結型証明ユニットの品質によって、下流のエージェントが各ステップを正常に形式化して証明できるかどうかが直接決まります。したがって、ToMap は、Formalizer と Prover を下流の実行者として扱い、Decomposer の改良にテスト時の計算を効率的に集中させます。この改良は、GEPA に触発されたループに従い、候補分解に対するプロンプトを進化させ、形式的検証の進行状況と意味論的証明ルーブリックを併用して、次の分解更新を導くパレート フロンティアを定義します。 ProofFlowBench での実験では、構文の正しさと意味の忠実さの両方で評価した場合、ToMap が以前の最良の方法よりも 19.0% 向上し、同時にテスト時間のコストが低いことがわかりました。スケーリング分析により、ほとんどのゲインは分解の進化の数回の反復内で出現し、テスト時の予算の選択の指針となることが示されています。

原文 (English)

Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization

Full-proof autoformalization bridges extensive mathematical proofs in natural language with formally validated reasoning, offering a pathway to elevate the ceiling of verifiable mathematical reasoning. Unlike statement-level formalization, proof autoformalization is a long-horizon challenge requiring coordination of claims, contexts, and dependencies across many proof steps, yet has only recently come under focused study. Current approaches either rely on costly model training or apply excessive, unguided repair at inference time. To this end, we introduce ToMap, a multi-agent framework that structures proof autoformalization as a Decomposer-Formalizer-Prover pipeline with efficient test-time optimization guided by formal verification and semantic rubrics for proof quality. Rather than distributing test-time compute across all agents, we perform bottleneck analysis and identify the Decomposer as the critical bottleneck: the quality of its atomic, self-contained proof units directly determines whether downstream agents can successfully formalize and prove each step. ToMap therefore treats the Formalizer and Prover as downstream executors and efficiently focuses test-time compute on Decomposer refinement. This refinement follows a loop inspired by GEPA, evolving prompts over candidate decompositions and using formal verification progress together with semantic proof rubrics to define a Pareto frontier that guides the next decomposition update. Experiments on ProofFlowBench show that ToMap improves over the best previous method by 19.0% when evaluated by both syntactic correctness and semantic faithfulness, while requiring lower test-time cost. Scaling analysis shows that most gains emerge within a few iterations of decomposition evolution, guiding test-time budget selection.

13:00 JST研究/論文QwenDeepSeek

量子化推論モデルの調整された e-CUSUM デコーディング: トークンの対数確率がデコーディング モニターの観測値として不適切である理由

低ビット量子化により、小規模な推論モデルの導入が安価になりますが、思考連鎖が低下する可能性があります。これにより、生成が信頼できなくなったときに介入するデコーダ側のモニターが起動されます。自然な候補である中心トークンの対数確率増分 $\log p(w_t)+H_t$ が、この目的には間違った観測値であることを示します。モデル自体のサンプリング法則の下では、構造上平均ゼロのマーチンゲールであるため、軌道の健全性ではなくサンプリングの自己一貫性を測定し、$\log p(w_t)$ とエントロピーの両方がゼロに近い自信のある反復中はほぼ沈黙します。我々は、(i) トークンの不確実性と明示的な逐語的反復を融合した縮退対応アラーム スコアと、(ii) 校正された電子プロセスにインスピレーションを得たシーケンシャル検出器を組み合わせた、トレーニング不要のデコード コントローラーを紹介します。生の製品プロセスは条件付き平均 null の下で Ville 有効ですが、スコアは履歴に依存し、自己相関があるため、展開された CUSUM フロア統計は経験的変化検出器として扱われます。 FP16 および INT4 の DeepSeek-R1-Distill-Qwen-1.5B を備えた GSM8K では、キャリブレーションにより、世代の 93 ~ 95% で起動するモニターが、障害のあるトレースの選択的検出器に変わります (0.38 ベース レートに対して $\phi \約 0.3$、精度 $\約 0.6$)。このパイロットでは、コントローラーは測定された逐語的縮退信号を削減し、28% のトークン予算コストで、63% から 69% (ペアの McNemar $p=0.18$、$n=100$) への肯定的だが統計的に決定的ではない INT4 精度の変化をもたらしました。また、GSM8K ではループではなく非終了が主な障害モードであることもわかりました。主な貢献は方法論です。中心トークンの対数確率がデコーダの監視に不適切である理由の説明と、調整され慎重に評価された代替案です。

原文 (English)

Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors

Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought. This motivates decoder-side monitors that intervene when generation becomes unreliable. We show that a natural candidate, the centered token log-probability increment $\log p(w_t)+H_t$, is the wrong observable for this purpose. Under the model's own sampling law it is a mean-zero martingale by construction, so it measures sampling self-consistency rather than trajectory health and is nearly silent during confident repetition, where both $\log p(w_t)$ and entropy are close to zero. We introduce a training-free decoding controller that combines (i) a degeneration-aware alarm score fusing token uncertainty with explicit verbatim repetition and (ii) a calibrated e-process-inspired sequential detector. The raw product process is Ville-valid under a conditional-mean null, while the deployed CUSUM-floored statistic is treated as an empirical change detector because the score is history-dependent and autocorrelated. On GSM8K with DeepSeek-R1-Distill-Qwen-1.5B in FP16 and INT4, calibration turns a monitor that fires on 93--95% of generations into a selective detector of failing traces ($\phi \approx 0.3$, precision $\approx 0.6$ against a 0.38 base rate). In this pilot, the controller reduces measured verbatim-degeneration signals and yields a positive but statistically inconclusive INT4 accuracy change from 63% to 69% (paired McNemar $p=0.18$, $n=100$), at a 28% token-budget cost. We also find that non-termination, rather than looping, is the dominant failure mode on GSM8K. The main contribution is methodological: an explanation of why centered token log-probability is inadequate for decoder monitoring and a calibrated, cautiously evaluated replacement.

13:00 JST研究/論文

検証者がガイドする 12 音構成: 象徴的な音楽生成のための生成、検証、修復のハーネス

大規模な言語モデルは、表面的には正当な 12 音スコアを生成し、それが崩壊して劣化したテクスチャになる可能性があります。シンボリック検証を備えた生成-検証-修復-トレース ループで言語モデル プロポーザーをラップするニューロシンボリック ハーネスを導入します。完全なパイプラインにより、全体の合法性を主張することなく、イベントローカルの一貫性が向上します。 40 の制御タスクと 4 つのペアモデルにわたって、監査済みの配信歩留まりは、生生成の場合の 13.3% から、ハーネスを使用した場合の 48.1% まで上昇しました。ハーネスを使用した場合は、それ以外の場合は明示的に抑制されます。より狭い衝突とシリアル化の一貫性チェックの合格率は 33.5% から 58.3% に上昇しますが、探索的な敵対的プロンプト下を含め、縮退は 0.05 近くにとどまります。 5 人の専門家による盲検評価でも、遵守、認識された合法性、一貫性、および全体的な品質において、生の生成よりもハーネス候補の記述的な総合的な優先順位が示されています。

原文 (English)

Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation

Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate-verify-repair-trace loop with symbolic verification. The complete pipeline improves event-local consistency without claiming whole-piece legality. Across 40 controlled tasks and four paired models, audited delivery yield rises from 13.3% under raw generation to 48.1% with the harness, which explicitly abstains otherwise. The pass rate of a narrower collision and serialisation-consistency check rises from 33.5% to 58.3%, while degeneracy remains near 0.05, including under exploratory adversarial prompting. A blinded evaluation by five experts also shows a descriptive aggregate preference for harness candidates over raw generation in adherence, perceived legality, coherence, and overall quality.

13:00 JST研究/論文

AutoVSR: 回路図からのシンボリック式生成のためのビジュアルからシンボリックへの自動推論

シンボリック式は回路の動作を効果的に特徴付けて予測できますが、回路図から直接それを導き出すのは困難です。このプロセスでは、画像からの回路構造のビジュアルからシンボリックへの正確な構築と、複数ステップのシンボリック導出の正確さが必要であり、どちらも厳密な正確性要件を課します。この研究では、ビジョン言語モデル (VLM) を使用して回路式をビジュアルからシンボリックに生成するための自動フレームワークである AutoVSR を提案します。 AutoVSR は、回路図を実行可能な中間表現 (Executable IR) に再構築し、推論にシンボリック ソルバーを利用することにより、シンボリック式生成の精度を大幅に向上させます。 AutoVSR は 2 つの主要な革新を導入しています。1 つはコンポーネント ルールの取得と検証ベースのフィードバックによって導かれる IR 構築方法、もう 1 つは信頼性の高いマルチステップ導出のためのシンボリック ツール ライブラリを備えたプランニング エージェントとして実装されたシンボリック ソルバーです。エンドツーエンドの VLM アプローチおよびメインのシンボリック式生成タスクにおける特殊な手法と比較して、AutoVSR はそれぞれ 30.01 ~ 59.45% および 41.96 ~ 51.84% の精度向上を達成します。さらに、AutoVSR は、推論コストと計算効率の点で、クローズドソースの最先端の VLM を上回っています。コードは https://github.com/LongfeiLi1/AutoVSR で入手できます。

原文 (English)

AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic

Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. This process requires accurate visual-to-symbolic construction of circuit structure from images and correct multi-step symbolic derivation, both of which impose strict correctness requirements. This work proposes AutoVSR, an automated framework for visual-to-symbolic generation of circuit expressions using Vision Language Models (VLMs). By reconstructing circuit diagrams into an executable intermediate representation (Executable IR) and leveraging a symbolic solver for reasoning, AutoVSR significantly improves the accuracy of symbolic expression generation. AutoVSR introduces two key innovations: an IR construction method guided by component rule retrieval and verification-based feedback, and a symbolic solver implemented as a planning agent equipped with a symbolic tool library for reliable multi-step derivation. Compared with end-to-end VLM approaches and specialized methods on the main symbolic expression generation task, AutoVSR achieves accuracy improvements of 30.01--59.45% and 41.96--51.84%, respectively. Moreover, AutoVSR surpasses closed-source state-of-the-art VLMs in inference cost and computational efficiency. Code is available at https://github.com/LongfeiLi1/AutoVSR.

13:00 JSTLLM/生成AIエージェント

コンパイルしてページング: 実行可能 SOP プログラムと手続き型 LLM エージェントの機能ゲート型ランタイム

企業エージェントは、長期にわたる条件付きの安全性を重視した標準運用手順 (SOP) に従う必要があります。機械可読な SOP 制約を実行可能な疑似コードにコンパイルし、LLM がセマンティック実行を実行している間にアクティブなフレームをページングするプログラムガイド付き (PG) スタック マシンで実行します。 6 つのモデルにわたる 3 アーム SOPBench の調査では、表現と実行時が分離されています。コンパイルされたテキストは決して大幅に損なうことはなく、公式の散文がパフォーマンスを下回る場合でも最大 16.0 ポイント向上します。ランタイム ガイダンスは機能ゲート型です。 2 つの強力なモデルは独立して、正の 7 ドメイン PG コントラスト (58:19 および 75:31 の不一致ペア) を示しますが、弱いモデルは損傷を受けています。フルプログラムのカーソルアブレーション (最初にアクティブなフレーム、完全なプログラムを保持) では、強力なモデルの拒否ゲインの多くが回復します。可視性を選択すると、多少の改善が加えられます。プローブと監査のペアの測定により、この分裂は、再構築能力ではなく自発的な状態規律に基づいて追跡されます。バンクでは、3 つの主要なアームが 70.4、86.4、92.8 に上昇し、100% の拒否精度が得られます。実践的なガイダンス: 最初にコンパイルします。モデルレベルの規律チェックの後にのみアクティブフレームページングを有効にします。

原文 (English)

Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the active frame while an LLM performs semantic execution. A three-arm SOPBench study across six models separates representation from runtime: compiled text never significantly hurts and gains up to 16.0 points where official prose underperforms. Runtime guidance is capability-gated. Two strong models independently show positive seven-domain PG contrasts (58:19 and 75:31 discordant pairs), whereas weak models are harmed. A full-program cursor ablation (active frame first, complete program retained) recovers much of the strong-model refusal gain; selective visibility adds a smaller improvement. Paired probe and audit measurements track this divide to spontaneous state discipline rather than reconstruction ability. On Bank the three primary arms rise from 70.4 to 86.4 to 92.8, with 100% refusal correctness. Practical guidance: compile first; enable active-frame paging only after a model-level discipline check.

13:00 JST研究/論文

From Neural Network Decisions to Training Cases: An Exact Account via Case-Based Decision Theory

Neural networks increasingly guide decisions in high-stakes domains such as medical diagnosis, credit approval, and energy bidding. Audit i…

13:00 JSTビジネス/資金調達

OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis

Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experie…

13:00 JSTLLM/生成AIエージェント

StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure

Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for comp…

13:00 JSTエージェント

Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA

Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web…

13:00 JST研究/論文

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it en…

13:00 JST研究/論文

Enhancing Query Efficiency for d-DNNF Representations Through Preprocessing

In this paper, we investigate preprocessing techniques aimed at improving the efficiency of accessing models of propositional formulas repr…

13:00 JST研究/論文

Comparative Analysis of GAT and BERT for Human-Like Playtesting

Accurately modeling and understanding player experience is crucial for designing engaging puzzle games. To achieve this, a common approach…

13:00 JST研究/論文

Learning Residual Kinematic Corrections for Continuous Neural Decoding via Reinforcement Learning

Decoding continuous three-dimensional (3D) motor imagery (MI) using non-invasive electroencephalography (EEG)-based brain--computer interfa…

13:00 JSTLLM/生成AIハードウェア/半導体

HCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference

Mixture-of-Experts (MoE) large language models (LLM) activate only a small number of experts during inference, but token routing introduces…

13:00 JST研究/論文

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of conte…

13:00 JST研究/論文

Interaction Scaling: Grounding the Third Axis of Test-Time Compute

There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share…

13:00 JSTエージェント

Auditing the Risk Claims of Distributional Reinforcement Learning

Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability…

13:00 JSTLLM/生成AIハードウェア/半導体

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language mode…

13:00 JSTLLM/生成AI

Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling

Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality.…

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction

Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its…

13:00 JSTLLM/生成AIGPT / ChatGPT

Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement

Large language models (LLMs) are rapidly reshaping workplace communication, yet whether AI-assisted writing changes how recipients actually…

13:00 JST研究/論文

Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts

The NIS-2 Directive increases the need for continuous, auditable compliance evidence and motivates a shift from document-based compliance t…

13:00 JST研究/論文

Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-st…

13:00 JST研究/論文

ECG-LDC: A Hardware-Efficient Low-Dimensional Computing Framework for ECG Arrhythmia Classification

Continuous cardiac monitoring in wearable devices demands classifiers that are simultaneously accurate, energy-efficient, and deployable on…

13:00 JST研究/論文

Ablation, Statistical Inference, and Validation for KV-Cache Compression

This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rota…

13:00 JST研究/論文

SciML in the Wild: A Diagnostic Study of When Structural Priors Help and When They Hurt

Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs…

13:00 JSTエージェント

Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems

Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions…

13:00 JSTエージェント

What Context Does a Coding Agent Actually Need to Act?

A modern coding agent can hold an entire repository in its context window. Most of its reading is wasted -- and the interesting question is…

13:00 JSTLLM/生成AI

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning

Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requ…

13:00 JST研究/論文

Mitigating Early Training Collapse in CTR Models

Deep neural models for click-through rate prediction often exhibit a sharp decline in validation performance immediately after the first tr…

13:00 JSTハードウェア/半導体

Model Collapse: On Recursion, Noise, and Uncharted Machine Visions

Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that pro…

13:00 JST研究/論文

The Ramanujan Challenge For AI

To help evaluate the mathematical skills of current AI systems, we present a set of formulas for fundamental mathematical constants. These…

13:00 JST研究/論文

The Universal Language of CSI:Unifying Wireless Sensing Across Devices and Environments

WiFi sensing based on Channel State Information (CSI) promises ubiquitous, device-free perception, yet current research remains trapped in…

13:00 JSTエージェントロボティクス

SWIFT: A Small-World Interaction Framework for Flow-Aware Trajectory Prediction in Autonomous Driving

Accurate trajectory prediction in autonomous driving hinges on modeling dynamic and context-dependent interactions among traffic agents. Ho…

13:00 JST研究/論文

MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms

Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data,…

13:00 JST研究/論文

Data-Driven Forward and Inverse Modeling of V-Beam Thermal Sensors

This paper presents a machine learning framework for data-driven inverse design of V-beam thermal sensors. The goal is to determine the opt…

13:00 JST画像/動画生成

Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that…

13:00 JST画像/動画生成

Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization

Large-scale, multi-subject widefield calcium imaging provides unprecedented access to brain-wide cortical dynamics. However, the high dimen…

13:00 JST画像/動画生成

RSLoRA: Training-free Rank Allocation for LoRA via Representational Sensitivity Probing

Low-Rank Adaptation (LoRA) has become a cornerstone of parameter-efficient fine-tuning (PEFT); however, the conventional practice of unifor…

13:00 JST画像/動画生成

ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams

Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-st…

13:00 JST研究/論文

Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting

Radio frequency fingerprint identification (RFFI) uses transmitter-specific hardware imperfections as a physicallayer identity cue for Inte…

13:00 JST画像/動画生成

Knowledge-Constrained Shape Optimization with a Mixture-of-Experts Neural Operator for High-Confidence Design

Engineering shape optimization faces challenges in both expert-dependent problem setup and surrogate-model reliability. In practical aerody…

13:00 JSTエージェントロボティクスビジネス/資金調達

OmniSCS: Omni Safety-Critical Scenario Synthesis for Autonomous Driving via a Fully Editable Driving World

The synthesis of safety-critical scenarios (SCS) and their evaluation through closed-loop simulations are crucial for developing robust aut…

13:00 JST研究/論文

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content…

13:00 JSTロボティクス

Maximizing Human Efficiency in Large-Scale Robot Post-Training via VLAC-Cut Guided Pipeline

When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post training are required because a single round…

13:00 JST画像/動画生成ロボティクス

Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models

Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- requ…

13:00 JSTエージェントロボティクスビジネス/資金調達

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation

Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prio…

13:00 JSTLLM/生成AIロボティクス

Large Multimodal Model-Based Environment-Aware Mobility Management

Recently, large language models (LLMs) have been successfully adopted in various fields, including wireless communications, robotics, and a…

13:00 JST研究/論文

JEPA for AI-Native 6G: Predictive Representations and Open Challenges

Sixth-generation (6G) networks are moving toward AI-native operation, where learning modules are embedded across the radio access network (…

13:00 JSTLLM/生成AIGoogle

Trivial Prompt Reframing Bypasses Safety Guardrails in Google\'s MedGemma-4B

Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model car…

13:00 JST画像/動画生成

A Unified Model for Highly Accurate ECG-Free Dynamic Coronary Roadmapping Using Spatio-Temporal Transformers

Percutaneous Coronary Intervention (PCI) is a minimally invasive procedure used to restore coronary blood flow obstructed by atheroscleroti…

13:00 JSTエージェント

An Autonomous Scientific Knowledge Generation Framework for AI-Driven Scientific Discovery

Artificial intelligence (AI) is transforming scientific discovery, but its effectiveness is fundamentally limited by the availability of st…

13:00 JST画像/動画生成ロボティクス

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute co…

13:00 JSTLLM/生成AI画像/動画生成エージェント

Memory-Conditioned Tool Calling for Camera-First Visual Agents

Recognition tells an agent what is in an image; personal memory affects what is worth looking up next. In a camera-first setting the user c…

13:00 JSTロボティクス

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning

Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix t…

13:00 JST画像/動画生成

Towards Objective Dysgraphia Detection: A Multi-Branch Deep Learning Approach for Online Handwriting Analysis

Dysgraphia is a specific learning disability that is prevalent among school-age children. It affects handwriting coherence, quality, fluenc…

13:00 JSTロボティクス

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reprod…

13:00 JSTLLM/生成AI画像/動画生成

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large v…

13:00 JSTLLM/生成AI研究/論文

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irr…

13:00 JSTLLM/生成AI

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention

Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimen…

13:00 JST研究/論文

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

Audio deepfake detection models determine whether speech is genuine or artificially generated, but high overall accuracy can mask substanti…

13:00 JST画像/動画生成

Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling

We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride predi…

13:00 JSTエージェント

Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code

Agentic coding tools are increasingly used to make autonomous repository-level changes to real-world projects. Prior work has largely evalu…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their t…

13:00 JST研究/論文

SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification

Cybersecurity systems must adapt rapidly to emerging threats. However, labeled data for new threat categories is unavailable when those thr…

13:00 JSTエージェント

A Knowledge-Based Multi-Agent Framework for Security Control Recommendation

Hardening IT on-premises environments can be a daunting task for teams without access to adequate cybersecurity expertise. In this regard,…

13:00 JST研究/論文

A Foundation Model for Multimodal Event Sequences in Financial Applications

Predictive modeling is a core component of modern financial services, where a wide range of tasks are traditionally addressed using separat…

13:00 JSTLLM/生成AIハードウェア/半導体GPT / ChatGPTGoogle

Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-…

13:00 JST研究/論文

Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization

Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but…

13:00 JSTビジネス/資金調達

A Production-Oriented Framework for Evaluation of SFX Generation

Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity…

13:00 JSTLLM/生成AIエージェント

An LLM-powered Agentic Recommendation System for Connected TV Content Discovery

Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse c…

13:00 JST研究/論文

Beyond Bayesian Nash: Learning Minimax-Regret Equilibria for Adversarial Team Games under Asymmetric Information

Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs,…

13:00 JSTLLM/生成AI

Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, withi…

13:00 JST研究/論文

Geometric mean-based pairwise comparison method with the reference values -- statistical approach

For many years, the pairwise comparison method has been widely used for decision-making involving experts. The best-known example of this m…

13:00 JSTエージェントビジネス/資金調達Claude

Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circ…

13:00 JSTLLM/生成AI

Efficiently Adapting Spoken Language Models for the Singaporean Context

Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially wh…

13:00 JST研究/論文

Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference

Deploying large-scale transformer models on resource-constrained edge devices remains a challenge due to the high energy and memory overhea…

13:00 JSTLLM/生成AI研究/論文

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful complian…

13:00 JSTLLM/生成AIQwen

Cost of Reasoning in non-English Languages: A Case Study on Japanese

Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented…

13:00 JST研究/論文

When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

We study robust generalization under spurious correlations: tasks where a shortcut feature is correlated with the true label in training bu…

13:00 JST研究/論文

A Large-Scale Dataset of MCP Implementations on GitHub

The rapid emergence of the Model Context Protocol (MCP) has introduced a new standard for connecting large language models to external tool…

13:00 JST研究/論文

ML in a Box: Analyzing Containerization Practices in Open Source ML Projects

Containerization has become increasingly essential in the machine learning (ML) domain, providing reproducibility, portability, and environ…

13:00 JSTLLM/生成AI

GAE: Graph-Augmented Evolution for Scientific Discovery via Reinforcement Optimization

Evolutionary program search guided by Large Language Models (LLMs) has emerged as a powerful paradigm for automated scientific discovery. H…

13:00 JST研究/論文

SALT-GNN: Handling Dense Neighborhoods in Anti-Money Laundering Graphs via Statistics-Aware Attention

Money laundering threatens financial stability and exposes institutions to penalties, motivating automated detection. Because laundering sc…

13:00 JSTLLM/生成AI

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each…

13:00 JST画像/動画生成

FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion

Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation. Iterative methods, exemp…

13:00 JSTLLM/生成AI画像/動画生成

EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation

Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the ta…

13:00 JST研究/論文

Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers

It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not…

13:00 JSTLLM/生成AI

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization

Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip…

13:00 JSTロボティクス研究/論文

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception…

13:00 JSTLLM/生成AIハードウェア/半導体

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory…

13:00 JST画像/動画生成

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unr…

13:00 JST研究/論文

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimi…

13:00 JSTLLM/生成AI

Instruction Set and Language for Hypergraphs

We present IsalHG, a method for representing the structure of any finite, connected hypergraph of bounded hyperedge arity as a string over…

13:00 JST研究/論文

Comparing Socially-Equitable Renewable Energy Budget Allocation MDP Policies in Mature and Emerging Economies

Equitable renewable-energy planning is a sequential decision problem, but the decision variables available to a public planner differ sharp…

13:00 JST研究/論文Google

When Does Depth Survive Composition? Compute--Quality Regimes in Latent World Models

Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better…

13:00 JSTロボティクス

Source-Lifted Flow Matching for Intervenable Multimodal Imitation

Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stoc…

13:00 JST研究/論文

Exploratory Analysis of Deep Learning Models for Forecasting Meteorological Parameters in the Agricultural Sector

Accurate meteorological forecasting is essential for agricultural planning, irrigation management, and environmental decision support. This…

13:00 JST画像/動画生成

PhenoEmbed: Self-Supervised Multispectral UAV Time-Series Embeddings for Individual Tree Crown Phenology

Tree crowns are a challenging target for resilient AI because they are not static objects: their spectral response, internal texture, trans…

13:00 JSTLLM/生成AI

Partial Contracts Suffice: Sound, LLM-Inferred Regression Verification

Software evolves continuously, yet ensuring that a patch preserves intended behavior without re-verifying an entire codebase remains diffic…

13:00 JSTエージェント

Program-Synthesis-Driven Autodesign of Universal Unitary Operators

We demonstrate that AI-driven program synthesis can autonomously discover fundamental strategies for decomposing unitary matrices in photon…

13:00 JSTLLM/生成AI

Polarization Detection: A Hybrid Approach with AfroXLMR-Social and DeBERTa for Low- and High-Resource Settings

The rapid proliferation of online polarization threatens social cohesion, necessitating robust automated detection systems that operate eff…

13:00 JSTLLM/生成AI研究/論文

Neutralizing Structural Inequality in the Nigerian FinTech Sector

Algorithmic decision systems in financial services often rely on data proxies that inadvertently encode structural inequalities. This paper…

13:00 JST研究/論文Gemini

From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement

AI visibility measurement is comparative: practitioners want to know which domains generative search engines cite most often and whether ob…

13:00 JST画像/動画生成

GRC-ProbNet: Uncertainty-aware Feature Extraction for Cardiovascular Disease Classification

The automatic detection and classification of cardiovascular disease (CVD) from computed tomography (CT) images plays an important role in…

13:00 JST画像/動画生成研究/論文

Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remai…

13:00 JST研究/論文

A Hyperbolic Neural Closure for M1 Radiation Transfer

In radiation transfer simulations, an M1 method achieves substantial computational savings by replacing the full angular transport equation…

13:00 JSTロボティクス

VINE: Taming Generative Control Policies for Reinforcement Learning

Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from nois…

13:00 JST画像/動画生成ロボティクス

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse…

13:00 JSTLLM/生成AI

Structured Thoughts For Improved Reasoning And Context Pruning

Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient…

13:00 JST研究/論文

The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy

Artificial intelligence (AI) is transforming electron microscopy by enabling quantitative analysis of increasingly large and complex datase…

13:00 JSTLLM/生成AIエージェント

A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization

Although large language models (LLMs) have shown promising potential in news summarization tasks, their performance on long-document summar…

13:00 JST画像/動画生成研究/論文

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MML…

13:00 JSTLLM/生成AI

Large Language Models in Misinformation Ecosystems: Misuse, Defense, and Vulnerability

Large language models (LLMs) have transformed misinformation from a primarily content-centric problem into a broader ecosystem-level securi…

13:00 JST研究/論文

Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control

Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting user…

13:00 JSTLLM/生成AI

Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts

Large Language Models (LLMs) are increasingly used for code smell detection tasks due to their ability to interpret program semantics. Howe…

13:00 JST研究/論文

Learning the Brain's Dynamics as a Port-Hamiltonian System

We model human motor cortex during a wrist-extension BCI task as a port-Hamiltonian system (pHS): a conservative interconnection (gyroscopi…

13:00 JSTLLM/生成AI

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams

Context engineering decides what information a model carries forward, and current designs meter it in tokens: compressing the past into a b…

13:00 JST画像/動画生成

Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer

Layout-based 3D scene synthesizers place each object using two human-annotated channels: a categorical class label and a canonical-pose con…

13:00 JSTLLM/生成AI

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards

Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers…

13:00 JSTLLM/生成AI

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process…

13:00 JSTLLM/生成AIエージェント

Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents

LLM agents can commit durable effects from authority evidence that was valid earlier in execution: a DOM snapshot, approval epoch, version…

13:00 JSTLLM/生成AIエージェント研究/論文

Confining Nondeterminism: AI-Driven Research Systems as DBMSs for Reliable, Non-Wasteful, Transparent, and Collaborative Research [Vision]

LLM agents that conduct research (proposing ideas, writing and running code, analyzing results) can already carry a study from research que…

13:00 JST研究/論文

Conditional Optimal Bridge for Riemannian Activation Steering

Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many exi…

13:00 JSTLLM/生成AI画像/動画生成エージェント

Towards Autonomous and Auditable Medical Imaging Model Development

Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugg…

13:00 JSTLLM/生成AI

Motif: Discovering and Automating Personal Web Workflows

Recent advances in LLMs and existing work on programming by demonstration have made it possible for end users to create automations by expl…

13:00 JSTLLM/生成AI

Tool-Adaptive LLM Reranker

Generative Large Language Models (LLMs) have revolutionized information retrieval, yet their strictly parametric nature frequently leads to…

13:00 JSTエージェントClaudeOpenAI

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation

Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the fiel…

13:00 JST研究/論文

Learning from Local Walks on Dynamic Graphs with Bandit Feedback

We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In t…

13:00 JST画像/動画生成

DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders

AI models are increasingly trained on personal images scraped from social media and public platforms, often without consent, leading to ser…

13:00 JSTLLM/生成AIエージェント

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference

Large language model (LLM) agents accumulate heterogeneous context, including system instructions, plans, user turns, retrieved documents,…

13:00 JST画像/動画生成

WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen

Efficient waste segregation is critical for sustainable urban management and environmental governance. Existing automated systems are limit…

13:00 JSTLLM/生成AI

Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation

We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source, an…

13:00 JSTエージェントロボティクス

World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning

Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepres…

13:00 JST研究/論文

Auditing Construct Overlap in Explainable Machine Learning: Evidence from Burnout-Depression Prediction Across Student Cohorts

Explainable machine learning (XML) pipelines applied to composite mental health outcomes can produce apparently-robust, cross-population-st…

13:00 JSTロボティクス

Coverage Path Planning: Classical Foundations, Recent Advances, and Future Directions

Coverage path planning (CPP) is a fundamental problem in robot motion planning, whose aim is to produce robot trajectories that provide com…

13:00 JSTLLM/生成AI

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks. However, t…

13:00 JST画像/動画生成GPT / ChatGPT

Answer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs

Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled…

13:00 JSTLLM/生成AICopilot

Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications

As AI code tools become integrated into programming environments, students increasingly describe intended behavior in natural language and…

13:00 JST研究/論文

Learning to Fine-tune Foundation Models under Resource Limitations

We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device. At each tim…

13:00 JST画像/動画生成ロボティクス

Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification

The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently…

13:00 JST研究/論文

MDQEC-QAS: Meta-Decoding for Quantum Error Correction with Hardware-Aware VQC Search and Confidence-Gated Recovery

We propose a unified meta-decoding framework for quantum error correction that learns syndrome-to-recovery mappings across multiple stabili…

13:00 JSTLLM/生成AI

PromptGraph: Graph-Guided Prompt Sanitization for Balancing Privacy and Utility in LLM Inference

Large Language Model (LLM) services introduce a fundamental privacy challenge. Sensitive information may be inferred not only from explicit…

13:00 JST研究/論文ClaudeGPT / ChatGPTGemini

Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud

Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required…

13:00 JSTLLM/生成AI

A Corpus of Persuasion Techniques in Slavic Languages

Persuasion techniques are powerful rhetorical devices used to sway public opinion in a wide range of media. We present a new corpus of pers…

13:00 JSTLLM/生成AIエージェント

To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning

Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved ne…

13:00 JST研究/論文

Multi-Scale Convolution with Optimal Transport Attention Effect on Multivariate Time Series

The analysis of Multivariate Time Series (MTS) plays an important role in a lot of real-world practical applications, but it still remains…

13:00 JST研究/論文

Lightning Fast Matching Dependency Discovery with Desbordante

Matching dependency is a generalization of the functional dependency concept, which allows users to apply custom similarity functions for m…

13:00 JST研究/論文

LSTrans: Efficient Knowledge Transfer for Lightweight and Automated ECG Classification

Deploying deep learning models for automated electrocardiogram classification on resource-constrained wearable devices remains challenging…

13:00 JSTLLM/生成AI

Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs

Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and i…

13:00 JSTLLM/生成AI

Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation

Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE…

13:00 JST研究/論文

Diachronic Sample Integration: Robust Tail-Risk Estimation with Generative Models

Deep generative models are increasingly used as simulators for downstream decision-making under data scarcity, but in risk-sensitive applic…

13:00 JSTLLM/生成AIエージェント

Distributed Agent System: Fault-Tolerant Collaboration Among Embodied Agents

AI engineering is shifting from passive text generation by large language models (LLMs) to agent-driven task execution, creating new reliab…

13:00 JSTLLM/生成AIエージェント

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games

Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent d…

13:00 JSTLLM/生成AI

Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization

Opinionated text - spanning product reviews, hotel feedback, and social posts - captures rich signals about user experiences, preferences,…

13:00 JST画像/動画生成ビジネス/資金調達

3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliabi…

13:00 JSTLLM/生成AIエージェント

How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study

The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks wi…

13:00 JSTLLM/生成AI

The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions

In the age of large language models, Natural Language to SQL (NL2SQL) translation remains an open problem with many useful applications. We…

13:00 JST研究/論文

The Singularity Space: A Generative Diffusion Framework for Signal Representation

Generative models often represent signals as dense grids of amplitudes, blurring sharp transients that are crucial for the correctness of p…

13:00 JSTエージェントハードウェア/半導体NVIDIA

Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles

Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringe…

13:00 JST研究/論文

Efficient Online Proportional Sampling with Applications to Smoothed Online Learning

We study the problem of efficient online proportional sampling from a high-dimensional domain under a $\sigma$-smoothed adversary, where th…

13:00 JST研究/論文

CGS: Configurable Graph Summarization with Bounded Neighborhood Loss and Query Support

Given a large graph, how to generate a compact summary graph that is configurable by the user and supports multiple graph queries with eith…

13:00 JST画像/動画生成

EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion

Existing Stochastic 3D Human Motion Prediction models are fundamentally constrained by hard-coding the skeleton kinematics, severely limiti…

13:00 JST画像/動画生成

MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI

Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demand…

13:00 JST画像/動画生成ロボティクス

Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies

As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human…

13:00 JST画像/動画生成

LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI

Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating…

13:00 JSTロボティクス

Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion

We present a manipulation planning system based on affordance recognition and action effect prediction. The system reasons through possible…

13:00 JST研究/論文

Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the…

13:00 JST画像/動画生成

SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remark…

13:00 JST研究/論文

Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration

Young job seekers frequently turn to social media to compare themselves with peers and make sense of career possibilities. However, passive…

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

Large language models (LLMs) are increasingly used in agentic coding settings, where they can inspect files, execute commands, run tests, o…

13:00 JSTLLM/生成AI

Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry

Fruitful collaborations rely on cooperative communications, including of contextual cues to incorporate into reasoning. The increasing use…

13:00 JSTLLM/生成AI画像/動画生成Gemini

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, i…

13:00 JST画像/動画生成

Controlling Motion Transfer in Diffusion Transformers via Attention Heads

Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to mot…

13:00 JSTLLM/生成AIエージェント

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its descripti…

13:00 JST研究/論文

The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved e…

13:00 JST研究/論文

MusicMark: A Robust Generative Watermarking Framework for Music Generation

AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attr…

13:00 JSTエージェントロボティクスClaude

VIA: Visual Interface Agent for Robot Control

Robot manipulation is a complex task that requires visual understanding, physical reasoning, planning, and closed-loop control. General-pur…

13:00 JST研究/論文

BeatEdit: Symbolic Music Generation as Explicit Editing

Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete s…

13:00 JSTLLM/生成AIビジネス/資金調達

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating r…

13:00 JSTロボティクス

Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation

Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D…

13:00 JST研究/論文

RepTran: Search-Based Repair of Transformer Models

To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and…

13:00 JSTLLM/生成AI

ProgramTab: Boosting Table Reasoning of LLMs via Programmatic Paradigm

Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabula…

13:00 JST画像/動画生成

HandFlow: Fully Generative 4D Hand Recovery with Flow Matching

Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce…

13:00 JST研究/論文

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. E…

13:00 JST研究/論文

An Empirical Study for GUI Test Migration from Android to OpenHarmony System

To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existin…

13:00 JSTLLM/生成AIエージェント

Multi-Agent LLMs Fail to Explore Each Other

Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can…

13:00 JSTLLM/生成AI

Enhancing LLMs through human feedback: a journey towards self-improvement

In the rapidly evolving landscape of information retrieval systems, the ability to adapt and improve through user feedback is paramount. Th…

13:00 JSTロボティクス

Towards Predictive, Aligned, and Scalable Robot Learning

Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities…

13:00 JSTLLM/生成AIエージェント

Automated Textbook Auditing with Multi-Agent LLM Systems

Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, dom…

13:00 JST画像/動画生成ビジネス/資金調達

A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation

Comprehensive quantification of cardiac structures from computed tomography (CT) remains limited not by data availability but by the scalab…

13:00 JSTエージェント

Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation

We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, vers…

13:00 JSTLLM/生成AILlama

The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students

As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. Thi…

13:00 JST研究/論文

Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis

The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science educati…

13:00 JSTLLM/生成AI

PRISM Edit: One Vector for All Temporal Answers

Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing lo…

13:00 JST研究/論文

Fail-Aware and Explainable Test Oracle Prediction

Despite their central role in fault detection, test oracles remain challenging to construct effectively. Recent learning based methods addr…

13:00 JST画像/動画生成

Longitudinal Multi-View Breast Cancer Risk Prediction

Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detec…

13:00 JST研究/論文Copilot

Understanding the Impact of AI Code Assistants on Security API Usage: An Empirical Study

AI code assistants are transforming software development, but their implications for software security remain a major concern, particularly…

13:00 JSTLLM/生成AI

Characterising AI Models for Cataloguing

The creation of digital collections involves not only the digitisation of content, but also the creation of catalogue records for it. This…

13:00 JSTLLM/生成AIビジネス/資金調達

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task th…

13:00 JST研究/論文

BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation

Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI.…

13:00 JSTロボティクス

A Glimpse into Long-term Physical Coexistence with Intelligent Robots

Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must suppo…

13:00 JSTLLM/生成AIエージェント

Agentic Routing: The Harness-Native Data Flywheel

Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, con…

13:00 JST画像/動画生成

Uncertainty Quantification for EO Regression Tasks: Building Height, Tree Canopy Height and Above-ground Biomass Estimation

Earth Observation regression tasks such as building height, canopy height, and above-ground biomass estimation underpin critical applicatio…

13:00 JSTLLM/生成AI研究/論文

A Multimodal Dataset for Large Language Model Applications in the Energy Domain

This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications i…

13:00 JSTLLM/生成AI画像/動画生成

LightMem-Ego: Your AI Memory for Everyday Life

Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, a…

13:00 JSTLLM/生成AIエージェントDeepSeek

Agentic Skill Optimization over Lie Algebroids

Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and trace…

13:00 JST研究/論文

IG-GAN: A Generative Adversarial Network for Aerodynamic Data Generation Based on Intrinsic Geometry

Existing generative models learn data distributions in flat Euclidean space. However, most data in our real world are manifolds embedded in…

13:00 JSTロボティクス

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in…

13:00 JSTLLM/生成AI

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization…

13:00 JST研究/論文

CDFM: Towards a General-Purpose Causal Discovery Foundation Model

Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientifi…

13:00 JST研究/論文

Toward Inclusive Avatar Design with Limb Differences Through Artificial Intelligence

As extended reality becomes more popular for social interaction and entertainment, 3D avatars must represent the full diversity of body typ…

13:00 JST画像/動画生成

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a…

13:00 JST研究/論文

AutoMatBench: An Automatic Optimization Toolkit for the Acceleration of Material Properties Prediction Benchmarking

Material property prediction (MPP) infers key properties from chemical composition and structure, accelerating the discovery and optimizati…

13:00 JST画像/動画生成エージェント

Technical Report on the CVPR 2026@AdvML Workshop Challenge

Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report…

13:00 JSTエージェント

Heuristic Learning for Active Flow Control Using Coding Agents

Active flow control involves nonlinear dynamics, partial observations, and computationally expensive simulations, making controller design…

13:00 JST研究/論文

Structure-Feature Aligned Graph Learning via Alternating Constrained Optimization

We introduce a constrained two-view framework for node prediction that aligns structure-conditioned GNN embeddings with a structure-free fe…

13:00 JST画像/動画生成

DiffEEG: A Self-Supervised Denoising Diffusion Model for Learning EEG Generic Representations

Deep learning for EEG-based seizure detection faces critical challenges: severe annotation scarcity and extreme class imbalance, where icta…

13:00 JSTLLM/生成AI

Extending LLM Context via Associative Recurrent Memory

Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain…

13:00 JST画像/動画生成ロボティクスGPT / ChatGPT

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodi…

13:00 JST研究/論文

Closing the Loop: An Access-Control Architecture for Automated, Anomaly-Driven Network Revocation in IoT Deployments

Network-based anomaly detection for IoT devices has matured to the point of reporting strong detection accuracy, yet most published systems…

13:00 JSTLLM/生成AI

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct kn…

13:00 JSTエージェントロボティクス

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-ac…

13:00 JSTLLM/生成AIエージェント研究/論文Claude

Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming

Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety fai…

13:00 JSTLLM/生成AI研究/論文

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represen…

13:00 JSTエージェント

An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory

Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams of…

13:00 JST研究/論文

Active Offline-to-Online Reinforcement Learning

Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subs…

13:00 JST研究/論文

Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories

Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flex…

13:00 JST研究/論文

Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal

Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regulated domains. Although…

13:00 JST画像/動画生成

StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and sto…

13:00 JST研究/論文

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, des…

13:00 JSTLLM/生成AI研究/論文

Introducing Human-Centeredness in AI-Assisted Lexicography

This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers sign…

13:00 JST画像/動画生成エージェントビジネス/資金調達研究/論文

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a state…

13:00 JST研究/論文NVIDIA

Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search

Neural Architecture Search (NAS) has automated the design of deep learning models but traditionally requires massive computational resource…

13:00 JST画像/動画生成研究/論文

LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments

This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthca…

13:00 JSTLLM/生成AI画像/動画生成

Evidence-Backed Video Question Answering

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual ans…

13:00 JSTLLM/生成AIハードウェア/半導体

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and…

13:00 JSTロボティクス

A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, the…

13:00 JST研究/論文

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous…

13:00 JSTLLM/生成AI

Metacognition in LLMs: Foundations, Progress, and Opportunities

Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication,…

13:00 JST研究/論文Claude

Measuring AI Ability to Complete Long Software Tasks

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of A…

13:00 JSTLLM/生成AIGoogleGemini

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

The rise of conversational interfaces has greatly enhanced LLM usability by leveraging dialogue history for sophisticated reasoning. Howeve…

13:00 JSTLLM/生成AI

LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning

Atomic commits, which address a single development concern, are a best practice in software development. In practice, however, developers o…

13:00 JSTエージェント

InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation

Collaborative partnerships play a crucial role in inquiry-oriented education. However, most learning partners are currently assigned throug…

13:00 JST研究/論文

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions

Despite rapid progress in artificial intelligence, current systems struggle with the interconnected challenges that define real-world decis…

13:00 JSTLLM/生成AI

Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs

Personalized large language models are often expected to follow explicit style instructions, yet we find that such instructions can undermi…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文GPT / ChatGPTDeepSeek

BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation

Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely depe…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating…

13:00 JSTエージェントビジネス/資金調達

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide ri…

13:00 JSTエージェント

A Model-Free Universal AI

In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using enviro…

13:00 JST研究/論文

Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions

Applying computational methods to law at scale requires separating genuine legal reasoning from surface similarity. We study this through a…

13:00 JST研究/論文

When Sensing Varies with Contexts: Context Probing for Tactile Few-Shot Class-Incremental Learning

Few-shot class-incremental learning (FSCIL) aims to recognize novel classes from only a few labeled samples while retaining previously lear…

13:00 JSTLLM/生成AIエージェント

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents

Tool-integrated LLMs retrieve information, perform computations, and take real-world actions, but their reliability depends on both tool-us…

13:00 JSTエージェントGoogleGemini

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms…

13:00 JSTエージェントGPT / ChatGPTGrok

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs

We present the Bayesian Linguistic Forecaster (BLF), an agentic system for binary forecasting that achieves state-of-the-art performance on…

13:00 JSTLLM/生成AI

Algorithm Selection with Zero Domain Knowledge via Text Embeddings

We propose a feature-free approach to algorithm selection: instead of hand-crafted instance features, we use pretrained text embeddings. Ou…

13:00 JSTLLM/生成AI

Ideological Bias in LLMs' Economic Causal Reasoning

Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly…

13:00 JSTLLM/生成AIエージェント

Recursive Multi-Agent Systems

Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over lat…

13:00 JSTLLM/生成AIエージェント

A Low-Latency Fraud Detection Layer for Detecting Adversarial Interaction Patterns in LLM-Powered Agents

Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning.…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPTNVIDIA

2.5-D Decomposition for LLM-Based Spatial Construction

Autonomous systems that build structures from natural-language instructions need reliable spatial reasoning, yet large language models (LLM…

13:00 JSTエージェントGPT / ChatGPT

EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents

Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since em…

13:00 JST研究/論文

Learning Developmental Scaffoldings to Guide Self-Organisation

From subcellular structures to entire organisms, many natural systems generate complex organisation through self-organisation: local intera…

13:00 JSTエージェント

MindClaw: 精密な介入のための閉ループの具体化された精神状態推論

Theory of Mind (ToM) を使用すると、エージェントは他のアクターの信念、目標、意図について推論することができます。これは人間中心の身体的支援に不可欠です。既存の ToM ベンチマークは高度なテキスト認識とマルチモーダルな精神状態認識を備えていますが、主にオフラインの質問応答や最終的な行動の予測を評価します。これらは、具体化されたエージェントが変化する環境とのつながりを維持できるかどうか、行為者固有の信念を更新できるかどうか、推論が必要な場合を判断できるかどうか、助けが役立つ場合にのみ介入できるかどうかを完全にテストしていません。 MindPower を基盤として、ロボット中心の ToM 推論をリアルタイムの閉ループ設定に拡張し、精密な介入を伴う身体化された精神状態推論のためのフレームワークである MindClaw を導入します。 MindClaw は、マルチソース入力、信念記憶、身体化された認知トリガー スキル、精神的推論、およびアクション生成を接続し、エージェントが介入が不要な場合は沈黙を保ちながら、適切なタイミングで役立つアクションを出力できるようにします。実験によれば、直接的な VLM ベースラインはタスクの認識と介入の調整に苦労する一方、MindClaw は最高の全体的なパフォーマンスを達成し、閉ループで組み込まれた ToM 支援におけるトリガー スキルの最適化の重要性を示しています。

原文 (English)

MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

Theory of Mind (ToM) enables an agent to reason about another actor's beliefs, goals, and intentions, which is essential for human-centered embodied assistance. Existing ToM benchmarks have advanced text and multimodal mental-state recognition, but they mostly evaluate offline question answering or final action prediction. They do not fully test whether an embodied agent can stay connected to a changing environment, update actor-specific beliefs, decide when reasoning is needed, and intervene only when help is useful. Building on MindPower, we extend robot-centric ToM reasoning to a real-time closed-loop setting and introduce MindClaw, a framework for embodied mental-state reasoning with precision intervention. MindClaw connects multi-source inputs, belief memory, an embodied cognitive trigger skill, mental reasoning, and action generation, allowing the agent to output helpful actions at the right time while remaining silent when intervention is unnecessary. Experiments show that direct VLM baselines struggle with task awareness and intervention calibration, while MindClaw achieves the best overall performance, demonstrating the importance of trigger-skill optimization for closed-loop embodied ToM assistance.

13:00 JSTLLM/生成AI

Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation

Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language mod…

13:00 JSTLLM/生成AI

代表団が過半数を上回るのはいつですか?マルチサンプル LLM 推論のための委任ベースのアグリゲーター

サンプルされた回答に対する多数決は、マルチサンプル LLM 推論の支配的な教師なしアグリゲーターです。各サンプルが運ぶシグナルを委任ベースのアグリゲーター (伝播代理投票、PPV) にパイプすると、MMLU-Pro の多数派を全体で +1.5 pp、非自明なサブセットで +2.24 pp 上回る教師なしコンセンサス ルールが得られることを示します (ペアの McNemar p ~ 1.0e-14、n = 8,099)。マジョリティは、各サンプルが持つ 2 つの自由信号、つまりグループ内の文字エントロピーとグループ間の推論ジオメトリを破棄します。 PPV は、WHEN (有権者が自分の選択にどの程度の重みを保持するか) と WHOM (残りをピア間でどのように分割するか) というシグナルを正確に消費する 2 つの投票者ごとのレバーを公開します。文字エントロピーを使用して WHEN を駆動し、質問ごとの中心の埋め込みコサインを使用して WHOM を駆動します。この方法にはゴールド ラベルや補助トレーニングは必要ありません。質問ごとに、128 のサンプリングされた世代を 16 のグループに分割し、各グループの文字レベルの意味論的エントロピーと推論埋め込みセントロイドを計算し、その両方を定常分布がコンセンサス回答を選択する確率的委任行列に入力します。 PPV が間違った文字で明らかに 10 対 6 の過半数を覆す例を見ていきます。10 人の投票者の多数派クラスターは幾何学的に一貫性がありません (クラスター内平均コサイン -0.02) が、6 人の投票者の少数派は緊密 (+0.26) であるため、エントロピーだけでは多数派が優位に保たれるにもかかわらず、伝播された代表団の集団は少数派の回答に集中します。さらに、教師なし LLM 集約の設計空間を制約する否定的な結果を伴う委任戦略を報告します。信頼モードの質問内アンサンブルがオラクル ギャップを埋めることはありません。

原文 (English)

When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference

Majority voting is the default unsupervised aggregator for multi-sample LLM inference, but it discards two signals: within-group answer entropy and between-group reasoning geometry. We aggregate by delegation instead (Propagational Proxy Voting, PPV): each group of samples keeps weight on its own answer in proportion to its entropy-based confidence (When) and routes the rest to peers by reasoning-embedding similarity (Whom); the stationary distribution of the resulting delegation matrix picks the consensus answer. This requires neither gold labels nor training. On MMLU-Pro with 128 samples per question, delegation beats majority by +1.5 pp overall and +2.24 pp on non-trivial questions (McNemar p ~ 1.0e-14, n = 8,099), overturning wrong majorities whose answer cluster is geometrically incoherent while the correct minority is tight. We then characterize exactly when delegation overturns majority: a two-option model gives a closed-form flip condition on each option's confidence and the weight it routes to the other, with a do-no-harm corollary for near-unanimous questions. The condition calls the realized winner on 96.5% of non-trivial questions, and its predicted mass gap tracks the realized gap at r = 0.97. We did not find any other unsupervised ensemble methods that close the oracle gap.

13:00 JSTエージェント

TouchThinker: 大規模なデータとアクションを意識した表現を使用して、触覚的常識推論をオープンワールドに拡張する

接触は、肉体を持ったエージェントが物理世界を理解するための重要なモダリティです。最近の研究では、触覚常識推論のための言語システムに触覚信号が組み込まれていますが、そのようなシステムを現実的なオープンワールド設定に拡張することは、2 つの重要なボトルネックのため依然として困難です。(1) 現在の触覚推論データセットは形式と規模が制限されたままであり、触覚観察から物理的常識への推論に対する監視が不十分であり、伝達可能な触覚常識の学習を妨げています。 (2) 触覚信号は本質的に冗長でアクション固有ですが、既存の方法ではこれらの特性が見落とされることが多く、その結果、意味表現力が限られた非効率な表現が生じます。これらの制限に対処するために、私たちは、データと表現の両方の観点から触覚の常識的推論をオープンワールドに拡張する触覚言語フレームワークである TouchThinker を提案します。まず、\textbf{415} オブジェクト、\textbf{8} シナリオ、\textbf{7} センサー タイプをカバーする百万規模のマルチソース触覚推論データセットである TouchThinker-1M を構築し、オープンワールドの一般化のための強固なデータ基盤を提供します。さらに、より現実的で多様なタスクを備えたオープンワールドのベンチマークである TouchThinker-Bench を紹介します。次に、触覚表現の効率を向上させ、効率的な推論を可能にするアクション認識モデリングメカニズムを提案します。実験結果は、TouchThinker が複数のデータセットにわたって最先端のモデルに対して競争力のあるパフォーマンスを達成することを示しています。私たちのコードとデータセットは、https://github.com/lvkailin0118/TouchThinker で利用できるようになります。

原文 (English)

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) Tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering \textbf{415} objects, \textbf{8} scenarios, and \textbf{7} sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTGemini

NormAct: 身体化された計画における隠れた社会規範遵守のベンチマーク

マルチモーダル大規模言語モデル (MLLM) は、自己中心的な環境で具体化されたプランナーとして導入されることが増えています。タスクを成功させるには、指示された目標を達成するだけでなく、社会的に適切な方法で行動する必要があります。明示的な目標によって特定の行動が最適化される場合もありますが、暗黙の社会規範によって隠れた制約が課せられることがよくあります。既存の評価は通常、明示的な目標達成または直接的な規範知識に焦点を当てており、計画者がアクション シーケンス内のこれらの隠れた制約を推測して適用できるかどうかを評価することはほとんどありません。目標達成、規範遵守、全体的なタスクの成功に関する計画を評価する、具体化された社会規範の相互作用のベンチマークである NormAct を紹介します。 NormAct は通常のタスク内に隠れた規範を独自に埋め込み、明示的な指示なしにモデルがそれらを実現できるかどうかをテストします。最先端の MLLM (GPT-5.4、Claude Opus 4.7、Gemini 3 Pro) を用いた実験では、大きなギャップが明らかになりました。モデルは 67.3\% のケースで明示的な目標を達成しましたが、隠れた基準に準拠したのは 26.4\% のみでした。合図条件の実験によると、このギャップは一般的な社会知識の欠如ではなく、文脈の中で関連する規範を活性化し、基礎づける際の課題から生じていることが示されています。これに対処するために、計画前にシーン関連の規範を推測するコンテキスト条件付きキュー ジェネレーターである NormPerceptor を提案し、タスクの成功率を 24.2\% から 46.7\% に向上させます。私たちの結果は、身体化されたエージェントが隠れた規範を積極的に検出し、視覚的な証拠に基づいて行動計画の制約として統合できるようにすることの重要性を強調しています。私たちのベンチマークは https://huggingface.co/datasets/Caleb196x/NormAct で公開されています。

原文 (English)

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task success requires not only achieving instructed goals but also acting in socially appropriate ways. While explicit goals may render certain actions optimal, implicit social norms often impose hidden constraints. Existing evaluations typically focus on explicit goal achievement or direct norm knowledge, seldom assessing whether planners can infer and apply these hidden constraints within action sequences. We introduce NormAct, a benchmark for embodied social-norm interactions that evaluates plans on Goal Achievement, Norm Compliance, and overall Task Success. NormAct uniquely embeds hidden norms within ordinary tasks, testing whether models can realize them without explicit instruction. Experiments with state-of-the-art MLLMs (GPT-5.4, Claude Opus 4.7, Gemini 3 Pro) reveal a significant gap: models achieve explicit goals in 67.3\% of cases, but comply with hidden norms in only 26.4\%. Cue-condition experiments indicate that this gap stems not from a lack of general social knowledge, but from challenges in activating and grounding relevant norms in context. To address this, we propose NormPerceptor, a context-conditioned cue generator that infers scene-relevant norms prior to planning, increasing Task Success from 24.2\% to 46.7\%. Our results underscore the importance of enabling embodied agents to proactively detect hidden norms, ground them in visual evidence, and integrate them as action-planning constraints. Our benchmark is publicly available at https://huggingface.co/datasets/Caleb196x/NormAct.

13:00 JSTLLM/生成AIエージェント

大規模言語モデルのエージェントワークフローの特徴付け: N8n エコシステムに関する研究

大規模言語モデル (LLM) は、ローコードおよびノー​​コードの自動化プラットフォームで急速に採用されており、専門家以外のユーザーが自然言語の理解を外部サービスや API と組み合わせたワークフローを設計します。 LLM エージェントは、複雑な複数ステップのタスクを推論、計画し、自律的に実行するためのコア「頭脳」として LLM を使用する LLM システムです。このペーパーでは、ローコード自動化プラットフォームにおける LLM エージェント ワークフローに関する最初の大規模な実証研究を紹介します。私たちは、公開されている 6,000 を超える n8n ワークフローを分析し、その設計の 4 つの側面 (タスク分散、構造およびツールの使用パターン、信頼性メカニズム、自律性レベル) を調査します。私たちの分析では、LLM ワークフローが単なる即時応答パイプラインではないことが示されています。代わりに、LLM は通常、制御ロジック、外部ツール、通信サービス、ストレージ システム、人間によるレビュー ポイントを含む、より広範な自動化構造内に組み込まれます。さらに、多くのワークフローには、LLM 実行後の軽量の後処理またはルーティング ロジックが含まれている一方で、構造化されたフォールバック パス、修復ループ、障害固有のアラート、人間による承認ゲートなどの明示的な信頼性メカニズムは、依然として比較的一般的ではないことがわかりました。これらの結果は、実際の自動化エコシステムにおける LLM エージェントの導入の増加と、信頼性、安全性、ガバナンスに対する限られたエンジニアリング サポートとの間にギャップがあることを明らかにしています。全体として、私たちの研究は、現実世界の LLM エージェント ワークフローを理解して改善しようとしている研究者、プラットフォーム開発者、実践者に 10 の実証結果と 5 つの研究の成果を提供します。

原文 (English)

Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem

Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs. LLM agents are LLM systems that use LLMs as a core "brain" to reason, plan, and autonomously execute complex, multi-step tasks. In this paper, we present the first large-scale empirical study of LLM agentic workflows in low-code automation platforms. We analyze more than 6,000 publicly available n8n workflows and examine four aspects of their design: task distribution, structural and tool use patterns, reliability mechanisms, and autonomy levels. Our analysis shows that LLM workflows are not merely prompt response pipelines. Instead, LLMs are commonly embedded within broader automation structures involving control logic, external tools, communication services, storage systems, and human review points. We further find that while many workflows include lightweight post-processing or routing logic after LLM execution, explicit reliability mechanisms such as structured fallback paths, repair loops, failure-specific alerts, and human approval gates remain relatively uncommon. These results reveal a gap between the increasing deployment of LLM agents in practical automation ecosystems and the limited engineering support for reliability, safety, and governance. Overall, our study provides ten empirical findings and five research takeaways for researchers, platform developers, and practitioners seeking to understand and improve real-world LLM agentic workflows.

13:00 JST研究/論文

利害関係のない AI 予測器における正直さによる安全性

AI システムの能力が高まるにつれて、下流の結果を最適化するトレーニング手順では、暗黙的なエージェンシー、つまり設計者が指定したことのない目標指向の行動が導入される危険性があります。我々は、「認識論的に文脈化された」自然言語ステートメントのデータセットに基づいて条件付けされたベイズ事後分布を近似するように訓練された、Scientist AI (SAI) Predictor の正式な安全性の議論を提示します。私たちは、そのような予測器は、それ自体が目標を達成するために出力を選択するエージェントではなくても、エージェント、アクション、およびその結果を正直に予測できると主張します。これはデータ表現とトレーニング手順に基づいています。テキストの認識論的文脈化は、潜在的な事実の主張とコミュニケーション行為を区別するため、目標の表現は、モデルの採用を推進するものではなく、説明される証拠として扱われます。事後探索型のトレーニング目標により、これは、Predictor を調整された慎重な予測に向けて駆動することを目的としています。トレーニングは進行するため、予測の展開による下流の効果が報酬シグナルとして機能することはありません。システムが必要とするあらゆるエージェンシーは、ガードレールによって制約された明示的な足場によって提供されます。トレーニングのダイナミクスと、議論された危険な予測子の希薄さに関する仮定の下で、トレーニングによって、指定されたしきい値を超える残留害をもたらす保護された展開を持つ予測子が生成される確率は小さいことを証明します。危険な予測子は、多くのクエリにわたって調整された方法で害を過小評価する必要がありますが、そのような調整されたパターンは初期化分布の下ではまれであり、直接トレーニング信号を受け取りません。正確性を確保するための制約は、調整された欺瞞を高コストにする制約と同じであるため、このフレームワークでは安全性と正確性が同時にサポートされています。 Predictor 自体の内部から生じる不整合やエージェンシーに対するこれらの保証は、Predictor をエージェント システムの一部として使用することを妨げるものではありません。

原文 (English)

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.

13:00 JSTエージェント研究/論文

FARS: 大規模に導入された完全に自動化された研究システム

最近の自動研究システムは、言語モデル エージェントが仮説を生成し、実験を実行し、完全な原稿を書くことができることを示していますが、ほとんどの証拠は依然として、選択された例、人間が組み立てたトピック、またはいくつかの事前定義された研究タスクから得られます。私たちは、研究テーマ全体にわたって大規模に動作するように設計された完全に自動化された AI 対 AI 研究システムである FARS (Fully Automated Research System) を紹介します。 FARS は、提案、コード、ログ、結果、原稿を記録する共有ワークスペースを通じて調整された段階固有のエージェントを使用して、アイデア出し、計画、実験、執筆を通じてプロジェクトを自律的に生成し、推進します。最初の公開展開で、FARS は 67 のきめ細かい AI/ML トピックにまたがる 166 の完全な研究論文を作成しましたが、中間成果物は厳選された成功セットではなく、監査可能なコーパスとして保存されました。このコーパスは、全体の評価、サブスコア、完全性チェック、LLM 使用の開示を含む、140 件の論文をカバーするボランティアの査読者による 282 件の構造化レビューによって評価されます。レビューでは、FARS が大規模な公共展開においてレビューに値する、場合によっては強力な AI/ML 研究成果物を生成できる一方で、狭い実験範囲、方法論的な制限、完全性の問題で繰り返される失敗モードも明らかにすることが示されています。

原文 (English)

FARS: A Fully Automated Research System Deployed at Scale

Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale. FARS autonomously generates and advances projects through ideation, planning, experimentation, and writing, using stage-specific agents coordinated through a shared workspace that records proposals, code, logs, results, and manuscripts. In its first public deployment, FARS produced 166 complete research papers spanning 67 fine-grained AI/ML topics while preserving intermediate artifacts as an auditable corpus rather than a curated set of successes. We evaluate this corpus with 282 structured reviews from volunteer reviewers covering 140 papers, including overall ratings, sub-scores, integrity checks, and LLM-use disclosure. The reviews indicate that FARS can produce review-worthy and occasionally strong AI/ML research artifacts in a large-scale public deployment, while also exposing recurring failure modes in narrow experimental scope, methodological limitations, and integrity issues.

13:00 JSTLLM/生成AIエージェント

決定論的で自己拡張的な反応分類のための検証可能なルールをエージェント的に生成

コンピューター支援合成計画では、各変換に決定論的で解釈可能なラベルを割り当てる反応ルールの大きなライブラリを使用して、標的分子をアクセス可能な前駆体に分割します。しかし、化学はロングテールであるため、手動エンコーディングが困難であり、既存のツールは新しい化学に適応できない固定ルールセットに依存しています。ここでは、大規模言語モデル (LLM) のマルチエージェント フレームワークが反応を分類し、665,901 件の米国特許反応にわたってルール自体を記述し、コーパスに対してテストする検証ループの下で各ルールを生成する、完全に自動化されたパイプラインを紹介します。人間によるキュレーションを行わずに、標準分類を 68 クラスから 14,073 クラスに拡張します。軽量の指紋分類器を使用して、目に見えない反応の 97.7% を分類し、主要な独自の分類器と一致しながら、化学をより細かく分解し、オンデマンドでトレーニング ディストリビューション外の化学に拡張します。その結果、生きた反応性データベースと、生成モデルを信頼性の高い自己拡張型のシンボリック システムに変えるための一般的なルートが得られます。

原文 (English)

Agentic generation of verifiable rules for deterministic, self-expanding reaction classification

Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encoding intractable, and existing tools rely on fixed rulesets that cannot adapt to new chemistries. Here we present a fully automated pipeline in which a multi-agent framework of large language models (LLMs) classifies reactions and writes the rules themselves across 665,901 US patent reactions, generating each rule under a verification loop that tests it against the corpus. It expands a standard taxonomy from 68 to 14,073 classes without human curation. With a lightweight fingerprint classifier, it classifies 97.7\% of unseen reactions, matching a leading proprietary classifier while resolving chemistry more finely and extending on demand to chemistry outside its training distribution. The result is a living reactivity database and a general route to turning generative models into reliable, self-expanding symbolic systems.

13:00 JSTエージェント

Raw-ECG-Replay-Free 継続的 ECG 導入における自律的なソース推論から専門家の保持を分離する

マルチソース ECG 展開では、以前の生の ECG を保持または再生できない場合、モデルに新しいデータ ソースを組み込む必要がある場合があります。事前トレーニングされたバックボーンをフリーズし、各ソースに分離された分類子を割り当てることでパラメータの干渉を防ぐことができますが、ソースのメタデータが利用できない場合でも展開には専門家を選択する必要があります。私たちは、凍結された 1024 次元の ECGFounder 機能に基づいて構築された増分エキスパート バンクである \ours{} を通じて、この違いを研究しています。到着するドメインごとにバランスの取れたソフトマックス線形エキスパートが追加されますが、軽量ルーターは、これまでに観察されたソースからの保持されたトレーニング特徴とドメイン ラベルにのみ適合します。検証によって調整されたマージン ルールは、単一のルーティングされたエキスパートにコミットするのではなく、最も可能性の高い 2 人のエキスパートを融合します。 CPSC、PTB-XL、ジョージア、およびチャップマン-紹興では、ソースを意識した専門家の選択は $0.7915\pm0.0036$ マクロ F1 に達し、一致するオフラインの独立ヘッド参照は $0.7885\pm0.0009$ に達し、強力なソースを認識した専門家の保持をサポートしています。ソース ID がない場合、MLP ルーターは $0.7756\pm0.0027$ に達し、トップ 2 マージン フュージョンは $0.7782\pm0.0022$ に達します。ハード MLP ルーティングに対する上位 2 のゲインは小さく ($+0.0026$)、ペアのブートストラップからの 95\% 信頼区間にはゼロが含まれます。 3 つのドメイン注文全体で、トップ 2 とオラクルの差は依然として 0.0111 ドルから 0.0133 ドルであり、自律的なソース推論が残りの主なボトルネックであることがわかります。生の ECG は再生されませんが、凍結されたトレーニング機能はルーターの更新のために保持されます。したがって、このメソッドはメモリフリーではありません。

原文 (English)

Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment

In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freezing a pretrained backbone and assigning each source an isolated classifier prevents parameter interference, but deployment still requires selecting an expert when source metadata are unavailable. We study this distinction through IRFE-ECG, an incremental expert bank built on frozen 1024-dimensional ECGFounder features. Each arriving domain adds a balanced-softmax linear expert, while a lightweight router is fitted only on retained training features and domain labels from sources observed so far. A validation-calibrated margin rule fuses the two most likely experts instead of committing to a single routed expert. On CPSC, PTB-XL, Georgia, and Chapman-Shaoxing, source-aware expert selection reaches $0.7915\pm0.0036$ Macro-F1 and a matched offline independent-head reference reaches $0.7885\pm0.0009$, supporting strong source-aware expert retention. Without source IDs, an MLP router reaches $0.7756\pm0.0027$ and top-2 margin fusion reaches $0.7782\pm0.0022$. The top-2 gain over hard MLP routing is small ($+0.0026$), with a 95\% confidence interval from paired bootstrap that includes zero. Across three domain orders, the top-2-to-oracle gap remains $0.0111$--$0.0133$, identifying autonomous source inference as the main remaining bottleneck. No raw ECGs are replayed, but frozen training features are retained for router updates; the method is therefore not memory-free.Code is available at https://github.com/yufanlu221/IRFE-ECG.

13:00 JST研究/論文

SUNTA: サプライズベースのチャンキングによる階層型ビデオ予測

階層状態空間モデル (HSSM) は、シーケンスを時間的なチャンクに分割することにより、長期予測への有望なアプローチを提供します。ただし、そのパフォーマンスはチャンク境界がどのように決定されるかによって決まります。従来の HSSM は通常、固定長のチャンキングまたは類似性に基づく境界検出に依存していますが、これらの方法はデータの固有の時間構造と一致しないことがよくあります。私たちは、チャンク化は予測誤差によって駆動されるべきだと主張します。予測誤差は、より直接的に、長距離のコンテキストが必要になる時期を示します。それにもかかわらず、サプライズベースのチャンキングを HSSM に統合すると、エンドツーエンドのトレーニング中の階層の崩壊や、開ループ予測中のサプライズ信号の欠如など、重大な課題が生じます。これらの問題に対処するために、我々はサプライズベースの入れ子型時間抽象化 (SUNTA) を提案します。これは、サプライズ信号を保存するために分離されたトレーニング戦略を採用し、内部の不一致をトップダウンのサプライズメトリクスとして使用して、想像上のロールアウト内のチャンク境界を決定する方法です。 2D および 3D 環境でのビデオ予測タスクの実験では、SUNTA がベースラインを上回り、250 タイムステップにわたって正確な予測を独自に維持する一方、最初の 10 タイムステップ以内にすべてのベースラインが低下することが実証されました。

原文 (English)

SUNTA: Hierarchical Video Prediction with Surprise-based Chunking

Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how chunk boundaries are determined. While prior HSSMs typically rely on fixed-length chunking or similarity-based boundary detection, these methods often misalign with the intrinsic temporal structure of the data. We argue that chunking should instead be driven by prediction errors, which more directly indicate when longer-range context becomes necessary. Nevertheless, integrating surprise-based chunking into HSSMs introduces critical challenges, including hierarchical collapse during end-to-end training and the absence of surprise signals during open-loop prediction. To address these issues, we propose Surprise-based Nested Temporal Abstraction (SUNTA), a method that employs a decoupled training strategy to preserve surprise signals and uses internal inconsistency as a top-down surprise metric to determine chunk boundaries within imagined rollouts. Experiments on video prediction tasks in 2D and 3D environments demonstrate that SUNTA outperforms baselines, uniquely maintaining accurate predictions over 250 timesteps, whereas all baselines degrade within the first 10 timesteps.

13:00 JSTエージェントDeepSeek

エージェント ステップ値: 状態接地 LLM エバリュエーターによる状態遷移測定

ほとんどのエージェント評価では、複数ステップのトレースが最終的な回答、成功フラグ、または軌跡レベルのスコアにまとめられます。これらの集計では、開発者が最も必要とする診断の質問、つまりどのアクションが状態を有益な方向に変更したのかがわかりにくくなります。我々は、状態遷移測定フレームワークであるエージェント ステップ値 (ASV) を導入します。これは、観察された各アクションを、固定された候補結果に対する状態に基づいた評価者の分布に誘発する変化によってスコア付けします。 ASV は、編集された前後の状態予測をレンダリングし、ステートレス LLM エバリュエーターを使用して候補ログ スコアを割り当て、ゴールドフリーの信念診断とオフライン オラクル検証メトリクスの両方をレポートします。ラベルフリーの理論的パスにより、評価者の審議がワントークンオプションのスコアリングから分離され、リークやフロアスコアイベントを明らかにしながら候補の可能性が維持されます。ライブ PubMed 検索、部分的にライブ DeepSeek アクター、および DeepSeek 対数確率スコアリングを使用した 100 件のレビュー済みオープン QA 証拠探索タスクについて、ASV は 1,100 のステップと 2,200 の状態を評価します。固定レイアウトの根拠条件付きプロトコルの下では、平均ゴールドマージンゲインは -2.335 (軌道ブートストラップ 95\% CI [-3.395, -1.272])、エントロピーの動きは 0.000、平均ベイジアンサプライズは 2.693 です。したがって、ASV は、最終回答スコアとエントロピーのみのステップ メトリクスが見逃している建設的および破壊的な信念ピボットを特定します。スタンドアロンの ASV Eval ツールキットをリリースします。

原文 (English)

Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces

When evaluator-derived step rewards are pooled or compared across scoring channels, their sign is treated as transportable. Yet the same frozen transition can change sign with the scoring channel. Process rewards vary agent states, while evaluator audits vary scoring configurations; neither first difference isolates their interaction. We define Agent Step Value (ASV) as channel-indexed target-margin gain and identify the missing state-by-channel interaction by replaying complete faces. Of 1,100 PubMed open question-answering transitions, 1,004 were complete across four cyclic layouts. Their mean update is positive under direct scoring (+0.163 [0.102, 0.218]) and negative with an externally generated view (-0.160 [-0.244, -0.079]), giving an interaction of -0.323 [-0.418, -0.232]. Across paired transitions, 507/1,004 cross zero in one direction or the other. Matched templates isolate a generated-minus-quote interaction of -1.121 [-1.534, -0.703]. Its direction remains negative after the readout and evaluator-stack changes, and quote yields a higher paired area under the receiver operating characteristic curve (AUC) for stored success at all three bridge vertices. Because the two views preserve the same task information conditional on the retained state, a representation-invariance null predicts equal responses and success rankings across them. ASV provides a transport audit for evaluator-derived step measurements; causal actor credit lies outside its estimand.

13:00 JST研究/論文

MoP-JEPA: 確率的 JEPA 世界モデル用のハード割り当て予測子混合物

JEPA ワールド モデルは、潜在回帰によってトレーニングされた単一の決定論的予測子を使用して、次の潜在状態を予測します。環境が確率的である場合、これは構造的に失敗することを示します。分岐遷移では、回帰最適予測子は後続埋め込みの条件付き平均、つまりどの状態にも対応しない真の次の状態の間の点を出力します。我々は、決定論的でゲートされた専門家の混合予測器についてこの崩壊を証明し、MoP-JEPA のハード割り当てされた予測器が代わりに遷移分布の量子化器に収束することを証明します。つまり、後続モードごとに 1 つのヘッドであり、単一の前方パスで列挙可能であり、プランナが使用するインターフェイスです。リークのない評価による公式の OGBench オフライン データでは、単一予測子のロールアウトに関する計画のパフォーマンスは低く ($0.02$ ~ $0.09$ 成功)、一方、予測モードによる計画のパフォーマンスは最大 $0.85$ に達し、すべてのタスクで決定論的予測子、ゲート MoE および変分予測子を上回っています。多重予測評価はカバレッジ フリーローディングを引き起こすため、検証プロトコルはメソッドの一部です。入力に依存しないコードブック コントロール、シャッフル コンテキスト テスト、ルーター ゲートによる読み出し、遷移精度ガード、モデルが遷移グラフをブラインドで提案し、グランド トゥルースは結果をチェックするためだけに使用される検証済みルート基準です。この基準の下では、私たちの方法は 3 つの迷路すべて ($2$--$5\times$) で最も強力なソフト代替案を上回っており、プロトコルは、そのベースラインの生スコアの残りのギャップを、存在しない予測遷移を通るルートとして特定します。同じモデルが実際の環境で実行され、最も難しい迷路で公開されている OGBench ベースラインに対して 7 件中 2 位に位置します。 JEPA 世界モデルがそもそも計画できるかどうかは、マルチモーダルなダイナミクスによって決まります。予測子とハード割り当てを混合することは、最小限の検証可能な修正です。

原文 (English)

MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models

JEPA world models commonly predict the next latent state with one regressor. Under stochastic transitions, squared and cosine regression return the conditional mean and its normalized direction, respectively: a single compromise that may match no valid successor. MoP-JEPA instead uses $K$ hard-assigned heads and a context-only router to produce a finite candidate set in one pass. On held-out OGBench transitions, graph search with single-output predictors succeeds on $0.02$--$0.09$ of queries, whereas MoP-JEPA reaches $0.85$. To distinguish useful successors from indiscriminate coverage, we also measure verified-route success (\emph{realroute}), which checks after graph construction whether the proposal contains a path of real transitions. MoP-JEPA leads this same-protocol metric on all three mazes; an MDN attains high raw coverage but predicts many nonexistent edges.

13:00 JST研究/論文

預言者が予測市場で利益を得るのはいつですか?

予測市場は、分散した信念を価格に集約し、不確実な出来事の確率的予測として機能します。古典的な理論では、予測精度と取引利益の間に明確な同等性が確立されていますが、それは特定の自動マーケット メーカー (AMM) 設計に限られます。しかし、今日の最大規模の取引所は中央の指値注文帳に基づいており、情報に基づいた予測担当者は日常的に損失を出しますが、情報に基づいていない戦略は単純なヒューリスティックで利益を得ることができます。私たちは、予測精度と収益性の形式的な同等性を確立することで、この矛盾を解決します。厳密に適切なスコアリング ルール $S$ の場合、予測者の予測 $\mathbf{p}$ と市場価格 $\mathbf{q}$ のみに依存する「適切な」賭け戦略を示し、$S$ の下で $\mathbf{p}$ が $\mathbf{q}$ を上回り、市場に十分な流動性がある場合には必ずプラスの期待利益を獲得します。さらに、この適切な賭けは、本質的に、これほど確実な収益性が保証される唯一の戦略です。この証明は、古典的な AMM 保証を厳密に一般化する期待利益の分解に基づいており、精度の優位性がなくても戦略がどのように利益を得ることができるかを説明します。経験的に、AI モデルによる何千もの予測にわたって、正確さを確実に利益に変える唯一の戦略は適切なベッティングであり、さらに体系的な予測ペルソナを特定し、最適な適切な戦略がどのように異なるかを示します。 Kalshi での 1 か月間にわたるライブ デプロイメントにより、シャープ レシオ 3.35 ドルで $+80.33\%$ の投資収益率が達成されました。

原文 (English)

When do prophets profit in prediction markets?

Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establishes a clean equivalence between forecasting accuracy and trading profit, but only for the specific automated market maker (AMM) design. However, the largest exchanges today are based on central limit order books in which informed forecasters routinely lose money while uninformed strategies can profit on simple heuristics. We resolve this discrepancy by establishing a formal equivalence between predictive accuracy and profitability. For any strictly proper scoring rule $S$, we exhibit a "proper" betting strategy that depends only on the forecaster's prediction $\mathbf{p}$ and the market price $\mathbf{q}$, and earns positive expected profit whenever $\mathbf{p}$ outperforms $\mathbf{q}$ under $S$ and the market has sufficient liquidity. Moreover, this proper betting is essentially the only strategy with such robust profitability guarantee. The proof rests on a decomposition of expected profit that strictly generalizes the classical AMM guarantee and also explains how strategies can profit without an accuracy edge. Empirically, across thousands of forecasts by AI models, proper betting is the only strategy that reliably converts accuracy into profit, and we further identify systematic forecasting personas and show how the optimal proper strategy varies across them. A month-long live deployment on Kalshi achieves $+80.33\%$ return on investment with a Sharpe ratio of $3.35$.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

理由は少なく、もっと検証してください: 決定論的ゲートは、ツールを使用する LLM エージェントのサイレント ポリシー違反障害モードを回復します

ツールを使用する LLM エージェントは、タスクを正常に完了したように見えながら、導入されたポリシーそのものに違反する可能性があります。ポリシーが許可された環境では、対応する状態遷移がドメイン ポリシーによって禁止されている場合でも、ツールは適切な形式の呼び出しを実行できます。その結果、ツールもエージェントの自己報告も明らかにしない、静かな間違った状態 (予約がキャンセルされ、乗客数が変更され、検証なしで請求が処理された) が生じます。私たちはこの故障モードを $\tau^2$-bench 航空会社ドメインで研究します。低予算エージェントでは、観察された障害の 78% は、ツール エラーのないサイレントな間違った状態の障害であり、総障害率は、サンプリング ノイズではなく、ばらばらのシード間で再現可能です。次に、軽量介入を評価します。これは、書き込みを許可する前に、提案された呼び出しと現在の状態を検査する決定論的な読み取り専用の実行前ゲートです。 4 ゲート スイートでは、gpt-4o-mini でのフルベンチマークの成功率が 29.6% から 42.0% に上昇し (+12.4pp; ペアのタスクレベル ブートストラップ P=0.0012)、リフトは素の 15 シード セットで再現されました (+12.3pp; P=0.0008)。効果はゲートが発火する場所に集中します。26/50 発火タスクでは成功率が +19.2pp 上昇しますが、24 発の非発火タスクでの動きはゼロを除外しません。 2 つのネガティブ コントロール (自己強制的な小売ドメインと BFCL) がこのメカニズムを制限します。ゲートは、ツールがポリシーに寛容な場合に役立ちますが、ツールが既に自己強制している場合にはほとんど追加されません。中心的な主張ではなく示唆的な証拠として、同じ障害モードがフロンティアで持続します。デフォルトの推論では gpt-5.2 は依然としてポリシー違反の書き込みを試行し、同じスイートでは成功率が 61.2% から 71.6% に向上します (+10.4pp; P=0.020; n=5、レプリケーションなし)。貢献は、制限された評価と信頼性の結果です。決定論的ゲートはタスクの成功を保証しませんが、アクションの境界で既知のクラスのサイレント ポリシー違反書き込みを決定論的に阻止できます。

原文 (English)

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transition is forbidden by domain policy. The result is a silent wrong state (a booking cancelled, a passenger count changed, a claim acted on without verification) that neither the tool nor the agent's self-report exposes. We study this failure mode in the $\tau^2$-bench airline domain. On a budget agent, 78% of observed failures are silent wrong-state failures with no tool error, and the aggregate failure rate is reproducible across disjoint seeds, not sampling noise. We then evaluate a lightweight intervention: deterministic, read-only pre-execution gates that inspect the proposed call and current state before allowing a write. A four-gate suite raises full-benchmark success from 29.6% to 42.0% on gpt-4o-mini (+12.4pp; paired task-level bootstrap P=0.0012), and the lift reproduces on a disjoint 15-seed set (+12.3pp; P=0.0008). The effect is concentrated where the gates fire: on the 26/50 firing tasks, success rises by +19.2pp, while movement on the 24 non-firing tasks does not exclude zero. Two negative controls (a self-enforcing retail domain and BFCL) bound the mechanism: gates help when tools are policy-permissive and add little where tools already self-enforce. As suggestive evidence, not a central claim, the same failure mode persists at the frontier: gpt-5.2 at default reasoning still attempts policy-violating writes, and the same suite improves success from 61.2% to 71.6% (+10.4pp; P=0.020; n=5, no replication). The contribution is a bounded evaluation and reliability result: deterministic gates do not guarantee task success, but they can deterministically prevent a known class of silent policy-violating writes at the action boundary.

13:00 JSTエージェント研究/論文

Long-Horizo​​n-terminal-bench: 高密度の報酬ベースの評価を使用して、長期期間のターミナル タスクにおけるエージェントの限界をテストする

AI エージェントは、短く明確に指定されたタスクを自律的に完了できるようになりました。ただし、既存の端末ベンチマークは主に、数分以内に終了する単純な問題に焦点を当てており、最終的な結果によってのみ評価されます。この設定では、中間の進捗状況や部分的な解決策が見落とされ、報酬シグナルがまばらになり、エージェントの能力の不完全な全体像が得られます。実験の再現、ソフトウェア エンジニアリング、マルチモーダル解析、インタラクティブ ゲーム、科学技術コンピューティングなど、9 つのカテゴリにわたる 46 の長期的タスクのターミナル ベンチマークである Long-Horizo​​n-terminal-bench を紹介します。各タスクは、リファレンス ソリューションまたはシミュレーション エンジンを使用したターミナル ベンチ スタイルのセットアップに従いますが、さらにきめの細かい段階的なサブタスクに分解されます。この設計により、高密度の中間報酬と部分的なクレジットが可能になり、エージェントが最終目標に到達したかどうかだけでなく、オープンエンドのワークフローでどこまで進んだかも評価で把握できるようになります。 Long-Horizo​​n-terminal-Bench のタスクは通常、数百のエピソードと数分から数時間の実行を必要とし、単発の問題解決ではなく、長期計画、長期コンテキスト管理、反復的なデバッグに重点を置きます。 15 のフロンティア モデルを評価したところ、エージェントはタスクあたり平均 990 万のトークンを消費し、実行あたり約 231 のエピソードと 85.3 分の実行時間を要し、Long-Horizo​​n-terminal-bench は以前の端末ベースのベンチマークよりも要求が厳しくなっていることがわかりました。最も強力なテスト済みモデルでも、部分報酬しきい値 0.95 で 15.2% の合格@1 を達成し、完全報酬しきい値 1.0 で 10.9% を達成します。一方、モデル全体の平均合格率は、2 つのしきい値の下でそれぞれ 4.3% と 1.7% です。これらの結果から、改善の余地があることがわかります。私たちは障害モードとエラー パターンをさらに分析し、Long-Horizo​​n-ターミナル エージェントの将来の進歩をサポートするために Long-Horizo​​n-terminal-Bench をリリースします。

原文 (English)

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.

13:00 JSTエージェントビジネス/資金調達研究/論文

LongMedBench: 長期的な臨床意思決定のための医薬品のベンチマーク

この研究では、長期的な臨床意思決定のための実際の EHR ベースのベンチマークである LongMedBench を紹介します。 LLM ベースの医療薬剤のこれまでの評価では、主にショートコンテキスト知識の QA とツールの使用が重視されてきました。しかし、実際の医療は本質的に長期的なものであり、臨床医は繰り返しの訪問、検査、進化する治療にわたる証拠を集約する必要があります。したがって、現実的な評価には長期的な相互作用が不可欠です。 LongMedBench は、MIMIC-IV 入院記録と臨床ノートを時系列イベント ストリームとロングコンテキスト メモリ データセットに統合する再現可能なパイプラインを介して構築されており、エージェントと臨床環境の間で長期にわたるマルチセッションの対話を可能にします。患者数は 335 名で、患者 1 人当たりの入院件数は平均 19.72 件、1 件当たりの医療イベント数は 44.91 件です。長期的な意思決定プロセスに基づいて、事実に基づく QA、時間的推論、長期的な意思決定という 3 つのスイートによる評価分類法を提案します。この分類法は、エージェントが長期にわたって過去の患者情報をどのように理解し、活用しているかを測定します。私たちの実験によると、最近の LLM は明示的なタイムスタンプをうまく利用できますが、暗黙的な時間推論には課題があることがわかりました。 RAG とエージェント メモリ システムは、情報検索タスクのパフォーマンスを向上させることができますが、意思決定タスクのパフォーマンスはモデルの直接のコンテキストに大きく依存します。

原文 (English)

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world medical care is inherently longitudinal, and clinicians must aggregate evidence across repeated visits, tests, and evolving treatments. Therefore, long-horizon interaction is essential for realistic assessment. LongMedBench is constructed via a reproducible pipeline that integrates MIMIC-IV admission records and clinical notes into time-series event streams and long-context memory datasets, enabling long-horizon, multi-session interactions between agents and a clinical environment. It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit. Guided by the long-horizon decision process, we propose an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making. This taxonomy measures how agents understand and leverage historical patient information over extended horizons. Our experiments show that while recent LLMs can make good use of explicit timestamps, they have challenges in implicit time inference; The RAG and agent memory system can improve the performance of information retrieval tasks, but the performance of decision-making tasks is highly dependent on the model's immediate context.

13:00 JST研究/論文

Research on Cross-media Science and Technology Information Data Retrieval

Since the era of big data, the Internet has been flooded with all kinds of information. Browsing information through the Internet has becom…

13:00 JST研究/論文

Research on Intellectual Property Resource Profile and Evolution Law

In the era of big data, intellectual property-oriented scientific and technological resources show the trend of large data scale, high info…

13:00 JST研究/論文

Profiling and Evolution of Intellectual Property

In recent years, with the rapid growth of Internet data, the number and types of scientific and technological resources are also rapidly ex…

13:00 JSTLLM/生成AI

HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA

Retrieval-augmented generation (RAG) has rapidly advanced the language model field, particularly in question-answering (QA) systems. By int…

13:00 JST研究/論文

Constrained Reinforcement Learning for Safe Heat Pump Control

Constrained Reinforcement Learning (RL) has emerged as a significant research area within RL, where integrating constraints with rewards is…

13:00 JSTエージェント

Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs

In the zero-shot policy transfer (ZSPT) setting for contextual Markov decision processes (CMDP), agents train on a fixed, finite set of con…

13:00 JST画像/動画生成研究/論文

On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets…

13:00 JST画像/動画生成

Asynchronous Perception Machine For Efficient Test-Time-Training

In this work, we propose Asynchronous Perception Machine (APM), a computationally-efficient architecture for test-time-training (TTT). APM…

13:00 JST画像/動画生成

Training-Free, Identity-Preserving Image Editing for Fashion Pose Alignment and Normalization

Diffusion models have recently unlocked new possibilities in editing images of real-world objects. Yet, transforming objects in non-rigid w…

13:00 JSTLLM/生成AIGPT / ChatGPT

Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers

Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Count…

13:00 JST研究/論文

Hyperflux: Pruning Reveals Importance

Network pruning is used to reduce inference latency and power consumption in large neural networks. However, most methods focus on empirica…

13:00 JST研究/論文

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-as…

13:00 JSTエージェントロボティクス

Adaptive Reinforcement Learning for Unobservable Random Delays

In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov de…

13:00 JSTハードウェア/半導体ビジネス/資金調達研究/論文

On the Necessity of Output Distribution Reweighting for Effective Class Unlearning

In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause inf…

13:00 JSTLLM/生成AIGemmaQwen

Can Argus Judge Them All? Comparing VLMs Across Domains

Vision-Language Models (VLMs) are increasingly used in industry VLM applications such as retrieval systems, content generation platforms, a…

13:00 JSTLLM/生成AI

Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI

Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs…

13:00 JSTLLM/生成AIエージェント研究/論文

SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, par…

13:00 JSTLLM/生成AI

CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search

Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retriev…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

Beyond Na\"ive Prompting: Strategies for Improved Context-aided Forecasting with LLMs

Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual f…

13:00 JSTLLM/生成AI

Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts

Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine…

13:00 JSTLLM/生成AI

PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains

Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning models (LRMs) by generating multiple candidate…

13:00 JST研究/論文Claude

TENET: One Step Toward Test-Driven Development for Repository-Level Code Generation

Test-Driven Development (TDD) is a widely adopted practice that requires developers to create and execute tests alongside implementation. W…

13:00 JSTLLM/生成AI

Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm

The pretrain-transfer paradigm, which underpins the success of large language models (LLMs), has demonstrated the immense power of creating…

13:00 JSTLLM/生成AIGPT / ChatGPT

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge

Large Language Models (LLMs) excel at multi-hop reasoning in distribution, yet fail on unseen compositions, a phenomenon known as the curse…

13:00 JSTエージェントロボティクス

Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning

In endovascular surgery, endovascular interventionists push a thin tube called a catheter, guided by a thin wire to a treatment site inside…

13:00 JST研究/論文

People use fast and flat simulation to reason about new games

Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on ex…

13:00 JSTLLM/生成AI

Improving Topic Modeling of Social Media Short Texts with Rephrasing: A Case Study of COVID-19 Related Tweets

Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID…

13:00 JSTLLM/生成AIエージェント

Enabling Agents to Communicate Entirely in Latent Space

While natural language is the de facto communication medium for LLM-based agents, it presents a fundamental constraint. The process of down…

13:00 JST研究/論文

Enhancing Adversarial Transferability through Block Stretch and Shrink

Input transformation-based attacks improve adversarial transferability by aggregating gradients over transformed inputs. Existing analyses…

13:00 JST画像/動画生成研究/論文

SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet e…

13:00 JST研究/論文

On the Condition Number Dependency in Bilevel Optimization

Bilevel optimization minimizes an objective function, defined by an upper-level problem whose feasible region is the solution of a lower-le…

13:00 JSTLLM/生成AI研究/論文NVIDIA

CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning

In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically opt…

13:00 JST研究/論文

The Theory of Strategic Evolution: Games with Endogenous Players and Strategic Replicators

Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. This paper provides th…

13:00 JST研究/論文

Graph-Based Bayesian Optimization for Quantum Circuit Architecture Search with Uncertainty Calibrated Surrogates

Quantum circuit design is a key bottleneck for practical quantum machine learning on complex, real-world data. We present an automated fram…

13:00 JSTLLM/生成AI画像/動画生成

Emotion Recognition in Signers

Recognition of signers' emotions suffers from one theoretical challenge and one practical challenge, namely, the overlap between grammatica…

13:00 JST画像/動画生成研究/論文

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture

This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the…

13:00 JST画像/動画生成

SwinIFS: Landmark Guided Swin Transformer For Identity Preserving Face Super Resolution

Face super-resolution aims to recover high-quality facial images from severely degraded low-resolution inputs, but remains challenging due…

13:00 JSTLLM/生成AI

BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts

Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where t…

13:00 JST研究/論文

Stable On-Policy Distillation through Adaptive Target Reformulation

Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models;…

13:00 JSTLLM/生成AI

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

While Mixture-of-Experts (MoE) scales capacity via conditional computation, Transformers lack a native primitive for knowledge lookup, forc…

13:00 JSTロボティクス研究/論文

PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour

Parkour tasks for quadrupeds have emerged as a promising benchmark for agile locomotion. While human athletes can effectively perceive envi…

13:00 JST研究/論文

Referential Regimes: Transformation-Invariant Identity for Neutral Substrates

Data systems increasingly operate under persistent legal, political, and analytic disagreement, where no single interpretive authority can…

13:00 JST研究/論文

SFO: Learning PDE Operators via Spectral Filtering

Partial differential equations (PDEs) govern complex systems, yet neural operators often struggle to efficiently capture the long-range, no…

13:00 JSTビジネス/資金調達

Rethinking Zero-Shot Time Series Classification: From Task-specific Classifiers to In-Context Inference

The zero-shot evaluation of time series foundation models (TSFMs) for classification typically uses a frozen encoder followed by a task-spe…

13:00 JSTLLM/生成AILlamaDeepSeek

Disentangling Intrinsic Importance from Emergent Structure in Multi-Expert Orchestration

Multi-expert systems, where multiple Large Language Models (LLMs) collaborate to solve complex tasks, are increasingly adopted for high-per…

13:00 JSTエージェント

Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM)

Generative social agents (GSAs) use artificial intelligence to autonomously communicate with human users in a natural and adaptive manner.…

13:00 JSTLLM/生成AI研究/論文

SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data

Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. Current LLM-based SAE bench…

13:00 JST画像/動画生成

Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision

Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context. Task-driven hard-attention mode…

13:00 JSTLLM/生成AI

BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking

Chain-of-Thought (CoT) reasoning has significantly improved LLMs' mathematical problem-solving capabilities, but distilling such capabiliti…

13:00 JST研究/論文Qwen

Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers

Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step inform…

13:00 JSTLLM/生成AIエージェント

A General Equilibrium Theory of Orchestrated AI Agent Systems

We establish a general equilibrium theory for systems of large language model (LLM) agents operating under centralized orchestration. The f…

13:00 JSTLLM/生成AI

MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models

Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each…

13:00 JST画像/動画生成ロボティクス研究/論文

RVN-Bench: A Benchmark for Reactive Visual Navigation

Safe visual navigation is critical for indoor mobile robots operating in cluttered environments. Existing benchmarks, however, often neglec…

13:00 JSTエージェントロボティクス

VehAnchor: Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery

Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telem…

13:00 JSTLLM/生成AI

Context-Dependent Affordance Computation in Vision-Language Models

We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Our primary study uses Qwen3-V…

13:00 JST研究/論文

Local Message-Passing for Discrete Graph Generation

Discrete graph generation has emerged as a powerful paradigm for modeling graph-structured data, yet state of the art models often rely on…

13:00 JSTLLM/生成AI

MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a compre…

13:00 JST研究/論文

Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning

The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain acc…

13:00 JSTビジネス/資金調達研究/論文

ECoLAD: Selecting Anomaly Detectors for Automotive Deployment via Compute-Reduction Evaluation

Automotive anomaly detectors are often selected from accuracy only benchmarks on workstation class hardware, whereas in-vehicle monitoring…

13:00 JST研究/論文

Evolutionarily Stable Stackelberg Equilibrium

We present a new solution concept called evolutionarily stable Stackelberg equilibrium (SESS). We study the Stackelberg evolutionary game s…

13:00 JST画像/動画生成

Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchic…

13:00 JST研究/論文

Critical Damping as a Momentum Schedule: Multi-Seed Validation, a Hybrid Recipe, and an Exhaustive Negative Result on Surgical Layer Selection

The critical damping condition of the damped harmonic oscillator model of SGD with momentum (Qian, 1999) yields a momentum schedule with no…

13:00 JSTLLM/生成AI

StanceMoE: Mixture-of-Experts Architecture for Stance Detection

Actor-level stance detection aims to determine an author expressed position toward specific geopolitical actors mentioned or implicated in…

13:00 JSTLLM/生成AI

How Annotation Trains Annotators: Competence Development in Social Influence Recognition

Human data annotation, especially when involving experts, is often treated as an objective reference. However, many annotation tasks are in…

13:00 JST研究/論文GPT / ChatGPT

From Paper to Program: Knowledge Externalization and Bottleneck Diagnosis in AI-Assisted Quantum Many-Body Programming

Large language models can write scientific code, but direct paper-to-program translation remains fragile when correctness depends on tacit…

13:00 JSTロボティクス

Pickalo: Leveraging 6D Pose Estimation for Low-Cost Industrial Bin Picking

Bin picking in real industrial environments remains challenging due to severe clutter, occlusions, and the high cost of traditional 3D sens…

13:00 JSTLLM/生成AI

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) mitigates hallucinations in Multimodal Large Language Models (MLLMs), yet existing systems struggle wi…

13:00 JSTLLM/生成AI

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various medi…

13:00 JSTLLM/生成AI

Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation

Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing…

13:00 JSTLLM/生成AI研究/論文

Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces

Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone…

13:00 JST画像/動画生成

SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation

Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and cli…

13:00 JSTLLM/生成AIエージェント

Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning

Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. Whe…

13:00 JSTLLM/生成AI画像/動画生成

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging

Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce c…

13:00 JST研究/論文

Graph Construction and Matching for Imperative Programs using Neural and Structural Methods

Reusing verification artefacts requires identifying structural and semantic similarities across programs and their specifications. In this…

13:00 JST画像/動画生成

Toward a Scientific Discovery Engine for Weather and Climate Data: A Visual Analytics Workbench for Embedding-Based Exploration

Earth system science is producing increasingly large, high-dimensional datasets from both physics-based and AI-driven models. While embeddi…

13:00 JSTLLM/生成AI画像/動画生成ロボティクス

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human de…

13:00 JSTLLM/生成AIビジネス/資金調達

Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA

Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive…

13:00 JST研究/論文GPT / ChatGPT

Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build

How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning o…

13:00 JSTLLM/生成AI研究/論文

A Multi-Model Metric-based Selection Framework for Abstractive Text summarization

Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents…

13:00 JST画像/動画生成ロボティクス

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target…

13:00 JST画像/動画生成

The Cross-Architecture Substrate: A Domain-Transcendent, Calibration-Surviving Geometric Invariant of Modern Vision Encoders

Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly diff…

13:00 JSTロボティクスGPT / ChatGPTGemini

Embodied-R1.5: 身体化された基盤モデルによる物理的知性の進化

私たちは、身体的認知、タスク計画、修正、ポインティングに及ぶ包括的な身体的推論機能を、一般的な身体的知能に向けた単一のアーキテクチャ内に統合する、統合された身体的基盤モデル (EFM) である Embodied-R1.5 を紹介します。 3 つの自動データ構築パイプラインを活用して重要な機能のデータ範囲を大幅に拡大し、150 億トークンを超える大規模なデータ システムを構築し、異種タスクの競合を軽減するマルチタスクのバランスのとれた RL レシピを設計します。さらに、単一のモデルが長期的なタスクにわたって自律的に実行および自己修正できるようにする Planner-Grounder-Corrector (PGC) 閉ループ フレームワークを導入します。 Embodied-R1.5 は、わずか 8B のパラメーターで、24 のエンボディド VLM ベンチマークのうち 16 で SOTA を達成し、Gemini-Robotics-ER-1.5 や GPT-5.4 などの主要モデルを上回っています。 Embodied-R1.5 は、内部化されたエンボディド機能の利点を活用して、少量のデータのみで VLA に微調整でき、4 つの一般的な操作ベンチマーク スイート全体で $\pi_{0.5}$ などの主要な VLA モデルを上回るパフォーマンスを発揮します。さらに、大規模なゼロショットの実際のロボット実験を実施し、命令追従、アフォーダンス グラウンディング、多関節オブジェクトの操作、および長期にわたる複雑なタスクのパフォーマンスを検証し、物理世界への強力な一般化を実証します。 EFM における将来の研究を促進するために、モデルの重み、データセット、トレーニング コード、および具体化されたタスクに合わせた評価フレームワークである EmbodiedEvalKit をオープンソースにしています。

原文 (English)

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $\pi_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We open-source model weights, datasets, training code, and EmbodiedEvalKit, an evaluation framework tailored for embodied tasks, to facilitate future research in EFMs.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

ISE: マルチターン OS エージェントの軌跡のための実行ベースのレシピ

有能な OS エージェントをトレーニングするには、構造化されたユーザーの意図、複数ターンのタスク委任、および根拠のあるツールの実行を同時にキャプチャするデータが必要ですが、これらのプロパティは既存のデータセットには存在しません。我々は、これらのギャップに共同で対処する 3 段階の合成パラダイムである ISE (Intent -> Simulate -> Execute) を提案します。ステージ 1 では、4D フレームワーク (ペルソナ x ドメイン x タスク x 複雑さ) を介して約 50,000 の構造化インテントを構築します。重複排除後、プールには 43956 個の一意のインテントが含まれ、mpnet-base-v2 埋め込み (コサイン カーネル、q=1) のプール全体で 61.57 の Vendi スコアを達成しました。ステージ 2 では、ロールロックされたユーザー シミュレータを介してマルチターンのユーザー エージェント インタラクションを推進し、各ユーザー ターンを実際の実行結果に基づいて実行し、平均 8.12 ユーザー ターンと合計 68.24 のダイアログ ターンに相当する 23132 の完全な軌跡を生成します。ステージ 3 では、ライブの分離された OS ワークスペース内ですべてのツール呼び出しが実行され、シミュレートされた応答ではなく、本物の障害回復ダイナミクスが生成されます。 ISETrace の微調整により、標準プロトコルのエージェント ツール使用タスクで Qwen3-8B を使用し、ClawEval pass@1 が 19.3 から 37.7 に改善されました。この結果は、ゼロショット GPT-4o や 4 倍大きい Qwen3-32B ベース モデルよりも優れています。ステージ 2 のアブレーションは、マルチターン シミュレーションがパフォーマンス向上の大部分をもたらすことを証明しています。すべてのソース コードとデータセットは https://github.com/Valiere01/ISE-Trace でリリースされます。

原文 (English)

ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories

Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded tool execution--properties absent from existing datasets. We propose ISE (Intent -> Simulate -> Execute), a three-stage synthesis paradigm that addresses these gaps jointly. Stage 1 constructs roughly 50000 structured intents via a 4D framework (Persona x Domain x Task x Complexity); after deduplication the pool contains 43956 unique intents and attains a Vendi Score of 61.57 over the entire pool on mpnet-base-v2 embeddings (cosine kernel, q=1). Stage 2 drives multi-turn user-agent interaction through a role-locked user simulator that grounds each user turn in actual execution outcomes, producing 23132 complete trajectories averaging 8.12 user turns and 68.24 total dialogue turns. Stage 3 runs every tool call inside a live, isolated OS workspace, generating authentic failure-recovery dynamics instead of simulated responses. Fine-tuning on ISETrace improves ClawEval pass@1 from 19.3 to 37.7 using Qwen3-8B on agent tool-use tasks with a standard protocol. This result outperforms zero-shot GPT-4o and the larger Qwen3-32B base model which is four times bigger. An ablation on Stage 2 proves multi-turn simulation brings a large portion of the performance gain. We release all source code and dataset at https://github.com/Valiere01/ISE-Trace.

13:00 JSTロボティクス研究/論文

Pipette: An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics

Wet-lab robots can improve the reproducibility, throughput, and safety of biomedical experiments, but scaling their learning requires custo…

13:00 JSTLLM/生成AI画像/動画生成

Gefen: 最適化された確率的オプティマイザー

AdamW は最新の深層学習のデフォルトのオプティマイザーですが、その第一モーメント状態と第二モーメント状態により、トレーニング メモリにおよそ 2 つのパラメーター サイズのバッファーが追加されます。私たちは、パラメータ ブロック間で 2 番目のモーメントの推定値を自動的に共有し、学習されたコードブックを使用して最初のモーメントを量子化する、メモリ効率の高いオプティマイザである Gefen を提案します。これにより、AdamW のメモリ フットプリントを同じパフォーマンスを維持しながら最大 8 倍削減できます。これは、10 億パラメータあたり 6.5 GiB の削減に相当します。この方法は、大規模な混合ヘッセ行列のエントリが二乗勾配の比率を 1 に向けて制約することを示す理論的結果によって動機付けられており、ヘッセ行列に整列されたパラメーターが二次モーメント統計を共有するための自然な候補であることを示唆しています。ヘッセ行列の計算は大規模には非現実的であるため、Gefen は初期の 2 乗勾配からブロック構造を推測し、AdamW のデフォルトを超えるアーキテクチャ固有のメタデータやハイパーパラメーターを必要としません。 Gefen は、正確なヒストグラムベースの動的プログラミング量子化コードブックを学習し、最初の瞬間のスケーリングに同じブロックを再利用します。さまざまな実験を通じて、Gefen は、AdamW レベルのパフォーマンスを維持しながら、比較した AdamW のような手法の中で最も低いピーク オプティマイザー メモリを実現しました。 FSDP および DDP トレーニングでは、メモリ フットプリントの削減により、より大きなマイクロバッチが可能になり、AdamW よりもスループットが大幅に向上します。これにより、メモリ使用量が減り、スループットが向上し、より大きなモデルのトレーニングやより大きなバッチ サイズの使用が可能になる実用的なドロップイン置換が提供されます。融合された CUDA カーネルを含む完全な Python 実装を https://github.com/ndvbd/Gefen で提供します。

原文 (English)

Gefen: Optimized Stochastic Optimizer

AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first moment using a learned codebook, thereby reducing AdamW's memory footprint by ~8x while maintaining the same performance, corresponding to a reduction of 6.5 GiB per billion parameters. The method is motivated by a theoretical result showing that large mixed Hessian entries constrain the ratio of squared gradients toward one, suggesting that Hessian-aligned parameters are natural candidates for sharing second-moment statistics. Since computing Hessians is impractical at scale, Gefen infers block structure from the initial squared gradients, requiring no architecture-specific metadata or hyperparameters beyond AdamW defaults. Gefen learns an exact histogram-based dynamic-programming quantization codebook and reuses the same blocks for first-moment scaling. Across diverse pretraining experiments, Gefen achieves the lowest peak optimizer memory among the compared AdamW-like methods while maintaining AdamW-level performance. In single-machine or distributed training, the reduced memory footprint enables larger microbatches and improves throughput significantly over AdamW, providing a practical drop-in replacement with lower memory usage that can increase throughput and enable training larger models or using larger global batch sizes. We provide the complete Python implementation, including fused CUDA kernels at https://github.com/ndvbd/Gefen

13:00 JST画像/動画生成エージェント

RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos

Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect and reproduce at scal…

13:00 JST研究/論文

RankGraph-2: 推奨事項における 10 億ノードのグラフ学習のためのライフサイクル協調設計

10億ノード規模のグラフベースの検索には、グラフ構築、表現学習、リアルタイム処理という3つの密結合した問題を共同で解決する必要があるが、既存の作業はそれぞれを個別に解決している。我々は、Meta に導入されたフレームワークである RankGraph-2 を紹介します。これは、類似性に基づく検索 (U2U2I および U2I2I) の 3 つのライフサイクル ステージすべてを共同設計し、各ステージの要件が他のステージの要件を形成します。サービスを提供するには、高価なオンライン KNN を回避するために共学習されたクラスター インデックスが必要です。これにより、インデックスの共トレーニングがトレーニング目標に組み込まれます。トレーニングでは、類似性に基づく検索が事前計算された近傍を許容し、オンライン グラフ インフラストラクチャが不要になるという観察から恩恵を受けます。これには、自己完結型データを生成するための構築が必要です。構築では、項目範囲の時間レベルの更新もサポートする必要があります。これらのカスケード要件に基づいて、RankGraph-2 は、人気バイアス補正を備えたサブサンプリングによって数百兆のエッジを数千億に削減し、パーソナライズされた PageRank によってマルチホップ近傍を事前計算し、サービスの計算コストを 83% 削減する残差量子化クラスター インデックスを共同学習します。このライフサイクルの共同設計により、シンプルなアーキテクチャで、二部グラフでは GAT + Deep Graph Infomax モデルよりも 3.8 倍高い再現率、項目検索では PyTorch-BigGraph よりも 2.1 倍高い再現率を達成できます。 RankGraph-2 は最大 +0.96% の CTR と +2.75% の CVR を実現し、主要なサーフェス全体で 20 回以上の検索起動を実現しました。

原文 (English)

RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation

Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation learning, and real-time serving -- yet existing work addresses each in isolation. We present RankGraph-2, a framework deployed at Meta that co-designs all three lifecycle stages for similarity-based retrieval (U2U2I and U2I2I), where each stage's requirements shape the others. Serving requires a co-learned cluster index to avoid expensive online KNN -- this pushes index co-training into the training objective. Training benefits from the observation that similarity-based retrieval tolerates pre-computed neighborhoods, eliminating online graph infrastructure -- this requires construction to produce self-contained data. Construction must also support hour-level refresh for item coverage. Acting on these cascading requirements, RankGraph-2 reduces hundreds of trillions of edges to hundreds of billions via subsampling with popularity bias correction, pre-computes multi-hop neighborhoods via personalized PageRank, and co-learns a residual-quantization cluster index that reduces serving computational cost by 83%. This lifecycle co-design enables a simple architecture to achieve 3.8 x higher recall than a GAT + Deep Graph Infomax model on a bipartite graph and 2.1 x higher than PyTorch-BigGraph on item retrieval. RankGraph-2 delivers up to +0.96% CTR and +2.75% CVR, and has powered 20+ retrieval launches across major surfaces.

13:00 JSTエージェント

FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving

Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficie…

13:00 JSTLLM/生成AILlama

Small edits, large models: How Wikipedia advocacy shapes LLM values

Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia ap…

13:00 JSTビジネス/資金調達

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…

13:00 JSTLLM/生成AI

What Does It Mean to Break a Distillation Defense?

Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a studen…

13:00 JSTロボティクス

Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning

Underwater vehicles operate from a fixed onboard energy budget that propulsion rapidly depletes, so a controller that completes its task wh…

13:00 JSTLLM/生成AILlama

役に立ちます: トレーニング中期の思いやりの値はトレーニング後にドメインに依存して低下します

標準的なポストトレーニング パイプラインは、教師あり微調整 (SFT) と強化学習 (RL) を適用して言語モデルを有用にしますが、これらのプロセスは、トレーニング前に注入された値を誤って低下させる可能性があります。動物危害ベンチマーク (AHB 2.2) と MORU ベンチマークで評価された SFT (Dolly-15k による有用性と Magicoder-110K によるコーディング) と GRPO (RLHFlow による有用性と Magicoder によるコーディング) の両方を使用して、思いやり指向の合成データで中間トレーニングされた Llama 3.1 8B モデルにおける動物の思いやりの値の保持に、トレーニング後のデータのドメインが差動的に影響を与えるかどうかを調査します。 (不確実性の下での道徳的推論)。有用性トレーニングは、AHB でのコーディング トレーニングと比較して動物の思いやりを大幅に低下させます (SFT: 35.7% 対 65.2%、GRPO: 18.7% 対 32.0%)。これは 2 つの独立した有用性データセットと 2 つのトレーニング パラダイムにわたって再現されています。英語のMORU項目では、有用性トレーニングは一般的な道徳的推論を25.5パーセントポイント(46.4%対71.9%)低下させ、その大きさは同情効果に匹敵する顕著な差でした。ただし、この効果は言語を越えて伝わりません。多言語の MORU ベンチマークでは、ドメイン効果は消失します (SFT: 52.3% 対 51.2%)。対照的に、動物の思いやりの効果は言語間で一貫して伝わり、Magiccoder の基本モデルに対する AHB パーセンテージ ポイントの増加は、英語以外の項目では英語の項目よりも 4.5 倍大きくなっています。この乖離は、トレーニング中に教え込まれた価値観が、ドメイン固有のトレーニング後の改善を推論するよりも深く、言語を超えてコード化されていることを示唆しています。これらの結果は、価値を満載したトレーニング途中で構築するラボの場合、トレーニング後の有用性よりも、トレーニング後のコーディング ドメインの方が、一般的な推論能力を損なうことなく、トレーニング途中の値をよりよく保存できる可能性があることを示唆しています。

原文 (English)

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, using both SFT (helpfulness via Dolly-15k vs. coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs. coding via Magicoder), evaluated on the ANIMA 2.2 benchmark and MORU benchmark (Moral Reasoning Under Uncertainty). Helpfulness training significantly degrades animal compassion relative to coding training on ANIMA (SFT: 35.7% vs. 65.2%; GRPO: 18.7% vs. 32.0%), replicating across two independent helpfulness datasets and two training paradigms. On English MORU items, helpfulness training degrades general moral reasoning by 25.5 percentage points (46.4% vs. 71.9%), a striking gap that rivals the compassion effect in magnitude. However, this effect does not transfer cross-lingually: on the multilingual MORU benchmark, the domain effect disappears (SFT: 52.3% vs. 51.2%). In contrast, the animal compassion effect transfers consistently across languages, with Magicoder's ANIMA percentage-point gain over the base model 4.5 times larger on non-English items than English items. This divergence suggests that values instilled through mid-training are encoded more deeply and cross-lingually than reasoning improvements from domain-specific post-training. These results suggest that, for labs building on value-laden mid-training, coding-domain post-training may better preserve mid-trained values than helpfulness post-training without harming general reasoning capabilities.

13:00 JSTLLM/生成AILlama

主張し、説明しないでください: 動物福祉に関する LLM の推論を変える言語的特徴

動物愛護活動家たちは多くの著作物を作成しており、その著作物が言語モデルを訓練し、その後何百万人もの人々が動物福祉について尋ねるようになっています。提示された動物福祉ベンチマークで語彙を一致させたスタンスコントラストプローブを使用して、10の言語的特徴のそれぞれが、微調整データとして使用された場合にラマ-3.2-1Bの動物福祉推進推論に対する好みをどのように変化させるかを測定します。 10 個の特徴のうち 8 個で、統計的に有意な変化が生じます。 7 つは、断定的な確実性、明確な道徳的語彙、感情的な言葉、評価的主張、物語の構造、描写された危害の深刻度、即時的な時間的枠組みなど、モデルをより強力な動物愛護推進の推論に向けて移行させています。 2 つはそれを逆方向に動かします。ヘッジされた言葉と具体的な感覚的説明は両方とも動物愛護推進の立場を薄めます。一人称視点には統計的に有意な効果はありません。 LLM トレーニング コーパスに組み込まれる可能性のある動物福祉に関するテキストを執筆する人に対する実際的な推奨事項: シーンを中立的に説明するのではなく、立場を主張することです。モデルを変える特徴は、作家の立場を明確にするものです。それを弱める特徴は動物愛護の内容を保持しますが、スタンスを保留します。

原文 (English)

Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B's preference for pro-animal-welfare reasoning when used as fine-tuning data. Eight of the ten features produce statistically significant shifts. Seven move the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two move it the other way: hedged language and concrete sensory description both dilute the pro-animal-welfare stance. First-person perspective has no statistically significant effect. The practical recommendation for anyone writing animal-welfare text that may end up in LLM training corpora: assert a position rather than describe a scene neutrally. The features that shift the model are the ones that make the writer's position explicit; the features that dilute it hold animal-welfare content but withhold stance.

13:00 JST画像/動画生成Sora

NaviCache: ビデオ生成のためのテスト時の自己調整キャッシング

ビデオ拡散モデル (VDM) は、膨大な計算コストによる制約を受けます。オフライン キャリブレーション ベースの高速化には、キャリブレーション データの依存性、法外なキャリブレーション期間、分布シフトの影響を受けやすいという問題がありますが、オフライン キャリブレーションを必要としない方法では、これらの障害が解消されます。ただし、入力と出力の差の間のマッピングがリアルタイムで変化する瞬間的なゼロ次近似に依存しているため、観測ノイズの影響を受けやすく、拡散軌跡内の固有運動量が無視されます。この論文では、機能進化を慣性航法システム (INS) 問題として再概念化する、プラグ アンド プレイのテスト時自己校正手法である NaviCache を提案します。 NaviCache は、入力と出力の変動の間の相対的な結合をモデル化することで、基本的なドメイン ギャップと拡散の非定常的な性質を橋渡しします。特徴変化率とその潜在ドリフトを適応的に追跡するデュアルステート推定アーキテクチャを導入し、特殊な初期調整フェーズを通じて初期化します。 NaviCache は、時間依存のノイズ スケジュールを不確実性を考慮した測定更新メカニズムと統合することにより、誤差制限のある計算スキップのための理論に基づいたメカニズムを提供します。 HunyuanVideo、Wan、Open-Sora シリーズでの広範な実験により、NaviCache が計算スキップに対してより正確なエラー判定を示し、優れた総合パフォーマンスを達成することが実証されました。

原文 (English)

NaviCache: Test-Time Self-Calibration Caching for Video Generation

Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from calibration data dependency, prohibitive calibration duration, and susceptibility to distribution shifts, offline calibration-free methods eliminate these hurdles. However, since they rely on instantaneous zero-order approximations where the mapping between input and output differences varies in real-time, they are susceptible to observational noise and ignore the intrinsic momentum within the diffusion trajectory. In this paper, we propose NaviCache, a plug-and-play test-time self-calibration method re-conceptualizing feature evolution as an Inertial Navigation System (INS) problem. NaviCache bridges the fundamental domain gap and the non-stationary nature of diffusion by modeling the relative coupling between input and output variations. We introduce a dual-state estimation architecture that adaptively tracks the feature change ratio and its latent drift, initialized via a specialized Initial Alignment phase. By integrating a time-dependent noise schedule with an uncertainty-aware Measurement Update mechanism, NaviCache provides a theoretically grounded mechanism for error-bounded computation skipping. Extensive experiments on the HunyuanVideo, Wan, and Open-Sora series demonstrate that NaviCache exhibits more accurate error judgment for computation skipping and achieves outstanding comprehensive performance.

13:00 JSTLLM/生成AI

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test constru…

13:00 JST研究/論文

CMSL: Constructive Multi-Sequence Learning for Recommendation Systems

Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Mode…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agen…

13:00 JSTエージェントロボティクス

Freeform Preference Learning for Robotic Manipulation

Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where spa…

13:00 JSTロボティクス

FLYNN: Fly Brain トポロジーを使用したロボット ナビゲーションのための堅牢なニューラル ネットワーク

深層学習モデルは複雑なタスクで最先端のパフォーマンスを実現しますが、新しい環境や感覚遮断に直面すると脆弱なままです。対照的に、生体系はこれらの課題に対して顕著な耐性を示します。私たちは、ショウジョウバエのシナプス分解能の脳コネクトームから直接派生したアーキテクチャをもつリカレント ニューラル ネットワーク (RNN) を開発することで、この脆弱性に対処します。我々は、MuJoCo でビジョンベースのナビゲーションを実行するためにフライ コネクトーム ニューラル ネットワーク (FLYNN) をトレーニングし、同様のパラメーター数の最新の手作りネットワークに匹敵するパフォーマンスを達成する実現可能性を実証します。重要なことは、FLYNN は、さらなるトレーニングを行わなくても、分布外 (OOD) データに対する優れた耐性と感覚喪失に対する耐性を示します。完全な視力喪失下でも機能を維持しましたが、手作りのネットワークは、カメラのドロップアウトで特別に訓練された場合でも、ほとんど機能しませんでした。 FLYNN の内部状態の主成分分析 (PCA) は、FLYNN が特に高度な表現モジュール性を示していることを示唆しており、これがその堅牢性に関連している可能性があります。私たちの研究は、生物学的な脳のトポロジーに従って弾力性のある人工エージェントを設計するための新しい方向性を提供します。

原文 (English)

FLYNN: Robust Neural Network for Robot Navigation using Fly Brain Topology

While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or sensory deprivation. In contrast, biological systems exhibit remarkable tolerance to these challenges. We address this vulnerability by developing a recurrent neural network (RNN) whose architecture is directly derived from the synaptic-resolution brain connectome of the fruit fly Drosophila melanogaster. We demonstrate the feasibility of training the fly connectome neural network (FLYNN) to perform vision-based navigation in MuJoCo, achieving performance comparable to modern hand-crafted networks of similar parameter counts. Crucially, FLYNN exhibits superior resistance to out-of-distribution (OOD) data and tolerance to sensory loss without further training. It remained functional even under total vision loss while hand-crafted networks largely failed, even when specifically trained with camera dropout. Principal Component Analysis (PCA) of the internal state of FLYNN suggests that it exhibits a particularly high degree of representational modularity, which might be related to its robustness. Our work provides a new direction for designing resilient artificial agents following the topology of biological brains.

13:00 JST研究/論文

Diffusion-GR2: 拡散生成推論リランカー

生成推論の再ランカーは、候補リストを並べ替える前に思考の連鎖を発行することで強力な推奨精度を実現しますが、推論が遅くなります。自己回帰 (AR) デコーダは推論トークンごとに 1 回の連続した前方パスを費やし、推論トレースは生成されるランキングをはるかに上回ります。このコストを削減するために、ブロック拡散言語モデルは、いくつかのノイズ除去ステップで多くの位置を並行してデコードし、大幅に高速になりますが、単純に AR リランカーを 1 つに変換すると、2 つの精度ギャップが生じます。 (1) 構造的なギャップ: 回答位置は並行してノイズ除去され、独立してスコアリングされるため、デコーダーは無効なランキング (重複、欠落、またはセット外の識別子) を生成しますが、AR はこれを左から右のマスキングによって回避します。 (2) 分布ギャップ: 固定教師軌道上で変換されたモデルを微調整することは、推論時の独自のデコードと比較してポリシーから外れており、精度ギャップが残ります。高速化を維持しながら両方のギャップを埋めるために、AR 推論リランカー (GR2) をブロック拡散リランカーに変換するレシピである \textbf{Diffusion-GR2} を提案します。まず、変換微調整 (CFT) は、AR で初期化された拡散モデルを適応させて、外部の制約付きデコーダーを使用せずに、独自に答えを有効な置換にノイズ除去します。次に、オンポリシー蒸留 (OPD) が、AR 教師からの高密度のトークンごとのターゲットを使用して、独自のデコードされた軌道でモデルを監視します。最後に、OPD のポリシーに関するポリシーに加えて、再ランキング報酬に対して強化学習 (RL) ステージを適用します。 Amazon Beauty での実験では、Diffusion-GR2 が AR リランカーとほぼ同等に回復し、ブロック並列デコードにより、モデルの推論出力長でデコード スループットが $2.4$ ~ $3.5\times$ 向上することが実証されました。アブレーションにより、CFT がコンバージョン ギャップのほとんどを回復し、ポリシーに基づいた蒸留により AR リファレンスにさらに近づくことが示されています。

原文 (English)

Diffusion-GR2: Diffusion Generative Reasoning Re-ranker

Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces. To reduce this cost, block-diffusion language models decode many positions in parallel over a few denoising steps and are substantially faster, yet naively converting an AR re-ranker into one opens two accuracy gaps: (1) a structural gap: answer positions are denoised in parallel and scored independently, so the decoder emits invalid rankings (duplicated, dropped, or out-of-set identifiers) that AR avoids through left-to-right masking; and (2) a distributional gap: fine-tuning the converted model on fixed teacher trajectories is off-policy relative to its own decoding at inference, leaving a residual accuracy gap. To close both gaps while keeping the speedup, we propose \textbf{Diffusion-GR2}, a recipe that converts our AR reasoning re-ranker (GR2) into a block-diffusion re-ranker. First, conversion fine-tuning (CFT) adapts the AR-initialized diffusion model to denoise the answer into a valid permutation on its own, without an external constrained decoder. Next, on-policy distillation (OPD) then supervises the model on its own decoded trajectories with dense per-token targets from the AR teacher. Finally, we apply a reinforcement-learning (RL) stage against a re-ranking reward on top of OPD's on-policy policy. Experiments on Amazon Beauty demonstrate that Diffusion-GR2 recovers to near-parity with the AR re-ranker, while block-parallel decoding raises decode throughput by $2.4$--$3.5\times$ at the model's reasoning output length. Ablations show that CFT recovers most of the conversion gap, and that on-policy distillation further closes it to the AR reference.

13:00 JSTLLM/生成AI

DemoPSD: Disagreement-Modulated Policy Self-Distillation

On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single mo…

13:00 JSTLLM/生成AI画像/動画生成

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-…

13:00 JST研究/論文

x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability

Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function eva…

13:00 JST画像/動画生成

AnchorPrune: ビジュアル トークン プルーニングのための関連性に基づいたコンテキスト拡張

高解像度の入力では数千のビジュアル トークンが導入され、その多くは特定のクエリに対して冗長であるため、大規模なビジョン言語モデルにはかなりの推論コストがかかります。既存の枝刈り手法では、クエリの関連性とトークンの多様性を組み合わせることがよくありますが、これらの目的は、積極的な圧縮の下では矛盾する可能性があります。関連性主導の選択では、相関する局所的な証拠に予算が集中しすぎる可能性がありますが、多様性主導の選択では、不可欠なトークンが抑制されたり、明確ではあるが情報のない領域が保持されたりする可能性があります。最初に保護された関連性アンカーを構築し、次にそれを補完的な視覚的コンテキストで拡張する、トレーニング不要のフレームワークである AnchorPrune を紹介します。 AnchorPrune は、関連性でランク付けされたトークンのノベルティ プロファイルからアンカー サイズを適応的に決定し、クエリクリティカルな証拠のコンパクトなセットを保存し、重要度に重み付けされたノベルティを通じて残りの予算を割り当て、アンカーに関連する有益で冗長でないコンテキストを回復します。この順序付けされたデザインにより、コンテキストの拡張によって不可欠なクエリ キューが置き換えられるのを防ぎ、全体的な視覚的範囲が向上します。 AnchorPrune は軽量でアーキテクチャを認識しており、再トレーニングもモデルの変更も必要ありません。画像およびビデオの視覚言語モデルとベンチマーク全体で、特に厳しい圧縮下で、トレーニング不要のベースラインと比較して精度と効率のトレードオフを一貫して改善します。 LLaVA-NeXT-7B では、AnchorPrune は 2,880 個のビジュアル トークンのうち 160 個のみを使用して、フルトークンのパフォーマンスの 97.6% を維持します。これらの結果は、効率的なマルチモーダル推論のための効果的な原理として、関連性にアンカーされた文脈拡張を確立します。コードは https://github.com/MULTI-cau/AnchorPrune で入手できます。

原文 (English)

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and token diversity, yet these objectives can conflict under aggressive compression: relevance-driven selection may overconcentrate the budget on correlated local evidence, while diversity-driven selection may suppress indispensable tokens or retain distinct but uninformative regions. We introduce AnchorPrune, a training-free framework that first constructs a protected relevance anchor and then expands it with complementary visual context. AnchorPrune adaptively determines the anchor size from the novelty profile of relevance-ranked tokens, preserving a compact set of query-critical evidence, and allocates the remaining budget through importance-weighted novelty to recover informative, non-redundant context relative to the anchor. This ordered design prevents contextual expansion from displacing indispensable query cues while improving overall visual coverage. AnchorPrune is lightweight, architecture-aware, and requires neither retraining nor model modification. Across image and video vision-language models and benchmarks, it consistently improves the accuracy-efficiency trade-off over training-free baselines, particularly under severe compression. On LLaVA-NeXT-7B, AnchorPrune preserves 97.6% of full-token performance using only 160 of 2,880 visual tokens. These results establish relevance-anchored contextual expansion as an effective principle for efficient multimodal inference. Code is available at https://github.com/MULTI-cau/AnchorPrune.

13:00 JST研究/論文

LieBN: リー群に対するバッチ正規化

多様体値の測定は、さまざまな機械学習タスクで普及しています。最近の進歩により、ディープ ニューラル ネットワーク (DNN) が多様体上で動作するように拡張され、さまざまな形状に合わせて調整された正規化技術 (総称してリーマン正規化と呼ばれます) が併用されています。ただし、既存のリーマン正規化法のほとんどは、特定の多様体向けに設計されているか、多様体値の標本分布を効果的に正規化できません。これらの制限に対処するために、リー群に対するリーマン バッチ正規化 (RBN) のフレームワークである LieBN を提案します。私たちのアプローチは、すべてのリー群に自然に存在する理論的に便利な左右不変計量を活用し、リーマン平均と分散を制御するための理論的保証を提供します。 9 つの異なるジオメトリにわたって LieBN をインスタンス化します。そのうちの 4 つは対称正定 (SPD) 多様体上に、1 つは回転行列のグループ上に、4 つはフルランク相関行列の多様体上にあります。特に、SPD 計量の中で、新しい右不変計量を導入し、行列累乗変形を介して 3 つの既存のリー群構造を拡張します。さまざまな多様体に対する広範な実験により、フレームワークの有効性が検証されています。コードは https://github.com/GitZH-Chen/LieBN.git で入手できます。

原文 (English)

LieBN: Batch Normalization over Lie Groups

Manifold-valued measurements are prevalent in various machine learning tasks. Recent advances have extended Deep Neural Networks (DNNs) to operate on manifolds, accompanied by normalization techniques tailored to different geometries, collectively referred to as Riemannian normalization. However, most existing Riemannian normalization methods are either designed for specific manifolds or fail to effectively normalize manifold-valued sample distributions. To address these limitations, we propose LieBN, a framework for Riemannian Batch Normalization (RBN) over Lie groups. Our approach leverages the theoretically convenient left- and right-invariant metrics, which naturally exist in every Lie group, and provides theoretical guarantees for controlling the Riemannian mean and variance. We instantiate LieBN across nine distinct geometries: four on the Symmetric Positive Definite (SPD) manifold, one on the group of rotation matrices, and four on the manifold of full-rank correlation matrices. Notably, among the SPD metrics, we introduce a novel right-invariant metric and extend three existing Lie group structures via matrix power deformation. Extensive experiments on different manifolds validate the effectiveness of our framework. The code is available at https://github.com/GitZH-Chen/LieBN.git.

13:00 JST研究/論文

EHR-MPC: 生成患者デジタル ツインを使用した敗血症治療のための推論時間制御

敗血症は死亡の主な原因ですが、最適な治療方針については依然として議論が続いています。既存の強化学習 (RL) アプローチは、敗血症治療のための固定戦略を学習するため、推論中に変化する臨床目的への適応性が制限されます。私たちは、生成電子医療記録 (EHR) モデルの形式で患者のデジタル ツインをトレーニングすることで、患者のダイナミクスの学習と治療の最適化を切り離すフレームワークである EHRMPC を提案します。デジタル ツインは介入中の臨床経過を予測し、モデル予測制御 (MPC) を可能にして、シミュレーションによる推論時間計画を通じて治療を最適化します。我々は、ポリシー外の重要度サンプリングとポリシー上のシミュレーションベースの評価の両方を使用して、マサチューセッツジェネラルブリガム医療システムの8つの病院にわたる多施設ICU敗血症コホートでEHR-MPCを評価します。 RL ベースラインと比較して、EHR-MPC は同等のオフポリシー パフォーマンスと改善されたシミュレーション パフォーマンスを実現します。 RL とは異なり、この作業は敗血症治療の最適化を学習された患者の動態に対する推論時間の制御として枠組み化し、生成臨床モデルを使用した意思決定のための一般的な枠組みを確立します。

原文 (English)

EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins

Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during inference. We propose EHRMPC, a framework that decouples learning patient dynamics from optimizing treatment by training a patient digital twin in the form of a generative electronic health record (EHR) model. The digital twin predicts clinical trajectories under interventions and enables model predictive control (MPC) to optimize treatments via inference-time planning over simulations. We evaluate EHR-MPC on a multicenter ICU sepsis cohort spanning 8 hospitals in the Mass General Brigham health system using both off-policy importance sampling and on-policy simulation-based evaluation. Relative to RL baselines, EHR-MPC achieves comparable off-policy performance and improved simulation performance. Unlike RL, this work frames sepsis treatment optimization as inference-time control over learned patient dynamics, establishing a general framework for decision making with generative clinical models.

13:00 JSTロボティクス

腕全体の操作のための触覚および視覚条件付き接触中心制御

アーム全体の操作には環境との直接接触が含まれ、ロボットは接触の形成、スライド、切断に応じて複数のリンクに接触を分散することでタスクを完了します。この設定は、多くの学習ベースの操作パイプラインにおける一般的な暗黙の前提を打ち破ります。つまり、アーム構成は動きと接触の力を密接に結び付け、接触状態はオクルージョン下で部分的に観察されます。また、純粋に学習されたロールアウトは、多くのマルチリンク接触構成がデータ内でまばらに表現されるため、分布シフトの下では物理的に不一致になる可能性があります。これに対処するために、腕全体を操作するための後退水平コントローラーである TACTIC (Tactile and Vision Conditioned Contact-Centric Control) を提案します。 TACTIC は、RGB-D、分散型触覚センシング、コンパクトな 2D 近接表現を組み合わせた接触中心のハイブリッド予測モデルを使用します。このモデルは、学習されアクション条件付けされた潜在力学モデルと接触ヤコビアンを介した解析運動学を結合し、将来の接触構成と相互作用力のロールアウトを可能にします。 TACTIC は、これらのロールアウトを、接触を意識したアクション サンプリングを備えたサンプリング ベースの MPC プランナーに統合します。接触ヤコビアン ベースの投影は、サンプリングされたアクション シーケンスを力を調整する方向に導き、予測された近接力と相互作用力に対して定義された目標は、タスクの進行状況と腕全体の力の調整をトレードします。当社は、最先端のモデルベースおよびモデルフリーの手法に対してシミュレーションで TACTIC を評価し、各設計選択の寄与を分離するアブレーションを実行します。 TACTIC は他の手法よりも常に優れたパフォーマンスを発揮します。さらに、複数の接触軌道を必要とする 3 つの腕全体の操作タスク (マネキンの裏返しと位置変更、および 3D ダイナミック迷路でのゴール到達) にわたる分散触覚センシングを備えたロボットの現実世界のパフォーマンスを実証します。ウェブサイト: https://emprise.cs.cornell.edu/tactic

原文 (English)

TACTIC: Tactile and Vision Conditioned Contact-Centric Control for Whole-Arm Manipulation

Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break. This setting breaks common implicit assumptions in many learning-based manipulation pipelines: arm configuration tightly couples motion and contact forces, contact state is partially observed under occlusion, and purely learned rollouts can become physically inconsistent under distribution shift because many multi-link contact configurations are sparsely represented in the data. To address this, we propose TACTIC (Tactile and Vision Conditioned Contact-Centric Control), a receding-horizon controller for whole-arm manipulation. TACTIC uses a contact-centric hybrid predictive model that combines RGB-D, distributed tactile sensing, and a compact 2D proximity representation. The model couples a learned, action-conditioned latent dynamics model with analytical kinematics through contact Jacobians, enabling rollouts of future contact configurations and interaction forces. TACTIC integrates these rollouts into a sampling-based MPC planner with contact-aware action sampling: contact Jacobian-based projections steer sampled action sequences toward force-modulating directions, and objectives defined over predicted proximity and interaction forces trade task progress against whole-arm force regulation. We evaluate TACTIC in simulation against state-of-the-art model-based and model-free methods, and perform ablations that isolate the contribution of each design choice. TACTIC consistently outperforms other methods. We further demonstrate real-world performance on a robot with distributed tactile sensing across three whole-arm manipulation tasks that require multi-contact trajectories: turning over and repositioning a manikin, and goal-reaching in a 3D dynamic maze. Website: https://emprise.cs.cornell.edu/tactic

13:00 JSTLLM/生成AI

ドイツ語と英語のための主権のあるオープンソース基盤モデル

私たちは、ドイツ語と英語向けのソブリンのオープンソース Mixture-of-Experts (MoE) ハイブリッド Mamba Transformer 基礎モデルである Soofi S 30B-A3B を紹介します。そのハイブリッド設計は、トークンごとに 30B パラメーターのうち 3B のみをアクティブにし、コンテキストが増加しても推論キャッシュをほぼ一定に保つため、長いコンテキスト、高同時実行の展開において、高密度モデルよりも決定的なスループットの利点をもたらします。意図的に重み付けされたドイツ語を使用して約 27 兆のトークンで事前トレーニングされた Soofi S は、英語とドイツ語の集約ベンチマークで高密度の 14 ~ 27B モデルに匹敵し、17 のオープンベースモデルの中で両方の言語で最高のコード集約を達成し、アクティブパラメータがはるかに大きいものも含め、比較においてすべてのヨーロッパのソブリンベースラインを上回っています。フルオープンモデルの中で、Soofi S は Olmo 3 32B や Apertus 70B を抑えて、英語とドイツ語で最高の評価スコアを獲得しています。 Soofi S は、ミュンヘンのドイツテレコムが運用する主権 HPC スケールの AI インフラストラクチャである German Industrial AI Cloud 上にエンドツーエンドで構築されました。 Soofi S は、重み付け、選択された中間チェックポイント、完全なソースごとのデータ アカウンティング、ハイパーパラメータ、トレーニングおよび評価コードなど、非常に寛容なオープンアクセス条件に基づいてリリースされます。ソースライセンスが許可する場合、データ構築アーティファクトは寛容なライセンスの下でリリースされます。商業的にライセンスされた情報源は、集計統計と正確な混合物の計算とともに文書化されています。

原文 (English)

A Sovereign, Open-Source Foundation Model for German and English

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.

13:00 JSTLLM/生成AI

Conceptual Networks for Cross-Linguistic Idiomatic Expressions: A Feature-Based Graph Approach

We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse la…

13:00 JST画像/動画生成エージェント

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millime…