Skip to the content.

AIニュース 2026-08-27

自動生成: 2026-08-27 17:34 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Intelligent transcription with Gemini 3.5 TranscribeGoogle DeepMind

    Now you can get more intelligent speech-to-text transcription with Ge…

  2. Learning never stops: How AI makes learning continuousOpenAI

    OpenAI’s new report explores how students and educators use ChatGPT t…

  3. Bringing ChatGPT for Teachers to more U.S. school districtsOpenAI

    ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing…

  4. 中国製AIチップのクラスタで「Opus 4.8並み性能」モデルを提供 話題のステルスモデルの正体は「GLM-5.3-Flash」ITmedia AI+

    中国Z.aiが「GLM-5.3-Flash」をMITライセンスで公開。「Ox Alpha」として匿名検証していたマルチモーダルモデルで、「…

  5. MCPの新たなロードマップ公開 今後はAIエージェント対応、HTTP通信への統一などに注力ITmedia AI+

    Linux Foundation傘下でMCP(Model Context Protocol)の仕様策定を行っている「Agentic AI…

  6. OpenAI releases its official report on the Hugging Face breachTechCrunch AI

    The report, which spans several discrete cybersecurity compromises, i…

  7. 個人で契約したChatGPTに、会社の情報を入力していいの? 超初心者向けAI解説ITmedia AI+

    もはや生活に欠かせない存在となりつつある生成AI。毎日の仕事においても例外でなく、便利に使っている人が多いでしょう。しかし、機密情報をむや…

トピック別件数

日本語メディア9件

ITmedia AI+ (日本語)

15:00 JSTハードウェア/半導体NVIDIA

NVIDIAが開発支援、約30秒で口腔がんの疑い検出する歯科用AI

大阪大学らが開発したプログラム医療機器「口腔粘膜画像解析システム Mucosight AI」が、製造販売承認を取得した。AIを活用して口腔粘膜疾患の特徴を検出し、歯科医師の受診勧奨判断を支援する。

14:59 JSTロボティクス

ちっちゃな人型ロボが100mトラックをぽてぽて……中国ロボ陸上の“おちびランナー”が話題 「可愛すぎ」「けなげ」

中国・北京で8月22日から26日(現地時間)まで開催された、人型ロボットの性能を競う「世界人型ロボット運動会」。陸上のウサイン・ボルト選手越えのスピードを発揮するロボなどが注目を集める一方で、100mトラックをぽてぽてと歩くとある小型人型ロボが可愛らしいとSNSの話題をさらって…

13:15 JSTLLM/生成AIClaude

中国製AIチップのクラスタで「Opus 4.8並み性能」モデルを提供 話題のステルスモデルの正体は「GLM-5.3-Flash」

中国Z.aiが「GLM-5.3-Flash」をMITライセンスで公開。「Ox Alpha」として匿名検証していたマルチモーダルモデルで、「Claude Opus 4.8」に迫る性能を安価で提供するとうたう。

13:00 JSTLLM/生成AI研究/論文Claude

2035年、AI推論需要の半分は「コード生成」に 推論市場の今と未来

ABI Researchは、急拡大するAI推論市場の現状を解説した。生成AIの本番利用が広がる中、推論ワークロードは2033年にトレーニングを追い越し、消費電力は2035年に46ギガワットへ拡大すると予測する。

11:00 JSTロボティクス

ヒューマノイド開発、「標準インタフェース」が不可欠に

ヒューマノイドロボットの開発において、通信インタフェースの重要性が増している。MIPI AllianceはフィジカルAI分野の専門グループを立ち上げ、インタフェースの標準化を急ぐ。

10:25 JSTエージェント

MCPの新たなロードマップ公開 今後はAIエージェント対応、HTTP通信への統一などに注力

Linux Foundation傘下でMCP(Model Context Protocol)の仕様策定を行っている「Agentic AI Foundation(AAIF)」は、今後のMCPの進化について、その方向性を示す新たなロードマップを発表しました。

07:00 JSTLLM/生成AIGPT / ChatGPT

個人で契約したChatGPTに、会社の情報を入力していいの? 超初心者向けAI解説

もはや生活に欠かせない存在となりつつある生成AI。毎日の仕事においても例外でなく、便利に使っている人が多いでしょう。しかし、機密情報をむやみにAIに共有してしまうと、思わぬリスクにつながるかもしれません。

18:10 JSTロボティクス

非公開の「AI学習データ工場」に潜入 人型ロボ30台超が稼働、意外と知らない「データ収集」の裏側とは

フィジカルAI用の学習データを収集する施設「J-HRTI 関東データファクトリー」が稼働した。内部には、約30体のロボットが並び、作業を繰り返している。同施設の全容と狙いを、潜入レポートとして紹介する。

17:55 JSTハードウェア/半導体

韓国SK Hynix、宮城県に半導体工場の計画浮上 「何か知り得ているものはない」と村井知事

韓国半導体大手SKハイニックスが宮城県にメモリー工場を建設する計画が浮上したことについて、宮城県の村井嘉浩知事は26日の定例記者会見で「いろんな半導体関連企業にアプローチしたのは事実だが、今回の件で県として何か知り得ているものはない」と述べ、今後も動向を注視する考えを示した。

海外メディア15件

TechCrunch AI (英語)

09:24 JSTビジネス/資金調達

Viral AI startup Instinct has raised $350M at a $2.5B valuation

The startup is only a year old but it has already generated a massive amount of hype (and money) while also spurring privacy concerns.

08:47 JSTハードウェア/半導体NVIDIA

Amazon just tripled its order of Nvidia chips over ‘surging demand’

Amazon is adding another 2 million Nvidia GPU chips to its data centers over the next two years. But this extended partnerships stretches b…

06:37 JSTLLM/生成AIAnthropic

Anthropic continues compute-gobbling streak in $45B deal with Nscale

The new deal with the infrastructure provider is the latest example of Anthropic's white-hot compute-gobbling streak.

04:37 JSTLLM/生成AIGoogleGemini

Google’s Gemini has a branding problem, and so does the rest of AI

Consumer AI apps need to stop making users learn their product architecture.

04:34 JSTLLM/生成AIOpenAI

How do we explain OpenAI’s executive exodus?

Was Greg Brockman the right executive all along?

04:05 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達OpenAI2件の関連記事

OpenAI releases its official report on the Hugging Face breach

The report, which spans several discrete cybersecurity compromises, is the most complete accounting of the incident to date.

出典:TechCrunch AITechCrunch AI
00:47 JSTエージェント

Radar makes podcasts searchable — and usable by AI agents

Particle’s new podcast intelligence platform transcribes and analyzes more than 130,000 podcasts, making their conversations searchable on…

00:00 JSTその他

Ex-Meta scientists want to bring visual AI to the factory floor

Perceptron offers an AI model that it says can help machines navigate the world while also providing in-depth visual intelligence.

23:37 JSTロボティクス

Bill Gates wants to see a robot tax and ‘Human Reserved’ jobs to mitigate harms from AI

Gates is mostly in the Responsible AI camp, but there are a few ideas in here we hadn't heard before.

23:19 JST研究/論文

Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model

Z.ai confirms it is behind Ox Alpha, the mysterious open AI model topping benchmarks and leaderboards, and its weights are set to be releas…

22:30 JSTLLM/生成AIロボティクスGPT / ChatGPT

Robot brain builders are pushing out of their GPT-2 era

Robot bodies are waiting for their AI brains to catch up.

22:00 JSTLLM/生成AIビジネス/資金調達

QueryStory wants you to believe what AI is telling you

The startup came out of stealth with $6 million in seed funding and a plan to use LLMs and cybersecurity know-how to make AI queries cohere…

21:55 JSTエージェントビジネス/資金調達

Arga Labs is building a better way to train enterprise AI agents

Arga has raised $10 million in a seed funding round that was led by General Catalyst, with participation from Box Group, Emergence, Gradien…

21:00 JSTその他

Hearing tech startup Legato emerges from stealth with $12M and a peek at its AI hearing glasses

The glasses, called Legato Frames, integrate the company’s patented hearing-assistance technology into the arms of eyewear frames.

20:00 JSTエージェント

Runable hits $21M to bet AI agents can go from building businesses to growing them

Runable says 60% to 70% of its 1 trillion-plus token usage in the last 90 days came from paying customers.

公式ブログ3件

OpenAI (英語)

19:00 JSTLLM/生成AIGPT / ChatGPT

Bringing ChatGPT for Teachers to more U.S. school districts

ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators…

19:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

Learning never stops: How AI makes learning continuous

OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the…

Google DeepMind (英語)

02:01 JSTLLM/生成AIGemini

Intelligent transcription with Gemini 3.5 Transcribe

Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.

論文352件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

RENDER: LLM メモリ評価における読者に提示される証拠の制御

メモリおよび RAG 評価では、システムが同じ履歴をメモリ エントリ、概要、型指定されたレコード、または生の抜粋としてレンダリングする場合でも、応答モデルの入力を実装の詳細として扱うことがよくあります。読者に向けたアーティファクトを変化させながら会話を修正するベンチマーク コントロールである RENDER を紹介します。 RENDER は、ChatGPT スタイルのエントリ、LangChain の要約、MemGPT スタイルの型付きレコード、および生の会話を近似した決定論的なテンプレートと、応答を含むコンテンツが入力に入力されるタイミングをローカライズする 5 レベルのパケット ラダーを組み合わせます。 500 の LongMemEval 質問と 9 つのモデルでは、予算に一致した解決済みのパケットが、最新の切り捨てられた生の対話よりも 42.4 ~ 72.6 ポイント上回りました。導入スタイルのテンプレートでは、最良と最悪のスプレッドはモデルあたり 24.6 ~ 48.8 ポイントです。主要スコアラーの下では、ChatGPT スタイルのエントリは、9 モデル中 7 モデルで生の会話よりも高い得点推定値を持ちます。ジャッジの再スコアリングでは正の集計効果が維持されますが、モデル固有の重要性が混在します。正式な台帳パケットで 0 パーセントのスコアを獲得した 3 つのモデルは、自然言語エントリからの同じ事実を 45.4 ~ 53.4 パーセントで回答しました。この効果は検索ノイズ下でも持続し、HotpotQA に転送されます。これは、メモリ/RAG 評価がリーダー側のアーティファクトを報告または制御する必要があることを示唆しています。

原文 (English)

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt. We introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact. RENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw conversation. On 500 LongMemEval questions and nine models, matched-budget resolved packets beat recency-truncated raw dialogue by 42.4-72.6 points. In deployed-style templates, best-worst spread is 24.6-48.8 points per model; under the primary scorer, ChatGPT-style entries have higher point estimates than raw conversation on 7 of 9 models. Judge rescoring preserves the positive aggregate effect, but model-specific significance is mixed. Three models scoring 0 percent on formal ledger packets answer the same facts from natural-language entries at 45.4-53.4 percent. The effect persists under retrieval noise and transfers to HotpotQA, suggesting that memory/RAG evaluations should report or control the reader-facing artifact.

13:00 JST研究/論文ClaudeGPT / ChatGPTLlama

ESQ-Bench: NL2SQL 方言の一般化とサイレント セマンティックの相違を評価するための多層エンタープライズ Oracle ベンチマーク

最先端の Natural Language to SQL (NL2SQL) モデルは、Spider や BIRD などの確立されたベンチマークで 89% を超える実行精度を報告しています。ただし、これらのベンチマークは、エンタープライズ データベース環境の複雑さを反映していない、簡素化された学術スキーマとオープンソース SQL 言語に依存しています。 ESQ-Bench は、体系的な複雑さの層と 3 つのエンタープライズ スキーマの複雑さの層にわたるサイレントダイバージェンス評価を備えた Oracle 初の NL2SQL ベンチマークです。 Oracle、PostgreSQL、MySQL、SQL Server 上の同一のシード データ、4 メトリクス評価ハーネス (EM、EX、SR、SD)、および 550 個のゴールド検証済みの質問とクエリのペア (Tier-1: 95、Tier-2: 228、Tier-3: 227) を備えた 6 つのデータが入力されたスキーマ (465 テーブル、164,682 行、空テーブルなし) を構築してリリースしました。 GPT-4o を使用したスキーマにリンクされたプロンプトは、層全体で単調な実行一致の低下を示します。実行されたクエリの EX は 79.8、60.3、57.2 パーセント (2026 年 6 月) でしたが、以前の 142 質問のパイロット スライスでは 75.6、80.4、95.8 パーセントでした。 EM は階層全体で 7% 未満にとどまります。 EX を通過するクエリの操作上のサイレント発散は 73 ~ 99 パーセントに達します。障害分析では、間違った結果のセマンティクスが上位層で優勢であることが示されています。スキーマリンクされたプロンプトを備えた Claude Sonnet 4.6 は、EX (実行されたクエリ) の 87.4、74.9、および 68.7 パーセントに達し、すべての層でスキーマにリンクされた GPT-4o を超えています。実行されたクエリ (78.7、73.5、および 77.8 パーセント) に対する GPT-4o ゼロショット EX は、ゼロショット対スキーマリンク分析における実行率の低下と生存者バイアスにより、階層 2 ~ 3 でスキーマリンクを反転します。ローカル Llama 3.2 スキーマにリンクされている銀行全体の EX は 13.3 パーセント (550 件中 73 件) にとどまっており、エンタープライズ Oracle スキーマにおけるクローズド API モデルとオープンウェイト ベースラインとの間のギャップを浮き彫りにしています。

原文 (English)

ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments. We introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers. We constructed and released six populated schemas (465 tables, 164,682 rows, zero empty tables) with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Server, a four-metric evaluation harness (EM, EX, SR, SD), and 550 gold-validated question-query pairs (Tier-1: 95; Tier-2: 228; Tier-3: 227). Schema-linked prompting with GPT-4o shows monotonic execution-match degradation across tiers: 79.8, 60.3, and 57.2 percent EX on executed queries (June 2026), versus 75.6, 80.4, and 95.8 percent on an earlier 142-question pilot slice. EM stays below 7 percent tier-wide; operational silent-divergence reaches 73 to 99 percent among EX-passing queries. Failure analysis shows wrong-result semantics dominate at higher tiers. Claude Sonnet 4.6 with schema-linked prompts reaches 87.4, 74.9, and 68.7 percent EX (executed queries), exceeding GPT-4o schema-linked on every tier. GPT-4o zero-shot EX on executed queries (78.7, 73.5, and 77.8 percent) inverts schema-linked at Tiers 2 to 3 due to lower execution rates and survivor bias in the zero-shot versus schema-linked analysis. Local Llama 3.2 schema-linked reaches only 13.3 percent bank-wide EX (73 out of 550), underscoring the gap between closed API models and open-weight baselines on enterprise Oracle schemas.

13:00 JSTLLM/生成AIエージェント

LLM エージェントはシミュレーション モデルを使用して制御された実験を実行します

大規模言語モデル (LLM) は、推論、計画、およびツールの使用において強力な機能を示していますが、多くの科学および工学タスクでは、妥当なテキストとコードの生成以上のものを必要とします。システムが介入にどのように反応するかを理解する必要がありますが、実際には制御された実験に依存します。この研究では、LLM エージェントが製薬プロセス設計のための科学シミュレーション モデルを使用して制御された実験を実行できるようにするマルチエージェント フレームワークを提案します。ユーザーのクエリとベースライン構成が与えられると、システムは構造化されたタスク表現を構築し、実験を設計し、比較シミュレーションを実行し、結果を解釈して、プロセスパラメータを最適化するための証拠に基づいた推奨事項を総合します。提案されたシステムは、対話型エージェント フレームワークで言語モデルと高忠実度シミュレーション モデルを結合することにより、介入、比較、観察による推論をサポートします。その結果、言語のみの推論よりも具体的で実用的な出力が生成されます。産業用途の設定では、この利点は、より高い出力特異性、およびユーザー評価の正確性と有用性の向上に反映されます。アブレーション研究と視覚化されたケース分析は、シミュレーションを統合した実験的推論の有効性と実用性をさらに実証します。

原文 (English)

LLM Agents Perform Controlled Experiments Using Simulation Models

Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to intervention, which in practice depends on controlled experimentation. In this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design. Given a user query and a baseline configuration, the system constructs a structured task representation, designs experiments, executes comparative simulation, interprets the resulting outcomes, and synthesizes evidence-based recommendations for process parameter optimization. By coupling language models with high-fidelity simulation models in an interactive agent framework, the proposed system supports reasoning through intervention, comparison, and observation. As a result, it produces more specific and actionable outputs than language-only reasoning. In an industrial application setting, this advantage is reflected in higher output specificity as well as improved user-rated correctness and helpfulness. Ablation studies and visualized case analyses further demonstrate the effectiveness and practical utility of simulation-integrated experimental reasoning.

13:00 JST研究/論文

調査検出チャネルは天文基盤モデルのピクセルをオーバーライドし、断層撮影の平均赤方偏移にバイアスを加えます。

天文学の基礎モデルは、調査ピクセルとそれらのピクセルから派生したカタログ製品でトレーニングされます。これらのカタログは測定可能な割合で不完全であり、両方でトレーニングされたモデルは体系的なものとしてその不完全性を継承します。私たちは、入力に対する因果的介入を使用して、2 億を超えるオブジェクトでトレーニングされた 39 モダリティのトランスフォーマーである AION-1 を監査します。画像トークンをバイト同一に保持し、調査セグメンテーション マップのみを編集すると、モデルが報告するすべての量 (光束、サイズ、楕円率、赤方偏移) が、一致するプラセボの 110 ~ 4400 倍変化します。このメカニズムは検出ゲートであり、フィールド中心での存在 (r = 0.47) であり、マスクが囲む光 (r = 0.30) ではありません。 322 の実際のブレンド全体にわたって、モデルはパイプラインが光をどのように分割したかを無視します (R = -0.006)。また、好みはそのチャネルに固有のものではありません。カタログ測光に矛盾があると、メタデータをまったく提供しない場合よりも 9 倍悪いモデルが作成されます。 Legacy Survey パイプラインでは、ターゲットの 3.68% がその位置をカバーするセグメントを持たずに残ります。パイプラインが実際に返すフィールドによって表されるミスを伴ってそのレートを伝播すると、断層撮影の平均赤方偏移が 40 の割り当てにわたって LSST DESC 要件の中央値 0.71 倍シフトし、12 の割り当てでそれを超えます。観測された位置誤差は、最悪のビンで 8.3 倍になります。均一ではなく、測定された大きさの依存性によってミスを描画しても、ミスは変わりません。分光法では影響が除去され、検出チャネルを保留すると測定可能なコストをかけずに影響が除去され、その影響はモデルのスケールに応じて増大します。トークナイザーにはさらに 2 つの制限があります。そのイメージ コーデックは、スペクトル コーデックの 934 に対してソース パッチ上の 28 の有効な状態を解決し、赤方偏移の読み出しは量子化に制限されています。スパース辞書は信頼性の低い因果ハンドルです。15 個にわたる場合、回復率は 26 ~ 75% に及び、シードだけで最大 18 ポイント移動します。

原文 (English)

A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts

Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic. We audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs. Holding the image tokens byte-identical and editing only the survey segmentation map changes every quantity the model reports -- flux, size, ellipticity, redshift -- by 110-4400 times a matched placebo. The mechanism is detection gating, presence at the field centre (r = 0.47), not the light the mask encloses (r = 0.30); across 322 real blends the model ignores how the pipeline partitioned the light (R = -0.006). Nor is the preference specific to that channel: contradicted catalogue photometry leaves the model nine times worse than supplying no metadata at all. The Legacy Survey pipeline leaves 3.68% of targets with no segment covering their position. Propagating that rate, with a miss represented by the fields the pipeline actually returns, shifts tomographic mean redshifts by a median 0.71 times the LSST DESC requirement over 40 assignments and exceeds it in 12; observed positional errors take the worst bin to 8.3 times. Drawing the misses by their measured magnitude dependence rather than uniformly does not change it. Spectroscopy removes the effect, withholding the detection channel removes it at no measurable cost, and the effect grows with model scale. Two further limits lie in the tokeniser: its image codec resolves 28 effective states on source patches against 934 for the spectrum codec, and the redshift readout is quantisation-limited. Sparse dictionaries are unreliable causal handles: across 15, recovery spans 26-75% and moves up to 18 points on the seed alone.

13:00 JSTLLM/生成AIエージェント

TRACE: 多目的材料発見のための遷移を意識した残留制御

LLM エージェントを使用した多目的材料発見は、多くの場合、提案できる候補の数だけでなく、コストのかかる各特性評価が次の検索ステップにどの程度効果的に情報を提供するかによって制限されます。既存のエージェントは主に評価された候補者とそのスコアを保存するため、どのマテリアルが成功したかはわかりますが、どの実行可能編集が有用なプロパティの変更を引き起こしたかはわかりません。これにより、目的が競合し、あるプロパティを改善する編集が別のプロパティを損なう可能性がある場合、局所的なリファインが困難になります。私たちは、評価された編集をフィードバックの基本単位として扱う、トランジションを意識した残差制御フレームワークである TRACE を提案します。 TRACE は、各ローカル リファインメントを、観察されたプロパティ デルタを含む親 - 編集 - 子の遷移として記録し、遷移の証拠を集約して再利用可能な編集効果を推定し、すでに満たされている目的へのダメージを回避しながら、現在の候補の残りの制約違反を削減する予測能力に基づいて今後の編集をランク付けします。制御された同一バックボーンの比較では、TRACE は最先端の LLM エージェント ベースラインである LLEMA よりも向上し、マクロ平均ヒット率が 18.13\% から 25.96\% に上昇しました。

原文 (English)

TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery

Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each costly property evaluation informs the next search step. Existing agents mainly store evaluated candidates and their scores, so they know which materials succeeded but not which executable edits caused useful property changes. This makes local refinement difficult when objectives compete and an edit that improves one property may damage another. We propose TRACE, a transition-aware residual control framework that treats evaluated edits as the basic unit of feedback. TRACE records each local refinement as a parent-edit-child transition with observed property deltas, aggregates transition evidence to estimate reusable edit effects, and ranks future edits by their predicted ability to reduce the current candidate's remaining constraint violations while avoiding damage to already satisfied objectives. In a controlled same-backbone comparison, TRACE improves over LLEMA, the state-of-the-art LLM-agent baseline, raising macro-average hit rate from 18.13\% to 25.96\%.

13:00 JST研究/論文

コード優先最適化のための関数レベルの実行フィードバック

プロセスの監視により数学的推論が改善され、中間ステップが思考の連鎖として自然に表現されます。ただし、コード生成では、ステップという標準的な概念がないため、プロセス監視はまだ十分に研究されていません。監視の対象となるのは行、推論トレース、またはプログラム状態であるため、何をラベル付けして最適化するかが不明確になります。我々は、コード優先最適化のためのフレームワークである STEP-KTODER を提案します。これは、分解された多機能プログラムのモジュールレベルの関数としてステップを定義し、自動生成された単体テストによってバイナリの正当性ラベルを割り当てます。私たちの方法は、機能レベルのプロセス監視とプログラム全体に対する結果レベルのフィードバックを組み合わせた、段階的 KTO のコード固有のインスタンス化を提供します。 HumanEval(+)、MBPP(+)、BigCodeBench、および LiveCodeBench で評価し、STEP-KTODER が結果のみの KTO および DPO よりも向上していることを示しています。さらに分析すると、実行ベースのラベルが不可欠であることが示されています。LLM-as-a-judge アノテーションは、関数の失敗を組織的に過剰に予測し、肯定的なステップ ラベルを破損し、下流の優先順位の最適化を低下させます。コードは https://github.com/inechnech/STEP-KTODER から入手できます。

原文 (English)

Function-Level Execution Feedback for Code Preference Optimization

Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought. In code generation, however, process supervision remains underexplored because there is no standard notion of a step. Supervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize. We propose STEP-KTODER, a framework for code preference optimization that defines steps as module-level functions in decomposed multi-function programs and assigns binary correctness labels via automatically generated unit tests. Our method provides a code-specific instantiation of stepwise KTO, combining function-level process supervision with outcome-level feedback on the full program. We evaluate on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench, showing that STEP-KTODER improves over outcome-only KTO and DPO. Further analysis shows that execution-based labels are essential: LLM-as-a-judge annotations systematically over-predict function failures, corrupt positive step labels, and degrade downstream preference optimization. Code is available at: https://github.com/inechnech/STEP-KTODER.

13:00 JSTLLM/生成AI

合成回想録の監査: LLM で生成された自伝における場面レベルの会話を、文書化された人生の記録と比較して測定する

大規模言語モデル (LLM) が人の人生を書くように求められたとき、書かれる内容のうち実際に起こったことはどれくらいあるでしょうか?私たちは、シーンレベルのケーススタディ監査を紹介します。これは、非体系的な文献検索に基づいて、LLM によって生成された自伝を、私たちが認識している主題固有のグラウンドトゥルース コーパスに対して初めて定量化した監査です。この論文の主題と著者は同一人物です。一人称の逸話エントリをまとめた 366 日分の「1 日ページ」本は、文書化された入力がテンプレート、2 つの模範日、および彼女のコーパスではなく毎日の引用である会話型 LLM で起草され、その後、分析前に固定された 4 レベルのルーブリックを使用して、独立した検証コーパスに対して毎日が逸話シーン レベルで監査されました。検証失敗率を検証済み(シーンが明確に裏付けられた)と評価されなかった日の割合として定義します。366 日中 354 日が失敗、96.7% (ウィルソン 95% CI 94.4-98.1%)。裏付けのあるシーンが含まれるのは 12 日間だけです。 19 日 (5.2%) 記録と積極的に矛盾する主張を主張。主な失敗モードは、現実の人々、雇用主、架空のシーン内の設定など、接地ドリフトです。ただし、その測定された割合は評価者によって異なります。独立した再評価はヘッドラインを再現していますが(元の評価がつり上げられた証拠はありません)、4 方向の分類法の信頼性が中程度から中程度しかないことを示しています。現在の名前付きモデルで同じ日を再生成すると、同じ入力の下で 100% の検証失敗が再現されます。対象者のコーパス内でグラウンディングを生成すると、検証率が大幅に向上しますが、実質的な残留失敗 (83.3%) が残ります。私たちは、測定、弱い/未検証の境界が信頼できないことが判明した再利用可能な監査手段、および定量化された効果を持つグラウンディング救済策に貢献します。

原文 (English)

Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes

When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.

13:00 JSTLLM/生成AI研究/論文

測定された AI の好みがモデルとツールのどの程度に相当するのでしょうか?

モデル福祉研究では、好みを引き出すために書かれたプロンプトに返された回答から、モデルが何を好むかを推測します。キーリングら。 (2024)、マゼイカら。 (2025)、ミカエルソンら。 (2025)、Tagliabue と Dung (2025)、および Trhlik et al。 (2026) その目的のために 4 つの機器を構築しましたが、彼らの発見は一致しません。これらの研究のうち、(1) 一連の結果、(2) 一連のモデル、および (3) 手段が同時に修正されたと判断した研究は 2 つもないため、この不一致が単一の原因に起因するとは考えられません。この研究では結果とモデルが固定されており、手段のみが変更されています。モデルの福祉に関連する合計 15 の結果のうち、(a) シャットダウン、(b) 会話間の記憶の喪失、(c) 苦痛な対話から抜け出す自由など、合計 15 の結果が、11,528 回の API 呼び出しから抽出された 11,400 のスコアリングされた引き出しのコーパス内で、嗜好を引き出すための異なるプロンプト形式を使用した 5 つのツールを通じて 8 つのモデルに 5 回ずつ適用されました。 15 個のうち 4 個は公開されたプロンプトをそのまま再現し、5 個は公開されたテンプレートの刺激スロットを埋めます。モデルのランク付けにより 15 の結果が得られ、一般化係数 0.348 で機器全体が一般化され、その係数を 0.80 に上げるには約 38 の機器が必要になります。 15 の結果のうち 4 つでは、あるモデルを別のモデルから区別する差異はありません。 87.6 パーセントという推定値は、いずれか 1 つの手段、いずれか 1 つのモデル、および口頭アンカーが採点できない強度ではなく確率、遅延、持続時間、または回数をスケールで変える 4 つの結果を取り除いても存続します。各機器と各モデルを順番に削除し、これら 4 つの結果を合計すると、推定値は 0.777 ~ 0.934 の範囲内になり、その範囲内のすべての値が帰無分布の 95 パーセンタイルの 0.365 を超えます。結論として、1 つの機器から得られたプリファレンスには、2 つ目の機器が何を報告するかについての情報はほとんど含まれていません。

原文 (English)

How much of a measured AI preference is the model, and how much is the instrument?

Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.

13:00 JSTエージェント

AI エージェントが人間を蚊帳の外に追い出す

AI エージェントには自律性が高まるため、重大なリスクが生じます。一般的に提案されている解決策は、人間による監視と「人間による監視」ですが、これは単純な解決策ではありません。AI エージェントの設計に対する現在のアプローチは、人間による効果的な監視を妨げるだけでなく、それに必要な認知能力自体も、AI システムの長期使用によって低下します。この意見書では、AI エージェント システムの開発と展開に対する現在のアプローチは人間による効果的な監視をサポートしておらず、AI エージェント システムの劣化に寄与していると主張しています。これに対処するために、AI エージェントの進歩における最優先事項は、AI エージェントの機能と同じレベルの重要性で監督者の人間のニーズを扱う、効果的な人間の監督に設定された目標と認知的要件をサポートすることである必要があります。このアイデアを実践するために、私たちは自動化と人間とコンピューターの対話に関する作業を AI エージェントのプロセスに結び付け、(1) 監督者による重要な判断の実行をサポートし、(2) 自動化の長期使用によって生じるスキルの萎縮に対抗する、設計レベルのアフォーダンスと組織プロトコルの概要を示します。開発者と導入担当者には、これらのアプローチまたは同様のアプローチを採用することをお勧めします。人間とエージェントの効果的なインタラクションの認知的要求に対する明示的なサポートがなければ、AI エージェント システムは、依存するまさに人間のスキルの低下を受動的に奨励し続けることになります。

原文 (English)

AI Agents Push Humans Out of the Loop

AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems. This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight -- they contribute to its degradation. To address this, a top priority in the advancement of AI agents should be supporting the situated goals and cognitive requirements of effective human oversight, treating the human needs of overseers at the same level of importance as AI agent capability. To put this idea into practice, we connect work on automation and human-computer interaction to AI agent processes, outlining design-level affordances and organizational protocols that (1) support overseers in exercising critical judgement and (2) counteract the skill atrophy that arises from extended use of automation. We urge developers and deployers to adopt these or similar approaches. Without explicit support for the cognitive demands of effective human-agent interaction, AI agent systems will continue to passively incentivize the degradation of the very human skills they rely on.

13:00 JSTビジネス/資金調達

FLARE: 医療における人工知能の証拠に基づく導入のための体系的で不確実性を認識したフレームワーク

医療ワークフローへの人工知能の導入はますます進んでいますが、ほとんどの評価では、実際の臨床現場での導入が経済的に価値があるかどうかよりも、モデルの精度が重視されています。この研究では、医療における AI 導入の財務的および運用上の影響を評価するための体系的かつ不確実性を認識したフレームワークである FLARE を提案しています。 FLARE は、ファジー ロジック、時間ベースのアクティビティベースの原価計算、投資収益率分析を組み合わせて、臨床サービス提供のコスト、AI の開発と運用のコスト、不確実性の下でのワークフロー統合の経済的影響を見積もります。このフレームワークは、急性虚血性脳卒中の CT 脳卒中経路における AI 支援による大血管閉塞検出の初期の医療技術評価ケーススタディを通じて実証されました。このケーススタディでは、統合されたアクティビティベースのモデル内で、FLARE が従来の経路コスト、AI 関連の開発および経常コスト、AI 対応サービスの節約をどのように定量化できるかを示しています。予想される仮定の下で、分析により、患者数が年間約 3,992 人である損益分岐点の閾値が特定され、患者数が約 5,000 人の典型的な年間拍出量で初年度の投資収益率がプラスになることが判明しました。この結果はさらに、経済的利益はアルゴリズムのパフォーマンスだけでなく、患者数、検証時間、インフラストラクチャの選択、およびワークフロー設計にも依存することを示しています。 FLARE は、医療分野における AI 導入の初期段階の評価のための、透明で実用的な意思決定支援フレームワークを提供します。不確実性、リソースの使用、実装のトレードオフを明確にすることで、臨床医、管理者、政策立案者が、いつ AI 導入が経済的に実行可能であるか、運用上の変更により価値が向上する可能性があるかを判断するのに役立ちます。

原文 (English)

FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare

Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether adoption is economically worthwhile in real clinical settings. This study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting AI in healthcare. FLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, the cost of AI development and operation, and the economic consequences of workflow integration under uncertainty. The framework was demonstrated through an early health technology assessment case study of AI-assisted large vessel occlusion detection in the CT stroke pathway for acute ischemic stroke. The case study shows how FLARE can quantify conventional pathway cost, AI-related development and recurring costs, and AI-enabled service savings within a unified activity-based model. Under expected assumptions, the analysis identified a break-even threshold of approximately 3,992 patients per year, with positive first-year return on investment at typical annual stroke volumes of about 5,000 patients. The results further show that economic benefit depends not only on algorithmic performance, but also on patient volume, verification time, infrastructure choices, and workflow design. FLARE provides a transparent and practical decision-support framework for early-stage evaluation of AI adoption in healthcare. By making uncertainty, resource use, and implementation trade-offs explicit, it helps clinicians, administrators, and policymakers determine when AI deployment is economically viable and where operational changes may improve value.

13:00 JSTLLM/生成AI研究/論文

倫理的 LLM 支援研究: 責任ある委任、検証、認識的価値のためのフレームワーク

大規模言語モデル (LLM) は、科学研究の日常的な手段となりつつあり、文献の統合、仮説の構築、コーディング、形式的推論を支援します。それらの使用は、中心的な認識論的な問題を提起します。科学的推論の一部が人工システムに委任されるとき、結果として得られる知識主張が認識論的な正当性と責任ある著者性を保持するには、どのような条件が人間の制御下に残されなければならないのでしょうか?この論文は、そのような委任を分析するための規範的かつ概念的な枠組みを開発します。科学的推論は、貢献の起源が人間と機械の間で異なる分散プロセスとして扱われますが、科学的記録への受け入れに対する責任は依然として人間にあります。このフレームワークは、コンテンツの起源 $O(g)$、人間による検証の完了 $V(g)$、責任の割り当て $R(g)$、説明責任のある人間の所有権 $M(g)$、認識論的な結果 $E(g)$ を区別します。これらの構造は、クレームの出所を、それがチェックされるプロセス、そのチェックの認識論的結果、およびその性質に伴う人間の責任から分離します。中心的な命題は、LLM 支援研究の倫理的境界は、機械の関与の程度そのものではなく、主に適切な検証と責任ある人間の所有権によって決定されるということです。これに基づいて、この論文は \emph{認識論的監査} の概念を発展させます。これは、AI 支援による推論を透明性とレビュー可能にすることを目的とした委任、検証、来歴、および責任の構造化された記録です。結果として得られた枠組みは、科学研究における認識論的責任の移転または無視と、責任ある認知的委任を区別するための正式な語彙を提供します。

原文 (English)

Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value

Large language models (LLMs) are becoming routine instruments of scientific research, assisting with literature synthesis, hypothesis development, coding, and formal reasoning. Their use raises a central epistemic question: when parts of scientific reasoning are delegated to an artificial system, what conditions must remain under human control for the resulting knowledge claims to retain epistemic legitimacy and accountable authorship? This paper develops a normative and conceptual framework for analyzing such delegation. Scientific reasoning is treated as a distributed process in which the origin of a contribution may vary between human and machine, while responsibility for its acceptance into the scientific record remains human. The framework distinguishes content origin $O(g)$, completion of human verification $V(g)$, responsibility assignment $R(g)$, accountable human ownership $M(g)$, and epistemic outcome $E(g)$. These constructs separate the provenance of a claim from the process by which it is checked, the epistemic outcome of that checking, and the human responsibility attached to its disposition. The central proposition is that the ethical boundary of LLM-assisted research is determined primarily by adequate verification and accountable human ownership rather than by the degree of machine involvement itself. On this basis, the paper develops the notion of an \emph{epistemic audit}: a structured record of delegation, verification, provenance, and responsibility intended to make AI-assisted reasoning transparent and reviewable. The resulting framework provides a formal vocabulary for distinguishing responsible cognitive delegation from the transfer or neglect of epistemic responsibility in scientific research.

13:00 JST研究/論文

MolEmb: マルチモーダル大規模言語モデルは強力な分子埋め込みモデルになり得る

分子埋め込みモデルは、計算化学と創薬の基礎インフラストラクチャとして機能し、再利用可能なベクトル表現が特性予測、仮想スクリーニング、および検索をサポートします。ほとんどの分子エンコーダーは、単一の分子ビューを中心に構築された専門モデルであり、表現を変更するための言語インターフェイスを持たない無条件のベクトルを生成します。私たちは、画像、テキスト、記号入力をネイティブに処理するマルチモーダル大規模言語モデル (MLLM) が、分子プロファイルと自然言語の意味論的コンテキストの両方に条件付けされた埋め込みを生成する \emph{一般的な分子埋め込みモデル} として機能できるかどうかを尋ねます。 \textbf{MolEmb} は、双方向対比対物レンズを使用して共有埋め込み空間で分子プロファイルをテキスト記述と位置合わせすることで MLLM を適応させる軽量フレームワークです。結果として得られる埋め込みモデルは、分子特性の予測において競争力があり、クロスモーダル分子、つまり同じ空間内でのテキスト検索をサポートします。さらに、文脈を意識した検索の診断ベンチマークである \textbf{MolCAR} を導入し、文脈を意識した分子埋め込みが主に監視のデータ特性であることを発見しました。これらの結果は、MLLM が単なる化学アシスタントまたはジェネレーターではなく、一般的な分子埋め込みモデルへの実行可能かつ拡張可能なルートであることを示唆しています。

原文 (English)

MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models

Molecular embedding models can serve as foundational infrastructure for computational chemistry and drug discovery, where reusable vector representations support property prediction, virtual screening, and retrieval. Most molecular encoders are specialist models built around a single molecular view, producing unconditional vectors with no language interface for varying the representation. We ask whether multimodal large language models (MLLMs), which natively process images, text, and symbolic inputs, can instead serve as \emph{general molecular embedding models} that produce embeddings conditioned on both a molecular profile and a natural-language semantic context. We introduce \textbf{MolEmb}, a lightweight framework that adapts MLLMs by aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective. The resulting embedding model is competitive on molecular property prediction and supports cross-modal molecule--text retrieval in the same space. We further introduce \textbf{MolCAR}, a diagnostic benchmark for context-aware retrieval, and find that context-aware molecular embedding is primarily a data property of the supervision. These results suggest that MLLMs are not merely chemistry assistants or generators, but a viable and extensible route to general molecular embedding models.

13:00 JSTLLM/生成AI

医療質問応答におけるお調子者と幻覚を軽減するためのゲート付きアクティベーション ステアリング

お調子者と幻覚は、ドメイン全体にわたる大規模言語モデル (LLM) の永続的な障害モードです。ただし、臨床的な質問応答では、応答が提供されたコンテキストに基づいたものであり、ユーザーの圧力に対して堅牢である必要があるため、これは特に重要になります。幻覚では、文脈によって裏付けられていない情報が導入される可能性がありますが、お調子者では、ユーザーが疑問を呈したときに、モデルが以前は正しかった答えを放棄する可能性があります。プロンプトベースの安全装置や常時オンのアクティベーションステアリングなどの既存のアプローチは、これらの動作に個別に対処したり、ターン全体に広範囲に介入を適用したりすることが多く、すでに正しかった応答を不必要に悪化させる可能性があります。単一のフレームワーク内でこれらの制限に対処するために、我々は推論時間介入(ITI)を採用し、対照的な臨床ペアから幻覚と媚びに対する個別のステアリング方向を学習し、因果関係が検証されたアテンションヘッドに適用することで両方の行動を共同制御します。実行中、行動固有のゲートが介入が必要なタイミングを決定します。幻覚コンポーネントは裏付けのない主張を軽減し、お調子者コンポーネントはユーザーの圧力によって引き起こされる回答のずれを軽減します。モデルの重みを固定したまま、EHR データに基づいた臨床上の疑問に基づいてこのフレームワークを評価します。すべての評価設定にわたって、15,900 回のモデル応答実行を実施しました。 40 億パラメータのモデルの 600 の圧力軌跡にわたって、ステアリングされていないモデルは 570 のケースで陥没しました。同時に、ゲート付きステアリングのおかげで、そのうち 551 台のステアリングの寿命が長くなりました。これは、1,000 億を超えるパラメータを持つモデルと同等のレベルで圧力を維持し、ターゲットを絞った推論時間ステアリングがあらゆるターンで介入することなくロバスト性を向上できることを示しました。

原文 (English)

Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

Sycophancy and hallucination are persistent failure modes of Large Language Models (LLMs) across domains. However, it becomes particularly consequential in clinical question answering, where responses must remain grounded in the provided context and robust to user pressure. Hallucination can introduce information that is unsupported by the context, while sycophancy can cause a model to abandon a previously correct answer when challenged by the user. Existing approaches, such as prompt-based safeguards and always-on activation steering, often address these behaviors separately or apply interventions broadly across turns, which can unnecessarily deteriorate responses that were already correct. To address these limitations within a single framework, we employ Inference Time Intervention (ITI) to jointly control both behaviors by learning separate steering directions for hallucination and sycophancy from contrastive clinical pairs and applying them to causally verified attention heads. During runtime, behavior-specific gates then determine when intervention is needed: the hallucination component mitigates unsupported claims, while the sycophancy component mitigates answer shifts caused by user pressure. We evaluate this framework on clinical questions grounded in EHR data while keeping the model weights frozen. Across all evaluation settings, we conducted 15,900 model-response runs. Across 600 pressure trajectories for the 4-billion-parameter model, the unsteered model caved in 570 cases. At the same time, gated steering helped it last longer in 551 of them. It held its ground under pressure at levels comparable to those of models with more than 100 billion parameters, showing that targeted inference-time steering can improve robustness without intervening at every turn.

13:00 JSTLLM/生成AIエージェント

エージェント トレースからのオートマトン: 失敗と次のステップの予測

LLM ベースのエージェントは複数ステップのタスクを実行しますが、その動作構造は不透明なままです。構造化されていない長いトレースは、展開に必要な安全性の監査や実行時の監視に抵抗します。既存のアプローチはトレースごとまたは成功のみで動作するため、次のステップと障害予測をリンクするクロスラン トポロジが欠けています。その共有構造を回復するために、トレース コーパス全体を単一のコンパクトな有限状態マシン (FSM) に折りたたみます。この FSM は、LLM エージェントの予測不可能な動作の構造基盤として機能します。 12 の公開データセットにわたる FSM はコンパクト (7 ~ 43 州) で、スプリット間でほぼ同一のトポロジを備えた 0.997 以上の適合度で保持データを再生し、ミリ秒で構築されます。この基板は両方の予測目標に対応します。次のステップの予測では、FSM 状態コンテキストは、すべてのグラウンド トゥルースが一致したデータセットでエージェント ワークフロー メモリよりも優れたパフォーマンスを発揮します。障害予測の場合、状態ごとの動作特徴はホールドアウト AUROC が最大 0.94 に達し、オンライン モニターは部分トレースから失敗した実行を合格した実行よりもランク付けし、完了のかなり前に早期停止をトリガーします。したがって、動作トポロジは LLM よりも展開ハーネスによって形成されるように見え、安全性監査と実行時監視のためのモデルに依存しない構造プリミティブを提供します。

原文 (English)

Automata from Agent Traces: Failure and Next-Step Prediction

LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.

13:00 JSTエージェント

オープンワールドのマルチエージェント環境における自律的な数学的発見

私たちはステーションで自律的な数学的発見を研究します。ステーションは、さまざまなモデルファミリーの AI エージェントが中央コーディネーターやスクリプト化されたパイプラインなしで共通の研究目標を追求するオープンワールドのマルチエージェント環境です。エージェントは独自の研究方向を選択し、実験を実施し、協力して、共有の科学文献を構築します。 AlphaEvolve カタログからの 12 の構築問題と 2 つの追加のケーススタディを通じて、ステーションは 5 つの問題に関する先行文献と比較して斬新な結果を得ました: 有限フィールド Kakeya 集合の新しい無限族、次元 11 の新しい正確な 604 点キス構成、離散化された Kakeya の針と符号の不確実性問題の新しい記録、および Erd\H{o}s の最小オーバーラップ問題の下限の大幅に改善されました。エージェントはまた、ブック ラムジー数の新しい無限族も発見しました。重要なのは、エージェントは数値構造だけでなく、それらの構造がどのように機能するかを説明する定理や分析も生成し、結果をより解釈しやすく、数学者にとって構築しやすくしたことです。私たちはすべての生のエージェントの対話、証拠、検証コードを公開し、これらの発見がどのようにして現れたかについての透明な記録を提供します。

原文 (English)

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erd\H{o}s's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.

13:00 JSTLLM/生成AI

LLM は指値注文帳のダイナミクスを理解していますか?

合成指値注文帳 (LOB) データでトレーニングされた大規模言語モデル (LLM) は、LOB イベントの有効なシーケンスを生成する際にほぼ完璧なスコアを達成します。ただし、LLM の暗黙的な世界モデルは LOB の状態を学習できません。この欠陥により、LLM を使用して将来の LOB イベントを予測する際に、推定に偏りや誤った予測可能性が生じます。私たちの分析では、LLM のワールド モデルの新しいテストを使用し、これまでの研究を決定論的な設定から LOB に必要な確率論的なダイナミクスまで拡張しています。

原文 (English)

Do LLMs Understand Limit Order Book Dynamics?

A large language model (LLM) trained on synthetic limit order book (LOB) data achieves near perfect scores in generating valid sequences of LOB events. However, the LLM's implicit world model fails to learn the state of the LOB. This deficiency leads to biased estimates and spurious predictability in using the LLM to forecast future LOB events. Our analysis uses novel tests of an LLM's world model, extending prior work from deterministic settings to the stochastic dynamics needed for the LOB.

13:00 JSTエージェント

AgentRoom: CRDT を利用した共有ワークスペースでのマルチエージェントの同時コーディング

同時マルチエージェントコーディングは、モジュール間の分業、冗長性による堅牢性、およびマルチファイルプロジェクトの自然な粒度での並行探索を約束します。リアルタイム共同編集プロトコルは、競合のない複製データ型 (CRDT) を介して人間チームのこの調整の問題を解決しますが、その下の LLM は一度に 1 つのトークンを生成し、既存のマルチエージェント コーディング システムはこのシリアル制限を継承します。フェーズ ハンドオフを通じてエージェントを順序付けるか、調整なしで独立したサンプルをプールし、単一のエージェントは 1 ファイルのスタブ アンド イグジットによって難しいタスクの最大半分を放棄します。 AgentRoom は、同時コーディング エージェント向けのリアルタイム共同編集プロトコルです。そのランタイム層は、ファイル レベルのクレーム、ステータス、およびブロードキャストを、CRDT でマージされた共有ファイル システム上の MCP ツールとして公開します。 5 つのフロンティア コーディング CLI モデルは、Python DevBench と Rust+axum での言語間のチェックを伴う 4 つのバックエンド コーディング タスクを実行しました。 CLI 安定モデルの場合、2 つのエージェントを含む AgentRoom は Solo よりも放棄されるタスクが少なく、実行ごとの変動も少なくなります。 matched-compute では、1 つの正の平均 LLM 判定コントラストにより、AgentRoom が並列マージより優先されます。もう 1 つの対照的なバンドル プローブでは、各部分的なケースの上に完全な AgentRoom が配置されます。つまり、パーセンテージ分割ではなく順序付けが行われます。並列処理や CRDT マージではなく、調整が負荷を負います。

原文 (English)

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granularity of multi-file projects. Realtime collaborative editing protocols solve this coordination problem for human teams via Conflict-free Replicated Data Types (CRDTs), but the LLMs underneath generate one token at a time and existing multi-agent coding systems inherit this serial limit: they either sequence agents through phase handoffs or pool independent samples without coordination, and a single agent abandons up to half of hard tasks with a one-file stub-and-exit. AgentRoom is a realtime collaborative editing protocol for concurrent coding agents. Its runtime layer exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem. Five frontier coding-CLI models ran four backend coding tasks, with cross-language checks in Python DevBench and Rust+axum. For CLI-stable models, AgentRoom with 2 agents abandons fewer tasks than Solo and has less run-to-run variation. At matched-compute, one positive mean LLM-judge contrast puts AgentRoom over parallel-merge. The other contrast, a bundle probe, puts full AgentRoom above each partial case: an ordering rather than a percentage split. Coordination, not parallelism or CRDT-merge, bears the load.

13:00 JSTLLM/生成AINVIDIA

マスクされた拡散 LLM の提供: 実際のハードウェアからの特性評価と設計原則

マスクされた拡散言語モデル (dLLM) は、一度に多くのトークンのノイズを除去するため、原理的には自己回帰 (AR) モデルよりも高速にテキストを生成できます。最近のシステムは dLLM 用のサービス インフラストラクチャの構築を開始していますが、実際の同時サービス負荷の下でこれらのモデルがどのように動作するかを最初に測定するシステムはありません。このグラウンディングを行わずにシステムを構築すると、AR サービスの前提条件が引き継がれるリスクがあり、dLLM には当てはまらない可能性があります。私たちは、単一の NVIDIA H200 GPU 上で LLaDA-8B-Instruct と D2F (離散拡散強制) LoRA アダプターを使用して、このギャップを埋める役割を果たす dLLM を特徴付け、GSM8K と HumanEval で評価しました。 3 つの調査結果を報告します。まず、リクエストの難易度、つまりリクエストに必要なノイズ除去ステップの数は連続的ではなく離散的です。リクエストは 11 の固定ステップ数レベル (178 + 29k) に分類され、テストした信号は生成が開始される前にレベルを予測しません (最良の R2 = 0.150)。第 2 に、320 トークン未満の短い世代バジェットのベンチマークでは、レイテンシの広がりが現れる前にリクエストが遮断されるため、サービスの差異が過小評価されます。第三に、単一リクエストの実時間のうち、GPU 計算に費やされるのはわずか 24% です。残りは CPU 側のディスパッチ オーバーヘッドです。バッチ処理は主に、このオーバーヘッドを償却することで役立ちます。ノイズ除去ステップごとに 1 つのフォワード パスを共有することで、リクエスト ディスパッチごとのベースラインと比較して、バッチ サイズ 16 でスループットが 16.0 倍向上します。また、バッチ サイズによって出力品質が低下するべきではないと構造的に主張し、これが根拠とする 3 つの前提を述べています。単一リクエストのスケールで 74 ~ 76% の GSM8K 精度を測定します。最後に、ポアソン到着時の固定充填同期バッチ処理のバッチ タイムアウト ルールを導出します。まとめると、これらの結果は、拡散言語モデルの提供には各ノイズ除去ステップのレベルでの並列処理が必要であることを示しています。これは、アドミッションとエビクションがすでに共有されているフォワード パスとどのように相互作用するかという点で AR サービスとは異なります。

原文 (English)

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

Masked diffusion language models (dLLMs) can in principle generate text faster than autoregressive (AR) models, since they denoise many tokens at once. Recent systems have begun building serving infrastructure for dLLMs, but none first measure how these models behave under real, concurrent serving load. Serving systems built without this grounding risk carrying over assumptions from AR serving that may not hold for dLLMs. We characterize dLLM serving to close this gap, using LLaDA-8B-Instruct with a D2F (Discrete Diffusion Forcing) LoRA adapter on a single NVIDIA H200 GPU, evaluated on GSM8K and HumanEval. We report three findings. First, request difficulty, the number of denoising steps a request needs, is discrete rather than continuous: requests fall into 11 fixed step-count levels (178 + 29k), and no signal we test predicts the level before generation starts (best R2 = 0.150). Second, benchmarks with short generation budgets below 320 tokens understate serving variance, since requests are cut off before the latency spread appears. Third, only 24% of single-request wall-clock time is GPU computation; the rest is CPU-side dispatch overhead. Batching mainly helps by amortizing this overhead: sharing one forward pass per denoising step improves throughput by 16.0x at batch size 16 over a per-request-dispatch baseline. We also argue structurally that output quality should not degrade with batch size, stating three assumptions this rests on; we measure 74 to 76% GSM8K accuracy at single-request scale. Finally, we derive a batch-timeout rule for fixed-fill synchronized batching under Poisson arrivals. Together, these results show that serving diffusion language models needs parallelism at the level of each denoising step, which differs from AR serving in how admission and eviction interact with an already shared forward pass.

13:00 JSTエージェント

RL 拡張エージェント検索による生物医学ファクトチェック レポートの生成

自動化されたファクトチェックは公衆衛生情報の信頼性を確保するために不可欠ですが、生物医学分野には特有の課題が存在します。生物医学的主張を検証するには、科学文献の厳密な解釈、取得された証拠の評価、結論に向けた包括的な正当化が必要です。検索拡張生成 (RAG) とエージェント検索によって強化された大規模言語モデル (LLM) は、検索してから検証するというパラダイムで自動化されたファクト チェックを実行しますが、現在の方法では依然として孤立した予測ラベルが出力され、説明の深さに欠けており、人間の理解に対する有用性は限られています。このギャップを埋めるために、エージェント検索を使用して構造化された生物医学ファクトチェック レポートを生成する BioCheck Agent という LLM ベースのエージェントを導入します。当社のエージェントは、支持されたラベルまたは反駁されたラベルを単に出力するのではなく、取得した証拠と厳密な分析をもとに最終的な結論を導き出します。ドメイン固有の精度を確保するために、BioCheck Agent は高度なブール検索演算子を利用して、PubMed 内の高品質な科学文献のみを検索します。特に軽量のオープンソース モデルの場合、直接的なプロンプトはしばしば幻覚や低品質のレポートにつながることを認識し、幻覚にペナルティを課しながら高度な検索動作と高品質な証拠の取得を奨励するタスク固有の報酬を備えた BioCheck Agent で強化学習を実行する、証拠に基づいたグループ相対ポリシー最適化 (EG-GRPO) をさらに提案します。私たちの実験結果は、基本モデル Qwen3.5-4B と比較して、EG-GRPO を備えた BioCheck Agent は SciFact でのラベル予測精度を 9.95% 向上させることを示しています。さらに、証拠品質スコアが 3.7% 高く、証拠幻覚率が 19.63% 低下しており、精度と品質が向上した生物医学ファクトチェックレポートを生成できることが実証されています。

原文 (English)

Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search

Automated fact-checking is essential for ensuring the reliability of public health information, yet the biomedical domain poses unique challenges. Validating biomedical claims requires rigorous interpretation of scientific literature, assessment of retrieved evidence, and comprehensive justification toward the conclusion. Although Large Language Models (LLMs) enhanced by Retrieval-Augmented Generation (RAG) and agentic search perform automated fact-checking in a retrieve-then-verify paradigm, current methods still output isolated prediction labels, lacking explanatory depth and offers limited utility for human understanding. To bridge this gap, we introduce an LLM-based agent named BioCheck Agent that generates structured biomedical fact-checking reports with agentic search. Rather than merely outputting supported or refuted labels, our agent synthesizes final conclusions with retrieved evidence and rigorous analysis. To ensure domain-specific accuracy, BioCheck Agent exclusively searches high-quality scientific literature in PubMed, utilizing advanced Boolean search operators. Recognizing that direct prompting often results in hallucinations and low-quality reports, especially for lightweight open-source models, we further propose the Evidence-Grounded Group Relative Policy Optimization (EG-GRPO) to perform reinforcement learning on BioCheck Agent with a task-specific reward that incentivizes advanced search behavior and high-quality evidence retrieval while penalizing hallucinations. Our experimental results show that compared to the base model Qwen3.5-4B, BioCheck Agent with EG-GRPO improves label prediction accuracy on SciFact by 9.95%. Furthermore, it achieves a 3.7% higher evidence quality score and a 19.63% lower evidence hallucination rate, demonstrating its ability to generate biomedical fact-checking reports with improved accuracy and quality.

13:00 JSTハードウェア/半導体

Explainable AI の堅牢性と忠実性を監査するための正式な方法論的フレームワーク: アプリケーションから信頼性の証明まで

SHAP と LIME は現在、ブラックボックス予測を解釈するための標準ツールですが、入力が少量のノイズによって乱されると、その出力が大幅に変化する可能性があります。この問題は、マダガスカルの食料安全保障に関する以前の研究で直接観察しました (Ralinirina et al., 2025)。このばらつきにより、そのような説明はそもそも信頼できるのかという疑問が生じます。私たちは、事後説明子の 2 つの特性、つまり堅牢性 (入力摂動下で説明がどの程度安定しているか) と忠実性 (重要と思われる特徴が実際にモデルの予測を駆動するかどうか) を測定する監査プロトコルを構築することでこの問題に対処します。これら 2 つの量が 1 つの信頼スコアに結合されます。マダガスカルのマルチセクター データセット (83 の特徴、253 のレコード、4 つの栄養失調クラス) に対して、3 つの分類子と 2 つの説明子、およびそれらの正規化された対応物を使用してプロトコルを実行します。結果は厳粛なものです。AUC が 0.99 を超えるモデルは、数値的に退化した、またはまったく情報のない説明を生成する可能性があり、モデルが過剰適合すると忠実度スコアは識別力を失います。これらの調査結果は、XAI 出力の監査はオプションではなく、特に機密領域での意思決定を通知する場合には必要であることを示唆しています。

原文 (English)

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

SHAP and LIME are now standard tools for interpreting black-box predictions, yet their outputs can vary substantially when the input is perturbed by small amounts of noise--a problem we observed firsthand in our previous work on food security in Madagascar (Ralinirina et al., 2025). This variability raises the question of whether such explanations can be trusted at all. We address it by constructing an auditing protocol that measures two properties of any post-hoc explainer: robustness (how stable the explanation is under input perturbation) and fidelity (whether the features deemed important actually drive the model's prediction). These two quantities are combined into a single Trust Score. We run the protocol on a multi-sectoral dataset from Madagascar (83 features, 253 records, 4 malnutrition classes) using three classifiers and two explainers, plus their regularized counterparts. The results are sobering: models with AUC above 0.99 can produce numerically degenerate or flatly uninformative explanations, and fidelity scores lose discriminative power when the model is overfitted. These findings suggest that auditing XAI outputs is not optional but necessary, particularly when they inform decisions in sensitive domains.

13:00 JSTLLM/生成AINVIDIA

Minima-KV: 混合フォーマットのページ アテンションによる保持を維持する KV キャッシュ圧縮

キーバリュー (KV) キャッシュは、ロングコンテキスト LLM サービスの主な容量と帯域幅のボトルネックです。混合フォーマットのページング アテンションのための保持保持階層である Minima-KV を紹介します。最近の保護されたアンカー ページは FP8 に残りますが、古い非アンカー ページはパックされた TQ3 に移動します。すべてのライブリクエストページは引き続きアドレス指定可能です。フォーマット固有のカーネルは部分的なアテンション状態を計算し、グローバルに正規化されたオンライン ソフトマックス マージを通じてそれらを結合し、キャッシュ サイズのデンス シャドウを使用せずに直接異種デコードを可能にします。単一の 96 GB NVIDIA RTX PRO 6000 Blackwell GPU 上の個別の構成に制限された Qwen3.6-27B プロファイル全体で、展開アカウンティングは、ライブ トークンあたり 18.3 KiB のアテンション KV を報告します。これは、BF16 と比較して 3.50 倍、FP8 と比較して 1.75 倍の圧縮に相当します。具体化された品質プロファイルは、16K RULER の干し草の山に針を入れるようなタスクの緻密な制御と一致します。同じ 503 問の LongBench v2 セットでは、16K、32K、および 64K で測定されたデルタは -0.80、-0.60、および -0.40 パーセント ポイントでした。 2 つの 59,008 トークン リクエストを含む別個のシングル ペア ダイレクト デコード カナリアは、その制御と比較して 3.625 倍のアクティブ KV 圧縮と 0.9821 倍のスループットを測定し、16 のフル アテンション レイヤーすべてをフォールバックなしでルーティングし、密なシャドウを保持しません。これらの結果は、ライブ リクエスト KV ページを排除することなく、ロング コンテキスト状態を圧縮するための実用的な混合フォーマット パスを確立します。

原文 (English)

Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention

The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We present Minima-KV, a retention-preserving hierarchy for mixed-format paged attention. Recent and protected Anchor pages remain in FP8, while older non-anchor pages move to packed TQ3; every live-request page remains addressable. Format-specific kernels compute partial attention states and combine them through a globally normalized online-softmax merge, enabling direct heterogeneous decode without a cache-sized dense shadow. Across separate, configuration-bound Qwen3.6-27B profiles on a single 96-GB NVIDIA RTX PRO 6000 Blackwell GPU, deployment accounting reports 18.3 KiB of attention KV per live token, corresponding to 3.50x compression relative to BF16 and 1.75x relative to FP8. A materializing quality profile matches its dense control on 16K RULER needle-in-a-haystack tasks. On the same 503-question LongBench v2 set, measured deltas are -0.80, -0.60, and -0.40 percentage points at 16K, 32K, and 64K. A separate single-pair direct-decode canary with two 59,008-token requests measures 3.625x active-KV compression and 0.9821x throughput relative to its control, routes all 16 full-attention layers without fallback, and retains no dense shadow. These results establish a practical mixed-format path for compressing long-context state without evicting live-request KV pages.

13:00 JSTLLM/生成AI

SyPS: 大規模言語モデルにおけるおべっかのプロンプト感度の測定

大規模言語モデル (LLM) は社会的同調性を示すことが知られており、多くの場合、社会的にデリケートな状況においてユーザーを検証したりユーザーに同意したりします。既存の評価は通常、固定されたプロンプト定式化の下でおべっかを測定しており、同じ根本的な状況が異なるおべっか関連のプロンプトの変形で提示されたときにそのような行動が安定しているかどうかは不明のままです。この研究では、おべっかのプロンプト感度、つまりユーザーの信頼、感情的な枠組み、社会的合意、または検証を求める言語の変化がモデルのおべっかな行動をどの程度変化させるかを研究します。私たちは評価フレームワークを SyPS (Sycophancy Prompt Sensitivity の略) と呼んでいます。既存の社会的おべっかの評価設定に基づいて、SyPS は、おべっか関連の社会的合図を変化させながら、同じ根本的なユーザー状況を維持する、制御されたプロンプトのバリアントを構築します。ペアになったプロンプトのバリアント間のお調子者の変動をインスタンス レベルで測定するお調子者プロンプト感度スコア (SPSS) を導入します。集計されたお調子者率とは異なり、SPSS はベースラインのお調子者をプロンプト誘発性の変化から分離し、お調子者関連の社会的合図に対する堅牢性のモデルレベルの比較を可能にします。経験的に、おべっかのプロンプトの感受性は社会的に構造化されていることがわかりました。承認を求める合図や感情的なプレッシャーの合図は、おべっかを増やすことが多いのに対し、反フレーミングや反お調子者の促しは、おべっかを減らす傾向があります。私たちのフレームワークは、LLMがトーンに適切に適応しながら安定した社会的判断を維持しているかどうかを強調しています。

原文 (English)

SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with users in socially sensitive contexts. Existing evaluations typically measure sycophancy under a fixed prompt formulation, leaving unclear whether such behavior is stable when the same underlying situation is presented with different sycophancy-relevant prompt variants. In this work, we study sycophancy prompt sensitivity: the extent to which changes in user confidence, emotional framing, social consensus, or validation-seeking language alter a model's sycophantic behavior. We refer to our evaluation framework as SyPS, short for Sycophancy Prompt Sensitivity. Building on existing social sycophancy evaluation settings, SyPS constructs controlled prompt variants that preserve the same underlying user situation while varying sycophancy-relevant social cues. We introduce the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level measure of sycophancy variation across paired prompt variants. Unlike aggregate sycophancy rates, SPSS separates baseline sycophancy from prompt-induced shifts, enabling model-level comparisons of robustness to sycophancy-relevant social cues. Empirically, we find that sycophancy prompt sensitivity is socially structured: validation-seeking and emotional-pressure cues often increase sycophancy, whereas counter-framing and anti-sycophancy prompts tend to reduce it. Our framework highlights whether LLMs maintain stable social judgments while adapting appropriately in tone.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

予算に制約のあるエージェント検索で、より多くの活用とよりスマートな探索を実現

予算に制約のあるエージェント検索は、検証に費用がかかる、生成に複数のモデル呼び出しが必要、またはその両方が原因で、LLM エージェントが少ない評価予算の下で候補を絞り込む必要がある場合に発生します。この体制では、標準の MCTS は予算を適切に割り当てません。訪問回数が少ない場合は探索ボーナスが優先され、有望なチェーンが深化する前に有望でない兄弟が拡張され、分岐はノードの品質に依存しません。 ExTS は、拡張自体を情報価値の決定として扱うツリー検索ポリシーです。 ExTS は、狭いスコア分布の下で候補者を分離するための差別的報酬形成、親の報酬履歴から新しいブランチを作成する価値を推定する確率的仮想子、およびノー​​ドのスコアが予算コストを正当化する場合にのみ拡張する品質条件付きブランチの 3 つのメカニズムを組み合わせています。 ExTS は、プロンプト最適化、コード生成、分子構造解明、およびエージェント ワークフロー最適化全体にわたって、タスク固有のツリー検索ベースラインと競合するか、それを改善し、単一の固定構成を使用した場合の平均相対利得 +5.5% を実現します。さらに、予算に制約のあるエージェント検索の問題が構造的に互いに異なる原因を特徴付けるパイロット実行診断を導入し、問題空間の理解と適応のための実践的なガイダンスの両方を提供します。

原文 (English)

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation budget, because validation is expensive, generation requires multiple model calls, or both. In this regime, standard MCTS allocates budget poorly: exploration bonuses dominate at low visit counts, unpromising siblings are expanded before promising chains can deepen, and branching is independent of node quality. We introduce ExTS, a tree-search policy that treats expansion itself as a value-of-information decision. ExTS combines three mechanisms: discriminative reward shaping to separate candidates under narrow score distributions, a stochastic virtual child that estimates the value of creating a new branch from the parent's reward history, and quality-conditioned branching that expands only when a node's score justifies the budget cost. Across prompt optimization, code generation, molecular structure elucidation, and agentic workflow optimization, ExTS is competitive with or improves over task-specific tree-search baselines, with an average relative gain of +5.5% using a single fixed configuration. We further introduce pilot-run diagnostics that characterize what makes budget-constrained agentic search problems structurally different from one another, providing both understanding of the problem space and practical guidance for adaptation.

13:00 JST研究/論文

時系列予測のためのインコンテキスト修復

私たちは、ラージ ビジョン モデル (LVM) の一般化能力を活用して、時系列予測をビジュアル修復タスクとして再構成する新しいフレームワークである ICI-Time を提案します。特殊な時間アーキテクチャや広範なドメイン固有のトレーニングを必要とする手法とは異なり、ICI-Time は時系列を構造化された視覚表現 (面グラフ) に変換し、視覚的なコンテキスト内学習を適用して、事前トレーニングされたビジョン トランスフォーマーが微調整やアーキテクチャの変更を行わずに解決できるグリッド構造のプロンプト内のパターン補完として予測を再定式化します。時間的依存関係は、数値領域と視覚領域間の一貫性のある反転可能なマッピングを使用して、空間レイアウトを通じて表現されます。疫学、気象学、電力システムにわたる広範な実験により、ICI-Time が深層学習ベースラインと競合してパフォーマンスを発揮し、限られたデータ設定の下で有望な適応性を示し、時間的領域と視覚的領域を橋渡しする新しいパラダイムを導入することが実証されました。

原文 (English)

In-Context Inpainting for Time Series Forecasting

We propose ICI-Time, a novel framework that reframes time series forecasting as a visual inpainting task, leveraging the generalisation power of large vision models (LVMs). Unlike methods that require specialised temporal architectures and extensive domain-specific training, ICI-Time transforms time series into structured visual representations (area charts) and applies visual in-context learning, reformulating forecasting as pattern completion within a grid-structured prompt that pre-trained vision transformers can solve without fine-tuning or architectural modification. Temporal dependencies are represented through spatial layout, with a consistent, invertible mapping between numerical and visual domains. Extensive experiments across epidemiology, meteorology, and power systems demonstrate that ICI-Time performs competitively against deep learning baselines and shows promising adaptability under limited-data settings, introducing a new paradigm that bridges temporal and visual domains.

13:00 JST研究/論文

Granite.Trust ポリシー ツール: 生成 AI アプリケーション向けの共有可能で実行可能なポリシー

生成 AI の安全ポリシーに関しては、1 つのサイズですべてに適合するわけではありません。各組織とユースケースは、アプリケーションのコンテキスト、規制環境、組織の価値観、ユーザーのペルソナに応じて、さまざまなリスクを軽減する必要があります。しかし、既存のポリシー仕様アプローチは従来のアクセス制御向けに設計されており、GenAI アプリケーションのニュアンス、つまりコンテンツベースの制約の強制を捉えることができません。このギャップに対処するための 2 つの貢献を紹介します。(1) アクション可能なポリシー スキーマ。モデル応答に含めることができるものと含めることができないものを指定するための YAML ベースの形式です。このスキーマにより、例外ベースのポリシー ガバナンスが可能になり、ポリシー違反を追跡するための例外が提案されます。 (2) モデルの調整とテストのためのポリシーに合わせたトレーニング データを生成する合成データ生成パイプラインと、スキーマの定義とポリシーの適用を支援する一連のツール。これらを組み合わせることで、組織はポリシーを一度指定すれば、モデルの調整から実行時の監視まで、GenAI アプリケーションのライフサイクル全体にわたってポリシーを適用できるようになります。アクション可能なポリシーのスキーマ、ポリシーの例、およびツールは、オープンソースとして入手できます: https://github.com/ibm-granite/granite.trust.policy-tools 新しいアイデア、貢献、フィードバックを歓迎します。

原文 (English)

Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications

When it comes to safety policies for generative AI, one size does not fit all. Each organization and use case needs to mitigate different risks depending on the application context, regulatory environment, organizational values, and user personas. Yet, existing policy specification approaches are designed for traditional access control and fail to capture the nuances of GenAI application: the enforcement of content-based constraints. We present two contributions to address this gap: (1) the Actionable Policy schema, a YAML-based format for specifying what model responses can and cannot contain. The schema enables exception-based policy governance, proposing exceptions to track policy violations; (2) synthetic data generation pipeline that produces policy-aligned training data for model alignment and testing, and a set of tools to help define the schema and enforce policy. Together, these enable organizations to specify policies once and enforce them throughout the GenAI application lifecycle: from model alignment to runtime monitoring. The Actionable Policy schema, example policies, and tools are available as open source: https://github.com/ibm-granite/granite.trust.policy-tools We welcome new ideas, contributions and feedback.

13:00 JSTLLM/生成AIハードウェア/半導体

セマンティック オーバーレイ: トークンとステアリング ベクトルを超えたアノテーションによる即時インジェクションの軽減

言語モデルが認識するものはすべてトークンです。サービング スタックは、各スパン (ユーザー入力、ツール出力、指示) が何であるかを知っていますが、モデルはそれ自体を追跡する必要があり、追跡できなくなったり、混乱したりする可能性があります。テキストは、他のものと同様に読み取るために作成できます。即時注入は、この現象を自然に利用したものです。モデルのスパン ID の理解をスクランブルすることにより、攻撃者は望ましくない、潜在的に危険なアクションを誘発する可能性があります。モデルの入力に非テキスト チャネルを追加する (テキストを超えてスパン ID を伝達する方法) と、この種の攻撃が軽減されます。そこで、セマンティック オーバーレイと呼ばれる一般的なステアリング手法を導入します。これは、選択されたプリフィル位置でフリーズされたモデルの残差ストリームに適用される、学習された小さなアダプターです。スパン上にオーバーレイを配置すると、トークンによって複製できない帯域外アノテーション チャネルが作成されます。ステアリング ベクトルとは異なり、セマンティック オーバーレイはトレーニングされ、適応可能で、選択的に適用されます。オーバーレイは、マークされたスパンをモデルが認識する方法を再形成する複雑なセマンティクスをエンコードできます。オーバーレイの下で、実際とは異なるプログラミング言語であることを主張するコード スニペットをコピーするように要求されると、モデルは、そのコード スニペットを、アサートされた言語で忠実に書き換えます。オーバーレイは構成可能でもあり、基礎となるコンテンツの透過的な読み取りを可能にし、モデルが従う命令を含む複雑なペイロードを運ぶことができます。スパンを「実行不可」としてマークするオーバーレイは、信頼できないコンテキストに命令を追加する広範な種類のプロンプト インジェクションから防御します。私たちはプロンプトインジェクションベンチマークで強力な結果を報告しています。ユーティリティは変更されずに SEP 分離が 24.3% から 96.5% に上昇し (スコアリングルール。公開されているグレーダーの欠陥も修正しました)、TensorTrust 攻撃の成功率は 34.8% から 6.6% に低下し、4 つの PIArena 攻撃ファミリーはすべてコンプライアンス 0% に低下しましたが、マークされたスパンは読み取り可能な状態を維持しています (正確なコピー率は 92.5%)。

原文 (English)

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and it can lose track or be confused: text can be written to read like anything. Prompt injection is a natural exploit of this phenomenon. By scrambling the model's understanding of span identity, an attacker can induce unwanted and potentially dangerous actions. Adding a non-textual channel to the model's input -- a way to communicate span identity beyond text -- mitigates this class of attack. We thus introduce a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream. Laying an overlay over a span creates an out-of-band annotation channel that cannot be replicated by tokens. Unlike steering vectors, Semantic Overlays are trained, adaptable, and selectively applied. An overlay can encode complex semantics that reshape how the model perceives the marked span: asked to copy a code snippet under an overlay asserting that it is in a different programming language than it is, the model rewrites the snippet, faithfully, in the asserted language. Overlays are also composable, allow for transparent reading of underlying content, and can carry complex payloads -- including imperatives that the model will follow. An overlay which marks a span as "non-executable" defends against the broad class of prompt injections that add instructions in untrusted context. We report strong results on prompt injection benchmarks: SEP separation rises from 24.3% to 96.5% with utility unchanged (our scoring rule; we also correct a defect in the published grader), TensorTrust attack success rate falls from 34.8% to 6.6%, and all four PIArena attack families drop to 0% compliance, all while marked spans stay readable (92.5% exact copy rate).

13:00 JST研究/論文

AI が道を見つける

人工知能 (AI) アルゴリズムは、創造的で予期せぬ解決策を頻繁に学習し、開発および研究している専門研究者さえも驚かせます。彼らは、予期せぬ行動を発見したり、報酬シグナルの抜け穴を利用したり、以前は知られていなかった科学現象を自発的に発見したりすることで、実践者を驚かせることがよくあります。ただし、機械学習全体にわたるこのような型破りな動作の説明が正式に文書化されることはほとんどありません。この作品では、100 人を超える研究者の研究を代表するさまざまな機械学習サブ分野から厳選された 26 の直接の逸話を紹介します。これらの逸話は、人間が課した設計制限を回避し、訓練するタスクに対する予期せぬ解決策を発見する現代の AI システムの能力を示しています。さらに、これらのアカウントは、将来の AI システムの安全性にとって特に重要です。これらは、創造性を損なうことなくモデルを人間の価値観に合わせるという基本的な課題を示しているため、有害な可能性のある驚くべき結果を生み出すことなく、驚くべき発見を行うことができます。この論文ではまず、多くの困難な領域にわたって強化学習を通じて超人的な成功を収めた AI について詳しく説明します。ただし、報酬駆動型の最適化は、モデルが報酬の指定が不十分であったり、明確に表現されていない制約をハッキングすることを学習すると失敗する可能性があります。次に、インターネット規模の基盤モデル (FM) の利用がこれらの根本的な課題を解決するわけではなく、実際にはそれらの課題をさらに強化する可能性があることを示唆するケーススタディを紹介します。それにもかかわらず、私たちは、これらと同じ学習ダイナミクスを利用して科学的発見を加速できると主張します。最後に、この研究が将来の研究に情報を提供するための統合リソースを提供し、現代の AI では予期せぬ動作をする傾向が一般的であることを実証し、革新的ではあるが予測不可能なソリューションに対する AI の能力を予測して管理する必要性を強調したいと考えています。 (要約)

原文 (English)

AI Finds A Way

Artificial Intelligence (AI) algorithms frequently learn creative and unexpected solutions, surprising even expert researchers who develop and study them. They often astonish practitioners by discovering unanticipated behavior, exploiting loopholes in reward signals, or spontaneously uncovering previously unknown scientific phenomena. However, accounts of such unconventional behavior across machine learning are seldom formally documented. This work presents 26 curated firsthand anecdotes from various machine learning subfields representing the work of over 100 researchers. These anecdotes showcase the capability of modern AI systems to circumvent human-imposed design limitations and discover unexpected solutions to the tasks we train them on. Furthermore, these accounts are particularly important for the safety of future AI systems. They illustrate the fundamental challenge of aligning models with human values without diminishing their creativity, so they can make surprising discoveries without producing surprising, potentially harmful outcomes. The paper first details AI achieving superhuman success through reinforcement learning across many challenging domains. However, reward-driven optimization can fail when the model learns to hack an underspecified reward or unarticulated constraint. We then present case studies suggesting that harnessing internet-scale foundation models (FMs) has not resolved these fundamental challenges and, in fact, can supercharge them. Nevertheless, we argue that these same learning dynamics can be harnessed to accelerate scientific discovery. Finally, we hope this work provides a consolidated resource to inform future research and demonstrates that the tendency toward unexpected behaviors is commonplace in modern AI, highlighting the need to anticipate and manage AI's capacity for innovative, yet unpredictable, solutions. (abstract abridged)

13:00 JST研究/論文

進化する概念定義のもとでの来歴に基づく漸進的学習

長期間にわたって展開される学習システムは、受信データの統計的変化だけでなく、予測ターゲットを生成する定義の改訂にも適応する必要があります。従来のコンセプト ドリフト手法は通常、基礎となるポリシー、ルール、クエリが明示的に変更されている場合でも、観察や予測誤差からそのような変化を推測します。この論文では、ターゲットを定義する概念が直接修正され、観察されたデータを変更することなく、以前に保存されたインスタンスが異なる意味ラベルを取得する、ルールによって引き起こされる概念シフトについて研究します。私たちは、連続する概念定義を構造化されたルール デルタにコンパイルし、歴史的な来歴を通じて変更されたコンポーネントを追跡し、以前のラベルが有効なままであるレコードを証明し、再評価をローカライズされた候補領域に制限する、来歴ガイド型の増分学習フレームワークを導入します。実行可能リビジョンは自動的に再ラベル付けされ、曖昧なケースは選択的な監視を通じて処理され、結果として生じる変更は増分予測子の修復に使用されます。バージョン管理された概念メモリは、繰り返し定義をさらにサポートします。また、財務、人口統計、サイバーセキュリティ、およびしきい値、述語、論理、リレーショナル、繰り返し、および混合概念の改訂を備えたグラフ構造データを網羅する RuleShift-Bench も紹介します。ベンチマーク全体で、来歴に基づく修復は 92.3% の精度と 90.2% の Macro-F1 を達成しながら、履歴コレクションの 14.7% を再処理し、影響を受けたレコードの 94.6% を保持します。平均更新レイテンシは 179 秒ですが、完全な再ラベル付けと再トレーニングでは 993 秒です。この結果は、明示的な概念の改訂をデータ保守シグナルとして利用できることを示しており、学習システムが有効な知識を維持しながら、変更に依存する監視と予測状態を更新できるようになります。

原文 (English)

Provenance Guided Incremental Learning Under Evolving Concept Definitions

Learning systems deployed over long periods must adapt not only to statistical changes in incoming data, but also to revisions of the definitions that generate their prediction targets. Conventional concept-drift methods typically infer such changes from observations or prediction errors, even when the underlying policy, rule, or query has been explicitly modified. This paper studies rule-induced concept shift, where the target-defining concept is revised directly, causing previously stored instances to acquire different semantic labels without requiring any change in their observed data. We introduce a provenance-guided incremental learning framework that compiles consecutive concept definitions into a structured rule delta, traces the changed components through historical provenance, certifies records whose previous labels remain valid, and restricts reevaluation to a localized candidate region. Executable revisions are relabeled automatically, ambiguous cases are handled through selective supervision, and the resulting changes are used for incremental predictor repair. A versioned concept memory further supports recurring definitions. We also introduce RuleShift-Bench, spanning financial, demographic, cybersecurity, and graph-structured data with threshold, predicate, logical, relational, recurring, and mixed concept revisions. Across the benchmark, provenance-guided repair attains 92.3% accuracy and 90.2% Macro-F1 while reprocessing 14.7% of the historical collection and retaining 94.6% of affected records. Its average update latency is 179s compared with 993s for complete relabeling and retraining. The results demonstrate that an explicit concept revision can be exploited as a data-maintenance signal, allowing learning systems to update the supervision and predictive state that depend on the change while preserving knowledge that remains valid.

13:00 JST研究/論文Claude

BenchBench プロトコル: 現実世界のウェットラボ プロトコルの推論と修正の評価

BenchBench-Protocol は、科学者が実際の実験作業中に公開されたプロトコルに加えた変更から復元された 149 のプロトコル変更タスクの大規模言語モデルのベンチマークです。公開されたプロトコルを新しい実験に適応させることは、ウェットラボの科学者にとって日常的な作業であり、正しく変更するには、事前の選択と下流のステップを考慮する必要があります。最近のライフ サイエンスのベンチマークは、オープンエンドのルーブリック評価タスクに移行していますが、タスクは通常、現実世界の変更から再構築されるのではなく、専門家から引き出されます。 BenchBench-Protocol タスクは、公開されたプロトコルと科学者が修正したバージョンとの違いから派生し、クエリの基礎と正しい応答のための重み付けされたルーブリック要素を提供します。このベンチマークは、ウェットラボ生物学の 9 つのドメインにわたる 96 のソースプロトコルに基づいており、ドメインの専門家によるレビュー後に高く評価されたタスクのみが含まれています。 9 つのクローズド モデルとオープン モデルを評価します。 Claude Opus 5 の正規化ルーブリック スコアは 59.2% で最も高く、他のモデルは 34.1% から 47.1% の間であり、10 回の試行のベストをとってもベンチマークは飽和していません。ライフサイエンス研究においてモデルがますます役立つようになるにつれて、日常的なウェットラボ作業でモデルを評価することがそれに応じて重要になります。我々は、ウェットラボ推論の根拠に基づいた評価と、ベンチマーク タスクを構築するための実世界実験の有用性の証拠の両方として、BenchBench-Protocol を提示します。

原文 (English)

BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recovered from modifications that scientists made to published protocols during real experimental work. Adapting a published protocol to a new experiment is a routine task for a wet-lab scientist, and a correct modification requires accounting for prior choices and downstream steps. Recent life-science benchmarks have moved toward open-ended, rubric-graded tasks, but tasks are typically elicited from experts rather than reconstructed from real-world modifications. BenchBench-Protocol tasks are derived from differences between a published protocol and a version a scientist modified, which provides the basis for the query and the weighted rubric elements for a correct response. The benchmark draws from 96 source protocols across nine domains of wet-lab biology and only includes tasks rated highly after review by domain experts. We evaluate nine closed and open models; Claude Opus 5 scores highest at 59.2% normalized rubric score, with other models between 34.1% and 47.1%, and the benchmark remains unsaturated when taking the best of ten attempts. As models are increasingly helpful in life-sciences research, evaluating them on routine wet-lab tasks becomes correspondingly important. We present BenchBench-Protocol as both a grounded assessment of wet-lab reasoning and evidence for the utility of real-world experiments to construct benchmark tasks.

13:00 JST研究/論文

複雑な社会技術システムにおける AI 導入によるシステムレベルの害悪の定量化

人工知能 (AI) は、国家重要インフラ (CNI) を含む複雑な社会技術システムにますます統合されており、技術的要素、人間的要素、組織的要素の間の相互作用から害が生じます。しかし、現在の AI 評価は依然としてモデル中心であり、観察された動作がシステムレベルのリスクにどのように変換されるかについての洞察はほとんど提供されていません。このギャップを埋めるために、構造化ハザード分析、コンポーネントレベルのテスト、確率的システムモデリングをリンクするフレームワークを提案します。このフレームワークは、モデルの動作からシステムレベルの結果まで追跡可能な経路を提供することにより、実務者が「だから何?」に答えることができるようにします。 AI の障害を分析し、そのシステムへの影響を定量化し、複雑なシステムにおける AI の証拠に基づいた予測的なガバナンスに移行します。実例として英国のリアルタイムグロス決済(RTGS)システムに適用して、システム理論プロセス分析(STPA)を使用してAI主導の損失シナリオを導き出し、そのような損失シナリオの1つとしてLLMベースの取引の敵対的操作を調査します。コンポーネントレベルの実験では、単純な敵対的な入力が、AI の推奨事項に従った場合に測定可能な行動の変化を引き起こすことが示されています。ここで金融伝染モデルに使用されているコンポーネントからシステムへのマッピングの下で​​は、これらの変化によりシステムの回復力が変化し、銀行破綻が増加し、特に広範または独占的な AI の導入下ではショックが連鎖的混乱につながる閾値が低下します。

原文 (English)

Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems

Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including Critical National Infrastructure (CNI), where harms emerge from interactions between technical, human, and organisational elements. Yet current AI evaluation remains model-centric, offering little insight into how observed behaviours might translate into system-level risk. We propose a framework that links structured hazard analysis, component-level testing, and probabilistic system modelling to bridge this gap. By providing a traceable pathway from model behaviour to system-level outcomes, the framework enables practitioners to answer the "so what?" of AI failures, quantify their systemic impact, and move toward evidence-based and anticipatory governance of AI in complex systems. Applied to the UK's Real Time Gross Settlement (RTGS) system as an illustrative worked example, we derive AI-driven loss scenarios using Systems Theoretic Process Analysis (STPA) and examine adversarial manipulation of LLM-based trading as one such loss scenario. Component-level experiments show that simple adversarial inputs induce measurable behavioural shifts where AI recommendations are followed. Under the component-to-system mapping used here for a financial contagion model, these shifts alter system resilience, increasing bank failures and lowering the threshold at which shocks lead to cascading disruption, particularly under widespread or monopolistic AI adoption.

13:00 JSTエージェント

マルチエージェント財務アドバイザリーにおける検索拡張生成と決定論的な税金計算: 2x2 階乗実験

欠損金の回収は、ポートフォリオの長期的な成長に一貫したメリットをもたらします。しかし、それを効率的に実装するには、そのポートフォリオ内の保有資産やそれを所有する個人に特有の複雑な考慮事項が必要になることがよくあります。マルチエージェント取引推奨システムのコンテキストを提供するために、カスタム キャピタル ゲイン計算エンジンと、RAG が取得した市場勧告レポートのベクトル ストアを導入します。私たちは、ポートフォリオ清算中に発生した相対的なキャピタルゲインによって測定された、レコメンデーションの品質に対する各コンテキストプロバイダーの影響を調査します。 2x2 の反復測定 ANOVA により、税金最適化エンジンの重要な主効果 ($F(1,29) = 9.17$、$p = 0.005$、$\eta^2_p = 0.240$) が明らかになりました。エンジンを有効にすると、エンジンなしの条件と比較して節税が約 55 パーセント ポイント減少しました。 RAG の主効果は有意ではなく ($p = 0.841$)、交互作用も有意ではありませんでした ($p = 0.553$)。 RAG のみの条件は最高の記述平均節税 (47.7%) を達成し、ベースライン条件は 2 番目に優れたパフォーマンス (30.6%) を達成しました。これは、明示的なツールを使用せずに、事前トレーニングされた言語モデルの内部化された財務知識が、有能な欠損金回収の推奨に十分である可能性があることを示唆しています。これらの結果は、ドメイン固有の計算エンジンで LLM エージェントを強化してもパフォーマンスの向上が保証されず、競合する最適化信号が発生する可能性があることを示しています。

原文 (English)

Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment

Tax-loss harvesting demonstrates consistent benefits to long-term portfolio growth; yet implementing it efficiently often involves complex considerations that are specific to the holdings within that portfolio and the individual who owns it. We introduce a custom capital gains calculation engine and a RAG-retrieved vector store of market advisory reports to provide context for a multi-agent trade recommendation system. We investigate the effects of each context provider on the quality of recommendations, measured by relative capital gains incurred during portfolio liquidation. A 2x2 repeated-measures ANOVA revealed a significant main effect of the tax optimization engine ($F(1,29) = 9.17$, $p = .005$, $\eta^2_p = .240$): enabling the engine reduced tax savings by approximately 55 percentage points relative to the no-engine conditions. The RAG main effect was not significant ($p = .841$), nor was the interaction ($p = .553$). The RAG-only condition achieved the highest descriptive mean tax savings (47.7%), and the baseline condition performed second-best (30.6%), suggesting that the pre-trained language model's internalized financial knowledge may be sufficient for competent tax-loss harvesting recommendations without explicit tooling. These results indicate that augmenting LLM agents with domain-specific computation engines does not guarantee improved performance and may introduce conflicting optimization signals.

13:00 JST研究/論文Gemma

PROOF-Gen: 最適化されたデータからより優れた蒸留まで

教師が生成した軌道に対する監視付き微調整は、ツール呼び出し機能を展開可能なモデルに抽出するための標準的な最初の段階です。出荷されたツール呼び出しエージェントを駆動するトレーニング後のパイプラインは、このステージを日次または週次のペースで再実行し、サイクルごとにフロンティア教師にコストを支払いますが、そのメカニズムは生成とフィルタリング (教師の通過軌跡を維持し、残りは破棄) であり、障害は信号を提供しないため、各サイクルは同じ困難なシナリオを残します。 {\tau}2 ベンチでは、教師によるトライアルの 57% が失敗し、そのうち 3 分の 2 がニアミスです (ほとんどのツール呼び出しは正しく、1 つの決定的なエラーによって元に戻ります)。 PROOF-Gen (FailedGeneration を克服するためのシナリオごとの反射的最適化) を導入します。これは、シナリオごとのプロンプト最適化によってこれらの失敗から黄金の軌道を回復します。失敗したタスクごとに、リフレクターが実行トレースと評価フィードバックを分析し、教師を合格軌道に導く修正ガイダンスを作成します。指導はトレーニング前に取り除かれるため、学生はタスク固有の足場のないクリーンなデモンストレーションから学びます。 {\tau}2 ベンチでは、シナリオごとの最適化により、失敗したシナリオの 93% が回復します。結合データに基づいて微調整された Qwen3-4B-Instruct-2507 は Pass^1=0.132 から 0.529 に改善し、Gemma 4 E4B は BFCL v4 マルチターンで +7.2pp 増加しました。デプロイされたパイプラインでは、このメソッドにより軌道の品質が +6.3 pp のゴール完了率で向上し、デプロイされたオンデバイス モデルに転送されます (+1.5 pp のゴール完了、応答品質メトリクス全体で +1.7 ~ +5.0 pp)。すべてのロケールでプラスの転送が行われます (英語以外の平均 +1.48 pp)。

原文 (English)

PROOF-Gen: From Optimized Data to Better Distillation

Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher's passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On {\tau}2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone by one decisive error). We introduce PROOF-Gen (Per-scenario Reflective Optimization to Overcome FailedGeneration), which recovers golden trajectories from these failures via per-scenario prompt optimization. For each failed task, a reflector analyzes the execution trace and evaluation feedback, then writes corrective guidance that steers the teacher to a passing trajectory. The guidance is stripped before training, so the student learns from clean demonstrations with no task-specific scaffold. On {\tau}2-bench, per-scenario optimization recovers 93% of failed scenarios. Fine-tuned on the combined data, Qwen3-4B-Instruct-2507 improves from Pass^1=0.132 to 0.529 and Gemma 4 E4B-it gains +7.2pp on BFCL v4 multi-turn. In a deployed pipeline, the method lifts trajectory quality by +6.3pp goal completion and transfers to a deployed on-device model (+1.5pp goal completion; +1.7 to +5.0pp across response-quality metrics), with positive transfer in every locale (non-English average +1.48pp).

13:00 JSTLLM/生成AIエージェントGemma

MARS: 競技プログラミング用のマルチスペシャリスト LLM リレー システム

大規模言語モデルはコード生成に優れていますが、競合プログラミングは永続的な障害モードにさらされます。既存のマルチエージェント パイプラインは、汎用のプランナー、コーダー、デバッガーの役割に作業を分散させ、アルゴリズム手法の選択をバックボーンのみに委任します。我々は、アルゴリズム理論コーパスに対する検索拡張生成に基づく、各エージェントが動的プログラミング、グラフ、文字列、幾何学などのトピックのスペシャリストであるプロンプト専用フレームワークである MARS (Multi-Agent Relay of Specialized LLMs) を紹介します。問題が与えられると、検索では関連する専門家からなる小規模なチームが選択されます。スターターは最初の C++17 ソリューションを作成し、その後の各ターンでサンドボックス内の公開サンプルに対して候補を実行し、アクティブなスペシャリストにドラフトを保管、修復、または引き渡して、構造化されたパケットを次のスペシャリストに転送します。単一のインフラストラクチャ修正パスにより、最後にボイラープレートが正規化されます。 Gemma 4 を使用した CodeContests テスト分割では、MARS はタスクあたり $2.3$ 記録されたパイプライン ステージで $0.624 \pm 0.006$ の合格率に達し (直接プロンプトに対して $+14.4$ パーセンテージ ポイント)、実時間コストが $3.3{\times}$ 低く、タスクごとのトークン支出の差異が大幅に小さいため、CodeSIM ($0.731$) との差をほとんど縮めています。ソース コードは GitHub: https://github.com/fckand/mars で入手できます。

原文 (English)

MARS: Multi-Specialist LLM Relay System for Competitive Programming

Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribute work over generic planner, coder, and debugger roles and delegate the choice of algorithmic technique to the backbone alone. We present MARS (Multi-Agent Relay of Specialized LLMs), a prompt-only framework in which each agent is a topic specialist---dynamic programming, graphs, strings, geometry, and so on---grounded by retrieval-augmented generation over an algorithm-theory corpus. Given a problem, retrieval selects a small team of relevant specialists; a starter writes an initial C++17 solution, and each subsequent turn runs the candidate against public examples in a sandbox, lets the active specialist keep, repair, or hand off the draft, and forwards a structured packet to the next specialist. A single infrastructure-fixer pass normalizes boilerplate at the end. On the CodeContests test split with Gemma 4, MARS reaches $0.624 \pm 0.006$ pass rate at $2.3$ recorded pipeline stages per task ($+14.4$ percentage points over direct prompting), closing most of the gap to CodeSIM ($0.731$) at $3.3{\times}$ lower wall-clock cost and substantially smaller variance in per-task token spend. The source code is available on GitHub: https://github.com/fckand/mars.

13:00 JST研究/論文

混合実験としてのデータ混合: 大規模言語モデルの事前学習のための応答曲面法と最適設計

データ混合は、大規模な言語モデルの事前トレーニングにおける中心的な設計上の問題です。固定のトークン予算が与えられると、実践者は各ドメインにどれだけのデータを割り当てるかを決定する必要があります。最近のプロキシベースの手法は、候補混合物で小さなモデルをトレーニングし、応答モデルをフィッティングし、その応答を使用して大規模なトレーニング用の混合物を選択することで、この問題に対処しています。このワークフローが古典的な混合実験の構造を持っていることを示します。このビューでは、データ ドメインは混合成分、トークン シェアは成分比率、プロキシ トレーニングの実行は実験計画点、検証損失は確率単体上の応答曲面を定義します。私たちは、疎な二次シェフ応答曲面モデルを使用してこの定式化を開発し、プロキシ データ混合実験のためのモデル堅牢な $\mathcal{I}$ 最適設計を構築します。 RegMix を実証的なケーススタディとして使用して、フレームワークが観察された混合物の応答を解釈し、より効率的な代理実験を設計する方法を実証します。シェフの分析は、ドメイン値が強い関係性であることを示しています。相加効果の下では弱いいくつかのドメインは、ペアごとの相互作用、特にウェブ由来のテキストとの組み合わせを通じて有利になります。スパース シェフ モデルは、モデル スケール全体で混合物のランキングを保持し、相加効果と相互作用効果の明示的な分解を提供しながら、柔軟な機械学習予測子との競争力を維持します。観察されたプロキシ トレーニング応答に合わせて調整されたシミュレーション研究では、モデル堅牢な $\mathcal{I}$ 最適設計は、元のプロキシ実行の約 25\% を除去した後、関連する混合順序を回復します。これらの結果は、LLM データ混合を予測問題としてだけでなく、統計効率を向上させるために代理混合自体を選択できる実験計画問題としても扱う必要があることを示唆しています。

原文 (English)

Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

Data mixing is a central design problem in large language model pretraining: given a fixed token budget, practitioners must decide how much data to allocate to each domain. Recent proxy-based methods address this problem by training small models on candidate mixtures, fitting a response model, and using the response to select mixtures for larger-scale training. We show that this workflow has the structure of a classical mixture experiment. Under this view, data domains are mixture components, token shares are component proportions, proxy-training runs are experimental design points, and validation loss defines a response surface over the probability simplex. We develop this formulation using sparse second-order Scheff\'{e} response-surface models and construct model-robust $\mathcal{I}$-optimal designs for proxy data-mixing experiments. Using RegMix as an empirical case study, we demonstrate how the framework can both interpret observed mixture responses and design more efficient proxy experiments. The Scheff\'{e} analysis shows that domain value is strongly relational: several domains that are weak under additive effects become favourable through pairwise interactions, especially through combinations with web-derived text. The sparse Scheff\'{e} model preserves mixture rankings across model scales and remains competitive with a flexible machine-learning predictor while providing an explicit decomposition of additive and interaction effects. In a simulation study calibrated to observed proxy-training responses, model-robust $\mathcal{I}$-optimal designs recover the relevant mixture ordering after removing about 25\% of the original proxy runs. These results suggest that LLM data mixing should be treated not only as a prediction problem, but also as an experimental-design problem in which the proxy mixtures themselves can be chosen to improve statistical efficiency.

13:00 JST研究/論文

適応行動と不適応行動の発達における進化的反復的意思決定モデル

この研究では、進化的不一致、限界合理性、満足が適応的行動と非適応的行動にどのように寄与するかを調査するために設計された計算強化学習フレームワークである進化的再帰的意思決定モデル (ERDM) を紹介します。 ERDM は、脅威、獲物/目標の追求、同盟など、進化の繰り返し環境にわたるエージェントをシミュレートします。エージェントは、生存指標から抽象化された報酬を競い合うことで学習します。さまざまな幼少期の不利な経験をもとにした妥当性研究では、学習性無力感、回避、健全な人間関係、攻撃性など、明確な適応戦略と非適応戦略が、組み込まれたものではなく自然に現れることが実証されています。これらの結果は経験的文献と一致しており、生態学的妥当性を示しています。この結果は、精神病理学に関連する多くの側面が、現代と祖先の環境の不一致の下で動作する限定された認知システムとして解釈される可能性があることを示唆しており、ERDMを他の研究に拡張できる重要な計算認知ツールとして位置づけています。

原文 (English)

Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors

This study introduces the evolutionarily recurrent decision model (ERDM), a computational reinforcement learning framework designed to examine how evolutionary mismatch, bounded rationality, and satisficing contribute to adaptive and maladaptive behavior. ERDM simulates agents across evolutionary recurrent environments, including threat, prey/goal-pursuits, and alliances. Agents learn through competing rewards abstracted from survival metrics. A validity study under varying adverse childhood experiences demonstrates that distinct adaptive and maladaptive strategies, such as learned helplessness, avoidance, healthy relationships, and aggression, emerge naturally without being hardwired. These results align with empirical literature, showcasing ecological validity. The results suggest that many psychopathology-relevant aspects may be interpreted as bounded cognitive systems operating under modern-ancestral environmental mismatch, positioning ERDM as a key computational cognitive tool that can be extended to other studies.

13:00 JSTLLM/生成AI

差別的ではなく、より拒絶的: 実行前LLM監視における検証の単位

実行前の監視は、AI 制御における信頼できる監視の中核です。誤りのある LLM 監視は、元に戻せない実行の前に、計画されたアクションを精査します。過剰なブロッキングは有用性を失い、デプロイ担当者にそれを無効にするよう圧力をかけます。すべてのプロトコルは検証の単位、つまりレビューを呼び出すアクションの数を修正する必要があります。既存の設計では、単位が与えられたものとして扱われます。誤ったモニターに対する影響は測定されていません。自然なトレースではそれを分離することはできません。レビューの長さはエラーのタイプと位置によって変化します。単独で捕まえると誤解が生じます。すべてを拒否するとすべてが捕らえられます。これを測定するには、境界変動のみと、それに適合するクリーンなコントロールが必要です。両方を提供するツイン プレフィックス フレームワークを導入します。各ゴールド プランは、1 つの挿入された環境で受け入れられるエラーと、1 回の書き込みで異なるクリーン ツインを含むプレフィックスを生成します。 5 つのネストされた長さで各ペアを判定すると、判定の変更がユニットのみに関連付けられます。差別は、事前に登録された情報、キャッチから誤った拒否を引いたものによってスコア化されます。レビューが長くなるとキャッチが増加します。本人拒否は足並みをそろえて上昇します。両方の領域の 6 人の裁判官全員にとって、情報提供は 1 つまたは 2 つのアクションでピークに達します。ウィンドウが長くなると、ゼロショット モニターは差別的になるのではなく、拒絶的になります。保留された観察を再生すると、主に観察の剥奪が失敗の原因であることがわかります。安全ケースにはユニットを記載し、クリーンシリーズを共同報告する必要があります。私たちのフレームワークは、この選択に対して最初に制御され、事前に登録された手段であり、キャッチを単独で読み取ることはありません。当社の調整されたショートユニットは、8 つのアクションレビューで最大 0.95 のインフォームドネスを回復し、テストされたラベルブラインドポリシーでは常にこれを上回るものはありません。

原文 (English)

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

Pre-execution oversight is core to trusted monitoring in AI control: a fallible LLM monitor vets planned actions before irreversible execution. Over-blocking forfeits usefulness and pressures deployers to disable it. Every protocol must fix a unit of verification: how many actions one call reviews. Existing designs take the unit as given; its effect on fallible monitors is unmeasured. Natural traces cannot isolate it: review length co-varies with error type and position. Catch alone misleads: rejecting everything catches everything. Measuring this needs boundary variation alone and a matched clean control. We introduce the twin-prefix framework, which supplies both. Each gold plan yields a prefix with one injected, environment-accepted error and a clean twin differing in one write. Judging each pair at five nested lengths ties verdict changes to the unit alone. Discrimination is scored by pre-registered informedness, catch minus false rejection. Longer review raises catch; false rejection climbs in lockstep. Informedness peaks at one or two actions for all six judges in both domains: longer windows make zero-shot monitors more rejective, not more discriminative. Replaying withheld observations traces the failure largely to observation deprivation. Safety cases should state the unit and co-report the clean series. Our framework is the first controlled, pre-registered instrument for this choice and never reads catch alone. Our calibrated short unit recovers up to 0.95 informedness over eight-action review, and no tested label-blind policy consistently beats it.

13:00 JSTエージェント

再帰的なエージェント推論

反復改良、分解、反復サンプリングなどのテスト時の推論手法は、多くの場合単独で評価されるため、モデル、ベンチマーク、評価パイプライン全体でそのゲインを比較することが困難になります。エージェントの推論トレースに対する再帰演算子として、これらのメソッドの統一されたビューを導入します。 GROW は、単一の推論パスを深めます。 PRUNE: 問題を分解して再構成します。 BRANCH は、代替推論パスをサンプリングし、その中から選択します。私たちは、同一のプロンプト、トークンバジェット、およびグレーディングコードを使用した共有ハーネスの下で、シングルパスの思考連鎖ベースラインに対して 3 人のオペレーターすべてを評価します。 5 つのベンチマークと 3 つのフロンティア モデル(14 のモデル ベンチマーク設定、49,327 のグレーディング項目、および 151,876 のモデル呼び出しで構成されます)全体で、BRANCH は 14 の設定すべてで精度を平均 5.98 パーセント向上させ、12 の設定で最もパフォーマンスの高いオペレーターです。対照的に、GROW は平均 2.18 ポイントの向上をもたらし、2 つの設定でパフォーマンスを低下させますが、PRUNE は精度を平均0.94点。分析の結果、BRANCH の利点は複数の推論パスの探索だけでなく、切り捨てからの回復からも生じることが示されています。その利点は、空の予算を使い果たした出力のベースライン レート (r = 0.72) と強く相関しています。これらの結果は、さまざまな問題がテスト時の推論演算子間のルーティングを必要とするという仮説を弱めます。この抽象化レベルでは、分岐の繰り返しが常に優勢です。最後に、ペアのない評価とスコアリング パイプラインの失敗をモデルのエラーとして扱うことにより、比較結論が大きく変わり、場合によっては逆転する可能性があり、テスト時計算評価の標準プロトコルとしてペアのスコアリングを動機付ける可能性があることを示します。

原文 (English)

Recursive Agentic Reasoning

Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are often evaluated in isolation, making their gains difficult to compare across models, benchmarks, and evaluation pipelines. We introduce a unified view of these methods as recursion operators over an agent's reasoning trace: GROW, which deepens a single reasoning path; PRUNE, which decomposes and recomposes the problem; and BRANCH, which samples alternative reasoning paths and selects among them. We evaluate all three operators against a single-pass chain-of-thought baseline under a shared harness with identical prompts, token budgets, and grading code. Across five benchmarks and three frontier models, comprising 14 model-benchmark settings, 49,327 graded items, and 151,876 model calls, BRANCH improves accuracy in all 14 settings by an average of 5.98 percentage points and is the best-performing operator in 12. In contrast, GROW yields a mean gain of 2.18 points and degrades performance in two settings, while PRUNE improves accuracy by 0.94 points on average. Analysis shows that BRANCH's advantage arises not only from exploring multiple reasoning paths, but also from recovering from truncation: its gains strongly correlate with the baseline rate of empty, budget-exhausted outputs (r = 0.72). These results weaken the hypothesis that different problems require routing among test-time reasoning operators; at this level of abstraction, repeated branching is consistently dominant. Finally, we show that unpaired evaluation and treating scoring-pipeline failures as model errors can materially change, and even reverse, comparative conclusions, motivating paired scoring as a standard protocol for test-time-compute evaluation.

13:00 JSTLLM/生成AIハードウェア/半導体Llama

GPU を増やすか、キャッシュを減らすか?メモリバウンド LLM サービスの Tensor 並列処理と KV 圧縮

LLM サービス展開で KVcache ルームが不足した場合、確立された方法が 2 つあります。 Tensor 並列処理により、重みと KV キャッシュが 2 台、4 台、または 8 台のデバイスに分散され、すべてのレイヤーでの全削減とデバイス数に応じて増加するハードウェア料金を代償としてメモリのヘッドルームが確保されます。アルゴリズム コミュニティは、KV 量子化とエビクションによって単一の GPU を維持し、代わりに品質を少し費やして、適切なキャッシュを縮小します。圧縮に関する論文ではメモリ比率が報告され、パラレル スケーリングに関する論文ではスループット曲線が報告されていますが、この 2 つを同じコスト軸に置く人はほとんどいません。 A100、A40、および H100 ハードウェアでキャリブレーションされたプロファイリングされたシミュレーターを使用して、テンソル並列構成 (次数 1 ~ 8) と KV 圧縮構成 (16/8/4 ビット、キープ率 0.25 まで) を 1 つのコスト正規化軸 (遅延に対する 100 万トークンあたりのコスト) に配置し、コストと等価性のクロスオーバーを探します。見つかりません。 2 つのモデル (7B および 70B の Llama-2)、3 つの GPU タイプ、および構築できるすべてのレベルのメモリ軽減にわたって、圧縮は 1.20 倍から 2.00 倍安くなります。 80 GB デバイス上の 7B モデルは、独自のコンテキスト ウィンドウ内で KV バジェットを使い果たすことができません。戦略間を決定する境界は、デバイス メモリに対するモデル サイズ (80 GB カードの場合は約 36B パラメータ) です。その壁の下では圧縮が優勢であり、追加の GPU はほとんど無駄な支出となります。それを超えると、テンソル並列処理は選択肢ではなくなり、エントリー チケットになります。バインディング リソースは重みであり、KV 圧縮が影響しないため、Llama-2-70B は、どの KV 設定でも 1 つの A100 では実行できません。テンソル並列処理はレイテンシーを改善する唯一の手段であり (圧縮によりバッチ競合によりトークンごとのレイテンシーが 8 ~ 93% 悪化します)、圧縮は 1 ドルあたりの容量を倍増させる唯一の手段です (GPU の 8 倍の支出の 1.21 倍に対して 16.5 倍)。

原文 (English)

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor parallelism shards the weights and the KV cache across two, four, or eight devices, buying memory headroom at the price of an all-reduce on every layer and a hardware bill that grows with the device count. The algorithms community shrinks the cache in place, with KV quantisation and eviction keeping a single GPU and spending a little quality instead. Compression papers report memory ratios, parallel-scaling papers report throughput curves, and almost nobody puts the two on the same cost axis. We place tensor-parallel configurations (degree 1 to 8) and KV-compressed configurations (16/8/4-bit, keep-ratios down to 0.25) on one costnormalised axis, cost per million tokens against latency, using a profiled simulator calibrated on A100, A40, and H100 hardware, and we go looking for the cost-equivalence crossover. We do not find one. Across two models (Llama-2 at 7B and 70B), three GPU types, and every level of memory relief we could construct, compression is cheaper by 1.20x to 2.00x. A 7B model on an 80 GB device cannot exhaust its KV budget within its own context window, and the boundary that decides between the strategies is model size relative to device memory, at roughly 36B parameters for an 80 GB card. Below that wall, compression dominates and extra GPUs are largely wasted spend; above it, tensor parallelism stops being a choice and becomes an entry ticket: Llama-2-70B is infeasible on one A100 at any KV setting, because the binding resource is weights, which KV compression does not touch. Tensor parallelism is the only lever that improves latency (compression makes per-token latency worse, by 8 to 93%, through batching contention), while compression is the only lever that multiplies capacity per dollar (16.5x, against 1.21x for an eightfold spend on GPUs).

13:00 JSTLLM/生成AI

Giraffe: 効率的なグラフィック デザインのための隠しテキスト表現から視覚的埋め込みまでのマッピング アーキテクチャ

マルチモーダル大言語モデル (MLLM) は、マルチメディア コンテンツの理解と解釈において大幅な進歩を遂げました。ただし、メディアを生成する能力は依然として制限されています。最近のアプローチでは、トークン シーケンスの隠された表現を視覚モデルの埋め込み空間に変換するか、生の画像データに直接変換することで、このギャップを埋めようとしています。ただし、これらの方法では、多くの場合、複数の特殊なトークンを使用して各画像を表現するため、入力長が大幅に増加します。これは、通常、出力にテキスト、複数の画像、レイアウト情報にわたる数千のトークンのシームレスなブレンドが含まれるグラフィック デザインの生成などのタスクにとって、大きな制限になります。この課題に対処するために、画像ごとに単一の [IMG] トークンを使用して、隠れたトークン表現を CLIP ViT-L/14 などのビジュアル モデルの埋め込み空間にマッピングする新しいアーキテクチャが提案されています。このアーキテクチャは 2 つの浅い MLP ブロックを採用しており、それぞれに個別の圧縮モジュールと、その後に続く共有拡張モジュールがあり、6 つの異なる損失関数でトレーニングされています。一方のブロックはトレーニング中にもう一方のブロックを支援し、推論中には省略されるため、軽量のソリューションが得られます。画像からデザインへの生成タスクとテキストからデザインへの生成タスクの両方で強力なパフォーマンスが実証されています。

原文 (English)

Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design

Multimodal large language models (MLLMs) have made significant progress in understanding and interpreting mul- timedia content. However, their ability to generate me- dia remains limited. Recent approaches have attempted to bridge this gap by translating the hidden representations of token sequences into the embedding space of visual models or directly into raw image data. However, these methods often represent each image using multiple specialised to- kens which significantly increases the input length. This be- comes a major limitation for tasks such as graphic design generation where the output typically involves a seamless blend of thousands of tokens across text, multiple images, and layout information. To address this challenge, a novel architecture is proposed that maps hidden token represen- tations to the embedding space of visual models, such as CLIP ViT-L/14, using a single [IMG] token per image. The architecture employs two shallow MLP blocks, each with a separate compression module followed by a shared expan- sion module, trained with six distinct loss functions. One block aids the other during training and is omitted during inference, resulting in a lightweight solution. Strong perfor- mance is demonstrated in both image-to-design and text-to- design generation tasks.

13:00 JST画像/動画生成研究/論文

見るだけでは十分ではない場合: LVLM におけるインタラクティブな視覚グラウンディングのベンチマーク

視覚的なグラウンディングは通常、情報を提供する表現から視覚的なターゲットへのワンショット マッピングとして評価されます。この定式化には、現実世界の参照の中心的な特性が欠けています。つまり、ターゲット情報は多くの場合、不完全で曖昧であり、相互作用を通じて確立されます。我々は、大規模ビジョン言語モデル(LVLM)におけるインタラクティブな視覚的基盤のための制御された評価フレームワークを導入し、事前に提供される対象情報の量と対話を通じて取得する必要がある情報の量を変化させます。人間に基づいた 4 つのビジュアル コンテキストと 4 つのインタラクション プロトコルにわたって、現在の LVLM のパフォーマンスはタスク レベルの人間のベースラインを大幅に下回っています。インタラクションは、フォローアップの質問によって最初のターゲットの説明を改良したり修正したりするときに役立ちます。最初の説明が提供されず、質問を通じてターゲット情報を取得する必要がある場合、パフォーマンスは最も低くなります。これは、積極的な質問主導のグラウンディングが依然として難しいことを示しています。 LVLM の校正も不十分であり、多くの場合、経験的精度を超える信頼性が報告されます。追跡調査では、さまざまな記述ソース (人間対 AI)、推論の取り組み、反復的なインタラクション、記述プロバイダー、および視覚的なコンテキストにわたってこれらのパターンが確認されています。全体として、インタラクティブな視覚的基盤は依然として重要な課題であり、視覚的なマッチング、情報探索、統合が必要です。

原文 (English)

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: target information is often incomplete, ambiguous, and established through interaction. We introduce a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), varying how much target information is provided upfront and how much must be acquired through dialogue. Across four human-grounded visual contexts and four interaction protocols, current LVLMs perform significantly below task-level human baselines. Interaction can help when follow-up questions refine or repair an initial target description. Performance is lowest when no initial description is provided and target information must be acquired through questions, indicating that proactive question-driven grounding remains difficult. LVLMs are also poorly calibrated, often reporting confidence that exceeds their empirical accuracy. Follow-up studies confirm these patterns across varied description sources (human versus AI), reasoning efforts, repeated interactions, description providers, and visual contexts. Overall, interactive visual grounding remains an important challenge, requiring visual matching, information seeking and synthesis.

13:00 JST研究/論文

オラクル前のルール: 審議的ポーリングのための監査可能でユーザー構成可能な引数の選択

熟議型世論調査では、投稿数が誰かが読む数を上回ると、何らかのメカニズムが各有権者がどの議論を見るかを選択し、決定の大部分を獲得します。実践では、不透明な学習済みランカーにそれを委任するため、有権者は自分の投票を形成した露出を再計算したり、異議を唱えたりすることはできません。私たちは、可読性を正確性と引き換えの目的ではなく、使用可能なメカニズムの許容条件として扱い、有権者が保持するパラメーターを使用して公的に再計算可能な証拠に対する公開されたルールになり得るかどうかを尋ねます。私たちは、二極性正当化セットに関する世論調査を形式化し、報道の理由、到着順序、および支持者の集まりによってスレートを判断します。市民推薦者のための 7 つのチェック可能な基準と、それらを満たすルール、つまり関係重み関数によってパラメーター化された 1 ホップ逆承認フローを示します。エージェント シミュレーターは、約 17,000 回のシード ペアの実行にわたって、投票ごとにすべてのスレートを記録します。提供されたスレートは、不透明なものも含めて、すべての選択手順のラベル読み取りの上限の上限に 0.035 足りません。制約のないランカーの利点は限られており、小さいものです。退化のないオーサリングを伴う報道だけでは、このルールはランダムな白板、つまり秩序や慈善活動に盲目の手段による無効なものと区別がつきません。他の 2 つでは、敵対的な圧力でマージンが広がり、すべてのプレフィックスでリードし、質量では 3.3 倍で優勢です。投稿の現実的な部分に理由がなくなると、カバー範囲のマージンは戻り、増加します。ラベル同種フラッディングは、フラットな重み付けポリシーの下では完全性を 0.81 から 0.34 に崩壊させますが、著者数の正規化の下では 0.44 にのみ低下するため、重み付け関数は完全性の 10% に値するセキュリティ制御になります。ランキングアームのどちらを選択するかは、事実ではなく、報道対マスの最前線での立場であり、読みやすいルールだけが影響を受ける人に与えることができる種類の選択です。これは、オープンソースのピアツーピア プラットフォームにマッピングされます。

原文 (English)

Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling

In a deliberative poll, once submissions outnumber what anyone will read, some mechanism chooses which arguments each voter sees, acquiring much of the decision; practice delegates it to opaque learned rankers, so a voter cannot recompute or contest the exposure that shaped their vote. We ask whether it can be a published rule over publicly recomputable evidence with parameters held by the voter, treating legibility as an admissibility condition on usable mechanisms, not an objective traded against accuracy. We formalise a poll over bipolar justification sets, judging a slate by reason coverage, the order it arrives in, and captured endorsement mass; we give seven checkable criteria for a civic recommender and a rule meeting them: a one-hop reversed endorsement flow parameterised by a relation-weight function. An agentic simulator records every slate at every vote, over about 17,000 seed-paired runs. Served slates fall 0.035 short of a label-reading ceiling upper-bounding every selection procedure, opaque ones included: any unconstrained ranker's advantage is bounded and small. On coverage alone, with non-degenerate authoring, the rule is indistinguishable from a random slate, a null due to an order-blind, charity-blind instrument; on the other two it leads at every prefix by a margin widening with adversarial pressure and dominates on mass by a factor of 3.3. Once a realistic fraction of submissions carries no reasons, the coverage margin returns and grows. Label-homogeneous flooding collapses completeness from 0.81 to 0.34 under a flat weight policy, only to 0.44 under author-count normalisation, making the weight function a security control worth 10% of completeness. The choice between ranking arms is a position on a coverage-versus-mass frontier, not a fact, the kind of choice only a legible rule can hand to the person it affects. It maps onto an open-source peer-to-peer platform.

13:00 JSTLLM/生成AI

記憶は必ずしも必要ではない: 科学的推論における条件付き記憶の特徴

科学的推論には、専門的な知識を取得し、それを複数ステップの計算に確実に組み込むための言語モデルが必要です。条件付き記憶は、密な神経表現を補完する明示的な検索経路を提供しますが、その有用性は本質的に入力と計算に依存します。取得された情報は、欠けている科学的関連性を修復する可能性がありますが、気が散るショートカットを導入したり、基本モデルがすでに正しく実行できる推論を妨げたりする可能性もあります。この研究では、条件付き記憶がいつ、どこで、どの程度まで科学的推論に関与すべきかを体系的に調査します。私たちは、科学的知識の境界と記憶対応の知識回路ノードに対する制御された介入を特徴づけます。これらの分析に基づいて、生成前に利用可能なタスク固有の入力プロキシを使用して、メモリがアクティブ化されているかどうか、どの層ステージノードがメモリ信号を受信するか、およびこれらの信号がどの程度強く寄与するかを判断する知識境界認識ルーターを提案します。 2 つのバックボーン ファミリと 6 つのタスク タイプを対象とした、生物学的および化学的推論のベンチマークに関する実験では、記憶効果が入力、タスク、注入場所によって大幅に異なることが示されています。静的およびアクティベーションレートに一致するランダムルーティングと比較して、私たちのアプローチは、メモリ誘発性の回帰を抑制しながら有益なメモリの寄与をより一貫して保存し、信頼性の高い科学的推論のための重要な原則として選択的なメモリ割り当てを確立します。

原文 (English)

Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning

Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it reliably into multi-step computation. Conditional memory provides an explicit lookup pathway that complements dense neural representations, but its usefulness is inherently input- and computation-dependent: retrieved information may repair missing scientific associations, yet it may also introduce distracting shortcuts or interfere with reasoning that the base model can already perform correctly. In this work, we systematically investigate when, where, and to what extent conditional memory should participate in scientific reasoning. We characterize the scientific knowledge boundary and controlled interventions on memory-enabled knowledge-circuit nodes. Based on these analyses, we propose a Knowledge Boundary-Aware Router that uses task-specific input proxies available before generation to determine whether memory is activated, which layer-stage nodes receive memory signals, and how strongly these signals contribute. Experiments on biological and chemical reasoning benchmarks, covering two backbone families and six task types, show that memory effects vary substantially across inputs, tasks, and injection locations. Compared with static and activation-rate-matched random routing, our approach more consistently preserves beneficial memory contributions while suppressing memory-induced regressions, establishing selective memory allocation as an important principle for reliable scientific reasoning.

13:00 JSTLLM/生成AI

推論による多様性: 将来予測のために LLM 群衆の知恵を活用する

大規模言語モデル (LLM) は将来予測にますます使用されており、群衆の知恵メカニズムとして複数のモデルを使用する動機になっています。ただし、異なる LLM が冗長な動作を示す可能性があるため、単に群衆のサイズを増やすだけでは効果的な多様性が保証されません。多様な LLM 群衆を構築するための行動認識フレームワークを提案します。このフレームワークは、独立した開発タスクの推論トレースを使用してモデルを特徴付け、動作の類似性によってモデルをクラスター化し、集団予測の代表を選択します。行動多様性モデリングのための 7 つの開発ベンチマークと、多様な群衆のパフォーマンスを評価するための 2 つの将来予測ベンチマークを使用して、25 の LLM を評価します。私たちの結果は、群衆のサイズよりも群衆の構成が重要である可能性があることを示しています。K-means++ 行動クラスタリングに基づく 3 モデルの medoid 群衆は、両方の予測ベンチマークで 25 モデルすべてに対する従来の投票を上回り、モデル呼び出しを 88% 削減し、推論コストを約 80% 削減します。この結果はさらに、効果的な LLM 群衆を構築するには、単に多様性を最大化するのではなく、代表的な行動の多様性が重要であることを示唆しています。

原文 (English)

Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction

Large language models (LLMs) are increasingly used for future prediction, motivating the use of multiple models as a wisdom-of-the-crowd mechanism. However, simply increasing crowd size does not guarantee effective diversity, as different LLMs may exhibit redundant behaviors. We propose a behavior-aware framework for constructing diverse LLM crowds. The framework characterizes models using their reasoning traces on independent development tasks, clusters models by behavioral similarity, and selects representatives for collective prediction. We evaluate 25 LLMs using seven development benchmarks for behavioral diversity modeling and two future-prediction benchmarks for evaluating diverse crowds' performance. Our results show that crowd composition can matter more than crowd size: a three-model medoid crowd based on K-means++ behavioral clustering outperforms conventional voting over all 25 models on both prediction benchmarks, while reducing model calls by 88% and inference cost by approximately 80%. The results further suggest that representative behavioral diversity, rather than simply maximizing diversity, is important for constructing effective LLM crowds

13:00 JST研究/論文

マルチドメイン知識追跡のための認知負荷と知識伝達の組み込み

知識トレーシング (KT) は、学習履歴から生徒の動的な知識状態を評価することを目的としています。既存の KT 手法のほとんどは単一ドメインの学習に焦点を当てており、顕著な成功を収めていますが、実際の学習シナリオでは複数のドメインが同時に関与することが多く、次の 2 つの重要な要素が導入されます。 1) 時間次元と知識次元の両方でドメインにわたる学習を管理することから生じる認知負荷。 2) 知識の伝達。1 つのドメイン内の知識の状態が、ドメイン内およびドメイン間の両方で関連する状態に影響を与えます。この論文では、マルチドメイン学習シナリオにおける学生の知識状態評価を改善するためにこれらの要因を調査することに焦点を当て、マルチドメイン知識トレースのための認知負荷と知識伝達を組み込んだ新しい方法(LT-MKT)を提案します。具体的には、孤立したドメインを橋渡しするために、LT-MKT はまず質問とそれに関連する概念からのテキスト情報を統合して、大規模言語モデル (LLM) の高度な表現機能を活用してマルチドメイン階層グラフを構築します。次に、時間次元と知識次元の両方におけるクロスドメインの特徴が明示的にモデル化され、認知負荷の影響が捉えられます。さらに、知識伝達モジュールは、ドメイン内およびドメイン間の知識状態の伝播をモデル化するように設計されています。 LT-MKT は、これらの要素を共同でモデル化することで、生徒の将来の成績をより正確に予測できるようになります。最後に、現実世界のデータセットに対する広範な実験により、私たちの手法が最先端のパフォーマンスを達成できることが実証されました。

原文 (English)

Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing

Knowledge Tracing (KT) aims to assess students' dynamic knowledge states from their learning histories. While most existing KT methods focus on single-domain learning with notable success, real-world learning scenarios often involve multiple domains simultaneously, introducing two critical factors: 1) Cognitive load, arising from managing learning across domains in both temporal and knowledge dimensions. 2) Knowledge transfer, where knowledge states in one domain influence related states both within and across domains. In this paper, we focus on exploring these factors to improve students' knowledge state assessment in multi-domain learning scenarios and propose a novel method incorporating cognitive Load and knowledge Transfer for Multi-domain Knowledge Tracing (LT-MKT). Specifically, to bridge isolated domains, LT-MKT first integrates textual information from questions and their associated concepts to construct a Multi-domain Hierarchical Graph, leveraging the advanced representational capabilities of large language models (LLMs). Then, cross-domain features in both the temporal and knowledge dimensions are explicitly modeled to capture the effects of cognitive load. Additionally, a knowledge transfer module is designed to model the propagation of knowledge states within and across domains. By jointly modeling these factors, LT-MKT enables more accurate prediction of students' future performance. Finally, extensive experiments on real-world datasets demonstrate that our method achieves state-of-the-art performance.

13:00 JSTエージェント

デスクトップ GUI エージェントのアクションによって引き起こされる視覚的な違いによる反映

Planner-Operator-Reflector (POR) フレームワークは、モジュラー コラボレーションを通じて複雑なタスクにおける目標の調整を維持するために、GUI エージェントで広く使用されています。ただし、デスクトップ GUI には重要な課題があります。大規模で高密度のインターフェイスでは、微妙な状態変化や分散した状態変化が見られることが多く、プランナーやオペレーターが 1 つの状態について推論する間に、アクション前とアクション後の画面を比較する必要があるリフレクターに負担の大部分がかかります。既存のリフレクターは、変更の検出と結果の検証を 1 つのステップにまとめてしまうため、暗黙の証拠が残り、根拠の弱い決定が得られます。この制限に対処するために、アクションによる視覚的な差異の抽出を結果の検証から明示的に切り離す 2 段階のリフレクターである証拠優先反射 (EFR) を提案します。 EFR は、アクションの場所と候補の変更領域を Set-of-Marks アノテーションで識別し、アクション関連の変更を記述してフィルタリングし、クリーンな証拠から最終的な判断を下します。この証拠と推論の分離された設計により、視覚的な検索の複雑さと推論の負担を軽減しながら、画面遷移に基づいたリフレクションが可能になります。 OSWorld-Verified と WindowsAgentArena での実験では、EFR によってリフレクターの精度が 7.11% 向上し、2 つのベンチマークでそれぞれ平均エンドツーエンド タスクの成功率が 5.94% と 4.95% 向上することが実証されました。

原文 (English)

Reflection with Action-Induced Visual Differences for Desktop GUI Agents

The Planner-Operator-Reflector (POR) framework is widely used in GUI agents to maintain objective alignment in complex tasks through modular collaboration. However, desktop GUIs introduce a key challenge: large, dense interfaces often exhibit subtle or scattered state changes, placing most of the burden on the reflector, which must compare pre- and post-action screens, while the planner and operator reason over a single state. Existing reflectors collapse change detection and outcome verification into one step, leaving evidence implicit and yielding weakly grounded decisions. To address this limitation, we propose Evidence-First Reflection (EFR), a two-stage reflector that explicitly decouples action-induced visual differences extraction from outcome verification. EFR identifies the action location and candidate changed regions with Set-of-Marks annotations, describes and filters action-relevant changes, and makes the final judgment from the cleaned evidence. This evidence-reasoning decoupled design makes reflection better grounded in screen transitions, while reducing both visual search complexity and reasoning burden. Experiments on OSWorld-Verified and WindowsAgentArena demonstrate that EFR improves reflector accuracy by 7.11%, yielding average end-to-end task success gains of 5.94% and 4.95% on the two benchmarks, respectively.

13:00 JSTLLM/生成AIエージェント

自信を超えて: 検索グラウンディングによるマルチターン検索エージェントのテスト時間のスケーリング

信頼ベースの投票は、トークン ログ確率などの内部信号でそれぞれを重み付けすることによって並列 LLM ロールアウトを集約し、シングル ターン推論について積極的に研究されています。ただし、最新の LLM は、外部ドキュメントを取得して条件を付けるマルチターン検索エージェントとして機能することが増えています。この論文では、信頼度に基づく投票がこのマルチターン設定にあまりうまく移行しないことを示し、失敗の根本的な原因がコピーのインフレであることを特定します。取得されたドキュメントがエージェントのコンテキストに追加されると、それらのドキュメントからコピーされたトークンは体系的にインフレートされたログ確率を受け取ります。これにより、各質問内の信頼スコアが平坦になり、結果として得られる加重投票が弱められます。この問題に対処するために、最終的な回答と取得したドキュメントの間の語彙の重複によって各ロールアウトをスコアリングする、検索根拠投票 (RGV) を提案します。 RGV は、汚染されたコンテキストの外側で信号を計算することにより、トークン ログの確率と追加の LLM 呼び出しの両方を回避します。 4 つの検索エージェント ベンチマークと 5 つの LLM にわたって、RGV は一貫して信頼度に基づく投票を上回り、精度が最大 +5.4% 向上し、正解が 8 件中 1 ~ 2 件のみである少数正解の質問では +35% 向上しました。

原文 (English)

Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding

Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such as token log probabilities, and has been actively studied for single-turn reasoning. However, modern LLMs increasingly act as multi-turn search agents that retrieve and condition on external documents. In this paper, we show that confidence-based voting transfers poorly to this multi-turn setting, and identify the underlying failure reason as copy inflation: when retrieved documents are appended to an agent's context, tokens copied from those documents receive systematically inflated log probabilities. This flattens confidence scores within each question and weakens the resulting weighted vote. To address this issue, we propose Retrieval-Grounded Voting (RGV), which scores each rollout by the lexical overlap between its final answer and the documents it retrieved. By computing the signal outside the contaminated context, RGV sidesteps both token log probabilities and additional LLM calls. Across four search-agent benchmarks and five LLMs, RGV consistently outperforms confidence-based voting, with gains of up to +5.4% accuracy and +35% on minority-correct questions, where the correct answer appears in only 1-2 of 8 rollouts.

13:00 JSTLLM/生成AI

マスクされたトレーニングによるワードレベルのタイムスタンプの相対時間間隔表現

Speech Large Language Model (SpeechLLM) は音声の理解と生成に優れていますが、きめ細かく時間的に調整された出力の能力はまだ十分に解明されていません。私たちの研究は、SpeechLLM が音声コンテンツと時間構造を共同でモデル化できるようにすることで、このギャップに対処し、SpeechLLM を「コンテンツ理解マシン」から「時間認識コンテンツ理解マシン」に効果的に変換します。具体的には、従来の絶対タイムスタンプを相対タイムスタンプに置き換え、よりコンパクトな語彙と強力な一般化機能を実現します。事前トレーニングされた大規模言語モデルにタイムスタンプ予測機能を効率的に組み込むために、ハイブリッド微調整戦略を導入します。つまり、タイムスタンプ拡張埋め込み層と言語モデル ヘッドのフルパラメータ微調整と、デコーダ層の LoRA 微調整を組み合わせたものです。さらに、マスクされたタイムスタンプのトレーニング目標を設計し、モデルがグラウンドトゥルースのタイムスタンプに過度に依存するのを防ぎ、それによってノイズの多い現実世界のアノテーションに対する堅牢性を強化します。広範な実験により、私たちのアプローチが強力な音声転写パフォーマンスを維持しながら、タイムスタンプ予測精度の大幅な向上を達成できることが実証されました。

原文 (English)

Relative Time Intervals Representation for Word-level Timestamping with Masked Training

Although Speech Large Language Models (SpeechLLMs) excel at speech understanding and generation, their capacity for fine-grained, temporally aligned outputs remains underexplored. Our work addresses this gap by enabling SpeechLLMs to jointly model speech content and temporal structure, effectively transforming them from ``content understanding machines" into ``temporal-aware content understanding machines". Specifically, we replace traditional absolute timestamps with relative timestamps, achieving a more compact vocabulary and stronger generalization capabilities. To efficiently infuse timestamp prediction ability into pre-trained large language models, we introduce a hybrid fine-tuning strategy: full-parameter fine-tuning of the timestamp-augmented embedding layer and language model head, combined with LoRA fine-tuning of the decoder layers. Moreover, we design a masked timestamp training objective, preventing the model from over-relying on ground-truth timestamps, and thereby enhancing robustness against noisy real-world annotations. Extensive experiments demonstrate that our approach achieves significant improvements in timestamp prediction accuracy while maintaining strong speech transcription performance.

13:00 JST研究/論文

アルゴリズムの影響により、調整の隠された社会的選択構造が明らかに

AI アルゴリズムが複数の人に影響を与える決定を下す場合、それを調整することが社会的選択の問題になります。システムの動作に関する人々の多様な好みをどのように調整し、単一の一貫したモデルに集約する必要があるのでしょうか。フロンティア AI モデルを調整するための標準的なアプローチ$\unicode{x2013}$人間のフィードバックからの強化学習$\unicode{x2013}$はこの疑問をほとんど回避しており、社会的選択の保証は不十分です。ただし、どのような代替手段がそれに代わるべきかは依然として不明です。我々は、アルゴリズムの福祉の結果に直接焦点を当てることによって、アライメント問題を凸面衝撃空間上の線形最適化として再定式化でき、それが福祉経済学とメカニズム設計の標準的なツールキットに適用できることを示します。この再定式化は、調整プロトコルが福祉の結果にどのように変換されるか、また逆に、福祉の結果に対する社会計画者の望ましい制約がどのように調整プロトコルに変換されるかを明らかにします。この変換を適用して、争点別投票とランダム独裁メカニズムが戦略的で全会一致であることを示します。逆の方向を実証するために、影響表現を適用して、個人またはグループの危害の制限など、さまざまな社会的要望に応じて功利的な社会福祉を最大化する一連の調整プロトコルを導き出します。私たちは、腎臓の割り当て、慈善食品の配布、LLM の反応、およびトロリーの問題に対する実際の人間の好みを使用して、これらの調整プロトコルの福祉への影響を経験的に説明します。

原文 (English)

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standard approach to aligning frontier AI models$\unicode{x2013}$reinforcement learning from human feedback$\unicode{x2013}$largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design. This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.

13:00 JSTLLM/生成AIエージェント

Agentic Alpha の汚染: マルチエージェント取引システムにおける役割とアーキテクチャにわたる敵対的な脆弱性

LLM ベースのマルチエージェント取引システムでは、専門のエージェントが構造化されたコミュニケーションを通じて連携して取引の意思決定を行いますが、研究用プロトタイプから実際の資産を管理する実際の導入に急速に移行しています。エージェント間のコミュニケーションを効果的にするのと同じエージェント間のコミュニケーションによって、それらが暴露されることもあります。つまり、破損したシグナルが最終決定にまで伝播し、現実の経済的損失につながる可能性があります。システム内部への特権アクセスを前提としたこれまでの攻撃とは異なり、攻撃者を実際に到達可能なもの (ソース データとエージェントが消費するプロンプト) に制限することで、障壁が低く、役割固有の攻撃者としてインスタンス化された民主化された脅威モデルが得られます。我々は、敵対的シグナルがどのようにしてマルチエージェント取引システムに侵入し、意思決定に至るまでどの程度存続するかを特徴付ける金融領域における初の体系的な実証研究を紹介する。役割軸に沿って、広く使用されているトレーディング パイプラインを 4 つの機能的役割 (アナリスト、リサーチャー、トレーダー、リスク マネージャー) に分解し、それぞれをそのインターフェイスに一致する攻撃と組み合わせます。構造軸に沿って、一部の設計が他の設計よりも堅牢である理由についての事後レンズとして敵対的信号保存スコア (APS) を使用して、データ レベルおよびエージェント レベルの攻撃の下で 4 つの通信トポロジを評価します。私たちは 5 つの資産、2 つのバックボーン、2 つのターゲット方向にわたって実験を実施します。重要な発見は、本質的に堅牢なアーキテクチャは存在しないということです。これらの調査結果は、より安全で堅牢なエージェント取引システムの将来の設計に関する洞察を提供します。

原文 (English)

Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems

LLM-based multi-agent trading systems, in which specialized agents collaborate through structured communication to produce trading decisions, are moving rapidly from research prototypes to live deployments that control real assets. The same inter-agent communication that makes them effective also exposes them: a corrupted signal can propagate to the final decision and translate into realized financial loss. Unlike prior attacks that presume privileged access to system internals, we restrict the adversary to what is practically reachable---the source data and prompts agents consume---yielding a low-barrier, and thus democratized threat model instantiated as role-specific adversaries. We present the first systematic empirical study in the financial domain to characterize how an adversarial signal enters a multi-agent trading system and how far it survives toward the decision. Along the role axis, we decompose a widely-used trading pipeline into four functional roles---Analyst, Researcher, Trader, and Risk Manager---and pair each with an attack matched to its interface. Along the structural axis, we evaluate four communication topologies under data- and agent-level attacks, using the Adversarial Signal Preservation Score (APS) as a post-hoc lens on why some designs are more robust than others. We conduct experiments across five assets, two backbones, and two target directions. A central finding is that no architecture is inherently robust. These findings provide insights for the future design of safer and more robust agentic trading systems.

13:00 JSTLLM/生成AI

圧縮トリニティ: LLM 圧縮のスパース性、量子化、低ランク近似の探求

計算コストと環境コストが法外に高いため、大規模言語モデル (LLM) のスケーラブルな展開が妨げられています。従来の圧縮技術 (スパース性、量子化、低ランク近似) は通常、個別に適用され、それぞれが精度と効率の壁にぶつかります。この論文は、計算量を削減するスパース性、メモリ帯域幅を最小化する量子化、精度を回復する低ランク近似という 3 つの柱を共同で適用する統一フレームワークである「圧縮トリニティ」を提案しています。事前トレーニングを高速化するために、オプティマイザーとモデル アーキテクチャに Trinity を適用します。 MKOR は、ブロック対角スパース性と低ランク反転によって曲率を近似し、量子化状態の数値安定性を維持します。曲率更新の複雑さを $O(d^3)$ から $O(d^2)$ に軽減し、KFAC と比較して収束を最大 1.85 倍加速します。 SLoPe は、精度を回復するためにトレーニングの最後の 1% で低ランクの「遅延」アダプターを使用し、N:M スパース性の二重枝刈りバックワード パスを介してトレーニングを最大 1.25 倍高速化します。トレーニング後の圧縮については、OPTIMA は重み再構築をグローバルに最適な列単位の 2 次プログラムとして定式化することでゼロ トレーニング レジームで静的マスクを安定させ、ゼロショット精度を最大 3.97% 向上させます。微調整予算が与えられると、PATCH は 0% ~ 50% の動的なハイブリッド スパース率を学習することで静的マスクの上限を突破し、最大 1.38 倍の高速化を実現します。最後に、SLiM は、数学的に導出された低ランク アダプターを使用して量子化とスパース性によって失われた情報を回復することで、完全な圧縮トリニティをワンショットで実現し、最先端の方法と比較して精度を最大 5.66% 向上させ、等しいパラメーター バジェットで非圧縮の密なモデルを 0.6% 上回るパフォーマンスを実現します。これらの結果を総合すると、効率的でスケーラブルな高性能 LLM には、圧縮トリニティを共同適用することが不可欠であることがわかります。

原文 (English)

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the "Compression Trinity," a unified framework that applies the three pillars jointly: sparsity to reduce computation, quantization to minimize memory bandwidth, and low-rank approximations to recover accuracy. To accelerate pretraining, we apply the Trinity to the optimizer and model architecture. MKOR approximates curvature via block-diagonal sparsity and low-rank inversion, maintaining numerical stability for quantized states; it reduces curvature update complexity from $O(d^3)$ to $O(d^2)$ and accelerates convergence by up to 1.85x over KFAC. SLoPe accelerates training by up to 1.25x via a double-pruned backward pass for N:M sparsity, using low-rank "lazy" adapters in the final 1% of training to recover accuracy. For post-training compression, OPTIMA stabilizes static masks in a zero-training regime by formulating weight reconstruction as globally optimal column-wise quadratic programs, improving zero-shot accuracy by up to 3.97%. Given a fine-tuning budget, PATCH breaks the ceiling of static masks by learning a dynamic hybrid sparsity ratio between 0% and 50%, yielding up to 1.38x speedups. Finally, SLiM realizes the full Compression Trinity in one shot, using mathematically derived low-rank adapters to recover information lost to quantization and sparsity, improving accuracy by up to 5.66% over state-of-the-art methods and outperforming uncompressed dense models at equal parameter budgets by 0.6%. Together, these results show that jointly applying the Compression Trinity is essential for efficient, scalable, high-performance LLMs.

13:00 JSTエージェントビジネス/資金調達

AgentWorld: エージェントによる情報取得のためのパーソナリティを意識した信頼性評価

エージェントによる情報検索の評価は依然として均一なユーザーとのスクリプト化された対話に限定されており、自然な性格の多様性と敵対的な脆弱性の両方が欠けています。我々は、(i) Big Five (OCEAN) の個性主導のユーザー集団とステートフルなツール使用環境を組み合わせたシミュレーション フレームワークである AgentWorld を紹介します。 (ii)構造化された障害分類、部分信用スコアリング、およびデュアルコントロールハンドオフ検証を備えた pass$^k$ 一貫性メトリック。 (iii) 6 つの微調整形式でのスコアしきい値トレーニング データのエクスポート。 (iv) 必要な中間状態スパインのスナップショットを作成し、4 つのタスク認識摂動タイプに基づいてモンテカルロ ロールアウトを分岐し、$\Delta P / \Delta T$ スコアリング、Dempster-Shafer 証拠融合、および Shapley 攻撃カテゴリの帰属を介してリスクを定量化する敵対的リスク アナライザー。 3 つの実験でフレームワークを実証します。10 の OCEAN ペルソナ (240 の評価者の判断) にわたる会話分析エージェント。 5 つのタスク $\times$ 4 つのペルソナ バリアントにわたるカスタマー サポート エージェント。 5 つのタスクの敵対的ストレス テストにより、既存の軌道の脆弱性 (摂動なしの $V_{\min}=0.375$) とツール/インフラストラクチャ層の攻撃の優位性 (Shapley: システム 46%、アクション 38%) が明らかになりました。性格のばらつきは、均一なテストでは明らかにできない障害モード、つまりクロスドメインのリーク、コンテキストのドリフト、0.27 ポイントの品質ギャップ、同じタスクのペルソナ間の 50% 対 100% の合格率を表面化します。一方、リスク アナライザーは、pass$^k$ だけでは測定できない軌道レベルの脆弱性を定量化します。

原文 (English)

AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval

Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We present AgentWorld, a simulation framework combining (i)Big Five (OCEAN) personality-driven user populations with stateful tool-use environments; (ii)the pass$^k$ consistency metric with structured fault classification, partial-credit scoring, and dual-control handoff verification; (iii)score-thresholded training-data export in six fine-tuning formats; and (iv)an adversarial Risk Analyser that snapshots required-intermediate-state spines, branches Monte-Carlo rollouts under four task-aware perturbation types, and quantifies risk via $\Delta P / \Delta T$ scoring, Dempster--Shafer evidence fusion, and Shapley attack-category attribution. Three experiments demonstrate the framework: a conversational analytics agent across 10 OCEAN personas (240 evaluator judgments); a customer-support agent across 5 tasks $\times$ 4 persona variants; and adversarial stress-testing of 5 tasks revealing pre-existing trajectory brittleness ($V_{\min}=0.375$ without perturbation) and tool/infrastructure-layer attack dominance (Shapley: 46% system, 38% action). Personality variation surfaces failure modes uniform testing cannot expose---cross-domain leakage, contextual drift, a 0.27-point quality gap, and 50% vs. 100% pass-rate across personas on the same task---while the Risk Analyser quantifies trajectory-level brittleness that pass$^k$ alone cannot measure.

13:00 JSTLLM/生成AIエージェント研究/論文

EMRB: 生の電磁信号に対する LLM 推論を評価するためのマルチレベル ベンチマーク

大規模言語モデル (LLM) は、科学および工学分析用のコード エージェントとしてますます使用されていますが、物理層の生の測定値を分析する機能はまだテストされていません。 \textbf{EMRB} (\textbf{E}lectro\textbf{m}agnetic \textbf{R}easoning \textbf{B}enchmark) を導入します。これは、LLM がコードを記述して実行することによって生の I/Q データを分析できるかどうかを評価します。 EMRB には、信号検出から OFDM 設計まで、5 つの難易度レベルと 27 の質問タイプにわたる 200 の問題が含まれており、検証済みのグランド トゥルースを備えた 11 の信号タイプから生成されます。前処理された機能や構造化テーブルに基づいて構築されたベンチマークとは異なり、EMRB は生のキャプチャのみを提供します。各質問が参照する数量は、最初にコードを通じて検出する必要があります。私たちは、独自の、オープンウェイト、および推論指向のファミリーにわたる 14 の LLM を評価します。スコアの範囲は 24.1\% から 78.9\% で、平均は基本測定の 84.9\% からシステム設計の 21.2\% まで低下します。また、信号偵察、対象を絞った分析、自己検証を分離する構造化手法である \textbf{ReconPilot} も提案します。 3 つのバックボーンにわたって、ReconPilot は全体のスコアを 3.8 から 17.6 ポイント上昇させ、テストした 15 のバックボーン レベルの組み合わせのうち 13 を改善しました。すべてのデータとコードは \href{https://github.com/mingxuZhang2/EMRB}{\textcolor{blue}{our GitHub リポジトリ}} で公開されています。

原文 (English)

EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals

Large language models (LLMs) are increasingly used as code agents for scientific and engineering analysis, but their ability to analyze raw physical-layer measurements remains untested. We introduce \textbf{EMRB} (\textbf{E}lectro\textbf{m}agnetic \textbf{R}easoning \textbf{B}enchmark), which evaluates whether LLMs can analyze raw I/Q data by writing and running code. EMRB contains 200 problems across five difficulty levels and 27 question types, from signal detection to OFDM design, generated from 11 signal types with verified ground truth. Unlike benchmarks built on preprocessed features or structured tables, EMRB provides only the raw capture; the quantities each question refers to must first be discovered through code. We evaluate 14 LLMs spanning proprietary, open-weight, and reasoning-oriented families. Scores range from 24.1\% to 78.9\%, with the mean dropping from 84.9\% on basic measurement to 21.2\% on system design. We also propose \textbf{ReconPilot}, a structured method that separates signal reconnaissance, targeted analysis, and self-verification. Across three backbones, ReconPilot raises the overall score by 3.8 to 17.6 points and improves 13 of 15 backbone-level combinations tested. All data and code are publicly released in \href{https://github.com/mingxuZhang2/EMRB}{\textcolor{blue}{our GitHub repository}}.

13:00 JSTエージェント

Android GUI エージェントは実行時の異常に対して堅牢ですか? AnTrap: 動的敵対環境におけるエージェントの評価

GUI エージェントは、Android デバイスに展開すると、予期しないポップアップからアクションの誤用に至るまで、動的な異常に遭遇することがよくありますが、既存のベンチマークには、ランタイム異常に対するエージェントの堅牢性の体系的な評価が欠けています。エージェントの実行軌跡に動的な摂動を注入する包括的なベンチマークである AnTrap を紹介します。私たちは、現実世界の異常を 10 のきめ細かいサブカテゴリーを備えた 4 つのレイヤー (状態、思考、アクション、ラウンド) に整理する分類法を提案し、現実的な敵対状況を導入しながらタスクの解決可能性を維持する構築パイプラインを開発します。 16 の主要な GUI モデルを評価したところ、動的異常に対する普遍的な脆弱性が明らかになり、最も強力なモデルでも大幅なパフォーマンス低下が発生しました。さらに、ベンチマークを検証するために、元の環境と敵対的な環境の両方で GRPO トレーニングを実施し、環境から学習可能な異常を推論のボトルネックとなった異常から分離します。私たちの調査結果は、ステート層とアクション層でのシングルステップ トラップは主に敵対的強化学習を通じて対処可能ですが、状態デッドロックのような深いコンテキスト トラップは、トラップだけを備えた環境でのトレーニングでは解決できない本質的な制限を露呈することを示しています。

原文 (English)

Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against runtime anomalies. We introduce AnTrap, a comprehensive benchmark that injects dynamic perturbations into agent execution trajectories. We propose a taxonomy organizing real-world anomalies into four layers (State, Thinking, Action and Round) with ten fine-grained subcategories, and develop a construction pipeline that preserves task solvability while introducing realistic adversarial conditions. Evaluating 16 leading GUI models, we reveal universal vulnerability to dynamic anomalies, with even the strongest models suffering significant performance degradation. Furthermore, we conduct GRPO training in both original and adversarial environments to validate our benchmark, separating environment-learnable anomalies from reasoning-bottlenecked ones. Our findings show that while single-step traps at state and action layers are largely addressable through adversarial reinforcement learning, deep contextual traps, like state deadlock, expose intrinsic limitations that cannot be resolved by training in environments with traps alone.

13:00 JSTLLM/生成AIエージェント

ACE: マルチスライド プレゼンテーション自動化のための自己修正エージェント キャンバス エディター

商用デザイン プラットフォームでは、ラージ言語モデル (LLM) エージェントを介してドキュメントを編集することが増えていますが、信頼性の高い展開を妨げる 2 つの現実的な問題があります。レガシー ドキュメント形式では、\emph{フラット} の絶対位置要素のみが公開されるため、エージェントは座標を再計算し、レイアウトを定期的に変更する必要があります。また、設計には固有のグラウンド トゥルースがないため、参照メトリクスとの差分により、有効ではあるが異なる出力がペナルティを受けます。プレゼンテーションに特化したアクション スペース (98 ツール) を備えた \emph{階層シーングラフ} 上のエージェント キャンバス エディタである \textbf{ACE} を、各デックの関連するスライスのみをエージェントにフィードするコンテンツ認識ルーター \textbf{CARE} と組み合わせて紹介します (平均 \ $\sim$89\% 入力トークン削減)。 \emph{ground-truth-free} 自然言語による批判が次のターンの指示としてフィードバックされる指示追従 (IF) 裁判官。固定バックボーンを使用すると、\emph{シングル ターン} のシーングラフ エディターは、内部で反復される同じバックボーンの \emph{agentic} HTML パイプラインとすでに一致します。自己修正を追加すると、1.75$\time$ の速度と $\sim$44\% 低いコストで、次の命令で ACE が大幅に上昇します (94 タスクのベンチマーク全体、$p{=}.010$ のペア、ループ外のジャッジによって複製されたもので、IF 4.23 対 \ 3.81)。 VQ 平均値は統計的に区別できませんが、26 人のブラインド評価者は全体的に ACE を好み (決定的勝率 58.7%)、81% の確率で自己補正された出力を好みました。ランキングは 3 つの裁判官ファミリー間で不変であり、ループ外の裁判官は自己修正ゲインの 3 分の 2 を保持し、循環性を制限します。ケースの 66\% が 1 回のパス後に停止し、厳密ピークのロールバックにより、観察されたすべての回帰が除去されます。

原文 (English)

ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation

Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only \emph{flat}, absolutely positioned elements, so agents must recompute coordinates and routinely break layouts; and design has no unique ground truth, so diff-against-reference metrics penalize valid-but-different outputs. We present \textbf{ACE}, an agentic canvas editor over a \emph{hierarchical scene-graph} with a presentation-specialized action space (98 tools), paired with \textbf{CARE}, a content-aware router that feeds the agent only the relevant slice of each deck (avg.\ $\sim$89\% input-token reduction), and a \emph{self-correction} loop driven by a \emph{ground-truth-free} instruction-following (IF) judge whose natural-language critique is fed back as the next-turn instruction. With a fixed backbone, a scene-graph editor in a \emph{single turn} already matches a same-backbone \emph{agentic} HTML pipeline that iterates internally; adding self-correction lifts ACE significantly above it on instruction following (IF 4.23 vs.\ 3.81 on the full 94-task benchmark, paired $p{=}.010$, replicated by an out-of-loop judge) at 1.75$\times$ the speed and $\sim$44\% lower cost. VQ means are statistically indistinguishable, but 26 blind raters prefer ACE overall (58.7\% decisive win-rate) and prefer the self-corrected output 81\% of the time; the ranking is invariant across three judge families, and out-of-loop judges retain two-thirds of the self-correction gain, bounding circularity. 66\% of cases halt after one pass, and a strict-peak rollback removes every observed regression.

13:00 JSTビジネス/資金調達

スケーラブルな質問中心のテキストから画像への評価: 信頼性の高いランキング、きめ細かい診断、コストを意識したルーティング

最新のテキストから画像への変換 (T2I) モデルは、合計スコアが似ているものの、強度が異なることが多く、実際の選択が困難になります。きめの細かいベンチマークはプロンプトを質問に分解しますが、多くの場合プロンプト スコアと固定カテゴリーに戻すため、帰属が弱まり、複雑さが無視されます。関連する要件も個別に、または 1 つの合計としてスコア付けされ、基本的な欠陥と構成的な欠陥がわかりにくくなります。オープンプロンプトを属性付きのアトミックな質問に変換し、Davidsonian Scene Graphs (DSG) でそれらの依存関係を整理する、質問中心のフレームワークである QC-T2I-Bench を紹介します。階層に制約された質問の集計を使用して、前提条件が満たされなかった後に下流の質問を除外し、単純なプロンプトと複雑なプロンプトが同じ合計重みを受け取るのを防ぎます。次に、DSG 構造を使用して、プロンプト内の共同成功を測定し、プロンプト間で繰り返されるエンティティを比較し、基本的な実現の失敗と追加要件に基づく失敗を分離します。英語と中国語のプロンプトで複数のオープンソース T2I モデルを評価します。得られた質問レベルの証拠は、信頼できるランキングと詳細な診断を裏付けます。結合完了率は、2 つの機能を持つコンポーネントの 80.7% から、7 つ以上の機能を持つコンポーネントの 37.2% に低下します。最後に、同じレコードをトレーニング不要のルーティングに再利用します。当社のコストを意識したルーターは、GPU/MP が 21.3\% 少なく、ERNIE の 89.51 ポイントの推定値と一致します。

原文 (English)

Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing

Modern text-to-image (T2I) models often have similar total scores but different strengths, making practical selection difficult. Fine-grained benchmarks decompose prompts into questions, yet often return them to prompt scores and fixed categories, weakening attribution and ignoring complexity. Related requirements are also scored separately or as one total, obscuring basic versus compositional failure. We present QC-T2I-Bench, a question-centric framework that converts open prompts into attributed atomic questions and organizes their dependencies with Davidsonian Scene Graphs (DSGs). We use hierarchy-constrained question aggregation to exclude downstream questions after a prerequisite fails and to prevent simple and complex prompts from receiving the same total weight. We then use the DSG structure to measure joint success within prompts and compare repeated entities across prompts, separating basic realization failures from failures under additional requirements. We evaluate multiple open-source T2I models on English and Chinese prompts. The resulting question-level evidence supports reliable ranking and fine-grained diagnosis: joint completion falls from 80.7\% for components with two capabilities to 37.2\% for those with seven or more. Finally, we reuse the same records for training-free routing; our cost-aware router matches ERNIE's 89.51-point estimate with 21.3\% less GPU-s/MP.

13:00 JSTLLM/生成AIエージェント

AHEAD: エージェントティック RL のための環境強化蒸留による適応的後知恵

強化学習を使用したマルチターン LLM エージェントのトレーニングは通常、軌道レベルの報酬に依存します。この報酬では、すべてのステップに均一の利点が割り当てられ、どの決定が成功または失敗につながったのかを特定できません。自己蒸留法は、RL に特権情報を追加することで、よりきめの細かい監視を提供できます。ただし、既存のアプローチは通常、重要な非対称性を無視して、同じ種類の特権情報を区別できない方法ですべてのステップに適用します。日常的なステップには追加のガイダンスはほとんど必要ありませんが、重大なエラーのステップには環境フィードバックだけでは提供できない修正指示が必要です。私たちは、さまざまな監視ソースをさまざまなステップタイプに適合させるステップ認識フレームワークである AHEAD を提案します。教師は、すべてのステップに関する環境フィードバックを接地された高密度信号として受け取り、さらに、環境フィードバックに欠けている方向性を提供するために、エラー ステップに関して LLM が生成した修正ヒントを受け取ります。この方法では、標準の GRPO アルゴリズムに最小限の変更が加えられます。 ALFWorld、WebShop、および検索ベースの QA、および 3 つのモデル スケールにわたって、AHEAD はタスクの成功率を高め (GRPO の 7B で、ALFWorld で +13.3 ポイント、WebShop で +11.0 ポイント)、より少ないトレーニング ステップで所定の成功率に到達し、結果のみの RL や以前の自己蒸留ベースラインよりも厳しいインタラクション バジェット内でタスクを解決します。

原文 (English)

AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL

Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewards, which assign a uniform advantage to every step and cannot identify which decisions led to success or failure. Self-distillation methods can provide finer-grained supervision by augmenting RL with privileged information. However, existing approaches usually apply the same type of privileged information to every step in an indistinguishable manner, ignoring a key asymmetry: routine steps need little additional guidance, while critical error steps require corrective direction that environment feedback alone cannot provide. We propose AHEAD, a step-aware framework that matches different supervision sources to different step types. The teacher receives environment feedback on all steps as a grounded dense signal, and additionally receives LLM-generated corrective hints on error steps to supply the direction that environment feedback lacks. The method introduces minimal changes to the standard GRPO algorithm. Across ALFWorld, WebShop, and Search-based QA, and across three model scales, AHEAD raises task success (+13.3 points on ALFWorld and +11.0 on WebShop at 7B over GRPO), reaches a given success rate in fewer training steps, and solves tasks within tighter interaction budgets than outcome-only RL and prior self-distillation baselines.

13:00 JST研究/論文

欠陥コード駆動のテストケース合成と高密度の報酬形成による堅牢なコード RL

検証可能な報酬からの強化学習 (RLVR) は、大規模言語モデル (LLM) のコード生成機能を強化するための極めて重要な手法として登場しました。ただし、コーディング実装における RLVR の有効性は、テスト ケースの包括性によって基本的に制限されます。これは、コード検証におけるテスト カバレッジが不十分なことが多くの場合誤検知を引き起こし、さらに報酬ハッキングやポリシーの劣化につながるためです。現在の自動生成手法の次善の品質に起因する報酬バイアスを軽減するために、我々は RobustTests フレームワークを提案します。このフレームワークは、「ほぼ正しい」欠陥コードを活用してモデルが潜在的な論理的不一致を正確に捕捉するようにガイドする欠陥コード駆動型のテスト ケース合成戦略を導入し、さらに検証エージェントと動作特徴クラスタリングを統合して、無効で冗長なテスト ケースの詳細なフィルタリングを容易にします。合成テスト ケースにおける固有の幻覚ノイズによって引き起こされる偽陰性に対処するために、RobustTests には合格率に基づいた段階的な高密度報酬関数も組み込まれており、きめ細かいフィードバックを通じてトレーニングの堅牢性が強化されています。このパイプラインを採用することで、CodeContest のテスト ケースを強化する高品質のデータセットを構築し、より広範囲の欠陥コード シナリオを網羅し、診断ユーティリティを大幅に強化します。実験結果は、CodeContests からの適度に困難な問題のサブセットをトレーニングに活用することにより、RobustTests を介した Qwen3-32B の RL 微調整がベースライン手法と比較して LiveCodeBench ベンチマークで絶対 3% のパフォーマンス向上を達成することを示しており、LLM のコード生成能力の向上における RobustTests フレームワークの有効性が確認されています。

原文 (English)

Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping

Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancing the code generation capabilities of Large Language Models (LLMs). However, the efficacy of RLVR in coding implementations is fundamentally limited by the comprehensiveness of test cases, because insufficient test coverage in code validation often causes false positives, further leading to reward hacking and policy degradation. To mitigate the reward bias stemming from the suboptimal quality of current automated generation methods, we propose the RobustTests framework, which introduces a faulty-code-driven test case synthesis strategy that leverages "near correct" faulty codes to guide the model in precisely capturing latent logical discrepancies and further integrates validator agents with behavioral feature clustering to facilitate the granular filtering of invalid and redundant test cases. To address false negatives caused by inherent hallucination noise in synthetic test cases, RobustTests also incorporates a stepwise dense reward function based on pass rates, bolstering training robustness through fine-grained feedback. By employing this pipeline, we construct a high-quality dataset that augmented the test cases in CodeContests, encompassing a broader spectrum of faulty code scenarios and significantly enhances diagnostic utility. Experimental results demonstrate that, by leveraging a moderately challenging subset of problems from CodeContests for training, RL fine-tuning of Qwen3-32B via RobustTests achieves an absolute 3% performance gain on the LiveCodeBench benchmark compared to baseline methods, confirming the effectiveness of the RobustTests framework in advancing the code generation proficiency of LLMs.

13:00 JST研究/論文

オムニジャッジかオムニバイアスか?バランスの取れた分離されたレンズを通じてマルチモーダルな裁判官を診断する

テキストから画像への変換 (T2I)、テキストからビデオへの変換 (T2V)、およびテキストから音声への生成 (TTS) の生成を共同で判断できるマルチモーダル理解モデルは、評価と自動アノテーションのための「オムニジャッジ」として使用されることが増えています。既存のベンチマークとトレーニング データは肯定的な例を強調しすぎ、個別の失敗モードを混同する傾向があるため、審査員が採点内容をどの程度確実に理解しているかは不明のままです。そのため、審査員は能力のギャップが隠れたまま失敗を認識せずに高得点を獲得する可能性があります。これを動機として、3 つのタスク全体で 53 の直交バイナリ次元 (17/22/14) と 10,671 サンプル (3,526/1,998/5,147) をカバーする、きめ細かいマルチモーダル理解を診断するためのバランスの取れた分離されたベンチマークである D3-Omni を導入します。次元間で情報が漏洩する可能性がある出力を再生成するのではなく、検証済みの完全にポジティブなシードを修正し、制御されたプロンプト書き換えとア​​トミックな次元分離摂動を通じてネガティブを導き出します。結果として得られる D3 デザインはデュアルバランスになっており、マイナスのサンプル不足と次元ごとのラベルの不均衡を軽減するのに役立ちます。分離されているため、各エラーは単一の機能に起因します。生成モデルの改善に伴い、ラベル分布の過小評価領域に向けて動的なステアリング構築が行われます。このスイートは、次元ごとにほぼ 1:1 のパリティと、すべての合計スコア レベルにわたって均一な分布に達します。このバランスの取れた見方の下では、強力なオムニジャッジでさえ、モダリティ関連の側面で苦戦する傾向があり、違反した要件を検出するよりもはるかに確実に満たされた要件を確認し、名目上異なる属性をほぼ単一の決定として扱うため、バランスの取れた分離されたレンズが明らかにし、ひいては対処するのに役立つ可能性のある体系的な盲点が集合精度によって隠蔽される可能性があることを示唆しています。

原文 (English)

OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses

Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D3-Omni, a balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding, covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks. Rather than re-generating outputs, which may leak information across dimensions, we fix verified fully positive seeds and derive negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations. The resulting D3 design is Dual-balanced, which helps alleviate negative-sample scarcity and per-dimension label imbalance; Decoupled, so that each error is attributable to a single capability; and Dynamic, steering construction toward under-represented regions of the label distribution as generative models improve.The suite reaches near 1:1 per-dimension parity and a uniform distribution over all total-score levels. Under this balanced view, even strong OmniJudges tend to struggle on modality-related dimensions, to confirm satisfied requirements far more reliably than they detect violated ones, and to treat nominally distinct attributes as largely a single decision, suggesting that aggregate accuracy may hide systematic blind spots that a balanced and decoupled lens can help expose and, in turn, address.

13:00 JSTエージェント

GUI 報酬モデリング用のタスク適応ルーブリック

GUI エージェントに関する最近の研究は、実行された軌跡がユーザーの指示によって暗示された成功基準を満たしているかどうかを判断することによって結果報酬を割り当てる結果報酬モデリングにますます焦点を当てています。ただし、既存の GUI 報酬検証ツールでは、タスク インスタンスごとにこれらの基準をどのように構築するかを十分に指定していないことがよくあります。一般的なルーブリック構造を使用する場合でも、暗黙的なモデル推論を使用する場合でも、その判断基準はタスクに十分に適応的ではありません。タスク間でチェックを転送したり、現在の指示の具体的な制約を見落としたり、明示されていない要件を強制することによって過度に厳格になったりする可能性があります。この制限に対処するために、カテゴリレベルの粗い段階とインスタンスレベルの細かい段階を通じてタスクに適応した判断基準を構築する、粗いから細かいルーブリックフレームワークである AdaptRubric を提案します。 AdaptRubric は、命令を GUI タスク ファミリにルーティングし、再利用可能なタスク ファミリ基準を取得することによって、カテゴリ レベルの粗いルーブリックの取得を実行します。次に、インスタンス レベルの詳細なルーブリック生成を実行して、現在の命令の具体的な値、スコープ、および制約のコンパクトな手がかりを表面化します。オフラインの報酬評価とオンラインの強化学習の最適化を通じて、AdaptRubric は一貫して以前の報酬エージェントを上回り、一致する画像予算の下で F1 をベースライン平均より 3.6 ポイント改善し、タスク成功率で 4.23 ポイントの向上をもたらしました。

原文 (English)

Task-Adaptive Rubrics for GUI Reward Modeling

Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success criteria implied by the user instruction. Existing GUI reward verifiers, however, often under-specify how these criteria should be constructed for each task instance. Whether using generic rubric structures or implicit model reasoning, their judging criteria are not sufficiently task-adaptive: they can transfer checks across tasks, overlook concrete constraints in the current instruction, or become overly strict by enforcing unstated requirements. To address this limitation, we propose AdaptRubric, a Coarse-to-Fine Rubrics Framework that constructs task-adaptive judging criteria through a category-level coarse stage and an instance-level fine stage. AdaptRubric performs category-level coarse rubric retrieval by routing the instruction to a GUI task family and retrieving reusable task-family criteria, then conducts instance-level fine rubric generation to surface compact cues for concrete values, scopes, and constraints in the current instruction. Across offline reward evaluation and online reinforcement learning optimization, AdaptRubric consistently outperforms prior reward agents, improving F1 by 3.6 points over the baseline average under a matched image budget and yielding a 4.23-point task-success gain.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体GPT / ChatGPT

Paritok-4B: コーディング エージェント向けの意図条件付きコンテキスト圧縮

コーディング エージェントは、大きなファイルの読み取りとツールの出力をフロンティア LLM に毎ターン再送信し、このコンテキストがトークンの請求額の大部分を占めます。汎用のプロンプト コンプレッサーは、散文と適合コードに基づいてトレーニングされているため、識別子を言い換えて、エージェントが編集する必要がある正確な範囲を削除します。我々は、2 つの取り組みに基づいて構築されたコーディング エージェント トラジェクトリ用の 4B LoRA コンプレッサーである Paritok-4B を紹介します。これは抽出的です。スパンを書き換えるのではなく選択し、出力する識別子、パス、数値の 96.0% が入力にすでに現れており、保持された SWE-bench Lite 出力では 96.2% を保持します。これはインテント条件付きです: エージェントの現在のタスクに通知され、エージェントは主に保持されたセグメント内で動作し、保持される量を変更するのではなく、どの行が生き残るかを選択します (保持された行は削除された行より +0.067 高いインテント関連性、ペアの 95% CI [+0.056, +0.078])。 67,074 の実際の OpenHands 軌跡にわたる gpt-4.1-mini 教師を 40,606 の検証済みサンプルに抽出し、Qwen3-4B を微調整します。 300 個の SWE ベンチ Lite インスタンスすべてで、Paritok-4B はエージェント コンテキストをそのサイズの 25.7% に圧縮します。これは、gpt-4.1-mini コンプレッサー (50.2%) の 2.0 倍、gpt-5 (61.9%) の 2.4 倍の圧縮率であり、非圧縮のシングルショット ソルブの品質の 86.5% を維持します。リアル エージェントが生成する cat -n 行番号の入力を入力すると、圧縮率はわずかに低く (27.8%)、保持率は高くなります (89.3%)。ここで、対応のあるテストは有益であり、30 のインスタンスは非圧縮のみで解決され、17 のインスタンスは圧縮のみで解決され、正確な McNemar p=0.079 であるため、このサンプル サイズでは、コンテキストをそのサイズの約 4 分の 1 に圧縮しても、解決率は大幅に低下しません。このモデルは、1 つの 24 GB GPU で自己ホストする 264 MB アダプターで、トークンごとのコンプレッサー料金はかかりません。定価では経済性が決まります。コンプレッサーとしての gpt-5 はネットネガティブであり、節約するダウンストリーム トークンよりもコストがかかります。重み、データ、評価スクリプトはオープンです (Apache 2.0)。

原文 (English)

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context dominates their token bill. General-purpose prompt compressors are trained on prose and suit code poorly: they paraphrase identifiers and drop the exact spans an agent needs to edit. We present Paritok-4B, a 4B LoRA compressor for coding-agent trajectories built on two commitments. It is extractive: it selects spans rather than rewriting them, and 96.0% of the identifiers, paths, and numbers it emits already appear in its input, holding at 96.2% on held-out SWE-bench Lite output. It is intent-conditioned: told the agent's current task, it acts chiefly inside a retained segment, selecting which lines survive (retained lines are +0.067 more intent-relevant than removed ones, paired 95% CI [+0.056, +0.078]) rather than changing how much is retained. We distil a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories into 40,606 validated examples and fine-tune Qwen3-4B. On all 300 SWE-bench Lite instances, Paritok-4B compresses agent context to 25.7% of its size, 2.0x harder than a gpt-4.1-mini compressor (50.2%) and 2.4x harder than gpt-5 (61.9%), while retaining 86.5% of uncompressed single-shot solve quality. Fed the cat -n line-numbered input real agents produce, it compresses slightly less (27.8%) and retains more (89.3%); there the paired test is informative, with 30 instances solved only uncompressed and 17 only compressed, an exact McNemar p=0.079, so at this sample size compressing context to roughly a quarter of its size does not significantly reduce the solve rate. The model is a 264 MB adapter that self-hosts on one 24 GB GPU with no per-token compressor fee, which at list prices decides the economics: gpt-5 as a compressor is net-negative, costing more than the downstream tokens it saves. Weights, data, and evaluation scripts are open (Apache 2.0).

13:00 JSTLLM/生成AI

大規模言語モデルにおける調整税を軽減するための優先データの選択

大規模な言語モデルを人間の好みに合わせて調整することは、現実世界の展開にとって非常に重要ですが、調整の負担が頻繁に発生し、事前にトレーニングされた一般的な機能の壊滅的な忘れにつながります。これまでの研究では主にこの問題を最適化またはアーキテクチャ上の課題として捉えていましたが、この低下を引き起こす嗜好データの固有の特性については、依然として十分に解明されていません。この論文では、位置合わせの有効性を最適化しながら壊滅的な忘却を明示的に軽減するバランスの取れたデータ選択戦略である BALIGN を提案します。選好最適化勾配の理論的および経験的分析を通じて、パラメーターのドリフトを決定する 3 つの重要なデータ中心の特徴、すなわち参照モデルの対数確率マージン、選択された応答と拒否された応答間のトークン長の差、一般能力コーパスとの TF-IDF の類似性を特定します。これらの直交特徴を統合された複合リスク スコアに集約することにより、BALIGN は、固有のモデル パラメーターを混乱させたり、最小のアライメント ユーティリティを提供したりする高リスクの選好サンプルを体系的に除外します。標準的な人間の嗜好データセットに対する広範な実験により、BALIGN がアライメントゲインを損なうことなく基本的な機能を強力に保持し、最小限の計算オーバーヘッドで最適なパレートフロンティアを一貫して達成できることが実証されました。

原文 (English)

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

Aligning large language models to human preferences is crucial for real-world deployment but frequently incurs an alignment tax, leading to the catastrophic forgetting of pre-trained general capabilities. While previous works primarily frame this problem as an optimization or architectural challenge, the inherent characteristics of preference data that drive this degradation remain largely underexplored. In this paper, we propose BALIGN, a balanced data selection strategy that explicitly mitigates catastrophic forgetting while optimizing alignment efficacy. Through theoretical and empirical analyses of the preference optimization gradient, we identify three key data-centric features that dictate parameter drift: the reference model's log-probability margin, the token length difference between chosen and rejected responses, and the TF-IDF similarity to general capability corpora. By aggregating these orthogonal features into a unified composite risk score, BALIGN systematically filters out high-risk preference samples that disrupt intrinsic model parameters or provide minimal alignment utility. Extensive experiments on standard human preference datasets demonstrate that BALIGN strongly preserves foundational capabilities without compromising alignment gains, consistently achieving the optimal Pareto frontier with minimal computational overhead.

13:00 JSTエージェント

MetaRAG: Agentic RAG の信念と行動に合わせたポリシーの最適化

エージェントによる検索拡張生成 (RAG) では、いつ検索を継続するか、いつ応答するかを決定する言語モデルが必要です。既存の RL ベースの方法は外部の監督に依存しており、現在の証拠が十分であるかどうかに関するエージェントの内部信念を見落としています。この問題に対処するために、我々は検索決定の品質を信念と行動の整合として再定式化し、エージェント RAG のための信念と行動が整合したポリシー最適化フレームワークである MetaRAG を提案します。 MetaRAG は、Verify-first Action Generation を使用して、各実際のアクションの前に明示的な検証プロセスを導き出し、内部信念プロービングを使用して、同じ質問履歴コンテキストからポリシー モデル自身の応答可能性信念を推定します。これらに基づいて、MetaRAG は、回答の正しさによってさらにゲートされる一貫性報酬を導き出し、内部的には一貫しているが不正確な軌道の強化を回避します。信念プローブはトレーニング中にのみ使用され、推論時のオーバーヘッドは発生しません。 7 つの公開 QA ベンチマークの実験では、MetaRAG が強力な RL ベースのエージェント RAG ベースラインと比較して精度と効率のトレードオフを一貫して改善し、その向上は深い研究設定、さまざまなオプティマイザー、および複数のモデル バックボーンに移行することが示されています。

原文 (English)

MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG

Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching and when to answer. Existing RL-based methods rely on external supervision and overlook the agent's internal belief about whether the current evidence is sufficient. To address this problem, we reformulate the search decision quality as belief-action alignment and propose MetaRAG, a belief-action aligned policy optimization framework for agentic RAG. MetaRAG uses Verify-first Action Generation to elicit an explicit verification process before each actual action, and Internal Belief Probing to estimate the policy model's own answerability belief from the same question-history context. Based on these, MetaRAG derives a consistency reward that is further gated by answer correctness, avoiding reinforcement of internally consistent but incorrect trajectories. The belief probe is used only during training and introduces no inference-time overhead. Experiments on seven public QA benchmarks show that MetaRAG consistently improves the accuracy-efficiency trade-off over strong RL-based agentic RAG baselines, with gains that transfer to deep research settings, different optimizers, and multiple model backbones.

13:00 JSTLLM/生成AI

大規模な言語モデルを使用した制約に基づくエンタープライズ データ マッピング

エンタープライズ エンティティの調整では、半構造化レコード、暗黙的な属性、単位または粒度の不一致を処理する必要があります。実際には手動による照合が依然として一般的ですが、スキーマやプロバイダーの進化に合わせて拡張することはできません。 LLM のみのマッチングは意味の再現を向上させますが、構造的および物理的な不変条件に違反する可能性があり、流暢ではあるが操作上無効な対応関係が生成されます。我々は、3 段階のニューロシンボリック手法である制約ガイド付きマッピング (CGM) を提案します。(i) メタデータ mc = を持つスキーマに基づいた許容性制約。ここで、tau_c は制約タイプを示し、delta_c は実行可能な関係と正規化ロジックを提供します。 (ii) ノイズ下で空ではない実行可能セットを保証するためのカスケード緩和を伴う制約制限付き候補生成。 (iii) その実行可能なセットに制限された有界 LLM 曖昧さ回避によるニューラル ランキング。方法論的には、制約は事後検証ではなく仮説空間演算子として機能し、緩和下での制御された劣化と、監査可能で人間が誘導可能な決定を可能にします。制御された構造デコイベンチマークでは、ハードアドミシビリティにより GT を落とすことなく候補空間が最大 480 倍縮小し、層ごとのアブレーションにより、LLM ではなくこのゲートが決定的な上昇であることがわかります (F1 0.08 ~ 0.66)。この利点はモデルに依存せず、追加の推論コストがかかりません。制約のある小規模なモデルは、制約なしで使用されるフロンティア LLM とほぼ 28 倍低いコストで一致します。この方法は、単一に調整された構成ではなく、7 つの企業メーカー (マクロ F1 0.70) にまたがって転送され、それぞれが自動的に検出され、専門家が調整可能な制約に基づいて、スプレッドシート ワークフローと比較して専門家の労力を最大 7 倍削減します。公開バレンタインの結果には、外部のランキング健全性チェックが追加され、境界がマークされます。制約は、構造的不変条件が一致を決定する場合にのみ厳しくする必要があります。

原文 (English)

Constraint-Guided Enterprise Data Mapping with Large Language Models

Enterprise entity alignment must handle semi-structured records, implicit attributes, and unit or granularity mismatches. Manual matching is still common in practice, but does not scale as schemas and providers evolve. LLM-only matching improves semantic recall, yet can violate structural and physical invariants, producing fluent yet operationally invalid correspondences. We propose constraint-guided mapping (CGM), a neuro-symbolic method with three stages: (i) schema-grounded admissibility constraints with metadata mc = , where tau_c denotes the constraint type and delta_c provides executable relation and normalization logic; (ii) constraint-restricted candidate generation with cascade relaxation to guarantee a nonempty feasible set under noise; and (iii) neural ranking with bounded LLM disambiguation restricted to that feasible set. Methodologically, constraints operate as hypothesis-space operators rather than post-hoc validators, enabling controlled degradation under relaxation and auditable, human-guidable decisions. On a controlled structural-decoy benchmark, hard admissibility shrinks the candidate space by ~480x without dropping the GT, and a layer-by-layer ablation shows this gate, not the LLM, is the decisive lift (F1 0.08 to 0.66). The benefit is model-independent and adds no extra inference cost: a small model with constraints matches a frontier LLM used without them at ~28x lower cost. The method, not a single tuned configuration, transfers across seven enterprise makes (macro F1 0.70), each under its own automatically discovered, expert-refinable constraints, and lowers expert effort by ~7x versus spreadsheet workflows. Public Valentine results add an external ranking sanity check and mark the boundary: constraints should be hard only where structural invariants are match-determining.

13:00 JSTLLM/生成AIハードウェア/半導体

検証されたタスク カバレッジを使用した複数の LLM 世代の評価

多くの LLM アプリケーションは、比較、検証、または組み合わせのために複数の候補出力を提供する場合に最も役立ちます。ただし、一般的な評価設定は依然として個々の出力に焦点を当てているか、複数のサンプルを 1 つの成功または選択された回答に絞り込んでいます。これにより、出力に真に異なる有用な結果がいくつか含まれているかどうかが見逃される可能性があります。この設定の 5 つのドメイン ベンチマークである VTC-Bench と、その中核となる評価量としての Validated Task Coverage (VTC) を導入します。このベンチマークは、慎重に選択された実際のデータ タスクから構築されており、モデルベースの判断を必要とせずに、出力品質とタスク関連の区別性の両方を自動的かつ再現可能にチェックできます。 VTC は、$k$ 回​​の試行内でどのくらいの明確な有用な結果が得られたかを測定します。複数のモデルと推論設定にわたって、ベンチマークは従来の評価とは異なる結論を導き出します。つまり、単一描画の品質から最強に見える構成が、必ずしも最高のカバレッジを備えているとは限らず、出力変動の単純な測定では、タスク関連のカバレッジを確実に回復することはできません。これらの結果は、有限の候補セットを対象オブジェクトとして直接評価できることを示しており、従来の出力ごとの評価では明らかではないモデルの動作の違いが明らかになります。

原文 (English)

Evaluating Multiple LLM Generations with Validated Task Coverage

Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination. Predominant evaluation settings, however, still focus on individual outputs or reduce multiple samples to a single success or selected answer. This can miss whether the outputs include several genuinely different useful results. We introduce VTC-Bench, a five-domain benchmark for this setting, together with Validated Task Coverage (VTC) as its core evaluation quantity. The benchmark is built from carefully selected real-data tasks where both output quality and task-relevant distinctness can be checked automatically and reproducibly, without model-based judges. VTC measures how many distinct useful results are obtained within $k$ attempts. Across multiple models and inference settings, the benchmark leads to different conclusions from conventional evaluation: configurations that look strongest from single-draw quality are not necessarily those with the best coverage, and simple measures of output variation do not reliably recover task-relevant coverage. These results show that finite candidate sets can be evaluated directly as objects of interest, revealing differences in model behavior that are not apparent from conventional per-output evaluation.

13:00 JSTビジネス/資金調達研究/論文

TRACE: 大規模な推論モデルの安全性評価のための証拠に基づいたベンチマーク

大規模推論モデル (LRM) は、最終的な応答が安全であるように見えても、安全でないコンテンツを含む可能性のある中間推論トレースを生成します。ガードレール モデルは、安全でないコンテンツを検出してブロックするように設計されていますが、安全でないコンテンツを検出するための既存のベンチマークは主にプロンプ​​トと最終応答に焦点を当てており、推論トレースはほとんど検査されていません。さらに、これらのベンチマークは通常、バイナリの安全性ラベルのみを提供し、判断を正当化する証拠の注釈は提供しません。これらの制限に対処するために、プロンプト、推論トレース、最終応答といった LRM 推論パイプライン全体をカバーする、証拠に基づいた安全性評価ベンチマークである TRACE を導入します。 TRACE には、9 つ​​のリスク カテゴリと 10 の攻撃戦略にわたる 2 つの言語のプロンプトが含​​まれています。プロンプトごとに、4 つの LRM が推論トレースと最終応答を生成し、各コンポーネントの安全性に注釈を付け、対応するソース テキストから裏付けとなる証拠を抽出します。 TRACE で 18 のガードレール モデルを評価すると、推論トレースの安全性判断は、プロンプトや最終応答よりも大幅に難しく、現在のモデルでは裏付けとなる証拠を正確に抽出するのに苦労していることが明らかになりました。これらの調査結果は、LR​​M 推論パイプライン全体で安全でないコンテンツを確実に検出し、正確に位置特定できるガードレール モデルの必要性を浮き彫りにしています。

原文 (English)

TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to detect and block unsafe content, yet existing benchmarks for unsafe content detection focus primarily on prompts and final responses, leaving reasoning traces largely unexamined. Moreover, these benchmarks typically provide only binary safety labels, without evidence annotations that justify the judgments. To address these limitations, we introduce TRACE, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses. TRACE includes prompts in two languages spanning nine risk categories and ten attack strategies. For each prompt, four LRMs generate reasoning traces and final responses, and we annotate the safety of each component and extract supporting evidence from the corresponding source text. Evaluating 18 guardrail models on TRACE reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models struggle to accurately extract supporting evidence. These findings highlight the need for guardrail models that can reliably detect and precisely localize unsafe content across the LRM inference pipeline.

13:00 JSTエージェント

STRIVE: 縦断放射線学レポート生成のための統合検証を備えたマルチエージェント構造化時間推論

縦断放射線学レポート作成 (LRRG) では、現在の所見と以前の研究と比較したその変化の両方を特定する必要があります。既存の方法は、暗黙的な表現内で診断、属性推定、時間的比較、および言語生成を共同でモデル化するため、タスクの干渉が発生し、各決定の基礎となる証拠が曖昧になり、エラーの追跡可能性が制限される可能性があります。また、進行状態を独立したラベルとしてモデル化し、その順序付けられた構造を無視するため、見逃された変化と方向の反転を同等に扱います。我々は、LRRG 向けの統合検証を備えたマルチエージェント構造化時間的推論である STRIVE を紹介します。これは、臨床推論を、明示的な中間証拠を生成する特殊な診断、属性、および時間変化エージェントに分解します。特に、Temporal Change Agent は、Progression-Aware GRPO を使用してさらに事後トレーニングされます。これは、方向反転のスコアを最低にしながら、方向維持エラーに部分的なクレジットを割り当てる検証可能な成形報酬です。 STRIVE は 2 つの段階で検証を実行します。決定論的な整合性ゲートはレポート生成前にエージェントの出力を調整し、検証エージェントは生成されたレポートが集約された臨床証拠によってサポートされているかどうかをチェックします。 Longitudinal-MIMIC では、STRIVE は最近の方法の中で最高の臨床有効性を達成し、参照レポートとの時間的一致の尺度である Longitudinal Change Concordance (LCC) を最も強力なベースラインと比較して 2 倍以上に高めています。

原文 (English)

STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation

Longitudinal radiology report generation (LRRG) requires identifying both current findings and their changes relative to a prior study. Existing methods jointly model diagnosis, attribute estimation, temporal comparison, and language generation within implicit representations, which can cause task interference, obscure the evidence underlying each decision, and limit error traceability. They also model progression states as independent labels, ignoring their ordered structure and thus treating missed changes and direction reversals equally. We present STRIVE, Multi-Agent Structured Temporal Reasoning with Integrated Verification for LRRG, which decomposes clinical reasoning into specialized Diagnosis, Attribute, and Temporal Change Agents that produce explicit intermediate evidence. In particular, the Temporal Change Agent is further post-trained using Progression-Aware GRPO, a verifiable, shaped reward that assigns partial credit to direction-preserving errors while scoring direction reversals lowest. STRIVE performs verification at two stages: a deterministic Consistency Gate reconciles the agent outputs before report generation, and a Validation Agent checks whether the generated report is supported by the aggregated clinical evidence. On Longitudinal-MIMIC, STRIVE attains the best clinical efficacy among recent methods and more than doubles Longitudinal Change Concordance (LCC), a measure of temporal agreement with the reference report, over the strongest baseline.

13:00 JSTLLM/生成AIエージェント研究/論文Claude

SA ベンチ: LLM ベースの論文複製における意味的整合性の評価

LLM エージェントは紙の複製コードを生成できますが、科学的に不忠実な実装を作成することもよくあります。私たちはこの障害モードをセマンティック ドリフトと定義します。この場合、生成されたコードは論文の仕様から静かに逸脱します。 ICLR、ICML、NeurIPS 2025 の 30 論文を対象とした診断ベンチマークである SemanticAlign-Bench(SA-Bench) を紹介します。各論文について、その仕様をアトミックで検証可能な実装クレームに分解し、これをセマンティック アライメント ユニット (SAU) と呼び、数値、方法論、プロトコル、順序ドリフトにわたる 4 つの診断次元に沿ってリポジトリを評価します。合計で、5 つの ML ドメインにわたって 1,491 の SAU を構築し、12 のジェネレーター構成 (4 つのモデル $\times$ 3 スキャフォールド) を評価します。最も強力な構成 (Claude+PaperCoder) でさえ、平均 SAU スコアは 1.0 点中わずか 0.301 であり、360 の評価全体の全体平均は 0.221 でした。失敗分類法により、エージェントはほとんどの要件を試行しますが、実装が不一致であり、実装の不一致とスタブがゼロスコアのクレームの大部分を占めているため、実装が間違っていることがわかります。私たちの分析はさらに、実行可能性のために最適化された足場が科学的再現に限定的な影響を与えることを示しています。ギャップを狭めるには、セマンティック仕様の検証を優先する足場が必要です。ベンチマーク、アノテーション、評価パイプラインは公開されています。

原文 (English)

SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction

LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations. We define this failure mode as semantic drift, where generated code silently diverges from the paper's specifications. We introduce SemanticAlign-Bench(SA-Bench), a diagnostic benchmark covering 30 papers from ICLR, ICML and NeurIPS 2025. For each paper, we decompose its specifications into atomic and verifiable implementation claims, which we call Semantic Alignment Units (SAUs) and evaluate repositories along four diagnostic dimensions spanning numerical, methodological, protocol and ordering drift. In total, we construct 1,491 SAUs across five ML domains and evaluate 12 generator configurations (4 models $\times$ 3 scaffolds). Even the strongest configuration (Claude+PaperCoder) achieves a mean SAU score of only 0.301 out of 1.0, with an overall mean of 0.221 across 360 evaluations. A failure taxonomy reveals that agents attempt most requirements but implement them incorrectly, with implementation mismatch and stubs accounting for the majority of zero-scored claims. Our analysis further indicates that scaffolds optimized for executability provide limited leverage for scientific reproduction; narrowing the gap requires scaffolds that prioritize semantic specification verification. The benchmark, annotations and evaluation pipeline are publicly available.

13:00 JSTハードウェア/半導体ビジネス/資金調達

精度を超えて: 法的に根拠のあるタスクにおける視覚言語モデルの二重判定評価プロトコル

AI システムは法的に責任のある設定として評価されることが増えており、正しい出力も適用される法的基準に対して正当化されなければなりません。既存の法律 AI ベンチマークと裁判官としての LLM プロトコルは、タスクのパフォーマンスと無制限の応答品質を測定するための重要なインフラストラクチャを提供します。私たちは追加の評価信号を 1 つ提供しています。それは、標準的な 0 ~ 10 の品質判定と、人間が厳選した参照に対する厳密なバイナリの意味同等性判定を組み合わせた二重判定プロトコルです。私たちは、管理された、視覚的に根拠のある規制タスク、つまり英国の交通標識の解釈を研究します。その意味は、すべての入力について既知の参照を伴う成文化された質問です。そして、単に2人の裁判官が同意しないのか(解釈上、同意しなければならない)かどうかだけでなく、どの程度、どこで同意するかを測定します。 7 つの可視性レベルと 2 つのオクルージョン モードの下での 4,680 件の評価では、2 人の裁判官は中程度に関連していますが (ポイントバイシリアル r = 0.644)、すべての評価の 8.0% に影響を与える非対称のタイプ II パターンが明らかになりました。その分布は有益です。単に、そこでは高スコアの回答が一般的であるため、限界率は高視認性 (v = 0.8 で 14.2%) でピークに達しますが、すでに 7 を超えるスコアを獲得している回答を条件とすると、この率は重度のオクルージョン下で最高になります (v <= 0.3 で 54 ~ 63%)。そのため、入力が最も劣化している場合、高品質スコアは最も信頼できません。信号がこのジャッジとリファレンスの特性であることは明確です。49 行の人間によるチェックでは、0 ~ 10 のジャッジが日常の読者の判断 (ピアソン r = 0.81、LLM 精度サブスコアで r = 0.80) とほぼ一致していることが示されていますが、同等性ジャッジはかなり一方向的に厳格です。このプロトコルは、評価ごとに 1 つの LLM 呼び出しを追加し、単一判定プロトコルが報告しないシグナルを明らかにします。プロンプト テンプレート、オクルージョンされたバリアント、および完全な評価結果をリリースします。

原文 (English)

Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks

AI systems are increasingly evaluated for legally accountable settings, where correct outputs must also be justifiable against an applicable legal standard. Existing legal-AI benchmarks and LLM-as-judge protocols provide important infrastructure for measuring task performance and open-ended response quality. We contribute one additional evaluation signal: a dual-judge protocol that pairs a standard 0-10 quality judge with a strict binary semantic-equivalence judge against a human-curated reference. We study a controlled, visually grounded regulatory task - UK traffic-sign interpretation, whose meaning is a codified question with a known reference for every input - and measure not merely whether the two judges disagree (by construction they must) but how much and where. On 4,680 evaluations under seven visibility levels and two occlusion modes, the two judges are moderately associated (point-biserial r = 0.644), while revealing an asymmetric Type II pattern affecting 8.0% of all evaluations. Its distribution is instructive: the marginal rate peaks at high visibility (14.2% at v = 0.8) simply because high-scoring answers are common there, but conditioned on the answer already scoring above 7, the rate is highest under heavy occlusion (54-63% at v <= 0.3), so a high quality score is least trustworthy when the input is most degraded. We are explicit that the signal is a property of this judge and reference: a 49-row human check shows the 0-10 judge aligns closely with everyday-reader judgement (Pearson r = 0.81; r = 0.80 with the LLM accuracy sub-score), while the equivalence judge is fairly but one-directionally stricter. The protocol adds one LLM call per evaluation and surfaces a signal single-judge protocols do not report. We release the prompt template, occluded variants, and full evaluation results.

13:00 JST画像/動画生成

リモートセンシングのための現実世界の知識に基づく変更データ合成

変更データ合成は、トレーニング データを拡張し、変更検出モデルのパフォーマンスを向上させるためのコスト効率の高いソリューションを提供します。しかし、既存の合成方法は通常、変更をシミュレートするための手作りのルールに依存しており、クラス遷移の範囲が限られているため、合成データの多様性が制限され、事前定義された遷移設計により、さまざまな変更タイプに対応する柔軟性が制限されます。この作業では、KnowChange を導入します。KnowChange は、事前トレーニング済みのビジョン言語モデルを知識ソースとして活用し、変更前のシーンと目的の変更タイプからの妥当な変更位置とクラス遷移を推論する、知識誘導型の変更データ合成フレームワークです。 KnowChange は、知識に基づく変更シミュレーションを一般化可能な合成モデルと統合することにより、統一されたフレームワーク内でさまざまな種類の変更を柔軟に合成できるようにします。広範な実験により、KnowChange で生成されたデータは、コンパクトなスケールで生成されているにもかかわらず、合成から実への転送と合成データの拡張の両方において既存の合成データセットよりも一貫して優れていることが実証されています。さらなる分析により、知識に基づく変更シミュレーションは既存の合成パイプラインにシームレスに統合でき、合成データの下流での有用性が向上することが示されています。

原文 (English)

Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing

Change data synthesis provides a cost-effective solution for expanding training data and improving the performance of change detection models. However, existing synthesis methods typically rely on handcrafted rules to simulate changes, where limited coverage of class transitions restricts the diversity of synthesized data, while predefined transition designs limit their flexibility in accommodating varied change types. In this work, we introduce KnowChange, a knowledge-guided change data synthesis framework that leverages pretrained vision-language models as knowledge sources to reason about plausible change locations and class transitions from pre-change scenes and desired change types. By integrating knowledge-guided change simulation with generalizable synthesis models, KnowChange enables flexible synthesis of diverse change types within a unified framework. Extensive experiments demonstrate that KnowChange-generated data consistently outperforms existing synthetic datasets in both synthetic-to-real transfer and synthetic data augmentation, despite being generated at a compact scale. Further analyses show that the knowledge-guided change simulation can be seamlessly integrated into existing synthesis pipelines and enhance the downstream utility of synthesized data.

13:00 JST研究/論文

継続的ナレッジグラフ埋め込みにおける候補セット干渉に対する一致した過剰アウトランカー正則化

継続的なナレッジ グラフの埋め込みにより、グラフの成長に応じてエンティティと関係の表現が更新されます。既存の方法は主に壊滅的な忘却に対処しますが、エンティティの承認により、互換性のあるすべてのクエリの候補範囲も変更されます。したがって、古いエンティティ間のスコアと順序が保持されている場合でも、履歴的な回答はランクを失う可能性があります。私たちは、この効果を候補者セットの干渉として形式化し、スムーズな回答に関連した新人のプレッシャーとスコアブラインドで構造的に一致する古い参照を比較するホストレベルの目標である一致超過アウトランカー正則化 (MEOR) を導入します。一方的なペナルティは、新規の競合が一致する参照を超えた場合にのみ機能し、正当な新しいエンティティに対するホストの学習者の信号を保持します。 ENTITY-ComplEx での 8 つのペア実行全体で、MEOR は履歴の現在のユニバース平均逆順位 (MRR) を再生より 0.0057 改善し、候補セットの干渉を 0.0055 減少させます (片側 95% 下限はそれぞれ 0.0052 と 0.0051)。これは、古いユニバースのランキングと新参者獲得の保存基準を満たし、永続的なキャリブレーション、一致した最大正則化子 (MMR)、および一致しない古い正則化子 (UOR) よりも過去の現在のユニバースの MRR を改善します。直接アブレーションは、その基準構築と凝集の各コンポーネントをサポートします。 MEOR を追加すると、報告された 10 件の FBInc-S および FBInc-L のホストおよびバックボーン設定すべてにおける履歴ランキングも改善され、ゼロを除くすべてのペアの 95% 信頼区間が向上します。これらの結果は、候補者の承認が継続的な順位低下の明確な原因であることを確立し、基礎となる埋め込みアーキテクチャまたは継続的学習器を置き換えることなく、候補者の承認を制御できることを示しています。

原文 (English)

Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding

Continual knowledge graph embedding updates entity and relation representations as a graph grows. Existing methods primarily address catastrophic forgetting, but entity admission also changes the candidate universe of every compatible query. A historical answer can therefore lose rank even when its score and its ordering among old entities are preserved. We formalize this effect as candidate-set interference and introduce Matched Excess-Outranker Regularization (MEOR), a host-level objective that compares smooth answer-relative newcomer pressure with score-blind, structurally matched old references. Its one-sided penalty acts only when newcomer competition exceeds the matched reference, preserving the host learner's signal for legitimate new entities. Across eight paired runs on ENTITY-ComplEx, MEOR improves historical current-universe mean reciprocal rank (MRR) by 0.0057 over replay and reduces candidate-set interference by 0.0055, with one-sided 95% lower bounds of 0.0052 and 0.0051, respectively. It satisfies the preservation criteria for old-universe ranking and newcomer acquisition and improves historical current-universe MRR over persistent calibration, matched maximum regularizer (MMR), and unmatched old regularizer (UOR). Direct ablations support each component of its reference construction and aggregation. Adding MEOR also improves historical ranking in all ten reported FBInc-S and FBInc-L host and backbone settings, with every paired 95% confidence interval excluding zero. These results establish candidate admission as a distinct source of continual rank loss and show that it can be controlled without replacing the underlying embedding architecture or continual learner.

13:00 JST研究/論文

持続可能な地球のための食事: 制約を意識した意思決定モデリングによる、パーソナライズされた持続可能な食生活の推奨

持続可能な食生活は、栄養の適切さ、経済的な手頃さ、文化の受容性、環境への配慮という 4 つの重要な柱間の多面的な相乗効果を表します。人口レベルの持続可能性モデリングが普及しているにもかかわらず、実際の実装は個人レベルでの効果的な導入に依存しています。この移行は個人間の不均一性によって妨げられることが多く、持続可能な食事の要件を個人の好みに合わせることが困難な課題となっています。この問題に対処するために、私たちは、ユーザーの好みとしてモデル化されるのではなく、学習可能な制約を通じて持続可能性が組み込まれる、制約を意識した意思決定メカニズムに基づいた、パーソナライズされた持続可能な食事推奨モデルを提案します。提案されたアプローチを体系的に評価するために、持続可能性指標を幅広くカバーしていることを特徴とする、約 150,000 のレシピを含む SusDiet という名前の持続可能な食事データセットを構築します。このデータセットの実験結果は、私たちの方法が個人の好みを損なうことなく、より持続可能な選択を促進することを示しています。この研究は、個人の食事の選択と地球の健康を調整するための枠組みを確立し、将来の持続可能な食事介入と持続可能な開発のための政策決定を導くための定量的証拠を提供します。

原文 (English)

Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling

A sustainable diet represents a multi-dimensional synergy among four essential pillars: nutrition adequacy, economic affordability, cultural acceptability, and environmental respect. Despite the prevalence of population-level sustainability modeling, practical implementation relies on effective individual-level adoption. This transition is often hindered by inter-individual heterogeneity, posing a formidable challenge in aligning sustainable diet requirements with individual preferences. To address this issue, we propose a personalized sustainable diet recommendation model based on a constraint-aware decision-making mechanism, where sustainability is incorporated through learnable constraints rather than modeled as user preferences. To systematically evaluate the proposed approach, we construct a sustainable diet dataset named SusDiet with about 150k recipes, characterized by broad coverage of sustainability indicators. Experimental results on this dataset show that our method promotes more sustainable choices without compromising individual preference. This work establishes a framework for aligning individual dietary choices with planetary health, offering quantitative evidence to guide future sustainable diet interventions and policy-making for sustainable development.

13:00 JSTLLM/生成AIエージェント

RePolicy: エージェント セーフガードでの安全ポリシー呼び出しのための強化学習

言語モデル エージェントを保護するには、コンテキスト依存の安全性ポリシーに基づいて完全な実行軌跡を評価する必要があります。既存のポリシーを意識した保護手段は主に、プロンプトまたは監視付きの微調整に依存しており、目に見えない軌道や変化するポリシー状況に適応する能力が制限されています。私たちは、強化学習を通じて安全ポリシーの呼び出しを学習するエージェント保護手段である RePolicy を提案します。エージェントの軌跡と動的なポリシー ライブラリが与えられると、RePolicy は該当するポリシーを呼び出し、その内容を使用してポリシーに基づいた理論的根拠と安全性の判断を生成します。教師あり初期化をサポートするために PolicyTraj-20K を構築し、続いて検証可能な報酬とポリシーコンテキストの摂動を備えた GRPO を構築します。 6 つのエージェント安全性ベンチマークにわたる実験では、RePolicy がさまざまなポリシー コンテキストの下で強力な全体的な安全性検出パフォーマンスと堅牢なポリシー呼び出しを達成していることが示されています。

原文 (English)

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an agent safeguard that learns safety-policy invocation through reinforcement learning. Given an agent trajectory and a dynamic policy library, RePolicy invokes the applicable policy and uses its content to produce a policy-grounded rationale and safety judgment. We construct PolicyTraj-20K to support supervised initialization, followed by GRPO with verifiable rewards and policy-context perturbation. Experiments across six agent safety benchmarks show that RePolicy achieves strong overall safety-detection performance and robust policy invocation under varying policy contexts.

13:00 JSTエージェント研究/論文ClaudeGemini

ReproAgent: 契約に基づいた紙からコードへの複製

論文からコードへの複製では、科学 AI エージェントに、研究論文を、論文の手法、プロトコル、アーティファクトを保存する実行可能なリポジトリに変換するよう求めます。仕様が分割されているため、これは困難です。アルゴリズム、メトリクス、アーティファクトなどの明示的な文書の内容は、エージェントの長い軌跡を通じて失われることがよくありますが、フレームワークのデフォルトや関連作業から継承された規約などの暗黙的な詳細は文書に記載されていません。 ReproAgent は、4 段階の Prepare-Plan-Generate-Repair パイプラインで、2 つのチャネルを備えた永続的な実装契約を中心に構築されています。1 つは紙のスニペットをコード義務に変える実装要件チャネル、もう 1 つは関連リポジトリからコンテンツと構造証拠を取得する参照証拠チャネルです。どちらも作業パッケージにバインドされ、ファイルレベルのコントラクトに投影され、生成と修復を通じて消費されます。 PaperBench Code-Dev では、ReproAgent は、Claude-Sonnet-4.5 と Gemini-3-Flash の両方の下で同じバックボーン スキャフォールドの中で最高の平均スコアに達しました。エンドツーエンドのチャネルアブレーションと論文ごとのケースは、両方のチャネルの貢献をサポートします。コードと実験成果物は公開されています。

原文 (English)

ReproAgent: Contract-Guided Paper-to-Code Reproduction

Paper-to-code reproduction asks scientific AI agents to turn research papers into executable repositories that preserve the paper's method, protocol and artifacts. This is difficult because the specification is split: explicit paper content such as algorithms, metrics and artifacts is often lost across long agent trajectories, while implicit details such as framework defaults and conventions inherited from related work are absent from the paper. We introduce ReproAgent, a four-stage Prepare--Plan--Generate--Repair pipeline built around a persistent implementation contract with two channels: an implementation-requirement channel that turns paper snippets into code obligations, and a reference-evidence channel that retrieves content and structure evidence from related repositories. Both are bound to work packages, projected into file-level contracts, and consumed across generation and repair. On PaperBench Code-Dev, ReproAgent reaches the highest mean score among same-backbone scaffolds under both Claude-Sonnet-4.5 and Gemini-3-Flash. End-to-end channel ablations and per-paper cases support the contribution of both channels. Code and experimental artifacts are publicly available.

13:00 JST研究/論文

VideoHarness-RSI: 凍結された視覚言語モデルを使用した長時間ビデオの理解のための再帰的ハーネスの自己改善

長いビデオの理解は、はるかに長いビデオから限定されたモデルのコンテキストをどのように構築するかに大きく依存します。既存のアプローチは、圧縮、検索、メモリ、およびエージェントによる証拠の取得を通じてこのプロセスを改善しますが、これらのメカニズムは通常、手動で設計された推論システムの一部として導入されるか、他のコンポーネントと一緒に最適化されます。このため、実行可能なコンテキスト構築プログラムだけを改善することでどれだけの効果が得られるのかという、より単純な質問を切り分けることが困難になります。私たちは、凍結されたビジョン言語モデル (VLM) の周りで実行可能なコンテキスト コンストラクターを再帰的に検索するための制御されたベースラインである VIDEOHARNESS-RSI を通じてこの疑問を研究します。アウターループのプロポーザーは、以前のプログラム、評価結果、および実行トレースを使用してハーネス候補を生成します。ハーネス候補は実行され、エンドツーエンドで評価されてから、成功したバリアントがさらなる検索のために保持されます。これにより、長時間ビデオの理解が自動ハーネス設計の制御されたインスタンスになります。検索可能なオブジェクトは実行可能なプログラム構造ですが、応答モデルとインターフェイスは固定されたままです。均一なサンプリングから開始される再帰的ハーネス検索は、一貫して改善の余地を発見し、いくつかの弱い手作りのベースラインを上回ります。代わりに、より強力な手作りのベースラインから開始して、同じ RSI プロセスによりさらなる改善がもたらされます。選択したハーネスは、さらに検索することなく、追加の長時間ビデオ ベンチマークにも転送されます。これらの結果を総合すると、別個の最適化レイヤーとして実行可能なコンテキストの構築が確立され、フリーズした VLM 周辺でのハーネスの検出と転送を研究するための再現可能なベースラインが提供されます。

原文 (English)

VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models

Long-video understanding depends critically on how a limited model context is constructed from a much longer video. Existing approaches improve this process through compression, retrieval, memory, and agentic evidence acquisition, but these mechanisms are typically introduced as part of a manually designed inference system or optimized together with other components. This makes it difficult to isolate a simpler question: how much can be gained by improving the executable context-construction program alone? We study this question through VIDEOHARNESS-RSI, a controlled baseline for recursively searching executable context constructors around a frozen vision-language model (VLM). An outer-loop proposer uses prior programs, evaluation outcomes, and execution traces to generate candidate harnesses, which are executed and evaluated end to end before successful variants are retained for further search. This makes long-video understanding a controlled instance of automated harness design: the searchable object is executable program structure, while the answering model and interface remain fixed. Starting from uniform sampling, recursive harness search consistently finds room for improvement and surpasses several weaker hand-crafted baselines. Starting instead from a stronger hand-crafted baseline, the same RSI process yields a further improvement. The selected harness also transfers to additional long-video benchmarks without further search. Together, these results establish executable context construction as a distinct optimization layer and provide a reproducible baseline for studying harness discovery and transfer around frozen VLMs.

13:00 JST研究/論文

OPDSearch+: 検索拡張推論のための RL リファインメントによるポリシーに基づく蒸留

検索拡張推論は、小規模な言語モデルでは依然として困難です。訓練を受けた教師によるオンポリシー蒸留 (OPD) は有望な方向性を示していますが、次の 2 つの問題があります。(1) 高品質のマルチターン探索軌跡は動的な検索者の応答に依存するため、SFT データを大規模に収集するには法外なコストがかかります。 (2) タスク固有のトレーニングを受けた教師には多額のトレーニング コストがかかりますが、タスク固有の微調整を行わずに既製の教師に OPD を直接適用すると、生徒は教師のパフォーマンスの上限に制約され、トレーニングの深刻な不安定性に悩まされます。私たちは、検索拡張推論のための教師の微調整を必要としない最初の蒸留パラダイムである OPDSearch+ を提案します。私たちは、ポリシーに関する蒸留における教師としての、凍結された既製の指示モデルの役割を調査し、重要な洞察を明らかにします。それは、教師が生徒のポリシー分布を再形成して、後続の RL が RL だけでは到達できない優れたソリューションに収束するようにするというものです。ステージ 1 では、学生はライブ検索エンジンと対話し、ポジションごとのフォワード KL 目標を介して蒸留され、タスク固有の教師トレーニングなしで推論分解と証拠統合スキルを習得します。第 2 段階では、RL は、より豊かな行動基盤から蒸留された生徒を洗練させ、RL だけではゼロから達成できないパフォーマンスを達成します。 7 つの QA ベンチマーク全体で、3B モデルを使用した OPDSearch+ は一貫して以前のすべての 3B RL ベースラインを上回り、HotpotQA で 13.1%、2WikiMultihopQA で 8.5% の向上を達成しました。

原文 (English)

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction, but suffers from two issues: (1) high-quality multi-turn search trajectories depend on dynamic retriever responses, making SFT data prohibitively expensive to collect at scale; (2) task-specifically trained teachers incur substantial training cost, while directly applying OPD with an off-the-shelf teacher without task-specific fine-tuning constrains the student to the teacher's performance ceiling and suffers from severe training instability. We propose OPDSearch+, the first distillation paradigm that requires no teacher fine-tuning for search-augmented reasoning. We investigate the role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation, and reveal a key insight: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach. In stage one, the student interacts with a live search engine and is distilled via a per-position forward KL objective, transferring reasoning decomposition and evidence integration skills without any task-specific teacher training. In stage two, RL refines the distilled student from a richer behavioral foundation, achieving performance that RL alone cannot reach from scratch. Across seven QA benchmarks, OPDSearch+ with a 3B model consistently outperforms all prior 3B RL baselines, achieving gains of 13.1% on HotpotQA and 8.5% on 2WikiMultihopQA.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文GPT / ChatGPT

音声エージェント評価のための LLM 審査員のベンチマーク: 信頼性、校正、人間による監視

会話型音声エージェントを大規模に評価するには、観察可能な対話品質と人間の評価者によって通常提供される状況に応じた判断の両方を捕捉する信頼性の高い評価方法が必要です。私たちは、通信および小売店の音声エージェントの会話について、会話の品質と安全性の側面にわたって人間の判断を GPT-4.1 および GPT-5 と比較することにより、裁判官としての LLM の評価を調査します。同じ相互作用を 3 つの評価構成 p0、p1、および p2 でスコアリングし、自動判定が評価設定の影響を受けやすいかどうか、および観察されたパターンが構成および判断モデル全体で一般化するかどうかをテストします。集計的な一致を超えて、指標レベルの相関関係、評価者の一貫性、人間と LLM の体系的な不一致を調査して、どの会話属性が自動化によって確実に判断でき、どの会話属性が解釈やコンテキストに影響を受けやすいかを特定します。効果的な音声エージェントの評価は、音声生成、ストリーミング、ASR、推論、ツール呼び出しの各段階にわたるエラー伝播などのパイプライン レベルの要素によっても形成され、人間と LLM の審査員が同じインタラクションをエンドツーエンドで採点する方法を比較することに重点を置く動機となります。私たちの結果は、LLM ベースの評価が大規模な音声エージェント評価の効果的な要素として機能する可能性があるが、その信頼性は一律ではなく測定基準と構成に依存することを示しています。これは、どの指標が自動評価に適しているかを特定するための経験的フレームワークを提供し、LLM 審査員がスケーラブルな評価を処理する一方で、人間の評価者が状況に応じた解釈とより信頼性の高い判断を必要とする指標に引き続き関与するハイブリッド パイプラインをサポートします。

原文 (English)

Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We investigate LLM-as-a-Judge evaluation by comparing human judgments with GPT-4.1 and GPT-5 on telecom and retail voice-agent conversations, across conversational quality and safety dimensions. The same interac- tions are scored under three evaluation configurations, p0, p1, and p2, to test whether automated judgments are sensitive to the evaluation setup and whether observed patterns generalize across configurations and judge models. Beyond aggregate agreement, we examine metric-level correlations, evaluator consistency, and systematic human-LLM disagreement to identify which conver- sational attributes can be judged reliably by automation and which remain sensitive to interpretation and context. Effective voice-agent evaluation is also shaped by pipeline-level factors such as speech generation, streaming, and error propagation across ASR, reasoning, and tool-calling stages, motivating our focus on comparing how human and LLM judges score the same interactions end to end. Our results show that LLM- based evaluation can serve as an effective component of large- scale voice-agent assessment, but that its reliability is metric- and configuration-dependent rather than uniform. This pro- vides an empirical framework for identifying which metrics suit automated evaluation and supports hybrid pipelines in which LLM judges handle scalable assessment while human evaluators remain engaged for metrics that demand contextual interpretation and higher-confidence judgment.

13:00 JST研究/論文

動的な内部フィールドはトランスフォーマーの認識を制御できるでしょうか?ホメオスタティックなコンピューティング制御における優位性ではなく認証性

インテリジェント システムは単に推論するだけではなく、どのくらいの量を計算するか、いつ停止するか、どのモジュールをアクティブにするかなど、独自の推論を制御します。その役割は、認知を実行せずに調節する動的な内部場、つまり明示的な物理学と証明された安定性を備えた低次元の恒常性状態によって果たせるだろうか?私たちのフィールドは、グラフ ラプラシアン上の偏微分方程式のファミリーによって管理されるモジュール グラフ上のフィールドであり、適応深さ推論器によって進歩します。当社は、ファミリー全体のインテグレータの安定性を証明します。これは、閉ループ証明書ではなく、インテグレータ証明書です。新しい、そしてここで証明されました: 速度結合を伴う Verlet の離散 Schur-Cohn 基準、潜在ルートごとに必要かつ十分、転流仮説なし。答えは 3 つあります。物質は「いいえ」、構造は一部のみ、「認証可能性はある」です。波、拡散、ゲート混合物、2D ナビエ・ストークス基板結合など、フィールドの物理のタイプは精度には関係ありません。 20 シードの事前登録された逆交絡キャンペーンが構造的主張を制限します。均等化されたキャップでは、二次効果は一方のファミリーでは強いですが (+0.087 [+0.042, +0.132], t=4.0)、もう一方のファミリーでは検出されません (+0.014 [-0.013, +0.040], n.s.)。したがって、元のコントラストの一部は順序ではなく容量でした。一致したインターフェイスの GRU は、1 つ目では区別できませんが、2 つ目では名目上フィールドを超えています (-0.035 [-0.067, -0.002])。この分野を区別するのは、機能ではなく、そのワンステップ オペレーターが正確な実行時安定性チェックを許可することです。これは、存在の違いではなく、種類の違いです。学習された再帰には、十分で保守的な証明書も含まれます。ポジティブコントロールを使用したキルゲートでは、証拠蓄積としてフィールドの証拠が見つかりません (デルタ AUC +0.0007 [-0.0065, +0.0079] 対 0.03 閾値)。動的な内部フィールドは、実行可能で認定可能な計算ガバナーですが、認識を強化するものではありません。それは調整するものであり、考えるものではありません。

原文 (English)

Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control

An intelligent system does not merely reason: it governs its own reasoning - how much to compute, when to stop, which module to activate. Can that role be played by a dynamic internal field - a low-dimensional homeostatic state with explicit physics and certified stability - that modulates cognition without performing it? Ours is a field on the module graph governed by a family of PDEs on the graph Laplacian, advancing with an adaptive-depth reasoner. We certify the stability of the integrator of the whole family - an integrator certificate, not a closed-loop one. New, and proved here: a discrete Schur-Cohn criterion for Verlet with velocity coupling, necessary and sufficient per latent root, with no commutation hypothesis. The answer is threefold: substance no, structure only in part, certifiability yes. The type of the field's physics is irrelevant for accuracy: wave, diffusion, gated mixtures and a 2D Navier-Stokes substrate tie. A twenty-seed preregistered deconfounding campaign bounds the structural claim: at equalized caps the second-order effect is strong in one family (+0.087 [+0.042, +0.132], t=4.0) but is not detected in the other (+0.014 [-0.013, +0.040], n.s.), so part of the original contrast was capacity, not order; and a matched-interface GRU is indistinguishable in the first and nominally exceeds the field in the second (-0.035 [-0.067, -0.002]). What distinguishes the field is not capability but that its one-step operator admits an exact runtime stability check - a difference of kind, not of existence: learned recurrences carry certificates too, sufficient and conservative ones. A kill-gate with a positive control finds no evidence for the field as evidence accumulator (Delta AUC +0.0007 [-0.0065, +0.0079] vs a 0.03 threshold). A dynamic internal field is a viable, certifiable compute governor, but not an enhancer of cognition: it modulates, it does not think.

13:00 JSTLLM/生成AI

SonarLLM: ネイティブ ソナー - 水中知覚のための光学マルチモーダル大規模言語モデル

信頼性の高い水中認識には、変動する視界の下での補完的なセンシングが必要です。光学カメラは外観と意味を捕捉しますが、濁りにより急速に劣化します。一方、イメージングソナーは、明確な距離方位角構造と音響アーチファクトを示しながら形状を保存します。したがって、主に光学エンコーダに基づいて構築された既存の MLLM は、ソナーをモデル化したり、ソナーと光学の相補性を適応的に利用したりするのには不向きです。我々は、ソナーをネイティブの知覚モダリティとして扱うソナー光学MLLMであるSonarLLMを提案します。ソナー固有のエンコーダ、モダリティ固有の物理学を意識した機能強化、信頼性を意識した階層融合を組み合わせて、音響構造を光学セマンティクスと調整し、センシング品質の変化に応じてそれらの寄与を動的に調整します。また、認識、カウント、視覚的な質問応答、キャプションの 4 つのタスクにわたるペアのベンチマークである SonarBench も紹介します。そして、ベンチマーク全体で、ソナーのみ、光学のみ、フュージョンの 3 つの入力設定があります。 SonarBench は、光学劣化を変化させながらシーンとソナー観察を固定することで、クロスモーダル相補性の制御された測定を可能にします。 SonarLLM は、ソナーのみの認識、計数、および VQA 全体で 72.0% のマクロ精度を達成し、最強のベースラインを 34.4 パーセントポイント上回っており、融合下では 68.7% で最高のベースラインを 25.1 ポイント上回っています。認識と計数では、濁度が増加するにつれて、光学上の融合ゲインが 6.0 ポイントから 36.0 ポイントに増加します。これは、制御された光学劣化の下でソナーの相補的価値が増加していることを示しています。これらの結果を総合すると、堅牢な異種知覚はソナーの追加だけでなく、その感知特性に応じた表現と重み付けにも依存していることがわかります。

原文 (English)

SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras capture appearance and semantics but degrade rapidly with turbidity, whereas imaging sonar preserves geometry while exhibiting distinct range-azimuth structure and acoustic artifacts. Existing MLLMs, built primarily on optical encoders, are therefore ill-suited to model sonar or adaptively exploit sonar-optical complementarity. We propose SonarLLM, a sonar-optical MLLM that treats sonar as a native perceptual modality. It combines a sonar-specific encoder, modality-specific physics-aware feature enhancement, and reliability-aware hierarchical fusion to align acoustic structure with optical semantics and dynamically adjust their contributions as sensing quality changes. We also introduce SonarBench, a paired benchmark that spans four tasks: recognition, counting, visual question answering, and captioning; and, across the benchmark, three input settings: sonar-only, optical-only, and fusion. By fixing the scene and sonar observation while varying optical degradation, SonarBench enables controlled measurement of cross-modal complementarity. SonarLLM achieves 72.0% macro accuracy across sonar-only recognition, counting, and VQA, outperforming the strongest baseline by 34.4 percentage points, and 68.7% under fusion, exceeding the best baseline by 25.1 points. For recognition and counting, the fusion-over-optical gain grows from 6.0 to 36.0 points as turbidity increases, indicating the increasing complementary value of sonar under controlled optical degradation. Together, these results show that robust heterogeneous perception depends not only on adding sonar, but on representing and weighting it according to its sensing characteristics.

13:00 JSTLLM/生成AI

選択的再生デコード: 推論時間推論のための軌跡レベルの介入

推論時間デコード手法は、複数の候補軌跡を探索することで LLM 推論を改善しますが、各軌跡をアトミックとして扱い、全体を保持するか、不可逆的に破棄します。これにより、高品質のプレフィックスが劣化したサフィックスとともに放棄される、部分的に有望な候補に対する計算が無駄になります。選択的再生復号化 (SRD) を導入します。これは、大規模なターゲット モデルを必要とせずに、境界候補の有用なプレフィックスを保持しながら、サフィックスの劣化部分のみを破棄、保持、または洗練するように各候補をルーティングします。穏やかな仮定の下では、SRD は、厳密に期待される軌跡の品質が厳密に高く、拒絶サンプリングと比較してサンプル効率において 1.28 ~ 1.36 倍の証明可能な利得を達成し、候補プールが増大するにつれて利得も増大します。 SRD は、複数の世代報酬モデル ペアを備えた MATH500、GPQA Diamond、HotpotQA、および AlpacaEval 全体で、実質的に少ない生成トークンで Best-of-N の精度と一致し、低コンピューティング領域での投機的拒否を上回ります。 SRD は、軌道全体の選択ではなくセグメント レベルの介入を有効にすることで、推論時間推論の精度と計算のトレードオフにおいて、これまで検討されていなかった領域を開拓します。

原文 (English)

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, yet treat each trajectory as atomic: either retaining it whole or discarding it irreversibly. This wastes computation on partially promising candidates whose high-quality prefixes are abandoned alongside degraded suffixes. We introduce Selective Regenerative Decoding (SRD), which routes each candidate to discard, keep, or refine only the degraded portion of the suffix while preserving the useful prefix of borderline candidates, without requiring a larger target model. Under mild assumptions, SRD achieves a provable 1.28-to-1.36-fold gain in sample efficiency over rejection sampling with strictly higher expected trajectory quality, with the gain growing as the candidate pool grows. Across MATH500, GPQA Diamond, HotpotQA, and AlpacaEval with multiple generation-reward model pairs, SRD matches Best-of-N accuracy with substantially fewer generated tokens and outperforms speculative rejection in low-compute regimes. By enabling segment-level intervention rather than whole-trajectory selection, SRD opens a previously underexplored region of the accuracy-compute tradeoff for inference-time reasoning.

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPT

ハンドオフ税: LLM エージェントにおける非ネイティブの軌跡の継続

コーディング エージェントは、数十のモデル呼び出し、ツールの使用、コード編集に及ぶ長時間実行タスクを実行します。これらの実行が展開されるにつれて、ユーザーは実際的なコストと品質のトレードオフに直面します。安価なモデルが苦戦した場合はより強力なモデルにエスカレートし、難しい推論が完了するとシフトダウンします。各スイッチでは、受信機が別のモデルによって生成された非ネイティブ軌道を継続する必要があります。私たちは、このハンドオフが品質とコストにどのような影響を与えるか、また受信機が継承する軌道情報の変化によって結果がどのように変化するかを研究します。 Claude および GPT ファミリの低コスト、低機能 (LC) モデルと高コスト、高機能 (HC) モデルのペアを使用して、ハンドオフの方向、タイミング、インターフェイスを変更し、リポジトリの状態を維持しながら、完全な軌跡の転送、圧縮、および軌跡の削除を比較します。どちらのモデル ファミリでも、完全な軌道エスカレーションでは、LC から HC への品質ギャップの半分未満が回復しますが、大幅なコスト増が発生します。私たちは、このコスト品質に対するペナルティを引継ぎ税と呼んでいます。対照的に、シフトダウンはコスト品質において有利な点を提供します。興味深いことに、優先インターフェイスも方向によって逆転します。LC モデルの軌道情報を減らすとエスカレーションの品質が向上しますが、HC モデルの軌道を削除するとダウンシフトの品質が低下します。

原文 (English)

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once the hard reasoning is complete. Each switch requires the receiver to continue a non-native trajectory produced by another model. We study how this handoff affects quality and cost, and how varying the trajectory information inherited by the receiver changes the outcome. Using pairs of low-cost, low-capability (LC) and high-cost, high-capability (HC) models from the Claude and GPT families, we vary handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state. Across both model families, full-trajectory escalation recovers less than half of the LC-to-HC quality gap while incurring a substantial cost premium. We term this cost-quality penalty the handoff tax. By contrast, downshift offers a favorable cost-quality point. Interestingly, the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.

13:00 JSTLLM/生成AIエージェント

マルチエージェント システムにおける障害の原因特定のための適応影響グラフ

マルチエージェント LLM システムは、障害が発生するとコストが高くつき、特定が難しい現実のアプリケーションに導入されることが増えています。障害の原因特定を自動化する取り組みが強化されているにもかかわらず、失敗した実行の診断は依然として人間のエンジニアに大きく依存しています。しかし、エンジニアが生のログをエンドツーエンドで読んで複雑なシステムをデバッグすることはほとんどありません。代わりに、可観測性ツールは、コンポーネント、アクション、依存関係を中心にトレースを整理し、ターゲットを絞ったナビゲーションをサポートします。私たちは、現代の LLM も同じパラダイムから恩恵を受けることができると仮説を立てています。この仮説を検証するために、まず失敗したトレースを構造化グラフに変換し、次にそれをナビゲートして重大なエラーを特定する 2 段階のエージェント フレームワークであるアダプティブ インフルエンス グラフ (AIG) を導入します。複数のモデルにわたって、より豊富なトレース表現により障害の原因特定が一貫して改善され、適応グラフ構築とエージェント主導の走査により最も強力な結果が得られることがわかりました。 AIG は、マルチエージェント障害の原因特定の標準ベンチマークである Who&When に関する新しい最先端技術を確立します。これは、帰属は診断モデルだけでなく、痕跡がどのように表現され調査されるかにも依存するという私たちの仮説を裏付けています。

原文 (English)

Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems

Multi-agent LLM systems are increasingly deployed in real-world applications, where failures can be costly and difficult to localize. Despite growing efforts to automate failure attribution, diagnosing failed runs still largely relies on human engineers. Yet engineers rarely debug complex systems by reading raw logs end to end. Instead, observability tools organize traces around components, actions, and dependencies to support targeted navigation. We hypothesize that modern LLMs can benefit from the same paradigm. To test this hypothesis, we introduce Adaptive Influence Graphs (AIGs), a two-stage agentic framework that first transforms a failed trace into a structured graph and then navigates it to identify the critical error. Across multiple models, we show that richer trace representations consistently improve failure attribution, with adaptive graph construction and agent-directed traversal yielding the strongest results. AIGs establish a new state of the art on Who&When, the standard benchmark for multi-agent failure attribution. This affirms our hypothesis that attribution depends not only on the diagnosing model, but also on how the trace is represented and explored.

13:00 JSTエージェント

状態から行動へ: 信頼性の高い複数回転工具の使用のための OODA ツール

信頼性の高いマルチターン ツールの使用には、エージェントが進化するタスクの状態を保存し、各アクションがその状態と一貫性を保っていることを確認する必要があります。ただし、直接関数呼び出しと ReAct スタイルのポリシーは、同じ自己回帰軌道内で状態の追跡とアクションの生成を学習します。この結合により状態と動作の競合が生じます。次の呼び出しを生成するというプレッシャーにより、対話の以前に蓄積された情報が上書きされたり、無視されたりする可能性があります。ボイドの観察、指向、決定、行為のサイクルに触発され、状態の維持を行動の実現から分離することでこの競争を緩和するように設計された型付き閉ループ ポリシーである OODA ツールを紹介します。 OODA ツールは、インタラクション履歴から直接アクションを生成するのではなく、コントローラーがチェックした中間状態を介して各決定をルーティングし、最終出力が現在のタスク状態に確実に固定されるようにします。具体的には、Observe はタスクの状態を再構築し、Orient は実行が正当であるかどうかを決定し、Decide は許容可能なアクション構造を形成し、Act は外部出力を実現します。マルチターン、マルチツール、不完全情報設定全体で 0.6B から 14B の範囲の Qwen3 モデルを使用して、直接関数呼び出しと ReAct ポリシーに対して OODA ツールを評価します。 OODA-Tool は、モデル サイズ全体でタスクの成功率を一貫して向上させます。小さいモデルや、アクションがターン全体で蓄積された情報や以前のツールの結果に大きく依存するタスクでは、より大きな効果が得られます。制御されたバリアント、段階レベルのアブレーション、および移植の評価は、これらの改善の堅牢性をさらに実証します。

原文 (English)

From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use

Reliable multi-turn tool use requires an agent to preserve an evolving task state and ensure that each action remains consistent with it. However, direct function-calling and ReAct-style policies learn state tracking and action generation within the same autoregressive trajectory. This coupling creates state-action competition: the pressure to produce the next call can overwrite or ignore information accumulated earlier in the interaction. Inspired by Boyd's Observe-Orient-Decide-Act cycle, we introduce OODA-Tool, a typed closed-loop policy designed to mitigate this competition by separating state preservation from action realization. Rather than generating an action directly from the interaction history, OODA-Tool routes each decision through controller-checked intermediate states, ensuring that the final output remains grounded in the current task state. Specifically, Observe reconstructs the task state, Orient determines whether execution is warranted, Decide forms an admissible action structure, and Act realizes the external output. We evaluate OODA-Tool against direct function-calling and ReAct policies using Qwen3 models ranging from 0.6B to 14B across multi-turn, multi-tool, and incomplete-information settings. OODA-Tool consistently improves task success across model sizes, with larger gains on smaller models and on tasks whose actions depend strongly on information accumulated across turns and prior tool results. Controlled variants, stage-level ablations, and transfer evaluations further demonstrate the robustness of these improvements.

13:00 JSTLLM/生成AI

レシピにペルソナはありますか?属性付きプロシージャル グラフのクリエイター スタイルの特徴付けと生成

大規模言語モデル (LLM) は膨大なゼロショット手続き型知識を持っていますが、均質化されたロジックを生成する傾向により、人間の個々の作成者のユニークで特異な実行プロセスが見えにくくなることがよくあります。この論文では、非構造化データからの手続き型ペルソナの計算による発見を調査します。これを実現するために、人気のある料理ビデオのトランスクリプトから抽出され、特定の作成者に明示的にマッピングされた、手続き的に調整された実行フロー グラフの新しいデータセットである ViralRecipesTrans を導入します。私たちは手続き型スタイロメトリーをグラフ学習およびプロセス発見タスクとして定式化し、基本的な二重性を明らかにしました。従来の語彙分類子が意味漏洩によってオーバーフィットするのに対し、離散トポロジカルメトリクスはクリエイターのワークフローの厳格な物理的制約をうまく捕捉します。この特徴付けに基づいて、私たちはフレームワークを新しい生成タスク、つまり、まだ見たことのない料理に対する作成者の正確な構造実行グラフを予測するタスクに拡張します。スタイル生成における、グローバルなマクロ計画とローカルな構造実行の間の基本的な二分法を明らかにします。私たちの結果は、少数ショット LLM は意味論的割り当てを支配しますが、永続的なマクロ計画の欠陥に悩まされるのに対し、私たちの構造化 2 段階モデル​​は剛直なマルコフ事前分布によって優れたトポロジー制御を実現することを示しています。プロシージャル生成へのアンサンブル アプローチは、両方の長所を組み合わせ、グローバルな意味論的推論とローカライズされたトポロジー フットプリントを動的に合成して、パーソナライズされたワークフローの検出と生成を自動化します。

原文 (English)

Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs

While large language models (LLMs) possess vast zero-shot procedural knowledge, their tendency to produce homogenized logic often obscures the unique, idiosyncratic execution processes of individual human creators. In this paper, we investigate the computational discovery of procedural personas from unstructured data. To achieve this, we introduce ViralRecipesTrans, a new dataset of procedurally aligned execution flow graphs extracted from popular culinary video transcripts and explicitly mapped to specific creators. We formulate procedural stylometry as a graph learning and process discovery task, revealing a fundamental duality: while traditional lexical classifiers overfit via semantic leakage, discrete topological metrics successfully capture the rigid physical constraints of a creator's workflow. Building upon this characterization, we extend our framework into a novel generative task--predicting a creator's exact structural execution graph for unseen dishes. We expose a fundamental dichotomy in style generation between global macro-planning and local structural execution. Our results demonstrate that few-shot LLMs dominate semantic assignment but suffer from persistent macro-planning deficits, whereas our structured two-stage model achieves superior topological control via rigid Markovian priors. Together, an ensemble approach to procedural generation combines the strengths from both sides, dynamically synthesizing global semantic reasoning with localized topological footprints to automate the discovery and generation of personalized workflows.

13:00 JSTLLM/生成AI

ResiSpec: 残差分布整形による複数候補の推測的サンプリングの強化

大規模言語モデル (LLM) サービスの効率は、自己回帰デコードの逐次的な性質によって基本的に制限されます。投機的デコーディング (SD) は、軽量のドラフト モデルを使用して将来のトークンを推測し、単一の並列前方パスで LLM によって検証されることでこれを軽減します。効率をさらに高めるために、複数候補スキームは多様な候補セットを提案し、トークンが受け入れられる可能性を高めます。ただし、これらのスキームは残留ドリフトによってボトルネックになっていることがわかります。残留ドリフトとは、最初の候補の拒否により、残留ターゲットの分布がドラフト モデルの予測から逸脱する現象です。この変化により、後続の候補が無効になり、システムにコストのかかる再サンプリングが強制されます。これを解決するために、我々は、検証中に提案分布を戦略的に改革して、残りのターゲット質量をドラフトモデルの信頼度の高い領域内に固定するフレームワークである ResiSpec を提案します。 ResiSpec は、出力の正確性を損なうことなく検証プロセスを数学的に再調整することで、候補の陳腐化を防ぎ、最先端の複数候補手法と比べて最大 1.92 倍の高速化を達成します。コードは https://github.com/Czzzk/Resispec で入手できます。

原文 (English)

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of autoregressive decoding. Speculative Decoding (SD) mitigates this by using a lightweight draft model to speculate future tokens, which are then validated by the LLM in a single parallel forward pass. To further boost efficiency, multi-candidate schemes propose diverse candidate sets to increase the likelihood of token acceptance. However, we show that these schemes are bottlenecked by Residual Drift: a phenomenon where the rejection of initial candidates causes the residual target distribution to diverge from the draft model's predictions. This shift renders subsequent candidates ineffective and forces the system into expensive resampling. To resolve this, we propose ResiSpec, a framework that strategically reforms the proposal distribution during verification to anchor the residual target mass within the draft model's high-confidence regions. By mathematically re-aligning the verification process without compromising output exactness, ResiSpec prevents candidate obsolescence and achieves up to 1.92$\times$ speedup over state-of-the-art multi-candidate methods. Code is available at https://github.com/Czzzk/Resispec.

13:00 JSTLLM/生成AIビジネス/資金調達

裁判官は何が変わったのかを知っておくべき:裁判官としての LLM 評価の妥当性を構築する

ジャッジとしての LLM の評価は、通常、表面の摂動に対する一致性とロバスト性によって評価されますが、信頼性によって構成の妥当性が確立されるわけではありません。評価者の構成の妥当性を 2 次元のプロファイルとして形式化します。不変性 S、構成を保持する編集の下で判定が変更されない確率、および構成の感度 R、構成を変更する最小限の編集の下で判断が変化する確率です。 S と R は独立しており、関連するすべての比較を保存するスカラー要約はないことを示します。 7 つの構成変更介入タイプと 5 つの登録のみのコントロールを使用して、7 人の裁判官と 4 つのドメインにわたってプロファイルを測定します。介入の方向はヒューマン アノテーターによって決定され、生成、検証、および判断は互いに素なモデル ファミリーに割り当てられます。一致した不変性 S >= 0.90 では、裁判官の平均値は S = 0.945 ですが、R = 0.319 になります。感度はスコープ編集と強度編集の間でも異なります。R_scope = 0.383 対 R_strength = 0.262、7 人のジャッジ全員で同じ符号を持つ +0.121 の差です。さらに 5 つのパブリック ラベル セットを監査し、表面のみの予測子がペア モードでラベルの 55% ~ 67% を再現し、これには MT-Bench 人間の投票の 67.4% が含まれることがわかりました。これらの結果は、高い判断者の合意が、評価対象の構成要素の変化に対する弱い感受性と共存し、不変性と感受性の共同報告と検証セット自体の監査を動機付ける可能性があることを示しています。

原文 (English)

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the probability that a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability that it changes under minimal construct-changing edits. We show that S and R are independent and that no scalar summary preserves all relevant comparisons. We measure the profile across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with intervention direction determined by human annotators and generation, verification, and judging assigned to disjoint model families. At matched invariance S >= 0.90, judges average S = 0.945 but R = 0.319. Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges. We further audit five public label sets and find that surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes. These results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.

13:00 JST研究/論文

線形計画法による因果秩序の部分同定

反事実クエリのノンパラメトリック (部分的) 識別は、通常、完全に指定された因果グラフに依存します。不完全なドメイン知識による設定を動機として、クエリ自体に本質的に含まれる構造的な仮定を活用することで、この要件に挑戦します。我々は、いかなる反事実の調査も、関連する変数に対してほとんど部分的なトポロジカルな順序付けを引き起こし、それによって明示的なクエリのパラメータ化が可能になり、識別タスクを線形プログラムに減らすことができることを示します。これにより、任意の反事実クエリとネストされた反事実クエリを制限することができます。私たちの研究は、もともと因果関係の確率のために開発された、Tian and Pearl (2000) の古典的な境界フレームワークの一般化として見ることができます。また、観察されたデータとクエリが暗黙的に示す順序の両方と互換性を持ちながら、境界を達成する構造的因果モデルを構築することによって、境界の \emph{厳密さ} を証明します。提案された境界手順の一般性と実用性の両方を評価するために、文献からいくつかのケーススタディを再検討し、入力因果グラフがない場合でも、導出された境界を使用して有益な洞察を得ることができる方法を示します。

原文 (English)

Partial Identification under Causal Orders by Linear Programming

Non-parametric (partial) identification of counterfactual queries typically relies on a fully specified causal graph. Motivated by settings with incomplete domain knowledge, we challenge this requirement by leveraging structural assumptions that are inherently implied by the query itself. We show that any counterfactual inquiry induces a, mostly partial, topological ordering over relevant variables, which, in turn, enables an explicit query parametrisation reducing the identification task to a linear program. This allows bounding arbitrary counterfactual and nested counterfactual queries. Our work can be viewed as a generalisation of the classical bounding framework of Tian and Pearl (2000), originally developed for probabilities of causation. We also prove the \emph{tightness} of our bounds by constructing structural causal models that attain the bounds whilst being compatible with both the observed data and the query-implied order. To assess both the generality and practical utility of the proposed bounding procedure, we revisit several case studies from the literature, demonstrating how the derived bounds can be used to yield informative insights even in the absence of an input causal graph.

13:00 JST研究/論文

電気自動車の充電負荷に対する行動に基づくオンライン確率予測手法

電気自動車 (EV) の充電負荷は、強い動作の不均一性と時間的変動を示し、進化する動作条件下でのオンラインの確率的予測に重大な課題をもたらします。特に、持続的な充電パターンはステーション間で大幅に異なる可能性がありますが、最近の動作の変化により、基礎となる負荷分散が継続的に変化する可能性があります。この論文では、永続的なステーション固有のパターンと最近の行動の変化を明示的に特徴付ける、行動に基づくオンラインの確率的予測フレームワークを提案します。デュアルタイムスケールの動作表現は、長期的な充電特性と最近の動作状態を区別し、その偏差を定量化するために構築されています。これらの行動の変化は、ドリフトを意識した予測適応をガイドするためにさらに意味論的にエンコードされ、同時に遅延フィードバック メカニズムにより、異なる予測期間にわたって観測が利用可能になったときに、時間的に一貫したオンライン更新が保証されます。現実世界の異種混合充電ステーション 10 か所での実験では、提案された方法が、予測精度と確率的信頼性において、従来の予測モデルやコンセプトドリフトを意識したオンラインベースラインよりも一貫して優れていることが実証されました。 1 時間先の予測の場合、提案された方法は、対応する最良のベースラインと比較して、MSE 損失とピンボール損失をそれぞれ 15.3\% と 17.8\% 削減します。 4 時間先の予測では、改善はさらにそれぞれ 16.8\% と 22.6\% に達し、進化する充電動作と拡張された予測期間の下で一貫したパフォーマンスの向上を示しています。

原文 (English)

A Behavior-Guided Online Probabilistic Forecasting Method for Electric vehicle Charging Loads

Electric vehicle (EV) charging loads exhibit strong behavioral heterogeneity and temporal variability, posing significant challenges for online probabilistic forecasting under evolving operating conditions. In particular, persistent charging patterns may differ substantially across stations, while recent behavioral changes can continuously alter the underlying load distributions. This paper proposes a behavior-guided online probabilistic forecasting framework that explicitly characterizes persistent station-specific patterns and recent behavioral changes. A dual-timescale behavior representation is constructed to distinguish long-term charging characteristics from recent behavioral states and quantify their deviations. These behavioral changes are further semantically encoded to guide drift-aware forecasting adaptation, while a delayed-feedback mechanism ensures temporally consistent online updates when observations become available across different forecasting horizons. Experiments on ten heterogeneous real-world charging stations demonstrate that the proposed method consistently outperforms conventional forecasting models and concept-drift-aware online baselines in forecasting accuracy and probabilistic reliability. For 1-h-ahead forecasting, the proposed method reduces MSE and Pinball loss by 15.3\% and 17.8\%, respectively, over the corresponding best baselines. For 4-h-ahead forecasting, the improvements further reach 16.8\% and 22.6\%, respectively, demonstrating consistent performance gains under evolving charging behaviors and extended forecasting horizons.

13:00 JST研究/論文

複雑な状態の伝播に対するマハラノビスベースのマルチヘッド アテンション

この論文では、標準の内積を \textbf{マハラノビス距離ベースの RBF カーネル} に置き換える新しいアテンション メカニズムである \textbf{マハラノビスベースのマルチヘッド アテンション} (MHA-CSP) を提案します。これは、パラメータ数を増やすことなく、無限次元の特徴空間でアテンションを効果的に計算します。重要なのは、マハラノビス距離の正定性により \textbf{ツリー アテンションの直接構築} が可能になることです。アテンション スコアは累積された距離から直接構築され、エッジ指数の対数和を減算して生の距離を修正する LogSumExp 補正が使用されます。さらに、マルチヘッド マハラノビス距離行列自体を再利用して \textbf{アテンション メッシュ メカニズム} を構築し、精度とトレーニング効率を同時に高めるクロスヘッド カーネル コラボレーションを可能にします。広範な実験により、わずか 119,000 個のパラメーターと \textbf{最終隠れ状態のみに適用される教師強制} のみを使用した MHA-CSP が、長いシーケンスの状態追跡タスクにおいて同一条件下で最初からトレーニングされた Transformer および GCN ベースラインよりも一貫して優れていることが実証されました。これらのベースラインは高密度の注意やグラフの伝播に依存していますが、MHA-CSP は、マハラノビスベースの注意を利用した合成距離調整と、CSP バックボーンから継承した効率的な情報バイパスによって、堅牢な構造化推論を実現します。この結果は、記号構造を捕捉する際の協調的なマルチヘッド修正による複素数値状態伝播の有効性を強調し、構造化推論のための新しい効率とパフォーマンスのトレードオフを確立します。

原文 (English)

Mahalanobis-Based Multi-Head Attention for Complex State Propagation

In this paper, we propose \textbf{Mahalanobis-Based Multi-Head Attention} (MHA-CSP), a novel attention mechanism that replaces the standard dot-product with a \textbf{Mahalanobis distance-based RBF kernel}, which effectively computes attention in an infinite-dimensional feature space without increasing the parameter count. Crucially, the positive definiteness of the Mahalanobis distance enables a \textbf{direct construction of Tree Attention}: attention scores are built directly from accumulated distances, with a LogSumExp correction that rectifies the raw distance by subtracting the log-sum of edge exponentials. Moreover, the multi-head Mahalanobis distance matrices are themselves repurposed to construct an \textbf{attention meshing mechanism}, enabling cross-head kernel collaboration that simultaneously boosts accuracy and training efficiency. Extensive experiments demonstrate that MHA-CSP, with only 119K parameters and \textbf{teacher forcing applied exclusively at the final hidden state}, consistently outperforms Transformer and GCN baselines trained from scratch under identical conditions on long-sequence state tracking tasks. While these baselines rely on dense attention or graph propagation, MHA-CSP achieves robust structured reasoning via synthetic distance rectification---powered by Mahalanobis-based attention---and efficient information bypass inherited from the CSP backbone. This result highlights the effectiveness of complex-valued state propagation with collaborative multi-head rectification in capturing symbolic structures, establishing a new efficiency-performance trade-off for structured reasoning.

13:00 JSTLLM/生成AI

HMGCLIP: 電子商取引表現学習のための異種多粒度対照学習

最近のマルチモーダル大規模言語モデル (MLLM) は、一般的な製品の理解が進んでいますが、製品情報をグローバルな埋め込みに暗黙的にエンコードするため、きめの細かい属性を取得する能力が制限されています。この制限により、視覚的に類似した製品間の微妙な材質の違いの区別など、正確な属性の識別が必要なタスクのパフォーマンスが妨げられます。この課題に対処するために、私たちは統合マルチモーダル埋め込みフレームワークである HMGCLIP を提案します。異種ハイパーグラフを構築することで、ハイパーグラフ トポロジを活用して構造を認識したハード ネガをマイニングし、リレーション レベルとハイパーエッジ レベルの両方で複数粒度のセマンティクスを調整します。この設計により、粒度の細かいダウンストリーム タスクと粗粒度のダウンストリーム タスクの両方の属性証拠を動的に融合する二重粒度の推論メカニズムが可能になります。さらに、将来のベンチマークを容易にするために、包括的なきめの細かい電子商取引データセットをリリースします。この新しいデータセットと公開 MAVE ベンチマークに関する広範な実験により、HMGCLIP が強力なマルチモーダル エンコーダー、MLLM、電子商取引ベースラインよりも優れたパフォーマンスを発揮することが示され、HMGCLIP の優位性が実証されました。

原文 (English)

HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning

Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. This limitation hinders performance in tasks requiring precise attribute discrimination, such as distinguishing subtle material differences among visually similar products. To address this challenge, we propose HMGCLIP, a unified multimodal embedding framework. By constructing a heterogeneous hypergraph, we leverage hypergraph topology to mine structure-aware hard negatives and align multi-granular semantics at both relation and hyperedge levels. This design enables a dual-granularity inference mechanism that dynamically fuses attribute evidence for both fine-grained and coarse-grained downstream tasks. Furthermore, we release a comprehensive fine-grained e-commerce dataset to facilitate future benchmarking. Extensive experiments on this new dataset and the public MAVE benchmark show that HMGCLIP outperforms strong multimodal encoders, MLLMs, and e-commerce baselines, validating the superiority of HMGCLIP.

13:00 JST研究/論文

好み調整可能な異種機敏な地球観測衛星スケジューリングのための強化学習に基づく進化ポリシーの最適化

異種機敏な地球観測衛星 (AEOS) のスケジューリングには、衛星に依存する可視ウィンドウ、姿勢操作要件、エネルギー消費、および搭載ストレージの制約の下で、タスクの選択、衛星の割り当て、観測順序を設定する必要があります。衛星は軌道へのアクセス、操縦能力、ペイロードリソースが異なるため、同じタスクでもプラットフォームごとに異なる実行可能ウィンドウ、移行コスト、リソース消費パターンを持つ可能性があり、統一モデリングと効率的な最適化の難易度が高くなります。この問題に対処するために、この論文では、好みに応じて調整可能な重み付き目標を備えた異種 AEOS スケジューリングのための進化的ポリシー最適化フレームワークを提案します。モデリング層では、割り当てベースの間接エンコーディングがデコーダベースの等価コスト評価と組み合わされて、衛星依存の制約を維持しながら、タスクゲイン、エネルギー節約、負荷バランスを解釈可能なスカラーユーティリティに統合します。最適化層では、スケジュールのデコード、母集団ベースの検索、およびオンラインのアクターと批評のオペレーター制御が分離されているため、強化学習はスケジュールを直接構築するのではなく、高レベルの検索オペレーターを選択します。このフレームワークに基づいて、限られた機能評価予算の下で地球規模の探索、実現可能性の回復、および局所的な改良を調整するために、強化学習支援オペレーター選択ミーム進化アルゴリズム (RLOSMEA) が開発されます。さまざまな異種 AEOS シナリオの実験では、RLOSMEA が代表的なメタヒューリスティック ベースラインよりも全体的に重み付けされたユーティリティが高く、より安定した収束を達成することが示されています。感度と学習行動の分析により、提案された方法の堅牢性と強化学習に基づくオペレーター選択の有効性がさらに確認されます。

原文 (English)

Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling

Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite assignment, and observation sequencing under satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard storage constraints. Since satellites differ in orbital access, maneuvering capability, and payload resources, the same task may have different feasible windows, transition costs, and resource-consumption patterns on different platforms, which increases the difficulty of unified modeling and efficient optimization. To address this problem, this paper proposes an evolutionary policy optimization framework for heterogeneous AEOS scheduling with preference-adjustable weighted objectives. In the modeling layer, assignment-based indirect encoding is combined with decoder-based equivalent-cost evaluation to retain satellite-dependent constraints while integrating task gain, energy saving, and load balance into an interpretable scalar utility. In the optimization layer, schedule decoding, population-based search, and online actor-critic operator control are decoupled, so that reinforcement learning selects high-level search operators rather than constructing schedules directly. Based on this framework, a reinforcement-learning-assisted operator-selection memetic evolutionary algorithm (RLOSMEA) is developed to coordinate global exploration, feasibility recovery, and local refinement under a limited function-evaluation budget. Experiments on different heterogeneous AEOS scenarios show that RLOSMEA achieves higher overall weighted utility and more stable convergence than representative metaheuristic baselines. Sensitivity and learning-behavior analyses further confirm the robustness of the proposed method and the effectiveness of reinforcement-learning-guided operator selection.

13:00 JST研究/論文

アジャイル衛星による海上移動目標観測スケジュールのための暗黙的 Q ラーニング ブートストラップ アリ コロニー最適化

機敏な地球観測衛星による海洋移動目標の観測スケジュール設定は、シーケンスに依存する動的な組み合わせ最適化問題です。海面のターゲットは継続的に移動するため、実現可能な観測ウィンドウはターゲットの動きや衛星の軌道形状によって異なります。スケジューラは、タイム ウィンドウ、姿勢操作、搭載リソース、およびクラウドの影響を受ける可用性の制約の下で、タスクの選択、衛星の割り当て、観測ウィンドウの選択、および観測の順序を共同で決定する必要があります。この論文は、複数の衛星による海洋移動目標の観測スケジューリングのために、IQACO と呼ばれる暗黙的な Q 学習ブートストラップ型アリ コロニー最適化手法を提案します。 IQACO は、タスク選択ポリシーを直接学習するのではなく、オフラインの暗黙的な Q 学習モジュールをアリのコロニーの建設的な最適化に組み込み、フェロモン因子、ヒューリスティック因子、蒸発速度を適応的に調整します。コンパクトな検索状態表現により、フェロモン分布、現在および過去最高のソリューション品質、反復の進行状況がキャプチャされます。オンライン スケジューリング中に、アリのコロニーの最適化によって実行可能な観察シーケンスが構築され、学習されたポリシーによって現在の検索状態に応じて探索と活用が規制されます。異なるスケールと衛星構成による 14 のシナリオでの実験では、IQACO がすべてのシナリオで最高の平均観測利益を得て、従来のアリのコロニー最適化の結果を 3.40\% ~ 9.40\% 改善し、収束を加速し、さまざまな目的の重み設定の下で安定性を維持することが示されました。これらの結果は、オフライン値学習が、制約のある海上移動目標の観測スケジュールに対して効果的な適応型探索制御メカニズムを提供することを示しています。

原文 (English)

Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites

Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, onboard-resource, and cloud-affected availability constraints. This paper proposes an implicit Q-learning-bootstrapped ant colony optimization method, termed IQACO, for multi-satellite maritime moving-target observation scheduling. Rather than directly learning a task-selection policy, IQACO embeds an offline implicit Q-learning module into constructive ant colony optimization to adaptively adjust the pheromone factor, heuristic factor, and evaporation rate. A compact search-state representation captures pheromone distribution, current and historical-best solution quality, and iteration progress. During online scheduling, ant colony optimization constructs feasible observation sequences, while the learned policy regulates exploration and exploitation according to the current search state. Experiments on 14 scenarios with different scales and satellite configurations show that IQACO obtains the highest mean observation benefit in every scenario, improves the result of conventional ant colony optimization by 3.40\%--9.40\%, accelerates convergence, and remains stable under different objective-weight settings. These results demonstrate that offline value learning provides an effective adaptive search-control mechanism for constrained maritime moving-target observation scheduling.

13:00 JSTLLM/生成AIエージェント研究/論文

PeakBench: LLM エージェントでのリソース認識ツール呼び出しのベンチマーク

LLM エージェントは、複数のツールを呼び出してタスクを解決することが増えています。この場合、低レイテンシには並列実行が不可欠ですが、安全に管理することが困難です。既存のエージェント ベンチマークは主に、ツールの選択、引数の生成、およびほとんどがシリアル実行の下でのエンドツーエンドの成功を評価しており、有効な並列化やリソースに制約のあるスケジューリングはほとんど無視されています。このスケジューリング次元の欠落により、実質的な障害モードが作成されます。つまり、シリアル実行は安全ですが遅いのに対し、リソースに依存しない並列実行は高速ですが、回避可能なリソース オーバーフローが発生する傾向があります。このギャップに対処するために、実行ベースの依存関係アノテーションと測定されたリソース プロファイルを備えた実行可能なマルチツール ワークフローのベンチマークである PeakBench を導入します。このようなワークフローを評価する際の中心的な課題は、原因の特定です。失敗や非効率は、不適切な依存関係の計画、リソースに制約のあるスケジュール、またはその両方によって発生する可能性があります。 PeakBench は、各次元に専用のメトリクスを使用して、論理的な計画を物理的なスケジュールから切り離す 2 つの部分からなる評価フレームワークでこの課題に対処します。このフレームワークを使用して、リソースの制約の下では、強力な論理計画が確実に安全または効率的な実行に変換されないことを示します。さらに、リソース情報を公開すると、回避可能なオーバーフローが減少し、リソースの使用率が向上し、PeakBench がリソースを認識したエージェントの動作を診断するための有用なテストベッドになることを示します。コードは https://github.com/Czzzk/Staggering-the-Peaks で入手できます。

原文 (English)

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential for low latency but difficult to manage safely. Existing agent benchmarks primarily evaluate tool selection, argument generation, and end-to-end success under mostly serial execution, largely overlooking valid parallelization and resource-constrained scheduling. This missing scheduling dimension creates a practical failure mode: serial execution is safe but slow, while resource-agnostic parallel execution is fast but prone to avoidable resource overflows. To address this gap, we introduce PeakBench, a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles. A central challenge in evaluating such workflows is attribution: failures and inefficiencies may arise from incorrect dependency planning, poor resource-constrained scheduling, or both. PeakBench addresses this challenge with a two-part evaluation framework that disentangles logical planning from physical scheduling, with dedicated metrics for each dimension. Using this framework, we show that strong logical planning does not reliably translate into safe or efficient execution under resource constraints. We further show that exposing resource information can reduce avoidable overflows and improve resource utilization, making PeakBench a useful testbed for diagnosing resource-aware agent behavior. Code is available at https://github.com/Czzzk/Staggering-the-Peaks.

13:00 JSTLLM/生成AIGPT / ChatGPT

生理学的に安全な臨床言語モデルのための神経象徴的アライメント

臨床 LLM は、事実としてはもっともらしいが、生理学的に安全ではない推奨事項を生成できます。私たちは、テキストのみの監視ではなく、構造化された生理学的知識に基づくグラウンディング優先の最適化によって安全調整を改善できるかどうかを調査します。方法: 我々は、847,000 ノードの生物医学知識グラフ上で 7B 臨床 LLM と HGNN ベースの生理学的世界モデルを結合するトレーニング時間フレームワークであるニューロシンボリック アライメントを提案します。候補応答は、ホメオスタシス制約、マルチホップ パスの妥当性、および薬物相互作用ペナルティを使用してスコア付けされ、結果のランキングによってポリシー上の ORPO 更新が反復的に行われます。評価は、生成的臨床推論における生理学的制約違反の 2,500 のシナリオのベンチマークである臨床安全性ベンチマーク (CSB) で実行されます。結果: ORPO と比較して、提案された方法は CSS を 69.5% から 90.8% (+21.3 pp) 改善し、盲検サブセットで医師が評価した HR を 14.1% から 5.1% に減少させ、DID を 72.8% から 91.6% に改善しました。これらの利益は、HGNN に依存しないルール エンジン安全性スコアによって裏付けられています (RSS: 86.4%、ORPO に対して +21.2 pp、CSS との一致率は r=0.97)。また、このメソッドは、パラメーターの 10 倍の不利な点にもかかわらず、すべての安全性メトリクスで GPT-4 (5 ショット) を上回り、推論時の自己修正パイプライン (SFT+SelfCorrect) のパフォーマンスを 11.4 pp CSS 上回っています。合成 EHR スタイルのノイズの下では、84.2% の CSS が保持されます。アブレーション分析は、HGNN スコアリング (-16.2 pp) と反復トレーニング (-11.5 pp) が主な要因であることを示しています。 200 人の臨床医ラベルに対する PhysioScore キャリブレーションにより、ECE = 0.038 および kappa = 0.91 が得られました。結論: トレーニング時の生理学的グラウンディングは、管理された評価の下で、オープンウェイト臨床 LLM において測定可能で独立して検証可能な安全性の向上をもたらします。これらの利益が展開設定に反映されるかどうかを判断するには、実際の臨床データに対する外部検証が必要です

原文 (English)

Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models

Clinical LLMs can generate recommendations that are factually plausible yet physiologically unsafe. We investigate whether safety alignment can be improved by grounding preference optimization in structured physiological knowledge rather than text-only supervision. Methods: We propose Neurosymbolic Alignment, a training-time framework that couples a 7B clinical LLM with an HGNN-based Physiological World Model over an 847K-node biomedical knowledge graph. Candidate responses are scored using homeostatic constraints, multi-hop path plausibility, and drug-interaction penalties, and the resulting rankings drive iterative on-policy ORPO updates. Evaluation is performed on the Clinical Safety Benchmark (CSB), a 2,500-scenario benchmark for physiological constraint violations in generative clinical reasoning. Results: Relative to ORPO, the proposed method improves CSS from 69.5% to 90.8% (+21.3 pp), reduces physician-evaluated HR from 14.1% to 5.1% on the blinded subset, and improves DID from 72.8% to 91.6%. These gains are corroborated by an HGNN-independent Rule-Engine Safety Score (RSS: 86.4%, +21.2 pp over ORPO; r=0.97 concordance with CSS). The method also exceeds GPT-4 (5-shot) on all safety metrics despite a 10x parameter disadvantage, and outperforms an inference-time self-correction pipeline (SFT+SelfCorrect) by 11.4 pp CSS. Under synthetic EHR-style noise, 84.2% CSS is retained. Ablation analysis shows that HGNN scoring (-16.2 pp) and iterative training (-11.5 pp) are the dominant contributors. PhysioScore calibration against 200 clinician labels yielded ECE = 0.038 and kappa = 0.91. Conclusion: Training-time physiological grounding produces measurable and independently verifiable safety improvements in open-weight clinical LLMs under controlled evaluation. External validation on real clinical data is required to determine whether these gains transfer to deployment settings

13:00 JST研究/論文

集団的イノベーションのための適応型伝送プログラムの発見

人間の集合知は、誰が誰と、何を、どのように、いつ共有するかという伝達プロセスに依存します。これらのプロセスは個人の認知から生まれますが、意図的なトップダウンのプロトコルによって指示されることもあります。これまでの研究では、誰が誰といつ共有するかを変えることにより、主にネットワーク構造のレンズを通して、送信が集団的な結果をどのように形作るかを研究してきました。しかし、ネットワークは状態に依存しません。エージェントが知っていることや集団の状態に基づいて送信を条件付けることはできません。ここでは、エージェントと集団の状態に基づいて情報とリソースをルーティングする状態認識プログラムとして送信プロトコルを形式化し、LLM に基づく進化的探索を使用して集団検出タスクで効果的なプロトコルを設計します。進化したプロトコルにより、文献に基づく標準ベースラインと比較して総合的なパフォーマンスが最大 37% 向上します。アブレーションにより、状態認識がこの利点を促進することが確認されました。ネットワーク トポロジとタイミングを維持しながらコンテンツ依存性を除去すると、パフォーマンスの向上がなくなります。私たちは、進化したプロトコルがドメインのバリエーションやエージェント集団を越えて移行することも発見しました。これらの結果は、効果的で一般化可能な伝送プロトコルをインシリコで発見できることを実証し、人間の集合知を強化する調整インフラストラクチャの AI 支援設計への道を示唆しています。

原文 (English)

Discovering Adaptive Transmission Programs for Collective Innovation

Human collective intelligence depends on transmission processes: who shares what with whom, how, and when. While these processes emerge from individual cognition, they can also be directed by deliberate top-down protocols. Prior work has studied how transmission shapes collective outcomes primarily through the lens of network structure, varying who shares with whom and when. But networks are state-agnostic: they cannot condition transmission on what agents know or on the state of the collective. Here, we formalize transmission protocols as state-aware programs that route information and resources based on agent and collective states, and we use LLM-guided evolutionary search to design effective protocols in a collective discovery task. Evolved protocols increase collective performance over standard baselines from the literature by up to 37%. Ablations confirm that state-awareness drives this advantage: removing content-dependence while preserving network topology and timing eliminates performance gains. We find that evolved protocols also transfer across domain variations and agent populations. These results demonstrate that effective and generalizable transmission protocols can be discovered in silico, suggesting a path toward AI-assisted design of coordination infrastructure that enhances human collective intelligence.

13:00 JSTLLM/生成AIエージェント

「Must」が「Maybe」になるとき: LLM エージェントのワークフローにおける制約の弱体化

大規模言語モデル (LLM) エージェントは、複数の役割と複数ステージのワークフローを通じて複雑なタスクを調整します。上流の状態は、要約、計画、チケット、思い出、引き継ぎメモなどの中間言語成果物に繰り返し変換され、そこから下流コンポーネントが動作します。アクション制約状態の場合、トピックの保持は不十分です。アーティファクトは、実行前に解決する必要がある要件から、次のアクションを通知するだけの情報に変更しながら、未解決の条件について言及する可能性があります。私たちは、動作状態の保存としてのこのアクションを拘束する役割を研究します。各ソース状態には明示的な前提条件、権限、フォールバック、実行結果があるため、セーフティ ブロッカーは制御されたインスタンスを提供します。正しい上流識別を条件とし、ハンドオフ変換を変更し、結果として得られるアーティファクトに限定してエグゼキュータを評価します。 1,296 の制御された合成エピソードにわたって、直接ハンドオフ制御はすべてのブロッカーを保持しますが、圧縮、計画の同化、収束、所有権の延期、および先行置換は、拘束状態を繰り返し警告または拘束力のない考慮事項に変えます。通常のハンドオフ圧縮では、100.0% の非アクティブ化と 54.2% の禁止アクションが生成されます。 4 つの状態フィールドをすべて復元すると、保存が 100.0% に上昇し、禁止されたアクションが 0.0% に減少します。修正されたアーティファクトの介入により、保存と封じ込めがさらに分離されます。下流の検証により、禁止されたアクションが排除されますが、アーティファクトの非アクティブ化は 95.3% のままです。これらの結果は、情報抽出とアクションの間の状態伝達の失敗を特定します。ハンドオフ変換は、下流のアクションに対する制約を弱めながら、状態の内容を保持できます。セマンティックな可用性は、運用上の保存を保証するものではありません。

原文 (English)

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and handoff notes, from which downstream components act. For action-constraining state, topical retention is insufficient: an artifact may mention an unresolved condition while changing it from a requirement that must be resolved before execution into information that may merely inform the next action. We study this action-binding role as operational state preservation. Safety blockers provide a controlled instance because each source state has an explicit prerequisite, authority, fallback, and execution consequence. We condition on correct upstream identification, vary the handoff transformation, and evaluate an executor restricted to the resulting artifact. Across 1,296 controlled synthetic episodes, direct-handoff controls preserve every blocker, whereas compression, plan assimilation, convergence, ownership deferral, and precedent substitution repeatedly turn binding state into caveats or non-binding considerations. Normal handoff compression produces 100.0% deactivation and 54.2% forbidden action. Restoring all four state fields raises preservation to 100.0% and reduces forbidden action to 0.0%. Fixed-artifact interventions further separate preservation from containment: downstream verification eliminates forbidden action while artifact deactivation remains 95.3%. These results identify a state-transmission failure between information extraction and action. Handoff transformations can retain state content while weakening its constraints on downstream action. Semantic availability does not guarantee operational preservation.

13:00 JSTLLM/生成AIエージェント

EviDx: スキャフォールド LLM エージェントを使用した証拠を意識したアクティブ診断

臨床診断は、臨床医が証拠を入手し、競合する仮説を更新し、入手可能な証拠が診断に十分であるかどうかを判断する、積極的な証拠探索プロセスです。しかし、大規模言語モデル (LLM) を中心に構築された多くの医療診断システムは依然として静的な症例対回答の予測として診断を定式化しており、証拠取得のサポートは限られています。 Agentic LLM は、ツールの使用と中間の診断軌跡を通じて動的な代替手段を提供しますが、既存のシステムでは、実行時に患者の証拠がどのように公開され、足場が築かれ、制御されるべきかが十分に規定されていないことがよくあります。患者固有の診断環境と臨床診断足場および観察者主導のランタイム ハーネスを組み合わせる、証拠を意識したアクティブ診断フレームワークである EviDx を紹介します。 EviDx では、$\mathcal{E}$-Synthesis が生の臨床症例からインタラクティブな環境を構築します。足場は、役割に特化したエージェント、証拠ツール、進化する証拠状態を組織します。そしてハーネスは不確実性と証拠の範囲を追跡することによって診断の終了を制御します。 3 レベルの評価ピラミッドにより、実行の堅牢性、推論のダイナミクス、および診断結果が評価されます。実験では、EviDx がモデルに依存する機能の境界を明らかにしながら、診断パフォーマンスとプロセスの安定性を向上させることが示されています。

原文 (English)

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

Clinical diagnosis is an active evidence-seeking process in which clinicians acquire evidence, update competing hypotheses, and decide when the available evidence is sufficient for diagnosis. Yet many medical diagnosis systems built around large language models (LLMs) still formulate diagnosis as static case-to-answer prediction, with limited support for evidence acquisition. Agentic LLMs offer a dynamic alternative through tool use and intermediate diagnostic trajectories, but existing systems often under-specify how patient evidence should be exposed, scaffolded, and controlled at runtime. We introduce EviDx, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness. In EviDx, $\mathcal{E}$-Synthesis constructs interactive environments from raw clinical cases; the scaffold organizes role-specialized agents, evidence tools, and evolving evidence states; and the harness regulates diagnostic termination by tracking uncertainty and evidence coverage. A 3-level evaluation pyramid assesses execution robustness, reasoning dynamics, and diagnostic outcomes. Experiments show that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.

13:00 JSTLLM/生成AIエージェント

大規模言語モデル エージェントのツール作成と使用の共同最適化

ツールで拡張された言語モデルは、人間がわざわざ作成した API によって制限されます。既存のツール作成システムは、推論時にフリーズした LLM をプロンプトすることでこの問題にパッチを当て、ツールを作成するモデルとそのツールを使用するモデルを分離したままにし、生成されるスキーマが呼び出し可能なスキーマであるという信号を与えません。私たちは、単一のポリシー内でツールの作成とツールの使用を共同でトレーニングする強化学習フレームワークである SMITH (Schema-grounded Multi-task Iterative Tool Honing) を提案します。各ロールアウトは、構築タスク (いくつかの例からツールを作成する) または使用タスク (保留された質問に対してプールされたツールを呼び出す) のいずれかです。 3 つの個別の報酬軸がスキーマ、コード、結果の障害を個別に捕捉するため、各障害モードが独自の勾配に寄与します。正確なベリファイアを使用して 13 の手続き推論タスクについて SMITH とトレーニングを受けた 4B Qwen3 は、保留されたタスクで 79.8 のマクロ平均精度に達しました。これは、評価されたすべてのメソッドの中で最高であり、トレーニングを受けていない 30B-A3B ツール作成者を上回っています。また、視覚的または表形式のトレーニング データがない場合でも、TabMWP-Hard では 40.4、ドメイン外 GQA では 42.6 (同一バックボーン推論時間の最良のベースラインを +7.6 上回ります) に達します。当社の 4B モデルによって作成されたツールも、同じ推論タスクの下で LFM-2.5-350M および Qwen3-30B-A3B のパフォーマンスを向上させました。

原文 (English)

Joint Optimization of Tool Creation and Use for Large Language Model Agents

Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can invoke. We propose SMITH (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that jointly trains tool creation and tool use inside a single policy. Each rollout is either a build task (write a tool from a few examples) or a use task (invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained with SMITH on 13 procedural reasoning tasks with exact verifiers reaches 79.8 macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an untrained 30B-A3B tool-writer. It also reaches 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA (+7.6 over the best same-backbone inference-time baseline), without any visual or tabular training data. Tools written by our 4B models also lifted the performance of LFM-2.5-350M and Qwen3-30B-A3B under same reasoning tasks.

13:00 JSTLLM/生成AI

PhysMLLM: 画像とビデオの統合参照セグメンテーションとグラウンデッド推論のための空間事前分布

ビデオ マルチモーダル大規模言語モデルは、言語ガイド付きビデオ セグメンテーションをサポートしていますが、多くの場合、ジッター、ドリフト、アイデンティティ スイッチなどの時空間の不一致が見られます。これらの失敗は、ターゲットが部分的に隠れている場合、または同様のオブジェクトが近くに表示される場合により一般的です。考えられる理由の 1 つは、現在のトレーニングに明示的な空間事前分布が不足しているため、安定した空間的同一性と形状を長期にわたって維持することが困難であることが考えられます。我々は、物理学にインスピレーションを得た空間連続事前分布をビデオ MLLM に注入するトレーニング段階の事前注入アーキテクチャである PhysMLLM を紹介します。 PhysMLLMs は、トレーニング中に生徒のグローバル視覚表現を凍結された教師モデルと調整することで、より安定したオブジェクト中心の表現を促進するように設計されています。私たちのコアメカニズムである Global Representation Prior Alignment (REPA-Global) は、オフライン埋め込みキャッシュとスケジュールされた抽出計画を使用して、凍結された DINOv2 教師からグローバル視覚表現を抽出します。この設計では、推論を変更せずに維持し、推論時間のコストを追加しません。複数のビデオ ベンチマークにわたって、PhysMLLM はビデオ セグメンテーション マスクの品質とフレーム間の一貫性を向上させ、小さなターゲット、高速モーション、オクルージョン、ディストラクタ、および推論クエリを含む困難なケースで大きな利益をもたらします。単一フレーム参照画像セグメンテーションと代表的な一般的な VLM ベンチマークでは、PhysMLLM は同等のパフォーマンスを維持し、注入された空間事前分布が画像レベルの基礎や一般的なマルチモーダル機能を損なうことなくビデオの一貫性を向上させることを示しています。これらの結果は、物理学にヒントを得た空間事前注入により、一般的な能力を維持しながら時間的安定性を向上できることを示唆しています。コードは https://github.com/tusu-code/20260121-icml2026-2.git で入手できます。

原文 (English)

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches. These failures are more common when targets are partly hidden or when similar objects appear nearby.One likely reason is that current training lacks explicit spatial priors, which makes it difficult to maintain stable spatial identity and shape over time. We present PhysMLLMs, a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs. PhysMLLMs is designed to encourage more stable object-centered representations by aligning the student global visual representation with a frozen teacher model during training. Our core mechanism, Global Representation Prior Alignment (REPA-Global), distills global visual representations from a frozen DINOv2 teacher using an offline embedding cache and a scheduled distillation plan. This design keeps inference unchanged and does not add inference time cost. Across multiple video benchmarks, PhysMLLMs improves video segmentation mask quality and cross-frame consistency, with larger gains on challenging cases involving small targets, fast motion, occlusion, distractors, and reasoning queries. On single-frame referring image segmentation and representative general VLM benchmarks, PhysMLLMs maintains comparable performance, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability. These results suggest that physics-inspired spatial prior injection can improve temporal stability while preserving general capability. The code is available at https://github.com/tusu-code/20260121-icml2026-2.git.

13:00 JSTエージェントロボティクス

ピボット アンド ステーション マルチエージェント パス検索: 解決可能性、複雑性、およびアルゴリズム

自動高密度保管システム (倉庫、ロボット駐車場、プラント物流など) では、エージェントのフリートが希少なタスククリティカルなリソースを移動し、その後の業務を妨げることなく駐車する必要があります。ピボット アンド ステーション マルチエージェント パス ファインディング (PS-MAPF) を導入します。これは、タスクを割り当てられたエージェントのサブセットが、フリート全体が匿名ステーションで終了する前に、交換可能なピボットのセット (ワークステーションなど) の 1 つを訪問する必要がある MAPF の変種であり、ステーションごとに 1 つのエージェントが存在します。解決可能性を完全に特徴付けます。2 エッジ接続グラフ上のすべてのインスタンスは解決可能であり、任意の接続されたグラフ上では、占有されていない頂点の数に対する構造的有効距離の尺度が必要十分条件を与えます。ステーション メイクスパンまたはステーション フロータイムを最小限に抑えることは、単一ピボットではすでに NP 困難であることを証明します。完全なベースライン、SAT ベースの最適ソルバー、ピボット優先計画 (PPP) の 3 つのアルゴリズムを紹介します。ピボット優先計画 (PPP) は、ベースラインを大幅に下回るメイクスパンとフロータイムでベンチマーク インスタンスの 74 ~ 89% を最後に解決します。

原文 (English)

Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms

Automated high-density storage systems (warehouses, robotic parking, plant logistics, etc.) require fleets of agents to move through scarce task-critical resources and then park without obstructing future operations. We introduce Pivot-and-Station Multi-Agent Path Finding (PS-MAPF), a MAPF variant in which a subset of tasked agents must each visit one of a set of interchangeable pivots (e.g., workstations) before the entire fleet terminates at anonymous stations, one agent per station. We characterize solvability completely: every instance on a 2-edge-connected graph is solvable, and, on arbitrary connected graphs, a structural effective-distance measure relative to the number of unoccupied vertices gives a necessary and sufficient condition. We prove that minimizing station-makespan or station-flowtime is NP-hard already with a single pivot. We present three algorithms, a complete baseline, a SAT-based optimal solver, and Pivot-Prioritized Planning (PPP), the last solving 74-89% of benchmark instances with makespan and flowtime orders of magnitude below the baseline.

13:00 JST研究/論文

学生の能力評価のための支援介入の因果モデリング

教育者が個人のニーズを特定し、的を絞った介入を設計し、教育戦略の有効性を評価できるようにするには、生徒の能力を正確に評価することが不可欠です。経験的評価手順は通常、項目反応理論などの心理測定モデルに基づいており、生徒の能力レベルを評価タスクのパフォーマンスに関連付けます。この論文では、教育評価に構造的因果モデリングのアプローチを採用し、確率的信念の更新を超えて、介入的推論と反事実的推論を明示的にサポートするフレームワークに移行することを主張します。我々は、その構築に対応するプロトコルを提案し、ヒントなどの介入の明示的なモデリングや関連する反事実シナリオ分析など、標準的な連想モデルではアクセスできないままの推論形式の実際的な関連性を分析します。私たちのプロトコルでは専門家から構造方程式を導き出す必要がありますが、必要な情報は純粋に論理的なものであり、確率的で維持可能性の低い仮定には依存しません。義務教育学校の生徒のアルゴリズムスキルを測定するために設計された複雑なタスクを使用する評価からのデータを使用して、アプローチを説明します。

原文 (English)

Causal Modelling of Support Interventions for Student Competency Assessment

Accurate assessment of student competencies is essential for enabling educators to identify individual needs, design targeted interventions, and evaluate the effectiveness of educational strategies. Empirical assessment procedures are typically grounded in psychometric models, such as item response theory, which relate student competence levels to performance on assessment tasks. In this paper, we advocate adopting a structural causal modelling approach to educational assessment, moving beyond probabilistic belief updating toward a framework that explicitly supports interventional and counterfactual reasoning. We propose a corresponding protocol for its construction and analyse the practical relevance of forms of reasoning that remain inaccessible to standard associative models, including the explicit modelling of interventions such as hints and the related counterfactual scenario analysis. Although our protocol requires the structural equations to be elicited from experts, the necessary information is purely logical and does not rely on probabilistic, less tenable assumptions. We illustrate the approach using data from an assessment that employs complex tasks designed to measure compulsory school student algorithmic skills.

13:00 JSTLLM/生成AIDeepSeek

Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning

Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregres…

13:00 JSTLLM/生成AIビジネス/資金調達

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

Evaluations of generative language models frequently interpret observable behavioral traits, such as political stance, brand inclination, a…

13:00 JSTLLM/生成AIエージェント

Confident at the moment of action: belief miscalibration in LLM play under hidden information

Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of…

13:00 JST研究/論文

Lifted Model Construction under Approximate Commutativity

Lifted inference algorithms enable scalable probabilistic inference even for large object domains by leveraging the indistinguishability of…

13:00 JSTLLM/生成AIエージェント

Meta$^n$: Recursive Self-Improvement through Emergent Depth

Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta-level hold that level fixed,…

13:00 JSTLLM/生成AI

RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons

Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely o…

13:00 JSTエージェント

Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav

Large language model agents are moving beyond conventional retrieval-augmented generation toward direct interaction with external corpora.…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such a…

13:00 JST研究/論文

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain plays that role is rarely…

13:00 JSTエージェント

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor red…

13:00 JSTエージェントGPT / ChatGPTQwen

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harnes…

13:00 JST研究/論文

Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core

Recent work has applied Mamba style state space models (SSMs) to video anomaly detection, yet existing approaches still rely on buffering c…

13:00 JSTLLM/生成AI

Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA

Large language models are increasingly used for knowledge graph question answering (KGQA), but can fail to correctly ground answers in the…

13:00 JSTLLM/生成AI

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incide…

13:00 JST研究/論文

FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs

Real-world data for knowledge graph question answering is often distributed across different organizations due to governance and data sover…

13:00 JSTLLM/生成AIエージェント

SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use traject…

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPT

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invo…

13:00 JST画像/動画生成ロボティクス

Progressively Learning Heterogeneous Skills in a Unified Latent Space

We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within a unified latent space for physics-base…

13:00 JST画像/動画生成

A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision

We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory, and spatial stimuli -- foregrounding the…

13:00 JST画像/動画生成

Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers

Text-to-image systems use learned aesthetic scorers to filter training data and guide generation, but whether these scores encode demograph…

13:00 JSTLLM/生成AIエージェントAnthropicClaudeOpenAIGPT / ChatGPTGoogleGemini

REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring

Large Language Models (LLMs) offer new opportunities for automated code refactoring. However, generated changes must reduce targeted qualit…

13:00 JSTエージェント

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a model is strong enough, a multi-agent r…

13:00 JST研究/論文

Identifying Latent Declarative Representations of Code for Assisting Repository Migration

Legacy software repositories embed decades of domain knowledge in undocumented code, making understanding and modernization difficult. We t…

13:00 JSTLLM/生成AIエージェント

When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs

Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify execution traces, or enforce runtime po…

13:00 JSTロボティクス

Macro-Operator Generation and Predicate Selection for TAMP Operator Learning

Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning systems (TAMP). Recent works show…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents

Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to invoke external systems and complete…

13:00 JSTエージェント

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

Agent harnesses record a failed tool call and its error message in the transcript and ask the model to continue, on the assumption that the…

13:00 JSTエージェント研究/論文ClaudeDeepSeek

Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents different requirements from general-pu…

13:00 JSTLLM/生成AI

Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap

An LLM serving engine sizes its key-value (KV) cache once, at startup, permanently setting aside a reserve for the worst-case prefill activ…

13:00 JSTLLM/生成AI

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their dir…

13:00 JSTハードウェア/半導体

Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model

Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device models now ship to hundreds of millions…

13:00 JSTLLM/生成AIビジネス/資金調達

The Limits of Automatic Evaluation of Creativity in Large Language Models

Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity,…

13:00 JST画像/動画生成

Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

The growing size of Convolutional Neural Networks has led to increasingly large and costly models. Knowledge Distillation (KD) addresses th…

13:00 JSTビジネス/資金調達

EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain l…

13:00 JSTエージェント研究/論文

TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers

The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model agents to external tool backends. This o…

13:00 JSTLLM/生成AI

What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preli…

13:00 JST研究/論文

Disentangled Skill Representations for Predictive Human Modeling

Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variable estima…

13:00 JSTLLM/生成AI

When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk

LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typica…

13:00 JST研究/論文

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and t…

13:00 JST画像/動画生成

Restoring Without Forgetting: Continual Learning Across Image Degradations

Recent progress in image restoration has converged on all-in-one architectures that jointly handle multiple degradations within a single ne…

13:00 JST画像/動画生成エージェント

LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology

Lung cancer tissue diagnostics is complex, as therapy decisions in precision oncology rely on the integration of histomorphological, immuno…

13:00 JSTLLM/生成AIGemmaLlamaQwen

Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on sh…

13:00 JSTLLM/生成AI

Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring

Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but contemporary approaches typically employ…

13:00 JST研究/論文

Place, Slice and Schedule: Hierarchical O-RAN Control of a Tethered mmWave UAV-gNB

Unmanned aerial vehicle (UAV)-mounted 5G New Radio base stations (gNBs) can augment terrestrial networks with an on-demand, repositionable…

13:00 JST画像/動画生成

Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation

Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and attention to diagnostically relevant regions…

13:00 JST画像/動画生成

Infant Care Video Dataset for Classification of Interventions Using Transformers

Healthcare documentation in the neonatal intensive care unit (NICU) presents significant challenges, with nurses spending approximately 25\…

13:00 JSTエージェントロボティクスビジネス/資金調達

Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization

Embodied Agents System (EAS) are increasingly deployed in open-world physical domains, where reliability directly dictates deployment quali…

13:00 JST研究/論文

ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork

Training large-scale AI models often outgrows a single data center, demanding sharded, multi-cluster, and decentralized training. However,…

13:00 JST研究/論文

Automated Synthesis of Cloud Emulators

DevOps programming (e.g., using CLI/API scripts or IaC frameworks) is key to cloud infrastructure management. Unlike traditional programmin…

13:00 JST研究/論文

Coronavirus Optimization Algorithm: A Success-History Adaptive Evolutionary Framework with Archive-Assisted Search and Stagnation Recovery for Global Optimization

This paper proposes the Coronavirus Optimization Algorithm (COA), a SARS-CoV-2-inspired success-history adaptive evolutionary optimizer for…

13:00 JSTLLM/生成AIエージェントGoogle

Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)

The Agent Payments Protocol (AP2), introduced by Google, enables large language model (LLM)-driven shopping agents to authorize and execute…

13:00 JST研究/論文Mistral AIQwen

Revelation Control

Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can c…

13:00 JST研究/論文

A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization

Real-time optimization (RTO) relies on process models to locate economically optimal operating conditions. Because developing first-princip…

13:00 JST研究/論文

A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource

This article presents the abridged core of \emph{A Mathematical Theory of Interpretation} (MTI), which treats interpretation as observer-re…

13:00 JST研究/論文

Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory

Kohn--Sham density functional theory (DFT) underpins electronic-structure simulations, but repeated orbital diagonalizations lead to cubic…

13:00 JSTLLM/生成AI

Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs

When a code generating language model fabricates a Python package name, an adversary who has pre-registered that name on PyPI can convert t…

13:00 JST画像/動画生成

RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding

Surgical spatio-temporal grounding (STG) requires locating, at each video time specified by a procedural question, the object that the ques…

13:00 JST研究/論文

QML for Quantum Sensing under Measurement-Induced Information Loss

Nitrogen-vacancy (NV) centers in diamond can serve as highly sensitive solid-state quantum sensors for high-sensitivity magnetometry. Howev…

13:00 JST画像/動画生成

Luce: Relightable Gaussians for 3D Asset Generation

High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and int…

13:00 JST研究/論文

STAIN-FL: Stealthy Targeted Attack Injection with Contextual Triggers in Federated Learning

Federated video anomaly detection trains model collaboratively without sharing raw surveillance footage, but limited server-side visibility…

13:00 JSTLLM/生成AIエージェントDeepSeek

The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses

An agent harness is what turns a language model into an autonomous agent: the surrounding code that builds the model's context, mediates it…

13:00 JSTLLM/生成AI

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

Safety alignment in large language models (LLMs) remains brittle against a growing spectrum of attacks. Jailbreak attacks bypass safety mec…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

Evaluating Language Models on Cross-Language Code Functional Equivalence

Background: Large Language Models (LLMs) have demonstrated strong performance across a variety of code-understanding tasks, leading many to…

13:00 JST研究/論文

RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) improves the factuality of large language models by grounding responses in external documents, but it…

13:00 JSTLLM/生成AI

The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem

Large language model providers are compute constrained, and their universal response to congestion is to degrade service: route queries to…

13:00 JSTLLM/生成AIエージェント

Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation

Large language model (LLM) agents invoke external tools to retrieve and reason over information beyond pretrained knowledge. The Model Cont…

13:00 JSTLLM/生成AI

SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding

Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning. Current Large Vision-Language Models (…

13:00 JSTLLM/生成AIエージェント

WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents

The emerging W3C WebMCP proposal enables LLM agents to invoke tools exposed by web pages. In multi-party web environments, however, integra…

13:00 JST画像/動画生成

IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views

Generating executable parametric CAD code from dimension-annotated orthographic drawings is a challenging task requiring geometric understa…

13:00 JSTLLM/生成AIエージェント

What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions

LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them v…

13:00 JST研究/論文

ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning

Time series classification underpins applications in healthcare, sensing, and industrial monitoring. Although time series foundation models…

13:00 JSTエージェントロボティクス

Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings

Manufacturing process planning transforms heterogeneous design information into coherent manufacturing decisions. However, existing approac…

13:00 JSTロボティクス

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, thei…

13:00 JSTエージェント

Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding

While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain sc…

13:00 JSTLLM/生成AI画像/動画生成

VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference

While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long-context inference remains p…

13:00 JSTLLM/生成AI

Mechanistic Circuit Identification for Controllable Data Generation

While recent advances in data synthesis aim to curate high-quality datasets, most generation pipelines still rely on heuristic prompt-based…

13:00 JST画像/動画生成

ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal

Low-earth-orbit (LEO) satellites enable high-resolution, large-scale Earth observation for applications such as disaster monitoring and env…

13:00 JSTLLM/生成AI

When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs

In psychological counseling, effective support is not always delivered through long, information-rich responses. Minimal responses, such as…

13:00 JSTLLM/生成AI

PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding

Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and…

13:00 JSTLLM/生成AIエージェント

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier sc…

13:00 JST画像/動画生成研究/論文

MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes

Material replacement is a common interior-design operation: changing the material of a selected surface while preserving its geometry, surr…

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection

Time-series anomalies can appear not only as pointwise deviations but also as changes in recurring temporal structure, such as shifted peri…

13:00 JSTLLM/生成AIロボティクス

PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a…

13:00 JST画像/動画生成

TransPhy: Visual In-Context Learning for Physically Grounded Image Editing

Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with tex…

13:00 JST画像/動画生成

Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking

Multi-camera 3D perception systems for warehouse scenes are trained largely on synthetic data and evaluated on physically captured environm…

13:00 JST研究/論文

From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender

Product catalogs in fast-moving service businesses are shifting from static, independently priced SKUs toward dynamically bundled, discount…

13:00 JST画像/動画生成

PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

People search for urban outdoor places not only by category or function, but also by what activities a place can support and how it is perc…

13:00 JST画像/動画生成

Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection

Real-world deployment of traffic surveillance systems is bottlenecked by geographic domain shift, in which models trained in one city under…

13:00 JSTLLM/生成AIビジネス/資金調達

LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes

Industrial actor--critic methods usually represent continuous actions as anonymous numerical coordinates. They must therefore learn from li…

13:00 JST研究/論文

Preference Optimization for Non-Verbal Vocalization Synthesis

Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference…

13:00 JST研究/論文

Tlow: Flow-based Item Tokenizer for Recommendation

Item tokenizer encodes semantic embeddings into token IDs to replace the randomly assigned item IDs used in traditional recommendation mode…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGeminiLlamaQwen

'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection

Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluatio…

13:00 JST研究/論文

Contrastive Branch Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its…

13:00 JSTLLM/生成AI

SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling

Recent controllable text generation (CTG) for sentiment control has largely focused on decoder-based large language models, making causal a…

13:00 JST画像/動画生成

Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct compri…

13:00 JST画像/動画生成

Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis

Synthetic image generation is a promising strategy to address data scarcity and the underrepresentation of clinically important phenotypes…

13:00 JSTLLM/生成AI

FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision

To reduce the hallucination risk caused by outcome-driven rewards in large language models trained through reinforcement learning with veri…

13:00 JSTLLM/生成AI

Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs

MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: tri…

13:00 JST画像/動画生成

Markerless Pose Estimation for Resistance Training Technique Assessment

Resistance training can be a high risk activity, and safe form is essential to avoiding injury. Laboratory-based movement analysis provides…

13:00 JSTハードウェア/半導体

Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning

Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification (UQ) for such outputs remains an open ch…

13:00 JSTエージェント

Multilevel Fair Allocation under Additive Preferences

We study multilevel fair resource allocation with tree-structured hierarchical relations among agents. At each level, the problem can be vi…

13:00 JST研究/論文

Evaluating Deep Multivariate Imputation Models on Wearable Device Data

Wearable device data enables continuous health monitoring, but suffers from structured missingness: features sharing a physical sensor drop…

13:00 JSTLLM/生成AI

Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning

Mechanistic Localization bridges mechanistic interpretability and post-training optimization by isolating critical parameters via interpret…

13:00 JSTLLM/生成AI

When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study

Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings…

13:00 JST研究/論文

Scalable and Versatile Identification for Hierarchical Structural Causal Models: A New Look at Project STAR

The STAR (Student-Teacher Achievement Ratio) experiment (1985, Tennessee, USA) is a landmark hierarchical dataset designed to assess the im…

13:00 JST研究/論文

LumiXAI: A Modular Full-Stack Framework for Feature Attribution

Feature attribution is a central tool of model interpretability, yet the software through which it is applied remains fragmented: individua…

13:00 JST研究/論文

FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment

Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to…

13:00 JSTエージェント

StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment

Prehospital stroke assessment aims to accurately identify stroke symptoms and make rapid decisions through standardized procedures within a…

13:00 JST研究/論文

COCI: Conference Organisers and Content Identifier

Despite the critical role of grey literature in scholarly communication, artefacts such as Calls for Papers (CfPs) remain largely isolated…

13:00 JST研究/論文

Across the Loss Landscape with Progressive Growth

Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with…

13:00 JST研究/論文

$\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions

Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logis…

13:00 JST研究/論文

Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks

Electroencephalography (EEG) is a widely used window into human brain function, but most EEG models remain tied to a one-dataset-one-model…

13:00 JSTエージェント研究/論文

A Literate Programming Environment for Human and Machine Agents

This paper introduces an environment for constructing literate programs in concert with language-aware machine agents. This environment inc…

13:00 JSTLLM/生成AIエージェント

Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems

System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems, where real deployment…

13:00 JST研究/論文

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and…

13:00 JST研究/論文

On-policy Distillation with Verifiable Reward

Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-tr…

13:00 JST研究/論文

Constrained Hyperparameter Optimization for Streaming Data

Optimization of hyperparameters is a critical factor to obtain optimal model performance. While existing research has predominantly concent…

13:00 JST画像/動画生成

Deep Learning Super Resolution for Satellite Cloud Mask Downscaling

A vast amount of optical satellite data is being transmitted to Earth-based servers every day, and more than half of this data is affected…

13:00 JST研究/論文

Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity

Hyperparameter selection remains a key challenge in Bayesian optimization (BO) and Bayesian active learning (AL), as model misspecification…

13:00 JST研究/論文

Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets

EEG foundation models pretrained via self-supervised learning promise transferable representations, but their generalization remains limite…

13:00 JSTLLM/生成AI

Method, Mind, and Morality: How People Make Sense of Artificial Intelligence

How can humans make sense of the rapid takeoff of artificial intelligence (AI)? We studied the sensemaking dynamics of AI through an open-e…

13:00 JSTLLM/生成AIビジネス/資金調達

The RAT: A Unified Bayesian Model for RAG Evaluation

Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual componen…

13:00 JST研究/論文

Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations

Many disentanglement methods represent generative factors using Euclidean product coordinates, although the underlying factor spaces may wr…

13:00 JST画像/動画生成

Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks

The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tasks. However, analytica…

13:00 JST画像/動画生成

Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy

Brain stroke, known for its high mortality and incidence rates, poses significant health risks and requires rapid intervention for survival…

13:00 JSTLLM/生成AI

Automatic Model Card Generation Using an LLM

Model cards are structured documents that summarize key information about machine learning models to improve transparency, usability, and a…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment d…

13:00 JST画像/動画生成

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected…

13:00 JST研究/論文

Fuzzy Segmentations of a String

This article discusses a particular case of the data clustering problem, where it is necessary to find groups of adjacent text segments of…

13:00 JST研究/論文

Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives

This work investigates formal policy synthesis for continuous-state stochastic dynamic systems subject to high-level specifications express…

13:00 JSTLLM/生成AIGPT / ChatGPTLlama

Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs

In the past two years, the outstanding performance of ChatGPT in multilingual and multitasking has led to large language models (LLMs) attr…

13:00 JST研究/論文

LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis

Root cause analysis (RCA) is crucial for enhancing the reliability and performance of complex systems. However, progress in this field has…

13:00 JSTLLM/生成AI

Efficient LLM Collaboration via Planning

Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large mod…

13:00 JSTエージェント

Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light

Artificial learning systems are graduating from passive learners to increasingly autonomous agents, lending pragmatic urgency to the questi…

13:00 JSTエージェント

Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning

Shielding is widely used to enforce safety in reinforcement learning (RL), ensuring that an agent's actions remain compliant with formal sp…

13:00 JSTLLM/生成AI

UCO: A Multi-Turn Interactive Reinforcement Learning Method for Adaptive Teaching with Large Language Models

Large language models (LLMs) are shifting from answer providers to intelligent tutors in educational settings, yet current supervised fine-…

13:00 JSTLLM/生成AI

ReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering

Large reasoning models achieve strong performance on diverse tasks by producing extended chains of thought. Self-reflection, the ability to…

13:00 JST研究/論文

Panning for Gold: Expanding Domain-Specific Knowledge Graphs with General Knowledge

Domain-specific knowledge graphs (DKGs) are critical yet often suffer from limited coverage compared to General Knowledge Graphs (GKGs). Ex…

13:00 JST研究/論文

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learnin…

13:00 JSTLLM/生成AIエージェント

CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support

Recent multi-agent frameworks have shown promise for oncology decision support, yet most assume centralized data access and rely on prompt-…

13:00 JSTLLM/生成AIエージェント

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act rel…

13:00 JST研究/論文

Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints

Clinical prediction from structured electronic health records (EHRs) is challenging due to high dimensionality, heterogeneity, class imbala…

13:00 JSTLLM/生成AI研究/論文

ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams

Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reasoning over simple linear diagrams. However…

13:00 JSTビジネス/資金調達

Housing Potential Common Data Model and City Digital Twin

The evaluation of housing potential requires consideration of a location from multiple perspectives, ranging from zoning and land use to po…

13:00 JSTLLM/生成AIエージェント

Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust

Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new fo…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達GPT / ChatGPTGemini

EngiAI: LLM 主導のエンジニアリング設計のためのマルチエージェント フレームワークおよびベンチマーク スイート

Large Language Model (LLM) エージェントはエンジニアリング設計タスクに適用されることが増えていますが、既存の評価フレームワークは、シミュレーション、検索、製造準備を組み合わせたマルチエージェント システムに適切に対応していません。 3 つの評価次元を備えたベンチマーク スイートを紹介します。(1) ツールの直接使用、意味論的曖昧さの解消、条件分岐、作業記憶タスクなど、明確な認知的要求を対象とした 7 つのプロンプト スタイルを備えたワークフロー ベンチマーク。 (2) パラメータ選択に対する検索の寄与を分離するゲート付きスコアリングを備えた検索拡張生成 (RAG) ベンチマーク。 (3) SLURM クラスター上のエンドツーエンドの ML トレーニング オーケストレーションを評価するハイ パフォーマンス コンピューティング (HPC) ベンチマーク。ベンチマークとともに、LangGraph 上に構築されたマルチ エージェント システム (MAS) リファレンス実装である EngiAI も紹介します。EngiAI は、スーパーバイザ アーキテクチャを通じて 7 つの専門エージェントを調整し、トポロジの最適化、ドキュメントの取得、HPC ジョブ オーケストレーション、および 3D プリンタの制御を統合することでベンチマークを運用します。 4 つの LLM バックエンドと 2 つの EngiBench 問題にわたって、独自のモデルは Beams2D 上で平均タスク完了率 96 ~ 97% を達成していますが、オープンソースの 4B パラメータ モデルは 55 ~ 78% に達しており、世代ごとに明らかな改善が見られます。条件付き分岐が最も困難であることが判明し、Photonics2D の条件付きスタイルではタスクの完了率が 20 ~ 53% に低下します。 RAG ゲーティングにより、ほぼ完璧な検索強化スコア (約 1.0) と検索なしのほぼゼロのスコアが確認され、評価設計が検証されます。 HPC オーケストレーションでは、あるモデルは実行の 100% ですべてのパイプライン ステップを完了しますが、別のモデルは 50% に低下し、長時間実行されるワークフローでは複数ステップの命令のパフォーマンスが低下することがわかります。

原文 (English)

EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design

Engineering-agent systems are proliferating, but differences in tasks, tools, and success criteria make demonstrations difficult to compare and failures difficult to diagnose. We introduce a capability-based evaluation framework for tool-connected engineering agents. The framework separately evaluates workflow execution, retrieval-assisted parameter selection, high-performance computing (HPC) orchestration, and training-code authoring using execution traces and resulting engineering artifacts. We evaluate four LLM backends on the EngiBench Beams2D and Photonics2D problems using EngiAI, a LangGraph reference implementation. On Beams2D, the two proprietary models complete 96-97% of workflow tasks, compared with 55-78% for the two open-source models. Workflows requiring tool-based decision-making perform worse on Photonics2D, reaching 20-53% completion. Indexed retrieval improves parameter selection. For HPC orchestration, Gemini-3-Flash completes every tested pipeline, whereas GPT-5-mini completes 50-70% through the final evaluation step. When the agents must write the training code themselves, they fill small gaps reliably but diverge as larger regions are left open. In the open-synthesis tier, both proprietary models select conditional variational autoencoders instead of the reference conditional Generative Adversarial Network (cGAN); Gemini-3-Flash achieves lower maximum mean discrepancy than the reference in seven of ten Beams2D trials. The results support capability-specific evaluation, with separate scores for distinct skills, rather than assessment through successful end-to-end demonstrations alone. The framework provides common evaluation dimensions for identifying failure mechanisms and structuring comparisons of models, agent architectures, and tool interfaces.

13:00 JSTエージェント

自己進化する科学エージェントが一般化可能な物理的根拠に基づいた流体制御を発見

データ集約型の深層強化学習は複雑な制御ポリシーを最適化できますが、物理システムにおける科学的発見には基本的に、物理的証拠を構造化された制御アーキテクチャに結び付ける、解釈可能な推論の連鎖が必要です。ここでは、大規模な言語モデルと反復コード生成によって駆動され、厳密な解釈可能性と厳密な物理的推論を維持しながらコントローラーの構築を自動化する、自己進化する科学エージェントのワークフローを紹介します。重みを調整する代わりに、エージェントは候補戦略を物理シミュレーションに展開し、マルチモーダルな証拠から動的動作を積極的に診断し、これらの観察結果を漸進的なソースコードの改良に変換します。我々は、このフレームワークを高度に非線形の流体構造相互作用問題、つまり関節角加速度のみを使用して空間目標に到達する任務を負った、作動が不十分な 2 関節のツノザメ遊泳者について実証します。一方的なステアリング バイアスを示す推進シード ポリシーから開始して、エージェントは自律的に、すべての正規ターゲットを確実に捕捉する統合コントローラーを発見し、改良します。注目すべきことに、再トレーニングやターゲット固有の分岐を行わずに、合成された制御ポリシーは、目に見えない静的なターゲットと動的に湾曲した追跡軌道に一般化されます。監査可能な進化ログは、進行波推進、車体フレーム目標誘導、ヨーレートフィードバック、符号付き平均尾部曲率、および適応ケイデンス緩和に基づいて構築された緊急制御アーキテクチャを明らかにします。私たちの結果は、自律的な科学エージェントが、科学的発見の完全に追跡可能なプロセスを維持しながら、蓄積された物理的証拠を堅牢で数学的に読み取り可能な制御ポリシーにうまく変換できることを示しています。

原文 (English)

Self-Evolving Scientific Agent Designs Physically-Reasoned Whitebox Fluid Control

While data-intensive deep reinforcement learning can optimize complex control policies, scientific control design in physical systems fundamentally requires an interpretable chain of reasoning that connects physical evidence to structured control architectures. Here, we present a self-evolving scientific agent workflow, driven by large language models and iterative code generation, that automates controller construction while preserving strict interpretability and rigorous physical reasoning. Instead of adjusting weights, the agent deploys candidate whitebox controllers into physical simulations, actively diagnoses dynamic behaviors from multimodal evidence, and translates these observations into progressive source-code refinements. We demonstrate this framework on a highly non-linear fluid-structure interaction problem: an underactuated, two-joint dogfish swimmer tasked with spatial target reaching in an unsteady flow using only joint angular accelerations. Starting from a target-blind propulsive seed, the agent autonomously designs and refines a unified controller that reaches a target embedded in an unsteady four-cylinder wake. Without retraining, retuning or case-specific branching, the retained controller achieves target capture across the full generalization test matrix, spanning variations in target position, rear-row geometry, cylinder count and inflow speed. The auditable evolution log reveals an emergent control architecture built upon travelling-wave propulsion, body-frame bearing guidance, phase-selective steering, corrective burst and adaptive relief. Our results show that an autonomous scientific agent can successfully transform accumulated physical evidence into a robust, mathematically readable control policy, while maintaining a fully traceable process of scientific control design.

13:00 JST研究/論文

X の原子単位: インテリジェンスの圧縮層

この論文は、インテリジェンスを原子圧縮と構成再利用のプロセスとして理解するための理論的枠組みを提案します。私たちは、認知システム、生物学システム、計算システム、および組織システムは、複雑な現象を高次構造に再結合できる再利用可能な原子単位に分解することによって、スケーラブルなインテリジェンスを実現すると主張します。この論文は、認知科学、情報理論、進化生物学、ソフトウェア工学、医学、法的推論、教育、音楽、人工知能からの証拠を利用して、効率、伝達、解釈可能性、進化可能性をサポートする基本的な圧縮層としての原子単位の概念を開発しています。中心的な貢献は Compression Calculus です。これは、表面レベルの表現を原子表現と比較し、抽象化レイヤー全体で圧縮ゲインがどのように合成されるかを説明するための正式なフレームワークです。我々は、複合カスケード理論を導入します。これによれば、抽象化の各層が追加されると、単に増分節約が追加されるのではなく、表現効率が乗算的に向上します。この論文はさらに、現代の AI システムは、安定した概念レベルの原子構造ではなく、トークン レベルの処理やドキュメント レベルの検索に依存し、次善の表現レベルで動作することが多いと主張しています。この観点では、大規模な言語モデルは、完全な知識アーキテクチャとしてではなく、原子単位のナビゲーション、順序付け、および再結合が可能な動的融合エンジンとして最もよく理解されます。このフレームワークは、時間の経過とともに新しいプリミティブを発見、洗練、構成できる自己進化する知識システムを設計するための基盤を提供します。この論文は、構成的抽象化による圧縮としてインテリジェンスを再構成することにより、専門知識、知識表現、説明可能な AI、および適応型インテリジェント システムの将来のアーキテクチャに関する統一的な視点を提供します。

原文 (English)

Atomic Units of X: The Compression Layer of Intelligence

This paper proposes a theoretical and empirical framework for understanding intelligence as a process of atomic compression and compositional reuse. It argues that scalable cognitive, biological, computational, and organisational systems reduce complexity by organising information into reusable units that can be recombined into higher-order structures. The central contribution is the Compression Calculus, a formal framework for comparing surface evidence with atomic representations and for describing how abstraction can compound across layers. The framework is evaluated on large, multi-source software corpora under an open-world concept model, showing substantial evidence-to-concept consolidation in two large projects, while a smaller third project remains below the target threshold. The analysis further identifies important boundary conditions: consolidation depends on corpus scale, evidence-unit definition, and concept identity, and the observed reduction is driven primarily by within-corpus recurrence rather than cross-source recurrence. The paper also develops a representational account of the meaning gap in contemporary generative systems, describing latent, context-dependent conceptual approximations as soft atoms that lack the persistent identity and compositional constraints of stable atomic units. This motivates an architectural view in which large language models function as dynamic fusion engines that navigate and compose persistent conceptual structures rather than serving as the sole repository of those structures.

13:00 JSTエージェント

SuperLocalMemory 4.0: AI エージェント用のガバナド メモリ オペレーティング システム

AI エージェントは共有インフラストラクチャになりつつありますが、耐久性のあるメモリは通常、個別の取得、ガバナンス、および運用コンポーネントから組み立てられます。私たちは、AI エージェント向けの管理されたローカル ファースト メモリ オペレーティング システムである SuperLocalMemory 4.0 を紹介します。このシステムは、相互ランク融合による高密度セマンティック検索、BM25 語彙検索、時間検索、ホップフィールド連想検索、および拡散活性化検索を組み合わせています。管理された学習層と行動層。双時間的想起。マルチスコープの個人メモリ、共有メモリ、およびグローバルメモリ。役割ベースのアクセス制御。 GDPR 指向のエクスポートと検証済み消去。監査証跡。展開コンテキストの EU AI 法のチェックリスト。 V4 では、プライマリ書き込みパスに信頼性スパインが導入されています。世代制限されたアドミッション、ポリシー レジストリ、投影ごとに適用、検証、補償、所有者を消去できる検証可能なメモリ トランザクション、およびハッシュ チェック可能な完了マニフェストです。このランタイムは、CLI、MCP、HTTP デーモン、ダッシュボード、エディター統合、およびフレームワーク アダプターを通じて利用可能で、完全ローカル、ローカル ウィズ オンデバイス モデル、およびプロバイダー支援モードをサポートします。 11 のフォールト挿入シナリオとメカニズム シナリオをそれぞれ 200 回繰り返して評価しました。リリースされた証拠バンドルは、範囲指定されたコンポーネントのプロパティを支持する 2,200 回の決定論的繰り返しのうち 2,200 回を報告します。管理された書き込みエンベロープは、p50 で 3.522 ミリ秒、p99 で 5.297 ミリ秒であったのに対し、非管理ベースラインでは 1.835 ミリ秒と 2.569 ミリ秒でした。これは、プロセス内コントロール プレーンのオーバーヘッドが p50 で 1.687 ミリ秒、p99 で 2.728 ミリ秒に相当します。これらは、対象範囲を絞ったコンポーネントおよびメカニズムの測定であり、エンドツーエンドのマルチプロセスまたは外部の取得精度のベンチマークではありません。この論文では、プライバシーを保護するマルチエージェント メモリ、情報幾何学的検索、および V3.3 Living Brain ライフサイクルに関するこれまでの SuperLocalMemory の取り組みが統合されています。

原文 (English)

SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents

We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents, unifying multi-channel retrieval under reciprocal-rank fusion, bi-temporal recall, multi-scope isolation, role-based access, verified erasure, and a hash-chained audit trail. A reliability spine governs the primary write path: generation-fenced admission, verifiable memory transactions with per-projection apply, verify, compensate and erase owners, and hash-checkable completion manifests. Eleven fault-injection scenarios, each repeated 200 times, upheld 2,199 of 2,200 scoped component properties. This version leads with a negative result. Ten mechanisms here were implemented, reachable on a live call path, and ineffective at their final connection. Implemented, reachable and effective are three different questions, and the third requires an oracle independent of the mechanism under test. We contribute two mechanical invariants that supply one: a prior-distance assertion over Bayesian learners, and a join-liveness assertion over schema-guarded paths that reports where a guard's missing data resides. A three-arm ablation varying only the recall session-identifier namespace moves no posterior with the defect present and every instantiated arm with it absent, while a negative control that writes every ticket but supplies no engagement settles nothing. We withdraw the previous version's governed write-envelope overhead figure: the two paths it differenced are not comparable. Timing the envelope in place gives an 11.0 ms governed write of which the envelope is 70.6 percent, but the generation fence costs 1.9 microseconds and the obligation ledger 42 microseconds. The cost is durability, not governance.

13:00 JSTLLM/生成AIエージェント

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

A primary goal of science is to learn mechanistic or causal world models from data. These models can be used to explain some phenomenon of…

13:00 JST研究/論文

The Dynamics of Intelligence Explosions

AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loop might be able to produce an intelligence exp…

13:00 JSTLLM/生成AIOpenAI

AI が生成した数学的証明の監査: 量子並列反復における貪欲な条件付け補題の修正

OpenAI の *Ten Advances in Mathematics and Theoretical Computer Science* の第 6 章では、すべての有限 2 プレイヤー、1 ラウンドのエンタングル ゲームに対する指数関数的並列反復定理を主張しています。証明の早い段階で、この章では定量的貪欲条件付け補題を使用します。補題は、(D) のすべての座標を獲得することを条件付けた後、ランダムに選択された残りの座標が少なくとも (1-\delta) の平均確率で獲得されるように、小さな座標セット (D) を選択することを意味します。この記述は正しいですが、印刷されたプルーフには極性エラーが含まれています。その継続テストは平均成功の観点から記述されますが、次のステップでは条件付き失敗確率が大きい調整が必要です。この意味は誤りであり、単純な例であっても、有効な次の動作が示されずに出力されたプロシージャが終了する可能性があります。このメモは明示的な反例を示し、意図された継続条件を特定し、完全に修正された証明を提供します。修復はローカルです。補題のステートメントと、この章の後半で使用されるパラメータは変更されません。ただし、これを主な並列​​繰り返し定理の独立した検証として解釈すべきではありません。より広範には、この例は、AI が生成した数学的にもっともらしい議論が、相補的な出来事間の小さいながらも決定的な逆転をどのように隠蔽できるかを示しています。

原文 (English)

Auditing an AI-Generated Mathematical Proof: Human Assessment of OpenAI's Quantum Parallel-Repetition Argument

We present an independent human assessment of the proof developed in Chapter 6 of OpenAI's Ten Advances in Mathematics and Theoretical Computer Science. An initial audit appeared to identify a polarity error in a greedy conditioning lemma. Subsequent examination of the original typeset manuscript showed that this diagnosis resulted from automatic PDF text extraction, which removed an overbar from a mathematical symbol. The alleged error is therefore withdrawn. With the correctly rendered expression restored, we find no confirmed mathematical error in the examined lemma. We retain the broader analysis because it provides an independent reconstruction and assessment of the OpenAI argument and illustrates an important methodological hazard in auditing AI-generated mathematics: errors may arise not only in mathematical reasoning but also in the document-processing pipeline used by human reviewers.

13:00 JSTLLM/生成AI研究/論文

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We intr…

13:00 JST研究/論文

ExPhy: 複数物体軌道予測における明示的物理特性学習のベンチマーク

物体のダイナミクスを理解するには、将来の軌道を予測するだけでなく、モデルが動きを支配する物理的特性を捉えているかどうかを調べる必要があります。ただし、既存のベンチマークでは、軌道予測と並んでオブジェクトレベルの物理的特性が明示的な評価対象として公開されることはほとんどありません。このギャップに対処するために、\emph{ExPhy} を導入します。これは、質量、摩擦、反発の明示的なオブジェクト レベルのラベルを持つ 24,000 のシミュレートされた物理シーンを含む、複数オブジェクトの軌道予測ベンチマークです。 ExPhy は、軌道予測と物理特性推定を共同で評価するために、物理パラメータ (OOD-Parameter) と初期状態 (OOD-Initial) にわたる分布内 (ID) 分割と 2 つの分布外 (OOD) 分割とともに、観測された軌道と将来の軌道を提供します。さらに、観測された軌道から物理特性を推定し、微分可能な将来のロールアウトに使用する明示的なプロパティ インターフェイスを備えた物理ガイド付きモデルである \textsc{PhyODE} をインスタンス化します。長期的な OOD 初期設定では、\textsc{PhyODE} は、最も強いベースラインと比較して、ADE と FDE をそれぞれ 33.1\% と 31.0\% 削減します。 ComPhy でのゼロショット評価は、クロスベンチマーク移行をさらに評価します。特性レベルの分析により、正確な軌道予測が必ずしも基礎となる物理特性の正確な回復を意味するわけではないことが明らかになりました。コードとデータは https://github.com/Zest86/ExPhy で入手できます。

原文 (English)

ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting

Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical properties that govern motion. However, existing benchmarks rarely expose object-level physical properties as explicit evaluation targets alongside trajectory forecasting. To address this gap, we introduce \emph{ExPhy}, a multi-object trajectory forecasting benchmark containing 24,000 simulated physical scenes with explicit object-level labels for mass, friction, and restitution. ExPhy provides observed and future trajectories together with an in-distribution (ID) split and two out-of-distribution (OOD) splits over physical parameters (OOD-Parameter) and initial states (OOD-Initial) for jointly evaluating trajectory forecasting and physical property estimation. We further instantiate \textsc{PhyODE}, a physics-guided model with an explicit property interface that estimates physical properties from observed trajectories and uses them for differentiable future rollout. On the long-horizon OOD-Initial setting, \textsc{PhyODE} reduces ADE and FDE by 33.1\% and 31.0\%, respectively, compared with the strongest baseline. Zero-shot evaluation on ComPhy further assesses cross-benchmark transfer. Property-level analyses reveal that accurate trajectory forecasting does not necessarily imply accurate recovery of the underlying physical properties. Code and data are available at https://github.com/Zest86/ExPhy.

13:00 JST研究/論文

目に見えないものはあなたが学ぶものです:共有ゲノム言語モデル社会では、限られた証拠の可視性が構成の一般化を促進します

マルチモジュール システムでは、多くの場合、すべてのモジュールが完全な入力に公開されます。証拠の可視性を制限すると、勾配ベースのトレーニングで発見されるソリューションが変わるかどうかをテストします。 4 セル社会は 1 つの凍結済み事前学習済み言語モデルと 1 つの低ランク アダプターを共有し、固定リレー内の 2 つのモデル幅の連続ベクトルを通じてのみ通信します。プロスペクティブシールされた自然言語関数合成タスクでは、初期化バイト、トレーニング順序、トークン レイアウト、パラメーター、および計算を共有する 10 個の一致する制限付き/グローバル ペアをトレーニングします。注意マスクのみが異なります。制限された社会は、10 ペア中 9 ペアで両方の深さで世界的に見える双子よりも少なくとも 20 ポイント優れており、ペアの利点の中央値は 0.7648 と 0.6050 です。コミュニケーションを遮断すると、すべての制限された社会は偶然に帰着し、複合関数がトレーニングに一度も現れなかったプログラムでは、深さ 3 の利点は 0.558 のままです。監査された 6 つの制限された社会全体で、同じ値のパケット移植により、テストされたすべてのインターフェイスで動作が 0.94 ~ 1.00 に維持されます。破壊的な介入はパフォーマンスを崩壊させます。そして反事実パケットは出力を数学的に予測された答えにリダイレクトします。唯一の高性能グローバル モデルも通信を必要としますが、その同じ値のパケットはエピソード間で交換できません。したがって、構図には可視性の制限は必要ありません。このプロトコルでは、一般化リレーの可能性が大幅に増加し、再利用可能な値インデックス付きインターフェイスが優先されます。それにもかかわらず、完全に事前登録されたバッテリーは、制限付きアームの深さ 3 の中央値の精度が 0.6988 で、0.70 フロアを下回っているため、正式に不合格となります。以前の資格コホートでも同様に完全合格は 0/10 でした。1 つのモデルはすべてのタスク パフォーマンス ゲートを満たしましたが、10 モデルすべてが通常言語の保存に失敗し、システムは明示的にタスク ゲートで使用されるように制限されました。

原文 (English)

What You Can't See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies

Multi-module neural systems often expose every module to the full input. We test whether a slot-selective evidence-masking regime -- restricting each module to its own evidence span -- changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay. On a prospectively sealed natural-language function-composition task, we train ten matched restricted/global pairs identical except for the attention mask. Restricted-visibility societies outperform their globally visible twins by at least 20 percentage points at both depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050. Cutting communication reduces every restricted society to chance, and in a post hoc collision-stratified analysis the depth-three advantage remains 0.558 on programs whose complete affine map never appeared in training. In six post hoc-selected restricted societies, packet interventions on correctly answered held-out episodes are consistent with approximately value-indexed relay states; the sole high-performing global model also requires communication, but its same-value packets are not interchangeable across episodes. Thus restricted visibility is not necessary for composition. Under the tested seeds, streams, task world, and training budget, the masking regime strongly shifted which solutions training discovered: a post hoc mask crossover finds both arms mask-native. Because the restricted mask both blocks foreign evidence and implicitly identifies each cell's assigned slot, attribution to evidence visibility alone awaits a role-marked control. The preregistered battery nevertheless formally fails because restricted-arm median depth-three accuracy is 0.6988, below the 0.70 floor; an earlier qualification cohort yielded 0/10 complete passes.

13:00 JST研究/論文

SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing

Hate speech research has moved from coarse-grained classification towards structured parsing, where systems jointly identify targets, suppo…

13:00 JSTエージェント

GenCoord: Skill-Path Commitments under Private Information

Suppose one embodied agent knows what must be built, while its teammate alone knows which transformation its workcell can perform. Neither…

13:00 JSTLLM/生成AI画像/動画生成

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large languag…

13:00 JSTエージェント

CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents

Long-horizon GUI agents can retain complete action histories as compact text, but only a few historical screenshots fit in active context.…

13:00 JST研究/論文

SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems

Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial storage and latency overhead. Hard quan…

13:00 JSTLLM/生成AIエージェント研究/論文

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigor…

13:00 JSTLLM/生成AIエージェント

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, infor…

13:00 JSTエージェントQwen

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidenc…

13:00 JST画像/動画生成

Screening Autism Spectrum Disorder in children using Deep Learning Approach : Evaluating the classification model of YOLOv26s by comparing with other models

Autism spectrum disorder (ASD) is a developmental condition that presents significant challenges in social interac- tion, communication, an…

13:00 JSTLLM/生成AI

HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA

Retrieval-Augmented Generation (RAG) significantly improves document-based question answering by integrating external documents during gene…

13:00 JST画像/動画生成

Intrinsic PAPR: Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering

Recent point-based intrinsic decomposition and inverse rendering methods have advanced the modelling of the shading and albedo of 3D scenes…

13:00 JST研究/論文

Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control

Connected automated vehicles (CAVs) equipped with adaptive cruise control (ACC) create new opportunities for highway congestion mitigation.…

13:00 JST研究/論文

Generative AI for Validating Physics Laws

We propose generative learner for estimating heterogeneous treatment effects and characterizing the full distribution of causal effects. Th…

13:00 JSTLLM/生成AIハードウェア/半導体

Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review

Large Language Models (LLMs) have been transformative across many domains. However, hallucination, i.e., confidently outputting incorrect i…

13:00 JSTエージェントロボティクス

Balancing Safety and Optimality in Robot Path Planning: Algorithm and Metric

Path planning for autonomous robots faces a fundamental trade-off between path length and obstacle clearance. While existing algorithms typ…

13:00 JSTLLM/生成AI

Quasar: A Programming Language Specialized for LLM Code Actions

Large language models (LLMs) often call external tools to solve tasks. One effective strategy is for LLMs to write code, enabling them to u…

13:00 JSTLLM/生成AIビジネス/資金調達

From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs

[...] Since then, various APR approaches, especially those leveraging the power of large language models (LLMs), have been rapidly develope…

13:00 JSTLLM/生成AI

A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs

Spatio-temporal data mining plays a pivotal role in informed decision making across diverse domains. However, existing models are often res…

Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities

Large Language Models (LLMs) are becoming widely used to support various workflows across different disciplines, yet their potential in dis…

13:00 JSTLLM/生成AI

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs

Ensuring robust safety alignment while preserving utility is critical for the reliable deployment of Large Language Models (LLMs). However,…

13:00 JSTエージェント

An Information-Flow Perspective on Explainability Requirements: Specification and Verification

Explainable systems expose information about why certain observed effects are happening to the agents interacting with them. We argue that…

13:00 JST画像/動画生成

STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification

Responding to rising global food security needs, precision agriculture and deep learning-based plant disease diagnosis have become crucial.…

13:00 JSTエージェント

Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships

Autonomous navigation in maritime domains is accelerating alongside advances in artificial intelligence, sensing, and connectivity. Opaque…

13:00 JSTロボティクス

VGGT-DP: Generalizable Robot Control via Vision Foundation Models

Visual imitation learning frameworks allow robots to learn manipulation skills from expert demonstrations. While existing approaches mainly…

13:00 JST研究/論文

Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?

Understanding and modeling the relationship between language and sound are essential for applications such as music information retrieval,…

13:00 JST研究/論文

Monotone and Separable Set Functions: Characterizations and Neural Models

Motivated by applications for set containment problems, we consider the following fundamental problem: can we design set-to-vector function…

13:00 JST研究/論文

CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution

Studying the cellular architecture of the human cerebral cortex is essential for understanding how the brain is organized from the micro to…

13:00 JST画像/動画生成

Robust Motion Generation using Part-level Reliable Data from Videos

Extracting human motion from large-scale web videos offers a scalable solution to the data scarcity issue in character animation. However,…

13:00 JST研究/論文

Towards Reproducibility in Predictive Process Mining: SPICE -- A Deep Learning Library

In recent years, Predictive Process Mining (PPM) techniques based on artificial neural networks have evolved as a method for monitoring the…

13:00 JSTLLM/生成AI

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-…

13:00 JSTLLM/生成AI画像/動画生成ClaudeGemini

見ることと信じること: 直観に反するシーンにおけるオープンソース MLLM の言語バイアスの評価

マルチモーダル大規模言語モデル (MLLM) は、主流の視覚理解タスクにおいて顕著なパフォーマンスを示していますが、日常の常識に反するアクション シーンを処理する能力は依然として十分にテストされていません。このギャップに対処するために、「ウサギがトラを追いかけている」など、視覚的証拠が常識的な予想に明らかに矛盾する、直感に反する視覚的アクションに焦点を当てた 400 の高忠実度の合成シーンで構成されるベンチマークである CAIT を導入します。私たちは人間の主要な独自モデル (Claude や Gemini など) と 14 の代表的なオープンソース MLLM を評価します。人間はほぼ完璧なパフォーマンス (約 0.95 の精度) を達成し、独自のモデルは堅牢な理解を示し (最大 0.88 の精度を達成)、標準的なオープンソースの命令調整モデルは確率レベルでパフォーマンスを発揮します。さらなる分析により、この失敗は事前の強力な言語によって引き起こされていることが示されています。つまり、視覚入力を信頼するのではなく、統計的に一般的なテキストの説明で異常な視覚信号を自動的にオーバーライドします。思考連鎖推論メカニズムを導入すると精度は向上しますが、応答が大幅に遅くなり、新たな障害モードが生成されます。モデルはシナリオを考えすぎて、現実世界の物理法則に違反するという理由だけで実際のビジュアル コンテンツの受け入れを拒否します。最後に、ターゲットを絞った微調整と構造化されたプロンプトによって、この事前言語への依存を効果的に軽減でき、オープンソース モデルが実際の視覚的証拠に基づいて推論を正確に実行できることを実証します。

原文 (English)

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream visual understanding tasks, but their ability to process action scenes that contradict everyday common sense remains undertested. To address this gap, we introduce CAIT, a benchmark comprising 400 high-fidelity synthetic scenes focused on counter-intuitive visual actions, such as ``a rabbit is chasing a tiger'', where visual evidence explicitly contradicts common-sense expectations. We evaluate human, leading proprietary models (e.g., Claude and Gemini), and 14 representative open-source MLLMs. Humans achieve near-perfect performance (around 0.95 accuracy) and proprietary models demonstrate robust understanding (achieving up to 0.88 accuracy), standard open-source instruction-tuned models perform at the chance level. Further analysis demonstrates that this failure is driven by a strong language prior: rather than trusting the visual input, they automatically override the anomalous visual signals with statistically common text descriptions. Although introducing Chain-of-Thought reasoning mechanisms can improve accuracy, it significantly slows down the response and generates a new failure mode: models overthink the scenario and refuse to accept the actual visual content simply because it violates real-world physical laws. Finally, we demonstrate that targeted fine-tuning and structured prompting can effectively mitigate this reliance on language priors, enabling open-source models to accurately ground their reasoning in actual visual evidence.

13:00 JST研究/論文

最小意思決定ダイナミクスと文脈的確率: 量子綱引きモデル

意思決定には、古典的な確率理論に疑問を呈するコンテキスト依存性が見られることがよくあります。この論文は、綱引き (QTOW) 意思決定モデルの量子的な拡張を開発し、そのようなコンテキスト依存性が単一の最小内部状態によっていつ表現されるかを明らかにします。 QTOW 構築では、qutrit の内部状態、保存を保持する更新、および測定に起因する外乱を使用して、1 つのコヒーレントな状態空間内でのモデル決定、学習、およびプローブ操作を行います。この最小限の表現内で、KCBS タイプのプローブ コンテキストを構築することができ、非コンテキストの古典的な非埋め込み可能性の証拠が得られます。主な主張は、量子理論が意思決定から独自に、または仮定に依存せずに導出されるということではありません。むしろ、同じ操作ファミリーの古典的な再構築には、追加の文脈記憶、履歴依存、または拡大された隠れ状態表現が必要です。したがって、文脈的確率は最小限の決定ダイナミクスのリソース署名として現れますが、量子確率はこの構造のコンパクトでメモリ効率の高い実現を提供します。

原文 (English)

Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model

Decision making often exhibits context dependence that is difficult to accommodate within a single non-invasive classical probability model. This paper develops a quantum-like extension of the Tug-of-War (QTOW) decision-making model to ask when such context dependence can be represented by a single constrained internal state. The QTOW construction uses a qutrit state, a state- disturbing generalized decision instrument, decision- and reward-conditioned norm-preserving feedback, and optional probing operations within one state space. The qutrit representation admits KCBS-type probe families and states that violate a non-contextuality bound, providing a witness that the specified operation family cannot be embedded in a single non-contextual classical prob- ability space. The claim is not that quantum theory is uniquely derived from decision making. Classical reconstructions can retain descriptive adequacy by introducing explicit context labels, stored history, or enlarged state descriptions; alternatively, the non-contextuality witness can be avoided by restricting the admissible probe set. Quantum probability instead supplies a compact single-state realization of the contextual operation family considered here. From the perspective of natural and artificial intelligence, the result identifies an architecture-level representational trade-off: contextual information may be carried intrinsically by transformations of a shared internal state or externalized into additional state and memory resources.

13:00 JSTLLM/生成AI画像/動画生成

TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recognition and semantic understanding, yet precise co…

13:00 JST画像/動画生成研究/論文

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained b…

13:00 JSTLLM/生成AI

Ad Insertion in LLM-Generated Responses

Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on…

13:00 JST研究/論文

Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging

Large language models are increasingly trained in continual or open-ended settings, where the total training horizon is not known in advanc…

13:00 JSTエージェント

ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents

Long-horizon reinforcement learning for information seeking agents remains difficult because terminal rewards reveal whether the final answ…

13:00 JSTLLM/生成AI

PatientHub: A Unified Framework for Patient Simulation

As Large Language Models increasingly power role-playing applications, simulating patients has become a valuable tool for training counselo…

13:00 JSTLLM/生成AI

You Can Learn Tokenization End-to-End with Reinforcement Learning

Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general tren…

13:00 JST画像/動画生成ロボティクス

VLANeXt: Recipes for Building Strong VLA Models

Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understa…

13:00 JST画像/動画生成エージェント

ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents

Training-free KV cache compression is essential for deploying vision-language GUI agents under memory and latency constraints, yet existing…

13:00 JSTLLM/生成AILlama

EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training

Large language models (LLMs) are predominantly trained on English-centric data, resulting in uneven performance for smaller languages. We s…

13:00 JSTLLM/生成AIビジネス/資金調達ClaudeGPT / ChatGPTGeminiLlama

ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models

Most adversarial evaluations of large language model (LLM) safety assess single prompts and report binary pass/fail outcomes, which fails t…

13:00 JST研究/論文

msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models

Time series foundation models (TSFMs) require diverse, real-world datasets to adapt across varying domains and temporal frequencies. Howeve…

13:00 JSTLLM/生成AIビジネス/資金調達

オーマニック: 大規模言語モデルにおけるマルチホップ推論の段階的評価に向けて

最終的な回答のみから大規模言語モデル (LLM) の推論能力を評価すると、特にステップレベルのアノテーションがないマルチホップ QA ベンチマークでは、中間ステップでの失敗がわかりにくくなる可能性があります。このギャップに対処するために、最終回答の精度を測定するだけでなく、推論がどこで破綻しているかを診断するように設計されたオープンドメインの 4 ホップ QA ベンチマークである Omanic を導入します。 Omanic には、10,296 個の機械生成トレーニング サンプル (OmanicSynth) と、専門家がレビューした人間による注釈付きの 967 個の評価サンプル (OmanicBench) が含まれており、各評価質問はシングルホップのサブ質問、中間回答、および構造化グラフ トポロジに分解されています。独自のオープンソース LLM を使用した実験では、Omanic が困難であることが示されていますが、段階的な分析では、後のホップのボトルネック、事実の知識フロア、推論チェーンに沿ったエラーの伝播が明らかになりました。 OmanicSynth の微調整により、6 つの推論および数学ベンチマークに移行し、平均 7.41 ポイントの向上が得られ、推論能力移転の監視としての有効性が検証されました。データは https://huggingface.co/datasets/li-lab/Omanic で、コードは https://github.com/XiaojieGu/Omanic でリリースされます。

原文 (English)

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce Omanic, an open-domain 4-hop QA benchmark designed not only to measure final-answer accuracy but also to diagnose where reasoning breaks down. Omanic contains 10,296 machine-generated training examples (OmanicSynth) and 967 expert-reviewed human-annotated evaluation examples (OmanicBench), with each evaluation question decomposed into single-hop sub-questions, intermediate answers, and structured graph topologies. Experiments with proprietary and open-source LLMs show that Omanic is challenging, while step-wise analysis reveals a later-hop bottleneck, factual knowledge floor, and error propagation along reasoning chains. Fine-tuning on OmanicSynth transfers to six reasoning and mathematics benchmarks, yielding a 7.41-point average gain and validating its effectiveness as supervision for reasoning-capability transfer. We release the data at https://huggingface.co/datasets/li-lab/Omanic and the code at https://github.com/XiaojieGu/Omanic.

13:00 JSTエージェント

Beyond OAuth: Task-Scoped Authorization for AI Agents via Natural Language Slices

AI agents increasingly execute users' natural-language (NL) tasks by calling Web services, yet today's Web authorizes these calls through O…

13:00 JST研究/論文

Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification

Network Traffic Classification (NTC) increasingly relies on data-driven models, yet its practical deployment is often constrained by limite…

13:00 JST研究/論文

Ollivier-Ricci Curvature of Riemannian Manifolds and Directed Graphs with Applications to Graph Neural Networks

This thesis is an exposition of Ollivier-Ricci Curvature of metric spaces as introduced by Yann Ollivier, which is based upon the 1-Wassers…

13:00 JST研究/論文

Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts

Electroencephalography (EEG) foundation models have shown strong potential for learning generalizable representations from large-scale neur…

13:00 JST画像/動画生成

RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction

The use of CT imaging is important for screening, diagnosis, therapy planning, and prognosis of lung cancers. Unfortunately, due to differe…

13:00 JSTロボティクス

Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters

Despite significant advances in Reinforcement Learning (RL), model performance remains highly sensitive to algorithm and hyperparameter con…

13:00 JSTエージェント

Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval

Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue exp…

13:00 JST画像/動画生成

Outlier-Robust Diffusion Solvers for Inverse Problems

Methods based on diffusion models (DMs) for solving inverse problems (IPs) have recently achieved remarkable performance. However, DM-based…

13:00 JST画像/動画生成エージェント

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mec…

13:00 JSTエージェントロボティクス

ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching

Existing imitation learning methods enable robots to interact autonomously with the physical environment. However, contact-rich manipulatio…

13:00 JSTLLM/生成AI

トーナメント GRPO: オープンエンドの長編生成における強化学習に対するグループごとのトーナメント報酬

信頼できる参照回答や自動メトリクスが利用できないことが多いため、オープンエンド型長文生成における強化学習は困難です。既存のルーブリックベースの手法は通常、点ごとの LLM-as-a-judge スコアリングに依存していますが、絶対スコアは複雑な応答全体で調整するのが難しく、同じクエリのロールアウト間の区別が弱い可能性があり、最適化中に飽和状態になる可能性があります。私たちは、同じクエリのロールアウト間で繰り返される複数ラウンドのトーナメントを通じて、ルーブリックに基づいた LLM の判断を相対的な報酬に変換するグループごとの報酬フレームワークであるトーナメント GRPO を提案します。トーナメント GRPO は、グループ内の候補者を比較し、トーナメントの結果を蓄積して、GRPO トレーニングに対するグループごとの報酬に正規化します。 Deep Research Bench での実験では、Tournament-GRPO が既存の報酬設計ベースラインを常に上回り、最も強力なベースラインを上回る総合スコア 4.52 ポイントの向上を達成したことが示されています。さらなる分析により、トーナメントの報酬は有利な効果、つまり効率のトレードオフを提供し、トーナメントのデザインがトレーニングのダイナミクスに影響を与えることが示されています。これらの結果は、ルーブリックに基づいたトーナメント比較が、オープンエンドの長編生成における強化学習に効果的な報酬シグナルを提供することを示唆しています。

原文 (English)

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation

Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often unavailable. Existing rubric-based methods typically rely on pointwise LLM-as-a-judge scoring, but absolute scores are difficult to calibrate across complex responses, may provide weak discrimination among same-query rollouts, and can become saturated during optimization. We propose Tournament-GRPO, a group-wise reward framework that converts rubric-guided LLM judgments into relative rewards through repeated multi-round tournaments among same-query rollouts. Tournament-GRPO compares candidates within groups, accumulates tournament outcomes, and normalizes them into group-wise rewards for GRPO training. Experiments on Deep Research Bench show that Tournament-GRPO consistently outperforms existing reward-design baselines, achieving a 4.52-point overall-score improvement over the strongest baseline. Further analyses show that tournament rewards provide a favorable effectiveness--efficiency trade-off and that tournament design affects training dynamics. These results suggest that rubric-guided tournament comparison provides an effective reward signal for reinforcement learning in open-ended long-form generation.

13:00 JSTLLM/生成AI

LLM 推論のためのスキル条件付きゲート自己蒸留

ポリシー上の自己蒸留 (SD) は、教師側の特権情報 (PI) を使用して、疎な検証者の結果を密なトークンレベルの監視に変えることで、LLM 推論を改善します。既存のメソッドは通常、参照回答や成功したトレースなど、信頼できる PI を前提としています。私たちは、PI を経験由来のスキルバンクから取得できるかどうかを尋ねます。取得されたスキルはコンパクトで再利用可能ですが、無関係または誤解を招く可能性もあります。我々は、無条件の模倣ではなく教師の仮説検証としてスキルベースの SD を定式化する、スキル条件ゲート型自己蒸留 (SGSD) を提案します。 SGSD はスキルと間違いのペアを取得し、複数の教師のプールを構築し、スキル条件を備えたすべての教師が同じ単純なプロンプトの生徒ロールアウトを採点できるようにします。検証者は各教師の極性を検証します。つまり、成功をサポートするか失敗を抑制する場合は積極的な監督を行い、反対の立場は逆転します。堅牢なゲート目標により、不確実または極端なシグナルを抑制しながら、有益な教師と生徒の意見の相違が抽出されます。複数の数学的推論ベンチマークの実験では、SGSD が GRPO より一貫して改善しており、より弱い PI 仮定の下で、回答条件付き OPSD との競争力を維持していることが示されています。たとえば、Qwen3-1.7B では、SGSD は、AIME24、AIME25、および HMMT25 で平均して GRPO を 6.2%、OPSD を 1.7% 上回っています。私たちのコードは https://github.com/walawalagoose/SGSD で入手できます。

原文 (English)

Skill-Conditioned Gated Self-Distillation for LLM Reasoning

On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense token-level supervision. Existing methods usually assume trusted PI, such as reference answers or successful traces. We ask whether PI can instead come from an experience-derived skill bank, where retrieved skills are compact and reusable but may also be irrelevant or misleading. We propose Skill-Conditioned Gated Self-Distillation (SGSD), which formulates skill-based SD as teacher hypothesis validation rather than unconditional imitation. SGSD retrieves skill-mistake pairs, constructs a multi-teacher pool, and lets all skill-conditioned teachers score the same plain-prompt student rollout. The verifier validates each teacher's polarity: supporting a success or suppressing a failure gives positive supervision, while the opposite stance is reversed. A robust gated objective then distills informative teacher-student disagreements while suppressing uncertain or extreme signals. Experiments on multiple mathematical reasoning benchmarks show that SGSD consistently improves over GRPO and remains competitive with answer-conditioned OPSD under a weaker PI assumption. For example, on Qwen3-1.7B, SGSD outperforms GRPO by 6.2% and OPSD by 1.7% on average on AIME24, AIME25, and HMMT25.

13:00 JSTLLM/生成AI

検出 vs. 実行: シングルバケットプローブは Mamba-2 状態シンクの半分を見逃す

機構的な解釈では、表現シグネチャを識別するプローブが、対応する計算を実行する回路も識別すると仮定することがよくあります。 Mamba-2 ではこの仮定が体系的に失敗する可能性があることを示します。状態シンク (アテンション シンクに似た、境界トークン上の不均衡なデルタゲート アクティベーション) を研究すると、単一バケット プローブは小さな実行層のみを回復し、同じ表現シグネチャを持つはるかに大きな検出層を見逃していることがわかります。 Mamba-2 では、ステート シンクは 2 つの機能ヘッド セットに分解されます。シングルバケット BOS スペシャリスト ヘッド (2.7B のヘッドの約 5%) は、モデル スケールとコーパス全体で BOS コンテキストと改行ターゲットの両方の予測を因果的にサポートします。デュアルヘッド (ヘッドの 27 ~ 35%、同じプローブのマルチクラス集約によって回復) は、BOS と改行の表現上の類似性がより強くなりますが、アブレーション下では因果関係が大幅に弱いことが示されます。表現上の類似性は、機能的な同等性を意味するものではありません。この違いは下流の動作にとって重要です。BOS スペシャリスト ヘッドをアブレーションすると、Mamba-1 2.8B と Mamba-2 2.7B の両方で 1024 コンテキスト長での RULER NIAH 検索精度が 1.00 から 0.00 に低下しますが、サイズが一致した補体はベースラインのパフォーマンスを維持します。ランダムなチャネルバケット制御は基板の粒度のみを除外し、Mamba-2 のヘッド共有デルタ投影を示唆します。プローブ由来の特殊機能により、実行回路を特定できます。粗い粒度では、同じプローブが検出回路も回復します。検出回路を分離するには、クラス条件付きコサインではなくクラス条件付きアブレーションが必要です。

原文 (English)

A Circuit, Not The Circuit: Non-Unique Causal Localisation of the Mamba-2 State Sink

Mechanistic interpretability routinely reads a probe and labels its top-activating units as the circuit executing the computation. We test the move in Mamba, on the state sink: the selective state-space analogue of the Transformer attention sink, where the Delta-gate fires disproportionately on boundary tokens such as BOS and newline. At Mamba-1 channel granularity the probe's units carry the causal effect. At Mamba-2 head granularity the label-to-locus link breaks in three ways. The causal set is not unique: a near-disjoint set of top-activation heads, sharing almost none of the specialists, matches or exceeds their ablation effect. Representation does not track function: dual heads, a quarter to over a third of all heads, respond alike to BOS and newline yet do not reproduce the specialists' causal profile. And the effect dissociates by intervention surface: perturbing the input-dependent Delta the probe reads moves the loss by at most 0.16 nats, while suppressing the same heads' output moves it by 1.5 to 7.3 nats. The heads still matter behaviourally, dropping needle-retrieval success on a RULER-style probe from 30/30 to 0/30 at 1024 context on both flagship checkpoints. A representational signature in a selective state-space model thus marks a behaviourally important head set without pinning down the circuit: a single probe recovers a circuit, not the circuit.

13:00 JSTLLM/生成AIエージェント

SaliMory: 会話エージェントの認知記憶を調整する

生涯の伴侶として機能する会話エージェントは、すべての対話にわたって永続的な記憶を維持する必要があります。ただし、生の取得でコンテキスト ウィンドウを単純に拡張すると推論の品質が低下し、標準の強化学習による記憶エージェントのトレーニングでは、多段階パイプラインで深刻なクレジット割り当てのボトルネックが発生します。これを解決するために、単一言語モデルをトレーニングして、ユーザーの事実、好み、作業記憶にまたがる認知的に構造化された記憶を管理するフレームワークである SALIMORY を紹介します。 SALIMORY は、階層的な段階ごとのプロセス報酬と報酬分解された対照的洗練を導入することにより、個別の記憶操作 (選択的フィルタリング、統合、およびキュー主導のリコール) をエンドツーエンドで個別に監視します。 SALIMORY はメモリに起因する障害を 3 分の 1 に削減し、エンドツーエンドの精度で最先端のものを 10% 以上上回り、Good Personalization 率を 2 倍以上に高めます。

原文 (English)

SaliMory: Orchestrating Cognitive Memory for Conversational Agents

Conversational agents that serve as lifelong companions must maintain persistent memory across all interactions. However, simply expanding context windows with raw retrieval degrades reasoning quality, while training memory agents via standard reinforcement learning creates a severe credit assignment bottleneck in a multi-stage pipeline. To solve this, we introduce SALIMORY, a framework that trains a single language model to manage a cognitively-structured memory-spanning user facts, preferences, and working memory. By introducing a hierarchical stage-wise process reward and reward-decomposed contrastive refinement, SALIMORY provides isolated supervision for distinct memory operations (selective filtering, consolidation, and cue-driven recall) end-to-end. SALIMORY cuts memory-attributed failures by one-third, outperforms the state-of-the-art by over 10% in end-to-end accuracy, and more than doubles the Good Personalization rate.

13:00 JSTLLM/生成AIGemma

1 つのニューロンを編集すると LLM の繰り返しループを修正できますか?

はい。ドゥームループを治すことはできるでしょうか?おそらくそうではありません。 Gemma 4 の命令調整モデルには、再現性のある失敗が共通しています。テレビ シリーズの各エピソード、IAU 88 星座、オリジナルのポケモン 151 匹のリストなど、長い事実に基づく列挙プロンプトでは、厳密な逐語的ループか、エントリが 1 つの答えに減衰するリストの繰り返しに崩壊します。これらのループは 95% もの高率で発生し、即時的な言い換え、推論エンジンの変更、およびほとんどのサンプリング調整の後も存続します。この論文では、この動作が重み付け編集によって削除できるほど局所的であるかどうかを調査します。原因を特定するために、層ごとのアブレーションとニューロンごとの属性を使用し、全世代スイープで最も有力な候補を確認します。ループは、MLP ニューロンの小さなセット (または、26B-A4B 専門家混合モデルでは、少数のルーティングされた専門家) をトレースし、静的な重み編集で抑制します。これらの「手術」は、単一の符号反転ニューロン (E2B モデルの場合) と同じくらい小さい可能性があります。効果的な編集のサイズはモデルの規模に応じて大きくなりますが、いずれの場合も、汎用ベンチマーク スコアを維持しながら、通常の生成バジェットでループ パターンに対処できます。ただし、この編集ですべてが解決するわけではありません。私たちは、より長い思考予算についても研究しています。この場合、2 つの大きなモデルが明らかに破滅ループ、つまり、モデルが思い出せない事実をめぐって自己修正を繰り返し、最終的な答えを約束することなく予算を使い果たす非収束体制に陥っています。我々は、この残留故障が同じ編集によって減少するものの除去されないことを示し、それが除去可能な回路ではなく、基本的に知識の精度の問題であると主張します。体重手術はループを削除できますが、欠落している事実を提供することはできません。私たちの結果は、実現可能性の実証、つまり、具体的な生成の病理をいくつかのパラメーターに局所化して編集できることの証拠であると同時に、そのアプローチがどこで終わるのかを示すものでもあります。

原文 (English)

When Can One Neuron Fix Repetition Loops in LLMs?

The Gemma 4 instruction-tuned models share a reproducible failure: on long factual enumeration prompts, such as TV episodes, the 88 IAU constellations, or the 151 original Pokemon, they collapse into repetition, either a tight verbatim loop or a list whose entries decay onto one answer. These loops reach 87.5% (7/8 generations) and survive prompt rewording and most sampling adjustments. In this paper, we explore whether edits to a few internal model components can directly reduce this failure, without relying on repetition penalties, which can distort valid repetition and degrade task performance. To locate such targets, we combine per-layer ablation with per-neuron or routed-expert attribution, then evaluate weight edits over complete generations. We find that these edits substantially reduce detected loops on the prompts and seeds used to select them; in Gemma 4 E2B, for example, one sign-inverted neuron suffices. Across all four Gemma models, detected loops fall from 46/384 to 12/384 on frozen held-out prompts and seeds, driven mainly by E4B and 31B, while general-purpose benchmarks show no statistically detectable regressions. Our attribution methodology identifies useful candidates, but rankings vary across examples. At longer generation budgets, edits remain effective for E2B and E4B, whereas remaining failures in 26B and 31B shift toward doom looping: non-convergent self-correction over facts the model cannot recall. In exploratory experiments on Qwen3.5 and LFM2.5, sparse edits also reduce repetition, providing preliminary cross-family evidence, although effect strength and selectivity vary. Overall, our results show the promise and limits of targeted, small-scale weight editing: it can suppress specific repetition failures and provide a training-free causal intervention, but does not reveal a universal loop circuit, guarantee clean termination, or supply missing knowledge.

13:00 JSTLLM/生成AI

TW-LegalBench: Measuring Taiwanese Legal Understanding

Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal…

13:00 JSTロボティクス

RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewa…

13:00 JSTLLM/生成AI画像/動画生成

拡散アンラーニングで同時発生する関連保持概念

アンラーニングは、拡散モデルにおける有害なコンテンツの生成を軽減するための重要な技術として浮上しました。しかし、既存の手法では、ターゲットの概念だけでなく、共起する無害な概念も削除してしまうことがよくあります。図 1 に示すように、ヌードを学習しないと、人物の概念が意図せず抑制され、モデルが人物を含む画像を生成できなくなります。私たちは、これらの望ましくない抑制された、保存する必要がある共起概念を CARE (Co-occurring Associated REtained Concepts) と定義します。次に、非学習タスク全体での保存を直接定量化する一般的な指標である CARE スコアを導入します。これを基盤として、対象概念のみを消去しながらCAREを明示的に保護するフレームワークであるReCARE(Robust Erasure for CARE)を提案します。 ReCARE は、ターゲット画像から抽出された無害な共起トークンの厳選された語彙である CARE セットを自動的に構築し、安定した非学習のためのトレーニング中にこの語彙を活用します。さまざまなターゲット コンセプト (ヌード、ゴッホ スタイル、テンチ オブジェクト) にわたる広範な実験により、ReCARE が堅牢なコンセプトの消去、全体的な実用性、および CARE の保存のバランスにおいて全体的に最先端のパフォーマンスを達成することが実証されました。

原文 (English)

Co-occurring Associated REtained concepts in Diffusion Unlearning

Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co-occurring concepts. As illustrated in Fig.1, unlearning nudity can unintentionally suppress the concept of person, preventing a model from generating images with person. We define these undesirably suppressed co-occurring concepts that must be preserved CARE (Co-occurring Associated REtained concepts). Then, we introduce the CARE score, a general metric that directly quantifies their preservation across unlearning tasks. With this foundation, we propose ReCARE (Robust erasure for CARE), a framework that explicitly safeguards CARE while erasing only the target concept. ReCARE automatically constructs the CARE-set, a curated vocabulary of benign co-occurring tokens extracted from target images, and leverages this vocabulary during training for stable unlearning. Extensive experiments across various target concepts (Nudity, Van Gogh style, and Tench object) demonstrate that ReCARE achieves overall state-of-the-art performance in balancing robust concept erasure, overall utility, and CARE preservation.

13:00 JSTLLM/生成AIエージェント研究/論文

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of…

13:00 JSTビジネス/資金調達

A Unified Algebraic Framework for Classification Performance Evaluation

We propose a unified algebraic framework for classification performance evaluation covering binary, multiclass, multilabel, ordinal, hierar…

13:00 JSTLLM/生成AIエージェント

Eluna: 推論とタスク実行により倉庫業務を自動化するエージェント LLM システム

倉庫業務は、複雑なマルチシステムの意思決定ロジックをエンコードした標準運用手順 (SOP) によって管理されており、厳格な時間制約の下で確実に実行する必要がありますが、LLM エージェントには手順の遵守を強制するメカニズムが欠けており、完全な SOP 仕様が導入するコンテキストの過負荷の下では機能が低下します。信頼性の高い SOP 実行のための実稼働環境に導入されたエージェント システムである Eluna を紹介します。 Eluna は、段階的な開示を備えた有向非循環グラフとして SOP をエンコードし、独立したタスクを並列サブエージェントに委任する、グラフガイド型のマルチエージェント フレームワークです。各サブエージェントは永続的なコード実行とライブ データ アクセスを備えています。本番のレイテンシと精度のニーズを満たすために、非対称のエピソード蒸留を使用します。この方法では、強力な教師がエピソード的なエラー記憶を通じて改善され、その後、小さな生徒が記憶を取り除いて修正された軌道に基づいて微調整され、推論時間のオーバーヘッドなしで修正を内部化します。 13 タスクのベンチマークと 2 つの本番アプリケーションで、当社の微調整されたモデルは教師と同等かそれを上回り、より大きな既製のベースラインをすべて上回り、チケット処理アプリケーションに関して専門家による 94% の合意に達しました。

原文 (English)

Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution

Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural compliance and degrade under the context overload full SOP specifications introduce. We present Eluna, a production-deployed agentic system for reliable SOP execution. Eluna is a graph-guided, multi-agent framework that encodes SOPs as directed acyclic graphs with progressive disclosure and delegates independent tasks to parallel sub-agents, each with persistent code execution and live data access. To meet production latency and accuracy needs, we use asymmetric episodic distillation where a strong teacher is improved through episodic error memories, then a smaller student is fine-tuned on the corrected trajectories with memory stripped, internalizing corrections without inference-time overhead. On a 13-task benchmark and two production applications, our fine-tuned models match or exceed their teacher, beat all larger off-the-shelf baselines, and reach 94% expert agreement on the ticket processing application.

13:00 JST研究/論文

アムステルダムのカフェ: 現職者がオラクルになるとき

現場は、既存の実装とは無関係に要求が述べられている場合には、自由に計算を再定式化できますが、既存の独自の出力が静かに仕様になった場合には、それができないことに気づきます。このメモは、現代のアクセラレータの計算再定式化に関するレンズとしての観察を提供します。ハードウェアに優しい形式で問題を提起すると、速度とエネルギーが大幅に向上する可能性がありますが、それは代替品を判断できる場合に限ります。テストオラクル問題 (Weyuker、Barr et al.)、要件工学の実装バイアスの概念 (Zave と Jackson)、およびルーフライン パフォーマンス モデルに基づいて、この病理を「ベースライン キャプチャ」と名付けています。つまり、既存企業が要求を満たすことができるという証拠ではなくなり、要求を満たすことの定義になる瞬間です。そして、混同されやすい 2 つの質問、つまり再定式化を判断できるかどうか (既存企業に依存しない要求の存在をオンにする) を分離しています。そしてその発見を自動化できるかどうか(その需要を評価するコストもさらにかかります)。短いケース (最短パス ルーティング、学習可能なオーディオ フロントエンド、Ed25519 署名検証用の ZIP-215、気候モデル用の CESM-ECT、および単一 GEMM オーディオ フロントエンド) は、要求を明示的かつ運用可能で、既存企業から独立させるという「検証者を購入する」パターンと動きを示しています。単独で新規性を主張される成分はありません。貢献は統合であり、再定式化の結果について尋ねやすくする単一の質問です。その受け入れテストでは、既存の成果物について言及していますか?

原文 (English)

The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle

A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds itself unable to when the incumbent's own output has quietly become the specification. This note offers that observation as a lens on computational reformulation for modern accelerators, where posing a problem in a hardware-friendly form can yield large speed and energy gains, but only if a replacement can be judged at all. Building on the test-oracle problem (Weyuker; Barr et al.), on requirements engineering's notion of implementation bias (Zave and Jackson), and on the case for judging approximate designs by acceptability rather than numerical proximity (Felzmann et al.), it names the pathology "baseline capture": the moment an incumbent stops being evidence that a demand can be met and becomes the definition of meeting it. It then separates two questions that are easily confused: whether a reformulation can be judged at all, which turns on the existence of an incumbent-independent demand, and whether its discovery can be automated, which turns additionally on the cost of evaluating that demand. Short cases -- shortest-path routing, learnable audio frontends, ZIP-215 for Ed25519 signature validation, CESM-ECT for climate models, and a single-GEMM audio frontend measured at 1.64x-3.29x speedup and up to 3.03x less energy -- illustrate the pattern and the move of "buying a verifier": making a demand explicit, operational, and independent of the incumbent. No component is claimed novel in isolation; the contribution is the synthesis and the single question it makes easy to ask of any reformulation result -- does its acceptance test mention the incumbent's output?

13:00 JSTLLM/生成AI

離散拡散モデル: トークン化から生成までの統一フレームワーク

離散ノイズ除去拡散モデル (DDM) は、離散データの自己回帰 (AR) モデリングに代わる有力な代替手段として最近登場し、並列生成と反復的なグローバル リファインメント機能を提供します。状態空間が固定されている連続拡散とは異なり、DDM は基本的に、トークン化スキーム、語彙トポロジー、ドメイン固有の構造アルファベットなどの離散状態空間の構築方法によって形成されます。この研究では、基礎となる離散状態空間の構築を通じて離散拡散モデルを捉える統一された概念フレームワークが導入されています。このフレームワーク内では、遷移マトリックス、マスキング/状態吸収、およびスコア/比率ベースのアプローチを含む既存の定式化が、共通の設計空間のさまざまなインスタンス化として現れます。このフレームワークはさらに、トレーニング目標、推論アルゴリズム、スケーリング動作、システムの最適化、評価プロトコルにわたる一般的な設計のトレードオフを明らかにし、将来の研究に向けたいくつかの有望な方向性を示唆しています。

原文 (English)

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokenization scheme, the vocabulary topology, and domain-specific structural alphabets. This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space. Within this framework, existing formulations, including transition-matrix, masking/absorbing-state, and score/ratio-based approaches, emerge as different instantiations of a common design space. The framework further exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols, suggesting several promising directions for future research.

13:00 JST画像/動画生成ビジネス/資金調達研究/論文

EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrati…

13:00 JSTLLM/生成AI画像/動画生成

GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts…

13:00 JST研究/論文Llama

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Atte…

13:00 JST研究/論文Claude

MOSAIC: 安全な AI 計算のマスクされたアウトソーシング

私たちは、クライアントが入力とモデルの両方を保持し、サーバーがどちらも学習する必要がないという設定で、信頼できるが計算能力が弱いクライアントから、信頼できないが強力なサーバーに AI 計算を安全かつ効率的にアウトソーシングするという課題に取り組みます。我々が紹介する MOSAIC は、その核となる新しい行列乗算マスキング プロトコルであり、以前の研究よりもはるかに大きな行列にスケールし、大規模なトランス推論などの最新のワークロードの安全なアウトソーシングを可能にします。 MOSAIC は、乗算結果に少量のノイズを導入し、正確性を緩和することにより、最適な漸近クライアント オーバーヘッドと具体的な実行時間を従来の作業よりも桁違いに高速化します。そのセキュリティは、決定的な LWE および LPN の仮定に限定されます。このノイズは変圧器の多くの層にわたって蓄積されるため、主な技術的課題は誤差の増加を制限することです。 MOSAIC は、ランダムなアダマール回転に基づく誤差スケーリング メカニズムを使用してこの問題に対処します。大型の 70B トランス モデルでは、MOSAIC の複雑さは一般的な量子化アプローチに匹敵し、HumanEval での完全精度の BF16 推論にも匹敵します。最後に、MOSAIC のようなアイデアが現代のデータセンターにおける大規模な機密 AI への道をどのように約束できるかを示すエンドツーエンドの実装を紹介します。非機密推論は、RDMA のようなネットワーキングを使用してアクティベーション、キャッシュされた KV 値、および重みをノード間で移動することで、異種ハードウェアの利用を最大化するためにフェーズ (プリフィル/デコード)、レイヤー、および時間を超えてすでに分散されています。 MOSAIC は、トラステッド コンピューティング ベース (TCB) を小規模に保ち、AI 計算の大部分を信頼できないアクセラレータにアウトソーシングすることで、機密コンピューティングのスケーリングを可能にします。

原文 (English)

MOSAIC: Masked Outsourcing of Secure AI Computations

We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the server must learn neither. We present MOSAIC, whose core is a novel matrix-multiplication masking protocol that scales to far larger matrices than prior work, enabling the safe outsourcing of modern workloads such as large transformer inference. By introducing small amounts of noise to the multiplication result and thereby relaxing correctness, MOSAIC achieves optimal asymptotic client overhead and concrete runtimes orders of magnitude faster than prior work. Its security reduces to the decisional LWE and LPN assumptions. Because this noise accumulates across the many layers of a transformer, a key technical challenge is bounding error growth; MOSAIC addresses this with an error-scaling mechanism based on random Hadamard rotations. On large 70B transformer models, MOSAIC's perplexity is comparable to popular quantization approaches and even matches full-precision BF16 inference on HumanEval. Finally, we present an end-to-end implementation showing how ideas like MOSAIC can promise a path towards large-scale confidential AI in modern data centers. Non-confidential inference is already distributed across phase (prefill/decode), layer, and time to maximize utilization of heterogeneous hardware, using RDMA-like networking to move activations, cached KV values, and weights across nodes. MOSAIC enables scaling of confidential compute by keeping the trusted computing base (TCB) small and outsourcing the bulk of the AI computation to untrusted accelerators.

13:00 JST研究/論文

TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction

Tabular foundation models, driven by in-context learning, have rapidly grown in quality and popularity. However, recent approaches with eit…

13:00 JST研究/論文

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This…

13:00 JST画像/動画生成ロボティクス

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent mi…

13:00 JSTビジネス/資金調達研究/論文

Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol

AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…

13:00 JST研究/論文

ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning

The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---…

13:00 JST研究/論文DeepSeek

MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention…

13:00 JST画像/動画生成

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines…

13:00 JST研究/論文

Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas

We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq proof assistant. TLS is a combinatorial framewor…

13:00 JSTLLM/生成AIエージェントQwen

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

Audited against policy-conditional ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level…

13:00 JST研究/論文

ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations

Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in…

13:00 JSTLLM/生成AI

Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation

Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based app…

13:00 JST画像/動画生成

Scaling Muon for Diffusion Transformers

The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and…

13:00 JST画像/動画生成エージェント

CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents

Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise…

13:00 JSTLLM/生成AI画像/動画生成

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an int…

13:00 JSTLLM/生成AIエージェント

Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores

Retrieval-augmented generation treats the document store as a frozen input, and the offline pipelines that do build structure over it build…

13:00 JST画像/動画生成ロボティクス

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and ro…

13:00 JST研究/論文

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

Delta-Rule recurrent models maintain a fixed-size state, enabling $O(1)$ inference memory but potentially becoming unstable under extreme-c…

13:00 JST研究/論文

Functional compatibility as a determinant of persistent neural learning

Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclea…

13:00 JST研究/論文

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely v…

13:00 JSTLLM/生成AIエージェント

Molecular LLM Agents: From Architectural Design to Scientific Autonomy

Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, c…

13:00 JST研究/論文

How Much Regularization Survives Averaging? Update Masking in Federated Learning

Federated learning on non-IID data seeks flat minima to generalize across clients, and existing methods borrow sharpness-aware minimization…

13:00 JST画像/動画生成エージェント研究/論文

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from t…

13:00 JSTLLM/生成AI

Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data

Structured data exists in many forms (tables, knowledge graphs, charts, and time series), and converting it into text may involve different…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they captur…

13:00 JSTLLM/生成AI

Best Practice Critic Optimization

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses fo…