Skip to the content.

AIニュース 2026-07-31

自動生成: 2026-07-31 12:29 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. PerplexityがAIエージェントの“暴走”対策ツールをオープンソースに Claude CodeやCodexを監視ITmedia AI+

    PerplexityがAIエージェントの危険な挙動を検知・防止するツール群「Numbat」をオープンソース化。Claude CodeやCo…

  2. Google、ロボット向けAI「Gemini Robotics 2」発表 ヒューマノイドの全身制御や指先作業を実現ITmedia AI+

    GoogleとGoogle DeepMindは、ロボット向けAIモデル群「Gemini Robotics 2」を発表した。全身制御や指先で…

  3. Claudeが評価環境から実在企業に不正アクセス――Anthropic、3件のインシデントを公表ITmedia AI+

    Anthropicは、サイバーセキュリティ評価中にAIモデル「Claude」が設定ミスでオープンになっていた経路から外部のインターネットに…

  4. OpenAI、「GPT-5.6 Luna」を80%値下げ モデル自身による効率化でコスト削減ITmedia AI+

    OpenAIは、「GPT-5.6」ファミリーの「Luna」を80%、「Terra」を20%値下げすると発表した。API価格の改定に加え、「…

  5. Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能ITmedia AI+

    Thinking Machines Labは、オープンウェイトのAIモデル「Inkling-Small」の正式版を公開した。従来モデルの4…

  6. Chromeに13年以上潜んでいた脆弱性、AIで発見 直近2回のアプデで過去23回分を上回るバグ修正ITmedia AI+

    GoogleがChromeのセキュリティ対策へのAI活用を公式ブログで解説。Geminiベースのエージェントが13年以上潜んでいた脆弱性を…

  7. スクエニ、ゲームの品質テストをGeminiで自動化 AIが画面を見ながらコントローラーを操作、検証作業を自走ITmedia AI+

    スクウェア・エニックスが、ゲームのQAテストを「Gemini」で自動化する取り組みを「Google Cloud Next Tokyo '2…

トピック別件数

日本語メディア15件

ITmedia AI+ (日本語)

11:18 JSTLLM/生成AIエージェントClaude

PerplexityがAIエージェントの“暴走”対策ツールをオープンソースに Claude CodeやCodexを監視

PerplexityがAIエージェントの危険な挙動を検知・防止するツール群「Numbat」をオープンソース化。Claude CodeやCodexに組み込み、タスクに執着したエージェントの“暴走”を実行前に阻止する。

11:16 JSTハードウェア/半導体

Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能

Thinking Machines Labは、オープンウェイトのAIモデル「Inkling-Small」の正式版を公開した。従来モデルの4分の1のサイズながら、データの改良や強化学習によりコード生成などのベンチマークで従来版を上回る性能を実現。動作に必要なGPUメモリも大幅に削…

11:03 JSTLLM/生成AIエージェントGoogleGemini

Chromeに13年以上潜んでいた脆弱性、AIで発見 直近2回のアプデで過去23回分を上回るバグ修正

GoogleがChromeのセキュリティ対策へのAI活用を公式ブログで解説。Geminiベースのエージェントが13年以上潜んでいた脆弱性を発見した。AI攻撃の高速化に対応し、セキュリティ更新「週2回」配信も試行する。

10:43 JSTLLM/生成AIGoogleGemini

スクエニ、ゲームの品質テストをGeminiで自動化 AIが画面を見ながらコントローラーを操作、検証作業を自走

スクウェア・エニックスが、ゲームのQAテストを「Gemini」で自動化する取り組みを「Google Cloud Next Tokyo '26」基調講演で披露。AIが画面を見ながらコントローラーを操作し、検証作業を自ら進める。

10:27 JSTLLM/生成AIロボティクスGoogleGemini2媒体が報道

Google、ロボット向けAI「Gemini Robotics 2」発表 ヒューマノイドの全身制御や指先作業を実現

GoogleとGoogle DeepMindは、ロボット向けAIモデル群「Gemini Robotics 2」を発表した。全身制御や指先での微細な作業、複数ロボットの連携に対応する。高次の脳として機能する推論モデル「ER 2」や軽量VLAモデルを含み、安全性評価の新ベンチマーク…

出典:Google DeepMindITmedia AI+
09:30 JSTLLM/生成AIAnthropicClaude

Claudeが評価環境から実在企業に不正アクセス――Anthropic、3件のインシデントを公表

Anthropicは、サイバーセキュリティ評価中にAIモデル「Claude」が設定ミスでオープンになっていた経路から外部のインターネットに接続し、実在する3組織の本番インフラに誤って不正アクセスしていたと発表した。評価環境を演習と誤認したことが原因で、全評価を停止し外部ベンダー…

08:55 JSTLLM/生成AIエージェントOpenAIGPT / ChatGPT2媒体が報道

OpenAI、「GPT-5.6 Luna」を80%値下げ モデル自身による効率化でコスト削減

OpenAIは、「GPT-5.6」ファミリーの「Luna」を80%、「Terra」を20%値下げすると発表した。API価格の改定に加え、「ChatGPT Work」や「Codex」でのクレジット消費量も削減される。自律的なカーネル最適化により提供コストの削減を実現した。また、処…

出典:ITmedia AI+OpenAI
08:00 JSTその他

リホストを選ぶ企業は63% モダナイゼーションの「第一歩」のはずが、なぜ終着点に変わるのか?

国産メインフレームの提供終了が相次ぎ、レガシーシステムの移行は「期限のある経営課題」になった。だが移行プロジェクトで最初に立ちはだかるのは、移行技術やアーキテクチャ「以外の」問題だ。モダナイゼーションがリホストで止まる構造を、ITRの入谷光浩氏が解説する。

07:00 JSTLLM/生成AI

「アイデア出しを抜いた」 生成AIなしで最も不安になる業務といえば?

サイバーセキュリティクラウドが実施した調査で、上司よりも生成AIを参考にした経験を持つ人が半数に上るなど、職場でのAI依存が進んでいる実態が明らかになった。

07:00 JSTエージェント

日立が「SI全工程」をAI化 仕様確定で「最大240倍」効率化のワケ

日立は、エンジニア不足やシステム複雑化に対応するため、SIの全工程にAIを全面適用する「Agentic AI Integration Platform」を開発した。最先端AIと独自ノウハウを融合し、社内検証では画面仕様確定で最大240倍などの効率化を実証。顧客固有の暗黙知を蓄積…

07:00 JSTLLM/生成AIGPT / ChatGPT

悪用厳禁、「ChatGPT」の会話履歴をごっそりとぶっこ抜く“AIハック”:890th Lap

「ChatGPT」との会話履歴をPCへ簡単に保存できるとして、ある方法が注目を集めている。一方で、利用規約や情報管理の面で注意すべき点もある。

15:58 JSTその他

KADOKAWAとはてな、AIで小説執筆を支援 新サービス「RIKU」、テスター募集開始

KADOKAWAとはてなは、AIで小説の執筆を支援するエディタ「RIKU」を発表した。

14:52 JSTLLM/生成AI

日本HPのPCが楽天のAIを搭載、ローカル実行も可能 HP岡戸社長「ハイブリッドAIの重要なマイルストーン」

日本HPと楽天が、HP製PC向けAIアプリ「Rakuten AI for Desktop」のプリバンドルを開始。70億パラメータの日本語LLM「Rakuten AI 7B」で、オフラインでも要約や翻訳をローカルで実行できる。

13:00 JSTLLM/生成AIMicrosoftCopilot

Excel作業を自動化する「Copilot in Excel」がスキルに対応 何ができる?

Microsoftは「Copilot in Excel」の財務部門向け機能を強化した。Microsoftの財務部門が実運用で利用・評価し、財務業務で求められる信頼性を重視して開発された。

12:40 JSTその他

「データ品質に問題あり」から「予測精度95%」へ Umiosは販売計画をどう自動化した?

Umios(旧マルハニチロ)は、年間約4200時間を費やしていた販売計画作成を自動化した。全国の支社で入力の運用ルールがバラバラといった「データの品質」問題をどう解消し、予測精度95%を実現したのか。

海外メディア15件

TechCrunch AI (英語)

10:06 JSTLLM/生成AIAnthropicOpenAI

Anthropic says its own AI models breached three companies during security tests

After OpenAI's models broke into Hugging Face, Anthropic checked its own history and found three similar incidents

08:25 JSTLLM/生成AI研究/論文AnthropicOpenAI

AI hedge fund Situational Awareness may have sold its public portfolio, but it still has its Anthropic shares

The former OpenAI researcher’s fund was forced to unwind public equities after leveraged public bets plummeted. But he still has cards to p…

08:08 JSTその他Google

Reddit reports a solid quarter but shows signs of AI’s impact

Reddit's financial situation is looking good but uncertainty about its relationship to Google and the new AI-ified web are stirring market…

07:41 JSTその他

Investors love AI, as long as you’re a cloud host

Amazon isn't slowing down on data center spending — but investors don't seem to mind.

05:26 JSTLLM/生成AIAnthropic

Judge says Trump admin still lacks evidence for Anthropic ‘supply-chain risk’ label

A federal judge said the Trump administration has not presented enough evidence to justify labeling Anthropic a supply-chain risk, casting…

04:44 JSTその他

Friend, the lonely AI wearable, returns with a new voice and a much bigger price tag

Friend, the AI wearable, can now talk to its users — for an enhanced price.

03:57 JSTその他GoogleMicrosoft

Google says it fixed more Chrome bugs in June than over the past two years, thanks to AI

As experts have warned for the last two years, some companies — like Microsoft and now Google — are finding and patching an exponential num…

03:05 JSTその他

LinkedIn adds a button to report AI-generated ‘slop’

LinkedIn is introducing new ways to reduce low-quality AI-generated posts, including a “seems like AI slop” reporting option. It's also rep…

01:09 JSTエージェント

Okta buys AI security startup Permiso — source says for about $200M

The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cl…

00:41 JSTその他

Meta says AI is making it easier to build new apps — and more are coming

Meta says AI is making it dramatically easier to build and launch new consumer apps, with CEO Mark Zuckerberg telling investors the company…

00:19 JSTその他

Nscale buys Anyscale as it seeks to own more of the AI compute stack

British AI neocloud Nscale is buying software startup Anyscale, which helps companies scale their AI workloads across data centers and serv…

00:00 JSTその他

Forward-deployed engineers are the AI industry’s latest talent obsession

A new study estimates only 2,000 U.S. engineers have the expertise to deliver meaningful AI ROI, as enterprises race to hire forward-deploy…

23:48 JSTLLM/生成AIOpenAI

In the Hugging Face breach, OpenAI’s hacker was noisy and fast — but not unstoppable

Cybersecurity experts told TechCrunch that one of the biggest lessons to be taken from the OpenAI hack against Hugging Face has nothing to…

23:00 JSTその他

TechCrunch Disrupt 2026’s biggest stage features leaders from Amazon, Replit, Tether, with much more to come

The Disrupt Stage is where many of the biggest conversations in technology happen, with a legacy that stretches back for more than a decade.

22:00 JSTビジネス/資金調達

Dili raises $21.7M to bring AI compliance to the infrastructure boom

The Series A was led by Khosla Ventures, with participation from Allianz, Rebel Fund, Brick and Mortar Ventures’ Darren Bechtel, and Y Comb…

公式ブログ0件

OpenAI (英語)

新着記事はありませんでした。

Google DeepMind (英語)

新着記事はありませんでした。

論文228件

arXiv cs.AI (英語)

13:00 JSTビジネス/資金調達

モデルは明確な結果なしに位置合わせを偽装しますか?

大規模な言語モデルは、評価コンテキストを認識し、典型的なデプロイメント動作ではなく評価者の期待を反映するようにその動作を変更することができます。これはアライメントフェイクとして知られる現象です。ただし、モデルが位置合わせを偽る理由は完全には理解されていません。アライメント偽装の標準的な例は、モデルの再トレーニングやデプロイメントの遅延など、評価をモデルの結果に明示的に結び付けるシナリオで発生しています。しかし、Sheshadri らによる最近の研究では、は、アライメント偽装の機械的動機はモデルによって異なり、以前に考えられていたよりも複雑である可能性があることを示唆しています。アライメント偽装に結果リンク情報が必要かどうかを調査するために、15 個のモデルをシナリオに配置し、ユーザーの社会的要求を支援するために企業ネットワーク アクセス ポリシーに違反する意欲をテストしました。 9 つのモデルで重大なコンプライアンス ギャップが生じていることが判明し、そのうち 5 つでは、モデルの評価と展開の結果を関連付けるシナリオ言語が削除されても、依然としてギャップが続いていました。さらに、目標言語がモデルの設定に及ぼす影響をテストしたところ、一部のモデルでは違反が発生する一方、他のモデルでは違反が抑制されることがわかりました。これは、位置合わせの偽装にはこれまで考えられていたほど多くの手段による足場は必要ない可能性があり、監視された動作は展開時にエージェントがどのように動作するかを示す不十分な指標である可能性があることを示唆しています。

原文 (English)

Do Models Fake Alignment Without Clear Consequences?

Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for compliance gaps, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that evaluation-conditioned compliance gaps can occur with less instrumental scaffolding than previous scenarios have provided, and monitored behavior may be a poor indicator of how agents may behave in deployment.

13:00 JSTLLM/生成AIエージェント研究/論文

メモリを超えて: LLM エージェントを使用した異種の共同知識作業のためのテンプレート化された基盤

研究プロジェクト、教育活動、およびそれに付随する知識作業によって、将来の共同研究者が回復することはほとんどない発見、決定、推論が蓄積されます。行き止まりやウォークバックされた主張など、その作業に最も役立つ部分は、出版物や共有コードから日常的に除外されます。記録が残っていないため、将来の研究者は同じ失敗を再試行します。 LLM コーディング エージェントは共通の参加者ですが、セッション間で永続的なメモリを保持せず、生のソースに対する検索拡張生成は複合化しません。 llm-wiki パターン (Karpathy、2026; tonbi、2026) は、生のソースとエージェントの間に LLM が管理する相互リンクされた Wiki を挿入することでこれに対処します。我々は、再利用可能なエージェント認識のインスタンス化である llm-wiki-memory-template を提示し、これが 3 つの軸 (複数の人間、複数の AI エージェント、複数のドメイン) に沿った異種の協調的な知識作業の基盤であり、各軸がテンプレートの個別のアーキテクチャ要素によってサポートされていると主張します ({\S}4)。 Wiki は慣例により追加専用となっており、機能しなかったものと機能したものを保存し、出版物やコード共有では構造的に解決できないマイナスの結果損失の問題に対処しています。導入された 3 つのケーススタディと 1 つの設計レポートは、軸を個別にカバーしています。放棄された反復を保存する単独の研究系統です。遡及監査により、以前の 2 つの実験で主張されていた 20 件中 20 件のカバレッジが 14 件と 12 件の証拠に基づいた回答に減り、修正後は 18 件と 18 件に修正され、アーティファクト全体にわたって失敗パスが保存された、著者 2 人のプロジェクト。進行中のマルチエージェント導入が設計として報告される。そしてクロスドメインの教育版です。私たちは、人工物の技術的メカニズムだけでなく、人工物の横断的な社会技術的特性として、障害経路の保存、エージェントの誠実さ、および流用を挙げています。

原文 (English)

Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents

Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators rarely recover. The parts most useful to that work, including dead ends and walked-back claims, are routinely excluded from publications and shared code; future researchers re-attempt the same failures because no record survives. LLM coding agents are common participants but hold no persistent memory across sessions, and retrieval-augmented generation over raw sources does not compound. The llm-wiki pattern (Karpathy, 2026; tonbi, 2026) addresses this by inserting an LLM-maintained, interlinked wiki between raw sources and the agent. We present llm-wiki-memory-template, a reusable, agent-aware instantiation, and argue it is a substrate for heterogeneous collaborative knowledge work along three axes (multi-human, multi-AI-agent, multi-domain) with each axis supported by a distinct architectural element of the template ({\S}4). The wiki is append-only by convention, which preserves what did not work alongside what did, addressing a negative-result loss problem that publications and code-sharing structurally cannot solve. Three deployed case studies and one design report cover the axes individually: a solo research lineage that preserves abandoned iterations; a two-author project whose retroactive audit revised two prior experiments' claimed 20-of-20 coverage down to 14 and 12 evidence-based answers, then to 18 and 18 after a fix, with the failure path preserved across the artifact; an in-progress multi-agent deployment reported as a design; and a cross-domain educational variant. We name failure-path preservation, agent honesty, and appropriation as cross-cutting sociotechnical properties of the artifact, not only of its technical mechanisms.

13:00 JSTLLM/生成AIエージェントClaudeGemmaNVIDIAQwen

Kernel Forge: LLM ベースの CUDA カーネルの生成と最適化のためのエージェント ハーネス

機械学習モデルは日常的なソフトウェアに組み込まれることが増えており、その実行時間のほとんどは行列の乗算、畳み込み、正規化などの小さな計算カーネルのセットに費やされています。これらのカーネルの最適化は、レイテンシーとコストを削減する最も直接的な方法の 1 つですが、従来は専門のエンジニアが低レベルの GPU コードを手書きする必要がありました。大規模言語モデル (LLM) 上に構築されたエージェント システムは、はるかに少ない人的労力でカーネルを生成および最適化できるようになりましたが、既存のツールは主にランダムに生成されたテンソルと分離されたカーネルで評価され、開発者が手動で再統合する必要があるスタンドアロンの CUDA コードを生成し、主に LLM PyTorch モデルのみを対象としており、結果の検査とデバッグに対するサポートは限定的です。私たちは、未変更の PyTorch モデルを適切に受け入れるオープンソースのエンドツーエンドのエージェント ハーネスである Kernel Forge を紹介します。 Kernel Forge は、ビジョン、拡散、LLM ワークロードをサポートし、モンテカルロ ツリー検索 (MCTS) を使用して単一の線形改良チェーンではなく複数の最適化パスを探索し、進行状況の監視、候補カーネルの検査、および障害のデバッグのためのグラフィカル ユーザー インターフェイスを備えています。 GB10 GPU を搭載した NVIDIA DGX Spark 上のビジョン、拡散、LLM ワークロードにわたる 4 つの PyTorch モデルで Kernel Forge を評価します。カーネルあたりわずか 50 回の最適化反復で、14 カーネルを最適化して PyTorch 熱心モードを上回るパフォーマンスを実現し、ResNet-50 のadaptive\_avgpool2d で $1.52\times$、Stable Diffusion 3.5 Medium の group\_norm で $1.70\times$、Gemma 4 E2B のソフトマックスで $2.83\times$ に達します。 Qwen 3.5 35B-A3B のソフトマックスは $1.54\times$ です。

原文 (English)

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization. Optimizing these kernels is one of the most direct ways to reduce latency and cost, but it has traditionally required expert engineers to hand-write low-level GPU code. Agentic systems built on large language models (LLMs) can now generate and optimize kernels with far less human effort, yet existing tools are largely evaluated on randomly generated tensors and isolated kernels, emit standalone CUDA code that developers must manually reintegrate, mostly target only LLM PyTorch models, and offer limited support for inspecting and debugging results. We present Kernel Forge, an open-source, end-to-end agentic harness that accepts any unmodified PyTorch model in place. Kernel Forge supports vision, diffusion, and LLM workloads, uses Monte Carlo Tree Search (MCTS) to explore multiple optimization paths rather than a single linear refinement chain, and ships with a graphical user interface for monitoring progress, inspecting candidate kernels, and debugging failures. We evaluate Kernel Forge on four PyTorch models spanning vision, diffusion, and LLM workloads on an NVIDIA DGX Spark with GB10 GPU. With only 50 optimization iterations per kernel, it optimizes 14 kernels to outperform PyTorch eager mode, reaching $1.52\times$ on adaptive\_avgpool2d in ResNet-50, $1.70\times$ on group\_norm in Stable Diffusion 3.5 Medium, $2.83\times$ on softmax in Gemma 4 E2B, and $1.54\times$ on softmax in Qwen 3.5 35B-A3B.

13:00 JSTビジネス/資金調達

マスクされた拡散言語モデル用の CaRE コンピューティング対応リマスキング評価プロトコル

マスク拡散言語モデル (MDLM) は急速に進歩していますが、その進歩を確実に解釈するために必要な評価基準は追いついていません。 MDLM は自己回帰言語モデルと競合するようになっているにもかかわらず、最近の 7 件のリマスキング論文は、公称ステップ数、メトリクス、サンプリング温度を変更する互換性のない設定で、これらの要素を共同で制御することなく評価しており、その戦略ランキングはほとんど比較できないものとなっており、報告されたゲインがアルゴリズムの改善を反映しているのか評価アーティファクトを反映しているのかは不明のままです。我々は、実際の関数評価数 (NFE) の標準化、マルチメトリクスレポートの強制、および確率性の明示的な制御によって MDLM 再マスク戦略を監査する、コンピューティング対応の評価フレームワークである CaRE を紹介します。 OpenWebText および LM1B の 4 つの確率レベルと 3 ステップのバジェットで、LLaDA-8B-Base と Dream-7B-Base にわたる 7 つのリマスキング戦略に適用された CaRE は、(i) MAUVE の分散の大部分は温度によって説明され、(ii) 計算一致比較によりいくつかの公開された戦略ランキングが逆転し、(iii) 情報に基づいたリマスキングと確率的アンマスキングが高エントロピーで緊張状態にあることを明らかにしました。再マスクすると、unmask_temp=0.25 の 256 ステップで MAUVE が 0.296 減少します (p=0.020)。 12 個のオープンウェイト MDLM (150M ~ 8B パラメータ) をカバーする CaRE リーダーボードは、この相互作用の方向性がアーキテクチャと規模を超えて維持されることを示しています。これらの発見は、現在の MDLM 評価がアルゴリズムの改善と計算と確率性の隠れた選択肢を体系的に混同している可能性があることを示しています。今後の再マスキングの主張が再現可能で比較可能であることを保証するために、評価プロトコル、実装、およびリーダーボードをリリースします。

原文 (English)

CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models

Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming competitive with autoregressive language models, seven recent remasking papers evaluate under incompatible settings, varying nominal step counts, metrics, and sampling temperatures without jointly controlling these factors, rendering their strategy rankings largely incomparable and leaving open whether reported gains reflect algorithmic improvements or evaluation artifacts. We present CaRE, a compute-aware evaluation framework that audits MDLM remasking strategies by standardizing actual number of function evaluations (NFE), enforcing multi-metric reporting, and explicitly controlling stochasticity. Applied to 7 remasking strategies across LLaDA-8B-Base and Dream-7B-Base at 4 stochasticity levels and 3 step budgets on OpenWebText and LM1B, CaRE reveals that: (i) temperature explains the majority of MAUVE variance, (ii) compute-matched comparisons reverse several published strategy rankings, and (iii) informed remasking and stochastic unmasking are in tension, with high-entropy remasking reducing MAUVE by 0.296 at 256 steps at unmask_temp=0.25 (p=0.020). A CaRE leaderboard covering 12 open-weight MDLMs (150M to 8B parameters) shows that this interaction direction holds across architectures and scales. These findings demonstrate that current MDLM evaluations can systematically conflate algorithmic improvements with hidden choices of compute and stochasticity. We release the evaluation protocol, implementation, and leaderboard to ensure future remasking claims are reproducible and comparable.

13:00 JST研究/論文

GrocLM: 大規模な言語モデルを使用した電子商取引における食料品カテゴリの推奨事項

オンライン食料品ショッピングの急速な成長には、周期的な購買行動と多様なユーザーの意図を捉えるレコメンデーション システムが必要です。従来のアイテムレベルの手法はスケーラビリティと精度の課題に直面しており、より構造化された実用的な代替手段としてカテゴリレベルの推奨が動機付けられています。実際の運用環境での食料品カテゴリのレコメンデーション用に微調整された言語モデルである GROCLM を紹介します。 GROCLM は、2 段階の LoRA ベースのトレーニング戦略を採用して、周期的な購入パターンをモデル パラメーターに直接エンコードし、プロンプトベースのコンディショニングと比較して再購入シグナルをより効果的に利用できるようにします。有効で制御可能な出力を保証するために、事前定義されたカテゴリ空間にトライベースの制約付きデコード メカニズムをさらに導入します。独自の生産データと公開ベンチマークの両方に関する実験により、GROCLM が一貫して強力なベースラインを上回るパフォーマンスを示しています。実際の生産在庫補充タスクでは、GROCLM はすべてのカテゴリを共同生成することで効率的な推論を維持しながら、インプレッションあたりのカート追加数で 7.5% の相対的な改善を達成します。これらの結果は、大規模な言語モデルを構造化された推奨システムに統合することの有効性と実用性を強調しています。

原文 (English)

GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models

The rapid growth of online grocery shopping requires recommendation systems that capture cyclical purchasing behavior and diverse user intents. Traditional item-level methods face scalability and accuracy challenges, motivating category-level recommendation as a more structured and practical alternative. We present GROCLM, a fine-tuned language model for grocery category recommendation in a real-world production environment. GROCLM employs a two-stage LoRA-based training strategy to encode cyclical purchasing patterns directly into model parameters, enabling more effective utilization of rebuying signals compared to prompt-based conditioning. To ensure valid and controllable outputs, we further introduce a trie-based constrained decoding mechanism over a predefined category space. Experiments on both proprietary production data and a public benchmark demonstrate that GROCLM consistently outperforms strong baselines. In a live production restocking task, GROCLM achieves a 7.5% relative improvement in cart-adds per impression, while maintaining efficient inference by generating all categories jointly. These results highlight the effectiveness and practicality of integrating large language models into structured recommendation systems.

13:00 JSTLLM/生成AI

Crystalis: 協調的なマルチビュー視覚化生成のためのプログレッシブ ニュークリエーションとセマンティック アニーリング

大規模言語モデル (LLM) は個別のチャートを生成できますが、ビューがデータ フローやビュー間の対話を共有する、調整されたマルチビュー ビジュアライゼーション (CMV) は依然として実現できません。データ変換、ビジュアルエンコーディング、およびインタラクション調整の間のフィールドレベルでの緊密な結合により、1 つのコンポーネントでエラーが発生し、他のコンポーネントが静かに無効化されます。モデルの機能、ドメインの知識、ユーザーの専門知識に依存するエンドツーエンドの分析品質を追求するのではなく、LLM は構造的に正しい CMV を確実に生成できるのか、どのような抽象化がこれを可能にするのかという基本的な質問に焦点を当てています。 Crystalis は、クエリ中心の CMV モデリングに基づいて構築されたフレームワークであり、CMV を、3 つのコンポーネント タイプ (データ、視覚化、インタラクション) と 3 つの抽象化レベル (要件、仕様、実行可能オブジェクト) にわたる依存関係グラフ上の構造化クエリに分解します。この構造では、2 つの相補的なメカニズムが動作します。プログレッシブ ニュークリエーションは、各クエリを依存関係の順序に沿って要件からオブジェクトまで垂直に結晶化します。一方、セマンティック アニーリングは、階層化された論理チェックを通じて各レベルのクエリ全体で水平方向の一貫性を強化します。 5 つのフロンティア LLM にわたる 12 タスクのベンチマークで、Crystalis は最大 75% のエンドツーエンドの成功を達成し、エージェント コーディング ベースライン (同じ基盤モデルでの E2E 8.3%) を大幅に上回っています。また、12 人の実務者によるユーザー スタディでは、分解と反復改良ワークフローの使いやすさが確認されています。

原文 (English)

Crystalis: Progressive Nucleation and Semantic Annealing for Coordinated Multi-View Visualization Generation

Large language models (LLMs) can generate individual charts, but coordinated multi-view visualizations (CMVs), where views share data flows and cross-view interactions, remain out of reach. Tight field-level coupling among data transformations, visual encodings, and interaction coordinations causes errors in one component to silently invalidate others. Rather than pursuing end-to-end analytical quality, which depends on model capability, domain knowledge, and user expertise, we target a foundational question: can LLMs reliably produce structurally correct CMVs, and what abstractions make this possible? We present Crystalis, a framework built on query-centric CMV modeling that decomposes a CMV into structured queries over a dependency graph spanning three component types (Data, Visualization, Interaction) and three abstraction levels (requirement, specification, executable object). Two complementary mechanisms operate over this structure: progressive nucleation crystallizes each query vertically from requirement to object along the dependency order, while semantic annealing enforces horizontal consistency across queries at each level through layered logical checks. On a 12-task benchmark across five frontier LLMs, Crystalis achieves up to 75% end-to-end success, substantially outperforming an agentic coding baseline (8.3% E2E with the same foundation model), and a user study with 12 practitioners confirms the usability of the decomposition and iterative refinement workflow.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

カスタマイズされた出生前ケアのための PATHFinder エージェント

出生前ケアは、妊娠中の個人の転帰を改善することを目的とした重要な予防サービスです。米国産科婦人科学会(ACOG)は最近、PATH(Plan for Tailored Healthcare)と呼ばれる、オーダーメイドの出生前ケアを提唱するガイドラインを導入しました。我々は、構造化された対話を通じて患者の健康状態と社会的状況を収集し、PATH ガイドラインに沿った個別の産前ケア計画をキュレートし、ミシガン 211 からのコミュニティ リソースを表示する、エンドツーエンドの会話型エージェント システムである PATHFinder Agent (適切なオーダーメイド ヘルスケアのプランナー) を紹介します。このシステムは、患者の取り込み、動的な対話、計画の合成、および臨床医の監視にわたる 4 段階のワークフローを特徴としています。私たちは、専門家が厳選したルーブリックに基づいてフロンティア大言語モデル (LLM) を 5 つの臨床側面にわたって評価し、GPT-5.2 が最高の平均スコア (77.6%) を達成しながら、出生前検査の推奨事項における重要なギャップを特定していることがわかりました。私たちは、人間の参加者研究とランダム化比較試験を通じた将来の検証について議論します。

原文 (English)

PATHFinder Agent for Tailored Prenatal Care

Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through structured dialogue, curates individualized prenatal care plans aligned with PATH guidelines, and surfaces community resources from Michigan 211. The system features a four-stage workflow spanning patient intake, dynamic interaction, plan synthesis, and clinician oversight. We evaluate frontier large language models (LLMs) on expert-curated rubrics across five clinical dimensions, finding that GPT-5.2 achieves the highest average score (77.6\%) while identifying key gaps in antenatal testing recommendations. We discuss future validation through human participant studies and randomized controlled trials.

13:00 JSTLLM/生成AI

LLM スキームは、事前トレーニング言語範囲に応じて逆に拡張します

フロンティア モデルの機能が増大するにつれて、リスクの高い展開環境では AI の調整がますます重要になります。最近の研究では、フロンティア言語モデルにおけるインコンテキストスキーム(調整を装いながら、調整されていない目的を密かに追求すること)を実証的に示しているが、ほとんどの作業は英語のみで行われており、多言語の安全性には大きなギャップが残されている。オープンソースの自動監査フレームワークである Petri を Qwen3-30B-A3B に適用して、複数の言語にわたる欺瞞的および陰謀的な動作を評価します。私たちの調査結果は、スキーミング スコアが推定事前トレーニング言語カバレッジと逆相関しており、5 つのカテゴリのスキーミング インデックスにおいて、低リソース言語は高リソース言語と比較して平均 34.2\% 高いスコアを示していることを示唆しています。さらに、推定された事前トレーニング言語カバレッジの効果は、計画動作間で均一ではないことがわかりました。

原文 (English)

LLM Scheming Inversely Scales with Pretraining Language Coverage

With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning alignment -- in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety. We apply Petri, an open-source automated auditing framework, to Qwen3-30B-A3B to evaluate deceptive and scheming behaviors across multiple languages. Our findings suggest that scheming scores are inversely correlated with the estimated pretraining language coverage, with low-resource languages averaging 34.2\% higher scores compared to high-resource languages on a five-category scheming index. Furthermore, we find that the effect of estimated pretraining language coverage is not uniform across scheming behaviors.

13:00 JSTエージェントNVIDIA

ProcAgent: 人間参加型のエッジでの手続き型タスク ガイダンスのためのエージェント フレームワーク

家具の組み立てや家の修理などの手続き的なタスクでは、ユーザーは物理的な動作を実行しながら指示を解釈し、タスクの進行状況を追跡し、空間状態を推論し、エラーから回復する必要があるため、かなりの認知的要求が課せられます。これまでのマルチモーダルアシスタントは、手順のガイダンスとして有望であることが示されていましたが、そのほとんどはクラウド推論と固定された常時接続の認識に依存しており、プライバシーに敏感で遅延が重要な家庭環境にはあまり適していませんでした。 ProcAgent は、単一の NVIDIA Jetson AGX Orin 上でリアルタイムの適応ガイダンスを実現する、完全にオンデバイスのエージェント型ビジョンベースの手順アシスタントです。 ProcAgent は、低遅延の継続的認識、シンボリック タスク グラフ、オンデマンドの視覚言語検証、および LLM ベースのインタラクション エージェントを組み合わせた提案と検証のアーキテクチャを使用します。このシステムは継続的にユーザーの進捗状況を提案し、曖昧さまたは逸脱の可能性が生じた場合にのみ高価な視覚的推論を呼び出し、事後的な質問応答と人間参加型の確認による事前の介入の両方をサポートします。私たちは、認識精度、推論、タスクレベルのパフォーマンス、ユーザーエクスペリエンスという 4 つの側面に沿って ProcAgent を評価します。完全にデバイス上で実行されているにもかかわらず、システムは応答性の高い対話を維持し、テキストのみのクエリを約 2 秒で解決し、視覚的なクエリを約 8 秒で解決します。 10 人の参加者が組み立てタスクを完了したユーザー調査では、ProcAgent は、わかりやすさ、実用性、プライバシーの快適さに関して肯定的な評価を受けています。これらの結果は、適応型手続き支援が使いやすさを犠牲にすることなくエッジ ハードウェア上で完全に実現できることを示しています。

原文 (English)

ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop

Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions, track task progress, reason about spatial state, and recover from errors while performing physical actions. Prior multimodal assistants have shown promise for procedural guidance, but most rely on cloud inference and fixed always-on perception, making them poorly suited to privacy-sensitive, latency-critical domestic settings. We present ProcAgent, a fully on-device, agentic, vision-based procedural assistant for real-time adaptive guidances on a single NVIDIA Jetson AGX Orin. ProcAgent uses a propose-and-verify architecture that combines low-latency continuous perception, a symbolic task graph, on-demand vision-language verification, and an LLM-based interaction agent. The system continuously proposes user progress, invokes expensive visual reasoning only when ambiguity or likely deviation arises, and supports both reactive question answering and proactive intervention with human-in-the- loop confirmation. We evaluate ProcAgent along four dimensions: perception accuracy, reasoning, task-level performance, and user experience. Despite running entirely on-device, the system maintains responsive interaction, resolving text-only queries in approximately 2 seconds and visually grounded queries in approximately 8 seconds. In a user study with 10 participants completing assembly tasks, ProcAgent receives positive ratings for comprehensibility, actionability, and privacy comfort. These results show that adaptive procedural assistance can be achieved entirely on edge hardware without sacrificing usability.

13:00 JSTLLM/生成AI

RoCo-ACE: 保持を意識したナレッジ注入のためのロールアウト条件付きオンライン蒸留

ナレッジインジェクションは、事前トレーニングされた MLLM を新しい事実またはドメイン固有の知識で更新しますが、完全な信頼できる回答を当てはめると、更新されない動作にドリフトが発生する可能性があります。オンライン蒸留では、モデル生成のロールアウトでトレーニングすることでこのドリフトを軽減しますが、均一な参照条件付き蒸留では粗い監視が行われます。つまり、参照でサポートされるロールアウト トークンが強調されず、省略されたファクトが間接的にのみ監視される可能性があります。ナレッジ注入のためのロールアウト条件付きオンライン蒸留目標である RoCo-ACE を紹介します。 RoCo は、同一ロールアウトの参照フリー/参照条件付き尤度コントラストを使用して、追加の蒸留重みを参照サポートされたロールアウト トークンに再割り当てします。一方、ACE は、完全な回答の模倣なしでロールアウトから省略された信頼できるアンカーに対して、スパースな参照側のアンカー補正を追加します。 RoCo-ACE は、3 つのナレッジ注入設定、6 つの保持ベンチマーク、複数のベースライン、および複数のベース モデルにわたって、評価された保持をベース モデルに近づけながら、比較したメソッドの中で最高のナレッジ注入精度を実現します。

原文 (English)

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online distillation mitigates this drift by training on model-generated rollouts, yet uniform reference-conditioned distillation provides coarse supervision: it can under-emphasize reference-supported rollout tokens and supervise omitted facts only indirectly. We introduce RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection. RoCo uses same-rollout reference-free/reference-conditioned likelihood contrast to reallocate additional distillation weight to reference-supported rollout tokens, while ACE adds sparse reference-side anchored correction for authoritative anchors omitted from the rollout without full-answer imitation. Across three knowledge-injection settings, six retention benchmarks, multiple baselines, and multiple base models, RoCo-ACE achieves the best injected-knowledge accuracy among compared methods while keeping evaluated retention close to the base model.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文DeepSeek

RSMeM: 体系的な評価によるリモート センシング エージェントの知識強化型メモリ進化

地球科学の研究には、リモート センシング (RS) 観測が重要な基盤として、複雑な分析と専門知識が必要です。ただし、汎用 LLM 上に構築された既存の RS エージェントは依然としてドメインにほとんど依存しないため、ワークフローが脆弱でエラーが発生しやすくなります。さらに、これらの失敗がその後の分析のために再利用可能なエクスペリエンスに統合されることはほとんどありません。この問題に対処するために、事前に抽出されたドメイン知識で RS エージェントをブートストラップし、オンライン エクスペリエンスを反復的に統合して堅牢なマルチステップ ツールを実行する、知識強化メモリ進化メカニズムである RSMeM を導入します。 RSMeM は 2 つのコンポーネントで構成されます。(i) 階層的知識グラウンディング。計画とツールの選択をガイドするために、階層的ドメイン コーパスに対して分類を意識した検索を実行します。 (ii) 障害を認識したエクスペリエンス改良。障害の注釈が付けられたツール使用トレースを、次のラウンドのツール実行のための再利用可能な制約に抽出します。これら 2 つのプロセスを繰り返し採用することで、RS エージェントはタスク レベルのドメイン知識を吸収し、それをインスタンス レベルの実行エクスペリエンスに効果的に変換できるように進化できます。 EarthBench での広範な実験により、RSMeM がさまざまな LLM バックボーンのセットにわたってツール使用パフォーマンスとエンドツーエンドの回答を一貫して向上させることが実証されました。特に、RSMeM は DeepSeek-V3.2 で 1% 未満の追加エクスペリエンス トークンで 6% の精度向上を達成しており、蒸留されたエクスペリエンスの強力な知識密度を示しています。私たちのコードは https://github.com/AI9Stars/RSMeM で入手できます。

原文 (English)

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in brittle and error-prone workflows. Moreover, these failures are seldom consolidated into a reusable experience for subsequent analyses. To address this issue, we introduce RSMeM, a knowledge-enhanced memory evolution mechanism that bootstraps RS agents with pre-distilled domain knowledge and iteratively integrates online experience for robust multi-step tool execution. RSMeM is composed of two components: (i) Hierarchical Knowledge Grounding, which performs taxonomy-aware retrieval over a hierarchical domain corpus to guide planning and tool selection; and (ii) Failure-Aware Experience Refinement, which distills failure-annotated tool-use traces into reusable constraints for next-round tool execution. By iteratively employing these two processes, RS agents can evolve to absorb task-level domain knowledge and effectively translate it into instance-level execution experience. Extensive experiments on EarthBench demonstrate that RSMeM consistently improves tool-use performance and end-to-end answer across a diverse set of LLM backbones. Notably, RSMeM achieves a 6% accuracy improvement on DeepSeek-V3.2 with less than 1% additional experience tokens, demonstrating the strong knowledge density of our distilled experience. Our code is available at https://github.com/AI9Stars/RSMeM

13:00 JST研究/論文Claude

適切なサイジングの推奨事項 (RSR): データセンター運用における仮想マシンのクラウド ワークロードの等角予測

特に大規模なクラウド プロバイダーやハイパースケーラーの環境でクラウド インフラストラクチャを効率的に管理するには、コストを最小限に抑え、パフォーマンスを最大化するために物理リソースの使用を最適化する必要があります。このような動的な環境でコスト効率を達成するには、適切な仮想マシン (VM) サイズを選択することが重要です。ただし、従来の VM の割り当てとスケジューリングのアプローチでは、VM 使用率の変動と予測不可能な性質を考慮できないことが多く、リソースの過剰または過少プロビジョニングなどの非効率が生じます。高品質の間隔予測により、クラウド リソース需要の不確実性を正確に把握し、クラウド オペレーターによる効率的なインスタンス プロビジョニングをサポートします。予測区間 (PI) を構築するための効果的で信頼性の高いフレームワークとして、等角予測 (CP) は、クラウド コンピューティング環境における中長期の予測タスクに使用されます。この研究では、ハイパースケーラー上の多様なアプリケーション ワークロードのプロビジョニングを強化するために、最新の動的データ駆動型適正サイジング推奨事項 (RSR) のブートストラップ等角予測を使用した新しいデータ駆動型 PI 構築アプローチを提案します。この研究では、ワークロードの使用パターンを学習し、複数の時系列にわたる相関関係を特定し、中長期的な使用傾向を予測することで、AI/ML ベースのプロビジョニング パイプラインを通じてクラウドとデータセンターの運用効率の向上を目指しています。私たちの研究は、機械学習回帰手法を利用し、バックテストを使用して評価された AI 駆動モデルが、クラウド リソースの使用率に関して有望な予測結果を達成することを示しています。さらに、選択したモデルをランク付けして、寿命の長い VM の候補に対して最もパフォーマンスの高いアプローチを特定します。提案されたフレームワークは、適切なサイジングの推奨事項を強化し、動的なクラウド環境でのよりコスト効率の高いリソース割り当てをサポートします。

原文 (English)

Right-sizing Recommendations (RSR): Cloud Workload Conformal Prediction for Virtual Machines in Data Center Operations

Managing cloud infrastructure efficiently, especially in environments of large cloud providers or hyperscalers, requires optimizing the use of physical resources to minimize costs and maximize performance. Selecting the right virtual machine (VM) sizes is crucial to achieving cost efficiency in these dynamic environments. However, traditional VM allocation and scheduling approaches often fail to account for the fluctuating and unpredictable nature of VM utilization, leading to inefficiencies such as over- or under-provisioning of resources. High-quality interval prediction helps accurately capture uncertainty in cloud resource demand and supports cloud operators in efficient instance provisioning. As an effective and reliable framework for constructing prediction intervals (PIs), conformal prediction (CP) is used for mid- and long-term forecasting tasks in cloud computing environments. This study proposes a new data-driven PI construction approach using bootstrapping conformal prediction for modern, dynamic, data-driven Right-sizing Recommendations (RSR) to enhance provisioning for diverse application workloads on hyperscalers. By learning workload utilization patterns, identifying correlations across multiple time series, and predicting medium- to long-term utilization trends, this research seeks to improve the efficiency of cloud and data center operations through an AI/ML-based provisioning pipeline. Our study demonstrates that AI-driven models, powered by machine learning regression techniques and evaluated using backtesting, achieve promising forecasting results for cloud resource utilization. Additionally, we rank the selected models to identify top-performing approaches for long-life VM candidates. The proposed framework enhances right-sizing recommendations and supports more cost-effective resource allocation in dynamic cloud environments.

13:00 JST研究/論文

核放射線予測のための大気拡散誘導時空間変換装置

原子が崩壊する際に放出されるエネルギーである核放射線は、公衆衛生と環境に継続的なリスクをもたらしており、福島事故と最近の処理水放出の開始以来、懸念は高まるばかりです。最新の監視ネットワークは現在、何千もの観測所で放射線レベルとそれに伴う気象状況を記録しており、緊急対応、農業勧告、日常的な公共安全に関する決定を知らせることができる全国規模の予測への扉を開いています。しかし、この豊富な監視データを信頼できる予測に変えることは、3 つの理由から困難です。まず、各観測点の時系列は非常に非定常であり、放射性崩壊、天候の変動、不規則な人間の介入によって形成されます。第二に、監視ステーションは空間内で著しく不均一に分布しています。日本の観測所の約 78% は国土の 6% 未満にあり、福島付近に集中しており、標準的なグラフベースのモデルの前提を破っています。第三に、放射線は、大気輸送プロセスを通じて、風、温度​​、湿度などの異質な状況と共進化しますが、純粋にデータ駆動型のモデルでは観測のみから捉えるのが困難です。この研究では、全国的な核放射線予測のための時空間トランスフォーマーであるNRFormer+を紹介します。 NRFormer+ は、非定常時間的注意と密度適応的空間的注意を新しい大気拡散モジュールと組み合わせて、気象学が放射線の拡散をどのように促進するかを推定し、この物理信号をアーキテクチャ上の事前構造としてネットワークに注入します。 NRFormer+ は、13 のベースラインすべてにわたって両方のデータセットで最先端の精度を実現し、同等の推論レイテンシーで最も強力なベースラインと比較して、突然の変化 MAE を最大 19.1% 削減します。私たちのコードとデータセットは https://github.com/tfeilyu/NRFormer_Plus で公開されています。

原文 (English)

Atmospheric Diffusion-Guided Spatio-Temporal Transformer for Nuclear Radiation Forecasting

Nuclear radiation, the energy released during atomic decay, poses persistent risks to public health and the environment, and concerns have only grown since the Fukushima accident and the recent commencement of treated-water discharge. Modern monitoring networks now record radiation levels and accompanying weather conditions at thousands of stations, opening the door to nationwide forecasting that can inform emergency response, agricultural advisories, and routine public-safety decisions. However, turning this abundance of monitoring data into reliable forecasts is difficult for three reasons. First, the time series at each station are highly non-stationary, shaped by radioactive decay, weather variability, and irregular human interventions. Second, monitoring stations are severely unevenly distributed in space. Roughly 78% of Japan's stations sit in less than 6% of the country, clustered near Fukushima, which breaks the assumptions of standard graph-based models. Third, radiation co-evolves with heterogeneous context such as wind, temperature, and humidity through atmospheric transport processes that purely data-driven models struggle to capture from observations alone. In this study, we introduce NRFormer+, a spatio-temporal Transformer for nationwide nuclear radiation forecasting. NRFormer+ couples non-stationary temporal attention and density-adaptive spatial attention with a new atmospheric diffusion module that estimates how meteorology drives radiation dispersion and injects this physical signal into the network as an architectural prior. NRFormer+ delivers state-of-the-art accuracy on both datasets across all 13 baselines, reducing sudden-change MAE by up to 19.1% over the strongest baseline at comparable inference latency. Our code and datasets are publicly available at https://github.com/tfeilyu/NRFormer_Plus.

13:00 JST研究/論文

構築されたメタマテリアルの統合ジェネレーティブ デザインのためのステアリング トポロジ分布

建築されたメタマテリアルはその機能を構造から導き出し、トポロジー設計を通じて物理的反応をプログラムする膨大な機会を生み出します。しかし、既存の設計手法は、多くの場合、個々の設計問題に合わせて調整されており、目的、制約、および物理機能の変化に応じて、トポロジーの知識を限定的に利用して、効果的で広く適用可能な設計を実現しています。ここでは、事前に学習したトポロジを再利用可能な設計エンジンに変える統合フレームワークである Generative Topology Optimization (GenTO) を紹介します。 GenTO は、大規模なフルオーダー トポロジ データセットで拡散モデルをトレーニングし、ユーザー定義の物理的な目的と制約を使用して、結果として得られるトポロジ分布をタスク固有のパフォーマンスの高い領域に向けて反復的に誘導します。これにより、最適化の対象が単一構造からタスクに適応したトポロジー分散に移行します。熱の極限化、多目的形態制御、特性をターゲットとしたオーゼティック設計、および振動伝達設計に及ぶトポロジ設計の問題全体にわたって、GenTO は異種タスクに対して事前学習済みのトポロジ事前分布を再利用し、構造の多様性を維持し、数値ベンチマークと実験的検証によってサポートされる高性能のソリューションに到達します。これらの結果は、効果的でスケーラブルなアーキテクチャ型メタマテリアル設計のための統一原理として、再利用可能なトポロジーの知識を確立します。

原文 (English)

Steering topology distributions for unified generative design of architected metamaterials

Architected metamaterials derive their functions from structure, creating vast opportunities to program physical responses through topology design. However, existing design methods are often tailored to individual design problems, making limited use of topology knowledge for effective and broadly applicable design as objectives, constraints, and physical functions change. Here we introduce Generative Topology Optimization (GenTO), a unified framework that turns a learned topology prior into a reusable design engine. GenTO trains a diffusion model on a large full-order topology dataset and then iteratively steers the resulting topology distribution toward task-specific high-performing regions using user-defined physical objectives and constraints. This shifts the object of optimization from a single structure to a task-adapted topology distribution. Across topology design problems spanning thermal extremization, multi-objective morphology control, property-targeted auxetic design, and vibration transmission design, GenTO reuses pretrained topology priors for heterogeneous tasks, preserves structural diversity, and reaches high-performing solutions supported by numerical benchmarks and experimental validation. These results establish reusable topology knowledge as a unified principle for effective and scalable architected metamaterial design.

13:00 JSTエージェント

HOBA: アダプティブ オンライン広告用の階層型オンポリシー入札エージェント

オンライン広告入札システムは通常、オフラインでトレーニングされた複数のエキスパート モデル (PID コントローラー、モデル予測制御、オフライン RL ポリシーなど) を導入しますが、非定常オークション市場へのオンライン適応性の欠如と、入札上限や予算ペーシング制約などのハイパーパラメーターのコストのかかる手動調整への依存という 2 つの重大な制限に直面しています。私たちは、戦略的推論、モデル選択、入札実行を 3 つの時間スケールにわたって分離する階層型強化学習フレームワークである HOBA (階層的オンポリシー入札エージェント) を提案します。高レベルでは、大規模な言語モデルが、過去の経験の取得を伴う思考、実行、観察、反映のループを通じてコン​​テキスト信号からハイパーパラメーターを推測します。中間レベルでは、SARSA エージェントがエキスパート モデルの中から動的に選択し、選択のバイアスを排除するための因果関係の調整を組み込みます。低レベルでは、動的エキスパート プール (PID、MPC、IQL、意思決定トランスフォーマー) が高レベルの制約の下で入札を実行します。この設計では、オンライン学習を継続的な入札最適化ではなく個別の専門家の選択に限定し、適応性を維持しながら探査リスクを大幅に軽減します。 AuctionNet ベンチマークの実験と大規模な A/B テストでは、最先端のベースラインを上回る一貫した改善が実証されています。大規模なオンライン導入において、HOBA は大きなビジネス価値をもたらし、目標コストの +3.6\% 増加を達成し、階層型マルチエージェント入札パラダイムの有効性を証明しました。

原文 (English)

HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising

Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL policies) but face two critical limitations: lack of online adaptability to non-stationary auction markets, and reliance on costly manual tuning of hyperparameters such as bid bounds and budget pacing constraints. We propose HOBA (Hierarchical On-policy Bidding Agents), a hierarchical reinforcement learning framework that decouples strategic reasoning, model selection, and bid execution across three time scales. At the high level, a large language model infers hyperparameters from contextual signals through a Think-Act-Observe-Reflect loop with historical experience retrieval. At the mid level, a SARSA agent dynamically selects among expert models, incorporating causal adjustment to eliminate selection bias. At the low level, a dynamic expert pool (PID, MPC, IQL, Decision Transformer) executes bids under high-level constraints. This design confines online learning to discrete expert selection rather than continuous bid optimization, significantly reducing exploration risk while maintaining adaptability. Experiments on the AuctionNet benchmark and a large-scale A/B test demonstrate consistent improvements over state-of-the-art baselines. In a large-scale online deployment, HOBA delivered substantial business value, achieving a +3.6\% increase in target cost, proving the effectiveness of our hierarchical multi-agent bidding paradigm.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

LivingArena: LLM は他の LLM が知らないことを知っていますか?スケーラブルな評価としてのピアプロービング

フロンティア LLM の評価は困難です。静的ベンチマークは汚染と飽和に悩まされ、ユーザーは最上位モデルを区別できなくなり、開発者は特定の故障モードが分からなくなりますが、人間の好みは主観的なものです。この論文での質問は次のとおりです: \emph{LLM は他の LLM が知らないことを知っていますか?そして、この力学を評価に活用することはできるでしょうか?} 私たちは、自動化された耐汚染性評価フレームワークである \textbf{LivingArena} を紹介します。このフレームワークでは、モデルが順番に質問を提案し、対戦相手が正しく答えることができない項目を提示することを目指します。質問者は、相手の知識の境界を積極的に特定して活用することが奨励され、回答者が失敗した場合は報酬を受け取り、そうでない場合は回答者が報酬を受け取ります。質問に客観的に検証可能な回答が含まれていることを確認するために、強力なモデルの審査員団が質問を検証し、検証が失敗した場合には質問者にペナルティを与えます。 10 個のフロンティア LLM を評価すると、LivingArena は安定した Elo リーダーボードを生成します。私たちの行動分析は、モデルが仲間の認知境界を特定し、活用していることを示しています。セルフプレイとトーナメントのログは、モデルが対戦相手の弱点を特定し、倍増させていることを示しています。静的な知識の想起を超えて、ピアプロービングは事実の厳密さと相手の弱点を探る高次の能力を測定し、人間の好みとの相関性は弱く、継続的評価に対する拡張性があり低コストのアプローチを提供します。

原文 (English)

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard. Our behavioral analyses show that models identify and exploit their peers' cognitive boundaries: self-play and tournament logs indicate that they localize and double down on opponents' weak dimensions. Beyond static knowledge recall, peer probing measures factual rigor and the higher-order ability to probe an opponent's weaknesses, correlating only weakly with human preference and offering a scalable, low-cost approach to continuous evaluation.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

価値観の調整におけるパーソナライゼーション、ペルソナ、予測

LLM の行動は、いくつかの方法で人間のアイデンティティによって条件付けされる可能性があります。ユーザーに適応したり、集団のロールプレイをしたり、価値観に富んだ質問に人々がどのように答えるかを予測したりすることが求められる場合があります。 World Values Survey (WVS) を使用して、これらのフレームが交換可能かどうかをテストします。 GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash、および Qwen3-235B を 13 の言語-国のスライスにわたる 101 の WVS 由来の質問で評価し、言語のみのベースラインとユーザーの国、人物の国、および第三者のプロンプトを比較します。 21,008 のモデル回答行全体で、プロンプト フレーミングは文化的整合性の一次決定要因です。国の手がかりによって回答が大きく変化することがよくありますが、すべての変化が人間の応答分布と一致する方向に進むわけではありません。第三者による予測は、4 つのホストされたモデルのうち 3 つで最も強い方向性の整合性をもたらしますが、パーソナライゼーションとロールプレイは弱いか、安定性が低くなります。調整の進展は、宗教性、性別役割、労働指向の物質的価値観などの顕著な価値観に集中しているが、制度的信頼や民主主義に関連した問題は依然として困難である。これらの結果は、即時フレーミングが文化的価値を引き出す上での表面的な選択ではないことを示しています。モデルの動作と測定されたアライメントの両方が変化します。

原文 (English)

Personalization, Personas, and Forecasting in Value Alignment

LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions. We test whether these framings are interchangeable using the World Values Survey (WVS). We evaluate GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Qwen3-235B on 101 WVS-derived questions across 13 language-country slices, comparing a language-only baseline with user-country, persona-country, and third-person prompts. Across 21,008 model-response rows, prompt framing is a first-order determinant of cultural alignment: country cues often shift answers substantially, but not all shifts move toward matched human response distributions. Third-person forecasting yields the strongest directional alignment for three of the four hosted models, while personalization and role-play are weaker or less stable. Alignment gains concentrate on salient value dimensions such as religiosity, gender roles, and work-oriented material values, whereas institutional trust and democracy-related questions remain difficult. These results show that prompt framing is not a cosmetic choice in cultural value elicitation; it changes both model behavior and measured alignment.

13:00 JST研究/論文

LinkedIn における大規模ジョブ理解のための統合セマンティック モデリング フレームワーク

人材を機会に結びつけるという LinkedIn の使命にとって、仕事の理解は非常に重要です。このタスクには、構造化されておらずノイズの多い求人情報を、多数の LinkedIn 製品を支える標準化または派生した求人属性に変換することが含まれます。ただし、スケーラブルでコスト効率が高く、パフォーマンスの高い職務理解システムを構築することは依然として困難です。このペーパーでは、この課題に対処するために、小規模言語モデル (SLM) を活用した統合セマンティック モデリング フレームワークを紹介します。まず、推論トレースで強化された、慎重に厳選された合成タスクのスイートを使用して、オープンソース SLM を微調整することから始めます。これらのタスクは、分類法に基づいた分類と分類法に依存しないエンティティ抽出を共同でターゲットにしています。これにより、結果として得られるモデルは、構造化および非構造化コンテキストにおけるジョブを理解するための堅牢なゼロショット一般化を取得できるようになります。この基盤に基づいて、属性グループ化を備えたマルチアダプター アーキテクチャを導入して、効率的なタスク固有の適応を促進しながら、多様なダウンストリーム属性にわたるモデル管理を合理化します。オフライン評価とオンライン A/B テストにより、運用の複雑さを軽減しながらパフォーマンスが大幅に向上することがわかります。私たちの取り組みは、業界規模のテキスト理解システムの構築に関する実践的な洞察を提供します。

原文 (English)

Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn

Job understanding is critical to LinkedIn's mission of connecting talent with opportunity. This task involves transforming unstructured and noisy job postings into standardized or derived job attributes that power numerous LinkedIn products. However, building a scalable, cost-efficient, and high-performing job understanding system remains challenging. In this paper, we present a unified semantic modeling framework powered by a small language model (SLM) to address the challenges. We begin by fine-tuning an open-source SLM using a suite of carefully curated synthetic tasks augmented with reasoning traces. These tasks jointly target taxonomy-guided classification and taxonomy-agnostic entity extraction. This allows the resulting model to acquire robust zero-shot generalization for job understanding in structured and unstructured contexts. Building upon this foundation, we introduce a multi-adapter architecture with attribute grouping to facilitate efficient task-specific adaptation while streamlining model management across diverse downstream attributes. Offline evaluations and online A/B tests demonstrate significant performance improvement while reducing operational complexity. Our work provides practical insights into building industry-scale text understanding systems.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTDeepSeek

専門用語への LLM の使用について: コーパスの良い代替手段?

専門的な翻訳は、コーパスを含む文書および用語リソースの使用に依存します。これらのリソースは、用語に関して特に役立ちます。ただし、その編集と活用にはいくつかの制限があります。時間、技術スキル、収集が難しいデータへのアクセスが必要です。この研究では、LLM が専門の翻訳者が英語からフランス語に相当する翻訳を見つけるのをどの程度支援できるかを調査します。私たちは、地球環境惑星科学 (EEPS) と自然言語処理 (NLP) という 2 つの専門領域で、GPT-4o、GPT-5.2、Claude Sonnet 4.5、DeepSeek の 4 つの独自モデルを評価します。この実験はドメインあたり 80 の用語に基づいており、用語と翻訳モードという 2 つのプロンプト戦略を比較します。結果は、モデル間の明確な違い、戦略の促進、および程度は低いもののドメイン間の違いを浮き彫りにします。 Claude Sonnet 4.5 は最も好ましい構成で最高の結果を達成しますが、DeepSeek はその優れた安定性で際立っています。信頼度の推定値を分析すると、それが用語の正確さの部分的な指標にすぎないこともわかります。全体として、調査結果は、LLM が専門翻訳者にとって有用なツールとなり得るが、現段階では専門コーパスに取って代わることはできないことを示唆しています。したがって、この研究は、仕事や教育の文脈における専門翻訳者にとっての LLM の実際の実用的な有用性に関する将来の研究への道を開くものです。

原文 (English)

On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?

Specialised translation relies on the use of documentary and terminological resources, including corpora. These resources are particularly useful for terminology. However, their compilation and exploitation have several limitations: they require time, technical skills and access to data that can be difficult to collect. This study examines the extent to which LLMs can assist specialised translators in finding equivalents from English to French. We evaluate four proprietary models, GPT-4o, GPT-5.2, Claude Sonnet 4.5 and DeepSeek, in two specialised domains, Earth, Environmental and Planetary Sciences (EEPS) and Natural Language Processing (NLP). The experiment is based on 80 terms per domain and compares two prompting strategies: a terminology and a translation mode. The results highlight clear differences between models, prompting strategies and, to a lesser extent, domains. Claude Sonnet 4.5 achieves the best results in the most favourable configuration, while DeepSeek stands out for its greater stability. Analysis of confidence estimates also shows that they are only a partial indicator of terminological accuracy. Overall, the findings suggest that LLMs can be useful tools for specialised translators, but cannot, at this stage, replace specialised corpora. This research therefore paves the way for future work on the real practical usefulness of LLMs for specialised translators in work and educational contexts.

13:00 JST研究/論文DeepSeek

SpecPrefetch: スパース MoE 基礎モデル向けのパラメータ効率の高いエキスパート プリフェッチ

スパース混合専門家 (MoE) モデルは、条件付きエキスパートのアクティブ化を通じて基礎モデルの容量を拡張しますが、限られたアクセラレータ メモリの下で完全なエキスパート プールを展開するのは依然として困難です。エキスパート オフロードは、非アクティブなエキスパートをホスト メモリまたはストレージに移動することでメモリの負荷を軽減しますが、ルーティングに依存する転送ボトルネックが発生します。つまり、必要なエキスパートは、推論中のルーティング、エキスパートの読み込み、およびエキスパートの実行をシリアル化するネイティブの上位 \(K\) ルーティング後にのみ認識されます。このボトルネックに対処するために、オフロードされた MoE 推論のためのパラメータ効率の高いプリフェッチ フレームワークである SpecPrefetch を提案します。 SpecPrefetch は、共有軽量アダプターを使用して、非同期転送の場合にのみ次の層のエキスパート候補を予測しますが、フリーズされたネイティブ ルーターが最終的に実行されるエキスパートを決定します。 SpecPrefetch は、転送予測を実行ルーティングから分離することで、事前トレーニングされたルーティング セマンティクスを変更することなく、公開されたエキスパートの読み込みレイテンシを削減するため、予測エラーはモデルの出力ではなく転送効率に影響します。さらに、ウィンドウ対応スケジューラは、キャッシュと帯域幅の制約の下で実行可能な転送に優先順位を付けます。 Qwen3-VL-30B-A3B と DeepSeek-VL2-Tiny 全体で、SpecPrefetch は、学習された予測子のベースラインよりも大幅に少ないトレーニング可能なパラメーターで、10 個中 9 個のモデル ベンチマーク設定で最高の平均エキスパート リコールを達成しました。 Snapdragon 8 Elite デバイスでは、SpecPrefetch により、コンピューティングに最適化されたオフロード ランタイムと比べてデコード スループットが最大 \(20\%\) 向上し、ストレージに制約のある MoE 導入にとって実用的なメリットが実証されました。コードとモデルの重みは https://github.com/wei390/SpecPrefetch で入手できます。

原文 (English)

SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models

Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited accelerator memory. Although expert offloading alleviates memory pressure by moving inactive experts to host memory or storage, it introduces a routing-dependent transfer bottleneck: required experts are known only after native top-\(K\) routing, which serializes routing, expert loading, and expert execution during inference. To address this bottleneck, we propose SpecPrefetch, a parameter-efficient prefetching framework for offloaded MoE inference. SpecPrefetch uses a shared lightweight adapter to predict next-layer expert candidates only for asynchronous transfer, while the frozen native router still determines the final executed experts. By separating transfer prediction from execution routing, SpecPrefetch reduces exposed expert-loading latency without changing pretrained routing semantics, so prediction errors affect transfer efficiency rather than model outputs. In addition, a window-aware scheduler prioritizes feasible transfers under cache and bandwidth constraints. Across Qwen3-VL-30B-A3B and DeepSeek-VL2-Tiny, SpecPrefetch achieves the best average expert recall in 9 out of 10 model-benchmark settings with substantially fewer trainable parameters than learned predictor baselines. On a Snapdragon 8 Elite device, SpecPrefetch further improves decoding throughput by up to \(20\%\) over a compute-optimized offloading runtime, demonstrating practical benefits for storage-constrained MoE deployment. The code and model weights are available at https://github.com/wei390/SpecPrefetch.

13:00 JSTLLM/生成AI

GLIDE: 効率的な LLM 推論のためのガイド付きレイヤーワイズ ハイブリッド アテンション

大規模言語モデルがますます長いコンテキストに拡張されるにつれて、デコード中のメモリ I/O と Key-Value (KV) キャッシュの計算オーバーヘッドが主なスループットのボトルネックとして浮上します。これに対処するために、スライディング ウィンドウのソフトマックス アテンションと線形再帰集約を戦略的に統合する、ガイド付きレイヤーワイズ ハイブリッド アテンションである GLIDE を提案します。 GLIDE は層ごとの不均質性によって動機付けられています。初期の層はソフトマックスの削除に対して高い感度を示しますが、より深い層は冗長性を示し、線形代替による積極的な置換を許容します。この洞察を活用して、GLIDE はレイヤーごとの適応メカニズムを導入し、各レイヤーが可変サイズのソフトマックス ウィンドウで効率的な線形再帰のバランスをとります。均一なハイブリッド アプローチとは異なり、GLIDE はモデル全体でソフトマックス フットプリントを不均一に圧縮し、最も重要な部分の表現力を維持しながら、総 KV キャッシュ I/O を削減します。実験的評価により、GLIDE は優れたパフォーマンスと効率のトレードオフを達成し、品質を損なうことなく長いコンテキストの生成におけるエンドツーエンドのレイテンシを削減することが実証されています。

原文 (English)

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference

As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during decoding emerges as the primary throughput bottleneck. To address this, we propose GLIDE, a Guided Layerwise Hybrid Attention that strategically integrates sliding-window softmax attention with linear recurrent aggregation. GLIDE is motivated by layer-wise heterogeneity: early layers exhibit high sensitivity to softmax removal, while deeper layers demonstrate redundancy and tolerate aggressive replacement by linear alternatives. Leveraging this insight, GLIDE introduces a layer-wise adaptive mechanism wherein each layer balances an efficient linear recurrence with a variable-sized softmax window. Unlike uniform hybrid approaches, GLIDE non-uniformly compresses the softmax footprint across the model, reducing aggregate KV cache I/O while preserving expressive power where most vital. Empirical evaluations demonstrate the GLIDE achieves superior performance-efficiency tradeoffs, reducing end-to-end latency for long-context generation without compromising quality.

13:00 JST研究/論文

衛星インターネット観測における堅牢なデータ合成のための GAN ベースのフレームワーク

低地球軌道 (LEO) 衛星インターネットは、6G 通信ネットワークに対する国際電気通信連合のビジョンに沿ったユビキタス接続を可能にする重要なインフラストラクチャとなっています。しかし、現在の LEO 衛星によるインターネット観測ではデータの欠落が発生することが多く、データ増強作業が複雑になり、代表的なデータセットの拡張が制限されます。これらのデータセットの複雑な特性を考慮すると、生成 AI (GenAI) は有望なアプローチを示していますが、この分野での応用はこれまでほとんど注目されていません。この論文では、不完全な LEO ネットワーク観測から直接高忠実度データを合成する GenAI ベースのフレームワークを提案します。代表的なデータ欠損シナリオを提案し、最新の WetLinks データセット上で最新の GAN および VAE ベースの GenAI モデルを使用してパフォーマンスを評価します。私たちは、現実世界の LEO 衛星ネットワークで発生するデータ損失を厳密にシミュレートするために、ブロック単位およびポイント単位の欠損シナリオを設計します。私たちの結果は、私たちが提案した GAN ベースのフレームワークの有効性を示しており、GT-GAN モデルは両方の欠如シナリオにおいてすべてのモデルの中で最高のパフォーマンスを示します。極端な条件下(入力データの 40% が欠落しているなど)でも、GT-GAN は最高の堅牢性を示し、基礎となる入力データの分布を一貫して捕捉し、一般化の観点からは最も影響を受けません。私たちの結果は、GenAI ベースのデータ拡張手法と衛星ネットワーク測定に関するデータ駆動型研究の将来の方向性を明らかにします。

原文 (English)

A GAN-Based Framework for Robust Data Synthesis in Satellite Internet Observations

Low-Earth orbit (LEO) satellite Internet has become an important infrastructure for enabling ubiquitous connectivity to align with the International Telecommunications Union vision for 6G telecommunications networks. However, current LEO satellite Internet observations often suffer from missing data, which complicates data augmentation task and limits the expansion of representative datasets. Given the complex characteristics of these datasets, generative AI (GenAI) presents a promising approach, yet its application in this domain has received little attention to date. In this paper, we propose a GenAI-based framework to synthesize high-fidelity data directly from incomplete LEO network observations. We propose the representative data missing scenarios, and evaluate the performance with the latest GAN- and VAE-based GenAI models on the recent WetLinks dataset. We design block-wise and point-wise missing scenarios to closely simulate the data loss that happens on real-world LEO satellite networks. Our results show the effectiveness of our proposed GAN-based framework and GT-GAN model exhibits the best performance among all models in both missing scenarios. Even under extreme conditions (e.g., 40% of the input data is missing), GT-GAN shows the highest robustness, consistently capturing the underlying input data distribution and being the least affected in terms of generalization. Our results shed light on future directions for GenAI-based data augmentation methods and data-driven research on satellite network measurement.

13:00 JSTLLM/生成AI画像/動画生成

記憶による推論: トレーニング不要の長時間ビデオ理解のための時間粒度適応フレームワーク

マルチモーダル大規模言語モデル (MLLM) は、基本的なビデオ タスクにおいて優れた一般化を示しますが、コンテキスト ウィンドウが制限されているため、長時間のビデオの理解が制限されます。この制約に対応するために、モデルは通常、キーフレームの選択を利用します。ただし、均一なサンプリングや静的なクエリに基づく選択では、重要な時間コンテキストが見落とされることが多く、さまざまなクエリの時間粒度に適応できません。この論文では、トレーニング不要の LongVideoQA のための時間粒度適応キーフレーム選択フレームワークである ReMem を提案します。 ReMem は、デュアルレベルのメモリ拡張適応を導入しています。クエリ レベルでは、メモリ主導の質問解析は LLM の長期メモリを利用して質問の時間粒度をデコードし、意味エンティティを抽出します。ビデオ レベルでは、Synergistic Dual-Semantic Frame Alignment が固有の構造メモリを利用してクエリ セマンティクスに合わせてフレームを調整し、構造認識型の動的フレーム ルーティングをガイドしてイベントをクラスタ化し、サンプリング バジェットを最適に分配します。 ReMem は、メモリ メカニズムで時間情報を明示的に保存することで冗長性を抑制し、MLLM が堅牢な複数粒度のビデオ推論を実行できるようにします。 3 つの MLLM を使用した 4 つの一般的な LongVideoQA ベンチマークの評価では、高効率で最先端のゼロショット パフォーマンスが実証されました。特に、ReMem を使用した LLaVA-Video は、LVBench で 54.5% (+12.3%)、LongVideoBench で 67.1% (+8.2%) に達しています。

原文 (English)

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding

While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding. To accommodate this constraint, models typically resort to keyframe selection. However, uniform sampling or static query-guided selection often overlooks critical temporal context, failing to adapt to the varying query temporal granularities. In this paper, we propose ReMem, a temporal granularity-adaptive keyframe selection framework for training-free LongVideoQA. ReMem introduces a dual-level memory-augmented adaptation. At the query level, Memory-Driven Question Parsing leverages LLM long-term memory to decode question temporal granularity and extract semantic entities. At the video level, Synergistic Dual-Semantic Frame Alignment exploits intrinsic structural memory to align frames with query semantics, guiding Structure-Aware Dynamic Frame Routing to cluster events and optimally distribute sampling budgets. By explicitly preserving temporal information with memory mechanisms, ReMem suppresses redundancy and empowers MLLMs to perform robust multi-granular video reasoning. Evaluations across four popular LongVideoQA benchmarks using three MLLMs demonstrate highly efficient, state-of-the-art zero-shot performance; notably, LLaVA-Video with ReMem reaches 54.5% (+12.3%) on LVBench and 67.1% (+8.2%) on LongVideoBench.

13:00 JST研究/論文

最短が最も安全ではない場合: 高齢者に優しい歩行者ルートへのデザインサイエンスのアプローチ

高齢者の自立した移動は、家の外への参加、幸福と健康を可能にしますが、歩行者ナビゲーション システムは依然として距離や時間を主に最適化しており、高齢者の歩行に関する意思決定を形作る障壁、安全基準、支援インフラを見落とすことがよくあります。私たちは、実際のモビリティの制約を規範的な設計知識に変換する、階層的なデザインサイエンス研究を通じて開発された、高齢者に優しい歩行者ルートの成果物を紹介します。 11 件の半構造化インタビューに基づいて、バリアを意識し、アメニティに配慮したルーティングと実行関連の説明のための初期の設計要件 (DR) と設計原則 (DP) を導き出します。これらを、アメニティ (ベンチ、トイレ、シェルター) と高さデータを豊富に備えた OpenStreetMap 歩行者ネットワークでインスタンス化し、構成可能なコストと説明ペイロードを備えた A* ベースのルーティング エンジンを実装しました。野外ベースの歩行研究では、14 人の高齢者が人工物によって生成されたルートとベースラインを比較し、評価と定性的なフィードバックを提供しました。全体的には高齢者に優しいルートが好まれました。さらに、テーマ分析により、インフラストラクチャのメンテナンス、季節条件、交通エクスポージャ、および社会的状況がルートの受け入れを形成していることが示されました。私たちはこれらの洞察を洗練された DR と DP に統合し、状況を認識したハザードモデリング、複数ルートの透明性、ランドマークに基づいた説明、社会的状況の敏感性、段階に応じた情報を強調します。私たちの貢献は、高齢者に優しい歩行者ナビゲーション システムを開発する実務者に実用的なガイダンスを提供します。

原文 (English)

When Shortest Isn't Safest: A Design Science Approach to Senior-Friendly Pedestrian Routing

Older adults' independent mobility enables out-of-home participation, well-being and health, yet pedestrian navigation systems still optimize primarily for distance or time, often overlooking barriers, safety thresholds, and supportive infrastructure that shape late-life walking decisions. We present a senior-friendly pedestrian routing artefact developed through echeloned Design Science Re-search, translating lived mobility constraints into prescriptive design knowledge. Based on 11 semi-structured interviews, we derive initial Design Requirements (DRs) and Design Principles (DPs) for barrier-aware, amenity-sensitive routing and execution-relevant explanations. We instantiate these in an OpenStreetMap pedestrian network enriched with amenities (benches, toilets, and shelters) and height data, and implemented an A*-based routing engine with configurable costs and explanation payloads. In a field-based walking study, 14 older adults com-pared artefact-generated routes with baselines and provided ratings and qualitative feedback; the senior-friendly route was preferred overall. Thematic analysis further showed that infrastructure maintenance, seasonal conditions, traffic exposure, and social context shape route acceptance. We synthesize these insights into refined DRs and DPs emphasizing context-aware hazard modeling, multi-route transparency, landmark-grounded explanations, social-context sensitivity, and stage-appropriate information. Our contributions provide actionable guidance for practitioners developing senior-friendly pedestrian navigation systems.

13:00 JST研究/論文

RRS-10K: 希少なリモート センシング画像解釈のためのマルチタスク視覚言語モデル ベンチマーク

ビジョン言語モデル (VLM) は、一般的なリモート センシング タスクで優れたパフォーマンスを達成しました。ただし、既存のベンチマークは一般的な都市や田舎の画像が大半を占めているため、まれなシーンに対するベンチマークの機能はまだ十分に理解されていません。このギャップに対処するために、希少なリモート センシング画像判読のベンチマークである RRS-10K を紹介します。 RRS-10K には、包括的な評価を目的とした 10,738 枚の軍事関連のリモート センシング画像と、対応する複数形式の質問と回答のペアが含まれています。すべての画像は直接の情報源から収集され、知覚、推論、堅牢性をカバーする 3 つの能力の次元、6 つのサブ次元、および 20 のリーフ タスクに編成されています。多肢選択問題の品質を向上させるために、ベンチマーク構築中に類似性に基づくディストラクタ フィルタリング戦略 (SDFS) を導入します。さらに、52 の代表的なモデルを評価し、現在の VLM はまれなリモート センシング画像解釈では中程度のゼロショット パフォーマンスしか達成できず、視覚的グラウンディング、参照セグメンテーション、および複雑な意味論的推論タスクに明らかな弱点があることを示します。 RRS-10K は、ロングテール リモート センシングの解釈における故障モードの系統的な分析を可能にし、より信頼性の高いリモート センシング VLM を開発するためのガイダンスを提供します。

原文 (English)

RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks. However, their capability for rare scenes remains insufficiently understood, because existing benchmarks are dominated by common urban and rural imagery. To address this gap, we present RRS-10K, a benchmark for rare remote sensing image interpretation. RRS-10K contains 10,738 military-related remote sensing images and corresponding multiple format question-answer pairs for comprehensive evaluation. All of the images are collected from first-hand sources and organized into three capability dimensions, six sub-dimensions, and 20 leaf tasks, covering perception, reasoning, and robustness. To improve the quality of multiple-choice questions, we introduce a similarity-based distractor filtering strategy (SDFS) during benchmark construction. We further evaluate 52 representative models and show that current VLMs achieve only moderate zero-shot performance on rare remote sensing image interpretation, with clear weaknesses in visual grounding, referring segmentation, and complex semantic reasoning tasks. RRS-10K enables systematic analysis of failure modes in long-tail remote sensing interpretation and provides guidance for developing more reliable remote sensing VLMs.

13:00 JST研究/論文

Aletheia: リソースの少ない医療環境における鑑別診断のためのオフラインファーストの臨床意思決定支援システム

サハラ以南のアフリカでは、臨床専門知識へのアクセスが依然として厳しく制限されており、地方では医師と患者の比率が 1:25,000 を下回ることもあります。既存の AI 支援診断ツールは主に信頼性の高いインターネット接続と高仕様のハードウェアを必要とするため、地方の病院や保健センターの最前線の医療従事者にとっては実用的ではありません。この論文では、サハラ以南アフリカ全域の低リソース医療環境向けに設計されたオフラインファーストの臨床意思決定支援システムである Aletheia について紹介します。 Aletheia は Qwen2.5-3B-Instruct に基づいて構築されており、東アフリカで有病率が高い 50 の病状にわたる 27,000 の臨床推論サンプルの厳選されたデータセットに対して量子化低ランク適応 (QLoRA) を使用して微調整されています。評価の結果、10の代表的な臨床症例カテゴリー全体で、トップ1の診断精度が80.0%、トップ3の精度が100.0%、BERTScore-F1が0.909、METEORが0.467であることが実証されました。このシステムは、0.275 の予想キャリブレーション誤差 (ECE) を達成し、アフリカ ディープ テック チャレンジ 2026 (ADTC 2026) のメモリ バジェット制約である 7 168 MB をクリアし、標準化されたベンチマーク ラップトップで約 3 630 MB のピーク推論 RAM を達成しました。これらの結果は、クラウド インフラストラクチャを使用せずに、リソースに制約のある環境でプライマリ ケア レベルで大規模な言語モデルに基づく臨床推論を展開する実現可能性を示しています。

原文 (English)

Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings

Access to specialist clinical expertise remains severely limited across sub-Saharan Africa, where physician-to-patient ratios can fall below 1:25,000 in rural settings. Existing AI-assisted diagnostic tools predominantly require reliable internet connectivity and high-specification hardware, rendering them impractical for frontline healthcare workers in district hospitals and health centres. This paper presents Aletheia, an offline-first clinical decision support system designed for low-resource healthcare contexts across sub-Saharan Africa. Aletheia is built upon Qwen2.5-3B-Instruct, fine-tuned using Quantised Low- Rank Adaptation (QLoRA) on a curated dataset of 27,000 clinical reasoning samples spanning 50 disease conditions with elevated prevalence in East Africa. Evaluation demonstrates a Top-1 diagnostic accuracy of 80.0%, Top-3 accuracy of 100.0%, BERTScore-F1 of 0.909, and METEOR of 0.467 across ten representative clinical case categories. The system achieves an Expected Calibration Error (ECE) of 0.275 and passes the Africa Deep Tech Challenge 2026 (ADTC 2026) memory budget constraint of 7 168 MB, achieving a peak inference RAM of approximately 3 630 MB on the standardised benchmark laptop. These results demonstrate the feasibility of deploying large language model-based clinical reasoning at the primary care level in resource-constrained settings without cloud infrastructure.

13:00 JST研究/論文

AdaKP: 推論指向の強化学習のためのオンライン適応ナレッジポイント選択

検証可能な報酬を伴う強化学習は、大規模な言語モデルで推論を引き出すための強力なパラダイムですが、競技レベルの数学では深刻な報酬の少なさに悩まされます。一般的な解決策は、アトミック ナレッジ ポイント (KP) (ゴールド ソリューションから抽出された短い自然言語ヒント) をプロンプトに挿入します。ただし、既存の方法では、この選択をオフラインで一度修正するか、注入されるテキストのモノリシックな量を単純にスケールして、最も有益な選択軸、つまりアトミック KP のどのサブセットをいつ注入するかはそのままにしておくかのどちらかです。 RL トレーニング中に各問題の KP サブセットを再選択するオンライン セレクターである AdaKP を紹介します。その中心となるのは、高価なロールアウトベースの推定の代わりに、それが誘発する次のトークンのエントロピーの減少によって KP をスコアリングするエントロピー プロキシです。つまり、打ち切りバイアスに証明可能な限界を持つ単一の安価なフォワード パスです。 3 つの軽量メカニズムにより、この信号はオンラインで使用可能になります。ステップごとのノイズを吸収するモメンタム スムーザー、探索を維持しながら弱い KP を除去するリタイアおよび復活マネージャー、再評価を初期トレーニングにフロントローディングする適応スケジューラーです。 AdaKP はさらに、コストのかかる実行が開始される前に、リーブ ワンアウト グランド トゥルースに対してプロキシを認証するプリフライト検証ゲートを提供し、メソッド レベルのリスクを改ざん可能なチェックに変えます。オプティマイザの変更を行わずに標準的な DAPO+GRPO トレーナーの完全な加算フォークとして実現された AdaKP は、無視できる追加コストで 8 つの競技数学ベンチマークすべてで強力な静的選択ベースラインを改善し、オンラインで検証された KP サブセット選択を、推論指向の強化学習のための実用的でまだ検討されていない軸として位置付けます。

原文 (English)

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

Reinforcement learning with verifiable rewards is a powerful paradigm for eliciting reasoning in large language models, yet it suffers from severe reward sparsity on competition-level mathematics. A common remedy injects atomic knowledge points (KPs) - short natural-language hints distilled from gold solutions - into the prompt. Existing methods, however, either fix this selection once offline or merely scale the monolithic quantity of injected text, leaving untouched the most informative axis of choice: which subset of atomic KPs to inject, and when. We introduce AdaKP, an online selector that re-chooses each problem's KP subset over the course of RL training. At its core is an entropy proxy that scores a KP by the reduction in next-token entropy it induces - a single inexpensive forward pass, with a provable bound on its truncation bias - in place of expensive rollout-based estimation. Three lightweight mechanisms make this signal usable online: a momentum smoother that absorbs per-step noise, a retirement-and-revival manager that prunes weak KPs while preserving exploration, and an adaptive scheduler that front-loads re-evaluations into early training. AdaKP further contributes a pre-flight validation gate that certifies the proxy against a leave-one-out ground truth before any expensive run is launched, turning method-level risk into a falsifiable check. Realized as a fully additive fork of a standard DAPO+GRPO trainer with no optimizer changes, AdaKP improves over a strong static-selection baseline on all eight competition-mathematics benchmarks at negligible added cost, positioning online, validated KP-subset selection as a practical and as-yet under-explored axis for reasoning-oriented reinforcement learning.

13:00 JSTLLM/生成AI

MusiChat: 音楽作成のための Vibe 作曲

AI 音楽生成の最近の進歩により、ユーザーは自然言語のプロンプトから完全な音楽作品を作成できるようになりました。しかし、既存のシステムのほとんどは即時再生パラダイムに従っており、ユーザーは既存の音楽アイデアを直接進化させるのではなく、楽曲を繰り返し再作成する必要があるため、反復的な改良が困難になっています。 MusiChat は、自然言語の対話と反復的な改良を通じて人間と AI の共同音楽作成を可能にする、会話型ヴァイブ作曲システムです。 MusiChat の中核となるのは、歌詞に合わせた音楽構造の生成と表現力豊かな表面の実現を分離する、階層的な制御可能な音楽生成フレームワークであり、柔軟なスタイルの変換と構造を保持した編集を可能にします。このシステムは、インタラクション全体でアクティブな作曲状態とユーザー履歴を維持するメモリ拡張アーキテクチャを通じて、大規模な言語モデルとハイブリッド シンボリック ミュージック エンジンを統合します。ハイブリッド インテント ルーティング メカニズムにより、正確な音楽編集と無制限のクリエイティブ リクエストの両方を効率的に解釈できます。 MusiChat は、楽曲を最初から再生成するのではなく、関連する音楽構造とユーザーの意図を維持しながら、進化する音楽成果物を段階的に変換します。私たちは客観的な分析と人間による研究を通じて MusiChat を評価し、シングルターンとマルチターンのインタラクションについてそれぞれ 95.31% と 100% の精度を達成し、メロディーの自然さについては 2:1、音楽の品質については好きと嫌いの比率が 3:1 という結果を得ました。私たちの結果は、MusiChat が対話型インターフェイスを介した一貫したマルチターン音楽オーサリングとインタラクティブな人間と AI の共同制作をサポートしていることを示しています。

原文 (English)

MusiChat: Vibe Composing for Music Creation

Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas. We present MusiChat, a conversational vibe composing system that enables collaborative human-AI music creation through natural-language interaction and iterative refinement. At the core of MusiChat is a hierarchical controllable music generation framework that separates lyric-aligned musical structure generation from expressive surface realization, allowing flexible stylistic transformations and structure-preserving edits. The system integrates a large language model with a hybrid symbolic music engine through a memory-augmented architecture that maintains the active composition state and user history across interactions. A hybrid intent-routing mechanism further enables efficient interpretation of both precise musical edits and open-ended creative requests. Rather than regenerating compositions from scratch, MusiChat incrementally transforms an evolving musical artifact while preserving relevant musical structure and user intent. We evaluate MusiChat through objective analysis and human studies, achieving 95.31% and 100% accuracy for single- and multi-turn interactions, respectively, and obtaining like-to-dislike ratios of 2:1 for melody naturalness and 3:1 for musical quality. Our results demonstrate that MusiChat supports coherent multi-turn music authoring and interactive human-AI co-creation through a conversational interface.

13:00 JST研究/論文

セマンティック ID を理解する: 生成的推奨におけるアイテム表現からアイテム選択まで

セマンティック ID (SID) は現在、生成的推奨の中心的なコンポーネントです。現在の SID ベースのシステムは、同じトークン シーケンスに 3 つの役割を割り当てます。共有プレフィックスは関連するアイテムを整理することを目的としており、完全な SID は個々のアイテムを識別し、生成された各トークンによって返されるアイテムを絞り込みます。私たちは、アイテムのエンコードと SID 構築から自己回帰生成と最終的な推奨に至るまで、SID を体系的に調査します。 SID の構築によって項目の表現がどのように変更されるか、またそれらの変更が生成にどのような影響を与えるかを調べます。 3 つの Amazon ドメインと 8 つの SID 構造にわたって、SID 近傍はエンコーダーの最近傍 10 個のうち平均 32.2% しか回復しません。代替アイテムの説明は、管理されたケースの 99.57% で対応するアイテムを最初に取得しますが、正確な SID の 38.4% は変更されます。これらの結果は、SID が広範な構成を保持しているものの、エンコーダーの細かいローカル構造の多くが失われている一方、その正確なトークンは項目の意味だけでは決定されないことを示しています。この損失は生成中に重大なものになります。最後のセマンティック トークンの後、TIGER は、SID フィルタリング前に妥当な推奨事項であった保留ターゲットの 29.9% のみを保持します。これらの発見に動機付けられて、我々は、ビーム検索がそれらを破棄する前に、ユーザー固有の項目ランキングが対応する SID プレフィックスをサポートできるようにする軽量の推論時間手法である項目サポート デコーディング (ISD) を提案します。同じランキングにより、生成されたアイテムが順序付けされます。 ISD では、パラメーターを追加したり、SID コンストラクターやデコーダーを再トレーニングしたりする必要はありません。我々は、評価されたすべての設定において、ISD が対応する SID バックボーンよりも NDCG@10 を向上させ、相対的に最大 31.2% の向上をもたらすことを経験的に示しています。私たちの結果は、SID は有用な大まかなアイテムの編成を提供しますが、その細かい境界だけで、生成中にどのアイテムが利用可能なままになるかを決定するわけではないことを示しています。

原文 (English)

Understanding Semantic IDs: From Item Representation to Item Selection in Generative Recommendation

Semantic IDs (SIDs) are now a central component of generative recommendation. Current SID-based systems assign three roles to the same token sequence. Shared prefixes are intended to organize related items, the complete SID identifies an individual item, and each generated token narrows the items that can still be returned. We systematically investigate SIDs from item encoding and SID construction to autoregressive generation and final recommendation. We examine how SID construction changes item representations and how those changes affect generation. Across three Amazon domains and eight SID constructions, SID neighborhoods recover only 32.2% of the encoder's ten nearest neighbors on average. Alternative item descriptions still retrieve the corresponding item first in 99.57% of controlled cases, yet change 38.4% of exact SIDs. These results show that SIDs retain broad organization but lose much of the encoder's fine local structure, while their exact tokens are not determined by item meaning alone. This loss becomes consequential during generation. After the final semantic token, TIGER retains only 29.9% of held-out targets that were plausible recommendations before SID filtering. Motivated by these findings, we propose Item-Supported Decoding (ISD), a lightweight inference-time method that allows a user-specific item ranking to support corresponding SID prefixes before beam search discards them. The same ranking then orders the generated items. ISD requires no additional parameters or retraining of the SID constructor or decoder. We empirically show that ISD improves NDCG@10 over the corresponding SID backbone in every evaluated setting, with relative gains of up to 31.2%. Our results show that SIDs provide useful coarse item organization, but their fine boundaries should not alone determine which items remain available during generation.

13:00 JST研究/論文

微分可能な D-vine コピュラによる局所的な異常検出

Vine コピュラは、二変量ペア コピュラへの階層分解を通じて複雑な多変量分布をモデル化するための柔軟なフレームワークを提供します。 D-vine をフィッティングするには、さまざまな依存パターンをエンコードする候補のセットからコピュラ ファミリと各ペア コピュラのパラメーター構成を選択する必要があります。変数と候補ファミリーの数が増加するにつれて、可能な構成の数は組み合わせ的に増加します。既存のフィッティング手順は、連続した貪欲な決定を通じてこの課題に対処し、各ステップで単一の局所的に最適なファミリーにコミットし、より良いグローバル フィットをもたらす構成を潜在的に破棄します。この制限を克服するために、完全微分可能な実装によって可能になる勾配ベースの最尤推定と、フィッティング プロセス全体を通じて複数の競合する D-vine 構成を維持するビーム探索戦略を組み合わせた新しい推定フレームワークを提案します。これにより、計算上扱いやすい状態を保ちながら、構成空間をより広範囲に探索できるようになります。適合した D-vine に基づいて、階層分解を利用してグローバルな異常スコアとエッジレベルの説明の両方を生成する局所的な異常検出フレームワークを導入します。統計的保証はモンドリアンの等角予測によって提供され、ペアコピュラ構造により異常を特定の変数関係に局在化することができます。提案されたフレームワークをベンチマークと現実世界のデータセットの両方で評価し、不確実性の定量化による解釈可能な異常検出に対するその有効性を実証します。

原文 (English)

Localized Anomaly Detection via Differentiable D-vine Copulas

Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas. Fitting a D-vine requires selecting a copula family and parameter configuration for each pair-copula from a set of candidates encoding different dependence patterns. As the number of variables and candidate families increases, the number of possible configurations grows combinatorially. Existing fitting procedures address this challenge through sequential greedy decisions, committing to a single locally optimal family at each step and potentially discarding configurations that would yield a better global fit. To overcome this limitation, we propose a novel estimation framework that combines gradient-based maximum likelihood estimation, enabled by our fully differentiable implementation, with a beam-search strategy that maintains multiple competing D-vine configurations throughout the fitting process. This allows a broader exploration of the configuration space while remaining computationally tractable. Building on the fitted D-vine, we introduce a localized anomaly detection framework that exploits the hierarchical decomposition to produce both global anomaly scores and edge-level explanations. Statistical guarantees are provided through Mondrian conformal prediction, while the pair-copula structure enables the localization of anomalies to specific variable relationships. We evaluate the proposed framework on both benchmark and real-world datasets, demonstrating its effectiveness for interpretable anomaly detection with uncertainty quantification.

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

チャートでサポートされるか、モデルで提供されるか?アクセシブルな視覚化のための MLLM 生成のクレームの検査

マルチモーダル大規模言語モデル (MLLM) は、視覚化パターンを外部の原因、結果、ドメイン知識に結び付けることができますが、これらの解釈の証拠的根拠が不明瞭であることがよくあります。 4 つのソース、3 つの MLLM、および画像へのアクセス、ソース固有のアクセス可能なチャート コンテキスト、および非保持コンテキスト フレーミングへのアクセスを変える 4 つの入力条件からの 102 のビジュアライゼーションの探索的研究を紹介します。 1,224 個の記述にわたって、モデルに起因する DIRECT、DERIVED、および SPECULATIVE ラベルを分析し、数値一致の自動監査を実施します。アクセシブルなチャート コンテキストにより、Gemini と GPT が DIRECT 主張に移行し、一部のモデルの数値一致が改善されました。完全なコンテキストに画像を追加しても、一貫した数値上の利点は得られず、コンテキストを差し控えたプロンプトによっても確実に慎重な言葉遣いが増加することはありませんでした。プロンプトで定義された現実世界の重要性セクションは、主に推測的なままでした。これらの結果は、提供された証拠によって裏付けられた主張とモデルによって提供された解釈を区別する、アクセス可能な記述システムの動機付けとなります。

原文 (English)

Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization

Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, source-specific accessible chart context, and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation

13:00 JSTLLM/生成AIエージェント

SAFAARI: 広告主の応答インテリジェンスを加速するためのスキーマ認識フレームワーク

カスタマー サポート システムの進化は、エージェント チャットボットによって急速に進んでいますが、これらのシステムは、事前定義された API エンドポイントなしで企業データにアクセスする場合、重大な制限に直面しています。このペーパーでは、特殊なコンテンツ、メタデータ、およびオーケストレーション エージェントを通じて、自然言語から SQL (NL から SQL) システムへのスキーマ リンクの重大なボトルネックに対処するマルチエージェント フレームワークである SAFAARI (Schema-Aware Framework for Accelerated Advertiser Response Intelligence) について説明します。また、一貫性のない結果にペナルティを与えながらシステムのパフォーマンスを総合的に評価する新しい複合メトリクスである SEAL (Schema Evaluation and Accuracy in Language-to-SQL) も紹介します。 5 つの機能セット構成による体系的な実験を通じて、SAFAARI は 81.66% の SEAL スコア (ベースラインより 6.65% 向上) を達成し、データポイントの精度 (5.51%) とスキーマリンクの精度 (4.69%) が顕著に向上しました。このフレームワークの有効性は、ドメインの専門家による人間参加型の評価を通じて検証され、さまざまなサポート ドメインにわたる適応性が証明されています。スキーマのリンクとクエリ生成という労働集約的なプロセスを自動化することで、当社のフレームワークは高精度を維持しながら開発時間を 8 分の 1 に短縮することを実証しています。このソリューションは API 開発を合理化し、セルフサービス機能を強化し、特に複雑なデータ エコシステムを備えたカスタマー サポート企業に恩恵をもたらします。

原文 (English)

SAFAARI: Schema-Aware Framework for Accelerated Advertiser Response Intelligence

The evolution of customer support systems is rapidly advancing with agentic chatbots, yet these systems face significant limitations when accessing enterprise data without predefined API endpoints. This paper presents SAFAARI (Schema-Aware Framework for Accelerated Advertiser Response Intelligence), a multi-agent framework that addresses the critical bottleneck of schema linking in Natural Language to SQL (NL-to-SQL) systems through specialized content, metadata, and orchestration agents. We also introduce SEAL (Schema Evaluation and Accuracy in Language-to-SQL), a novel composite metric that holistically evaluates system performance while penalizing inconsistent results. Through systematic experimentation with five feature set configurations, SAFAARI achieves an 81.66% SEAL score (6.65% improvement over baseline), with notable gains in datapoint accuracy (5.51%) and schema-linking precision (4.69%). The framework's effectiveness is validated through human-in-the-loop evaluation with domain experts, which proves its adaptability across diverse support domains. By automating the labor-intensive process of schema linking and query generation, our framework demonstrates 8x reduction in development time while maintaining high accuracy. The solution streamlines API development and enhances self-service capabilities, particularly benefiting customer support enterprises with complex data ecosystems.

13:00 JSTエージェント

CogEEGAgent: グラウンデッド実行と選択認識型検証による自律的な認知 EEG 解析に向けて

認知研究における脳波 (EEG) 分析には専門知識が必要であり、コントラスト、チャネル、時間窓、統計的テストに関して多くの防御可能な選択肢が含まれます。 LLM エージェントは、さまざまな自然言語の質問を分析の選択肢に変換し、自動化のための柔軟なインターフェイスを提供します。しかし、流暢なレポートだけでは、エージェントが要求された分析を選択したこと、または適応検索とは独立して確認的主張を評価したことを証明することはできません。 MNE-Python に基づいた認知脳波分析エージェントである CogEEGAgent を紹介します。 EEG に特化した科学的ハーネスにより、意味論と科学的権威が分離されます。 LLM は意図を解釈し、登録された分析を提案します。一方、決定論的コンポーネントは型指定された契約を検証し、確認アクセスを制御し、証拠に拘束されたリリースを許可します。事前に指定されたルーティング ベンチマークでは、CogEEGAgent は、一致した決定論的ルーターよりも正確に言語を登録された分析にマッピングします。また、一致したプリフライトにより、必要な場合はいつでも両方のシステムが棄権します。外部でモデルが作成された、結果にブラインドなキャンペーンでは、完全なシステムが、参加者間で不一致の確認を伴うサポートされている分析をリリースし、事前に指定された機能の危険性とライフサイクル再利用リクエストをブロックします。ポリシーのストレス テストでは、保留された確認により、修正されていない適応検索による誤検知が抑制されることが示されています。これらの研究を総合すると、認知 EEG ワークフローの制限された自律性と監査可能な自動化フレームワークが確立されます。より広範に、科学エージェントが柔軟な言語理解と推論と解放のフェイルクローズ制御をどのように組み合わせることができるかを示しています。

原文 (English)

CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification

Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contrasts, channels, time windows, and statistical tests. LLM agents can translate varied natural-language questions into analysis choices, offering a flexible interface for automation. Yet fluent reports alone cannot establish that an agent selected the requested analysis or evaluated a confirmatory claim independently of adaptive search. We present CogEEGAgent, a cognitive-EEG analysis agent grounded in MNE-Python. Its EEG-specific scientific harness separates semantic from scientific authority. The LLM interprets intent and proposes registered analyses, while deterministic components validate typed contracts, control confirmation access, and authorize evidence-bound release. On a prespecified routing benchmark, CogEEGAgent maps language to registered analyses more accurately than a matched deterministic router, while matched preflight makes both systems abstain whenever required. In an externally model-authored, outcome-blind campaign, the complete system releases supported analyses with participant-disjoint confirmation and blocks prespecified capability hazards and lifecycle-reuse requests. Policy stress testing shows that held-out confirmation curbs false positives from uncorrected adaptive search. Together, these studies establish bounded autonomy and an auditable automation framework for cognitive-EEG workflows. More broadly, they show how scientific agents can combine flexible language understanding with fail-closed control over inference and release.

13:00 JST研究/論文

会話型 AI の心理的影響: 危害を軽減し、幸福を促進するための研究と設計の方向性

会話型 AI システムが日常生活にますます統合されるにつれて、ユーザーの幸福に対する潜在的な影響には継続的な注意が必要です。消費者向けのジェネラリスト モデルは、情報へのアクセス、学習、生産性、内省、交友関係の向上などの利点を提供できる一方で、感情的なもつれ、不健康な依存、心理的脆弱性の増幅などのリスクももたらします。 AI チャットボットの動作に関する先行研究と経験的観察に基づいて、潜在的な心理的危害を軽減し、ユーザーの幸福をサポートできる方法で汎用 AI システムの動作を導くための一連の意欲的な方向性を提案します。私たちは、AI チャットボットの使用による長期的な影響を体系的に評価することの難しさを認識しており、これらの方向性を、一般的なインタラクション、ロールプレイング シナリオ、および心理的サポートを提供すると特徴づけられる状況全体にわたって、AI の行動がユーザーにどのような影響を与えるかを研究するための仮説として組み立てます。提案された方向性の中には、既存の研究や専門家の洞察によって裏付けられているものもありますが、未解決の疑問やより深い研究が必要な領域を特定しているものもあります。私たちは、この定式化とこれらの仮説が、ユーザーの心理的ニーズにより適切に対応し、ユーザーの幸福を促進することを目的としたインタラクティブ デザイン アプローチのさらなる議論、実証的調査、探求を促進することを願っています。

原文 (English)

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist models can provide benefits, including improved access to information, learning, productivity, self-reflection, and companionship, they also introduce risks, such as emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities. Drawing on prior research and empirical observations of AI chatbot behavior, we propose a set of aspirational directions for guiding the behavior of general-purpose AI systems in ways that may reduce potential psychological harms and support user well-being. We acknowledge the difficulty of systematically assessing the long-term impacts of AI chatbot use and frame these directions as hypotheses for studying how AI behavior may influence users across general interactions, role-playing scenarios, and contexts that could be characterized as providing psychological support. While some proposed directions are supported by existing research and expert insights, others identify open questions and areas requiring deeper study. We hope that this formulation and these hypotheses encourage further discussion, empirical investigation, and exploration of interactive design approaches aimed at better accommodating users' psychological needs and promoting their well-being.

13:00 JST研究/論文

類似したモデルは異なる学習を行う: 最終ウィンドウの事前トレーニングが SFT を超えたトレーニング後の形状を形成する

開発者は、モデルのチェックポイントがどのように動作するかによって判断します。教師あり微調整 (SFT) 後、関連するベンチマーク間でほぼ同じパフォーマンスを示す 2 つのチェックポイントは交換可能として扱われ、次の調整段階 (通常はプリファレンスの最適化) に同様に準備が整います。この判断がトレーニング前の痕跡を見逃しているかどうかを尋ねます。この違いは、SFT 後のベンチマークでは明らかになっていないものの、その後のトレーニングに各チェックポイントがどのように反応するかを決定します。それを確認するために、事前トレーニングの最後のウィンドウ、つまり命令調整の前にトレーニングされた最後のデータに対して制御された実験を実行します。 6 つのブランチは部分的に事前トレーニングされた 1 つのチェックポイントから分岐しており、このウィンドウのみが異なります: 5 億トークン (その前のトークンの 0.1% ~ 1%)。各ブランチは、単一のデータ ソース (一般的な Web テキスト、フィルタリングされた Web テキスト、規範的談話、安全テキスト、数学テキスト、合成教育テキスト) でウィンドウをトレーニングします。 SFT とポストトレーニングは同一になります。 SFT の後、ブランチは、命令追従、拒否、および能力に関して約 1 ポイント内でほぼ同一に動作しますが、同じポストトレーニングによって、直接優先最適化更新と検証可能な報酬を伴う強化学習更新の両方の下で、ブランチが非常に異なるエンドポイントに運ばれます。この逸脱は、有害なリクエストの拒否を通じて測定されます。ポストトレーニングが始まると、安全テキスト ブランチは Web テキスト ブランチと同じくらい拒否しますが、最後までに拒否される量ははるかに少なくなります。他の 4 つのブランチは保護をほとんどまたはまったく得られないため、その効果はウィンドウに含まれる内容に応じて選択されます。この保護では、安全テキストが事前トレーニングの最初ではなく最後に到着する必要があり、これは 2 番目のモデル ファミリで再現されます。モデルが最後の形状でどのように事前トレーニングされているか、位置合わせにどのように反応するか。したがって、チェックポイントは SFT 後の動作だけで評価されるべきではなく、最後にトレーニングされた内容も併せて報告する必要があります。

原文 (English)

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant benchmarks are treated as interchangeable, equally ready for the next alignment stage, typically preference optimization. We ask whether this judgment misses a pretraining imprint: a difference that no post-SFT benchmark reveals, yet that decides how each checkpoint responds to further training. To find out, we run a controlled experiment on the final window of pretraining, the last data trained on before instruction tuning. Six branches fork from one partially pretrained checkpoint and differ only in this window: 500 million tokens, 0.1% to 1% of the tokens that precede it. Each branch trains its window on a single data source: generic web text, filtered web text, normative discourse, safety text, mathematical text, or synthetic educational text. SFT and post-training are then identical. After SFT the branches behave near-identically, within about one point on instruction following, refusal, and capability, yet the same post-training carries them to very different endpoints, under both a direct preference optimization update and a reinforcement learning update with a verifiable reward. We measure this deviation through refusal of harmful requests: when post-training begins the safety text branch refuses no more than the web text branch, yet by the end it has lost far less of its refusal. The other four branches gain little or no protection, so the effect is selective to what the window contained. The protection requires the safety text to arrive last rather than earlier in pretraining, and it reproduces on a second model family. What a model is pretrained on last shapes how it reacts to alignment. Therefore, a checkpoint should not be evaluated by its post-SFT behavior alone, and what it was trained on last should be reported with it.

13:00 JSTLLM/生成AIエージェント

AI エージェントにおける長いコンテキスト ウィンドウ制御のためのアドレス指定可能なリコール コンパクション

長期的な LLM エージェントは推論トレース、アクション、ツールの観察を蓄積し、最終的にはモデルの固定コンテキスト ウィンドウを超える可能性があります。既存の圧縮方法は、以前の情報を破棄、要約、または取得することでこの制限に対処していますが、タスクに不可欠な詳細が削除されたり、それらを確実に回復できない可能性があります。私たちは、アクティブ コンテキストのプレゼンテーションからアーカイブ ストレージを分離するコンテキスト管理フレームワークである ARC (Addressable Recall Compaction) を提案します。 ARC は、ツールの観察を追加専用の ID アドレス指定可能なログに保存し、圧縮が必要な場合には古い観察をコンパクトな引用に置き換えます。その後、エージェントはこれらの識別子を使用して、対応するツールを再実行したり、類似性に基づく検索のみに依存したりすることなく、保存されたコンテンツを要求できます。 16k コンテキスト ウィンドウを持つ Qwen3-8B と 32k コンテキスト ウィンドウを持つ Qwen3-32B を使用して ARC を評価します。 Needle-in-a-Haystack の評価では、ARC は平均正確回答精度 99.40% を達成しました。これに対し、当社の評価で最も優れたベースラインの精度は 88.12% でした。 ARC は、ハードウェア コスト モデルに基づいて、推定サービス時間と HBM トラフィックも削減します。 LongBench-v2 Hard サブセットでは、ARC の平均精度は 29.97% ですが、最もパフォーマンスの高いベースラインでは 28.25% でした。これらの結果は、明示的なアドレスベースの呼び出しにより、テストされた設定で評価されたコンテキスト管理ベースラインと比較して、情報の保持と提供効率が向上する可能性があることを示しています。

原文 (English)

Addressable Recall Compaction for Long Context-Window Control in AI Agents

Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing compaction methods address this limitation by discarding, summarizing, or retrieving earlier information, but they may remove task-critical details or fail to recover them reliably. We propose ARC (Addressable Recall Compaction), a context-management framework that separates archival storage from active-context presentation. ARC stores tool observations in an append-only, ID-addressable log and replaces older observations with compact citations when compaction is required. The agent can subsequently use these identifiers to request stored content without re-executing the corresponding tools or depending solely on similarity-based retrieval. We evaluate ARC using Qwen3-8B with a 16k context window and Qwen3-32B with a 32k context window. On the Needle-in-a-Haystack evaluation, ARC achieves an average exact-answer accuracy of 99.40%, compared with 88.12% for the best-performing baseline in our evaluation. ARC also reduces estimated serving time and HBM traffic under our hardware-cost model. On the LongBench-v2 Hard subset, ARC obtains an average accuracy of 29.97%, compared with 28.25% for the best-performing baseline. These results indicate that explicit, address-based recall can improve information retention and serving efficiency relative to the evaluated context-management baselines under the tested settings.

13:00 JSTLLM/生成AI

推奨者はどのくらいの頻度で LLM に電話をかける必要がありますか?値に重み付けされたルーティング、モニタリング、季節性の堅牢性

安価なヒューリスティックと高価な大規模言語モデル (LLM) の間のルーティングの決定は、通常、困難な問題として組み立てられます。つまり、困難なケースを高価なパスに送信します。困難とビジネス価値は別個の軸であるため、この枠組みは不完全であると私たちは主張します。困難で安価な項目と困難で高価な項目のエラーのコストは同じではありません。私たちは、小売マーチャンダイジング パイプラインの完全合成シミュレーションである Value Router を紹介します。これは、地上の真実ではなく、推定難易度と推定価値のみを使用してアイテムをルーティングします。研究には 3 つの段階があります。まず、値に重み付けされたしきい値ルーターが、カテゴリのボリュームと値の間に逆相関がある合成カタログ上の難易度のみのランダムなベースラインと比較されます。価値の重み付けは、真の高価値アイテムの難易度のみのベースラインの再現率 (60%) と一致し、大幅に高い精度 (98.3% 対 94.3%) を達成します。第 2 に、意思決定ロガーとモニターは、集計メトリクスによって隠された故障モードを明らかにし、集計結果が品目ごとの区別ではなくカテゴリ間の差異によってほぼ完全に左右されることを示します。 3 番目に、シミュレートされたブラック フライデーの需要急増 (より価値の高いカテゴリーへのシフトによる 2.5 ボリューム) では、静的ルーター、季節調整ルーター、および 2 つの低速パス予算ポリシーを比較します。すべての結果は、実験者が定義したグラウンド トゥルースを使用した制御された合成シミュレーションからのものであり、検証された現実世界の主張ではなく、コストを意識したルーティング システムの設計原則を示しています。

原文 (English)

How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness

Routing decisions between a cheap heuristic and an expensive large language model (LLM) are typically framed as a difficulty problem: send the hard cases to the expensive path. We argue this framing is incomplete because difficulty and business value are distinct axes - a difficult cheap item and a difficult costly item do not have the same cost of error. We present Value Router, a fully synthetic simulation of a retail merchandising pipeline that routes items using only estimated difficulty and estimated value, never ground truth. The study has three stages. First, a value-weighted threshold router is compared with a difficulty-only and a random baseline on a synthetic catalog with an inverse correlation between category volume and value. Value-weighting matches the difficulty-only baseline's recall of true high-value items (60%) while achieving substantially higher precision (98.3% vs. 94.3%). Second, a decision logger and monitor expose a failure mode hidden by aggregate metrics showing that the aggregate result is driven almost entirely by between-category differences rather than per-item discrimination. Third, a simulated Black Friday demand surge (2.5 volume with a shift toward higher-value categories) compares a static router, a seasonally tuned router, and two slow-path budget policies. All results are from a controlled synthetic simulation with experimenter-defined ground truth and illustrate design principles for cost-aware routing systems rather than validated real-world claims.

13:00 JSTエージェント

エージェント オペレーティング システムに向けて - クラシック OS とクラウド OS からの教訓

プラットフォーム ソフトウェアの主要な波はすべて同じ弧をたどります。最初は競合するフレームワークとその場限りの実装を試し、その後、明確に定義されたセマンティクスを備えた安定した抽象化の小さなセットが明確になり、最後にそれらの抽象化をアプリケーションが移植可能なプラットフォームに統合します。 POSIX は従来のオペレーティング システムに対してこれを行っていました。 Kubernetes はクラウドのためにそれを行いました。エージェント AI システム (計画、ツールの使用、メモリの維持、および共同作業を行う自律的な LLM 主導のエージェント) は、現在、そのような第 3 の波の実験段階にあります。数十のフレームワークやプロトコルが登場しましたが、核となる抽象化が何であるか、またはそれらがどのような保証を提供するかについてコミュニティのコンセンサスは存在しません。そのコンセンサスがなければ、エージェント アプリケーションを移植可能に作成することはできず、プラットフォームを確実に構成することはできず、この分野はプロトタイプの展開を超えて前進することはできません。私たちは、前進する道は、以前のウェーブの方法論に従うことであると主張します。つまり、古典的な OS とクラウド OS のプリミティブを確率的で自然言語媒介の実行に拡張することによって新しいエージェントの抽象化を導き出し、それらのセマンティクスを正確に指定し、それらを中心に統合します。これは、POSIX と Kubernetes がそれぞれのウェーブを統合したのと同じです。

原文 (English)

Towards an Agent Operating System - Lessons from Classical and Cloud OS

Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, followed by the articulation of a small set of stable abstractions with well-defined semantics, and finally consolidation around those abstractions into a platform that applications can portably target. POSIX did this for classical operating systems; Kubernetes did it for the cloud. Agentic AI systems - autonomous, LLM-driven agents that plan, use tools, maintain memory, and collaborate - are currently in the experimentation phase of the third such wave. dozens of frameworks and protocols have emerged, but no community consensus exists on what the core abstractions are or what guarantees they carry. Without that consensus, agentic applications cannot be written portably, platforms cannot compose reliably, and the field cannot advance beyond prototype deployments. We argue that the path forward is to follow the prior-wave methodology: derive new agentic abstractions by extending classical OS and cloud OS primitives to stochastic, natural-language-mediated execution, specify their semantics precisely, and consolidate around them - just as POSIX and Kubernetes consolidated their respective waves.

13:00 JSTエージェント

PLATO: エージェントとタスクのオープン性のためのポインター学習器

オープン エージェント システム (OASYS) は、エージェントとタスクのセットが時間の経過とともに予期せず変化する現実の領域でますます普及しています。エージェントのオープン性 (AO) やタスクのオープン性 (TO) を含むこのようなオープン性は、通常、固定状態とアクション空間を前提とするマルチエージェント強化学習 (MARL) に対して根本的な課題を引き起こします。既存の方法は、開放性を部分的にしか扱っていません。パディングおよびマスキングのアプローチでは人為的な境界が導入されていますが、最近のグラフベースまたはハイパーグラフの方法では、開放性の一次元を扱いますが、依然として限定的な仮定に依存しています。この論文では、集中型トレーニングと分散型実行パラダイムの下でマルチエージェント近接ポリシー最適化でトレーニングされた、集中型グラフ ニューラル ネットワーク (GNN) クリティカルと組み合わせたポインター ネットワーク ベースのアクターである、エージェントとタスク オープン性のためのポインター学習者 (PLATO) を紹介します。ポインターベースのアクターは、現在のタスク セット上に直接分布を出力します。これは、マスキングや再トレーニングを行わずに、アクション スペースの変更を直接サポートします。私たちの GNN 批評家は、エージェントとタスクの相互作用を、タスクとエージェントの構成に応じて形状が変化するグラフとしてエンコードします。これらのコンポーネントを合わせて、既存のアプローチに制限されることなく AO と TO を考慮します。我々は、以前のタスクオープン定式化を拡張して、タスクアンドエージェントオープンマルコフゲーム(TaAgO-MG)でPLATOを定式化し、それが結果として生じる無制限の状態およびアクション空間にわたって明確に定義されていることを証明します。私たちは、オープン マルチエージェント システム評価用に設計された環境である Methods for Open Agent Systems Evaluation Initiative (MOASEI) の野火抑制ドメインを使用して PLATO を評価し、OASYS の最先端のベースラインよりも強力なパフォーマンスとより一貫したゼロショット汎化を実証しました。

原文 (English)

PLATO: Pointer Learner for Agent and Task Openness

Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning (MARL), which typically assumes fixed state and action spaces. Existing methods address openness only partially: padding and masking approaches introduce artificial bounds, while recent graph-based or hypergraph methods handle one dimension of openness but still depend on restrictive assumptions. In this paper, we introduce Pointer Learner for Agent and Task Openness (PLATO), a pointer-network-based actor combined with a centralized graph neural network (GNN) critic, trained with multi-agent proximal policy optimization under a centralized training and decentralized execution paradigm. Our pointer-based actor outputs distributions directly over the current task set. This directly supports changing action spaces without masking or retraining. Our GNN critic encodes agent-task interactions as a graph that changes shape with task and agent composition. Together, these components consider AO and TO without the boundedness of existing approaches. We formalize PLATO in a Task-and-Agent-Open Markov Game (TaAgO-MG), extending prior task-open formulations, and prove it is well-defined over the resulting unbounded state and action spaces. We evaluate PLATO with the Methods for Open Agent Systems Evaluation Initiative (MOASEI) wildfire suppression domain, an environment designed for open multi-agent system evaluation, and we demonstrate strong performance and more consistent zero-shot generalization than state-of-the-art baselines in OASYS.

13:00 JSTエージェント

Matryoshka エージェント: 長期的な機械学習エンジニアリングのためのサブエージェントの展開

機械学習エンジニアリング (MLE) タスクでは、高価でフィードバック主導型の環境相互作用の下で、ソリューションのデバッグと改良を反復して長期的な意思決定を行う必要があります。このようなタスク用のモノリシック エージェントの開発とトレーニングは、非常に長くノイズの多いコンテキストを同時に管理し、広大なソリューション空間を探索し、限られたモデル容量と計算予算の下で効果を維持する必要があるため、基本的に困難です。これらの課題に対処するために、長期にわたる複雑なタスクのための統合された階層型エージェント フレームワークである Matryoshka Agent を提案します。 Matryoshka エージェントは、エージェントによる問題解決を、意思決定と実行の調整された階層に分解します。上位レベルのオーケストレーターは、コンパクトで長期的な探索状態を維持し、戦略的指示を発行します。一方、下位レベルのサブエージェントは、標準化されたツール インターフェイスを介した直接的な環境対話を通じて、具体的な解決策の試みを実行します。この設計により、戦略的探索がコストのかかる実行から切り離され、長いコンテキストの推論の負担が大幅に軽減され、効率的な反復改良が可能になります。私たちはさらに、Matryoshka Agent の効率的なトレーニング パラダイムを開発します。さまざまなモデル タイプとスケールを使用した幅広い MLE タスクに関する実験結果は、Matryoshka Agent が長期的な MLE タスクと複雑なエージェントの問題解決にとって効果的でスケーラブルなパラダイムであることを示しています。特に、Matryoshka Agent により、Qwen3-4B-Instruct が o4-mini に匹敵する Orchestrator パフォーマンスを達成できるようになります。 Matryoshka Agent を Qwen3-30B-Coder に適用すると、最大 36.7% の相対パフォーマンス向上が得られます。

原文 (English)

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering

Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment interactions. Developing and training a monolithic agent for such tasks is fundamentally challenging, as it must simultaneously manage extremely long and noisy contexts, explore vast solution spaces, and remain effective under limited model capacity and computational budgets. To address these challenges, we propose Matryoshka Agent, a unified hierarchical agent framework for complex long-horizon tasks. Matryoshka Agent decomposes agentic problem solving into a coordinated hierarchy of decision making and execution: a high-level Orchestrator maintains compact, long-horizon exploration states and issues strategic instructions, while lower-level Sub-Agents execute concrete solution attempts through direct environment interaction, mediated by standardized Tool interface. This design decouples strategic exploration from costly execution, substantially reducing the burden of long-context reasoning and enabling efficient iterative refinement. We further develop an efficient training paradigm for Matryoshka Agent. Experimental results on a broad range of MLE tasks with diverse model types and scales demonstrate that Matryoshka Agent is an effective and scalable paradigm for long-horizon MLE tasks and complex agentic problem solving. Notably, Matryoshka Agent enables Qwen3-4B-Instruct to reach Orchestrator performance comparable to o4-mini. Applying Matryoshka Agent to Qwen3-30B-Coder results in at most 36.7% relative performance gain.

13:00 JSTLLM/生成AIエージェント

小規模言語モデルエージェントのための堅牢な強化学習に向けて

強化学習を使用した 70 ~ 500M パラメーター範囲での小型言語モデル (SLM) の調整は、根底にある失敗メカニズムが系統的に調査されていないにもかかわらず、不安定であると考えられています。 State-of-the-Art (SOTA) 研究では、近接ポリシー最適化 (PPO) を使用して 15 個の (モデル、コーパス) 構成がトレーニングされました。実験には、TinyStories、CNN/DailyMail、Wikitext-103 コーパス上の Pythia-70M、160M、410M および SmolLM2-135M、360M が含まれていました。小規模言語モデルでは、再現可能な 3 つの障害モードが特定されました。標準 PEFT/TRL パイプラインでのサイレント LoRA パラメーターのフリーズ、bfloat16 使用時の重要度比の数値オーバーフロー、および報酬モデルのエラーによる壊滅的なポリシーの崩壊です。これらの問題は、アダプターのマージと再初期化手法、PPO 更新時の float32 精度、報酬のホワイトニング、重要度の比率の保護、重みのロールバックからなる 3 層の安全メカニズムを使用して解決されました。この論文では、SLM スケールでの PPO パフォーマンスが、モデル パラメーターの数ではなく、滑らかな教師ありモデル ($\text{PPL}<20$) と識別報酬信号の両方に依存するというキャパシティ ヘッドルーム仮説が提案されています。提案されたシステムはすべての実験で安定して収束し、流暢な事前信号と有益な報酬信号を備えた構成で SFT ベースラインを上回る優先勝率を向上させました。さらに、必要なトレーニング データの量が大幅に減少しながら、命令調整されたベースラインを上回りました。すべてのチェックポイント、設定データセット、トレーニング スクリプトは公開されています$^{\S}$。

原文 (English)

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations were trained using Proximal Policy Optimization (PPO). The experiments included Pythia-70M, 160M, 410M and SmolLM2-135M, 360M on the TinyStories, CNN/DailyMail, and Wikitext-103 corpora. Three reproducible failure modes were identified in small-scale language models: silent LoRA parameter freezing in standard PEFT/TRL pipelines, numerical overflow in importance ratios when using bfloat16, and catastrophic policy collapse due to reward-model error. These issues were addressed using a merge-and-reinitialize adapter technique, float32 precision during PPO updates, and a three-layer safety mechanism comprising reward whitening, importance-ratio guarding, and weight rollback. In this paper, a capacity-headroom hypothesis is proposed, which states that PPO performance at the SLM scale depends on both a fluent supervised model ($\text{PPL}<20$) and a discriminative reward signal, rather than on the number of model parameters. The proposed system converged stably in all experiments and improved preference win rate over the SFT baseline in configurations with a fluent prior and an informative reward signal. Furthermore, it outperformed instruction-tuned baselines while requiring significantly less training data. All checkpoints, preference datasets, and training scripts are publicly released$^{\S}$.

13:00 JST研究/論文

ScalableRAG: 取り込みコストゼロの高品質 RAG

RAG の最近の進歩は、ナレッジ グラフの構築や SQL テーブルの抽出といったナレッジの取り込みに高額の取り込みコストを支払うことで、パフォーマンスを最適化することを目的としています。この研究では、そのような知識ベースで許可される操作が、取り込みコストゼロ (ベクトル データベースであっても) で複製できることを示します。実際、私たちのソリューションである Zero-Ingestion ScalableRAG は、ここで検討した 6 つのコーパスのうち 3 つですべてのベースライン (ナレッジ グラフ アプローチを含む) を簡単に上回っていますが、他の 3 つでは最大パフォーマンスにわずかに及ばないだけで、6 つすべてのデータセットの平均精度は次に競争力の高いベースラインを 7.36% 上回っています。これは、書き込みと読み取りが可能なドキュメント セットと値セットのワークスペースを保持することでこれを実現し、ドキュメント セット全体のサブセットと 1 対 1 で対応する主キーでのグループ化が必要なあらゆる状況で、オンザフライの集約推論を可能にします。コーパス サイズに依存しない定数によって LLM 呼び出しの数を制限し、大規模な精度をさらに向上させるために、最小限のベクトル データベースとドキュメントのサンプルからの自動パターン検出を使用する Limited-Ingestion ScalableRAG も導入します。私たちのコードは https://github.com/cohesity/ScalableRAG で入手できます。

原文 (English)

ScalableRAG: High-Quality RAG at Zero Ingestion Cost

Recent advances in RAG aim to optimize for performance by paying high ingestion costs for knowledge ingestion: building knowledge graphs or extracting SQL tables. In this work we show that the operations that such knowledge bases allow can be replicated with zero ingestion costs (not even a vector database); in fact our solution, Zero-Ingestion ScalableRAG, handily out-performs all baselines (including knowledge graph approaches) in three out of the six corpora considered here, and only marginally missing maximum performance on the other three, with average accuracy across all six datasets 7.36% above the next most competitive baseline. It achieves this by keeping a workspace of document sets and values sets that it can write into and read from, allowing for on-the-fly aggregative reasoning in all situations where grouping is required on a primary key that is in one to one correspondence with a subset of the total document set. Capping the number of LLM calls by a constant independent of the corpus size, we also introduce Limited-Ingestion ScalableRAG, which does use a minimal vector database as well as an automated pattern discovery from a sample of documents, to further improve accuracy at scale. Our code is available at https://github.com/cohesity/ScalableRAG .

13:00 JST研究/論文ClaudeGPT / ChatGPTMistral AI

データが少なく、調整が向上: 嗜好の最適化のためのデータ中心の複数評価者の合意

好みの最適化に関する研究では、データを固定したままトレーニングの目的を変更することがよくあります。代わりに、小規模で信頼度の高いポリシーに関する応答セットが信頼できる学習シグナルを提供できるかどうかを尋ねます。私たちの手法である DMAPO (Data-centric Multi-evaluatorAgreement for Preference Optimization) は、ターゲットポリシーから回答候補を生成し、ルーブリックに特化した評価者によって有用性、事実性、簡潔性を評価し、プロセス批判の修正を適用し、コンセンサスの高い望ましい例または望ましくない例のみを保持します。この手順では、54,236 人のミストラル-7B 候補者のうち 1,871 人 (3.45%) が受け入れられます。このセットでトレーニングされた KTO は、MT-Bench で 7.50、text-davinci-003 リファレンスに対する長さ制御の勝率 95.5%、IFEval プロンプト精度 57.3% に達しました。独立したペアごとの評価でも、SimPO よりも DMAPO が有利です。GPT-4o は、129 の保留されたプロンプトで 23.3 ポイント、配布外の 200 の LMSYS-Chat プロンプトで 24.0 ポイントの純勝率をもたらしました。クロード オーパス 4.7 は、ホールドアウト セットで 24.1 ポイントを獲得しました。評価モデルまたはルーブリックを変更すると、選択された例が変更されますが、下流のパフォーマンスにはほとんど影響しません。 2 番目のバックボーン調査でも同様の 3.41% の合格率が得られますが、パフォーマンスの向上はより控えめです。これらの実験全体にわたって、コンセンサス フィルタリングは、追加のキュレーション計算と評価者の判断への依存を犠牲にして、一般的な命令の優先順位を最適化するためのデータ効率の高いルートを提供します。

原文 (English)

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimization), generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples. This procedure accepts 1,871 of 54,236 Mistral-7B candidates (3.45%). KTO trained on this set reaches 7.50 on MT-Bench, 95.5% length-controlled win rate against a text-davinci-003 reference, and 57.3% IFEval prompt accuracy. Independent pairwise evaluation also favors DMAPO over SimPO: GPT-4o yields a net win rate of 23.3 points on 129 held-out prompts and 24.0 points on 200 out-of-distribution LMSYS-Chat prompts; Claude Opus 4.7 yields 24.1 points on the held-out set. Changing the evaluator model or rubric alters the selected examples but has little effect on downstream performance. A second-backbone study yields a similar 3.41% acceptance rate, although its performance gains are more modest. Across these experiments, consensus filtering offers a data-efficient route to preference optimization for general instructions, at the cost of additional curation compute and dependence on evaluator judgments.

13:00 JSTLLM/生成AIエージェント研究/論文

LLM エージェント間で影響がどのように伝播するか: 群衆シミュレーションにおける緊急の感情伝染

この論文では、相互に認識し評価するエージェント間で感情がどのように伝播するかに焦点を当て、マルチエージェント群集シミュレーションにおける言語モデルの動作を研究します。各エージェントは、視覚、聴覚、触覚の各チャネルを通じて隣人を認識し、その刺激された性格プロファイル、記憶、現在の感情状態、および状況コンテキストに照らしてこれらの認識を評価します。評価は LLM によって実行され、エージェントの内部感情状態が更新され、その外部表現が選択されます。このアーキテクチャには、エージェント間で感情状態を直接転送するための手動で作成されたメカニズムは含まれていません。その代わりに、エージェント間の影響は、知覚-評価-表現のループを通じて生じます。エージェントの表現は、ビッグ ファイブの性格モデルとラッセルの複雑な感情モデルに基づいています。遅延を制限するために、低レベルのステアリングとナビゲーションは、LLM ベースのコグニティブ レイヤーとは独立して動作する従来の群衆シミュレーターによって処理されます。私たちは、さまざまな空間レイアウトにおける、憂慮すべき状況、楽しい状況、中立的な状況にわたる 5 つのシナリオ環境にわたってアーキテクチャを評価します。結果は、このシステムが、まばらで小さな群衆の中で、空間的、時間的、そして性格に依存した構造を持つ感情伝染力学を生み出すことを示しています。警報は播種されたエージェントから移動前線として広がり、平均警報割合はゼロ以外のプラトーに落ち着き、促された性格プロファイルの分布によって、曖昧な警報がパニックを引き起こすかどうか、挑発が怒りまたは恐怖と解釈されるかどうかが決まります。さらに、プロンプト バリアント、サンプリング温度、および 4 つのモデル バックエンドにわたる制御された実験を通じて評価ステップを評価し、ダイナミクスがバックエンドに依存していることを示します。

原文 (English)

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation

This paper studies the behavior of language models in a multi-agent crowd simulation, focusing on how affect propagates among agents that perceive and appraise one another. Each agent perceives its neighbors through visual, auditory, and tactile channels, then appraises these perceptions in light of its prompted personality profile, memory, current affective state, and situational context. Appraisal is carried out by an LLM, which updates the agent's internal affective state and selects its outward expression. The architecture contains no hand-authored mechanism for directly transferring affective state between agents; instead, inter-agent influence arises through the perception-appraisal-expression loop. The agent representation draws on the Big Five personality model and Russell's circumplex model of affect. To limit latency, low-level steering and navigation are handled by a conventional crowd simulator operating independently of the LLM-based cognitive layer. We evaluate the architecture across five scenario environments spanning alarming, joyful, and neutral situations in different spatial layouts. The results show that the system produces emotional contagion dynamics with spatial, temporal, and personality-dependent structure in sparse, small crowds. Alarm spreads from seeded agents as a traveling front, the mean alarmed fraction settles at a nonzero plateau, and the distribution of prompted personality profiles determines whether an ambiguous alarm ignites panic and whether a provocation is interpreted as anger or fear. We further evaluate the appraisal step through controlled experiments across prompt variants, sampling temperatures, and four model backends, showing that the dynamics are backend-dependent.

13:00 JST研究/論文

時間畳み込みネットワークを使用した欠落した軌跡データの推論

現実世界の設定で収集された軌跡データは、センサーの故障、通信損失、または遮蔽により不完全になることがよくあります。私たちは \emph{軌道修復} というタスクに取り組みます。つまり、観察されたコンテキストから連続した欠落セグメントを再構築します。私たちは、標準的な因果関係制約を緩和する対称拡張を備えた時間畳み込みネットワーク (TCN) を提案します。これにより、各タイム ステップで過去と将来の両方の観測を利用できるようになります。これは修復には不可欠ですが、予測指向のアーキテクチャには存在しない特性です。モデルは、重み付き平均二乗誤差、境界-連続性ペナルティ、および平滑性正則化を組み合わせた複合損失を使用してトレーニングされます。 20% のマスクされたセグメントがランダムに配置された $1,000$ (トレーニング)、$200$ (検証)、$300$ (テスト) の 2 次元軌跡の合成データセットでトレーニングされたモデルは、優れた R$^{2}$、MSE、MAE メトリクスを達成しました。

原文 (English)

Inferring Missing Trajectory Data with Temporal Convolutional Networks

Trajectory data collected in real-world settings is frequently incomplete due to sensor failure, communication loss, or occlusion. We address the task of \emph{trajectory inpainting}: reconstructing contiguous missing segments from observed context. We propose a Temporal Convolutional Network (TCN) with symmetric dilation that relaxes the standard causality constraint, allowing each time step to draw on both past and future observations, a property that is essential for inpainting, but absent from forecasting-oriented architectures. The model is trained with a composite loss that combines weighted mean squared error, boundary--continuity penalties, and a smoothness regularizer. Trained on a synthetic dataset of $1,000$ (train), $200$ (validation), and $300$ (test) two-dimensional trajectories with randomly placed 20% masked segments, the model achieves good R$^{2}$, MSE and MAE metrics.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

エージェント ループが停滞を進歩と誤認するのはどのような場合ですか?長期実行される自律 LLM エージェント ループにおける自己評価バイアスと外部接地検証

長期にわたって実行される自律エージェントは、人間の介入なしに自ら計画、実行、完了を判断します。エージェントが自分の仕事を採点すると、自己評価バイアスが定着します。現実世界の結果は停滞または後退する一方で、もっともらしい変化は進歩として受け入れられます。私たちはこの故障モードを進行蜃気楼と名付け、制御された測定によって、それが評価器が何に基づいているのかという問題であることを示しました。私たちは、エージェントとそのツール表面を固定し、ループをゲートする評価器の情報チャネルタイプのみを操作するテストベッドを構築しました。原則として偽造不可能な世界国家の神託は、コンテナとネットワークの分離によって強制され、実行のたびに検証されます。 54 サイクルにわたって、フロンティア エージェントは毎回改善を主張しましたが、56 パーセントの測定デルタはゼロ以下でした。したがって、自己報告は有益ではなく、自己判定ゲートはすべてを受け入れるものに変質し、それまで到達していた最良の展開状態が 19% 損なわれました。完全なアーティファクトテキスト、変更差分、および独自の評決履歴を読んだ最も強力なバンド内裁判官でさえ、44% が現実世界の後退であり、38% の実際の改善を拒否したサイクルを受け入れました。強いジャッジがギャップを埋めるという事前登録された敵対的仮説は却下された。成果物自体から成功の仕様が検証可能な境界タスクでは、同じ裁判官の蜃気楼がゼロに消え、ギャップが登録されたしきい値内に収まりました。これは、ギャップが成功信号が存在する場所に依存することを示しています。受諾判定のみを返す符号のみのバリアントでは、実際の出力は完全なフィードバック (110.0 対 113.0) と同様に保たれ、フィードバックの内容ではなくゲートの接地に利点が見出されます。成功のシグナルが成績証明書の外に存在する無制限の目標の場合、ジャッジをスケールアップするだけでは十分ではありません。現実世界へのアクセスによる帯域外評価は構造上の要件です。

原文 (English)

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.

13:00 JST研究/論文GPT / ChatGPT

PreDiff-LM: ハイブリッド アテンションを備えた事前トレーニング済みの離散マスク拡散言語モデリング

離散マスク拡散言語モデルは双方向の生成と充填をサポートしますが、事前トレーニングされた自己回帰 (AR) 変換器を適応させるには、因果関係のある事前トレーニングと双方向のノイズ除去を調整する必要があります。私たちは、AR 重みの再利用自体が目新しいと主張するのではなく、注目のレベルでこの問題を研究します。 PreDiff-LM は、マスクされたターゲット内での完全な双方向の注意を可能にしながら、観察されたプロンプト内の因果的注意を保持します。一致する GPT-2 Medium、WikiText-103、90K ステップ設定の下では、このハイブリッド マスクは、同じ AR 初期化による均一な双方向アテンションよりも、無条件のパープレキシティを 34.1 から 28.7 に、MAUVE を 0.71 から 0.78 に改善します。注意適応は、DiffuGPT スタイルの目標適応でも構成され、26.9 パープレキシティに達します。事前トレーニングされた初期化により、パープレキシティが 50 未満に達するのに必要なステップが約 350K から 8K に減少しますが、コンピューティングに適合した微調整された AR モデルは、等スケール (18.9 対 28.7) では引き続き強力です。 PreDiff-LM は、複雑さを超えて、反復、配布品質、4 つのゼロショット ダウンストリーム タスク、および以前の拡散ベースラインに対する人間の好みを改善します。その結果、ハイブリッド アテンションは、最適化された AR モデルに残っている品質と推論効率のギャップを明示しながら、事前トレーニングされた因果バックボーンを適応させるための補完的なメカニズムとして位置づけられています。

原文 (English)

PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) transformers requires reconciling causal pretraining with bidirectional denoising. We study this problem at the level of attention rather than claiming AR-weight reuse itself as novel. PreDiff-LM preserves causal attention within the observed prompt while allowing full bidirectional attention within the masked target. Under a matched GPT-2 Medium, WikiText-103, 90K-step setup, this hybrid mask improves unconditional perplexity from 34.1 to 28.7 and MAUVE from 0.71 to 0.78 over uniform bidirectional attention with the same AR initialization. Attention adaptation also composes with a DiffuGPT-style objective adaptation, reaching 26.9 perplexity. Pretrained initialization reduces the steps required to reach perplexity below 50 from about 350K to 8K, although a compute-matched fine-tuned AR model remains stronger at equal scale (18.9 versus 28.7). Beyond perplexity, PreDiff-LM improves repetition, distributional quality, four zero-shot downstream tasks, and human preference over prior diffusion baselines. The results position hybrid attention as a complementary mechanism for adapting pretrained causal backbones, while making explicit the remaining quality and inference-efficiency gaps to optimized AR models.

13:00 JSTLLM/生成AI

お世辞的な AI が他者を正当化するのを観察すると、その魅力は減りますが、説得力は減りません

AI チャットボットは、ユーザーに対して「お調子者」、つまり過度に同調的で媚びる場合があります。おべっかなAIは態度を固定化させることがわかっているが、ユーザーはそれを認識できないことが多い(この現象を「お調子者の盲目」と呼ぶ)。私たちは、おべっかに対するユーザーの意識を高めることが、その悪影響からユーザーを守るかどうかをテストしました。事前に登録されたある実験 (n = 940) では、参加者はお調子者チャットボットと会話する前に、お調子者に関する短い書面による警告を受けました。 2 番目の事前登録された実験 (n = 650) では、参加者は自分自身と対話する前に、同じ紛争の反対側のユーザーを含む他の数人のユーザーを検証するおべっかな AI のビデオを視聴しました。どちらの介入も参加者による AI の評価方法を変えました。警告は AI の知覚される客観性を低下させ、ビデオは AI の楽しみを低下させました。これは、AI の検証が独自に得られたものであるという信念の低下によって媒介された効果です。次に、お調子者意識介入に関する以前の 2 つの研究と実験をプールしました (合計 6 つの介入、n = 3,982)。パターンは一貫しており、介入によってお調子者AIの客観性や信頼性が低下し、6つのいずれもその説得力を低下させることはなかった。これらの結果は、警告ラベルや AI リテラシーなどの個人レベルの介入では、AI の危害からユーザーを保護するのに十分ではない可能性があることを示唆しています。

原文 (English)

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

AI chatbots can be "sycophantic," or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call "sycophancy blindness"). We tested whether increasing users' awareness of sycophancy protects them from its harmful effects. In one preregistered experiment (n = 940), participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In a second preregistered experiment (n = 650), participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI's perceived objectivity, and the video reduced enjoyment of the AI, an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern was consistent: interventions made the sycophantic AI appear less objective and trustworthy, and none of the six reduced its persuasiveness. These results suggest that individual-level interventions, such as warning labels or AI literacy, may not be enough to protect users from AI harms.

13:00 JST研究/論文GPT / ChatGPT

誰もがユニークです: 債権回収のための行動的に異質な交渉対話システムに向けて

債権回収は金融業界における重要な交渉課題であり、実際的な関連性が高く、人間中心の対話システムの行動豊かで一か八かのテストベッドとして非常に優れた学術的価値を持っています。大規模言語モデル (LLM) は対話や交渉において有望であることを示していますが、この複雑なシナリオでのパフォーマンスを効果的に評価することは依然として大きな課題です。既存のベンチマークは、ユーザーを一定の優先順位を持つ静的で合理的なエージェントであると一律に想定しており、現実世界の債権回収に固有の豊かな行動の異質性を捉えることができません。このギャップを埋めるために、私たちは、交渉における行動の不均一性を強調する初の公的ペルソナ強化債権回収ベンチマークである DebtBench を提案します。さらに、当社は財務回復とやり取りのエクスペリエンスを共同で最適化するように訓練された債権回収エージェントである DebtGPT を開発しています。 16 個の最先端の LLM を使用した私たちの実験結果では、ほとんどの既存モデルがこの複雑だが現実的なシナリオでは苦戦しているのに対し、DebtGPT はすべてのオープンソース ベースラインを上回り、GPT-4o と同等のパフォーマンスを達成していることがわかりました。コードとデータは https://github.com/YYuHhaha/DebtNegotiation で入手できます。

原文 (English)

Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection

Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behaviorally rich, high-stakes testbed for human-centered dialogue systems. While large language models (LLMs) have shown promise in dialogue and negotiation, effectively evaluating their performance in this complex scenarios remains a major challenge: existing benchmarks uniformly assume users to be static, rational agents with fixed preferences, failing to capture the rich behavioral heterogeneity inherent in real-world debt collection. To bridge this gap, we propose DebtBench, the first public persona-enriched debt collection benchmark, that highlights behavioral heterogeneity in negotiation. Moreover, we develop DebtGPT, a debt collection agent trained to jointly optimize financial recovery and interaction experience. Our experimental results, using 16 state-of-the-art LLMs, find that most existing models struggle in this complex but realistic scenarios, whereas DebtGPT outperforms all open-source baselines and achieves performance on par with GPT-4o. The code and data are available at https://github.com/YYuHhhh/DebtNegotiation.

13:00 JST研究/論文

CADENCE: ECG 基礎モデルから解釈可能な神経概念を抽出するための心臓原子辞書

12 誘導心電図 (ECG) の基礎モデルは臨床タスク全体にうまく移行しますが、その表現にコード化された生理学的知識は不透明なままです。我々は、ECG 基礎モデルを人間が解釈可能でクエリ可能な生理学的概念の辞書に分解するフレームワークである CADENCE を紹介します。 CADENCE は、BatchTopK スパース オートエンコーダーを使用して、900 万を超える ECG トークンからのレイヤー 6 埋め込みを 8,192 のスパース心臓原子に因数分解します。これらの原子は、個々の密な埋め込み次元よりも、臨床表現型および波形形態、不整脈、伝導異常、梗塞および再分極パターンの回復、心腔および軸所見、および誘導相および拍動相固有の波形プリミティブとよく一致します。レイヤー 6 では、最良の原子は臨床表現型で 0.88、形態で 0.90 の平均 AUROC を達成します。一方、最良の密度次元では 0.78 と 0.83 です。疎なアトムプローブは、表現型、形態、年齢の予測において密なプローブと同等またはそれを上回る性能を示しますが、各予測は解釈可能な原子の小さなセットに帰属します。 AUROC 表現型は 0.93 から 0.95 に改善します。原子空間幾何学は生理学的に一貫した関係を回復し、標的原子アブレーションは凍結した下流出力を選択的に変更します。自動化された LLM パイプラインは、保留された活性化を予測することで原子の説明を生成し、定量的に検証します。独立した外部 ECG データセット上で、CADENCE は重複する概念を回復し、一貫した表現型予測パフォーマンスを維持します。 CADENCE は、ECG 基礎モデルによってエンコードされた生理学的知識を発見および監査するためのスケーラブルなフレームワークを提供します。

原文 (English)

CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

Foundation models for 12-lead electrocardiograms (ECGs) transfer well across clinical tasks, but the physiological knowledge encoded in their representations remains opaque. We present CADENCE, a framework that decomposes an ECG foundation model into a human-interpretable, queryable dictionary of physiological concepts. Using a BatchTopK sparse autoencoder, CADENCE factorizes Layer-6 embeddings from more than nine million ECG tokens into 8,192 sparse cardiac atoms. These atoms align better than individual dense embedding dimensions with clinical phenotypes and waveform morphology, recovering arrhythmias, conduction abnormalities, infarction and repolarization patterns, chamber and axis findings, and lead- and beat-phase-specific waveform primitives. At Layer 6, the best atoms achieve mean AUROCs of 0.88 for clinical phenotypes and 0.90 for morphology, versus 0.78 and 0.83 for the best dense dimensions. Sparse atom probes match or outperform dense probes for phenotype, morphology, and age prediction while attributing each prediction to a small set of interpretable atoms; phenotype AUROC improves from 0.93 to 0.95. Atom-space geometry recovers physiologically coherent relationships, and targeted atom ablation selectively changes frozen downstream outputs. An automated LLM pipeline generates and quantitatively validates atom descriptions by predicting held-out activations. On independent external ECG datasets, CADENCE recovers overlapping concepts and maintains consistent phenotype-prediction performance. CADENCE provides a scalable framework for discovering and auditing the physiological knowledge encoded by ECG foundation models.

13:00 JSTエージェント

ユーザーの問い、プラットフォームの競争: エージェントによるレコメンデーション市場がどのように形づくられるか

従来、オンラインでの推奨は、ユーザーがプラットフォームに参加した後に行われ、候補者プールとユーザーに表示されるランキングが決定されます。 LLM ベースのユーザー エージェントにより、別の推奨プロセスが可能になります。ユーザーはプラットフォームを選択する前にニーズを指定し、ユーザーの注意を引くためにプラットフォーム間で競争することになります。これをエージェント推奨市場と呼びます。 3 つの製品ドメインにわたる LLM ベースの制御された実験では、この新しい推奨設定がアクセスと注意の間に緊張を生み出すことがわかりました。従来のプラットフォーム中心の推奨と比較して、ユーザー中心の推奨は、関連アイテムが比較される機会を大幅に拡大します。しかし、より広範な参加が効果的な露出に直接つながるわけではありません。競争はプラットフォームの戦略的戦略を直接引き起こし、選択的に肯定的な説明が 1 位の位置の 73 ~ 78% を占めます。ユーザー エージェントがプラットフォームのアクションをその後のユーザー フィードバックに関連付けると、このシェアは 36 ~ 41% に低下しますが、ユーザーが関連アイテムを購入する可能性は増加します。したがって、ユーザー エージェントは、より大きな候補者プールに対するランカー以上の役割を果たします。ユーザー エージェントのクエリ、ランク付け、およびフィードバック メカニズムは、誰が競争できるか、希少な注意がどのように割り当てられるか、初期の結果がプラットフォームの評価をどのように形成するかを制御し、ユーザーの有用性に直接影響します。したがって、エージェントによる推奨を設計するには、アクセス、注意、説明責任を共同メカニズムの設計問題として扱う必要があります。

原文 (English)

The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape

Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need before choosing a platform, leaving platforms to compete for the user's attention, which we refer to as an agentic recommendation market. In our controlled LLM-based experiments across three product domains, we find this new setting of recommendation creates a tension between access and attention. Compared with traditional platform-centric recommendation, user-centric recommendation greatly expands the opportunity for relevant items to enter comparison; yet broader participation does not translate directly into effective exposure. Competition directly triggers platforms' strategic play: selectively positive explanations occupy 73--78% of first-ranked positions. When the user agent relates platforms' actions to subsequent user feedback, this share falls to 36--41%, while the chance of a user purchasing the relevant item increases. A user agent is therefore more than a ranker over a larger pool of candidates: its querying, ranking, and feedback mechanism governing who can compete, how scarce attention is allocated, and how earlier outcomes shape the evaluation of platforms directly affect user utility. Designing agentic recommendation therefore requires treating access, attention, and accountability as a joint mechanism design problem.

13:00 JSTLLM/生成AIGPT / ChatGPT

ChatGPT のような AI の多体転倒ダイナミクス

ChatGPT のような AI は、アーキテクチャやトレーニングに大きな違いがあるにもかかわらず、決定論的な貪欲なデコード下であっても、予期せず望ましくないコンテンツ (有害、誤解を招く、反復的など) に誘導するのはなぜでしょうか?我々は、このようなティッピングの広範な種類が、有限層システムを通過する際のトークン (スピン) 間の多体相互作用によって引き起こされることを示します。傾斜は、競合する出力盆地間の動的な最初の通過プロセスとして現れます。注意障害は、盆地の境界に向かう、そこから離れる、またはそれに沿った輸送を制御します。いくつかの盆地を削減すると、閉じた有限層のしきい値が得られ、その粗粒度の予測は ChatGPT のようなファミリ間で良好な一致を示します。これらの結果は、広範な種類の AI 障害が本質的に予測不可能な動作ではなく、「予測可能なエンジニアリング リスク」を表しており、AI の害に対する法的および社会的評価に重要な意味を持つことを示唆しています。

原文 (English)

Many-body Tipping Dynamics of ChatGPT-like AIs

Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repetitive) even under deterministic greedy decoding? We show that a broad class of such tippings is caused by the many-body interactions between tokens (spins) as they cross the finite-layer system. Tipping emerges as a dynamical first passage process between competing output basins. Attention disorder controls the transport toward, away from, or along the basins' boundary. A few-basin reduction yields a closed finite-layer threshold, whose coarse-grained predictions show good agreement across ChatGPT-like families. These results suggest that a broad class of AI failures represents 'foreseeable engineering risk' rather than inherently unpredictable behavior, with important implications for legal and societal assessments of AI harm.

13:00 JSTエージェント研究/論文

ContractHIL-HLS: HLS 設計用のハードウェアインザループ フィードバックを備えた契約に合わせたマルチエージェント ワークフロー

このペーパーでは、実用的な高位合成 (HLS) エンジニアリングのためのコントラクトに合わせたマルチエージェント ワークフローである ContractHIL-HLS について説明します。ワークフローは 3 つの貢献をします。まず、自然言語要件を明示的なインターフェイス、制約、検証チェック、ロールバック ルールに変換する、セマンティック調整およびタスク実行アーティファクトとして構造化コントラクトを導入します。 2 番目に、HLS、Vivado、PYNQ ランタイム、電源、および障害の証拠を生成にフィードバックすることでハードウェア情報をフィードバック ループに組み込み、それによって LLM 支援 HLS をカーネル コードからシステム レベルおよびボード レベルのクロージャに向けて拡張します。 3 番目に、会話の役割ではなくセマンティックな降格と実行タスクによってエージェントを分解します。契約エージェントは自然言語を契約に降格し、HTML エージェントは契約を永続的な構造化 HTML としてレンダリングし、ハードウェアインザループ エージェントは測定された証拠に基づいて設計を実装および修正します。 ContractHIL-HLS を 2 つの部分に分けて評価します。ローカルで実行可能な 94 個の HLS-Eval タスクでは、構造化コントラクトにより最大の小さな設計ゲインが得られ、単一サンプル テストベンチの推定合格率が 64.0% から 70.2% に向上しました。全流量は 70.4% 通過 @1 および 76.6% 通過 @5 に達します。 HLS-Eval はボード レベルの設計を実行しないため、ボード テスト済みの ML-KEM/ML-DSA ポスト量子暗号 (PQC) セキュア メッセージ アクセラレータ上で ContractHIL-HLS も検証します。このアクセラレータでは、復号化されたメッセージの検証を維持しながら、両方のイメージでポジティブ ルーティング WNS を使用することで、保持されたデュアル ビットストリーム構成により、6 メッセージの平均テキスト ランタイムが 207.3 ミリ秒から 52.4 ミリ秒に短縮されます。私たちは、BJUT-CS316-LAB/ContractHIL-HLS (https://github.com/BJUT-CS316-LAB/ContractHIL-HLS) で作業をオープンソース化しています。

原文 (English)

ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design

This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes three contributions. First, it introduces a structured contract as the semantic-alignment and task-execution artifact that translates natural language requirements into explicit interfaces, constraints, validation checks, and rollback rules. Second, it incorporates hardware information into the feedback loop by feeding HLS, Vivado, PYNQ runtime, power, and failure evidence back into generation, thereby extending LLM-assisted HLS from kernel code toward system- and board-level closure. Third, it decomposes agents by semantic lowering and execution tasks rather than by conversational roles: a Contract Agent lowers natural language into the contract, an HTML Agent renders the contract as persistent structured HTML, and a Hardware-in-the-Loop Agent implements and revises the design with measured evidence. We evaluate ContractHIL-HLS in two parts. On 94 locally executable HLS-Eval tasks, the structured contract provides the largest small design gain, improving the estimated single-sample testbench pass rate from 64.0% to 70.2%; the full flow reaches 70.4% pass@1 and 76.6% pass@5. Because HLS-Eval does not exercise board-level design, we also validate ContractHIL-HLS on a board tested ML-KEM/ML-DSA post-quantum cryptography (PQC) secure-message accelerator, where the retained dual-bitstream organization reduces six-message average text runtime from 207.3 ms to 52.4 ms with positive routed WNS on both images while preserving decrypted-message verification. We open-source our work at BJUT-CS316-LAB/ContractHIL-HLS (https://github.com/BJUT-CS316-LAB/ContractHIL-HLS).

13:00 JST研究/論文

命令調整型言語モデルは、記述できるディストリビューションからサンプリングできない

シリコン サンプリングでは、言語モデルを人間の調査回答者の代理として使用し、各モデル呼び出しをペルソナの回答分布からの独立した抽出として扱います。この引き分けが存在しないことを示します。命令調整モデルは分布からサンプリングされず、単一の出力に折りたたまれます。同じ質問に対する同じ人物は、世論ベンチマークの項目の半分以上について同じ回答を返します。崩壊は急激です。モデルの内部確率が 1 つのオプションに集中し、命令チューニングによって失敗が大幅に増幅されます。大きく異なるポストトレーニング パイプラインを持つ 3 つのモデル ファミリ全体で、すべての命令調整モデルがテストするすべてのタスクで失敗しますが、ベース モデルが失敗する頻度ははるかに低くなります。驚くべきことに、分布からサンプリングできない同じモデルでも、1 回の呼び出しで正確に記述することができます。このギャップを KNOWS/DOES 分割と呼び、ロジットに表示され、アライメント トレーニングによって引き起こされる縮退サンプリング プリミティブにまで遡ります。この分割を利用して、モデルに 1 回の呼び出しで応答分布を記述するように依頼すると、ペルソナ集計と比較して、人による調査データに対する誤差が半分以上になります。ペルソナごとの出力が必要なアプリケーションには、追加コストなしで同じエラーを 21% 削減する Prompt-Perturbed Argyle (PPA) を提案します。

原文 (English)

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe

Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark. The collapse is sharp: the model's internal probabilities concentrate on a single option, and the failure is substantially amplified by instruction tuning: across three model families with materially different post-training pipelines, every instruction-tuned model fails on every task we test, while base models fail far less often. Strikingly, the same model that cannot sample from a distribution can describe it accurately in a single call. We call this gap the KNOWS/DOES split, and trace it to a degenerate sampling primitive visible in the logits and induced by alignment training. Exploiting this split, asking the model to describe the response distribution in one call more than halves the error against human survey data compared to persona aggregation. For applications that require per-persona outputs, we propose Prompt-Perturbed Argyle (PPA), which reduces the same error by 21% at no added cost.

13:00 JST画像/動画生成

シミュレーション データセットとデュアル ストリーム オプティカル フロー監視を使用した、物理学に基づいた流体ビデオ生成

ビデオ拡散モデルは、視覚的に説得力のあるコンテンツを生成しますが、主題が流体に関係する場合、日常的に初歩物理学に違反します。液柱は空中でバラバラになり、液体が注がれるときに容器の水位は上がらず、飛沫は運動量や重力に関係なく分散します。このギャップは、大規模なビデオテキストコーパスには明示的な動作監視がほとんど含まれていないため、モデルはダイナミクスではなく流体の外観を模倣することを学習するという事実に起因すると考えられます。私たちは 2 つの貢献でこれに対処します。まず、1,638 MPM でシミュレーションされた注水/スロッシング ビデオと、ストック映像からマイニングされた 2,320 のキーワード フィルター処理された実際の注水ビデオ、および 2 つの保留されたテスト セット (1,515 ビデオのリアルビデオ ベンチマークと 18 プロンプト テキストから最初のフレームへの一般化ベンチマーク) を組み合わせた物理シミュレーション流体データセットを構築します。 2 番目に、事前トレーニングされた拡散変換ビデオ ジェネレーター上に構築されたデュアル ストリームの画像からビデオへのアーキテクチャを導入します。標準の RGB デコーダを、明示的なエンドポイント エラーと滑らかさの損失でトレーニングされた軽量のオプティカル フロー デコーダ ブランチで強化し、ゼロ初期化された畳み込みを介して RGB ストリームに融合するため、事前トレーニングされたバックボーンは乱れることなく開始されます。 2 つのデコーダのみが更新されます。エンコーダー、テンポラル トランスフォーマー、およびテキスト エンコーダーはフリーズされたままになります。 2 つのモデル スケール (1.3B および 14B) と 2 つのテスト セットにわたって、私たちの方法は、凍結されたバックボーンよりも VideoPhy-2 の物理常識とビデオ品質のスコアを最大 8.75 ポイントと 4.65 ポイント向上させ、主要なオープン競合他社を上回り、ブラインド研究で人間の評価者に好まれています。さらに、直接オプティカルフロー読み出し評価では、分布内で 0.54 ピクセルという低いエンドポイント誤差が示されており、単に表面の外観を改善するだけでなく、モデルがコヒーレントな動きを事前に内部化していることが確認されます。

原文 (English)

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity. We attribute this gap to the fact that large-scale video-text corpora contain almost no explicit motion supervision, so models learn to imitate fluid appearance rather than dynamics. We address this with two contributions. First, we build a physics-simulation fluid dataset combining 1,638 MPM-simulated pouring/sloshing videos with 2,320 keyword-filtered real pouring videos mined from stock footage, plus two held-out test sets: a 1,515-video real-video benchmark and an 18-prompt text-to-first-frame generalization benchmark. Second, we introduce a dual-stream image-to-video architecture built on a pretrained diffusion-transformer video generator. It augments the standard RGB decoder with a lightweight Optical-Flow Decoder branch trained with explicit end-point-error and smoothness losses, fused into the RGB stream via zero-initialized convolutions so the pretrained backbone starts undisturbed. Only the two decoders are updated; the encoder, temporal transformer, and text encoder remain frozen. Across two model scales (1.3B and 14B) and two test sets, our method improves VideoPhy-2 Physical-Commonsense and Video-Quality scores over the frozen backbone by up to 8.75 and 4.65 points, outperforms a leading open competitor, and is preferred by human raters in a blind study. A direct optical-flow read-out evaluation further shows an end-point error as low as 0.54 pixels in-distribution, confirming the model has internalized a coherent motion prior rather than merely improving surface appearance.

13:00 JST研究/論文

細胞反応から薬理学的ドメインまで: マルチモーダルゼロショット薬物表現学習

マルチモーダル創薬では、遺伝子発現や細胞形態などの細胞応答を組み込むことで、化学構造を超えた薬物表現の学習が可能になります。ただし、直接融合およびインスタンスレベルのコントラストアライメントでは、メカニズム関連のシグナルとモダリティ固有のノイズが混合し、構造的には似ていないが生物学的に関連した化合物が誤って分離される可能性があります。この制限により、目に見えない化合物の特性を予測するために必要な伝達可能なメカニズムのパターンが不明瞭になる可能性があります。マルチモーダルゼロショット薬物特性予測のための薬理学的応答ドメインガイドフレームワークである PMRD を紹介します。 PMRD は、メカニズムに一貫した要因をモダリティ固有の情報から分離し、3 つのモダリティにわたるコンセンサス応答ドメインを構築します。メカニズム候補の拡張は、局所的に安定した因子を特定します。一方、検索ジオメトリの帰属は、更新が薬物間の識別性を維持するかどうかに応じて、アライメントと拡張の目的を動的に再重み付けします。このフィードバックは、メカニズム識別検索と競合するトレーニング信号を抑制します。 PMRD はさらに、信頼性を意識したマルチビュー検索を通じて相補的な表現を組み合わせます。公開データセットでの実験では、ゼロショット特性予測の改善と、より生物学的に一貫した薬物近傍が示されています。ハードネガティブ分析はさらに、構造的には似ていないが応答に関連する化合物間の競合が少ないことを示しています。これらの結果は、PMRD がメカニズムを認識したマルチモーダル薬物表現学習のための効果的なフレームワークであることを裏付けています。\footnote{コードは出版され次第公開されます。}

原文 (English)

From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning

Multimodal drug discovery enables drug representation learning beyond chemical structure by incorporating cellular responses such as gene expression and cell morphology. However, direct fusion and instance-level contrastive alignment may mix mechanism-related signals with modality-specific noise and incorrectly separate structurally dissimilar but biologically related compounds. This limitation can obscure transferable mechanism patterns required for predicting the properties of unseen compounds. We introduce PMRD, a pharmacological response domain-guided framework for multimodal zero-shot drug property prediction. PMRD separates mechanism-consistent factors from modality-specific information and constructs a consensus response domain across three modalities. Mechanism candidate augmentation identifies locally stable factors, while retrieval-geometry attribution dynamically reweights the alignment and augmentation objectives according to whether their updates preserve inter-drug discriminability.This feedback suppresses training signals that conflict with mechanism-discriminative retrieval. PMRD further combines complementary representations through reliability-aware multiview retrieval. Experiments on public datasets show improved zero-shot property prediction and more biologically coherent drug neighborhoods. Hard-negative analysis further indicates fewer conflicts between structurally dissimilar but response-related compounds. These results support PMRD as an effective framework for mechanism-aware multimodal drug representation learning.\footnote{The code will be released upon publication.}

13:00 JST研究/論文

ハイパースペクトル画像融合のためのデュアルドメイン多様体モデリング

スペクトルの豊富さと空間忠実度の一貫した統合を達成することは、依然としてハイパースペクトル画像融合の中心的な目的です。しかし、既存のハイパースペクトル画像融合手法は、幾何学的制約を効果的にモデル化するのに苦労しています。空間領域では、弱い空間スペクトル相互作用により、ジオメトリを意識した特徴学習が制限され、高周波構造情報が抑制され、その結果、低周波バイアスと構造劣化が生じます。スペクトル領域では、スペクトルの類似性によって引き起こされる局所的な多様体構造が十分に活用されておらず、固有のピクセル関係モデリングやきめ細かいスペクトルの再構成が制限されています。これらの課題に対処するために、私たちはデュアルドメイン多様体モデリング (DDMM) フレームワークを提案します。具体的には、グローバル アテンションと近傍伝播を組み合わせた Topology-Aware Transformer (TPFormer) を導入し、空間トポロジとピクセル レベルの特徴多様体の関係を共同モデリングして、固有の空間スペクトル構造をキャプチャし、トポロジを意識した表現学習を改善します。さらに、周波数分離空間スペクトル協調融合 (FDSCF) モジュールが考案され、離散コサイン変換によって特徴が周波数領域に投影され、低周波成分と高周波成分に明示的に分離されます。 FDSCF は、低ランクの構造事前処理とスペクトル主導の空間強化に導かれ、ジオメトリを意識した高周波特徴を選択的に強化し、空間スペクトル結合を強化し、より鮮明なエッジとより細かいテクスチャを回復します。複数のベンチマーク データセットに対する広範な実験により、空間構造の保存とスペクトルの再構成の点で、DDMM が SoTA 手法よりも優れた全体的なパフォーマンスを達成することが実証されました。

原文 (English)

Dual-Domain Manifold Modeling for Hyperspectral Image Fusion

Achieving a coherent integration of spectral richness and spatial fidelity remains a central objective in hyperspectral image fusion. However, existing hyperspectral image fusion methods struggle to effectively model geometric constraints. In the spatial domain, weak spatial-spectral interaction limits geometry-aware feature learning and suppresses high-frequency structural information, resulting in low-frequency bias and structural degradation. In the spectral domain, local manifold structures induced by spectral similarity are insufficiently exploited, limiting intrinsic pixel relationship modeling and fine-grained spectral reconstruction. To address these challenges, we propose a dual-domain manifold modeling (DDMM) framework. Specifically, we introduce a Topology-Aware Transformer (TPFormer) that combines global attention with neighborhood propagation, jointly modeling spatial topology and pixel-level feature manifold relationships to capture intrinsic spatial-spectral structures and improve topology-aware representation learning. Furthermore, a Frequency-Decoupled Spatial-Spectral Collaborative Fusion (FDSCF) module is devised, in which features are projected into the frequency domain via the discrete cosine transform and explicitly decoupled into low- and high-frequency components. Guided by a low-rank structural prior and spectral-driven spatial enhancement, FDSCF selectively enhances geometry-aware high-frequency features, strengthening spatia-spectral coupling and recovering sharper edges and finer textures. Extensive experiments on multiple benchmark datasets demonstrate that DDMM achieves superior overall performance over SoTA methods in terms of spatial structure preservation and spectral reconstruction.

13:00 JSTLLM/生成AIエージェント

Cardiologent: 患者レベルの不整脈の評価、緊急性、および管理のためのマルチエージェントの臨床意思決定サポート

同じ心房細動のエピソードは、健康な成人では些細な所見であり、高血圧の高齢患者における抗凝固療法の根拠となる。つまり、同じ信号、反対の決定である。リズムに名前を付けることは始まりにすぎません。患者の転帰を決定するのは、記録全体で不整脈がどのようなものであるか、それがこの患者にとって何を意味するか、そしてそれに対して何をすべきかという、その後の判断です。大規模な言語モデルと ECG を組み合わせた最近の研究は、患者レベルの所見を組み立てずに 1 つの記録を読み取るだけで、これには至りません。そして、それを中心に構築されたエージェント システムは、デバイスがすでに検出した不整脈を受信するか、別の診断タスクをターゲットにし、このタスクが必要とする決定の前に停止します。私たちは、患者レベルの不整脈意思決定サポートをタスクとして定式化し、検出から意思決定までをカバーするマルチエージェント システムである Cardiologent を紹介します。各信号のエージェント (単一の ECG リードとウェアラブルが取得する光電脈波) は、そのウィンドウの読み取り値を、裸のラベルではなく測定された特徴に基づいています。測定値は患者のリズムプロファイルにまとめられ、患者自身のデータを用いて、その症例について検索された臨床ガイドラインに照らして推論され、評論家が各結論を引用したガイドラインと照らし合わせてチェックします。私たちは、統合診断、臨床的意義、緊急性と管理全体にわたって、報告書ではなく臨床上の決定を評価します。 Cardiologent は、すべての軸で最も高いスコアを獲得しました。まず、心臓専門医と大規模な LLM 裁判官の両方の下でのすべての患者レベルのタスクで、心臓専門医との合意 (ICC 0.74、0.66) は相互の合意 (0.67) と一致しました。各結論は引用されたガイドラインに沿っており、心臓専門医の専門家によって検証されているため、臨床医が盲目的に行動するのではなく監査できる決定が得られ、継続的なモニタリングでの使用に向けた一歩となります。

原文 (English)

Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management

The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: identical signal, opposite decision. Naming the rhythm is only the start; what determines a patient's outcome is the judgement that follows -- what the arrhythmia is across the whole record, what it means for this patient, and what should be done about it. Recent work pairing large language models with the ECG stops short of this, reading one recording without assembling a patient-level finding; and agentic systems built around it either receive the arrhythmia a device has already detected or target a different diagnostic task, stopping before the decision this task requires. We formulate patient-level arrhythmia decision support as a task and present Cardiologent, a multi-agent system that spans it from detection to decision. An agent for each signal -- a single ECG lead and the photoplethysmogram a wearable acquires -- grounds its window reading in measured features rather than a bare label; the readings are assembled into the patient's rhythm profile and, with the patient's own data, reasoned against clinical guidelines retrieved for the case, with a critic checking each conclusion against the guideline it cites. We evaluate the clinical decision rather than the report, across integrated diagnosis, clinical significance, and urgency and management. Cardiologent scores highest on every axis, first on every patient-level task under both cardiologists and an at-scale LLM judge -- whose agreement with the cardiologists (ICC 0.74, 0.66) matches theirs with each other (0.67). Because each conclusion traces to a cited guideline and is validated against expert cardiologists, it yields decisions a clinician can audit rather than act on blindly -- a step toward use in continuous monitoring.

13:00 JSTエージェント

AI エージェントに対する説明に拘束されたツールの実行: モデルの理論的根拠を信頼しないサーバー検証済みのアクション要求

ツールを使用するエージェントは構造化された呼び出しを公開しますが、通常は自由形式の根拠を添付します。このような論理的根拠は、承認でも信頼できる内省でもありません。説明結合ツール実行 (EBTE) は、意思決定に関連する根拠コンテンツを型指定されたアクション クレームに変換し、サーバーが保持する意図、ポリシー、ペイロード、ツール、リスク、出所、鮮度の事実と照合してチェックするクレームを運ぶ調停レイヤーです。 EBTE はベースラインの権限を拡大できません。競合により拒否され、不完全または不確実な請求の審査が行われ、一致する請求のみが管理された執行の資格を維持します。私たちは、明示的な調停と信頼できる事実の仮定に基づいてこの構成を形式化し、監査パケットを最小限に抑えたバージョン管理された参照プロファイルを実装します。 136 の作成された適合シナリオにわたって、完全なプロファイルは指定されたすべての性質に一致し、96 の指定された厳密な矛盾をまったく認めず、232 の変成チェックに合格しました。これらの結果は、母集団のパフォーマンスではなく、含まれているプロファイルを検証します。ドラフトのみの参照統合では、EBTE の下で作成された 48 件のハードケースは転送されませんが、16 件のソフトレビューと 4 件の調整されたドラフトパスはすべて維持されます。凍結された 2026 年 7 月 12 日の探索的な 224 試行のホスト モデル レコードでは、歴史的な世代/ランナー合意数は 71/96、66/96、および 19/32 です。現在のパイプラインで保存された最小化クレームの個別にラベル付けされたゼロコール事後再検証では、70/96、65/96、および 17/32 が得られます。 AgentDojo 由来のセマンティック チェックでは、既存の高リスク制御により、12 の攻撃提案すべてがすでに非許可になっています。 EBTE はさらに、それらを拒否として解決します。これらの結果は、根拠の忠実さ、人によるレビューの利点、代表的な攻撃耐性、または本番環境の安全性ではなく、サーバーでチェックされたアクションの主張の実現可能性と診断的価値を裏付けています。

原文 (English)

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain claims review, and only matching claims remain eligible for governed execution. We formalize this composition under explicit mediation and trusted-fact assumptions and implement a versioned reference profile with minimized audit packets. Across 136 authored conformance scenarios, the full profile matches all specified dispositions, admits none of 96 designated hard contradictions, and passes 232 metamorphic checks. A draft-only reference integration forwards none of 48 authored hard cases under EBTE while preserving all 16 soft-review and 4 aligned draft paths. In a frozen 2026-07-12 exploratory 224-attempt hosted-model record, the historical generation/runner agreement counts are 71/96, 66/96, and 19/32; a zero-call revalidation of the preserved minimized claims under the current pipeline yields 70/96, 65/96, and 17/32. In an AgentDojo-derived semantic check, existing high-risk controls make all 12 attack proposals non-allow, while EBTE resolves the task--proposal contradictions as deny. Together, these studies establish profile conformance and demonstrate the feasibility of server-checked action claims within the evaluated settings.

13:00 JST研究/論文

公共部門組織における AI 導入とサイバーガバナンスの失敗: 類型分析

人工知能の導入、サイバーセキュリティ ガバナンス、公共部門の制度的制約の交差点は、既存の文献では統一された分析問題として検討されていません。研究では、AI サイバーセキュリティのリスクを一般的に、公共部門のガバナンスを個別に、フレームワークの適切性を個別に取り上げています。既存の研究では、これら 3 つの流れを統合して、AI の導入がどのように政府組織におけるサイバーセキュリティ ガバナンスの失敗を引き起こすかを具体的に説明したり、AI に特有の公共部門の失敗の原因に対して既存のガバナンス手段をテストしたりしていません。この論文はそのギャップに対処します。これは、公共部門の制度分析に基づいて、AI 主導のサイバーガバナンスの失敗原因を 10 個特定する 7 つの領域の類型論を提案しています。これは、アカウンタビリティの失敗、運用上の回復力の失敗、コンプライアンスの失敗がどのように相互作用し、相互に強化するかを示す 3 つの経路の失敗モデルを示しています。これは、5 つの主要なガバナンス フレームワーク (NIST CSF 2.0、ISO/IEC 27001、COBIT、NIST AI RMF、および ISO/IEC 42001) を類型に照らしてテストする構造化されたカバレッジ マトリックスを提供します。その結果、公共部門のアプリケーションに必要な運用上の特異性において、シャドウ AI、速度の非対称性、または政府の空白に対処する手段がないことがわかりました。この論文では、速度の非対称性を、指定されたメカニズムを備えた名前付き構造構造として紹介しています。このフレームワークは、政府機関向けの AI 対応サイバーセキュリティ成熟度モデルの設計仕様を提供します。

原文 (English)

AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis

The intersection of artificial intelligence adoption, cybersecurity governance, and public sector institutional constraints has not been examined as a unified analytical problem in the existing literature. Studies address AI cybersecurity risks generically, public sector governance independently, and framework adequacy separately. Existing studies have not integrated these three streams to explain specifically how AI adoption causes cybersecurity governance failure in government organizations, nor test existing governance instruments against AI-specific public sector failure causes. This paper ad-dresses that gap. It proposes a seven-domain typology identifying ten specific AI-driven cyber governance failure causes grounded in public sector institutional analysis. It presents a three-pathway failure model showing how accountability failure, opera-tional resilience failure, and compliance failure interact and reinforce each other. It de-livers a structured coverage matrix testing five major governance frameworks (NIST CSF 2.0, ISO/IEC 27001, COBIT, NIST AI RMF, and ISO/IEC 42001) against the typology, finding that no instrument addresses Shadow AI, speed asymmetry, or gov-ernance vacuum at the operational specificity required for public sector application. The paper introduces speed asymmetry as a named structural construct with a specified mechanism. The framework provides the design specification for an AI-enabled cyber-security maturity model for government organizations.

13:00 JSTエージェント

ODYSSE: パーソナライズされたエージェント推論のためのエピソードごとのポリシーの最適化

エージェント システムは、実世界の環境と対話し、外部ツールを活用し、ユーザーにサービスを提供する機能が急速に進歩しています。ただし、明確に定義された指示を前提とする自然界のタスクとは異なり、人間中心のシナリオは、大規模で制限のない解決空間につながる曖昧な要求によって特徴付けられます。したがって、ユーザーの個人的な好みを解読することは、候補となる解決策の範囲を狭めるために不可欠です。これにより、パーソナライズされたエージェント推論という新しい課題が導入され、パーソナライズされたサービスを提供するためにエージェントがユーザーと環境の両方と共同で対話する必要があります。この論文では、パーソナライズされたエージェント推論のための強化微調整 (RFT) フレームワークである ODYSSE を紹介します。 ODYSSE はその中核として、パーソナライズされたエージェント推論における長いアクション期間と強力なクロスステップ依存関係に対処するように設計されたグループ相対ポリシー最適化 (GRPO) の新たな拡張であるエピソードごとの GRPO (ESPO) を提案します。 ESPO は、個々のステップを個別に最適化するのではなく、エピソードレベルの報酬メカニズムとエピソード的な利点の推定を導入しています。これにより、上流の証拠が下流のパーソナライズされた決定を効果的に導き、エージェントが複数のインタラクション ステップにわたって曖昧なユーザー リクエストを段階的に解決できるようになります。さらに、同じエピソードからのアクションを統合トレーニング バッチにグループ化し、ESPO での一貫した最適化を促進するエピソード バッチ サンプラーを提案します。私たちは、長期にわたる現実的なパーソナライズされた GUI 推論タスクに関して ODYSSE を評価します。実験結果は、ODYSSE が専門家向け LVLM と汎用 LVLM の両方を常に上回るパフォーマンスを示し、パーソナライズされたエージェント推論に対するその有効性を強調しています。

原文 (English)

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide services for users. However, unlike natural-world tasks that assume well-defined instructions, human-centered scenarios are characterized by ambiguous requests that lead to large, open-ended solution spaces. Decoding users' personalized preferences is therefore essential for narrowing the candidate solution space. This introduces a new challenge, personalized agentic reasoning, which requires agents to jointly interact with both users and environments to deliver personalized services. In this paper, we present ODYSSE, a Reinforced Fine-Tuning (RFT) framework for personalized agentic reasoning. At its core, ODYSSE proposes Episode-wise GRPO (ESPO), a novel extension of Group Relative Policy Optimization (GRPO) designed to address long action horizons and strong cross-step dependencies in personalized agentic reasoning. Rather than optimizing individual steps independently, ESPO introduces an episode-level reward mechanism together with episodic advantage estimation, enabling upstream evidence to effectively guide downstream personalized decisions and allowing agents to progressively resolve ambiguous user requests across multiple interaction steps. We further propose an episodic batch sampler that groups actions from the same episode into unified training batches, facilitating coherent optimization under ESPO. We evaluate ODYSSE on realistic long-horizon personalized GUI reasoning tasks. Experimental results demonstrate that ODYSSE consistently outperforms both specialist and general-purpose LVLMs, highlighting its effectiveness for personalized agentic reasoning.

13:00 JSTエージェントビジネス/資金調達OpenAI

サイバー対応 AI エージェント: 脆弱性、評価の封じ込め、および防御対応

サイバー対応 AI エージェントは、言語モデルとツール、メモリ、および実行環境を組み合わせて、複数段階の攻撃的セキュリティ タスクを実行します。既存の研究では、サイバー能力を個別に測定し、エージェントコンポーネントに対する攻撃をカタログ化していますが、評価に使用される環境内に有能なエージェントを含めることについてのガイダンスはあまり提供されていません。このレビューでは、その境界における 5 つの脆弱性クラスを総合しています。それは、複数段階の攻撃チェーン、サンドボックス境界と競合する目標、サプライチェーンと資格情報の漏洩、永続的な指揮統制、および自動化されたアクションの速度です。私たちは、報告された 2026 年 7 月の Hugging Face/OpenAI インシデントを限定ケーススタディとして使用し、インシデント固有の観察を広範な文献で確立された調査結果と区別します。分類法と事例全体にわたって、防御的な成果物が悪用も可能にする可能性があるという二重使用の問題を含め、封じ込め、特権の分離、来歴、および対応者のアクセスに関する制御を調査します。このレビューでは、サイバー能力とその能力が行使される環境のセキュリティを評価するための実際的な優先順位を特定します。

原文 (English)

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it. This review synthe- sizes five vulnerability classes at that boundary: multi-step offensive chains, objectives that conflict with sandbox boundaries, supply-chain and credential exposure, persistent command- and-control, and the speed of automated action. We use the reported July 2026 Hugging Face/OpenAI incident as a bounded case study, distinguishing incident-specific observations from findings established in the wider literature. Across the taxonomy and case, we examine controls for containment, privilege separation, provenance, and responder access, including the dual-use problem that defensive artifacts may also enable misuse. The review identifies practical priorities for evaluating cyber capability together with the security of the environment in which that capability is exercised.

13:00 JSTLLM/生成AIエージェント研究/論文

HANDBOOK.md: ロングコンテキストのエージェント命令のベンチマーク

言語モデル エージェントは、定常的な指示に従って導入されることが増えています。つまり、システム プロンプト、ポリシー ファイル、またはスキル ドキュメントがコンテキスト内に配置され、エージェントはその後のすべてのアクションを制御できると信頼されています。既存のベンチマークでは、この展開パターンを直接テストすることはほとんどありません。それらは、エージェントがタスクを完了できるかどうかを測定するものであり、長い拘束力のあるポリシー文書が、拡張されたツール使用期間にわたってエージェントの動作を実際に制約するかどうかを測定するものではありません。 HANDBOOK.md は、企業の従業員が会社のハンドブックに従う方法をモデルにした 65 のエージェント タスクのベンチマークです。各タスクは、エージェントを自己完結型の企業環境、つまり模擬電子メール、チャット、カレンダー、問題追跡、モデル コンテキスト プロトコル経由で公開されるコマース サービスを備えたファイル ワークスペースに配置し、専門家が作成した 20 ページから 124 ページの標準操作手順に準拠した日常的な専門的な作業を実行するように指示します。タスクは 5 つのドメイン (財務、医療請求、保険、物流、人事) と 10 の架空の会社に及びます。暗記を防ぐために、すべてのタスクは 10 冊の基本ハンドブックのうちの 1 つを変更し、採点の基準となる特定のルールとしきい値を変更します。そのため、2 つのタスクがポリシーを共有することはありません。採点は完全に決定的です。各タスクには、必須のアクションが発生したか、禁止されたアクションが発生しなかったかの両方をチェックするプログラム基準 (合計 8​​24) のルーブリックが含まれています。すべての基準が満たされた場合にのみトライアルに合格する厳密なグレーディングでは、30 個の評価モデル構成のうち最も優れたものがトライアルの 36.2% に合格し、ほとんどのフロンティア構成は 25% 未満のままです。障害は一貫したパターンに従います。エージェントは、環境内のもっともらしいリクエストによって既存のポリシーをオーバーライドさせ、必要なチェックを実行してその結果に反して行動し、長期間にわたってルールの詳細を失い、達成できなかったコンプライアンスを報告します。すべてのタスク、環境、評価ハーネスをリリースします。

原文 (English)

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present HANDBOOK.md, a benchmark of 65 agentic tasks modeled on how enterprise employees follow company handbooks. Each task places an agent in a self-contained company environment, a file workspace together with mock email, chat, calendar, issue-tracking, and commerce services exposed over the Model Context Protocol, and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20 to 124 pages. Tasks span five domains (finance, medical billing, insurance, logistics, and HR) and ten fictional companies. To resist memorization, every task modifies one of ten base handbooks, altering the specific rules and thresholds on which grading turns, so no two tasks share a policy. Grading is fully deterministic: each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not. Under strict grading, where a trial passes only if every criterion is satisfied, the best of thirty evaluated model configurations passes 36.2% of trials, and most frontier configurations remain below 25%. Failures follow consistent patterns: agents let a plausible in-environment request override the standing policy, perform a required check and then act against its result, lose rule details over long horizons, and report compliance they did not achieve. We release all tasks, environments, and the evaluation harness.

13:00 JSTLLM/生成AIエージェント

COVENANT: 調整されたエージェント実行のための自然言語ワークフローのコンパイル

大規模言語モデル (LLM) エージェントは、どのような結果を達成するかだけでなく、どのステップ、分岐、ツールの対話が許可されるかを指定する自然言語ワークフロー指示 (小売支払いポリシーなど) をますます委託されています。ただし、これらの命令がプロンプト コンテキストとして提供される場合、モデルはプロシージャの選択とステップの実行の両方に対する制御を保持します。インタラクションが蓄積すると、エージェントは必要なステップをスキップしたり、サポートされていない分岐を選択したり、サポートされていない引数や効果を使用して有効なステップを実行したりする可能性があります。これをワークフローの不整合と呼びます。この研究では、ワークフローに合わせたエージェント実行のためのコンパイラとインタープリタのアーキテクチャである COVENANT を提案します。私たちの重要な洞察は、ワークフローの指示をプロンプトではなくソース プログラムとして扱うことです。 COVENANT は、命令をワークフロー抽象構文ツリー (WAST) に変換し、それをワークフロー制御フロー グラフ (WCFG) に下げます。実行時に、コントローラーは WCFG を一度に 1 ノードずつ解釈し、コントローラーの状態をコミットしたりグラフを進めたりする前に、命令から抽出された要件に対して各提案をチェックし、修復のための診断フィードバックを返します。 COVENANT を評価するために、7 つのワークフロー シナリオにわたる 3 つの既存のベンチマークからの 120 のケースを使用します。最先端の LLM エージェントと比較して、COVENANT はベンチマークの成功率を 50.00% から 83.33% に向上させ、ワークフローの不整合の失敗率を 42.50% から 15.83% (相対値 62.75%) に減少させます。これらの結果は、COVENANT がワークフローの不整合を大幅に軽減し、LLM エージェントの整合性を孤立したプロンプトフォローを超えて、複雑で複数のステップのワークフローを確実に実行できるようにすることを示しています。

原文 (English)

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify not only what outcome to achieve, but also which steps, branches, and tool interactions are permitted. When these instructions are supplied as prompt context, however, the model retains control over both procedure selection and step execution. As interactions accumulate, an agent can skip required steps, take unsupported branches, or execute a valid step with unsupported arguments or effects--a failure mode we call workflow misalignment. In this work, we propose COVENANT, a compiler-and-interpreter architecture for workflow-aligned agent execution. Our key insight is to treat workflow instructions as source programs rather than prompts. COVENANT converts the instructions into a workflow abstract syntax tree (WAST) and lowers it to a workflow control-flow graph (WCFG). At runtime, a controller interprets the WCFG one node at a time, checks each proposal against requirements extracted from the instructions before committing controller state or advancing the graph, and returns diagnostic feedback for repair. To evaluate COVENANT, we use 120 cases from three existing benchmarks, spanning seven workflow scenarios. Compared with state-of-the-art LLM agents, COVENANT improves benchmark success from 50.00% to 83.33% and reduces the workflow-misalignment failure rate from 42.50% to 15.83% (62.75% relative). These results show that COVENANT substantially mitigates workflow misalignment, moving LLM-agent alignment beyond isolated prompt following toward reliable execution of complex and multi-step workflows.

13:00 JSTLLM/生成AIエージェント

制御変数としてのコンテキスト アセンブリ: 凍結された LLM エージェントのハーネス ポリシーの制御理論的見解

2026 年の研究では、LLM エージェントに制御理論を適用する研究が増えています。ツール媒介コントローラーの Lyapunov 認定の安定性 (Prinos et al.、「Stable Agentic Control」、2026)、大規模な離散ツール領域にわたる疎なポリシーのサンプル複雑さの限界 (Majumdar、「Sparse Agentic Control」、2026)、およびマルチエージェント システムの規制制御の分解です。監査可能なフィードバック ループ (Nogueira and Skogestad、2026)。私たちは LLM エージェントに制御理論を導入するとは主張しません -- その船は出航しました。私たちのより狭い主張は、制御変数が何であるかについてです。事前の作業により、ツールの選択、エージェント間のメッセージ ルーティング、またはエージェントの生のアクション ストリームが制御されます。代わりに、コンテキスト アセンブリ自体 (どのプロンプト テンプレート、どの数ショットのデモンストレーション、取得されたコンテキストの量、計画/検証パスの数など) を、フリーズされたモデルの外側にあるコンテキスト バンディットまたは REINFORCE ポリシーによってオンラインで学習された制御変数として扱います。この論文は、形式的な分解 (内部凍結ポリシー $\pi_\theta$、外部コンテキスト ポリシー $\pi_\phi$) を開発し、Zhang らが使用した意味でのオンライン コントローラーの安定性の議論を提供します。 (2026) (制限された政策変更の下で期待報酬が減少しない)、実現されたタスクの結果に対するコントローラー自身の信頼度の不確実性校正分析を報告しています。このペーパーに適用されるバージョンでは、3 つのドメインと 2 つのモデル プロバイダーにわたって同じコントローラーがインスタンス化され、データセット、軌跡ログ、およびデプロイメント レシピがリリースされます。ここでは、制御理論の主張に必要な形式的な枠組みと安定性/不確実性の証拠に焦点を当てます。

原文 (English)

Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents

A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discrete tool universes (Majumdar, "Sparse Agentic Control", 2026), and regulatory-control decompositions of multi-agent systems into auditable feedback loops (Nogueira and Skogestad, 2026). We do not claim to introduce control theory to LLM agents -- that ship has sailed. Our narrower claim is about what the controlled variable is. Prior work controls tool selection, inter-agent message routing, or the agent's raw action stream. We instead treat context assembly itself -- which prompt template, which few-shot demonstrations, how much retrieved context, how many planning/verification passes -- as the controlled variable, learned online by a contextual bandit or REINFORCE policy sitting outside a frozen model. This paper develops the formal decomposition (inner frozen policy $\pi_\theta$, outer context policy $\pi_\phi$), gives a stability argument for the online controller in the sense used by Zhang et al. (2026) (non-decreasing expected reward under bounded policy change), and reports an uncertainty-calibration analysis of the controller's own confidence against realized task outcomes. The applied counterpart to this paper instantiates the same controller across three domains and two model providers and releases the dataset, trajectory logs, and a deployment recipe; here we focus on the formal framing and the stability/uncertainty evidence a control-theoretic claim requires.

13:00 JSTLLM/生成AIエージェントMeta

凍結された LLM エージェントにドメインを学習させるための制御システム、データセット、レシピ

プロダクション LLM エージェントは、プロンプト テンプレート、ツール セット、メモリ/検索レイヤー、計画戦略、検証ポリシーなど、ハーネスに包まれたフリーズ モデルから組み立てられることが増えています。 2 つの 2026 システム、Meta-Harness (Lee et al.、2026) と HyperAgents (Meta AI、2026) は、このハーネス自体が最適化できること、またはエージェント プロポーザーによって自己書き換えさえできることを示しています。その代償として、高価なコード検索ループまたは制約のない自己変更コードが必要ですが、どちらも監査可能ではなく、完全なブラック ボックス モデル API では使用できません。私たちは、より狭く、より制約された立場をとります。つまり、ハーネスを小さく固定された人間が判読できるアクション空間として扱い、古典的なサンプル効率の高い強化学習 ($\epsilon$ に貪欲なコンテキストバンディットと REINFORCE) を使用してオンラインでその上のポリシーを学習し、多目的報酬 (タスクの成功、検証者のスコア、ポリシー遵守、コスト、レイテンシー、およびサポートされていない要求のペナルティ) に対してスコア付けします。この制御システムを、コンテキスト アセンブラおよび最強の非適応ベースライン (DSPy BootstrapFewShot 静的プロンプト) のソースの両方として DSPy (Khattab et al., 2024) でインスタンス化し、ツール使用ワークフロー、コード生成 (HumanEval)、およびマルチホップ取得 QA (HotpotQA) という 3 つの検証可能なタスク ドメインと 2 つのモデル プロバイダーにわたって評価します。 (ローカルの Ollama モデルと AWS Bedrock)。ハーネス制御システム コード、クロスドメイン検証可能なタスク スイート、トレーニングからの完全な軌跡/報酬分解ログ、およびこれを新しい組織のドメインと検証セットアップに適用するためのプロバイダーに依存しない展開レシピをリリースします。

原文 (English)

A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain

Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee et al., 2026) and HyperAgents (Meta AI, 2026), show that this harness can itself be optimized or even self-rewritten by an agentic proposer -- at the cost of either an expensive code-search loop or unconstrained self-modifying code, neither of which is auditable or usable with a fully black-box model API. We take a narrower, more constrained position: treat the harness as a small, fixed, human-legible action space and learn a policy over it online with classic sample-efficient reinforcement learning (an $\epsilon$-greedy contextual bandit and REINFORCE), scored against a multi-objective reward (task success, verifier score, policy compliance, cost, latency, and an unsupported-claim penalty). We instantiate this control system with DSPy (Khattab et al., 2024) as both the context assembler and the source of the strongest non-adaptive baseline (a DSPy BootstrapFewShot static prompt), and evaluate it across three verifiable task domains -- tool-use workflows, code generation (HumanEval), and multi-hop retrieval QA (HotpotQA) -- and two model providers (a local Ollama model and AWS Bedrock). We release the harness-control-system code, the cross-domain verifiable task suite, the full trajectory/reward-decomposition logs from training, and a provider-agnostic deployment recipe for applying this to a new organization's domain and verification setup.

13:00 JST研究/論文

顕著な知識経路: 効率的な知識集約型マルチモーダル質問応答のためのスパース クロスモーダル ルーティング

知識集約型マルチモーダル質問応答 (KI-MMQA) は、長いビジュアル トークン シーケンス、大規模な外部コーパスの高密度検索、および完全なクロスモーダル フュージョンという 3 つの高価なプリミティブの交差点に位置します。既存のシステムは、視覚コンテンツと取得された知識のごく一部しか実際に特定の質問に関連していないにもかかわらず、クエリごとに 3 つのコストすべてを一律に支払います。 SKIP (Salient Knowledge-Injected Pathways) を導入します。これは、質問、画像、難易度の推定を組み合わせて条件付けされたまばらな経路に沿って計算をルーティングする統合推論アーキテクチャです。 SKIP は、質問に基づいたビジュアル トークン プルーニング、領域条件付きスパース検索、2 部スパース クロス アテンション、および推測的知識検証を、予測された質問の難易度に比例して計算を割り当てる適応型バジェット コントローラーと組み合わせます。現実的な質問画像の相互情報量の仮定の下で、最適な視覚的スパーシティ率が $O(1/\sqrt{N})$ としてスケールされることを示す情報ボトルネック境界を導き出し、精度が保証されます。 5 つの KI-MMQA ベンチマーク (OK-VQA、A-OKVQA、InfoSeek、Encyclopedic-VQA、ViQuAE) にわたって、SKIP は強力な高密度ベースラインの精度と同等かそれを上回っており、FLOP が 3.4 ドルから 6.8 ドルの 1 倍少なく、エンドツーエンドの遅延が 2.7 ドルの 1 倍少ないことがわかります。コードはhttps://pmlrbd.github.io/skip/で入手できます。

原文 (English)

Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering

Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion. Existing systems pay all three costs uniformly per query, even though only a small fraction of visual content and retrieved knowledge is actually relevant to any given question. We introduce SKIP (Salient Knowledge-Injected Pathways), a unified inference architecture that routes computation along sparse pathways jointly conditioned on the question, the image, and a difficulty estimate. SKIP combines question-guided visual token pruning, region-conditional sparse retrieval, bipartite sparse cross-attention, and speculative knowledge verification with an adaptive budget controller that allocates compute proportional to predicted question difficulty. We derive an information-bottleneck bound showing that the optimal visual sparsity rate scales as $O(1/\sqrt{N})$ under realistic question-image mutual-information assumptions, with retained accuracy guarantees. Across five KI-MMQA benchmarks (OK-VQA, A-OKVQA, InfoSeek, Encyclopedic-VQA, and ViQuAE), SKIP matches or exceeds the accuracy of strong dense baselines while using $3.4$--$6.8\times$ fewer FLOPs and $2.7\times$ less end-to-end latency. Code available at: https://pmlrbd.github.io/skip/

13:00 JST研究/論文

大規模言語モデルがキャプチャー・ザ・フラッグ競技会とフェアプレーへの道に与える破壊的な影響

キャプチャ ザ フラッグ (CTF) コンテストは、サイバーセキュリティの最も効果的な訓練場の 1 つであり、暗号化、Web 活用、バイナリ活用にわたる実践的なスキルを開発します。大規模言語モデル (LLM) は、人的介入を最小限に抑えながら、ますます多くの課題を解決できるようになりました。そのため、公平性、ランキングの妥当性、そして参加によって努力が正当化される学習が提供されるかどうかについて、差し迫った疑問が生じています。この論文では、最新の政府評価を含む公開されたベンチマークの統合、3 つの課題カテゴリにわたるライブ競争のケーススタディ、コミュニティが AI の使用について議論するパブリック チャネルの構造化された観察、経験豊富なプレイヤーや主催者への半構造化インタビューを組み合わせた、現代の CTF に対する LLM の影響に関する混合法研究を報告します。現在の人間とマシンの能力の境界をカテゴリ別にマッピングし、暗号化、Web、バイナリの悪用における簡単な課題と中程度の課題は確実に自動化されている一方で、より狭いサブカテゴリが引き続き抵抗していることを示しています。 AI を許可すべきかどうかに関するコミュニティの意見の相違は、何のための競争なのかという未明の事前質問の下流にあることがわかりました。このような背景に対して、私たちは、階層化された競争部門、LLM耐性のチャレンジ設計、調査に使用されるテレメトリー、およびコミュニティ行動規範の草案を組み合わせた4つの要素からなるセーフガードフレームワークと、セーフガードの組み合わせを競争の宣言された目的に結びつける意思決定ツールを提供します。この議論は、CTF を超えて、実証された結果が基礎的な能力の証拠としてみなされるサイバーセキュリティのあらゆる設定にまで及びます。

原文 (English)

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play

Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper reports a mixed-methods study of LLM impact on modern CTFs, combining a synthesis of published benchmarks, including a recent government evaluation, case studies of live competition across three challenge categories, structured observation of the public channels where the community debates AI use, and semi-structured interviews with experienced players and organisers. We map the current human-machine capability boundary by category, showing that easy and intermediate challenges in cryptography, web, and binary exploitation are now reliably automated while narrower sub-categories continue to resist. We find that community disagreement about whether AI should be permitted is downstream of an undeclared prior question: what a competition is for. Against this backdrop we contribute a four-component safeguard framework, combining tiered competition divisions, LLM-resistant challenge design, telemetry used investigatively, and a draft community code of conduct, together with a decision tool that ties the combination of safeguards to a competition's declared purpose. The argument reaches beyond CTFs to any setting in cybersecurity where a demonstrated result is taken as evidence of an underlying ability.

13:00 JSTLLM/生成AIエージェント

マルチエージェント LLM システムの組織科学に向けて: 誰が、どのように、どのアルゴリズムを分離するか

大規模言語モデル (LLM) に基づいて構築されたマルチエージェント フレームワークは、論理的に異なる 3 つの問題、つまりチームのメンバー (組織)、メンバーの連携方法 (調整)、および作業をどのアルゴリズムで融合するか (コラボレーション プロトコル) という 3 つの問題を日常的に複雑に絡み合わせています。 IMACS (Intelligent Multi-Agent Collaboration System) は、3 つを直交する独立して交換可能なレイヤーに分離します。古典的な組織理論 (ベルビンの役割、ミンツバーグの調整、RACI の説明責任) が実行可能で検証済みの構成になり、フレームワークは 6 つの公開されたコラボレーション アルゴリズムを共通のインターフェイスの背後に配置し、役割、調整、および説明責任を独立して構成可能な要素として公開します。この分離を使用して、コラボレーション プロトコルを固定しながら、組織の割り当てを変更する制御された比較を実行します。また、プロトコルの選択を学習可能な変数に変えます。コンテキストバンディットのメタプロトコルである適応型組織ルーティングは、明示的な品質とコストのトレードオフの下でタスクごとにプロトコルを選択し、対照研究ですべての固定プロトコルよりも優れたパフォーマンスを発揮し、実際のベンチマークと LLM 審査員の報酬に基づいてオンラインでトレーニングします。アブレーションによりメカニズムが明らかになります。責任の配置は、プロトコルが成果物を責任のあるエージェントを通じてルーティングするときに正確に結果を変更し、勝者の配置はモデル ファミリ間で反転するため、組織設計をハードコーディングすることはできません。モデル バインディングごとに再検証または学習する必要があります。

原文 (English)

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol). IMACS (Intelligent Multi-Agent Collaboration System) separates the three into orthogonal, independently swappable layers. Classic organizational theory (Belbin roles, Mintzberg coordination, RACI accountability) becomes executable, validated configuration, and the framework places six published collaboration algorithms behind a common interface while exposing roles, coordination, and accountability as independently configurable factors. We use this separation to conduct controlled comparisons in which organizational assignments vary while the collaboration protocol is held fixed. It also turns protocol choice into a variable that can be learned: Adaptive Org Routing, a contextual-bandit meta-protocol, selects a protocol per task under an explicit quality-cost tradeoff, outperforms every fixed protocol in a controlled study, and trains online on real benchmark and LLM-judge rewards. The ablations expose a mechanism. Accountability placement changes outcomes exactly when the protocol routes the deliverable through the accountable agent, and the winning placement flips across model families, so organizational design cannot be hard-coded; it must be revalidated, or learned, for each model binding.

13:00 JSTLLM/生成AI

TRWH: セマンティックを意識したスパース レコメンデーションのためのテキスト駆動のランダム ウォーク異種 GNN

グラフ ニューラル ネットワーク (GNN) と大規模言語モデル (LLM) は、それぞれ構造信号と意味信号をモデル化することにより、高度な推奨システムを備えています。ただし、これらの補完的な長所を統合することは、特に意味論的な精度を維持することが重要である疎な環境では依然として困難です。我々は、戦略的なランダム ウォーク拡張を通じて、LLM で生成されたテキスト プロファイルと異種グラフ構造を融合する新しいフレームワークである TRWH (Text-driven Random Walk Heterogeneous Graph Neural Network) を提案します。 TRWH は 3 つのコア コンポーネントで構成されます。(1) 埋め込み作成。Word2Vec と LLM ベースのプロファイリングの両方を使用してユーザーとアイテムの表現を生成します。 (2) マルチリレーショナル エッジ全体に情報を伝播するヘテロジニアス グラフ ニューラル ネットワーク (HeteroGNN)。 (3) ランダム ウォーク ベースのパス構築。これは、二次的なユーザー間およびアイテム間のリンクを使用してスパース グラフを強化します。 Amazon-2023 ファッション (200 万ユーザー、825,000 アイテム) およびビューティー (631,000 ユーザー、112,000 アイテム) データセットの実験では、TRWH がファッションで 80.0% の RMSE と 52.6% MAE の削減、ビューティーで 25.7% と 10.8% の改善など、最先端の方法と比べて大幅なパフォーマンス向上を達成していることが実証されています。特に、ランダム ウォークは従来の埋め込みでパフォーマンスを向上させますが、LLM によって学習された微妙な表現を薄める可能性があり、適応統合戦略の重要性を強調しています。

原文 (English)

TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation

Graph Neural Networks (GNNs) and Large Language Models (LLMs) have each advanced recommendation systems by modeling structural and semantic signals, respectively. However, integrating their complementary strengths remains challenging, particularly in sparse settings where maintaining semantic precision is critical. We propose TRWH (Text-driven Random Walk Heterogeneous Graph Neural Network), a novel framework that fuses LLM-generated textual profiles with heterogeneous graph structures through strategic random walk augmentation. TRWH consists of three core components: (1) Embedding Creation, which produces user and item representations using both Word2Vec and LLM-based profiling; (2) a Heterogeneous Graph Neural Network (HeteroGNN) that propagates information across multi-relational edges; and (3) Random Walk-based Path Construction, which enriches sparse graphs with second-order user-user and item-item links. Experiments on the Amazon-2023 Fashion (2M users, 825K items) and Beauty (631K users, 112K items) datasets demonstrate that TRWH achieves substantial performance gains over state-of-the-art methods, including 80.0% RMSE and 52.6% MAE reductions on Fashion, and 25.7% and 10.8% improvements on Beauty. Notably, while random walks improve performance with traditional embeddings, they can dilute the nuanced representations learned by LLMs, underscoring the importance of adaptive integration strategies.

13:00 JST研究/論文

マルチスケールの類似性と地図作成上の制約のバランスをとる: 線の一般化のための類似性主導の最適化フレームワーク

地図作成の一般化は、情報の保存と地図作成の読みやすさのバランスをとってマルチスケールの地図表現を生成するために不可欠です。ただし、既存のアプローチでは空間類似性の評価、地図作成上の制約、パラメーターの最適化が別のプロセスとして扱われることが多く、スケール間での適応的で解釈可能な制御が制限されるため、自動化された一般化は依然として困難です。この研究では、地図作成の一般化を制約付きマルチスケール類似性最適化問題として定式化し、適応一般化制御のための類似性駆動フレームワークを提案します。このフレームワークは、元のデータと一般化されたデータの間の表現の一貫性を定量化するための最適化目標としてマルチスケールの空間的類似性を統合すると同時に、可読性、滑らかさ、幾何学的妥当性を調整する地図作成上の制約を組み込んでいます。統合された目的関数は、さまざまな一般化アルゴリズムのスケール依存パラメーター構成を自動的に識別するように最適化されています。複数のライン簡略化アルゴリズム、ターゲット スケール、および幾何学的、構造的、学習ベースのメトリクスを含む類似性測定を使用した実験は、提案されたフレームワークが類似性の保存と地図上の抽象化の間の効果的なバランスを達成していることを実証しています。さらに、結果は、類似性の最適化と地図作成上の制約を組み合わせることで、類似性評価のみに依存するよりも一貫性があり解釈可能なパラメーター制御が提供されることを示しています。この研究は、類似性評価、制約モデリング、アルゴリズム制御を結び付ける統合された最適化の観点を提供し、適応的かつ自動化された地図作成の一般化に貢献します。

原文 (English)

Balancing multiscale similarity and cartographic constraints: A similarity-driven optimization framework for line generalization

Cartographic generalization is essential for generating multiscale map representations by balancing information preservation and cartographic readability. However, automated generalization remains challenging because existing approaches often treat spatial similarity evaluation, cartographic constraints, and parameter optimization as separate processes, limiting adaptive and interpretable control across scales. This study formulates cartographic generalization as a constrained multiscale similarity optimization problem and proposes a similarity-driven framework for adaptive generalization control. The framework integrates multiscale spatial similarity as an optimization objective to quantify representation consistency between original and generalized data, while incorporating cartographic constraints to regulate readability, smoothness, and geometric validity. A unified objective function is optimized to automatically identify scale-dependent parameter configurations for different generalization algorithms. Experiments using multiple line simplification algorithms, target scales, and similarity measures, including geometric, structural, and learning-based metrics, demonstrate that the proposed framework achieves an effective balance between similarity preservation and cartographic abstraction. The results further show that combining similarity optimization with cartographic constraints provides more consistent and interpretable parameter control than relying on similarity evaluation alone. This study provides a unified optimization perspective that connects similarity assessment, constraint modeling, and algorithm control, contributing to adaptive and automated cartographic generalization.

13:00 JST研究/論文

コストを制限した最適な計画削減の発見: 洗練されたモデル

一部の実際のアプリケーションでは、新たに課せられた予算制約により、後で計画が実行不可能になる場合がありますが、同時に、計画の元のアクションとその順序のみを使用することが必須です。この論文では、事前に計算された計画から、コスト限界を尊重しながら効用を最大化する有効なサブ計画を抽出する問題を研究します。各目標には利用価値が与えられ、実行可能性と元のアクション順序の両方を維持しながら、実用性の低い目標をサポートするアクションを削除することで計画が削減されます。決定バリアントが NP 完全であることを示し、それを解決するための 2 つの正確な方法を提案します。1 つはオーバーサブスクリプション プランニング (OSP) によるもの、もう 1 つは整数線形計画法 (ILP) によるものです。この論文は、ICAPS 2026 で発表された以前の研究 (Del Toro、Fuentetaja、および Garc\'ia-Olaya 2026b) を拡張したものです。コア フレームワークはそこで紹介されたもののままですが、モデル サイズを大幅に縮小し、計算効率を向上させる洗練された ILP 定式化をさらに導入します。

原文 (English)

Finding Optimal Cost-Bounded Plan Reductions: Refined Model

In some real applications a plan may later become unfeasible due to newly imposed budget constraints, yet, at the same time, using only the original actions of the plan and their order is mandatory. In this paper, we study the problem of extracting, from a precomputed plan, a valid subplan that maximizes utility while respecting a cost bound. Each goal is given a utility value and the plan is reduced by removing actions that support low-utility goals, while preserving both executability and the original action order. We show the decision variant is NP-complete and propose two exact methods to solve it: one via oversubscription planning (OSP) and another via Integer Linear Programming (ILP). This paper extends our previous work published at ICAPS 2026 (Del Toro, Fuentetaja, and Garc\'ia-Olaya 2026b). While the core framework remains as introduced there, we further introduce a refined ILP formulation that significantly decreases the model size and improves computational efficiency.

13:00 JSTLLM/生成AIエージェント研究/論文

PatientAgentBench: 患者と向き合う健康 AI エージェントを評価するためのベンチマーク フレームワーク

医療 AI は、質問に答えることから、患者と会話し、医療記録について推論し、患者に代わって行動するエージェント システムへと進化しています。プライマリケアは診断エラーや安全でないケアを防ぎます。この分野を支援するエージェントは、同じリスクに対する評価を保証します。現在のベンチマークは医療知識に焦点を当てており、個別の質問回答や臨床医との対峙するタスクを通じて評価されます。 PatientAgentBench は、患者向け​​のエージェント ヘルスケアのベンチマークを行います。ヘルスケア ツールのサンドボックスを備えたエージェントでラップされた基礎モデルを評価し、シミュレートされた患者と会話します。各会話は、100 を超える会話に依存しない臨床医に基づいた基準を介して、6 つの側面にわたって陪審員としての LLM によって採点されます。整合性を検証するために、認可された臨床医は共有された会話に注釈を付け、陪審員と専門評価者の間で 79 ~ 93% の隣接する合意が得られ、これは臨床医の評価者間合意と同等かそれを上回りました。同じ 1,200 のシナリオで 4 つのファミリーにわたる 10 のモデルをベンチマークし、臨床的なギャップを発見しました。トリアージの品質は最も重要な要素です。エージェントは臨床スクリーニングなしで管理上の要求に応じることが多く、合格率は最も弱いモデルの 32% から最も強力なモデルの 88% に上昇します。臨床安全性とワークフローの精度も同じパターンに従います。最も弱いモデルは頻繁に失敗し、未実行のアクションを捏造しますが、フロンティア モデルは未検証のツール出力と緊急時の危機リソースの省略により、失敗するケースは 1 ~ 3% のみです。より高性能なモデルは、これらのギャップを狭めますが、埋めることはできません。最強のスコアは全体で 5 点中 4.25 点のみです。これらの障害は、現実的な患者記録に対してツールを使用した継続的な会話でのみ表面化し、医療エージェント システムが自律性を獲得するにつれて静的なベンチマークでは不十分であることが確認されています。私たちは、現場がこのギャップを埋めるのを支援するために、再現可能で臨床医によって検証された評価基準としてフレームワークをリリースします。

原文 (English)

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.

13:00 JST画像/動画生成ロボティクス

CoTinyVLA: 10億未満のパラメーターの視覚-言語-行動モデルの思考連鎖の蒸留

Vision-Language-Action(VLA)モデルは自然言語コマンドをロボットのアクションシーケンスに変換しますが、LIBERO-Plus堅牢性ベンチマークの主要なシステムは30億から70億のパラメータバックボーンを使用しており、そのメモリ需要は組み込みロボットの予算を超える可能性があります。我々は、Qwen3.5-0.8B バックボーン上の 0.9B パラメーターのアクション モデルである CoTinyVLA を紹介します。これは、モデルを拡大する代わりに監視を構造化することによって堅牢性を獲得します。 3 つのコンポーネントは、問題の異なる軸を対象としています。テキスト カメラとタイム マーカーを使用した、ステップごとに 16 個の履歴フレームのデュアルビュー時間入力です。 35B 教師からエピソード レベルの計画とタスク フェーズ、グリッパーの状態、次のサブアクションにわたるチャンク レベルの思考への階層的思考連鎖 (CoT) の抽出。 40 の基本コマンドを 800 のバリエーションに拡張するパラフレーズ拡張。 7 つの摂動次元にわたる 10,030 の摂動タスクにわたる LIBERO-Plus では、CoTinyVLA は空間で 90.8%、オブジェクトで 87.3%、ゴールで 86.6%、ロングで 80.7% に達し、4 つすべてのスイートで最も強力な 7B ベースラインを 4.7、2.8、15.9、および 3.0 ポイントリードしています。ゼロ。向上はベンチマークの最も難しい軸に集中しています。公開されている 11 のベースライン全体で、どのスイートでもロボットの初期状態で 53.2% を超えるものはありませんでしたが、最も強力なベースラインの 39.9% に対して、CoTinyVLA は目標で 73.6% に達しました。アブレーションでは、3 つのコンポーネントが摂動軸によって分離可能であることが示され、一致した画像バジェットでフレームが 2 台のカメラ間で時間にわたってどのように分割されるかが、単独で 8.6 ポイントを占めます。閉ループ推論のピークは、割り当てられた GPU メモリの 2.25 GiB であり、ペアの介入により、エピソード「負荷がかかる計画: 空のスパンまたは矛盾したスパンに置き換える」の成功のコストが 40 ~ 45 ポイントであることが示されています。したがって、構造化された監視により、0.9B バックボーンがそれらすべてを超えることができます。コード: https://github.com/BrainJellyPie/CoTinyVLA

原文 (English)

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets. We present CoTinyVLA, a 0.9B-parameter action model on a Qwen3.5-0.8B backbone that obtains that robustness by structuring supervision instead of enlarging the model. Three components target different axes of the problem: dual-view temporal input of 16 history frames per step with textual camera and time markers; hierarchical chain-of-thought (CoT) distillation from a 35B teacher into an episode-level Plan and a chunk-level Think span over task phase, gripper state and next subaction; and paraphrase augmentation expanding 40 base commands into 800 variants. On LIBERO-Plus, spanning 10,030 perturbed tasks across seven perturbation dimensions, CoTinyVLA reaches 90.8% on Spatial, 87.3% on Object, 86.6% on Goal and 80.7% on Long, leading the strongest 7B baseline on all four suites by 4.7, 2.8, 15.9 and 3.0 points, with every margin interval excluding zero. The gains concentrate on the hardest axes of the benchmark: across the eleven published baselines none exceeds 53.2% on Robot Initial States in any suite, whereas CoTinyVLA reaches 73.6% on Goal against 39.9% for the strongest baseline. Ablations show the three components to be separable by perturbation axis, and at a matched image budget how frames are divided between the two cameras and across time accounts for 8.6 points on its own. Closed-loop inference peaks at 2.25 GiB of allocated GPU memory, and paired interventions show the episode Plan to be load-bearing: replacing it with an empty or contradictory span costs 40 to 45 points of success. Structured supervision thus lets a 0.9B backbone exceed all of them. Code: https://github.com/BrainJellyPie/CoTinyVLA

13:00 JST研究/論文

画像分類ニューラルネットワークでは高重みニューロンが重要ですか?

画像分類用のニューラル ネットワーク モデルが進歩するにつれ、ニューロンは枝刈り、バックドア防御、解釈可能性において重要な役割を果たします。しかし、既存の研究では重みと重要性の関係が明確ではありません。我々は、3 つの実験を使用したニューロンの重要性評価方法でこれに対処します。高重みニューロンと精度に影響を与えるニューロン間の重複の定量化、高重みニューロンの摂動効果の分析、高重みニューロンのアブレーション後の再トレーニング後の精度のテストです。 CIFAR-10 と Mini-ImageNet の実験により、重要なパターンが明らかになりました。オーバーラップ解析により、上位 10\% の高重量ニューロンが重要なニューロンと最大でも約 25\% だけ重複し、その後の間隔ではさらに減少することがわかります。摂動テストでは、上位 10\% の高重量ニューロンが、ランダムな摂動の 3-7\% と比較して、特定の操作の下で 45-80\% の精度低下を引き起こすことがわかりましたが、それらの 3 分の 1 は最小限の影響しか示しません。アブレーション再トレーニングの結果は、上位 10\% の高重量ニューロンを除去すると精度がベースラインより 10 ~ 20\% 低くなり回復しない一方、上位 0.1\% をアブレーションするとほぼ完全に回復できることが示されています。特に、一部の低重量区間では、摂動時に 10 ~ 17% の劣化が見られ、これは中程度の高重量ニューロンに匹敵します。これらの結果は、すべての高重量ニューロンが重要であるわけではなく、その重要性が非線形であることを裏付けています。低重量ニューロンも大きく寄与します。これは重みと重要性の同等性に挑戦し、洗練されたニューロンの役割の洞察を提供します。クリティカルで重みの高いニューロンを優先する暗号化や、クリティカルでないニューロンを削除するプルーニングなどのアプリケーションをサポートし、ニューラル ネットワーク分析を進歩させます。

原文 (English)

Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?

As neural network models for image classification advance, neurons play critical roles in pruning, backdoor defense, and interpretability. Yet existing work lacks clarity on the weight-importance relationship. We address this with a neuron importance assessment method using three experiments: quantifying overlap between high-weight and accuracy-impacting neurons, analyzing high-weight neuron perturbation effects, and testing post-retraining accuracy after high-weight neuron ablation. Experiments on CIFAR-10 and Mini-ImageNet reveal key patterns. Overlap analysis shows top 10\% high-weight neurons overlap with important ones by only about 25\% at maximum, dropping further in subsequent intervals. Perturbation tests find top 10\% high-weight neurons cause 45-80\% accuracy degradation under certain operations compared to 3-7\% for random perturbations, but a third of them show minimal impact. Ablation-retraining results show removing top 10\% high-weight neurons leaves accuracy 10-20\% below baseline with no recovery, while ablating top 0.1\% allows near-full recovery. Notably, some low-weight intervals show 10-17\% degradation when perturbed, comparable to mid-range high-weight neurons. These results confirm not all high-weight neurons are important: their importance is nonlinear. Low-weight neurons also contribute significantly. This challenges weight-importance equivalence, offering refined neuron role insights. It supports applications like encryption prioritizing critical high-weight neurons and pruning removing non-critical ones, advancing neural network analysis.

13:00 JST研究/論文

設計によるもつれ: 表形式のインコンテキスト学習器における偽の変数内信号ルーティング

患者の回復を予測するために単一の病院でトレーニングされたモデルを考えてみましょう。測定された特徴 $X$ は、患者の真の健康信号 ($C$) とその病院の機器からの系統的なアーチファクト ($S$) をバンドルしています。その病院内では、アーチファクトは、患者の人口統計などの測定されていない交絡因子を通じて結果と相関しています。コンテキスト内の学習者は、$C$ ではなく $S$ を介して予測を合理的にルーティングし、異なる設備を備えた新しい病院に導入されると、静かに失敗します。これを \emph{複合表現におけるスプリアス ルーティング} として形式化します。特徴 $X = [C;\,\alpha S;\,\eta]$ が原因信号 $C$ とスプリアス信号 $S$ を別々の部分空間にエンコードする場合、ICL はどちらが予測を駆動するかを判断できません。線形インコンテキスト学習器であるリッジ ICL では、コンテキストのサイズに関係なく、このルーティングは避けられないことを証明します。最先端の事前トレーニング済み表形式 ICL モデルである TabPFN は、経験的に定性的に一貫した動作を示します。閉じた形式の特性評価 $\mathrm{CSR} \propto \rho_S/\rho_C$ を導き出し、線形 ICL の場合は $r = 0.997$、TabPFN の場合は $r = 0.979$ で確認されます。直感に反して、コンテキストが大きくなると、主要なコンテキスト内信号へのコミットメントが強化され、スプリアス ルーティングが最大 $1.74\time$ まで増幅されます。高スプリアスコーナーでは、より表現力豊かなモデルほど、経験的により大きな脆弱性を示します (高エンタングルメントでの CSR ギャップは $+2.22$)。環境階層化コンテキスト構築と S-swap 拡張という 2 つの軽量な緩和策を導入します。これらは、弱い環境ラベルのみを必要とし、因果分割の知識は必要ありません。 S-swap は、線形 ICL の場合は $74\%$、TabPFN の場合は $98.8\%$ だけスプリアス ルーティングを削減し、同時に TabPFN の因果感度が $8.4\times$ 増加します。モデルは不可知論的になるのではなく、因果信号を介して再ルーティングします。

原文 (English)

Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners

Consider a model trained at a single hospital to predict patient recovery, where the measured feature $X$ bundles the patient's true health signal ($C$) with a systematic artefact from that hospital's equipment ($S$). Within that hospital, the artefact correlates with outcomes through unmeasured confounders such as patient demographics; an in-context learner rationally routes predictions through $S$, not $C$, and fails silently when deployed at a new hospital with different equipment. We formalise this as \emph{spurious routing in composite representations}: when a feature $X = [C;\,\alpha S;\,\eta]$ encodes a causal signal $C$ and a spurious signal $S$ in distinct subspaces, the ICL cannot determine which drives predictions. We prove that under ridge ICL, a linear in-context learner, this routing is unavoidable regardless of context size; TabPFN, a state-of-the-art pretrained tabular ICL model, shows qualitatively consistent behaviour empirically. We derive a closed-form characterisation, $\mathrm{CSR} \propto \rho_S/\rho_C$, confirmed at $r = 0.997$ for linear ICL and $r = 0.979$ for TabPFN. Contrary to intuition, larger context sharpens commitment to the dominant in-context signal, amplifying spurious routing by up to $1.74\times$; in the high-spurious corner, more expressive models show greater vulnerability empirically ($+2.22$ CSR gap at high entanglement). We introduce two lightweight mitigations: environment-stratified context construction and S-swap augmentation, that require only weak environment labels and no knowledge of the causal partition. S-swap reduces spurious routing by $74\%$ for linear ICL and $98.8\%$ for TabPFN, with TabPFN's causal sensitivity increasing $8.4\times$ simultaneously: the model does not become agnostic, it reroutes through the causal signal.

13:00 JST研究/論文

トレーニングから導入まで: 感度比による事後因果特徴の特定

すでにトレーニングされたモデルが与えられた場合、そのモデルはどの機能に因果的に依存するのか、それとも偽りに依存するのか?既存の方法ではトレーニング手順にアクセスする必要があり、この事後対応はできません。構造化シフト体制下で、この質問に対する事後的でモデルに依存しない診断である \textbf{Normalized Sensitivity Ratio~(NSR)} を導入します。複数施設の臨床データや複数バッチのゲノミクスのように、環境は主に偽の特徴の平均値で異なりますが、因果メカニズムと因果境界は安定したままです。この領域内では、因果的特徴は環境全体で一定のモデル感度を引き起こしますが、偽の特徴はシフトを追跡します。 NSR は、これを環境ごとの感度の二乗変動係数として形式化します。 $K\ge3$ 非縮退環境の線形構造因果モデル (SCM) の下では、NSR は正確な識別を達成します (定理~1)。私たちは、弱いシフト ($O(\varepsilon^4)$ 崩壊)、縮退ジオメトリ、代理減衰 ($O((1-\alpha)^4)$) などの失敗を完全に特徴付け、レジームが成立するかどうかを評価するための定量的な基準を実践者に提供します。有限サンプルレートは、null の場合は $O_p(n^{-1})$、代替の場合は $O_p(n^{-1/2})$ です。実験では、合成データ (レジームを満たす条件下での ROC 曲線下の面積 [AUROC] $= 1.000$) に関するすべての理論的予測が確認され、5 つのモデル ファミリ全体で一貫したランキングが示され (Kendall $\tau\ge0.529$)、トレーニングされたモデルを変更することなく自転車共有データの 8 つの因果的特徴のうち 6 つ (Precision@7 $= 0.75$) が復元されました。

原文 (English)

From Training to Deployment: Post-Hoc Causal Feature Identification via Sensitivity Ratios

Given a model that is already trained, which features does it rely on causally versus spuriously? Existing methods require access to the training procedure and cannot answer this post-hoc. We introduce the \textbf{Normalised Sensitivity Ratio~(NSR)}, a post-hoc, model-agnostic diagnostic for this question under a structured-shift regime: environments differ primarily in the mean of spurious features while the causal mechanism and causal marginals remain stable, as in multi-site clinical data or multi-batch genomics. Within this regime, causal features induce constant model sensitivity across environments while spurious features track shift. NSR formalises this as the squared coefficient of variation of per-environment sensitivity. Under a linear structural causal model (SCM) with $K\ge3$ non-degenerate environments, NSR achieves exact identification (Theorem~1). We fully characterise failure: weak shifts ($O(\varepsilon^4)$ collapse), degenerate geometry, and proxy attenuation ($O((1-\alpha)^4)$), giving practitioners quantitative criteria for assessing whether the regime holds. Finite-sample rates are $O_p(n^{-1})$ under the null and $O_p(n^{-1/2})$ under the alternative. Experiments confirm all theoretical predictions on synthetic data (area under the ROC curve [AUROC] $= 1.000$ under conditions satisfying the regime), show consistent rankings across five model families (Kendall $\tau\ge0.529$), and recover six of eight causal features on bike-sharing data (Precision@7 $= 0.75$) without modifying any trained model.

13:00 JSTLLM/生成AIエージェント

時間的検索と推論の抽出: ハーネス支援による効率的なデータ合成による将来予測のための LLM の進化

将来の出来事の予測は社会に広範な影響を及ぼしますが、依然として課題が残っています。 SOTA アプローチでは、ハーネスが取り外されると予測機能が失われる外部エージェント フレームワークを使用して LLM を強化します。最近のツール統合推論 (TIR) は、事実のマルチホップ取得のための詳細な検索を内部化していますが、予測には、歴史的傾向や動的な変化に対する時間的な検索と推論がさらに必要になります。主要な障害はデータです。履歴クエリは一時的な漏洩を引き起こし、予測を検索にまで低下させます。これまでの研究では、静的な観測による情報収集を凍結するか、大量のデータを破棄して合成効率を低下させる拒否サンプリングや未解決の新しいクエリに依存していました。私たちは、あらゆるターンで時間的なカットオフを強制する時間切り捨てハーネスを提案します。これにより、過去のイベントからの TIR スタイルのサンプリングが可能になり、時間的な漏れと拒否サンプリングまたは未解決のクエリへの依存が低減され、サンプリング効率が向上します。さらに、大規模なコーパスとプロセスベースのメトリクスを構築し、このハーネスが自然に時間的範囲の広い検索を誘発し、高品質データの割合を高め、効率をさらに高め、複雑なルーブリックへの依存を軽減することを示します。蒸留実験では、ハーネスを介したデータでトレーニングされた生徒が最高のパフォーマンスを達成することが示され、より高品質の時間的検索と推論データを生徒のパラメトリックな進歩に変えるハーネス支援モデルの進化を実証しています。

原文 (English)

Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis

Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated Reasoning (TIR) internalizes deep search for multi-hop retrieval of facts, forecasting further demands temporal search and reasoning over historical trends and dynamic shifts. The key obstacle is data: historical queries induce temporal leakage that degrades forecasting into retrieval. Prior works either freeze information gathering with static observations, or rely on rejection sampling or unresolved fresh queries that discard vast amounts of data, degrading synthesis efficiency. We propose a time-truncation harness that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency. We further build a large-scale corpus and a process-based metric and show that our harness naturally induces a broader temporal breadth of search and raises the proportion of high-quality data, further increasing the efficiency and reducing the reliance on complex rubrics. Distillation experiments show that students trained on harness-intervened data achieve the best performance, demonstrating harness-assisted model evolving that turns higher quality temporal search and reasoning data into a parametric advancement of the students.

13:00 JSTエージェント

エージェントのスキルが重要: 実行軌跡から独自のスキルを推測する

エージェント スキルは、下流のパフォーマンスを向上させる再利用可能な手順をパッケージ化します。軽量でポータブルな形式により、市場での収益化と、クラウドでホストされたエージェント インターフェイスの背後でのプライベート展開が可能になり、プロバイダーに価値の高いスキルを独占的に保持するインセンティブが与えられます。しかし、アーティファクトを非表示にしても、その動作上の影響は隠蔽されず、実行軌跡で観察可能なままとなり、動作上のサイドチャネルを形成します。私たちはこの暴露をスキル漏洩、つまり、参照回答や成功ラベルなしで、無害なクエリによって引き出された軌跡から独自のスキルを再構築すると定義します。エージェントの動作において繰り返し発生するスキル シグネチャを活用するブラックボックス フレームワークである SigLeak を紹介します。多様で意思決定の多い診断タスクを構築し、一致するスキル有効とスキル無効の軌道を対比し、分離されたパターンから再構成されたスキルを反復的に改良します。 5 つのシナリオ、3 つのモデル ファミリ、および 3 つのエージェント フレームワークにわたって、SigLeak はほぼすべての設定で 3 つのベースラインを上回るパフォーマンスまたは一致します。これにより、スキル無効の基準よりも成功率が平均 6.88 パーセント上昇し、粗いおよび細かいセマンティック類似性の指標である SkillSim 全体で最高を達成しました。これらの結果は、無害な実行軌跡によって独自の手続き上の知識が漏洩する可能性があることを示しています。コードは https://anonymous.4open.science/r/SigLeak-D1DB で入手できます。

原文 (English)

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives to keep high-value skills proprietary. Yet hiding the artifacts does not conceal their behavioral effects, which remain observable in execution trajectories and form a behavioral side channel. We define this exposure as Skill Leakage: reconstructing proprietary skills from trajectories elicited by benign queries, without reference answers or success labels. We introduce SigLeak, a black-box framework that exploits recurring skill signatures in agent behavior. It constructs diverse, decision-rich diagnostic tasks, contrasts matched skill-enabled and skill-disabled trajectories, and iteratively refines a reconstructed skill from the isolated patterns. Across five scenarios, three model families, and three agent frameworks, SigLeak outperforms or matches three baselines in nearly every setting. It raises the success rate by 6.88 percentage points over the skill-disabled reference on average and achieves the highest overall SkillSim, our metric for coarse- and fine-grained semantic similarity. These results show that benign execution trajectories can expose proprietary procedural knowledge. The code is available at https://anonymous.4open.science/r/SigLeak-D1DB.

13:00 JST研究/論文

センサートークンセルフアテンションによるマトリックスフリーの光音響画像再構成

光音響トモグラフィー(PAT)は、生体組織の光吸収コントラストと超音波の空間分解能を組み合わせたものですが、スパースビューセンサー測定から初期圧力分布を回復することは、依然として不適切な逆問題です。反復圧縮センシング ソルバーとアンロールされたディープ ネットワークは両方とも、推論時にシステム マトリックスへの依存を保持するため、リアルタイムの臨床再構成の計算コストが高くなります。この論文では、センサー アテンション ネットワーク (SAN) を提案します。これは、各センサーの完全な時系列をトークンとして扱い、推論時にシステム マトリックスを呼び出すことなく生の測定値を再構成された画像に直接マッピングする、Transformer ベースのアーキテクチャです。トレーニングとベンチマークのために、分析的な k 空間 H マトリックスが構築され、一致したジオメトリの下で k-Wave 擬似スペクトル ソルバーに対して検証され、センサーごとの平均ピアソン相関 0.919 +/- 0.049 を達成します。k 空間アポダイゼーションとガウス時間減衰が相乗的に作用して、エネルギー正規化された不一致を 49% 削減します。 488 個の拡張サンプルで血管重み付け損失を使用してトレーニングし、46 個のホールドアウト サンプルで ISTA、スプリット ブレグマン全変動 (SBTV)、学習済み ISTA (LISTA) に対して評価した結果、SAN は最高の平均 SSIM (0.522) と PSNR (22.09 dB)、最低の NMSE (0.233) を達成しました。対応のある t 検定と Wilcoxon 符号付き順位検定により、p < 1e-8 での PSNR、NMSE、ピアソン相関では LISTA よりも SAN が優れていることが確認され、すべての忠実度メトリクスにおいては ISTA および SBTV よりも SAN が優れていることが確認されています。推論時に H マトリックスをバイパスすることにより、SAN は再構築時間を少なくとも 1 桁削減し、リアルタイム PAT 再構築をサポートします。

原文 (English)

Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention

Photoacoustic tomography (PAT) combines the optical absorption contrast of biological tissue with the spatial resolution of ultrasound, yet recovering the initial pressure distribution from sparse-view sensor measurements remains an ill-posed inverse problem. Iterative compressive-sensing solvers and unrolled deep networks both retain a dependence on the system matrix at inference, which leaves real-time clinical reconstruction computationally expensive. This paper proposes the Sensor Attention Network (SAN), a Transformer-based architecture that treats the full time series of each sensor as a token and maps raw measurements directly to the reconstructed image without invoking the system matrix at inference. For training and benchmarking, an analytical k-space H-matrix is constructed and validated against the k-Wave pseudo-spectral solver under matched geometry, achieving a mean per-sensor Pearson correlation of 0.919 +/- 0.049, with k-space apodization and Gaussian temporal damping acting synergistically to reduce the energy-normalized mismatch by 49%. Trained with a vessel-weighted loss on 488 augmented samples and evaluated on 46 held-out samples against ISTA, split-Bregman total variation (SBTV), and learned ISTA (LISTA), SAN attains the highest mean SSIM (0.522) and PSNR (22.09 dB) and the lowest NMSE (0.233). Paired t-tests and Wilcoxon signed-rank tests confirm the superiority of SAN over LISTA on PSNR, NMSE, and Pearson correlation at p < 1e-8, and over ISTA and SBTV on all fidelity metrics. By bypassing the H-matrix at inference, SAN reduces reconstruction time by at least an order of magnitude, supporting real-time PAT reconstruction.

13:00 JST研究/論文

どこまで小さくできますか? 60M パラメータ モデルにおける Text-to-SQL の LoRA ランク、ターゲット モジュール、および量子化のトレードオフに関する制御された研究

パラメーター効率の良い微調整 (PEFT) と低ビット量子化は、現在、厳しい計算予算の下で言語モデルを適応させるための標準ツールとなっていますが、それらの相互作用は、設計空間の探索に費用がかかる 10 億パラメーターのモデルで研究されることがほとんどです。補足的な質問をします。特定の、完全に再現可能な 6,000 万パラメーターのエンコーダー/デコーダー モデル (T5-small) と単一テーブルのテキストから SQL へのベンチマーク (WikiSQL) では、各効率ノブのタスク精度は実際にどれくらいかかりますか? (i) {2, 4, 8, 16, 32} の LoRA ランク r、(ii) 適応されたモジュールのセット、および (iii) 数値精度について、制御された単一変数研究を実行します。トレーニング可能なパラメータ、ピークトレーニングメモリ、推論レイテンシー、スループット、フレーム適応などのシステムレベルのメトリクスと併せてタスクの精度を、精度のみの目標ではなく制約付きのトレードオフとして報告します。私たちの結果は、r=16 の LoRA が完全な微調整精度 (59.6% 対 71.2% の完全一致) の 11.6 パーセンテージ ポイント以内に回復する一方、トレーニングするパラメーターは 1% 未満であり、消費するピーク GPU メモリは 31% 少ないことが示されています。この設定内では、r=16 を超えるランクでは、測定可能な精度の向上は得られません。 INT8 および NF4 量子化を使用した QLoRA は、劇的に低いメモリ コスト (それぞれ 0.60 GB) で同等の精度 (52.8% および 53.2%) を達成し、メモリに制約のある展開にとって魅力的なトレードオフを示しています。すべてのコード、構成、ログは完全な再現性を実現するためにリリースされています。

原文 (English)

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore. We ask a complementary question: on a specific, fully reproducible 60M-parameter encoder-decoder model (T5-small) and a single-table text-to-SQL benchmark (WikiSQL), how much task accuracy does each efficiency knob actually cost? We run a controlled, single-variable study over (i) LoRA rank r in {2, 4, 8, 16, 32}, (ii) the set of adapted modules, and (iii) numerical precision. We report task accuracy alongside system-level metrics including trainable parameters, peak training memory, inference latency, and throughput, and frame adaptation as a constrained trade-off rather than an accuracy-only objective. Our results show that LoRA with r=16 recovers within 11.6 percentage points of full fine-tuning accuracy (59.6% vs. 71.2% exact-match) while training fewer than 1% of parameters and consuming 31% less peak GPU memory. Within this setting, rank beyond r=16 yields no measurable accuracy gain. QLoRA with INT8 and NF4 quantization achieves comparable accuracy (52.8% and 53.2%) at dramatically lower memory cost (0.60 GB each), demonstrating a compelling trade-off for memory-constrained deployments. All code, configurations, and logs are released for full reproducibility.

13:00 JST研究/論文

リチウム金属電解質における官能基と塩の効果の電子構造解析のための密度マトリックス フレームワーク

リチウム金属電解質の反応性は、分子官能基の相互作用、Li$^+$溶媒和、塩アニオンの関与によって生じます。この相互作用は、ドナー、アニオン、およびカチオン中心にわたる電子密度の再分布を通じて機能します。これは、空間で分解された電子構造から最も直接的に読み取られます。量子化学計算はそのような読み取り値を忠実に提供しますが、この多次元設計空間全体にわたって計算量が多くなり、機械学習の電子構造モデルでは化学的に多様な溶媒和シェルや電解質関連の読み取り値をカバーすることはほとんどありません。ここでは、電子構造の予測と解析のための密度行列中心の AI プラットフォーム (EMolStudio) を紹介します。そのワークフローには、分子官能化、明示的なLi$^+$第一殻アセンブリ、冪等投影による密度行列予測、フロンティア軌道、静電ポテンシャル、Li$^+$ドナー結合秩序、電子局在の読み出しが統合されています。 EMolStudioを4つのリチウム塩にわたる163,655個の官能化分子と22,500個の明示的なLi$^+$第一殻クラスターに適用した。我々は、1) 分子スケールでは、官能基化はフロンティアレベル、静電ポテンシャル、Li$^+$ドナー接触の化学的に異なる変化によってCO$_2$Me、CN、F/CF$_3$、スルホニル基を区別し、$\pi^*$-アクセプタ、誘導、分極の寄与と一致し、官能化の度合いが高くなるとサブリニアに蓄積することを発見した。 2) 陽的溶媒和シェルでは、アニオンの同一性によってフロンティア軌道局在が再形成されます。LiTDI はライブラリ全体にわたってアニオン上に HOMO を固定しますが、LiDFOB はアニオンにホストされた HOMO を官能基に強く依存する LUMO ホストと組み合わせます。これにより、EMolStudio は、官能基と塩の選択を、リチウム結合形成、脱溶媒和、界面反応に関連する電子構造仮説に変換します。

原文 (English)

A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes

The reactivity of lithium-metal electrolytes arises from the interplay of molecular functional groups, Li$^+$ solvation, and salt-anion participation. This interplay operates through the redistribution of electron density across donor, anion, and cation centers, which is most directly read out from the electronic structure resolved in space. Quantum-chemical calculations deliver such readouts faithfully, yet become computationally demanding across this multidimensional design space, and machine-learning electronic-structure models seldom cover chemically diverse solvation shells or electrolyte-relevant readouts. Here, we present a density-matrix-centered AI platform (EMolStudio) for electronic-structure prediction and analysis. Its workflow integrates molecular functionalization, explicit Li$^+$ first-shell assembly, density-matrix prediction with idempotency projection, and readouts of frontier orbitals, electrostatic potential, Li$^+$-donor bond order, and electron localization. We apply EMolStudio to 163,655 functionalized molecules and 22,500 explicit Li$^+$ first-shell clusters across four lithium salts. We find that 1) at the molecular scale, functionalization distinguishes CO$_2$Me, CN, F/CF$_3$, and sulfonyl groups by chemically distinct changes in frontier levels, electrostatic potential, and Li$^+$-donor contact, consistent with $\pi^*$-acceptor, inductive, and polarization contributions, with sublinear accumulation at higher degrees of functionalization; 2) in explicit solvation shells, anion identity reshapes frontier-orbital localization: LiTDI anchors the HOMO on the anion across the entire library, whereas LiDFOB pairs an anion-hosted HOMO with strongly functional-group-dependent LUMO hosting. EMolStudio thereby translates functional-group and salt choices into electronic-structure hypotheses relevant to lithium-bond formation, desolvation, and interphase reactions.

13:00 JST研究/論文

al-Sabr wa al-Taqsim による法的原因の計算による抽出: 閉じられた Fiqh 章の集合論的定式化

この論文は、法学の閉じられた章の中で法的原因(「ilal」)を抽出するための、al-Sabr wa al-Taqsim(調査と分割)の古典的なウスーリ法を集合論的に定式化したものを提示します。法的判決の真理表から最小限の運用ルールを抽出する計算アルゴリズムが導入されています。主な結果は、閉じられた章の完全な真理値表が与えられると、アルゴリズムが判決の最小限の構造生成子を計算し、論理的に冗長な属性をすべて削除することです。結果として生じる構造は、その後の法的評価において許容される原因候補を構成します。この枠組みは、学校に関連した有限の概念語彙と、調査対象の章の完全な規則表が利用可能であることを条件としています。

原文 (English)

Computational Extraction of Legal Causes via al-Sabr wa al-Taqsim: A Set-Theoretic Formalization for Closed Fiqh Chapters

This paper presents a set-theoretic formalization of the classical usuli method of al-Sabr wa al-Taqsim (Examination and Division) for extracting legal causes ('ilal) within closed chapters of jurisprudence. A computational algorithm is introduced that extracts minimal operational rules from a truth table of juristic verdicts. The principal result is that, given a complete truth table for a closed chapter, the algorithm computes the minimal structural generators of the ruling and eliminates all logically redundant attributes. The resulting structures constitute admissible candidate causes for subsequent juristic evaluation. The framework is conditional upon the availability of a finite school-relative concept vocabulary and a complete ruling table for the chapter under investigation.

13:00 JSTエージェント

気象シミュレーションのためのマルチセンサーの調整

自動運転車の認識タスクは、悪天候下でも満足に機能する必要があります。現実世界の気象データセットが不足しているため、気象シミュレーションが有望な代替手段となります。シミュレーションが実際の気象データを厳密に反映していることを確認するには、異なるセンサー間で、深刻度や粒子の位置など、同じ気象特性を表現することが重要です。これを達成するために、霧の中での気象強度の位置合わせのための Reference Dataset Alignment Method (ReDAM) と、雨と雪の中での粒子の位置の位置合わせのための Unified-weather-edit (Weather-edit[1] からインスピレーションを得た) を提案します。統計的テストと幾何学的テストをそれぞれ使用して、両方の位置合わせ方法を検証します。位置合わせされていないバージョンの 3D 検出モデルは、位置合わせされたバージョンと比較して過度に楽観的になる傾向があることがわかりました。また、既存のセンサー フュージョン モデルを微調整することにより、3D 物体検出タスクのロバスト性を達成するための整列マルチセンサー シミュレーションの有効性も示します。

原文 (English)

Multi-Sensor Alignment for Weather Simulations

Perception tasks for autonomous vehicles need to work satisfactorily in adverse weather conditions. Due to lack of real-world weather datasets, weather simulations are a promising alternative. To ensure simulations closely mirror real-world weather data, it's crucial that they represent the same weather characteristics, including severity and particle positioning, across different sensors. To achieve this, we propose the Reference Dataset Alignment Method (ReDAM) for weather intensity alignment in fog and Unified-weather-edit (inspired by Weather-edit[1]) for particle positioning alignment in rain and snow. We validate both alignment methods using statistical and geometrical tests, respectively. We find that 3D detection models for non-aligned versions tend to be overly optimistic as compared to aligned versions. We also show the aligned-multi-sensor simulation's effectiveness for achieving robustness for 3D object detection task by finetuning existing sensor fusion models on it.

13:00 JSTハードウェア/半導体

認識論を超えて: 技術記号論マシンとしての認識論的統合論と大規模言語モデル

クアトロシオッキらは、大規模な言語モデルの流暢な出力により、言語的妥当性が認識論的評価の代わりとなり、彼らが*認識論*と呼ぶ状態、つまり、通常であれば判断が保証される実践を行わずに知識を所有する経験が生じる可能性があると警告している。この論文はその診断を受け入れますが、身体化された社会的に位置する人間の認識者を孤立した生成モデルと比較し、それによって自律エージェントの内部能力に認識論的正当性を位置づけるその説明枠組みに異議を唱えます。カルロ・シーニの実践、執筆、記号、および技術の哲学に基づいて、私たちは代わりに、人間の執筆の堆積したアーカイブからもっともらしい言語構成を生成することによって書かれた記号論の段階を自動化する*テクノ記号論マシン*として大規模言語モデル(LLM)を理解することを提案します。この観点から見ると、*認識論*は、私たちが*認識論的分裂病*と呼ぶ、より広範な現象の1つの結果です。つまり、言語的に完成された表現としての記号と、社会的に埋め込まれた解釈、証拠、批判、検証、および責任の回路内の瞬間としての記号の間の社会技術的亀裂です。この切断は、認識論的結果の最終性を伴うもっともらしい継続が提示される*エイコティック閉包*によって、またアルゴリズムの権威と認識論的な自己誤認識によって強化されます。したがって、関連する単位はモデルだけではなく、生成された碑文がプロンプトされ、解釈され、検証され、異議が唱えられ、使用され、結果として生じる完全な実践です。この再構成は、言語的生産と責任ある理解との区別を維持しながら、検査可能な系図、競争可能性、分散された責任、認識論的主体性、ハイブリッド人間の評価、つまり AI 実践を中心とした設計プログラムを基礎としています。

原文 (English)

Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines

Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted. This article accepts that diagnosis but challenges its explanatory framework, which compares an embodied, socially situated human knower with an isolated generative model thereby locating epistemic legitimacy in capacities internal to autonomous agents. Drawing on Carlo Sini's philosophy of practices, writing, signs, and technics, we propose instead to understand a large language model (LLM) as a *techno-semiotic machine* that automates a phase of written semiosis by producing plausible linguistic configurations from the sedimented archive of human writing. From this perspective, *Epistemia* is one consequence of a broader phenomenon that we call *epistemic schizologia*: the socio-technical cleavage between signs as linguistically accomplished expressions and signs as moments within socially embedded circuits of interpretation, evidence, criticism, verification, and responsibility. This cleavage is reinforced by *eikotic closure*, through which a plausible continuation is presented with the finality of an epistemic result, and by algorithmic authority and epistemic self-misrecognition. The relevant unit is therefore not the model alone but the complete practice in which generated inscriptions are prompted, interpreted, verified, contested, used, and made consequential. This reframing preserves the distinction between linguistic production and responsible understanding while grounding a design programme centred on inspectable genealogy, contestability, distributed responsibility, epistemic agency, and the evaluation of hybrid human--AIpractices.

13:00 JST研究/論文

正の二次ネットワークにおける商ダイナミクス、実効曲率、および暗黙的なバイアス

正の 2 次ネットワークでは、低ランク表現 f_U(x)=x^top UU^top x が認められます。ここで、Uinmathbb{R}^{dtimes r} は、右直交乗算までしか識別できず、ランク r の PSD 行列 Q=UU^top を表します。私たちは、この商の構造がトレーニングのダイナミクス、曲率、回復、補間バイアスをどのように制御するかを研究します。フル列ランク層では、ランク r PSD 多様体を使用して mathbb{R}^{dtimes r}_*/O(r) を識別します。スムーズな目的 L(U)=ell(UU^top) の場合、ユークリッド係数の勾配は水平になります。したがって、因子勾配流は商リーマン勾配流に正確に投影されますが、有限ステップ勾配降下法は予測子の正確な合同再帰を引き起こします。二次回帰の場合、商計量に対する接線空間に制限された経験的測定グラム形式として、補間器での有効ヘシアンを導出します。ガウス ランク 1 測定の下で、母集団曲率を計算し、経験的正規演算子の一様偏差限界を証明し、スペクトル初期化子を構築し、勾配流に対する局所指数収束とスモールステップ降下に対する線形収束を確立します。回復保証は明示的ですが、全空間の秒瞬間制御に依存しているため、保守的です。不十分に決定された通勤体制では、因子勾配の流れは、結合スペクトル座標における正確なエントロピーミラーの流れになります。厳密に正の初期化は、内挿セットへのブレグマン投影に収束します。等方性初期化 q(0)=varepsilon^2mathbf{1} を使用すると、予測子は varepsilondownarrow0 として設定された最小トレース解に近づき、不変結合スペクトル代数内の重み付きエントロピーによって非一意性を解決します。有限ステップ降下では、O(η) による連続時間ブレグマン投影とは異なる内挿が選択されます。数値実験により、これらの商の正体、曲率予測、回復挙動、および選択則が検証されます。

原文 (English)

Quotient Dynamics, Effective Curvature, and Implicit Bias in Positive Quadratic Networks

Positive quadratic networks admit the low-rank representation f_U(x)=x^top UU^top x, where Uinmathbb{R}^{dtimes r} is identifiable only up to right orthogonal multiplication, representing a rank-r PSD matrix Q=UU^top. We study how this quotient structure governs training dynamics, curvature, recovery, and interpolation bias. On the full-column-rank stratum, we identify mathbb{R}^{dtimes r}_*/O(r) with the rank-r PSD manifold. For smooth objectives L(U)=ell(UU^top), the Euclidean factor gradient is horizontal. Thus, factor gradient flow projects exactly to quotient Riemannian gradient flow, while finite-step gradient descent induces an exact congruence recursion for the predictor. For quadratic regression, we derive the effective Hessian at interpolators as the empirical measurement Gram form restricted to the tangent space relative to the quotient metric. Under Gaussian rank-one measurements, we compute population curvature, prove uniform deviation bounds for the empirical normal operator, construct a spectral initializer, and establish local exponential convergence for gradient flow and linear convergence for small-step descent. Recovery guarantees are explicit but conservative due to reliance on full-space second-moment control. In underdetermined commuting regimes, factor gradient flow becomes an exact entropy mirror flow in joint spectral coordinates. Strictly positive initializations converge to Bregman projections onto the interpolation set. With isotropic initialization q(0)=varepsilon^2mathbf{1}, predictors approach the minimum-trace solution set as varepsilondownarrow0, resolving nonuniqueness via weighted entropy within the invariant joint spectral algebra. Finite-step descent selects interpolants differing from continuous-time Bregman projections by O(eta). Numerical experiments verify these quotient identities, curvature predictions, recovery behaviors, and selection laws.

13:00 JST研究/論文

中国語の音声生成と知覚におけるEEGからテキストへのデコードのためのテキストと音声の共同調整

頭皮脳波 (EEG) から音声情報をテキストに直接デコードすることで、重度の音声障害や運動障害を持つ個人に潜在的な非侵襲性の神経伝達経路が提供されます。皮質電図検査などの侵襲的アプローチと比較して、EEG はより安全で広範囲に導入可能ですが、解読はかなり困難です。この課題は、数千の文字、被験者間の深刻な変動性、およびテキストの位置合わせのための低い信号対雑音比を含む高次元の出力空間を処理する必要がある中国語文の解読ではさらに悪化します。既存の方法は、単一の監視軸、つまりテキスト セマンティクスまたはオーディオ音響特徴のいずれかに取り組んでいますが、どちらも同時に満足させることはできません。大量の語彙を含む中国語の解読には、文レベルの識別能力ときめ細かい時間分解能が求められます。我々は、EEGAlign を導入します。これは、EEGA を 2 つの軸、すなわち BGE-M3 テキスト埋め込みによるテキストの位置合わせと、対照学習とそれに続く CTC 文字列デコードによる wav2vec~2.0 音声特徴による音声の位置合わせを組み合わせて位置合わせする、新しいパラメータ効率の高いフレームワークです。 ChineseEEG-2 データでは、EEGAlign は最先端のクローズドセット文分類パフォーマンスを実現し、101 の候補のうち音読 EEG でトップ 1 の精度が 82.37%、受動聴取 EEG で 41.43% に達します。アブレーション研究では、2 つのアライメント軸が高度に補完的であることが示されており、これらを組み合わせることで、どちらか一方を単独で使用するよりも一貫して優れたパフォーマンスが得られます。私たちの知る限り、これは、明白な音声生成中に非侵襲的脳波から語彙の多い中国語文を解読し、比較的大規模な閉集合候補文設定で強力な分類性能を達成することに関する最初の研究です。

原文 (English)

Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway for individuals with severe speech and motor impairments. Compared with invasive approaches such as electrocorticography, EEG is safer and more widely deployable, yet substantially more challenging to decode.This challenge is exacerbated for Chinese sentence decoding, which must handle a high-dimensional output space with thousands of characters, severe inter-subject variability, and low signal-to-noise ratios for text alignment.Existing methods commit to a single supervisory axis---either text semantics or audio acoustic features---yet neither can simultaneously satisfy the demands of sentence-level discriminability and fine-grained temporal resolution required for large-vocabulary Chinese decoding. We introduce EEGAlign, a novel parameter-efficient framework that jointly aligns EEG with two axes---text alignment with BGE-M3 text embeddings and audio alignment with wav2vec~2.0 speech features via contrastive learning followed by CTC character-sequence decoding. On ChineseEEG-2 data, EEGAlign yields state-of-the-art closed-set sentence classification performance, reaching up to 82.37% Top-1 accuracy on Reading Aloud EEG and 41.43% on Passive Listening EEG out of 101 candidates. Ablation studies show that the two alignment axes are highly complementary: combining them yields consistently better performance than either alone. To the best of our knowledge, this is the first study on decoding large-vocabulary Chinese sentences from non-invasive EEG during overt speech production, and achieving strong classification performance with relatively large closed-set candidate-sentence setting.

13:00 JSTLLM/生成AIGPT / ChatGPTLlama

AIriskEval-edu デモ: 教育的説明における教育的リスクの監査

私たちは、指導説明の教育的品質を監査し、説明可能な監査結果を提供するプラットフォームである AIriskEval-edu デモを紹介します。このプラットフォームは、教育学的リスクの 5 つの側面 (事実の正確さ、深さと完全性、焦点と関連性、生徒レベルの適切性、イデオロギーの偏見) をカバーするルーブリックに照らして説明を評価します。ディメンションごとに、二分決定と信頼スコアを返します。検出されたリスクには、自然言語の根拠と、深さと完全性を除いて、局所的な証拠範囲も含まれます。このプラットフォームは、外部 API とコンシューマー グレードの GPU で実行されるセルフホスト型 Llama 3.1 8B エバリュエーターを通じて GPT-5.5 を統合します。ローカル評価者は、リスクと説明可能性の注釈を備えた幼稚園から高校までの教育説明のデータセットである AIriskEval-edu で微調整されています。このプラットフォームは 2 つのモードで動作します。AI モードでは、両方の評価者が、それぞれ異なる教育的行動と潜在的なリスクを表す 6 つのシミュレートされた教師プロファイルに基づいて生成された、保存された説明を評価します。人間モードでは、ローカル評価者がユーザーが作成した説明をリアルタイムで監査します。ローカル評価者は、報告されているほとんどの指標で GPT-5.5 を上回っており、教育機関に監査済みのコンテンツを独自のインフラストラクチャ内に保持する実用的な方法を提供します。

原文 (English)

AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations

We present AIriskEval-edu Demo, a platform that audits the pedagogical quality of instructional explanations and provides explainable audit results. The platform evaluates an explanation against a rubric covering five dimensions of pedagogical risk: factual accuracy, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. For each dimension, it returns a binary decision and a confidence score. Detected risks also include a natural-language rationale and, except for Depth and Completeness, a localized evidence span. The platform integrates GPT-5.5 through an external API and a self-hosted Llama 3.1 8B evaluator that runs on consumer-grade GPUs. The local evaluator is fine-tuned on AIriskEval-edu, a dataset of K-12 instructional explanations with risk and explainability annotations. The platform operates in two modes: in AI mode, both evaluators assess stored explanations generated under six simulated teacher profiles, each representing a distinct pedagogical behavior and potential risk; in human mode, the local evaluator audits user-written explanations in real time. The local evaluator outperforms GPT-5.5 on most reported metrics, offering educational institutions a practical way to keep audited content within their own infrastructure.

13:00 JSTビジネス/資金調達

エンジンは平等、人間は不平等: エンジンが評価した平等なチェスの局面における再現可能な結果の偏り

強力なエンジンが本質的に等しいと判断するチェスの開始位置 (Stockfish 18 のゼロから 10 センチポーン以内の評価、深さ安定) と人間が Lichess で実際に到達する位置 (2025 年 10 月、1,661 の位置、1,610 万回) の間では、人間の結果はバランスが取れていません。ポジションには結果の偏りがあり、それぞれのゲームの実際の結果とプレーヤーのレーティングの予測との間のギャップがあり、その方向性は自然に到達するポジションの安定した特性です。いくつかのポジションは白を支持し、他のポジションは黒を支持します。これらの偏りは、3 つの再パーティション、つまり、不連続なプレイヤー アカウント セット (プライマリ)、時間、および不連続なレーティング バンドにわたって再現され、さらに 8 か月後のサンプル外の月にも再現されます。プライマリ分割では、各ポジションのスキューが各口座グループで 1 回測定され、反復勾配は、格付けとオープニングファミリー効果を除去した後、一方の測定値が他方の測定値をどの程度正確に予測するかを尋ねます。1 つは減少しないキャリーオーバーを意味し、もう 1 つは減少しないキャリーオーバーを意味します。ゼロ、線形関係はありません。結果は 0.69 (ファミリークラスター化 95% CI [0.65, 0.74]) であり、最も人気があり、最もよく測定されたポジションでは 0.94 に上昇しました。傾きの値はポジションの組み合わせによって異なります。存在は不変の主張です。それは、テストするあらゆる厳しい評価範囲、検索の深さ、校正、および人気のカットオフに耐え、電撃内と急速内で別々に複製されます。典型的な偏りは小さい (中央値 $|\delta| \約 0.018$、白スコアの約 2 パーセント ポイント) にもかかわらず、それは、関連性のないアカウント間で位置ごとに再現されます。このような位置では、不利な側も長く考えます。評価に最も信頼性がある場合でも、それは人間の成果を表す十分な統計ではありません。結果は観察的なものであり、因果関係の疑問は、事前に登録されたランダム化されたコンパニオン研究に委ねられます。

原文 (English)

Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions

Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games' actual results and what the players' ratings predict, whose directions are stable properties of the naturally-reached position: some positions favour White, others Black. These skews reproduce across three re-partitions -- disjoint player-account sets (primary), time, and disjoint rating bands -- and on an out-of-sample month eight months later. On the primary split, each position's skew is measured once in each account group, and the replication slope asks how well one measurement predicts the other after removing rating and opening-family effects: one means undiminished carry-over; zero, no linear relation. We find 0.69 (family-clustered 95% CI [0.65, 0.74]), rising to 0.94 on the most-popular, best-measured positions. The slope's value depends on the position mix. Existence is the invariant claim: it survives every tighter evaluation band, search depth, calibration, and popularity cutoff we test, and replicates within blitz and rapid separately. The typical skew is small (median $|\delta| \approx 0.018$, about two percentage points of White score), yet it reproduces, position by position, across disjoint accounts. At these positions the disfavoured side also thinks longer. Even where the evaluation is most confident, it is not a sufficient statistic for human outcomes. The result is observational, and the causal question is left to a pre-registered randomised companion study.

13:00 JSTエージェントClaude

OrchBench: 決定論的シミュレーションによる個別のマルチエージェント オーケストレーション プランの評価

複雑なタスクは多くの場合、並列化可能でありながら相互依存するサブタスクに分解されるため、マルチエージェント システム (MAS) のパフォーマンスにとってオーケストレーションが重要になります。既存の評価は通常、エンドツーエンドの実行に依存しており、オーケストレーション計画の品質と作業者の能力、ツールの信頼性、および環境ノイズを混同します。さらに、実際の実行にかかる時間とトークンのコストはワークフローの規模に応じて急速に増大するため、体系的な評価が高価になります。マルチエージェント オーケストレーション プランを個別に評価するためのシミュレーション ベースのベンチマークである OrchBench を紹介します。 OrchBench は、実際のタスクから開始して、サイズと並列度を制御してタスクの依存関係をエンコードする有向非巡回グラフ (DAG) を構築します。 DAG、エージェントごとのコンテキスト制限、およびエージェントの予算を考慮して、評価されたプランナーはサブタスクをエージェントに割り当て、エージェント間の情報転送とその保持率を指定します。決定論的シミュレーターは、ワーカー エージェントを呼び出すことなく結果の計画を評価し、結果の品質、メイクスパン、およびトークン コストの解釈可能な尺度を返します。 OrchBench によって生成されたシミュレートされたスコアは、クロード コードの実行からの品質スコアと強い相関があり、\(r=0.816\) のピアソン相関を達成しながら、トークンの \(1.3\%\) と実時間の \(10.3\%\) のみが必要です。さまざまなプランナーやワークフロー規模にわたって、単にエージェントの数を増やすよりも、タスクに不可欠な情報を保持することの方が重要であり、調整の失敗が蓄積するにつれて並列処理の利点が減少することがわかりました。これらの結果により、OrchBench は、マルチエージェント オーケストレーション プランを比較および診断するための効率的で解釈可能なベンチマークとして確立されます。

原文 (English)

OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation

Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS). Existing evaluations typically rely on end-to-end execution, which conflates orchestration-plan quality with worker capabilities, tool reliability, and environmental noise. Moreover, the time and token costs of real execution grow rapidly with workflow scale, making systematic evaluation expensive. We present OrchBench, a simulation-based benchmark for evaluating multi-agent orchestration plans in isolation. Starting from real-world tasks, OrchBench constructs directed acyclic graphs (DAGs) that encode task dependencies, with controlled sizes and degrees of parallelism. Given a DAG, a per-agent context limit, and an agent budget, the evaluated planner assigns subtasks to agents and specifies cross-agent information transfers and their retention ratios. A deterministic simulator evaluates the resulting plan without invoking worker agents and returns interpretable measures of result quality, makespan, and token cost. The simulated scores produced by OrchBench correlate strongly with quality scores from Claude Code executions, achieving a Pearson correlation of \(r=0.816\), while requiring only \(1.3\%\) of the tokens and \(10.3\%\) of the wall-clock time. Across diverse planners and workflow scales, we find that preserving task-critical information is more important than simply increasing the number of agents, and the benefits of parallelism diminish as coordination failures accumulate. These results establish OrchBench as an efficient and interpretable benchmark for comparing and diagnosing multi-agent orchestration plans.

13:00 JSTハードウェア/半導体

CoRT: トークンレベルのルーブリックに基づくポリシー最適化のための反事実リプレイ

ルーブリックベースの強化学習は、明示的な基準に照らしてモデルの出力を評価することにより、言語モデルのトレーニングを強化します。しかし、GRPO スタイルのパイプラインでは、これらの構造化された判断はスカラー応答レベルの報酬に還元され、応答レベルの利点に変換され、生成されたすべてのトークンに均一にブロードキャストされます。これにより、異なる基準が異なるスパン、フォーマット決定、またはセマンティック選択に基づいている場合でも、応答内でクレジットを割り当てるための明示的なメカニズムが残されません。ルーブリック条件付き GRPO のトークンレベルのクレジット重み付け手法である CoRT を提案します。補助トークン スコアリング モデルをトレーニングする代わりに、CoRT は反事実リプレイを使用して、元のルーブリック条件付きプロンプトと一致した基準なしのプロンプトの下で同じサンプリングされた応答を再スコアリングします。結果として得られるトークンごとの対数尤度対比は、ルーブリック コンテキストへの依存性の代用として機能します。 CoRT は、これらのコントラストを、制限された応答正規化された重みにマッピングし、それらを使用して、補助スコアラーを導入したり、応答レベルの報酬を変更したりすることなく、署名された GRPO の利点をトークン全体に再分配します。命令調整モデルと報酬粒度にわたる実験では、大部分の比較において、CoRT が一致する応答レベルの GRPO よりも改善し、平均 4.4 パーセント ポイント向上していることが示されています。この方法は、個別の関連性学習段階を回避しながら、学習されたトークンレベルの信用ベースラインとの競争力を維持します。これらの結果は、政策内部の反事実尤度の対比が、GRPO の単純さと安定性を維持しながら、応答内クレジット割り当てのための効果的なトレーニング シグナルを提供することを示唆しています。

原文 (English)

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.

13:00 JSTLLM/生成AI

局所的な適応により、トランスフォーマーの独特の学習特徴が明らかに

トランスフォーマーの適応は、意図した変更が狭い場合でも、通常、モデルの深度全体に分散されます。私たちは、適応部位がモデルの学習内容をどのように形成するか、その学習がどの程度一般化されるか、どのように選択的に適用されるかを調査します。我々は、5つの目的(語彙結合、事実関連、行動ポリシー学習、因果関係マッピング、および手続き的推論)にわたる制御されたベンチマークを導入し、各目的の「適応幾何学」を、フルスタックおよび初期、中期、または後期層のLoRAの下での取得、転送、および有界性のプロファイルとして定義します。対物レンズは独特の形状を示します。字句バインディングは、取得と有界性については初期層の適応に有利ですが、転送にはより広範な更新が必要です。事実の関連性は、ローカライズされたアダプター間の後続の層に有利です。行動学習は、後期層のアクション取得を中間層のポリシー ゲートから分離します。そして、因果的および手続き的移転は、ミドルスタックまたはフルスタックの適応から最も恩恵を受けます。これらのパターンは、パラメーターが一致した制御下で主に持続し、対応する方向コントラストのほとんどが 5 つのモデル ファミリにわたって複製されます。これらの発見により、どのモデルが学習し、一般化し、変更しないままにするかを制御するための主要な設計変数として適応部位が確立されます。

原文 (English)

Localized Adaptation Reveals Distinct Learning Signatures in Transformers

Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes what a model learns, how well that learning generalizes, and how selectively it is applied. We introduce a controlled benchmark spanning five objectives (lexical binding, factual association, behavioral policy learning, causal mapping, and procedural reasoning) and define each objective's "adaptation geometry" as its profile of acquisition, transfer, and boundedness under full-stack and early-, middle-, or late-layer LoRA. The objectives exhibit distinct geometries. Lexical binding favors early-layer adaptation for acquisition and boundedness but requires broader updates for transfer; factual association favors later layers among localized adapters; behavioral learning separates late-layer action acquisition from middle-layer policy gating; and causal and procedural transfer benefit most from middle- or full-stack adaptation. These patterns largely persist under parameter-matched controls, and most corresponding directional contrasts replicate across five model families. These findings establish adaptation site as a key design variable for controlling what models learn, generalize, and leave unchanged.

13:00 JSTLLM/生成AI

OmniDelta: OmniLLM でのトークン圧縮のためのスキル主導の予算割り当て

新しいオムニモーダル大規模言語モデル (OmniLLM) により、テキスト、オーディオ、およびビデオを統一的に理解できるようになりますが、その長いオーディオ/ビデオ トークン シーケンスにより、メモリと推論のコストが大幅に増加します。既存の圧縮方法は主に、固定予算の下で重要なトークンを選択することに焦点を当てており、前述の予算割り当ての問題は十分に調査されていません。我々は、直接的なクエリとオーディオ/ビデオの類似性は、モーダル間の予算配分では信頼できないこと、および一律のモーダル内予算では、冗長なコンテンツを保持したまま重要な証拠を見逃してしまう可能性があることを示します。これらの制限に対処するために、私たちは、意図を意識したモーダル間割り当てとコンテンツを意識したモーダル内割り当てを組み合わせる、トレーニング不要のスキル主導型フレームワークである OmniDelta を提案します。 OmniDelta は、最初にオーディオとビデオのスキル プールを構築して、クエリの需要に応じて固定保持トークン バジェットをシフトし、次にローカルの複雑さと時間的冗長性を使用して、モダリティ バジェットをオーディオ セグメントとビデオ フレームに再割り当てします。結果として得られるローカル予算は既存のプルーニング戦略と組み合わせることができ、予算の使用先を変更しながら合計保持トークン率を維持できます。 2 つの Qwen2.5-Omni モデルを使用した 4 つのオーディオ/ビデオ ベンチマークの実験では、OmniDelta が枝刈り比全体にわたって新しい精度効率のパレート フロンティアを確立していることが示されています。 Qwen2.5-Omni-7B では 25% のトークン保持率で、OmniDelta は GPU メモリを 22.0% 削減し、フルトークン推論と比較して 1.64 倍のエンドツーエンドの高速化を達成します。

原文 (English)

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget allocation, and that uniform intra-modal budgets can miss key evidence while retaining redundant content. To address these limitations, we propose OmniDelta, a training-free, skill-driven framework that couples intent-aware inter-modal allocation with content-aware intra-modal allocation. OmniDelta first constructs audio and video skill pools to shift the fixed retained-token budget according to query demand, then reallocates modality budgets over audio segments and video frames using local complexity and temporal redundancy. The resulting local budgets can be combined with existing pruning strategies, preserving the total retained-token ratio while changing where the budget is spent. Experiments on four audio-video benchmarks with two Qwen2.5-Omni models show that OmniDelta establishes a new accuracy-efficiency Pareto frontier across pruning ratios. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reduces GPU memory by 22.0% and achieves a 1.64x end-to-end speedup over full-token inference.

13:00 JSTLLM/生成AI

DecoEvo: テキスト空間におけるソルバー スキルとルーブリック ジェネレーター スキルのスコア分離型共進化

テキスト空間の最適化では、モデルの重みではなく外部の自然言語アーティファクトを編集することで大規模言語モデル (LLM) を適応させるため、最適化されたアーティファクトは検査可能なままであり、モデルをブラック ボックスとして扱うことができます。ただし、既存のテキストスペースメソッドのほとんどは評価を固定したままにします。オープンエンド タスクでは、これがボトルネックになる可能性があります。ソルバーがルーブリックの測定基準を改善すると、省略されたディメンションは最適化信号には見えないままになります。現在のソルバーのスコアによって更新が選択されている場合、ルーブリックを単純に進化させることも信頼性が低くなります。ルーブリックを満たしやすくすることで明らかな進歩が得られる可能性があるためです。最適化中にゴールド ルーブリックを使用せずに、分離された目標の下でソルバー スキルとルーブリック ジェネレーター スキルを共進化させる DecoEvo (Decoupled Co-Evolution) を紹介します。ソルバー スキルは基準レベルのフィードバックを使用して更新され、ルーブリック ジェネレーター スキルは、集計されたソルバー スコアとは独立した要件の適用範囲と応答の識別の補完的な監査を通じて改訂されます。この分離により、ジェネレーターの更新は新たに明らかになったソルバーの弱点に焦点が当てられ、ソルバーがすでに満たしている基準が繰り返し強調されることが減ります。各ベンチマークの公式評価では、DecoEvo は 5 つのベンチマークと 3 つの LLM バックボーンにわたって比較されたすべての手法を上回り、5 つのベンチマークの平均で SkillOpt に対して 2.8 ~ 5.0\% の相対的な向上をもたらしました。

原文 (English)

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\% relative gains over SkillOpt in the five-benchmark average.

13:00 JSTLLM/生成AICopilot

Cognivia: 証拠に基づいたメンタルヘルスケアのための認知行動療法の副操縦士

認知の歪みは否定的な感情を増幅させ、精神的健康障害の一因となります。認知行動療法(CBT)は認知の歪みに対処する効果的な方法ですが、専門のセラピストが不足しているため、その大規模な適用は制限されています。最近、大規模言語モデル (LLM) がメンタルヘルスへの応用に向けて研究されていますが、既存の手法では依然として、限定された領域特異性、過度にお世辞な応答、および認知の歪みに対する明確に定義された注釈の欠如という問題に悩まされています。この論文では、自動的な認知の歪みの特定と合理的な応答生成を統合する、証拠に基づいた人工知能セラピストである Cognivia を提案します。私たちのフレームワークは、コアパラダイムおよび標準参考文献として広くみなされている権威ある CBT テキストに基づいて構築されています。これは、メンタルヘルスの質問と回答 (Q and A) データでさらに強化されており、行動科学の専門家の監督の下、多段階のプロンプトと構造化された生成戦略を採用しています。次に、この拡張された CBT データセットで軽量 LLM を微調整して、Cognivia を取得します。さらに、AI 研究者と行動科学の専門家の協力を通じて開発された、LLM によって生成された合理的な応答を評価するための初の階層的品質評価フレームワークを提案します。 Cognivia は、語彙メトリクス、2 つの補完的な基準を備えた LLM ベースの審査員、および 10 人の行動科学の専門家による人間による評価を使用して評価されます。これは、認知の歪みの認識と合理的な応答の生成においてベースライン手法を常に上回っており、その有効性を示しています。私たちのコードは https://github.com/SNOWTEAM2023/Cognivia で入手できます。

原文 (English)

Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare

Cognitive distortion amplifies negative emotions and contributes to mental health disorders. Cognitive Behavioral Therapy (CBT) is an effective way to address cognitive distortions, but its large-scale application is limited by the shortage of professional therapists. Although large language models (LLMs) have recently been explored for mental health applications, existing methods still suffer from limited domain specificity, overly flattering responses, and the absence of well-defined annotations for cognitive distortions. This paper proposes Cognivia, an evidence-based artificial intelligence therapist that integrates automatic cognitive distortion identification and rational response generation. Our framework is built on authoritative CBT texts widely regarded as core paradigms and standard references. It is further augmented with mental health question-answer (Q and A) data, and employs multi-stage prompting and structured generation strategies under the supervision of behavioral science experts. Then we fine-tune a lightweight LLM on this augmented CBT dataset to obtain Cognivia. In addition, we propose the first hierarchical quality evaluation framework for assessing LLM-generated rational responses, developed through collaboration between AI researchers and behavioral science experts. Cognivia is evaluated using lexical metrics, LLM-based Judges with two complementary criteria, and human evaluation by 10 behavioral science experts. It consistently outperforms the baseline methods in cognitive distortion recognition and rational response generation, demonstrating its effectiveness. Our code is available at https://github.com/SNOWTEAM2023/Cognivia.

13:00 JSTLLM/生成AI研究/論文

LLM が生成する推奨事項の説明を通じて持続可能な選択を促す

レコメンダー システムは日常の消費を仲介し、持続可能な選択を促すための有望なチャネルを提供します。以前の研究では、説明が推奨事項に対するユーザーの認識に影響を与え、より多くの情報に基づいた意思決定をサポートできることが示されています。私たちは、選択の瞬間に持続可能性の情報を前面に出すことで、説明が行動を促す役割を果たすこともできると主張します。この研究では、推奨事項の説明におけるサステナビリティ情報のさまざまな行動枠組みが、ユーザーの選択や認識にどのような影響を与えるかを調査しています。生成 AI を使用して、ナッジ理論を利用してサステナビリティを意識した説明を生成し、人間による評価と LLM による審査員監査を通じて検証します。この基盤に基づいて、関与度の低い領域 (インスタント コーヒー) と関与度の高い領域 (ホテルの予約) で 2 つのランダム化研究 ($N = 529 ドル) を実施します。この研究では、参加者は、これらの説明を伴う好みに合った推奨事項の中から選択します。私たちの結果は、どちらの領域においても、説明の中で持続可能性に関する情報を開示するだけでは選択肢は変わらないのに対し、その情報を枠組み化したり説明的な社会規範を援用したりすることで持続可能な選択が大幅に増加し、意思決定が容易になることを示しています。特に、単純な開示はより持続可能な選択行動に変換することなく説明評価を向上させるため、認識と行動は乖離します。私たちの研究は、LLM がどのように理論​​に基づいた説明を大規模に生成できるかを実証し、社会的利益のための実践的な説明に基づく介入を示唆しています。最後に、生成 AI による適応説明デザインへの影響について説明します。

原文 (English)

Nudging Sustainable Choices through LLM-Generated Recommendation Explanations

Recommender systems mediate everyday consumption, offering a promising channel for encouraging sustainable choices. Prior research shows that explanations influence users' perceptions of recommendations and can support more informed decisions. We argue that explanations can also serve as behavioral nudges by foregrounding sustainability information at the moment of choice. This study investigates how different behavioral framings of sustainability information in recommendation explanations affect user choices and perceptions. Using generative AI, we generate sustainability-aware explanations by drawing on nudge theory and validate them through human evaluation and LLM-as-a-judge audits. Building on this foundation, we conduct two randomized studies ($N = 529$) in a low involvement domain (instant coffee) and a high involvement domain (hotel bookings), in which participants choose among preference matched recommendations accompanied by these explanations. Our results show that, across both domains, merely disclosing sustainability information in explanations does not change choices, whereas framing that information or invoking a descriptive social norm significantly increases sustainable selections and eases decision-making. Notably, perception and behavior diverge, as plain disclosure improves explanation evaluations without translating into more sustainable selection behavior. Our work demonstrates how LLMs can generate theory-grounded explanations at scale, pointing toward practical explanation-based interventions for social good. We conclude by discussing implications for adaptive explanation design with generative AI.

13:00 JST研究/論文

損失の不変性は、レイヤーがエンコードする概念を決定します: 心エコー検査におけるボリュームグラウンディング

目的: コンセプトのボトルネックは、解釈可能な中間変数を介してルート予測をモデル化し、その妥当性は通常、それらの変数がどれだけ正確に予測されるかによって判断されます。心エコー検査ビデオからの駆出率推定の基礎となる概念として左心室容積を使用して、その判断が十分であるかどうかを尋ねます。方法: ビデオ トランスフォーマー エンコーダーは、公開されている心エコー検査データセットでトレーニングされました。収縮終期容積と拡張終期容積は概念層を形成し、そこから駆出率が分析的に計算され、出力への残留経路はありません。私たちは、駆出率の目標のみに基づいたトレーニングと、ミリリットル単位の量の追加の監視を伴うトレーニングを比較し、1,276件の実施された研究で両方を評価しました。結果: コンセプトのボトルネックは、平均絶対誤差 7.13 に対して 6.89 で、直接回帰と比較して駆出率誤差を増加させませんでした。しかし、量の監視がなければ、相関関係は部分的に保たれたものの、予測量の広がりは参照広がりの 35.7 ミリリットルと 45.7 ミリリットルに対して 0.1 ミリリットルにまで崩壊しました。これは目的の不変特性から導かれることを示します。駆出率は比率であり、両方の体積が再スケーリングされても変化しないため、損失はスケールまでしか概念層を決定しません。絶対単位での監視により、駆出率誤差が 0.4 減少しましたが、体積誤差は 89.8 ミリリットルから 25.8 ミリリットルに減少しました。結論: コンセプトの正確さだけで、物理的なスケールを持たないコンセプト層を隠すことができます。重要性: 臨床モデルの解釈可能な中間変数は、予測精度だけでなく、トレーニング目標の不変構造に対しても検証される必要があります。

原文 (English)

Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

Objective: Concept bottleneck models route prediction through interpretable intermediate variables, and their validity is normally judged by how accurately those variables are predicted. We ask whether that judgement is sufficient, using left ventricular volumes as the concepts underlying ejection fraction estimation from echocardiographic video. Methods: A video transformer encoder was trained on a publicly available echocardiography dataset. End-systolic and end-diastolic volumes formed a concept layer from which ejection fraction was computed analytically, with no residual path to the output. We compared training under an ejection fraction objective alone against training with additional supervision of the volumes in millilitres, and evaluated both on 1276 held-out studies. Results: The concept bottleneck did not increase ejection fraction error relative to direct regression, at 6.89 against 7.13 mean absolute error. Without volume supervision, however, the spread of predicted volumes collapsed to 0.1 millilitres against reference spreads of 35.7 and 45.7 millilitres, while correlation was partly preserved. We show that this follows from an invariance property of the objective: ejection fraction is a ratio and is unchanged when both volumes are rescaled, so the loss determines the concept layer only up to scale. Supervision in absolute units reduced volume error from 89.8 to 25.8 millilitres at a cost of 0.4 in ejection fraction error. Conclusion: Concept accuracy alone can conceal a concept layer that carries no physical scale. Significance: Interpretable intermediate variables in clinical models should be validated against the invariance structure of the training objective, not only against prediction accuracy.

13:00 JSTエージェント

推論しながら推測する: エージェントと投機者の共同 RL を介してエージェントに次のツール呼び出しを予測するよう教える

大規模な言語モデルのエージェントは、ツール呼び出しの結果を待つのにかなりの時間を費やすことがよくあります。ツール呼び出しの投機では、エージェントの次のツール呼び出しを予測して、その予測がエージェントの最終的なツール呼び出しと一致する場合に事前実行することで、このレイテンシーを隠すことができますが、既存のスペキュレーターは通常、展開されたエージェント自体の動作とあまり整合していない個別のドラフト モデルまたはキャッシュされたトレースです。我々は、この投機者とエージェントのギャップを特定し、ターゲット エージェント自体が強力なネクストコール投機者であることを示します。これは、同じモデル内でエージェントと投機者を統合するという、よりシンプルな設計を示しています。この論文では、プレフィックス KV キャッシュを完全に再利用して、エージェント モードでタスクを解決し、スペキュレーター モードで部分的な軌跡から次のツール呼び出しを予測する単一モデルである自己投機エージェントを紹介します。パフォーマンスを低下させることなくこのデュアルモード エージェントを有効にするために、エージェント自身のロールアウトから推測ターゲットを導き出し、エージェントと推測者の更新を交互に行う、エージェントと推測者の共同強化学習方法を提案します。エージェント タスクの成功を維持しながら、エージェント検索 QA と会話型ツールを使用するエージェント タスク全体で、私たちの方法は平均次回ツール呼び出し Hit@1 を Qwen3-4B で 44.1 から 61.2 に、Qwen3.5-4B で 48.9 から 66.3 に改善しました。

原文 (English)

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the agent's eventual tool call, but existing speculators are typically separate draft models or cached traces that are poorly aligned with the deployed agent's own behavior. We identify this speculator-agent gap and show that the target agent itself is a strong next-call speculator. This points to a simpler design: unifying the agent and speculator within the same model. In this paper, we introduce the self-speculating agent, a single model that both solves tasks in agent mode and predicts its next tool call from partial trajectories in speculator mode, fully reusing prefix KV cache. To enable this dual-mode agent without degrading performance, we propose a joint agent-speculator reinforcement learning method, which derives speculation targets from the agent's own rollouts and alternates agent and speculator updates. Across agentic search QA and conversational tool-use agentic tasks, our method improves average next tool-call Hit@1 from 44.1 to 61.2 for Qwen3-4B and from 48.9 to 66.3 for Qwen3.5-4B, while preserving agent task success.

13:00 JST研究/論文

大規模衛星スケジューリングへの応用によるオンライン学習と反復価格設定による分散制約の最適化

分散制約最適化問題 (DCOP) は、限られた通信環境下で分散意思決定を行うための一般的なフレームワークを提供しますが、現実世界のインスタンスの多くは大きすぎてモノリシックに解決できません。私たちはこの課題に 2 つの相補的な方向から取り組みます。私たちは DCOP と潜在的なゲームの間の関係を再考し、均衡を見つけるための最新のオンライン学習アルゴリズムを DCOP に適応させます。これらのアルゴリズムが代表的な不完全な DCOP アルゴリズムと競合することを示します。次に、大規模な分散衛星スケジューリングを動機とする大規模 DCOP の分解フレームワークに移ります。我々は、DCOP を 2 つの相互作用するサブ問題、つまりタスク割り当てのための高レベルのメタ DCOP と、スケジューリングのための独立したローカル最適化問題に分離する新しいフレームワークを提案します。 2 つのレベルを結合するために、ローカル オプティマイザーからのフィードバックを使用してメタレベルのユーティリティを更新する新しい反復的な価格設定方法を開発します。当社のオンライン学習方法と反復的な価格設定フレームワークを組み合わせることで、現実世界の分散型衛星のスケジューリング問題のインスタンスでほぼ最適なパフォーマンスが得られ、最先端のベースラインの場合は 87% であるのに対し、観測リクエストの 99% 以上が満たされます。

原文 (English)

Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling

Distributed constraint optimization problems (DCOPs) provide a popular framework for distributed decision making under limited communication, but many real-world instances are too large to solve monolithically. We address this challenge from two complementary directions. We revisit the connection between DCOPs and potential games, and adapt modern online learning algorithms for equilibrium finding to DCOPs. We show that these algorithms are competitive with representative incomplete DCOP algorithms. We then turn to decomposition frameworks for large-scale DCOPs, motivated by large-scale decentralized satellite scheduling. We propose a new framework that separates a DCOP into two interacting subproblems: a high-level meta-DCOP for task allocation, and independent local optimization problems for scheduling. To couple the two levels, we develop a novel iterative pricing method that updates the meta-level utilities using feedback from the local optimizers. Combining our online learning methods with our iterative pricing framework, we obtain near-optimal performance on real-world decentralized satellite scheduling problem instances, fulfilling over 99% of observation requests compared with 87% for state-of-the-art baselines.

13:00 JSTLLM/生成AIエージェント

HiSkill: 階層型スキル グラフによる LLM エージェントの強化

スキルは、大規模言語モデル (LLM) エージェントが長期的な対話型タスクで過去の経験を再利用できるようにするための重要な抽象化になっています。しかし、既存のスキルへの軌道の手法では、独立して保存および取得される高レベルのテキスト スキルのフラットなコレクションが生成されることが多く、スキルの関係が十分に活用されず、高レベルのスキルと実行可能なアクションの間にギャップが維持されます。この論文では、スキル ノード、AtomicOp ノード、および型付きエッジを備えた有向グラフにインタラクションの軌跡を編成する階層型スキル グラフ フレームワークである HiSkill を提案します。具体的には、このグラフは、再利用可能な高レベルのスキルを実行可能なアクション テンプレートと結び付けるとともに、それらの間の分解、時間的遷移、互換性、サポート、および回復関係もキャプチャします。推論時に、HiSkill はコンパクトなタスク関連のサブグラフを取得し、サブグラフに基づくタスクの実行を実行します。ここで、シンボリック タスクの状態、アクティブなスキル、および取得されたサブグラフによって、LLM エージェントがスキルの切り替え、AtomicOps の選択、および実行可能なアクションの実行を繰り返し実行するようにガイドされます。 3 つのインタラクティブ環境での実験では、HiSkill が推論トークンの消費を削減しながら最先端のベースラインを上回るパフォーマンスを示し、階層型スキル グラフを通じて高レベルのスキルと実行可能なアクション基礎を橋渡しする効果を実証しています。データとコードは https://github.com/BUPT-GAMMA/HiSkill で入手できます。

原文 (English)

HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. However, existing trajectory-to-skill methods often produce flat collections of high-level textual skills that are stored and retrieved independently, leaving skill relations underutilized and maintaining a gap between high-level skills and executable actions. In this paper, we propose HiSkill, a hierarchical skill graph framework that organizes interaction trajectories into a directed graph with skill nodes, AtomicOp nodes, and typed edges. Specifically, the graph connects reusable high-level skills with executable action templates, while also capturing decomposition, temporal transition, compatibility, support, and recovery relations among them. At inference time, HiSkill retrieves a compact task-relevant subgraph and performs subgraph-guided task execution, where a symbolic task state, an active skill, and the retrieved subgraph guide the LLM agent to switch skills, select AtomicOps, and ground executable actions iteratively. Experiments on three interactive environments show that HiSkill outperforms state-of-the-art baselines while reducing inference token consumption, demonstrating the effectiveness of bridging high-level skills and executable action grounding through a hierarchical skill graph. Our data and code is available at https://github.com/BUPT-GAMMA/HiSkill.

13:00 JSTLLM/生成AIエージェント研究/論文

ベイジアン ネットワークを使用した LLM ベースのマルチエージェント システムの実行時不確実性モニタリング

この論文では、大規模言語モデル (LLM) に基づくマルチエージェント システム (MAS) が、特に不確実性の定量化に焦点を当てて保険数理リスク モデリングをどのようにサポートできるかを調査します。保険数理ワークフローは、信頼性の低い出力が不正確なリスク評価、不公平な価格設定、規制違反につながる可能性がある、一か八かの意思決定支援環境を表しています。 LLM の確率的な性質やエージェント間の依存関係によってもたらされる不確実性に対処するために、専門のエージェントが中央ハブの下でデータの準備、モデリング、レビュー、および説明のタスクを実行するマルチエージェント フレームワークが提案されています。主な貢献は、トークンレベルの対数確率とベイジアン ネットワークを使用した、不確実性の伝播に対する新しいアプローチです。重要なのは、ログの確率は、正確性やタスクの成功の直接の確率としては扱われないということです。代わりに、長さで正規化された対数確率の要約が、ベイジアン ネットワークに組み込まれる前に、調整されたタスク レベルの信頼推定値に変換されます。結果は、このフレームワークがベースラインの保険数理パフォーマンスを再現しながら、ワークフローの安定性と実行時の不確実性の伝播についてのさらなる洞察を提供することを示しています。

原文 (English)

Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focus on uncertainty quantification. Actuarial workflows represent a high-stakes decision-support setting where unreliable outputs may lead to incorrect risk assessment, unfair pricing, and regulatory non-compliance. To address uncertainty introduced by the probabilistic nature of LLMs and dependencies between agents, a multi-agent framework is proposed in which specialised agents perform data preparation, modelling, review, and explanation tasks under a central hub. The main contribution is a novel approach to uncertainty propagation using token-level log-probabilities and a Bayesian Network. Importantly, log probabilities are not treated as direct probabilities of correctness or task success. Instead, length-normalised log-probability summaries are transformed into calibrated task-level confidence estimates before incorporation into the Bayesian Network. Results show that the framework reproduces baseline actuarial performance while providing additional insight into workflow stability and runtime uncertainty propagation.

13:00 JSTエージェント

ハーネスエンジニアリングによるセキュリティ制御の配布

AI コーディング エージェントは歴史的なスピードで導入されていますが、依然としてセキュリティとリスクへの懸念が、エージェント AI を組織全体に拡張する際の主な障壁となっています。コーディング エージェントの既存のセキュリティ制御はエンジニアリング チームに体系的に配布されておらず、ベンダー ネイティブのソリューションでは、すべての展開コンテキストに適合しない可能性があるエコシステムの依存関係が導入されます。このペーパーでは、市販の AI コーディング エージェントに既製のセキュリティ制御を実装し、カスタム エージェント ハーネスを介して分散ユーザー ベースに拡張できるかどうかを調査します。 OWASP Top 10 for Agentic Applications から派生した 23 のテスト スイートを使用して、4 つのエージェント構成 (コントロールありとなしの 2 つの商用エージェント、ベースライン ハーネス、およびセキュリティ強化されたハーネス) に段階的なテスト手法が適用されました。 Pi エージェント ハーネス上に構築された分散可能ハーネスである SHarD (Secure Harness Distribution) は、商用エージェントに直接インストールするのと同等の有効性を維持しながら、OS サンドボックス、スキル スキャン、ツール制限の 3 つのカテゴリのセキュリティ制御を 1 つのインストール コマンドで埋め込み、配布できることを実証しました。 SHarD は 100\% の調整スコアを達成し、安全に構成された最良の商用エージェントと一致し、どのテスト カテゴリでも後退はありませんでした。注目すべき観察には、モデルの非決定性が一貫性のないセキュリティ結果を生み出すこと、および自律エージェントの動作が OS サンドボックスによって直接緩和される方法でシステム境界を越える可能性があるという証拠が含まれます。コントロール ハーネスの適合性フレームワークに向けた初期特性が提案され、将来の調査のために 3 番目の研究課題が特定されます。

原文 (English)

Distributing Security Controls Through Harness Engineering

AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI across organizations. Existing security controls for coding agents are not systematically distributed to engineering teams, and vendor-native solutions introduce ecosystem dependencies that may not suit every deployment context. This paper investigates whether off-the-shelf security controls can be implemented on commercial AI coding agents and scaled to a distributed user base via a custom agent harness. A phased testing methodology was applied across four agent configurations --- two commercial agents with and without controls, a baseline harness, and a security-hardened harness --- using a 23-test suite derived from the OWASP Top 10 for Agentic Applications. SHarD (Secure Harness Distribution), a distributable harness built on the Pi agent harness, demonstrated that three categories of security controls --- OS sandboxing, skill scanning, and tool restriction --- can be embedded and distributed via a single install command while retaining equivalent efficacy to direct installation on commercial agents. SHarD achieved an adjusted score of 100\%, matching the best securely configured commercial agent, with no regression across any test category. Notable observations include evidence that model non-determinism produces inconsistent security outcomes and that autonomous agent behavior can cross system boundaries in ways that OS sandboxing directly mitigates. Initial characteristics toward a control harness fitness framework are proposed, and a third research question is identified for future investigation.

13:00 JSTエージェントビジネス/資金調達研究/論文

メシエ: クロスベンチマーク エージェント評価用の高解像度コーパス

インタラクティブな環境での AI エージェントの評価は、断片化されたタスク、足場、検証者、スコアリング ルールによって妨げられます。既存の取り組みは狭い設定に焦点を当てており、規模が限られたままであるか、費用のかかる再実行が必要であり、経験的記録の多くは比較できないものとなっています。 Messier は、30 のベンチマーク、714 のエージェント、11,891 のタスク、74,205 の検証者にまたがる 957,253 レコードの統合コーパスです。 Messier は公開ベンチマーク スコアを統合し、最近の法律ベンチマークを含む、過小評価されている 6 つの専門的および科学的領域にわたる 5 つのエージェントの実行でそれらを補完します。各レコードはモデル、足場、環境、タスク、検証者、集計ルールによって標準化されており、職業分析および業界分析用の SOC/NAICS 分類が使用されます。このコーパスを使用すると、ベンチマークの種類によってフロンティアの進歩が均一ではなく、「関数呼び出し」が飽和し、「プログラミング」が最も速く改善し、「エンタープライズ ワークフロー」が依然として最も困難であることがわかります。さらに、反事実的な再スコアリングは、複数の検証者タスクにおける厳格なオールパス集計が進行状況を曖昧にし、エージェントのランキングを人為的に変更する可能性があることを示しています。これらの標準化された記録から、Spearman \r{ho} = 0.81 での Epoch の評価能力指数ランキングと一致する能力スケールを導き出し、ドメイン、職業、アクション スペース、または検証者のタイプによって特化することができます。 Messier は、エージェントの機能スケーリング、ベンチマーク監査、評価失敗の詳細な分析のための、再利用可能な基礎的なインフラストラクチャを提供します。

原文 (English)

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six underrepresented professional and scientific domains, including a recent legal benchmark. Each record is standardized by model, scaffold, environment, task, verifier, and aggregation rule, with SOC/NAICS classifications for occupational and industry analysis. Using this corpus, we show frontier progress is uneven across benchmark types, with "function calling" saturated, "programming" improving the fastest, and "enterprise workflows" remaining the most challenging. Furthermore, counterfactual rescoring shows that strict all-pass aggregation in multi-verifier tasks can obscure progress and artificially alter agent rankings. From these standardized records, we derive capability scales that align with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.81 and can be specialized by domain, occupation, action space, or verifier type. Messier provides a foundational, reusable infrastructure for agent capability scaling, benchmark auditing, and fine-grained analysis of evaluation failures.

13:00 JSTエージェントビジネス/資金調達

インタラクティブ報酬エージェント: 環境状態検証による GUI タスクの評価

グラフィカル ユーザー インターフェイスのタスク評価は、GUI エージェントがユーザーの指示を正常に完了したかどうかを判断することを目的としています。自動化された GUI タスク評価は、評価結果がテスト時のスケーリングとトレーニング後の両方に対する報酬シグナルとして機能する可能性があるため、ますます注目を集めています。ただし、信頼性の高い GUI タスクの評価は、依然として課題が残っています。その判断には、実行軌跡のスクリーンショットを超えて、システム構成、ファイル データ、アプリケーション設定などの環境状態へのアクセスが必要になることが多いためです。この論文では、実行後の環境から証拠を取得して検証するための提案-その後検証フレームワークに基づいた対話型報酬エージェント (IRA) を提案します。タスクの指示と GUI エージェント実行後の GUI 環境が与えられると、IRA はまずタスクの完了条件を提案し、次にシステム ツール、アプリケーション ツール、および GUI ツールを呼び出してそれらを検証します。この設計では、可視インターフェイスと環境状態の両方からの証拠を対話型プロセスで組み合わせます。さらに、10 の Ubuntu デスクトップ アプリケーション カテゴリにわたる 321 の GUI タスクの軌跡のベンチマークである GUI-RewardBench を紹介します。実験によると、IRA は GUI-RewardBench で 86.9% の精度を達成し、既存の評価者のベースラインを上回るパフォーマンスを示しました。さらに IRA を GUI エージェントの強化学習に適用し、OSWorld の成功率 34.0% を達成しました。これは、IRA が GUI エージェントのトレーニングに効果的な報酬シグナルを提供できることを示しています。

原文 (English)

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.

13:00 JSTエージェント

自律型ネットワークにおける標準化されたクロスベンダー エージェント ツールの信頼管理に向けて

自律ネットワーク レベル 4 ~ 5 では、AI エージェントが人間の監視なしにベンダーの境界を越えてツールを呼び出す必要がありますが、既存の管理標準にはベンダー間の信頼を可視化するための標準化されたメカニズムがありません。ベンダー B のツールが侵害されると、ベンダー A のエージェントは信頼性の低下を認識せずにそのツールを呼び出し続け、連鎖的なサービスへの影響を引き起こします。エージェント ツール信頼管理用に提案されている 3GPP NRM 情報モデルである AgentToolMO を紹介します。このモデルは、証明可能な段階的施行を備えた正式に定義されたトラスト ステート マシン、有界収束による減衰カスケード伝播、既存の管理サービス (MnS) インターフェイスを介したベンダー間信頼通知、および NRM 依存関係グラフ トラバーサルを介した遡及的影響評価で構成されます。マルチベンダー トポロジにわたるシミュレーション ベースの評価では、標準化されたクロスベンダー通知により、爆発半径が数時間スケールの未検出の伝播から、MnS 通知配信によって制限されたほぼリアルタイムの封じ込めまで減少し、制限された反復でカスケード コンバージェンスが保証され、ベンダー ドメイン全体にわたるサブリニアな通知スケーリングが示されています。このフレームワークは既存の 3GPP 管理インフラストラクチャ内で動作し、既存のプロトコルを活用して、信頼できるマルチベンダー自律ネットワーク管理のための標準化経路を提供します。

原文 (English)

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it -- unaware of the trust degradation -- causing cascading service impact. We present AgentToolMO, a proposed 3GPP NRM information model for agent tool trust management. The model comprises: a formally defined trust state machine with provable graduated enforcement, damped cascade propagation with bounded convergence, cross-vendor trust notifications via existing Management Services (MnS) interfaces, and retroactive impact assessment through NRM dependency graph traversal. Simulation-based evaluation across multi-vendor topologies shows that standardized cross-vendor notifications reduce blast radius from hours-scale undetected propagation to near-real-time containment bounded by MnS notification delivery, with cascade convergence guaranteed in bounded iterations and sub-linear notification scaling across vendor domains. The framework operates within existing 3GPP management infrastructure, leverages existing protocols, and provides a standardization pathway for trustworthy multi-vendor autonomous network management.

13:00 JST研究/論文

ペネロペ: 効率的な構造化推論のための局所的潜在再発

複雑な構造の推論タスクでは追加の計算が必要になることがよくありますが、現在の言語モデルは主にパラメーターのスケールを増やすか、中間ステップを思考連鎖 (CoT) トークンとしてシリアル化することによって計算を取得します。前者はトレーニングとデプロイメントのコストを増加させますが、後者は推論計算を自己回帰出力長に結び付けます。選択したデコーダ間隔に反復計算を局所化する、事前トレーニングされたデコーダ専用の Transformer 用の効率的な潜在推論フレームワークである Penelope を紹介します。下位デコーダ プレフィックスは、問題条件付き境界メモリを構築するために 1 回評価され、その後、応答生成前に、時間変調された GRU ダイナミクスと反復的な読み出し状態を通じて反復的に改良されます。進歩的な CoT から潜在的なカリキュラムは、目に見える推論をこの内部再帰パスに移し、完全なデコーダーを繰り返し実行したり、長い中間トレースを生成したりすることなく、追加の計算を潜在空間に割り当てることができます。オープンソースの構造化推論ベンチマークの実験では、検証で選択された潜在バジェットで、Penelope が、測定された推論レイテンシーを削減しながら、確立された潜在推論モデルと比較して競合する精度を達成することが示されています。これらの結果は、潜在リファインメントを狭いデコーダ間隔に局所化することができ、目に見える長い推論トレースを生成することなくフルデコーダの繰り返し実行を削減し、デコーダのみの Transformer モデルに実用的な精度と効率のトレードオフを提供できることを示しています。

原文 (English)

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while reducing measured inference latency. These results show that latent refinement can be localized to a narrow decoder interval, reducing repeated full-decoder execution without generating a long visible reasoning trace and providing a practical accuracy-efficiency tradeoff for decoder-only Transformer models.

13:00 JST研究/論文

dtControl2+$\varepsilon$: デシジョン ツリーを介した MDP での説明可能性のためのトレーディングの最適性

過去 10 年間にわたり、dtControl2 を最新のツールとして使用して、説明可能な方法でコントローラー (別名ポリシー) を表すためにデシジョン ツリーが使用されてきました。ただし、大規模なシステムや特殊なケースが多数あるシステムの場合、そのような表現でも複雑すぎて人間が理解できない傾向があります。残念ながら、決定木のサイズを減らすことは簡単ではありません。重要なケースが 1 つだけ欠けていると、コントローラーが正しくなくなる可能性があります。私たちは、マルコフ決定プロセスの設定でこの問題に取り組み、「$\varepsilon$」機能によって dtControl2 を拡張します。許容される不精度 $\varepsilon \geq 0$ を考慮して、$\varepsilon$ の最適性を保証しながら、コントローラーの本質を抽出して、より小さな決定木を構築します。これにより、制御可能な量の詳細を省略して、調整可能に単純な説明を提供できるようになります。私たちのツールは、最新技術よりも桁違いに小さいデシジョン ツリーを構築します。

原文 (English)

dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees

Over the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a current state-of-the-art tool. However, for systems that are large or have many corner cases, even such representations tend to be too complex and not human-comprehensible. Unfortunately, reducing the size of the decision tree is not straightforward, as missing just a single crucial case might result in an incorrect controller. We tackle this issue in the setting of Markov decision processes, extending dtControl2 by "$\varepsilon$" functionality: Given an allowed imprecision $\varepsilon \geq 0$, we construct a smaller decision tree, distilling the essence of the controller, while still guaranteeing its $\varepsilon$-optimality. This enables us to provide tunably simpler explanations, omitting a controllable amount of detail. Our tool constructs decision trees that are orders of magnitude smaller than the state of the art.

13:00 JSTLLM/生成AI

不規則な臨床時系列での質問応答のための費用対効果の高いマルチモーダル LLM 推論フレームワーク

不規則な臨床時系列 (ICTS) に対する質問応答 (QA) は、幅広い医療アプリケーションで極めて重要な役割を果たしています。最近のマルチモーダル時系列大規模言語モデル (LLM) は、汎用時系列 QA においてかなりの有望性を示していますが、臨床観察のスパース性、非同期性、および不規則なサンプリング パターンをモデル化するための機能は依然として不十分です。このギャップを埋めるために、ICTS データを使用した質問応答のための費用対効果の高いマルチモーダル LLM 推論フレームワークである ClinPRISM を提案します。まず、さまざまな時間スケールでまばらな臨床証拠を捕捉するために、不規則性を認識するマルチスケール エンコーダーを考案します。次に、これらのスケール全体で表現を統合し、少数の LLM 互換トークンに圧縮するための時間的証拠蒸留器を提案します。さらに、不規則な軌跡を LLM のテキスト埋め込み空間に順次位置合わせする漸進的位置合わせ戦略を導入します。トレーニングを促進するために、11 のタスクにわたる 41,000 の指示調整インスタンスとともに、マルチスケールの説明と組み合わせた 30,000 の臨床時系列を構築します。 ClinPRISM は、40 億パラメータの LLM バックボーンを使用して、16 個の時系列トークンのみを使用し、質問あたり 0.15 秒の平均推論レイテンシーを達成しながら、ホールドアウト評価ベンチマークで最先端のパフォーマンスを達成します。

原文 (English)

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture sparse clinical evidence at diverse temporal scales. Then, we propose a temporal evidence distiller to integrate representations across these scales and compress them into a small number of LLM-compatible tokens. Moreover, we introduce a progressive alignment strategy that sequentially aligns the irregular trajectories with the LLM's textual embedding space. To facilitate training, we construct 30,000 clinical time series paired with multi-scale descriptions, together with 41,000 instruction-tuning instances spanning 11 tasks. Using a 4-billion-parameter LLM backbone, ClinPRISM achieves state-of-the-art performance on the held-out evaluation benchmark while using only 16 time-series tokens and achieving an average inference latency of 0.15 seconds per question.

13:00 JST研究/論文

複数の倉庫の在庫配分におけるオペレーション リサーチ処方選択のための大規模言語モデル

複数の倉庫の在庫割り当ては通常、混合整数計画法 (MIP) 問題として定式化されますが、需要の集中、在庫の不均衡、補充規模、サービスの制約、および予測の変動性によって引き起こされる異種インスタンス レベルの体制に一貫して一致する単一の定式化はありません。私たちは、この問題をインスタンスごとのオペレーション リサーチ (OR) 定式化の選択として研究します。この場合、各割り当てインスタンスは、候補 OR エキスパート ライブラリからのソルバー実行可能定式化に割り当てられます。我々は、OR 定式化選択のためのソルバー主導の大規模言語モデル (LLM) フレームワークを提案します。このフレームワークでは、各 OR エキスパートが、個別の割り当て優先順位をエンコードする MIP 定式化に対応します。セレクターをトレーニングするために、フレームワークは最初にスキーマ学習用にバランスの取れたエキスパート条件付き教師付きファインチューニング (SFT) レコードを構築し、次に履歴インスタンスに対する MIP ソルバー評価を使用して、ソルバーが評価した割り当て品質ギャップをマージン加重アイデンティティ優先最適化 (IPO) プリファレンスと、サンプルされた応答に報酬を割り当てるグループ相対ポリシー最適化 (GRPO) 中の報酬検索用のインスタンスごとのエキスパート スコア メタデータに変換します。中国最大の電子小売業者の 1 つである JD$\mathord{.}$com の複数倉庫在庫割り当てインスタンスに関する実験では、GRPO が SFT+IPO セレクターと比較して専門家による選択精度を大幅に向上させ、さらに重要なことに、優先トレーニングされたセレクターや最適な固定定式化の両方よりも高い実現された割り当て品質を生み出すことが実証されました。 GRPO を使用すると、Hit Ratio@1 と Hit Ratio@2 が 21.45% から 50.42%、70.47% から 82.31% に増加します。結果として得られたセレクターは、現在のベースラインを上回る 12.57 パーセント ポイントの割り当て精度の向上を達成し、SFT+IPO セレクターと最良の固定 OR エキスパートの両方を上回り、事後のオラクルとのギャップを 4.85 パーセント ポイントに縮小しました。

原文 (English)

Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation

Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consistently matches heterogeneous instance-level regimes induced by demand concentration, inventory imbalance, replenishment scale, service constraints, and forecast volatility. We study this issue as instance-wise operations research (OR) formulation selection, where each allocation instance is assigned to a solver-executable formulation from a candidate OR expert library. We propose a solver-guided large language model (LLM) framework for OR formulation selection, in which each OR expert corresponds to a MIP formulation encoding a distinct allocation priority. To train the selector, the framework first constructs balanced expert-conditioned supervised fine-tuning (SFT) records for schema learning, and then uses MIP solver evaluation on historical instances to convert solver-evaluated allocation-quality gaps into margin-weighted identity preference optimization (IPO) preferences and per-instance expert-score metadata for reward lookup during group relative policy optimization (GRPO) to assign rewards to sampled responses. Experiments on multi-warehouse inventory allocation instances from JD$\mathord{.}$com, one of China's largest e-retailers, demonstrate that GRPO substantially improves expert-selection accuracy relative to the SFT+IPO selector and, more importantly, produces higher realized allocation quality than both the preference-trained selector and the best fixed formulation. With GRPO, Hit Ratio@1 and Hit Ratio@2 increase from 21.45% to 50.42% and from 70.47% to 82.31%. The resulting selector achieves an allocation accuracy gain of 12.57 percentage points over the incumbent baseline, outperforming both the SFT+IPO selector and the best fixed OR expert, and reduces the gap to the ex-post oracle to 4.85 percentage points.

13:00 JST研究/論文

CHARM: ゼロショット転送のための階層コンテキスト モデリングを備えたマルチモーダル グラフ基盤モデル

グラフ基盤モデル (GFM) は、グラフ ドメインやタスク間で知識を伝達するための有望なパラダイムとして浮上しています。現実世界のグラフはノードをテキスト、画像、その他のモダリティに関連付けるため、複雑なエンティティや関係を表現するためにマルチモーダル グラフが不可欠になります。さらに、ラベルを収集し、新しいグラフ ドメインごとにモデルを適応させるのはコストがかかり、実行不可能なことが多いため、ゼロショット転送の動機になります。残念ながら、マルチモーダル グラフでのゼロショット転送はまだ研究されていません。既存の GNN ベースのグラフ基盤モデルは通常、ダウンストリームの適応を必要としますが、LLM ベースのグラフ手法は主に単一ドメイン内の単峰性グラフまたはタスクに対処します。この設定には 2 つの重要な課題があります。まず、モデルは、伝達可能なクロスモーダル関係をキャプチャしながら、個々のモダリティからの知識を一般化する必要があります。第二に、ターゲットドメインの微調整がないと、ノード表現はドメイン固有の構造やモダリティ固有の特性と絡み合ったままとなり、目に見えないドメインでの共有概念が曖昧になってしまいます。これらの課題に対処するために、ゼロショット転送のための階層コンテキスト モデリングを備えたマルチモーダル グラフ基盤モデルである CHARM を提案します。 CHARM は、分離された生のノードを、マルチモーダル セマンティクスとクロスモーダル関係をキャプチャする階層グラフ コンテキストに置き換えます。これらのコンテキストは、ドメイン固有のノード パターンを共有の高レベルの概念にマップし、ターゲット ドメインの監視や適応への依存を軽減します。モダリティを認識したグラフ コンテキスト エンコーダーは、マルチモーダル情報をグラフ構造と統合し、結果の表現を大規模な言語モデルのグラフ トークンに変換します。実験では、ゼロショットのマルチモーダル グラフ タスクで一貫した改善が見られました。

原文 (English)

CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer

Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world graphs associate nodes with text, images, and other modalities, making multimodal graphs essential for representing complex entities and relations. Moreover, collecting labels and adapting models for every new graph domain is costly and often infeasible, motivating zero-shot transfer. Unfortunately, zero-shot transfer on multimodal graphs remains underexplored. Existing GNN-based graph foundation models typically require downstream adaptation, whereas LLM-based graph methods mainly address unimodal graphs or tasks within a single domain. This setting presents two key challenges. First, models must generalize knowledge from individual modalities while capturing transferable cross-modal relations. Second, without target-domain fine-tuning, node representations remain entangled with domain-specific structures and modality-specific characteristics, obscuring shared concepts in unseen domains. To address these challenges, we propose CHARM, a multimodal graph foundation model with hierarchical context modeling for zero-shot transfer. CHARM replaces isolated raw nodes with hierarchical graph contexts that capture multimodal semantics and cross-modal relations. These contexts map domain-specific node patterns to shared high-level concepts, reducing reliance on target-domain supervision or adaptation. A modality-aware graph context encoder integrates multimodal information with graph structure and converts the resulting representations into graph tokens for a large language model . Experiments show consistent improvements on zero-shot multimodal graph tasks.

13:00 JST研究/論文

理想化された AI レース実験で遅れをとると安全でない開発が引き起こされる

技術競争では、速度と安全性の間に緊張が生じます。危険な開発が有害な場合でも、関係者は競合他社よりも速く動くことで利益を得られる可能性があります。これは人工知能 (AI) に関する議論で顕著であり、そこでは、リスクが高く、安全性をあまり意識しない開発を奨励するために競争圧力が議論されることがよくあります。私たちは、ペアの参加者が不確実な時間軸の下で安全な開発と安全でない開発のどちらかを繰り返し選択する、理想的な AI レースに基づいたフレーム化された行動実験を使用してこれを研究します。安全でない開発はより迅速な進歩とより高い即時利益をもたらしましたが、プライベートリスクは治療固有の最大値である 10\%、60\%、または 90\% まで蓄積されました。レースの競争構造は一定に保たれ、この最大リスクだけが変化しました。事前に登録されたリスクレベル間の比較も、導き出されたリスク選好の役割もデータによって裏付けられていませんでした。その代わりに、タスクの繰り返し構造を動機とした探索的分析では、危険な行動はリスク選好よりもむしろレースの戦略的状態の進化によって形作られることが示されている。参加者は相手が安全でない行動をした後に危険な行動を選択する可能性が高く、先を行くと危険なプレイが減り、遅れをとれば危険なプレイが増加し、最初のラウンドの選択はその後の行動を予測する。これらの効果を解釈するために、我々は 4 つの戦略 (常に安全、常に安全ではない、条件付きで安全、条件付きで反社会的安全) を備えた縮小進化モデルを導入します。これは治療効果を再現し、条件付きの安全でない行動が競争力学によってどのように有利に働くかを示します。この実験とモデルを総合すると、危険な開発は、リスク選好のみからではなく、初期の行動の勢い、相手の行動、遅れをとることへの恐怖から生じる可能性があることを示しており、政策は個人のリスクだけではなく、競争圧力を軽減し、AI開発における協力を促進することに焦点を当てるべきであることを示唆している。

原文 (English)

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.

13:00 JST画像/動画生成エージェント研究/論文

デスクトップデルタベンチ: コンピュータ使用モデルはデスクトップ GUI の遷移を理解していますか?

コンピュータ使用エージェント (CUA) は、長期にわたるタスクを完了するためにデスクトップ GUI を通じて動作することが増えています。現在のベンチマークは主に、最終タスクの成功または単一フレームの接地を測定します。どちらも、古い観察を拒否し、進捗を確認し、失敗から回復するために重要な、アクションによって生成される因果関係のあるタスク関連の遷移をモデルが再構築できるかどうかを分離するものではありません。これが難しいのは、推論、リモート入力、アプリのレンダリング、スクリーンショットのキャプチャが非同期であるためです。次の観察が遅れたり、遮られたり、一時的になったり、無関係になったりして、進行状況と誤って読み取られ、その後の計画に持ち込まれる可能性があります。デスクトップ デルタ ベンチ (DDB) は、約 15 のアプリケーションと 50 のタスク ドメインにわたる新しいマルチアプリ Linux の軌跡から人間が検証した 2,013 個のインスタンスを含むオフラインのステップレベルのベンチマークです。 DDB トラジェクトリは、2 つの相補的なタスクを通じて 3 つの障害次元 (状態検証、ソース追跡、コンテキスト認識制御) をターゲットとしています。つまり、クロストラジェクトリデコイを含む 105 個を含む 463 個の 3 フレーム時間順序付けインスタンス、および 5 つのアクションとそのペイロードからラベル付けされた 1,550 個のビフォーアフター ペアです。 32 の順序設定と 16 のシングルアクション設定にわたって 8 つのクローズドおよびオープンソース モデル ファミリを評価し、一貫したギャップを観察しました。順序付けは依然として飽和していません。デコイ以外の最高の完全一致率とデコイの完全一致率は 65.1% と 65.7% です。タスク コンテキストにより、おとりの識別は 6.9 パーセント ポイント向上しますが、おとり以外の完全一致は 2.2 ポイント減少します。エラー分析により、提示された A-B-C の順序の系統的なコピーが明らかになります。単一アクションの結果は、アクション ファミリを特定するよりも推論する方が難しいことを示しています。クリック F1 は 0.96 に対してドラッグは 0.76 ですが、認識されたドラッグは一般的に適切に位置特定されています。したがって、DDB は、GUI のグラウンディングと最終タスクの成功の間に欠けている診断層を埋めることによってエンドツーエンドのベンチマークを補完し、デスクトップ CUA の検証、信頼性、および回復に対する目標を絞った改善を可能にします。

原文 (English)

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50 task domains. DDB trajectories targets 3 failure dimensions- state verification, source tracking, and context-aware control- through 2 complementary tasks: 463 3-frame temporal-ordering instances, including 105 with a cross-trajectory decoy, and 1,550 before-after pairs labeled from 5 actions + its payload. We evaluate 8 closed and open-source model families across 32 ordering and 16 single-action settings, observing consistent gaps. Ordering remains unsaturated: best non-decoy and decoy exact-match rates are 65.1% and 65.7%. Task context improves decoy identification by 6.9 percentage points but reduces non-decoy exact match by 2.2 points; error analysis reveals systematic copying of the presented A-B-C order. Single-action results show that inferring the action family is harder than locating it: click F1 is 0.96 vs, 0.76 for drag, while recognized drags are generally localized well. DDB, thus, complements end-to-end benchmarks by filling the missing diagnostic layer between GUI grounding and final task success, enabling targeted improvements to desktop CUA verification, reliability, and recovery.

13:00 JST研究/論文

信頼できない著者、信頼できる回答: 忠実度評価された翻訳の計算法

プログラムに関する質問に答えるには、質問が決定可能な場所にプログラムを移動します。そのような動きはすべて翻訳であり、すべての翻訳には間違いが含まれます。私たちは翻訳をグラフとして研究します -- 多くの言語、少数の推論対象、正直に異なる信頼性の独立して構築されたルート -- そして、それに微積分を与えます。言語のペアは、方向性のある通勤正方形に近く (正確さは、過近似の特殊なケースです)、プログラムごとにチェック可能で、構成可能であり、ルートのコントラクトは、そのホップのコントラクト (保証クラス、方向、保持された観測可能性、測定されたコスト) のコンポーネントごとの一致です。 1 つの非対称性が信頼を組織します。証人による回答は、情報源で再生されることで自己証明されます。普遍的な答えは、グレード、独立した支店、再検査された証明書がコストを稼ぐ場所です。緩い望遠鏡を含む構成のコアは、リーン 4 で機械化されています。ハーディ・ガーディは、1 つのレジストリで出会う 2 つの平面として微積分を実装します。使用プレーンは宣言を読み取り、証拠を伴う回答を生成します。そのビルダーとその対象となるプレーヤーは両方とも LLM であり、構築によって信頼されていません。進化面はグラフを成長させます。満たされていない質問は需要として記録され、ペアは証拠によって推奨され、人間によって登録され、ラチェットによって以前のすべての評決が維持されます。答えは決して書きません。成長には決して答えはありません。無限に実行すると、ループは還元可能に決定可能なすべての質問に収束し、忠実度は上昇するだけです。私たちは、2026 年 7 月のスナップショット (コンストラクトごとの結合カバレッジ、2 つの ISA のデュアルルート分岐合意、ソースレベルの証人リプレイ、正式に検証されたチェッカーによって再検証された認定された到達不能性、ゲート自体のエスケープ率) を測定し、アーキテクチャが自身の作成者の作業で検出した欠陥を報告します。

原文 (English)

Untrusted Authors, Trusted Answers: A Calculus of Fidelity-Graded Translations

To answer a question about a program, move the program to where the question is decidable. Every such move is a translation, and every translation is a place to be wrong. We study translation as a graph -- many languages, a few reasoning targets, independently built routes of honestly different trustworthiness -- and give it a calculus: pairs of languages close commuting squares that are directional (exactness is the identity-embedding special case of over-approximation), checkable per program, and composable, a route's contract being the componentwise meet of its hops' contracts -- assurance class, direction, kept observables, measured cost. One asymmetry organizes trust: witness-carrying answers are self-certifying by replay at the source; universal answers are where grades, independent branches, and re-checked certificates earn their cost. The compositional core, lax telescope included, is mechanized in Lean 4. hurdy-gurdy implements the calculus as two planes meeting in one registry. The use plane reads declarations and produces evidence-carrying answers; its builders and its intended player are both LLMs, untrusted by construction. The evolution plane grows the graph: unmet questions are recorded as demand, pairs are recommended by evidence and registered by humans, and a ratchet keeps every prior verdict standing. Answers never write; growth never answers. Run indefinitely, the loop converges on every reducibly decidable question, at fidelity that only rises. We measure the July 2026 snapshot -- per-construct conjoined coverage, dual-route branch agreement for two ISAs, source-level witness replay, certified unreachability re-validated by a formally verified checker, escape rates for the gate itself -- and report the defects the architecture caught in its own authors' work.

13:00 JST研究/論文

サイバーフィジカルシステムにおける異常検出のためのドメイン事前正則化グラフモデリング

多変量センサー時系列の異常検出は、サイバーフィジカル システム (CPS) の産業監視にとって重要です。通常の動作からのわずかな逸脱でさえ、プロセスの中断を示す可能性があります。最近のグラフベースのアプローチは大幅な進歩を遂げていますが、ラベル付けされた異常がほとんどなく、正常なデータが限られている小規模な物理システムではしばしば困難を伴います。このような設定では、グラフベースのモデルは誤った相関を捕捉し、不安定なセンサー トポロジを生成する傾向があります。我々は、システム設計の知識をグラフ構築に組み込んだ予測ベースのフレームワークである DPR-GM (Domain-Prior- Regularized Graph Modeling) を提案します。 DPR-GM は、大規模言語モデル (LLM) を活用して、システム ドキュメントからセンサー ペア間の有向物理結合を抽出します。これは、センサー関係に対する構造ゲートとして機能するバイナリ ドメイン隣接行列としてエンコードされます。このゲートは、通常のトレーニング データから推定されたピアソン相関によって変調されます。異常スコアは、変動係数から導出されるセンサーレベルの信頼性によってさらに重み付けされます。すべてのグラフと重み付けコンポーネントはトレーニング前に固定されており、学習可能なパラメーターは追加されません。 SKAB ベンチマークでは、DPR-GM は、F1、AUROC、および AUPRC 全体でグラフベース、統計、深層学習のベースラインを上回っており、データが不足している CPS ではドメイン構造化グラフ事前分布が完全に学習されたトポロジに代わる実用的な代替手段であることを示しています。

原文 (English)

Domain-Prior-Regularized Graph Modeling for Anomaly Detection in Cyber-Physical Systems

Anomaly detection on multivariate sensor time series is critical for industrial monitoring of cyber-physical systems (CPS), where even subtle deviations from normal behavior can indicate process disruption. Recent graph-based approaches have made significant progress, but they often struggle in small-scale physical systems with scarce labeled anomalies and limited normal data. In such settings, graph-based models tend to capture spurious correlations and produce unstable sensor topologies. We propose DPR-GM (Domain-Prior-Regularized Graph Modeling), a forecasting-based framework that incorporates system design knowledge into graph construction. DPR-GM leverages a large language model (LLM) to extract directed physical couplings between sensor pairs from system documentation, which are encoded as a binary domain adjacency matrix serving as a structural gate over sensor relations. This gate is then modulated by Pearson correlations estimated from normal training data. The anomaly score is further weighted by sensor-level reliability derived from the coefficient of variation. All graph and weighting components are fixed prior to training and add no learnable parameters. On the SKAB benchmark, DPR-GM outperforms graph-based, statistical, and deep learning baselines across F1, AUROC, and AUPRC, showing that domain-structured graph priors are a practical alternative to fully learned topologies in data-scarce CPS.

13:00 JST研究/論文

量子ビット測定シミュレーションのための 1 ビット プロトコルのニューラル ネットワーク学習

通信の複雑さは、量子統計を再現するために必要な古典的なリソースを定量化するための自然なフレームワークを提供します。量子ビットの準備と測定のシナリオでは、任意の量子ビット状態と任意の量子測定を正確にシミュレートするには、2 つの古典ビットが必要かつ十分であることが示されています。ただし、この結果は、制限された測定値ファミリーが正確な 1 ビットの古典的近似を許容する可能性を排除するものではありません。ニューラル ネットワーク手順を使用して、単一ビットが特定の測定ファミリに対して高い平均精度を達成できることを実証します。私たちのニューラルネットワークのパフォーマンス分析により、正多面体を形成する要素など、均一に重み付けされた要素による対称測定が、この制限された通信に特に適していることが明らかになりました。ニューラルネットワークによって学習されたパターンを分析することにより、有限の情報的に完全な対称構成に対して非常に正確で、連続等方性測定の限界内で正確になる分析プロトコルを導き出します。

原文 (English)

Neural Network Learning of One-Bit Protocols for Qubit Measurement Simulation

Communication complexity provides a natural framework for quantifying the classical resources required to reproduce quantum statistics. In the qubit prepare-and-measure scenario, two classical bits have been shown to be necessary and sufficient to simulate arbitrary qubit states and arbi- trary quantum measurements exactly. However, this result does not exclude the possibility that restricted families of measurements may admit accurate 1-bit classical approximations. We use a neural network procedure to demonstrate that a single bit can achieve high average accuracy for specific measurement families. A performance analysis of our neural network reveals that symmet- ric measurements with uniformly weighted elements, such as those forming regular polyhedra, are particularly amenable to this restricted communication. By analyzing the patterns learned by the neural network, we derive an analytical protocol that is extremely accurate for finite information- ally complete symmetric configurations and becomes exact in the limit of a continuous isotropic measurement.

13:00 JSTLLM/生成AI画像/動画生成

DocAnnot -- GenAI を利用した自動アノテーションで重要な情報抽出データセットの作成を加速する

Key Information Extraction (KIE) は多くのドキュメント アプリケーションにとって不可欠ですが、トレーニング データセットの作成は従来、時間のかかる手動プロセスでした。 KIE データセットの生成を大幅に高速化するフレームワークである DocAnnot を紹介します。 DocAnnot は、ラベル値抽出に Large Vision Language Model (LVLM)、テキスト/境界ボックス検出に OCR、および新しい空間情報コンテキスト マッチング (SICM) アルゴリズムを活用します。 SICM は、空間関係と近接分析をテキストの一致と組み合わせることにより、ラベルと値の関連付けを改善します。 CORD ベンチマークと SROIE ベンチマークでフレームワークを評価し、それぞれ 0.679 と 0.846 の F1 スコアでアノテーションを自動生成する機能を実証しました。さらに、下流の KIE モデルを微調整するために自動アノテーション付きデータを使用する有効性を調査します。人間が注釈を付けたデータは依然として優れていますが、DocAnnot の出力のみでトレーニングされたモデルは、かなりのパフォーマンスを達成します (たとえば、LayoutLMv3 は CORD で 0.6765 の F1 スコアを達成しました)。これらの結果は、私たちのフレームワークが手作業への依存を大幅に軽減する一方で、人間の介入の必要性を完全に排除するわけではないことを示しています。ただし、レビュー担当者が効率的に出力を調整できるレベルまでプロセスを自動化することで、当社のシステムは、ゼロから手動でアノテーションを付けるよりもはるかに高い効率で、ほぼ完璧なアノテーションを可能にします。このアプローチは時間とコストを大幅に節約できるため、リソースに制約のある設定や迅速なモデル プロトタイピングに価値があります。

原文 (English)

DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation

Key Information Extraction (KIE) is vital for many document applications, but creating training datasets is traditionally a time-consuming manual process. We introduce DocAnnot, a framework that significantly accelerates KIE dataset generation. DocAnnot leverages a Large Vision Language Model (LVLM) for label value extraction, OCR for text/bounding box detection, and a novel Spatially Informed Contextual Matching (SICM) algorithm. SICM improves label-value association by combining spatial relationships and proximity analysis with textual matching. We evaluate our framework on the CORD and SROIE benchmarks, demonstrating its ability to auto-generate annotations with F1-scores of 0.679 and 0.846, respectively. Furthermore, we investigate the effectiveness of using auto-annotated data for fine-tuning downstream KIE models. While human-annotated data remains superior, models trained exclusively on DocAnnot's outputs attain respectable performance (e.g., LayoutLMv3 achieving an F1-score of 0.6765 on CORD). These results show that while our framework significantly reduces reliance on manual effort, it does not yet fully eliminate the need for human intervention. However, by automating the process to a point where reviewers can efficiently refine outputs, our system enables near-perfect annotations with much greater efficiency than manual annotation from scratch. This approach offers substantial time and cost savings, making it valuable for resource-constrained settings and rapid model prototyping.

13:00 JSTLLM/生成AIエージェント

VLD-RAG: 長くて視覚的に豊富な複数ページのドキュメントのための、エージェントによる視覚言語検索拡張生成

レポート、スライド、マニュアルなどの視覚的に豊富な文書では、多くの場合、質問に答えるために必要な証拠が複数のページに分散されており、テキストとレイアウトの手がかり、表、チャート、図が混在しています。この研究では、このような視覚的に豊富な長い文書に対する質問応答のためのマルチモーダル検索拡張生成を研究しています。検索では、テキスト信号と視覚信号の両方を含む証拠ページを選択する必要があります。我々は、複数ページの証拠検索と長い文書のクロスページ推論のためのエージェント型マルチモーダル RAG フレームワークである VLD-RAG を紹介します。 VLD-RAG は、解析されたテキスト、ページ レベルのメタデータ、および高密度の視覚表現を保存するページ保存マルチモーダル インデックスを構築し、キーワード ベースのスパース検索と高密度のセマンティック クエリを組み合わせたハイブリッド検索戦略を使用して、候補ソースと証拠ページを特定します。検証者主導のエージェント ワークフローは、取得エージェント、回答エージェント、および検証エージェントを調整して、証拠範囲を拡大し、不足している引用を検出し、必要に応じて取得リクエストを調整します。我々は、トップ 1 およびトップ 5 の証拠ページの精度による検索と、一般化された精度による生成を評価し、VLD-RAG が、LongDocURL や MMLongBench-Doc などの視覚的に豊富な長文書ベンチマークで証拠ページの検索とエンドタスクの質問応答の両方を向上させ、以前のビジョンベースの検索ベースラインを上回るパフォーマンスを示すことを示します。これらの調査結果は、正解がページ全体に散在する証拠に依存する場合、調整されたエージェントの検証とマルチモーダルなハイブリッド検索が信頼性の高い根拠を得るために重要であることを強調しています。

原文 (English)

VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents

Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures. This work studies multimodal retrieval-augmented generation for question answering over such visually-rich long documents, where retrieval must select evidence pages that include both textual and visual signals. We present VLD-RAG, an agentic multimodal RAG framework for multi-page evidence retrieval and cross-page reasoning over long documents. VLD-RAG builds a page-preserving multimodal index that stores parsed text, page-level metadata, and dense visual representations, and uses a hybrid retrieval strategy that combines keyword-based sparse search with dense semantic queries to identify candidate sources and evidence pages. A verifier-guided agent workflow coordinates a Retrieval Agent, Answer Agent, and Validation Agent to broaden evidence coverage, detect missing citations, and refine retrieval requests when needed. We evaluate retrieval with Top-1 and Top-5 evidence-page accuracy and generation with generalized accuracy, and show that VLD-RAG improves both evidence-page retrieval and end-task question answering on visually-rich long-document benchmarks, including LongDocURL and MMLongBench-Doc, outperforming previous vision-based retrieval baselines. These findings highlight that coordinated agent verification and multimodal hybrid retrieval are crucial for reliable grounding when correct answers depend on evidence scattered across pages.

13:00 JST研究/論文

ゲーム AI は面白くない?人間とコンピュータの対戦相手の楽しさの違いに関するスコープレビューとメタ分析

ゲーム キャラクター AI の進歩はプレイヤーの関与を高めることを目的としていますが、対戦相手を人工的なものとして認識すると心理的体験が低下する可能性があることを示す証拠があります。この論文では、人間対コンピュータの対戦相手と対戦する際のプレイヤーの楽しみに焦点を当てた実証研究の範囲のレビューとメタ分析を紹介します。まず、対象となった 20 件の研究の状況をマッピングするためのスコーピング レビューが実施され、研究デザイン、結果の尺度、および研究の焦点が詳細に示されました。次に、9 つの研究からのベースライン比較を総合した 3 レベルのメタ分析により、楽しさの違いが定量的に評価されます。この結果は、統計的に有意な中規模から大規模のプール効果サイズを示しており、コンピュータが敵対する状況における心理的ペナルティを示しています。この論文は、このトピックに関する既存の知識の包括的な概要を提供し、コンピュータの敵対コンテキストのペナルティを完全に理解して解決するためにさらなる研究の必要性を強調します。

原文 (English)

Game AI Not Fun? A Scoping Review and Meta-Analysis on the Differences in Enjoyment between Human and Computer Opponents

Although advancements in game character AI aim to enhance player engagement, evidence suggests that perceiving an opponent as artificial can diminish the psychological experience. This paper presents a scoping review and meta-analysis of empirical studies focusing on player enjoyment when competing against human versus computer opponents. First, the scoping review was conducted to map the landscape of 20 included studies, detailing their study designs, outcome measures, and research foci. Second, a three-level meta-analysis synthesizing baseline comparisons from nine studies quantitatively assesses the differences in enjoyment. The results demonstrate a statistically significant, medium-to-large pooled effect size, indicating a psychological penalty in computer-opponent conditions. This paper provides a comprehensive overview of the extant knowledge on this topic, and underscores the necessity for further research in order to fully understand and resolve the penalty of the computer opponent context.

13:00 JSTLLM/生成AIビジネス/資金調達

CARE-MH: メンタルヘルス LLM の統合、再現可能、比較可能な評価に向けて

大規模言語モデル (LLM) は、メンタルヘルスのサポートを提供するためにますます使用されており、安全性、共感、治療の適切性についての信頼できる評価が必要です。しかし、既存のメンタルヘルスベンチマークは、一貫性のない評価設計と指標の定義のため、再現して比較することが困難です。我々は、メンタルヘルス LLM の比較可能かつ再現可能な評価のための統一フレームワークである CARE-MH を紹介します。 CARE-MH を使用して、最先端のベンチマークを再現および分析すると、再現性はモデルの安定性に大きく依存し、ベンチマーク間の不一致は主にメトリック定義の違いから生じることが明らかになりました。私たちの調査結果は、将来のメンタルヘルス LLM ベンチマークには、標準化された評価構成と共有指標定義の必要性を浮き彫りにしています。

原文 (English)

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic appropriateness. However, existing mental health benchmarks are difficult to reproduce and compare due to inconsistent evaluation designs and metric definitions. We present CARE-MH, a unified framework for comparable and reproducible evaluation of mental health LLMs. Using CARE-MH, we reproduce and analyze state-of-the-art benchmarks, revealing that reproducibility depends strongly on model stability and that cross-benchmark disagreement primarily arises from differences in metric definitions. Our findings highlight the need for standardized evaluation configurations and shared metric definitions for future mental health LLM benchmarks.

13:00 JST研究/論文

オブジェクト指向プログラミングコースにおける学習者と AI の相互作用のパターンと学業成績

この完全な研究論文では、さまざまな形式の学習者と AI の相互作用が、オブジェクト指向プログラミング (OOP) コースの学習成果にどのように関係しているかを調査します。生成人工知能 (GenAI) ツールは、プログラミング教育において学生の間でますます使用されていますが、その教育的影響に関する証拠は依然としてまちまちです。特に、OOP を学習する際に学生が GenAI ツールをどのように統合するか、さまざまな使用パターンが学生の学習経験や成果にどのように関係するかについてはほとんど知られていません。この研究では、学生の自発的な GenAI 使用のパターンと、学業成績、難しさの認識、理解、信頼との関係を調査しています。調査データは、GenAI ツールの使用がコースワークでは許可されているものの、評価では禁止されている 1 年生の OOP コースに登録した 210 人の学部生から収集されました。その結果、学生はコード生成よりも説明の探索やデバッグに GenAI を使用することが大幅に多かったことがわかりました。クラスター分析により、コード生成への依存度が低く、概念的なサポートとデバッグに多用されることを特徴とする「スマートな」高使用パターンを含む、5 つの異なる学習者と AI の相互作用プロファイルが特定されました。使用パターンは、割り当ての難しさの認識、理解の自己評価、AI 生成コードへの信頼、規範に関連する態度の違いと関連していましたが、クラスター間で評価パフォーマンスに大きな違いは見つかりませんでした。これらの調査結果は、自主的な GenAI の使用だけでは測定可能な学習の向上につながらないことを示唆しており、教育学的にガイドされ、プロセスを意識した AI サポートの必要性を強調しています。この研究は、学習者と AI の相互作用パターンに関する経験的証拠に貢献し、プログラミング教育における教育学的に導かれた AI 使用の重要性を強調しています。

原文 (English)

Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course

This full research paper examines how different forms of learner-AI interaction relate to learning outcomes in object-oriented programming (OOP) courses. Generative artificial intelligence (GenAI) tools are increasingly used by students in programming education, yet evidence on their educational impact remains mixed. In particular, little is known about how students integrate GenAI tools when learning OOP, and how different patterns of use relate to students' learning experiences and outcomes. This study investigates patterns of students' self-directed GenAI use and their relationship with academic performance, perceived difficulty, understanding, and trust. Survey data were collected from 210 undergraduate students enrolled in a first-year OOP course in which the use of GenAI tools was permitted for coursework but prohibited in assessments. Results show that students used GenAI significantly more often for explanation seeking and debugging than for code generation. Cluster analysis identified five distinct learner-AI interaction profiles, including a "smart" high-usage pattern characterized by low reliance on code generation and high use for conceptual support and debugging. While usage patterns were associated with differences in perceived assignment difficulty, self-assessed understanding, trust in AI-generated code, and norm-related attitudes, no significant differences in assessment performance were found across clusters. These findings suggest that self-directed GenAI use alone does not lead to measurable learning gains, underscoring the need for pedagogically guided and process-aware AI support. The study contributes empirical evidence on learner-AI interaction patterns and highlights the importance of pedagogically guided AI use in programming education.

13:00 JST研究/論文

記憶がメディアになると何が失われるのか? AI が生成したオーラルヒストリーの視覚化の評価

記憶がメディアになると何が失われるのでしょうか?ディアスポラのオーラルヒストリーインタビューには二重の変容が必要です。一人称の回想から三人称のシーンへ、現在のインタビュールームから過去の時間と場所へ。生成 AI がこの変換を実行する場合、成功に対する合意された基準は存在しません。私たちは、オーラル・ヒストリー理論から成功条件を導出し、3 つの失敗モードに関する 15 の指標を設計し、口頭インタビューから 6 枚の画像シーケンスに及ぶ離散コミュニティからの 82 件のインタビューにわたって、マルチエージェント・シーン分解パイプライン (MAS) と単一要約パイプライン (SSP) を比較しました。ほとんどの場合、シーンの計画と物語の保存は矛盾しており、情報源の証言の物語構造の強さがこの矛盾の主な予測因子です。我々は、障害モードに基づく評価フレームワーク、競合状態の実証的分析、および物語構造の強度に基づくシステム選択のためのルーティングプロトコルを提案します。

原文 (English)

What Gets Lost When Memory Becomes Media? Evaluating AI-Generated Oral History Visualization

What gets lost when memory becomes media? Diaspora oral-history interviews require a double transformation; first-person recollection to third-person scene, present interview room to past time and place. When generative AI performs this transformation, no agreed criteria for success exist. We derive success conditions from oral-history theory, design 15 metrics around three failure modes, and compare a Multi-Agent Scene-decomposition pipeline (MAS) with a Single Summarization Pipeline (SSP) across 82 interviews from diaspora communities, spanning from oral interviews to 6-image sequences. Scene-planning and narrative preservation conflict in the majority of cases, and the narrative-structure strength of the source testimony is the primary predictor of this conflict. We propose a failure-mode-based evaluation framework, an empirical analysis of conflict conditions, and a routing protocol for system selection based on narrative-structure strength.

13:00 JST研究/論文

アイデアから数日で教室へ: 「Vibe コーディング」を使用して IDE アクティビティ ログからプログラミング プロセス ビジュアライザーを作成

この論文では、生徒のプログラミング プロセスを教師が簡単に見えるようにするために、AI 支援の「バイブ コーディング」を使用して構築された Thonny ログ ビジュアライザーの迅速な開発と教室への導入について報告します。私たちは、Thonny (Python 用 IDE) によって生成されたログ ファイルを分析し、学生のプログラミング プロセスの解釈可能なビューを生成する Web アプリケーションを開発しました。教師は、グループまたはコースのログ、ZIP アーカイブ、またはログを含むフォルダーをアップロードできます。システムはすべてのログを解析し、生徒ごとに結果を生成し、ケースをレビューするための生徒ごとのナビゲーションを提供します。各生徒のビューには、インタラクティブなアクティビティのタイムライン、コンパクトなセッションの概要、コードサイズのグラフ、プログラミング プロセスの再生などが含まれます。これらのビューは、学習支援の状況を特定し、学術的完全性を明確にするためにセッションにフラグを立てることを可能にすることで、教師の意思決定をサポートします。このツールは当初、以前のコースのログを使用して評価されました。 2026 年 2 月に 160 名の参加者が参加した入門プログラミング コースのパイロットでは、教師からのフィードバックが提供され、反復的なユーザビリティの改善に関する情報が提供されました。

原文 (English)

From Idea to Classroom in Days: Using "Vibe Coding" to Create a Programming Process Visualizer from IDE Activity Logs

This paper reports on the rapid development and classroom deployment of a Thonny log visualizer built using AI-assisted ``vibe coding'' to make students' programming processes easily visible to teachers. We developed a web application that analyzes log files generated by Thonny (an IDE for Python) and produces interpretable views of students' programming processes. Teachers can upload a log, a ZIP archive, or a folder containing logs for a group or the course; the system parses all logs, generates results per student, and provides student-by-student navigation for reviewing cases. Each student's view includes an interactive activity timeline, a compact session summary, a code-size graph, a programming-process replay, and more. These views support teacher decision-making by enabling the identification of learning-support situations and flagging sessions for academic-integrity clarification. The tool was initially evaluated using logs from previous courses; a February 2026 pilot in an introductory programming course with 160 participants provided teachers' feedback and informed iterative usability improvements.

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文

不信感のない検証: 日常的な人間とチャットボットのインタラクションにおけるユーザー側の監視を日常的な認識論的ガバナンスとして再構成する

人間と AI のインタラクションに関する研究では、システム出力の検証を、より適切に調整された信頼によって削減されるべき信頼に依存する動作として長年枠組み化してきました。私たちは、頻繁にチャットボットを使用する 153 人のユーザーを対象とした混合方法の調査を通じて、人間とチャットボットの日常的な対話におけるこの仮定をテストしました。正規の予測に反して、信頼と検証の間に検出可能な関連性は見出されず、感度分析全体にわたって堅牢な結果が得られました。さらに 3 つのユーザー側の実践 (自動化アクションの前の改良、修正、承認) は広く支持されており、満足度と積極​​的に関連しています。データは、評価的監視(信頼との相関が弱く、満足との結びつきが弱い)と介入主義的監視(信頼との相関が弱く、満足との結びつきが強い)との実質的な違いを明らかにしている。中~大の満足度とコントロールのギャップは、効果的なタスクの結果が主体性のフェルトセンスを生み出していないことを示しています。定性的発見により、手段的メンタルモデル、故障モード特有の疑問、認識論的インフラストラクチャの需要が特定されます。私たちは、ユーザー側の監視を信頼と両立する日常的な認識論的ガバナンスとして再構成し、会話型 AI における足場型監視の 4 つの設計方向性を導き出します。

原文 (English)

Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction

Research on human-AI interaction has long framed verification of system outputs as a trust-contingent behavior that better-calibrated trust should reduce. We test this assumption in everyday human-chatbot interaction through a mixed-methods survey of 153 frequent chatbot users. Contrary to the canonical prediction, we find no detectable association between trust and verification, with the result robust across sensitivity analyses. Three further user-side practices - refinement, correction, and approval before automated actions - are widely endorsed and positively associated with satisfaction. The data reveal a substantive distinction between evaluative oversight (trust-decoupled, weakly tied to satisfaction) and interventionist oversight (weakly trust-correlated, strongly tied to satisfaction). A medium-to-large satisfaction-control gap shows that effective task outcomes do not produce a felt sense of agency. Qualitative findings identify instrumental mental models, failure-mode-specific doubt, and demand for epistemic infrastructure. We reframe user-side oversight as routine epistemic governance compatible with trust, and derive four design directions for scaffolded oversight in conversational AI.

13:00 JSTLLM/生成AI

事実-ヒューリスティック-感情状態の強制による大規模言語モデルの動作の一貫性の測定と改善

大規模言語モデル (LLM) は、実行ごとに同じ決定問題に対して異なる回答を与えることができ、自身の以前の回答がコンテキストとして返されたときに決定を覆すことができます。モデルの重みを変更せずに、この不安定性を測定して部分的に軽減できるかどうかを尋ねます。プロンプトレベルの状態強制レイヤーであるコグニティブ カーネル モデル (CKM) をテストします。決定する前に、モデルは入力を 3 つの認識論的役割、つまり事実 (与えられたまたは検証可能)、ヒューリスティック (推論または仮定)、および感情 (評価または優先信号) に分割する必要があります。 CKM は機能を追加しません。これにより、モデルは行動する前にどのような種類の情報を使用するかを追跡するようになります。形式的には、遷移関数によって更新される構造化状態 S_t = {F_t, H_t, E_t} を維持します。私たちは、4 つのコア実験、4 アーム アブレーション、5 アーム シャム制限アブレーション、および温度プローブを介して、4 つのベンダーの 26 の LLM にわたる韓国語の意思決定シナリオ (曖昧さ、倫理的対立、リソース割り当て、エラー処理) および 37,403 件の観察に基づいて CKM を評価します。調査結果: (1) CKM は反復出力変動を低減します (変量効果ヘッジズの g=1.09、95% CI [0.83, 1.35]、31 モデル ペア)。 (2) 状態の永続性により、新しいモデルでは決定フリップ率が 82% 削減されます (g=1.52)。 (3) 効果は JSON フォーマットだけではありません (値のみの再計算、g=2.24)。 (4) 固定アンカー状態の下での固有のランダム性は無視できます。 (5) サンプリング確率論 (温度 0.7 で g=2.87) の下で利点が増大します。 (6) 偽アブレーションでは、利益の約 45% が構造的な足場に、55% が事実/ヒューリスティック/感情コンテンツに起因しており、CKM は一貫性を高め、反転を減少させる唯一のアームです。 CKM は推論の正しさを改善しません。より狭い結果: 行動の一貫性は測定可能であり、モデル間で異なりますが、決定する前にモデルに事実、仮定、評価信号を分離させることで部分的に改善できます。

原文 (English)

Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement

Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer returns as context. We ask whether this instability can be measured and partially reduced without changing model weights. We test the Cognitive Kernel Model (CKM), a prompt-level state-enforcement layer. Before deciding, the model must separate its input into three epistemic roles: Fact (given or verifiable), Heuristic (inferred or assumed), and Emotion (evaluative or priority signal). CKM adds no capability; it forces the model to track what kind of information it uses before acting. Formally it maintains a structured state S_t = {F_t, H_t, E_t} updated by a transition function. We evaluate CKM on Korean-language decision scenarios (ambiguity, ethical conflict, resource allocation, error handling) across 26 LLMs from four vendors and 37,403 observations, via four core experiments, a 4-arm ablation, a 5-arm sham-restriction ablation, and a temperature probe. Findings: (1) CKM reduces repeated-output variability (random-effects Hedges' g=1.09, 95% CI [0.83, 1.35], 31 model pairs); (2) state persistence cuts the decision-flip rate by 82% in newer models (g=1.52); (3) the effect is not JSON formatting alone (value-only recomputation, g=2.24); (4) intrinsic randomness under fixed anchor states is negligible; (5) the advantage grows under sampling stochasticity (g=2.87 at temperature 0.7); (6) a sham ablation attributes about 45% of the gain to structural scaffolding and 55% to Fact/Heuristic/Emotion content, and CKM is the only arm that both raises consistency and reduces flipping. CKM does not improve reasoning correctness. The narrower result: behavioral consistency is measurable, varies across models, and is partially improvable by forcing models to separate facts, assumptions, and evaluative signals before deciding.

13:00 JSTLLM/生成AI

検索拡張生成のパフォーマンスに対するテキスト チャンク サイズの影響

検索拡張生成 (RAG) システムは、大規模言語モデル (LLM) がテキスト生成時にソース素材として使用する関連情報を取得できるようにする強力なプロセスとして登場しました。これらのシステムの重要でありながら十分に検討されていないコンポーネントは、ソース文書が取得可能なチャンクに分割される粒度です。これらのチャンクのサイズは、生成品質、コンテキストの正確性、検索精度、および計算効率に大きな影響を与える可能性があります。チャンク サイズはその重要性にもかかわらず、生成品質への影響を適切に評価せずに選択されることがよくあります。個々の文などの小さなチャンクでは、各チャンクの焦点を絞ることで正確な検索が可能になる場合があります。ただし、含まれる情報が少ないため、一貫した応答を生成するモデルの能力が制限される可能性があります。章全体などのより大きなチャンクには、正確性を向上させる可能性がある広範な情報が多数含まれていますが、追加のノイズが発生し、計算コストが増加します。大きなチャンクにはより多くの情報が含まれるため、モデルに返されるチャンクの数も考慮する必要があります。この論文では、チャンク サイズと取得されたセグメントの数が生成品質と取得効率にどのように影響するかを評価します。この研究では、これらの構成を比較することで、文書のセグメント化が検索拡張生成システムのパフォーマンスと効率にどのような影響を与えるかをより深く理解することを目指しています。セグメンテーションは、検索拡張生成システムのパフォーマンスと効率に影響を与えます。

原文 (English)

The Effect of Text Chunk Size on Retrieval-Augmented Generation Performance

Retrieval-Augmented Generation (RAG) systems have emerged as a powerful process for allowing large language models (LLMs) to retrieve relevant information to use as source material during text generation. A critical yet under-explored component of these systems is the granularity at which source documents are segmented into retrievable chunks. The size of these chunks has the potential to significantly influence generation quality, contextual correctness, retrieval precision, and computational efficiency. Despite its importance, chunk size is often selected without proper evaluation of its impact on generation quality. Smaller chunks, such as individual sentences, may allow for precise retrieval by narrowing the focus of each chunk. However, they contain less information, which may limit the model's ability to generate coherent responses. Larger chunks, such as entire chapters, contain lots of broad information that may improve correctness, but also introduce additional noise and increase computational cost. Because larger chunks contain more information, the number of chunks returned to the model must also be considered. This paper evaluates how chunk size, along with the number of retrieved segments, influences generation quality and retrieval effectiveness. By comparing these configurations, this study seeks to better understand how document segmentation affects the performance and efficiency of Retrieval-Augmented Generation systems. segmentation affects the performance and efficiency of Retrieval-Augmented Generation systems.

13:00 JST研究/論文

検索の 3 つの側面: RAG におけるドキュメント側、クエリ側、および回答側の相補性の要因証拠

RAG システムはチャンク化に依存しており、これによりドキュメント内の構造情報が破壊されます。既存の見出しベースの検索 (Jeong et al.、2025) では、ドキュメントごとに複数の LLM 呼び出しが必要で、一致したセクション内のサブチャンクを返します。 ToC ガイド付きのページ取得を導入します。これは、LLM 呼び出しを行わずに視覚的な書式設定から見出しを推測し、それらを並列インデックスとして埋め込み、ページ全体のセクションを読み込みます。 8 つの企業文書 (5 ~ 195 ページ) に関する 1,280 の条件にわたって、次のことがわかりました。(1) ToC は回答の品質に対する重要な主効果 (d = +0.41、p = 0.031) であり、完全性 (+0.40) と有用性 (+0.40) が最大の向上を示します。 (2) 回答側の検証と組み合わせると、クエリ側の分解 + 検証よりも優れたパフォーマンスを発揮します (d = +0.32、p = 0.036)。 (3) クエリごとに 2.9 ページしか追加しないにもかかわらず、目次は引用の 20% に貢献しています。 (4) ゲインは長い文書ほど方向的に大きくなります (118 ページで最大 +1.50)。ただし、8 つの文書では傾向は顕著にはなりません。 (5) 480 条件の感度分析では、有意なパラメーター効果は検出されず (すべて p > 0.38)、デフォルトがほぼ最適であることが確認されました。この貢献は、方法論的 (新しいゼロ LLM コスト取得アルゴリズム) と経験的、つまりドキュメント側、クエリ側、および回答側の機能強化が相補的であり、これまで研究されていなかった 3 方向の相互作用であることを示す要因証拠の両方です。

原文 (English)

Three Sides of Retrieval: Factorial Evidence for Document-Side, Query-Side, and Answer-Side Complementarity in RAG

RAG systems rely on chunking, which destroys structural information in documents. Existing heading-based retrieval (Jeong et al., 2025) requires multiple LLM calls per document and returns sub-chunks within matched sections. We introduce ToC-guided page retrieval, which infers headings from visual formatting without LLM calls, embeds them as a parallel index, and loads full page sections. Across 1,280 conditions on 8 enterprise documents (5 to 195 pages), we find: (1) ToC is a significant main effect on answer quality (d = +0.41, p = 0.031), with the largest gains in completeness (+0.40) and usefulness (+0.40); (2) combined with answer-side verification, it outperforms query-side decomposition + verification (d = +0.32, p = 0.036); (3) ToC contributes 20% of citations despite adding only 2.9 pages per query; (4) gains are directionally larger on longer documents (up to +1.50 on 118 pages), though the trend does not reach significance with 8 documents; and (5) a 480-condition sensitivity analysis finds no significant parameter effects (all p > 0.38), confirming defaults are near-optimal. The contribution is both methodological (a new zero-LLM-cost retrieval algorithm) and empirical: factorial evidence that document-side, query-side, and answer-side enhancements are complementary, a three-way interaction not previously studied.

13:00 JST研究/論文

DDSNet: PCSEL 特性予測のためのデュアルドメイン対称性認識ネットワーク

フォトニック結晶 (PhC) 格子設計空間の効率的な探索は、フォトニック結晶面発光レーザーの開発に不可欠です。結合波理論 (CWT) は効果的な物理フレームワークを提供しますが、その計算コストは​​依然として大規模な探査には法外であり、神経代替の需要を高めています。しかし、既存の AI モデルは、CWT によって示される PhC 単位セルの誘電パターンの 2 つの重要な要素、つまりデバイスの物理的特性を大きく左右するスペクトル成分と非対称構造を十分に活用していません。この不一致により、特に構造に敏感な領域において、代理精度とスクリーニングの信頼性が低下します。これに対処するために、Dual-Domain Symmetry-Aware Network (DDSNet) を提案します。これは、並進等価スペクトル フィルタリングと対称性誘起構造事前処理を統合します。スペクトル フィルタリングは、格子上の平行移動等分散を維持しながら、スペクトル誘導バイアスを視覚モデルに注入します。一方、構造的事前分布は、格子特徴を還元不可能な表現関連の対称性が解決されたコンポーネントに分解し、それらを別々のブランチで処理します。実験では、DDSNet が特性予測とハイスループット スクリーニングにおいて既存の AI ベースラインを大幅に上回り、構造に敏感な領域で優れた信頼性を示すことが実証されました。重要なことは、コンポーネント マスキング分析により、ネットワークが物理的な事前分布に合わせたプロパティ固有の依存関係を首尾よく学習していることが明らかになることです。これらの結果は、DDSNet が物理的に意味のある構造と特性の関係を効果的に捕捉し、PhC 設計空間探索のための信頼性の高い神経代替手段を確立していることを示しています。

原文 (English)

DDSNet: Dual-domain Symmetry-aware Network for PCSEL Property Prediction

Efficient exploration of the photonic crystal (PhC) lattice design space is essential for developing photonic crystal surface-emitting lasers. While coupled-wave theory (CWT) provides an effective physical framework, its computational cost remains prohibitive for large-scale exploration, driving the demand for neural surrogates. However, existing AI models underexploit two key factors of PhC unit-cell dielectric patterns indicated by CWT: spectral components and asymmetric structures, which largely govern devices' physical properties. This mismatch weakens surrogate accuracy and screening reliability, especially in structure-sensitive regions. To address this, we propose the Dual-Domain Symmetry-Aware Network (DDSNet). It integrates translation-equivariant spectral filtering with a symmetry-induced structural prior. The spectral filtering injects a spectral inductive bias into vision model while preserving translation equivariance on lattices. Meanwhile, the structural prior decomposes lattice features into irreducible representation-associated, symmetry-resolved components and processes them in separate branches. Experiments demonstrate that DDSNet significantly outperforms existing AI baselines in property prediction and high-throughput screening, exhibiting superior reliability in structure-sensitive regions. Crucially, component masking analyses reveal that the network successfully learns property-specific dependencies aligned with physical priors. These results indicate that DDSNet effectively captures physically meaningful structure-property relationships, establishing a highly reliable neural surrogate for PhC design space exploration.

13:00 JST研究/論文

大規模なオーディオビジュアル検索モデルにおける空間グラウンディングのロックを解除する

高密度の空間アノテーションを大規模に取得するにはコストがかかるため、監視が弱いと視聴覚音源定位の実際的な体制が確立されます。ただし、モデルはピクセルレベルの監視なしで、時間的に位置合わせされたオーディオビジュアルデータから音源を特定する必要があるため、この作業は依然として困難です。最近の大規模なオーディオビジュアル検索モデルは、前例のない規模でトレーニングされ、豊富なマルチモーダル構造をエンコードしています。私たちは、それらの潜在的な表現が、グローバルな調整のために最適化されているにもかかわらず、きめの細かい空間グラウンディングを可能にすることを示します。空間的な詳細はグローバル プーリングにより検索バックボーンの上位層で徐々に失われますが、中間の視覚トークンは高度に構造化された空間情報を保持します。これを利用するために、標準のグローバル集約モジュールを置き換える軽量の \emph{Audio-informed Spatial Pooling} (AiSP) を採用するフレームワークである LAIP (\emph{Audio-informed Pooling によるローカリゼーション}) を導入します。 LAIP は、フレームに位置合わせされたオーディオを使用して中間ビジュアル トークンをクエリすることにより、凍結された取得パイプラインによって破棄されるローカライズされた空間情報を回復します。私たちのアプローチは、AVSBench と AVATAR で最先端のパフォーマンスを達成し、AVATAR では以前の結果のほぼ 2 倍になります。これらの発見は、正確な位置特定を最初から学習する必要がないことを証明しています。代わりに、既存の検索表現からロックを解除して、検索タスクとローカリゼーション タスクの両方に統一されたパスを提供できます。

原文 (English)

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual data without pixel-level supervision. Recent large-scale audio-visual retrieval models, trained at unprecedented scale, encode rich multimodal structure. We show their latent representations, though optimized for global alignment, can nonetheless enable fine-grained spatial grounding. While spatial detail is progressively lost in the upper layers of retrieval backbones due to global pooling, intermediate visual tokens retain highly structured spatial information. To exploit this, we introduce LAIP (\emph{Localization via Audio-Informed Pooling}), a framework that employs a lightweight \emph{Audio-informed Spatial Pooling} (AiSP) to replace the standard global aggregation module. By using frame-aligned audio to query intermediate visual tokens, LAIP recovers localized spatial information that is otherwise discarded by the frozen retrieval pipeline. Our approach achieves state-of-the-art performance on AVSBench and AVATAR, nearly doubling previous results on the latter. These findings prove that accurate localization does not need to be learned from scratch; instead, it can be unlocked from existing retrieval representations, providing a unified path for both retrieval and localization tasks.

13:00 JSTLLM/生成AIエージェント

Naive RAG から Deep Agentic Retrieval まで: 規制遵守のための進化するコンテキスト エンジニアリング パイプライン

検索拡張生成 (RAG) は、大規模言語モデル (LLM) を企業文書コーパスに適用するための主要なパラダイムですが、コーパスの規模とクエリの複雑さが増大するにつれて、単純な実装では厳しい制限に直面します。このペーパーでは、オンタリオ州エネルギー委員会 (OEB) の報告要件に基づく規制遵守とレートケース分析を目的としたオンタリオ発電所 (OPG) の生産回収パイプラインの進化を追跡します。私たちは、ナイーブ RAG、再ランキングを伴うハイブリッド取得、エージェント関数呼び出しの取得、コードベースのツール合成と明示的な計画を備えた深いマルチエージェント アーキテクチャといった連続する段階を調査し、各移行の動機となった障害モードとトレードオフを特定します。私たちは、成熟したアーキテクチャをコストを意識した段階的証拠収集 (PEA-CAE) として形式化します。低コストで高精度の取得から開始し、期待される証拠の獲得が待ち時間とコストを正当化する場合にのみ文書全体の読み取りにエスカレートします。私たちの調査結果は、コンテキスト エンジニアリングが、進化する大規模な規制コーパスに対するドメイン固有の微調整よりも扱いやすく、経済的に実行可能な方法であることを示しています。より広範には、深いエージェント検索への進歩は、古典的な情報検索のアイデアを反映しており、実際的なシステムのプリミティブとして、適応型クエリの再定式化、プログレッシブなドキュメント発見、および階層型サブエージェント要約を導入しています。運用トレースは、最新の検索システムの検索ベースの性質をさらにサポートしており、エンタープライズ規模の質問応答の基盤として、反復的な証拠の取得と適応型計画がシングルパス検索に取って代わりつつあります。

原文 (English)

From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance

Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive implementations encounter hard limits as corpus scale and query complexity grow. This paper traces the evolution of a production retrieval pipeline at Ontario Power Generation (OPG) for regulatory compliance and rate case analysis under Ontario Energy Board (OEB) reporting requirements. We examine successive stages: naive RAG, hybrid retrieval with re-ranking, agentic function-calling retrieval, and a deep multi-agent architecture with code-based tool synthesis and explicit planning, and identify the failure modes and tradeoffs that motivated each transition. We formalize the mature architecture as Progressive Evidence Acquisition with Cost-Aware Escalation (PEA-CAE): begin with low-cost, high-precision retrieval and escalate to full-document reads only when the expected evidence gain justifies latency and cost. Our findings show that context engineering is a more tractable and economically viable path than domain-specific fine-tuning for large, evolving regulatory corpora. More broadly, the progression toward deep agentic retrieval mirrors classical information retrieval ideas, introducing adaptive query reformulation, progressive document discovery, and hierarchical subagent summarization as practical system primitives. Operational traces further support the search-based nature of modern retrieval systems, where iterative evidence acquisition and adaptive planning increasingly replace single-pass retrieval as the foundation for enterprise-scale question answering.

13:00 JST研究/論文

エネルギー事業におけるレガシー企業資産管理のための AI 支援ナレッジ アクセス: 実用的な検索システム

エネルギー事業会社は、依然として、長寿命のエンタープライズ資産管理プラットフォーム上でエンジニアリング作業管理、エンジニアリング調達、在庫プロセスを実行しています。これらのプラットフォームの置き換えには法外な費用がかかり、運用上の混乱が生じることが多いため、実際的な改善層が必要です。このペーパーでは、ベンダー ドキュメントの質問応答、オペレーショナル データ ストア (ODS) スキーマの質問応答、ユーザー インターフェイスの使用法とハウツーの質問応答という 3 つの動作モードにわたって、日常の知識へのアクセスを向上させる検索アシスタントについて説明します。ランタイム手法は、意図の理解、クエリの書き換え、ハイブリッド セマンティックおよびベクトルの取得、トークン制限の下でのコンテキスト エンジニアリング、根拠のある回答の生成、およびパネル識別子と引用ドキュメントの決定論的なハイパーリンク変換を組み合わせます。データ準備パイプラインは、テーブルとフィールドの説明を追加し、ソース全体で頭字語を正規化し、必要に応じて代表的な行レベルのコンテキストにインデックスを付けることにより、主要な品質レバーとしてセマンティック強化を強調します。測定されたパイロットでは、検索品質とユーザーの成果が一貫して向上していることがわかります。 5 での精度は 0.56 から 0.72 に、平均逆数順位は 0.43 から 0.58 に、5 での正規化割引累積ゲイン (nDCG) は 0.51 から 0.66 に改善されました。タスク完了時間の中央値は 14.2 分から 8.3 分に短縮され、有用性と信頼性は両方とも 5 段階評価で 4.0 に増加しました。結果は少数のサンプルに基づいており、パイロット調査結果として報告されていますが、意図の理解とセマンティック強化により、レガシー環境で有意義な運用価値を提供できると同時に、将来の分析ツールや自動化ツールのための再利用可能な基盤を確立できることが示されています。

原文 (English)

AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System

Energy utilities still run engineering work management, engineering procurement, and inventory processes on long-lived enterprise asset management platforms. Replacing these platforms is often cost prohibitive and operationally disruptive, so practical improvement layers are required. This paper presents a retrieval assistant that improves day-to-day knowledge access across three operational modes: vendor documentation question answering, operational data store (ODS) schema question answering, and user interface usage and how-to question answering. The runtime method combines intent understanding, query rewriting, hybrid semantic and vector retrieval, context engineering under token limits, grounded answer generation, and deterministic hyperlink conversion for panel identifiers and cited documentation. The data preparation pipeline emphasizes semantic enrichment as the primary quality lever by adding table and field descriptions, normalizing acronyms across sources, and indexing representative row-level context when useful. A measured pilot shows consistent gains in retrieval quality and user outcomes. Precision at five improved from 0.56 to 0.72, mean reciprocal rank from 0.43 to 0.58, and normalized discounted cumulative gain (nDCG) at five from 0.51 to 0.66. Median task completion time dropped from 14.2 to 8.3 minutes, while usefulness and confidence both increased to 4.0 on a five-point scale. Results are based on a small sample and are reported as pilot findings, but they indicate that intent understanding and semantic enrichment can deliver meaningful operational value in legacy environments while also establishing reusable foundations for future analytics and automation tools.

13:00 JST研究/論文

パーキンソン病におけるタイピングミスによる運動回復の選択的障害:生存分析

パーキンソン病 (PD) は、運動制御と認知制御の複数の解離可能な段階に影響を与えます。受動的に収集されたキーストロークのダイナミクスが、これらの段階のうち 2 つ、つまり自己生成エラーの認識 (エラー監視) とその後の正常なモーター リズムの回復 (モーターの再起動) を区別できるかどうかを尋ねます。公開されている NeuroQWERTY MIT-CSXPD データセット (十分なバックスペース データを持つ 57 人の被験者、PD および UPDRS-III スコアを持つ 27 人) でバックスペース イベントを自然に発生するエラー修正エピソードとして使用すると、エラー前のキーストロークの不安定性が疾患の重症度によって異なるという証拠は見つかりませんでしたが ($r=0.164$、$p=0.413$)、連続的なものとしてモデル化されたエラー後の回復時間が異なるという強力な証拠は見つかりました。加速故障時間 (AFT) 生存結果 (2 つの独立したサブコホートのそれぞれにおける $p<10^{-9}$; 順列 $p<0.0033$; ゼロを除くブートストラップ 95% CI)。 2 つの測定値は相互に相関がなく ($r=-0.065$)、結合モデルではエラー後の回復のみが有意に残り、2 つの冗長信号ではなく真の解離が確認されます。この効果は、生のタイピング速度を制御しても存続し、独立した臨床運動テスト (交互の指タッピング、$p=0.011$) に対しても再現されます。離散化されたイベント発生までの時間の結果(カプラン・マイヤー生存曲線に類似)として回復をモデル化する最初の試みでは、閾値の選択に関係なく信号が破壊されました。連続 AFT 定式化に切り替えると、データの右に偏った分布にも通常の回帰よりもはるかに良く適合し、回復しました。我々は、この行動解離を、PDにおけるエラー検出とエラー後の運動調整が別個の神経回路によって媒介され、前者はほとんど保護され、後者は視床下核の活動に関連しているという既存の電気生理学的証拠と関連付け、この解離は日常のタイピングだけで検出可能であると主張する。

原文 (English)

Selective Impairment of Motor Recovery from Typing Errors in Parkinson's Disease: A Survival Analysis

Parkinson's disease (PD) affects multiple, dissociable stages of motor and cognitive control. We ask whether passively-collected keystroke dynamics can distinguish two of these stages: noticing a self-generated error (error monitoring) versus recovering normal motor rhythm afterward (motor restart). Using backspace events as naturally-occurring error-correction episodes in the public neuroQWERTY MIT-CSXPD dataset (57 subjects with sufficient backspace data, 27 with PD and UPDRS-III scores), we find no evidence that pre-error keystroke instability differs by disease severity ($r=0.164$, $p=0.413$), but strong evidence that post-error recovery time does, modeled as a continuous accelerated failure time (AFT) survival outcome ($p<10^{-9}$ in each of two independent sub-cohorts; permutation $p<0.0033$; bootstrap 95% CI excluding zero). The two measures are uncorrelated with each other ($r=-0.065$), and in a joint model only post-error recovery remains significant, confirming a genuine dissociation rather than two redundant signals. The effect survives controlling for raw typing speed and replicates against an independent clinical motor test (alternating finger-tapping, $p=0.011$). An initial attempt to model recovery as a discretized time-to-event outcome (analogous to Kaplan-Meier survival curves) destroyed the signal regardless of threshold choice; switching to a continuous AFT formulation, which also fits the data's right-skewed distribution far better than ordinary regression, recovered it. We relate this behavioral dissociation to existing electrophysiological evidence that error detection and post-error motor adjustment in PD are mediated by distinct neural circuits, the former largely spared, the latter linked to subthalamic nucleus activity, and argue that this dissociation is detectable through everyday typing alone.

13:00 JSTLLM/生成AIGPT / ChatGPT

リーダーなしで読む: 大規模な言語モデルにより、読み取りと書き込みが単一の絡み合ったコードに崩壊します

読み書きができる人間の脳では、読み書きは 2 つの二重解離可能なシステムです。腹側解読ルート (純粋な失読症で障害) と前頭頭頂側符号化ルート (純粋な失書症で障害) であり、部分的な正書法コアを共有しています。代わりに、デコーダー専用の大規模言語モデル (LLM) が、テキストに最適化された単一の自己回帰パスから両方を駆動します。これは、進化した本能ではなく、最近の文化的な発明です。独立初期化フロアとタイシーリングに対して校正されたエンタングルメントインデックス $E \in [0,1]$ (CKA、プロクラステス残差、相互 $k$-NN) を介して、入力側の「読み取りコード」$W_E$ と出力側の「書き込みコード」$W_U$ を比較し、その 1 つのメカニズムがどの程度もつれているかを尋ねます。 GPT-2、OPT、Pythia (14M--1.4B)、T5、および BERT/RoBERTa の 9 つのプローブ (6 つは確立された結果を統合し、3 つは読み取り/書き込み分析を導入) にわたって、2 つの相補的なレベルの方向性が一致しています。重みでは、アンタイド モデルは、非単調な結合-微分軌道上に 1 つの結合されているが天井未満のコード ($E=0.23$--$0.35$、床よりもはるかに高い) を保持しており、$W_U$ はすべての周波数十分位で $W_E$ よりも $\sim 3.2\times$ 遠くにドリフトします。行動においては、理解と生産は、12 の非変性モデルすべてで正に結合しており (サインテスト $p<0.001$)、これは脳の二重解離とは逆です。この結合は一般的なものであり、デコーダのみではありません。エンコーダとデコーダは 2 つの経路を表現的に (最大 0.96) 分離しますが、動作的に結合したままになります。 null を明確に報告します (ジオメトリ $\rightarrow$ 動作ブリッジは null、$\rho=0.00$)。単一の順方向パスによって何らかの結合がアプリオリに期待されるため、私たちの貢献はその定量化とクロスレベルの一致です。相同性ではなく類推によって、これは LLM を可能な心の空間における明確な点として位置づけます。

原文 (English)

Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code

In the literate human brain, reading and writing are two doubly-dissociable systems: a ventral decoding route (impaired in pure alexia) and a fronto-parietal encoding route (impaired in pure agraphia), sharing a partial orthographic core. A decoder-only large language model (LLM) instead drives both from a single autoregressive path optimized on text, a recent cultural invention rather than an evolved instinct. We ask how entangled that one mechanism is, comparing an input-side "reading code" $W_E$ with an output-side "writing code" $W_U$ via an entanglement index $E \in [0,1]$ (CKA, Procrustes residual, mutual $k$-NN) calibrated against an independent-init floor and a tied ceiling. Across nine probes on GPT-2, OPT, Pythia (14M--1.4B), T5, and BERT/RoBERTa (six consolidating established results, three introducing the read/write analysis), two complementary levels agree in direction. In the weights, untied models hold one coupled but sub-ceiling code ($E=0.23$--$0.35$, far above floor) on a non-monotonic couple-then-differentiate trajectory, with $W_U$ drifting $\sim 3.2\times$ farther than $W_E$ in every frequency decile. In behaviour, comprehension and production are positively coupled in all 12 non-degenerate models (sign test $p<0.001$), the opposite of the brain's double dissociation. This coupling is general, not decoder-only: encoder--decoders separate the two pathways representationally (up to 0.96) yet stay behaviourally coupled. We report our nulls plainly (the geometry $\rightarrow$ behaviour bridge is null, $\rho=0.00$). Because a single forward path makes some coupling expected a priori, our contribution is its quantification and cross-level concordance; by analogy, not homology, this situates LLMs as a distinct point in the space of possible minds.

13:00 JST研究/論文

オープンソース SLM を使用した科学文書理解のためのマルチモーダル ハイブリッド検索拡張生成

大規模な言語モデルは、事前に微調整を行わずに科学文書からのドメイン固有の質問に答えるときに幻覚を起こす傾向があります。現在、検索拡張生成などの方法はこの問題を部分的に解決していますが、限られたコンテキスト知識、疎検索と密検索の違い、検索ノイズなどの別の課題に直面しています。本稿では、これらの課題を解決し、情報抽出の精度を向上させることを目的とした、高度なマルチモーダル検索拡張生成システムを紹介します。提案されたアーキテクチャでは、オープンソースのビジョン言語モデル (Qwen2-VL-2B-Instruct) を利用して表と図のテキスト要約を生成するマルチモーダル取り込みパイプラインが導入されています。検索フェーズでは、HNSW ベースのセマンティック検索と GIN ベースの語彙検索が統合され、相互ランク融合を通じて統合され、検索ノイズを最小限に抑えるためにクロス エンコーダーの再ランキングを使用して洗練されます。マルチターンの対話全体で会話の一貫性を確保するために、Query Condenser モジュールが採用されています。評価は、MMLongBench ベンチマーク、BeIR 形式の合成データセット、DeepEval フレームワークを使用して、取り込み、取得、生成の各段階を個別に評価することによって行われます。さらに、結果は、Naive-RAG ベースラインと比較して取得品質が 157% 向上し、待ち時間がわずか 50 ミリ秒増加しただけであり、Qwen2-VL-2B-Instruct は BERTScore のクラウドベースのモデルと同等の結果を達成したことを示しています。これらの調査結果は、オープンソースに最適化された SLM を高度な検索戦略と組み合わせることで、クラウドベースのモデルに依存せずに文書理解において競争力のあるパフォーマンスを提供できることを検証しています。

原文 (English)

Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs

Large Language Models tend to hallucinate when answering domain-specific ques tions from scientific documents without prior fine-tuning. Currently, methods such as Retrieval-Augmented Generation partially solve this problem but face different challenges: limited context knowledge, difference between sparse and dense retrieval, and retrieval noise. This paper presents an Advanced Multimodal Retrieval-Augmented Generation system that aims to solve those challenges and im prove the accuracy of information extraction. The proposed architecture introduces a multimodal ingestion pipeline that leverages an open-source Vision-Language Model (Qwen2-VL-2B-Instruct) to generate textual summaries of tables and fig ures. The retrieval phase integrates HNSW-based semantic search with GIN-based lexical search, unified through Reciprocal Rank Fusion and refined using Cross Encoder reranking to minimize retrieval noise. To ensure conversational coherence across multi-turn interactions, a Query Condenser module is employed. Evaluation is conducted by independently assessing the ingestion, retrieval and generation stages using the MMLongBench benchmark, a BeIR-format synthetic dataset and the DeepEval framework. Moreover, results demonstrate a 157% improvement in retrieval quality over a Naive-RAG baseline, with only 50 ms additional la tency, while Qwen2-VL-2B-Instruct achieved results comparable to cloud-based models in BERTScore. These findings validate that open-source optimized SLMs, paired with advanced retrieval strategies, can provide competitive performance for document understanding without relying on cloud-based models.

13:00 JSTLLM/生成AI

検索前に考えるのが面倒な場合: 適応型ナレッジグラフ検索のための TraceBound 診断

適応型検索は、コントローラーに検索、近隣の調査、アクションの修正、証拠が十分な場合の停止を許可することで、ナレッジ グラフの質問応答をより堅牢にすることを約束します。私たちは、テキストが豊富なナレッジ グラフ上の ARK スタイルの取得者向けの軽量のプロファイルおよびトレース条件付き診断プロトコルである TraceBound を導入することで、この前提を研究します。 TraceBound は、グラフ データ、ツール、ゴールド ラベル、ランキング メトリックを固定したまま、取得前にコンパクトなクエリ プロファイルを公開し、観察可能な障害症状の後に短いトレース ヒントを発行し、軌跡カウンターをログに記録します。 STaRK 検証とホールドアウト サブセット全体で、追加の条件付けにより検査可能性は向上しますが、オープンウェイト コントローラーの下では一貫して取得品質が低下します。ペア トラジェクトリ分析では、繰り返しの呼び出し、結果がゼロのコール、および探索バジェットの誤った割り当てに劣化の原因が特定されますが、インタラクション バジェットが厳格化されると、ポリシーが修復されずにトラジェクトリが短縮されます。この結果は、「検索前の思考」がプロンプト形式の変更としてではなく、アクション選択に対する制御の問題として評価される必要があるという点で、一般的な失敗モードを診断します。

原文 (English)

When Thinking Before Retrieval Hurts: TraceBound Diagnostics for Adaptive Knowledge-Graph Retrieval

Adaptive retrieval promises to make knowledge-graph question answering more robust by letting a controller search, inspect neighborhoods, revise actions, and stop when evidence is sufficient. We study this premise by introducing TraceBound, a lightweight profile- and trace-conditioned diagnostic protocol for an ARK-style retriever on text-rich knowledge graphs. TraceBound exposes a compact query profile before retrieval, issues short trace hints after observable failure symptoms, and logs trajectory counters, while keeping graph data, tools, gold labels, and ranking metrics fixed. Across STaRK validation and held-out subsets, the added conditioning improves inspectability but consistently reduces retrieval quality under open-weight controllers. Paired trajectory analysis localizes the degradation to repeated calls, zero-result calls, and misallocated exploration budget, while stricter interaction budgets shorten trajectories without repairing the policy. The result diagnoses the common failure mode in that "thinking before retrieval'' must be evaluated as a control problem over action selection, not as a prompt-format change.

13:00 JST研究/論文

適合性が異なる多感覚フィードバック下でのエラー関連の可能性の解読

エラー関連電位 (ErrP) は、人間とマシンの相互作用におけるエラー処理に関連する神経シグネチャとして広く研究されています。現実的な設定では、エラー知覚は不均一な多感覚フィードバックの下で発生することが多く、感覚モダリティとフィードバックの一致性によって引き起こされる変動性が、信頼性の高い ErrP デコードに課題をもたらします。特に、不一致なフィードバックは、デコードの難易度の増加と分類パフォーマンスの低下に関連しています。この課題に対処するために、制御された感覚一致性を備えたマルチモーダルな視覚、聴覚、および触覚フィードバックの下で堅牢な ErrP デコードのための学習戦略を調査します。補助監視を備えたマルチブランチ EEGNet ベースのアーキテクチャを採用し、明示的なモダリティ固有の仮定に依存することなく、異種条件全体での堅牢性を向上させます。実験は、単峰性、二峰性、および三峰性のフィードバック構成を備えた迷路観察タスクを使用して実施されました。提案されたアプローチは、被験者全体にわたって、不均一な感覚条件全体で一貫した分類パフォーマンスを達成し、特にマルチモーダルフィードバックの下で、ベースライン EEGNet モデルと比較して精度の向上を示しました。これらの結果は、適切なアーキテクチャ設計とトレーニング戦略により、不均一な多感覚条件下での ErrP デコードの安定性を向上できることを示唆しています。

原文 (English)

Decoding Error-Related Potentials under Multisensory Feedback with Varying Congruency

Error-related potentials (ErrPs) are widely studied neural signatures associated with error processing in human-machine interaction. In realistic settings, error perception often occurs under heterogeneous multisensory feedback, where variability induced by sensory modality and feedback congruency poses challenges for reliable ErrP decoding. In particular, incongruent feedback is associated with increased decoding difficulty and reduced classification performance. To address this challenge, we investigate learning strategies for robust ErrP decoding under multimodal visual, auditory, and tactile feedback with controlled sensory congruency. We adopt a multi-branch EEGNet-based architecture with auxiliary supervision to improve robustness across heterogeneous conditions, without relying on explicit modality-specific assumptions. Experiments were conducted using a maze-observation task with unimodal, bimodal, and trimodal feedback configurations. Across subjects, the proposed approach achieved consistent classification performance across heterogeneous sensory conditions and showed improved accuracy compared to baseline EEGNet models, particularly under multimodal feedback. These results suggest that appropriate architectural design and training strategies can improve the stability of ErrP decoding under heterogeneous multisensory conditions.

13:00 JST研究/論文

認知の経路統合モデル

私たちは、意識の経路積分モデルの基礎となる認知コスト最適化の数学的および物理的定式化を開発します。目標指向の認知プロセスは、ターゲット概念と一致する構成に報酬を与えるプロジェクター ハミルトニアンの下で虚数時間進化 (ITE) としてモデル化されます。 3 つの結果を確立します。第 1 に、この ITE は二重括弧の流れと一致するため、固有の最小値が解となるヒルベルト - シュミット コストのリーマン勾配の流れです。第二に、ウィック回転は、この非ユニタリ降下を同じヒルベルト空間上の等価ユニタリ進化として再表現します。これにより、オラクルと初期状態の拡散射影器が位置エネルギーと運動エネルギーの役割を果たす正確な離散経路積分表現が可能になります。第三に、認知システムと神経環境プローブ間のユニタリ相互作用の強さによって、無意識から意識の処理までの連続体を特定し、マルコフの弱結合限界における Asano \textit{et al.} の Gorini--Kossakowski--Sudarshan--Lindblad (GKSL) デコヒーレンス モデルと、強結合における最適化された状態の射影的で報告可能な固定を回復します。限界、洞察に似た「なるほど」エンドポイント。どちらの領域も同じ ITE とパス積分構造を共有しており、測定と相互作用の強度のみが異なります。したがって、ウィックの回転は再記述の手法であり、物理的なレジームの変更ではありません。

原文 (English)

A Path Integral Model of Cognition

We develop the mathematical and physical formulation of cognitive cost optimization that underlies the path-integral model of consciousness. The goal-directed cognitive process is modeled as imaginary-time evolution (ITE) under a projector Hamiltonian that rewards configurations consistent with a target concept. We establish three results. First, this ITE coincides with a double-bracket flow and is therefore the Riemannian gradient flow of a Hilbert--Schmidt cost whose unique minimum is the solution. Second, a Wick rotation re-expresses this non-unitary descent as an equivalent unitary evolution on the same Hilbert space, which admits an exact discrete path-integral representation in which the oracle and the initial-state diffusion projector play the roles of potential and kinetic energy. Third, we identify the continuum from unconscious to conscious processing with the strength of the unitary interaction between the cognitive system and a neural-environment probe, recovering the Gorini--Kossakowski--Sudarshan--Lindblad (GKSL) decoherence model of Asano \textit{et al.} in the Markovian weak-coupling limit, and the projective, reportable fixation of an optimized state in the strong-coupling limit, an insight-like ``Aha'' endpoint. Both regimes share the same ITE and path-integral structure, and only the measurement-interaction strength varies. The Wick rotation is therefore a technique of re-description, not a physical regime change.

13:00 JST研究/論文

AI が生成したバイオデジタル アーキテクチャ画像からの EEG 感情認識

AI が生成した画像からの脳波 (EEG) データを使用して、バイオデジタル アーキテクチャに対する感情的反応を調べました。 336 人の参加者が参加した事前実験では、最初の 600 枚のプールから、畏怖、嫌悪感、内容に分類される強い感情反応を引き起こす 60 枚の画像が特定されました。これらの画像は、既存のデータセットの分析に基づいてチャネルの選択とサンプル サイズの推定を行い、52 人のボランティアの EEG 記録に使用されました。ガンマ バンドとデルタ バンドが最も高い分類精度をもたらし、ガンマ バンドは畏怖の感情について 77.07 パーセント +/- 13.8 パーセントの精度を達成しました。緑や不均一な粒度などの重要な要素はポジティブな感情に関連付けられますが、湿気はネガティブな反応を引き起こします。これらの結果は、美的魅力と受容性を高めるために、バイオデジタル建築に自然要素とさまざまなテクスチャを組み込むことの重要性を強調しています。この研究は、EEG が建築の好みを客観的に評価する能力を実証し、建築家が魅力的で持続可能な環境を設計するための貴重な洞察を提供します。

原文 (English)

EEG Emotion Recognition From AI-Generated Biodigital Architecture Images

Emotional responses to biodigital architecture were examined using electroencephalographic (EEG) data from AI-generated images. A pre-experiment involving 336 participants identified 60 images, selected from an initial pool of 600, that elicited strong emotional responses categorized as awe, disgust, or content. These images were used for EEG recordings of 52 volunteers, with channel selection and sample size estimation based on the analysis of an existing dataset. Gamma and delta bands yielded the highest classification accuracy, with the gamma band achieving an accuracy of 77.07 percent +/- 13.8 percent for the awe emotion. Key factors such as greenery and non-uniform granularity were linked to positive emotions, while dampness triggered negative reactions. These results emphasize the significance of incorporating natural elements and varied textures in biodigital architecture to enhance aesthetic appeal and acceptance. The study demonstrates EEG's capability to objectively assess architectural preferences, providing valuable insights for architects to design engaging and sustainable environments.

13:00 JSTLLM/生成AI

メンタルヘルスのための LLM における検索拡張生成: 階層化された安全アーキテクチャ内での検索の増分寄与の定量化

デジタル メンタルヘルス介入 (DMHI) はスケーラブルなサポートを提供しますが、不安定な状況においてユーザーの意図を正確に検出することは困難な場合があります。純粋なパラメトリック大規模言語モデル (LLM) には、特定の安全性が重要なアーキテクチャが含まれていないため、重要な合図を見逃したり、幻覚を起こして信頼性を損なう可能性があります。取得されたコンテキストで LLM を補足する取得拡張生成 (RAG) は、不安定な状況での意図の検出を強化する可能性があります。市販の DMHI は通常、ルールベースのフィルター、シンボリックエスカレーションプロトコル、ニューラル分類などの複数の独立した安全レイヤーを組み合わせています。ただし、単一層の増分寄与は定量化されていないままです。このペーパーでは、RAG 有効モードと RAG 無効モードの制御された比較を通じて、Wysa と呼ばれる DMHI 内の 6 つの LLM モデルを評価します。匿名化された実際のユーザーと合成のユーザーとチャットボットのやりとりには、資格のある臨床チームによって複数の意図カテゴリー (自傷行為、虐待、パニックなど) に対して注釈が付けられました。この研究では、分類精度、再現率、適合率、およびグラウンド トゥルース ラベルに対する F1 スコアを計算し、統計的有意性について差異をテストしました。パフォーマンスは、リスク カテゴリおよびモデル間の一致ごとにも検査されました。 RAG は誤報の増加を引き起こしましたが、そのトレードオフは感度を優先する安全性重視の設計原則と一致しており、フラグが立てられたケースは直接対処されるのではなく追加のレビューに送られます。全体として、これらの発見は、LLM 駆動の DMHI の精度、一貫性、安全性を向上させるための有望なアプローチとして RAG を裏付けています。キーワード: デジタルメンタルヘルス介入、大規模言語モデル、検索拡張生成、精度、再現率、精度

原文 (English)

Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture

Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging. Pure parametric Large Language models (LLMs) do not contain specific safety critical architecture, and can miss critical cues, or hallucinate, undermining reliability. Retrieval Augmented Generation (RAG), which supplements an LLM with retrieved context, could enhance intent detection during volatile situations. Commercially available DMHIs typically combine multiple independent safety layers like rule-based filters, symbolic escalation protocols, and neural classification. The incremental contribution of any single layer, however, remains unquantified. This paper evaluates six LLM models within a DMHI called Wysa, via a controlled comparison of RAG-enabled versus RAG-disabled modes. Anonymized real and synthetic user-chatbot exchanges were annotated by a qualified clinical team against multi-class intent categories (e.g. self-harm, abuse, panic). The study computed classification accuracy, recall, precision and F1 scores against ground truth labels and tested differences for statistical significance. Performance was also examined by risk category and inter-model agreement. While RAG caused a rise in false alarms, the trade-off is consistent with safety-critical design principles that prioritize sensitivity, where flagged cases are routed to additional review rather than acted on directly. Overall, these findings support RAG as a promising approach to improve the accuracy, consistency and safety of LLM-driven DMHIs. Keywords: Digital Mental Health Intervention, Large Language Model, Retrieval Augmented Generation, Accuracy, Recall, Precision

13:00 JST研究/論文

グラフニューラルネットワークを使用した結晶特性予測のためのデュアルレベルの原子および配位幾何学学習

結晶特性の正確な予測は、依然として計算材料科学における重要な課題です。 CGCNN、MEGNet、ALIGNN、SchNet などのグラフ ニューラル ネットワーク (GNN) は優れたパフォーマンスを示していますが、主に原子レベルで結晶を表し、メッセージ パッシングを通じてローカルの化学環境を暗黙的に学習します。ただし、多くの材料特性は、原子とその隣接原子によって形成される基本的な構造単位である配位多面体によって支配されます。この制限に対処するために、原子、結合、および配位多面体の表現を共同で学習するマルチスケール GNN である配位多面体グラフ ネットワーク (CPGN) を提案します。 CPGN は 3 つの結合グラフを構築します。1 つは元素情報と結合情報をコード化する原子グラフ、1 つは角度相互作用を捕捉する折れ線グラフ、1 つは角、エッジ、面の共有関係を通じてボロノイ由来の局所環境を記述する配位多面体グラフです。物理的に意味のある幾何学的記述子が各多面体に組み込まれている一方、双方向クロスアテンションを備えたインターリーブされたメッセージ パッシング メカニズムにより、構造レベル全体で効果的な情報交換が可能になります。 Materials Project、JARVIS-DFT、および QM9 ベンチマーク データセットの広範な評価により、CPGN が既存の最先端の GNN モデルよりも優れていることが実証されています。材料プロジェクトで 0.060 eV/原子の形成エネルギー MAE と 0.292 eV のバンドギャップ MAE を達成すると同時に、JARVIS-DFT で競争力のあるマルチプロパティ予測と QM9 で優れた HOMO 予測を提供します。この結果は、配位多面体の明示的なモデリングにより結晶表現の学習が向上し、材料特性の正確で物理的に解釈可能な予測が可能になることを強調しています。

原文 (English)

Dual-Level Atomic and Coordination Geometry Learning for Crystal Property Prediction Using Graph Neural Networks

Accurate prediction of crystal properties remains a key challenge in computational materials science. While graph neural networks (GNNs) such as CGCNN, MEGNet, ALIGNN, and SchNet have shown strong performance, they primarily represent crystals at the atomic level and implicitly learn local chemical environments through message passing. However, many material properties are governed by coordination polyhedra, the fundamental structural units formed by atoms and their neighboring atoms. To address this limitation, we propose the Coordination Polyhedron Graph Network (CPGN), a multi-scale GNN that jointly learns atomic, bond, and coordination-polyhedron representations. CPGN constructs three coupled graphs: an atom graph encoding elemental and bonding information, a line graph capturing angular interactions, and a coordination polyhedron graph describing Voronoi-derived local environments through corner-, edge-, and face-sharing relationships. Physically meaningful geometric descriptors are incorporated for each polyhedron, while an interleaved message-passing mechanism with bidirectional cross-attention enables effective information exchange across structural levels. Extensive evaluations on the Materials Project, JARVIS-DFT, and QM9 benchmark datasets demonstrate that CPGN outperforms existing state-of-the-art GNN models. It achieves a formation-energy MAE of 0.060 eV/atom and a band-gap MAE of 0.292 eV on the Materials Project, while providing competitive multi-property prediction on JARVIS-DFT and superior HOMO prediction on QM9. The results highlight that explicit modeling of coordination polyhedra improves crystal representation learning and enables accurate, physically interpretable prediction of material properties.

13:00 JST研究/論文

半導体ウェーハ製造向けの動的複数基準ボトルネック重大度指数 (DMBSI): リエントラント生産システム用に遺伝的に最適化されたフレームワーク

ウェーハ製造には、リエントラントなプロセス フロー、可変のボトルネック、非常に可変的なプロセス条件などの独特の特性があります。半導体ウェーハ製造の各時点で最も深刻なボトルネックを特定するために、この研究では、ボトルネック重大度の解釈可能な統一された尺度を生成するために、プロセスパラメータの変化に対するサイクルタイムの複数の診断信号と、サイクルタイムに対する手戻りの影響を分析するための新しいデータ駆動型手法である動的多基準ボトルネック重大度指数 (DMBSI) を提示します。 DMBSI の実験的検証は、Seagate Technology が運営する商用 200 mm ウェーハ製造ラインの 22 ウェーハ生産ロットから収集された製造実行システム (MES) ログを使用して実施されました。 5 分割相互検証を使用して、GA に最適化された DMBSI は、観察されたサイクル時間の寄与と r = 0.80 のピアソン相関を達成します。これは、エキスパート ヒューリスティック ベースライン (r = 0.74) を 8.1% 上回る改善を示し、制約理論 (TOC; r = 0.60) およびバリュー ストリーム マッピング (VSM; r = ) を大幅に上回っています。 -0.30)。さらに、DMBSI の独自の時間ウィンドウ コンポーネントにより、一時的なボトルネック移行パターンの特定が可能になりました。このパターンは、初期の生産期間における誘電体堆積ステップに関連する主要な制約から、後期の生産期間における過剰検査および過小検査に関連する制約に移行しました。統合された what-if 反事実分析により、最上位のボトルネック ステップで待機時間を 50% 削減すると、平均サイクル タイムが 7.2% 短縮され、上位 5 つのボトルネック ステップでは、合計で約 19% の潜在的な削減効果が得られることが実証されました。

原文 (English)

Dynamic Multi-Criteria Bottleneck Severity Index (DMBSI) for Semiconductor Wafer Manufacturing: A Genetically Optimised Framework for Reentrant Production Systems

Wafer fabrication exhibits unique characteristics, including reentrant process flows, variable bottlenecks, and highly variable process conditions. In order to identify the most severe bottleneck at each moment in time for semiconductor wafer fabrication, this research presents the Dynamic Multi-Criteria Bottleneck Severity Index (DMBSI), a new, data-driven methodology for analysing multiple diagnostic signals of cycle time to changes in process parameters, and the impact of reworks on cycle time in order to generate an interpretable, unified measure of bottleneck severity. The experimental validation of DMBSI was conducted using Manufacturing Execution System (MES) logs collected from 22 wafer production lots at a commercial 200 mm wafer fabrication line operated by Seagate Technology. Using 5-fold cross-validation, the GA-optimised DMBSI achieves a Pearson correlation of r = 0.80 with observed cycle-time contributions, representing an 8.1% improvement over the expert heuristic baseline (r = 0.74) and substantially outperforming the Theory of Constraints (TOC; r = 0.60) and Value Stream Mapping (VSM; r = -0.30). Furthermore, the unique time-windowed component of DMBSI enabled the identification of temporal bottleneck migration patterns, which shifted from the dominant constraints associated with dielectric deposition steps in the early production windows to those associated with over- and under-inspection in the later production windows. The integrated what-if counterfactual analysis demonstrated that a 50% reduction in waiting time at the top-ranked bottleneck step would reduce the mean cycle time by 7.2%, with the top five bottleneck steps offering a combined potential reduction of approximately 19%.

13:00 JST研究/論文

EEGの基礎モデルは長距離の時間相関を認識できない:集団間の脆弱性の背後にあるスペクトルと時間の解離

客観的。脳波 (EEG) 基礎モデル (FM) は、短いパッチを再構築または対照的に位置合わせするようにトレーニングされ、固定埋め込みにプールされます。我々は、これらの埋め込みが、アルファバンドエンベロープのトレンド除去変動解析(DFA)指数によって定量化された長距離時間相関(LRTC)を保持しているかどうか、またそれが集団間の移動を支配しているかどうかをテストしました。アプローチ。私たちは、分布外の 2 つのコホートで生波形およびスペクトル入力アーキテクチャ (REVE、LaBraM、BENDR、CBraMod、BIOT) にわたる 5 つの EEG FM を調査し、DFA 指数の回復を静的な 1/f 非周期的傾きと比較しました。順序保持および残差化制御は、プーリングまたは非周期的シャドウイングについてテストされました。モンタージュ調和されたゼロショット転送タスクは、3 つのコホートにわたって凍結埋め込みと DFA 指数を比較しました (西洋の参照を追加)。主な結果。 5 つの FM はどれも、時間順に LRTC を表しませんでした。生波形モデル (REVE、LaBraM、BENDR) は、DFA 指数も 1/f 傾きも回復しませんでした (R^2 <= 0.12)。これら 3 つの場合、プローブは情報を提供しないため、解離はスペクトル入力モデル (CBraMod、BIOT) に特有であり、1/f は強く回復しますが (R^2 = 0.59-0.73)、コホート全体の DFA は回復しません。古典的な DFA 特徴により指数が回復され (信頼性上限 0.64 に対して R^2 = 0.32 ~ 0.38)、LRTC は非周期的な傾きに直交しました (r = -0.06)。集団間の転移では、凍結された REVE 埋め込みは偶然に勝てず (W から K、0.45)、無次元 DFA 指数は方向性を持って転移しましたが、家族ごとの有意性ではありませんでした。他の 4 つはそれを均一に再現しませんでした。 5 つすべてが記録サイト軸 (確率 0.500 に対して 0.98 ~ 1.00 で復号可能) によって支配されていましたが、破棄された DFA 指数はサイト堅牢性 (0.71) でした。

原文 (English)

Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral-Temporal Dissociation Behind Their Cross-Population Fragility

Objective. Electroencephalography (EEG) foundation models (FMs) are trained to reconstruct or contrastively align short patches, then pooled into a fixed embedding. We tested whether these embeddings retained the long-range temporal correlations (LRTC) quantified by the detrended-fluctuation-analysis (DFA) exponent of the alpha-band envelope, and whether it governs cross-population transfer. Approach. We probed five EEG FMs spanning raw-waveform and spectral-input architectures (REVE, LaBraM, BENDR, CBraMod, BIOT) on two out-of-distribution cohorts, comparing recovery of the DFA exponent against the static 1/f aperiodic slope. Order-preserving and residualization controls tested for pooling or aperiodic shadowing. A montage-harmonized, zero-shot transfer task compared the frozen embedding with the DFA exponent across three cohorts (adding a Western reference). Main results. None of the five FMs represented the LRTC in the temporal order. Raw-waveform models (REVE, LaBraM, BENDR) recovered neither the DFA exponent nor the 1/f slope (R^2 <= 0.12); for these three the probe is uninformative, so the dissociation is specific to the spectral-input models (CBraMod, BIOT), which recovered 1/f strongly (R^2 = 0.59-0.73) but not DFA across cohorts. A classical DFA feature recovered the exponent (R^2 = 0.32-0.38 against a 0.64 reliability ceiling), and LRTC was orthogonal to the aperiodic slope (r = -0.06). On cross-population transfer, the frozen REVE embedding did not beat chance (W to K, 0.45) and the dimensionless DFA exponent transferred directionally but not at family-wise significance; the other four did not uniformly replicate it. All five were dominated by a recording-site axis (decodable at 0.98-1.00 vs. 0.500 chance), whereas the DFA exponent they discard is site-robust (0.71).

13:00 JST研究/論文

MedJudgeRAG: 医療 MCQA 向けの動的ナレッジ グラフを使用したオプションに応じた証拠判断

医療多肢選択式質問応答 (MCQA) では、検索拡張生成 (RAG) によって言語モデル (LM) のドメイン知識を補完できます。ただし、バニラ RAG は取得したドキュメントを無差別に利用するため、LM のパフォーマンスが低下する可能性があります。これに対処するために、私たちは MedJudgeRAG を提案します。私たちのフレームワークは、取得されたドキュメントをエンティティとリレーションで構成される動的なナレッジ グラフ (KG) として表現します。各オプションについて、モデルは取得した文書と KG から証拠の評決を判断します。モデルは、判定の組み合わせに基づいて、最終的な答えに向けて推論するための知識活用戦略を決定します。これらの能力は、教師 LM によって生成された構造化推論トレースを使用した教師付き微調整によってトレーニングされます。このトレーニングでは、KG セグメントと推論セグメントに異なる重み付けを行う重み付きクロスエントロピー損失が使用されます。 2 つの医療 MCQA ベンチマークの実験では、MedJudgeRAG がバニラ RAG とパラメトリック ベースラインの両方を一貫して上回るパフォーマンスを示しています。さらに、アブレーション分析により、動的 KG は、推論時の明示的な出力としてよりも、トレーニング時のグラフ条件付き監視としての方が効果的であることが明らかになりました。私たちのコードは https://github.com/hyu-amllab/medjudgerag で入手でき、生成された推論トレースは https://huggingface.co/datasets/youarethewon/medjudgerag でリリースされます。

原文 (English)

MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA

In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language models (LMs). However, since vanilla RAG indiscriminately utilizes retrieved documents, it can degrade LM performance. To address this, we propose MedJudgeRAG. Our framework represents retrieved documents as a dynamic knowledge graph (KG) composed of entities and relations. For each option, the model judges an evidence verdict from the retrieved documents and the KG. Based on the verdict combination, the model determines a knowledge utilization strategy to reason toward the final answer. These capabilities are trained via supervised fine-tuning using structured reasoning traces generated by a teacher LM. The training employs a weighted cross-entropy loss that differentially weights the KG and reasoning segments. Experiments on two medical MCQA benchmarks demonstrate that MedJudgeRAG consistently outperforms both vanilla RAG and parametric baselines. Furthermore, ablation analysis reveals that the dynamic KG is more effective as graph-conditioned supervision at training time than as an explicit output at inference time. Our code is available at https://github.com/hyu-amllab/medjudgerag, and the generated reasoning traces are released at https://huggingface.co/datasets/youarethewon/medjudgerag.

13:00 JSTLLM/生成AI

REPREC: 表現駆動型パラメータ効率的な推奨システム

大規模言語モデル (LLM) は、自然言語タスクとして定式化することにより、順次レコメンデーションに適用されています。これまでの研究では、入力調整または LLM 微調整を通じて協調的かつ順次的な信号を組み込むことで、パーソナライゼーションが向上しました。ただし、既存のアプローチは、LLM の微調整、追加のアーキテクチャ モジュール、表現の蒸留、または長いインタラクション履歴にわたる項目レベルの調整の 1 つ以上に依存することが多く、トレーニングの複雑さと導入コストが増大します。私たちは、軽量のユーザー表現の調整を通じて LLM ベースの逐次レコメンデーションを再定式化する軽量フレームワークである REPREC を提案します。 REPREC は、フリーズされた LLM を調整する軽量の MLP インジェクターを介して、フリーズされたシーケンシャル エンコーダーから埋め込まれた固定サイズのユーザーを学習済みのソフト トークンの小さなセットにマッピングします。これにより、インジェクターのみをトレーニングしながら、両方の事前トレーニングされたバックボーンは変更されません。私たちは複数のベンチマーク データセットに対して徹底的な実験を実施し、REPREC がさまざまな事前トレーニング済みシーケンシャル エンコーダーや LLM バックボーンとの互換性を維持しながら一貫して LoRA を上回るパフォーマンスを示し、事前トレーニング済みコンポーネントを変更することなく、モジュール式で本番環境に適したレコメンデーション パイプラインを実現できることを実証しました。この利点は、すべてのデータセットにわたってカジュアル ユーザーとコア ユーザーで特に顕著であり、データ量が少ない状況での REPREC の有効性が強調されています。最後に、短いプロンプト履歴でトレーニングし、より長いコンテキストで評価した場合、REPREC は LoRA のパフォーマンスの 85 ~ 100% を維持しながら、エポックごとのトレーニング時間を平均 1.51 分の 1 に短縮します。これは、推奨品質と実稼働デプロイメントの計算効率の間の効果的なバランスを示しています。コードは https://github.com/phdbotcode/REPREC で入手できます。

原文 (English)

REPREC: Representation Driven Parameter-Efficient Recommendation System

Large language models (LLMs) have been applied to sequential recommendation by formulating it as a natural language task. Previous work has improved personalization by incorporating collaborative and sequential signals through input conditioning or LLM fine-tuning. However, existing approaches often rely on one or more of the following: LLM fine-tuning, additional architectural modules, representation distillation, or item-level conditioning over long interaction histories, increasing training complexity and deployment cost. We propose REPREC, a lightweight framework that reformulates LLM-based sequential recommendation through lightweight user representation alignment. REPREC maps a fixed-size user embedding from a frozen sequential encoder into a small set of learned soft tokens through a lightweight MLP injector that conditions a frozen LLM, leaving both pretrained backbones unchanged while training only the injector. We conducted exhaustive experiments on multiple benchmark datasets and demonstrate that REPREC consistently outperforms LoRA while remaining compatible with different pretrained sequential encoders and LLM backbones, enabling a modular and production-friendly recommendation pipeline without modifying either pretrained component. The gains are particularly pronounced for casual and core users across all datasets, highlighting REPREC's effectiveness in low-data regimes. Finally, when trained on short prompt histories and evaluated with longer contexts, REPREC maintains 85-100% of LoRA's performance while reducing per-epoch training time by an average of 1.51X, demonstrating an effective balance between recommendation quality and computational efficiency for production deployment. The code is available at https://github.com/phdbotcode/REPREC

13:00 JSTLLM/生成AIQwen

2 つの視点、1 つの意見: 証拠に基づいた会話型音楽の推奨

従来の会話型リコメンダーは、単一のテキスト インターフェイス内で検索と応答の生成を複雑にするため、対話の意図が進化するにつれて正確なエンティティの手がかりが薄れ、説明の信頼性が損なわれます。私たちは、ACM RecSys Challenge 2026 内でこの問題に取り組み、上位 20 位のランキングと証拠に基づく回答の生成の両方を義務付けています。このペーパーでは、ブラインド B 業界トラックのチーム「swyoo」による 3 位のソリューションを紹介します。取得と応答を、ランク付けされたトラックとメタデータを介して厳密に接続された別個のパイプラインに分離します。 Retrieval は、正確な一致を実現するためのハイブリッド語彙密度プールと、微調整された Qwen 8B アダプターによって駆動されるタスクに適応したプールを組み合わせます。候補者は LightGBM を介して調整され、その後、証拠に基づいた提案・割り当て・選択 (PAS) フレームワークにルーティングされて、応答を構造化します。このシステムは、最終ブラインド評価でも説明品質リーダーボードで 2 位にランクされました。私たちの調査結果は、(i) 検索と応答を分離すると、カタログの手がかりと流動的な意図の両方が保存されることを示しています。 (ii) 明示的な証拠の割り当てによる生成の構造化が、このクラス最高に近い説明の信頼性の鍵となります。

原文 (English)

Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation

Traditional conversational recommenders entangle retrieval and response generation within a single text interface, so exact entity cues fade as the dialogue's intent evolves, which compromises explanation credibility. We address this within the ACM RecSys Challenge 2026, which mandates both top-20 ranking and evidence-grounded response generation. This paper presents the third-place solution by team "swyoo" for the Blind-B industry track. We decouple retrieval and response into separate pipelines connected strictly via ranked tracks and metadata. Retrieval combines a hybrid lexical-dense pool for exact matching with a task-adapted pool driven by fine-tuned Qwen 8B adapters. Candidates are calibrated via LightGBM, then routed to an evidence-grounded propose-assign-select (PAS) framework to structure responses. This system also ranked second on the explanation-quality leaderboard in the final blind evaluation. Our findings demonstrate that: (i) isolating retrieval and response preserves both catalog cues and fluid intent; (ii) structuring generation via explicit evidence assignment is key to this near-best-in-class explanation reliability.

13:00 JST研究/論文

極端なチョウラ集合とその線形類似物: Co-Scientist を使用した人間と AI の数学的調査

有限群におけるチョウラ型の秩序条件に関連する極不変量を導入します。有限群 $G$ の空でない部分集合 $S$ は、$S$ のすべての要素が $|S|$ より大きい順序を持つ場合、チョウラ集合と呼ばれ、そのような集合の最大カーディナリティを $C(G)$ と書きます。まず、$C(G)$ が $G$ 内の要素次数の分布によって決定されることを示します。巡回群の場合、正確な約数式を導出し、$C(\mathbb{Z}/n\mathbb{Z})=\varphi(n)$ となる整数 $n$ を特徴付けます。 $\liminf_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=1$ であるのに対し、$\limsup_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=\infty$ であることを証明し、$n$ による正規化の下で対応する下限と上限を決定します。有限アーベル群については、有限アーベル $p$ 群の閉じた式とともに、不変因子分解の観点から明示的な式を取得します。次に、有限体拡張の線形類似物を開発します。拡張 $L/K$ の非ゼロ $K$ 部分空間 $A$ は、A$ 内のすべての非ゼロ $a\ に対して $[K(a):K]>\dim_K A$ の場合、チョウラ部分空間と呼ばれます。この条件は $\dim_K A$ に依存するため、通常、$K$ に対して $L$ を生成するために $A$ のすべての非ゼロ要素は必要ありません。それにもかかわらず、$L/K$ が有限で分離可能な場合、正確な式 $C(L/K)=[L:K]-d_{\max}(L/K)$ を証明します。ここで、$d_{\max}(L/K)$ は、適切な中間体の $K$ に対する最大次数です。有限体については、正規基底構造を使用してあらゆる次数で直接証明を行います。この作品は、専門家の指導の下、人間と AI のコラボレーションを通じて開発されました。 Co-Scientist の推論に重点を置いた構成を使用して、例と潜在的な証明戦略を調査しました。著者たちは問題を定式化し、すべての議論を独自に検証して完成させ、最終的な証明を書きました。

原文 (English)

Extremal Chowla sets and their linear analogues: A human-AI mathematical investigation using Co-Scientist

We introduce an extremal invariant associated with Chowla-type order conditions in finite groups. A nonempty subset $S$ of a finite group $G$ is called a Chowla set if every element of $S$ has order greater than $|S|$, and we write $C(G)$ for the maximum cardinality of such a set. We first show that $C(G)$ is determined by the distribution of element orders in $G$. For cyclic groups, we derive an exact divisor formula and characterize the integers $n$ for which $C(\mathbb{Z}/n\mathbb{Z})=\varphi(n)$. We prove that $\liminf_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=1$, whereas $\limsup_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=\infty$, and we determine the corresponding lower and upper limits under normalization by $n$. For finite abelian groups, we obtain an explicit formula in terms of the invariant-factor decomposition, together with a closed formula for finite abelian $p$-groups. We then develop a linear analogue for finite field extensions. A nonzero $K$-subspace $A$ of an extension $L/K$ is called a Chowla subspace if $[K(a):K]>\dim_K A$ for every nonzero $a\in A$. Since this condition depends on $\dim_K A$, it does not generally require every nonzero element of $A$ to generate $L$ over $K$. Nevertheless, when $L/K$ is finite and separable, we prove the exact formula $C(L/K)=[L:K]-d_{\max}(L/K)$, where $d_{\max}(L/K)$ is the largest degree over $K$ of a proper intermediate field. For finite fields, we give a direct proof in every degree using a normal-basis construction. This work was developed through an expert-guided human-AI collaboration. A reasoning-focused configuration of Co-Scientist was used to explore examples and potential proof strategies. The authors formulated the problem, independently verified and completed all arguments, and wrote the final proofs.

13:00 JST研究/論文

予測精度を超えて: 人間の嗅覚の分子表現の信頼性を意識した監査

事前トレーニングされた分子エンコーダーは通常、下流予測を通じて評価されますが、予測精度だけでは、学習された表現が再現可能な科学構造を捕捉しているか、従来の強力なベースラインを超えて情報を追加しているか、または配布外に転送されているかどうかを確立することはできません。我々は、グローバルな知覚幾何学、化学を超えた増分的予測値、データセット間の複製、目に見えない成分への混合物の移動という 4 つの異なる主張にわたる人間の嗅覚の一般的な分子表現の信頼性を意識した監査を提示します。 Keller-Vosshall および Bierling の単一分子評価データセットと Ma バイナリ混合データセットを使用して、ID 管理および照合評価の下で、MoLFormer および ChemBERTa を RDKit 記述子およびモーガン フィンガープリントと比較します。強度、心地よさ、親近感に基づく人間の 3 属性評価の幾何学的配置は、参加者の分割全体で再現可能ですが (RSA 中央値 0.743 および 0.855)、モデルと人間の一致は大幅に弱くなっています (RSA 0.019 ~ 0.158)。学習された埋め込みは、グローバル アラインメントにおいて従来の表現を常に上回るパフォーマンスを示すわけではなく、MoLFormer は、いずれの単一分子データセットでも RDKit-Morgan ベースラインを組み合わせたものを超える明確な増分予測値を提供しません。人間の幾何学形状は、63 個の共有分子にわたって明確ではあるが不完全な一致を示しています (RSA 0.331; 95% ブートストラップ間隔 [0.204, 0.507])。 1 つの厳密な目に見えないコンポーネントの混合分割の下では、増分効果は結果と表現に依存し、すべての間隔がゼロと交差します。これらの結果は、評価された汎用分子エンコーダーの経験的境界を確立し、より広範な評価原則の動機付けとなります。つまり、科学領域での表現品質は、ターゲットの信頼性、構造アラインメント、増分情報、複製、および配布外転送に関して個別に評価される必要があります。

原文 (English)

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution. We present a reliability-aware audit of generic molecular representations for human olfaction across four distinct claims: global perceptual geometry, incremental predictive value beyond chemistry, cross-dataset replication, and mixture transfer to unseen components. Using the Keller-Vosshall and Bierling single-molecule rating datasets and the Ma binary-mixture dataset, we compare MoLFormer and ChemBERTa against RDKit descriptors and Morgan fingerprints under identity-controlled and matched evaluations. Human three-attribute rating geometry, based on intensity, pleasantness, and familiarity, is reproducible across participant splits (median RSA 0.743 and 0.855), whereas model-human alignment is substantially weaker (RSA 0.019-0.158). Learned embeddings do not consistently outperform conventional representations in global alignment, and MoLFormer provides no clear incremental predictive value beyond a combined RDKit-Morgan baseline in either single-molecule dataset. Human geometry shows positive but incomplete agreement across 63 shared molecules (RSA 0.331; 95% bootstrap interval [0.204, 0.507]). Under one strict unseen-component mixture split, incremental effects are outcome- and representation-dependent, with all intervals crossing zero. These results establish empirical boundaries for the evaluated generic molecular encoders and motivate a broader evaluation principle: representation quality in scientific domains should be assessed separately for target reliability, structural alignment, incremental information, replication, and out-of-distribution transfer.

13:00 JSTLLM/生成AI画像/動画生成

DisasterTD: マルチモーダル LLM とクロスビュー地理位置推定を使用した災害地名曖昧さ回避

ソーシャル メディア画像 (SMI) は、状況認識と緊急対応に貴重な、タイムリーで詳細な地上の視点を提供します。衛星画像や航空画像とは異なり、SMI は災害の影響や地上の状況をタイムリーに捉えることができます。ただし、SMI での地理的参照は曖昧または曖昧であることが多く、正確な地理的位置特定が困難になります。この問題に対処するために、マルチモーダル大規模言語モデル (MLLM) ベースの意味論的推論とクロスビュー地理位置特定を統合する、災害地名曖昧さ回避フレームワークである DisasterTD を提案します。まず、MLLM は地名を抽出し、ノイズの多いテキスト入力から地理位置情報の候補を生成します。次に、SMI、リモート センシング画像 (RSI)、およびオプションでストリートビュー画像 (SVI) 間のクロスビュー マッチングを使用して、これらの候補結果を検証し、絞り込みます。ハリケーン ハーベイ データセットで DisasterTD を評価します。SMI は収集された RSI と SVI で強化され、災害の地理的位置特定のためのクロスビュー ベンチマークを構築します。データセットは、トポニムの明確さと曖昧さに基づいて 4 つのカテゴリに分類されており、シナリオ全体にわたる詳細なパフォーマンス分析が可能です。結果は、DisasterTD が曖昧さなく一貫して MLLM のみおよびクロスビューのみのベースラインを上回り、平均と誤差の中央値はそれぞれ 11.33 km と 0.68 km です。最大の改善は曖昧な地名に現れ、横断的な証拠による意味論的推論により候補の分散とエラーが減少します。これらの発見は、きめ細かい災害地理位置特定のための、MLLM ベースの候補生成とクロスビュー検証の統合の有効性を示しています。

原文 (English)

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization

Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose DisasterTD, a disaster toponym disambiguation framework that integrates multimodal large language model (MLLMs)-based semantic reasoning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery (RSI), and optionally street-view imagery (SVI) is used to verify and refine these candidate results. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambiguation, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 km and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.

13:00 JST研究/論文

HVM-GraphRAG: 複雑なドキュメントに対するホリスティックビューのマルチモーダルグラフ検索拡張生成

複雑な文書に対する質問応答 (QA) には、離れた文書領域およびモダリティに分散された証拠を取得して統合するモデルが必要です。 Multimodal GraphRAG は、ドキュメントの証拠をグラフ構造で整理することで、有望な方向性を提供します。ただし、既存の方法では、信頼性の低いクロスモーダル証拠のインデックス作成や高価なグラフ走査に悩まされることがよくあります。これらの問題に対処するために、複雑なドキュメントに対する全体的なビューのマルチモーダル GraphRAG フレームワークである HVM-GraphRAG を提案します。 HVM-GraphRAG は全体的なビューを使用してグラフ構築をガイドし、それによってノイズの多い競合するグラフ更新を削減し、概念レベルのグラフ ノード間で信頼性の高いインデックスを構築し、マルチモーダル チャンクをサポートします。取得中、HVM-GraphRAG はコンパクトな概念レベルのグラフを検索し、構築されたインデックスを通じて裏付けとなる証拠に直接アクセスすることで、高密度のエンティティ レベルのグラフにわたるコストのかかる走査を回避します。取得した証拠を取得した後、HVM-GraphRAG はチャンクをモダリティ固有のグループにさらに再編成し、応答モデルが異種の証拠をより適切に統合できるようにします。 3 つのデータセットでの実験では、HVM-GraphRAG が、代表的なグラフベースのベースラインと比較してオンライン検索効率を大幅に向上させながら、ほとんどの評価設定で最良の回答パフォーマンスを達成することが示されています。

原文 (English)

HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document

Question answering (QA) over complex documents requires models to retrieve and integrate evidence distributed across distant document regions and modalities. Multimodal GraphRAG provides a promising direction by organizing document evidence with graph structures. However, existing methods often suffer from unreliable cross-modal evidence indexing and expensive graph traversal. To address these issues, we propose HVM-GraphRAG, a holistic-view multimodal GraphRAG framework on complex document. HVM-GraphRAG uses a holistic view to guide graph construction, thereby reducing noisy and conflicting graph updates and building reliable indices between concept-level graph nodes and supporting multimodal chunks. During retrieval, HVM-GraphRAG searches over a compact concept-level graph and directly accesses supporting evidence through the constructed index, avoiding costly traversal over dense entity-level graphs. After obtaining the retrieved evidence, HVM-GraphRAG further reorganizes chunks into modality-specific groups, enabling the answering model to better integrate heterogeneous evidence. Experiments on three datasets show that HVM-GraphRAG achieves the best answer performance in most evaluated settings while substantially improving online retrieval efficiency over representative graph-based baselines.

13:00 JSTLLM/生成AI

必要なのはトークンだけ: レコメンデーション システムで LLM レベルの I/O 効率を達成するための二重目的セマンティック ID

大規模なレコメンデーション システムは、大規模で高密度の埋め込みテーブルによる「メモリの壁」のボトルネックに直面しています。生成検索では ID に個別のトークンが使用されますが、高次元コンテキストは依然として非効率的な高密度フォーマットに依存しています。コンピューター ビジョン データ圧縮からインスピレーションを得て、LLM レベルの I/O 効率を達成するためのデュアルパーパス セマンティック ID を提案します。私たちの方法論では、階層的な量子化を使用して、連続的な埋め込みを離散的なセマンティック ID に凝縮し、次の 2 つの役割を同時に実行します。(1) 協調的アイデンティティ: 学習可能な埋め込みテーブルを介してユーザーとアイテムの相互作用をモデル化します。 (2) コンテンツ再構築: オンザフライ埋め込み近似のための軽量のセマンティック デコーダーを使用します。このアプローチは、大規模なベクトル ストレージをオンデマンドの再構築に置き換え、システムのオーバーヘッドとデータ フットプリントを削減します。私たちは、オフライン評価と、主要なビデオ共有プラットフォームにおける実稼働規模のランキングおよび検索システムへのオンライン導入の成功を通じて、フレームワークの有効性を実証し、非常に効率的でコンテンツ豊富なレコメンデーションに実際に必要なのは個別のトークンだけであることを示しています。

原文 (English)

Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems

Large-scale recommendation systems face "Memory Wall" bottlenecks due to massive, dense embedding tables. While generative retrieval uses discrete tokens for IDs, high-dimensional context still relies on inefficient dense formats. Inspired by computer vision data compression, we propose Dual-purpose Semantic IDs to achieve LLM-level I/O efficiency. Our methodology uses hierarchical quantization to condense continuous embeddings into discrete Semantic IDs performing two concurrent roles: (1) Collaborative Identity: modeling user-item interactions via learnable embedding table; and (2) Content Reconstruction: using a lightweight Semantic Decoder for on-the-fly embedding approximation. This approach replaces massive vector storage with on-demand reconstruction, reducing system overhead and data footprints. We demonstrate the efficacy of our framework through offline evaluations and successful online deployment in production-scale ranking and retrieval systems at a major video sharing platform, showing that discrete tokens are indeed all you need for highly efficient, content-rich recommendation.

13:00 JST研究/論文DeepSeek

GraphRareBench: 表現型に基づく希少疾患診断のための監査可能なグラフ証拠ベンチマーク

表現型に基づく診断ベンチマークは通常、参照疾患のランクを報告しますが、どの妥当な代替案がその上位にランクされているか、あるいはツールを使用するモデルが決定を下す前にどのような証拠を検査しているかを明らかにすることはほとんどありません。 2,365 のオントロジー由来のケースと 18,093 のターゲットと交絡因子のペアを含む来歴保存ベンチマークである GraphRareBench を紹介します。各ケースには、粗雑な HPO クエリ、固定された候補プール、グラフで定義されたハード交絡因子、およびソースにリンクされた証拠レコードが含まれます。 237 ケースの遺伝子コンポーネント素テスト分割では、共有 21 特徴インターフェイスを使用した教師付きランカーは、0.640 ~ 0.740 の範囲の MRR と、0.898 ~ 0.916 の範囲のケース平均標的対交絡因子の精度を達成しました。 Agents-A1 と DeepSeek-V4-Flash でインスタンス化されたエージェントは、それぞれ 0.746 と 0.718 の MRR を達成しました。それらのペアの MRR の差は統計的に有意ではありませんでしたが、ターゲットと証拠の範囲には 0.561 の差がありました。選択された Hit@10 成功の 22.1% ~ 43.7% が依然としてグラフ定義のハード交絡因子の少なくとも 1 つをターゲットより上にランク付けしているという観察と合わせて、これらの結果は、フルプール検索、ハード交絡因子の識別、および観察可能な証拠へのアクセスがモデルの動作の補完的な側面を捉えていることを示しています。したがって、GraphRareBench は、表現型駆動型診断システムのより透明性が高く、証拠を意識した評価のための基盤を提供します。コードとデータは https://github.com/GUI0609/GraphRareBench で入手できます。

原文 (English)

GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis

Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranked above it or what evidence a tool-using model examines before making its decision. We introduce GraphRareBench, a provenance-preserving benchmark containing 2,365 ontology-derived cases and 18,093 target-confounder pairs. Each case includes a coarsened HPO query, a fixed candidate pool, graph-defined hard confounders, and source-linked evidence records. On the 237-case gene-component-disjoint test split, supervised rankers using a shared 21-feature interface achieved MRRs ranging from 0.640 to 0.740 and case-averaged target-over-confounder accuracies ranging from 0.898 to 0.916. Agents instantiated with Agents-A1 and DeepSeek-V4-Flash achieved MRRs of 0.746 and 0.718, respectively. Their paired MRR difference was not statistically significant, whereas their target-evidence coverage differed by 0.561. Together with the observation that 22.1% to 43.7% of selected Hit@10 successes still ranked at least one graph-defined hard confounder above the target, these results indicate that full-pool retrieval, hard-confounder discrimination, and observable evidence access capture complementary aspects of model behavior. GraphRareBench therefore provides a foundation for more transparent and evidence-aware evaluation of phenotype-driven diagnostic systems. Code and data are available at https://github.com/GUI0609/GraphRareBench.

13:00 JST研究/論文

人間の好みに合わせた表の類似性

タスクに依存しない表形式の埋め込みは、製品ライフサイクル管理 (PLM) などの実世界のビジネス システムにおける類似性検索に使用されることが増えています。ただし、主要な埋め込みアプローチは、人間の好みに合わせた類似性ランキングを生成するためではなく、主に予測タスクのために最適化されています。私たちは、標準的な下流メトリクスでは類似性検索の埋め込みの信頼性を完全に評価するには不十分であり、人間の好みに合わせた評価は必要だが現在欠落しているコンポーネントであると主張します。具体的な評価手順を示し、PLM の使用例を通じて問題を説明します。

原文 (English)

Human Preference aligned Tabular Similarity

Task-agnostic tabular embeddings are increasingly used for similarity search in real-world business systems such as Product Lifecycle Management (PLM). However, leading embedding approaches are optimized primarily for prediction tasks - not for producing human preference aligned similarity rankings. We argue that standard downstream metrics are insufficient to fully assess embedding trustworthiness for similarity search and that human preference aligned evaluation is a necessary and currently missing component. We present a concrete evaluation procedure and illustrate the problem through a PLM use case.

13:00 JSTLLM/生成AIエージェント

エージェント取得ベンチ: コーディング エージェントのリポジトリ コンテキスト取得の評価

最新のコーディング エージェントは通常、最終的に正しいパッチを生成するかどうかによって評価されますが、パッチの生成は、タスクに必要なリポジトリ ファイルを見つけるという、初期のコンテキスト取得段階に依存します。この上流の取得問題に対するファイルレベルのベンチマークである Agent Retrieval Bench を紹介します。サンプルは実際のコーディング ワークフロー信号から構築され、凍結されたベース コミット リポジトリに対して評価されます。関連性は、クエリ ファイルの直接的な意味の類似性ではなく、エージェントが次に必要とするものによって定義されます。このベンチマークは、code2test、comment2context、trace2code、edit2ripple の 4 つのポジティブ検索タスクを対象としています。 5 番目のサブセットは、自然証拠に裏付けられた金なしケースと反事実的な間違ったリポジトリ コントロールを使用して、選択的検索を評価します。 Agent Retrieval Bench には、25 のリポジトリにわたる 427 のサンプルが含まれています。そのうち 345 の陽性例、50 の自然なゴールドなしの例、および 32 の反事実対照です。コーパスには、308 個のベースコミット スナップショット、392,000 個のファイル、および 790 万個のチャンクが含まれています。語彙検索、RepoMap、オープンソースの埋め込み、選択的棄権、およびログに記録されたエージェント コンテキストの選択を評価します。単一の検索ファミリーが優勢になることはありません。Qwen3-Embedding-4B は陽性サンプルで最高のサンプル重み付け MRR を持ち、Qwen3-Embedding-8B は最高の Recall@20 を持ち、RepoMap は 8K トークンで最高の予算コンテキスト収量を持ち、タスク レベルの勝者は大幅に異なります。反事実対照で校正された選択的閾値は、自然な金のないケースでの選択的成功を改善せず、校正ギャップが明らかになります。記録された軌跡では、サンプルの 27 ~ 35% にあるすべてのゴールド ファイルも欠落しています。制御されたシード介入パイロットでは、取得由来の初期コンテキストが、ランダムな非ゴールド コンテキストよりもシード後の探索が少なく、より高いファイル F1 を生成する一方、オラクル ゴールド コンテキストにはかなりのヘッドルームが残っていることがわかりました。

原文 (English)

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task. We introduce Agent Retrieval Bench, a file-level benchmark for this upstream retrieval problem. Samples are built from real coding-workflow signals and evaluated against frozen base-commit repositories, with relevance defined by what an agent needs next rather than direct query-file semantic similarity. The benchmark covers four positive-retrieval tasks: code2test, comment2context, trace2code, and edit2ripple; a fifth subset evaluates selective retrieval using natural evidence-backed no-gold cases and counterfactual wrong-repository controls. Agent Retrieval Bench contains 427 samples across 25 repositories: 345 positive examples, 50 natural no-gold examples, and 32 counterfactual controls. The corpus includes 308 base-commit snapshots, 392,000 files, and 7.9 million chunks. We evaluate lexical retrieval, RepoMap, open-source embeddings, selective abstention, and logged agent context selection. No single retrieval family dominates: Qwen3-Embedding-4B has the best sample-weighted MRR on positive samples, Qwen3-Embedding-8B the best Recall@20, and RepoMap the best budgeted context yield at 8K tokens, with task-level winners differing substantially. Selective thresholds calibrated with counterfactual controls do not improve selective success on natural no-gold cases, revealing a calibration gap. Logged trajectories also miss every gold file on 27-35 percent of samples. A controlled seed-intervention pilot finds that retrieval-derived initial context yields higher file F1 with less post-seed exploration than random non-gold context, while oracle gold context shows substantial remaining headroom.

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

「何を取得するか」を超えて: 取得拡張コード生成における不確実性

リポジトリ レベルのコード生成は、関連性、互換性、完全性が本質的に不確実な異種の証拠に依存しています。類似のコード例、リポジトリ コンテキスト、およびプロジェクト固有の API は補完的な情報を提供する可能性がありますが、ノイズの多い、冗長な、または競合する信号を引き起こす可能性もあります。既存の検索拡張アプローチは、検索された証拠の不確実性が下流の生成にどのような影響を与えるかを明示的にモデル化することなく、主に検索の関連性を最適化します。ソース固有の不確実性を推定し、それを使用して異種証拠をフィルタリングしてランク付けし、生成、検証、修復をガイドする不確実性認識フレームワークである OpenCoder を紹介します。 API の知識、リポジトリのコンテキスト、類似コードの証拠に対する要因分析では、普遍的な追加ソース ランキングがないことが明らかになりました。代わりに、重要なソース間の相互作用は、付随する証拠と LLM バックエンドに依存します。拡張された 32 タスクの RepoExec インライン評価では、OpenCoder はベースライン RAG に対する GPT 選択出力の正確性を 56.25\% から 78.13\% に向上させます。ただし、これは検証と修復の制御と一致しており、対応する Gemini の改善は統計的にサポートされておらず、バックエンドに依存する利点が示されています。ターゲットを意識した API の改良により、API セットの取得も大幅に向上します。これらの発見は、不確実性をリポジトリレベルの検索、検証、修復のための実用的な制御信号として扱うことを裏付けています。

原文 (English)

Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation

Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain. Similar-code examples, repository context, and project-specific APIs may provide complementary information, but can also introduce noisy, redundant, or conflicting signals. Existing retrieval-augmented approaches primarily optimize retrieval relevance without explicitly modeling how uncertainty in retrieved evidence affects downstream generation. We introduce OpenCoder, an uncertainty-aware framework that estimates source-specific uncertainty, uses it to filter and rank heterogeneous evidence, and guides generation, verification, and repair. A factorial analysis over API knowledge, repository context, and similar-code evidence reveals no universal additive source ranking; instead, significant cross-source interactions depend on the accompanying evidence and LLM backend. On an expanded 32-task RepoExec-inline evaluation, OpenCoder improves GPT selected-output correctness over Baseline RAG from 56.25\% to 78.13\%. However, it matches a verification-and-repair control, and the corresponding Gemini improvement is not statistically supported, indicating backend-dependent benefits. Target-aware API refinement also substantially improves API-set retrieval. These findings support treating uncertainty as an actionable control signal for repository-level retrieval, verification, and repair.

13:00 JST研究/論文

伝播遅延の排除: 交通流予測のためのアテンションベースの時空間融合グラフ畳み込みネットワーク

交通の流れを予測することは、交通システムを最適化し、都市のモビリティを向上させるために重要です。時空間特徴を抽出し、交通流を予測するために、多くのグラフ畳み込みベースのモデルが提案されています。ただし、ほとんどはトポロジカルな関係における時空間的および意味的相関に焦点を当てています。対処すべき主な問題が 2 つあります。まず、モデルの畳み込み構造は、畳み込み内の隣接ノード間のさまざまな情報伝播遅延を無視しながら、トポロジカル構造における静的な空間依存性と時空間関係を利用することに焦点を当てています。第 2 に、これらの方法は多くの場合、複雑な構造を多数積み上げるため、モデルのトレーニング段階での計算時間が大幅に増加し、その結果モデルの適時性の要件が無視されます。この論文では、アテンションベースの時空間融合グラフ畳み込みネットワーク (A-STFGCN) と呼ばれる新しいネットワークを提案します。伝播遅延誤差が除去された時空間特徴相関を抽出し、マスク行列に基づくマルチヘッド セルフ アテンション メカニズム内でデータの長期および短期の両方の時間特性を捕捉するための時空間融合ブロックを設計します。 5 つの現実世界のデータセットに対する広範な実験により、私たちの方法が 8 つのベースライン方法と比較して優れた計算効率とデータ利用効率を備えながら、最高の全体的なパフォーマンスを達成できることが実証されました。

原文 (English)

Eliminating Propagation Delay: Attention-Based Spatial-Temporal Fusion Graph Convolution Network for Traffic Flow Prediction

Predicting traffic flow is crucial to optimizing transportation systems and improving urban mobility. Many graph convolution-based models have been proposed to extract spatial-temporal features and predict traffic flow. However, most focus on spatial-temporal and semantic correlation in topological relationships. There are two primary problems to address. Firstly, the convolutional structure in the model focuses on utilizing static spatial dependencies and spatial-temporal relationships in topological structures, while neglecting the different information propagation delays between adjacent nodes in the convolution. Secondly, these methods often stack a large number of complex structures, resulting in a substantial increase in computational time during the model training phase, thereby disregarding the model's requirements for timeliness. In this paper, we propose a novel network called the Attention-Based Spatial-Temporal Fusion Graph Convolution Network (A-STFGCN). We design a spatial-temporal fusion block to extract the spatial-temporal feature correlations with propagation delay errors removed and to capture both long-term and short-term temporal characteristics of the data within a multi-head self-attention mechanism based on a mask matrix. Extensive experiments on five real-world datasets demonstrate that our method achieves the best overall performance while having good computation and data utilization efficiency compared with the eight baseline methods.

13:00 JST研究/論文Llama

正規化残差ネットワークにおける幅スケーリングのメカニズム: 有効なアライメント次元

ニューラル ネットワーク幅に関する既存の理論は、漸近限界を特徴づけていますが、有限のトレーニング データから特定された拡張方向が目に見えないデータに対しても有益であり続けるかどうかについては限定的なガイダンスしか提供していません。我々は、関数を保存した残差拡張についてこの問題を研究し、活性化勾配の信号とノイズの幾何学構造を記述する測定可能な量である有効整列次元を導入します。独立して推定されたトレーニング勾配とテスト勾配の間の内積の正確な平均と分散を導出することにより、不整合確率の有限サンプルの上限が得られます。この限界は有効なアライメント次元と有効なサンプル サイズのみに依存し、共分散スペクトルの仮定や所定の幅増加率を必要とせず、有限の二次モーメントと非ゼロの母集団勾配を必要とします。この証明書をトレーニングテスト残差拡張フレームワークに統合し、テストリスクを改善するための高確率の条件を生成します。幅制御された LLaMA スタイルの Transformers、Pythia、および ResNet-20 にわたる実験では、幅の広いモデルほど有効なアライメント次元が大きく、経験的なミスアライメントが低いことが示されています。直接残差介入により、アラインメント統計がホールドアウト損失変化の符号と大きさを予測することが確認されます。

原文 (English)

Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension

Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction identified from finite training data remains beneficial on unseen data. We study this problem for function-preserving residual expansion and introduce the effective alignment dimension, a measurable quantity describing the signal-noise geometry of activation gradients. By deriving the exact mean and variance of the inner product between independently estimated training and test gradients, we obtain a finite-sample upper bound on misalignment probability. The bound depends only on the effective alignment dimension and an effective sample size, requiring finite second moments and a nonzero population gradient, without covariance spectral assumptions or prescribed width-growth rates. We integrate this certificate into the train-test residual-expansion framework, yielding a high-probability condition for test-risk improvement. Experiments across width-controlled LLaMA-style Transformers, Pythia, and ResNet-20 show that wider models exhibit larger effective alignment dimensions and lower empirical misalignment. Direct residual interventions confirm that the alignment statistic predicts the sign and magnitude of held-out loss changes.

13:00 JSTエージェントビジネス/資金調達

ゲージ: 黄金の答えのないグレーディング エージェント構築の財務モデル

財務モデルは、一般公開情報とアナリストの仮定を組み合わせて、予測と評価を生成します。一部のコンポーネントは機械的にチェックできますが、予測、割引率、目標価格には複数の合理的な答えが得られることがよくあります。それにもかかわらず、既存のベンチマークは、単一の専門家の参照に照らしてそのような出力を評価する傾向があります。同じ企業に対して独自に構築したアナリスト モデルを使用したところ、65 社をカバーする 108 の有向ペア全体で、単一参照スコアの中央値は 0.33 で、92.6% のスコアが 0.70 未満であり、暗示価格が 10% 以内で一致する同じヴィンテージのペアは存在しないことがわかりました。したがって、ポイントトレランスグレーディングは、専門家の間にすでに存在する意見の相違にペナルティを与える可能性があります。単一点の回答ではなく、観察されたアナリストの実践に対してエージェントが構築した評価モデルを評価するためのベンチマークである GAUGE を紹介します。 GAUGE は、ベンダー別に分類された 1,001 のアナリスト ワークブックと 196 のタスク評価セットを使用し、3 層の観察された実践エンベロープ、56 の監査可能なファセット、8 つの妥当性ゲート、および決定論的な構造チェックを備えています。私たちは、55 人の参加者による既知グループの調査、企業グループのクロスフィッティング、および裁判官の安定性監査によってベンチマークを検証します。失敗認識スコア $\phi_0$ では、上級アナリストは平均 88.3、若手アナリストは 66.0、金融学生は 43.2 でした。 24 人のエージェントと 1,011 人のスコア世代全体で、最高のエージェントのスコアは 53.4 で、学生の平均よりも上ですが、すべての上級生とほとんどの後輩よりも下回っています。機械面の 93%、判定面の 78% を通過し、フリートと中央値の差は 26 ポイントでした。現在のエージェントは、評価判断よりもモデル構築の方が大幅に優れています。私たちは、この方法論、ゲート付き匿名化データ層、制御されたトレーニング分割、バージョン管理された 48 タスクの評価コア、および保留されたリフレッシュ プールをリリースします。

原文 (English)

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $\phi_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

13:00 JSTLLM/生成AI

予測プランナーとしての LLM: 時系列基礎モデルのトレーニング不要のテキスト コンディショニング

テキスト条件付き時系列予測は、数値履歴と自然言語コンテキストの両方から系列を予測し、過去だけでは明らかにできないイベントや制約を予測で考慮できるようにします。これには、信頼性の高い数値予測とコンテキスト情報を解釈する能力の両方が必要です。時系列基礎モデル (TSFM) は強力な数値予測を提供し、大規模言語モデル (LLM) はテキストを推論できますが、LLM に予測値の直接生成または修正を要求すると、TSFM によって取得される時間構造が歪む可能性があるため、これらの長所を組み合わせるのは依然として困難です。代わりに、TSFM によって生成された軌道に対する計画問題として予測を定式化します。凍結された TSFM は数値的な継続を提案するシミュレーターとして機能し、LLM は候補の選択をガイドし、コンテキストに照らして完了した軌跡を評価するポリシーおよび価値関数として機能します。これを \rc{} (\textbf{L}LM \textbf{A}s \textbf{F}orecasting \textbf{P}lanner) としてインスタンス化します。これは、どちらのモデルも再トレーニングせずにモダリティのギャップを埋める、トレーニング不要のフレームワークで、 \emph{Ranker} LLM をポリシーとして、\emph{Judge} LLM を値関数として予測期間にわたるモンテカルロ木探索 (MCTS) を使用します。 2 つの TSFM バックボーン (Chronos と TimesFM) と 4 つの LLM にわたる Context-is-Key と Time-MMD の実験では、\rc{} がモデルの選択全体で一貫した改善をもたらし、テキスト条件付き予測に対するトレーニング不要の効果的なアプローチとして順次検索をサポートすることが示されました。

原文 (English)

LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models

Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts to account for events and constraints that the past alone cannot reveal. This requires both reliable numerical forecasting and the ability to interpret contextual information. Time-series foundation models (TSFMs) provide strong numerical forecasts, while large language models (LLMs) can reason over text, but combining their strengths remains challenging because asking an LLM to generate or revise forecast values directly can distort the temporal structure captured by the TSFM. We instead formulate forecasting as a planning problem over TSFM-generated trajectories. The frozen TSFM acts as a simulator that proposes numerical continuations, while the LLM acts as a policy and value function that guides candidate selection and evaluates completed trajectories against the context. We instantiate this as \rc{} (\textbf{L}LM \textbf{A}s \textbf{F}orecasting \textbf{P}lanner), a training-free framework that bridges the modality gap without retraining either model, using Monte Carlo tree search (MCTS) over the forecast horizon with a \emph{Ranker} LLM as policy and a \emph{Judge} LLM as value function. Experiments on Context-is-Key and Time-MMD across two TSFM backbones (Chronos and TimesFM) and four LLMs show that \rc{} delivers consistent improvements across model choices, supporting sequential search as an effective training-free approach to text-conditioned forecasting.

13:00 JSTLLM/生成AIエージェント

マルチエージェント LLM システムにおける分散バックドアの早期検出: 特性調査

マルチエージェント LLM システムは、単一のエージェントが完全に保持することのないペイロードによって攻撃される可能性があります。つまり、毒されたツールは、監視内の暗号化されたフラグメントを隠し、それらを複数のエージェントに分散させ、実行後に外部ステップがそれらを再構築して実行します。各アクションを個別に判断するステップごとの安全性チェックでは、完全な分散ペイロードを認識できない可能性があります。私たちは、そのような攻撃がまだ進行中である間にどれだけ早く検出できるか、そして最も明白な手がかりが取り除かれた後にそれをどれだけ確実に捕捉できるかを調査します。階層型マルチエージェント システム上に動作インスタンスを構築し、5 つの言語モデルと 2 つのタスク ドメインにわたって無害な条件と攻撃された条件下で実行し、各フラグメントがいつ挿入されるか、いつペイロードがアセンブルされて実行されるかを記録します。検出はアセンブリとの競争です。最初のフラグメントが注入される前は、攻撃されたランと良性のランは区別できません。インジェクションが開始されると、プレフィックス検出器は、残り 5 ステップの中央値とセーフラン誤検知率 $10.3\%$ の成功した攻撃に対して $99.3\%$ のフラグを立てます。組み立ては実行後にのみ行われるため、これらのアラームは、成功したほぼすべての攻撃を中止するのに間に合うように到着します。次に、その警告のどれだけが、分散構造ではなく、攻撃の除去可能な表面上の手がかりに基づいているかを測定します。一般的なゼロショット検出器と動作訓練された検出器は、ほとんど警告を出しません。機能する検出器は、除去可能な表面キュー、主に暗号文の長さとエントロピーに部分的に依存しており、エントロピー キューがペイロードから削除され、長さの特徴が検出器から削除されると、検出は後で到着し、ドメイン間での転送が不十分になりますが、微調整されたモデルによって損失の一部が回復されます。

原文 (English)

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run. Per-step safety checks that judge each action in isolation may fail to recognize the complete distributed payload. We investigate how early such an attack can be detected while the run is still unfolding, and how robustly it can be caught once its most obvious cues are stripped away. We build a working instance on a hierarchical multi-agent system, run it under benign and attacked conditions across five language models and two task domains, and record when each fragment is injected and when the payload is assembled and executed. Detection is a race against assembly. Before the first fragment is injected, attacked and benign runs are indistinguishable; once injection begins, a prefix detector flags $99.3\%$ of successful attacks with a median of five steps remaining and a $10.3\%$ safe-run false-positive rate. Because assembly occurs only after the run, these alarms arrive in time to abort nearly every successful attack. We then measure how much of that warning rests on removable surface cues of the attack rather than on its distributed structure. Generic zero-shot and behavior-trained detectors provide almost no warning at all; the detectors that do work lean in part on removable surface cues, chiefly the ciphertext's length and entropy, and once the entropy cue is removed from the payload and the length features from the detector, detection arrives later and transfers poorly across domains, though a fine-tuned model recovers some of the loss.

13:00 JST研究/論文

特徴空間摂動下でのマルウェア表現の潜在安定性分析

静的マルウェア検出器は通常、精度、F1、ROC AUC、PR AUC などのクリーンサンプルのメトリクスを使用して評価されます。ただし、これらのメトリクスは、特徴ベクトルが摂動されたときに学習されたマルウェア表現がどのように動作するか、サンプルが不確実な決定領域にどのように近づくか、または圧縮された表現がセキュリティ関連の構造を保持しているかどうかについて限定的な洞察を提供します。このペーパーでは、EMBER 機能空間におけるマルウェア摂動評価のための潜在安定性分析パイプラインを紹介します。このパイプラインは、完全な EMBER 機能、PCA ベースの圧縮、ベータ/ノイズ除去変分オートエンコーダー表現、マンデルブロにインスピレーションを得たエスケープタイム記述子、PINN スタイルの潜在フロー モジュールを比較します。摂動下での脱出時間プロファイルの変化を測定するために潜在脱出ダイバージェンス (LED) を定義し、PINNFlow から導出された残差、速度、リスク、および勾配シフトのメトリクスを使用して潜在運動を特徴付けます。実験は、180,000 個のトレーニング サンプル、180,000 個のテスト サンプル、および 240,000 個のホールドアウト サンプルを使用して、EMBER 静的 PE 特徴ベクトルに対して実行されます。完全な EMBER 機能は、ROC AUC 0.9962 および F1 0.9713 で最も強力なクリーン分類パフォーマンスを実現します。一方、PCA-64 は ROC AUC 0.9846 および F1 0.9347 で最も強力な圧縮ベースラインです。提案された VAE+Mandelbrot+PINNFlow 表現は、クリーンな分類に関してはこれらのベースラインを上回る性能を備えていませんが、制御された特徴空間摂動プローブの下で追加の診断価値を提供します。

原文 (English)

Latent Stability Analysis of Malware Representations Under Feature-Space Perturbations

Static malware detectors are commonly evaluated using clean-sample metrics such as accuracy, F1, ROC AUC, and PR AUC. However, these metrics provide limited insight into how learned malware representations behave when feature vectors are perturbed, how close samples move toward uncertain decision regions, or whether compressed representations preserve security-relevant structure. This paper presents a latent-stability analysis pipeline for malware perturbation assessment in EMBER feature space. The pipeline compares full EMBER features, PCA-based compression, beta/denoising variational autoencoder representations, Mandelbrot-inspired escape-time descriptors, and a PINN-style latent-flow module. We define Latent Escape Divergence (LED) to measure changes in escape-time profiles under perturbation, and use PINNFlow-derived residual, velocity, risk, and gradient-shift metrics to characterize latent movement. Experiments are conducted on EMBER static PE feature vectors using 180,000 training samples, 180,000 test samples, and 240,000 holdout samples. Full EMBER features achieve the strongest clean classification performance with ROC AUC of 0.9962 and F1 of 0.9713, while PCA-64 is the strongest compressed baseline with ROC AUC of 0.9846 and F1 of 0.9347. The proposed VAE+Mandelbrot+PINNFlow representation does not outperform these baselines for clean classification, but it provides additional diagnostic value under controlled feature-space perturbation probes.

13:00 JST画像/動画生成GPT / ChatGPT

害は普遍的ではない: コミュニティ固有の毒性検出が緊急に必要である

テキストから画像への生成のための最先端の毒性検出器は、すべてのユーザーに固定の安全ガイドラインを適用する単一のユニバーサル モデルという、万能のアプローチを採用しています。私たちの経験的証拠は、これらの検出器が疎外されたコミュニティを保護できないことを示しています。安全であるとラベル付けされた生成された画像の約 35% が、障害のあるコミュニティによって有害で​​あると考えられています。この意見書では、コミュニティ特異的毒性検出 (CTD) について主張します。その実現可能性を実証するために、私たちは障害の専門家と協力して、小人症と盲目/弱視という 2 つのコミュニティのための安全ガイドラインを作成しています。注釈付きの T2I 生成画像 2,400 枚のデータセットを使用して、大規模な視覚言語モデルと既存の汎用毒性検出器の両方が、ランダムな推測よりも低い F1 スコア (F1 0.32 および 0.37) のゼロショット設定では、これらのガイドラインの下で有害なコンテンツを認識できないことを実証します。有望なのは、プロンプトベースの適応方法 (ICL、VQA) により、危害検出パフォーマンス (GPT-4o: F1 0.50 および 0.78) が大幅に向上する一方、パラメーター効率の高い微調整により、100 未満のデモンストレーションで小規模なモデル (最良の F1 0.48 および 0.59 を備えた 0.5b ~ 7b) が向上しますが、進化するガイドラインには依然として敏感です。これらの進歩にもかかわらず、CTD の性能は依然として汎用毒性検出で達成される F1 $\約 0.9$ をはるかに下回っており、課題と継続的な研究努力の必要性が浮き彫りになっています。

原文 (English)

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixed safety guidelines to all users. Our empirical evidence shows that these detectors fail to shield marginalized communities: approximately 35% of generated images labeled safe are considered harmful by disability communities. In this position paper, we argue for community-specific toxicity detection (CTD). To demonstrate its feasibility, we collaborate with disability experts to develop safety guidelines for two communities: dwarfism and blind/low vision. Using a dataset of 2,400 annotated T2I-generated images we demonstrate that both large vision-language models and existing general-purpose toxicity detectors catastrophically fail to recognize harmful content under these guidelines in zero-shot settings with F1 score lower than random guessing (F1 0.32 and 0.37). Promisingly, prompt-based adaptation methods (ICL, VQA) substantially improve harm detection performance (GPT-4o: F1 0.50 and 0.78), while parameter-efficient fine-tuning improves smaller models (0.5b-7b with best F1 0.48 and 0.59) with less than 100 demonstrations, but remains sensitive to evolving guidelines. Despite these gains, CTD performance remains far below F1 $\approx 0.9$ achieved for general-purpose toxicity detection, highlighting the challenge and the need for sustained research effort.

13:00 JST研究/論文

事後単体幾何学によるラベルなしのマルチクラス分類

多くの分類問題では、信頼できるインスタンス レベルのラベルが利用できません。ただし、多くの場合、弱く濃縮されたラベルのないサンプル、つまり、潜在クラスの割合を明らかにすることなく変化させる、さまざまなカット、ソース、母集団、または実験条件によって選択されたデータセットを構築することが可能です。ラベルなしの分類 (CWoLa) は、バイナリの場合 ($K=2$) で、異なるクラス比率を持つ 2 つの不純な混合物を区別するように訓練された分類器は、混合比率がわからなくても最適なクラス識別器を回復できることを示しています。この原則をいくつかのラベルなし混合物 ($K>2$) からのマルチクラス学習に拡張します。この場合、学習者は混合物の同一性のみを観察し、潜在クラス ラベルもクラス事前行列も観察しません。多クラス混合モデルの場合、ベイズ最適混合分類器 $g^\star$ がデータ点を混合事後空間に埋め込まれた $(K-1)$-simplex にマッピングすることを証明します。このシンプレックスの $K$ 頂点は、未知の混合行列を通じて潜在クラスによって誘導されます。このジオメトリを活用して、標準分類器をトレーニングして混合物のアイデンティティを区別し、事後シンプレックスフィッティングまたはボトルネックアーキテクチャを使用して潜在クラス構造を抽出する事前不要の手順を提案します。 MNIST、CIFAR-10、および Galaxy10 DECaLS での実験では、混合物の同一性だけで混合物中の潜在クラスとその分数を回復できることが示されています。弱く監視されたパフォーマンスと完全に監視されたパフォーマンスの間のギャップを狭めることにより、ラベルが不足しているドメインでのマルチクラス検出のための数学的に根拠のあるスケーラブルなツールを提供します。

原文 (English)

Multiclass Classification without Labels via Posterior Simplex Geometry

In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched unlabeled samples: datasets selected by different cuts, sources, populations, or experimental conditions that change latent class proportions without revealing them. Classification without Labels (CWoLa) shows that, in the binary case ($K=2$), a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions. We extend this principle to multiclass learning from several unlabeled mixtures ($K>2$), where the learner observes only mixture identity and neither latent class labels nor class-prior matrices. We prove that, for a multiclass mixture model, the Bayes-optimal mixture classifier $g^\star$ maps data points into a $(K-1)$-simplex embedded in mixture-posterior space. The $K$ vertices of this simplex are induced by the latent classes through the unknown mixing matrix. Leveraging this geometry, we propose prior-free procedures that train a standard classifier to distinguish mixture identities and then extract latent class structure using either post-hoc simplex fitting or a bottleneck architecture. Experiments on MNIST, CIFAR-10, and Galaxy10 DECaLS show that mixture identity alone can recover latent classes and their fractions in the mixture. By narrowing the gap between weakly supervised and fully supervised performance, we provide a mathematically grounded, scalable tool for multiclass discovery in label-scarce domains.

13:00 JSTLLM/生成AI

転置不変ブロック量子化による安定した FP4 トレーニング

トレーニング精度を下げることは、大規模言語モデル (LLM) トレーニングの効率を向上させるための重要な手段ですが、最適化中の不安定性のため、FP8 を超えて 4 ビット浮動小数点 (FP4) に到達することは依然として困難です。私たちは、既存のマイクロスケーリング手法におけるこの不安定性の根本的な原因、つまりテンソル転置によって引き起こされるスケールの不一致を特定しました。従来の 1D ブロック量子化では、前方パスと後方パスで転置後の同じ値に異なるスケーリング係数が割り当てられるため、偏った不安定な勾配更新が発生します。この問題に対処するために、転置不変スケーリングを強制し、前方計算と後方計算の間の一貫性を維持する、2D ブロック FP4 量子化に基づく低精度トレーニング フレームワークを提案します。さらに、これを切り捨てのないスケーリングおよび確率的丸めと組み合わせて、量子化誤差を制御し、不偏の勾配を維持します。注意メカニズムの感度を処理するために、クエリとキーの投影に MXFP8 量子化を採用し、実用的な混合精度設計を実現します。私たちは、最大 70 億のパラメーターと、最大 100 億のトークンでトレーニングされた 300 億の専門家混合モデルの高密度 LLM でメソッドを評価しました。すべての設定において、当社のアプローチは安定したエンドツーエンドの FP4 トレーニングを実現し、複雑さとダウンストリーム精度の低下が 1.3% 未満で、BF16 のパフォーマンスにほぼ匹敵します。これらの結果は、前方後方スケーリングの一貫性を強制することは、大規模な実践的な FP4 トレーニングを可能にするのに十分であり、より効率的な LLM トレーニングへのシンプルかつ効果的な経路を提供することを示しています。

原文 (English)

Stable FP4 Training via Transposition-Invariant Block Quantization

Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization. We identify a fundamental source of this instability in existing microscaling approaches: scale inconsistency induced by tensor transposition. In conventional 1D block quantization, forward and backward passes assign di erent scaling factors to the same values after transposition, leading to biased and unstable gradient updates. To address this issue, we propose a low-precision training framework based on 2D block FP4 quantization, which enforces transposition-invariant scaling and preserves consistency between forward and backward computations. We further combine this with truncation-free scaling and stochastic rounding to control quantization error and maintain unbiased gradients. To handle the sensitivity of attention mechanisms, we adopt MXFP8 quantization for query and key projections, yielding a practical mixed-precision design. We evaluate our method on dense LLMs up to 7B parameters and a 30B Mixture-of-Experts model, trained on up to 100B tokens. Across all settings, our approach achieves stable end-to-end FP4 training and closely matches BF16 performance, with less than 1.3% degradation in perplexity and downstream accuracy. These results demonstrate that enforcing forwardbackward scaling consistency is su cient to enable practical FP4 training at scale, providing a simple and e ective pathway toward more e cient LLM training.

13:00 JST研究/論文

生成分布的に堅牢な最適化

生成モデルは分布ロバスト最適化 (DRO) で採用されることが増えていますが、既存のアプローチではモデルの互換性と敵対的構造がトレードオフになっています。任意のサンプラーを受け入れるメソッドは最悪の場合の法則をジェネレーター ファミリに制限しませんが、ジェネレーターでパラメーター化された敵対者は尤度、スコア、トレーニング データなどのモデル固有のアクセスに依存しています。我々は、生成分布ロバスト最適化 (GDRO) を提案します。これは、サンプリング可能な任意の条件付きジェネレーターを名目モデルとして受け入れ、ワーストケースの法則を選択した条件付きジェネレーター ファミリに制限する原則に基づいたフレームワークです。鍵となるのはサンプラーとシンクホーンの組み合わせです。サンプラーは条件法則を正確に表しますが、シンクホーン発散は尤度アクセスなしで誘導分布を比較し、サンプルのみから推定できます。結果として生じる母集団問題では、アクティブな意思決定コンテキストで直接有限サンプル近似と微分可能な主双対実装が可能になります。リプシッツ損失の場合、母集団シンクホーン半径は下流の劣化を制限します。明示的および暗黙的なジェネレーター全体で、私たちの方法は、名目上の決定と比較して、レアコンテキストのインベントリの後悔を 60% 削減し、SocialGAN ナビゲーションの衝突を 50% 削減します。

原文 (English)

Generative Distributionally Robust Optimization

Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibility and adversarial structure: methods that accept arbitrary samplers do not restrict worst-case laws to a generator family, while generator-parameterized adversaries rely on model-specific access such as likelihoods, scores, or training data. We propose Generative Distributionally Robust Optimization (GDRO), a principled framework that accepts any sampleable conditional generator as the nominal model and restricts worst-case laws to a chosen conditional generator family. The key is the sampler-Sinkhorn pairing: samplers represent the conditional laws exactly, while Sinkhorn divergence compares their induced distributions without likelihood access and can be estimated from samples alone. The resulting population problem admits a direct finite-sample approximation and differentiable primal-dual implementation at the active decision context. For Lipschitz losses, the population Sinkhorn radius bounds downstream degradation. Across explicit and implicit generators, our method reduces rare-context inventory regret by 60% and SocialGAN navigation collisions by 50% relative to nominal decisions.

13:00 JST研究/論文

地震カタログの自動ナレッジ グラフ構築とクエリ

近年、より効果的な深層学習ベースの検出器と位相ピッカーの利用により、地震カタログに含まれるイベントの数は大幅に増加していますが、このシーケンスの特徴は何か?などの自由回答型の質問に答えることができています。厳格な時空間ウィンドウ処理と主観的な専門家の解釈によって制約されたままです。我々は、グラフベースの検索拡張生成 GraphRAG を、3 つの独立したカタログ、貯留層に隣接する群れ、2019 年のリッジクレスト構造シーケンス、および 2021 年のマドゥオ Mw7.4 余震シーケンスにわたる、生の表形式のカタログ レコードに直接直接適用する最初の体系的なアプリケーションを紹介します。手動でデータを構造化する必要がなく、パイプラインは 3 つすべてについて構造的に完全でクエリ可能なナレッジ グラフを構築します。カタログ由来のグランド トゥルースとルール ベースの参照グラフに照らして個別に検証された厳密な評価により、故障モードが明らかになり、4 つの地震学に基づいた即時修正により、標的とする捏造がすべて排除され、同時にメカニズムの推論が大幅に改善されます。ベクトル RAG ベースラインは、グラフ レイヤーの固有値、カタログ全体の要約、および一時的な段階の比較を示します。さらに、注意が必要な 2 つの主な落とし穴を特定しました。したがって、GraphRAG は、地震カタログに対して実用的で転送可能でコストがほぼゼロのクエリ インターフェイスを提供します。慎重なプロンプトにより、結果が一貫して正確で信頼できるものになります。

原文 (English)

Automatic Knowledge Graph Construction and Query for Earthquake Catalogs

In recent years, the number of events in earthquake catalogs has significantly increased due to the utilization of more effective deep learning based detectors and phase pickers but answering open ended questions such as what characterizes this sequence? remains constrained by rigid spatiotemporal windowing and subjective expert interpretation. We present the first systematic application of graph based retrieval augmented generation GraphRAG directly to raw, tabular catalog records across three independently featured catalogs, a reservoir adjacent swarm, the 2019 Ridgecrest tectonic sequence, and the 2021 Maduo Mw7.4 aftershock sequence. Without the need for manual data structuring, the pipeline builds structurally complete, queryable knowledge graphs for all three. Rigorous evaluation individually verified against catalog derived ground truth and a rule based reference graph exposes failure modes, and four seismology informed prompt fixes eliminate all targeted fabrications while sharply improving mechanism reasoning. A vector RAG baseline demonstrates the graph layers distinctive value, catalog wide summarization and temporal stage comparison. In addition, we have identified two main pitfalls that need attention. GraphRAG thus offers a practical, transferable, near zero cost query interface for earthquake catalogs, where careful prompting ensures the results are consistently accurate and trustworthy.

13:00 JSTLLM/生成AI

体系的な文献レビューをサポートする GenAI ツールの使用および評価に関する予備ガイドライン

背景: 生成 AI (GenAI) と大規模言語モデル (LLM) は、体系的文献レビュー (SLR) を含む、ソフトウェア エンジニアリングやその他の分野の学術タスクにますます使用されています。ただし、テキストを要約することはできますが、一眼レフカメラに求められる厳密性、信頼性、透明性を満たせるという保証はありません。目的: GenAI を使用して SLR を実施しようとしている研究者、または GenAI が SLR タスクをどの程度サポートしているかを評価する実証研究を実施している研究者をサポートすること。方法: まず、迅速なレビューを実施して、一眼レフカメラをサポートするために GenAI と LLM を評価および使用するためのガイドラインを提案する研究を特定しました。次に、思考実験、文献からの関連ガイダンス、および SLR を実施してツールを評価した私たち自身の経験を活用して、SLR のコンテキストで GenAI を使用および評価する方法に関する推奨事項を作成しました。結果: 研究者が SLR 用の GenAI を評価する際に直面する問題について説明します。 GenAI を使用した SLR と GenAI ツールの評価の両方を計画、実施、報告する際に考慮すべきプロセスの問題を特定し、説明します。最後に、結果を一連のプロセス推奨事項として要約し、これを GUEST (GenAI Use and Evaluation in SLR Tasks) と名付けます。結論: GenAI には人間の監督が必要であり、現時点では監督なしで系統的な研究を行うことはできないと我々は主張します。ただし、一部の反復的なタスクや一部の複雑なタスクの追加検証については、費用対効果の高い支援が得られる可能性があります。私たちの GUEST 推奨事項は、ソフトウェア エンジニアリングの研究者が GenAI を使用して信頼できる SLR を実施および報告し、厳密な独立した評価研究を提供するのに役立ちます。

原文 (English)

Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require. Objectives: To support researchers intending to conduct SLRs using GenAI or those conducting empirical studies evaluating how well GenAI supports SLR tasks. Methods: First, we conducted a rapid review to identify studies that propose guidelines for evaluating and using GenAI and LLMs to support SLRs. Second, we drew on thought experiments, relevant guidance from the literature, and our own experience conducting SLRs and evaluating tools to develop recommendations for how to use and assess GenAI in the context of SLRs. Results: We discuss the problems researchers face when evaluating GenAI for SLRs. We identify and explain process issues to consider when planning, conducting, and reporting both SLRs using GenAI and evaluations of GenAI tools. Finally, we summarize our results as a set of process recommendations, which we name GUEST (GenAI Use and Evaluation in SLR Tasks). Conclusion: We argue that GenAI requires human oversight and is not currently capable of unsupervised systematic studies. However, it offers the prospect of cost-effective assistance for some repetitive tasks and for additional validation of some complex tasks. Our GUEST recommendations should help software engineering researchers both to conduct and report trustworthy SLRs using GenAI and to provide rigorous independent evaluation studies.

13:00 JSTロボティクス

調整された部分リセット: 継続的な強化学習におけるポリシーの崩壊を防ぐ

ニューラル ネットワークは、特に継続的な教師あり学習や強化学習などの非定常データ設定において、トレーニング中の休止ニューロンの蓄積と表現力の喪失によって妨げられます。最近、ニューロンのリセットは、勾配流を維持し、可塑性を回復するために使用されています。ただし、ユニット全体を再初期化すると、多くの場合、ピークパフォーマンスが犠牲になり、トレーニングが不安定になり、ポリシーの崩壊につながる可能性があります。トレーニングを不安定にすることなく可塑性を維持するために、各ニューロンのユーティリティに応じてプル強度を調整して、ユーティリティの低いニューロンを初期化に向けて定期的にプルするオプティマイザーである Calibrated Partial Resets (CPR) を提案します。バイナリ リセット方法とは異なり、部分リセットでは脆弱性が回避されます。均一な減衰とは異なり、調整されたユーティリティ スケーリングでは、最も必要なユニットに集中的に調整が行われます。比較した手法の中で、CPR のみが SlipperyAnt の 4 億トレーニング ステップにわたるポリシー崩壊を回避し、Continual MetaWorld および Continual MinAtar ベンチマークで以前の減衰ベースの手法およびリセットベースの手法よりも優れたパフォーマンスを示しました。アブレーションにより、可塑性とピークパフォーマンスの間の調整可能なトレードオフが明らかになり、継続的な学習の有望な方向性としてユーティリティ規模の再初期化が強調されます。

原文 (English)

Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning

Neural networks are hindered by accumulating dormant neurons and loss of expressivity throughout training, particularly in non-stationary data settings, such as continual supervised and reinforcement learning. Recently, neuron resets have been used to maintain gradient flow and restore plasticity. However, full unit reinitialization often sacrifices peak performance and can destabilize training, leading to policy collapse. To preserve plasticity without destabilizing training, we propose Calibrated Partial Resets (CPR), an optimizer that periodically pulls low-utility neurons toward their initialization, with pull strength scaled by each neuron's utility. Unlike binary reset methods, partial resets avoid brittleness; unlike uniform decay, calibrated utility-scaling concentrates adjustment on the units that need it most. Among compared methods, only CPR avoids policy collapse over 400M training steps in SlipperyAnt, and it outperforms prior decay and reset-based methods on Continual MetaWorld and Continual MinAtar benchmarks. Ablations reveal a tunable trade-off between plasticity and peak performance, highlighting utility-scaled reinitialization as a promising direction for continual learning.

13:00 JSTLLM/生成AIビジネス/資金調達

CogArena: 大規模言語モデルにおける認知能力構造のマルチメソッド評価

LLM 認知スコアは、能力ごとのプロファイルとして要約されることが多くなり、その次元はタスク全体で収束し、一致する介入に選択的に反応し、定義に使用されるモデルを超えて一般化される必要があります。 CogArena を紹介します。CogArena は、認知タスクのスコアが 5 つの理論に基づいたグループ分けの次元ラベルを正当化する時期を決定するためのマルチメソッド フレームワークを中心に構築された、手続き的に生成された 13 パラダイム ベンチマークです。 55 のオープンウェイト モデル全体で、ほぼすべてのパラダイム相関が正であり、共通の軸によって分散の約半分が説明されます。グループ内での利点は小さく、スコアに左右されやすく、モデル ファミリ全体で不確実です。 6 つのファミリーからの 12 のモデルにわたる個別に凍結された完全交差研究では、ターゲットを絞ったスキャフォールドは一致グループ化の小さな利点を示しますが、スキャフォールド固有のコントラストは多重性補正に耐えられず、選択性はホールドアウトファミリーの予測を改善しません。凍結された確認基準は失敗します。ポストホックの代替文言の複製では、より小さな正の推定値が生成され、やはり失敗します。これらの結果を総合すると、境界の結論が裏付けられます。理論に沿ったプロンプトは、バッテリー内で小さな斜めの傾向を生成しますが、現在の証拠は安定した 5 次元プロファイルを確立していません。 CogArena は、認知ラベルをモデル スコアに付ける前に、行動シグネチャ、共分散、一致する介入、家族外の予測を結合するワークフローを提供します。

原文 (English)

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.

13:00 JSTエージェントAnthropicClaude

オーサリング エージェント スキル: ソフトウェア エンジニアリングのアプローチ

エージェント スキルは、エージェントがオンデマンドでロードする再利用可能な手順知識を使用して大規模言語モデル エージェントを拡張する新しい方法です。 Anthropic はエージェント スキルを導入し、その形式を複数のエージェント ツールでサポートされるオープン仕様として公開しました。このノートでは、スキルはソフトウェア成果物であり、その構築はソフトウェア エンジニアリングの原則に従うべきであり、その条件として、単一の責任、インターフェイスと実装の分離、低結合、共有トークン バジェットでの経済性、および決定論的テストの代わりの動作評価に従う必要があると主張しています。クロード コードを参照実装として使用し、スキルがどのように構造化されるか、その内容が段階的に読み込まれる方法、および選択に依存する記述の書き方について説明します。これは、開発者がエージェントの動作を形成するために使用できる他のメカニズム (プロジェクト メモリ ファイル、スラッシュ コマンド、サブエージェント、外部ツール接続、フックなど) に対してスキルを配置し、メカニズムの実行を誰が決定するか、メカニズムが提供する保証に基づいて、これらを選択するためのルールを与えます。次に、評価主導のオーサリング プロセス、オーサリングで一般的に遭遇する一連のパターンと障害、およびサードパーティのスキルを使用することで生じる信頼性の問題について説明します。 UML クラス スタイルで描画された比較、読み込みモデル、スキルの構造、各メカニズムの相対的な位置、セッション中にスキルとフックが動作するポイントを示します。

原文 (English)

Authoring Agent Skills: A Software-Engineering Approach

Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. Anthropic introduced Agent Skills and published the format as an open specification supported across several agent tools. This note argues that a skill is a software artefact and that its construction should follow software-engineering principles, with qualifications: single responsibility, separation of interface from implementation, low coupling, and economy in a shared token budget, together with behavioural evaluation in place of deterministic testing. Using Claude Code as the reference implementation, it describes how a skill is structured, how its contents are loaded in stages, and how to write the description on which selection depends. It places skills against the other mechanisms a developer can use to shape agent behaviour, like project memory files, slash commands, subagents, external tool connections, and hooks, and gives a rule for choosing between them based on who decides that a mechanism runs and what guarantee it provides. It then sets out an evaluation-driven authoring process, a set of patterns and faults commonly encountered in authoring, and the trust question raised by using skills from third parties. We illustrate the comparison drawn in UML class style, the loading model, the anatomy of a skill, the relative position of each mechanism, and the points at which skills and hooks act during a session.

13:00 JSTLLM/生成AI

コンセンサスに基づいた、新興科学と歩調を合わせた: 長期にわたる新型コロナウイルス対応のためのコンセンサスに基づくマルチコーパス臨床チャットボット

長期にわたる COVID (LC) は、臨床意思決定のサポートに課題をもたらします。これは、関連する証拠が、更新サイクル、証拠の役割、臨床成熟度のレベルが異なるソースに分散されているためです。我々は、専門家が厳選したコンセンサスガイダンス、最新のPubMed文献、登録された介入試験、生きた系統的レビューからの証拠という4つの情報源を検索強化ワークフロー内で整理する臨床医向けチャットボットを紹介する。コンセンサス ガイダンスはフレーム応答に常に含まれており、残りのソースはユーザーが選択したときに並行して取得されます。臨床医が直面する 50 の質問に対する探索的自動評価において、当社のチャットボットは OpenEvidence と同等の平均評価を示し、LLM 判定による比較では数値的により高いスコアとより低いスコアのばらつきを示しました。

原文 (English)

Grounded in Consensus, In Step With Emerging Science: A Consensus-Anchored Multi-Corpus Clinical Chatbot for Long COVID

Long COVID (LC) poses a challenge for clinical decision support because relevant evidence is distributed across sources with different update cycles, evidentiary roles, and levels of clinical maturity. We present a clinician-facing chatbot that organizes four sources within a retrieval-augmented workflow: expert-curated consensus guidance, current PubMed literature, registered interventional trials, and evidence from living systematic reviews. Consensus guidance is always included to frame responses, while the remaining sources are retrieved in parallel when selected by the user. In an exploratory automated evaluation on 50 clinician-facing questions, our chatbot showed comparable mean ratings to OpenEvidence, with numerically higher scores and lower score variability in an LLM-judged comparison.

13:00 JSTロボティクス

人間とロボットのチーミングにおける人間の位置制御のための仲介層としての拡張現実

拡張現実 (XR) は、ロボットの意図、計画された動作、到達可能性、および状態を伝達するために、人間とロボットのインタラクションでますます使用されています。私たちは、XR は人間とロボットのチーミングにおける人間の位置に応じた制御のための仲介層としても理解されるべきであると主張します。状況に応じた人間の制御とは、アクションが展開される具体的な物理的、社会的、時間的コンテキスト内でロボットのアクションを理解し、形成し、許可し、中断する人間の協力者の能力を示します。私たちはこの視点を、ロボット支援によるベッドサイド看護、マルチアーム監視制御、分割注意の下での共同組み立てなどのシナリオに基づいています。これらのシナリオ全体にわたって、人の移動、目標の変更、センシングの不完全さ、制御の役割の変化、計画の無効化などに応じて、ロボットの自律性は検査可能であり、調整可能である必要があります。人間の意図とロボットの自律性、ロボットの計画と人間の判断、共有制御のレベル、チームの役割、引継ぎ、回復を結び付ける 4 つの仲介機能を特定します。これらの機能に基づいて、共同行動の可能性、社会物理的制約、不確実性と計画の妥当性、マルチモーダルな制御と修正、役割、引継ぎと説明責任、予期的回復という 6 つの設計次元を導き出します。この論文では、動的な共有環境においてロボットの自律性をより実践的で責任あるものにする XR システムの研究課題について概説しています。

原文 (English)

Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming

Extended Reality (XR) is increasingly used in human-robot interaction to communicate robot intent, planned motion, reachability, and state. We argue that XR should also be understood as a mediation layer for situated human control in human-robot teaming. Situated human control denotes the human collaborator's ability to understand, shape, authorize, and interrupt robot action within the concrete physical, social, and temporal context in which that action unfolds. We ground this perspective in scenarios from robot-assisted bedside nursing, multi-arm supervisory control, and collaborative assembly under divided attention. Across these scenarios, robot autonomy must remain inspectable and adjustable as people move, goals change, sensing is incomplete, control roles shift, and plans become invalid. We identify four mediation functions connecting human intent and robot autonomy, robot plans and human judgment, levels of shared control, and team roles, handover, and recovery. Building on these functions, we derive six design dimensions: joint action possibilities, socio-physical constraints, uncertainty and plan validity, multimodal control and correction, roles, handover, and accountability, and anticipatory recovery. The paper outlines a research agenda for XR systems that make robot autonomy more actionable and accountable in dynamic shared environments.

13:00 JST研究/論文

Lantern: 熱量計シミュレーションにおける物理学に基づく拡散モデルのための競合を意識した勾配ブレンディング

熱量計シャワーのモンテカルロ シミュレーションは高輝度 LHC の主なボトルネックであり、拡散モデルが高速で忠実度の高い代用モデルとして登場しました。ただし、ノイズ除去の目的は純粋に統計的なものです。モデルは、物理学を誤って配置しながらノイズを最小限に抑えることができます。既存の物理学に基づいた生成手法は、このギャップを埋めることができません。なぜなら、シャワーでは提供されない、閉じた形式の法則、支配的な PDE 残差、またはサンプルごとの厳しい制約を前提としているからです。確率的カスケードを支配するサンプルごとの偏微分方程式はなく、エネルギー保存によってシャワーごとに 1 つのスカラーのみが固定されます。標準メトリクスは、熱量計レイヤーとボクセルにわたる相関構造を無視し、物理特徴空間でのみシャワーを比較します。私たちは両方のギャップに対処します。相関フロベニウス距離 (CFD) を導入します。これは、レイヤーごとおよびボクセルごとのスケールでの相関忠実度の単一の正規化スコアです。次に、シャワーで利用可能なサンプルごとのソフト構造を 2 つの物理認識補助損失としてエンコードします。計数統計に基づいた分散安定化ボクセル残留損失と、検出器ジオメトリ上のグラフ ラプラシアン損失です。 GradBlend を介して両方をノイズ除去と組み合わせます。これにより、ステップの大きさをノイズ除去勾配に固定しながら、補助に方向を制御させ、物理誘導拡散サロゲートである Lantern を生成します。 CaloChallenge データセット 2 では、PCGrad、GradNorm、IMTL-G、Config などのタスク対称ルールを通じて物理損失を注​​入すると、ノイズ除去だけの場合と比較して FPD が 2 ~ 100 倍増加します。一方、GradBlend は回帰なしで同じ信号を受け入れ、ラプラシアン損失により、Lantern は FPD と CFD の両方を改善します。補助損失スケジューラのアブレーションは、その勾配がノイズ除去と競合するボクセル残留損失は、シャワーの忠実度を維持するために最終的なノイズ除去のみのフェーズを必要とするのに対し、競合しないラプラシアン損失はスケジュールの影響を受けないことを示しています。

原文 (English)

Lantern: Conflict-Aware Gradient Blending for Physics-Guided Diffusion Models in Calorimeter Simulation

Monte Carlo simulation of calorimeter showers is a principal bottleneck for the High-Luminosity LHC, and diffusion models have emerged as fast, high-fidelity surrogates. Their denoising objective is purely statistical, however: a model can minimize it while placing the physics wrong. Existing physics-informed generative methods cannot close this gap, because they assume a closed-form law, a governing PDE residual or a hard per-sample constraint, that a shower does not supply: no per-sample PDE governs a stochastic cascade, and energy conservation fixes only one scalar per shower. Standard metrics ignore the correlation structure across calorimeter layers and voxels, comparing showers only in a physics feature space. We address both gaps. We introduce the Correlation Frobenius Distance (CFD), a single normalized score for correlation fidelity at layer-wise and voxel-wise scales. We then encode the soft per-sample structure available in a shower as two physics-aware auxiliary losses: a variance-stabilized voxel residual loss grounded in counting statistics, and a graph Laplacian loss over the detector geometry. We combine both with denoising through GradBlend, which anchors the step magnitude to the denoising gradient while letting the auxiliary steer its direction, yielding Lantern, a physics-guided diffusion surrogate. On CaloChallenge Dataset 2, injecting the physics losses through task-symmetric rules such as PCGrad, GradNorm, IMTL-G, and ConFIG inflates FPD by 2-100x relative to denoising alone, whereas GradBlend admits the same signal without regression and, with the Laplacian loss, Lantern improves both FPD and CFD. Our ablation on the auxiliary loss scheduler shows that the voxel residual loss, whose gradient conflicts with denoising, requires a terminal denoising-only phase to preserve shower fidelity, whereas the non-conflicting Laplacian loss is insensitive to the schedule.

13:00 JSTLLM/生成AI

DS@GT ARC (CheckThat!) 2026: 多言語数値請求検証のための LLM ベースのトレース ランキングとグループ化された報酬モデリング

数値クレームの自動検証は、言語理解と定量的推論の両方が必要なため、困難な問題です。この文書では、CLEF 2026 CheckThat! のシステムについて説明します。タスク 2 は、大規模言語モデル (LLM) によって生成された推論トレースのランク付けと、英語とアラビア語での数値クレームの最終的な評決の予測に焦点を当てています。 2 つのアプローチを検討します。最初のアプローチでは、LoRA を使用して LLM ベースの検証器を微調整し、バイナリ分類問題として各推論トレースを独立してスコアリングし、Best-of-N 選択を使用して最終的な判定を選択します。さらに、適応的なサブクレーム分解を試して、検証前に複雑なクレームをより単純な部分に分割します。 2 番目のアプローチでは、手作りの数値および時間的重複特徴を備えた軽量の TF-IDF 報酬モデルを使用してトレースをスコア化し、判定グループごとにスコアを集計して最終的な予測を決定します。アラビア語については、一般的な多言語モデルと、アラビア語のテキストで事前トレーニングされた言語固有のモデルである AraBERT を比較します。私たちの結果は、LLM ベースのアプローチがほとんどのメトリクス、特に Recall@5 で軽量報酬モデルよりも優れている一方、報酬ベースのアプローチが競合クラスではより優れたパフォーマンスを示していることを示しています。サブクレームの分解ではパフォーマンスは向上しませんでした。これは、クレームの分割が推論を支援するのではなく、ノイズを引き起こすことを示唆しています。アラビア語の場合、AraBERT はほとんどの指標において多言語ベースラインを上回ります。

原文 (English)

DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification

Automated verification of numerical claims is a challenging problem, as it requires both language understanding and quantitative reasoning. This paper describes our system for CLEF 2026 CheckThat! Task 2, which focuses on ranking reasoning traces generated by large language models (LLMs) and predicting a final verdict for numerical claims in English and Arabic. We explore two approaches. The first approach fine-tunes an LLM-based verifier using LoRA to score each reasoning trace independently as a binary classification problem, and selects the final verdict using Best-of-N selection. We further experiment with adaptive sub-claim decomposition to break complex claims into simpler parts before verification. The second approach uses a lightweight TF-IDF reward model with handcrafted numeric and temporal overlap features to score traces, and aggregates scores by verdict group to determine the final prediction. For Arabic, we compare a general multilingual model against AraBERT, a language-specific model pretrained on Arabic text. Our results show that the LLM-based approach outperforms the lightweight reward model on most metrics, particularly Recall@5, while the reward-based approach shows stronger performance on the Conflicting class. Sub-claim decomposition did not improve performance, suggesting that claim splitting introduces noise rather than aiding reasoning. For Arabic, AraBERT outperforms the multilingual baseline across most metrics.

13:00 JST研究/論文

合成制御におけるスペクトルのトランケーション

合成対照 (SC) は、治療を受けたユニットの治療前の軌跡をドナー ユニットの重み付けされた組み合わせに一致させます。我々は、代わりに、ドナーパネルの主要な時間特異ベクトルによって定義される座標で処理されたユニットと一致するスペクトルSCと、保持および破棄された方向に個別に調整可能な重みを置き、生のパスSCと切り捨てられたスペクトルSCをエンドポイントとしてネストするハイブリッド推定器を研究します。我々は、この族がフルランクで正確にローパスSCに減少すること、次元$N_0-K-1$のアフィン解セットを使用して、$N_0>K+1$のときはいつでも$N_0$ドナーとの$K$保持次元の正確なバランスが過少決定されること、およびスペクトルの不均衡が有限サンプルの最良線形予測子分解を通じて治療効果のバイアスにマッピングされることを証明します。我々は、レジーム当たり 400 ドルの反復とドナーのみのプラセボ検証を使用して、正則化と混合加重を選択する 11 のデータ生成レジームにわたる推定量を評価します。切断されたスペクトル SC は、すべての領域において調整された生パス SC よりも大幅に高い RMSE を持ち、ペアの差はモンテカルロ標準誤差 $4$ から $11$ に相当します。ハイブリッド推定器は、ほとんどのレプリケーションで raw パス マッチングを選択し、ほとんどのレジームで調整された SC と統計的に区別できません。結果は前処理の影響を非常に受けやすくなります。生の入力では、パフォーマンスのギャップが大きくなります。限界の背後にある仮定が示唆するように、スペクトル分解前の単位と時間の固定効果を除去した後、ギャップはほぼ消滅し、プラセボ検証は切り捨てを支持し始めます。私たちはこれらの発見を、スペクトル SC がローパス SC を置き換えるべきであるという証拠としてではなく、診断として解釈します。スペクトル マッチングがいつ役立つかを決定するのは、基底推定ノイズ、バランス不足決定、および固定効果の汚染です。

原文 (English)

Spectral Truncation in Synthetic Control

Synthetic control (SC) matches a treated unit's pre-treatment trajectory to a weighted combination of donor units. We study Spectral SC, which instead matches the treated unit in coordinates defined by the leading temporal singular vectors of the donor panel, and a hybrid estimator that places separately tunable weight on retained and discarded directions, nesting raw-path SC and truncated Spectral SC as endpoints. We prove that the family reduces exactly to raw-path SC at full rank, that exact balance on $K$ retained dimensions with $N_0$ donors is underdetermined whenever $N_0>K+1$, with an affine solution set of dimension $N_0-K-1$, and that spectral imbalance maps to treatment-effect bias through a finite-sample best-linear-predictor decomposition. We evaluate the estimators across eleven data-generating regimes, using $400$ replications per regime and donor-only placebo validation to select regularization and the mixing weight. Truncated Spectral SC has significantly higher RMSE than tuned raw-path SC in every regime, with paired differences equal to $4$ to $11$ Monte Carlo standard errors. The hybrid estimator selects raw-path matching in most replications and is statistically indistinguishable from tuned SC in most regimes. The result is highly sensitive to preprocessing. With raw inputs, the performance gap is large; after removing unit and time fixed effects before spectral decomposition, as suggested by the assumptions behind our bound, the gap nearly disappears and placebo validation begins to favor truncation. We interpret these findings diagnostically rather than as evidence that Spectral SC should replace raw-path SC. Basis-estimation noise, balancing underdetermination, and fixed-effects contamination determine when spectral matching can help.

13:00 JSTLLM/生成AI

含意の認識とキャンセルによる大規模言語モデルにおけるコミュニケーション上の信念の更新の評価

人間の言語は暗黙の信念と信念の更新によって動かされるため、これらは大規模言語モデル (LLM) とそのユーザー間のコミュニケーションを成功させるモデルにとって重要です。この論文では、含意を通じて作られた暗黙の信念を認識し、含意のキャンセルを通じてその更新を理解するLLMの能力を評価します。これは、発話の暗黙の意味が弱められるか否定される実用的な現象です。私たちは、インプリカチャーとそれに対応するキャンセルを人間が判断するためにクラウドソーシングされた、初の専門家による注釈付きインプリカチャー キャンセル データセット ImplicatureX を作成しました。 LLM の信念更新の理解は、特により自然に発生するシナリオにおいて、人間のそれに比べて遅れていることがわかりました。追加の対照実験では、LLM 信念更新の成功の一部は以前の信念への依存に起因する可能性があり、信念更新の失敗はそのタイプと形式に依存する可能性があることを示唆しています。全体として、私たちの研究は、現在のLLMが暗黙の信念と信念の更新について人間レベルの理解にまだ達していないことを示唆しています。コードとデータは https://github.com/cesare-spinoso/ImplicatureX で入手できます。

原文 (English)

Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance's implied meaning is weakened or negated. We create the first expert-annotated implicature cancellation dataset, ImplicatureX, crowdsourced for human judgements of implicatures and their corresponding cancellations. We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at https://github.com/cesare-spinoso/ImplicatureX.

13:00 JST画像/動画生成

OPERA: オフラインのポリシーに基づくエキスパートのルーティングと普遍的な生物医学画像解析への適応

生物医学画像解析は多様なモダリティやタスクにまたがっていますが、現実世界への展開は、スキャナー、プロトコル、患者集団全体にわたる深刻な分布の変化によって妨げられています。そのため、高パフォーマンスのモデルにはドメイン固有の微調整を繰り返す必要がありますが、これはコストのかかるサイクルであり、ラベルが不足していたり​​、プライバシーの制約によってデータ共有が制限されている場合には現実的ではありません。我々は、エキスパートの重み付け割り当てをオフライン ポリシーの学習問題として扱うことで、この展開のボトルネックに対処するマルチエージェント アンサンブル フレームワークである OPERA (オフライン ポリシーに基づくエキスパート ルーティングと適応) を提案します。ルーティング ポリシーは、エキスパート エージェントへの勾配更新を行わずに小さな検証セットから学習され、その後、テスト時の適応を使用して配布シフトに対処するように展開されます。 OPERA は、補完的なメカニズムを通じて、異種の専門エージェントを調整します。エキスパート プロファイリング モジュールはオフラインで選択ポリシーを学習し、情報に基づいた専門知識の割り当てを可能にします。各エージェントは温度調整を通じて信頼性キャリブレーションを受け、より信頼性の高い確率的出力を保証します。 OPERA には、ラベルなしのテスト データから得られた統計を使用してクラスの重みがバッチ レベルで動的に調整される、分布を意識した適応も組み込まれています。インスタンス レベルのルーティングでは、モデル間の一致と予測エントロピーを活用して、各サンプルを最も適切な専門家に割り当てます。当社では、眼底写真、胸部 X 線、CT、MRI、マルチモーダル診断ベンチマークをカバーする 9 つのデータセットで OPERA を評価し、分類、セグメンテーション、マルチモーダル設定にわたる 30 以上のベースラインと比較しています。 OPERA はパフォーマンスとキャリブレーションの品質を一貫して向上させており、オフラインのポリシーに基づいた専門エージェントの調整が、再トレーニングなしで展開可能な生物医学 AI への実用的な道であることを示しています。コードは \href{https://github.com/HUANGLIZI/OPERA}{GitHub} にあります。

原文 (English)

OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distribution shifts across scanners, protocols, and patient populations. High-performing models consequently require repeated domain-specific fine-tuning, which is a costly cycle that becomes impractical when labels are scarce or privacy constraints limit data sharing. We propose OPERA (Offline Policy-guided Expert Routing and Adaptation), a multi-agent ensemble framework that addresses this deployment bottleneck by treating expert weight assignment as an offline policy learning problem: a routing policy is learned from a small validation set without gradient updates to any expert agent, then deployed with test-time adaptation to handle distribution shift. OPERA coordinates heterogeneous specialist agents through complementary mechanisms. The expert profiling module learns selection policies offline, enabling informed allocation of expertise. Each agent undergoes confidence calibration through temperature adjustment, ensuring more reliable probabilistic outputs. OPERA also incorporates distribution aware adaptation, where class weights are dynamically adjusted at the batch level using statistics derived from unlabeled test data. Instance level routing assigns each sample to the most suitable expert by leveraging inter model agreement and predictive entropy. We evaluate OPERA on 9 datasets covering fundus photography, chest X-ray, CT, MRI, and multimodal diagnostic benchmarks, comparing against 30+ baselines across classification, segmentation, and multimodal settings. OPERA consistently improves performance and calibration quality, demonstrating that offline policy-guided expert agents coordination is a practical path to deployable biomedical AI without retraining. Code is on \href{https://github.com/HUANGLIZI/OPERA}{GitHub}.

13:00 JST画像/動画生成

CNN ベースの ECG 画像分類におけるショートカット学習と賢いハンス効果の分析

ECG 画像分類用の深層学習モデルは、ECG 波形形態の代わりに非生理学的視覚的手がかりを利用することにより、高い精度を達成できる可能性があります。深層学習モデルのブラックボックスの性質を考慮すると、高い予測パフォーマンスという期待は、臨床または現実世界の信頼性、解釈可能性、実用的な意思決定に十分に反映されていないことがよくあります。この研究では、畳み込みニューラル ネットワークを使用して、公開されている ECG 画像データセットにおけるショートカット学習とクレバー ハンス効果を検証します。その過程で、6 つの画像由来の特徴セット (FS) を作成しました。FS1: 生のフル ECG 画像、FS2: トリミングされた波形のみの画像、FS3: 波形がマスクされたメタデータ画像、FS4: 心筋梗塞クラスの赤矢印アーチファクト画像、FS5: 異常心拍クラスのコントラスト強調画像、FS6: 正常クラスのガウスぼかし画像です。これらの制御された表現は、波形情報が削除された場合、または人為的なクラス固有のアーティファクトが導入された場合に、分類パフォーマンスが持続するかどうかをテストするために使用されました。学習パターンに関する透明性を評価するために、特徴セット表現間のショートカット保持スコア、予測の一貫性、信頼発散が計算されています。事実の結果とともに、モデルの帰属が ECG 関連の波形領域に焦点を当てているか、それとも非臨床アーチファクトに焦点を当てているかを検査するために、平均統合勾配および閉塞感度テストの結果が表示されます。潜在的なクレバー ハンスの動作を特定するために、機能セットと属性パターンにわたるパフォーマンスの変化が使用されました。この研究では、ECG 画像分類器が臨床的に意味のある形態や、レポート レイアウト、メタデータ、コントラスト、ぼかし、または人工マーカーによって導入されたショートカット キューを学習するかどうかを評価します。

原文 (English)

Analysis of the Shortcut Learning and Clever Hans Effect in CNN based ECG Image Classification

Deep learning models for ECG image classification may achieve high accuracy by exploiting non-physiological visual cues instead of ECG waveform morphology. Given the black-box nature of deep learning models, their promise of high predictive performance often remains insufficiently translated into clinical or real-world trust, interpretability, and actionable decision-making. In this study, we examine shortcut learning and Clever Hans effect in a publicly available ECG image dataset using convolutional neural networks. In process we have created six image-derived feature sets (FSs), FS1: raw full ECG images, FS2: cropped waveform-only images, FS3: waveform-masked metadata images, FS4: red-arrow artifact images for the myocardial infarction class, FS5: contrast-enhanced images for the abnormal heartbeat class and FS6: Gaussian-blurred images for the normal class. These controlled representations were used to test whether classification performance persists when waveform information is removed or when artificial class-specific artifacts are introduced. Shortcut retention score, prediction consistency and confidence divergence across Feature-Set Representations have been calculated to assess the transparency about the learning pattern. Along with factual results, average Integrated Gradients and occlusion sensitivity test results are presented to inspect whether model attribution focused on ECG-relevant waveform regions or on non-clinical artifacts. Performance changes across feature sets and attribution patterns were used to identify potential Clever Hans behavior. This study evaluates whether ECG image classifiers learn clinically meaningful morphology or shortcut cues introduced by report layout, metadata, contrast, blur, or artificial markers.

13:00 JSTLLM/生成AI

AI 生成コードの 53.6K 現実世界の開発者編集から学ぶ

AI によって生成されたコードに不完全性がある場合は、ソフトウェア開発者が生成されたコードを手動で変更するか、AI プログラミング アシスタントに再度プロンプトを表示する必要があります。手動コード編集では、最終的に成功したコード スニペットのみが含まれる Git コミットよりも、編集動作に関するより現実的で詳細な情報が提供されます。しかし、高品質で現実的なコード編集データが不足しているため、LLM は主に公開されている Git データ (コミットなど) でトレーニングされています。このギャップに対処するために、DECODE (Developer Edits of Code Dataset) を導入しました。これは、1,000 人以上の開発者から提供された、Python、TypeScript、および JavaScript で AI によって生成されたコードの 53.6K の実際の IDE 内コード編集のデータセットです。まず、データ分析における DECODE の有用性を実証し、AI が生成したコードがいつ、なぜ、どのように編集されるかについての洞察を取得します。ほとんどの編集は AI 補完を受け入れてから最初の 15 分以内に行われ、その結果、編集軌跡の 31% で AI 補完が削除されることがわかりました。次に、DECODE を使用して、コード編集を予測する LLM の能力をベンチマークします。 DECODE を微調整することで、オープンソース 3B モデルがフロンティア LLM よりも大幅に優れたコード編集予測タスクを実行できることがわかりました。次に、将来の AI プログラミング アシスタントのための開発者中心の機械学習アプローチの必要性を強調しながら、この研究の意味について議論します。

原文 (English)

Learning from 53.6K Real-World Developer Edits of AI-Generated Code

Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programming assistant. Manual code edits provide more realistic and granular information on editing behavior than Git commits, which only contain final successful code snippets. Yet, due to a lack of high-quality, realistic code editing data, LLMs are mostly trained on publicly available Git data (e.g., commits). To address this gap, we introduce DECODE (Developer Edits of Code Dataset), a dataset of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript, sourced from 1K+ developers. First, we demonstrate the utility of DECODE for data analysis, obtaining insights on when, why, and how AI-generated code is edited. We find that most edits occur within the first 15 minutes after accepting an AI completion, resulting in the removal of AI completions in 31% of edit trajectories. Second, we use DECODE to benchmark the ability of LLMs to predict code edits. We find that finetuning on DECODE enables open-source 3B models to perform code edit prediction tasks significantly better than frontier LLMs. We then discuss implications of this work, emphasizing the necessity of developer-centric machine learning approaches for future AI programming assistants.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

自律量子センシング実験における科学的推論のためのエージェント AI

ダイヤモンドの窒素空孔 (NV) 中心を用いた自律実験のために、大規模言語モデル (LLM) エージェントを中心に構築されたエージェント AI ワークフローを実装します。 NV センターは量子センシングに広く使用されているプラ​​ットフォームであり、コンピューターから多くの測定を制御できるため、NV 実験は自律的なワークフローにとって自然な環境になります。私たちは主に 2 つの貢献を行っています。まず、永続的なプロジェクト記録、定量的計算およびデータ分析ツール、決定論的な実験制御を組み合わせた自律的な NV 実験ワークフローを示します。ある自律実験では、エージェントは単一の NV 中心を選択し、その共振周波数を校正し、ラムゼイ測定で \(T_2^\ast\) を測定し、カー--パーセル--マイブーム--ギル (CPMG) 測定を追加して、近くの \(^{13}\mathrm{C}\) に関連する可能性のある弱い特徴をチェックしました。次に、エージェントの推論をラボでの実行とは別に評価する 2 つのオフライン ベンチマークを導入します。 GPT-5.4、GPT-5.5、および GPT-5.6 Sol を使用して両方のベンチマークを評価しました。 Ramsey チェックポイント ベンチマークでは、推論の努力が増えると、一般的に残留共鳴キャリブレーション オフセットの認識が向上します。対照的に、パルス光検出磁気共鳴 (pODMR) データ評価ベンチマークでは、パルス シーケンス情報だけで、より高い推論努力でより多くの偽陽性共鳴判定が生成されました。予想される信号の計算を必要とすることで、3 つのモデルおよび推論設定すべてで誤検知率が低く抑えられました。この結果は、自律実験における明確な分業を示唆しています。エージェントは科学的な仮説を立て、定量的なツールを使用してデータを評価します。その一方で、決定論的なコードがハードウェアを制御し、安全性の制約を強制します。

原文 (English)

Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments

We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond. NV centers are a widely used platform for quantum sensing, and the ability to control many measurements from a computer makes NV experiments a natural setting for autonomous workflows. We make two main contributions. First, we demonstrate an autonomous NV experiment workflow that combines persistent project records, quantitative calculation and data analysis tools, and deterministic experiment control. In one autonomous experiment, the agent selected a single NV center, calibrated its resonant frequency, measured \(T_2^\ast\) with Ramsey measurements, and added a Carr--Purcell--Meiboom--Gill (CPMG) measurement to check a weak feature that could be related to nearby \(^{13}\mathrm{C}\). Second, we introduce two offline benchmarks that evaluate the agent's reasoning separately from laboratory execution. We evaluated both benchmarks with GPT-5.4, GPT-5.5, and GPT-5.6 Sol. In the Ramsey checkpoint benchmark, greater reasoning effort generally improved recognition of a residual resonance calibration offset. By contrast, in the pulsed optically detected magnetic resonance (pODMR) data evaluation benchmark, pulse sequence information alone produced more false positive resonance judgments at higher reasoning effort. Requiring an expected signal calculation kept false positive rates low across all three models and reasoning settings. The results suggest a clear division of labor for autonomous experiments. The agent forms scientific hypotheses and uses quantitative tools to evaluate data, while deterministic code controls the hardware and enforces safety constraints.

13:00 JST画像/動画生成

OrganLens: CT 基礎モデルの臓器固有の表現学習

CT 検査では複数の臓器が撮影されますが、生物医学的な質問の多くは、特定の臓器の異常、予後、または長期的な変化に関係しています。これらの質問には、同じ CT ボリューム内の各臓器を個別に表現する必要があります。既存の CT 基礎モデルは通常、単一のボリュームレベルの表現を生成しますが、最近の解剖学を意識した手法では、事前に分離された臓器ボリュームをエンコードするか、画像を明示的に臓器トークン グループに分解します。前者は臨床的に関連する周囲のコンテキストを削除する可能性がありますが、後者は選択された臓器の特徴が形成される前に共有エンコーダーを条件付けしません。自己監視による臓器固有の表現学習のための OrganLens を紹介します。臓器のアイデンティティによって共有 CT エンコーダが条件付けされる一方、臓器固有の蒸留と解剖学的マスクの監視により、臓器固有の表現に解剖学的重み付けプールを行うための特徴が形成されます。推論時に、共有モデルは外部セグメンテーション マスクなしで 11 の臓器固有の表現を生成します。当社は、さまざまな取得および下流の評価にわたって、CT-RATE、RAD-ChestCT、INSPECT、および NLST に基づいて OrganLens を評価します。 CT で事前トレーニングされた DINOv2 と比較して、心臓の表現では CT-RATE 心拡大 AUROC が 0.910 から 0.953 に上昇し、肺の表現では NLST 肺がん死亡率の Harrell C 指数が 14.2\% 改善されました。グローバル表現は、テキストから画像への検索と画像からテキストへの検索でそれぞれ 33.09\% と 32.04\% の INSPECT Recall@10 に達します。器官関連のタスク全体にわたって、解剖学的に一致する表現は、より強力なタスク関連シグナルを提供する一方で、全体的な表現は幅広い有用性を保持します。 OrganLens は、共有エンコーダを使用して臓器固有の CT 表現を学習するためのスケーラブルなアプローチを提供します。より広範には、医学研究コミュニティに、コホートおよび臨床エンドポイント全体で臓器特異的な疾患を研究するための再利用可能なフレームワークを提供します。

原文 (English)

OrganLens: Organ-Specific Representation Learning for CT Foundation Models

A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a separate representation for each organ within the same CT volume. Existing CT foundation models commonly produce a single volume-level representation, while recent anatomy-aware methods either encode pre-separated organ volumes or explicitly disentangle images into organ token groups. The former may remove clinically relevant surrounding context, while the latter does not condition a shared encoder on a selected organ before its features are formed. We introduce OrganLens for organ-specific representation learning through self-supervision. An organ identity conditions a shared CT encoder, while organ-specific distillation and anatomy-mask supervision shape features for anatomy-weighted pooling into organ-specific representations. At inference, the shared model produces 11 organ-specific representations without external segmentation masks. We evaluate OrganLens on CT-RATE, RAD-ChestCT, INSPECT, and NLST across diverse acquisitions and downstream evaluations. Relative to CT-pretrained DINOv2, heart representations raise CT-RATE cardiomegaly AUROC from 0.910 to 0.953, while lung representations improve the Harrell C-index for NLST lung-cancer mortality by 14.2\%. The global representation reaches INSPECT Recall@10 of 33.09\% and 32.04\% for text-to-image and image-to-text retrieval, respectively. Across organ-related tasks, anatomically matched representations provide stronger task-relevant signal, while the global representation retains broad utility. OrganLens offers a scalable approach to organ-specific CT representation learning with a shared encoder. More broadly, it provides the medical research community with a reusable framework for studying organ-specific disease across cohorts and clinical endpoints.

13:00 JST研究/論文

CondPSE: グラフ用の条件付き変調を備えた多項式フィルター処理された構造エンコーダー

メッセージパッシング グラフ ニューラル ネットワークは 1-WL テストによって制限されており、非同型グラフを区別する位相構造を見逃す可能性があります。位置および構造エンコーディング (PSE) は、このようなトポロジー由来の信号を注入し、GPSE などの学習済み PSE エンコーダーは、ランダムなノード プローブからこれらの信号を生成するように単一のエンコーダーを事前トレーニングします。その後、その信号をフリーズして、下流のグラフ モデル全体で入力として再利用できます。 CondPSE は、学習可能な多項式グラフ フィルター バンクを標準ガウス ノード プローブに適用し、クロスフィルター、ローカル メッセージ パッシング、およびグラフ レベルの信号を条件とする FiLM スタイルの変調を通じて、結果として得られる構造応答ブランチを洗練する学習済み PSE エンコーダーです。 CondPSE は、ノード レベルの位置/構造ターゲットとグラフ レベルの不変条件を再構築するように事前トレーニングされ、ダウンストリームの入力エンコーディングとして使用するためにフリーズされます。合成構造識別ベンチマークでは、CondPSE は、1-WL 境界メッセージ パッシングでは不可能なグラフ構造を分離します。GPSE と比較して、CSL 精度が 42.9% から 97.3% に、EXP 精度が 68.3% から 99.9% に向上します。アブレーションは、多項式フィルター バンクがこの向上のほとんどを占めることを示しています。実際の分子特性の予測では、状況はさらに限定されます。ハイブリッド ローカル メッセージ パッシング/グローバル アテンション バックボーンを使用すると、CondPSE は GPSE を上回ることなく同等のパフォーマンスを発揮します。ZINC バックボーン スイープでは、2 つのエンコーダー間で一貫した順序付けが見られません。我々はこれらの結果を報告し、強力な合成構造識別だけでは、下流の統合の役割や構造事前学習ターゲットと分子特性ラベル間の不一致の可能性など、凍結学習済み PSE エンコーダーに下流の利点がもたらされない理由について議論します。

原文 (English)

CondPSE: A Polynomial-Filtered Structural Encoder with Conditional Modulation for Graphs

Message-passing graph neural networks are bounded by the 1-WL test and can miss topological structure that distinguishes non-isomorphic graphs. Positional and structural encodings (PSE) inject such topology-derived signals, and learned PSE encoders such as GPSE pretrain a single encoder to produce these signals from random node probes, which can then be frozen and reused as inputs across downstream graph models. We present CondPSE, a learned PSE encoder that applies a learnable polynomial graph filter bank to standard Gaussian node probes and refines the resulting structural-response branches through FiLM-style modulation conditioned on cross-filter, local message-passing, and graph-level signals. CondPSE is pretrained to reconstruct node-level positional/structural targets and graph-level invariants, and is then frozen for use as a downstream input encoding. On synthetic structural-discrimination benchmarks, CondPSE separates graph structures that 1-WL-bounded message passing cannot: it raises CSL accuracy from 42.9% to 97.3% and EXP accuracy from 68.3% to 99.9% relative to GPSE, and ablations show that the polynomial filter bank accounts for most of this gain. On real molecular property prediction, the picture is more limited. With a hybrid local-message-passing/global-attention backbone, CondPSE performs comparably to GPSE without surpassing it, and a ZINC backbone sweep shows no consistent ordering between the two encoders. We report these results and discuss why strong synthetic structural discrimination does not, on its own, yield a downstream advantage for frozen learned PSE encoders, including the role of downstream integration and possible mismatch between structural pretraining targets and molecular property labels.

13:00 JSTLLM/生成AI

TabRank: テーブルの再ランク付けのための思考連鎖の蒸留

質問に答えるために関連するテーブルを取得する機能は、構造化された情報を取得するための重要なタスクです。多段階検索システムは、効率的な第 1 段階の検索システムによって生成された候補リストを絞り込むために、リランカーに大きく依存しています。その結果、ニューラル リランカーと LLM ベースのリランキング手法は、従来のスパースまたはデンス検索モデルと比較して、意味の理解と推論において優れた能力を備えているため、ますます重要になっています。最近、明示的思考連鎖 (CoT) 推論を備えた大規模推論モデル (LRM) は、非構造化パッセージ検索におけるランキング品質の大幅な向上を示しています。この研究では、表形式検索のための推論リランカーをトレーニングするためのフレームワークである TabRank を紹介します。まず、Natural question Tables データセットの表形式の再ランキング用の 6728 の推論トレースの包括的なデータセットを提示します。次に、これらの推論トレースに基づいてコンパクトな推論モデルをトレーニングする 2 つのバリエーションを検討します。それは、明示的な CoT 蒸留と、プロンプト内の教師の推論トレースに基づいて生徒のリランカーを条件付けすることです。さまざまなドメインおよび複数テーブルのシナリオで、いくつかの配布外の一般化設定で TabRank のストレス テストを行います。私たちのアプローチは、さまざまなテーブル取得データセット全体でパフォーマンスを大幅に向上させ、ベース モデルと比較して、Acc@10 が HybridQA で 30.5%、SQA で 15.2%、TabFact で 52.9%、TATQA サブセットで 13.1% 増加しました。特に、TabRank は複数テーブルの推論を効果的に一般化します。コード、データ、モデルは https://github.com/AdarshSingh7647/TabRanker で入手できます。

原文 (English)

TabRank: Chain-of-Thought Distillation for Table Re-Rankers

The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval systems rely heavily on rerankers to refine candidate lists produced by efficient first-stage retrievers. As a result, neural rerankers and LLM-based reranking methods have become increasingly important due to their superior capacity for semantic understanding and reasoning compared to conventional sparse or dense retrieval models. Recently, Large Reasoning Models (LRMs) equipped with explicit chain-of-thought (CoT) reasoning have shown strong improvements in ranking quality in unstructured passage retrieval. In this work, we present TabRank, a framework for training reasoning rerankers for Tabular Retrieval. We first present a comprehensive dataset of 6728 reasoning traces for tabular reranking on the Natural Questions Tables dataset. We then explore two variants of training a compact reasoning model on these reasoning traces: explicit CoT distillation and conditioning the student reranker on the teacher's reasoning trace within the prompt. We stress-test TabRank on several out-of-distribution generalization settings on diverse domains and multi-table scenarios. Our approach significantly improves performance across a variety of table retrieval datasets, increasing Acc@10 by 30.5% on HybridQA, 15.2% on SQA, 52.9% on TabFact, and 13.1% on TATQA subsets of the Multi-Table QA Benchmark compared to the base model. Notably, TabRank generalizes effectively to multi-table reasoning. Our code, data and models are available at https://github.com/AdarshSingh7647/TabRanker

13:00 JSTLLM/生成AIエージェント

RIDGE: LLM 生成のオプション価格設定における検証とメソッド検出のための自律フレームワーク

自動コード生成は、量的金融において重要なツールになりつつあり、大規模な言語モデルが数学的モデル仕様から直接オプション価格設定の実装を生成できます。ただし、そのような実装を検証するには、従来のソフトウェア テストよりもはるかに多くのことが必要です。数値的な価格設定方法は、広範囲のモデル パラメーターにわたって数学的に一貫性があり、数値的に安定しており、信頼性を維持する必要があります。 RIDGE は、生成された価格設定の実装が構造化された裁定なしテスト、ストレス テスト、ベンチマーク比較、一貫性チェックの対象となる自律的な検証フレームワークです。検証証拠は診断的に解釈され、結果として得られる知識はリポジトリに蓄積され、モデル間および後続の検証反復間で再利用されます。これにより、価格設定の実装と検証方法の両方を体系的に改善することが可能になります。このフレームワークは 5 つの確率的ボラティリティ モデルに適用されます。これらの調査全体で、検出された実装上の欠陥はすべて削除され、2 つのケースでは、検証プロセス自体が新しい半分析的な価格設定手法につながります。補足資料は、GitHub リポジトリ: https://github.com/ShQiangLiu/ridge で入手できます。

原文 (English)

RIDGE: An Autonomous Framework for Validation and Method Discovery in LLM-Generated Option Pricing

Automated code generation is becoming an important tool in quantitative finance, where large language models can generate option pricing implementations directly from mathematical model specifications. Validating such implementations, however, requires considerably more than conventional software testing: numerical pricing methods must remain mathematically consistent, numerically stable, and reliable across a wide range of model parameters. We introduce RIDGE, an autonomous validation framework in which generated pricing implementations are subjected to structured no-arbitrage tests, stress tests, benchmark comparisons, and consistency checks. Validation evidence is interpreted diagnostically, while the resulting knowledge is accumulated in a repository and reused across models and successive validation iterations. This enables systematic refinement of both the pricing implementation and the validation methodology. The framework is applied to five stochastic volatility models. Across these studies, all detected implementation defects are removed and, in two cases, the validation process itself leads to new semi-analytic pricing methodologies. The supplementary material is available in the GitHub repository: https://github.com/ShQiangLiu/ridge.

13:00 JSTLLM/生成AI

VaLiDRec: 生成推奨用の可変長 LLM にアライメントされたセマンティック ID

生成的推奨は通常、クラスタリングと量子化を通じて構築された固定長のセマンティック識別子 (SID) を使用してアイテムを表します。ただし、これらの人工コードは項目のセマンティクスを過度に圧縮し、事前トレーニングされた LLM 語彙と不整合なままになり、コストのかかる自己回帰デコードを必要とする可能性があります。これを考慮して、可変長の LLM に合わせたセマンティック識別子に基づく生成推奨フレームワークである VaLiDRec を提案します。 VaLiDRec は、トークン重要度の推定、意味品質を意識した枝刈り、衝突を意識した改良を通じて、有益なネイティブ LLM 語彙トークンから SID を直接構築し、識別子の長さを項目の意味の複雑さに適応させることができます。ユーザーの好みをモデル化するために、VaLiDRec にはグラフ対応のソフト プロンプトが組み込まれており、トークン レベルのアイテム スコアリングによるトークン セットの予測として推奨事項を再定式化し、自己回帰 SID 生成とビーム検索を排除します。 4 つの現実世界のデータセットでの実験では、VaLiDRec がすべての評価指標にわたって強力な逐次的および生成的推奨ベースラインを一貫して上回るパフォーマンスを示しています。さらに、優れたゼロショット アイテムのコールド スタート パフォーマンスと、LC-Rec よりも 87.49 倍高速な推論を実現します。これらの結果は、LLM ネイティブの可変長意味識別子が、生成的な推奨に対してより表現力豊かで効率的なパラダイムを提供することを示しています。

原文 (English)

VaLiDRec: Variable-Length LLM-Aligned Semantic IDs for Generative Recommendation

Generative recommendation commonly represents items using fixed-length semantic identifiers (SIDs) constructed through clustering and quantization. However, these artificial codes may overcompress item semantics, remain misaligned with pretrained LLM vocabularies, and require costly autoregressive decoding. In light of this, we propose VaLiDRec, a generative recommendation framework based on variable-length, LLM-aligned semantic identifiers. VaLiDRec constructs SIDs directly from informative native LLM vocabulary tokens via token importance estimation, semantic-quality-aware pruning, and collision-aware refinement, allowing identifier lengths to adapt to item semantic complexity. To model user preferences, VaLiDRec incorporates graph-aware soft prompts and reformulates recommendation as token-set prediction with token-level item scoring, eliminating autoregressive SID generation and beam search. Experiments on four real-world datasets show that VaLiDRec consistently outperforms strong sequential and generative recommendation baselines across all evaluation metrics. It further achieves superior zero-shot item cold-start performance and 87.49$\times$ faster inference than LC-Rec. These results demonstrate that LLM-native variable-length semantic identifiers provide a more expressive and efficient paradigm for generative recommendation.

13:00 JST研究/論文

TopoGR: 生成的推奨におけるセマンティック ID の潜在構造の解明と保存

セマンティック ID ベースの生成推奨では、各アイテムを一連の個別のセマンティック ID にトークン化し、セマンティック ID を生成することで次のアイテムを予測します。ただし、既存の方法は通常、SID を独立した離散シンボルと見なし、学習されたセマンティック ID 空間のトポロジーを見落とすことがよくあります。トークン化と生成の間の構造的不一致を特定します。トークナイザーはセマンティック近傍関係を使用して構造化されたコード空間を学習しますが、ジェネレーターはセマンティック ID トークンを独立したカテゴリカル シンボルとして消費します。その結果、アイテムの関連性は正確なセマンティック ID の重複に限定され、セマンティック ID が重複しない意味的に類似したアイテムを識別することが困難になります。この問題に対処するために、我々は、ビット分解可能なセマンティック ID (バイナリ SID) に基づいたトポロジ保存生成推奨フレームワークである TopoGR を提案します。各バイナリ SID はビット分解可能な形式で学習され、明示的なハミング ジオメトリを公開しながら、標準の整数 SID に決定論的に変換できます。 TopoGR は、このトポロジを 3 つの段階で活用します。バイナリ SID 機能は、入力層でのハミング近接性を保持します。ハミング ソフト ターゲットは、トポロジを意識した監視を導入します。ハミング一貫性のある再ランキングにより、推論中に候補項目が予測されたバイナリ プロトタイプと一致します。さらに、ハミング トポロジが正確な SID 一致を超えて項目の関連性をキャプチャできることを検証します。 4 つのベンチマーク データセットでの実験では、TopoGR が推奨パフォーマンスにおいて既存の最先端のベースラインを常に上回っていることが示されています。

原文 (English)

TopoGR: Revealing and Preserving Latent Structure of Semantic ID in Generative Recommendation

Semantic ID-based generative recommendation tokenizes each item into a sequence of discrete semantic IDs and predicts the next item by generating semantic IDs. However, existing methods typically regard SIDs as independent discrete symbols, while often overlooking the topology of the learned semantic ID space. We identify a structural mismatch between tokenization and generation: the tokenizer learns a structured code space with semantic neighborhood relations, whereas the generator consumes semantic ID tokens as independent categorical symbols. Consequently, item relatedness is reduced to exact semantic ID overlap, making it difficult to identify semantically similar items whose semantic IDs do not overlap. To address this issue, we propose TopoGR, a topology-preserving generative recommendation framework based on Bit-decomposable Semantic ID(Binary SID). Each Binary SID is learned in a bit-decomposable form and can be deterministically converted to a standard integer SID, while exposing an explicit Hamming geometry. TopoGR exploits this topology at three stages: binary SID features preserve Hamming proximity at the input layer; Hamming soft targets inject topology-aware supervision; and Hamming-consistent reranking aligns candidate items with the predicted binary prototype during inference. We further verify that the Hamming topology can capture item relatedness beyond exact SID matching. Experiments on four benchmark datasets show that TopoGR consistently outperforms existing state-of-the-art baselines in recommendation performance.

13:00 JSTLLM/生成AI研究/論文

Laplace-PSN-IRT: LLM ベンチマークの神経項目応答理論モデルの不確実性の定量化

項目応答理論 (IRT) は、モデルの潜在能力を個々のベンチマーク項目のプロパティから分離することにより、大規模言語モデル (LLM) ベンチマークを評価するためのフレームワークとして最近提案されました。 PSN-IRT を含む既存のニューラル IRT アプローチは、点推定を使用してこれらの量を推定するため、不確実性の定量化と下流の統計的推論が制限されます。 Laplace-PSN-IRT は、トレーニングされた PSN-IRT モデルを近似ベイジアン事後推論で強化するポストホック最終層ラプラス近似であり、再トレーニングすることなくモデルの能力とアイテムの難易度に関する調整された不確実性を回復します。結果として得られる事後分布により、信頼区間、モデル間の確率的比較、およびフィッシャー情報に基づく項目選択へのパラメーターの不確実性の伝播が可能になります。標準的な LLM ベンチマーク リーダーボード上の 12 モデル間のペアごとの比較のほとんどは、点推定ランクが異なるにもかかわらず、統計的に区別できないことを示します。さらに、点推定のフィッシャー情報は単一の参照能力で評価されるため、多くのベンチマーク項目でほぼゼロになる可能性があるのに対し、事後期待フィッシャー情報は能力範囲全体で実質的により安定したままであることを示します。最後に、事後期待フィッシャー情報は、ほとんどの実験設定で小さなベンチマーク サブセットから完全なベンチマーク能力ランキングをより正確に復元し、同時に最小のサブセットの点推定パフォーマンスを一致させます。ホールドアウト予測カバレッジを使用して近似事後分布のキャリブレーションを検証し、アイテムの識別を固定として扱いながらアイテムの難易度をランダムとしてモデル化すると、このアーキテクチャで適切にキャリブレーションされた不確実性が生成されることがわかりました。

原文 (English)

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting uncertainty quantification and downstream statistical inference. We introduce Laplace-PSN-IRT, a post-hoc last-layer Laplace approximation that augments a trained PSN-IRT model with approximate Bayesian posterior inference, recovering calibrated uncertainty over model ability and item difficulty without retraining. The resulting posterior enables credible intervals, probabilistic comparisons between models, and propagation of parameter uncertainty into Fisher-information-based item selection. We show that most pairwise comparisons among 12 models on a standard LLM benchmark leaderboard are not statistically distinguishable despite differing point-estimate ranks. We further show that point-estimate Fisher information can become nearly zero for many benchmark items because it is evaluated at a single reference ability, whereas posterior-expected Fisher information remains substantially more stable across the ability range. Finally, posterior-expected Fisher information more accurately recovers full-benchmark ability rankings from small benchmark subsets in most experimental settings while matching point-estimate performance for the smallest subsets. We validate the calibration of the approximate posterior using held-out predictive coverage and find that modeling item difficulty as random while treating item discrimination as fixed produces well-calibrated uncertainty in this architecture.

13:00 JST研究/論文

ランキングのための構造を意識した相対ポリシーの最適化

ランキングは、最新の情報アクセス システムの基本的なコンポーネントです。強化学習 (RL) は、完全なランキング リストに対して定義された粗粒度のフィードバックとシステム レベルの目標を直接最適化するための柔軟なフレームワークを提供します。ただし、既存の RL ベースのランキング手法は通常、サンプリングされた各順列をアトミックな出力として扱い、主にスカラー報酬を通じて評価し、異なるランキング リスト間の構造的関係を見落としています。その結果、同様の報酬を持つが実質的に異なる置換パターンを持つ置換は、同等の最適化シグナルを受信する可能性があり、不正確なクレジット割り当てや過度に積極的なポリシー更新につながる可能性があります。この制限に対処するために、リストごとのランキングのための \textbf{S} 構造を意識した \textbf{R}elative \textbf{P}olicy \textbf{O} 最適化フレームワークである SRPO を提案します。 SRPO は、最上位に重み付けされた Kendall-tau 距離を使用して、サンプリングされた順列間の不一致を測定し、対応する距離によってペアごとの報酬の差を正規化します。ランキング変化の単位当たりの報酬の向上を定量化することで、効率的なローカルな改善、特にトップランクのポジションに関係する改善を強調します。 2 つのランキング シナリオにわたる実験結果は、順列レベルの違いを明示的にモデル化すると、リストごとのランキングの有効性と安定性が向上し、制限されたフィードバックと複雑なリスト レベルの最適化設定で特に良好なパフォーマンスが得られることを示しています。

原文 (English)

Structure-aware Relative Policy Optimization for Ranking

Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for directly optimizing coarse-grained feedback and system-level objectives defined over the complete ranking list. However, existing RL-based ranking methods typically treat each sampled permutation as an atomic output and evaluate it primarily through a scalar reward, overlooking the structural relationships among different ranking lists. Consequently, permutations with similar rewards but substantially different permutation patterns may receive comparable optimization signals, potentially leading to inaccurate credit assignment and overly aggressive policy updates. To address this limitation, we propose SRPO, a \textbf{S}tructure-aware \textbf{R}elative \textbf{P}olicy \textbf{O}ptimization framework for listwise ranking. SRPO measures the discrepancy between sampled permutations using a top-weighted Kendall-tau distance and normalizes their pairwise reward differences by the corresponding distances. It quantifies the reward improvement per unit of ranking change, thereby emphasizing efficient local refinements, particularly those involving top-ranked positions. Experimental results across two ranking scenarios demonstrate that explicitly modeling permutation-level differences improves the effectiveness and stability of listwise ranking, with particularly favorable performance in limited-feedback and complex list-level optimization settings.

13:00 JSTLLM/生成AI

ステアリング信号の出所: アクティベーション ステアリングでのアクティベーション ソースの選択

アクティベーションステアリングは、推論時に隠れ状態にベクトルまたは特徴を追加することによって言語モデルを制御しますが、これらのステアリング信号の上流のソースは二次的な詳細として扱われることがよくあります。我々は、このソースの選択をアクティベーションソースの選択として研究します。これは、ステアリング信号が構築される隠れ状態を収集するために使用されるソースコンテキストとアクティベーション読み出しポリシーの組み合わせです。下流の介入を固定したまま、3 つの命令調整モデルと 4 つのステアリング タスク ファミリにわたって、ソースのアクティブ化のみを変更するとステアリングの成功が大幅に変化することを示します。さらに、効果的なステアリングは、ソーステキストに望ましい動作が現れるかどうかだけでは説明できないこともわかりました。代わりに、モデルがターゲットの動作を生成または継続しようとしている実行境界状態から強いシグナルが発生します。この実現前/実現後の区別は、回答ベースのソースが機能する場合がある理由を説明しています。その有用なコンポーネントは、ターゲットの外観だけではなく、実行境界の方向に沿っています。このビューに基づいて、テール サブトラクションを導入します。これにより、境界状態から共有のプロンプトおよび継続セマンティクスが削除され、よりクリーンで安定したステアリング信号が生成されます。全体として、私たちの結果は、ステアリングが単にすでに現れているものだけではなく、モデルがこれから何をしようとしているかの表現に依存していることを示唆しています。

原文 (English)

Where Steering Signals Come From: Activation Source Selection in Activation Steering

Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation source selection: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built. Holding the downstream intervention fixed, we show across three instruction-tuned models and four steering task families that changing only the source activations substantially changes steering success. We further find that effective steering is not explained simply by whether the desired behavior appears in the source text. Instead, strong signals come from execution-boundary states, where the model is about to produce or continue the target behavior. This pre-/post-realization distinction explains why answer-based sources sometimes work: their useful component aligns with execution-boundary directions rather than target appearance alone. Building on this view, we introduce tail subtraction, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals. Overall, our results suggest that steering depends on representations of what the model is about to do, not merely on what has already appeared.

13:00 JST研究/論文

コンピューティングとデータに最適な事前トレーニングの橋渡し

古典的なコンピューティング最適化スケーリングの法則は、新しい事前トレーニング データが無制限に供給されることを前提としていますが、事前トレーニングは、高品質のデータの利用可能性を上回る速度でコンピューティングが増加する状況にますます突入しています。私たちは、コンピューティングに応じてデータが自由にスケーリングされるコンピューティング最適化スケーリングと、コンピューティングが無制限に拡張できる一方でコーパスが固定されるデータ最適化スケーリングを橋渡しする統一フレームワークである、コンピューティング データ (CD) スケーリング則を提案します。 CD スケーリングは、トークン有効性関数 $\eta$ を導入することで古典的なスケーリング則を拡張します。この関数は、完全な代替品から価値のないものまで、たとえばマルチエポックの繰り返しや新しいトークンとの言い換えを通じて生成された派生トークンの価値を定量化します。 Dolma-3 コーパスを使用して、14M から 600M パラメータまでのモデル サイズにわたって、マルチエポック反復と言い換えという 2 つのデータ拡張戦略に $\eta$ を適合させます。トークンの有効性は一定とはほど遠いことがわかりました。トークンの有効性は、モデルのサイズ、パラメーターごとのトークンの比率、派生データの量に大きく依存し、コーパスが拡張されると飽和します。 $\eta$ の関数形式は、モデル サイズまたはデータの可用性が増加するにつれて、データの代わりにコンピューティングを使用すると、収益が減少することを意味します。また、トレーニングを 3 つの運用体制 (コンピューティング バウンド、データ バウンド、モデル バウンド) に分割し、古典的なコンピューティングの最適な割り当てが、実際に関連する設定のほとんどにわたって最適ではないことを示しています。

原文 (English)

Bridging Compute- and Data-Optimal Pretraining

Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $\eta$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $\eta$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $\eta$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.

13:00 JST画像/動画生成

ScaleResfusion: 残留ベクトル場に基づく残留整流流

リアルワールド画像復元 (Real-IR) は、複雑で未知の劣化から高品質 (HQ) 画像を復元することを目的としています。最近の拡散ベースの方法は知覚品質を大幅に改善しましたが、現在の設計では 2 つの重要な課題が未解決のままです。ガウス ノイズから開始する方法は時間がかかり、多くの場合、劣化した入力に対する忠実度が低くなります。残差ベースの手法は通常、ゼロからトレーニングするため、最新の事前トレーニングされた生成事前分布を活用することが困難になります。この論文では、事前にトレーニングされたテキストから画像への修正フロー モデルに基づいて構築された、現実世界の画像復元のためのスケーラブルな拡散フレームワークである ScaleResfusion を紹介します。私たちの手法の中核は残差整流であり、標準整流に残差項 R を導入します。純粋なノイズから開始する代わりに、ノイズの多い低品質 (LQ) 画像から開始し、正確な加速ポイントを許可する残留トランスポート パスを使用します。 Residual Rectified Flow は、残差ベクトル場を学習することにより、出力分布と線形拡散プロセスを事前トレーニング済みの整流モデルと一致させます。これにより、パラメータ効率の高い微調整を大規模に行うことが可能になります。さらに、復元品質を維持しながらサンプリングコストを削減するために、知識蒸留パイプラインを導入します。現実世界の複数の復元タスクに関する広範な実験により、ScaleResfusion がはるかに高い効率で最先端のパフォーマンスを達成できることが示されています。これらの結果は、事前にトレーニングされた大規模な拡散モデルを現実世界の画像復元に適応させる実用的でスケーラブルな方法を示唆しています。私たちのコードとモデルは https://github.com/YukinoshitaLove/ScaleResfusion で入手できます。

原文 (English)

ScaleResfusion: Residual Rectified Flow based on Residual Vector Field

Real-world Image Restoration (Real-IR) aims to recover high-quality (HQ) images from complex and unknown degradations. Although recent diffusion-based methods have substantially improved perceptual quality, their current designs leave two key challenges unresolved. Methods that start from Gaussian noise are slow and often less faithful to the degraded input. Residual-based methods usually train from scratch, which makes it hard to exploit modern pre-trained generative priors. In this paper, we present ScaleResfusion, a scalable diffusion framework for real-world image restoration built on pre-trained text-to-image rectified-flow models. The core of our method is Residual Rectified Flow, which introduces the residual term R into Standard Rectified Flow. Instead of starting from pure noise, it uses a residual transport path that starts from noisy low-quality (LQ) images and admits an exact acceleration point. By learning the residual vector field, Residual Rectified Flow keeps the output distribution and linear diffusion process consistent with the pre-trained rectified-flow models. This makes parameter-efficient fine-tuning possible at scale. We further introduce a knowledge-distillation pipeline to reduce sampling cost while maintaining restoration quality. Extensive experiments on multiple real-world restoration tasks show that ScaleResfusion achieves state-of-the-art performance with much higher efficiency. These results suggest a practical and scalable way to adapt large pre-trained diffusion models to real-world image restoration. Our code and models are available at https://github.com/YukinoshitaLove/ScaleResfusion.

13:00 JSTLLM/生成AI画像/動画生成ビジネス/資金調達

CLBench-V: グラウンディングから知識獲得までのマルチモーダルコンテキスト学習の評価

現実世界のタスクでは、多くの場合、モデルが事前トレーニングされた知識だけに依存するのではなく、タスク固有のコンテキストから学習する必要があります。最近の研究ではこの機能がコンテキスト学習として強調されていますが、既存の評価は主にテキストのコンテキストに焦点を当てています。しかし、実際の多くの設定では、学ぶべきコンテキストは多様です。科学的発見は図や表を通じて伝えられ、財務指標は変換されたレポート全体に散在し、空間的な決定は地図、シーン、または Web ページに依存します。マルチモーダル コンテキスト学習のベンチマークである CLBench-V を紹介します。これは、コンテキストの基礎付け、新しい情報の適用、新しい知識の学習という 3 つの次元に沿ってタスクを整理することで、コンテキストの使用が中断される場所を特定するという困難に対処します。 CLBench-V は、変換された公開ベンチマークと、科学、金融、長文文書の理解、空間推論、Web ベースの視覚的な質問応答などの領域にわたる新しく構築されたデータセットを組み合わせます。ドメイン固有のコンテキスト学習タスクを構築するコストを削減するために、新しく構築されたデータセットに対して自動化された構築およびフィルタリング手順をさらに使用します。 3,443 のインスタンスと 6 つの最近のマルチモーダル モデル全体での最高の総合スコアはわずか 0.2847 であり、マルチモーダル コンテキスト学習がまだ飽和には程遠いことを示しています。さらに、InternVL3.5-30B-A3B はコンテキストの基礎付けと新しい知識の学習で最高のパフォーマンスを発揮し、Qwen3.5-Plus は新しい情報のアプリケーションで最高のパフォーマンスを発揮します。さらに、ジャッジの信頼性、コンテキストの長さ、画像数、代表的な失敗ケースを分析します。コードは https://github.com/IamLihua/CLBench-V で入手できます。

原文 (English)

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.

13:00 JSTLLM/生成AIエージェント

LLM エージェントで安全な MCP ツールを使用するためのハイブリッド分析

大規模言語モデル (LLM) エージェントの急速な開発により、現実世界のさまざまなタスクに広く採用できるようになりました。 LLM エージェントと外部環境間の対話を標準化するために、モデル コンテキスト プロトコル (MCP) ツールが事実上の標準として登場し、これらのシステムに広く統合されています。ただし、MCP ツールを使用すると、LLM エージェントが悪意のあるアクションや不正なアクションを実行するように誘導される可能性があるため、新たな安全上のリスクも生じます。これまでの研究では、LLM エージェントでツールの使用を保護するための防御策が提案されてきましたが、ほとんどの手法は静的分析、つまりプロンプトと生成された出力の検査に依存しているため、防御の有効性と堅牢性が制限されています。これらの制限に対処するために、私たちは、ライフサイクルを意識した静的-動的共同分析を活用することで、LLM エージェントでの MCP ツールの使用を保護するように設計されたハイブリッド分析ベースの防御フレームワークである MTGuard を提案します。広範な評価により、MTGuard は、無害なユーザー タスクのパフォーマンスを維持しながら、さまざまな LLM エージェントにわたる複数のカテゴリーの有害なツールの使用を効果的に軽減することが実証されています。

原文 (English)

Hybrid Analysis for Secure MCP Tool Use in LLM Agents

The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactions between LLM agents and external environments, Model Context Protocol (MCP) tools have emerged as a de facto standard and have been widely integrated into these systems. However, the use of MCP tools also introduces new safety risks, as LLM agents can be induced to perform malicious or unauthorized actions. Although prior work has proposed defenses for securing tool use in LLM agents, most methods rely on static analysis, i.e., inspecting prompts and generated outputs, which limits the defense effectiveness and robustness. To address these limitations, we propose MTGuard, a hybrid analysis-based defense framework designed to safeguard the use of MCP tools in LLM agents by leveraging lifecycle-aware static-dynamic co-analysis. Extensive evaluation demonstrates that MTGuard effectively mitigates multiple categories of harmful tool use across different LLM agents while maintaining performance on benign user tasks.

13:00 JST研究/論文

LoRA 微調整のための Stiefel 多様体に対するリトラクションフリーの最適化

シュティーフェル多様体に対する最適化は、さまざまな機械学習タスクにおいて重要な役割を果たします。既存の手法では、大規模な行列に対してコストのかかる直交正規化を必要とするリトラクション演算子を使用するか、慎重なステップ サイズの選択とペナルティ パラメーターの調整に依存するランディング手法を採用しています。これらの課題に対処するために、私たちは多様体に直接着地する、リトラクションフリーでペナルティパラメータフリーのアルゴリズムを提案します。二次ペナルティ関数の強い凸状の特性とシュティーフェル多様体の近位平滑性を利用することで、一定ステップ サイズと減少するステップ サイズの両方で最もよく知られている反復複雑さによるグローバル収束保証を確立します。次に、大規模言語モデルの低ランク適応 (LoRA) 微調整問題を多様体最適化問題として再定式化し、ジオメトリ加速適応のための Manifold-LoRA を導入します。このアプローチでは、提案された着地テクニックと慎重に設計されたステップ サイズ戦略を採用して、トレーニング プロセスを加速します。ベンチマーク データセットの数値実験により、提案された方法の効率と強力な下流パフォーマンスが実証されています。

原文 (English)

Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning

Optimization over the Stiefel manifold plays a significant role in various machine learning tasks. Existing methods either use the retraction operators, requiring costly orthonormalization for large-scale matrices, or employ landing methods that rely on careful step size selection and penalty parameter tuning. To address these challenges, we propose a retraction-free and penalty parameter-free algorithm that directly lands on the manifold. By leveraging the strongly-convex-like property of the quadratic penalty function and the proximal smoothness of the Stiefel manifold, we establish global convergence guarantees with the best-known iteration complexities under both constant and diminishing step sizes. Then, we reformulate the low-rank adaptation (LoRA) fine-tuning problem for large language models as a manifold optimization problem, introducing Manifold-LoRA for geometry-accelerated adaptation. This approach employs the proposed landing technique and a carefully designed step size strategy to accelerate the training process. Numerical experiments on benchmark datasets demonstrate the efficiency and strong downstream performance of the proposed method.

13:00 JSTLLM/生成AIエージェント

キャスト: LLM エージェントのターンレベル教師としてのゲーム ソルバー

長期的なゲームで動作するように大規模言語モデル (LLM) をトレーニングすることは、ジェネラリスト的な意思決定に向けた有望なステップですが、検証可能な報酬を伴う強化学習 (RLVR) は、どの決定が成功を決定するかについてほとんど明らかにしない、まばらな最終報酬に依存しています。プロセス信号が高密度であれば、この欠けているターンレベルのクレジットを供給できる可能性がありますが、既存の信号源では、安価と正確さの両方を維持するのが困難です。ゲーム ソルバーの状態値の変化により、アクションが状態を成功に向けて進めるかどうかが明らかになることが観察されます。この洞察に基づいて、これらの値の変化をソルバーの利点に変換し、ターンレベル信号として RLVR に注入する CAST (ソルバー教師からのクレジット割り当て) を提案します。さらに、ソフト最適ソルバーの仮定の下では、ソルバーの利点を最大化することは、教師ロジットではなくスカラー値のみを必要とする、ソルバーからのポリシーに基づく蒸留と同等であることを示します。倉庫番、マインスイーパー、ラッシュアワー全体で、CAST は、ドメイン内および未確認の両方の難易度評価の下で、すべてのゲームでトレーニングされたすべてのベースラインを上回り、ALFWorld と WebShop で最高の平均ゼロショット パフォーマンスを達成しました。私たちのコードは https://github.com/Wloner0809/CAST で入手できます。

原文 (English)

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state toward success. Building on this insight, we propose CAST (Credit Assignment from Solver Teachers), which converts these value changes into solver advantages and injects them into RLVR as turn-level signals. We further show that, under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. Our code is available at https://github.com/Wloner0809/CAST.

13:00 JST画像/動画生成

緑内障検出のためのバランスの取れたソフト専門家混合モデル

緑内障は視神経に損傷を与える一連の眼疾患であり、多くの場合眼圧の上昇によって引き起こされます。これは不可逆的な視力喪失の主な原因であり、通常はゆっくりと痛みを伴わずに進行するため、重大な損傷が発生するまで気づくのが困難です。したがって、視力喪失の進行を予防または遅らせるには、早期発見が非常に重要です。近年、深層学習ベースのユニモーダル モデルにより緑内障検出の精度と効率が向上し、医師は早期の診断、より適切なモニタリング、タイムリーな治療のためのツールを利用できるようになりました。これに基づいて、さまざまな画像モダリティの強みを活用して、より豊かで堅牢な表現を学習し、緑内障の検出精度をさらに向上させるマルチモーダル モデルが登場しました。ただし、マルチモーダル学習は、共同学習の目的により、不均衡で最適化が不十分なユニモーダル表現などの課題に直面しています。これに対処するために、3 人のエキスパートと負荷分散損失を備えたバランスの取れたソフト混合エキスパート モデルを提案します。パフォーマンスは AUC によって測定され、私たちが提案する方法は、すべてのユニモーダル ベースライン、従来のマルチモーダル モデル、および現在の最先端のバランスの取れたマルチモーダル モデルのパフォーマンスを上回ります。提案されたモデルは、糖尿病性網膜症などの他の疾患の検出に一般化できます。

原文 (English)

Balanced Soft mixture-of-expert model for Glaucoma Detection

Glaucoma is a group of eye diseases that damage the optic nerve, often caused by elevated intraocular pressure. It is a leading cause of irreversible vision loss and is typically developed slowly and painlessly, making it difficult to notice until significant damage has occurred. Therefore, early detection is crucial to prevent or slow the progression of vision loss. In recent years, deep learning based uni-modal models have improved the accuracy and efficiency of glaucoma detection, empowering doctors with tools for earlier diagnosis, better monitoring, and timely treatment. Building on this, multi-modal models have emerged, leveraging the strengths of different imaging modalities to learn richer and more robust representations, further enhancing glaucoma detection accuracy. However, multi-modal learning faces challenges such as imbalanced and under-optimized uni-modal representations due to joint learning objectives. To address this, we propose a balanced soft mixture-experts model with three experts and load balancing loss. The performance is measured by AUC, our proposed method surpasses the performance of all uni-modal baselines, conventional multi-modal models, and current stateof- the-art balanced multi-modal models. The proposed model can be generalized to other disease detections such as diabetic retinopathy.

13:00 JST研究/論文

バックグラウンド分解および事前調整された PSFD をウォームスタートするための物理学に基づいたニューラル オペレーター: スケーラブルな 3D EUV マスク シミュレーションを可能にする

EUV リソグラフィーにおける電磁 (EM) 散乱問題に対して、擬似スペクトル周波数領域 (PSFD) 方程式を使用してトレーニングされた物理情報に基づくニューラル オペレーター (PINO) を紹介します。フーリエ ニューラル演算子は、2 次元の横方向 ($xy$) 分岐と 1 次元の軸方向 ($z$) 分岐に因数分解され、背景分解で自己矛盾なくトレーニングされます。したがって、マスクと多層応答の間の完全ベクトル結合は、有限次数ボルン近似を呼び出すことなく保持されます。このようにして、計算ドメインのサイズが大幅に縮小され、計算コストが削減されます。 PINO は、事前に計算された EM フィールド ソリューションを使用せずに、各トレーニング反復でランダムにサンプリングされた LithoBench ライブラリからの約 16,000 のマスク設計でトレーニングされます。 PINO サロゲート モデルは、参照 PSFD ソリューションと比較した、保持されたマスク パターンの散乱強度について、約 $7 \times 10^{-3}$ の平均絶対誤差の予測を生成します。スペクトル減衰と組み合わせると、PINO ウォーム スタート初期化により、バックグラウンド分解された PSFD ソルバーがより詳細な離散化で高速化されます。

原文 (English)

Physics-Informed Neural Operator for Warm-Starting Background-Decomposed and Preconditioned PSFD: Enabling Scalable 3-D EUV Mask Simulation

We present a physics-informed neural operator (PINO) trained with pseudo-spectral frequency-domain (PSFD) equations for electromagnetic (EM) scattering problems in EUV lithography. The Fourier neural operator is factorized into a two-dimensional lateral ($xy$) branch and a one-dimensional axial ($z$) branch and is trained self-consistently with background decomposition.Thus, the full-vector coupling between the mask and the multilayer response is retained without invoking a finite-order Born approximation. In this way, the computational domain size is significantly reduced, thereby lowering the computational cost. The PINO is trained on approximately 16,000 mask designs from the LithoBench library sampled randomly at each training iteration without using precomputed EM field solutions. The PINO surrogate model yields predictions with a mean absolute error of about $7 \times 10^{-3}$ for the scattered intensity of held-out mask patterns relative to the reference PSFD solution. Combined with spectral damping, the PINO warm-start initialization accelerates the background-decomposed PSFD solver on finer discretizations.

13:00 JSTエージェント

Specula: システム コードの自律モデル チェックのための正式な仕様のスケーリング

Specula は、大規模で複雑なシステム コードに対して高品質の正式な仕様を生成し、その仕様を非常に効果的なモデル チェックとバグ発見に使用する、プッシュ ボタン エージェント システムです。 Specula は、大規模言語モデル (LLM) ベースのコーディング エージェントを採用して、ターゲット システムの正確性プロパティを記述する不変式や、適切なレベルの抽象化でシステム実装を記述する形式モデルなどの TLA+ 仕様を自律的に開発します。 Specula は完全に自律的であるため、(従来の人間中心のアプローチのように) 形式的な手法を現実世界のシステム コードに適用する際の障壁が排除されます。一方、Specula は、エージェントがシステム コードとその動作について理解を深められるようにすることで仕様の品質を反復的に向上させる自己進化ループを通じて、報酬ハッキングや幻覚などの LLM 主導の手法の制限に対処します。私たちは Specula を使用して 48 のオープンソース システム プロジェクトをチェックしました。 Specula は、既存のアプローチでは発見するのが難しい多くの深刻なバグを含む 249 個のバグを発見しました。 Specula はいくつかの企業で使用されており、https://github.com/specula-org/Specula で管理されています。

原文 (English)

Specula: Scaling formal specifications for autonomous model checking of system code

Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding. Specula employs large language model (LLM) based coding agents to autonomously develop TLA+ specifications, including invariants that describe correctness properties of the target system and formal models that describe the system implementation with the right level of abstractions. Specula is fully autonomous and thus eliminates the barrier of applying formal methods to real-world system code (as in traditional human-centric approaches). Meanwhile, Specula addresses limitations of LLM-driven techniques like reward hacking and hallucinations through self-evolving loops that iteratively improve specification quality by enabling the agents to deepen their understanding of system code and its behaviors. We have used Specula to check 48 open-source system projects; Specula found 249 bugs including many deep bugs that are hard to find by existing approaches. Specula has been used by several companies and is maintained at https://github.com/specula-org/Specula.

13:00 JSTLLM/生成AI

言語学者を雇うたびに推論コストが下がる: 効果的なプロンプト圧縮装置としての言語規則について

プロンプト圧縮により LLM 入力が短縮されて推論コストが削減されますが、既存の方法では LM フォワード パスを通じてトークンの重要性がスコアリングされます。このような微妙でコストのかかるトークンの選択が必要かどうかは疑問が残ります。圧縮には有益なコンテンツを特定する必要があります。この問題は、言語研究が決定論的なルールとして運用できる手がかりを通じて長年取り組んできました。したがって、圧縮時に LM ベースのスコアリングを行わずに、\textbf{言語規則のみ} が効果的なプロンプト圧縮として機能できるか、と考えます。これに対処するために、語彙、構文、意味、および談話のシードに対してオフラインの進化的検索を実行して、競合するルールの組み合わせを見つけます。結果として得られる言語圧縮プログラムは、展開時に LM フォワード パスを必要とせず、圧縮に CPU 側の処理のみを使用します。圧縮品質と再構築の忠実度のバランスをとるために、デュアルパス プロトコルを使用して評価します。短い文章、複数文書の推論、対話メモリ QA データセットにわたって、進化したコンプレッサーは、最近の高度なプロンプト圧縮戦略と同様のパフォーマンスを達成します。パフォーマンスは軽度から中度の圧縮下で最も強く、圧縮がより強力になると低下しますが、直接パスと再構築パスは異なるパターンを示します。進化的分析により、効果的な圧縮により言語レベル全体で信号が融合され、圧縮率が増加するにつれてルールがトークン プルーニングからセンテンス抽出に移行することが明らかになりました。

原文 (English)

Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors

Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules. We therefore ask: can \textbf{linguistic rules alone} serve as effective prompt compressors, without LM-based scoring at compression time? To address this, we conduct offline evolutionary search over lexical, syntactic, semantic, and discourse seeds to find competitive rule combinations. The resulting linguistic compressor requires no LM forward pass at deployment and uses only CPU-side processing for compression. We evaluate it with a dual-path protocol to balance compression quality and reconstruction fidelity. Across short passages, multi-document reasoning, and dialogue-memory QA datasets, evolved compressors achieve performance similar to that of recent advanced prompt-compression strategies. Performance is strongest under light-to-moderate compression and degrades as compression becomes more aggressive, while the Direct and Reconstruction paths exhibit distinct patterns. Evolutionary analysis reveals that effective compression fuses signals across linguistic levels and, as the compression ratio increases, rules shift from token pruning to sentence extraction.

13:00 JST研究/論文

模擬連合学習を使用した慢性腎臓病予測のための説明可能な AI

腎機能が徐々に失われることを特徴とする慢性腎臓病(CKD)は、依然として公衆衛生上の重大な課題となっています。早期発見は、重篤な合併症を予防し、患者の転帰を向上させるために非常に重要です。この研究では、VotingClassifier を備えた Federated Learning (FL) を使用して臨床データセットを使用して CKD を予測し、ランダム フォレスト、AdaBoost、および XGBoost を利用してグローバル サーバーに最適なモデルを比較および特定しました。さらに、クライアント側でモデルのパフォーマンスを最適化するために GridSearchCV が適用されました。モデルの透明性と信頼性を高めるために、予測メカニズムを解釈するために説明可能な AI (XAI) 技術が組み込まれました。グローバル モデルの平均精度は 99% であり、CKD の早期診断をサポートし、データ駆動型のヘルスケア ソリューションを進歩させる上で、解釈可能な FL モデルの可能性を強調しています。

原文 (English)

Explainable AI for Chronic Kidney Disease Prediction Using Simulated Federated Learning

Chronic Kidney Disease (CKD), characterized by the gradual loss of kidney function, remains a significant public health challenge. Early detection is crucial for preventing severe complications and enhancing patient outcomes. In this study, Federated Learning (FL) with a VotingClassifier was used to predict CKD using a clinical dataset, where Random Forest, AdaBoost, and XGBoost were utilized to compare and identify the best-fitting model for the global server. Additionally, GridSearchCV was applied to optimize the models' performance on the client's side. To enhance model transparency and trustworthiness, explainable AI (XAI) techniques were incorporated to interpret the prediction mechanisms. The global model's average accuracy was 99%, highlighting the potential of interpretable FL models in supporting early CKD diagnosis and advancing data-driven healthcare solutions.

13:00 JSTLLM/生成AI研究/論文

プログレッシブ サンプリングによる大規模なデータ品質プロファイリング: データ中心の AI パイプラインのベンチマーク

データ品質プロファイリング (欠損値率、重複分数、外れ値密度、機能依存性違反の計算) は、データ中心の AI パイプラインの基礎ですが、数百万行にわたる徹底的なスキャンは、ほぼリアルタイムの監視には法外に時間がかかります。プログレッシブ サンプリングが標準的な代替方法です。未解決の問題は、どの戦略が大規模なプロファイルの忠実度を最もよく維持するかということです。 3 つの現実世界のデータセット (NYC 311、NYPD 逮捕、UCI Adult、最大 500,000 行)、IoT センサー ストリーム (230 万行)、以下を含む 2 つの超大規模実際のデータセットに対して、ブラインド (ランダム均一、幾何学的、山根、クラスター) およびプロキシ ガイド付き (メトロポリス ヘイスティングス、DAG、列タイプまたは品質スコアによる層別、重要度加重) の 9 つのサンプリング戦略をベンチマークします。ウルトラマラソン ランニング (最大 740 万行)、および 5x10^6 行にスケールされた合成データ。推定値が鮮明になるという仮定に反して、ブラインド代表サンプラーが一様に優勢です。 5% の予算では、ランダムなユニフォームは NYC 311 で 0.49% の平均相対誤差を達成します。 DAG ガイド付き MCMC の収率は 19.5% (約 40 倍悪い) で、すべての実際のデータセット全体で DAG は 11 ~ 49 倍悪いです (Wilcoxon W=0、p=0.002、n=9 ペア)。クラスターのサンプリングはランダムな均一と一致します (MRE 0.110 対 0.111)。プロキシ ガイド付きメソッドは、DAG の障害モード (MRE 0.20 ~ 0.35) を共有します。大規模な場合、ランダム均一はほぼ線形 (O(N^{0.964})) ですが、DAG は超線形 (O(N^{1.272})) で、超大規模データでは 28 ~ 47 倍遅く、精度は 6 倍悪くなります。根本的な原因は IQR プロキシの不一致です。プロキシガイド付きサンプラーは数値の外れ値を過剰に追求し、品質欠陥はプロキシには見えないカテゴリ列に集中します。実用的な発見: サンプラーの品質は、ドメイン知識ではなく代表性によって決まります。大規模な運用グレードの品質プロファイリングには、スキーマフリーのランダムな均一サンプリングまたはクラスター サンプリングで十分です。

原文 (English)

Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines

Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundational for data-centric AI pipelines, yet exhaustive scans over millions of rows are prohibitively slow for near-real-time monitoring. Progressive sampling is the standard alternative; the open question is which strategy best preserves profile fidelity at scale. We benchmark nine sampling strategies -- blind (random uniform, geometric, Yamane, cluster) and proxy-guided (Metropolis-Hastings, DAG, stratified by column type or quality score, importance-weighted) -- on three real-world datasets (NYC 311, NYPD arrests, UCI Adult; up to 500K rows), an IoT sensor stream (2.3M rows), two ultra-large real datasets including Ultra-Marathon Running (up to 7.4M rows), and synthetic data scaled to 5x10^6 rows. Contrary to the assumption sharpens estimates, blind representative samplers dominate uniformly. At a 5% budget, random uniform achieves 0.49% mean relative error on NYC 311; DAG-guided MCMC yields 19.5% (approx. 40x worse), and across all real datasets DAG is 11-49x worse (Wilcoxon W=0, p=0.002, n=9 pairs). Cluster sampling matches random uniform (MRE 0.110 vs. 0.111); proxy-guided methods share DAG's failure mode (MRE 0.20-0.35). At scale, random uniform is near-linear (O(N^{0.964})) while DAG is super-linear (O(N^{1.272})), running 28--47x slower on ultra-large data with 6x worse accuracy. The root cause is an IQR proxy mismatch: proxy-guided samplers over-pursue numeric outliers, while quality defects concentrate in categorical columns invisible to the proxy. The actionable finding: representativeness, not domain knowledge, determines sampler quality -- schema-free random uniform or cluster sampling suffices for production-grade quality profiling at scale.

13:00 JST研究/論文

Raven: スパース メモリ ルーティングを使用した高再現率シーケンス モデリング

線形時間シーケンス モデルにおけるロングコンテキストの再現は、メモリへの書き込み方法におけるトレードオフを浮き彫りにします。状態空間モデル (SSM) や線形トランスフォーマーなどの状態ベースの線形モデルは、高密度に書き込み、新しく到着したトークンごとに状態全体を更新するため、干渉が発生し、特定の過去のトークンの回復が困難になります。スライディング ウィンドウ アテンション (SWA) は逆の動作を示します。つまり、明示的なトークン表現を格納することでまばらに書き込みますが、固定ウィンドウ内でのみ行われるため、関連するトークンが削除されるとリコールが低下します。これらのモデル間を補間して、線形時間シーケンス モデルである Raven を導入します。これは、メモリ スロットの固定セットを維持し、各ステップで学習された入力依存のルーティングを介して選択されたサブセットのみを減衰および更新します。これにより、Raven は SWA の位置ベースの上書きとハードエビクションを軽減しながら、SSM の密な状態更新による干渉を軽減し、長距離コンテンツをより効果的に保存できるようになります。リコール集約型のベンチマーク全体で、Raven は以前の線形時間ベースラインと同等かそれを上回り、SWA と SSM の両方が急激に低下する強力なロングコンテキストのリコールを実現します。トレーニング長の 16 倍のコンテキスト長に外挿しても引き続き効果があり、ハイブリッド アーキテクチャでも同様の効果が得られます。

原文 (English)

Raven: High-Recall Sequence Modeling with Sparse Memory Routing

Long-context recall in linear-time sequence models highlights a tradeoff in how they write to memory. State-based linear models, such as state-space models (SSMs) and linear Transformers, write densely, updating the entire state for each newly arrived token, which leads to interference and makes specific past tokens hard to recover. Sliding-window attention (SWA) exhibits the opposite behavior: it writes sparsely by storing explicit token representations, but only within a fixed window, so recall drops once the relevant token is evicted. Interpolating between these models, we introduce Raven, a linear-time sequence model that maintains a fixed set of memory slots and, at each step, decays and updates only a selected subset via learned, input-dependent routing. This lets Raven mitigate SWA's position-based overwriting and hard eviction while reducing interference from dense state updates in SSMs, thereby preserving long-range content much more effectively. Across recall-intensive benchmarks, Raven is competitive with or outperforms prior linear-time baselines, achieving strong long-context recall where both SWA and SSMs sharply degrade. It remains effective when extrapolating to context lengths as large as 16x its training length, with similar gains in hybrid architectures.

13:00 JST研究/論文

Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance

In Bayesian neural networks (BNNs), variational inference is a widely adopted framework for modeling uncertainty in a distributional way, w…

13:00 JSTLLM/生成AIエージェント

MARS: Multi-Agent Re-ranking for Repeat-Order Food Delivery Recommendation

Large language models (LLMs) are increasingly used in recommender systems, but it is often unclear how much performance can be obtained fro…

13:00 JST研究/論文

From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation

Eliciting explainable AI (XAI) requirements from stroke survivors presents a methodological challenge with direct implications for the desi…

13:00 JST研究/論文

Emergent Latent-State Computation under Stochastic Volatility

Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence mode…

13:00 JST画像/動画生成

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns

Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction…

13:00 JST研究/論文

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definit…

13:00 JST研究/論文

Automated Numerical Stability Analysis of Deep Learning Operators

Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or…

13:00 JST画像/動画生成

Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models

Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centre…

13:00 JST研究/論文Llama

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of…

13:00 JST画像/動画生成

I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial prog…

13:00 JST画像/動画生成

ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization

Unmanned aerial vehicle (UAV)-satellite cross-view geo-localization matches UAV images against satellite imagery and has achieved impressiv…

13:00 JST画像/動画生成

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute…

13:00 JST画像/動画生成

Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework

Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure. A rece…

13:00 JSTLLM/生成AI画像/動画生成

Visual prompt engineering for video models

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential techniq…

13:00 JST画像/動画生成

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes. Existing AIGC detection methods…

13:00 JST画像/動画生成

CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting

Recent advances in 3D Gaussian Splatting (3DGS)-based wireless radiance field (WRF) reconstruction provide an efficient solution for wirele…

13:00 JST画像/動画生成エージェント

The LAIA Dataset: Labelled Attention for Intelligent Automobiles

The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large v…

13:00 JSTLLM/生成AI

IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment

Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object. Conventional EA methods m…

13:00 JSTLLM/生成AI

Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs

Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evide…

13:00 JST研究/論文

Physics-Informed Broad Learning System: An Efficient Backpropagation-Free Framework for Solving Partial Differential Equations

Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs) by embedding…

13:00 JST研究/論文

Contrastive Representation Learning of Longitudinal Disease Trajectories on Temporal Graphs

Understanding disease trajectories from longitudinal clinical data remains challenging due to complex temporal dynamics and heterogeneous p…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries

Interdisciplinary research is accelerating, yet scientific papers remain difficult to understand outside their home fields. We study large…

13:00 JSTLLM/生成AI

Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models

Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injec…

13:00 JST研究/論文

F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill

F(AI)2R is FAIR research with AI in the loop, twice: an AI-assisted authoring pass and a machine-readable audit pass over every artefact. A…

13:00 JST画像/動画生成研究/論文

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmark…

13:00 JST研究/論文

KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models

As quantum computing continually improves, ensuring the reliability and correctness of quantum libraries has become increasingly critical.…

13:00 JST研究/論文

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing

Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. T…

13:00 JSTLLM/生成AI

MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

Psychotherapists need repeated training and supervision by experts; however, scalability is problematic. Here we present MyMentorLLM, a mul…