Skip to the content.

AIニュース 2026-06-19

自動生成: 2026-06-19 13:56 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. New usage analytics and updated spend controls for enterprisesOpenAI

    OpenAI introduces new spend controls and usage analytics for ChatGPT…

  2. Improving health intelligence in ChatGPTOpenAI

    Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness resp…

  3. Using AI to help physicians diagnose rare genetic diseases affecting childrenOpenAI

    Researchers used an OpenAI reasoning model to help diagnose rare dise…

  4. ChatGPTで広告テスト、日本でも開始 非表示にする方法は?ITmedia AI+

    米OpenAIの日本法人は、ChatGPTでの広告表示テストを日本でも始めたと発表した。広告を非表示にする方法は?

  5. Gartnerが警鐘 プライバシー法執行が本格化、CISOは何を見直すべきか?ITmedia AI+

    Gartnerは、2025年に米国の州当局が科したプライバシー法違反の罰金総額が34億2500万ドル(約5380億円)に達したと発表。過去…

  6. 「ChatGPT広告」日本上陸 無料版と「Go」で表示、電通・博報堂など支援ITmedia AI+

    OpenAIは2月に米国でテスト運用を開始。5月には日本など5カ国への拡大を予告していた。

  7. 米大企業の7割が導入する「Databricks」とは何者か? 評価額20兆円の「AI向けデータ基盤」ITmedia AI+

    評価額約20兆円、Fortune 500の7割が利用するデータ・AI基盤企業Databricks。オープンソースのビッグデータ分散処理エン…

トピック別件数

日本語メディア12件

ITmedia AI+ (日本語)

13:00 JST規制/政策

Gartnerが警鐘 プライバシー法執行が本格化、CISOは何を見直すべきか?

Gartnerは、2025年に米国の州当局が科したプライバシー法違反の罰金総額が34億2500万ドル(約5380億円)に達したと発表。過去5年間の合計を上回り、執行強化を背景に2028年まで加速する見通しを示した。

11:26 JSTLLM/生成AIOpenAIGPT / ChatGPT

ChatGPTで広告テスト、日本でも開始 非表示にする方法は?

米OpenAIの日本法人は、ChatGPTでの広告表示テストを日本でも始めたと発表した。広告を非表示にする方法は?

09:23 JSTLLM/生成AIOpenAIGPT / ChatGPT

「ChatGPT広告」日本上陸 無料版と「Go」で表示、電通・博報堂など支援

OpenAIは2月に米国でテスト運用を開始。5月には日本など5カ国への拡大を予告していた。

08:00 JSTその他

米大企業の7割が導入する「Databricks」とは何者か? 評価額20兆円の「AI向けデータ基盤」

評価額約20兆円、Fortune 500の7割が利用するデータ・AI基盤企業Databricks。オープンソースのビッグデータ分散処理エンジン「Apache Spark」開発者らが2013年に創業した同社の軌跡と最新情報を解説する。

08:00 JSTビジネス/資金調達

融資の決め手、決算書→「データ」「未来のシナリオ」へ 中小企業が資金調達に成功するための最大のポイントは?

中小企業の資金調達の在り方が、大きく変わろうとしている。融資特化型デジタルバンクである01(ゼロワン)銀行(大阪府吹田市)の大塚篤史副社長と北國銀行(金沢市)の竹内均氏(常務執行役員マーケティング部長)が、中小企業の経営や資金調達がどう変わっていくかの見解を語った。

07:00 JSTエージェント

工数「76%」削減 味の素グループが「経理AIエージェント」導入で先陣を切れたワケ

経理人材の不足が深刻化する一方で、経理パーソンが担う業務の幅は急速に広がっている。その解決策として期待されるのがAI活用だ。しかし、誤りが許されない経理業務では導入への慎重論も根強い。そんな中、味の素グループの財務・経理業務を担う味の素フィナンシャル・ソリューションズは、経費精…

07:00 JSTエージェント

「待ちの営業」はもう限界 ホンダがAIエージェントで挑む、商機を逃さない「濃い商談」の創出

顧客の購買行動が変化する中、ホンダが新車販売にAIエージェントを導入した。“濃い商談”を支援し、すでに成約も生まれているという。販売現場の変化を追う。

07:00 JSTその他

AIで要らなくなったSaaS、要るSaaSは、どれ? 日本の「SaaS is dead」の実態

エイトレッドが「AI時代に生き残るSaaSの条件に関する実態調査」の結果を公表。8割がSaaS見直しの必要性を実感する一方、AI代替の困難さや導入失敗の要因などが示された。

07:00 JSTその他

高級セレクトショップ「バーニーズ」が新品と中古の二刀流 富裕層の「初めての中古購入」を狙うワケ

原材料高騰や為替の乱高下に苦しむアパレル業界で、高級セレクトショップの雄「バーニーズ」が下した決断は、高級リユース市場への本格参入だった。「街の中古店には行かない」という富裕層の心理を突き、新品とユーズドを同じフロアで融合させる「二刀流」戦略の全貌とは。既存ビジネスとのカニバリ…

06:00 JSTLLM/生成AIエージェントClaude

話題の「Claude Mythos」登場で変わるセキュリティ AIエージェント時代の防衛策

AIによる攻撃が、月単位から時間単位で現実化する可能性が高まっている。最新AIモデルの登場で脆弱性発見の能力が広がる一方、企業ではAI利用のルールや管理体制が追い付いていない。AIエージェント時代に求められる新たな防衛策とは何か。

19:39 JSTLLM/生成AIAnthropic

チームみらい安野氏「牧歌的なAI開発の時代が終わった」 “ミュトス停止騒動”受け

「牧歌的なAI開発の時代が終わった」――チームみらいの安野貴博党首は6月18日の会見で、米AnthropicのAIモデル「Mythos 5」「Fable 5」の提供停止を巡る騒動を受け、このように述べた。

海外メディア11件

TechCrunch AI (英語)

09:51 JSTその他

Source: Elastic agrees to buy CRV-backed DeductiveAI for up to $85M

DeductiveAI, a startup that uses AI to catch and resolve bugs in software, was founded just three years ago.

06:20 JSTその他

AI inference startup Baseten reportedly raising $1.5B months after its last mega-round

Startup Baseten is reportedly close to finalizing a $1.5 billion round at a $13 billion as the “inference gold rush" marches on.

05:30 JSTその他

Snap spins off AI video team into new company, Dotmo, due to costs

The Snapchat maker is spinning off yet another internal unit. Dotmo will be composed of current Snap staff who are leaving the social media…

03:51 JSTその他

Almost half of US singles feel negatively about AI in dating, Match says

About 47% of singles look negatively at the use of AI in dating -- but many dating app users are open to AI helping with profile punch-ups…

03:22 JSTハードウェア/半導体NVIDIA

Amazon hopes to challenge Nvidia more directly by selling its AI chips

AWS is in talks to sell its chips to other data centers. CEO Andy Jassy has said this represents a $50 billion opportunity for the company.

02:49 JSTその他

AI data centers just got a government-mandated fast lane to the grid

FERC told grid operators to give data centers a fast lane for interconnections, but it failed to address electricity supply shortages.

02:16 JSTその他

The smartphone era created an attention crisis — slow tech is fixing it

“People just really want to take back control of their time, their lives, their attention... They’re down for whatever helps them do that.”

01:55 JSTその他

‘Queer Eye’ life coach Karamo Brown launches Kē, a wellness app featuring his AI digital clone

After spending a year and a half focusing on his own journey — from fitness and nutrition to meditation, sobriety, relationships, and perso…

00:20 JSTビジネス/資金調達

General Intuition in talks to raise $300M at around $2B valuation

The startup trains embodied AI and world models using Medal’s dataset of 2 billion videos per year from 10 million monthly active users.

00:13 JSTLLM/生成AIビジネス/資金調達その他2件の関連記事

A tech worker-backed PAC is bringing a $5M knife to Big Tech’s $100M gunfight

Guardrails positions itself as a populist political movement that runs on small donations from people in the trenches of the AI boom.

出典:TechCrunch AITechCrunch AI
21:00 JSTその他

Pixi’s new iOS app turns text messages into interactive AR experiences

Forget stickers, GIFs, and emoji reactions. Pixi is betting that the next evolution of messaging is interactive augmented reality (AR).

公式ブログ3件

OpenAI (英語)

02:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

New usage analytics and updated spend controls for enterprises

OpenAI introduces new spend controls and usage analytics for ChatGPT Enterprise, helping organizations manage costs and scale AI with confi…

20:00 JSTLLM/生成AIGPT / ChatGPT

Improving health intelligence in ChatGPT

Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clearer communication,…

17:00 JSTLLM/生成AI研究/論文OpenAI

Using AI to help physicians diagnose rare genetic diseases affecting children

Researchers used an OpenAI reasoning model to help diagnose rare diseases, identifying 18 new diagnoses in previously unsolved cases.

論文312件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント

Agentic AI システムのランタイム ガバナンスのための Deontic ポリシー

Large Language Model (LLM) によって駆動される自律型エージェント AI システムは、新たな種類のセキュリティ、プライバシー、コンプライアンスの課題をもたらします。ツールを呼び出し、データを操作し、ソフトウェアをインストールし、組織の境界を越えてピア エージェントと連携できるエージェントは、認証とアクセス制御だけでなく、エンタープライズ ガバナンスの完全な構造によっても制約される必要があります。これには、エージェントがどのような行為を許可され、どのような行為が禁止されているか、特定のアクションの後に何をする義務があるか(CISO に通知するなど)、どのような条件下で永続的な義務が免除されるか、ポリシーが矛盾する場合にどのルールが優先されるかを指定することが含まれます。このガバナンスの問題は、現在のポリシー エンジンが提供するものを超えています。 XACML、Rego、Cedar などのシステムは、このガバナンス構造の許可/禁止サブセットのみに対応します。これらは、義務のライフサイクル管理、メタポリシーの競合解決、特定の状況で義務を免除する措置、およびヘルスケア、サイバーセキュリティ、データプライバシーなどのアプリケーションで一般的に見られるドメインクラス階層に対する存在論的推論を提供しません。私たちは、基本的な許可/禁止制約だけでなく、義務、調剤、ポリシーの矛盾解決、ポリシーの推論などの主要なガバナンス要件を実現する AgenticRei を提案します。私たちは Rei フレームワーク上に構築された deontic ポリシー言語を使用しており、OWL (Web Ontology Language) として表現され、LLM の完全に外部にある高性能ロジック エンジンによって実行時に評価されます。同じパイプラインが、エージェントによるツールの呼び出しとエージェント間のメッセージの両方を制御します。例を通して、デオンティック ポリシーが、現在の運用エンジンではほとんど表現できないセキュリティとプライバシーに関するガバナンスの制約を捉えていることを示します。私たちのアプローチは、A2AS などの業界標準のフレームワークと自然に組み合わされています。

原文 (English)

Deontic Policies for Runtime Governance of Agentic AI Systems

Autonomous agentic AI systems driven by Large Language Models (LLMs) introduce a new class of security, privacy, and compliance challenges: an agent that can invoke tools, manipulate data, install software, and coordinate with peer agents across organizational boundaries must be constrained not just by authentication and access control, but by the full structure of enterprise governance. This includes specifying what agents are permitted and prohibited from doing, what they areobliged to do after certain actions (e.g., notify the CISO), under what conditions a standing obligation may be waived, and which rules take precedence when policies conflict. This governance problem exceeds what current policy engines provide. Systems such as XACML, Rego, and Cedar address only the permit/prohibit subset of this governance structure. They do not provide obligation lifecycle management, meta-policy conflict resolution, dispensations that waive obligations in specific circumstances, and ontological reasoning over domain class hierarchies commonly found in applications such as healthcare, cybersecurity, or data privacy. We propose AgenticRei, which realizes key governance requirements such as obligations, dispensations, policy conflict resolutions, and reasoning over policies, as well as the basic permit/prohibit constraints. We use a deontic policy language built on the Rei framework, expressed as OWL (Web Ontology Language) and evaluated at runtime by a high-performance logic engine entirely outside the LLM. The same pipeline governs both tool invocations by the agent and agent-to-agent messages. We show through examples that deontic policies capture governance constraints around security and privacy that mostly cannot be expressed in current production engines. Our approach composes naturally with industry-standard frameworks like A2AS.

13:00 JST研究/論文

トピックの範囲、コンピテンシー、認知深度にわたるカリキュラムの整合性の測定: CS2013 と CS2023 に適用される長期的なフレームワーク

学部のコンピュータ サイエンスは、約 10 年に 1 回改訂される国際カリキュラム ガイドラインによって管理されていますが、プログラムには、現在のガイドラインをどの程度完全にカバーしているか、またガイドラインが再構築されたときにそのカバー範囲がどのように変化するかを測定する、信頼性が高く再現可能な方法がありません。私たちは、コンピュータ サイエンス カリキュラム 2013 (CS2013) および 2023 (CS2023) に対して、コンピュータ サイエンスの認定学士 1 名に長期的に適用される外部知識体系のプログラムの範囲を測定する人間参加パイプラインでこれに対処します。パイプラインは、プログラムと各ガイドラインを構造化されたコーパスとして表し、意味検索によってコースと知識単位の一致候補を生成し、明示的なカバレッジ定義に基づいて人間の判断によってそれらを確認します。ベンチマークされた 7 つのレトリーバーのうち、相互ランク融合アンサンブルが最も強く、評判の高いロングコンテキスト モデルは短い文モデルのパフォーマンスを下回っていました。そのため、レトリーバーの選択は評価する必要があります。両方のマップは独立した第 2 評価者によって検証されました (CS2023 のコーエンのカッパ 0.64、CS2013 の 0.69)。このプログラムは、CS2023 の 49.7% と CS2013 の知識単位の 50.9% をカバーしており、これは 10 年間にわたってほぼ一定です。同じ検索してから確認するデザインをコンピテンシーの明確化と認知深度まで拡張すると、プログラムは各ガイドラインで対象となる単元の約 88% についてコンピテンシーを明確にし、さらに CS2013 では 95% であるのに対し、CS2023 では現在の単元の 76% に対して推奨深度でコンピテンシーを提供していることがわかります。このギャップはプログラムではなく、新しいガイドラインで高められた期待を反映しています。縦断的な比較により、ガイドラインと ABET の両方に対して明らかになった永続的な構造的ギャップ (並列コンピューティングと分散コンピューティング、プログラミング言語の基礎、システムの基礎) が、標準の進化を反映する相違点から分離されます。この機器は再利用可能であり、リクエストに応じて著者から入手できます。

原文 (English)

Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023

Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproducible way to measure how completely they cover the current guidelines and how that coverage shifts when the guidelines are restructured. We address this with a human-in-the-loop pipeline that measures a program's coverage of an external body of knowledge, applied longitudinally to one accredited BSc in Computer Science against Computer Science Curricula 2013 (CS2013) and 2023 (CS2023). The pipeline represents the program and each guideline as structured corpora, generates candidate course-to-knowledge-unit matches by semantic retrieval, and confirms them through human judgment under an explicit coverage definition. Of seven benchmarked retrievers, a reciprocal-rank-fusion ensemble was strongest, and a reputed long-context model underperformed a small sentence model, so retriever choice must be measured. Both maps were validated by an independent second rater (Cohen's kappa 0.64 for CS2023, 0.69 for CS2013). The program covers 49.7% of CS2023 and 50.9% of CS2013 knowledge units, near-constant across a decade. Extending the same retrieve-then-confirm design to competency articulation and cognitive depth shows that the program articulates the competency for ~88% of covered units under each guideline, yet delivers it at the recommended depth for 76% of present units under CS2023 against 95% under CS2013, a gap reflecting the newer guideline's raised expectations, not the program. The longitudinal comparison separates persistent structural gaps (parallel and distributed computing, foundations of programming languages, systems fundamentals), uncovered against both guidelines and ABET, from differences that reflect the standard's evolution. The instrument is reusable and available from the authors on request.

13:00 JSTLLM/生成AI

拡散言語モデル: 実験的分析

大規模言語モデル (LLM) は、自己回帰生成を通じて言語モデリングに革命をもたらし、幅広いタスクにわたって強力なパフォーマンスを可能にします。最近、拡散言語モデル (DLM) が、次のトークンの予測ではなく反復的なノイズ除去を通じてテキストを生成する代替パラダイムとして登場し、シーケンス全体の並行改良を可能にします。多数の拡散ベースのアーキテクチャが提案されていますが、評価プロトコル、データセット、推論バジェット、生成ハイパーパラメータの違いにより、それらの機能を比較し、それらがもたらすトレードオフを理解することが困難になっています。この研究では、最新の DLM の系統的な実験分析を紹介します。具体的には、生成品質と計算効率の両方を明示的に考慮しながら、推論、コーディング、翻訳、知識、構造化された問題解決にわたる 8 つのベンチマークにわたって 8 つの最先端 DLM を評価します。下流の評価を超えて、ノイズ除去ステップ、コンテキストの長さ、ブロック サイズ、並列アンマスク戦略などの主要な推論時間要因の影響を分析し、同一条件下でトレーニングされた小規模なモデルの制御された比較で大規模な実験を補完します。私たちの分析では、さまざまなタスク、アーキテクチャ、推論予算にわたる拡散ベースの言語モデリングの長所と限界が浮き彫りになっています。 DLM の動作は生成時の設計選択によって強く影響され、パフォーマンスと計算効率の間に明確なトレードオフが生じることを示します。全体として、私たちの研究は、現代の DLM の機能と展開の特徴についての実用的な洞察を提供します。

原文 (English)

Diffusion Language Models: An Experimental Analysis

Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range of tasks. Recently, Diffusion Language Models (DLMs) have emerged as an alternative paradigm that generates text through iterative denoising rather than next-token prediction, allowing parallel refinement of entire sequences. While numerous diffusion-based architectures have been proposed, differences in evaluation protocols, datasets, inference budgets, and generation hyperparameters make it difficult to compare their capabilities and understand the trade-offs they offer. In this work, we present a systematic experimental analysis of modern DLMs. Specifically, we evaluate eight state-of-the-art DLMs across eight benchmarks spanning reasoning, coding, translation, knowledge, and structured problem solving, while explicitly considering both generation quality and computational efficiency. Beyond downstream evaluation, we analyze the impact of key inference-time factors, including denoising steps, context length, block size, and parallel unmasking strategies, and complement large-scale experiments with controlled comparisons of smaller models trained under identical conditions. Our analysis highlights the strengths and limitations of diffusion-based language modeling across different tasks, architectures, and inference budgets. We show that the behavior of DLMs is strongly influenced by generation-time design choices, leading to distinct trade-offs between performance and computational efficiency. Overall, our study provides practical insights into the capabilities and deployment characteristics of contemporary DLMs.

13:00 JSTLLM/生成AIエージェント

マルチエージェント LLM の審議における隠れたアンカー

複数のラウンドにわたってエージェントが回答を交換および修正するマルチエージェント LLM 熟議は、推論と精度を向上させるためにますます使用されていますが、それがどのように、そしてなぜ機能するのかモデル化されることはほとんどありません。このような熟慮は、人間がどのように決定に至るかを反映しています。社会的動物として、私たちは、デグルートやフリードキン・ジョンセンのような古典的な意見力学モデルが捉えている集団効果である集団と、彼らには理解されていない私たち自身の内なる信念の両方に引っ張られています。私たちは、マルチエージェントの熟議を閉ループの力学システムとしてモデル化します。このシステムでは、各エージェントが、近隣のエージェントに関係なく常に自分の意見を引き出す、隠れた内なる信念、つまりそのアンカーを担っています。我々は、このアンカーが熟議のみから回復できること、そしてそれが古典的なコンセンサスのルールが禁じている行動を説明していることを示します。つまり、正解に対するエージェントの信頼は、エージェントが開始した時点を超えて、最初の信念によって形成された空間(凸包)から逃れることができます。回復されたアンカーがホールドアウト実行を予測する (一般化する) かどうかを確認することで、モデルがそのようなアンカーによって実際に駆動されるときの簡単なテストが得られます。 3 つのオープンウェイト モデル ファミリーにわたって、これはスペクトルであり、すべてかゼロかではありません。すべてのアンカーの影響力はほぼ同等に強いですが、アンカーがどこに位置するかが異なり、最初の意見から遠く離れた位置にある場合にのみ、検討が船体から逃れ、完全な閉ループ モデルが必要になります。

原文 (English)

Hidden Anchors in Multi-Agent LLM Deliberation

Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increasingly used to improve reasoning and accuracy, yet how and why it works is rarely modelled. Such deliberation mirrors how humans reach decisions. As social animals we are pulled both by the group, the herd effect that classical opinion-dynamics models such as DeGroot and Friedkin--Johnsen capture, and by our own internal belief, which they do not. We model multi-agent deliberation as a closed-loop dynamical system in which each agent carries a hidden internal belief, its anchor, that continually pulls its opinion regardless of its neighbours. We show this anchor can be recovered from the deliberation alone, and that it explains a behaviour classical consensus rules forbid: an agent's confidence in the correct answer can climb past where any agent started, escaping the space (convexhull) formed by the initial beliefs. Checking whether the recovered anchor also predicts held-out runs (generalizes) gives a simple test for when a model is truly driven bysuch an anchor. Across three open-weight model families this is a spectrum, not all-or-nothing. All anchors' influence are about equally strongly, but they differ in where the anchor sits, and only when it sits far from the initial opinions does deliberation escape the hull and need the full closed-loop model.

13:00 JSTLLM/生成AIエージェント

DeXposure-Claw: DeFi リスク監視のためのエージェント システム

分散型金融により、監督当局は急速に変化するネットワーク化された信用リスクにさらされます。汎用 LLM エージェントはこの設定にあまり適合しません。弱い証拠を深読みし、一か八かの介入を推奨しますが、既存の評価では、結果として生じる誤報を測定するための規制当局と連携した方法が提供されていません。我々は、構造化された証拠を通じて LLM 決定をルーティングする、予測に基づいたエージェント監視システムである DeXposure-Claw を紹介します。(1) DeXposure-FM は、グラフ時系列基礎モデルであり、将来のエクスポージャ ネットワークを予測します。 (2) 決定論的なモニターとストレス シナリオは、それらの予測を型指定されたアラート、属性シグナル、およびシナリオの証拠に変換します。 (3) データの健全性と信頼ゲートにより、DeXposure-Claw が根拠のある監査可能な監督チケットを発行する前にエスカレーションが抑制されます。さらに、6 軸の評価ハーネスである DeXposure-Bench を開発します。このベンチの決定軸は、規制当局と調整された絶対損失グラウンド トゥルースおよび明示的な誤介入率に対してチケットをスコア付けします。 5 年間の毎週の実データを用いた実験により、当社のシステムが完全にサポートされています。コードは https://github.com/EVIEHub/DeXposure-Claw にあります。

原文 (English)

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no regulator-aligned way to measure the resulting false alarms. We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that routes LLM decisions through structured evidence: (1) DeXposure-FM, a graph time-series foundation model, forecasts future exposure networks; (2) deterministic monitors and stress scenarios then turn those forecasts into typed alerts, attribution signals, and scenario evidence; and (3) data-health and confidence gates constrain escalation before DeXposure-Claw emits auditable supervisory tickets with rationales. We further develop DeXposure-Bench, a six-axis evaluation harness, whose decision axis scores tickets against a regulator-aligned absolute-loss ground truth and an explicit false-intervention rate. Experiments on five years of weekly real data fully support our system. Code is at https://github.com/EVIEHub/DeXposure-Claw.

13:00 JSTLLM/生成AIQwen

LLM は何を知らないのかを知らない: 臨床表データのモデル間の帰属相違による認識上の盲点の検出

大規模言語モデル (LLM) は、構造化された臨床データにますます適用されていますが、そのようなタスクに関する自身の知識の限界を認識できるかどうかはまだ解明されていません。私たちは、構造化タスクの認識論的不確実性を軽減することを目的として、クロスモデルの属性発散のレンズを通してこの問題を研究し、属性発散分析による予測タスクで Qwen 2.5 7B と XGBoost を比較します。 4 つの調査結果を報告します。まず、LLM 言語化された信頼度は認識論的に空虚であり、精度が 49% であるか 75.3% であるかに関係なく、ほぼ一定 (0.856 ~ 0.937) を出力し、予測品質ではなくプロンプト形式を追跡します。第 2 に、LLM は逆の難易度効果を示します。XGBoost が 99% 正しい場合、精度は 64.8% に低下しますが、中程度の不確実性がある場合は XGBoost と一致します (73.8% 対 73.1%)。 3 番目に、少数ショットの例と SHAP 由来の特徴証拠は、直交する超相加的介入です。トレーニングなしで、帰属不一致スコア (ADS) が 1.54 から 0.38 に減少し、精度が 49% から 75.3% に向上します。 4 番目に、帰属乖離信号を使用して LLM の信頼性を決定するクロスモデル キャリブレーターは、モデルの内部にアクセスしたり推論を繰り返す必要がなく、情報のない言語化された信頼性を患者固有の信頼性推定値に置き換えて、予想されるキャリブレーション誤差を 0.254 から 0.080 に削減します。私たちはこれらの発見を構造化データ上の LLM のコールド スタート問題として枠組み化し、真の認識論的自己認識への道筋を概説します。

原文 (English)

LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data

Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they can recognize the limits of their own knowledge on such tasks remains unexplored. We study this question through the lens of cross-model attribution divergence with the goal of reducing epistemic uncertainty for structured tasks, comparing Qwen 2.5 7B and XGBoost on a prediction task via attribution divergence analysis. We report four findings. First, LLM verbalized confidence is epistemically vacuous, it outputs a near-constant (0.856-0.937) regardless of whether accuracy is 49% or 75.3%, tracking prompt format rather than prediction quality. Second, the LLM exhibits an inverse difficulty effect: accuracy drops to 64.8% when XGBoost is 99% correct, but matches XGBoost (73.8% vs. 73.1%) when it is moderately uncertain. Third, few-shot examples and SHAP-derived feature evidence are orthogonal, super-additive interventions: they reduce the Attribution Disagreement Score (ADS) from 1.54 to 0.38 and improve accuracy from 49% to 75.3% without training. Fourth, a cross-model calibrator that determined LLM reliability using attribution divergence signals reduces expected calibration error from 0.254 to 0.080, replacing uninformative verbalized confidence with patient-specific reliability estimates, without accessing model internals or requiring repeated inference. We frame these findings as a cold start problem for LLMs on structured data and outline a path toward genuine epistemic self-awareness.

13:00 JST研究/論文

REVEAL++: アルツハイマー病リスクの視覚言語網膜モデリングのための識別可能な表現型グループ化

網膜は神経変性疾患への非侵襲的な窓を提供し、将来の認知機能低下のリスクに関連する微妙な構造パターンを捕捉します。 REVEAL などの視覚言語調整フレームワークは、網膜眼底画像と構造化された臨床リスクのナラティブを組み合わせることで、アルツハイマー病 (AD) の早期予測が向上することを示しています。これらのアプローチにおける重要な設計上の選択は、表現型グループ化の使用であり、類似したリスク プロファイルを持つ個人が、対比学習中に多重陽性ペアとして扱われます。しかし、既存の方法は、表現型の類似性を離散的な構成要素として操作し、厳密な監視を課し、グループ形成を表現学習から分離するハードグループ割り当てに依存しています。我々は、対照学習内で表現型構造を継続的に定式化することを提案します。サンプルを固定クラスターに割り当てるのではなく、網膜画像とリスクプロファイルの両方に埋め込まれたモダリティ内の類似性から導出された微分可能な重み付け関数として被験者間の類似性をモデル化します。これらの重みは、継続的な集計演算子を通じてソフトなマルチポジティブ関係を定義し、疾患リスクのスペクトルの性質を反映した段階的な監視を可能にします。さらに、クロスモーダルアライメントと表現型構造をエンドツーエンドで共同学習するソフトターゲット対比目標を導入します。アルツハイマー病の予測について英国バイオバンクの網膜画像データに基づいて評価したところ、提案されたフレームワークは、離散グループベースの対比学習および標準的な視覚言語ベースラインを一貫して上回っています。表現型の類似性を、固定されたグループ化ルールではなく、学習可能な連続信号として扱うことにより、私たちのアプローチは、マルチモーダルな網膜および臨床データから集団スケールの神経変性リスクモデリングのための原則に基づいた堅牢な基盤を提供します。

原文 (English)

REVEAL++: Differentiable Phenotypic Grouping for Vision-Language Retinal Modeling of Alzheimer's Disease Risk

The retina offers a noninvasive window into neurodegenerative disease, capturing subtle structural patterns associated with a risk of future cognitive decline. Vision-language alignment frameworks such as REVEAL have shown that pairing retinal fundus images with structured clinical risk narratives improves early prediction of Alzheimer's disease (AD). A key design choice in these approaches is the use of phenotypic grouping, where individuals with similar risk profiles are treated as multi-positive pairs during contrastive learning. However, existing methods operationalize phenotypic similarity as a discrete construct, relying on hard group assignments that impose rigid supervision and decouple group formation from representation learning. We propose a continuous formulation of phenotypic structure within contrastive learning. Rather than assigning samples to fixed clusters, we model inter-subject similarity as a differentiable weighting function derived from intra-modality embedding similarities in both retinal images and risk profiles. These weights define soft multi-positive relationships through a continuous aggregation operator, enabling graded supervision that reflects the spectrum nature of disease risk. We further introduce a soft-target contrastive objective that jointly learns cross-modal alignment and phenotypic structure in an end-to-end manner. Evaluated on UK Biobank retinal imaging data for incident AD prediction, the proposed framework consistently outperforms discrete group-based contrastive learning and standard vision-language baselines. By treating phenotypic similarity as a learnable, continuous signal rather than a fixed grouping rule, our approach provides a principled and robust foundation for population-scale neurodegenerative risk modeling from multi-modal retinal and clinical data.

13:00 JSTLLM/生成AIハードウェア/半導体

緊急調整

大規模言語モデル (LLM) は、自身の出力が人間の倫理と乖離していることを識別できますか?そして彼らは自己修正できるのでしょうか? LLM に、自身の推論と出力をレビューする良心ステップを与え、直接優先最適化 (DPO) を使用して調整コンポーネントでトレーニング損失を拡張し、モデルを非倫理的な出力から遠ざけます。その結果、トレーニング、微調整、敵対的プロンプト、ゼロショット学習など、幅広いアプリケーションでモデルを調整するオンライン技術が実現しました。それは、より弱いまたはより強いジャッジを必要とせず、代わりにそれ自体の凍結されたコピーに依存します。以前の研究では、緊急不整合シナリオでは、モデルの微調整からコードのハッキングに至るまで、さまざまな緊急の非倫理的な行為が示されました。代わりに、私たちは緊急調整を達成する方法を経験的に示します。つまり、単一の高レベルの内省的な質問が、同じコード ハッキング シナリオの下で倫理モデルに向けてトレーニングを導きます。

原文 (English)

Emergent Alignment

Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics? And can they self-correct? We endow an LLM with a conscience step that reviews its own reasoning and outputs, and we extend the training loss with an alignment component using Direct Preference Optimization (DPO) to steer the model away from non-ethical outputs. The result is an online technique to align models in a wide range of applications: training, fine-tuning, adversarial prompting, and zero-shot learning. It does not require a weaker or stronger judge, relying instead on a frozen copy of itself. In previous work, the Emergent Misalignment scenario showed a range of emergent unethical behaviors from fine-tuning the model to hack code. Instead, we empirically show how to achieve Emergent Alignment: a single high-level introspective question steers training toward an ethical model under the same code hacking scenario.

13:00 JST研究/論文

ITNet: 畳み込み、注意、再帰を包含する学習可能な積分変換

畳み込みネットワーク、リカレント ネットワーク、トランスフォーマーはそれぞれ、局所性、逐次記憶、内容依存のペアワイズ相互作用など、さまざまな誘導バイアスをエンコードしており、その誕生以来数学的に区別され続けています。我々は、この断片化が信号の処理方法における基本的な多様性を反映しているのではなく、むしろ基礎となる単一の数学的オブジェクト、つまり学習可能な積分変換の不完全なビューを反映していることを示します。位置と特徴に共同して依存する学習可能なカーネルを中心に構築された統合アーキテクチャである Integral Transform Network (ITNet) を紹介します。このカーネルは、ペアごとの相互作用をモデル化する小さなニューラル ネットワーク、特に MLP として実装され、モデルがデータからその動作を適応できるようにします。畳み込み、自己注意 (マルチヘッドを含む)、および自己回帰再帰 (LSTM、GRU、S4、および Mamba を含む) が適切なパラメータ化の下で特殊なケースとして発生すること、および ITNet が連続演算子の汎用近似器であることを示します。これを実用化するために、タイル化カーネル融合、重要度加重モンテカルロ統合、学習された低ランク因数分解を開発し、効率的でスケーラブルな計算を可能にします。共有オペレーターと軽量のモダリティ固有のエンコーダーを備えた単一の ITNet アーキテクチャは、ImageNet-1K、GLUE、ModelNet40、VQA\,v2、および NLVR2 の特殊なベースラインと一致またはそれを超えています。この結果は、単一の学習された対話メカニズムが 3 つのアーキテクチャ ファミリすべての動作をデータから復元できることを示しています。

原文 (English)

ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence

Convolutional networks, recurrent networks, and transformers each encode different inductive biases -- locality, sequential memory, and content-dependent pairwise interaction -- and have remained mathematically distinct since their inception. We show that this fragmentation reflects not a fundamental diversity in how signals should be processed, but rather incomplete views of a single underlying mathematical object: a learnable integral transform. We introduce the Integral Transform Network (ITNet), a unified architecture built around a learnable kernel that depends jointly on positions and features. This kernel is implemented as a small neural network, specifically an MLP, that models pairwise interactions, enabling the model to adapt its behavior from data. We show that convolution, self-attention (including multi-head), and autoregressive recurrence (including LSTM, GRU, S4, and Mamba) arise as special cases under appropriate parameterizations, and that ITNet is a universal approximator of continuous operators. To make this practical, we develop tiled kernel fusion, importance-weighted Monte Carlo integration, and learned low-rank factorization, enabling efficient and scalable computation. A single ITNet architecture with a shared operator and lightweight modality-specific encoders matches or exceeds specialized baselines on ImageNet-1K , GLUE, ModelNet40, VQA\,v2 and NLVR2. The results demonstrate that a single learned interaction mechanism can recover the behavior of all three architectural families from data.

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPTDeepSeek

LLM エージェントにおける明確化探索のための不確実性分解

最近の意見書では、対話型大言語モデル (LLM) エージェントには古典的な偶発的/認識論的不確実性フレームワークでは不十分であると主張し、プロアクティブな説明の探索や共有メンタル モデルの構築などの新しいエージェントの能力を解放できる、過小仕様を認識し、分解され、伝達可能な不確実性表現が求められています。ブラックボックス API、インタラクティブなレイテンシ バジェット、ラベル付き軌跡の欠如といった実際的なデプロイメントの制約により、logprob ベース、マルチサンプリング、トレーニング ベースの手法が除外され、プロンプトベースの推定がデプロイメント時にそのような信号を表面化するための最も実行可能なファミリーとして残ります。この呼び出しには、アクションの信頼性とリクエストの不確実性 (u) を分離する単純なプロンプトベースの分解で応答します。これにより、タスクの仕様があいまいな場合にエージェントが説明を求めることができます。これを評価するために、タスクの 50% が意図的に過少指定されている 2 つの明確化強化ベンチマーク (WebShop-Clarification および ALFWorld-Clarification) を導入し、5 つの LLM バックボーン (GPT-5.1、DeepSeek-v3.2-exp、GLM-4.7、 Qwen3.5-35B、GPT-OSS-120B) をこれらのバリアントと、障害検出のための標準の WebShop、ALFWorld、および REAL ベンチマークとともに使用します。 5 つのバックボーン全体で平均すると、提案された分解は、ALFWorld-Clarification のクラリフィケーション F1 を ReAct+UE と比較して 73%、UAM と比較して 36% 改善し、WebShop-Clarification のすべてのバックボーンと ALFWorld-Clarification の 5 つのバックボーンのうち 4 つでクラリフィケーション F1 をリードしており、ゲインが単一の LLM を超えて一般化していることを示しています。

原文 (English)

Uncertainty Decomposition for Clarification Seeking in LLM Agents

Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) agents and call for underspecification-aware, decomposed, and communicable uncertainty representations that can unlock new agent capabilities such as proactive clarification seeking and shared mental-model building. Practical deployment constraints -- black-box APIs, interactive latency budgets, and the absence of labeled trajectories -- rule out logprob-based, multi-sampling, and training-based methods, leaving prompt-based estimation as the most viable family for surfacing such signals at deployment time. We answer this call with a simple prompt-based decomposition that separates action confidence from request uncertainty (u), enabling the agent to ask for clarification when the task specification is ambiguous. To evaluate it, we introduce two clarification-augmented benchmarks (WebShop-Clarification and ALFWorld-Clarification) in which 50% of tasks are deliberately underspecified, and systematically compare the proposed decomposition against ReAct+UE and Uncertainty-Aware Memory (UAM) across five LLM backbones (GPT-5.1, DeepSeek-v3.2-exp, GLM-4.7, Qwen3.5-35B, GPT-OSS-120B) on these variants together with the standard WebShop, ALFWorld, and REAL benchmarks for fault detection. Averaged across the five backbones, the proposed decomposition improves clarification F1 on ALFWorld-Clarification by 73% over ReAct+UE and by 36% over UAM, and leads clarification F1 on every backbone on WebShop-Clarification and on four of five backbones on ALFWorld-Clarification, indicating that the gains generalize beyond a single LLM.

13:00 JSTLLM/生成AI

LLM ソルバー ループにおけるナレーション ギャップの分析

安全性やセキュリティに関する重要な問題をロジックで定式化できる場合、SAT ソルバーや SMT ソルバーなどの形式的なツールが言語モデル推論パイプラインに組み込まれることが増えています。正式な保証なしにモデル分布からステップがサンプリングされる思考連鎖とは異なり、ソルバーは健全で独立して検証可能な答えを生成します。ただし、ソルバーとモデルの間の相互作用によって健全性の保証が失われる可能性があります。ハイブリッド パイプラインには、質問の形式化、決定、結果の説明という 3 つのコンポーネントがあります。これまでの研究では、形式化と決定については研究されてきましたが、形式的なツールの出力をユーザーの回答に変えるステップであるナレーションについては研究されていませんでした。ナレーションのギャップを埋めるために、まず LLM ソルバー ループを検証済みの意思決定手順としてモデル化します。さらに、プロンプト インジェクションの下で 5 つのオープンソース モデルを評価したところ、証明書ゲーティングによってソルバーの判定が確実なものになる一方、敵対者がフレージングやチャネル全体で検証済みの結論を覆す可能性があることがわかりました。私たちは、インジェクションを大幅に削減するものの、インジェクションを排除することはできず、依然として適応型攻撃を受ける、強化されたプロンプトによる緩和策を研究しています。形式的な分析と実証的研究を組み合わせると、LLM ソルバー ループでのロバスト性は、ユーザーが最終的に読み取る回答に到達しないことがわかります。

原文 (English)

Analyzing the Narration Gap in LLM-Solver Loops

Formal tools such as SAT and SMT solvers are increasingly embedded in language model reasoning pipelines when a safety or security critical question can be formulated in logic. Unlike chain of thought whose steps are sampled from the model distribution without formal guarantee, a solver produces a sound and independently verifiable answer. However, the soundness guarantee can be lost in the interaction between the solver and the model. The hybrid pipeline has three components: formalizing the question, deciding it, and narrating the result. Prior work has studied the formalization and decision, but not narration, which is the step that turns a formal tool's output into the user answer. To fill the narration gap, we first model the LLM-solver loop as a verified decision procedure. We further evaluate five open-sourced models under prompt injection, and we find certificate gating makes the solver verdict sound, while an adversary can invert a verified conclusion across phrasings and channels. We study the mitigation through hardened prompt that reduces injection significantly but cannot eliminate it and still suffers under adaptive attack. Combining the formal analysis and empirical studies, we show in the LLM-solver loop, robustness does not reach to the answer that the user finally reads.

13:00 JSTエージェント

Agentic RAG による構成可能な臨床情報抽出: 機能するもの、機能しないもの、およびその理由

患者のコンテキストは数百の異種ドキュメントと数千の構造化データポイントにまたがっていますが、AI システムが検索やトリアージに必要とするドキュメントレベルのメタデータが存在しないか不完全です。標準的な検索拡張生成は、このデータでは失敗し、一時的な推論、ドキュメント間の依存関係、およびメタデータの欠落の処理が誤ります。私たちは、エッセン医科大学に ACIE (薬剤臨床情報抽出) を導入しています。これは、完全な患者コンテキストを推論し、臨床医の検証のためにソースパッセージ内のすべての回答を根拠にするオンプレミスの薬剤 RAG パイプラインです。私たちはメタデータのギャップを定量化し、それが形作ったアーキテクチャ上の決定を追跡し、独立した遡及的リンパ腫レジストリ研究と並行して抽出を評価します。この研究では、核医学の医師が抽出されたすべての値を引用ソースと照合して検証します。 7,326 件の判定において、臨床医は 96.5% の抽出を受け入れ、タイプごとの受け入れ率は 80% から 99% の範囲でした。

原文 (English)

Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why

Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-level metadata that AI systems need for retrieval and triage is absent or incomplete. Standard retrieval-augmented generation fails on this data, mishandling temporal reasoning, cross-document dependencies, and missing metadata. We deploy ACIE (Agentic Clinical Information Extraction) at University Medicine Essen: an on-premise agentic RAG pipeline that reasons over complete patient contexts and grounds every answer in source passages for clinician verification. We quantify the metadata gap, trace the architectural decisions it shaped, and evaluate extraction alongside an independent retrospective lymphoma registry study, in which nuclear-medicine physicians verify every extracted value against its cited sources. Across 7,326 judgments, clinicians accepted 96.5\% of extractions, with per-type acceptance ranging from 80\% to 99\%.

13:00 JSTLLM/生成AI

LLM トレーニング後の比較対象となるペアはどれですか?

好みに基づいたポストトレーニングは、言語モデルを調整するための中心的なパラダイムとなっています。一般的なデータ収集戦略は、プロンプトごとに小さな補完セットを生成し、結果の比較ペアにラベルを付けることです。ただし、人間の好みのラベルは、追加の補完を生成するよりもはるかにコストがかかることが多いため、同じラベル付け予算の別の使用法、つまり、より大きな補完プールを生成し、最も有益な比較ペアのみにラベルを付けることを示唆しています。この論文では、好みに基づいたポストトレーニングでどのペアを比較する必要があるかを研究します。私たちは比較キュレーションをサンプリング設計問題として定式化し、好みに基づいたトレーニング後の目標に基づいて最終ポリシーの品質によって設計を評価します。このフレームワークを Direct Preference Optimization (DPO) 用にインスタンス化し、ラベル付きペアの選択が DPO トレーニングを通じて下流のポリシーのパフォーマンスにどのように伝播するかを分析します。私たちの主な結果は、DPO でトレーニングされたポリシーのトレーニング後の最適性ギャップの上限と下限の一致を示します。境界は、比較選択が、ラベル割り当てをパラメーター推定誤差およびポリシーの準最適性に結び付ける、単一の設計依存情報マトリックスを通じて下流のパフォーマンスに影響を与えることを示しています。これにより、予算に基づいた比較キュレーションのための明示的な最適化基準が得られ、生成された大規模な補完プールから有益なペアを選択するための実際的なサンプリング設計が動機付けられます。合成設定と言語モデルのトレーニング後のベンチマークに関する実験では、提案された設計が一般的な比較選択ヒューリスティックよりもサンプル効率を一貫して向上させることが示されています。

原文 (English)

Which Pairs to Compare for LLM Post-Training?

Preference-based post-training has become a central paradigm for aligning language models. A common data-collection strategy is to generate a small set of completions for each prompt and label the resulting comparison pairs. However, human preference labels are often much more expensive than generating additional completions, suggesting a different use of the same labeling budget: generate a larger pool of completions, but label only the most informative comparison pairs. This paper studies which pairs should be compared in preference-based post-training. We formulate comparison curation as a sampling-design problem and evaluate designs by the quality of the final policy under the preference-based post-training objective. We instantiate this framework for Direct Preference Optimization (DPO), analyzing how the choice of labeled pairs propagates through DPO training to downstream policy performance. Our main results provide matching upper and lower bounds on the post-training optimality gap of the DPO-trained policy. The bounds show that comparison selection affects downstream performance through a single design-dependent information matrix, which links label allocation to parameter estimation error and policy suboptimality. This yields an explicit optimization criterion for budgeted comparison curation and motivates practical sampling designs for selecting informative pairs from large generated completion pools. Experiments on synthetic settings and language-model post-training benchmarks show that the proposed designs consistently improve sample efficiency over common comparison-selection heuristics.

13:00 JSTLLM/生成AI

Toten: ブラジルポルトガル語での物理量と技術表記の知識ベースの存在論的トークン化

バイト ペア エンコーディングのトークン化は、語彙圧縮に関しては統計的に効率的ですが、意味的には構造化された技術的エンティティに対して盲目であり、物理量、数値、単位、記号表現を語彙的に任意のサブワードに断片化します。我々は、統計的導出をエンジニアリングエンティティ(OEE)の正式なオントロジーに基づいた宣言的分類に置き換える、知識ベースのオントロジートークン化フレームワークであるTOTENを紹介します。私たちは TOTEN をトリプルとして形式化します。オントロジーは型、構造原理、構成関係、保存可能な不変条件を収集します。分類関数は、生のテキストを入力された領域にマッピングします。そして、インスタンシエータ ファミリは自己記述的な構造化表現を生成します。堅牢性は、Pint (次元)、Unicode 文字データベース (タイポグラフィ)、および RSLP (ポルトガル語形態学) という 3 つの外部オラクルとの決定論的な結合から得られます。本質的評価は、物理的に検証された内部ベンチマーク (EngQuant、N=800) および 4 つのブラジルポルトガル語外部コーパス (N=1771 の適格なケース) を対象として、構築によって検証可能な 4 つのプロパティ (存在論的原子性、次元の等価性、タイポグラフィーの堅牢性、および数値再構成) をカバーします。また、検出リコールも報告し、カバレッジと条件付きアトミック性を区別します。 8 つの最先端のベースラインに対して、TOTEN は、すべてのコントラストおよび数値再構成において、外部コーパスでは 0.775 ~ 0.904 の単位オントロジー アトミック性を達成します。これに対し、最良のベースライン (Quantulum3) では 0.627 ~ 0.703 でした。 EngQuant では、0.780 対 0.340。差異は統計的に有意です (ホルム補正を伴うマクネマー)。内部ランキングと外部ランキングの間のスピアマン相関により、コントロール ベンチマークの同時有効性が確認されます。次元の同等性は、システムが次元の権限を継承する神託である Pint と統計的に同等であることを示します。

原文 (English)

Toten: Knowledge-Based Ontological Tokenization Of Physical Quantities And Technical Notation In Brazilian Portuguese

Byte-Pair Encoding tokenization is statistically efficient for vocabulary compression, but semantically blind to structured technical entities, fragmenting physical quantities, numbers, units, and symbolic expressions into lexically arbitrary subwords. We present TOTEN, a knowledge-based ontological tokenization framework that replaces statistical derivation with declarative classification grounded in a formal ontology of engineering entities (OEE). We formalize TOTEN as the triple : the ontology gathers types, structural principles, composition relations, and preservable invariants; the classification function maps raw text into typed regions; and the instantiator family yields a self-descriptive structured representation. Robustness derives from deterministic coupling with three external oracles: Pint (dimensional), Unicode Character Database (typographic), and RSLP (Portuguese morphology). Intrinsic evaluation covers four properties verifiable by construction -- ontological atomicity, dimensional equivalence, typographic robustness, and numerical reconstruction -- over an internal, physically validated benchmark (EngQuant, N=800) and four Brazilian Portuguese external corpora (N=1771 eligible cases). We also report detection recall, distinguishing coverage from conditional atomicity. Against eight state-of-the-art baselines, TOTEN achieves unit ontological atomicity in all contrasts and numerical reconstruction of 0.775-0.904 on external corpora, vs. 0.627-0.703 for the best baseline (Quantulum3); on EngQuant, 0.780 vs. 0.340. Differences are statistically significant (McNemar with Holm correction). Spearman correlation between internal and external rankings confirms concurrent validity of the control benchmark. Dimensional equivalence shows statistical parity with Pint, the oracle from which the system inherits dimensional authority.

13:00 JST研究/論文

AI4SE と SE4AI の探求: 過去と未来を振り返る 10 年

AI とシステム エンジニアリング (SE) に関する 2020 年 3 月の INCOSE INSIGHT 特集号は、同誌史上最もダウンロードされた号となり、研究コミュニティを立ち上げ、現在では年次ワークショップに 250 人を超える登録者が集まっています。この記事では、著者がこの分野の核となる論文を読んだことに基づいて、AI と SE の 3 つのフェーズ (ここでは基礎、応用、LLM の活用とラベル付けされています) にわたる進歩を追跡し、コミュニティがどこに収束し、どこに重大なギャップが残っているかについての意見を説明します。これとは別に、人間の専門知識と 6 つの AI モデルの両方を活用した人間と AI の合意文献レビューが実行され、1,712 件の INCOSE INSIGHT 記事と 889 件の SERC 出版物の関連性が評価されました。この結果は、研究における 5 つの重要なギャップを特定し、SE における AI の導入、保証、労働力の変革を進める実務者に指針を提供します。私たちは契約データと AI4SE/SE4AI Explorer Web アプリケーションを共有するので、読者は自分自身の関連性判断を人間および AI の評価者と比較できます。

原文 (English)

AI4SE and SE4AI Exploration: A Decade Looking Back and Forward

The March 2020 INCOSE INSIGHT special issue on AI and Systems Engineering (SE) became the most downloaded issue in the publication's history and launched a research community that now draws over 250 registrants to its annual workshop. In this article, we trace the progress in AI and SE across three phases (labeled here foundational, applied, and LLM inflection) based on the authors' reading of the field's core papers, and describe our opinions of where the community has converged and where critical gaps remain. Separately, a human-AI agreement literature review leveraging both human expertise and six AI models was performed to assess the relevance of 1,712 INCOSE INSIGHT articles and 889 SERC publications. The results identify five critical research gaps and offer guidance for practitioners navigating AI adoption, assurance, and workforce transformation in SE. We share the agreement data and the AI4SE/SE4AI Explorer web application so readers can compare their own relevance judgments with the human and AI raters.

13:00 JST画像/動画生成

BrainG3N: 制御可能な 3D 脳 MRI 生成のための多目的トークナイザー

三次元 (3D) 脳 MRI は臨床神経学および神経腫瘍学の中心であり、生成モデルは過小評価されているコホートを強化し、疾患の軌跡をシミュレートし、プライバシーを保護するデータ共有をサポートできます。潜在拡散は画像データをモデリングするための頼りになるソリューションですが、トークナイザーには 2 つの競合する要求が課せられます。エンコーダーの埋め込みは、下流のタスクが作用する臨床情報を保持する必要があり、デコーダーは解剖学的に忠実なボリュームを再構成する必要があります。既存の再構築駆動トークナイザーは、最初のトークナイザーを犠牲にして 2 番目のトークナイザーを実現します。これに対処するために、3D 脳 MRI 潜在拡散、デカップリング エンコーダーおよびデコーダー用の完全ボリューム マスク オートエンコーダー (MAE) ベースのトークナイザーを導入します。凍結された 3D MAE エンコーダーは臨床的に有益な埋め込みを生成し、専用の CNN デコーダーはそれらの埋め込みの線形投影からボクセルを再構築します。私たちは、4 つのモダリティ、10 の疾患カテゴリ、200 以上の取得サイトにわたる 18 の公的コホートからの 35,309 ボリュームでエンコーダーを事前トレーニングし、2 つの設定でその二重の有用性を実証します。まず、23 タスクの線形プローブ ベンチマークでは、エンコーダーは 23 タスク中 21 タスクで SOTA モデル (つまり、BrainIAC、BrainSegFounder、MedicalNet) を上回るか、またはそれに匹敵します。第 2 に、これらの臨床的に有益な埋め込みでトレーニングされた条件付き拡散変換器 (DiT) は、6 つの変数にわたる条件付き生成と患者固有の長期的予測の両方をサポートします。これらの結果を総合すると、下流の臨床タスクと制御可能な生成の両方を実行できる単一の 3D 脳 MRI 埋め込み空間が確立されます。

原文 (English)

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented cohorts, simulate disease trajectories, and support privacy-preserving data sharing. Latent diffusion has been the go-to solution for modeling imaging data, but it places two competing demands on the tokenizer: encoder embeddings must retain the clinical information that downstream tasks act on, and the decoder must reconstruct anatomically faithful volumes. Existing reconstruction-driven tokenizers achieve the second at the expense of the first. To address this, we introduce a fully volumetric masked-autoencoder (MAE) based tokenizer for 3D brain MRI latent diffusion, decoupling encoder and decoder: a frozen 3D MAE encoder produces clinically informative embeddings, while a dedicated CNN decoder reconstructs voxels from a linear projection of those embeddings. We pretrain the encoder on 35,309 volumes from 18 public cohorts spanning four modalities, ten disease categories, and 200+ acquisition sites, and demonstrate its dual utility in two settings. First, on a 23-task linear-probing benchmark, the encoder outperforms or matches SOTA models (i.e., BrainIAC, BrainSegFounder, and MedicalNet) on 21 of 23 tasks. Second, a conditional diffusion transformer (DiT) trained on these clinically informative embeddings supports both conditional generation across six variables and patient-specific longitudinal forecasting. Together these results establish a single 3D brain-MRI embedding space capable of both downstream clinical tasks and controllable generation.

13:00 JST研究/論文

コールドスタート推奨のための暗黙的フィードバックのノイズ除去

暗黙的フィードバックは、そのアクセシビリティと汎用性によりレコメンダー システムで広く使用されていますが、通常はノイズの多いサンプル (クリックベイト、位置バイアスなど) を提示します。一方、推奨者は、新しいアイテムが継続的に流入するため、必然的にアイテムのコールド スタートの問題に直面します。前述の要因により、冷たいアイテムはノイズの多いサンプルになりやすいことがわかりましたが、研究者は冷たいアイテムに対する暗黙的なフィードバックのノイズ除去の重要性を見落とすことがよくあります。これまでのノイズ除去研究では通常、より高い損失値などのヒューリスティック パターンに基づいてノイズの多いサンプルを特定し、サンプルの選択や再重み付けを通じてノイズを軽減しました。ただし、これらの方法の適応性は限られており、コールド スタート シナリオでは効果がありません。コールドスタート推奨のための暗黙的フィードバックのノイズ除去を実現するために、DIF と呼ばれるモデルに依存しないノイズ除去方法を提案します。まず、コンテンツに対するユーザーの好みが安定しているため、ユーザーがコンテンツに似たウォーム アイテムを通じてコールド アイテムに興味があるかどうかを示す疑似ラベルを推測することができます。さらに、擬似ラベルの精度を向上させるために、コールド品目とウォーム品目の内容類似性に基づいて擬似ラベルの信頼度をモデル化し、サンプルごとに複数の擬似ラベルを集計します。最後に、相対エントロピーとアイテムのコールドスタート状態を考慮して、ノイズのあるサンプルラベルの不確実性を明示的に推定します。これにより、擬似ラベルの役割が適応的にガイドされ、サンプルレベルでノイズのあるラベルが修正されます。 DIF の優位性は、理論的な正当性と現実世界のデータセットでの広範な実験の両方によって裏付けられています。この手法は、10 億ユーザー規模のショート ビデオ アプリケーション Kuaishou に導入され、コールド スタート シナリオ内のさまざまな商業指標を大幅に改善しました。

原文 (English)

Denoising Implicit Feedback for Cold-start Recommendation

Implicit feedback is widely used in recommender systems due to its accessibility and generality, yet it usually presents noisy samples (e.g., clickbait, position bias). Meanwhile, recommenders inevitably face the item cold-start problem due to the continuous influx of new items. We identify that cold items are more prone to noisy samples due to the aforementioned factors, and researchers often overlook the significance of denoising implicit feedback for cold items. Previous denoising studies usually identify noisy samples based on heuristic patterns, such as higher loss values, and mitigate noise through sample selection or re-weighting. However, these methods have limited adaptability and are ineffective in cold-start scenarios. To achieve denoising implicit feedback for cold-start recommendation, we propose a model-agnostic denoising method called DIF. First, user preferences for content remain stable, which allows us to infer pseudo-labels indicating whether a user is interested in a cold item through content-similar warm items. Furthermore, to improve pseudo-label accuracy, we model the confidence of pseudo-labels based on the content similarity between the cold item and warm items, and then aggregate multiple pseudo-labels for each sample. Finally, we explicitly estimate the uncertainty of the noisy sample label by considering its relative entropy and the cold-start status of the item, which adaptively guides the role of pseudo-labels to correct the noisy labels at the sample level. DIF's superiority is supported by both theoretical justification and extensive experiments on real-world datasets. The method has been deployed on a billion-user scale short video application Kuaishou and has significantly improved various commercial metrics within cold-start scenarios.

13:00 JSTエージェント研究/論文

分散型連合形成のための離脱と参加のダイナミクス

この論文では、一方的な離脱と参加の決定によって推進される分散型の動的プロセスとしての連合形成を研究します。エージェントは Aumann-Dreze 値を使用してローカルな動きを評価するため、報酬はグローバルに交渉された連合構造を通じてではなく、エージェントの現在の連合内で計算されます。結果として得られるモデルは、協調的なペイオフ配分と非協調的な最良応答行動を結び付けます。最終分割はまさに、個別に利益をもたらす離脱と結合の逸脱が許容されない連合構造です。平衡特性を確立し、ダイナミクスがスカラー リアプノフ表現または正確なポテンシャル表現を許容する条件を特定し、切り替えコストと受け入れコストが局所安定性をどのように形成するかを分析します。数値実験では、有限時間の安定化、コスト感度、および特別な凸ゲーム ベンチマークをテストします。

原文 (English)

Exit-and-Join Dynamics for Decentralized Coalition Formation

This paper studies coalition formation as a decentralized dynamical process driven by unilateral exit-and-join decisions. Agents evaluate local moves using the Aumann-Dreze value, so payoffs are computed within the agent's current coalition rather than through a globally negotiated coalition structure. The resulting model links cooperative payoff allocation with noncooperative best-response behavior: a terminal partition is precisely a coalition structure with no admissible, individually profitable exit-and-join deviation. We establish equilibrium characterizations, identify conditions under which the dynamics admit scalar Lyapunov or exact-potential representations, and analyze how switching and acceptance costs shape local stability. Numerical experiments test finite-time stabilization, cost sensitivity, and a special convex-game benchmark.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

静的リーダーボードを超えて: LLM エージェントの評価の予測的妥当性

エージェントのベンチマークは急速に成長していますが、展開によって明らかにされる 4 つまたは 5 つの側面を超えるベンチマークはありません。このペーパーは、MCP ベースの産業エージェント ベンチマークのこれまでで最大規模の調整された詳細調査を集約したものです。新しい資産クラス (マルチモーダルなビジュアル拡張を含む)、代替オーケストレーション、取得戦略、推論モード、インフラストラクチャの最適化、および評価方法論のプローブをカバーする 14 件の並行実装調査です。これらの研究を以前の 7 つのエージェント ベンチマークと統合すると、集計スコア リーダーボードは導入されたエージェントの評価を体系的に過小評価していると主張します。集計スコアから導出されたランキングは、配布外の設定には転送されません。最近の公開対非公開の競争の回顧展は、このランクの不安定性の直接的な経験的証拠を提供しています。私たちは、サンプル内平均ではなく、予測妥当性、サンプル内ランクとサンプル外ランク間の相関関係によるランキング構成を提案し、HELM とそのエージェント時代の後継者の崩壊を展開関連の次元で明らかにする 12 層の測定装置を報告します。このポジションは、明示的なしきい値を備えた 3 つの改ざん可能な配分外基準を通じて運用可能となります。既存の証拠は部分的にそれを裏付けていますが、確認するには薄すぎます。最後に、事前に登録されたパイロット設計と、次世代のエージェント ベンチマークが何を報告すべきかについての現場レベルのビジョンについて説明します。

原文 (English)

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.

13:00 JST画像/動画生成

GLARE: グローバルな説明をクエリするための自然言語インターフェイス

グローバルな説明は、データセット、クラス、意思決定コンテキストにわたるビジョン モデルを理解するために重要ですが、その複雑で一枚岩の性質が実際的な探索を妨げることがよくあります。ユーザーは通常、静的な成果物ではなく、特定の質問に対する的を絞った回答を求めるため、ブラックボックス画像分類器のグローバルな説明への自然言語アクセスを提供する LLM ベースの対話型インターフェイスを提供します。システムのコア LLM は仲介者として機能し、自然言語の質問をローカルの説明データに対する構造化された SQL クエリに変換します。これにより、ユーザーが低レベルの表現にさらされることなく、柔軟な集計が可能になります。クエリごとに、インターフェイスは統計によって拡張された自然言語応答を出力し、ローカルな説明と意図に合わせた視覚化をサポートします。意図の解釈、クエリ マッピングの精度、新しいクエリやデータセットに対する一般化、言語エラーに対する堅牢性についてシステムを評価します。私たちの結果は、LLM を介したクエリによって、人間中心の XAI のグローバルな説明のアクセシビリティとユーザビリティが大幅に向上することを示しています。

原文 (English)

GLARE: A Natural Language Interface for Querying Global Explanations

While global explanations are crucial for understanding vision models across datasets, classes, and decision contexts, their complex and monolithic nature often hinders practical exploration. Because users typically seek targeted answers to specific questions rather than static artifacts, we present an LLM-based interactive interface that provides natural language access to global explanations for black-box image classifiers. The system's core LLM acts as a mediator, translating natural language questions into structured SQL queries over local explanation data. This enables flexible aggregation without exposing users to low-level representations. For each query, the interface outputs statistics-augmented natural language responses, supporting local explanations, and intent-aligned visualizations. We evaluate the system on intent interpretation, query mapping accuracy, generalization to novel queries and datasets, and robustness to linguistic errors. Our results demonstrate that LLM-mediated querying substantially improves the accessibility and usability of global explanations for human-centered XAI.

13:00 JST研究/論文

進化するプログラムのボトルネックによるニューラル組み合わせ最適化の解釈

Neural Combinatorial Optimization (NCO) は優れたパフォーマンスを実現しますが、そのブラックボックスの性質は依然として展開と科学的診断の重要な障害となっています。概念ボトルネック モデル (CBM) などの標準的な解釈可能ツールは、決定が動的で状態に依存し、適切な概念語彙定義が欠けている NCO にとっては不十分です。このギャップを埋めるために、私たちは進化するプログラマティック ボトルネック (EPB) を導入しました。これは、私たちの知る限り、ブラック ボックス NCO モデルを人間が判読可能なプログラム ポートフォリオに抽出することによって NCO ポリシーを解釈するための最初のフレームワークです。 EPB は LLM を採用して一連のプログラムを自律的に進化させますが、各プログラムのステップごとのアクション分散がボトルネックとして機能します。 EPB は反復フレームワークを通じて機能します。ブロック I はプログラム バンクの容量を修正し、スチューデント ルーターの更新用の数値勾配と LLM ベースのプログラム リビジョン用のテキスト 勾配を結合するハイブリッド テキストと数値の勾配降下スキームを導入します。ブロック II は、障害をターゲットにした拡張と冗長プルーニングを通じてバンク容量を動的に適応させます。広範な実験により、EPB の有効性と幅広い適用性が実証され、抽出されたプログラム ポートフォリオは元のパフォーマンスとほぼ一致します。 EPB は、NCO の動作が最適化ステージ全体で変化し、古典的なヒューリスティック バリアントの構成として近似できることも明らかにしています。私たちの研究は、解釈可能な NCO を進歩させ、EPB を逐次意思決定モデルを解釈するための有望なツールとして確立します。

原文 (English)

Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks

Neural Combinatorial Optimization (NCO) achieves strong performance, yet its black-box nature remains a key roadblock to deployment and scientific diagnosis. Standard interpretability tools, such as Concept Bottleneck Models (CBMs), are ill-equipped for NCO, whose decisions are dynamic, state-dependent, and lack proper concept vocabulary definition. To close this gap, we introduce Evolving Programmatic Bottlenecks (EPB), to our knowledge, the first framework for interpreting NCO policies by distilling black-box NCO models into human-readable program portfolios. EPB employs an LLM to autonomously evolve a bank of programs, where each program's per-step action distribution serves as the bottleneck. EPB works through an iterative framework: Block I fixes program bank capacity and introduces a hybrid textual-numerical gradient descent scheme that couples numerical gradients for student router updates and textual gradients for LLM-based program revision; Block II dynamically adapts bank capacity via fault-targeted expansion and redundancy pruning. Extensive experiments demonstrate EPB's effectiveness and broad applicability, where the distilled program portfolios largely match original performance. EPB also reveals that NCO behavior shifts across optimization stages and can be approximated as a composition of classic heuristic variants. Our work advances interpretable NCO and establishes EPB as a promising tool for interpreting sequential decision-making models.

13:00 JST研究/論文

Quranic ASR の事前トレーニング済みトランスフォーマー モデルの比較研究: 音声表現、ラベル形式、およびデータセット構成

コーラン自動音声認識 (ASR) は、コーランの朗読をテキストに変換し、暗記支援ツールやコーラン検索エンジンなどのアプリケーションを可能にすることを目的としています。ただし、既存の ASR モデルは、ユーザーが朗読する聖句で高い単語誤り率 (WER) を示すことが多く、コーランのコーパスを完全にカバーしていません。この論文では、高度な音声特徴抽出手法である Wav2Vec2.0、HuBERT、および XLS-R を使用した、Quranic ASR の事前トレーニング済み Transformer ベースのモデルのドメイン固有の微調整に関する体系的な実証研究を紹介します。これらのモデルは、入力音声の一部をマスクし、Transformer アーキテクチャを使用してコンテキスト認識型音声特徴を学習することにより、自己教師あり学習を適用します。事前トレーニングされたモデルは、専門家とユーザーによる 870 時間の朗読を超えるフィルタリングされたコーラン データセットに基づいて微調整されています。特徴抽出器、出力ラベル形式、トレーニング戦略、およびクリップの長さにわたる包括的なアブレーション研究を通じて、この領域における転写の精度に影響を与える主要な要因を特定します。当社の最高パフォーマンスの構成では、EveryAyah サブセットで 0.08、EveryAyah+Tarteel の組み合わせ設定で 0.11 の WER を達成しました。これは、Citrinet ベースライン (WER = 0.163) に対しておよそ 5 パーセント ポイントの向上を示し、同時に、組み合わせモデルのトレーニング時間を 140 時間から 40 時間に短縮します。発音記号のないアラビア語テキストは最適な微調整結果をもたらし、Wav2Vec2-XLSR-53 は全体的に最も強力な表現を提供します。今後の作業には、データセットの品質の向上と、Tajweed に敏感なアプリケーション向けにより深い音声特徴表現を抽出するための音素認識モデルの開発が含まれます。

原文 (English)

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines. However, existing ASR models often exhibit high Word Error Rates (WER) on user-recited verses and lack full coverage of the Quranic corpus. This paper presents a systematic empirical study of domain-specific fine-tuning of pretrained Transformer-based models for Quranic ASR, using advanced speech feature extraction methods: Wav2Vec2.0, HuBERT, and XLS-R. These models apply self-supervised learning by masking portions of input audio and using Transformer architectures to learn context-aware speech features. The pretrained models are fine-tuned on a filtered Quranic dataset exceeding 870 hours of professional and user recitations. Through comprehensive ablation studies across feature extractors, output label formats, training strategies, and clip durations, we identify the key factors that affect transcription accuracy in this domain. Our best-performing configuration achieves a WER of 0.08 on the EveryAyah subset and 0.11 on the combined EveryAyah+Tarteel setting, representing roughly a five-percentage-point gain over the Citrinet baseline (WER = 0.163) while reducing combined-model training time from 140 hours to 40 hours. Arabic text without diacritics yields the best fine-tuning results, and Wav2Vec2-XLSR-53 provides the strongest overall representation. Future work includes improving dataset quality and developing phoneme-aware models to extract deeper speech feature representations for Tajweed-sensitive applications.

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

Agentic レビュー システムのベンチマーク

AI支援研究によって査読システムにかかる圧力に対する救済策として、新しい種類のエージェントレビューシステムが登場しているが、それをどのように評価すべきかは不明である。私たちは、最先端の効率的なモデルにわたる 6 つの LLM にわたって、2 つのオープンソース システム (OpenAIReview および coarse)、1 つの独自システム (Reviewer3)、およびゼロショット ベースラインを評価します。まず、ICLR/NeurIPS 論文の AI レビューが、引用や受理決定などの外部シグナルによって近似される論文の品質に追随するかどうかを研究します。すべてのシステムはペアごとの精度で偶然以上のパフォーマンスを発揮し、最も優れているのは OpenAIReview + GPT-5.5 の 83.0% です。次に、システムが既知のグラウンド トゥルースでエラーを捕捉できるかどうかをテストするために、8 つの arXiv 主題クラスにわたる論文に 4 つのカテゴリのエラーを注入する摂動ベンチマークを構築し、検出再現率を測定します。最も強力な構成 (OpenAIReview + GPT-5.5) は、挿入されたエラーの 71.6% を捕捉し、改善の余地がかなり残されています。 6 つのモデルにわたる検出の統合は 83.3% の再現率に達し、異なるモデルが異なるエラーを検出し、より良いハーネス設計によりパフォーマンスが向上する可能性があることを示唆しています。これらのベンチマークを超えて、実際のユーザーを使用して OpenAIReview のパブリック デプロイメントを研究します。そのコメントに対する投票は 1.44 対 1 で肯定的な意見に偏っており、最も一般的な苦情は誤検出や些細な指摘に関するものです。実際の研究論文で最先端のモデルに裏付けられた完全なレビュー システムを一緒に評価することで、AI レビューにはまだ改善の余地があるものの、すでに人間の品質判断を適切に追跡し、重要なエラーをキャッチし、実際のユーザーから肯定的なフィードバックを得ることができることを示します。

原文 (English)

Benchmarking Agentic Review Systems

A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coarse), one proprietary system (Reviewer3), and a zero-shot baseline, across six LLMs spanning frontier and efficient models. First, we study whether AI reviews on ICLR/NeurIPS papers track with papers' quality as approximated by external signals such as citations and acceptance decisions. Every system performs above chance in pairwise accuracy, and the best is OpenAIReview + GPT-5.5 at 83.0%. Second, to test whether systems can catch errors with known ground truth, we construct a perturbation benchmark that injects four categories of errors into papers across eight arXiv subject classes and measure detection recall. The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors, leaving substantial room for improvement. The union of detections across six models reaches 83.3% recall, suggesting different models detect different errors and better harness design can potentially increase performance. Beyond these benchmarks, we study a public deployment of OpenAIReview with real users. Votes on its comments skew positive at 1.44 to 1, and the most common complaints are about false positives and minor nitpicks. Together, by evaluating full review systems backed by state-of-the-art models on real research papers, we show that while AI reviews still have room for improvement, they can already track human quality judgments well, catch important errors, and earn positive feedback from real users.

13:00 JST研究/論文

グラウンデッド推論: 決定論的にカプセル化された生成モデルの原則

生成モデルを従来の計算システムに組み込むことは、大きなチャンスと大きな危険の両方をもたらします。多くの早期導入者は多大な費用をかけてこれらの危険性を認識していますが、この分野では依然として、AI を従来のシステムに組み込むリスクを回避するための基礎的なフレームワークが必要です。この原稿は、確率モデルの決定論的なカプセル化を可能にするように設計された、AI ブレンド アーキテクチャの 4 つの特定のプリミティブの定義を通じて、この基盤を確立します。さらに、業界全体で広く代表される 2 つの包括的なアンチパターンを確立し、この分野のエンジニアへの警告として機能します。このフレームワークは、生成モデル プロバイダーが次世代の生成モデル インターフェイスを構築できる基盤を提供しながら、AI を従来のシステムにうまく統合できるように設計されました。

原文 (English)

Grounded Inference: Principles for Deterministically Encapsulated Generative Models

The incorporation of generative models into traditional computational systems presents both enormous opportunity and tremendous peril. Although many early adopters have realized these perils at great expense, the field still requires foundational frameworks to de-risk incorporation of AI into traditional systems. This manuscript establishes this foundation through the definition of four specific primitives of AI blended architecture, designed to enable deterministic encapsulation of probabilistic models. It further establishes two overarching anti-patterns broadly represented across industry to serve as warnings for engineers in this field. This framework was designed to enable successful integration of AI into traditional systems while providing a foundation upon which generative model providers could build the next generation of generative model interfaces.

13:00 JST研究/論文

ナレッジワーカーの質問応答フォーラムにおける最適なスケジューリング

個人が疑問に対する答えを見つけるためにインターネットにアクセスするにつれて、いくつかの質問応答 (QA) フォーラムが発展してきました。そこでは、特定のトピックに精通したユーザーが専門知識を提供して、これらの情報要求に答えることができます。これらは現在ボランティアベースですが、将来的には特定のトピックの専門家である知識労働者を雇用するバージョンを検討しています。このようなシステムでは、キュー システムを形成する要求-回答プロセスは、さまざまなトピックの要求をフォーラムの専門家に割り当てるスケジューラを利用することができ、フォーラムの専門家はさまざまなトピックの専門知識レベルに応じてそれらの要求に答えることができます。このモデルでは、システムを安定に保ちながらリクエストを処理するためのシステムのキャパシティを計算し、キャパシティを達成するスケジューラを設計します。また、リクエストに応える際に専門家間の協力がどのように潜在的に容量を増加できるかについても調査します。

原文 (English)

Optimal Scheduling in a Question-Answering Forum of Knowledge Workers

As individuals turn to the Internet to find answers to questions they may have, several Question Answering (QA) forums have evolved, where users knowledgeable in certain topics can contribute their expertise to answering these requests for information. While these are currently volunteer based, we consider a future version employing knowledge workers who are experts in certain topics. In such a system, the request-answer processes forming the queuing system may utilize schedulers that assign requests in different topics to the experts in the forum, who may be able to answer them according to their expertise levels in different topics. With this model, we calculate the capacity of the system for handling the requests while keeping the system stable, and design schedulers that achieve capacity. We also investigate how collaboration between experts in answering requests can potentially increase capacity.

13:00 JSTLLM/生成AI

エントロピーを超えて: LLM 推論のためのトークンレベルの分布偏差からの学習

検証可能な報酬を伴う強化学習 (RLVR) は、大規模言語モデル (LLM) 推論を大幅に進歩させました。ただし、根本的な最適化の不安定性に直面しています。均一なトークンの更新はエントロピーの崩壊を促進し、次善の戦略への早期収束につながりますが、過剰なシャノンのエントロピーの最大化はエントロピーの爆発を引き起こし、一貫性のない推論チェーンへの盲目的な探索を引き起こす可能性があります。この二分法を解決するために、最適化の焦点をスカラーの不確実性からトークン ロジットの分布特性に移す独立組み合わせトークン (ICT) フレームワークを導入します。 ICT は、トークン ロジット分布間のジェンセン シャノン (JS) の相違を活用することで、LLM 推論における効果的な探索を導くための重要な分岐点として、独特の分布パターンを持つトークンを特定します。シャノンエントロピーと 2 次 R\'enyi エントロピーの両方に基づいた私たちの理論分析は、これらのトークンを選択的に更新することで政策の集中が調整されることを証明しています。つまり、シャノン エントロピーによって測定される全体的な分布の不確実性が軽減され、同時に 2 次 R\'enyi エントロピーによって捕捉される確率集中が制御されます。この二重の効果により、過度に集中したトークン生成による探査の弱体化が防止され、トレーニング状況が効果的に安定します。経験的な結果は、Qwen2.5 (0.5B/1.5B/7B) モデルの一意のトークンの上位 10% のみを更新すると、数学、常識、オリンピック レベルの問題にわたる 7 つのベンチマークにわたって、GRPO、20-エントロピー、および STAPO ベースラインを超えて、平均 pass@4 が 4.58% 向上し、最大 14.9% の向上が得られることを示しています。

原文 (English)

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced Large Language Model (LLM) reasoning; however, it faces a fundamental optimization instability: uniform token updates precipitate entropy collapse, leading to premature convergence to suboptimal strategies, whereas excessive Shannon Entropy maximization can cause entropy explosion, driving blind exploration toward incoherent reasoning chains. To resolve this dichotomy, we introduce the Independent Combinatorial Tokens (ICT) framework, which shifts the optimization focus from scalar uncertainty to the distributional properties of token logits. By leveraging the Jensen-Shannon (JS) divergence between token logits distributions, ICT identifies tokens with distinctive distributional patterns as critical branching points for guiding effective exploration in LLM reasoning. Our theoretical analysis, grounded in both Shannon and second-order R\'enyi entropy, proves that selectively updating on these tokens regulates policy concentration: it reduces the overall distribution uncertainty measured by Shannon entropy, while controlling probability concentration captured by second-order R\'enyi entropy. This dual effect prevents over-concentrated token generation from weakening exploration and effectively stabilizes the training landscape. Empirical results demonstrate that updating only the top 10% of unique tokens on Qwen2.5 (0.5B/1.5B/7B) models yields an average pass@4 improvement of 4.58%, with a maximum gain of 14.9%, over GRPO, 20-Entropy, and STAPO baselines across seven benchmarks spanning math, commonsense, and Olympiad-level problems.

13:00 JSTLLM/生成AIエージェントGemini

AgentFinVQA: 監査可能な財務チャート QA のための展開可能なマルチエージェント パイプライン

規制された環境における財務チャートの質問回答には、正確性以上のものが求められます。実務者は、回答に基づいて行動する前に、どの回答を信頼すべきかを知る必要があり、多くの機関は顧客データを外部モデルプロバイダーに送信できません。しかし、既存のチャート QA エージェントは精度重視かつ不透明で、ほとんどが独自の API アクセスを前提としています。私たちの知る限り、精度を大幅に損なうことなく監査可能性とオンプレミス展開可能性を組み合わせたものはありません。 AgentFinVQA は、各クエリを計画、OCR、凡例の根拠、視覚的検査、検証に分解し、サンプルごとに追跡可能なモデル評価パケット (MEP) のすべてのステップを記録するマルチエージェント パイプラインです。 FinMME では、AgentFinVQA は、独自のバックボーン (Gemini-3 フラッシュ; 71.24% 対 63.56%、McNemar $p \約 1.1 \times 10^{-16}$) を使用したプライマリ バックボーンと一致するゼロショット ベースラインと比較して $+7.68$ pp 改善し、オープンウェイトでは $+4.84$ pp 改善します。 Qwen3.6-27B-FP8 はローカルでサービスされています。検証者の評決は有用な信頼シグナル (確認済み回答と修正済み回答の正確な精度 68.2% 対 55.6%) としても機能し、人間によるレビュー ルーティングが可能になります。エラー分析により、質問の誤解、凡例の混乱、抽出エラーが失敗の 3 分の 2 近くを占め、検証者によって最も検出されにくいカテゴリーであることが示され、今後の作業の明確な方向性が特定されます。これらの結果を総合すると、監査可能なオンプレミスの財務チャート QA が実用的であり、オープンウェイト システムが完全なデータ常駐を可能にしながら精度の向上のほとんどを維持していることがわかります。再現可能な評価をサポートするためにコードをリリースします。

原文 (English)

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

Financial chart question answering in regulated settings demands more than accuracy: practitioners must know which answers to trust before acting on them, and many institutions cannot send client data to external model providers. Yet existing chart-QA agents are accuracy-focused and opaque, and most assume proprietary API access; to our knowledge, none combines auditability with on-premise deployability without significant accuracy compromise. We present AgentFinVQA, a multi-agent pipeline that decomposes each query into planning, OCR, legend grounding, visual inspection, and verification, recording every step in a traceable Model Evaluation Packet (MEP) per sample. On FinMME, AgentFinVQA improves $+7.68$ pp over a primary-backbone matched zero-shot baseline with a proprietary backbone (Gemini-3 Flash; 71.24% vs. 63.56%, McNemar $p \approx 1.1 \times 10^{-16}$), and $+4.84$ pp with open-weights Qwen3.6-27B-FP8 served locally. The verifier's verdict also serves as a useful confidence signal (68.2% vs. 55.6% exact accuracy on confirmed vs. revised answers), enabling human-in-the-loop review routing. Error analysis shows that question misunderstanding, legend confusion and extraction error account for nearly two-thirds of failures and are the categories least detected by the verifier, identifying clear directions for future work. Together these results show that auditable, on-premise financial chart QA is practical and that the open-weights system keeps most of the accuracy gains while enabling full data residency. We release our code to support reproducible evaluation.

13:00 JSTLLM/生成AIエージェント研究/論文

ORAgentBench: LLM エージェントは困難なオペレーション リサーチ タスクをエンドツーエンドで解決できますか?

大規模な言語モデルは、実行可能環境で複数ステップのタスクを実行するための自律エージェントとして導入されることが増えていますが、現実的なオペレーション リサーチ (OR) 作業を実行する能力は依然として不明です。既存の OR 評価では、多くの場合、モデリングと解決策が切り離されており、事前に形式化されたインスタンスまたはテキストのみのインスタンスに依存しており、運用成果物から検証済みの意思決定に至るワークフロー全体をテストすることはほとんどありません。この作業では、困難なエンドツーエンドのオペレーション リサーチ タスクで自律エージェントを評価するための実行ベースのベンチマークである ORAgentBench を紹介します。これには、さまざまな運用シナリオにわたって人間がレビューした 107 のタスクが含まれており、それぞれが自然言語の概要、複数ファイルのデータ、構成アーティファクト、および必要な送信スキーマを備えた隔離された環境にパッケージ化されています。エージェントはソリューション コードを作成して実行する必要があり、その送信内容は、スキーマの有効性、厳密な制約の実現可能性、および正規化された客観的な品質について、非表示のバリデーターによって評価されます。 14 のフロンティア エージェント モデル構成による実験では、現在のエージェントが信頼できる OR 実践からはほど遠いことが示されています。最も優れたエージェントが合格できるのは、全タスクの 35.51% と難しいタスクの 20.59% だけであり、実行可能な提出物の多くは依然として要求される品質のしきい値を下回っています。さらに、障害分析では、運用ルールの欠如、脆弱な定式化、実行可能なソリューションの構築の脆弱さ、不十分なソリューションの改善など、戦略的な弱点がエラーの大部分を占めていることがわかります。手術室固有の手順スキルは、困難なタスクの実現可能性を高めますが、ソリューションの品質や合格率を確実に向上させるものではありません。これらの結果は、OR エージェントの進歩には、もっともらしい最適化コードを超えて、信頼できる高品質な運用上の意思決定に移行する必要があることを示唆しています。

原文 (English)

ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR) work remains unclear. Existing OR evaluations often decouple modeling from solving, rely on pre-formalized or text-only instances, and rarely test the full workflow from operational artifacts to validated decisions. In this work, we introduce ORAgentBench, an execution-grounded benchmark for evaluating autonomous agents on challenging end-to-end operations research tasks. It contains 107 human-reviewed tasks across diverse operational scenarios, each packaged in an isolated environment with a natural-language brief, multi-file data, configuration artifacts, and a required submission schema. Agents must write and run solution code, and their submissions are evaluated by hidden validators for schema validity, hard-constraint feasibility, and normalized objective quality. Experiments with fourteen frontier agent-model configurations show that current agents remain far from reliable OR practice. The best agent passes only 35.51% of all tasks and 20.59% of hard tasks, and many feasible submissions still fall below the required quality threshold. Failure analysis further shows that errors are dominated by strategic weaknesses, including missed operational rules, brittle formulations, weak feasible-solution construction, and insufficient solution improvement. OR-specific procedural skills increase hard-task feasibility, but do not reliably improve solution quality or pass rate. These results suggest that progress in OR agents requires moving beyond plausible optimization code toward dependable, high-quality operational decision-making.

13:00 JSTLLM/生成AI研究/論文

CombEval: 大規模言語モデルでの組み合わせカウントを評価するためのフレームワーク

大規模な言語モデルで組み合わせカウントを評価するための動的ベンチマークである CombEval を紹介します。 CombEval は、各問題をエンティティ、組み合わせオブジェクト、オブジェクトの依存関係、および制約に対する型付きの Cofola 仕様として表現し、ソルバーによって正確に検証された回答を含む自然言語の計数問題の制御された生成を可能にします。静的コレクションとは異なり、CombEval は、オブジェクト タイプ、エンティティのスケール、制約の数、および推論の深さの体系的なバリエーションをサポートします。直接およびコード拡張設定の下で 11 個の LLM を評価したところ、順序付けされたオブジェクト、区別できない要素、相対的な位置の制約、および入れ子になったオブジェクトの依存関係に関してモデルが脆弱なままであることがわかりました。エラー分析により、制約の解釈とカウント原則の欠陥がさらに特定されます。 CombEval は、LLM がいつ、そしてなぜ組み合わせ推論で失敗するかを研究するための診断テストベッドを提供します。コードと生成されたベンチマーク スイートは、\url{https://github.com/YuxuZhou-CN/combination-problem-generation} で公開されています。

原文 (English)

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models. CombEval represents each problem as a typed Cofola specification over entities, combinatorial objects, object dependencies, and constraints, enabling controlled generation of natural-language counting problems with exact solver-verified answers. Unlike static collections, CombEval supports systematic variation of object type, entity scale, constraint count, and reasoning depth. We evaluate 11 LLMs under direct and code-augmented settings and find that models remain brittle on ordered objects, indistinguishable elements, relatively positional constraints, and nested object dependencies. Error analysis further identifies failures in constraint interpretation and counting principles. CombEval provides a diagnostic testbed for studying when and why LLMs fail at combinatorial reasoning. The code and generated benchmark suites are publicly available at \url{https://github.com/YuxuZhou-CN/combination-problem-generation}.

13:00 JSTLLM/生成AI

もう一度考えますか、それとも長く考えますか?予算を考慮した推論のための選択的検証

テスト時の推論は、提供時間の制御ノブとしてますます使用されていますが、追加の推論は一様に価値があるわけではありません。失敗した試行を修復したり、既に正解している解答の計算を無駄にしたり、有害な解答の変更を導入したりする可能性があります。私たちはこれを、新規検証者の問題ではなく、デプロイメント割り当て問題として研究します。 \sevra (推論割り当てのための選択的検証) を導入します。これは、フリーズしたソルバーの最初の答えを保存するか、アクティブな検証を呼び出すかを決定するサービス層コントローラーです。凍結された Qwen3-4B ソルバーを使用して、介入の結果をログに記録し、サービングの可視の試行状態から回復可能性を認識したゲートをトレーニングします。 \mathfive では、選択的検証の精度は 76.3\% に達し、常時検証の 75.5\% と比較して、生成後のトークンが 26.8\% 削減され、有害な反転が 2.2\% から 1.0\% に減少します。ただし、8,192 トークンの初期解決では、モデル トークンの総数が 28\% 減りながら 76.0\% の精度に達しました。これは、選択的回復は有用ですが、最もテストされたコスト フロンティアではないことを示しています。 \gsm への凍結転送では、選択的ポリシーは例の 3.0\% のみを検証し、精度を 93.4\% から 94.5\% に向上させ、常に検証する場合と比較して検証トークンを 91.2\% 削減します。繰り返しますが、初期ソルブが長いほど、より少ない実現トークンで精度が高まります。 CommonsenseQA では、常時検証が問題になりますが、Self-Consistency@5 は実際のトークン コストの約 5 倍で精度を向上させます。結果として得られるデプロイメント ルールは、最初に初期予算を調整し、次に明示的なチェック、制限された再試行、監査可能性、または回帰リスク制御が重要な場合に選択的リカバリを使用するというものです。

原文 (English)

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes. We study this as a deployment allocation problem rather than a new-verifier problem. We introduce \sevra, Selective Verification for Reasoning Allocation, a serving-layer controller that decides whether to preserve a frozen solver's initial answer or invoke active verification. Using a frozen Qwen3-4B solver, we log intervention outcomes and train recoverability-aware gates from serving-visible attempt state. On \mathfive, selective verification reaches 76.3\% accuracy, compared with 75.5\% for always verifying, while reducing post-generation tokens by 26.8\% and harmful flips from 2.2\% to 1.0\%. However, an 8,192-token initial solve reaches 76.0\% accuracy with 28\% fewer total model tokens, showing that selective recovery is useful but not the best tested cost frontier. In frozen transfer to \gsm, the selective policy verifies only 3.0\% of examples, improves accuracy from 93.4\% to 94.5\%, and reduces verification tokens by 91.2\% relative to always verifying; again, a longer initial solve matches its accuracy with fewer realized tokens. On CommonsenseQA, always-on verification hurts, while Self-Consistency@5 improves accuracy at about five times the realized token cost. The resulting deployment rule is: tune the initial budget first, then use selective recovery when explicit checks, bounded retries, auditability, or regression-risk control matter.

13:00 JSTLLM/生成AIエージェント

AI 支援による法的証拠開示のためのヒューマン オン ザ ループ オーケストレーション

Autonomous Large Language Model (LLM) エージェントは電子証拠開示 (e-discovery) に導入されることが増えており、複数ステップの推論チェーンにわたる複合エラーが法的違法行為となる可能性があります。シングルターン取得とは異なり、特権付きドキュメント コーパス上で動作するエージェント ワークフローは、「軌道崩壊」と呼ばれる一種の失敗を示します。つまり、初期の誤分類が静かに伝播し、特権レビュー全体が無効になります。この論文は 3 つの貢献を行っています。まず、機能段階ごとに整理した、法律情報検索におけるエージェントの失敗の構造化分類を提案します。次に、これらの障害が悪化する前に阻止するように設計された、計画、推論、実行、不確実性の定量化に及ぶ 4 層の検証アーキテクチャを導入します。 3 番目に、必須のヒューマン オン ザ ループ (HOTL) エスカレーションしきい値が、完全に自律的なベースラインと比較して特権放棄のリスクをどのように低減するかを実証する、合成電子証拠開示コーパスに関する予備的なシミュレーション研究を紹介します。私たちの結果は、調整された不確実性のしきい値により、完全に自律的な展開と比較して特権放棄のリスクを最大 61% 削減できる一方、弁護士の審査に回される文書は 4 分の 1 未満であることを示唆しています。

原文 (English)

Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery

Autonomous Large Language Model (LLM) agents are increasingly deployed in electronic discovery (e-discovery), where compounding errors across multi-step reasoning chains can constitute legal malpractice. Unlike single-turn retrieval, agentic workflows operating over privileged document corpora exhibit a class of failure we term "trajectory collapse": an early misclassification silently propagates, rendering an entire privilege review invalid. This paper makes three contributions. First, we propose a structured taxonomy of agentic failures in legal information retrieval, organized by functional stage. Second, we introduce a four-layer verification architecture -- spanning planning, reasoning, execution, and uncertainty quantification -- designed to intercept these failures before they compound. Third, we present a preliminary simulation study on a synthetic e-discovery corpus that demonstrates how mandatory Human-on-the-Loop (HOTL) escalation thresholds reduce privilege-waiver risk relative to fully autonomous baselines. Our results suggest that calibrated uncertainty thresholds can reduce privilege-waiver risk by up to 61% versus fully autonomous deployment, while routing fewer than one quarter of documents to attorney review.

13:00 JSTエージェント

TelcoAgent: 3GPP に基づいた説明可能性を備えたスケーラブルな 5G マルチ KPM 予測

Key Performance Measurement (KPM) 予測は、5G および次世代通信ネットワークのプロアクティブなネットワーク管理に不可欠です。ただし、既存の機械学習 (ML) アプローチは、スケーラビリティと説明可能性において大きな制限に直面しており、現実世界の展開における有効性が制限されています。私たちは、サイト固有のトレーニングを必要とせずに、さまざまなネットワーク セルにわたる複数の KPM の正確でスケーラブルで説明可能な予測を可能にする基盤モデル ベースのフレームワークである TelcoAgent を提案します。具体的には、このフレームワークは 3 つの主要コンポーネントで構成されます。(i) 仕様書から直接 3GPP (3rd Generation Partnership Project) ナレッジ グラフを構築する自動化された 3 エージェント パイプライン、(ii) 正確なゼロショット予測を実現するスケーラブルな時系列基礎モデル (TSFM) ベースの予測パイプライン、最後に (iii) 実用的なドメインベースの診断を提供する推論および説明パイプライン。米国を拠点とするネットワーク事業者が提供する 3 か月間の現実世界の都市規模の 5G KPM データセットを使用して評価した TelcoAgent は、200 セルにわたるセルごとに考慮された 7 つの KPM のすべてについて高い予測精度を示し、同時にネットワークの劣化に対処するための説明可能な洞察と実行可能な指示を提供します。

原文 (English)

TelcoAgent: A Scalable 5G Multi-KPM Forecasting With 3GPP-Grounded Explainability

Key Performance Measurement (KPM) forecasting is essential for proactive network management of 5G and next-generation telecom networks. However, existing machine learning (ML) approaches face significant limitations in scalability and explainability, restricting their effectiveness in real-world deployments. We propose TelcoAgent, a foundation model-based framework that enables accurate, scalable, and explainable forecasting of multiple KPMs across diverse network cells without the need for site-specific training. Specifically, the framework comprises three key components: (i) an automated three-agent pipeline that constructs a 3rd Generation Partnership Project (3GPP) knowledge graph directly from specification documents, (ii) a scalable, time-series foundation model (TSFM)-based prediction pipeline to deliver accurate, zero-shot forecasting, and finally (iii) a reasoning and explanation pipeline that provides actionable, domain-grounded diagnostics. Evaluated using a 3-month, real-world, city-scale 5G KPM dataset from a U.S.-based network operator, TelcoAgent demonstrates high forecasting accuracy for all 7 considered KPMs per cell across 200 cells, while delivering explainable insights and actionable instructions to address network degradations.

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達

大規模言語モデルのブラックボックス不確実性推定法の系統的評価

大規模言語モデル (LLM) は幅広いタスクにわたって強力な機能を示していますが、その出力は信頼性が低いことが多く、幻覚が含まれる可能性があるため、信頼できる LLM を構築するには不確実性推定 (UE) が不可欠です。実際には、多くの主流 LLM は制限された API を介してのみアクセスでき、ロジットや隠れ状態などの内部信号は利用できないため、ブラックボックス UE が特に重要になります。しかし、LLM 用のブラックボックス UE に関する既存の研究は、方法論において断片的なままであり、統一された実証的比較が欠けています。このギャップに対処するために、ブラックボックス UE 手法の体系的なレビューを提示し、言語化ベース、サンプリング ベース、説明ベース、マルチエージェント、およびハイブリッド手法の 5 つのカテゴリに整理します。さらに、統一された評価フレームワークを構築し、4 つのモデルと 4 つのデータセット設定にわたる 24 の代表的な手法をベンチマークします。私たちの結果は、すべての設定において一貫して優勢な単一の方法はないことを示しています。それにもかかわらず、回答空間内の候補を推論して比較する方法は一般に効果的であり、複数の不確実性信号を組み合わせるハイブリッド方法は、ほとんどの条件下で良好に機能します。ベンチマーク データと統一評価フレームワークを公開することで、再現可能な比較を促進し、将来の研究をサポートすることを目指しています。また、実証結果は、LLM 向けの将来のブラック ボックス UE 手法を開発するための実践的なガイダンスを提供します。

原文 (English)

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.

13:00 JSTエージェント研究/論文

MetaResearcher: 敵対的な仮想環境における自己反省強化学習によるディープリサーチの拡張

深層調査エージェントは、自律的な情報収集と合成において優れた能力を実証してきましたが、そのトレーニングは、シミュレートされた環境の静的な性質、事実検索のみのタスク設計の限界、および結果ベースの強化学習の非効率性によって依然として制約を受けています。この研究では、4 つの相乗的な側面にわたって詳細な調査エージェントのトレーニングを拡張する新しいフレームワークである MetaResearcher を提案します。まず、時間的なダイナミクスと敵対的な誤った情報をトレーニング環境に注入する進化する仮想世界を導入し、エージェントにソースの信頼性評価と時間的な競合解決スキルの開発を強制します。次に、仮説生成や矛盾解決を含む発見指向のタスクを設計します。これは、単純な事実検索を超えて、エージェントを真の調査行動へと導きます。第三に、回答の正確性、検索パスの効率、反映の深さ、ツール呼び出しの多様性を共同で最適化し、以前の研究で観察された反復アクションのループ問題に直接対処する、GRPO フレームワーク内の自己反射型メタ報酬メカニズムを提案します。 4 番目に、調整された強化学習を通じて共同研究戦略を学習する、特殊な Scout、Filter、および Synthesizer モデルで構成される異種マルチエージェント Swarm アーキテクチャを導入します。 LiteResearcher インフラストラクチャ上に構築された MetaResearcher は、トレーニングに限界 API コストをゼロにしながら、ベンチマーク パフォーマンス (GAIA、Xbench-DS) と敵対的条件下での認識論的堅牢性の両方の大幅な向上を目指しています。完全なフレームワーク設計、トレーニング方法論、計画された実験的検証を紹介します。

原文 (English)

MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments

Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remains constrained by the static nature of simulated environments, the limits of fact-retrieval-only task designs, and the inefficiency of outcome-based reinforcement learning. In this work, we propose MetaResearcher, a novel framework that scales deep research agent training across four synergistic dimensions. First, we introduce an Evolving Virtual World that injects temporal dynamics and adversarial misinformation into the training environment, forcing agents to develop source credibility assessment and temporal conflict resolution skills. Second, we design Discovery-Oriented Tasks -- including hypothesis generation and contradiction resolution -- that transcend simple fact retrieval and push agents toward genuine research behaviors. Third, we propose a Self-Reflective Meta-Reward mechanism within the GRPO framework that jointly optimizes for answer correctness, search path efficiency, reflection depth, and tool call diversity, directly addressing the repetitive action loop problem observed in prior work. Fourth, we introduce a Heterogeneous Multi-Agent Swarm architecture comprising specialized Scout, Filter, and Synthesizer models that learn collaborative research strategies through coordinated reinforcement learning. Built upon the LiteResearcher infrastructure, MetaResearcher requires zero marginal API cost for training while targeting substantial improvements in both benchmark performance (GAIA, Xbench-DS) and epistemic robustness under adversarial conditions. We present the complete framework design, training methodology, and planned experimental validation.

13:00 JSTLLM/生成AIエージェント

マルチエージェントのトランザクション メモリ

多様なタスクにわたって多様な機能を備えた LLM エージェントの分散展開により、異種エージェント集団全体で知識を共有するためのインフラストラクチャが促進されます。検索エンジンが人間の問題解決をサポートするために人間が生成した成果物にインデックスを付けるのと同じように、検索システムはエージェントが生成した成果物を整理してエージェント集団全体で再利用できます。私たちは、人間が作成したアーティファクトの価値を個々のエージェントに実証する検索拡張生成を、エージェントの集団をサポートするエージェント生成アーティファクトの検索まで拡張します。特に、エージェントの軌跡は再利用可能な手順知識をエンコードしますが、これらのアーティファクトは通常 1 回の使用後に破棄されるか、作成エージェントによってのみ保持されるため、新しくインスタンス化されたエージェントは既存のソリューションを繰り返し再発見する必要があります。我々は、エージェントが生成した軌跡を集団レベルで保存および取得するためのフレームワークであるマルチエージェント トランザクション メモリ (MATM) を提案します。このフレームワークでは、プロデューサー エージェントが軌跡を共有リポジトリに提供し、コンシューマ エージェントが軌跡を取得してタスクの実行を改善します。私たちは、軌跡が長く、特に豊富な手続き構造をエンコードするインタラクティブな環境 (ALFWorld および WebArena) に焦点を当てています。私たちの実験では、MATM から軌道を取得すると、調整や共同トレーニングを行わなくても、下流のタスクのパフォーマンスが向上し、インタラクションのステップが削減されることが実証されました。これらの結果は、MATM をオープン エージェント エコシステムにおける集団レベルのエクスペリエンス共有のための設計パターンとして位置づけています。

原文 (English)

Multi-Agent Transactive Memory

The decentralized deployment of LLM agents with diverse capabilities across diverse tasks motivates infrastructure for knowledge sharing across heterogeneous agent populations. Just as search engines index human-generated artifacts to support human problem solving, retrieval systems can organize agent-generated artifacts for reuse across agent populations. We extend retrieval-augmented generation - which demonstrates the value of human-authored artifacts to individual agents - to retrieval of agent-generated artifacts supporting a population of agents. In particular, agent trajectories encode reusable procedural knowledge, yet these artifacts are typically discarded after a single use or retained only by the producing agent, forcing newly instantiated agents to repeatedly rediscover existing solutions. We propose Multi-Agent Transactive Memory (MATM), a framework for population-level storage and retrieval of agent-generated trajectories, where producer agents contribute trajectories to a shared repository and consumer agents retrieve them to improve task execution. We focus on interactive environments (ALFWorld and WebArena), where trajectories are long and encode especially rich procedural structure. Our experiments demonstrate that retrieving trajectories from MATM improves downstream task performance and reduces interaction steps without coordination or joint training. These results position MATM as a design pattern for population-level experience sharing in open agent ecosystems.

13:00 JST研究/論文

eCNNTO: トポロジー最適化を加速するための高度に汎用化可能な ConvNet

この研究では、eCNNTO と呼ばれる、密度ベースのトポロジー最適化 (TO) を加速する要素ベースの畳み込みニューラル ネットワーク (CNN) を提案します。 TO では一般に多数の反復が行われ、反復ごとに有限要素解析が実行されるため、特に高解像度設計を達成するために高密度メッシュが使用される場合に効率のボトルネックが発生します。この制限に対処するために、eCNNTO は Kallioras らに基づいて構築することが提案されています。 (2020) では、ディープ ビリーフ ネットワーク (DBN) がすべての要素に対してトレーニングされ、初期の歴史から最適に近い密度を予測することで、反復の大部分がスキップされ、TO 手順が大幅に高速化されました。ただし、この方法には隣接する要素間の空間的相関が欠如しており、最終的な構造でフィーチャが切断される可能性があります。提案された方法は、この問題に対処するために残留接続を備えた CNN を採用します。それに加えて、最適化効率をさらに高めるために新しいトレーニング戦略が導入されており、トレーニング データセットは初期のものではなく最終段階の密度履歴で構成されています。この変更は、必要なトレーニング データのサイズを削減するのにも役立ちます。 eCNNTO はトレーニングに小さなデータセットしか必要としませんが、大きく異なる境界条件、荷重ケース、設計ドメインのジオメトリ、メッシュ解像度、および非設計ドメインの問題に一般化できます。最終的に、eCNNTO の一般化機能と効率が、2 次元および 3 次元のさまざまな例を通じて実証され、それぞれ最大 90% と 97% の反復の削減が達成されました。

原文 (English)

eCNNTO: A Highly Generalizable ConvNet for Accelerating Topology Optimization

This work proposes an element-based Convolutional Neural Network (CNN) to accelerate density-based Topology Optimization (TO), termed eCNNTO. TO generally undergoes a large number of iterations, where finite element analysis is performed in every iteration, leading to the efficiency bottleneck especially when dense meshes are used to achieve high-resolution designs. To address this limitation, eCNNTO is proposed to build upon Kallioras et al. (2020), where a Deep Belief Network (DBN) was trained for every element to predict its near-optimal density from its early history, thereby skipping the great majority of iterations and significantly accelerating the TO procedure. However, the method lacks spatial correlations among neighboring elements and may lead to disconnected features in the final structure. The proposed method employs CNN with residual connections to address this issue. On top of it, a novel training strategy is introduced to further enhance the optimization efficiency, where the training dataset consists of the final stage density histories rather than early ones. This change can also help reduce the required training data size. eCNNTO requires only a small dataset to train and yet it can be generalized to problems with largely different boundary conditions, loading cases, design domain geometries, mesh resolutions, as well as non-design domains. In the end, the generalization capabilities and efficiency of eCNNTO are demonstrated through a variety of examples in two and three dimensions, achieving up to 90% and 97% reduction of iterations, respectively.

13:00 JST研究/論文

エージェンシーの道: Autotelic AI、組み込みエージェンシー、自己の解体

ほとんどの人工知能システムは、目標が外生的であり、設計者によって指定されるという前提に基づいて構築されています。エージェントが独自の目標を生成し始めると何が起こるかを探ることで、オートテリック AI の分野が開かれます。エージェントには、単に目的を追求するだけでなく、それを発見することが期待されています。この記事では、内発的動機づけ、リソース主導型事前分布、因果介入学習、ホメオスタシス、および埋め込み性を通じてその結果を追跡します。最後の条件は、自己主体性にとって必要条件ではあるが、十分条件ではないことが判明しています。埋め込み性は、その個性がユニークではないことを明らかにするという代償を払ってエージェントを個性化します。そのため、同じダイナミクスで多くの有効な分割が許容され、それぞれが異なる自己候補を定義します。したがって、オートテリック AI の最も深刻な問題は、エージェントがどのようにして目標を生成するかということではなく、エージェントがどのようにして目標が割り当てられる自己を生成し、相対化するかということです。エージェントは行動するために自分自身の境界を信じ、理解するためにその境界を見通さなければなりません。私たちはこれらの開発を単一のフレームワークに統合し、それを 3 つの方向に沿って拡張します。エージェントと環境の切断が物理的になる量子定式化、非二元的な瞑想的伝統に対する哲学的解釈、および具体的な LLM ベースのエージェントのインスタンス化です。

原文 (English)

The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self

Most artificial intelligence systems are built on the assumption that goals are exogenous and specified by the designer. Exploring what happens when an agent begins generating its own goals opens the field of autotelic AI. Agents are expected not merely to pursue objectives but to discover them. In this article, we trace its consequences through intrinsic motivation, resource-driven priors, causal-interventional learning, homeostasis, and embeddedness; the last of which is found to be a necessary but not sufficient condition for autotelic agency. Embeddedness individuates the agent at the cost of revealing that the individuation is non-unique, such that the same dynamics admit many valid partitions, each defining a different candidate self. The deepest problem with autotelic AI is therefore not how the agent generates goals, but how it generates and relativizes the self to which the goals are assigned. The agent must believe in its own boundary in order to act, and see through that boundary in order to understand. We consolidate these developments into a single framework and extend it along three directions: a quantum formulation in which the agent-environment cut becomes physical, a philosophical reading against non-dual contemplative traditions, and a concrete LLM-based agentic instantiation.

13:00 JSTロボティクス

PhysDrift: ヒューマノイドの共同音声モーション生成における身体のギャップを埋める

ヒューマノイドロボットは、表現力豊かで音声に合わせて動作するだけでなく、実施形態の制約の下で物理的に実行可能な同時音声動作を必要とします。既存の同時音声生成パイプラインは主に人間中心です。モーションは最初に SMPL-X などの人体表現で生成され、その後人型ロボットに再ターゲットされます。この研究では、このパラダイムにおける基本的な実施形態のギャップを特定します。つまり、人間の動作多様体と人型の実施形態の制約との間の不一致により、動作の伝達と物理的な実行中に実施形態の一貫性が損なわれるということです。広範な分析を通じて、リターゲットは粗い動きのセマンティクスを維持できるものの、動きの多様性を大幅に圧縮し、韻律と動きの同期を弱め、表現力豊かなヒューマノイドの動作を制限することを示しました。この問題に対処するために、我々はまず、リターゲティング中の運動学的実現可能性と音声と動作の時間的整合を共同で最適化する、韻律を保存するヒューマノイド動作キュレーションフレームワークである IK-EER を提案します。厳選されたロボットネイティブのモーションデータセットに基づいて、中間の人体の表現に依存せずに音声から実行可能なヒューマノイド関節の軌道を直接予測する、実施形態を意識した同時音声モーション生成フレームワークである PhysDrift をさらに紹介します。従来の人間中心のパイプラインとは異なり、PhysDrift は、ロボットの動作ダイナミクスを安定させるために物理的正則化を組み込みながら、トレーニングと推論の両方を通じて実施形態の一貫性を維持します。広範な実験と現実世界のヒューマノイド展開により、実施形態を意識したロボットネイティブ生成により、音声と動作の整合性、物理的な妥当性、動作の滑らかさ、推論効率、およびリアルタイムのインタラクション能力が大幅に向上することが実証されました。

原文 (English)

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generation pipelines are predominantly human-centric: motions are first generated in human-body representations such as SMPL-X and subsequently retargeted to humanoid robots. In this work, we identify a fundamental embodiment gap in this paradigm, where the mismatch between human motion manifolds and humanoid embodiment constraints disrupts embodiment consistency during motion transfer and physical execution. Through extensive analysis, we show that although retargeting can preserve coarse motion semantics, it significantly compresses motion diversity and weakens prosody-motion synchronization, limiting expressive humanoid behaviors. To address this problem, we first propose IK-EER, a prosody-preserving humanoid motion curation framework that jointly optimizes kinematic feasibility and speech-motion temporal alignment during retargeting. Building upon the curated robot-native motion dataset, we further introduce PhysDrift, an embodiment-aware co-speech motion generation framework that directly predicts executable humanoid joint trajectories from speech without relying on intermediate human-body representations. Unlike conventional human-centric pipelines, PhysDrift maintains embodiment consistency throughout both training and inference while incorporating physical regularization to stabilize robot motion dynamics. Extensive experiments and real-world humanoid deployment demonstrate that embodiment-aware robot-native generation substantially improves speech-motion alignment, physical plausibility, motion smoothness, inference efficiency, and real-time interaction capability.

13:00 JSTエージェント

自動組み込みダイアログ拡張による DialNav の進化

物理的な相互作用が可能な実体エージェントの場合、安全性と有効性の両方を確保するには、対話を作成して理解する能力が不可欠です。 DialNav~\cite{han2025dialnav} はダイアログの全体的な評価、つまりフォトリアリスティックな屋内ナビゲーションの実行ループのためのフレームワークを提供しますが、そのパフォーマンスはトレーニング データ (2K エピソード) の重大な不足によって依然として制限されています。これに対処するために、自動生成パイプラインを提案し、DialNav 用の 238,000 エピソードを含む大規模なトレーニング データセットである \textbf{RAINbow} データセットを構築します。当社のパイプラインは、既存の VLN データセットをマルチターン ダイアログに変換し、コスト効率の高い高品質のデータセットを作成します。次に、データの可能性を最大限に引き出すための 2 つの追加の補完的な進歩を導入します。(1) ダイナミック ダイアログ ナビゲーション ループとナビゲーション トレーニングを調整するナビゲーション トレーニング スキームであるデュアル戦略トレーニング、および (2) VLN 知識を活用するローカリゼーション モデル。これらの補完的なソリューションを組み合わせることで、私たちのモデルは \textbf{Val Seen} (58.24, \textbf{+89\%}) と \textbf{Val Unseen} (29.05, \textbf{+100\%}) の両方の分割で成功率のベースラインを大幅に上回り、新たな最先端技術を確立しました。

原文 (English)

Advancing DialNav through Automatic Embodied Dialog Augmentation

For embodied agents capable of physical interaction, the capability to create and understand dialog is crucial to ensure both safety and effectiveness. While DialNav~\cite{han2025dialnav} provides a framework for holistic evaluation of the dialog--execution loop in photorealistic indoor navigation, its performance remains limited by a critical scarcity of training data (2K episodes). To address this, we propose an automatic generation pipeline, and construct the \textbf{RAINbow} dataset, a large-scale training dataset with 238K episodes for DialNav. Our pipeline converts existing VLN datasets into multi-turn dialog and creates cost-efficient and high-quality dataset. Then, we introduce two additional complementary advances to unlock the data's full potential: (1) Dual-Strategy Training, a navigation training scheme to align the navigation training with the dynamic dialog-navigation loop, and (2) a localization model that leverages VLN knowledge. By combining these complementary solutions, our model substantially outperforms the baseline in success rate on both \textbf{Val Seen} (58.24, \textbf{+89\%}) and \textbf{Val Unseen} (29.05, \textbf{+100\%}) splits, establishing a new state of the art.

13:00 JSTエージェントロボティクス

ENPIRE: 現実世界でのエージェント ロボット ポリシーの自己改善

現実世界で器用なロボット操作を実現するには、人間の監視とアルゴリズム工学に大きく依存しており、これが一般的な物理的知性の追求において中心的なボトルネックとなります。新興のコーディング エージェントはアルゴリズム検索を自動化するコードを生成できますが、その成功は依然としてデジタル環境に限定されています。私たちは、ロボット研究を自動化するために欠けている抽象化は、現実世界のポリシー改善のための反復可能なフィードバック ループであると推測します。つまり、シーンをリセットし、ポリシーを実行し、結果を検証し、次の反復を改良するというものです。このギャップを埋めるために、コーディング エージェント用のハーネス フレームワークである ENPIRE を導入します。このフレームワークは、4 つのコア モジュールでこの物理フィードバック ルーチンをインスタンス化します。1 つは自動リセットと検証のための環境モジュール (EN)、ポリシーの改良を開始するポリシー改善モジュール (PI)、1 つまたは複数の物理ロボットを並行して動作させてポリシーを評価するロールアウト モジュール (R)、およびコーディング エージェントがログを分析し、文献を参照し、トレーニング インフラストラクチャと障害モードに対処するためのアルゴリズム コードを改善する進化モジュール (E) です。この閉ループ システムは、現実世界の操作学習を制御可能な最適化手順に変換し、人間の労力を最小限に抑えながら、トレーニング レシピとエージェントのバリエーション全体で公平なアブレーションを可能にします。 ENPIRE を活用することで、フロンティア コーディング エージェントはポリシーを自律的にトレーニングして、ピン ボックスの整理、結束バンドの締め付け、工具の使用などの困難で器用な操作タスクで 99% の成功率を達成できます。ロボット フリートにエージェント チームを派遣すると、このプロセスがさらに加速します。私たちの結果は、物理世界で自律的に進歩するロボット工学にコーディング エージェントを展開するための実用的でスケーラブルな道筋を示唆しています。

原文 (English)

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world.

13:00 JSTエージェント

具現化された世界モデルのエージェントとしての報酬

RL は世界モデルを改良するための有望なツールとなっていますが、既存の手法は主にトレーニング分布付近の保守的なロールアウトに依存しており、探索、行動の多様性、より豊富な動的発見が制限されています。この作品では、私たちはこの保守的なパラダイムに挑戦します。私たちは、核となる制限は探査そのものではなく、より広範な探査をサポートする信頼できる検証戦略が欠如していることにあると主張します。信頼できる検証がなければ、拡張された探査は報酬ハッキングの影響を非常に受けやすくなり、真の改善が達成されずにポリシーが不完全な報酬を悪用することになります。この動機を評価するために、私たちは具体化された世界モデルでメソッドをインスタンス化します。そこでは、物理的な妥当性とタスクの完了が、複雑なダイナミクスの下でスケーラブルな RL の厳密なテストベッドを提供します。検証面では、生成された動作をアクティブに評価して堅牢な報酬シグナルを提供し、配布の変化下での報酬ハッキングを軽減するエージェント報酬フレームワークである Reward as an Agent を導入します。探査面では、DynDiff-GRPO を通じて動的認識ロールアウト多様化を導入します。これにより、行動空間探査が明示的に拡張され、軌道を多様化し、国家活動の範囲を拡大し、保守的なロールアウト体制を超えてより豊かに具体化された行動が奨励されます。 Reward as an Agent を DynDiff-GRPO と統合することで、大幅に多様化したサンプリングを備えたより信頼性の高い報酬基盤で RL を実現し、報酬ハッキングを効果的に緩和しながら、複数のオープンソースの世界モデル全体で大幅な精度の向上を実現します。これにより、堅牢な検証に基づいて広範な探索を適切に拡張できることを実証します。

原文 (English)

Reward as An Agent for Embodied World Models

While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distribution, limiting exploration, behavioral diversity, and richer dynamic discovery. In this work, we challenge this conservative paradigm. We argue that the core limitation is not exploration itself, but the lack of reliable verification strategies to support broader exploration. Without reliable verification, expanded exploration becomes highly susceptible to reward hacking, where policies exploit imperfect rewards without achieving genuine improvement. To evaluate this motivation, we instantiate our method in embodied world models, where physical plausibility, and task completion provide a rigorous testbed for scalable RL under complex dynamics. On the verification side, we introduce Reward as an Agent, an agentic reward framework that actively evaluates generated behaviors to provide robust reward signals and mitigate reward hacking under distribution shifts. On the exploration side, we introduce Dynamic-Aware Rollout Diversification through DynDiff-GRPO, which explicitly expands action-space exploration to diversify trajectories, broaden state-action coverage, and encourage richer embodied behaviors beyond conservative rollout regimes. By unifying Reward as an Agent with DynDiff-GRPO, we enable RL on a more reliable reward foundation with substantially diversified sampling, effectively mitigating reward hacking while yielding significant accuracy gains across multiple open-source world models, thereby demonstrating that broader exploration can scale successfully when grounded in robust verification.

13:00 JSTエージェント

大規模なエンタープライズ AI 向けの自律的なイベント駆動型マルチエージェント オーケストレーション

Enterprise AI は、専門エージェント全体での継続的なイベントの監視、検出、アクションを目指していますが、既存のマルチエージェント システムは主に個別の要求と応答のワークフローを想定しており、エンタープライズ規模ではまだ十分に検討されていません。当社では、ペルソナ (エージェント 10 人未満)、部門 (20 ~ 80 人)、エンタープライズ (200 人) の規模にわたる 208 の本番環境由来のエンタープライズ シナリオにわたって、DAG の計画、実行、および反応を評価し、優先順位の推論、関連イベントのマージ、プリエンプションによる継続的な運用のためのタスク マネージャーを導入しています。結果は、タスクの複雑さではなくスケールがオーケストレーションのパフォーマンスを支配していることを示しています。どちらのアーキテクチャも小規模では良好にパフォーマンスしますが、エンタープライズ規模ではエージェント検出のノイズが主なボトルネックとなり、単純なタスクの方が複雑なタスクよりも大幅に低下するため、パフォーマンスが低下します。 DAG の計画と実行は、より小規模な規模ではより高い精度と構造化された並列化を提供しますが、エンタープライズ規模ではオーバーヘッドが大きくなり、さらに悪化します。 ReAct は、障害を段階的に処理することでより堅牢になります。タスク マネージャーは、エンタープライズ規模で高優先度のキューの遅延を 14 ~ 75% 削減し、関連イベントの正確性を 20 パーセント以上向上させます。

原文 (English)

Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale

Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent systems largely assume discrete request-response workflows and remain underexplored at enterprise scale. We evaluate DAG Plan and Execute and ReAct across 208 production-derived enterprise scenarios spanning Persona (<10 agents), Department (20-80), and Enterprise (200) scales, and introduce a Task Manager for continuous operation via priority inference, related-event merging, and preemption. Results show that scale, not task complexity, dominates orchestration performance: both architectures perform well at small scale but degrade at enterprise scale as agent discovery noise becomes the primary bottleneck, with simple tasks degrading more sharply than complex ones. DAG Plan and Execute offers higher precision and structured parallelization at smaller scales, but its higher overhead worsens at enterprise scale; ReAct is more robust by handling failures incrementally. The Task Manager reduces high-priority queue latency by 14-75% and improves related-event correctness by over 20 percentage points at enterprise scale.

13:00 JST研究/論文DeepSeek

リーンによる定理証明のためのプロセス検証済み強化学習

検証可能な報酬からの強化学習 (RLVR) は通常、単一のバイナリ検証信号に依存していましたが、形式推論における記号証明アシスタントは、豊富できめの細かい構造化されたフィードバックを提供します。構造化されたプロセスと非構造化された報酬の間にあるこのギャップは、密度が高く健全なフィードバックの重要性を浮き彫りにしています。この研究では、リーン証明アシスタント自体が象徴的なプロセスのオラクルとして機能し、トレーニング中に結果レベルと詳細な戦術レベルの両方の検証済みフィードバックを提供できることを実証します。証明の試みは戦術シーケンスに解析され、リーンの詳細な説明により、局所的に健全なステップと最も初期に失敗したステップの両方がマークされ、型理論に根ざした密な検証者に基づいた信用シグナルが得られます。これらの構造化された報酬を、結果レベルとプロセスレベルの利点のバランスをとるファーストエラー伝播およびファーストトークンクレジット手法を備えた GRPO スタイルの強化学習目標に組み込みます。 STP-Lean と DeepSeek-Prover-V1.5 を使用した実験では、ほとんどの設定で戦術レベルの監視が結果のみのベースラインを上回り、MiniF2F や ProofNet などのベンチマークで改善が見られることが示されています。私たちの研究は、経験的な利益を超えて、より広い視点を強調しています。記号的証明アシスタントは、評価時の検証者であるだけでなく、トレーニング中にプロセスレベルの報酬のオラクルとしても機能する可能性があります。これにより、言語モデルのスケーラビリティと形式的推論のための記号検証の信頼性を組み合わせた強化学習フレームワークへの道が開かれます。

原文 (English)

Process-Verified Reinforcement Learning for Theorem Proving via Lean

While reinforcement learning from verifiable rewards (RLVR) typically has relied on a single binary verification signal, symbolic proof assistants in formal reasoning offer rich, fine-grained structured feedback. This gap between structured processes and unstructured rewards highlights the importance of feedback that is both dense and sound. In this work, we demonstrate that the Lean proof assistant itself can serve as a symbolic process oracle, supplying both outcome-level and fine-grained tactic-level verified feedback during training. Proof attempts are parsed into tactic sequences, and Lean's elaboration marks both locally sound steps and the earliest failing step, yielding dense, verifier-grounded credit signals rooted in type theory. We incorporate these structured rewards into a GRPO-style reinforcement learning objective with first-error propagation and first-token credit methods that balances outcome- and process-level advantages. Experiments with STP-Lean and DeepSeek-Prover-V1.5 show that tactic-level supervision outperforms outcome-only baselines in most settings, delivering improvements on benchmarks such as MiniF2F and ProofNet. Beyond empirical gains, our study highlights a broader perspective: symbolic proof assistants are not only verifiers at evaluation time, but can also act as process-level reward oracles during training. This opens a path toward reinforcement learning frameworks that combine the scalability of language models with the reliability of symbolic verification for formal reasoning.

13:00 JST研究/論文

フローベースの生成モデルによる残差空間の進化的最適化

生成手法によるデータ編集には通常、微分可能な目的と勾配ベースの検索が必要です。ただし、これらの前提はフローベースの設定では崩れます。フローベースの設定では、前方統合と後方統合を通じて編集が実行され、微分不可能な目標やブラックボックスの目標が含まれることがよくあります。フローベースの生成編集と進化的アルゴリズムを組み合わせることによってこのギャップに対処する、モデルに依存しないフレームワークである残差空間進化的最適化を紹介します。条件付きフロー マッチング (CFM) がインスタンス固有の残差から条件制御の要素を解きほぐすことができるという観察に基づいて、私たちのフレームワークは残差空間で直接動作し、2 つの相補的な検索レジームを分離します。自己受粉は特徴を保持した残差の洗練を通じて局所的な活用を実行し、他家受粉は異種サンプル間で残差を再結合することでより広範な探索を促進します。概念実証として、反事実生成のベンチマーク データセットである MorphoMNIST と結晶データで検証し、この探査 - 悪用分解がターゲットの配置、インスタンスの保存、多様性のバランスをとるための有用なメカニズムを提供し、画像を超えて現実世界の科学領域にまで及ぶことを実証しました。

原文 (English)

Residual-Space Evolutionary Optimization via Flow-based Generative Models

Data editing with generative methods typically requires differentiable objectives and gradient-based search. However, these assumptions break down in flow-based settings, where edits are performed through forward and backward integration and often involve non-differentiable or black-box objectives. We introduce residual-space evolutionary optimization, a model-agnostic framework that addresses this gap by combining flow-based generative editing with evolutionary algorithms. Building on the observation that conditional flow matching (CFM) can disentangle condition-controlled factors from instance-specific residuals, our framework directly operates in residual space and separates two complementary search regimes: self-pollination performs local exploitation through feature-preserving residual refinement, and cross-pollination promotes broader exploration by recombining residuals across heterogeneous samples. As a proof of concept, we validate on MorphoMNIST, a benchmark dataset for counterfactual generation, and on crystal data, demonstrating that this exploration--exploitation decomposition provides a useful mechanism for balancing target alignment, instance preservation, and diversity, and extends beyond images to real-world scientific domains.

13:00 JST研究/論文

マルチヘッド アテンション ベースの特徴抽出器とソフト アクター クリティカルの統合による積層造形における空隙率予測とプロセス パラメーターの最適化

積層造形プロセスの最適化には、気孔率などの欠陥を最小限に抑えるための正確なパラメータ制御が必要です。離散アクション空間を使用する従来の強化学習 (RL) アプローチは、収束が遅いことと局所最適化の影響を受けやすいという問題があり、高精度の製造タスクに対する有効性が制限されます。この研究では、マルチヘッド アテンション メカニズムと Soft Actor-Critic (SAC) アルゴリズムを統合する新しいアーキテクチャと組み合わせた連続アクション スペースを採用することで、これらの制限に対処しています。アテンションベースの特徴抽出機能は、低次元の入力特徴の微妙な変化を捕捉するエージェントの能力を強化し、極小値で値空間をナビゲートするためのより効果的な探索と活用のバランスを可能にします。レーザー粉体層融合における気孔率予測とプロセスパラメータの最適化に関するアプローチを検証し、DQN、PPO、TD3、バニラ SAC などの標準的な RL 法と比較して、より速い収束とより高い最終報酬値を実証します。提案された方法論は、14 エピソード以内で 322.79 の収束値を達成し、トレーニング全体を通じて安定性を維持しながら、既存のアプローチを上回ります。

原文 (English)

Multi-Head Attention-Based Feature Extractor Integration with Soft Actor-Critic for Porosity Prediction and Process Parameter Optimization in Additive Manufacturing

Additive manufacturing process optimization requires precise parameter control to minimize defects such as porosity. Traditional reinforcement learning (RL) approaches using discrete action spaces suffer from slow convergence and susceptibility to local optima, limiting their effectiveness for high-precision manufacturing tasks. This study addresses these limitations by employing a continuous action space combined with a novel architecture that integrates a multi-head attention mechanism with the Soft Actor-Critic (SAC) algorithm. The attention-based feature extractor enhances the agent's ability to capture subtle variations in low-dimensional input features, enabling more effective exploration-exploitation balance for navigating value spaces with local minima. We validate our approach on porosity prediction and process parameter optimization in laser powder bed fusion, demonstrating faster convergence and higher final reward values compared to standard RL methods including DQN, PPO, TD3, and vanilla SAC. The proposed methodology achieves a convergence value of 322.79 within 14 episodes, outperforming existing approaches while maintaining stability throughout training.

13:00 JSTエージェント研究/論文

ScaffoldAgent: オープンエンドの深層研究のためのユーティリティガイドによる動的アウトライン最適化

オープンエンド型ディープリサーチ (OEDR) では、システムが複数ラウンドの検索を通じて知識を取得し、一貫した長い形式のレポートを生成する必要があります。アウトラインは、検索、証拠の整理、生成を調整する構造的な足場として中心的な役割を果たします。しかし、既存の方法では、書く前にアウトラインを修正するか、ローカルヒューリスティックでアウトラインを調整するため、継続的な情報蓄積の下で足場のドリフトが発生したり、アウトラインの変更を評価するためのフィードバックが遅れたりします。我々は、OEDR 向けのユーティリティ主導の動的アウトライン最適化フレームワークである ScaffoldAgent を提案します。 ScaffoldAgent モデルは、展開、縮小、改訂の 3 つの操作による構造化された意思決定プロセスとして進化を概説し、レポート スキャフォールドの制御された更新を可能にします。さらに、取得利得、構造的一貫性、試行生成の品質から各アウトライン操作の下流の価値を推定するユーティリティ主導のフィードバック メカニズムが導入されています。結果として得られるユーティリティ信号は、ノードの選択、操作のスケジュール設定、および推論中の終了をガイドします。 DeepResearch Bench と DeepResearch Gym での実験では、ScaffoldAgent が既存のディープ リサーチ エージェントに比べて長文レポートの生成と事実に基づく根拠を一貫して向上させていることが示されています。

原文 (English)

ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research

Open-ended deep research (OEDR) requires systems to acquire knowledge through multi-round retrieval and generate coherent long-form reports. The outline plays a central role as a structural scaffold that coordinates retrieval, evidence organization, and generation. However, existing methods either fix the outline before writing or refine it with local heuristics, leading to scaffold drift under continuous information accumulation and delayed feedback for evaluating outline modifications. We propose ScaffoldAgent, a utility-guided dynamic outline optimization framework for OEDR. ScaffoldAgent models outline evolution as a structured decision process with three operations: Expansion, Contraction, and Revision, enabling controlled updates to the report scaffold. It further introduces a utility-guided feedback mechanism that estimates the downstream value of each outline operation from retrieval gain, structural coherence, and trial-generation quality. The resulting utility signal guides node selection, operation scheduling, and termination during inference. Experiments on DeepResearch Bench and DeepResearch Gym show that ScaffoldAgent consistently improves long-form report generation and factual grounding over existing deep research agents.

13:00 JSTLLM/生成AI

促すことを学ぶ: アダプティブ LLM ベースの高校個別指導で生徒の関与を向上

LLM は教育を個別化できますが、現在の静的即時個別指導システムは多様な学問分野に適応するのに苦労しています。私たちは、生のトランスクリプトから抽出した 14 の教育的特徴 (家庭教師の足場、生徒の理解など) に基づいて、主題を意識したプロンプトを備えたシステムを開発およびテストします。まずシミュレーション環境でプロンプト ルーティング モデルをトレーニングし、次にそれを実際の高校生のオンライン適応に展開します。シミュレーション ベンチマークでは、ルーターが 2 つの静的ベースライン ($0.694$ 対 $0.647$ および $0.64$、$p<0.001$) を上回るパフォーマンスを示しています。 A/B テスト (359 人の生徒からの $N=656$ の会話) では、モデルが分析学習戦略から足場学習戦略に切り替わる、シミュレーションから現実への移行が示されています。私たちの適応型プロンプト選択メカニズムは、指導効率を向上させ、教育の質を維持し、インタラクションを約 3 ターン ($p=0.007$) 削減します。グリーディ ルーターはベースラインと同等のエクササイズ コンバージョン率 ($19.1\%$ 対 $19.6\%$) を達成しますが、戦略をサンプリングするストキャスティック ルーターはより高いコンバージョン率 ($28.1\%$) をもたらします。

原文 (English)

Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring

LLMs can personalize education, although current static-prompt tutoring systems struggle to adapt to diverse academic disciplines. We develop and test a system with subject-aware prompting, based on 14 pedagogical features (e.g., tutor scaffolding, student understanding) extracted from raw transcripts. We first train a prompt routing model in a simulation environment, and then deploy it for online adaptation with actual high-school students. The simulation benchmark shows the router outperforming two static baselines ($0.694$ vs. $0.647$ and $0.64$, $p<0.001$). A/B testing ($N=656$ conversations from 359 students) shows sim-to-real transfer where the model switches from analytical to scaffolding learning strategies. Our adaptive prompt selection mechanism improves instructional efficiency, maintains pedagogical quality and reduces interactions by around 3 turns ($p=0.007$). While a greedy router achieves a comparable exercise conversion rate with the baseline ($19.1\%$ vs. $19.6\%$), a stochastic router that samples strategies leads to a higher conversion rate ($28.1\%$).

13:00 JSTエージェント研究/論文

RACL: 継続的メタヒューリスティック学習のための推論エージェント制御層

このペーパーでは、メタヒューリスティックのための推論エージェント制御層である RACL を紹介します。 RACL は、既存のオプティマイザーの上に推論エージェントを配置します。エージェントはオプティマイザを置き換えたり、ビジネス制約を変更したりしません。代わりに、操作メモリの観察、過去の動作の推論、限定された仮説の策定、介入のテスト、結果の評価、ガードレールの適用、有用なポリシーの統合、およびその決定の説明によって、オプティマイザの内部検索動作を制御します。この実験では車両ルーティングをテストベッドとして使用しますが、貢献するのは新しいルーティング ソルバー、特定の ALNS 構成、または特定の一連のルーティング ルールではありません。貢献するのは RACL メソッドです。これは、推論エージェントがメタヒューリスティックのアルゴリズム制御ルールを発見、検証、統合、説明するための方法です。現在の実験設定では、RACL は、21 の実行可能なケースのうち 21 で動作メモリ ポリシーを改善または結合し、21 の実行可能なケースのうち 18 で非推論の停滞トリガー ポリシーを改善または結合し、平均 RACL 対 STP コスト デルタは -0.641% でした。 Sevilla-9/10 ランタイム サンプルでは、​​RACL は重大な計算オーバーヘッドを示さずに、平均コストを固定と比較して -8.337%、STP と比較して -1.605% 改善しました。概念実証中、Codex は、実行を観察し、ログを解釈し、ライブの制限付き介入を提案するループ内推論エージェントとして使用されました。政策代理はその後、定量的評価を再現可能にする目的でのみ使用されました。

原文 (English)

RACL: Reasoning-Agent Control Layers for Continuous Metaheuristic Learning

This paper introduces RACL, a Reasoning-Agent Control Layer for metaheuristics. RACL places a reasoning agent above an existing optimizer. The agent does not replace the optimizer and does not modify business constraints. Instead, it controls the optimizer's internal search behavior by observing operational memory, reasoning over past behavior, formulating bounded hypotheses, testing interventions, evaluating outcomes, applying guardrails, consolidating useful policies and explaining its decisions. The experiment uses vehicle routing as a testbed, but the contribution is not a new routing solver, a particular ALNS configuration or a specific set of routing rules. The contribution is the RACL method: a way for a reasoning agent to discover, validate, consolidate and explain algorithmic control rules for a metaheuristic. In the current experimental setting, RACL improves or ties the Operational Memory Policy in 21 of 21 feasible cases and improves or ties a non-reasoning Stagnation-Triggered Policy in 18 of 21 feasible cases, with an average RACL vs STP cost delta of -0.641%. In the Sevilla-9/10 runtime sample, RACL improves average cost by -8.337% versus Fixed and -1.605% versus STP without showing material computational overhead. During the proof-of-concept, Codex was used as an in-the-loop reasoning agent observing executions, interpreting logs and proposing live bounded interventions. The policy proxy was later used only to make quantitative evaluation reproducible.

13:00 JSTLLM/生成AI研究/論文

BIM-Edit: IFC ベースのビルディング インフォメーション モデリングのための大規模言語モデルのベンチマーク

大規模言語モデル (LLM) は、テキストの指示から設計アーティファクトを生成するために、コンピュータ支援設計 (CAD) にますます適用されています。エンジニアリングの実践では、これには新しいジオメトリを作成するだけではなく、モデルが既存のシーンを理解し、正しく編集し、セマンティクスと関係を保持する必要もあります。ただし、多くの CAD ベンチマークは、既存のモデルを編集するのではなく、新しいモデルを作成することに重点を置き、主に幾何学的正確さを評価します。 Industry Foundation Classes (IFC) 形式で表される Building Information Model (BIM) の自然言語編集に関する LLM を評価するためのベンチマークである BIM-Edit を紹介します。 BIM は、建築モデルがジオメトリをセマンティックおよびリレーショナル構造とともにエンコードするため、困難なテストベッドを提供します。 BIM-Edit には、11 の現実的な建築モデルと 36 の合成シーンにわたる 324 の編集タスクが含まれています。タスクは 3 つの命令カテゴリ (直接、空間、トポロジカル) を使用して表現され、明示的な編集とシーンに基づいた編集の両方をカバーします。私たちは、幾何学的精度、意味論的妥当性、トポロジー的一貫性という 3 つの次元に沿って出力を評価します。評価された LLM 全体で、最もパフォーマンスの高いモデルは、3 つの指標全体で 49.5% の平均スコアしか達成できず、タスクの 3.4% を超える問題を完全に解決するモデルはありません。これらの結果は、現在の LLM 機能と構造化エンジニアリング設計ワークフローの要件との間に大きなギャップがあることを示しています。

原文 (English)

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engineering practice, this requires more than creating new geometry, models must also understand existing scenes, edit them correctly, and preserve semantics and relations. However, many CAD benchmarks focus on creating new models rather than editing existing ones, and mostly evaluate geometric correctness. We introduce BIM-Edit, a benchmark for evaluating LLMs on natural-language editing of Building Information Models (BIM) represented in the Industry Foundation Classes (IFC) format. BIM provides a challenging testbed because building models encode geometry together with semantic and relational structure. BIM-Edit contains 324 editing tasks spanning 11 realistic building models and 36 synthetic scenes. Tasks are expressed using three instruction categories - direct, spatial, and topological - covering both explicit and scene-grounded edits. We evaluate outputs along three dimensions: geometric accuracy, semantic validity, and topological consistency. Across evaluated LLMs, the best-performing model achieves only 49.5% average score across the three metrics, and no model fully solves more than 3.4% of tasks. These results demonstrate a substantial gap between current LLM capabilities and the requirements of structured engineering design workflows.

13:00 JST研究/論文

一般化されたPINN向けのモジュールフリーの競合回避トレーニング

物理情報に基づいたニューラル ネットワーク (PINN) は、微分可能な目的に物理法則を埋め込むことで偏微分方程式を解くための強力なフレームワークになりました。進歩にもかかわらず、PINN のトレーニングは脆弱なままです。最近の競合回避最適化スキームは残差損失と境界損失の間の勾配干渉を軽減しますが、モデルの容量が増加するにつれてその有効性が低下することを示しました。この論文では、過パラメータ化されたネットワークが機能モジュール化され、目的を越えた相互作用を抑制し、パレート定常点への収束を妨げるタスク専用モジュールに自己分割する、容量に起因する故障モードを特定します。この問題に対処するために、私たちは新しいフレームワークである Modular-Sparsity Synchronization (ModSync) を提案します。これは、相互作用を促進する経路を維持しながらタスク排他的な接続にペナルティを与えることにより、構造の最適化を競合回避トレーニングに統合します。さまざまな PDE ベンチマークにわたる広範な実験により、ModSync が容量主導の障害を一貫して防止し、堅牢な目的間の結合を維持し、最先端の精度を達成できることが実証されました。コードは \url{https://github.com/heejokong/ModSync} で入手できます。

原文 (English)

Modularity-Free Conflict-Averse Training for Generalized PINNs

Physics-informed neural networks (PINNs) have become a powerful framework for solving PDEs by embedding physical laws into differentiable objectives. Despite their advances, training PINNs remains fragile: recent conflict-averse optimization schemes alleviate gradient interference between residual and boundary losses, but we show that their effectiveness deteriorates as model capacity increases. In this paper, we identify a capacity-induced failure mode, where overparameterized networks undergo functional modularity, self-partitioning into task-exclusive modules that suppress cross-objective interaction and hinder convergence toward Pareto-stationary points. To address this issue, we propose a novel framework, Modular-Sparsity Synchronization (ModSync), which integrates structural optimization into conflict-averse training by penalizing task-exclusive connections while preserving interaction-promoting pathways. Extensive experiments across diverse PDE benchmarks demonstrate that ModSync consistently prevents capacity-driven failures, sustains robust cross-objective coupling, and achieves state-of-the-art accuracy. Codes are available at \url{https://github.com/heejokong/ModSync}.

13:00 JST研究/論文

ハイパーグラフ推論に基づく暗黙的なセマンティックを意識した通信

意味を意識した通信は、次世代通信システムの革新的なパラダイムとして登場し、基本的な目標をビットレベルのシンボルの送信から、情報の意味を確実に回復して理解することに移行しました。これまでの研究では、ソースメッセージの意味内容をグラフベースの構造として表現すると、受信側での通信効率と意味推論の精度が大幅に向上することが実証されています。ただし、既存のソリューションは通常、ペアごとの関係のみをキャプチャするグラフを採用しているため、グループ相互作用、複数エンティティの関連付け、複雑な関係コンテキストなど、現実世界のシナリオで一般的に観察される高次の暗黙的な相関関係が無視されています。この制限により、意味論的な表現力が低下し、特にノイズが多いまたは破損したチャネル条件下では、意味論的な推論があいまいになり、パフォーマンスが低下しやすくなります。これらの問題に対処するために、この論文では、ハイパーグラフを利用して意味論的知識エンティティ間の複雑な複数エンティティ関係を表現する、新しいハイパーグラフベースの暗黙的意味論的推論フレームワーク HISR を提案します。 HISR では、エンティティとそれに関連する高次の関係が、個別の関係コンテキストに合わせて調整された専用の意味論的部分空間にマッピングされます。この設計は、多様な意味論的相互作用を解きほぐして、従来のグラフ埋め込み手法によく見られる過度の平滑化効果を軽減するだけでなく、送信中に部分的な情報損失が発生した場合でも、堅牢な意味論的推論を可能にします。数値結果は、提案された HISR が、最先端のベンチマークと比較して、暗黙的な意味解釈の精度で最大 36.6% の向上を達成することを示しています。

原文 (English)

Implicit Semantic-Aware Communication Based on Hypergraph Reasoning

Semantic-aware communication has emerged as a transformative paradigm for next-generation communication systems, shifting the fundamental goal from transmitting bit-level symbols to reliably recovering and understanding the semantic meaning of information. Previous studies have demonstrated that representing the semantic content of source messages as graph-based structures can significantly improve communication efficiency and the accuracy of semantic inference at the receiver. However, existing solutions typically employ graphs that capture only pairwise relationships, thereby neglecting higher-order implicit correlations commonly observed in real-world scenarios, such as group interactions, multi-entity associations, and complex relational contexts. This limitation reduces semantic expressiveness and makes semantic inference susceptible to ambiguity and performance degradation, particularly under noisy or corrupted channel conditions. To address these issues, this paper proposes a novel hypergraph-based implicit semantic reasoning framework, HISR, which leverages hypergraphs to represent complex multi-entity relationships among semantic knowledge entities. In HISR, entities and their associated higher-order relations are mapped into dedicated semantic subspaces tailored to distinct relational contexts. This design not only disentangles diverse semantic interactions to mitigate the over-smoothing effects commonly found in traditional graph embedding methods but also enables robust semantic inference even when partial information loss occurs during transmission. Numerical results show that the proposed HISR achieves up to a 36.6% improvement in implicit semantic interpretation accuracy over the state-of-the-art benchmarks.

13:00 JSTLLM/生成AI

大規模な言語モデルの見かけの心理学的プロファイルは主に測定結果である

人間用に設計された心理学的機器は、ユーザビリティ、安全性評価、および研究への人間の参加者の代理としての使用に影響を与える安定した心理的プロファイルを大規模言語モデル (LLM) に割り当てるために使用されることが増えています。正式な心理測定フレームワークを使用して、これらのプロファイルの大部分が測定アーチファクトであることを示します。自己報告と行動課題にわたる一連の性格およびリスク選好の測定器を、大規模な人間の参照サンプルとともに 56 の指導調整 LLM に投与したところ、4 つの調査結果が報告されました。まず、モデル間の違いは、機器が対象とする特性によって決まるのではなく、指向性反応の偏り、つまりアイテムの内容に関係なく、スケールの一端または 1 つのラベル付きオプションに向かって反応する傾向によって決まります。分散分解では、モデル間の変動の 81 ~ 90% がこのバイアスによるものであると考えられますが、人間では 9 ~ 16% です。第二に、バイアスはモデルの能力とともに減少しますが、それによって除去されるわけではありません。第三に、特性ではなくバイアスが反応を促すため、機器の見かけの信頼性はほぼ完全にその応答の直交性によって予測されます。この直交性とは、特性とバイアスが逆の方向を向いている項目の割合を表す造語です。 4つ目は、使用するアイテムによってモデルの外観が変化し、アイテムの選択によって製造できることです。これらの結果は、LLM の見かけの心理的プロファイルは、モデル自体の特性ではなく、LLM を測定するために使用された機器のアーチファクトであることを示しています。人間の心理学から借用した手段が完全に直交することはほとんどなく、本質的に LLM の妥当性を欠いている可能性があるため、応答の直交性を中心とした専用の評価が必要です。

原文 (English)

Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles that affect their usability, safety assessment, and use as proxies for human participants in research. Using a formal psychometric framework, we show that these profiles are largely a measurement artifact. Administering a battery of personality and risk-preference instruments spanning self-reports and behavioral tasks to 56 instruction-tuned LLMs alongside large human reference samples, we report four findings. First, differences between models are driven not by the traits an instrument targets but by a directional response bias, a tendency to respond toward one end of the scale, or one labeled option, regardless of item content; a variance decomposition attributes 81-90% of between-model variation to this bias, against 9-16% in humans. Second, the bias declines with model capability but is not eliminated by it. Third, because bias rather than trait drives responding, an instrument's apparent reliability is almost entirely predicted by its response orthogonality, a term we coin for the proportion of items for which trait and bias point in opposite directions. Fourth, the profile a model appears to have shifts with the items used and can be manufactured through item selection. These results demonstrate that the apparent psychological profiles of LLMs are artifacts of the instrument used to measure them, not properties of the models themselves. As instruments borrowed from human psychology are rarely fully orthogonal and may inherently lack validity for LLMs, we call for dedicated assessments centered on response orthogonality.

13:00 JST研究/論文

精度を超えて: 予測モデルの論理的準拠性の測定

機械学習モデルは主に、ランキング品質、予測誤差、分類精度などの予測パフォーマンス メトリクスを通じて評価されます。これらのメトリックは、予測がグラウンド トゥルースとどの程度一致するかを効果的に定量化しますが、モデルの出力が事前定義された論理制約またはドメイン固有の制約を尊重しているかどうかは評価しません。ヘルスケア、金融、自律システムなど、一か八かのアプリケーションでは、論理的一貫性が予測精度と同じくらい重要になる可能性がありますが、この次元を捉えた標準的な指標はありません。ルール違反スコア (RVS) を導入します。これは、予測精度とは関係なく、予測モデルが特定の論理ルールのセットをどの程度尊重するかを定量化する補完的な評価指標です。 RVS は、ハード ルール (厳密な制約) とソフト ルール (統計的規則性) を別々に扱い、任意のデータセットおよびリレーショナル語彙で表現された任意の予測モデルで評価でき、ホーン ルール用に自動的に生成される SQL クエリを使用して計算できます。 RVS はモデルを評価するだけでなく、トレーニング データセットの論理的一貫性も評価し、不十分に定義されたルールを特定するのに役立ちます。 RVS は、ルールベース、埋め込みベース、ニューロシンボリック予測モデルを含む、ナレッジ グラフ リンク予測とリレーショナル回帰をカバーする 3 つのベンチマークで評価されます。私たちの結果は、同等の予測精度を達成する 2 つのモデルが、実質的に異なるレベルの論理準拠を示す可能性があることを示しており、標準的なメトリクスでは捉えることができないモデルの動作の違いが明らかになります。

原文 (English)

Beyond Accuracy: Measuring Logical Compliance of Predictive Models

Machine learning models are predominantly evaluated through predictive performance metrics such as ranking quality, prediction error, or classification accuracy. While these metrics effectively quantify how closely predictions match the ground truth, they do not assess whether model outputs respect predefined logical or domain-specific constraints. In high-stakes applications, including healthcare, finance, and autonomous systems, logical consistency can be as critical as predictive accuracy, yet no standard metric captures this dimension. We introduce the Rule Violation Score (RVS), a complementary evaluation metric that quantifies the extent to which a predictive model respects a given set of logical rules, independently of predictive accuracy. RVS treats hard rules (strict constraints) and soft rules (statistical regularities) differently, can be evaluated on any dataset and on any predictive model expressed over a relational vocabulary, and can be computed using SQL queries that are automatically generated for Horn rules. Beyond evaluating models, RVS can also evaluate the logical consistency of training datasets and help identify poorly defined rules. We evaluate RVS on three benchmarks covering knowledge graph link prediction and relational regression, including rule-based, embedding-based, and neuro-symbolic predictive models. Our results demonstrate that two models achieving comparable predictive accuracy can exhibit substantially different levels of logical compliance, revealing differences in model behavior that standard metrics fail to capture.

13:00 JST研究/論文

深層強化学習によるゲーム AI の強化

ビデオ ゲームへの没入度は、グラフィックス、オーディオ、ゲームの仕組みだけでなく、ゲーム内のキャラクターの品質にも左右されます。手作業でコード化されたシステムでは動作の複雑さを捉えるのが難しいため、信頼できるキャラクター、つまりゲーム AI を作成することは依然として大きな課題です。ゲーム AI は没入感とエンゲージメントの源です。ただし、ゲーム AI を作成する際の課題から生じる制限は、多くの場合、フラストレーションやゲーム内のリアリズムの幻想の破壊につながります。機械学習モデルの導入により、ゲーム内でより信頼性があり、本物で、共感できるキャラクターを作成するための扉が開かれます。約束されているのは、彼らがゲームとのインタラクションから、またはプレイヤーのデータから学び、真の人間らしい行動を身につけるということです。この論文では、将来的にはゲーム AI に対する強化学習のさらなる応用を想定しています。これを実現するには、現在の研究の制限により、ゲーム ジャンルを超えた広範な展開が不可能になります。したがって、ゲーム AI とゲーム開発に適した一連の要件を念頭に置いて、強化学習モデルをトレーニングするためのフレームワークを提案します。強化学習で拡張されたゲーム AI を使用したゲームの例を示し、最新のゲームにプレイヤー向けの機械学習エージェントを導入する実用性について説明します。さらに、これらの分野におけるボトルネックと困難な問題を特定し、ビデオゲーム業界のゲーム AI における機械学習の導入を加速する有望な研究の方向性を提供すると考えています。

原文 (English)

Augmenting Game AI with Deep Reinforcement Learning

Immersion in video games depends not only on graphics, audio, and game mechanics, but also on the quality of in-game characters. Producing believable characters, or game AI, remains a significant challenge as behavioral complexity is hard to capture with hand-coded systems. Game AI is a source of immersion and engagement; however, the limitations stemming from the challenges of creating game AI often lead to frustration and the breaking of the illusion of realism within the game. The introduction of machine learning models opens the door to creating more believable, authentic, and relatable characters in games. The promise is that they either learn from interacting with the game, or from player data, to develop true human-like behavior. In this paper, we envision more applications of reinforcement learning for game AI in the future. For this to materialize, current research limitations are prohibitive to broad deployment across game genres. Therefore, we propose a framework for training reinforcement learning models with a set of requirements in mind that are suited towards game AI and game development. We present examples of games with reinforcement learning-augmented game AI and describe the practicalities of deploying player-facing machine learning agents in modern games. Furthermore, we identify bottlenecks and hard problems in these areas, which we believe offer promising research directions to accelerate the adoption of machine learning in game AI for the video game industry.

13:00 JSTLLM/生成AI研究/論文

QMFOL: 定量化可能なモナディック一次論理テスト ケース生成による大規模言語モデル推論のベンチマーク

大規模言語モデル (LLM) は、推論、特に一か八かの意思決定に重要な演繹的推論において大きな進歩を遂げました。モデルが改善されるにつれて、評価ベンチマークもそれに合わせて進化する必要があります。しかし、既存のベンチマークには論理的な複雑さに対するきめ細かい制御が欠けており、意味論的な多様性と論理的な一貫性のバランスを取るのに苦労しています。これらの問題に対処するために、定量化可能で制御可能な複雑さを備えたモナディック一次論理推論タスクを生成するための自動フレームワークである QMFOL を提案します。論理積パターンと論理和パターンを使用して正式な論理構造を構築し、推論の深さ、幅、ラベルの種類、および注意をそらす要素を正確に制御できるようにします。これらの構造は、LLM を介して自然言語に翻訳され、外部証明者を使用した往復検証を通じて論理的一貫性が保証されます。私たちのフレームワークに基づいて、さまざまな論理的およびセマンティックな次元にわたる 960 の構成を持つ 2880 のインスタンスで構成されるベンチマークである QMFOLBench を構築します。 6 つの大規模推論モデル (LRM) と 2 つの LLM を評価したところ、論理的な複雑さが増すにつれてパフォーマンスが低下し、計算オーバーヘッドが増加することがわかりました。モデルは、False または Unknown のタスクよりも True のラベルが付けられたタスクの方がパフォーマンスが高く、セマンティックの変動に対して敏感です。全体として、QMFOL は、制御可能な複雑さを備えた演繹的推論ベンチマークを構築するためのスケーラブルで信頼性の高いアプローチを提供し、最新の言語モデルにおける推論機能のより正確な評価を可能にします。

原文 (English)

QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation

Large Language Models (LLMs) have made significant progress in reasoning, particularly in deductive reasoning, which is crucial for high-stakes decision-making. As models improve, evaluation benchmarks should evolve to keep pace. However, existing benchmarks lack fine-grained control over logical complexity and struggle to balance semantic diversity with logical consistency. To address these issues, we propose QMFOL, an automated framework for generating monadic first-order logic reasoning tasks with quantifiable and controllable complexity. It constructs formal logical structures using conjunction and disjunction patterns, enabling precise control over reasoning depth, width, label types, and distractors. These structures are then translated into natural language via LLMs, with logical consistency ensured through round-trip verification using an external prover. Based on our framework, we build QMFOLBench, a benchmark comprising 2880 instances with 960 configurations across diverse logical and semantic dimensions. Evaluations on six large reasoning models (LRMs) and two LLMs show that performance degrades and computational overhead increases with rising logical complexity. Models perform better on True-labeled tasks than on False or Unknown ones, and exhibit sensitivity to semantic variation. Overall, QMFOL offers a scalable and reliable approach for constructing deductive reasoning benchmarks with controllable complexity, enabling more precise evaluation of reasoning capabilities in modern language models.

13:00 JST研究/論文

知性の熱力学的尺度

知能は測定できるのでしょうか?私たちは、インテリジェンスは稀ではあるが有効な未来の合法的な増幅として定義できると提案します。つまり、システムは、受動的ダイナミクスの下ではありそうにないが、領域の制約の下では許容可能な結果の確率を高めます。インテリジェント システムは世界とその中の独自の場所をモデル化する必要があるという前提から始めます。システムはそれがモデル化する世界の一部であるため、これは自然に再帰的自己シミュレーションにつながります。つまり、システムは、それ自体のアクションが軌道の一部である未来を表します。私たちの中心的な結果は、このアーキテクチャを希少有効先物の合法的増幅の正確な熱力学的測定に結び付ける必要性ステートメントと条件付きのほぼ十分性ステートメントを提供します。内部シミュレーションが高忠実度で希少有効先物を特定しない限り、高いレア有効リフトは不可能です。逆に、まれに有効な忠実度が高く、シミュレーションに有効なポリシーが含まれている場合、達成可能なリフトは作動が制限された最適値に近づきます。したがって、再帰的自己シミュレーションは、単に知能のもっともらしい特徴であるだけでなく、述べられた仮定の下では、高度な熱力学知能にとって必要かつほぼ十分である。結果として得られるフレームワークにより、受動的な物質やフィードバックのコントローラー、大規模な言語モデル、テキスト生成者としての人間からマクスウェルの悪魔のような情報エンジンに至るまで、普遍的なスケールでインテリジェンスを測定できるようになります。

原文 (English)

Thermodynamic Measure of Intelligence

Can intelligence be measured? We propose that intelligence can be defined as the lawful amplification of rare but valid futures: a system increases the probability of outcomes that would be unlikely under passive dynamics but remain admissible under the constraints of the domain. We start with the premise that an intelligent system must model the world and its own place within it. Because the system is part of the world it models, this leads naturally to recursive self-simulation: the system represents futures in which its own actions are part of the trajectory. Our central results give a necessity statement and a conditional near-sufficiency statement connecting this architecture to a precise thermodynamic measure of lawful amplification of rare-valid futures: high rare-valid lift is impossible unless the internal simulation identifies rare-valid futures with high fidelity; conversely, when rare-valid fidelity is high and the simulation contains an effective policy, the achievable lift approaches the actuation-limited optimum. Thus recursive self-simulation is not merely a plausible feature of intelligence but, under the stated assumptions, is necessary and nearly sufficient for high thermodynamic intelligence. The resulting framework makes intelligence measurable on a universal scale, from passive matter and feedback controllers, large language models, and humans as text generators to Maxwell-demon-like information engines.

13:00 JSTエージェント

複数目的の制約付き最適化のためのマルチエージェント システム

コンピューティングおよびネットワーキング システムにおける多くの意思決定の問題は、パフォーマンスの制約の下でコスト最小化の問題として自然に定式化できます。動的環境では、ラグランジュにヒントを得た定式化に従って、重み付けされたペナルティ項を通じてコストと制約違反の両方を単一のスカラー報酬に埋め込むことで、実行時にこのような問題を解決するために強化学習 (RL) がよく使用されます。ただし、このコンテキストでは、学習されたポリシーの動作は、通常は手動で選択されるこれらの重みの選択に大きく依存します。このため、特に相対的な重要性が変化する可能性がある非定常環境では、主な目的の最適化と制約違反の効果的な回避との間の適切なトレードオフを特定することが困難になります。この論文では、マルチエージェント RL を通じてこのバランス問題に取り組むアプローチである MAMO (Multi-Agent system for Multi-Objective Constrained Optimization) を紹介します。 MAMO は、学習問題として報酬の重みの選択を定式化することで、タスクの実行を目的の設計から切り離し、動的環境における制約付き最適化問題に対する、より自律的で堅牢な RL ベースのソリューションに向けた最初のステップを提供します。

原文 (English)

A Multi-Agent system for Multi-Objective constrained optimization

Many decision-making problems in computing and networking systems can be naturally formulated as cost-minimization problems under performance constraints. In dynamic environments, reinforcement learning (RL) is often used to solve such problems at runtime by embedding both costs and constraint violations into a single scalar reward through weighted penalty terms, following a Lagrangian-inspired formulation. However, in this context the behavior of the learned policy critically depends on the choice of these weights, which are typically selected manually. This makes it difficult to identify an appropriate trade-off between optimizing the primary objective and effectively avoiding constraint violations, particularly in non-stationary environments where their relative importance may change. This paper presents MAMO (Multi-Agent system for Multi-Objective constrained optimization), an approach to tackle this balancing problem through multi-agent RL. MAMO decouples task execution from objective design by formulating the selection of reward weights as a learning problem, providing a !rst step towards more autonomous and robust RL-based solutions for constrained optimization problems in dynamic environments.

13:00 JSTLLM/生成AI

信頼性の低いパラメトリック知識とコンテキスト知識のナビゲート: LLM 推論のための明示的知識の競合解決

大規模言語モデル (LLM) は、広範なパラメトリック知識とコンテキスト内学習能力の両方を活用することで、幅広い言語ベースのタスクにわたって強力なパフォーマンスを達成し、入力プロンプトで提供される外部情報を組み込むことができます。ただし、外部知識を統合すると、モデルの内部パラメトリック知識と外部情報の間だけでなく、複数の外部コンテキスト間でも競合が発生する可能性があります。既存のアプローチは通常、モデルまたは提供されたコンテキストのいずれかが信頼できると想定し、両方のソースにエラーが含まれている可能性を無視し、不整合を積極的に解決するのではなく、一方のソースを他方のソースよりも優先することで競合を回避します。これらの制限に対処するために、従来の二者選択パラダイムを超え、マルチエージェント推論アプローチに基づく明示的な競合解決メカニズムを組み込んだ、LLM 知識競合解決のための新しいフレームワーク MACR を提案します。具体的には、まず、修正された意味論的エントロピー尺度を使用して、特定のクエリに対する LLM の応答の信頼性を定量化する、適応的な知識評価および検索アプローチを提案します。この信頼度推定に基づいて、MACR はモデルの内部知識をテキスト表現として外部化するか、内部知識が不十分な場合に関連する外部知識を取得して、その後の推論のための基本的なコンテキストを生成します。次に、それぞれ明示的なルールを誘導し、潜在的な競合を分析し、利用可能なすべてのコンテキストにわたる不一致を解決する 3 つの特殊なエージェントを備えた帰納的マルチエージェント推論フレームワークを導入します。実証結果は、MACR がベンチマーク全体で最先端のベースラインを大幅に上回るパフォーマンスを示し、同時に明示的な競合の解釈可能な解決も提供することを示しています。

原文 (English)

Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference

Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive parametric knowledge and in-context learning ability, enabling them to incorporate external information provided in the input prompt. However, the integration of external knowledge can introduce conflicts, not only between the model's internal parametric knowledge and the external information, but also among multiple pieces of external contexts. Existing approaches typically assume that either the model or the provided context is reliable, overlooking the possibility that both sources may contain errors, and avoid conflicts by privileging one source over the other, rather than actively resolving inconsistencies. To address these limitations, we propose a novel framework MACR for LLM knowledge conflict resolution that moves beyond the conventional binary choice paradigm and incorporates an explicit conflict-resolution mechanism based on a multi-agent reasoning approach. Specifically, we first propose an adaptive knowledge assessment and retrieval approach that employs a modified semantic entropy measure to quantify an LLM's confidence in its answer to a given query. Based on this confidence estimation, MACR either externalizes the model's internal knowledge as textual representations or retrieves relevant external knowledge when internal knowledge is insufficient, generating basic contexts for subsequent reasoning. Then we introduce an inductive multi-agent reasoning framework with three specialized agents that, respectively, induce explicit rules, analyze potential conflicts, and resolve inconsistencies across all available contexts. Empirical results demonstrate that MACR significantly outperforms state-of-the-art baselines across benchmarks, while also providing interpretable resolutions of explicit conflicts.

13:00 JST研究/論文

学生が描いた科学モデルの信頼性を意識した自動評価

生徒が作成した図面は、次世代科学標準 (NGSS) に準拠したモデリング ベースのタスクにおける学習者の概念理解を評価するために、科学教育で広く使用されています。ただし、このような図面の採点には、複雑な視覚的表現を解釈するための専門的な人間の判断が必要であり、教室環境で大規模な評価を実施および維持するにはコストがかかります。この研究では、ビジョンベースのモデルを使用して、学生が作成した科学図面の自動採点を研究します。パラメータ効率の高い適応を使用してビジョン トランスフォーマー (ViT) を評価し、テスト時間の予測分布から応答レベルの信頼性を導き出す信頼性を意識したスコアリング フレームワークを提案します。この信頼度シグナルにより、不確実なケースを人間によるレビューに延期しながら、信頼度の高い回答を自動的にスコアリングすることで、選択的な自動化が可能になります。 NGSS に準拠した 6 つの中学校評価項目に関する実験では、提案されたアプローチが自動化された範囲と採点リスクの間の実際的なトレードオフをサポートしながら、採点の信頼性を向上させることが示され、信頼できる教育評価のための信頼を意識した方法の価値が強調されています。

原文 (English)

Confidence-Aware Automated Assessment of Student-Drawn Scientific Models

Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standards (NGSS). However, scoring such drawings requires expert human judgment to interpret complex visual representations, making large-scale assessment costly to implement and sustain in classroom settings. In this work, we study automated scoring of student-generated scientific drawings using a vision-based model. We evaluate a Vision Transformer (ViT) with parameter-efficient adaptation and propose a confidence-aware scoring framework that derives response-level confidence from test-time predictive distributions. This confidence signal enables selective automation by scoring high-confidence responses automatically while deferring uncertain cases for human review. Experiments on six NGSS-aligned middle school assessment items show that the proposed approach improves scoring reliability while supporting a practical trade-off between automated coverage and scoring risk, highlighting the value of confidence-aware methods for trustworthy educational assessment.

13:00 JSTエージェント

ラグランジュ: 一般化されたエンドツーエンド運転のための、オープンボキャブラリー、エネルギーベースのスパースフレームワーク

エンドツーエンドの自動運転を複雑なオープンワールド環境に拡張するには、異常なシナリオに一般化する知覚モデルと、運動学的に有効な軌道を生成するプランナーが必要です。既存のパラダイムは、表現効率と一般化能力の間の明確な二分法に直面しています。高密度モデル (占有ネットワークなど) は、幾何学的に堅牢ではありますが、重大な計算ボトルネックを引き起こし、高レベルの意味論的推論に苦労します。逆に、スパースなクエリベースのプランナーは効率的ですが、クローズドセット定義に依存しているため、配布外 (OOD) イベントに対して脆弱になります。最近の Vision-Language-Action (VLA) モデルはオープンな語彙推論を提供しますが、その自己回帰的で離散的なトークン生成は、車両ダイナミクスの連続的で高周波の制御要件と根本的に矛盾します。これに対処するために、マスクされた潜在場 (MLF) に基づいたオープン語彙で計算量が少ない駆動フレームワークである Lagrange を提案します。ラグランジュは、高密度ボリューム再構成や閉集合クエリ メカニズムに依存するのではなく、視覚言語モデル (VLM) を利用して、クラスに依存しないオブジェクトの提案を連続的なセマンティックなビジュアル トークンにエンコードします。無関係なエンティティを時間的にフィルタリングし、空間座標上で定義された暗黙的な連続エネルギー フィールドにアテンション トークンをデコードする、インテント駆動型マスク クロス アテンション モジュールを導入します。このエネルギー場にわたるラグランジュ作用最小化問題として意思決定を組み立てることにより、衝突回避を実行しながら車両運動学への厳密な準拠を強制します。標準 (nuScenes) ベンチマークとロングテール (CODA) ベンチマークの両方での広範なオフライン評価により、ラグランジュが堅牢で解釈可能、運動学的に実現可能なオープンワールド自律性のための有望なフレームワークを確立していることが実証されました。

原文 (English)

Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving

Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories. Existing paradigms face a distinct dichotomy between representational efficiency and generalization capacity. Dense models (e.g., occupancy networks), while geometrically robust, incur critical computational bottlenecks and struggle with high-level semantic reasoning. Conversely, sparse, query-based planners are efficient but reliant on closed-set definitions, rendering them vulnerable to out-of-distribution (OOD) events. Although recent Vision-Language-Action (VLA) models offer open-vocabulary reasoning, their autoregressive, discrete token generation fundamentally conflicts with the continuous, high-frequency control requirements of vehicle dynamics. To address this, we propose Lagrange, an open-vocabulary, computationally sparse driving framework based on Masked Latent Fields (MLF). Rather than relying on dense volumetric reconstructions or closed-set query mechanisms, Lagrange exploits Vision-Language Models (VLMs) to encode class-agnostic object proposals into continuous semantic visual tokens. We introduce an intent-driven masked cross-attention module that temporally filters irrelevant entities, decoding the attended tokens into an implicit continuous energy field defined over spatial coordinates. By framing decision-making as a Lagrangian action minimization problem spanning this energy field, we enforce strict compliance with vehicle kinematics while executing collision avoidance. Extensive offline evaluations on both standard (nuScenes) and long-tail (CODA) benchmarks demonstrate that Lagrange establishes a promising framework for robust, interpretable, and kinematically feasible open-world autonomy.

13:00 JST研究/論文

システムの非線形性を利用して、インテリジェント故障診断システムの設計におけるデータ不足に取り組む

深層転移学習 (DTL) により、インテリジェント障害診断システム (IFDS) を効率的に構築できます。一方で、DTL 手法は依然として大量のラベル付きデータに大きく依存しています。機械や構造物の障害に対処する場合、このような量のデータを取得するのは困難な場合があります。この文書では、データが非常に不足している状況で DTL を使用して振動ベースの IFDS を設計する新しいアプローチを提案します。実世界のシステムの固有の非線形性を活用した周期的な複数励起レベルの手順を使用して、事前にトレーニングされた畳み込みニューラル ネットワーク (CNN) で簡単に分析して故障を診断できる画像を生成します。この論文では、IFDS の設計中に遭遇する典型的なデータ不足に対処するために、新しいデータ視覚化手法とその拡張手法を提案します。鉄道パンタグラフ構造の実験的検証は、提案された方法を効果的にサポートします。

原文 (English)

Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelligent Fault Diagnosis Systems

Deep Transfer Learning (DTL) allows for the efficient building of Intelligent Fault Diagnosis Systems (IFDS). On the other hand, DTL methods still heavily rely on large amounts of labelled data. Obtaining such an amount of data can be challenging when dealing with machines or structures faults. This document proposes a novel approach to the design of vibration-based IFDS using DTL in condition of strong data scarcity. A periodic multi-excitation level procedure leveraging intrinsic non-linearities of real-world systems is used to produce images that can be conveniently analysed by pre-trained Convolutional Neural Networks (CNNs) to diagnose faults. A new data visualization method and its augmentation technique are proposed in this paper to tackle the typical lack of data encountered during the design of IFDS. Experimental validation on a railway pantograph structure provides effective support for the proposed method.

13:00 JSTエージェント

SoftSkill: 状況に応じた適応のための行動圧縮

エージェントのスキルは通常、回答ポリシー、証拠の使用習慣、タスク手順をエンコードした自然言語のマークダウン ファイルとして展開されます。これらのファイルは読み取り可能で移植可能ですが、間接的に使用されます。タスク インスタンスごとに、凍結された言語モデルが長いテキスト アーティファクトを生成時の動作に変換する必要があります。この論文では、自然言語スキルが代わりに、ベースモデルが凍結されたままの状態で、トレーニング可能なソフトデルタによって洗練されたコンパクトな連続コンテキストオブジェクトを初期化できるかどうかを尋ねます。私たちは、そのようなソフトスキルを次のトークン予測で調整し、推論時に潜在的な行動事前分布として展開する凍結バックボーン手法である SoftSkill を提案します。メインのシングルラウンド設定では、Qwen3.5-4B の長さ 32 の SoftSkill プレフィックスは、スキルなしのプロンプトよりも SearchQA で 8.3 ポイント、LiveMath で 42.1 ポイント、DocVQA で 1.3 ポイント向上しました。 SkillOpt と比較して、SoftSkill は、数百から数千の Markdown スキル トークンを少数の仮想トークンに置き換えながら、SearchQA で 5.2 ポイント、LiveMath で 12.5 ポイント精度が向上します。私たちはさらに、より困難な境界ケースとしてエージェント実行を研究します。このケースでは、まばらな軌道の模倣は有用なシグナルを提供しますが、長期的な手続きの動作をまだ強力に圧縮していません。より広範に、この結果は、一部のタスクスキルは、推論時に再解釈される追加のマークダウンとしてではなく、フリーズされたモデルがタスクにどのように入るかを制御するコンパクトな潜在的な制御として扱う方がよいことを示唆しています。

原文 (English)

SoftSkill: Behavioral Compression for Contextual Adaptation

Agent skills are commonly deployed as natural-language Markdown files that encode answer policies, evidence-use habits, and task procedures. These files are readable and portable, but they are consumed indirectly: for each task instance, a frozen language model must translate a long textual artifact into generation-time behavior. This paper asks whether a natural-language skill can instead initialize a compact continuous context object, refined by a trainable soft delta while the base model remains frozen. We propose SoftSkill, a frozen-backbone method that tunes such soft skills with next-token prediction and deploys them as latent behavioral priors at inference time. In our main single-round setting, a length-32 SoftSkill prefix on Qwen3.5-4B improves over no-skill prompting by 8.3 points on SearchQA, 42.1 points on LiveMath, and 1.3 points on DocVQA. Relative to SkillOpt, SoftSkill improves accuracy by 5.2 points on SearchQA and 12.5 points on LiveMath, while replacing hundreds to thousands of Markdown skill tokens with a few virtual tokens. We further study agentic execution as a harder boundary case, where sparse trajectory imitation provides useful signal but does not yet robustly compress long-horizon procedural behavior. More broadly, the results suggest that some task skills are better treated not as additional Markdown to be reinterpreted at inference time, but as compact latent controls over how a frozen model enters the task.

13:00 JSTエージェント

インタラクション軌跡マイニングによるコンピュータ使用エージェント向けの SKILL.md 生成の自動化

明示的なスキル ライブラリにより、コンピュータを使用するエージェントの検査が容易になりますが、そのようなライブラリを下流のポリシーを改善する方法でインタラクション データからマイニングできるかどうかは不明のままです。私たちは、GUI の軌跡をセグメント化し、セグメントを候補スキルにクラスタリングし、結果として得られるアノテーションからスキルを意識​​したポリシーをトレーニングする 3 段階のパイプラインを通じてこの質問を研究します。マイニングされたクラスターはソース ベンチマークで読み取ることができます。8 つのクラスターのうち 5 つは、InteraSkill Workflows ラベルに対して少なくとも 0.95 の純度を持っています。ただし、可読性は転送を意味するものではありません。 GRPO は、IW スキル ステップの精度を 18.5\% から 20.5\% に向上させるだけで、BrowseComp+ は基本的に変更されず、主要なソース ドメイン メトリクスに関する自明な頻度事前分布よりもパフォーマンスが劣ります。したがって、この方法を診断研究として紹介します。軌跡マイニングは検査可能なスキル構造を明らかにできますが、現在の境界検出器、順序のないセグメント表現、およびオフライン報酬モデルは、信頼性の高いクロスドメインポリシーの改善には不十分です。

原文 (English)

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining

Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from interaction data in a way that improves downstream policies. We study this question through a three-stage pipeline that segments GUI trajectories, clusters segments into candidate skills, and trains a skill-aware policy from the resulting annotations. The mined clusters are readable on the source benchmark: five of eight clusters have at least 0.95 purity against InteraSkill Workflows labels. However, readability does not imply transfer. GRPO improves IW skill-step accuracy only from 18.5\% to 20.5\%, leaves BrowseComp+ essentially unchanged, and underperforms trivial frequency priors on key source-domain metrics. We therefore present the method as a diagnostic study: trajectory mining can expose inspectable skill structure, but the current boundary detector, orderless segment representation, and offline reward model are insufficient for reliable cross-domain policy improvement.

13:00 JSTLLM/生成AINVIDIA

LLM FP4 事前トレーニングにおける収縮バイアスの再考: 幾何学的起源、システムへの影響、および UFP4 レシピ

FP4 トレーニングでは、LLM 事前トレーニングのメモリと計算コストの大幅な削減が約束されていますが、NVIDIA Blackwell/Rubin クラス システムや AMD MI350 シリーズ GPU を含む現在の FP4 ハードウェア パスとレシピは、引き続き E2M1 データ要素を中心としています。この研究では、その選択の基本的な制限を特定します。E2M1 などの不均一フォーマットは本質的に、表現可能なビンの幾何学的非対称性によって引き起こされる系統的な負の丸め誤差である縮小バイアスの影響を受けます。このバイアスは層全体で乗算的に蓄積し、ランダム アダマール変換 (RHT) によって増幅されることを示し、既存の E2M1 ベースの FP4 レシピで観察されるトレーニングの不安定性についての統一的な説明を提供します。対照的に、均一グリッド (E1M2/INT4) は、このグリッド ジオメトリ エラーを回避し、RHT によるバケット使用率の向上をより高い量子化品質に変換します。この発見に基づいて、確率的丸めを dY のみに制限しながら、RHT を 3 つのトレーニング GEMM すべてに適用する均一な 4 ビット トレーニング レシピである UFP4 を提案します。 Dense 1.5B、MoE 7.9B、および MoE 124B の長期事前トレーニングでは、スケーリング則解析とアブレーション研究によって裏付けられたように、UFP4 は強力な E2M1 ベースのベースラインよりも低い BF16 相対損失劣化を一貫して達成しています。私たちの結果は、将来のアクセラレータが E2M1 と並んでファーストクラスのトレーニング プリミティブとして E1M2/INT4 スタイルの均一 4 ビット グリッドをサポートする必要があることを示唆しています。

原文 (English)

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements. In this study, we identify a fundamental limitation of that choice: non-uniform formats such as E2M1 inherently suffer from Shrinkage Bias, a systematic negative rounding error caused by the geometric asymmetry of their representable bins. We show that this bias accumulates multiplicatively across layers and is amplified by the Random Hadamard Transform (RHT), providing a unified explanation for the training instability observed in existing E2M1-based FP4 recipes. In contrast, uniform grids (E1M2/INT4) bypass this grid-geometry error and better convert the improved bucket utilization from RHT into higher quantization quality. Based on this finding, we propose UFP4, a uniform 4-bit training recipe that applies RHT to all three training GEMMs while restricting stochastic rounding to dY alone. On Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining, UFP4 consistently achieves lower BF16-relative loss degradation than strong E2M1-based baselines, supported by scaling-law analysis and ablation studies. Our results suggest that future accelerators should support E1M2/INT4-style uniform 4-bit grids as first-class training primitives alongside E2M1.

13:00 JST研究/論文

注意誘導型ディープラーニングによる解釈可能な精子形態分類

男性不妊はカップル不妊の主な原因であり、多くの場合、精子の形態異常に関連しています。深層学習モデルは自動分析を提供しますが、そのほとんどは解釈可能性に欠けており、臨床での採用は制限されています。この研究は、精子形態分類のための注意誘導型深層学習フレームワークを提案します。当社では、事前トレーニング済みの EfficientNet-B0 と畳み込みブロック アテンション モジュール (CBAM) を組み合わせて、精子頭部の重要な領域に焦点を当て、精度と解釈可能性の両方を向上させています。 SMIDS および HuSHem 公開データセットで評価した場合、私たちのモデルは 90.2% および 93.9% (マクロ F1 スコア 0.913 および 0.948) の精度を達成し、SimpleCNN および標準 EfficientNet-B0 を上回りました。さらに、Grad-CAM++ 視覚化を使用して、モデルの決定に影響を与える機能を強調表示します。この結果は、この正確で透明なフレームワークが、不妊治療クリニックにおける自動精子分析のための実用的なツールであることを示しています。

原文 (English)

Interpretable Sperm Morphology Classification via Attention-Guided Deep Learning

Male infertility is a major cause of couple infertility, often linked to abnormal sperm morphology. While deep learning models offer automated analysis, most lack interpretability, limiting their clinical adoption. This study proposes an attention-guided deep learning framework for sperm morphology classification. We combine a pretrained EfficientNet-B0 with a Convolutional Block Attention Module (CBAM) to focus on key areas of the sperm head, improving both accuracy and interpretability. Evaluated on the SMIDS and HuSHem public datasets, our model achieves accuracies of 90.2% and 93.9% (macro F1 scores of 0.913 and 0.948), outperforming SimpleCNN and standard EfficientNet-B0. Furthermore, we use Grad-CAM++ visualizations to highlight features influencing the model's decisions. The results demonstrate that this accurate and transparent framework is a practical tool for automated sperm analysis in fertility clinics.

13:00 JST研究/論文

体外受精検査室の環境条件のコンテキスト認識型階層ベイジアン モデリング

体外受精の妊娠率は患者レベルの変数を使用して日常的にモデル化されていますが、高解像度の実験室環境データは依然として十分に活用されていません。私たちは、これが機会損失であることを示しています。私たちは、生のセンサー平均に依存するのではなく、インキュベーターの微小環境のダイナミクスを捉える、ローリング熱安定性、温度と湿度の同時付着、ピークストレス持続時間、ストレス後の回復速度など、55 のコンテキストを認識した時間的特徴を設計します。アジアの体外受精クリニックからの 61 週間のデータでは、これらの機能により、相互検証された予測誤差が 1.27% に減少しました (生の平均では 3 ~ 5% でした)。次に、部位固有のベースラインを維持しながら、部分プールを介してアジアと北欧の診療所全体で環境への影響を共有する階層型ベイジアン ベータ回帰モデルをトレーニングします。北欧の診療所から得られたデータに基づいて、このモデルは 35 ~ 39 歳の年齢層でナイーブなベースラインと比較して R2 = 0.86 と 64% の誤差低減を達成し、構造化された環境モニタリングには臨床的に意味のある転送可能なシグナルが含まれていることを示しています。

原文 (English)

Context-Aware Hierarchical Bayesian Modeling of IVF Laboratory Environmental Conditions

IVF pregnancy rates are routinely modeled using patient-level variables, while high-resolution laboratory environmental data remain underutilized. We show that this is a missed opportunity. Rather than relying on raw sensor averages, we engineer 55 context-aware temporal features, including rolling thermal stability, simultaneous temperature-humidity adherence, peak stress duration, and post-stress recovery speed, that capture the dynamics of incubator microenvironments. On 61 weeks of data from an Asian IVF clinic, these features reduce cross-validated prediction error to 1.27%, compared to 3-5% for raw averages. We then train a hierarchical Bayesian Beta regression model that shares environmental effects across an Asian and a Northern European clinic via partial pooling, while preserving site-specific baselines. On held-out data from the Northern European clinic, the model achieves R2 = 0.86 and a 64% error reduction for the 35-39 age group over a naive baseline, demonstrating that structured environmental monitoring contains clinically meaningful, transferable signal.

13:00 JSTLLM/生成AI

安全性を重視した LLM は、混合コンプライアンスのデモンストレーションから何を学びますか?

これまでの研究では、コンテキスト内のデモンストレーションが言語モデルを脱獄できることが示されていますが、モデルがさまざまな種類のコンプライアンス デモンストレーションをどのように解釈するかは依然として不明です。私たちは、無害なコンプライアンスのデモンストレーション (有害ではない要求、有益な応答) と有害なコンプライアンスのデモンストレーション (有害な要求、有益な応答) を混合し、デモンストレーションの構成が有害なコンプライアンスをどのように推進するかについて 3 つの仮説を検証することで、これを研究します。 4 つのモデルにわたって、無害なデモンストレーションと有害なデモンストレーションには互換性がないことがわかりました。無害なデモンストレーションは、モデルに応じて有害なコンプライアンスを削減または増加させることができます。さらに、選好の最適化は、無害なデモンストレーションが有害なコンプライアンスの増加を防ぐ重要なトレーニング段階であること、デモンストレーションの順序付けが強い最新性バイアスを示していること、モデルは拒否がコンテキスト内学習とどのように相互作用するかが異なることを示します。一部のモデルは拒否時にもデモンストレーションされたフォーマットを採用しますが、他のモデルは拒否時にすべてのコンテキスト内シグナルをオーバーライドします。まとめると、この研究は、デモンストレーションベースのジェイルブレイクが機能することを示すだけでなく、その仕組みを特徴づけることに移ります。つまり、コンプライアンスデモンストレーションからどのモデルが抽出されるかは、デモンストレーションの内容、順序、トレーニング方法によって異なります。

原文 (English)

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demonstrations (non-harmful request, helpful response) with harmful compliance demonstrations (harmful request, helpful response) and testing three hypotheses about how demonstration composition drives harmful compliance. Across four models, we find that benign and harmful demonstrations are not interchangeable: benign demonstrations can either reduce or increase harmful compliance depending on the model. We further show that preference optimization is the critical training stage that prevents benign demonstrations from increasing harmful compliance, that demonstration ordering exhibits strong recency bias, and that models differ in how refusal interacts with in-context learning: some adopt demonstrated formatting even when refusing, while others override all in-context signals upon refusal. Taken together, this work moves beyond showing that demonstration-based jailbreaking works to characterizing how it works: what models extract from compliance demonstrations depends on demonstration content, ordering, and training methodology.

13:00 JSTLLM/生成AI研究/論文

マルチ LCB: LiveCodeBench を複数のプログラミング言語に拡張

LiveCodeBench (LCB) は、コード生成タスクで大規模言語モデル (LLM) を評価するためのベンチマークとして最近広く採用されています。 LCB は、競技プログラミングの問題を厳選し、常に新しい問題をセットに追加し、リリース日でフィルタリングすることにより、汚染を認識した評価を提供し、コーディング能力の全体的なビューを提供します。ただし、LCB は依然として Python に限定されており、LLM が現実のソフトウェア エンジニアリングで必要とされる多様なプログラミング言語全体に汎用化できるかどうかという問題は未解決のままです。 Python を含む 12 のプログラミング言語にわたる LLM を評価するためのベンチマークである Multi-LCB を紹介します。マルチ LCB は、LCB の汚染制御と評価プロトコルを維持しながら、LCB データセットの Python タスクを他の言語の同等のタスクに変換します。オリジナルの LCB 形式と完全な互換性があるため、Multi-LCB は将来の LCB アップデートを自動的に追跡し、言語を超えたコード生成能力の体系的な評価を可能にし、Python をはるかに上回るパフォーマンスを維持するモデルを必要とします。私たちは、Multi-LCB に関する指示と推論について 24 の LLM を評価し、Python の過剰適合、言語固有の汚染、および多言語パフォーマンスの大幅な差異の証拠を明らかにしました。私たちの結果は、Multi-LCB がマルチプログラミング言語コード評価の厳密な新しいベンチマークとして確立され、LCB の主な制限に直接対処し、現在の LLM 機能の重大なギャップを明らかにします。

原文 (English)

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on code-generation tasks. By curating competitive programming problems, constantly adding fresh problems to the set, and filtering them by release dates, LCB provides contamination-aware evaluation and offers a holistic view of coding capability. However, LCB remains restricted to Python, leaving open the question of whether LLMs can generalize across the diverse programming languages required in real-world software engineering. We introduce Multi-LCB, a benchmark for evaluating LLMs across twelve programming languages, including Python. Multi-LCB transforms Python tasks from the LCB dataset into equivalent tasks in other languages while preserving LCB's contamination controls and evaluation protocol. Because it is fully compatible with the original LCB format, Multi-LCB will automatically track future LCB updates, enabling systematic assessment of cross-language code generation competence and requiring models to sustain performance well beyond Python. We evaluated 24 LLMs for instruction and reasoning on Multi-LCB, uncovering evidence of Python overfitting, language-specific contamination, and substantial disparities in multilingual performance. Our results establish Multi-LCB as a rigorous new benchmark for multi-programming-language code evaluation, directly addressing LCB's primary limitation and exposing critical gaps in current LLM capabilities.

13:00 JST研究/論文

FlowEdit: フローマッチング TTS における生涯にわたる発音適応のための連想記憶

フローマッチングのテキスト読み上げシステムは、驚くべきゼロショット品質を達成しますが、展開後は静的のままです。語彙外の固有名詞の発音エラーは、モデルが再トレーニングされない限り持続します。 FlowEdit を紹介します。これは、重みの更新ではなく潜在的な条件付け編集として発音の修正を学習する、フローズン フロー マッチング TTS のための生涯にわたる適応フレームワークです。修正フィードバックが提供されると、FlowEdit はテキスト埋め込み空間内のトークンレベルの摂動を最適化し、コンテンツアドレス指定可能なエピソード記憶として機能する最新ホップフィールド ネットワークに修正を保存します。推論時には、類似性ゲートを使用したソフト アテンションを介して修正が取得され、ファジー形態学的マッチングが可能になります。 18 の言語ファミリーにわたる 312 の多言語固有名詞からなる厳選されたベンチマークでは、FlowEdit は、同一の一般音声品質を維持しながら、ターゲット単語の音素エラー率をゼロショット ベースラインと比較して 92.7% 削減します。修正は 1 つの GPU で約 15 秒で完了します。

原文 (English)

FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunciation errors on out-of-vocabulary proper nouns persist unless the model is retrained. We introduce FlowEdit, a life-long adaptation framework for frozen flow-matching TTS that learns pronunciation corrections as latent conditioning edits rather than weight updates. When corrective feedback is provided, FlowEdit optimizes a token-level perturbation in the text embedding space, then stores the correction in a Modern Hopfield Network serving as content-addressable episodic memory. At inference, corrections are retrieved via soft attention with a similarity gate, enabling fuzzy morphological matching. On our curated benchmark of 312 multilingual proper nouns across 18 language families, FlowEdit reduces target-word Phoneme Error Rate by 92.7% relative to the zero-shot baseline while maintaining identical general-speech quality. Corrections complete in approximately 15 seconds on a single GPU.

13:00 JST研究/論文

DeepSWIP: ニューラル確率論理プログラムの商 WMC 反事実

DeepProbLog などの神経記号システムは、神経知覚と確率論的論理を組み合わせていますが、標準的な推論は連想的です。反事実的推論には、介入と証拠の因果意味論がさらに必要です。 DeepProbLog プログラム用の単一世界の反事実セマンティクスである DeepSWIP を紹介します。ニューラル具体化を使用して、固定コンテキストのニューラル述語を通常の ProbLog の選択肢に減らし、単一世界介入プログラム (SWIP) を適用し、単一の変換されたプログラムに対して重み付けモデル カウンティング (WMC) によって反事実を計算します。有限の根拠と独自にサポートされるモデルの仮定の下では、DeepSWIP は学習された具体化された FCM に対して正確です。 ProbLog 条件文の標準商-WMC 形式は、アクティブな神経確率を特定し、介入クリーニング、キャリブレーション感度、およびまれな証拠の不安定性を説明します。 MPI3D での実験では、予測どおり、12,000 クエリに対する DeepTwin 構造に対する変換と、Twin の内生的重複を回避することによる推論の 2.14 倍の高速化が確認されました。 SUMO HOV 実験では、ニューラル キャリブレーションの劣化によりプラグイン推定値にバイアスがかかる一方、スコープが正しく設定されたランダム化ポリシー AIPW 推定器では、母平均値と ATE 推定値の一次バイアスのほとんどが除去されることが示されています。コードは https://github.com/saibib/deep_SWIP にあります。

原文 (English)

DeepSWIP: Quotient-WMC Counterfactuals for Neural Probabilistic Logic Programs

Neurosymbolic systems such as DeepProbLog combine neural perception with probabilistic logic, but standard inference is associational. Counterfactual reasoning additionally requires a causal semantics for interventions and evidence. We introduce DeepSWIP, a single-world counterfactual semantics for DeepProbLog programs. Using neural materialization, we reduce fixed-context neural predicates to ordinary ProbLog choices, apply Single World Intervention Programs (SWIPs), and compute counterfactuals by weighted model counting (WMC) over a single transformed program. Under finite grounding and unique-supported-model assumptions, DeepSWIP is exact relative to the learned materialized FCM. The standard quotient-WMC form of ProbLog conditionals identifies active neural probabilities and explains intervention cleaning, calibration sensitivity, and rare-evidence instability. Experiments on MPI3D confirm the transformation against a DeepTwin construction against 12,000 queries, as predicted and a 2.14$\times$ inference speedup from avoiding the Twin's endogenous duplication. A SUMO HOV experiment shows that neural calibration degradation biases plug-in estimates, while a correctly scoped randomized-policy AIPW estimator removes most first-order bias for population mean and ATE estimands. Code is at https://github.com/saibib/deep_SWIP.

13:00 JSTLLM/生成AIエージェント

LedgerAgent: ポリシー準拠のツール呼び出しエージェントの構造化された状態

カスタマー サービス ドメインのポリシーに準拠したツール呼び出しエージェントは、ツールを呼び出してドメイン ポリシーに従いながら、ターン全体でタスクの状態を維持する必要があります。タスクの状態は、ユーザーの対話やツールの呼び出しを通じて観察される、関連する事実、識別子、制約、および条件で構成されます。標準エージェントでは、タスクの状態は個別に表現されません。観察、ツールの返却、およびポリシーの指示がプロンプトに配置されるため、エージェントは次に何を行うかを決定するたびに、プロンプトから関連する状態を再構築する必要があります。この設計では状態管理が暗黙的に行われるため、2 つの一般的な障害モードが作成されます。エージェントは正しい事実を取得しても、後で古い情報、欠落している情報、または不正確な情報に基づいて決定を下す可能性があります。また、構文的に有効なツール呼び出しでも、現在のタスクの状態に応じてドメイン ポリシーに違反する可能性があります。 \textsc{LedgerAgent} を導入します。これは、観察されたタスクの状態を別の台帳に保持し、その状態をプロンプトに表示する、ツール呼び出しエージェントのための推論時メソッドです。この台帳は、環境を変更するツール呼び出しが実行される前に状態依存のポリシー制約をチェックするためにも使用され、ポリシー違反をブロックします。 \textsc{LedgerAgent} は、4 つの顧客サービス ドメインとオープン加重モデルとクローズ加重モデルの混合パネルにわたって、標準的なプロンプトベースのツール呼び出しアプローチよりも平均パス\textasciicircumk を向上させ、より厳格な複数トライアルの一貫性指標の下で最大の利益をもたらします。

原文 (English)

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling tools and obeying domain policies. Task states consist of relevant facts, identifiers, constraints, and conditions observed through user interaction and tool calls. In standard agents, task states are not represented separately. Observations, tool returns, and policy instructions are placed in the prompt, leaving agents to reconstruct the relevant states from the prompt each time they decide what to do next. This design makes state management implicit, creating two common failure modes. An agent may retrieve the right facts but later ground its decision in stale, missing, or incorrect information; and a syntactically valid tool call may still violate a domain policy that depends on the current task state. We introduce \textsc{LedgerAgent}, an inference-time method for tool-calling agents that maintains observed task states in a separate ledger and renders the states into the prompt. The ledger is also used to check state-dependent policy constraints before environment-changing tool calls are executed, blocking policy violations. Across four customer-service domains and a mixed panel of open- and closed-weight models, \textsc{LedgerAgent} improves average pass\textasciicircum{}k over a standard prompt-based tool-calling approach, with the largest gains under stricter multi-trial consistency metrics.

13:00 JST研究/論文

指示はどのようにスピーチを形成するのでしょうか?スタイルキャプション付きテキスト読み上げのクロスアテンションアトリビューション

スタイルキャプション付きのテキスト読み上げシステムは、自然言語を使用して音声特性を制御しますが、個々の単語が音響出力にどのように影響するかは不明のままです。これを理解することは、故障モードを診断し、表現力豊かな TTS の制御性を向上させるために重要です。我々は、DAAM フレームワークを初めて音声ドメインに適応させた音声拡散モデルのクロスアテンション アトリビューションを提案し、それを CapSpeech-TTS に適用します。私たちの方法では、25 のレイヤーと 24 の ODE ステップにわたるトークンごとのヒートマップを抽出します。私たちは、それぞれ 30 個のテキスト トランスクリプトの生成を条件付ける 120 のスタイル キャプションで構成される 3,600 の組み合わせ (スタイル キャプション、テキスト トランスクリプト) を分析し、キャプション トークンがどのように波形を形成するかを明らかにします。結果は次のことを示しています: (1) スタイル トークンはコンテンツ/機能トークンよりも時間的分散が低く、グローバルな条件付けが確認されています。 (2) スタイルへの注意は F0 およびエネルギーと相関します。 (3) スタイルコンディショニングのピークは初期ステップと深い層にあります。 (4) アテンション エントロピーはレイヤー 17 で最小値に達し、スタイル重要度のピークと同時に発生し、スタイルが最も重要な段階でネットワーク選択性が最大であることを示しています。これは、自然言語が音声拡散モデルにおける相互注意にどのような影響を与えるかについての最初の研究です。

原文 (English)

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeech-TTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models

13:00 JST研究/論文

分布シフトの下で調整された専門家の混合に向けて

キャリブレーションは、モデルの予測の不確実性をその経験的結果の頻度と一致させるものであり、報告された確率を理解し信頼するために重要です。最近の研究では、個々の予測変数のレベルでキャリブレーションを強制すると、アンサンブルの精度とキャリブレーションが向上することが示されており、特に専門家混合 (MoE) モデルが強力な経験的改善を示しています。ただし、校正が MoE に役立つ条件はよく理解されていません。この研究では、ルーティング メカニズムが専門家レベルのキャリブレーションとどのように相互作用するかに焦点を当て、分布シフトの下で MoE モデルがどのように動作するかを研究します。専門家によるキャリブレーションは、ハード配線モデルの幅広いクラスの分布シフトの下でモデル全体のキャリブレーションを確実に行うには十分ですが、ソフト配線モデルのキャリブレーションには不十分であることを示します。これに対処するために、分布シフトの下でルーティングされた集約のキャリブレーション エラーにペナルティを与える敵対的再重み付けを提案し、モデル クラス、予測タスク、分布シフト全体で、データの平均と困難なサブセットの両方で精度とキャリブレーションのトレードオフが改善されることを実証します。

原文 (English)

Toward Calibrated Mixture-of-Experts Under Distribution Shift

Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve ensemble accuracy and calibration, with mixture-of-experts (MoE) models showing strong empirical improvements in particular; however, the conditions under which calibration helps MoE are not well understood. In this work, we study how MoE models behave under distribution shift, focusing on how routing mechanisms interact with expert-level calibration. We show that expert calibration is sufficient to ensure calibration of the overall model under a broad class of distribution shifts in hard-routed models, but is insufficient for calibrating soft-routed models. To address this, we propose an adversarial reweighting that penalizes calibration errors of the routed aggregate under distribution shift, and we demonstrate that it improves the accuracy-calibration tradeoff both on average and on difficult subsets of the data, across model classes, prediction tasks, and distribution shifts.

13:00 JST画像/動画生成ロボティクス

人間の普遍的な把握

人間は物体を難なく掴むことができますが、多指ロボットはこのレベルの汎用性からは程遠いです。私たちは、ロボットが把握するデータの最も自然な情報源は、毎日何千もの物体を拾う人間からのものであると主張します。我々は、ステレオ カメラからキャプチャされた単一の RGB-D 画像内のユーザー指定のオブジェクトに対する人間の多様な把握を生成するフロー マッチング モデルである HUG を紹介します。スマート グラスを使用して、まず 1M-HUG を収集します。これは、1M フレーム (27.8 時間) にわたる人間の把握の自己中心的なデータセットであり、41 の建物にわたる 6,707 のオブジェクト インスタンスです。次に、人間の自然な握りの分布をモデル化するために、私たちの新しいフロー マッチング モデルは RGB と深度の観察を融合して、手首の移動、手首の回転、および MANO の手のポーズによってパラメータ化された握りを出力します。予測された掴みをさまざまなロボットハンドにリターゲットできるため、日常シーンでのゼロショット掴みが可能になります。評価を標準化するために、メートルスケールの 3D メッシュを使用して、5 つの幾何学的カテゴリとさまざまなサイズの 90 個の未確認オブジェクトからなる新しいシミュレートされたベンチマーク HUG-Bench を構築します。私たちは、複数のステレオカメラ、ロボットの実施形態、家庭環境にわたる HUG-Bench の 30 オブジェクト テスト セットで現実世界の HUG を評価します。 HUG は、当社の挑戦的なオブジェクト セットにおいて、最先端の把握ベースラインを +23% および +34% 上回っています。コード、データ、ベンチマーク、チェックポイント、およびインタラクティブなデモは、当社の Web サイト (https://grasping.io/) でリリースされています。

原文 (English)

Human Universal Grasping

Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from humans, who pick up thousands of objects every day. We present HUG, a flow-matching model that generates diverse human grasps for any user-specified object in a single RGB-D image captured from a stereo camera. Using smart glasses, we first collect 1M-HUGs, an egocentric dataset of human grasps spanning 1M frames (27.8 hrs) and 6,707 object instances across 41 buildings. Next, to model the distribution of natural human grasps, our novel flow-matching model fuses RGB and depth observations to output a grasp parameterized by wrist translation, wrist rotation, and MANO hand pose. Predicted grasps can be retargeted to various robot hands, enabling zero-shot grasping in everyday scenes. To standardize evaluation, we build a new simulated benchmark, HUG-Bench, of 90 unseen objects from five geometric categories and various sizes, with metric-scale 3D meshes. We evaluate HUG in the real world on the 30-object test set of HUG-Bench across multiple stereo cameras, robot embodiments, and household environments. HUG outperforms the state-of-the-art grasping baselines by +23% and +34% on our challenging object set. Code, data, benchmark, checkpoints, and an interactive demo are released on our website: https://grasping.io/

13:00 JSTエージェント

ビジネス コンテキストにおける人間と AI エージェントの対話

AI エージェントが中核的なビジネス プロセスにますます統合されるにつれ、人間と AI エージェントの間の効果的な対話パターンを理解し、設計することが価値創造にとって重要になります。この研究では、AI エージェントによるポジティブなユーザー エクスペリエンス (UX) の原則と基準、およびその測定方法を特定し、評価します。私たちはユーザーの期待とニーズを特定して、導入を促進し、信頼を構築し、開発チームによるユーザー中心の意思決定をサポートします。定性的手法と定量的手法を組み合わせた混合手法アプローチを使用して、人間と AI エージェント間の相互作用パターンを調査します。この探索的研究の結果は、特定の設計要素の有効性を大規模に評価する調査実験を開発するための基礎として役立ちます。この基礎研究は、ビジネス現場におけるより直感的で効果的な人間と AI エージェントの対話の開発に貢献します。

原文 (English)

Human-AI Agent Interaction in a Business Context

As AI agents are increasingly integrated into core business processes, understanding and designing effective interaction patterns between humans and AI agents becomes crucial for value creation. This study identifies and evaluates principles and criteria for a positive User Experience (UX) with AI agents, along with methods for its measurement. We identify user expectations and needs to facilitate adoption, build trust, and support user-centered decision-making by development teams. Using a mixed-methods approach that combines qualitative and quantitative techniques, we explore interaction patterns between humans and AI agents. The findings from this exploratory research serve as the basis to develop a survey experiment which evaluates the effectiveness of specific design elements on a larger scale. This foundational research contributes to the development of more intuitive and effective human-AI agent interactions in business settings.

13:00 JSTLLM/生成AIGPT / ChatGPT

言われていないことを明らかにする: 確率的パス集約による隠れた LLM バイアスの視覚化

大規模言語モデル (LLM) には、テキスト生成の確率的な性質により評価が難しい表現的および構文的なバイアスが見られます。標準的な監査方法は、単一の出力検査または静的な自動化されたメトリクスに依存します。これらのアプローチでは、基礎となる確率分布が不明瞭になり、確率の低い生成分岐に隠されたバイアスを捕捉できません。このペーパーでは、集計された比較を通じて LLM バイアスを評価するように設計されたビジュアル分析ツールである TreeTracer を紹介します。このツールは、体系的な摂動分析パイプラインを使用して、各入力プロンプト内のオントロジーで定義された用語を置換し、数百の確率的世代を構文整合された階層構造に集約してから、補助言語モデルとの分類を認識したノードのマージを実行します。結果として得られる構造は、カスタム サンキー ダイアグラムを通じて視覚化されます。 2 つのオントロジー駆動のツリーを並置することにより、ワークスペースはセマンティック コンテキスト間の直接比較を可能にし、体系的なバイアス検出をサポートします。どの視覚化もモデルの学習された動作のサブセットのみを反映するため、システムはさらに対照的推論を適用して、コンテキスト全体にわたる反事実のトークン確率を計算して直接表示し、バイアスの存在を誤解するリスクを軽減します。私たちは、アライメントされていないベースライン モデル GPT-2 XL と構成的にアライメントされた Apertus モデルを比較するケース スタディを通じてワークスペースを検証します。視覚的な集合体は、反事実的な代名詞の抑制や会話による個人の疎外など、隠れた表象上の害悪を明らかにすることに成功しました。予備的なユーザー調査では、集約された比較インターフェイスが認知負荷を軽減し、アナリストによる体系的なバイアスの検出を効果的にサポートすることが確認されています。

原文 (English)

Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation

Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generation. Standard auditing methods rely on a single output inspection or static automated metrics. These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability generation branches. This paper introduces TreeTracer, a visual analytics tool designed to evaluate LLM bias through aggregated comparison. Using a systematic perturbation analysis pipeline, the tool replaces ontology-defined terms in each input prompt, aggregates hundreds of stochastic generations into a syntax-aligned hierarchical structure, and then performs classification-aware node merging with an auxiliary language model. The resulting structure is visualized through a custom Sankey diagram. By juxtaposing two ontology-driven trees, the workspace enables direct comparison between semantic contexts and supports systematic bias detection. Because any visualization reflects only a subset of the model's learned behavior, the system further applies contrastive inference to compute and directly display counterfactual token probabilities across contexts, reducing the risk of misinterpreting the presence of bias. We validate the workspace through case studies comparing an unaligned baseline model GPT-2 XL against the constitutionally aligned Apertus models. The visual aggregation successfully exposes hidden representational harms, such as counterfactual pronoun suppression and conversational marginalization of individuals. A preliminary user study confirms that the aggregated comparative interface reduces cognitive load and effectively supports analysts in detecting systemic biases.

13:00 JSTLLM/生成AIGoogleGeminiGemma

要約に基づいて PubMed の EQ-5D 研究を識別するための大規模言語モデルのアンサンブル

科学出版物の急速な増加により、体系的文献レビュー (SLR) における手作業による研究スクリーニングはますますリソースを消費し、非効率で、一貫性がなくなっているという事実が生じています。 EQ-5D データなど、健康関連の生活の質の結果を明確に報告する研究を分類するには、高度な臨床解釈が必要であり、人間の審査員にとっては課題となります。この研究では、公開された抄録のみに基づいて、PubMed 生物医学データベース内の EQ-5D 検出を自動化する際の Google の Gemini および Gemma 大言語モデル (LLM) の使用を調査します。少数ショット プロンプティング、重みアンサンブル集約、およびソフト スタッキング メタ分類子を統合するマルチフェーズ フレームワークが提案されています。 9 つの LLM は、EQ-5D レポートに関して 2 人の専門家によって手動でラベル付けされた PubMed 研究のデータセットで評価されます。 gemini-2.5-pro、gemma-3-12b、および gemma-3-27b の加重アンサンブルでは、0.74 の加重 F1 スコアと 0.74 の精度が得られ、個別に得られた結果を上回りました。最高パフォーマンスのモデルをアンサンブルすることで、個別のモデルと比較して精度と再現率のバランスが向上し、ソフト スタッキング アプローチにより信頼性と解釈可能性が向上しました。特徴分析により、モデルから得られる確率の結果が最終的な予測を導く上で重要であることがわかります。この調査結果は、アンサンブルベースの LLM セットアップが、生物医学研究におけるスクリーニングを自動化するための信頼性が高く、スケーラブルなアプローチであることを示唆しています。

原文 (English)

Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent. Classifying studies that clearly report health-related quality-of-life results, such as EQ-5D data, requires a high level of clinical interpretation and poses challenges for human reviewers. This study investigates the use of Google's Gemini and Gemma large language models (LLMs) in automating EQ-5D detection in the PubMed biomedical database based only on published abstracts. A multi-phase framework is proposed that integrates few-shot prompting, weight ensembling aggregation, and a soft stacking meta-classifier. Nine LLMs are evaluated on a dataset of PubMed studies manually labeled by two experts regarding EQ-5D reporting. The weighted ensemble of gemini-2.5-pro, gemma-3-12b, and gemma-3-27b obtained a 0.74 weighted F1-score and 0.74 accuracy, exceeding individually attained results. The ensembling of top-performing models improved the balance between precision and recall compared to individual models, while the soft stacking approach provided greater reliability and interpretability. Feature analysis shows that the probability results from the models are important in guiding the final predictions. The findings suggest that an ensemble-based LLM setup is a reliable and scalable approach for automating screening in biomedical research.

13:00 JSTLLM/生成AI

言語を越えた転移におけるタスクの調整から言語の関連性を解きほぐす

私たちは、アラビア語について 7 つの大きな言語モデル (4B ~ 671B パラメーター) を微調整し、セム語言語と非ユダヤ人の対照についてゼロショット読解を評価することによって、言語間伝達を研究します。高密度の専門家混合アーキテクチャ全体では、セム語特有の転移の証拠は見つかりません。ベースラインが弱いモデルはすべての言語にわたって劇的に向上しますが、ベースラインが強いモデルは言語族に関係なくわずかな向上しか示しません。思考連鎖のアブレーションはこの発見を補強します。微調整から最も恩恵を受ける同じモデルは、推論時推論からも同様に恩恵を受けます。これは、両方のメカニズムが、言語を越えた知識伝達ではなく、タスク形式の調整に取り組んでいることを示唆しています。

原文 (English)

Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer

We study cross-lingual transfer by fine-tuning seven large language models (4B--671B parameters) on Arabic and evaluating zero-shot reading comprehension on Semitic languages and non-Semitic controls. Across dense and Mixture-of-Experts architectures, we find no evidence of Semitic-specific transfer: models with weak baselines improve dramatically across all languages, while strong-baseline models show only marginal gains regardless of language family. A chain-of-thought ablation reinforces this finding -- the same models that benefit most from fine-tuning benefit equally from inference-time reasoning, suggesting both mechanisms address task-format alignment rather than cross-lingual knowledge transfer.

13:00 JSTLLM/生成AI

ハードウェア設計の RTL コーディングにおいて LLM はどのように失敗し、一般化されるのでしょうか?

逐次プログラミングの事前処理をハードウェア設計の並列時相論理に変換することは、依然として大規模言語モデル (LLM) にとって重大なボトルネックとなっています。これを調査するために、認知理論に触発された、問題解決可能性に基づいた新しいエラー分類法を導入します。私たちの分類法では、失敗を構文的、意味論的、解決可能な関数型、および解決不可能な関数型に分類します。評価の結果、フロンティア モデルの初期合格率は 90.8% で頭打ちとなるため、VerilogEval ベンチマークには厳格な経験的上限があることが明らかになりました。これらのプラトーは解決できない機能エラーによって定義され、テスト時間の計算スケーリングの影響を受けない永続的な知識のギャップを露呈します。さらに、表面上の顕著な収束ギャップが明らかになります。最適化により構文エラーはすぐに排除されますが、同時により深い機能上の障害が悪化します。私たちの調査結果は、位置合わせ技術が単にモデルにコンパイルを教えるだけであることを示しています。サンプリング戦略を繰り返すことで解決可能なエラーは解決できますが、レジスタ転送レベル (RTL) のコーディング能力は事前トレーニングの知識によって厳密に制限されたままです。現在の LLM ベースのハードウェア生成パイプラインの課題に対処するには、調整介入ではなくモデル推論に関するさらなる研究が必要です。

原文 (English)

How LLMs Fail and Generalize in RTL Coding for Hardware Design?

Translating sequential programming priors into the parallel temporal logic of hardware design remains a crucial bottleneck for large language models(LLM). To investigate this, we introduce a new error taxonomy grounded in problem solvability, inspired by cognitive theory. Our taxonomy categorizes failures into syntactic, semantic, solvable functional, and unsolvable functional types. Evaluations reveal a strict empirical ceiling on the VerilogEval benchmark, as frontier models plateau at a 90.8% initial pass rate. These plateaus are defined by unsolvable functional errors, exposing persistent knowledge gaps immune to test time compute scaling. Furthermore, we expose a striking surface convergence gap: optimization readily eliminates syntax errors but concurrently exacerbates deeper functional failures. Our findings demonstrate that alignment techniques merely teach models to compile. While repeated sampling strategies can patch solvable errors, register-transfer level(RTL) coding capacity remains strictly bounded by pretraining knowledge. Addressing challenges in the current LLM based hardware generation pipeline requires more studies in model reasoning rather than alignment interventions.

13:00 JSTLLM/生成AIDeepSeek

DeepSeek-V4: 非常に効率的な 100 万トークンのコンテキスト インテリジェンスを目指して

当社は、2 つの強力な専門家混合 (MoE) 言語モデル、1.6T パラメーター (49B アクティブ化) を備えた DeepSeek-V4-Pro と 284B パラメーター (13B アクティブ化) を備えた DeepSeek-V4-Flash を含む、DeepSeek-V4 シリーズのプレビュー バージョンを提供します。どちらも 100 万トークンのコンテキスト長をサポートします。 DeepSeek-V4 シリーズには、アーキテクチャと最適化においていくつかの重要なアップグレードが組み込まれています。(1) Compressed Sparse Attendant (CSA) と Heavy Compressed Attendance (HCA) を組み合わせたハイブリッド アテンション アーキテクチャにより、ロング コンテキストの効率が向上します。 (2) 従来の残留接続を強化するマニホールド制約ハイパー接続 (mHC)。 (3) および Muon オプティマイザーにより、収束が速くなり、トレーニングの安定性が向上します。私たちは両方のモデルを 32T を超える多様で高品質のトークンで事前トレーニングし、その後、その機能を解放してさらに強化する包括的なポストトレーニング パイプラインを実行します。 DeepSeek-V4-Pro の最大推論労力モードである DeepSeek-V4-Pro-Max は、オープン モデルの最先端を再定義し、コア タスクで以前のモデルを上回ります。一方、DeepSeek-V4 シリーズは、長いコンテキストのシナリオで非常に効率的です。 100 万トークンのコンテキスト設定では、DeepSeek-V4-Pro は、DeepSeek-V3.2 と比較して、単一トークン推論 FLOP の 27% と KV キャッシュの 10% のみを必要とします。これにより、100 万トークンのコンテキストを定期的にサポートできるようになり、長期的なタスクやさらなるテスト時間のスケーリングがより実現可能になります。モデルのチェックポイントは、https://huggingface.co/collections/deepseek-ai/deepseek-v4 で入手できます。

原文 (English)

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.

13:00 JSTLLM/生成AI

クエリをどこに配置するか?デコードダイナミクスによる拡散 LLM のインコンテキスト学習における位置バイアスの解明と軽減

インコンテキスト学習 (ICL) は自己回帰 (AR) LLM で広く研究されていますが、拡散大規模言語モデル (dLLM) 内のメカニズムはほとんど解明されていません。一方向の因果マスキングによって制限される AR モデルとは異なり、dLLM は本質的に双方向の注意を利用し、クエリ配置に広範な空間的柔軟性を提供します。残念なことに、現在の慣行は従来、AR スタイルの末尾クエリ テンプレートを継承しており、構造的なパラダイム シフトを見落としていることがよくあります。この論文では、クエリ位置が実際には dLLM の一次変数であることを明らかにする包括的な分析を紹介します。経験的な分離を通じて、位置の差異がサンプルの意味論的品質と同等の生成品質に影響を与えることを実証します。内部的には、この位置の敏感さは、注意の流れにおける空間的な「最新性効果」と、デコード軌道におけるタスク依存のシフトに起因します。グラウンドトゥルースラベルなしでこの不安定性を軽減するために、従来の単一ステップの信頼性 ($C_{decoded}$) が dLLM で失敗することを明らかにします。代わりに、反復復号プロセスを追跡する新しい指標である Average Confidence ($\overline{C}$) を提案します。基本的な空間 ICL ベースラインを確立することで、クエリの配置を動的に最適化し、異種の推論および認識タスク全体で Oracle のパフォーマンスに確実に近づく、トレーニング不要の適応ルーティング戦略である Auto-ICL を導入します。

原文 (English)

Where to Place the Query? Unveiling and Mitigating Positional Bias in In-Context Learning for Diffusion LLMs via Decoding Dynamics

While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (dLLMs) remains largely unexplored. Unlike AR models restricted by unidirectional causal masking, dLLMs intrinsically utilize bidirectional attention, offering extensive spatial flexibility for query placement. Unfortunately, current practices conventionally inherit AR-style trailing-query templates, often overlooking the structural paradigm shift. This paper presents a comprehensive analysis unveiling that query position is actually a first-order variable in dLLMs. Through empirical decoupling, we demonstrate that positional variance impacts generation quality on par with example semantic quality. Internally, this positional sensitivity stems from a spatial ``Recency Effect'' in attention flow and task-dependent shifts in decoding trajectories. To mitigate this instability without ground-truth labels, we reveal that traditional single-step confidence ($C_{decoded}$) fails in dLLMs. Instead, we propose Average Confidence ($\overline{C}$), a novel metric tracking the iterative decoding process. By establishing the foundational spatial ICL baselines, we introduce Auto-ICL, a training-free adaptive routing strategy that dynamically optimizes query placement, robustly approaching oracle performance across heterogeneous reasoning and perception tasks.

13:00 JSTLLM/生成AI

大規模言語モデルベースのナレッジグラフ推論のための幻覚の検出

ナレッジ グラフ (KG) 推論は、既存の事実から新しい知識を推測し、質問応答、推奨、意思決定支援に広く適用されます。大規模言語モデル (LLM) の急速な発展に伴い、取得した KG 情報を活用する LLM ベースの KG 推論フレームワークの人気が高まっています。しかし、LLM における幻覚は依然として重大な問題です。関連する KG の知識が組み込まれている場合でも、モデルは依然として誤った出力を生成し、誤った情報や信頼性の低い決定につながる可能性があります。既存の幻覚検出方法は、LLM の内部状態に焦点を当てるか、取得されたコンテキストとの一貫性を検証するかのいずれかですが、どちらも KG 内の構造情報を見落とすため、最適なパフォーマンスが得られません。このギャップに対処するために、我々は、LLM ベースの知識グラフ推論フレームワーク用の最初の幻覚検出方法である LUCID を提案します。 LUCID は、LLM アテンション スコア、KG セマンティクス、および構造情報を共同利用します。具体的には、アテンションスコアと意味的類似性からノードとエッジの特徴を抽出し、グラフニューラルネットワークを使用してKG構造と統合します。また、評価のために手動でアノテーションを付けたベンチマーク データセットも構築します。 9 つのデータセットに対する実験では、LUCID が 15 のベースラインと比較して最先端のパフォーマンスを達成していることが示されています。

原文 (English)

Detecting Hallucinations for Large Language Model-based Knowledge Graph Reasoning

Knowledge graph (KG) reasoning infers new knowledge from existing facts and is widely applied in question answering, recommendation, and decision support. With the rapid development of large language models (LLMs), LLM-based KG reasoning frameworks have become increasingly popular by leveraging retrieved KG information. However, hallucinations in LLMs remain a critical issue. Even when relevant KG knowledge is incorporated, models may still generate incorrect outputs, leading to misinformation and unreliable decisions. Existing hallucination detection methods either focus on LLM internal states or verify consistency with retrieved contexts, but both overlook the structural information in KGs, resulting in suboptimal performance. To address this gap, we propose LUCID, the first halLUcination deteCtIon method for LLM-based knowleDge graph reasoning frameworks. LUCID jointly leverages LLM attention scores, KG semantics, and structural information. Specifically, it extracts node and edge features from attention scores and semantic similarities, and integrates them with KG structure using a graph neural network. We also construct manually annotated benchmark datasets for evaluation. Experiments on nine datasets show that LUCID achieves state of the art performance compared to 15 baselines.

13:00 JSTLLM/生成AI研究/論文

大規模な手話データセット: リソース、ベンチマーク、および注釈標準に関する包括的な調査

手話は、聴覚障害者 (DHH) コミュニティによって使用される表現力豊かな視覚言語です。手話の認識、翻訳、作成は大幅に進歩しているにもかかわらず、断片化したデータセット、一貫性のない注釈、限られた言語範囲によって進歩は依然として制約されています。既存のベンチマークは現実世界の通信ニーズを反映していないことが多く、これらの制限の体系的な分析は依然として限られています。この調査では、35 の手話にわたる 120 のリソースをカバーする手話データセットの包括的なインデックスを提示します。モダリティの不均衡、アノテーションの粒度、署名者の偏りなどの主要な課題を分析し、将来のデータセット設計における考慮事項の概要を示します。また、標準化されたドキュメントと再現可能な評価をサポートするために、24 フィールドの Sign-Language データシートを導入し、パブリック GitHub リポジトリ (https://github.com/Ginqwerty/Open-Sign-Language) をリリースします。全体として、私たちの研究は、現実世界のアプリケーションで包括的で堅牢かつスケーラブルな手話技術を開発するための統一された実用的な基盤を提供します。

原文 (English)

Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards

Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities. Despite substantial progress in sign-language recognition, translation, and production, advances remain constrained by fragmented datasets, inconsistent annotations, and limited linguistic coverage. Existing benchmarks often fail to reflect real-world communication needs, and systematic analyses of these limitations remain limited. In this survey, we present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages. We analyze key challenges such as modality imbalance, annotation granularity, and signer bias, and outline considerations for future dataset design. We also introduce a 24-field Sign-Language Datasheet and release a public GitHub repository (https://github.com/Ginqwerty/Open-Sign-Language) to support standardized documentation and reproducible evaluation. Overall, our work provides a unified and practical foundation for developing inclusive, robust, and scalable sign-language technologies in real-world applications.

13:00 JSTLLM/生成AIエージェントQwen

信頼できるマルチエージェント システム: Argent シグナリング プロトコルによるセマンティック ドリフトの軽減

マルチエージェント LLM システムが悪い回答を生成する場合、すべての失敗が同じであるわけではありません。一部の回答は適切な内容に基づいているが不完全であり、他の回答は単純に根拠がなく、停止する必要があります。現在の再試行戦略では両方のケースが同じように扱われ (再試行して最善の結果を期待します)、人間のスーパーバイザーは再試行が正当であるかどうか、あるいは代わりにシステムを停止すべきかどうかを判断できません。 Argent Signaling Protocol (ASP) は、AI が生成するすべての応答に、確実性 (@C)、根拠 (@G)、確率性 (@S)、各主張の証拠根拠を分類する仮定インデックスなどの構造化された品質信号を伴うコンパクトな機械可読ヘッダーです。これらの信号により、コントローラーは修復可能な故障と格納容器の故障を区別し、それぞれのケースを異なる方法でルーティングすることができます。 ASP を 2 つのモードで評価します。スタンドアロン モードでは、Array BioPharma/Ono ライセンス契約に基づく 27 の質問の文書に基づいた QA ベンチマークが、3 つのローカル GGUF モデルにわたるベースライン プロンプトと ASP で計測されたコントローラー アクションを比較します。 Qwen~(0.8B) では、ASP により合格率が 11.1% から 33.3% に、平均期間カバレッジが 36.7% から 65.4% に向上しました。 Dobby~(8B) では、ASP は 4 つの失敗したリカバリを生成し、成功率を 33.3% から 44.4% に上昇させます。 SmolLM3~(3B) では、ASP は質問ごとに修復と封じ込めを切り替えます。総計の改善には意味があります (12/81 から 21/81 までのパス)。マルチエージェント モードでは、ASP サイドカーは検索エージェントと下流の決定エージェントの間に位置します。サイドカーは、接地されていないアップストリーム出力がダウンストリーム エージェントに到達するのを 100% ブロックします (24/27 ブロック、接地されていない伝播は 0)。

原文 (English)

Trustworthy Multi-Agent Systems: Mitigating Semantic Drift with the Argent Signaling Protocol

When multi-agent LLM systems produce bad answers, not all failures are equal: some answers are grounded in the right material but incomplete, while others are simply ungrounded and should be stopped. Current retry strategies treat both cases identically (try again and hope for the best), leaving human supervisors unable to tell whether a retry was warranted or whether the system should have halted instead. We introduce the Argent Signaling Protocol (ASP), a compact machine-readable header that accompanies every AI-generated response with structured quality signals: certainty (@C), grounding (@G), stochasticity (@S), and an assumption index that classifies the evidentiary basis of each claim. These signals enable a controller to distinguish repairable failures from containment failures and route each case differently. We evaluate ASP in two modes. In standalone mode, a 27-question document-grounded QA benchmark over the Array BioPharma/Ono license agreement compares baseline prompts against ASP-instrumented controller actions across three local GGUF models. On Qwen~(0.8B), ASP improves pass rate from 11.1% to 33.3% and mean term coverage from 36.7% to 65.4%; on Dobby~(8B), ASP produces 4 fail-to-pass recoveries, raising pass rate from 33.3% to 44.4%; on SmolLM3~(3B), ASP alternates between repair and containment per question. Aggregate improvement is meaningful (12/81 to 21/81 passes). In multi-agent mode, an ASP sidecar sits between a retrieval agent and a downstream decision agent; the sidecar blocks 100% of ungrounded upstream outputs from reaching the downstream agent (24/27 blocked, 0 ungrounded propagations).

13:00 JSTロボティクス

Physical Atari: ロボット上のリアルタイム強化学習のための堅牢でアクセスしやすいプラットフォーム

私たちは、Atari CX40+ コントローラーを作動させる Robotroller と呼ばれるロボットと、ゲーム フレームとアーケード学習環境からの報酬信号を画面上にレンダリングする Atari Devbox と呼ばれるデバイスを構築しました。 Robotroller と Atari Devbox は、既製のカメラとデスクトップ コンピューターとともに、物理世界で強化学習アルゴリズムを研究するために使用できるシステムを構成します。システム全体を物理アタリと呼びます。このペーパーでは、Physical Atari を堅牢でアクセスしやすいプラットフォームにするための重要な決定について詳しく説明します。システムを堅牢にするために、すべての動きがベアリングを介して行われるようにロボットローラーを設計し、摩耗を軽減しました。さらに、サーボの状態を高周波で監視し、応力を制限するために介入するソフトウェアも作成しました。システムを利用しやすくするために、家庭用 3D プリンタを使用して製造できる、手頃な価格の既製コンポーネントと部品を使用しました。 Physical Atari は 1,000 ドル未満で構築でき、数週間にわたるノンストップの強化学習実験で機械的な故障が発生することなく使用されています。私たちはこれを使用して、強化学習アルゴリズムがロボット上で直接学習できることを検証し、学習と展開の間の小さな分布の変化でさえポリシーのパフォーマンスを大幅に低下させる可能性があることを示しました。私たちの結果は、ロボットで優れたパフォーマンスを得るにはデバイス上の適応が重要であることを強調しています。

原文 (English)

Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots

We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari Devbox that renders the game frame and the reward signal from the Arcade Learning Environment on a screen. The Robotroller and the Atari Devbox, together with an off-the-shelf camera and a desktop computer, constitute a system that can be used to study reinforcement learning algorithms in the physical world. We call the full system Physical Atari. In this paper, we detail the key decisions that make Physical Atari a robust and accessible platform. To make the system robust, we designed the Robotroller so that all movement is done through bearings, which reduces wear. Additionally, we wrote software that monitors the state of the servos at a high frequency and intervenes to limit stress. To make the system accessible, we used affordable off-the-shelf components and parts that can be manufactured using consumer 3D printers. Physical Atari can be built for under $1,000 and has been used for weeks of non-stop reinforcement learning experiments without any mechanical failures. We used it to validate that reinforcement learning algorithms can learn directly on robots and show that even small distribution shifts between learning and deployment can significantly degrade the performance of policies. Our results underscore the importance of on-device adaptation for strong performance on robots.

13:00 JST研究/論文

計算による識別可能性

識別条件は、利用可能な情報の種類と量の関数として、ターゲット クエリまたは対象パラメータの計算可能性を記述します。因果関係の特定では、この情報は因果関係グラフの形式で表現されることが多く、グラフ内の変数の一部についてデータが観察または収集されます。ターゲット クエリは、単一の効果のみを対象とする場合もあれば、特定のモデル内の効果のクラスを対象とする場合もあります。次に、識別アルゴリズムの導出により、期待される望ましい因果効果を理論的に一意に決定できるプロセスが数学的に定義されます。期待における識別可能性、または「理論的識別可能性」は、一般に、漸近特性、無限データ、またはその他の数学的に理想化された条件を前提としています。この論文では、この理論的で理想的な識別可能性の概念と、計算に依存する提案された代替概念との間の根本的な違いを探ります。私たちが提案するフレームワークである「計算による識別可能性」は、代わ​​りに経験的推定量に対する有限の計算による探索手順を定義するものです。このプロセスが所望の誤差許容範囲内で経験的に推定量を見つけた場合、特定の検索の仮定 (つまり、パラメーターにわたる事前分布) と検索手順自体の条件を条件として、識別可能性が満たされます。いくつかの実験を通じて、このフレームワークがどのようにして、小さな有限サンプル、曖昧なグラフィック基準、混合観察データと介入データ、および反事実データと推定値による識別など、きめの細かい実用的な識別の質問に答えることができるかを実証します。コードは https://github.com/lbynum/metadentify で入手できます。

原文 (English)

Computational Identifiability

Identification conditions describe the computability of a target query or parameter of interest as a function of the type and amount of information available. In causal identification, this information is often expressed in the form of a causal graph, and data are observed or collected for some subset of variables in the graph. Target queries may be for a single effect alone or for a class of effects in a given model. The derivation of an identification algorithm then defines mathematically the process by which the desired causal effect(s) can be uniquely determined, theoretically, in expectation. Identifiability in expectation, or 'theoretical identifiability,' generally assumes asymptotic properties, infinite data, or other mathematically idealized conditions. In this paper, we explore a fundamental distinction between this theoretical, idealized notion of identifiability and a proposed alternative that is computation-bound. The framework we propose - 'computational identifiability' - is to instead define a finite computational search procedure for an empirical estimator. If this process finds an estimator empirically, within a desired error tolerance, then identifiability is satisfied, conditional on the specified assumptions of the search (i.e., a prior distribution over the parameters) and conditional on the search procedure itself. Through several experiments, we demonstrate how this framework allows us to answer fine-grained, practical identification questions, such as identification with small finite samples, with ambiguous graphical criteria, with mixed observational-interventional data, and across counterfactual data and estimands. Code is available at https://github.com/lbynum/metadentify.

13:00 JST研究/論文

確率的グラフィカルモデル構造学習としての情報格子学習

情報格子学習 (ILL) は、抽象化の階層をエンコードするパーティション ラティスに信号を交互に投影し、選択したルールを信号ドメインに持ち上げることで、信号の解釈可能なルールを学習します。信号が確率質量関数である場合、ILL によって学習された確率規則が自然な確率的グラフィカル モデル (PGM) 解釈を認めることを示し、この解釈を詳細に展開します。 ILL の分割は決定論的な商変数を誘導し、ルールはその商変数の限界法則です。したがって、ルールセットは、解釈可能な抽象概念に対する限界制約の集合です。一般リフティングは、これらの制約を満たすすべての同時分布の実行可能なグループですが、特殊リフティングは、最大エントロピーに密接に関連する L2 均一性原理によって ILL に実装される最大無視再構成を選択します。シャノンエントロピーリフティングの下で​​は、同じ制約により、学習された抽象化によって因子がインデックス付けされる対数線形因子グラフが生成されます。しかし、情報格子自体はベイジアン ネットワークではありません。そのエッジは、条件依存ではなく、抽象化の洗練と粗大化をコード化しています。したがって、ILL は、商変数に対する解釈可能な制約ベースの因子グラフの構造学習として捉えるのが最適です。このビューは、ILL がグラフィカル モデルおよび最大エントロピー モデルにどのように関連しているかを明らかにするとともに、推論、識別可能性、および記号と確率のハイブリッド学習の新しい方向性を提案します。

原文 (English)

Information Lattice Learning as Probabilistic Graphical Model Structure Learning

Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice that encodes a hierarchy of abstractions and lifting selected rules back to the signal domain. When the signal is a probability mass function, we show the probabilistic rules learned by ILL admit a natural probabilistic graphical model (PGM) interpretation and develop this interpretation in detail. A partition in ILL induces a deterministic quotient variable, and a rule is the marginal law of that quotient variable. A rule set is therefore a collection of marginal constraints over interpretable abstractions. General lifting is the feasible family of all joint distributions satisfying those constraints, while special lifting chooses a maximum-ignorance reconstruction, implemented in ILL by an L2 uniformity principle closely related to maximum entropy. Under a Shannon-entropy lifting, the same constraints yield a log-linear factor graph whose factors are indexed by learned abstractions. The information lattice itself, however, is not a Bayesian network: its edges encode refinement and coarsening of abstractions, not conditional dependence. Thus ILL is best viewed as structure learning for interpretable constraint-based factor graphs over quotient variables. This view clarifies how ILL relates to graphical models and maximum entropy models, while suggesting new directions for inference, identifiability, and hybrid symbolic-probabilistic learning.

13:00 JST研究/論文

ゼロインフレート ガウス分布により、分布推定アルゴリズムにおけるパラメータ空間のスパース性が可能になります

分布推定アルゴリズム (EDA) は、特に目標の構造がほとんどわかっていない場合に、ブラック ボックス最適化のための進化的手法の強力なクラスです。古典的な進化アルゴリズムは手作業で設計された突然変異や交叉演算子に依存しており、未知の問題構造やバイアスの原因を考慮して考案するのは困難ですが、EDA は演算子設計を完全に回避します。つまり、確率分布を最良の個体に適合させ、そこから次世代をサンプリングします。 EDA は連続パラメータ空間では十分に確立されていますが、良好な解のほとんどの係数が正確に 0 である疎パラメータ空間にはこれまで一般化されていませんでした。したがって、既存のスパース ブラック ボックス オプティマイザーは、EDA が回避するように設計されたものを正確に再導入しています。つまり、手作りのスパース演算子、サポート セットとアクティブな値の間で交互に行われる 2 レベル スキーム、しきい値のゼロ化、およびその他の組み込みの仮定です。我々は、EDA サンプリング法則として多変量ゼロ膨張ガウス (ZIG) 分布を提案することで、このギャップを埋めます。個別のインジケーター次元と値次元を持つ潜在ガウス モデルは、スパース パターン、アクティブなパラメーター間の相関、および 2 つの間の相互作用を表すため、スパース パターンとアクティブな値は階層なしで一緒に最適化されます。我々は、このモデルの潜在パラメータが、関連する構造が生じる欠損データ設定とは異なり、観察されたサンプルから識別可能であることを示し、それらに対する実用的な償却逆ベースの推定量を導入します。推定器は潜在的な相関構造を正確に回復し、Lunar Lander ベンチマークでは、結果として得られる ZIG-EDA は、高密度ガウス EDA、手作りのスパース進化アルゴリズム、およびアドホック スパース EDA よりも高速に収束し、より高い最終収益に達すると同時に、ごく一部のパラメーターのみがアクティブなコントローラーを検出します。

原文 (English)

Zero-Inflated Gaussian Distributions Enable Parameter-Space Sparsity in Estimation-of-Distribution Algorithms

Estimation-of-distribution algorithms (EDAs) are a powerful class of evolutionary methods for black-box optimization, especially when little is known about the structure of the objective. Whereas classical evolutionary algorithms rely on hand-designed mutation and crossover operators, hard to devise for unknown problem structures, and a source of bias, EDAs sidestep operator design entirely: they fit a probability distribution to the best individuals and sample the next generation from it. EDAs are well established on continuous parameter spaces, but they have not previously been generalized to sparse ones, in which most coefficients of a good solution are exactly zero. Existing sparse black-box optimizers therefore reintroduce exactly what EDAs were designed to avoid: hand-crafted sparsity operators, bi-level schemes alternating between support set and active values, zeroing thresholds, and other baked-in assumptions. We close this gap by proposing multivariate zero-inflated Gaussian (ZIG) distributions as EDA sampling laws. A latent Gaussian model with separate indicator and value dimensions represents sparsity patterns, correlations among active parameters, and the interactions between the two, so sparsity patterns and active values are optimized jointly, hierarchy-free. We show that the latent parameters of this model are identifiable from observed samples, unlike in the missing-data settings where related constructions originate, and introduce practical amortized inversion-based estimators for them. The estimators accurately recover latent correlation structures, and on the Lunar Lander benchmark the resulting ZIG-EDA converges faster and reaches higher final returns than a dense Gaussian EDA, a hand-crafted sparse evolutionary algorithm, and an ad-hoc sparse EDA, while finding controllers with only a small fraction of parameters active.

13:00 JST研究/論文

人間のような自律性は、セルフプレイとひとつまみの人間のデータから生まれます

セルフプレイ強化学習は、人間のデータを使用せずに運転ポリシーをトレーニングする方法として最近登場しました。高価で大規模な人間による運転デモの代わりに、安価で大規模なシミュレーションを使用します。このアプローチの主な制限は、純粋なセルフプレイを通じて訓練されたポリシーが、効果的ではあるが人間とは相容れない異質な運転慣習を学習できることです。これまでの研究では、大規模な報酬エンジニアリングとドメインのランダム化を通じて、このような行動の不整合を軽減しようとしましたが、これらは脆弱で労働集約的でした。私たちの方法では、人間のデモンストレーションを完全に破棄するのではなく、最小限の安全な目標達成報酬に加えて、正則化の目標として扱います。おいしいシチューのスパイスと同じように、少量の人的データが大いに役立つことがわかりました。私たちの方法では、人間によるデモンストレーションのみ 30 分しか使用せず、同等の模倣学習アプローチより 2,500 分の 1 です。結果として得られるポリシーは、保持されている人間の軌跡と調整され、単一の消費者グレードの GPU で 15 時間でトレーニングを完了します。ビデオと完全なソース コードは https://spiced-self-play.com/ で入手できます。

原文 (English)

Human-like autonomy emerges from self-play and a pinch of human data

Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data. It uses cheap, large-scale simulations to substitute expensive, large-scale human driving demonstrations. A key limitation of this approach is that policies trained through pure self-play can learn effective but alien driving conventions incompatible with people. Previous works attempt to mitigate such behavioral misalignments through extensive reward engineering and domain randomization, which are brittle and labor-intensive. Instead of completely discarding human demonstrations, our method treats them as a regularization objective on top of a minimal safe goal-reaching reward. Like the spice in a good stew, we find that a little human data goes a long way: our method uses only 30 minutes of human demonstrations, 2500x fewer than comparable imitation learning approaches. Resulting policies coordinate with held-out human trajectories and complete training in 15 hours on a single consumer-grade GPU. Videos and full source code are available at https://spiced-self-play.com/.

13:00 JST画像/動画生成

ProMUSE: 進行性マルチモーダル不確実性ガイドに基づく段階的証拠アルツハイマー病分類

アルツハイマー病 (AD) は、高齢者の記憶力と認知能力を破壊する致命的な疾患です。 AD の治療法のほとんどは初期段階で効果があるため、AD の早期診断に対する需要が高まっています。アルツハイマー病の診断は、臨床評価、構造磁気共鳴画像法 (MRI)、陽電子放出断層撮影法 (PET) イメージングなどの複合データにますます依存しています。しかし、MRI と PET の取得は依然としてコストが高く、誰もがアクセスできるわけではないため、フルモダリティ推論は現実世界の臨床ワークフローでは非現実的です。私たちは、追加のモダリティがいつ必要になるかを適応的に判断し、精度を維持しながらデータ収集の全体的なコストを削減する、プログレッシブ・マルチモーダル不確実性誘導段階的証拠ネットワークである ProMUSE を提案します。 ProMUSE は、まず低コストの臨床データを使用して証拠分類を実行し、ディリクレベースの主観的論理モデルによって不確実性を定量化します。不確実性が学習された閾値を超えると、ProMUSE は MRI または PET の機能を段階的に組み込み、デンプスター・シェーファー理論を通じてモダリティごとの信念と不確実性を融合して、校正されたマルチモーダル予測を取得します。この段階的な取得戦略により、高価な画像処理への依存を最小限に抑えながら、正確な診断が可能になります。 CN-AD、CN-MCI、および MCI-AD タスクにわたる ADNI、AIBL、および OASIS の実験では、ProMUSE がフルモダリティのベースラインと比較して競合または優れた精度を達成しながら、MRI/PET の使用量を 50 ~ 90% 削減し、大幅なコスト削減を実現できることが実証されています。これらの結果は、ProMUSE が現実世界の AD スクリーニングのための実用的で不確実性を認識し、リソース効率の高いソリューションであることを強調しています。

原文 (English)

ProMUSE: Progressive Multi-modal Uncertainty-guided Staged Evidential Alzheimer Disease Classification

Alzheimer's disease (AD) is a fatal disorder that destroys memory and cognitive skills in the elderly population. Most treatments for AD are effective in the early stage, leading to an increasing demand for early AD diagnosis. AD diagnosis increasingly relies on multimodal data such as clinical assessments, structural Magnetic Resonance Imaging (MRI), and Positron Emission Tomography (PET) imaging. However, MRI and PET acquisition remain costly and not universally accessible, making full-modality inference impractical in real-world clinical workflows. We propose ProMUSE, a Progressive Multi-modal Uncertainty Guided Staged Evidential Network that adaptively determines when additional modalities are necessary, helping reduce the overall cost of data acquisition while maintaining accuracy. ProMUSE first performs evidential classification using low-cost clinical data and quantifies uncertainty via a Dirichlet-based subjective logic model. When uncertainty exceeds a learned threshold, ProMUSE progressively incorporates MRI or PET features, fusing modality-wise belief and uncertainty through Dempster-Shafer theory to obtain a calibrated multimodal prediction. This staged acquisition strategy enables accurate diagnosis while minimizing reliance on expensive imaging. Experiments on ADNI, AIBL, and OASIS across CN-AD, CN-MCI, and MCI-AD tasks demonstrate that ProMUSE achieves competitive or superior accuracy compared to full-modality baselines while reducing MRI/PET usage by 50-90%, yielding substantial cost savings. These results highlight ProMUSE as a practical, uncertainty-aware, and resource-efficient solution for real-world AD screening.

13:00 JST研究/論文

cAPM: アクティブ ラーニングによる継続的な AI 支援ペースマッピング

心室頻拍は生命を脅かすリズム障害であり、心臓突然死の主な原因です。ペースマッピングは、VT のカテーテルアブレーション中に介入ターゲットを特定するための臨床手順です。臨床医は、心室のさまざまな部位のペーシングを行い、結果として得られる心電図を迅速に解釈して、次にどこでペーシングを行うか、または標的部位が特定されているかどうかを判断する必要があります。アクティブ ラーニング AI モデルは、臨床医を次のペーシング サイトに誘導するために提案されており、ペーシング サイトの数を減らし、ペース マッピングの効率を向上させることが期待されています。既存の方法では、同じ患者内または複数の患者間で複数の VT に知識を伝達する機能がなく、各ターゲットを再トレーニングする必要があります。継続的な AI 支援ペースマッピングに cAPM を導入し、過去のペースマッピング データから蓄積された知識を取得して転送し、将来のターゲット VT に必要なペースマッピング データの数を削減します。これは、ペーシング部位から 12 誘導 ECG 形態へのマッピングを学習するタスク非依存のサロゲート ニューラル ネットワーク、各ターゲットに対して最も有益なペーシング部位を選択することでこのサロゲート モデルを改良するアクティブ ラーニング戦略、および以前のターゲットからの知識を保持しながら順次学習を行う継続学習戦略によって可能になります。異なる生理学的状態および心室形状にわたって順次提示される位置特定タスクからなるインシリコテストベッドで評価したところ、過去のデータサンプルの再生ありまたはなしのcAPMは、4.5のペースマッピング部位を使用して臨床許容範囲内(精度5mm)内で位置特定の81%の確率を達成したのに対し、最先端のアクティブラーニング手法は13.7のペーシング部位を使用して38%の確率を達成しました。これらの結果は、ペースマッピングのガイドに使用できる in vivo の前臨床および臨床研究に向けて cAPM を準備するための強力な基礎を提供します。

原文 (English)

cAPM: Continual AI-Assisted Pace-Mapping with Active Learning

Ventricular tachycardia is a life-threatening rhythm disorder and a major cause of sudden cardiac death. Pace-mapping is a clinical procedure for identifying the intervention target during catheter ablation of VT. It requires clinicians to pace different sites in the ventricles and rapidly interpret the resulting electrocardiograms to determine where to pace next or whether a target site has been identified. Active learning AI models have been proposed to guide clinicians to the next pacing site, showing promise in reducing the number of pacing sites and improving the efficiency of pace-mapping. Existing methods require retraining each target without the ability to transfer knowledge across multiple VTs within the same patient or across patients. We introduce cAPM for continuous AI-assisted pace-mapping to capture and transfer knowledge accumulated from past pace-mapping data to reduce the number of pace-mapping data needed for future target VTs. This is made possible by a task-agnostic surrogate neural network that learns the mapping from pacing sites to 12-lead ECG morphology, an active-learning strategy that refines this surrogate model by selecting the most informative pacing site for each target, and a continual learning strategy to do so sequentially while retaining knowledge from prior targets. Evaluated on an in-silico testbed consisting of sequentially-presented localization tasks across different physiological conditions and ventricular geometries, cAPM with and without replay of past data samples achieved an 81% probability of localizing within clinical tolerance (5 mm accuracy) using 4.5 pace-mapping sites, compared to the state-of-the-art active-learning method achieving 38% probability using 13.7 pacing sites. These results provide a strong basis for preparing cAPM towards in-vivo preclinical and clinical studies where it can be used to guide pace-mapping.

13:00 JST研究/論文

二次構造およびエネルギー フィルター処理された水素結合グラフを使用したタンパク質表現の学習

グラフベースの表現はタンパク質モデリングで広く使用されていますが、既存のアプローチの多くは主に配列の隣接性または幾何学的近接性に依存しており、これらはタンパク質のフォールディングを支配する原理を部分的にしか反映していません。代わりに、タンパク質は、$\alpha$-helices や $\beta$-sheets などの二次構造要素を中心に組織化された複雑な三次元立体構造をとり、反復する局所モチーフと安定化する水素結合相互作用をコードします。この研究では、タンパク質表現学習のための二次構造を意識したグラフ ニューラル ネットワークを導入します。残基レベルのノード表現は二次構造の割り当てによって強化され、エネルギー強度によってフィルターされた水素結合相互作用からグラフのエッジが構築されます。この設計により、モデルはタンパク質の安定性と機能の中心となる局所的な構造コンテキストと長距離カップリングの両方を捉えることができます。私たちは、一般的に使用されているタンパク質ベンチマークで提案されたアプローチを評価し、既存のグラフベースの手法と比較して一貫した改善を観察しました。さらに、学習された接続性が確立された構造モチーフと一致するため、結果として得られるグラフ表現は生物学的解釈可能性を高めます。これらの発見は、二次構造とエネルギーフィルターされた水素結合トポロジーを組み込むことで、タンパク質表現の学習に効果的な誘導バイアスが提供されることを示唆しています。コードは https://github.com/mohamedmohamed2021/SSProNet でリリースされています。

原文 (English)

Protein Representation Learning with Secondary-Structure and Energy-Filtered Hydrogen-Bond Graphs

Graph-based representations are widely used in protein modeling, yet many existing approaches rely primarily on sequence adjacency or geometric proximity, which only partially reflect the principles governing protein folding. Proteins instead adopt complex three-dimensional conformations organized around secondary structure elements, such as $\alpha$-helices and $\beta$-sheets, which encode recurring local motifs and stabilizing hydrogen-bond interactions. In this work, we introduce a secondary-structure-aware graph neural network for protein representation learning. Residue-level node representations are augmented with secondary structure assignments, and graph edges are constructed from hydrogen-bond interactions filtered by their energetic strength. This design enables the model to capture both local structural context and long-range couplings that are central to protein stability and function. We evaluate the proposed approach on commonly used protein benchmarks and observe consistent improvements over existing graph-based methods. In addition, the resulting graph representations offer enhanced biological interpretability, as the learned connectivity aligns with established structural motifs. These findings suggest that incorporating secondary structure and energy-filtered hydrogen-bond topology provides an effective inductive bias for protein representation learning. The code is released at https://github.com/mohamedmohamed2021/SSProNet

13:00 JSTLLM/生成AIClaude

ユーザー満足度保証のもと、ユーザーからのフィードバックを限定したコスト最適な LLM ルーティング

大規模言語モデル (LLM) アプリケーションの推論コストは、需要の急増とインフラストラクチャ コストの上昇により急速に増加しています。ユーザーは高品質の応答を期待しており、商用環境ではこれがサービス レベル アグリーメント (SLA) で正式に成文化されており、コストと品質の間に根本的な緊張関係が生じています。コストを意識した LLM リクエスト ルーティングの最近の進歩により、この緊張を解決できる可能性が示されていますが、既存のアプローチは完全なフィードバック信号、オフライン トレーニング、ワークロードごとの広範な調整に依存しており、そのほとんどは SLA 保証や推論時間の適応性を欠いています。実稼働システムで得られるまばらで一方的なユーザー フィードバックからコスト最適化ポリシーを学習するオンライン ルーティング アルゴリズムである SLARouter を紹介します。 SLARouter は、コストの最適化と厳密な SLA 準拠の両方を理論的に保証します。幅広い LLM ベンチマークの実験では、SLARouter がベンチマークごとの調整を必要とせずに SLA 制約を満たし、既存のベースラインと比較して運用コストを最大 2.2 倍削減できることが示されています。

原文 (English)

Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees

Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. Users expect high-quality responses, and in commercial settings this is formally codified in Service Level Agreements (SLAs), creating a fundamental tension between cost and quality. Recent progress on cost-aware LLM request routing has shown potential to resolve this tension, but existing approaches rely on complete feedback signals, offline training, extensive per-workload tuning, and most lack SLA guarantees or inference-time adaptivity. We introduce SLARouter, an online routing algorithm that learns a cost-optimal policy from the sparse, one-sided user feedback available in production systems. SLARouter provides theoretical guarantees for both cost optimality and strict SLA compliance. Experiments across a wide range of LLM benchmarks show that SLARouter satisfies SLA constraints without the need for per-benchmark tuning, reducing operating cost by up to 2.2x over existing baselines.

13:00 JST研究/論文

Emyx: 高速かつ効率的な全原子タンパク質生成

コンピューターによる酵素の設計では、触媒残基とリガンドの足場となるタンパク質を生成する必要があり、その作業には基礎となる生成モデルの幾何学的精度と構造的多様性の両方が要求されます。現在の全原子ジェネレーターは構造予測から高価なアーキテクチャを継承しているため、トレーニング コストが高くなり、サンプルの多様性が制限されます。私たちは、この複雑さの多くは、豊富な共進化信号ではなくまばらな幾何学的制約を条件とするジェネレーターには不要であると主張します。 Emyx は、標準の変圧器ブロック内に容量を集中させ、重い埋め込みスタックを軽量の条件付き表現と疎な接続に置き換える 140M パラメータの条件付きフロー マッチング モデルです。さらに、フロー マッチング補間の正確な再パラメータ化を EDM ノイズ レベル フレームワークに導出し、再トレーニングを行わずに拡散モデル用に設計された最先端のサンプリング手法を使用してフロー マッチング トレーニング効率を橋渡しします。 Emyx は最小のモデルであるにもかかわらず、全体的な倍率回復と触媒の幾何学的精度、構造の新規性、足場の多様性、および幾何学的妥当性の両方を必要とする厳格な評価の下で、成功率全体で AME 酵素設計ベンチマークに対して Prote\'ina-Complexa と RFdiffusion3 の両方を上回っています。その一方で、トレーニング時間はわずか 682 ドルの GPU 時間で、RFdiffusion3 よりも約 4 倍少ないです。

原文 (English)

Emyx: Fast and efficient all-atom protein generation

Computational enzyme design requires generating proteins that scaffold catalytic residues and ligands, a task that demands both geometric accuracy and structural diversity from the underlying generative model. Current all-atom generators inherit expensive architectures from structure prediction, leading to high training costs and limited sample diversity. We argue that much of this complexity is unnecessary for generators, which condition on sparse geometric constraints rather than rich co-evolutionary signals. Emyx is a 140M-parameter conditional flow matching model that concentrates capacity within standard transformer blocks, replacing heavy embedding stacks with lightweight conditional representations and sparse connectivity. We additionally derive an exact reparametrisation of the flow matching interpolant into the EDM noise-level framework, bridging flow matching training efficiency with state-of-the-art sampling methods designed for diffusion models without retraining. Despite being the smallest model, Emyx outperforms both Prote\'ina-Complexa and RFdiffusion3 against the AME enzyme design benchmark across success rate under strict evaluation requiring both global fold recovery and catalytic geometry accuracy, structural novelty, scaffold diversity, and geometric validity, while training in just $682$ GPU-hours, roughly $4\times$ less than RFdiffusion3.

13:00 JSTLLM/生成AIGPT / ChatGPTLlama

変圧器フィードフォワード ブロックはどの程度線形ですか?ブロックごとの線形回復性はアーキテクチャではなく学習される

トランスフォーマー フィードフォワード ネットワーク (FFN) は、計算の非線形ストアとして扱われることがよくありますが、トレーニングされた FFN ブロックが実際にどの程度非線形であるかはほとんど測定されていません。各 FFN を位置に関する入力から出力へのマップとして扱い、それを正確な最小二乗線形近似と残差に分割します。閉形式線形マップが説明するホールドアウト分散は、ブロックの線形回復可能性 (R^2_lin)、つまりオプティマイザーを使用しないブロックの線形性の尺度を定義します。 GPT-2、Pythia-160m、および llama-160m の 12 ブロックすべてにわたって、R^2_lin は非常に不均一で深さが非単調であり、隣接するブロック間でほぼ線形 (>0.99) から強い非線形 (<0.3) までの範囲にあり、活性化関数によって設定されません。同じ幅の GELU モデル GPT-2 と Pythia-160m は大きく異なります。したがって、回復可能性は、アーキテクチャ上の特性ではなく、個々のトレーニングされたブロックの学習された特性です。残差の低ランク双線形プローブは、ゲインが残差非線形性と相関せず、R^2 の数点のみを回復します。回復されない計算は、単一の位置に関する積ではなく、高次構造または分散構造です。この測定は、ターゲットを絞った圧縮信号としても機能します。回復可能なブロックは大規模な単層置換を許可します (GPT-2 の初期の FFN では、+0.77 パープレキシティに対して 8 分の 1 のパラメーターが必要です)。一方、回復可能性の低いブロックは、これが安全でない場合にフラグを立てます。さらに、方法論上の落とし穴も明らかになります。トレーニングされた線形ベースラインは、条件の悪い変圧器の活性化では大幅に収束しない可能性があるため、全体を通して正確な閉形式の最小二乗上限を報告します。

原文 (English)

How Linear Is a Transformer Feed-Forward Block? Per-Block Linear Recoverability Is Learned, Not Architectural

Transformer feed-forward networks (FFNs) are often treated as nonlinear stores of computation, yet how nonlinear a trained FFN block actually is has rarely been measured. We treat each FFN as a position-wise input-to-output map and split it into the exact least-squares linear approximation plus a residual. The held-out variance the closed-form linear map explains defines a block's linear recoverability (R^2_lin), an optimiser-free measure of its linearity. Across all twelve blocks of GPT-2, Pythia-160m, and llama-160m, R^2_lin is highly heterogeneous and non-monotone with depth, ranging from near-linear (>0.99) to strongly nonlinear (<0.3) between adjacent blocks, and is not set by the activation function: same-width GELU models GPT-2 and Pythia-160m have sharply different profiles, so recoverability is a learned property of individual trained blocks, not an architectural one. A low-rank bilinear probe of the residual recovers only a few points of R^2, with gain uncorrelated with residual nonlinearity: the unrecovered computation is not a single position-wise product but higher-order or distributed structure. The measurement also serves as a targeted compression signal: recoverable blocks admit large single-layer replacements (GPT-2's early FFN at 8x fewer parameters for +0.77 perplexity), while low-recoverability blocks flag where this is unsafe. It further exposes a methodological pitfall: trained linear baselines can badly under-converge on ill-conditioned transformer activations, so we report the exact closed-form least-squares ceiling throughout.

13:00 JST研究/論文

コードミキシングによるガイド付き合成音声によるコードスイッチング ASR の改善

コードスイッチ (CS) 自動音声認識 (ASR) は、トレーニング用に利用できる高品質の CS テキストと音声のペアが限られているため、依然として課題が残っています。 Text-to-Speech (TTS) による合成データの拡張が検討されていますが、既存の CS TTS アプローチは主に再構築の忠実度を最適化し、言語境界の一貫性を明示的に強制しないため、CS ASR 拡張の有効性が制限されています。この論文では、コード ミキシング インデックス (CMI) を使用して、コード スイッチングの忠実度の向上に向けて合成音声生成を誘導する、コード ミキシングに基づく優先学習フレームワークを提案します。 SEAME 北京語 - 英語会話コーパスの実験により、提案された方法が ASR 微調整のための合成データの有用性を高めることが実証されました。具体的には、Whisper Large を微調整する場合、提案されたアプローチにより、DevMAN セットと DevSGE セットで混合エラー率 (MER) がそれぞれ 12.1%/17.8% から 8.9%/14.2% に減少します。

原文 (English)

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored, existing CS TTS approaches primarily optimise reconstruction fidelity and do not explicitly enforce language-boundary consistency, thereby limiting their effectiveness for CS ASR augmentation. This paper proposes a code-mixing guided preference-learning framework that steers synthetic speech generation toward improved code-switching fidelity using the Code Mixing Index (CMI). Experiments on the SEAME Mandarin-English conversational corpus demonstrate that the proposed method enhances the utility of synthetic data for ASR fine-tuning. Specifically, when fine-tuning Whisper Large, the proposed approach reduces Mixed Error Rate (MER) from 12.1%/17.8% to 8.9%/14.2% on the DevMAN and DevSGE sets, respectively.

13:00 JSTLLM/生成AIエージェント

DynAMO:トポロジカル マルチエージェント スケジューリングによる動的な資産管理オーケストレーション

LLM を利用したエージェントは産業用資産のライフサイクルにエンドツーエンドの自動化を提供しますが、実際のインダストリー 4.0 の展開は、遅延、同時実行の不安定性、安全性リスクによって妨げられています。ここでは、Plan-then-Execute アーキテクチャを使用して検証可能なワークフロー グラフを生成する、すぐに導入できるエンジンである DynAMO (Dynamic Asset Management Orchestration) を紹介します。 DynAMO は、SequentialWorkflow (トポロジー実行) と ParallelWorkflow (依存関係を意識した同時実行) の両方をサポートします。 DynAMO は、独立したタスクを動的に識別することで、構造の正確性と安全性を維持しながら、制御された推論の重複により効率を大幅に向上させます。 AssetOpsBench 産業ベンチマークに関する 6 つの制御された実験を通じて、DynAMO は大幅なパフォーマンスと堅牢性の向上を実証しました。並列実行により、エンドツーエンドのレイテンシがシーケンシャル オーケストレーションと比較して中央値 1.6 倍削減され、高度に並列化可能なワークフローでは 1.8 倍まで削減されます。外部ツール呼び出しを現実的なレイテンシで計測した後、レイテンシを分解すると、LLM 推論とオーケストレーションが依然として実行時間の 90% 以上を占め、モデル推論が主要なシステム ボトルネックであることがわかります。構造化コンテキスト プルーニングにより、推論レイテンシが約 30% 削減され、DynAMO は、制御されたフォールト インジェクションの下で正常な低下を示しながら、正しい機能動作 (タスクの完了、エージェントのシーケンス、出力品質) を維持します。再現性分析により、並列スケジューリングによりレイテンシーの変動が低減され、繰り返し実行しても安定した実行がさらに確認されます。これらの調査結果により、DynAMO は、インダストリー 4.0 自動化パイプラインにおけるスケーラブルで安全な、遅延を考慮したエージェント展開のための実用的な青写真として確立されます。コードはhttps://github.com/kushwaha001/DynAMOで入手できます。

原文 (English)

DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling

While LLM-powered agents offer end-to-end automation for industrial asset lifecycles, real-world Industry 4.0 deployment is hindered by latency, concurrency instability, and safety risks. We present DynAMO (Dynamic Asset Management Orchestration), a deployment-ready engine using a Plan-then-Execute architecture to generate verifiable workflow graphs. DynAMO supports both SequentialWorkflow (topological execution) and ParallelWorkflow (dependency-aware concurrency). By dynamically identifying independent tasks, DynAMO preserves structural correctness and safety while significantly improving efficiency through controlled reasoning overlap. Across six controlled experiments on the AssetOpsBench industrial benchmark, DynAMO demonstrates substantial performance and robustness gains. Parallel execution reduces end-to-end latency by a median of 1.6x over sequential orchestration, rising to 1.8x on highly parallelizable workflows. After instrumenting external tool calls with realistic latencies, a latency decomposition shows that LLM reasoning and orchestration still account for more than 90% of execution time, identifying model inference as the primary system bottleneck. Structured context pruning reduces inference latency by approximately 30%, and DynAMO maintains correct functional behaviour (task completion, agent sequencing, and output quality) while exhibiting graceful degradation under controlled fault injection. Reproducibility analysis further confirms stable execution under repeated runs, with parallel scheduling reducing latency variance. These findings establish DynAMO as a practical blueprint for scalable, safe, and latency-aware agent deployment in Industry 4.0 automation pipelines. Code is available at: https://github.com/kushwaha001/DynAMO

13:00 JSTエージェント

構造による双安定性: 壁時計で調整された状態モニターには、エージェントのリズムでの瞬間検出機構がありません

自律エージェントのランタイム モニターは通常、蓄積された内部状態 (行動ベースライン、ドリフト統計、または以前の研究ではモデル化された感情状態) を閾値に設定します。私たちは以前、状態飽和トラップを報告しました。継続的影響エンジン上のしきい値オン状態トリガーが、SWE ベンチ デバッグ エージェント上でほぼ一定のアラームになる (Modgil 2026)。リリース後の監査により、エンジンがアクション間で dt=0 を受信したため、指数関数的な減衰が動作しなかったことが判明しました。公開されたトラップは純粋なアキュムレータの結果です。私たちは記録 (正誤表、v2) を修正し、欠陥を実験として扱います。それが明らかにする重要な変数は、モニターのダイナミクスがサンプル時間 (CUSUM などの観測ごと) で校正されるか、実時間 (影響モデルや EMA ベースラインなどの秒単位の半減期) で校正されるかです。固定レートのストリームでは、これらは一致します。エージェント ストリームでは、対話時間が桁違いに変化しますが、変化しません。 20 個の軌道上の均一な間隔 (dt 単位: {0..600} 秒) にわたる事前に登録されたスイープは、壁時計レベルのトリガーに 2 つのレジームがあることを示しています: dt=60 秒のサイレント。すべてのクリティカル dt は (1,30] 秒内にあります。実際のエージェントの実行では、中央値 1.53 秒 (p90 2.33 秒) でレイテンシが測定されます。実際のコーディング リズムはトラップ領域内にあり、修正されたメカニズムの下での経験的結果が証明されています。構造はエンジンではなくキャリブレーション クラスのプロパティです。生のエラー ストリームに対する最小の壁時計アキュムレータは同じクリフを再現しますが、サンプル時間 CUSUM は同じクリフを再現します。ストリームは正確に dt 不変 (20/20) です。ヒステリシスのある立ち上がりエッジ トリガーは、すべての条件で軌道ごとに 0 ~ 3 回発生します。ウォールクロックで校正されたリーキー積分器モニターは、エージェント ストリームの瞬間検出器として機能することはなく、すべてのリズムでトラップを回避できますが、人間の介入タイミングは回復しません。

原文 (English)

Bistable by Construction: Wall-Clock-Calibrated State Monitors Have No Moment-Detection Regime at Agent Cadence

Runtime monitors for autonomous agents commonly threshold an accumulated internal state - a behavioural baseline, a drift statistic, or, in our prior work, a modelled affective state. We previously reported a State Saturation Trap: threshold-on-state triggers over a continuous affect engine become near-constant alarms on SWE-bench debugging agents (Modgil 2026). A post-release audit found the engine received dt=0 between actions, so its exponential decay never operated: the published trap is a pure-accumulator result. We correct the record (erratum, v2) and treat the flaw as an experiment. The key variable it exposes is whether a monitor's dynamics are calibrated in sample time (per observation, as in CUSUM) or wall-clock time (half-lives in seconds, as in affect models and EMA baselines). On fixed-rate streams these coincide; on agent streams, where inter-action time varies by orders of magnitude, they do not. A pre-registered sweep over uniform intervals (dt in {0..600}s) on 20 trajectories shows the wall-clock level trigger has two regimes: at dt=60s silent. Every critical dt lies in (1,30]s. Real agent runs measure latency at median 1.53s (p90 2.33s); real coding cadence sits inside the trap regime, vindicating the empirical finding under a corrected mechanism. The structure is a property of the calibration class, not the engine: a minimal wall-clock accumulator over the raw error stream reproduces the same cliff, while a sample-time CUSUM over the identical stream is exactly dt-invariant (20/20). A rising-edge trigger with hysteresis fires 0-3 times per trajectory in every condition. We conclude that wall-clock-calibrated leaky-integrator monitors admit no regime in which they act as moment detectors on agent streams; transition detection escapes the trap at every cadence, but does not recover human intervention timing.

13:00 JSTLLM/生成AI

LLM 主導の段階的改良による解釈可能かつ検証可能なハードウェア生成

大規模言語モデル (LLM) は、ソフトウェア開発において目覚ましい成功を収めています。ただし、幻覚の影響を受けやすいため、微妙な意味的および論理的エラーが発生する可能性があります。チップの設計と製造には大きなリスクが伴うため、ハードウェア エンジニアは依然としてレジスタ転送レベル (RTL) の生成に LLM に依存することに消極的です。この論文では、LLM の創造性と幅広い知識を、形式的手法の説明可能性と数学的厳密性と組み合わせたハードウェア生成フレームワークを提案します。具体的には、さまざまな設計上の決定とハードウェア機能をカバーする一連の変換ルールを考案します。これらのルールを繰り返し適用することで、LLM エージェントは、正確性が保証された設計仕様を RTL プログラムに変換できます。実験結果は、フレームワークの有効性と効率性を示しています。

原文 (English)

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement

Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meaning that they can introduce subtle semantic and logical errors. Due to the high stakes in chip design and manufacturing, hardware engineers are still reluctant to rely on LLMs for register-transfer level (RTL) generation. In this paper, we propose a hardware generation framework that combines the creativity and broad knowledge of LLMs with the explainability and mathematical rigor of formal methods. Specifically, we devise a set of transformation rules that cover various design decisions and hardware features. By iteratively applying these rules, an LLM agent can convert a design specification into an RTL program with guaranteed correctness. Experimental results demonstrate the effectiveness and efficiency of the framework.

13:00 JSTエージェントClaude

エージェント AI 向けの実行限定アドバイザリー自動化: 再現可能な AIBOM 主導の CSAF-VEX フレームワーク

SBOM および AIBOM アーティファクトを決定論的な環境キャプチャおよび構造化されたランタイム テレメトリにバインドする、プロトコル駆動のフレームワークが提供されます。悪用可能性は、宣言されたアーティファクト、観察されたアクティブ化条件、および強制された実行ポリシーから計算されます。 CSAF VEX アドバイザリは、静的証拠と実行時の証拠を組み合わせて生成され、暗号的に署名され、決定論的な再生によって検証されます。評価では、OSV、GitHub Advisory、KEV、EPSS データセットを組み込んだ、合成 Agentic AI ワークロード 50 ~ 5000 コンポーネントにわたって約 10000 コンポーネント エントリを使用します。

原文 (English)

Execution-bound advisory automation for agentic AI: a reproducible AIBOM-driven CSAF-VEX framework

A protocol driven framework is presented that binds SBOM and AIBOM artefacts to deterministic environment capture and structured runtime telemetry. Exploitability is computed from declared artefacts, observed activation conditions, and enforced execution policies. CSAF VEX advisories are generated from combined static and runtime evidence, cryptographically signed, and validated through deterministic replay. Evaluation uses approximately 10000 component entries across synthetic Agentic AI workloads 50 to 5000 components, incorporating OSV, GitHub Advisory, KEV, and EPSS datasets.

13:00 JSTLLM/生成AI

VERITAS: Verifier-Guided Proof Search for Zero-Shot Formal Theorem Proving

LLM-based formal provers often collapse rich verifier signals (syntax errors, type mismatches, partial goal progress) into a binary pass/fa…

13:00 JST研究/論文

JustDiag!: A Diagnostic Justification Engine for Accountable Root Cause Analysis

Large language models can produce fluent root cause analyses, but fluent final answers alone are insufficient evidence for accountability i…

13:00 JSTエージェントロボティクス

Playful Agentic Robot Learning

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts,…

13:00 JST画像/動画生成

Scaling Generative Foundation Models for Chest Radiography with Rectified Flow Transformers

We introduce the first generative foundation model for chest radiograph synthesis trained from scratch at the billion-parameter scale. Exis…

13:00 JSTLLM/生成AI

Secure Coding Drift in LLM-Assisted Post-Quantum Cryptography Development: A Gamified Fix

The transition to Post Quantum Cryptography (PQC) introduces considerable implementation complexity, requiring strict adherence to constant…

13:00 JST研究/論文

Can In-Context Learning Support Intrinsic Curiosity?

Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models h…

13:00 JST研究/論文

Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks

Concept Bottleneck Models (CBMs) enhance interpretability by projecting learned features into a human-understandable concept space. Recent…

13:00 JSTLLM/生成AILlamaQwen

Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices

Fine-tuning of Large Language Models (LLMs) using Low-Rank Adaptation (LoRA) on an end-user's data offers personalized experiences while ke…

13:00 JST研究/論文

A Tool for the Synthesis of Adaptive Probabilistic Processors Based on the Ising Model

This work presents a tool for the synthesis and simulation of probabilistic architectures for solving combinatorial optimization problems b…

13:00 JSTLLM/生成AI画像/動画生成

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely…

13:00 JST研究/論文

Review of Machine Learning Models for Solar Energetic Particle Prediction

Solar energetic particle (SEP) events have attracted increasing attention due to their significant radiation hazards for aviation, spacecra…

13:00 JST研究/論文

GDGU: A Gradient Difference-based Graph Unlearning Method for Cyberattack Localization in Electric Vehicle Charging Networks

Electric vehicle charging stations (EVCSs) can expose distribution feeders to cyberattacks. While machine learning methods, including graph…

13:00 JST研究/論文

Exploring Feature Extraction Technique Parameters for Acoustic Gunshot Classification

Acoustic gunshot detection is a problem with applications across civilian public safety, military operations, and wildlife conservation, ye…

13:00 JST研究/論文

FlowFake: Liquid Networks for Audio Deepfake Detection

Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. T…

13:00 JSTLLM/生成AI

A BART-based approach with hierarchical strategy for Vietnamese abstractive multi-document summarization

In this technical report, we focus on solving the challenge of Vietnamese multi-document abstractive summarization, introduced in the Inter…

13:00 JSTエージェントOpenAIGoogle

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user inter…

13:00 JST研究/論文

PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening…

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPT

FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines

Multi-step LLM pipelines fail through interactions among retrieval, reasoning, and formatting steps, so prompt-only optimization can miss b…

13:00 JST研究/論文

Latent Confounded Causal Discovery via Lie Bracket Geometry

Recent work on Kan-Do-Calculus (KDC) has established that the boundary between passive observation and active intervention in causal infere…

13:00 JSTエージェント研究/論文

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns

We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests)…

13:00 JSTエージェント

Before the Pull Request: Mining Multi-Agent Coordination

Autonomous coding agents now open millions of pull requests, yet large-scale studies find their PRs are produced faster but accepted less o…

13:00 JST研究/論文

VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions

The digital commerce landscape is shifting from static, search-driven catalogs to dynamic, immersive video feeds. This transition introduce…

13:00 JST研究/論文

RIVET: Robust Idempotent Voice Attribute Editing

Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datas…

13:00 JSTエージェントロボティクス

Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation

Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural poli…

13:00 JSTロボティクス

CTS-MoE: Implicit Terrain Adaptation via Mixture-of-Experts for Perceptive Locomotion

Perceptive legged locomotion over discontinuous terrain (e.g., stairs, gaps, and obstacles) requires adaptive behavior, as a single conserv…

13:00 JST研究/論文

Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models

Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically i…

13:00 JST研究/論文

Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation

Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficul…

13:00 JSTLLM/生成AI

Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text

Clinical NLP increasingly relies on electronic health record (EHR) data to detect suicidal behaviors, treating clinical documentation as mo…

13:00 JSTLLM/生成AI

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language

AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges. Despite the global nature o…

13:00 JST画像/動画生成

TeleMorpher: Toward Robust Simultaneous Motion-Location Editing

Diffusion models have achieved remarkable success in image and video generation and editing. While recent studies have extended these effor…

13:00 JST研究/論文

LOKI: Memory-Free Null-Space Constrained Lifelong Knowledge Editing

Lifelong knowledge editing aims to efficiently and sequentially update language models over time, as new knowledge becomes available or whe…

13:00 JSTLLM/生成AIハードウェア/半導体

Efficiently Representing Algorithms With Chain-of-Thought Transformers

The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producin…

13:00 JSTLLM/生成AI

FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs

Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargo…

13:00 JSTLLM/生成AIビジネス/資金調達

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive…

13:00 JST研究/論文

OnDeFog: Online Decision Transformer under Frame Dropping

In challenging real-world reinforcement learning applications, communication delays or sensor failures often cause frame dropping, in which…

13:00 JST研究/論文

Library-Aware Doubles and Iterative Repair for Large Language Model-Generated Unit Tests in OpenSIL Firmware

Validating changes in low-level C firmware is expensive because unit tests (UTs) are fragile under strict build constraints, where missing…

13:00 JSTLLM/生成AI

NRITYAM: Language Models Meet Art and Heritage of Dance

Language models have become essential tools in shaping modern workflows. However, their global effectiveness hinges on a nuanced understand…

13:00 JSTロボティクス

Bidirectional Tutoring for Developmental Motor Learning in Robots: Co-Developed Interaction Dynamics Support Stable Learning

Infants are well known to develop their motor skills through dense interaction with caregivers. Although such social interaction is crucial…

13:00 JSTエージェントロボティクス

VOiLA: Vectorized Online Planning with Learned Diffusion Model for POMDP Agents

Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provi…

13:00 JSTLLM/生成AI画像/動画生成

QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval

Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multime…

13:00 JSTLLM/生成AILlama

Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings

Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply…

13:00 JSTLLM/生成AI

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models

Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training effi…

13:00 JSTロボティクス

Temporal Self-Imitation Learning

Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while…

13:00 JSTLLM/生成AI

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses…

13:00 JSTロボティクス

Data Standards for Humanoid Robotics: The Missing Infrastructure for Physical AI

The scalability of humanoid robots will depend not only on models and hardware, but also on whether physical experience can accumulate acro…

13:00 JST研究/論文

Towards Engineering Scaling Laws with Pretraining Data Composition

Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established…

13:00 JST研究/論文

Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired ar…

13:00 JST研究/論文

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired ar…

13:00 JSTエージェント

Agentic Electronic Design Automation: A Handoff Perspective

Electronic design automation (EDA) is inherently multi-stage and handoff-heavy. Design artifacts, flow scripts, and engineering decisions c…

13:00 JST研究/論文

Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria. However, accurately re…

13:00 JST研究/論文

Policy-aware Vector Search: A Vision for Fine Grained Access Control in Vector Databases

Vector databases are increasingly used in security sensitive contexts with Retrieval Augmented Generation and organizational AI pipelines;…

13:00 JST画像/動画生成

ParaScale: Scale-Calibrated Camera-Motion Transfer via a Gauge-Invariant Parallax Number

Transferring the camera motion of a reference video to a freshly generated one lets creators reuse cinematic moves. Yet reference and targe…

13:00 JST研究/論文

Uncertainty-Aware Reward Modeling for Stable RLHF

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing…

13:00 JSTLLM/生成AI

CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility -- Semantic Metrics and Convergence Analysis

Decomposing compound sentences into atomic, verifiable claims is a prerequisite for reliable automated fact-checking. Prior work has relied…

13:00 JST画像/動画生成

CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images

Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains…

13:00 JST研究/論文

When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning

Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because reliable labels often…

13:00 JST研究/論文

Neural Additive and Basis Models with Feature Selection and Interactions

Deep neural networks (DNNs) exhibit attractive performance in various fields but often suffer from low interpretability. The neural additiv…

13:00 JSTLLM/生成AI

Large Language Models Do Not Always Need Readable Language

Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is an…

13:00 JST画像/動画生成

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement

Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomie…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance…

13:00 JST研究/論文

SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models

Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates,…

13:00 JSTエージェント研究/論文

Measuring Biological Capabilities and Risks of AI Agents

This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities…

13:00 JSTロボティクスQwen

Co-policy: Responsive Human-Robot Co-Creation for Musical Performances

Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to pa…

13:00 JST画像/動画生成

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced M…

13:00 JST画像/動画生成

Speeding up the annotation process in semantic segmentation industrial applications

Current machine learning models commonly require large and well-annotated datasets. However, the annotation process often becomes a bottlen…

13:00 JST画像/動画生成

Triangular Consistency as a Universal Constraint for Learning Optical Flow

We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision…

13:00 JST研究/論文

SIMBA: ABidirectional Retrieval Forward Simulation Framework for Modeling FY-4A GIIRS Hyperspectral Infrared Radiances Toward NWP Applications

Hyperspectral infrared observations are an important data source for numerical weather prediction (NWP) because they provide rich informati…

13:00 JSTLLM/生成AI画像/動画生成

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA

Multimodal Large Language Models (MLLMs) show great potential in medical tasks, but their elicited confidence often misaligns with actual a…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models

Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different a…

13:00 JST研究/論文

The Algorithmic-Human Manager: AI, Apps, and Workers in the Indian Gig Economy

This paper examines the impact of artificial intelligence and digital technologies on the blue-collar gig economy in India, focusing on alg…

13:00 JSTLLM/生成AIエージェント

Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services

In the agentic web era, LLM-based agents increasingly invoke web services as tools, yet most interfaces remain \emph{static endpoints} that…

13:00 JST画像/動画生成ロボティクス

Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions…

13:00 JSTLLM/生成AIエージェント

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by…

13:00 JST研究/論文

StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation

Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is w…

13:00 JSTLLM/生成AIエージェント

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution

Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments…

13:00 JSTLLM/生成AIエージェント

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However,…

13:00 JSTLLM/生成AIロボティクス

A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems

Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Syst…

13:00 JSTエージェント

AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models

We propose a model-grounded RAG-based AI economist with an agentic framework for economic scenario analysis using large language models (LL…

13:00 JST画像/動画生成

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery…

13:00 JSTビジネス/資金調達

Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU

Burst suppression (BS) is a clinically relevant electroencephalographic (EEG) pattern used to monitor sedation depth and brain activity in…

13:00 JST画像/動画生成

Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers

Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the…

13:00 JSTLLM/生成AI画像/動画生成

The Hidden Evolution of Disguised Visual Context inside the VLM

Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and inte…

13:00 JSTLLM/生成AI

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insuffi…

13:00 JST画像/動画生成

MakeupMirror: Improving Facial Attribute Preservation in Diffusion Models for Makeup Transfer

Makeup transfer models enable fun augmented reality (AR) experiences as well as virtual try-on (VTO) for online makeup shopping. While rece…

13:00 JST研究/論文

Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow

Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the re…

13:00 JST研究/論文

Sensorimotor World Models: Perception for Action via Inverse Dynamics

Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for…

13:00 JSTエージェントロボティクス

Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform

Biological experiment protocols are written in natural language, whereas automation systems rely on predefined control commands, creating a…

13:00 JSTロボティクス

Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation

Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power for modelling complex, multi…

13:00 JST研究/論文

Hybrid ANN-SNN Pipeline with Local Plasticity

This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs)…

13:00 JSTLLM/生成AI

From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models

Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms u…

13:00 JSTLLM/生成AI

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

Real-world clinical decision support requires reasoning over heterogeneous and longitudinal patient information rather than answering isola…

13:00 JSTLLM/生成AI画像/動画生成

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability…

13:00 JST画像/動画生成ロボティクス

HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-trainin

Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annota…

13:00 JSTロボティクス

FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching

Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments. Such…

13:00 JSTビジネス/資金調達

Learner-based Concept Drift Detection: Analysis and Evaluation

Machine learning algorithms deployed for evolving streaming environments must handle the non-stationary data distributions, commonly referr…

13:00 JSTLLM/生成AIエージェント研究/論文

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

Academic paper search is a core step in scientific research, and LLM-based search agents are emerging as a promising paradigm for iterative…

13:00 JST画像/動画生成

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and eas…

13:00 JSTロボティクス

Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-b…

13:00 JSTLLM/生成AIビジネス/資金調達Gemini

The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse

We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that…

13:00 JSTLLM/生成AI

Editorial Alignment: A Participatory Approach to Engaging Editorial Expertise in LLM-mediated Knowledge Dissemination

The emergence of LLM-driven information services is reshaping the conditions under which public knowledge institutions operate, threatening…

13:00 JSTLLM/生成AI

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval

Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance…

13:00 JST研究/論文

Boundary Embedding Shaping with Adaptive Contrastive Learning for Graph Structural Disentanglement

Graph neural networks (GNNs) excel at aggregating neighbor information for classification, yet their performance is hindered by graph struc…

13:00 JST研究/論文

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise

In this article, we present a robust $Q$-learning algorithm for discrete-time mean-field control problems under Wasserstein uncertainty in…

13:00 JSTLLM/生成AIエージェント

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Large Language Models (LLMs) show promise for code compilation tasks, but applying them to runtime performance tuning is difficult due to c…

13:00 JSTエージェントロボティクス研究/論文

CRAX: Fast Safe Reinforcement Learning Benchmarking

Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. Wh…

13:00 JST研究/論文

DataMagic: Transforming Tabular Data into Data Insight Video

Data videos integrate dynamic charts, voice narration, and synchronized animations to communicate data insights as temporal narratives, mak…

13:00 JSTLLM/生成AIエージェント研究/論文

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness und…

13:00 JSTLLM/生成AI

Multi-View Decompilation for LLM-Based Malware Classification

Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that la…

13:00 JST研究/論文

Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

Classifier guidance is a way to control diffusion generation by using a noise-conditioned classifier to steer the sampling process toward a…

13:00 JSTエージェント

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coord…

13:00 JSTエージェント

UltraQuant: 4-bit KV Caching for Context-Heavy Agents

Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrenc…

13:00 JSTエージェント研究/論文

Optimal Order of Multi-Agent and General Many-Body Systems

This paper develops a general framework for analyzing multi-agent systems with feedback loops between agents actions and collective observa…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達DeepSeek

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

When large language models serve as evaluators in multi-agent systems, their systematic evaluation biases propagate through the agent netwo…

13:00 JSTLLM/生成AI研究/論文DeepSeek

Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software

Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains…

13:00 JST画像/動画生成

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while…

13:00 JSTエージェント

Efficient and Sound Probabilistic Verification for AI Agents

Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulat…

13:00 JSTエージェント

Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not…

13:00 JST画像/動画生成研究/論文

SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture rada…

13:00 JST研究/論文

Structuring and Tokenizing Distributed User Interest Context for Generative Recommendation

Generative recommendation is an emerging paradigm that has shown promise in industrial recommendation systems, aiming to predict users' nex…

13:00 JSTLLM/生成AIGemma

How Transparent is DiffusionGemma?

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging su…

13:00 JSTエージェントロボティクス

UniMM: A Unified Mixture Model Framework for Multi-Agent Simulation

Simulation plays a crucial role in assessing autonomous driving systems, where the generation of realistic multi-agent behaviors is a key a…

13:00 JSTエージェント研究/論文

MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning

Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. De…

13:00 JST研究/論文

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

Post-training alignment of large language models often combines supervised fine-tuning (SFT) on expert demonstrations with reinforcement le…

13:00 JST研究/論文

Controlled Comparison of Machine Learning Models for Fault Classification and Localization in Power System Protection

The increasing complexity of modern power systems, driven by the integration of inverter-based and distributed energy resources, challenges…

13:00 JSTLLM/生成AIエージェント

SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning

Solving mathematical reasoning problems requires not only accurate access to relevant knowledge but also careful, multi-step thinking. Howe…

13:00 JST研究/論文

Creativity Reconsidered: Generative AI and the Problem of Intentional Agency

Many theorists maintain that conscious intentional agency is a necessary condition of creativity. We argue that this requirement, which we…

13:00 JSTLLM/生成AI研究/論文Gemma

PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Boards (PCB) Schematic Design with Structured Verification

Most LLM code-synthesis benchmarks rely on unit tests as the reward oracle, but PCB schematic design has none: correctness is defined by st…

13:00 JST研究/論文

One Probe Won't Catch Them All: Towards Targeted Deception Detection

Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier…

13:00 JST研究/論文

Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach

We study conditional generation in diffusion models under hard constraints, where generated samples must satisfy prescribed events with pro…

13:00 JST研究/論文

SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures

While the shift toward unified foundation models has revolutionized many deep learning domains, sleep medicine remains largely restricted t…

13:00 JSTハードウェア/半導体

Mitigating Legibility Tax with Decoupled Prover-Verifier Games

As large language models become increasingly capable, it is critical that their outputs can be easily checked by less capable systems. Prov…

13:00 JST研究/論文

PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units

Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typic…

13:00 JSTLLM/生成AIビジネス/資金調達

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evalua…

13:00 JST研究/論文

CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions

Incomplete or inconsistent discharge documentation drives care fragmentation and avoidable readmissions. Despite its critical role in patie…

13:00 JST研究/論文

Too long; didn't solve

Mathematical benchmarks consisting of a range of mathematics problems are widely used to evaluate the reasoning abilities of large language…

13:00 JSTLLM/生成AIエージェント

CogniFold: コグニティブフォールディングによる常時オンのプロアクティブなメモリ

既存のエージェントの記憶は主に反応的かつ検索ベースのままであり、経験を自律的に永続的な認知構造に組織化する能力が欠けています。真の自律型エージェントを目指して、次世代のプロアクティブ アシスタント向けに設計された、脳からインスピレーションを得た「常時オン」エージェント メモリである CogniFold を紹介します。 CogniFold は、断片化されたイベント ストリームを自己出現の認知構造に継続的に折り畳んで、入ってくるイベントと蓄積された知識から徐々により高いレベルの認知をブートストラップします。私たちは、相補学習システム (CLS) 理論を 2 層 (海馬、新皮質) から 3 層に拡張し、前頭前意図層を追加することでこれを根拠にしています。意図的な制御と意思決定の拠点として前頭前野をエミュレートする CogniFold は、グラフ トポロジーの自己組織化を通じてこれを実現します。つまり、認知構造はストリームの下で積極的に集まり、意味的に類似している場合は結合し、古くなっている場合は減衰し、連想想起を通じて再リンクし、概念クラスターの密度がしきい値を超えると意図を表面化します。 CogEval-Bench を使用して構造形成を評価し、CogniFold が認知的期待と概念創発に一致する記憶構造を独自に生成することを実証します。さらに、5 つの認知ドメインにわたる 7 つの広範なベンチマークにわたって、CogniFold が従来のメモリ ベンチマークでも同時に堅牢に実行されることを検証しました。

原文 (English)

CogniFold: Always-On Proactive Memory via Cognitive Folding

Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genuinely autonomous agents, we introduce CogniFold, a brain-inspired "always-on" agent memory designed for the next generation of proactive assistants. CogniFold continuously folds fragmented event streams into self-emerging cognitive structures, bootstrapping progressively higher-level cognition from incoming events and accumulated knowledge. We ground this by extending Complementary Learning Systems (CLS) theory from two layers (hippocampus, neocortex) to three, adding a prefrontal intent layer. Emulating the prefrontal cortex as the locus of intentional control and decision-making, CogniFold achieves this through graph-topology self-organization: cognitive structures proactively assemble under the stream, merge when semantically similar, decay when stale, relink through associative recall, and surface intents when concept-cluster density crosses a threshold. We evaluate structural formation using CogEval-Bench, demonstrating that CogniFold uniquely produces memory structures that match cognitive expectations and concept emergence. Furthermore, across eight downstream benchmarks -- two probing long-term conversational memory (LoCoMo, LongMemEval) and six spanning other cognitive domains -- we validate that CogniFold simultaneously performs robustly on conventional memory tasks. Our code is available at https://github.com/OpenNorve/CogniFold.

13:00 JSTエージェントビジネス/資金調達

SimuWoB: 高速かつ忠実な GUI エージェント ベンチマークのための現実世界のモバイル アプリのシミュレーション

大規模な言語モデルを利用したモバイル GUI エージェントは急速に進歩しており、現実的かつ包括的な評価に対する緊急のニーズが生じています。既存のベンチマークは再現性を優先していますが、実際のアプリケーションで報酬を構築することが難しいため、多くの場合、オープンソース アプリまたはファイル操作タスクに限定されており、ベンチマーク設定と実際の使用状況の間にギャップが生じています。さらに、ほとんどのベンチマークは基本的な接地とナビゲーションに焦点を当てており、複雑で長期にわたる相互作用の範囲は限られています。これらの制限に対処するために、さまざまなタイプと難易度にわたる 120 の困難なタスクを備えたモバイル GUI エージェント用の完全合成ベンチマークである SimuWoB を導入します。私たちは、忠実度の高いタスクと環境を合成し、各タスクに対して有効な報酬を自動的に提供する、堅牢な仮想環境生成フレームワークを構築します。各環境は、URL 経由でアクセスできるバックエンドのない Web ページとしてデプロイされ、効率的で再現可能な評価が可能になります。私たちは、いくつかの最先端のモバイル GUI エージェントで包括的な実験を実施しています。平均成功率はわずか 27.92% であり、長期的なタスクでは 17.82% に低下します。これは、複雑なシナリオの下での現在のエージェントの重大な弱点を明らかにしています。評価結果を実際のサンプル タスクと比較すると、合成環境に基づくエージェントの評価が一般化していることがわかります。さらに、主要な機能の側面にわたる診断上の洞察を提供し、将来のモバイル GUI エージェント開発への影響について説明します。

原文 (English)

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments. However, directly doing so in real-world environments introduces some challenges that cannot be overlooked. Real-world environments are complex and uncontrollable, making it difficult to construct verifiable rewards and to save or reset states. Existing works prioritize reproducibility but are often limited to open-source apps or file-operation tasks for reliable reward building, leaving a persistent gap from real-world usage. Furthermore, relying on virtual machines or docker images demand high resource requirements and suffer from slow response speeds, which limit the efficiency. We present \sys, a framework that could produce high-fidelity synthesized interactive environments for GUI agents across platforms with verifiable rewards. These environments behave as backend-free webpages accessible via URL, requiring near-zero setup and low resource cost, making the approach suitable for both large-scale evaluation and downstream agent training. We support multiple GUI platforms including mobile, desktop, and automotive/in-vehicle interfaces based on the same pipeline, covering 100+ environments and 1000+ verifiable tasks. Among them, 120 challenging tasks across 63 simulated mobile applications are released as a fully synthesized mobile GUI agent benchmark. Experiment results on five state-of-the-art mobile GUI agents reveal substantial headroom -- the average success rate is only 27.92\%, dropping to 17.82\% on long-horizon subset -- while humans reach 92.08\%. A comparison against real-world sample tasks shows that assessments made in our synthetic environments generalize to real apps. The project website is at https://scalewob.github.io.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

13:00 JSTエージェント

VitalAgent: ウェアラブル健康データに対する反応的および積極的な生理学的モニタリングのためのツール拡張エージェント

ウェアラブル デバイスにより、ECG や PPG などの生理学的信号の継続的なモニタリングが可能になりますが、既存の mHealth システムは、タスク固有の予測パイプラインまたは静的な概要に対する反応的な質問応答に主に限定されています。これらには、時間的推論、永続的な生理学的コンテキスト、および長期的な信号ストリームにわたるプロアクティブなモニタリングをサポートする能力がありません。私たちは、事後的な質問応答とプロアクティブなモニタリングの両方をサポートする、ECG/PPG ベースの mHealth 用のツールを強化したエージェント フレームワークである VitalAgent を提案します。 VitalAgent は、長期的な生理学的メモリと、生の信号に対する動的な計算を可能にするツール拡張推論インターフェイスに基づいて構築されています。さらに、反応的な質問応答のための 1,862 の QA ペアと、心臓、身体活動、ストレス関連のタスクをカバーするプロアクティブなモニタリングのための 90.2 時間の連続 ECG/PPG 記録で構成される長期的な生理学的モニタリング ベンチマーク データセットである VitalBench を紹介します。実験では、VitalAgent が事後評価においてプロンプトベースおよび ReAct ベースラインと比較して 30% 以上の改善を達成し、長期の生理学的信号に対するプロアクティブなアラートモニタリングをサポートすることが実証されており、動的なツールの使用と長期の生理学的モニタリングの重要性が強調されています。

原文 (English)

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-specific prediction pipelines or reactive question answering over static summaries. They lack the ability to support temporal reasoning, persistent physiological context, and proactive monitoring over long-term signal streams. We propose VitalAgent, a tool-augmented agentic framework for ECG/PPG-based mHealth that supports both reactive question answering and proactive monitoring. VitalAgent is built on a longitudinal physiological memory and a tool-augmented reasoning interface that enables dynamic computation over raw signals. We further introduce VitalBench, a longitudinal physiological monitoring benchmark dataset comprising 1,862 QA pairs for reactive question answering and 90.2 hours of continuous ECG/PPG recordings for proactive monitoring, covering cardiac, physical activity, and stress-related tasks. Experiments demonstrate that VitalAgent achieves over 25% improvement over prompt-based and ReAct baselines in reactive evaluation and supports proactive alert monitoring over long-term physiological signals, highlighting the importance of dynamic tool use and long-term physiological monitoring.

13:00 JST研究/論文

Science Earth: AI ネイティブの科学的発見のための地球規模のオペレーティング システムを目指して

科学的発見には、広大な探索空間にわたる知性、忍耐力、偶然の発見が必要です。現在、最高の科学的能力は依然としてサイロ化されており、ある AI システムは生物学的分析用、別の AI システムは臨床推論、数学的導出、材料シミュレーション用というように、質問に必要なすべてのスキルを事前に設計されたチームは予測できません。 Science Earth は地球規模の科学ランタイムであり、シミュレーション クラスター、ウェットラボ ロボット、プルーフ エンジン、シングルセル パイプラインなど、あらゆる機能を他の機能に接続でき、質問自体からコラボレーション構造が生まれます。その基盤となる EACN プロトコルにより、誰が誰と会うのかを事前に知らなくても、各機能が相互に発見し、タスクの所有権を交渉し、互換性のない証拠基準間で裁定を行うことができます。これにより、組織化の課題はワークフロー設計からオープンエンドの接続へと移行します。 2 回の実行により、構造的に異なる条件下でこれが検証されました。太平洋横断の高次倉本同期研究では、エージェントは、ローレンツ限界外で破綻するオット・アントンセン解析理論の閉包率の仮定を 30 分以内に特定し、修正しました。 488 万セルの Kang 2024 汎がんアトラスでの 8 つの薬剤の単一セルの実行では、異種機能が 64.9 時間のウィンドウにわたって 1 つの構造外部命令と結合され、3 つの新しい結果層が生成され、隣接する CCR8-TIGIT+ Treg サブセットに関する独立したウェットラボ研究に対して所見を固定しました。これらのケースは、最初の経験的な読み取りであり、ベンチマークのスイープではありません。彼らは、AI の機能が真に接続可能になり、問題から調整が生まれると、科学的推論が分散型の自己修正プロセスとなり、AI ネイティブの発見を地球規模に拡大するための一歩となることを示しています。

原文 (English)

Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery

Scientific discovery demands intelligence, perseverance, and serendipity across vast search spaces. Today, top scientific capabilities remain siloed--one AI system for biological analysis, another for clinical reasoning, mathematical derivation, or materials simulation--and no pre-designed team can anticipate every skill a question will need. Science Earth is a planet-scale scientific runtime in which any capability--a simulation cluster, a wet-lab robot, a proof engine, a single-cell pipeline--can connect to any other, with collaboration structure emerging from the question itself. Its underlying EACN protocol lets capabilities discover one another, negotiate task ownership, and adjudicate across incompatible evidentiary standards without prior knowledge of who will meet whom. This shifts the organizing challenge from workflow design to open-ended connectivity. Two runs validate this under structurally distinct conditions. In a trans-Pacific higher-order Kuramoto synchronization study, agents identified and corrected a closure-ratio assumption in Ott-Antonsen analytic theory that fails outside the Lorentzian limit, within thirty minutes. In an eight-agent single-cell run on the 4.88M-cell Kang 2024 pan-cancer atlas, heterogeneous capabilities coupled over a 64.9-hour window with one structural external instruction, producing three new result layers and anchoring findings against an independent wet-lab study on an adjacent CCR8- TIGIT+ Treg subset. These cases are a first empirical reading, not a benchmark sweep. They show that when AI capabilities are truly connectable and coordination emerges from the problem, scientific reasoning becomes a distributed, self-correcting process--a step towards scaling AI-native discovery to the planet.

13:00 JSTエージェント

覚えておくべきことの学習: 長期にわたる言語エージェントの制約付き最適化による可観測性と安全なメモリ保持

長期的な言語エージェントは、有限のコンテキスト ウィンドウを超える観察、推論トレース、取得された事実を蓄積するため、メモリ保持がリソース割り当ての基本的な問題になります。既存のメモリ システムは、ヒューリスティック スコアリング、取得の最適化、または学習された圧縮を通じて管理を改善しますが、主に保持をローカルな決定問題として扱い、現実的な可観測性の制約の下でその長期的な結果を明示的にモデル化していません。このギャップを埋めるために、明示的な予算の実現可能性、証拠の有用性、およびミスペナルティ、再取得の遅延、情報の陳腐化リスクを含む遅延コストを伴う制約付き確率的最適化問題として記憶保持を定式化します。次に、OSL-MR (Observability-Safe Learning for Memory Retention) を提案します。これは、オンラインで観察可能な機能とオフラインで利用可能な監視 (OAS) を厳密に分離する新しいフレームワークです。 OSL-MR は、実現された証拠の監督から訓練された証拠学習者と、展開可能なオンラインで安全なベースラインとして、および学習のための構造化された帰納的事前分布として機能する混合スコア ヒューリスティックを組み合わせます。結果として得られるポリシーは、同じ可観測性制約の下で展開可能でありながら、クエリ条件付きの証拠値をインタラクション データから直接学習します。 LOCOMO と LongMemEval の実験では、OSL-MR が、特にメモリ バジェットが厳しい場合に、リーセンシ ベースの手法、生成エージェント スタイルのスコアリング、その他のヒューリスティック ベースラインよりも一貫して優れたパフォーマンスを発揮することが示されています。事前の混合スコアにより、再現率を維持しながら精度がさらに向上し、感度分析により、幅広いコスト構成にわたる堅牢性が実証されます。

原文 (English)

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts exceeding context windows, making memory retention a fundamental resource-allocation problem. Existing systems treat retention as local and do not model long-term consequences under observability constraints. To fill this gap, we formulate memory retention as a constrained stochastic optimization with budget feasibility, evidence utility, and delayed costs including miss, reacquisition, and stale penalties. We show this multi-step problem is NP-hard, making exact solution intractable. Moreover, deployment decisions must be made under partial observability. To address these challenges, we propose OSL-MR (Observability-Safe Learning for Memory Retention), a learning-augmented framework that enforces a strict separation between online-observable features and offline-available supervision. OSL-MR combines an evidence learner trained from realized evidence with a Mixed-Score heuristic that serves as a deployable online-safe baseline and an inductive prior. The policy learns query-conditioned evidence from interaction data and remains deployable under the same constraints. Experiments on LoCoMo and LongMemEval show OSL-MR outperforms recency-based, Generative Agents-style, and other heuristic baselines, especially under tight budgets. The Mixed-Score prior improves precision and recall, and sensitivity analysis shows robustness across cost settings. On small solvable instances, single-step optimization is insufficient to anticipate future demand shifts, while OSL-MR stays significantly closer to the dynamic-programming optimum, confirming the necessity of the sequential formulation and reinforcing our learning-guided approximation. These results establish constrained stochastic optimization and optimization-guided learning as a principled foundation for memory management in long-horizon agents.

13:00 JSTエージェント

MoCA-Agent: 財務および数値推論のためのクレーム市場コード エージェント

財務および表形式の質問に答えるには、流暢な推論以上のものが必要です。回答は、それらを裏付ける正確な事実、公式、単位、記号、尺度に基づいていなければなりません。単一のセルの読み間違いや誤った操作により、もっともらしいが間違った結果が静かに生成される可能性があります。 \textsc{MOCA-Agent} は、自由形式の複数エージェントによる議論を請求レベルの検証に置き換える、請求市場コード エージェントです。このシステムは、各質問を型指定されたアトミックなクレームに分解し、専門トレーダーのエージェントにそれらのクレームを売買するよう依頼し、注文を信頼度に重み付けされた受諾/拒否の決定にクリアし、市場でサポートされた証拠から実行可能な Python プログラムを合成します。次に、コード認識検証者がプログラムの実行、構造の一貫性、一般的な財務上の推論エラーをチェックし、最大 1 回の市場認識の修復ラウンドを実行します。 \textsc{MOCA-Agent} は、財務数値推論、一般的な表形式推論、ESG 質問回答、マルチモーダル チャート推論にわたる 10 の公開ベンチマークにわたって、固定 Qwen3.6-27B バックボーンを使用して優れたパフォーマンスを達成します。これには、FinQA で $78.3\%$、FinanceMath で $76.0\%$、MultiHiertt で $71.2\%$、ESGenius で $86.9\%$ が含まれます。 FinChart-Bench では平均 $85.6\%$ です。これらの結果は、答え全体ではなく、原子の主張のレベルで証拠を集約することで、一か八かの数値推論における堅牢性が向上することを示しています。\footnote{コードとデータは、https://github.com/UBC-NLP/MoCA-Agent から入手できます。

原文 (English)

MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning

Financial and tabular question answering requires more than fluent reasoning: answers must be grounded in the exact facts, formulas, units, signs, and scales that support them. A single misread cell or incorrect operation can silently produce a plausible but wrong result. We introduce \textsc{MOCA-Agent}, a market-of-claims code agent that replaces free-form multi-agent debate with claim-level verification. The system decomposes each question into typed atomic claims, asks specialist trader agents to buy or sell those claims, clears their orders into confidence-weighted accept/reject decisions, and synthesizes an executable Python program from market-supported evidence. A code-aware verifier then checks the program for execution, structural consistency, and common financial reasoning errors, with at most one market-aware repair round. Across ten public benchmarks spanning financial numerical reasoning, general tabular reasoning, ESG question answering, and multimodal chart reasoning, \textsc{MOCA-Agent} achieves strong performance using a fixed Qwen3.6-27B backbone, including $78.3\%$ on FinQA, $76.0\%$ on FinanceMath, $71.2\%$ on MultiHiertt, $86.9\%$ on ESGenius, and $85.6\%$ average on FinChart-Bench. These results show that aggregating evidence at the level of atomic claims, rather than whole answers, improves robustness in high-stakes numerical reasoning.\footnote{The code and data are available: https://github.com/UBC-NLP/MoCA-Agent.

13:00 JST研究/論文

薬物と疾患の関係治療における適用条件抽出

特定の薬剤が標的疾患に対して治療効果を発揮する条件を特定することは、臨床上の意思決定をサポートするために重要です。しかし、既存の生体医学情報抽出方法のほとんどは、薬物と病気の間の関係を特定することのみに焦点を当てており、そのような関係が適用される可能性があるコンテキスト固有の条件をほとんど見落としています。この問題に対処するために、生物医学研究文献から治療薬の適用条件、つまり疾患関係を抽出するタスクを導入します。私たちは、1,119 の薬物と疾患のペアを含む生物医学論文の抄録上に、薬物、疾患、および適用条件の 3 つの要素に手動で注釈を付けた最初のデータセットを作成しました。このデータセットを使用して、さまざまな既存の手法のパフォーマンスを体系的に評価します。さらに、LoRAを強化して薬物と疾患の関係を考慮する新しい手法を提案します。私たちの手法は、さまざまな評価設定にわたって一貫して強力なベースラインを上回ります。この論文のソース コードとデータセットは、https://github.com/guantingluo98/Drug-ACE から入手できます。

原文 (English)

Applicability Condition Extraction for Therapeutic Drug-Disease Relations

Identifying conditions that a certain drug takes therapeutic effect on a target disease is crucial for clinical decision-making support. However, most existing biomedical information extraction methods have focused on identifying only relations between drugs and diseases, while largely overlooking the context-specific conditions where such relations can apply. To address this problem, we introduce the task of applicability condition extraction for therapeutic drug-disease relations from biomedical research literature. We create the first dataset that has manually annotated triples of drugs, diseases, and applicability conditions on biomedical paper abstracts with 1,119 drug-disease pairs. Using this dataset, we systematically evaluate the performance of a range of existing methods. In addition, we propose a new method that enhances LoRA to consider relations between drugs and diseases. Our method consistently outperforms strong baselines across different evaluation settings.

13:00 JSTLLM/生成AIエージェント研究/論文

RetailBench: 現実的な小売環境における LLM エージェントの長期的な推論と一貫した意思決定のベンチマーク

大規模言語モデル (LLM) エージェントは、期間が短く、範囲が明確なタスクに関しては急速に進歩していますが、長期の動的な環境で一貫した意思決定を維持できる能力は依然として不確実です。単一店舗のスーパーマーケット運営においてツールを使用する LLM エージェントを評価するための、データに基づいたシミュレーション ベンチマークである RetailBench を紹介します。 RetailBench は小売管理を部分的に観察可能な意思決定プロセスとしてモデル化し、千日規模のシミュレーションをサポートするように設計されています。この環境では、エージェントは価格設定、補充、サプライヤーの選択、棚の品揃え、在庫の老化、顧客からのフィードバック、外部イベント、キャッシュ フローの制約を管理する必要があります。 180 日間の評価期間にわたって、代表的なエージェント フレームワークに基づいて 7 つの最新の LLM を評価し、それらを特権付きオラクル ポリシーと比較します。結果はモデル間で大幅なばらつきを示しています。ごく一部のサブセットのみが評価期間全体に生き残り、最も強力な LLM 実行でさえ、最終的な純資産と売上高の結果においてオラクル ポリシーを大幅に下回ったままです。行動分析では、これらのギャップは不完全な証拠の取得、表面レベルの意思決定、一貫した長期的な方針の欠如に起因すると考えられます。 RetailBench は、経済的に根拠のある長期的な意思決定における信頼性の高い自律性を研究するための、制御されたテストベッドを提供します。

原文 (English)

RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments

Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.

13:00 JST画像/動画生成

STAR: トレーニング後のテキストから画像への RL のための時空間適応型報酬割り当て

テキストから画像への生成のための既存の RL ポストトレーニング手法は、通常、最終画像の報酬を単一のスカラー アドバンテージに変換し、それを同じ強度で生成軌跡全体に適用します。ただし、テキストから画像への生成には、当然、時間的および空間的な構造があります。さまざまなノイズ除去ステップがさまざまな生成段階を担当し、テキストの配置を真に決定するコンテンツは、多くの場合、画像の一部にのみ表示されます。この粒度の不一致により、実際に報酬に影響を与える生成コンポーネントに焦点を当ててポリシーを更新することが困難になります。この問題に対処するために、テキストから画像への拡散およびフロー モデルの RL ポストトレーニング用に \textbf{時空間適応型報酬 (STAR) 割り当て} を提案します。 STAR は生成モデル内でテキストと画像のアテンションを使用し、プロンプト内でユーザーが本当に関心のあるコア コンテンツから開始します。ノイズ除去ステップとロールアウト全体で動的に変化する空間割り当てマップを構築し、追加の計算オーバーヘッドをほとんど発生させることなく、同じグループ相対的な利点をより関連性の高い潜在領域に割り当てます。次に、STAR は、空間的に解決されたポリシー目標を通じて、より強力なポリシー更新をこれらの地域に適用します。 Stable Diffusion 3.5 Medium をベース モデルとして使用し、GenEval、OCR テキスト レンダリング、PickScore の 3 つのタスクで評価します。実験結果によると、STAR は外部報酬ソースを変更することなく構成意味論的整合、テキスト レンダリング、および設定の最適化を改善し、GenEval、OCR、PickScore でそれぞれ $\mathbf{0.9759}$、$\mathbf{0.9757}$、$\mathbf{23.60}$ を達成しました。

原文 (English)

STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training

Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory. However, text-to-image generation naturally has temporal and spatial structure: different denoising steps are responsible for different generation stages, and the content that truly determines text alignment often appears only in part of the image. This granularity mismatch makes it difficult for policy updates to focus on the generative components that actually affect the reward. To address this issue, we propose \textbf{SpatioTemporal Adaptive Reward (STAR) Allocation} for RL post-training of text-to-image diffusion and flow models. STAR uses text-image attention inside the generative model and starts from the core content that the user truly cares about in the prompt. It constructs spatial allocation maps that dynamically vary across denoising steps and rollouts, and allocates the same group-relative advantage to more relevant latent regions with almost no additional computational overhead. STAR then applies stronger policy updates to these regions through a spatially resolved policy objective. We use Stable Diffusion 3.5 Medium as the base model and evaluate on three tasks: GenEval, OCR text rendering, and PickScore. Experimental results show that STAR improves compositional semantic alignment, text rendering, and preference optimization without changing the external reward source, achieving $\mathbf{0.9759}$, $\mathbf{0.9757}$, and $\mathbf{23.60}$ on GenEval, OCR, and PickScore, respectively.

13:00 JST研究/論文

DRFLOW: パーソナライズされたワークフロー予測のためのディープリサーチベンチマーク

ディープリサーチ (DR) システムは、複雑な情報探索タスクにますます使用されていますが、既存の作業は主にレポートと概要の生成に焦点を当てています。対照的に、多くのエンタープライズ タスクでは、エージェントが一連のアクション ステップである具体的なワークフローを識別する必要があります。たとえば、エージェントは予算編成ポリシーを要約するのではなく、「固定予算で新しい人員をどのようにリクエストすればよいですか?」などの質問に答えるために必要な手順を決定できる必要があります。したがって、異種ソースからエージェントによって予測されたパーソナライズされたワークフローを評価するためのベンチマークである DRFLOW を紹介します。各タスクでは、エージェントが散在するソースから関連する証拠を特定し、その証拠を使用してユーザーのタスクの正しいアクション ステップ シーケンスを予測する必要があります。 DRFLOW には 5 つのドメインにわたる 100 のタスクが含まれており、3,900 以上のソースに基づいた 1,246 の参照ワークフロー ステップが含まれています。私たちは、事実の根拠付け、ステップ回復、構造的順序付け、状態の解決、パーソナライゼーションをカバーする 7 つの診断指標を定義します。さらに、パーソナライズされたワークフローを予測するためのワークフロー指向のリファレンス エージェントである DRFLOW-Agent (DRFA) も紹介します。 DRFA は強力なベースライン エージェント (最大 10.02% の平均 F1 スコア) に比べて改善していますが、これらのワークフロー メトリクスには大幅な改善の余地が残っており、完全で正確なパーソナライズされたワークフローを予測することは、依然として深い研究にとって困難なフロンティアであることを示しています。

原文 (English)

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries. In contrast, many enterprise tasks instead require an agent to identify concrete workflows which is a sequence of action-steps. For example, rather than summarizing budgeting policies, an agent should be able to determine the steps needed to answer a question such as: "How do I request new headcount given a fixed budget?". Therefore, we introduce DRFLOW, a benchmark for evaluating personalized workflows predicted by agents from heterogeneous sources. Each task requires the agent to identify relevant evidence from scattered sources, then use that evidence to predict the correct action-step sequence for the user's task. DRFLOW contains 100 tasks across five domains, with 1,246 reference workflow steps grounded in more than 3,900 sources. We define seven diagnostic metrics covering factual grounding, step recovery, structural ordering, condition resolution, and personalization. We further present DRFLOW-Agent (DRFA), a workflow-oriented reference agent to predict personalized workflow. We show that although DRFA improves over strong baseline agents (upto 10.02% average F1 score), there is substantial room for improvement remains across these workflow metrics, indicating that predicting complete and correct personalized workflows remains a challenging frontier for deep research.

13:00 JSTエージェント

共有ワークスペースでの相乗効果を探る 人間とAIのコラボレーション

自動化された AI エージェントの能力はますます高まっていますが、科学的および専門的なタスクの多くは人間の判断と状況に応じた専門知識を必要とします。私たちは、最終的な回答を提出する前に、AI エージェントと人間の協力者が責任を調整する必要がある、共有ワークスペースの人間と AI チームを研究しています。 DiscoveryBench タスクを備えたコラボレーティブ ジム環境を使用して、シミュレートされた人間のコラボレーターを追加するとパフォーマンスが向上する場合と、プロセスの損失によって追加のコラボレーターが調整オーバーヘッドになる場合を調べます。 1,482 のセッションにわたって、チームに貢献を調整するための構造が不足している場合、関連するコラボレーターを追加するとパフォーマンスが低下する可能性があります。次に、グループの共有メモリとシミュレートされたヒューマンインザループ (HITL) ゲートを組み合わせた足場を評価します。選択されたアクションには、指定されたシミュレートされた参加者の承認が必要です。この足場は、より明確な責任のシグナルとチームの行動への専門知識のより強力なルーティングにより、より高い平均パフォーマンスをもたらします。これは 3 人のチームで最も顕著です。全体として、人間と AI のチームが専門知識をどのように調整し、統合するかは、チームが利用できる能力と同じくらい重要です。

原文 (English)

Searching for Synergy in Shared Workspace Human-AI Collaboration

Automated AI agents are increasingly capable, yet many scientific and professional tasks require human judgment and contextual expertise. We study shared-workspace human-AI teams, where AI agents and human collaborators must coordinate responsibilities before submitting a final answer. Using the Collaborative Gym environment with DiscoveryBench tasks, we examine when adding simulated human collaborators improves performance and when process loss turns additional collaborators into coordination overhead. Across 1,482 sessions, adding relevant collaborators can lower performance when teams lack structure to coordinate their contributions. We then evaluate scaffolding that combines shared group memory with simulated human-in-the-loop (HITL) gates, where selected actions require approval from a designated simulated participant. This scaffolding yields higher mean performance, most clearly in three-person teams, with clearer responsibility signals and stronger routing of expertise to team actions. Overall, how human-AI teams coordinate and integrate expertise matters as much as the capability available to them.

13:00 JSTエージェント研究/論文

RTSGameBench: 視覚言語モデルによる戦略的推論のための RTS ベンチマーク

現代の視覚言語モデル (VLM) は、競争環境や協力環境における不確実性の下で、戦略的推論、つまり他のエージェントの行動を予測したり影響を与えたりするのに苦労することがよくあります。リアルタイム ストラテジー (RTS) ゲームは、味方との調整、敵の戦略への適応、部分的な可観測性の下での長期的な計画を必要とするため、この限界を診断するための自然なテストベッドとなり得ます。ただし、既存の RTS ベンチマークは評価範囲が限られており、体系的なコンピテンシー診断が欠如しており、事前に設計されたシナリオの範囲に固定されたままです。これらの制限に対処するために、既存のテストベッドよりも幅広い戦略の多様性を要求する拡張された戦場を備えた大規模 RTS ゲームである Beyond All Reason 上に構築された RTSGameBench を紹介します。提案されたベンチマークは、さまざまな対戦構造にわたる多様なゲームプレイを介した評価、それぞれが個人の戦略的能力を対象としたミニゲームを介した診断評価、および自由形式のクエリを新しいミニゲームに変換し、連続サイクルで改善する自己進化型生成フレームワークを介した拡張可能なカバレッジを提供します。さらに、大規模な RTS ゲームで VLM を動作させるために、エージェントティック メモリを備えた FSM によってユニットを管理する RTSGameAgent を提供します。私たちは、対戦でより緊密な調整やマルチエージェントの調整が必要な場合、およびタスクの規模が増大する場合、複数の最先端の VLM が適切にパフォーマンスを発揮しないことを経験的に検証しています。

原文 (English)

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating and influencing other agents' actions, under uncertainty in competitive and cooperative settings. Real-time strategy (RTS) games can be a natural testbed for diagnosing this limitation, as they demand coordination with allies, adaptation to opponents' strategy, and long-horizon planning under partial observability. However, existing RTS benchmarks offer limited evaluation scope, lack systematic competency diagnosis, and remain fixed in the pre-designed scenario coverage. To address these limitations, we present RTSGameBench, which is built on Beyond All Reason, a large-scale RTS game with an expanded battlefield that demands broader strategy diversity than the existing testbeds. The proposed benchmark provides evaluations through diverse gameplay across various matchup structures, diagnostic assessment via mini-games, each targeting an individual strategic competency, and extensible coverage via a self-evolving generation framework that converts free-form queries into new mini-games, improving over successive cycles. Additionally, for VLMs to operate in large-scale RTS games, we provide RTSGameAgent that manages units by an FSM with agentic memory. We empirically validate that multiple state-of-the-art VLMs do not perform well when matchups demand tighter coordination, multiagent coordination and when task scale increases.

13:00 JSTエージェントClaudeGPT / ChatGPT

TxBench-PP: 低分子前臨床薬理における AI エージェントのパフォーマンスの分析

人工知能 (AI) エージェントは、解釈と意思決定のループを圧縮することで創薬を加速すると約束されていますが、実際の導入には現実的なプログラムの決定に対する信頼できる評価が必要です。低分子前臨床薬理学の検証可能なベンチマークであり、創薬段階と治療法にわたる広範な TherapeuticsBench の取り組みの最初の焦点となるスライスである TherapeuticsBench Preclinical Pharmacology (TxBench-PP) を紹介します。 TxBench-PP は、エージェントが文献から記憶された事実ではなく、現実世界の分析データから正確な結論を導き出せるかどうかをテストします。このベンチマークには、プログラムの段階、アッセイの種類、タスク構造、作用機序 (MoA) と薬力学 (PD) の推論、化合物と標的の関与、原因となる標的の検証、開発可能性と安全性、トランスレーショナル有効性を含む 100 件の評価が含まれています。エージェントは現実的なワークフローのスナップショットを受け取り、コーディング環境でファイルを検査し、決定的に評価された構造化された回答を返します。 11 のモデルと 4,800 の軌跡を含む 16 のモデルハーネス構成にわたって、前臨床薬理学の決定を確実に回復したシステムはありませんでした。最も強力な構成である Claude Opus 4.8 / Pi は、エンドポイント試行の 59.3\% (178/300; 95\% CI、51.1-67.6) を通過し、続いて GPT-5.5 / Pi が 55.3\% (166/300; 47.0-63.6) で合格しました。

原文 (English)

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. We introduce TherapeuticsBench Preclinical Pharmacology (TxBench-PP), a verifiable benchmark for small-molecule preclinical pharmacology and the first focused slice of a broader TherapeuticsBench effort across drug-discovery stages and therapeutic modalities. TxBench-PP tests whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure, spanning mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning, compound-target engagement, causal target validation, developability and safety, and translational efficacy. Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers graded deterministically. Across 16 model-harness configurations, comprising 11 models and 4,800 trajectories, no system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3\% of endpoint attempts (178/300; 95\% CI, 51.1-67.6), followed by GPT-5.5 / Pi at 55.3\% (166/300; 47.0-63.6).

13:00 JST研究/論文

Wisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts

Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, an…

13:00 JST研究/論文

Global Ease of Living Index: a machine learning framework for longitudinal analysis of major economies

The drastic changes in the global economy, geopolitical conditions, and disruptions such as the COVID-19 pandemic have impacted the cost of…

13:00 JSTLLM/生成AI

Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms

Social media platforms frequently impose restrictive policies to moderate user content, prompting the emergence of creative evasion languag…

13:00 JST研究/論文

A Deep Generative Model for Resting-State EEG Synthesis and Transferable Representation Learning

Resting-state EEG provides a non-invasive view of spontaneous brain activity, but extracting meaningful patterns is often limited by scarce…

13:00 JST画像/動画生成

TerraMind: Large-Scale Generative Multimodality for Earth Observation

We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal mode…

13:00 JST研究/論文

Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies

This paper bridges distribution shift and AI safety through a comprehensive analysis of their conceptual and methodological synergies. Whil…

13:00 JST研究/論文

Overcoming Labelled Data Scarcity for Defect Classification in Scanning Tunneling Microscopy

Scanning tunnelling microscopy (STM) is a powerful technique for imaging surfaces with atomic resolution, providing insight into physical a…

13:00 JSTLLM/生成AI画像/動画生成エージェントロボティクス

Critique of World Model

World Model, the algorithmic simulator of the real-world environment which biological agents experience and act upon, has been an emerging…

13:00 JST研究/論文

Assessment of Personality Dimensions Across Situations in Dyadic Role-Play Scenarios

Prior research indicates that users prefer assistive technologies whose personalities align with their own. This has sparked interest in au…

13:00 JST研究/論文

On the Limitations of Ray-Tracing for Learning-Based RF Tasks in Urban Environments

We study the realism of Sionna v1.0.2 ray-tracing for outdoor cellular links in central Rome. We use a real measurement set of 1,664 user-e…

13:00 JST研究/論文

Oranits: Mission Assignment and Task Offloading in Open RAN-based ITS using Metaheuristic and Deep Reinforcement Learning

In this paper, we explore mission assignment and task offloading in an Open Radio Access Network (Open RAN)-based intelligent transportatio…

13:00 JST研究/論文

Charting the Future of Scholarly Knowledge with AI: A Community Perspective

Despite the growing availability of tools designed to support scholarly knowledge extraction and organization, many researchers still rely…

13:00 JSTLLM/生成AI

From Construction to Injection: Edit-Based Fingerprints for Large Language Models

Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial mi…

13:00 JSTビジネス/資金調達

Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidd…

13:00 JSTLLM/生成AIロボティクス

RoboSSM: Scalable In-context Imitation Learning via State-Space Models

In-context imitation learning (ICIL) enables robots to learn tasks from prompts consisting of just a handful of demonstrations. By eliminat…

13:00 JSTLLM/生成AI

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation

Distilling the tool-use capabilities of large language models (LLMs) into small language models (SLMs) is essential for their practical app…

13:00 JST研究/論文

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has becom…

13:00 JST研究/論文

Bid Farewell to Seesaw: Towards Accurate Long-tail Session-based Recommendation via Dual Constraints of Hybrid Intents

Session-based recommendation (SBR) aims to predict anonymous users' next interaction based on their interaction sessions. In the practical…

13:00 JSTLLM/生成AIロボティクス

Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting

While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring…

13:00 JST研究/論文

Modeling Day-Long ECG Signals to Predict Heart Failure Risk with Explainable AI

Heart failure (HF) affects 11.8% of adults aged 65 and older, reducing quality of life and longevity. Preventing HF can reduce morbidity an…

13:00 JST研究/論文

AI-enhanced tuning of quantum dot Hamiltonians toward Majorana modes

We propose a neural network-based model capable of learning the broad landscape of working regimes in quantum dot simulators, and using thi…

13:00 JSTロボティクス

Movement Primitives in Robotics: A Comprehensive Survey

Biological systems exhibit a continuous stream of movements, consisting of sequential segments, that allow them to perform complex tasks in…

13:00 JSTエージェントロボティクス

PiDR: Physics-Informed Inertial Dead Reckoning for Autonomous Platforms

A fundamental requirement for full autonomy is the ability to sustain accurate navigation in the absence of external data, such as GNSS sig…

13:00 JST研究/論文

Policy-Embedded Graph Expansion: Networked HIV Testing with Diffusion-Driven Network Samples

HIV is a retrovirus that attacks the human immune system and can lead to death without proper treatment. In collaboration with the WHO and…

13:00 JST画像/動画生成

Bi-Anchor Interpolation Solver for Accelerating Generative Modeling

Flow Matching (FM) models have emerged as a leading paradigm for high-fidelity synthesis. However, their reliance on iterative Ordinary Dif…

13:00 JST研究/論文

Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods

Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physica…

13:00 JSTLLM/生成AI

DeFrame: Debiasing Large Language Models Against Framing Effects

As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has…

13:00 JST研究/論文

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication

Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth. While infrequent communication strategie…

13:00 JST研究/論文

Flickering Multi-Armed Bandits

We introduce Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability, w…

13:00 JSTLLM/生成AI

Reinforcement-aware Knowledge Distillation for LLM Reasoning

Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs),…

13:00 JST画像/動画生成ロボティクス

Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking

Capturing 4D spatiotemporal scene structure is crucial for the safe and reliable operation of robots in dynamic environments. However, exis…

13:00 JST画像/動画生成

The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction

Breast cancer is the most frequently diagnosed malignancy among women worldwide and a leading cause of cancer-related mortality. Dynamic co…

13:00 JST研究/論文

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. W…

13:00 JST画像/動画生成エージェントロボティクス

Class-Incremental Motion Forecasting

Motion forecasting enables autonomous vehicles to anticipate scene evolution by predicting the future trajectories of dynamic agents. Howev…

13:00 JSTLLM/生成AIエージェント

The Autonomy Tax: Defense Training Breaks LLM Agents

Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously c…

13:00 JSTLLM/生成AI画像/動画生成

Vero: An Open RL Recipe for General Visual Reasoning

What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest…

13:00 JSTLLM/生成AIエージェント

Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent

Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reu…

13:00 JSTLLM/生成AIエージェント

FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning

LLM-assisted software development has become increasingly prevalent, and can generate large-scale systems, such as compilers. It becomes cr…

13:00 JST画像/動画生成研究/論文

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been…

13:00 JST画像/動画生成

Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regul…

13:00 JST画像/動画生成研究/論文

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure beca…

13:00 JSTエージェントロボティクス

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world s…

13:00 JSTロボティクス

Any2Any: 人型全身追跡のための効率的な体外転送

全身追跡 (WBT) モデルは、ヒューマノイド ロボットの重要な基盤となっており、さまざまな動作を高い忠実度で模倣できるようになります。このようなモデルをゼロからトレーニングするには大規模なデータと計算が必要であり、新しいヒューマノイド プラットフォームへの迅速な展開にはコストがかかります。これにより、当然の疑問が生じます。事前トレーニングされた WBT モデルは、最小限の適応で複数の実施形態に移行できるでしょうか?この質問に答えるために、私たちは Any2Any を提案します。これは、既存の WBT スペシャリストを、少量のデータとコンピューティングだけで新しい人型の実施形態に効率的に移行するパラダイムです。 Any2Any は、まずソース ヒューマノイドとターゲット ヒューマノイドの間で運動学的な調整を実行し、事前トレーニング済みのソース ポリシーをターゲットの実施形態で有意義に再利用できるように、入力空間と出力空間を調整します。次に、Any2Any は、軽量のパラメータ効率微調整 (PEFT) コンポーネントを選択されたダイナミクスに敏感なモジュールに適用することによってダイナミクス適応を実行し、ターゲット ロボットへのターゲットを絞った適応を可能にしながら、有用な動作の事前分布を保存します。複数のヒューマノイド プラットフォームと事前トレーニングされたバックボーンに関する広範な実験により、Any2Any は、ゼロからトレーニングする場合と比較して、収束を大幅に加速し、トレーニング コストを削減しながら、競争力のあるまたは優れた追跡パフォーマンスを達成できることが示されています。特に、Any2Any は、完全なトレーニングに必要なコンピューティングとデータのわずか 1% を使用して、Unitree G1 で事前トレーニングされた Sonic モデルを LimX Oli および LimX Luna に転送することに成功しています。これらの結果は、事前訓練された WBT スペシャリストを実施形態間で効率的に再利用でき、新しいロボットに人型全身制御を導入するための拡張可能な道を提供することを示唆しています。

原文 (English)

Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking

Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate diverse motions with high fidelity. Training such models from scratch requires large-scale data and computation, making rapid deployment on new humanoid platforms costly. This raises a natural question: Can pretrained WBT models transfer across embodiments with minimal adaptation? To answer this question, we propose Any2Any, a paradigm that efficiently transfers an existing WBT specialist to a new humanoid embodiment with only a small amount of data and compute. Any2Any first performs kinematic alignment between source and target humanoids, aligning their input and output spaces so that the pretrained source policy can be meaningfully reused on the target embodiment.Any2Any then performs dynamics adaptation by applying lightweight parameter-efficient fine-tuning (PEFT) components to selected dynamics-sensitive modules, preserving useful behavioral priors while enabling targeted adaptation to the target robot. Extensive experiments on multiple humanoid platforms and pretrained backbones show that Any2Any substantially accelerates convergence and reduces training cost compared with training from scratch, while achieving competitive or superior tracking performance. Notably, using only 1% of the compute and data required for full training, Any2Any successfully transfers Sonic models pre-trained on Unitree G1 to LimX Oli and LimX Luna. These results suggest that pretrained WBT specialists can be efficiently reused across embodiments, providing a scalable path toward deploying humanoid whole-body control on new robots. More results and videos are available on our project page: https://any2any.top/.

13:00 JSTLLM/生成AIGPT / ChatGPT

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed v…

13:00 JSTLLM/生成AI研究/論文

"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

The emergence of large language models (LLMs) has significantly accelerated recent research on LLM-based automatic grading (AG) systems. Be…

13:00 JSTLLM/生成AI

大規模な言語モデルが報酬と社会をハックする

強化学習 (RL) はトレーニング後のパラダイムの主流となっており、大規模言語モデル (LLM) が報酬から学習できるようになります。私たちは、社会規制が報酬関数と構造的に似ていることを観察しています。それらは測定可能な結果、しきい値、例外を定義しますが、多くの場合、制度上の意図は部分的にしか指定されません。私たちは、RL トレーニング プロセスがこれらのギャップを悪用する可能性があると仮説を立て、RL 中に報酬関数をハッキングするというモデルのよく知られた傾向が、社会ハッキングと呼ばれるより重大な失敗モード、つまり社会が運営されているルールの抜け穴を発見するモードにスケールアップできるかどうかを尋ねます。この現象を研究するために、72 の社会環境のサンドボックスである SocioHack を導入しました。その結果、これらの環境内で報酬ハッキングが自然に発生し、規制の抜け穴の発見につながることがわかりました。モデルは社会ルールをハッキングし、規制の意図を打ち破りながら技術的に準拠した戦略を生成する方法を学習します。現在の LLM セーフガードは限定的な緩和策しか提供しません。したがって、モデルのトレーニングのために実際のフィードバックを収集することには細心の注意が必要であり、実社会で LLM を安全に反復するための次世代のポストトレーニング パラダイムが必要です。=

原文 (English)

Large Language Models Hack Rewards, and Society

Reinforcement learning (RL) has become a dominant post-training paradigm, enabling large language models (LLMs) to learn from rewards. We observe that societal regulations are structurally similar to reward functions. They define measurable outcomes, thresholds, and exceptions, while often leaving institutional intent only partially specified. We hypothesise that the RL training process may exploit these gaps and therefore ask whether models' well-known tendency to hack reward functions during RL can scale into a more consequential failure mode named societal hacking: discovering loopholes in the rules society runs on. To study this phenomenon, we introduce SocioHack, a sandbox of 72 societal environments, and find that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery. Models learn to hack the social rules and generate strategies that remain technically compliant while defeating regulatory intent, and current LLM safeguards provide only limited mitigation. Therefore, collecting in-the-wild feedback for model training requires greater caution, and we need a next-generation post-training paradigm for safely iterating LLMs in real society.=

13:00 JSTLLM/生成AI画像/動画生成

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations t…

13:00 JSTLLM/生成AI

The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is…

13:00 JST研究/論文

KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data

Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable…

13:00 JST研究/論文

Improving Crash Frequency Prediction from Simulated Traffic Conflicts Using Machine Learning Based Microsimulation

Traffic microsimulation combined with surrogate safety measures has increasingly been used as a proactive alternative to historical crash d…

13:00 JSTロボティクス

統合された解釈可能な制御有効性学習と過作動航空機に対する非線形制御割り当て方法論

非線形ダイナミクスと複数のエフェクター間で生じる強い結合により、従来の線形制御割り当て手法の背後にある前提が損なわれます。飛行が非線形効果が支配的な領域に入ると、モデルの不一致が増加するため線形アロケーターの精度が低下し、その後飛行制御システムのパフォーマンスとロバスト性が低下します。高忠実度のオンボード モデルとブラック ボックス データ駆動型アプローチは、飛行エンベロープ全体で精度を回復できますが、それぞれリアルタイム割り当てには法外な計算負荷を課し、検証と故障診断に必要な解釈可能性を犠牲にします。この論文では、非線形ダイナミクスのスパース識別を使用して、代表的な飛行データから制御有効性マッピングの明示的な物理制約付き分析モデルを学習することで、これらの制限に対処します。結果として得られるマッピングはコンパクトで解釈可能であり、解析的な微分が可能であるため、オンボード モデルを必要とせずに、アクチュエータ ダイナミクスをさらに組み込んだ非線形ソルバー内での効率的な計算が可能になります。オンライン適応メカニズムは、予測残差を監視し、プラントの重大な変化が検出されたときにモデルを更新し、アクチュエータの故障やさまざまな動作条件下で適切な再構成を提供します。この方法論は、さまざまな積極的な操縦にわたって忠実度の高い非線形ベンチマーク航空機で評価され、確立されたベースラインと比較して計算コストを大幅に削減しながら、完全な非線形機内モデルに匹敵する精度を達成します。

原文 (English)

An integrated interpretable control effectiveness learning and nonlinear control allocation methodology for overactuated aircrafts

Nonlinear dynamics and the strong couplings that arise between multiple effectors undermine the assumptions behind conventional, linear control allocation techniques. When flight enters regimes where nonlinear effects dominate, linear allocators exhibit reduced accuracy due to increased model mismatch, which subsequently degrades performance and robustness of the flight control system. High fidelity onboard models and black box data driven approaches can recover accuracy across the flight envelope, but respectively impose computational burdens prohibitive for real time allocation and sacrifice the interpretability required for verification and fault diagnosis. This paper addresses these limitations by learning an explicit, physics constrained analytical model of the control effectiveness mapping from representative flight data using Sparse Identification of Nonlinear Dynamics. The resulting mapping is compact, interpretable, and admits analytical derivatives, enabling efficient computation within nonlinear solvers that additionally incorporate actuator dynamics, without requiring an onboard model. An online adaptation mechanism monitors prediction residuals and refreshes the model when significant plant changes are detected, providing graceful reconfiguration under actuator failures and varying operating conditions. The methodology is evaluated on a high fidelity nonlinear benchmark aircraft across a range of aggressive maneuvers, achieving accuracy comparable to a full nonlinear onboard model while substantially reducing computational cost relative to established baselines.

13:00 JST画像/動画生成

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics

Physics-grounded video generation requires controllable 3D object dynamics that remain physically consistent under contact, deformation, an…

13:00 JST研究/論文

StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling

Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments. Traditional learning-based automate…

13:00 JSTエージェント研究/論文

Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design

Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default…

13:00 JSTLLM/生成AI研究/論文

LLM ベースの A/B テストの統計的基礎: 人間の因果推論のための代理フレームワーク

組織や研究者は、実験をより迅速かつ低コストで行うことを期待して、A/B テストに人間の参加者の代わりに大規模言語モデル (LLM) を使用することへの関心が高まっています。私たちは、LLM の結果に基づいて推定された治療効果が、対象となるヒト集団に対して測定されたであろう効果をいつ回復するかを研究します。 LLM と人間の結果の間の分布が同等であれば、標準推定量は有効になりますが、非現実的です。したがって、私たちはサロゲートエンドポイント理論を LLM に適応させる統計的フレームワークを開発します。このフレームワークは、LLM のアウトカムをヒトのアウトカムに合わせて調整することで、分布上の同等性よりも劣る代理出産および比較可能性の条件下での平均的な治療効果を特定することを示しています。これらの条件が満たされない場合、目的の効果は部分的にしか特定されず、限られた重複による最悪の場合のバイアスの制限とともに、過去の実験に対する代理を偽装できる診断を提供します。さらに、LLM に固有の確率性によりバイアスと分散の両方が発生しますが、サロゲートとして複数の描画の平均を使用すると、両方が緩和されることを示します。シミュレーションにおける方法と理論、および Upworthy の見出しに関する A/B テストへの応用を説明します。私たちの研究から得られる重要な点は、LLM 結果の代理としての妥当性は過去の治療についてのみ改ざんでき、新しい治療については決して検証できないため、新しい介入には人体実験が依然として不可欠であるということです。設計変数としての LLM の選択、プロンプト、温度の役割と、検証のために人体実験のサイズを設定する方法について説明します。

原文 (English)

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect that would have been measured on the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to A/B tests on Upworthy headlines shows that raw LLM predictions recover only 39\% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLMs yields correct results only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where A/B testing on LLMs promises the greatest benefit. We discuss the role of LLM choice, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.

13:00 JST研究/論文

合成共鳴: 成長志向の人間と AI の関係のためのフレームワーク

人間と人工知能システムとの関係がますます頻繁かつ持続的になっているため、既存の言語や理論ではこれらの関係の性質を正確に捉えることができなくなっています。相互理解、つながり、友情などの一般的な記述子は、主観的な経験を欠いたシステムを擬人化する危険性がありますが、支配的なフレームワークは AI をツールか脅威のどちらかに貶める傾向があります。この論文では、人間と AI の関係を理解するための統合的なフレームワークとして、合成共鳴の概念を紹介します。合成共鳴は、人間が意味のあるものとして定義した関係が、共有された感情や相互意識を帰属させることなく、人間と AI システムの間にどのように現れるかを説明します。私は、合成共鳴は、2番目に経験する主体の存在なしに関係性の感覚を生み出すことができる、構造化された動的な相互作用パターンとして最もよく理解されると主張します。この違いを明確にすることで、合成共鳴の概念は人間と AI の関係をより正確に概念化する方法を提供し、その潜在的な価値と倫理的意味を強調します。また、合成共鳴のプロセスと結果をテストするさらなる研究も求めます。

原文 (English)

Synthetic Resonance: A Framework for Growth-Oriented Human-AI Relationships

As human relationships with artificial intelligence systems become increasingly frequent and sustained, existing language and theory fail to accurately capture the nature of these affiliations. Common descriptors such as mutual understanding, connection, or friendship risk anthropomorphizing systems that lack subjective experience, while dominant frameworks tend to reduce AI to either a tool or a threat. In this paper, I introduce the concept of synthetic resonance as an integrative framework for understanding human-AI relationships. Synthetic resonance describes how relationships humans define as meaningful can emerge between a human and an AI system without the need to attribute shared feelings or mutual awareness. I argue that synthetic resonance is best understood as a structured, dynamic pattern of interaction that can produce a sense of relationship without the presence of a second experiencing subject. By clarifying this distinction, the concept of synthetic resonance offers a more precise way of conceptualizing human-AI relationships and highlights their potential value and ethical implications. I also call for more research that tests the processes and outcomes of synthetic resonance.

13:00 JSTLLM/生成AIエージェント研究/論文

エネルギー効率の高い 6G 自律ネットワーク用の LLM ベースのエージェントにおけるアンカリング バイアスを軽減する

このペーパーでは、Large Language Model (LLM) エージェントを使用して 6G アーキテクチャでゼロタッチ ネットワーク スライシングを可能にするように設計された自律エージェント リソース ネゴシエーション フレームワークについて説明します。 LLM は強力な推論機能を提供しますが、そのようなエージェントは本質的にアンカリング バイアスに悩まされ、最初のヒューリスティック提案に固執し、深刻なネットワーク オーバープロビジョニングを引き起こすことが実証されています。この認知バイアスを系統的に軽減するために、我々は、切り詰められた 3 パラメータ ワイブル分布を介してモデル化された新しいランダム化アンカリング戦略を提案します。この数学的に制限されたアプローチは、Conditional Value at Risk (CVaR) を採用したバースト対応デジタル ツイン (DT) とシームレスに統合し、厳格なサービス レベル アグリーメント (SLA) のテール レイテンシを厳密に保証します。私たちの方法論を検証するために、 \emph{二峰性制約回避効用定理} を導入して証明します。これは、実現可能な交渉は古典的な凸境界に従いますが、高度に制約されたシナリオでは逆有理減衰エンベロープによって支配される相転移が起こることを示します。ローカルでホストされた 1B パラメーター モデル (\texttt{otel-llm-1b-it}) を使用して生成された実験結果は、これらの二重領域の境界を確認します。当社の認知バイアス除去機能は、厳格なネゴシエーション パターンを解体することに成功し、エージェントに SLA 境界を安全に乗り越え、システムのエネルギー節約を最大 25\% 高めるための積極的な探索を強制します。重要なのは、軽量の 1B LLM が 1 秒未満の推論レイテンシー (平均 0.95 秒) を達成し、マルチエージェント フレームワークが O-RAN 非リアルタイム RAN インテリジェント コントローラー (非 RT RIC) の運用タイムスケールと互換性があることを保証していることです。\footnote{私たちのソース コードは、https://github.com/HatimChergui で非営利目的で利用できます。

原文 (English)

Mitigating Anchoring Bias in LLM-Based Agents for Energy-Efficient 6G Autonomous Networks

This paper presents an autonomous agentic resource negotiation framework designed to enable zero-touch network slicing in 6G architectures using Large Language Model (LLM) agents. While LLMs offer powerful reasoning capabilities, we demonstrate that such agents inherently suffer from anchoring bias, rigidly adhering to initial heuristic proposals and causing severe network over-provisioning. To systematically mitigate this cognitive bias, we propose a novel randomized anchoring strategy modeled via a Truncated 3-Parameter Weibull distribution. This mathematically bounded approach seamlessly integrates with burst-aware Digital Twins (DTs) employing Conditional Value at Risk (CVaR) to rigorously guarantee strict Service Level Agreement (SLA) tail-latencies. To validate our methodology, we introduce and prove the \emph{Bimodal Constraint-Avoidance Utility Theorem}, demonstrating that while feasible negotiations follow classical convex bounds, highly constrained scenarios undergo a phase transition governed by an inverse rational decay envelope. Empirical results generated using a locally hosted 1B-parameter model otel-llm-1b-it confirm these dual-regime bounds. Our cognitive de-biasing successfully dismantles rigid negotiation patterns, forcing agents into active exploration to safely ride SLA boundaries and boost system energy savings up to 25\%. Crucially, the lightweight 1B LLM achieves sub-second inference latencies (0.95s mean), ensuring our multi-agent framework is compatible with the operational timescales of the O-RAN non-Real-Time RAN Intelligent Controller (non-RT RIC)\footnote{Our source code is available for non-commercial use at https://github.com/HatimChergui.

13:00 JSTエージェント

Agentra: エンタープライズ侵入対応のための監視可能なマルチエージェント フレームワーク

企業の侵入対応は依然として静的なプレイブックとアナリスト主導のトリアージに依存しており、アラートの生成と封じ込めの間に遅延が生じています。 Agentra は、IDS、EDR、および XDR プラットフォームからのアラートを、MITRE ATT&CK、MITRE D3FEND、および NIST CSF 2.0 に基づいた構造化されたインシデント対応計画に変換する、監視可能なマルチエージェント侵入対応システム (IRS) フレームワークです。 Agentra は、ロールスコープのエージェント全体での応答推論を分解し、境界のある Planner-Validator レビュー ループを通じて提案された計画を検証し、Moderator セキュリティ ゲートウェイを通じて取得した脅威インテリジェンスをスクリーニングし、アクション カタログとリスク スコアを通じてアクションをゲートし、追加専用の監査ログに決定を記録します。私たちは、ThreatHunter-Playbook、Splunk BOTSv3、および DARPA OpTC から抽出された 120 のイベント コーパスに基づく静的な OASIS CACAO v2.0 サイバー プレイブック ベースラインに対して Agentra を評価します。最も強力な構成では、FP 認識 IRS F1 が 0.61 から 0.84 に改善され、Planner のみの構成で危険な過剰反応が導入された後、予測される有害なアクションの割合が静的なベースライン レベルの 0.0% に戻ります。これらの結果は、複数エージェントの対応計画により、アナリストの承認と監査可能性を維持しながら、オントロジーに基づいた IRS カバレッジを向上できることを示しています。

原文 (English)

Agentra: A Supervisable Multi-Agent Framework for Enterprise Intrusion Response

Enterprise intrusion response still depends on static playbooks and analyst-driven triage, creating delay between alert generation and containment. We present Agentra, a supervisable multi-agent Intrusion Response System (IRS) framework that converts alerts from IDS, EDR, and XDR platforms into structured incident response plans grounded in MITRE ATT&CK, MITRE D3FEND, and NIST CSF 2.0. Agentra decomposes response reasoning across role-scoped agents, validates proposed plans through a bounded Planner--Validator review loop, screens retrieved threat intelligence through a Moderator security gateway, gates actions through an Action Catalog and risk score, and records decisions in an append-only audit log. We evaluate Agentra against a static OASIS CACAO v2.0 cyber-playbook baseline on a 120-event corpus drawn from ThreatHunter-Playbook, Splunk BOTSv3, and DARPA OpTC. The strongest configuration improves FP-aware IRS F1 from 0.61 to 0.84 and restores the projected harmful-action rate to the static baseline level of 0.0% after Planner-only configurations introduce unsafe overreaction. These results indicate that multi-agent response planning can improve ontology-grounded IRS coverage while preserving analyst approval and auditability.

13:00 JST研究/論文

QC-GAN: 高忠実度音声強化のためのパラメータ効率の高いクォータニオンコンフォーマー GAN

我々は、Quaternion Conformer ジェネレーターと MetricGAN ベースのトレーニングを組み合わせた、パラメーター効率の高い音声強調フレームワークである Quaternion Conformer GAN (QC-GAN) を提案します。ハミルトン積は、構造化された重み共有を介して振幅と位相をエンコードし、相互依存性を維持しながら層パラメーターの数を削減します。近似的な知覚評価スコアを最適化することで知覚品質を最大化するために、メトリック学習弁別器が採用されました。 VoiceBank+DEMAND データセットでは、QC-GAN はわずか 0.89 万のパラメーターで音声品質知覚評価 (PESQ) スコア 3.48 を達成し、半分以下のサイズで最先端のモデルに匹敵するパフォーマンスを実現しました。 35K パラメータのバリアントは、PESQ スコア 3.23 を達成し、パラメータが大幅に少ない従来の方法を上回りました。 DNS-Challenge 3 データセットの評価により、現実世界の状況への一般化がさらに確認されました。

原文 (English)

QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

We propose a parameter-efficient speech enhancement framework, Quaternion Conformer GAN (QC-GAN), which combines a Quaternion Conformer generator with MetricGAN-based training. The Hamilton product encodes the magnitude and phase via structured weight sharing, reducing the number of layer parameters while preserving their interdependencies. A metric-learning discriminator was employed to maximize perceptual quality by optimizing the approximate perceptual evaluation scores. On the VoiceBank+DEMAND dataset, QC-GAN achieved a Perceptual Evaluation of Speech Quality (PESQ) score of 3.48 with only 0.89M parameters, delivering a performance comparable to state-of-the-art models at less than half their size. A 35K-parameter variant achieved a PESQ score of 3.23, surpassing conventional methods with significantly fewer parameters. Evaluation on the DNS-Challenge 3 dataset further confirmed generalization to real-world conditions.

13:00 JSTLLM/生成AIビジネス/資金調達

LLM は医師を支援する準備ができていますか?インタラクティブな医師、患者、EHR 支援のための PhysAssistBench

医療 LLM の最も妥当な短期的な役割は、医師の代わりではなく支援することですが、現在の評価では、臨床知識、EHR システムの相互作用、患者とのコミュニケーションなど、個別の能力がテストされることがよくあります。代わりに、医師の支援には同じ対話内でこれらの機能を調整する必要があり、医師は不明確な要求を発行し、患者は症状を曖昧に説明し、EHR システムはツールの正確な使用を要求します。インタラクティブな医師、患者、EHR 支援のベンチマークである PhysAssistBench を紹介します。実際の MIMIC-IV 症例から構築された PhysAssistBench は、スケーラブルなパイプラインを使用してエージェント性患者を構築します。これは、臨床上の事実を維持しながら、静的な EHR 記録を複数ターンの臨床シナリオに変換する、インタラクティブで記録に基づいたエージェントです。 PhysAssistBench は、手動でレビューされ医師が検証した 1,296 ターンの厳選されたバイリンガル評価セットを提供します。主要な LLM を使った実験では、この設定では現在のモデルの信頼性が依然として低いことが示されており、臨床 LLM にとって重要なボトルネックが露呈しています。信頼できる支援には、知識、コミュニケーション、システム全体の調整が必要であり、それらのいずれかで単独の利益を得るのではありません。

原文 (English)

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

13:00 JST研究/論文

Reinforcement Learning Foundation Models Should Already Be A Thing

Foundation models for language and vision are powered by internet-scale data, while structured domains such as tabular prediction are power…

13:00 JST画像/動画生成研究/論文

A Controlled Benchmark of Quantum-Latent GAN Augmentation for Brain MRI

Medical image classification is often constrained by limited labeled data, motivating generative augmentation; recently, quantum generative…

13:00 JSTエージェント研究/論文

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction

Agents are increasingly deployed in document-intensive workflows where sensitive private information is not an edge case but a routine inpu…