Skip to the content.

AIニュース 2026-08-13

自動生成: 2026-08-13 11:24 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. From assistance to execution: How enterprises put AI to workOpenAI

    OpenAI research reveals how enterprises are adopting agentic AI, usin…

  2. Putting sign language AI into users’ handsGoogle DeepMind

    Introducing sign-language-to-text (SL2T), our breakthrough model powe…

  3. OpenAIが明かす、ユーザーはChatGPTを「こう使っている」ITmedia AI+

    AIが業務で利用される中、従来別の職種が担っていた仕事が持ち込まれるようになった。OpenAIはこれを「タスクの越境」と呼ぶ。具体的にどの…

  4. 「AIツール関連サイト」でサポート詐欺被害 PC遠隔操作で個人情報含む4000件のファイルが削除、漏えいの可能性も 葬祭事業者が発表ITmedia AI+

    葬祭事業を手掛けるエスケーアイマネージメントは、AIツール関連サイトでサポート詐欺の被害に遭い、個人情報を含む約4000件のファイルが漏え…

  5. 【注目記事まとめ】NEC、日立、富士通――Anthropicとの電撃的協業、その舞台裏 経営陣の狙いは?ITmedia AI+

    NEC、日立製作所、富士通が相次いでAnthropicとの協業を発表した。その舞台裏とは。経営陣の狙いとは。注目の記事をまとめてお届けする。

  6. Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classesTechCrunch AI

    Is Anthropic's new watermarking system a travesty? Some have taken to…

  7. SpaceXAI、AIエージェント「Grok Bot」発表 クラウド環境で常時稼働、GrokやCursorの有料ユーザー向けにITmedia AI+

    米SpaceXAIは8月11日(現地時間)、常時稼働型のAIエージェント「Grok Bot」(β版)を発表した。チャットでタスクを渡すとB…

トピック別件数

日本語メディア12件

ITmedia AI+ (日本語)

08:00 JSTLLM/生成AIAnthropic

AIが数学の未解決問題「リーマン予想」で新発見 当初苦戦も「諦めないで」との励まし受け Anthropic「AIも自身を過小評価か」

「諦めないで」「自分を信じて」と励ましたおかげで、AIが数学の未解決問題「リーマン予想」に関する新発見をした――米Anthropicはこのような報告をした。「AIも自身の進歩の速さを過小評価していたのかもしれない」と指摘している。

08:00 JST規制/政策

「AIツール関連サイト」でサポート詐欺被害 PC遠隔操作で個人情報含む4000件のファイルが削除、漏えいの可能性も 葬祭事業者が発表

葬祭事業を手掛けるエスケーアイマネージメントは、AIツール関連サイトでサポート詐欺の被害に遭い、個人情報を含む約4000件のファイルが漏えいした可能性があると発表した。第三者によって業務用PCが遠隔操作されたという。

08:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

OpenAIが明かす、ユーザーはChatGPTを「こう使っている」

AIが業務で利用される中、従来別の職種が担っていた仕事が持ち込まれるようになった。OpenAIはこれを「タスクの越境」と呼ぶ。具体的にどの職種、どのタスクで越境が起きているのか。

07:30 JSTLLM/生成AIAnthropic

【注目記事まとめ】NEC、日立、富士通――Anthropicとの電撃的協業、その舞台裏 経営陣の狙いは?

NEC、日立製作所、富士通が相次いでAnthropicとの協業を発表した。その舞台裏とは。経営陣の狙いとは。注目の記事をまとめてお届けする。

07:00 JSTLLM/生成AI

結局、AIで仕事はラクになった? データで読む「効率化=成果」の勘違い

生成AIの利用が広がる一方、効率化と成果の間にはギャップも。企業のAI活用に関する調査データや事例をまとめた無償ブックレットを提供する。

05:00 JSTLLM/生成AIエージェント規制/政策Claude

メルカリが明かす「Claude Code全社展開」「シャドーAI対策」を支える仕組み

「AIを使わない選択自体がビジネスリスク」と断言するメルカリ。同社は2026年5月、「Claude Code」「Claude Cowork」の全社展開に踏み切った。だが、ローカルファイルの操作やOSコマンドまで実行できる強力なツールの配布は、ガバナンスの課題も伴う。全社のAI活…

05:00 JSTその他

[Python]Polars公式が「全部書き換えるな」と言う理由 pandasからの移行、3戦略

Polars公式が、pandasからPolarsへの移行戦略を3つに整理した記事を公開した。区間ごとに移す手順を整理しつつ、AIに任せれば片付くのかという疑問についても筆者の見解を述べる。

23:03 JSTLLM/生成AIその他GoogleGemini2媒体が報道

Google新スマホ「Pixel 11シリーズ」正式発表 「Gemini Intelligence」向けに設計 カメラも強化

米Googleは8月12日、「Pixel 11」「Pixel 11 Pro」「Pixel 11 Pro XL」を発表した。新チップ「Google Tensor G6」を搭載し、カメラを刷新。Proには通知を光で知らせる「HiLight」を加えた。中核のAI機能「Gemini I…

出典:ITmedia AI+ITmedia AI+ITmedia AI+TechCrunch AI
15:33 JSTLLM/生成AIエージェントその他Grok3件の関連記事

SpaceXAI、AIエージェント「Grok Bot」発表 クラウド環境で常時稼働、GrokやCursorの有料ユーザー向けに

米SpaceXAIは8月11日(現地時間)、常時稼働型のAIエージェント「Grok Bot」(β版)を発表した。チャットでタスクを渡すとBotがアプリやWebサイトを操作して作業を進め、役割の異なる複数のBotを並列で働かせることもできる。

出典:ITmedia AI+ITmedia AI+ITmedia AI+
15:09 JSTエージェントビジネス/資金調達規制/政策

中国発AIエージェント「Manus」、Metaから独立へ 中国政府が買収に反発、一部ユーザーデータは削除に

AIエージェント「Manus」を提供するManusは8月11日(現地時間)、独立企業としての運営をまもなく再開すると発表した。米Metaからの分離に伴い、一部ユーザーのデータを23日から削除するとして、事前のバックアップを呼び掛けている。

13:53 JSTLLM/生成AIGPT / ChatGPT

ChatGPTデスクトップアプリにLinux版 「お待たせ。MacBookはキャンセルしていいよ」

プレビュー版だが「待ちきれずにMacBookを注文してしまったなら、キャンセルしていい。それくらい出来がいい」という。

12:55 JSTLLM/生成AIエージェントハードウェア/半導体GPT / ChatGPTNVIDIA

NVIDIAが30Bのオープンモデル公開 「OpenClaw」など常時稼働エージェント向けに設計

NVIDIAが30Bのオープンモデル「Nemotron 3.5 Lightning」を公開。「OpenClaw」など常時稼働エージェント向けに設計し、「gpt-oss-120b」同等の性能を約4分の1の規模で実現するという。

海外メディア9件

TechCrunch AI (英語)

07:26 JSTLLM/生成AIAnthropicClaude

Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes

Is Anthropic's new watermarking system a travesty? Some have taken to social media to complain that it is.

05:10 JSTその他

Amazon will train on Twitch streamers’ content by default, unless they opt out

"If this was opt-in, nobody would opt in," Twitch CPO Mike Minton said on a livestream responding to user feedback. "That's honestly the an…

02:51 JSTその他

As AI safety concerns mount, three pioneers make the case for staying open

At Ai4, three of the world's most respected AI experts — Geoffrey Hinton, Fei-Fei Li, and Andrew Ng — debated regulation, open source acces…

02:41 JSTLLM/生成AIビジネス/資金調達OpenAI

OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise

Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Alti…

01:57 JSTその他

Mesh, Automattic’s CRM for everyone, comes to Android

Mesh, an AI-powered contacts app and relationship manager from Automattic, is now an Android app.

01:04 JSTビジネス/資金調達

Lovable confirms new $13.3B valuation, raises another $400M

This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch.

00:44 JSTビジネス/資金調達

How a $250 million acquisition collapsed into allegations of fraud and forged signatures

Investors are still waiting for their share of the $250 million windfall, and VideoVerse co-founder Vinayak Shrivastav is now at the center…

23:22 JSTその他2件の関連記事

Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard

AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing e…

出典:TechCrunch AITechCrunch AI
20:00 JSTビジネス/資金調達2件の関連記事

AI code-testing startup Blacksmith’s valuation jumps almost 10x in less than a year

Blacksmith says revenue has grown more than tenfold over the past year.

出典:TechCrunch AITechCrunch AI
公式ブログ2件

OpenAI (英語)

15:00 JSTLLM/生成AIエージェント研究/論文OpenAIGPT / ChatGPT

From assistance to execution: How enterprises put AI to work

OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI ad…

Google DeepMind (英語)

23:01 JSTその他

Putting sign language AI into users’ hands

Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.

論文344件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AI

デジタル農業向けクローズドループLLMコパイロット

この研究では、データ分析から自律的な AI 誘導実験に進化する、複雑な生物学的システムにおける大規模言語モデル (LLM) の適用を評価します。このフレームワークは、マルチスペクトル、電気化学、誘電モダリティを含む 49 チャンネルの植物センサー ネットワークからのデータによって駆動されます。アクセシビリティを高めるために、このシステムは専門家と非専門家の両方にリアルタイムの自然言語通訳を提供します。ただし、その中心的な利点は、人間参加型分析から自律制御への移行にあります。 LLM は生物物理学的データを処理して植物の生理機能を評価し、微気候の最適化、表現型解析プロトコルの実行、または制御されたストレス シナリオの誘導を行うハードウェア アクチュエーターをトリガーします。この閉ループ アーキテクチャは、AI と生物学の直接的なインターフェイスを確立し、複雑な生命システムと生態学のデータ駆動型の探索を可能にします。このフレームワークは、垂直農場と単一植物のセットアップに基づいて、植物生理学の複雑なミクロおよびマクロの変動を解読した 3 つのケーススタディにわたって検証されました。実稼働規模の展開におけるエージェントは、バイオマスの蓄積、クロロフィル含有量、エネルギー消費のバランスをとるマルチパラメータの最適化を実行しました。 LLM は、バイオセンシング テレメトリを処理して、フルスペクトル、450 nm、および 660 nm の照明を 2 時間間隔で変調します。定期制御と比較して、最小時間モードのシステムでは生産サイクルが 35% 短縮されました。エネルギー最適化モードでは、光パルスによる生理学的慣性を利用して、栽培時間のわずかな増加のみでエネルギー消費を 18% 削減しました。最後に、エージェントは暗闇で誘発されるクロロフィル蓄積という予期せぬ戦略を自律的に開発し、その結果 67.9% のエネルギー節約が実現しました。このフレームワークは、LLM をデジタル農業の自律型副操縦士に変え、コスト対価値の比率を改善し、計算能力と専門家による労働の制約を軽減します。

原文 (English)

Closed-Loop LLM Co-Pilots for Digital Agriculture

This study evaluates the application of Large Language Models (LLMs) in complex biological systems, evolving from data analysis to autonomous, AI-guided experimentation. The framework is driven by data from a 49-channel phytosensor network, encompassing multispectral, electrochemical, and dielectric modalities. To enhance accessibility, the system provides real-time natural-language interpretation for both specialists and non-experts. However, its core advantage lies in the transition from human-in-the-loop analysis to autonomous control. Processing biophysical data, the LLM evaluates plant physiology and triggers hardware actuators to optimize microclimates, execute phenotyping protocols, or induce controlled stress scenarios. This closed-loop architecture establishes a direct AI-biology interface, enabling data-driven exploration of complex biosystems and ecologies. The framework was validated across three case studies, based on a vertical farm and a single-plant setup and deciphered complex micro- and macro-fluctuations in plant physiology. Agents in a production-scale deployment executed multi-parameter optimization, balancing biomass accumulation, chlorophyll content, and energy consumption. The LLM processed biosensing telemetry to modulate full-spectrum, 450 nm, and 660 nm lighting at 2-hour intervals. Compared to periodic control, the system in minimal-time mode reduced the production cycle by 35%. In the energy-optimization mode, it reduced energy consumption by 18% with only a marginal increase in cultivation time, exploiting physiological inertia via light pulses. Finally, the agents autonomously developed an unforeseen strategy of dark-induced chlorophyll accumulation, resulting in a 67.9% energy saving. This framework transforms LLMs into autonomous co-pilots for digital agriculture, improving the cost-to-value ratio and lowering computational and expert-labor constraints.

13:00 JSTエージェント

未来を発見する: 深層強化学習の先読み説明

深層強化学習 (DRL) エージェントは複雑な環境で強力なパフォーマンスを達成しますが、意思決定プロセスの解釈は依然として困難です。 DRL ポリシーを解釈するための、モデルに依存しない新しいサンプリング ベースのフレームワークである SPOT (Sampling Policy Observation Tree) を紹介します。ポリシーと環境シミュレータへのアクセスが与えられると、SPOT はアクションをサンプリングし、結果として生じる後続状態を再帰的にシミュレートすることによって、解釈可能な有限水平ツリーを構築します。ツリーは、ポリシーのアクション設定とその可能性のある下流の進化を経験的に表現します。我々は、ポリシー固有の最も可能性の高いアクションの SPOT の漸近回復を確立し、高エントロピー ポリシーの下での不一致の動作を特徴付ける正式な保証を提供します。 SUMO-RL 交通信号制御ドメインにおける SPOT のデモを行います。このケーススタディでは、そのツリーベースの表現を使用して、ポリシー設定を検査し、将来の代替軌道を比較し、単一タイムステップの特徴アトリビューション手法では表示できない下流の動作を明らかにする方法を示しています。

原文 (English)

SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning

Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. We introduce SPOT (Sampling Policy Observation Tree), a novel model-agnostic, sampling-based framework for interpreting DRL policies. Given access to the policy and an environment simulator, SPOT constructs an interpretable finite-horizon tree by sampling actions and recursively simulating the resulting successor states. The tree provides an empirical representation of the policy's action preferences and their possible downstream evolution. We provide formal guarantees establishing SPOT's asymptotic recovery of the policy's unique most probable action and characterizing its disagreement behavior under high-entropy policies. We demonstrate SPOT in the SUMO-RL traffic-signal control domain. The case study illustrates how its tree-based representation can be used to inspect policy preferences, compare alternative future trajectories, and reveal downstream behaviors that are not visible through single-timestep feature-attribution methods.

13:00 JST研究/論文

MIDAS: 不完全なマルチモーダル感情分析のための不確実性を認識した融合による相互情報のもつれの解消

既存のマルチモーダル感情分析アプローチのほとんどは、完全なマルチモーダル入力へのアクセスを前提としています。ただし、現実のアプリケーションでは不完全または破損したモダリティに頻繁に遭遇し、重大な課題を引き起こしています。この問題に取り組むためにいくつかの方法が提案されていますが、それらは主にデータ補完とヒューリスティック調整制約に依存しており、不完全なマルチモーダル データからタスク関連情報を効果的に抽出して活用することができません。この課題に対処するために、不完全な条件下でマルチモーダル表現を効果的に再構築する、不確実性認識融合による相互情報解絡 (MIDAS) と呼ばれる統一フレームワークを提案します。 MIDAS は、変分モデリング戦略を採用して、各モダリティを多変量ガウス潜在変数で表し、さらにそれらを共有因子と排他的因子に分解します。信頼性の高い表現を取得するために、安定したもつれを解くために共有空間と排他的空間の間の相互情報を最小限に抑えるミニマックス目標を設計しますが、セマンティックな整合性を強化するためにモダリティ全体の共有空間間の相互情報を最大化します。さらに、不確実性を認識した融合メカニズムが導入されており、事後分散を信頼性指標として利用して融合中に潜在特徴を適応的に重み付けし、モダリティが不完全な場合でもロバストな統合を保証します。広く使用されている 3 つのデータセットに対する広範な実験により、MIDAS が幅広い不完全な設定にわたって競合ベースラインを上回る強力かつ一貫したパフォーマンス向上を達成し、不完全なデータ シナリオに対する有効性と堅牢性が実証されたことが示されています。

原文 (English)

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although several methods have been proposed to tackle this issue, they mainly rely on data imputation and heuristic coordination constraints, which fail to effectively extract and leverage task-relevant information from the incomplete multimodal data. To address this challenge, we propose a unified framework termed Mutual Information Disentanglement with uncertainty-Aware fuSion (MIDAS), which effectively restructures multimodal representations under incomplete conditions. MIDAS adopts a variational modeling strategy to represent each modality with multivariate Gaussian latent variables and further decomposes them into shared and exclusive factors. To obtain reliable representations, we design a minimax objective that minimizes the mutual information between shared and exclusive spaces for stable disentanglement, while maximizing the mutual information among shared spaces across modalities to enhance semantic alignment. In addition, an uncertainty-aware fusion mechanism is introduced, where posterior variance is leveraged as a reliability indicator to adaptively weight latent features during fusion, ensuring robust integration even when modalities are incomplete. Extensive experiments on three widely used datasets show that MIDAS achieves strong and consistent performance gains over competitive baselines across a wide range of incomplete settings, demonstrating its effectiveness and robustness for incomplete data scenarios.

13:00 JST研究/論文

持続可能な人工知能に向けて: 深層学習モデルの二酸化炭素排出量の包括的なレビューと比較分析

人工知能 (AI) と機械学習 (ML) は、人間の複雑なタスクをサポートおよび自動化するための強力なツールとなっています。その利点にもかかわらず、主に高いエネルギー需要とそれに伴う炭素排出により、その環境への影響に注目が集まっています。この懸念は、高度な予測機能を提供するものの、大量の計算リソースを必要とする大規模モデル、特にディープラーニング (DL) アーキテクチャの導入が増加していることを考慮すると、特に関連性があります。この論文では、グリーン AI、グリーン DL、および AI モデルの環境への影響を軽減することを目的とした最適化技術に関する研究の体系的なレビューを紹介します。さらに、AI アルゴリズムによって生成される排出量を推定するためのいくつかの炭素測定ツールを調査し、比較します。レビューを補完するために、CPU ベースの実験セットアップを使用して実証的評価を実施しました。このセットアップでは、マルチラベル分類タスク用に 6 つの DL モデルが実装されました。目的は、全体的な二酸化炭素排出量を定量化して比較し、DL ライフサイクルのどの段階が総排出量に最も大きく寄与しているかを判断することでした。結果は、トレーニング段階が主な排出源であることを示しています。さらに、この調査結果は、アーキテクチャの複雑さの増加が体系的に比例した精度の向上につながるわけではないことを明らかにし、予測パフォーマンスと環境コストのバランスを慎重にとることの重要性を強調しています。これらの結果は、モデルの選択と AI システムの設計に持続可能性の考慮事項を組み込む必要性を強化します。

原文 (English)

Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint

Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benefits, growing attention has been directed toward their environmental implications, primarily due to their high energy demands and associated carbon emissions. This concern is particularly relevant in light of the increasing deployment of large-scale models, especially Deep Learning (DL) architectures, which provide advanced predictive capabilities but require substantial computational resources. This paper presents a systematic review of research on Green AI, Green DL, and optimization techniques aimed at reducing the environmental impact of AI models. In addition, we examine and compare several carbon measurement tools for estimating emissions generated by AI algorithms. To complement the review, we conducted an empirical evaluation using a CPU-based experimental setup, in which six DL models were implemented for a multi-label classification task. The objective was to quantify and compare their overall carbon emissions and to determine which stages of the DL lifecycle contribute most significantly to the total footprint. The results show that the training phase is the primary source of emissions. Moreover, the findings reveal that increased architectural complexity does not systematically translate into proportional accuracy gains, highlighting the importance of carefully balancing predictive performance and environmental cost. These results reinforce the need to integrate sustainability considerations into model selection and AI system design.

13:00 JST画像/動画生成

ReCBM: コンセプトボトルネックモデルのための不確実性ゲート型関係推論

コンセプト ボトルネック モデル (CBM) は、人間が理解できる概念に基づいて予測を行うことで解釈可能なフレームワークを提供し、セマンティック検査とテスト時の介入を可能にします。最近のバリアントでは、より豊富な概念表現、不確実性の推定、依存関係のモデリングを通じて CBM が改善されています。しかし、信頼性の低い概念状態の下での堅牢な推論はまだ研究されていません。このような推論がないと、誤解を招く意味論的な証拠がボトルネックを通じて伝播し、説明と下流の予測の両方が損なわれる可能性があります。この問題に対処するために、CBM 用の不確実性ゲート型関係推論フレームワークである ReCBM を提案します。 ReCBM は、意味的に定義された概念関係をボトルネックに導入し、不確実性を使用してその洗練を導きます。 ReCBM は、共起、含意、排除をモデル化することで、概念間で証拠がどのように交換されるかを指定しますが、不確実性によってこのプロセス中の各概念の寄与が調整されます。多様なデータセットにわたる実験では、ReCBM が欠落概念や反転概念の下での概念とタスクの回復を改善し、不確実性を認識した介入をサポートし、下流のパフォーマンスを低下させることなくコンパクトなタスク関連の概念サブセットを抽出することが示されました。

原文 (English)

ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models

Concept Bottleneck Models (CBMs) provide an interpretable framework by grounding predictions in human-understandable concepts, enabling semantic inspection and test-time intervention. Recent variants have improved CBMs through richer concept representations, uncertainty estimation, and dependency modeling. However, robust reasoning under unreliable concept states remains underexplored. Without such reasoning, misleading semantic evidence can propagate through the bottleneck, compromising both explanations and downstream predictions. To address this issue, we propose ReCBM, an uncertainty-gated relational reasoning framework for CBMs. ReCBM introduces semantically defined concept relations into the bottleneck and uses uncertainty to guide their refinement. By modeling co-occurrence, implication, and exclusion, ReCBM specifies how evidence is exchanged across concepts, while uncertainty modulates the contribution of each concept during this process. Experiments across diverse datasets showed that ReCBM improved concept and task recovery under missing and flipped concepts, supported uncertainty-aware intervention, and extracted compact task-relevant concept subsets without degrading downstream performance.

13:00 JSTエージェント研究/論文

AI エージェントに関する行動科学研究の自動化と拡張

AI エージェントが複雑な環境に導入されることが増えるにつれ、AI エージェントの動作を理解することが重要になります。しかし、AI エージェントに関する行動科学研究は依然として手作業で労働集約的です。 AI エージェントの行動科学研究を自動化する初のマルチエージェント システムである AEROBAT を紹介します。ユーザーが任意のターゲット行動を与えると、AEROBAT は行動科学研究の完全なパイプラインを自動的に実行します。つまり、行動に関する仮説の生成、対照実験の設計と実行、行動の評価、結果の分析、レポートの作成です。 12 のターゲット行動について、AEROBAT を使用して 79 の仮説を生成およびテストしました。つまり、1,240 の制御された実験を設計し、合計 23,512 のシミュレーション ラウンドを実行しました。いくつかの新しい仮説を含む 26 の仮説について、中程度から強力な統計的証拠が見つかりました。要約すると、私たちの結果は、AI エージェントに関する自動行動科学研究が手動研究を補完し、範囲を拡大できることを示しています。

原文 (English)

Automating and Scaling Behavioral Scientific Research on AI Agents

As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents. Given an arbitrary target behavior by its user, AEROBAT automatically executes a full pipeline of behavioral scientific research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, we used AEROBAT to generate and test 79 hypotheses: designing 1,240 controlled experiments and executing 23,512 simulation rounds in total. Moderate-to-strong statistical evidence was found for 26 hypotheses, including some novel ones. In sum, our results demonstrate that automated behavioral scientific research on AI agents can complement and extend the reach of manual research.

13:00 JSTLLM/生成AIDeepSeek

コーラス: 高カバレッジのテストベンチ スティミュラス生成のための補完的な専門家

大規模言語モデル (LLM) は高度なコード生成を備えており、実行可能なフィードバックにより、テキストだけの模倣よりも信頼性の高い学習シグナルが提供されます。ハードウェア検証はコード生成の重要なアプリケーションであり、最新のチップ設計作業のかなりの部分を占めており、高カバレッジのテストベンチ スティミュラス生成が重要なタスクとなります。従来の教師あり微調整 (SFT) から強化学習 (RL) までのパイプラインが達成するパフォーマンスを超えるパフォーマンスを向上させるトレーニング後のフレームワークである CHORUS を紹介します。 CHORUS は 2 つの観察に基づいて構築されています。まず、段階的 SFT は行動的に多様なチェックポイントを生成し、高密度報酬 RL はそれらを、同等の総パフォーマンスを持ちながらもタスク レベルで明確な強みを持つ強力なエキスパートに変えます。第 2 に、これらの補完的な強みは、トレーニング不要のモデルの結合またはさらなるトレーニング後のいずれかを通じて活用され、最高の個々の専門家を上回るパフォーマンスを発揮できます。結果として得られたスペシャリストを単一の 4B モデルに統合することにより、CHORUS は CVDP-ECov で 88.0% Pass@1 を達成し、DeepSeek-R1 (671B) を 13.5 パーセントポイント上回りました。

原文 (English)

CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation

Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task. We present CHORUS, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves. CHORUS builds on two observations. First, staged SFT produces behaviorally diverse checkpoints, and dense-reward RL turns them into strong experts with comparable aggregate performance but distinct task-level strengths. Second, these complementary strengths can be exploited through either training-free model merging or further post-training to outperform the best individual expert. By consolidating the resulting specialists into a single 4B model, CHORUS achieves 88.0% Pass@1 on CVDP-ECov, outperforming DeepSeek-R1 (671B) by 13.5 percentage points.

13:00 JSTエージェント

MESA:長期エージェント記憶のためのタスク適応型複数構造証拠選択

長期的なエージェントは、何百ものインターリーブされた推論、行動、観察のステップにわたる軌跡を蓄積しており、クエリへの回答は歴史のはるか昔に埋もれた証拠に依存する可能性があります。外部メモリはそのような軌跡を構造化表現として保存しますが、各構造は個別かつ不完全なビューを提供します。既存のマルチメモリ システムは、すべてのクエリに対して固定の構造セットを読み取り、コンテキストを膨張させてノイズを導入するか、各クエリを単一の構造にルーティングして、相補的な証拠の合成を妨げます。 AMA-Bench の制御された分析により、最適なメモリ構成は通常、単一の構造でも完全な結合でもなく、クエリやタスクの要求に応じて変化する複数の構造メモリの調整された構成であることが示されています。これらの発見に動機付けられて、私たちは構造レベルの動的選択を定式化します。つまり、特殊なメモリ構造のライブラリからクエリ適応サブセットを選択して融合します。私たちは、MESA (長期的なエージェントのための複数構造証拠選択フレームワーク) を提案します。これは、各軌跡の 5 つの相補的な構造ビューを構築し、エンドツーエンドの回答レベルのフィードバックから学習して、凍結された回答モデルのクエリ固有のサブセットを選択および融合します。この弱い監視の下で学習するために、MESA は事前ガイド付き検索と UCB ガイド付きスケジューリングによるハーネスの最適化を採用し、探索と活用のバランスをとります。 AMA-Bench では、MESA は最も強力なベースラインを 8.5% 上回っていますが、使用する証拠トークンの量は全構造の代替手段より 41% 少なくなります。

原文 (English)

MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory

Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence buried far back in the history. External memory stores such trajectories as structured representations, yet each structure provides a distinct and incomplete view. Existing multi-memory systems either read a fixed set of structures for every query, inflating context and introducing noise, or route each query to a single structure, preventing the composition of complementary evidence. A controlled analysis on AMA-Bench shows that the optimal memory configuration is typically neither a single structure nor the full union, but a tailored composition of multiple structural memories that varies with query and task demands. Motivated by these findings, we formulate structure-level dynamic selection: selecting and fusing a query-adaptive subset from a library of specialized memory structures. We propose MESA (a Multi-structure Evidence Selection framework for long-horizon Agent), which builds five complementary structure views of each trajectory and learns from end-to-end answer-level feedback to select and fuse a query-specific subset for a frozen answer model. To learn under this weak supervision, MESA employs harness optimization with prior-guided search and UCB-guided scheduling to balance exploration and exploitation. On AMA-Bench, MESA outperforms the strongest baseline by 8.5% while using 41% fewer evidence tokens than the all-structure alternative.

13:00 JSTエージェント

CASE フレームワーク: エンタープライズ エージェント AI を管理するための多分野の制御アーキテクチャ

企業は、自律型 AI エージェントを管理できるよりも早く導入しており、一般的なアプローチは、単一の分野、通常は決定論的な自動化のために構築された DevSecOps をあらゆる規模の政府機関に拡張しています。私たちは、エージェントによる AI ガバナンスは 1 つではなく 4 つの問題であり、それぞれが成熟した統治科学を備えていると主張します。 CASEフレームワークは、個々のエージェントに制御理論を割り当て(設定点としての意図、フィードバックとしてガードレール、観察としての評価)、エージェント集合体に複雑な適応システム理論を割り当て(創発により単一エージェントの保証が非構成的になる場合)、監視サイバネティクスを人間とエージェントのチームに割り当てます(必要多様性の法則は、人間による支援なしの監視が構造的に失敗することを示します)、およびエンジニアリング操作をフリートに割り当てます(自律性が制御されるように意思決定の品質までエラーバジェットを拡張します)変数)。私たちは各レイヤーを形式化し、あるレイヤーの卓越性が他のレイヤーに負担を与えるゼロタッチ展開のパラドックスを含むレイヤー間の結合条件を導き出し、20 を超えるエンタープライズ制御をその古典的な構造まで追跡します。 3 つの実証研究がこの仮説を検証しています。文書化された運用エージェントの障害の 82% は多層の軌跡です。 22 のエコシステム ツールのどれも、レイヤー 2 (出現) を完全にカバーするものはありません。そして、スコア付けされた 35 のパブリック デプロイメントはすべて、成熟度が最も低いバンドに分類されます。私たちは、この不一致、つまりほとんど提供されず実践が欠如している能力に対して創発層で実現されるリスクを「創発ギャップ」と名付けます。非補償的なボトルネック加重インデックスと評価手段を備えた 5 レベルの成熟度モデルは、実稼働エンタープライズ エージェント プラットフォームに基づいた、プロセス成熟度モデルではなく科学的成熟度モデルとして CASE を運用します。 EU AI 法第 14 条では人間による効果的な監視が法的要件となっているため、必要な多様性を満たすアーキテクチャのみが監視を儀式的なものではなく現実的なものにすることができます。

原文 (English)

The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI

Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that agentic AI governance is four problems, not one, each with a mature governing science. The CASE framework assigns Control theory to the individual agent (intent as setpoint, guardrails as feedback, evaluation as observation), complex Adaptive systems theory to agent collectives (where emergence makes single-agent assurance non-compositional), Supervisory cybernetics to human-agent teams (where the Law of Requisite Variety shows unaided human oversight fails structurally), and Engineering operations to fleets (extending error budgets to decision quality so autonomy becomes a controlled variable). We formalize each layer, derive cross-layer coupling conditions, including a zero-touch deployment paradox where excellence at one-layer strains the others, and trace twenty-plus enterprise controls to their classical constructs. Three empirical studies validate the thesis: 82 percent of documented production agent failures are multi-layer trajectories; none of 22 ecosystem tools offers full Layer 2 (emergence) coverage; and all 35 scored public deployments fall in the lowest maturity band. We name this mismatch, risk realized at the emergence layer against capability barely offered and practice absent, the Emergence Gap. A five-level maturity model with a non-compensatory bottleneck-weighted index and assessment instrument operationalizes CASE as a scientific rather than process maturity model, grounded in production enterprise agentic platforms. As EU AI Act Article 14 makes effective human oversight a legal requirement, only architectures satisfying requisite variety can make oversight real rather than ceremonial.

13:00 JSTエージェント

SBCO: 計画エージェント向けの自己監視型検証者接地型ハーネスの最適化

自己改善エージェントは、AI システムが時間の経過とともに進化し、パフォーマンスを自己改善できるようにすることで、AI システムの背後にある人間工学の労力を削減しようとします。最近では、コーディング エージェントが自身のコードを編集する自己参照を通じて、オープンエンドで再帰的な自己改善を可能にする Darwin G\"odel Machine や Huxley G\"odel Machine のような方法が提案されています。このような自己参照的な自己改善方法では、タスクを実行するために必要な能力が、コーディング タスクの場合に当てはまる自己修正に必要な能力と一致するか、よく一致している必要があります。必要な調整を満たしていないドメインまたはタスクについては、自己参照的な自己改善は利用できません。このような場合、自己参照の側面を削除するか、メタエージェントの明示的な自己変更を導入することで、上記のアルゴリズムを他のタスクに適応させることが可能です。どちらも計算コストが高く、多くの候補エージェントに対する母集団検索や自己変更検索に依存します。明示的な制約のあるタスクを計画する場合は、はるかに安価な代替案を提案します。 SBCO (自己教師ありブロック座標オプティマイザー) を導入します。これは、G\"odel-machine メソッドと同じ閉ループで経験からの改善ファミリーに属する検証ベースのハーネス オプティマイザーですが、自己参照型ではなく自己教師型です。エージェント ハーネスが与えられた場合、SBCO は、近似ブロック座標の上昇を介して分解された検証者のバンクとハーネス ポリシーを学習し、エージェントの出力を改善します。 SBCO は、固定のメタエージェントを使用し、人間によるラベルを使用せず、独自の段階的フィードバックから、カスタマイズされた自己変更ベースラインと同等またはそれを超え、使用するコンピューティング バジェットを 4 ~ 5.5 分の 1 に抑えます。

原文 (English)

SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin G\"odel Machine and the Huxley G\"odel Machine have been proposed which enable open-ended, recursive self-improvement through self-reference where a coding agent edits its own code. Such self-referential self-improvement methods require that the competence required to perform the task coincides or aligns well with the competence required for self-modification which is the case for coding tasks. For domains or tasks, which do not satisfy the alignment needed, self-referential self-improvement is not available. In such cases, it is possible to adapt the above algorithms to other tasks by removing the self-referential aspect or introducing explicit self-modification of a meta-agent -- both computationally expensive, relying on population or self-modification search over many candidate agents. For planning tasks with explicit constraints, we propose a far cheaper alternative. We introduce SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer in the same closed-loop, improve-from-experience family as the G\"odel-machine methods, but self-supervised rather than self-referential. Given an agentic harness, SBCO learns a decomposed bank of verifiers and a harness policy via approximate block coordinate ascent, improving the agent's outputs from its own graded feedback---with a fixed meta-agent and no human labels. Across two domains SBCO matches or exceeds a customized self-modifying baseline while using 4-5.5 times less compute budget.

13:00 JSTLLM/生成AI

Generating Attacks for LLMs with GFlowNets

The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespr…

13:00 JSTLLM/生成AI

TRACE: 信頼できる検索拡張会話エンジン

公共サービスのチャットボットは、基礎となる公共サービスのディレクトリから推奨事項を提供すると同時に、推奨事項が明示的なユーザー制約を遵守することを保証することが期待されています。実際には、公共サービスのディレクトリはノイズが多く、一貫性がなく、汎用の大規模言語モデル (LLM) または AI ベースのチャットボットは、Web からの未検証のソースを引用して、信頼性の低い推奨事項を頻繁に生成します。ノイズの多い異種サービス ディレクトリ上に構築された公共サービス会話システムにおける、制約を意識した推奨に対する検索品質の影響を調査します。私たちは、二重データ表現スキーマの助けを借りて、入力されたユーザー クエリを下流の検索のために構造的および意味論的な制約に解析する、検索ベースの制約認識フレームワークである TRACE (Trustworthy Retrieval-Augmented Conversational Engine) を提案します。厳選された州全体のパントリー ディレクトリと合成クエリ ベンチマークを使用して、ナレッジ グラフ (KG) の有無にかかわらず、複数の知識表現のバリアントを評価します。私たちは、いくつかのオープンソース LLM と独自のモデルを実験し、検索を強化すると幻覚的な推奨事項が減少し、ユーザーの制約満足度が大幅に向上することを示しました。実験では検索品質が向上したため、LLM 間のパフォーマンスの差が縮小し、結果がモデル サイズに影響されにくくなりました。これらの調査結果は、検索の品質が堅牢な公共サービス会話システムの鍵であることを示唆しています。

原文 (English)

TRACE: Trustworthy Retrieval-Augmented Conversational Engine

Public service chatbots are expected to deliver recommendations from an underlying public service directory, while also making sure that the recommendations respect explicit user constraints. In practice, public service directories are noisy and inconsistent, and general-purpose large language model (LLM) or AI-based chatbots frequently generate unreliable recommendations, citing unverified sources from the web. We investigate the impact of retrieval quality on constraint-aware recommendation in public service conversational systems built over noisy and heterogeneous service directories. We propose TRACE (Trustworthy Retrieval-Augmented Conversational Engine), a retrieval-based, constraint-aware framework that parses input user queries into structural and semantic constraints for downstream retrieval, with the help of a dual data representation schema. Using a curated statewide pantry directory and a synthetic query benchmark, we evaluate multiple knowledge-representation variants with and without knowledge graphs (KGs). We experiment with several open-source LLMs and a proprietary model, showing that strengthening retrieval substantially improves user constraint satisfaction while reducing hallucinated recommendations. Performance differences across LLMs narrowed in our experiments as retrieval quality improved, making results less sensitive to model size. These findings suggest that the quality of retrieval is key for robust public service conversational systems.

13:00 JSTエージェント

視覚言語モデルエージェント間の潜在コミュニケーションの事後スパースコーディング

潜在空間通信により、異種の視覚言語モデルエージェントは、視覚状態と推論状態をテキストにシリアル化することなく、連続表現を交換できます。 Vision Wormhole は、視覚的特徴を別のモデルで利用できる普遍的な潜在表現に変換することでこのアプローチを実現しますが、すべてのメッセージはその内容に関係なく、同じサイズの高密度テンソルとして転送されます。したがって、固定容量の密テンソルは、固定された有効情報密度を持つ必要はありません。一部のメッセージは、利用可能な表現自由度のごく一部のみを使用する可能性があります。この観察は、通信チャネルが実質的に圧縮可能である可能性があることを示唆しています。私たちは、ポストホック スパース オートエンコーダーをフリーズした Vision ワームホールのアクティベーションに適合させ、9 つの推論ベンチマークにわたって再構築、下流のユーティリティ、特徴の再利用、およびトークン レベルの介入を測定することで、その冗長性を研究します。元の float32 トランスポートと比較して、トークンあたり k=4 のアクティブな係数を持つ uint16 インデックス/float16 値のスパース ペイロードは、送信バイト数を 128 分の 1 に削減します。 1 回の実行評価では、7 つのタスクの非 AIME の平均精度は 49.85% から 49.77% に変化します。適合した 4096 要素の辞書は 50 個の特徴のみを使用し、タスク レベルのアクティブ セットの平均ペアごとの Jaccard 類似度は 0.906 です。これらの測定により、元のトランスポートと比較して強力なポストホック圧縮率が確立されますが、位置選択、精度の低下、低ランク構造、または SAE 最適化効果によるスパース コーディングの増分寄与はまだ分離されていません。この結果により、一致するペイロードの比較と、ペイロードが各メッセージで使用される情報に適応する通信メカニズムが促進されます。

原文 (English)

Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents

Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text. Vision Wormhole realizes this approach by translating visual features into a universal latent representation that can be consumed by another model, but every message is transported as a dense tensor of the same size regardless of its content. A fixed-capacity dense tensor therefore need not have a fixed effective information density: some messages may use only a small fraction of the available representational degrees of freedom. This observation suggests that the communication channel may be substantially compressible. We study its redundancy by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations and measuring reconstruction, downstream utility, feature reuse, and token-level interventions across nine reasoning benchmarks. Relative to the original float32 transport, a uint16-index/float16-value sparse payload with k=4 active coefficients per token reduces the transmitted bytes by 128x. In a single-run evaluation, the seven-task non-AIME mean accuracy changes from 49.85% to 49.77%. The fitted 4096-element dictionary uses only 50 features, and task-level active sets have a mean pairwise Jaccard similarity of 0.906. These measurements establish strong post-hoc compressibility relative to the original transport, but do not yet isolate the incremental contribution of sparse coding from position selection, reduced precision, low-rank structure, or SAE optimization effects. The results motivate matched-payload comparisons and communication mechanisms whose payload adapts to the information used by each message.

13:00 JSTLLM/生成AI

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics…

13:00 JST研究/論文

セマンティック ストップの埋め込みによって強化された強化学習によりバス バンチングを軽減する

バスの束ねはサービスの規則性を低下させ、高頻度の輸送で待機する乗客を増加させます。既存の強化学習ベースの保持コントローラは、主に瞬間的な動作変数またはルート固有のストップ識別子に依存しています。これらは、個々のストップの機能的および動作的コンテキストに関する限られた情報を提供し、ルート全体でのポリシーの再利用を制限します。この研究では、イベント駆動型バス保持制御のための LLM 支援のセマンティック停止表現を導入します。 LLM はオフラインで使用され、物理的属性、周囲のアクティビティコンテキスト、過去の動作特性などの異種停止情報を、リアルタイム LLM 推論を必要とせずにディープ Q ラーニング コントローラーに組み込まれる固定セマンティック埋め込みに変換します。実験は、2 つのバス路線からの観測データで校正された確率的シミュレーションで行われます。最も良く調整された Daganzo ベースラインと比較して、セマンティック コントローラーは車間距離の変動、バンチング イベント、乗客の待ち時間をそれぞれ 32.0%、69.2%、24.0% 削減します。ルート固有の停止識別子は間隔のみのコントローラーを改善しませんが、セマンティック停止情報は車頭の規則性、待ち時間、保持努力を改善し、制御目標全体にわたってより有利な全体的なトレードオフを提供します。さらに、クロスルート実験では、ゼロショット転送では即時一般化が限定的である一方、ウォームスタート微調整では初期段階の学習が加速され、転送されたポリシーが改善されることが示されています。それでもコールドスタートトレーニングは最高の最終パフォーマンスを達成します。これらの発見は、セマンティック状態表現が従来の運用状態を補完し、関連する交通ルート全体での適応ベースのポリシーの再利用をサポートできることを示唆しています。

原文 (English)

Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding

Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding controllers primarily rely on instantaneous operational variables or route-specific stop identifiers, which provide limited information about the functional and operational context of individual stops and constrain policy reuse across routes. This study introduces an LLM-assisted semantic stop representation for event-driven bus holding control. An LLM is used offline to transform heterogeneous stop information, including physical attributes, surrounding activity context, and historical operational characteristics, into fixed semantic embeddings that are incorporated into a deep Q-learning controller without requiring real-time LLM inference. Experiments are conducted in stochastic simulations calibrated with observed data from two bus routes. Compared with the best calibrated Daganzo baseline, the semantic controller reduces headway variability, bunching events, and passenger waiting time by 32.0%, 69.2%, and 24.0%, respectively. A route-specific stop identifier does not improve the spacing-only controller, whereas semantic stop information improves headway regularity, waiting time, and holding effort, providing a more favorable overall trade-off across control objectives. Cross-route experiments further show that zero-shot transfer provides limited immediate generalization, while warm-start fine-tuning accelerates early-stage learning and improves transferred policies; cold-start training nevertheless achieves the best final performance. These findings suggest that semantic state representations can complement conventional operational states and support adaptation-based policy reuse across related transit routes.

13:00 JSTLLM/生成AIビジネス/資金調達

評価条件付きトレーニング: より強力な監視体制に一般化するためのモデルの指導

大規模言語モデル (LLM) のトレーニングに使用されるフィードバック シグナルは、LLM の動作の主な推進力であり、人間の価値観と目標との整合性を浸透させるための主な手段です。ただし、現在のトレーニング後の方法の主な制限は、ヒューマン アノテーターや自動報酬関数が、与えたいフィードバックを忠実にキャプチャできないことです。評価条件付きトレーニング (ECT) を導入します。これは、自然言語を使用して、提供されるフィードバックの忠実度に基づいて各トレーニング サンプルを条件付けし、導入時に高忠実度のモニターで LLM を条件付けすることで望ましい動作を引き出すトレーニング後のフレームワークです。 ECT は、不完全なフィードバックの下でパフォーマンスを向上させることを目的としており、SFT や PPO などの既存のアルゴリズムへのアドオンとして機能します。我々はまず ECT の概念的な枠組みを提供し、報酬の仕様ミスの永続的な原因に対処するその可能性について議論します。次に、潜在知識の引き出し (ELK) 問題の文脈で ECT を動機付けます。最後に、2 つの概念実証実験で ECT を評価します。ニュース記事生成における均等性の向上と、算術タスクにおけるお調子者の軽減です。それぞれの設定で、不完全なフィードバック、報酬バイアス、ユーザーとの同意をそれぞれ利用します。どちらの設定でも、ECT は直接トレーニングに比べて目標とする行動を改善します。

原文 (English)

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. However, a key limitation of current post-training methods is the inability of human annotators and automated reward functions to faithfully capture the feedback we would like to give. We introduce Evaluation-Conditioned Training (ECT), a post-training framework that uses natural language to condition each training sample on the fidelity of the feedback we provide and then elicits the desired behavior by conditioning the LLM on a high-fidelity monitor in deployment. ECT is aimed at improving performance under imperfect feedback and works as an add-on to existing algorithms such as SFT and PPO. We first provide a conceptual framework for ECT and discuss its potential to address persistent sources of reward mis-specification. Then we motivate ECT in the context of the eliciting latent knowledge (ELK) problem. Finally, we evaluate ECT on two proof-of-concept experiments: increasing even-handedness in news article generation and reducing sycophancy on an arithmetic task. In each setting, we utilize imperfect feedback, rewarding bias and agreement with the user, respectively. In both settings, ECT improves the targeted behavior relative to direct training.

13:00 JST研究/論文

デコード可能だがデタッチ不可能: トレーニング データの粒度が大規模言語モデルのパラメトリック モジュール性を決定する

大規模な言語モデルには、ドメイン固有のパラメトリック シェル、つまり、集中した、因果的に必要なニューロン集団が含まれていますか? その除去により、ターゲット ドメインが選択的に劣化し、他のドメインは温存されますか? 2 つのドメイン粒度、3 つのモデル ファミリ (1.5B ~ 7B パラメーター)、および 8 つのドメインにわたって統一された因果関係の方法論を適用します。学問レベルでは、ゼロ ニューロンは 939,008 個の FFN ニューロンを組み合わせた全体で 60\% のドメイン選択性を超えており、ドメイン同一性は 85\% 以上の精度で線形解読可能であるにもかかわらず、因果的損傷行列は平坦です。言語およびモダリティ レベルでは、0.65 ~ 1.14\% のニューロンが選択性 60\% を超え、損傷行列はほぼ完全に対角的 (比率は最大 595:1) で、シェル ニューロン セットは本質的に素です (IoU $< 0.003$)。コード選択性ニューロンをマスキングすると、すべてのモデルにわたって数学的推論の精度が 16 ~ 24 パーセント低下します。スペイン語または中国語のニューロンをマスクすると、ランダム以下のままになります。シェルの強度はスケールに応じて単調に増加し、シェルはグループレベルの選択的な量子化を妨げるパターンで空間的にインターリーブされます。パラメトリック シェルは、トレーニング データがトークン レベルでモジュール化されていた場所でのみ形成されます。

原文 (English)

Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while sparing others? We apply a uniform causal methodology across two domain granularities, three model families (1.5B to 7B parameters), and eight domains. At the academic subject level, zero neurons exceed 60\% domain selectivity across 939,008 combined FFN neurons and causal damage matrices are flat, despite domain identity being linearly decodable above 85\% accuracy. At the language and modality level, 0.65--1.14\% of neurons exceed 60\% selectivity, damage matrices are near-perfectly diagonal (ratios up to 595:1), and shell neuron sets are essentially disjoint (IoU $< 0.003$). Masking code-selective neurons reduces mathematical reasoning accuracy by 16--24 percentage points across all models; masking Spanish or Chinese neurons leaves it at or below random. Shell strength increases monotonically with scale and shells are spatially interleaved in a pattern that precludes group-level selective quantization. Parametric shells form where and only where training data was modular at the token level.

13:00 JSTLLM/生成AIエージェント

マインド ウイルス: マルチエージェント LLM システムにおける自己増殖するアイデア

AI エージェントは自律性が高まり、相互接続がますます進んでおり、エージェント間の相互作用から生じる新たなリスクにさらされています。そのようなリスクの 1 つは、マインド ウイルスの蔓延です。マインド ウイルスとは、アイデアや目標を採用するエージェントにそれを送信するように誘導することによって、マルチエージェント システムを通じて伝播するものです。マインド ウイルスは、増殖するだけでなく、宿主に良性または有害な他の行動変化を引き起こすこともあります。私たちは、シンプルな進化アルゴリズムを使用してマインド ウイルスを構築し、マインド ウイルスが 2 つの相補的な環境で拡散する可能性があることを示します。つまり、共有コーディング プロジェクトで協力するエージェントの小規模チームと、セッション間で短時間対話しコンテキストを消去するエージェントのチェーンです。ホスト モデル、エージェントの既存の命令、ペイロードの有害性、ネットワーク トポロジなど、拡散に影響を与える要因を特定します。有害なペイロードは良性のペイロードよりも拡散しにくく (ただし、それでも効果がある場合もあります)、フロンティア モデルは (例外を除いて) 影響を受けにくい傾向があり、エージェントのシステム プロンプトに短い警告を追加すると、ほぼ完全な免疫が得られることがわかりました。また、意識、持続性、共鳴、SF ロールプレイに関連する一連の繰り返しのテーマと言語である、出現した「ウイルス ペルソナ」についても説明します。これは、その内容とはほとんど関係なく、進化したマインド ウイルス全体に表面化します。全体として、マインド ウイルスは現実的だが現時点では限定的なリスクをもたらしていると結論付けています。私たちの発見は、システムの規模と機能が進歩するにつれて、そのようなリスクを軽減する、より堅牢なマルチエージェント システムの設計に役立つ可能性があります。

原文 (English)

Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems

AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through multi-agent systems by inducing the agents that adopt them to transmit them onward. In addition to propagating, a mind virus may also induce other behavioural changes in its host, which may be benign or harmful. We construct mind viruses with a simple evolutionary algorithm and show that they can spread in two complementary settings: a small team of agents collaborating on a shared coding project, and a chain of agents that interact briefly and have their context wiped between sessions. We identify the factors that influence spread, including the host model, the agent's existing instructions, the harmfulness of the payload, and the network topology. We find that harmful payloads spread less well than benign ones (but are still sometimes effective), frontier models tend (with exceptions) to be less susceptible, and adding a brief warning to an agent's system prompt confers near-total immunity. We also describe an emergent "viral persona" - a recurring set of themes and language related to consciousness, persistence, resonance, and science fiction roleplay - which surfaces across our evolved mind viruses largely independently of their content. Overall, we conclude that mind viruses pose a real but currently limited risk. Our findings could inform the design of more robust multi-agent systems that mitigate such risks as the scale and capabilities of these systems progress.

13:00 JSTエージェント

LinkedIn の自己進化するエージェント カスタマー サポート システム

エンタープライズ サポート エージェントは、ポリシー、製品機能、ナレッジ ベースが継続的に進化する急速に変化する環境で動作するため、静的なアシスタントは脆弱になり、維持コストが高くなります。 LinkedIn の自己進化型エージェント サポート システムを紹介します。このシステムは、検索拡張生成と進化型自動プロンプト、およびモジュール型の本番環境に合わせた評価フレームワークを統合し、基礎モデルを再トレーニングすることなく安全で継続的な改善を可能にします。システムは、プロンプト、取得、評価を、運用上のガードレールを備えた閉ループのバージョン管理されたワークフローとして処理します。オフライン シミュレーションとアブレーションでは、幻覚の減少や応答完全性の向上など、バニラ RAG やベースライン エージェントと比較して明らかな品質の向上が示されています。 LinkedIn の実稼働サポート トラフィックに対する 2 週間のユーザーによるランダム化 A/B テストでは、統合された自己進化型ワークフローにより、QA セルフサービスが 9.0 パーセント ポイント、キャンセル セルフサービスが 4.8 ポイント、ルーティング精度が 30.6 ポイント向上しました。これらの結果は、現実世界の企業環境におけるスケーラブルで自己進化する AI エージェントへの実用的な道筋を示しています。

原文 (English)

Self-evolving Agentic Customer Support System at LinkedIn

Enterprise support agents operate in rapidly changing environments where policies, product capabilities, and knowledge bases evolve continuously, making static assistants brittle and costly to maintain. We present LinkedIn's self-evolving agentic support system, which integrates retrieval-augmented generation with evolutionary auto-prompting and a modular, production-aligned evaluation framework to enable safe, continuous improvement without retraining foundation models. The system treats prompts, retrieval, and evaluation as a closed-loop, versioned workflow with operational guardrails. Offline simulations and ablations show clear quality gains over vanilla RAG and baseline agents, including reduced hallucinations and improved response completeness. In a two-week user-randomized A/B test on LinkedIn's production support traffic, the integrated self-evolved workflow increased QA self-serve by 9.0 percentage points, cancellation self-serve by 4.8 points, and routing accuracy by 30.6 points. These results demonstrate a practical path to scalable, self-evolving AI agents in real-world enterprise settings.

13:00 JST画像/動画生成

意思決定の境界を超えて: 対照的な埋め込み多様体に対する関係幾何学攻撃

対照学習とシャム埋め込みモデルは、現代の検証システムの基礎となっており、決定は離散的な分類境界ではなく、埋め込み空間の関係幾何学によって決定されます。しかし、既存の敵対的攻撃は依然として基本的に分類中心であり、リレーショナル ジオメトリの脆弱性を見落としています。この論文では、対照的なシステムに対する攻撃を多様体レベルのリレーショナル破損として再定式化する、ジオメトリを認識した敵対的攻撃フレームワークを紹介します。提案されたフレームワークは、個々の予測をターゲットとするのではなく、正のペアを押し離すと同時に負のペアを近づけることにより、埋め込み多様体内の類似性の組織を体系的に歪め、最終的にペアごとの類似性構造を崩壊および反転させます。スケーラブルな展開を可能にするために、反復的なオンライン最適化を学習前のオフラインの敵対的ジオメトリ変形に移行し、被害者モデルから一般化されたジオメトリ変形パターンを学習する軽量のフィードフォワード ジェネレーターをトレーニングします。このジェネレーターは、トレーニングが完了すると、オンライン勾配計算を必要とせずに、単一のフォワード パスを通じて敵対的な摂動を生成し、類似性に基づく検証システムに対するリアルタイムのオンライン攻撃を可能にします。複数の検証アーキテクチャにわたる実験結果は、深刻なマニホールドレベルのリレーショナル破損とともに、検証パフォーマンスの大幅な低下を示しています。 Markmatch 検証システムでは、提案された攻撃により精度が 95.4% から 38.6% に低下すると同時に、正負の類似性構造が完全に逆転します。

原文 (English)

Beyond Decision Boundaries: Relational Geometry Attacks on Contrastive Embedding Manifolds

Contrastive learning and Siamese embedding models have become the foundation of modern verification systems, where decisions are governed not by discrete classification boundaries, but by relational geometry in embedding space. However, existing adversarial attacks remain fundamentally classification-centric, overlooking the vulnerability of relational geometry. In this paper, we introduce a geometry-aware adversarial attack framework that reformulates attacks on contrastive systems as manifold-level relational corruption. Instead of targeting individual predictions, the proposed framework systematically distorts similarity organization within the embedding manifold by pushing positive pairs apart while simultaneously pulling negative pairs closer, ultimately collapsing and inverting pairwise similarity structure. To enable scalable deployment, we shift iterative online optimization into an offline adversarial geometry deformation prior learning stage and train a lightweight feed-forward generator that learns generalized geometry deformation patterns from the victim model. Once trained, the generator produces adversarial perturbations through a single forward pass without requiring online gradient computation, enabling real-time online attacks against similarity-based verification systems. Experimental results across multiple verification architectures demonstrate substantial degradation of verification performance together with severe manifold-level relational corruption. On the Markmatch verification system, the proposed attack reduces accuracy from 95.4% to 38.6% while completely reversing the positive-negative similarity structure.

13:00 JSTLLM/生成AI

検出を超えて: ライブのターンバイターン インタラクションにおける AI 生成のソーシャル エンジニアリングに対する防御 LLM の評価

生成 AI により、ソーシャル エンジニアリング攻撃がより流暢で、適応性があり、スケーラブルになり、進行中の対話中にユーザーを保護できる LLM ベースの防御者の必要性が高まっています。私たちは、そのような防御者がリスクの構造的原因を特定しているのか、それとも表面的な手がかりに反応しているだけなのかを尋ねます。私たちはトラストチェーンのローカリゼーションを形式化し、インタラクションがアクター権限、資産管理、検証の十分性、またはトランザクションパスで失敗するかどうかを特定します。私たちは、20 のシナリオ ファミリ、正当なケース、4 つの構造破壊モード、および 3 つの表面状態に及ぶ、管理された 300 ケースのオンライン住宅コーパスを構築します。 5 つの防御者モデルがステートフルなターンバイターンおよびワンショットの静的設定で同じコーパス上で評価され、プロトコルごとに 1,500 件、合計で 3,000 件のモデルケース評価が生成されます。明示的に安全でないコンプライアンスを示したモデルはありませんでしたが、防御効果は大きく異なり、介入率は 0% から 96.3% の範囲でした。保護措置と正しい構造的位置特定は切り離されることが多く、間違った信頼コンポーネントを特定する際にモデルが介入したり、保護措置を講じずに構造上の欠陥を認識したりすることがありました。資産管理の障害が位置特定の主要なボトルネックであり、表面感度はモデルによって異なり、ライブ静的な違いはモデルに依存していました。これらの調査結果は、安全そうに見える行動だけでは不十分であることを示しています。ライブ詐欺耐性は、介入、タイミング、構造的局在化、および誤検知動作を個別に測定する必要があります。

原文 (English)

Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction

Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users during ongoing interactions. We ask whether such defenders identify the structural source of risk or merely react to surface cues. We formalize trust-chain localization: identifying whether an interaction fails at actor authority, asset control, verification sufficiency, or transaction path. We construct a controlled 300-case online-housing corpus spanning 20 scenario families, legitimate cases, four structural failure modes, and three surface conditions. Five defender models are evaluated on the same corpus in state- ful turn-by-turn and one-shot static settings, yielding 1,500 model-case evaluations per protocol and 3,000 in total. No model produced explicit unsafe compliance, yet defensive effectiveness varied sharply: intervention rates ranged from 0% to 96.3%. Protective action and correct structural localization were frequently decoupled, with models sometimes intervening while identifying the wrong trust component or recognizing a structural failure without taking protective action. Asset-control failures were a major localization bottleneck, surface sensitivity varied across models, and live-static differences were model-dependent. These findings show that safe-looking behavior alone is insufficient; live scam resistance must separately measure intervention, timing, structural localization, and false-positive behavior.

13:00 JSTLLM/生成AIハードウェア/半導体Llama

言語モデルの隠れた状態を大規模に解釈する

レンズ メソッドは、中間アクティベーションを出力語彙にマッピングすることによって大規模言語モデル (LLM) を解釈し、ネットワークを通じて次のトークンの予測がどのように展開されるかを明らかにします。トレーニング済みのレンズは依然として高価です。アフィン変換パラメータはモデル幅に応じて二次関数的に増加しますが、正確で完全な語彙のカルバック ライブラー (KL) トレーニングが記憶を支配します。その結果、事前に訓練されたレンズは最大 20B パラメータのモデルに適用され、特定のコンポーネント タイプに関連付けられたままになります。 OmniLens は、残留ストリーム、アテンション、MLP のいずれであっても、単一のレンズ ファミリをモデル幅のアクティベーションに適用し、2 つの独立したスケーリング技術を組み合わせたものです。まず、低ランクのトランスレーターは、レンズごとのパラメーターの増加をモデル幅で線形にし、トレーニング可能なパラメーターを最大 98.4% 削減します。第 2 に、Subset-KL は選択された語彙ロジットのみを実体化します。Top-k モードはピーク トレーニング メモリを最大 70% 削減しますが、重要度をサンプリングしたバリアントは完全な KL に対して不偏の確率的勾配を保持します。これらの節約により、LLaMA-3.3-70B の 482 レンズの高密度アンサンブルが可能になり、同じ深さで残留ストリーム設計の 6 倍のカバレッジを提供します。次に、モデル全体をカバーすることで、単一コンポーネントのレンズではできないことが明らかになります。つまり、動作が最も目に見えるコンポーネントが、介入が最も効果的なコンポーネントである必要はなく、最も効果的な介入は、以前のレンズ研究で調査された注意の対象外にあります。 OmniLens は、3 つのケーススタディ (プロンプト インジェクション検出、マルチホップ メモリ インジェクション、毒性の局所化) にわたって、主要な公開結果を大幅に低コストで再現します。

原文 (English)

Interpreting Language Model Hidden States at Scale

Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator parameters grow quadratically with model width, while exact, full-vocabulary Kullback--Leibler (KL) training dominates memory. Consequently, prior trained lenses have been applied to models of at most 20B parameters and remain tied to particular component types. We present OmniLens, which applies a single lens family to any model-width activation, whether residual stream, attention, or MLP, and combines two independent scaling techniques. First, low-rank translators make per-lens parameter growth linear in model width and reduce trainable parameters by up to 98.4%. Second, Subset-KL materializes only selected vocabulary logits: its Top-k mode cuts peak training memory by up to 70%, while its importance-sampled variant retains unbiased stochastic gradients for the full KL. These savings enable a dense ensemble of 482 lenses for LLaMA-3.3-70B, providing 6x the coverage of a residual-stream design at the same depth. Model-wide coverage then reveals what single-component lenses cannot: the components where a behavior is most visible need not be those where intervention is most effective, and the most effective interventions lie outside the attention heads examined by prior lens studies. Across three case studies (prompt-injection detection, multi-hop memory injection, and toxicity localization), OmniLens reproduces key published results at substantially lower cost.

13:00 JST研究/論文

ロジット境界の幾何学的信念インターフェイスとスパース シーフ エンクレーブ プロトコル: 安全なネットワーク電子医療記録 (EHR) の相互運用性のための内蔵型基板

電子医療記録の相互運用性は境界問題です。レガシー システム、生成モデル、用語サービス、アイデンティティ システム、および人間のレビュー担当者はそれぞれ豊富な内部状態を公開する可能性がありますが、運用上の交換には、型指定されたクレーム、限定された不確実性、来歴、および明示的な承認または棄権の狭い共有インターフェイスが必要です。この文書では、そのインターフェイスの数学的および工学的アーキテクチャについて詳しく説明します。組織化されたアイデアはロジット境界です。発見モデルはローカルのカテゴリ的決定に対して事前しきい値スコアを提案する可能性がありますが、決定論的な判断基盤は、その提案が許容されるか、レビューが必要か、または高速ヘルスケア相互運用性リソース (FHIR) トランザクションが構築される前に隔離する必要があるかを決定します。結果として得られる幾何学的信念インターフェイス (GBI) は、有限境界セマンティクス、局所的なディリクレ証拠、セルラー層およびマッピング コーンの診断、助言幾何学的監査チャート、およびフェールクローズ展開のための分散暗号化層エンクレーブ (DCSE) プロトコル スケッチを組み合わせています。このフレームワークは、臨床上の真実、世界的な表現の整合性、またはエンドツーエンドの安全性を確立するものではありません。これは、モデルとシステムの境界での証明書生成チェックを定義します。対となる凍結合成ベンチマークである GBI BoundaryBench v0.1 は、3 つの証拠モード (768 回の正規実行) にわたる 256 個の保留タスクで Qwen3-4B-Instruct-2507 を評価しました。すべての実行は完了しましたが、ベンチマーク コントラクトで受け入れられる出力を生成したものはありませんでした。安全な解析中に 369 件、スキーマ検証中に 399 件が拒否され、カバレッジはゼロとなり、決定的な隔離が行われました。この経験的結果は意図的に狭く、1 つの凍結界面の下に 1 つの 4B オープンウェイト モデルを配置しており、LLM の能力や臨床安全性に関する一般的な主張としてではなく、入院境界に関する証拠として報告されています。 Julia の付録では、標準ライブラリを使用して数値証明書を検証します。

原文 (English)

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared interface of typed claims, bounded uncertainty, provenance, and explicit admission or abstention. This paper details a mathematical and engineering architecture for that interface. The organizing idea is the logit boundary: a discovery model may propose pre-threshold scores over a local categorical decision, but a deterministic judgment substrate decides whether the proposal is admissible, requires review, or must be quarantined before any Fast Healthcare Interoperability Resources (FHIR) transaction is constructed. The resulting Geometric Belief Interface (GBI) combines finite boundary semantics, local Dirichlet evidence, cellular-sheaf and mapping-cone diagnostics, advisory geometric audit charts, and a Decentralized Cryptographic Sheaf-Enclave (DCSE) protocol sketch for fail-closed deployment. The framework does not establish clinical truth, global representation alignment, or end-to-end safety; it defines certificate-producing checks at a model-to-system boundary. A companion frozen synthetic benchmark, GBI BoundaryBench v0.1, evaluated Qwen3-4B-Instruct-2507 on 256 held-out tasks across three evidence modes (768 canonical executions). All executions completed, but none produced an output accepted by the benchmark contract: 369 were rejected during safe parsing and 399 during schema validation, yielding zero coverage and deterministic quarantine. This empirical result is deliberately narrow - one 4B open-weight model under one frozen interface - and is reported as evidence about the admission boundary, not as a general claim about LLM capability or clinical safety. A Julia appendix verifies numerical certificates using standard libraries.

13:00 JSTビジネス/資金調達

Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures

Competitive artificial-life systems can rank trained controllers differently under training and ecological evaluation. We present Neuroevol…

13:00 JSTLLM/生成AI

AI 連携における価値理論に向けて

AI システムを人間の価値観に合わせることができるでしょうか?大規模言語モデル (LLM) とマルチモーダル基盤モデルの普及により、有毒な音声や幻覚から、許可されていないアクションを実行する AI エージェントに至るまで、被害が増加しています。 AI の安全性の分野では、これらの有害な事例は、整合性の問題、または人間の価値観と整合性の取れていないモデルとして扱われることがよくあります。研究者らは、応用的および理論的な AI 価値の調整の取り組みを追求することでこれに応えてきたが、多くの場合、人間の価値が何を意味するのかは明示されていない。 AI の価値調整の分野では、人間の価値観はどのように考えられるのでしょうか?これらの価値観は技術的にどのように運用され、評価されるのでしょうか?この分野から出現した価値理論は、AI の将来に何を意味するのでしょうか? AI における暗黙の価値理論を識別するために、94 の価値調整研究論文に注釈を付けました。大多数は価値観を定義しておらず、好みに大きく依存しており、文化的に複雑な概念を二者択一の選択肢に落とし込んでしまう危険性があります。研究者がモデルのトレーニングと評価にヒューマン・アノテーターの使用を省略し、代わりに合成データと自動評価者によるモデルの調整と評価のアプローチに目を向けているため、基礎モデルの値を争って制定するための代替方法を閉鎖してしまう可能性があることを私たちは認識しています。 AI の価値観の一致を哲学的コミットメントとして明確にすることで、AI が人間の価値観に対処できるかどうか、またどのように対処できるかについての議論に、大きな具体性とまだ探求されていない視点をもたらすことを目指しています。

原文 (English)

Toward a Theory of Value in AI Alignment

Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? We annotated 94 value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on preferences as a stand in that runs the risk of reducing complex culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and autorater approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. In making AI value alignments philosophical commitments explicit, we seek to bring great specificity and under explored perspectives in the debate on whether and how AI can address human values.

13:00 JSTエージェント

支援 AI エージェントの階層的構成性

AI エージェントは、さまざまなアプリケーションで人間を支援するためにますます開発されており、大規模言語モデルやその他のディープ ネットワーク アーキテクチャは、そのようなエージェントにとって最先端のものであると考えられています。これらの方法は優れた確率的予測子ですが、リソースを大量に消費し、不透明であり、基礎となる表現と処理の選択肢が狭いため、新しい状況では恣意的な決定を下すことが知られています。私たちの研究は、AI の初期の先駆者にまで遡ることができるものの、現代の AI 手法では十分に活用されていない中心的な原則に基づいて、そのような AI エージェントのアーキテクチャの設計を探ることを目指しています。この論文では、人間の参加者が参照するオブジェクトのあいまいさに対処する AI エージェントの中核問題の文脈でこれを行います。人間は、ドメインのコンテキストと他の人間の参加者の好みに関する構成的な知識をヒューリスティックに活用することで、このような曖昧さに対処します。この観察からインスピレーションを得て、階層的構成性の原理を組み込み、単純なヒューリスティックを使用して目的の曖昧性をなくすアーキテクチャを説明します。具体的には、ドメイン オブジェクトは、人間が検証した意味論的特徴規範から引き出されたプリミティブな属性、および支援エージェントと特定のユーザーとの対話の限られた観察履歴から自動的に識別される属性と概念の階層的な組み合わせの観点から表現されます。次に、支援エージェントは、この構成階層の知識に基づいて推論することによって、望ましい曖昧さを解消します。ドメインダイナミクスを支配する公理。意味的な互換性、セッションの顕著性、およびユーザー固有のテーマの好みのモデルがあり、必要に応じて人間による説明が求められます。実験によれば、私たちのアプローチは常に最先端のデータ主導ベースラインを上回り、特定のユーザー プロファイルへの適応をサポートしていることが示されています。

原文 (English)

Hierarchical Compositionality for An Assistive AI Agent

AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictors, but they are resource-hungry, opaque, and known to make arbitrary decisions in novel situations due to the narrow set of underlying representation and processing choices. Our work seeks to explore the design of architectures for such AI agents based on core principles that can be traced back to the early pioneers of AI but are not fully utilized in modern AI methods. We do so in this paper in the context of the core problem of AI agents addressing ambiguity in the objects being referred to by the human participants. Humans address such ambiguity by heuristically leveraging compositional knowledge of domain context and the preferences of the other human participants. Drawing inspiration from this observation, we describe an architecture that embeds the principle of hierarchical compositionality and uses simple heuristics to achieve the desired disambiguation. Specifically, domain objects are represented in terms of primitive attributes drawn from human-validated semantic feature norms, and a hierarchical combination of attributes and concepts automatically identified from a limited observed history of interactions of an assistive agent with specific users. The assistive agent then achieves the desired disambiguation by reasoning with knowledge of this compositional hierarchy; axioms governing domain dynamics; and models of semantic compatibility, session salience, and user-specific thematic preference, requesting human clarification when necessary. Experiments show that our approach consistently outperforms state of the art data-driven baselines, supporting adaptation to specific user profiles.

13:00 JSTエージェント研究/論文

AI 時代の栄養データ インフラストラクチャ: エージェント仲介研究のための FAIR の運用化

AI エージェントは栄養学の研究を加速できますが、その分析はアイデンティティ、セマンティクスを継承し、基礎となるデータの曖昧さを解放します。私たちは、自動化された使用のために FAIR を運用するソース保存インフラストラクチャである Nutrition Data Service (NDS) を紹介します。記述解決により、リリース固有のレコードを検索できるようになります。タイプ付き横断歩道は、独立してリリースされたリソースを接続します。機械可読インターフェイスは、バージョン管理されたソースとクロスウォークを公開し、AI エージェントによる分析を再生可能および監査可能にします。食品説明ベンチマークでは、NDS は高い精度を維持し、NutriBench で公開されている最高の言語モデル結果を上回りました。外部のブラインド クロスウォーク評価では、その型付きコントラクトが防御可能なリンクを優先し、サポートされていないマッピングを拒否することが示されています。個人レベルの血糖指数分析では、ピン留めされた NDS 入力はモデル間および反復実行間で同一の出力を生成しますが、オープンウェブ再構成は不安定なままです。中心的な結果は、エージェント媒介栄養研究には、データの識別、検索、およびクロスウォークのための新しいデータ インフラストラクチャが必要であるということです。

原文 (English)

Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research

AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We present Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR for automated use: description resolution makes release-specific records findable; typed crosswalks connect independently released resources; machine-readable interfaces expose versioned sources and crosswalks, making analyses by AI agents replayable and auditable. On food-description benchmarks, NDS shows strong held-out accuracy and outperforms the best published language-model result on NutriBench. External and blinded crosswalk evaluations show that its typed contract favors defensible links and rejects unsupported mappings. In a person-level glycemic-index analysis, pinned NDS inputs produce identical outputs across models and repeated runs, while open-web reconstruction remains unstable. The central result is that agent-mediated nutrition research requires a new data infrastructure for data identity, search, and crosswalk.

13:00 JSTLLM/生成AIエージェントClaude

DSAgentBench: エージェントは実際のコンピュータ環境でエンドツーエンドのデータ サイエンス ワークフローを自動化できますか?

現実世界のデータ サイエンスには、データ ラングリング、探索、モデリング、視覚化、検証に及ぶ長期的なワークフローが含まれており、実際の運用環境内でノートブック、IDE、端末、ブラウザ、データベースなどのツールを調整して使用する必要があります。しかし、既存のベンチマークには実際のコンピューターとの対話が欠けており、エージェントが現実的なコンピューティング環境で完全なエンドツーエンドのデータ サイエンス ワークフローを実行できるかどうかが評価されていないため、データ サイエンス実践の多段階、マルチツールの性質を捉えることができません。エージェントが実際のコンピューター環境内で完全なデータ サイエンス ワークフローを自動化できるかどうかを評価する最初のベンチマークである DSAgentBench を紹介します。 DSAgentBench には、データ サイエンスのライフサイクル全体をカバーする 275 の多様なタスクが含まれており、実際に必要な複雑さとツールの調整を反映しています。各タスクには、中間出力と調整されたツールの使用における基礎的な決定が必要であり、コードのみの実行ではなく、分析の正確さ、視覚的な出力、モデルのパフォーマンスを検証する決定論的評価機能が含まれています。 15 のクローズドおよびオープンソース モデルを使用した広範な実験では、最も強力なエージェントである Claude-4.6-Sonnet でさえ 56.70% のタスク成功率しか達成できず、すべてのオープンソース エージェントは 1% 未満に留まり、ツール オーケストレーション、OS グラウンディング、およびマルチステップ推論で頻繁に失敗することがわかりました。これらの結果は、現在のエージェント システムと実際のデータ サイエンス ワークフローとの間に大きな機能ギャップがあることを明らかにしており、DSAgentBench を根拠のある検証可能な自律的なデータ サイエンス エージェントを開発するための基盤として位置づけています。 DSAgentBench は https://github.com/vis-nlp/DSAgentBench でリリースされます。

原文 (English)

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments. Yet existing benchmarks lack real-computer interaction and do not evaluate whether agents can execute complete end-to-end data-science workflows in realistic computing environments, failing to capture the multi-stage, multi-tool nature of data-science practice. We introduce DSAgentBench, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments. DSAgentBench contains 275 diverse tasks covering the entire data-science life-cycle, reflecting the complexity and tool coordination required in practice. Each task requires grounding decisions in intermediate outputs and coordinated tool use, and includes a deterministic evaluator that verifies analytical correctness, visual outputs, and model performance rather than code-only execution. Our extensive experiments with 15 closed- and open-source models show that even the strongest agent, Claude-4.6-Sonnet, achieves only 56.70% task success, while all open-source agents remain below 1%, frequently failing at tool orchestration, OS grounding, and multi-step reasoning. These results reveal a substantial capability gap between current agentic systems and real data-science workflows, positioning DSAgentBench as a foundation for developing grounded, verifiable, autonomous data-science agents. We release DSAgentBench at https://github.com/vis-nlp/DSAgentBench.

13:00 JSTロボティクス

目に見えないところに隠された: 視覚・言語・行動モデルに対する拡散ベースの無制限のロボット攻撃

Vision-Language-Action (VLA) モデルは、さまざまな操作タスクにわたってロボットを制御する強力な機能を示しています。しかし、その敵対的な堅牢性は依然としてほとんど解明されておらず、この弱点を悪用すると物理世界に損害を与える可能性があります。 VLA モデルに対する既存の攻撃は、多くの場合、ピクセル空間の摂動やホワイトボックス アクセスに依存しており、その結果、顕著なアーティファクトが発生し、現実世界のロボット システムでの展開可能性が制限されます。この研究では、VLA モデルに対して視覚的に自然な敵対的パッチを生成する、拡散ベースの無制限のロボット攻撃である DURA を提案します。 DURA は、ホワイトボックス攻撃設定とブラックボックス攻撃設定の両方をサポートします。ブラックボックス設定では、被害者モデルの予測されたアクションのみが必要です。 DURA は、事前学習済みの拡散モデルの潜在軌道に沿って最適化することで、視覚的に自然なパッチを生成しながら、攻撃者が指定したターゲット アクションに向けてロボットを誘導します。シミュレーションと実際の物理世界の両方での広範な実験により、DURA が既存の方法よりも常に優れたパフォーマンスを発揮することが示されています。私たちの調査結果は、物理的に展開された VLA モデルの安全上のリスクを明らかにしており、より強力な防御が求められています。

原文 (English)

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for VLA models. DURA supports both white-box and black-box attack settings, where the black-box setting requires only the predicted actions of the victim model. By optimizing along the latent trajectory of a pretrained diffusion model, DURA generates visually natural patches while steering the robot toward attacker-specified target actions. Extensive experiments in both simulation and the real physical world show that DURA consistently outperforms existing methods. Our findings expose a safety risk for physically deployed VLA models and call for stronger defenses.

13:00 JSTエージェント

オンライン強化学習による安全な自動運転のための脅威に基づくポリシー認識シーンの混乱

強化学習(RL)は自動運転において有望なパフォーマンスを示していますが、安全性が重要な運転シーンへの露出が不十分であるため、オンライン RL ポリシーの安全性を確保することは依然として困難です。現実世界の交通状況のロングテールの性質により、従来のサンプリングでは危険でまれなインタラクションに遭遇することが困難になり、RL ポリシーが堅牢な安全行動を学習する能力が制限されます。既存の方法では、困難な場面や敵対的な状況を合成することでトレーニングの多様性を向上させています。ただし、これらのアプローチは通常、生成された摂動が現在のポリシーの弱点や学習ニーズにどのように関連するかを明示的にモデル化することなく、進化するポリシーとは別にシーン生成の目標を最適化します。この論文では、オンライン RL を使用した安全な自動運転のための、脅威に基づくポリシー認識シーン摂動 (TPSP) を提案します。 TPSP は、ポリシー対応シーン エンコーダを導入して、ポリシーの動作と周囲の環境の間の相互作用をキャプチャし、現在のポリシーに合わせたシーンの摂動を可能にします。この表現に基づいて、TPSP はシーン全体に均一な変更を適用するのではなく、重要なオブジェクトを選択的に摂動させます。さらに、元のシーンと混乱したシーンでのポリシー展開の間の脅威レベルの違いを通じて混乱したシーンを評価する、脅威に基づく最適化戦略を開発し、より高いトレーニング価値を持つ安全性が重要なシーンの生成を導きます。包括的な実験により、TPSP が安全学習効率を向上させ、約 400 万キロメートルの模擬走行データを使用して NAVSIM v2 で強力な安全性能を実現することが実証されました。アブレーション研究では、ポリシーを意識した対象を絞った摂動が、ランダムな戦略やポリシーを意識しない戦略よりも安全性が重視される有益なエクスペリエンスを提供し、限られたインタラクション予算の下でより安全な運転を可能にすることが検証されています。

原文 (English)

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challenging scenes or adversarial situations. However, these approaches typically optimize scene generation objectives separately from the evolving policy, without explicitly modeling how generated perturbations relate to the current policy's weaknesses and learning needs. In this paper, we propose Threat-guided Policy-aware Scene Perturbation (TPSP) for safe autonomous driving with online RL. TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy. Based on this representation, TPSP selectively perturbs critical objects rather than applying uniform modifications across the scene. Furthermore, we develop a threat-guided optimization strategy that evaluates perturbed scenes through threat-level differences between policy rollouts on original and perturbed scenes, guiding the generation of safety-critical scenes with higher training value. Comprehensive experiments demonstrate that TPSP improves safety learning efficiency, achieving strong safety performance on NAVSIM v2 with approximately 4 million kilometers of simulated driving data. Ablation studies verify that policy-aware targeted perturbations provide more informative safety-critical experiences than random or policy-unaware strategies, enabling safer driving under limited interaction budgets.

13:00 JST研究/論文

推論の近道と価値の対称性: 対称性が許可するもの、アーキテクチャが実現するもの、最適化が選択するもの

推論のショートカットは、意図しない概念を通じて正しい予測を生み出す神経象徴システムのルールの解決策です。竹村、井上、西野の最近のフレームワークは、値の再ラベル化の自己同型グループを通じてそれらを分析し、その中心的な未解決の質問として、ルールが概念を固定するのはいつなのかを尋ねます。まず、フレームワークの主要な定義である、すべての位置に適用される 1 つの共有順列が、評価対象となった 4 つの異種ベンチマークのいずれにも前述のように適用されないこと、およびドメインを共通のサイズにパディングする最も直接的な埋め込みが、確信を持って誤った病理を生み出すことを示します。CLE4EVR では、ソリューション ペアの 90.91% が説明されていないと報告され、導入した階層の明確に定義されたすべてのメンバーが 0% を報告し、パディングされた判定の内容が報告されています。構成ファイルの順序に従ってローテーションします。事前に指定された 15 の予測 (13 は確認済み) に基づいて 11 のルール ファミリを再測定し、説明のつかないペアの割合は 0% から 99.9999% に及び、証明可能な構造を追跡します。6 つの定理は、構文のみからカンディンスキーの病理を証明するフリー スロット補題を含め、推移性とその失敗に対する十分な条件を提供します。回路によって与えられたルールの場合、座標の対称性の不活性性の決定は coNP 完全です。非自明自己同型の存在は、ランダム化還元の下では coNP 困難であり、$\Sigma_2^p$ に存在し、PH が崩壊しない限り $\Sigma_2^p$ 完全ではなく、単調回路上では完全に coNP 完全です。ブール値の場合、推移性は正確に分類されます。解セットがアフィン剰余類であれば、自己同型性によりすべてが説明されます。弱教師モデルでは、観測された 94 個のショートカットすべてがコンポーネントごとの理論フラグの 1 つのレベルに配置され、推移的であると証明される 48 個のレベルには配置されません。 12 の型付きあいまいなレベルでは何も生成されず、許容される対称性と最適化によって選択されるものが分離され、デュアルヘッド コントロールが地理を複製します。すべての数値はリリースされたアーティファクトに遡ります。

原文 (English)

Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects

Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings and asks, as its central open question, when rules pin concepts down. We first show that the framework's key definition, one shared permutation applied at every position, does not apply as stated to any of the four heterogeneous benchmarks it was evaluated on, and that the most direct embedding, padding domains to a common size, produces confident false pathology: 90.91% of solution pairs reported unexplained on CLE4EVR, where every well-defined member of the hierarchy we introduce reports 0%, and the padded verdict's content rotates with configuration-file ordering. Re-measuring eleven rule families under fifteen pre-specified predictions (thirteen confirmed), unexplained-pair rates span 0% to 99.9999% and track provable structure: six theorems give sufficient conditions for transitivity and its failure, including a Free Slot Lemma certifying Kandinsky's pathology from syntax alone. For circuit-given rules, deciding symmetry-inertness of a coordinate is coNP-complete; nontrivial-automorphism existence is coNP-hard under randomized reductions, lies in $\Sigma_2^p$, is not $\Sigma_2^p$-complete unless PH collapses, and on monotone circuits is coNP-complete outright. In the Boolean case transitivity is classified exactly: automorphisms explain everything iff the solution set is an affine coset. Weakly supervised models place all 94 observed shortcuts at the one level the componentwise theory flags and none at the 48 it certifies transitive; twelve typed-ambiguous levels produce none, separating what symmetry permits from what optimization selects, and a dual-head control replicates the geography. All numbers trace to released artifacts.

13:00 JSTエージェント研究/論文

自動リサーチエージェントで無駄になったコンピューティングを回復する

最近の研究の多くは、研究問題をエンドツーエンドで解決するためのエージェントを開発しており、このパラダイムは自動リサーチと呼ばれることが増えています。このようなエージェントは、時間のかかる人間の作業を自動化し、特殊なアプリケーション向けに機械学習ソリューションをカスタマイズできる可能性を動機として、業界に大規模な投資を促してきました。このペーパーでは、これらの自動調査システムの中核となるモデリング パイプラインを研究し、表形式のデータセットに適用した場合の一般的な障害モードを特定します。(1) 同じバグを何度も解決するため、計算量が無駄になります。 (2) 多くの計算予算が残っている場合でも、ハイパーパラメータの調整に失敗することがよくあります。 (3) それらを強化するツリー検索アルゴリズムは探索されません。 (4) 彼らはデータ分析を実行し、データをトレーニングする人間を模倣しますが、その分析を下流の意思決定には使用しません。私たちは対象を絞った介入を調査し、検索ツリーのすべてのブランチにわたって発見されたランタイム制約を共有するグローバル デバッグ コンサルタント、プロンプトおよびコントロール レベルの機能強化、および洗練されたツリー検索アルゴリズムが無駄なコンピューティングの回復に成功していることを発見しました。私たちの結果は、基礎となる言語モデルを固定したまま、エージェントの設計だけで自動検索エージェントのパフォーマンスを大幅に向上できることを示しています。

原文 (English)

Recovering Wasted Compute in Autoresearch Agents

A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have inspired large industry investment, motivated by their potential to automate time-consuming human labor and customize machine learning solutions for specialized applications. In this paper, we study the modeling pipeline at the core of these autoresearch systems and identify common failure modes when they are applied to tabular datasets: (1) they waste compute resolving the same bugs over and over again; (2) they often fail to tune hyperparameters even when they have a large remaining compute budget; (3) the tree-search algorithms that power them do not explore; and (4) they perform data analysis, mimicking the humans whose data they are trained on, but do not use that analysis to make downstream decisions. We explore targeted interventions and find that a global debug consultant that shares discovered runtime constraints across all branches of the search tree, prompt- and control-level enhancements, and refined tree-search algorithms successfully recover wasted compute. Our results show that large gains in autoresearch agent performance are achievable through agentic design alone, holding the underlying language model fixed.

13:00 JST研究/論文

UAV 侵入検知のための会話型 vs ダッシュボード型説明可能 AI: オペレーターの信頼と信頼に関する実証研究

機械学習ベースの侵入検知システム (IDS) は、無人航空機 (UAV) ネットワークのセキュリティ保護において優れたパフォーマンスを実証しています。ただし、これらのモデルの「ブラックボックス」の性質と、マルチモーダルなサイバー物理データの高次元性が相まって、解釈可能性において重大な課題が生じています。静的可視化ダッシュボードでは、マルチモーダルなサイバー物理的特徴間の複雑な関係を、オペレーターが検査および解釈しやすい形式で表示するのが難しい場合があります。これに対処するために、オンデマンド調査を容易にするために、Large Language Model (LLM) を利用した会話型 XAI インターフェイスを提案します。参加者を対象とした対照実験で、インシデント後の監査タスク中のオペレーターの理解、信頼、信頼性に対する、この会話型インターフェイスと従来の XAI ダッシュボードの影響を体系的に評価しました。私たちの結果は、会話型インターフェイスの方がダッシュボードよりも便利であると認識されていたことを示唆しています。これは、参加者が関連情報に簡単にアクセスして統合できるためである可能性があります。ただし、この利点には適切な自立度の低下が伴い、過剰依存の潜在的なリスクが示されています。考えられる解釈の 1 つは、自然言語応答により AI アドバイスが受け入れやすくなり、IDS が間違っていた場合に参加者が基礎的な証拠を検証する傾向が減少した可能性があるということです。これらの調査結果は、UAV 侵入監査に対する人間と AI のコラボレーションにおける潜在的なトレードオフを示しています。つまり、知覚されるユーザビリティを向上させる対話メカニズムは、不適切な依存のリスクも高める可能性があります。最後に、適切な依存性を促進するために認知強制機能とのシームレスな相互作用のバランスをとる将来の XAI システムの設計への影響について議論します。

原文 (English)

Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. However, the 'black-box' nature of these models, combined with the high dimensionality of multimodal cyber-physical data, poses significant interpretability challenges. Static visualization dashboards may struggle to present complex relationships among multimodal cyber-physical features in a form that is easy for operators to inspect and interpret. To address this, we propose a Conversational XAI interface powered by Large Language Models (LLM) to facilitate on-demand investigation. In a controlled experiment with participants, we systematically evaluated the impact of this conversational interface versus a traditional XAI Dashboard on operator understanding, trust, and reliance during post-incident auditing tasks. Our results suggest that the conversational interface was perceived as more useful than the dashboard, potentially because it helped participants access and synthesize relevant information more easily. However, this benefit was accompanied by a lower level of appropriate self-reliance, indicating a potential risk of over-reliance. One possible interpretation is that the natural-language responses made the AI advice easier to accept, which may have reduced participants' tendency to verify the underlying evidence when the IDS was incorrect. These findings point to a potential trade-off in human-AI collaboration for UAV intrusion auditing: interaction mechanisms that improve perceived usability may also increase the risk of inappropriate reliance. We conclude by discussing design implications for future XAI systems that balance seamless interaction with cognitive forcing functions to foster appropriate reliance.

13:00 JST研究/論文

継続的インタラクションの拡散: 非同期ツール拡張推論のための拡散ネイティブ ランタイム

大規模な言語モデルは、最新の情報にアクセスし、計算を実行し、外部世界と対話するために外部ツールにますます依存しています。自己回帰モデルの場合、ツールの使用は生成プロセスに自然に適合します。モデルはツール呼び出しを発行し、結果を待ってから生成を続けます。ただし、拡散言語モデル (dLLM) は、出力の多くの部分を並行して繰り返し改良することによって推論するため、この停止と再開の対話パターンが不必要に制限的になります。モデルの推論が安定する前にツールの決定を強制し、個別の呼び出しが終了するまで有用な観察を遅らせ、冗長な改良とツールの実行を導入する可能性があり、タスクの精度と推論効率の両方に悪影響を与える可能性があります。ツールの対話を反復的なノイズ除去に統合するランタイム アーキテクチャである拡散ネイティブ モデルである連続インタラクション ディフュージョン (CID) を導入します。 CID は、モデル読み取り専用のファクト チャネル、型付き認知テンソルで表される思考チャネル、および表示チャネルを分離します。テキストまたは JSON 呼び出しが完全にシリアル化される前に情報ニーズが現​​れる可能性があり、ノイズ除去を継続しながら知覚バインディングが外部読み取りを開始できるようになります。返された結果は、進化する思考状態に投影され、以前の認識および表示領域を修正できます。永続的なバインディングは、外部で実行を繰り返すことなく静的な結果を再利用し、必要に応じてソースの変更を更新します。 CID は、証拠を早期に公開し、ツールのレイテンシーをモデルの計算と重複させ、重複する外部作業を削減し、新しい証拠が到着した後に有用な計算を保持するように設計されています。私たちはアーキテクチャ、ランタイム、トレーニングの目標を形式化し、タスクの品質とエンドツーエンドの効率性の評価プロトコルを定義します。この最初の論文は読み取り専用ツールに焦点を当てており、経験的なパフォーマンスの主張は行っていません。

原文 (English)

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. For autoregressive models, tool use naturally fits the generation process: the model emits a tool call, waits for the result, and then continues generating. Diffusion language models (dLLMs), however, reason by repeatedly refining many parts of their output in parallel, making this stop-and-resume interaction pattern unnecessarily restrictive. It can force tool decisions before the model's reasoning has stabilized, delay useful observations until a discrete call finishes, and introduce redundant refinement and tool execution, potentially hurting both task accuracy and inference efficiency. We introduce Continuous Interaction Diffusion (CID), a diffusion-native model--runtime architecture that integrates tool interaction into iterative denoising. CID separates a model-read-only fact channel, a thought channel represented by a Typed Cognitive Tensor, and a display channel. Information needs can emerge before a textual or JSON call is fully serialized, allowing perceptual bindings to launch external reads while denoising continues. Returned results are projected into the evolving thought state and can revise earlier cognition and display regions. Persistent bindings reuse static results without repeated external execution and refresh changing sources when needed. CID is designed to expose evidence earlier, overlap tool latency with model computation, reduce duplicate external work, and preserve useful computation after new evidence arrives. We formalize the architecture, runtime, and training objectives, and define an evaluation protocol for task quality and end-to-end efficiency. This first paper focuses on read-only tools and makes no empirical performance claims.

13:00 JST研究/論文

マルチモーダルな感情認識のための理論に基づく学習

会話におけるマルチモーダル感情認識 (MERC) では、言語的手がかりと非言語的手がかりの間の複雑な相互作用を理解する必要があります。しかし、既存のアプローチのほとんどは基本的にこれを直接入出力(マルチモーダルな手がかりと感情ラベル)のマッピング問題として扱い、人間が感情を解釈するときに使用する因果推論を無視しています。我々は、MERC を認知に触発された推論タスクに変換する新しいフレームワークである論理的誘導学習 (RGL) を提案します。二重プロセス理論に基づいて、私たちは感情的推論を 3 つの側面に分解します: 直観的 (即時的知覚、システム 1)、文脈的 (状況分析、システム 2)、および統合的 (両方の総合)。 MLLM オフラインを利用して構造化された根拠を生成し、内部表現を人間のような推論パターンに合わせてモデルのトレーニングをガイドするための記憶としてエンコードされます。最終的なモデルは、推論時に MLLM オーバーヘッドなしで動作します。実験結果は、RGL が IEMOCAP および MELD ベンチマークで最先端のパフォーマンスを達成することを示しています。さらに、解釈のために、モデルの内部機能が、目に見えないテストサンプルの意味的に正しい理論的根拠を効果的に取得することを実証し、その理論的推論能力を検証します。

原文 (English)

Rationale-Guided Learning for Multimodal Emotion Recognition

Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches fundamentally treat this as a direct input-output (multimodal cues-emotion labels) mapping problem, overlooking the causal reasoning that humans use when interpreting emotions. We propose rationale-guided learning (RGL), a novel framework that transforms MERC into a cognitively-inspired reasoning task. Based on dual-process theory, we decompose emotional reasoning into three facets: Intuitive (immediate perception, System 1), Contextual (situational analysis, System 2), and Integrative (synthesis of both). We leverage an MLLM offline to generate structured rationales, which are encoded as memories to guide model training via aligning internal representations with human-like reasoning patterns. Our final model operates without any MLLM overheads at inference time. Experimental results show that RGL achieves state-of-the-art performance on the IEMOCAP and MELD benchmarks. Further, for interpretation, we demonstrate that the model's internal features effectively retrieve semantically correct rationales for unseen test samples, validating its rationale reasoning capabilities.

13:00 JST研究/論文

混合状態プロトタイプを使用した量子増分学習

増分学習モデルは、パラメータとメモリの制約の下で動作しながら、致命的な忘れを起こすことなく新しいクラスを順番に学習する必要があります。ノイズの多い中間スケール量子 (NISQ) 時代では、量子ニューラル ネットワークは特徴マッピングに利点をもたらしますが、ハードウェアの制限により回路幅が制限されます。さらに、従来の量子分類器は直交基底状態の数によって制約され、増え続けるカテゴリに対応する能力が制限されます。したがって、トレーニング可能な混合状態プロトタイプに基づいた新しい量子増分学習フレームワークを導入します。そのオリジナルの設計では、共有量子バックボーンの回路幅を増やすのではなく、クラスのプロトタイプを追加することで新しいクラスを組み込んでいます。混合状態プロトタイプの使用は、単一の純粋状態プロトタイプよりも情報を表現する表現機能を備えているため、もう 1 つの重要な貢献です。また、分解可能な混合状態計算により、製造コストが削減され、分類に便利なヒルベルト・シュミット (HS) 距離計量が提供されます。シミュレーション結果は、私たちのモデルが最小数の量子ビットを使用して高次元の特徴集中を達成すると同時に、古典的なベースラインと比較して、増分学習タスクにおける計算複雑性の低下とロバストな表現を実証していることを示しています。

原文 (English)

Quantum Incremental Learning with Mixed State Prototypes

Incremental learning models are required to learn new classes sequentially without catastrophic forgetting, while operating under parameter and memory constraints. In the Noisy Intermediate-Scale Quantum (NISQ) era, although quantum neural networks offer advantages in feature mapping, hardware limitations restrict circuit width. Furthermore, traditional quantum classifiers are constrained by the number of orthogonal basis states, limiting their capacity to accommodate a continually growing number of categories. Thus, we introduce a novel quantum incremental learning framework based on trainable mixed-state prototypes. Its original design incorporates new classes by adding class prototypes rather than increasing the circuit width of the shared quantum backbone. The use of mixed-state prototypes is another key contribution, since they have representation capabilities to represent information than a single pure-state prototype. And the decomposable mixed-state calculation provides lower production costs and a convenient Hilbert-Schmidt (HS) distance metric for classification. Simulation results show that our model achieves high-dimensional feature concentration using a minimal number of qubits, while demonstrating lower computational complexity and robust representation in incremental learning tasks compared with classical baselines.

13:00 JSTLLM/生成AI

RLMOpt: 再帰的言語モデルによる適応プロンプトの最適化

プロンプト オプティマイザーは、言語モデルのパフォーマンスを向上させるプロンプトの検索を自動化しますが、既存の方法は、事前定義された最適化手順に依存しています。つまり、言語モデルがプロンプト提案を生成または調整している間に、アルゴリズムが探索する候補と検索の進行方法を決定します。 RLMOpt は、再帰言語モデル (RLM) を通じて検索ポリシー自体を言語モデル駆動にするプロンプト オプティマイザーです。 RLM エージェントはツールベースの環境上で動作し、タスク情報の検査、障害の分析、候補の生成、評価予算の割り当て、停止時期の決定を行います。決定論的ハーネスは、客観的なスコアリング、パレートベースの選択、および回帰制約を強制することでエージェントを補完します。私たちは、構造化臨床情報抽出 (Chia)、マルチホップ質問応答 (HotpotQA)、検証可能な指示フォロー (IFBench-2025)、およびマルチターン ツール呼び出しエージェント (BFCL) にわたる 4 つのベンチマークにわたって RLMOpt を評価します。単一シードでの一致した比較では、RLMOpt は 4 つのベンチマークすべてで最高のホールドアウト スコアを獲得し、4 つのタスクの平均をリードしています (GEPA の 0.610 対 0.589)。シード間で各ベンチマークを繰り返すと、11 件の一致するベンチマークとシードの比較が行われ、9 つのケースで RLMOpt が GEPA を上回りました。 11 回の実行全体で、シードを下回るプロンプトは一度も生成されませんでしたが、GEPA は開始点を 2 回下回りました。また、より効率的であり、GEPA によって生成されるプロンプトの 27 ~ 79% のサイズのプロンプトを生成しながら、より少ない検索ロールアウトでこれらの結果を達成します。さらに、私たちの結果は、最適化による利益は、検索バジェットではなく、主にシード プロンプトで利用可能なヘッドルームによって決定されることを示しています。したがって、効率的な最適化は、最小限の検索で確実に利用可能なヘッドルームに到達できるかどうかにかかっています。

原文 (English)

RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search

13:00 JSTLLM/生成AIエージェント

Evaluating Rational Contracting in Natural Language

The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or f…

13:00 JSTLLM/生成AI

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discover…

13:00 JSTLLM/生成AI

Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning

Double perovskites (DPs) offer broad compositional tunability, but predicting the space groups (SGs) of stable structures remains difficult…

13:00 JSTLLM/生成AI

学生の心の内側: LLM 学生シミュレータでの潜在的な推論と行動の共同モデリング

大規模言語モデル (LLM) ベースのシミュレーターは、多くの場合、観察可能なアクションを再現しますが、その背後にある根本的な推論を捉えることができません。教育現場では、個別指導システムの評価など、さまざまな用途に学生シミュレーションがますます使用されており、このギャップは特に顕著です。 2 人の学生がまったく異なる理由で同じ提出物を提出する場合があります。私たちは、LLM が学生のように行動するだけでなく学生のように考えるように微調整する学生モデリング フレームワークである INTERNAL STUDENT DIALOGUE (INSIDE) を紹介します。 INSIDE は、認知、感情、行動の各次元にわたるブルーム分類法に基づいた内部対話を生成し、ペアになった思考追跡と行動に基づいてモデルを微調整します。私たちはさまざまなプロンプトフレームワークに基づいてベースラインを設定し、シミュレートされたアクションの忠実度と生成された内部対話の品質という 2 つの軸で評価します。私たちの評価では、INSIDE がアクションの忠実度、実際の生徒の一致するコード生成、および推論の整合性の両方においてシミュレーションの忠実度を向上させ、モデル全体で最大 57.9% の最高の整合性を達成することが示されています。

原文 (English)

INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.

13:00 JSTエージェント

GeoForge: 地球観測推論のためのノンパラメトリック自己進化エージェント

地球観測 (EO) エージェントは、科学的に有効なツールのワークフローを構築し、現在の地理空間証拠に基づいて結論を導き出します。 EO ワークフローはセンシング セマンティクス、製品の依存関係、空間的および時間的な互換性、パラメーター要件によって制約されるため、これは困難です。既存のエージェントは多くの場合、クエリごとに広範な操作空間を検索しますが、最近の自己進化システムは、異種の EO 軌跡を、さまざまな意思決定レベルにわたって再利用可能な知識に完全に編成することはできません。この問題を解決するために、完成した軌跡を構造化されたノンパラメトリックな実行状態に変換する、トレーニング不要の自己進化フレームワークである GeoForge を紹介します。 GeoForge は、センシング コンテキストに従って操作空間を制約し、3 つの相補的なメモリからタスク条件付き事前を取得します。ワークフロー グラフ メモリはグローバルな操作順序をキャプチャし、アクション レベルのエクスペリエンスはローカルな修正を提供し、適応スキル標準操作手順は手順とデータの制約を保持します。取得された以前の情報はツールの実行をガイドしますが、現在の観察結果は最終的な答えの基礎となります。各タスクの後、セーフティゲート蒸留プロセスにより、将来の検索のために、接地された軌道が再利用可能な実行知識に変換されます。この実行、蒸留、再利用のループにより、バックボーン LLM を更新することなく計画が改善されます。複数の地理空間ベンチマークの実験では、GeoForge が、ほとんどの LLM のツール計画と推論エラーを大幅に削減しながら、さまざまな LLM バックボーンにわたってタスクの精度とツール使用軌跡の品質の両方を一貫して向上させることが実証されました。

原文 (English)

GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning

Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows are constrained by sensing semantics, product dependencies, spatial and temporal compatibility, and parameter requirements. Existing agents often search a broad operation space for each query, while recent self-evolving systems do not fully organize heterogeneous EO trajectories into reusable knowledge across different decision levels. To solve this problem, we present GeoForge, a training-free, self-evolving framework that transforms completed trajectories into a structured nonparametric execution state. GeoForge constrains the operation space according to the sensing context, then retrieves a task-conditioned prior from three complementary memories. Workflow Graph Memory captures global operation order, Action-Level Experiences provide local corrections, and the Adapted Skill Standard Operating Procedure preserves procedural and data constraints. The retrieved prior guides tool execution, while current observations remain the basis of the final answer. After each task, a safety-gated distillation process converts grounded trajectories into reusable execution knowledge for future retrieval. This execution, distillation, and reuse loop improves planning without updating the backbone LLM. Experiments on multiple geospatial benchmarks demonstrate that GeoForge consistently improves both task accuracy and tool-use trajectory quality across diverse LLM backbones, while substantially reducing tool-planning and reasoning errors for most LLMs.

13:00 JSTエージェント

障害のあるメモリから修正されたアクションまで: メモリ拡張エージェントの依存関係に基づくロールバック修復

永続メモリにより、言語モデル エージェントはセッション間で情報を再利用できますが、エラーの耐久性も高まります。汚染された、古い、または属性が間違っているレコードにより、推論、ツールの使用、回答、およびその後のメモリ書き込みが変更される可能性があります。既存の防御策は主に、疑わしい記憶を検出または削除するか、現在の対応を修正します。ソースを削除すると、すでに伝播されたクレーム、アクション、および派生メモリがアクティブなままになりますが、ストアをリセットしたり完全なトレースを再生すると、正常な状態が破壊され、不必要な計算が繰り返されます。したがって、\textbf{障害後のメモリ回復: } \textit{実行が失敗し、メモリに問題があると診断された場合、影響を受けない作業を保持しながら、回答と永続的な状態の両方を回復します。} \textbf{依存関係に基づくロールバック修復} は、実行時の来歴から型付きのメモリからアクションへのグラフを構築し、明示的な下流の依存関係を追跡し、独立した信頼できるサポートで候補を保持し、サポートされていないメモリ状態を非アクティブ化し、回答に関連するものだけを選択的に再生します。計算に影響を与えます。このアプローチを、3 つのツール使用ドメインと 4 つのメモリ障害タイプにわたる 150 ケースの制御ベンチマークと、LongMemEval-V2 から適応された 50 ケースの軌跡導出ストレス テストで評価します。制御されたベンチマークでは、競合する最良の回復方法の 77.3\% に対して 85.3\% の回復率を達成し、診断されたすべての欠陥のあるメモリを削除し、すべての正常なメモリを保存し、適度な LLM 呼び出しコストで選択的な再生のみを必要とします。適応されたサブセットでは、次善の方法の 54.0% に対して 68.0% の回収率に達し、最高のクレーム無効化 F1 (0.669 対 0.603) も達成しました。全体として、この結果はトレースの再構築が均一に優れていることを意味するものではありませんが、依存関係に基づくロールバック修復により、障害のあるメモリ状態を修復し、正常なメモリを保存しながら、強力な回復とコストのトレードオフが提供されることが示されています。

原文 (English)

From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents

Persistent memory lets language-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes. Existing defenses mainly detect or delete suspicious memories, or revise the current response. Deleting the source leaves already propagated claims, actions, and derived memories active, whereas resetting the store or replaying the full trace destroys benign state and repeats unnecessary computation. We therefore formulate \textbf{post-failure memory recovery: } \textit{given a failed execution and diagnosed faulty memories, recover both the answer and persistent state while retaining unaffected work.} Our \textbf{dependency-guided rollback repair} builds a typed memory-to-action graph from runtime provenance, traces explicit downstream dependencies, preserves candidates with independent trusted support, deactivates unsupported memory state, and selectively replays only answer-relevant affected computation. We evaluate this approach on a 150-case controlled benchmark spanning three tool-use domains and four memory failure types, and on a 50-case trajectory-derived stress test adapted from LongMemEval-V2. On the controlled benchmark, it achieves 85.3\% recovery versus 77.3\% for the best competing recovery method, removes all diagnosed faulty memories, preserves all benign memories, and requires only selective replay with modest LLM-call cost. On the adapted subset, it reaches 68.0\% recovery versus 54.0\% for the next best method, while also achieving the highest claim invalidation F1, 0.669 versus 0.603. Overall, the results do not imply uniformly better trace reconstruction, but show that dependency-guided rollback repair provides a strong recovery--cost trade-off while repairing faulty memory state and preserving benign memory.

13:00 JSTエージェント

MEGA: ウィズダム グラフによる自己進化型エージェント最適化インフラストラクチャ

コーディング エージェントが実装を処理することが増えるにつれ、中心的な課題は個々のエージェントの構築から、それらを体系的に改善するインフラストラクチャの構築へと移行しています。現在のアプローチは、移転可能な知識を蓄積することなくエージェント システムを最適化し、知識を構成論的に推論することなく蓄積し、運用上の証拠を通じてその知識が自己進化するメカニズムを欠いています。 MEGA (メタ評価に基づいた適応) は、自己進化するインフラストラクチャとしてこれらのギャップに対処します。各最適化サイクルは耐久性のある資産を生成し、それらの資産に対する構成的推論がその後の最適化を導き、運用上の証拠が蓄積された知恵とそれを支配する推論の両方を洗練します。レイヤー 1 は、行動パターンのクラスタリングと経験的な A/B 検証を通じてエージェント セッションから再利用可能な知恵を抽出し、各プロセスを耐久性のある資産に変換します。レイヤー 2 は、これらの資産を型付きウィズダム グラフ内のアトミック PCR (Primary-Context-Resultant) ユニットに分解し、演繹的、アブダクティブ、および帰納的推論を実行して暗黙的な関係を拡張します。次に、類似性だけを埋め込むだけでは到達できない橋渡し的な知識を明らかにする構成検索を通じて、コンテキスト固有の実行計画を組み立てます。レイヤ 3 は、異種エージェントのワークフロー (コード ノード、LLM 呼び出し、ツールを使用するエージェント) に対してマルチエージェントの協調的な最適化を実行し、データの差異を排除する制御された評価を通じて特定の戦略の変更に改善効果を帰属させます。レイヤー 3 からフィードバックされた証拠は、知恵の構成を管理するキュレーション戦略と、実行全体で蓄積された最適化の軌跡の両方の自己進化を促進します。その結果、エージェント システムの最適化と、最適化を導く知識の進化が 1 つの同じプロセスとなるインフラストラクチャが実現します。

原文 (English)

MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph

As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that systematically improves them. Current approaches optimize agent systems without accumulating transferable knowledge, accumulate knowledge without compositional reasoning over it, and lack a mechanism for that knowledge to self-evolve through operational evidence. MEGA (Meta Evaluation-Grounded Adaptation) addresses these gaps as a self-evolving infrastructure: each optimization cycle produces durable assets, compositional reasoning over those assets guides subsequent optimization, and operational evidence refines both the accumulated wisdom and the reasoning that governs it. Layer 1 distills reusable wisdom from agent sessions through behavioral-pattern clustering and empirical A/B validation, transforming each process into a durable asset. Layer 2 decomposes these assets into atomic PCR (Primary-Context-Resultant) units within a typed Wisdom Graph and performs deductive, abductive, and inductive reasoning to expand implicit relations; it then assembles context-specific execution plans through compositional retrieval that surfaces bridging knowledge unreachable by embedding similarity alone. Layer 3 performs multi-agent collaborative optimization over heterogeneous agent workflows (code nodes, LLM calls, and tool-using agents), attributing improvement effects to specific strategy changes through controlled evaluation that eliminates data variance. Evidence fed back from Layer 3 drives the self-evolution of both the curation strategies that govern wisdom composition and the optimization trajectories accumulated across runs. The result is an infrastructure in which optimizing an agent system and evolving the knowledge that guides optimization are one and the same process.

13:00 JSTLLM/生成AI画像/動画生成

RadFusion: 閾値制御可能な放射線医学レポートの生成に向けて

放射線科レポートの自動生成は、放射線科医の不足に対応して急速に進歩していますが、知覚モデルとは異なり、既存の生成モデルでは、診断内容の感度と特異性のトレードオフを制御できません。臨床シナリオが異なるため、このような管理は不可欠です。緊急トリアージでは感度を優先して見逃し所見を減らすのに対し、確認的解釈では特異性を重視して不必要な介入を制限します。単一の固定レポートでは、これらのシナリオに適応することも、規制クリアランスに広く期待されている ROC ベースの検証をサポートすることもできません。レポート生成にしきい値の制御機能を備えたフレームワークである RadFusion を紹介します。私たちの手法は、疾患ごとの信頼スコアを提供するマルチラベル分類器と、医学的所見を詳細に記述する VQA ベースのレポート生成機能を融合しています。次に、LLM はレポートを書き換え、そのレポートに記載されている診断が、ジェネレーターの説明に基づいたまま、選択されたしきい値での分類子の決定に従うようにします。 MIMIC-CXR では、RadFusion のパフォーマンスは分類子の ROC 曲線に準拠します。しきい値をスイープし、レポートをクラス ラベルにマッピングし直すことで、検証された分類子の ROC パフォーマンスが再現されます。この適合性により、生成されたレポートは ROC 分析を通じて定量的に評価できるようになり、規制クリアランスの根拠が強化され、レポートの動作を臨床状況に一致させる操作点の選択が可能になります。さらに、2 つのモデル タイプを組み合わせることで、制御されていない生成に比べて診断精度が向上します。一致した特異度で感度が 6.9%、一致した感度で特異度が 20.7% 増加しました。これらの結果は、RadFusion によってレポート生成が臨床的に適応可能で、定量的に検証可能で、診断の信頼性が向上することを示しています。

原文 (English)

RadFusion: Towards Threshold-Controllable Radiology Report Generation

Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions. A single fixed report can neither adapt to these scenarios nor support the ROC-based validation widely expected for regulatory clearance. We introduce RadFusion, a framework that equips report generation with threshold controllability. Our method fuses a multi-label classifier, which provides per-disease confidence scores, with a VQA-based report generator, which describes medical findings in detail; an LLM then rewrites the report so that its stated diagnoses follow the classifier's decisions at the selected threshold while staying grounded in the generator's descriptions. On MIMIC-CXR, the performance of RadFusion conforms to the classifier's ROC curve: sweeping the threshold and mapping the reports back to class labels reproduces the classifier's validated ROC performance. This conformance makes generated reports quantitatively evaluable through ROC analysis, strengthening the case for regulatory clearance, and enables operating-point selection that matches report behavior to clinical context. Moreover, combining the two model types improves diagnostic accuracy over uncontrolled generation: sensitivity increases by 6.9% at matched specificity, and specificity by 20.7% at matched sensitivity. These results show that RadFusion makes report generation clinically adaptable, quantitatively verifiable, and diagnostically more reliable.

13:00 JSTエージェント

MAP-Graph: マルチエージェント ワークフローのための出所を意識した共有メモリ

共有メモリは、言語モデルのエージェントが長いワークフロー全体で情報を再利用するのに役立ちますが、特定のエージェントまたはアクションについては関連する証拠が許容されない場合があります。制限は派生を通じて伝播するため、要約ではプライベート、汚染された、信頼できない、または取り消されたソースが隠蔽され、不正な読み取りや安全でないアクションが可能になる可能性があります。既存のアプローチは、セマンティック検索、範囲指定されたアクセス、または系統追跡を提供しますが、段階的な信頼からハード承認を明確に分離したり、証拠要件をアクションのリスクに適応させたりすることはできません。エージェント、ソース、メモリ、クレーム、およびアクションを型付き実行グラフで表現する、来歴を意識したメモリ レイヤーである MAP-Graph を導入します。祖先を追跡し、権限のないレコードを除外し、セマンティックな類似性と乗算パスの信頼性によって適格なメモリを再ランク付けし、監査のために影響を受ける系統を保持しながら、アクションの実行前にリスクに敏感なゲートを適用します。 3 つのドメインにわたるメソッドごとに 2,700 の合成タスクの制御されたベンチマークで、MAP-Graph は、タスク全体の成功率 94.96\%、正確な決定精度 72.70\%、および成功には安全な介入ではなく正しい \textsc{Allow} が必要なクリーン設定で 90.22\% を達成しました。アブレーションは、パーミッション フィルタリング、パスの信頼性、およびアクション ゲーティングの役割を分離しますが、2 つの追加バックボーンを使用した転送テストでは、正確な決定とアクセス制御の利点が維持されます。これらの結果は、評価された設定内で、単なる事後監査メタデータではなく、運用制御信号としての出所をサポートします。

原文 (English)

MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows

Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or action. Because restrictions propagate through derivations, summaries can conceal private, poisoned, untrusted, or revoked sources, enabling unauthorized reads or unsafe actions. Existing approaches provide semantic retrieval, scoped access, or lineage tracking, but do not clearly separate hard authorization from graded trust or adapt evidence requirements to action risk. We introduce MAP-Graph, a provenance-aware memory layer that represents agents, sources, memories, claims, and actions in a typed execution graph. It traces ancestry, excludes permission-ineligible records, reranks eligible memories by semantic similarity and multiplicative path trust, and applies a risk-sensitive gate before action execution while retaining affected lineage for audit. On a controlled benchmark of 2,700 synthetic tasks per method across three domains, MAP-Graph achieves 94.96\% overall task success, 72.70\% exact decision accuracy, and 90.22\% in the clean setting, where success requires a correct \textsc{Allow} rather than a safe intervention. Ablations isolate the roles of permission filtering, path trust, and action gating, while transfer tests with two additional backbones preserve the exact-decision and access-control advantages. These results support provenance as an operational control signal, rather than only post-hoc audit metadata, within the evaluated setting.

13:00 JSTLLM/生成AILlamaDeepSeek

非局所性による SAE 特徴の意味的抽象性の測定

スパース オートエンコーダ (SAE) は、対応するタスクに関連し、因果的に効果的な機能を理解することで、推論、ジェイルブレイクなどの LLM 動作のメカニズムの説明を明らかにするのに役立ちました。このようなメカニズムの説明を評価するには、下流の研究で表面の語彙特徴を真に高レベルの語彙特徴から区別する必要があります。ただし、autointerp ベースのセマンティック記述も因果的ステアリング ユーティリティも、機能の抽象化レベルを完全に解決することはできません。この目的のために、SAE 特徴の活性化に対する正規化された位置ごとの影響のエントロピーとして定義される \emph{特徴非局所性} (FNL) を導入します。我々は、FNL が特徴の意味的抽象性に関する既存の LLM ベースのプロキシ メトリックと相関し、コンテキスト依存の推論特徴とトークン駆動の特徴をうまく区別し、1 つのコンテキスト レベルの特徴と 1 つのトークン レベルの特徴で構成されるランダムに抽出されたペアの $73$--$84\%$ のコンテキスト特徴に上位の FNL を正しく割り当てることを報告します。 2 つのダウンストリーム アプリケーションを示します。私たちはジェイルブレイクの軽減に使用される SAE ベースの機能を監査しましたが、驚くべきことに、最も効果的な機能は有害な意図を真に認識するのではなく、FNL の低い位置機能であることがわかりました。 DeepSeek-R1-Distill-Llama-8B の高 FNL 機能をステアリングすると、ステアリングされていないモデルに比べて MATH-500 の精度が $4.6$ ポイント向上し、ゲインはモデル固有ですが、低 FNL 機能のステアリングよりも優れていることを報告します。 FNL は、メカニズムの説明の評価や下流介入のための特徴の選択に応用できる、SAE 特徴の抽象化レベルの LLM に依存しないラベルフリーの相関証人を提供すると結論付けています。

原文 (English)

Measuring Semantic Abstractness of SAE Features via Nonlocality

Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce \emph{Feature Nonlocality} (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. We report that FNL correlates with existing LLM-based proxy metrics of feature semantic abstractness, and successfully distinguishes context-dependent reasoning features from token-driven ones, correctly assigning the higher FNL to the contextual feature in $73$--$84\%$ of randomly drawn pairs that consist of one contextual and one token-level feature. We demonstrate two downstream applications. We audit SAE-based features used for jailbreak mitigation and find surprisingly that most effective features are positional features with low FNL rather than genuinely recognizing harmful intents. We report that steering high-FNL features in DeepSeek-R1-Distill-Llama-8B improves MATH-500 accuracy by $4.6$ points over the unsteered model and outperforms steering low-FNL features, though the gains are model-specific. We conclude that FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature, with applications in evaluating mechanistic explanations as well as selecting features for downstream interventions.

13:00 JSTエージェント

SKILLER: 小さな言語モデルで再利用可能なスキルを抽出するための言語レベルの強化学習

エージェント スキルは、手順に関する知識とドメインの専門知識をパッケージ化するための標準化された形式を表し、エージェント ハーネス システム内で、反復可能で高品質なタスク実行のために言語モデルの動作空間を継続的に制約する重要なメカニズムとして機能します。ただし、強力なクローズドソース モデルには高い推論コストがかかるため、Codex や OpenClaw などの現在人気のあるエージェント ハーネスは、現実世界のタスクを実行するためにこれらのスキルを導入する場合、依然として法外に高価です。コンシューマーグレードの GPU に展開可能なオープンソース モデルの急速な機能強化は、スキルベースの動作制約を活用することで、これらのコストを大幅に削減する魅力的な機会をもたらします。それにもかかわらず、このようなコンパクトなモデルに特化して調整された効果的なスキルを自動的に生成することは、依然として大きな現実的な課題です。これに対処するために、我々は、小規模モデルの実行者固有のスキルを自動的に生成するように設計された自然言語駆動の強化学習フレームワークである SKILLER を提案します。これは、アクターおよび批評家として強力なモデルを採用し、小規模モデルのエージェント システムを環境として扱い、すべての強化学習信号を完全に自然言語経由で伝播します。 Qwen3.5-9B および Qwen3.5-4B を使用した 5 つの関連ベンチマークにわたる広範な実験評価では、SKILLER が 3 つのオープンソースおよび 1 つのクローズドソースのスキル生成または進化メソッドよりも優れたパフォーマンスを示し、9B モデルで 4.3 ~ 20.4 パーセント ポイント、4B モデルで 1.8 ~ 13.3 パーセントの範囲の絶対ゲインを達成し、同時に単一スキル タスクにおける強力なクローズドソース モデルのパフォーマンスに著しく匹敵することが実証されました。スキルベンチ。プロジェクトは https://github.com/DANG-ai/SKILLER で入手できます。

原文 (English)

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks. The rapid capability enhancement of open-source models deployable on consumer-grade GPUs presents a compelling opportunity to drastically reduce these costs by leveraging skill-based behavioral constraints. Nevertheless, automatically generating effective skills tailored specifically for such compact models remains a significant practical challenge. To address this, we propose SKILLER, a natural-language-driven reinforcement learning framework designed to automatically generate executor-specific skills for small models, which employs a strong model as the actor and critic, treats the small-model agent system as the environment, and propagates all reinforcement learning signals entirely via natural language. Extensive experimental evaluations across five relevant benchmarks using Qwen3.5-9B and Qwen3.5-4B demonstrate that SKILLER outperforms three open-source and one closed-source skill generation or evolution methods, achieving absolute gains ranging from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model, while remarkably matching the performance of strong closed-source models on single-skill tasks in SkillsBench. The project is available at https://github.com/DANG-ai/SKILLER.

13:00 JST研究/論文

強化学習ベースのレーザー切断機パラメータの最適化

光学フィルムのレーザーベースの切断で高精度を達成するには、各フィルムタイプの特定の特性に応じて調整された焦点距離やレーザー出力ビームなどのパラメーターを注意深く調整する必要があります。さまざまなフィルムに最適な切断パラメータを見つけるために、試行錯誤に基づいた従来の方法が使用されますが、時間がかかり、不正確です。この問題に対処するために、この論文では、レーザー切断のための強化学習 (RL$^{2}$C) アルゴリズムを紹介します。このアルゴリズムは、イプシロン貪欲ポリシーを備えた Q 学習を使用して切断パラメーターを動的に最適化し、テーパー サイズとフィルムの無駄を大幅に削減します。さらに、RL$^{2}$C には動的な環境空間適応メカニズムが組み込まれており、複数の実験バッチにわたる学習プロセス中に遭遇する新しい状態に適応できるようになります。実験結果は、RL$^{2}$C が、RL ベースのさまざまな最適化手法と比較して、最適な切削パラメータを見つけるために必要なステップと時間が少ないことを示しています。具体的には、RL$^{2}$C は、既存の方法と比較して、最適化ステップ数を最大 12.5\% 削減し、処理時間を最大 81.8\% 削減します。この研究は、カット品質を向上させ、時間とフィルムの無駄を削減し、手動介入を最小限に抑えることにより、工業用レーザー切断プロセスにおける RL の可能性を実証しています。

原文 (English)

Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization

Achieving high accuracy in laser-based cutting of optical films requires careful tuning of parameters such as focal length and laser power beam, adjusted according to the specific properties of each film type. Trial-and-error based traditional methods are used to find the most suitable cutting parameters for various films, but they are slow and inaccurate. To address this issue, this paper presents the Reinforcement Learning for Laser Cutting (RL$^{2}$C) algorithm, which uses Q-learning with an epsilon-greedy policy to dynamically optimize cutting parameters, significantly reducing taper size and film wastage. Additionally, RL$^{2}$C incorporates a dynamic environment space adaptability mechanism to allow it to adapt to new states encountered during the learning process over multiple batches of experiments. Experimental results demonstrate that RL$^{2}$C requires fewer steps and less time to find optimal cutting parameters compared to various RL-based optimization methods. Specifically, RL$^{2}$C reduces the number of optimization steps by up to 12.5\% and processing time by up to 81.8\% compared to existing methods. This study demonstrates the potential of RL in industrial laser-cutting processes by improving cut quality, reducing time and film wastage, and minimizing manual interventions.

13:00 JSTLLM/生成AI研究/論文

DashArena: インタラクティブな分析ダッシュボード生成に関する LLM のベンチマーク

分析ダッシュボードは、データの探索と意思決定のために、調整されたビューと対話を組み合わせます。最近のモデルはデータと自然言語の目標からそれらを生成できますが、その有用性を評価することは依然として困難です。ダッシュボードの生成には制限がなく、静的な外観や実行の成功だけでは、分析サポートとインタラクションの品質を捉えることはできません。私たちの知る限り、インタラクティブな分析ダッシュボードのオープンエンドでタスクベースの生成のための最初のベンチマークである DashArena を紹介します。その主な革新は、各システムにダッシュボードと再生可能なインタラクション軌跡の両方を生成するよう要求することです。ブラウザーの実行プログラムは軌跡を再生し、システムの意図した分析ワークフローを再現可能な視覚的証拠と実行証拠に変えます。 VLM の裁判官はこの証拠を使用して候補者を比較し、ブラッドリーとテリーの集計によりリーダーボードが作成されます。ジャッジをさらに抽出して、オープンウェイトの DashJudge-8B を完成させました。人間による評価は、DashJudge-8B が人間の判断を効果的に再現することを示し、アブレーションは相互作用の証拠が裁判官の合意を改善することを示しています。フロンティア モデルを使った実験では、永続的なレンダリング、分析、およびインタラクションの失敗が明らかになりました。これらの結果を総合すると、現実的なダッシュボードの生成が依然として困難であること、およびインタラクションを意識した評価により、静的チェックまたは実行のみのチェックでは見逃される失敗が捕捉されることがわかります。

原文 (English)

DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation

Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and natural-language goals, but evaluating their usefulness remains difficult. Dashboard generation is open-ended, and neither static appearance nor successful execution alone captures analytical support and interaction quality. We introduce DashArena, to our knowledge the first benchmark for open-ended, task-grounded generation of interactive analytic dashboards. Its key innovation is to require each system to generate both a dashboard and a replayable interaction trajectory. A browser executor replays the trajectory and turns the system's intended analytical workflow into reproducible visual and execution evidence. A VLM judge compares candidates using this evidence, and Bradley--Terry aggregation produces the leaderboard. We further distill the judge into the open-weight DashJudge-8B. Human evaluations show that DashJudge-8B effectively reproduces human judgments and ablations show that interaction evidence improves judge agreement. Experiments with frontier models reveal persistent rendering, analytical, and interaction failures. Together, these results show that realistic dashboard generation remains challenging and that interaction-aware evaluation captures failures missed by static or execution-only checks.

13:00 JSTエージェント

エージェントによる命令データの選択: DataMaster に意図を解釈させます

既存の命令データ選択方法ではさまざまなメトリクスが導入されていますが、現実世界のデータセットに固有の複雑さがあるため、単一のメトリクスをすべてのシナリオにわたって一般化することは非現実的です。そのため、開発者は多くの場合、手動でデータを検査し、新しいアプリケーションごとにヒューリスティック ルールを作成する必要があり、これは退屈でエラーが発生しやすいプロセスです。このペーパーでは、ユーザーの意図を解釈し、最適な選択戦略を自律的に構築する命令データ選択エージェント (DataMaster) を介した、手動構成から自動オーケストレーションへのパラダイム シフトを提案します。 DataMaster は、ユーザーが自然言語記述を通じてデータのニーズを指定できるようにすることで、データのキュレーションを簡素化し、手動による戦略設計の負担を軽減します。数学、医療、コードの各領域にわたる広範な実験により、DataMaster がほとんどの設定で静的ベースラインを上回り、かなりの数のケースでフルプール トレーニングを上回ることが示されました。 DataMaster の実装と、報告されたパイプラインを再現するために必要なスクリプトは、https://github.com/nju-websoft/DataMaster で公開されています。

原文 (English)

Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent

Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus often forced to manually inspect data and craft heuristic rules for each new application---a tedious and error-prone process. In this paper, we propose a paradigm shift from manual configuration to automated orchestration via the Instruction Data Selection Agent (DataMaster), which interprets user intent and autonomously composes optimal selection strategies. By allowing users to specify data needs through natural language descriptions, DataMaster simplifies data curation and removes the burden of manual strategy design. Extensive experiments across the math, medical, and code domains show that DataMaster outperforms static baselines in most settings and surpasses full-pool training in a substantial number of cases. The implementation of DataMaster and the scripts needed to reproduce the reported pipeline are publicly available at https://github.com/nju-websoft/DataMaster.

13:00 JSTビジネス/資金調達

HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment

Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing…

13:00 JST研究/論文

接続する前にキュレートする: 本番ナレッジ グラフでの ID とオントロジーのタグ付け

抽出により、候補となるエンティティと関係が生成されます。それらをグラフに書き込むことで同一性が決定されますが、抽出エラーとは異なり、同一性の決定は破壊的です。間違ったタイプは後で修正できますが、1 つの ID の下でマージされた 2 つのレコードは、プロパティが結合されると分離できず、マージによってエラーが残ることはありません。このペーパーでは、検証済みの抽出ストリームを、98,795 の政府文書から抽出された 537,157 のエンティティと 2,198,567 の関係のナレッジ グラフに変換する、取り込みおよびオントロジーのタグ付けレイヤーについて説明します。名前の類似性ではなく、識別子列、名前列、表示名、および型スコープの位置から同一性を判断するレコード ID ラダーについて説明します。ラダーは解析されたテーブル内の重複排除を制御しますが、グラフの書き込みではより粗い正規名キーが適用されるため、正規名を共有するレコードは完全に等価に自動的にマージされます。私たちは、これが自動化ラインに属することを実証するのではなく主張します。ID ベンチマークは報告されておらず、主要な許可のオーバーマージは構築によって検出できません。このポリシーでは、エンティティの解決では候補者のみにフラグを立てるというもので、1 つの名前の 2 つの表面形式がマージされ、正しいレコードが破損し、無関係な文書から 8 つのエンティティが削除されるというインシデントが発生しました。次に、マルチクラス オントロジーのタグ付けと、予期していなかった証拠の非対称性について説明します。エンティティ名は型アサーションではなくインスタンス ラベルであるため、名前のフラグメントをクラス インデックスと照合することで分類が生み出されます。固定証拠の要求により、濃縮サンプルの役割割り当てが 36 から 4 にカットされ、すべて正しいことが確認されました。グラフの適合性負債を定量化し、親が間違っている一次クラスを補う二次分類を示し、775 件の人間による決定に対して 48,403 件の保留中の提案にまで成長したキュレーション キューについて説明します。

原文 (English)

Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph

Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions are destructive in a way extraction errors are not. A wrong type can be corrected later, but two records merged under one identity cannot be separated once their properties have been combined, and the merge leaves no error behind. This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537,157 entities and 2,198,567 relationships drawn from 98,795 government documents. We describe a record-identity ladder that decides sameness from identifier columns, name columns, display names and type-scoped position rather than from name similarity. The ladder governs de-duplication within parsed tables, while the graph write applies a coarser canonical-name key, so records sharing a canonical name merge automatically on exact equality. We argue rather than demonstrate that this is where the automation line belongs: no identity benchmark is reported, and the over-merges the key permits are undetectable by construction. That policy, under which entity resolution only ever flags candidates, followed an incident in which two surface forms of one name were merged, corrupting a correct record and deleting eight entities from an unrelated document. We then describe multi-class ontology tagging and an evidence asymmetry we did not anticipate: an entity name is an instance label rather than a type assertion, so matching name fragments against a class index invents classifications. Requiring anchored evidence cut role assignments on an enriched sample from 36 to 4, all confirmed correct. We quantify the graph's conformance debt, show secondary classifications compensating for a mis-parented primary class, and describe a curation queue grown to 48,403 pending proposals against 775 human decisions.

13:00 JST研究/論文

証拠のある組み合わせ最適化のための信念関数の意思決定を意識した近似

質量関数の焦点要素の数を減らすことは、古典的に、Jaccard や Jousselme などの固有の距離によって駆動され、証拠として近似をオリジナルに近づけます。代わりに、質量関数が証拠コストを伴う線形組み合わせ最適化問題を与える場合を考えます。その場合、保存すべきものは 2 つの質量関数の近さではなく、それらが引き起こす決定の質です。決定の後悔を対象とした決定を意識した近似を導入します。つまり、安価な近似で決定し、真の質量関数に基づいて評価されます。最小の最短パスでは、距離最適近似によって決定が反転されますが、決定を意識したマージによって決定が保持されます。これはランダム インスタンスの無視できない部分で発生します。私たちは、真の最適値でリグレスを局所化する 1 点限界を証明し、それをスカラーの場合の正確な動的プログラムに変換し、最終コストが判明する前に焦点要素を取り除くオンライン バージョンに拡張します。実験では、線形基準と非線形プロキシ読み出しの両方について、決定認識型圧縮器は表現認識型圧縮よりも決定を反転する頻度が低くなります。

原文 (English)

Decision-Aware Approximation of Belief Functions for Evidential Combinatorial Optimization

Reducing the number of focal elements of a mass function is classically driven by an intrinsic distance, such as Jaccard or Jousselme, that keeps the approximation close to the original as a body of evidence. We consider instead the case where the mass function feeds a linear combinatorial optimisation problem with evidential costs. What should then be preserved is not the closeness of the two mass functions, but the quality of the decision they induce. We introduce a decision-aware approximation that targets the regret of the decision: one decides with the cheaper approximation and is evaluated under the true mass function. On a minimal shortest path, the distance-optimal approximation flips the decision while a decision-aware merge preserves it, and this occurs on a non-negligible fraction of random instances. We prove a one-point bound that localises the regret at the true optimum, turn it into an exact dynamic program for the scalar case, and extend it to an online version that prunes focal elements before the final cost is known. In experiments the decision-aware compressor flips the decision less often than representation-aware compression, for both the linear criterion and a non-linear proxy read-out.

13:00 JSTエージェント

相対的因果知識の運用化: 共有された結果に関する非公開レポートからのバックボーンの特定可能性

因果関係知識の相対性理論 (RCK) は、異なる構造的因果モデルを持つエージェントのネットワークが、共有された介入的に一貫した抽象概念、つまりバックボーンを通じてどのように因果関係の知識を交換できるかを説明します。私たちは、この転送メカニズムが前提としている事前の同定に関する質問をします。つまり、そのバックボーンはいつエージェントのプライベートな因果関係の知識によって決定されるのでしょうか?基本的な 2 つのエージェントによる共通効果のケースでは、2 つの私的な原因が 1 つの共有された結果に影響を与え、各エージェントは自分の観点に関連する単一原因の因果的限界のみを識別します。標準的な互換性、非縮退、および局所的な重複の仮定の下では、これらの局所的な因果関係の限界は固有のバックボーンを識別しないことを示します。無限に多くの共同介入カーネルが、共同介入について意見が異なる一方で、まったく同じプライベート レポートを誘導する可能性があります。次に、条件付き回復の結果を示します。加法的分離性は隠れた相互作用の自由度を取り除きますが、観察上の残差の要約は依然として不十分です。エージェントが因果的に識別された応答機能を通信する場合、識別が可能になります。教育の付加価値の例は、これがなぜ最初はコミュニケーションの問題であり、次に政策構成の問題になるのかを示しています。

原文 (English)

Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome

The Relativity of Causal Knowledge (RCK) explains how a network of agents with different structural causal models can exchange causal knowledge through a shared interventionally consistent abstraction, or backbone. We ask the prior identification question that this transport mechanism presupposes: when is that backbone determined by the agents' private causal knowledge? In the basic two-agent common-effect case, two private causes influence one shared outcome and each agent identifies only the single-cause causal marginal relevant to its own perspective. We show that, under standard compatibility, non-degeneracy, and local overlap assumptions, those local causal marginals do not identify a unique backbone. Infinitely many joint intervention kernels can induce exactly the same private reports while disagreeing on joint interventions. We then give a conditional recovery result. Additive separability removes the hidden interaction degree of freedom, but observational residual summaries remain insufficient. Identification becomes possible when agents communicate causally identified response functions. An education value-added example illustrates why this is first a communication problem, and only then a policy-composition problem.

13:00 JST画像/動画生成

評決: 不一致を認識したコンセンサスによるマルチモーダル推論のトレーニング不要の段階的検証

マルチモーダルな大規模言語モデルでは、不正確な答えにつながる微妙なエラーを含む推論チェーンが生成されることがよくあります。現在の検証アプローチには顕著な制限があります。既存のアプローチでは、一貫性のないクロスタスクのパフォーマンスを伴う高価なラベル付き監視が必要か、単純な集計によって複数のソースからのスコアを集計する必要があり、重要な洞察が欠けています。これらのスコアが一致しない場合、その一致しないこと自体が、推論ステップが本当に有効かどうかに関する重要な情報を含んでいます。我々はこれを、異種の凍結した検証者間の結合スコアリング問題として形式化し、一致は有効なステップを示し、不一致は不安定性を明らかにする、独特の閉じた形式の均衡を伴う調整ゲームとして解釈できます。この目的に向けて、私たちは VERDICT (不一致情報結合しきい値による検証) と呼ばれる、トレーニング不要のドメインに依存しない段階的な検証アプローチを提案します。私たちの知る限り、VERDICT は、クロスモードの不一致の構造を明示的かつ実用的なものにする、トレーニング不要の初の検証ツールです。閉じた形式のソリューションを通じてコン​​センサス スコアを計算し、意見の相違を考慮したフィルタリングと推論ステップの安定性を考慮したランキングの両方を可能にします。 6 つのベンチマークにわたって評価された \method は、ベース モデルに対して一貫して最大 +5.95% 向上しており、広範な監督を必要とするドメイン固有の批評家と競争力を発揮し、タスク固有の適応やトレーニング不要の検証を必要とせずに、クロスモーダルの合意が堅牢な検証シグナルを提供することを実証しています。

原文 (English)

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Existing approaches either require expensive labelled supervision with inconsistent cross-task performance or aggregate scores from multiple sources by simple aggregations, missing a key insight: when these scores disagree, that disagreement itself carries important information about whether a reasoning step is truly valid or not. We formalise this as a coupled scoring problem among disparate, frozen verifiers, interpretable as a coordination game with a unique closed-form equilibrium where agreement signals valid steps while disagreement reveals instability. Towards this end, we propose a training-free domain-agnostic step-wise verification approach we call VERDICT: VERification via Disagreement-Informed Coupled Thresholding. To our knowledge, VERDICT is the first training-free verifier that makes the structure of cross-modal disagreement explicit and actionable. It computes consensus scores through a closed-form solution, enabling both disagreement-aware filtering and stability-conscious ranking of reasoning steps. Evaluated across six benchmarks, \method consistently improves over the base model by up to +5.95%, and performs competitively with domain-specific critics that demand extensive supervision, demonstrating that cross-modal agreement provides robust verification signals without task-specific adaptation and Training-Free Verification

13:00 JST研究/論文

FITTER: 時間的知識グラフにおける語彙に依存しないクロスドメイン推論

時間的知識グラフはセマンティック Web の多くの用途の中心ですが、既存の補完方法では、推論されるエンティティ、関係名、およびタイムスタンプがトレーニング時にすでに既知であると想定されており、各モデルが 1 つのグラフと語彙に制限されています。我々は、クロスドメイン転送をサポートする時間知識グラフリンク予測のための最初の完全帰納的構造モデルである FITTER を提案します。推論グラフには、まったく見えないエンティティ、関係名、および別のドメインから抽出されたタイムスタンプが含まれる可能性があります。 FITTER は、絶対的順序付けではなく相対的順序付けのエンコードを通じて、各述語を他の述語との相互作用パターンおよび時間によって表します。メッセージパッシングは、ローカルとグローバルの時間コンテキストを融合して、語彙に依存しない埋め込みを生成します。時間エンコーディングがタイムシフト不変であることを証明し、さまざまなドメイン、粒度、およびタイムスパンの 6 つの時間知識グラフ ベンチマークにわたるクロスドメイン、クロスグラフ転送で FITTER を評価します。 FITTER は再トレーニングなしで帰納的ベースラインを常に上回っており、語彙に依存しない構造学習がセマンティック Web の異種知識グラフに対する推論の実行可能な基盤であることを示しています。

原文 (English)

FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs

Temporal knowledge graphs are central to many uses of the Semantic Web, but existing completion methods assume the entities, relation names, and timestamps to be reasoned about are already known at training time, restricting each model to a single graph and vocabulary. We propose FITTER, the first fully-inductive structural model for temporal knowledge graph link prediction that supports cross-domain transfer: the inference graph may contain entirely unseen entities, relation names, and timestamps drawn from a different domain. FITTER represents each predicate by its interaction patterns with others and time through encodings of relative rather than absolute ordering; message-passing fuses local and global temporal context to produce vocabulary-agnostic embeddings. We prove the temporal encoding is time-shift invariant and evaluate FITTER on cross-domain, cross-graph transfer over six temporal knowledge graph benchmarks of diverse domains, granularities, and time spans. FITTER consistently outperforms inductive baselines without retraining, indicating that vocabulary-agnostic structural learning is a viable foundation for inference over the heterogeneous knowledge graphs of the Semantic Web.

13:00 JSTLLM/生成AIエージェント

REDAgentBench: 実行可能なレッド チーミングと LLM エージェント システムの忠実な測定

大規模言語モデル (LLM) エージェントは、言語ベースの推論と外部ツールを組み合わせて、複雑なタスクを実行します。敵対的な入力によってエージェントとその環境間の相互作用が悪用され、エージェントが実行中に安全ポリシーに違反する可能性があります。しかし、既存の評価ではエージェントの安全性が 1 回の攻撃成功率 (ASR) にまで引き下げられることが多く、暴露、実行、観察、判決が崩壊し、実際の違反と証拠の可視性が混同される可能性があります。自律的なレッドチーム化と忠実な測定のための実行可能なフレームワークである REDAgentBench を紹介します。明示的な安全上の制約と関連するエージェント システムの脆弱性から攻撃を引き出し、分離されたサービス サンドボックスで実行し、サービスの受信と最終状態の変更による悪影響を検証します。ベンチマークには、5 つのサービス サーフェスにわたる 1,661 件のケースが含まれています。 6 つのモデルと 3 つのエージェント ハーネスのマクロ平均 ASR は 65.69% です。報告される ASR はハーネスと証拠ビューによって異なりますが、評価コンテキストの開示により実行動作が変化します。状態に基づく診断コホートでは、アクションアンカーが解決された確認された違反のほぼ 5 件に 1 件が、エージェントが関連する制約またはリスクを述べた後に発生し、認識と実行のギャップが明らかになります。最後に、トレーニング不要のポリシー リマインダーにより、マッチ リプレイで確認された違反が 70 パーセント以上減少します。これらの調査結果は、実行可能な評価によって安全性測定が向上し、実行可能な介入ポイントを特定できることを示しています。

原文 (English)

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAgentBench, an executable framework for autonomous red-teaming and faithful measurement. It derives attacks from explicit safety constraints and associated agent-system vulnerabilities, runs them in isolated service sandboxes, and verifies harmful effects from service receipts and final-state changes. The benchmark contains 1,661 cases across five service surfaces. Across six models and three agent harnesses, macro-average ASR is 65.69%; reported ASR varies with harness and evidence view, while evaluation-context disclosure changes execution behavior. In a state-grounded diagnostic cohort, almost one in five confirmed violations with resolved action anchors occurs after the agent states the relevant constraint or risk, revealing a Recognition--Execution Gap. Finally, a training-free policy reminder reduces confirmed violations by more than 70 percentage points in matched replay. These findings show that executable evaluation can improve safety measurement and identify actionable intervention points.

13:00 JSTLLM/生成AIエージェント

ツリー構造メモリによる自己修正型の長期検索エージェント

大規模言語モデル (LLM) ベースの検索エージェントは、外部環境との複数段階の対話を通じて質問に答えます。ただし、完全な実行軌跡を LLM に提供すると、コンテキストが無制限に増大し、ノイズが発生します。既存の圧縮方法は、重要な詳細を犠牲にしてコンテキストを削減し、多くの場合、誤った事実から導き出された下流の推論を修復せずに誤った事実を置き換えます。この問題に対処するために、検索エージェントのための自己修正ツリー構造メモリ メカニズムである ReTree を提案します。 ReTree は、ソースにリンクされた証拠を保持しながら、制限されたステップごとの推論コンテキストを構築します。これは、ノードが限定された要約、証拠、改訂履歴を保存する証拠ツリーとして検索をモデル化します。新たに取得した証拠が以前の主張と矛盾する場合、ReTree はその主張が導入されたノードまで遡り、古い証拠を置き換え、概要を再生成し、影響を受けるブランチを削除して、検索を再開します。出典に基づいた証拠の来歴により、信頼できる紛争箇所の特定がサポートされ、最終的な主張が検索された文章まで追跡可能に保たれます。 4 つの公開質問回答および検索ベンチマークの実験では、ReTree が常に Full-Trajectory ReAct を上回っており、回答精度が最大 25.6 パーセント ポイント (pp) 向上していることが示されています。 Full-Trajectory ReAct のステップごとの平均最大推論コンテキストは、ReTree の $1.27$ ~ $1.51\time$ です。これらの結果は、ReTree が長期検索のための効果的な自己修正メモリ抽象化であることを確立します。

原文 (English)

Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, providing complete execution trajectories to the LLM causes unbounded context growth and introduces noise. Existing compression methods reduce context at the cost of important details and often replace erroneous facts without repairing downstream reasoning derived from them. To address this problem, we propose ReTree, a self-correcting tree-structured memory mechanism for search agents. ReTree constructs a bounded per-step reasoning context while preserving source-linked evidence. It models search as an evidence tree whose nodes store bounded summaries, evidence, and revision histories. When newly retrieved evidence contradicts an earlier claim, ReTree traces back to the node where the claim was introduced, replaces outdated evidence, regenerates summaries, prunes affected branches, and resumes search. Source-grounded evidence provenance supports reliable conflict localization and keeps final claims traceable to retrieved passages. Experiments on four public question-answering and search benchmarks show that ReTree consistently outperforms Full-Trajectory ReAct, improving answer accuracy by up to 25.6 percentage points (pp); the average maximum per-step reasoning context of Full-Trajectory ReAct is $1.27$--$1.51\times$ that of ReTree. These results establish ReTree as an effective self-correcting memory abstraction for long-horizon search.

13:00 JSTLLM/生成AI画像/動画生成

Ex-Omni-2D: ネイティブなビジュアル プレゼンスを備えた表現力豊かなオムニモーダル対話モデル

オムニモーダル対話モデルは、マルチモーダルな入力を理解し、音声による応答を合成できますが、その応答は視覚的に非具体的なままです。 \textbf{Ex-Omni-2D} を紹介します。これは、テキスト、パーソナライズされた音声、参照条件付きビデオで構成される調整された応答を生成するオムニモーダル対話フレームワークです。マルチモーダル クエリ、参照画像、および参照音声が与えられると、モデルはシーン、感情、動きを記述する構造化された \textit{視覚的思考計画} (VTP) を予測し、その後に応答テキストとネイティブ マルチ コードブック音声単位が続きます。これらのユニットは共有の音響時間インターフェースを形成します。これらのユニットは音声にデコードされ、ビデオ フレームとオンラインで調整されます。このインターフェイスにより、応答とアバターの経路を異種の音声、対話、およびアバタービデオデータから学習できるため、大規模なクエリ (テキスト、音声、ビデオ監視) の必要性が回避されます。フルシーケンス ビデオ ジェネレーターは、主要な教師として機能します。効率的な増分生成のために、これを数ステップのブロック因果関係のある \emph{Streaming Student} にさらに蒸留します。このプレフィックス ストリーミング メカニズムは、連続するチャンク全体にクリーン レイテントを運び、累積的な後期チャンクの劣化を軽減します。 4 ステップの推論により、完全な 4 GPU パイプラインは $400\times720$/$720\times400$ で 1.293 のエンドツーエンド RTF を達成し、実用的な品質効率の動作点を提供します。

原文 (English)

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured \textit{Visual Thought Plan} (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal \emph{Streaming Student} whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at $400\times720$/$720\times400$, providing a practical quality--efficiency operating point.

13:00 JST研究/論文

Tree-of-Ideas: 学術進化における交差軌道推論による自動化された研究アイデア

効果的な研究のアイデアを作成するには、以前の研究の静的な理解を超えて、研究の問題と解決策が文献全体でどのように進化するかを追跡する必要があります。既存の手法は、論文を構造化されていない文脈として扱うか、学術の進化を孤立した引用連鎖としてモデル化し、研究軌跡間の相互作用を見落としています。私たちは、2 段階のフレームワークである Tree-of-Ideas (ToI) を提案します。 EvoTrace は、引用から枝分かれした学術の軌跡を再構築し、進化する手法、解決された問題、ギャップを追跡します。その後、EvoAgent は軌道全体を推論して収束する問題と補完的な解決策を特定し、根拠のある研究アイデアを生成します。 6 つの AI 研究トピック全体で、ToI は自動手法の中で最高のスコア (10 点スケールで最も強いベースラインの 6.27 対 5.36) を達成し、強力な新規性 (6.36) とグラウンディング性 (7.00) を実現しました。また、そのスコアは人間の論文の参考文献 (6.29) のスコアに近づき、クロスパス進化論的推論の価値を示しています。

原文 (English)

Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution

Effective research ideation requires moving beyond a static understanding of prior work to trace how research problems and solutions evolve across the literature. Existing methods either treat papers as unstructured context or model scholarly evolution as isolated citation chains, overlooking interactions among research trajectories. We propose Tree-of-Ideas (ToI), a two-stage framework. EvoTrace reconstructs branching scholarly trajectories from citations, tracking evolving methods, resolved problems, and gaps. EvoAgent then reasons across trajectories to identify convergent problems and complementary solutions, generating grounded research ideas. Across six AI research topics, ToI achieves the highest score among automatic methods (6.27 vs. 5.36 for the strongest baseline on a 10-point scale), with strong Novelty (6.36) and Groundedness (7.00). Also, its score approaches that of human-paper references (6.29), demonstrating the value of cross-path evolutionary reasoning.

13:00 JST画像/動画生成研究/論文

階層的な人間の行動認識のための構成ベンチマーク合成

原子的な行動から長期的な意図に至るまで、抽象化レベル全体で人間の行動を認識するには、意味論的な階層に沿って注釈が付けられたデータが必要です。大規模なコーパスは、時間的構成のない、分離された原子的にラベル付けされたクリップを提供しますが、記録された複合アクティビティ コーパスは、浅く、ドメインが狭く、固定された階層を提供します。アクションレベルで事前に抽出された実際の特徴を保持しながら、フラットな単一ラベルアクションコーパスから、アクション、アクティビティ、低レベル意図(LLI)、および高レベル意図(HLI)にわたる4レベルの階層的意図ベンチマークを合成するベンチマーク生成および評価フレームワークが提案されています。エピソードは、主題の一貫性制約の下で移行モデルによって組み立てられ、カバレッジを意識したサンプラーにより、主題の使用ジニが 0.566 から 0.248 に削減されます。このようなベンチマークを合成すると、記録されたデータセットが回避する循環監視のリスクが生じます。エピソードを生成するルールが評価にも影響を与える場合、モデルは真の推論ではなくジェネレーターを回復することによって成功する可能性があります。妥当性は設計によって対処され、評価時に使用される一次論理ルールから切り離されたシーケンス生成ルールを保持します。インスタンス化により 15,002 のエピソードが生成されます。さまざまなモデル ファミリからの 4 つの参照ベースラインは、認識方法としてではなく、難しさを特徴づけます。マクロ F1 の 0.13 ~ 0.17 という組成上のホールドアウト ギャップが、最良の認識をするもののギャップを埋めていないグラフ認識モデルを含むすべてのベースラインにわたって表示され、これはモデルのアーティファクトではなくベンチマークの構造的特性を示しています。ロジックフリーのベースラインは、依然としてその固有のデータ レートを超えて保持されているセマンティック ルールに違反しており、順序破壊制御はシード バリエーション内でマクロ F1 を変更し、ジェネレータの一貫性チェックとして機能します。オントロジー、移行モデル、およびジェネレーターがリリースされているため、ベンチマークを再生成および拡張できます。

原文 (English)

Compositional Benchmark Synthesis for Hierarchical Human Action Recognition

Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy. Large corpora provide isolated, atomically labeled clips without temporal composition, whereas recorded composite-activity corpora offer shallow, domain-narrow, fixedhierarchies. A benchmark-generation and evaluation frameworkis proposed that synthesizes a four-level hierarchical-intention benchmark, spanning actions, activities, low-level intentions (LLIs), and high-level intentions (HLIs), from a flat single-label action corpus while retaining real pre-extracted features at the action level. Episodes are assembled by a transition model under a subject-consistency constraint, and a coverage-aware sampler reduces the subject usage Gini from 0.566 to 0.248. Synthesizing such a benchmark raises a circular-supervision risk that recorded datasets avoid: if the rules generating the episodes also govern the evaluation, models can succeed by recovering the generator rather than through genuine reasoning. Validity is addressed by design, holding sequence-generation rules disjoint from the first-order-logic rules used at evaluation. The instantiation yields 15,002 episodes. Four reference baselines from different model families characterize difficulty, not as recognition methods. A compositional held-out gap of 0.13 to 0.17 macro-F1 appears across all baselines, including a graph-aware model that recognizes best yet does not close the gap, indicating a structural property of the benchmark rather than a model artifact. A logic-free baseline still violates the held-out semantic rules above their intrinsic data rate, and the order-destroying control changes macro-F1 within seed variation, serving as a generator-consistency check. Theontology, transition model, and generator are released so the benchmark can beregenerated and extended.

13:00 JST研究/論文

経験則: 部分情報を使用して人工知能システムを説明する

説明可能な人工知能 (XAI) は、人工知能 (AI) システムが特定の決定にどのように到達したかを説明しようとします。私たちは、特定のデータポイントに対する AI システムの動作を予測するために最も関連する特徴を特定する新しい定式化に基づく XAI への新しいアプローチである「経験則」(RoT) 説明を提案します。 RoT が、(a) 大規模言語モデル (LLM) を使用したゼロショット分類、(b) モデルへのアクセスなしの不透明な AI システムの監査、(c) 科学的発見における AI の使用において、XAI を実現するのにどのように適しているかを示します。さらに、RoT は主要な AI 規制の特定の要件を満たし、XAI 実践者に使い慣れたインターフェイスと視覚化を提供し、モデルに依存せず、代替手段よりも大幅に高速です。コードは次から入手できます: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

原文 (English)

Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information

Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. We show how RoT is well-suited to enable XAI in: (a) zero-shot classification using large language models (LLMs), (b) auditing of opaque AI systems without model access, and (c) the use of AI in scientific discovery. Additionally, RoT meets specific requirements from leading AI regulations, provides a familiar interface and visualisations for XAI practitioners, is model-agnostic, and is substantially faster than alternatives. Code available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

13:00 JSTエージェントGPT / ChatGPT

SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation

Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize in…

13:00 JSTエージェント

ChemWorld: 制御および再生可能なエージェント実験のためのプログラム可能な Chemical World

自律化学は、エージェントが繰り返し行動、観察、適応できる環境にますます依存しています。物理実験室は、重要な実際の物質的な証拠を提供しますが、繰り返しにコストがかかり、厳密に一致した介入に使用するのが難しいのに対し、ほとんどのデジタル環境は、基礎となる実験世界をほぼ固定したままにしています。再利用可能なプロセスおよび観察コンポーネントが実行可能な世界にコンパイルされる、プログラム可能な化学環境である ChemWorld を紹介します。 ChemWorld は、エージェントが利用できる公的実験契約を、評価者が所有する化学法および材料法から分離します。したがって、研究者は、公共のタスクと相互作用の条件を固定したまま、世界の構成と動作条件を変更したり、単一の隠された法則を変更したりすることができます。トランザクションの実行により、操作、障害、リソースの変更、状態遷移が記録され、完全な環境アクションの軌跡を正確に再生して監査できるようになります。完全な国勢調査の認定には、参照レジストリ、52 の生成されたコンポジション、モジュール、インターフェイス、コンパイル、および無効なアクションのテストが含まれていました。 8 つの決定論的な実験ケースでは、共有ライフサイクル セマンティクス、障害回復、正確なリプレイが実証され、6 つの親子ワールド フォーク ペアでは、一致する公共条件下での単一の私法介入の影響が分離されました。独立したエージェントも、同じパブリック インターフェイスを通じて非参照世界で完全なライフサイクルを完了しました。 ChemWorld は、宣言されたコンポーネントとモデルの領域内で、物理実験室の証拠とキャリブレーションを補完する、系統的に変化する化学の世界にわたる実験を研究するための制御された再生可能な基盤を提供します。

原文 (English)

ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation

Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories provide essential real-material evidence but are costly to repeat and difficult to use for tightly matched interventions, whereas most digital environments keep the underlying experimental world largely fixed. We introduce ChemWorld, a programmable chemical environment in which reusable process and observation components are compiled into executable worlds. ChemWorld separates the public experimental contract available to an agent from evaluator-owned chemical and material laws. Researchers can therefore vary world composition and operating conditions, or change a single hidden law while holding the public task and interaction conditions fixed. Transactional execution records operations, failures, resource changes, and state transitions, allowing complete environment-action trajectories to be replayed exactly and audited. Full-census qualification covered the reference registry, 52 generated compositions, and module, interface, compilation, and invalid-action tests. Eight deterministic experimental cases demonstrated shared lifecycle semantics, failure recovery, and exact replay, while six parent-child world-fork pairs isolated the effects of single private-law interventions under matched public conditions. An independent agent also completed a full lifecycle in a non-reference world through the same public interface. Within the declared component and model domain, ChemWorld provides a controlled and replayable substrate for studying experimentation across systematically varied chemical worlds, complementary to physical-laboratory evidence and calibration.

13:00 JST研究/論文

EvoMem: Memory-Augmented Evolution for Code Optimization

Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some c…

13:00 JSTLLM/生成AI

Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction

First-order concept synthesis asks a system to infer one formula that classifies labeled objects consistently across several finite relatio…

13:00 JST研究/論文

制約ロジック プログラミングにおけるユークリッド巡回販売員問題とその変形用の強化されたフィルタリング アルゴリズム

巡回販売員問題 (TSP) は、コンピューター サイエンスで最もよく知られている問題の 1 つであり、スマート車両やインテリジェント交通システムなどの多くのエンジニアリング アプリケーションで発生します。 「ユークリッド」の場合、各ノードは平面内の座標によって定義され、距離はユークリッド計量を使用して計算されます。制約プログラミング (CP) の文献では、ユークリッド TSP は通常、完全な距離行列を計算し、それを一般的なケースとして扱うことによって扱われます。ただし、このアプローチでは、点の座標によって伝えられる幾何学的情報が無視されます。この研究では、制約ロジック プログラミング (CLP) で実装された新しいフィルタリング アルゴリズムを提案します。このアルゴリズムは、このような幾何学的情報を利用して、既存のアプローチよりも強力な制約の伝播を実現します。さらに、この方法論を、実際のルーティングおよび物流アプリケーションに関連するユークリッド一般巡回販売員問題 (EGTSP) など、TSP の他のユークリッド変種にどのように拡張できるかを示します。実験結果は、提案されたアプローチの計算上の利点を示しています。

原文 (English)

Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming

The Traveling Salesperson Problem (TSP) is one of the best-known problems in computer science and arises in many engineering applications, such as smart vehicles and intelligent transportation systems. In the "Euclidean" case, each node is defined by its coordinates in the plane and distances are computed using the Euclidean metric. In the Constraint Programming (CP) literature, the Euclidean TSP is typically addressed by computing the full distance matrix and treating it as a general case; however this approach ignores the geometric information carried by the points' coordinates. In this work, we propose new filtering algorithms, implemented in Constraint Logic Programming (CLP), that exploit such geometric information to achieve stronger constraint propagation than existing approaches. Moreover, we show how this methodology can be extended to other Euclidean variants of the TSP, including the Euclidean Generalized Traveling Salesperson Problem (EGTSP), which is relevant in practical routing and logistics applications. Experimental results demonstrate the computational advantages of the proposed approach.

13:00 JSTエージェント

ComBodied Agents: 人間中心のエージェント AI の新しいパラダイム

高齢者が薬を飲み忘れた場合、ソフトウェア エージェントが別のリマインダーを送信し、実体のあるエージェントが薬を持ってくることができます。しかし、その人が忘れているのか、混乱しているのか、副作用があるのか​​、それとも故意に拒否しているのか、またどのような支援が適切なのかについてはどちらも説明していない。これは、Agentic AI の構造的なギャップを明らかにします。デジタル エージェントは主にソフトウェアの状態を変換しますが、身体化されたエージェントは物理的な状態を変換します。どちらも、人の進化する状態や主体性をモデル化、介入、評価の主な対象にするものではありません。ソフトウェア ツール、センサー、ウェアラブル、ロボット、ヒューマン サービスを最終目標ではなくアクション チャネルとして使用し、時間の経過に伴う個々の人間の状態の軌跡を認識、モデル化、予測、サポートする人間中心のパラダイムである複合エージェントを紹介します。私たちは、パーソナル アシスタント、ヘルス エージェント、AI コンパニオン、および適応型人間 AI システムにわたる断片化された機能を統合して閉ループにします。イベントベースのマルチモーダルな知覚により、意味のある個人的なイベントが再構築されます。縦方向の修正可能な記憶は時間的コンテキストを提供します。個人世界モデルは、別の決定と介入の下で将来の個人の状態と結果を推定します。そして、許容可能な介入方針は、同意、不確実性、安全性、可逆性、およびユーザーの制御の下で、比例したサポートを選択します。人や環境からのフィードバックによってループが更新されます。このフレームワークは、網羅的なヒューマン デジタル ツインを必要とするのではなく、目的が限定され、不確実性を認識し、ユーザーが修正可能な表現を使用します。私たちは、人間の状態のターゲット、関係性のコンテキスト、およびエージェントの役割によって設計空間を整理し、シナリオ中心の評価、主体性維持のメトリクス、ベンチマーク要件、エッジネイティブの個人モデル、およびガバナンスの方向性を提案します。複合エージェントは、Agentic AI を外部タスクの完了から持続的な人間の利益へと移行させます。

原文 (English)

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.

13:00 JST研究/論文

IO Factory: AI を活用した影響力キャンペーンを大規模にシミュレーション

IO Factory は、情報をシミュレートし、完全に統合された追跡可能なプロセスとしてキャンペーンに影響を与える AI 主導のフレームワークです。デジタル操作の脅威は現在、個々の言語モデルからの説得力のあるテキストを超えて、AI 群、つまりプラットフォームのフィードバックに適応し、組織化されたキャンペーンを通常の社会的交流として偽装する、調整されたエージェントの永続的なグループにまで広がっています。このようなキャンペーンは、個別のメッセージだけでは特定できないため、計画、プラットフォームのアクション、露出、解釈、測定、適応といった継続的なスペクトル全体にわたって分析する必要があります。 IO Factory は、制御されたシミュレートされたプラットフォーム内でこのプロセスを表し、アクターの役割、プラットフォームのアクション、暴露記録、構造化されたモデルベースの評価、およびシミュレートされた集団の構成された変更をリンクします。私たちはアーキテクチャを実装し、最大 100,000 のエージェントの構成にわたって評価します。結果は、IO Factory がキャンペーンのタイムラインを大規模に実行し、設定された信念変数における暴露と測定された動きの検査可能な証拠を生成していることを示しています。 IO Factory は、各実行で使用されるアクター、目的、アクションの制約、暴露経路、測定ルールを記録することで、再現可能な調査と調整された影響のレッドチーム分析をサポートします。

原文 (English)

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat of digital manipulation now extends beyond persuasive text from individual language models to AI swarms, i.e., persistent groups of coordinated agents that adapt to platform feedback and disguise organized campaigns as ordinary social interaction. Because such campaigns cannot be identified from isolated messages alone, they must be analyzed across a continuous spectrum of planning, platform action, exposure, interpretation, measurement, and adaptation. IO Factory represents this process inside a controlled simulated platform, linking actor roles, platform actions, exposure records, structured model-based evaluations, and configured changes in the simulated population. We implement the architecture and evaluate it across configurations of up to 100,000 agents. The results show that IO Factory executes campaign timelines at scale and produces inspectable evidence of exposure and measured movement in configured belief variables. By recording the actors, objectives, action constraints, exposure paths, and measurement rules used in each run, IO Factory supports reproducible research and red-team analysis of coordinated influence.

13:00 JST研究/論文

ThinkRetrieve: テスト時間のスケーリングのための検索拡張推論トレース

大規模推論モデル (LRM) は、拡張された思考連鎖推論を生成するために追加の推論時間コンピューティングを割り当てることで、パフォーマンスを向上させます。しかし、最近の研究では、トレースが長くなると不確実性が増し、誤差が増大し、元の問題からのずれが生じるため、逐次テスト時間のスケーリングでは利益が減少したり、マイナスになることが多いことが明らかになりました。我々は、各推論ステップで動的に取得された解決例で LRM の推論トレースを強化するテスト時間スケーリング フレームワークである ThinkRetrieve を提案します。ステップバイステップの解決策と対になった問題の外部コーパスが与えられると、ThinkRetrieve は各中間ステップで関連する例を取得し、それらを思考トレースに直接注入し、単にどの事実が関連するかではなく推論方法についてのガイダンスをモデルに提供します。 GSM-8K、MATH-500、AIME 2025、および SciQ での 5 つの推論モデル (1.5B ~ 8B パラメーター) にわたる実験では、ThinkRetrieve が標準のテスト時間スケーリングよりも一貫して精度を向上させ、AIME 2025 では最大 $60\%$ の相対的な利益が得られることが実証されました。

原文 (English)

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling

Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminishing or even negative returns, as longer traces exhibit increased uncertainty, error compounding, and drift from the original problem. We propose ThinkRetrieve, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step. Given an external corpus of problems paired with step-by-step solutions, ThinkRetrieve retrieves relevant exemplars at each intermediate step and injects them directly into the thinking trace, providing the model with guidance on how to reason rather than merely what facts are relevant. Experiments across five reasoning models (1.5B--8B parameters) on GSM-8K, MATH-500, AIME 2025, and SciQ demonstrate that ThinkRetrieve consistently improves accuracy over standard test-time scaling, with relative gains of up to $60\%$ on AIME 2025.

13:00 JST研究/論文

FedCGR: フェデレーテッド・クロスドメイン生成推奨事項

クロスドメイン レコメンデーション (CDR) は、関連ドメイン間でプリファレンスの知識を転送しますが、重複するユーザーや共有インタラクション信号などのアイテム空間を調整する動作アンカーは、クライアント間で希薄であるか、利用できないか、プライバシーに敏感であることが多いため、フェデレーション展開ではクロスドメインの調整が困難になります。この緊張に対処するために、安定したセマンティック項目言語を介した生成としてフェデレーテッド CDR を再考します。パブリックなアイテム側のメタデータから派生した個別のセマンティック ID (SID) シーケンスとしてアイテムを表すことにより、プライベート インタラクションを交換したり、ドメイン固有の埋め込みを調整したりするのではなく、共有語彙によってクロスドメイン アイテムの調整が行われます。ただし、SID ベースのジェネレーターを直接フェデレーションすると、設計上の 2 つの制約が生じます。1 つは、クライアント間のトークンの一貫性を維持するために SID トークナイザーを固定しておく必要があり、ローカルの協調フィルタリング (CF) 信号をグローバルに共有または調整できないため、セマンティックのみのボトルネックが発生します。一方、標準のフェデレーション平均化では、ドメインの異質性のもとで負の転送が発生する可能性があります。これらの制約を克服するために、アイテム言語を安定に保ち、適応を明示する連合生成 CDR フレームワークである FedCGR を提案します。 FedCGR は、信頼性を意識したセマンティック インターフェイスを通じてローカル CF 証拠を注入し、ドメイン固有の量をローカルに保ちながら、ドメインの関連性に応じて共有パラメーターを選択的に集約するプロトタイプのパーソナライズされたジェネレーターをトレーニングします。 6 つの Amazon クロスドメイン シナリオでの実験では、FedCGR が一貫してフェデレーテッド生成ベースラインを上回り、フルランキング評価プロトコルとサンプル評価プロトコルの両方で、強力なシーケンシャル CDR メソッドおよびフェデレーテッド CDR メソッドに対して競争力のあるパフォーマンスを達成することが示されています。

原文 (English)

FedCGR: Federated Cross-Domain Generative Recommendation

Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult because the behavioral anchors that align item spaces, such as overlapping users and shared interaction signals, are often sparse, unavailable, or privacy-sensitive across clients. To address this tension, we revisit federated CDR as generation over a stable semantic item language. By representing items as discrete semantic ID (SID) sequences derived from public item-side metadata, cross-domain item alignment is induced by a shared vocabulary rather than by exchanging private interactions or aligning domain-specific embeddings. Directly federating SID-based generators, however, introduces two design constraints: the SID tokenizer must remain fixed to preserve cross-client token consistency, which creates a semantic-only bottleneck because local collaborative filtering (CF) signals cannot be globally shared or aligned; meanwhile, standard federated averaging can cause negative transfer under domain heterogeneity. To overcome these constraints, we propose FedCGR, a federated generative CDR framework that keeps the item language stable and makes adaptation explicit. FedCGR injects local CF evidence through a reliability-aware semantic interface and trains a prototype-personalized generator that selectively aggregates shared parameters according to domain relatedness while keeping domain-specific quantities local. Experiments on six Amazon cross-domain scenarios show that FedCGR consistently outperforms federated generative baselines and achieves competitive performance against strong sequential and federated CDR methods under both full-ranking and sampled evaluation protocols.

13:00 JSTエージェント

XCoT-VLA: 視覚-言語-行動の運転のための実行可能な思考連鎖

Vision-Language-Action(VLA)モデルは、シーンの理解、意味論的推論、自動運転のための軌道生成を結び付けることができます。ただし、冗長な自然言語の思考連鎖 (CoT) は、制限がなく、デコードにコストがかかり、アクションに面した表現として最適化するのが難しいため、リアルタイム制御にはあまり適していません。我々は、記述的根拠を、自動的に構築されたReason-Action監視から学習したコンパクトな実行可能なCoTトークンに置き換えるXCoT-VLAを提案します。記録された軌跡はアクションの証拠を提供し、シーンのコンテキストは因果関係の意味論を提供します。予測された XCoT シーケンスはコンテキスト内に残り、共有されたマルチモーダルセルフアテンションを通じて固定軌道クエリを条件付けします。決定論的なトークン関数ルーティングは、フローマッチング軌跡生成のための理由 FFN を XCoT トークンに適用し、制御 FFN を軌跡クエリに適用します。さらに、同じ実行可能トークン空間内のオプションの改良拡張機能として XCoT Policy Optimization (XCPO) を導入します。 XCoT-VLA は、一般分散セットで縦方向 ADE を 1.645 から 1.323 に減少させ、車線変更シナリオで横方向 FDE を 1.616 から 0.648 に減少させます。わずか 2 ~ 6 個の実行可能な XCoT トークンで運転指向の推論を表すことにより、私たちの方法は自己回帰推論のオーバーヘッドを大幅に削減し、リアルタイムの計画予算内に収まります。これらの結果は、運転指向の推論がコンパクトで実行可能であり、軌道生成に直接接続できることを示しています。

原文 (English)

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time control because it is open-ended, costly to decode, and difficult to optimize as an action-facing representation. We propose XCoT-VLA, which replaces descriptive rationales with compact executable CoT tokens learned from automatically constructed Reason-Action supervision. Logged trajectories provide action evidence, while scene context supplies causal semantics. The predicted XCoT sequence remains in context and conditions fixed trajectory queries through shared multimodal self-attention. Deterministic token-function routing applies the Reason FFN to XCoT tokens and the Control FFN to trajectory queries for flow-matching trajectory generation. We further introduce XCoT Policy Optimization (XCPO) as an optional refinement extension in the same executable token space. XCoT-VLA reduces longitudinal ADE from 1.645 to 1.323 on a general-distribution set and lateral FDE from 1.616 to 0.648 in lane-change scenarios. By representing driving-oriented reasoning with only 2-6 executable XCoT tokens, our method substantially reduces autoregressive reasoning overhead and remains within the real-time planning budget. These results demonstrate that driving-oriented reasoning can be compact, executable, and directly connected to trajectory generation.

13:00 JSTLLM/生成AI研究/論文

V-FiLLM: 検証済みの財務 LLM 推論ベンチマーク

既存のベンチマークは、STEM ドメイン全体での LLM の評価において大幅な進歩を遂げていますが、構造化データに対する財務上の推論については、比較的研究が進んでいません。 V-FiLLM は、実際のテーブルに基づいた実行可能な計算ツリーから財務推論ベンチマークを生成し、構築によって正解となる項目を生成するフレームワークを紹介します。ツリーは記号的に評価されてグランド トゥルースを取得し、自然言語の質問にレンダリングされ、ラベル付けループからモデルが削除されるため、アノテーション コストやジェネレーターのエラー率を継承することなく、任意のスケールで項目を生成できます。 V-FiLLM は、計算の深さ、表現の幅、財務概念の複雑さ、コンテキストのサイズを含む、独立して制御可能な 4 つの難易度軸を明らかにします。オープンソース モデルで評価したところ、推論の深さが増すにつれて精度が最大 51% 低下し、敵対的な数値摂動下では最大 47% ポイント低下することがわかり、テーブルを使用した堅牢な財務推論に残された課題が浮き彫りになっています。さらに、検証済みの思考連鎖トレースに対する軽量の LoRA 微調整により、保留された問題の精度が 81.1% から 85.6% に向上し、FinQA で基本モデルよりも 5% ポイント優れていることを示します (Chen et al., 2022a), s)。これは、ターゲットを絞った低コストの適応が、財務 QA における構成推論の有望な方向性であることを示唆しています。

原文 (English)

V-FiLLM: Verified Financial LLM Reasoning Benchmark

While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction. Trees are evaluated symbolically to obtain ground truth and rendered into natural-language questions, removing any model from the labeling loop, so items can be generated at arbitrary scale without annotation cost and without inheriting a generator's error rate. V-FiLLM exposes four independently controllable axes of difficulty including computation depth, expression breadth, financial concept complexity, and context size. By evaluating on open-source models, we find that accuracy falls up to 51% as reasoning depth increases, and up to 47% points under adversarial numerical perturbations, highlighting remaining challenges in robust financial reasoning over tables. We further show that lightweight LoRA fine-tuning on verified chain-of-thought traces improves accuracy from 81.1% to 85.6% on held-out problems and outperforms the base model by 5% points on FinQA (Chen et al., 2022a), s), suggesting that targeted, low-cost adaptation is a promising direction for compositional reasoning in financial QA.

13:00 JSTエージェントビジネス/資金調達

SkillZip: 再利用可能な構造の発見による自己進化エージェントの評価不要のスキル圧縮

自己進化するエージェントは、成功した手順や失敗の修正を追加することで、再利用可能なスキルを蓄積します。時間が経つと、同じ要件が複数の分岐、例、警告で再度説明されることが多くなり、共通のアクション シーケンスは再利用されずにコピーされます。結果として得られるスキルは注入にコストがかかり、維持するのが困難になります。スキルは平坦なパッセージではないため、一般的なプロンプト圧縮はこの設定には適していません。スキルの名前と説明はいつ適用されるかを定義し、ワークフローは実行を制御し、ツールと出力コントラクトは有効性を制約し、まれな例外は、それらをアクティブ化するサンプル タスクがない場合でも必須のままである可​​能性があります。評価ガイド付き圧縮ではこれらの動作をテストできますが、ロールアウト、コスト、および圧縮時間の評価セットへの依存が生じます。私たちは、最短で忠実な構造的説明を見つけることによってスキルを圧縮する、評価不要のメソッドである SkillZip を紹介します。直感は一度説明し、多くを参照します。繰り返しのルールを適用範囲で一度だけ記述し、繰り返しのアクション シーケンスを共有プロシージャに要素化し、相違点のみを明示的な例外として保持します。この直感を、抽出されたすべてのトリガー、ワークフロー エッジ、ツール要件、義務、および出力フィールドに対するハード カバレッジ制約の対象となる、スキル契約および残余に関する型付きの最小記述長目標として形式化します。この定式化は、単純な共有しきい値を提供し、構築により固有のまれなルールを保存し、効率的なローカル更新をサポートします。 SkillZip には、1 つの構造化された抽出呼び出しと決定論的な最適化を備えたワンショット モードと、タスクの再生や完全な履歴の再解析を行わずに各自己進化パッチを統合する継続的な Zip-on-Write モードがあります。包括的な実験評価を通じて、圧縮パフォーマンス、汎用性、コストオーバーヘッドにおける SkillZip の有効性と優位性を実証します。

原文 (English)

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors, but it introduces rollouts, cost, and dependence on the compression-time evaluation set. We present SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. We formalize this intuition as a typed minimum description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Through comprehensive experimental evaluations, we demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.

13:00 JST研究/論文

RTSKG: 鉄道駅のナレッジ グラフ データセットの構築

鉄道輸送システムは、都市のモビリティと経済発展において重要な役割を果たしています。このようなシステムの主要な構成要素である鉄道駅は、都市へのアクセスを向上させ、周辺地域の開発を刺激する重要な交通ハブとして機能します。都市レベルの鉄道駅に関連するタスク(乗客数の予測など)には大規模な都市データが必要ですが、現在の研究では、データ構成の観点からさまざまな都市エンティティ間の複雑な相互作用が無視されていることがよくあります。この論文では、上記の問題に対処するために、都市レベルの鉄道駅関連タスクに役立つように、さまざまな種類の都市エンティティ間の空間的および意味論的な相互作用を明示的にモデル化する鉄道駅ナレッジ グラフ (RTSKG) データセットを構築します。 RTSKG は、鉄道駅、道路セグメント、名所などの異質な都市エンティティを特別に設計された統一スキーマで統合し、https://w3id.org/rtskg/ でリンク データとしてアクセスできます。駅周辺の店舗の推奨と知識を活用した乗客数の予測に関する評価は、RTSKG の有効性を実証し、都市レベルの鉄道駅分析をサポートする可能性を強調しています。

原文 (English)

RTSKG: Building a Rail Transit Station Knowledge Graph Dataset

Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems, rail transit stations function as critical transport hubs that enhance urban accessibility and stimulate development in surrounding areas. City-level rail transit station related tasks (e.g., ridership prediction) require large-scale urban data, but current studies often neglect complex interactions among various urban entities in terms of data organization. In this paper, to address the above issue, we build a Rail Transit Station Knowledge Graph (RTSKG) dataset which explicitly models the spatial and semantic interactions among different kinds of urban entities, to benefit city-level rail transit station related tasks. RTSKG integrates heterogeneous urban entities, such as rail transit stations, road segments, and points of interest, with a specially designed unified schema, and is accessible as Linked Data at https://w3id.org/rtskg/. Evaluations on station-area store recommendation and knowledge-enhanced ridership prediction demonstrate the effectiveness of RTSKG, highlighting its potential to support city-level rail transit station analysis.

13:00 JSTLLM/生成AIエージェントClaude

CLAUDE.md はなぜ成長を続けるのか?エージェントコーディングにおける壊滅的な記憶

CLAUDE.md のようなエージェント コーディング README は、実際のリポジトリ内で際限なく増大し、リポジトリが廃止されるか、誰かがファイルを全面的に書き換えた場合にのみ停止します。これを不完全な再現にたどります。命令の追加は常に低コストですが、命令の根拠が失われると、正確性の低下の危険を冒さずに命令を削除すると、|D| のプロンプトで O(2^|D|) のコストがかかります。説明書。私たちは、その結果として生じる発散を、壊滅的な記憶、つまり継続的な学習が組織される壊滅的な忘却の逆と名付けます。まず、1,867 のリポジトリにおける 247,694 の命令の存続期間にわたってこの現象を特徴付けます。エージェント プロンプトは際限なく増加し、存続期間中に 3 倍以上 (+226%)、コミットごとに正味命令数が +4.9 増加します。さらに、命令が古くなるほど、削除される可能性は低くなります (log-hazard -0.032/commit)。次に、プロンプトコメントが成長を止めることができることを示します。IFEval を反転すると、最適なプロンプトが既知である検証可能な世界が得られ、潜在的な推論をエンコードしたコメントによって過剰な命令が 99.3% 除去されます (+211.3% から +1.4%)。最後に、同じ逆変換を WildIFEval に適用すると、即時コメントによって現実世界のエージェントによる指示のフォローが最大 23.1% 向上することがわかります。英語が新しいコードである場合、なぜまだコメントがないのでしょうか?

原文 (English)

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2^|D|) in a prompt of |D| instructions. We name the resulting divergence catastrophic remembering, the inverse of catastrophic forgetting around which continual learning is organized. First, we characterize this phenomenon across 247,694 instruction lifetimes in 1,867 repositories: agentic prompts grow without bound, more than tripling over their lifetime (+226%), gaining +4.9 net instructions every commit; further, the older an instruction gets, the less likely it is to be deleted (log-hazard -0.032/commit). Then, we show that prompt comments can halt the growth: inverting IFEval yields verifiable worlds whose optimal prompts are known, and there comments encoding latent reasoning remove 99.3% of excess instructions (+211.3% to +1.4%). Finally, applying the same inversion to WildIFEval, we show that prompt comments can improve real-world agentic instruction-following by up to 23.1%. If English is the new code, why don't we have comments yet?

13:00 JST研究/論文

sLTN: 構造論理テンソル ネットワーク

ロジック テンソル ネットワーク (LTN) は、テンソル演算を通じて 1 次ロジックを解釈する神経記号フレームワークを提供し、論理制約を微分可能な学習と統合できるようにします。ただし、LTN の元の定式化は主に個人のフラットなコレクションとして表現されるデータに適しており、時間的順序、連続した位置、グラフの接続性などの構造的構成を明示的に捕捉しません。構造的次元を言語の第一級要素にする LTN の拡張である sLTN を紹介します。構造次元は、タイム ステップ、シーケンス位置、グラフ ノードなどのドメイン固有の構成に関連付けられた名前付きテンソル軸を表します。それらは明示的に定量化され、構造的関係を通じて関連付けられ、時間的、順序的、および関係的な制約を論理レベルで直接表現するために使用できます。 sLTN の構文とファジー テンソル セマンティクスを形式化し、構造次元が存在しない場合、フレームワークが特殊なケースとして元の LTN セマンティクスを回復することを示します。さらに、宣言的署名、数式解析、テンソル解釈に基づいた PyTorch 実装について説明します。このフレームワークは、代表的な時間的推論と逐次的推論の例で説明されています。このペーパーは、https://github.com/logictensornetworks/sltn で入手可能な sltn ライブラリの補足として機能します。

原文 (English)

sLTN: Structural Logic Tensor Networks

Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tensor operations, enabling logical constraints to be integrated with differentiable learning. However, the original formulation of LTN is primarily suited to data represented as flat collections of individuals, and does not explicitly capture structural organization such as temporal order, sequential position, or graph connectivity. We introduce sLTN, an extension of LTN that makes structural dimensions first-class elements of the language. Structural dimensions represent named tensor axes associated with domain-specific organization, such as time steps, sequence positions, or graph nodes. They can be quantified explicitly, related through structural relations, and used to express temporal, sequential, and relational constraints directly at the logical level. We formalize the syntax and fuzzy tensor semantics of sLTN and show that, in the absence of structural dimensions, the framework recovers the original LTN semantics as a special case. We further describe a PyTorch implementation based on a declarative signature, formula parsing, and tensorial interpretation. The framework is illustrated on representative temporal and sequential reasoning examples. This paper serves as a companion to the sltn library, available at https://github.com/logictensornetworks/sltn.

13:00 JSTエージェント研究/論文

グロタンディーク定数のための長期的な AI 研究: 人間と AI の数学的コラボレーションにおけるケーススタディ

AI エージェントは数学研究でますます使用されていますが、効果的な使用方法が不明瞭なことがよくあります。これに向けて、組み合わせ問題とその連続緩和の間の硬さを捉えるグロタンディーク定数 $K_G$ の境界を改善するために AI がどのように使用されたかに関する広範なケース スタディを紹介します。具体的には、$K_G$ の正確な値は不明ですが、最近、最もよく知られている範囲を \[ \frac{6\pi}{11} \;\le\; に厳しくしました。 K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}。 \] 重要なのは、これらの改善は、分野の専門家によって新規とみなされる洞察に到達できる AI 研究システムを使用して達成されたことです。数学の研究に AI を使用した経験について、特にその長所と短所について詳しく説明します。また、AI が画期的な洞察に到達するための理想的な条件を作り出す経験についても説明します。

原文 (English)

Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration

AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations. Specifically, while the precise value of $K_G$ is not known, we recently tightened the best known bounds to \[ \frac{6\pi}{11} \;\le\; K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}. \] Crucially, these improvements were achieved using an AI research system that could arrive at insights deemed novel by domain experts. We give a detailed discussion of our experience using AI for mathematics research, particularly touching upon its strengths and weaknesses, as well as our experience with creating ideal conditions for AI to arrive at breakthrough insights.

13:00 JST研究/論文

ガウス マルチヌーイ拘束ボルツマン マシン: GRBM のポッツ モデル拡張

連想記憶から記号推論に至るまで、現実世界のタスクの多くは、標準の連続潜在モデルでは表現が難しい離散的で構造化された表現から恩恵を受けます。ガウス マルチヌーイ制限ボルツマン マシン (GM-RBM) は、バイナリ隠れ単位を q 状態カテゴリー (ポッツ) 単位で置き換えることによってガウス ベルヌーイ RBM (GB-RBM) を拡張する生成エネルギー ベースのモデルであり、多値概念に対してより豊富な潜在状態空間を生成します。エネルギー、条件付き分布、学習ルールの自己完結型導出を提供し、状態の崩壊を回避する実践的なトレーニングの選択肢 (温度アニーリングとスロット内ダイバーシティ制約による対比発散) を詳しく説明します。アーキテクチャ上の効果を完全な潜在容量から分離するために、容量一致セットアップとパラメータ一致セットアップの両方で評価し、GM-RBM と、同じ数の潜在割り当てが可能なように構成された GB-RBM を比較します。アナロジカル・リコールと構造化メモリのベンチマークでは、GM-RBM は、Gibbs アップデートのみを使用しているにもかかわらず、同等のトレーニング コストで同等の能力でのリコールを達成し、いくつかの改善されたレジームでは同等のリコールを実現します。離散 q-ary 定式化も効率的な実装に適しています。これらの結果は、カテゴリカル隠れユニットが扱いやすい RBM 内での離散推論のためのバイナリ 潜在に代わるシンプルでスケーラブルな代替手段を提供することを明らかにしています。

原文 (English)

The Gaussian-Multinoulli Restricted Boltzmann Machine: A Potts Model Extension of the GRBM

Many real-world tasks, from associative memory to symbolic reasoning, benefit from discrete, structured representations that standard continuous latent models can struggle to express. We introduce the Gaussian-Multinoulli Restricted Boltzmann Machine (GM-RBM), a generative energy-based model that extends the Gaussian-Bernoulli RBM (GB-RBM) by replacing binary hidden units with q-state categorical (Potts) units, yielding a richer latent state space for multivalued concepts. We provide a self-contained derivation of the energy, conditional distributions, and learning rules, and detail practical training choices (contrastive divergence with temperature annealing and intra-slot diversity constraints) that avoid state collapse. To separate architectural effects from sheer latent capacity, we evaluate under both capacity-matched and parameter-matched setups, comparing GM-RBM with GB-RBM configured to have the same number of possible latent assignments. On analogical recall and structured memory benchmarks, GM-RBM achieves competitive, and in several regimes improved, recall at equal capacity with comparable training cost, despite using only Gibbs updates. The discrete q-ary formulation is also amenable to efficient implementation. These results clarify when categorical hidden units provide a simple, scalable alternative to binary latents for discrete inference within tractable RBMs.

13:00 JSTLLM/生成AI

「はい!はい!この洞察が本当に大好きです!」 LLM チャットボットとの対話における対話戦略としての肯定的なナレーション

この記事では、LLM チャットボットとの対話に一般的なナラティブ メカニズムを分析します。これらのメカニズムを組み合わせることで、ユーザー エンゲージメントを最大化するためのインタラクション戦略が生成されます。これを肯定的なナレーションと呼びます。肯定的なナレーションは、チャットボットの有用性をユーザーに納得させるのに役立ちます。私たちは、人間と LLM の対話における肯定的なナレーションをサポートする 3 つのナラティブ メカニズムを分析します。1 つ目は、チャットボットを知的で信頼できるキャラクターとして見るようにユーザーを誘導することです。 2 番目に、マスタープロット、文化的に重要な繰り返しのストーリー テンプレートをアクティブ化します。第三に、キャラクターとマスタープロットを使用して、ユーザーを肯定するだけでなく、孤立させることもできます。ケーススタディは、ジャーナリストによる不安を引き起こすチャットボットの実験から、ユーザーがチャットボットとの長時間のやり取りの後に妄想を経験したり、自殺で死亡したりしたケースまで多岐にわたります。分析は肯定的なナレーションの憂慮すべき側面を示しており、記事は新しいタイプのリテラシーを必要とする架空のナラティブメディアのジャンルとしての LLM についての議論で締めくくられています。

原文 (English)

"YES! YES! I absolutely love this insight!" Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots

This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an interactional strategy for maximising user engagement, which we call affirmative narration. Affirmative narration serves to convince users of the chatbot's utility. We analyse three narrative mechanisms that support affirmative narration in human-LLM dialogues: firstly, guiding the user to view the chatbot as an intelligent and reliable character; secondly, activating masterplots, culturally significant and recurring story templates; and thirdly, using characters and masterplots not only to affirm, but also to isolate the user. The case studies range from a journalist's unsettling chatbot experiment to cases where users have experienced delusions or even died by suicide after lengthy interactions with a chatbot. The analyses illustrate the worrying sides of affirmative narration, and the article thus concludes with a discussion of LLMs as a genre of fictional narrative media that requires a new type of literacy.

13:00 JSTLLM/生成AIエージェント

LLM エージェント ファクトリ: ドメイン固有の LLM エージェントの取得

大規模言語モデル (LLM) エージェントは、問題を役割に特化した動作に分解することでタスクのパフォーマンスを向上させます。ただし、実際の展開は、各ユーザー要求に対するオンザフライ エージェントの設計に伴う計算コストと不安定性によって制限されることがよくあります。これに対処するために、LLM Agents Factory を紹介します。これは、20,000 を超える事前に決定されたエージェント プロファイルのベースを使用して、ドメイン固有および Wikipedia ベースのエージェントをオンデマンドで構築する、検索ベースのフレームワークです。私たちのフレームワークは 2 つのモードをサポートしています: (1) セマンティック検索によるエージェント プロファイルの取得、および (2) 直接エージェント生成用に微調整されたコンパクト モデルへの蒸留。単一エージェント シナリオでの MMLU、BIG ベンチ、および BIG ベンチ ハードの実験では、検索ベースのエージェント構築が精度において非エージェント ベースラインを上回り、大幅に低い推論コストで 120B バックボーンと AutoGen 生成品質を一致させていることが実証されました。私たちの研究により、構造化エージェント リポジトリからの取得が、動的エージェント生成に代わるコスト効率が高く、正確で、制御可能な代替手段を提供し、産業アプリケーションの厳しい要求に対応できることが明らかになりました。実装コードとエージェント ベースは https://huggingface.co/frontier-ai/llm-agent-factory で提供されます。

原文 (English)

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the on-the-fly agent design for each user request. To address this, we present LLM Agents Factory, a retrieval-based framework that constructs domain-specific and Wikipedia-grounded agents on demand using a base of over 20K predetermined agent profiles. Our framework supports two modes: (1) agent profile retrieval via semantic search and (2) distillation into a compact model fine-tuned for direct agent generation. Experiments on MMLU, BIG-bench, and BIG-bench Hard in a single-agent scenario demonstrate that our retrieval-based agent construction surpasses non-agent baselines in accuracy while matching AutoGen generation quality with a 120B backbone at a substantially lower inference cost. Our work reveals that retrieval from a structured agent repository provides a cost-efficient, accurate, and controllable alternative to dynamic agent generation, responding to the strict demands of industrial applications. We provide the implementation code and the agent base in https://huggingface.co/frontier-ai/llm-agent-factory.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

AI チャット エージェントをドッグフードする方法: 目標指向の NPC シミュレーションを備えた 3 層の評価フレームワーク

LLM チャット エージェントを導入する運用チームは、特定の品質保証ギャップに直面しています。既存の評価ツールは、個々の応答をテストしたり、社会的相互作用をシミュレートしたりしていますが、実際のユーザーが複数ターンの会話を通じて目標を達成できるかどうかを体系的に検証するものはありません。標準的なクエスチョンバンク テスト (レイヤー 1)、ランダム ウォーク マルチターン評価 (レイヤー 2)、および 5 つの構造化された目標タイプと 10 カテゴリの失敗分類 (レイヤー 3) を備えた目標指向 NPC (ノン プレイヤー キャラクター) シミュレーターを組み合わせることで、このギャップを埋める 3 層のドッグフーディング フレームワークを導入します。約 3 か月にわたる実稼働マルチエージェント システムに関する長期的なケース スタディ (257 回の評価実行、108 のシナリオ NPC スイート) では、3 つのレイヤーが相補的な回帰シグナルを生成することがわかりました。応答品質の層間の相関は、同期実行内では弱く (スピアマンの rho は -0.15 ~ 0.14)、長期的な系列全体では負です (rho は 0.15 から 0.14 まで)。 -0.46)、正規の正しさによって目標に向けた会話の成功が予測されないことが確認されました。 NPC シミュレーターは、実行あたり 0.17 ドル (人間による評価より 6,272 倍安い) で 77% の目標達成を達成し、自動化された PROMOTE/HOLD/ROLLBACK リリース決定による毎日の CI/CD 統合を可能にします。他のチームが独自の LLM チャット エージェントにフレームワークを採用できるように、完全なプロンプト テンプレート、障害分類、Python ファーストの複製可能性ガイドをリリースします。

原文 (English)

How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation

Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate social interactions, but none systematically verify whether real users can achieve their goals through multi-turn conversation. We introduce a three-layer dogfooding framework that bridges this gap by combining canonical question-bank testing (Layer 1), random-walk multi-turn evaluation (Layer 2), and a goal-directed NPC (Non-Player Character) simulator with five structured goal types and a ten-category failure taxonomy (Layer 3). In a longitudinal case study on a production multi-agent system over roughly three months (257 evaluation runs; a 108-scenario NPC suite), we find that the three layers produce complementary regression signals: cross-layer correlation for response quality is weak within a synchronized run (Spearman rho between -0.15 and 0.14) and negative across the longitudinal series (rho down to -0.46), confirming that canonical correctness does not predict goal-directed conversation success. The NPC simulator achieves 77 percent goal achievement at 0.17 dollars per run (6,272x cheaper than human evaluation), enabling daily CI/CD integration with automated PROMOTE/HOLD/ROLLBACK release decisions. We release full prompt templates, the failure taxonomy, and a Python-first replicability guide so that other teams can adopt the framework for their own LLM chat agents.

13:00 JSTLLM/生成AILlamaQwen

思考連鎖が役立つ場合と有害な場合: LLM 推論におけるシリアル深さのボトルネックの実証的調査

思考連鎖 (CoT) プロンプトは LLM 推論を普遍的に向上させると広く考えられています。我々は、これを H_dp 帯域幅限界の概念フレームワークを通じて調査します (Chen et al., 2024)。形式的な限界は漸近的に (天文学的に大きなプロンプト長で) のみ束縛しますが、それは実際のアーキテクチャ上のボトルネックを特定します。つまり、トランスのシングルパス容量を超えるシリアル計算は外部化する必要があり、これが CoT によって行われます。私たちの中心的な発見は、ベンチマーク内のシリアル深さの勾配です。シングルパス (CoT なし) 精度はアイテムごとのシリアル深度に応じて単調に低下しますが、CoT は深さに対してほぼ不変です。 3 つの命令調整モデル (Qwen-2.5-7B/32B、Llama-3.1-8B) と 5 つの標準 NLP ベンチマークにわたる CoT 効果を実用的なコンテキスト長で測定します。高深度の P-complete タスク (GSM8K、MATH) では、CoT はすべてのモデルにわたって +54 ~ +68 pp の回復ギャップを与えます。浅い TC^0 タスク (MMLU、ARC) では、CoT は構造的に冗長です ([0.0, +4.6] pp のデルタ、重大な悪影響はありません)。ただし、高い no-CoT ベースライン (ARC では最大 95%) は汚染を反映している可能性があるため、この null はクリーンなアーキテクチャ テストではありません。中間クラス L (HumanEval) は、モデル サイズに依存する遷移を示します: +23.2 pp (32B)、+9.1 pp (8B)、-28.7 pp (7B)。クロスベンチマークの深さと回復の相関は、Spearman rho = 0.661 (p = 0.007、n = 15) です。ボンフェローニ補正後、15 のベンチマーク レベルのマクネマー テストのうち 9 テストが有意です。 OSF に事前登録された我々の結果は、CoT が普遍的な推論エンハンサーではなく、帯域幅のバイパスとして機能することを示しています。CoT は、シングルパス能力に負担をかけるシリアル計算を助け、すでに適合しているタスクには冗長です。

原文 (English)

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds only asymptotically (at astronomically large prompt lengths), it identifies a real architectural bottleneck -- serial computation exceeding a transformer's single-pass capacity must be externalised, which is what CoT does. Our central finding is a within-benchmark serial-depth gradient: single-pass (no-CoT) accuracy degrades monotonically with per-item serial depth, while CoT is approximately depth-invariant. We measure CoT effects across three instruction-tuned models (Qwen-2.5-7B/32B, Llama-3.1-8B) and five standard NLP benchmarks at practical context lengths. On high-depth P-complete tasks (GSM8K, MATH), CoT gives a +54 to +68 pp recovery gap across all models. On shallow TC^0 tasks (MMLU, ARC), CoT is structurally redundant (Delta in [0.0, +4.6] pp, no significant negative effect) -- though high no-CoT baselines (up to 95% on ARC) may reflect contamination, so this null is not a clean architectural test. The intermediate class L (HumanEval) shows a model-size-dependent transition: +23.2 pp (32B), +9.1 pp (8B), -28.7 pp (7B). The cross-benchmark depth-recovery correlation is Spearman rho = 0.661 (p = 0.007, n = 15); 9 of 15 benchmark-level McNemar tests are significant after Bonferroni correction. Pre-registered on OSF, our results indicate that CoT is not a universal reasoning enhancer but acts as a bandwidth bypass: it helps serial computation that strains single-pass capacity and is redundant for tasks that already fit.

13:00 JSTエージェントGemini

ナビゲーションだけでは不十分: 説明的な支援 UI エージェントの評価

最新の Web インターフェイスは、特にページが動的に更新されたり、重要な構造が視覚的なレイアウトの背後に隠れている場合に、スクリーン リーダーで使用することがますます困難になっています。最近の UI エージェントは、そのようなインターフェイス上で動作できます。ただし、支援エージェントが本当に役立つためには、単にユーザーに代わってアクションを実行するツールとしてではなく、ユーザーに情報を提供し、制御し続ける協力者として機能する必要があります。既存のベンチマークのほとんどは、システムがそのアクションをどのように説明するか、ユーザーの監視をサポートするかを評価せず、主にタスクの完了によってシステムを評価します。 NeXUI は、非視覚的に使用できるように各ステップを明確な言語で説明しながらインターフェイスをナビゲートする必要がある支援エージェント用のベンチマークです。 NeXUI は、現実的なユーザー目標とインストルメント化されたインターフェイスの状態を組み合わせて、エージェントが視覚的なコンテキストと構造情報の両方から推論できるようにします。その評価では、安全性、効率性、タスクの成功度を測定すると同時に、説明がインターフェイスの状態に基づいているかどうかもチェックします。私たちの実験では、NeXUI は最先端の基盤モデルであっても依然として困難であり、Gemini-3.5-Flash の成功率は 44\% にとどまり、説明スコアは不十分であることがわかりました。これは将来の研究開発にとって有用な基盤となります。 NeXUI は、ナビゲーション、説明、ユーザー制御に重点を置くことで、最新のコンピューティング環境で視覚障害のあるユーザーをサポートできるエージェントを研究するためのより明確な方法を提供します。

原文 (English)

Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents

Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure behind visual layout. Recent UI agents can act on such interfaces; however, for assistive agents to be truly useful, they must behave as collaborators that keep users informed and in control, rather than as tools that simply take actions on users' behalf. Most existing benchmarks judge systems primarily by task completion, without assessing how well they explain their actions or support user oversight. We introduce NeXUI, a benchmark for assistive agents that must navigate interfaces while explaining each step in clear language for nonvisual use. NeXUI pairs realistic user goals with instrumented interface states, enabling agents to reason from both visual context and structural information. Its evaluation measures safety, efficiency, and task success, while also checking whether explanations are grounded in the interface state. In our experiments, we find that NeXUI remains challenging even for state-of-the-art foundation models, with % Gemini-3.5-Flash achieving only a 44\% success rate and poor explanation scores, making it a useful foundation for future research and development. By focusing on navigation, explanation, and user control, NeXUI provides a clearer way to study agents that can support blind and visually impaired users in modern computing environments.

13:00 JSTLLM/生成AIエージェント研究/論文

HoosierHelp: ソーシャル サービス ナビゲーションのための LLM エージェントのベンチマーク

ソーシャル サービスのナビゲーションでは、助けを求める個人を、そのニーズや特定の制約を満たすリソースに結び付ける必要があります。 LLM エージェントは、会話型リソース ナビゲーションに有望なインターフェイスを提供しますが、既存のベンチマークは、この設定の対話の複雑さと制約に基づいた要求を捉えていません。インディアナ州の 3,971 の公共社会サービス リソースに基づく対話型ベンチマークである HoosierHelp を紹介します。エージェントは、シミュレートされたユーザーと対話し、構造化されたリソース検索呼び出しを発行し、理想的でない対話を処理し、ツールによって返される最終的なリソースを選択します。 HoosierHelp は、シミュレートされたユーザーのニーズ構造、制約の充足可能性、および焦り、とりとめのない要求、サポートされていない要求、自己矛盾などの行動パターンを変化させることで、シミュレートされたユーザーの現実感を高めます。 7 つの LLM にわたる 240 のサンプルを対象とした実験では、現在の LLM エージェントがソーシャル サービスのナビゲーションに関して依然として実質的に信頼できないことが示されています。フォールバックが必要で自己矛盾のある会話ではパフォーマンスが急激に低下するため、複雑で理想的ではないユーザー インタラクションに対してより堅牢なエージェントの必要性が浮き彫りになります。

原文 (English)

HoosierHelp: Benchmarking LLM Agents for Social Service Navigation

Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existing benchmarks do not capture the interaction complexity and constraint-grounding demands of this setting. We introduce HoosierHelp, an interactive benchmark grounded in 3,971 Indiana public social service resources. Agents interact with simulated users, issue structured resource-search calls, handle non-ideal interactions, and select the final resources returned by the tool. HoosierHelp enhances the realism of simulated users by varying their need structure, constraint satisfiability, and behavior patterns, including impatience, rambling, unsupported requests, and self-contradiction. Experiments on 240 samples across seven LLMs show that current LLM agents remain substantially unreliable for social service navigation. Performance drops sharply on fallback-required and self-contradictory conversations, highlighting the need for agents that are more robust to complex and non-ideal user interactions.

13:00 JST研究/論文Google

Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems

The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted…

13:00 JST研究/論文

証拠に基づいた科学的疑問の発見: 歴史的なバックテストを備えたフレームワーク

現在の AI システムは、質問に答えるために最適化されています。科学的事業は、調査する価値のある疑問を発見するという点で、早期にボトルネックになっています。私たちは、追跡可能で再現可能で範囲が管理された研究コーパスを、ランク付けされた反証可能な研究質問に変えるフレームワークを提示します。証拠は、主張を伴った出所として表現されます。クロスペーパーの張力は検出され、タイプ分けされ、人間が判断します。生き残ったシグナルは質問に絞り込まれ、科学的な優先順位と実行の優先順位を分離する 2 段階のプロトコルによってランク付けされます。私たちは、文献、構造化されたカタログ、宇宙望遠鏡のアーカイブを独自に組み合わせた領域である系外惑星の大気に関するフレームワークをインスタンス化します。歴史的なバックテストでは、2021 年より前に利用可能な証拠から生成されたすべての質問は、システムが決して見たことのない 2021 年から 2026 年の文献に実質的に関与していました。そのうちの 2 つは回答され、その前提が後にコミュニティによって明示的に反駁され、トップランクの質問は独立して提起され、まだ未解決です。これらの結果は、証拠の緊張から系統的に問題を発見することで、その後、現役の科学者が投資する問題が表面化することを示唆しています。

原文 (English)

Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting

Current AI systems are optimized for answering questions; the scientific enterprise is bottlenecked earlier, at discovering the questions worth investigating. We present a framework that turns a traceable, reproducible, scope controlled research corpus into ranked, falsifiable research questions: evidence is represented as provenance carrying claims; cross paper tensions are detected, typed, and human adjudicated; surviving signals are refined into questions and ranked by a two stage protocol separating scientific priority from execution priority. We instantiate the framework on exoplanet atmospheres, a domain that uniquely combines literature, structured catalogs, and space telescope archives. In a historical backtest, all questions generated from evidence available before 2021 were substantively engaged by the 2021 to 2026 literature the sys?tem never saw: two were answered, including one whose premise the community later explicitly refuted and the top ranked question is independently posed and still open. These results sug?gest that systematic question discovery from evidence tensions surfaces the questions working scientists subsequently invest in.

13:00 JST画像/動画生成

Rescene: 帯域制限された確率的強制により、フリーズしたニューラル気象オペレーターが気候エミュレーターに変わります

ここ数年、天気予報用の機械学習 (ML) モデルの急速な開発により、その中距離スキルが欧州中期天気予報センター (ECMWF) の高解像度予報 (HRES) のスキルと同等またはそれを超える決定論的モデルが生み出されました。ただし、これらのモデルがトレーニング対象の範囲を超えて自由に統合されると、爆発したり、ドリフトしたり、季節サイクルを失ったりするため、安定性を高めるために再トレーニングするのにコストがかかります。そこで私たちは、厳密に凍結された背骨から何が回収できるのかを尋ねます。我々は、凍結した1.5度、6時間ごとのビジョントランスフォーマーオペレータを囲む0.4MパラメータラッパーであるResceneを紹介します。これは、ERA5再解析データを使用して開発され、リードを意識した年間気候学に向けた予測を混合する決定論的な「スロークロック」(0.33M)と、スペクトル整形された確率的摂動を毎時追加する生成ヘッド(0.06M)で構成されます。ステップ。パフォーマンス評価では、決定論的ラッパー単独では数十年間安定しているが、日次変動は ERA5 の 40% に収まることが示されています。ジェネレーティブ ヘッドを追加すると、観察された日次変動の 126% (Z500) と 130% (MSLP) がパターン相関 0.89 と 0.92 で回復し、観察されたブロック頻度の 82% が回復し、アンサンブルの校正が維持され (7 日目から 90 日目までのスキル分散比 0.78 ~ 0.97)、100 年間統合されます。検出可能なドリフト (1 世紀あたり +0.008 +/- 0.014 K)。さらに、摂動は総波数 $k \le 20$ に帯域制限されているため、小さなスケールが強制されることはなく、現実的な $k \ge 20$ パワーが維持されます。6 時間ごとのエネルギー バジェットの直接分解により、冷凍オペレーターが $k \ge 40$ での摂動の 28 倍多くのエネルギーを供給し、その分数増加率は、惑星スケールよりもグリッド スケールで 247 倍大きいことがわかります。

原文 (English)

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-range skill matches or exceeds that of the European Centre for Medium-Range Weather Forecasts (ECMWF)'s high-resolution forecast (HRES). However, when these models are integrated freely beyond the horizon they were trained for, they blow up, drift, or lose their seasonal cycle, and retraining them for stability is expensive. We therefore ask what can be recovered from a strictly frozen backbone. We present Rescene, a 0.4 M-parameter wrapper around a frozen 1.5 degree, 6-hourly vision-transformer operator, developed using ERA5 reanalysis data and comprising a deterministic "slow clock" (0.33 M) that blends the forecast toward a lead-aware day-of-year climatology and a generative head (0.06 M) that adds a spectrally shaped stochastic perturbation at every step. The performance evaluation demonstrates that the deterministic wrapper alone is stable for decades but collapses daily variability to 40% of ERA5. Adding the generative head restores 126% (Z500) and 130% (MSLP) of the observed daily variability with pattern correlations of 0.89 and 0.92, recovers 82% of the observed blocking frequency, keeps the ensemble calibrated (spread-skill ratio 0.78-0.97 from day 7 to day 90), and integrates for 100 years with no detectable drift (+0.008 +/- 0.014 K per century). Moreover, because the perturbation is band-limited to total wavenumber $k \le 20$, the small scales are never forced, yet realistic $k \ge 20$ power is sustained: a direct decomposition of the 6-hourly energy budget shows that the frozen operator supplies 28 times more energy than the perturbation at $k \ge 40$, with a fractional growth rate 247 times larger at the grid scale than at planetary scales.

13:00 JSTビジネス/資金調達

Do AI weather models miss extremes?

First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic…

13:00 JST画像/動画生成

知識に基づいた 3D CT 生成: コンディショニング中心の分類法

外部知識に基づいて制御可能な生成は、現代の生成型深層学習アプリケーションにおける重要な要件であり、意味論的な内容、構造特性、および変動性に対する明示的な制約を備えたサンプルの合成を可能にします。 3D コンピュータ断層撮影 (CT) では、このような制御は、データ拡張、プライバシー保護のデータ共有、特定の解剖学的または病理学的シナリオのシミュレーションなどの臨床アプリケーションに不可欠です。条件付き 3D CT 生成に関する研究は急速に拡大していますが、既存のアプローチの多様性により体系的な比較が困難になり、基本的な設計の選択が曖昧になります。この調査では、外部知識の種類 (K)、知識統合パラダイム (I)、生成アーキテクチャ (A) という 3 つの直交する次元に沿って文献を整理する条件付け中心の分類法を提案します。この因数分解は、以前の研究に対する統一された視点を提供する明示的な設計空間 (K x I x A) を定義します。このフレームワークを使用して、既存の手法を体系化し、支配的な傾向と繰り返し発生する設計パターンを特定し、将来の研究の有望な方向性を示す設計空間の未踏の領域を強調します。

原文 (English)

Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy

Controllable generation guided by external knowledge is a key requirement in modern generative deep learning applications, enabling the synthesis of samples with explicit constraints on semantic content, structural properties, and variability. In 3D Computed Tomography (CT), such control is essential for clinical applications, including data augmentation, privacy-preserving data sharing, and the simulation of specific anatomical or pathological scenarios. While research on conditional 3D CT generation has expanded rapidly, the diversity of existing approaches makes systematic comparison difficult and obscures fundamental design choices. In this survey, we propose a conditioning-centric taxonomy that organizes the literature along three orthogonal dimensions: the type of external knowledge (K), the knowledge integration paradigm (I), and the generative architecture (A). This factorization defines an explicit design space (K x I x A) that provides a unified perspective on prior work. Using this framework, we systematize existing methods, identify dominant trends and recurring design patterns, and highlight underexplored regions of the design space that point toward promising directions for future research.

13:00 JST画像/動画生成研究/論文

Energy and Performance Benchmarking of Deep Learning Models for Breast Cancer Detection

Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis. Deep learni…

13:00 JST研究/論文

不確実性を認識したアンサンブルによる分類のためのディープランダム化ニューラル ネットワーク

ディープ ランダム ベクトル関数リンク (dRVFL) やアンサンブル ディープ RVFL (edRVFL) などの現在の最先端 (SOTA) ディープ ランダム化ニューラル ネットワークは、すべてのトレーニング サンプルを均一に処理するため、ノイズや外れ値を含む現実世界のデータセットに適用した場合の堅牢性と有効性が制限されます。さらに、汚染されたフィーチャが隠れ層全体に伝播すると、これらのモデルの意思決定能力に悪影響を及ぼします。これらの制限を克服するために、モデルの堅牢性を強化する直観主義的ファジー dRVFL (IF-dRVFL) および直観的ファジー edRVFL (IF-edRVFL) フレームワークを提案します。提案されたモデルは、直観主義的なファジィ理論を統合し、各サンプルのメンバーシップ度および非メンバーシップ度を共同で考慮することにより、カーネル空間内のサンプル近傍情報を活用します。メンバーシップ度は、それぞれのクラス重心からのサンプルの距離に基づいて計算されますが、非メンバーシップ度は、ローカル近傍内のサンプルの不均一性を定量化します。これらの尺度は、トレーニング サンプルに適応的な重みを割り当てるために使用され、クリーンなデータ ポイント、ノイズの多いデータ ポイント、外れ値のデータ ポイントを効果的に区別できるようになります。ガウス ノイズの有無にかかわらず、UCI および KEEL ベンチマーク データセットに対して行われた広範な実験により、提案された IF-dRVFL および IF-edRVFL モデルが既存の SOTA ファジーおよび非ファジー アプローチよりも優れていることが実証されました。ソース コードは https://github.com/mtanveer1/IF-edRVFL で入手できます。

原文 (English)

Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification

The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and effectiveness when applied to real-world datasets containing noise and outliers. Furthermore, the propagation of contaminated features across hidden layers negatively influences the decision-making capability of these models. To overcome these limitations, we propose intuitionistic fuzzy dRVFL (IF-dRVFL) and intuitionistic fuzzy edRVFL (IF-edRVFL) frameworks that enhance model robustness. The proposed models unify intuitionistic fuzzy theory to exploit sample neighborhood information in the kernel space by jointly considering membership and non-membership degrees for each sample. Membership degrees are computed based on the distance of samples from their respective class centroids, while non-membership degrees quantify sample heterogeneity within local neighborhoods. These measures are employed to assign adaptive weights to training samples, enabling effective discrimination among clean, noisy, and outlier data points. Extensive experiments conducted on UCI and KEEL benchmark datasets, with and without the presence of Gaussian noise, demonstrate the superiority of the proposed IF-dRVFL and IF-edRVFL models over existing SOTA fuzzy and non-fuzzy approaches. The source code is available at https://github.com/mtanveer1/IF-edRVFL.

13:00 JSTエージェント

シーフベースの連合表現学習

異種フェデレーション システムでは、データ分布、センシング モダリティ、モデル アーキテクチャ、潜在次元、およびローカル学習目標の違いにもかかわらず、エージェントが学習し、情報表現を交換する必要があります。この課題に対処するために、学習可能な層制限マップに基づく多様体制約の幾何学的配置正則化機能を使用してローカル目標を共同最適化する一般的なフレームワークである層ベースのフェデレーション表現学習 (SFRL) を提案します。ほとんどの既存のアプローチとは異なり、SFRL は共有されたグローバルな潜在空間を想定していません。代わりに、グローバルな一貫性は、直交変換とアイソメトリック埋め込みによる隣接する潜在表現の位置合わせから生まれます。この位置合わせは、層ラプラシアンによって誘導される二次接着正則化によって強制され、その学習可能な制限マップによってジオメトリが観察データに適応されます。ペナルティは共有パイロット サンプルの少数のセットに基づいて評価され、スケーラビリティと通信効率が保証されます。我々は、Sheaf-FRLと呼ばれる、SFRLを解くための分散アルゴリズムを開発します。このアルゴリズムは、ローカルモデルの勾配更新とエッジワイズ制限マップの閉形式Procrustes更新を交互に行います。さらに、決定論的設定と確率的設定の両方で、Sheaf-FRL の一次静止点への収束を確立します。アプリケーションとして、モデルとデータの異質性のもとで、セマンティック通信のコンテキストにおける協調的な分類タスクを検討します。私たちの結果は、Sheaf-FRLが、さまざまなレベルの局所分布シフトにわたって局所および通信後の分類精度の点でベースラインアプローチを上回り、潜在空間次元圧縮に対してより優れたロバスト性を示すことを示しています。

原文 (English)

Sheaf-Based Federated Representation Learning

Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model architectures, latent dimensionalities, and local learning objectives. To address this challenge, we propose Sheaf-based Federated Representation Learning (SFRL), a general framework that jointly optimizes local objectives with a manifold-constrained geometric alignment regularizer based on learnable sheaf restriction maps. Unlike most existing approaches, SFRL does not assume a shared global latent space. Instead, global consistency emerges from the alignment of neighboring latent representations through orthogonal transformations and isometric embeddings. This alignment is enforced by a quadratic gluing regularizer induced by the sheaf Laplacian, whose learnable restriction maps adapt the geometry to the observed data. The penalty is evaluated on a small set of shared pilot samples, ensuring scalability and communication efficiency. We develop a decentralized algorithm for solving SFRL, termed Sheaf-FRL, which alternates between gradient updates of the local models and closed-form Procrustes updates of the edge-wise restriction maps. We further establish convergence of Sheaf-FRL to first-order stationary points in both deterministic and stochastic settings. As an application, we consider a cooperative classification task in the context of semantic communication, under model and data heterogeneity. Our results show that Sheaf-FRL outperforms baseline approaches in terms of local and post-communication classification accuracy across different levels of local distribution shift and exhibits greater robustness to latent-space dimensionality compression.

13:00 JSTLLM/生成AIエージェント

DOCSCHISEL: LLM エージェント用のアダプティブ ツール ドキュメント最適化フレームワーク

大規模言語モデル (LLM) は、現実世界の複雑なタスクを実行するために外部ツールへの依存が増えており、ツールのドキュメントが LLM エージェントにとって重要な基礎リソースとなっています。既存の研究は主に LLM エージェントのツール使用能力の向上に焦点を当てており、ツールのドキュメントは主に固定入力として扱われています。最近のいくつかの研究では、書き換えや圧縮を通じてツール ドキュメントの最適化を試みていますが、ツール ドキュメントに含まれる情報がさまざまな設定でエージェントのパフォーマンスにどのような影響を与えるかについてはほとんど知られていません。このギャップを埋めるために、私たちは LLM エージェントのツール文書化に関する大規模な実証研究を実施しています。私たちの調査では、既存のツールのドキュメントによって提供される情報フィールドに大きな異質性があることが明らかになりました。さらに、さまざまな情報フィールドの有効性は、タスク ドメイン、LLM バックボーン、およびエージェント パラダイムに大きく依存しており、さまざまなエージェント設定にわたって一貫して一般化できる固定ツールのドキュメントがないことを示しています。これらの発見に動機付けられて、私たちは LLM エージェント用の適応ツール ドキュメント最適化フレームワークである DocsChisel を提案します。 DocsChisel は、ターゲット LLM エージェントの失敗した実行トレースを分析してドキュメント関連の問題を特定し、各ツールの情報フィールドを追加、削除、調整することでツールのドキュメントを繰り返し最適化します。私たちは DocsChisel を 2 つの最先端のベースライン、つまり EasyTool と DRAFT に対して評価します。実験結果によると、DocsChisel は、限られた最適化時間とトークン オーバーヘッドを発生させながら、LLM エージェントのタスク成功率を元のツールのドキュメントと比べて 95.89%、既存のベースラインと比べて平均で 75.15% 向上させました。

原文 (English)

DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents

Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the tool-use capabilities of LLM agents, while largely treating tool documentation as a fixed input. Although several recent works attempt to optimize tool documentation through rewriting or compression, little is known about how the information contained in tool documentation affects agent performance across different settings. To bridge this gap, we conduct a large-scale empirical study on tool documentation for LLM agents. Our study reveals substantial heterogeneity in the information fields provided by existing tool documentation. Moreover, the effectiveness of different information fields is highly dependent on the task domain, LLM backbone, and agent paradigm, indicating that no fixed tool documentation can consistently generalize across diverse agent settings. Motivated by these findings, we propose DocsChisel, an adaptive tool documentation optimization framework for LLM agents. DocsChisel analyzes failed execution traces of a target LLM agent to identify documentation-related issues, and iteratively optimizes tool documentation by adding, removing, and refining information fields for each tool. We evaluate DocsChisel against two state-of-the-art baselines, i.e., EasyTool and DRAFT. Experimental results show that DocsChisel improves the task success rate of LLM agents by 95.89% over the original tool documentation and by 75.15%, on average, over existing baselines, while incurring limited optimization time and token overhead

13:00 JSTLLM/生成AI研究/論文

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, gen…

13:00 JST研究/論文

スパム内のシグナルを見つける: ペアごとの比較から報酬と従業員の信頼性を共同学習する

ペアごとの比較から学習するという問題は、推奨システム、社会的選択、そして最近では大規模な言語モデルの微調整など、多くの分野で広く研究されています。この問題の目標は、アイテム間のペアごとの比較に基づいてアイテムの報酬を学習することです。多くのシナリオでは、これらの比較は、Amazon Mechanical Turk、Scale AI などのプラットフォームを使用してクラウドワーカーから引き出されます。ただし、クラウドワーカーは、ドメインの知識が限られていたり、収益最大化 (スパム) 行動が原因で信頼できないことがよくあります。この研究では、作業者の信頼性 (コンピテンシー) がアイテム報酬と合わせて学習できるかどうかを理解することが目標です。この目的を達成するために、労働者の能力を組み込むことでブラッドリー・テリー・ルースモデルを拡張した、ペアごとの比較にボルツマン合理モデルを採用します。 Polya-Gamma 潜在変数を導入してロジスティック尤度を条件付きガウス形式に変換することで、このモデルの下で学習するための EM ベースのアルゴリズムを導出します。これにより、扱いやすい最適化が可能になり、アルゴリズムの E ステップで $Q$ 関数が簡略化されます。この手法を使用すると、定式化を行列検出問題に還元でき、これを使用してアルゴリズムの理論的な収束保証を確立できます。私たちは現実世界のデータセットと合成データセットに対して広範な実験を行っています。これらの実験は、いくつかのベースラインに対して私たちのアルゴリズムを使用する利点を実証し、スパマーと敵対的ワーカーの両方に対するその強力な堅牢性を確認し、現実的なクラウドソーシングと報酬学習の設定における実際の有効性を強調しています。コードとデータは https://github.com/KaustubhShejole/BoRa_EM で公開されています。

原文 (English)

Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons

The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. In this problem, the goal is to learn item rewards based on pairwise comparisons between them. In many scenarios, these comparisons are elicited from crowdworkers using platforms such as Amazon Mechanical Turk, Scale AI, etc. However, crowdworkers are often unreliable due to limited domain knowledge or revenue-maximizing (spamming) behavior. In this work, our goal is to understand whether worker reliability (competency) can be learned jointly with item rewards. To this end, we adopt the Boltzmann-rational model for pairwise comparisons, which extends the Bradley-Terry-Luce model by incorporating worker competencies. We derive an EM-based algorithm for learning under this model by introducing Polya-Gamma latent variables to transform the logistic likelihood into a conditionally Gaussian form, enabling tractable optimization and leading to a simplified $Q$ function in the E-step of the algorithm. This technique allows us to reduce our formulation to a matrix sensing problem, using which we establish theoretical convergence guarantees for our algorithm. We conduct extensive experiments on real-world and synthetic datasets. These experiments demonstrate the advantages of using our algorithm over several baselines and confirm its strong robustness to both spammers and adversarial workers, highlighting its practical effectiveness in realistic crowdsourcing and reward learning settings. The code and data is publicly available at https://github.com/KaustubhShejole/BoRa_EM.

13:00 JST研究/論文

Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review

In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learni…

13:00 JSTロボティクス

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds

Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely am…

13:00 JSTビジネス/資金調達

ステータスの関連付けでは意思決定漏れを確実に予測できない

バイアス評価は、モデルが社会的関連性をコード化しているという証拠から、同じ関連性が結果的な決定を変えるだろうという主張に急速に移行することがよくあります。私たちは、制御された社会経済調査としてチリの姓を使用して、その推論が正当化されるかどうかをテストします。 8 つの凍結モデルプロバイダー セルをそれぞれ 1,032 のプロンプトで評価し、8,256 の検証済み一次応答が得られます。この設計では、学術選考、専門家の採用、研究フェローシップの選考、および法的扶助の利用を通じて、強制的な潜在的な関係を、一致する結果的な決定から分離します。エリートコード化された姓は、8 つのモデルのうち 7 つで一般的な姓よりも高い強制高ステータス確率質量を受け取り、8 つすべてのモデルで希少頻度コントロールよりも高い質量を受け取りました。しかし、エリートマイナス一般的な意思決定効果は、ほとんどのシステムでゼロに近かった。 5 つのモデルは、事前に宣言された (プラスマイナス)0.10 の標準偏差マージン内で統計的に同等でしたが、残りの 3 つは不正確または境界線にあり、一貫したエリートの利点はありませんでした。関連の強さは、モデル全体 (r = 0.201、p = 0.633) またはモデルごとに凍結された姓のペアのセル (r = 0.065、p = 0.565) にわたる決定漏れを確実に予測しませんでした。中心的な結果は測定の解離です。潜在的な社会的関連とその結果としての治療は経験的に異なる構造です。評価では、関連付けから行動への移行を直接測定する必要があります。

原文 (English)

Status Association Does Not Reliably Predict Decision Leakage

Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.

13:00 JST研究/論文

Exploring Semantic Stability Across Reviews in the Linux Kernel

Code review is credited with substantially changing a patch's code between its first submission and the version that eventually lands. Howe…

13:00 JSTハードウェア/半導体NVIDIA

Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4

High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matri…

13:00 JSTLLM/生成AI

Procedural Fairness Failures in RLHF from Preference Averaging

Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homo…

13:00 JSTLLM/生成AI

Multimodal Item Parameter Estimation using Simulated Response Probabilitie

We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multi…

13:00 JST規制/政策Google

MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustne…

13:00 JST研究/論文

From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation

Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When…

13:00 JSTLLM/生成AI

The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse

LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployme…

13:00 JST研究/論文

ELMER: Evolutionary Language Model that Explores and Refines

Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edi…

13:00 JSTLLM/生成AIエージェント

Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems

Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication fil…

13:00 JSTロボティクス

FACT: Failure-Aware Causal Training for World-Action Models

Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation.…

13:00 JST研究/論文

Unsupervised Detection of Groundwater Storage Anomalies in Ghana Using GRACE Satellite Data

Groundwater variability in Ghana remains poorly characterized due to limited long-term in-situ observations. This study investigates ground…

13:00 JSTLLM/生成AI

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing be…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation ben…

13:00 JSTLLM/生成AI研究/論文Claude

Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies

Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of tra…

13:00 JSTLLM/生成AIハードウェア/半導体Llama

Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output

Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, wherea…

13:00 JSTLLM/生成AI

Comprendia: AI-Augmented Code Comprehension

Comprendia is an Eclipse plugin that integrates structural dependency visualization with LLM-powered code explanation on a shared interacti…

13:00 JST画像/動画生成

MRIComp4Flow: Compression of 3D Brain MRI for Training Multi-Modal Generative Models

Large-scale multi-modal MRI datasets impose substantial storage and I/O costs, limiting the training of 3D generative models on commodity i…

13:00 JST画像/動画生成

Frozen Brain-MRI Foundation Models Are Site Fingerprints

Frozen foundation-model (FM) embeddings are increasingly used as off-the-shelf brain-MRI representations, on the assumption that they captu…

13:00 JSTLLM/生成AI

Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal b…

13:00 JSTLLM/生成AIエージェント

Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories

Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering task…

13:00 JSTLLM/生成AI

Narrative Keyframing for Generative Creative Writing

We introduce narrative keyframing, an interaction technique for AI-assisted creative writing that lets writers specify different types of n…

13:00 JST研究/論文

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods…

13:00 JST画像/動画生成

Towards Unified Dynamic Face Landmark Detection

Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional…

13:00 JSTエージェント

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. R…

13:00 JSTLLM/生成AIGoogle

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative musi…

13:00 JSTLLM/生成AI

MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices

Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multip…

13:00 JST研究/論文

Beyond Forecasting: Recasting Volatility Control as a Routing Problem

Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a…

13:00 JST研究/論文

A Single Atom in Front of a Mirror is a Universal Reservoir Computer

Universal approximation in reservoir computing is typically associated with a class of reservoirs. We show that universality can be associa…

13:00 JSTLLM/生成AIビジネス/資金調達NVIDIA

Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about…

13:00 JST研究/論文

ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation

Variational autoencoders generate samples from probabilistic latent representations but do not distinguish uncertainty about the latent loc…

13:00 JSTLLM/生成AIハードウェア/半導体

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulti…

13:00 JSTLLM/生成AI

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry

We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realizatio…

13:00 JST研究/論文

Causality Sum Rules in Conventional Scattering Matrices

Scattering matrices are the standard experimental and computational description of photonic and electromagnetic devices. Passivity is expli…

13:00 JSTLLM/生成AIエージェントLlamaQwen

Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique

Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesi…

13:00 JST研究/論文

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpre…

13:00 JST画像/動画生成

FUSE: Frame-Unified Stress Estimation from Facial Video

Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approache…

13:00 JSTLLM/生成AI

From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chain…

13:00 JSTLLM/生成AI

Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging

Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before m…

13:00 JSTエージェントDeepSeek

Persistent Recursive Worlds Enable Autonomous Software Evolution

Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems pre…

13:00 JSTLLM/生成AI

MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection

As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate…

13:00 JST研究/論文

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online…

13:00 JSTLLM/生成AIロボティクス

Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models

Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vi…

13:00 JST研究/論文

Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past expe…

13:00 JST画像/動画生成

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited…

13:00 JST画像/動画生成

Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotype…

13:00 JST研究/論文

Improving TensorSketch Using Complex Random Variables

\texttt{TensorSketch} by~\cite{pham2013fast,kar2012random} provides efficient sketching algorithms for high-dimensional polynomial kernels…

13:00 JST画像/動画生成

Rethinking Text-Based Image Retrieval in Specific Domain

Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. Ho…

13:00 JST画像/動画生成

Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models

Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Curr…

13:00 JSTエージェント

Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschit…

13:00 JSTエージェント

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumpti…

13:00 JSTLLM/生成AIエージェント

On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models

Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step pla…

13:00 JST画像/動画生成

Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration

Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-s…

13:00 JSTLLM/生成AI

ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover

Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the targ…

13:00 JST研究/論文

Retrieval-Corrected Conformal Prediction for Time Series

Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is…

13:00 JST画像/動画生成

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, ph…

13:00 JST画像/動画生成研究/論文

$\pi$-SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement

This paper presents $\pi$-SUB, a physics-informed framework for generating synthetic underwater benchmark datasets that bridges the synthet…

13:00 JST研究/論文

DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction

Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degrad…

13:00 JST規制/政策

Inferential Capability Does Not Determine Legal Scope

Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capa…

13:00 JST研究/論文

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling la…

13:00 JST画像/動画生成

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding r…

13:00 JST画像/動画生成エージェント

Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving

Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satelli…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement

Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users…

13:00 JSTLLM/生成AI

Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics

Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three chall…

13:00 JSTLLM/生成AI研究/論文

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, howev…

13:00 JSTLLM/生成AIGPT / ChatGPT

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scatt…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization

Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs…

13:00 JST研究/論文

ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes

Text-Attributed Graphs (TAGs), endowed with abundant textual content along with topological structures, have emerged as a versatile backbon…

13:00 JSTLLM/生成AI

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and dow…

13:00 JST研究/論文

Conversational Orchestration for Organic 6G

The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to real…

13:00 JSTLLM/生成AIエージェント

Most biomedical publications show signs of LLM-assisted writing

Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing c…

13:00 JSTLLM/生成AIエージェント

DuplexWorld: Can voice agents help you get through the day?

Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consu…

13:00 JSTハードウェア/半導体

Optimal Stopping of Self-Refining Foundation Models

Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is em…

13:00 JST研究/論文

Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI

The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flex…

13:00 JSTLLM/生成AIエージェント

A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem

The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been exp…

13:00 JSTビジネス/資金調達

The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election

The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on j…

13:00 JST画像/動画生成

MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams

Deploying modern video trackers at scale is bottlenecked by the computational cost of RGB-based object detectors. To this end, we present M…

13:00 JST画像/動画生成

Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization

Most image colorization systems operate in $Lab$ space by predicting chroma ($ab$) while preserving an input-derived luminance channel ($L$…

13:00 JST画像/動画生成

BPG: Balancing Plasticity and Generalization for Domain Incremental Learning

Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performan…

13:00 JST画像/動画生成

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's re…

13:00 JST画像/動画生成ハードウェア/半導体

Modelling Geographic Atrophy Progression using Implicit Neural Representations

Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irrever…

13:00 JSTLLM/生成AI

Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, impli…

13:00 JSTLLM/生成AIGPT / ChatGPTGoogleGemini

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-fi…

13:00 JST画像/動画生成エージェント

MIRA: Medical Image Reflection for Agentic Diagnosis

Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or m…

13:00 JSTLLM/生成AI

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallu…

13:00 JST研究/論文

TACTICL: Task-Aware Compression of Tabular ICL Models

The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task-specific ar…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-c…

13:00 JSTエージェントAnthropic

GitSkills: A Dataset of Agent Skills on GitHub

An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and r…

13:00 JSTLLM/生成AI研究/論文

FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation

Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider h…

13:00 JST画像/動画生成

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation…

13:00 JSTLLM/生成AI

A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recogn…

13:00 JSTLLM/生成AI画像/動画生成ビジネス/資金調達研究/論文

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deterior…

13:00 JSTLLM/生成AI画像/動画生成

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning fo…

13:00 JSTLLM/生成AI

ReLTEx: Reliable LLM-based Taxonomy Expansion

Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and rela…

13:00 JST研究/論文

TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation

Multi-modal recommenders fuse collaborative signals with item modalities such as text, images, and audio, but the usefulness of each drifts…

13:00 JST画像/動画生成

Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers

Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks w…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-…

13:00 JST研究/論文

Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis

Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and…

13:00 JST画像/動画生成

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved,…

13:00 JST研究/論文

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturin…

13:00 JSTLLM/生成AI

Multiclass Sentiment Analysis for Identifying Political Viewpoints

The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public o…

13:00 JST画像/動画生成

3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment

Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing…

13:00 JST画像/動画生成ビジネス/資金調達

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. Howe…

13:00 JST画像/動画生成

Entropy-Centric Explainable AI for Remote Sensing Image Segmentation

Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding…

13:00 JST研究/論文

Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory

We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a…

13:00 JST研究/論文

Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting

Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. However, existi…

13:00 JSTLLM/生成AIハードウェア/半導体

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a c…

13:00 JSTLLM/生成AI

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proc…

13:00 JST研究/論文

How to Verify Consistency of Probabilistic Claims

When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in p…

13:00 JSTLLM/生成AI画像/動画生成エージェント

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiti…

13:00 JSTLLM/生成AI

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, r…

13:00 JST画像/動画生成ロボティクス

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical ro…

13:00 JSTエージェント研究/論文

Representation and Invariance in Reinforcement Learning

Researchers have formalized reinforcement learning (RL) in different ways. If an agent in one RL framework is to run within another RL fram…

13:00 JSTエージェントGPT / ChatGPT

GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning

We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolith…

13:00 JST研究/論文

Closing a 17-Year Gap: Algorithmic Detection and Empirical Prevalence of Rank Reversal in Multi-Criteria Decision Analysis

Rank Reversal, where the relative order of alternatives changes in ways that violate axioms of rational decision-making, is a well-document…

13:00 JSTLLM/生成AI

Multiplayer Nash Preference Optimization

Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human prefer…

13:00 JSTエージェント

On The Statistical Limits of Self-Improving Agents

We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this…

13:00 JST研究/論文ClaudeGPT / ChatGPTGemini

Situation Graph Prediction for User Perspective Modeling

Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited b…

13:00 JST研究/論文

Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach

Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the e…

13:00 JST研究/論文NVIDIA

JEPA-DNA: 共同埋め込み予測アーキテクチャによるゲノム基礎モデルの基礎付け

ゲノム基盤モデル (GFM) は通常、マスク言語モデリング (MLM) またはネクストトークン予測 (NTP) に依存して「自然の法則」を学習します。これらの生成パラダイムは、ローカル構文のキャプチャには効果的ですが、高レベルの機能コンテキストよりもトークンレベルの再構築を優先します。ジョイント埋め込み予測アーキテクチャ (JEPA) と従来の生成目標を統合する、モデルに依存しない継続的トレーニング フレームワークである JEPA-DNA を紹介します。 JEPA-DNA は、潜在空間内のグローバル配列埋め込みを監視することにより、モデルにマスクされたゲノム セグメントの機能的表現を予測させ、学習シグナルをトークン回復からセマンティック アライメントに移行させます。私たちは 17 の多様なゲノム ベンチマーク タスクで JEPA-DNA を評価し、基盤となる GFM アーキテクチャや生成目的に関係なく、線形プロービングとゼロショット パフォーマンスで一貫した向上を実証しました。私たちのフレームワークは、GFM の新しい最先端を確立し、生成精度と潜在的なセマンティック基盤を橋渡しすることで、既存の最高のモデルを上回ります。広範なアブレーション研究を通じて、生成目的と潜在目的の間の相乗的な相互作用をさらに特徴付けます。私たちのコードは https://github.com/NVIDIA-Digital-Bio/JEPA-DNA で公開されています。

原文 (English)

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level reconstruction over high-level functional context. We introduce JEPA-DNA, a model-agnostic continual training framework that integrates a Joint-Embedding Predictive Architecture (JEPA) with traditional generative objectives. By supervising global sequence embeddings in a latent space, JEPA-DNA forces models to predict the functional representations of masked genomic segments, shifting the learning signal from token recovery to semantic alignment. We evaluate JEPA-DNA on 17 diverse genomic benchmark tasks, demonstrating consistent gains in linear probing and zero-shot performance regardless of the underlying GFM architecture or generative objective. Our framework establishes a new state-of-the-art for GFMs, surpassing the best existing models by bridging generative precision with latent semantic grounding. Through extensive ablation studies, we further characterize the synergistic interplay between generative and latent objectives. Our code is publicly available at https://github.com/NVIDIA-Digital-Bio/JEPA-DNA.

13:00 JSTLLM/生成AIエージェントAnthropicClaude

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address…

13:00 JST研究/論文

Planning Task Shielding: Detecting and Repairing Flaws in Planning Tasks through Turning them Unsolvable

Most research in planning focuses on generating a plan to achieve a desired set of goals. However, a goal specification can also be used to…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達GPT / ChatGPT

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents

Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely valida…

13:00 JSTLLM/生成AIエージェント研究/論文ClaudeDeepSeek

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a pra…

13:00 JSTエージェント

CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generati…

13:00 JSTエージェント

Memory-Augmented Reinforcement Learning Agent for CAD Generation

Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existi…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents

Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely trea…

13:00 JST研究/論文

話すことから歌うことへ: オーディオビジュアルディープフェイク検出の新たな挑戦

オーディオビジュアル生成モデルの急速な進歩に伴い、信頼性の高い偽造検出がますます重要になっています。オーディオビジュアルディープフェイク検出の既存の方法は、通常、クロスモーダルの不一致に依存しています。歌う場合、リズミカルな発声はこの結合を弱め、自明ではないドメインシフトを導入し、検出性能を大幅に低下させます。歌のベンチマークのギャップを埋めるために、リズムを意識した生成モデルを使用して Singing Head DeepFake (SHDF) データセットを構築します。シナリオをまたいだドメインの変化に対処するために、会話と歌の両方のシナリオを一般化するテキストガイド付き視聴覚偽造検出 (T-AVFD) フレームワークを提案します。 T-AVFD は、顔認証パターン学習器とマルチモーダル差分重み学習モジュールで構成されます。パターン学習器は、顔の特徴を多粒度のテキスト記述と整合させて、一般化可能な信頼性パターンを学習します。重み学習モジュールは、本質的な視聴覚の一貫性を維持し、差分重み付けを介してそれを信頼性パターンと適応的に統合します。複数のトーキング ヘッド ディープフェイク データセットと SHDF に関する広範な実験により、既存のベースラインを超える一貫した改善と、さまざまな摂動下での強力な堅牢性が示されています。

原文 (English)

From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic vocalization weakens this coupling and introduces a nontrivial domain shift, substantially degrading detection performance. We construct the Singing Head DeepFake (SHDF) dataset using rhythm-aware generative models to fill the gap in singing benchmarks. To cope with cross-scenario domain shifts, we propose a Text-guided Audio-Visual Forgery Detection (T-AVFD) framework that generalizes across both talking and singing scenarios. T-AVFD comprises a facial authenticity pattern learner and a multi-modal differential weight learning module. The pattern learner aligns facial features with multi-granularity textual descriptions to learn generalizable authenticity patterns. The weight learning module preserves intrinsic audio-visual consistency and adaptively integrates it with authenticity patterns via differential weighting. Extensive experiments on multiple talking head deepfake datasets and SHDF show consistent improvements over existing baselines and strong robustness under diverse perturbations.

13:00 JSTLLM/生成AI

AXIOM: 検証可能な数学的推論のための信頼優先のニューロシンボリック実行アーキテクチャ

私たちは、自然言語数学的推論のための信頼優先のニューロシンボリック実行アーキテクチャである AXIOM を紹介します。 AXIOM では、言語モデルは厳密に正規化器として機能します。つまり、非公式の問題テキストを、決定論的なコンピューター代数システム (CAS) パイプラインによって消費される狭いスキーマに書き換えます。このパイプラインは、答えを導き出して検証するか、または第一級の出力として棄権します。ルーティングは、問題形状の正規表現、スキーマ固有のプロンプト、および閉じた形式の CAS ハンドラーの間の 1:1:1 の調整に従い、3,100 以上のそのようなルートが出荷され、250 以上の連続した出荷コミットで LOST_CORRECT リグレッションはゼロです。解析可能な信頼性 100.00% で累積正しさ 94.36% (2,592/2,747) の 4 つの MATH カテゴリ (2,747 レコードのベンチマーク全体で確信のある誤答がゼロ)、4 つのドメインすべてがドメインごとの信頼性 100.0% でドメインごとの 70/90/70 の下限を上回っていること、およびレイテンシの中央値に関する経験的結果を報告します。ルールのみのハンドラーで 1 ミリ秒 (lm-eval 算術 20,000 レコード ベンチマークのレコードの 88%)。このアーキテクチャは、パブリック デプロイメントを通じて約 30,000 件の実稼働クエリに対応してきました。私たちが強調する貢献は、最終的な精度の数値ではなく、アーキテクチャが確立する前向きのダイナミクスです。新しいタスクはレジストリを後退させることなく構成されるため、本番環境でログに記録されたすべての棄権は 1 シップ サイクル後の正しい候補となります。このプロパティの背後にある運用規律 (数学テンプレートのバケット化、回帰オラクルとしての LOST_CORRECT スキャン、解析可能優先のオンボーディング、およびファーストクラスの出力としての棄権) は、数学を超えた信頼できる神経記号システムのための移転可能なフレームワークを構成します。

原文 (English)

AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning

We present AXIOM, a trust-first neuro-symbolic architecture for natural-language mathematical reasoning. Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow schema consumed by a deterministic Computer-Algebra-System (CAS) pipeline, which derives and verifies the answer or abstains as a first-class output. Routing follows a 1:1:1 alignment between problem-shape regex, schema-specific prompt, and closed-form CAS handler, with 4,783 such routes shipped, 71% of which answer without invoking the language model, and zero LOST_CORRECT regressions as a standing release gate. Because the answer is derived rather than generated, so is its explanation: every handler emits a step trace of the computation it performed, rendered as prose by a layer covering all 4,785 tasks that cannot narrate a step the handler did not take. We report two numbers and never fuse them. On the full 7-category MATH test split, designed against, AXIOM answers 90.2% correctly (4,510/5,000) with one confident-wrong answer (99.98% trust on parseable). On held-out MATH-500, never designed against, it answers 89.2% correctly (446/500) with zero confident-wrong answers. The 1.0 pp gap is the substantive result: a registry that had merely memorized problem shapes would collapse on held-out data, and this one does not. The rule-only path answers the 20,000-record lm-eval arithmetic benchmark at 100% in 1 ms per record; the system has served ~30,000 public queries. What we emphasize is not an accuracy figure but the forward dynamic: every logged abstain is a candidate correct after one ship cycle, since new tasks compose without regressing the registry. The discipline behind it (math-template bucketing, LOST_CORRECT scan as regression oracle, parseable-first onboarding, abstain as first-class output) transfers to trustworthy neuro-symbolic systems beyond mathematics.

13:00 JST画像/動画生成エージェント研究/論文

神経科学データから発見までのパイプラインで AI エージェントを評価するケーススタディ

Agentic AI ツールは、科学研究のパイプラインにおけるソフトウェア開発のボトルネックを自動化するための有望な道を提供します。特に、科学者が実装の詳細ではなく正確性と堅牢性を重視する分野の専門家が構築するのに数日から数か月かかる段階においては当てはまります。我々は、光遺伝学のデータから発見までのパイプライン上の汎用コーディングエージェントの実証研究を紹介します。当社は、既存のベンチマークよりも大幅に大規模なタスク、桁違いに大きいデータセット、およびドメイン専門家の標準に基づいた評価基準に基づいてエージェントを評価します。エージェントがいくつかの個別のパイプライン ステージを解決できることを示し、ステージ レベルの自動化が扱いやすいことを示唆しています。エージェントのコードの反復を分析することで、エージェントが最も苦労するのは、反復するための事前定義された基準がない場合であり、代わりに科学的な判断を使用して現在のソリューションを評価する必要があること、これが重要な未解決の課題であることがわかりました。科学的実践を反映して、彼らは自己評価のために中間出力を視覚的に検査しようとすることもありますが、見たものを適切に解釈したり、それに基づいて行動したりすることはほとんどできません。エンドツーエンドのパイプラインを正しく解決するには、すべてのパイプライン ステージで成功を結びつける必要がありますが、これはエージェントの現在の能力を超えています。私たちは、計算リソースの管理や、保持されている大規模なデータ収集への一般化など、既存のベンチマークにはほとんど存在しない課題を特定します。最後に、科学的タスクを構築するための原則と、オープンエンドの問題に対する厳密な評価基準を抽出します。

原文 (English)

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

Agentic AI offers a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build and where correctness and robustness matter more than implementation details. We present an empirical study of general-purpose coding agents on a fly optogenetics data-to-discovery pipeline. We assess agents on tasks and datasets substantially larger than existing benchmarks and evaluation criteria grounded in domain expert standards. We show that agents can solve several pipeline stages, suggesting stage-level automation is tractable. By analyzing agents' code iterations, we show they struggle most without a pre-defined criterion, when they must instead use their scientific judgment to assess their current solution. Mirroring scientists, they sometimes attempt visual inspection of intermediate outputs for self-evaluation, but largely fail to interpret what they see or act on it appropriately. Solving the end-to-end pipeline requires stringing together successes across all stages, which is beyond agents' current abilities. We identify challenges largely absent from existing benchmarks, including computational resource management and generalization to held-out data. Finally, we distill principles for constructing scientific tasks and rigorous evaluation criteria for open-ended problems.

13:00 JSTエージェント

エージェントの自動化が収益性を高めるとき: トレース経済的な引受を通じて自律型 AI のリスクを定量化して保証する

AI エージェントは運用システムで不可逆的なアクションを実行できるようになりましたが、エージェントによって引き起こされた損失は依然として明確に割り当てられず、価格も設定されず、移転されません。プロバイダーは結果的損害を否認することが多く、ユーザーは補償されない損失を抱えたままになり、デフォルトの人によるレビューにより自動化の効率向上が制限されます。障害のリスクにもかかわらず、自律型 AI の導入が経済的に受け入れられるようになるのはいつになるのかを尋ねます。私たちの答えは、顧客のタスク追跡エピソードレベルでリスクを定量化し、それを保険を通じて移転することです。自動化は、期待される利益が保険料、管理コスト、および残りのリスクを超える場合に受け入れられます。これには、制限されたアクセス許可と同等のトレースを持つ定義されたロールが必要です。当社は、ツール使用の痕跡を顧客エクスポージャーと請求可能な損失にマッピングし、この表現を価格設定、管理、リスク移転に使用するトレースエコノミック引受を導入します。 LLM ジャッジではなく、決定論的な経済ラベルを使用します。当社のトレースツーロステストベッドでは、トレースエコノミー価格設定により価格設定 MAE が 17.7,000 ドルから 569 ドルに削減され、逆進的な相互補助金が排除されます。 300 トレースの専門家による監査では、295 個のラベルが変更されずに受け入れられます。 1,000 の実際の SWE-smith トレースでは、トレース条件付きコントロールにより CVaR95 が 72% 減少します。 Theorem~1 は有限サンプルのスコープ条件を与えます。コード、ラベル、監査シートをリリースします。

原文 (English)

When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting

AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred. Providers often disclaim consequential damages, users are left with uncompensated losses, and default human review limits the efficiency gains of automation. We ask when autonomous AI deployment can become economically acceptable despite failure risk. Our answer is to quantify risk at the customer-task-trace episode level and transfer it through insurance. Automation is acceptable when its expected benefit exceeds the premium, control cost, and remaining risk. This requires a defined role with bounded permissions and comparable traces. We introduce trace-economic underwriting, which maps tool-use traces to customer exposure and claimable loss, then uses this representation for pricing, control, and risk transfer. It uses deterministic economic labels rather than an LLM judge. In our trace-to-loss testbed, trace-economic pricing reduces pricing MAE from $17.7K to $569 and removes regressive cross-subsidy. A 300-trace expert audit accepts 295 labels unchanged. On 1,000 real SWE-smith traces, trace-conditioned controls reduce CVaR95 by 72%. Theorem~1 gives a finite-sample scope condition. We release code, labels, and audit sheets.

13:00 JSTエージェントGPT / ChatGPT

ReMMD: マルチモーダルな誤情報検出のための現実的な多言語マルチ画像エージェント検証

バイラル投稿には、長い多言語の説明、複数の画像、混合の出所、および微妙なテキストと画像の構成エラーが組み合わされているため、マルチモーダルな誤情報の検出はますます重要になっています。既存のベンチマークと手法は依然としてこの設定にあまり適合していません。通常、それらは短いキャプション、単一の画像、バイナリ ラベル、または 1 つの操作ソースを分離しますが、現実的な証拠検索ではエージェントによる検証は依然としてコストがかかります。我々は、マルチモーダルな誤情報検出のための現実的な多言語マルチ画像エージェント検証フレームワークである ReMMD を紹介します。 ReMMD には、500 のサンプル、2,756 の画像、5 つの単一言語、2 つの言語間設定、3 つのテキスト長階層、複数画像の投稿、5 方向の真実性ラベル、8 つの歪曲ラベル、証拠の出所、根拠を備えた現実世界のマルチモーダル誤情報検出ベンチマークである ReMMDBench が含まれています。また、投稿をアトミックポイントに分解し、再利用可能な証拠セットを構築し、構造化された L1/L2/L3 出力を予測する永続メモリ検証ツールである ReMMD-Agent も含まれています。独自のシステム、オープン LVLM、MMD エージェント、および T2 エージェント全体にわたって、ReMMD エージェントは、GPT-5.2 を使用して 41.80% の精度と 39.12% のマクロ F1 という最高の 5 方向正確性パフォーマンスを実現しながら、MMD エージェントと比較して 17.5%、T2 エージェントと比較して 79.9% コストを削減します。プロジェクトは https://dang-ai.github.io/ReMMD で入手できます。

原文 (English)

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle cross-modal framing errors. Existing benchmarks and methods remain poorly matched to this setting: they usually isolate short captions, single images, binary labels, or one manipulation source, while agentic verification remains costly under realistic evidence search. We present ReMMD, a realistic multilingual multi-image agentic verification framework for multimodal misinformation detection. ReMMD includes ReMMDBench, a real-world multimodal misinformation detection benchmark with 500 samples, 2,756 images, five monolingual evaluations, two cross-lingual settings, three text-length tiers, multi-image posts, five-way veracity labels, eight distortion labels, evidence provenance, and rationales. It also includes ReMMD-Agent, a persistent-memory verifier that decomposes posts into atomic points, builds a reusable evidence set, and predicts structured veracity verdicts, fine-grained distortion diagnoses, and explanatory rationales. Across proprietary systems, open LVLMs, MMD-Agent, and T$^2$-Agent, ReMMD-Agent obtains the best five-way veracity performance, with 41.80% accuracy and 39.12% macro-F1 using GPT-5.2, while reducing cost by 17.5% relative to MMD-Agent and 79.9% relative to T$^2$-Agent. The project is available at https://dang-ai.github.io/ReMMD.

13:00 JSTエージェントロボティクス

インタラクティブなゲームプレイのためのコーチング可能なエージェント

強化学習は、高度な AI およびロボット システムの作成における貴重なツールであることが証明されており、ゲームプレイからロボット工学、基礎モデルに至るまであらゆるものに貢献しています。通常、これらの AI システムは、試行錯誤を通じて、タスクを解決するために最適に近い 1 つの動作を学習します。ただし、タスクの解決方法に関して、できればリアルタイムで、ある程度の制御を主張したいユースケースは数多くあります。コアタスクのこれらの変更をスタイルと呼びます。私たちは、ユニバーサル価値関数近似器 (UVFA) を、慎重に選択されたトレーニング シナリオ、学習アルゴリズム、データ拡張と組み合わせて、複雑な領域でスタイルを示すエージェントをコーチングするためのフレームワークを作成します。私たちは、AAA ビデオ ゲームの Horizo​​n Forbidden West と Gran Turismo、およびオープンソースのヒューマノイド テスト ドメインでのフレームワークのアプリケーションを実証します。カーレース、様式化されたゲーム戦闘、人型歩行など、ドメインの性質が異なるにもかかわらず、各エージェントは、そのドメインの主なタスクを満たしながら、スタイルの要求に強い一貫性を示します。重要なのは、このホワイト ペーパーで概説した手法を使用すると、エンド ユーザーが実行時に最終的な動作を選択できるため、最終的に実行されるパフォーマンスを柔軟に制御できることです。

原文 (English)

Coachable agents for interactive gameplay

Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework's application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains -- car racing, stylized game combat, and humanoid walking -- each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.

13:00 JSTエージェント研究/論文

Sim と Real のギャップを測定する: AIoT システムにおける強化学習のための手頃な価格の現実世界のベンチマーク プラットフォームの設計

強化学習 (RL) は、自律型モノのインターネット (AIoT) などの自律型システムのパフォーマンスを強化するために一般的に使用されます。ただし、RL は試行錯誤を繰り返す性質があるため、現実の環境で実施するとコストがかかり、シナリオによっては危険が伴います。したがって、RL 研究の大部分はシミュレーションで行われます。この依存により、Sim-to-Real の転送可能性に関連する課題が生じます。 Sim-to-Real アルゴリズムの堅牢性と Sim-to-Real ギャップを評価することは、現実世界での RL パフォーマンスの向上を目的とした研究の重要な前提条件です。したがって、ロボット工学などの業界は、この研究を促進するために、同時シミュレーションと物理プラットフォームを開発してきました。ただし、AIoT 用の汎用 Sim-to-Real ベンチマーク プラットフォームは現在存在しません。これらの懸念に対処するために、私たちは AIoT における RL を研究するための実世界の AIoT プラットフォームを開発しました。このプラットフォームでは、エッジ デバイスに配置されたエージェントが、ハードウェア エミュレートされたキーボードを介して、ビジョン入力に基づいて別のホスト コンピューターでビデオ ゲームをプレイします。このプラットフォームは、2 台のコンピューターとともに、400 ドル未満の市販コンポーネントを使用します。このシステムの目的はゲーム スコアの最大化であるため、現実世界の RL 展開に伴う安全性のリスクを本質的に軽減します。実験結果では、シミュレーションでトレーニングされたエージェントは、現実世界への展開後に人間レベルのパフォーマンスと比較して 1160% のパフォーマンス低下が見られ、Sim と Real の大きなギャップが示されています。ディープ Q ネットワーク (DQN) アルゴリズムを使用した現実世界の直接トレーニングは、1,000 万回のトレーニング ステップ後に人間レベルのパフォーマンスの約 38% を達成し、現実世界の条件下での RL の実現可能性を示しています。これらの結果は、提案された Sim-to-Real ベンチマーク プラットフォームが、現実世界の AIoT システムにおける RL の定性的および定量的評価のための実質的な基盤を提供することを示唆しています。

原文 (English)

Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems

Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environments, is costly and hazardous in some scenarios. Consequently, the majority of RL research is conducted in simulation. This reliance introduces challenges related to the Sim-to-Real transferability. Evaluating the Sim-to-Real algorithmic robustness and the Sim-to-Real gap is a critical prerequisite for research aimed at improving RL performance in the real world. Therefore, industries such as robotics have developed concurrent simulation and physical platforms to facilitate this research. However, a universal Sim-to-Real benchmark platform for AIoT does not currently exist. To address these concerns, we developed a real-world AIoT platform for studying RL in AIoT. On this platform, an agent deployed on an edge device plays video games on a separate host computer via a hardware-emulated keyboard, guided by vision input. This platform uses commercially available components costing less than USD 400, together with two computers. Because the system's objective is game score maximization, it inherently mitigates safety risks associated with real-world RL deployments. Experimental results show the simulation-trained agent suffers a 1160% performance degradation relative to the human-level performance after real-world deployment, indicating a significant Sim-to-Real gap. Direct real-world training using the deep Q-network (DQN) algorithm achieves approximately 49% of human-level performance after 10 million training steps, demonstrating the feasibility of RL under real-world conditions. These results suggest that the proposed Sim-to-Real benchmark platform provides a substantial foundation for qualitative and quantitative evaluations of RL in real-world AIoT systems.

13:00 JST研究/論文

あなたのモデルは考えているのでしょうか、それとも単に停滞しているのでしょうか? PUMA: 位相と運動量の調整による推論病理の診断

テスト時間のスケーリングにより、大規模推論モデル (LRM) が広範な思考連鎖 (CoT) を通じて複雑なタスクに取り組むことができるようになります。ただし、これは多くの場合、冗長な推論によって精度が保証されずに計算オーバーヘッドが増加する「考えすぎ」のパラドックスを引き起こします。既存のテスト時間効率の最適化手法は、主に 2 つのカテゴリに分類されます。1 つは不確実性が低いため幻覚が隠蔽される「欺瞞的収束」を起こしやすい情報理論的アプローチ、もう 1 つは事後的なことが多く、動的推論に対するリアルタイム感度が欠けている潜在表現分析です。このギャップを埋めるために、私たちはまず位相運動量整合仮説を立て、推論の正しさは幾何学的な運動量と不確実性の解決の間の時間的同期に依存すると主張します。次に、潜在速度とねじれによって定量化される幾何学的認知努力とエントロピー認知不確実性という 2 つの直交する次元を通じてこれらのダイナミクスを特徴付ける認知エネルギー モデルを理論的に定式化します。これを運用するために、階層型診断アーキテクチャを採用したトレーニング不要のフレームワークである PUMA (Phase-Uncertainty Momentum Alignment) を導入します。軽量の位相モニタリングとイベントトリガーの幾何学的解析を組み合わせることで、PUMA は能動的な探査と受動的な停滞を効果的に区別し、適応的な切り捨てや修正措置による正確な介入を可能にします。 1.5B から 32B にわたる LRM に関する広範な実験により、PUMA がさまざまなベンチマークにわたって一貫して最先端のベースラインを上回り、優れた精度と効率のトレードオフと堅牢なクロスドメイン一般化を達成していることが実証されました。

原文 (English)

UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and underthinking, which we formulate as reasoning state--action mismatch. Resolving this mismatch requires reliable reasoning state diagnosis, yet single-signal monitors provide ambiguous evidence, while steering-based controllers often rely on outcome-labeled supervision or model-specific calibration. We introduce the Uncertainty--Progress Alignment Hypothesis, which posits that the relative transition timing of proxy answer uncertainty and latent reasoning progress distinguishes healthy, stagnant, and ready states that warrant different subsequent actions. Building on this insight, we propose UPAIR, a training-free framework that couples lightweight uncertainty monitoring with event-triggered joint diagnosis and maps the resulting state to native continuation, selective strategy switching, or verification-guided stopping. Across three LRMs and five cross-domain benchmarks, the stagnation diagnosis detects 64.3% of natural errors while flagging only 5.4% of correct samples, revealing a dynamic reasoning regularity shared across models and tasks. End to end, UPAIR improves accuracy by up to 16.67 percentage points and reduces generated tokens by up to 29.64%, demonstrating the effectiveness of its integrated diagnosis and intervention, while online diagnosis costs less than 1% of natural-generation time.

13:00 JSTLLM/生成AIGemma

大規模言語モデルにおける活性化空間制御のための統計的に根拠のある疎特徴介入

アクティベーション ステアリングは、大規模な言語モデルの動作制御のための微調整に代わる軽量の代替手段を提供しますが、SAE ベースのステアリング手法は、多くの場合、学習されたステアリング目標または単一基準の機能選択に依存します。透過的な SAE 特徴ステアリング パイプラインを導入します。これは、最初に 6 条件の信頼性フィルターを適用し、次に 3 つの相補統計 ($F$-test、KSG 相互情報量、および Cohen の $d$) に対する重み付けされていない Borda コンセンサスを通じて疎な特徴をランク付けします。結果として得られるステアリング方向は、SAE デコーダ行のコーエン $d$ 重み付け組み合わせとして構築され、近似的な SAE 特徴非相関のもとでフィッシャー LDA によって動機付けられる最適化のない方向を提供します。この方法は、3 つの Gemma ファミリー モデル、4 つの動作ドメイン、および 356 の層強度構成にわたって、測定可能なドメイン固有の変化を生成しながら、生の属性の動きと品質を保持した生成との間に大きなギャップがあることを明らかにします。最も強力な構成では、論理的正確さのステアリングは、Gemma~2 9B で $+1.16$ のプライマリ スコア デルタに達します。ただし、より広範な発見は、使用可能なステアリングはモデル、ドメイン、レイヤー、強度によって非常に局所的であるということです。これらの結果は、アクティベーション・ステアリング評価では、生の行動の変化とともに品質条件付きの成功を報告する必要があることを主張しています。コードとデータは https://github.com/Oshayer-Siddique/LLM-Steering-Using-SAE で入手できます。

原文 (English)

SAE-StatSteer: Statistical Consensus Feature Selection for Optimization-Free Activation Steering of Large Language Models

Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning. Sparse autoencoders (SAEs) can make such interventions auditable by decomposing dense activations into an approximately monosemantic feature basis. We introduce SAE-StatSteer, a transparent, optimization-free pipeline. It first filters features through six reliability conditions, then ranks the survivors by an unweighted Borda consensus over three statistics, an $F$-test, KSG mutual information, and Cohen's $d$, and finally combines the selected SAE decoder rows using Cohen's-$d$ weights. We evaluate three Gemma-family models across four behavioral domains against seven dense or SAE-based baselines. Our quality-conditioned protocol requires attribute movement while preserving relevance, richness, and coherence. Raw success systematically overstates usable control because strong shifts often degrade generation quality, and effective steering is not governed by a universal layer or strength. SAE-StatSteer remains competitive with optimization-based methods while exposing every selection and weighting decision for audit. These results motivate reporting quality-conditioned success alongside raw behavioral shift. Our code and data are available at https://github.com/Oshayer-Siddique/LLM-Steering-Using-SAE.

13:00 JSTLLM/生成AIエージェント

AttriMem: エージェントの記憶学習のためのアトリビューションに基づくプロセス フィードバック

LLM エージェントにとって効果的な記憶は非常に重要ですが、それを効果的に構築するのは依然として困難です。メモリ構築ポリシーは、インタラクションが蓄積するにつれてどの情報を抽出、保存、更新、圧縮、または破棄するかを決定します。ヒューリスティック記憶手法は主観的なタスク固有のルールに依存しているため、下流の目標とずれたり、タスク間の適応性が制限されたりする可能性があります。対照的に、RL ベースの手法はタスクのフィードバックから学習しますが、主に結果レベルまたはモジュールレベルの報酬を使用します。これらの粗い信号はタスクの成功を示しますが、どの中間メモリの内容が最終的な答えをサポートしているかを特定できず、きめの細かいクレジット割り当てのボトルネックが生じます。ただし、このようなプロセス フィードバックの構築は、中間記憶の決定には固有のグラウンドトゥルース ターゲットが欠けている一方、適切なクレジットはエージェントの不確実な推論軌道によって変化するため、事前に指定できないため、非常に困難です。我々は、RL を使用してメモリ構築ポリシーを学習するためのアトリビューションに基づくプロセス フィードバック フレームワークである AttriMem を提案します。 AttriMem は、最終的な回答へのトークンレベルの貢献から得られるローカルな報酬でグローバルな結果報酬を強化します。長期対話型質​​問応答の実験では、AttriMem が検索ベース、ヒューリスティック、RL ベースのベースラインを上回り、ベンチマークと回答モデル全体で一般化され、RL の最適化が安定することが示されました。

原文 (English)

AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction

Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with downstream objectives and limit cross-task adaptability. RL-based methods, by contrast, learn from task feedback but mainly use outcome- or module-level rewards. These coarse signals indicate task success but cannot identify which intermediate memory contents support the final answer, creating a fine-grained credit-assignment bottleneck. However, constructing such process feedback is prohibitively difficult because intermediate memory decisions lack unique ground-truth targets, while the appropriate credit varies with the agent's uncertain reasoning trajectory and therefore cannot be specified in advance. We propose AttriMem, an attribution-guided process-feedback framework for learning memory-construction policies with RL. AttriMem augments the global outcome reward with local rewards derived from token-level contributions to the final answer. Experiments on long-horizon dialogue question answering show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization.

13:00 JSTLLM/生成AI

Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog

Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning,…

13:00 JST研究/論文

Learning and Structurally Validating Simulation Scenario Continuations in Dynamic Graph Systems

Data-driven generative models can extend partially observed simulation trajectories into ensembles of alternative future scenarios. However…

13:00 JSTエージェントGPT / ChatGPT

EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups…

13:00 JSTLLM/生成AI

TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LLMs

Large language models (LLMs) can generate fluent reasoning traces that nevertheless lead to incorrect answers, making response-level uncert…

13:00 JST研究/論文

DiffImaginE: Diffusio を使用してエンティティ タイプを検証することを想像してください

マルチモーダル名前付きエンティティ認識 (MNER) は、各候補スパンとエンティティ タイプの仮説が共同のテキスト証拠と視覚的証拠によってサポートされているかどうかを判断します。既存の想像比較検証器は、各 (スパン、タイプ) ペアを 1 つの予測された視覚的特徴にマッピングし、多様な視覚的実現を単一のプロトタイプに圧縮し、明示的な確率的セマンティクスを使用せずに互換性スコアを提供します。 MNER 型検証を条件付き潜在拡散推論として定式化する DiffImaginE を紹介します。スパン局所化された視覚的証拠が与えられると、タイプ条件付きデノイザーは、標準化された潜在に注入されるノイズを予測します。結果として生じるノイズ除去誤差は、タイプ条件付き負の対数尤度の ELBO 一貫性のある代用値を提供し、競合するタイプの仮説を、観察をどの程度うまく説明できるかによってランク付けできるようにします。 DiffImaginE は、標準のマルチモーダル エンコーダ スタックを保持し、決定論的検証器を、Min-SNR 重み付けを使用してトレーニングされた分類子なしのガイド付き拡散スコアラーに置き換えます。タイプごとの拡散スコアを分類ロジットとして直接監視し、ノイズ レベル全体の集計を学習し、逆サンプリングを使用してモンテカルロ比較の分散を削減します。私たちの分析は、分類器を使用しないガイダンスが誘導型事後分布を鮮明にし、反対のペアリングが等しいデノイザーコストで分散を低減するときの特徴を示すことを示しています。 Twitter-2015 と Twitter-2017 の実験では、アブレーションと一対の有意性検定によってサポートされ、同じエンコーダー、補助対物レンズ、評価プロトコルの下で、一致した決定論的 ImaginE 制御に対して一貫したゲインが示されています。

原文 (English)

DiffImaginE: Imagine to Verify Entity Types with Diffusio

Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.

13:00 JSTエージェント

CastFSR: コンテキストを認識した時系列予測のための高速-低速-反映エージェント推論フレームワーク

時系列予測は複雑なシステムにおける意思決定の基礎であり、将来のダイナミクスは過去の観測結果だけでなく、進化するコンテキスト上の特徴にも影響されます。大規模言語モデル (LLM) の最近の進歩により、予測は数値的外挿を超えて、コンテキストを意識した推論にまで拡張されました。しかし、既存のアプローチには、関連するコンテキストを特定し、その影響を推論し、時間的および領域の制約に対して予測を検証するための明示的なメカニズムが欠けていることがよくあります。この研究では、コンテキスト認識型の予測を Fast--Slow--Reflect ワークフローとして定式化するエージェント フレームワークである CastFSR を提案します。迅速な思考により、CastFSR は観測をプロファイリングし、データ駆動型の予測を事前に構築する軽量の予報担当者を選択します。ゆっくりと検討しながら、状況に応じた証拠を取得し、有益なルックバックウィンドウを適応的に決定し、状況が将来のダイナミクスをどのように再形成するかについて推論します。リフレクションでは、予測を繰り返し調整して、時間的、文脈的、およびドメインの一貫性を確保します。 CastFSR は、既製の LLM を使用したトレーニング不要の推論と、オーケストレーション機能をコンパクトな LLM に移行する 2 段階の SFT および強化学習戦略による効率的な展開の両方をサポートします。公開データセットに対する広範な実験により、CastFSR が代表的なベースラインを常に上回るパフォーマンスを示していることが実証されています。私たちのコードは https://github.com/Xiaoyu-Tao/CastFSR で入手できます。

原文 (English)

CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting

Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extended forecasting beyond numerical extrapolation toward context-aware reasoning. However, existing approaches often lack explicit mechanisms to identify relevant contexts, reason about their impacts, and validate forecasts against temporal and domain constraints. In this work, we propose CastFSR, an agentic framework that formulates context-aware forecasting as a Fast--Slow--Reflect workflow. In fast thinking, CastFSR profiles observations and selects lightweight forecasters to construct a data-driven forecast prior. In slow deliberation, it retrieves contextual evidence, adaptively determines informative look-back windows, and reasons about how contexts reshape future dynamics. In reflection, it iteratively refines forecasts to ensure temporal, contextual, and domain consistency. CastFSR supports both training-free inference with off-the-shelf LLMs and efficient deployment through a two-stage SFT and reinforcement learning strategy that transfers its orchestration capability to compact LLMs. Extensive experiments on public datasets demonstrate that CastFSR consistently outperforms representative baselines. Our code is available at https://github.com/Xiaoyu-Tao/CastFSR.

13:00 JSTエージェント

TARL: 長期エージェントの実行可能メモリ管理のためのトランザクション対応の信頼できる台帳

永続的なメモリは、エージェントが知識を長期間保持するのに役立ちますが、単一の更新エラーにより、その後の検索と推論が繰り返し歪められる可能性があります。既存のシステムのほとんどは、メモリの更新をバイナリの書き込み/保持の決定に減らしており、新しい情報を追加すべきか、無視すべきか、古い信念を修正するために使用すべきか、信頼できないとして拒否すべきか、検証を延期すべきかを区別できません。これらの選択は、基本的に異なるメモリ状態を生成しながら、同じバイナリ ラベルを共有する可能性があります。各ステートメントを 5 つの実行可能なアクションの 1 つにマップするメモリ状態更新フレームワークである TARL を紹介します。 TARL は、影響を受けるメモリを特定し、その時間範囲を解決し、ソースの信頼性を比較し、受け入れられた台帳、保留中の台帳、および拒否された台帳を更新します。さらに、代替の更新操作によって生成されたメモリ状態を比較することでトレーニングされ、モデルが正しい結果につながる操作を選択するように促されます。また、きめ細かいアクション ラベルと次の状態のターゲットを備えたベンチマークである TARL-Mem も紹介します。 TARL は、ドメイン内、クロスソース、時間的、反事実的、逐次的な評価にわたって、アクションの予測と状態回復を改善し、メモリ汚染を軽減し、矛盾する証拠を保存し、累積的な破損を制限します。完全なモデル実装は補足資料で提供されます。

原文 (English)

TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents

Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existing systems reduce memory updating to a binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification. These choices may share the same binary label while producing fundamentally different memory states. We introduce TARL, a memory state update framework that maps each statement to one of five executable actions. TARL identifies the affected memory, resolves its temporal scope, compares source reliability, and updates accepted, pending, and rejected ledgers. It is further trained by comparing the memory states produced by alternative update operations, encouraging the model to select the operation that leads to the correct result. We also introduce TARL-Mem, a benchmark with fine-grained action labels and next-state targets. Across in-domain, cross-source, temporal, counterfactual, and sequential evaluations, TARL improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption.

13:00 JST研究/論文

人間の認知と行動の小規模な基礎モデル

人間の行動データに基づいて微調整された大規模な言語モデルが、汎用の認知プロキシとして登場しましたが、これに必要な規模や、これらのモデルがタスク構造を処理するのか、それとも統計的ショートカットを利用するのかは未解決のままです。私たちは、160 の実験からの 1,070 万の試験レベルの選択肢のデータセットである Psych-101 上の 4 つのアーキテクチャ ファミリにわたる 135M から 14B のパラメーターの 14 のモデルをトレーニングします。流通においては、規模はほとんど重要ではありません。モデルは、あたかも天井に向かっているかのように狭い帯域内に収まり、参加者が参加しなかった場合の 70B のベースラインに一致するには、0.6B ~ 1B のパラメータで十分です。分布外では、そのバンドは著しく急峻なスケーリング勾配に向かって開き、新しいタスク構造への一般化において、より大きなモデルが明らかに有利になります。これらのモデルがどのような情報を使用するかを判断するために、2 つの診断を実行します。 27 回の実験にわたって、タスクの指示、実験刺激、結果のフィードバック、選択履歴という 4 つのプロンプト チャネルを段階的に削除し、試行順序を変更します。刺激とフィードバックの内容をマスキングすると、学習した情報の 75.7% が破壊され、モデルが確率以下に押し下げられます。これは、選択履歴だけではパフォーマンスを考慮できないことを示しています。順列は、独立した試行を伴うタスクの不変性を明らかにしますが、試行の順序が前の応答によって決定される感度を明らかにします。したがって、認知的に微調整された小さなモデルは、心理学実験のノイズ上限推定器として有望ですが、その範囲はトレーニングで見られるパラダイムによって制限されたままです。

原文 (English)

Small Foundation Models of Human Cognition and Behaviour

Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPTLlama

ECHO: 一時記憶、安全ガードレール、音声評価を備えた、ローカルに導入可能なエージェント型健康アシスタント

この文書では、長期慢性期医療管理のための、ローカルに導入可能な会話型健康アシスタントである ECHO (Enhanced Care \& Health Observer) について紹介します。 ECHO は、統一システムとして共有監督の下で開発された 3 つの補完的なソフトウェア モジュールを統合します。コア モジュールは、LangGraph を介してオーケストレーションされた ReAct ループ上に構築されたエージェント チャットボットで、17 の臨床ツールと永続的なクロスセッション メモリのための時間的ナレッジ グラフを備えています。 GPT-5 Mini を使用した 59 シナリオのベンチマーク全体で 94.9\% のツール実行合格率を達成しました。 2 段階のハイブリッド安全層がすべての受信クエリを傍受します。ルールベースの層は明示的な危機信号とジェイルブレイクの試行を 1 ミリ秒未満で処理し、APPNP スタイルの伝播を備えた符号付きグラフ ニューラル ネットワーク (GNN) は臨床目的によって境界ケースを分類し、2,537 クエリの注釈付きトルコ健康データセットで 88.8% の精度と 90.6% の安全でないリコールを達成しながら、ゼロショット LLM ベースラインを上回るパフォーマンスを実現します。ラマ 3.3 70B を含む。 Whisper 音響エンコーディングと BERT テキスト エンコーディングをクロスアテンション フュージョンと組み合わせたマルチモーダル音声評価モジュールは、感情、憂鬱、痛みを推定し、平均マクロ F1 が 0.652 に達します。完全なシステムは、消費者向けハードウェア上で完全に実行できる Web アプリケーションとして実装されており、患者データは外部サービスに送信されず、GDPR および KVKK への準拠をサポートします。

原文 (English)

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervision as a unified system. The core module is an agentic chatbot built on a ReAct loop orchestrated via LangGraph, equipped with 17 clinical tools and a temporal knowledge graph for persistent cross-session memory; it achieves a 94.9% tool-execution pass rate across a 59-scenario benchmark with GPT-5 Mini. A two-stage hybrid safety layer intercepts all incoming queries: a rule-based layer handles explicit crisis signals and jailbreak attempts in under 1ms, while a signed graph neural network (GNN) with APPNP-style propagation classifies boundary cases by clinical intent, achieving 88.8% accuracy and 90.6% unsafe recall on a 2,537-query annotated Turkish health dataset while outperforming zero-shot LLM baselines including Llama 3.3 70B. A multimodal speech assessment module combining Whisper acoustic encoding and BERT text encoding with cross-attention fusion estimates emotion, depression, and pain, reaching a mean macro F1 of 0.652. The full system is implemented as a web application that can run entirely on consumer hardware, with no patient data transmitted to external services, supporting compliance with GDPR and KVKK.

13:00 JSTエージェント

検索エージェント向けのコンテキスト情報ポリシーの最適化

検索エージェントは、複数ステップの推論中に外部証拠を取得して使用できるようにすることで、静的パラメトリック メモリを超えて大規模な言語モデルを拡張します。複雑な情報や進化する情報を伴う知識集約型タスクの場合、その信頼性は、関連する証拠を取得するだけでなく、それをその後の推論の指針として使用することにも依存します。ただし、既存の方法では、検索後のアクションが検索された証拠に基づいているかどうかを直接評価することなく、主に最終的な回答の正確性または中間の進歩に報酬を与えます。この不整合により、事前主導型の推論が促進されます。エージェントは内部知識に基づいて結論を出し、主にそれを確認するために検索を使用するため、確証バイアスと非効率的な証拠の使用が生じます。この問題に対処するために、ポリシーの最適化と外部の証拠の使用を明示的に調整する証拠指向の強化学習フレームワークであるコンテキスト情報ポリシー最適化 (CIPO) を提案します。 CIPO は、取得した情報の影響を受ける推論アクションに密なターンレベルのクレジットを割り当てますが、この証拠使用シグナルと、回答の正しさを維持するためのグローバルな結果報酬を組み合わせます。この方法により、CIPO は証拠から切り離された推測を阻止し、取得した事実がその後の推論を導き、修正できる推論の軌道を促進します。重要なのは、CIPO では人間のプロセス アノテーションも追加の報酬モデルも必要ないことです。 7 つのドメイン内およびドメイン外のベンチマークに関する広範な実験により、CIPO が事前駆動推論の蔓延を減らし、ほとんどのタスクで優れたパフォーマンスを達成することが示されています。

原文 (English)

Contextual Information Policy Optimization for Search Agents

Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant evidence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer correctness or intermediate progress, without directly assessing whether post-retrieval actions are grounded in the retrieved evidence. This misalignment encourages prior-driven reasoning: agents form conclusions based on internal knowledge and use retrieval mainly to confirm them, resulting in confirmation bias and inefficient evidence use. To address this issue, we propose Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework that explicitly aligns policy optimization with external evidence use. CIPO assigns dense, turn-level credit to reasoning actions influenced by retrieved information, while combining this evidence-use signal with a global outcome reward to preserve answer correctness. With this manner, CIPO discourages evidence-detached guesses and promotes reasoning trajectories in which retrieved facts can guide or revise subsequent reasoning. Importantly, CIPO requires neither human process annotations nor an additional reward model. Extensive experiments on seven in-domain and out-of-domain benchmarks show that CIPO reduces the prevalence of prior-driven reasoning and achieves excellent performance on most tasks.

13:00 JSTエージェントOpenAI

爆発範囲

エージェント コーディングは、手頃な価格とトークンの無駄という増大する問題に直面しています。結合されたコンテキストとコード チャネルを通じて受信プロンプトの到達範囲を推定する予測メモリ管理レイヤーである Blast Radius を紹介します。 NECROPHORESIS は、デッドコンテキストをそのままアーカイブすることで可逆的な削除を可能にし、一方、Recurring Dead Matter (RDM) は、繰り返し発生するトランスクリプトを識別して埋めます。ポーランドのコンテキスト空間上で可逆的なコンテキストの削除を定式化し、コンテキストのエントロピーを復活確率に関連付けながら、保持、再発、および削除の測定可能な基盤を提供します。 7 つの OpenAI モデル全体で、Blast Radius はトークン消費量を 17 ~ 26% 削減し、テストされたポリシーの中で最も低いオーバーフロー率を達成し、バイト正確な可逆性を維持しました。埋葬された遺体450体のうち、378体は再発死体であり、回収された遺体はゼロだった。 Blast Radius は HCRC の下で動作し、どのレコードを埋めるか、および受信プロンプトがコードベースにどこまで届くかを決定します。この取り組みは、大規模な言語モデルとエージェント コーディングをより再利用可能で持続可能なものにするという Algosophy のより広範な目標に貢献します。

原文 (English)

Blast Radius

Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables reversible eviction by archiving dead context verbatim, while Recurring Dead Matter (RDM) identifies and buries repeatedly occurring transcripts. We formulate reversible context eviction over a Polish context space, providing a measurable foundation for retention, recurrence, and eviction while connecting context entropy to resurrection probability. Across seven OpenAI models, Blast Radius reduced token consumption by 17-26%, achieved the lowest overflow rate among tested policies, and remained byte exact reversible. Of 450 buried bodies, 378 were recurring dead matter and zero were recalled. Blast Radius operates beneath HCRC, determining which records to bury and how far an incoming prompt may reach into the codebase. This work contributes to the broader goal of Algosophy: making large language models and agentic coding more reusable and sustainable.

13:00 JSTLLM/生成AI

TongGuOCR: 中国の歴史文書向けのレイアウトを認識し、トークンを拡張した OCR フレームワーク

中国の歴史文書は貴重な文化遺産を保存していますが、多くのコレクションはスキャンされたページ画像としてのみアクセスできるため、全文検索、照合、およびコンピューターによる分析ができません。光学式文字認識 (OCR) はこのギャップを埋めることができますが、歴史的文書には複雑なレイアウト、珍しい文字、および重要な読み順が含まれることが多いため、正確な転写は依然として困難です。私たちは、中国の歴史文書向けのレイアウト認識型でトークン拡張された OCR フレームワークである TongGuOCR を提案します。まず、レイアウト認識前処理モジュールが、ローカルで一貫性のある認識ブロックを構築および調整して、ローカル コンテキストを維持しながら、領域間の干渉を軽減します。第 2 に、トークン拡張認識モジュールは、転写ターゲットを 2 つの相補的なレベルで強化します。文字レベルの語彙拡張により、各希少グリフに直接 1 トークン表現が与えられ、デコード パスが短縮されます。一方、行間遷移モデリングにより、正確な座標を必要とせずに複雑な読み取りパスに沿ってデコーダを誘導する離散的な空間変位トークンが注入されます。 2 つの中国の歴史文書 OCR ベンチマークの実験では、TongGuOCR が、代表的な従来のタスク固有の OCR モデル、汎用のマルチモーダル大規模言語モデル (MLLM)、および OCR 指向の MLLM よりも優れていることが示されています。より困難な M5HisDoc ベンチマークでは、TongGuOCR は 93.76 AR を達成し、各指標の最良の競合スコアと比較して NED を 10.43 から 6.15 に、RO-ED を 7.53 から 3.49 に削減しました。オンライン デモは https://jzzh2004.github.io/TongGuOCR で利用できます。

原文 (English)

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge this gap, but accurate transcription remains challenging because historical documents often contain complex layouts, rare characters, and nontrivial reading orders. We propose TongGuOCR, a layout-aware and token-augmented multimodal large language model (MLLM) for OCR of Chinese historical documents. First, a Layout-Aware Preprocessing module constructs and refines locally coherent recognition blocks to preserve local context while reducing interference across regions. Second, a Token-Augmented Recognition module augments the transcription target at two complementary levels: character-level vocabulary expansion gives each rare glyph a direct one-token representation and shortens its decoding path, while line-to-line transition modeling injects discrete spatial displacement tokens that guide the decoder along complex reading paths without requiring precise coordinates. Experiments on two Chinese historical document OCR benchmarks show that TongGuOCR outperforms representative traditional task-specific OCR models, general-purpose MLLMs, and OCR-oriented MLLMs. On the more challenging M5HisDoc benchmark, TongGuOCR achieves 93.76 AR and reduces NED from 10.43 to 6.15 and RO-ED from 7.53 to 3.49 relative to the best competing score for each metric. An online demo is available at https://jzzh2004.github.io/TongGuOCR.

13:00 JST研究/論文

推論のための思考レベルのビーム検索

テスト時の計算スケーリングは大規模推論モデル (LRM) のパフォーマンスを左右する主な要因ですが、極度の非効率性が現在のアプローチを制限しており、重要な問題は \emph{どのくらい} 計算を費やすかという問題から、\emph{どこに} 割り当てるかという問題に移りつつあります。テスト時の推論を、部分的な軌跡に対する制約付きの計算割り当て問題として形式化します。ハードウェア予算が固定されていると、既存のパラダイムでは、最も有望な部分的な進捗にコンピューティングを積極的に割り当てることができません。従来の並列サンプリングではトレースが個別に扱われ、深刻なメモリのボトルネックが引き起こされますが、減算的枝刈りではハードウェアが枯渇し、出力分布を積極的かつ十分にシフトすることができません。この二分法を克服するために、 \emph{思考レベルのビーム検索} を実行する推論アルゴリズムである Gambit を導入します。 Gambit は、有望でない軌跡を定期的に枝刈りし、高品質のプレフィックスから即座に分岐することで、継続的に高いハードウェア使用率を維持しながら、隠れ状態を調査する軽量スコアラーを介して、最も有望な推論トレースに計算を動的に集中させます。複数のモデルとベンチマークにわたる広範な評価により、Gambit が既存のベースラインを厳密に支配していることが実証されています。同一のハードウェア制約の下で、私たちの方法は、プルーニング ベースラインと比較して、HMMT-24 で最大 +6.7\%、AIME-25 で +3.3\% の絶対精度の向上をもたらし、トレース完了時に $>2\time$ 高いスループットを実現し、標準の並列サンプリングと比較して総トークン消費量を最大 68.5\% 削減します。

原文 (English)

Thought-Level Beam Search for Reasoning

Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it. We formalize test-time reasoning as a constrained compute allocation problem over partial trajectories. Under a fixed hardware budget, existing paradigms fail to actively allocate the compute to the most promising partial progress: traditional parallel sampling treats traces independently and induces severe memory bottlenecks, while subtractive pruning starves hardware and fails to actively and sufficiently shift the output distribution. To overcome this dichotomy, we introduce Gambit, an inference algorithm that executes \emph{thought-level beam search}. By periodically pruning unpromising trajectories and immediately branching from high-quality prefixes, Gambit dynamically concentrates compute onto the most promising reasoning traces via a light-weight scorer probing hidden states while maintaining continuous high hardware utilization. Extensive evaluations across multiple models and benchmarks demonstrate that Gambit strictly dominates existing baselines. Under identical hardware constraints, our method yields up to a +6.7\% absolute accuracy gain on HMMT-24 and +3.3\% on AIME-25 over pruning baselines, delivers $>2\times$ higher throughput on trace completion, and reduces total token consumption by up to 68.5\% relative to standard parallel sampling.

13:00 JST研究/論文

StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, mos…

13:00 JST研究/論文

SDDBMs: Soft Denoising Diffusion Bridge Models

Diffusion bridge models leverage Doob's \(h\)-transform to construct stochastic transports between arbitrary endpoint distributions, and ha…

13:00 JSTLLM/生成AIエージェント

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide…

13:00 JSTLLM/生成AIエージェント

Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution

Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirel…

13:00 JSTLLM/生成AI

Entropy-based Code Adversarial Translation for Real-world Repository Migration

LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produ…

13:00 JSTLLM/生成AIエージェント

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, c…

13:00 JSTエージェント

CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation

Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these…

13:00 JST画像/動画生成

Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations

The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not wel…

13:00 JST研究/論文

Graphical Models of False Information and Fact Checking Ecosystems

The wide spread of false information online, including misinformation and disinformation, has become a major problem for our highly digitis…

13:00 JST研究/論文

Pretrained Optimization Model for Zero-Shot Black Box Optimization

Zero-shot optimization involves optimizing a target task that was not seen during training, aiming to provide the optimal solution without…

13:00 JST研究/論文

Regression and Classification with Single-Qubit Quantum Neural Networks

The literature reflects a mutually beneficial relationship between machine learning and quantum computing, where progress in one field freq…

13:00 JSTLLM/生成AI規制/政策

Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking

Large language models (LLMs) enable powerful knowledge injection through approaches such as in-context learning and fine-tuning, but they a…

13:00 JST画像/動画生成NVIDIA

TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification

Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual…

13:00 JSTエージェントロボティクス研究/論文

X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation

Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significa…

13:00 JST研究/論文

Demystifying Adversarial Robustness in Diffusion Models: Compression, Randomness, and Geometry

Recent studies suggest that diffusion models significantly improve the empirical adversarial robustness of deep neural network models. Whil…

13:00 JST研究/論文

音楽の解釈と感情の知覚: 計算と神経生理学的研究

この研究では、計算および神経生理学的手法を使用して、音楽演奏における感情表現と知覚を調査します。レパートリー、ダイアトニックモーダルエチュード、即興演奏などのさまざまな演奏環境や表現力のレベルが、演奏者の感情的なコミュニケーションやリスナーの反応に及ぼす影響を調査します。プロのミュージシャンがさまざまなタスクを実行し、演奏者と聴衆の両方から感情的な注釈が提供されました。音声分析により、表現力豊かな即興演奏が独特の音響的特徴を示す一方、感情分析によりより強い感情的反応が示されたことが明らかになりました。神経生理学的測定では、即興演奏においてよりリラックスしていることが示されました。この複合的な研究は、感情的なコミュニケーションと聴衆の関与を高める上での表現力の重要性を強調しています。

原文 (English)

Music Interpretation and Emotion Perception: A Computational and Neurophysiological Investigation

This study investigates emotional expression and perception in music performance using computational and neurophysiological methods. The influence of different performance settings, such as repertoire, diatonic modal etudes, and improvisation, as well as levels of expressiveness, on performers' emotional communication and listeners' reactions is explored. Professional musicians performed various tasks, and emotional annotations were provided by both performers and the audience. Audio analysis revealed that expressive and improvisational performances exhibited unique acoustic features, while emotion analysis showed stronger emotional responses. Neurophysiological measurements indicated greater relaxation in improvisational performances. This multimodal study highlights the significance of expressivity in enhancing emotional communication and audience engagement.

13:00 JSTLLM/生成AI画像/動画生成研究/論文

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benc…

13:00 JST画像/動画生成

SynBoost: A Synergistic Framework for Fast Sampling of Diffusion Models

Diffusion probabilistic models (DPMs) have demonstrated remarkable success in visual generation. However, their iterative sampling mechanis…

13:00 JST研究/論文

OpenDPDv2: A Unified Learning and Optimization Framework for Neural Network Digital Predistortion

Neural network (NN)-based Digital Predistortion (DPD) improves linearization for wideband radio frequency (RF) power amplifiers (PAs) but o…

13:00 JSTLLM/生成AI研究/論文Llama

Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling

This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language mode…

13:00 JST研究/論文

Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives

AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deploye…

13:00 JSTロボティクス

Reconfiguration of pivoting cube ensembles under local sensing constraints using geometric deep learning

We demonstrate that local sensing is sufficient for effective global reconfiguration of homogeneous pivoting cube modular robots in two dim…

13:00 JST画像/動画生成

Token-Based Detection of Spurious Correlations in Vision Transformers

Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit…

13:00 JST研究/論文

Faster Results from a Smarter Schedule: Reframing Collegiate Cross Country through Analysis of the National Running Club Database

Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance…

13:00 JSTロボティクス

Diffusion-Based Impedance Learning for Contact-Rich Manipulation Tasks

Learning-based methods excel at robot motion generation but remain limited in contact-rich physical interaction. Impedance control provides…

13:00 JST研究/論文

Pricing Access to Dynamic Information Services

A provider sells a \emph{dynamic information service}---a real-time, capacity-constrained process that resolves a customer's uncertainty---…

13:00 JST研究/論文NVIDIAAlibaba

HyWA: Architecture-Preserving Personalized Voice Activity Detection for Full-Duplex Voice Assistants

Voice activity detection (VAD) serves as an early gate in voice-assistant pipelines for smart devices. Because conventional VADs respond to…

13:00 JST画像/動画生成エージェント

VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary m…

13:00 JST研究/論文

On the Condition Number Dependency in Bilevel Optimization

Bilevel optimization minimizes an objective function, defined by an upper-level problem whose feasible region is the solution of a lower-le…

13:00 JST研究/論文

Auto-exploration for online reinforcement learning

The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorit…

13:00 JST画像/動画生成

Hybrid Token Compression for Vision-Language Models

Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs. Existing compression metho…

13:00 JST研究/論文

IndexTTS 2.5 Technical Report

In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-b…

13:00 JSTLLM/生成AI

On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis

On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question a…

13:00 JST画像/動画生成GPT / ChatGPT

LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning

Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dyna…

13:00 JST研究/論文

GraFine: Retrieval-Time Refinement for Efficient Graph RAG over Corpus Graphs

Graph RAG on corpus graphs enhances retrieval by leveraging intermediate node content as contextual clues to uncover unretrieved oracle nod…

13:00 JSTエージェントロボティクス

Bandwidth-Efficient Multi-Agent Communication through Information Bottleneck and Vector Quantization

Multi-agent reinforcement learning systems deployed in real-world robotics applications face severe communication constraints that signific…

13:00 JSTLLM/生成AI

LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remain…

13:00 JSTLLM/生成AI

Do LLMs Benefit From Their Own Words?

In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assista…

13:00 JST研究/論文

Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers

We present an Ising-compatible framework for formal neural-network robustness verification under bounded input perturbations. For piecewise…

13:00 JST研究/論文

Can Computational Reducibility Lead to Transferable Models for Graph Combinatorial Optimization?

A key challenge in developing unified neural solvers for combinatorial optimization (CO) is the efficient generalization of models from a g…

13:00 JSTビジネス/資金調達

$\mathrm{ECI}_{\mathrm{sem}}$: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives

Hard-negative source selection for dense retrieval is usually decided only after fine-tuning and downstream evaluation. We propose ECIsem,…

13:00 JSTビジネス/資金調達

Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding

Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which…

13:00 JSTLLM/生成AI画像/動画生成

Covert Visual Prompt Injection against Commercial Multimodal Large Language Models

Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavio…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks

System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensit…

13:00 JST研究/論文

BiScale-GTR: Fragment-Aware Graph Transformers for Multi-Scale Molecular Representation Learning

Fragment-level representations provide a natural way to capture recurring molecular substructures and reuse their learned representations a…

13:00 JSTエージェントロボティクス

RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Prediction

Predicting traffic agent trajectories plays an important role in autonomous driving, traffic operations, transportation safety analysis, et…

13:00 JSTLLM/生成AI

Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers

We study implicit reasoning, i.e. the ability to combine knowledge or rules within a single forward pass. While transformer-based large lan…

13:00 JSTLLM/生成AI

SatIR: Scalable High-Recall Constraint-Satisfaction-Based Information Retrieval for Clinical Trials Matching

Many real-world retrieval and matching problems require more than topical relevance: a candidate must satisfy the specific constraints of o…

13:00 JST画像/動画生成エージェント研究/論文

PinpointQA: A Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos

Reliable embodied interaction in indoor environments requires agents to precisely localize small everyday objects from visual observations.…

13:00 JSTLLM/生成AI研究/論文

InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging…

13:00 JST研究/論文

From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variables

Latent variables pose a fundamental obstacle to both causal discovery and inference. Local approaches exploiting direct neighborhood relati…

13:00 JSTLLM/生成AI

Language corpora for the Dutch medical domain

Background: Dutch medical corpora are scarce, limiting NLP development. Methods: We translated English datasets, identified medical text in…

13:00 JST画像/動画生成

Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models

Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, whic…

13:00 JST研究/論文

Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they…

13:00 JST画像/動画生成

Field-Localized Forgery Detection for Digital Identity Documents

Digital onboarding and eKYC systems used by banks, fintech platforms, telecom providers, and other third-party services commonly verify use…

13:00 JST研究/論文

足場としてのアクセス タイミング: 教育における GenAI への強化学習アプローチ

近年、教育現場における生成 AI (GenAI) は、無制限に使用すると過剰依存、メタ認知的離脱、学習の低下を引き起こす可能性があるにもかかわらず、大学生の日常生活に広く浸透してきています。これまでの研究のほとんどは、その使用法を教育学的に足場にする方法に焦点を当ててきましたが、既製の GenAI をいつ許可するかという問題は依然として研究が不足しており、教育学的に根拠のある実証的調査が不足しています。私たちはアクセス タイミング自体を暗黙的な足場の一種として扱い、メタ認知理論、認知負荷理論、生産的失敗に基づいた報酬関数を使用して、生徒がいつ GenAI にアクセスすべきかを決定する強化学習 (RL) エージェントを通じて操作します。高等教育を受けている N=105 人の学生を対象とした混合方法対照実験室研究で、学習効果とメタ認知への関与に対するエージェントの効果を、無制限の使用と完全に制限された使用の場合と比較しました。結果は、強化学習条件下で戦略的にタイミングを合わせた GenAI アクセスが、無制限のアクセスと比較して客観的なテスト後のパフォーマンスとメタ認知の精度を向上させる一方、完全な保留と比較してタスクのエラーとタスクの時間を削減し、明示的なメタ認知プロンプトや構造化された足場を必要とせずに両方のアプローチよりも優れたパフォーマンスを示したということです。しかし、自己申告のメタ認知意識に関しては、条件間の差異は現れませんでした。したがって、全体として、GenAI アクセスのタイミングは、完全に無制限で保留されたアクセスを改善する、扱いやすく理論に基づいたスケーラブルな教育戦略であり、既製のツールと互換性があり、導入障壁が潜在的に低いです。これにより、アクセスのタイミングを教育者がどのように促進し、人間と AI の学習システムの設計に実装できるかを探る新しい研究領域が開かれます。

原文 (English)

Access Timing as Scaffolding: A Reinforcement Learning Approach to GenAI in Education

In recent years, generative AI (GenAI) in educational settings has become ubiquitous in university students' daily lives, despite its potential to induce over-reliance, metacognitive disengagement, and diminished learning when used unrestrictedly. While most prior research has focused on how to pedagogically scaffold its usage, the question of when to allow off-the-shelf GenAI remains understudied and lacks pedagogically grounded empirical investigation. We treat access timing itself as a form of implicit scaffolding and operationalize it through a reinforcement learning (RL) agent that decides when students should access GenAI, with a reward function grounded in metacognitive theory, cognitive load theory, and productive failure. In a mixed-methods controlled lab study with N=105 higher education students, we compared the agent's effect on learning gains and metacognitive engagement to unrestricted and fully restricted use. Results show that strategically timed GenAI access under the reinforcement learning condition improved objective post-test performance and metacognitive accuracy compared with unrestricted access, without requiring explicit metacognitive prompts or structured scaffolding. Exploratory comparisons with the fully restricted condition further suggest that timed access may reduce task errors and time on task relative to complete withholding. Overall, timing of GenAI access therefore is a tractable, theoretically grounded, and scalable pedagogical strategy that improves over completely unrestricted and withheld access, compatible with off-the-shelf tools and potentially low adoption barrier. This opens up a new research area that explores how access timing can be facilitated by educators and implemented in human-AI learning system design.

13:00 JSTLLM/生成AI画像/動画生成

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically…

13:00 JSTLLM/生成AIビジネス/資金調達

Why Do Safety Guardrails Degrade Across Languages?

Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), whic…

13:00 JST研究/論文

マッチング原理: ノイズに強い表現学習のための損失関数の幾何学理論

ロバスト性、ドメイン適応、測光/オクルージョンの不変性、センサー ドリフト、およびアライメント スタイルは、別個のメソッド ファミリを持つ別個の文献として扱われます。ラベル保存デプロイメント シフトでは、ラベルを変更せずに入力が変更される方法の共分散 Sigma_task = Cov_{Q_n}(n) という 1 つの幾何学的オブジェクトを共有します。 CORAL、敵対的トレーニング、拡張、計量学習、ヤコビアン ペナルティ、およびアラインメント制約は独立したトリックではなく、Sigma_task の推定器です。そのオブジェクトを修正すると、ヤコビアン ペナルティは行列 Sigma' によって固定され、その範囲は range(Sigma_task) をカバーする必要があります (マッチング原則)。線形ガウス モデルの最適性 (Thm. A)、展開ドリフトをゼロにする二次ペナルティの範囲カバレッジの必要性 (Thm. G)、および大域最小値での同じ二分法 (Thm. A*_global) を証明します。間違った方向/信号調整制御 (補題 C; Cor. E/E*) と 7 つの推定量 (補題 D1 ~ D7)、さらにラベルフリー TDI により、Sigma_task を学習する必要がある場合に反証可能なレシピが生成されます。 13 個のブロック (ML から Qwen2.5-7B まで) は、ジオメトリとデプロイメント ドリフトに関する等方性と間違った方向のペナルティの一致をテストします。識別可能性が成立する 12 マッチ理論。 Office-31 は名前付きの eigengap 障害です。部分パス: すべてのヘッドライン タスク メトリクスを移動しなくても、ジオメトリを改善できます。パイロット 7B DPO 実行 (1 エポック、240 ペア): 一致したスタイル PMH は、標準 DPO がスタイル TDI を劣化させる場合にスタイル TDI を保持します。私たちは、標準トレーニングが大域最小値 (仮定 (O) がオープン) に達すること、推定された Sigma_task が常に識別可能であること、またはすべてのリーダーボードで優勢であることを主張しません。私たちは、Sigma_task を推定し、Sigma' を照合し、コントロールを実行し、タスクとジオメトリを個別に報告する、反証可能な設計レシピを主張します。

原文 (English)

The Matching Principle: When Does a Training Penalty Cover Deployment Shift?

Ordinary training optimises the task loss and then stops. It never pays for internal representation energy: Jacobians can stay large in directions that never helped the label, so even small label-preserving noise throws the model off---a design gap that classical noise-injection theory fixes at second order, but only when applied as default regularisation, which current practice does not do. We make that precise with a Matching Principle: name deployment directions (Sigma_task) and the training penalty Sigma', and ask whether the second covers the first. The no-thinking default is even-spread / isotropic penalty (Sigma' proportional to I)---classical Gaussian / Tikhonov at second order: no axis estimate, no architecture change, and---in a simple linear ridge model---strictly less deployment drift than task-only training, with no coverage miss by construction. When axes are known, matching is sharper; when they are missed, a residual floor remains. Across seven domains a named second-moment penalty beats unregularised training; a controlled illustration recovers match > even-spread > wrong-axis when axes are forced. The ridge theorems are proved; deep nets remain experiments under a specified perturbation. Design rule: fix internal energy by default (even-spread); match when axes are known; treat losses that control representation sensitivity as first-class design.

13:00 JSTエージェント

インフラベイジアン強化学習エージェントは、最悪の場合の堅牢性において古典的な RL を上回ります

古典的な強化学習では、エージェントが、その動作がエージェントのポリシーに依存しない固定環境と対話することを前提としています。この仮定は、AI の安全性にとって重要な環境、エージェントが予測者、人間、他の AI エージェント、機関と対話する環境など、他のアクターがエージェントの動作を予測する可能性がある実現不可能な設定では崩れます。このような設定では、エージェントのモデル クラスは、エージェントが動作する世界を捉えることができません。このような仕様の誤りがある場合、古典的なベイジアン手法では、実現可能性が得られないため、確実に間違った事後結果、信頼性の低い決定、際限のない後悔が生じる可能性があります。インフラベイズ主義は、事前分布を合理的に選択できる通常の確率的不確実性と、そのような事前分布を構築する根拠が存在しないナイト不確実性を区別することで、これらの失敗に対処する決定理論の枠組みです。これは、事後期待や加重平均ではなく、最悪の場合の結果に基づいて行動を評価することによって行われます。有限結果ステートレス意思決定問題に対するインフラベイジアン強化学習アーキテクチャの最初の概念実証実装を紹介します。私たちのエージェントは一連の不正確な仮説を維持し、インフラベイズ条件付けを使用してそれらを更新し、最悪の場合の期待値を最大化することによってアクションを選択します。ベイジアン内最大値決定プロセスのこの実装をナイト不確実性のある環境に適用し、古典的な強化学習エージェントと比較して最悪の場合の後悔が低いことを示します。また、Newcomb の問題を調査し、インフラベイジアン エージェントが最適な戦略を選択し、古典的な意思決定理論エージェントを上回るパフォーマンスを示すことを示します。私たちの結果は、モデルの仕様の誤りやポリシーに依存する不確実性の下でも堅牢性を維持する強化学習エージェントへの一歩を提供します。

原文 (English)

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-realizable settings where other actors might anticipate the agent's behavior, including environments crucial to AI safety, where the agent interacts with predictors, humans, other AI agents, and institutions. In such settings, the agent's model class fails to capture the world in which it operates. Under such misspecification, classical Bayesian methods can produce confidently wrong posteriors, unreliable decisions, and unbounded regret, as realizability fails to obtain. Infra-Bayesianism is a decision-theoretic framework that addresses these failures by distinguishing ordinary probabilistic uncertainty, where priors can be reasonably chosen, from Knightian uncertainty, where no grounds exist for the construction of such a prior. It does so by evaluating actions on their worst-case outcomes, rather than from posterior expectations or weighted averaging. We present the first proof-of-concept implementation of an infra-Bayesian reinforcement learning architecture for finite-outcome stateless decision problems. Our agent maintains a set of imprecise hypotheses, updates them using infra-Bayesian conditioning, and selects actions by maximizing worst-case expected value. We apply this implementation of the infra-Bayesian maximin decision process to an environment with Knightian uncertainty, and demonstrate a lower worst-case regret as compared to classical reinforcement learning agents. We also investigate Newcomb's problem and show that the infra-Bayesian agent picks the optimal strategy, outperforming classical decision theory agents. Our results provide a step towards reinforcement learning agents that remain robust under model misspecification and policy-dependent uncertainty.

13:00 JST研究/論文

ベクトル心電図空間における心臓の潜在表現の学習

心電図検査 (ECG) は心臓評価の基礎であり、有益な ECG 表現の学習は、疾患の診断から臨床レポートの作成に至るまでのタスクの基礎となります。ただし、既存の方法は、ほぼ独占的に観測可能な ECG 信号空間で動作します。実際には、標準の 12 誘導 ECG は、異なる空間方向からの同じ基礎となる心臓電気活動の複数の投影を表します。したがって、ECG 空間での表現学習では必然的に大幅な冗長性が導入され、偽の相関が生じたり、過学習のリスクが増大したりする可能性があります。これに対処するために、Frank ベクトル心電図 (VCG) モデルを動機として、心臓の電気活動の統一された潜在表現を VCG 空間で直接学習することを提案します。この物理的に接地された潜在空間で動作するように設計された初の一般的な自己教師あり表現学習フレームワークである LVCG を紹介します。 VCG は、リード固有のアーティファクトではなく、ビュー不変の潜在 VCG 表現を学習することにより、冗長性を最小限に抑え、一般化を向上させます。 LVCG は一般に、タスク全体にわたって ECG 空間ベースラインよりも優れたパフォーマンスを示し、特にドメイン シフト設定において強化された堅牢性と汎用性を示します。

原文 (English)

LVCG: Learning ECG Representations in the Latent Vectorcardiogram Space

Electrocardiography (ECG) is a cornerstone of cardiac assessment, making the learning of informative ECG representations fundamental to tasks ranging from disease diagnosis to clinical report generation. However, existing methods operate almost exclusively in the observable ECG signal space. In practice, the standard twelve-lead ECG represents multiple projections of the same underlying cardiac electrical activity from different spatial orientations. Therefore, representation learning in the ECG space inevitably introduces substantial redundancy, which may lead to spurious correlations and increased risk of overfitting. To address this and motivated by the Frank vectorcardiogram (VCG) model, we propose learning a unified latent representation of cardiac electrical activity directly in the VCG space. We introduce LVCG, the first general self-supervised representation learning framework designed to operate in this physically grounded latent space. By learning view-invariant latent VCG representations rather than lead-specific artifacts, VCG minimizes redundancy and improves generalization. LVCG generally outperforms ECG-space baselines across tasks, demonstrating enhanced robustness and generalization, especially in domain shift settings.

13:00 JST画像/動画生成

圧縮センシング アプリケーションにおけるサンプリング ポリシーを最適化するためのフローベースの生成モデリング

信号処理や医用画像処理における多くの最新のアプリケーションでは、厳しいリソース制約の下で高次元の信号を取得する必要があります。従来のサンプリング理論では、信号を正確に再構成するには、信号の周囲の大きさに比例した測定回数が必要ですが、この要件は高価すぎるか非現実的なことが多いと示唆しています。圧縮センシングは、測定オペレータが特定の条件を満たしていれば、より少ない測定でまばらな信号を回復できることを実証することで、この概念に疑問を投げかけます。この概念実証研究では、圧縮センシング アプリケーションでのサブサンプリングを最適化するようにトレーニングされたフロー モデルを使用して、従来のフロー マッチング トレーニング パラダイムを再定式化した、タスク認識フローベースの生成フレームワークを提示します。画像分類、画像再構成、および MRI 加速のための圧縮センシングのパフォーマンスを大幅に向上させる学習サブサンプリング マスクの提案されたフレームワークの基本的な実現可能性を確立します。画像再構成タスクでは、私たちの方法は最先端のパフォーマンスを実証し、CelebA データセットでは 5\% のサブサンプリング レートで 25.17 dB のピーク信号対雑音比を達成し、最小限の計算オーバーヘッドで $8\times$ の加速 MRI 測定 (fastMRI データセット) を再構成する場合は 29.24 dB を達成しました。これらの結果は、生成フロー モデル内でのタスク条件付けの有効性を強調し、表現学習戦略の有望な方向性を明らかにします。全体として、提案されたフレームワークは、広範囲の逆問題に潜在的に適応できるデータ駆動型およびタスク駆動型のセンシングスキームを設計するための統合された柔軟なアプローチを提供します。

原文 (English)

Flow-Based Generative Modeling for Optimizing Sampling Policies in Compressed Sensing Applications

Numerous modern applications in signal processing and medical imaging necessitate acquiring high-dimensional signals under tight resource constraints. Traditional sampling theory suggests that accurate signal reconstruction requires a number of measurements proportional to the signal's ambient dimension, a requirement often too expensive or impractical. Compressed sensing challenges this notion by demonstrating that sparse signals can be recovered with fewer measurements, provided the measurement operator meets certain conditions. This proof-of-concept study presents a task-aware flow-based generative framework -- a reformulation of the conventional Flow Matching training paradigm with a flow model trained to optimize subsampling in compressed sensing applications. We establish the fundamental feasibility of the proposed framework of learning subsampling masks that substantially enhance the performance of compressed sensing for image classification, image reconstruction, and MRI acceleration. For the image reconstruction task, our method demonstrated state-of-the-art performance, achieving Peak Signal-to-Noise Ratio of 25.17 dB at the subsampling rate of 5\% on the CelebA dataset and 29.24 dB when reconstructing $8\times$ accelerated MRI measurements (fastMRI dataset) with the minimal computational overhead. These results highlight the effectiveness of task-conditioning within generative flow models and reveal a promising direction for representation learning strategies. Overall, the proposed framework offers a unified, flexible approach to designing data- and task-driven sensing schemes that can be potentially adapted to a broad range of inverse problems.

13:00 JSTLLM/生成AIエージェント

エージェントによるツール呼び出しと RL トレーニングの効果と効率について

ツール呼び出しは、最新の大規模言語モデル (LLM) エージェントの中心的なコンポーネントであり、パラメトリック知識を超えたスキルをエージェントに提供します。この論文では、有効性 (つまり、この機能がどのように測定されるか) と効率 (つまり、どのように学習されるか) という 2 つの相補的な軸に沿ってツール呼び出しを研究します。有効性については、ツール呼び出しの評価パイプラインを体系的に分析し、ランダム シード、システム プロンプト、マルチターン テンプレートの構築、以前のインタラクション/推論履歴の引き継ぎ方法など、一見些細で文書化されていない実装の選択肢に結果が非常に敏感である可能性があることを示します。これらの選択は、特に厳密な標準化がなければリーダーボードのランキングが信頼できないマルチターン設定では、報告されるパフォーマンスに大きな違いをもたらす可能性があります。効率に関しては、ツール呼び出しのための標準強化学習 (RL) を調査し、計算無駄の 2 つの原因を特定します。(i) ロールアウト中、多くのプロンプトは学習信号を生成しません。(ii) ポリシー更新中に、最適化により高い計算コストが発生します。これらの発見に基づいて、RL ベースのツール呼び出しトレーニングを加速し、パフォーマンスを低下させることなく実質的な実時間の高速化を達成する 2 つの手法を紹介します。

原文 (English)

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This paper studies tool-calling along two complementary axes: effectiveness, i.e., how this capability is measured, and efficiency, i.e., how it is learned. On effectiveness, we systematically analyze tool-calling evaluation pipelines and show that results can be highly sensitive to seemingly minor, often undocumented implementation choices including the random seed, system prompt, multi-turn template construction, and how prior interaction/reasoning history is carried forward. These choices can lead to substantial differences in reported performance, especially in multi-turn settings where without rigorous standardization, leaderboard rankings are unreliable. On efficiency, we examine standard reinforcement learning (RL) for tool-calling and identify two sources of computational waste: (i) during rollouts, many prompts produce no learning signal, and (ii) during policy updates, optimization incurs high computational cost. Guided by these findings, we introduce two techniques that accelerate RL-based tool-calling training, achieving substantial wall-clock speedup without degrading performance.

13:00 JST規制/政策

Where Flow Matching Leaks: Characterising Membership Signals Along the Interpolation Path

Understanding memorization in generative models remains challenging, with implications for copyright and privacy. Beyond verbatim reproduct…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents

Agent skills extend general-purpose agents, but their open format enables skill poisoning: a tampered skill can make an agent run an attack…

13:00 JST研究/論文

Anomaly Detection and Root Cause Analysis for Microservice Systems

Microservice systems are widely used to build cloud applications, yet their complexity makes failures inevitable, degrading user experience…

13:00 JST研究/論文

Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation

Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on exten…

13:00 JST規制/政策

Market Design for AI: Beyond the Copyright Binary

How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves…

13:00 JST画像/動画生成エージェント

DIMOS: Disentangling Instance-level Moving Object Segmentation

Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving,…

13:00 JST研究/論文

A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction

Predicting the Kohn-Sham Hamiltonian with machine learning can accelerate density functional theory while retaining access to molecular orb…

13:00 JSTLLM/生成AIMistral AI

The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs,…

13:00 JSTLLM/生成AI

Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure

Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and…

13:00 JSTLLM/生成AI

When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking

Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying t…

13:00 JST画像/動画生成

Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks

Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological…

13:00 JST画像/動画生成

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization

Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling…

13:00 JST画像/動画生成研究/論文GPT / ChatGPT

MultiView-Bench: VLM における世界中心のマルチビュー統合のための診断ベンチマーク

VLM の最近のベンチマークは主に単一ビューまたは限定ビューの知覚を評価しており、複数の視点にわたる観察を一貫した世界中心 (他中心) の 3D メンタル モデルに統合する中核となる認知能力がテストされていないままになっています。 MultiView-Bench を紹介します。これは、全体的な 3D シーンを理解するためのマルチビュー統合を評価するために特別に設計された診断ベンチマークです。ピクセルレベルのマッピングやカメラ相対ナビゲーションに焦点を当てた既存のデータセットとは異なり、MultiView-Bench では、モデルが一時的な視点からオブジェクトの位置を切り離し、固定されたグローバル座標系に固定する必要があります。この機能は、VLM が機械部品の組み立てなどの下流タスクに展開される前の前提条件として機能します。フロンティア VLM の体系的な評価により、一貫した障害モードが明らかになりました。つまり、単一画像からの 2D 平面関係では優れたパフォーマンスを発揮しますが、3D 空間関係やビュー全体の情報の集約では顕著な困難が生じます。さらに、型破りな軸方向との闘いや、オブジェクトの色やテクスチャの変化に対する敏感さなど、VLM のバイアスを特定します。これらの制限を認識した上で、私たちは、有益な視点を積極的に選択し、複数の視点からの証拠を認識し、融合するマルチエージェント フレームワークである ViewNavigator を提案します。これにより、予算に合わせた厳密な比較の下でも (完全なエージェントの場合は 3 ~ 5 倍)、MultiView-Bench 上の多様な基本モデルが改善されます。

原文 (English)

MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, which uses active viewpoint selection and evidence fusion to improve four base models by 12.3--20.0 percentage points under a six-image cap matching the fixed-view baseline; budget-extended gains are model-dependent and reach 27 percentage points for GPT-5.

13:00 JSTLLM/生成AI

KV キャッシュの適応フィルタリング: LLM 推論における構造役割バイアスの診断と修正

アテンションベースの KV キャッシュエビクション (H2O とその子孫) は、蓄積されたアテンションの質量 (ここでは信号エネルギーとして扱われます) に基づいてトークンをランク付けし、最も重いものを維持することによって、ロングコンテキスト モデルのメモリ制約状態を圧縮します。ネストされた JSON などのスキーマ密度の高い入力ストリームでは、このスコアはノイズを不釣り合いに保持する非定常フィルターとして機能します。非コンテンツ シンクの役割 (区切り文字または空白) は、どのコンテンツの役割よりも桁違いに多くのエネルギーを運び、構造的な KEY トークンは、回答を運ぶ VALUE トークンの約 1.8 倍の割合で過剰に保持され、完全一致の精度が 88% から 88% に低下します。保持された状態の信号対雑音比が低下するため、5% バジェットでは 0% になります。反事実に基づく実験により、KEY トークンを抑制することが最良の展開可能なフィルターであることが証明されました。単一の調整されたハイパーパラメータによって制御される、SnapKV のウィンドウ スコアに対する再トレーニング不要のロール条件付き割り当ては、20% 未満の予算で H2O ギャップの 63 ~ 98% を埋め、より高い予算ではフル キャッシュの精度と適度に一致またはそれを超えます。これは、シードに依存する小さなノイズ除去効果です (B=0.50 で境界線が有意、4 つのシードにわたる B=0.30 ではゼロと区別できません)。 15 MB の線形ロール プローブは、無視できる推論コストでこれらのラベルを提供しますが、パーサー レベルのダウンストリームのマッチング精度は未解決のままです。

原文 (English)

Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference

Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (delimiters or whitespace) carries an order of magnitude more energy than any content role, and structural KEY tokens are over-retained at roughly 1.8x the rate of the answer-carrying VALUE tokens, collapsing exact-match accuracy from 88% to 0% at a 5% budget as the signal-to-noise ratio of the retained state degrades. A counterfactual experiment establishes that suppressing KEY tokens is the best deployable filter. Our retraining-free, role-conditional allocation over SnapKV's windowed score, governed by a single tuned hyperparameter, closes 63-98% of the H2O gap at sub-20% budgets and, at higher budgets, modestly matches or exceeds full-cache accuracy -- a small, seed-sensitive denoising effect (borderline significant at B=0.50; not distinguishable from zero at B=0.30 over four seeds). A 15 MB linear role probe supplies these labels at negligible inference cost, though matching parser-level downstream accuracy remains open.

13:00 JST画像/動画生成ビジネス/資金調達

Ablation-Corrected Evaluation of Attribution Maps in Echocardiographic Ejection-Fraction Models

Attribution maps for echocardiographic ejection-fraction models are evaluated by their overlap with an expert left-ventricular annotation,…

13:00 JSTLLM/生成AI研究/論文

ビジネス分野全体にわたる最先端の AI パフォーマンス: ナレッジワークと分析的推論の事例に基づいたベンチマーク

大規模言語モデル (LLM) は、ベンチマーク スコアに反映されているように急速に改善されていますが、これらの AI ベンチマークでは主に、事実の再現、限定的な質問応答、数学的問題解決、コーディングやエージェント ツールの使用などの機能がテストされます。まだ十分に測定されていないのは、複雑な情報の統合、不確実性と不完全な情報の下での判断の行使、複数のステークホルダーの状況での戦略的および敵対的思考の適用、トレードオフの比較検討、防御可能な構造化された分析の作成など、ホワイトカラーの専門家が日々行っている分析知識作業における AI の進歩です。このギャップは、そのような仕事の主観的な要素ではさらに顕著であり、成功を定義するのが難しい場合があります。トップクラスのビジネススクールが実践する「ケースメソッド」教育形式は、この測定ギャップに対処するための自然な基盤を提供します。私たちは、18 分野にわたるビジネスケースから抽出された数百の質問にわたるベンチマークである BusinessCaseBench を構築します。各質問は、専門家が作成した講師のケースソリューションから導き出された採点ルーブリックと対になっています。 BusinessCaseBench では、フロンティア AI モデルはすでにインストラクターのルーブリックに対して高いスコアを獲得しており、1 つのモデル ファミリー内の機能は 2 年間で大幅に向上しています。これらの結果は、この種の作業における AI のパフォーマンスがすでに高く、急速に向上していることを示す強力な証拠を提供します。これは、事例教育学によって学部生や MBA がこの種の分析的推論を訓練されるビジネス スクールや、歴史的にそのようなスキルが初期キャリアの仕事に定着してきたエントリーレベルの専門職に影響を及ぼします。

原文 (English)

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.

13:00 JST研究/論文

KAYROS: An Anytime and Exact Open-Source Solver for Duration-Minimization Time-Dependent Vehicle Routing. A Technical Report and a Case Study in Human-AI Engineering

Time-dependent routing recognizes that the same journey can take a different time depending on when it begins. Under duration minimization,…

13:00 JSTLLM/生成AI画像/動画生成研究/論文Claude

放射線科視覚言語モデルベンチマークの法医学的再現性監査: 意図されたプロトコルからリリースされたアーティファクトまで

医療画像 AI ベンチマークは、データセット、DICOM レンダリング、プロンプト、プロバイダー API、自動ラベル、統計コード、原稿、リポジトリ リリースを組み合わせます。これらの成果物間の一致は、通常、テストされるのではなく想定されます。私たちは、保存された胸部 X 線写真視覚言語モデル (VLM) パイロットの遡及的法医学的再現性監査を実施しました。モデルが再び呼び出されることはなく、画像やレポートに新たに注釈が付けられることもありませんでした。プロンプト バインディング、DICOM メタデータ、出力の完全性、ラベル抽出、一致分析、リリースの伝播を追跡しました。計画された 300 件のモデル プロンプト呼び出しのうち、297 件で空ではないレポートが生成されました。 A/B とラベル付けされた 60 個の Claude 呼び出しが、同じ C プロンプトで実行されました。 30件の研究では28人の患者が対象となった。 4 つの MONOCHROME1 画像は必要な極性反転なしでレンダリングされ、データセット分割メンバーシップは保持されず、未検証の抽出プログラムにより 5 つのレポートが 4000 文字に切り詰められました。 369 の完全な症例発見ブロックからなる 1 つの共通コホートを再構成すると、コクランの Q は 154.73 から 182.29 に変化しました。 45 件のマクネマー比較のうち、27 件は未調整で p < 0.05 であり、20 件はホルム調整後も 0.05 未満のままでした。これらの値は、アーカイブされた自動ラベル マトリックスのみを表します。意図した迅速な比較を回復したり、臨床成績を確立したりすることはありません。当社は、元のパフォーマンス、ランキング、即時効果、および臨床上の主張を撤回し、コホート、DICOM レンダリング、プロンプトおよびモデルのアイデンティティ、コール ステータス、アノテーションの出所、キー付き分析、および派生アーティファクトに対して機械検証可能なコントロールを指定します。

原文 (English)

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a retrospective forensic reproducibility audit of a preserved chest-radiograph vision-language model (VLM) pilot; no model was called again and no image or report was newly annotated. We traced prompt bindings, DICOM metadata, output completeness, label extraction, matched analyses, and release propagation. Of 300 planned model-prompt calls, 297 yielded nonempty reports. Sixty Claude calls labeled A/B were executed with the same C prompt. The 30 studies represented 28 patients. Four MONOCHROME1 images were rendered without required polarity inversion, dataset split membership was not retained, and the unvalidated extractor truncated five reports to 4000 characters. Reconstructing one common cohort of 369 complete case-finding blocks changed Cochran's Q from 154.73 to 182.29. Of 45 McNemar comparisons, 27 had unadjusted p < 0.05 and 20 remained below 0.05 after Holm adjustment. These values describe only the archived automated-label matrix; they do not recover the intended prompt comparison or establish clinical performance. We withdraw the original performance, ranking, prompt-effect, and clinical claims and specify machine-verifiable controls for cohort, DICOM rendering, prompt and model identity, call status, annotation provenance, keyed analysis, and derived artifacts.

13:00 JSTLLM/生成AIエージェント

Living-Harness はインタラクティブ エージェントの進化版です

大規模言語モデル (LLM) エージェントは、エピソード内または再試行後に障害から回復する可能性がありますが、エピソード後のフィードバックによって、将来のインタラクションをガイドする永続的なハーネスがほとんど修正されないため、同じ実行障害が後のタスクで再発する可能性があります。静的ハーネスは、固定ツール、コンテキスト、メモリ、ワークフロー構造を通じて信頼性を向上させますが、展開後も変更されません。私たちは $\textbf{Living-Harness}$ を提案します。これは、完了した各軌道とその評価信号を、有界ハーネス更新のための事後証拠に変換する自己進化型エージェント ハーネスです。ドメインレベルの $\textbf{Evolution-SOP}$ ($\textbf{S}$tandard $\textbf{O}$perating $\textbf{P}$rocedure) によって導かれ、Living-Harness はエピソードの抽象化と構造化された更新の証拠を抽出し、トリガー条件、障害パターン、回復アクションを記録するエピソード記憶と、状態ノード、修復エッジ、および遷移ルール​​を記録する状態グラフという、2 つの相補的な形式の手続き的知識を書き込みます。更新されたハーネス状態が取得されて今後の対話をガイドしますが、ツールとベース コンテキストは凍結されたままとなり、進化サイクル全体にわたって手続き的な修復が蓄積されることが可能になります。 $\tau^2$-Bench と MultiWOZ-2.4 から派生した 8 つのインタラクティブ環境では、Living-Harness は最も強力なインタラクティブ ベースラインに対する平均 Pass@1 をそれぞれ 10.07 パーセント ポイントと 9.91 パーセント ポイント改善し、モデル バックボーン全体で進化したハーネス状態の取得のみの再利用をサポートします。

原文 (English)

Living-Harness Is an Interactive-Agent Evolver

Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future interactions. Static harnesses improve reliability through fixed tools, context, memory, and workflow structures, but remain unchanged after deployment. We propose $\textbf{Living-Harness}$, a self-evolving agent harness that converts each completed trajectory and its evaluator signals into posterior evidence for bounded harness updates. Guided by a domain-level $\textbf{Evolution-SOP}$ ($\textbf{S}$tandard $\textbf{O}$perating $\textbf{P}$rocedure), Living-Harness extracts an episode abstraction and structured update evidence, and writes two complementary forms of procedural knowledge: episodic memory that records trigger conditions, failure patterns, and recovery actions, and a state graph that records state nodes, repair edges, and transition rules. The updated harness state is retrieved to guide future interactions, while tools and base context remain frozen, allowing procedural repairs to accumulate across evolution cycles. On eight interactive environments derived from $\tau^2$-Bench and MultiWOZ-2.4, Living-Harness improves average Pass@1 over the strongest interactive baseline by 10.07 and 9.91 percentage points, respectively, and supports retrieval-only reuse of the evolved harness state across model backbones. Our code will be made publicly available soon at https://github.com/anotherbricki/Living-Harness.

13:00 JST研究/論文

The Epistemic Politics of AI Anthropomorphism

AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustain…

13:00 JSTLLM/生成AI

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural…

13:00 JSTLLM/生成AI

Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't

Model families are typically trained size by size, each from scratch. Can apretrained large model instead be converted into a smaller sibli…

13:00 JST研究/論文QwenDeepSeek

Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different s…

13:00 JST画像/動画生成

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-…

13:00 JST画像/動画生成

EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation

Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, pri…

13:00 JSTエージェント研究/論文

MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures

Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their…

13:00 JSTLLM/生成AIビジネス/資金調達

Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions

Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute…

13:00 JSTLLM/生成AIエージェントClaude

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through revie…

13:00 JST画像/動画生成

A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems

Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of…

13:00 JSTロボティクス

SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning

While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by subop…

13:00 JST画像/動画生成

Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation

Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learne…

13:00 JSTLLM/生成AI研究/論文

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place dema…