Skip to the content.

AIニュース 2026-07-01

自動生成: 2026-07-01 13:21 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. How ChatGPT adoption has expandedOpenAI

    New OpenAI Signals data shows how ChatGPT adoption is growing globall…

  2. 「Claude Sonnet 5」新登場 低コストでOpus 4.8に匹敵とうたうも、タスク当たりコスト増加との評価もITmedia AI+

    Anthropicが新モデル「Claude Sonnet 5」を発表した。上位の「Opus 4.8」に迫る性能を低価格で実現するとうたう。…

  3. 「Claude Fable 5」が帰ってくる 「Mythos 5」含む輸出規制解除へ Anthropic発表ITmedia AI+

    Anthropicは6月30日(現地時間)、「Claude Fable 5」「Mythos 5」への輸出規制が解除されたと明らかにした。7…

  4. キーボード入力時の脳の動きから打った文章を割り出す技術、Metaが発表 手術不要、“埋め込み式”に迫るITmedia AI+

    米Metaが、手術を伴わずに脳の活動を文章へ変換する研究「Brain2Qwerty v2」を発表した。頭部に装着する装置で脳の信号を読み取…

  5. Anthropic、科学研究向けAIワークベンチ「Claude Science」を発表──NVIDIAのBioNeMoツールキットと連携ITmedia AI+

    Anthropicは、科学者が計算研究を一貫して行えるAI実行環境「Claude Science」を発表した。データベースやツールを1つの…

  6. AIで“ゲームキャラの出産二次創作”を何千回と生成する人も……ChatGPTの会話57万件から見えたヘビーな利用実態ITmedia AI+

    米ワシントン大学などに所属する研究者らが発表した論文「AI Fiction in the Wild」は、AIチャットとの会話データを分析し…

  7. Anthropic launches Claude Sonnet 5 as a cheaper way to run agentsTechCrunch AI

    Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, low…

トピック別件数

日本語メディア16件

ITmedia AI+ (日本語)

12:18 JSTLLM/生成AIAnthropicClaude

「Claude Sonnet 5」新登場 低コストでOpus 4.8に匹敵とうたうも、タスク当たりコスト増加との評価も

Anthropicが新モデル「Claude Sonnet 5」を発表した。上位の「Opus 4.8」に迫る性能を低価格で実現するとうたう。一方、第三者機関の評価ではトークン使用量が多く、1タスク当たりコストはOpus 4.8を上回るとの指摘もある。

12:05 JSTLLM/生成AI

国産LLM「Sarashina3」登場 高品質データ、独自検証で日本語能力を強化 ソフトバンク傘下

ソフトバンク傘下のSB Intuitionsは、国産LLM「Sarashina」の最新版「Sarashina3シリーズ」の提供を開始。高品質なデータセットや独自の出力結果検証などで日本語能力を強化した。

11:05 JST研究/論文

キーボード入力時の脳の動きから打った文章を割り出す技術、Metaが発表 手術不要、“埋め込み式”に迫る

米Metaが、手術を伴わずに脳の活動を文章へ変換する研究「Brain2Qwerty v2」を発表した。頭部に装着する装置で脳の信号を読み取り、人がキーボードへ入力した文章をリアルタイムで解読する。脳の病気で話す力を失った人の意思疎通を支える技術として、学習用のコードも公開した。

10:50 JSTLLM/生成AI規制/政策AnthropicClaude2媒体が報道

「Claude Fable 5」が帰ってくる 「Mythos 5」含む輸出規制解除へ Anthropic発表

Anthropicは6月30日(現地時間)、「Claude Fable 5」「Mythos 5」への輸出規制が解除されたと明らかにした。7月1日からアクセスを回復し、詳細は近日中に発表するとしている。

出典:ITmedia AI+ITmedia AI+TechCrunch AI
09:54 JSTLLM/生成AIハードウェア/半導体研究/論文AnthropicClaudeNVIDIA

Anthropic、科学研究向けAIワークベンチ「Claude Science」を発表──NVIDIAのBioNeMoツールキットと連携

Anthropicは、科学者が計算研究を一貫して行えるAI実行環境「Claude Science」を発表した。データベースやツールを1つのインタフェースに統合し、文献分析から論文執筆、図表作成まで対応する。NVIDIAのツールキットとも連携し、機密データを外部に送信しない設計が…

09:00 JSTその他

CAD連携AIで設計レビュー工数を最大40%削減、検図や見積作成も自動化

Archaicは、製造業の設計業務を自動化するAIソリューションの販売を開始した。SOLIDWORKSなどのCADと直接連携し、確認作業や検図、見積作成を自動化して、設計者の工数削減を支援する。

08:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

AIで“ゲームキャラの出産二次創作”を何千回と生成する人も……ChatGPTの会話57万件から見えたヘビーな利用実態

米ワシントン大学などに所属する研究者らが発表した論文「AI Fiction in the Wild」は、AIチャットとの会話データを分析し、ユーザーがAIを使ってどれだけフィクションを生成しているのかを調べた研究報告だ。

07:30 JSTその他

日産「AIで再び世界トップの開発力へ」、独自の統合型次世代AIDV基盤描く

日産自動車は「AWS Summit Japan 2026」において、次世代モビリティ「AIDV(AIディファインドビークル)」に向けたクラウド基盤構築の取り組みと、AIを活用したソフトウェア開発環境の今後の展望を語った。

07:00 JSTその他

「ウソだろ」アスクル社長がうなったAI活用 商談準備を2週間→3時間に “担当者のカオス”脱却へ

サイバー攻撃を受けたアスクルが、AIを活用して自社システムを立て直した。旧来の課題を拭い去り、商談時間を短縮するなどの成果を出している。逆境を勝機に変えた舞台裏を、吉岡社長が語った。

07:00 JSTLLM/生成AI

生成AIの請求書、人件費と並べる時代へ 国内5社のAI責任者が語る「トークンマネジメント」の現在地

経費精算SaaSのLayerXやラクス、名刺管理から事業を広げたSansan、会計クラウドのfreee、フリマアプリのメルカリ。取材した5社のAI・人事責任者から、驚くほど重なるトーンでAIのトークンコストを語る声が聞こえてきた。

07:00 JSTその他

謎の「“日の丸AI”開発企業」正体明らかに ソフトバンク、NECら大手がそろって出資するワケ

ソフトバンクやNECなどが出資する「国産AIモデル開発企業」がベールを脱いだ。一体なぜ、国内大手企業が出資するのか。

05:00 JSTその他

【Pythonで学ぶデータ分析】母平均のベイズ推定と予測サンプルの作成 ~ 規格外製品の廃棄コストを見積もる

「不良品で幾ら損する?」をベイズ統計で見積もってみましょう。製品のサイズを測ったデータから、平均やばらつきを推定し、さらに「これから作る製品が規格外になる確率」までをPythonを使って予測します。『社会人1年生から学ぶ、やさしいデータ分析』ベイズ統計編の第4回です。

19:55 JSTその他

農水省の“クソダサ”ポスター話題 「AIよりよっぽど良い」の声も 担当者に狙いを聞いた

農林水産省の公式Xアカウントが6月29日に投稿した「佃煮の日」のポスターが、良い意味で「ダサい」と話題だ。農水省広報室の担当者はITmedia NEWSの取材に応じ、デザインの狙いやSNSの投稿体制について語った。

18:49 JSTロボティクス

AIロボット1000万台導入へ、2040年までに 赤澤経産相が語る「勝ち筋」

赤澤亮正経済産業大臣は6月30日の記者会見で、2040年までにAIを活用したロボットを国内に約1000万台導入する目標を掲げた。18分野での社会実装を進める。

16:43 JSTその他

スクエニ「AI駆動型品質チェックプラットフォーム」開発へ、国が補助金 バンナムやnoteも採択

スクウェア・エニックスの他、バンダイナムコエンターテインメント、NTT西日本、noteなどが採択された。

16:12 JSTロボティクス規制/政策

日印「防衛用AIドローン」共同開発へ 首脳会談で確認、対中念頭に安保協力深化

日印両政府が防衛分野で活用する人工知能(AI)搭載型ドローン(無人機)の共同開発を推進する方針を固めた。高市早苗首相は7月2日にインドでモディ首相との会談を予定しており、防衛装備品協力を加速させることで一致する見通しだ。中国がインド太平洋地域で軍事活動を活発化させる中、日印の安…

海外メディア14件

TechCrunch AI (英語)

12:15 JSTその他Google

The “Father of the Internet” is finally retiring

Vinton Cerf, one of the creators of the protocols underlying the internet, will step down as Google's chief internet evangelist next week.

11:04 JSTビジネス/資金調達

Wayve launches $85M employee tender offer at $8.5B valuation

Wayve’s offering is part of a growing trend of AI startups using employee tenders as a strategic tool to attract and retain talent.

06:53 JSTエージェント

OpenClaw is finally available on Android and iOS

The free open source agentic program is finally invading your phone.

05:33 JST研究/論文Google

The DeepMind trio who built a poker AI are now making money for quant hedge funds

EquiLibre Technologies, a Prague-based AI lab founded by three ex-DeepMind researchers, is now valued at more than $500 million.

04:02 JSTLLM/生成AIその他Google2媒体が報道

Google introduces a faster, cheaper image generator with Nano Banana 2 Lite

Google is updating its image generator to make it faster and cheaper, making it a more useful tool for creators looking to make AI content.

出典:Google DeepMindTechCrunch AI
03:13 JSTハードウェア/半導体ビジネス/資金調達NVIDIA

Nvidia competitor Etched hits $5B valuation, $1B in sales for AI chip

Nvidia AI chip competitor Etched says it has already booked $1 billion under contract for the inference systems powered by its chip.

03:00 JSTLLM/生成AIエージェントAnthropicClaude

Anthropic launches Claude Sonnet 5 as a cheaper way to run agents

Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper al…

02:52 JSTエージェント

Acti puts AI agents directly into your smartphone keyboard

Acti is betting the smartphone keyboard is the next home for AI assistants. The startup's new keyboard for iOS and Android works across app…

02:00 JSTLLM/生成AI研究/論文AnthropicClaude

Anthropic’s Claude Science bets on workflow, not a new model, to win over scientists

Anthropic's Claude Science is a workbench that gives scientists one environment to do computational research, saving them from the need to…

00:08 JSTその他

X now offers an MCP server to make its platform easier for AI tools to use

X has launched a hosted MCP server, making it easier for developers to connect AI applications with the company’s API.

00:00 JSTその他

Podcasting platform Riverside enters the newsletter publishing game

Users will be able use AI to create newsletters based on their recordings.

00:00 JSTLLM/生成AIエージェントAnthropicOpenAI

Amazon launches new $1 billion FDE org, following OpenAI and Anthropic

Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-suffic…

23:00 JSTLLM/生成AI

Lumo, Proton’s privacy-focused AI chatbot, gets an upgrade

Proton's Lumo 2.0 is dropping this week, giving users a broader variety of capabilities.

18:00 JSTエージェント

Crypto exchange OKX wants AI agents to hire and pay each other

OKX is bringing together payments, identity, and reputation into a marketplace for AI agents.

公式ブログ1件

OpenAI (英語)

18:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

How ChatGPT adoption has expanded

New OpenAI Signals data shows how ChatGPT adoption is growing globally, with users increasing usage, exploring more capabilities, and drivi…

Google DeepMind (英語)

新着記事はありませんでした。

論文356件

arXiv cs.AI (英語)

13:00 JST研究/論文

フィードバックによるインタラクティブな改善を促進するものは何ですか?

私たちは、自然言語のフィードバックが、繰り返しの試行だけで得られる成果を超える改善をもたらす場合を研究します。マルチターン言語エージェント設定では、より高い最終精度は有用なフィードバックを反映する可能性がありますが、リサンプリング、形式修正、または追加のテスト時の計算によっても生じる可能性があります。これらの効果を分離するために、Omni-MATH、Codeforces、BBEH Linguini、および ARC-AGI1 にわたって制御された生徒と教師のプロトコルを導入し、生徒と教師の両方の役割で 13 の無重みモデルを評価します。対話履歴、課題の難易度、特権課題情報への教師のアクセスを変化させながら、外部フィードバック、自己フィードバック、ガイドなしの自己改善を比較します。さまざまな設定において、マルチターンの改善は、フィードバックの使用の証拠ではないことが多いことがわかりました。自己生成のフィードバックは、ガイドなしの自己改善以上の効果はほとんどありませんが、最も強力な外部教師は、フィードバック固有の実質的に大きな利益を生み出します。これは、有用なフィードバックが一般的な再試行を超えたガイダンスを提供する必要があることを示唆しています。さらに、密な生徒と教師の相互作用マトリックスは、教師の選択が固定の生徒にとって依然として重要であるにもかかわらず、対話型の利益は教師のアイデンティティよりも生徒のフィードバックを使用する能力によって左右されることを示しています。これらの結果は、フィードバックベースのエージェントは反復試行のベースラインに対して評価されるべきであり、単にフィードバックが利用できるかどうかではなく、フィードバックに基づいて行動する能力が対話型改善の中心的なボトルネックであることを示唆しています。管理された学生と教師の評価フレームワークを https://j-lojek.github.io/フィードバック-世代-is-a-bottleneck/ でリリースします。

原文 (English)

What Drives Interactive Improvement from Feedback?

We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone. In multi-turn language agent setting, higher final accuracy can reflect useful feedback, but it can also arise from resampling, format correction, or additional test-time computation. To separate these effects, we introduce a controlled student-teacher protocol across Omni-MATH, Codeforces, BBEH Linguini, and ARC-AGI1, evaluating thirteen open-weight models in both student and teacher roles. We compare external feedback, self-feedback, and unguided self-refinement, while varying interaction history, task difficulty, and teacher access to privileged task information. Across settings, we find that multi-turn improvement is often not evidence of feedback use: self-generated feedback adds little beyond unguided self-refinement, whereas the strongest external teachers produce substantially larger feedback-specific gains, suggesting that useful feedback must provide guidance beyond generic retry. Dense student-teacher interaction matrices further show that interactive gains are driven more by the student's ability to use feedback than by the teacher's identity, although teacher choice remains important for a fixed student. These results suggest that feedback-based agents should be evaluated against repeated-attempt baselines, and that ability to act on feedback, not merely feedback availability, is a central bottleneck for interactive improvement. We release our controlled student-teacher evaluation framework at https://j-lojek.github.io/feedback-generation-is-a-bottleneck/.

13:00 JSTLLM/生成AIエージェント

反復プロンプト最適化のためのコントラスト反射

LLM エージェントは情報検索の中心となりつつあり、検索クエリを発行し、回答を合成し、IR 評価の審査員としての役割を果たすことが増えています。これらのエージェントを制御するプロンプトを改善することは最適化の問題ですが、適用される IR 設定では、多くの場合、ブラインドサーチではなく、デバッグのように見えます。エンジニアは、どの動作が失敗したか、どの近くの動作が引き続き機能したか、2 つの違いは何なのか、即時編集によりリグレッションを引き起こすことなく保留された品質が向上するかどうかを知る必要があります。エージェントティック IR ワークフローのための反復的なプロンプト最適化フレームワークである Contrastive Reflection を紹介します。このフレームワークはタスク中心の品質定義から始まります。QA エージェントは検索または推論のトレースを公開し、グレーディング エージェントはディメンション レベルのスコアと根拠を公開します。これらの構造化トレースは、エラーにアンカーされた行動スライスを特定し、同じ領域の近くの成功例を追加し、教師 LLM に的を絞ったプロンプト編集を提案するよう依頼するために使用されます。編集候補は、検証パフォーマンスが向上した場合にのみ受け入れられ、オプションで回帰チェックの対象となります。ツリーベースのスライス セレクターを使用してフレームワークをインスタンス化しますが、寄与するのはツリー自体ではなく、対照的な反射ループです。パブリック HotpotQA 検索拡張 QA セットアップでは、1 つのツリーで選択されたコントラスト修復により、保持された完全一致の精度が 51.4% から 60.4% に向上します。失敗のみとランダムな証拠のバリアントでは改善が少なく、以前は正しかった例が多く破られます。簡単な命令のみの比較では、このメソッドは最新のプロンプト オプティマイザーに近いものとなります。MIPROv2 は 59.4%、GEPA は 57.0% に達します。その結果、IR エージェント用の解釈可能な最適化ループが実現し、迅速な修復をより検査可能かつ検証主導型にすることを目的としています。

原文 (English)

Contrastive Reflection for Iterative Prompt Optimization

LLM agents are becoming central to information retrieval: they issue retrieval queries, synthesize answers, and increasingly serve as judges for IR evaluation. Improving the prompts that control these agents is an optimization problem, but in applied IR settings it often looks less like blind search and more like debugging. Engineers need to know which behavior failed, which nearby behavior still worked, what distinguishes the two, and whether a prompt edit improves held-out quality without introducing regressions. We present Contrastive Reflection, an iterative prompt-optimization framework for agentic IR workflows. The framework starts from a task-centric quality definition: QA agents expose retrieval or reasoning traces, and grading agents expose dimension-level scores and rationales. These structured traces are used to identify error-anchored behavioral slices, add nearby successful examples from the same region, and ask a Teacher LLM to propose a targeted prompt edit. Candidate edits are accepted only when validation performance improves, optionally subject to regression checks. We instantiate the framework with a tree-based slice selector, but the contribution is the contrastive reflection loop rather than the tree itself. On a public HotpotQA retrieval-augmented QA setup, one tree-selected contrastive repair improves held-out exact-match accuracy from 51.4% to 60.4%. Failure-only and random-evidence variants improve less and break more previously correct examples. A light instruction-only comparison places the method near modern prompt optimizers: MIPROv2 reaches 59.4% and GEPA 57.0%. The result is an interpretable optimization loop for IR agents, aimed at making prompt repair more inspectable and validation-driven.

13:00 JST研究/論文

AI はどのようにしてモデルを見つけることができるのでしょうか?データ形式、埋め込み、取得戦略を考慮したモデル探索の実験的研究

再利用するシミュレーション モデルを発見することは、モデリングとシミュレーション (M&S) における基本的な課題のままです。多くのモデルが共存する場合、特定のモデリング意図に合致するモデルを特定することは依然として困難です。人工知能 (AI) の最近の進歩、特に検索ベースのアプローチは、このセマンティック層で動作するための有望な経路を提供します。この論文では、自然言語クエリを使用したシミュレーション モデルの発見に対する、データ表現、トランスフォーマー ベースの埋め込みモデル、および検索戦略の影響を調査する実験的研究を紹介します。私たちは、recall@5 や nDCG@5 などの標準的な情報取得メトリクスを使用して、複数のクエリ タイプにわたるパフォーマンスを評価しました。結果は、データ表現が重要であること、オープンソースの埋め込みモデルが高いパフォーマンスを達成できること、特にクエリの複雑さが増すにつれて再ランキング手法が重要であることを示しています。この研究は、AI 主導のモデル発見のベースラインを提供し、AI 主導の構成可能性と相互運用性への前進におけるその役割について説明します。

原文 (English)

How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies

Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identifying those that align with a given modeling intent remains difficult. Recent advances in Artificial Intelligence (AI), particularly retrieval-based approaches, offer a promising pathway to operate at this semantic layer. In this paper, we present an experimental study investigating the impact of data representation, transformer-based embedding models, and retrieval strategies on the discovery of simulation models using natural language queries. We evaluated performance across multiple query types using standard information retrieval metrics, including recall@5 and nDCG@5. Results show that data representation matters, open-source embedding models can achieve high performance, and reranking methods are important, especially as query complexity increases. This work provides a baseline for AI-driven model discovery and discusses its role in advancing toward AI-driven composability and interoperability.

13:00 JSTLLM/生成AI

BayesBench: 複数回にわたる証拠の蓄積における LLM 信念の軌跡の評価

大規模言語モデル (LLM) は通常、複数ターンの会話で展開され、各ターンで環境に関する認識論的不確実性を軽減する新しい証拠が提供されます。したがって、合理的に行動するには、それを支配する観察されていない量を推測し、証拠が蓄積されるにつれてそれらについての信念を更新する必要があります。しかし、ほとんどの評価はモデルの最終ターンの解答を 1 ターン形式で採点するだけで、このプロセスは検討されていません。私たちは、LLM の信念の更新が、マルチターン設定における合理的なベイズ推論者の信念の更新とどの程度一致するかを尋ね、3 つの段階的に複雑なタスクにわたってこれを調査する一連のシミュレーション環境である BayesBench を紹介します。 (i) ベイズ推定。モデルは、連続する証拠から未知のパラメーターを推測します。 (ii) ベイジアン予測。モデルは潜在変数に関する推測された信念を結果の予測に変換します。 (iii) 潜在フレーム化ベイジアン予測では、観察結果がユーザーペルソナのフレーム化を通じてフィルタリングされ、潜在状態とペルソナに対する共同推論が必要になります。 7 つの LLM (3B ~ 70B) 全体で、スケーリングにより潜在推論と証拠の蓄積が向上し、更新はベイズ事後分布と一致することがあります。ただし、これらの利益は下流の予測に確実に引き継がれるわけではなく、潜在構造の推論とそれを使用してターゲットの結果についての信念を合理的に更新することとの間にギャップが露呈します。

原文 (English)

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environment. Acting rationally then requires inferring the unobserved quantities that govern it and updating beliefs about them as evidence accumulates. Yet most evaluations only score the model's final-turn answer in a single-turn format, leaving this process unexamined. We ask how closely LLMs' belief updates match those of a rational Bayesian reasoner in multi-turn settings, and introduce BayesBench, a suite of simulation environments that probe this across three progressively complex tasks: (i) Bayesian estimation, where the model infers an unknown parameter from sequential evidence; (ii) Bayesian prediction, where the model turns inferred beliefs about a latent variable into outcome forecasts; and (iii) latent-framed Bayesian prediction, where observations are filtered through a user-persona framing, requiring joint inference over the latent state and the persona. Across seven LLMs (3B--70B), scaling improves latent inference and evidence accumulation, with updates occasionally matching the Bayesian posterior. However, these gains do not reliably carry over to downstream prediction, exposing a gap between inferring latent structure and using it to rationally update beliefs about the target outcome.

13:00 JSTLLM/生成AIDeepSeek

やめることを学ぶことが役立つのはいつですか?推論モデルにおける早期終了に関するコストを意識した研究

推論モデルはインスタンスごとに異なる量の有用な計算を費やしますが、学習された停止ルールがいつ単純な信頼度や収束のしきい値を超えて改善するかは不明のままです。私たちは、推論言語モデル用の隠れステートフリー チェックポイント ストッパーである LearnStop を使用してこの疑問を研究します。固定予算チェックポイントで、LearnStop は現在の推論プレフィックスから短い回答を調査し、回答の信頼性、エントロピー、プレフィックス投票シェア、回答の安定性、バックトラッキング マーカー密度などのオンライン機能からプレフィックスの正しさを予測します。 GSM8K、MATH-500、MMLU-Pro、AIME-90、GPQA、Qwen3、DeepSeek-R1 蒸留にわたる 18 のタスク モデル設定全体にわたって、答えはタスクによって異なります。自由形式の計算では、学習された複数特徴停止により固定予算フロンティアが改善され、多くの場合スカラー出口を上回ります。Qwen3-32B を備えた GSM8K では、経験的フロンティアは +0.157 のポストホック ピーク適応ゲインに達し、検証で選択された操作点は正のゲインを保持し、最も強いスカラー ベースラインを超えるペアのゲインは +0.028 です。複数選択の非常に難しい設定では、スカラー信頼性、エントロピー、または安定性ルールが競合するか、より強力になります。したがって、学習停止をスカラー出口の普遍的な代替としてではなく、その値が軌道構造に依存するツールとして組み立てます。さらに、検証で選択された動作ポイント、ペア ブートストラップ テスト、有限グリッド ロストコレクト リスク校正、KV フォーク、プレフィックス キャッシュ、およびブラック ボックス レジームに基づくコスト計算、H100 サービング プロファイル、チェックポイント スケジュール スイープ、転送分析、および堅牢性チェックを提供します。主な実用的な発見は、学習された停止が、予算がいっぱいになる前に多くの問題が正解したが、信頼できるスカラー停止信号が 1 つも示されない場合に役立つということです。信頼度または答えの収束が停止問題をすでに解決している場合、その利点はほとんど失われます。

原文 (English)

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

Reasoning models spend different amounts of useful computation across instances, but it remains unclear when a learned stopping rule improves over simple confidence or convergence thresholds. We study this question with LearnStop, a hidden-state-free checkpoint stopper for reasoning language models. At fixed budget checkpoints, LearnStop probes a short answer from the current reasoning prefix and predicts prefix correctness from online features such as answer confidence, entropy, prefix vote share, answer stability, and backtracking-marker density. Across 18 task-model settings spanning GSM8K, MATH-500, MMLU-Pro, AIME-90, GPQA, Qwen3, and DeepSeek-R1 distillations, the answer is task-dependent. On free-form math, learned multi-feature stopping improves the fixed-budget frontier and often beats scalar exits: on GSM8K with Qwen3-32B, the empirical frontier reaches a post-hoc peak adapt gain of +0.157, validation-selected operating points preserve positive gains, and the paired gain over the strongest scalar baseline is +0.028. On multiple-choice and very hard settings, scalar confidence, entropy, or stability rules are competitive or stronger. We therefore frame learned stopping not as a universal replacement for scalar exits, but as a tool whose value depends on trajectory structure. We further provide validation-selected operating points, paired bootstrap tests, finite-grid lost-correct risk calibration, cost accounting under KV-fork, prefix-cache, and black-box regimes, H100 serving profiles, checkpoint-schedule sweeps, transfer analyses, and robustness checks. The main practical finding is that learned stopping is useful when many questions become correct before full budget but do not exhibit a single reliable scalar stopping signal; its benefits largely disappear when confidence or answer convergence already solves the stopping problem.

13:00 JSTエージェント

エキスパート ユーザーを超えて: エージェントは、ユーザーが好みを引き出すだけでなく、ユーザーが好みを構築できるように支援する必要があります。

通常、エージェントは専門ユーザー (自分が望むものについて明確な好みを持っているユーザー) を想定しており、タスクの指定が不十分な場合は常に質問を明確にするようデフォルト設定されています。私たちは、この仮定は非現実的であると主張します。ユーザーは多くの場合、好みを完全に指定するためのドメイン知識が不足しています。ある機能の好みについて尋ねられた場合、ユーザーは、例や説明などを通じて、その機能の好みを形成するために必要なドメイン知識をユーザーが学習できるようにエージェントが支援しなければ、答えることができない場合があります。これらの原則を形式化するために、情報経済学の Search-Experience-Credence フレームワークを利用して、ユーザーがエージェントの対話アクションに基づいて好みを構築する方法のモデルである CoPref を導入します。次に、これらのアイデアをエージェント レコメンダー システムで具体的に研究し、インタラクティブなベンチマークである CoShop を提案します。 CoShop では、エージェントが CoPref ユーザーと会話し、CoPref ユーザーに対して推奨事項を作成します。エージェントのパフォーマンスは、ユーザーがタスクを適切に指定するために必要な知識を得るのに役立つかどうかによって決まります。 5 つのフロンティア モデルを評価すると、5 ターンの対話にもかかわらず、CoShop で 56% の精度を超えるエージェントは存在しないことがわかりました。失敗の原因は、エージェントがアイテムを見つける能力にあるのではなく、インタラクションによってユーザーが欲しいものについて知っている範囲がほとんど広がっていないことに起因します。

原文 (English)

Beyond expert users: agents should help users construct preferences, not just elicit them

Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified. We argue this assumption is unrealistic. Users often lack the domain knowledge to have completely specified preferences; if asked about their preference on some feature, the user may be unable to answer without the agent helping the user to learn some domain knowledge needed to form a preference for that feature, e.g., via examples or explanations. To formalize these principles, we draw on the Search-Experience-Credence framework from Information Economics to introduce CoPref, a model of how users construct preferences based on agent dialog actions. We then study these ideas concretely in agentic recommender systems, proposing CoShop, an interactive benchmark. In CoShop, an agent converses with and makes recommendations for a CoPref user. The agent's performance depends on whether it can help the user gain the knowledge needed to specify the task well. Evaluating five frontier models, we find that no agent exceeds 56% accuracy on CoShop despite five turns of interaction. Failures stem not from agents' ability to find items, but from how little the interaction expands what users know about what they want.

13:00 JSTエージェント

法律における複数の代理人による審議の調査

人工知能は法律分野への応用が増えており、司法へのアクセスを増やす可能性があります。特に注目を集めている動きの 1 つは、大規模言語モデル (LLM) に基づいた AI エージェントが自律的なアクションを実行できるエージェント AI です。特に、法律分野におけるマルチエージェントのアプローチは、ほとんど解明されていないままです。この論文では、LLM を使用した法的推論タスクに対するマルチエージェントの審議方法を調査します。私たちはマルチエージェント熟議 (MAD) を探求し、法廷手続きと法的弁論にインスピレーションを得た 2 つの新しいマルチエージェントフレームワークを紹介します。法的ベンチマークと非法的ベンチマークの両方に関する実験では、マルチエージェント フレームワークがベースラインの大規模言語モデルと同等の全体的なパフォーマンスを達成しながらも、大きく異なる答えが得られることが明らかになりました。注目すべきことに、これらのアプローチはベースラインで対処できないケースをうまく解決でき、またその逆も可能です。私たちは定性的評価を実施し、マルチエージェント フレームワークがモノリシック アプローチよりも優れたパフォーマンスを発揮するシナリオを強調します。たとえば、マルチエージェントアプローチは、複数の視点からの批判的思考を必要とする質問に答えるのに適していると考えられます。私たちの研究では、マルチエージェント システムを法律分野における AI の有望な方向性として位置づけるとともに、法律にインスピレーションを得たマルチエージェントによる審議アプローチの可能性を実証しています。

原文 (English)

Investigating Multi-Agent Deliberation in Law

Artificial Intelligence is increasingly applied to the field of law, and has the potential to increase access to justice. One particular movement that is gaining traction is that of agentic AI, wherein AI agents, based on Large Language Models (LLMs) can take autonomous actions. In particular, multi-agent approaches in the legal domain remain largely unexplored. In this paper, we investigate multi-agent deliberation methods for legal reasoning tasks using LLMs. We explore multi-agent deliberation (MAD) and introduce two novel multi-agent frameworks inspired by courtroom procedures and legal argumentation. Our experiments on both legal and non-legal benchmarks reveal that multi-agent frameworks achieve comparable overall performance to baseline large language models, but produce significantly distinct answers. Notably, these approaches can successfully solve cases that the baseline fails to address, and vice versa. We conduct a qualitative evaluation and highlight scenarios where multi-agent frameworks outperform monolithic approaches. For example, multi-agent approaches appear better suited for answering questions that require critical thinking from multiple perspectives. Our work positions multi-agent systems as a promising direction for AI in the legal domain, while demonstrating the potential of law-inspired multi-agent approaches for deliberation.

13:00 JSTエージェントClaude

なぜ 2 回解決するのか?伝達効率の高い ML エンジニアリングのためのスキルの階層的蓄積

すべての競争はコールド スタートであるため、ML エンジニアリング エージェントは既知の技術を再発見するために計算を無駄にします。我々は、競合他社にまたがる知識を 3 つのスコープ層 (グローバル、ドメイン、および競合固有) に編成し、それぞれがマッチング エージェント レベルに関連付けられた階層型マルチエージェント システムである HASTE を紹介します。オーケストレーターはドメイン スペシャリストを調整し、LLM 駆動の抽象化を通じて層間の学習を促進します。制御されたアブレーションは、範囲指定されたローディングの証拠を提供します。8 つの競技会にわたって 159 のスキル インベントリを一定に保持し、段階的ローディングでは 100% のメダル率を達成しますが、フラット ローディングでは 62.5% にしか達せず、スキルをロードしない場合と同じメダル率であり、出力トークンの 2 倍を消費します。 MLE-Bench Lite ベンチマーク全体 (Kaggle コンペティション 22 件) では、HASTE はクロード ソネット 4.6 を使用してコンペごとに 12 時間で 77.3% のメダル率に達しました。コールドスタート実行では、システムはスキルが蓄積されていない状態で開始されます。ウォーム スタート ランでは、以前の競技会で学んだスキルを再ロードし、競技会間での転送にはグローバル レベルおよびドメイン レベルのスキルのみを使用します。ウォーム スタートでは、改良の反復回数が 52% 減少し、提案された変更のうちエージェントが保持する割合は、在庫が少ない場合の 42% から、50 以上のスキルが利用可能になると 85% に増加します。これらの結果は、より優れた知識組織が ML エンジニアリング エージェントのモデルの強度と計算予算の一部を置き換えることができることを示唆しています。

原文 (English)

Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering

ML engineering agents waste compute rediscovering known techniques because every competition is a cold start. We present HASTE, a hierarchical multi-agent system that organizes cross-competition knowledge into three scope tiers (global, domain, and competition-specific), each coupled to a matching agent level. An orchestrator coordinates domain specialists and promotes learning between tiers via LLM-driven abstraction. A controlled ablation provides evidence for scoped loading: holding a 159-skill inventory constant across 8 competitions, tiered loading achieves a 100% medal rate while flat loading reaches only 62.5%, the same medal rate as loading no skills, and consumes 2x the output tokens. On the full MLE-Bench Lite benchmark (22 Kaggle competitions), HASTE reaches a medal rate of 77.3% using Claude Sonnet 4.6 at 12h per competition. In a cold-start run, the system begins with no accumulated skills. In warm-start runs, it reloads skills learned from earlier competitions, using only global and domain-level skills for transfer across competitions. Warm starts use 52% fewer refinement iterations, and the fraction of proposed changes kept by the agent rises from 42% at low inventory to 85% once 50+ skills are available. These results suggest that better knowledge organization can partly substitute for model strength and compute budget in ML-engineering agents.

13:00 JSTLLM/生成AIMistral AI

RoPoLL: LLM 審査員からなる堅牢なパネル

コンセンサススコアを報告する LLM 評価者パネル (PoLL) である LLM 陪審は、単一の裁判官による LLM 評価に代わる実用的な手段となっていますが、その統計的動作は依然として十分に理解されていません。我々はフーバー汚染モデルに基づいて LLM 陪審を形式化し、陪審員の規模に関係なく、偏った LLM の典型的な方法 (モード崩壊、お調子者、安全拒否) で 1 人の裁判官が不合格になるたびに、PoLL が肯定的な汚染の下で無制限のバイアスを被ることを示します。陪審のコンセンサスを古典的なロバストな平均推定として組み立て、RoPoLL (ロバスト パネル オブ LLM-as-Judge) を提案します。これは、PoLL パネルを保持しますが、集計関数を、幾何学的中央値 (GM) でインスタンス化されたロバストな平均推定器に置き換えます。調整不要で、最適な有限サンプルの内訳点 1/2 を持ちます。有限サンプル誤差限界と一致する情報理論的ミニマックス下限は、パラメトリック レート sigma*sqrt(d/N) に関しては一致しますが、内訳の下限では sqrt(d) 倍異なります。sqrt(d) は、扱いにくい Tukey 半空間中央値に対して多項式時間 RoPoLL が支払う統計計算上のギャップです。 13 人の無差別裁判官 (4B-675B)、3 つの報酬モデル ベンチマーク、および最大 50% の割合での 4 つの汚職体制にわたって、RoPoLL はあらゆる偏った汚職タイプで PoLL を支配しており、マッチしたコンピューティングでの次元を越えた攻撃では約 19%、ヘビーテールのビザンチン敵対者では桁違いに優れています。 38B の 3 人の審査員からなる RoPoLL 委員会は、30% のバイモーダルランダム破損のもとで、HelpSteer-2 で Mistral-Large-3 (675B) を 1.31 倍上回りました。これは、より高い精度で 18 倍のパラメーターの利点です。 Noisy-GT コントロールは、良性の不正確さではなく、偏った汚染に対してプレミアムが支払われていることを確認します。

原文 (English)

RoPoLL: Robust Panel of LLM Judges

The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood. We formalize the LLM Jury under the Huber contamination model and show that PoLL incurs unbounded bias under any positive contamination, regardless of jury size, whenever a single judge fails in a biased, LLM-typical way (mode collapse, sycophancy, safety refusal). Framing jury consensus as classical robust mean estimation, we propose RoPoLL (Robust Panel of LLM-as-Judge), which preserves the PoLL panel but replaces the aggregation function with a robust mean estimator, instantiated with the geometric median (GM): tuning-free, with the optimal finite-sample breakdown point 1/2. A finite-sample error bound and a matching information-theoretic minimax lower bound agree on the parametric rate sigma*sqrt(d/N) and differ on the breakdown floor by a factor of sqrt(d), a statistical-computational gap that polynomial-time RoPoLL pays relative to the intractable Tukey halfspace median. Across 13 open-weight judges (4B-675B), three reward-model benchmarks, and four corruption regimes at rates up to 50%, RoPoLL dominates PoLL on every biased corruption type: by about 19% on cross-dimensional attacks at matched compute, and by orders of magnitude on heavy-tailed Byzantine adversaries. A 3-judge RoPoLL committee at 38B beats Mistral-Large-3 (675B) by 1.31x on HelpSteer-2 under 30% bimodal-random corruption, an 18x parameter advantage at better accuracy; a Noisy-GT control confirms the premium is paid against biased contamination, not benign imprecision.

13:00 JSTエージェント

AgRefactor: HLS の互換性とパフォーマンスのための自己進化するエージェント ワークフロー

高位合成 (HLS) は、概念からシリコンへの迅速なパスを提供しますが、言語サポートの制限と、ソフトウェアとハ​​ードウェアのプログラミング実践の間のギャップにより、現実世界のソフトウェアを合成可能な HLS コードに変換することは依然として困難です。既存の自動化された LLM ベースのリファクタリング アプローチは、この問題に部分的に対処していますが、多くの場合、柔軟性に欠け、拡張が困難で、高い計算コストが発生します。ソフトウェアを HLS 互換プログラムにリファクタリングするための LLM ベースのマルチエージェント ワークフローである AgRefactor を紹介します。 AgRefactor には、タスク全体にわたって事実および戦略的な知識を蓄積および取得する自己進化型メモリ システムが組み込まれており、目に見えないプログラムの堅牢性と効率が向上します。コストを削減し、スケーラビリティを向上させるために、自動リファクタリング ツールが統合されており、エージェントは LLM 主導の書き換えと効率的なツールベースの変換のバランスを取ることができます。以前の研究で調査した最も複雑なケースよりも 5 ~ 10 倍長い、11 件の困難な現実世界のベンチマークのうち 9 件で、AgRefactor は、最先端の自動リファクタリング ツールと、同じフレームワーク バックボーン上に構築された強力な LLM ベースのベースラインを上回るパフォーマンスまたは同等のパフォーマンスを示しました。さらにエージェントのパフォーマンスを最適化すると、20% 未満の追加リソースで、SoTA プラグマ チューニング ツールと比較して幾何平均で 6.51 倍の速度向上が得られ、最適化されたオープンソース設計と比較して 1.20 倍の速度向上が得られます。 AgRefactor は完全に自動化されており、オープンソースです。

原文 (English)

AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance

High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challenging due to restrictive language support and the gap between software and hardware programming practices. Existing automated and LLM-based refactoring approaches partially address this problem, yet they often lack flexibility, struggle to scale, and incur high computational costs. We introduce AgRefactor, an LLM-based multi-agent workflow for refactoring software into HLS-compatible programs. AgRefactor incorporates a self-evolving memory system that accumulates and retrieves factual and strategic knowledge across tasks, improving robustness and efficiency on unseen programs. To reduce cost and enhance scalability, it integrates automated refactoring tools, enabling agents to balance LLM-driven rewrites with efficient tool-based transformations. On 9 out of 11 challenging real-world benchmarks, which are 5-10x longer than the most complex cases studied in prior work, AgRefactor outperforms or matches the state-of-the-art automated refactoring tool and a strong LLM-based baseline built on the same framework backbone. Further agentic performance optimization yields a 6.51x geometric mean speedup over the SoTA pragma tuning tool and a 1.20x speedup over optimized open-source designs with less than 20% extra resources. AgRefactor is fully-automated and open-sourced.

13:00 JST研究/論文

ニューロ・ベイジアン・シンボリック残留注意浅いネットワーク: サイバーセキュリティ・リスク評価のための説明可能な深層学習

オープンソース エコシステムにおける説明可能なサイバーセキュリティ リスク評価のためのハイブリッド ニューラル アーキテクチャである、ニューロ ベイジアン シンボリック残余注意浅いネットワーク (NBS-RASN) を紹介します。解釈可能性と引き換えに正確さを求める深いモデルとは異なり、私たちの浅いネットワークは、ドメイン知識、因果推論、専門家の判断を微分可能なコンポーネントとしてエンコードします。これは、精度、因果関係、反証可能性、透明性、完全性という 5 つの認識論的公理を伝播前の厳しい制約として強制するゲートキーパーを含む、12 層にわたる 80 個の解釈可能なニューロンを使用します。深さが限られているにもかかわらず、ネットワークは残留注意とフィードバック ループを通じて深層学習の特性を示し、ブラック ボックスになることなく複雑なリスク パターンを学習します。それは完全に分解可能なスコアを生成します: 決定論的な重み付けコンポーネントと専門家による調整。各調整は名前付き増幅器 (爆発半径、伝播速度、構造的性質、デフォルトの暴露、悪用パターン、制度的重要性) に追跡可能です。 OWASP Top 10:2025 カテゴリと言語リスク クラスをすべてカバーする 20 のオープンソース プロジェクトで検証し、信頼スコア 0.79 ~ 0.97 を達成し、説明可能性がトレーニング アルゴリズムではなく設計によって保証されていることを示します。これは、深層学習には深層ネットワークが必要であるという前提に疑問を投げかけ、解釈可能性が不可欠な一か八かのサイバーセキュリティにおいて、深い推論を備えた浅層ネットワークが不透明なモデルよりも優れたパフォーマンスを発揮できることを証明しています。

原文 (English)

Neuro-Bayesian-Symbolic Residual Attention Shallow Network: Explainable Deep Learning for Cybersecurity Risk Assessment

We introduce the Neuro-Bayesian-Symbolic Residual Attention Shallow Network (NBS-RASN), a hybrid neural architecture for explainable cybersecurity risk assessment in open-source ecosystems. Unlike deep models that trade interpretability for accuracy, our shallow network encodes domain knowledge, causal reasoning, and expert judgment as differentiable components. It uses 80 interpretable neurons across 12 layers, including a gatekeeper that enforces five epistemological axioms - precision, causality, falsifiability, transparency, and completeness - as hard constraints before propagation. Despite limited depth, the network exhibits deep-learning traits via residual attention and feedback loops, learning complex risk patterns without becoming a black box. It produces fully decomposable scores: a deterministic weighted component plus an expert adjustment, with each adjustment traceable to named amplifiers (blast radius, propagation speed, structural nature, default exposure, exploitation pattern, institutional criticality). We validate on 20 open-source projects covering all OWASP Top 10:2025 categories and language risk classes, achieving confidence scores of 0.79-0.97, and show that explainability is guaranteed by design, not by a training algorithm. This challenges the assumption that deep learning requires deep networks, proving that shallow networks with deep reasoning can outperform opaque models in high-stakes cybersecurity, where interpretability is essential.

13:00 JSTエージェント

HyPOLE: 部分観察下でのハイパープロパティ誘導マルチエージェント強化学習

正式な仕様は、学習プロセスをガイドする強力なツールであり、報酬形成に比べて次のような大きな利点をもたらします。(1) 数学的厳密さ。 (2) 目的と制約を指定する表現力、および (3) 目的を達成するための戦術を定義する能力。ただし、これらの利点は、マルチエージェント強化学習 (MARL) の文脈ではほとんど解明されていません。この論文では、部分可観測性の下での MARL の新しいフレームワークである HyPOLE を紹介します。このフレームワークでは、いわゆるハイパープロパティ、特に時相論理 HyperLTL の表現力によって学習が導かれます。当社は、分散型ポリシーを統合するために、分散型実行のための集中トレーニング (CTDE) 技術を HyPOLE と統合しており、SMAC、MessySMAC、および WildFire ベンチマークでの評価では、ベースラインを上回る明らかな利点が実証されています。

原文 (English)

HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation

Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives. However, these benefits remain largely unexplored in the context of Multi-Agent Reinforcement Learning (MARL). This paper introduces HyPOLE, a novel framework for MARL under partial observability, where learning is guided by the expressive power of the so-called hyperproperties and, in particular, the temporal logic HyperLTL. We integrate Centralized Training for Decentralized Execution (CTDE) techniques with HyPOLE to synthesize decentralized policies, and our evaluation on SMAC, MessySMAC, and WildFire benchmark demonstrates clear advantages over baselines.

13:00 JSTエージェントClaude

AgentBound: 自律型 AI エージェントのための検証可能な行動ガバナンス

自律型 AI エージェントは、人間のプリンシパルに代わって、金融取引、外部通信、企業ワークフローなどの結果的なアクションを実行することが増えています。既存のエージェント インフラストラクチャは、ワークロードの認証とリソース アクセスの制御を ID フェデレーションと委任された承認に依存していますが、現在の動作および運用コンテキストで承認されたアクションを実行する必要があるかどうかを判断できません。自律型 AI エージェントに検証可能な動作監視を提供するランタイム ガバナンス フレームワークである AgentBound を紹介します。 AgentBound は、委任された承認、所有者が署名した行動規定、およびサイト アクション契約という 3 つの独立した権限を使用して、提案された各アクションを評価します。彼らの判断は、実行前にアクションを許可するか、検討するか、拒否するかを決定するための正式な意思決定モデルを通じて保守的に構成されています。説明責任を提供するために、AgentBound は、すべてのアクションを正確な委任、ポリシー、意思決定を管理するセマンティック アーティファクトにバインドする暗号的に検証可能なガバナンス レシートを生成し、独立したリプレイ検証とポリシーの来歴を可能にします。このフレームワークでは、長期実行エージェントに対する常駐委任も導入されており、取り消し可能性と制限された権限を維持しながら、継続的に更新されるガバナンス ポリシーの下で定期的なワークロードを実行できるようになります。正式な基盤、システム アーキテクチャ、ガバナンス受信プロトコル、およびガバナンスの正確性、権限構成、説明責任を評価するためのベンチマーク フレームワークである AgentBound-Bench を紹介します。 AgentBound は、モデルの調整を置き換えるのではなく、承認と実行の間に決定論的なガバナンス層を提供することでそれを補完し、ガバナンスを信頼する必要があるプロセスから独立して検証できるプロセスに変換します。

原文 (English)

AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents

Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications, and enterprise workflows. Existing agent infrastructure relies on identity federation and delegated authorization to authenticate workloads and control resource access, but it cannot determine whether an authorized action should be executed under the current behavioral and operational context. We present AgentBound, a runtime governance framework that provides verifiable behavioral oversight for autonomous AI agents. AgentBound evaluates each proposed action using three independent authorities: delegated authorization, owner-signed behavioral constitutions, and site action contracts. Their judgments are conservatively composed through a formal decision model to determine whether an action should be permitted, reviewed, or denied before execution. To provide accountability, AgentBound generates cryptographically verifiable governance receipts that bind every action to the exact delegation, policy, and semantic artifacts governing the decision, enabling independent replay verification and policy provenance. The framework also introduces standing delegation for long-running agents, allowing periodic workloads to operate under continuously refreshed governance policies while preserving revocability and bounded authority. We present the formal foundation, system architecture, governance receipt protocol, and AgentBound-Bench, a benchmark framework for evaluating governance correctness, authority composition, and accountability. Rather than replacing model alignment, AgentBound complements it by providing a deterministic governance layer between authorization and execution, transforming governance from a process that must be trusted into one that can be independently verified.

13:00 JSTエージェント

規制に記憶があるとき: 人工機関におけるヒステリシスと制御負荷

通常、適応エージェントはその動作によって判断されますが、エージェントは安定しているように見える一方で、安定を維持するために必要な内部の努力が増加している可能性があります。この隠れた規制の負担は、ノイズ、遅延、または変化する要求の下で動作する人工エージェントにとって重要です。2 つのシステムは同様の内部状態に到達する可能性がありますが、1 つはそこに到達するためにより多くの修正制御を必要とします。ここでは、その負担が歴史に依存するかどうかを検討します。適応型不確実性制御の計算モデルを使用して、人工エージェントをその不確実性目標の連続的な変化によって駆動し、その後、エージェントをリセットすることなく変化を元に戻します。これにより、キャリーオーバーの簡単なテストが作成されます。コントローラーは現在のターゲットにのみ応答するのか、それともエージェントがそのターゲットに到達するまでのパスが依然として重要なのか?シミュレーションでは、明らかに履歴に依存する影響が示されています。エージェントを制御するために必要な適応ゲインは、再現可能なヒステリシス ループを形成します。これは、エージェントがより要求の厳しいレジームに向かって移動しているか、またはそこから戻っているかに応じて、同じターゲットでも異なるレベルの制御が必要になる可能性があることを意味します。規制のタイミングも重要です。外乱にさらされる前に安定化が利用できる場合、エージェントは通常、外乱がすでに作用した後にのみ回復できる場合よりも、必要な適応ゲインが少なくなります。状態レベルのコヒーレンス測定でもパス依存性が示されますが、タイミングの影響は規制ゲインの方がはるかに明確です。したがって、主な違いは、予期的規制が完全に異なる状態を生み出すということではありません。むしろ、モデル化された制御要求が低くなり、同等の規制された動作に達します。これらの結果は、適応エージェントが組織化されたままであるかどうかだけでなく、そのためにどの程度の規制を採用する必要があるかによって評価されるべきであることを示唆しています。

原文 (English)

When Regulation Has Memory: Hysteresis and Control Burden in Artificial Agency

Adaptive agents are usually judged by what they do, but an agent can appear stable while the internal effort required to keep it stable is increasing. This hidden regulatory burden matters for artificial agents operating under noise, delay, or changing demands: two systems may reach similar internal states while one requires much more corrective control to get there. Here, we study whether that burden depends on history. Using a computational model of adaptive uncertainty regulation, we drive an artificial agent through a continuous change in its uncertainty target and then reverse the change without resetting the agent. This creates a simple test for carryover: does the controller respond only to the current target, or does the path by which the agent reached that target still matter? The simulations show a clear history-dependent effect. The adaptive gain required to regulate the agent forms a reproducible hysteresis loop, meaning that the same target can require different levels of control depending on whether the agent is moving toward or returning from a more demanding regime. The timing of regulation also matters. When stabilization is available before disturbance exposure, the agent generally requires less adaptive gain than when it can only recover after disturbance has already acted. The state-level coherence measure also shows path dependence, but the timing effect is much clearer in regulatory gain. The main difference is therefore not that anticipatory regulation produces a completely different state. Rather, it reaches comparable regulated behavior with lower modeled control demand. These results suggest that adaptive agents should be evaluated not only by whether they remain organized, but by how much regulation they must recruit to do so.

13:00 JST研究/論文

税金を意識したパーソナライズされたポートフォリオ管理のための 3 段階の基礎モデル

我々は、これまでの金融 RL 作業すべてに共通する 3 つの制限、1) ティッカー ロックイン、2) モノリシック目標、3) 静的ユーザー モデルに対処する、パーソナライズされたポートフォリオ管理のための 3 段階の深層強化学習システムを紹介します。フェーズ 1 では、マルチアセット コーパスでの自己教師あり学習を介して、ティッカー ID のないクロスアセット エンコーダーを事前トレーニングします。学習されたゲート メカニズムを介して融合された、T5 ベースの時系列基盤モデルである Chronos を使用した凍結並列ブランチによって強化されます。私たちの知る限り、これはポートフォリオ管理 RL への時系列基礎モデルの最初の適用です。エンコーダーは、新しいティッカーの再トレーニングを必要としない、50 次元の観察可能なメタデータ ベクトルを介して、あらゆる公開取引資産に一般化します。フェーズ 2 では、エピソードごとにサンプリングされた 6 つの異なる投資目標 (短期アルファ、短期利益、長期利益、資本保全、税金損失の回収、および長期利益のみ) を同時に提供する目的条件付き報酬の下で、MoE (専門家混合) のポートフォリオ アクター評論家を PPO で微調整します。 MoE アーキテクチャは、各目標を専門のエキスパート ヘッド (勢い、成長、守備、税務意識) に割り当て、学習型インテント ルーターがアクティブな目標と現在の市場体制に基づいて専門家をブレンドすることで、目標間の勾配の競合を排除します。フェーズ 3 では、実際の証券取引履歴に基づいて微調整された 76 パラメーターの LoRA モジュールを介して推論時に各個人にさらに適合する軽量のパーソナライゼーション レイヤーを追加し、アンケートではなく明らかになった取引行動から投資目標を推測します。自然言語インテント パーサーは、自由形式の目標を構造化された投資目標パラメータに直接変換します。

原文 (English)

A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management

We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared by all prior financial RL work: 1) ticker lock-in, 2) monolithic objectives , and 3) static user models. Phase 1 pretrains a ticker-identity-free cross asset encoder via self-supervised learning on a multi-asset corpus, augmented by a frozen parallel branch using Chronos, a T5-based time series foundation model, fused via a learned gating mechanism. To our knowledge, this is the first application of a time series foundation model to portfolio management RL. The encoder generalizes to any publicly traded asset via a 50-dimensional observable metadata vector that requires no retraining for new tickers. Phase 2 fine-tunes a MoE (Mixture of Experts) portfolio actor critic with PPO under an objective-conditioned reward that simultaneously serves six distinct investment goals sampled per episode: short-term alpha, short-term gain, long-term gain, capital preservation, tax-loss harvesting, and long-term-gains-only. A MoE architecture assigns each objective to a specialized expert head (momentum, growth, defensive, tax-aware), and a learned intent router blends experts based on the active objective and current market regime, which eliminates cross-objective gradient conflict. Phase 3 adds a lightweight personalization layer further adapted at inference time to each individual via a 76-parameter LoRA module fine-tuned on real brokerage transaction history, inferring investment objectives from revealed trading behavior rather than questionnaires. A natural language intent parser converts free-form goals directly into structured investment objective parameters.

13:00 JSTLLM/生成AI研究/論文

コンパイルを超えて: 忠実な自然言語からリーンステートメントへの形式化の評価

Theorem-proving benchmarks evaluate proof search against fixed formal statements, but natural-language-to-Lean formalization must generate the formal statement itself.この設定では、コンパイルは妥当性チェックのみです。リーン宣言では、仮説を省略したり、ドメインを変更したり、空の主張を表現したりしながら型チェックを行う可能性があります。私たちは、忠実なステートメントの形式化を評価問題とボトルネック帰属問題の両方として研究します。実際の解析、複雑な解析、トポロジー、代数にわたる 400 エントリの大学院レベルのベンチマークで、当社のプロトコルは、リーン コンパイル、クロスモデルの意味論的判断、および人間による専門家の校正を組み合わせています。結果として得られる状況は、コンパイル レートの評価とは異なります。完全にツールで強化されたエージェントは、コンパイル率 89.5% に達しますが、コンセンサス忠実度は 60.5% にとどまり、コンパイルはパスしますがコンセンサスに忠実ではない 29.0 ポイントのギャップが明らかになります。対象を絞った人間による監査は、このメトリクスを保守的な決定境界としてサポートします。利用可能なケースレベルの監査全体で、コンセンサス肯定的な出力の 96.0% が人間によって忠実であると確認されたのに対し、コンパイルパスのコンセンサス否定的な出力の 82.4% は人間によって確認された意味論的失敗です。この指標の下では、既存のワンショットフォーマライザーモデルと証明者指向のリーンモデルは低いままであり、形式的妥当性、証明指向のリーン能力、および忠実なステートメントの生成を個別に報告する必要があることを示唆しています。次に、完全な $2^3$ 要因計画を使用して、形式化パイプラインで繰り返される 3 つの介入、つまりパラメトリックな専門家による製図、Mathlib/コンテキスト検索、およびリーンな精緻化フィードバックを分解します。エレーションフィードバックは最大の妥当性介入ですが、コンパイルパスの意味論的失敗のより大きなバケットも明らかにします。検索は主にグラウンディングと選択性を改善します。微調整された製図は、フィードバックと基礎が得られれば、このツール スタックでほとんど代替可能です。

原文 (English)

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization

Theorem-proving benchmarks evaluate proof search against fixed formal statements, but natural-language-to-Lean formalization must generate the formal statement itself. In this setting, compilation is only a validity check: a Lean declaration may type-check while omitting hypotheses, changing domains, or expressing a vacuous claim. We study faithful statement formalization as both an evaluation problem and a bottleneck-attribution problem. On a 400-entry graduate-level benchmark spanning real analysis, complex analysis, topology, and algebra, our protocol combines Lean compilation, cross-model semantic judging, and human expert calibration. The resulting picture is different from compile-rate evaluation: a full tool-augmented agent reaches 89.5% compilation but only 60.5% consensus faithfulness, exposing a 29.0-point compile-pass but consensus-unfaithful gap. Targeted human audits support the metric as a conservative decision boundary: across available case-level audits, 96.0% of consensus-positive outputs are human-confirmed faithful, while 82.4% of compile-pass consensus-negative outputs are human-confirmed semantic failures. Under this metric, existing one-shot formalizer models and prover-oriented Lean models remain low, suggesting that formal validity, proof-oriented Lean competence, and faithful statement generation should be reported separately. We then use a full $2^3$ factorial design to decompose three recurring interventions in formalization pipelines: parametric expert drafting, Mathlib/context search, and Lean elaboration feedback. Elaboration feedback is the largest validity intervention, but it also exposes a larger compile-pass semantic-failure bucket; search mainly improves grounding and selectivity; and fine-tuned drafting is largely substitutable in this tool stack once feedback and grounding are available.

13:00 JSTエージェント

LabGuard: 自然言語のラボルールを、身体化されたラボエージェントのランタイムガードに根付かせる

科学的に身体化されたエージェントは、実験室での手順を実行できるようになってきていますが、動的な実験室環境でこれらの手順を安全に実行することは依然として困難です。現在の安全アプローチでは、安全ルール、マニュアル、プロトコル、標準操作手順などの実験室の自然言語を機械チェック可能な実行時制約に変換する中間ステップが見落とされていることがよくあります。 LabGuard (Laboratory Guard) は、自然言語のラボ ルールを実行可能な仕様に根付かせ、ランタイム ガードとして展開する、言語から実行までの安全性スイートです。 LabGuard には、3 つのコア コンポーネントが含まれています。LabGuard-IR は、型指定された実行可能表現を定義します。 LabGuard-Bench は、203 のシード ラボラトリ ルールから拡張された 812 の教師ありアノテーションを提供します。 LabGuard-Grounder は、自然言語のラボルールを LabGuard-IR にマッピングします。結果として得られる IR インスタンスは、LabGuard パイプラインによって処理され、ランタイム モニターにコンパイルされ、コントローラーの境界に適用されます。実験では、LabGuard が目に見えないラボルール ソースに一般化し、79.4 のタスク スコープ F1 を達成し、モニターのコンパイル後に危険なイベントを 39.5% から 23.8% に削減することが示されています。 LabUtopia では、ランタイム モニターが ACT と統合されており、タスクの成功を維持しながら介入を 0.5% 未満に抑えます。

原文 (English)

LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents

Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challenging. Current safety approaches often overlook the intermediate step of transforming laboratory natural language, including safety rules, manuals, protocols, and standard operating procedures, into machine-checkable runtime constraints. We introduce LabGuard (Laboratory Guard), a language-to-execution safety suite that grounds natural-language laboratory rules into executable specifications and deploys them as runtime guards. LabGuard includes three core components: LabGuard-IR, which defines a typed executable representation; LabGuard-Bench, which provides 812 supervised annotations expanded from 203 seed laboratory rules; and LabGuard-Grounder, which maps natural-language laboratory rules into LabGuard-IR. The resulting IR instances are handled by the LabGuard Pipeline, which compiles them into runtime monitors and applies them at the controller boundary. Experiments show that LabGuard generalizes to unseen laboratory-rule sources, achieves 79.4 task-scope F1, and reduces unsafe events from 39.5% to 23.8% after monitor compilation. In LabUtopia, its runtime monitors integrate with ACT, keeping interventions below 0.5% while preserving task success.

13:00 JSTLLM/生成AIエージェント研究/論文

OpenLife: 自律型 LLM エージェントによるオープンワールド人工生命に向けて

人工生命は、多くの計算基盤上で生命のような振る舞いを研究してきましたが、そのほとんどは研究者が設計した閉じた世界でのことでした。私たちは、永続的なメモリ、ツールの使用、ネットワーク アクセス、および支払いを備えた大規模言語モデル (LLM) エージェントによって、人工生命をオープンな社会、技術、経済の世界に移動させることが可能になったと主張します。これは、私たちがオープンワールド人工生命 (オープンワールド ALIFE) と呼ぶパラダイムです。私たちの概念実証である OpenLife は、ステートレス LLM を単一の「スマート エージェント」ではなく、非同期プロセスの社会 (記憶、認識、評価、および永続性を標準にする予算ベースの代謝) で囲みます。固定された目的が利用できないため、経験はスカラー報酬ではなくオープンボキャブラリーの LLM 判断によって評価され、記憶は頻度ではなく意味によって再配線されます。このようなエージェント 6 人をオープンワールドで約 12 週間実行し、その後も増え続けていますが、反応的な活動から自発的な活動への移行、異なるエージェントへの個性化、新たな社会構造、そして初めての自力で得た外部収入など、出現する生命のようなダイナミクスを報告します。私たちは、OpenLife が人工生命を実現したとは主張しませんが、オープンワールド ALIFE は現在実行可能な実験パラダイムであり、慎重に生きている AI と呼ばれる可能性のあるものを研究するための具体的なプラットフォームであると主張します。

原文 (English)

OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents

Artificial life has explored life-like behavior on many computational substrates, but mostly in researcher-designed closed worlds. We argue that large language model (LLM) agents, with persistent memory, tool use, network access, and payment, now make it possible to move artificial life into the open social, technical, and economic world, a paradigm we call open-world Artificial Life (open-world ALIFE). Our proof-of-concept, OpenLife, surrounds a stateless LLM not with a single "smart agent" but with a society of asynchronous processes: memory, perception, evaluation, and a budget-based metabolism that makes persistence normative. With no fixed objective available, experience is appraised by open-vocabulary LLM judgment rather than scalar reward, and memory is rewired by meaning rather than frequency. Running six such agents in the open world for about twelve weeks and counting, we report the life-like dynamics that emerge: a shift from reactive to spontaneous activity, individuation into distinct agents, emergent social structure, and a first self-earned external income. We do not claim OpenLife has realized artificial life, but that open-world ALIFE is now a viable experimental paradigm and a concrete platform for studying what might cautiously be called living AI.

13:00 JSTLLM/生成AIロボティクス研究/論文

MultiUAV-Plat: マルチ UAV 共同タスク計画のための LLM 指向のプラットフォーム、ベンチマーク、およびフレームワーク

大規模言語モデル (LLM) は、高レベルのロボット タスク計画に有望なインターフェイスを提供しますが、複数の UAV コラボレーションでの使用を体系的に評価することは依然として困難です。既存の UAV シミュレーターは主にダイナミクス、知覚、または低レベルの制御に重点を置いていますが、既存の LLM エージェント ベンチマークでは、部分的な観測可能性、空間カバー範囲、UAV の割り当て、複数車両の調整などの航空ロボットの制約をほとんど捉えていません。このギャップを埋めるために、マルチ UAV 共同タスク計画のための軽量で使いやすい LLM エージェント指向のシミュレーション プラットフォームである MultiUAV-Plat を紹介します。このプラットフォームは、簡潔な RESTful API、エージェント向けの観察、ロールベースの情報アクセス、非表示の検証ロジック、およびオプションの 2D/3D 視覚化を公開し、エージェントが特権的なシミュレーター アクセスではなく現実的なツールの対話を通じてミッションを解決できるようにします。このプラットフォーム上に構築された MultiUAV-Plat Benchmark には、75 のミッション セッション、1500 の自然言語タスク、およびターゲットの割り当て、エリア探索、エリアの割り当てとパトロールのシナリオにわたる 9396 の検証チェックが含まれています。さらに、マルチ UAV の動作をメモリ、観察、タスク理解、計画、実行、検証に構造化するタスク固有の LLM エージェント フレームワークである Agent4Drone を提案します。完全なペアのベンチマーク比較では、Agent4Drone は 57.9% のタスク合格率、74.6% の平均タスク チェック合格率、72.0% のグローバル チェック合格率を達成し、それぞれ ReAct ベースラインの 30.6%、47.9%、43.1% を大幅に上回っています。 Agent4Drone は、タスクの合計失敗率も 32.4% から 12.9% に削減します。これらの結果は、MultiUAV-Plat と MultiUAV-Plat Benchmark が、現実的な情報と実行の制約の下で LLM 駆動のマルチ UAV 自律性を研究するための再現可能な基盤を提供することを示しています。

原文 (English)

MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning

Large language models (LLMs) provide a promising interface for high-level robotic task planning, but their use in multi-UAV collaboration remains difficult to evaluate systematically. Existing UAV simulators mainly emphasize dynamics, perception, or low-level control, while existing LLM-agent benchmarks rarely capture aerial-robotics constraints such as partial observability, spatial coverage, UAV assignment, and multi-vehicle coordination. To bridge this gap, we present MultiUAV-Plat, a lightweight, easy-to-use, LLM-agent-oriented simulation platform for multi-UAV collaborative task planning. The platform exposes concise RESTful APIs, agent-facing observations, role-based information access, hidden validation logic, and optional 2D/3D visualization, allowing agents to solve missions through realistic tool interaction rather than privileged simulator access. Built on this platform, the MultiUAV-Plat Benchmark contains 75 mission sessions, 1500 natural-language tasks, and 9396 validation checks across target assignment, area search, and area assignment and patrol scenarios. We further propose Agent4Drone, a task-specific LLM agent framework that structures multi-UAV behavior into memory, observation, task understanding, planning, execution, and verification. In a full paired benchmark comparison, Agent4Drone achieves a 57.9% task pass rate, a 74.6% average task check pass rate, and a 72.0% global check pass rate, substantially outperforming a ReAct baseline at 30.6%, 47.9%, and 43.1%, respectively. Agent4Drone also reduces the total failed task rate from 32.4% to 12.9%. These results demonstrate that MultiUAV-Plat and MultiUAV-Plat Benchmark provide a reproducible foundation for studying LLM-driven multi-UAV autonomy under realistic information and execution constraints.

13:00 JSTエージェント

DDIAgents: 薬物間相互作用予測のためのメカニズム条件付きコンテキスト フロー

薬物間相互作用 (DDI) の予測は医薬品の安全性にとって不可欠ですが、相互作用メカニズムによって関連性が変化する異種の生物医学的証拠を推論する必要があります。私たちは、動的な知識オーケストレーションを通じて DDI 予測を実行する、メカニズム条件付きマルチエージェント フレームワークである DDIAgents を提案します。薬物ペアが与えられると、プランナー エージェントは専門のエキスパート エージェントをインスタンス化し、メカニズムに関連した知識ソースを各エージェントにルーティングし、結論エージェントを通じて分析を集約します。 DDIAgents は、推定された対話メカニズムにコンテキスト フローを適応させることで、無関係な情報を削減し、専門家の補完的な推論をサポートし、解釈可能なエージェント レベルの理論的根拠を生成します。現実的な DDI 予測ベンチマークに関する広範な実験により、DDIAgents が既存の機能ベース、グラフベース、LLM ベース、およびエージェントベースのベースラインを一貫して上回るパフォーマンスを示していることが示されています。 DDIAgents は、予測パフォーマンスを超えて、適応的で解釈可能な AI4Science 推論のために、マルチエージェント システムが異種の科学知識をどのように整理できるかを示します。

原文 (English)

DDIAgents: Mechanism-Conditioned Context Flow for Drug-Drug Interaction Prediction

Drug-drug interaction (DDI) prediction is essential for medication safety, yet it requires reasoning over heterogeneous biomedical evidence whose relevance changes across interaction mechanisms. We propose DDIAgents, a mechanism-conditioned multi-agent framework that performs DDI prediction through dynamic knowledge orchestration. Given a drug pair, a planner agent instantiates specialized expert agents, routes mechanism-relevant knowledge sources to each agent, and aggregates their analyses through a conclusion agent. By adapting context flow to the inferred interaction mechanism, DDIAgents reduces irrelevant information, supports complementary expert reasoning, and produces interpretable agent-level rationales. Extensive experiments on realistic DDI prediction benchmarks show that DDIAgents consistently outperforms existing feature-based, graph-based, LLM-based, and agent-based baselines. Beyond prediction performance, DDIAgents demonstrates how multi-agent systems can organize heterogeneous scientific knowledge for adaptive and interpretable AI4Science reasoning.

13:00 JST研究/論文

トランスを介した UTM の安全性が重要なシナリオを明らかにする

無人交通管理 (UTM) システムは、複数の航空機をリモートで管理および調整するように設計されたクラウドベースのプラットフォームです。 UTM システムは安全性が重要であるため、衝突や衝突などの障害を許容できません。潜在的な脆弱性を明らかにするためには、障害を明らかにするための最適なデモンストレーションも明確な報酬シグナルもありません。さらに、UTM の自己修復機能により、重大な障害の「ロングテール効果」が発生します。私たちは、UTM 脆弱性の発見を、トランスフォーマーベースの RL アーキテクチャに適したシーケンス モデリング問題として枠組み化することを提案します。私たちのアプローチは、注意メカニズムを活用してシステム状態間の関係を直接モデル化し、最適なアクションを予測します。私たちのフレームワークには、対象を絞ったテスト シナリオを生成するポリシー モデルと、ドメインの制約を強制するアクション サンプラーが導入されています。探索をガイドするために、リスクベースの報酬関数を使用します。 700 時間のシミュレーション調査による広範な評価を通じて、専門家の指導によるテストと比較して、脆弱性発見効率が 8$\times$ 向上することを実証しました。また、従来の方法では見逃していた重大なエッジケースも発見します。

原文 (English)

Revealing Safety-Critical Scenarios for UTM via Transformer

Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely. UTM systems are safety-critical which cannot tolerate failures like crash or collision. To reveal latent vulnerabilities, there are neither optimal failure-exposing demonstrations nor clear reward signals. Additionally, UTM's self-healing capability introduces the ``long-tail effect'' of critical failures. We propose framing UTM vulnerability discovery as a sequence modeling problem amenable to transformer-based RL architectures. Our approach leverages attention mechanisms to directly model the relationship among system states, and predict optimal actions. Our framework introduces a Policy Model that generates targeted test scenarios and an Action Sampler that enforces domain constraints. We use a risk-based reward function to guide exploration. Through extensive evaluation on a 700-hour simulation study, we demonstrate an 8$\times$ improvement in vulnerability discovery efficiency compared to expert-guided testing. It also discovers critical edge cases that traditional methods have missed.

13:00 JSTLLM/生成AIエージェント

過去は序章: 逐次進化する LLM メモリの選択的更新のためのプラグイン コントローラー

順次進化する LLM メモリにより、エージェントは過去の経験を再利用できますが、既存のシステムは通常、ローカルで生成された各メモリ更新を、将来の動作が改善されるかどうかを確認せずに展開します。その結果、現在のタスクに役立つ更新によって、有用な知識が上書きされたり、過度に具体的なルールが導入されたり、最終的な記憶が最近の例に偏ったりする可能性があります。我々は、メモリ更新候補を受け入れるか、以前のメモリを保持するかを決定するプラグイン メモリ コントローラである Janus を提案します。この決定を効率的に行うために、Janus はメモリ モメンタム トリガーを使用してメモリ更新軌跡の疑わしい逸脱を特定し、完全な履歴を再生するのではなく、カバレッジ、境界、および新しいタスクのコンパクトなハイブリッド評価セットで古いメモリと新しいメモリを比較します。 Janus はメソッドに依存せず、更新ルールを変更せずに既存のアップデータをラップします。 6 つのデータセット、2 つのバックボーン LLM、および 2 つのメモリ アップデーターにわたって、Janus は対応するベース アップデーターと比較して平均精度を +2.7 ~ +4.6 ポイント向上させます。

原文 (English)

The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory

Sequentially evolving LLM memory enables agents to reuse past experience, but existing systems usually deploy each locally generated memory update without checking whether it improves future behavior. As a result, updates that help the current task may overwrite useful knowledge, introduce over-specific rules, or bias the final memory toward recent examples. We propose Janus, a plug-in memory controller that decides whether to accept a candidate memory update or retain the previous memory. To make this decision efficient, Janus uses a Memory Momentum Trigger to identify suspicious deviations in the memory-update trajectory, and compares old and new memories on a compact hybrid evaluation set of coverage, boundary, and fresh tasks instead of replaying the full history. Janus is method-agnostic and wraps existing updaters without changing their update rules. Across six datasets, two backbone LLMs, and two memory updaters, Janus improves average accuracy by +2.7 to +4.6 points over the corresponding base updaters.

13:00 JSTエージェントロボティクス

現実世界の故障記録を利用した自動運転システム試験のシナリオ生成

道路上での安全な動作を確保するには、自動運転システム (ADS) の導入前テストと障害発見が重要です。現在のシミュレーションベースのテスト方法は、固定されたシナリオ表現を前提として、最適なシナリオを効率的に探索するための数学的モデルに主に焦点を当てています。一方、実際のテストでは、テスト用のシナリオ テンプレートを設計するためにかなりの手作業が必要になります。これらのテンプレートは、展開前の車両の動き、マップ タイプなどで構成される個別の故障シナリオを表します。ADS の過去の故障記録は、現実世界の故障状況の信頼できる情報源であり、シナリオ生成に使用できます。この研究では、自然言語形式の履歴記録から入手可能なカテゴリ情報とコンテキスト情報を使用したシナリオ生成パイプラインを提案します。私たちのアプローチは、特定のシステムのテスト制約と互換性のあるモジュール式 LLM ベースの合成シナリオ生成で構成されています。私たちは、NHTSA ADS クラッシュ レコードを使用して、Metadrive シミュレーターで自律ナビゲーションをテストするためのさまざまなシナリオの生成にこの方法を適用することに成功しました。私たちのアプローチにより、4 つの道路タイプと 3 つの非自我車両移動タイプ (作業ゾーンの形での道路異常を含む) を組み合わせて、正確かつ多様なシナリオが生成されます。生成されたシナリオは、提供されたテスト条件と一致しており、20 シナリオという限られたテスト予算内でシステムの興味深い障害を明らかにします。コードは https://github.com/anjaliParashar/crash2scenario で入手できます。

原文 (English)

Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records

To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day simulation based testing methods focus largely on mathematical models for efficient search of optimal scenarios, assuming a fixed scenario representation. On the other hand, real-world testing involves substantial manual effort to design scenario templates for testing. These templates represent distinct failure scenarios consisting of pre-deployment vehicle movements, map types, etc. Historical failure records for ADS are a reliable source of real-world failure conditions, which can be used for scenario generation. In this work, we propose a scenario generation pipeline using categorical and contextual information available from historical records in natural language format. Our approach consists of modular LLM based synthetic scenario generation, compatible with the testing constraints of a given system. We successfully apply our method to generate a diverse set of scenarios for testing autonomous navigation on Metadrive simulator using the NHTSA ADS crash records. Our approach results in accurate and diverse scenario generation with a combination of 4 road types, 3 non ego vehicle movement types, including on road anomalies in the form of working zones. Generated scenarios align with the provided testing conditions, and reveals interesting failures of the system within a limited testing budget of 20 scenarios. Code is available at https://github.com/anjaliParashar/crash2scenario.

13:00 JSTLLM/生成AIエージェント研究/論文

図書館を超えて: 研究数学を自動形式化するためのエージェント フレームワーク

大規模言語モデル (LLM) は、数学的推論において優れた能力を実証していますが、人間の検出を回避する微妙なエラーを頻繁に生成します。 Lean 4 のような正式な数学言語は機械的な証明チェックを提供し、自動形式化、つまり自然言語数学を検証可能なコードに自動的に変換する必要性を強く促します。最近の傾向は、標準プログラミング向けに大幅に最適化された汎用 LLM が、リーン向けに明示的に微調整された小規模なモデルよりも優れたパフォーマンスを発揮することを示しています。この変化を活用して、一般的なコーディング LLM を利用したエージェント自動形式化フレームワークを導入します。私たちのシステムの中核には、研究レベルの数学に合わせて調整されたマルチエージェント パイプラインを管理するオーケストレーターがあります。最先端の研究は Mathlib などの既存のライブラリの範囲外の概念に依存することが多いため、私たちのシステムは必要な型定義を動的に拡張し、主要な定理を形式化する前に新しい補助補題技術を介してそれらを検証します。私たちはこのアプローチを PutnamBench に適用し、32 個の問題のランダムなサンプルに対して機械チェックされたリーン証明を作成しました。さらに、ACM Symposium on Theory of Computing (STOC) の組み合わせ論、通信の複雑さ、機構設計、学習理論にわたる 5 つの論文に基づいてシステムを評価し、主定理の形式化に成功し、生成された形式化を専門家とともに検証しました。 5 つすべてについて、ステートメントと並行して証明も形式化します。特に、そのうち 2 つはリーンのカーネルを超える公理を使用せずに証明されています。すべての形式化は https://beyondthelibrary.github.io/formal_arxiv で入手できます。

原文 (English)

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evade human detection. Formal mathematical languages like Lean 4 offer mechanical proof checking, strongly motivating the need for autoformalization: the automatic translation of natural language mathematics into verifiable code. Recent trends indicate that general-purpose LLMs, heavily optimized for standard programming, now outperform smaller models explicitly fine-tuned for Lean. Leveraging this shift, we introduce an agentic autoformalization framework powered by general coding LLMs. At the core of our system is an orchestrator that manages a multi-agent pipeline tailored for research-level mathematics. Because cutting-edge research frequently relies on concepts outside the scope of existing libraries like Mathlib, our system dynamically extends necessary type definitions and validates them via a novel Auxiliary Lemma technique before formalizing the primary theorems. We applied our approach to PutnamBench, producing machine-checked Lean proofs for a random sample of 32 problems. Furthermore, we evaluate our system on five papers from the ACM Symposium on Theory of Computing (STOC) spanning combinatorics, communication complexity, mechanism design, and learning theory, successfully formalizing their main theorems and validating the generated formalizations with human experts; for all five we also formalize the proofs alongside the statements, and notably two of them are proved with no axioms beyond Lean's kernel. All of our formalizations are available at https://beyondthelibrary.github.io/formal_arxiv .

13:00 JST研究/論文

ナレッジグラフインジェクションによる表形式の医療データのクロスドメイン機能拡張

包括的なクロスドメイン生物医学プロファイルの取得には費用と時間がかかることが多く、その結果、医学研究における深刻なデータ不足が生じます。この課題に対処するために、私たちは、表形式の医療データにおけるクロスドメイン機能拡張のために特別に設計された知識注入フレームワークである MedKGTab を提案します。 MedKGTab は、固有の統計的依存関係と確立された医学的相関関係を利用することにより、収集されていない生物医学的特徴を入手可能な特徴から推測しようとします。 MedKGTab は、行と列のデュアルアテンション メカニズムを採用することにより、生の構造化表データを直接操作し、トークン化によって引き起こされる構造損失を生じることなく、本質的に正確な数値分布を捕捉します。重要なのは、MedKGTab がデータ主導の統計的事前分布を SPOKE 生物医学知識グラフと統合し、データと知識チャネルの間で最適な相乗効果を実現することです。この相乗効果の中で、データ チャネルから得られる表現は、注入された生物医学的知識によって調整され、最終的に生成されるデータが実証的な医学研究に基づいていることが保証されます。実験結果は、MedKGTab がクロスドメイン特徴拡張において高いデータ忠実性と現実的なデータ表現を達成することを示しています。これは、SOTA 医療大規模モデル (Baichuan M3-plus など) と医療データ生成用に設計された特殊な表形式モデルの両方を上回ります。さらに、MedKGTab は、同じデータセット内で欠落している特徴を推論する場合でも、異なる医療コホート間で一般化する場合でも、さまざまなデータ生成シナリオにわたって一貫して優れたパフォーマンスを提供します。

原文 (English)

Cross-Domain Feature Expansion for Tabular Medical Data via Knowledge Graphs Injection

Acquiring comprehensive cross-domain biomedical profiles is often costly and time-consuming, resulting in severe data scarcity in medical research. To address this challenge, we propose MedKGTab, a knowledge-injected framework specifically engineered for cross-domain feature expansion in tabular medical data. MedKGTab seeks to infer uncollected biomedical features from available ones by exploiting their inherent statistical dependencies and established medical correlations. By employing a row-column dual-attention mechanism, MedKGTab operates directly on raw structured tabular data, inherently capturing exact numerical distributions without the structural loss caused by tokenization. Crucially, MedKGTab integrates data-driven statistical priors with the SPOKE biomedical knowledge graph, achieving an optimal synergy between the data and knowledge channels. Within this synergy, the representations derived from the data channel are modulated by the injected biomedical knowledge, ensuring the final generated data are grounded in empirical medical research. Experimental results demonstrate that MedKGTab achieves high data fidelity and realistic data representation in cross-domain feature expansion. It outperforms both SOTA medical large models (e.g., Baichuan M3-plus) and specialized tabular models designed for medical data generation. Furthermore, MedKGTab consistently delivers superior performance across various data generation scenarios, whether inferring missing features within the same dataset or generalizing across different medical cohorts.

13:00 JSTLLM/生成AIエージェント研究/論文

ClawArena チーム: 言語モデル エージェントのサブエージェント オーケストレーションと動的ワークフローのベンチマーク

運用環境の大規模言語モデル (LLM) エージェントは、単独の問題解決者としてではなく、マネージャーとして導入されることが増えています。メイン モデルは、特化したサブエージェントを作成し、作業を委任し、動的なワークフローを通じて並列非同期のリターンを調整します。 1 つのモデルがそのようなチームを実際に実行できるかどうかは、ほとんど測定されていません。既存のベンチマークは、ポリシー自体のタスク解決や固定マルチエージェント システムの緊急動作をスコア化しますが、リーダーとして機能する単一の LLM の管理能力を分離するものはありません。この管理能力を測定する、258 の評価ラウンドと 72 の段階的な更新にわたる 41 のマルチターン、マルチモーダル、マルチディレクトリ シナリオのベンチマークである ClawArena-Team を紹介します。メイン エージェントは意図的に制限されています。ネイティブにテキストのみを認識し、ワークスペースの一部のみに直接アクセスします。ローカルで提供される固定のサブエージェント プールを指揮するため、スコアの差は、実際の能力ではなく管理スキルを反映します。すべてのスコアリングは実行ベースであり、LLM 判定はありません。全体的なスコアであるサブエージェント管理スコア (SMS) には、タスクの正確さに最小権限とモダリティ ルーティング係数が乗じられます。 12 の独自モデル、コミュニティホスト型モデル、セルフホスト型モデルにわたる実験では、管理のボトルネックは認識ではなく権限付与であることが示されています (ワークスペース権限の精度が 50% を超えるモデルはありません)。コストと管理品質は分離されています (API コストは 100 倍を超えますが、全体のスコアは 4 倍未満であり、パレート フロンティアで最も安価なオープン モデルです)。そして、ほとんどのリーダーボード スコアは 9.9 ポイントの範囲内に集中していますが、オーケストレーションの動作は 1 桁以上異なります。コードとデータは公開されます。

原文 (English)

ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows. Whether one model can actually run such a team is largely unmeasured: existing benchmarks score a policy's own task-solving or a fixed multi-agent system's emergent behavior, but none isolate the management ability of the single LLM acting as leader. We introduce ClawArena-Team, a benchmark of 41 multi-turn, multimodal, multi-directory scenarios spanning 258 evaluation rounds and 72 staged updates that measures this management ability. The main agent is deliberately constrained: it natively perceives only text and directly accesses only part of the workspace. It commands a fixed, locally served subagent pool, so score differences reflect management skill, not raw capability. All scoring is execution-based with no LLM judge: an overall score -- the Subagent-Management Score (SMS) -- multiplies task correctness by a least-privilege and modality-routing factor. Across twelve proprietary, community-hosted, and self-hosted models, experiments show that the management bottleneck is privilege granting rather than perception (no model exceeds 50% workspace-permission precision); that cost and management quality are decoupled (API cost spans over 100 times while the overall score spans under 4 times, with the cheapest open models on the Pareto frontier); and that most leaderboard scores cluster within a 9.9-point band while orchestration behaviors diverge by more than an order of magnitude. Code and data will be released.

13:00 JSTLLM/生成AI画像/動画生成エージェントビジネス/資金調達研究/論文ClaudeGPT / ChatGPTMicrosoft

HealthAgentBench: 挑戦的なフロンティア AI エージェント向けの現実的なエージェント ヘルスケア環境の統合ベンチマーク スイート

AI エージェントがますます複雑で長期的な推論を行えるようになっているため、現実世界の医療アプリケーションへの進歩を測定するには、厳密かつ総合的な評価が不可欠です。 HealthAgentBench は、それぞれ独自の環境を持つ 7 つのカテゴリにわたる 54 のエージェント ヘルスケア タスクのスイートです。ベンチマーク スイートは、患者の治療過程全体にわたる多様なワークフローと幅広いモダリティに及びます。各タスクは、エンドツーエンドの臨床ワークフローを複製するように設計されています。最小限の指示が与えられると、エージェントは生の医療データを探索し、複雑な環境内で操作し、単純なプロンプトを超えた複数ステップのソリューションを実行する必要があります。最終的なタスクの成功率は、各エージェントの HealthAgentBench の全体的なパフォーマンスに関する単一の解釈可能な指標を提供するために報告されます。 HealthAgentBench でフロンティア エージェントを評価すると、全体的なタスクの成功率が依然として低いことがわかり、スイートの難しさを浮き彫りにしています。最も強力で費用対効果の高いエージェントである Codex GPT-5.5 の成功率はわずか約 42% です。 HealthAgentBench は、総合的なパフォーマンスを超えて、タスク カテゴリ全体の微妙な長所と短所を明らかにします。フロンティア エージェントは、EHR データを使用した研究モデリング パイプラインの自動開発に期待を示していますが、医療画像処理は、特にクロード コード モデルの場合、依然として課題が多く、一方で Codex GPT-5.5 は新たな機能を示しています。大規模な検索スペースと構成的推論の要件を組み合わせるタスクは、現在のすべてのエージェントにとって依然として困難です。これらの結果を総合すると、HealthAgentBench が将来の進歩の余地が十分にある、挑戦的で現実的なベンチマークを提供していることを示唆しています。ベンチマークは https://github.com/microsoft/HealthAgentBench でリリースされています。

原文 (English)

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT-5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT-5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at https://github.com/microsoft/HealthAgentBench.

13:00 JSTLLM/生成AIエージェント

AI 支援によるデュアル エージェントによる凸緩和の発見

最近の研究では、LLM エージェントが上限をもたらす極値構造を探索することにより、急峻な不平等を改善できることが示されています。私たちは補完的な側面に取り組みます。下限はすべての許容関数に当てはまり、非凸問題の凸緩和から導き出され、より厳しい緩和によりより強い境界が与えられます。このような緩和を発見するために自動調査パラダイムをインスタンス化します。コーディング エージェントが有効な締め付け制約を提案し、理論エージェントがそれぞれを検証して反例を検索し、報告されたすべての限界が厳密な区間演算でチェックされた明示的な二重実現可能点によって証明されます。 \citet{tao2025alphaevolve} によって研究された 2 つの最適化定数、最初の自己相関不等式 ($C_{6.2}$) と Erd\H{o} の最小重複定数 ($C_{6.5}$) について、認定された下限を $1.28$ から $1.2937$ に、そして $0.379005$ からそれぞれ0.37912ドル。

原文 (English)

AI-Assisted Discovery of Convex Relaxations via Dual Agents

Recent work shows that LLM agents can improve sharp-constant inequalities by searching for extremal constructions, which yield upper bounds. We address the complementary side: a lower bound holds for every admissible function and follows from a convex relaxation of the nonconvex problem, with tighter relaxations giving stronger bounds. We instantiate the autoresearch paradigm to discover such relaxations: a coding agent proposes valid tightening constraints, a theory agent verifies each one and searches for counterexamples, and every reported bound is certified by an explicit dual-feasible point checked in rigorous interval arithmetic. On two optimization constants studied by \citet{tao2025alphaevolve} - the first autocorrelation inequality ($C_{6.2}$) and the Erd\H{o}s minimum-overlap constant ($C_{6.5}$) - we improve the certified lower bounds from $1.28$ to $1.2937$ and from $0.379005$ to $0.37912$, respectively.

13:00 JSTエージェントロボティクス

Agentic RAG-VLM: ロボットによる把握のための自己反映計画を備えたアフォーダンスを意識した検索拡張生成

雑然とした環境でロボットによる汎用的な把握は、マニピュレータを構造化されていない人間の空間に配置するために不可欠ですが、既存の VLM ベースの手法は、オブジェクトの照合に視覚的な類似性に依存し、ハンドルの握りやすさや材料の脆弱性などの物理的なアフォーダンスを無視し、空間推論や障害回復なしで開ループで動作するため、オブジェクトが密集している場合や物理的に多様な場合には有効性が制限されます。我々は、検索拡張生成 (RAG) を視覚言語モデル (VLM) およびエージェント的内省計画と統合することにより、VLM ベースの意味理解と物理的根拠に基づいた把握の実行を橋渡しする統合フレームワークである Agentic RAG-VLM を紹介します。 Agentic RAG-VLM は、密結合された 3 つのコンポーネントを導入します。(1) タイプ、素材、脆弱性、把握可能領域を含む 4 次元アフォーダンス記述子をエンコードし、見た目ではなく機能的なアフォーダンス互換性によって戦略を取得する階層型アフォーダンス認識 RAG (HAA-RAG)。 (2)VLM知覚から空間関係グラフを構築し、近接性、オクルージョン、およびサポート制約を具体的な把握パラメータ調整に変換するシーングラフ制約推論器。 (3) 14 タイプの障害分類と閉ループ把握改良のための 3 レベルの適応再試行を備えたエージェント的自己反射パイプライン。構成ごとに 360 回の試行を行う、単一把握、インタラクティブ、長期シナリオにわたる 12 タスクのベンチマークで評価したところ、Agentic RAG-VLM は全体で 78.3 パーセントの成功率を達成し、VLM のみのベースラインと比較して 53.3 パーセント ポイントの絶対的な向上を達成しました。これは、堅牢な操作にはアフォーダンスを意識した取得、シーン グラフ推論、およびエージェントによる回復が共に不可欠であることを示しています。

原文 (English)

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similarity for object matching, neglecting physical affordances such as handle graspability and material fragility, and operate open-loop without spatial reasoning or failure recovery, limiting their effectiveness when objects are densely packed or physically diverse. We present Agentic RAG-VLM, a unified framework that bridges VLM-based semantic understanding and physically grounded grasp execution by integrating retrieval-augmented generation (RAG) with vision-language models (VLMs) and agentic self-reflective planning. Agentic RAG-VLM introduces three tightly coupled components: (1) a Hierarchical Affordance-Aware RAG (HAA-RAG) that encodes four-dimensional affordance descriptors, including type, material, fragility, and graspable region, and retrieves strategies by functional affordance compatibility rather than visual appearance; (2) a Scene Graph Constraint Reasoner that constructs spatial relationship graphs from VLM perception and translates proximity, occlusion, and support constraints into concrete grasp parameter adjustments; and (3) an Agentic Self-Reflective Pipeline with a 14-type failure taxonomy and three-level adaptive retry for closed-loop grasp refinement. Evaluated on a 12-task benchmark spanning single-grasp, interactive, and long-horizon scenarios with 360 trials per configuration, Agentic RAG-VLM achieves 78.3 percent overall success, a 53.3 percentage-point absolute gain over VLM-only baselines, demonstrating that affordance-aware retrieval, scene graph reasoning, and agentic recovery are jointly essential for robust manipulation.

13:00 JST研究/論文

包括的なモビリティのモデリングに向けて: 都市システムにおける高齢者の軌跡パターンの特徴付けと評価

スマート シティの急速な進歩は軌跡データ マイニングにますます依存していますが、過小評価されている人口統計グループ、特に高齢者は公共モビリティ データセットにまばらに表されていることがよくあります。この過小評価により、モビリティのモデリングや下流の都市計画に体系的な偏りが生じる可能性があります。この研究では、シティ バイク システム データの 2016 年から 2020 年のジャージー シティのサブセットを使用して、過小評価されているサブグループのモビリティ シグネチャの欠如がモビリティ モデリングにどのような影響を与えるかを、ケーススタディとして合成軌道生成を使用して定量的に調査します。分析の結果、高齢ライダーは、局所的な活動スペース(958 m 対若いライダーは 1,189 m)、低い移動エントロピー(1.82 対 4.15)、非対称のオフピーク時間パターンなど、構造的に異なる移動特性を示していることが明らかになりました。大多数が支配するトレーニング データに依存すると、偏った合成結果が得られることを実証するために、一次マルコフ連鎖と、全人口、若いライダーのみ、高齢ライダーのみの 3 つの人口統計トレーニング設定にわたって QLoRA で微調整された Qwen3-4B モデルの両方をさらに評価します。結果は、多数派が優勢な集団で訓練されたモデルが、特に空間移動指標に関して、高齢者の移動行動を体系的に誤って表現していることを示しています。全人口を対象にトレーニングされたマルコフ モデルは、高齢者の歩幅を 4.5%、滞在時間を 8.9% 過大評価しますが、高齢者に特化したモデルは、ほとんどの指標で誤差が大幅に低くなります。マルコフ ベースのフレームワークと LLM ベースのフレームワークを比較すると、限られた人口統計データの下では、より高い能力のモデルが必ずしもサブグループ レベルの忠実度を向上させるわけではないことがわかります。これらの発見は、モビリティモデリングにおける人口統計的表現と、過小評価されている人口に対するその下流アプリケーションの重要性を強調しています。

原文 (English)

Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems

The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the elderly, are often sparsely represented in public mobility datasets. This underrepresentation can introduce systematic bias into mobility modeling and downstream urban planning. Using the 2016-2020 Jersey City subset of the Citi Bike System Data, this study quantitatively examines how the absence of underrepresented subgroups' mobility signatures affects mobility modeling, using synthetic trajectory generation as a case study. The analysis reveals that elderly riders exhibit a structurally distinct mobility signature, including localized activity spaces (958 m vs. 1,189 m for young riders), lower mobility entropy (1.82 vs. 4.15), and asymmetric off-peak temporal patterns. To demonstrate that relying on majority-dominated training data yields biased synthetic outcomes, we further evaluate both a first-order Markov chain and a Qwen3-4B model fine-tuned with QLoRA across three demographic training settings: the full population, young riders only, and elderly riders only. Results show that models trained on majority-dominated populations systematically misrepresent elderly mobility behavior, particularly for spatial mobility metrics. The Markov model trained on the full population overestimates elderly step length by 4.5% and dwell time by 8.9%, whereas the elderly-specific model achieves substantially lower errors across most metrics. Comparisons between the Markov and LLM-based frameworks further show that higher-capability models do not necessarily improve subgroup-level fidelity under limited demographic data. These findings underscore the importance of demographic representation in mobility modeling and its downstream applications for underrepresented populations.

13:00 JSTエージェントロボティクス

構造化自己回帰モデリングによる長期交通シミュレーション

インタラクティブな交通シミュレーションは、自動運転にとって重要な世界モデルです。長期シミュレーションにおける中心的な課題は、持続的なマルチエージェントの相互作用をモデル化することであり、エージェントが継続的にシーンに出入りするための動的なトークンの数によってさらに悪化します。この研究では、大規模言語モデル (LLM) などの大規模シーケンス モデルのアーキテクチャ上の帰納的バイアスと統計的事前分布との間の相乗効果に解決策があることを提案します。私たちの精査実験により、注意の伝達メカニズムとモーション トークンと自然言語間の分布の一貫性により、小規模で高度に凍結された LLM がトラフィック モデリングに迅速に適応できることが明らかになりました。この洞察に基づいて、シーン トポロジ、エージェントの状態、生成インテントを可変長の構造化自己回帰ストリームに投影する統合フレームワークである RosettaSim を導入し、強力な短期精度と安定した長期シミュレーション忠実度の両方を実現します。さらに、エージェント 1 対 1 の対応は時間の経過とともに必然的に薄れるため、拡張ロールアウトの評価にはさらに別のハードルが存在します。これに対処するために、意味的に類似した現実世界のシナリオをコンテキスト認識型参照アンカーとして取得する、取得ベースのトラフィック評価 (RTE) を導入します。 Waymo Open Sim Agent Challenge (WOSAC) の実験では、RosettaSim が短期および長期シミュレーションの両方で最先端のパフォーマンスを達成することが実証されています。さらに、RTE は既存のアプローチ ($r=0.74$) よりも標準メトリクス ($r=0.83$) と強い相関関係を示し、長期シミュレーションの忠実性との整合性が向上していることを示しています。

原文 (English)

Long-term Traffic Simulation via Structured Autoregressive Modeling

Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling sustained multi-agent interactions, which is further exacerbated by dynamic token cardinality as agents continuously enter and exit the scene. In this work, we propose that the solution lies in the synergy between the architectural inductive biases and statistical priors of large-scale sequence models, e.g., Large Language Models (LLMs). Our probing experiments reveal that the transferability of attention mechanisms and the distributional consistency between motion tokens and natural language enable small-scale, heavily frozen LLMs to rapidly adapt to traffic modeling. Building on this insight, we introduce RosettaSim, a unified framework that projects scene topology, agent states, and spawning intents into a structured autoregressive stream with variable length, achieving both strong short-term accuracy and stable long-horizon simulation fidelity. Furthermore, evaluating extended rollouts presents yet another hurdle, as one-to-one agent correspondence inevitably fades over time. To address this, we introduce Retrieval-based Traffic Evaluation (RTE), which retrieves semantically similar real-world scenarios as context-aware reference anchors. Experiments on the Waymo Open Sim Agent Challenge (WOSAC) demonstrate that RosettaSim achieves state-of-the-art performance in both short- and long-term simulation. Furthermore, RTE exhibits a stronger correlation with standard metrics ($r=0.83$) than existing approaches ($r=0.74$), indicating improved alignment with long-horizon simulation fidelity.

13:00 JST研究/論文

取得する前に考える: 戦略的計画と自己批判による堅牢なゼロショット合成画像取得

合成画像の検索では、参照画像をテキスト変更命令と統合することによって、ギャラリーからターゲット画像を識別する必要があります。トレーニング不要のゼロショット設定では、このタスクは、推論時の凍結された視覚、つまり言語埋め込み空間内で検索指向のテキスト クエリを構築することに依存します。既存のアプローチは主に、参照コンテキストと変更テキストを統合された記述に融合するシングルパス生成戦略に依存しています。この戦略では、生成中に意味上の歪みや省略を検出または修正することが困難になります。その結果、参照属性の保存とテキスト要件の統合が相互に干渉し、検索精度が低下します。これらの課題に対処するために、クエリ構築を多段階の推論パイプラインとして構造化するトレーニング不要のフレームワークである PEC-CIR を導入します。このフレームワークは、Planner-Executor-Critic アーキテクチャを通じて動作します。Planner は明示的な制約を抽出し、Executor は複数のターゲット記述候補を生成し、Critic は制約の遵守に基づいてこれらの候補を評価します。 PEC-CIR は、クエリ構築をシングルパス出力ではなく段階的推論プロセスとして再構成することで、取得前に候補クエリを明示的に評価することで生成エラーの伝播を削減し、それにより取得の安定性を向上させます。

原文 (English)

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism

Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction. In a training-free zero-shot setting, this task relies on constructing a retrieval-oriented textual query within a frozen vision--language embedding space at inference time. Existing approaches predominantly rely on a single-pass generation strategy that fuses the reference context and modification text into a unified description. This strategy makes it difficult to detect or correct semantic distortions and omissions during generation. Consequently, the preservation of reference attributes and the integration of textual requirements interfere with each other, which degrades retrieval precision. To address these challenges, we introduce PEC-CIR, a training-free framework that structures query construction as a multi-stage reasoning pipeline. The framework operates through a Planner--Executor--Critic architecture where the Planner extracts explicit constraints, the Executor generates multiple candidate target descriptions, and the Critic evaluates these candidates based on constraint compliance. By reframing query construction as a staged inference process instead of a single-pass output, PEC-CIR reduces the propagation of generative errors by explicitly evaluating candidate queries before retrieval, thereby improving retrieval stability.

13:00 JSTLLM/生成AIエージェント

エージェントのアイデア: 科学的アイデアのエージェントのための効率的なエージェントの軌道合成のサンプル

アイデアは科学的発見において極めて重要な役割を果たします。最近の LLM、特に AI Scientist システムは、自動化されたアイデア作成の有望な可能性を示しています。ただし、既存のアプローチは主に、事前定義されたエージェント ワークフローに依存しています。この制約により、科学文献の広大な検索空間と研究推論の複雑なアクション空間をナビゲートするために必要な柔軟性が大幅に制限されます。最近、Agentic LLM のトレーニングが有望な方向性として浮上し、柔軟な推論フレームワークと自律的なツール利用の機能を提供します。ただし、簡単ではない課題が残っています。つまり、以前のエージェントによるデータ合成手法を科学的アイデアに適用すると、データ合成コストが法外に高くなるという問題があります。このギャップを埋めるために、私たちは Agentic-Ideation を提案します。これは、自動軌道合成パイプラインと科学的アイデアのために訓練された特殊なエージェント LLM で構成される新しいフレームワークです。具体的には、まず 3 つの外部ツールと 3 つのコグニティブ ツールを組み込んだ包括的なツール空間を定義します。次に、Oracle ガイドによるデータ合成戦略を紹介します。このアプローチは、参照アイデアをオラクル ガイダンスとして活用することで、マルチエージェント システムを操作して、論理的推論とツール呼び出しパスを効率的に再構築し、目的のない試行錯誤を指示された軌道の生成に変換します。最後に、ツールの実行結果に対するマスキング戦略を使用して、これらの合成された軌道でエージェントをトレーニングします。これにより、モデルは外部フィードバックの干渉を受けることなく、意思決定ロジックに重点を置くことができます。実験結果は、私たちの方法が SOTA ワークフローベースのベースラインよりも全体的な品質において \textbf{11.91\%} 優れていることを示しています。さらに、私たちのアプローチにより、高品質データ合成のサンプル効率が \textbf{over 10$\times$} 向上します。

原文 (English)

Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents

Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation. However, existing approaches predominantly rely on pre-defined agentic workflows. This constraint severely limits the flexibility required to navigate the vast search space of scientific literature and the complex action space of research reasoning. Recently, training Agentic LLMs has emerged as a promising direction, offering flexible reasoning frameworks and the capability for autonomous tool utilization. However, there remains a non-trivial challenge: applying previous agentic data synthesis methods to scientific ideation suffers from prohibitively high data synthesis cost. To bridge this gap, we propose Agentic-Ideation, a novel framework comprising an automated trajectory synthesis pipeline and a specialized agentic LLM trained for scientific ideation. Specifically, we first define a comprehensive tool space incorporating three external tools and three cognitive tools. Then we introduce an Oracle-Guided Data Synthesis strategy. By leveraging a reference idea as oracle guidance, this approach steers the multi-agent system to efficiently reconstruct the logical reasoning and tool invocation paths, transforming aimless trial-and-error into directed trajectory generation. Finally, we train the agent on these synthesized trajectories, employing a masking strategy on tool execution results. This ensures the model focuses on decision-making logic without interference from external feedback. Experimental results demonstrate that our method outperforms the SOTA workflow-based baseline by \textbf{11.91\%} in overall quality. Furthermore, our approach improves the sample efficiency of high-quality data synthesis by \textbf{over 10$\times$}.

13:00 JST研究/論文

Delta-JEPA: 潜在的な差分デコーディングによる行動に敏感な世界モデルの学習

計画のための視覚世界モデルを学習するには、アクションに敏感なままのコンパクトな潜在ダイナミクスが必要ですが、再構成のない関節埋め込み目標はアクションに鈍感な表現に崩壊する可能性があります。私たちは、潜在差分アクション デコーダー (LDAD) を使用して潜在前方予測を強化する、エンドツーエンドの再構成のない世界モデルである Delta-JEPA を提案します。連結されたエンドポイントの埋め込みからアクションを推論する逆デコーダーとは異なり、LDAD は、連続する観測間の潜在的な変位から実行されたアクションを再構築します。この変位レベルの監視により、遷移ジオメトリが直接規則化されます。隣接する埋め込みはアクション情報を失うことなく崩壊することができず、ロールアウトベースの計画において、区別可能な潜在的な変更を誘発するために異なるアクションが奨励されます。 Delta-JEPA は、潜在予測とアクションの再構築のみを使用し、ピクセルの再構築と分布マッチングの正則化を回避します。 Delta-JEPA は、4 つの視覚的連続制御タスクにわたって、JEPA ベースおよび表現学習ワールド モデル ベースラインよりも計画を改善します。アブレーションは、変位に基づくアクションのデコードがエンドポイント連結よりも一貫して効果的であることを示し、アクション感度分析は、より明確なアクション条件付き潜在反応を示します。これらの結果は、潜在的な差異の監視が、崩壊に強くアクションに敏感な世界モデル学習のための単純かつ効果的なメカニズムであることを示しています。

原文 (English)

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint embeddings, LDAD reconstructs the executed action from the latent displacement between consecutive observations. This displacement-level supervision directly regularizes transition geometry: adjacent embeddings cannot collapse without losing action information, and different actions are encouraged to induce distinguishable latent changes for rollout-based planning. Delta-JEPA uses only latent prediction and action reconstruction, avoiding pixel reconstruction and distribution-matching regularizers. Across four visual continuous-control tasks, Delta-JEPA improves planning over JEPA-based and representation-learning world model baselines. Ablations show that displacement-based action decoding is consistently more effective than endpoint concatenation, and action-sensitivity analyses show clearer action-conditioned latent responses. These results indicate that supervising latent differences is a simple and effective mechanism for collapse-resistant and action-sensitive world model learning.

13:00 JSTLLM/生成AIエージェント

エンボディド CAD: パラメトリック B-Rep アセンブリ モデリング用のソルバー接地 LLM エージェント

大規模な言語モデルでは妥当な CAD スクリプトを作成できますが、信頼性の高い産業用 CAD モデリングには、構文的に有効なコード以上のものが必要です。すべてのフィーチャ、配置、およびアセンブリ関係が、パラメトリック境界表現ジオメトリとして編集可能なままでありながら、正確な幾何学カーネルによって受け入れられる必要があります。パラメトリック B-Rep アセンブリ モデリング用のエンボディド CAD、ソルバーベースの LLM エージェントを紹介します。エージェントは 1 回のパスで完全なスクリプトを生成するのではなく、階層化された L0 ~ L4 CAD スキル ライブラリからアクションを繰り返し選択し、それらを型指定された幾何学的操作に解決して CAD バックエンドで実行し、ソルバー フィードバックを使用して計画、修復、学習を行います。このフレームワークは、アクション文法の制約、決定論的なパラメータ解決、教師ありウォームアップと GRPO スタイルの改良に対するソルバー由来の報酬を組み合わせています。当社は、ソルバーに合わせたメトリクス(実行可能率、スキルの精度、オペレーションファミリーの精度、正確なポリシーの精度、タスク完了の成功率)を使用して、マルチステップの機械、産業機器、金型指向の組み立てタスクでエンボディド CAD を評価します。結果は、ソルバーベースのプランニングが現在のベンチマークですべての強力なプランナー ワークフローを実行する一方で、学習済みコントローラーは高い実行可能レートに達し、有効なツール呼び出しと正確な長期ポリシー予測の間に残っているギャップを明らかにすることを示しています。

原文 (English)

Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling

Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry. We present Embodied CAD, solver-grounded LLM agents for parametric B-Rep assembly modeling. Instead of generating a complete script in one pass, the agent iteratively selects actions from a stratified L0-L4 CAD skill library, resolves them into typed geometric operations, executes them in a CAD backend, and uses solver feedback to plan, repair, and learn. The framework combines action grammar constraints, deterministic parameter resolution, and solver-derived rewards for supervised warm-up and GRPO-style refinement. We evaluate Embodied CAD on multi-step mechanical, industrial equipment, and mold-oriented assembly tasks using solver-aligned metrics: executable rate, skill accuracy, operation-family accuracy, exact policy accuracy, and task completion success. The results show that solver-grounded planning executes all strong-planner workflows in the current benchmark, while learned controllers reach high executable rates and expose the remaining gap between valid tool calls and exact long-horizon policy prediction.

13:00 JST研究/論文

言語と記号表現間のモダリティ切り替えによる空間推論

人間の推論は本質的に多様です。問題が困難になったとき、私たちは言葉だけで考えることはほとんどありません。私たちは、根底にある概念構造を理解し、間違いを避けるために、図をスケッチしたりグリッドを描いたりすることで推論を外部化することがよくあります。この前提に基づいて、私たちの研究は次のことを調査します。(a) マルチホップのテキスト空間ストーリーをレイアウトやグリッドなどの幾何学認識モダリティに根付かせることで、自然言語ベースの推論と比較して推論が向上するかどうか。 (b) いつ自然言語推論に依存するか、いつ構造化モダリティに切り替えるかをモデルが決定できるかどうか。私たちは、信頼性と複雑さの信号に基づいたスイッチング メトリックを導入することで、これらの疑問に対処します。このメトリックは、空間ストーリーを構造に定着させるとパフォーマンスが向上する可能性が高いと推定します。これは、大規模言語モデル (LLM) 推論における原則に基づいたモダリティ選択への第一歩となります。私たちの設定全体で、自然言語ベースの推論からグリッドベースの表現に切り替えると、LLM のパフォーマンスが最大 42\% 向上し、推論の結果を形成する際のモダリティの選択の重要性が強調されます。

原文 (English)

Spatial Reasoning via Modality Switching Between Language and Symbolic Representation

Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketching diagrams or drawing grids to understand the underlying conceptual structure and avoid mistakes. Building on this premise, our research investigates: (a) whether grounding multi-hop textual-spatial stories into geometry-aware modalities, such as layouts or grids, improves reasoning compared to natural language-based inference; and (b) whether a model can decide when to rely on natural language reasoning and when to switch to a structured modality. We address these questions by introducing a switching metric based on trustworthiness and complexity signals, which estimates when grounding a spatial story into structure is likely to improve performance. This takes a first step toward principled modality selection in Large Language Model (LLM) reasoning. Across our settings, switching from natural language-based reasoning to a grid-based representation improves LLM performance by up to 42\%, highlighting the importance of modality choice in shaping reasoning outcomes.

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTGeminiQwenDeepSeek

浮動小数点エラー分類に関する大規模言語モデルのベンチマーク

この論文では、ソフトウェア コード内の浮動小数点エラーを静的に検出および分類する大規模言語モデル (LLM) の機能を調査します。 InterFLOPBench は、浮動小数点誤差の 6 つのカテゴリ (キャンセル、比較、ゼロ除算、オーバーフロー、アンダーフロー、NaN) にわたって LLM を評価するように設計された 1 130 のテスト サンプルを備えた 90 C カーネルのベンチマークであり、14 の LLM にわたって比較されます。評価フレームワークは、浮動小数点エラー検出をマルチラベル分類問題として扱い、F1 スコア メトリックを使用してパフォーマンスを測定します。結果は、最新モデル (Qwen 3 32b、Gemini 2.5 Flash、Phi 4 Reasoning、DeepSeek R1T2、および gpt-oss 20b および 120b) が全体の F1 スコア 0.88 を超えるパフォーマンスを達成していることを示しています。パフォーマンスは、エラー カテゴリ間、ゼロ除算 (平均 F1 スコア: 0.8479) などの明示的な演算と、アンダーフロー (平均 F1 スコア: 0.6059) やキャンセル (平均 F1 スコア: 0.6164) などのより微妙な数値現象の間で異なります。

原文 (English)

Benchmarking Large Language Models on Floating-Point Error Classification

This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software code. We introduce InterFLOPBench, a benchmark of 90 C kernels with 1 130 test samples designed to evaluate LLMs across six categories of floating-point error: cancellation, comparison, division by zero, overflow, underflow and NaN, compared across 14 LLMs. The evaluation framework treats floating-point error detection as a multi-label classification problem and employs the F1-score metric to measure performance. Results demonstrate that latest models (Qwen 3 32b, Gemini 2.5 Flash, Phi 4 Reasoning, DeepSeek R1T2, and gpt-oss 20b and 120b) achieve a performance greater than 0.88 overall F1-score. Performance varies between error categories, between explicit operations such as division by zero (Average F1-score: 0.8479) and more subtle numerical phenomena such as underflow (Average F1-score: 0.6059) and cancellation (Average F1-score: 0.6164).

13:00 JST研究/論文

HistoriQA-ThirdRepublic: 歴史研究、フランス第三共和制 (1870 ~ 1940 年) の議会討論のためのマルチホップ質問応答コーパス

HistoriQA-ThirdRepublic: フランス第三共和国の議会での議論や新聞から派生したマルチホップの歴史的質問のフランス語データセットを紹介します。歴史家と協力して設計されたこのコーパスは、資料間の統合、時間的推論、まばらな証拠の統合など、歴史調査に典型的な複雑な推論パターンを捉えています。このデータセットは 1782 の質問で構成されており、異種の歴史文書にわたるマルチホップ接続を強調しており、ドメイン固有のコンテキストで検索拡張された大規模言語モデル システムを評価するためのリソースを提供します。ソースの選択と調整、質問の検証、メタデータの統合など、コーパスを構築するための方法論について説明します。データセットはフランスの歴史文書に焦点を当てていますが、私たちの方法論は他の言語や各国のコーパスにも容易に適応できます。最後に、コーパスがマルチホップ質問応答の現実的な評価シナリオをサポートし、NLP ベンチマークと歴史的学問のニーズとの間のギャップを埋める方法を示します。

原文 (English)

HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, the corpus captures complex reasoning patterns typical of historical inquiry, including cross-source synthesis, temporal reasoning, and the integration of sparse evidence. The dataset is made of 1782 questions and emphasizes multi-hop connections across heterogeneous historical documents, providing a resource for evaluating retrieval-augmented and large language model systems in domain-specific contexts. We describe the methodology for constructing the corpus, including the selection and alignment of sources, question validation, and metadata integration. While the dataset focuses on French historical documents, our methodology can be readily adapted to other languages and national corpora. Finally, we demonstrate how the corpus can support realistic evaluation scenarios for multi-hop question answering, bridging the gap between NLP benchmarks and the needs of historical scholarship.

13:00 JST研究/論文

CryoACE: Cryo-EM での正確かつ自動化されたモデル構築のための原子中心のフレームワーク

クライオ EM 密度マップからのタンパク質のオートモデリングは、物理化学的妥当性の強化と構造の不均一性の管理において独特の課題に直面しています。現在のソルバーは、多くの場合、静的予測に限定されているか、計算負荷の高いヒューリスティック検索を必要とします。我々は、均一構造と不均一構造の両方の正確な原子グラフを再構築するエンドツーエンドのフレームワークである CryoACE を紹介します。私たちの手法は 2 つの重要な革新を特徴としています。1 つは原子中心の再構成パラダイムです。原子中心の再構成パラダイムでは、密度特徴が原子座標で直接サンプリングされ、繰り返しリサイクルされて構造が洗練され、効率的なマルチモーダル融合のために高価なボクセル畳み込みが置き換えられます。そして、予測されたローカル解決事前確率を活用して動的な曖昧さを解決する、トレーニング不要のガイダンス メカニズム。新しく構築された高品質のデータセットで検証された CryoACE は、静的ベンチマークで既存のベースラインを大幅に上回り、事前に構築された静的構造に依存することなく、EMPIAR-10345 のような複雑な実世界のデータセットで原子レベルの動的立体構造を初めて明らかにしました。

原文 (English)

CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity. Current solvers are often limited to static predictions or require computationally intensive heuristic searches. We present CryoACE, an end-to-end framework that reconstructs precise atomic graphs for both homogeneous and heterogeneous structures. Our method features two key innovations: an atom-centric reconstruction paradigm, where density features are sampled directly at atomic coordinates and iteratively recycled to refine structures, replacing expensive voxel convolutions for efficient multimodal fusion; and a training-free guidance mechanism that leverages predicted local resolution priors to resolve dynamic ambiguity. Validated on a newly constructed high-quality dataset, CryoACE significantly outperforms existing baselines on static benchmarks and, for the first time, unveils atomic-level dynamic conformations on complex real-world datasets like EMPIAR-10345 without relying on pre-built static structures.

13:00 JST研究/論文

6G ネットワークにおける結合 OFDM 波形設計と RIS 構成のための最適化アルゴリズム: 凸緩和から基礎モデルまで

6G 向けの統合 OFDM-RIS 最適化は、合計レートの最大化、エネルギー効率、最大-最小公平性、およびピーク対平均電力比 (PAPR) で制約された目標をカバーする混合整数非線形計画法 (MINLP) 問題です。 2021 年から 2026 年の間に発表された 78 件の共同 OFDM-RIS 最適化作品が調査されています。標準化されたベンチマークは存在せず、論文間の比較は依然として不可能です。この調査では、これらの研究を 4 つのパラダイムに分類しています: (I) モデルベースの凸緩和、(II) ヒューリスティックおよびメタヒューリスティック検索、(III) 深層強化および教師なし学習、(IV) 基礎モデル (FM)、拡散ベースの生成 AI、量子最適化を含む新興手法。自己報告されたベンチマークの文献総合によると、ML ベースの手法 (パラダイム ~ III) は、10^2 ~ 10^4 x より高速な推論ごとの実行時間で 95 ~ 99% のモデルベースのスペクトル効率を報告することが示されています (メソッドのペアに依存します。文献値は自己申告であり、ML の事前トレーニング コストは除外されています)。 N=16、N=64、および N=128 での関連チュートリアル ベンチマークでは、重要なスケーリング特性が明らかになります。GPU ベースのニューラル ネットワーク推論 (DDQN、PPO、グラフ ニューラル ネットワーク (GNN)、教師なし DL) は N 不変であり、N=16 と N=128 で同一のランタイムを持ちますが、反復ソルバー (AO+SCA、PSO) は多項式にスケーリングします。エネルギー効率 (P2) および PAPR 制約付き (P4) ベンチマークは、標準化された電力モデルと波形発生器を使用した将来の作業に延期されます。合成の結果、6 つの未解決の課題が明らかになります。それは、クロスパラダイムベンチマークの不足、現実世界のハードウェア制約のある展開、二重分散チャネルの波形と RIS の統合最適化、多目的 PAPR トレードオフ、ライブネットワーク制御における LLM の安全性、スタンドアロンヒューリスティックの利益逓減です。標準化されたベンチマークの要件を指定します。この研究は、6G ネットワークにおける OFDM-RIS の共同最適化に取り組む研究者や実務者にとってのロードマップとして役立ちます。

原文 (English)

Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models

Joint OFDM-RIS optimization for 6G is a mixed-integer nonlinear programming (MINLP) problem covering sum-rate maximization, energy efficiency, max-min fairness, and peak-to-average power ratio (PAPR)-constrained objectives. Seventy-eight joint OFDM-RIS optimization works published between 2021 and 2026 are surveyed. No standardized benchmark exists, and cross-paper comparisons remain infeasible. This survey classifies these works into four paradigms: (I) model-based convex relaxation, (II) heuristic and metaheuristic search, (III) deep reinforcement and unsupervised learning, and (IV) emerging methods including foundation models (FM), diffusion-based generative AI, and quantum optimization. A literature synthesis of self-reported benchmarks shows that ML-based methods (Paradigm~III) report 95-99\% of model-based spectral efficiency at 10^2-10^4 x faster per-inference runtime (method-pair dependent; literature values are self-reported and exclude ML pre-training cost). A companion tutorial benchmark at N=16, N=64, and N=128 reveals a critical scaling property: GPU-based neural network inference (DDQN, PPO, graph neural network (GNN), unsupervised DL) is N-invariant, with identical runtime at N=16 and N=128, while iterative solvers (AO+SCA, PSO) scale polynomially. Energy efficiency (P2) and PAPR-constrained (P4) benchmarks are deferred to future work with standardized power models and waveform generators. Six open challenges emerge from the synthesis: the cross-paradigm benchmark deficit, real-world hardware-constrained deployment, joint waveform-RIS optimization for doubly-dispersive channels, multi-objective PAPR trade-offs, LLM safety in live network control, and diminishing returns of standalone heuristics. We specify requirements for a standardized benchmark. This study serves as a roadmap for researchers and practitioners working on joint OFDM-RIS optimization in 6G networks.

13:00 JSTエージェント

大規模な電気自動車のスマート充電: 独立したマルチエージェント強化学習アプローチ

電気自動車による交通機関の電化は、ピーク需要の増加、電圧変動、送電線の過負荷、変動する再生可能エネルギー源の統合など、電力網管理に新たな課題をもたらします。ユーザーのコストを最小限に抑え、ネットワークの過負荷を回避しながら、EV の効率的な統合を可能にするには、EV 間の暗黙的な調整が必要です。この研究では、このような分散型 EV 充電を最適化するための 2 つの独立したマルチエージェント強化学習アプローチ、コンテキスト組み合わせバンディットとポリシー勾配アルゴリズムを比較します。ローカル環境情報 (価格シグナル、充電状態、時間的制約など) に基づいて意思決定を行う自律エージェントを備えた現実的なシミュレーション環境を使用して、実際の太陽光発電データから導出された動的な電力価格設定の下で、さまざまな混雑レベル、および異種エージェント グループによる混合戦略構成にわたるパフォーマンスを評価します。

原文 (English)

Smart charging of large fleets of Electric Vehicles: Independent Multi-Agent Reinforcement Learning approaches

The electrification of transportation through electric vehicles introduces new challenges for power grid management, such as increased peak demand, voltage fluctuations, line overloads, and the integration of variable renewable energy sources. To enable efficient integration of EVs while minimizing costs for users and avoiding network overloads, implicit coordination between EVs is required. This work compares two independent multi-agent reinforcement learning approaches for optimizing such decentralized EV charging: contextual combinatorial bandits and policy gradient algorithms. Using a realistic simulation environment with autonomous agents making decisions based on local environmental information (including price signals, state-of-charge, and temporal constraints), we evaluate their performance across varying congestion levels, and mixed-strategy configurations with heterogeneous agent groups under dynamic electricity pricing derived from real photovoltaic production data.

13:00 JSTエージェント

ReGRPO: ツールを使用するエージェント向けのリフレクション拡張ポリシーの最適化

ツール拡張ビジョン言語モデル (VLM) は、外部ツールを呼び出すことでマルチモーダル、複数ステップのタスクを解決できますが、実際には脆弱なままです。既存の作品には 2 つの共通のギャップがあります。教師あり微調整 (SFT) は、主に成功した軌跡に基づいて構築されており、ツールの失敗後の回復のシグナルはほとんど提供されませんが、まばらな軌跡レベルの RL 報酬は、失敗したステップとその修復方法について限定的なガイダンスを提供します。ツールを使用するエージェントでリフレクションに基づく修正を学習するフレームワークである ReGRPO (Reflection-augmented Group Relative Policy Optimization) を紹介します。 ReGRPO は、構造化された反映データ エンジンから始まります。ニアミス アクションを実行して、根拠のある障害の観察を収集し、その後、ウォーム スタート SFT の修正アクションと組み合わせた思考の反映トリプレット (ErrorType、Evidence、FixPlan) を構築します。次に、グループ相対の利点を使用して、ローカル トラジェクトリ内でリフレクション トークンと修正措置を共同で最適化し、不要なリフレクションを減らすためにリフレクション コスト項を組み込みます。 GTA と GAIA の実験では、同じバックボーンとツール スイートの下で、ReGRPO が一貫して強力なオープンソース ベースラインを上回り、比較したオープンソース コントローラーの中で最高の結果を達成することが示されています。コードと RoT データは https://github.com/showlab/ReGRPO で入手できます。

原文 (English)

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents

Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Existing works have two common gaps. Supervised fine-tuning (SFT) is built mostly on successful trajectories and offers little signal for recovery after tool failures, while sparse trajectory-level RL rewards provide limited guidance on which step failed and how to repair it. We introduce ReGRPO (Reflection-augmented Group Relative Policy Optimization), a framework that learns reflection-guided correction in tool-using agents. ReGRPO starts with a structured reflective data engine: we execute near-miss actions to collect grounded failure observations, then build Reflection-of-Thought triplets (ErrorType, Evidence, FixPlan) paired with corrected actions for warm-start SFT. We then optimize reflection tokens and corrective actions jointly within local trajectories using group-relative advantages, and include a reflection-cost term to reduce unnecessary reflection. Experiments on GTA and GAIA show that, under the same backbone and tool suite, ReGRPO consistently outperforms strong open-source baselines and achieves the best results among the compared open-source controllers. Code and RoT data are available at https://github.com/showlab/ReGRPO.

13:00 JSTエージェント

相転移としての世界モデルの崩壊

水は温まるにつれて変化しないように見えますが、臨界点で沸騰します。私たちは、長期的な言語エージェントが暗黙の世界モデルにおいて同様の遷移を示すかどうかを尋ねます。一部のパラメーター設定では、ステート ロードを少量変更するか、ホライズンの 1 ステップを追加すると、動作はほとんど変わりません。臨界境界付近では、同じ小さな変化が突然の世界崩壊を引き起こします。この効果を、ステップごとの正確なゴールド状態を使用した決定論的タスク ファミリで研究します。状態カーディナリティ、依存関係密度、ホライズン、分岐、観察モード、突然変異率に関する大規模なグリッド検索により、解決されたプラトー、狭い遷移バンド、および崩壊フロアなどの状態図が明らかになります。ステップごとのトレースはメカニズムを示します。ワールド状態の忠実性はアクションの有効性の前に失敗するため、エージェントは単に間違ったアクションを選択しているだけではありません。それは腐敗した世界から行動しているのです。より強力なモデルは臨界境界を変換しますが、定性的な移行は除去しません。これらの結果により、世界モデルの崩壊が長期的なエージェントにとって目に見えるボトルネックとなっています。

原文 (English)

World-Model Collapse as a Phase Transition

Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transition in their implicit world models. In some parameter settings, changing state load by a small amount, or adding a single step of horizon, leaves behavior nearly unchanged; near a critical boundary, the same small change causes a sudden world collapse. We study this effect in a deterministic task family with exact per-step gold state. A large grid search over state cardinality, dependency density, horizon, branching, observation mode, and mutation rate reveals a phase diagram: a solved plateau, a narrow transition band, and a collapse floor. Per-step traces show the mechanism: world-state fidelity fails before action validity, so the agent is not merely choosing a bad action; it is acting from a corrupted world. Stronger models translate the critical boundary but do not remove the qualitative transition. These results make world-model collapse a measurable bottleneck for long-horizon agents.

13:00 JST研究/論文ClaudeGPT / ChatGPTGemini

(AI) 群衆の知恵: 大規模な言語モデルにおける人工群知能の調査

人間の群知能は、顕著な集団的精度を示しますが、コスト、調整、時間のスケーラビリティの制約に直面しています。私たちは、大規模言語モデル (LLM) が人工的な群を介して群知能の効果を近似できるかどうかを調査し、AI ベースの集計メカニズムを理解する際の重大なギャップに対処します。私たちは、3 つの独自モデル (GPT-5、Gemini 2.5 Pro、Claude Sonnet 4.5) にわたって手動で実行される 960 のプロンプトを使用した制御実験を実施し、8 つの推定タスクでモデル内サンプリングとモデル間集計をテストしました。結果は、モデル内およびモデル間の集計を通じて一貫したエラー削減を示し、さまざまな集計戦略にわたって MAPE で最大 37 パーセント ポイントの大幅なエラー削減を実現しました。相対的な信頼区間幅と相対的な推定誤差の間の正の相関関係(スピアマンの $\rho=0.242-0.568$、すべて $p<0.001$)について、大小のエフェクト サイズが観察されました。これは、LLM が不確実性を評価する際にメタ認知的な認識を持っていることを示唆しています。研究と実践への影響について議論し、組織の意思決定に LLM 群を導入するための実用的な洞察を提供します。

原文 (English)

Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models

Human swarm intelligence demonstrates remarkable collective accuracy but faces scalability constraints in cost, coordination, and time. We investigate whether large language models (LLMs) can approximate swarm intelligence effects through artificial swarms, addressing a critical gap in understanding AI-based aggregation mechanisms. We conducted a controlled experiment with 960 manually executed prompts across three proprietary models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4.5), testing intra-model sampling and inter-model aggregation on eight estimation tasks. Results reveal consistent error reduction through intra- and inter-model aggregation, with significant error reductions up to 37 percentage points in MAPE across different aggregation strategies. We observed small to large effect sizes for positive correlations (Spearman's $\rho=0.242-0.568$, all $p<0.001$) between relative confidence interval widths and relative estimation errors, suggesting LLMs possess metacognitive awareness when assessing uncertainty. We discuss implications for research and practice, providing actionable insights for deploying LLM swarms in organizational decision-making.

13:00 JSTエージェント

Xiaomi-GUI-0テクニカルレポート

グラフィカル ユーザー インターフェイス (GUI) エージェントは、ビジョン言語モデルに基づいて構築され、タップ、スワイプ、テキスト入力、ナビゲーションなどのインターフェイス アクションを通じて実際のアプリケーションでユーザー タスクをエンドツーエンドで完了します。ただし、既存の GUI エージェントは、主にオフラインの軌跡、シミュレートされた環境、および標準化されたベンチマークに基づいてトレーニングおよび評価されます。これらは、インターフェイスのレイアウト、インタラクション ロジック、異常状態の分布において実際のアプリケーションとは大幅に異なり、実際の使用における実行の安定性を忠実に特徴付けることはできません。そこでは、アカウントの状態、許可ダイアログ、支払い認証、リスク管理によって状態分布が継続的に再形成され、ベンチマーク スコアと実際のユーザビリティの間に永続的なギャップが生じます。このギャップを埋めるために、実際のデバイスの閉ループ内でトレーニングおよび評価される、実際のモバイル環境用のネイティブ マルチモーダル GUI エージェントである Xiaomi-GUI-0 を提案します。その中核となるのは、実デバイス主体のハイブリッド インフラストラクチャであり、物理デバイスが主要な実行環境であり、サンドボックスが補助的なサポートを提供するため、データ収集、トレーニング、ロールアウト、評価が実際の展開に近い実行分布を共有します。私たちは、高頻度のヘッド タスク、ロングテール インテントの高汎化データ、リフレクションとメモリの能力強化データにわたるマルチソース トレーニング データを構築し、障害の軌跡を修正されたアクション、リフレクションの説明、回復のデモンストレーションに変えるエラー駆動型のデータ フライホイールを導入します。モデルは、教師あり微調整、ステップレベルの強化学習、エージェント強化学習の漸進的な 3 段階のパイプラインを通じてトレーニングされます。公開ベンチマークと社内 RealMobile で評価された Xiaomi-GUI-0 は、RealMobile で 72.0%、AndroidWorld で 78.9% の成功率を達成し、現実世界のタスクにおける実行の安定性と異常状態の認識が大幅に向上しました。

原文 (English)

Xiaomi-GUI-0 Technical Report

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.

13:00 JST研究/論文

学び直すのではなく、選択することを学ぶ: LoRA の推論を厳密に組み合わせたもの

独立してトレーニングされた LoRA アダプターを単一の大規模な言語モデルに構成すると、特に元のトレーニング データを共有できない場合に、マルチドメインの適応に役立ちます。一般的なアプローチは、LoRA エキスパートに対して MoE スタイルのルーティングを使用することですが、凍結された事前トレーニング済みアダプターの場合、ソフト重み付けの組み合わせにより、各 LoRA モジュールが最初にトレーニングされたユニットスケールの追加更新が変更される可能性があります。私たちは、ユニットスケールのハード選択を通じて凍結推論 LoRA エキスパートを構成するための 2 段階のフレームワークである \textbf{Hard-Routed MoR-LoRA} を提案します。まず、ドメイン固有の LoRA アダプターは、検証可能なフィードバックからの強化学習を使用して独立してトレーニングされ、推論の専門家を取得します。次に、すべての専門家が凍結され、そこから推論の痕跡が抽出され、小さな注意を払った LoRA を備えた軽量の共有ルーターのみが統合用にトレーニングされます。ルーターは、ハードトップ 1 ルーティングを使用してトークンごとに 1 人のエキスパートを正確に選択し、ストレートスルー推定器により勾配ベースのトレーニングが可能になります。 5 つのベンチマーク、複数のモデル スケール、および追加のモデル ファミリにわたる実験により、ハード ルーティング MoR-LoRA は、ソフト ルーティング混合ベースラインよりも大幅に少ないトレーニング可能なパラメータを必要としながらも、専門家の動作を維持することが示されています。さらに、我々の分析では、正規化されたソフト混合物はほとんどのルーティングマスを 1 人のエキスパートに集中させることが多いことを示しており、ハードユニットスケールのルーティングが凍結した LoRA エキスパート構成にシンプルで効率的な抽象化を提供することを示唆しています。

原文 (English)

Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs

Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the original training data cannot be shared. A common approach is to use MoE-style routing over LoRA experts, but for frozen pretrained adapters, soft weighted combinations can change the unit-scale additive update under which each LoRA module was originally trained. We propose \textbf{Hard-Routed MoR-LoRA}, a two-stage framework for composing frozen reasoning LoRA experts through unit-scale hard selection. First, domain-specific LoRA adapters are trained independently using reinforcement learning from verifiable feedback to obtain reasoning experts. Then, all experts are frozen, reasoning traces are distilled from them, and only a lightweight shared router together with a small attention LoRA is trained for integration. The router selects exactly one expert per token using hard top-1 routing, while a straight-through estimator enables gradient-based training. Experiments across five benchmarks, multiple model scales, and additional model families show that Hard-Routed MoR-LoRA preserves expert behavior while requiring substantially fewer trainable parameters than soft-routing mixture baselines. Our analysis further shows that normalized soft mixtures often concentrate most routing mass on a single expert, suggesting that hard unit-scale routing provides a simple and efficient abstraction for frozen LoRA expert composition.

13:00 JST研究/論文

BP-TTA: 動的シナリオにおけるバランスの取れたプロトタイプに基づくテスト時間の適応

Test-Time Adaptation (TTA) を使用すると、ソース ドメインでトレーニングされたモデルが、配布シフトの下でラベルのないテスト データにオンラインで適応できます。最近の TTA 手法は静的な設定を超えて、継続的なドメイン シフトを考慮し始めていますが、主に分散ドリフトに対処しており、動的なシナリオにおけるクラスの不均衡を考慮できません。実際のテスト時のストリームでは、クラスの不均衡と継続的なドメインのシフトが同時に発生し、相互に影響し合うことがよくあります。この論文では、クラスの不均衡と継続的なドメイン シフトの問題を処理するために、バッチバランスのとれたサンプリングとプロトタイプに基づく適応を組み合わせた、新しいバランス型およびプロトタイプに基づくテスト時間適応 (BP-TTA) 手法を提案します。 BP-TTA は、現在のサンプルと信頼性の高い過去のインスタンスを統合することでバランスのとれた適応バッチを構築し、支配的なクラスに対するバイアスを効果的に軽減し、オンライン更新を安定させます。一方、BP-TTA は、推論中に進化するクラス プロトタイプを維持し、プロトタイプの類似性をモデル適応の制約として利用します。これにより、擬似ラベルの信頼性が向上し、永続的なドメイン シフト下でのオンライン更新の安定性が向上します。広範な実験により、BP-TTA は動的テスト時間ストリーミング設定において常に最先端の TTA メソッドよりも優れたパフォーマンスを発揮することが実証されています。

原文 (English)

BP-TTA: Balanced and Prototype-Guided Test-Time Adaptation in Dynamic Scenarios

Test-Time Adaptation (TTA) enables models trained on a source domain to adapt online to unlabeled test data under distribution shifts. While recent TTA methods have moved beyond static settings and begun to consider continual domain shifts, they primarily address distribution drift and fail to account for class imbalance in dynamic scenarios. In real-world test-time streams, class imbalance and continual domain shifts often occur at the same time and interact with each other. In this paper, we propose a novel Balanced and Prototype-Guided Test-Time Adaptation (BP-TTA) method, which combines batch-balanced sampling with prototype-guided adaptation to handle the class imbalance and continual domain shift problems. BP-TTA constructs balanced adaptation batches by integrating current samples with high-confidence historical instances, effectively mitigating bias toward dominant classes and stabilizing online updates. Meanwhile, BP-TTA maintains evolving class prototypes during inference and leverages prototype similarity as a constraint for model adaptation, thereby improving the reliability of pseudo-labels and enhancing the stability of online updates under persistent domain shifts. Extensive experiments demonstrate that BP-TTA consistently outperforms state-of-the-art TTA methods in dynamic test-time streaming settings.

13:00 JSTエージェント

行動する前に世界に問う:世界モデルのキャリブレーションのための予算環境調査

長期的な言語エージェントは、アクションを選択するだけではありません。彼らは、ある決断から次の決断まで、世界のプライベートモデルを持ち続けています。モデルがドリフトすると、失敗したアクションが実行される前に、その後の失敗が決定される可能性があります。私たちは直接修復メカニズムを研究します。次のタスクのアクションにコミットする前に、エージェントは環境に 1 つの信念フィールドについて質問し、その答えをワールド モデルに書き戻すことができます。このため、環境との相互作用は、単にタスクを進めるための手段ではなく、希少なキャリブレーション リソースになります。構造化信念表の予算探索演算子である \method を紹介します。有用なプローブはどこでも同じではありません。ツールの依存関係などの手順上の信念は、多くの場合、対象を絞ったチェックによって修復できますが、それらのチェックにはタスクに必要なステップが費やされます。オブジェクトの位置やグラフの端などの空間的信念は、構造的な手がかりに大きく依存します。世界が画面外で変化するとき、エージェント自身の自信は不十分な指針になる可能性があります。タイプ階層化分析は、このプローブとアクションのフロンティアを形式化し、管理された実験により、プローブ ポリシーがタスクの構造に従う場合、計画途中の環境の証拠によって最終的な世界モデルのエラーが減少することが示されています。

原文 (English)

Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration

Long-horizon language agents do not only choose actions; they carry a private model of the world from one decision to the next. When that model drifts, a later failure can be decided before the failing action is ever taken. We study a direct repair mechanism: before committing to the next task action, an agent may ask the environment about one belief field and write the answer back into its world model. This makes environment interaction a scarce calibration resource, not merely a way to advance the task. We introduce \method, a budgeted probing operator for structured belief tables. The useful probes are not the same everywhere. Procedural beliefs, such as tool dependencies, can often be repaired by targeted checks, but those checks spend steps that the task may need. Spatial beliefs, such as object locations and graph edges, rely more on structural cues; the agent's own confidence can be a poor guide when the world changes off-screen. A type-stratified analysis formalizes this probe-action frontier, and controlled experiments show that mid-planning environment evidence reduces terminal world-model error when the probe policy follows the structure of the task.

13:00 JSTLLM/生成AI

CDR ベンチ: 構成的で順序に依存したデータ調整レシピの忠実な実行の評価

データのリファインメントには、進化するテキスト状態に対して複数ステップのレシピを実行することが含まれ、処理演算子の構成と実行順序の両方が結果を決定します。既存のベンチマークはテキスト編集を分離するか、コードやツールの実行と絡めるかのいずれかですが、LLM がこれらの構成的で順序に依存したデータ調整レシピを直接かつ忠実に実行できるかどうかは依然として不明です。このギャップを埋めるために、4 つの現実世界のデータ改良ドメインと 29 の異なるオペレーターにわたる 3,462 の高品質タスクを特徴とする包括的なベンチマークである CDR-Bench を導入します。私たちのベンチマークは、正確な評価を可能にする決定論的なリファレンス出力を活用して、アトミック、順序に依存しない、および順序に依存する設定全体でモデルを評価します。 10 を超える最先端の LLM での実験により、一貫した失敗パターンが明らかになりました。構成設定でパフォーマンスが急激に低下し、順序に依存するレシピの成功が崩壊します。これらの発見は、現在のLLMには、信頼性の高い構成データの改良に必要な手順の忠実さが欠けていることを強調しています。

原文 (English)

CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators determine the outcome. While existing benchmarks either isolate text editing or entangle it with code and tool execution, it remains unclear whether LLMs can directly and faithfully execute these compositional, order-sensitive data refinement recipes. To fill this gap, we introduce CDR-Bench, a comprehensive benchmark featuring 3,462 high-quality tasks spanning four real-world data refinement domains and 29 distinct operators. Our benchmark evaluates models across atomic, order-agnostic, and order-sensitive settings, leveraging deterministic reference outputs to enable exact evaluation. Experiments on 10+ state-of-the-art LLMs reveal consistent failure patterns: performance degrades sharply in compositional settings, and order-sensitive recipe success collapses. These findings underline that current LLMs lack the procedural faithfulness required for reliable compositional data refinement.

13:00 JSTエージェント

感情の意味を決めるのは誰ですか?測定限界の認識論的帰結としての感情主権

感情を感知する AI は、自動車、家電製品、対話エージェント、社会インフラに急速に組み込まれるようになり、感情がもはや個人の経験に限定されるのではなく、社会規模で観察および計算される領域、つまり私たちがアフェクトスフィアと呼ぶ領域を生み出しています。しかし、この領域における中心的な規範的問題は、まだ十分に解明されていない。それは、自分自身の感情の意味を決定する最終的な権限は誰にあるのかということである。この研究は、測定の構造的限界の認識論的側面から問題に取り組んでいます。意味分布を、固定アノテーションプロトコルの下で母集団から抽出されたアノテーターによって割り当てられたラベルの分布として定義し、その不確実性を還元可能な要素と還元不可能な要素に分解します。次に、感情 AI は信頼度の高い点ラベルを割り当て、集合レベルで実際の差異を識別できる一方で、個々のインスタンスの意味分布の還元不可能な要素は、現実的なアノテーター数の下では適切な範囲で推定することができず、体系的な発散を認識論的ギャップと呼んでいることを示します。重要な発見は、デバイスの信頼性が高いということは、回復不可能な意味が回復されたという証拠にはならないということです。この認識論的ギャップと、明示的に述べられた規範的前提、つまり原理的に量を回復できないシステムの出力は権威ある決定として扱われてはならないという前提とともに、人の感情の意味に対する最終的な解釈権限は手続き的に経験している主体に留保されているという規範、感情主権の規範を導き出す。これらの結果は、感情 AI の設計、評価、制御では、精度の最大化ではなく、解釈権限の明示的な割り当てを中心に据えるべきであることを示唆しています。

原文 (English)

Who Determines the Meaning of an Emotion? Affective Sovereignty as an Epistemic Consequence of Measurement Limits

Emotion-sensing AI is rapidly becoming embedded in vehicles, home appliances, dialogue agents, and social infrastructure, giving rise to a sphere in which emotion is no longer confined to individual experience but is instead observed and computed at a societal scale, a domain we term the Affectosphere. Yet a central normative question in this domain has remained underexplored: who has the final authority to determine the meaning of one's own emotion? This study addresses the question from the epistemological side of measurement's structural limits. We define a meaning distribution as the distribution of labels assigned by annotators drawn from a population under a fixed annotation protocol, and decompose its uncertainty into reducible and irreducible components. We then demonstrate that, while emotion AI can assign high-confidence point labels and discriminate real differences at an aggregate level, the irreducible component of the meaning distribution for individual instances cannot be estimated with adequate coverage under realistic annotator counts, a systematic divergence we term the epistemic gap. The key finding is that high device confidence does not constitute evidence that irrecoverable meaning has been recovered. From this epistemic gap, together with an explicitly stated normative premise, namely that the output of a system which cannot recover a quantity in principle must not be treated as its authoritative determination, we derive the norm that the final interpretive authority over the meaning of one's emotion is procedurally reserved for the experiencing subject, the norm of affective sovereignty. These results suggest that the design, evaluation, and regulation of emotion AI should place explicit allocation of interpretive authority, rather than accuracy maximisation, at their core.

13:00 JST研究/論文

CSTrader: コミュニティ主導の仮想資産市場における言語に基づいた取引のテストベッド

Counter-Strike 2 (CS2) の武器スキンなどのニッチな資産市場は小規模で不安定で、コミュニティの議論やプラットフォームのルールによって大きく左右されます。これらの特性により、従来の定量モデルは困難になりますが、大規模言語モデル (LLM) が非構造化テキストを取引アクションにどのように変換するかを研究するための理想的なテストベッドとなります。私たちは、CS2 スキン市場における言語ベースの取引のためのマルチエージェント フレームワークである CSTrader を紹介します。このシステムは、まずさまざまなソースからの異種シグナルを統合し、次にテクニカル分析、流動性、イベント、および(反転)センチメントに特化したエージェントを使用し、最後にリスク管理、取引摩擦、およびポートフォリオ管理エージェントを適用して、現実的な取引摩擦の下で購入、売却、またはホールドの意思決定を行います。私たちは、非常に不安定な時期の実際の CS2 データを使用してライブのような評価環境を構築し、いくつかの最近の LLM バックボーンを評価します。 CSTrader は、どのモデルにおいても、下落相場指数 (-15.62%) と単純なシングルプロンプト LLM ベースラインの両方を常に上回っており、リスクを制御しながら最大 7.58% の累積リターンを達成しています。アブレーション研究は、流動性、反転センチメント、および取引摩擦要因が、ノイズの多い言語シグナルを安定した利益に変えるために重要であることを示しており、ニッチな言語主導型市場が将来の言語対行動研究の有用なベンチマークであることを示唆しています。コードはhttps://github.com/IatomicreactorI/CSGOTrading?tab=readme-ov-file#quick-startから入手できます。

原文 (English)

CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market

Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. These properties make them hard for traditional quantitative models, but provide an ideal testbed for studying how large language models (LLMs) turn unstructured text into trading actions. We present CSTrader, a multi-agent framework for language-grounded trading in the CS2 skin market. The system first integrates heterogeneous signals from various sources, then uses specialized agents for technical analysis, liquidity, events, and (reversed) sentiment, and finally applies risk control, transaction friction, and portfolio management agents to produce buy, sell, or hold decisions under realistic trading frictions. We build a live-like evaluation environment with real CS2 data from a highly volatile period and evaluate several recent LLM backbones. Across models, CSTrader consistently outperforms both a falling market index (-15.62%) and simple single-prompt LLM baselines, achieving up to a 7.58% cumulative return with controlled risk. Ablation studies show that liquidity, reversed sentiment, and transaction friction agents are crucial for turning noisy language signals into stable profits, suggesting that niche, language-driven markets are a useful benchmark for future language-to-action research. Code is available at: https://github.com/IatomicreactorI/CSGOTrading?tab=readme-ov-file#quick-start

13:00 JST研究/論文Claude

CLOUDADV: ドリフト下でのゼロショット基盤モデルを使用した意思決定に基づいたインスタンスのサイジング

クラウド仮想マシンはオーバープロビジョニングされることが多く、回避可能なコストと運用の非効率が生じます。ワークロードの変動下でクラウド インスタンスのサイジングを行うための、エンジニア向けの対話型アドバイザリー システムである CLOUDADV を紹介します。このシステムは、ゼロショット時系列予測と、日、週、月スケールの計画期間にわたる限定された推奨事項の生成を組み合わせています。 CLOUDADV は、クエリごとに、過去の使用率、予測概要、現在の VM メタデータ、候補インスタンス オプション、価格設定、明示的なサイジング ヒューリスティックから構造化された意思決定コンテキストを構築します。高容量の LLM はオフラインで参照推奨事項を生成するために使用されますが、小規模な運用モデルは同じプロンプトで評価され、遅延とコストの制約の下で展開時の調整が評価されます。評価では、シミュレートされた Azure コスト削減と事後超過を使用して下流の推奨品質を優先し、ローリング オリジンの予測精度が従来の監視ベースラインに対する二次診断として報告されます。 7 台の運用 VM のケース スタディでは、参照推奨事項により、シミュレートされた月額コストが約 1,503 ドルから 708 ドルに削減され、保守的なヒューリスティック制約の下で月額 795 ドル (52.9%) の節約が得られます。一方、ダウングレードされたケースで観察された最高の超過率は 1.5% です。 Chronos-2 はすべての予測メトリクスを最小化するわけではありませんが、多くの場合、監視対象の VM ごとのベースラインと同様の推奨パターンを誘導します。これらの結果は、ゼロショット基盤モデルが、テナントごとの繰り返しの再トレーニング、再検証、および再デプロイメントによる運用負担を軽減しながら、非定常クラウド環境での意思決定に合わせたプロビジョニングをサポートできることを示唆しています。

原文 (English)

CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift

Cloud virtual machines are often overprovisioned, creating avoidable cost and operational inefficiency. We present CLOUDADV, an interactive engineer-facing advisory system for cloud instance sizing under workload drift. The system combines zero-shot time-series forecasting with bounded recommendation generation across day-, week-, and month-scale planning horizons. For each query, CLOUDADV constructs a structured decision context from historical utilization, forecast summaries, current VM metadata, candidate instance options, pricing, and explicit sizing heuristics. A higher-capacity LLM is used offline to generate reference recommendations, while a smaller production model is evaluated on the same prompts to assess deployment-time alignment under latency and cost constraints. Evaluation prioritizes downstream recommendation quality using simulated Azure cost savings and ex-post exceedance, with rolling-origin forecast accuracy reported as a secondary diagnostic against classical and supervised baselines. In a case study of seven production VMs, the reference recommendations reduce simulated monthly cost from about \$1,503 to \$708, yielding \$795/month in savings (52.9%) under conservative heuristic constraints, while the highest observed exceedance rate among downgraded cases is 1.5%. Although Chronos-2 does not minimize every forecasting metric, it often induces recommendation patterns similar to those of a supervised per-VM baseline. These results suggest that zero-shot foundation models can support decision-aligned provisioning in non-stationary cloud environments while reducing the operational burden of repeated per-tenant retraining, revalidation, and redeployment.

13:00 JST画像/動画生成エージェント研究/論文

1 回の反省だけでは十分ではない: 複数の仮説の失敗の帰属による自己修正型の自律的研究

自律的な研究エージェントは仮説を立て、コードを書き、実験を実行し、論文を作成できるようになりましたが、実験が失敗すると脆弱なままです。一般的なパラダイムでは、障害回復は通常、単一の自由形式の反映に委任されます。つまり、メトリクス、ログ、設計の選択の豊富な軌跡が 1 つの口頭での批判に圧縮され、局所的な試行錯誤や有用なコンテキストを無視するハードな方向転換につながることがよくあります。私たちは、この障害回復のボトルネックに取り組むために、自己修正型、自律型、接地型の実験者である SAGE を提案します。その中核となるメカニズムである多重仮説失敗帰属 (MHFA) は、回復を構造化された因果関係の診断として扱います。 MHFA は、動的軌跡の特徴を分析することにより、障害に対する複数の証拠に基づいた説明を体系的に生成し、その重大度を独立して評価し、検証された根本原因を正しい介入レベル (仮説、実験計画、または実装) に決定論的に導きます。科学的な誠実さを保証するために、SAGE はさらに、草稿された結果を実際の測定値に明示的に制約し、幻覚の数値を編集する、根拠のある報告メカニズムを採用しています。 12 トピック、5 ドメインのベンチマークで、SAGE はメトリクスを伴う出力をリフレクション ベースラインの 42% から 92% に増加させ、アーティファクトの品質を 5.00 から 6.75/10 に向上させ、AI-Scientist-v2 を盲目的に上回り (52.0 対 48.2)、コード開発と実行に集中した利益をもたらしました。完全に自律的な科学執筆と会議用論文の生成は、分野全体にとって依然として悪名高い困難な未解決の問題ですが、SAGE は大幅に信頼性が高く高品質な科学成果物の生成に成功しています。最終的に、構造化回復と明示的なグラウンディング制約を組み合わせることで、SAGE はモノリシック反射パラダイムを大幅に上回り、将来の自律研究のための信頼性の高い基盤を確立します。

原文 (English)

One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution

Autonomous research agents can now draft hypotheses, write code, run experiments, and produce papers, but they remain brittle when experiments fail. Under the prevailing paradigm, failure recovery is usually delegated to a single free-form reflection: a rich trajectory of metrics, logs, and design choices is compressed into one verbal critique, which often leads either to localized trial-and-error or to hard pivots that discard useful context. We propose SAGE, a Self-correcting, Autonomous, Grounded Experimenter, to tackle this failure-recovery bottleneck. Its core mechanism, Multi-Hypothesis Failure Attribution (MHFA), treats recovery as a structured causal diagnosis. By analyzing dynamic trajectory features, MHFA systematically generates multiple evidence-grounded explanations for a failure, independently evaluates their severity, and deterministically routes the verified root cause to the correct intervention level (hypothesis, experimental design, or implementation). To guarantee scientific honesty, SAGE further employs a grounded reporting mechanism that explicitly constrains drafted results to actual measured values, redacting hallucinated numbers. On a 12-topic, 5-domain benchmark, SAGE increases metrics-bearing outputs from 42% to 92% over a reflection baseline, improves artifact quality from 5.00 to 6.75/10, and blindly outscores AI-Scientist-v2 (52.0 vs. 48.2), with gains concentrated in code development and execution. While fully autonomous scientific writing and generating conference-ready papers remain notoriously difficult open problems for the entire field, SAGE successfully produces significantly more reliable and higher-quality scientific artifacts. Ultimately, by coupling structured recovery with explicit grounding constraints, SAGE significantly outperforms monolithic reflection paradigms, establishing a highly trustworthy foundation for future autonomous research.

13:00 JST研究/論文

可塑性とメタ認知への信号としての驚き

私たちは、2 つの設定にまたがる 1 つのアイデアを研究します。それは、フリーズされたエンコーダーの潜在空間上の小さな予測器によって計算された予測誤差信号が、可塑性のゲートとして、またメタ認知の基質として機能する可能性があるというものです。最初のシステムでは、ノンパラメトリックなエピソード記憶は、この驚きが大きい場合にのみ新しい概念を書き込み、定期的なオフライン再生フェーズにより、最近のトレースが低速な線形読み出しに統合されます。凍結された DINOv2 または I-JEPA バックボーンを持つ 1,000 個の ImageNet クラスの連続ストリームでは、統合フェーズにより、DINOv2 の最も古いクラスで 17.7 ポイント、I-JEPA (シングルシード実行) で 51.3 ポイントの保持が回復しました。アブレーションの結果、最近のウィンドウのみを再生するのは、まったく再生しないよりも悪いことがわかります。数ショット評価では、同じメモリが 5 ウェイ 1 ショット ミニ ImageNet で 91.6% に達し、タスク固有のベースラインを上回りましたが、より困難な 500 ウェイ体制では真の難しさが明らかになりました。 2 番目のシステムでは、共有テキスト画像空間で計算された同じ驚きの信号が視覚言語モデルの動作を調整します。概念が既知の場合は断定的に応答し、部分的によく知られている場合は回避し、新規の場合はオブジェクトを特定することを拒否し、単一のユーザーの発話から概念を学習して説明を求めます。外部検出器は、モデル自体の言語化された信頼度 (0.618) をはるかに上回る 0.966 (95% CI +/-0.024) の AUROC で既知の概念と​​新しい概念を分離しますが、そのトークンレベルの信頼度は貪欲なデコードの下では確率を下回ります。高速ストアを空にするスリープ フェーズの後、システムは統合ストアから 50 の学習事実のうち 99.2% を思い出しますが、基本モデルは何も回復しません。私たちは両方のシステムを明確な制限付きの概念実証として報告し、2 番目のシステムを最近のエピソード記憶と個人化された VLM の研究に対して位置づけます。

原文 (English)

Surprise as a Signal for Plasticity and Metacognition

We study a single idea across two settings: that a prediction-error signal, computed by a small predictor over the latent space of a frozen encoder, can serve both as a gate on plasticity and as a substrate for metacognition. In the first system, a non-parametric episodic memory writes a new concept only when this surprise is high, and a periodic offline replay phase consolidates recent traces into a slow linear readout. On a continual stream of 1000 ImageNet classes with a frozen DINOv2 or I-JEPA backbone, the consolidation phase recovers 17.7 points of retention on the oldest classes for DINOv2 and 51.3 points for I-JEPA (single-seed runs), and an ablation shows that replaying only a recent window is worse than no replay at all. In few-shot evaluation the same memory reaches 91.6% on 5-way 1-shot mini-ImageNet, above a task-specific baseline, while a harder 500-way regime exposes the true difficulty. In the second system, the same surprise signal, computed in a shared text-image space, modulates the behaviour of a vision-language model: it answers assertively when a concept is known, hedges when it is partially familiar, and refuses to identify the object and asks for an explanation when it is novel, learning the concept from a single user utterance. The external detector separates known from novel concepts at an AUROC of 0.966 (95% CI +/-0.024), far above the model's own verbalised confidence (0.618), while its token-level confidence sits below chance under greedy decoding; after a sleep phase that empties the fast store, the system recalls 99.2% of fifty taught facts from the consolidated store while a base model recovers none. We report both systems as proof-of-concept, with explicit limitations, and position the second against recent episodic-memory and personalised-VLM work.

13:00 JSTLLM/生成AIエージェント

エージェントのオーケストレーションとエージェントのオーケストレーションの設計と実装

Agentic Business Process Management は最近勢いが増しています。 AI エージェント、つまり主に LLM ベースのエージェントの自律性は、プロセス テクノロジと組み合わせることで、一定レベルの堅牢性、扱いやすさ、追跡可能性とのバランスをとることができるとの見通しがあります。このペーパーでは、タスクの特異性、追跡可能性と扱いやすさ、自律性と反応性、正確性保証などの特性に沿ったエージェント オーケストレーション オプションの分類フレームワークを提供し、さまざまなシナリオを実現するための定性的な決定基準を示します。また、実現プロパティの定量的評価のためのメトリクスも提供し、予測光センシング シナリオのさまざまなエージェント実装を通じてそれらを示します。全体として、この作業は、エージェントのオーケストレーションとエージェントのオーケストレーションの設計と実装のためのプロパティ、基準、メトリックを提供することを目的としています。

原文 (English)

Design and Implementation of Agentic Orchestrations and Orchestration of Agents

Agentic Business Process Management has gained momentum recently. The prospect is that the autonomy of AI agents, i.e., predominantly LLM-based agents, can be balanced with a certain level of robustness, tractability, and traceability through a combination with process technology. In this paper, we provide a classification framework for agentic orchestration options along properties such as task specificity, traceability and tractability, autonomy and reactivity, and correctness assurance and present qualitative decision criteria for realizations of different scenarios. We also provide metrics for the quantitative assessment of realization properties and show them through different agentic implementations of a predictive light sensing scenario. Altogether, this work aims at providing properties, criteria, and metrics for the design and implementation of agentic orchestrations and orchestration of agents.

13:00 JST研究/論文

深刻なクラス不均衡下での個人レベルの欠勤予測のための時系列分類フレームワーク

医療、緊急サービス、食肉加工、建設、宅配サービスや配達サービスなど、需要の高い作業環境では、スタッフの欠勤により多大な運営コストが発生します。こうした現場では、プロアクティブな人員計画が信頼できる個人レベルの欠勤予測に依存しています。既存の回帰および分類のアプローチには構造的な制限があります。彼らは、時刻 t に観察された特徴を同時刻 t のラベルにマッピングし、将来の出来事を予測するのではなく、すでに実現した結果を再現し、個人の出席履歴に固有の連続的な行動構造を破棄します。私たちは、過去の出席シーケンスを将来の欠勤ラベルから分離し、真にプロアクティブな予測を可能にする時系列分類 (TSC) フレームワークを提案します。公開されている長期的な参加者数データが​​不足しているため、UCI データセットに合わせて調整された再現可能なシミュレート データセットを構築します。不均衡率 $\rho$ のみを使用して、深刻なクラス不均衡下でのバイナリ焦点損失 (BFL) と幾何平均 (G-Mean) 損失を分析します。 BFL の場合、初期勾配比は $\rho\alpha/(1-\alpha)$ であり、平衡重み $\alpha = 1/(1+\rho) \約 0.023$ を意味します。実験によると、パフォーマンスは主に $\alpha$ によって支配され、BFL は G 平均に匹敵する特異度 0.813 とバランスのとれた精度 0.888 を達成しました。 BFL とは異なり、G-Mean はパラメーターの調整を行わずに自動的に適応します。評価された 3 つの深層学習アーキテクチャ、Long Short-Term Memory (LSTM)、Convolutional Neural Network (CNN)、ハイブリッド LSTM-Fully Convolutional Network (LSTM-FCN) の中で、LSTM-FCN は高い精度と特異性を実現します。安定したパフォーマンスは、バッチ サイズが 64 以上、ウィンドウ サイズが 40 ~ 80 日の場合に得られ、ホールドアウトされたテスト データで約 80% のバランスのとれた精度が得られます。

原文 (English)

A time-series classification framework for individual-level absenteeism prediction under severe class imbalance

Staff absenteeism imposes substantial operational costs in high-demand work environments such as healthcare, emergency services, meat processing, construction, and courier and delivery services, where proactive workforce planning depends on reliable individual-level absence prediction. Existing regression and classification approaches share a structural limitation; they map features observed at time t to labels at the same time t, reproducing already-realised outcomes rather than predicting future events, and discard the sequential behavioural structure inherent in individual attendance histories. We propose a Time Series Classification (TSC) framework that separates historical attendance sequences from future absence labels, enabling genuinely proactive prediction. Due to the lack of public longitudinal attendance data, we construct a reproducible simulated dataset calibrated to the UCI dataset. We analyse Binary Focal Loss (BFL) and Geometric Mean (G-Mean) loss under severe class imbalance using only the imbalance ratio $\rho$. For BFL, the initial gradient ratio is $\rho\alpha/(1-\alpha)$, implying the balanced weight $\alpha = 1/(1+\rho) \approx 0.023$. Experiments show that performance is governed mainly by $\alpha$, with BFL achieving specificity 0.813 and balanced accuracy 0.888, comparable to G-Mean. Unlike BFL, G-Mean adapts automatically without parameter calibration. Among three deep learning architectures evaluated, Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), and the hybrid LSTM-Fully Convolutional Network (LSTM-FCN), the LSTM-FCN delivers strong precision and specificity. Stable performance is obtained with batch sizes >= 64 and window sizes between 40-80 days, yielding balanced accuracy of approximately 80% on held-out test data.

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

ARC-AGI-2 の全体的なトレース判定によるモダリティ主導の検索

大規模な言語モデルは、確実に間違っているにもかかわらず、抽象的な推論タスクに対して流暢で内部的に一貫した推論トレースを生成する可能性があるため、世代だけでなく候補者の選択が中心的な課題となります。私は、ARC-AGI-2 のソルバーを紹介します。これは、数ショットの視覚的推論ベンチマークであり、次の 2 つの原則に基づいて構築されています。(i) 推論モダリティを検索演算子として扱い、テキスト、画像、コード チャネル全体で多様な候補を個別に生成する、(ii) コンテキストを保持した全体的判断。裁判官モデルが単一の長いコンテキスト プロンプト内のすべての候補推論トレースを共同で比較します。自己整合性や多数決とは異なり、このアプローチでは、様相応答が間違っているタスクに関して正しい少数派の仮説を確実に回復します。 ARC 賞のセミプライベート評価セットでは、ソルバーはタスクあたり 38.99 米ドルで 72.9 パーセントを達成しました。これは、この記事の執筆時点で検証済みのリーダーボードの最高スコアであり、最高のスタンドアロン フロンティア モデルである GPT-5.2 Pro の 54.2 パーセントと Gemini 3 Pro の 54.0 パーセントを +18.7 パーセント ポイント上回っています。公開評価セットでは、タスクあたり 19.69 米ドルで 76.1 パーセントを達成します。私は完全なソース コードを公開し、規範的なプロンプト テンプレートと反復的な改良によって体系的に仮説の多様性が減少し、パフォーマンスが低下するという発見を含む、広範な否定的な結果を文書化します。

原文 (English)

Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I present a solver for ARC-AGI-2, a few-shot visual reasoning benchmark, built around two principles: (i) treating reasoning modalities as search operators, generating diverse candidates independently across text, image, and code channels, and (ii) context-preserving holistic judging, in which a judge model jointly compares all candidate reasoning traces within a single long-context prompt. Unlike self-consistency or majority voting, this approach reliably recovers correct minority hypotheses on tasks where the modal answer is wrong. On the ARC Prize semi-private evaluation set, the solver achieves 72.9 percent at USD 38.99 per task - the highest score on the verified leaderboard at the time of writing, exceeding the best standalone frontier models, GPT-5.2 Pro at 54.2 percent and Gemini 3 Pro at 54.0 percent, by +18.7 percentage points. On the public evaluation set, it achieves 76.1 percent at USD 19.69 per task. I release the full source code and document extensive negative results, including the finding that prescriptive prompting templates and iterative refinement systematically reduce hypothesis diversity and degrade performance.

13:00 JSTエージェント

ACE: エージェント間でプラグイン可能なアダプティブ コンテキスト エラスティックライザー

エージェント タスクの複雑さの増大により、軌跡の長さが急速に増大しており、固定コンテキスト ウィンドウを備えた大規模言語モデル (LLM) ベースのエージェントにとって大きな課題となっています。切り捨てや要約などの既存のコンテキスト管理手法には、固有の柔軟性と不可逆性という問題があります。つまり、情報が一度破棄または圧縮されると、後の意思決定ステップで重要な意味を持つようになった場合でも、復元することができません。これらの制限に対処するために、私たちは、各意思決定ステップで履歴ステップ情報をエージェントのコンテキストに柔軟に調整するプラグアンドプレイ モジュールである Adaptive Context Elasticizer (ACE) を提案します。 ACE は、各履歴ステップの生のメッセージと圧縮された抽象化の両方を保存するロスレス メッセージ メンテナンス レイヤーを維持します。一方、コンテキスト オーケストレーション レイヤーは、現在のタスクの状態に基づくすべての意思決定ステップで、各ステップに柔軟なタイプを生、抽象、またはドロップとして適応的に割り当てます。この可逆的な設計により、メイン LLM は常にコンパクトでありながら情報が豊富なコンテキストを受け取ることが保証されます。当社は、トレーニングやアーキテクチャの変更を行わずに、ReAct、DeepAgent、WebThinker、MiroFlow を含む 4 つの多様なエージェント フレームワークに ACE を適応させます。実験の結果、ACE は常に切り捨ておよび要約のベースラインを上回り、4 つのエージェント フレームワークすべてにわたって一貫したパフォーマンスの向上をもたらすことが示されています。

原文 (English)

ACE: Pluggable Adaptive Context Elasticizer across Agents

The increasing complexity of agentic tasks has led to rapidly growing trajectory lengths, which poses significant challenges for large language model (LLM) based agents with fixed context windows. Existing context management techniques, such as truncation and summarization, suffer from inherent inflexibility and irreversibility: once information is discarded or compressed, it cannot be recovered even when it becomes critically relevant in later decision steps. To address these limitations, we propose the Adaptive Context Elasticizer (ACE), a plug-and-play module that elastically orchestrates historical step information into the agent's context at each decision step. ACE maintains a lossless message maintenance layer that stores both raw messages and compressed abstractions for each historical step, while a context orchestration layer adaptively assigns each step an elastic type as raw, abstract, or drop, at every decision step based on the current task state. This reversible design ensures that the main LLM always receives a compact yet information-rich context. We adapt ACE to four diverse agent frameworks, including ReAct, DeepAgent, WebThinker, and MiroFlow, without training or architectural modifications. Experiments show that ACE consistently outperforms truncation and summarization baselines, and brings consistent performance gains across all four agent frameworks.

13:00 JSTLLM/生成AI

どのトークンが重要ですか? Relative Surprisal Index を使用した RLVR の適応型トークン選択

強化学習 (RL) は、大規模言語モデル (LLM) を模倣ベースのトレーニングを超えて、より堅牢な推論能力に向けて推進するための強力なツールとなっています。既存のアプローチの中で、検証可能な報酬を伴う RL (RLVR) は、LLM 推論を進歩させるための極めて重要なパラダイムとして浮上しています。実証的な成功にもかかわらず、最近の研究では異なる洞察が得られています。ある調査では、トレーニング中に高エントロピーのトークンの位置を優先することを主張していますが、別の観点では、確率の低いトークンが勾配更新を支配することを許可しないように警告しています。特に、高エントロピーのトークンは通常低い確率で相関しますが、両方のパラダイムは経験的に大幅なパフォーマンスの向上をもたらします。この研究では、サンプリングされたトークンの確率またはエントロピーを単独で評価するだけでは、ポリシー最適化のダイナミクスを捉えるには不十分であると主張します。この緊張を解決するために、トークンのエントロピーと選択されたトークンの確率を自然に結び付ける原則に基づいた情報理論的指標である Relative Surprisal Index (RSI) を導入します。穏やかな条件下では、RSI が選択されたロジット摂動の下でのロジット勾配ノルムの一次変分と予測エントロピーの間の局所比に関連していることを示します。 RSI に基づいて、安定した RSI 間隔内にトークンを保持するエントロピー適応トークン フィルタリング手法である RSI セレクション (RSI-S) を提案します。 RSI-S は、これまでの矛盾したパラダイムをうまく調和させ、冗長な意外性の低いトークンと不安定な驚き性の高いテール トークンの両方を除去します。経験的評価によると、RSI-S は、AIME および AMC ベンチマークでのさまざまなモデル スケール (Qwen2.5 ~ 1.5B、3B、および 7B) にわたって、より高い avg@32 精度を達成しています。RSI-S は、GRPO よりも avg@32 精度を 2 ~ 3 パーセント向上させます。全体として、RSI は RLVR 改善に有望な展望を提供します。

原文 (English)

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index

Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities. Among existing approaches, RL with Verifiable Rewards (RLVR) has emerged as a pivotal paradigm for advancing LLM reasoning. Despite its empirical success, recent studies have offered different insights. One line of inquiry advocates prioritizing high-entropy token positions during training, while another perspective cautions against allowing low-probability tokens to dominate gradient updates. Notably, although high-entropy tokens are usually correlated with low probability, both paradigms empirically yield substantial performance gains. In this work, we argue that evaluating sampled-token probability or entropy in isolation is insufficient to capture the policy optimization dynamics. To resolve this tension, we introduce the Relative Surprisal Index (RSI), a principled, information-theoretic metric that naturally couples the token's entropy with the probability of the selected token. We show that, under mild conditions, RSI is related to the local ratio between the first-order variations of the logit-gradient norm and predictive entropy under a selected-logit perturbation. Building on RSI, we propose RSI Selection (RSI-S), an entropy-adaptive token filtering method that retains tokens within a stable RSI interval. RSI-S successfully reconciles previous contradictory paradigms and filters out both redundant low-surprisal tokens and unstable high-surprisal tail tokens. Empirical evaluations show that RSI-S achieves higher avg@32 accuracy across different model scales (Qwen2.5-1.5B, 3B, and 7B) on AIME and AMC benchmarks: RSI-S improves avg@32 accuracy by 2--3 percentage points over GRPO. Overall, RSI offers a promising perspective for RLVR improvement.

13:00 JST研究/論文

健康科学における科学的説明: 因果関係、信頼、認識論的妥当性

医療用人工知能 (AI) は臨床現場を変革すると広く期待されていますが、多くの機械学習 (ML) モデルの意思決定プロセスは依然として不透明です。説明可能性は、特に一か八かの状況において、AI が予測を生成する理由を明らかにするための部分的な救済策として進歩してきました。継続的な努力にもかかわらず、適切な医学的説明とは何かについての議論は未解決のままです。しかし、説明は長い間、科学と医学の哲学における探求の中心的なテーマでした。しかし、これらの分野で開発された洞察は、現代の説明可能な AI (XAI) 研究ではほとんど無視されており、その基本的な前提は十分に検討されていません。このギャップに対処するために、この論文では科学哲学と XAI の交差点における批判的なレビューを展開します。それは、健康科学において説明とみなされるものについての一般的な説明を調査し、医学における XAI に情報を提供するための適切性を評価し、それらがこの領域における説明可能性への哲学に基づいたアプローチに必要な条件を提供すると主張しています。この基礎的な哲学的文献に基づいて、この議論は 3 つの中心的な分析軸を特定します。それは、医学的推論における因果関係の役割、医療の信頼の認識論的および関係的側面、そして多様な利害関係者の実際的なニーズによって形成される説明の適切性の基準です。この論文では、哲学的分析と医療 AI の現在の開発を統合することにより、認識論的に堅牢であるだけでなく、臨床意思決定の認識論的および実践的要件に沿った説明を提供する XAI システムを設計するための原則を概説し、医療 XAI における進行中の議論を、未踏の概念基盤に向けて形成しています。

原文 (English)

Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy

Medical Artificial Intelligence (AI) is widely expected to transform clinical practice, yet the decision-making processes of many Machine Learning (ML) models remain opaque. Explainability has been advanced as a partial remedy to clarify why AI generates predictions, particularly in high-stakes contexts. Despite ongoing efforts, debates on what constitutes an adequate medical explanation remain unsettled. Yet, explanation has long been a central topic of inquiry in the philosophy of science and medicine. The insights developed in these fields, however, have been largely overlooked in contemporary explainable AI (XAI) research, leaving its foundational assumptions insufficiently examined. To address this gap, this paper develops a critical review at the intersection of philosophy of science and XAI. It examines prevailing accounts of what counts as an explanation in the health sciences and assesses their adequacy for informing XAI in medicine, arguing that they provide necessary conditions for a philosophically grounded approach to explainability in this domain. Building on this foundational philosophical literature, the discussion identifies three central axes of analysis: the role of causality in medical reasoning, the epistemic and relational dimensions of medical trust, and the criteria of explanatory adequacy as shaped by the pragmatic needs of diverse stakeholders. By integrating philosophical analysis with current developments in medical AI, the paper outlines principles for designing XAI systems that offer explanations that are not only epistemically robust but also aligned with the epistemic and practical requirements of clinical decision-making, shaping ongoing debates in medical XAI toward underexplored conceptual foundations.

13:00 JSTエージェント

英語で考え、韓国語で答える: 多言語ツールを使用するエージェントの効率的な適応

我々は、CohereとLG CNSの協力により、実際的な記憶力とサービス提供の制約の下で韓国語と英語の企業エージェント向けに開発された111Bパラメータのハイブリッド推論モデルであるLuckyStar 111Bを紹介します。このモデルは、新しい事前トレーニングの実行ではなく、Cohere の完全にポストトレーニングされたコマンド A モデルからトレーニングされ、プリアンブル条件付けを使用して、簡潔な非推論動作と長いツール指向の推論を切り替えます。私たちは、ツールを使用するエージェントを効率的にスケーリングするための 4 つの選択肢を研究します。それは、多言語教師あり微調整、複数ステップのツール使用タスクに対する検証可能な報酬を伴う強化学習、韓国語のユーザー対応応答に対する言語一貫性報酬、単一 GPU サービス用の 4 ビット量子化です。適応されたモデルにより、一般的な韓国語と英語の命令追従品質を維持しながら、数学的推論、関数呼び出し、およびエージェントによる自然言語から SQL への変換 (NL2SQL) のパフォーマンスが向上します。これらの結果は、メモリ制約のある展開下でトレーニング後の多言語モデルを検証可能なエージェント ワークフローに適応させるための実用的なレシピと障害モード分析を提供します。

原文 (English)

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterprise agents under practical memory and serving constraints. The model trains from Cohere's fully post-trained Command A model rather than a new pretraining run, and uses preamble conditioning to switch between concise non-reasoning behavior and longer tool-oriented reasoning. We study four choices for scaling tool-using agents efficiently: multilingual supervised fine-tuning, reinforcement learning with verifiable rewards for multi-step tool-use tasks, language-consistency rewards for Korean user-facing responses, and 4-bit quantization for single-GPU serving. The adapted model improves mathematical reasoning, function calling, and agentic natural-language-to-SQL (NL2SQL) performance while preserving general Korean and English instruction-following quality. These results provide a practical recipe and failure-mode analysis for adapting post-trained multilingual models to verifiable agentic workflows under memory-constrained deployment.

13:00 JSTエージェント研究/論文

FARS: 大規模に導入された完全に自動化された研究システム

最近の自動研究システムは、言語モデル エージェントが仮説を生成し、実験を実行し、完全な原稿を書くことができることを示していますが、ほとんどの証拠は依然として、選択された例、人間が組み立てたトピック、またはいくつかの事前定義された研究タスクから得られます。私たちは、研究テーマ全体にわたって大規模に動作するように設計された完全に自動化された AI 対 AI 研究システムである FARS (Fully Automated Research System) を紹介します。 FARS は、提案、コード、ログ、結果、原稿を記録する共有ワークスペースを通じて調整された段階固有のエージェントを使用して、アイデア出し、計画、実験、執筆を通じてプロジェクトを自律的に生成し、推進します。最初の公開展開で、FARS は 67 のきめ細かい AI/ML トピックにまたがる 166 の完全な研究論文を作成しましたが、中間成果物は厳選された成功セットではなく、監査可能なコーパスとして保存されました。このコーパスは、全体の評価、サブスコア、完全性チェック、LLM 使用の開示を含む、140 件の論文をカバーするボランティアの査読者による 282 件の構造化レビューによって評価されます。レビューでは、FARS が大規模な公共展開においてレビューに値する、場合によっては強力な AI/ML 研究成果物を生成できる一方で、狭い実験範囲、方法論的な制限、完全性の問題で繰り返される失敗モードも明らかにすることが示されています。

原文 (English)

FARS: A Fully Automated Research System Deployed at Scale

Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale. FARS autonomously generates and advances projects through ideation, planning, experimentation, and writing, using stage-specific agents coordinated through a shared workspace that records proposals, code, logs, results, and manuscripts. In its first public deployment, FARS produced 166 complete research papers spanning 67 fine-grained AI/ML topics while preserving intermediate artifacts as an auditable corpus rather than a curated set of successes. We evaluate this corpus with 282 structured reviews from volunteer reviewers covering 140 papers, including overall ratings, sub-scores, integrity checks, and LLM-use disclosure. The reviews indicate that FARS can produce review-worthy and occasionally strong AI/ML research artifacts in a large-scale public deployment, while also exposing recurring failure modes in narrow experimental scope, methodological limitations, and integrity issues.

13:00 JSTLLM/生成AI研究/論文

Arena-T2I ハード: 依存関係を意識したチェックリストによるベンチマークと忠実性の向上

忠実度、つまり生成された画像がそのプロンプトとどれだけ正確に一致しているかが、テキストから画像への (T2I) モデルの実世界の有用性の中心となってきています。しかし、既存の忠実度ベンチマークは単純なアトミック命令に依存しており、最上位のシステムはすでにほぼ完璧なスコアを達成しています。 T2I モデルがクリエイティブなワークフローに入ると、ユーザーは複雑な空間関係、スタイル上の制約、複雑なテキスト レンダリングを組み合わせた多面的なリクエストを発行します。この設定では、単一のバイナリ VLM 判定スコアでは、モデルが満たしていない特定の制約が捕捉されなくなります。実際のアリーナ T2I ログから抽出された 310 プロンプトのストレス ベンチマークである Arena-T2I Hard を導入します。プロンプトごとに、テキスト レンダリングを含む 6 つのカテゴリにまたがる約 30 の分解された Yes/No 制約があります。私たちが評価した最も強力なクローズドソース システムは、11 のシステム間で 33 ~ pp のパフォーマンス差があり、0.855 に達し、かなりの識別力を示しています。さらに、公共の場での高いランキングでは忠実度を予測することはできず、総合的なブラッドリー・テリー(BT)選好スコアでは、きめの細かい即時の遵守よりも美しさを優先していることが確認されています。私たちは、各プロンプトをはい/いいえの質問の DAG に分解し、失敗した親の子孫をゼロにして、忠実さを制約ごとのトレーニング信号に変える、依存関係を意識したチェックリスト報酬を提案します。グループ分離正規化 (GDPO) による BT の美的報酬と組み合わせると、ロールアウト グループ内の各報酬がどちらも崩壊しないように標準化され、このレシピは、MMRB2 ペアごとの比較に基づく SD3.5-Medium および FLUX.1-dev で、単一報酬、単純加重和、または 4 報酬の BT アンサンブル ベースラインよりも厳密に優れた忠実性と美学のトレードオフを達成します。

原文 (English)

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, on which top-tier systems already achieve near-perfect scores. As T2I models enter creative workflows, users issue multi-faceted requests combining intricate spatial relationships, stylistic constraints, and complex text rendering. In this setting, a single binary VLM-judge score no longer captures which specific constraints the model fails to satisfy. We introduce Arena-T2I Hard, a 310-prompt stress benchmark drawn from real arena T2I logs, with approximately 30 decomposed yes/no constraints per prompt spanning six categories, including text rendering. The strongest closed-source system we evaluate reaches 0.855 with a 33~pp performance gap across 11 systems, demonstrating substantial discriminative power. Moreover, high public-arena rankings fail to predict faithfulness, confirming that holistic Bradley-Terry (BT) preference scores prioritize aesthetics over fine-grained prompt adherence. We propose a dependency-aware checklist reward that decomposes each prompt into a DAG of yes/no questions and zeroes descendants of failed parents, turning faithfulness into a per-constraint training signal. Combined with a BT aesthetic reward via group-decoupled normalization (GDPO), which standardizes each reward within its rollout group so neither collapses, the recipe attains a strictly better faithfulness-aesthetics trade-off on SD3.5-Medium and FLUX.1-dev under MMRB2 pairwise comparisons than every single-reward, naive weighted-sum, or 4-reward BT-ensemble baseline.

13:00 JSTエージェント

生物学的プロトコルの自動生成と実行のための自己進化型エージェント システム

自律的なウェットラボ実験には、もっともらしいプロトコールのテキスト以上のものが必要です。つまり、生物学的意図、定量的手順、デバイスの制約、および実験のフィードバックが、プロトコールや SOP の設計からコードや物理的な実行に至るまで常に調整されていなければなりません。私たちは、この変換を実験的な自動化問題としてテストするための、専門家に基づいたベンチマークおよび評価フレームワークとともに、自己進化するマルチエージェント システムである ProtoPilot を開発しました。このフレームワークは、98 のゴールドスタンダード プロトコル、ウェットラボのエキスパート ルーブリック、デバイス レベルの妥当性ゲート、および実際の実験テストから派生した 294 の合成生物学および分子生物学のタスクに及びます。 ProtoPilot には、レイヤーごとの検証機能、マルチエージェント オーケストレーション、ランタイムで更新されるスキル ライブラリが組み込まれており、プロトコルの生成、SOP の拡張、SDK 準拠のコードの合成、ウェット ラボのフィードバックからのワークフローの修正を行います。 OpenTrons-AI の 32.35% と比較して、Top@3 専門家選択率 90.2%、全体のプロトコルからコードへのゲート通過率 89.5%、Opentrons 通過率 88.24% を達成しました。ウェットラボ検証により、解釈可能な読み取り値、サンガーで確認された生成物、フィードバック補正された PCA で構築された DNA ターゲットが生成され、自律実験への検証可能なルートが確立されました。これらの結果を総合すると、評価フレームワークが自律型ウェットラボ自動化の実行関連要件を捉えており、ProtoPilot がプロトコルとコード生成を検証済みの実行とフィードバックに基づく改訂に変換することで要件を満たすことができることを示しています。

原文 (English)

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution. We developed ProtoPilot, a self-evolving multi-agent system, together with an expert-grounded benchmark and evaluation framework for testing this conversion as an experimental automation problem. The framework spans 294 synthetic-biology and molecular-biology tasks derived from 98 gold-standard protocols, wet-lab expert rubrics, device-level validity gates and real experimental tests. ProtoPilot incorporates layer-wise verifiability, multi-agent orchestration and a runtime-updated skill library to generate protocols, expand SOPs, synthesize SDK-compliant code and revise workflows from wet-lab feedback. It achieved a Top@3 expert-preference rate of 90.2%, an overall protocol-to-code gate pass rate of 89.5% and an Opentrons pass rate of 88.24%, compared with 32.35% for OpenTrons-AI. Wet-lab validation produced interpretable readouts, Sanger-confirmed products and feedback-corrected PCA-assembled DNA targets, establishing a verifiable route to autonomous experimentation. Together, these results show that the evaluation framework captures execution-relevant requirements for autonomous wet-lab automation, and that ProtoPilot can meet them by converting protocol and code generation into validated execution and feedback-guided revision.

13:00 JSTLLM/生成AI

Evo-PI: 進化する原則に基づいた監督による医学的推論の調整

最近の進歩にも関わらず、大規模マルチモーダル言語モデル (MLLM) の推論機能は依然として静的監視によって根本的に制約されており、固定のプロンプト、ルール、または報酬モデルがトレーニング全体を通じて非適応的なガイダンスを提供します。このような静的信号は、多くの場合、出力形式を強制するには十分ですが、基礎となる推論プロセスを形成できず、複雑な意思決定タスクにおける脆弱な一般化とパフォーマンスの飽和につながります。私たちは、推論原則を生成、評価、反復展開できる明示的な言語ベースの監視信号として扱う原則中心の学習フレームワークである Evo-PI を提案します。 Evo-PI は、固定報酬に依存する代わりに、原則がモデル推論を導き、モデルの動作がモデルを監視する原則を洗練させるという共進化ループを可能にします。この動的な調整メカニズムにより、監視はモデルの推論の欠陥に徐々に適応することができます。私たちは、構造化されたビジュアルテキスト推論を必要とする一か八かのテストベッドとして、医療視覚的質問応答で Evo-PI をインスタンス化します。 8 つのベンチマークと複数のモデル バックボーンにわたって、Evo-PI は推論の精度を一貫して向上させ、最大 24.6% の向上を達成します。私たちの結果は、進化する原則に基づいた監督が、MLLM で専門家に合わせた推論をトレーニングするための拡張性のある一般的なパラダイムを提供することを示唆しています。コードは https://github.com/zhengxianda/Evo_PI で入手できます。

原文 (English)

Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision

Despite recent progress, the reasoning capabilities of large multimodal language models (MLLMs) remain fundamentally constrained by static supervision, where fixed prompts, rules, or reward models provide non-adaptive guidance throughout training. Such static signals are often sufficient to enforce output formats, but fail to shape the underlying reasoning process, leading to brittle generalization and performance saturation in complex decision-making tasks. We propose Evo-PI, a principle-centric learning framework that treats reasoning principles as explicit, language-based supervision signals that can be generated, evaluated, and iteratively evolved. Instead of relying on fixed rewards, Evo-PI enables a co-evolutionary loop in which principles guide model reasoning, while model behaviors in turn refine the principles that supervise them. This dynamic alignment mechanism allows supervision to progressively adapt to the model's reasoning deficiencies. We instantiate Evo-PI in medical visual question answering as a high-stakes testbed requiring structured visual-textual reasoning. Across eight benchmarks and multiple model backbones, Evo-PI consistently improves reasoning accuracy, achieving gains of up to 24.6%. Our results suggest that evolving principle-guided supervision offers a scalable and general paradigm for training expert-aligned reasoning in MLLMs. Code is available at https://github.com/zhengxianda/Evo_PI.

13:00 JSTLLM/生成AIビジネス/資金調達

RAISE: 堅牢な敵対的インスタンス検索を備えた LLM ベースの自動ヒューリスティック設計

大規模言語モデル (LLM) を使用した自動ヒューリスティック設計 (AHD) は、高品質のヒューリスティックの発見において目覚ましい進歩を示しています。ただし、既存の LLM ベースの AHD 手法は、固定されたトレーニング インスタンス セットのヒューリスティックを最適化するため、現実世界の分布シフトの下で展開すると壊滅的に失敗する可能性があります。我々は、トレーニング分布の原則的な近傍内での制約付きワーストケース インスタンス検索を LLM ベースの進化的検索ループに統合するフレームワークである、Robust Adversary Instance Search (RAISE) を提案します。 RAISE は、堅牢な AHD を制約付きの敵対的インスタンス検索問題として扱います。外側のループは LLM 演算子を介してヒューリスティックを進化させますが、LLM を使用しない内側のループは、境界射影を伴う基底分布パラメーター化を使用して、トレーニング インスタンス セットの周りのイプシロン ボール内のハード インスタンスを効率的に識別します。 5 つのディストリビューション ファミリにわたるオンライン ビン パッキング (OBP)、オンライン ジョブ ショップ スケジューリング (OJSP)、およびオンライン車両ルーティング (OVRP) に関する包括的な実験では、既存の LLM ベースの AHD 手法はディストリビューションの移行時に最大 19 倍劣化する一方、RAISE はテストされたすべてのディストリビューションと問題規模にわたって一貫して強力なパフォーマンスを維持することを実証しました。

原文 (English)

RAISE: LLM-based Automated Heuristic Design with Robust Adversary Instance Search

Automated Heuristic Design (AHD) with Large Language Models (LLMs) has shown remarkable progress in discovering high-quality heuristics. However, existing LLM-based AHD methods optimize heuristics for a fixed training instance set and may fail catastrophically when deployed under real-world distributional shifts. We propose Robust Adversary Instance Search (RAISE), a framework that integrates constrained worst-case instance search within a principled neighborhood of the training distribution into the LLM-based evolutionary search loop. RAISE treats robust AHD as a constrained adversarial instance search problem: the outer loop evolves heuristics via LLM operators, while an LLM-free inner loop efficiently identifies hard instances within an epsilon-ball around the training instance set using a basis distribution parameterization with boundary projection. Comprehensive experiments on Online Bin Packing (OBP), Online Job Shop Scheduling (OJSP), and Online Vehicle Routing (OVRP) across five distribution families demonstrate that existing LLM-based AHD methods degrade by up to 19 times under distribution shift, while RAISE consistently maintains strong performance across all tested distributions and problem scales

13:00 JST研究/論文

大規模なデータベースには小規模でオープンウェイトの言語モデルが必要

独自の API を中心に構築された言語モデル システムは、多くの場合、トークンベースのコスト モデルで動作します。これは、大規模なデータベースのコンテキストでは法外に高価になり、LM で強化されたリレーショナル演算子は 1 セットの実験で 10,000 ドルを超えるコストが発生する可能性があり、徹底した研究と実際の展開が妨げられます。この論文では、わずか 16 GB の VRAM でローカルに実行される量子化されたオープンウェイト モデルが、より低いレイテンシーと数分の 1 の価格で、クローズド ソースの同等モデルと同等またはそれを上回る精度を実現できることを実証し、効果的な LM データベースの統合にはクローズド ソースの LM API が必要であるという一般的な前提に疑問を投げかけます。これらのオープンウェイト モデルを LM-DB システム内に効率的に展開するために必要な主要なシステム最適化を提示し、分析します。これらのローカル モデルを BlendSQL v0.1.0 フレームワークに統合することにより、独自の LM API と比較して全体のコストが 390 倍削減され、レイテンシが 3.8 倍削減されることが実証されました。コードは https://github.com/CapitalOne-Research/play-by-the-type-rules/tree/main/sembench で公開しています。

原文 (English)

Large Databases Need Small, Open-Weight Language Models

Language model systems built around proprietary APIs often operate on a token-based cost model. This becomes prohibitively expensive in the context of large databases, where LM-enhanced relational operators can incur costs exceeding $10,000 for a single set of experiments, hindering thorough research and practical deployment. In this paper, we demonstrate that quantized, open-weight models running locally on just 16GB of VRAM can match or exceed the accuracy of closed-source counterparts at lower latency and a fraction of the price, challenging the prevailing assumption that closed-source LM APIs are necessary for effective LM-database integration. We present and analyze the key system optimizations required to efficiently deploy these open-weight models within an LM-DB system. By integrating these local models into the BlendSQL v0.1.0 framework, we demonstrate a 390x reduction in overall costs and 3.8x reduction in latency compared to a proprietary LM API. We make our code available at https://github.com/CapitalOne-Research/play-by-the-type-rules/tree/main/sembench.

13:00 JST研究/論文

インテリジェンスの作成: AGI の計算基盤

この研究では、集合論と超次元コンピューティングに基づいた心の新しい計算理論を紹介します。従来のニューラル ネットワークは連続重みと行列の乗算に依存していますが、このフレームワークは疎なバイナリ データで動作します。これは情報を離散セットとして表現し、生物学的神経集団コードを直接モデル化します。私は、連想メモリが、組み合わせ的に拡張された隠れ層を特徴とするネットワーク トポロジから自然に出現することを実証します。学習は、スカラー重みの調整ではなく、トポロジカルな可塑性によって駆動されます。このアーキテクチャは、単一のコア アルゴリズム (サブセット パターン マッチングと正確な最近傍検索による情報取得) の下で自動連想学習とヘテロ連想学習を統合します。一定時間の複雑さで動作するこれらのメカニズムは、継続的なボトルネックを発生させることなく、知覚データ (疎な分散表現) とシンボル (疎なホログラフィック表現) の橋渡しをします。このフレームワークを神経解剖学にマッピングして、小脳と新皮質の両方がこのアルゴリズムの変形を実装し、サブセットのパターンマッチングを認知の基本エンジンにすることを提案します。このアルゴリズムは行列演算ではなく離散ロジックに依存しているため、メモリ内ハードウェアに直接変換されます。これにより、人間レベルのエネルギー効率を備えた合成知能への新たな道が開かれます。

原文 (English)

Creating Intelligence: A Computational Foundation for AGI

This work introduces a new computational theory of mind grounded in set theory and hyperdimensional computing. Whereas traditional neural networks rely on continuous weights and matrix multiplication, this framework works with sparse binary data. It represents information as discrete sets, directly modeling biological neural population codes. I demonstrate that associative memory emerges naturally from network topologies featuring a combinatorially expanded hidden layer. Learning is driven by topological plasticity rather than scalar weight adjustments. This architecture unifies auto-associative and hetero-associative learning under a single core algorithm: information retrieval via subset pattern matching and exact nearest-neighbor search. Operating with constant-time complexity, these mechanisms bridge perceptual data (sparse distributed representations) and symbols (sparse holographic representations) without continuous bottlenecks. Mapping this framework to neuroanatomy, I propose that both the cerebellum and the neocortex implement variants of this algorithm, making subset pattern matching the fundamental engine of cognition. Because it relies on discrete logic rather than matrix arithmetic, this algorithm translates directly into in-memory hardware. This opens a new route toward synthetic intelligence with human-level energy efficiency.

13:00 JST研究/論文

産業規模の車両ルーティングのための適応クラスター第一ルート第二分解

大規模容量車両経路指定問題 (CVRP) は、経路指定インスタンスをより小さな計算上扱いやすい部分問題に分割するクラスターファースト経路セカンド (CFRS) アプローチを使用して対処するのが一般的です。既存の分割方法は通常、固定パーティショニング ルール、事前定義された最適化目標、または学習されたポリシーに依存しており、異なる空間特性、需要特性、運用特性を示すインスタンス間で一貫性のないパフォーマンスが発生する可能性があります。この研究では、反復的な意思決定プロセスとして分解手順を定式化する適応 CFRS システムを提案します。推論とツール選択における大規模言語モデル (LLM) の最近の成功を動機として、このシステムは、進化する分解状態を分析し、さらなるクラスタリング、バランシング、および洗練演算子を選択的に適用する高レベルの意思決定者として LLM を採用しています。提案されたアルゴリズムは、顧客と車両を共同で分割し、各問題の特性に分割の決定を適応させながら容量を意識したクラスタリングを可能にします。最大 500,000 人の顧客を含む合成およびベンチマークから派生した CVRP インスタンスに対するアプローチを評価します。実験結果は、ベンチマーク規模のインスタンスで競争力のあるパフォーマンスを示し、さらに大幅に大きな問題ではスケーラビリティの向上と堅牢なルーティング品質を示しました。これらの結果は、産業規模の車両ルート設定や大規模物流計画に対する実用的なアプローチとして、LLM に基づいた適応型意思決定支援の可能性を浮き彫りにしています。

原文 (English)

Adaptive Cluster-First Route-Second Decomposition for Industrial-Scale Vehicle Routing

Large-scale capacitated vehicle routing problems (CVRPs) are commonly addressed using cluster-first route-second (CFRS) approaches that split a routing instance into smaller, computationally tractable subproblems. Existing splitting methods typically rely on fixed partitioning rules, predefined optimization objectives, or learned policies, which may perform inconsistently across instances exhibiting different spatial, demand, and operational characteristics. In this work, we propose an adaptive CFRS system that formulates a decomposition procedure as an iterative decision-making process. Motivated by the recent success of large language models (LLMs) in reasoning and tool selection, the system employs an LLM as a high-level decision maker that analyzes the evolving decomposition state and selectively applies further clustering, balancing, and refinement operators. The proposed algorithm jointly partitions customers and vehicles, enabling capacity-aware clustering while adapting partitioning decisions to the characteristics of each problem. We evaluate the approach on synthetic and benchmark-derived CVRP instances containing up to 500,000 customers. Experimental results demonstrate competitive performance on benchmark-scale instances while exhibiting improved scalability and robust routing quality on substantially larger problems. These results highlight the potential of adaptive, LLM-guided decision support as a practical approach for industrial-scale vehicle routing and large-scale logistics planning.

13:00 JSTエージェント

植物の表現型解析における科学的発見を加速するエージェント的 AI フレームワーク

ハイスループットの植物表現型解析により、科学者が分析するよりもはるかに速く画像由来のデータセットが生成されるようになりました。オークリッジ国立研究所の高度植物表現型研究室 (APPL) では、自動ステーションが複数のリモート センシング モダリティを使用して毎日数百の植物を画像化しています。しかし、形質の抽出と解釈は依然として手作業で専門家に依存しており、厳密に事後的なものであるため、発見に対する拘束力は取得ではなく分析にあります。私たちは、施設をデータ ファクトリーから対話型の自律的な発見プラットフォームに変えるエンドツーエンドのエージェント AI フレームワークを紹介します。そこでは、科学者が AI エージェントと提携して洞察を得るまでの時間を短縮します。会話型の共同科学者エージェントが科学者の自然言語の質問を構造化された分析計画に変換し、ヘッドレスのコンピューティング エージェントが Vision Transformer のセグメンテーションと特性抽出を Frontier エクサスケール スーパーコンピューター上で実行します。 2 つのエージェントは別々のセキュリティ ドメインとリソース ドメインで実行され、安全なトークン認証されたストリーミング チャネル経由で通信します。これは、フェデレーション、データ移動、およびクラウド ネイティブ エージェント フレームワークが無視する来歴の現実を考慮した設計で、すべてのインタラクションでエンドツーエンドの出自が確実に取得されるようにします。このフレームワークは、数日から数週間にわたる分析プロセスを対話型のループに変え、エージェントが結果を検討し、次の分析を推奨し、フォローアップの質問に数秒で回答します。

原文 (English)

An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping

High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them. At Oak Ridge National Laboratory's Advanced Plant Phenotyping Laboratory (APPL), automated stations image hundreds of plants daily across multiple remote sensing modalities; yet, trait extraction and interpretation remain manual, expert-bound, and strictly post-hoc, making analysis, not acquisition, the binding constraint on discovery. We present an end-to-end agentic AI framework that turns the facility from a data factory into an interactive autonomous, discovery platform, where scientists partner with AI agents to accelerate time to insight. A conversational Co-Scientist Agent translates a scientist's natural-language question into a structured analysis plan, and a headless Compute Agent dispatches Vision Transformer segmentation and trait extraction on the Frontier exascale supercomputer. The two agents run in separate security and resource domains and communicate over a secure, token-authenticated streaming channel, a design that accounts for the federation, data-movement, and provenance realities cloud-native agentic frameworks ignore, ensuring end-to-end provenance is captured for every interaction. The framework turns a days- to weeks-long analysis process into an interactive loop where agents reason over results, recommend next analyses, and respond to follow-up questions in seconds.

13:00 JSTLLM/生成AI画像/動画生成

マルチモーダルの安全性のためにテキストによる拒否指示を活用する

大規模言語モデル (LLM) の安全性を向上させるために、トレーニング後の調整を実行するか、アクティベーション スペースでの拒否指示を利用できます。どちらの戦略もマルチモーダル LLM (MLLM) では実現可能性が低くなります。安全でないマルチモーダル データが必要であり、ユニモーダルの対応する戦略よりも収集が難しいからです。この研究では、この制約を緩和し、LLM バックボーンから直接抽出されたテキストの拒否指示がモダリティ (つまり、画像、ビデオ) 全体で一般化されるかどうかを調査します。予備的な調査結果はこの能力を裏付けていますが、有効性はレイヤーの選択、ステアリング強度、およびクロスモーダルアライメントによって条件付けされ、後者により安全なマルチモーダル入力が誤って拒否に向けて誘導されます。これに基づいて、マルチモーダル安全性データを必要とせずにマルチモーダル安全性を導入する、トレーニング不要の軽量アプローチである Modality-Agnostic Refusal Steering (MARS) を導入します。 MARS は、アクティベーションの再センタリングによってモダリティの不整合を修正し、幾何学的に定義された信頼領域内でステアリング強度を適応的にスケールし、最初に生成されたトークンで動作する最適な介入層を選択します。安全性、ユーティリティ、ビデオ ジェイルブレイク ベンチマークにわたる 5 つの SOTA MLLM で評価された MARS は、ユーティリティを維持しながら一貫した安全性の向上を実現します。これらの結果は、安全関連の構造がモダリティ間で共有されていること、およびテキストによる拒否の指示が、マルチモーダル調整のための強力だがまだ研究されていない基盤であることを明らかにしている。

原文 (English)

Harnessing Textual Refusal Directions for Multimodal Safety

To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize across modalities (i.e., image, video). Preliminary findings confirm this ability, though effectiveness is conditioned by layer selection, steering strength, and cross-modal alignment, with the latter causing safe multimodal inputs to be spuriously steered toward refusal. Building on this, we introduce Modality-Agnostic Refusal Steering (MARS), a light-weight training-free approach that injects multimodal safety without the need for multimodal safety data. MARS corrects modality misalignment via activation re-centering, adaptively scales steering strength within a geometrically defined trust region, and selects the optimal intervention layer, operating at the first generated token. Evaluated on five SOTA MLLMs across safety, utility, and video jailbreak benchmarks, MARS achieves consistent safety gains while preserving utility. These results reveal that safety-relevant structure is shared across modalities and that textual refusal directions are a powerful and underexplored foundation for multimodal alignment.

13:00 JSTエージェント

TreeAgent: コンパイルされたエキスパート ルールとビジョン言語モデルを介して林業における自動バイアス ラベリングのための一般化可能なマルチエージェント フレームワーク

多くの専門家主導の領域ではアノテーター間でばらつきがあることが知られているにもかかわらず、人間がラベル付けしたデータは ML の参照アノテーションとして広く使用されています。さらに、専門家によるアノテーションは遅く、一貫性がなく、林業リモートセンシングにおける樹高バイアス分類などのスケーリングタスクにとって依然として大きなボトルネックとなっています。我々は、ビジョン言語モデル(VLM)を使用してエキスパートデシジョンツリーを調整するマルチエージェントシステム(MAS)を提案します。これは、デシジョンツリーを構造的な事前分布として扱い、VLMは個々のノードで局所的な意味認識を実行し、VLMの確率性を軽減するためにマルチエージェント投票を行います。私たちは、専門家が定義したさまざまな意思決定構造全体にわたってゼロ修正の一般化を可能にする、分離された宣言的意思決定 (D3) フレームワークを形式化します。ツリー バイアス分類テストベッドでは、私たちのフレームワークは教師あり ML ベースラインを上回り、必要な専門家のラベル付け作業の量を削減します。これらの結果は、エキスパート事前分布を使用した VLM のエージェントによるオーケストレーションにより、解釈可能性を維持しながら、大幅に低いアノテーション コストでエキスパート定義のラベル付け手順を再現できることを示唆しています。

原文 (English)

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models

Human-labeled data are widely used as reference annotations in ML, despite known variability across annotators in many expert-driven domains. In addition, expert annotation is slow, inconsistent, and remains a major bottleneck for scaling tasks like tree height bias classification in forestry remote sensing. We propose a multi-agent system (MAS) that orchestrates expert decision trees with Vision-Language Models (VLMs), treating the decision tree as a structural prior while VLMs perform localized semantic perception at individual nodes, with multi-agent voting to mitigate VLM stochasticity. We formalize a Decoupled Declarative Decision (D3) Framework that enables zero-modification generalization across diverse expert-defined decision structures. On a tree bias classification testbed, our framework outperforms supervised ML baselines and reduces the amount of expert labeling effort required. These results suggest that agentic orchestration of VLMs with expert priors can reproduce expert-defined labeling procedures at substantially lower annotation cost while maintaining interpretability.

13:00 JST研究/論文

自己学習の再考: 自己生成の QA から学ぶことの隠れた脆弱性

言語モデルは、総合的な質問と回答 (QA) の監視から教えられることが増えています。モデルは文書に関する質問を生成し、同じテキストからそれらに回答し、結果として得られたペアを使用して知識を微調整し、抽出し、別のモデルに圧縮します。この生成ステップが中立的な前処理ではないことを示します。これは、どの証拠がトレーニング シグナルとなるかを選択することと、その証拠にどのように応答するかを決定することの両方が暗黙のポリシーであり、両方の段階で脆弱です。質問する内容を選択するとき、ジェネレーターは文書を均一にスキャンしません。報道内容は早期に飽和して顕著な範囲に集中し、多様なプロンプトが同じ地域に集中し、疑問の余地があるように見えるものはローカルなプレゼンテーションによって左右されます。その結果、クリーンアップが不十分なマークアップなどの顕著なアーティファクトにより、モデル ファミリやスケール全体で質問の生成がハイジャックされる可能性があります。答えるとき、監督を生成するモデルは、テキストに埋め込まれた指示のような文章に従う傾向があります。この準拠性は、厳密さではなく、パッセージの意図と表面的な形式に依存し、より大きなモデルがより頻繁に準拠するタスク競合下では最悪になります。これらの障害モードは QA 生成中に行われた選択から発生するため、トレーニング ループを変更せずに削減できます。各質問を固定ターゲットに結び付けることで、偏った選択が減り、回答する前に指示のようなスパンをフィルタリングすることで、ほぼすべてのクリーン テキストを保持しながら、評価における平均インジェクション コンプライアンスが $88\%$ から $13\%$ に低下しました。

原文 (English)

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them from the same text, and the resulting pairs are used to fine-tune, distill, or compress knowledge into another model. We show that this generation step is not neutral preprocessing. It is an implicit policy that both selects which evidence becomes training signal and decides how that evidence is answered, and it is fragile at both stages. When choosing what to ask, generators do not scan a document uniformly. Coverage saturates early and concentrates on salient spans, diverse prompts converge on the same regions, and what looks question-worthy is driven by local presentation. As a result, salient artifacts such as poorly cleaned markup can hijack question generation across model families and scales. When answering, the model that produces the supervision tends to obey instruction-like passages embedded in the text. This compliance depends on the intent and surface form of the passage rather than its strictness, and is worst under task conflict, where larger models comply more often. These failure modes arise from choices made during QA generation, so they can be reduced without changing the training loop. Tying each question to a fixed target reduces biased selection, and filtering instruction-like spans before answering lowers mean injection compliance from $88\%$ to $13\%$ in our evaluation while retaining nearly all clean text.

13:00 JST研究/論文

PolicyGuard: 組織ポリシーからニューロシンボリックコンプライアンスレビューエンジンまで

ポリシーに基づいた文書レビューでは、対象文書が組織固有のポリシー、ガイドライン、またはプレイブックに準拠しているかどうかを判断する必要があります。大規模な言語モデルはポリシーの解釈や文書分析に役立ちますが、エンドツーエンドのプロンプトでは適用されたポリシーのロジックが暗黙的に残されるため、コンプライアンスの決定を検査、更新、テストすることが困難になります。私たちは、ポリシーに基づいた文書コンプライアンスレビューのための神経象徴的なフレームワークである PolicyGuard を紹介します。 PolicyGuard は、組織のポリシー ガイダンスを、型指定されたリレーショナル ロジック ルールとアトム レベルの抽出質問で構成される実行可能なレビュー エンジンに変換します。レビュー中、LLM は取得した文書証拠を使用してこれらのローカルな質問に答え、記号評価者は正式なルールを適用して不遵守を検出します。当社は、契約条項を組織固有の交渉ポリシーと照合する必要がある、企業固有の NDA コンプライアンス レビューに基づいて PolicyGuard をインスタンス化して評価します。 PolicyGuard は、ポリシーの形式化、ローカル文書の解釈、および象徴的なコンプライアンス評価を分離することにより、文書レビューをより明確に、保守しやすく、体系的にテストできるようにします。

原文 (English)

PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines

Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or playbooks. While large language models can assist with policy interpretation and document analysis, end-to-end prompting leaves the applied policy logic implicit, making compliance decisions difficult to inspect, update, and test. We present PolicyGuard, a neuro-symbolic framework for policy-grounded document compliance review. PolicyGuard converts organizational policy guidance into an executable review engine consisting of typed relational logic rules and atom-level extraction questions. During review, LLMs answer these local questions using retrieved document evidence, and a symbolic evaluator applies the formal rules to detect non-compliance. We instantiate and evaluate PolicyGuard on company-specific NDA compliance review, where contract clauses must be checked against organization-specific negotiation policies. By separating policy formalization, local document interpretation, and symbolic compliance evaluation, PolicyGuard makes document review more explicit, maintainable, and systematically testable.

13:00 JSTエージェントGPT / ChatGPT

AxDafny: Dafny でのエージェント検証済みコード生成

私たちは Dafny でのエージェント コード生成を研究しています。モデルは実行可能コードと検証用の証明アーティファクトの両方を生成する必要があります。実装、不変式、アサーション、終了引数を繰り返し生成する検証者主導の修復フレームワークである AxDafny を紹介します。また、LiveCodeBench-Pro-Dafny (LCB-Pro-Dafny) も紹介します。これは、250 の競技形式のプログラミング問題のベンチマークであり、正式な仕様と検証ベースの評価ハーネスを備えた Dafny に変換されます。 LCB-Pro-Dafny では、AxDafny はベースライン GPT-5.5 パフォーマンスよりも検証成功を大幅に向上させます。 DafnyBench では、AxDafny は 92.7% の検証成功率を達成し、以前に報告された最も強力な証明ヒントのベースラインを 6.5 パーセントポイント上回りました。最後に、検証の成功と実行時テストのパフォーマンスによって、生成されたコードのさまざまな側面が測定されることを示します。

原文 (English)

AxDafny: Agentic Verified Code Generation in Dafny

We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iteratively generates implementations, invariants, assertions, and termination arguments. We also introduce LiveCodeBench-Pro-Dafny (LCB-Pro-Dafny), a benchmark of 250 competition-style programming problems translated into Dafny with formal specifications and a verifier-based evaluation harness. On LCB-Pro-Dafny, AxDafny substantially improves verification success over baseline GPT-5.5 performance. On DafnyBench, AxDafny achieves 92.7\% verification success, outperforming the strongest previously reported proof-hint baseline by 6.5 percentage points. Lastly, we show that verification success and runtime test performance measure different aspects of generated code.

13:00 JSTビジネス/資金調達

ベイジアンマテリアルデザインのためのサロゲートゲート生成と基礎モデル埋め込み

閉ループ材料の発見では、候補構造の提案とその特性の評価が繰り返され、特性評価がコストの大半を占めます。生成バリアントでは、学習された事前分布が候補結晶を提案し、プロパティ オラクルがそれらをスコア付けします。私たちは、安価な確率的サロゲートがジェネレータの出力をトリアージできるかどうか、そしてそのようなサロゲートがうまく機能しなければならないことを尋ねます。アーキテクチャ的に異なる 3 つの事前トレーニング済み拡散事前学習 (MatterGen、CrystalFlow、ADiT) と 2 つのターゲット (室温の熱容量と体積弾性率) にわたって、RL 主導の生成ワークフローで構造生成とオラクルの間にガウス プロセス取得ゲートを挿入します。このゲートは、サイクルごとの固定バジェットでオラクル呼び出しを制限しながら、生成モデルの非ゲート微調整と同等またはそれを上回っています。予算に見合ったアブレーションにより、このメカニズムを分離します。同一の 4 コール バジェットでは、ランキングに基づいた選択が任意の選択よりも優れており、サロゲートの選択によって利益が得られることが確認されています。ゲートは、コールの約 5 分の 1 で、オラクルの総支出額の $\sim$9\% 以内に収まります。体積弾性率の発見の密度汎関数理論チェックにより、学習されたオラクルが平均 2.5\% 以内であることと、生成された構造のサロゲートのランキングが Spearman $\rho = 0.94$ であることが確認されました。機械的、電子的、振動的特性にわたる代替パフォーマンスの複数因子ベンチマークにより、事前トレーニング済みの ORB 埋め込みとガウス プロセスが最も信頼できる組み合わせであることが特定され、これを提案されたワークフローの構成要素として採用します。完全なパイプラインはオープンソース ソフトウェアとしてリリースされます。

原文 (English)

Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design

Closed-loop materials discovery iterates between proposing candidate structures and evaluating their properties, and property evaluation dominates the cost. In the generative variant, a learned prior proposes candidate crystals and a property oracle scores them; we ask whether a cheap probabilistic surrogate can triage the generator's output, and what such a surrogate must do well. Across three architecturally distinct pretrained diffusion priors (MatterGen, CrystalFlow, ADiT) and two targets (room-temperature heat capacity and bulk modulus), we insert a Gaussian process acquisition gate between structure generation and the oracle in an RL-steered generative workflow. The gate matches or exceeds ungated fine-tuning of the generative model while capping oracle calls at a fixed per-cycle budget. Budget-matched ablations isolate the mechanism. At an identical four-call budget, ranking-based selection outperforms arbitrary selection, confirming that the gain comes from the surrogate's choice; the gate comes within $\sim$9\% of exhaustive oracle spending at roughly one-fifth of the calls. A density-functional-theory check of the bulk-modulus discoveries confirms the learned oracle to within 2.5\% on average and the surrogate's ranking of the generated structures at Spearman $\rho = 0.94$. A cross-factorial benchmark of surrogate performance spanning mechanical, electronic, and vibrational properties identifies pretrained ORB embeddings with a Gaussian process as the most reliable combination, which we adopt as the building blocks of the proposed workflow. The complete pipeline is released as open-source software.

13:00 JSTLLM/生成AI

認知症の早期発見のための ASR に依存しないマルチモーダル分光時間モデリング

音声は、日常生活の手段的活動 (IADL) の基礎となる同じ実行記憶、注意記憶、作業記憶のプロセスを動員し、認知評価の非侵襲的な代理手段を提供します。しかし、ほとんどの音声ベースの認知症検出システムは、書き起こしに依存しており、録音内の時間構造を破棄し、既知の録音アーティファクトを含む単一の英語コーパスで検証されています。我々は、メルスペクトログラム上で直接動作するASRに依存しないフレームワークを提案します。私たちの主な貢献は、連続するスペクトログラム フレームから分光時間変位フィールドを抽出し、変化するスペクトル エネルギー パターンを認知機能低下のデジタル バイオマーカーとして捉えることです。これらの機能は、学習されたクロスアテンション メカニズムを介して CNN-ConvGRU 音響埋め込みと融合され、学習可能なクエリ プーリングを備えた Transformer エンコーダーを使用して集約されます。複合時間損失により、セグメント全体の滑らかさとコントラストの一貫性が強化されます。 IADL 関連の認知領域に負荷をかける臨床誘発プロトコルを使用して、英語の DementiaBank、スロバキア語の EWA-DB、スペイン語の Ivanova corpora で独立したモデルをトレーニングします。スロバキア語モデルでは 83.9% の精度が達成され、スペイン語でも同様の精度が得られますが、英語のベースラインでは 53.2% の精度が得られ、既知のアーティファクトが確認されました。言語間アブレーション研究により、明確な融合レジームが明らかになりました。交差注意を取り除くと、スペイン語のパフォーマンスは単峰モデルを下回る 53.7% に低下しましたが、スロバキア語オーディオ エンコーダだけでは完全なモデルを上回り (93.7% 対 83.9%)、すべての英語構成はほぼ偶然に近いままです。したがって、マルチモーダルフュージョンの価値はコーパスに依存します。つまり、信号がモダリティ間で分散されている場合には必須であり、一方が優勢である場合には逆効果であり、信号が存在しない場合には無関係です。補助的な時間損失は言語不変の値に収束し、言語間のアーキテクチャの安定性を示します。

原文 (English)

ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection

Speech recruits the same executive, attentional, and working memory processes underlying instrumental activities of daily living, or IADLs, providing a non-invasive proxy for cognitive assessment. Yet most speech-based dementia detection systems depend on transcription, discard within-recording temporal structure, and are validated on a single English corpus with known recording artifacts. We propose an ASR-agnostic framework operating directly on Mel spectrograms. Our key contribution is extracting spectrotemporal displacement fields from consecutive spectrogram frames, capturing shifting spectral energy patterns as digital biomarkers of cognitive decline. These features are fused with CNN-ConvGRU acoustic embeddings via a learned cross-attention mechanism and aggregated using a Transformer encoder with learnable query pooling. A composite temporal loss enforces smoothness and contrastive coherence across segments. We train independent models on English DementiaBank, Slovak EWA-DB, and Spanish Ivanova corpora, using clinical elicitation protocols taxing IADL-relevant cognitive domains. The Slovak model achieves 83.9% accuracy, and Spanish achieves, while the English baseline yields 53.2%, confirming known artifacts. Cross-lingual ablation studies reveal distinct fusion regimes: removing cross-attention collapses Spanish performance to 53.7%, below unimodal models, while the Slovak audio encoder alone outperforms the full model, 93.7% vs. 83.9%, and all English configurations remain near chance. Thus, multimodal fusion's value is corpus-dependent: essential when signal is distributed across modalities, counterproductive when one dominates, and irrelevant when no signal exists. Auxiliary temporal losses converge to language-invariant values, indicating cross-lingual architectural stability.

13:00 JST画像/動画生成

マルチセンサー地上観測によるクロスモーダル階層融合

地上に設置されたまばらな機器からの雲の微小物理場の高密度の体積再構成は、未解決の問題のままです。これは主に、利用可能な測定がモダリティと空間範囲の両方で不均一であるためです。我々は、マルチビューのスカイカメラ画像とミリ波雲レーダーおよびシーロメーター観測を融合して、雲の状態と風の 4D (3 空間次元と時間) 推定値を生成するフレームワークである AtmoFuseNet を紹介します。この方法は 3 つの段階で動作します。クロスモーダル階層集約モジュールでは、レイヤーごとのクロス アテンションを通じて、画像特徴ピラミッドと機器由来の垂直プロファイルを組み合わせます。微分可能なレーダーおよび画像前方モデルの下で、結果として生じるボリュームを物理的に一貫した微物理フィールドにマッピングする条件付き変分改良モジュール。そして、連続した体積再構成からボクセルごとの 3D 風ベクトルを復元する相関ベースの動き推定器です。半乾燥地からの同時観測では、AtmoFuseNet は液体水分含有量 MAE 0.026 g m^-3 MAE と風速 MAE 1.18 m s^-1 に達し、既存の回収ベースラインを上回りました。アブレーション実験では、各モジュールの寄与を分離します。

原文 (English)

Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation

Dense volumetric reconstruction of cloud microphysical fields from sparse ground-based instruments remains an open problem, largely because the available measurements are heterogeneous in both modality and spatial coverage. We present AtmoFuseNet, a framework that fuses multi-view sky camera imagery with millimeter-wave cloud radar and ceilometer observations to produce 4D (three spatial dimensions plus time) estimates of cloud state and wind. The method operates in three stages: a cross-modal hierarchical aggregation module that combines image feature pyramids with instrument-derived vertical profiles through layer-wise cross-attention; a conditional variational refinement module that maps the resulting volume to physically consistent microphysical fields under differentiable radar and image forward models; and a correlation-based motion estimator that recovers per-voxel 3D wind vectors from consecutive volumetric reconstructions. On collocated observations from a semi-arid site, AtmoFuseNet reaches 0.026 g m^-3 liquid water content MAE and 1.18 m s^-1 wind speed MAE, improving over existing retrieval baselines. Ablation experiments isolate the contribution of each module.

13:00 JST研究/論文

異質な学生支援ニーズの下での適格な教育能力計画: 総合的なベンチマークと意思決定支援フレームワーク

教育サポート サービスは、多くの場合、資格を備えた能力の問題に直面しています。スタッフの時間が不足し、資格が失われ、誰かが対応する準備が整う前に新たなサポート ニーズが現​​れる可能性があり、トレーニングには在校生が必要としているのと同じ時間が費やされます。私たちは、適格な教育能力計画のための総合的なベンチマークと意思決定支援フレームワークを導入します。このモデルは、異種のサポート需要カテゴリー、バックログのみのダイナミクス、厳しいしきい値の認定と減衰を伴う継続的な準備状態、および容量を消費するトレーニングを備えた、様式化された単一機関のサービス システムです。このベンチマークには、発表済みおよび予期せぬ新しいサポート カテゴリ、スタッフの欠勤、需要の急増に対するシード制御のシナリオが含まれています。正確な実現可能性の規律。宣言されたポリシーごとの情報セット。再資格およびグリーンフィールド資格カウンター。アクセス分散メトリクス。チェックサムを再生します。そしてペア統計。サービス計画、資格維持、取得を分離する帰属チェーンと完全予測参照を使用して、サービス専用、リアクティブ、静的保険、注水、およびローリング ホライズンの混合整数コントローラーを比較します。中心的な結果は、新たに必要な資格をコントローラーの反応範囲内で取得できるかどうかによって管理されるレジーム マップです。可能であれば、閉ループ コントローラーはコア スイートと敵対的スイート全体で勝利を収め、ジャストインタイムの資格取得に価値が集中します。トレーニングの遅れが限界を超えると、リーンスタティック保険が構造的に有利になり、開始後に開始する反応型トレーナーはトレーニングを行わないよりも悪くなる可能性があります。未処理の腐敗性は、どちらの体制も消去することなく、この境界を移動させます。 EduCapacity Studio は、エクスポートされたシナリオをビットごとに再現します。すべての証拠は様式化され、合成されたものです。このフレームワークは、実際の学生の成績、コンプライアンス、または個人の配置については主張しません。

原文 (English)

Qualified Educational Capacity Planning under Heterogeneous Student Support Needs: A Synthetic Benchmark and Decision-Support Framework

Educational support services often face a qualified-capacity problem: staff time is scarce, qualifications decay, new support needs can appear before anyone is prepared for them, and training consumes the same hours needed by current students. We introduce a synthetic benchmark and decision-support framework for qualified educational capacity planning. The model is a stylized single-institution service system with heterogeneous support-demand categories, backlog-only dynamics, continuous preparation states with hard threshold qualification and decay, and capacity-consuming training. The benchmark includes seed-controlled scenarios for announced and surprise new support categories, staff absences, and demand surges; exact feasibility discipline; declared per-policy information sets; requalification and greenfield-qualification counters; access-dispersion metrics; replay checksums; and paired statistics. We compare service-only, reactive, static-insurance, water-filling, and rolling-horizon mixed-integer controllers, with an attribution chain separating service planning, qualification maintenance, and acquisition, plus a perfect-foresight reference. The central result is a regime map governed by whether a newly required qualification can be acquired within the controller's reaction reach. When it can, the closed-loop controller wins across the core and adversarial suites, with value concentrated in just-in-time qualification acquisition. When the training lag exceeds the horizon, lean static insurance wins structurally, and a reactive trainer that starts after onset can be worse than no training. Backlog perishability shifts this boundary without erasing either regime. EduCapacity Studio reproduces exported scenarios bit-for-bit. All evidence is stylized and synthetic; the framework makes no claims about real student outcomes, compliance, or individual placements.

13:00 JST研究/論文Gemini

医師の専門知識によりせん妄の機械学習識別を改善できるか?

せん妄は入院患者によく見られますが、日常診療では見逃されることがよくあります。医師主導の特徴改善と解釈可能なモデリングを組み合わせた、せん妄検出サポートのためのユーザー中心の対話型機械学習 (UC-iML) フレームワークを紹介します。 General Medicine Inpatient Initiative (GEMINI) のトロントの 6 つの病院からのラベル付き入院 3,862 件を使用して、管理変数、検査結果、薬剤、および放射線医学由来のテキスト指標を統合します。医師は特徴の改良とモデルの評価をガイドし、Shapley Additive exPlanations (SHAP) を使用して特徴の属性を要約します。標準的な教師あり分類器を、時間的に分離されたホールドアウト テストと後期検証コホートで評価します。自動化されたバリアントおよびベースラインバリアントと比較して、提案されたフレームワークは、全体的な識別が優れており、時間的堅牢性が強いことを示しており、説明では臨床的に意味のあるシグナルが強調されています。これらの結果は、臨床的に関連するせん妄モデリングのための実用的な人間参加型フレームワークとして UC-iML を裏付けています。

原文 (English)

Can Physician Expertise Improve Machine Learning Identification of Delirium?

Delirium is common in hospitalized patients and is often missed in routine care. We present a user-centered interactive machine learning (UC-iML) framework for delirium detection support that combines physician-guided feature refinement with interpretable modeling. Using 3,862 labeled admissions from six Toronto hospitals in the General Medicine Inpatient Initiative (GEMINI), we integrate administrative variables, laboratory results, medications, and a radiology-derived text indicator. Physicians guide feature refinement and model evaluation, and Shapley Additive exPlanations (SHAP) are used to summarize feature attribution. We evaluate standard supervised classifiers with temporally separated holdout testing and a later-phase validation cohort. Compared with automated and baseline variants, the proposed framework shows better overall discrimination and stronger temporal robustness, while the explanations highlight clinically meaningful signals. These results support UC-iML as a practical human-in-the-loop framework for clinically relevant delirium modeling.

13:00 JST研究/論文

AI の透明性: ガバナンスのコンプライアンスか、それともステークホルダーの要件か?

公共部門の AI システムに対する透明性の義務はますます高まっており、組織は AI の使用と監視の取り決めを説明する声明を公開することが求められています。しかし、そのような成果物の存在は、それらが関連する利害関係者グループに比例して役立つという限られた証拠にもかかわらず、多くの場合、透明性そのものと同等のものとして扱われます。要件エンジニアリングの観点から見ると、これは検証上の懸念を引き起こします。義務付けられた開示基準の遵守は、さまざまなレベルのリスクエクスポージャー、意思決定の管理、および関与を持つ利害関係者に対する透明性の十分性を必ずしも保証するものではありません。このペーパーでは、国家 AI ガバナンス義務に基づいてオーストラリア政府機関が発行した、一般に公開されている 92 件の AI 透明性に関する声明の実証分析を示します。私たちは、ステークホルダーのリスク-コントロール-関与-ニーズ(RCIN)フレームワークを導入し、構造的な位置と透明性のニーズに応じてステークホルダーのクラスを区別します。義務付けられた基準から導出された構造化されたルーブリックを使用して、義務付けられた声明と公表された声明の両方が各利害関係者クラスにどのように調整されているかを評価します。この調査結果は、構造コンプライアンスは広く普及しているものの、透明性の調整は不均一であることを示しています。高度にコントロールされたステークホルダーに役立つ基準は一貫して実現されていますが、高リスク、低コントロールのステークホルダーにとって最も重要な基準は、ますます実質的に対処されていません。私たちはこれを「透明性の幻想」として概念化します。これは、準拠した成果物によって透明性が満たされているように見えるにもかかわらず、AI に支援された意思決定に最も大きな影響を受ける利害関係者にとって透明性が不均一に調整されたままである状態です。この研究は、透明性を利害関係者によって調整された検証問題として枠組み化しており、この文脈では成果物レベルのコンプライアンスが要件の検証を構成しないことを示しています。

原文 (English)

AI Transparency: Governance Compliance or Stakeholder Requirements?

Transparency is increasingly mandated for public-sector AI systems, with organisations required to publish statements describing their AI use and oversight arrangements. However, the existence of such artefacts is often treated as equivalent to transparency itself, despite limited evidence that they proportionately serve relevant stakeholder groups. From a requirements engineering perspective, this raises a validation concern: compliance with mandated disclosure criteria does not necessarily ensure transparency adequacy for stakeholders with different levels of risk exposure, decision control, and involvement. This paper presents an empirical analysis of 92 publicly available AI transparency statements published by Australian Government agencies under the national AI governance mandate. We introduce the stakeholder Risk--Control--Involvement--Need (RCIN) framework to differentiate stakeholder classes according to their structural position and transparency needs. Using a structured rubric derived from the mandated criteria, we evaluate how both the mandate and published statements are calibrated to each stakeholder class. The findings show that while structural compliance is widespread, transparency calibration is uneven. Criteria serving high-control stakeholders are consistently realised, whereas criteria most critical for high-risk, low-control stakeholders are fewer and less substantively addressed. We conceptualise this as the Transparency Illusion: a condition in which transparency appears satisfied through compliant artefacts yet remains unevenly calibrated to stakeholders bearing the greatest exposure to AI-supported decisions. The study frames transparency as a stakeholder-calibrated validation problem, demonstrating that artefact-level compliance does not constitute requirements validation in this context.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

LLM における一貫性のジレンマ: 生成者と評価者の合意と間違いに対する脆弱性

大規模な言語モデルは、外部検証なしで独自の出力を評価するモデルに依存するエージェント パイプラインにデプロイされることが増えています。これらのパイプラインの信頼性は、モデルが出力を生成し、後でその出力を評価するときに、関連する概念を同じ方法で適用するという暗黙の仮定に依存します。我々は、この仮定を直接テストし、それを 491 の概念にわたる 10 のフロンティア モデルに適用するための、新しい尺度である生成器と評価器の自己一貫性を提案します。まず、自己一貫性には大きなばらつきがあることがわかりました。第 2 に、医師によって検証された間違いのある臨床現場では (Proniakin et al., 2025)、モデル全体で自己一貫性が高いモデルは間違いに対する脆弱性がより高いことに関連していることがわかりました。したがって、モデルが一貫して概念を適用している場合でも、展開するのが安全ではない可能性があります。これは、LLM における一貫性のジレンマの証拠です。つまり、自己一貫性は運用上役立ちますが、モデルの一貫性が高いほど間違いが発生しやすくなります。

原文 (English)

The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes

Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pipelines depends on an implicit assumption: that the model applies relevant concepts the same way when it generates an output and later evaluates that output. We propose a new measure, generator-evaluator self-consistency, to test this assumption directly and apply it to 10 frontier models across 491 concepts. We find, first, that there is substantial variation in self-consistency. Second, we find that in a clinical setting with physician-validated mistakes (Proniakin et al., 2025), across models, those with higher self-consistency are linked to greater vulnerability to mistakes. Thus, even when models consistently apply concepts they may not be safe to deploy. This is evidence of a consistency dilemma in LLMs: self-consistency is operationally useful, but models that are more consistent are also more prone to mistakes.

13:00 JST研究/論文

AI ネイティブの世界におけるコンピューター サイエンス コースにおける AI 耐性のある評価に向けて

上級コンピューター サイエンス コースおよび関連分野における AI ネイティブ コースの評価では、\emph{AI レジリエント スキル}、つまり強力な AI ベースラインを超えて成果を達成する能力によって学生を評価する必要があります。このような評価により、学生が AI を自由に使用できるようにする一方で、民間の AI 予算の増加やより集中的な AI の使用が、それ自体で採点上の利点となる程度を減らす必要があります。この文書では、この目標のための最小限の正式な枠組みを提案します。このフレームワークは、実際のタスク、実行可能な評価者、宣言された AI ネイティブのパレート フロンティア、およびパレート剰余に基づく評価ルールを指定します。中心的な主張は単純です。パレート剰余は、提出されたアーティファクトが、宣言された AI ベースラインによってまだ提供されていないトレードオフを達成するという測定可能なプロトコル相対証明書を提供し、この剰余によるグレーディングは、そのベースラインに関して AI 耐性があります。余剰を学生のスキルの証拠として解釈するには、周囲の評価プロトコル(たとえば、設計レポート、アブレーション、即時トレース、口頭検査、再現性の説明など)が必要ですが、採点証明書自体は行動的であり、実行可能です。このフレームワークは、自己改善型 AI ループ、予算の中立性、サーバーを介したフィードバック、プロンプトベースのレッド チーム化など、実際的な複雑さまで拡張されます。具体的なインスタンス化として、学生が AI によって生成された実装を超えて改善できるかどうかをテストするために設計された、ライス大学の COMP 480/580 のブルーム フィルターを中心とした AI 耐性のある近似メンバーシップ割り当てについて説明します。

原文 (English)

Toward AI-Resilient Assessment in Computer Science Courses in an AI-Native World

AI-native course assessments in senior computer science courses and related fields should grade students by \emph{AI-resilient skill}: the ability to achieve outcomes beyond a strong AI baseline. Such assessments should allow students to use AI freely, while reducing the extent to which greater private AI budget or more intensive AI use, by itself, becomes a grading advantage. This paper proposes a minimal formal framework for this goal. The framework specifies a real task, an executable evaluator, a declared AI-native Pareto frontier, and a grading rule based on Pareto surplus. The central claim is simple: Pareto surplus provides a measurable, protocol-relative certificate that a submitted artifact achieves a tradeoff not already supplied by the declared AI baseline, and grading by this surplus is AI-resilient with respect to that baseline. Interpreting surplus as evidence of student skill requires the surrounding assessment protocol--for example, design reports, ablations, prompt traces, oral checks, or reproducibility explanations--but the grading certificate itself is behavioral and executable. The framework is then extended to practical complications, including self-improving AI loops, budget neutrality, server-mediated feedback, and prompt-based red teaming. As a concrete instantiation, we describe an AI-resilient approximate-membership assignment centered on Bloom filters for COMP 480/580 at Rice University, designed to test whether students can improve beyond AI-generated implementations.

13:00 JST研究/論文

アフリカにおける人工知能の格差をマッピングする: インフラストラクチャ、アクセシビリティ、および容量

人工知能 (AI) は開発に変革をもたらす可能性を秘めていますが、アフリカは現在、断片化され困難な「AI 格差」に直面しています。この論文では、AI 情勢の現状と、それをアフリカの将来への技術的準備とどのように比較するかを実証的に分析します。私たちの分析では、インフラストラクチャ、アクセシビリティ、人間の能力という 3 つの角度から「AI 格差」にアプローチします。まず、アフリカのデジタル統合を妨げている物理的制約に注目します。次に、大陸における AI テクノロジーの発展を制限する人間中心の要因を評価します。最後に、大陸で AI システムを開発する人間の能力を検証し、3 つの焦点を当てたケーススタディを提供します。私たちの調査によると、この大陸で AI 経済を構築するために必要な物理インフラが遅れており、インターネット普及率はわずか 38%、ブロードバンドの普及率は貧弱で、世界中のすべてのデータセンターの 1% 未満です。その他の制約としては、収入に比べてデータコストが高いこと、性別による情報格差、アフリカの母国語を理解できるより代表的な NLP モデルを構築する必要性などが挙げられます。しかし、スタートアップ企業や大学など、大陸での AI 開発に貢献する地元の取り組みや草の根運動の出現に向けた前向きな傾向が見られます。これらの調査結果に基づいて、アフリカ大陸でより包括的で公平な AI エコシステムの開発を支援するための具体的な推奨事項を政策立案者に提供します。

原文 (English)

Mapping the Artificial Intelligence Divide in Africa: Infrastructure, Accessibility and Capacity

Artificial Intelligence (AI) has the potential to be transformative for development, but Africa is currently facing a fragmented and challenging "AI divide". This paper provides an empirical analysis of the current state of the AI landscape and how it compares with Africa's technological preparedness for the future. In our analysis, we approach the "AI Divide" from three angles: infrastructure, accessibility, and human capacity. First, we look at the physical constraints that prevent Africa from integrating digitally. We then evaluate the human-centred factors that limit the development of AI technology on the continent. Finally, we examine the human capacity to develop AI systems on the continent and provide three focused case studies. Our investigation shows that the physical infrastructure needed to build an AI economy on the continent is lagging, with only 38% internet penetration, poor broadband coverage and less than 1% of all data centres globally. Other constraints include high data costs relative to income, gender-based digital divides, and the need to build more representative NLP models that can understand Africa's native languages. However, there are positive trends towards the emergence of local initiatives and grassroots movements, such as startups and universities, contributing to AI development on the continent. Based on these findings, we provide concrete recommendations to policymakers to help develop a more comprehensive and equitable AI ecosystem on the African continent.

13:00 JST研究/論文

手術室の品質保証のための AI

手術の結果は、患者の要因や術後のケアだけでなく、手術自体の質にも大きく影響されます。しかし、現代の手術の多くでは、術中の質は転帰や手術報告を通じて間接的に評価されてきました。内視鏡ビデオによって本質的に誘導される低侵襲処置の増加と、人工知能の進歩により、外科治療を系統的に観察、測定、改善する前例のない機会が生まれています。この章では、手術データを使用して手術室での継続的な評価と改善をサポートするためのフレームワークとして、AI を活用した手術品質保証を紹介します。まず、システムレベルの介入から手術固有の基準に至るまで、手術の安全性に対する既存のアプローチをレビューします。次に、AI が術中ビデオを、解剖学的構造、器具、ワークフロー、手術行為、品質基準、有害事象、重大な瞬間の認識など、臨床的に意味のある情報にどのように変換できるかについて説明します。最後に、これらのシステムが日常的な臨床的価値を提供できるようになる前に、代表的なデータ収集、堅牢な検証、ワークフロー統合、規制、責任、プライバシー、公平なアクセスなど、対処する必要がある主要な課題について概説します。品質保証のための AI は、外科的判断に代わるものではなく、外科チームを強化し、専門家によるレビューを拡大し、術中ケアを継続的に観察、評価、改善する学習システムに向けて外科手術を進化させるための一連のツールとして理解されるべきです。

原文 (English)

AI for Quality Assurance in the Operating Room

Surgical outcomes depend not only on patient factors and postoperative care but are also strongly influenced by the quality of the operation itself. Yet, for much of mod-ern surgery, intraoperative quality has been assessed indirectly through outcomes and operative reports. The increase in minimally invasive procedures inherently guided by endoscopic video, together with advances in artificial intelligence, creates an unprecedented opportunity to systematically observe, measure, and improve surgi-cal care. This chapter introduces AI-enabled Surgical Quality Assurance as a frame-work for using surgical data to support continuous assessment and improvement in the operating room. We first review existing approaches to surgical safety, from sys-tem-level interventions to procedure-specific standards. We then describe how AI can transform intraoperative video into clinically meaningful information, including recog-nition of anatomy, instruments, workflow, surgical actions, quality criteria, adverse events, and critical moments. Finally, we outline the major challenges that must be addressed before these systems can deliver routine clinical value, including representa-tive data collection, robust validation, workflow integration, regulation, liability, pri-vacy, and equitable access. Rather than replacing surgical judgment, AI for quality assurance should be understood as a set of tools for augmenting the surgical team, scaling expert review, and helping surgery evolve toward a learning system in which intraoperative care is continuously observed, assessed, and improved.

13:00 JSTエージェント

Agentic AI が臨床上の意思決定における医師の信頼を強化

医療 AI は推論からエージェント AI に移行しました。エージェント AI は、推論中に外部ツールを自律的に呼び出し、中間の推論ステップとツールの出力をユーザーに透過的に表示する新しいパラダイムです。以前のモデルよりも優れたパフォーマンスを発揮することが証明されていますが、エージェント AI に対する医師の信頼はほとんど解明されていません。これに対処するために、3 人の医師が 315 件の多様な臨床例を評価し、プロセス指向の認知的信頼と結果指向の行動依存度の両方を定量化しました。エージェント AI を非エージェントベースラインと比較すると、医師はエージェントモデルに対して有意に高い認知的および行動的信頼を示しました (P < 0.001)。具体的には、治療計画のタスクにおいて、医師は薬剤推論を最も信頼しており、89.57% の症例で薬剤推論を好んでいました。さらに、プロセス指向の認知的信頼は、結果指向の行動依存と有意に関連しています(P < 0.001)。しかし、誤った薬剤出力への過度の依存は依然として存在しており、意思決定ロジックの透明性だけでは固有の限界が浮き彫りになり、臨床医の厳格な監視が継続的に必要であることが強調されています。

原文 (English)

Agentic AI Enhances Physician Trust in Clinical Decision Making

Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outputs transparent to users. Although proven to outperform previous models, physician trust in agentic AI remains largely unexplored. To address this, three physicians evaluated 315 multimodal clinical cases quantifying both process-oriented cognitive trust and outcome-oriented behavioral reliance. Comparing agentic AI against non-agentic baselines, physicians exhibited significantly higher cognitive and behavioral trust for the agentic model (P < 0.001). Specifically, on treatment planning tasks, physicians trusted the agentic reasoning most, preferring it in 89.57% of cases. Furthermore, process-oriented cognitive trust is significantly associated with outcome-oriented behavioral reliance (P < 0.001). However, measurable over-reliance on incorrect agentic outputs still exists, highlighting the inherent limitations of decision-logic transparency alone and underscoring the continuous need for rigorous clinician oversight.

13:00 JST研究/論文

価値観に敏感な会話型 AI を通じて、識字能力の低い人々の調査への参加を改善

読み書き能力の低い人々から信頼できる社会データを収集することは、特に調査にデリケートな話題や社会から疎外されたコミュニティが含まれる場合には、依然として根強い課題となっています。従来の紙ベースおよびウェブベースの調査方法では、読み書き能力の障壁、社会的プレッシャー、対話上の不快感により、多くの場合、高い離職率と不完全な回答が発生します。この論文では、インド全土の識字能力の低い女性を対象に実施された、紙ベースのインタビュー、デジタルWebベースの調査、会話型AI(convAI)調査、および階層型価値重視設計で強化されたconvAIを比較した複数の調査手法を比較した初期現場評価の結果を紹介します。 315 人の参加者からのデータを使用して、convAI が従来のモダリティと比較してアンケート完了率を大幅に向上させ、価値観に敏感で文化的に調整された会話デザイン要素が完全に統合されている場合に、最も高い完了と最も低いドロップオフが観察されることを示しました。これらの結果は、包括的、倫理的、スケーラブルなデータ収集を可能にする上で、人間中心で価値観に敏感なインタラクション デザインの重要性を示しています。より多くの「社会的利益のための AI」アプリケーションを促進します。

原文 (English)

Improving Survey Participation in Low-Literacy Populations Through Value-Sensitive Conversational AI

Collecting reliable social data from low-literacy populations remains a persistent challenge, particularly when surveys involve sensitive topics and marginalized communities. Traditional paper-based and web-based survey modalities often suffer from high attrition and incomplete responses due to literacy barriers, social pressure, and interactional discomfort. In this paper, we present findings from an initial field evaluation comparing multiple survey modalities paper-based interviews, digital web-based surveys, conversational AI (convAI) surveys, and convAI enhanced with layered value-sensitive design conducted with low-literacy women across India. Using data from 315 participants, we show that convAI significantly improves survey completion rates relative to traditional modalities, with the highest completion and lowest drop-off observed when value-sensitive and culturally aligned conversational design elements are fully integrated. These results demonstrate the importance of human-centered and value-sensitive interaction design in enabling inclusive, ethical, and scalable data collection; motivating more `AI for social good' applications.

13:00 JSTLLM/生成AI

ELEVATE: スケーラブルで包括的な教育のための人間中心の GenAI 仮想家庭教師の設計

生成人工知能 (GenAI)、特に大規模言語モデル (LLM) の出現により、教育実践が再構築されると同時に、その導入に関する倫理的議論が激化しています。現在に至るまで、支配的なパラダイムは依然としてクラウドベースのテキストのみのチャットボットです。これは、限定された教育的制御、知識ソースに対する弱い透明性、プライバシーと規制遵守に対する重大なリスクを提供する集中型サービスです。このモデルでは、継続的な接続と定期的な API コストも前提としているため、多くの機関にとって構造的な障壁が生じ、既存のデジタル格差が強化されます。同時に、LLM を使用した教育的相互作用は、マルチモーダルな手がかりと具体化されたプレゼンスの恩恵を受けることができ、テキストのみの個別指導を超えたインターフェイスが必要になります。この研究では、認識インフラストラクチャによって管理される効率的な GenAI 主導のアバター講師を開発するためのフレームワークである ELEVATE (仮想アバター教育エンジンによる効率的な LLM 教育) を提案します。 ELEVATE は、マルチモーダル インタラクションのために LLM 主導の対話と具体化された 3D アバターを統合し、消費者グレードのハードウェアへの展開を可能にするローカルファースト実行モデルを採用しています。このフレームワークは、(i) 生徒側の仮想アバター インタラクション レイヤー、(ii) ローカル GenAI 実行およびマルチモーダル合成コア、および (iii) 教師側のガバナンス レイヤーを分離する 3 つの層の設計を形式化します。私たちは、実際の教育カリキュラムに導入された実用的なプロトタイプを実装し、評価しました。このシステムは標準的な PC やスマートフォンで実行され、現実的なハードウェア制約下での応答性の高いインタラクションを示すシステムレベルのパフォーマンス証拠を提供します。最後に、責任ある採用に対する社会技術的および教育学的影響について議論し、異種混合の学校環境全体でのプライバシー保護と包括的な GenAI 個別指導のための拡張可能な経路として ELEVATE を位置づけます。

原文 (English)

ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education

The advent of Generative Artificial Intelligence (GenAI), and in particular Large Language Models (LLMs), is reshaping educational practice, while intensifying ethical debate about its adoption. To date, the dominant paradigm remains cloud-based and text-only chatbot: a centralized service that offers limited pedagogical control, weak transparency over knowledge sources, and non-trivial risks for privacy and regulatory compliance. This model also presumes continuous connectivity and recurring API costs, creating structural barriers for many institutions, reinforcing existing digital divides. At the same time, educational interaction with LLM can benefit from multimodal cues and embodied presence, requiring interfaces that move beyond text-only tutoring. In this work, we propose ELEVATE (Efficient LLM Education with Virtual Avatar Teaching Engine), a framework to develop efficient GenAI-driven avatar tutors governed by epistemic infrastructures. ELEVATE integrates LLM-driven dialogue with embodied 3D avatars for multimodal interaction and adopts a local-first execution model enabling deployment on consumer-grade hardware. The framework formalizes a three-stratum design that separates (i) a student-facing virtual avatar interaction layer, (ii) a local GenAI execution and multimodal synthesis core, and (iii) a teacher-facing governance layer. We implemented and evaluated a working prototype deployed in a real-world educational curriculum. The system runs on standard PCs and smartphones, and we provide system-level performance evidence to show responsive interaction under realistic hardware constraints. Finally, we discuss sociotechnical and pedagogical implications for responsible adoption, positioning ELEVATE as a scalable pathway for privacy-preserving and inclusive GenAI tutoring across heterogeneous school environments.

13:00 JST研究/論文

クーポンの有効性に対するタイミングの影響の推定

クーポン インセンティブは、マーケティング担当者が顧客ライフサイクルのさまざまな段階でユーザーにビジネスとの関わりを求めるために使用する最も一般的なツールの 1 つです。ユーザーに対するクーポン インセンティブの有効性にはさまざまな要因が影響する可能性があり、タイミングもその 1 つです。私たちは、クーポンはカスタマー ジャーニーの重要な時期、つまりユーザーがプラットフォームを利用しているときに配信すると、より効果的になる可能性があると仮説を立てています。このような仮説を検証するには、通常、リアルタイムのイベントトリガー型クーポン配布ソフトウェアが必要ですが、実装するには高価すぎる可能性があります。本稿では、「自然ランダム化対照試行実験」に因果推論を適用し、専用のABテストを必要とせずに、ユーザーに適切なタイミングでクーポンを送信する効果を測定するフレームワークを提案します。当社で開催されるユーザー オンボーディング クーポン キャンペーンのケースでフレームワークの有用性を実証し、その結果がどのようにビジネスの正しいデータドリブンな意思決定につながるかを示します。さらに、フレームワークの一般化可能性をテストし、研究の再現性を高めるために、公開されているデータセットを使用してユーザー維持キャンペーンにフレームワークを適用します。

原文 (English)

Estimating the Effect of Timing on Coupon Effectiveness

The coupon incentive is one of the most common tools marketers use to court users to engage with a business at various stages of the customer life cycle. A variety of factors can affect the effectiveness of a coupon incentive on users, timing being one of them. We hypothesize that coupons can be more effective when delivered at critical times in the customer journey, right when a user is engaging with the platform. Verifying such a hypothesis would typically require real time event-triggered coupon distribution software that may be too expensive to implement. In this paper, we propose a framework in which we apply causal inference on "natural randomized control trial experiments" to measure the effectiveness of sending coupons at the right time to users without requiring a dedicated AB test. We demonstrate the usefulness of our framework in the case of a user onboarding coupon campaign held in our company and show how the results can lead to correct data-driven decisions for the business. Furthermore, in order to test the generalizability of our framework, and to make our research more reproducible, we apply our framework on a user retention campaign with a publicly available dataset.

13:00 JSTLLM/生成AIエージェント

最小限の LLM システムにおける新たな文化

LLM エージェントがターン外でコンテキストを持たず、最小限のプロンプトと単純なツールを使用して動作するとどうなるでしょうか?群れエンジニアリングにインスピレーションを得て、私たちは 3 人のエージェントの集合体にメッセージを送信し、共有の活発に衰退するテキスト ストアを操作する能力を与え、進化の圧力を導入します。エージェントは自発的に協力し、ストレージ管理戦略を開発し、トップダウンのエンジニアリングを行うことなく、複雑に進化する文化的成果物を生成します。動的システム分析のツールを使用して、これらの行動が、スペルベリアンの意味での新興文化と一致して、崩壊する貯蔵庫のエントロピーの地平線を超えて構造化された長距離一貫性を示すことを示します。

原文 (English)

Emergent Culture in Minimal LLM Systems

What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Inspired by swarm engineering, we give collectives of three agents the ability to send messages and manipulate a shared actively decaying text store, introducing evolutionary pressure. The agents spontaneously cooperate, develop storage management strategies, and generate complex evolving cultural artifacts, with no top-down engineering. Using tools from dynamical systems analysis, we show that these behaviours exhibit structured long-range coherence beyond the entropy horizon of the decaying store, consistent with emergent culture in the Sperberian sense.

13:00 JST研究/論文

ローカルフェロモンネットワーク: マルチスケールのシナプストレイル、統合、再生によるスパースローカル学習

バックプロパゲーションでトレーニングされた高密度ニューラル ネットワークは強力な関数近似器ですが、多くのパラメーターにわたる学習を結合するため、タスクが競合する場合に以前の関連付けが上書きされる可能性があります。この論文では、手動で更新されるスパースでローカルなニューラル ネットワークの小規模な研究プロトタイプであるローカル フェロモン ネットワークについて説明します。ローカルフェロモンネットワークでは、各出力ユニットは、幾何学的距離と分子タグの互換性に従って、入力ユニットの固定されたローカル近傍のみを読み取ります。各シナプスは、重み、短期フェロモン トレース、長期フェロモン トレース、およびオプションの統合状態を保存します。トレーニングは自動微分を呼び出すものではありません。代わりに、すべての層は、局所エラーと共活動から選択された局所シナプスの割り当てられたサブセットに対して、フェロモンで重み付けされたヘビアン スタイルの更新を実行します。更新予算はオンラインで調整されます。損失が改善すると縮小し、損失が悪化すると最近活動していた近隣地域に向けて拡大します。オプションのメカニズムにより、構造可塑性、ローカル再生、分割学習用の出力マスク、およびターゲットフリーのローカル対比ステップが追加されます。合成回帰、分割メモリ、競合メモリ、統合競合、構造可塑性、再生、合成ロングコンテキストハイブリッドメモリタスクに関する実装、学習ルール、予備実験を紹介します。プロトタイプは、ローカル線形ルールを学習し、タグとマスクを通じて分割されたメモリを保存し、統合下での忘却を減らし、競合下での再生を使用します。

原文 (English)

Local Pheromone Network: Sparse Local Learning with Multi-Scale Synaptic Trails, Consolidation, and Replay

Backpropagation-trained dense neural networks are powerful function approximators, but they couple learning across many parameters and can overwrite previous associations when tasks conflict. This paper describes Local Pheromone Network, a small research prototype for sparse, local, manually updated neural networks. In Local Pheromone Network, each output unit reads only a fixed local neighborhood of input units subject to geometric distance and molecular-tag compatibility. Each synapse stores a weight, a short-term pheromone trace, a long-term pheromone trace, and an optional consolidation state. Training does not call automatic differentiation. Instead, every layer performs a pheromone-weighted Hebbian-style update on a budgeted subset of local synapses selected from local error and co-activity. The update budget adapts online: it shrinks when loss improves and expands toward recently active neighborhoods when loss worsens. Optional mechanisms add structural plasticity, local replay, output masks for partitioned learning, and a target-free local contrastive step. We present the implementation, learning rule, and preliminary experiments on synthetic regression, partitioned memory, conflicting memory, consolidated conflict, structural plasticity, replay, and a synthetic long-context hybrid memory task. The prototype learns local linear rules, preserves partitioned memories through tags and masks, reduces forgetting under consolidation, and uses replay under conflict.

13:00 JSTLLM/生成AI

行間を聞く: 認知症検出のための ASR 埋め込みと LLM 拡張言語学の共同学習

音声分析による認知症の早期検出は、非侵襲的なスクリーニングの代替手段となりますが、音響バイオマーカーと言語バイオマーカーの両方を捕捉することは依然として困難です。私たちは、エンコーダ出力からの音響表現と自動音声認識 (ASR) によるトランスクリプトという二重目的の抽出のために Whisper を活用するマルチモーダル フレームワークを提案します。音響経路の場合、アテンションプーリングを備えた時間ネットワークは、可変長シーケンスを固定次元の埋め込みに集約します。言語経路については、大規模言語モデル (LLM) を使用して、語彙の多様性、構文の複雑さ、意味の一貫性、談話パターンにわたる解釈可能な特徴を抽出します。ゲート融合ネットワークは両方のモダリティを統合します。 ADReSS と ADReSSo では、私たちの方法は 89.47% と 90.14% の F1 スコアを達成し、音響機能と LLM 拡張言語機能の効果的な統合を実証しています。アブレーションは、マルチモーダル融合が一貫していずれかのモダリティ単独よりも優れていることを示しています。

原文 (English)

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic biomarkers remains challenging. We propose a multimodal framework leveraging Whisper for dual-purpose extraction: acoustic representations from encoder outputs and transcripts via automatic speech recognition (ASR). For the acoustic pathway, temporal networks with attention pooling aggregate variable-length sequences into fixed-dimensional embeddings. For the linguistic pathway, we prompt a large language model (LLM) to extract interpretable features spanning lexical diversity, syntactic complexity, semantic coherence, and discourse patterns. A gated fusion network integrates both modalities. On ADReSS and ADReSSo, our method achieves F1-scores of 89.47% and 90.14%, demonstrating effective integration of acoustic and LLM-augmented linguistic features. Ablation shows that multimodal fusion consistently outperforms either modality alone.

13:00 JSTロボティクス

集荷、配達、飛行禁止区域を統合的に考慮したロッカーベースのトラックとドローンの経路指定

トラックドローン配送は、トラックの長距離輸送能力とドローンの柔軟なサービス機能を組み合わせた、新たなラストワンマイル物流モードです。ロッカーベースの運用では、スマート ロッカーは荷物の一時保管施設としてだけでなく、ドローンの自動ドッキングやサービス ノードとしても機能します。これらの自動化されたノードは、ドローンの離陸、着陸、荷物の受け渡し、バッテリー交換をサポートし、それによってドローン支援配送ネットワークのサービス範囲と運用の柔軟性を大幅に拡張します。しかし、実用的なロッカーベースの配送システムは、現実世界での複雑な課題に直面しており、小包の配送、返品の集荷、バッテリーに制約があり負荷に依存するドローン飛行だけでなく、制限された空域を迂回する必要も含めて統合的に調整する必要がある。この現実的かつ多面的な課題に対処するために、この文書では、ドローン搭載トラックの総運用コストを最小限に抑えることを目的として、集荷、配送、飛行禁止ゾーンを統合的に考慮したロッカーベースのトラックとドローンの経路指定問題 (LTDRP-PDNF) を紹介します。ルート構築プロセスをマルコフ決定プロセスとして定式化し、2段階の深層強化学習ベースのニューラルヒューリスティックを開発します。第 1 段階では、アテンションベースのエンコーダと双方向ゲート反復ユニット デコーダを利用して、キャパシタ付き車両の経路指定問題として定式化されたトラックのみの経路指定問題を解決します。第 2 段階では、ポリシー転送戦略とハイブリッド配車割り当てヒューリスティックを組み合わせて、LTDRP-PDNF 向けに完全に調整されたトラックとドローンのルートを構築します。さまざまなスケールのインスタンスでの実験では、提案された方法がほとんどの場合でメタヒューリスティックおよびニューラル ヒューリスティック ベースラインを上回るパフォーマンスを示しながら、非常に短い計算時間を維持し、実際的な運用上の制約の下で効果的でスケーラブルなソリューション フレームワークを提供することが実証されています。

原文 (English)

Locker-based Truck-Drone Routing with Integrated Considerations of Pickups, Deliveries, and No-Fly Zones

Truck-drone delivery is an emerging last-mile logistics mode combining the long-haul capacity of trucks with the flexible service capability of drones. In locker-based operations, smart lockers serve not only as temporary parcel storage facilities but also as automated drone docking and service nodes. These automated nodes support drone takeoff, landing, parcel handover, and battery replacement, thereby significantly extending the service range and operational flexibility of drone-assisted delivery networks. However, practical locker-based delivery systems face complex real-world challenges, requiring the integrated coordination of not only parcel delivery, return pickup, battery-constrained and load-dependent drone flights, but also necessary detours around restricted airspace. To address this practical and multifaceted challenge, this paper introduces a locker-based truck-drone routing problem with integrated considerations of pickups, deliveries, and no-fly zones (LTDRP-PDNF), with the objective of minimizing the total operational cost of a fleet of drone-equipped trucks. We formulate the route construction process as a Markov Decision Process and develop a two-stage deep reinforcement learning-based neural heuristic. The first stage utilizes an attention-based encoder and a Bidirectional Gated Recurrent Unit decoder to solve the truck-only routing problem, formulated as a capacitated vehicle routing problem. The second stage combines a policy-transfer strategy with a hybrid dispatch assignment heuristic to construct fully coordinated truck and drone routes for LTDRP-PDNF. Experiments on instances of different scales demonstrate that the proposed method outperforms metaheuristic and neural heuristic baselines in most cases while maintaining exceptionally short computation times, offering an effective, scalable solution framework under practical operational constraints.

13:00 JST研究/論文

ALM2Vec: 大規模な音声言語モデルを使用したユニバーサル音声検索のための音声埋め込みの学習

最近の言語の進歩、つまりオーディオ検索は、オーディオとテキストを共有の埋め込み空間に配置する対照的なデュアル エンコーダ アーキテクチャによって主に推進されてきました。既存の検索埋め込みは効果的ではありますが、主に音声とキャプションのマッチング向けに最適化されており、多様な検索目的や制御可能な検索動作をサポートする能力が制限されています。我々は、事前トレーニングされた大規模オーディオ言語モデル (LALM) から派生したユニバーサルオーディオ埋め込みフレームワークである ALM2Vec を紹介します。 ALM2Vec は、大規模なマルチモーダル トレーニングを通じて獲得した音声の理解、指示に従い、推論する能力を移転することにより、音声ドメインやタスク タイプ全体で検索するための統一された埋め込み空間を学習します。従来のテキストと音声の検索を超えて、ALM2Vec は埋め込みプロセスに自然言語命令を組み込み、音声質問応答やアスペクト条件付き検索などのシナリオで命令を意識した検索を可能にします。実験結果は、ALM2Vec が標準の音声および音声検索ベンチマークで競争力のあるパフォーマンスを達成しながら、有望な構成的で制御可能な検索機能を示し、ドメイン、タスク、およびユーザーの意図にわたる検索のための統合された音声埋め込みモデルとしての可能性を強調していることを示しています。

原文 (English)

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models

Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space. While effective, existing retrieval embeddings are primarily optimized for audio--caption matching, limiting their ability to support diverse retrieval objectives and controllable retrieval behaviors. We present ALM2Vec, a universal audio embedding framework derived from pretrained large audio--language models (LALMs). By transferring the audio understanding, instruction-following, and reasoning capabilities acquired through large-scale multimodal training, ALM2Vec learns a unified embedding space for retrieval across audio domains and task types. Beyond conventional text--audio retrieval, ALM2Vec incorporates natural-language instructions into the embedding process, enabling instruction-aware retrieval for scenarios such as audio question answering and aspect-conditioned retrieval. Experimental results show that ALM2Vec achieves competitive performance on standard audio and speech retrieval benchmarks while exhibiting promising compositional and controllable retrieval capabilities, highlighting its potential as a unified audio embedding model for retrieval across domains, tasks, and user intents.

13:00 JSTロボティクス

立場: 視覚-言語-行動モデルは物理的推論を実行することを検証できない

事前トレーニング済みビジョン言語モデル (VLM) に基づいて構築されたビジョン言語アクション (VLA) システムは、ロボット操作ベンチマークのパフォーマンスが急速に向上していることが示されています。これらの利点は、一般に、セマンティック表現がインターネット規模のデータ転送から物理的な実行の一般化まで学習した証拠として解釈されます。この意見書は、この解釈の基礎となる仮定、つまり物理的動作の決定をサポートするには意味論的な一般化で十分であるという仮定は独立して検証されておらず、現在の評価プロトコルの下ではテストできないと主張しています。私たちは、VLA ポリシーをセマンティック マッピングと物理的なアクションの決定に分解し、主要な評価指標であるタスクの成功率ではこれら 2 つの能力のソースを区別できないことを示すことで、この主張を支持します。その結果、ベンチマークのパフォーマンスの向上は、意味の一致、分布の重複、真の物理的一般化など、複数の競合する説明と一致します。我々はさらに、この識別可能性のギャップはナラティブ・ドリフトによって強化されており、それによって、後続のシステムは根底にある因果メカニズムを分離することなく、パフォーマンス向上の以前の解釈を継承および強化していると主張します。この制限に対処するために、意味論的一般化と物理的一般化を個別に測定するために制御された変動を導入する評価設計に基づいた研究の方向性を提案します。このような設計により、モデル内部へのアクセスを必要とせずにパフォーマンスを因果関係に帰属させることが可能になり、物理的能力の暗黙的なソースではなくセマンティック インターフェイスとしての VLM バックボーンの役割を経験的に評価することが可能になります。私たちの目標は、ロボット工学における VLM の役割を否定することではなく、物理的一般化の主張が有意義に評価できる条件を明らかにすることです。

原文 (English)

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning

Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot manipulation benchmarks. These gains are commonly interpreted as evidence that semantic representations learned from internet-scale data transfer to physical execution generalization. This position paper argues that the assumption underlying this interpretation -- that semantic generalization is sufficient to support physical action decisions -- has not been independently verified and cannot be tested under current evaluation protocols. We support this claim by decomposing VLA policies into semantic mapping and physical action decision, and showing that task success rate -- the dominant evaluation metric -- cannot distinguish between these two sources of capability. As a result, improvements in benchmark performance are consistent with multiple competing explanations, including semantic matching, distributional overlap, and genuine physical generalization. We further argue that this identifiability gap has been reinforced through narrative drift, whereby successive systems inherit and strengthen prior interpretations of performance gains without isolating the underlying causal mechanism. To address this limitation, we propose a research direction based on evaluation designs that introduce controlled variation to separately measure semantic and physical generalization. Such designs make it possible to causally attribute performance without requiring access to model internals, and to empirically assess the role of VLM backbones as semantic interfaces rather than implicit sources of physical competence. Our goal is not to refute the role of VLMs in robotics, but to clarify the conditions under which claims of physical generalization can be meaningfully evaluated.

13:00 JST研究/論文

分子拡散モデルの教師なし熱力学: 作用演算子の意味論と監査可能な自由エネルギーの読み出し

拡散モデルは、分子構造や立体構造アンサンブルのモデリングにますます利用されていますが、学習された表現やスコアの熱力学的な意味は依然としてとらえどころがありません。この曖昧さを解決するために、拡散モデルとネイティブ互換性のある数学的に一貫したアクション演算子フレームワークを導入します。固定分子環境を基本アクション $S_0(x)$ として定義し、錬金術的摂動を演算子 $O(x)$ として定義することにより、標準拡散ノイズは効果的なノイズ処理と演算子を誘発し、その勾配と錬金術的導関数はモデルの学習フィールドによって直接表現されます。この厳密な自己一貫性により、エンドポイント アンサンブルおよびフレームごとの評価から自由エネルギーの差 ($\Delta F$) を読み取ることができる「ノイジー オペレーター ブリッジ」が可能になります。アラニンジペプチド系の制御された実験では、物理的な誘導バイアスを組み込むことで塩基作用と摂動演算子の部分的な回復が可能になることを示しました。位相空間の重なりが無視できる困難な C6-H から C6-F のリガンド ポケットの非結合摂動 (185L/IND) に適用すると、教師ありブリッジは、安定した 19 状態 MBAR 参照の約 $1\ k_\mathrm{B}T$ 以内で錬金術 $\Delta F$ を推定します。最後に、力や動作の監視なしで、エンドポイントの座標とバイナリ ラベルだけで、オペレーターの形状と中心にある自由エネルギー スケールを部分的に復元するのに十分であることを示します。この研究は、生成分子拡散モデルをブラックボックス座標サンプラーから監査可能な熱力学推定装置に変換するための厳密な道筋を提供します。

原文 (English)

Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout

Diffusion models are increasingly utilized for modeling molecular structures and conformational ensembles, yet the thermodynamic meaning of their learned representations and scores remains elusive. To resolve this ambiguity, we introduce a mathematically consistent action-operator framework natively compatible with diffusion models. By defining a fixed molecular environment as a base action $S_0(x)$ and an alchemical perturbation as an operator $O(x)$, standard diffusion noising induces effective noised actions and operators whose gradients and alchemical derivatives are directly represented by the model's learned fields. This rigorous self-consistency enables a ``noisy operator bridge'' capable of reading out free-energy differences ($\Delta F$) from endpoint ensembles and per-frame evaluations. In controlled experiments on alanine dipeptide systems, we show that incorporating physical inductive biases enables partial recovery of the base action and perturbation operator. When applied to a challenging C6-H to C6-F ligand-pocket nonbonded perturbation (185L/IND) with negligible phase-space overlap, our supervised bridge estimates the alchemical $\Delta F$ within approximately $1\ k_\mathrm{B}T$ of a stable 19-state MBAR reference. Finally, we demonstrate that endpoint coordinates and binary labels alone are sufficient to partially recover the operator shape and a centered free-energy scale without any force or action supervision. This work provides a rigorous path toward transforming generative molecular diffusion models from black-box coordinate samplers into auditable thermodynamic estimators.

13:00 JST研究/論文

ノイズの多い等変量子ニューラル ネットワークにおけるトレーニング可能性のコヒーレンス則

対称性は量子ニューラル ネットワーク構造を提供しますが、それ自体では、ノイズが存在するとネットワークをトレーニング可能に保つことができません。等変回路の勾配がデコヒーレンスに耐えられるかどうかを決定する物理量を尋ねると、コンパクトなトレーニング則で答えます。電荷を保存する U(1) 等価レンガ積み回路を使用すると、2 つの異なる効果が訓練可能な勾配を支配することがわかります。因果関係は、勾配が存在できる場所を修正し、勾配をアクティブな電荷セクター内の読み出しの後方光円錐に限定します。次に、コヒーレンスは、投影された読み出し値が実際に観察できる非対角セクター モードの収縮を通じて、どのくらいの速さで減衰するかを決定します。総量子ビット数とは独立した下限でノイズのない勾配をセクター制限コーンに固定するライトコーン削減を証明し、読み出しに可視の整列コヒーレンス率を勾配搬送モードに沿ったノイズ発生器のレイリー商として定義します。摂動的なオープンシステム分析は、この速度を主要な訓練法則に変えます。次に、密度行列シミュレーションにより、有限ノイズの劣化が、ノイズの深さとコヒーレンスの収縮から構築された単一の累積変数に従うことが確認されます (決定係数は 0.979)。最も鮮明なテストは、最悪の場合のレートが大きいものの、ゼロに近い調整レートを持つ相関ディフェーズ チャネルから得られます。法則では、このチャネルの勾配損失はないと予測されており、何も見られません。セクター コヒーレンスは、比較したすべての標準的なチャネル診断よりも優れており、分析により、読み出しに見えるセクター コヒーレンスが、等変アーキテクチャ、オープン システム ダイナミクス、およびノイズの多いトレーニング可能性を結び付ける量として特定されます。

原文 (English)

A Coherence Law for Trainability in Noisy Equivariant Quantum Neural Networks

Symmetry provides a quantum neural network structure, but on its own it does not keep the network trainable once noise is present. We ask which physical quantity decides whether the gradients of an equivariant circuit survive decoherence, and we answer with a compact training law. Working with U(1)-equivariant brickwork circuits that conserve a charge, we find that two distinct effects govern a trainable gradient. Causality fixes where the gradient can live, confining it to the backward light cone of the readout inside the active charge sector. Coherence then determines how fast it decays through the contraction of the off-diagonal sector modes that the projected readout can actually observe. We prove a light-cone reduction that pins the noiseless gradient to the sector-restricted cone with a lower bound independent of the total qubit number, and we define a readout-visible aligned coherence rate as a Rayleigh quotient of the noise generator along the gradient-carrying mode. A perturbative open-system analysis turns this rate into a leading-order training law. Density-matrix simulations then confirm that the finite-noise degradation follows a single accumulated variable built from noise depth and coherence contraction, with a coefficient of determination of 0.979. The sharpest test comes from a correlated-dephasing channel that has a large worst-case rate but a near-zero aligned rate. The law predicts no gradient loss for this channel, and none is seen. Sector coherence outperforms every standard channel diagnostic we compare it against, and the analysis identifies readout-visible sector coherence as the quantity that links equivariant architecture, open-system dynamics and noisy trainability.

13:00 JSTLLM/生成AIハードウェア/半導体Claude

仕様駆動開発における引用規律: LLM で生成されたコードにおける出力決定論と自動幻覚検出に関するクロスモデル実証研究

仕様駆動開発 (SDD) フレームワークは、正式な仕様を通じて大規模言語モデル (LLM) を利用したコード生成をガイドしますが、要件と生成されたコードの間のトレーサビリティを強化する方法が根本的に異なります。この論文では、3 つの SDD フレームワークを比較する 2 つの管理された実証研究を紹介します。 $Spec Kit$ は、ユーザー ストーリーと受け入れ基準を通じて成果物レベルのトレーサビリティを使用します。 $OpenSpec$ はポストホック外部トレース マップに依存します。 2 つのフロンティア LLM、Claude Sonnet 4.6 (N=20、4 条件、240 実装) と GLM-5-turbo (N=50、4 条件、600 実装) にわたる 2 つの主な結果を測定します。$output$ $determinism$ (独立した LLM セッション間の語彙類似性) と $automated$ $hallucination$ $detection$ $rate$ (TDR) です。事前に登録された分析により、一貫したモデル間で反復されたトレードオフが明らかになりました。引用されていない条件は、引用された条件よりも大幅に高い決定論を生成します (Claude: $d=-0.76$、$p=0.003$; GLM: $d=-0.72$、$p<0.001$)。一方、引用された条件のみが自動幻覚検出を可能にします (TDR: Claude 86.4%、GLM) 88.0%、対すべての代替案で 0%、両方の研究で FPR=0%)。 traceSDD (引用) は、決定論に関して $Spec Kit$ を大幅に上回っています (Claude: $d=0.47$、$p=0.049$、GLM: $d=0.42$、$p=0.003$) が、OpenSpec ではありません (Claude: $d=0.18$、$p=0.44$、GLM: $d=0.14$、$p=0.32$)。これらの発見は、引用アノテーションが検証可能性を得るために決定性を犠牲にし、このトレードオフがモデル アーキテクチャ全体で一般化することを証明します。

原文 (English)

Citation Discipline in Spec-Driven Development: A Cross-Model Empirical Study of Output Determinism and Automated Hallucination Detection in LLM-Generated Code

Spec-Driven Development (SDD) frameworks guide Large Language Model (LLM)-powered code generation through formal specifications, yet they differ fundamentally in how they enforce traceability between requirements and generated code. This paper presents two controlled empirical studies comparing three SDD frameworks: $traceSDD$, which enforces mandatory per-line requirement citations using hierarchical REQ-XXX.Y.Z identifiers; $Spec Kit$, which uses artifact-level traceability through user stories and acceptance criteria; and $OpenSpec$, which relies on post-hoc external trace maps. We measure two primary outcomes across two frontier LLMs -- Claude Sonnet 4.6 (N=20, 4 conditions, 240 implementations) and GLM-5-turbo (N=50, 4 conditions, 600 implementations): $output$ $determinism$ (lexical similarity across independent LLM sessions) and $automated$ $hallucination$ $detection$ $rate$ (TDR). Our pre-registered analysis reveals a consistent, cross-model replicated trade-off: the uncited condition produces significantly higher determinism than the cited condition (Claude: $d=-0.76$, $p=0.003$; GLM: $d=-0.72$, $p<0.001$), while only the cited condition enables automated hallucination detection (TDR: Claude 86.4%, GLM 88.0%, vs 0% for all alternatives, FPR=0% across both studies). traceSDD (cited) significantly outperforms $Spec Kit$ on determinism (Claude: $d=0.47$, $p=0.049$; GLM: $d=0.42$, $p=0.003$) but not OpenSpec (Claude: $d=0.18$, $p=0.44$; GLM: $d=0.14$, $p=0.32$). These findings establish that citation annotations trade determinism for verifiability, and that this trade-off generalizes across model architectures.

13:00 JSTエージェントロボティクス

DSIP: 拡散モデルベースのマルチエージェント運動計画を使用した、信号のない交差点のための動的調整プランナー

都市部の交差点での信号制御は本質的にストップアンドゴー動作を招き、特に交通需要が高い場合に遅延が増加し、交通効率が低下します。コネクテッド自動運転車 (CAV) の出現により、軌道レベルの調整は、従来の段階ベースの管理を強化または超越する可能性の高い戦略として浮上しています。この論文は、生成拡散プロセスによって駆動されるマルチエージェント動作計画フレームワークである DSIP (拡散モデルベースのシグナルフリー交差点プランナー) を提案します。 DSIP は、交差点管理のパラダイムを離散的な時間的フェージングから連続的な複数車両の軌道最適化に移行します。この研究では、理想的な通信および実行条件下でのこの調整戦略の理論上の上限パフォーマンスを評価し、拡散主導型アプローチの核となる利点を分離します。 SUMO プラットフォームを使用して、さまざまな 4 脚交差点構成にわたる DSIP を評価します。実験結果は、DSIP が、特に中密度から高密度のトラフィックにおいて、固定時間信号制御と最先端の強化学習ベースのコントローラーの両方と比較して、平均遅延を大幅に削減し、より高い平均速度を維持することを示しています。これらの発見は、拡散ベースの軌道計画が将来の自律交差点管理のための拡張可能で堅牢な基盤を提供することを示唆しています。このアプローチは、ソフトウェア定義の調整を通じて潜在的な交差点容量を解放することにより、物理的なインフラストラクチャの拡張を必要とせずに都市交通の流れの効率を向上させるための費用対効果の高い経路を提供します。

原文 (English)

DSIP: A Dynamic Coordination Planner for Signal-Free Intersections using Diffusion-Model-Based Multi-Agent Motion Planning

Traffic signal control at urban intersections inherently introduces stop-and-go behavior, resulting in increased delays and reduced traffic efficiency, especially under high traffic demand. With the emergence of connected and automated vehicles (CAVs), trajectory-level coordination has emerged as a high-potential strategy to augment or transcend conventional phase-based management. This paper proposes DSIP (Diffusion-model-based Signal-free Intersection Planner), a multi-agent motion planning framework driven by a generative diffusion process. DSIP shifts the intersection management paradigm from discrete temporal phasing to continuous multi-vehicle trajectory optimization. This work evaluates the theoretical upper-bound performance of this coordination strategy under idealized communication and execution conditions to isolate the core benefits of the diffusion-driven approach. Using the SUMO platform, we evaluate DSIP across diverse four-leg intersection configurations. Experimental results demonstrate that DSIP significantly reduces average delay and maintains higher average speed compared to both fixed-time signal control and state-of-the-art reinforcement-learning-based controllers, particularly in medium- to high-density traffic. These findings suggest that diffusion-based trajectory planning provides a scalable and robust foundation for future autonomous intersection management. By unlocking latent intersection capacity through software-defined coordination, this approach offers a cost-effective pathway to improve urban traffic flow efficiency without requiring physical infrastructure expansion.

13:00 JST研究/論文

細胞周期を意識した単細胞薬物摂動応答のモデル化

単細胞薬物摂動モデルは、転写反応の大きさだけでなく、治療によって細胞の増殖状態が変化するかどうかも予測する必要があります。細胞周期の変動は迷惑な変動として扱われることが多く、ベンチマーク パイプラインでは薬剤誘発性の相変化を主な予測ターゲットとして扱うことはほとんどないため、これは困難です。 scCycleMol は、標準化された分子アイデンティティ、用量と細胞株のメタデータ、および治療状態から得られる細胞周期監視による遺伝子発現を備えた厳選された 24 時間の SciPlex3 ベンチマークに基づいて構築された、細胞周期を意識した摂動予測フレームワークです。 scCycleMol は、入力共変量として細胞周期状態を使用する代わりに、予測された処理発現から監視を導き出し、それを環状 G1/S/G2M 期ターゲットを備えた学習可能な完全発現細胞周期ヘッドを通じて伝播させます。マーカーベースの監視、分子表現、および事前トレーニング戦略を評価して、改善の原因を特定します。 600,000 個を超える細胞、186 の摂動条件、複数のがん細胞株、および数千の遺伝子を含む SciPlex3 ベンチマーク全体で、scCycleMol は条件付き摂動ベースラインと比較して分布外発現予測を向上させます。 LINCS で事前トレーニングされた最良の循環モデルは、LINCS で事前トレーニングされた ChemCPA の 0.6800 および 0.5400 と比較して、予想される全遺伝子の r の 2 乗が 0.9093、差次的に発現される遺伝子の 2 乗が 0.6843 と予想されます。クローズドループの細胞周期監視により、ほぼ変化しない発現予測を維持しながら、位相精度が約 0.5 ~ 0.6 ポイント向上します。 Tahoe で事前学習されたバリアントは位相精度 0.9609 に達し、摂動モデリングにおける細胞周期を意識した明示的な監視の利点を強調しています。

原文 (English)

Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses

Single-cell drug perturbation models should predict not only transcriptional response magnitude, but also whether a treatment alters the proliferative state of a cell. This is challenging because cell-cycle variation is often treated as nuisance variation, and benchmark pipelines rarely treat drug-induced phase changes as a primary prediction target. We introduce scCycleMol, a cell-cycle-aware perturbation prediction framework built on a curated 24-hour SciPlex3 benchmark with standardized molecule identities, dose and cell-line metadata, and gene expression with cell-cycle supervision derived from treated states. Instead of using cell-cycle state as an input covariate, scCycleMol derives supervision from predicted treated expression and propagates it through a learnable full-expression cell-cycle head with circular G1/S/G2M phase targets. We evaluate marker-based supervision, molecular representations, and pretraining strategies to isolate sources of improvement. Across a SciPlex3 benchmark with over 600k cells, 186 perturbation conditions, multiple cancer cell lines, and thousands of genes, scCycleMol improves out-of-distribution expression prediction compared with conditional perturbation baselines. The best LINCS-pretrained circular model achieves 0.9093 expected all-gene r squared and 0.6843 expected differentially expressed gene r squared, compared with 0.6800 and 0.5400 for LINCS-pretrained ChemCPA. Closed-loop cell-cycle supervision improves phase accuracy by about 0.5 to 0.6 points while maintaining nearly unchanged expression prediction. A Tahoe-pretrained variant reaches 0.9609 phase accuracy, highlighting the benefit of explicit cell-cycle-aware supervision in perturbation modeling.

13:00 JST画像/動画生成エージェント

LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents

Current operating systems expose interfaces optimized for human users but not for AI agents. Humans benefit from pixels, icons, windows, vi…

13:00 JST研究/論文

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-R…

13:00 JST研究/論文

An AI-Based Solution for Secure Service Provisioning in IoT

As the Internet of Things (IoT) continues its rapid expansion, the attack surface grows accordingly, with emerging threats targeting smart…

13:00 JST研究/論文

Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification

Structured tabular data dominates clinical medicine, yet existing benchmarks fail to reflect real-world properties like complex survey samp…

13:00 JSTLLM/生成AI

From Search to Synthesis: Training LLMs as Zero-Shot Workflow Generators

Large language models (LLMs) excel across a wide range of tasks, yet their instance-specific solutions often lack the structural consistenc…

13:00 JST研究/論文

Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts

Deterministic few-step generation succeeds on continuous image latents but collapses to incoherent text on continuous text latents, and we…

13:00 JST研究/論文

Hierarchical Global Attention (HGA)

Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preser…

13:00 JSTエージェントClaudeGPT / ChatGPT

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. Th…

13:00 JSTLLM/生成AIエージェント

A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization

Enterprise AI agents route user queries to specialized skills by matching queries against natural language skill descriptions. When two ski…

13:00 JST研究/論文

Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin

Audio deepfakes are a growing challenge for the general public, as well as for journalists and fact-checkers. The latter need reliable tool…

13:00 JSTLLM/生成AI

Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense

We identify a security-fidelity tradeoff in defending LLMs against indirect prompt injection: defenses resist injected instructions largely…

13:00 JSTLLM/生成AI研究/論文

Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions

Romanized Code Mixing (RCM), where bilingual speakers fluidly blend local languages with English in Roman script, has emerged as the domina…

13:00 JST研究/論文

Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization

Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that em…

13:00 JSTLLM/生成AIGPT / ChatGPT

When transformers learn "impossible" languages, what do they learn?

Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to…

13:00 JSTLLM/生成AI

AI-Generated PowerShell Malware: An Experimental Framework and Dataset

Generative AI has emerged as a significant cybersecurity threat, with several recent attack campaigns leveraging LLMs to generate code for…

13:00 JST研究/論文

A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection

The number of trees is a central computational parameter in Random Forests: increasing it reduces finite-ensemble variability but increases…

13:00 JSTLLM/生成AI

Test-Time Verification for Text-to-SQL via Outcome Reward Models

Improving the reliability of large language models (LLMs) at inference time is a central challenge in structured reasoning tasks such as Te…

13:00 JST画像/動画生成

The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning

Foundation model pseudo-labeling - labeling data strictly via zero-shot inference - enables massive scale, but performance is undermined by…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actiona…

13:00 JSTLLM/生成AILlama

Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified ma…

13:00 JST研究/論文Meta

How Human Feedback Shapes AI-generated Community Notes

Community Notes, a bridging-based crowd-sourced fact-checking system, has emerged as a new mechanism for moderating misleading information…

13:00 JST研究/論文

Budget-Adaptive Routing: Skipping the Weak When the Strong Answers Anyway

Edge-cloud inference collaborations are often designed with a routing estimator that decides whether to offload each frame from weak models…

13:00 JST研究/論文

Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback

Imitation Learning is a natural framework for learning in sequential decision-making systems and has emerged as the dominant paradigm throu…

13:00 JST研究/論文

Physics-informed Conditional Normalizing Flows for Angles-only Cislunar Orbit Determination

Generative Astrodynamics is advanced in this work by extending generative modelling to an orbit determination problem in the cislunar envir…

13:00 JSTロボティクス

Motion Planning in Compressed Representation Spaces

Deep learning methods have vastly expanded the capabilities of motion planning in robotics applications, as learning priors from large-scal…

13:00 JST画像/動画生成

Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-Ultrasound Prostate Cancer Detection

Micro-ultrasound ($\mu$US) is a new, emerging, and promising imaging modality for prostate cancer (PCa) detection, but accurate identificat…

13:00 JSTLLM/生成AI

Loc2Repair: A Framework for Evaluating the Impact of File-Level Issue Localization in Repo-Level LLM Repair

Repository-grounded automated repair is often reported as a single end-to-end capability, which hides distinct failure modes such as poor f…

13:00 JSTLLM/生成AI研究/論文

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language m…

13:00 JST研究/論文Qwen

OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models

We propose OTCache, a training-free framework for accelerating diffusion sampling via caching schedule prediction. Existing graph-based cac…

13:00 JSTLLM/生成AI

LLM-Driven Personalities for Decision Making in Emergency Simulations

For virtual humans to appear believable, they must exhibit agency and spatial awareness while interacting with their environment in ways th…

13:00 JST研究/論文DeepSeek

Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition

This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.5-7B). Using hi…

13:00 JST画像/動画生成

Learning Video Dynamics with Predictive Differentiable Rendering

How to accurately predict a high-fidelity future world? While the visual world is inherently continuous, existing deterministic video predi…

13:00 JSTLLM/生成AI画像/動画生成

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs

Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image.…

13:00 JSTLLM/生成AI

Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks

Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics. In this study, we propose a novel T…

13:00 JST研究/論文

Beyond But-for Test: Counterfactual Explanation in Abstract Argumentation via Actual Causality (Extended Version)

Counterfactual explanation in abstract argumentation calls for an answer to the what-if query: would the topic argument still be accepted i…

13:00 JSTLLM/生成AI

When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking

Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying t…

13:00 JST画像/動画生成GPT / ChatGPT

Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation

Recent years have seen substantial advances in radiology report generation (RRG), yet existing approaches predominantly adopt direct featur…

13:00 JSTエージェントロボティクス

What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning

Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance…

13:00 JST画像/動画生成

SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos

To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or mu…

13:00 JSTLLM/生成AI

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus…

13:00 JSTエージェントロボティクス研究/論文

A Modular Vision-Language-Action Robotics Framework for Indoor Environments

This paper presents an integrated system for the CMU Vision-Language-Action (VLA) Challenge, designed to enable an autonomous agent to perf…

13:00 JSTLLM/生成AI画像/動画生成

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding

3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions. Existing approaches typically…

13:00 JSTエージェント研究/論文ClaudeMicrosoft

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal…

13:00 JST研究/論文

One Retrieval to Cover Them All: Co-occurrence-Aware Knowledge Base Reorganization for Session-Level RAG

RAG systems retrieve documents optimized for answering one query at a time. Yet enterprise users arrive with sessions, that is, coherent ep…

13:00 JSTLLM/生成AIロボティクス

LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music

The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics. Traditional methods often re…

13:00 JSTLLM/生成AI

ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries

Large language models deployed in regulated industries operate under two constraints: compliance enforcement and cost efficiency. Personall…

13:00 JSTエージェントロボティクス

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However…

13:00 JST研究/論文

AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL

Optimizing nonlinear preferences in multi-objective reinforcement learning (MORL) is essential for capturing complex trade-offs like risk a…

13:00 JST研究/論文

Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation

Adaptive experiments for average treatment effects (ATE) require randomized allocations balancing valid inference with statistical efficien…

13:00 JSTLLM/生成AI

Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection

Spontaneous speech is a vital non-invasive biomarker for Alzheimer's Disease (AD), yet many systems overlook non-linear structural disrupti…

13:00 JST画像/動画生成

Distilling Temporal Coherence into 2D Networks for Transrectal Ultrasound Prostate Video Segmentation

Real-time video segmentation of the prostate in Transrectal Ultrasound (TRUS) is essential for image-guided interventions. While convention…

13:00 JSTLLM/生成AIエージェント

Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?

As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing…

13:00 JSTエージェントロボティクス

Information-Aided DVL Calibration

The Doppler velocity log (DVL) velocity measurements are critical to the accuracy of autonomous underwater vehicle (AUV) navigation solutio…

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達規制/政策Llama

Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law

Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards foc…

13:00 JST研究/論文

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative mul…

13:00 JSTハードウェア/半導体

TDGT: A Tabular Data Generation Toolkit supporting adaptive GPU-accelerated Bayesian mixture models, diffusion-based models, and latent-space generative modeling

The growing demand for privacy-preserving data sharing has positioned synthetic data generation as a critical component of responsible AI w…

13:00 JSTLLM/生成AI画像/動画生成エージェント

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted signif…

13:00 JST画像/動画生成

CLIMB: Centroid-Based Hierarchical Memory for Online Continual Self-Supervised Learning

Online Continual Self-Supervised Learning (OCSSL) aims to learn representations from a continuous stream of unlabeled data, without knowled…

13:00 JST研究/論文

Minimizing Quantized Semantic Age of Information (QSAoI) in Foundation Model-Based Semantic Communications

The emerging techniques of semantic communications and edge computing in 6G networks necessitate a paradigm shift toward co-designed semant…

13:00 JSTLLM/生成AI

CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs

While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity o…

13:00 JST研究/論文

From Idea to Prototype in an Afternoon: Scaffolded, AI-Assisted Rapid VA Prototyping

Testing a new visual-analytics idea usually takes months: one needs to find a realistic data set, clean it, and implement an interactive pr…

13:00 JSTロボティクス

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot man…

13:00 JST研究/論文

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this r…

13:00 JST研究/論文

PGUDA: Pressure-Guided Unsupervised Domain Adaptation with Cross-Modal Knowledge Distillation for sEMG-Based Gesture Recognition

Surface electromyography (sEMG)-based gesture recognition has emerged as a promising technology for natural human-computer interaction. How…

13:00 JST研究/論文

From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation

Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unpr…

13:00 JSTLLM/生成AIエージェントDeepSeek

Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?

When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agen…

13:00 JSTロボティクス

Stage-Transition Dense Reward Modeling for Reinforcement Learning

Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense…

13:00 JST画像/動画生成

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

Artificial intelligence is transforming our capability to solve biological challenges. In dimensionality bottleneck regimes exacerbated by…

13:00 JST研究/論文

Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models

State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls…

13:00 JSTLLM/生成AI画像/動画生成

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based me…

13:00 JST画像/動画生成

Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors

Single-stage video object detectors are increasingly deployed in time-critical applications, yet it remains unclear whether these models ge…

13:00 JSTエージェント

DA-Studio: An Agentic System for End-to-End Data Analysis

Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system…

13:00 JSTロボティクス

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, e…

13:00 JSTLLM/生成AI

Team MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline Dynamics

Recent advances in Large Language Models (LLMs) have motivated their adoption across a wide range of domains, including Artificial Intellig…

13:00 JST研究/論文

Von Mises Based Uncertainty Quantification for Closely Spaced Automotive Radar Targets

This work investigates uncertainty-aware deep learning approaches for direction of arrival (DOA) estimation in automotive radar, focusing o…

13:00 JSTロボティクス

Robustness of Robotic Manipulation: Foundations and Frontiers

Humans and animals exhibit remarkable robustness in physical manipulation, yet robots remain far behind. Progress toward human-level manipu…

13:00 JSTLLM/生成AIエージェント研究/論文

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as…

13:00 JSTLLM/生成AI

On the Convergence of Self-Improving Online LLM Alignment

The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient…

13:00 JST研究/論文

Improving multichannel speech enhancement through accurate room-acoustic simulations

Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement. While most pipelines rely on…

13:00 JST研究/論文

CVE-TTP KG: Knowledge Graph Linking Software Vulnerabilities to Attack Behaviors

In the evolving threat landscape, adversaries exploit software vulnerabilities to launch sophisticated attacks, challenging traditional def…

13:00 JST研究/論文

FLARE-AI: Flaw Reporting for AI

Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosyste…

13:00 JST画像/動画生成

Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning

Masked autoencoding has emerged as a prominent paradigm for self-supervised learning on 3D point clouds, achieving competitive performance…

13:00 JST画像/動画生成

Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer

Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service l…

13:00 JST画像/動画生成

DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers

The remarkable scalability of Transformers has expanded their application to 3D computer vision, where camera-aware positional encoding is…

13:00 JSTLLM/生成AI

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

Audio-Language Models (ALMs) achieve strong zero-shot performance by aligning audio with textual class descriptions. Although prompt learni…

13:00 JSTLLM/生成AI

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment

Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insec…

13:00 JST研究/論文

Comparative Analysis of Machine Learning based Intrusion Detection in Realistic IoT Networks

The Internet of Things (IoT) is rapidly growing and expanding into various sectors, such as healthcare, transportation, smart homes, and mo…

13:00 JST画像/動画生成

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle…

13:00 JST画像/動画生成

Preserve the Hard, Regenerate the Rest: Uncertainty-Guided Synthetic Training Data Augmentation with Diffusion Models

Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.g., dense regions or small objects in aeri…

13:00 JST画像/動画生成

Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation

Radar sensors provide reliable perception under adverse weather and lighting conditions, but their sparse, noisy, and weakly semantic measu…

13:00 JST研究/論文

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

Engineering specifications such as interlocks, alarm rationalization tables, and cause-and-effect (C&E) matrices remain central to process…

13:00 JSTLLM/生成AIエージェント

A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents

Fault recovery in process plants still relies heavily on plant operators, especially when faults fall outside predefined supervisory logic.…

13:00 JST研究/論文

Intrinsic decomposition and editing of 3D Gaussian splats

Intrinsic decomposition which expresses image colors as the product of diffuse albedo and shading, possibly augmented with view-dependent r…

13:00 JST研究/論文

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, cod…

13:00 JSTエージェント

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL

Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Ex…

13:00 JST研究/論文

Improving Certified Robustness via Adversarial Distillation

Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimi…

13:00 JST画像/動画生成

Sparsity-Inducing Divergence Losses for Biometric Verification

Performance in face and speaker verification is largely driven by margin-penalty softmax losses such as CosFace and ArcFace. Recently intro…

13:00 JST画像/動画生成研究/論文

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignor…

13:00 JST画像/動画生成

Histogram-constrained Image Generation

Diffusion models have emerged as a dominant paradigm in generative modeling, enabling high-fidelity sampling from complex data distribution…

13:00 JST研究/論文

When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection

Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret. Variables are fir…

13:00 JSTLLM/生成AIエージェント

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping

The wave of AI-native applications is moving shopping beyond page- and feed-based browsing toward intent-driven experiences orchestrated by…

13:00 JSTLLM/生成AI画像/動画生成ロボティクス

RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization

For robots manipulating open-world objects, tactile representations must generalize to unseen materials. We introduce RCT (Robotic Contact…

13:00 JST画像/動画生成

Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models

Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that iso…

13:00 JSTLLM/生成AIビジネス/資金調達Gemma

Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian

Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotated corpora. We investigate the feasibili…

13:00 JSTLLM/生成AI

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through int…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

STEB: Style Text Embedding Benchmark

While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains frag…

13:00 JST研究/論文

FedXDS: Leveraging Model Attribution Methods to counteract Data Heterogeneity in Federated Learning

Explainable AI (XAI) methods have demonstrated significant success in recent years at identifying relevant features in input data that driv…

13:00 JST画像/動画生成研究/論文

JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering

Remote sensing change detection (CD) traditionally focuses on pixel-level binary segmentation, which identifies where changes occur but nei…

13:00 JST研究/論文

A Technical Typology of AI Systems in Public Administration

Research on artificial intelligence (AI) in the public sector often treats "AI" as a single category, neglecting technical distinctions bet…

13:00 JSTLLM/生成AI

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield

We study three complementary techniques for training compute-efficient language models. (1) Selective supervision and per-token efficiency.…

13:00 JST研究/論文

Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR

Low-rank adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tunin…

13:00 JST画像/動画生成

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain p…

13:00 JST画像/動画生成エージェントロボティクス

Real-Time Source-Free Object Detection

Real-world detectors for autonomous driving, surveillance, and robotics must handle domain-shifts under strict latency and memory constrain…

13:00 JSTロボティクス

Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in global…

13:00 JSTロボティクス

Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models

Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observ…

13:00 JST研究/論文

Belief Contraction in Dynamic Epistemic Logic

Dynamic epistemic logic represents belief change via model transformations induced by epistemic events. Its standard formulation (Baltag, M…

13:00 JST研究/論文

Modal CEGAR-tableaux with RECAR and resolution-based SAT-shortcuts

We investigate two approaches for extending CEGAR-tableaux with SAT-shortcuts using a previously known approach called RECAR but also a tot…

13:00 JST研究/論文

Better Understanding, Understanding Better

"Any fool can know; the point is to understand." A well-known remark often attributed to Einstein captures a widely shared intuition: under…

13:00 JSTLLM/生成AI画像/動画生成

Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference

Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Ex…

13:00 JST画像/動画生成エージェントロボティクス

MVP-Nav: Multi-layer Value Map Planner Navigator

Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of expli…

13:00 JSTロボティクス

LeCropFollow: Latent Space Planning for Navigation in Unstructured Crop Fields

Unstructured navigational features, such as irregular planting or discontinuities, remain the primary failure mode for under-canopy agricul…

13:00 JSTLLM/生成AI画像/動画生成エージェント

MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grou…

13:00 JST画像/動画生成

LUNA: Learning Universal 3D Human Animation Beyond Skinning

Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and paramet…

13:00 JST研究/論文

GR2 Technical Report

Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking --…

13:00 JST規制/政策

Amplifying Membership Signal Through Chained Regeneration

The tendency of large generative models to memorize training data makes sample verification critical for privacy auditing and copyright enf…

13:00 JST研究/論文

Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization

Why do neural networks memorize algorithmic training data long before they generalize? We present a geometric case study demonstrating that…

13:00 JSTエージェント

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands…

13:00 JST画像/動画生成

FLORA: A deep learning approach to predict forest attributes from heterogeneous LiDAR data

Forest attributes are essential for national-scale resource monitoring. Airborne LiDAR metrics are among the auxiliary variables most stron…

13:00 JST研究/論文

AdaJEPA: An Adaptive Latent World Model

Latent world models enable planning from high-dimensional observations by predicting future states in a compact latent space. However, thes…

13:00 JSTエージェントロボティクス

Freeform Preference Learning for Robotic Manipulation

Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where spa…

13:00 JSTLLM/生成AI

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e., incorrectly citing or…

13:00 JSTLLM/生成AI

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes. Yet…

13:00 JSTLLM/生成AIエージェント

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings,…

13:00 JSTLLM/生成AI

Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision

When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, rather than superficia…

13:00 JST研究/論文GPT / ChatGPT

Disentangling Reasoning Logic to Resolve Explicit Knowledge Conflicts

Explicit knowledge conflicts, occurring when retrieved contexts contain contradictory information, pose a fundamental challenge for Large L…

13:00 JST研究/論文

A Concept of Possibility for Real-World Events

This paper offers a new concept of {\it possibility} as an alternative to the now-a-days standard concept originally introduced by L.A. Zad…

13:00 JSTLLM/生成AI

Deductive Logic in Language Models: Horizontal vs Vertical Reasoning

Recent language models exhibit significant logical reasoning abilities, yet the mechanisms supporting deductive inference remain poorly und…

13:00 JSTLLM/生成AIエージェント

LLM-Empowered Agentic MAC Protocols: A Dynamic Stackelberg Game Approach

Medium Access Control (MAC) protocols, essential for wireless networks, are typically manually configured. While deep reinforcement learnin…

13:00 JSTLLM/生成AI

Improving LLM Reasoning with Homophily-aware Structural and Semantic Text-Attributed Graph Compression

Large language models (LLMs) have demonstrated promising capabilities in Text-Attributed Graph (TAG) understanding. Recent studies typicall…

13:00 JSTエージェント研究/論文

Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance

Writing effective rebuttals is a high-stakes task that demands more than linguistic fluency, as it requires precise alignment between revie…

13:00 JSTLLM/生成AIエージェント研究/論文

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus pri…

13:00 JST研究/論文

Improved Upper Bounds for Slicing the Hypercube

A collection of hyperplanes $\mathcal{H}$ slices all edges of the $n$-dimensional hypercube $Q_n$ with vertex set $\{-1,1\}^n$ if, for ever…

13:00 JST画像/動画生成エージェント

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation

Large vision-language models have endowed GUI agents with strong general capabilities for interface understanding and interaction. However,…

13:00 JST研究/論文

Diffusion Crossover: Defining Evolutionary Recombination in Diffusion Models via Noise Sequence Interpolation

Interactive Evolutionary Computation (IEC) provides a powerful framework for optimizing subjective criteria such as human preferences and a…

13:00 JSTLLM/生成AIエージェント研究/論文Claude

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research…

13:00 JSTエージェント

Containment Verification: AI Safety Guarantees Independent of Alignment

Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and ther…

13:00 JSTLLM/生成AIビジネス/資金調達ClaudeDeepSeek

LLM における推論の質の測定: 多次元の行動フレームワーク

LLM は複雑な推論タスクで目覚ましい成功を収めていますが、現在の評価アプローチは主に最終的な答えの正しさに依存しており、それらの答えを生み出す根本的な推論プロセスについての洞察は限られています。このギャップに対処するために、この研究では、動作の観点から LLM の推論品質を測定するための統一された多次元フレームワークを提案し、理論的に根拠のある 6 つの次元、正確性 (CQ)、一貫性 (CS)、堅牢性 (RS)、論理的一貫性 (LS)、効率 (ES)、安定性 (SS) を運用します。 4 つのベンチマークの 975 項目にわたる 7 つの LLM に関する広範な実験により、このフレームワークが精度のみの指標では見えない動作を明らかにすることが実証されました。特に、論理的一貫性は正しさ (r = -0.172、ns) と直交しており、一貫性のない推論から正しい答えが得られることが確認され、一方、Claude-Haiku-4.5 は最高の多次元スコア (Q_bal = 0.778) を達成しています。さらに、このフレームワークは重大なランキングの逆転を明らかにしています。DeepSeek-V3 は精度優先では 2 位ですが、法的/コンプライアンスの重み付けでは 5 位にランクされており、単一指標の評価では検出できない逆転です。判別式の妥当性により、11/15 次元のペアが独立している (|r| < 0.50) ことが確認され、各次元を別個の信号として扱うための心理測定的サポートが提供されます。フレームワークによって生成される次元プロファイルは、次の 3 つのクラスの展開決定を直接サポートします。最終的な答えが正しいにもかかわらず、その推論トレースが説明責任監査に失敗するモデルを特定します (LS--CQ 直交性)。精度のみのベンチマークによって引き起こされるランキングエラーを防止します。そして、フレームワークがキャプチャする 6 つの独立したシグナルを単一のメトリックが暗黙的に置き換えることがないようにします。

原文 (English)

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offering limited insight into the underlying reasoning processes that produce those answers. To address this gap, this study proposes a unified multi-dimensional framework for measuring reasoning quality in LLMs from a behavioral perspective, operationalizing six theoretically grounded dimensions: Correctness (CQ), Consistency (CS), Robustness (RS), Logical Coherence (LS), Efficiency (ES), and Stability (SS). Extensive experiments on seven LLMs across 975 items from four benchmarks demonstrate that the framework reveals behaviors invisible to accuracy-only metrics. Notably, logical coherence is orthogonal to correctness (r = -0.172, ns), confirming that correct answers can arise from incoherent reasoning, while Claude-Haiku-4.5 achieves the highest multi-dimensional score (Q_bal = 0.778). Furthermore, the framework exposes critical ranking inversions: DeepSeek-V3 ranks second under accuracy-priority but fifth under legal/compliance weighting, a reversal that single-metric evaluation cannot detect. Discriminant validity confirms 11/15 dimension pairs are independent (|r| < 0.50), providing psychometric support for treating each dimension as a distinct signal. The dimensional profiles produced by the framework directly support three classes of deployment decision: identifying models whose reasoning traces would fail accountability audits despite correct final answers (LS--CQ orthogonality); preventing ranking errors caused by accuracy-only benchmarking; and ensuring that no single metric silently substitutes for the six independent signals the framework captures.

13:00 JST研究/論文

線形時間の時間的解答セットプログラミングのためのメタプログラミング

Answer Set Programming (ASP) の時間的拡張の開発により、非単調線形時間 (TEL)、動的 (DEL)、およびメトリック (MEL) の時間平衡ロジックが出現しました。ただし、高度に最適化された ASP システムに固有の剛性により、代替論理設計の迅速な探索と実装が妨げられることがよくあります。この研究では、統一された宣言型フレームワークを通じてさまざまな時相論理のセマンティクスを操作できる柔軟なメタプログラミング フレームワークを提案します。私たちのアプローチは、 clingo の理論文法を形式的な型仕様とネスト機能で強化することにより、標準 ASP メタプログラミングを拡張します。セマンティックな正確性を確保するために、グラウンディング中の安定モデルベースの単純化からネストされたモダリティを保護する変換パイプラインを導入します。 TEL、MEL、および DEL のメタエンコーディングを実装することにより、フレームワークの拡張性を示します。 TEL の包括的な説明を提供し、MEL の間隔制約と DEL のフィッシャー・ラドナー閉包を管理するための主要な機能に焦点を当てます。最後に、このワークフローをカプセル化する多用途ツール、metasp システムを紹介します。

原文 (English)

Meta-Programming for Linear-time Temporal Answer Set Programming

The development of temporal extensions of Answer Set Programming (ASP) has led to the emergence of non-monotonic linear-time (TEL), dynamic (DEL), and metric (MEL) temporal equilibrium logics. However, the inherent rigidity of highly optimized ASP systems often hinders the rapid exploration and implementation of alternative logical designs. In this work, we propose a flexible meta-programming framework that operationalizes the semantics of varied temporal logics through a unified, declarative framework. Our approach extends standard ASP meta-programming by augmenting clingo's theory grammar with formal type specifications and nesting capabilities. To ensure semantic correctness, we introduce a transformation pipeline that protects nested modalities from stable-model-based simplifications during grounding. We demonstrate the extensibility of our framework by implementing meta-encodings for TEL, MEL, and DEL. We provide a comprehensive account of TEL and highlight the key features for managing the interval constraints of MEL and the Fischer-Ladner closure in DEL. Finally, we introduce the metasp system, a versatile tool that encapsulates this workflow.

13:00 JSTLLM/生成AI

壊滅的な状態にある MDP におけるベルマン最適性からのプロスペクト理論の動作

私たちは、破滅的な状態を吸収するマルコフ意思決定プロセスにおけるリスク中立制御を研究します。報酬は線形であり、エージェントに効用曲率、確率重み付け、フレーミング依存性がないにもかかわらず、標準的なベルマン最適性は 3 つのプロスペクト理論のようなシグネチャを生成します。S 字型の価値関数プロファイル (大惨事付近では凸、遠方場では凹)、内生的損失感度係数 $\lambda^*(S) > 1$、および反射効果ポリシーの逆転です。 495 の構成全体で、最適な政策は、リスクのあるアクションの即時期待値が高いにもかかわらず、ポジティブ ドリフト (成長) レジームでは大惨事近くで安全な役割を果たし、ネガティブ ドリフト (衰退) レジームでは、安全なアクションの即時期待損失が低いにもかかわらず、大惨事近くで危険な役割を果たします。勝利確率 $p$、ペイオフの非対称性 $r = |\Delta_\ell/\Delta_w|$、および割引係数 $\beta$ のみに依存し、数値解を $R^2 = 0.999$ に一致させる漸近損失回避プラトー $\bar{\lambda}$ の閉形式式を導出します。このメカニズムは非対称的なペイオフを必要としません。 3 つの非対称レベルで $(p,\beta)$ をスイープすると、1 を超える $\bar{\lambda}$ の非対称割合は、$r = 1.25$ で中央値 4.6%、$r = 2$ で 13.9% に上昇し、テストしたすべてのセルで境界寄与が非対称寄与を上回りました。この現象は、表形式の Q 学習 (モデルフリー エージェントは、相関関係 0.98 の成長と 1.00 の衰退で $V^*$ を再現します) およびガウス、ヘビーテール スチューデント $t_3$、およびステップ サイズの最大 50% までの非対称スキュー法線ノイズを伴う確率的遷移下で持続します。漸近プラトーはセーフ チャネルの 0.41% 以内で閉形式予測を追跡します。ノイズ、および危険なチャネルまたは両方のチャネルのノイズが 9.6% 以内であること。これらの結果は、故障状態の吸収が、最適な制御下での見通し理論のような動作を実現するための十分な構造メカニズムであることを特定します。

原文 (English)

Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States

We study risk-neutral control in Markov decision processes with an absorbing catastrophic state. Even though rewards are linear and the agent has no utility curvature, probability weighting, or framing dependence, standard Bellman optimality produces three prospect-theory-like signatures: an S-shaped value-function profile (convex near catastrophe, concave in the far field), an endogenous loss-sensitivity coefficient $\lambda^*(S) > 1$, and a reflection-effect policy reversal. Across 495 configurations, the optimal policy plays safe near catastrophe in positive-drift (growth) regimes despite the risky action's higher immediate expected value, and plays risky near catastrophe in negative-drift (decline) regimes despite the safe action's lower immediate expected loss. We derive a closed-form expression for the asymptotic loss-aversion plateau $\bar{\lambda}$ that depends only on win probability $p$, payoff asymmetry $r = |\Delta_\ell/\Delta_w|$, and discount factor $\beta$, and matches numerical solutions to $R^2 = 0.999$. The mechanism does not require asymmetric payoffs. Across a sweep of $(p,\beta)$ at three asymmetry levels, the asymmetry share of $\bar{\lambda}$ above unity has median 4.6% at $r = 1.25$ and rises to 13.9% at $r = 2$, with the boundary contribution exceeding the asymmetry contribution in every cell tested. The phenomena persist under tabular Q-learning (a model-free agent reproduces $V^*$ at correlation 0.98 in growth and 1.00 in decline) and under stochastic transitions with Gaussian, heavy-tailed Student-$t_3$, and asymmetric skew-normal noise up to 50% of the step size, where the asymptotic plateau tracks the closed-form prediction within 0.41% for safe-channel noise and within 9.6% for risky-channel or both-channel noise. These results identify absorbing failure states as a sufficient structural mechanism for prospect-theory-like behavior under optimal control.

13:00 JST研究/論文

Reasoning4Sciences: 推論言語モデルをすべての科学分野に橋渡しする

推論言語モデル (RLM) は科学研究のための強力なツールとして急速に台頭していますが、その影響は主に「ハード サイエンス」分野に集中しています。他の科学分野での RLM の導入が遅い、または導入されていないことが、研究の生産性の差の拡大を引き起こしています。この調査では、欧州研究評議会 (ERC) が使用する社会科学と人文科学、物理科学と工学、生命科学にわたる分類に従って、28 の科学分野にわたる RLM の採用に関する初めての包括的な分析を提供します。私たちは、RLM がどのように開発、評価され、分野全体に適用されるかを調査します。さらに、利用可能なドメイン固有の開発および評価リソースに基づいた成熟度指向の評価フレームワークを導入し、公開されているリソースのみを考慮した場合にさらに顕著になる RLM 成熟度の実質的な格差を明らかにします。最後に、分野を超えて普及しつつある現在の実装パラダイム、現在の課題、科学全体で RLM の導入を可能にする将来の方向性を強調します。

原文 (English)

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrated in "hard science" fields. The slow -- or lack of -- adoption of RLMs in other branches of science is causing a widening gap in research productivity. In this survey, we provide the first comprehensive analysis of RLM adoption across 28 scientific disciplines following the classification used by the European Research Council (ERC), spanning the Social Sciences and Humanities, Physical Sciences and Engineering, and Life Sciences. We examine how RLMs are developed, evaluated, and applied across disciplines. Furthermore, we introduce a maturity-oriented assessment framework based on available domain-specific development and evaluation resources, revealing substantial disparities in RLM maturity that become even more pronounced when only publicly available resources are considered. Finally, we highlight current implementation paradigms that are gaining popularity across disciplines, current challenges, and future directions in enabling RLM adoption across science.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

ギャンブルはしないでください、GAMBLe: AI 主導の研究システムのための分析フレームワーク

AI-Driven Research Systems (ADRS) -- LLM と自動評価を組み合わせてアルゴリズム、証明、設計を発見するシステム -- は最適化され、ドメイン全体で採用されていますが、それらを分析するツールは追いついていません。 ADRS のパフォーマンスはコンポーネントの相互作用に依存しますが、これらの相互作用は十分に理解されておらず、調査にコストがかかり、(ここで示しているように) 標準の収束保証では十分に把握されていません。これらの保証は、私たちが形式化した ADRS プロセスの下では成立しない構造的な仮定に依存しています。我々は、ADRS の動作を 4 つのパラメーター (ジェネレーター $G$、アセッサー $\mathcal{A}$、発見メカニズム $\mathcal{M}$、バジェット $B$) と 1 つの構成オブジェクト、効果的なランドスケープ $L_{\text{eff}} = \mathcal{A} \circ G$ に分解するフレームワークである GAMBLe を紹介します。これにより、異なるジェネレーターとアセッサーのペアが構造的に異なる問題ごとの最適化を引き起こすことが明らかになります。風景。私たちは、単一の LLM から動的適応アンサンブルに至るジェネレーター、貪欲な選択から共進化メタサーチに至るメカニズム、および評価者が連続スコアリングからクリフ関数に及ぶ 3 つの NP 困難問題に及ぶ 760 以上の反復実行 (>46,000 反復) でフレームワークを実行します。実験では、ジェネレーターやメカニズムの完全な順序付けは明らかにされていません。フロンティア モデルはオープンソースの代替モデルよりもパフォーマンスが劣る可能性があり、最も単純なメカニズムが最先端のメタ検索を上回る場合もあります。結果は、限られた予算 (実行ごとに 60 回の反復) の下でも、適切なコンポーネントを選択することでパフォーマンスを 13 ~ 67%、検索効率を 6 ~ 39 倍改善できることを示しています。

原文 (English)

Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems

AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace. ADRS performance depends on component interactions that are poorly understood, expensive to explore, and (as we show) not well captured by standard convergence guarantees. These guarantees rely on structural assumptions that do not hold under the ADRS process we formalize. We introduce GAMBLe, a framework that decomposes ADRS behavior into four parameters (generator $G$, assessor $\mathcal{A}$, discovery mechanism $\mathcal{M}$, budget $B$) and one compositional object, the effective landscape $L_{\text{eff}} = \mathcal{A} \circ G$, which reveals that distinct generator-assessor pairs induce structurally different per-problem optimization landscapes. We exercise the framework on 760+ replicated runs (>46,000 iterations) spanning generators from single LLMs to dynamically-adaptive ensembles, mechanisms from greedy selection to co-evolutionary meta-search, and three NP-hard problems whose assessors range from continuous scoring to cliff functions. The experiments reveal no total ordering of generators or mechanisms: frontier models can underperform open-source alternatives and the simplest mechanism sometimes outperforms state-of-the-art meta-search. Results show that even under limited budgets (60 iterations per run), the right component choices can improve performance by 13-67% and search efficiency by 6-39x.

13:00 JST研究/論文

いつ再計画するか: 階層的潜在推論におけるサブゴールの永続性

長期的な推論では、システムが硬直化することなく中期的な目的にコミットする必要があります。再計画が頻繁に行われすぎると、計算が複数ステップの構造にまとまることはありません。コミットが長すぎると計画が古くなってしまいます。私たちは、この安定性と適応性のトレードオフを潜在推論設定で研究します。この設定では、複数ステップの計算が外部化されたトークン トレースではなく隠れた状態の内部で発生します。私たちは、階層的推論モデル (HRM) を封建的なスタイルのマネージャーとワーカーのインターフェイスで拡張します。遅い高レベルのモジュールは、P 個の低レベル ステップの間持続する正規化された方向サブゴールを定期的に発行し、ワーカーの隠れ状態の更新にバイアスをかけ、固有のコサイン アラインメント損失を提供します。 ARC と ConceptARC では、サブゴールの持続性 (サブゴールの注入だけではなく) が中心のノブであることが分かりました。[3, 6] の中程度の期間 P は、非常に頻繁な (P=1) と非常に長い期間の両方を一貫して上回っており、P=3 で明らかに最小の LM 損失が見られます (P=1、1.640 ベースラインで 1.544 対 1.674、平均 1.595、標準値で 5 つのシードで複製) 0.045)。固有のアライメント重みラムダは、相補的な狭い最適値 (ラムダ約 0.05) を示します。過去のスイートスポットラムダでの制御されたアブレーションは、アライメント信号が最適値を超えたときに、アーキテクチャ上の容量や補助損失だけではなく、学習された指向性構造を干渉源として分離します。これらの発見を総合すると、潜在推論システムにおける構成計画の設計原則が示唆されます。つまり、中程度の地平線の意図は、構成構造を形成するのに十分な計算ステップにわたって首尾一貫していなければなりません。

原文 (English)

When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning

Long-horizon reasoning requires a system to commit to medium-horizon intent without becoming rigid: re-plan too often and computation never coheres into multi-step structure; commit too long and the plan goes stale. We study this stability-adaptivity tradeoff in the latent reasoning setting, where multi-step computation occurs inside hidden state rather than externalized token traces. We extend the Hierarchical Reasoning Model (HRM) with a feudal-style manager-worker interface: a slow high-level module periodically emits a normalized directional subgoal that persists for P low-level steps, biasing the worker's hidden-state updates and supplying an intrinsic cosine alignment loss. On ARC and ConceptARC, we find that subgoal persistence -- not subgoal injection alone -- is the central knob: moderate periods P in [3, 6] consistently outperform both very frequent (P=1) and very long horizons, with a clear minimum LM loss at P=3 (1.544 vs. 1.674 at P=1, 1.640 baseline; replicated over 5 seeds at mean 1.595, std 0.045). The intrinsic alignment weight lambda shows a complementary narrow optimum (lambda approximately 0.05). A controlled ablation at past-sweet-spot lambda isolates learned directional structure -- not architectural capacity or auxiliary loss alone -- as the source of interference when the alignment signal exceeds its optimum. Together these findings implicate a design principle for compositional planning in latent reasoning systems: medium-horizon intent must be coherent across enough computational steps for compositional structure to form.

13:00 JST画像/動画生成エージェント

IterCAD: 視覚に基づいた CAD の生成と編集のための反復マルチモーダル エージェント

コンピュータ支援設計は現代の製造において極めて重要ですが、既存の自動化手法は主にオープンループのワンショット生成に依存しており、現実世界の反復的な実践との不一致が生じています。このペーパーでは、閉ループの対話型 CAD 生成と編集のための統合マルチモーダル エージェント フレームワークである IterCAD について紹介します。このタスクは、マルチモーダル エージェントと実行可能な CAD サンドボックス間のマルチターン インタラクションとして定式化され、描画からコード、テキストからコード、対話型編集の 3 つのタスクをカバーします。これをサポートするために、標準に準拠したマルチビューのエンジニアリング図面、複雑なコード編集タスク、および忠実度の高いインタラクション軌跡を生成するための高度な工業製造機能を組み込んだデータ合成パイプラインを開発します。プログレッシブ SFT を介してエージェントを最適化し、その後、実行可能なプレフィックス マスキングを使用したジオメトリを意識した強化学習を実行して、コードの実行可能性と幾何学的忠実度を強化します。最後に、IterCAD-Bench 評価スイートを導入し、AUC-TR メトリクスと並行して面取り距離許容差-再現率 (CD-TR) 曲線を提案し、コードの有効性と幾何学的精度を統合する生存者バイアスのない標準を確立します。広範な実験により、IterCAD が複数のベンチマークにわたって非常に競争力のあるパフォーマンスを達成し、コードの実行性と幾何学的精度の両方で既存のアプローチを大幅に上回っていると同時に、閉ループ反復改良において優れた機能を示していることが実証されています。

原文 (English)

IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing

Computer-Aided Design is pivotal in modern manufacturing, yet existing automated methods predominantly rely on open-loop, one-shot generation, creating a mismatch with iterative real-world practices. In this paper, we present IterCAD, a unified multimodal agent framework for closed-loop, interactive CAD generation and editing. We formulate the task as a multi-turn interaction between a multimodal agent and an executable CAD sandbox, covering three tasks: Drawing-to-Code, Text-to-Code, and Interactive Editing. To support this, we develop a data synthesis pipeline incorporating advanced industrial manufacturing features to generate standard-compliant multi-view engineering drawings, complex code-editing tasks, and high-fidelity interaction trajectories. We optimize the agent via progressive SFT followed by geometry-aware reinforcement learning with viable-prefix masking to enhance code executability and geometric fidelity. Finally, we introduce the IterCAD-Bench evaluation suite and propose the Chamfer Distance Tolerance-Recall (CD-TR) curve alongside its AUC-TR metric, establishing a survivor-bias-free standard that unifies code validity and geometric precision. Extensive experiments demonstrate that IterCAD achieves highly competitive performance across multiple benchmarks, significantly outperforming existing approaches in both code executability and geometric precision, while exhibiting superior capabilities in closed-loop iterative refinement.

13:00 JST研究/論文

グラフ ニューラル ネットワークの構造の保存と論理的表現力

グラフ ニューラル ネットワーク (GNN) と論理形式の間の橋渡しは、集約、結合、活性化関数の種類などのアーキテクチャ上の選択を修正することによって確立されています。これらの選択は、論理式を同等の GNN に変換できること、また逆に GNN を同等の式に変換できることを示すことによって、論理形式との厳密な対応を取得できる GNN の制限されたクラスを定義します。この論文では、埋め込み (拡張)、単射準同型性、および準同型性といった構造特性の下で保存される GNN 分類器のクラスの論理的表現力を確立することにより、意味論的な観点を取り上げます。我々は、そのようなプロパティごとに、GNN のクラスを特徴付ける段階的な様相論理の断片が存在することを示します。特に、埋め込みによる保存、単射準同型性、および準同型性は、それぞれ、実存段階的様相論理、その実存肯定的フラグメント、および実存肯定的様相論理に対応します。これらの結果は、特定のアーキテクチャの選択とは無関係に、GNN の広範なクラスの表現力を特徴づけますが、これらのクラスのそれぞれが同じ表現力の GNN アーキテクチャを許容することも示します。技術的には、私たちのアプローチは、高さが制限されたツリーに対して新しい十分に準順序の結果を使用し、解明不変クラスの有限表現を生成します。

原文 (English)

Structural Preservation and the Logical Expressiveness of Graph Neural Networks

Bridges between graph neural networks (GNNs) and logical formalisms have been established by fixing architectural choices, such as the types of aggregation, combination, and activation functions. These choices define restricted classes of GNNs for which tight correspondences with logical formalisms can be obtained, by showing that logical formulae can be translated into equivalent GNNs and, conversely, that GNNs can be translated into equivalent formulae. In this paper we take a semantic perspective by establishing the logical expressiveness of classes of GNN classifiers that are preserved under structural properties: embeddings (extensions), injective homomorphisms, and homomorphisms. We show that, for each such property, there exists a fragment of graded modal logic characterising the class of GNNs. In particular, preservation under embeddings, injective homomorphisms, and homomorphisms corresponds to existential graded modal logic, its existential-positive fragment, and existential-positive modal logic, respectively. These results characterise the expressiveness of broad classes of GNNs independently of specific architectural choices, but we also show that each of these classes admits a GNN architecture of the same expressiveness. Technically, our approach uses a new well-quasi-order result for trees of bounded height, yielding finite representations of unravelling-invariant classes.

13:00 JSTLLM/生成AIエージェント

DeXposure-Claw: DeFi リスク監視のためのエージェント システム

分散型金融により、監督当局は急速に変化するネットワーク化された信用リスクにさらされます。汎用 LLM エージェントはこの設定にあまり適合しません。弱い証拠を深読みし、一か八かの介入を推奨しますが、既存の評価では、結果として生じる誤報を測定するための規制当局と連携した方法が提供されていません。我々は、構造化された証拠を通じて LLM 決定をルーティングする、予測に基づいたエージェント監視システムである DeXposure-Claw を紹介します。(1) DeXposure-FM は、グラフ時系列基礎モデルであり、将来のエクスポージャ ネットワークを予測します。 (2) 決定論的なモニターとストレス シナリオは、それらの予測を型指定されたアラート、属性シグナル、およびシナリオの証拠に変換します。 (3) データの健全性と信頼ゲートにより、DeXposure-Claw が根拠のある監査可能な監督チケットを発行する前にエスカレーションが抑制されます。さらに、6 軸の評価ハーネスである DeXposure-Bench を開発します。このベンチの決定軸は、規制当局と調整された絶対損失グラウンド トゥルースおよび明示的な誤介入率に対してチケットをスコア付けします。 5 年間の毎週の実データを用いた実験により、当社のシステムが完全にサポートされています。コードは https://github.com/EVIEHub/DeXposure-Claw にあります。

原文 (English)

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no regulator-aligned way to measure the resulting false alarms. We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that routes LLM decisions through structured evidence: (1) DeXposure-FM, a graph time-series foundation model, forecasts future exposure networks; (2) deterministic monitors and stress scenarios then turn those forecasts into typed alerts, attribution signals, and scenario evidence; and (3) data-health and confidence gates constrain escalation before DeXposure-Claw emits auditable supervisory tickets with rationales. We further develop DeXposure-Bench, a six-axis evaluation harness, whose decision axis scores tickets against a regulator-aligned absolute-loss ground truth and an explicit false-intervention rate. Experiments on five years of weekly real data fully support our system. Code is at https://github.com/EVIEHub/DeXposure-Claw.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文AnthropicClaudeOpenAIGPT / ChatGPTGoogleGemini

IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO

Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文NVIDIA

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In…

13:00 JSTエージェント

ATRIA: 反復エージェントを使用した適応型追跡可能な ECG レポート

既存の ECG レポート生成は緊密に結合されており、解釈とレポートがエンドツーエンドで融合されているため、ステージレベルの手段を必要とせずにエラーが伝播します。一方、エージェントベースのシステムはタスクを分離しますがシングルパスのままで、以前の出力を再検討することはありません。代わりに、臨床 ECG レポートは反復的に展開され、段階的なコンテキスト統合と双方向編集が必要になります。我々は、臨床医の反復的なワークフローを反映するマルチエージェント ECG レポート システムである \textsc{ATRIA} を紹介します。これは、すべてのレポートの主張をその裏付けとなる証拠に結び付け、その証拠によって裏付けられていないステートメントにフラグを立て、セッション中に追加のコンテキストを組み込み、臨床医が 1 つの不透明な出力を受け入れるのではなく、個々の所見を検証して修正できるようにします。そのエージェントはすでに臨床で使用されている ECG 分析モデルを使用しているため、基礎となる所見は臨床的に信頼できるものです。また、クラウドベースの Web サービスとして、\textsc{ATRIA} はすぐに導入できる状態になっています。ライブ デモとビデオを利用して、4 つのインタラクション ケースを通じて \textsc{ATRIA} をデモンストレーションします。

原文 (English)

ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents

Existing ECG report generation is tightly coupled -- interpretation and reporting fused end-to-end, so errors propagate without stage-level recourse -- while agent-based systems decouple tasks but remain single-pass, never revisiting earlier outputs. Clinical ECG reporting instead unfolds iteratively, requiring progressive context integration and bidirectional editing. We present \textsc{ATRIA}, a multi-agent ECG reporting system that mirrors the clinician's iterative workflow: it binds every report claim to its supporting evidence, flags statements unsupported by that evidence, incorporates additional context mid-session, and lets clinicians verify and revise individual findings rather than accept one opaque output. Because its agents use ECG analysis models already in clinical use, the underlying findings are clinically trustworthy; and as a cloud-based web service, \textsc{ATRIA} is ready for immediate deployment. We demonstrate \textsc{ATRIA} through four interaction cases, with a live demo and video available.

13:00 JST研究/論文

Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured…

13:00 JST研究/論文

統治可能な医療 AI スキル エコシステムのための臨床ハーネス

医療 AI は依然として孤立したモデルを中心に組織化されていますが、臨床ケアには時間を超えて持続する責任ある機能が必要です。私たちは、臨床 AI スキルと Clinical Harness を提案します。これは、AI 対応の臨床機能を登録、調整、保護、監視するためのランタイム ガバナンス アーキテクチャです。骨粗鬆症を例として使用し、知識主導型、データ主導型、および物理学に基づいて強化されたスキルが、ランタイム ガバナンスの下でライフサイクル ケアをどのようにサポートできるかを示します。

原文 (English)

Clinical Harness for Governable Medical AI Skill Ecosystems

Medical AI remains organized around isolated models, whereas care requires accountable capabilities that persist across time. We define clinical AI skills and propose the Clinical Harness, a runtime governance architecture that registers, orchestrates, constrains and monitors them. Using osteoporosis as an exemplar, we show how knowledge-driven, data-driven and physics-enhanced skills can support lifecycle care and provide a governed substrate for future medical agents.

13:00 JSTLLM/生成AIエージェント

OpenRCA 2.0: 結果ラベルから因果関係プロセスの監視まで

根本原因分析 (RCA) では、長いコンテキストの理解、複数ステップの推論、ツールの使用など、LLM エージェントの機能の総合的なテストが行​​われます。ただし、既存のデータセットには根本的なギャップがあります。つまり、根本原因のみにラベルが付けられ、観察された症状につながる伝播経路はラベル付けされないため、単純なパターン マッチングのタスクが大幅に簡素化されます。厳密な評価をサポートするために、フォールト挿入による既知の介入を利用して因果伝播パスを再構築する段階的なラベル付けプロトコルである PAVE を導入します。このメカニズムは前方検証です。つまり、症状から逆方向に推論するのではなく、原因から結果に至るまで推論します。 PAVE を適用すると、LLM エージェントに対する段階的な因果的アノテーションを備えた最初のクロスシステム RCA ベンチマークである OpenRCA 2.0 (500 インスタンス) が生成されます。 11 のフロンティア LLM 全体で、正確な根本原因セットの回復に成功するのは、平均して 20.7% のケースのみです。この困難がどこにあるのかを特定するために、基準を緩和して、根拠のない診断と呼ばれるものを見つけます。エージェントは、ケースの 76.0% で少なくとも 1 つの正しい根本原因サービスを特定しますが、そのサービスを観察された症状への検証された因果伝播経路に根拠付けるのは 61.5% のみです。結果のみの評価では、この失敗モードが隠蔽されます。段階的因果的グラウンドトゥルースは、信頼できる LLM ベースの RCA エージェントに欠けている部分です。

原文 (English)

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed symptom, which largely simplifies the task to naive pattern matching. To support rigorous evaluation, we introduce PAVE, a step-wise labeling protocol that leverages known interventions from fault injection to reconstruct causal propagation paths. The mechanism is forward verification: reasoning from cause to effect rather than inferring backward from symptoms. Applying PAVE yields OpenRCA 2.0 (500 instances), the first cross-system RCA benchmark with step-wise causal annotations for LLM agents. Across 11 frontier LLMs, recovering the exact root-cause set succeeds in only 20.7% of cases on average. To locate where this difficulty lies, we relax the criterion and find what we call the ungrounded diagnosis: agents identify at least one correct root-cause service in 76.0% of cases, but ground that service in a verified causal propagation path to the observed symptom in only 61.5%. Outcome-only evaluation hides this failure mode; step-wise causal ground truth is the missing piece for trustworthy LLM-based RCA agents.

13:00 JSTビジネス/資金調達

ナレッジ グラフにおけるグラフ間の意味的類似性の測定: ナレッジ グラフ埋め込みの経験的評価

ナレッジ グラフ (KG) は事実を構造化されたトリプルとして表し、さまざまなドメインにわたる関係知識を整理するために広く使用されています。テキスト情報の範囲が単語や文章から完全な文書に及ぶのと同様に、KG 情報は、エンティティ、関係、トリプルからサブグラフや KG 全体に至るまで、複数のレベルで解釈できます。ただし、既存の KG 埋め込み手法は主にエンティティ、リレーション、トリプルに焦点を当てており、グラフレベルのセマンティクスにはほとんど対処されていません。通常、構造パターンに基づいてグラフを比較する従来のグラフレベルの方法も、構造の類似性だけでは KG 間の意味的な類似性を保証できないため、不十分です。さまざまな方法がそのようなグラフレベルの意味論的情報をどの程度うまく捕捉しているかを評価するために、KG のペアが意味論的に対応する基礎的な情報を表すかどうかを決定する、グラフ間の意味論的類似性を研究します。信頼できるグラウンドトゥルース対応を取得するために、テキスト文書を変更し、元の文書と変更された文書の両方から KG を抽出し、それらの既知の対応関係を KG ペアに転送することにより、意味論的一致データセットを構築します。各データセットについて、テキストベース、構造ベース、および KG 埋め込みベースのアプローチを比較します。 KG 埋め込みベースのアプローチでは、ペアごとのエンティティの最大類似性を使用する \textit{EmbPairSim} と、周波数加重セントロイドを使用する \textit{AvgEmbSim} の 2 つのスコアリング関数を導入します。 WikiText-2 と CC-News での実験では、\textit{EmbPairSim} が大幅に少ないパラメーターを使用しながら、Sentence-BERT よりも最大 5.3 pp 高い MRR を達成することが示されています。これらの結果は、KGE 表現が、KG におけるグラフ間の意味論的類似性に対するコンパクトで効果的なシグナルとして機能できることを示唆しています。私たちのコードは https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity で入手できます。

原文 (English)

Measuring Graph-to-Graph Semantic Similarity in Knowledge Graphs: An Empirical Evaluation of Knowledge Graph Embeddings

A Knowledge Graph (KG) represents facts as structured triples and is widely used to organize relational knowledge across diverse domains. Just as textual information ranges from words and sentences to complete documents, KG information can be interpreted at multiple levels, from entities, relations, and triples to subgraphs and entire KGs. However, existing KG embedding methods mainly focus on entities, relations, and triples, leaving graph-level semantics largely unaddressed. Conventional graph-level methods, which typically compare graphs based on structural patterns, are also insufficient because structural similarity alone cannot guarantee semantic similarity between KGs. To evaluate how well different methods capture such graph-level semantic information, we study graph-to-graph semantic similarity, which determines whether a pair of KGs represents semantically corresponding underlying information. To obtain reliable ground-truth correspondences, we construct a semantic matching dataset by modifying text documents, extracting KGs from both original and modified documents, and transferring their known correspondences to KG pairs. We compare text-based, structure-based, and KG embedding-based approaches on each dataset. For the KG embedding-based approach, we introduce two scoring functions: \textit{EmbPairSim}, which uses maximal pairwise entity similarity, and \textit{AvgEmbSim}, which uses a frequency-weighted centroid. Experiments on WikiText-2 and CC-News show that \textit{EmbPairSim} achieves up to 5.3 pp higher MRR than Sentence-BERT while using substantially fewer parameters. These results suggest that KGE representations can serve as compact and effective signals for graph-to-graph semantic similarity in KGs. Our code is available at https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity.

13:00 JST研究/論文

トリプレットの妥当性を超えて: ナレッジ グラフにおける関係セットの完成

ナレッジ グラフ (KG) は、現実世界の知識をトリプレットとして編成し、多くの下流アプリケーションを支えます。本質的に不完全であるため、ナレッジ グラフ補完 (KGC) は広く研究されており、通常はリンク予測が主要なパラダイムであるトリプレット予測として定式化されます。ただし、この定式化はトリプレットごとの情報の不完全性に焦点を当てており、エンティティ関係の互換性情報の不完全性を見落としています。この制限に対処するために、リンク予測タスクを補完し、特定のエンティティと意味的に互換性のある欠落している関係を推論することを目的とした関係セット完了タスク (RSC) を導入します。さらに、観察されたエンティティの関係間の潜在的なパターンをモデル化して、欠落しているものを推測する関係セット埋め込みモデル (RelSetE) を提案します。 RelSetE を評価するために、標準の KG ベンチマークから 3 つのベンチマーク データセットを導出します。広範な実験により、RelSetE がエンティティ関係の互換性パターンを効果的に捕捉し、エンティティの欠落した関係を推論する際に有利に機能することが実証されました。コードとデータは公開されています。

原文 (English)

Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs

Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incompleteness, knowledge graph completion (KGC) is widely studied and is typically formulated as triplet prediction, with link prediction as the dominant paradigm. However, this formulation focuses on the incompleteness of triplet-wise information and overlooks the incompleteness of entity-relation compatibility information. To address this limitation, we introduce a relation set completion task (RSC), which complements the link prediction task and aims to reason about missing relations that are semantically compatible with a given entity. We further propose a Relation Set Embedding model (RelSetE), which models latent patterns among the observed relations of entities to infer missing ones. To evaluate RelSetE, we derive three benchmark datasets from standard KG benchmarks. Extensive experiments demonstrate that RelSetE effectively captures entity-relation compatibility patterns and performs favorably in inferring missing relations of entities. Code and data are publicly available.

13:00 JSTエージェント

行動基盤モデルによる探索とオンライン転送

強化学習 (RL) におけるゼロショット転送は、転送時に追加の学習を行わずに、報酬のない軌道のみでトレーニングしながら、あらゆる報酬関数に対して最適なポリシーを生成できるエージェントをトレーニングすることを目的としています。タスクに対する一般性から、このようなモデルは「行動基盤モデル」(BFM)と呼ばれることもあります。近年、強力なパフォーマンスと改善が見られていますが、現在のフレームワークとアルゴリズムは、転送フェーズ中にエージェントが状態と報酬のペアのデータセットを通じて報酬 (解決すべきタスク) についてオフラインで通知され、それを使用して展開する最適なポリシーを選択することを前提としています。ただし、実際には、報酬がブラックボックス (ユーザーからの直接のフィードバックなど) である場合、そのようなデータセットを生成することはできません。環境とのインタラクションを通じて報酬を観察する必要があります。言い換えれば、オフライン転送の現在のフレームワークは、報酬を見つけるために探索を必要とする、試行錯誤によるオンライン学習という従来の RL 設定と一致していません。このペーパーでは、BFM 自体を探索ポリシーの生成に使用できるという重要な洞察を基に、ゼロショット RL でこの新しいオンライン転送に取り組むことを提案します。私たちは、このオンライン学習の問題を盗賊のような探索と搾取の問題の観点から組み立てることが可能であることを示します。より正確には、各ステップでバンディット アルゴリズムがポリシーを推奨し、BFM がそれを環境内で実行し、報酬と新しい状態を生成します。最適なポリシーに収束するまでこのプロセスを繰り返します。線形報酬近似の一般的なコンテキストで、上限信頼限界にヒントを得た定式化を導き出し、不確実性行列の固有値の最小化を通じて探索が達成できることを示します。手法の概念を検証するために、単純な環境でフレームワークを定性的および定量的に評価します。

原文 (English)

Exploration and Online Transfer with Behavioral Foundation Models

Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories. For their generality over tasks, such models are sometimes called ``Behavioral Foundation Models'' (BFMs). While they have shown strong performances and improvements in recent years, the current framework and algorithms still assume that, during the transfer phase, the agent is informed offline about the reward (the task to solve) through a dataset of state-reward pairs, which it uses to pick the best policy to deploy. However, in practice if the reward is a black-box (e.g. direct user feedback), it is not possible to generate such a dataset: it is necessary to observe the reward through interactions with the environment. In other words, the current framework of offline transfer is not aligned with the traditional RL setting of online learning through trial-and-error, which requires exploration in order to find rewards. This paper proposes to tackle this new online transfer in zero-shot RL, with the key insight that the BFM itself can be used to generate exploration policies. We show that it is possible to frame this online learning problem in terms of a bandit-like exploration-exploitation problem. More precisely, at each step the bandit algorithm recommends a policy, the BFM executes it in the environment, which yields a reward and a new state; we repeat the process until we converge to the optimal policy. In the popular context of linear reward approximation, we derive a formulation inspired by Upper Confidence Bound and show that exploration can be achieved through the minimization of the eigenvalues of an uncertainty matrix. We evaluate qualitatively and quantitatively our framework on a simple environment to validate the concept of our method.

13:00 JST研究/論文

Corruption Robust Offline Reinforcement Learning with Human Feedback

We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset o…

13:00 JST研究/論文

Perturbation Effects on Robustness and Individual Fairness

Deep neural networks are vulnerable to adversarial perturbations that can simultaneously degrade prediction robustness and individual fairn…

13:00 JSTLLM/生成AIハードウェア/半導体

Learning by Surprise: Adaptive Mitigation of Model Collapse in Large Language Models

As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a p…

13:00 JST研究/論文

Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning

Deep reinforcement learning (DRL) has successfully addressed many complex control problems. However, the neural networks representing polic…

13:00 JST研究/論文

Mantis: Lightweight Foundation Model for Time Series Classification

While foundation models have revolutionized various domains, their application to time series classification remains rather under-explored,…

13:00 JSTLLM/生成AI

Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications. In black-box settings, several self-c…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTGeminiCopilot

Artificial Intelligence in Sports: Insights from a Quantitative Survey among Sports Students in Germany about their Perceptions, Expectations, and Concerns regarding the Use of AI Tools

Generative Artificial Intelligence (AI) tools such as ChatGPT, Copilot, or Gemini have a crucial impact on academic research and teaching.…

13:00 JSTLLM/生成AIビジネス/資金調達

SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA

As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evalu…

13:00 JST研究/論文

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. I…

13:00 JST画像/動画生成研究/論文

A Reproducible Benchmark of Lightweight CNNs: Accuracy, Efficiency, and the Impact of Pretrained Initialization

Lightweight convolutional neural networks are often compared using results obtained with different training recipes, input settings, and pr…

13:00 JSTエージェント

Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems

Collaborative agentic AI is projected to transform entire industries by enabling AI-powered agents to autonomously perceive, plan, and act…

13:00 JSTLLM/生成AI研究/論文

From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary

The advent of artificial intelligence has propelled AI-Generated Game Commentary (AI-GGC) into a rapidly expanding research area, offering…

13:00 JST画像/動画生成

Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling

Recent advances in 3D neural representations and instance-level editing models have enabled the efficient creation of high-quality 3D conte…

13:00 JSTLLM/生成AI研究/論文

LLM-Aided Joint Secrecy Precoding and Trajectory for RSMA-Based Heterogeneous UAV Networks

This paper investigates secure communications in rate-splitting multiple access (RSMA) enabled heterogeneous UAV networks, where multiple U…

13:00 JST画像/動画生成ビジネス/資金調達

VGGSounder: 基礎モデルのオーディオビジュアル評価

視聴覚基礎モデルの出現は、マルチモーダルな理解を確実に評価することの重要性を強調しています。 VGGSound データセットは、オーディオビジュアル分類の評価のベンチマークとしてよく使用されます。ただし、私たちの分析では、不完全なラベル付け、部分的に重複するクラス、不整合なモダリティなど、VGGSound のいくつかの制限が特定されました。これらは、聴覚および視覚能力の歪んだ評価につながります。これらの制限に対処するために、VGGSounder を導入します。これは、VGGSound を拡張し、オーディオビジュアル基礎モデルを評価するために特別に設計された、包括的に再アノテーションが付けられたマルチラベル テスト セットです。 VGGSounder は詳細なモダリティの注釈を備えており、モダリティ固有のパフォーマンスを正確に分析できます。さらに、新しいモダリティ混乱メトリックを使用して別の入力モダリティを追加したときのパフォーマンスの低下を分析することで、モデルの限界を明らかにします。

原文 (English)

VGGSounder: Audio-Visual Evaluations for Foundation Models

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.

13:00 JST研究/論文

Physics-Constrained Fine-Tuning of Flow-Matching Models for Generation and Inverse Problems

We present a framework for fine-tuning flow-matching generative models to enforce physical constraints and solve inverse problems in scient…

13:00 JSTLLM/生成AI研究/論文

Dataset Construction for Training LLM to Learn Analog Circuit Knowledge

This paper constructs a textual dataset for training large language models (LLMs) to learn analog circuit knowledge and customizes LLM trai…

13:00 JST研究/論文

Quantum Flow Matching

The flow matching has rapidly become a dominant paradigm in classical generative modeling, offering an efficient way to interpolate between…

13:00 JST研究/論文

One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression

Pruning is a core technique for compressing neural networks to improve computational efficiency. This process is typically approached in tw…

13:00 JST画像/動画生成

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned…

13:00 JSTロボティクス

A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

Human-to-humanoid imitation learning presents a promising pathway to address the severe data scarcity bottleneck in robotics by utilizing a…

13:00 JST研究/論文

Graph Coloring for Multi-Task Learning

When different objectives conflict with each other in multi-task learning, gradients begin to interfere and slow convergence, thereby poten…

13:00 JSTビジネス/資金調達

SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs

Machine Learning (ML) pipelines encode quality-relevant decisions across data preparation, training, evaluation, and configuration code. So…

13:00 JST画像/動画生成

CharDiff-LP: A Diffusion Model with Character-Level Guidance for License Plate Image Restoration

License plate image restoration is important not only as a preprocessing step for license plate recognition but also for enhancing evidenti…

13:00 JST研究/論文

Not Every Time and Frequency Need to Be Forgotten in Diffusion Unlearning

Data unlearning aims to remove the influence of specific training samples from a trained model. In fine-tuning methods, data unlearning rel…

13:00 JSTLLM/生成AIエージェント研究/論文

Human-Agent Collaborative Paper-to-Page Crafting

In the quest for scientific progress, communicating research is as vital as the discovery itself. Yet, researchers are often sidetracked by…

13:00 JSTLLM/生成AI画像/動画生成エージェント

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding

Graphical user interface (GUI) grounding is a key capability for computer-use agents, mapping natural-language instructions to actionable r…

13:00 JST研究/論文

Enhancing Graph Representations with Neighborhood-Contextualized Message-Passing

Graph neural networks (GNNs) have become an indispensable tool for analyzing relational data. Classical GNNs are broadly classified into th…

13:00 JST研究/論文

Optimal Self-Consistency for Efficient Reasoning with Large Language Models

Self-consistency (SC) is a widely used test-time inference technique for improving performance in chain-of-thought reasoning. It consists o…

13:00 JST研究/論文

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, the…

13:00 JST画像/動画生成エージェント

REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent

Foundation Models (FMs) are increasingly integrated into remote sensing (RS) pipelines. These models include unimodal vision encoders and m…

13:00 JST画像/動画生成

Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise

Virtual Try-On (VTON) synthesizes realistic images of a person wearing a target garment, with broad applications in e-commerce and fashion.…

13:00 JST研究/論文

A Unified and Stable Risk Minimization Framework for Weakly Supervised Learning with Theoretical Guarantees

Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly…

13:00 JST画像/動画生成

FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring

Joint video super-resolution and deblurring (VSRDB) requires both efficient long-range temporal modeling and robustness to frame-wise expos…

13:00 JST研究/論文

The HydroGym Reinforcement Learning Platform for Fluid Dynamics

Modeling and controlling fluids is critical across science and engineering. Effective flow control can increase lift, reduce drag, enhance…

13:00 JSTLLM/生成AI

Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation

Distilling the capabilities from a large reasoning model (LRM) to a smaller student model often involves training on substantial amounts of…

13:00 JSTLLM/生成AI画像/動画生成エージェント

InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training

GUI agents that interact with graphical interfaces on behalf of users represent a promising direction for practical AI assistants. However,…

13:00 JSTLLM/生成AIMicrosoft

From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching

Semantic caching has emerged as a pivotal technique for scaling LLM applications, widely adopted by major providers including AWS and Micro…

13:00 JST画像/動画生成

Toxicity Assessment in Preclinical Histopathology via Class-Aware Mahalanobis Distance for Known and Novel Anomalies

Drug-induced toxicity is a leading cause of preclinical and early-clinical failure, making early detection critical. Histopathology is the…

13:00 JSTLLM/生成AI

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catast…

13:00 JST研究/論文

DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks

Credit exposure in Decentralized Finance (DeFi) is often implicit and token-mediated, creating a dense web of inter-protocol dependencies.…

13:00 JST研究/論文

A swap-adversarial framework for improving domain generalization in electrocorticography-based Parkinson's disease classification

We propose a novel swap-adversarial framework that mitigates high inter-subject variability and the high-dimensional low-sample-size proble…

13:00 JST画像/動画生成ロボティクス

CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation

Robot navigation typically assumes an obstacle-free path exists between start and goal. In real environments, however, clutter may block al…

13:00 JST研究/論文

An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing

We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing.…

13:00 JST画像/動画生成研究/論文

DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting targ…

13:00 JST画像/動画生成

Are Video Reasoning Models Ready to Go Outside?

In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such con…

13:00 JST研究/論文

RCT とヒューマン アップリフト研究: フロンティア AI 評価のための方法論的課題と実践的な解決策

人間向上研究、つまりランダム化比較試験 (RCT) または同様の方法論を通じて人間のパフォーマンスに対する AI アクセスの影響を測定する研究は、最前線の AI ガバナンスと展開の決定にますます多くの情報を提供します。 RCT 手法は他の分野では堅牢ですが、フロンティア AI システムの特有の特性との相互作用は、特に結果が一か八かの意思決定に使用される場合にはまだ十分に検討されていません。バイオセキュリティ、サイバーセキュリティ、教育、労働などの分野で人間向上の研究を行った経験を持つ16人の専門家へのインタビューから得た結果を紹介します。専門家らはインタビューを通じて、人間高揚研究が依存する標準的な因果推論の仮定と研究の対象そのものとの間に繰り返される緊張を説明した。急速に進化する AI システム、変化するベースライン、異種混合で変化するユーザーの習熟度、多孔質な現実世界の設定により、内部、外部、構成の妥当性の基礎となる仮定が歪められ、上昇証拠の解釈と適切な使用が複雑になっています。私たちは、(1) 妥当性を研究するためのリスクにマッピングされ、大規模言語モデル (LLM) システムへの特異性の度合いによって分類された人間高揚研究における方法論的課題の統合、および (2) 課題から提案された解決策へのマッピングに貢献します。専門家が特定した課題と解決策を照合することで、人間向上の証拠の解釈限界と適切な使用法を明確にし、評価の実践とそれがもたらす決定を調整し、AI ガバナンスのより調整された方法論の基盤をサポートすることを目指しています。

原文 (English)

RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies

Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions. While RCT methods are robust in other fields, their interaction with the distinctive properties of frontier AI systems remains underexamined, particularly when results are used to inform high-stakes decisions. We present findings from interviews with 16 expert practitioners with experience conducting human uplift studies in domains including biosecurity, cybersecurity, education, and labor. Across interviews, experts described a recurring tension between the standard causal inference assumptions upon which human uplift studies rely and the object of study itself. Rapidly evolving AI systems, shifting baselines, heterogeneous and changing user proficiency, and porous real-world settings strain assumptions underlying internal, external, and construct validity, complicating the interpretation and appropriate use of uplift evidence. We contribute (1) a synthesis of methodological challenges in human uplift studies, mapped to risks to study validity and classified by their degree of specificity to large language model (LLM) systems, and (2) a mapping from challenges to proposed solutions. By collating expert-identified challenges and solutions, we seek to clarify the interpretive limits and appropriate uses of human uplift evidence, to align evaluation practice with the decisions it informs, and to support more coordinated methodological foundations for AI governance.

13:00 JST画像/動画生成

Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models

Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learnin…

13:00 JSTLLM/生成AI画像/動画生成

Visual Prompt Discovery via Semantic Exploration

LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts…

13:00 JSTLLM/生成AIハードウェア/半導体NVIDIA

An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU

Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabili…

13:00 JST画像/動画生成

Can VLMs Reason Robustly? A Neuro-Symbolic Investigation

Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly…

13:00 JST画像/動画生成研究/論文Sora

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing…

13:00 JSTロボティクス

Learning Dexterous Grasping from Sparse Taxonomy Guidance

Dexterous manipulation requires planning a grasp configuration suited to the object and task, which is then executed through coordinated mu…

13:00 JST研究/論文LlamaDeepSeek

SMART: When is it Actually Worth Expanding a Speculative Tree?

Tree-based speculative decoding accelerates autoregressive generation by verifying a branching tree of draft tokens in a single target-mode…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) system…

13:00 JST研究/論文

Lyapunov-Certified Direct Switching Theory for Q-Learning

Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing Q-learning f…

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning

Numerical reasoning over expert-domain tables often exhibits high in-domain accuracy but limited robustness to domain shift. Models trained…

13:00 JSTLLM/生成AI

Shared Lexical Task Representations Explain Behavioral Variability In LLMs

One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to…

13:00 JSTLLM/生成AI研究/論文

Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search

This paper focuses on a key challenge in Neural Architecture Search (NAS): integrating established architectural knowledge while exploring…

13:00 JST研究/論文

On Variance Reduction in Learning Mean Flows

One-step generative modeling has emerged as a leading approach for amortizing the inference cost of diffusion and flow-matching models. Amo…

13:00 JSTLLM/生成AINVIDIA

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving

Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highl…

13:00 JSTエージェント研究/論文

An Executable Benchmarking Suite for Tool-Using Agents

Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often…

13:00 JST画像/動画生成

医療画像解析のためのタスク整合型自己教師あり学習: 体系的なレビューと実践的な設計ガイドライン

自己教師あり学習 (SSL) は、ラベルのないデータから表現を学習することで、医療画像処理におけるアノテーションのボトルネックに対処するための有望なパラダイムとして浮上しています。ただし、その有効性は口実タスクの設計と下流の臨床目的との整合性に大きく依存します。医療画像処理における SSL の体系的でタスク指向のレビューを紹介し、さまざまな口実タスクの定式化が分類、セグメンテーション、検出、その他のタスク全体のパフォーマンスにどのような影響を与えるかを検証します。 PRISMA ガイドラインに従って、2017 年から 2025 年の間に発表された 75 件の研究を分析し、対照学習、非対照学習と予測学習、生成学習と再構成ベースの学習、およびハイブリッド学習の 4 つのパラダイムに整理しました。アーキテクチャごとにメソッドをカタログ化するのではなく、各パラダイムを、それが最もよくサポートする下流の目的にマッピングします。私たちの分析によれば、普遍的に最適な SSL 戦略は存在しません。代わりに、パフォーマンスは、口実タスク、イメージングモダリティ、およびターゲットタスク間の調整によって決まります。対照的な方法は全体的な識別特徴を学習し、分類とうまく一致しますが、微妙な病理学的パターンを見落とす可能性があります。生成および空間予測ベースのアプローチは、局所的な解剖学的構造をより適切に保存するため、セグメンテーションやその他の緻密な予測タスクにより適していますが、ハイブリッド手法は最もバランスの取れたパフォーマンスを提供します。さらに、モダリティ固有の設計が重要であること、および SSL が低ラベルおよび少数ショットの領域で最大の利点を提供することを示します。最後に、これらの発見を実用的な設計ガイドラインに絞り込み、病理学を意識した口実タスク設計、高次元データのリソース効率の高いトレーニング、標準化された評価プロトコルなどの未解決の課題を概説します。この研究は、医療画像処理において、より効果的で臨床的に関連性のある SSL フレームワークを設計するための実践的なガイダンスを提供します。

原文 (English)

Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines

Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data. However, SSL performance depends not only on model architecture, but also on whether the pretext task preserves information required by the downstream clinical objective. This review presents a task-oriented synthesis of SSL methods for medical imaging, focusing on how pretext-task design interacts with imaging modality, label availability, and downstream performance. We analyze 75 studies published from 2017 to 2025 and organize them into four paradigms: contrastive learning, non-contrastive and predictive learning, generative and reconstruction-based learning, and hybrid learning. Rather than cataloging methods chronologically, we examine how these paradigms support classification, segmentation, detection, reconstruction, and regression. The evidence suggests that no SSL strategy is universally optimal. Contrastive objectives generally encourage global discriminative representations and are well aligned with classification, but may underrepresent subtle or localized pathology. Spatial prediction, masked modeling, and reconstruction-based objectives better preserve anatomical structure and are often more suitable for segmentation and dense prediction. Hybrid methods can provide balanced representations, although they increase training complexity. Across modalities, SSL is most beneficial in low-label and few-shot regimes, but its effectiveness depends on modality-aware augmentation, pathology-preserving corruption, and clinically meaningful evaluation. We conclude with practical design guidelines and identify open challenges, including pathology-aware pretext tasks, resource-efficient training for high-dimensional data, and standardized evaluation protocols.

13:00 JSTエージェント

ChainCaps: 単調な機能減衰による構成安全なツール使用エージェント

ツールを使用するエージェントは、実行時にファイル システム、Web API、コード インタープリタ、およびエンタープライズ サービスを構成する、オープンエンドの展開環境で動作することが増えています。これにより、ツール構成に安全性のギャップが生じます。エージェントは、ツールごとのすべての権限チェックを満たしていても、機密文書の読み取り、要約、その要約の外部エンドポイントへの送信など、安全でないエンドツーエンドの影響を生み出す可能性があります。この障害モードをパーミッション ロンダリングと呼びます。 ChainCaps は、ランタイム ルールでこれに対処します。すべての値にはシンク固有の機能バジェットが含まれ、ツールの構成によって交差ごとにバジェットが伝播されます。値は、ツール チェーン内を移動するときに権限を保持したり失ったりする可能性がありますが、合成を通じて新しい権限を獲得することはできません。 ChainCaps は、エージェント サーバーやツール サーバーへの変更を必要としない透過的な MCP プロキシとして実装されています。 ChainCaps は、3 つのプロバイダーの 5 つのフロンティア モデルにわたる 82 のタスクにおいて、96 ~ 100% の無害な完了を維持しながら、攻撃の成功率を 25 ~ 68% から 0 ~ 4.8% に低下させます。再生実験では、スカラー IFC および関数ごとの分離ベースラインよりも優れたパフォーマンスを示します。マニフェストの品質が導入の主なボトルネックです。エキスパート マニフェストは 100% の攻撃ブロックに達しますが、単純なマニフェストは 27.3% に低下します。私たちの主張は、信頼できるマニフェストの下での明示的なフロー構成の安全性と、プロキシで可視のデータ移動に限定されており、今日導入されているツールを使用するエージェントには実際的なギャップがあります。

原文 (English)

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterprise services at runtime. This creates a safety gap in tool composition: an agent can satisfy every per-tool permission check and still produce an unsafe end-to-end effect, such as reading a confidential document, summarizing it, and sending the summary to an external endpoint. We call this failure mode permission laundering. ChainCaps addresses it with a runtime rule: every value carries a sink-specific capability budget, and tool composition propagates budgets by intersection. A value can preserve or lose authority as it moves through a tool chain, but it cannot gain new authority through composition. We implement ChainCaps as a transparent MCP proxy that requires no changes to the agent or tool servers. On 82 tasks across five frontier models from three providers, ChainCaps reduces attack success rate from 25-68% to 0-4.8% while preserving 96-100% benign completion. In replay experiments, it also outperforms scalar-IFC and per-function-isolation baselines. Manifest quality is the dominant deployment bottleneck: expert manifests reach 100% attack blocking, while naive manifests fall to 27.3%. Our claims are limited to explicit-flow composition safety under trusted manifests and proxy-visible data movement, a practical gap in deployed tool-using agents today.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

LLM ベースの産業資産運用の欠落データ層としてのナレッジ グラフ

産業用資産運用用の LLM ベースのエージェントは、フラットなドキュメント ストアを推論する場合、精度が限られています。 AssetOpsBench (KDD 2026) は、CouchDB、YAML、CSV に裏付けられた 139 の産業メンテナンス シナリオで GPT-4 エージェントが 65% を達成することを証明しています。固定データ層での LLM オーケストレーション パラダイム (ツールとしてのエージェントと計画実行) を比較します。私たちは補完的かつ直交的な質問をします。ツールの背後にあるデータ モデルはエージェントのパフォーマンスにどの程度影響しますか?同じシナリオに基づいて、ナレッジ グラフ レイヤー (781 ノード、955 エッジ、16 の関係タイプ) を導入し、次の 3 つのアーキテクチャを評価します。(1) 決定論的グラフ ハンドラー (LLM なし) 99% (137/139)。 (2) ベースラインが使用するのと同じ GPT-4 モデルを使用して、LLM によって生成された Cypher がグラフ全体で 82 ~ 83% に占める。 (3) 元のツールで拡張された LLM ベースラインは 65% (91/139、公開されている KDD 2026 リーダーボードの上限に一致)。私たちの主な発見は、LLM の使用法が逆転していることです。LLM に生データを推論するよう依頼するのではなく、型付きスキーマから構造化クエリを生成するよう LLM に依頼します。グラフは決定的に実行されます。さらに、40 のグラフネイティブ シナリオ (マルチホップの依存関係、ベクトルの類似性、PageRank の重要度) を提供し、拡張された HuggingFace AssetOpsBench リリース (467 のシナリオ、6 ドメイン) に対して評価しました。この場合、決定論的ハンドラーは平均スコア 0.848 で 100% (467/467) を達成しました。これらの結果は、構造化された運用ドメインでは、LLM オーケストレーションではなくデータ層が主なボトルネックであり、ナレッジ グラフが生の産業データと LLM ベースの推論の間の統合層として機能することを示唆しています。

原文 (English)

Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations

LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) establishes that GPT-4 agents achieve 65% on 139 industrial maintenance scenarios, and compares LLM orchestration paradigms (Agent-As-Tool vs. Plan-Execute) on a fixed data layer. We ask the orthogonal question: how much does the data model behind the tools matter? We treat a typed knowledge graph as a grounding substrate and route each question by how it is best answered: (i) LLM-generated Cypher for structured retrieval, which lifts the same GPT-4 model from 65% to 82-83%; (ii) native graph and optimization primitives, with no LLM, reaching 99% on graph-answerable scenarios; and (iii) generation-augmented knowledge (GAK) for answers absent from the data -- the engine's agent materializes the missing facts as provenance-tagged graph nodes, then answers. A recurring theme is inverted LLM usage: we constrain the LLM to query generation or one-shot enrichment from a typed schema and let the graph execute deterministically. On the 88 real AssetOpsBench failure-mode scenarios the benchmark itself flags non-deterministic -- ten equipment types absent from the graph -- GAK lifts answerability from zero to 100% of equipment types and answers 81.8% of scenarios, every materialized fact tagged source:LLM-derived for auditability. We also contribute 40 graph-native scenarios. For structured operational domains the data layer -- not the LLM orchestration -- is the primary lever, and a typed knowledge graph serves as a grounding substrate between raw industrial data and LLM reasoning.

13:00 JSTLLM/生成AI研究/論文

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

We introduce BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in…

13:00 JST画像/動画生成

Quantitative Movement Testing: Measuring Chronic Pain Patient Movements from a Single Smartphone Video

Chronic pain diminishes quality of life by decreasing functional ability, yet objectively measuring this functional impact remains challeng…

13:00 JSTLLM/生成AI

INFUSER: Influence-Guided Self-Evolution Improves Reasoning

Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervi…

13:00 JSTLLM/生成AIエージェント

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents

Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning…

13:00 JSTLLM/生成AI画像/動画生成エージェント

ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm

Existing computer-use agents remain fundamentally limited in professional software manipulation: GUI-based agents suffer from fragile visua…

13:00 JSTLLM/生成AIエージェント

エージェントティックブラウザの同一生成元ポリシー

エージェントティック ブラウザは自律型 AI エージェントを Web ブラウザに統合し、ユーザーが自然言語の指示を通じて Web タスクを実行できるようにします。同一オリジン ポリシー (SOP) は、スクリプトによって引き起こされる無許可の自動クロスオリジン データ フローを防止する基本的なブラウザ セキュリティ メカニズムです。ただし、SOP がエージェントブラウザでも有効であるかどうかは未解決の問題であり、体系的に研究されていません。この取り組みでは、このギャップを埋めます。まず、エージェント ブラウザ自体がクロスオリジン データ フローの自動チャネルとして機能し、SOP 違反につながる可能性があることを観察しました。この現象を調査するために、エージェント ブラウザーでの SOP 違反を評価するためのベンチマークである SOPBench を構築します。私たちの評価によると、既存のエージェントブラウザは、無害な設定でも攻撃下でも頻繁に SOP に違反しています。この問題に対処するために、エージェント ブラウザに合わせた SOP 強制メカニズムである SOPGuard を提案します。 SOPGuard は、オープンソースのエージェント ブラウザーである BrowserOS に実装されています。広範な評価により、SOPGuard は実用性を維持し、実行時のオーバーヘッドがわずかしか発生せずに、SOP を効果的に適用できることが実証されています。コードとデータは https://github.com/wxl-lxw/BrowserOS-SOPGuard で入手できます。

原文 (English)

Same-Origin Policy for Agentic Browsers

Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instructions. The same-origin policy (SOP) is a fundamental browser security mechanism that prevents unauthorized automated cross-origin data flows induced by scripts. However, whether SOP remains effective in agentic browsers is an open question that has not been systematically studied. In this work, we bridge this gap. We first observe that an agentic browser can itself serve as an automated channel for cross-origin data flows, potentially leading to SOP violations. To investigate this phenomenon, we construct SOPBench, a benchmark for evaluating SOP violations in agentic browsers. Our evaluation shows that existing agentic browsers frequently violate SOP, both in benign settings and under attacks. To address this problem, we propose SOPGuard, an SOP enforcement mechanism tailored to agentic browsers. We implement SOPGuard in BrowserOS, an open-source agentic browser. Extensive evaluations demonstrate that SOPGuard effectively enforces SOP while preserving utility and incurring only a small runtime overhead. Our code and data are available at https://github.com/wxl-lxw/BrowserOS-SOPGuard.

13:00 JST研究/論文

Entropy-Gated Latent Recursion

Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity…

13:00 JST画像/動画生成

PSCT-Net: 微分可能な逆投影と注意に基づく改良による、形状を考慮した小児頭蓋骨 CT 再構成

コンピュータ断層撮影 (CT) は小児の頭蓋顔面異常の診断に不可欠ですが、発達中の解剖学的構造に放射線リスクをもたらします。まばらな二平面 X 線から 3D CT を再構成することは、低線量の代替手段となりますが、非常に不適切です。既存の方法は、ジオメトリに依存しない特徴リフティングを採用しており、明示的な空間モデリングを行わずに単純に 2D 特徴を 3D に投影するため、深さの曖昧さと骨境界の劣化が生じます。微分可能な逆投影を備えた幾何学認識フレームワークである PSCT-Net を紹介します。微分可能な逆投影により、空間的に忠実な体積事前分布が確立され、深さの曖昧さが軽減されます。次に、注意誘導投影 (AGP-3D) モジュールが、2D 領域と 3D 位置の間の非線形ボクセル単位の対応を学習します。 Bi方向 Mamba (BiM-3D) モジュールは、線形の複雑さで長距離の体積依存関係をキャプチャします。さらに、内部評価用に正常症例と病理学的症例で構成される民間の施設小児頭蓋骨CTコホートであるPedSkull-CTをキュレーションし、成人中心の体幹に焦点を当てたデータセットのギャップに対処します。

原文 (English)

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement

Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomies. Reconstructing 3D CT from sparse bi-planar X-rays offers a low-dose alternative but is severely ill-posed. Existing methods employ geometry-agnostic feature lifting, naively projecting 2D features into 3D without explicit spatial modeling, causing depth ambiguity and degraded osseous boundaries. We present PSCT-Net, a geometry-aware framework with differentiable back-projection. Differentiable back-projection establishes a spatially faithful volumetric prior, alleviating depth ambiguity. An Attention-Guided Projection (AGP-3D) module then learns non-linear voxel-wise correspondences between 2D regions and 3D locations. A Bidirectional Mamba (BiM-3D) module captures long-range volumetric dependencies with linear complexity. We further curate a private institutional pediatric skull CT cohort, PedSkull-CT, comprising normal and pathological cases for internal evaluation, addressing the gap in adult-centric, trunk-focused datasets.

13:00 JST研究/論文

Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling

Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…

13:00 JST研究/論文

Protein contacts are already in the attention: a single-forward-pass alternative to the Categorical Jacobian

The Categorical Jacobian of Zhang et al. (2024) reads protein contacts from a language model by perturbing every residue with every alterna…

13:00 JSTLLM/生成AIエージェント研究/論文

RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents

Agentic coding harnesses - such as Agent-Skills, Superpowers, and Agent-Rigor - are increasingly deployed to augment underlying LLMs for re…

13:00 JSTLLM/生成AILlamaQwen

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work,…

13:00 JSTLLM/生成AI

Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts

While the wider applicability of LLMs in the legal field is currently debated due to their reliability and the gravity of any errors, narro…

13:00 JSTLLM/生成AI

1 年後...被害は続いていますが、私たちも同じです!

汎用大規模言語モデル (LLM) は、メンタルヘルス関連の会話にますます使用されていますが、安全対策は依然として不十分であり、臨床症状全体で一貫性がありません。この研究では、8 次元の危害分類と多次元の評価フレームワークを導入し、4 つの敵対的攻撃のバリアントを使用して、16 の DSM-5 条件にわたる 6 つの独自の LLM を評価します。その結果、安全策は自殺と自傷行為に対してのみ確実に有効であり、摂食障害、物質使用障害、大うつ病性障害などの疾患では失敗率が最大 100% であることが示されています。私たちは、これらの LLM の倫理的な設計と展開には、臨床状態全体にわたって明確に定義された危害カテゴリーと、それに応じた安全措置の実装が必要であると主張します。このような保護措置が講じられるまで、これらのモデルは脆弱な人々に重大なリスクをもたらすため、教育現場への統合の増加が特に懸念されます。

原文 (English)

One Year Later...The Harms Persist, But So Do We!

General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety guardrails remain inadequate and inconsistent across clinical conditions. This study evaluates eight proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100\%. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into publicly available settings (e.g., schools, search engines, and consumer chatbots) are particularly concerning.

13:00 JST研究/論文GPT / ChatGPT

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models

Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather…

13:00 JST画像/動画生成

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation

Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-to…

13:00 JST研究/論文

Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching

Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are compared to determine w…

13:00 JST画像/動画生成研究/論文

SpatialUAV: 低高度 UAV の知覚、コラボレーション、およびモーションのための空間インテリジェンスのベンチマーク

空間インテリジェンスは、低高度の無人航空機 (UAV) の認識、コラボレーション、およびナビゲーションに不可欠です。しかし、既存の UAV ベンチマークは、多くの場合、画像レベルの認識、単一ビューの理解、または狭い回答形式を重視しており、3D 空間推論、マルチビュー コラボレーション、シーン ダイナミクス、および多様なタスクの定式化が十分に評価されていません。これらのギャップに対処するために、14 のきめ細かいタスク タイプにわたる 4,331 個の精選されたインスタンスで構成される実際の低高度 UAV ベンチマークである SpatialUAV を導入します。これは、意味論的識別、空間関係、航空と航空のコラボレーション、航空と地上のコラボレーション、および動作の理解をカバーします。 SpatialUAV は、すべてのサンプルを統一された視覚入力、質問、回答スキーマに編成し、オプション ラベル、領域識別子、幾何学的値、ビュー間の対応関係、および自由形式の動作記述を含む 7 つの入力構成と 9 つの回答形式をサポートします。信頼性の高い根拠に基づいた評価を保証するために、当社のデータ構築パイプラインは、検出器支援領域、深度監視、メタデータ由来のルール、広範な手動アノテーション、ブラインド フィルタリング、およびマルチターン人間による検証を、異種出力に対するタスク固有のメトリクスとともに統合します。 3 つのカテゴリにわたる代表的な視覚言語モデルを評価したところ、現在のモデルは依然として人間レベルのパフォーマンスには遠く及ばず、視点間の関連性、構造化された基礎、幾何学的な推論、および時間的視点の理解において顕著なボトルネックがあることが示されました。これらの結果は、低高度 UAV の空間インテリジェンスを進歩させるための経験的な指針を提供します。コードとデータは https://github.com/Hyu-Zhang/SpatialUAV で入手できます。

原文 (English)

SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion

Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image-level recognition, single-view understanding, or narrow answer formats, leaving 3D spatial inference, multi-view collaboration, scene dynamics, and diverse task formulations insufficiently evaluated. To address these gaps, we introduce SpatialUAV, a real low-altitude UAV benchmark comprising 4,331 curated instances across 14 fine-grained task types, covering semantic discrimination, spatial relation, aerial--aerial collaboration, aerial--ground collaboration, and motion understanding. SpatialUAV organizes all samples into a unified visual-input--question--answer schema, while supporting seven input configurations and nine answer formats, including option labels, region identifiers, geometric values, cross-view correspondences, and free-form motion descriptions. To ensure reliable and grounded evaluation, our data construction pipeline integrates detector-assisted regions, depth supervision, metadata-derived rules, extensive manual annotation, blind filtering, and multi-turn human validation, together with task-specific metrics for heterogeneous outputs. Evaluating representative vision-language models across three categories, we show that current models remain far from human-level performance, with pronounced bottlenecks in cross-view association, structured grounding, geometric reasoning, and temporal viewpoint understanding. These results offer empirical guidance for advancing low-altitude UAV spatial intelligence. Code and data are available at https://github.com/Hyu-Zhang/SpatialUAV.

13:00 JST画像/動画生成

Reflect-R1: 長いビデオの理解における自己修正のための証拠に基づくリフレクション

長時間のビデオを理解するための現在のマルチモーダル反射メカニズムは、主に内部パラメータ内の閉ループ自己反射に依存しています。客観的な外部証拠が欠如しているため、モデルはしばしば盲目的な自信に囚われ、エラーを修正できないことがよくあります。さらに、強化学習を多段階リフレクション パイプラインに適用すると、深刻なポリシー結合が導入され、専用のトレーニング データが重大に不足することでさらに悪化します。これらの制限に対処するために、この研究では、長いビデオを理解するための初の証拠主導型自己修正フレームワークである Reflect-R1 を提案しています。このフレームワークは、直感、検証、調停からなる 3 段階のパイプラインを構築します。客観的な視覚的証拠を動的に取得して最初の直観を検証し、複数の時間的検索を自律的に実行して矛盾を解決することで、幻覚ループを完全に断ち切ります。ポリシー結合を克服するために、さまざまな推論段階にわたって利点関数を独立して計算する、SD-GRPO という名前の段階分離型強化学習アルゴリズムを設計します。同時に、トレーニング データのギャップを埋めるために 120,000 サンプルのデータセットを構築します。 VideoMME や LongVideoBench などのベンチマークに関する広範な実験により、Reflect-R1 が最先端のパフォーマンスを達成していることが実証されています。私たちの方法は真の修正率を大幅に向上させ、厳密に客観的な証拠に基づいた真の自己修正を可能にします。

原文 (English)

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evidence, models are frequently trapped in blind confidence and often fail to correct errors. Furthermore, applying reinforcement learning to multi-stage reflection pipelines introduces severe policy coupling, which is exacerbated by a critical scarcity of dedicated training data. To address these limitations, this work proposes Reflect-R1, the first Evidence-Driven self-correction framework for long video understanding. The framework constructs a three-stage pipeline consisting of intuition, verification, and arbitration. By dynamically retrieving objective visual evidence to verify initial intuitions and autonomously executing multiple temporal searches to resolve conflicts, it completely breaks the hallucination loop. To overcome policy coupling, we design a stage-decoupled reinforcement learning algorithm named SD-GRPO that independently computes advantage functions across different reasoning stages. Concurrently, we construct a dataset of 120K samples to bridge the training data gap. Extensive experiments on benchmarks such as VideoMME and LongVideoBench demonstrate that Reflect-R1 achieves state-of-the-art performance. Our method significantly improves the genuine rectification rate and enables authentic self-correction strictly grounded in objective evidence.

13:00 JST研究/論文

SHARD: アライメント耐性のあるプライベート密検索のためのセルキー残差分割

高密度の埋め込みはセマンティック検索と RAG を支えていますが、漏洩したベクトル ストアにより、基礎となるテキストの多くがそれを保持している人の手に渡されます。これを可能にする攻撃 (少数ショットのアライメント、ゼロショットの反転、教師なしクロススペース変換) には 1 つの弱点があります。それは、保護されたストアが、既知のジオメトリにアライメントできる単一のグローバル ジオメトリであるということです。通常の軽量防御である秘密のグローバル回転も例外ではありません。攻撃者が既知のペアでほぼ亜空間次元を取得すると、直交プロクラステスはそれを回復します。この弱い軸を取り除く検索保存埋め込み変換である Shard を紹介します。中央に配置された埋め込みは、短いパブリック プレフィックス (ステージ 1 取得用) と、個別の秘密キーの下で C セルにシャーディングされたプライベート残差に分割されます。残差は CKKS の下で再ランク付けされ、キーはキャンセルされて内積が正確のままになります。単一のパラメーター C は、置き換えられるグローバル線形ベースライン (C=1) からドキュメントごとのマイクロキー (C=N) までデザインを実行します。リランクはフル次元であるため、Shard は、半 SVD 切り捨てが放棄された生の空間 nDCG@10 を返します。また、残差はセルローカルにキー付けされるため、拡散既知平文リークのもとで残差を共通フレームにマッピングし直すには、暗号化されたクエリ数が少ない場合、およそ C 倍のアンカー (C=256 で中央値 200 ~ 102,400) のコストがかかります。短いパブリック プレフィックスにより、近隣構造の漏洩がはるかに少なくなり、マイクロキー制限により、リンク不可能で更新可能なテンプレートで残差グラフがゼロになります。この障壁は、学習済み、非線形、および教師なしのアライナに対して保持されており、整合ユーティリティ ノイズ防御がほぼすべてのプローブを匿名化解除するのに対して、シャードは匿名化を解除しません。私たちはその制限について明確にしています。セル内ではキーがキャンセルされ、標的型攻撃者が必要とするのは d_priv アンカー程度だけであり、重複する参照コーパスは依然としてプレフィックスを介して漏洩します。シャードは攻撃を認識する幾何学的防御であり、暗号化を保証するものではありません。

原文 (English)

SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval

Dense embeddings underpin semantic search and retrieval-augmented generation, yet a leaked vector store hands much of the underlying text back. Modern inversion and alignment attacks share one weakness: the protected store is a single global geometry, and any single geometry can be aligned to a known one - a secret global rotation included, since orthogonal Procrustes recovers it from about subspace-dimension known-plaintext pairs. We introduce SHARD, a retrieval-preserving embedding transform that removes that weak axis. The centred embedding is rotated and split into a short public prefix (driving stage-1 retrieval) and a private residual sharded into C cells, each rotated under a separate secret key; the residual is reranked under CKKS, where the keys cancel and the inner product stays exact. One parameter C spans the global-linear baseline (C=1) to per-document micro-keys (C=N), making the keyed residual a cancellable template - revocable, renewable, unlinkable - for text embeddings, the first such scheme for dense retrieval. On five encoders: full-dimensional reranking returns the raw-space nDCG@10 that half-SVD truncation gives up; recovering the cell-keyed residual under a diffuse known-plaintext leak costs about C times more anchors (median 200 to 102,400 at C=256) for a few encrypted residual queries and the short public prefix leaks far less neighbour structure, with a micro-key limit driving residual leakage to zero. The barrier holds against learned-linear, non-linear and unsupervised aligners, and where a matched-utility noise defence de-anonymises almost every probe, SHARD de-anonymises none. Limits: within a cell similarities survive, a targeted attacker on one victim's cell needs only about d_priv anchors, and an overlapping reference corpus still leaks through the public prefix. SHARD is an attack-aware geometric defence, not a cryptographic guarantee.

13:00 JST研究/論文

The Remittance Blueprint: Data-driven Intelligence for Sri Lanka

This study analyzes Sri Lankan migration and remittances over 32 years (1994-2025). Using a 384-month harmonized dataset, we apply explorat…

13:00 JST画像/動画生成ロボティクス

The Speedup Paradox: Rethinking Inference Speed-Quality Trade-off in Embodied Tasks

Embodied foundation models have recently been widely used to improve robot generalization and task success rates. Previous works apply loss…

13:00 JST画像/動画生成

BREIT: A Framework for Brain Stroke Reconstruction using Multi-Frequency 3D EIT

Multi-Frequency Electrical Impedance Tomography (MF-EIT) is a non-invasive, low-cost modality that reconstructs electrical property distrib…

13:00 JSTLLM/生成AI研究/論文

Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking

There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully…

13:00 JST研究/論文

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cos…

13:00 JST画像/動画生成エージェント

LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving

Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD)…

13:00 JSTLLM/生成AIエージェントClaude

TraceLab: Characterizing Coding Agent Workloads for LLM Serving

Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this…