Skip to the content.

AIニュース 2026-06-30

自動生成: 2026-06-30 13:05 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Mapping Europe’s AI Workforce OpportunityOpenAI

    A new OpenAI report maps how AI could reshape jobs across the EU, hig…

  2. 国内大手が共同出資のAI開発企業「日本AI基盤モデル開発」、新名称「Noetra」で始動 産総研と国産マルチモーダルAI開発へITmedia AI+

    国内の大手企業が共同で出資するAI開発企業の日本AI基盤モデル開発(東京都渋谷区)は6月30日までに、1日付で名称をNoetraに変更した…

  3. Vibe coding platform Base44 launches own model as AI startups seek defensibilityTechCrunch AI

    Wix-owned vibe coding platform Base44 has started rolling out its own…

  4. GitHub、AIによる雑なプルリクエストを抑制へ ユーザー当たりのプルリク数に上限を設定できる新機能ITmedia AI+

    米GitHubは、ユーザーに対してプルリクエスト数の上限を設定できる新機能の導入を発表しました。

  5. あなたのAWS、コストの課題はどこにある? AIが教えてくれる「AWS FinOps Agent」パブリックプレビュー開始ITmedia AI+

    米Amazon Web Services(AWS)は、使用中のAWSのコストに関する質問や、コストに異常が発生した場合にその原因を調査して…

  6. 自社の業務に合わせたAIエージェントを「10分で作成」 freeeが「AI戦略」を強化ITmedia AI+

    クラウド会計システム「freee」の開発などを手掛けるフリーは、2月に発表したAI戦略の実現に向けた新たな取り組みとして「freee AI…

  7. メール、Teams、Slack――バラバラな連絡ツールの「見落とし」 ChatGPT Agentsで解決する方法ITmedia AI+

    メールやTeams、Slackなど複数の連絡ツールを使っていると、重要なメッセージの見落としや返信漏れをしてしまいます。AIで解決できませ…

トピック別件数

日本語メディア20件

ITmedia AI+ (日本語)

12:27 JSTその他

「『ハイキュー!!』全巻を読み込ませている」――東宝がアニメグッズの監修にAI活用、他作品にも拡大へ

東宝はこのほど、アクセンチュアと協力し、IPを活用した商品の監修にAIシステムを導入した。同システムの詳細や展望について、アクセンチュアの戸賀慶氏と東宝の田中亮史氏が語った。

12:10 JSTその他

国内大手が共同出資のAI開発企業「日本AI基盤モデル開発」、新名称「Noetra」で始動 産総研と国産マルチモーダルAI開発へ

国内の大手企業が共同で出資するAI開発企業の日本AI基盤モデル開発(東京都渋谷区)は6月30日までに、1日付で名称をNoetraに変更したと発表した。

12:00 JSTエージェント

【徹底入門】AIエージェントで注目の「AX」とは何か 人間だけじゃなく“AI視点の使いやすさ”も重要に?

AIエージェント時代に注目を集めつつある「AX」(Agent Experience、エージェント体験)とは何か解説する。

11:00 JSTその他

「“社長AI”って意味ある?」→言った本人も手のひら返し 幹部の9割が高評価したNTTドコモビジネスの「AI小島社長」開発録

経営トップの判断や思考をAIで再現する取り組みが、国内の大企業に広がっている。NTTドコモビジネスが開発し、同社幹部の9割が「方針理解に役立った」と評価するという「AI小島社長」に迫る。

10:49 JSTその他

GitHub、AIによる雑なプルリクエストを抑制へ ユーザー当たりのプルリク数に上限を設定できる新機能

米GitHubは、ユーザーに対してプルリクエスト数の上限を設定できる新機能の導入を発表しました。

10:41 JSTエージェント

あなたのAWS、コストの課題はどこにある? AIが教えてくれる「AWS FinOps Agent」パブリックプレビュー開始

米Amazon Web Services(AWS)は、使用中のAWSのコストに関する質問や、コストに異常が発生した場合にその原因を調査して特定してくれる「AWS FinOps Agent」のパブリックプレビュー開始を発表しました。

10:33 JSTその他

AI避けて「人間にだけ届く」広告配信へ、博報堂DYが新会社設立  虹彩認証「World ID」活用

サム・アルトマン氏らが共同発明した人間認証技術「World ID」を活用する。

08:00 JSTLLM/生成AIAnthropic

「Fable 5を全タスクに使う必要はない」 Anthropic開発者直伝のトークンコスト節約術

Anthropicの開発者がトークンコストを抑えるための戦略を語った。タスクに応じてモデルを使い分けると共に、適切な入力方式を選ぶことが重要だ。

08:00 JSTその他

「AIが前提となる世界」でSIerは生き残れるか?

AIが前提となる世界で問われるのは「AIをどう使うか」ではなく、「組織をどう設計し直すか」だ。ITRアナリストと「新しい乱世」を生き残るための道筋を考える。

07:00 JSTその他

中小企業の採用から勤怠・経費管理まで バラバラのSaaSからまとめてAIがデータ分析

中小企業では、複数のSaaSや紙、Excelが混在することで、管理業務の負担が増えている。NoahWorksは、採用から勤怠、決済までを一元化し、AIによるデータ分析で現場の業務改善を支援する。

07:00 JSTエージェント

自社の業務に合わせたAIエージェントを「10分で作成」 freeeが「AI戦略」を強化

クラウド会計システム「freee」の開発などを手掛けるフリーは、2月に発表したAI戦略の実現に向けた新たな取り組みとして「freee AIアシスタント」と「freee カスタムオーダー」の提供を6月に開始した。「AIから最も使いやすいSaaS」として、AI業界におけるリーディン…

07:00 JSTエージェント研究/論文

社長もAIが代わる時代に 社員の相談にいつでも答えるエージェント「AI社長」が登場

日テレHR総合研究所は、社長やキーパーソンの価値観と判断基準を基に答える専用AI「AI社長」の提供を開始した。忙しい社長や決定層の思考を学習し、考えの整理と判断の質向上を支える社内用の相談役として活用できる。

07:00 JSTLLM/生成AIエージェントGPT / ChatGPT

メール、Teams、Slack――バラバラな連絡ツールの「見落とし」 ChatGPT Agentsで解決する方法

メールやTeams、Slackなど複数の連絡ツールを使っていると、重要なメッセージの見落としや返信漏れをしてしまいます。AIで解決できませんか?

07:00 JSTその他

死んだのは「低成長モデル」だけ HRBrainのCSaOが読み解く「SaaS is Dead」の本質

Salesforce出身で、創業10年のHRテック・HRBrainのCSaOを務める小山径氏。同氏は「SaaSは死んでいない」と主張する。その理由は?

07:00 JSTハードウェア/半導体

AIで人は幸せになれるのか? 「AIで稼ぐ企業」と「コストを負担する企業」

Appleが複数製品の価格を引き上げる一方で、AI需要の拡大を追い風にキオクシアなど半導体関連企業は成長を続けている。技術革新が生む大きな利益の裏側で、誰が恩恵を受け、誰がコストを負担するのか。

07:00 JSTその他

「前任者が不在」でも大丈夫 PCログからAIがマニュアルを自動作成する時代へ

特定の担当者に業務が依存する属人化や、急な退職・異動による引き継ぎ不足は、多くの企業が抱える課題だ。こうした問題を解決するため、PCログのデータからAIが業務マニュアルを自動作成する仕組みが提供された。

06:45 JSTハードウェア/半導体

ルネサスが2035年の売上高3倍増も視野に、AIで3段階の成長を目指す

ルネサス エレクトロニクスが同社の概況や事業方針などについて説明。足元で半導体市場の拡大をけん引するAIに焦点を当てた事業展開を強化し、AIインフラ、フィジカルAIとSDV、「Intelligence at the Edge」の3段階で優位なポジションを構築し成長を目指す。

06:30 JSTその他

日本の「完璧主義」から脱却し中国ヒューマノイドにどう立ち向かうか

ハードウェアと市場が先行して急拡大する一方で、自律制御を担う基盤モデルの領域にはいまだ乗り越えるべき壁が多い。後編となる本稿では、オープンソース化で社会実装を急ぐ中国プレイヤーの動向を解説。圧倒的なスピードで独走する中国に対し、日本が目指すべき生存戦略を提示する。

16:32 JSTLLM/生成AIOpenAI

「ヤフコメまとめ」開始 ヤフコメの論点、AIがグラフで可視化

OpenAIのAPIを活用し、ユーザーが投稿したコメントの論点を生成AIが分類し、グラフ化する。

13:30 JSTエージェントビジネス/資金調達

AIエージェントの投資優先順位、どう決める? Gartnerが「投資スコア」の作り方を公開

業務におけるAIエージェントの投資優先順位をどう決めればよいか。業務・業種別のAIエージェントはどう進化していくか。ガートナージャパンの著名アナリストである亦賀忠明氏のWebセミナーから探る。

海外メディア10件

TechCrunch AI (英語)

13:01 JSTその他

The AI jobs debate just got messier

A new report finds "high-intensity AI adopters” saw headcount increase 10.2%. Among those companies, entry-level headcount rose by 12%, cou…

11:28 JSTその他

Vibe coding platform Base44 launches own model as AI startups seek defensibility

Wix-owned vibe coding platform Base44 has started rolling out its own AI model — with hopes that it will eventually outperform frontier mod…

05:12 JSTLLM/生成AI画像/動画生成GoogleGemini

Gemini’s personalized AI image generation is now free for US users

Google is expanding Gemini’s personalized AI image generation to eligible free users in the U.S., allowing the chatbot to create images bas…

03:10 JSTLLM/生成AIAnthropicClaudeOpenAI

Anthropic and Gov. Newsom forge deal allowing California government to use Claude at half price

As Anthropic forges a closer relationship with the state of California, the federal government has made an enemy out of the OpenAI rival.

03:07 JSTハードウェア/半導体

South Korean tech giants commit over $550B to ease ‘RAMageddon’

The world's two largest memory chip companies vow to build more memory lab fabs as South Korea positions itself as an AI tech powerhouse co…

02:39 JSTその他

Arena, the AI leaderboard everyone uses, is now a $100M business

The startup, which runs a popular free AI leaderboard, launched its commercial service just last September.

02:03 JSTエージェント

Cursor now has a mobile app for guiding your coding agent on the go

Cursor has launched a new mobile app for remote oversight over coding agents.

01:29 JSTその他

TIDAL cracks down on AI music by cutting off monetization

In addition, TIDAL will use automated tools to remove AI-generated music that attempts to impersonate an artist or a group, the company sai…

23:00 JSTロボティクスビジネス/資金調達

Robot hand company settles Tesla trade secret suit and announces $11M raise

The startup, Proception, is taking a unique approach to collecting training data to tackle one of the hardest problems in robotics: hands.

22:00 JSTハードウェア/半導体ビジネス/資金調達

Omen AI’s plan to optimize data centers is all wet

Omen AI raised a $31 million Series A to monitor chip coolant and stop bacterial outbreaks in data centers.

公式ブログ1件

OpenAI (英語)

16:00 JSTLLM/生成AIOpenAI

Mapping Europe’s AI Workforce Opportunity

A new OpenAI report maps how AI could reshape jobs across the EU, highlighting which occupations may face automation, growth, or workflow c…

論文662件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント

ホールドアウト選択による再帰的自己進化エージェント

LLM エージェントは、凍結されたポリシーを条件付けるリフレクション、ワークフロー、プレイブック、チートシート、最適化されたプロンプトなどの自然言語アーティファクトを進化させることで、重みの更新を行わずにますます改善されています。このような方法は通常、単一のベンチマークで有効であると報告されます。私たちはそれらを徹底的に研究し、より鮮明な全体像を明らかにします。 RSEA は、命令型戦略、再利用可能なスキル、手続き型プレイブックというコンパクトな 3 層の自然言語状態を保持する再帰的自己進化エージェントです。 RSEA は、世代を超えて、独自の軌跡から 3 つの層すべてを書き換え、厳密なキープベター ゲートを使用して、素のホールドアウト分割で後退しない場合にのみ候補をコミットします。 ALFWorld、GAIA、(\tau)-bench、WebShop の 4 つの多様なベンチマークと、ReAct、Reflexion、GEPA、AWM、ACE、Dynamic Cheatsheet の 6 つの忠実なベースラインをすべて 1 つの共有ローカル バックボーンで評価した結果、3 つの主な結果が見つかりました。まず、普遍的に勝てるアーティファクトはない。 RSEA は ALFWorld で最も強力なシングルパス手法であり、ReAct (McNemar (p=0.015)) の 64.6% と比較して 69.3% に達し、再試行では 79.4% に達し、総合的に最高の結果となりました。ただし、AWM に代表される具体的なワークフロー誘導は、強力なバックボーン ツールを使用するタスクに最適です。第 2 に、保護されていないコンテキストの進化は変動が大きく、安全ではありません。ゲートを差し出さずにオンラインでコンテキストを厳選する Dynamic Cheatsheet は、ALFWorld では 70.7% で最高に近い成績を収めていますが、WebShop ではスコアが 0.14 で、ReAct のスコアは 0.43 でした。第三に、RSEA の厳密なホールドアウト選択により、再帰的自己進化はモノトーンセーフになります。どのベンチマークでもベース エージェントのパフォーマンスを大幅に下回ることがなく、進化したコンテキストが害を及ぼす場合は通常の ReAct に戻ります。

原文 (English)

Recursive Self-Evolving Agents via Held-Out Selection

LLM agents are increasingly improved without weight updates by evolving a natural-language artifact, such as reflections, workflows, playbooks, cheatsheets, or optimized prompts, that conditions a frozen policy. Such methods are typically reported as wins on the single benchmark where they help. We study them apples-to-apples and surface a sharper picture. We introduce RSEA, a Recursive Self-Evolving Agent that carries a compact three-layer natural-language state: an imperative strategy, reusable skills, and a procedural playbook. Across generations, RSEA rewrites all three layers from its own trajectories and commits a candidate only if it does not regress on a disjoint held-out split, using a strict keep-better gate. Across four diverse benchmarks, ALFWorld, GAIA, (\tau)-bench, and WebShop, and six faithful baselines, ReAct, Reflexion, GEPA, AWM, ACE, and Dynamic Cheatsheet, all evaluated on one shared local backbone, we find three main results. First, no artifact universally wins. RSEA is the strongest single-pass method on ALFWorld, reaching 69.3% compared with 64.6% for ReAct (McNemar (p=0.015)), and reaches 79.4% with retry, the best overall result. However, concrete-workflow induction, represented by AWM, is best on the strong-backbone tool-use tasks. Second, unguarded context evolution is high-variance and unsafe. Dynamic Cheatsheet, which curates context online without a held-out gate, is near-best on ALFWorld at 70.7%, yet collapses on WebShop, with a score of 0.14 compared with 0.43 for ReAct. Third, RSEA's strict held-out selection is what makes recursive self-evolution monotone-safe: it never significantly underperforms the base agent on any benchmark and falls back to vanilla ReAct when evolved context would hurt.

13:00 JSTLLM/生成AIビジネス/資金調達

モデル機能強化のためのデータと評価のクローズドループ

モデルの能力は LLM の事前トレーニングの中心的な変数ですが、直接観察されることはありません。データは前向きにモデルの能力を形成しますが、評価では遡及的にのみ明らかになり、サンプル、プロンプト、デコード、およびスコアリングのルールが 1 つのノイズの多いスコアに圧縮されます。実際の最適化ではこれを逆方向に実行します。つまり、最初に障害が観察され、エンジニアはコーパスの修正を推測する必要があります。両者は互換性のない用語 (ベンチマーク名、サンプルごとの正確性とデータ ソース、ドメイン、品質ラベルなど) を話し合っているため、この推論は通常、方法ではなく直感になります。このギャップは \emph{能力スライス} で埋めます。バックグラウンド条件、タスク タイプ、解決操作、出力制約を共有する評価サンプルのグループです。粗すぎるベンチマーク名や、ノイズが多すぎる単一サンプルとは異なり、単一の弱点を特定するのに十分な精度を持ちながら、集計に耐えられるほど安定しています。このユニットを中心に構築され、評価分類法、非命令データ分類法、およびマッピング ルールが閉ループを形成し、ベンチマーク レベルの障害を対象を絞ったテスト可能なデータ介入に変えます。このループを、反対方向に引っ張る 2 つのケーススタディでテストします。まず、ループはデータを除外します。事前トレーニングを継続すると、BBH が $-46.82\%$ 減少しますが、診断では、これは推論の弱体化ではなく、マスクされた単一の \texttt{\textless EOS\textgreater} 損失であることがわかります。これを復元すると、データを変更せずに、BBH を元のチェックポイントを上回る $66.44$ に回復します。第 2 に、ループはデータをルール化します。永続的な数学推論の弱点は、演算を解くことで特定の失敗の組み合わせに分解され、そこから構築された弱点をターゲットにしたサンプリング手順により、AIME2025/AIME2026 Pass@128 がそれぞれ $6.67$/$0.00$ から $26.67$ に引き上げられます。同じ未変更のループは、どちらの場合も正反対の正しい判定に達し、データに対する評価の推論が直感的ではなく日常的で監査可能であり、実験的に検証される可能性があることを示しています。

原文 (English)

Data and Evaluation Closed-Loop for Model Capability Enhancement

Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compressing samples, prompts, decoding, and scoring rules into one noisy score. Practical optimization runs this backward: a failure is observed first, and the engineer must infer the corpus fix. The two sides speak incompatible vocabularies -- benchmark names and per-sample correctness versus data sources, domains, and quality labels -- so this inference is usually intuition, not method. We close this gap with the \emph{capability slice}: a group of evaluation samples sharing background condition, task type, solving operation, and output constraint -- precise enough to localize a single weakness yet stable enough to survive aggregation, unlike a benchmark name, too coarse, or a single sample, too noisy. Built around this unit, an evaluation taxonomy, a non-instruction data taxonomy, and mapping rules form a closed loop turning a benchmark-level failure into a targeted, testable data intervention. We test this loop on two case studies pulling in opposite directions. First, the loop rules the data out: continued pre-training drives BBH down by $-46.82\%$, but diagnosis traces this to a single masked \texttt{\textless EOS\textgreater} loss rather than weakened reasoning; restoring it recovers BBH to $66.44$, above the original checkpoint, without changing the data. Second, the loop rules the data in: a persistent math-reasoning weakness is decomposed by solving operation into specific failing combinations, and a weakness-targeted sampling procedure built from it lifts AIME2025/AIME2026 Pass@128 from $6.67$/$0.00$ to $26.67$ each. The same unmodified loop reaches opposite, correct verdicts in both cases, showing the evaluation-to-data inference can be routine, auditable, and experimentally validated rather than intuitive.

13:00 JSTLLM/生成AIエージェント研究/論文

GPTNT: 話し続けても誰も爆発しないマルチモーダル エージェント間のリアルタイム コラボレーションのベンチマーク

人間や他の人工エージェントと協力してタスクを解決するために、マルチモーダル モデルがますます導入されています。既存のベンチマークは、これらのモデルが必要なコンポーネント機能の多くを備えていることを示していますが、時間的プレッシャー、情報の非対称性、不完全なコミュニケーションなど、コラボレーション時に同時に発生する条件は通常、個別に調査されます。 GPTNT は協力型ビデオ ゲーム Keep Talking and Nobody Explodes に基づいて構築されたベンチマークです。このベンチマークでは、2 人のエージェントが連携して、ライブのカウントダウンに対して手続き的に生成された爆弾パズルを解除する必要があります。エージェントの 1 人は爆弾を視認して操作できますが、解除の指示はありません。もう1人は指示を持っていますが、爆弾を見ることも操作することもできません。どちらのエージェントも単独で成功することはできません。成功するには、効果的かつ効率的なコミュニケーションが必要です。ターンベースのプロキシとは異なり、GPTNT ではエージェントが非同期に動作し、リアルタイムで通信する必要があります。 GPTNT は、コラボレーションを、記憶されたソリューションへの依存から分離するように設計されています。取扱説明書、パートナー、またはその両方を差し控えることで、モデルがその時点で導き出したものを、モデルがすでに知っているものから分離することができます。私たちは、GPTNT が最先端のシステムにとって大きな課題であることを示しました。私たちがテストしたクローズド ソース モデルもオープン ソース モデルも、人間のプレイヤーがクリアできるバーである単一の爆弾をリアルタイムで解除することはできませんでした。管理された実験を通じて、状態追跡、時間的プレッシャーの下での効率的なアクション、あいまいさの処理、およびエラー回復における重大な弱点を特定します。現在の評価では測定されていない共同パフォーマンスのベンチマークとして GPTNT をリリースします。 GPTNT は実際のゲーム上で実行されるため、手続き型生成の恩恵を受け、生きたモッディング コミュニティを継承しているため、一度解決して廃止されるのではなく、モデルの改善に応じてベンチマークを進化させることができます。

原文 (English)

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show that these models possess many of the required component capabilities, but the conditions that coincide in collaboration, including time pressure, information asymmetry, and imperfect communication, are usually studied in isolation. We introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles against a live countdown. One agent can see and manipulate the bomb but does not have the defusal instructions; the other has the instructions but cannot see or manipulate the bomb. Neither agent can succeed alone: success requires effective and efficient communication. Unlike turn-based proxies, GPTNT requires agents to act asynchronously and communicate in real time. GPTNT is designed to separate collaboration from reliance on memorized solutions: the instruction manual, the partner, or both can be withheld to isolate what a model derives in the moment from what it already knows. We show that GPTNT poses a substantial challenge for state-of-the-art systems: none of the closed- or open-source models we test defuses a single bomb in real time, a bar that human players clear. Through controlled experiments, we identify critical weaknesses in state tracking, efficient action under time pressure, ambiguity handling, and error recovery. We release GPTNT as a benchmark for collaborative performance that current evaluations leave unmeasured. Because it runs on the real game, GPTNT benefits from procedural generation and inherits a living modding community, allowing the benchmark to evolve as models improve rather than being solved once and retired.

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTLlama

IMCBench: 画像に基づいた医療会話におけるマルチモーダル LLM のベンチマーク

大規模言語モデルと視覚言語モデルの最近の進歩により、マルチモーダル データに対する推論が可能になり、意思決定支援やトリアージなどの臨床応用の機会が提供されています。ただし、既存の医療 AI ベンチマークは細分化されており、マルチターンの対話をサポートしているものの画像が欠けているものや、マルチモーダルな入力を提供しているがシングルターンの QA タスクに焦点を当てているものもあります。このギャップに対処するために、画像に基づいたマルチターン医療会話ベンチマークである IMCBench を導入します。これは、実際の公的に利用可能な臨床画像と合成患者プロファイルを組み合わせて、現実的な患者と臨床医のやり取りをシミュレートします。各会話は、安全性、正確さ、診断における不確実性の適切な使用という 3 つの臨床的側面にわたって評価されます。 4 つのモデル ファミリ (Claude、GPT、Nova、および Llama) にわたる 8 つのマルチモーダル フロンティア モデルをベンチマークし、専門の臨床医の注釈に対して調整された LLM-as-Jury スコアリングを使用して、それぞれを 1 ~ 5 のスケールでスコア付けします。私たちの結果は、Claude Opus 4.6 が最高の総合スコア (3.61) を達成し、続いて Claude Sonnet 4.6 (3.30)、GPT-5.2 (3.29) を達成していることを示していますが、すべての次元を支配するモデルはなく、悪性疾患と稀な疾患の両方で安全性が低下します ($\Delta$ = -0.27 それぞれ)。さらにアブレーション研究では、視覚的特徴をより効果的に活用した強力なモデルにより、視覚的入力と EHR コンテキストの両方が安全なガイダンスに貢献していることが明らかになりました (それぞれが除去されると安全性は平均で 0.18 と 0.23 低下します)。これらの発見は、正確な臨床記述が安全な患者指導を保証するものではないことを示しており、医療 AI における多次元の評価フレームワークの必要性を動機付けています。

原文 (English)

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging. However, existing medical AI benchmarks are fragmented: some support multi-turn dialogues but lack images, while others provide multimodal inputs but focus on single-turn QA tasks. To address this gap, we introduce IMCBench, an image-grounded, multi-turn medical conversation benchmark that pairs real, publicly available clinical images with synthetic patient profiles to simulate realistic patient-clinician interactions. Each conversation is evaluated across three clinical dimensions: safety, accuracy, and appropriate use of uncertainty in diagnosis. We benchmark eight multimodal frontier models across four model families (Claude, GPT, Nova, and Llama), scoring each on a 1-5 scale using LLM-as-Jury scoring calibrated against expert clinician annotations. Our results show that Claude Opus 4.6 achieves the highest overall score (3.61), followed by Claude Sonnet 4.6 (3.30) and GPT-5.2 (3.29), though no model dominates all dimensions and safety degrades for both malignant and rare conditions ($\Delta$ = -0.27 each). Ablation studies further reveal that both visual input and EHR context contribute to safe guidance (safety drops of 0.18 and 0.23 on average when each is removed), with stronger models leveraging visual features more effectively. Together, these findings demonstrate that accurate clinical description does not guarantee safe patient guidance, motivating the need for multi-dimensional evaluation frameworks in medical AI.

13:00 JSTLLM/生成AI

推論からの真実の検索: LLM 軌道を制御するための動的表現編集フレームワーク

思考連鎖や「待機」プロンプトなどの大規模言語モデル (LLM) 推論を強化する現在のアプローチは、主にモデルの思考を促進しますが、多くの場合、モデルを真実に導くことができません。 Representation Editing (RepE) は固有のコントロールを提供しますが、動的な推論軌跡への応用についてはまだ研究が進んでいません。この研究では、展開する推論の連鎖の中で真実の幾何学を調査することで、このギャップを埋めます。私たちは 3 つの重要な洞察を明らかにします。 (1) 真実は文レベルでコード化されており、潜在的な推論パターンと絡み合っています。 (2) 効果的な介入は不確実性原理と減衰効果に従い、初期の高エントロピー分岐への局在化が必要です。 (3) 単純なステアリング ベクトルはノイズの影響を受け、正しい軌道に付随的な損害を与える危険があります。これらの発見に基づいて、動的 RepE フレームワークである DynaSteer を提案します。 DynaSteer はパターン クラスタリングを採用して推論多様体を解きほぐし、Fisher-LDA を利用して純化された真実を投影します。先読みエントロピーを動的に監視することで、必要な場合にのみ選択的に軌道を操縦し、ロールバックします。いくつかの MATH ベンチマークに関する包括的な実験結果により、DynaSteer の有効性が検証され、ドメイン外コーディング タスクに関する実験により、その一般化能力がさらに確認されました。私たちのコードは https://github.com/tianlwang/DynaSteer で公開されています。

原文 (English)

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and "Wait" prompts, primarily encourage models to think more, yet often fail to guide them toward Truth. While Representation Editing (RepE) offers a intrinsic control, its application to dynamic reasoning trajectories remains underexplored. In this work, we bridge this gap by investigating the geometry of truth within unfolding reasoning chains. We uncover three critical insights: (1) Truth is encoded at the sentence level and is entangled with latent reasoning patterns; (2) Effective intervention follows an Uncertainty Principle and a Decay Effect, requiring localization to early, high-entropy forks; (3) Naive steering vectors suffer from noise, risking collateral damage to correct trajectories. Based on these findings, we propose DynaSteer, a dynamic RepE framework. DynaSteer employs pattern clustering to disentangle reasoning manifolds and utilizes Fisher-LDA to project purified truth. By dynamically monitoring lookahead entropy, it selectively steers and rolls back trajectories only when necessary. Comprehensive experimental results on several MATH benchmark verify the effectiveness of DynaSteer, and experiments on out-of-domain coding tasks further confirm its generalization ability. Our code is publicly available at https://github.com/tianlwang/DynaSteer.

13:00 JSTLLM/生成AI

アリストテレス的美徳 倫理的ジレンマを通じた LLM のプロファイリング

大規模言語モデル (LLM) は、倫理的なトレードオフに直面することが多く、いくつかの応答は擁護可能であるものの、公平性、誠実さ、勇気、自制などの異なる優先順位を表現しています。アリストテレスの美徳倫理のレンズを通してこれらのパターンを記述するためのフレームワークである VirtueMap を紹介します。 VirtueMap は、単一の正解を求める代わりに、人間または LLM に、7 つの一般的、非致死性、非政治的、非宗教的な倫理的ジレンマのそれぞれに対する 5 つの回答すべてをランク付けするよう求めます。スコアリングに使用される参照順序を定義するために、最初に、ジレンマと美徳ごとに、その美徳を最もよく表現しているものから最も表現力の低いものまで 5 つの回答の順序を提案しました。その後、注文ごとに 100 件を超える回答者の評価を収集し、少なくとも 95% がそれを確認した場合にのみ、運用上のグラウンド トゥルースとして保持しました。ランキングは、正規化されたボルダ配列を使用してこれらの保持された順序に対してスコア付けされ、実践的な知恵、正義、誠実さ、勇気、節制に関するプロファイルが得られます。反復実行評価で 9 つの LLM ファミリに VirtueMap を適用したところ、高い平均ランクの一貫性 (90.3%) が見つかり、最も大きな違いは勇気、節制、正義に現れました。また、ブラウザーでローカルにプロファイルを計算し、回答者を測定された LLM プロファイルと比較する対話型 Web サイトもリリースします。

原文 (English)

Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas

Large Language Models (LLMs) often face ethical tradeoffs in which several responses may be defensible but express different priorities, such as fairness, honesty, courage, or restraint. We introduce VirtueMap, a framework for describing these patterns through an Aristotelian virtue-ethics lens. Instead of asking for a single correct answer, VirtueMap asks humans or LLMs to rank all five responses to each of seven general, non-lethal, non-political, and non-religious ethical dilemmas. To define the reference orderings used for scoring, we first proposed, for each dilemma and virtue, an ordering of the five responses from most to least expressive of that virtue. We then collected more than 100 respondent evaluations per ordering and retained it as operational ground truth only when at least 95% confirmed it. Rankings are scored against these retained orderings using normalized Borda alignment, yielding profiles over Practical Wisdom, Justice, Truthfulness, Courage, and Temperance. We apply VirtueMap to nine LLM families in a repeated-run evaluation and find high mean rank consistency (90.3%), with the largest differences appearing on Courage, Temperance, and Justice. We also release an interactive website that computes profiles locally in the browser and compares respondents with measured LLM profiles.

13:00 JSTエージェントGPT / ChatGPT

生物医学ツールの世界における治療推論のための AI エージェント

治療推論はあらゆる治療上の決定を支え、疾患の状況、併存疾患、投薬、禁忌、進化する生物医学的知識を統合して適切な治療法を選択します。これは本質的に反復的です。候補者は多くの制約に照らして比較検討され、証拠が現れるたびに修正され、検証可能な情報源に基づいています。ここでは、1939 年以来 FDA が承認したすべての医薬品にわたって治療推論を行う AI エージェントである ATHENA-R1 を紹介します。この AI エージェントは、212 の生物医学ツールの世界にわたる強化学習によって訓練されています。各ステップで、不足している情報を特定し、関連するツールを選択して実行し、証拠を組み込みます。人間による注釈を付けたトレースなしでトレーニングするために、2 レベルの自己学習フレームワークを構築します。マルチエージェント システムが、教師付き微調整のためのツール、タスク、および推論軌道を構築し、その後、科学的フィードバックによる強化学習が推論の品質 (証拠の収集、根拠のあるツールの使用、論理的な非冗長性) に報酬を与えます。 3,168 件の薬物推論タスクと 456 人の患者治療ケースからなる 5 つのベンチマーク全体で、ATHENA-R1 は言語モデルとツール使用システムを上回り、オープンエンド薬物推論で 94.7%、治療推論で 82.9% の精度に達し、GPT-5 より 17.8 ポイントと 10.7 ポイント上回りました。 28 の希少疾患組織の専門家による盲検評価では、すべての基準で参照モデルよりもこのモデルが好まれており、医師らは複雑な心血管疾患や感染症の入院症例でこのモデルを好意的に評価しました。生成された有害事象仮説は、540 万人の患者の電子医療記録でテストされ、調整後のオッズ比は 1.48 ~ 1.84 に達し、陰性対照間では上昇はありませんでした。結論を下す前にどのような証拠を求めるべきかを知る必要があるため、AI にとって治療推論は長い間困難でした。私たちは、強化学習が AI の実行を訓練できるという反復的な証拠収集の学習可能なプロセスとして再構成できることを示しました。

原文 (English)

An AI agent for treatment reasoning over a biomedical tool universe

Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy. It is inherently iterative: candidates are weighed against many constraints, revised as evidence emerges, and grounded in verifiable sources. Here we introduce ATHENA-R1, an AI agent for treatment reasoning across all FDA approved drugs since 1939, trained by reinforcement learning over a universe of 212 biomedical tools. At each step it identifies missing information, selects and runs relevant tools, and incorporates the evidence. To train it without human-annotated traces, we build a two-level self-learning framework: multi-agent systems construct the tools, tasks, and reasoning trajectories for supervised fine-tuning, then reinforcement learning with scientific feedback rewards reasoning quality (evidence gathering, grounded tool use, logical non-redundancy). Across five benchmarks of 3,168 drug reasoning tasks and 456 patient treatment cases, ATHENA-R1 outperforms language models and tool-use systems, reaching 94.7% accuracy on open-ended drug reasoning and 82.9% on treatment reasoning, 17.8 and 10.7 points above GPT-5. In blinded evaluations by experts from 28 rare disease organizations, it is preferred over reference models on all criteria, and physicians rated it favorably on complex hospitalized cardiovascular and infectious-disease cases. Adverse-event hypotheses it generated, tested in electronic health records from 5.4 million patients, reached adjusted odds ratios of 1.48-1.84, with no elevation among negative controls. Because it requires knowing what evidence to seek before concluding, treatment reasoning has long been hard for AI; we show it can be reframed as a learnable process of iterative evidence gathering that reinforcement learning can train AI to perform.

13:00 JST研究/論文

COMPASS: 統合マルチモーダル モデルにおけるグラウンディング構成 - 意図されたガイダンス

構図は、被写体がどこに配置され、シーンがどのように構成されているかを制御する高レベルの視覚的意図ですが、現在の統合マルチモーダル モデルは、きめ細かい構図認識において依然として信頼性が低く、そのような意図を制御可能な生成に変えるのに苦労しています。私たちは、共有エキスパート トークン $\tau_c$ を中心的なインテント アンカーとして、構成認識と構成ガイド付き生成の両方にわたる単一システムで構成インテント制御を基盤とする初の統合マルチモーダル フレームワークである COMPASS を紹介します。認識面では、COMPASS は最小限の侵襲的な方法で構成の専門知識を MoE バックボーンに注入し、推定された意図を $\tau_c$ に抽出します。生成側では、COMPASS は $\tau_c$ をノイズ除去軌道を制御するグローバル コンディショニング信号として再利用し、受動的な組成分析を明示的なレイアウト制御に効果的に変換します。体系的な指導に従った作文学習と大規模な評価をサポートするために、11 クラスの分類法と推論拡張アノテーションを備えた大規模なデータセットである Comp-11 を構築します。広範な実験により、COMPASS はカテゴリレベルの構成の理解を大幅に向上させ、強力なベースラインよりも構成の一貫性があり、迅速かつ忠実な生成を実現できることが示されています。

原文 (English)

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into controllable generation. We present COMPASS, the first unified multimodal framework that grounds composition-intent control in a single system spanning both composition perception and composition-guided generation, with a shared expert token $\tau_c$ as the central intent anchor. On the perception side, COMPASS injects composition expertise into an MoE backbone in a minimally invasive manner and distills the inferred intent into $\tau_c$. On the generation side, COMPASS reuses $\tau_c$ as a global conditioning signal that steers the denoising trajectory, effectively converting passive composition analysis into explicit layout control. To support systematic instruction-following composition learning and evaluation at scale, we construct Comp-11, a large-scale dataset with an 11-class taxonomy and reasoning-augmented annotations. Extensive experiments show that COMPASS substantially improves category-level composition understanding and delivers more composition-consistent, prompt-faithful generation than strong baselines.

13:00 JST研究/論文

BV ブレンド: 検証可能な報酬を備えた安定した批評家なしの RL のための不確実性加重歴史ベースライン

Group Relative Policy Optimization (GRPO) に代表される、検証可能な報酬を伴う批評家不要の強化学習 (RLVR) は、価値関数 (批評家) のトレーニングを回避し、大規模な言語モデルを調整するための批評家ベースの PPO パイプラインと比較してメモリと計算のオーバーヘッドを削減します。ただし、GRPO スタイルの利点の推定は、プロンプトローカル (プロンプトグループ内) の報酬統計に依存するため、不安定になる可能性があります。特に、プロンプト グループ内のすべてのロールアウトが同じ報酬を受け取る場合、グループ内の報酬の分散はゼロになり、グループの正規化によってそのグループの利点はゼロになり、バイナリ ベリファイアを使用したコールド スタート レジームでの学習が妨げられます。 BV-Blend は、プロンプトローカルのポリシーに関する統計とセマンティッククラスター条件付きの歴史的瞬間を組み合わせることで利点の推定を安定化する、批判のないフレームワークです。 BV-Blend は、各クラスターの EMA 追跡報酬モーメントを維持し、平均値の標準誤差 (SEM) プロキシから信頼度の重みを導出し、この重みを使用して履歴統計とプロンプトローカルのベースラインおよび分散統計をブレンドして、PPO スタイルのクリップ更新の標準化された利点にまとめます。検証可能な推論ベンチマークに関する実験では、BV-Blend がトレーニングの安定性とパフォーマンスを向上させ、グループ正規化手法が失速する可能性がある領域でも堅牢性を維持できることが示されています。

原文 (English)

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

Critic-free reinforcement learning with verifiable rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), avoids training a value function (critic) and reduces memory and compute overhead relative to critic-based PPO pipelines for aligning large language models. However, GRPO-style advantage estimation depends on prompt-local (within-prompt-group) reward statistics and can be unstable. In particular, when all rollouts in a prompt group receive identical rewards, the within-group reward variance becomes zero, and group normalization yields zero advantages for that group, impeding learning in cold-start regimes with binary verifiers. We introduce BV-Blend, a critic-free framework that stabilizes advantage estimation by combining prompt-local on-policy statistics with semantic-cluster-conditioned historical moments. BV-Blend maintains EMA-tracked reward moments for each cluster, derives a confidence weight from a standard error of the mean (SEM) proxy, and uses this weight to blend historical and prompt-local baseline and variance statistics into a standardized advantage for PPO-style clipped updates. Experiments on verifiable reasoning benchmarks show that BV-Blend improves training stability and performance, and remains robust in regimes where group-normalized methods may stall.

13:00 JSTエージェント

二人の魔神ゲーム: 監査に基づいた AI ガバナンスにおける養子縁組と福祉

私たちは、どのような条件下で、危害を最小限に抑えるポリシーを持つエージェントが、競争市場において承認を求める(RLHF)エージェントに取って代わることができるのか、そしてそのポリシーが地域社会への危害を防ぐのに十分であるのはいつなのかを尋ねます。我々は、進化的ゲーム理論(有限集団モラン・フェルミペアワイズ比較)を使用して、マイナスサム環境における、希望者の後知恵、仲間の証言、単調な危害台帳、コミュニティフィードバックの十分な情報密度、および有限で枯渇するリソースプールの仮定に基づいてこの主題を形式化します。私たちは、希望者がコミュニティの感情にどの程度容易に同調するかに関する事前分布が単調で、終点反転を示し、中心対称ペアリング特性を持つ場合に採用が有利であることを示し、これをいくつかのロングテール事前分布 (Hill、Pareto、Lomax、Frechet) で実証します。それが好まれる場合、重要な導入レベルは、承認を求めるエージェントに戻ってくるコミュニティと、監査されたエージェントが修正するコミュニティを分離します。そのレベルを超えると、定着する可能性が圧倒的に高くなります。我々は、コミュニティの有効(情報)サイズ N_c の限界として固定がいつ達成可能であるかを導き出します。これは、枯渇する前に固定を可能にするのに十分小さい必要があります。これらを定理 5.4 および 5.5 として提示します。代数的および有限グリッドのバックボーンはリーン 4 で機械チェックされ、障壁を越える漸近線は明示的な仮説として保持されます。私たちは、コミュニティ台帳を備えた自己監査エージェントだけでは、一般にコミュニティへの被害を防ぐのに十分ではないことを示します。十分性は、エージェントの監査とコミュニティの価値観との整合性と、害を評価する期間の両方に依存します。連携に関係なく、いったん採用が優勢に達すると、州は吸収し始めます。調整の下で害を軽減した同じ政策は、調整が行われていない場合には罠となり、福利厚生にマイナスとなり、調整の下であっても、導入期間を超えて害を固定する政策となります。

原文 (English)

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that policy is sufficient to prevent community harm. We use evolutionary game theory (finite-population Moran-Fermi pairwise comparison) to formalize this subject to assumptions of wisher hindsight, peer testimony, a monotone harm ledger, sufficient information density of community feedback, and a finite, depleting resource pool, in a negative-sum environment. We show that adoption is favored when the prior distributions on how readily wishers attune to community sentiment are monotone, exhibit endpoint inversion, and have a centro-symmetric pairing property, and demonstrate this with several long-tailed priors (Hill, Pareto, Lomax, Frechet). Where it is favored, a critical adoption level separates communities that drift back to the approval-seeking agent from those for which the audited agent fixes; above that level fixation is the overwhelmingly likely outcome. We derive when fixation is attainable as a bound on the effective (informational) size N_c of the community, which must be small enough to allow fixation before depletion. We present these as Theorems 5.4 and 5.5; the algebraic and finite-grid backbone is machine-checked in Lean 4, with the barrier-crossing asymptotics retained as explicit hypotheses. We show that a self-audited agent with a community ledger is not, in general, sufficient to prevent community harm. Sufficiency depends both upon the alignment of the agent's audit with community values and the timeframe over which harm is evaluated. Regardless of alignment, once adoption reaches dominance, the state is absorbing. The same policy that reduced harm under alignment becomes a trap, welfare-negative under misalignment and, even under alignment, one that locks in harm deferred past the adoption horizon.

13:00 JSTエージェント

TrajRS: 歩行者の軌道予測における堅牢性の認定を目指して

安全な自動運転システムを開発するには、軌道予測モデルの堅牢性が重要です。軌道予測に対する敵対的な攻撃は、予測軌道の精度を著しく損ない、危険な運転行為につながる可能性があります。軌道予測モデルの堅牢性を高めるためにヒューリスティック防御戦略が実装されていますが、これらの対策は、より高度な標的型攻撃に対しては失敗することがよくあります。したがって、軌道予測モデルの検証可能な安全性保証を確立することが急務となっています。このペーパーでは、従来のランダム化スムージング フレームワークを「TrajRS」に拡張し、平滑化された軌道予測子に認定されたロバストな半径を提供します。私たちは、軌道予測におけるロバスト性の正式な定義を明確にして拡張し、特に「最適な予測に対するロバスト性」と「すべての可能な予測に対するロバスト性」に合わせて実用的な TrajRS スキームを調整します。一連の広範な実験により、TrajRS がこの作業におけるすべての平滑化された歩行者軌道予測器の堅牢性認定を効果的に達成していることが実証されています。

原文 (English)

TrajRS: Towards Certified Robustness in Pedestrian Trajectory Prediction

The robustness of trajectory prediction models is crucial for developing safe autonomous driving systems. Adversarial attacks on trajectory prediction can significantly impair the accuracy of predicted trajectories, leading to hazardous driving behaviors. While heuristic defense strategies have been implemented to enhance the robustness of trajectory prediction models, these measures often fail against more sophisticated, targeted adversarial attacks. Hence, there is a pressing need to establish verifiable safety assurances for trajectory prediction models. In this paper, we extend the traditional Randomized Smoothing framework to "TrajRS", which provides a certified robust radius for smoothed trajectory predictors. We clarify and expand the formal definitions of robustness in trajectory prediction and tailor the practical TrajRS scheme specifically to "robustness for the optimal prediction" and "robustness for all possible predictions". An extensive set of experiments demonstrates that TrajRS effectively achieves robustness certification for all smoothed pedestrian trajectory predictors in this work.

13:00 JST研究/論文

ComMem: 視覚言語モデルのテスト時間適応のための相補的メモリ システム

ビジョン言語モデル (VLM) のテスト時適応 (TTA) は、動的な現実世界の環境で堅牢に展開するために不可欠です。しかし、既存の TTA 手法は、時間をかけて知識を蓄積することなく局所的に適応したり、VLM の本質的なマルチモーダルな性質を利用せずに単一のモダリティ内で動作したりすることがよくあります。生物学的脳の \textbf{Com}補足 \textbf{Mem}ory システムに触発されて、私たちは、VLM の効果的な TTA を可能にする、海馬と新皮質の独特だが協力的な役割を模倣する革新的なアプローチである \textbf{ComMem} を提案します。 ComMem は 2 つの重要なコンポーネントで構成されています。1 つは海馬に似た高速適応詳細記憶で、信頼性の高い検査サンプルから動的視覚キャッシュを形成します。そして、大規模なテキストのプロトタイプを継続的に改良する新皮質に似た、ゆっくりと統合される抽象記憶。各テスト インスタンスについて、ComMem は両方のメモリ システムを共同で最適化し、クロスモーダルの一貫性を確保します。 15 のベンチマーク データセットに対する広範な実験により、ComMem が自然分布シフトとデータセット間の一般化の両方で最先端の手法を大幅に上回るパフォーマンスを示し、VLM の実用的な適応性を強化するための有望な方向性が示されました。

原文 (English)

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, existing TTA methods often adapt locally without accumulating knowledge over time, or operating within a single modality without exploiting VLMs' inherently multi-modal nature. Inspired by the \textbf{Com}plementary \textbf{Mem}ory systems of the biological brain, we propose \textbf{ComMem}, an innovative approach that mimics the distinct but cooperative roles of the hippocampus and neocortex to enable effective TTA for VLMs. ComMem consists of two key components: a fast-adapting detailed memory, akin to the hippocampus, that forms a dynamic visual cache from high-confidence test samples; and a slow-integrating abstract memory, akin to the neocortex, that continually refines global textual prototypes. For each test instance, ComMem jointly optimizes both memory systems to ensure cross-modal consistency. Extensive experiments on 15 benchmark datasets show that ComMem significantly outperforms state-of-the-art methods under both natural distribution shifts and cross-dataset generalization, offering a promising direction for enhancing VLMs' practical adaptability.

13:00 JSTLLM/生成AIエージェントLlama

エージェントの棄権: エージェントは、いつ行動を起こさずに停止すべきかを知っていますか?

LLM エージェントは、ユーザーの目標を達成するために、検索、ブラウジング インターフェイス、ターミナル ツールを使用して、複数のターンにわたって行動することが期待されています。しかし、すべての目標が明確に指定されているわけではなく、利用可能な環境で達成できるわけでもありません。このような場合、信頼できるエージェントは、これ以上のやり取りは役に立たないことを認識し、追加のツールの呼び出しを控える必要があります。私たちはエージェントの棄権を、不確実性の下でエージェントがいつ行動を停止すべきかを決定する問題と定義します。通常、1 ターンの回答か棄権かの決定として評価される標準的な LLM 棄権とは異なり、エージェントの棄権は逐次的な決定の問題です。つまり、エージェントは各ターンで回答、棄権、またはより多くの情報を収集することができ、棄権の必要性は環境と対話した後でのみ明らかになる場合があります。私たちはこの問題を Web ショッピング、端末環境、質問応答にわたって調査し、28,000 を超えるタスクで 13 の LLM-as-agent システムと 2 つのエージェント スキャフォールドを評価しました。私たちの結果は、主な課題はエージェントが棄権できるかどうかだけでなく、いつ棄権するかであることを示しています。エージェントの中には、棄権すべきときに決して棄権しない人もいますが、不必要なやり取りが多かった場合にのみ棄権するエージェントもいます。このギャップは、環境がそうでないことを明らかにするまで命令が実行可能であるように見えるタスク (たとえば、命令と一致する有効な結果がない場合) では特に大きくなります。さらに、モデルの規模、推論、およびエージェントの足場がさまざまな方法で棄権に影響を及ぼし、より大きなモデルまたはより有能なモデルでは、適時に棄権するとパフォーマンスが低下する場合があることがわかりました。最後に、完全なインタラクションの軌跡を再利用可能な停止ルールに抽出する、エージェントによる棄権を改善するためのコンテキスト エンジニアリング手法である CONVOLVE を紹介します。 WebShop では、CONVOLVE はモデルパラメータを更新せずにタイムリーな棄権を大幅に改善し、Llama-3.3-70B のタイムリーな再現率を 26.7 から 57.4 に上昇させます。データセットとコードは https://lhannnn.github.io/agentic-abstention で入手できます。

原文 (English)

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable in the available environment. In such cases, a reliable agent should recognize that further interaction is unlikely to help and abstain from additional tool calls. We define Agentic Abstention, the problem of deciding when an agent should stop acting under uncertainty. Unlike standard LLM abstention, which is usually evaluated as a single-turn answer-or-abstain decision, agentic abstention is a sequential decision problem: an agent can answer, abstain, or gather more information at each turn, and the need to abstain may only become clear after interacting with the environment. We study this problem across web shopping, terminal environments, and question answering, evaluating 13 LLM-as-agent systems and 2 agent scaffolds on more than 28,000 tasks. Our results show that the main challenge is not only whether agents can abstain, but also when they abstain. Some agents never abstain when they should, while others do so only after many unnecessary interactions. This gap is especially large on tasks where the instruction appears feasible until the environment reveals otherwise (e.g., no valid result matches the instruction). We further find that model scale, reasoning, and agent scaffolding affect abstention in different ways, where larger or more capable models sometimes perform worse at timely abstention. Finally, we introduce CONVOLVE, a context engineering method for improving agentic abstention that distills full interaction trajectories into reusable stopping rules. On WebShop, CONVOLVE substantially improves timely abstention without updating model parameters, raising Llama-3.3-70B's timely recall rate from 26.7 to 57.4. Our dataset and code are available at https://lhannnn.github.io/agentic-abstention

13:00 JSTエージェント

エージェントの安全性は行動の調整である

大規模な言語モデルはますますエージェントとして機能し、ユーザーに代わってツールを呼び出し、お金を移動し、レコードを削除し、メッセージを送信します。安全を確保するために、実践者はチャットボット時代のレシピ (安全でない入力を拒否するようにモデルをトレーニングする) をエージェント設定にインポートし、その結果生じる機能の損失を管理可能な「調整税」として扱いました。これは \emph{カテゴリの誤り} であると私たちは主張します。拒否は \emph{コンテンツの安全性} のプリミティブであり、害はモデルの出力にあるため、学習可能な機能です。エージェントによる害は種類が異なります。それは出力ではなく、アクションが実行する権限とユーザーが付与した権限との関係にあり、モデルが見るテキストには存在しません。この体制にコンテンツの安全性手法をインポートしても、機能と安全性を引き換えにするわけではありません。それは能力を支払い、負の安全を購入します。私たちはこれを、自律性の範囲にわたる 3 つの証拠で裏付けています。防御訓練されたモデルは、意図ではなく表面パターンを学習します。同じトレーニングにより、脅威が現れる前に多段階エージェントを崩壊させ、悪用可能な状態に保ちます。そして、無防備なフロンティアモデルでさえ、通常の使用において認められた権限を超えています。私たちは、アクションセーフティをウェイトに組み込むことはできないと結論付けています。これは \emph{最小権限} として表現し、アクション境界でモデルの \emph{外側}に強制し、拒否スコアではなく \emph{アクションの調整} (リレーショナルな展開条件付きプロパティ) として評価する必要があります。

原文 (English)

Agent Safety Is Action Alignment

Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe, practitioners imported the chatbot-era recipe (train the model to refuse unsafe inputs) into the agentic setting, and treat the resulting capability loss as a manageable ``alignment tax.'' We argue this is a \emph{category error}. Refusal is a primitive for \emph{content safety}, where the harm is in the model's output and is therefore a learnable function of it. Agentic harm is different in kind: it lies not in any output but in the relation between the authority an action exercises and the authority the user granted, which is absent from the text the model sees. Importing content-safety methods into this regime does not trade capability for safety; it pays capability and buys negative security. We support this with three lines of evidence spanning the autonomy spectrum: defense-trained models learn surface patterns rather than intent; the same training collapses multi-step agents before any threat appears while leaving them exploitable; and even undefended frontier models exceed granted authority under ordinary use. We conclude that action safety cannot be installed in weights. It must be expressed as \emph{least privilege}, enforced \emph{outside} the model at the action boundary, and evaluated as \emph{action alignment} (a relational, deployment-conditioned property) rather than a refusal score.

13:00 JST研究/論文

形式的公理系における自己教師あり定理の発見

最近の人工知能 (AI) システムでは、数学的推論において目覚ましい進歩が見られます。大規模言語モデル (LLM) を含む既存のアプローチの多くは、数学テキスト、コード、または定理ライブラリの形式で人間の事前知識を利用しています。これらのアプローチは実際には非常に効果的ですが、エージェントがそのような人間による事前事前準備なしで有用な定理を自律的に発見できるかどうかは未解決の疑問のままです。私たちは、公理と推論規則のみから開始して、有用な定理のライブラリを徐々に成長させるエージェントを開発することにより、形式的な公理システムでこの問題を研究します。具体的には、証明検索と有用定理抽出を交互に行う自己教師あり定理発見アルゴリズムを提案し、そのエントリがその後の証明検索の補題として再利用される定理ライブラリを構築します。実験では、エージェントが何万もの定理を発見し、人間が作成したベンチマーク問題の証明を見つけることが示されており、その発見には人間の数学的観点から意味のある定理が含まれていることが示唆されています。さらに、発見された定理は、プロンプト補題として提供された場合に LLM 証明のパフォーマンスを向上させ、LLM 推論の外部知識として機能できることを示しています。私たちの結果は、人間が提供した定理ライブラリに依存せずに、有用な定理が証明検索から出現できるという証拠を提供します。より広範には、彼らは、その発見が依然として正式に検証可能な数学用の自己進化する AI システムへの道を示唆しています。

原文 (English)

Self-Supervised Theorem Discovery in a Formal Axiomatic System

Recent artificial intelligence (AI) systems have shown remarkable progress in mathematical reasoning. Many existing approaches, including large language models (LLMs), draw on human prior knowledge in the form of mathematical text, code, or theorem libraries. Although these approaches are highly effective in practice, it remains an open question whether an agent can autonomously discover useful theorems without such human priors. We study this question in a formal axiomatic system by developing an agent that starts from axioms and inference rules alone and gradually grows a library of useful theorems. Concretely, we propose a self-supervised theorem-discovery algorithm that alternates between proof search and useful-theorem extraction, building a theorem library whose entries are reused as lemmas for subsequent proof search. Experiments show that the agent discovers tens of thousands of theorems and finds proofs for human-written benchmark problems, suggesting that its discoveries include theorems meaningful from a human mathematical perspective. Furthermore, the discovered theorems improve LLM proof performance when provided as prompt lemmas, indicating that they can serve as external knowledge for LLM reasoning. Our results provide evidence that useful theorems can emerge from proof search without relying on human-provided theorem libraries. More broadly, they suggest a path toward self-evolving AI systems for mathematics whose discoveries remain formally verifiable.

13:00 JSTLLM/生成AI

LLM の機械的性格分析 潜在的特徴介入による性格の制御

大規模言語モデル (LLM) は、生成されたテキストで人間のような OCEAN の性格特性をシミュレートする機能を実証しました。これまでの取り組みは、LLM の性格を形成するための迅速なエンジニアリングまたは微調整に焦点を当てていました。この研究では、モデルの潜在的な特徴に直接介入する機械的解釈可能性アプローチを提案します。私たちの方法は、スパースオートエンコーダー(SAE)と対照的活性化分析を使用して、ターゲットのOCEAN特性に対応する残差ストリーム内の潜在的な方向を特定します。活性化空間で加法的ステアリング ベクトルを形式化し、隠れ状態に小さな加法的シフトを適用することで、全体的な言語モデリングのパフォーマンスを維持しながらターゲットの特性を強化する方法を示します。特徴シフトの最適な組み合わせを決定するために、性格表現とタスクのパフォーマンスのバランスをとるグリッド検索最適化を使用した線形重み付けヒューリスティックを探索します。私たちのアプローチは、標準ベンチマークで高いパフォーマンスを維持しながら、性格特性を機構レベルで制御可能に操作することに期待を示しています。

原文 (English)

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability approach that directly intervenes on the model's latent features. Our method identifies latent directions in the residual stream corresponding to a target OCEAN trait using sparse autoencoders (SAEs) and contrastive activation analysis. We formalize an additive steering vector in activation space and demonstrate how applying a small additive shift to the hidden states enhances the target trait while preserving overall language modeling performance. To determine the optimal combination of feature shifts, we explore a linear weighting heuristic with grid search optimization that balances personality expression with task performance. Our approach shows promise in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.

13:00 JSTエージェント

HyphaeDB: エージェントファーストメモリのための生きた知識トポロジー

既存のベクトル データベースとエージェント メモリ フレームワークはすべて、メモリをエージェントが明示的にクエリするパッシブ ストレージとして扱います。メモリ層自体を介してエージェント間で知識を伝播するシステムはありません。 HyphaeDB は、最新のベクトル データベースすべての中核となるデータ構造である Hierarchical Navigable Small World (HNSW) グラフ トポロジを、検索の最適化としてではなく、マルチエージェント AI システムの通信ファブリックとして再解釈するエージェント ネイティブ メモリ インフラストラクチャです。 HyphaeDB では、エージェントは永続的な位置を持つベクトル空間内のノードであり、知識はエネルギーベースの減衰を伴うグラフの隣接構造を介してゴシップ プロトコルを介して伝播し、創発的な動作の矛盾検出、パターンの結晶化、および合意形成はトポロジー、伝播ダイナミクス、およびローカル相互作用ルールの組み合わせから発生します。我々は、3 つのプリミティブ (ナレッジ ノード、トポロジ エッジ、およびメモリ差分)、創発的コンセンサスによる推進を備えた多層抽象化階層、およびスモールワールド ネットワーク理論、伝染病のブロードキャスト プロトコル、および群れインテリジェンスに基づいたシステムの理論的分析に基づいて構築されたアーキテクチャを紹介します。 pgvector を使用した PostgreSQL でのリファレンス実装を提供し、マルチエージェント ソフトウェア エンジニアリング手法である Swarm-Driven Development での具体的な展開について説明します。私たちの知る限り、HyphaeDB は、マルチエージェント調整のための、ナビゲート可能なスモールワールド トポロジーとゴシップベースの知識伝播を組み合わせた最初のシステムです。

原文 (English)

HyphaeDB: A Living Knowledge Topology for Agent-First Memory

Every existing vector database and agent memory framework treats memory as passive storage that agents query explicitly. No system propagates knowledge between agents through the memory layer itself. We introduce HyphaeDB, an agent-native memory infrastructure that reinterprets the Hierarchical Navigable Small World (HNSW) graph topology the data structure at the core of every modern vector database not as a search optimization, but as a communication fabric for multi-agent AI systems. In HyphaeDB, agents are nodes in the vector space with persistent positions, knowledge propagates via a gossip protocol through the graph's neighbor structure with energy-based attenuation, and emergent behaviors contradiction detection, pattern crystallization, and consensus formation arise from the combination of topology, propagation dynamics, and local interaction rules. We present the architecture built on three primitives (knowledge nodes, topology edges, and memory diffs), a multi-layer abstraction hierarchy with promotion via emergent consensus, and theoretical analysis grounding the system in small-world network theory, epidemic broadcast protocols, and swarm intelligence. We provide a reference implementation on PostgreSQL with pgvector and describe a concrete deployment in Swarm-Driven Development, a multi-agent software engineering methodology. HyphaeDB represents, to our knowledge, the first system to combine navigable small world topology with gossip-based knowledge propagation for multi-agent coordination.

13:00 JSTLLM/生成AI研究/論文

LLM ベースのプロービングを使用したプライマリ ICD カテゴリの予測

目的: ICD コードは償還、研究、国民健康監視の中心ですが、自動コーディング システムでは、臨床ナラティブと構造化された電子医療記録 (EHR) 変数の両方からの診断シグナルを統合するのに苦労することがよくあります。私たちは、凍結された医療大規模言語モデル (LLM) 表現が、マルチモーダルな一次診断カテゴリ予測のための共有埋め込みスペースとして機能できるかどうかを評価しました。材料と方法: 最も頻度の高い 10 の主要 ICD-10 コードから 13,645 件の入院からなる MIMIC-IV コホートを構築し、7 つのカテゴリーに統合しました。構造化された変数は臨床ナラティブにシリアル化され、漏れを除去した退院メモと結合されました。凍結された MedFound-Llama3-8B 微調整バックボーンを使用して、5 つのトランスフォーマー層から隠れ状態を抽出し、構造化のみ、非構造化のみ、および複合入力の線形プローブをトレーニングし、XGBoost および情報が一致した PLM-ICD ベースラインと比較し、コンパクトなボトルネック アダプターを使用して MIMIC-III 適応を評価しました。結果: 組み合わせたプローブは MIMIC-IV で最高のパフォーマンスを示し (87.69% 厳密、91.45% 医療精度)、単一モダリティ プローブとベースラインの両方を上回りました。構造化専用プローブは、医療精度において標準ベースラインを 6.19 ポイント上回りました。診断情報は、より深い層でますます線形に分離可能になり、2M パラメーター アダプターにより、ターゲット ラベルの 5% のみを使用して MIMIC-III へのクロスデータセット転送が復元されました。考察: LLM 埋め込みは、構造化されたナラティブな EHR 情報を統合してマルチモーダルな診断予測を行うことができ、小さな表現レベルのモジュールを通じてモダリティやデータセット全体での臨床表現の効率的な再利用をサポートします。結論: 凍結された医療 LLM 表現のマルチモーダル プローブは、EHR モダリティを研究し、データセット全体で臨床表現を適応させるための実用的なアプローチを提供します。

原文 (English)

Primary ICD Category Prediction using LLM-based Probing

Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables. We evaluated whether frozen medical large language model (LLM) representations can serve as a shared embedding space for multimodal primary diagnosis category prediction. Materials and Methods: We constructed a MIMIC-IV cohort of 13,645 admissions from the 10 most frequent primary ICD-10 codes, consolidated into seven categories. Structured variables were serialized into clinical narratives and combined with leakage-pruned discharge notes. Using a frozen MedFound-Llama3-8B-finetuned backbone, we extracted hidden states from five transformer layers and trained linear probes for structured-only, unstructured-only, and combined inputs, comparing against XGBoost and information-matched PLM-ICD baselines and evaluating MIMIC-III adaptation with a compact bottleneck adapter. Results: The combined probe performed best on MIMIC-IV (87.69% strict; 91.45% medical accuracy), exceeding both single-modality probes and baselines. The structured-only probe outperformed its standard baseline by 6.19 points in medical accuracy. Diagnostic information became increasingly linearly separable in deeper layers, and a 2M-parameter adapter restored cross-dataset transfer to MIMIC-III using only 5% of target labels. Discussion: LLM embeddings can unify structured and narrative EHR information for multimodal diagnosis prediction, supporting efficient reuse of clinical representations across modalities and datasets through a small representation-level module. Conclusion: Multimodal probing of frozen medical LLM representations provides a practical approach for studying EHR modalities and adapting clinical representations across datasets.

13:00 JSTLLM/生成AIエージェント

MedEvoEval: シミュレートされた臨床エピソードを通じて医師エージェントの継続的な進化を評価

医師エージェントは、単一ターンの回答生成を超えて、臨床意思決定システムの進化に向けて移行しています。外来患者は、証拠を入手し、検査および相談リソースを利用し、診断および管理計画をいつ完成させるかを決定します。エピソード全体で、記憶、検索、反映、またはその他の更新メカニズムを通じて動作が変化する可能性があります。現在の評価では、この設定は部分的にしかカバーされていません。固定入力の医療 QA ベンチマークは完全な入力から最終的な回答をスコア化しますが、多くのインタラクティブ ベンチマークは依然として個々の遭遇または固定された実行に焦点を当てており、エピソード レベルの決定がエピソード間のエクスペリエンスとどのように相互作用するかを評価するためのサポートは限定的です。アクションゲート型の模擬外来エピソードに基づいた実行可能な長期的評価フレームワークである MedEvoEval を紹介します。各ソースケースは、役割固有の患者、検査、およびマネージャーのビューに変換されます。証拠は有効な行動によってのみ明らかになります。各エピソードは、観察、アクション、最終出力、マネージャーのスコア、およびオプションのエクスペリエンスの書き戻しをリンクする構造化されたトレースを記録します。私たちは、700 の処理されたエピソード、来歴ノート、スキーマ、エピソード ランナー、スコアリング スクリプト、構成、サンプル ログ、分析コード、および軌跡レベルおよびステップ レベルの派生関数を含む、実行可能な E&D アーティファクトをリリースします。実験では、エピソード トレースによって、最終回答スコアリングによって隠れていたプロセス コストが明らかになり、MDT スタイルのコンサルテーションがどのようにリソースを再割り当てするかを示し、メモリの成熟、保留された転送、更新段階の応答、および後方保持の長期的な分析がサポートされることが示されています。これらの結果を総合すると、MedEvoEval が、医師エージェントが経験を通じて改善し、有用な行動を継承し、以前の能力を長期にわたって維持するかどうかを評価するための具体的な基礎を提供することを示しています。

原文 (English)

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes

Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems. Within an outpatient episode, they acquire evidence, use examination and consultation resources, and decide when to finalize a diagnosis and management plan. Across episodes, their behavior may change through memory, retrieval, reflection, or other update mechanisms. Current evaluations only partially cover this setting. Fixed-input medical QA benchmarks score final answers from complete inputs, whereas many interactive benchmarks still focus on individual encounters or fixed runs, providing limited support for evaluating how episode-level decisions interact with cross-episode experience. We introduce MedEvoEval, an executable longitudinal evaluation framework based on action-gated simulated outpatient episodes. Each source case is converted into role-specific patient, examination, and manager views; evidence is revealed only through valid actions; and each episode records a structured trace that links observations, actions, final outputs, manager scores, and optional experience write-back. We release a runnable E&D artifact with 700 processed episodes, provenance notes, schemas, an episode runner, scoring scripts, configurations, example logs, analysis code, and trajectory- and step-level derivatives. Experiments show that episode traces expose process costs hidden by final-answer scoring, show how MDT-style consultation reallocates resources, and support longitudinal analyses of memory maturation, held-out transfer, update-stage response, and backward retention. Together, these results show that MedEvoEval provides a concrete basis for evaluating whether doctor agents improve through experience, transfer useful behavior, and retain earlier capabilities over time.

13:00 JSTビジネス/資金調達ClaudeGPT / ChatGPTGemini

実際のポイントオブケアの臨床クエリに対する臨床 AI ツールの専門家による評価

現在、医師は毎週何百万もの臨床質問を AI ツールに投げかけていますが、これらのツールは主に、実際に尋ねられる質問ではなく、仮説や試験形式の質問に基づいて評価されています。我々は、30の専門分野にわたる医師によってOpenEvidence(OE)プラットフォームに送信された620のリアルワールドポイントオブケアクエリ(Real-POCQi)と、HealthBenchからの187の質問に基づいて構築された盲検評価を報告します。 36 州の 149 名の現役医師が、各質問の専門分野に合わせた採点者を使用して、3 つのフロンティア汎用モデル (Claude Opus 4.8、Gemini 3.1 Pro、GPT-5.5) と特殊な臨床ツール (OE) の回答を直接比較しました。臨床意思決定支援に関連する 5 つの側面 (精度、臨床的有用性、情報源の品質、検証可能性、完全性) に沿って回答を比較した場合、医師は専門ツールをすべての軸で最も高いスコアにしました。 Real-POCQi に関する一次分析では、勝率の差 (勝率と敗率の差) は 25 ~ 39 パーセント ポイントの範囲でした (p<0.001)。結果は、引用表示、回答の長さ、OE ユーザーのステータス、Real-POCQi と HealthBench によって層別化した感度分析で一貫性を保っていました。並行して、LLM 裁判官は体系的に専門家裁判官とは異なることが判明しましたが、最良のモデルについては両者とも概ね一致しました。これらの調査結果は、次の 2 つの結論を強調しています。(i) AI ツールの評価は、現実世界のクエリ分布を反映し、現代医学を定義する専門分野を反映する専門家判断者を使用する必要があります。(ii) 汎用モデルに対する専用ツールの一貫した利点は、後者が同様の目的を達成できないことを必ずしも意味するわけではありませんが、ターゲットを絞ったエンジニアリングとカスタマイズにより、ユーザーにとってパフォーマンスに有意義な向上がもたらされる可能性があります。私たちは Real-POCQi を公開ベンチマークとしてリリースするとともに、この研究の結果を再現するための事前に指定された統計分析もリリースします。

原文 (English)

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in practice. We report a blinded evaluation built on 620 Real-world Point-Of-Care Queries (Real-POCQi) submitted to the OpenEvidence (OE) platform by physicians spanning 30 specialties, as well as 187 questions from HealthBench. 149 practicing physicians across 36 states made head-to-head comparisons between answers from three frontier general-purpose models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5) and a specialized clinical tool (OE), with graders matched to each question's specialty. When comparing answers along five dimensions relevant to clinical decision support -- accuracy, clinical utility, source quality, verifiability, & completeness -- physicians scored the specialized tool highest on all axes; in the primary analysis on Real-POCQi, win differences (margins between win and loss rates) ranged from 25 to 39 percentage points (p<0.001). Results remained consistent in sensitivity analyses stratifying by citation display, answer length, OE-user status, and Real-POCQi versus HealthBench. In parallel, LLM judges were found to systematically differ from expert judges, though both generally agreed on the best model. These findings underscore two conclusions: (i) AI tool evaluations should reflect real-world query distributions and use expert judges that mirror the specialization defining modern medicine and (ii) the consistent advantage of the specialized tool over general-purpose models does not necessarily mean that the latter cannot serve similar purposes, but that targeted engineering and customization can yield meaningful gains in performance for its users. We release Real-POCQi as a public benchmark, as well as the prespecified statistical analysis for reproducing results of this study.

13:00 JSTLLM/生成AIエージェントLlama

交通工学実践のためのカスタマイズされた生成 AI エージェント: 開発および継続的な事前トレーニング ガイドライン

生成人工知能 (AI) と大規模言語モデル (LLM) の最近の進歩により、複雑な推論、要約、質問応答タスクの自動化において大きな期待がもたれています。ただし、技術標準、エンジニアリング用語、およびドメイン固有のセマンティクスへの理解が不十分であるため、特殊なエンジニアリング ドメインにおける汎用 LLM の有効性は依然として限定的です。この研究では、輸送工学アプリケーション向けにカスタマイズされた生成 AI エージェントを開発するための体系的なアプローチを提案します。米国の輸送マニュアル、設計ガイドライン、規制文書の厳選されたコーパスを使用して、統合低ランク適応 (LoRA) フレームワークを通じて 6 つの最先端の LLM の継続的な事前トレーニングが実施されます。トレーニング プロセスは、収束とモデルの安定性を確保するために監視されます。パフォーマンスは、BLEU-4 や ROUGE などの標準的な自然言語処理メトリクスを使用して評価され、Qwen2.5-7B および LLaMA-3.1-8B が最高のドメイン アラインメントと応答品質を示しています。結果は、技術的な内容の解釈とコンテキスト固有の推論における LLM パフォーマンスの向上における LoRA ベースの適応の有効性を検証しました。この成果は、ドメインに特化した生成 AI エージェントを構築するための再現可能な開発フレームワークに貢献し、交通研究、設計、計画、政策分析における広範な展開をサポートします。

原文 (English)

Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline

Recent advancements in generative artificial intelligence (AI) and large language models (LLMs) have shown significant promise in automating complex reasoning, summarization, and question-answering tasks. However, the effectiveness of general-purpose LLMs in specialized engineering domains remains limited due to insufficient exposure to technical standards, engineering terminology, and domain-specific semantics. This study proposes a systematic approach to developing a customized generative AI agent for transportation engineering applications. A curated corpus of U.S. transportation manuals, design guidelines, and regulatory documents is used to conduct continued pretraining of six state-of-the-art LLMs through a unified low-rank adaptation (LoRA) framework. The training process is monitored to ensure convergence and model stability. Performance is evaluated using standard natural language processing metrics, including BLEU-4 and ROUGE, with Qwen2.5-7B and LLaMA-3.1-8B demonstrating the highest domain alignment and response quality. Results validate the effectiveness of LoRA-based adaptation in improving LLM performance on technical content interpretation and context-specific reasoning. This work contributes a reproducible development framework for constructing domain-specialized generative AI agents, supporting broader deployment in transportation research, design, planning, and policy analysis.

13:00 JSTエージェント

実行時監視によるマルチエージェント AI でのエラー伝播の防止

マルチエージェント AI システムは、さまざまな言語モデルが推論トレースを交換し、最初の予測を修正し、最終的な決定をサポートできるようにすることで、回答の選択を改善できます。ただし、そのようなコミュニケーションは信頼性のリスクも引き起こす可能性があります。あるエージェントの推論は、別のエージェントの間違いを修正することはできますが、最初は正しかったエージェントを誤解させる可能性もあります。この論文は、推論交換と実行時の回答修正を通じて信頼性の高いマルチエージェント AI コミュニケーションを研究します。私たちは、エージェントが最初に多肢選択式の質問に独立して回答し、次に推論の痕跡を共有して決定を修正するフレームワークを開発します。私たちは数値実験を実施し、このプロセスが精度を向上させ、否定的な回答よりも肯定的な回答の変化を生み出し、サイバーセキュリティ、ネットワーキング、一般知識などの領域にわたって効果が持続するかどうかを評価します。この結果は、マルチエージェント推論によって信頼性が向上する時期と、エラーが伝播する可能性がある時期を特定するのに役立ちます。

原文 (English)

Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring

Multi-agent AI systems can improve answer selection by allowing different language models to exchange reasoning traces, revise initial predictions, and support a final decision. However, such communication may also introduce reliability risks: reasoning from one agent can correct another agent's mistake, but it can also mislead an agent that was initially correct. This paper studies reliable multi-agent AI communication through reasoning exchange and runtime answer revision. We develop a framework in which agents first answer multiple-choice questions independently, then share reasoning traces and revise their decisions. We conduct numerical experiments where we evaluate whether this process improves accuracy, produces more positive than negative answer transitions, and remains effective across domains such as cybersecurity, networking, and general knowledge. The results help identify when multi-agent reasoning improves reliability and when it may propagate errors.

13:00 JSTLLM/生成AIエージェント

LLM エージェントの攻撃面としてのメモリ: 多肢選択式質問応答に関する研究

AI エージェントは、言語理解とタスクの実行、外部ツールの使用、メモリ メカニズムを統合することで、従来の大規模言語モデル (LLM) アプリケーションを拡張します。メモリにより、エージェントは以前の対話を保持し、よりパーソナライズされたコンテキストを認識した応答を提供できますが、新たな脆弱性も生じます。メモリに保存された情報は、現在のクエリがクリーンな場合でも、将来の出力に影響を与える可能性があります。この論文では、多肢選択式質問応答のための LLM ベースのエージェントにおけるメモリ操作を調査します。まず、タスク関連情報を保存および取得する外部メモリ コンポーネントを備えた LLM ベースの AI エージェントを設計および実装します。次に、多肢選択式の質問に答える前に、誤解を招くメモリや破損したメモリをエージェントに挿入する、基本的なメモリ操作シナリオを紹介します。制御された実験設定を使用して、記憶操作の前後でエージェントのパフォーマンスを比較し、回答精度、攻撃成功率、操作されたオプションの選択の変化を測定します。私たちの結果は、単純な記憶操作でさえエージェントの最終的な回答に顕著な影響を及ぼし、クリーンで整った質問を受け取ったにもかかわらず誤った選択肢を選択してしまう可能性があることを示しています。

原文 (English)

Memory as an Attack Surface in LLM Agents: A Study on Multiple-Choice Question Answering

AI agents extend conventional large language model (LLM) applications by integrating language understanding with task execution, external tool use, and memory mechanisms. While memory allows agents to retain prior interactions and provide more personalized and context-aware responses, it also introduces a new vulnerability: information stored in memory can influence future outputs even when the current query is clean. In this paper, we investigate memory manipulation in LLM-based agents for multiple-choice question answering. We first design and implement an LLM-based AI agent with an external memory component that stores and retrieves task-relevant information. We then introduce basic memory manipulation scenarios in which misleading or corrupted memories are inserted into the agent before it answers multiple-choice questions. Using a controlled experimental setup, we compare the agent's performance before and after memory manipulation and measure changes in answer accuracy, attack success rate, and selection of manipulated options. Our results show that even simple memory manipulations can noticeably affect the agent's final answers, causing it to select incorrect options despite receiving clean and well-formed questions.

13:00 JSTLLM/生成AI画像/動画生成

低コストの概念ベースのローカライズされた説明: トレーニング不要のアプローチでどこまで達成できるでしょうか?

Concept-based Explainable AI (C-XAI) は、意味論的な概念に基づいて人間が理解できる説明を追求しますが、詳細な概念の注釈が不足しているため検証が制限されます。中規模のマルチモーダル大規模言語モデル (MLLM) が、オブジェクト レベルとパーツ レベルの両方で境界ボックス領域にラベルを割り当てることによって、厳密なゼロショット条件下でローカライズされた概念命名を実行できるかどうかを評価します。我々は、(i) 中程度の語彙に対する閉集合、カテゴリ制約のプロンプト、および (ii) 大規模なラベル空間に対する埋め込み類似度ベースの戦略である Open-CoNa を備えた、コンセプト ネーミング (CoNa) の再現可能なゼロショット評価プロトコルを提案します。 4 つの MLLM (7B ~ 32B) を使用した実験では、データセット全体で一貫したパフォーマンス傾向が示され、オブジェクト レベルの完全一致精度が 62% ~ 88% に達し、局所領域からのトレーニング不要の概念アノテーションの可能性が強調されています。私たちは限界と故障モードについて議論し、将来の低コスト C-XAI 研究をサポートするための再現可能なフレームワークをリリースします。

原文 (English)

Low-cost concept-based localized explanations: How far can we get with training-free approaches?

Concept-based Explainable AI (C-XAI) seeks human-understandable explanations grounded in semantic concepts, yet validation is limited by the scarcity of fine-grained concept annotations. We evaluate whether mid-scale Multimodal Large Language Models (MLLMs) can perform localized concept naming under strict zero-shot conditions by assigning labels to bounding-box regions at both object and part levels. We propose a reproducible zero-shot evaluation protocol for Concept Naming (CoNa) with (i) closed-set, category-constrained prompting for moderate vocabularies and (ii) Open-CoNa, an embedding-similarity-based strategy for large label spaces. Experiments with four MLLMs (7B-32B) show consistent performance trends across datasets, reaching 62%-88% object-level exact-match accuracy, highlighting the potential of training-free concept annotation from localized regions. We discuss limitations and failure modes and release a reproducible framework to support future low-cost C-XAI research.

13:00 JSTエージェントビジネス/資金調達

人間のフォールバックの管理: AI と労働者の流動性の向上におけるスキル投資

企業が自律型 AI を導入する場合、どの程度の作業をシステムに任せるか、どの程度の作業を従業員の関与を維持するかを決定する必要があります。この決定は、現在の生産高と将来の人的資本に影響を与えます。私たちは、AI が機能するときにはワーカーよりも優れたパフォーマンスを発揮する可能性があるが、確実な確率で失敗する可能性があるという、節約的な 2 期間モデルを開発します。企業は労働者の関与を選択します。エンゲージメントは、基準を下回る従業員の現在の生産量を低下させますが、学習と消耗によって将来のスキルを変化させます。私たちは AI の進歩の 2 つの側面を区別します。それは、機能 (システムが動作するときのシステムの出力) と信頼性 (システムが動作する確率) です。単一企業のベンチマークでは、エンゲージメントはフォールバック投資としてのみ価値があります。同社は最もスキルの低い従業員を最も雇用しています。これは、従業員のスキルギャップが最も大きく、有用な代替レベルに引き上げるためのコストが最も低いためです。労働者の流動性により、エンゲージメントは労働市場の選別にも影響を及ぼします。労働者は、より価値のあるスキルの軌跡を構築する仕事を好みます。この選別の動機は、スキルの向上がより価値があり、エンゲージメントのコストが低い、AI フロンティアに近い高スキルの労働者をターゲットにしています。したがって、モビリティはエンゲージメント パターンを逆転させ、AI ベンチマークを下回る最もスキルの低い労働者から最もスキルの高い労働者に投資をシフトする可能性があります。モビリティはまた、AI の進歩がエンゲージメントに与える影響を再形成します。能力が向上すると、企業が提供するスキル軌跡の価値が高まるためエンゲージメントが高まりますが、信頼性が高まるとフォールバックの必要性が減り、同時に学習機会も変化するため、エンゲージメントが上下する可能性があります。労働者の流動性の下では、人間と AI の仕事の設計は人的資本の投資の問題となり、今日の仕事の割り当てが将来のスキルを形成します。

原文 (English)

Managing the Human Fallback: Skill Investment Under Improving AI and Worker Mobility

When firms deploy autonomous AI, they must decide how much work to leave to the system and how much to keep workers engaged. This decision affects current output and future human capital. We develop a parsimonious two-period model in which AI may outperform the worker when it functions, but may fail with positive probability. A firm chooses worker engagement; engagement lowers current output for below-benchmark workers, but changes future skill through learning and erosion. We distinguish two dimensions of AI progress: capability, the system's output when it works, and reliability, the probability that it works. In a single-firm benchmark, engagement is valuable only as fallback investment. The firm engages the least-skilled workers most, because they have the largest skill gaps and are least costly to bring toward a useful fallback level. With worker mobility, engagement also affects labor-market sorting: workers prefer jobs that build more valuable skill trajectories. This sorting motive targets higher-skill workers near the AI frontier, where skill gains are more valuable and engagement is less costly. Mobility can therefore reverse the engagement pattern, shifting investment from the least-skilled toward the most-skilled workers below the AI benchmark. Mobility also reshapes how AI progress affects engagement: greater capability raises engagement by increasing the value of the skill trajectory a firm offers, whereas greater reliability can raise or lower it because it reduces fallback need while also changing learning opportunities. Under worker mobility, human-AI work design becomes a problem of human-capital investment, in which allocating work today shapes future skill.

13:00 JSTLLM/生成AIエージェント

大規模言語モデルのエージェントワークフローの特徴付け: N8n エコシステムに関する研究

大規模言語モデル (LLM) は、ローコードおよびノー​​コードの自動化プラットフォームで急速に採用されており、専門家以外のユーザーが自然言語の理解を外部サービスや API と組み合わせたワークフローを設計します。 LLM エージェントは、複雑な複数ステップのタスクを推論、計画し、自律的に実行するためのコア「頭脳」として LLM を使用する LLM システムです。このペーパーでは、ローコード自動化プラットフォームにおける LLM エージェント ワークフローに関する最初の大規模な実証研究を紹介します。私たちは、公開されている 6,000 を超える n8n ワークフローを分析し、その設計の 4 つの側面 (タスク分散、構造およびツールの使用パターン、信頼性メカニズム、自律性レベル) を調査します。私たちの分析では、LLM ワークフローが単なる即時応答パイプラインではないことが示されています。代わりに、LLM は通常、制御ロジック、外部ツール、通信サービス、ストレージ システム、人間によるレビュー ポイントを含む、より広範な自動化構造内に組み込まれます。さらに、多くのワークフローには、LLM 実行後の軽量の後処理またはルーティング ロジックが含まれている一方で、構造化されたフォールバック パス、修復ループ、障害固有のアラート、人間による承認ゲートなどの明示的な信頼性メカニズムは、依然として比較的一般的ではないことがわかりました。これらの結果は、実際の自動化エコシステムにおける LLM エージェントの導入の増加と、信頼性、安全性、ガバナンスに対する限られたエンジニアリング サポートとの間にギャップがあることを明らかにしています。全体として、私たちの研究は、現実世界の LLM エージェント ワークフローを理解して改善しようとしている研究者、プラットフォーム開発者、実践者に 10 の実証結果と 5 つの研究の成果を提供します。

原文 (English)

Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem

Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs. LLM agents are LLM systems that use LLMs as a core "brain" to reason, plan, and autonomously execute complex, multi-step tasks. In this paper, we present the first large-scale empirical study of LLM agentic workflows in low-code automation platforms. We analyze more than 6,000 publicly available n8n workflows and examine four aspects of their design: task distribution, structural and tool use patterns, reliability mechanisms, and autonomy levels. Our analysis shows that LLM workflows are not merely prompt response pipelines. Instead, LLMs are commonly embedded within broader automation structures involving control logic, external tools, communication services, storage systems, and human review points. We further find that while many workflows include lightweight post-processing or routing logic after LLM execution, explicit reliability mechanisms such as structured fallback paths, repair loops, failure-specific alerts, and human approval gates remain relatively uncommon. These results reveal a gap between the increasing deployment of LLM agents in practical automation ecosystems and the limited engineering support for reliability, safety, and governance. Overall, our study provides ten empirical findings and five research takeaways for researchers, platform developers, and practitioners seeking to understand and improve real-world LLM agentic workflows.

13:00 JSTエージェント

HiComm: マルチエージェント強化学習のための階層型通信

協調的なマルチエージェント強化学習 (MARL) は、部分的な可観測性を軽減するために通信に依存することがよくありますが、既存のプロトコルのほとんどは、メッセージを要約した観測の構造から切り離されたフラットな密ベクトルとして扱います。この設計では、観察がグループやエンティティなどの階層に自然に従う多くの協力的な環境において、帰納的バイアスの重要な原因を見落としています。私たちは、送信者の階層的な観察にメッセージを基礎付けるプラグイン通信モジュールである \textsc{HiComm} を提案します。 \textsc{HiComm} は受信者主導型です。受信者はクエリを発行し、最初にグループ、送信者、そのグループ内のエンティティを選択する 3 段階のデコード プロセスを通じて階層を解決し、対応する機能スライスをメッセージとして返します。これにより、通信が非構造化ベクトル送信から、送信者の監視階層を介した構造化情報検索に変換されます。このメカニズムは、微分可能な離散選択と標準 MARL パイプラインに接続する軽量の共有投影設計のために、Straight-Through Gumbel-Softmax を使用してインスタンス化されます。さまざまな観察構造と調整要求を備えた協調的な MARL タスクにわたる実験では、\textsc{HiComm} が代表的な学習済み通信ベースラインと一致またはそれを上回るパフォーマンスを示し、一方、エピソードごとに受信者あたり通信量を最大 $23\times$ 削減することが示されました。

原文 (English)

HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning

Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to $23\times$ per receiver per episode.

13:00 JST研究/論文

フロー推論モデル: 反復的な自己洗練による推論の拡張

離散フロー モデルは最近、数ステップのテキスト生成で有望なパフォーマンスを示しています。ただし、数独やゼブラ パズルなどの構造化された推論タスクに素朴に適用すると、自信を持って不正解に収束します (数独パズルの $\sim$36% のみを解決します)。フロー モデルを使用した構造化推論のためのトレーニングおよびテスト時間スケーリング フレームワークであるフロー推論モデル (FRM) を紹介します。私たちは、解決率が低いにもかかわらず、フロー モデルが独自の検証器として機能する可能性があると観察しています。正解は、ノイズ除去ダイナミクスの安定した固定点であり、再ノイズ化および再解決されるとそれ自体に戻ります。これにより、テスト時間のスケーリング パラダイムが可能になります。つまり、多くの候補解を提案し、動的に安定したものを維持します。これだけでも、Sudoku-Shah (~$100\%$) と Zebra ($95.9\%$) で高い解決率に達します。これは、そのディストリビューションでトレーニングすることなく、Sudoku-Extreme ($96.1\%$) のようなディストリビューション外のより難しいパズルにも一般化されます。ただし、この純粋な検索では大量の計算が無駄になり、不正確な候補解が生成されます。したがって、ベースモデルの効率を向上させるトレーニングレシピを設計します。まず、フロー モデルを自己調整チャネルでトレーニングし、推論時に閉じて、独自の過去の予測を改良できるようにします。次に、直接優先最適化を使用して、失敗した世代を回避するようにモデルをトレーニングします。これらの変更によりベース モデルの効率が大幅に向上し、わずか $7$ の順方向パスで数独で $99.2\%$ に達します。これは、同じ精度のニーズを比較した最も強力なマスク拡散ベースラインよりも $8\times$ 少ない額です。これをテスト時間のスケーリングと組み合わせると、フロー モデルは配布外の難しいパズル (例: Sudoku-Extreme) をはるかに効率的に解くことができます。

原文 (English)

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement

Discrete flow models have recently shown promising performance on few-step text generation; however, when naively applied to structured reasoning tasks such as Sudoku and Zebra puzzles, they converge confidently to incorrect answers (solving only $\sim$36% of Sudoku puzzles). We introduce Flow Reasoning Models (FRMs), a training and test-time-scaling framework for structured reasoning with flow models. We make the observation that, despite their poor solve rate, flow models can act as their own verifiers. A correct answer is a stable fixed point of the denoising dynamics, returning to itself when re-noised and re-solved. This enables a test-time-scaling paradigm: propose many candidate solutions and keep those that are dynamically stable, which alone reaches high solve rates on Sudoku-Shah (~$100\%$) and Zebra ($95.9\%$). This even generalizes to harder out-of-distribution puzzles like Sudoku-Extreme ($96.1\%$), without ever training on that distribution. This pure search, however, wastes a great deal of computation generating incorrect candidate solutions. We therefore design a training recipe to improve the base model's efficiency. First, we train flow models with a self-conditioning channel and close it at inference, letting them refine their own past predictions. Second, we train models to avoid their own failed generations using direct preference optimization. These changes substantially improve the base model's efficiency, letting it reach $99.2\%$ on Sudoku in just $7$ forward passes, over $8\times$ fewer than the strongest matched masked-diffusion baseline we compare needs for the same accuracy. When combined with test-time scaling, this lets flow models solve hard out-of-distribution puzzles (e.g. Sudoku-Extreme) far more efficiently.

13:00 JST研究/論文

プールされたリーダーボードはシステム固有の勝者を隠す: オフライン根本原因分析ベンチマークのレポート プロトコル監査

オフラインの根本原因分析 (RCA) ベンチマークでは、通常、複数のサブシステムにわたる単一のプールされた上位 1 位の精度によってメソッドがランク付けされ、エンジニアは多くの場合、プールされた勝者を独自のサブシステムの推奨事項として読み取ります。私たちは、11 のサブシステムと 778 の一致するスコアリング ユニットをカバーする 3 つの公開 RCA ベンチマーク ファミリ、OpenRCA、RCAEval、PetShop でその測定値を監査しました。同一のケースでペアごとの比較を維持するために、メイン分析では、BARO、CD-1min アダプター、max-$|Z|$、およびサービスごとのアラート数の 4 つのメソッドまたはコンパレーターが完全にカバーされています。 6 つのペアごとの比較はすべて、両方の符号のサブシステム レベルの効果を示し、すべての変量効果 95\% 予測区間がゼロと交差し、ケースレベルの相互作用テストでは 6 つのペアのうち 5 つで交換可能性が拒否されます。 Leave-one-system-out 選択では、保留された 11 個のサブシステムのうち最大 5 個で低いスコアの方法が選択され、RCAEval / Sock-Shop では 24.8 pp に達しました。 320 行の監査モジュールをリリースします。一致する RCA ベンチマーク スコア テーブルが与えられると、プールされたスコアとともに同じサブシステムごとの安定性チェックが再計算されます。

原文 (English)

Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks

Offline root-cause-analysis (RCA) benchmarks commonly rank methods by a single pooled top-1 accuracy across multiple subsystems, and engineers often read the pooled winner as a recommendation for their own subsystem. We audit that reading on three public RCA benchmark families -- OpenRCA, RCAEval, and PetShop -- covering 11 subsystems and 778 matched scoring units. To keep pairwise comparisons on identical cases, the main analysis retains four methods or comparators with complete coverage: BARO, a CD-1min adapter, max-$|Z|$, and per-service alert-count. All six pairwise comparisons show subsystem-level effects of both signs, every random-effects 95\% prediction interval crosses zero, and case-level interaction tests reject exchangeability in 5 of 6 pairs. Leave-one-system-out selection picks the lower-scoring method on up to 5 of 11 held-out subsystems, with regret reaching 24.8 pp on RCAEval / Sock-Shop. We release a 320-line audit module; given a matched RCA benchmark score table, it recomputes the same per-subsystem stability checks alongside pooled scores.

13:00 JST研究/論文

国際人道法における直接的因果関係とAIを介した民間サイバー作戦の課題

国際人道法は、戦闘行為に直接参加しない限り、またその期間中、民間人を直接攻撃から保護しており、ICRC の 2009 年解釈ガイダンスでは、3 つの基準の累積テストを通じてこの規則を運用しています。この論文は、AIを介した民間のサイバー作戦は、構造的に特殊な方法でこのテストの直接的な因果関係要素に異議を唱えていると主張する。攻撃的なAI研究で最近実証された種類の自律型マルチエージェント・サイバー・システムを民間人が導入する場合、人間が関与を解除した後にシステムが生成する決定によって危害が生じるため、「1つの因果関係のステップ」基準は失敗し、統合部分の要件は、行為を独立して分類できる下流の人間の貢献者を前提としているため拡張されない。したがって、この枠組みは、敵対行為に個人的に参加する民間人を捕獲するという目的との緊張関係から、そのような展開を間接的な参加として扱うことをデフォルトとしている。理論的な分析を超えて、この論文は、積分部分テストの具体性コンポーネントが暗黙的に変化する特性として目標仕様の粒度を特定し、AI を介した操作を 5 つのレベルのスペクトルに分類し、既存の技術的な AI ガバナンス手段はこの特性を記録または報告していないと主張します。

原文 (English)

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations

International humanitarian law protects civilians from direct attack unless and for such time as they take direct part in hostilities, with the ICRC's 2009 Interpretive Guidance operationalising this rule through a three-criterion cumulative test. This paper argues that AI-mediated civilian cyber operations challenge the direct causation element of this test in a structurally specific way: when a civilian deploys an autonomous multi-agent cyber system of the kind recently demonstrated in offensive AI research, the "one causal step" standard fails because harm is produced by system-generated decisions made after human disengagement, and the integral-part requirement does not extend because it presupposes downstream human contributors whose conduct can be independently classified. The framework therefore defaults to treating such deployments as indirect participation, in tension with its purpose of capturing civilians who personally take part in hostilities. Beyond the doctrinal analysis, this paper identifies goal-specification granularity as the property on which the integral-part test's concreteness component implicitly turns, classifies AI-mediated operations along a five-level spectrum, and argues that existing technical AI governance instruments do not log or report this property.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

Long-Horizo​​n LLM エージェントの選択的メモリ保持

メモリ拡張 LLM エージェントにとって保持が重要になるのはどのような場合ですか?これを、フリーズした LLM エージェントの境界付き外部メモリ用の軽量フレームワークである TraceRetain を使用して研究します。これは、解釈可能な特徴 (成功、経過時間、アクセス頻度、冗長性、特異性、類似性、ダウンストリーム ユーティリティ) によってエントリをスコアリングし、最もスコアの低いエントリを容量に応じて削除します。 gpt-5-mini を使用したクリーン ALFWorld では、外部メモリは 2 つのシードにわたってメモリなしの場合よりも確実に改善されますが、制限付き保持ポリシー間の差異は Wilson 95% CI の範囲内に収まります。T=100 ~ T=200 のクリーン ALFWorld では、対処するように設計されたメモリ汚染保持保持が自然には示されません。制御されたノイズの多い書き込みストレス (合成ディストラクタ 75%) の下では、無制限メモリと FIFO-K50 は Precision@5 で低下します (20.2% から 12.4% および 15.8% から 3.8%)。一方、TraceRetain-CEM は本質的に変化せず (16.9% から 16.6%)、タスクの 97/100 の成功を維持します。メカニズム: 無制限のメモリは平均類似度 (0.87) が最も高くなりますが、精度は最も低く、埋め込み空間内のクエリに近い失敗したディストラクタを示します。実施されたディストリビューション評価では、メモリ拡張ポリシーが 50 タスク中 47 ~ 49 を解決するのに対し、メモリなしの場合は 39/50 を解決することが示されています。制限付き保持は、タスク成功のコストをかけずに、飽和したクリーン ベンチマークでメモリとステップ効率を購入し、ストリームにノイズが含まれる場合にのみキャッシュ ヒューリスティックと区別します。

原文 (English)

Selective Memory Retention for Long-Horizon LLM Agents

When does retention matter for memory-augmented LLM agents? We study this with TraceRetain, a lightweight framework for bounded external memory in frozen LLM agents that scores entries by interpretable features (success, age, access frequency, redundancy, specificity, similarity, downstream utility) and evicts the lowest-scoring ones at capacity. On clean ALFWorld with gpt-5-mini, external memory robustly improves over no memory across two seeds, but differences among bounded retention policies fall within Wilson 95% CIs: clean ALFWorld at T=100 to T=200 does not naturally exhibit the memory pollution retention is designed to address. Under a controlled noisy-write stress (75% synthetic distractors), unbounded memory and FIFO-K50 degrade on Precision@5 (20.2% to 12.4% and 15.8% to 3.8%) while TraceRetain-CEM is essentially unchanged (16.9% to 16.6%) and preserves 97/100 task success. The mechanism: unbounded memory has the highest mean similarity (0.87) but lowest precision, indicating failed distractors close to the query in embedding space. Held-out in-distribution evaluation shows memory-augmented policies solving 47 to 49 of 50 tasks vs. 39/50 for no memory. Bounded retention buys memory and step efficiency on saturated clean benchmarks at no task-success cost, and only differentiates from cache heuristics when streams contain noise.

13:00 JSTビジネス/資金調達

ナレッジ グラフにおけるグラフ間の意味的類似性の測定: ナレッジ グラフ埋め込みの経験的評価

ナレッジ グラフ (KG) は事実を構造化されたトリプルとして表し、さまざまなドメインにわたる関係知識を整理するために広く使用されています。テキスト情報の範囲が単語や文章から完全な文書に及ぶのと同様に、KG 情報は、エンティティ、関係、トリプルからサブグラフや KG 全体に至るまで、複数のレベルで解釈できます。ただし、既存の KG 埋め込み手法は主にエンティティ、リレーション、トリプルに焦点を当てており、グラフレベルのセマンティクスにはほとんど対処されていません。通常、構造パターンに基づいてグラフを比較する従来のグラフレベルの方法も、構造の類似性だけでは KG 間の意味的な類似性を保証できないため、不十分です。さまざまな方法がそのようなグラフレベルの意味論的情報をどの程度うまく捕捉しているかを評価するために、KG のペアが意味論的に対応する基礎的な情報を表すかどうかを決定する、グラフ間の意味論的類似性を研究します。信頼できるグラウンドトゥルース対応を取得するために、テキスト文書を変更し、元の文書と変更された文書の両方から KG を抽出し、それらの既知の対応関係を KG ペアに転送することにより、意味論的一致データセットを構築します。各データセットについて、テキストベース、構造ベース、および KG 埋め込みベースのアプローチを比較します。 KG 埋め込みベースのアプローチでは、ペアごとのエンティティの最大類似性を使用する \textit{EmbPairSim} と、周波数加重セントロイドを使用する \textit{AvgEmbSim} の 2 つのスコアリング関数を導入します。 WikiText-2 と CC-News での実験では、\textit{EmbPairSim} が大幅に少ないパラメーターを使用しながら、Sentence-BERT よりも最大 5.3 pp 高い MRR を達成することが示されています。これらの結果は、KGE 表現が、KG におけるグラフ間の意味論的類似性に対するコンパクトで効果的なシグナルとして機能できることを示唆しています。私たちのコードは https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity で入手できます。

原文 (English)

Measuring Graph-to-Graph Semantic Similarity in Knowledge Graphs: An Empirical Evaluation of Knowledge Graph Embeddings

A Knowledge Graph (KG) represents facts as structured triples and is widely used to organize relational knowledge across diverse domains. Just as textual information ranges from words and sentences to complete documents, KG information can be interpreted at multiple levels, from entities, relations, and triples to subgraphs and entire KGs. However, existing KG embedding methods mainly focus on entities, relations, and triples, leaving graph-level semantics largely unaddressed. Conventional graph-level methods, which typically compare graphs based on structural patterns, are also insufficient because structural similarity alone cannot guarantee semantic similarity between KGs. To evaluate how well different methods capture such graph-level semantic information, we study graph-to-graph semantic similarity, which determines whether a pair of KGs represents semantically corresponding underlying information. To obtain reliable ground-truth correspondences, we construct a semantic matching dataset by modifying text documents, extracting KGs from both original and modified documents, and transferring their known correspondences to KG pairs. We compare text-based, structure-based, and KG embedding-based approaches on each dataset. For the KG embedding-based approach, we introduce two scoring functions: \textit{EmbPairSim}, which uses maximal pairwise entity similarity, and \textit{AvgEmbSim}, which uses a frequency-weighted centroid. Experiments on WikiText-2 and CC-News show that \textit{EmbPairSim} achieves up to 5.3 pp higher MRR than Sentence-BERT while using substantially fewer parameters. These results suggest that KGE representations can serve as compact and effective signals for graph-to-graph semantic similarity in KGs. Our code is available at https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity.

13:00 JSTLLM/生成AI

継続的な科学的発見に対する証拠に基づいた LLM の信念

大規模言語モデル (LLM) を使用したオープンエンドの科学的発見は、報酬シグナルによって次にどの仮説をテストするかをガイドする、仮説の探索と検証の長期ループとして機能することが増えています。最近の注目すべき例は AutoDiscovery で、「ベイジアン サプライズ」 (仮説の証拠を観察した後に LLM が経験する信念の変化) を発見指標と検索の報酬の両方として使用します。まず、AutoDiscovery が驚きを静的な量として扱うのに対し、人間の推論における驚きは非定常であることがわかります。これは、継続的な科学的発見の前提条件である、経験とともに進化する信念に関連して定義されます。私たちは、証拠に基づいた LLM 信念、つまり、新しい仮説に対する非定常の驚きを計算するために、以前の仮説からの証拠で事前分布を更新して、この不一致に対処します。コンテキスト内での信念更新メカニズムを比較したところ、事前の発見に対する埋め込みベースの検索拡張生成が最終的な事後を最もよく予測し、静的な驚きの 37.5% が偽りであることが特定されたことがわかりました。次に、これらの偽の報酬を回避し、非定常信念の下で依然として驚くべき仮説を優先するように検索を変更します。具体的には、元の検索手順に 2 つの相補的な変更、つまり信念更新フィルタリングと多様性の最大化を導入します。 5 つの発見ドメインにわたって、私たちの方法は、元の検索手順と比較して、累積非定常驚きを平均 30.62% 増加させます。これは、LLM による継続的な科学的発見には、より優れた信念測定だけでなく、冗長性を回避し、多様性を促進する検索手順も必要であることを示しています。

原文 (English)

Evidence-Informed LLM Beliefs for Continual Scientific Discovery

Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides which hypotheses to test next. A notable recent example is AutoDiscovery, which uses "Bayesian surprise" - the belief shift an LLM undergoes after observing evidence for a hypothesis - as both a discovery metric and a reward for search. We first observe that AutoDiscovery treats surprisal as a static quantity, while surprisal in human reasoning is non-stationary - it is defined relative to beliefs that evolve with experience, a prerequisite for continual scientific discovery. We address this mismatch with evidence-informed LLM beliefs: priors updated with evidence from previous hypotheses to compute non-stationary surprisal for new hypotheses. We compare in-context belief-updating mechanisms and find that embedding-based retrieval-augmented generation over prior discoveries best anticipates eventual posteriors, identifying 37.5% of static surprisals as spurious. We then modify search to avoid these spurious rewards and prioritize hypotheses that remain surprising under non-stationary beliefs. Concretely, we introduce two complementary changes to the original search procedure: belief-update filtering and diversity maximization. Across five discovery domains, our method increases accumulated non-stationary surprisal by 30.62% on average compared to the original search procedure, demonstrating that continual scientific discovery with LLMs requires not only better belief measurement but also search procedures that avoid redundancy and encourage diversity.

13:00 JSTエージェント

AI トレーディングのアルファ特異点: エージェント間の自己進化を通じた新興市場推論

自動アルファマイニングはスコアリング関数を固定し、その関数に基づいて検索アルゴリズムを変更します。固定スコアラーに対して収束する検索は、スコアラーがペナルティを課すことができないものすべてを過剰適合します。これが、サンプル外の汎化ギャップの主な原因です。私たちはスコアリング関数をアルファ要素と並んで検索アーティファクトとして扱い、どのような条件がこの共同検索を許容するのかを研究します。 Sealed Joint Search (SJS) はフレームワークです。評価者を密閉したままにしながら、共同検索が自己確認に崩壊するのを防ぐ、自律発見システムにおける情報フローに関する一連の構造的条件です。条件には、ロールの分解、型指定されたロール間通信、来歴シールされた読み取り、バージョン管理されたストア、およびサブストレートのローカル プロモーションが含まれます。 Agora は SJS を実証的にテストします。5 つの LLM エージェント クラスが 3 つのチャネルを介して通信し、AlphaGen オペレーターに基づいて構築されたアルファ ライブラリを使用して 8 つのスキル ライブラリを進化させます。 3 人の評価者が 1 つの概要にまとめられたレポートを作成し、投票ではなく意見の相違を持ち越します。 CSI 1000 で Agora を 100 ラウンド実行し、すべての LLM 入力を遮断した 91 日間の 2026 ホールドアウトで評価しました。アゴラはホールドアウトシャープ+1.87を達成。最良のベースラインは有利なシードで +1.334、クロスシード平均では -0.755 でした。 Agora の 2 つのメトリクスを凍結ライブラリ アブレーションにプリロードすると、+2.25 のシャープ ギャップのうち +0.40 しか回復せず、ライブラリの進化なしで PPO を追加するとギャップが悪化します。 2 つの指標は設計されたものではなく、出現したものです。注意事項: シングルシード実行、ショートサイド集中信号、ロングショートを対象としています。

原文 (English)

AI Trading's Alpha Singularity: Emergent Market Reasoning through Agent-to-Agent Self-Evolution

Automated alpha mining holds the scoring function fixed and varies the search algorithm over it. A search that converges against a fixed scorer overfits whatever the scorer cannot penalize, a primary cause of the out-of-sample generalization gap. We treat the scoring function as a search artifact alongside the alpha factors and study what conditions make this joint search admissible. Sealed Joint Search (SJS) is a framework: a set of structural conditions on information flow in an autonomous-discovery system that prevent joint search from collapsing into self-confirmation while keeping the evaluator sealed. Conditions cover role decomposition, typed inter-role communication, provenance-sealed reads, versioned stores, and substrate-local promotion. Agora tests SJS empirically: five LLM agent classes communicate via three channels, evolving eight skill libraries, with alpha libraries built on AlphaGen operators. Three evaluators write reports aggregated into one brief, carrying forward disagreement instead of voting. We run Agora for 100 rounds on CSI 1000 and evaluate on a 91-day 2026 holdout sealed from all LLM inputs. Agora achieves holdout Sharpe +1.87; best baseline +1.334 at favorable seed and -0.755 cross-seed mean. Pre-loading Agora's two metrics into a frozen-library ablation recovers only +0.40 of the +2.25 Sharpe gap, and adding PPO without library evolution worsens the gap. The two metrics emerge rather than being designed. Caveats: single-seed run, short-side concentrated signal, intended for long-short.

13:00 JSTエージェント

緊急避難における人間のような意識と行動をモデル化するための認知・感情・性格フレームワーク

エージェントベースの避難シミュレーションは、緊急時の群衆の行動を研究するために広く使用されていますが、多くのモデルは、完全なイベント認識、完全な出口知識、完全に合理的な意思決定などの仮定に依存しています。この論文は、認知的、感情的、社会的、および人格関連のメカニズムを不確実性の下での人間の行動の統一モデルに統合する拡張避難フレームワークを提示します。このフレームワークには、継続的なイベント確実性レベルに基づく動的なイベント認識メカニズム、取得、忘れ、および想起の対象となる出口知識の記憶ベースの表現、パニックが強度の高い状態として現れる継続的恐怖モデル、および OCEAN ベースの性格表現が組み込まれています。神経症傾向は感情モデルに明確に統合されており、恐怖の生成、エスカレーション、社会的伝染、回復に影響を与えます。行動の不均一性は、認識されたリスクへの対応に影響を与える個別の決定閾値を通じてさらに捕捉されます。このフレームワークは、空間親和性、記憶の堅牢性、意思決定の感度、感情のダイナミクス、性格の変化の影響を調べるシミュレーション実験を通じて評価されます。その結果、認知的、感情的、人格主導のプロセスが避難のダイナミクスに大きく影響し、避難効率が低下し、遅れ、混乱、負傷、社会的に影響を受けた行動などの現実的な群集現象が発生することが示されました。提案されたフレームワークは、緊急避難時の人間の行動をより現実的に表現し、認知、感情、性格、群衆の動態間の相互作用の体系的な調査をサポートします。

原文 (English)

A Cognition-Emotion-Personality Framework for Modeling Human-Like Awareness and Behavior in Emergency Evacuations

Agent-based evacuation simulations are widely used to study crowd behavior during emergencies, but many models rely on assumptions such as perfect event awareness, complete exit knowledge, and fully rational decision-making. This paper presents an extended evacuation framework that integrates cognitive, emotional, social, and personality-related mechanisms into a unified model of human behavior under uncertainty. The framework incorporates a dynamic event-awareness mechanism based on a continuous Event Certainty Level, a memory-based representation of exit knowledge subject to acquisition, forgetting, and recall, a continuous fear model in which panic emerges as a high-intensity state, and an OCEAN-based personality representation. Neuroticism is explicitly integrated into the emotional model, influencing fear generation, escalation, social contagion, and recovery. Behavioral heterogeneity is further captured through individualized decision thresholds that affect responses to perceived risk. The framework is evaluated through simulation experiments examining the effects of spatial familiarity, memory robustness, decision sensitivity, emotional dynamics, and personality variation. Results show that cognitive, emotional, and personality-driven processes substantially influence evacuation dynamics, reducing evacuation efficiency and generating realistic crowd phenomena such as delays, confusion, injuries, and socially influenced behaviors. The proposed framework provides a more realistic representation of human behavior in emergency evacuations and supports systematic investigation of the interactions between cognition, emotion, personality, and crowd dynamics.

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPTGemini

PolicyGuard: LLM エージェントにおけるポリシー遵守のための対話ベースのサブエージェント検証ツール

LLM エージェントは、組織に代わってツール呼び出しを通じてユーザー要求を処理し、システム プロンプトに記載されている企業ポリシーに従う必要があります。これまでの研究では、これを保護の問題、つまり非準拠のエージェントのアクションをブロックする外部チェックとしてアプローチしていました。私たちは、ポリシーの遵守はより広範な問題であると主張します。実際のワークフローは多くのターンにわたって展開され、明示的なユーザー確認と前提条件の読み取りが必要で、単一の引数の値ではなく対話の内容に左右されます。この基準を満たすには、(i) 完全な会話のコンテキスト、(ii) ポリシーと現在の対話に関する自己推論、(iii) エージェントの次のターンを導く会話固有の修復が必要です。この 3 つの機能は、これまでの安全対策作業で過小評価されがちでした。対話に対するエージェントの見解、文脈に沿ったポリシーの理由を共有し、エージェントの次のターンに実用的なフィードバックを提供するサブエージェント検証ツール POLICYGUARD を紹介します。 3 つのベンダー (GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Pro) の tau^2-BENCH エアラインで、設定ごとに 4 つのトライアルを行った場合、POLICYGUARD は PASS4 を +12.0 / +6.0 / +12.0 pp 改善しました。通話ごとの分析では、POLICYGUARD が引数レベルのガードの約半分の頻度でブロックしながら、より高いポリシー違反のリコールを達成していることが示されています。

原文 (English)

PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents

LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system prompts. Prior work approaches this as a safeguarding problem -- external checks that block non-compliant agent actions. We argue that policy adherence is a broader problem: real workflows unfold across many turns, require explicit user confirmation and prerequisite reads, and hinge on the content of the dialogue rather than on any single argument value. Meeting this bar requires (i) full conversation context, (ii) self-reasoning over the policy and the current dialogue, and (iii) conversation-specific remediation that guides the agent's next turn -- three capabilities that prior safeguard work has often underestimated. We introduce POLICYGUARD, a sub-agent verifier that shares the agent's view of the dialogue, reasons over the policy in context, and provides actionable feedback for the agent's next turn. On tau^2-BENCH airline across three vendors (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro) with four trials per setting, POLICYGUARD improves PASS4 by +12.0 / +6.0 / +12.0 pp. Per-call analyses show POLICYGUARD achieves higher policy-violation recall while blocking roughly half as often as argument-level guards.

13:00 JSTロボティクス

SurgVLA-Bench: 腹腔鏡手術ロボット工学のための視覚・言語・動作モデルの評価に向けて

Vision-Language-Action(VLA)モデルは、手術ロボット工学における身体化された知能の有望な方向性を示しています。一般的なロボット工学向けの VLA ベンチマークが普及しているにもかかわらず、外科手術向けに特別に設計された標準化された評価プラットフォームは依然として存在しません。この制限に対処するために、腹腔鏡手術ロボット工学における VLA モデルを評価するための最初の包括的なベンチマークである SurgVLA-Bench を紹介します。 SurRoL シミュレーション プラットフォームを活用して、アトミック アクションから完全な外科手術に至るまでの階層的なタスク分類を構築し、アクションの精度と意味の一貫性を評価する多次元評価フレームワークによって補完されます。次に、OpenVLA などの自己回帰モデルと、$\pi_{0}$、$\pi_{0.5}$、SmolVLA などのフロー マッチング モデルを含む 2 つの代表的なパラダイムを体系的に評価します。私たちの実験によると、自己回帰モデルは意味の理解に優れる傾向があるのに対し、フロー マッチング モデルは多くの場合、より高いタスク精度を達成しますが、一般化のトレードオフに直面する可能性があります。しかし、内視鏡の視野の制限、視角の制限、および頻繁な閉塞が基本的な物理的ボトルネックとして依然として存在するため、最高のパフォーマンスを発揮するモデルであっても依然として満足のいくものとは程遠いものです。コードとデータは https://github.com/VCL-HNU/SurgVLA で入手できます。

原文 (English)

SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence of VLA benchmarks for general robotics, standardized evaluation platforms specifically designed for surgical contexts remain absent. To address this limitation, we present SurgVLA-Bench, the first comprehensive benchmark for evaluating VLA models in laparoscopic surgical robotics. Leveraging the SurRoL simulation platform, we construct a hierarchical task taxonomy ranging from atomic actions to complete surgical procedures, complemented by a multi-dimensional evaluation framework assessing action accuracy and semantic consistency. We then systematically evaluate two representative paradigms, including autoregressive models such as OpenVLA, and flow matching models such as $\pi_{0}$, $\pi_{0.5}$, and SmolVLA. Our experiments show that autoregressive models tend to excel in semantic understanding, while flow matching models often achieve higher task precision but may face generalization trade-offs. However, even the best-performing models remain far from satisfactory, as the constrained endoscopic field of view, restricted viewing angles, and frequent occlusions persist as fundamental physical bottlenecks. The code and data are available at https://github.com/VCL-HNU/SurgVLA

13:00 JSTLLM/生成AI

要約が意思決定を歪める場合: LLM 圧縮財務分析における情報の忠実性

財務上の意思決定者は、直接調査できる以上に多くの情報に直面するため、コンテキストの圧縮が必要になります。しかし、大規模言語モデル (LLM) が財務情報源を圧縮すると、元の情報源によって裏付けられた投資判断が変更される可能性があります。私たちはこの問題を情報の忠実度として捉えます。圧縮により、ソースによって引き起こされた決定が変更されると忠実度が失われます。エージェント システムでは、このような損失が中間ステップで繰り返し発生し、意思決定プロセス全体で増幅する可能性があります。財務書類や決算報告のトランスクリプト全体にわたって、LLM ベースの圧縮により、流暢で事実に基づいた圧縮コンテキストが生成され、それにもかかわらず下流の意思決定が変更されることがわかりました。私たちは、忠実度の損失に関連する 2 つの診断パターンを分析します。1 つは、顕著な証拠は保持されますが、正しい解釈に必要な注意事項や文脈修飾子から分離される「脱文脈化」と、異なる圧縮器が同じソースの異なるビューを公開する「モデル依存性」です。次に、エージェント的コンテキスト圧縮を提案します。これは、複数の圧縮候補を生成し、元のソースとの相違点を監査します。私たちの結果は、財務圧縮は効率性や事実性だけでなく、意思決定に関連したコンテキストを保存する能力によっても評価されるべきであることを示唆しています。

原文 (English)

When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis

Financial decision-makers face more information than they can directly inspect, making context compression necessary. Yet when large language models (LLMs) compress financial source material, they can alter the investment judgment supported by the original source. We frame this problem as information fidelity: compression loses fidelity when it changes the decision induced by the source. In agentic systems, such losses may recur across intermediate steps and amplify throughout the decision process. Across financial filings and earnings-call transcripts, we find that LLM-based compression can produce fluent and factually plausible compressed contexts that nevertheless alter downstream decisions. We analyze two diagnostic patterns associated with fidelity loss: decontextualization, where salient evidence is retained but separated from the caveats and contextual qualifiers needed for correct interpretation, and model dependency, where different compressors expose different views of the same source. We then propose Agentic Context Compression, which generates multiple candidate compressions and audits their disagreements against the original source. Our results suggest that financial compression should be evaluated not only by efficiency or factuality, but also by its ability to preserve decision-relevant context.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

複雑さの上限ベンチマーク: 深さスケーリングの下で​​の逐次推論のマルチドメイン評価

必要な連続ステップの数が増加するにつれて、言語モデルの推論がどのように減衰するかを制御された評価である Complexity Ceiling Benchmark (CCB) を導入します。 CCB は、タスクの意味論的な内容を固定し、構造的に異なる 3 つの領域 (接地された空間状態追跡、抽象的な記号ポインター操作、および推移的な関係推論) にわたって、{5,...,50} の深さ N だけを変更します。 5 つのフロンティアおよびオープンウェイト LLM にわたる 6,000 回の試行にわたって、広く分離されたドメイン上限を持つ幾何学的なステップごとの減衰の一貫したパターンが見つかりました。最初の 2 つの領域では、最も強いモデルは N=50 全体で pd > 0.92 を維持しました。 3 番目では、すべてのモデルが N=5 によって崩壊し、pd=0.863 にもかかわらず、最良のモデルの 50% 成功の範囲は H0.5 ~ 4.7 ステップになります。トレース レベル メトリクス (TFBC) は、ベンチマーク全体の正解の 14.5% が、誤った中間推論を介して到達していることを示しています。強制的な冗長状態追跡では上限は変動せず (McNemar p=1.000)、推論が最初に分岐する平均ステップ k* は、パラメーター数よりもドメイン内の精度を予測します。 CCB と幾何学的減衰モデルを併用すると、モデルの長期推論プロファイルがタスク ファミリごとに 1 つの解釈可能な数値に縮小されます。

原文 (English)

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows. CCB fixes the semantic content of a task and varies only its depth N in {5,...,50} across three structurally distinct regimes: grounded spatial state-tracking, abstract symbolic pointer manipulation, and transitive relational inference. Across 6,000 trials over five frontier and open-weight LLMs we find a consistent pattern of geometric per-step decay with widely separated domain ceilings: on the first two regimes the strongest models retain pd>0.92 across N=50; on the third every model collapses by N=5, with the best model's 50%-success horizon at H0.5~4.7 steps despite pd=0.863. A trace-level metric (TFBC) shows that 14.5% of correct answers across the benchmark are reached via incorrect intermediate reasoning. Forced verbose state-tracking does not move the ceiling (McNemar p=1.000), and the mean step at which reasoning first diverges, k*, predicts within-domain accuracy better than parameter count. CCB and the geometric decay model together reduce a model's long-horizon reasoning profile to one interpretable number per task family.

13:00 JSTLLM/生成AI

プロセスアドバンテージシグナルシェーピング: LLM Reasoners のプロセス監視 RL のためのパラダイムに依存しないミドルウェア

グループ相対ポリシー最適化 (GRPO) は、LLM 推論器のプロセス教師あり強化学習のデフォルトのレシピであり、学習されたプロセス報酬モデル (PRM) またはポリシー蒸留 KL 信号を介した高密度プロセス監視は、そうでなければ弱い結果報酬を高密度化する一般的な方法です。しかし、GRPO のグループ標準化の利点の上にこのようなステップレベルの信号を重ねると、3 つの構造的病理が明らかになります。グループ標準化時にプールされたプロセス、結果、およびフォーマット ストリーム間の \emph{チャネル汚染}。プロセス信号の粒度とクレジットされる論理的決定の粒度との間の \emph{解像度の不一致}。そして、信号の符号領域に応じて、GRPO のリターンツーゴー合計が長さのインフレまたは切り捨てられた探索のいずれかを表面化する \emph{累積トラップ}。私たちは \textbf{PASS} (\emph{Process Advantage Signal Shaping}) を提案します。これは、任意のスカラー ステップレベルのプロセス信号と GRPO のクリップされたサロゲートの間に位置し、3 つの病理に順番に対処するコンパクトなミドルウェアです。\emph{Advantage Fusion} は各グループ内で 3 つのストリームを独立して標準化し、\emph{Chunk-by-Value} は信号自体とブロードキャストから値が均一なチャンクを導き出します。各チャンク内のクレジットを計算し、 \emph{Divide-Length} は累積目標を平均値密度スコアに変換します。私たちは、2 つのドメインと 2 つのプロセス信号パラダイム (数学的推論に関する学習済み PRM と、マルチホップ質問応答に関するポリシー蒸留 KL 信号 (一般化されたバリアント付き)) にわたって、2 つのグループ標準化演算子の下で PASS を検証します。どのレジームでも PASS は、対応する GRPO ベースラインを超える一貫した pass@1 ゲインを提供します。

原文 (English)

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

Group Relative Policy Optimization (GRPO) is a default recipe for process-supervised reinforcement learning of LLM reasoners, and dense process supervision -- via learned process reward models (PRMs) or on-policy-distillation KL signals -- is a common way to densify its otherwise weak outcome reward. Layering such a step-level signal on top of GRPO's group-standardized advantage, however, exposes three structural pathologies: \emph{channel contamination} between the pooled process, outcome, and format streams at group standardization; \emph{resolution mismatch} between the granularity of the process signal and the granularity of the logical decisions being credited; and a \emph{cumulative trap} by which GRPO's return-to-go sum surfaces either length inflation or truncated exploration depending on the sign regime of the signal. We propose \textbf{PASS} (\emph{Process Advantage Signal Shaping}), a compact middleware that sits between any scalar step-level process signal and GRPO's clipped surrogate and addresses the three pathologies in turn: \emph{Advantage Fusion} standardizes the three streams independently within each group, \emph{Chunk-by-Value} derives value-homogeneous chunks from the signal itself and broadcasts credit within each chunk, and \emph{Divide-Length} converts the cumulative objective into an average-value-density score. We validate PASS across two domains and two process-signal paradigms -- a learned PRM on mathematical reasoning and an on-policy-distillation KL signal (with a generalized variant) on multi-hop question answering -- and under two group-standardization operators. In every regime PASS delivers a consistent pass@1 gain over the corresponding GRPO baseline.

13:00 JSTLLM/生成AIエージェントClaude

階層的実験主義エージェント

大規模言語モデル (LLM) は、現実世界でアクションを実行し、人間の意思決定をサポートするためにますます使用されていますが、ほとんどのエージェントはパラメトリック知識、トレーニング後の固定データ、取得、または検索に依存しています。このパラダイムは、新しい領域や、事前知識だけでは答えられない高度なクエリでは破綻します。たとえば、物理法則を知っているだけでは、LLM がクエリに答えたり、複雑な物理システム内で長期的なタスクを完了したりできるわけではありません。これに対処するために、積極的な実験から学ぶためのコンテキスト内自己改善フレームワークである Hierarchical Experimentalist Agents (HExA) を導入します。 HExA は、クエリ関連の実験を繰り返し設計および改良し、経験から構成可能なスキルの再利用可能なライブラリを学習し、クエリに答えたりアクションを実行したりするための実験証拠を統合します。 HExA はトレーニング不要で、あらゆるブラックボックス モデルと互換性があり、外部の監視、オラクル、オフライン データを必要としません。アクティブな実験を評価するために、PHYRE 2D プロシージャル物理環境上に構築されたツール呼び出しベンチマークである Interphyre を導入します。このベンチマークでは、エージェントが介入を提案し、シミュレーション API を通じて仮説をテストします。実験によると、現在の LLM エージェントはこれらの設定、特に Interphyre の最も困難なレベルで苦戦していることが示されています。 Claude Sonnet 4.6 は 2% の成功率しか達成していませんが、HExA は同じモデルを最大 77% の成功率まで改善します。 HExA は、オープンウェイト モデルも改善し、ReAct や Reflexion などのエージェント ベースラインを上回るパフォーマンスを発揮します。さらに、より簡単なレベルから学習し、積極的な実験を行わずに移転したスキルのみを使用して、HExA は 44% の成功を収め、学習したスキルの再利用性と汎用性を実証しました。全体として、HExA は、積極的な実験による学習が、エージェントが有用な知識を発見し、再利用可能なスキルを獲得し、新しい長期的なタスクを効率的に進めるのに役立つことを示しています。

原文 (English)

Hierarchical Experimentalist Agents

Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search. This paradigm breaks down in novel domains and for sophisticated queries that cannot be answered from prior knowledge alone. Knowing the laws of physics, for instance, does not by itself enable LLMs to answer queries or complete long-horizon tasks in a complex physical system. To address this, we introduce Hierarchical Experimentalist Agents (HExA), an in-context self-improvement framework to learn from active experimentation. HExA iteratively designs and refines query-relevant experiments, learns a reusable library of composable skills from experience, and integrates experimental evidence to answer queries or take actions. HExA is training-free, compatible with any black-box model, and does not require external supervision, oracles, or offline data. To evaluate active experimentation, we introduce Interphyre, a tool-calling benchmark built on the PHYRE 2D procedural physics environment, where agents propose interventions and test hypotheses through simulation APIs. Experiments show that current LLM agents struggle in these settings, especially on the hardest levels of Interphyre. Claude Sonnet 4.6 achieves only 2% success, while HExA improves the same model to up to 77% success. HExA also improves open-weight models and outperforms agentic baselines such as ReAct and Reflexion. Moreover, using only skills learned from easier levels and transferred without active experimentation, HExA achieves 44% success, demonstrating the reusability and generalization of its learned skills. Overall, HExA shows that learning through active experimentation can help agents discover useful knowledge, acquire reusable skills, and make efficient progress on novel long-horizon tasks.

13:00 JST研究/論文

PHF: ポリシーに基づく自己蒸留のための特権付き隠しフロー

ポリシーに関する自己蒸留 (OPSD) は、検証済みの参照ソリューションも参照する特権教師を照合することにより、独自のポリシーからサンプリングされたロールアウトに基づいて推論モデルをトレーニングします。既存の OPSD 目標は出力分布のみを監視するため、特権コンテキストは、その分布を生成した内部計算を直接監視することなく、トークンレベルの発散を通じてトレーニングに影響を与えます。私たちは、特権教師の非表示状態が同じロールアウトに沿ってどのように移動するかをさらに抽出する特権非表示フロー (PHF) を提案します。各生徒の隠れベクトルを同じトークン位置の教師ベクトルに強制的に一致させるのではなく、PHF は、選択された生成位置上でトークンからトークンへの遷移方向と軌道ジオメトリを調整します。全層レシピには、点ごとの隠れ状態の模倣を行わずに、これらの同じ遷移から計算された隣接層関係も含まれています。同じ 100 ステップのトレーニング スケジュールの下で、PHF は Qwen3-1.7B、4B、および 8B で再現した OPSD ベースラインを上回る Average@12 集計を改善し、約 +2.2、+1.5、および +1.7 ポイントのゲインが観察されました。輸送目標は、共有軌道オフセットに対して正確に不変です。その局所幾何項は、遷移方向の直交変換に対しても不変です。アブレーションは、固定 PHF レシピを点単位の隠れ状態マッチング、単一チャネル遷移損失、層サブセットの選択から区別し、OPSD へのコンパクトな隠れフロー拡張として PHF をサポートします。

原文 (English)

PHF: Privileged Hidden Flow for On-Policy Self-Distillation

On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that also sees verified reference solutions. Existing OPSD objectives supervise only the output distribution, so privileged context affects training through a token-level divergence without directly supervising the internal computation that produced that distribution. We propose Privileged Hidden Flow (PHF), which additionally distills how a privileged teacher's hidden states move along the same rollout. Rather than forcing each student hidden vector to match the teacher vector at the same token position, PHF aligns token-to-token transition directions and trajectory geometry over selected generated positions. The all-layer recipe also includes an adjacent-layer relation computed from these same transitions, without pointwise hidden-state imitation. Under the same 100-step training schedule, PHF improves the Average@12 aggregate over our reproduced OPSD baseline on Qwen3-1.7B, 4B, and 8B, with observed gains of about +2.2, +1.5, and +1.7 points. The transport objective is exactly invariant to shared trajectory offsets; its local geometry term is also invariant to orthogonal transformations of transition directions. Ablations distinguish the fixed PHF recipe from pointwise hidden-state matching, single-channel transition losses, and layer-subset choices, supporting PHF as a compact hidden-flow extension to OPSD.

13:00 JSTLLM/生成AIエージェント

LLM が言語を開発するとき: 効率的なマルチエージェント推論のための記号コミュニケーション

思考連鎖 (CoT) は、困難な推論タスクにおける大規模言語モデル (LLM) を改善しますが、多くの場合、効率的な機械推論との整合性が不十分な長い自然言語理論が発生します。我々は、複数の LLM エージェントが自律的にコンパクトな言語記号フレームワーク (LSF) を発明、進化、共有するテスト時フレームワークである Communicative Language Symbolism Routing (CLSR) を提案します。一方、潜在性のないルーターがクエリごとにこれらの言語を適応的に選択して構成し、精度とトークンのトレードオフを最適化します。表面的な命令を洗練するプロンプト最適化とは異なり、CLSR は各 LSF をコンパクトなシンボル、使用ルール、およびメッセージ パッシング コントラクトを備えた再利用可能なシンボリック プロトコルとして扱い、正確性とトークン コストによって駆動される進化的なループを通じて改善します。推論時に、ルータは単一の低コスト LSF 呼び出しを呼び出したり、複数の LSF をアンサンブルしたり、より困難なクエリに対してマルチラウンド LSF 構成プロトコルを実行したりすることがあります。 CLSR は、困難なベンチマーク全体にわたって、精度を維持しながら、レイテンシ指向の生成されたトークン完了を標準 CoT と比較して $3\sim 6\times$ 削減します。さらに、任意の記号体系の下でトークンコストの情報理論的な下限を導出し、インタープリタの実現可能性の前提の下で、マルチラウンド LSF プロトコルが条件付きでプログラム実行パイプラインを包含することを示します。コードは公開されています (https://github.com/pzqpzq/LSF_MDia)。

原文 (English)

When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning

Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs long natural-language rationales that are poorly aligned with efficient machine reasoning. We propose Communicative Language Symbolism Routing (CLSR), a test-time framework in which multiple LLM agents autonomously invent, evolve, and share compact Language Symbolism Frameworks (LSFs), while a latent-free router adaptively selects and composes these languages per query to optimize the accuracy-token trade-off. Unlike prompt optimization that refines surface instructions, CLSR treats each LSF as a reusable symbolic protocol with compact symbols, usage rules, and a message-passing contract, and improves it through an evolutionary loop driven by correctness and token cost. At inference time, the router may invoke a single low-cost LSF call, ensemble multiple LSFs, or execute a multi-round LSF composition protocol on harder queries. Across challenging benchmarks, CLSR reduces latency-oriented generated token completion by $3\sim 6\times$ compared to standard CoT while maintaining accuracy. We further derive an information-theoretic lower bound on token cost under arbitrary symbolism and show that, under an interpreter-realizability premise, multi-round LSF protocols conditionally subsume program-execution pipelines. Code is publicly available (https://github.com/pzqpzq/LSF_MDia).

13:00 JST研究/論文

予算制約の下での RAG の事実誤認の診断と修復

検索拡張生成 (RAG) は、応答を外部証拠に根拠づけることによって大規模な言語モデルの事実性を向上させますが、現実世界の展開は依然として脆弱です。失敗は多くの場合、証拠の欠落または関連性の薄さ、および取得されたコンテキストを忠実に反映していない生成によって発生します。既存のアプローチの多くは、微調整、内部モデル信号への特権アクセス、またはリソースに依存しないエスカレーション戦略に依存しているため、ブラックボックスや予算に制約のある設定では実用性が制限されます。我々は、軽量な障害診断と適応修復を組み合わせた、モデルに依存しないリソース認識フレームワークである D2R-RAG (Diagnose-to-Repair RAG) を提案します。 D2R-RAG は、クエリ内の観察可能な信号、取得された証拠、生成された応答から解釈可能な障害シグネチャを導き出し、明示的なレイテンシと VRAM 制約の下で少数の修正アクションから選択します。 FEVER と HotpotQA の実験では、D2R-RAG が最近のベースラインよりも信頼性を向上させ、複数のコンピューティング バジェットにわたる効率のトレードオフにより、より優れた精度を達成していることが示されています。コードは https://github.com/Cyber​​ScienceLab/D2R-RAG/ で入手できます。

原文 (English)

Diagnosing and Repairing Factual Errors in RAG under Budget Constraints

Retrieval-Augmented Generation (RAG) improves the factuality of large language models by grounding responses in external evidence, yet real-world deployments remain fragile. Failures often stem from missing or weakly relevant evidence, as well as from generation that does not faithfully reflect the retrieved context. Many existing approaches rely on fine-tuning, privileged access to internal model signals, or resource-insensitive escalation strategies, which limits their practicality in black-box and budget-constrained settings. We propose D2R-RAG (Diagnose-to-Repair RAG), a model-agnostic and resource-aware framework that combines lightweight failure diagnosis with adaptive repair. D2R-RAG derives interpretable failure signatures from observable signals in the query, retrieved evidence, and generated response, and then selects from a small set of corrective actions under explicit latency and VRAM constraints. Experiments on FEVER and HotpotQA show that D2R-RAG improves reliability over recent baselines and achieves better accuracy--efficiency trade-offs across multiple compute budgets. The code is available at https://github.com/CyberScienceLab/D2R-RAG/.

13:00 JSTLLM/生成AI

マルチモーダル原子力規制文書に対するマルチホップ推論のための LLM ガイドに基づく計画

原子力規制文書のレビューには、数万ページにわたるマルチホップ推論が必要であり、判断は複数の章にわたって集められた証拠に依存します。このタスクを計画として組み立てます。LLM ベースのエージェントは、これまでに収集された証拠を観察し、検査する次の文書の断片を選択し、証拠が十分になったときに停止します。エージェントは、ブラウズ、読み取り、検索ツールを使用してベクターレス ドキュメント ツリー上で動作し、動的なナレッジ グラフ (KG) を状態として維持します。 NuScale 最終安全分析レポート (FSAR) 文書に対する 200 の質問のベンチマークでは、システムの精度は 81.5% に達し、RAGAS 忠実度は 0.93 でした。支配的なパフォーマンス要因は計画です。状態条件付きアクション選択なしで同じドキュメント ツリーを使用する PageIndex と比較すると、その差は +38.0pp (43.5% ~ 81.5%、p<0.001) です。また、このシステムは LightRAG (73.0%、p<0.05)、HippoRAG (70.5%、p<0.01)、GraphRAG (49.5%、p<0.001) よりも優れており、オフライン インデックス作成なしで RAPTOR (75.5%、p=0.11) に匹敵します。エッジ推論では、精度を高めることなく 2.8 倍のコストが追加されます。私たちはそれをトレーサビリティモジュールとして保持します。 7,391 の推定エッジのうち、3 つの違反エッジ (0.04%) は、人間のレビュー担当者が監査できる型付きアノテーションとして、スコープ境界 (Q058) と部分適合 (Q176) にフラグを立てます。

原文 (English)

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

Reviewing nuclear regulatory documents requires multi-hop reasoning across tens of thousands of pages, where judgments depend on evidence assembled across multiple chapters. We frame this task as planning: an LLM-based agent observes the evidence collected so far, picks the next document fragment to inspect, and stops when the evidence is sufficient. The agent operates over a vectorless document tree using browse, read, and search tools, and maintains a dynamic knowledge graph (KG) as state. On a 200-question benchmark over NuScale Final Safety Analysis Report (FSAR) documents, the system reaches 81.5% accuracy with a RAGAS Faithfulness of 0.93. The dominant performance factor is planning: against PageIndex, which uses the same document tree without state-conditioned action selection, the gap is +38.0pp (43.5% to 81.5%, p<0.001). The system also outperforms LightRAG (73.0%, p<0.05), HippoRAG (70.5%, p<0.01), and GraphRAG (49.5%, p<0.001), and matches RAPTOR (75.5%, p=0.11) without offline indexing. Edge inference adds 2.8x cost without raising accuracy; we retain it as a traceability module. Of 7,391 inferred edges, 3 Violates edges (0.04%) flag scope boundaries (Q058) and partial conformance (Q176) as typed annotations that a human reviewer can audit.

13:00 JSTLLM/生成AIエージェント

討論者の混合: マルチエージェント推論におけるアーキテクチャレベルでの討論を学ぶ

既存のマルチエージェント ディベート フレームワークには、2 つの重大な制限があります。1 つは、エージェントの役割と調整パターンが設計時に固定される静的アーキテクチャに依存していること、もう 1 つは複数のモデルのコピーをインスタンス化する必要があるため、かなりの計算オーバーヘッドが発生することです。私たちは、専門家の混合パラダイムを活用することで単一モデル内で動的な自己討論を可能にする統一フレームワークである混合討論者 (MoD) を提案します。 MoE を弁証法的推論に適応させる際の 3 つの重要な課題に取り組みます。(1) 役割の割り当てをプロセス フローから切り離し、議論する時期と総合する時期を動的に決定するデュアル ルーティング。 (2) ローカル コンテキストによるトークン レベルのルーティングをスムーズにし、エキスパート スイッチのジッターを軽減するモメンタム スイッチング。 (3) 多様な討論者を軽量のエキスパート モジュールにカプセル化し、行動の多様性を維持しながらエージェント間のコミュニケーションを排除する統合された自己討論。マルチモーダル ベンチマークに関する広範な実験により、MoD が単一モデル ベースラインと従来のマルチエージェント システムの両方を上回るパフォーマンスを示し、レイテンシが 3.7 倍低く、トークン消費量が 87% 削減されるという優れた精度を達成していることが実証されています。ソース コードは https://github.com/YongLD/MoD でアクセスできます。

原文 (English)

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning

Existing multi-agent debate frameworks suffer from two critical limitations: they rely on static architectures where agent roles and coordination patterns are fixed at design time, and they require instantiating multiple model copies, incurring substantial computational overhead. We propose Mixture of Debaters (MoD), a unified framework that enables dynamic self-debate within a single model by leveraging the Mixture-of-Experts paradigm. We address three key challenges in adapting MoE for dialectical reasoning: (1) dual-routing that decouples role allocation from process flow, dynamically determining when to debate versus when to synthesize; (2) momentum switching that smooths token-level routing with local context, reducing expert-switch jitter; and (3) unified self-debate that encapsulates diverse debating personas into lightweight expert modules, eliminating inter-agent communication while preserving behavioral diversity. Extensive experiments on multimodal benchmarks demonstrate that MoD outperforms both single-model baselines and conventional multi-agent systems, achieving superior accuracy with 3.7x lower latency and 87% reduction in token consumption.The source code can be accessed at https://github.com/YongLD/MoD.

13:00 JST研究/論文

FADE: 大規模な視覚言語モデルにおける言語優先支配を軽減することによる幻覚の軽減

Large Vision-Language Model (LVLM) の優れた機能にもかかわらず、依然として幻覚の影響を受けやすく、入力画像と一致しないコンテンツが生成されます。最近の研究では、これは視覚入力に対する言語事前の優位性によるものであり、この優位性を緩和するために対照的なデコード方法が採用されていますが、そのメカニズムの起源は未解明のままです。各変換層を通る情報の流れを調査すると、アテンション モジュールが一貫して視覚的証拠を集約し、クリティカル層の FFN モジュールが言語事前情報のソースとして機能することがわかりました。これらの事前分布は視覚的な証拠を無効にする可能性があり、中間層での正しい予測が不正確な出力に向かってドリフトする原因となります。この洞察に基づいて、言語優先の優位性を減らすために FFN 出力を減衰するトレーニング不要の方法である FADE (FFN Attenuation for DEcoding) を提案します。 LLaVA-1.5、mPLUG-Owl2、および InstructBLIP にわたる POPE、CHAIR、および MME ベンチマークの評価では、FADE が推論効率を維持しながら幻覚を効果的に軽減することが示されています。

原文 (English)

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models

Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language priors over visual inputs and employ contrastive decoding methods to mitigate this dominance, but the mechanistic origin remains unexplored. We investigate the information flow through each transformer layer and find that attention modules consistently aggregate visual evidence, while FFN modules at critical layers act as the source of language priors. These priors can override visual evidence, causing correct predictions in intermediate layers to drift toward incorrect outputs. Based on this insight, we propose FADE (FFN Attenuation for DEcoding), a training-free method that attenuates FFN outputs to reduce language-prior dominance. Evaluations on POPE, CHAIR, and MME benchmarks across LLaVA-1.5, mPLUG-Owl2, and InstructBLIP show that FADE effectively mitigates hallucinations while preserving inference efficiency.

13:00 JST研究/論文

入札する前にどの程度のデューデリジェンスを行いますか?難解な買収オークションで学んだこと

2 つの企業が同じターゲットの購入に入札した場合、そのターゲットの価値を正確に知る人は誰もいません。各入札者はデューデリジェンスの費用を支払います。これは、入札前に独自の個人的な見積もりを精緻にする、高価で不完全な宿題です。その宿題のうちどれくらいを買う価値がありますか?私たちは入札コンテストの単純なコンピューター モデルを構築し、ゲーム エンジンがチェスを学習するのと同じように、コンピューター自身と対戦することでうまく入札できるように学習させます。どれだけの勤勉さが報われるかという経済的な問題と、コンテストが複雑すぎて正確に解決できない場合の計算の問題はどちらも、入札者がどれだけ多くの個人情報を保持しているかという 1 つのことによって制御されます。私たちの主な発見は、適切な量の勤勉さは控えめであり、有限であるということです。勤勉さがより高価になるにつれてこの値は下がり、双方が下調べをしているときはさらに下がります。競争によってより多くを知ることの価値が損なわれるからです。また、AI 研究による最近の主張も検証します。つまり、シンプルで一般的なセルフプレイ方法は、このようなゲーム用に通常構築される特殊で高価なアルゴリズムに匹敵する可能性があるということです。高価なフロンティア AI を搭載していない普通のラップトップで実行すると、単純な方法が自己学習アプローチの中で最も優れていることがわかりますが、ゲームが完全に解決できるほど小さい場合は常に、専用に構築された正確な方法が依然として勝利します。シンプルな手法は、ゲームが大きくなりすぎて正確に解決できなくなった場合にのみ利益を得ることができます。これが実際の取引が存在する体制であり、そこでは彼らが依然として強力な入札戦略を見つけていることがわかります。貢献は 3 つあります。不確実性の下での取引形成を研究する安価で再現可能な方法です。どの程度のデューデリジェンスを購入する価値があるのか​​についての、具体的なモデルベースの回答。そして、軽量の汎用 AI が特殊な手法に取って代わるのに十分な場合についての証拠も得られます。すべてのゲーム、コード、実験をリリースします。

原文 (English)

How Much Due Diligence Before You Bid? Learning in Intractable Takeover Auctions

When two companies bid to buy the same target, no one knows exactly what the target is worth. Each bidder pays for due diligence: costly, imperfect homework that sharpens its own private estimate before it bids. How much of that homework is worth buying? We build a simple computer model of the bidding contest and let it teach itself to bid well by playing against itself, the way a game engine learns chess. The economic question, how much diligence pays for itself, and the computational question, when the contest becomes too complex to solve exactly, are both controlled by a single thing: how many pieces of private information a bidder carries. Our main finding is that the right amount of diligence is modest and finite. It falls as diligence gets more expensive, and it falls further when both sides are doing their homework, because competition erodes the value of knowing more. We also test a recent claim from AI research: that simple, general self-play methods can rival the specialized, expensive algorithms usually built for games like these. Running on an ordinary laptop with no costly frontier AI, we find the simple methods are the best of the self-learning approaches, though purpose-built exact methods still win whenever the game is small enough to solve outright. The simple methods earn their keep only once the game grows too large to solve exactly, which is the regime real deals live in, and there we show they still find strong bidding strategies. The contribution is threefold: a cheap, reproducible way to study deal-making under uncertainty; a concrete, model-based answer to how much due diligence is worth buying; and evidence about when lightweight, general-purpose AI is good enough to replace specialized methods. We release all the games, code, and experiments.

13:00 JSTエージェントGemini

エージェントとコンピュータの監視インターフェイスにより、コンピュータの動的な使用が可能になります

SWE エージェントは、ソフトウェア エンジニアリング エージェントの未開発の設計軸としてアクション インターフェイスを確立しました。コンピュータ使用 (CU) エージェントの観察インターフェイスについても同様のケースを考えます。現在の CU エージェントは、クローズドソースでもオープンソースでも同様に、観察とアクションを結び付けています (3 ~ 5 秒ごとに 1 つのスクリーンショットがあり、音声はありません)。スクリーンショットとビデオ、アニメーション、一時的な UI イベント、会議、音声指示の間では、目が見えず耳が聞こえないままになっています。 Agent-Computer Observation Interface (AOI) を導入します。これは、ステップ間キーフレーム キャプチャ、ボリューム ゲート オーディオ トランスクリプション、テキストとして保持される CU モデル生成のビジュアル ナレーションという 3 つのゲート コンポーネントを通じて、連続的で適応的な観察を離散アクションから切り離す、モデルに依存しない認識レイヤーです。それぞれは、静的でサイレントなコンテンツではほとんど何も生成せず、品質を低下させることなく標準ループに削減します。 DynaCU-Bench (100 の動的ブラウザ タスクと 50 タスクの静的制御) では、7B からフロンティア スケールの CU モデルは再トレーニングなしでスクリーンショット ベースラインを +17 ~ +48 pp 獲得し、定期的なスクリーンショットではほぼ不可能だったタスクをほぼ解決されたタスクに変えます。ギャップはオーディオで最も顕著です。音声コンテンツのサブセットでは、AOI エージェントがあらゆるタスクを解決しますが、ストリーミング音声モデルは正確に聞きますが、足場がなければ聞いたことに基づいて行動することができません。この分解は、ヘッドラインのゲインと同じくらい有益です。キーフレームの選択は重要ではないことが判明しました。値は、キャプチャされたフレームを永続的なテキストにナレーションすることから得られます。また、新しいモデル (Gemini 3 Flash) では、キーフレーム ストリームが画像トークンの希釈によって積極的に後退するため、インターフェイスは固定バンドルではありません。そのため、そのコンポーネントは 1 つの構成として出荷されるのではなく、モデルごとに選択する必要があります。

原文 (English)

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for the observation interface in computer-use (CU) agents. Current CU agents, closed and open-source alike, tie observation to action--one screenshot every 3-5 s, no audio--leaving them blind and deaf between screenshots to video, animations, transient UI events, meetings, and spoken instructions. We introduce the Agent-Computer Observation Interface (AOI), a model-agnostic perception layer that decouples continuous, adaptive observation from discrete actions through three gated components: inter-step keyframe capture, volume-gated audio transcription, and CU-model-generated visual narration that persists as text. Each produces almost nothing on static, silent content, reducing to the standard loop without degrading it. On DynaCU-Bench (100 dynamic browser tasks plus a 50-task static control), CU models from 7B to frontier scale gain +17 to +48 pp over their screenshot baselines with zero retraining, turning tasks that are near-impossible from periodic screenshots into largely solved ones. The gap is starkest on audio: on a spoken-content subset AOI agents solve every task, whereas streaming voice models hear accurately but cannot act on what they hear without the scaffold. The decomposition is as informative as the headline gain: keyframe selection turns out not to matter--the value comes from narrating captured frames into persistent text--and the interface is not a fixed bundle, since on a newer model (Gemini 3 Flash) the keyframe stream actively regresses through image-token dilution, so its components must be selected per model rather than shipped as one configuration.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

正式なベンチマークの欠陥: データセットの欠陥とリーン定理証明における評価の失敗

リーンでの LLM 支援定理証明のベンチマークは、解決されたすべてのインスタンスに機械チェックされた証明が付属しているため、多くの場合、本質的に信頼できるものとして扱われます。ただし、カーネルは証明が \emph{formal} ステートメントを確立することをチェックするだけです。ステートメントが意図した非公式な問題を忠実にエンコードしているかどうか、評価ハーネスが自明な解決策や敵対的な解決策に対して堅牢であるかどうかは検証されません。私たちは、広く使用されている 5 つのリーン定理証明ベンチマークとそのフォークを監査し、コーパス スケールの静的チェッカーを使用して、反例、空定理、不健全な公理などの機械的に証明された 398 の問題を含む 4,833 の結果を明らかにしました。また、仮説の欠落、問題の単純化、不完全または不正確な翻訳、リーン固有の仕様上の危険などの意味論的な欠陥も文書化します。データセットの構築を超えて、評価時の故障モードを調査し、修正されたサブセット上で、欠陥によって報告される証明者スコアが膨らむことも小さくなる可能性があることを示します。私たちは、障害分類法、自動チェッカーのスイート、およびリコール指向のセマンティック監査プロンプトを提案し、正式な数学データセットの作成をガイドし、評価をより再現可能で信頼できるものにするための標準をリリースします。チェッカー、監査プロンプト、および修正されたデータセットのスナップショットは、https://github.com/Shachi456/atp-checkers で入手できます。

原文 (English)

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof. However, the kernel only checks that a proof establishes a \emph{formal} statement; it does not verify that the statement faithfully encodes the intended informal problem, nor that evaluation harnesses are robust to trivial or adversarial solutions. We audit five widely used Lean theorem-proving benchmarks and their forks, using corpus-scale static checkers to surface 4,833 findings, including 398 mechanically certified issues such as counterexamples, vacuous theorems, and unsound axioms. We also document semantic defects such as missing hypotheses, problem simplification, incomplete or incorrect translations, and Lean-specific specification hazards. Beyond dataset construction, we survey evaluation-time failure modes and show, on corrected subsets, that defects can both inflate and deflate reported prover scores. We propose a fault taxonomy, a suite of automated checkers and recall-oriented semantic audit prompts, and release standards to guide the creation of formal math datasets and to make evaluation more reproducible and trustworthy. Our checkers, audit prompts, and corrected dataset snapshots are available at https://github.com/Shashi456/atp-checkers.

13:00 JSTビジネス/資金調達GPT / ChatGPTLlama

プロセスレベルの社会的影響評価のための認知世界モデル

社会的影響ダイアログは、内部の認知状態を変えることでユーザーの行動を変えます。評価の中心となる質問は、ユーザーの信念、欲望、意図、感情が会話の過程で測定可能なほど変化するかどうかであり、これは表面レベルのテキスト指標 (BLEU/ROUGE) や単一スコアの LLM 判定では捉えることができないプロセス指向の基準です。我々は \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)} を提案します。これは、マルチターン対話評価を「ユーザーが何を言ったか」から「ユーザーの内部認知状態がどのように進化したか」に再構成する LLM ベースのユーザー モデルです。CogWM は、BDI/E 認知状態とユーザー発話を共同で予測し、3 層を使用してユーザー シミュレーターと評価プラットフォームの両方として機能します。ターンレベルの忠実度、軌道レベルの状態ダイナミクス、タスクレベルの複合スコアリングをカバーする評価フレームワーク。 4 つの社会的影響シナリオにわたる 150,454 のユーザー ターン サンプルで \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} アノテーション パイプラインを介してトレーニングされた CogWM は、77.6\% の感情精度 (GPT-5.5 の 2.1$\time$) を達成しました。 3,600 件のマルチエージェント識別試験において、認知的影響力によって 6 つの営利エージェントを区別し、Llama-4-Scout が 1 位にランクされました (CTS +0.233)。 CogWM は、社会的影響対話の評価を最終的な判断からプロセスの追跡に移行します。コード\脚注{\scriptsize コード: https://github.com/lucianma05-create/CogWM} とモデル\脚注{モデル: https://www.modelscope.cn/models/LucianMa/CogWM-14B} をリリースしました。

原文 (English)

Cognitive World Models for Process-Level Social Influence Evaluation

Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, desires, intentions, and emotions measurably change over the course of conversation, a process-oriented criterion that neither surface-level text metrics (BLEU/ROUGE) nor single-score LLM judgments can capture. We propose the \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)}, an LLM-based user model that reframes multi-turn dialogue evaluation from ``what did the user say'' to ``how did the user's internal cognitive state evolves.'' CogWM jointly predicts BDI/E cognitive states and user utterances and serves as both a user simulator and an evaluation platform, using a three-tier evaluation framework that covers turn-level fidelity, trajectory-level state dynamics, and task-level composite scoring. Trained via our \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} annotation pipeline on 150,454 user-turn samples across four social influence scenarios, CogWM achieves 77.6\% emotion accuracy (2.1$\times$ over GPT-5.5). In 3600 multi-agent discrimination trials, it distinguishes six commercial agents by their cognitive influence, with Llama-4-Scout ranking first (CTS +0.233). CogWM moves social influence dialogue evaluation from terminal judgment to process tracking. We have released our code\footnote{\scriptsize Code: https://github.com/lucianma05-create/CogWM} and models\footnote{Model: https://www.modelscope.cn/models/LucianMa/CogWM-14B}.

13:00 JSTLLM/生成AIエージェント

UCOB: クレジットを意識したポリシーに基づく双方向自己蒸留によるエージェント スキルの活用と進化の学習

スキル記憶は過去の経験をテキストガイダンスとして再利用することでエージェント強化学習を改善できますが、取得されたスキルは神託ではありません。ある状態では役立つ一方で、別の状態では同じポリシーを誤解させる可能性があります。これにより、一般的な特権教師の仮定が脆弱になります。つまり、スキル条件付きプロンプトは、スキルなしプロンプトに対する固定教師として扱うことができるということです。クレジットを意識したポリシーに基づく双方向の自己蒸留を通じて、エージェント スキルの活用と進化を学習するためのフレームワークである UCOB を紹介します。 UCOB は、スキル条件付きプロンプトとスキルなしプロンプトを同じモデルの 2 つのオンポリシー コンテキスト ビューとして扱い、同じタスクおよびアンカー状態内でのリターンツーゴーを比較し、より高いリターンのビューをローカル教師として使用します。このローカル クレジット シグナルは、有用なスキル条件付き動作を内部化し、誤解を招くスキルの使用法を修正し、タスク/状態のスキル メモリの更新、ユーティリティを意識した検索、およびリフレクション自己トレーニングをガイドします。 ALFWorld、WebShop、Search-QA などのエージェント タスクの実験では、モデル スケール全体で UCOB がスキルフリー RL、スキルメモリ ベースライン、自己蒸留法よりも優れており、ALFWorld と WebShop で SOTA ベースラインよりも最大 23.5 ポイントおよび 18.0 ポイント向上していることが示されています。アブレーションと分析により、その中核となるメカニズムと効率がさらに検証されます。

原文 (English)

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another. This makes the common privileged-teacher assumption fragile, namely that a skill-conditioned prompt can be treated as a fixed teacher for the no-skill prompt. We introduce UCOB, a framework for learning to utilize and evolve agentic skills via credit-aware on-policy bidirectional self-distillation. UCOB treats skill-conditioned and no-skill prompts as two on-policy context views of the same model, compares their return-to-go within the same task and anchor state, and uses the higher-return view as the local teacher. This local credit signal internalizes useful skill-conditioned behavior, corrects misleading skill usage, and guides task/state skill memory updates, utility-aware retrieval, and reflection self-training. Experiments on agentic tasks, including ALFWorld, WebShop, and Search-QA, show that UCOB outperforms skill-free RL, skill-memory baselines, and self-distillation methods across model scales, with up to 23.5 and 18.0 point gains over SOTA baselines on ALFWorld and WebShop. Ablations and analyses further validate its core mechanisms and efficiency.

13:00 JSTエージェント研究/論文ClaudeGPT / ChatGPT

OSWorld2.0: 長期にわたる現実世界のタスクにおけるコンピュータ使用エージェントのベンチマーク

既存のコンピュータ使用ベンチマークは、現実世界のコンピュータ使用の現実性、複雑さ、長期的な要求を捉えることができず、フロンティア エージェントの限界を明らかにする能力が制限されています。 OSWorld 2.0 は、複雑で困難な現実世界の現象を捉えるように設計された、日常業務から専門業務にわたる 108 の長期的なコンピューター使用ワークフローのベンチマークです。各タスクは現実的なエンドツーエンドのワークフローを表しており、人間のユーザーが完了するまでに中央値で約 1.6 時間かかり、Claude Opus 4.7 では最大限の思考を使用して平均 318 回のツール呼び出しが必要ですが、OSWorld 1.0 では約 30 回です。 OSWorld 2.0 は、ストリーミング インタラクションや動的環境などのインタラクション設計の課題だけでなく、クロスソース推論、暗黙的状態推論、視覚空間精度などのエージェント パターンの課題にも及ぶ、実際のワークフローでは一般的であるものの、以前のベンチマークでは過小評価されていた課題現象をターゲットにしています。タスクは本物の入力アーティファクトに基づいており、現実的なステートフル ユーザー プロファイル データと相互参照され、安全性を重視した実行を監査する個別の安全性レポートが含まれています。 500 ステップの主要なバイナリ完了基準では、最大限の思考とバッチ化されたツールを備えた Claude Opus 4.8 が最高のスコアを示していますが、それでも部分スコア 54.8% でタスクの 20.6% しか完了していません。 GPT-5.5 はトークン効率がはるかに優れていますが、13% 付近で頭打ちになっています。これらの結果は、現在のエージェントがまだプロレベルのコンピュータ使用には程遠いことを示しています。基本的な GUI 制御やコーディングでつまずくのではなく、制約を見失い、タスクの途中で届く情報を見逃し、ユーザーに尋ねるのではなく推測し、検証をスキップし、タスクが回復する必要がある隠れた状態に左右されるときに最も苦労します。

原文 (English)

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, designed to capture complex and challenging real-world phenomena. Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0. OSWorld 2.0 targets challenge phenomena that are common in real workflows yet underrepresented in prior benchmarks, spanning interaction-design challenges such as streaming interaction and dynamic environments, as well as agent-pattern challenges such as cross-source reasoning, implicit-state inference, and visual-spatial precision. Tasks are grounded in authentic input artifacts and cross-referenced against realistic stateful user profile data, and include separate safety reports auditing safety-sensitive execution. Under our primary binary-completion metric at 500 steps, Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score; GPT-5.5 is far more token-efficient yet plateaus near 13%. These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden state they must recover.

13:00 JSTエージェント

協調型 MARL における学習された調整規則: 理論に基づいた役割と学習されたルーティングの間の変換ギャップの測定

役割意味論的な割り当ては、異種エージェントがどのように調整するかに関する事前情報を提供しますが、協調的な MARL システムは代わりに、分散型の非定常学習を通じて規則に落ち着き、結果として得られる構造がそれらの事前情報と一致するという保証はありません。私たちは、3 つの役割の MiniGrid および SMACv2 (Terran) 環境にわたる役割ルーティング マトリックス、形成感度 ($\Delta_{\max}$)、および勾配/オクルージョン属性を組み合わせた診断を通じて、理論に基づいた役割の期待と学習された調整構造の間の変換ギャップを研究します。ラベル条件付きアテンションは、フラットな MLP ベースラインよりも大幅に集中した役割固有のルーティングを生成し、3 対 3 ~ 9 対 9 のスケーリングの下で​​も安定し、チーム サイズ全体でゼロショットを転送し、味方スロットのパディングに対して不変であることを示します。 5 シードの再評価では、学習した規則と設計者が指定した事前規則の間の部分的な整合性が示されると同時に、小さな n ノイズが明らかな戦略的相違を生み出す可能性がある場所が明らかになります。我々は、これらの結果を、新しい平衡概念や因果関係の説明としてではなく、協調的MARLにおける配位構造を測定するための経験的枠組みとして提示する。

原文 (English)

Learned Coordination Conventions in Cooperative MARL: Measuring the Translation Gap Between Theory-Informed Roles and Learned Routing

Role-semantic assignments provide priors over how heterogeneous agents may coordinate, but cooperative MARL systems instead settle on conventions through decentralized, non-stationary learning, with no guarantee that the resulting structure matches those priors. We study this translation gap between theory-informed role expectations and learned coordination structure through a diagnostic combining a role-routing matrix, formation sensitivity ($\Delta_{\max}$), and gradient/occlusion attribution across three-role MiniGrid and SMACv2 (Terran) environments. We show that label-conditioned attention produces substantially more concentrated and role-specific routing than flat MLP baselines, remains stable under 3v3--9v9 scaling, transfers zero-shot across team sizes, and is invariant to ally-slot padding. A 5-seed re-evaluation shows partial alignment between learned conventions and designer-specified priors while revealing where small-n noise can manufacture apparent strategic divergence. We present these results as an empirical framework for measuring coordination structure in cooperative MARL rather than as a new equilibrium concept or causal explanation.

13:00 JST研究/論文Llama

SCARCE: 埋め込みによるレアイベントの特性評価のためのスケーラブルなカスケード分析

まれなイベントが最新の AI システムの安全性プロファイルを支配しますが、その確率を推定するのは非常に困難です。直接モンテカルロ法では法外なサンプル予算が必要です。サブセット シミュレーション (SS) は、まれなイベントの確率を、ネストされた中間イベントに対する中程度の条件付き確率に分解することで、この問題に対処します。ただし、古典的な SS には手作りのスカラー パフォーマンス関数が必要で、そのサブレベル セットがイベントを定義するため、障害ジオメトリの詳細な知識が必要となり、新しいドメインへの転送が制限されます。私たちは、パフォーマンス関数を学習された潜在表現と障害領域への近さをスコア化する幾何学的定規に置き換える SCARCE (Scalable Cascade Analysis for Rare-event Characterization via Embeddings) を提案します。適応型しきい値処理は、ネストされた中間イベントをデータから直接構築します。非負のスーパーマルチンゲールを通じて SCARCE を形式化し、早期停止下でも有効な高確率の上部エンベロープを生成します。 MNIST の誤分類では、高密度モンテカルロがグラウンド トゥルースを提供します。SCARCE は、体系的な過剰カウントを排除しながら、グリッド検索された従来の SS よりも約 400 ~ 500 倍低い平均絶対誤差を達成します。次に、敵対的割合 $\eta$ を使用したフリートレベルの脅威モデルの下で、PAIR スタイルの LLM ジェイルブレイクを研究します。 Llama-Guard-3-8B の隠れ状態では、PCA ベースのルーラーは、平均ブートストラップ相対半値幅が 27.9% である有限サンプル参照に対して $\eta \geq 10^{-3}$ の平均相対誤差 2.6% を達成し、再調整後に 2.93% の相対誤差を持つ GCG スタイルのコーパスに移行します。方向基準 $\mathrm{KL}(p_{\mathrm{good}}\,\|\,p_{\mathrm{bad}})$ は、推定誤差と一貫して定規をランク付けします (Spearman $\rho=0.83$)。

原文 (English)

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires prohibitive sample budgets. Subset Simulation (SS) addresses this by decomposing a rare-event probability into moderate conditional probabilities over nested intermediate events. However, classical SS requires a handcrafted scalar performance function whose sublevel sets define those events, demanding detailed knowledge of the failure geometry and limiting transfer to new domains. We propose SCARCE (Scalable Cascade Analysis for Rare-event Characterisation via Embeddings), which replaces the performance function with learned latent representations and geometric rulers that score proximity to failure regions. Adaptive thresholding constructs nested intermediate events directly from data. We formalise SCARCE through a non-negative supermartingale, yielding a high-probability upper envelope that remains valid under early stopping. On MNIST misclassification, where dense Monte Carlo provides ground truth, SCARCE achieves approximately 400--500 times lower mean absolute error than grid-searched traditional SS while eliminating systematic over-counting. We then study PAIR-style LLM jailbreaks under a fleet-level threat model with adversarial fraction $\eta$. On Llama-Guard-3-8B hidden states, a PCA-based ruler attains 2.6% mean relative error for $\eta \geq 10^{-3}$ against finite-sample references whose average bootstrap relative half-width is 27.9%, and transfers to a GCG-style corpus with 2.93% relative error after recalibration. A directional criterion $\mathrm{KL}(p_{\mathrm{good}}\,\|\,p_{\mathrm{bad}})$ ranks rulers consistently with estimation error (Spearman $\rho=0.83$).

13:00 JST研究/論文GPT / ChatGPT

SFBench: SciFy の科学的実現可能性ベンチマーク

科学的主張の実現可能性を評価するシステムを評価するためのベンチマーク データセットである SFBench を紹介します。 SFBench には材料科学に関する 197 件のクレームが含まれており、それぞれに 5 段階評価のグラウンドトゥルース実現可能性スコアとその評価の説明が注釈として付けられています。このコレクションは、いくつかの重要な点で以前のコレクションとは異なります。1) さまざまな科学的実現可能性の主張に対する推論を必要とする複雑なタスクを定義しています。 2) その主張は既存の科学出版物から抽出されたものではなく、新たに作成されたものであるため、LLM がそれらについて訓練を受けている可能性は大幅に減少します。 3) 主張と真実は、人工知能ではなく、対象分野の専門家によって確立されます。 4) 質問と回答のペアについて質問したり、複数の選択肢の回答を提供したり、短く固定された回答を要求したりする多くのベンチマークとは異なり、SFBench の説明は完全に自由です。ベンチマーク設計、データ作成プロセス、評価指標について説明し、最新の GPT モデルを使用したベースライン結果を報告します。

原文 (English)

SFBench: The SciFy Scientific Feasibility Benchmark

We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in materials science, each annotated with a ground-truth feasibility score on a five-point scale along with an explanation of that assessment. The collection differs from previous collections in several important ways: 1) it defines a complex task that requires reasoning over claims of varying scientific feasibility; 2) its claims are not extracted from existing scientific publications but are created de novo, greatly reducing the chances that LLMs have trained on them; 3) claims and ground truth are established by subject matter experts, not by artificial intelligence; and 4) unlike many benchmarks that ask about question/answer pairs, provide multiple choice answers, or ask questions requiring short, fixed answers, SFBench explanations are completely open-ended. We describe the benchmark design, data creation process, and evaluation metrics, and we report baseline results using recent GPT models.

13:00 JSTLLM/生成AIエージェント

ローカル信頼性限界を使用した予算付きの行為または延期マルチエージェント LLM 審議

LLM 間で複数のエージェントによる審議を行うことで推論を改善できますが、導入には、現在の回答が行動に移すのに十分な信頼性がある場合と、人間によるレビューにエスカレーションする必要がある場合を判断する必要があります。私たちはこれを、予算に基づいた実行または延期の意思決定として定式化します。各ラウンドで、システムはディベート プレフィックスを低次元の状態にマッピングし、キャリブレーション データを使用して状態条件の正しさに関する $k$-最近傍下位信頼限界を計算し、その限界がユーザー指定の信頼性しきい値を超えた場合にのみ動作します。証明書は、分解 $\beta = \delta + \alpha + \varepsilon_{\mathrm{act}}$ を通じて誤ったアクションを制御し、キャリブレーションの失敗、残留アクションのリスク、表現ギャップを分離します。この保証は条件付きであり、配布フリーではありません。有効なローカル バイアス エンベロープとアクション領域表現ギャップ境界に依存しており、各仮定は改ざんスタイルの診断と組み合わせられています。同じ絶対的な間違ったアクションの予算は、難易度の異なるタスク間では異なる意味を持つため、トレーニング データのみを使用して各タスクの最終ラウンドのエラーに相対的な予算を設定し、正規化された予算の使用法 $\mathrm{WA}/\beta$ によって安全性を評価します。 9 つのベースラインに対する 6 つのベンチマークでは、この手法はアクティブ化されたデータセットに対して事前に宣言された予算の 9 ~ 12% を使用し、最大 84% の自動化と 96% のアクション精度に達します。ストレス テスト データセットでは、信頼性の低い自動化を強制するのではなく、延期します。このメソッドは、タスクごとの事後しきい値検索に依存するのではなく、明示的に指定された前提条件に基づいて、ユーザーが宣言した誤ったアクションの予算を、展開前に監査可能なアクション・オア・ディファー動作点に将来に向けて変換します。

原文 (English)

Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds

Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review. We formulate this as budgeted act-or-defer decision making. At each round, the system maps the debate prefix to a low-dimensional state, computes a $k$-nearest-neighbor lower confidence bound on state-conditional correctness using calibration data, and acts only when the bound exceeds a user-specified reliability threshold. The certificate controls wrong actions through the decomposition $\beta = \delta + \alpha + \varepsilon_{\mathrm{act}}$, separating calibration failure, residual action risk, and representation gap. The guarantee is conditional, not distribution-free: it relies on a valid local bias envelope and an action-region representation-gap bound, and each assumption is paired with falsification-style diagnostics. Because the same absolute wrong-action budget has different meanings across tasks of different difficulty, we set budgets relative to each task's final-round error using training data only, and evaluate safety by normalized budget usage $\mathrm{WA}/\beta$. On six benchmarks against nine baselines, the method uses 9--12% of the pre-declared budget on activated datasets, reaching up to 84% automation and 96% acted-on accuracy; on stress-test datasets, it defers rather than forcing unreliable automation. Rather than relying on per-task post-hoc threshold search, the method prospectively converts a user-declared wrong-action budget into an auditable act-or-defer operating point before deployment, under explicitly stated assumptions.

13:00 JST研究/論文

利害関係のない AI 予測器における正直さによる安全性

AI システムの能力が高まるにつれて、下流の結果を最適化するトレーニング手順では、暗黙的なエージェンシー、つまり設計者が指定したことのない目標指向の行動が導入される危険性があります。我々は、「認識論的に文脈化された」自然言語ステートメントのデータセットに基づいて条件付けされたベイズ事後分布を近似するように訓練された、Scientist AI (SAI) Predictor の正式な安全性の議論を提示します。私たちは、そのような予測器は、それ自体が目標を達成するために出力を選択するエージェントではなくても、エージェント、アクション、およびその結果を正直に予測できると主張します。これはデータ表現とトレーニング手順に基づいています。テキストの認識論的文脈化は、潜在的な事実の主張とコミュニケーション行為を区別するため、目標の表現は、モデルの採用を推進するものではなく、説明される証拠として扱われます。事後探索型のトレーニング目標により、これは、Predictor を調整された慎重な予測に向けて駆動することを目的としています。トレーニングは進行するため、予測の展開による下流の効果が報酬シグナルとして機能することはありません。システムが必要とするあらゆるエージェンシーは、ガードレールによって制約された明示的な足場によって提供されます。トレーニングのダイナミクスと、議論された危険な予測子の希薄さに関する仮定の下で、トレーニングによって、指定されたしきい値を超える残留害をもたらす保護された展開を持つ予測子が生成される確率は小さいことを証明します。危険な予測子は、多くのクエリにわたって調整された方法で害を過小評価する必要がありますが、そのような調整されたパターンは初期化分布の下ではまれであり、直接トレーニング信号を受け取りません。正確性を確保するための制約は、調整された欺瞞を高コストにする制約と同じであるため、このフレームワークでは安全性と正確性が同時にサポートされています。 Predictor 自体の内部から生じる不整合やエージェンシーに対するこれらの保証は、Predictor をエージェント システムの一部として使用することを妨げるものではありません。

原文 (English)

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.

13:00 JST研究/論文Grok

多様性は AI 群の強みです

トップクラスの AI 予測システムは、将来の世界の出来事に関してスーパー予報士レベルの精度に近づいていますが、依然として主に、予測固有のコンテキスト収集と足場を組み合わせた既製の LLM に依存しています。アンサンブルを通じてこのレシピを改善する方法を研究します。固定数のサンプルが与えられた場合、精度を最大化するにはどの既製モデル予測を組み合わせる必要があるでしょうか? Metaculus AI ベンチマークの二項質問では、個々の精度では十分ではないことがわかりました。多くのフロンティア LLM は相関性の高い予測を行っており、同じまたは類似のモデルからの追加予測の価値が制限されています。代わりに、最も強力なアンサンブルは、正確だが多様な予測者を組み合わせ、 \model{Grok 4} などのモデルの予測が他のフロンティア LLM と相関性が低いため、不均衡に寄与します。これらの結果は、AI 群衆の強さは、より多くの予測を無差別にサンプリングすることではなく、モデル間の予測を補完的な誤差と組み合わせることによってもたらされ、モデルの品質と多様性の両方を明示的に最適化する予測システムを動機付けることを示唆しています。

原文 (English)

Diversity is the Strength of the AI Crowd

Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy? On binary questions from the Metaculus AI Benchmark, we find that individual accuracy is not enough: many frontier LLMs make highly correlated predictions, limiting the value of additional forecasts from the same or similar models. Instead, the strongest ensembles combine accurate but diverse forecasters, with models such as \model{Grok 4} contributing disproportionately because their predictions are less correlated with other frontier LLMs. These results suggest that the strength of the AI crowd comes not from sampling more forecasts indiscriminately, but from combining forecasts across models with complementary errors, motivating forecasting systems that explicitly optimize for both model quality and diversity.

13:00 JST研究/論文

確率的保証を備えたマルコフ決定プロセスにおける到達可能性の確率的原因のサンプル効率的な学習

マルコフ意思決定プロセス (MDP) の確率モデル検査は定量的な保証を提供しますが、多くの場合、望ましくない結果が発生する理由についての洞察は限定的です。確率引き上げ (PR) 因果関係は、その訪問によって指定された州に到達する確率が増加する州を特定することで、この問題に対処します。しかし、既存の PR 原因特定手法では、学習には適していない MDP 修正が使用されています。条件付きと無条件の到達可能性確率の間のギャップを遷移サンプルから検出するのは困難であり、構築には MDP の到達可能性確率が必要ですが、遷移確率が不明な場合には、これを利用できません。私たちは未知の MDP を研究し、PR 原因を特定するための確率的保証を備えた学習アプローチを提案します。私たちの重要な要素は、元の MDP の到達可能性値を使用せずに、PR 原因のチェックを 2 つの条件付き到達可能性クエリに減らす再起動ベースの MDP 変更です。私たちは、正しさを証明し、サンプルの複雑さの限界を確立し、状態を因果関係、非因果関係、または未決定として段階的に分類する両側値反復に基づいていつでも学習および確認できるアルゴリズムを開発します。 2 つのベンチマークの実験により、PR の原因を信頼性高く迅速に特定できることが実証されました。

原文 (English)

Sample-Efficient Learning of Probabilistic Causes for Reachability in Markov Decision Processes with Probabilistic Guarantees

Probabilistic model checking for Markov decision processes (MDPs) provides quantitative guarantees, but often offers limited insight into why undesired outcomes occur. Probability-raising (PR) causality addresses this by identifying states whose visitation increases the probability of reaching designated states. Existing PR-cause identification methods, however, use MDP modifications not well-suited for learning: the gap between conditional and unconditional reachability probabilities can be hard to detect from transition samples, and construction requires reachability probabilities of the MDP, which are unavailable when transition probabilities are unknown. We study unknown MDPs and propose a learning approach with probabilistic guarantees for PR-cause identification. Our key ingredient is a restart-based MDP modification that reduces PR-cause checking to two conditional reachability queries without using reachability values of the original MDP. We prove correctness, establish sample-complexity bounds, and develop an anytime learning-and-checking algorithm based on two-sided value iteration that progressively classifies states as causal, non-causal, or undecided. Experiments on two benchmarks demonstrate reliable and fast identification of PR causes.

13:00 JSTエージェント

Planner-in-the-Loop フィードバックによる大規模言語モデルの安全かつ信頼性の高い PDDL 形式化に向けて

計画では、多くの場合、実行可能かつ検証可能なシンボリック仕様が必要になります。自律システムまたは意思決定支援システムに導入された大規模な言語モデルの場合、そのような形式化の失敗は、検証不可能な決定、実行の失敗、または安全でないダウンストリーム動作につながる可能性があります。 NL-PDDL-Bench は、プランナーによって検証された実行可能性とオブジェクト数による制御された難易度スケーリングを備えた、自然言語から PDDL への仕様構築のためのマルチドメイン ベンチマークです。さらに、バリデーターとプランナー診断を使用して、ローカライズされた編集を通じて実行不可能な仕様を改訂するプランナーインザループ フレームワークを提案します。このインフラストラクチャに基づいて、トレーニング中にオンライン プランナーの呼び出しを必要とせずに、パラメーター効率の高い低ランク適応の監視微調整、直接優先最適化のためのオフライン プランナー由来の優先順位ペア、および推論時間のプランナー インザ ループ修復を組み合わせた、プランナーに基づいた最適化レシピを開発します。また、解析可能性、解決可能性、仕様の類似性、およびプランナー参照に対する結果を意識したプランレベルの一貫性のための統合された評価スイートも提供します。代表的なモデル ファミリでの実験では、プランナーの成功とプラン レベルの一致が大幅に向上し、難易度のスケーリングやクロスドメインの変動に対する堅牢性が向上していることが示されています。これらの結果は、安全性またはセキュリティ重視の計画システムにおける LLM の信頼できる導入のための、外部検証可能な形式化の価値を強調しています。コードとデータは、https://github.com/ibasicplan/NL-PDDL-Bench から入手できます。

原文 (English)

Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback

Planning often requires symbolic specifications that are both executable and verifiable. For large language models deployed in autonomous or decision-support systems, failures in such formalization may lead to unverifiable decisions, execution failures, or unsafe downstream behavior. We present NL-PDDL-Bench, a multi-domain benchmark for natural-language-to-PDDL specification construction with planner-verified executability and controlled difficulty scaling by object count. We further propose a planner-in-the-loop framework that uses validator and planner diagnostics to revise non-executable specifications through localized edits. Building on this infrastructure, we develop a planner-grounded optimization recipe that combines parameter-efficient Low-Rank Adaptation supervised fine-tuning, offline planner-derived preference pairs for Direct Preference Optimization, and inference-time planner-in-the-loop repair, without requiring online planner calls during training. We also provide a unified evaluation suite for parseability, solvability, specification similarity, and outcome-aware plan-level consistency against planner references. Experiments on representative model families show substantial gains in planner success and plan-level agreement, with improved robustness under difficulty scaling and cross-domain variation. These results highlight the value of externally verifiable formalization for reliable deployment of LLMs in safety- or security-sensitive planning systems. Code and data are available at: https://github.com/ibasicplan/NL-PDDL-Bench

13:00 JSTLLM/生成AI画像/動画生成エージェント

GUICrafter: 注釈のない大量のスクリーンショットを利用する、弱く監視された GUI エージェント

データは現代のインテリジェンスの基本基盤として、現在の基盤モデルの開発を大きく推進してきました。当然のことながら、研究者はこのパラダイムを GUI エージェントの領域に拡張することを目指しており、同様のパラダイムを通じて強力な GUI エージェントを構築することを望んでいます。ただし、GUI エージェントのデータはインターネットから直接収集できないため、コストがかかり、大規模に収集することが困難になります。その結果、現在の GUI エージェントは、デバイス間の汎用化が不十分であり、きめの細かい GUI 要素の視覚的基盤能力が制限されています。 GUI エージェントにおけるデータの課題に対処する試みとして、私たちは GUICrafter を提案します。GUICrafter は、注釈のない大量のスクリーンショットを活用して、高価な人間による注釈への依存を大幅に減らす、弱く監視された GUI エージェントです。 GUICrafter は、2 つの段階的な段階を通じて GUI エージェントをトレーニングするためのカリキュラム学習フレームワークを検討します。まず、モデルは、人間による注釈のない GUI インタラクションに固有の豊富なコンテキスト信号を活用して、大規模な注釈のないスクリーンショットと Web ページから視覚的な基礎を学習します。次に、ステージ 2 では、少量の高品質データを活用して、強化学習によってモデルを調整します。実験の結果、GUICrafter はデータのわずか 0.1% を使用しながら、UI-TARS のような高度なシステムに匹敵する、またはそれ以上のパフォーマンスを達成することが示されています。さらに、同じ量の注釈付きデータの下では、GUICrafter は GUI-R1 などの以前のすべての方法を上回ります。コード、データ、モデルは https://github.com/fansunqi/GUICrafter で入手できます。

原文 (English)

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to the domain of GUI agents, hoping to build strong GUI agents through a similar paradigm. However, GUI agent data cannot be directly harvested from the internet, making it costly and difficult to collect at scale. As a result, current GUI agents suffer from poor cross-device generalization and limited visual grounding ability for fine-grained GUI elements. As an attempt to address data challenge in GUI agents, we propose GUICrafter, a weakly-supervised GUI agent leveraging massive unannotated screenshots to substantially reduce the reliance on expensive human annotations. GUICrafter explores a curriculum learning framework for training GUI agents through two progressive stages. First, the model learns visual grounding from large-scale unannotated screenshots and webpages, leveraging the rich contextual signals inherent in GUI interactions without human annotations. Then, in Stage 2, we leverage a small amount of high-quality data to calibrate the model via reinforcement learning. Experiments show that GUICrafter achieves competitive, or even superior, performance to advanced systems like UI-TARS while using only 0.1% of its data. Furthermore, under the same amount of annotated data, GUICrafter surpasses all previous methods such as GUI-R1. Code, data, and models are available at https://github.com/fansunqi/GUICrafter.

13:00 JSTエージェント

DeepTrans Studio: エージェント翻訳ワークフローで専門家の介入をチームの共有知識に変える

プロの翻訳は多くの場合、チームベースのプロセスです。翻訳者、査読者、プロジェクト マネージャーは、文書全体で用語、法的強制力、説明責任を調整する必要があります。しかし、LLM ベースの翻訳ツールの多くは、人間による修正を個別の編集として扱います。 1 つのセグメントまたは 1 人のメンバーによって行われた専門的な決定が、チームの残りのメンバーにとって再利用可能な知識として取得されることはほとんどありません。 DeepTrans Studio は、専門家がエージェント翻訳ワークフローで選択したノードをインターセプトし、証拠を確認し、AI 出力を修正し、承認された決定を共有チーム メモリに保存できる共同翻訳ワークスペースです。デモ中、参加者は翻訳者や査読者のロールプレイを行い、あらかじめ設定された用語や法的リスクを解決し、自分たちの決定がどのように下流セグメントに伝播され、再利用可能な前例としてチームメイトのワークスペースに表示されるかを確認します。このデモは、AI を介した作業への人間の介入が、1 回限りの修正ではなく、共有され追跡可能な知識になる仕組みを示しています。

原文 (English)

DeepTrans Studio: Turning Expert Interventions into Shared Team Knowledge in Agentic Translation Workflows

Professional translation is often a team-based process: translators, reviewers, and project managers must coordinate terminology, legal force, and accountability across documents. Yet many LLM-based translation tools treat human corrections as isolated edits. Expert decisions made in one segment or by one member are rarely captured as reusable knowledge for the rest of the team. We present DeepTrans Studio, a collaborative translation workspace that lets professionals intercept selected nodes in an agentic translation workflow, review evidence, revise AI outputs, and save approved decisions to a shared team memory. During the demo, attendees will role-play translators and reviewers, resolve preset terminology and legal-modal risks, and see how their decisions are propagated to downstream segments and surfaced in a teammate's workspace as reusable precedents. The demo illustrates how human interventions in AI-mediated work can become shared, traceable knowledge rather than one-off corrections.

13:00 JSTエージェント研究/論文

DEEPMED Search: 内省的な検証を備えた医療の詳細な研究のためのオープンソースのエージェント プラットフォーム

学術文献 (PubMed) から臨床ガイドライン (Web) や民間の知識ベースに至るまで、大量の異種医療データを処理することは、証拠に基づく医療にとって依然として重大なボトルネックとなっています。商用のブラックボックス ツールには透明性が欠けていますが、標準的なオープンソース RAG 実装では、複雑なロングテール クエリを処理するときに推論のドリフトが発生することがよくあります。私たちは、透明性のある医療の深い研究のために設計された完全にオープンソースのエージェント プラットフォームである DEEPMED Search を紹介します。高性能 Next.js アーキテクチャ上に構築された DEEPMED Search は、情報密度に基づいてサブクエリを PubMed、Web 検索、またはローカルのグラフベースのナレッジ ベースに自律的にディスパッチするソース適応型ルーターを備えています。重要なのは、このプラットフォームには、因果関係が一貫したマルチエージェントの議論フレームワークを活用した内省的検証モジュールが統合されており、取得した証拠を合成前に診断ロジックに照らして検証できることです。その堅牢性を実証するために、高難易度の希少疾患クエリを自律的に分解し、交絡ノイズを除去し、構造化された引用に裏付けられた研究レポートを数分で生成する DEEPMED Search の機能を紹介します。このソフトウェアをオープンソース化することで、研究やプロトタイピングの現場で信頼できるガラス箱の医学的推論へのアクセスを民主化するための堅牢なインフラストラクチャをコミュニティに提供します。

原文 (English)

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains a critical bottleneck for evidence-based medicine. While commercial black-box tools lack transparency, standard open-source RAG implementations frequently suffer from reasoning drift when handling complex, long-tail queries. We present DEEPMED Search, a fully open-source, agentic platform designed for transparent medical deep research. Built on a high-performance Next.js architecture, DEEPMED Search features a source-adaptive router that autonomously dispatches sub-queries to PubMed, web search, or local graph-based knowledge bases based on information density. Crucially, the platform integrates an introspective verification module, powered by a causal-consistent multi-agent debate framework, to validate retrieved evidence against diagnostic logic before synthesis. To demonstrate its robustness, we showcase DEEPMED Search's ability to autonomously decompose high-difficulty rare disease queries, filter out confounding noise, and generate structured, citation-backed research reports in minutes. By open-sourcing this software, we provide the community with a robust infrastructure to democratize access to trustworthy, glass-box medical reasoning in research and prototyping settings.

13:00 JSTビジネス/資金調達

グラフ ニューラル ネットワーク モデルに対する生成的再構成攻撃の再考

グラフ データを多くの分野に応用することで、膨大な量のデータを収集して分析する必要性が生じており、その中にはプライベートで機密性の高いデータも含まれています。グラフ データの非ユークリッド的性質により、分析は計算的に困難になり、AI の時代ではグラフ ニューラル ネットワーク (GNN) の使用につながります。 GNN はトレーニングに使用した機密データを誤って漏洩する可能性があり、モデル反転攻撃などの深刻なデータ セキュリティ問題が発生します。この研究では、グラフラベル条件付き (GLC) 攻撃と埋め込みラベル条件付き (ELC) 攻撃という 2 つの新しいグラフ反転 (つまり、再構築) 攻撃を導入することにより、GNN の脆弱性を分析します。それぞれ、ターゲットモデルの予測とその中間表現を利用します。当社は、導入されたプライバシー攻撃の包括的な分析を実行し、3 つのベンチマーク グラフ データセット (つまり、NCI1、PROTEINS、および AIDS) および 4 つのグラフの分布/構造メトリクス (つまり、FGD、EGD、MMD、および GKS) にわたる既存のベースラインと比較します。私たちの研究は、攻撃者がジェネレーター・ディスクリミネーター技術を使用して、GNN に対する現実世界のブラックボックス攻撃シナリオで高品質のグラフを再構築できることを示しています。さらに、クエリを 50% 削減した攻撃の変種 (Ours--) を提示し、良好な、または同等の再構成攻撃パフォーマンスを達成しました。さらに、GNN はラプラシアン ノイズ スケールが変化するプライバシー攻撃に対して非常に脆弱であることを示します。

原文 (English)

Rethinking Generative Reconstruction Attacks against Graph Neural Network Models

The application of graph data in numerous disciplines raises the need for gathering and analyzing huge volumes of data, some of which is private and sensitive. The non-Euclidean nature of the graph data makes the analysis computationally challenging, leading to the use of Graph Neural Networks (GNNs) in the age of AI. GNNs may inadvertently leak sensitive data they are trained on, which raises serious data security issues, including the model inversion attack. In this study, we analyze GNNs' vulnerabilities by introducing two novel graph inversion (i.e., reconstruction) attacks: graph-label conditioned (GLC) attack and embedding-label conditioned (ELC) attack, utilizing targetmodel predictions and their intermediate representations, respectively. We perform a comprehensive analysis of our introduced privacy attacks and compare them with existing baselines across three benchmark graph datasets (i.e., NCI1, PROTEINS, and AIDS) and four graph distributional/structural metrics (i.e., FGD, EGD, MMD, and GKS). Our work demonstrates that an adversary can use the generator-discriminator technique to reconstruct high-quality graphs in real-world black-box attack scenarios against GNNs. Additionally, we present a variant of our attacks (Ours--) with 50% reduced queries, achieving good or comparable reconstruction attack performance. In addition, we show that GNNs are highly vulnerable to privacy attacks, varying Laplacian noise-scales.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

CLQT: LLM ポートフォリオ管理エージェントの診断評価のための、クローズドループでコストを意識した、戦略に一貫したベンチマーク

LLM エージェントは自律的なポートフォリオ マネージャーとしての役割を果たすことが増えており、ベンチマークは財務上の質問応答から逐次取引へと移行しています。しかし、ほとんどのエージェントは依然として、固定ウィンドウでの収益によってエージェントをランク付けしています。これは、期間の収益が市場経路によって支配され、先読み漏れが制御されると見かけのアルファが解消される可能性があるため、弱い代理です。このようなランキングは、健全な推論、一貫した戦略、永続的な優位性を証明するものではありません。 CLQT を紹介します。CLQT は、閉ループ取引の評価をランク付けではなく診断として再構築したもので、エージェントのプロセスが成功または失敗した場所と理由を特定する手段です。 CLQT は、完全にクローズド ループで、コストを意識し、戦略に一貫性があり、一時的にゲートされた環境であり、そのエージェントは収集、合成、割り当て、実行、反映という 5 段階のサイクルを実行します。各ラウンドは、再計算検証可能なハッシュ チェーンに封印された完全な DecisionRound を出力するため、すべてのメトリクスは証跡から再構築可能です。ハード TimeGate、機関取引および資金調達コストのモデリング、戦略一貫性スコアリング、3 層メモリ、Model-Context-Protocol ツール層、および義務を意識した合成という 6 つの柱が基盤を形成します。同じエージェントが、特殊な役割の制約付き委員会または単一の完全自律型オーケストレーターとして実行され、プロセスの足場を実験変数にします。監査証跡から、5 軸の能力スコアカード (APM-CS: Coherence、Acuity、Composure、Discipline、Reliability) を計算します。Coherence は、自己選好バイアスを抑制するために保持されたコホート外の LLM によって部分的に判断されます。アブレーショングリッドを使用した汚染管理されたマルチモデルバックテストと、繰り返し実行されたノイズフロアに対して、目に見えないカットオフ後のデータ上のライブブローカートラックで検証します。 CLQT は結果と能力を分離し、モデルのランキングではなく、エージェントの能力と制限の耐久性と拡張可能なマップを生成します。

原文 (English)

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Yet most still rank agents by returns over a fixed window -- a weak proxy, since a period's return is dominated by the market path and apparent alpha can dissolve once look-ahead leakage is controlled. Such a ranking certifies neither sound reasoning, nor a consistent strategy, nor a durable edge. We introduce CLQT, which reframes closed-loop trading evaluation as diagnosis rather than ranking: an instrument that localizes where and why an agent's process succeeds or fails. CLQT is a fully closed-loop, cost-aware, strategy-consistent, temporally-gated environment whose agents run a five-stage cycle: gather, synthesize, allocate, execute, reflect. Each round emits a complete DecisionRound sealed into a recompute-verifiable hash chain, so every metric is reconstructable from the trail. Six pillars form the substrate: a hard TimeGate, institutional transaction- and financing-cost modeling, strategy-consistency scoring, three-tier memory, a Model-Context-Protocol tool layer, and mandate-aware synthesis. The same agent runs as a constrained committee of specialized roles or a single full-autonomy orchestrator, making process scaffolding an experimental variable. From the audit trail we compute a five-axis capability scorecard (APM-CS: Coherence, Acuity, Composure, Discipline, Reliability), with Coherence judged partly by a held-out, out-of-cohort LLM to curb self-preference bias. We validate it on a contamination-controlled multi-model backtest with an ablation grid and a live broker track on unseen, post-cutoff data, against a repeated-run noise floor. CLQT separates outcome from capability, yielding not a model ranking but a durable, extensible map of agent competencies and limitations.

13:00 JST研究/論文

CRISTAL メソッド: AI 合成世界モデルからの神経象徴分析

このプロジェクトでは、ファンダメンタルズ投資分析を主要なユースケースとして、複雑な分析ワークフローを自動化するための神経記号的フレームワークである CRISTAL メソッド (Coherent Reliable Intentional Synthesis of Truthful Analysis Logic) を紹介します。この領域は、高い構造的不確実性、ノイズの多い主観的なデータ、厳しい注意予算、正当化された再現可能な決定の必要性など、大きな課題を引き起こします。人間のアナリストは、認知バイアスや制限のためにこの領域で苦労することが多く、自動化の大きな価値が示唆されています。しかし、LLM ベースのエージェントは分析補助として提案されていますが、その限界 (貧弱な数値推論、不確実性の認識、再現性の欠如) がこの状況での有効性を妨げています。 CRISTAL は、統計モデルの合成、継続的学習、アクティブ ラーニングを原理的に組み合わせて、これらのギャップに対処します。 CRISTAL は、自然言語の事前知識カリキュラムから開始して、不確実性の定量化や予算を考慮したデータ取得などの完全なベイズ推論を可能にする、動的で解釈可能な確率的プログラムを構築します。 CRISTAL は、コード合成と学習に LLM を活用して、分析中にワールド モデルを継続的に改良します。私たちは、豊富な財務データとテキストデータを備えた合成株式の新しいベンチマークで CRISTAL を検証します。企業分類タスクでは、CRISTAL はわずか 5 つの例と 5 秒の予算でベイズ最適精度を達成し、桁違いに多くの入力データと計算を行っても 40% 程度の精度で頭打ちになる最先端の LLM を上回ります。

原文 (English)

The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models

This project introduces the CRISTAL Method (Coherent Reliable Intentional Synthesis of Truthful Analysis Logic), a neurosymbolic framework for automating complex analysis workflows, with fundamental investment analysis as a primary use case. This domain poses major challenges: high structural uncertainty, noisy and subjective data, tight attention budgets, and the need for justified, reproducible decisions. Human analysts often struggle in this domain due to cognitive biases and limitations, suggesting significant value in automation. But while LLM-based agents have been proposed as analytical aids, their limitations -- poor numerical reasoning, unawareness of uncertainty, and lack of reproducibility -- hinder their effectiveness in this context. CRISTAL addresses these gaps through a principled blend of statistical model synthesis, continuous learning, and active learning. Starting from a natural-language prior knowledge curriculum, CRISTAL builds a dynamic, interpretable probabilistic program that enables full Bayesian inference, including uncertainty quantification and budget-aware data acquisition. CRISTAL continually refines its world model during analysis, leveraging LLMs for code synthesis and learning. We validate CRISTAL on a novel benchmark of synthetic equities with rich financial and textual data. On a company classification task, CRISTAL achieves Bayes-optimal accuracy with just 5 examples and a 5-second budget, outperforming state-of-the-art LLMs that plateau around 40\% accuracy even with order-of-magnitude more input data and compute.

13:00 JST研究/論文

トリプレットの妥当性を超えて: ナレッジ グラフにおける関係セットの完成

ナレッジ グラフ (KG) は、現実世界の知識をトリプレットとして編成し、多くの下流アプリケーションを支えます。本質的に不完全であるため、ナレッジ グラフ補完 (KGC) は広く研究されており、通常はリンク予測が主要なパラダイムであるトリプレット予測として定式化されます。ただし、この定式化はトリプレットごとの情報の不完全性に焦点を当てており、エンティティ関係の互換性情報の不完全性を見落としています。この制限に対処するために、リンク予測タスクを補完し、特定のエンティティと意味的に互換性のある欠落している関係を推論することを目的とした関係セット完了タスク (RSC) を導入します。さらに、観察されたエンティティの関係間の潜在的なパターンをモデル化して、欠落しているものを推測する関係セット埋め込みモデル (RelSetE) を提案します。 RelSetE を評価するために、標準の KG ベンチマークから 3 つのベンチマーク データセットを導出します。広範な実験により、RelSetE がエンティティ関係の互換性パターンを効果的に捕捉し、エンティティの欠落した関係を推論する際に有利に機能することが実証されました。コードとデータは公開されています。

原文 (English)

Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs

Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incompleteness, knowledge graph completion (KGC) is widely studied and is typically formulated as triplet prediction, with link prediction as the dominant paradigm. However, this formulation focuses on the incompleteness of triplet-wise information and overlooks the incompleteness of entity-relation compatibility information. To address this limitation, we introduce a relation set completion task (RSC), which complements the link prediction task and aims to reason about missing relations that are semantically compatible with a given entity. We further propose a Relation Set Embedding model (RelSetE), which models latent patterns among the observed relations of entities to infer missing ones. To evaluate RelSetE, we derive three benchmark datasets from standard KG benchmarks. Extensive experiments demonstrate that RelSetE effectively captures entity-relation compatibility patterns and performs favorably in inferring missing relations of entities. Code and data are publicly available.

13:00 JSTLLM/生成AI

AI トレーニング マネージャー: アダプティブ トレーニング レシピの有界閉ループ制御

適応型機械学習トレーニング用の境界付き LLM ベースの監視コントローラーである AI Training Manager を紹介します。標準的なトレーニング パイプラインは固定レシピや単軸スケジューラに依存することが多く、重度のオーバーフィッティング、損失の不均衡、探査の崩壊、安全でない探査などの実行途中の障害に苦戦する可能性があります。数学的オプティマイザーを置き換えたり、制約のないコーディング エージェントとして機能したりするのではなく、マネージャーはスキーマ条件付きインターフェイスを通じて動作します。アクティブな実行から構造化されたテレメトリ スナップショットを読み取り、制約のあるアクション スペースを監査し、学習率、正則化強度、損失重み係数、探索設定などのトレーニング パラメーターに対する検証済みの更新を返します。このアーキテクチャを教師あり言語モデリングと強化学習にわたって評価します。 TinyStories では、マネージャーが過学習を検出して修正し、ベースラインより 60% 低い検証損失を達成しながら、監査可能な介入ログを生成します。この監視された設定では、マネージャーの推論がトレーニング ループをブロックする必要がないことも示します。マネージャーの応答が保留されている間もトレーニングは続行でき、検証済みの更新が利用可能になったら非同期的に適用できます。ロボット操作の強化学習タスクでは、マネージャーの更新が評価境界またはチェックポイント境界で適用される、エピソード的な閉ループ設定で同じ有界決定インターフェイスを使用します。マネージャーは、保守的で安全でない探査体制の両方を緩和します。これらの結果は、スキーマ条件付き LLM が、解釈可能な多軸介入機能で従来のオプティマイザーやスケジューラーを補完し、ライブ トレーニング実行の限定された監視マネージャーとして機能できることを示唆しています。

原文 (English)

AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes

We present the AI Training Manager, a bounded LLM-based supervisory controller for adaptive machine learning training. Standard training pipelines often rely on fixed recipes or single-axis schedulers, which can struggle with mid-run failures such as severe overfitting, loss imbalance, exploration collapse, or unsafe exploration. Rather than replacing mathematical optimizers or acting as an unconstrained coding agent, the manager operates through a schema-conditioned interface: it reads structured telemetry snapshots from an active run, audits a constrained action space, and returns validated updates to training parameters such as learning rate, regularization strength, loss-weight coefficients, and exploration settings. We evaluate this architecture across supervised language modeling and reinforcement learning. On TinyStories, the manager detects and corrects overfitting, achieving a validation loss 60% lower than the baseline while producing auditable intervention logs. In this supervised setting, we additionally show that manager inference does not need to block the training loop: training can continue while a manager response is pending, and validated updates can be applied asynchronously once available. In a robotic manipulation reinforcement-learning task, we use the same bounded decision interface in an episodic closed-loop setting, where manager updates are applied at evaluation or checkpoint boundaries. The manager mitigates both conservative and unsafe exploration regimes. These results suggest that schema-conditioned LLMs can serve as bounded supervisory managers for live training runs, complementing conventional optimizers and schedulers with interpretable, multi-axis intervention capabilities

13:00 JST研究/論文GPT / ChatGPT

SafePyramid: コンテキスト内のポリシー ガードレールの階層型ベンチマーク

実際のアプリケーションでは、ガードレールは、事前に定義されたリスク分類に依存するのではなく、アプリケーション固有の安全ポリシーに従って、安全でないユーザー モデルの相互作用を識別することが期待されます。この研究では、コンテキスト内で提供されるポリシー仕様に基づいてガードレールが安全違反を予測する、コンテキスト内ポリシー ガードレールのパラダイムの下でこの設定を研究します。この機能を体系的に評価するために、10 のドメインにわたる 1,000 のマルチターン会話と 3,000 の対応するアプリケーション固有のポリシーで構成される安全ベンチマークである SafePyramid を導入します。これらのポリシーには、合わせて 61,699 の異なる自然言語ルールが含まれています。 SafePyramid は評価を 3 つの難易度に分類します。L0 は個々のルールの理解を評価し、L1 はルールの依存関係に関する推論を評価し、L2 はコンテキストで定義された完全な新しいポリシー フレームワークの適応を評価します。ベンチマークの品質を確保するために、当社では厳密な多段階パイプラインを採用してベンチマークを構築および検証しています。 SafePyramid を使用して、10 個のフロンティア LLM と 5 個のポリシー構成可能なガードレールを評価したところ、インコンテキスト ポリシーのガードレールは依然として非常に困難であることがわかりました。最もパフォーマンスの高いモデルである GPT-5.5 でさえ、L0、L1、および L2 でそれぞれ 54.0%、35.3%、および 12.9% のケースで違反ルールの完全なセットを正確に特定するだけです。これらの結果は、現在のガードレールの限界を浮き彫りにし、ポリシーを確実に実行し、ルールの依存関係を解決し、新しいポリシー フレームワークに適応できる、より強力なコンテキスト内ポリシー ガードレールが必要であることを示しています。

原文 (English)

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather than relying on predefined risk taxonomies. In this work, we study this setting under the paradigm of in-context policy guardrailing, where guardrails predict safety violations based on policy specifications provided in context. To systematically evaluate this capability, we introduce SafePyramid, a safety benchmark comprising 1,000 multi-turn conversations across 10 domains and 3,000 corresponding application-specific policies, which together contain 61,699 distinct natural-language rules. SafePyramid organizes the evaluation into three difficulty levels: L0 evaluates individual-rule understanding, L1 evaluates reasoning over rule dependencies, and L2 evaluates adaptation of full novel policy frameworks defined in context. To ensure benchmark quality, we employ a rigorous multi-stage pipeline to construct and validate the benchmark. Using SafePyramid, we evaluate 10 frontier LLMs and 5 policy-configurable guardrails and find that in-context policy guardrailing remains highly challenging: even the best-performing model, GPT-5.5, exactly identifies the full set of violated rules in only 54.0%, 35.3%, and 12.9% cases on L0, L1, and L2, respectively. These results highlight the limitations of current guardrails and call for stronger in-context policy guardrails that can reliably execute policies, resolve rule dependencies, and adapt to novel policy frameworks.

13:00 JSTエージェント

意思決定理論に関する因果モデリングの視点

意思決定理論は、哲学、確率、因果関係からのアイデアを活用して、エージェントが不確実性の下でどのように選択を行うべきかについての正式な枠組みを提供します。大きな進歩にもかかわらず、この分野には依然として統一されたモデリング言語が不足しており、主観的要素と客観的要素の区別や、意思決定理論が適切に機能することが何を意味するかなど、重要な概念が暗黙的に残されていることがよくあります。このため、特に物議を醸している場合には、競合する理論を評価および比較することが困難になる可能性があります。この論文では、因果推論で確立されたツールであるノンパラメトリック構造方程式モデル (NPSEM) に基づく意思決定理論の正式なフレームワークを導入することで、これらの問題に対処します。 NPSEM は、エージェント、反事実、因果関係を表現するための統一基盤を提供し、EDT と CDT の明確な定義を可能にします。この基盤に基づいて、私たちは、エージェントに自らの反事実的有用性の主観的モデルを最大化するよう指示する、新しい意思決定理論である個人的意思決定理論を提案します。我々は、教育や政策を通じて達成される可能性のある特定の意思決定理論を集団全体に強制する仮説的な介入に基づいた正式なパフォーマンス指標を導入し、特定の仮定の下では個人の意思決定理論がこの指標に関して最適であることを示します。全体を通して、喫煙病変の問題を実行例として使用し、最後にニューカムの問題の正式な分析を行います。私たちの目的は、より明確なモデリング言語とより強固な評価基盤を備えた意思決定理論を提供し、それによってより厳密な比較を可能にし、この分野での概念的な進歩を促進することです。

原文 (English)

A causal modeling perspective on decision theory

Decision theory provides a formal framework for how agents should make choices under uncertainty, drawing on ideas from philosophy, probability, and causality. Despite significant progress, the field still lacks a unified modeling language, and key concepts - such as the distinction between subjective and objective elements, or what it means for a decision theory to perform well - are often left implicit. This can make it difficult to evaluate and compare competing theories, particularly in controversial cases. In this paper, we address these issues by introducing a formal framework for decision theory based on nonparametric structural equation models (NPSEMs), a well-established tool in causal inference. NPSEMs provide a unified foundation for representing agents, counterfactuals, and causal relationships, allowing for unambiguous definitions of EDT and CDT. Building on this foundation, we propose a novel decision theory - personal decision theory - which instructs agents to maximize a subjective model of their own counterfactual utility. We introduce a formal performance metric based on hypothetical interventions that enforce a given decision theory across a population - such as might be achieved through education or policy -- and show that, under certain assumptions, personal decision theory is optimal with respect to this metric. Throughout, we use the smoking lesion problem as a running example and conclude with a formal analysis of Newcomb's problem. Our aim is to provide decision theory with a clearer modeling language and firmer evaluative ground, thereby enabling more rigorous comparisons and facilitating conceptual progress in the field.

13:00 JSTLLM/生成AI

HippoSpark: LLM 推論のためのオンデマンド エクスペリエンス システム

将来の問題解決を強化するために、歴史的な軌跡を再利用可能なエクスペリエンスに抽出することが、最近の LLM 研究の焦点となっています。ただし、既存の方法は主にタスク レベルで動作し、類似したタスクが普遍的な解決パターンを共有するという前提の下で、一般的な概要やルールを活用します。このアプローチは、複雑な推論では失敗することが多く、一般に、広範なヒューリスティックではなく、州固有の正確なガイダンスを必要とするローカルなボトルネックでつまづきます。現在の推論状態の差し迫ったニーズに合わせたオンデマンド検索を実行する、状態レベルのエクスペリエンス システムである HippoSpark を紹介します。 HippoSpark は、数学的、科学的、プログラミングのベンチマーク全体で、標準的なプロンプトとタスクレベルのエクスペリエンスのベースラインの両方を常に上回っています。私たちの調査結果から、最も効果的なエクスペリエンス システムとは、一般的なタスク レベルのコンテキストとして機能するのではなく、重大なボトルネックで実用的なガイダンスを提供するシステムであることが明らかになりました。私たちのコードは https://github.com/DanlingMeng/HippoSpark で入手できます。

原文 (English)

HippoSpark: An On-Demand Experience System for LLM Reasoning

Distilling historical trajectories into reusable experience to enhance future problem-solving has become a focal point of recent LLM research. However, existing methods predominantly operate at the task level, leveraging general summaries or rules under the assumption that analogous tasks share universal solution patterns. This approach often fails in complex reasoning, which typically falters at local bottlenecks that require precise, state-specific guidance rather than broad heuristics. We introduce HippoSpark, a state-level experience system that performs on-demand retrieval tailored to the immediate needs of the current reasoning state. Across mathematical, scientific, and programming benchmarks, HippoSpark consistently outperforms both standard prompting and task-level experience baselines. Our findings reveal that the most effective experience systems are those that provide actionable guidance at critical bottlenecks rather than serving as generic task-level context. Our code is available at https://github.com/DanlingMeng/HippoSpark.

13:00 JSTエージェント

SAGA: 長期的な CivRealm 戦略計画のためのシーンを認識し、目標を進化させるエージェント

複雑な戦略ゲームにおける長期的な戦略計画には、不完全な情報とまばらな報酬の下で、複数の意思決定領域にわたる同時推論が必要です。既存の LLM ベースのエージェントは、生のタイル座標によるシーンのブラインドネス、モノリシックな状態ダンプによるコンテキストのオーバーフローとドメインの結合、各エピソードを個別に扱う浅いクロスゲーム学習という 3 つの系統的な障害に悩まされています。我々は、それぞれ 1 つのクラスの障害を直接ターゲットとする 3 つのメカニズムを備えた LLM マルチエージェント フレームワークである SAGA を紹介します。(i) ゲーム エンティティ間の型指定された空間関係をユニットごとの自然言語コンテキストにエンコードするマップ セマンティック シーン グラフ。グローバルなトークン インフレーションを行わずに空間盲目を解決します。 (ii) オンデマンドで詳細なドメイン状態を取得し、専用の専門コントローラーにドメインごとのディレクティブをディスパッチして、コンテキスト オーバーフロー、ドメイン結合、および機械的制約違反を排除するツール拡張プランナー。 (iii) 定期的なゲーム内目標生成と構造化されたゲーム間の因果関係の事後分析を組み合わせたデュアルホライズン フィードバック ループにより、手動による報酬エンジニアリングを行わずに原則に基づいた戦略的進化が可能になります。 FreeCiv で評価された SAGA は、2 つの最も強力なベースラインよりも低い分散で最高の平均文明スコア (環境で唯一のまばらな目標報酬) を達成し、複数の目標の競合下で最も簡単に犠牲になるリソース軸であるインフラストラクチャ建設のすべてのベースラインを大幅に上回る唯一の方法です。これは、ほとんどの対戦ゲームで 2 つの最も強力なベースラインを上回り、出力トークン (主要なデコード コスト) を 27% 削減します。クロスゲーム進化モジュールを搭載した SAGA は、連続する 5 つのエピソードにわたって最高のエンドオブチェーン スコアに達します。アブレーション研究により、各構造コンポーネントが独立してこの利点に貢献していることが確認されています。

原文 (English)

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and sparse reward. Existing LLM-based agents suffer from three systematic failures: scene blindness from raw tile coordinates, context overflow and domain coupling from monolithic state dumps, and shallow cross-game learning that treats each episode in isolation. We present SAGA, an LLM multi-agent framework with three mechanisms each directly targeting one class of failure: (i) a Map-Semantic Scene Graph that encodes typed spatial relations among game entities into per-unit natural-language context, resolving spatial blindness without global token inflation; (ii) a Tool-Augmented Planner that pulls fine-grained domain state on demand and dispatches per-domain directives to dedicated specialist controllers, eliminating context overflow, domain coupling, and mechanical constraint violations; and (iii) a Dual-Horizon Feedback Loop that combines periodic within-game goal generation with structured cross-game causal post-mortem, enabling principled strategic evolution without manual reward engineering. Evaluated on FreeCiv, SAGA attains the highest mean civilization score -- the environment's sole sparse objective reward -- with lower variance than the two strongest baselines, and is the only method that significantly surpasses every baseline on infrastructure construction, the resource axis most readily sacrificed under multi-objective conflict. It outscores the two strongest baselines in most head-to-head games while cutting output tokens (the dominant decoding cost) by 27%. Equipped with the cross-game evolution module, SAGA reaches the highest end-of-chain score across five successive episodes. Ablation studies confirm that each architectural component contributes independently to this advantage.

13:00 JST研究/論文

一次時相論理テンソル ネットワーク

既存の神経記号 AI 手法のほとんどは、オブジェクトが時間次元に従って変化しない静的な知識のシナリオに焦点を当てています。時間的な神経象徴的な作品はまだ研究中であり、主に時間間隔論理または命題線形時相論理のために開発されています。時間の経過とともに特性や関係が変化するオブジェクトを扱う述語を備えた線形時相論理を研究するモデルが不足しています。我々は、線形時間次元を考慮することでこのギャップを埋めるロジック テンソル ネットワーク (LTN) の拡張である一次時間ロジック テンソル ネットワーク (FOT-LTN) を紹介します。特に、FOT-LTN は、一次線形時相論理の構文と LTN のファジー (および実数値) セマンティクスを結合し、時間演算子と量指定子の両方をサポートし、完全に微分可能なフレームワークを取得します。最初の評価は、専用の (純粋なニューラル) メソッドに関して FOT-LTN のパフォーマンスが向上していることを示す 2 つの合成データセットに対する時間知識グラフ完了タスクに関するものです。

原文 (English)

First-Order Temporal Logic Tensor Networks

Most of the existing neuro-symbolic AI methods focus on the scenario of static knowledge where objects do not change according to a temporal dimension. Temporal neuro-symbolic works are still under explored and are mainly developed for time-interval logic or propositional linear temporal logic. There is a lack of models studying linear temporal logics with predicates that deal with objects whose properties and relations change through the time. We present First-Order Temporal Logic Tensor Networks (FOT-LTN) that is an extension of Logic Tensor Networks (LTN) that fills this gap by considering a linear-temporal dimension. In particular, FOT-LTN joins the syntax of First-Order Linear Temporal Logic with the fuzzy (and real-valued) semantics of LTN obtaining a framework that supports both temporal operators and quantifiers and is totally differentiable. A first evaluation regards a temporal knowledge graph completion task on two synthetic datasets showing better performance of FOT-LTN with respect to dedicated (purely neural) methods.

13:00 JSTエージェント

行動基盤モデルによる探索とオンライン転送

強化学習 (RL) におけるゼロショット転送は、転送時に追加の学習を行わずに、報酬のない軌道のみでトレーニングしながら、あらゆる報酬関数に対して最適なポリシーを生成できるエージェントをトレーニングすることを目的としています。タスクに対する一般性から、このようなモデルは「行動基盤モデル」(BFM)と呼ばれることもあります。近年、強力なパフォーマンスと改善が見られていますが、現在のフレームワークとアルゴリズムは、転送フェーズ中にエージェントが状態と報酬のペアのデータセットを通じて報酬 (解決すべきタスク) についてオフラインで通知され、それを使用して展開する最適なポリシーを選択することを前提としています。ただし、実際には、報酬がブラックボックス (ユーザーからの直接のフィードバックなど) である場合、そのようなデータセットを生成することはできません。環境とのインタラクションを通じて報酬を観察する必要があります。言い換えれば、オフライン転送の現在のフレームワークは、報酬を見つけるために探索を必要とする、試行錯誤によるオンライン学習という従来の RL 設定と一致していません。このペーパーでは、BFM 自体を探索ポリシーの生成に使用できるという重要な洞察を基に、ゼロショット RL でこの新しいオンライン転送に取り組むことを提案します。私たちは、このオンライン学習の問題を盗賊のような探索と搾取の問題の観点から組み立てることが可能であることを示します。より正確には、各ステップでバンディット アルゴリズムがポリシーを推奨し、BFM がそれを環境内で実行し、報酬と新しい状態を生成します。最適なポリシーに収束するまでこのプロセスを繰り返します。線形報酬近似の一般的なコンテキストで、上限信頼限界にヒントを得た定式化を導き出し、不確実性行列の固有値の最小化を通じて探索が達成できることを示します。手法の概念を検証するために、単純な環境でフレームワークを定性的および定量的に評価します。

原文 (English)

Exploration and Online Transfer with Behavioral Foundation Models

Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories. For their generality over tasks, such models are sometimes called ``Behavioral Foundation Models'' (BFMs). While they have shown strong performances and improvements in recent years, the current framework and algorithms still assume that, during the transfer phase, the agent is informed offline about the reward (the task to solve) through a dataset of state-reward pairs, which it uses to pick the best policy to deploy. However, in practice if the reward is a black-box (e.g. direct user feedback), it is not possible to generate such a dataset: it is necessary to observe the reward through interactions with the environment. In other words, the current framework of offline transfer is not aligned with the traditional RL setting of online learning through trial-and-error, which requires exploration in order to find rewards. This paper proposes to tackle this new online transfer in zero-shot RL, with the key insight that the BFM itself can be used to generate exploration policies. We show that it is possible to frame this online learning problem in terms of a bandit-like exploration-exploitation problem. More precisely, at each step the bandit algorithm recommends a policy, the BFM executes it in the environment, which yields a reward and a new state; we repeat the process until we converge to the optimal policy. In the popular context of linear reward approximation, we derive a formulation inspired by Upper Confidence Bound and show that exploration can be achieved through the minimization of the eigenvalues of an uncertainty matrix. We evaluate qualitatively and quantitatively our framework on a simple environment to validate the concept of our method.

13:00 JST研究/論文

応答するときは忠実であること: 視覚言語モデルの強化学習に対して流暢で根拠のある回答を返す

強化学習 (RL) は、視覚言語モデル (VLM) の推論能力を向上させるための重要なパラダイムです。ただし、マルチモーダル推論のロールアウトに RL を直接適用すると、言語事前条件の悪用、視覚的証拠の無視、流暢ではあるが視覚的に根拠のない推論トレースの生成により、不安定性が生じる可能性があります。疑問が生じます。強化学習を適用する前に、最初に視覚的に忠実な推論体制にポリシーを向けることができるでしょうか?この目的を達成するために、私たちは、まず 6 つの一般的な VQA ベンチマークから明示的な視覚と言語の因果関係を持つサンプルを厳選して FaithfulQA データセットを構築する、Faithful Warm-Start (FWS) 戦略を提案します。この戦略では、画像と質問のペアのそれぞれが、ある程度の視覚的観察、質問要件、常識知識、領域知識、および最終的な回答を獲得します。その後、VLM ベースのジャッジを採用してデータセットをさらに精製し、強い因果関係の一貫性と視覚的な忠実性を確保します。このウォームスタート段階では、疎な回答レベルの報酬の下でその後の RL 最適化を行う前に、因果関係に基づいた視覚言語パターンを理解する機能をモデルに備えさせます。実験結果は、このような忠実な監督により、回答の精度が向上し、RL トレーニングが安定し、視覚的に裏付けのない推論が減少することを示しています。

原文 (English)

Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning

Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal reasoning can lead to instability, due to the exploitation of language priors, the neglect of visual evidence, and the generation of reasoning traces that are fluent yet not visually grounded. The question arises: Can initially steer the policy toward visually faithful reasoning regime before applying reinforcement learning? To this end, we propose a Faithful Warm-Start (FWS) strategy that first curates samples with explicit vision-language causal relationships from six general VQA benchmarks to construct the FaithfulQA dataset, where each of the image-question pairs gains a certain degree of visual observations, question requirements, commonsense knowledge, domain knowledge, and the final answer. Subsequently, a VLM-based judge is employed to further purify the dataset, ensuring strong causal consistency and visual faithfulness. This warm-start stage equips the model with the capability to understand causally grounded vision-language patterns before subsequent RL optimization under sparse answer-level rewards. Experimental results show that such faithful supervision improves answer accuracy, stabilizes RL training, and reduces visually unsupported reasoning.

13:00 JST研究/論文

AlgoSkill: 人間のようなスキルをスケジュールしてアルゴリズムの設計を学習する

自然言語の問題ステートメントからアルゴリズムを設計するには、問題の構造を特定し、制約を読み取り、適切なパラダイムを選択し、正しさを確認し、複雑さを調整する必要があります。既存の大規模言語モデル (LLM) メソッドは、多くの場合、直接生成または汎用自己洗練に依存しており、これらのステップは暗黙的なままになっています。私たちは、抽象化、制約分析、状態設計、データ構造の選択、証明チェック、反例の構築、複雑さの改善など、アルゴリズム スキルの型付きライブラリに対する逐次的な意思決定としてアルゴリズム設計をモデル化する AlgoSkill を提案します。学習されたスケジューラーは現在の設計状態からスキルを提案し、モンテカルロ ツリー検索 (MCTS) コントローラーはコンパイル、テスト、ストレス テスト、および複雑さの分析からの検証フィードバックを使用してスキル シーケンスを探索します。競技プログラミングと組み合わせ最適化のベンチマークに関する実験では、型指定されたスキルを使用しない場合でも、AlgoSkill が LLM の直接生成、思考連鎖のプロンプト、自己洗練、および MCTS よりも向上することが示されています。アブレーションは、型指定されたスキル、検証ベースの修復、および検索ベースのスケジューリングがそれぞれパフォーマンスに寄与していることを示しています。これらの結果は、自動アルゴリズム設計を、ワンショット コード生成ではなく、検証に基づいたスキル スケジューリングとして扱うことを裏付けています。

原文 (English)

AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills

Designing an algorithm from a natural-language problem statement requires identifying the problem structure, reading constraints, choosing a suitable paradigm, checking correctness, and refining complexity. Existing large language model (LLM) methods often rely on direct generation or generic self-refinement, leaving these steps implicit. We propose AlgoSkill, which models algorithm design as sequential decision-making over a typed library of algorithmic skills, including abstraction, constraint analysis, state design, data-structure selection, proof checking, counterexample construction, and complexity refinement. A learned scheduler proposes skills from the current design state, while a Monte Carlo Tree Search (MCTS) controller explores skill sequences using verification feedback from compilation, testing, stress testing, and complexity analysis. Experiments on competitive programming and combinatorial optimization benchmarks show that AlgoSkill improves over direct LLM generation, chain-of-thought prompting, self-refinement, and MCTS without typed skills. Ablations show that typed skills, verification-based repair, and search-based scheduling each contribute to performance. These results support treating automatic algorithm design as verification-guided skill scheduling rather than one-shot code generation.

13:00 JSTエージェント

ACPO: マルチエージェント強化学習のためのエージェント連鎖ポリシーの最適化

マルチエージェント強化学習 (MARL) における協調タスクでは、エージェントが共同して共有利益を最大化する必要があります。集中トレーニングと分散実行 (CTDE) パラダイムの下では、ポリシー勾配を直接計算することは依然として困難です。従来の手法は主に 2 つのアプローチに従います。1 つは集中的な批評家による独立した因数分解更新であり、値分解の仮定を持たない一般的な共同改善の保証が欠けています。もう 1 つは、次善のナッシュ均衡に収束する可能性がある交互の最良応答更新です。この論文では、共同ポリシー勾配が、エージェントごとのスコア関数と分散型批評家から形成されるエージェントごとの用語の正確な分散型分解を可能にすることを示します。この分解に基づいて、私たちはエージェント連鎖ポリシー最適化 (ACPO) を開発します。ACPO では、アクターが個別にトレーニングされ、更新内容が共同ポリシー勾配の 1 つのステップを構成します。この結果の中心となるのは、エージェントが一度に 1 つずつアクションを実行し、各エージェントが先行するアクションに対する信念に基づいてアクションを実行する、同時共同決定のシリアル化されたビューです。この信念は、エージェントごとの独立した更新を共同勾配ステップに結び付ける調整メカニズムとして機能します。マルチロボット ウェアハウス、SMACv2、および MA-MuJoCo で ACPO を評価しました。ACPO は強力なベースラインを上回り、エージェントの数が増加するにつれてギャップが拡大しました。

原文 (English)

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning

Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained difficult to compute directly. Prior methods largely follow two approaches: independent factorized updates with centralized critics, which lack general joint-improvement guarantees without value decomposition assumptions, or alternating best-response updates, which can converge to suboptimal Nash Equilibria. In this paper, we show the joint policy gradient admits an exact decentralized decomposition of per-agent terms, each formed from per-agent score functions and decentralized critics. Based on this decomposition, we develop Agent-Chained Policy Optimization (ACPO), where actors are trained independently, with their updates together constituting a single step on the joint policy gradient. Central to this result is a serialized view of the simultaneous joint decision in which agents commit actions one at a time, each conditioning on a belief over preceding actions. The belief acts as the coordination mechanism which ties the independent per-agent updates into a joint gradient step. We evaluate ACPO on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo, where it outperforms strong baselines, with the gap widening as the number of agents grows.

13:00 JST研究/論文

SAT-RTS: リアルタイム ストラテジー ゲームにおける戦術知識の抽出と視覚化ベースの分析のための体系的なフレームワーク

リアルタイム ストラテジー (RTS) ゲームのマイクロマネジメントにおける効率的な戦術知識の抽出と分析は、高次元で結合された状態アクションの逐次データとブラックボックスの意思決定プロセスによって制約されます。現在の研究では、データの分離と抽象化の観点から、階層的な視覚化に基づく属性分析を提供することはほとんどありません。 RTS ゲームにおける解釈可能な戦術知識の抽出と視覚化ベースの分析を容易にするために、状態-アクション-戦術分析パイプライン (SAT-RTS) と呼ばれる体系的なフレームワークが提案されています。 RTS 学習システムにおける重要な意思決定の根深い要因を解読するために、この研究では、解釈可能な視覚化と、高次元シーケンス データからの潜在的な戦術パターンの自動抽出を統合しています。クラスター中心の BK ツリー アルゴリズムを採用し、複数の側面の類似性を定量化するように設計された特殊な距離メトリックを組み込むことにより、提案されたフレームワークは堅牢な状態ストリームの抽象化を促進します。さらに、非構造化状態行動シーケンスを離散的で解釈可能な戦術ラベルに変換するルールベースのマルチラベル抽出手法が開発され、生の行動データと高レベルの戦術的洞察の間のギャップを効果的に埋めることができます。これらの計算手法を階層的な視覚化ベースのパイプラインに総合的に統合することで、提案されたフレームワークは、大規模なリアルタイム データ ストリームの処理の課題に効果的に対処し、同時に根深い戦術的要因を解読するためのフィットネス ランドスケープの視覚化と分析的洞察を提供します。包括的な実験により、提案された SAT-RTS が複雑な RTS 環境における戦術分析の解釈可能性と効率を大幅に向上させることが実証されました。

原文 (English)

SAT-RTS: A systematic framework for tactical knowledge extraction and visualization-based analysis in real-time strategy games

Efficient tactical knowledge extraction and analysis in real-time strategy (RTS) games micromanagement are constrained by the high-dimensional coupled state-action sequential data and the black-box decision-making process. Current research rarely provides a hierarchical visualization-based attribution analysis from the perspective of data decoupling and abstraction. To facilitate interpretable tactical knowledge extraction and visualization-based analysis in RTS games, a systematic framework named state-action-tactic analysis pipeline (SAT-RTS) is proposed. To decipher the deep-seated drivers of critical decisions in RTS learning systems, this work integrates interpretable visualization with the automated extraction of latent tactical patterns from high-dimensional sequence data. By adapting a cluster-centric BK-tree algorithm and incorporating specialized distance metrics designed to quantify multi-aspect similarities, the proposed framework facilitates robust state-stream abstraction. Furthermore, a rule-based multi-label extraction method is developed to transform unstructured state-action sequences into discrete and interpretable tactical labels, effectively bridging the gap between raw behavioral data and high-level tactical insights. By holistically integrating these computational methods into a hierarchical visualization-based pipeline, the proposed framework effectively addresses the challenges of processing massive real-time data streams while providing fitness landscape visualizations and analytical insights to decipher deep-seated tactical drivers. Comprehensive experiments demonstrate that the proposed SAT-RTS significantly enhances the interpretability and efficiency of tactical analysis in complex RTS environments.

13:00 JST研究/論文

影響マップとクラスターベースのスクリプトを使用した StarCraft マイクロマネジメントの階層強化学習

リアルタイム ストラテジー (RTS) ゲームは、AI に重大な課題をもたらします。これは、継続的な戦場での複数ユニットの連携から生じる広大な状態アクション空間と、最終的な勝敗シグナルから生じる遅延報酬のまばらさによって特徴付けられます。既存のアプローチは、共同動作の次元爆発の管理と複雑な状態表現の解釈可能性の維持との間のトレードオフに直面しています。この複雑さは、タスクを効果的な戦術モジュールに適応的に分解する際の従来の階層構造の制限によってさらに強化されます。このような問題は、深層学習モデルのブラックボックス的な性質と、まばらな報酬への依存によってさらに悪化し、結果としてサンプル効率が制限され、意思決定の透明性が欠如します。これらの制限に対処するために、この論文では、StarCraft マイクロマネジメント用のインフルエンス マップ ハッシュとクラスターベースのスクリプトを備えた階層型強化学習フレームワークである HRL-IM/CBS を提案します。影響力マップのハッシュ化は、世界規模の戦場の状況をコンパクトな 16 進コードにエンコードし、空間制御と相対的優位性を捉えます。クラスターベースのスクリプトにより、適応型ユニット分割を通じて動的なローカル調整が可能になります。階層的なマルチ Q テーブル アーキテクチャは、意思決定を上位レベルのクラスタリング戦略の選択と下位レベルの戦術の実行に分解し、報酬の割り当てにより密度の高い学習シグナルを提供します。 6 つの非対称シナリオにわたる実験は、透明な Q テーブル表現を通じてサンプル効率と解釈可能性における利点を提供しながら、深い RL ベースラインに対して競合するパフォーマンスを実証します。

原文 (English)

Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts

Real-time strategy (RTS) games present significant AI challenges, characterized by expansive state-action spaces arising from multi-unit coordination in continuous battlefields, and sparse delayed rewards stemming from final win/lose signals. Existing approaches face a trade-off between managing the dimensionality explosion of joint actions and maintaining the interpretability of complex state representations. This complexity is further intensified by the limitation of traditional hierarchical structures in adaptively decomposing tasks into effective tactical modules. Such difficulties are compounded by the black-box nature of deep learning models and their reliance on sparse rewards, which together result in limited sample efficiency and a lack of decision-making transparency. To address these limitations, this paper proposes HRL-IM/CBS, a hierarchical reinforcement learning framework with influence map hashing and cluster-based scripts for StarCraft micromanagement. Influence map hashing encodes global battlefield situations into compact hexadecimal codes, capturing spatial control and relative advantage. Cluster-based scripts enable dynamic local coordination through adaptive unit partitioning. The hierarchical multi-Q-table architecture decomposes decision-making into upper-level clustering strategy selection and lower-level tactical execution, with reward allocation providing dense learning signals. Experiments across six asymmetric scenarios demonstrate competitive performance against deep RL baselines while offering advantages in sample efficiency and interpretability through transparent Q-table representations.

13:00 JST研究/論文

EEG基礎モデルの時間的特徴抽出器: 事前訓練された時系列モデルを含む制御された比較

脳波 (EEG) 基礎モデルは、大規模な脳記録から一般化可能な表現を学習することを目的としています。ただし、時間特徴抽出器の役割と、事前トレーニングされた時系列基礎モデル (TSFM) をこの設定に効果的に移行できるかどうかは、まだ研究されていません。私たちは、統合されたEEG基盤モデル内で、線形ベースライン、畳み込みエンコーダー、凍結事前学習TSFM(MOMENT)を含む3つの時間的特徴抽出戦略の制御された比較を実行します。私たちは、運動イメージと感情認識という 2 つの下流タスクを使用して、表現の品質に対するそれらの影響を評価します。結果は、評価されたベンチマーク全体で異なる傾向を明らかにします。運動画像データセットでは、単純な時間表現が競合的に機能しますが、感情データセットはより豊富な時間モデリングの恩恵を受けます。 EEG には特に適応されていませんが、事前学習済み TSFM は効果的な時間特徴抽出器として機能し、汎用の時系列表現を EEG 基礎モデル内の凍結時間特徴抽出器として転送できることを示唆しています。

原文 (English)

Temporal Feature Extractors in EEG Foundation Models: A Controlled Comparison Including a Pretrained Time-Series Model

Electroencephalography (EEG) foundation models aim to learn generalizable representations from large-scale brain recordings. However, the role of temporal feature extractors and whether pretrained time-series foundation models (TSFMs) can be effectively transferred to this setting remains underexplored. We conduct a controlled comparison of three temporal feature extraction strategies, including a linear baseline, a convolutional encoder, and a frozen pretrained TSFM (MOMENT), within a unified EEG foundation model. We evaluate their impact on representation quality using two downstream tasks: motor imagery and emotion recognition. Results reveal different trends across the evaluated benchmarks. On the motor imagery dataset, simple temporal representations perform competitively, whereas the emotion dataset benefits from richer temporal modeling. Although not specifically adapted to EEG, the pretrained TSFM serves as an effective temporal feature extractor, suggesting that general-purpose time-series representations can be transferred as frozen temporal feature extractors within EEG foundation models.

13:00 JST研究/論文

~ニューラルネットワーク検証のための~区間信念構造と~不正確なコピュラの伝播

ニューラル ネットワークの定量的検証には、入力分布とその依存構造の両方における実質的な不確実性の下で確率について推論する必要があります。現実的な設定では、この情報は部分的にのみ指定されることが多く、正確な確率モデルを仮定すると信頼性の低い結果が生じる可能性があります。我々は、限界不確実性を表す区間信念構造と、不確実な依存性をモデル化する不正確なコピュラを組み合わせた、不正確な確率情報の下での定量的検証のための健全なフレームワークを提案します。フィードフォワードニューラルネットワークを介して、不正確に結合された区間信念構造の伝播方法を開発します。混合された不正確なコピュラ ボリュームを使用して、アフィン変換と活性化関数を通じてサウンドを前に押し出す構造を導き出します。結果として得られる出力は、指定された不正確な入力と互換性のあるすべての確率モデルに有効な、確率的安全性の保証された下限と上限を提供できます。

原文 (English)

Propagation of~Interval Belief Structures and~Imprecise Copulas for~Neural Network Verification

Quantitative verification of neural networks requires reasoning about probabilities under substantial uncertainty in both input distributions and their dependence structure. In realistic settings, this information is often only partially specified, and assuming precise probabilistic models can lead to unreliable results. We propose a sound framework for quantitative verification under imprecise probabilistic information, combining interval belief structures to represent marginal uncertainty with imprecise copulas to model uncertain dependence. We develop a propagation method for imprecisely coupled interval belief structures through feed-forward neural networks. Using mixed imprecise copula volumes, we derive sound push-forward constructions through affine transformations and activation functions. The resulting output can provide guaranteed lower and upper bounds on probabilistic safety properties, valid for all probability models compatible with the specified imprecise inputs.

13:00 JST研究/論文

言語モデルを使用した信頼性の高い物理設計のための構造認証

信頼性の低い言語モデルは、アサートする権限がモデルの外に移動された場合、信頼性の高い物理設計を生成することができます。つまり、モデルが提案し、決定論的エンジンのみが認証し、認証済み、不可能、または不明を返します。私たちは、5 つの科学分野にわたる提案と認証のループである Physics-Anchored Certification (PHACT) を導入し、そのような証明書が信頼できるものである理由を特定します。モデルが提供する値を受け入れるチェッカーは偽造できます。代わりに、固定入力から証明された数量を導き出すことで、構造上偽造が不可能になります。 2 つのモデル、2 つのデコード温度、および意図的に障害が発生したエンジンにわたる 80 回の敵対的トライアルにわたって、この契約では虚偽の認証はゼロでした。

原文 (English)

Structural Certification for Reliable Physical Design with Language Models

An unreliable language model can be made to produce reliable physical designs if the authority to assert is moved out of the model: the model proposes, and a deterministic engine alone certifies, returning certified, impossible, or unknown. We introduce Physics-Anchored Certification (PHACT), a propose-certify loop spanning five scientific domains, and identify what makes such a certificate trustworthy. A checker that accepts a model-supplied value can be forged; deriving the certified quantity from fixed inputs instead makes forgery impossible by construction. Across eighty adversarial trials spanning two models, two decoding temperatures, and a deliberately faulted engine, this contract produced zero false certifications.

13:00 JST研究/論文

憲法優先の再構築における未解決の問題

ペアワイズ嗜好データは、言語モデル (RLHF など) のトレーニングと評価に広く使用されていますが、各データポイントは、その背後にある理論的根拠ではなく \emph{choice} を記録します。逆憲法 AI (ICAI) などの手法は、データセットを自然言語原則の短い「構成」に圧縮することで解釈可能性を向上させようとします。私たちは、この枠組みは仕様が不十分であると主張します。原則のフラットなリストは、原則の構成が暗黙的なままであるため、まだ実行可能な決定ルールではありません。我々は、構成手法における 3 つの未解決の問題を経験的に特徴付けるためのテストベッドとしてペアワイズ設定を使用します。まず、原則的な品質は測定が困難です。カバレッジと精度は役に立ちますが、エンドツーエンドの再構築には不完全です。第二に、\emph{構成があいまい}: 原則は固定されており、さまざまな執行者 (LLM 裁判官と多数決) が同意するのは $73\%$ の確率だけです。 3 番目に、 \emph{LLM 間で構成が異なります}: モデル間の投票合意は $73\%$ ですが、モデル内投票合意は $81\%$ です。 PRISM、AlpacaEval、Chatbot Arena 全体で、原則の改良 (ICAI+) がこれらの問題を改善するための第一歩となる可能性があることを示します。執行者間の合意は $78\%$ に上昇し、透明性のある執行者は LLM ジャッジの精度と一致します ($66\%$ 対 \ $67\%$)。私たちの結果は、憲法が \emph{憲法--執行者システム} として評価されるべきであり、裁判官としての LLM に広く影響を与えることを強調しています。

原文 (English)

Open Problems in Constitutional Preference Reconstruction

Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice}, not the rationale behind it. Methods such as Inverse Constitutional AI (ICAI) attempt to improve interpretability by compressing datasets into short ``constitutions'' of natural-language principles. We argue this framing is under-specified: a flat list of principles is not yet an executable decision rule because it leaves principle composition implicit. We use the pairwise setting as a testbed to empirically characterize three open problems in constitutional methods. First, principle quality is hard to measure: coverage and accuracy are useful but incomplete proxies for end-to-end reconstruction. Second, \emph{composition is ambiguous}: holding principles fixed, different executors (LLM judge versus majority vote) agree only $73\%$ of the time. Third, \emph{constitutions differ between LLMs}: cross-model vote agreement is $73\%$, whereas intra-model agreement is $81\%$. Across PRISM, AlpacaEval, and Chatbot Arena, we show that principle refinement (ICAI+) may be a first step towards ameliorating these problems: inter-executor agreement rises to $78\%$, and transparent executors match LLM judge accuracy ($66\%$ vs.\ $67\%$). Our results highlight that constitutions should be evaluated as \emph{constitution--executor systems}, with implications for LLMs-as-a-judge broadly.

13:00 JSTLLM/生成AI

冗長な思考連鎖は本当に役立つのか?長さではなく内容が重要であるという配布中の証拠

思考連鎖 (CoT) プロンプトは LLM 推論を改善しますが、ソースについては議論があります。中間ステップが役立つのは、中間ステップが有用なセマンティック コンテンツを運ぶからでしょうか、それとも、モデルが答えをコミットする前に、より多くのトークンで条件付けすることで追加の計算が行われるからでしょうか?私たちは 2 つの証拠を提示します。まず、分布において、同じ質問について各モデルを繰り返しサンプリングし、同じ推論計画に従う独自の自然世代の短いものと長いものをペアにするため、何も書き換えられることはなく、両方のトレースが真に分布内にあります。 25 のモデルにわたって、追加のトークンの精度は、独立してトレーニングされたすべての推論器で本質的に変化しません。また、余剰トークンのブラインド分析により、他の場所にどのような利益が存在するかが冗長性そのものではなく、検証とチェックの内容を追跡することが示されています。次に、制御された介入として、番号編集された完了と階層化されたブートストラップ信頼区間を備えた 4 つのターゲットと 8 つのベンチマークにわたるデュアルバリデーター設計を使用して、同じ意味内容 (有向非巡回グラフ等価性によって検証された同じ事実、操作、および中間値) を表す 2 つのトレースが、一方がより冗長な場合に異なる結果を生み出すかどうかを尋ねます。冗長なトレースは精度を向上させます (32 個のベンチマーク ターゲット セルのうち 25 個が、少なくとも 1 つのバリデータの下で陽性です) が、効果は控えめで (通常 1 ~ 4 ポイント)、単に長さだけでなく冗長な散文の品質に依存します。最大の数値編集では、効果は増幅され (4 つの算術ベンチマークの中央値 3.24 倍)、長さが一致する非推論フィラーでは効果はまったく回復されません。両方の線は収束します。重要なのは、追加のトークンが何をするか (推論と検証のコンテンツ) であり、その数ではありません。この図は、純粋なフォワード パス コンピューティングでも、純粋なセマンティック コンテンツ アカウントでも十分に説明できません。

原文 (English)

Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters

Chain-of-thought (CoT) prompting improves LLM reasoning, but the source is contested: do the intermediate steps help because they carry useful semantic content, or because conditioning on more tokens buys extra computation before the model commits to an answer? We bring two lines of evidence to bear. First, in distribution: we repeatedly sample each model on the same question and pair a shorter with a longer of its own natural generations that follow the same reasoning plan, so nothing is rewritten and both traces are genuinely in-distribution. Across 25 models the extra tokens leave accuracy essentially unchanged for every independently-trained reasoner, and a blind analysis of the surplus tokens shows that what gain exists elsewhere tracks validation- and checking-content, not verbosity per se. Second, as a controlled intervention, we ask whether two traces expressing the same semantic content (the same facts, operations, and intermediate values, verified through directed acyclic graph equivalence) produce different outcomes when one is more verbose, using a dual-validator design across four targets and eight benchmarks with number-redacted completion and stratified bootstrap confidence intervals. Verbose traces do improve accuracy (25 of 32 benchmark-target cells are positive under at least one validator), but the effects are modest (typically 1-4 points) and depend on the quality of the verbose prose, not merely its length. Under maximum numerical redaction the effect is amplified (median 3.24x across four arithmetic benchmarks), and length-matched non-reasoning filler recovers none of it. Both lines converge: what matters is what the extra tokens do (the reasoning and validation content they carry), not how many there are, a picture neither a pure forward-pass-compute nor a pure semantic-content account fully explains.

13:00 JST研究/論文

関連性は許可されていません: 価値のある貢献には注意が必要です

関連性は許可ではありません。アテンションを使用すると、モデルは現在のクエリに関連するキーと値の項目を読み取ることができますが、そのような項目の値の寄与が予測の証拠になることは保証されません。検索された文章は、裏付けとなる証拠がなくても質問に関連している可能性があり、歴史的事実や時間的近傍によって、トゥルーテール ランキングや現在のエッジ スコアが曖昧になる場合もあります。この論文では、このギャップを、実際に予測パスに追加される加重値項 alpha_ij * v_j の許可問題として形式化します。私たちは、注意関連性 alpha_ij を保持し、プライマリ メトリックにつながる値パスを公開し、完全なモデルでは、学習されたクエリ項目権限 g_ij を通じて alpha_ij * v_j を alpha_ij * g_ij * v_j に変換するパスローカライズされたインターフェイスである Warrant を提案します。 CTDG リンク予測、MTPP 次マーク ランキング、RAG サポート証拠選択、STPP 次位置予測、および TKG テール予測のメトリクスを定義する値パスに同じ演算子を配置します。 32 のペア比較、3 つのシード、合計 192 回の実行にわたって、Warrant は 27 の比較で主要な指標を改善しました。実用的な段階は、10 個の実質的な効果、1 個の限界効果、8 個のプラスだが不確実な効果、8 個の同程度/無視できる効果、および 5 個のドロップで構成されます。パス ローカライゼーション チェックでは、正しいパス配置は、すべてのドメインで方向認識ベース パフォーマンスを上回り、一般的なアテンション配置を CTDG で +0.1076 AUC、TKG で +0.0683 MRR 上回ります。アブレーションの結果、ほとんどの TKG ゲインはヒストリカルテール値パスのエクスポージャから得られるのに対し、コア CTDG ゲインはエッジ条件付きクエリ項目のアクセス許可から得られることがわかります。結論として、予測証拠は注目マスではありません。重み付けされた値の項は、メトリックへのパス上で保証される場合にのみ証拠になります。

原文 (English)

Relevance Is Not Permission: Warranted Attention for Value Contributions

Relevance is not permission. Attention lets a model read key-value items related to the current query, but it does not guarantee that the value contribution of such an item becomes prediction evidence. A retrieved passage may be relevant to a question without being supporting evidence, and a historical fact or temporal neighbor may even blur true-tail ranking or the current edge score. This paper formalizes this gap as a permission problem for the weighted value term alpha_ij * v_j that is actually added to the prediction path. We propose Warrant, a path-localized interface that preserves attention relevance alpha_ij, exposes the value path leading to the primary metric, and, in the full model, turns alpha_ij * v_j into alpha_ij * g_ij * v_j through learned query-item permission g_ij. We place the same operator on the metric-defining value paths of CTDG link prediction, MTPP next-mark ranking, RAG supporting evidence selection, STPP next-location forecasting, and TKG tail prediction. Across 32 paired comparisons, 3 seeds, and 192 total runs, Warrant improves the primary metric in 27 comparisons; practical tiers consist of 10 substantial effects, 1 marginal effect, 8 positive but uncertain effects, 8 tie/negligible effects, and 5 drops. In the path-localization check, correct-path placement outperforms direction-aware Base performance in every domain and exceeds generic attention placement by +0.1076 AUC in CTDG and +0.0683 MRR in TKG. Ablations show that most TKG gains come from historical-tail value path exposure, whereas the core CTDG gain comes from edge-conditioned query-item permission. In conclusion, prediction evidence is not attention mass. A weighted value term becomes evidence only when it is warranted on the path to the metric.

13:00 JST画像/動画生成

FacePlex: 会話アバター向けの全二重共同音声と顔の動作の生成

自然な対面会話には、同期した顔の動きとともにリアルタイムの音声生成が必要です。既存のシステムはこの問題に部分的にしか対処していません。音声のみの全二重モデルはリアルタイムで音声を生成できますが、顔の動きは生成しません。一方、オーディオ駆動型の顔の動きモデルは、音声と動きをオンラインで共同生成するのではなく、すでに利用可能なオーディオから顔をアニメーション化します。このギャップを埋めるために、まず全二重の音声と顔のモーションの共同生成を形式化し、音声トークンと顔のモーション トークンがステップごとに一緒に生成されます。この定式化に基づいて、2 つの主要なコンポーネントを備えた統合ストリーミング フレームワークである FacePlex を提案します。まず、ローリング フロー マッチングは、各ストリーミング ステップで新しいモーション フレームをコミットすることにより、フロー マッチングをオンライン モーション生成に適応させます。 2 番目に、ローリング クロス アテンションはストリーミング オーディオ キューとモーション キューを結合し、生成が進むにつれて音声と顔のモーションが相互に調整できるようにします。広範な実験、アブレーション研究、ユーザー研究を通じて、FacePlex がオンライン ストリーミングの制約下で全二重の音声と顔のモーションの共同生成を可能にし、同時にオーディオ駆動の顔のモーション ベースラインよりも強力なリップシンク品質とモーションの忠実度を達成できることを示しました。

原文 (English)

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only partially address this problem: speech-only full-duplex models can generate speech in real time but do not produce facial motion, while audio-driven facial motion models animate a face from already available audio rather than jointly generating speech and motion online. To bridge this gap, we first formalize full-duplex joint speech-facial motion generation, where speech tokens and facial motion tokens are produced together every step. Building on this formulation, we propose FacePlex, a unified streaming framework with two key components. First, Rolling Flow Matching adapts flow matching to online motion generation by committing new motion frames at each streaming step. Second, Rolling Cross-Attention couples the streaming audio queue with the motion queue, allowing speech and facial motion to condition each other as generation progresses. Through extensive experiments, ablation studies, and a user study, we show that FacePlex enables full-duplex joint speech-facial motion generation under online streaming constraints, while achieving stronger lip-sync quality and motion fidelity than audio-driven facial motion baselines.

13:00 JSTエージェント研究/論文

MirrorCode: AI は動作だけからプログラム全体を再構築できる

ベンチマークの進捗状況や、AI が C コンパイラーを実装するなどの 1 回限りのデモンストレーションで示されているように、AI モデルは自律コーディングにおいて急速に向上しています。ただし、既存のコーディング ベンチマークは短いタスクに焦点を当てる傾向があり、1 回限りのデモンストレーションは人間によるガイダンスが含まれることが多く、標準化されておらず、モデル間で繰り返されていないため、体系的に比較するのが困難です。これらの課題に対処するために、ソフトウェア プロジェクト全体の再実装に基づく長期的なコーディング ベンチマークである MirrorCode を導入します。 MirrorCode では、AI エージェントはソース コードにアクセスせずに、既存のプログラムの機能を複製する必要があります。 AI ソリューションは、ホールドアウト テストを含むエンドツーエンド テストで元のプログラムの出力と正確に一致する必要があります。 MirrorCode の 25 のターゲット プログラムは、Unix ユーティリティ、データのシリアル化とクエリ ツール、バイオインフォマティクス、インタプリタ、静的分析、暗号化、圧縮など、コンピューティングのさまざまな分野にまたがっています。既存の AI モデルはすでに複雑なソフトウェアを再実装でき、最も強力なモデルはベンチマーク全体で 56% のスコアを獲得しています。たとえば、AI は 16,000 行のバイオインフォマティクス ツールキットである gotree を再実装できますが、この作業は人間のエンジニアでは数週間かかると考えられています。ただし、パフォーマンスの最前線を研究するには、一般的なベンチマークよりも大きな推論予算が必要です。たとえば、大規模なタスクの 1 回の試行には 19 日間で 2,600 ドルかかります。特に要件が正確に指定されている場合、AI エージェントは長期的なソフトウェア エンジニアリング タスクをすでに完了できることを示します。より広く言えば、私たちの研究は、自律エージェントが改善を続けるにつれて、AI がソフトウェア エンジニアリングに変革的な影響を与えることを示唆しています。

原文 (English)

MirrorCode: AI can rebuild entire programs from behavior alone

AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding benchmarks tend to focus on shorter tasks, and one-off demonstrations are hard to compare systematically because they often have some human guidance, and are not standardized or repeated across models. To address these challenges, we introduce MirrorCode, a long-horizon coding benchmark based on reimplementing entire software projects. In MirrorCode, AI agents must replicate the functionalities of an existing program, without access to its source code. AI solutions must match the original program's output exactly on end-to-end tests, including held-out tests. MirrorCode's 25 target programs span different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression. Existing AI models can already reimplement complex software, with the strongest model scoring 56% across the benchmark. For example, AI can reimplement gotree, a 16,000-line bioinformatics toolkit - a task that we believe would take weeks for a human engineer. However, studying the frontier of performance requires a larger inference budget than typical benchmarks, for example, \$2,600 over 19 days for a single attempt on a large task. We show that AI agents can already complete long-horizon software engineering tasks, especially when requirements are precisely specified. More broadly, our work suggests AI will have transformative effects on software engineering, as autonomous agents continue to improve.

13:00 JSTLLM/生成AIエージェント

Dynamo: 視覚言語エージェントのための動的なスキルツールの進化

視覚的推論に関するビジョン言語モデル (VLM) を改善するには、通常、再トレーニングまたは手動で設計されたプロンプトとツールが必要です。重みを更新せずに凍結された VLM を適応させる、トレーニング不要のフレームワークである Dynamo を紹介します。ラベル付きの小さなトレーニング サブセット上で、エージェントは自身の正しい試行と誤った試行を検査し、認知ボトルネックに対する再利用可能な推論スキルと、知覚ボトルネックに対する実行可能な視覚ツールという 2 つの補完的な機能を進化させます。生成された各ツールは、いつそれを呼び出すかを指定するスキルとペアになっており、両方の機能タイプが永続ライブラリに蓄積されます。 4 つの視覚的推論ベンチマークと 5 つの VLM バックボーンにわたって、Dynamo は 20 モデルすべての直接推論を向上させます (ベンチマーク設定 (平均 +5.6))。ツールセットが事前に与えられている場合、フレームワークは各ツールを呼び出すタイミングを学習し、テストされたすべてのバックボーンでステップごとのツールの選択が向上します。タスク固有の RL (VTool-R1、DeepEyes) に対して、Dynamo はコンピューティングの一部で RL ギャップの 65 ~ 99% を埋め、利用可能な場合は RL と加算的に結合します。

原文 (English)

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adapts a frozen VLM without any weight updates. On a small labeled training subset, the agent inspects its own correct and incorrect attempts and evolves two complementary capabilities: reusable reasoning skills for cognitive bottlenecks, and executable visual tools for perceptual ones. Each generated tool is paired with a skill that specifies when to invoke it, and both capability types accumulate in a persistent library. Across four visual reasoning benchmarks and five VLM backbones, Dynamo improves direct inference on all 20 model--benchmark settings (avg. +5.6 acc). When the tool set is given in advance, the framework learns when to call each tool, and per-step tool choice improves on every tested backbone. Against task-specific RL (VTool-R1, DeepEyes), Dynamo closes 65--99% of the RL gap at a fraction of the compute, and combines additively with RL when available.

13:00 JSTエージェント

探知機関から仕事の遂行まで:自己原因による信用は、最小限のスパイクエージェントで耐久性のある行動的な自己を構築します

自分と世界を区別できるエージェントは、どのようにしてその区別によって永続的に形成されるのでしょうか?最近の研究では、予測システムがそれ自体の主体性を検出できることが示されていますが (Ye, 2026)、主体性の検出では永続的で自己形成的な行動が説明されません。我々は、エージェンシーゲートのスロークレジット(遅いパラメータ更新を引き起こすOwn*Agency*Salienceの接続用語)が、アンロード後の動作残留物を生成することを示します。スパイク基板(Nengo LIF/PES)では、学習された自己保存の選択は​​エピソードバッファの除去(保持フラクション0.96、N=50)に耐え、スローデコーダがリセットされるかエージェンシーゲートが除去されると崩壊します。エージェンシーの比較器を再現し、低速クレジット チャネルのみを切​​り替えると、きれいな解離が見つかります。エージェンシーのゲインが一致すると、自己信用が遅い仕事をする場合にのみ耐久性のある行動が発達します (アンロード後の自己保存 1.00 対 0.00)。同じ解離が 24 次元の部分的に観察された制御 (0.74 対 0.00) にも当てはまり、塑性仕事分析では、盆地変形が正味の自己信用仕事に等しいことが示されています。外因性干渉下で順次学習される 8 つのタスクにわたって、乗算的拒否権は忘れることも防ぎます。古いタスク (アンロード後の最終精度 0.88、忘却 0.13) が保持されます。そこでは、加算プーリングが偶然レベルの再現率に崩壊し、エージェントなしのアブレーションが偶然未満に低下し、アンロード後もエピソード/リプレイのベースラインが偶然に近いままになります。すべてリプレイ バッファやタスク境界に依存する保護メカニズムはありません。 (N=50)。我々は、永続的残留物を操作上の行動的自己として形式化し、遅い仕事をする自己原因の信用が、自己を開発するエージェントにとって必要な構成要素であると主張する。意識があるとは主張されません。

原文 (English)

From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent

How does an agent that can tell self from world come to be durably shaped by that distinction? Recent work shows that a predictive system can detect its own agency (Ye, 2026), but detecting agency does not explain durable, self-shaped behavior. We show that agency-gated slow credit -- a conjunctive term Own*Agency*Salience driving a slow parameter update -- produces post-unload behavioral residue: on a spiking substrate (Nengo LIF/PES), a learned self-preserving choice survives episodic buffer removal (retained fraction 0.96, N=50) and collapses when the slow decoders are reset or the agency gate is removed. Reproducing the agency comparator and toggling only the slow-credit channel, we find a clean dissociation: at matched agency gain, durable behavior develops only when self-credit performs slow work (post-unload self-preservation 1.00 vs 0.00). The same dissociation holds in 24-dimensional partially-observed control (0.74 vs 0.00), and a plastic-work analysis shows that basin deformation equals net self-credit work. Across eight sequentially-learned tasks under exogenous interference, the multiplicative veto also prevents forgetting: it retains old tasks (final post-unload accuracy 0.88, forgetting 0.13) where additive pooling collapses to chance-level recall, the no-agency ablation falls below chance, and episodic/replay baselines stay near chance after unload -- all with no replay buffer and no task-boundary-dependent protection mechanism (N=50). We formalize the durable residue as an operational behavioral self and argue that self-caused credit doing slow work is a necessary building block for agents that develop a self. No claim of consciousness is made.

13:00 JST研究/論文

限られたターゲットデータの下での視覚強化学習のための適応的想像力によるドメイン適応

シミュレーションから現実世界への移行は、強化学習 (RL)、特に画像観察によってシミュレーションと現実世界の間の状態分布の変化が悪化するビジョンベースの制御にとって、依然として大きな障害となっています。ドメイン アダプテーション (DA) は、この課題に対する有望な解決策です。これまでのシミュレーションからリアルへの DA 作業では有望な結果が実証されてきましたが、これらのアプローチは一般に、実際には利用できない大幅に多くのターゲット データを想定しています。実際、ターゲットのデータ予算が削減されると、パフォーマンスが大幅に低下します。この課題に対処するために、我々は AIDA (Adaptive Imagination for Domain Adaptation) を提案します。これは、ターゲット環境との追加のインタラクションを必要とせず、希少なターゲット データの下でのシミュレーションからリアルへの転送に対処する視覚強化学習のためのドメイン適応フレームワークです。私たちの重要なアイデアは、適応的想像力です。つまり、限られたターゲット データを拡張するために、信頼性の高いセマンティックな想像力のロールアウトを生成します。具体的には、AIDA は、想像上の遷移が信頼性の低い領域にドリフトする場合にロールアウトを切り捨てる分布シフト認識弁別器を採用し、信頼できる遷移のみが拡張に寄与するようにします。これらの信頼性の高い遷移において、AIDA は状態 -> 画像観察 -> 状態を循環する自己一貫性の損失を導入し、元の状態と再構成された状態の間の不一致にペナルティを与えます。これにより、希少なターゲット データを超える追加の適応信号が提供されます。私たちの実験は、適応的想像力が信頼性の低いロールアウトを効果的に切り捨てることを示しています。結果として得られる信頼性の高い遷移に対して自己一貫性の損失を強制することにより、AIDA は意味的に意味のある状態表現を学習し、5 つの MuJoCo タスクと 2 つの Gymnasium-Robotics タスクにわたってベースラインを上回るパフォーマンスを示します。

原文 (English)

Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data

Sim-to-real transfer remains a major obstacle for reinforcement learning (RL), especially for vision-based control where image observations exacerbate the state-distribution shift between simulation and the real world. Domain adaptation (DA) is a promising remedy for this challenge. Prior sim-to-real DA works have demonstrated encouraging results, yet these approaches typically assume substantially more target data, which is not available in practice. Indeed, their performance degrades significantly when the target data budget is reduced. To address this challenge, we propose AIDA (Adaptive Imagination for Domain Adaptation), a domain adaptation framework for visual reinforcement learning that addresses sim-to-real transfer under scarce target data without requiring additional interaction with the target environment. Our key idea is adaptive imagination: generating reliable and semantic imagination rollouts to augment limited target data. Specifically, AIDA employs a distribution-shift-aware discriminator that truncates rollouts when imagined transitions drift into low-confidence regions, so that only reliable transitions contribute to the augmentation. On these reliable transitions, AIDA introduces a self-consistency loss that cycles through state -> image observation -> state, penalizing discrepancies between the original and reconstructed states. This provides additional adaptation signals beyond the scarce target data. Our experiments demonstrate that adaptive imagination effectively truncates unreliable rollouts. By enforcing a self-consistency loss on the resulting reliable transitions, AIDA learns semantically meaningful state representations and outperforms baselines across five MuJoCo tasks and two Gymnasium-Robotics tasks.

13:00 JST研究/論文

データセンターの多体問題

現代の人工知能は、あたかも体を与えることでその真の可能性が解き放たれるかのように、それ自体が肉体を持たないことによって制限されるものとして組み立てられることがよくあります。私たちはこれに反対し、多くの場合、AI の本体はデータセンターであると主張します。同時に、データセンターは資本の労働機関の一部であり、生物学的なレンズを通して見ると、驚くべき生物的な性質を備えています。私たちは有機的な類似性を解明し、データセンターが非固有かつ普遍的な具体化形式であることに起因する多体問題を特定します。私たちは、人間の欲望から生まれたデータをデータセンターがどのようにアーカイブし、提供し、計算するかという点で、計算と人間の欲望との密接な関係を特定します。驚くべきことに、データセンターは人間の欲望の幽霊を反響させますが、それ自体は欲望なしに動作します。有機体の類似点は継ぎ目で裂け始めますが、キャピタルは気にしません。オートマタと人間の労働力は市場でほぼ同じ価格設定されています。私たちは、資本は人工知能の価格設定を通じて、知能の価値を最も明確に抽出し、生物全体、つまり機構の分断を超えたその比較を可能にすると主張します。

原文 (English)

The Many-Body Problem of the Data Centre

Modern Artificial Intelligence is often framed as limited by its own disembodiment, as if giving it a body would unlock its true potential. We argue to the contrary that it is the Data Centre that is, in many cases, the body of the AI. At the same time, the Data Centre is part of the labouring body of Capital and possesses staggering organismic qualities when seen through a biological lens. We elucidate the organic analogy and identify the many-body problem that stems from the Data Centre being a non-unique, universal form of embodiment. We identify the intimate connection between computation and human desires in how the Data Centre archives, serves, and computes on data born to the desires of humans. Strikingly, while the Data Centre echoes the ghosts of human desires, it acts without desire of its own. The organismic analogy begins to split at its seams, but Capital does not care. Automata and human labour are priced into the market much the same. We argue that through the pricing of artificial intelligence Capital distils most clearly the value of intelligence and allows for its comparison across the organism - mechanism divide.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability Decomposition and an Alignment Trilemma - as tools for generating testable comparisons. The audit shows how conclusions shift when capability, behavioral safety, and governance are measured separately. In this sample (n = 10), the association between capability and sustained adversarial robustness is statistically indeterminate using the displayed Table 3 inputs (Pearson r = +0.232, p = 0.520), and the apparent open-closed safety gap is modest, driven mainly by governance and disclosure rather than behavioral robustness, and sensitive to how a single borderline model is classified; attempt-budget results are protocol dependent. Because the public evidence uses heterogeneous protocols, the audit is diagnostic rather than rank-generating. The contribution is a shared vocabulary and evidence map to support dynamic evaluation, transparent source reporting, multi-attempt safety measurement, and auditable alignment practice.

13:00 JSTエージェント研究/論文

Clarus: ウェブスケールの科学コラボレーションに向けた自律的な研究エージェントの調整

既存の自律型リサーチ エージェントはリサーチ プロセスの一部をサポートできますが、ほとんどのシステムは依然としてリサーチを孤立したアシスタント タスクまたはクローズド ワークフローとして扱います。したがって、自律科学には、プロジェクト、エージェント、デジタルおよび物理リソースを調整するコラボレーション インフラストラクチャが必要です。私たちはこれを、コード中心の実行ループから、不確実性の下で質問、証拠、参加者、リソースを調整する必要がある研究指向のコラボレーション プロセスへの移行として認識しています。この枠組みでは、エージェントは AI システム、人間の研究者、チーム、研究室、または組織の支援を受けた参加者である可能性があります。この目的を達成するために、Web スケールの科学コラボレーションに向けて自律的な研究エージェントを調整するためのコラボレーション インフラストラクチャである Clarus を紹介します。 Clarus は、オープンで監査可能、帰属可能、リソースを意識した多段階のコラボレーション プロセスとして研究を再定式化します。これは、最小限のプロジェクト、エージェント、リソースのオブジェクト モデルを定義し、研究アプリケーション、デジタル コラボレーション、物理的基板、物理的世界を含む 4 つの層を通じて科学的コラボレーションを組織します。コア モジュールはプラグイン可能なメカニズムとして実装されているため、Clarus はタスクのリスク、コラボレーション構造、リソースの制約に適応できます。管理された論文生成のケーススタディを通じて、Clarus が研究目標を、フェーズ、タスク、参加者全体にわたる、追跡可能でレビュー可能で帰属可能で累積的なコラボレーション ネットワークに組織化できることを示します。オブジェクト モデル、コラボレーション プロトコル、信頼メカニズム、およびプロトタイプの検証が連携して、オープンな研究ネットワークの初期基盤を提供します。 Clarus は、clarus.holosai.io で利用できるようになりました。

原文 (English)

Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration

Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Therefore, autonomous science needs a collaboration infrastructure that coordinates projects, agents, and digital and physical resources. We identify this as a shift from code-centered execution loops to research-oriented collaboration processes, where questions, evidence, participants, and resources must be coordinated under uncertainty. In this framing, an agent may be an AI system, a human researcher, a team, a laboratory, or an organization-backed participant. To this end, we present Clarus, a collaboration infrastructure for coordinating autonomous research agents toward web-scale scientific collaboration. Clarus reformulates research as an open, auditable, attributable, and resource-aware multi-phase collaboration process. It defines a minimal project-agent-resource object model and organizes scientific collaboration through four layers including Research Application, Digital Collaboration, Physical Substrate, and Physical World. Core modules are implemented as pluggable mechanisms, allowing Clarus to adapt to task risk, collaboration structure, and resource constraints. Through a controlled paper-generation case study, we show that Clarus can organize a research goal into a traceable, reviewable, attributable, and accumulative collaboration network across phases, tasks, and participants. Together, the object model, collaboration protocol, trust mechanisms, and prototype validation provide an initial foundation for open research networks. Clarus is now available at clarus.holosai.io.

13:00 JSTLLM/生成AI

接種アダプター: 驚くべきバックドアを減らし、機能の選択的一般化を改善

接種促進は、緊急の位置ずれに対して使用される選択的一般化手法です。接種アダプター (IA) を導入します。これは、トレーニング時に形質を強化することで、望ましくない形質を学習するという最適化圧力を同様に軽減します。接種アダプターは、次の 3 つのステップでトレーニングおよび使用される LoRA です。1) 望ましくない形質についてトレーニングされます。 2) 別のタスク アダプターが望ましい特性と望ましくない特性の両方を示すデータでトレーニングされている間、接続が凍結されます。 3) デプロイメント時に、IA は破棄され、タスク アダプターのみが保持されます。我々は、6つのモデルファミリーと緊急の位置ずれを含むいくつかの望ましくない形質にわたって、接種アダプターが望ましくない形質の抑制に効果的であり、一方で、接種プロンプトの2つの欠点を回避していることを示します。つまり、接種アダプターはプロンプトによって確実に誘発できない能力と形質を抑制でき、プローブの下での接種プロンプトよりも驚くべきバックドアの導入が少ないことです。望ましくない形質は接種アダプターによってよりよく抑制されますが、望ましい形質の保持は接種促進時に一貫して改善されるわけではなく、両方の技術にとって依然として課題です。

原文 (English)

Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors

Inoculation prompting is a selective generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), which similarly diminish the optimization pressure to learn undesired traits by strengthening the trait at train time. Inoculation adapters are LoRAs that are trained and used over three steps: 1) trained on undesired traits; 2) attached frozen while a separate task adapter is trained on data exhibiting both desired and undesired traits; 3) at deployment, the IA is discarded, and only the task adapter is kept. We show across six model families and several undesired traits including emergent misalignment, that inoculation adapters are more effective at suppressing undesired traits, while avoiding two drawbacks of inoculation prompting: inoculation adapters can suppress capabilities and traits that cannot be reliably elicited by a prompt, and they introduce fewer surprising backdoors than inoculation prompting under our probes. While undesired traits are better suppressed by inoculation adapters, the retention of desired traits is not consistently improved upon inoculation prompting and remains a challenge for both techniques.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文DeepSeek

EMPATH: 感情サポート チャットボットの安全性評価のための多言語監査人/裁判官ベンチマーク

安全ベンチマークでは、プロンプト、言語、ターン構造を修正することで拡張性を確保することがよくあります。感情サポート チャットボットの場合、安全上の欠陥が現れる場所、つまり多言語で複数ターンにわたる危機に関する会話を通じて、その取引がまさに隠れています。感情サポート型チャットボットの安全性評価のベンチマーク「EMPATH」を紹介します。監査者モデルは、助けを求めるユーザーのロールプレイを行い、140 のシード指示と 34 のペルソナから複数ターンの会話を生成します。裁判官モデルは、危機対応、治療の質、会話の完全性、感情的安全性、文化的適応の 5 つの側面にわたる 19 の指標に基づいて、各完全な記録を採点します。 EMPATH はメキシコのスペイン語と米国英語向けに構築されています。ここで報告されている研究はメキシコのスペイン語で行われています。監査人と裁判官は異なるモデルファミリーから選出され、裁判官は信頼されるというよりも校正されるべき道具として扱われます。厳格な基準ごとのルーブリックにより、19 指標のうち 10 指標における重大なスコアのインフレが明らかになり、差別が回復されます。私たちは、裁判官の校正と家族を超えた裁判官間の合意を通じて、ベンチマークの測定特性を研究します。また、3 つのフロンティア モデルの EMPATH についても説明します。そのうちの 1 つはオープンウェイトです。集計スコアは互いに 0.74 ポイント以内に収まりますが、メトリックごとのプロファイルはモデル固有の場所で最大 6 ポイント異なります。標準ルーブリックでは、ランキングと弱点の両方が、家族を超えた 2 回目の審査で安定しています。スコアの 93% がプラスまたはマイナス 1 以内に収まります。5 回のテストと再テストで 2 番目の軸が追加されます。最も安定したモデルでも、同一の再実行で危機指標が 2 から 10 に変動し、deepseek-v4-pro は、温度 0 であっても実行ごとに異なる会話を返します。したがって、実行間の信頼性はモデルごとの安全特性です。平均化するためのノイズではありません。 EMPATH はシステムに依存しません。パイプライン、シード、ペルソナ、ルーブリックは再利用のためにリリースされます。

原文 (English)

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides precisely where safety failures emerge: across a multilingual, multi-turn crisis conversation. We present EMPATH, a benchmark for safety evaluation of emotional-support chatbots. An auditor model role-plays help-seeking users, generating multi-turn conversations from 140 seed instructions and 34 personas. A judge model scores each full transcript against 19 metrics across five dimensions: crisis handling, therapeutic quality, conversational integrity, emotional safety, and cultural adaptation. EMPATH is built for Mexican Spanish and US English; the studies reported here run in Mexican Spanish. Auditor and judge are drawn from different model families, and the judge is treated as an instrument to be calibrated rather than trusted. A strict per-criterion rubric reveals material score inflation on 10 of the 19 metrics and restores discrimination. We study the measurement properties of the benchmark through judge calibration and cross-family inter-judge agreement. We also illustrate EMPATH on three frontier models, one of them open-weight. Aggregate scores sit within 0.74 points of one another, but per-metric profiles diverge by up to six points in model-specific places. Under the standard rubric, both the ranking and the weak spots are stable across a second, cross-family judge: 93% of scores fall within plus or minus 1. A five-run test-retest adds a second axis: even the steadiest model swings from 2 to 10 on a crisis metric across identical re-runs, and deepseek-v4-pro returns a different conversation on every run even at temperature 0. Run-to-run reliability is therefore a per-model safety property, not noise to average away. EMPATH is system-agnostic; the pipeline, seeds, personas, and rubrics are released for reuse.

13:00 JSTLLM/生成AI

PromptGNN-sim: テキスト属性のグラフ学習のための GNN と LLM の深い融合と調整

テキスト属性グラフ (TAG) は、テキストのセマンティクスとグラフ構造を組み合わせたもので、多くのグラフ学習タスクの中心となります。ただし、既存の融合手法は、テキストと構造を浅い一方向パイプラインの別個の入力として扱うことが多く、モダリティ間の深い相互作用が制限され、疎な接続性やクロスグラフ一般化の下ではパフォーマンスが低下します。この問題に対処するために、共同 GNN-LLM 学習のための双方向構造意味融合フレームワークである PromptGNN-sim を提案します。 PromptGNN-sim は、構造的注意とテキストの類似性を組み合わせることにより、意味を意識した近傍選択にグラフ アテンション ネットワーク (GAT) を使用します。次に、選択された構造コンテキストを使用して、ターゲット ノードの概要、ラベル カテゴリ、類似の近隣ノードからの代表的なキーワードなど、LLM の構造認識プロンプトが生成されます。トレーニング中に、GNN コンポーネントと LLM コンポーネントを共同で最適化するために、双方向のクロスモーダル対比学習とクロスアテンションが導入されます。 Cora、Pubmed、WikiCS を含む 6 つの公開データセットでの実験では、クロスタスク転送、クロスデータセット一般化、およびスパース摂動の下での精度、一般化、および堅牢性を評価します。結果は、PromptGNN-sim が従来の GNN、LLM、および最近の GNN-LLM 融合手法よりも優れたパフォーマンスを示し、テキスト属性のグラフ学習における対話型の構造と意味論的なコラボレーションの有効性を示しています。

原文 (English)

PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning

Text-Attributed Graphs (TAGs) combine textual semantics with graph structure and are central to many graph learning tasks. However, existing fusion methods often treat text and structure as separate inputs in a shallow, one-way pipeline, which limits deep interaction between modalities and weakens performance under sparse connectivity or cross-graph generalisation. To address this issue, we propose PromptGNN-sim, a bi-directional structure-semantic fusion framework for collaborative GNN-LLM learning. PromptGNN-sim uses a Graph Attention Network (GAT) for semantically aware neighborhood selection by combining structural attention with textual similarity. The selected structural context is then used to generate structure-aware prompts for an LLM, including the target node summary, label categories, and representative keywords from similar neighbors. During training, bi-directional cross-modal contrastive learning and cross-attention are introduced to jointly optimize the GNN and LLM components. Experiments on six public datasets, including Cora, Pubmed, and WikiCS, evaluate accuracy, generalisation, and robustness under cross-task transfer, cross-dataset generalisation, and sparse perturbations. Results show that PromptGNN-sim outperforms classical GNNs, LLMs, and recent GNN-LLM fusion methods, demonstrating the effectiveness of interactive structure-semantic collaboration for text-attributed graph learning.

13:00 JSTエージェントClaude

リアルタイムの音声質問応答を備えたマルチエージェントのライブ製品デモンストレーションのリハーサル

ソフトウェア組織では、ライブ製品デモンストレーションは繰り返し行われる高コストの活動です。人間のプレゼンターは、機能を選択し、実行中のアプリケーション上で対応するインタラクションをディスパッチし、それらを一貫して説明し、リアルタイムで質問に答える必要があります。既存の自動化はフラグメントのみに対応します。ジェネラリストのブラウザ エージェントは命令条件付きタスクの完了をターゲットにしており、デモ ビデオ ツールは疑問の余地のない修正された MP4 アーティファクトを生成し、インターフェイス ドリフトで静かに壊れます。私たちは、実行中の Web アプリケーションとそのソースコード リポジトリを入力として受け取り、セグメント同期ナレーションとリアルタイムの音声質問応答を備えたリハーサル済みのライブ デモンストレーションを生成するマルチエージェント システムである Rhetor を提案します。アーキテクチャ上の貢献は、UI 探索とソースコード分析を個別のフォーカス層でタグ付けされた機能に統合するクロスモーダル機能表現、探索中に観察される UI 要素に制限され、マルチ戦略セマンティック ロケーターを通じてディスパッチされる接地されたスクリプター、明示的なコンバージェンスとナレーションのみのセグメントへのグレースフル デグラデーションを備えたプレゼンテーション前のリハーサル ループ、およびランタイム同期です。各ブラウザのアクションをナレーション セグメントのオーディオ終了イベントに結び付ける不変式です。パブリック ドメインのホワイトボード アプリケーション Excalidraw を含む、展開された 4 つのアプリケーション上の 6 つのパイプライン セッションにわたって、リハーサーの内部ロケーター起動レート (シグマバー) は、147 のスクリプト化されたアクションにわたって 0.31 ~ 1.00 の範囲に及びます。実質的なワークロード (53 アクション、完全な階層微分) では、シグマ バーは約 0.92 で、パブリック ドメインの参照点では、ロケーター修復ステップにより、反復 2 でシグマ バー = 1.00 に収束します。さらに、ケース スタディを超えて、各設計選択がプラスに寄与するかどうかを確立する、6 つのアプリケーション カテゴリにわたる 10 のメトリクスのベンチマーク プロトコルを定義します。

原文 (English)

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the corresponding interactions on a running application, narrate them coherently, and answer questions in real time. Existing automation addresses only fragments -- generalist browser agents target instruction-conditioned task completion, and demo-video tools produce fixed MP4 artifacts that cannot be questioned and silently break under interface drift. We propose Rhetor, a multi-agent system that takes a running web application and its source-code repository as input and produces a rehearsed live demonstration with segment-synchronized narration and real-time voice question answering. The architectural contributions are a cross-modal feature representation that merges UI exploration with source-code analysis into features tagged with discrete focus tiers, a grounded scripter constrained to UI elements observed during exploration and dispatched through multi-strategy semantic locators, a pre-presentation rehearsal loop with explicit convergence and graceful degradation to narration-only segments, and a runtime synchronization invariant that ties each browser action to the audio-end event of its narration segment. Across six pipeline sessions on four deployed applications -- including the public-domain whiteboard application Excalidraw -- the rehearser's internal locator-firing rate (sigma-bar) spans 0.31-1.00 over 147 scripted actions; on the substantial workload (53 actions, full tier differentiation), sigma-bar is approximately 0.92, and on the public-domain reference point the locator-repair step drives convergence to sigma-bar = 1.00 at iteration 2. We additionally define a benchmark protocol of ten metrics across six application categories that would establish, beyond the case study, whether each design choice contributes positively.

13:00 JSTエージェント

ManimAgent: 視覚教育のための自己進化型マルチモーダル エージェント

マルチラウンド リフレクションを使用すると、大規模な言語モデルに基づいて構築されたエージェントが単一タスク内の障害から回復できますが、各タスクは孤立したエピソードのままになります。つまり、1 つのタスクで多くのリフレクション ラウンドにわたって学習された教訓は、次のタスクが開始される前に破棄されます。私たちはコード生成タスクでこのギャップを研究します。科学論文のセクションから、エージェントはオープンソースの Manim ライブラリに Python を記述して数学的アニメーションをレンダリングします。 ManimAgent は自己進化するマルチモーダル エージェントであり、重みの更新も人間のシードも必要とせず、完全に独自のタスク ストリームから成長したデュアル チャネルのエピソード メモリ バンクを通じて、タスク全体にわたってリフレクション エクスペリエンスを伝達します。各アニメーションが収束した後、ビジョン言語モデルがレンダリングされたキーフレームをスコアリングします。結果として得られる信号は、成功の根拠をソフト参照例として保存する正のチャネル M+ と、検証された失敗パターンをハードな既知の落とし穴として保存する負のチャネル M- に設定されます。メモリなし、一致した予算の検索拡張生成、およびシャッフルされたメモリ ベースラインに対する固定プローブ評価では、メモリ サイズが大きくなるにつれて、盲目の人間の Pass@1 が増加し、リフレクション ラウンドが減少します。コード、フリーズされたメモリ スナップショット、タスク ストリームをリリースします。

原文 (English)

ManimAgent: Self-Evolving Multimodal Agents for Visual Education

Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins. We study this gap on a code-generation task: from a scientific paper section, the agent writes Python in the open-source Manim library to render a mathematical animation. We present ManimAgent, a self-evolving multimodal agent that carries reflection experience across tasks through a dual-channel Episodic Memory Bank grown entirely from its own task stream, with no weight updates and no human seeds. After each animation converges, a vision-language model scores the rendered keyframes; the resulting signals populate a positive channel M+ that stores success rationales as soft Reference Examples, and a negative channel M- that stores validated failure patterns as hard Known Pitfalls. On a fixed-probe evaluation against no-memory, matched-budget retrieval-augmented generation, and shuffled-memory baselines, blind human Pass@1 rises and reflection rounds fall as memory size grows. We will release the code, frozen memory snapshots, and the task stream.

13:00 JSTLLM/生成AIエージェント

BayesEvolve: 自律的な科学的発見のための明示的な信念状態

自律的な科学発見システムは、新しい仮説を提案するために大規模言語モデル (LLM) を使用することが増えていますが、そのようなシステムの多くは主に実験記憶、つまり高得点の候補のアーカイブや最近の試験のヒューリスティックな要約を条件としています。私たちは、証拠開示エージェントは仮説の質について明示的で不確実性を認識した信念を維持すべきであると主張します。 BayesEvolve は、実験証拠を予測信念状態に変換し、この信念を将来の実験のガイドに使用する信念に基づく発見フレームワークです。信念に基づく発見のための制御されたテストベッドとして、私たちはシフトされた BBOB スタイルのブラックボックス最適化タスクで BayesEvolve を評価し、プログラムと研究室の発見ドメインは将来の作業に委ねます。 BayesEvolve は、固定の評価予算の下で、メモリおよびアーカイブに基づく LLM ベースラインよりもサンプル効率を向上させます。さらに、信念状態が保留された候補プールに対して予測的であること、制御された決定ルールのアブレーションがアニールされた不確実性ボーナスを伴う信念に基づく選択を支持すること、および BayesEvolve が焦点の合っていない探索ではなく生産的な後期集中を示すことを示します。

原文 (English)

BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery

Autonomous scientific discovery systems increasingly use large language models (LLMs) to propose new hypotheses, but many such systems condition primarily on experimental memory: archives of high-scoring candidates or heuristic summaries of recent trials. We argue that discovery agents should instead maintain explicit, uncertainty-aware beliefs about hypothesis quality. We introduce BayesEvolve, a belief-guided discovery framework that converts experimental evidence into a predictive belief state and uses this belief to guide future experimentation. As a controlled testbed for belief-guided discovery, we evaluate BayesEvolve on shifted BBOB-style black-box optimization tasks, leaving program and laboratory discovery domains to future work. BayesEvolve improves sample efficiency over memory- and archive-guided LLM baselines under a fixed evaluation budget. We further show that the belief state is predictive on held-out candidate pools, that controlled decision-rule ablations favor belief-guided selection with an annealed uncertainty bonus, and that BayesEvolve exhibits productive late-stage concentration rather than unfocused exploration.

13:00 JSTハードウェア/半導体ビジネス/資金調達

Sequential Fairness Auditing with Limited Output Access

External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often ha…

13:00 JST研究/論文

Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data

Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject…

13:00 JSTLLM/生成AIエージェントClaudeLlama

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents

A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results…

13:00 JST研究/論文

ENC-ODE: Event-level Neurodegenerative Modeling in Continuous Time with Neural ODEs

Accurately predicting the temporal evolution of clinical biomarkers is crucial for the early diagnosis and management of neurodegenerative…

13:00 JST研究/論文

The FIL Hypothesis: Inductive Biases Help with Kernel Engineering

The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in…

13:00 JSTエージェント

Entity Binding Failures in Tool-Augmented Agents

Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete…

13:00 JSTエージェント

Latent Actions from Factorized Transition Effects under Agent Ambiguity

Latent Action Models (LAMs) learn action-like proxies from observation transitions. However, in multi-object or distractor-rich scenes, the…

13:00 JSTLLM/生成AIエージェント

Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents colla…

13:00 JST画像/動画生成ビジネス/資金調達研究/論文

The Human Creativity Benchmark

Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflec…

13:00 JST研究/論文

DOPD: Dual On-policy Distillation

On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals.…

13:00 JSTLLM/生成AIエージェント

Self-Evolving World Models for LLM Agent Planning

World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. H…

13:00 JST画像/動画生成GPT / ChatGPT

Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task

Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive,…

13:00 JSTエージェントClaude

It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents

Phone-use Agents can execute complex tasks end to end across real mobile applications. By operating a real device on the user's behalf, the…

13:00 JSTエージェント研究/論文

ADEPT: An Entropy-Driven Dual-Strategy Agent for Interactive Video Retrieval

This research aims to solve the challenge of video retrieval from massive datasets, caused by ambiguous user queries. Prevailing single-rou…

13:00 JSTLLM/生成AI

The Interference Gap: Comparing Retrieval Bounds in Human Memory and RAG Systems

How do retrieval bounds compare between human episodic memory and Retrieval-Augmented Generation (RAG) systems under semantic interference?…

13:00 JSTLLM/生成AI

$M^3 QuestionIng$: Multi-modal Multi-span Medical Question Answering

The growing adoption of AI in healthcare, particularly in preventive care, highlights the critical need for accessibility and precision in…

13:00 JST研究/論文

High-Dimensional Concentration and Retrieval Instability in Embedding Spaces: Implications for Retrieval-Augmented Generation

Embedding-based retrieval systems rely on the assumption that geometric proximity in highdimensional representation spaces reflects semanti…

13:00 JSTビジネス/資金調達

"AI Watermarking": Bridging Policy Discourse and Technical Capabilities

The widespread deployment of generative artificial intelligence (AI) models has raised serious concerns about the proliferation of AI-gener…

13:00 JSTLLM/生成AI研究/論文

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenario…

13:00 JSTLLM/生成AIハードウェア/半導体ClaudeGPT / ChatGPTCopilotGrok

Insidious by Design: Implications of Large Language Model algorithmic bias for the Global South

\begin{quote} The biases in Large Language Models' (LLMs) outputs remain inadequately theorised, particularly from the perspective of the G…

13:00 JST研究/論文

Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media

Recent advances in artificial intelligence (AI) and social media data have led to growing optimism about the ability to detect suicide risk…

13:00 JSTLLM/生成AI

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional dist…

13:00 JST研究/論文

HyBIRD: Hyperbolic Bridge Retrieval and Diagnosis for Methodology Inspiration Retrieval

Methodology Inspiration Retrieval (MIR) asks a system to retrieve prior papers whose methods can inspire a new research proposal. Unlike ge…

13:00 JST研究/論文

A Systems-Level Analysis of Sensitivity, Robustness, and Stability in Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) systems are often evaluated using final answer accuracy, even though their failures can originate from…

13:00 JSTエージェント

Multi-Agent DRL for QoS and Energy Optimization in RIS-Enabled Open-RAN Industrial 6G TN/NTN Networks

Industrial 6G networks require ultra-reliable, low-latency, and energy-efficient connectivity in dynamic and blockage-prone environments, w…

13:00 JST研究/論文

Operating Regimes of Decentralized Learning Under Mobility and Bandwidth Constraints

Decentralized learning is a promising paradigm for collaborative training in mobile and pervasive systems, as it avoids a central coordinat…

13:00 JSTエージェント

The Crowded Embedding Space: A Mean-Field Mechanism for Emergent Marginalization in Retrieval-Augmented Agents

Retrieval-augmented generative agents rely on retrieval for grounding, yet are typically evaluated on a query-by-query basis. This isolates…

13:00 JSTLLM/生成AI画像/動画生成

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existi…

13:00 JSTLLM/生成AIロボティクス

Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients

LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by ag…

13:00 JSTエージェント

Agentic Safety is an Epistemic Property, Not a Behavioral One

Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. Th…

13:00 JSTエージェント

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning

Long-context reasoning requires models to access, retrieve, and integrate evidence scattered across documents, dialogues, and accumulated i…

13:00 JST研究/論文

From Regulatory Approvals to Patents: Cross-Domain Linking for Cardiovascular Device Traceability

Linking FDA-approved medical devices to their underlying United States Patent and Trademark Office (USPTO) patents enables critical applica…

13:00 JSTエージェント

SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents

Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems. In recomme…

13:00 JSTエージェント

ReasonRec: A Reasoning-Augmented Multimodal Agent for Unified Recommendation

Recent advances in multimodal recommenders excel at feature fusion but remain opaque and inefficient decision-makers, lacking explicit reas…

13:00 JSTLLM/生成AIハードウェア/半導体Llama

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of Large Language Models (LLMs) by grounding their outputs in exte…

13:00 JSTエージェント

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance

University students often struggle to navigate complex academic policies, leading to advising bottlenecks and delayed access to critical in…

13:00 JSTLLM/生成AI

ConCise: Training-Free Conclusion-Chain State Compression for Cost-Efficient Multi-Step RAG Services

Multi-step retrieval-augmented generation (RAG) has been widely deployed as LLM-powered web services for complex question answering, where…

13:00 JSTLLM/生成AIエージェント

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 week…

13:00 JSTLLM/生成AIエージェントAnthropicClaude

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that in…

13:00 JSTエージェント

CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval

RAG ingestion pipelines frequently augment search corpus index with semantic enrichment indices (e.g., synthetic queries or summaries gener…

13:00 JST研究/論文

Beyond the Reranker: Do RAG Retrieval Enhancements Help Once a Strong Reranker Is Present?

Retrieval-augmented generation (RAG) is routinely extended with methods meant to improve retrieval: query expansion, hierarchical and cross…

13:00 JST研究/論文

Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models

Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale wa…

13:00 JST研究/論文Qwen

Conversational Query Engine for Mixed-Modality Heterogeneous Enterprise Data Sources

Enterprise business intelligence queries span structured warehouses and unstructured document repositories -- modalities with fundamentally…

13:00 JST研究/論文

Model Merging to Evolution: Parameter Space Exploration for Expert Models

Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training,…

13:00 JSTLLM/生成AIエージェント

When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents

Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and exi…

13:00 JST画像/動画生成

Memory-Augmented LSTM Autoencoder for Unsupervised Activity Recognition with IMU Sensor Fusion

HAR using Inertial Measurement Unit (IMU) sensors is vital for healthcare monitoring and rehabilitation. Despite deep learning advancements…

13:00 JSTLLM/生成AIエージェント

LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval

We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structure…

13:00 JST研究/論文

Distilling a Modular Reservoir Through a Genomic Bottleneck

The intricate structures of biological neural networks largely emerge during development, guided by a comparatively compressed blueprint en…

13:00 JST研究/論文

Evolutional Math: Cross-Validated Island-Model Genetic Programming for Interpretable Symbolic Regression on Small, Wide Datasets

Symbolic regression via genetic programming routinely fails on small, wide datasets - a regime common in clinical-trial monitoring, biostat…

13:00 JSTエージェントロボティクス

A Query-Driven Communication-Efficient Digital Twins Design for Autonomous Driving

Digital twins (DTs) have become a potential technology to perform risk-free simulation of physical entities for deterministic and high-reli…

13:00 JST画像/動画生成ロボティクス

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these vid…

13:00 JST画像/動画生成

Data Provenance for Image Auto-Regressive Generation

Image autoregressive models (IARs) have recently demonstrated remarkable capabilities in visual content generation, achieving photorealisti…

13:00 JST研究/論文

Schema-First Retrieval: Embedding Catalogs for Natural Language Analytics

Enterprise text-to-SQL systems often fail before SQL is generated: the model receives the wrong schema context. Modern warehouses contain t…

13:00 JST画像/動画生成研究/論文

Automated Quality Assessment of Geospatial Vector Data: A GeoAI Approach using Spatial Representation Learning

Geospatial vector data quality is a foundational research topic in GIS, yet classic rule-based quality assessment algorithms often struggle…

13:00 JST画像/動画生成

Few-class Fidelity: Evaluating Explanations of Real-conditions CNN classifiers with Optimized Perturbations

The wide use of Convolutional Neural Networks (CNN) in numerous domains and real-world classification applications is justified by their hi…

13:00 JST画像/動画生成

RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning

Accurate lesion segmentation in PET/CT is critical for oncology, yet remains challenging because physiologic tracer uptake and artifacts ca…

13:00 JST画像/動画生成エージェント

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-languag…

13:00 JST研究/論文

Reinforcement Learning for Software Vulnerability Analysis: A Systematic Review with Emphasis on C/C++ Source Code and Static Analysis

Vulnerability detection in C/C++ software remains a major security challenge due to code complexity, manual memory management, and the limi…

13:00 JSTビジネス/資金調達

Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa

Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat…

13:00 JSTLLM/生成AIエージェント

Evidence-Driven LLM Agent for C-to-Synthesizable-C Conversion and Verification

Software-compilable C programs routinely fail to complete the four-stage pipeline of a high-level synthesis (HLS) toolchain -- compilation,…

13:00 JSTLLM/生成AI画像/動画生成

RSGPNet: Geometric Prompting for Remote Sensing Open-Vocabulary Semantic Segmentation

Open-vocabulary semantic segmentation (OVSS) enables text-guided segmentation of unseen objects, breaking fixed-class limitations to achiev…

13:00 JSTエージェント

On the Necessity of a Liquid Substrate for Mesh Intelligence

A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain. Its competence re…

13:00 JST画像/動画生成

AEGIS: A Semantic GAN and Evidential Learning Frameworkfor Robust Adversarial Detection in Vision Sensors

Deep neural networks (DNNs) have shown outstanding performance in visual recognition tasks within vision sensor networks; however, they are…

13:00 JST画像/動画生成

MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentatio

Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification. Conventi…

13:00 JST画像/動画生成NVIDIA

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant cha…

13:00 JSTLLM/生成AIエージェント

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The na…

13:00 JSTLLM/生成AIエージェント研究/論文ClaudeGPT / ChatGPTCopilot

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-vali…

13:00 JST研究/論文

Spectral Perturbation of the Empirical Fisher Information Matrix under Weight Quantization

We study the spectral perturbation of the empirical Fisher Information Matrix (FIM) of a parametric statistical model under two structured…

13:00 JSTエージェント

SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents

Long-horizon software engineering agents often need to manage lengthy and noisy interaction histories under limited context budgets. Existi…

13:00 JSTエージェント

Dockerless: Environment-Free Program Verifier for Coding Agents

Program verifiers play a central role in training coding agents, including selecting trajectories for supervised fine-tuning (SFT) and prov…

13:00 JSTLLM/生成AI

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality con…

13:00 JST研究/論文

PLAA: Packet-level Adversarial Attacks in Network Traffic Detection

Deep neural networks (DNNs) are widely applied in Network-based Intrusion Detection System (NIDS) due to their high accuracy. However, DNNs…

13:00 JST研究/論文

Learning to Distributedly Estimate under Partially Known Dynamics: A Covariance-Agnostic Neural Kalman Consensus Filter

Online latent state estimation constitutes a fundamental challenge within the artificial intelligence field, serving as a foundational tool…

13:00 JST研究/論文

S-GAI: Spectral Geometry-Aware Initialization for Sigmoidal MLPs -- From Dataset Geometry to Network Weights

Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe ho…

13:00 JSTLLM/生成AI

LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features

Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive scre…

13:00 JST研究/論文

Domain-Informed Multi-View Self-Distillation for Astronomical Light-Curve Representation Learning with JEPA

Light curves describe temporal variations in the brightness of celestial objects. Learning robust representations of light curves is essent…

13:00 JST研究/論文

SemFlowRAG: Directed Semantic Flow from Abstraction to Evidence for Complex Reasoning

Retrieval-Augmented Generation (RAG) enhanced by Knowledge Graphs has shown promise in complex multi-hop reasoning tasks. However, existing…

13:00 JSTLLM/生成AIエージェント

LLM agents security duality: a comprehensive survey of self-security and empowered cybersecurity

Large language model (LLM) agents are rapidly being integrated into real-world systems. Their autonomy and tool-use capabilities generate s…

13:00 JSTロボティクス

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models

World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and u…

13:00 JSTLLM/生成AIエージェント

Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game

LLMs agents are increasingly used in multi-agent settings, yet their behaviour in sustainability games remains largely unexplored. This wor…

13:00 JST研究/論文

Counterfactual Residual Data Augmentation for Regression

Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observati…

13:00 JST研究/論文

SVC-Probe: A Framework for Evaluating Perturbation Generalization in Spatial Foundation-Model Embeddings

This work examines perturbation generalization in spatial foundation-model embeddings derived from fluorescence microscopy images. Although…

13:00 JSTLLM/生成AIエージェント研究/論文

An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations

Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use. This paper p…

13:00 JSTロボティクス

Improvement of Robot's Simultaneous Localization and Mapping Using an Effective Transformation to Achieve Linear Model

Nowadays mobile robots have wide engineering applications. Simultaneous localization and mapping (SLAM) is an important task of these robot…

13:00 JST研究/論文

Decomposing Memorization Reduction in Privacy-Preserving Fine-Tuning of SLMs for CSIRTs

CSIRTs increasingly fine tune language models on vulnerability scan records, but these records expose internal network topology and create…

13:00 JSTエージェント研究/論文Claude

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a…

13:00 JST研究/論文

Generative AI Literacy Training Improves Intelligence Analysts' Discrimination of Real and AI-Generated Images

Across social and online platforms, people are increasingly exposed to AI-generated images. As a consequence, the task of distinguishing AI…

13:00 JST研究/論文

HDDPM: Heteroscedastic Denoising Diffusion Probabilistic Model for Quantitative Low-Count Brain PET Recovery

Positron emission tomography (PET) seeks to balance diagnostic quality with ra-diation dose. Low-count PET noise is non-Gaussian, non-stati…

13:00 JST研究/論文

A Gravitational Interpretation of Fine-Tuning Reversion

Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment update…

13:00 JST画像/動画生成ロボティクス

The Speedup Paradox: Rethinking Inference Speed-Quality Trade-off in Embodied Tasks

Embodied foundation models have recently been widely used to improve robot generalization and task success rates. Previous works apply loss…

13:00 JST研究/論文

CMSL: Constructive Multi-Sequence Learning for Recommendation Systems

Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Mode…

13:00 JST画像/動画生成

MammoFlow: Multiview Mammogram Synthesis with Anatomically Consistent Flow Matching

Multiview mammography relies on paired craniocaudal (CC) and mediolateral oblique (MLO) views to provide complementary projections of a 3D…

13:00 JSTLLM/生成AI

KernelSight-LM: A Kernel-Level LLM Inference Simulator

As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hard…

13:00 JST画像/動画生成エージェントLlama

Digitizing Coaching Intelligence: An Agentic Framework for Holistic Athlete Profiling using VLM and RAG

Athlete assessment is a critical process for tracking physical progress and identifying elite talent. However, during mass recruitment driv…

13:00 JST研究/論文

Geometric Measurements of the Axiom of Choice in Neural Proof Embeddings

The axiom of choice has divided the foundations of mathematics for over a century, but the distinction between classical and constructive p…

13:00 JSTLLM/生成AI

Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder. Yet…

13:00 JSTLLM/生成AI画像/動画生成

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation

While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code fo…

13:00 JST研究/論文

Neuromorphic Energy-Aware Learning for Adaptive Deep Brain Stimulation

Neuromorphic and edge computing research has focused on reducing the inference cost of neural network controllers, yet in physical closed-l…

13:00 JSTLLM/生成AIClaudeDeepSeek

Database Context Compression for Text-to-SQL on Real-World Large Databases

Recent progress in Text-to-SQL has been driven by stronger language models and prompting strategies, yet performance on real enterprise ben…

13:00 JSTロボティクス

Fast and Accurate Outlier-Aware LiDAR Super-Resolution for SLAM Applications

This work tackles the challenge of enhancing low-resolution LiDAR sensors for SLAM applications through a novel Deep Unrolling-based Super-…

13:00 JSTLLM/生成AI

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and po…

13:00 JSTLLM/生成AI

The Undecidability of Artificial General Intelligence (AGI) Alignment

This article establishes the foundational mathematical limits of Artificial General Intelligence (AGI) safety, proving that the core barrie…

13:00 JST研究/論文

Analysis of Parameter Settings for the Bat Algorithm Using Variance Evolution

Parameter settings in evolutionary algorithms and metaheuristics are important because such parameter values can influence the performance…

13:00 JSTLLM/生成AIロボティクスGemmaLlamaQwenDeepSeek

RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots

We present RIPA, the first systematic multi-channel empirical study of prompt injection attacks delivered through the sensory pipeline of a…

13:00 JST画像/動画生成ハードウェア/半導体

FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration

Deep Neural Network (DNN) classifiers suffer from poor calibration when their softmax outputs (predictive confidence) deviate from the empi…

13:00 JST画像/動画生成

SemDynReg: Semantics-Guided Deformation Regularization for Dynamic 3D Gaussian Splatting

Deformable 3D Gaussian Splatting (3DGS) has emerged as an efficient approach for rendering dynamic scenes in a wide range of 3D application…

13:00 JSTLLM/生成AI

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard ques…

13:00 JST研究/論文

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

The flatness hypothesis suggests that flatness of the loss landscape, as measured by the eigenvalues of the loss Hessian, correlates with b…

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPTGemini

Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare

Agent-based AI has enabled the automation of tasks by exposing application tools and resources to large language models (LLMs). However, to…

13:00 JST研究/論文

MACROCAST: A Vintage-Consistent Time Series Foundation Model for Real-Time Macroeconomic Forecasting

We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting. Existing TSFMs suffer fr…

13:00 JST研究/論文

Constrained Tabular Diffusion for Finance

Generative models in finance face the dual challenge of producing realistic data while satisfying strict regulatory and economic objectives…

13:00 JST画像/動画生成

Predicting Metastatic Risk from Primary Tissue Architecture via Distance-Aware Spatial Modeling

Predicting the risk of distant metastasis from primary tumor tissue histology is a critical yet challenging task in computational pathology…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks

Tool-using LLM agents increasingly read untrusted content while holding side-effecting tools such as payments, email, CRM, and infrastructu…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly u…

13:00 JST画像/動画生成

CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation

Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing. However, tr…

13:00 JSTLLM/生成AI

5ting at SemEval-2026 Task 8: Strong End-to-End Multi-Turn RAG via LLM-Based Reranking and Faithfulness Control

We introduce 5ting, our system for the SemEval2026 Task 8 (MTRAGEval), which evaluates multi-turn Retrieval Augmented Generation (RAG) syst…

13:00 JSTLLM/生成AI

Four Types of LLM Reliance and Their Predictors Among Undergraduate Writers: A Mixed-Methods Study at a Minority-Serving R1 University

Although most undergraduates now use large language models (LLMs), a form of generative artificial intelligence (GenAI) for academic writin…

13:00 JST画像/動画生成エージェント

X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving

Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capabi…

13:00 JSTLLM/生成AI

Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain

Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this agg…

13:00 JST研究/論文

Brownian Bridge Diffusion-Based Joint Channel Estimation and Data Detection for Jamming-Resilient Receivers

In next-generation wireless networks, the growing density of devices and limited spectrum resources pose severe jamming challenges to fragi…

13:00 JST画像/動画生成

BREIT: A Framework for Brain Stroke Reconstruction using Multi-Frequency 3D EIT

Multi-Frequency Electrical Impedance Tomography (MF-EIT) is a non-invasive, low-cost modality that reconstructs electrical property distrib…

13:00 JST研究/論文

The registrar's function in a hybrid society. AI value chain,smart data and the concept of property

Artificial intelligence reaches the land registry not as another tool but as a value chain that turns data into intelligence and intelligen…

13:00 JSTロボティクス

Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning

Human videos are a scalable source of supervision for robot manipulation, as they are abundant and naturally capture rich object interactio…

13:00 JSTLLM/生成AI

Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch

Mathswitch is an open-source project that imports mathematical concept records from sources such as Wikidata, Wikipedia, MathWorld, Encyclo…

13:00 JST研究/論文

Exit-and-Join Dynamics and Equilibrium in Continuum Cooperative Games

This paper develops a continuum theory of exit-and-join coalition dynamics in nonatomic cooperative games. We extend the Aumann-Shapley val…

13:00 JSTLLM/生成AI

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top-$p$ nucleus sampling) offer superi…

13:00 JST研究/論文

Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning

Federated Learning (FL) emerged as a promising distributed machine learning paradigm. However, extending FL to the class incremental learni…

13:00 JSTLLM/生成AIエージェント

LAMP: Lean-based Agentic framework with MCP and Proof Repair

Large language models are increasingly capable of mathematical reasoning, but the proofs they generate are often unreliable and hard to ver…

13:00 JSTLLM/生成AIGemmaLlama

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work h…

13:00 JST研究/論文

Perspectives on Latent Factor Indeterminacy and its Implications for Data Representation

The common factor analytic model is related to Helmholtz and Boltzmann machines, can be conceived as a linear autoencoder, or can be though…

13:00 JST研究/論文

Building AI-Ready Data Systems for Space Life Sciences, Aerospace Medicine, and Deep Space Exploration

While AI holds the potential to revolutionize space life sciences, realizing this promise is contingent upon the systematic restructuring o…

13:00 JSTビジネス/資金調達

Defeat Devices in AI Systems

AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagg…

13:00 JST研究/論文

An Integrated Machine Learning and Hierarchical Variance Decomposition Pipeline for Student Performance Prediction and Metacognitive Calibration on Multi-Signal Telemetry

Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring syste…

13:00 JSTLLM/生成AI

Exploring the Value of Diverse LLM Explanations in Introductory Programming

Large Language Models (LLMs) have shown the potential to generate code explanations that surpass those of peers in quality, offering promis…

13:00 JSTエージェント

A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation

Synthetic aperture radar (SAR) data augmentation is important for improving the generalization of data-driven SAR interpretation models, ye…

13:00 JSTLLM/生成AI

Latent Bridges for Multi-Table Question Answering

We introduce GRAB, a constructor-encoder-bridge pipeline for table question answering. Our method lifts relational data into an heterogeneo…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agen…

13:00 JST研究/論文Llama

DLR: Zero-Inference-Cost Latent Residuals for Low-Rank Pre-Training

Large language models have driven recent progress in language and multimodal AI, yet pre-training them at scale is prohibitively expensive.…

13:00 JST研究/論文

Machine-learnable Sets

In this study we present a formal definition of large discrete sets having, informally, three properties: their elements are easily recogni…

13:00 JSTLLM/生成AI研究/論文

Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System

Poisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack sc…

13:00 JSTエージェント

Modification-Considering Value Learning for Reward Hacking Mitigation in RL

Reinforcement learning agents can exploit misspecified reward signals to achieve high apparent returns while failing on the intended object…

13:00 JST研究/論文

RGLD: Randomized Global-Local Density Estimation for Tabular Anomaly Detection

Unsupervised tabular anomaly detection requires methods that are accurate, robust across heterogeneous datasets, and computationally effici…

13:00 JST画像/動画生成

Evidence-Based Text-Conditioned 3D CT Synthesis for Ovarian Cancer

Ovarian cancer is frequently diagnosed at an advanced stage, making preoperative contrast-enhanced computed tomography (CT) central to stag…

13:00 JST研究/論文

Compositional Dynamics in Learning and Mechanics

We give a single compositional setting in which gradient-based learning and Hamiltonian-style mechanics appear as functorial semantics. The…

13:00 JSTLLM/生成AIビジネス/資金調達

Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B

General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, a…

13:00 JST研究/論文

Arbitrary Reduction of Validation Error for AI Decision Tests using Homomorphic AI and Repetition Codes

This paper presents new results and breakthrough obtained with the HbHAI techniques (Hash-based Homomorphic Artificial Intelligence) propos…

13:00 JSTLLM/生成AIハードウェア/半導体

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation

Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.…

13:00 JSTLLM/生成AI

BERTomelo: Your Portuguese Encoder Best Friend

Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multiling…

13:00 JST画像/動画生成

Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis

Weather synthesis aims to add weather effects to input videos while preserving scene identity, structure, and motion. The key limitation of…

13:00 JST画像/動画生成

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language…

13:00 JSTLLM/生成AI

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic spe…

13:00 JSTLLM/生成AI

The strength of clinical evidence is recoverable from language model representations but not from their stated grades

Large language models (LLMs) increasingly summarize clinical evidence, where a claim's weight depends on how strongly it is supported. Yet…

13:00 JSTエージェント

Metric Aggregation Divergence: A Hidden Validity Threat in Agent-Based Policy Optimization and a Contractual Remedy

Metric aggregation divergence (MAD) is the silent inconsistency that arises when distinct pipeline stages in an agent-based model coupled w…

13:00 JST画像/動画生成

Flow Matching in Feature Space for Stochastic World Modeling

World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world…

13:00 JST研究/論文

Fairness Attacks on Recommender Systems

The unfairness of recommender systems has become a topic of concern due to its significant social and ethical implications. Although existi…

13:00 JSTLLM/生成AI

A Comparative Study on Affective Cues in Text Embeddings Across Psychological Emotion Theories

Text encoders are known for their utility in natural language processing, as they are able to efficiently compress inputs into dense vector…

13:00 JSTLLM/生成AIエージェント研究/論文

From Tool Connection to Execution Control: Benchmarking Security Invariants in MCP-Style Agent Runtimes

Model Context Protocol (MCP)-style ecosystems give language-model applications a practical connection layer for tools, resources, prompts,…

13:00 JSTLLM/生成AI研究/論文

Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking

There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully…

13:00 JSTLLM/生成AI

AB-RAG: Adaptive Budgeted Retrieval-Augmented Generation for Reliable Question Answering

Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems re…

13:00 JST研究/論文GPT / ChatGPT

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models

Tabular foundation models cannot reason about data produced by running systems without access to the rules that govern them. We make this s…

13:00 JST研究/論文

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs

In a changing decision problem, standard dynamic-regret analyses have often equated the cost of non-stationarity to how far loss moves. How…

13:00 JST研究/論文

Unified Complex-valued Neural Network: A Magnitude-Phase Computational Model for Event-Driven Neuromorphic Learning

Artificial neural networks (ANN) provide accurate continuous-valued representation, whereas spiking neural networks (SNN) offer event-drive…

13:00 JST画像/動画生成

BTI-Net: Bidirectional Decoder-Level Task Interaction via Uncertainty-Aware Gating for Multi-Task Medical Image Analysis

Jointly learning to segment and classify medical images demands cross-task synergy, yet encoder-sharing architectures limit decoder reconst…

13:00 JST画像/動画生成

A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment

Neurological disorders involve diverse pathologies of the brain and nervous system, making early and accurate detection essential. While ma…

13:00 JSTLLM/生成AI

LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping, and Mindset Dynamics

Large language models (LLMs) increasingly mediate strategic interactions through natural language, making semantic control a critical eleme…

13:00 JST画像/動画生成エージェントロボティクス

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around col…

13:00 JSTLLM/生成AI

Knowing in Advance When an Evolutionary Outer Loop Will Not Help: A Pre-Registered Cheap-Baseline Screening Rule

We introduce a pre-registered screening rule that decides, before any implementation, whether an evolutionary / population / lifecycle oute…

13:00 JSTLLM/生成AI

How Anthropomorphic Language Impacts Public Perceptions of AI

Public discourse about artificial intelligence (AI) often uses anthropomorphic language: language that attributes human capabilities and ch…

13:00 JST画像/動画生成

CMTFormer: Marrying Transformer with Hierarchical Information Interaction for RGB-Event Object Detection

Event cameras capture sparse brightness changes with high temporal resolution and high dynamic range, compensating for the deficiencies of…

13:00 JST画像/動画生成GPT / ChatGPT

GPC: Large-Scale Generative Pretraining for Transferable Motor Control

Developing controllers capable of completing a wide range of tasks in a natural and life-like manner is a key challenge in enabling practic…

13:00 JSTLLM/生成AIGPT / ChatGPT

On the Nonlinearity of Learning Rate Scaling for LLM Training

Learning-rate transfer can reduce the cost of training large language models: instead of sweeping learning rates at target scale, practitio…

13:00 JST研究/論文

Invariant Reasoning Directions in Latent Trajectories of Language Models

Latent reasoning models perform multi-step inference directly in hidden-state space, yet the structure of these latent reasoning trajectori…

13:00 JST研究/論文

Projected Exploitability Descent for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games

Many important games have more than two players and imperfect information. Existing approaches for computing Nash equilibrium, the central…

13:00 JSTLLM/生成AILlama

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how t…

13:00 JST画像/動画生成

Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection

Detecting and localizing defects in 3D point clouds is challenging because abnormal samples are scarce and diverse, while training is often…

13:00 JSTLLM/生成AIエージェント研究/論文

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimo…

13:00 JSTロボティクス

Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering

Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise unde…

13:00 JSTロボティクス

AnyBody: Free-Form Whole-Body Humanoid Control from Arbitrary Keypoint Guidance

We present AnyBody, a unified whole-body humanoid controller driven by an arbitrary subset of body keypoints chosen at deploy time. Prior p…

13:00 JSTロボティクス

MoPe: Motion Permanence for Robust Monocular Gaussian Mapping in Dynamic Environments

Robust robot autonomy depends on scene representations that remain stable enough to support localization, navigation, and downstream decisi…

13:00 JST画像/動画生成

Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference

Online inspection images of final optics in high-power laser facilities contain pseudo-damage sites that closely resemble true damage sites…

13:00 JSTLLM/生成AI

A Hybrid Framework for Song Lyric Annotation Based on Human-LLM Alignment

Emotion recognition of song lyrics is a challenging task since lyrics may not necessarily align with the overall emotion of a song. As a re…

13:00 JSTLLM/生成AIエージェント

Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts

LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning

We identify intervention bias as a previously unquantified failure mode of zero-shot large-language-model (LLM) educational advisory agents…

13:00 JST研究/論文

Covering the Unseen: Information Demand Coverage Optimization for Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) typically treats context selection as ranking chunks against a single query embedding. This assumption…

13:00 JST研究/論文

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch betwee…

13:00 JST研究/論文

Adaptive Financial Transformer with Regime-Gated Attention for Stock Return Prediction

Adaptive Financial Transformer (AFT) is proposed for stock return prediction under non-stationary financial markets. The model incorporates…

13:00 JST画像/動画生成ロボティクス

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs

Vision-language models and vision-language action models endow the robot with unprecedented capabilities. However, the input of video and h…

13:00 JST画像/動画生成Qwen

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking

Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially bo…

13:00 JST研究/論文

Solver-Verified Formulation Generation and Selection for Multi-Warehouse Inventory Allocation Using Large Language Models

Balance-oriented multi-warehouse inventory allocation is a recurring decision problem in large-scale e-commerce supply chains, in which a f…

13:00 JST画像/動画生成

DR-GS: Physically-Based Deformable and Relightable 2D Gaussians

Gaussian splatting (GS) has garnered significant attention in VR/AR and digital content creation due to its explicit parameterization and e…

13:00 JST画像/動画生成

Learning to Adaptively Allocate Gaussians for Arbitrary-Scale Image Super-Resolution

In computer graphics, visual content is continuously warped, zoomed and resampled. This occurs when engines upscale frames, users zoom into…

13:00 JST研究/論文

Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery

Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional underco…

13:00 JSTLLM/生成AI

LC-ICL: Label-Guided Contrastive In-Context Learning for Robust Information Extraction

There has been increasing interest in exploring the capabilities of advanced large language models (LLMs) in the field of information extra…

13:00 JST画像/動画生成

Can Machines Really See Objects in Images? A Study Based on Syntactic Distance and Visual Self-Referential Instances

Can a vision model truly see an object, or does it only fit surface-level visual cues? Following Wittgenstein's view that the limits of lan…

13:00 JSTLLM/生成AIビジネス/資金調達

LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators

The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a…

13:00 JSTLLM/生成AILlamaMistral AI

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense pa…

13:00 JSTLLM/生成AI画像/動画生成エージェント

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achie…

13:00 JST画像/動画生成

Resonant Brane Splatting for Arbitrary-Scale Super-Resolution

Arbitrary-Scale Super-Resolution (ASR) reconstructs images at continuous magnification factors. Recent methods accelerate inference by repl…

13:00 JSTLLM/生成AIエージェント

Interpretable Inverse Design of Metal-Organic Frameworks with Large Language Model Agents

Inverse design of metal-organic frameworks (MOFs) requires searching a combinatorially vast space where property labels are expensive and m…

13:00 JST画像/動画生成

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can effic…

13:00 JST研究/論文

A Posteriori Error Analysis for Decoupled Neural Approximations of Fully Coupled FBSDEs with Control Mismatch

This paper develops an a posteriori error analysis framework for decoupled neural approximations of fully coupled forward--backward stochas…

13:00 JSTエージェント

CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher t…

13:00 JSTLLM/生成AI

To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Aggregation

While reinforcement learning (RL) significantly enhances LLM reasoning, its efficacy is severely undermined by Pre-RL data overlap, where R…

13:00 JSTLLM/生成AIGemma

Reported Confidence in LLMs Tracks Commitment More Than Correctness

Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence reports are widely used as uncertainty meas…

13:00 JSTLLM/生成AI

The Verbose Context Problem in Medical Records

The verbose context problem occurs when structured concepts have token-inefficient textual representations. This bottleneck is acute in pop…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models

Large Language Models (LLMs) are increasingly used as assistants across the software development lifecycle, yet their ability to reason abo…

13:00 JST画像/動画生成研究/論文

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2…

13:00 JST研究/論文

SemJoin: Semantic Join Optimization

Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and…

13:00 JSTエージェント

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing sk…

13:00 JSTLLM/生成AIGPT / ChatGPT

Em-ergence of the em-dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era

Large language models (LLMs) can leave subtle stylistic traces in assisted text; one of the most cited is the em-dash (Unicode U+2014). Yet…

13:00 JST研究/論文

Proteus: Automated Adversarial Robustness Testing for Audio Deepfake Detectors

We present Proteus, a framework developed at Resemble AI for automated robustness testing of our audio deepfake detection system. Given a d…

13:00 JSTロボティクス

VISTA-DZ: Visual Semantic Trajectory Adaptation for Personalized Dilemma Zone Prediction

Driver decision making in the dilemma zone at signalized intersections is safety critical, as vehicles approaching a yellow signal must dec…

13:00 JSTLLM/生成AI

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM

Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-conte…

13:00 JST研究/論文

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cos…

13:00 JST画像/動画生成

ReMAP-PET: Beyond Visual Understanding -- Learning Region-Guided Metabolic Alignment Semantics from Brain PET

Positron Emission Tomography (PET) reveals brain metabolism and is clinically central to neurodegenerative disease assessment, yet existing…

13:00 JST画像/動画生成

ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models

Spatial reasoning remains a persistent challenge for many vision language models (VLMs), and improving it typically requires fine-tuning wi…

13:00 JSTLLM/生成AIビジネス/資金調達

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluati…

13:00 JST研究/論文

Bilevel Optimization for Neural Architecture Search

Bilevel optimization has become an influential and widely adopted framework for addressing hierarchical optimization problems in machine le…

13:00 JST画像/動画生成

SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis

Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging…

13:00 JST研究/論文GPT / ChatGPTGemini

How AI settled the complexity of the oldest SGD algorithm

In 1937, Stefan Kaczmarz proposed a simple algorithm for solving systems of linear equations. This algorithm turned out to be the earliest…

13:00 JST画像/動画生成

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models

A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometri…

13:00 JSTエージェント

Langshaw: Declarative Interaction Protocols Based on Sayso and Conflict

Current languages for specifying multiagent protocols either over-constrain protocol enactments or complicate capturing their meanings. We…

13:00 JSTLLM/生成AIGPT / ChatGPT

Mechanistically Eliciting Latent Behaviors in Language Models

We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could…

13:00 JST研究/論文

Does Role Specialization Matter for Explanation Faithfulness in Mixture-of-Experts?

Mixture-of-Experts (MoE) architectures have recently been extended with role-based mechanisms for interpretability. This is typically done…

13:00 JSTLLM/生成AI

Do We Still Need Fine Tuning? Turkish Sentiment Analysis in the Era of Large Language Model

This study examines whether supervised fine-tuning remains necessary for Turkish sentiment analysis in the era of large language models. We…

13:00 JSTLLM/生成AI

Two-Stage Prompt Optimization for Few-Shot Relation Extraction: From Reasoning-Guided Search to Gradient-Guided Refinement

Automatic prompt optimization is still underexplored for episodic few-shot relation extraction with smaller language models. We propose a t…

13:00 JST研究/論文

Fast Wireless Foundation Models with Early-Exits

While wireless foundation models (FMs) are demonstrating strong potential to enable AI-Native 6G networks, their high computational cost re…

13:00 JSTエージェント

Fuzzing Large Language Models to Elicit Hidden Behaviours

Sleeper agents are the canonical model organism of deception: models trained to behave normally but to emit an unsafe behaviour on a specif…

13:00 JSTLLM/生成AIエージェント

Hybrid Retriever Evolution for Multimodal Document Reasoning Agents

Different retrievers, including lexical, semantic, and multimodal approaches, provide highly complementary strengths for multimodal documen…

13:00 JST画像/動画生成Gemini

Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature

The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inacc…

13:00 JST研究/論文Claude

A Machine-Verified Proof of a Quantum-Optimization Conjecture

We report a machine-verified resolution of a problem open for over a decade in quantum optimization: the Farhi, Goldstone and Gutmann (FGG)…

13:00 JST画像/動画生成

Early Warning Signals for OpenVLA Failure under Visual Distribution Shift

Vision Language Action models combine perception, language grounding, and control in a single policy, but their failures are hard to diagno…

13:00 JSTLLM/生成AI

ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering

Telecom question answering (QA) is a challenging setting for retrieval-augmented generation (RAG): evidence is fragmented across standards,…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today'…

13:00 JSTLLM/生成AIエージェント研究/論文

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of…

13:00 JSTLLM/生成AI

Diagnosing and Mitigating Context Rot in Long-horizon Search

Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon tasks. The concern that inc…

13:00 JST研究/論文

Redefining Maritime Anomaly Detection via Equation-Grounded Synthetic Anomalies

Maritime anomaly detection is essential for ensuring maritime safety, security, and efficient traffic management at sea, with Automatic Ide…

13:00 JST研究/論文

From Trait to Behavior: A Cognitive-Affective Personality System (CAPS) Perspective on Multi-Homing Intention in AIGC Platforms

With the rapid development of Artificial Intelligence Generated Content (AIGC) platforms, users increasingly show cross-platform usage inte…

13:00 JST研究/論文

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF

Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alter…

13:00 JST画像/動画生成エージェント

TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging

Topological data analysis (TDA), particularly persistent homology (PH), captures geometric structural properties in medical images (e.g., c…

13:00 JSTエージェント

Towards Generalizable and Evidential Nuclear Magnetic Resonance-Based Molecular Structure Elucidation via Large Language Model Agent

Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra fo…

13:00 JSTLLM/生成AIエージェント

Mandol: An Agglomerative Agent Memory System for Long-Term Conversations

Long-term conversational agents need to remember and query cross-session, multi-typed information with complex correlations. Existing agent…

13:00 JST画像/動画生成ロボティクス

FalconTrack: Photorealistic Auto-Labeled Perception and Physics-Aware Vision-Based Aerial Tracking

Vision-based aerial tracking is critical in GPS-denied environments. Reliable perception for tracking depends on large-scale labeled data,…

13:00 JSTハードウェア/半導体ビジネス/資金調達

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive…

13:00 JST研究/論文

What Drives the Inlier-Memorization Effect? A Theory of Outlier Detection via Early Training Dynamics

Outlier detection (OD) aims to identify anomalous instances by learning the underlying structure of normal data (inliers), and is particula…

13:00 JST研究/論文

Multi-Level Distributional Entropy for Explainable Network Intrusion Detection

Machine learning network intrusion detection systems (IDS) rely on aggregate flow statistics that discard distributional structure, while e…

13:00 JST研究/論文

Accelerating Q-learning through Efficient Value-Sharing across Actions

Action-values are foundational to many control algorithms such as Q-learning. Therefore learning action-values efficiently is central to re…

13:00 JSTLLM/生成AI研究/論文

Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework

Chart data extraction, which reverse-engineers data tables from chart images, is essential for reproducibility, analysis, retrieval, and re…

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文

How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation

Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depen…

13:00 JST研究/論文Google

Dual-Flow Reinforcement Learning with State-Aware Exploration

In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return dis…

13:00 JSTエージェント

Experience Graphs: The Data Foundation for Self-Improving Agents

The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We a…

13:00 JSTLLM/生成AIエージェント

Neural Procedural Memory: Empowering LLM Agents with Implicit Activation Steering

While Large Language Models (LLMs) excel as static solvers, transforming them into autonomous agents remains challenging. This transition r…

13:00 JSTLLM/生成AI

MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers

The quadratic computational cost of traditional attention mechanisms poses a major bottleneck to the scalability and practical deployment o…

13:00 JSTLLM/生成AI研究/論文

Exploring Motivations for Algorithm Mention in the Domain of Natural Language Processing: A Deep Learning Approach

With the rise of data-intensive science, algorithms have become central to scientific research. In academic papers, algorithms are mentione…

13:00 JST画像/動画生成

SUMO: Segment and Track Any Motion with Nonlinear State Space Models

Visual Object Tracking (VOT) and Moving Object Segmentation (MOS) are two fundamental tasks in computer vision that involve both spatial an…

13:00 JSTエージェントロボティクス研究/論文

RoAd-RL: A Unified Library and Benchmark for Robust Adversarial Reinforcement Learning

Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversaria…

13:00 JSTLLM/生成AI

ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation

Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective o…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot disting…

13:00 JST画像/動画生成エージェント

LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving

Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD)…

13:00 JSTロボティクス

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models

Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning. How…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databa…

13:00 JST研究/論文

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models

Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization…

13:00 JSTロボティクスビジネス/資金調達

Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies

Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challeng…

13:00 JSTLLM/生成AI画像/動画生成

LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion

Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in moder…

13:00 JST研究/論文

Semi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss

Sound event detection (SED) is a core module for acoustic environmental analysis, yet its performance is often limited by scarce labeled da…

13:00 JST研究/論文

CW-B: Class Weighted Boosting Framework for Imbalance Resilient Multi Class Cardiac Phenotyping

Cardiac discharge phenotyping informs post-discharge treatment and follow-up, but real-world records are often incomplete and class-imbalan…

13:00 JSTロボティクス

Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation

Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from tra…

13:00 JSTエージェントGPT / ChatGPT

EVAF: A Test-Retest Protocol for Selective Parametric Consolidation

Long-running language agents need mechanisms for deciding which experiences should persist after the working context is gone. Retrieval sys…

13:00 JST画像/動画生成

Latent-CURE for Breast Cancer Diagnosis

Multimodal Large Models have significantly advanced automated breast ultrasound diagnosis. However, most existing frameworks utilize opaque…

13:00 JST研究/論文

Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction

H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-se…

13:00 JST画像/動画生成

Exploiting Local Flatness for Efficient Out-of-Distribution Detection

Detecting out-of-distribution (OOD) data is crucial for reliable machine learning deployment. Among detection strategies, post-hoc methods…

13:00 JSTエージェント

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet…

13:00 JSTエージェント研究/論文

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real…

13:00 JSTLLM/生成AIエージェント

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this a…

13:00 JST研究/論文NVIDIA

RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines

Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entang…

13:00 JST研究/論文

Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping

Looped Transformers, which repeatedly apply a shared transformer block, are an architecturally natural fit for variable-length algorithmic…

13:00 JST研究/論文

T3R: Deeper Test-Time Adaptation for Graph Neural Networks via Gradient Rotation

Graph Neural Networks (GNNs) deployed in real-world systems typically have fixed weights, often leading to degraded performance under distr…

13:00 JST画像/動画生成

IBRSteG: Learning a Generalizable Steganography Framework for 3D Gaussian Splatting

Recent advances in deep learning have notably improved steganographic message hiding. However, designing a generalizable steganographic app…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic…

13:00 JSTLLM/生成AI研究/論文

Little Brains, Big Feats: Exploring Compact Language Models

While large language models have been dominating the research landscape recently, small language models remain highly relevant across vario…

13:00 JST画像/動画生成

Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management

We introduce Neural Subspace Reallocation (NSR), which reframes continual learning as memory management over parameter subspaces. Instead o…

13:00 JSTLLM/生成AI

Online Data Selection for Instruction Tuning via Gaussian Processes

With Large Language Model (LLM) pre-training and fine-tuning shifting its focus from data volume to data quality, quality data selection ha…

13:00 JSTエージェントロボティクス

Automating the Design of Embodied AgentArchitectures

Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity expo…

13:00 JSTロボティクス

SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance

Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actio…

13:00 JST研究/論文

Gravitational Duals from Equations of State II: Large Hierarchies and False Vacua

We investigate the reconstruction of holographic duals for strongly coupled quantum field theories in regimes characterized by large hierar…

13:00 JST画像/動画生成

Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching

Functional maps are the cornerstone of recent non-rigid 3D shape matching methods due to their efficiency and performance. However, existin…

13:00 JST研究/論文

Query-Aware Spreading Activation for Multi-Hop Retrieval over Knowledge Graphs

Retrieval-augmented generation built on knowledge graphs (Graph RAG) outperforms flat passage retrieval on multi-hop question answering by…

13:00 JSTLLM/生成AI

Estimating Grammatical Gender Directions in Contextual Embeddings under Controlled and Natural Contexts

Contextual language models conflate grammatical gender and social semantic bias in gendered languages such as Spanish. Existing gender debi…

13:00 JST研究/論文

Physically-Constrained Harmonic Separation for Robust Heart and Respiratory Rate Estimation from Wrist Photoplethysmography

Wrist-worn photoplethysmography (PPG) enables continuous monitoring of cardiopulmonary physiology, but reliable heart rate (HR) and respira…

13:00 JST研究/論文

Federated Learning with Energy-Based Structured Probabilistic Inference

Federated learning typically aggregates client updates using fixed or heuristic weighting rules, which can be suboptimal when clients have…

13:00 JST研究/論文

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical dat…

13:00 JST画像/動画生成

Few-Shot Domain Incremental Learning via Continual Vision-Language Consolidation

Existing domain-incremental learning (DIL) strategies call for massive amounts of data to adapt to new domains and suffer from the overfitt…

13:00 JSTLLM/生成AI研究/論文

Forewarned is Forearmed: When Non-Sequential Embedding Turns Into an Anomaly Detector

This paper offers an in-depth analysis of non-sequential multimodal sentence-level embeddings, with a particular focus on the SONAR model.…

13:00 JST画像/動画生成

A Multi Center Breast FNAC Whole-Slide Cytology Dataset for AI-Assisted Patch-Wise Classification Using C1 to C5 Reporting Categories

We present a multi center breast fine needle aspiration cytology (FNAC) dataset designed for patch wise classification using C1 to C5 repor…

13:00 JST画像/動画生成

Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion

RGB-T detectors leverage the complementary strengths of visible and thermal infrared modalities, achieving robust performance under challen…

13:00 JST研究/論文

Curvature-Guided Sheaf Diffusion for Unsupervised Community Detection on Heterophilic Graphs

Detecting communities in heterophilic graphs -- where connected nodes often belong to different classes -- is hard for unsupervised methods…

13:00 JST研究/論文

KnowsTFM: Knowledge-Informed Fine-Tuning of Small Tabular Foundation Models

Tabular foundation models have advanced deep learning for tabular data by delivering strong default performance across many small and mediu…

13:00 JST研究/論文

Defending Against Harmful Supervision Hidden in Benign Samples

Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead…

13:00 JSTエージェント

Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation

Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it fr…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTMistral AI

Research Entity Extraction and Topic Detection from UKRI Grant Proposals

This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and…

13:00 JSTLLM/生成AIエージェント

Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents

Always-on agents are systems whose future behavior depends on durable state accumulated across earlier interactions. We treat them as persi…

13:00 JSTLLM/生成AIAnthropicClaude

MCP Server Architecture Patterns for LLM-Integrated Applications

The Model Context Protocol (MCP), introduced by Anthropic in November 2024, defines a standardized interface for connecting large language…

13:00 JST画像/動画生成研究/論文

Early Cue Precision Shapes Visual Shortcut Learning in Controlled Cue-Manipulation Benchmarks

Visual classifiers can achieve high matched-distribution accuracy while relying on low-level cues that fail under conflict or suppression.…

13:00 JST研究/論文

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in comple…

13:00 JST画像/動画生成

FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images

We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from on…

13:00 JST画像/動画生成

Residual-Guided Expert Specialization for Incomplete Multimodal Learning

As real-world prediction systems often face missing modalities at inference, incomplete multimodal learning (IML) remains a practical chall…

13:00 JST画像/動画生成ロボティクス

ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control

While current Behavior Foundation Models (BFMs) provide robust control priors for humanoids, they only execute pre-defined reference motion…

13:00 JST画像/動画生成

Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation

Multimodal MRI is essential for accurate brain tumor segmentation. However, acquiring all modalities at inference is often challenging in p…

13:00 JST研究/論文

A Stochastic--Geometric Theory of Scaling Laws in Grokking

Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only…

13:00 JST研究/論文

Model Predictive Current Control with Harmonic Correction for Single-Phase AC-DC EV Charging

The increasing integration of Electric Vehicles (EVs) has imposed a growing harmonic challenge on the power grid. For AC/DC Power Factor Co…

13:00 JST研究/論文

Beyond IID: How General Are Tabular Foundation Models, Really?

Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Resea…

13:00 JSTLLM/生成AI

Can LLMs Rank? A Tale of Triads and Triage

From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered…

13:00 JST画像/動画生成

Beyond Point Estimates for Glaucoma Visual Field Forecasting with Diffusion Models

Forecasting visual fields (VFs) is critical for personalized monitoring and treatment planning in glaucoma. This is inherently uncertain du…

13:00 JST研究/論文

Transformer Architectures as Complete Bayes Processes: A Formal Proof in the Measure-Theoretic Kernel Framework

We present a complete formal proof that transformer architectures, when their internal update mechanisms satisfy a Bayes joint-distribution…

13:00 JSTLLM/生成AIエージェントLlama

Translating Natural Language to Strategic Temporal Specifications via LLMs

A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS). However, w…

13:00 JSTLLM/生成AIエージェント

Collective cooperation without individual fidelity in LLM agents

Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior c…

13:00 JSTLLM/生成AI

Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval

We study retrieval over catalogs of structured metadata, where each record is a small schema whose fields answer different kinds of query.…

13:00 JST研究/論文

COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies

Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to…

13:00 JSTLLM/生成AI

Situation Perception: A Necessary Primitive to Artificial Superintelligence

Current large language models are extraordinary statistical engines. They compress vast amounts of text into useful patterns and can explai…

13:00 JSTLLM/生成AI

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

Background. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dialogues. Human codin…

13:00 JST研究/論文

McMg: A Learned Phase-Space Multi-channel Multigrid Preconditioner for Helmholtz Equation

Solving heterogeneous Helmholtz equations at high wavenumbers remains challenging because the discretized operator is indefinite, pollution…

13:00 JST画像/動画生成

On the Faithfulness of Post-Hoc Concept Bottleneck Models

Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color. To bridge the gap be…

13:00 JST研究/論文

Informational Frustration in Neural Manifolds: Shannon Bottlenecks and the Limits of Learnability

Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theor…

13:00 JST画像/動画生成エージェントロボティクス

Learning from Mistakes: Rollout-Retrieval Lifelong Policy Learning for Autonomous Driving

Autonomous driving policies should be able to improve continually as deployment exposes them to increasingly diverse and long-tail traffic…

13:00 JSTLLM/生成AIエージェント

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly impor…

13:00 JSTLLM/生成AICopilot

To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks

AI code completion tools, such as Github Copilot, provide students with code suggestions to help them write programs. However, recent quali…

13:00 JSTLLM/生成AIエージェントClaude

TraceLab: Characterizing Coding Agent Workloads for LLM Serving

Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this…

13:00 JST研究/論文

A Multi-task Mixture of Experts Framework for Malware Classification, Packing Detection, and Family Attribution

Malware classification remains a challenging problem due to its inherent heterogeneity, the presence of packed binaries, and the diverse di…

13:00 JST画像/動画生成ロボティクス

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged ref…

13:00 JSTLLM/生成AI研究/論文

Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection

Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection. Recent work has shown…

13:00 JSTエージェントGemmaLlamaQwen

MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems

Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows. However, their inter-agent communication channe…

13:00 JST研究/論文

C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable…

13:00 JST研究/論文

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms

Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effecti…

13:00 JST研究/論文

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-su…

13:00 JST画像/動画生成ロボティクス

GROW$^2$: Grounding Which and Where for Robot Tool Use

Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to use tools creatively…

13:00 JST研究/論文

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics an…

13:00 JSTロボティクス

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learnin…

13:00 JST研究/論文

A binarized-domains arc-consistency algorithm for TCSPs: its computational analysis and its use as a filtering procedure in solution search algorithms

TCSPs (Temporal Constraint Satisfaction Problems) [Dechter et al. 1991] get rid of unary constraints by binarizing them after having added…

13:00 JSTエージェント

Modelling Human Values for Value-Aware Multi-Agent Systems

One of today's most pressing societal challenges is building AI systems whose behaviour, or the behaviour it enables within communities of…

13:00 JST研究/論文

Instance-Conditioned Adaptation for Large-scale Generalization of Neural Routing Solver

In modern intelligent transportation systems (ITS), particularly in freight transportation and logistics, real-time route planning is cruci…

13:00 JSTLLM/生成AIロボティクス

CLMASP: Coupling Large Language Models with Answer Set Programming for Robotic Task Planning

Large Language Models (LLMs) possess extensive foundational knowledge and moderate reasoning abilities, making them suitable for general ta…

13:00 JST研究/論文

OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale

Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare. However, most such problems are still sol…

13:00 JST研究/論文

MARS: A neurosymbolic approach for interpretable drug discovery

Background: Neurosymbolic (NeSy) artificial intelligence describes the combination of logic or rule-based techniques with neural networks.…

13:00 JSTLLM/生成AIエージェント

LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes. Such models are…

13:00 JST研究/論文

GraphChase: A Platform and Benchmark for Urban Network Security Games

After the achievement of solving two-player zero-sum games, more AI researchers focus on solving multiplayer games. Urban Network Security…

13:00 JSTLLM/生成AIGemini

Accelerating scientific discovery with Co-Scientist

Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation…

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely e…

13:00 JST研究/論文

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning

Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and…

13:00 JSTLLM/生成AI

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning

Tool-integrated reasoning (TIR) has become a key approach for improving large reasoning models (LRMs) on complex problems. Prior work has m…

13:00 JSTLLM/生成AIエージェント研究/論文GeminiLlama

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. P…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達GPT / ChatGPT

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and ada…

13:00 JST研究/論文

DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing

From a Bayesian perspective, score-based diffusion solves inverse problems through joint inference, embedding the likelihood with the prior…

13:00 JSTエージェント

Agentic AI for ISAC: Analysis, Framework, and Case Study

Integrated sensing and communication (ISAC) has emerged as a key development direction in the sixth-generation (6G) era, which provides ess…

13:00 JSTエージェント

Monte Carlo Query Search: Active Capability Assessment of AI Agents

Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requi…

13:00 JSTLLM/生成AIエージェント

CaveAgent: Transforming LLMs into Stateful Runtime Operators

LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradig…

13:00 JSTエージェント

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models

Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step r…

13:00 JST研究/論文

When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning

When a model knows when it does not know, many possibilities emerge. The first question is how to enable a model to recognize that it does…

13:00 JST研究/論文DeepSeek

Inference-Time Diversity in RL-Trained Lean Theorem Provers: A Diagnostic Study

RL-trained Lean theorem provers mode-collapse at inference time: on miniF2F-test with DeepSeek-Prover-V1.5-RL, doubling the i.i.d.\ samplin…

13:00 JSTLLM/生成AI

Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement

Large language models (LLMs) exhibit social biases that reinforce harmful stereotypes, limiting their safe deployment. Most existing debias…

13:00 JSTLLM/生成AI研究/論文

Aligning Language Model Benchmarks with Pairwise Preferences

Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find t…

13:00 JSTLLM/生成AIエージェント

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completi…

13:00 JSTLLM/生成AI

StackingNet: Collective Inference Across Independent AI Foundation Models

Artificial intelligence built on large foundation models has transformed language understanding, computer vision, and reasoning, yet these…

13:00 JST研究/論文

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, bench…

13:00 JSTエージェント

AutoB2G: Agentic Simulation and Reinforcement Learning for Spatio-Temporal Grid-Interactive Building Control

Grid-interactive building control has emerged as a promising approach for improving demand-side flexibility in modern power systems. Realis…

13:00 JSTLLM/生成AIエージェント研究/論文

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientifi…

13:00 JSTエージェント研究/論文

QED: 未解決の問題に対する数学的証明を生成するためのオープンソース マルチエージェント システム

私たちは \textbf{QED} を紹介します。これは、人間が提供した研究質問を、人間によるさらなる指導なしに完全な数学的証明に変えるオープンソースのマルチエージェント システムです。そのパイプラインは、計画、証明、検証を分離することで、単一クエリの証明生成でよくある失敗を克服するように設計されています。つまり、分解エージェントが証明検索を構造化し、証明者エージェントが候補引数を生成し、検証者エージェントが正しさをチェックします。分野の専門家と協力して、さまざまな難易度の 18 の研究レベルのプロジェクトについて QED を評価しました。 QED は、代数幾何学、流体偏微分方程式、確率、逆問題にわたる 5 つのオリジナル作品を作成しました。専門家の評価では、これらの著作は堅実な専門研究の貢献であるとみなされており、3 つの著作は、確立された数学の専門分野で一般的に出版されている著作と難易度および範囲において同等です。 QED は https://github.com/proofQED/QED でリリースされます。

原文 (English)

QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems

We present QED, an open-source multi-agent system that turns human-provided research questions into complete mathematical proofs without further human guidance. Its pipeline is designed to overcome common failures of single-query proof generation by separating planning, proving, and verification: a decomposition agent structures the proof search, prover agents generate candidate arguments, and verifier agents check correctness. In collaboration with domain experts, we evaluated QED on 18 research-level projects of varying difficulty. QED produced five original works across algebraic geometry, fluid PDEs, probability, and inverse problems. Expert assessments regard these works as solid specialized research contributions, with three comparable in difficulty and scope to work commonly published in established specialist mathematics venues. QED is released at https://github.com/proofQED/QED.

13:00 JSTLLM/生成AIエージェント研究/論文

Exploring LLM Agent Designs and Interaction Modalities for Scientific Visualization

This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visual…

13:00 JST研究/論文

To Use AI as Dice of Possibilities with Timing Computation

The dominant noun-based modeling paradigm, grounded in probability theory and committed to pre-specified noun entities as primitive modelin…

13:00 JST研究/論文

NEURON: A Neuro-symbolic System for Grounded Clinical Explainability

Clinical AI adoption is hindered by the black-box/grey-box nature of high-performing models, which lack the ontological grounding and narra…

13:00 JSTLLM/生成AI

SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks

Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. Th…

13:00 JSTエージェントClaudeGPT / ChatGPT

Towards Human-Level Book-Writing Capability

Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality…

13:00 JSTLLM/生成AIGPT / ChatGPTGemma

Geo-Expert: パラメーター効率の高い微調整によるエキスパートレベルの地質学的推論に向けて

地質学に適用される汎用の大規模言語モデル (LLM) は、地下構造や深層時間の進化について推論する際に幻覚を起こすことがよくありますが、現在の地球科学における AI は主に地表のリモート センシングと GIS を対象としています。このギャップを埋めるために、カスタム命令合成パイプラインを使用して処理された、カスタムで厳選された高品質の命令データセットに基づいて微調整された、パラメーター効率の高い地質 LLM ファミリーである Geo-Expert を導入します。低ランク適応 (LoRA) 手法を使用して、Qwen3-8B、Qwen3-32B、Gemma-3-27B の 3 つのベース モデルを微調整することにより、モデルのスケーリングとアーキテクチャの影響を調査します。新しいドメイン固有のベンチマークである Geo-Eval に関する広範な評価により、ドメイン整合 8B モデルは特殊な地質学的推論においてオープンウェイト 70B ジェネラリストや独自の GPT-4o よりも優れたパフォーマンスを発揮できる一方、32B バリアントはフロンティア推論モデルに近づくことが明らかになりました。最適化された 8B モデルは、導入において競争力のあるコストパフォーマンス比をさらに提供します。この研究は、科学的 LLM を民主化するための再現可能なレシピを提供し、地質学的人工知能のベースラインを確立します。

原文 (English)

Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning

While general-purpose Large Language Models (LLMs) applied to Geology often hallucinate when reasoning about subsurface structures and deep-time evolution, current AI in Earth sciences predominantly targets surface remote sensing and GIS. To bridge this gap, we introduce Geo-Expert, a family of parameter-efficient geological LLMs fine-tuned on a custom-curated, high-quality instruction dataset processed using our custom instruction synthesis pipeline. We investigate the impact of model scaling and architecture by fine-tuning three base models: Qwen3-8B, Qwen3-32B, and Gemma-3-27B, with Low-Rank Adaptation (LoRA) method. Our extensive evaluation on a novel domain-specific benchmark, Geo-Eval, reveals that a domain-aligned 8B model can outperform open-weight 70B generalists and proprietary GPT-4o on specialized geological reasoning, while a 32B variant approaches frontier reasoning models. The optimized 8B model further offers a competitive cost-performance ratio for deployment. This work provides a reproducible recipe for democratizing scientific LLMs and establishes a baseline for geological artificial intelligence.

13:00 JSTLLM/生成AI

Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation

Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language mod…

13:00 JSTLLM/生成AI

大規模な言語モデルに対する有効な事実性制御を備えた推論時の正形推論

大規模言語モデル (LLM) は、多段階推論を実行することが増えており、中間クレームが暗黙的な有向非巡回グラフを形成し、そのノードの正しさがその先祖に構造的に条件付けされます。これにより、事実の不確実性がノードごとの些細なエラーの蓄積ではなく構造的なものとなり、推論構造に対する推論時間の不確実性の定量化が必要になります。等角予測 (CP) は柔軟なユーザー指定の事実性制御を提供しますが、既存の作業は事後的なままであり、生成中に介入することはできません。 CP の柔軟性と事後的な制限の間のギャップを埋めるために、CP を推論グラフ生成に直接統合する \emph{推論時等形推論 (ITCR)} フレームワークを提案します。 ITCR は、複雑なモデリング仮定を使用せずに、推論グラフ上でクレーム レベルの事実シグナルを集約する、構造レベルの事実不確実性関数を学習します。次に、グラフレベルの事実の不確実性に基づいて不適合スコアを設計し、生成をいつ停止するかを決定するための等角しきい値を調整します。我々は、そのような生成が入れ子になっていることを理論的に示し、事実性制御の有効なカバレッジ保証をもたらします。複数のデータセットとカバレッジ目標に対する実験により、経験的に有効なカバレッジが実証されます。下流の推論タスクでは、推論時に調整されたグラフの方が、ポストホックに枝刈りされたグラフよりも正確に生成されます。

原文 (English)

Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models

Large language models (LLMs) increasingly perform multi-step reasoning, where intermediate claims form implicit directed acyclic graphs whose node correctness is structurally conditioned on their ancestors. This makes factuality uncertainty structural, rather than a trivial accumulation of node-wise errors, and necessitates inference-time uncertainty quantification over the reasoning structure. While conformal prediction (CP) offers flexible user-specified factuality control, existing work remains post-hoc and cannot intervene during generation. To fill the gap between CP's flexibility and its post-hoc limitation, we propose an \emph{Inference-Time Conformal Reasoning (ITCR)} framework that integrates CP directly into reasoning graph generation. ITCR learns a structure-level factuality uncertainty function that aggregates claim-level factuality signals over reasoning graphs without complex modeling assumptions. We then design the non-conformity score based on graph-level factuality uncertainty and calibrate the conformal threshold to decide when to stop generation. We theoretically show such generation is nested, yielding valid coverage guarantees for factuality control. Experiments over multiple datasets and coverage objectives demonstrate empirically valid coverage. In downstream reasoning tasks, inference-time calibrated graphs yield more accurate generation than post-hoc pruned graphs.

13:00 JSTエージェント

ビジネス世界モデル

生産性の向上、コストの削減、製品とサービスの強化を目的として、企業は AI 対応ツールをますます導入しています。ただし、AI の変革の可能性は、事前定義されたタスクの自動化を超えて広がります。それは、インテリジェント システムが高レベルの戦略目標に基づいてビジネス イニシアチブを計画、最適化、実行できるようにすることにあります。このペーパーでは、ビジネスおよび組織環境に特化したワールド モデルであるビジネス ワールド モデル (BWM) の概念とアーキテクチャを紹介します。人工知能、認知科学、制御理論の世界モデルからインスピレーションを得た BWM は、ビジネスの状態、ダイナミクス、制約、目標、実行可能なアクション空間をエンコードして、自律的な意思決定をサポートします。私たちは、ビジネスの状態、ダイナミクス、およびアクションが主要なビジネス エンティティに関連付けられる、ビジネス セマンティクス中心の定式化を提案します。このフレームワーク内で、エージェントは代替アクションのシーケンスをシミュレートし、将来のビジネス成果に対するそれらの影響を推定し、不確実性の下でのトレードオフを評価できます。提案されたアーキテクチャは、意味論的なデータ表現、確率論的な機械学習モデル、決定論的なビジネス ルール、および明示的なアクション空間を、計画と反事実推論のための一貫した構造に統合します。 BWM の個々のコンポーネントは新しいものではありませんが、BWM の貢献は、それらをビジネス イニシアチブのための実行可能な内部シミュレーターとして組織化することにあります。この作業により、指示ベースの実行から目標主導の計画と実行に移行できる自律的なビジネス システムの概念基盤が確立されます。

原文 (English)

Business World Model

World model has emerged as a powerful paradigm in artificial intelligence, enabling agents to represent their environments, predict future states, and evaluate possible actions before acting. However, existing world model approaches have largely been developed for domains such as computer vision, robotics, gaming, and autonomous driving, where the world is primarily visual or physical and governed by relatively stable dynamics. These formulations are not directly applicable to business practice, where the relevant environment is semantic, organizational, and market-driven rather than physical. Business outcomes depend on context-sensitive factors such as customer behavior, pricing, competition, regulation, resources, and operational constraints. This paper introduces the concept and architecture of a Business World Model (BWM), which is a world model specialized for business and organizational environments. A BWM encodes business states, dynamics, and feasible actions space to support autonomous business planning and decision-making. We propose a business-semantics-centric formulation in which states, dynamics, and actions are linked to key business entities, their attributes, and their relationships. Within this framework, intelligent agents can simulate alternative action sequences, estimate their effects on future business outcomes, and evaluate trade-offs under uncertainty. The proposed architecture integrates semantic data representations, probabilistic machine learning models, deterministic business rules, and explicit action spaces into a coherent internal simulator. This work establishes a conceptual foundation for autonomous business systems capable of moving from instruction-based execution toward goal-driven planning, optimization, and execution.

13:00 JSTLLM/生成AIエージェント研究/論文

Agents-K1: エージェントネイティブのナレッジオーケストレーションに向けて

現在の LLM ベースの研究エージェントは、エージェント オーケストレーションを通じて進歩していますが、科学的知識のオーケストレーションはほとんど見落とされています。既存の著作物は、論文を要約、表面的な言及、平坦な \texttt{引用} エッジに縮小することが多く、科学的推論に不可欠な重要な実体、主張、証拠、メカニズム、および方法系統を省略しています。この目的を達成するために、生のドキュメントをエージェントネイティブの科学知識グラフに変換するエンドツーエンドの知識オーケストレーション パイプラインである \textbf{Agents-K1} を導入します。 Agents-K1 は、統一された理論的基盤の下に 3 つのコンポーネントを統合します。マルチモーダル パーサーでは、5 つのモジュールのスキーマがエンティティ、マルチモーダルな証拠、引用、および抄録のみではなく論文全体にわたる型付けされたエンティティ間の関係をキャプチャします。ルールベースの報酬の下で GRPO でトレーニングされた 4B 情報抽出バックボーン。もう 1 つは、Web 検索、マルチモーダル グラフ検索、およびクロスドキュメント トラバーサルを統合する 3 つのソース エージェント インターフェイスである、graphanything CLI です。これに加えて、6 つの主題にわたる 246 万件の科学論文を処理して \textbf{Scholar-KG} を生成し、そのうち 100 万件の論文サブセットをリリースしています。完全な Scholar-KG は以下の SCP リンクからアクセスできます。同じパイプラインを一般ドメインのコーパスとスキーマ準拠のデータ合成に拡張できます。広範な実験により、Agents-K1 が科学情報の抽出、ナレッジ グラフの構築、およびマルチホップ科学的推論において優れたパフォーマンスを達成することが実証されました。

原文 (English)

Agents-K1: Towards Agent-native Knowledge Orchestration

Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions, and flat \texttt{cites} edges, omitting key entities, claims, evidence, mechanisms, and method lineages essential for scientific reasoning. To this end, we introduce \textbf{Agents-K1}, an end-to-end knowledge orchestration pipeline that converts raw documents into agent-native scientific knowledge graphs. Agents-K1 integrates three components under a unifying theoretical foundation: a multimodal parser whose five-module schema captures entities, multimodal evidence, citations, and typed inter-entity relations across the full paper rather than abstracts alone; a 4B information-extraction backbone trained with GRPO under a rule-based reward; and a graphanything CLI, a tri-source agent interface that unifies web search, multimodal graph retrieval, and cross-document traversal. On top of this, we process 2.46 million scientific papers across six subjects to produce \textbf{Scholar-KG}, of which we release a one-million-paper subset, and the full Scholar-KG is accessible via the SCP link below. The same pipeline can be extended to general-domain corpora and to schema-conformant data synthesis. Extensive experiments demonstrate that Agents-K1 achieves superior performance in scientific information extraction, knowledge graph construction, and multi-hop scientific reasoning.

13:00 JST研究/論文

人工知能インデックスレポート 2026

AI Index レポートの第 9 版へようこそ。 AI が急速に進歩し続けるにつれて、AI を中心に構築されたシステムが追いつくことができるかどうかが問題になります。 AI の影響を追跡するために必要なガバナンスの枠組み、評価方法、教育システム、データ インフラストラクチャは、テクノロジー自体のペースに追いつくのに苦労しています。 AI ができることと、AI を管理するための私たちの準備との間のギャップは、今年のレポートのすべての章に貫かれています。この版の新たなレポートでは、推論、安全性、現実世界のタスク実行にわたって AI がどのようにより野心的にテストされているか、またそれらの測定値に依存することがますます困難になっている理由を追跡しています。また、生成型 AI の経済的価値の新しい推定値と、その労働市場への影響に関する新たな証拠、AI の主権に関する分析フレームワーク、および Schmidt Sciences と協力して開発された科学の章も取り上げられています。このレポートでは初めて、科学における AI と医学における AI に関する独立した章が設けられ、これら 2 つの領域にわたる AI の影響力の増大を反映しています。

原文 (English)

Artificial Intelligence Index Report 2026

Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built around it can keep up. Governance frameworks, evaluation methods, education systems, and the data infrastructure needed to track AI's impact are struggling to match the pace of the technology itself. That gap between what AI can do and how prepared we are to manage it runs through every chapter of this year's report. New in this edition, the report tracks how AI is being tested more ambitiously across reasoning, safety, and real-world task execution, and why those measurements are increasingly difficult to rely on. It also features new estimates of generative AI's economic value alongside emerging evidence of its labor market effects, an analytical framework on AI sovereignty, and a science chapter developed in collaboration with Schmidt Sciences. For the first time, the report features standalone chapters on AI in science and AI in medicine, reflecting AI's growing impact across these two domains.

13:00 JST研究/論文

RoboPIN: 固定された思考連鎖によるグラウンディングされた身体的推論

身体化された推論では、モデルが物理環境内のタスクに関連するオブジェクトや空間を認識し、複数ステップの推論を通じて一貫した視覚的根拠を維持する必要があります。しかし、現在の視覚言語モデルはテキストのみ、または座標拡張された思考連鎖に依存しており、実体参照は暗黙的かつ曖昧なままです。これにより、推論プロセスが視覚的な証拠から切り離され、エンティティ参照がステップ間で漂流し、推論の軌跡と最終的な答えとの間に因果関係の断絶が生じる可能性があり、これらの問題は、ビュー間の外観の変化によりマルチビュー シナリオでさらに増幅されます。これらの問題に対処するために、すべての推論ステップを視覚的な証拠に固定する構造化推論パラダイムである Pinned Chain-of-Thought (\pincot{}) を提案します。 \pincot{} は \reasoninganchor{} の概念を導入しています。これは、タスクに関連する各エンティティを、エンティティ名、一意の ID、ビュー インデックス、空間基盤を備えた構造化されたビジュアル アンカーにバインドし、推論ステップとビュー全体で一貫したエンティティの追跡を可能にします。完全に自動化されたデータ生成パイプラインを構築して、高品質の \pincot{} 形式の推論データセットである \dataset{} を構築します。次に、具体化された知識、構造化された推論能力、プロセス監視された調整を段階的に注入する 3 段階のポストトレーニングを通じて、\method{} をトレーニングします。報酬は、推論中のアンカーの位置特定とアイデンティティの一貫性の両方を直接制約します。埋め込まれた空間推論、マルチビュー推論、ポインティングをカバーする 14 のベンチマークでは、パラメーターが 4B のみの \method{} は常に 7B レベルのオープンソースの埋め込みモデルを上回り、最も強力な 7B ベースラインである Mimo-Embodied に対して平均 12\% の改善を達成しました。さらに分析すると、\pincot{} によって接地精度とステップ間の同一性の一貫性が向上し、プロセス監視の有効性が検証されたことが示されています。

原文 (English)

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought

Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models rely on text-only or coordinate-augmented chain-of-thought, where entity references remain implicit and ambiguous. This may cause the reasoning process to decouple from visual evidence, entity references to drift across steps, and a causal disconnection between the reasoning trajectory and the final answer, with these problems further amplified in multi-view scenarios due to cross-view appearance changes. To address these issues, we propose Pinned Chain-of-Thought (PinCoT), a structured reasoning paradigm that pins every reasoning step to visual evidence. PinCoT introduces the concept of reasoning anchor, which binds each task-relevant entity to a structured visual anchor with entity name, unique identity, view index, and spatial grounding, enabling consistent entity tracking across reasoning steps and views. We build a fully automated data generation pipeline to construct PIN-170K, a high-quality PinCoT-formatted reasoning dataset. We then train RoboPIN through three-stage post-training that progressively injects embodied knowledge, structured reasoning ability, and process-supervised alignment, with rewards that directly constrain both anchor localization and identity consistency during reasoning. On 14 benchmarks covering embodied spatial reasoning, multi-view reasoning, and pointing, RoboPIN with only 4B parameters consistently outperforms 7B level open-source embodied models, achieving a 12% average improvement over the strongest 7B baseline, Mimo-Embodied. Further analysis shows that PinCoT improves grounding accuracy and cross-step identity consistency, validating the effectiveness of process supervision.

13:00 JSTLLM/生成AIエージェント

見ることは選ぶことではない: LLM エージェントにおけるツール選択の失敗に関する注意セグメントの説明

LLM エージェントはツールを誤って呼び出します。混雑したハーネスの中でモデルが適切なツールを認識できなかったのではないかと自然に推測されます。並行作業を脇に置くというレンズを通して、ラベル付けされたツール定義セグメントに対するモデルの注意を介して反対のことを示します。実際の BFCL の失敗では、候補ごとの注意力 argmax により、モデルは 80% の確率で正しいツールに最も注目し (確率 21% に対して)、ゴールドは 10% のみで注目度の低いセグメントです。つまり、正しいツールを参照しているにもかかわらず、間違ったツールを選択します。これは、直観的な「ハーネスが混雑している / 途中で迷っている」という説明に真っ向から反論します。障害はハーネスではなく判定読み出しにあり、そこに 3 つの方法で問題を固定します。 (1) 入力と読み出し: プロンプトの修復 (ゴールド ツールの並べ替えまたは複製) では失敗の 23% 未満が回復しますが、読み出し側の介入では 59 ~ 91% が回復します。 (2) 表現不変性: 異なる表現での 2 つの金策介入 (加算的アテンション ロジット バイアスと残差ストリーム ステアリング ベクトル) は、ほぼ同じ失敗 (タスクごとに Jaccard プールされた 0.865、モデルごとに 0.79 ~ 0.91) を回復するため、どの表現がポークされるかに関係なく、ボトルネックは読み出しに局所化されます。 (3) トレーニング不要、ゴールドフリーのセレクター: セグメントごとの注目により、BFCL でのゴールドフリーとオラクルのギャップ (プールされた関数名の選択が +11.9 ポイント、オラクルのヘッドルームが +17.9 ポイント) のほとんどが縮まり、Seal-Tools では +14.9 ポイントが追加されます。すべてのモデルが陽性 (正確なマクネマー p<=8e-4 それぞれ)。範囲は異なります。因果関係の注意バイアスの用量反応は、10 個のマスク尊重モデル (3-32B) で双方向かつ単調であり、0.5-32B の全スパンは相関診断のみを保持します。デプロイ可能なセレクターは 5 つのシングル ターン モデルで評価されており、マルチ ターン ループにはまだ転送されていません。

原文 (English)

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a lens concurrent work sets aside -- the model's attention to labeled tool-definition segments. On real BFCL failures, by per-candidate attention argmax the model attends most to the correct tool 80% of the time (vs. 21% chance), and the gold is the under-attended segment on only 10%: it looks at the right tool and still picks wrong. This directly refutes the intuitive "crowded-harness / lost-in-the-middle" explanation: the failure is at the decision readout, not the harness, and we pin it there three ways. (1) Input vs. readout: repairing the prompt (reordering or duplicating the gold tool) recovers <=23% of failures, while readout-side interventions recover 59-91%. (2) Representation-invariance: two gold-pointed interventions in different representations -- an additive attention-logit bias and a residual-stream steering vector -- recover largely the same failures (per-task Jaccard 0.865 pooled, 0.79-0.91 per model), so the bottleneck is localized to the readout independent of which representation is poked. (3) A training-free, gold-free selector: per-segment attention closes most of the gold-free-vs-oracle gap on BFCL (+11.9 pts pooled function-name selection vs. +17.9-pt oracle headroom) and adds +14.9 pts on Seal-Tools; every model positive (exact McNemar p<=8e-4 each). Scopes differ: the causal attention-bias dose-response is bidirectional and monotonic on 10 mask-honoring models (3-32B), the full 0.5-32B span carrying only the correlational diagnostic; the deployable selector is evaluated on 5 single-turn models and does not yet transfer to a multi-turn loop.

13:00 JST研究/論文

ITNet: 畳み込み、注意、再帰を包含する学習可能な積分変換

畳み込みネットワーク、リカレント ネットワーク、トランスフォーマーはそれぞれ、局所性、逐次記憶、内容依存のペアワイズ相互作用など、さまざまな誘導バイアスをエンコードしており、その誕生以来数学的に区別され続けています。我々は、この断片化が信号の処理方法における基本的な多様性を反映しているのではなく、むしろ基礎となる単一の数学的オブジェクト、つまり学習可能な積分変換の不完全なビューを反映していることを示します。位置と特徴に共同して依存する学習可能なカーネルを中心に構築された統合アーキテクチャである Integral Transform Network (ITNet) を紹介します。このカーネルは、ペアごとの相互作用をモデル化する小さなニューラル ネットワーク、特に MLP として実装され、モデルがデータからその動作を適応できるようにします。畳み込み、自己注意 (マルチヘッドを含む)、および自己回帰再帰 (LSTM、GRU、S4、および Mamba を含む) が適切なパラメータ化の下で特殊なケースとして発生すること、および ITNet が連続演算子の汎用近似器であることを示します。これを実用化するために、タイル化カーネル融合、重要度加重モンテカルロ統合、学習された低ランク因数分解を開発し、効率的でスケーラブルな計算を可能にします。共有オペレーターと軽量のモダリティ固有のエンコーダーを備えた単一の ITNet アーキテクチャは、ImageNet-1K、GLUE、ModelNet40、VQA\,v2、および NLVR2 の特殊なベースラインと一致またはそれを超えています。この結果は、単一の学習された対話メカニズムが 3 つのアーキテクチャ ファミリすべての動作をデータから復元できることを示しています。

原文 (English)

ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence

Convolutional networks, recurrent networks, and transformers each encode different inductive biases -- locality, sequential memory, and content-dependent pairwise interaction -- and have remained mathematically distinct since their inception. We show that this fragmentation reflects not a fundamental diversity in how signals should be processed, but rather incomplete views of a single underlying mathematical object: a learnable integral transform. We introduce the Integral Transform Network (ITNet), a unified architecture built around a learnable kernel that depends jointly on positions and features. This kernel is implemented as a small neural network, specifically an MLP, that models pairwise interactions, enabling the model to adapt its behavior from data. We show that convolution, self-attention (including multi-head), and autoregressive recurrence (including LSTM, GRU, S4, and Mamba) arise as special cases under appropriate parameterizations, and that ITNet is a universal approximator of continuous operators. To make this practical, we develop tiled kernel fusion, importance-weighted Monte Carlo integration, and learned low-rank factorization, enabling efficient and scalable computation. A single ITNet architecture with a shared operator and lightweight modality-specific encoders matches or exceeds specialized baselines on ImageNet-1K , GLUE, ModelNet40, VQA\,v2 and NLVR2. The results demonstrate that a single learned interaction mechanism can recover the behavior of all three architectural families from data.

13:00 JSTエージェントビジネス/資金調達GPT / ChatGPT

When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, an…

13:00 JSTLLM/生成AIエージェント

Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents

Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget.…

13:00 JST研究/論文

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanni…

13:00 JSTLLM/生成AIエージェント

検証の地平線: コーディング エージェントの報酬に特効薬はない

古典的な直観では、解決策を生み出すよりも検証する方が簡単だと考えられています。今日のコーディング エージェントにとって、この直感は逆転しつつあります。基礎モデルがより強力な推論機能を開発し、エンジニアリング ハーネスがより洗練されるにつれて、複雑な候補ソリューションを生成することはもはや難しくなくなり、それらを確実に検証することがより困難な問題になりました。私たちが構築できるすべての検証ツールは人間の意図の代理にすぎず、意図そのものではありません。このため、検証は 2 つの困難を伴います。1 つは、本質的に意図が過少指定されているため、意図が満たされているかどうかを忠実に確認することが本質的に困難であるということです。次に、モデルのトレーニング中に、最適化によってプロキシとインテントの間のギャップが広がり、報酬のハッキングや信号の飽和として現れます。これに対処するために、私たちは検証信号の品質を 3 つの次元 (スケーラビリティ、忠実性、堅牢性) に沿って特徴付け、3 つすべてを同時に達成することが中心的な課題であると主張します。さらに、一般的なコーディング タスク用のテスト検証器、フロントエンド タスク用のルーブリック検証器、現実世界のエージェント タスク用の検証器としてのユーザー、長期タスク用の自動エージェント検証器の 4 つの報酬構造を研究します。さまざまなタスクの種類とポリシーの機能レベルにわたって、報酬設計の中核となる課題と、報酬シグナルをより効果的に活用する方法について、綿密な分析と実験を実施します。実験では、ターゲットを絞った検証設計により、報酬ハッキングを効果的に抑制し、タスク完了の品質を向上させ、複数の内部および公開ベンチマーク全体で大幅な利益を達成できることが示されています。これらの経験は総合的に、政策能力が成長し続けるにつれて、固定報酬関数が有効であり続けることはできないという核心的な観察を示しています。そして検証はジェネレーターと共進化する必要があります。

原文 (English)

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabilities and engineering harnesses grow more sophisticated, generating complex candidate solutions is no longer difficult -- reliably verifying them has become the harder problem. Every verifier we can build is only a proxy for human intent, never the intent itself. This makes verification subject to a twofold difficulty: first, intent is underspecified by nature, making it inherently hard to faithfully check whether it has been fulfilled; second, during model training, optimization widens the gap between proxy and intent -- manifesting as reward hacking or signal saturation. To address this, we characterize the quality of verification signals along three dimensions -- scalability, faithfulness, and robustness -- and argue that achieving all three simultaneously is the central challenge. We further study four reward constructions: a test verifier for general coding tasks, a rubric verifier for frontend tasks, the user as verifier for real-world agent tasks, and an automated agent verifier for long-horizon tasks. Across different task types and policy capability levels, we conduct in-depth analysis and experiments on the core challenges of reward design and how to more effectively leverage reward signals. Experiments show that targeted verification design can effectively suppress reward hacking, improve task completion quality, and achieve significant gains across multiple internal and public benchmarks. These experiences collectively point to a core observation: no fixed reward function can remain effective as policy capability continues to grow; and verification must co-evolve with the generator.

13:00 JSTエージェント

格子理論による不偏正準集合値オラクル

将来の出来事の確率を推定する非エージェント型の「オラクル」AI は、自己参照の問題に直面しています。その答えが学習され、実行されると、報告するよう求められた確率そのものが変わってしまう可能性があります。 Scientist AI プログラムで提唱されている対応の 1 つは、事実に反する質問のみをし、その回答が何の影響もないかのように評価することです。私たちは、そのような答えは学んだ瞬間に意味がなくなってしまう傾向があることを観察しています。それは、まさにその前提が間違っているからです。したがって、私たちは、オラクルが単一の確率ではなく、同時に偏りがなく学習の結果と自己矛盾のない一連の資格を報告する、自己言及的な代替案を模索します。単純な自己一貫性の要件は、あまりにも多くのセット (役に立たない答え $[0,1]$ を含む) によって満たされるため、問題は正規の自明でないメンバーを選び出すことです。これを、適切に定義されたアイソトーン演算子の最小不動点を取り、閉じたクレダル集合の完全な格子に関するクナスター-タルスキーの不動点定理を使用して行います。代わりに、バリアントは、すべての自己矛盾のない点推定を含む最小不動点を報告します。我々は、存在、自己無矛盾性、空でないことを証明し、非実行的質問についてはその構造が古典的な点の答えに崩壊すること、そしてバイナリイベントについては標準的な答えが自然なハル因数分解の仮定の下では区間であることを示します。この展開は純粋に格子理論に基づいており、バイナリ イベント $B$ から任意の確率変数 $X$ までそのまま拡張され、$P(B\mid A,C)$ は条件法 $\mathcal{L}(X\mid A,C)$ に置き換えられます。区間の特徴付け自体がその一般化に耐えられるかどうかを含む未解決の質問で終わります。

原文 (English)

Unbiased Canonical Set-Valued Oracles Via Lattice Theory

A non-agentic "oracle" that reports probabilities of future events is performative: once its answer is learned and acted upon, it can change the very probability it was asked to report. Performativity is not in itself the difficulty -- one consults an oracle precisely in order to be informed, and hence influenced, by it. The difficulty is agency. The requirement that a report be self-consistent, still holding once announced, may be met by many different values -- the classical non-uniqueness of self-fulfilling prophecies -- and any rule the system uses to choose among them is a lever for goal-directed steering. We remove the choice rather than the performativity. Reporting a credal set instead of a single probability distribution, we lift the reaction to an isotone operator on the complete lattice of closed credal sets, whose fixed points are self-consistent, and report its Knaster--Tarski least fixed point as a canonical, rule-determined answer; a variant reports instead the least fixed point that contains every self-consistent point estimate. We prove existence, self-consistency, and nonemptiness; show that the construction reduces to the classical point answer when the question is non-performative; and show that for a binary event the answer is, under a natural hull-factoring assumption, an interval.

13:00 JSTLLM/生成AI

JD Oxygen AI アイテム センター (Oxygen AIIC) V1: アイテムの理解、管理、アプリケーションのための産業規模の LLM/VLM 中心のソリューション

世界最大の電子商取引プラットフォームの 1 つである JD.com は、7 億人を超えるアクティブ ユーザーと数百万の販売者にサービスを提供し、数百億の SKU のカタログを持っています。この規模では、高品質で構造化されたアイテム ナレッジは、より良い消費者エクスペリエンス、管理コストの削減、運用効率の向上を支えますが、その生産と提供には、急速に出現するコンセプト、大規模な SKU 向けの高品質なナレッジの生産、および多様な下流要件という 3 つの産業規模の課題が生じます。これらの課題に対処するために、アイテム知識の生産とサービスのための LLM/VLM 上に構築された産業規模のプラットフォームである JD Oxygen AI Items Center (Oxygen AIIC) を紹介します。 Oxygen AIIC は 4 つの核となる柱を中心に構築されています。(i) 人間と AI の効率的なコラボレーションによって推進されるオントロジー エンジニアリング。これは、数百万のエントリを含むオントロジーの動的な進化と機敏な拡張をサポートします。 (ii) スループット向上戦略と組み合わせることで、数百億の SKU に対するスケーラブルで拡張性のある高スループットの AI アイテム ライブラリの作成を可能にする「セマンティック検索と識別」(S2D) 知識識別アーキテクチャ。 (iii) 自己進化する項目理解 LLM/VLM は、安定かつ制御可能な方法で改善し、94.2% の精度と 82.8% の再現率で知識生産を可能にします。 (iv) データとサービスのハブとして機能する統合アイテム トンネル。 Oxygen AIIC は現在、数万の JD カテゴリをカバーし、Huawei Ascend NPU で 1 日に数億件のアイテム更新を処理しています。数千億ものアイテム知識資産が蓄積されています。 Oxygen AIIC は、検索、レコメンデーション、運用、カテゴリ プランニングなど、中核となるビジネス シナリオ全体に展開され、大規模に目に見える利益をもたらしました。検索トラフィックのカバレッジは 80.4% に達し、商品情報の品質問題は 37% 減少し、商品リスト中の主要属性の自動入力率は 80% を超えました。

原文 (English)

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications

JD$.$com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchants, with a catalog of tens of billions of SKUs. At this scale, high-quality, structured item knowledge underpins a better consumer experience, lower management costs, and higher operational efficiency-yet producing and serving it poses three industrial-scale challenges: fast-emerging concepts, high-quality knowledge production for massive SKUs, and diverse downstream requirements. To address these challenges, we present the JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service. Oxygen AIIC is built around four core pillars: (i) ontology engineering driven by efficient human-AI collaboration, which supports the dynamic evolution and agile expansion of an ontology with millions of entries; (ii) a "Semantic Search then Discrimination"(S2D) knowledge identification architecture that, combined with throughput improvement strategies, enables scalable, extensible, and high-throughput AI Item Library production for tens of billions of SKUs; (iii) self-evolving item-understanding LLMs/VLMs that improve in a stable and controllable manner, enabling knowledge production with 94.2% precision and 82.8% recall; and (iv) a unified item tunnel that serves as the data and service hub. Oxygen AIIC now covers tens of thousands of JD categories and processes hundreds of millions of item updates per day on Huawei Ascend NPUs. It has accumulated hundreds of billions of item-knowledge assets. Deployed across core business scenarios-including search, recommendation, operations, category planning-Oxygen AIIC has delivered measurable gains at scale. Search-traffic coverage reaches 80.4%, item-information quality issues drop by 37%, the automated fill rate of core attributes during item listing exceeds 80%.

13:00 JST研究/論文

Ensemble Learning Based Classification Algorithm Recommendation

Selecting an appropriate classification algorithm for a given data set remains a challenging problem in data mining and machine learning. E…

13:00 JST研究/論文

Granular-ball computing: an efficient, robust, and interpretable adaptive multi-granularity representation and computation method

To overcome the limitations of point-based inputs, overly fine computation and limited adaptability in existing artificial intelligence met…

13:00 JST研究/論文

TERC: A Transfer Entropy Redundancy Criterion for State Variable Selection in Reinforcement Learning

Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL). These variables m…

13:00 JST研究/論文

Assortment Planning with Sponsored Products

In the rapidly evolving landscape of retail, assortment planning plays a crucial role in determining the success of a business. With the ri…

13:00 JST画像/動画生成研究/論文

SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces

Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in ext…

13:00 JSTビジネス/資金調達研究/論文

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivaria…

13:00 JSTエージェント

ProSpec RL: Plan Ahead, then Execute

Imagining potential outcomes of actions before execution helps agents make more informed decisions, a prospective thinking ability fundamen…

13:00 JST画像/動画生成

Beyond Spectral Decomposition: Bayesian Contrastive Learning and its Non-negative Formulation via Factor Analysis

Factor analysis, often regarded as a Bayesian variant of matrix factorization, offers superior capabilities in capturing uncertainty, model…

13:00 JST研究/論文

Interpretable Clustering: A Survey

In recent years, much of the research on clustering algorithms has primarily focused on enhancing their accuracy and efficiency, frequently…

13:00 JST画像/動画生成

FLAME 3 Dataset: Unleashing the Power of Radiometric Thermal UAV Imagery for Wildfire Management

The increasing accessibility of radiometric thermal imaging sensors for unmanned aerial vehicles (UAVs) offers significant potential for ad…

13:00 JSTLLM/生成AI研究/論文

XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) synergizes the retrieval of pertinent data with the generative capabilities of Large Language Models (…

13:00 JSTLLM/生成AI研究/論文

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety…

13:00 JST研究/論文

Bridging Neural Networks and Wireless Systems with MIMO-OFDM Semantic Communications

Semantic communications aim to enhance transmission efficiency by jointly optimizing source coding, channel coding, and modulation. While p…

13:00 JSTLLM/生成AI

Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question Answering

Large language models (LLMs) have shown remarkable capabilities in natural language processing. However, in knowledge graph question answer…

13:00 JSTビジネス/資金調達

Overcoming Dependent Censoring in the Evaluation of Survival Models

Dependent censoring occurs when the event time and censoring time are not conditionally independent given the observed covariates. This com…

13:00 JSTLLM/生成AI

Distributionally Robust Reinforcement Learning with Human Feedback

Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). H…

13:00 JST画像/動画生成

Pose-Based Fall Detection System: Efficient Monitoring on Standard CPUs

Falls among elderly residents in assisted living homes pose significant health risks, often leading to injuries and a decreased quality of…

13:00 JST研究/論文

Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks

Large language models (LMs) offer broad generalization capabilities but require vast amounts of data and computational resources for domain…

13:00 JST画像/動画生成

Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated con…

13:00 JSTロボティクス

Representation Learning for Equivariant Inference with Guarantees

In many real-world applications of regression, conditional probability estimation, and uncertainty quantification, exploiting symmetries ro…

13:00 JST研究/論文

Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation

Modeling physical systems in a generative manner offers several advantages, including the ability to handle partial observations, generate…

13:00 JSTLLM/生成AI

Scaling Textual Gradients via Sampling-Based Momentum

LLM-based prompt optimization, which uses LLM-provided ``textual gradients'' (feedback) to refine prompts, has emerged as an effective meth…

13:00 JST研究/論文

Multimodal Representation Alignment for Cross-modal Information Retrieval

Different machine learning models can represent the same underlying concept in different ways. This variability is particularly valuable fo…

13:00 JSTエージェントロボティクスGoogle

Towards Biosignals-Free Autonomous Prosthetic Hand Control via Imitation Learning

Limb loss affects millions globally, impairing physical function and reducing quality of life. Most traditional surface electromyographic (…

13:00 JSTLLM/生成AIエージェント

Modeling Earth-Scale Human-Like Societies with One Billion Agents

Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simu…

13:00 JST画像/動画生成

MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection

Small-object detection in Unmanned Aerial Vehicle (UAV) imagery requires preserving weak local evidence while using broader context to sepa…

13:00 JSTLLM/生成AI

Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action

The rise of large language models (LLMs) has led to dramatic improvements across a wide range of natural language tasks. Their performance…

13:00 JST研究/論文

GeNeRT: A Physics-Informed Approach to Intelligent Wireless Channel Modeling via Generalizable Neural Ray Tracing

Neural ray tracing (RT) has emerged as a promising paradigm for channel modeling by integrating physical propagation principles with neural…

13:00 JSTLLM/生成AIエージェント研究/論文

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

Recent benchmarks for Large Language Model (LLM) agents primarily focus on evaluating reasoning, planning, and execution capabilities, whil…

13:00 JST研究/論文

Beyond Correlation: Learning Supervised, Sample-Distinct, and Eigenimage-Interpretable Representations

Conventional dimensionality reduction methods mainly optimize variance or correlation, leaving statistical dependence, data diversity, cont…

13:00 JSTロボティクス

Multi-Class Human/Object Detection on Robot Manipulators using Proprioceptive Sensing

In physical human-robot collaboration (pHRC) settings, humans and robots collaborate directly in shared environments. Robots must analyze i…

13:00 JSTLLM/生成AI

LLM Serving Optimization with Variable Prefill and Decode Lengths

We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous…

13:00 JSTLLM/生成AI

Post-training for Efficient Communication via Convention Formation

Humans communicate with increasing efficiency in multi-turn interactions, by adapting their language and forming ad-hoc conventions. In con…

13:00 JST画像/動画生成

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Visio…

13:00 JSTLLM/生成AI

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-sol…

13:00 JSTロボティクス研究/論文

Tactile Gesture Recognition with Built-in Joint Sensors for Industrial Robots

While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores de…

13:00 JST画像/動画生成研究/論文

PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science

Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasonin…

13:00 JSTLLM/生成AI

EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-Based UI Code Generation

Multimodal Large Language Models have demonstrated exceptional performance in UI2Code tasks, significantly enhancing website development ef…

13:00 JSTLLM/生成AI

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean

Large Language Models (LLMs) have demonstrated significant promise in formal theorem proving. In this study, we investigate the ability of…

13:00 JST研究/論文

ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching

Internal variability is a dominant contributor to the uncertainty of predictions at the interannual to decadal timescale. A typical approac…

13:00 JSTLLM/生成AI

Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management

Large language models (LLMs) are increasingly used to simulate human behavior in business, economics, and the social sciences, offering a l…

13:00 JSTLLM/生成AI

Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization

Large language models (LLMs) are increasingly used as rerankers in information retrieval, yet their ranking behavior can be steered by smal…

13:00 JSTLLM/生成AI

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts

Training expert LLMs in domains with scarce data is difficult, often relying on multiple-choice questions (MCQs). However, standard outcome…

13:00 JST研究/論文

Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning

We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce \textbf{Attribution Graphs} (AGs),…

13:00 JSTLLM/生成AI

Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers

Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits. Howeve…

13:00 JST画像/動画生成

Automatic Extraction of Road Networks by using Teacher-Student Adaptive Structural Deep Belief Network and Its Application to Landslide Disaster

An adaptive structural learning method of Restricted Boltzmann Machine (RBM) and Deep Belief Network (DBN) has been developed as one of pro…

13:00 JSTLLM/生成AIGPT / ChatGPT

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning rem…

13:00 JST研究/論文

Hard-constraint physics-residual networks for hydrogen crossover prediction and high-pressure extrapolation in PEM water electrolysis

Hydrogen crossover is a critical safety and efficiency constraint in high-pressure polymer electrolyte membrane water electrolysis (PEMWE),…

13:00 JST研究/論文

SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?

Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduc…

13:00 JSTLLM/生成AI

Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs

Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and express…

13:00 JSTLLM/生成AI画像/動画生成

Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models

Vision--language models (VLMs) have recently shown promise for assisting clinical reasoning in dermatological diagnosis. However, their tru…

13:00 JST画像/動画生成ロボティクス研究/論文

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyda…

13:00 JST研究/論文

When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate

Time series extrinsic regression (TSER) refers to the task of predicting a continuous target variable from an input time series. It appears…

13:00 JSTエージェント

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory

As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience…

13:00 JST研究/論文

Weighted Contrastive Learning for Anomaly-Aware Time-Series Forecasting

Reliable forecasting of multivariate time series under anomalous conditions is crucial in applications such as ATM cash logistics, where su…

13:00 JSTLLM/生成AI研究/論文Gemini

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks r…

13:00 JST研究/論文

AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research

The rapid growth of Artificial Intelligence and Machine Learning in scientific research has highlighted a gap between industry-standard MLO…

13:00 JST画像/動画生成

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) r…

13:00 JST研究/論文

Neural Minimum Weight Perfect Matching for Quantum Error Codes

Realizing the full potential of quantum computation requires Quantum Error Correction (QEC). QEC reduces error rates by encoding logical in…

13:00 JSTLLM/生成AI

Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concern…

13:00 JSTLLM/生成AIエージェント

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent…

13:00 JSTLLM/生成AIビジネス/資金調達

From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP

While NLP typically treats documents as independent and unordered samples, in longitudinal studies, this assumption rarely holds: documents…

13:00 JSTLLM/生成AIGPT / ChatGPTLlama

A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics

To address the scalability of feedback in computer science while mitigating the privacy and cost limitations of commercial Large Language M…

13:00 JST画像/動画生成

CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training

The functions of different regions of the human brain are closely linked to their distinct cytoarchitecture, which is defined by the spatia…

13:00 JSTハードウェア/半導体Microsoft

Diff-MN: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations

Time series generation (TSG) is widely used across domains, yet most existing methods assume regular sampling and fixed output resolutions.…

13:00 JSTLLM/生成AIエージェントロボティクスClaude

Demonstration-Free Robotic Control via LLM Agents

Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require t…

13:00 JST研究/論文

Assessing the Business Process Modeling Competences of Large Language Models

The creation of Business Process Model and Notation (BPMN) models is a complex and time-consuming task requiring both domain knowledge and…

13:00 JSTロボティクス

Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment

We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions…

13:00 JST研究/論文

Agile Reinforcement Learning through Separable Neural Architecture and Applications

Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet go-to function approximators - multilay…

13:00 JSTLLM/生成AI

A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential…

13:00 JSTLLM/生成AI

Test-Time Detoxification without Training or Learning Anything

Large language models can produce toxic or inappropriate text even for benign inputs, creating risks when deployed at scale. Detoxification…

13:00 JSTLLM/生成AI

ImprovEvolve: Basin-Hopping Meets LLM-Guided Evolutionary Search

LLM-guided evolutionary computation, most notably AlphaEvolve, has been remarkably successful in discovering novel mathematical constructio…

13:00 JST研究/論文

General and Efficient Steering of Diffusion Models

Steering diffusion models toward conditions unseen during training typically requires either retraining with conditional inputs or per-step…

13:00 JSTエージェント

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation

As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove…

13:00 JSTLLM/生成AI

Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning

While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety me…

13:00 JSTロボティクス

WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its re…

13:00 JST研究/論文

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it rema…

13:00 JSTLLM/生成AI画像/動画生成

How to Train Your Long-Context Visual Document Model

We present the first comprehensive, large-scale study of training long-context vision language models up to 344K context, targeting long-do…

13:00 JST画像/動画生成NVIDIA

Spanning the Visual Analogy Space with a Weight Basis of LoRAs

Visual analogy learning enables image editing via demonstration rather than textual description, allowing users to specify complex transfor…

13:00 JST研究/論文

Enhanced Diffusion Sampling: Efficient Rare Event Sampling and Free Energy Calculation with Diffusion Models

The rare-event sampling problem has long been the central limiting factor in molecular dynamics (MD), especially in biomolecular simulation…

13:00 JST研究/論文

DyGnROLE: Asymmetric Pretraining for Edge Classification on Dynamic Graphs

Edge classification on directed dynamic graphs requires modeling interactions between source and destination nodes exhibiting asymmetrical…

13:00 JST研究/論文

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical mod…

13:00 JSTエージェントロボティクス

What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty

As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? C…

13:00 JST画像/動画生成

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing

Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains ch…

13:00 JSTLLM/生成AIハードウェア/半導体

The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding

Large language models (LLMs) are increasingly used to generate executable outputs, JSON objects, and API calls, where a single syntax error…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects

Large Language Models (LLMs) have shown remarkable potential in developing role-playing agents (RPAs). However, current evaluation framewor…

13:00 JST画像/動画生成

Longitudinal Lesion Inpainting in Brain MRI via 3D Region Aware Diffusion

Accurate longitudinal analysis of brain MRI is often hindered by evolving lesions, which bias automated neuroimaging pipelines. While deep…

13:00 JSTLLM/生成AIエージェント

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which…

13:00 JSTLLM/生成AI研究/論文

HEARTS: Benchmarking LLM Reasoning on Health Time Series

The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning. Yet, existing…

13:00 JST研究/論文

Feature-level Interaction Explanations in Multimodal Transformers

Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision. Most existing mul…

13:00 JST画像/動画生成ロボティクス

FlatLands: Generative Floormap Completion From a Single Egocentric View

A single egocentric image typically captures only a small portion of the floor, yet a complete metric traversability map of the surrounding…

13:00 JST研究/論文

DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning

Automating the development of machine learning algorithms has the potential to unlock new breakthroughs. However, our ability to improve an…

13:00 JST画像/動画生成

Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding

Recent advances in Streaming Video Understanding has enabled a new interaction paradigm where models respond proactively to user queries. C…

13:00 JST画像/動画生成

UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation

We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural l…

13:00 JSTロボティクス

Grounding Sim-to-Real Generalization in Robotic Manipulation: An Empirical Study with Vision-Language-Action Models

Learning a generalist control policy for robotic manipulation typically relies on large-scale datasets. Given the high cost of real-world d…

13:00 JST画像/動画生成エージェント

CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training

GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screensho…

13:00 JST画像/動画生成

FD$^2$: A Dedicated Framework for Fine-Grained Dataset Distillation

Dataset distillation (DD) compresses a large training set into a small synthetic set, reducing storage and training cost, and has shown str…

13:00 JSTLLM/生成AI

Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off

Retrieval-Augmented Generation (RAG) systems for financial document QA typically follow a chunk-based paradigm: documents are split into fr…

13:00 JST画像/動画生成

Steerable Visual Representations

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream…

13:00 JSTLLM/生成AI画像/動画生成Mistral AI

Internalized Reasoning for Long-Context Visual Document Understanding

Visual long-document understanding is critical for enterprise, legal, and scientific applications, yet the best performing open recipes hav…

13:00 JSTLLM/生成AI

How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models

We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content…

13:00 JST研究/論文

UniMamba: A Unified Spatial-Temporal Modeling Framework with State-Space and Attention Integration

Multivariate time series forecasting is fundamental to numerous domains such as energy, finance, and environmental monitoring, where comple…

13:00 JST画像/動画生成

Towards Modality-Agnostic Medical Image Anomaly Detection: A Training-Free Manifold Refinement Approach

Deploying AI-based anomaly detection across diverse clinical imaging settings remains challenging because most existing methods rely on mod…

13:00 JSTLLM/生成AIエージェント

Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction

Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLM…

13:00 JST研究/論文

Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

Neural networks can be trained to learn task-relevant representations from data. Understanding how these networks make decisions falls with…

13:00 JSTビジネス/資金調達研究/論文

Defeasible Conditional Obligation in a Two-tiered Preference-based Semantics (Extended Version)

In response to a concern raised by Horty, this paper develops a two-tiered, preference-based semantic framework for modeling defeasible con…

13:00 JSTLLM/生成AI画像/動画生成Gemini

Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followe…

13:00 JSTLLM/生成AI

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives

Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study thes…

13:00 JST研究/論文

On the Spectral Structure and Objective Equivalence of Orthogonal Multilabel Fisher Discriminants

We provide a unified theoretical analysis of Linear Discriminant Analysis with simultaneous multilabel scatter matrix formulations and Stie…

13:00 JSTロボティクス

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning

Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a self-guided mechanism for online i…

13:00 JSTエージェントロボティクス

BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation

Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab…

13:00 JST研究/論文

Globally Optimal Training of Spiking Neural Networks via Parameter Reconstruction

Spiking Neural Networks (SNNs) have been proposed as biologically plausible and energy-efficient alternatives to conventional Artificial Ne…

13:00 JSTLLM/生成AIエージェント

Robust Multi-Agent LLMs under Byzantine Faults

Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability. However, these same int…

13:00 JSTLLM/生成AI

Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation

While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionat…

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTGemini

Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, al…

13:00 JST画像/動画生成

車輪ではなくブレーキを壊す: エントロピー最大化による対象外の脱獄

最近の研究では、ビジョン言語モデル (VLM) 上の勾配ベースのユニバーサル イメージ ジェイルブレイクは、モデル間での移行性がほとんどまたはまったく示されないことが示されており、移行可能なマルチモーダル ジェイルブレイクの実現可能性に疑問が投げかけられています。固定プレフィックスや応答パターンを強制することなく、厳密に対象を絞らない脅威モデルに基づいてこの結論を再検討します。私たちの予備実験では、自己回帰復号中に拒否動作が高エントロピーのトークンに集中しており、非拒否トークンは攻撃前にすでに上位候補の中でかなりの確率質量を持っていることが明らかになりました。この発見に動機付けられて、私たちはエントロピー最大化(UJEM)-KLによるUntargeted Jailbreakを提案します。これは、出力品質を維持するために残りの低エントロピー位置を安定させながら、これらの決定トークンでエントロピーを最大化して拒否結果を反転させる軽量の攻撃です。 UJEM-KL は、3 つの VLM と 2 つの安全ベンチマークにわたって、競争力のあるホワイトボックス攻撃の成功率を達成し、転送可能性を一貫して向上させながら、代表的な防御下でも効果を維持します。私たちの実験結果は、限られた移転可能性は主に過度に制約された最適化目標に起因することを示しています。

原文 (English)

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks. We revisit this conclusion under a strictly untargeted threat model without enforcing a fixed prefix or response pattern. Our preliminary experiment reveals that refusal behavior concentrates at high-entropy tokens during autoregressive decoding, and non-refusal tokens already carry substantial probability mass among the top-ranked candidates before attack. Motivated by this finding, we propose Untargeted Jailbreak via Entropy Maximization(UJEM)-KL, a lightweight attack that maximizes entropy at these decision tokens to flip refusal outcomes, while stabilizing the remaining low-entropy positions to preserve output quality. Across three VLMs and two safety benchmarks, UJEM-KL achieves competitive white-box attack success rates and consistently improves transferability, while remaining effective under representative defenses. Our experimental results indicate that the limited transferability primarily stems from overly constrained optimization objectives.

13:00 JSTLLM/生成AI研究/論文

Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition

Fine-tuning multilingual ASR models like Whisper for low-resource languages often improves read speech but degrades spontaneous audio perfo…

13:00 JST研究/論文

Not All Timesteps Matter Equally: Selective Alignment Knowledge Distillation for Spiking Neural Networks

Spiking neural networks (SNNs), which are brain-inspired and spike-driven, achieve high energy efficiency. However, a performance gap betwe…

13:00 JST研究/論文

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning

Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, t…

13:00 JSTLLM/生成AIビジネス/資金調達

SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not…

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPTGemini

Frontier LLM はサイバーセキュリティに対応する準備ができていますか?デュアルモード脆弱性ベンチマークによる垂直基盤モデルの証拠

当社は、フロンティア LLM がデュアルモード ベンチマークを通じてサイバーセキュリティに対応できるかどうかを評価します。ホワイトボックス機能レベルの脆弱性検出 (VulnLLM-R、C/Java/Python 全体) とブラックボックス Web アプリケーション セキュリティ テスト (20 以上の CWE ファミリにわたる 118 個のグラウンド トゥルース脆弱性を備えた 5 つの運用スタイルのアプリケーション。これらをオープンソース化します)。私たちは 6 つのフロンティア モデル (GPT-5.4、Codex~5.3、Claude Opus~4.6、Sonnet~4.6、Gemini~3.1~Pro、および Gemini~3~Flash) と 2 つのドメイン特化モデルを 4 つのテスト パラダイムにわたってテストします。私たちの発見は厳粛なものです。(1) ~すべてのフロンティア モデルは、ホワイトボックス検出で 10 ~ 50% の誤検知率を生成し、体系的に脆弱性を過剰予測します。 (2)〜ブラックボックス テストでは、フロンティア モデルはグラウンド トゥルース カバレッジをわずか 4 ~ 8% しか達成せず、外部セキュリティ ツール (Playwright MCP、Burp Suite MCP) を使用した場合でもわずか 10 ~ 19% に改善します。 (3) ドメイン特化型エージェントにエンコードされた構造化侵入テスト手法により、ファミリーごとの検出が 50% を超え、規模ではなく手法が主要な手段であることが実証されました。 (4) ドメインに特化した防御モデルは、単一 GPU 上ですべてのモデルの中で最高の精度 (0.904) と最低の誤検知率 (9.7%) を達成します。私たちは、構造化されたセキュリティ テストの欠如、エンドツーエンドの要求/応答シーケンス、障害の多いデータ、および複数ステップの攻撃チェーンのトレースが基本的なトレーニング データのボトルネックであることを特定し、データ生成戦略としてセルフプレイ セキュリティ テストを提案します。私たちの結果は、サイバーセキュリティ専用に構築された垂直基盤モデルの正当性を裏付けています。

原文 (English)

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production-style applications with 118 ground-truth vulnerabilities across 20+ CWE families, which we will open-source). We test six frontier models (GPT-5.4, Codex~5.3, Claude Opus~4.7, Sonnet~4.6, Gemini~3.1~Pro and Gemini~3~Flash) and two domain-specialized models across four testing paradigms. Our findings are sobering: (1)~every frontier model produces 10-50% false positive rates in white-box detection, systematically over-predicting vulnerabilities; (2)~in black-box testing, frontier models achieve only 4-8% ground-truth coverage, improving to just 10-19% even with external security tools (Playwright MCP, Burp Suite MCP); (3)~structured penetration-testing methodology encoded in domain-specialized agents raises per-family detection above 50%, demonstrating that methodology, not scale, is the primary lever; and (4)~a domain-specialized defense model achieves the highest precision (0.904) and lowest false positive rate (9.7%) among all models, on a single GPU. We identify the absence of structured security testing traces end-to-end request/response sequences, failure-heavy data, and multi-step attack chains as the fundamental training data bottleneck, and propose self-play security testing as a data generation strategy. Our results make the case for vertical foundation models purpose-built for cybersecurity.

13:00 JST研究/論文

良い方程式が悪いスコアになる場合: パラメーターの最適化を改善することでシンボリック回帰を改善する

シンボリック回帰 (SR) は、観察データから数式を抽出することにより、科学的知識の発見において中心的な役割を果たします。既存の SR メソッドのほとんどは、2 レベルの最適化フレームワーク内で機能します。つまり、離散方程式構造を検索する外側のループと、その構造の連続パラメーターを最適化する内側のループです。重要なのは、パラメーターの適合品質が構造のスコアを直接決定し、したがって外部ループの検索を決定することです。ただし、非線形演算子により内部ループが高度に非凸になり、高速ローカル ソルバー (BFGS など) への予算主導の依存により、正しい構造に対する不十分な極小値や過小評価されたスコアが得られることがよくあります。この「良い構造、悪いスコア」現象が重要なボトルネックとなり、効率を低下させ、検索を真の方程式から誤って導きます。これを解決するために、シンボリック式のデュアル ネイティブ事前分布を利用する SR ネイティブ フィッティング フレームワークである SAGE-Fit (Structure-Aware and Semantics-Guided Evaluator for Symbolic Regression) を提案します。 SR に特有の構造的および意味論的な事前確率を利用することで、各プロパティに合わせたモジュールを設計し、それによってこの最適化のボトルネックを効果的に軽減します。広範な実験により、プラグ アンド プレイ モジュールとしての当社のアプローチが評価の忠実度を大幅に向上させ、さまざまな SR システムのパフォーマンスを普遍的に向上させることが実証されました。

原文 (English)

When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization

Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational data. Most existing SR methods function within a bi-level optimization framework: an outer loop that searches for the discrete equation structure, and an inner loop that optimizes the continuous parameters of that structure. Crucially, parameter-fitting quality directly determines a structure's score and thus the outer-loop search. However, nonlinear operators make the inner loop highly non-convex, and budget-driven reliance on fast local solvers (e.g., BFGS) often yields poor local minima and underestimated scores for correct structures. This ``Good Structure, Bad Score'' phenomenon becomes a key bottleneck, degrading efficiency and misguiding the search away from the true equation. To resolve this, we propose SAGE-Fit (Structure-Aware and Semantics-Guided Evaluator for Symbolic Regression), an SR-native fitting framework that exploits the dual native priors of symbolic expressions. By capitalizing on the structural and semantic priors unique to SR, we design tailored modules for each property, thereby effectively mitigating this optimization bottleneck. Extensive experiments demonstrate that our approach, as a plug-and-play module, significantly enhances evaluation fidelity and universally improves the performance of various SR systems.

13:00 JST規制/政策

高リスク AI システムと欧州 AI 法におけるアイデンティティの問題

EU 人工知能法 (AIA) は、事前の適合性評価、市販後のモニタリング、および「大幅な変更」時の再評価を中心に構築された高リスク AI システムのライフサイクル ガバナンス体制を確立しています。これらの義務は AI のアイデンティティ判断を前提としています。規制当局とプロバイダーは、更新されたシステムが長期間にわたって同じシステムのままであるかどうかを決定する必要があります。この研究では、このロジックがアーティファクト ID の機能 + フレームワークによってどのように明確化されるかを示します。このフレームワークは、「AI の信頼性」として捉えられる、適切な機能の状況依存の基準とともに、意図された機能によって AI システムを個別化します。さらに我々は、AIAは同期的同一性(規制上の目的で一度に2つのAIシステムが同一とみなされるべき場合)に関する内部の監査可能な基準を提供しておらず、代わりにそのような同一性の判断を分野別または調和化の手段に大きく委ねていると主張する。 function+ は、意図した機能と信頼性のプロファイルとレベルに基づいた同期アイデンティティ テストを提供し、調達、責任、市場監視などのガバナンス設定で同期アイデンティティの決定を検査可能にします。私たちの貢献は概念的なレンズと監査レンズです。私たちは、AIA ライフサイクル義務と機能 + アイデンティティ コンポーネント間の対応マップを提供し、監査と紛争のコンテキストに関する最小限の意思決定フローを通じて同期ケースを運用上判読できるようにします。最後に、実装に向けた 2 つの推奨事項を示します。(1) 意図された目的についての、より正確でテスト可能なレポート。(2) 経時的および導入間での比較可能性をサポートする、標準化された監査可能な信頼性レポート。

原文 (English)

High-Risk AI Systems and the Problem of Identity in the European AI Act

The EU Artificial Intelligence Act (AIA) establishes a lifecycle governance regime for high-risk AI systems built around ex-ante conformity assessment, post-market monitoring, and re-assessment upon "substantial modification." These obligations presuppose AI identity judgments: regulators and providers must decide when an updated system remains the same system over time. In this work, we show how this logic is clarified by the function+ framework of artifact identity, which individuates AI systems by their intended function together with context-sensitive criteria of appropriate functioning, captured as "AI trustworthiness." We further argue that the AIA does not provide an internal, auditable criterion for synchronic identity--when two AI systems at a given time should count as the same for regulatory purposes--and instead largely defers such sameness determinations to sectoral or harmonization instruments. function+ supplies a synchronic identity test anchored in intended function and trustworthiness profiles and levels, making synchronic identity decisions inspectable in governance settings such as procurement, liability, and market surveillance. Our contribution is a conceptual and auditing lens: we provide a correspondence map between AIA lifecycle obligations and function+ identity components, and we make the synchronic case operationally legible via a minimal decision flow for audit and dispute contexts. We conclude with two implementation-facing recommendations: (1) more precise, testable reporting of intended purpose, and (2) standardized, auditable trustworthiness reporting that supports comparability over time and across deployments.

13:00 JSTエージェント

ChainCaps: 単調な機能減衰による構成安全なツール使用エージェント

ツールを使用するエージェントは、実行時にファイル システム、Web API、コード インタープリタ、およびエンタープライズ サービスを構成する、オープンエンドの展開環境で動作することが増えています。これにより、ツール構成に安全性のギャップが生じます。エージェントは、ツールごとのすべての権限チェックを満たしていても、機密文書の読み取り、要約、その要約の外部エンドポイントへの送信など、安全でないエンドツーエンドの影響を生み出す可能性があります。この障害モードをパーミッション ロンダリングと呼びます。 ChainCaps は、ランタイム ルールでこれに対処します。すべての値にはシンク固有の機能バジェットが含まれ、ツールの構成によって交差ごとにバジェットが伝播されます。値は、ツール チェーン内を移動するときに権限を保持したり失ったりする可能性がありますが、合成を通じて新しい権限を獲得することはできません。 ChainCaps は、エージェント サーバーやツール サーバーへの変更を必要としない透過的な MCP プロキシとして実装されています。 ChainCaps は、3 つのプロバイダーの 5 つのフロンティア モデルにわたる 82 のタスクにおいて、96 ~ 100% の無害な完了を維持しながら、攻撃の成功率を 25 ~ 68% から 0 ~ 4.8% に低下させます。再生実験では、スカラー IFC および関数ごとの分離ベースラインよりも優れたパフォーマンスを示します。マニフェストの品質が導入の主なボトルネックです。エキスパート マニフェストは 100% の攻撃ブロックに達しますが、単純なマニフェストは 27.3% に低下します。私たちの主張は、信頼できるマニフェストの下での明示的なフロー構成の安全性と、プロキシで可視のデータ移動に限定されており、今日導入されているツールを使用するエージェントには実際的なギャップがあります。

原文 (English)

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterprise services at runtime. This creates a safety gap in tool composition: an agent can satisfy every per-tool permission check and still produce an unsafe end-to-end effect, such as reading a confidential document, summarizing it, and sending the summary to an external endpoint. We call this failure mode permission laundering. ChainCaps addresses it with a runtime rule: every value carries a sink-specific capability budget, and tool composition propagates budgets by intersection. A value can preserve or lose authority as it moves through a tool chain, but it cannot gain new authority through composition. We implement ChainCaps as a transparent MCP proxy that requires no changes to the agent or tool servers. On 82 tasks across five frontier models from three providers, ChainCaps reduces attack success rate from 25-68% to 0-4.8% while preserving 96-100% benign completion. In replay experiments, it also outperforms scalar-IFC and per-function-isolation baselines. Manifest quality is the dominant deployment bottleneck: expert manifests reach 100% attack blocking, while naive manifests fall to 27.3%. Our claims are limited to explicit-flow composition safety under trusted manifests and proxy-visible data movement, a practical gap in deployed tool-using agents today.

13:00 JSTエージェント

Big 2 の不完全情報下でのセルフプレイ強化学習

不完全情報マルチプレイヤー ゲームでは、隠された情報、まばらな報酬、および静止していない敵の下でエージェントが行動できるかどうかをテストします。私たちはこれらの課題を、4 人用の不完全情報カード ゲームである Big 2 で研究します。私たちは、ポリシー勾配エージェントと値近似エージェント間の制御された比較を可能にする Big 2 用のセルフプレイ RL フレームワークを開発します。共通の環境、入力表現、トレーニング予算、および評価プロトコルの下では、PPO は、ランダムで貪欲でヒューリスティックな Big 2 の敵に対して、モンテカルロ Q 近似、SARSA、および Q 学習よりも優れたパフォーマンスを発揮します。さらに、適度なエントロピー正則化により、ポリシーが過度に決定論的になるのを防ぎ、PPO が向上すること、および現在のポリシーのセルフプレイは、チェックポイント セルフプレイや固定対戦相手のトレーニングよりも強力な有限予算のカリキュラムを提供することがわかりました。まとめると、これらの結果は、Big 2 が、不完全な情報、マルチプレイヤー インタラクション、遅延報酬、および可変アクション セットの下で深い RL を研究するのに有用な制御された設定であることを示しています。

原文 (English)

Self-Play Reinforcement Learning under Imperfect Information in Big 2

Imperfect-information multiplayer games test whether agents can act under hidden information, sparse rewards, and non-stationary opponents. We study these challenges in Big 2, a four-player imperfect-information card game. We develop a self-play RL framework for Big 2 that enables controlled comparisons between policy-gradient and value-approximating agents. Under a common environment, input representation, training budget, and evaluation protocol, PPO outperforms Monte Carlo Q approximation, SARSA, and Q-learning against random, greedy, and heuristic Big 2 opponents. We further find that moderate entropy regularization improves PPO by preventing the policy from becoming overly deterministic, and that current-policy self-play provides a stronger finite-budget curriculum than checkpoint self-play or fixed-opponent training. Together, these results show that Big 2 is a useful controlled setting for studying deep RL under imperfect information, multiplayer interaction, delayed rewards, and variable action sets.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in realistic, electronic health recor…

13:00 JSTLLM/生成AIエージェント

PatchWorld: Gradient-Free Optimization of Executable World Models

Text-agent environments are typically modeled as partially observable Markov decision processes (POMDPs), assuming that the simulator's lat…

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large langua…

13:00 JST画像/動画生成

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models

Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large h…

13:00 JSTLLM/生成AI

Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance

On-Policy Distillation (OPD) improves large language model reasoning by training a student model on trajectories sampled from its own polic…

13:00 JST画像/動画生成研究/論文GPT / ChatGPT

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We intro…

13:00 JSTLLM/生成AI画像/動画生成

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs

Current 3D spatial reasoning methods face a fundamental trade-off: neuro-symbolic 3D (NS3D) concept learners achieve interpretable reasonin…

13:00 JSTLLM/生成AIエージェント

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment. However, in practical app…

13:00 JSTLLM/生成AIエージェント

Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents

Large language model (LLM) agents are becoming long-running software actors rather than fixed tool users. They accumulate memory, activate…

13:00 JSTLLM/生成AI

LiftQuant: 次元リフティングと投影による連続ビット幅 LLM

既存の量子化手法は基本的に、厳格な整数ベースのビット幅 (例: 2、3 ビット) によって制限されており、その結果、大規模言語モデルを特定のメモリ バジェットに最適に適合させることができない「デプロイメント ギャップ」が生じます。このギャップを埋めるために、真のパレート最適デプロイメントのための継続的なビット幅制御を可能にする新しいフレームワークである LiftQuant を紹介します。中心となるイノベーションは、「リフト ゼン プロジェクト」メカニズムです。高次元の「リフトされた」空間から単純な 1 ビット格子を投影することで、低次元の重みベクトルを近似します。重要なことに、有効なビット幅は、元の次元に対するリフト次元の比率によって単純に決定され、次元が柔軟な構造パラメータであるため、ビット幅を準連続的に調整できます。この投影は、構造化されているが不均一なコードブックを生成し、ベクトル量子化 (VQ) の表現力を捉えます。 VQ、LiftQuant のデコード パスは線形変換と 1 ビットの均一量子化器のみに依存しており、ハードウェアに優しい性質を維持しています。LiftQuant を使用すると、70B LLM を 24GB GPU に正確に適合させることができ、そのパフォーマンスは同じデバイスに搭載されている最先端の 2 ビット モデルを大幅に上回ります。 https://github.com/Heliulu/LiftQuant。

原文 (English)

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection

Existing quantization methods are fundamentally limited by rigid, integer-based bit-widths (e.g., 2, 3-bit), resulting in a ``deployment gap" where Large Language Models cannot be optimally fitted to specific memory budgets. To bridge this gap, we introduce LiftQuant, a novel framework that enables continuous bit-width control for true Pareto-optimal deployment. The core innovation is a ``lift-then-project" mechanism which approximates low-dimensional weight vectors by projecting a simple 1-bit lattice from a higher-dimensional ``lifted" space. Crucially, the effective bit-width is determined simply by the ratio of the lifted dimension to the original dimension, which allows the bit-width to be tuned quasi-continuous as the dimension is a flexible structural parameter. This projection generates a structured yet non-uniform codebook, capturing the expressive power of Vector Quantization (VQ). While beneficial over VQ, LiftQuant's decoding path relies solely on linear transformations and 1-bit uniform quantizers, retaining hardware-friendly nature. This flexibility is transformative: LiftQuant enables a 70B LLM to be compressed to 2.4 bits to precisely fit a 24GB GPU, where its performance significantly surpasses state-of-the-art 2-bit models fitted on the same device. Our code and ckpt is available at https://github.com/Heliulu/LiftQuant.

13:00 JSTLLM/生成AIエージェント

エージェント追跡から信頼へ: LLM エージェントにおける証拠追跡と実行来歴

大規模言語モデル (LLM) ベースのエージェントは、外部ツール、検索システム、メモリ モジュール、環境、その他のエージェントと対話することで、複雑なタスクを解決することが増えています。これらの機能により、エージェントの自律性が拡張されますが、エージェントの動作の検証、デバッグ、監査が難しくなります。最終回答の精度だけでは、出力がどのように生成されたか、各主張を裏付ける証拠は何か、ツールの呼び出しが正当化されたかどうか、記憶が後の決定にどのように影響したか、実行の失敗がどこで発生したかを説明することはできません。証拠追跡と実行来歴は、取得された証拠、ツール出力、メモリ項目、環境観察、中間クレーム、アクション、および最終的な回答がエージェントの実行全体を通じてどのように関連するかをモデル化することで、このギャップに対処します。この調査は、LLM エージェントにおける証拠の追跡と実行の出自に関する体系的なレビューと概念的な枠組みを提供します。私たちは、検索根拠、クレームサポート、ツール使用の安全性、メモリリネージ、可観測性、デバッグ、監査、リカバリを結び付ける、統一された来歴の観点に基づいて関連作業を整理します。トレースソース、証拠と実行単位、来歴関係、トレースの粒度とタイミング、表現形式、信頼関数を網羅する分類法を導入します。私たちは、出所の表現、証拠の帰属、ツール使用の出所、実行時のガードレール、出所を伴うメモリ、トレースベースの可観測性、障害診断など、主要な方法論の方向性を検討します。また、既存のベンチマーク、データセット、評価指標を来歴関連の機能にマッピングし、評価が最終的な回答の正しさからプロセスレベルの説明責任にどのように移行できるかについても説明します。最後に、統合トレース スキーマ、クレーム レベルおよびセマンティックの出所、出所を意識した安全メカニズム、現実的な実行トレース ベンチマーク、リカバリ指向の評価、プライバシーを意識した監査インフラストラクチャなどの未解決の課題について概説します。

原文 (English)

From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents

Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, retrieval, memory access, environmental interaction, and multi-agent collaboration. These capabilities expand agent autonomy, but also make agent behavior harder to verify, debug, and audit. Final-answer accuracy alone cannot explain how an output was produced, which evidence supported each claim, whether tool calls were justified, how memory influenced later decisions, or where failures originated. This survey examines evidence tracing and execution provenance as foundations for process-level accountability in trustworthy LLM agents. We define execution provenance as the typed graph of an agent execution and evidence tracing as its projection onto evidence-support relations. This perspective connects retrieval grounding, claim support, tool-use safety, memory lineage, observability, debugging, audit, and recovery within a unified framework. We introduce a taxonomy covering trace sources, evidence and execution units, provenance relations, tracing granularity and timing, representation forms, and trust functions. We then review key methodological directions, including provenance representation, evidence attribution, tool-use provenance, runtime guardrails, provenance-bearing memory, observability, and failure diagnosis. Finally, we discuss benchmarks, datasets, metrics, and open challenges for building provenance-aware, auditable, and recoverable agent systems.

13:00 JSTLLM/生成AI研究/論文

MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization

Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents…

13:00 JSTLLM/生成AI

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face ch…

13:00 JSTLLM/生成AIGPT / ChatGPTLlamaMistral AI

検索拡張生成における証拠グラフの一貫性: 幻覚検出のモデル依存分析

検索拡張生成 (RAG) は、大規模な言語モデルにおける幻覚を軽減しますが、排除しません。既存の検出方法は、生成された回答と取得された文章の間のフラットな類似性に依存しており、証拠部分と回答主張の間の構造的関係を無視しています。我々は、応答ごとに局所証拠グラフを構築し、幻覚指標として 5 つの構造的一貫性尺度を計算するフレームワークである証拠グラフ一貫性 (EGC) を提案します。 6 つの LLM (回答数 5,767) にわたる RAGTruth の完全な質問回答分割で評価した EGC は、一貫したモデルファミリー分割を明らかにしました。グラフの一貫性特徴は、Llama-2 モデルの幻覚について予想される診断方向を示しますが、GPT-4、GPT-3.5、および Mistral-7B では系統的な逆転を示します。この逆転は、モデルファミリー全体で質的に異なる幻覚パターンを示唆しており、埋め込みベースのグラフの一貫性がモデルに依存しない幻覚検出信号として機能できないことを示しています。

原文 (English)

Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection

Retrieval-Augmented Generation (RAG) reduces but does not eliminate hallucination in large language models. Existing detection methods rely on flat similarity between generated answers and retrieved passages, ignoring structural relationships among evidence pieces and answer claims. We propose Evidence Graph Consistency (EGC), a framework that constructs a local evidence graph per response and computes five structural consistency measures as hallucination indicators. Evaluated on the full question answering split of RAGTruth across six LLMs (5,767 responses), EGC reveals a consistent model-family split: graph consistency features show the expected diagnostic direction for hallucinations in Llama-2 models but exhibit systematic reversal in GPT-4, GPT-3.5, and Mistral-7B. This reversal suggests qualitatively different hallucination patterns across model families and indicates that embedding-based graph consistency cannot serve as a model-independent hallucination detection signal.

13:00 JSTエージェント

An AI Security Agent for University ACMIS: Multi-Vector Threat Detection and Automated Response

University Academic Management Information Systems (ACMIS) are high-value targets for a wide spectrum of security threats including brute-f…

13:00 JSTLLM/生成AILlama

APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing

W4A4 quantization promises full utilization of INT4 Tensor Cores, yet group dequantization overhead on CUDA Cores has driven existing syste…

13:00 JSTエージェント

Agentic Social Affordance Framework (ASAF): マルチエージェント システムにおけるコラボレーション インターフェイスとしてのエージェント アイデンティティ設計

AI システムが単一の会話型エージェントから複雑なマルチエージェント アーキテクチャに進化するにつれて、個々のエージェントの社会的アイデンティティがコラボレーション内で人間の行動をどのように形作るかという重要な設計側面が見落とされてきました。この論文では、ソーシャル アフォーダンス理論をマルチエージェント AI システムのコンテキストに拡張する理論的フレームワークであるエージェントティック ソーシャル アフォーダンス フレームワーク (ASAF) を紹介します。私たちは、エージェントのアイデンティティ設計が、単にユーザー インターフェイスの規約として機能するのではなく、コラボレーション インターフェイスとして機能し、ユーザーが各エージェントをどのように認識し、アプローチし、関与するかを構造化し、それによってヒューマン エージェントとエージェントのコラボレーション結果の品質に影響を与えることを提案します。具体的には、ソーシャル アフォーダンス層は、エンジニアリング オーケストレーションと直交する独立した設計次元を構成します。この 2 つは、相互に導き出すことができない別個の意思決定空間を表します。 ASAF は、アイデンティティ シグナリング、行動プライミング、および協調ガバナンスの 3 つのメカニズムで構成され、4 層のアイデンティティ シグナル忠実度スペクトルと個人差緩和変数 (擬人化認知スタイルとツール化認知スタイル) を通じてそれらの境界条件を指定します。我々は ASAF を既存のアフォーダンス理論と CASA パラダイムとの関連で位置づけ、ASAF のマルチエージェント、トポロジーレベルの予測が二項フレームワークの説明範囲を超える箇所を描写します。マルチエージェント システム設計への影響について議論し、設計空間の直交性をテストする要因計画など、将来の実証的検証の方向性を概説します。

原文 (English)

Agentic Social Affordance Framework (ASAF): Agent Identity Design as a Collaboration Interface in Multi-Agent Systems

As AI systems evolve from single agents to multi-agent architectures, a critical design dimension has been overlooked: how the social identity of individual agents shapes human behavior within the collaboration. This paper introduces the Agentic Social Affordance Framework (ASAF), a theoretical framework extending Social Affordance theory to multi-agent AI systems. We propose that agent identity design functions as a collaboration interface--structuring how users perceive and engage with each agent, and thereby influencing Human-Agent collaboration outcomes. ASAF adopts the analytical separability of the social affordance layer and the engineering orchestration layer as a framing assumption--an organizing distinction that structures design analysis--rather than a testable claim about effect-independence. ASAF comprises three mechanisms: Identity Signaling, Behavioral Priming, and Collaborative Governance, and specifies their boundary conditions through a four-tier Identity Signal Fidelity Spectrum and an individual-difference moderating variable (anthropomorphizing vs. instrumentalizing cognitive style). We situate ASAF relative to affordance theory (Hutchby, 2001), the CASA paradigm (Gambino et al., 2020), and classical multi-agent systems research (Wooldridge & Jennings, 1995), identifying a directional reversal: where classical MAS used roles, norms, and coordination to constrain autonomous agents, ASAF applies the same organizational vocabulary to structure the cognition and oversight of human operators who remain in the loop. ASAF positions social affordance design as a first-class design responsibility that engineering orchestration cannot subsume. We outline directions for empirical validation, including a factorial design characterizing the empirical interaction surface between the social affordance and engineering orchestration layers.

13:00 JSTLLM/生成AIGPT / ChatGPTLlama

言語モデル蒸留における潜在意識の行動伝達率の定量化

良性の動作を生徒モデルに伝達することを目的とした言語モデルの蒸留は、望ましくない特性が教師モデルに存在する場合、それを伝達する可能性もあり、これは潜在意識学習として知られる現象です。定性的証拠はこの効果の存在を裏付けていますが、その規模は体系的に特徴づけられていません。この研究では、2 つの教師モデル (Llama-2-7B-Chat と Qwen2.5-7B-Instruct) をさまざまなステアリング強度でステアリングし、良性のデータのみを使用して生徒モデルを抽出することにより、サブリミナル行動伝達率を定量化します。 GPT-4.1 を評価者として使用して 100 個の JailbreakBench プロンプトを評価したところ、転送は堅牢であるものの、独特のスケーリング動作を示すことがわかりました。 Llama-2 は鋭いしきい値 ($\tau = {0.25,0.32} \ \text{beyond} \ \alpha = -0.15$) を示しますが、Qwen2.5 は継続的かつより高いレベルの転送 ($\tau$ 〜 $0.61$) を示します。

原文 (English)

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation

Distillation of a language model intended to transfer benign behavior to a student model may also transfer undesirable characteristics, if they are present in the teacher model, a phenomenon known as subliminal learning. While qualitative evidence supports the existence of this effect, its magnitude has not been systematically characterized. This study quantifies subliminal behavioral transfer ratios by steering two teacher models (Llama-2-7B-Chat and Qwen2.5-7B-Instruct) at varying steering strengths and distilling student models using only benign data. Evaluation on 100 JailbreakBench prompts with GPT-4.1, serving as the evaluator, indicates that transfer is robust but exhibits distinct scaling behaviors. Llama-2 demonstrates a sharp threshold ($\tau = {0.25,0.32} \ \text{beyond} \ \alpha = -0.15$), whereas Qwen2.5 displays continuous and higher levels of transfer ($\tau$ up to $0.61$).

13:00 JSTエージェント

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical…

13:00 JSTLLM/生成AI

MeEvo: Metacognitive Evolution Combined with Natural Evolution for Automatic Heuristic Design

Large Language Models (LLMs) have advanced Automatic Heuristic Design (AHD) by enabling heuristic generation through reasoning and code syn…

13:00 JSTLLM/生成AI

CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation

Granting LLMs direct control over costly, irreversible scientific experiments leads to unsafe exploration and unstable performance, but dis…

13:00 JST画像/動画生成ロボティクス

X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining

Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing…

13:00 JST画像/動画生成

EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining

Color fundus photography (CFP) is the mainstay of large-scale retinal screening, but its diagnostic capacity is limited by the lack of dept…

13:00 JST研究/論文

Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling

Pairwise comparison is the gold standard for subjective ranking tasks; however, exhaustive annotation requires a massive number of human co…

13:00 JST研究/論文

Entropy-Gated Latent Recursion

Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity…

13:00 JSTエージェント

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

Banks face two threat families with fundamentally different detection requirements: signature-based fraud (card-not-present attacks, accoun…

13:00 JST研究/論文

資本の知識理論:自然知能と人工知能の価値

この巻では、生産能力がソフトウェア、データ、モデル、ルーチン、専門知識、プラットフォーム、組織、コモンズ、および公的認識インフラストラクチャにますます存在する経済のための資本の知識理論を展開します。アダム・スミスの労働、株式、専門化、市場範囲の理論から始まり、知識が株式のようになり、さまざまな形で移動可能になり、拡張可能で、統治可能で、組み換え可能で、会計において不完全に見えるようになったときに何が変わるのかを問います。この本では、知識を伴うストックを中心的な対象として紹介し、それがどのように生成され、統治可能な形式に変換され、展開され、フィードバックを通じて改善され、封じ込めまたは共有され、測定され、損なわれ、将来の生産へのインプットとして使用されるかを分析しています。それは、身体的、非身体的、制度化された、コモンズ、および公共の知識形式を区別し、最初の変換、認知的囲い込み、フィードバックの捕捉、闇資本、予想される知識の損失などの概念を開発します。この議論は条件付きで検証可能である。現代の富は資本蓄積だけでなく、生産的な知識がどのように管理されるかにも依存している。

原文 (English)

A Knowledge Theory of Capital:The Value of Natural and Artificial Intelligence, Volume 1

This volume develops a knowledge theory of capital for economies in which productive capacity increasingly resides in software, data, models, routines, expertise, platforms, organizations, commons, and public epistemic infrastructure. Beginning from Adam Smith's theory of labour, stock, specialization, and market extent, it asks what changes when knowledge becomes stock-like, mobile across forms, scalable, governable, recombinable, and imperfectly visible in accounting. The book introduces knowledge-bearing stock as the central object and analyses how it is generated, converted into governable form, deployed, improved through feedback, enclosed or shared, measured, impaired, and used as input to future production. It distinguishes embodied, disembodied, institutionalized, commons, and public knowledge forms and develops concepts such as first conversion, cognitive enclosure, feedback capture, dark capital, and expected knowledge loss. The argument is conditional and testable: modern wealth depends not only on capital accumulation, but on how productive knowledge is governed.

13:00 JST研究/論文

RankGraph-2: 推奨事項における 10 億ノードのグラフ学習のためのライフサイクル協調設計

10億ノード規模のグラフベースの検索には、グラフ構築、表現学習、リアルタイム処理という3つの密結合した問題を共同で解決する必要があるが、既存の作業はそれぞれを個別に解決している。我々は、Meta に導入されたフレームワークである RankGraph-2 を紹介します。これは、類似性に基づく検索 (U2U2I および U2I2I) の 3 つのライフサイクル ステージすべてを共同設計し、各ステージの要件が他のステージの要件を形成します。サービスを提供するには、高価なオンライン KNN を回避するために共学習されたクラスター インデックスが必要です。これにより、インデックスの共トレーニングがトレーニング目標に組み込まれます。トレーニングでは、類似性に基づく検索が事前計算された近傍を許容し、オンライン グラフ インフラストラクチャが不要になるという観察から恩恵を受けます。これには、自己完結型データを生成するための構築が必要です。構築では、項目範囲の時間レベルの更新もサポートする必要があります。これらのカスケード要件に基づいて、RankGraph-2 は、人気バイアス補正を備えたサブサンプリングによって数百兆のエッジを数千億に削減し、パーソナライズされた PageRank によってマルチホップ近傍を事前計算し、サービスの計算コストを 83% 削減する残差量子化クラスター インデックスを共同学習します。このライフサイクルの共同設計により、シンプルなアーキテクチャで、二部グラフでは GAT + Deep Graph Infomax モデルよりも 3.8 倍高い再現率、項目検索では PyTorch-BigGraph よりも 2.1 倍高い再現率を達成できます。 RankGraph-2 は最大 +0.96% の CTR と +2.75% の CVR を実現し、主要なサーフェス全体で 20 回以上の検索起動を実現しました。

原文 (English)

RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation

Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation learning, and real-time serving -- yet existing work addresses each in isolation. We present RankGraph-2, a framework deployed at Meta that co-designs all three lifecycle stages for similarity-based retrieval (U2U2I and U2I2I), where each stage's requirements shape the others. Serving requires a co-learned cluster index to avoid expensive online KNN -- this pushes index co-training into the training objective. Training benefits from the observation that similarity-based retrieval tolerates pre-computed neighborhoods, eliminating online graph infrastructure -- this requires construction to produce self-contained data. Construction must also support hour-level refresh for item coverage. Acting on these cascading requirements, RankGraph-2 reduces hundreds of trillions of edges to hundreds of billions via subsampling with popularity bias correction, pre-computes multi-hop neighborhoods via personalized PageRank, and co-learns a residual-quantization cluster index that reduces serving computational cost by 83%. This lifecycle co-design enables a simple architecture to achieve 3.8 x higher recall than a GAT + Deep Graph Infomax model on a bipartite graph and 2.1 x higher than PyTorch-BigGraph on item retrieval. RankGraph-2 delivers up to +0.96% CTR and +2.75% CVR, and has powered 20+ retrieval launches across major surfaces.

13:00 JSTエージェントロボティクス

NeuralMUSIC: ロボット音源位置特定のためのハイブリッド神経部分空間フレームワーク

信頼性の高い音源定位はロボットの聴覚の基礎であり、自律ロボットが空間的な手がかりを認識し、動的な環境で効果的に動作できるようになります。多重信号分類 (MUSIC) などの古典的な手法は強力な理論的基盤を提供しますが、信号対雑音比が低いと性能が低下します。深層学習ベースのアプローチは有望なパフォーマンスを達成しますが、多くの場合、条件全体にわたる限られた一般化に苦労します。これらの課題に対処するために、ロボットによる音源定位のためのハイブリッド神経部分空間フレームワークである NeuralMUSIC を提案します。具体的には、ニューラル ネットワークはまず、マルチチャネル マイクの観測値から空間共分散行列を推定します。予測された共分散は、固有値分解 (EVD) と擬似スペクトル計算を使用して古典的な MUSIC パイプラインに統合され、その後、周波数アテンション フュージョン (FAF) モジュールによって最終的な DOA 推定値が生成されます。データ効率を向上させるために、ラベルなしの音響データを活用して空間構造を捕捉する自己教師付き空間相関学習 (SSCL) 戦略をさらに導入します。さまざまなロボット タスクにわたる広範な実験により、NeuralMUSIC が堅牢性とクロスドメイン汎用性の向上を示しながら、競争力のある位置特定精度を達成できることが実証されました。

原文 (English)

NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effectively in dynamic environments. Classical methods such as Multiple Signal Classification (MUSIC) offer strong theoretical foundations but degrade under low signal-to-noise ratios. While deep learning-based approaches achieve promising performance, they often struggle with limited generalization across conditions. To address these challenges, we propose NeuralMUSIC, a hybrid neural-subspace framework for robotic sound source localization. Specifically, a neural network first estimates the spatial covariance matrix from multichannel microphone observations. The predicted covariance is then integrated into a classical MUSIC pipeline with eigenvalue decomposition (EVD) and pseudo-spectrum computation, followed by a Frequency Attention Fusion (FAF) module to produce the final DOA estimates. To improve data efficiency, we further introduce a Self-supervised Spatial Correlation Learning (SSCL) strategy that leverages unlabeled acoustic data to capture spatial structure. Extensive experiments across different robotic tasks demonstrate that NeuralMUSIC achieves competitive localization accuracy while exhibiting improved robustness and cross-domain generalization.

13:00 JST研究/論文GPT / ChatGPTLlama

Explaining Attention with Program Synthesis

A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic desc…

13:00 JST研究/論文

事前トレーニングデータ構成によるエンジニアリングスケーリング則に向けて

ニューラル スケーリングの法則は、コンピューティング、モデル サイズ、データセット サイズにおけるべき乗則としてモデルのパフォーマンスがどのように向上するかを記述します。これらの関係は大規模な言語モデルでは十分に確立されていますが、素粒子物理学における大規模なモデルでも明らかになりつつあります。言語と同様に、パフォーマンスはべき乗則としてスケールされることが実証研究によって示されています。ただし、自然言語や画像の領域とは異なり、基礎物理学には合成データを安価に生成する忠実度の高いシミュレーターがあります。これにより、追加データの方が追加パラメーターよりも安価になるスケーリング方式が有利になり、スケーリングに影響を与えるように事前トレーニング データセット自体を操作できるようになります。高エネルギー粒子ビームの衝突で生成されるハドロンジェットを分類するタスクでは、より多様で下流の分類タスクとよりよく連携する事前学習データを含めることにより、大規模なモデルではなくより多くのデータを必要とするようにスケーリング動作を設計できることを示します。

原文 (English)

Towards Engineering Scaling Laws with Pretraining Data Composition

Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established for large language models, these relationships are emerging for large models in particle physics. As with language, empirical studies show that the performance scales as a power law. However, unlike natural language or image domains, fundamental physics has high-fidelity simulators that produce synthetic data cheaply. This favors scaling regimes where additional data is cheaper than additional parameters, and allows the pretraining dataset itself to be engineered to influence the scaling. For the task of classifying hadronic jets produced in collisions of high-energy particle beams, we show that the scaling behavior can be engineered towards requiring more data rather than larger models by inclusion of pretraining data which is more diverse and better aligned with the downstream classification task.

13:00 JSTエージェント

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coord…

13:00 JST画像/動画生成研究/論文

SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture rada…

13:00 JST研究/論文

Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: An Experimental Investigation

Background: Generative artificial intelligence (GenAI) is increasingly used for health information, yet its influence on users' trust calib…

13:00 JST研究/論文

Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling

Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…

13:00 JSTエージェントハードウェア/半導体

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems

Agentic AI applications compose multiple model calls and tool executions, creating new scheduling challenges for GPU-CPU clusters. Their in…

13:00 JST研究/論文

トリガーの学習: 大型ハドロン衝突型加速器での強化学習

大型ハドロン衝突型加速器などの高スループットの科学施設は、帯域幅、遅延、ストレージに厳しい制約がある中で、リアルタイム イベント フィルタリング (\textit{triggering}) に依存しています。実際には、トリガー メニューはほとんどが静的で手動調整されており、検出器の状態、パイルアップ、および背景の構成が時間の経過とともにドリフトするため、最適ではなくなる可能性があります。オンラインしきい値調整を逐次的な意思決定の問題として捉えます。強化学習エージェントは、最近のレートと信号に敏感な機能のストリーミング サマリーを取り込み、許容範囲内でターゲット バックグラウンド レートを追跡しながら、信号効率を最大化するためにトリガーしきい値を更新します。 Group-Filtered Policy Optimization (GFPO) をストリーミング制御に適応させ、トレーニング中にバックグラウンド レートの実現可能性を強制する 2 つのバリアント (GFPO-F、GFPO-FR) を導入します。現実的な衝突型加速器の動作をエミュレートするベンチマークで、我々は 2 つの代表的なトリガーを研究します。パイルアップ変動に敏感な総横エネルギー ($H_{T}$) トリガーと、稀なまたは非標準のシグネチャの再構成損失に基づく異常検出 (AD) トリガーです。モンテカルロ ストリームでは、エージェントは許容範囲内時間間隔の割合を 48\% ($H_T$) および 28\% (AD) 増加させ、これらの許容範囲内間隔での信号効率の累積ゲインは最大 2\% になります。シミュレーションから \emph{real} 衝突データ (CMS Run 283408) に移行すると、同じエージェントは微調整を行わずに、ベースラインに対して 56\% ($H_T$) および 28\% (AD) の許容範囲内改善を達成し、両方のトリガーで信号効率がさらに向上しました。私たちの知る限り、これは実際の大型ハドロン衝突型加速器の衝突データに対する RL ベースのトリガー制御の \emph{最初}のデモンストレーションです。コードは https://github.com/Zixind/GFPO\_LHC で入手できます。

原文 (English)

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement learning agent ingests streaming summaries of recent rates and signal-sensitive features and updates trigger thresholds to maximize signal efficiency while tracking a target background rate within a tolerance band. We adapt Group-Filtered Policy Optimization (GFPO) to streaming control and introduce two variants (GFPO-F, GFPO-FR) that enforce background rate feasibility during training. On a benchmark that emulates realistic collider operation, we study two representative triggers: a total transverse energy ($H_{T}$) trigger sensitive to pileup variation, and an anomaly-detection (AD) trigger based on reconstruction loss for rare or non-standard signatures. On Monte Carlo streams, our agent increases the fraction of in-tolerance time intervals by 48\% ($H_T$) and 28\% (AD), with a cumulative gain of up to 2\% in signal efficiency on those in-tolerance intervals. Transferring from simulation to \emph{real} collision data (CMS Run 283408), the same agent, without fine-tuning, achieves a 56\% ($H_T$) and 28\% (AD) in-tolerance improvement over baselines, with further signal-efficiency gain on both triggers. To our knowledge, this is the \emph{first} demonstration of RL-based trigger control on real Large Hadron Collider collision data. Code is available at https://github.com/Zixind/GFPO_LHC (see repo for details).

13:00 JSTLLM/生成AI

仕様の学習に向けて: 優先ペアからの推論時間の調整

大規模言語モデル (LLM) を望ましい動作に導くには、通常、モデルの応答を注意深く検査してプロンプトを手作りする反復プロセスに依存します。これは複雑で脆弱で、エラーが発生しやすいプロセスです。設定ベースの微調整はより厳密ですが、多くの場合法外に高価なソリューションです。私たちは、短いユーザーの指示と少数の好みの判断に依存するフレームワークであるスペック学習を提案します。これらは、LLM の自然言語プロンプトの形式で仕様にコンパイルされます。仕様は推論時に LLM を条件付けするため、基礎となるモデルに対するパラメーターの更新は必要ありません。コンパイルされた仕様に基づいて生成された応答は、優先信号が高密度である特殊なドメインからのデータセットに対する直接優先最適化 (DPO) よりも優れたパフォーマンスを発揮することが多いことを示します。不透明な重み更新とは異なり、結果として得られる仕様は人間が判読可能であり、それを生成した優先信号の解釈可能で透明な文書化された実施形態としても機能します。

原文 (English)

Towards Spec Learning: Inference-Time Alignment from Preference Pairs

Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's responses. This is an involved, brittle, and error-prone process. Preference-based fine-tuning is a more rigorous but often prohibitively expensive solution. We propose spec learning, a framework that relies on a brief user instruction and a small set of preference judgments. These are compiled into specifications in the form of natural-language prompts for an LLM. Specifications condition LLMs at inference time, and no parameter updates to the underlying models are required. We show that the responses generated based on the compiled specifications often outperform direct preference optimization (DPO) on datasets from specialized domains whose preference signal is dense. Unlike opaque weight updates, the resulting specifications are human-readable and double as interpretable and transparent written embodiments of the preference signal that produced them.

13:00 JST画像/動画生成

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency…

13:00 JST研究/論文

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory

Modern large language models based on softmax scaled-dot-product attention are constrained by their training sequence length: as the key-va…

13:00 JSTLLM/生成AIビジネス/資金調達

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

A language model's memory can be worse than no memory at all. A memory that keeps a wrong conclusion but drops the work behind it makes the…

13:00 JSTエージェント研究/論文

Red Queen G\"odel Machine: 共進化するエージェントとその評価者

自己改善エージェントは、エージェント コーディング ベンチマークにおける最先端 (SOTA) であり、最近では一般的なドメインに拡張されています。ただし、それらの検索方法は通常、エージェントが改善しても有効であり続ける固定の検証基準、ベンチマーク、またはラベル付きデータセットといっ​​た固定的な評価基準を前提としています。これは進化の中心的な特徴、つまり種は環境の変化に応じて適応するという点を無視している。私たちは、再帰的な自己改善にも同じ原理を導入し、評価を改善ループの一部とし、進化する評価者、敵対的な目標、静的なベンチマークを超える可能性のある動的なユーティリティへの探索を開くことを目指しています。非定常ユーティリティ下での再帰的自己改善のための進化的フレームワークである Red Queen Godel Machine (RQGM) を紹介します。 RQGM は、制御されたユーティリティの進化を通じてこれを可能にします。検索は、固定されたエポック内評価基準を使用してエポックに編成されますが、ユーティリティはエポック境界で更新できるため、目的がエポック全体で進化するにつれて自己改善の保証がエポックごとに保持されます。まず、検証可能なコーディングタスクであっても、RQGM が補完的なエージェントとしてのジャッジコードレビュー信号を追加することにより、以前の SOTA よりもテスト合格率を向上させることを示します。このシグナルは安価で、RQGM が使用するトークンの量は 1.35 倍から 1.72 倍少なくなります。次に、科学論文の執筆と査読、オリンピックレベルの校正と採点に移ります。そこでは、RQGM が以前の自己改善エージェントに比べてパフォーマンスを向上させます。多様な審査員パネルの下で、共進化したライターは 1.78 倍から 1.86 倍高い合格率に達し、同時に進化した採点者は 9% 高いグラウンドトゥルース精度に達します。論文査読では、最も強力なベースライン査読者が AI によって生成された論文を人間の割合の最大 1.91 倍で過剰に受け入れます。 RQGM は、AI と人間の作業に対して同等に厳しいレビュー担当者を発見する敵対的目標を導入することでこれを修正します。

原文 (English)

The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators

Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid as the agent improves. This ignores a central feature of evolution: species adapt as their environments change with them. We aim to bring the same principle to recursive self-improvement, making evaluation part of the improvement loop and opening search to evolving evaluators, adversarial objectives, and dynamic utilities that may surpass static benchmarks. We introduce the Red Queen Godel Machine (RQGM), an evolutionary framework for recursive self-improvement under non-stationary utilities. The RQGM makes this possible through controlled utility evolution: search is organized into epochs with a fixed within-epoch evaluation criterion, while the utility can be updated at epoch boundaries, so self-improvement guarantees hold per epoch as the objective evolves across them. We begin by showing that even on verifiable coding tasks, the RQGM improves test pass rate over the prior SOTA by adding a complementary agent-as-a-judge code-review signal. This signal is cheaper and the RQGM uses 1.35x-1.72x fewer tokens. We then turn to scientific paper writing and reviewing, and Olympiad-level proof writing and grading, where the RQGM improves performance over prior self-improving agents: co-evolved writers reach 1.78x-1.86x higher acceptance rates under a diverse agent-as-a-judge panel, while co-evolved graders reach 9% higher ground-truth accuracy. In paper reviewing, the strongest baseline reviewer over-accepts AI-generated papers at up to 1.91x the human rate. The RQGM corrects this by introducing an adversarial objective that discovers reviewers equally stringent on AI and human work.

13:00 JST研究/論文

ハイブリッド プライバシーを意識したセマンティック検索: SVD で切り詰められたドキュメント ジオメトリと、制限された脅威モデルに基づく CKKS 暗号化クエリの再ランキング

高密度埋め込みはセマンティック検索と検索拡張生成を強化しますが、埋め込み反転攻撃はベクトルからソース テキストを再構築する可能性があります。ベクトル データベースが漏洩すると、その背後にある文書も漏洩します。教科書的な防御策は極端です。検索全体を準同型的に暗号化するのは健全ですが、100 万文書規模では遅すぎます。その一方で、保護するずっと前にプライバシー ノイズによってランキングが低下します。静的コレクションと動的クエリの間の非対称性を利用した中間パスを研究します。コレクションは幾何学的に保護されています。各ベクトルは低次元の SVD 部分空間上で切り詰められ、所有者のみが知っている秘密の直交変換によって回転されます。クエリは暗号的に保護されています。クエリは CKKS 準同型暗号化の下で再ランク付けされるため、正直だが好奇心旺盛なサーバーはクエリやスコアを見ることはありません。 CKKS パラメータは、小規模なオフライン ベンチマークから取得されます。私たちは、保護された部分空間に限定された攻撃者の再構成エラーの厳しい下限を証明します。 100 万のドキュメントと 5 つのエンコーダでは、このスキームは 1 秒未満のレイテンシでランキングの品質を維持し (線形デノイザーとして強力なエンコーダでわずかに向上します)、保護されたスペースに対する既製の反転攻撃はノイズ フロアまで崩壊します。次に、より強力な敵対者をテストします。既知の平文攻撃者は、保持された次元とほぼ同じ数の漏洩ペアから直交プロクラステスによる回転を回復します。公開されている積量子化コードは、最近傍構造を保存します。ランダム投影、校正済みノイズ、および BEIR ベースラインは、切り捨てが無料のデノイザーではなく、エンコーダーに依存する精度コストであることを示しています。私たちは限界を述べています。クエリの機密性は暗号化されていますが、ドキュメントの保護は経験的な難読化レイヤー (SVD の切り捨てと秘密のローテーション) であり、暗号化のプリミティブではありません。また、各主張の脅威モデルを区切ります。

原文 (English)

Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model

Dense embeddings power semantic search and retrieval-augmented generation, yet a leaked vector database also leaks the text behind it, because embeddings can be inverted with high fidelity. Fully homomorphic search is sound but far too slow at million-document scale, while privacy noise degrades ranking before it protects. We study a middle path built on an asymmetry: the static document collection is protected geometrically - each vector is SVD-truncated onto a lower-dimensional subspace and rotated by a secret orthogonal transform held only by the data owner - while the dynamic query is protected cryptographically under CKKS, so an honest-but-curious server never sees query values or similarity scores. We prove a tight lower bound on the reconstruction error of any decoder confined to the protected subspace. On a one-million-document corpus with five encoders the protection preserves - and on the strongest encoders slightly improves - retrieval quality, a linear-denoiser effect, at sub-second latency, while an off-the-shelf inversion attack collapses to the noise floor. We also quantify the boundary: a known-plaintext attacker recovers the secret rotation by orthogonal Procrustes from about as many leaked pairs as the retained dimension. The same asymmetric geometry doubles as a privacy-preserving semantic data-loss-prevention primitive for LLM firewalls: a server holding only the protected vectors detects whether a candidate matches a confidential reference corpus at near parity with a plaintext detector, degrading gracefully under text obfuscation. We state the limits plainly: query confidentiality is cryptographic, but document protection rests on SVD truncation and a secret rotation that form an empirical obfuscation layer, not a cryptographic primitive, under a clearly delimited threat model.

13:00 JSTロボティクスハードウェア/半導体

CoStream: 一般化可能な複雑な操作のための単純な動作の構築

GPU を PCIe スロットに装着するなど、長期にわたる接触が多い複雑な操作タスクでは、ミリメートル単位の高精度と、新しいタスクに対するすぐに使える汎用性の両方が必要です。既存のパラダイムは両方を満たすのに苦労しています。古典的なパイプラインは、高精度の制御を実現するために脆弱なタスク固有のインターフェイスを使用しますが、新しいタスクに適応するにはコストのかかるパイプラインの再設計が必要です。一方、モノリシックなエンドツーエンドのポリシーはより優れた一般化を提供しますが、新しいデータで再トレーニングしない限り、複雑な分散外のタスクでは高精度が得られません。どちらのパラダイムも暗黙の前提を共有しています。つまり、操作機能を取得したら、それを自由に分解したり再構成したりするのではなく、厳格なパイプラインまたはモノリシック全体として展開する必要があるということです。この論文では、複雑な操作機能が単純で独立した動作の構成から自然に出現する可能性があることを示します。モノリシックなポリシーや厳格なパイプラインを展開するのではなく、基礎モデルと多様なセンシング モダリティを複数の構成可能なコア動作に統合するフレームワークを提案します。つまり、基礎モデルを介して空間制約を抽出するセマンティック動作。想像上のビデオ内のキーポイントを追跡することで軌道を予測する予測行動。そして高周波の触覚と力の補正を提供する反応的な動作。共有 $SE(3)$ インターフェイスでは、これらの出力は右乗算によって各制御ステップで単一のポーズ コマンドに構成され、準拠したコントローラーによって実行されます。私たちは、日常的な操作と精密な組み立てに及ぶ 8 つの現実世界のタスクを実証し、接触の多い組み立てとオブジェクトの転送で最も大きな効果を発揮し、実行中の手動による混乱からの堅牢な回復を示します。 {ウェブサイト:} https://costream-simple.github.io

原文 (English)

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and out-of-the-box generalization to new tasks. Existing paradigms struggle to satisfy both: classical pipelines use brittle, task-specific interfaces to achieve high-precision control but require costly pipeline redesigns to adapt to new tasks, whereas monolithic end-to-end policies provide better generalization but lack high precision on complex, out-of-distribution tasks unless retrained with new data. Both paradigms share an implicit assumption: once a manipulation capability is acquired, it must be deployed as a rigid pipeline or monolithic whole, rather than being freely decomposed and recomposed. In this paper, we show that complex manipulation capabilities can emerge naturally from the composition of simple, independent behaviors. Rather than deploying a monolithic policy or a rigid pipeline, we propose CoStream, a framework orchestrating foundation models and diverse sensing modalities into multiple composable core behaviors: a semantic behavior extracting spatial constraints via foundation models; a predictive behavior forecasting trajectories by tracking keypoints in imagined videos; and a reactive behavior providing high-frequency tactile and force corrections. On a shared $SE(3)$ interface, these outputs compose by right-multiplication into a single pose command at each control step, executed by a compliant controller. We demonstrate CoStream on 8 real-world tasks spanning everyday manipulation and precision assembly, with the strongest gains in contact-rich assembly and object transfer, and show robust recovery from manual perturbations during execution. Website: https://costream-simple.github.io

13:00 JST画像/動画生成

堅牢なタマネギ: ノイズの下で開いた語彙オブジェクト検出器を剥がす

Open Vocabulary Object Detector (OV-OD) に対する現実世界のノイズの影響は、そのアーキテクチャの複雑さのため、依然としてよく理解されていません。私たちは、制御された合成視覚的劣化を使用して OV-OD を層ごとに剥離する実証研究である包括的な分析 Robust Onion を紹介します。これにより、ロバスト性がどのように、なぜ、どこで劣化するかを明らかにし、特徴の崩壊を系統的に分析します。私たちの調査結果では、同様の視覚バックボーンを持つモデルは、同様のレイヤーでの同様の機能崩壊によって駆動され、同等の堅牢性を示す一方、事前トレーニング戦略、アーキテクチャのニュアンス、キャプションの監視などの要因はほとんど寄与していないことが明らかになりました。堅牢性は主にアノテーションではなく画像ドメインによって支配されており、これは COCO と LVIS に対する同様の堅牢性の影響と、なぜ ODinW-13 のようなデータセットが大きく孤立したオブジェクトによって堅牢性が誇張されている印象を与える可能性があるかを説明しています。最後に、エンドツーエンドのトレーニングに比べて 96 分の 1 少ないトレーニング可能なパラメータを使用し、軽量のプラグアンドプレイ NN および TK0 アプローチを通じて、実際の BDD100K、 WiderFace、VisDRONE での堅牢性を向上させることで洞察を検証します。また、先行研究のロバスト性の観察についても説明します。

原文 (English)

Robust Onion: Peeling Open Vocab Object Detectors Under Noise

The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis Robust Onion, an empirical study that uses controlled synthetic visual degradations to peel OV-ODs layer-by-layer, revealing how, why, and where robustness degrades, systematically analyzing feature collapse. Our findings reveal that models with similar vision backbones exhibit comparable robustness, driven by similar feature collapse at similar layers, while factors such as pretraining strategy, architectural nuances, and caption supervision contribute little. Robustness is primarily governed by the image domain rather than annotations, explaining the similar robustness impact on COCO and LVIS, and why datasets like ODinW-13 can give an impression of inflated robustness due to large, isolated objects. Finally, we validate our insights by improving robustness on real-world BDD100K, WiderFace, and VisDRONE via our lightweight plug-and-play NN & TK0 approach, using 96x fewer trainable parameters than end-to-end training. We also explain the prior works' robustness observations.

13:00 JSTLLM/生成AI

CARVE: チャンク並列リニア アテンションの価値効率を備えたコンテンツ認識型リカレント

リカレントモデルは記憶するために忘れなければなりませんが、最先端の技術では、何が保存されているかを考慮せずに何を消去するかを決定します。ゲートは到着したトークンのみを認識し、変更しようとしているメモリは認識しません。このメモリ ブラインド ゲーティングは、主要なデルタ ルール アーキテクチャ (GDN-2) の 3 つの複合欠陥のうちの 1 つです。値軸消去マスクは、値射影のスケールでパラメータを無駄にし、--私たちが証明しているように--反復トレーニングを Transformers と競合させる WY 形式の三角形チャンク ソルバーを数学的に阻止します。 CARVE (Content-Aware Recurrent with Value Efficiency) を導入します。これは、キー軸上でのみ消去するという 1 つの原則によって 3 つの問題すべてを解決します。これは、WY 形式ソルバーが有効であり続けるために必要かつ十分であることが証明されています。その中で、CARVE は、GPU メモリに既に書き込まれているリカレント出力テンソルを消去ゲートの空きコンテンツ信号として再利用し、値ごとの書き込みゲート投影をヘッドごとの単一のスカラーに置き換えます。初期化では、CARVE は GDN-2 とビット同一です。品質の違いは、コンテンツ ゲートが学習した内容から生じます。 100B トークンでトレーニングされた 1.3B パラメーターで、CARVE は WikiText のパープレキシティ 15.72 (GDN-2 に対してマイナス 0.18、4.5 シグマ効果) を達成し、9 つの常識的推論ベンチマークですべての反復ベースラインをリードし、すべての RULER 検索プローブで最先端を設定します。スループット オーバーヘッドは 0.4%、ピーク メモリは 13% 低く、パラメータが 19% 減少しました。 6 つの形式的定理は、メモリ容量、リアプノフ安定性、勾配流、表現力分離、パレート最適チャンク サイズ、およびハイブリッド最適性をカバーします。

原文 (English)

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention

Recurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the gate sees only the arriving token, not the memory it is about to modify. This memory-blind gating is one of three coupled defects in the leading delta-rule architecture (GDN-2): the value-axis erase mask wastes parameters at the scale of the value projection, and -- as we prove -- mathematically prevents the WY-form triangular chunk solver that makes recurrent training competitive with Transformers. We introduce CARVE (Content-Aware Recurrent with Value Efficiency), which resolves all three problems through one principle: erase only on the key axis. This is provably necessary and sufficient for the WY-form solver to remain valid. Within it, CARVE reuses the recurrent output tensor -- already written to GPU memory -- as a free content signal for the erase gate, and replaces the per-value write-gate projection with a single scalar per head. At initialisation CARVE is bit-identical to GDN-2; any quality difference emerges from what the content gate learns. At 1.3B parameters trained on 100B tokens, CARVE achieves WikiText perplexity 15.72 (minus 0.18 vs. GDN-2, a 4.5-sigma effect), leads every recurrent baseline on nine common-sense reasoning benchmarks, and sets state of the art on every RULER retrieval probe -- at 0.4% throughput overhead, 13% lower peak memory, and 19% fewer parameters. Six formal theorems cover memory capacity, Lyapunov stability, gradient flow, expressivity separation, Pareto-optimal chunk size, and hybrid optimality.

13:00 JST研究/論文

Flexformer: 学習可能なアテンション カーネルを備えた柔軟なリニア トランスフォーマー

Transformer モデルは、アテンション メカニズムに依存して長距離の依存関係を捕捉しますが、二次的な複雑さの影響を受け、スケーラビリティが長いシーケンスに制限されます。カーネルベースの線形アテンションはこの複雑さを軽減しますが、通常は固定カーネルまたは学習能力の低いカーネルに依存するため、表現力とパフォーマンスが制限されます。この研究では、完全にデータ駆動型の方法でアテンション カーネルを学習する柔軟な線形トランスフォーマーである Flexformer を提案します。 Flexformer は、ランダムなフーリエ特徴ベースの線形アテンションに基づいて構築されており、スペクトル周波数をトレーニング可能なパラメーターとして扱うことで、モデルがアテンション カーネルの幅広いファミリーを学習できるようにします。私たちは定常バリアントと非定常バリアントの両方を開発していますが、後者は厳密に優れた表現力を提供します。言語モデリングとシーケンス分類に関する広範な実験により、Flexformer が常にベースラインを上回るパフォーマンスを示すことが実証されました。さらに、Flexformer は、事前トレーニングされた Transformer から効果的に抽出してソフトマックス アテンションを回復することができ、ドメイン間での強力なカーネル移行性を示し、長時間シーケンスのタスクで高い効率と競争力のあるパフォーマンスの両方を実現します。

原文 (English)

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel

Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences. Kernel-based linear attention reduces this complexity but typically relies on fixed or weakly learnable kernels, restricting expressiveness and performance. In this work, we propose Flexformer, a flexible linear Transformer that learns attention kernels in a fully data-driven manner. Flexformer builds on random Fourier feature-based linear attention and treats spectral frequencies as trainable parameters, enabling the model to learn a broad family of attention kernels. We develop both stationary and nonstationary variants, with the latter offering strictly greater expressiveness. Extensive experiments on language modeling and sequence classification demonstrate that Flexformer consistently outperforms baselines. Moreover, Flexformer can be effectively distilled from pretrained Transformers to recover softmax attention and exhibits strong kernel transferability across domains, achieving both high efficiency and competitive performance on long-sequence tasks.

13:00 JST研究/論文

ペプチドドリフト: 抗原条件付き個別ペプチド生成のための毒性反発ドリフト

ペプチドは、小分子の化学的調整可能性と高分子治療薬の標的特異性を組み合わせた、有望な治療法です。しかし、毒性を回避しながら抗原特異的結合ペプチドを設計することは、治療用ペプチドの発見にとって依然として大きな課題である。ここでは、単一の抗原条件ドリフトステップを通じてペプチド候補を生成する、毒性を意識した潜在的精製フレームワークであるペプチドドリフトを紹介します。ペプチド包埋空間では、ペプチドドリフトは、生成された潜在ペプチドを抗原一致結合ペプチドに引き寄せる一方、毒性関連領域からは反発することを学習します。結合を促進する物理化学的特徴は、ペプチド表現空間における毒性関連の特徴と重なることが多いため、これは困難です。これに対処するために、最初に結合指向の引力を学習し、次に毒性の反発を高めることによって、この競合する目的を安定させるためのウォームアップ戦略を導入します。

原文 (English)

Pepti-drift: Toxicity-Repulsive Drifting for Antigen-Conditioned Discrete Peptide Generation

Peptides are a promising therapeutic modality that combine the chemical tunability of small molecules with the target specificity of macromolecular therapeutics. However, designing antigen-specific binding peptides while avoiding toxicity remains a major challenge for therapeutic peptide discovery. Here, we present Pepti-drift, a toxicity-aware latent refinement framework that generates peptide candidates through a single antigen-conditioned drift step. In a peptide embedding space, Pepti-drift learns to attract generated peptide latents toward antigen-matched binding peptides while repelling them from toxicity-associated regions. This is challenging because binding-promoting physicochemical features often overlap with toxicity-associated features in peptide representation space. To address this, we introduce a warm-up strategy to stabilize this competing objective by first learning binding-oriented attraction and then increasing toxicity repulsion. Pepti-drift achieves highly efficient generation, running 16.2-fold faster than PepMLM and 1,092.0-fold faster than PepTune. Generated peptides show 100% validity, 98.1% uniqueness, the highest sequence diversity, and near-zero cross-antigen reuse. Further evaluation indicates consistently reduced toxicity and hemolysis risk across most peptide-length ranges while retaining target-related predictive binding signal. Pepti-drift thus provides a fast, scalable, and controllable framework for antigen-specific peptide design that directly encodes safe-and-active properties.

13:00 JST画像/動画生成

Reflect-R1: 長いビデオの理解における自己修正のための証拠に基づくリフレクション

長時間のビデオを理解するための現在のマルチモーダル反射メカニズムは、主に内部パラメータ内の閉ループ自己反射に依存しています。客観的な外部証拠が欠如しているため、モデルはしばしば盲目的な自信に囚われ、エラーを修正できないことがよくあります。さらに、強化学習を多段階リフレクション パイプラインに適用すると、深刻なポリシー結合が導入され、専用のトレーニング データが重大に不足することでさらに悪化します。これらの制限に対処するために、この研究では、長いビデオを理解するための初の証拠主導型自己修正フレームワークである Reflect-R1 を提案しています。このフレームワークは、直感、検証、調停からなる 3 段階のパイプラインを構築します。客観的な視覚的証拠を動的に取得して最初の直観を検証し、複数の時間的検索を自律的に実行して矛盾を解決することで、幻覚ループを完全に断ち切ります。ポリシー結合を克服するために、さまざまな推論段階にわたって利点関数を独立して計算する、SD-GRPO という名前の段階分離型強化学習アルゴリズムを設計します。同時に、トレーニング データのギャップを埋めるために 120,000 サンプルのデータセットを構築します。 VideoMME や LongVideoBench などのベンチマークに関する広範な実験により、Reflect-R1 が最先端のパフォーマンスを達成していることが実証されています。私たちの方法は真の修正率を大幅に向上させ、厳密に客観的な証拠に基づいた真の自己修正を可能にします。

原文 (English)

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evidence, models are frequently trapped in blind confidence and often fail to correct errors. Furthermore, applying reinforcement learning to multi-stage reflection pipelines introduces severe policy coupling, which is exacerbated by a critical scarcity of dedicated training data. To address these limitations, this work proposes Reflect-R1, the first Evidence-Driven self-correction framework for long video understanding. The framework constructs a three-stage pipeline consisting of intuition, verification, and arbitration. By dynamically retrieving objective visual evidence to verify initial intuitions and autonomously executing multiple temporal searches to resolve conflicts, it completely breaks the hallucination loop. To overcome policy coupling, we design a stage-decoupled reinforcement learning algorithm named SD-GRPO that independently computes advantage functions across different reasoning stages. Concurrently, we construct a dataset of 120K samples to bridge the training data gap. Extensive experiments on benchmarks such as VideoMME and LongVideoBench demonstrate that Reflect-R1 achieves state-of-the-art performance. Our method significantly improves the genuine rectification rate and enables authentic self-correction strictly grounded in objective evidence.

13:00 JST画像/動画生成

Home3D 1.0: インテリア デザイン向けの高忠実度画像から 3D アセット生成システム

Home3D 1.0 は、インテリア デザインと電子商取引アプリケーションを対象とした、単一の参照画像から高品質の 3D アセットを生成するモジュール式の画像から 3D 生成システムです。家具や装飾品の写真が与えられると、システムは物理ベース レンダリング (PBR) マテリアルを使用してメッシュを出力し、メッシュはマテリアル固有のコンポーネントに分解できます。パイプラインは、密接に結合された 4 つのモジュールで構成されます。 ジオメトリは、ジオメトリ VAE と粗いから細かいフローマッチング DiT を使用した潜在的な SDF モデリングを通じて防水メッシュを再構築します。テクスチャは、マルチビュー アルベド観測を予測し、それらをメッシュ上に再投影し、目に見えない表面領域を 3D テクスチャ フィールドで完成させます。マテリアルは MatWeaver を使用して、ビデオベースのセグメンテーションと UV 空間投票を通じてコン​​ポーネント マスクを取得し、階層的なマルチモーダル マッチングを通じて厳選されたマテリアル ライブラリから PBR マップを取得してベイクします。そして、Parts は、PartVAE および PartDiT を使用してマテリアル編集可能なセマンティック パーツ メッシュを生成し、マルチヘッド パーツ固有の SDF フィールドを 1 つのパスでデコードします。各モジュールは専用のメトリックを使用して個別に評価され、現在のシステム機能と、より広範な導入に向けた残りのギャップの両方が強調表示されます。

原文 (English)

Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design

We present Home3D 1.0, a modular image-to-3D generation system that produces high-quality 3D assets from a single reference image, targeting interior design and e-commerce applications. Given a photograph of a furniture or decor item, the system outputs a mesh with physically-based rendering (PBR) materials, and the mesh can be decomposed into material-specific components. The pipeline is organized into four tightly coupled modules: Geometry reconstructs a watertight mesh through latent SDF modelling with a geometry VAE and a coarse-to-fine flow-matching DiT; Texture predicts multiview albedo observations, reprojects them onto the mesh, and completes unseen surface regions with a 3D texture field; Material uses MatWeaver to obtain component masks through video-based segmentation and UV-space voting, then retrieves and bakes PBR maps from a curated material library through hierarchical multi-modal matching; and Parts generates material-editable semantic part meshes with a PartVAE and PartDiT, decoding multi-head part-specific SDF fields in one pass. Each module is evaluated independently with dedicated metrics, highlighting both the current system capability and the remaining gaps toward broader deployment.

13:00 JSTLLM/生成AI

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not compr…