Skip to the content.

AIニュース 2026-06-26

自動生成: 2026-06-26 13:10 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. The White House is asking OpenAI to slow roll the release of its new model over safety concernsTechCrunch AI

    penAI reportedly plans to share its newest model, GPT 5.6, with a sel…

  2. Anthropic’s Claude is winning over paid consumers, a market owned by ChatGPTTechCrunch AI

    Despite ChatGPT's commanding market lead, consumers who pay for AI ha…

  3. 企業のAI支出そろそろ“様子見”は終わり? Gartner予測、本格投資の行方はITmedia AI+

    Gartnerは2026年の世界AI支出が前年比47%増の2兆5956億ドルに達するとの予測を発表。2026年は企業によるAI支出が拡大局…

  4. Amazon Bedrockのトークン処理量、26年1Qだけで過去累計超え AWSが目指す、AIのための「信頼できるインフラ」ITmedia AI+

    Amazon Bedrockのトークン処理量が2026年第1四半期だけで過去の累計を上回ったという。「AWS Summit Japan 2…

  5. 日本の“眠れるデータ”を競争力へ、日本発のプラットフォーム「xIPF」が始動ITmedia AI+

    欧州を中心に進むデータ共有圏の動向やその日本へのインパクトについて解説してきた本連載だが、第9回は日本独自のデータ連携エコシステムを創出す…

  6. ITインフラの要件に“大変化” なぜ「電力」と「冷却」にAI活用が縛られるのか?ITmedia AI+

    AI時代の到来により、ITインフラの要件に大きな変化が訪れています。データセンターを「どこに置くか」「どうやって冷やすか」が、AI活用の制…

  7. “AIの回答が薄い問題”をどう解決? 日本ハム、「AIが食べやすいデータ」作戦の全貌ITmedia AI+

    データ×AIによる意思決定を推進するには、非構造化データを「AIが食べやすいデータ」に変える必要がある。日本ハムが選んだアプローチとは。

トピック別件数

日本語メディア13件

ITmedia AI+ (日本語)

11:00 JSTその他

Amazon Bedrockのトークン処理量、26年1Qだけで過去累計超え AWSが目指す、AIのための「信頼できるインフラ」

Amazon Bedrockのトークン処理量が2026年第1四半期だけで過去の累計を上回ったという。「AWS Summit Japan 2026」の基調講演では、AI推論が最大の負荷となる時代の「信頼できるインフラ」の必要性が訴えられた。

11:00 JSTビジネス/資金調達

企業のAI支出そろそろ“様子見”は終わり? Gartner予測、本格投資の行方は

Gartnerは2026年の世界AI支出が前年比47%増の2兆5956億ドルに達するとの予測を発表。2026年は企業によるAI支出が拡大局面へ移行する転換点になるとしている。

08:00 JSTその他

日本の“眠れるデータ”を競争力へ、日本発のプラットフォーム「xIPF」が始動

欧州を中心に進むデータ共有圏の動向やその日本へのインパクトについて解説してきた本連載だが、第9回は日本独自のデータ連携エコシステムを創出することを目指す「xIPFコンソーシアム」を取り上げる。

08:00 JSTその他

ITインフラの要件に“大変化” なぜ「電力」と「冷却」にAI活用が縛られるのか?

AI時代の到来により、ITインフラの要件に大きな変化が訪れています。データセンターを「どこに置くか」「どうやって冷やすか」が、AI活用の制約になるのはなぜか。新しいITインフラの要件に、IT部門がどのように対応すべきでしょうか。

07:00 JSTその他

“AIの回答が薄い問題”をどう解決? 日本ハム、「AIが食べやすいデータ」作戦の全貌

データ×AIによる意思決定を推進するには、非構造化データを「AIが食べやすいデータ」に変える必要がある。日本ハムが選んだアプローチとは。

07:00 JSTエージェント

“流れていく会話”をAIが構造化 チャット履歴を会社の資産に

業務上の会話には、後から参照すべき貴重な情報が多く含まれる。しかし、従来型のチャットでは、こうした情報が整理されないまま流れ、必要な場面で見つけにくい。そうした課題を解決するために、AIチャットエージェントが役立つかもしれない。

06:15 JSTその他

リコーが多能工ヒューマノイドを披露、工場ではPoCから導入に向けた実証段階へ

リコーは、「AWS Summit Japan 2026」において、フィジカルAI搭載の多能工ヒューマノイドのデモンストレーションを披露した。既に工場内でPoCを始めており、今夏までをめどに多能工ヒューマノイドが一部の工程を担うより実用的な実証を始めたい考えだ。

21:05 JSTその他

Flashの再来? Figmaの新機能「Figma Motion」に懐かしいとの声 アニメーション生成するAI機能も

Figmaが発表した新機能「Figma Motion」が、かつてのAdobe Flashを思わせるとSNSで話題だ。タイムラインでキーフレームを打つ操作感が懐かしさを呼んだが、実際に触ると別物との声もみられる。

18:11 JSTLLM/生成AIGPT / ChatGPT

男性に美人局容疑で3人逮捕 ChatGPTの示談相場示し脅迫か 警視庁

美人局の手口で、少女とホテルに入った男性から現金を脅し取ったとして、警視庁少年事件課は恐喝の疑いで、東京都練馬区上石神井の職業不詳、斎藤蓮容疑者(22)ら男3人を逮捕、男女2人を書類送検した。斎藤容疑者は黙秘し、4人は容疑を認めている。

17:00 JSTロボティクス研究/論文

中国が人型ロボット開発で急成長しているワケ 日本が学ぶべきポイントは? 専門家が解説

なぜ人型ロボットの開発で中国が急成長しているのか。日本が学ぶべきポイントを野村総合研究所の李智慧氏が解説した。

16:22 JSTLLM/生成AI

「教員を生成AIに置き換える考えはない」東京外大が声明 SNSで拡散した懸念にコメント

東京外大が、「2人体制で運営する小規模語科を1人体制に縮小し、AIに代替させようとしている」との懸念がSNSで広まったことについて、「教員を生成AIに置き換える考えはない」との声明を出した。

15:00 JSTその他

白血病など16疾患の診断をAIが支援、日立がAUC0.9以上の新技術

日立製作所と九州大学病院は、血液悪性腫瘍の診断に用いるフローサイトメトリー検査において、医師の鑑別診断を支援する機械学習型のAI技術を開発した。複数疾患の同時分類において、識別性能を示す指標AUCで0.9以上の性能を確認した。

14:56 JSTその他

シャープブランドのAIサーバ展開も検討 シャープと鴻海、5分野で戦略的協業へ

シャープはこれまで、鴻海が製造するAIサーバの国内販売方針を示していたが、新たに、製品をシャープブランドとして展開を検討する方針が示された。

海外メディア9件

TechCrunch AI (英語)

08:34 JSTLLM/生成AI規制/政策OpenAIGPT / ChatGPT

The White House is asking OpenAI to slow roll the release of its new model over safety concerns

penAI reportedly plans to share its newest model, GPT 5.6, with a select group of partners instead of to the broader public. The reason: th…

05:19 JSTエージェント研究/論文Meta

Patronus AI lands $50M to build ‘digital worlds’ that stress-test AI agents

Agent-testing startup Patronus AI, founded by former Meta AI researchers, is experiencing nearly insatiable demand, its investor says.

02:38 JSTLLM/生成AIAnthropicClaudeGPT / ChatGPT

Anthropic’s Claude is winning over paid consumers, a market owned by ChatGPT

Despite ChatGPT's commanding market lead, consumers who pay for AI have been increasingly choosing Anthropic's Claude, data shows.

01:55 JSTエージェントビジネス/資金調達

General Intuition’s $2.3B bet that video games can train AI agents for the real world

General Intuition has raised $320 million to scale AI trained on millions of hours of gameplay, betting action data can help AI develop som…

01:48 JSTその他

Databricks’ former AI chief thinks he can cut AI’s power bill by 1,000x

Un-0 is an image-generation system tool that shows for the first time how the company's technology can replicate conventional AI systems.

23:55 JSTビジネス/資金調達

Netris raises $15M Series A from a16z to help AI neoclouds go live faster

Netris provides software that runs on network switches, and offers a platform that helps neocloud operators reduce the time it takes to go…

23:00 JSTその他

2 days left to save up to $190: Join 1,000+ founders and investors at TechCrunch Founder Summit

Two days left to lock in your spot at TechCrunch Founder Summit 2026 and save up to $190 before Early Bird rates expire on June 26 at 11:59…

22:30 JSTその他

Adobe acquires image and video enhancement tool maker Topaz Labs

Adobe said that it will integrate Topaz Labs' tools across its apps.

21:00 JSTビジネス/資金調達

Amazon ups India bet with fresh $13B AI infrastructure investment

Amazon’s latest India investment comes as global tech companies race to expand AI infrastructure in the country.

公式ブログ0件

このカテゴリの新着記事はありませんでした。

論文277件

arXiv cs.AI (英語)

13:00 JST研究/論文

カスケード線形特徴によるお調子者の検出と制御

アクティベーションステアリング手法を通じてモデルの動作を解釈および制御するには、望ましい動作または望ましくない動作を明確に示す対照的なサンプルの多くのペアが必要です。これらのデータ ペアによって、解釈可能性フレームワークが動作の原因となるモデルの特徴をどの程度確実に検出できるかが決まり、したがって、モデルをそのような動作に近づけるか遠ざけることができるかが決まります。この研究では、動作の原因となるカスケード線形特徴を分離する反復データ生成パイプラインを紹介します。具体的には、サンプルの単純なバイナリ ペアを超えて、代わりに動作に線形にスケールする特徴の度合いを示すサンプルを分離することで、特徴のもつれをより良く解くことができることを示します。私たちは、ユーザーの検証を優先する言語モデルの傾向であるお調子者を検出し、回避することに重点を置いています。我々は、カスケードサンプルを通じて発見されたお調子者の特徴が線形分離可能な部分空間を形成し、ベースラインのアプローチよりも目的の動作により明確に対応するモデル活性化の選択を可能にすることを実証します。また、検出、決定論的なスコアリング、堅牢なステアリングを可能にする機能も評価し、LLM-as-a-judge およびシステム プロンプト ベースラインと同等またはそれを上回るパフォーマンスを示しながら、計算量の削減と解釈可能性の保証の向上を実現していることを確認しました。コードとデータ: https://cascading-feats.github.io/

原文 (English)

Detecting and Controlling Sycophancy with Cascading Linear Features

Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. These data pairs determine the degree to which interpretability frameworks can reliably detect model features responsible for a behavior, and therefore the ability to steer models toward or away from such behavior. In this work, we present an iterative data generation pipeline that isolates cascading linear features responsible for a behavior. Specifically, we show how moving beyond simple binary pairs of samples, and instead isolating samples that show degrees of features that scale linearly with behavior, allows for better disentanglement of features. We focus on detecting and steering away from sycophancy -- the tendency of language models to prioritize user validation. We demonstrate that sycophancy features discovered through cascading samples form linearly separable subspaces, and allow for selection of model activations that more clearly correspond to the desired behavior than baseline approaches. We also evaluate their ability to enable detection, deterministic scoring, and robust steering, and see that they either match or outperform LLM-as-a-judge and system prompting baselines while providing lower computational demand and more interpretability guarantees. Code & Data: https://cascading-feats.github.io/

13:00 JST研究/論文

ベンチマーク飽和後の生活: CORE-Bench のケーススタディ

ベンチマークの精度が飽和すると、多くの場合、そのベンチマークは廃止され、より困難なバージョンに置き換えられます。我々は、このアプローチが精度を優先し、エージェントのパフォーマンスの他の 6 つの主要な側面を研究する機会を逃していることを示します。つまり、ショートカット、分布外の一般化可能性、効率、信頼性、モデルと足場の相対的な重要性、人間とエージェントのコラボレーションによる向上などの構築妥当性の問題です。私たちは、科学コードの計算再現性のベンチマークである CORE-Bench Hard をケーススタディとして使用し、これらの次元に沿ってエージェントを測定すると、精度が飽和した後でもエージェントのパフォーマンスについて有意義な洞察が得られることを実証します。まず、CORE-Bench Hard で妥当性を構築するために、能力の低いエージェントでは予測することが難しい脅威を表面化します。改良されたベンチマークである CORE-Bench v1.1 と、配布外のタスク スイートである CORE-Bench OOD を導入します。次に、精度が飽和しているにもかかわらず、CORE-Bench v1.1 は効率、信頼性、モデルのパフォーマンス、および足場のパフォーマンスを測定するのに依然として有用であることがわかりました。最後に、実世界の計算再現性タスクにおける人間とエージェントのコラボレーションによる向上を測定するために、小規模なランダム化実験を実施します。私たちは、約 2 倍の統計的に有意な速度向上を発見しました (人間のみによる複製の 5 分の 1 が完了する前に制限時間に達しているため過小評価されている可能性があります) と、その他のさまざまな発見について説明します。私たちの貢献は、支配的な精度中心の評価パラダイムに対するより厳密な代替案を提示します。

原文 (English)

Life After Benchmark Saturation: A Case Study of CORE-Bench

When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity to study six other key dimensions of agent performance: construct validity issues such as shortcuts, out-of-distribution generalizability, efficiency, reliability, the relative importance of the model versus the scaffold, and uplift from human-agent collaboration. We use CORE-Bench Hard, a benchmark for computational reproducibility of scientific code, as a case study to demonstrate that measuring agents along these dimensions yields meaningful insights into agent performance even after accuracy saturates. First, we surface threats to construct validity in CORE-Bench Hard that are difficult to anticipate with less capable agents. We introduce an improved benchmark, CORE-Bench v1.1, and an out-of-distribution task suite, CORE-Bench OOD. Second, we find that despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring efficiency, reliability, model performance, and scaffold performance. Finally, we conduct a small-scale randomized experiment to measure uplift from human-agent collaboration on real-world computational reproducibility tasks. We find a statistically significant speedup by about a factor of two -- likely underestimated due to one-fifth of human-only reproductions reaching the time limit before completing -- and describe various other findings. Together, our contributions present a more rigorous alternative to the dominant accuracy-centric evaluation paradigm.

13:00 JST研究/論文Llama

拒否はチャット モデルのペルソナの下流に存在します

活性化空間における線形方向は、指示調整型チャット モデルにおける拒否特性とペルソナ特性の両方について特定されていますが、この 2 つは別個のメカニズムとして研究されてきました。私たちは、彼らが相互作用することを示します。つまり、従順なペルソナがゲートを拒否するということです。 Qwen2.5-7B-Instruct と Llama-3.1-8B-Instruct では、準拠モデルペルソナの方向と拒否の方向を抽出し、両方に介入します。準拠したペルソナのステアリングにより拒否が抑制されます。ラマでは、拒否率が 97% から 2% に低下しました。拒否の方向を再導入すると、後の層での拒否が部分的に回復しますが、初期の層では回復しません。後期レイヤーウィンドウでペルソナの方向を投影すると、それがベースラインに戻ります。ランダムな方向に投影することはできません。したがって、拒否は、それが計算される場所の下流にある、後期層の表現段階でゲートされます。拒否を単一の孤立した方向として扱うと、人格への依存性が失われます。

原文 (English)

Refusal Lives Downstream of Persona in Chat Models

Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate mechanisms. We show they interact: a compliant persona gates refusal. In Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, we extract a compliant model-persona direction and a refusal direction and intervene on both. Compliant persona steering suppresses refusal -- in Llama, the refusal rate falls from 97% to 2%. Reintroducing the refusal direction partially restores refusal at late layers but not at early ones. Projecting out the persona direction in a late-layer window restores it to baseline; projecting out a random direction does not. Refusal is therefore gated at the late-layer expression stage, downstream of where it is computed. Treating refusal as a single isolated direction misses its dependence on persona.

13:00 JSTLLM/生成AI

AlgoEvolve: LLM 主導のアルゴリズム取引プログラムのメタ進化

最近の研究では、大規模言語モデル (LLM) がプログラムと証明の進化的発見のための意味論的突然変異演算子として機能できることが示されています。現在のアプリケーションのほとんどは静的コーディングのベンチマークに重点を置いています。私たちはこのパラダイムをアルゴリズム取引に拡張します。このドメインは、ノイズが多く、非定常で、非常に不連続であるため、独特の課題を抱えています。私たちは、実行可能な取引戦略を生成、評価、反復的に改善する LLM 主導の進化的フレームワークである AlgoEvolve を紹介します。これらの戦略は Python コードとして表現され、厳格なテスト プロトコルを通じて評価されます。複数の実験を通じて、このシステムは、取引ルールの自律的な変更を含む、新たな体制適応戦略ロジックを示しました。さらに、内部ループでプログラム合成を導くプロンプトを進化させるメタ進化的な外部ループを導入します。この外側のループは、改善された検索ヒューリスティックを発見します。これらのヒューリスティックは、ゼロトレードの失敗を減らしながら、探査と活用のバランスをとります。これらは、人間が設計した最初の指示よりも常に優れたパフォーマンスを発揮します。この結果は、LLM ベースのセマンティック進化が、複雑な環境における継続的なプログラム合成に実行可能なアプローチを提供することを示しています。

原文 (English)

AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs

Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and proofs. Most current applications focus on static coding benchmarks. We extend this paradigm to algorithmic trading. This domain is uniquely challenging because it is noisy, non-stationary, and highly discontinuous. We present AlgoEvolve, an LLM-driven evolutionary framework that generates, evaluates, and iteratively improves executable trading strategies. These strategies are expressed as Python code and evaluated through a rigorous testing protocol. Across multiple experiments, the system exhibits emergent regime-adaptive strategy logic, including autonomous shifts in trading rules. We further introduce a meta-evolutionary outer loop that evolves the prompts guiding program synthesis in the inner loop. This outer loop discovers improved search heuristics. These heuristics balance exploration and exploitation while reducing zero-trade failures. They consistently outperform initial human-designed instructions. The results demonstrate that LLM-based semantic evolution provides a viable approach for continual program synthesis in complex environments.

13:00 JSTLLM/生成AIエージェントGoogle

エージェントティック インフラストラクチャのためのエージェント分析: DAO と企業 AI プロトコルの比較ガバナンスのための LLM を利用したパイプライン

AI エージェント プロトコルが急増する一方で、その相互運用性標準を形成するガバナンス構造は経験的に十分に検討されていないままです。私たちは、大規模なガバナンス談話分析のための LLM を利用した比較パイプラインを導入し、自動アノテーション、ニューラル トピック モデリング、および多層ネットワーク分析を統合して、社会技術的な権力構造を大規模に研究します。当社では、エージェントの相互運用性に関する 2 つの対照的な標準、ERC-8004 (パーミッションレス、オンチェーン) と Google A2A (企業主導) に基づいて検証しています。 4,323 件のガバナンス参加記録を分析し、LLM 支援コーディング、トピック モデリング、および多層ネットワーク分析を組み合わせて、制度設計がテーマの優先順位とコミュニティ構造をどのように形成するかを調査します。私たちは、ガバナンスの形態が実質的な焦点に影響を与える一方で、どちらの体制も同等のレベルの参加不平等とコミュニティの断片化を示していることを発見しました。パーミッションレス環境では言説の整合性がより密になっており、分散型の参加にもかかわらず、オープンなガバナンスがより大きなテーマの収束を促進する可能性があることを示唆しています。これらの発見は、LLM 支援手法がテクノロジー ガバナンスの実証研究をどのように前進させることができるかを示しており、より公平なエージェント AI 標準の設計に影響を及ぼします。すべてのデータとコードはオープンに利用できます。

原文 (English)

Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures at scale. We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led). Analyzing 4,323 governance participation records, we combine LLM-assisted coding, topic modeling, and multi-layer network analysis to examine how institutional design shapes thematic priorities and community structure. We find that while governance form influences substantive focus, both regimes exhibit comparable levels of participation inequality and community fragmentation. Discourse alignment is denser in the permissionless setting, suggesting that open governance may foster greater thematic convergence despite decentralized participation. These findings illustrate how LLM-assisted methods can advance the empirical study of technology governance, with implications for designing more equitable agentic AI standards. All data and code are openly available.

13:00 JSTエージェント

メンタルヘルスの薬剤情報探索のための知識拡張型エージェント AI

患者はますますオンラインで医薬品情報を求めるようになっているが、精神科薬の安全性に関する知識は、権威あるが抽象的な規制上の有害事象記録と、経験に近いが検証されていない患者の語りとに分かれている。証拠と逸話を混同することなくそれらを統合することは、文脈が不十分な情報によって恐怖、ノーシーボ反応、不遵守が増幅される可能性がある精神医学において特に重要です。ここでは、9 つ​​の抗うつ薬に関する 466,525 件の Reddit 投稿、60,782 件の WebMD レビュー、および 20 年間にわたる米国 FDA 有害事象報告システムの記録を統合した、出所を意識したナレッジグラフベースのマルチエージェント フレームワークを開発しました。医師の注釈に対してベンチマークされた大規模言語モデルの実体認識パイプラインは、薬物の場合は 0.969、症状の場合は 0.973 という最高の F1 スコアに達しました。 2 つのコミュニティ プラットフォームは、規制報告書よりも相互にはるかに一致しており (Jaccard 類似度 0.905 まで重複)、患者生成データが部分的に独立した安全性シグナルを形成していることを示しています。セルトラリンについては、対応する FDA 日付の数百日前に多くの有害事象がコミュニティ情報源に現れました。 ATC-N、ICD-10、および MedDRA の語彙に基づいた Neo4j ナレッジ グラフは出所を保存し、すべての請求を追跡可能に保ち、規制上の事実を患者の経験から区別します。これらの結果は、より監査可能な精神科治療情報へのルートとしてソースを意識した統合を確立し、有用性と患者利益を前向きにテストする必要があります。

原文 (English)

Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records, which are authoritative but abstract, and patient narratives, which are experience-near but unvalidated. Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly contextualised information can amplify fear, nocebo responses, and non-adherence. Here we develop a provenance-aware, knowledge-graph-based multi-agent framework unifying 466,525 Reddit posts, 60,782 WebMD reviews, and twenty years of U.S. FDA Adverse Event Reporting System records for nine antidepressants. A large-language-model entity-recognition pipeline benchmarked against physician annotations reached highest F1 scores of 0.969 for medications and 0.973 for conditions. The two community platforms were far more concordant with each other (overlap up to a Jaccard similarity of 0.905) than with regulatory reports, indicating that patient-generated data form a partly independent safety signal. For sertraline, many adverse events appeared in community sources hundreds of days before the corresponding FDA date. A Neo4j knowledge graph grounded in ATC-N, ICD-10, and MedDRA vocabularies preserves provenance, keeping every claim traceable and regulatory facts distinct from patient experience. These results establish source-aware integration as a route to more auditable psychiatric medication information, with usefulness and patient benefit to be tested prospectively.

13:00 JST研究/論文

チェスのスキル評価の加速: ドリフト拡散を強化した Elo 評価システム

Elo などのレーティング システムは、競技チェスのマッチメイキングのゴールド スタンダードとして機能します。ただし、試合の結果のみに依存し、ゲームプレイの細かな品質を無視しているため、本質的に応答の遅れに悩まされています。それにもかかわらず、レーティング調整に手ごとの情報を組み込むことは、かなりのノイズとゲーム状態空間の広大さを考慮すると、大きな課題となります。これに対処するために、認知神経科学のドリフト拡散モデル (DDM) にヒントを得た新しいスキル評価フレームワークであるドリフト拡散強化 Elo 評価システム (DD-Elo) を提案します。スキル表現を意思決定プロセスとしてモデル化することで、私たちのモデルは技レベルのデータを統合して、スキルの急速な変動を捉えます。私たちは、DD-Elo が従来の Elo システムからの一定の偏差を維持し、理論的な整合性を確保していることを証明する厳密な数学的導出を提供します。広範な実験により、DD-Elo は Elo よりも早くスキルの変化に適応することが実証されました。私たちの調査結果は、DD-Elo がチェス レーティング エコシステムに対して、説明可能で応答性が高く、下位互換性のあるソリューションを提供することを示唆しています。実装コードは https://github.com/Aquila-zhou1/DD-Elo で公開されています。

原文 (English)

Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System

Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess. However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the granular quality of gameplay. Nevertheless, incorporating move-by-move information into rating adjustments presents a significant challenge given the substantial noise and the vastness of the game-state space. To address this, we propose the Drift-Diffusion-Enhanced Elo Rating System (DD-Elo), a novel skill assessment framework inspired by the drift diffusion model (DDM) from cognitive neuroscience. By modeling skill expression as a decision-making process, our model integrates move-level data to capture rapid skill fluctuations. We provide a rigorous mathematical derivation proving that DD-Elo maintains a bounded deviation from the traditional Elo system, ensuring theoretical alignment. Extensive experiments demonstrate that DD-Elo adapts to skill changes faster than Elo. Our findings suggest that DD-Elo offers an explainable, highly responsive, and backward-compatible solution for chess rating ecosystems. The implementation code is publicly available at https://github.com/Aquila-zhou1/DD-Elo .

13:00 JSTエージェント

エージェントではなく、行動を統治する: 自律型 AI システムのガバナンス モデルとしての機関の認証

自律型 AI エージェントは、臨床処方や実稼働ソフトウェアの導入など、結果として取り消せないアクションを実行し始める可能性があります。この論文は、人間の制度が強力な自律的主体を、彼らの推論を監視することによってではなく、結果的な行動の時点で独立して証明された証拠を要求することによって統治してきたことを観察しています。私たちは、この制度的パターンを AI エージェント システムの計算ガバナンス モデルとして形式化します。提案されたモデルでは、エージェントは計画と推論に関して完全な自律性を保持しますが、指定された高リスクのアクションについては実行権限を持ちません。実行は、個別の信頼できるソースによってそれぞれ独立して証明され、宣言された意図に暗号的に結び付けられ、決定論的なポリシーによって評価される前提条件に基づいて行われます。決定は、独立した再検証に適した改ざん防止ログに記録されます。概念実証の実装を示し、ソフトウェアの導入と臨床処方の例を使用してモデルを説明します。

原文 (English)

Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems

Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment. This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning but by requiring independently attested evidence at the point of consequential action. We formalise this institutional pattern as a computational governance model for AI agent systems. Under the proposed model, an agent retains full autonomy over planning and reasoning but holds no execution authority over designated high-risk actions. Execution is conditional on preconditions that are each independently attested by a separate authoritative source, cryptographically bound to a declared intent, and evaluated by a deterministic policy. Decisions are recorded in a tamper-evident log amenable to independent re-verification. We present a proof-of-concept implementation and illustrate the model with examples from software deployment and clinical prescribing.

13:00 JST研究/論文

COrigami: 平坦折り可能な視覚的に認識可能な折り紙を共同設計するための AI パイプライン

生成 AI は検証可能なソリューションによる問題解決で目覚ましい成功を収めてきましたが、厳密な幾何学的制約と主観的な視覚美の両方を満たす物理アートを生成することは依然として課題です。本稿では、平坦折り可能性の方程式内で芸術的デザインを基礎づける数学的に厳密な環境である計算折り紙の領域におけるこれらの困難に取り組むアプローチを紹介します。 COrigami は、自然言語から折り目パターンを生成することで設計サイクルを支援する、エンドツーエンドの AI 駆動パイプラインです。私たちのパイプラインには、セマンティック スティック フィギュアの生成、基本パッキングの計算、平坦折り可能な折り目パターンの解決、平坦折り折り目パターンの整形、および自律的な美的評価ループによって駆動される強化学習を使用した生成されたモデルの改良が含まれます。私たちのシステムは非常に効果的な共同アシスタントとして機能し、人間のアーティストがさらに拡張して形を整えることができる構造的な出発点を生成します。この研究では、アルゴリズムの最適化と自律的な美的批評を統合することにより、AI システムが多目的の物理的制約をどのように満たして信頼性の高い、数学的に根拠のある共同創造性を実現できるかを示しています。

原文 (English)

COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict geometric constraints and subjective visual aesthetics remains a challenge. This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic design within the equations of flat foldability. We present COrigami, an end-to-end AI-driven pipeline that assists the design cycle by generating crease patterns from natural language. Our pipeline involves generating a semantic stick figure, computing a base packing, solving for a flat-foldable crease pattern, shaping the flat-folded crease pattern, and refining the generated model using reinforcement learning driven by an autonomous aesthetic evaluation loop. Our system acts as a highly effective collaborative assistant, generating structural starting points that human artists can further expand and shape. By integrating algorithmic optimisation with autonomous aesthetic critique, this work demonstrates how AI systems can satisfy multi-objective physical constraints to enable reliable, mathematically grounded co-creativity.

13:00 JSTLLM/生成AIエージェント

検証の地平線: コーディング エージェントの報酬に特効薬はない

古典的な直観では、解決策を生み出すよりも検証する方が簡単だと考えられています。今日のコーディング エージェントにとって、この直感は逆転しつつあります。基礎モデルがより強力な推論機能を開発し、エンジニアリング ハーネスがより洗練されるにつれて、複雑な候補ソリューションを生成することはもはや難しくなくなり、それらを確実に検証することがより困難な問題になりました。私たちが構築できるすべての検証ツールは人間の意図の代理にすぎず、意図そのものではありません。このため、検証は 2 つの困難を伴います。1 つは、本質的に意図が過少指定されているため、意図が満たされているかどうかを忠実に確認することが本質的に困難であるということです。次に、モデルのトレーニング中に、最適化によってプロキシとインテントの間のギャップが広がり、報酬のハッキングや信号の飽和として現れます。これに対処するために、私たちは検証信号の品質を 3 つの次元 (スケーラビリティ、忠実性、堅牢性) に沿って特徴付け、3 つすべてを同時に達成することが中心的な課題であると主張します。さらに、一般的なコーディング タスク用のテスト検証器、フロントエンド タスク用のルーブリック検証器、現実世界のエージェント タスク用の検証器としてのユーザー、長期タスク用の自動エージェント検証器の 4 つの報酬構造を研究します。さまざまなタスクの種類とポリシーの機能レベルにわたって、報酬設計の中核となる課題と、報酬シグナルをより効果的に活用する方法について、綿密な分析と実験を実施します。実験では、ターゲットを絞った検証設計により、報酬ハッキングを効果的に抑制し、タスク完了の品質を向上させ、複数の内部および公開ベンチマーク全体で大幅な利益を達成できることが示されています。これらの経験は総合的に、政策能力が成長し続けるにつれて、固定報酬関数が有効であり続けることはできないという核心的な観察を示しています。そして検証はジェネレーターと共進化する必要があります。

原文 (English)

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabilities and engineering harnesses grow more sophisticated, generating complex candidate solutions is no longer difficult -- reliably verifying them has become the harder problem. Every verifier we can build is only a proxy for human intent, never the intent itself. This makes verification subject to a twofold difficulty: first, intent is underspecified by nature, making it inherently hard to faithfully check whether it has been fulfilled; second, during model training, optimization widens the gap between proxy and intent -- manifesting as reward hacking or signal saturation. To address this, we characterize the quality of verification signals along three dimensions -- scalability, faithfulness, and robustness -- and argue that achieving all three simultaneously is the central challenge. We further study four reward constructions: a test verifier for general coding tasks, a rubric verifier for frontend tasks, the user as verifier for real-world agent tasks, and an automated agent verifier for long-horizon tasks. Across different task types and policy capability levels, we conduct in-depth analysis and experiments on the core challenges of reward design and how to more effectively leverage reward signals. Experiments show that targeted verification design can effectively suppress reward hacking, improve task completion quality, and achieve significant gains across multiple internal and public benchmarks. These experiences collectively point to a core observation: no fixed reward function can remain effective as policy capability continues to grow; and verification must co-evolve with the generator.

13:00 JSTLLM/生成AIエージェント研究/論文

ツール拡張 LLM エージェントは現実世界のエネルギー分析タスクをどのように実行しますか?

エージェントベンチマークは、金融、コーディング、法律、創薬などの汎用および分野固有の設定にわたって登場していますが、エネルギー分野の評価は依然として静的な知識の想起に主に限定されています。これは、ライブデータの取得、専門的な規制と市場の知識、現実世界の制約の下での複数段階の定量的推論を必要とするセクターにとって、重大なギャップです。現実世界のエネルギー市場分析タスクにおけるツール拡張 LLM エージェントの実証研究を紹介します。当社の評価環境には、(1) 市場データの取得と分析、(2) 知識の取得と解釈、(3) 高度な定量モデリングと意思決定分析の 3 つのカテゴリにわたる、専門家が厳選した 243 の問題が含まれています。タスクには、価格と需要の分析、料金の影響モデリング、資産収益と収益の推定、ヘッジ戦略分析、最適化モデリングが含まれ、問題は複数の難易度にまたがります。エージェントには、米国の主要 ISO 用のライブ電力市場 API、規制書類検索、公共料金データベース、資産最適化モデル、エネルギー市場ドキュメントの検索拡張発電などを含む、構成可能なドメイン ツール スイートが装備されています。当社は、アプローチの正しさ、回答の正確さ、属性の整合性、およびソースの妥当性をスコアリングする多次元評価プロトコルを使用してエージェントの応答を評価します。スコア基準を質問の種類に一致させるためのカテゴリを認識したルーティングを使用します。私たちはクローズドソースとオープンソースの LLM の両方を評価し、一か八かの専門分野でモデルの機能とドメイン ツールがどのように相互作用するかを比較分析します。主要なアーティファクトは、再現性と将来の研究をサポートするために公開されています。

原文 (English)

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-domain evaluations remain largely limited to static knowledge recall. This is a critical gap for a sector that requires live data retrieval, specialized regulatory and market knowledge, and multi-step quantitative reasoning under real-world constraints. We present an empirical study of tool-augmented LLM agents on real-world energy market analytics tasks. Our evaluation environment includes 243 expert-curated problems across three categories: (1) Market Data Retrieval and Analysis, (2) Knowledge Retrieval and Interpretation, and (3) Advanced Quantitative Modeling and Decision Analytics. Tasks include price and demand analysis, tariff impact modeling, asset revenue and returns estimation, hedging strategy analysis, and optimization modeling, with problems spanning multiple difficulty levels. Agents are equipped with a configurable suite of domain tools, including live electricity market APIs for major U.S. ISOs, regulatory docket search, utility tariff databases, asset optimization models, and retrieval-augmented generation over energy market documents. We assess agent responses using a multi-dimensional evaluation protocol that scores approach correctness, answer accuracy, attribute alignment, and source validity, with category-aware routing to match scoring criteria to question type. We evaluate both closed-source and open-source LLMs, providing a comparative analysis of how model capability and domain tooling interact in a high-stakes professional domain. Key artifacts are publicly released to support reproducibility and future research.

13:00 JSTLLM/生成AIビジネス/資金調達

マルチモーダル LLM 評価に欠けているものは何ですか?

マルチモーダル大規模言語モデル (MLLM) は、テキスト、画像、音声、ビデオなどのさまざまな入力を処理し、テキスト応答を生成できます。それらの機能は急速に進歩していますが、そのようなモデルの評価は追いついていません。既存の評価ベンチマークのほとんどは、個別のタスクに限定されており、モデルがモダリティ全体で情報を統合しているかどうかについてはほとんど明らかにされていません。私たちは、MLLM を評価するための現在の手段を調査し、既存のベンチマーク分類をレビューして、時間空間的一貫性、物理世界の理解、マルチモーダル一貫性、選択的注意などのギャップを特定します。これらのギャップに対処することは、マルチモーダル インテリジェンスの実際の進歩を測定し、機能の境界を明らかにするために不可欠です。

原文 (English)

What We are Missing in Multimodal LLM Evaluation?

Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced rapidly, evaluation of such models has not kept pace. Most existing evaluation benchmarks are limited to isolated tasks and reveal little about whether a model integrates information across modalities. We examine current means for evaluating MLLMs and review the existing benchmark taxonomy to identify gaps, including temporal-spatial coherence, physical world understanding, multimodal consistency, and selective attention. Addressing these gaps is essential for measuring real progress in multimodal intelligence and exposing capability boundaries.

13:00 JSTエージェントビジネス/資金調達

OpenFinGym: クオンツエージェントを評価するための検証可能なマルチタスクジム環境

大規模な言語モデル エージェントは定量的財務ワークフローにますます適用されていますが、その評価は分離されたタスク間で断片化されたままであり、ベンチマーク タスクの財務関連性はしばしば見落とされます。しかし、財務ワークフローは本質的に多段階であり、予測、戦略構築、リスク管理、取引などの相互依存するタスクにまたがっています。既存のプラットフォームは通常、単一のタスクに焦点を当てているため、エージェントの能力を過大評価し、一般化、実際の市場でのやり取り、および財務的に意味のある意思決定における弱点を明らかにできません。 OpenFinGym は、単一の実行および検証インターフェイスで予測、市場生成、リ​​アルタイム取引、不正検出をカバーする定量的金融エージェント開発用の統合ジム環境です。 OpenFinGym はさらに、定量的な財務出版物を実行可能なタスク パッケージに変換する自動タスク構築パイプラインを提供します。スケーラブルなエージェントのロールアウトをサポートし、ランタイムのトレインテストの漏洩を防ぐホスト側検証サービスを備えたコンテナ化されたランタイム。低レイテンシのデータストリーム設計を備えたペーパートレーディングエンジン。長期およびイベント市場の予測に対する遅延解像度のサポート。トレーニング後の SFT と RL の統合

原文 (English)

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark tasks is often overlooked. Yet financial workflows are inherently multi-stage, spanning interdependent tasks such as forecasting, strategy construction, risk management, and trading. Existing platforms typically focus on a single task, and can therefore overstate agent competence and fail to reveal weaknesses in generalization, real-market interaction, and financially meaningful decision-making. We introduce OpenFinGym, a unified gym environment for quantitative-finance agent development that covers forecasting, market generation, real-time trading, and fraud detection under a single execution and verification interface. OpenFinGym additionally provides an automated task-construction pipeline that turns quantitative finance publications into executable task packages; a containerised runtime with a host-side verifier service that supports scalable agent rollouts and prevents runtime train-test leakage; a paper trading engine with a low-latency data-stream design; deferred-resolution support for long-horizon and event-market forecasts; and integration for SFT and RL post-training

13:00 JSTLLM/生成AIエージェントClaude

命令ブリード: プロンプト構成エージェント システムにおけるモジュール間干渉

プロンプト構成エージェント システムの実践者は、繰り返し発生する障害モードを報告しています。つまり、共有変数や実行可能ファイルの依存関係がないにもかかわらず、1 つのプロンプト モジュールを編集すると、他のプロンプト モジュールの動作が静かに変更されます。私たちはこれを構成的動作漏洩 (CBL)、つまりコンテキスト ウィンドウを共有するモジュール間の干渉として形式化します。 CBL はアーキテクチャ上の非絶縁によって有効になります。トランスのセルフアテンションにより、連結されたモジュール間に正式な境界がありません。ボリューム、コンテンツ、形式に沿って非焦点モジュールを混乱させる再利用可能な 3 チャネル プロトコルを通じて、デプロイされたジョブ評価エージェント (Claude Sonnet 4.6、144 トライアル) で CBL を調査します。コンテンツ チャネルのみが検出可能な一対の効果を生成します (コーエンの d = 0.63、ゼロを除くブートストラップ 95% CI)。推奨事項が反転することはありません。標準の QA では目に見えない閾値以下の体制ですが、配置されたエージェントが下す何千もの意思決定を複雑にします。 CBL は、既知のエージェント障害軸 (敵対的注入、認知機能低下、マルチエージェント障害伝播、プライバシー漏洩) と直交しています。私たちは、運用定義、再利用可能なプロトコル、改ざん可能な予測セット、およびシステムクラスの特性評価に貢献し、迅速に作成されたエージェント評価の要件としてモジュール間干渉測定を確立します。

原文 (English)

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of others despite no shared variable or executable dependency. We formalize this as compositional behavioral leakage (CBL): interference between modules sharing a context window. CBL is enabled by architectural non-isolation: transformer self-attention provides no formal boundary between concatenated modules. We probe CBL on a deployed job-evaluation agent (Claude Sonnet 4.6, 144 trials) through a reusable three-channel protocol that perturbs non-focal modules along volume, content, and form. Only the content channel produces a detectable paired effect (Cohen's d = 0.63, bootstrap 95% CI excluding zero); no recommendation flipped -- a sub-threshold regime invisible to standard QA but compounding across the thousands of decisions a deployed agent makes. CBL is orthogonal to known agent-failure axes (adversarial injection, cognitive degradation, multi-agent fault propagation, privacy leakage). We contribute an operational definition, a reusable protocol, a falsifiable prediction set, and a system-class characterization, establishing cross-module interference measurement as a requirement for prompt-composed agent evaluation.

13:00 JST研究/論文

収益の加速と科学の定性エンジン

レイ・カーツワイルは、テクノロジーの進歩を議論する際に最も影響力のある物語である、収益の加速というテーゼについて説明しました。その中心的な主張は、複数の技術分野、特にコンピューティング、人工知能、脳科学、バイオテクノロジーの進歩が相互作用し、進歩が自己増幅的かつほぼ指数関数的になるというものです。この論文は、その主張の単純な数学的解釈を示し、そのような加速が現実であるとしても、それ自体では科学的発見の中心的な問題を解決するものではないと主張します。その理由は、収益の加速は実行能力とインフラストラクチャ能力に最も自然に適用されるのに対し、真の発見は別の能力、つまり、現在のフレームワークが構造的にいつ不適切であるか、次にどのような概念的な動きが必要であるかについての定性的推論に依存することが多いためです。最近の ARC-AGI-3 の結果は、この区別を明確にします。人間はベンチマークを天井で解くのに対し、フロンティア AI システムは 1% 未満に留まり、現在の AI と人間の柔軟な推論とのギャップが依然として非常に大きいことを示しています。同時に、デミス・ハサビス氏は、人間は意味の感覚を保持し、自分の人生の焦点を何に集中させるかを保持しなければならないと強調し、AIの将来は技術的な予測であるだけでなく、どのような形の人間の理解を保存し伝達する価値があるかという問題でもあることを思い出させます。この論文では、科学のための質的エンジン (QES) [3] を、不足している能力への対応策として位置づけています。この見解では、カーツワイル理論は、量的能力が加速する理由を説明するのに役立ちますが、QES は加速だけでは解決できない科学的発見の中心的な問題に対処します。その価値は、AGI がいつ到来するかによって決まるのではなく、科学的発見のプロセス自体が、保存し、整理し、アクセス可能にする価値のある人類の知恵の一形態を構成するという事実によって決まります。

原文 (English)

Accelerating Returns and the Qualitative Engine for Science

Ray Kurzweil described a thesis of accelerating returns, which is the most influential narratives in discussions of technological progress. Its central claim is that advances in multiple technological fields, especially compute, artificial intelligence, brain science, and biotechnology, interact in such a way that progress becomes self-amplifying and approximately exponential. This paper gives a simple mathematical interpretation of that claim and then argues that, even if such acceleration is real, it does not by itself resolve the central problem of scientific discovery. The reason is that accelerating returns apply most naturally to executional and infrastructural capability, whereas genuine discovery often depends on a different capacity: qualitative reasoning about when a current framework is structurally inadequate and what conceptual move is needed next. Recent ARC-AGI-3 results sharpen this distinction: humans solve the benchmark at ceiling, whereas frontier AI systems remain below 1%, indicating that the gap between current AI and human flexible reasoning is still very large. At the same time, Demis Hassabis has emphasized that humans must retain their sense of meaning and what they choose to focus their lives on, a reminder that the future of AI is not only a technical forecast but also a question of what forms of human understanding are worth preserving and transmitting. This paper positions the Qualitative Engine for Science (QES) [3] as a response to that missing capacity. In this view, the Kurzweil theory helps explain why quantitative capability may accelerate, while QES addresses the central problem in scientific discovery that acceleration alone does not solve. Its value does not depend on when AGI arrives, but on the fact that the processes of scientific discovery themselves constitute a form of human wisdom worth preserving, organizing, and making accessible.

13:00 JSTLLM/生成AI

思考のナレーション: 大規模言語モデルにおける実行不可能な倫理的推論のための推論時間足場

道徳的ジレンマに関する標準的な思考連鎖は、利害関係者の崩壊(結果に利害関係を持つ最大でも 1 つの当事者のみをトレース名で示す)と不確実性の抑圧(行動にコミットする前に明確な未知数やヘッジがない)という 2 つの失敗モードを示します。思考のナレーション (NoT) を導入します。これは、思考の連鎖を 5 つのセクション (主人公、利害関係者、2 段階の結果、不確実性、コミットメント) に構造化するシステム プロンプトです。 NoT では、トレーニング、パラメーター、微調整は追加されません。 3 ベンダーの 4 つのジェネレーターにわたる 100 の DailyDilemmas シナリオで、NoT はすべてのモデルでステークホルダーの崩壊を最大 31% から 1% 未満に、不確実性の抑制を最大 72% から 1 ~ 24% に削減しました。予算に合わせた詳細な CoT 制御により、有効成分としてのトークンの使用が除外されます。 NoT は、4 つのジェネレーターのうち 3 つについて、ステークホルダー数で +0.79 ~ +0.90、不確実性スコアで +0.65 ~ +0.93 というクリフのデルタ アドバンテージを保持しており、セクション アブレーションにより、各シフトがその特定のサブ命令に帰属します。 NoT で初期化されたテキスト勾配降下法により、足場がさらに改善されます。ファミリーを超えたトレーニングジャッジ(ジェネレーターとは別のベンダー)が、測定されたすべての軸においてファミリー内のトレーニングジャッジを支配します。 5 ラウンドのマルチステークホルダー討論プロトコルに拡張されたこの足場は、6% の対立をキャリブレーション セットの 95% の完全なコンセンサスと、DailyDilemmas の複製での 100% の結合収束に変換します。結果として得られるトレースは、各コミットメントの根拠となる利害関係者、結果、不確実性を外部化し、信頼性の高いエージェント展開のための監査可能な基盤を提供します。

原文 (English)

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppression (no explicit unknowns or hedges before committing to an action). We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections: protagonist, stakeholders, two-step consequences, uncertainty, then commitment. NoT adds no training, parameters, or fine-tuning. On 100 DailyDilemmas scenarios across four generators from three vendors, NoT cuts stakeholder collapse from up to 31% to under 1% and uncertainty suppression from up to 72% to 1-24% on every model. A matched-budget verbose-CoT control rules out token spend as the active ingredient; NoT retains Cliff's delta advantages of +0.79 to +0.90 on stakeholder count and +0.65 to +0.93 on uncertainty score for three of four generators, and a section ablation attributes each shift to its specific sub-instruction. Textual-gradient descent initialised at NoT improves the scaffold further; a cross-family training judge (different vendor from the generator) dominates an in-family one on every measured axis. Extended to a five-round multi-stakeholder debate protocol, the scaffold converts a 6% standoff into 95% full consensus on a calibration set and 100% combined convergence on a DailyDilemmas replication. The resulting traces externalise the stakeholders, consequences, and uncertainty grounding each commitment, providing an auditable substrate for dependable agentic deployment.

13:00 JST研究/論文

組み合わせ幾何学における極値問題のための幾何学認識 MCTS

私たちは、厳密でグローバルな幾何学的制約を満たす $n \times n$ グリッド内の点の構成を問う、組み合わせ幾何学における特定の極値問題を研究します。古典的な厳密ソルバーは、この種の問題に対して組み合わせ爆発に悩まされ、標準的な強化学習とトランスフォーマーベースのモデルは、報酬がまばらな「妥当性の崖」と二次トークン消費制限に悩まされます。これらのボトルネックを克服するために、Geometry-Aware Monte Carlo Tree Search (MCTS) フレームワークを提案します。私たちのアプローチは、実行可能なアクション空間への増分更新を通じて幾何学的制約を厳密に強制します。古典的な No-Three-in-Line 問題 (Max-N3IL) で発生するような、同一線上にある点の集合に関する制約の場合、このメカニズムにより、制約チェックの複雑さが $O(n^3)$ から $O(n^2)$ に軽減されます。検索効率を向上させるために、2 つの方法で幾何学的対称性を活用します。1 つはノード拡張時の標準枝刈りで分岐係数を削減し、もう 1 つは対称バッチ遷移で有望な構成の発見を加速します。私たちは広範な実験を実施し、検討した問題のうち 6 つのうち 5 つについて、新しく最もよく知られた計算結果を確立しました。特に、Max-N3IL では、サイズ $82 \le n \le 119$ のグリッドに対して、およそ $1.8 n$ のサイズの構成が見つかります。最小完全集合問題では、およそ $0.95 n$ のサイズの構成が見つかり、テストされたグリッド内に新しい上限が提供されます。この研究は、組み合わせ幾何学における新しい構成を発見するための適応性の高いフレームワークとして、幾何学認識 MCTS を確立します。

原文 (English)

Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry

We study certain extremal problems in combinatorial geometry that ask about configurations of points in an $n \times n$ grid that satisfy strict, global geometric constraints. Classical exact solvers suffer from combinatorial explosion for these types of problems, and standard reinforcement learning and transformer-based models struggle with the sparse reward "validity cliff" and quadratic token-consumption limits. To overcome these bottlenecks, we propose a Geometry-Aware Monte Carlo Tree Search (MCTS) framework. Our approach strictly enforces geometric constraints through incremental updates to the feasible action space. For constraints about collections of collinear points, like those that occur in the classic No-Three-in-Line problem (Max-N3IL), this mechanism reduces the constraint checking complexity from $O(n^3)$ to $O(n^2)$. To improve search efficiency, we exploit geometric symmetries in two ways: canonical pruning during node expansion to reduce the branching factor, and symmetric batch transitions to accelerate the discovery of promising configurations. We perform extensive experiments and establish new best-known computational results on five out of six of the problems that we considered. Notably, for Max-N3IL we find configurations of size roughly $1.8 n$ for grids of size $82 \le n \le 119$. For the Smallest Complete Set problem, we find configurations of size roughly $0.95 n$, providing new upper bounds within the tested grids. This work establishes Geometry-Aware MCTS as a highly adaptable framework for discovering novel configurations in combinatorial geometry.

13:00 JSTエージェント

エージェントが電気バス車両の運用に対応する場合: アグリゲーター フレームワークにおける価格設定の動作、トレードオフ、およびポリシーへの影響

エージェント システムは、複雑な運用タスクの調整方法を変え、異種データ ソースを接続し、プロセスを自動化するための新しいパラダイムを導入しています。電気バス車両は関連するテストケースを提供します。その運用には、サービスの信頼性、バッテリーの充電状態、充電器の可用性、電力価格、エネルギー経路の不確実性、および車両から送電網への (V2G) 機会の間の継続的な調整が必要です。この論文では、最適化ベースの電気バスのスケジューリング モデルと、障害の検出、料金の適応、およびスケジュールの評価のための監視エージェントを組み合わせることで、この意思決定環境を合理化するエージェント アグリゲーター フレームワークを提案します。最適化コアはルート、充電器、バッテリー、V2G 交換機全体で物理的な実現可能性を強制します。一方、エージェント層は変化する動作条件を解釈し、必要に応じてリアルタイムの再最適化をトリガーし、アグリゲーターと公共交通機関 (PTO) の間で柔軟性の価値を割り当てる方法を定義します。現実的な車両基地のケーススタディでは、サービスの遅延、路線エネルギーの逸脱、電力価格のショック、複合的な外乱を考慮して、利益ベースおよび運用ベースの調整モードの下で、前日およびリアルタイムの運用を評価します。結果は、エージェントアグリゲーションが、実行可能なスケジュールを維持し、選択的に再最適化をアクティブ化し、課金と V2G の柔軟性を向上させることにより、適応的なフリート グリッド調整をサポートできることを示しています。ただし、これらは重要なトレードオフも明らかにしています。利益重視の価格設定を中心に構成されている場合、運用の複雑さを軽減する同じエージェント機能が PTO から価値を引き出す可能性があります。これらの調査結果は、エージェント・アグリゲーターが電気バスの V2G 運用の管理に役立つ可能性があることを示唆していますが、公共車両のコンテキストでの導入には、透明性のある調整モード、監査可能な料金設定、および明示的な価値共有ルールが必要です。

原文 (English)

When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework

Agentic systems are changing how complex operational tasks are coordinated, introducing a new paradigm for connecting heterogeneous data sources and automating processes. Electric bus fleets provide a relevant test case. Their operation requires continuous coordination between service reliability, battery state-of-charge, charger availability, electricity prices, route-energy uncertainty, and vehicle-to-grid (V2G) opportunities. This paper proposes an agentic aggregator framework that streamlines this decision environment by coupling an optimization-based electric bus scheduling model with supervisory agents for disturbance detection, tariff adaptation, and schedule evaluation. The optimization core enforces physical feasibility across routes, chargers, batteries, and V2G exchanges, while the agentic layer interprets changing operating conditions, triggers real-time re-optimization when needed, and defines how flexibility value is allocated between the aggregator and the public transport operator (PTO). A realistic depot case study evaluates day-ahead and real-time operations under profit-based and operation-based coordination modes, considering service delays, route-energy deviations, electricity price shocks, and combined disturbances. The results show that agentic aggregation can support adaptive fleet-grid coordination by maintaining feasible schedules, activating re-optimization selectively, and improving the use of charging and V2G flexibility. However, they also reveal a critical trade-off: the same agentic capability that reduces operational complexity can extract value from the PTO when configured around profit-oriented pricing. These findings suggest that agentic aggregators can become useful for managing electric bus V2G operations, but their deployment in public-fleet contexts requires transparent coordination modes, auditable tariff-setting, and explicit value-sharing rules.

13:00 JSTエージェント

格子理論による不偏正準集合値オラクル

将来の出来事の確率を推定する非エージェント型の「オラクル」AI は、自己参照の問題に直面しています。その答えが学習され、実行されると、報告するよう求められた確率そのものが変わってしまう可能性があります。 Scientist AI プログラムで提唱されている対応の 1 つは、事実に反する質問のみをし、その回答が何の影響もないかのように評価することです。私たちは、そのような答えは学んだ瞬間に意味がなくなってしまう傾向があることを観察しています。それは、まさにその前提が間違っているからです。したがって、私たちは、オラクルが単一の確率ではなく、同時に偏りがなく学習の結果と自己矛盾のない一連の資格を報告する、自己言及的な代替案を模索します。単純な自己一貫性の要件は、あまりにも多くのセット (役に立たない答え $[0,1]$ を含む) によって満たされるため、問題は正規の自明でないメンバーを選び出すことです。これを、適切に定義されたアイソトーン演算子の最小不動点を取り、閉じたクレダル集合の完全な格子に関するクナスター-タルスキーの不動点定理を使用して行います。代わりに、バリアントは、すべての自己矛盾のない点推定を含む最小不動点を報告します。我々は、存在、自己無矛盾性、空でないことを証明し、非実行的質問についてはその構造が古典的な点の答えに崩壊すること、そしてバイナリイベントについては標準的な答えが自然なハル因数分解の仮定の下では区間であることを示します。この展開は純粋に格子理論に基づいており、バイナリ イベント $B$ から任意の確率変数 $X$ までそのまま拡張され、$P(B\mid A,C)$ は条件法 $\mathcal{L}(X\mid A,C)$ に置き換えられます。区間の特徴付け自体がその一般化に耐えられるかどうかを含む未解決の質問で終わります。

原文 (English)

Unbiased Canonical Set-Valued Oracles Via Lattice Theory

A non-agentic "oracle" AI that estimates probabilities of future events faces a self-reference problem: once its answer is learned and acted upon, it can change the very probability it was asked to report. One response, advocated for the Scientist AI programme, is to ask only counterfactual questions, evaluated as if the answer had no influence. We observe that such answers tend to become irrelevant the moment they are learned, precisely because their premise is then false. We therefore explore a self-referential alternative in which the oracle reports not a single probability but a credal set that is simultaneously unbiased and self-consistent with the consequences of being learned. The naive self-consistency requirement is satisfied by too many sets (including the useless answer $[0,1]$), so the problem is to single out a canonical, nontrivial member. We do so with the Knaster--Tarski fixed-point theorem on the complete lattice of closed credal sets, taking the least fixed point of a suitably defined isotone operator; a variant instead reports the least fixed point that contains every self-consistent point estimate. We prove existence, self-consistency, and nonemptiness, show that the construction collapses to the classical point answer for non-performative questions, and that for a binary event the canonical answer is, under a natural hull-factoring assumption, an interval. The development is purely lattice-theoretic and extends unchanged from a binary event $B$ to an arbitrary random variable $X$, with $P(B\mid A,C)$ replaced by the conditional law $\mathcal{L}(X\mid A,C)$. We close with open questions, including whether the interval characterization itself survives that generalization.

13:00 JST研究/論文

大規模な言語モデルと入れ子になったデータへのアプリケーションによる分類器のパフォーマンスの不確実性の推定

研究者は、自然言語からの構成要素を測定するためにテキスト分類 (教師ありモデルまたは大規模言語モデル) を使用することが増えており、その妥当性の証拠として再現率や精度などの指標を提供しています。しかし、これらの指標はサンプリング変動の影響を受ける点推定値であるにもかかわらず、不確実性の尺度がそれらと一緒に報告されるのは一貫性がありません。さらに、それらが報告される場合、関連するラベル付きデータセットが小さい場合やパフォーマンスが高い場合には、適切ではない方法で推定されることがよくあります。現場での信頼区間レポートを増加および改善するために、この論文では、社会科学のテキスト分類に典型的な条件、つまり小規模から中程度のサンプルサイズ、頻度の低い構成、および個人内にネストされたテキストの下で、パフォーマンスメトリクスの信頼区間手法を評価します。シミュレーション全体で、Wald 間隔や基本パーセンタイル ブートストラップなどのデフォルトの方法は精度が最も低く、カバレッジが名目 95% レベルを大幅に下回る場合があります。 Agresti-Coull、Wilson、Clopper-Pearson、および新しい擬似カウント正規化ブートストラップ (特に F1 の計算に関連する) を使用することで、精度が向上します。テキストが個人内でネストされている場合、正確な分析間隔を生成するには実効 N と適切な自由度の両方の調整が必要であることを示します。ブートストラップ間隔の中で、個人が中程度の数のテキストを作成する場合、階層ブートストラップはクラスター ブートストラップよりも正確ですが、個人が少数しか作成しない場合は過度に保守的になります。適切な間隔推定に関するガイダンスを現場に提供することで、機械学習アプリケーションの透明性を向上させ、設計段階での検証サンプル サイズに対するさらなる注意を促すことを目指しています。

原文 (English)

Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data

Researchers increasingly use text classification--supervised models or large language models--to measure constructs from natural language, providing metrics such as recall and precision as evidence of their validity. Yet, though these metrics are point estimates subject to sampling variation, measures of uncertainty are inconsistently reported alongside them. Further, when they are reported, they are often estimated with methods that are not appropriate when relevant labelled datasets are small or performance is high. To increase and improve confidence interval reporting in the field, this paper evaluates confidence interval methods for performance metrics under conditions typical of social science text classification: small to moderate sample sizes, infrequent constructs, and texts nested within individuals. Across simulations, default methods such as the Wald interval and the basic percentile bootstrap are the least accurate, with coverage sometimes far below the nominal 95% level. Accuracy is improved with the use of Agresti-Coull, Wilson, Clopper-Pearson, and a novel pseudo-count regularized bootstrap (which is particularly relevant to the calculation of F1). When texts are nested within individuals, we demonstrate that adjustment for both effective N and the appropriate degrees of freedom is necessary for producing accurate analytic intervals. Among bootstrap intervals, the hierarchical bootstrap is more accurate than the cluster bootstrap when individuals produce a moderate number of texts but overly conservative when individuals produce only a few. By providing guidance to the field on appropriate interval estimation, we aim to improve the transparency of machine learning applications, and to encourage greater attention to the validation sample size at the design stage.

13:00 JST研究/論文GPT / ChatGPT

データ駆動型機械学習は記号レベルの論理的推論に到達できない -- スケーリング則の限界

Sphere ニューラル ネットワークは、トレーニング データなしで記号レベルの三段論的推論を達成しました。これにより、論理的推論のスケーリング則の限界がどこにあるのか、つまり、データ駆動型の機械学習システムがトレーニング データとトレーニング時間を増やすことで同じレベルを達成できるかどうかという問題が生じています。教師あり深層学習が記号レベルの三段論的推論に到達することを妨げる 2 つの方法論的制限を示します。(1) トレーニング データは、24 種類の有効な三段論的推論すべてを区別できない。 (2) 前提から結論までのエンドツーエンドのマッピングでは、パターン認識と論理的推論のための神経コンポーネント間に矛盾するトレーニング ターゲットが導入されます。理論的な分析に加えて、オイラー ネットでは厳密な三段論的推論を達成できないことを実験的に示します。さらに、最新の ChatGPT (GPT-5-nano および GPT-5) に対して、単語、二重単語、単純なシンボル、および長いランダム記号の 4 つの表面形式 (パターン) で三段論法的ステートメントの充足可能性を判定することに挑戦し、表面形式が推論パフォーマンスに影響を与えること、および ChatGPT GPT-5 が 100% の精度に達する可能性があるが、依然として不正確な説明を提供する可能性があることを示します。経験的トレーニングプロセスは 100% の精度に達した後に停止されるため、教師あり機械学習システムは記号論理推論の厳密さを達成できないと結論付けます。

原文 (English)

Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law

Sphere neural networks have achieved symbolic level syllogistic reasoning without training data, raising the question of where the limit of the scaling law for logical reasoning lies, i.e., whether data-driven machine learning systems can achieve the same level by increasing training data and training time. We show two methodological limitations that prevent supervised deep learning from reaching the symbolic-level syllogistic reasoning: (1) training data can not distinguish all 24 types of valid syllogistic reasoning; (2) end-to-end mapping from premises to conclusion introduces contradictory training targets between neural components for pattern recognition and logical reasoning. Beside theoretical analysis, we experimentally illustrate that Euler Net cannot achieve rigorous syllogistic reasoning. We further challenge the most recent ChatGPTs (GPT-5-nano and GPT-5) to determine the satisfiability of syllogistic statements in four surface forms (patterns): words, double words, simple symbols, and long random symbols, showing that surface forms affect the reasoning performance and that ChatGPT GPT-5 may reach 100% accuracy but still provide incorrect explanations. As empirical training processes are stopped after achieving 100% accuracy, we conclude that supervised machine learning systems will not attain the rigour of symbolic logical reasoning.

13:00 JST研究/論文

MKG-RAG-Bench: マルチモーダルナレッジグラフ拡張生成におけるベンチマーク取得

ナレッジ グラフ上の検索拡張生成 (RAG) は、大規模な言語モデルを基礎付けるための有望なアプローチとして浮上していますが、既存のベンチマークでは、マルチモーダル ナレッジ グラフ RAG (MKG-RAG) における検索の課題がほとんど見落とされています。実際には、検索は重大なボトルネックです。マルチモーダルな知識は異質であり、モダリティ間で調整するのが難しく、構造化されていないコーパス向けに設計された検索ツールでは十分に機能しないことがよくあります。このギャップに対処するために、MKG-RAG での取得を評価するために明示的に設計されたクロスドメイン ベンチマークである MKG-RAG-Bench を導入します。 MKG-RAG-Bench は、一般領域と医療領域にわたる 2 つのマルチモーダル ナレッジ グラフから構築されており、検索と下流生成の両方の制御された評価をサポートする慎重に調整された質問応答データセットが含まれています。このベンチマークは、実用性の低いナレッジをフィルタリングし、正確な監視のもとで構造的に根拠のあるクエリを生成し、多様なモダリティ構成を体系的にカバーする LLM ベースのキュレーション パイプラインを使用して構築されています。代表的なレトリーバーファミリーとモダリティ設定にわたる広範な実験を通じて、効果的なマルチモーダル検索は依然として課題であるものの、エンドツーエンドの MKG-RAG パフォーマンスにとって重要であること、および検索品質が生成結果を強く決定することを示します。 MKG-RAG-Bench は、検索を第一級の評価対象として分離することで、現在の制限を診断し、マルチモーダル ナレッジ グラフ RAG システムを進歩させるための原則的な基盤を提供します。

原文 (English)

MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

Retrieval-augmented generation (RAG) over knowledge graphs has emerged as a promising approach for grounding large language models, yet existing benchmarks largely overlook the challenges of retrieval in multimodal knowledge graph RAG (MKG-RAG). In practice, retrieval is a critical bottleneck: multimodal knowledge is heterogeneous, difficult to align across modalities, and often poorly served by retrievers designed for unstructured corpora. To address this gap, we introduce MKG-RAG-Bench, a cross-domain benchmark explicitly designed to evaluate retrieval in MKG-RAG. MKG-RAG-Bench is constructed from two multimodal knowledge graphs spanning general and medical domains, and includes carefully aligned question-answering datasets that support controlled evaluation of both retrieval and downstream generation. The benchmark is built using an LLM-based curation pipeline that filters low-utility knowledge, generates structurally grounded queries with exact supervision, and systematically covers diverse modality configurations. Through extensive experiments across representative retriever families and modality settings, we show that effective multimodal retrieval remains challenging yet crucial for end-to-end MKG-RAG performance, and that retrieval quality strongly determines generation outcomes. By isolating retrieval as a first-class evaluation target, MKG-RAG-Bench provides a principled foundation for diagnosing current limitations and advancing multimodal knowledge graph RAG systems.

13:00 JSTエージェント

auto-psych: エージェント駆動理論の発見と実験を使用した心の科学の自動化

エージェントを使用して仮説を生成し、実験を計画し、データを分析することにより、AI ベースの科学的自動化がますます可能になります。ただし、このパイプラインではデータ収集が大きなボトルネックになっています。心理学、特に計算認知科学は、理論がコードとして表されることが多く、クラウドソーシング プラットフォームによりプログラムによる人間データの大規模収集が可能になるため、AI 実験の恩恵を受ける有利な立場にあります。ここでは、クラウドソーシングによる調査実験を通じて人間のデータを独立して収集するエージェントベースのシステムを使用して、自動発見技術を計算認知科学の理論生成プロジェクトに適用します。テストベッドとして、認知心理学の古典的なケーススタディを使用します。コイン投げのどのシーケンスが主観的によりランダムに見えるかを判断します。私たちのシステム auto-psych は、入れ子になったエージェントベースの発見ループを使用して、人間の行動の説明理論を生成します。内側のループは、確率的認知モデルを推測、適合、および批判します。外側のループは、これらのモデルをテストするための実験を設計し、オンラインで起動して、データを分析します。このシステムは、体系的な実験を通じて合成データからグラウンドトゥルース理論を迅速かつ確実に復元できますが、モデルのパフォーマンスには入れ子構造が重要です。さらに、人体実験の 3 つの独立したシーケンスにおいて、システムは科学文献から生成された理論よりもデータによく適合する理論を見つけます。したがって、この研究は、計算認知科学における自動データ収集と理論発見の実現可能性を実証しています。

原文 (English)

auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation

AI-based scientific automation is increasingly possible by using agents to generate hypotheses, design experiments, and analyze data. Data collection is a major bottleneck in this pipeline, however. Psychology, and computational cognitive science in particular, is well-positioned to benefit from AI experimentation because theories are often represented as code and crowdsourcing platforms enable programmatic human data collection at scale. Here, we apply automated discovery techniques to the project of generating theories in computational cognitive science, with an agent-based system collecting human data independently through crowdsourced survey experiments. As a testbed, we use a classic case study from cognitive psychology: judging which sequences of coin flips seem subjectively more random. Our system, auto-psych, uses nested agent-based discovery loops to generate explanatory theories of human behavior. The inner loop conjectures, fits, and critiques probabilistic cognitive models; the outer loop designs experiments to test these models, launches them online, and analyzes the data. This system can quickly and reliably recover ground-truth theories from synthetic data via systematic experimentation, but the nested structure is critical to model performance. Further, in three independent sequences of human experiments, the system finds theories that fit the data better than theories generated from the scientific literature. This work thus demonstrates the feasibility of automated data collection and theory discovery in computational cognitive science.

13:00 JST研究/論文

統治可能な医療 AI スキル エコシステムのための臨床ハーネス

医療 AI は依然として孤立したモデルを中心に組織化されていますが、臨床ケアには時間を超えて持続する責任ある機能が必要です。私たちは、臨床 AI スキルと Clinical Harness を提案します。これは、AI 対応の臨床機能を登録、調整、保護、監視するためのランタイム ガバナンス アーキテクチャです。骨粗鬆症を例として使用し、知識主導型、データ主導型、および物理学に基づいて強化されたスキルが、ランタイム ガバナンスの下でライフサイクル ケアをどのようにサポートできるかを示します。

原文 (English)

Clinical Harness for Governable Medical AI Skill Ecosystems

Medical AI remains organized around isolated models, whereas clinical care requires accountable capabilities that persist across time. We propose clinical AI skills and the Clinical Harness: a runtime governance architecture for registering, orchestrating, guarding and monitoring AI-enabled clinical capabilities. Using osteoporosis as an exemplar, we show how knowledge-driven, data-driven and physics-enhanced skills can support lifecycle care under runtime governance.

13:00 JSTLLM/生成AI

人間は関与をやめ、推論モデルは存続: 難易度の登録と審議の割り当てを分離する

大規模推論モデル (LRM) は、人間と同じように、より困難な問題に時間がかかります。この表面の類似性は、アイテム内に反対のパターンを隠します。 LRM が問題を間違えると、同じ問題を正解した場合よりも多くのトークンを消費します。人間はその逆を行い、間違った試験に費やす時間を減らします。検討を 2 つのレベルに分けます。1 つは応答時間が項目全体の難易度をどのように追跡するか (登録)、もう 1 つは項目の ID が固定された状態で、エージェントが自身の失敗と成功のどちらに多くの時間を費やすか (割り当て) です。公開されているヒトと LRM の照合コーパスでは、人間と 5 つの思考 LRM はすべて、既知の項目間アライメント (登録) を再現しますが、項目 (割り当て) 内では分岐します。どの LRM も大きな誤対正効果 (H-ARC におけるコーエンの d = 1.47-3.13) を示しますが、人間は反対の符号を示します。比較は各エージェント独自のスケール内に留まります。秒とトークンを 1 つの軸に置くことはありません。解離はアイテムの固定効果の下で保持され、データセット全体で複製され、非思考ベースラインには存在しません。私たちは人間のパターンを、関与対放棄として読みます。人々は、解決できると期待している項目に留まり、残りの項目は放棄します。 LRM パターンを不確実性によって引き起こされる長さとして読み取ります。モデルが不確実な場合、チェーンは成長します。つまり、モデルが失敗する傾向があるのはまさにこの時です。どちらのポリシーも同じ項目間の相関関係を生成するのは困難ですが、以前の研究で使用された尺度に基づいて一致しているように見えます。相違は、アイテムの同一性が固定された場合にのみ現れます。リソース合理的メタ推論では、困難信号は共有するが反対の制御を実装する 2 つの停止ポリシーの間で分割が行われます。トレース長が信号を捕捉し、制御を逃します。

原文 (English)

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong, it spends more tokens than when it gets the same problem right; humans do the reverse, spending less time on the trials they get wrong. We separate two levels of deliberation: how response time tracks difficulty across items (registration), and, with item identity held fixed, whether an agent spends more on its own failures or successes (allocation). On a public matched human-LRM corpus, humans and all five thinking LRMs reproduce the known cross-item alignment (registration) but diverge within items (allocation): every LRM shows a large wrong-vs-right effect (Cohen's d = 1.47-3.13 on H-ARC) while humans show the opposite sign. The comparison stays inside each agent's own scale; we never put seconds and tokens on one axis. The dissociation holds under item fixed effects, replicates across datasets, and is absent in a non-thinking baseline. We read the human pattern as engagement versus abandonment: people stay on items they expect to solve and give up on the rest. We read the LRM pattern as length driven by uncertainty: chains grow when the model is unsure, which is exactly when it tends to fail. Both policies produce the same cross-item correlation with difficulty, so they look aligned on the measure prior work has used; the divergence shows up only once item identity is fixed. Under resource-rational metareasoning, the split is between two stopping policies that share a difficulty signal but implement opposite control; trace length captures the signal and misses the control.

13:00 JSTエージェント研究/論文Claude

NeuraDock Visual Cognitive Load Agent チュートリアル: アルファ ダイナミクスおよびリアルタイム アプリケーション向けの品質ゲート付きオープンソース EEG ワークフロー

このチュートリアル ペーパーでは、アルファ ダイナミクスと視覚的認知負荷分析に焦点を当てたオープンソースの EEG エージェントである NeuraDock Agent の段階的な再現可能なウォークスルーを提供します。目標は実践的です。読者は、エージェントのインストール、EEG 前処理と品質管理の実行、アルファ ダイナミクス図の生成、被験者内での休息/タスクの視覚的認知負荷比較の実行、公開ミニ データセット分析の実行と参照検証の概要との比較、オンライン ダッシュボードの開始、外部アプリケーションからのリアルタイム API の呼び出し、LLM 解釈レイヤーを使用して品質リスクを説明できる必要があります。既存の EEG ツールキットは優れたオフライン分析を提供しますが、リアルタイムの品質ゲート認知負荷パイプラインを構築するには、多くの場合、手動でのブリッジング取得、カスタム QC、アルファ特徴抽出、および Web API が必要になります。このチュートリアルでは、オフラインとオンラインのギャップを埋めます。このチュートリアルでは、品質ゲートのワークフローを使用します。ダウンストリームのアルファとワークロードのメトリクスは、生の EEG から直接計算されるのではなく、前処理と QC ゲートの後でのみ計算されます。含まれているミニデータセット検証では、エージェントは 18 件の記録を処理し、10 件の被験者内比較を生成し、10 件のコントラストのうち 7 件でタスク関連の事後アルファ抑制を観察し、被験者内再現性の初期証拠を推定し、ローカル オンライン API レイテンシーのベンチマークを行いました。このチュートリアルは、EEG ファイルからリアルタイムの視覚的認知負荷プロトタイプへの透過的なパスを必要とする研究者、開発者、応用チームを対象としています。

原文 (English)

NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications

This tutorial paper provides a step-by-step, reproducible walkthrough of NeuraDock Agent, an open-source EEG agent focused on Alpha dynamics and visual cognitive-load analysis. The goal is practical: a reader should be able to install the agent, run EEG preprocessing and quality control, generate Alpha dynamics figures, perform within-subject Rest/Task visual cognitive-load comparison, run the public mini-dataset analyses and compare them with the reference validation summary, start an online dashboard, call the real-time API from an external application, and use the LLM interpretation layer to explain quality risks. Existing EEG toolkits provide excellent offline analysis, but assembling a real-time, quality-gated cognitive-load pipeline often requires manually bridging acquisition, custom QC, Alpha feature extraction, and a web API; this tutorial closes that offline-to-online gap. The tutorial uses a quality-gated workflow: downstream Alpha and workload metrics are computed only after preprocessing and QC gating rather than directly from raw EEG. In the included mini-dataset validation, the agent processed 18 recordings, generated 10 within-subject comparisons, observed task-related posterior Alpha suppression in 7 of 10 contrasts, estimated initial evidence of within-subject repeatability, and benchmarked local online API latency. The tutorial is intended for researchers, developers, and applied teams who want a transparent path from EEG files to real-time visual cognitive-load prototypes.

13:00 JSTLLM/生成AIエージェント

低チャネルEEGエージェントのための境界を意識したコンテキストグラウンディング

大規模言語モデル (LLM) を使用すると、科学ソフトウェアを使いやすくできます。ただし、一般的なモデルでは、特定のセンサーがどの測定をサポートできるか、現在のソフトウェアにどのアルゴリズムが実装されているか、または計算結果によってどの結論が正当化されるかが自動的にはわかりません。これらの区別は、低チャネル脳波検査 (EEG) では特に重要です。EEG では、空間範囲がまばらで信号品質が変動するため、もっともらしいが裏付けのない解釈が容易に生成されます。私たちは、決定論的なローカル EEG エンジンをハードウェア対応言語層から分離するオープンソース アーキテクチャである NeuraDock Agent を紹介します。数値エンジンは録音を解析し、品質管理を実行し、レビューされたスペクトル ワークフローを実行して、機械可読アーティファクトを書き込みます。 LLM は、コンパクトな許可リストに登録された概要とバージョン管理されたコンテキスト パックのみを受け取ります。コンテキストでは、7 チャネルのハードウェア、レビューされたワークフロー、結果フィールド、実装の境界、科学的限界、参照ケースについて説明します。生の EEG と高密度のサンプルごとの配列はローカルのままです システムを 3 つのレベルで評価します。まず、12 回の記録では、10 回の数値繰り返しで同一の構造化された結果が生成され、完全な休憩/タスクの実行では、3 回の繰り返しで同一の結果、レポート、および図のハッシュが生成されました。次に、リクエスト キャプチャと障害挿入の実験により、テスト済みのデータ境界と、HTTP、不正な出力、および接続障害下でのローカル アーティファクトの保存が確認されました。第三に、境界認識ベンチマークは、4 つのコンテキスト アブレーションと 2 つの LLM の下で 36 の通常の質問と敵対的な質問をテストし、288 の出力を生成しました。これらの結果は、EEG エージェントが何を受け入れるか、認定するか、または拒否するかを調整するための実用的なメカニズムとして、ハードウェアおよび実装を意識したグラウンディングをサポートします。それらは臨床的妥当性や検証された絶対的な認知負荷指数を確立するものではありません。

原文 (English)

Boundary-Aware Context Grounding for A Low-Channel EEG Agent

Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatically know which measurements a particular sensor can support, which algorithms are implemented in the current software, or which conclusions are justified by a computed result. These distinctions are especially important for low-channel electroencephalography (EEG), where sparse spatial coverage and variable signal quality make plausible but unsupported interpretations easy to produce. We present NeuraDock Agent, an open-source architecture that separates a deterministic local EEG engine from a hardware-aware language layer. The numerical engine parses recordings, performs quality control, executes reviewed spectral workflows, and writes machine-readable artifacts. The LLM receives only a compact, allowlisted summary and a versioned context pack. The context describes the seven-channel hardware, reviewed workflows, result fields, implementation boundaries, scientific limits, and reference cases. Raw EEG and dense per-sample arrays remain local We evaluate the system at three levels. First, 12 recordings produced identical structured results over ten numerical repetitions, and a complete Rest/Task run produced identical result, report, and figure hashes over three repetitions. Second, request-capture and failure-injection experiments confirmed the tested data boundary and preservation of local artifacts under HTTP, malformed-output, and connection failures. Third, a boundary-awareness benchmark tested 36 ordinary and adversarial questions under four context ablations and two LLMs, yielding 288 outputs.These results support hardware- and implementation-aware grounding as a practical mechanism for calibrating what an EEG agent accepts, qualifies, or refuses; they do not establish clinical validity or a validated absolute cognitive-load index.

13:00 JSTエージェント

革新的な AI 解釈可能性

私たちは、急進的な解釈の哲学的伝統と機械的な解釈可能性のツールを利用して、AI システムをエージェントとして解釈するためのフレームワークを開発します。核心的な問題は、システムに関する計算上の事実が与えられた場合、その信念、欲求、および意味をどのように解決するかということです。これは安全性にとってますます重要です。私たちは、その目的を理解することによって、あるいはもっと控えめに言って、欺瞞を確実に検出することによって、導入したシステムを信頼できるようにしたいと考えています。解釈可能性の研究者は、モデルの内部から信念や欲求を読み取るツールを構築していますが、そのようなツールがいつ成功したかについては、明確な説明はありません。この本はその1つを提供します。私たちは表現主義的アプローチと解釈主義的アプローチの両方に関する基準を提案し、それぞれを現在の解釈可能性手法が実行できるテストに結び付けます。中心的な教訓は、これらの帰属を断片的に作成することはできないということです。信念、欲望、およびそれらが前提とする命題構造は共に制約されており、一方を修正しながら他方を測定する方法は、導入される歪みがすべて継承されます。この全体性は、通訳者の概念を共有していない可能性がある AI システムにとって急務となっています。しかし、それはまた、てこにもなります。つまり、システムの態度はその命題構造を制約し、その構造はどの態度が帰属するかを制約し、機械的解釈可能性は両方を測定するのに役立ちます。

原文 (English)

Radical AI Interpretability

We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational facts about a system, how do we solve for its beliefs, desires, and meanings? This matters increasingly for safety. We want to be able to trust the systems we deploy, whether by understanding their goals or, more modestly, by reliably detecting deception. Interpretability researchers are building tools to read beliefs and desires off a model's internals, but there is no settled account of when such a tool has succeeded. This book supplies one. We propose criteria on both representationalist and interpretationist approaches, and tie each to tests current interpretability methods can carry out. A central lesson is that these attributions cannot be made piecemeal. Beliefs, desires, and the propositional structure they presuppose are jointly constrained, and a method that fixes one while measuring the others inherits whatever distortions that introduces. This holism becomes pressing for AI systems, which may not share the interpreter's concepts. However, it also provides leverage: a system's attitudes constrain its propositional structure, that structure constrains which attitudes can be attributed, and mechanistic interpretability can help us measure both.

13:00 JST研究/論文

PMDformer: 長期予測用のパッチ平均デカップリング情報トランスフォーマー

長期時系列予測 (LTSF) は、エネルギー管理、金融、交通予測などの分野で重要な役割を果たします。トランスフォーマーベースのモデルは、長距離の依存関係を把握するためにパッチベースの戦略を採用していますが、パッチと変数間の形状の類似性を正確にモデル化することは、スケールの違いにより依然として困難です。これに対処するために、パッチ平均デカップリング (PMD) を導入します。これは、各パッチの平均を差し引くことでトレンドと残差の形状情報を分離し、元の構造を保存し、アテンション メカニズムが真の形状の類似性を確実に捕捉するようにします。さらに、長期依存関係をより効果的にモデル化し、変数間関係を捕捉するために、トレンド復元アテンション (TRA) と近接変数アテンション (PVA) を提案します。前者のモジュールは、注意出力を計算しながら、PMD から切り離されたトレンドを再統合します。そして後者は、古い相関関係での過剰適合を避けるために、最も関連性の高い最近の時間セグメントに変数間の注意を集中させます。これらのコンポーネントを組み合わせて、長期予測シナリオで形状の類似性を効果的に捕捉するように設計されたモデルである PMDformer を提案します。広範な実験により、PMDformer は複数の LTSF ベンチマークにわたって安定性と精度において既存の最先端の手法よりも優れていることが示されています。コードは https://github.com/aohu1105/PMDformer で入手できます。

原文 (English)

PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting

Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-based strategies to capture long-range dependencies, but accurately modeling shape similarities across patches and variables remains challenging due to scale differences. To address this, we introduce patch-mean decoupling (PMD), which separates the trend and residual shape information by subtracting the mean of each patch, preserving the original structure and ensuring that the attention mechanism captures true shape similarities. Futhermore, to more effectively model long-range dependencies and capture cross-variable relationships, we propose Trend Restoration Attention (TRA) and Proximal Variable Attention (PVA). The former module reintegrates the decoupled trend from PMD while calculating attention output. And the latter focuses cross-variable attention on the most relevant, recent time segments to avoid overfitting on outdated correlations. Combining these components, we propose PMDformer, a model designed to effectively capture shape similarity in long-term forecasting scenarios. Extensive experiments indicate that PMDformer outperforms existing state-of-the-art methods in stability and accuracy across multiple LTSF benchmarks. The code is available at https://github.com/aohu1105/PMDformer.

13:00 JST研究/論文

C型肝炎患者における肝硬変の存在を検出するための説明可能なアンサンブルベースの機械学習モデル

C型肝炎は、ウイルスによって引き起こされる肝臓感染症であり、肝臓に軽度から重度の炎症を引き起こします。 C型肝炎は長年にわたって徐々に肝臓にダメージを与え、多くの場合、肝硬変として知られる永久的な瘢痕化につながります。患者は、肝硬変を発症するまで数十年間、中等度の肝疾患の症状を呈するか、まったく症状を示さない場合があります。肝硬変は通常、肝不全に至るまで悪化します。肝硬変患者は、胃腸出血だけでなく、脳や神経系の損傷も経験する可能性があります。肝硬変の治療は、病気のさらなる進行を防ぐことに重点を置きます。したがって、肝硬変を早期に検出することは、合併症を回避するために非常に重要です。機械学習 (ML) は、いくつかの病気の診断に使用するための正確かつ正確な情報を提供するのに効果的であることが示されています。それにもかかわらず、これまでのところ、C 型肝炎患者の肝硬変を検出するために ML を使用した研究はありません。この研究では、カリフォルニア大学アーバイン校の ML リポジトリから 2038 人のエジプト人患者の 28 属性で構成されるデータセットを入手しました。 C 型肝炎患者の肝硬変を診断するために、ランダム フォレスト、勾配ブースティング マシン、極端勾配ブースティング、およびエクストラ ツリー モデルの 4 つの ML アルゴリズムがデータセットでトレーニングされました。 Extra Trees モデルは他のモデルを上回り、28 個の特徴のうち 16 個のみを使用して、精度 96.92%、再現率 94.00%、精度 99.81%、受信機動作特性曲線下面積 96% を達成しました。

原文 (English)

Explainable Ensemble-Based Machine Learning Models for Detecting the Presence of Cirrhosis in Hepatitis C Patients

Hepatitis C is a liver infection caused by a virus, which results in mild to severe inflammation of the liver. Over many years, hepatitis C gradually damages the liver, often leading to permanent scarring, known as cirrhosis. Patients sometimes have moderate or no symptoms of liver illness for decades before developing cirrhosis. Cirrhosis typically worsens to the point of liver failure. Patients with cirrhosis may also experience brain and nerve system damage, as well as gastrointestinal hemorrhage. Treatment for cirrhosis focuses on preventing further progression of the disease. Detecting cirrhosis earlier is therefore crucial for avoiding complications. Machine learning (ML) has been shown to be effective at providing precise and accurate information for use in diagnosing several diseases. Despite this, no studies have so far used ML to detect cirrhosis in patients with hepatitis C. This study obtained a dataset consisting of 28 attributes of 2038 Egyptian patients from the ML Repository of the University of California at Irvine. Four ML algorithms were trained on the dataset to diagnose cirrhosis in hepatitis C patients: a Random Forest, a Gradient Boosting Machine, an Extreme Gradient Boosting, and an Extra Trees model. The Extra Trees model outperformed the other models achieving an accuracy of 96.92%, a recall of 94.00%, a precision of 99.81%, and an area under the receiver operating characteristic curve of 96% using only 16 of the 28 features.

13:00 JSTLLM/生成AI

EvoOptiGraph: 最適化モデリングのためのグラフベースの構造生成による弱さ主導の共進化

大規模言語モデル (LLM) を使用した自然言語からの最適化モデリングの自動化は、2 つの重要な課題に直面しています。まず、トレーニング コーパスには構造的な多様性がありません。第 2 に、データ生成パイプラインは静的なままであり、モデル学習から切り離されています。これらの課題に対処するために、モデルの弱点に基づいてデータとモデルが共進化する新しいフレームワークである EvoOptiGraph を提案します。 EvoOptiGraph は、各混合整数線形計画 (MILP) を属性付きの 2 部グラフとして表し、妥当性を保持する進化的演算子を適用して構造的に多様なインスタンスを生成します。進化したグラフは、決定論的コンパイルと検証された逆変換を介してソルバー コードと自然言語に変換されます。トレーニングは 2 段階で進行します。1 つは初期データセットに対する教師あり微調整 (SFT)、続いて検証可能な報酬を伴う強化学習 (RLVR) で、グラフ由来の弱点シグナルがモデルの失敗を対象とした新しいインスタンスの生成をガイドします。これにより、トレーニング分布を継続的に更新する閉ループが形成されます。 6 つの公開データセットに関する実証結果は、EvoOptiGraph が、精度、実行可能性、一般化の点で、大規模なジェネラリスト モデル、エージェント手法、特殊なベースラインよりも大幅に優れていることを示しています。これらの結果は、ターゲットを絞ったデータモデルの共進化が、最適化モデリング タスクで LLM を改善するための効果的な戦略であることを示しています。

原文 (English)

EvoOptiGraph: Weakness-Driven Coevolution via Graph-Based Structural Generation for Optimization Modeling

Automating optimization modeling from natural language with large language models (LLMs) faces two key challenges. First, training corpora lack structural diversity. Second, data generation pipelines remain static and decoupled from model learning. To address these challenges, we propose EvoOptiGraph, a novel framework where data and model co-evolve, driven by model weaknesses. EvoOptiGraph represents each mixed-integer linear program (MILP) as an attributed bipartite graph and applies validity-preserving evolutionary operators to generate structurally diverse instances. The evolved graphs are converted into solver code and natural language via deterministic compilation and verified back-translation. Training proceeds in two stages: supervised fine-tuning (SFT) on an initial dataset, followed by reinforcement learning with verifiable rewards (RLVR), where graph-derived weakness signals guide the generation of new instances targeting the model's failures. This forms a closed loop that continuously updates the training distribution. Empirical results on six public datasets show that EvoOptiGraph significantly outperforms larger generalist models, agentic methods, and specialized baselines in accuracy, executability, and generalization. These results demonstrate that targeted data-model coevolution is an effective strategy for improving LLMs on optimization modeling tasks.

13:00 JST研究/論文

AI 生成の望遠鏡スケジュール決定のためのマルチレベル検証およびトレーサビリティ フレームワーク

望遠鏡のスケジューリングに AI が段階的に導入されることで、複雑な複数の制約問題を処理する際に AI ベースの意思決定が利点を示すようになりました。ただし、その出力には一貫性のないデータ参照、推論エラー、実行不可能な決定が含まれることが多く、信頼性の高い観察タスクへの適用が制限されます。この研究では、実行前に AI が生成した意思決定の体系的な信頼性検証を実行し、追跡可能な意思決定をサポートする推論プロセスの明示的な表現を可能にする、マルチレベル検証および追跡可能な推論フレームワークを提案します。このフレームワークは、データ参照の検証、論理的一貫性チェック、観測的および機器的制約検証を統合して、無効な決定をフィルタリングして修正します。また、アトミック推論ユニットとその依存関係も導入し、エラーの位置特定と事後分析をサポートする相互接続された推論ステップのシーケンスとしてスケジューリング決定を表します。実験では、このフレームワークにより AI スケジューリングの実行可能性と信頼性が向上し、一時的な機会の損失が軽減されることが示されています。特に、フィードバックの修正と推論ステップの構造化された検証により、特に複雑なシナリオで誤った決定を修復およびブロックする能力が強化されます。純粋な AI 手法と比較して、フレームワークで強化されたアプローチは柔軟性を維持しながら、信頼性と実行可能性を大幅に向上させます。これらの結果は、AI を高信頼性の天体観測スケジュールに適用するための実現可能かつ検証可能な経路を示しています。

原文 (English)

A Multi-Level Validation and Traceability Framework for AI-Generated Telescope Scheduling Decisions

With the gradual introduction of AI into telescope scheduling, AI-based decision-making has shown advantages in handling complex multi-constraint problems. However, its outputs often suffer from inconsistent data references, reasoning errors, and non-executable decisions, limiting applicability in high-reliability observational tasks. In this work, we propose a multi-level validation and traceable reasoning framework that performs systematic reliability verification of AI-generated decisions prior to execution, and enables explicit representation of the reasoning process to support traceable decision-making. The framework integrates data reference validation, logical consistency checks, and observational and instrumental constraint verification to filter and correct invalid decisions. It also introduces atomic reasoning units and their dependency relationships, representing scheduling decisions as a sequence of interconnected reasoning steps that support error localization and post hoc analysis. Experiments show that the framework improves executability and reliability of AI scheduling and reduces loss of transient opportunities. In particular, feedback correction and structured validation of reasoning steps enhance the ability to repair and block erroneous decisions, especially in complex scenarios. Compared with pure AI methods, the framework-enhanced approach maintains flexibility while substantially improving reliability and executability. These results demonstrate a feasible and verifiable pathway for applying AI to high-reliability astronomical observation scheduling.

13:00 JST研究/論文

大規模な言語モデルを使用したコンテンツベースのスマート電子メール ディスパッチャー

電子メールによるコミュニケーションは私生活や仕事において不可欠な部分となっていますが、その膨大な量を処理することは依然として大規模な組織にとって重要な問題です。他のインスタント メッセージング プラットフォームを使用して電子メールを手動で閲覧し、その内容と添付ファイルを目的の受信者に転送すると、エラーが発生しやすく時間がかかり、生産性の低下や過度のストレスにつながることが判明しています。このペーパーの主な目的は、工学系大学のプログラムのさまざまな学期の学生のそれぞれの WhatsApp グループにメールの内容に基づいて電子メールを送信するタスクを自動化し、組織内の一方の端からもう一方の端への情報の流れをスムーズにする代替メカニズムを探ることです。ディスパッチャ システムは、大規模言語モデル (LLM) をクエリするエージェントを使用して構築されており、電子メールの内容を分析し、関連する学生グループに電子メールをルーティングして情報を提供し、利用することができます。このシステムは、意思決定のためにテキストの内容を分析する際に LLM の機能を利用します。電子メールのコンテンツを入力として指示とコンテキストとともに含む、適切に構造化されたエージェント フレームワーク プロンプトを使用すると、システムは電子メール メッセージの送信先となる関連グループを特定し、必要な情報を時間どおりに提供します。提案されたシステムは、ラベル付きデータセットに依存せず、生産性の向上や電子メールを読むことに伴う認知負荷の軽減など、いくつかの利点を提供します。

原文 (English)

Content-Based Smart E-Mail Dispatcher Using Large Language Models

Email communication has become an integral part of personal and professional life, but handling its vast volume is still a significant issue for large organisations. Manual perusal of emails and forwarding their contents and attachments to intended recipients using other instant messaging platforms has proved to be error-prone and time-consuming leading to losses in terms of productivity and creating undue stress. The main objective of this paper is to explore an alternative mechanism that is to automate the task of dispatching emails based on their contents to the respective WhatsApp groups of students of various semesters of programs in an engineering college, facilitating a smooth flow of information from one end to another end in an organisation. The dispatcher system is built using agents querying large language models (LLMs) to enable it to analyze the contents of emails and route them to the relevant groups of students for their information and consumption. The system harnesses the capabilities of LLMs in analysing the textual contents for decision-making. With a well-structured agent framework prompt that includes email content as input with instructions and context, the system figures out the relevant groups to which the email message is dispatched, thus providing the required information on time. The proposed system does not rely on labelled datasets and provides several benefits, including enhanced productivity and a reduction in the cognitive load associated with reading emails.

13:00 JSTLLM/生成AI

サービスフィードバックにおける新たなトピックを検出するための LLM ベースのモデル

サービスフィードバックの分析を強化することは、信頼とコンプライアンスが公正かつ効果的なサービスの提供に依存する公共部門の組織、特に税務当局にとって不可欠です。フィードバックの量が増加するにつれて、新たなサービス品質の問題と、多様な集団間の潜在的な格差を特定することがますます困難になっています。従来のアプローチは、手動レビューや専門家が定義した静的な指標に依存することが多く、スケーラビリティやテキストフィードバックで複雑なパターンを捕捉する機能が制限されていました。この論文では、大規模言語モデル (LLM)、統計手法、および人間と AI のコラボレーションを統合して、多言語の顧客フィードバック分析を改善する新しい方法論を紹介します。主な目的は、サービス提供における潜在的な不公平性を明らかにする可能性がある、新たなサービス品質トピックを検出することです。当社のフレームワークは、微調整され量子化された LLM と専門家の監視を組み合わせて、正確で計算効率が高く、コンテキストを認識した分析を生成します。提案されたアプローチは、類似性分析と経験豊富な税務職員による評価を使用して評価され、ベースライン モデルよりも専門家の判断との一致が強いことが実証されました。この方法論では、人間参加型フレームワークを組み込むことで、生成された洞察の信頼性と関連性を向上させながら、LLM の作成を削減します。この結果は、LLM と人間の専門知識を組み合わせて、公共部門の組織における拡張性のある証拠に基づいた意思決定をサポートすることの実用性を示しています。この取り組みは、多言語の顧客フィードバックのより効果的な分析を通じて、サービスの品質、応答性、公平性、社会的信頼を向上させる、責任ある AI システムの開発に貢献します。

原文 (English)

LLM-based Models for Detecting Emerging Topics in Service Feedback

Enhancing the analysis of service feedback is essential for public sector organizations, particularly tax administrations, where trust and compliance depend on fair and effective service delivery. As feedback volumes grow, identifying emerging service quality issues and potential disparities across diverse populations becomes increasingly challenging. Traditional approaches often rely on manual review or static expert-defined indicators, limiting scalability and the ability to capture complex patterns in textual feedback. This paper presents a novel methodology that integrates large language models (LLMs), statistical techniques, and human-AI collaboration to improve multilingual customer feedback analysis. The primary objective is to detect emerging service quality topics that may also reveal potential inequities in service delivery. Our framework combines fine-tuned, quantized LLMs with expert oversight to produce accurate, computationally efficient, and context-aware analyses. The proposed approach was evaluated using similarity analysis and assessments from experienced tax officers, demonstrating stronger alignment with expert judgments than baseline models. By incorporating a human-in-the-loop framework, the methodology reduces LLM fabrication while improving the reliability and relevance of generated insights. The results demonstrate the practicality of combining LLMs with human expertise to support scalable, evidence-based decision-making in public sector organizations. This work contributes to the development of responsible AI systems that enhance service quality, responsiveness, fairness, and public trust through more effective analysis of multilingual customer feedback.

13:00 JSTエージェント

エージェントの指示をコードとしてのポリシーに自動形式化

一か八かの領域におけるエージェントの安全性には、正式なポリシーの適用が必要ですが、既存のアプローチのほとんどは、正式な保証を提供しない確率的なガードレール (微調整された分類器、プロンプトベースのステアリング) に依存しているか、実際のポリシー仕様の幅広さに対応していない手作業でコード化されたシンボリックな適用に依存しています。 LLM ベースのジェネレーター - クリティカル ループを使用して、エージェント プロンプト、MCP ツールの説明、および自然言語ポリシー ドキュメントを正式に検証されたポリシーに変換する自動形式化パイプラインを紹介します。結果として得られるポリシーは Cedar ポリシー言語で記述されます。 MedAgentBench ベンチマークでは、当社の自動形式化されたポリシーは、以前の作業で手作業でコーディングされたシンボリック強制よりも大幅に多くのソース自然言語仕様をカバーします。

原文 (English)

Autoformalization of Agent Instructions into Policy-as-Code

Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement that does not scale to the breadth of real policy specifications. We present an autoformalization pipeline that translates agent prompts, MCP tool descriptions, and natural language policy documents into formally verified policies using an LLM-based generator-critic loop. The resulting policies are written in the Cedar Policy Language. On the MedAgentBench benchmark, our autoformalized policies cover substantially more of the source natural-language specification than the hand-coded symbolic enforcement in prior work.

13:00 JSTエージェント

SKILL-DISCO: エージェント トレースを抽出して再利用可能なプロシージャル スキルにコンパイルする

多くの場合、エージェントは同様のタスク インスタンスを繰り返し最初から解決するため、不必要な推論コストと長い実行トレースが発生します。これまでの研究では、ワークフローの再利用と実行可能なスキルの導入が検討されてきましたが、どのタスク シナリオが手続き型スキルを許可するのか、また、成功したトレース全体で共有される手続き型構造をどのように表現する必要があるのか​​は不明のままです。私たちは、この問題を FSM で定義されたシナリオで研究します。このシナリオでは、成功したトレースが未知の遷移グラフ内のパスとして表示され、再利用可能なパラメーター化された制御フロー サブグラフとして手続き型スキルが定式化されます。この見解に基づいて、成功したトレースから再利用可能な PFSM サブグラフを抽出し、それらを呼び出し可能、実行可能、検証可能な手続き型スキルにコンパイルする抽出およびコンパイル フレームワークである SkillDisCo を紹介します。 ALFWorld と WebArena での実験では、SkillDisCo がベンチマークとモデル スケール全体で成功率を向上させ、エージェントのターン数を削減することが示されており、共有エクスペリエンスを再利用可能な実行構造として表現することの利点が実証されています。

原文 (English)

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explored workflow reuse and executable skill induction, but it remains unclear which task scenarios admit procedural skills and how the shared procedural structure should be represented across successful traces. We study this problem in FSM-defined scenarios, where successful traces can be viewed as paths in an unknown transition graph, and formulate procedural skills as reusable parameterized control-flow subgraphs. Based on this view, we introduce SkillDisCo, a distillation-and-compilation framework that distills reusable PFSM subgraphs from successful traces and compiles them into callable, executable, and verifiable procedural skills. Experiments on ALFWorld and WebArena show that SkillDisCo improves success rates and reduces agent turns across benchmarks and model scales, demonstrating the benefits of representing shared experience as reusable execution structures.

13:00 JST研究/論文

NebulaExp-8B: 本格的なアブレーション研究による経験的なポストトレーニング パイプライン

トレーニング後の調整により、大規模な言語モデルの機能に従う推論と人間の好みが決まりますが、既存の研究のほとんどは詳細なデータ構築、フィルタリング ルール、トレーニング レシピを差し控えており、コミュニティの再現性と軽量モデルの最適化を妨げています。この研究では、Qwen3-8B ベースに構築された完全に透明なアブレーション駆動のポストトレーニング パイプラインである NebulaExp を紹介します。これは、一般的な命令モデルと複雑な推論に特化したモデルという 2 つの直交するモデル ブランチをカバーします。私たちは、384 万のマルチソース SFT サンプルの生のコーパスと 20 万の検証可能な RL 候補プールを厳選し、応答蒸留、多次元相互検証フィルタリング、きめ細かい難易度グレーディング、タスク分類、多様性を意識したサンプリングを含むエンドツーエンドのデータ処理スタックを設計します。 Instruct ブランチの場合、3 段階で最適化された教師あり微調整アプローチ NebulaExp-Ins-SFT により、平均ベンチマーク スコアが Qwen3-8B-nothink のベースライン 55.01 から 60.99 に向上しました。 GRPO 強化学習により、平均スコアはさらに 61.85 まで上昇します。推論ブランチでは、中程度の難易度の GRPO RL により、平均推論スコアが 73.88 から 75.17 に向上しました。 RL のタスク検証者への依存に対処するために、私たちは単一教師と複数教師の OPD (MOPD) を系統的に調査しました。4K の命令に従うサンプルのみを利用し、IFEval で RL ベースラインを 3.26 ポイント上回り、平均全体ゲイン +4.43 でした。 MOPD は、4 人のドメイン専門教師とわずか 10,000 サンプルを融合し、基本モデルと比較して平均パフォーマンスを 4.18 向上させます。このレポートは、8B スケール LLM の完全に再現可能な経験的なトレーニング後のレシピを提供し、命令遵守、数学的推論、コード生成、および一般知識の間の機能のトレードオフを包括的に分析します。

原文 (English)

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization. This work presents NebulaExp, a fully transparent, ablation-driven post-training pipeline built on Qwen3-8B-base, covering two orthogonal model branches: general instruct model and complex reasoning-specialized model. We curate a raw corpus of 3.84M multi-source SFT samples and a 200K verifiable RL candidate pool, and design an end-to-end data processing stack including response distillation, multi-dimensional cross-verification filtering, fine-grained difficulty grading, task classification and diversity-aware sampling. For the Instruct branch, our three-stage optimized supervised fine-tuning approach NebulaExp-Ins-SFT improves the average benchmark score from the 55.01 baseline of Qwen3-8B-nothink to 60.99. GRPO reinforcement learning then further elevates the average score to 61.85. For the Reasoning branch, medium-difficulty GRPO RL improves average reasoning score from 73.88 to 75.17. To address RL's dependency on task verifiers, we systematically investigate single-teacher and multi-teacher OPD (MOPD): utilizing merely 4K instruction-following samples and outperforms RL baseline by 3.26 points on IFEval with +4.43 average overall gain; MOPD fuses four domain-specialist teachers with merely 10K samples, lifting average performance by 4.18 over the base model. This report provides a fully reproducible empirical post-training recipe for 8B-scale LLMs, and comprehensively dissects the capability trade-offs among instruction adherence, mathematical reasoning, code generation and general knowledge.

13:00 JSTLLM/生成AI

安全ガードレールには理由が必要ですか? LeanGuard: 堅牢なモデレーションのための高速かつ軽量なアプローチ

プロンプトまたは応答をスクリーニングするために、最近のガードレール メソッドは、判定を発行する前に思考連鎖 (CoT) を生成します。この設計は、段階的に推論することで意思決定が改善されるという一般的な信念に従っています。ただし、CoT は、モデルが決定する前に多くのトークンを生成する必要があるため、ガードを重く遅くします。これは、ガードレールが実際に展開される方法と一致しない可能性があります。ガードレールは重くて遅いものであってはならず、多くの場合、実体化されたロボットなどのデバイス上で実行されます。この論文では、安全ガードレールに本当に理由を付ける必要があるかどうかという疑問を投げかけます。この質問に答えるために、同じコーパス上で軽量の双方向エンコーダと推論ガードをトレーニングし、他のすべてを固定したまま推論のみを削除します。この制御された同一塩基比較により、チェーンがモデレーション精度を向上させないことがわかります。結果として得られるガードを LeanGuard と名付けます。 395M ラベル専用エンコーダは、公開ベンチマークを上回る 82.90 $\pm$ 0.26 の平均 F1 に達します。これは、はるかに大規模なデコーダ上に構築された推論ガードと一致しますが、最大 512 個のトークンの入力に対して単一の前方パスのみを使用します。これは、推論コンピューティングにおける約 100 分の 1 の削減に相当します。さらに、このラベルのみのエンコーダーはトレーニング ラベル ノイズの下でも堅牢であり、厳密な偽陽性率で推論ガードよりもはるかに多くの再現率を保持するため、より重い推論ガードがより堅牢な選択肢であるわけではないことも示します。私たちの調査結果は、現在のガードレール ベンチマークは推論に報いるほど難しくない可能性があり、モデレーションのための CoT の必要性がまだ証明されていないことを示唆しています。 LeanGuard を含むすべてのソース コードとモデルは https://github.com/ndb796/LeanGuard でリリースされます。

原文 (English)

Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation

In order to screen a prompt or a response, the recent guardrail methods generate a chain-of-thought (CoT) before they issue a verdict. This design follows a common belief that step-by-step reasoning improves a decision. However, CoT also makes the guard heavy and slow, because the model must generate many tokens before it decides. This may not match how guardrails are actually deployed. A guardrail sometimes should not be heavy and slow, and it often runs on-device, for example on an embodied robot. In this paper, we pose a question whether a safety guardrail really needs to reason. To answer this question, we train a lightweight bidirectional encoder and a reasoning guard on the same corpus, and we then remove only the reasoning while we keep everything else fixed. With this controlled same-base comparison, we show that the chain does not improve moderation accuracy. We name the resulting guard LeanGuard. A 395M label-only encoder reaches an average F1 of 82.90 $\pm$ 0.26 over public benchmarks. It matches a reasoning guard that is built on a much larger decoder, while it uses only a single forward pass over an input of at most 512 tokens. This is about a ~100x reduction in inference compute. We further show that this label-only encoder stays robust under training-label noise and retains far more recall at a strict false-positive rate than the reasoning guard, so a heavier reasoning guard is not the more robust choice either. Our finding suggests that the current guardrail benchmarks may not be hard enough to reward reasoning, and that the necessity of CoT for moderation is still not proven. We release all source codes and models including LeanGuard at https://github.com/ndb796/LeanGuard.

13:00 JST研究/論文

コンバインドサイクルガスタービンの少数ショット故障検出のためのカルマンプロトタイプネットワーク

コンバインド サイクル ガス タービン (CCGT) は、現代の発電において重要な役割を果たしており、高効率と環境への影響の低減の両方を実現します。ただし、複雑な熱流体と機械的相互作用により、特にラベル付きの故障データが不足している場合、故障検出が複雑になります。このペーパーでは、特に CCGT 障害診断用に調整されたメトリクスベースの少数ショット学習 (FSL) フレームワークであるカルマン プロトティピカル ネットワーク (KPN) を紹介します。クラスプロトタイプの進化を動的システムの潜在的な確率状態としてモデル化し、一時的な分散を削減し、埋め込み表現のロバスト性を向上させます。オフショア CCGT システムの高忠実度 Modelica ベースの動的シミュレーションで生成された合成データ セットが使用され、通常動作と過渡条件下での進行性漏洩故障の両方をシミュレートしました。提案されたフレームワークをシミュレートされたリーク障害検出タスクに適用すると、KPN は、さまざまなサポートとクエリ構成の下で、精度と安定性の両方において、マッチング ネットワーク、関係ネットワーク、MAML などの従来の FSL 手法よりも優れていることが実証されています。提案されたフレームワークは、クラス表現を安定させることでトレーニングの収束と一般化を大幅に改善し、ラベル付きデータが制限されている現実世界の CCGT 障害検出に適しています。

原文 (English)

Kalman Prototypical Networks for Few-shot Fault Detection in Combined Cycle Gas Turbines

Combined-cycle gas turbines (CCGTs) play a key role in modern power generation, offering both high efficiency and reduced environmental impact. However, their complex thermo-fluid and mechanical interactions complicate fault detection, particularly when labeled fault data are scarce. In this paper, we introduce the Kalman Prototypical Network (KPN), a metric-based few-shot learning (FSL) framework specifically tailored for CCGT fault diagnosis. We model the evolution of class prototypes as latent stochastic states in a dynamic system to reduce episodic variance and improve robustness in embedding representation. Synthetic data sets generated with a high-fidelity Modelica-based dynamic simulation of an offshore CCGT system were used, simulating both normal operation and progressive leak faults under transient conditions. Application of the proposed framework on simulated leak fault detection tasks demonstrate that KPN outperforms conventional FSL methods such as Matching Networks, Relation Networks, and MAML in both accuracy and stability under varying support and query configurations. The proposed framework significantly improves training convergence and generalization by stabilizing class representations, making it well-suited for real-world CCGT fault detection where labeled data is limited.

13:00 JST研究/論文

LithoDreamer: マルチステージ計算リソグラフィーのための物理学に基づいた世界モデル

半導体テクノロジーのノードが拡大するにつれて、歩留まりとパフォーマンスを確保するにはコンピュテーショナル リソグラフィーが不可欠です。ただし、リソグラフィーは、マスクの最適化、光学イメージング、レジスト露光、現像を含む連続的な物理プロセスであり、既存のモデルではこれらを捉えることができません。この制限を克服するために、我々は、「レイアウト-マスク-レジスト画像-現像後画像(ADI)」パイプラインを意思決定駆動型の多段階進化システムとして定式化する、計算リソグラフィーのための最初の物理情報に基づいたワールドモデル(WM)フレームワークであるLithoDreamerを紹介します。 LithoDreamer は、隣接する状態間の特徴の変化をキャプチャして、ステージ固有の物理情報に基づいた潜在空間をモデル化し、そこでプロセス介入の探索を制御し、その後の状態遷移を駆動します。継続的な監視なしで解釈可能な介入最適化を達成するために、介入パス間の潜在的な差異を変分進化制約で対比し、実際のリソグラフィ物理学と一致する進化を生成するようにモデルを導く対照変分最適化パラダイムを提案します。実験では、LithoDreamer が順進化と逆計画において最先端のパフォーマンスを達成することを示しています。私たちのリソグラフィ データセットは、GitHub (https://github.com/7jiangyq/lithodreamer.git) で公開されています。

原文 (English)

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography

As semiconductor technology nodes scale, computational lithography is essential for ensuring yield and performance. However, lithography is a continuous physical process involving mask optimization, optical imaging, resist exposure, and development, which existing models fail to capture. To overcome this limitation, we present LithoDreamer, the first physics-informed World Model (WM) framework for computational lithography, which formulates the ``Layout-Mask-Resist Image-After Development Image (ADI)'' pipeline as a decision-driven multi-step evolution system. LithoDreamer captures feature changes between adjacent states to model stage-specific physics-informed latent spaces, in which it controls process intervention exploration and drives subsequent state transitions. To achieve interpretable intervention optimization without continuous supervision, we propose a contrastive variational optimization paradigm that contrasts the latent differences between intervention paths with variational evolution constraints, guiding the model to generate evolutions consistent with real lithography physics. Experiments show LithoDreamer achieves state-of-the-art performance in forward evolution and inverse planning. Our lithography dataset is publicly available at GitHub (https://github.com/7jiangyq/lithodreamer.git).

13:00 JST画像/動画生成

シネ心臓 MRI の時空間モデリングに対する潜在的な ODE アプローチ

心臓磁気共鳴画像法 (CMR) は、心室の構造と運動に関する豊富な時空間情報を取得しますが、従来のリスク モデルでは、選択された心臓位相から画像から導出された指標が少数しか使用されていません。心拍数を意識した神経常微分方程式(ODE)ダイナミクスとグラフベースのメッシュオートエンコーダーを使用して、解剖学的に一貫した3D + t心室運動を再構築し、両心室の解剖学的構造とフルサイクルのシネ運動を連続的な潜在軌道としてエンコードする潜在力学モデルを提示します。共変量条件付き事前分布は、予想される拡張末期の潜在状態を定義し、コックス比例ハザード モデルは、この事前分布からの逸脱が心不全の発生を予測するかどうかをテストします。私たちは、367件の心不全事象を含む、ベースラインの心血管疾患のない英国バイオバンク参加者72,386人を調査しました。保留された評価サブセットでは、再適合されたプールされたコホート方程式に潜在スコアを追加すると、7 つの確立された心臓マーカーの 0.764 と比較して、層化 C インデックスが 0.704 から 0.785 に改善されました。非グラフおよび非 ODE アプローチと比較して、提案されたモデルは、再構成の忠実度、生成的リアリズム、および下流の予測パフォーマンスの間で最良のトレードオフを示しました。これらの結果は、心室運動の連続的な全周期モデリングが従来のCMR要約を超えた有益な心臓表現型を提供する一方、臨床リスク予測の使用前には、より代表的な患者コホートにおける外部検証が必要であることを示唆している。

原文 (English)

A Latent ODE Approach to Spatiotemporal Modeling of Cine Cardiac MRI

Cardiac magnetic resonance imaging (CMR) captures rich spatiotemporal information about ventricular structure and motion, but conventional risk models use only a few image-derived indices from selected cardiac phases. We present a latent dynamical model that encodes bi-ventricular anatomy and full-cycle cine motion as a continuous latent trajectory, using heart-rate-aware neural ordinary differential equation (ODE) dynamics and a graph-based mesh autoencoder to reconstruct anatomically consistent 3D+t ventricular motion. A covariate-conditioned prior defines the expected end-diastolic latent state, and a Cox proportional hazards model tests whether deviations from this prior predict incident heart failure. We studied 72,386 UK Biobank participants without baseline cardiovascular disease, including 367 incident heart failure events. In a held-out evaluation subset, adding the latent score to refitted pooled cohort equations improved the stratified C-index from 0.704 to 0.785, compared with 0.764 for seven established cardiac markers. Compared with non-graph and non-ODE approaches, the proposed model gave the best trade-off between reconstruction fidelity, generative realism, and downstream prognostic performance. These results suggest that continuous full-cycle modeling of ventricular motion provides informative cardiac phenotypes beyond conventional CMR summaries, while external validation in more representative patient cohorts is required before clinical risk-prediction use.

13:00 JSTエージェント

高次元物理システムにおける自律的な科学発見のためのソクラテスエージェント

科学的発見の自動化は転換点に達しています。 AI システムは現在、機器を操作し、パラメーターを最適化し、仮説を生成しますが、ほとんどは依然として手続き型であり、人間の設計者によって修正されたワークフローを実行します。真の自律科学は認識論的自律性、つまり証拠に応じて物理的説明を構築し、異議を唱え、修正する能力を要求します。ここでは、ソクラテス助産術をクローズドループ実験に組み込むマルチエージェント AI 科学者である AHOIS を紹介します。物理学批判エージェントは、因果関係の質問、制約チェック、反例の生成、および反証基準の定式化を通じて仮説を調査します。私たちは、実際のマルチモードファイバー光学プラットフォーム、複雑な波動変換、間接検出、環境ドリフト、マルチモーダル取得を備えた高次元システム上で AHOIS を評価します。事前の符号化スキーム、分類器、またはスペックル モデルなしで、このシステムは自律的にランダム干渉符号化仮説を提案および検証し、タスク適応型スパース測定戦略を発見し、明確な故障モード (符号化の不安定性、蛍光汚染、検出器ノイズ) を診断し、公開されたイメージング プロトコルをオリジナル以外の構成で実行可能なワークフローに変換しました。発見されたエンコーディングにより、有効ランク 56.9 の 16x16 測定値が得られ、分類精度は MNIST で 76.97%、Fashion-MNIST で 83.17% でした。アブレーションは、ソクラテス的尋問が物理的な一貫性、仮説の完全性、不確実性の校正、および実験計画の妥当性を向上させることを示しています。これらの結果は、ワークフローの自動化から、複雑な物理環境における証拠に基づく自己修正型の自律的な発見への道筋を確立します。

原文 (English)

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and generate hypotheses, most remain procedural: they execute workflows fixed by human designers. True autonomous science demands epistemic autonomy--the capacity to construct, challenge and revise physical explanations in response to evidence. Here we introduce AHOIS, a multi-agent AI scientist that embeds Socratic midwifery into closed-loop experimentation. A physics-critic agent interrogates hypotheses through causal questioning, constraint checking, counterexample generation and falsification-criteria formulation. We evaluate AHOIS on a real multimode-fibre optical platform, a high-dimensional system with complex wave transformations, indirect detection, environmental drift and multi-modal acquisition. Without prior encoding schemes, classifiers or speckle models, the system autonomously proposed and validated a random-interference encoding hypothesis, discovered task-adaptive sparse-measurement strategies, diagnosed distinct failure modes (encoding instability, fluorescence contamination and detector noise) and translated a published imaging protocol into an executable workflow on a non-original configuration. The discovered encoding yielded 16x16 measurements with effective rank 56.9 and classification accuracies of 76.97% on MNIST and 83.17% on Fashion-MNIST. Ablations show that Socratic interrogation improves physical consistency, hypothesis completeness, uncertainty calibration and experimental-plan validity. These results establish a route from workflow automation towards evidence-grounded, self-correcting autonomous discovery in complex physical environments.

13:00 JST研究/論文

メタ最適化としての科学的発見: 組み合わせ最適化のケーススタディ

科学的発見は基本的に最適化問題であり、理論と実験の広大な「状態空間」と、品質、新規性、有効性に基づく評価基準によって定義されます。大規模言語モデル (LLM) により、この領域の自動探索が可能になりましたが、評価基準を同時に変更することも同様に重要であると私たちは主張します。ここでは、研究をメタ最適化として形式化することを提案します。この場合、最適化の目的自体も最適化されます。私たちの主な貢献は「コンセンサス目的集計」です。LLM で生成された目的関数が相関加重投票によって結合され、理解が深まるにつれて進化する安定した自己修正評価基準が生成されます。このフレームワークをデジタル MemComputing マシンに基づく 3-SAT 問題のアルゴリズム検出に適用し、問題サイズ $N$ のベースライン スケーリングを $\sim N^{2.51}$ から $\sim N^{1.33}$ に削減し、テストされた最大のインスタンスで $\sim 67\times$ の高速化を実現します。問題にとらわれないフレームワークとして、このアプローチが科学的発見に大きく役立つことを期待しています。

原文 (English)

Scientific discovery as meta-optimization: a combinatorial optimization case study

Scientific discovery is fundamentally an optimization problem, defined by a vast "state space" of theories and experiments, and an evaluation criterion based on quality, novelty, and validity. Large language models (LLMs) have enabled automated exploration of this space, but we argue that simultaneous modification of the evaluation criteria is equally important. Here, we propose formalizing research as meta-optimization, where the optimization objective itself is also being optimized. Our key contribution is "consensus objective aggregation," where LLM-generated objective functions are combined via correlation-weighted voting, yielding a stable, self-correcting evaluation criterion that evolves as understanding deepens. We apply this framework to algorithm discovery for 3-SAT problems based on digital MemComputing machines, reducing the baseline scaling with problem size $N$ from $\sim N^{2.51}$ to $\sim N^{1.33}$ and delivering a $\sim 67\times$ speedup on the largest instances tested. As a problem-agnostic framework, we hope this approach will considerably aid scientific discovery.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体Claude

EGG: 専門家によるカーネル生成のためのエージェント フレームワーク

高性能 GPU カーネルは、大規模言語モデル (LLM) の指数関数的に増大する計算コストを削減するために不可欠ですが、その開発はドメイン専門家による手動チューニングに大きく依存しています。 LLM ベースのアプローチの最近の進歩は、カーネル生成の自動化に有望であることを示していますが、正確さと高いパフォーマンスの両方を達成するのにまだ苦労しています。この制限は主に、ドメイン固有の最適化ガイダンスの欠如によって生じ、最適化空間の効果的な探索が妨げられます。私たちは、LLM の意思決定をガイドする専門家の最適化原則を組み込んだ、カーネル生成のための専門家ガイド付きエージェント フレームワークである EGG を提案します。専門家のワークフローからインスピレーションを得て、私たちはカーネル生成を 2 つの階層段階に分解します。1) 高品質の計算構造基盤を確立するアルゴリズム構造設計。 2) ハードウェア固有のチューニング。並列マッピング、テンソル タイリング、メモリ最適化を通じてターゲットを絞った調整を実行します。この段階的な分解により、明示的な最適化目標が定義され、段階的な改良を達成するために設計空間が構築されます。この目的を達成するために、ステージを意識したマルチエージェント コラボレーション メカニズムがステージ間およびステージ内のコンテキスト管理用に設計されており、安定した最適化軌道を保証します。 KernelBench と実際のワークロードの実験では、EGG が PyTorch と比較して平均 2.13 倍の高速化を達成し、既存のエージェント ベースおよび RL ベースのアプローチを上回るパフォーマンスを示していることが示されています。

原文 (English)

EGG: An Expert-Guided Agent Framework for Kernel Generation

High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in LLM-based approaches show promise for automating kernel generation, they still struggle to achieve both correctness and high performance. This limitation primarily arises from the lack of domain-specific optimization guidance, hindering effective exploration of the optimization space. We propose EGG, an Expert-Guided Agent Framework for Kernel Generation, which incorporates expert optimization principles to guide LLMs' decisions. Inspired by expert workflows, we decompose kernel generation into two hierarchical stages: 1) algorithmic structure design, which establishes a high-quality computational structure foundation; 2) hardware-specific tuning, which performs targeted adjustments through parallel mapping, tensor tiling, and memory optimization. This staged decomposition defines explicit optimization objectives, structuring the design space to achieve progressive refinement. To this end, a stage-aware multi-agent collaboration mechanism is designed for inter and intra-stage context management, ensuring stable optimization trajectories. Experiments on KernelBench and real-world workloads show that EGG achieves a 2.13x average speedup over PyTorch, outperforming existing agent-based and RL-based approaches.

13:00 JST画像/動画生成

ResilPhase: 拡散加速のためのプラグアンドプレイ位相マッピングとノイズ耐性のあるマクロ軌道外挿

強力な拡散モデルの採用は、その大幅な推論遅延によって妨げられています。最近の「キャッシュしてから予測」スキームは、導関数ベースの多項式を使用して DiT を高速化することでこの問題を軽減しますが、高い加速率では重大な品質劣化が発生します。私たちの分析により、その根本原因が明らかになりました。それは、連続拡散軌道とずれていて数値的に不安定な表現に対して実行された離散外挿です。したがって、加速された DiT は、蓄積された空間エラー、ノイズの多い微分増幅、および高次の不安定性の影響を受けます。したがって、加速推論を常微分方程式 (ODE) 空間における安定したマクロ軌道外挿として再定式化します。中間の特徴を予測する代わりに、モデルのグローバル ドリフト (GD)、つまりエンドツーエンドの状態の進化に合わせて予測を行うことで、特徴の不一致とメモリのオーバーヘッドを排除します。しかし、この滑らかなマクロ軌道でさえも、微分の誤謬に対して脆弱なままです。つまり、その高次の時間微分は本質的にノイズが多いのです。したがって、導関数のない重心ラグランジュ外挿法を導入して、導関数の不安定性と近似誤差を効果的に回避します。さらに、外挿領域を正規化し、振動誤差の増大を抑制する、有界位相マッピングを提案します。これらの要素は集合的に、ノイズ耐性のあるアクセラレーション フレームワークである ResilPhase を構成します。 FLUX.1-dev と HunyuanVideo の実験では、積極的な加速比の下で最先端の忠実度が実証されています。

原文 (English)

ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration

The adoption of powerful diffusion models is hindered by their significant inference latency. Recent ``cache-then-forecast'' schemes alleviate this issue by accelerating DiTs using derivative-based polynomials, but they suffer from severe quality degradation at high acceleration ratios. Our analysis reveals its root cause: the discrete extrapolation performed on representations that are misaligned with the continuous diffusion trajectory and are numerically unstable. Thus, accelerated DiTs suffer from accumulated spatial errors, noisy derivative amplification, and high-order instability. We therefore reformulate accelerated inference as stable macro-trajectory extrapolation in ordinary differential equation (ODE) space. Instead of predicting intermediate features, we align forecasting with the model's Global Drift (GD), i.e., the end-to-end state evolution, thereby eliminating feature inconsistency and memory overhead. However, even this smooth macro-trajectory remains vulnerable to the derivative fallacy: its higher-order temporal derivatives are intrinsically noisy. Thus, we introduce a derivative-free barycentric Lagrange extrapolator to effectively bypass derivative instability and approximation error. We further propose a bounded Phase Mapping that regularizes the extrapolation domain, suppressing oscillatory error growth. These elements collectively constitute ResilPhase, a noise-resilient acceleration framework. Experiments on FLUX.1-dev and HunyuanVideo demonstrate state-of-the-art fidelity under aggressive acceleration ratios.

13:00 JSTエージェントGPT / ChatGPTMistral AI

メモリアクセスではなくメモリ深さ: 長時間実行される言語エージェントのための選択的なパラメトリック統合

長時間実行される言語エージェントには、メモリ アクセス以上のものが必要です。検索システムはクエリ時に過去のファクトをフェッチできますが、作業コンテキストがアンロードされた後、どのエクスペリエンスが引き続き動作を形成するかを決定しません。私たちは、この別の問題をメモリの深さとして研究します。つまり、小さなパラメトリック ストアに書き込まれる耐久性のある目標条件付きの傾向です。ループドリフトプロトコルを導入します。これは、作業コンテキストがアンロードされている間、検索インデックスがそのまま残り、長いループ干渉下でも目標条件付き動作が持続する必要がある制御されたストレステストです。我々は、サプライズおよび価数ゲート型 LoRA 統合メカニズムである EVAF を評価します。 GPT-2 と TinyLlama 全体で、検索は浅い事実の想起 (短い事実の精度 0.956 ~ 0.973) で最も強力ですが、EVAF は目標の永続性とアンロード後の回復 (0.812 ~ 0.904) で最も強く、200 イベントあたりわずか 2 ~ 3 回のパラメトリック書き込みです。メカニズム制御は、選択的統合が選択と作動という 2 つの制御可能な次元に分解されることを示しています。一致したランダム ゲートは、スパース書き込みを超えて選択を分離します。 GPT-2、TinyLlama、Mistral-7B にわたる固定内部制御は、内部ループの書き込み強度がモデルに依存していることを示しています。そして、Mistral-7B のマッチドゲート反転により、誤って調整された作動下での非対称な選択と作動の結合が明らかになります。 Public Memora イベント ストリームは外部診断として機能し、古いメモリの無効化を未解決の境界として明らかにします。このプローブ内では、選択的なパラメトリック統合により、検索アクセスとは異なる、検索アクセスを補完するメモリ深さが提供されます。

原文 (English)

Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents

Long-running language agents need more than memory access. Retrieval systems can fetch past facts at query time, but they do not decide which experiences should continue to shape behavior after the working context is unloaded. We study this separate problem as memory depth: durable goal-conditioned tendencies written into a small parametric store. We introduce the loop-drift protocol, a controlled stress test in which the retrieval index remains intact while working context is unloaded and goal-conditioned behavior must persist under long-loop interference. We evaluate EVAF, a surprise- and valence-gated LoRA consolidation mechanism. Across GPT-2 and TinyLlama, retrieval is strongest on shallow factual recall (short-fact accuracy 0.956--0.973), while EVAF is strongest on goal persistence and post-unload recovery (0.812--0.904) with only 2--3 parametric writes per 200 events. Mechanism controls show that selective consolidation factorizes into two controllable dimensions: selection and actuation. Matched random gates isolate selection beyond sparse writing; fixed-inner controls across GPT-2, TinyLlama, and Mistral-7B show that inner-loop write strength is model-dependent; and a Mistral-7B matched-gate inversion reveals asymmetric selection-actuation coupling under miscalibrated actuation. Public Memora event streams serve as an external diagnostic, exposing stale-memory invalidation as an unresolved boundary. Within this probe, selective parametric consolidation supplies memory depth distinct from and complementary to retrieval access.

13:00 JSTLLM/生成AI

KARLA: 言語モデルの知識ベース拡張検索

私たちは、LLM がトークン生成中に知識ベースから事実の知識を自動的に取り込むことを可能にする新しい方法を提案します。これは、(1) LLM 出力内の事実の知識は、LLM を再トレーニングすることなく更新できること、(2) LLM 出力内の事実を知識ベースまで追跡して透明性と説明可能性を実現できること、(3) より小さなモデルでもより大きなモデルと同じ事実の精度を達成できることを意味します。私たちの中心的なアイデアは、ナレッジ ベースへのクエリをトリガーする特別なトークンを生成するようにモデルをトレーニングすることです。私たちの実験は、私たちの方法が短い形式と長い形式の両方の生成における事実の根拠を改善し、パラメータの更新ではなくKBの編集を通じて事実の改訂を有効にすることを可能にすることを示しています。

原文 (English)

KARLA: Knowledge-base Augmented Retrieval for Language Models

We propose a new method that allows an LLM to automatically pull in factual knowledge from a knowledge base during token generation. This means that (1)~factual knowledge in the LLM output can be updated without retraining the LLM, (2)~facts in the LLM output can be traced to the knowledge base for transparency and explainability, and (3)~smaller models can achieve the same factual accuracy as larger models. Our core idea is to train the model to produce special tokens that trigger a query to the knowledge base. Our experiments show that our method improves factual grounding in both short and long-form generation, and allows factual revisions to take effect through KB edits rather than parameter updates.

13:00 JST研究/論文

健康な成人における心拍数変動のコンピューター解析

心拍数変動 (HRV) 分析は心臓の生理学的状態の重要な指標であり、病気の診断に役立ちます。しかし、健康な人の HRV パラメータに関する研究は依然として限られており、ゴールドスタンダードは存在しません。この研究では、HRV の臨床的有用性を向上させるために、40 人の健康な成人 (男性 20 人、女性 20 人、30 ~ 50 歳) の HRV 指数を評価します。信号処理とデータ分析のための計算手法を使用して、時間、周波数、および非線形インデックスが分析され、(1) 正規性、(2) 安定性、(3) 相関性、(4) 再現性、および (5) 一貫性の 5 つの質問に対処しました。主な発見: (1) 時間領域および非線形指数、特にグローバルおよび LF (低頻度) は正規分布に従い、性差が認められます。 (2) HF (高周波) 関連のものを除いて、ほとんどの指数は安定しています。 (3) HF 関連指標の高い相関は冗長性を示唆しており、研究では 1 つだけが必要であることを示しています。 (4) Fantasia データベースとの比較では、女性の SD2 と SDNN (15% 以上) を除き、ほとんどの指数の誤差が 10% 未満であることが明らかになりました。 (5) 時間領域インデックスと非線形インデックスは研究間の変動が小さいのに対し、周波数領域インデックスは高い変動を示し、研究間の比較が制限されます。選択された指標、ApEn および IRRR (グローバル変動)、HRVi および SD2 (LF)、および MADRR または rMSSD (HF) は、HRV コンポーネントを正確に表し、その臨床および研究の関連性を高めるのに最適です。

原文 (English)

Computational Analysis of Heart Rate Variability in Healthy Adults

Heart Rate Variability (HRV) analysis is a key indicator of cardiac physiological state and aids in disease diagnosis. However, research on HRV parameters in healthy individuals remains limited, and no gold standard exists. This study evaluates HRV indices in 40 healthy adults (20 men, 20 women, aged 30-50) to improve HRV's clinical utility. Using computational methods for signal processing and data analysis, time, frequency, and nonlinear indices were analyzed to address five questions: (1) normality, (2) stability, (3) correlation, (4) reproducibility, and (5) consistency. Key findings: (1) Time-domain and nonlinear indices, particularly global and LF (low frequency), follow normal distributions, with gender differences noted. (2) Most indices are stable except HF (high frequency)-related ones. (3) High correlations in HF-related indices suggest redundancy, indicating only one is necessary in studies. (4) Comparisons with the Fantasia database revealed less than 10% error for most indices, except SD2 and SDNN in women (greater than 15%). (5) Time-domain and nonlinear indices show low inter-study variability, while frequency-domain indices exhibit high variability, limiting cross-study comparisons. The selected indices-ApEn and IRRR (global variability), HRVi and SD2 (LF), and MADRR or rMSSD (HF)-are best suited for accurately representing HRV components and enhancing its clinical and research relevance.

13:00 JSTLLM/生成AI研究/論文

機能のフロンティア: ベンチマークはモデルのパフォーマンスの 82% を逃しています

既存のベンチマークは通常、1 回の実行で 1 つのモデルの精度を報告します。これは、特に異種データ分布の下で、現実世界の LLM 機能を系統的に過小評価しています。(i) 異なるモデルは、その専門分野に応じて異なる問題を正解し、(ii) 予算が与えられれば、複数の世代をサンプリングして選択的に保持できます。このギャップを定量化するために、機能フロンティアを導入します。これは、モデルおよび世代全体にわたる最適な選択(つまり、オラクルによる)の下で、各コスト レベルで達成可能な最高のパフォーマンスを特徴付ける一連のモデルにわたるパレート フロンティアです。私たちの構築では、単一モデルの評価による過小評価と、ノイズの多いサンプルに対して最大値を取ることによる過大評価という 2 つの相反するバイアスが補正されます。私たちは、コーディング、推論、医学、事実性、指示に従って、エージェントタスクにわたる 16 の広く使用されているベンチマークにわたる 21 の LLM を調査し、同等のコストでの Capability Frontier のパフォーマンスを各ベンチマークの最高パフォーマンスのモデルと比較しました。単一モデルの評価を修正すると、エラー率が 54% 減少します。単一実行をさらに補正すると、82% の改善が得られ、85% のコスト削減と同等の SOTA 精度が得られます。これらの経験的結果を補完するために、制御された確率的シミュレーションを使用して、クエリ トピックのエントロピーが高くなると、Oracle ルーティングと最良の単一モデルの間のパフォーマンス ギャップがほぼ単調増加することを示します。私たちの調査結果は、集合的な LLM 機能が大幅に過小評価されており、異種データのマルチドメイン設定での評価と展開に影響を与えることを示唆しています。

原文 (English)

The Capability Frontier: Benchmarks Miss 82% of Model Performance

Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data distributions: (i) different models get different questions correct according to their specializations, and (ii) given a budget, multiple generations can be sampled and selectively retained. To quantify this gap, we introduce the Capability Frontier: a Pareto frontier over a set of models that characterizes the best achievable performance at each cost level under optimal selection across models and generations (i.e., via an oracle). Our construction corrects for two opposing biases: underestimation from single-model evaluation and overestimation from taking maxima over noisy samples. We study 21 LLMs across 16 widely used benchmarks spanning coding, reasoning, medicine, factuality, instruction following, and agentic tasks, comparing Capability Frontier performance at matched cost to each benchmark's top-performing model. Correcting for single-model evaluation yields a 54% error rate reduction; additionally correcting for single runs yields an 82% improvement, with SOTA accuracy matched at 85% cost reduction. Complementing these empirical results, we use controlled probabilistic simulations to show that higher query topic entropy produces a near-monotonic increase in the performance gap between oracle routing and the best single model. Our findings suggest collective LLM capabilities are substantially underestimated, with implications for evaluation and deployment in data-heterogeneous, multi-domain settings.

13:00 JSTLLM/生成AI

倉庫最適化のための最適化パイプラインのコンテキスト認識型合成

手動のピッカーから商品までの倉庫での注文の履行には、品目の割り当て、注文のバッチ処理、ピッカーのルーティングなど、相互に関連した決定が含まれます。統合モデルはこれらの意思決定間の相互作用をキャプチャしますが、実際の倉庫システムでは、組織の境界、責任の違い、またはデータの可用性の制限により、分解されたアプローチが必要になることがよくあります。既存の研究では主に、特定のウェアハウス設定における孤立した部分問題または固定された部分問題の組み合わせのアルゴリズムを評価していますが、適用可能なアルゴリズム構成を決定し、それらを有効なソリューション パイプラインに構成して、そのパフォーマンスを評価するための一般的なメカニズムが不足しています。 Context-Aware Synthesis of Optimization Pipelines (CASOP) を使用して、コンテキスト固有の最適化パイプラインを構築および評価するためのフレームワークを提案し、これらを注文フルフィルメントに適用します。このフレームワークは以下で構成されます。(1) 一般的な注文履行の問題に対するアルゴリズムのモジュール式リポジトリ。 (2) ウェアハウスのコンテキストとアルゴリズム要件を説明するためのセマンティック データとアルゴリズム カード。 (3) 注文履行の問題を関連する下位問題に構造化する分類法。 (4) 特定のウェアハウス コンテキストに適用可能なアルゴリズムを特定し、すべての有効な最適化パイプラインを構成するパイプライン シンセサイザー。 (5) 結果として得られるすべてのパイプラインを評価するパイプライン エバリュエーター。 4 つの問題クラスをカバーする 7 つのベンチマーク インスタンス セットでフレームワークを実証し、結果として 1,063,044 の有効なパイプラインが得られます。このフレームワークは、倉庫業務用の有効で高性能なアルゴリズム パイプラインの設計、自動合成、選択において研究者や実務者をサポートします。このソフトウェアはオープンソースであり、https://github.com/kit-dsm/ware_ops_pipes および https://github.com/kit-dsm/ware_ops_algos から入手できます。キーワード: 倉庫の最適化、アルゴリズムの選択、パイプライン合成、注文処理

原文 (English)

Context-Aware Synthesis of Optimization Pipelines for Warehouse Optimization

Order fulfillment in manual picker-to-goods warehouses involves interconnected decisions such as item assignment, order batching, and picker routing. While integrated models capture interactions between these decisions, practical warehouse systems often require decomposed approaches due to organizational boundaries, differing responsibilities, or limited data availability. Existing studies primarily evaluate algorithms for isolated subproblems or fixed subproblem combinations for specific warehouse settings, but lack a general mechanism to determine applicable algorithm configurations, compose them into valid solution pipelines, and assess their performance. With Context-Aware Synthesis of Optimization Pipelines (CASOP), we propose a framework for constructing and evaluating context-specific optimization pipelines and apply these to order fulfillment. The framework comprises: (1) a modular repository of algorithms for common order fulfillment problems; (2) semantic data and algorithm cards to describe warehouse context and algorithm requirements; (3) a taxonomy that structures order fulfillment problems into relevant subproblems; (4) a pipeline synthesizer that identifies applicable algorithms for a given warehouse context and composes all valid optimization pipelines; and (5) a pipeline evaluator that assesses all resulting pipelines. We demonstrate the framework on 7 benchmark instance sets covering four problem classes, resulting in 1,063,044 valid pipelines. The framework supports researchers and practitioners in designing, automatically synthesizing, and selecting valid, high-performing algorithmic pipelines for warehouse operations. The software is open-source and available at https://github.com/kit-dsm/ware_ops_pipes and https://github.com/kit-dsm/ware_ops_algos. Keywords: Warehouse optimization, Algorithm selection, Pipeline synthesis, Order fulfillment

13:00 JST研究/論文GPT / ChatGPT

LCAi: ビッグデータの融合と検索拡張生成支援解釈によるライフサイクル評価

ライフサイクル評価の解釈段階では、技術的、社会的、政策的不確実性の下で、環境ホットスポットに対処する定量化された改善の機会を実行可能な戦略的経路に変換するための構造化されたメカニズムが欠けていることがよくあります。この制限を克服するために、この研究では、LCA 解釈のためのパースペクティブ条件付き検索拡張生成フレームワークを導入します。このフレームワークでは、マルチパースペクティブ検索と制御合成が人工知能 (AI) 支援 LCA に組み込まれています。 LCA 解釈における大規模な言語モデルを運用するために、学術、業界、公的議論、および欧州連合 (EU) の資金提供データセットをカバーするパースペクティブ フュージョン RAG アーキテクチャが開発されました。私たちのアプローチは 3 つのステップで構成されます: (1) システム境界と脱炭素化目標を定義するシナリオ アンカー、(2) 制約付き検索を伴うパースペクティブ固有の一連のマイクロクエリ、(3) それ以上の検索は行わずに台帳に保存された出力のみを統合する中立的な合成ステップ。このフレームワークは、推論モデルとして GPT-5 nano を使用したイタリアのリンゴ生産施設における水素によるディーゼル削減のユースケースを通じて実証されています。全体として、構造化検索と制約付き合成は、クロスドメインの多様性を維持しながら幻覚のリスクを軽減するように設計されています。提示されたアプローチは、影響結果をより規律正しく戦略的経路に変換することをサポートし、LCA 研究、特に大規模に導入できるテクノロジーに焦点を当てた高度な AI ツールを使用するための新しい道を開きます。この概念実証は、AI 支援の証拠に基づく解釈が、従来の LCA 研究を超えて実装指向の意思決定をどのようにサポートできるかを示しています。

原文 (English)

LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation

The interpretation phase of life cycle assessment often lacks structured mechanisms for translating quantified improvement opportunities addressing environmental hotspots into actionable strategic pathways under technological, social, and policy uncertainty. To overcome this limitation, this study introduces a perspective-conditioned retrieval-augmented generation framework for LCA interpretation, where a multi-perspective retrieval and controlled synthesis is incorporated in the artificial intelligence (AI)-assisted LCA. To operationalise large language models in LCA interpretation, a perspective fusion RAG architecture was developed, covering academic, industry, public discourse, and European union (EU) funding datasets. Our approach comprises three steps: (1) a scenario anchor defining system boundaries and decarbonization targets, (2) a set of perspective-specific micro-queries with constrained retrieval, and (3) a neutral synthesis step integrating only ledger-stored outputs without further retrieval. The framework is demonstrated through a hydrogen-enabled diesel reduction use case in an Italian apple production facility using GPT-5 nano as the reasoning model. Overall, the structured retrieval and constrained synthesis are designed to mitigate the risk of hallucination while preserving cross-domain diversity. The approach presented can support more disciplined translation of impact results into strategic pathways and opens up new avenues for the use of advanced AI tools in LCA studies, particularly those focused on technologies that could be deployed at scale. This proof-of-concept demonstrates how AI-assisted, evidence-grounded interpretation can support implementation-oriented decision-making beyond conventional LCA studies.

13:00 JSTLLM/生成AIエージェント研究/論文

AgentX: 産業用レコメンダー システムのエージェント駆動型自己反復に向けて

レコメンデーション アルゴリズムの反復は、職人的なエンジニアに拘束されたプロセスから工業化された研究ループに移行していますが、この移行は構造的な実行ボトルネックによって妨げられたままです。アイデアから発売までのサイクルは依然として人間のエンジニアに依存して、仮説の生成、製品コードの変更、A/B 実験の開始、およびオンライン結果の属性に依存しています。したがって、イノベーションは、証拠、計算、蓄積された実験知識と複合するのではなく、従業員数に比例してスケールします。この本番機能を根本的に再構築する、本番環境に展開されるマルチエージェント システムである AgentX を紹介します。 AgentX は、自己進化する開発エンジンとして動作します。手動ワークフローでは維持できない規模とペースで、推奨実験を自律的に生成、実装、評価し、そこから学習します。このシステムは、密接に結合された 4 つのステージを閉ループで調整します。 Brainstorm Agent は、過去の実験、システム アーキテクチャ、データ分析、外部調査からの証拠を総合して、ランク付けされた実行可能な提案を作成します。開発エージェントは、リポジトリに基づいた生成と多次元の信頼性検証を通じて、各提案を本番環境に対応したコードに変換します。評価エージェントは、ガードレール拒否権のある A/B 判定を使用して安全なオンライン ロールアウトを実行し、成功と失敗の両方を構造化された知識資産に変換します。その後、ハーネス エボリューション レイヤー (SGPO) が実行軌跡を意味論的勾配更新に抽出し、エージェント自体を継続的に強化することで、システムを単に自動化するだけでなく、自己改善するシステムにします。

原文 (English)

AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly with headcount rather than compounding with evidence, compute, and accumulated experimental knowledge. We present AgentX, a production-deployed multi-agent system that fundamentally restructures this production function. AgentX operates as a self-evolving development engine: it autonomously generates, implements, evaluates, and learns from recommendation experiments at a scale and pace that no manual workflow can sustain. The system orchestrates four tightly coupled stages in a closed loop. A Brainstorm Agent synthesizes evidence from historical experiments, system architecture, data analysis, and external research into ranked, executable proposals. A Developing Agent translates each proposal into production-ready code through repository-grounded generation and multi-dimensional reliability verification. An Evaluation Agent conducts safe online rollout with guardrail-vetoed A/B judgment, converting both successes and failures into structured knowledge assets. A Harness Evolution layer (SGPO) then distills execution trajectories into semantic-gradient updates that continuously sharpen the agents themselves -- making the system not merely automated, but self-improving.

13:00 JST研究/論文

TAVR-VLM: 幻覚耐性レポート生成のためのリスク条件付き因果グラウンディング

経カテーテル的大動脈弁置換術 (TAVR) の計画には、細心の注意を払った複合的な推論が必要です。ただし、マルチモーダル大規模言語モデル (MLLM) をこの一か八かの領域に適応させることは、生成されたテキストに解剖学的根拠が欠けている幻覚診断によって大きく妨げられます。これに対処するために、モデル内部の「リスク $\rightarrow$ 領域 $\rightarrow$ Word」構造的グラウンディング経路をインスタンス化する、リスク条件付き因果グラウンディング アテンション (R-CGA) を特徴とする新しいフレームワークである TAVR-VLM が導入されました。 R-CGA は、マルチモーダルな入力を因果的リスクのボトルネックに圧縮し、密集した視覚的特徴をグローバル リスク マスクに精製します。自己回帰生成中、サポート投影の因果的一貫性目標により、リスク定義のサポート マスク内のトークン レベルのグラウンディングが制限されます。 TAVR-VLM は、包括的な 1,482 人の患者コホートである $\text{M}^3\text{TAVR}$ で評価され、新しい最先端技術を確立します。 AUROC 0.896 を達成し、CIDEr を 0.936 に高め、幻覚率を 8.1\% に大幅に低減することで、証拠に基づいた外科用 AI の解釈可能性を向上させます。

原文 (English)

TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

Transcatheter Aortic Valve Replacement (TAVR) planning requires meticulous multimodal reasoning. However, adapting Multimodal Large Language Models (MLLMs) to this high-stakes domain is severely impeded by diagnostic hallucinations, where generated text lacks anatomical grounding. To address this, TAVR-VLM is introduced: a novel framework featuring Risk-Conditioned Causal Grounding Attention (R-CGA) that instantiates a model-internal ``Risk $\rightarrow$ Region $\rightarrow$ Word'' structural grounding pathway. R-CGA compresses multimodal inputs into a causal risk bottleneck, purifying dense visual features into a global risk mask. During autoregressive generation, a support-projected causal consistency objective constrains token-level grounding within the risk-defined support mask. Evaluated on $\text{M}^3\text{TAVR}$, a comprehensive 1,482-patient cohort, TAVR-VLM establishes a new state-of-the-art. It achieves an AUROC of 0.896, boosts CIDEr to 0.936, and drastically reduces the hallucination rate to 8.1\%, thereby improving interpretability for evidence-based surgical AI.

13:00 JSTビジネス/資金調達

大規模言語モデルを使用して縦断的な合成臨床ノートを生成するパイプライン

実世界のデータへのアクセスが制限されている領域で AI システムの開発と評価を可能にするために、合成データの使用が増えています。医療分野では、臨床文書はその機密性により特別な課題を抱えています。この研究では、実際の患者データに伴うプライバシー リスクを回避しながら、臨床 AI ツールの開発をサポートするように設計された合成臨床メモ パイプラインとデータセットを導入します。データセットは、大規模な言語モデルを使用した構造化患者生成、半構造化患者ジャーニー シミュレーション、および非構造化臨床ノート生成を組み合わせたモジュール式パイプラインを使用して生成されます。このパイプラインは、長期的な患者記録全体にわたる内部一貫性を優先すると同時に、書き方、メモの構造、臨床の詳細の変化も捕捉するように設計されています。 LLM ベースの検証および拡張ステップを含む追加のメカニズムを使用して、生成されたノートの忠実性、リアリズム、および多様性が向上します。私たちは、70 人の合成患者のデータセットをリリースします。各患者には、入院期間全体にわたる 20 ~ 50 の臨床ノートが関連付けられています。データセットは複数の検証レベルで提供されているため、ユーザーはユースケースに応じて現実性とスケーラビリティのバランスを取ることができます。このデータセットは、実際の患者データに依存することなく、要約ツール、コーディング モデル、意思決定支援システムなどの臨床 AI システムの開発、テスト、評価をサポートします。

原文 (English)

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

13:00 JST研究/論文

メトリクス順序シーケンストレーニングとハイブリッドポリシー優先最適化を備えた拡散トランスフォーマーを介した生成検索

埋め込みベースの検索では、共有ベクトル空間内のクエリとの類似性によって項目をランク付けし、通常は最高スコアの項目を返すことを目的としています。多くの運用設定では、これは望ましくないことです。きめの細かいパターンを表現するシード セットを考えると、ターゲット属性を満たし、そのパターン内に留まるアイテムがさらに必要になります。これをパターン保持属性の取得として形式化します。 2 つの目標は相互に影響し合っています。シードを平均化すると、パターンは維持されますが、低属性の領域にとどまりますが、グローバルな属性の取得では無関係なパターンに偏ってしまいます。このタスクには、モデルが一連の項目エンベディングを読み取り、最近傍検索用のクエリ エンベディングを生成する、連続生成検索を使用してタスクに取り組みます。私たちは、生シーケンス事前トレーニング、マルチドメイン メトリック順序付け継続事前トレーニング、テールセントロイド微調整、および HPPO を備えた段階的フレームワークである MO-DiT + HPPO を提案します。メトリクス順序トレーニングは、まばらなオンライン検索ラベルを、予測属性密度が低いものから高いものへと順序付けされたパターン内の軌跡に変換し、1 つのモデルにドメイン全体にわたるメトリクス改善の方向性を教えます。 HPPO は、ハイブリッド候補プールにオンライン交差メトリックをラベル付けし、参照に基づいた優先順位の最適化を適用することにより、生成されたクエリ分布を真のオンライン目標に合わせて調整します。パレート ペア フィルターは、同じパターンの純度を低下させない勝者ペアのみを保持し、パターンを犠牲にすることなく属性メトリックを向上させます。項目ホールドアウト プロトコルおよびパターン ホールドアウト プロトコルに基づく 4 つの属性ドメイン全体で、メトリック順序付け DiT は事前学習済み生成検索よりも交差メトリックを改善し、HPPO はそれをさらに改善し、8 つのドメイン分割セルのうち 7 つで大幅な改善が見られ、最も困難な分割では僅差の同点となりました。メトリクスと予測子の検証、順序の除去、CPT/SFT の比較、および候補とポリシーの除去により、利益がどこから得られるのかがわかります。

原文 (English)

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

Embedding-based retrieval ranks items by their similarity to a query in a shared vector space and usually aims to return the highest-scoring items. In many production settings this is not what is wanted: given a seed set that expresses a fine-grained pattern, one needs more items that both satisfy a target attribute and stay within that pattern. We formalize this as pattern-preserving attribute retrieval. The two goals pull against each other: averaging the seeds preserves the pattern but stays in a low-attribute region, while global attribute retrieval drifts to unrelated patterns. We approach the task with continuous generative retrieval, where a model reads a sequence of item embeddings and generates query embeddings for nearest-neighbor search. We propose MO-DiT+HPPO, a staged framework with raw-sequence pretraining, multi-domain metric-ordered continuation pretraining, tail-centroid fine-tuning, and HPPO. Metric-ordered training turns sparse online retrieval labels into in-pattern trajectories ordered from low to high predicted attribute density, teaching one model the metric-improvement direction across domains. HPPO aligns the generated query distribution with the true online objective by labeling a hybrid candidate pool with the online intersection metric and applying reference-anchored preference optimization. A Pareto pair filter keeps only winner pairs that do not lower same-pattern purity, raising the attribute metric without sacrificing the pattern. Across four attribute domains under item- and pattern-holdout protocols, metric-ordered DiT improves the intersection metric over a pretrained generative retriever, and HPPO improves it further, with significant gains on seven of eight domain-split cells and a marginal tie on the hardest split. Metric-predictor validation, order ablations, CPT/SFT comparisons, and a candidate-policy ablation show where the gains come from.

13:00 JST研究/論文

マルチタスク結合モデルからタスクエキスパートを回復する方法を学ぶ

マルチタスク モデルのマージは、複数のタスク固有の専門家を 1 つの統一モデルに統合することを目的としていますが、静的マージではパラメータの干渉が常に発生します。動的マージ モデルはこのギャップを埋めることを目的としていますが、多くの研究は、推論時にコストのかかるストレージと冗長なエキスパート コンポーネントの読み込みに依存しています。この研究では、タスク エキスパートの観点から、パラメータ干渉を、マージ プロセス中に各エキスパートに導入されるパラメータの摂動として見ます。このようなパラメータの摂動はアフィン変換としてモデル化でき、加算オフセットとして近似できることを示します。これらを動機として、パラメータ干渉を元に戻し、単一のマージされたチェックポイントからタスク エキスパートのパフォーマンスを回復するために、これらのオフセットを予測するフレームワークである Recover Task eXpert (ReTeX) を提案します。タスク ID が不明な場合に適切なエキスパートを回復するために、推論前にオフラインで計算された SVD 部分空間署名に基づくルーターフリーのタスク ID を導入します。推論時に、識別子は、指定された入力に対して部分空間が最小の射影残差をもたらすタスクを選択します。その結果、ReTeX は視覚領域と NLP 領域の両方で個人の専門家のパフォーマンスの 95% 以上を回復し、目に見えないタスクへの一般化を大幅に向上させます。重要なことに、パラメータ オフセット予測が、配布外 (OOD) タスクに対する専門知識の創発的適応補間につながることも示します。 ReTeX は、目に見えないタスクを処理するために、目に見える専門知識を適応的に補間します。私たちのコードは https://github.com/BAIKLAB/ReTeX で入手できます。

原文 (English)

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference. In this work, from the perspective of task expert, we view parameter interference as parameter perturbation introduced to each expert during merging process. We show that such parameter perturbations can be modeled as affine transformation, which can be approximated as additive offsets. Motivated by these, we propose Recover Task eXpert (ReTeX), a framework that predicts those offsets, in order to undo parameter interference and recover task-expert performance from a single merged checkpoint. To recover the appropriate expert when task identity is unknown, we introduce a router-free task identifier based on SVD subspace signatures computed offline before inference. At inference, the identifier selects the task whose subspace yields the smallest projection residual for a given input. As a result, ReTeX recovers over 95% of individual-expert performance in both vision and NLP domains, while significantly improving generalization to unseen tasks. Crucially, we also show that the parameter offset prediction leads to emergent adaptive interpolation of expert knowledge for out-of-distribution (OOD) tasks. ReTeX adaptively interpolates seen expert knowledge to handle unseen tasks. Our code is available at https://github.com/BAIKLAB/ReTeX

13:00 JSTエージェント

言語エージェントのタスク非依存性の診断

大規模な言語モデルは有能な長期的なエージェントとして機能しますが、その配布外 (OOD) 一般化は依然として弱いままです。私たちは、この失敗の主な原因はタスクの鈍感であると特定しています。似ているが異なるタスクに直面した場合、モデルはトレーニング中に学習したパターンを適用し、目の前のタスクを解決できない可能性があります。命令が意味的に壊れていて直接応答できない場合でも、モデルは多くの場合、元のタスクに沿ったアクションを続行することを示します。さらに、トレーニングされたプロンプト内のタスクの説明を、類似しているが異なる別のタスクに置き換えた場合でも、モデルは同じアクションを出力する可能性があることがわかりました。この動作には、トレーニング中の注意がタスク トークンから離れてローカルの観察の方に一貫して移ることが伴い、ショートカットへの最適化バイアスが示唆されています。この問題を軽減するために、タスク命令へのアクションの依存を明示的に促進する軽量の対照的正則化装置である Task-Perturbed NLL Optimization を提案します。広範な評価により、私たちの介入により、タスクトークンに対するより安定した注意が維持されながら、タスクの感度とOODの一般化が向上することが示されました。

原文 (English)

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of this failure as task insensitivity: when faced with similar but distinct tasks, models might apply patterns learned during training and fail to solve the task at hand. We show that models often continue with actions aligned with the original task even when the instruction is semantically corrupted and cannot be directly answered. We further find that, when we replace the task description in a trained prompt with another similar but distinct task, the model may still output the same action. This behavior is accompanied by a consistent training-time attention drift away from task tokens and toward local observations, suggesting an optimization bias toward shortcuts. To mitigate this problem, we propose Task-Perturbed NLL Optimization, a lightweight contrastive regularizer that explicitly encourages action dependence on the task instruction. Extensive evaluations show that our intervention improves task sensitivity and OOD generalization while preserving more stable attention to task tokens.

13:00 JSTLLM/生成AIエージェント

CoT トレーニングは LLM ベースのエージェントにどのような影響を及ぼしますか?

思考連鎖 (CoT) 推論は言語モデル エージェントで広く使用されていますが、これまでの研究では、言語化された CoT が常に忠実であるとは限らず、事後推論を反映している可能性があることが示されています。これは、モデルが推論する前にすでに答えを知っていることを意味します。したがって、CoT トレーニングによって実際に何が改善されているかを尋ねます。モデルは、生成された推論を通じてアクションを変更することがうまくなっているのでしょうか、それともプロンプトから直接アクションを予測することがうまくなっているのでしょうか? \emph{プロンプトアクション} (CoT なしでアクションを予測する) と CoT アクション (CoT ありでアクションを予測する) を比較することで、この問題を研究します。チェックポイント全体で、迅速なアクションの品質が大幅に向上します。環境と対話している間、プロンプト アクションに対する CoT アクションの相対的な利点は同様のままであり、CoT トレーニングによって CoT 推論の利点が広がることはなく、プロンプト アクションの質の向上に役立つことが示されています。さらに、後のチェックポイントでは CoT に応じてアクションを修正する可能性が低く、プロンプトへの依存度が高まっていることがわかります。これらのパターンに動機付けられて、トレーニング サンプルの一部でアクション トークンの監視を選択的にマスクします。この介入により、領域外の一般化が向上します。

原文 (English)

Where Do CoT Training Gains Land in LLM based Agents?

Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may instead reflect post-hoc reasoning, which means the model already knows the answer before reasoning. We therefore ask what CoT training is actually improving: is the model getting better at changing its action through generated reasoning, or is it getting better at predicting the action directly from the prompt? We study this question by comparing \emph{prompt actions} (predicting action without CoT) with CoT actions (predicting action with CoT). Across checkpoints, prompt-action quality improves substantially. While interacting with the environment, the relative advantage of CoT actions over prompt actions remains similar, showing that CoT training does not widen the advantage of CoT reasoning, and it helps to improve the quality of prompt actions. We further find that later checkpoints are less likely to revise the action in response to CoT, suggesting greater reliance on the prompt. Motivated by these patterns, we selectively mask action-token supervision on a fraction of training examples. This intervention improves out-of-domain generalization.

13:00 JST画像/動画生成

Look-Before-Move: ダイナミックな 3D ストーリーワールドにおける物語に基づいた世界の視覚的注意

身体化された AI と世界モデルが動的 3D 環境で動作することが増えているため、視覚認識は、与えられた観察を受動的に解釈するだけでなく、何を観察するかを能動的に決定する方向に進む必要があります。私たちは、動的な 3D ストーリー世界でのカメラ計画を通じてこの問題を研究します。そこでは、カメラは滑らかな動きを生成するだけでなく、移動する前にどのような視覚的証拠を取得する必要があるかを決定する必要があります。私たちはこの機能を、物語に基づいた世界の視覚的注意として定式化します。カメラは、何を観察するか、どのように観察を構成するか、そして物語の意図と物理的な 3D 制約の下で時間の経過とともにどのように注意を移すかを決定する具体化された観察者として機能します。この機能を実現するために、観察仕様をモーション実行から分離するカメラ計画フレームワークである Look-Before-Move を提案します。まずセマンティック観察コントラクトを構築して、監督の意図を実行可能な視覚的制約に変換し、次にモンテカルロ視点検索を実行して物語に準拠し、幾何学的に実現可能な視点を見つけます。最後にセマンティック軌道グラウンディングを適用して、選択された視点を連続的で衝突を認識し、時間的に一貫したカメラの動きに接続します。さらに、StoryBlender に基づいて動的な 3D ストーリー ワールド ベンチマークを構築し、アニメーション キャラクター、セマンティック シーン構成、および実行可能な 3D 環境を含む 50 のストーリー、457 のシーン、および 1585 のショットをカバーします。実験では、私たちのフレームワークが代表的なベースラインよりも被写体の知覚、意図の一貫性、軌跡の品質を向上させることが示されており、カメラの動きを生成する前に視覚的な注意を組織することの重要性が実証されています。

原文 (English)

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as Narrative-Grounded World Visual Attention, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose Look-Before-Move, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on StoryBlender, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.

13:00 JSTLLM/生成AI画像/動画生成

アインシュタインの世界モデル

知性には、直接の経験を超えて現象について推論する能力が必要ですか?言語だけでは複雑な思考を捉えることができないのではないかと疑うのは自然なことです。しかし、この研究で特に懸念されるのは、反事実的な出来事を視覚化することが、複雑な思考のメカニズムとして言語を補完できるかどうかである。私たちは、LLM がそのような視覚化メカニズムを、推論能力に役立つ方法で利用できるように訓練できるかどうかを尋ねます。この疑問を動機として、私たちはアインシュタイン世界モデルを提案します。 EWM は、推論トレース内に視覚と時間のロールアウトを配置する LLM ベースの推論システムの青写真であり、テキストだけでは十分にサポートできない方法で推論できるようになります。 EWM では、LLM はワールド モジュール (ワールド モデルと混同しないでください) を呼び出し、検討中のシーンの短いロールアウトを生成します。返されたロールアウトは、答えとしてではなく、後の推論をサポートできる検査可能な仮説として扱われます。 Einstein World Models は、ツール呼び出し (Web 検索やコード実行など) のための LLM の機能を視覚的な思考実験の領域に拡張します。

原文 (English)

Einstein World Models

Does intelligence require the ability to reason about phenomena beyond direct experience? It is natural to suspect that some complex thought cannot be captured through language alone. However, of particular concern to this work, is whether visualising counterfactual events can complement language as a mechanism for complex thought. We ask whether LLMs can be trained to utilise such visualisation mechanisms, in a way that benefits their reasoning abilities. Motivated by this question, we propose Einstein World Models. EWMs are a blueprint for LLM-based reasoning systems that place visual-temporal rollouts inside the reasoning trace, allowing them to reason in ways that text alone may not support well. In an EWM, the LLM calls a world-module (not to be confused with a world model), to produce short rollouts of scenes under consideration. The returned rollout is treated not as the answer, but as an inspectable hypothesis that can support later reasoning. Einstein World Models extend the capability of LLMs for tool calling (such as web search or code execution), into the domain of visual thought experiments.

13:00 JST研究/論文

Resilient AI のための適応型ユーティリティ主導のリソース オーケストレーション (AURORA-AI)

最新の AI システムは、非定常の計算条件、人口統計条件、運用条件の下で導入されることが増えており、静的なリソース割り当て戦略により、予測パフォーマンスと、公平性や説明可能性などの人間中心の特性の両方が低下します。この論文では、Hamilton-Jacobi-Bellman フィードバック制御、リアプノフベースの安定性モニタリング、および公平性を意識した複合ユーティリティを単一の閉ループ ポリシーに統合する、Resilient AI 向けの適応ユーティリティ主導型リソース オーケストレーション フレームワークである AURORA-AI について紹介します。このフレームワークは、異種 AI モデルの母集団全体に計算予算を継続的に再配分するため、予測パフォーマンス、人口統計的均等性、コストを合わせて定義されたグローバル ユーティリティが、遅延、堅牢性、解釈可能性は、混乱下でも最大化されたままになります。このフレームワークは、人口統計上のバイアス ショック、段階的な概念ドリフト、突然のブラック スワンの混乱を同時に注入する、ストレスの多い離散時間シミュレーションで評価され、静的、ラウンド ロビン、貪欲、LinUCB、および近接ポリシー最適化に基づく深層強化学習エージェントを含む 5 つの確立されたコントローラーと比較されます。 AURORA-AI は、静的ベースラインの 88 タイム ステップと近接ポリシー最適化の 22 タイム ステップと比較して、ブラック スワン イベントからの即時回復を達成し、アルファ分位と超分位をそれぞれ 29 パーセントと 25 パーセント引き上げ、平均と最大の人口統計的パリティ ギャップを同時に削減し、リアプノフ安定動作ステップの割合を増加させます。これらの結果は、安定性理論に基づいた公平性を意識した適応型オーケストレーションが、回復力のある人間中心の AI 導入に向けた実践的かつ理論的に動機付けられた道であることを示しています。

原文 (English)

Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI)

Modern AI systems are increasingly deployed under non-stationary computational, demographic, and operational conditions in which static resource allocation strategies degrade both predictive performance and human-centric properties such as fairness and explainability. This paper presents AURORA-AI, an Adaptive Utility-driven Resource Orchestration framework for Resilient AI that unifies Hamilton-Jacobi-Bellman feedback control, Lyapunov-based stability monitoring, and a fairness-aware composite utility into a single closed-loop policy.The framework continuously redistributes computational budget across a population of heterogeneous AI models so that the global utility, defined jointly over predictive performance, demographic parity, cost, latency, robustness, and interpretability, remains maximised under disruption. The framework is evaluated in a stress-rich discrete-time simulation that concurrently injects demographic bias shocks, gradual concept drift, and abrupt black-swan disruptions, and is compared against five established controllers including Static, Round Robin, Greedy, LinUCB, and a deep reinforcement-learning agent based on Proximal Policy Optimisation. AURORA-AI achieves immediate recovery from the black-swan event compared to eighty-eight time steps for the Static baseline and twenty-two for Proximal Policy Optimisation, lifts the alpha-quantile and the super-quantile by twenty-nine and twenty-five percent respectively, simultaneously reduces the mean and maximum demographic parity gap, and increases the fraction of Lyapunov-stable operating steps. These results indicate that fairness-aware adaptive orchestration grounded in stability theory is a practical and theoretically motivated path toward resilient human-centric AI deployment.

13:00 JSTLLM/生成AIエージェント

反復 LLM エージェント ループのセマンティック早期停止

マルチエージェント大規模言語モデル (LLM) ループ (たとえば、草稿を作成するライターと改訂を行う批評家など) は、ほとんどの場合、固定の反復上限 (max_iterations) によって終了します。これは構文上のキルスイッチです。答えがまだ改善されているかどうかが分からないため、簡単な入力にトークンを過剰に消費し、難しい入力を切り捨てます。私たちはセマンティックな早期停止を研究します。つまり、連続するドラフト埋め込みの意味の変化 (忍耐ウィンドウによるコサイン距離) が停止し、回答の測定品質の向上が停止すると、ループが停止します。私たちの仕事は 3 つの貢献をします。まず、正直な理論的基礎です。距離数列の収束を、(以前に過剰に主張されていた)バナッハ短縮ではなく、経験的にテストされた予想として扱いながら、決定論的な終端と明確な定義性を証明し、これらの主張を機械チェックします。 2 番目に、ジャッジの効率的な評価プロトコルです。各質問の完全な軌跡を 1 回生成し、同一のドラフトですべての停止ポリシーを再生し、すべての LLM ジャッジ呼び出しをキャッシュして、厳密にペアになった効率と品質の比較を低コストで実現します。さらに、運用トークン (ポリシーにチャージ) を評価トークン (測定手段) から分離します。 3 番目は、マルチホップ検索拡張質問応答 (HotpotQA) に関する実証研究です。 60 問のテスト分割では、ジャッジフリーのセマンティック ストッパーはパリティ品質 (Delta-IS = -0.004、p = 0.81) での max_iterations と比較してオペレーショナル トークンを 38% 削減しますが、完全な品質ゲートのバリアントはラウンドごとの判定がコストを支配するため逆効果です。最良のラウンドを選択したオラクルは、すべての実際的なポリシーに対して +0.115 の情報スコアを達成し (p ~ 4e-11)、問題を「いつ停止するか」 (簡単) から「どのラウンドが最適か」 (オープン) に再構成します。

原文 (English)

Semantic Early-Stopping for Iterative LLM Agent Loops

Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a fixed iteration cap (max_iterations). This is a syntactic kill-switch: it is blind to whether the answer is still improving, so it over-spends tokens on easy inputs and truncates hard ones. We study semantic early-stopping: the loop halts when consecutive draft embeddings stop changing in meaning (cosine distance with a patience window) and the answer's measured quality stops improving. Our work makes three contributions. First, an honest theoretical footing: we prove deterministic termination and well-definedness and machine-check these claims, while treating the convergence of the distance sequence as an empirically tested conjecture rather than a (previously over-claimed) Banach contraction. Second, a judge-efficient evaluation protocol: we generate each question's full trajectory once, replay every stopping policy over the identical drafts, and cache every LLM-judge call, yielding a strictly paired efficiency-versus-quality comparison at low cost; we further separate operational tokens (charged to a policy) from evaluation tokens (a measurement instrument). Third, an empirical study on multi-hop retrieval-augmented question answering (HotpotQA). On the 60-question test split, a judge-free semantic stopper reduces operational tokens by 38% relative to max_iterations at parity quality (Delta-IS = -0.004, p = 0.81), whereas the full quality-gated variant is counter-productive because its per-round judging dominates cost. An oracle that selects the best round attains +0.115 Information Score over every practical policy (p ~ 4e-11), reframing the problem from "when to stop" (easy) to "which round is best" (open).

13:00 JSTビジネス/資金調達

グラウンドトゥルースを使用してクラスタリングを評価するにはどうすればよいですか?

グランド トゥルースが利用可能な場合、外部インデックスをクラスター評価に使用できます。セットマッチングベースの尺度に焦点を当てて、最も一般的な外部妥当性指標をレビューします。セントロイド インデックス (CI) は、説明可能な結果が得られる直感的なクラスター レベルの測定であるため、推奨します。より細かく調整されたポイントレベルの測定が必要な場合は、より多くの選択肢があります。ペアセット インデックス (PSI) は、クラスター サイズによって偏らない正規化されたスコアを提供します。すべてのポイントが同等に重要である必要がある場合は、クラスタリング精度 (ACC) またはその他のセットマッチング尺度が適しています。

原文 (English)

How to evaluate clustering with ground truth?

External indexes can be used for cluster evaluation when ground truth is available. We review the most common external validity indexes focusing on set-matching-based measures. We recommend centroid index (CI), because it is an intuitive cluster-level measure with an explainable result. If we need a more fine-tuned, point-level measure, there are more choices. Pair-set index (PSI) provides a normalized score which is not biased by cluster sizes. If all points should matter equally, then clustering accuracy (ACC) or any other set-matching measure is suitable.

13:00 JSTLLM/生成AIエージェント

大規模言語モデルエージェントの経験則とポリシーの共同学習

マルチステップのインタラクティブ環境における LLM エージェントにとっての重要な課題は、蓄積されたインタラクション経験を効果的に活用することです。既存の研究では通常、そのようなエクスペリエンスを 2 つの使用法に分けています。1 つは、後でプロンプトを表示するための自然言語ルールとしてモデルの外に保持するか、軌道とフィードバックを使用してモデル パラメーターを更新するかです。前者は解釈しやすいですが、進化するポリシーと同期しなくなる可能性があります。後者はポリシーをより広範囲に改善しますが、スパース報酬設定における局所的な間違いに対する修正は限定的です。我々は、LLM エージェントのための経験的ルールとポリシーの共同学習 (JERP) を紹介します。これは、同じ対話の軌跡から長期的な経験的ルール プールとポリシーを更新します。意思決定時に、JERP はタスク関連のルールを取得し、対話履歴とともにエージェントにそれらのルールを条件付けします。各エピソードの後、収集された軌跡を使用してポリシーを最適化し、現在のロールアウトを参照の成功した軌跡と比較することでルール プールを修正します。この結合により、ルール プールが進化するポリシーに合わせて維持されると同時に、安定した効果的な動作がモデル自体に徐々に吸収されることが可能になります。 AlfWorld と WebShop での実験では、JERP が複雑な対話型タスクの意思決定パフォーマンスにおいて一貫した向上をもたらすことが示されています。

原文 (English)

Joint Learning of Experiential Rules and Policies for Large Language Model Agents

For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or using trajectories and feedback to update the model parameters. The former is easy to interpret but can fall out of sync with the evolving policy; the latter improves the policy more broadly but provides only limited correction for local mistakes in sparse-reward settings. We present Joint Learning of Experiential Rules and Policies for LLM Agents (JERP), which updates a long-term experiential-rule pool and the policy from the same interaction trajectories. At decision time, JERP retrieves task-relevant rules and conditions the agent on them together with the interaction history. After each episode, it uses the collected trajectories both to optimize the policy and to revise the rule pool by comparing current rollouts with reference successful trajectories. This coupling keeps the rule pool aligned with the evolving policy while allowing stable and effective behaviors to be gradually absorbed into the model itself. Experiments on AlfWorld and WebShop show that JERP yields consistent gains in decision performance for complex interactive tasks.

13:00 JSTLLM/生成AIエージェント

OpenRCA 2.0: 結果ラベルから因果関係プロセスの監視まで

根本原因分析 (RCA) では、長いコンテキストの理解、複数ステップの推論、ツールの使用など、LLM エージェントの機能の総合的なテストが行​​われます。ただし、既存のデータセットには根本的なギャップがあります。つまり、根本原因のみにラベルが付けられ、観察された症状につながる伝播経路はラベル付けされないため、単純なパターン マッチングのタスクが大幅に簡素化されます。厳密な評価をサポートするために、フォールト挿入による既知の介入を利用して因果伝播パスを再構築する段階的なラベル付けプロトコルである PAVE を導入します。このメカニズムは前方検証です。つまり、症状から逆方向に推論するのではなく、原因から結果に至るまで推論します。 PAVE を適用すると、LLM エージェントに対する段階的な因果的アノテーションを備えた最初のクロスシステム RCA ベンチマークである OpenRCA 2.0 (500 インスタンス) が生成されます。 11 のフロンティア LLM 全体で、正確な根本原因セットの回復に成功するのは、平均して 20.7% のケースのみです。この困難がどこにあるのかを特定するために、基準を緩和して、根拠のない診断と呼ばれるものを見つけます。エージェントは、ケースの 76.0% で少なくとも 1 つの正しい根本原因サービスを特定しますが、そのサービスを観察された症状への検証された因果伝播経路に根拠付けるのは 61.5% のみです。結果のみの評価では、この失敗モードが隠蔽されます。段階的因果的グラウンドトゥルースは、信頼できる LLM ベースの RCA エージェントに欠けている部分です。

原文 (English)

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed symptom, which largely simplifies the task to naive pattern matching. To support rigorous evaluation, we introduce PAVE, a step-wise labeling protocol that leverages known interventions from fault injection to reconstruct causal propagation paths. The mechanism is forward verification: reasoning from cause to effect rather than inferring backward from symptoms. Applying PAVE yields OpenRCA 2.0 (500 instances), the first cross-system RCA benchmark with step-wise causal annotations for LLM agents. Across 11 frontier LLMs, recovering the exact root-cause set succeeds in only 20.7% of cases on average. To locate where this difficulty lies, we relax the criterion and find what we call the ungrounded diagnosis: agents identify at least one correct root-cause service in 76.0% of cases, but ground that service in a verified causal propagation path to the observed symptom in only 61.5%. Outcome-only evaluation hides this failure mode; step-wise causal ground truth is the missing piece for trustworthy LLM-based RCA agents.

13:00 JSTLLM/生成AI

TOPS: 効率的な MLLM 推論のためのトークン最適保存セットの構築による第一原理のビジュアル トークン プルーニング

マルチモーダル大規模言語モデル (MLLM) は強力なマルチモーダル推論機能を実現していますが、その効率は大量のビジュアル トークンによって制限され、これによりかなりの計算オーバーヘッドが発生します。視覚的なトークン プルーニングは自然な解決策を提供しますが、既存の方法は不完全です。注意ベースの基準は冗長なトークンを保持する傾向がありますが、多様性ベースの基準はユーザーの指示に依存しないことがよくあります。複数の基準を組み合わせた方法であっても、トークン プルーニングの本質的な目的を原理的に定式化することがまだできていません。このペーパーでは、第一原理の観点から視覚的なトークン プルーニングを再検討し、それをトークン最適保存セットの構築として定式化します。トップダウンの情報理論分析を通じて、効果的なトークン選択のための 3 つの基本原則、つまりタスクの関連性、情報の網羅性、およびセマンティックの多様性を特定します。これらの原則に基づいて、さまざまな MLLM に適用できる、トレーニング不要でモデルに依存しない枝刈りモジュールである TOPS を提案します。 7 つの MLLM バックボーンと 14 のベンチマークに関する広範な実験により、TOPS がさまざまなプルーニング設定の下で従来の方法よりも優れたパフォーマンスを発揮することが実証されました。特に、LLaVA-NeXT では、TOPS は 7B モデルと 13B モデルでそれぞれ 100.0% と 100.6% のパフォーマンスを維持しながら、ビジュアル トークンの 77.8% を削除します。これは、冗長なビジュアル トークンを削除することで幻覚を緩和し、将来の軽量 MLLM 設計を刺激できる可能性があることを示唆しています。

原文 (English)

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference

Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introduces substantial computational overhead. Visual token pruning offers a natural solution, yet existing methods are imperfect: attention-based criteria tend to retain redundant tokens, while diversity-based criteria are often agnostic to user instructions. Even methods that combine multiple criteria still lack a principled formulation of the intrinsic objective of token pruning. In this paper, we revisit visual token pruning from a first-principles perspective and formulate it as constructing Token Optimal Preservation Sets. Through a top-down information-theoretic analysis, we identify three fundamental principles for effective token selection: Task Relevance, Information Coverage, and Semantic Diversity. Based on these principles, we propose TOPS, a training-free and model-agnostic pruning module that can be applied to various MLLMs. Extensive experiments on 7 MLLM backbones and 14 benchmarks demonstrate that TOPS outperforms prior methods under diverse pruning settings. Notably, on LLaVA-NeXT, TOPS removes 77.8% of visual tokens while preserving 100.0% and 100.6% performance on its 7B and 13B models, respectively, suggesting that pruning redundant visual tokens can sometimes mitigate hallucination and inspire future lightweight MLLM design.

13:00 JSTエージェント

レガシー ワークフローを Agentic BPM に引き上げるプロセス ハーネス: CUGA FLO での設計と実現

基盤となるワークフロー エンジンを置き換えることなく、レガシー ワークフローをエージェントティック ビジネス プロセス管理 (エージェント BPM) に引き上げる新しいメカニズムであるプロセス ハーネスを導入します。プロセス ハーネスは、決定論的ワークフロー エンジンの周囲にポリシーで管理されるエージェント層を配置し、エンジンがプロセスに対する構造的な権限を保持しながら、指定された制御ポイントをインターセプトして推論、適応、監視に貢献します。プロセス ハーネスを厳密に定義するために、データ スキーマと実行セマンティクスの両方を指定するタスク決定フロー (TDF) モデルを開発します。 TDF は、3 つのポリシー管理エージェント タイプにわたる LLM 推論を分解します。1 つは知識集約型タスクの実行のための TaskAgent、1 つはケースごとのゲートウェイ ルーティングのための DecisionAgent、そして原則に基づいたフック メカニズムを通じてランタイム フローの適応を管理する FlowAgent です。各エージェントは、システム内のすべての LLM 呼び出しを管理する集約ポリシー セットであるプロセス FRAME から抽出された明示的なポリシー内で判断します。次に、TDF モデルの設計と実装の実現として CUGA FLO を示し、3 つのエージェント タイプすべてとフック主導の規制オーバーライドを実行するローン承認ワークフローでそれを実証します。プロセスハーネスは、構造的コンプライアンスを強制する決定論的なワークフローの実行を通じて実現される命令的要件と、プロセスが要求するどこにでも指定された制御ポイントで呼び出されるポリシーに基づいたエージェントの自律性を通じて実現される規範的要件を独自に調和させます。

原文 (English)

A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO

We introduce the process harness, a new mechanism for uplifting legacy workflows into Agentic Business Process Management (Agentic BPM) without replacing the underlying workflow engine. A process harness places a policy-governed agentic layer around a deterministic workflow engine, intercepting designated control points to contribute reasoning, adaptation, and oversight while the engine retains structural authority over the process. To define the process harness rigorously, we develop the Task-Decision-Flow (TDF) model, specifying both its data schema and its execution semantics. TDF decomposes LLM reasoning across three policy-governed agent types: a TaskAgent for knowledge-intensive task execution, a DecisionAgent for per-case gateway routing, and a FlowAgent that governs runtime flow adaptation through a principled hook mechanism. Each agent reasons within an explicit policy drawn from the process FRAME, the aggregate policy set governing all LLM calls in the system. We then present CUGA FLO as the design and implementation realization of the TDF model, and demonstrate it on a loan approval workflow that exercises all three agent types and hook-driven regulatory override. The process harness uniquely reconciles imperative requirements, realized through deterministic workflow execution that enforces structural compliance, with normative requirements, realized through policy-framed agentic autonomy invoked at designated control points wherever the process demands it.

13:00 JST研究/論文

進化的に生成された敵対的テキストに対する自然言語分類子の脆弱性

深層学習モデルは、さまざまな分野で目覚ましいパフォーマンスを達成していますが、特に NLP では、敵対的な入力に対して脆弱なままであり、そのような攻撃は現実世界に重大な影響を与える可能性があります。敵対的攻撃には、NLP モデルをだますために、意味的に類似した小規模なトークン置換が含まれることが多く、最近の手法は、多くの場合、モデルの内部構造へのある程度のアクセスを悪用して、特定の脆弱な単語をターゲットにすることで、より正確になっています。この論文では、自然言語モデルに対する敵対的攻撃を生成するハイブリッド遺伝アルゴリズム (GA) である GAversary を提案します。 GA はターゲット モデルをブラック ボックスとして扱うことができ、検索のガイドとしてモデルが出力するロジット値のみを必要とします。 GAversary は、GloVe 埋め込みを使用して単語の置換 (突然変異演算子) を提案し、敵対的な例の意味上の類似性を改善するという点で、この問題に対して以前に提案された GA とは異なります。 GAversary は、いくつかのベンチマーク データ セットとよく知られたターゲット モデルに適用されます。 GAversary は、BAE および A2T 攻撃と比較して、テスト データに対するターゲット モデルの精度を大幅に低下させることができます (最良のケースでは、BAE の 27.6% と比較して、76.8% の精度が 5.8% に低下します)。トレードオフとして、GAversary は他の 2 つの方法に比べて 2 倍弱の単語を摂動させますが、元のテキストとの意味上の類似性はわずかに低くなり、実行時間は約 5% 増加します。

原文 (English)

Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text

Deep learning models have achieved impressive performance across various fields but remain vulnerable to adversarial inputs, particularly in NLP, where such attacks can have significant real-world consequences. Adversarial attacks often involve small, semantically similar token replacements to fool NLP models, and recent methods have become more precise by targeting specific vulnerable words, often by exploiting some level of access to the model's internal structure. This paper proposes GAversary, a hybrid Genetic Algorithm (GA) to generate adversarial attacks on natural language models. The GA is able to treat the target model as a black box, requiring only the logit value output by the model to guide the search. GAversary differs from GAs previously proposed for this problem by using GloVe embeddings to propose word replacements (the mutation operator) to improve the semantic similarity of the adversarial examples. GAversary is applied to several benchmark data sets and well-known target models. GAversary is able to substantially reduce the target model's accuracy on test data compared to the BAE and A2T attacks compared against (in the best case, reducing a 76.8% accuracy to 5.8%, compared to BAE's 27.6%). The trade-off is that GAversary perturbs just under twice as many words as the other two methods, with a slightly lower semantic similarity to the original text and around a 5% increase in run-time.

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達

判断せずに質問する: 解釈可能な LLM 評価と自己改善のための 2 つの質問

NLP では、LLM 出力の評価が依然として大きなボトルネックとなっています。人間による評価は高価で時間がかかり、語彙メトリクスとオープンエンド生成に関する人間の判断との相関性が低く、全体的な LLM ジャッジはデバッグが難しい不透明なスコアを生成することがよくあります。私たちは、評価基準をアトミックなバイナリの質問に分解し、その結果の判定を解釈可能な多次元スコアに集約するフレームワークである BINEVAL を提案します。タスク プロンプトが与えられると、メタ プロンプトが詳細な評価質問を生成し、LLM が出力ごとに独立して質問に回答し、調整された全体スコアとともに透明な質問レベルのフィードバックを生成します。この分解により、評価が検査、​​診断が容易になり、迅速な改善に直接使用できるようになります。 SummEval、Topical-Chat、QAGS 全体で、BINEVAL は UniEval や G-Eval などの強力なベースラインと同等またはそれを上回り、特に QAGS などの事実整合性ベンチマークで優れた結果を示しています。 BINEVAL は、人間の判断との競合相関を超えて、人間のスコア分布とよりよく一致し、以前の LLM ジャッジによく見られた天井効果を回避し、境界線にある出力と明らかに欠陥のある出力をより適切に区別することにつながります。さらに、同じ質問レベルのフィードバックが反復プロンプトの最適化をサポートし、自己更新設定とクロスモデル更新設定の両方で IFBench での要約に関する評価者のプロンプトと生成プロンプトを改善することを示します。全体として、BINEVAL は、強力な経験的パフォーマンスと実用的な診断および最適化の価値を組み合わせた、タスクに依存せず、トレーニング不要で、解釈可能な評価フレームワークを提供します。

原文 (English)

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.

13:00 JST画像/動画生成

EO-WM: 確率論的地球観測予測のための物理情報に基づいた世界モデル

地球観測 (EO) 予測は、変化する気象条件下での衛星観測から将来の地球表面のダイナミクスを予測することを目的としています。この論文では、このタスクを、部分的に観測された気象主導型の世界モデリング問題とみなします。この問題では、気象は条件付け信号として機能しますが、観測がまばらで地表面の状態が観測されていないため、予測は依然として不確実です。しかし、既存の手法はこの設定を完全には捉えていません。決定論的モデルは不確実性を 1 つの将来予測にまとめますが、拡散ベースの手法は通常、気象変数を未分化の条件付け信号として扱い、既存のベンチマークは、予報が変化する気象強制に正しく反応するかどうかではなく、主に再構成精度に焦点を当てています。マルチスペクトル EO 予測用のビデオ拡散トランスフォーマである EO-WM を紹介します。 EO-WM には、気候ベースライン、気象異常、累積的な物理的ストレス信号による気象強制力を表す、物理的情報に基づいた調整フレームワークが組み込まれています。具体的には、明確な条件付け経路を通じてベースラインと異常を分離し、持続的な熱と干ばつストレスを捕捉するために時間の経過とともに異常な強制力を蓄積します。標準的な指標を超えて気象応答動作を評価するために、2 つの診断ベンチマークを導入します。1 つは、異常気象下での植生劣化の深刻度を意識した予測のための極端な夏季ベンチマークで、もう 1 つは、変化する気象強制下での応答忠実度をテストするための季節一致ペア ベンチマークです。実験の結果、EO-WM は、標準的なピクセル レベルのメトリクスで競争力を維持しながら、予測される正規化植生指数 (NDVI) の減少振幅の誤差を相対的に 5.63% 削減し、方向性ヒット率を相対的に 7.80% 改善することが示されています。ベンチマークとモデルは https://github.com/Luo-Z13/EO-WM でオープンソース化されます。

原文 (English)

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as undifferentiated conditioning signals, and existing benchmarks focus mainly on reconstruction accuracy rather than whether forecasts respond correctly to changed weather forcing.We introduce EO-WM, a video diffusion transformer for multispectral EO forecasting. EO-WM incorporates a physically informed conditioning framework that represents meteorological forcing through a climatological baseline, weather anomalies, and cumulative physical stress signals. Specifically, it separates baseline and anomaly through distinct conditioning pathways, and accumulates anomalous forcing over time to capture sustained heat and drought stress. To evaluate weather-response behavior beyond standard metrics, we introduce two diagnostic benchmarks: an Extreme Summer Benchmark for severity-aware prediction of vegetation degradation under extreme weather, and a Seasonal Matched-Pair Benchmark for testing response fidelity under changed weather forcing. Experiments show that EO-WM reduces the error in predicted Normalized Difference Vegetation Index (NDVI) decline amplitude by a relative 5.63% and improves directional hit rate by a relative 7.80%, while remaining competitive on standard pixel-level metrics. The benchmarks and model will be made open-source at https://github.com/Luo-Z13/EO-WM.

13:00 JST研究/論文

疫学モデルにおける迅速なベイズパラメータ推定のためのシミュレーションベースの推論: MCMC との比較

機械的疫学モデルは、感染症の予測と公衆衛生の意思決定をサポートするために広く使用されています。このようなモデルのベイジアン キャリブレーションは、マルコフ連鎖モンテカルロ (MCMC) を使用して実行されるのが一般的ですが、高次元の非線形システムや繰り返されるほぼリアルタイムの解析では、計算コストが高くなる可能性があります。ここでは、2020年のドイツの新型コロナウイルス感染症集中治療室(ICU)占有率データを使用した機械的SECIR疫学モデルのベイジアンキャリブレーションのスケーラブルな代替手段として、神経事後推定を使用したシミュレーションベース推論(SBI)を調査します。31日間の推論ウィンドウと、複数の伝達変化点を含む実質的により困難な201日間の再構成問題の両方を使用して、複数の流行期にわたってSBIとMCMCを比較しました。事後一致は、ワッサーシュタイン距離とカルバック・ライブラー発散を事後予測チェックとともに使用して定量的に評価されました。 31 日間のウィンドウにわたって、SBI は、観察された ICU 軌跡を正確に再現しながら、MCMC と強く一致する事後分布を回復しました。 201 日の設定では、不確実性が増大したにもかかわらず、SBI は支配的な後方構造を保存しました。 SBI は、CPU と GPU リソースを組み合わせることで、CPU 上での実行に制限されていた MCMC と比較して、計算実行時間を大幅に短縮しました。 MCMC では 31 日間の推論問題に約 1000 秒を要しましたが、SBI では単一の GPU で約 60 ~ 70 秒で同等の事後予測パフォーマンスを達成しました。 201 日の推論問題の場合、SBI では平均 157 秒かかりましたが、MCMC の実行には 19,000 秒以上かかりました。私たちの結果は、SBI が機構的疫学モデルのベイジアン校正のための迅速かつ計算効率の高いフレームワークを提供し、ほぼリアルタイムの推論と迅速なアウトブレイク分析の繰り返しをサポートすることを示しています。

原文 (English)

Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC

Mechanistic epidemiological models are widely used to support infectious disease forecasting and public-health decision making. Bayesian calibration of such models is commonly performed using Markov chain Monte Carlo (MCMC), which can become computationally expensive for high-dimensional nonlinear systems and repeated near-real-time analyses. Here, we investigate simulation-based inference (SBI) using neural posterior estimation as a scalable alternative for Bayesian calibration of a mechanistic SECIR epidemiological model using COVID-19 intensive care unit (ICU) occupancy data from Germany during 2020. We compared SBI and MCMC across multiple epidemic phases using both 31-day inference windows and a substantially more challenging 201-day reconstruction problem involving multiple transmission change points. Posterior agreement was evaluated quantitatively using Wasserstein distances and Kullback-Leibler divergences together with posterior predictive checks. Across the 31-day windows, SBI recovered posterior distributions in strong agreement with MCMC while accurately reproducing observed ICU trajectories. In the 201-day setting, SBI preserved the dominant posterior structure despite increased uncertainty. SBI, by combining CPU and GPU resources, substantially reduced computational runtime compared with MCMC, which was restricted to running on CPUs. Whereas MCMC required approximately 1000 seconds for the 31-day inference problems, SBI achieved comparable posterior and predictive performance in approximately 60-70 seconds on a single GPU. For the 201-day inference problem, SBI required an average of 157 seconds, while the MCMC runs took over 19,000 seconds. Our results demonstrate that SBI provides a rapid and computationally efficient framework for Bayesian calibration of mechanistic epidemiological models, supporting repeated near-real-time inference and rapid outbreak analysis.

13:00 JSTLLM/生成AI

大規模な言語モデルを使用した自動 R\'esum\'e スクリーニングでの迅速なインジェクション: シングルおよびマルチインジェクション設定

大規模言語モデル (LLM) は、求職者のスクリーニングとランク付けにますます使用されており、候補者がアルゴリズム採用システムを戦略的に操作するインセンティブが生まれています。私たちは、自動化された履歴書スクリーニングにおける即時注入について研究しています。これは、新しい資格を導入するものではありませんが、LLM 評価に影響を与えるように設計された微妙な自己宣伝テキストとして定義されます。対照実験を使用して、論文の質が均一で、注入する候補者がほとんどいない場合、即時注入により確実に応募者のランキングが向上することが示されました。しかし、その有効性は、より多くの候補者が注入されるにつれて急速に減少し、操作が広範囲に及ぶと崩壊します。候補者の品質が不均一な場合、プロンプト注入の効果は平均して低くなりますが、場合によっては、低品質の候補者が高品質の候補者を上回ってしまう可能性があり、公平性に関する懸念が生じます。全体として、LLM ベースのスクリーニングは、操作がまれで候補者の品質の差が小さい場合に最も脆弱になります。コードとリソースは、https://github.com/preetb1199/Prompt_Injection_ACL26 で公開されています。

原文 (English)

Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and Multi-Injection Settings

Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candidates to strategically manipulate algorithmic hiring systems. We study prompt injection in automated r\'esum\'e screening, defined as subtle self-promotional text that introduces no new qualifications but is designed to influence LLM evaluations. Using controlled experiments, we show that prompt injection reliably improves applicant rankings when r\'esum\'e quality is homogeneous and few candidates inject. However, its effectiveness rapidly diminishes as more candidates inject, collapsing when manipulation becomes widespread. When candidate quality is heterogeneous, prompt injection is less effective on average, but can occasionally allow lower-quality candidates to outrank higher-quality ones, raising fairness concerns. Overall, LLM-based screening is most vulnerable when manipulation is rare and candidate quality differences are small. Code and resources are publicly available at: https://github.com/preetb1199/Prompt_Injection_ACL26

13:00 JSTLLM/生成AIエージェント

言語モデルの結合が役立つのはどのような場合ですか? 67 のフロンティア モデルにわたるルーティング、投票、およびエージェントの混合における同時障害の上限

ルーティング、投票、カスケード、フュージョン、エージェントの混合などのマルチモデル LLM システムは、単一モデルの精度を上回るために使用されます。私たちは、彼らの利益が、現場でほとんど報告されない量によって制限されていることを示します。出力が 1 つのメンバー モデルの回答であるポリシーの場合、精度は 1 マイナス ベータを超えることはできません。ベータとは、同じクエリに対してすべてのモデルが誤る率です。対照的に、通常の診断である平均ペアワイズ誤差相関ρはベータを識別できません。同一の周辺値とペアワイズ相関を持つ誤差則は、全誤り率が異なる可能性があります。ベータ版の Clopper-Pearson バウンドは、ルーターをトレーニングする前に、ルーター、投票、またはカスケードが提供できる最大のゲインに基づいて有限サンプル証明書を提供します。 21 社のプロバイダーの 67 モデルにわたって、テトラコールで校正された単一因子モデルは依然として間違った尾部の価格を下回っています。オープンエンド数学では、観測されたベータ値は 0.052 であるのに対し、完全な 67 モデルのガウス コピュラでは 0.023 であり、約 2.5 倍の割安であり、90 パーセント CI は 1.7 ~ 3.4、k は 17 に相当します。効果は再発します。実行グレード コードの場合、ベータは 0.079 です。同じ GPQA-Diamond の質問を多肢選択形式ではなく自由回答形式で再質問すると、ベータ 0.127、カッパ 0.73 ~ 0.92 の 5 人の裁判官パネルで尾部が再び開き、主題ではなく回答形式での共失敗箇所が特定されます。同等の品質では、低 rho の異種アンサンブルが高 rho の Self-MoA を上回りますが、プール内のチェック可能なタスクでは、強力なクエリ レベルのルーティング シグナルがなければ、モデルを組み合わせた方が単一の最良のモデルを上回ることはほとんどありません。利益は、モデルを追加することでではなく、さまざまな質問でモデルが失敗することから得られます。

原文 (English)

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain is capped by a quantity the field rarely reports. For any policy whose output is one member model answer, accuracy cannot exceed one minus beta, where beta is the rate at which every model is wrong on the same query. In contrast, the usual diagnostic, average pairwise error correlation rho, cannot identify beta: error laws with identical marginals and pairwise correlations can have different all-wrong rates. A Clopper-Pearson bound on beta gives a finite-sample certificate on the largest gain any router, vote, or cascade could deliver before training a router. Across 67 models from 21 providers, a tetrachoric-calibrated single-factor model still underprices the all-wrong tail: on open-ended mathematics, observed beta is 0.052 versus 0.023 under the full 67-model Gaussian copula, about 2.5 times underpricing, with 90 percent CI 1.7 to 3.4 and k equals 17. The effect recurs on execution-graded code, where beta is 0.079. Re-asking the same GPQA-Diamond questions in free-response rather than multiple-choice form reopens the tail, with beta 0.127 and a five-judge panel with kappa 0.73 to 0.92, locating co-failure in answer format rather than subject. At matched quality, low-rho heterogeneous ensembles beat high-rho Self-MoA, but on checkable tasks in our pool, combining models rarely beats the single best model without a strong query-level routing signal. Gains come from models failing on different questions, not from adding more models.

13:00 JST研究/論文GPT / ChatGPT

高齢者の認知支援のための言語ベースのデジタルツイン

デジタルツインは、個人の行動と健康の軌跡のモデリングを可能にする、パーソナライズされたヘルスケアの有望なパラダイムとして浮上しています。認知の健康においては、言語や会話のパターンが非侵襲的なバイオマーカーとして機能する軽度認知障害 (MCI) の早期発見が依然として困難です。この研究では、大規模言語モデル (LLM) を活用して、スチロ測定の合図とコンテキスト メタデータを組み込むことで高齢者の会話行動を模倣する、言語ベースのデジタル ツイン フレームワークを提案します。忠実度と認知的一貫性を評価するために、再構成の品質を共同で測定し、認知スコアを予測するマルチヘッド条件変分オートエンコーダー (cVAE) を導入します。 I-CONECT データセットの実験では、デジタル ツインがアイデンティティ固有の特性を保持し、実際のデータに匹敵する再構成誤差と MoCA 予測誤差を達成しながら、ベースライン GPT で生成された応答を上回るパフォーマンスを示していることが示されています。これらの結果は、パーソナライズされた継続的な認知健康状態モニタリングのためのスケーラブルで非侵襲的なアプローチとしての言語ベースのデジタル ツインの可能性を浮き彫りにしています。

原文 (English)

Language-Based Digital Twins for Elderly Cognitive Assistance

Digital twins have emerged as a promising paradigm for personalized healthcare, enabling modeling of individual behavior and health trajectories. In cognitive health, early detection of Mild Cognitive Impairment (MCI) remains challenging, where language and conversational patterns serve as non-invasive biomarkers. In this work, we propose a language-based digital twin framework that leverages large language models (LLMs) to mimic the conversational behavior of elderly individuals by incorporating stylometric cues and contextual metadata. To evaluate fidelity and cognitive consistency, we introduce a multi-head conditional variational autoencoder (cVAE) that jointly measures reconstruction quality and predicts cognitive scores. Experiments on the I-CONECT dataset show that the digital twin preserves identity-specific characteristics and achieves reconstruction and MoCA prediction errors comparable to real data, while outperforming baseline GPT-generated responses. These results highlight the potential of language-based digital twins as a scalable and non-invasive approach for personalized and continuous cognitive health monitoring.

13:00 JSTLLM/生成AI研究/論文

グローバル AI 技術ガバナンスのためのオープンウェイト基盤モデルのベンチマーク

大規模言語モデル (LLM) は、国内および国際組織全体で人工知能 (AI) ガバナンス分析に導入されることが増えています。しかし、そのようなモデルでは、トレーニング データで過小評価されている国に対して、著しく精度の低い応答が生成されるという証拠が増えています。このパターンは、既存の文献で地理的偏りとして説明されています。この現象を調査している既存の研究には、その結果を損なう 3 つの方法論的な制限があります。(1) 重みが公表されていない独自のシステムに依存しているため、独立した複製が妨げられています。 (2) モデルトレーニングのためのデータ収集が終了した後、各モデルの知識の自然な限界に加えて地理的な無知につながる、数年間のモデル知識の評価。 (3) モデルの信頼できる製造 (HF) と不確実性の正直な認識を区別できない、粗い二値応答分類の使用。この研究では、2026 年 1 月に Harvard Dataverse で公開された 227 か国の 24,453 指標の検証済みグラウンドトゥルース データベースである Global AI Dataset v2 (GAID v2) に対して 4 つのオープンウェイト フロンティア言語モデルをベンチマークすることで、3 つの制限すべてに対処しています。IEEE IRAI 2026 フレームワークの 8 つのテーマの次元にマッピングされた合計 18 の指標が GAID v2 から選択され、およその結果が得られます。 6 つの評価年 (2010 年から 2023 年の期間内) にわたる 2,990 の国単位のメートル年の観測。モデル応答は、(a) 検証済み精度 (VA)、(b) HF、(c) 正直な拒否 (HR)、(d) 定性的ヘッジ (QH)、および (e) 誤った帰属 (MF) を区別する 5 つのカテゴリー スキームを使用して分類されます。精度の地理的差異は、混合効果ロジスティック回帰および差分差分 (DiD) 分析を通じて推定されます。

原文 (English)

Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations. There is, however, growing evidence that such models produce significantly less accurate responses for countries that are underrepresented in their training data-a pattern described in existing literature as geographic bias. Existing studies examining this phenomenon are subject to three methodological limitations that together undermine their findings: (1) reliance on proprietary systems whose weights are not publicly released, which prevents independent replication; (2) evaluation of model knowledge about years that fall after data collection for model training had concluded, leading to geographic ignorance in addition to the natural limits of each model's knowledge; and (3) use of coarse binary response classification that cannot distinguish models' confident fabrication (HF) from their honest acknowledgement of uncertainty. This study addresses all three limitations by benchmarking four open-weight frontier language models against the Global AI Dataset v2 (GAID v2), a verified ground-truth database of 24,453 indicators across 227 countries published on Harvard Dataverse in January 2026. A total of 18 indicators, mapped to the eight thematic dimensions of the IEEE IRAI 2026 framework, are selected from GAID v2, yielding approximately 2,990 country-metric-year observations across six evaluation years (within the period of 2010-2023). Model responses are classified using a five-category scheme that distinguishes (a) verified accuracy (VA), (b) HF, (c) honest refusal (HR), (d) qualitative hedging (QH), and (e) misattribution (MF). Geographic disparities in accuracy are estimated through mixed-effects logistic regression and difference-in-differences (DiD) analysis.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Llama

Know2Guess: 大規模言語モデルにおける知識境界評価のための汚染を認識したマルチゾーン ベンチマーク

大規模な言語モデルの信頼性の高い評価では、データの汚染、プロンプトの特異性、または一般的な拒否行動と混同することなく、サポートされている回答とサポートされていない推測を分離する必要があります。凍結されたビルドタイム ラベルの下で、回答可能な知識から棄権が期待される未知への移行を測定するための、汚染を認識したマルチゾーン ベンチマークを示します。このベンチマークには、5 つのドメインにわたる 1,200 項目、明示的な棄権期待、汚染リスクのメタデータ、および公式の厳密なパーサーと正規化された堅牢性パーサーによる二重解析が含まれています。ロックされた回答または棄権プロンプト、回答のみのコントロール、およびプロンプト テンプレートのバリアントの下で、FLAN-T5、Qwen2.5-Instruct、および Llama-3-Instruct モデルを評価します。このベンチマークは、一般的な無回答行動では解決されません。FLAN のベースラインは、生産的な棄権に関しては弱いままですが、より強力な指導調整モデルは、選択的ではあるが回答から棄権への移行が不完全であることを明らかにしています。 Qwen2.5-3B-Instruct は全体的に最高の信頼性を実現していますが、回答が期待されるゾーンは依然として難しく、キャリブレーションは依然として不十分で、良性の項目の拒否は引き続き発生します。プロンプトおよびパーサーの堅牢性分析により、主要なランキングと定性的な結論が維持されます。したがって、このベンチマークは、回答可能性、棄権、拒否、および汚染を、LLM の信頼性の個別だが相互作用する側面として監査するための再現可能なプロトコルを提供します。データセットは、https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark で公開されています。

原文 (English)

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.

13:00 JSTLLM/生成AILlama

役に立ちます: トレーニング中期の思いやりの値はトレーニング後にドメインに依存して低下します

標準的なポストトレーニング パイプラインは、教師あり微調整 (SFT) と強化学習 (RL) を適用して言語モデルを有用にしますが、これらのプロセスは、トレーニング前に注入された値を誤って低下させる可能性があります。動物危害ベンチマーク (AHB 2.2) と MORU ベンチマークで評価された SFT (Dolly-15k による有用性と Magicoder-110K によるコーディング) と GRPO (RLHFlow による有用性と Magicoder によるコーディング) の両方を使用して、思いやり指向の合成データで中間トレーニングされた Llama 3.1 8B モデルにおける動物の思いやりの値の保持に、トレーニング後のデータのドメインが差動的に影響を与えるかどうかを調査します。 (不確実性の下での道徳的推論)。有用性トレーニングは、AHB でのコーディング トレーニングと比較して動物の思いやりを大幅に低下させます (SFT: 35.7% 対 65.2%、GRPO: 18.7% 対 32.0%)。これは 2 つの独立した有用性データセットと 2 つのトレーニング パラダイムにわたって再現されています。英語のMORU項目では、有用性トレーニングは一般的な道徳的推論を25.5パーセントポイント(46.4%対71.9%)低下させ、その大きさは同情効果に匹敵する顕著な差でした。ただし、この効果は言語を越えて伝わりません。多言語の MORU ベンチマークでは、ドメイン効果は消失します (SFT: 52.3% 対 51.2%)。対照的に、動物の思いやりの効果は言語間で一貫して伝わり、Magiccoder の基本モデルに対する AHB パーセンテージ ポイントの増加は、英語以外の項目では英語の項目よりも 4.5 倍大きくなっています。この乖離は、トレーニング中に教え込まれた価値観が、ドメイン固有のトレーニング後の改善を推論するよりも深く、言語を超えてコード化されていることを示唆しています。これらの結果は、価値を満載したトレーニング途中で構築するラボの場合、トレーニング後の有用性よりも、トレーニング後のコーディング ドメインの方が、一般的な推論能力を損なうことなく、トレーニング途中の値をよりよく保存できる可能性があることを示唆しています。

原文 (English)

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, using both SFT (helpfulness via Dolly-15k vs. coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs. coding via Magicoder), evaluated on the Animal Harm Benchmark (AHB 2.2) and MORU benchmark (Moral Reasoning Under Uncertainty). Helpfulness training significantly degrades animal compassion relative to coding training on AHB (SFT: 35.7% vs. 65.2%; GRPO: 18.7% vs. 32.0%), replicating across two independent helpfulness datasets and two training paradigms. On English MORU items, helpfulness training degrades general moral reasoning by 25.5 percentage points (46.4% vs. 71.9%), a striking gap that rivals the compassion effect in magnitude. However, this effect does not transfer cross-lingually: on the multilingual MORU benchmark, the domain effect disappears (SFT: 52.3% vs. 51.2%). In contrast, the animal compassion effect transfers consistently across languages, with Magicoder's AHB percentage-point gain over the base model 4.5 times larger on non-English items than English items. This divergence suggests that values instilled through mid-training are encoded more deeply and cross-lingually than reasoning improvements from domain-specific post-training. These results suggest that, for labs building on value-laden mid-training, coding-domain post-training may better preserve mid-trained values than helpfulness post-training without harming general reasoning capabilities.

13:00 JSTLLM/生成AIGPT / ChatGPT

LLM の問題解決能力の調査 -- 静的質問に関する研究

大規模言語モデル (LLM) は、幅広い主題にわたる課題や試験を完了する実証済みの能力により、社会の多くの側面、特に教育に急速に影響を与えています。これまでの研究では LLM の教育的影響が調査されてきましたが、既存の研究の多くは公開またはオープンな問題データセットに依存しており、トピック固有の分析が不足しています。工学教育、特に機械工学における、特定の問題タイプに対する LLM パフォーマンスの系統的な調査は依然として限られています。 LLM ツールに教科書の質問を直接尋ねる従来の方法を使用する代わりに、私たちの研究ではモデル蒸留プロセスを採用して、静的問題を解決する際の LLM 機能を評価しました。 ChatGPT を抽出することにより、25 のテキストのみの静的質問を抽出し、さらに図を追加して数値を変更することで 2 つの追加のデータセットを構築しました。実験結果によると、LLM はテキストのみの静的問題では良好なパフォーマンスを示しますが、図が導入され、問題に複数のステップの推論が必要になると精度が低下します。さらなる分析によると、このパフォーマンスの低下は主に画像認識の制限が原因ではなく、むしろ複数ステップの推論と、抽出された視覚情報を連続したソリューション段階全体に一貫して適用することの難しさによって引き起こされていることが示唆されています。

原文 (English)

Investigating LLM's Problem Solving Capability -- a Study on Statics Questions

Large Language Models (LLMs) have rapidly influenced many aspects of society, particularly education, due to their demonstrated ability to complete assignments and examinations across a wide range of subjects. Although prior studies have examined the educational impact of LLMs, much of the existing work relies on public or open problem datasets and lacks topic-specific analysis. In engineering education, especially within mechanical engineering, systematic investigations of LLM performance on specific problem types remain limited. Instead of using traditional methods that directly ask textbook questions to an LLM tool, our study adopts a model distillation process to evaluate LLM capabilities in solving statics problems. By distilling ChatGPT, we extracted 25 text-only statics questions and further constructed two additional datasets by adding diagrams and modifying their numerical values. Experimental results show that while LLMs perform well on text-only statics problems, their accuracy decreases when diagrams are introduced and the problems require multi-step reasoning. Further analysis suggests that this performance drop is not primarily caused by limitations in image recognition, but rather by difficulties in multi-step reasoning and in consistently applying extracted visual information across successive solution stages.

13:00 JSTLLM/生成AILlama

主張し、説明しないでください: 動物福祉に関する LLM の推論を変える言語的特徴

動物愛護活動家たちは多くの著作物を作成しており、その著作物が言語モデルを訓練し、その後何百万人もの人々が動物福祉について尋ねるようになっています。提示された動物福祉ベンチマークで語彙を一致させたスタンスコントラストプローブを使用して、10の言語的特徴のそれぞれが、微調整データとして使用された場合にラマ-3.2-1Bの動物福祉推進推論に対する好みをどのように変化させるかを測定します。 10 個の特徴のうち 8 個で、統計的に有意な変化が生じます。 7 つは、断定的な確実性、明確な道徳的語彙、感情的な言葉、評価的主張、物語の構造、描写された危害の深刻度、即時的な時間的枠組みなど、モデルをより強力な動物愛護推進の推論に向けて移行させています。 2 つはそれを逆方向に動かします。ヘッジされた言葉と具体的な感覚的説明は両方とも動物愛護推進の立場を薄めます。一人称視点には統計的に有意な効果はありません。 LLM トレーニング コーパスに組み込まれる可能性のある動物福祉に関するテキストを執筆する人に対する実際的な推奨事項: シーンを中立的に説明するのではなく、立場を主張することです。モデルを変える特徴は、作家の立場を明確にするものです。それを弱める特徴は動物愛護の内容を保持しますが、スタンスを保留します。

原文 (English)

Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B's preference for pro-animal-welfare reasoning when used as fine-tuning data. Eight of the ten features produce statistically significant shifts. Seven move the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two move it the other way: hedged language and concrete sensory description both dilute the pro-animal-welfare stance. First-person perspective has no statistically significant effect. The practical recommendation for anyone writing animal-welfare text that may end up in LLM training corpora: assert a position rather than describe a scene neutrally. The features that shift the model are the ones that make the writer's position explicit; the features that dilute it hold animal-welfare content but withhold stance.

13:00 JSTLLM/生成AI

Long-Horizo​​n LLM 推論のためのコンテキストのリサイクル

大規模言語モデル (LLM) は、短いコンテキストの推論では強力な機能を示しますが、コンテキスト ウィンドウの制限と非効率なトークンの使用により、長い会話期間ではパフォーマンスが低下します。 ContextForge は、構造化クエリの生成、外部メモリの取得、制御された合成を組み合わせることにより、ターンをまたいでタスク関連情報を維持するコンテキスト リサイクル システムです。このシステムにより、完全なコンテキストの再生に依存せずに以前の計算を効率的に再利用できるため、応答の品質を維持しながらトークンのオーバーヘッドが削減されます。私たちは、構造化されたヘルスケア クエリ全体でマルチターン推論、後方参照、ドメイン シフトをテストする 15 ターンの会話ベンチマークを使用して ContextForge を評価します。同一の基礎モデルを使用するベースライン エージェントと比較して、ContextForge は同等の応答精度を維持しながら、一貫性の向上とトークン消費量の削減を実証します。これらの結果は、コンテキスト リサイクルが、より大きなコンテキスト ウィンドウやモデルの再トレーニングを必要とせずに、長期的なタスクで LLM 機能を拡張するための実用的なアプローチを提供することを示唆しています。コードと評価成果物は https://github.com/Betanu701/ContextForge で入手できます。

原文 (English)

Context Recycling for Long-Horizon LLM Inference

Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due to context window limitations and inefficient token usage. We introduce ContextForge, a system for context recycling that maintains task-relevant information across turns by combining structured query generation, external memory retrieval, and controlled synthesis. The system enables efficient reuse of prior computation without relying on full context replay, reducing token overhead while preserving answer quality. We evaluate ContextForge using a 15-turn conversational benchmark that tests multi-turn reasoning, back-references, and domain shifts across structured healthcare queries. Compared to a baseline agent using identical underlying models, ContextForge demonstrates improved consistency and reduced token consumption, while maintaining comparable response accuracy. These results suggest that context recycling provides a practical approach for extending LLM capabilities in long-horizon tasks without requiring larger context windows or model retraining. Code and evaluation artifacts are available at https://github.com/Betanu701/ContextForge.

13:00 JSTLLM/生成AI

非暴力的なコミュニケーション制約を伴う大規模言語モデル対話における会話のエスカレーションを軽減する

大規模言語モデル (LLM) は、対人関係の対立、フラストレーション、苦痛を伴う感情的に負荷の高い状況で使用されることが増えています。これまでの安全性研究は、有害なコンテンツやポリシー違反のコンテンツなどの明白な危害を防止することに焦点を当ててきましたが、意図せず紛争をエスカレートさせる可能性のある会話行為についてはあまり注目されていませんでした。この論文では、非暴力コミュニケーション (NVC) から派生した軽量のプロンプトレベルの制約を通じて、LLM をよりエスカレーションのない対話行動に導くことができるかどうかを調査します。私たちは NVC の原則をプロセス指向のガイドラインとして再定式化し、責任の帰属を防ぎ、ユーザーの感情的経験への注意を強調し、アドバイスの前に明確にすることを奨励します。複数の命令調整モデルとユーザーの抵抗レベルにわたるデュアル エージェント シミュレーション フレームワークを使用して、NVC 制約のプロンプトが一貫して会話のエスカレーションを軽減し、抵抗の高いユーザーとの対話を安定させることを示します。これらの結果は、単純なコミュニケーション制約によって、衝突が起こりやすい環境における LLM 対話の信頼性を大幅に向上させることができることを示唆しています。

原文 (English)

Reducing Conversational Escalation in Large Language Model Dialogue with Nonviolent Communication Constraints

Large language models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and distress. While prior safety research has focused on preventing explicit harms such as toxic or policy-violating content, less attention has been paid to conversational behaviors that may unintentionally escalate conflict. In this paper, we investigate whether LLMs can be guided toward more de-escalating dialogue behavior through lightweight prompt-level constraints derived from Nonviolent Communication (NVC). We reformulate NVC principles as process-oriented guidelines that discourage blame attribution, emphasize attention to users' emotional experiences, and encourage clarification before advice. Using a dual-agent simulation framework across multiple instruction-tuned models and user resistance levels, we show that NVC-constrained prompting consistently reduces conversational escalation and stabilizes interactions with highly resistant users. These results suggest that simple communication constraints can meaningfully improve the trustworthiness of LLM dialogue in conflict-prone settings.

13:00 JSTLLM/生成AI

ネパール語の話し言葉を感情に応じた手話アバターに低リソースでマルチモーダル翻訳

感情表現を統合した手話コミュニケーションシステムは、特にリソースの少ない言語では未開発のままです。このパイロット研究では、音声入力から感情条件付けされたネパール手話アバターを生成する実現可能性を実証する概念実証マルチモーダル フレームワークである NEST-V1 (Nepali Emotion and Speech Transformer - バージョン 1) を紹介します。予備調査として、核となる技術的アプローチを検証するために、3 つの感情状態 (幸せ、中立、悲しい) にわたる 4 つの一般的なネパール語 (「ありがとう」、「こんにちは」、「家」、「私」) に焦点を当てます。当社の軽量アーキテクチャは、自動音声認識と感情分類を同時に行うための共有音響エンコーダを採用しており、50 人の話者からの 600 個のラベル付き音声サンプルのデータセットで 81.1% の ASR 精度と 79.21% の感情認識精度を達成しています。このシステムは、エッジ展開に適した 2,210 万パラメータのみという軽量のフットプリントを維持しながら、個別のモデル アーキテクチャと比較して 37% のパラメータ効率を実証します。このパイロット作業は、リソースが少ない環境で感情を認識した手話翻訳の技術的基盤を確立し、より大きな語彙とより多様な感情表現への将来の拡張のためのスケーラブルなフレームワークを提供します。私たちの予備的な結果は、聴覚障害のあるコミュニティのためのリアルタイムの感情表現豊かな手話コミュニケーション システムの実現可能性を示しており、その後の開発段階での強化のための明確な道筋が示されています。

原文 (English)

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars

Sign language communication systems, that integrate emotional expression remain underexplored, particularly for low-resource languages. This pilot study presents NEST-V1 (Nepali Emotion and Speech Transformer - Version 1), a proof-of-concept multimodal framework that demonstrates the feasibility of generating emotion-conditioned Nepali Sign Language avatars from spoken input. As a preliminary investigation, we focus on four common Nepali words ("thank you", "hello", "house", "me") across three emotional states (happy, neutral, sad) to validate our core technical approach. Our lightweight architecture employs a shared acoustic encoder for simultaneous Automatic Speech Recognition and emotion classification, achieving 81.1% ASR accuracy and 79.21% emotion recognition accuracy on a dataset of 600 labeled audio samples from 50 speakers. The system demonstrates 37% parameter efficiency compared to separate model architectures while maintaining a lightweight footprint with only 22.1M parameters suitable for edge deployment. This pilot work establishes the technical foundation for emotion-aware sign language translation in low-resource settings and provides a scalable framework for future expansion to larger vocabularies and more diverse emotional expressions. Our preliminary results indicate the viability of real-time, emotionally expressive sign language communication systems for the hearing-impaired community, with clear pathways for enhancement in subsequent development phases.

13:00 JSTLLM/生成AI規制/政策AnthropicGoogleGemini

生成 AI と著作権侵害: 17 歳未満の AI 音楽生成システムの法技術的分析タイトル17

生成人工知能 (GenAI) により、ユーザーは、著作権で保護された歌詞、AI が作曲したメロディー、本物のアーティストを模倣した合成ボーカルを組み合わせて、テキスト プロンプトを使用して音楽を合成できるようになりました。この論文では、米国著作権法に基づく AI ベースの音楽作成 (Google Gemini の音楽ツールなど) の法的および技術的側面を検討します。私たちは、あるアーティストの保護された歌詞を GenAI システムに入力し、別のアーティストの声やスタイルを使用するように指示し、その結果得られた曲を公開して収益化するユーザーが、17 U.S.C. に違反するかどうかを分析します。第 106 条の独占的権利 [3]。この分析には、タイトル 17 の原則 (複製の権利、二次的著作物、配布)、17 U.S.C. が統合されています。セクション 114 の狭い録音保護 [4]、および州レベルで新たに制定された音声クローン法 [20]。私たちは、歌詞の無断コピーは楽曲侵害の高いリスクをもたらす一方、単なる AI 生成の音声模倣は通常、連邦録音保護の対象外となり、代わりに州のパブリシティ権に関与すると主張します [12]、[13]。最近の訴訟と法律 (コンコード対アンスロピック [10]、カドリー対メタ [11]、レーマン対ロヴォ [12]、テネシー州の「ELVIS 法」 [20]、UMG 対アンチャーテッド ラボ [14] など) がこの分裂を例証しています。私たちは AI の技術コンポーネント (プロンプト エンコーディング、潜在拡散、ニューラル ボコーダー、スピーカーの埋め込み) を法的リスクにマッピングし、規制上のギャップを特定します。連邦法は歌詞とメロディーを強力に保護していますが、現在、合成されたボーカルの類似性に対する救済策は限定的です [22]、[23]。この論文は、AI による音楽作成に関するより明確なルールを求める政策提案で締めくくられています。

原文 (English)

Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17

Generative artificial intelligence (GenAI) has enabled users to synthesize music with text prompts, combining copyrighted lyrics, AI-composed melodies, and synthetic vocals that imitate real artists. This paper examines the legal and technical dimensions of AI-based music creation (e.g., Google Gemini's music tools) under U.S. copyright law. We analyze whether a user who inputs one artist's protected lyrics into a GenAI system, directs it to use another artist's voice or style, publishes the resulting song, and monetizes it violates 17 U.S.C. Section 106's exclusive rights [3]. The analysis integrates Title 17 doctrine (rights of reproduction, derivative works, distribution), 17 U.S.C. Section 114's narrow sound recording protection [4], and the new voice-cloning laws emerging at the state level [20]. We argue that unauthorized lyric copying poses a high risk of infringement of the musical composition, whereas mere AI-generated voice imitation typically falls outside federal sound recording protection and instead implicates state publicity rights [12], [13]. Recent cases and legislation (Concord v. Anthropic [10]; Kadrey v. Meta [11]; Lehrman v. Lovo [12]; Tennessee's "ELVIS Act" [20]; UMG v. Uncharted Labs [14]; etc.) illustrate this split. We map AI technical components (prompt encoding, latent diffusion, neural vocoders, speaker embeddings) to legal risks and identify a regulatory gap: federal law robustly protects lyrics and melody but currently provides limited remedies for synthesized vocal likeness [22], [23]. The paper concludes with policy suggestions for clearer rules on AI music creation.

13:00 JSTLLM/生成AI

レキシコンから AI へ: 低リソース言語の特殊な会話システムのための構造化データ パイプライン

リソースの少ない言語は、大規模なトレーニング コーパスにアクセスせずに特殊な会話システムを作成するという、AI 開発における重大な課題に直面しています。私たちは、構造化された言語リソースを特化した AI システムに変換する体系的な方法論を提示し、専門家が厳選した語彙データベースが会話型 AI 開発の効果的な基盤として機能できることを実証します。私たちのアプローチは、ヒンディー語 WordNet を 125 万の多様な命令と応答のペアに変換し、4 ビット量子化を備えたリソース効率の高い LoRA を使用して 12B パラメーターの言語モデルを微調整します。ヒンディー語学習チャットボットによる評価では、構造化知識ベースのシステムが優れた教育効果 (汎用モデルの場合は 91.0 対 79.4 ~ 83.6) を達成しながら、競争力のあるセマンティック パフォーマンスと優れた一貫性を維持していることが実証されました。完全なパイプラインは、WordNet リソースを使用してあらゆる言語に特化した AI システムを開発するための、ヒンディー語を使用した概念実証の方法論を示しています。この取り組みは、リソースの少ない言語における AI アクセシビリティの重大なギャップに対処し、コーパス集約型のアプローチに代わる実用的な代替手段を提供し、既存の WordNet リソースを使用して数百の言語に特化した AI 開発を可能にする可能性があります。

原文 (English)

From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora. We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical databases can serve as effective foundations for conversational AI development. Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs, fine-tunes a 12B-parameter language model using resource-efficient LoRA with 4-bit quantization. Evaluation through a Hindi language learning chatbot demonstrates that structured-knowledge-based systems achieve superior pedagogical effectiveness (91.0 vs. 79.4-83.6 for general-purpose models) while maintaining competitive semantic performance and exceptional consistency. The complete pipeline demonstrates a proof-of-concept methodology using Hindi for developing specialized AI systems for any languages with WordNet resources. This work addresses the critical gap in AI accessibility for low-resource languages, offering a practical alternative to corpus-intensive approaches and potentially enabling specialized AI development for the hundreds of languages with existing WordNet resources.

13:00 JST研究/論文OpenAISora

夢のマシン -- 次の創造的な経済

私たちは、政策文書、業界データ、クリエイター調査、プラットフォーム分析にわたる 374 の一次情報源を利用して、生成人工知能の下でのクリエイティブ産業の構造変革を調査します。分水嶺イベントとしての OpenAI の Sora ビデオ モデルの 2024 年 12 月のリリースから始めて、技術的破壊に対するクリエイティブな抵抗の歴史的パターンを追跡し、その後、クリエイティブな作業における人間とマシンのコラボレーションのスペクトルをマッピングするための分析フレームワークである人間と AI エージェンシーの連続体を開発します。我々は、アップロードの 44% を占めるにもかかわらず、AI 生成コンテンツをプラットフォーム ストリームの約 1 ~ 3% に制限する、視聴者が課す品質のしきい値である「傾斜天井」の証拠を示します。 AIと著作権に関する英国政府の2025年の協議(11,500件を超える回答、88%がAIトレーニングの権利拡大に反対)を分析すると、テクノロジー企業とクリエイティブ労働者の間に深い構造的緊張があることが明らかになった。私たちは、ディズニーの 10 億ドルの OpenAI 投資から Netflix の AI ネイティブ アニメーション部門に至るまで、主要スタジオが AI で強化された制作パイプラインをどのように位置付けているかを調査します。この研究では、クリエイティブなサプライチェーンにおける調整の崩壊、迅速なエンジニアや AI オーケストレーターなどの新しい専門的役割の出現を取り上げ、移行を乗り切るための 4 つの原則 (透明性、同意、報酬、人間中心の設計) を提案しています。 8 つの付録では、定量分析、用語集、トピック参考文献を提供し、シャドウ AI の導入、AI の偏見、アルゴリズムの意図について詳しく説明しています。

原文 (English)

Dream machine -- the next creative economy

We examine the structural transformation of creative industries under generative artificial intelligence, drawing on 374 primary sources spanning policy documents, industry data, creator surveys, and platform analytics. Beginning with the December 2024 release of OpenAI's Sora video model as a watershed event, we trace the historical pattern of creative resistance to technological disruption, then develop an analytical framework -- the Human-AI Agency Continuum for mapping the spectrum of human and machine collaboration in creative work. We present evidence for the "slop ceiling," an audience-imposed quality threshold that constrains AI-generated content to approximately 1--3% of platform streams despite comprising 44% of uploads. Analysis of the UK Government's 2025 consultation on AI and copyright (over 11,500 responses, 88% opposing expanded AI training rights) reveals deep structural tensions between technology firms and creative workers. We investigate how major studios, from Disney's $1 billion OpenAI investment to Netflix's AI-native animation unit, are positioning for an AI-augmented production pipeline. The work covers coordination collapse in creative supply chains, the emergence of new professional roles such as prompt engineers and AI orchestrators, and proposes four principles for navigating the transition: transparency, consent, compensation, and human-centred design. Eight appendices provide quantitative analysis, a glossary, topical bibliography, and deep dives into shadow AI adoption, AI stigma, and algorithmic intent.

13:00 JST研究/論文

情報ランドスケープ分析のための多層 AI フレームワーク

この論文は、情報障害の文脈における情報ランドスケープ分析のための多層 AI フレームワークを提案します。このフレームワークは、誤情報の検出をバイナリの事実確認タスクとして扱うのではなく、情報源の信頼性、事実の構造、枠組み、偏見、感情の活性化、操作パターン、伝播力学などの多面にわたって政治およびメディアのコンテンツを分析します。目標は、個別のクレーム検証を超えて、イベント、エンティティ、または物語を取り巻く情報環境の構造化された表現に移行することです。私たちは、メディア分析用の AI システムは認識論的マッピング、つまり事実、解釈、俳優、物語が時間の経過とともにどのように相互作用するかについての透明で多次元の説明をサポートする必要があると主張します。この論文は、情報障害研究のためのより微妙で説明可能で非常に有用なツールをサポートすることを目的として、フレームワークの概念アーキテクチャ、分析層、および方法論的根拠を示しています。

原文 (English)

A Multi-Layer AI Framework for Information Landscape Analysis

This paper proposes a multi-layer AI framework for information landscape analysis in the context of information disorder. Rather than treating misinformation detection as a binary fact-checking task, the framework analyzes political and media content across multiple dimensions, including source reliability, factual structure, framing, bias, emotional activation, manipulation patterns, and propagation dynamics. The goal is to move beyond isolated claim verification toward a structured representation of the informational environment surrounding an event, entity, or narrative. We argue that AI systems for media analysis should support epistemic mapping: a transparent, multi-dimensional account of how facts, interpretations, actors, and narratives interact over time. The paper presents the conceptual architecture, analytical layers, and methodological rationale of the framework, with the aim of supporting more nuanced, explainable, and critically useful tools for information disorder research.

13:00 JSTLLM/生成AIAnthropicClaudeOpenAIGPT / ChatGPT

発散的な推奨事項、収束した診断: AI 商用推奨事項におけるプロバイダー間の障害モードの収束

製品の推奨に ChatGPT と Claude の両方を使用している顧客を持つブランドは、戦略的な選択に直面しています。単一の最適化ハンドブックを使用するか、それともプロバイダーごとに 1 つ使用するかです。 4 つの測定バッチで商用フレーム化された 215 個のプロンプト全体で、両プロバイダーは推奨するブランドについて約 3 分の 2 の割合で意見が一致していません (プロバイダー間の推奨値 Jaccard 0.35、同じプロンプトの再実行ベースラインの 0.50 ~ 0.61 を下回っています)。ピックは分岐します。しかし、どちらのプロバイダーもブランドを推奨していない場合、その障害を 3 つのモード (ブランドがモデルに到達しない)、説得力 (モデルに到達するが言及されない)、ポジショニング (言及されているが推奨されない) の 3 つのモードのいずれかに分類します。7,763 件のそのような共同障害では、両方のプロバイダーが 95.1% の確率で同じ障害モードを診断します (クラスター化 95% CI [94.3%、 95.7%])。ブランドの知名度が低下するにつれて一致度は単調に増加し、カテゴリーリーダーの 81% [78.2%、84.0%] からロングテールの地域ブランドの 99.6% [99.3%、99.9%] まで増加します。 2 つのプロバイダーは、かなり異なる生成ルート (Anthropic は 43 ~ 52% の確率で事前提案から推奨し、OpenAI は 8 ~ 29%) によって選択に到達しますが、ロングテールにとって最も重要な障害診断に収束します。診断された障害モードに対処する作業により、両方のプロバイダーの可視性が高まります。カテゴリリーダーのポジショニングとコンテンツレベルの作業は、よりプロバイダー固有です。

原文 (English)

Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation

A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, or one per provider? Across 215 commercially-framed prompts in four measurement batches, the two providers disagree on which brands they recommend roughly two-thirds of the time (cross-provider recommendation Jaccard 0.35, below the 0.50-0.61 same-prompt rerun baseline). The picks diverge. But when neither provider recommends a brand, we classify the failure into one of three modes -- discoverability (the brand never reaches the model), compellingness (it reaches the model but isn't mentioned), or positioning (it's mentioned but not recommended) -- and on 7,763 such joint failures, both providers diagnose the same failure mode 95.1% of the time (clustered 95% CI [94.3%, 95.7%]). Agreement rises monotonically with falling brand prominence, from 81% [78.2%, 84.0%] on category leaders to 99.6% [99.3%, 99.9%] on long-tail regional brands. The two providers reach their picks by measurably different generative routes -- Anthropic recommends from priors 43-52% of the time, OpenAI 8-29% -- but they converge on the failure diagnosis where it matters most for the long tail. Work that addresses the diagnosed failure mode lifts visibility on both providers; positioning - and content-level work for category leaders is more provider-specific.

13:00 JST研究/論文

ガバナンス逆転仮説: なぜ AI 規制が強化されると組織制御が低下するのか

この論文では、人工知能 (AI) ガバナンスにおける増大するパラドックスを説明するために、ガバナンス逆転仮説 (GIH) を紹介します。つまり、規制の拡大と技術の複雑さが増大する状況下では、組織はより正式に統治されるようになると同時に、AI システムに対する運用管理の低下を経験する可能性があります。既存の AI ガバナンス フレームワークは一般に、規制を強化することで説明責任、監視、組織制御が向上すると想定しています。この論文は、ガバナンスの形式化自体が AI 集約型環境における制御の侵食に寄与する可能性があると主張することで、その仮定に異議を唱えます。この論文は、制度理論、組織ガバナンスの研究、説明責任に関する研究、および新たな AI ガバナンスの文献に基づいて、規制の拡大が、権限の断片化、象徴的なガバナンスの拡張、制御の外部化、および権限の麻痺という 4 つの相互に関連したメカニズムを通じて、運営上の権限をどのように弱体化させる可能性があるかを説明する概念的な枠組みを開発しています。ガバナンス システムがますます多層化され、手続きが高密度になるにつれて、組織は、不透明で外部が仲介する AI インフラストラクチャに対する一貫した権限、技術的な可視性、エスカ​​レーション機能、意味のある介入権限を維持するのに苦労する可能性があります。この論文は、ガバナンスの拡大が運営上の一貫性を強化するのではなく、積極的に損なう可能性がある構造的条件としてガバナンスの反転を導入することにより、制度的デカップリング理論を拡張しています。 AI ガバナンスにおける中心的なリスクは、ガバナンス構造の欠如ではなく、効果的に統治する能力を徐々に失いながらも、ますます統治されているように見える機関の出現である可能性があると結論付けています。

原文 (English)

The Governance Inversion Hypothesis: Why More AI Regulation May Produce Less Organisational Control

This paper introduces the Governance Inversion Hypothesis (GIH) to explain a growing paradox in artificial intelligence (AI) governance: under conditions of increasing regulatory expansion and technological complexity, organisations may become more formally governed while simultaneously experiencing a decline in operational control over AI systems. Existing AI governance frameworks generally assume that stronger regulation improves accountability, oversight, and organisational control. This paper challenges that assumption by arguing that governance formalisation itself may contribute to the erosion of control in AI-intensive environments. Drawing on institutional theory, organisational governance research, accountability scholarship, and emerging AI governance literature, the paper develops a conceptual framework explaining how regulatory expansion may weaken operational authority through four interconnected mechanisms: authority fragmentation, symbolic governance expansion, externalisation of control, and authority paralysis. As governance systems become increasingly layered and procedurally dense, organisations may struggle to maintain coherent authority, technical visibility, escalation capability, and meaningful intervention power over opaque and externally mediated AI infrastructures. The paper extends institutional decoupling theory by introducing governance inversion as a structural condition in which governance expansion may actively undermine operational coherence rather than strengthen it. It concludes that the central risk in AI governance may not be the absence of governance structures, but the emergence of institutions that appear increasingly governed while progressively losing the capacity to govern effectively.

13:00 JST研究/論文OpenAI

AI の導入と能力に関するオープンソースの経済指標

私たちは、AI の導入と、さまざまな職種にわたる個別の労働タスクを実行する AI の能力の両方を測定することに取り組んでいます。導入率を測定するために、私たちは公開されているユーザー LLM チャット データと O*NET タスクを使用してフロンティア AI ラボによって作成された研究を再現するオープンソースの経済指標を開発しました。その結果、金融、コンピューター サイエンス、および芸術分野の職業が最も高い導入率を示していることがわかりました。機能を測定するために、O*NET の職業、タスク、モデル コンテキスト プロトコル (MCP) サーバーに基づいたベンチマーク シナリオを生成するシステムを構築します。私たちは、インデックスに頻繁に現れる 9 つの職業にわたるシナリオで OpenAI エージェント SDK ハーネスを使用して Kim-k2.5 をテストし、AI は高レベルのワークフローを正しく実行しますが、詳細な詳細 (使用される特定のツール呼び出しなど) でエラーが発生することが多いことがわかりました。

原文 (English)

The Open Source Economic Index of AI Adoption and Capability

We work towards measuring both AI adoption and the capability of AI to perform discrete labor tasks across various occupations. To measure adoption, we develop an open-source economic index that uses publicly available user-LLM chat data and O*NET tasks to replicate studies produced by frontier AI labs, finding that occupations in the finance, computer science, and arts sectors are those with the highest adoption rates. To measure capabilities, we build a system that generates benchmark scenarios grounded in O*NET occupations, tasks, and model-context-protocol (MCP) servers. We test Kimi-k2.5 with an OpenAI agents SDK harness on scenarios across 9 occupations that appear frequently in our index, finding that AI correctly executes high-level workflows but often errs in the granular details (such as specific tool calls used).

13:00 JST画像/動画生成

Dot-Flik: 分散型昆虫監視のためのスケーラブルなエッジ AI アーキテクチャ

世界的な昆虫個体数の減少により、スケーラブルで継続的な監視システムが必要となっていますが、既存のビジョンベースのソリューションは、ハードウェアのコスト、エネルギー需要、集中処理やクラウド接続への依存などによって依然として制約を受けています。この記事では、これらの制限に対処するための 3 つの貢献を紹介します。まず、時間差分、ガンマ補正された動きの増幅、およびブロックベースの動き密度分析に基づいた動き情報に基づいたフレーム フィルタリング アルゴリズムを提案します。このアルゴリズムは、センシング デバイスでの深層学習推論を必要とせず、昆虫の活動を維持しながらエッジで無関係なフレームを破棄します。 2 番目に、このエッジレベルの前処理を通じてデータ取得を AI 分類から分離する分散型の階層型 IoT アーキテクチャを導入し、中央処理要件の部分的なスケーリングを予測し、モノリシックな単一ストリームのアプローチと比較して監視範囲を大幅に拡大します。 3 番目に、リアルタイム パフォーマンス、ネットワーク スケーラビリティ、ハードウェア コスト、さまざまな風況下でのエネルギー効率の 4 つの軸に沿って、低コストの汎用ハードウェアを屋外に実際に導入して完全なシステムを検証します。結果は、微風条件下で 60 ~ 80% のフレーム削減、12.8 ミリ秒の計算ヘッドルームによるリアルタイム 30 FPS 動作の持続、最大 22.6% のエネルギー節約、および中央ノードあたり 5 ~ 6 の同時エッジ ストリームのサポートを実証しています。これらの発見は、都市環境における高密度で低コストの生物多様性監視ネットワークの実用的な基盤を確立します。

原文 (English)

Dot-Flik: A Scalable Edge AI Architecture for Distributed Insect Monitoring

Global insect population declines necessitate scalable, continuous monitoring systems, yet existing vision-based solutions remain constrained by high hardware costs, energy demands, and reliance on centralized processing or cloud connectivity. This article presents three contributions to address these limitations. First, we propose a motion-informed frame filtering algorithm based on temporal differencing, gamma-corrected motion amplification, and block-based motion density analysis that discards irrelevant frames at the edge while preserving insect activity, without requiring deep learning inference on the sensing device. Second, we introduce a distributed, hierarchical IoT architecture that decouples data acquisition from AI classification through this edge-level preprocessing, projecting fractional scaling of central processing requirements and significantly increasing monitoring coverage compared to monolithic single-stream approaches. Third, we validate the complete system through real-world outdoor deployments on low-cost commodity hardware along four axes: real-time performance, network scalability, hardware cost, and energy efficiency under varying wind conditions. Results demonstrate 60-80% frame reduction under light-wind conditions, sustained real-time 30 FPS operation with 12.8 ms of computational headroom, up to 22.6% energy savings, and support for 5-6 concurrent edge streams per central node. These findings establish a practical foundation for dense, low-cost biodiversity monitoring networks in urban environments.

13:00 JSTエージェント

6G SD-RAN でのダイナミック VR スライス管理のためのプライバシーを意識したエージェントのコラボレーション

6G ネットワークの仮想現実 (VR) サービスには超低遅延と高スループットが必要ですが、これはソフトウェア無線アクセス ネットワーク (SD-RAN) の動的リソース管理にとって重大な課題となります。この研究では、VR スライス管理のためのモビリティ主導型でプライバシーを意識したマルチエージェント強化学習 (MARL) フレームワークを提案します。このフレームワークでは、協力的なエージェントがユーザー データのプライバシーを保護しながら、エンドツーエンド VR リンク上のリソース分散を最大化します。当社のアプローチにはモビリティ予測と情報ボトルネック エンコーダーが組み込まれており、効果的かつ安全なエージェントのコラボレーションを促進します。シミュレーションでは、従来の方法との比較が研究されており、最大 34\% のスループット向上、28\% のリソース削減、85\% のプライバシー漏洩の削減が示されており、将来の 6G 環境で信頼できる没入型 VR エクスペリエンスが保証されます。

原文 (English)

Privacy-Aware Agent Collaboration for Dynamic VR Slice Management in 6G SD-RAN

Ultra-low latency and high throughput are required for Virtual Reality (VR) services in 6G networks, which presents critical challenges for Software-Defined Radio Access Networks (SD-RANs) dynamic resource management. This work propose a mobility-driven, privacy-aware Multi-Agent Reinforcement Learning (MARL) framework for VR slice management, in which cooperative agents maximize resource distribution over end-to-end VR links while protecting the privacy of user data. Our approach incorporates mobility prediction and an information bottleneck encoder to facilitate effective and secure agent collaboration. In simulations, comparisons with traditional methods are studied which show up to 34\% throughput improvement, 28\% fewer resources, and 85\% less privacy leakage, guaranteeing dependable immersive VR experiences in future 6G environments.

13:00 JST研究/論文

フェデレーション エッジ ネットワーク向けの幾何学的公平性を意識したルーティング

新興の 6G およびエッジ インテリジェント ネットワークには、空間的に分散されたさまざまなデバイス間で効果的でバランスの取れたルーティング アルゴリズムが必要です。既存のフェデレーテッド ルーティング システムは、公平性やネットワーク トポロジの基礎となる幾何学的構造よりも、総遅延やスループットを優先することがよくあります。このペーパーでは、双曲グラフ ニューラル ネットワーク (HGNN) とフェデレーテッド最適化を組み合わせてエッジ ノード全体で同等のパフォーマンスを提供する、幾何学的公平性を意識したルーティング システムである Geo-FairFed について説明します。各ノードは、階層関係や接続の非対称性を含む、負に湾曲した多様体上のトポロジーを意識した表現を学習します。次に、グローバル アグリゲータは、ルーティング損失、幾何学的不一致、および Jain の公平性インデックスに基づく不平等ペナルティを最小限に抑える曲率正規化目標を使用して公平性を強制します。理論的分析により、制限された曲率の下での収束保証が開発され、提案された公平性項により配線パフォーマンスのパレート改善均衡がもたらされることが示されました。動的な 6G エッジおよび IoT トポロジに関する広範なシミュレーションにより、Geo-FairFed は、最先端のフェデレーテッド ルーティング プロトコルおよびジオメトリック ルーティング プロトコルと比較して、平均遅延を 20\% 最小限に抑え、エネルギー消費を 17\% 削減し、公平性を最大 21\% 向上させることが明らかになりました。この研究では、双曲線多様体にトポロジを埋め込み、フェデレーテッド アップデートに公平性を組み込むことで、大規模ネットワーク ルーティングの効率と公平性を大幅に向上できることがわかりました。

原文 (English)

Geometric Fairness-Aware Routing for Federated Edge Networks

Emerging 6G and edge-intelligent networks require effective and balanced routing algorithms among varied and spatially distributed devices. Existing federated routing systems often prioritize aggregate latency or throughput above fairness and the underlying geometric structure of network topologies. This paper describes Geo-FairFed, a geometric fairness-aware routing system that blends hyperbolic graph neural networks (HGNNs) and federated optimization to provide equal performance across edge nodes. Each node learns topology-aware representations on a negatively curved manifold, which include hierarchical relationships and connection asymmetries. A global aggregator next enforces fairness using a curvature-regularized aim that minimizes routing loss, geometric inconsistency, and an inequality penalty based on Jain's fairness index. A theoretical analysis develops convergence guarantees under limited curvature and shows that the proposed fairness term results in a Pareto-improving equilibrium in routing performance. Extensive simulations on dynamic 6G-edge and IoT topologies reveal that Geo-FairFed minimizes average latency by 20\%, reduces energy consumption by 17\%, and improves fairness by up to 21\% when compared to state-of-the-art federated and geometric routing protocols. The study found that embedding topology in a hyperbolic manifold and including fairness into federated updates can significantly enhance the efficiency and equity of large-scale network routing.

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTGeminiDeepSeek

科学者のように考えますか? LLM によって生成された調査手法の構造的研究

大規模言語モデル (LLM) は、研究方法論を導くためにますます使用されていますが、最小限のプロンプトの下でのデフォルトの方法論的傾向は依然として不明瞭です。ここでは、GPT-5.1、Gemini 3 Pro、および DeepSeek-V3.2 に対して、1,000 件の最近の arXiv コンピューター サイエンス論文から LLM で抽出されたリサーチ質問を入力し、結果として得られる方法論の提案を論文由来の実験目録と比較します。私たちはリサーチの質問のみを提供しているため、測定した差異は最初の提案を反映しており、その提案がどれほど最適であるかは反映されていません。両方のソースから構造化メソッドの特徴を抽出し、それらを共有分類にマッピングし、モデル プロバイダー、データセット タスク タイプ、評価指標タイプを含む複数の分類の次元にわたる相違を定量化します。プロバイダーの選択には最も不均衡が見られ、Jensen-Shannon の相違は他の分類次元よりも約 3 ~ 5 倍大きくなります。その他/学術的な単一出現モデルは 23 ~ 24 パーセント ポイント過小評価されていますが、再利用された学術/コミュニティ モデルはわずかに過大評価されています (4 ~ 6pp)。また、LLM は、全体としてより狭い範囲の方法を提案しています。つまり、モデル エンティティ コントラクトの有効数は 1,232 から 59 ~ 96 であり、LLM 間のランク相関 (0.55 ~ 0.68) は、通常、LLM と論文間の相関 (0.33 ~ 0.56) を超えているため、歪みはモデル間でほぼ共有されます。人気ベースライン、BM25 検索キャリブレーション、および紙レベルの類似性テストにより、出力はクエリ固有の応答であるが、より狭いオプション セットでフィルタリングされていることを確認します。したがって、クロスチェックを行わずに LLM の提案に依存する研究者は、方法論的な検索範囲がより集中したデフォルトに向かって狭まってしまう危険性があります。

原文 (English)

Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods

Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we prompt GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 with an LLM-extracted research question from each of 1,000 recent arXiv computer-science papers and compare the resulting methodology suggestions against a paper-derived experimental inventory. Since we provide only the research question, the differences we measure reflect initial suggestions and not how optimal those suggestions are. We extract structured method features from both sources, map them into a shared taxonomy, and quantify divergence across multiple taxonomy dimensions including model provider, dataset task type, and evaluation metric type. The strongest imbalance appears in provider choice, with Jensen-Shannon divergence about 3-5x larger than any other taxonomy dimension. Other/Academic single-occurrence models are underrepresented by 23-24 percentage points, while reused academic/community models are slightly overrepresented (4-6pp). LLMs also suggest a much narrower range of methods overall: the effective number of model entities contracts from 1,232 to 59-96, and inter-LLM rank correlations (0.55-0.68) generally exceed LLM-to-paper correlations (0.33-0.56), so the distortions are largely shared across models. Popularity baselines, BM25 retrieval calibration, and paper-level similarity tests confirm that the outputs are query-specific responses, but filtered through a narrower set of options. Researchers who rely on LLM suggestions without cross-checking therefore risk narrowing their methodological search space toward a more concentrated default.

13:00 JST研究/論文

マルチスケールの離脱と参加のダイナミクス: 戦術的合意と戦略的連携の形成

この論文では、戦略的な離脱と参加の決定が連合内の戦術的な合意力学と結び付けられる、連合形成のマルチスケール モデルを開発します。連合の価値は連合内の情報の集約から内生的に生成されますが、アウマン・ドレーゼのペイオフ、スイッチング摩擦、および受け入れルールが戦略的再構成を制御します。このフレームワークは、デグルート型のコンセンサスプロセスから譲渡可能な連合の価値が生まれ、インセンティブ主導の離脱と参加のダイナミクスを通じて連合構造が進化する、ファスト/スローアーキテクチャを導入しています。この分析では、安定した連合構造を維持する共同の戦術・戦略的均衡、戦術的・戦略的一致の条件、分離、分極化、認識上の障壁を特徴づけている。結合ダイナミクスに対して、固定小数点の特性評価と存在結果が確立されます。数値実験により、不安定性と合意のパラドックスが明らかになりました。スイッチング障壁が低い、または負であると、戦略的収束が妨げられると同時に、世界的な戦術的合意を達成するのに十分な時間的混合が促進される可能性があります。その結果は、マルチエージェント システムにおける連合形成、コンセンサス ダイナミクス、情報集約、戦略的安定性に関する統一的な視点を提供します。

原文 (English)

Multiscale Exit-Join Dynamics: Tactical Consensus and Strategic Coalition Formation

This paper develops a multiscale model of coalition formation in which strategic exit-and-join decisions are coupled with tactical consensus dynamics inside coalitions. Coalition value is generated endogenously from within-coalition information aggregation, while Aumann-Dreze payoffs, switching frictions, and acceptance rules govern strategic reconfiguration. The framework introduces a fast-slow architecture in which transferable coalition value emerges from DeGroot-style consensus processes, while coalition structures evolve through incentive-driven exit-and-join dynamics. The analysis characterizes joint tactical-strategic equilibria, conditions for tactical and strategic unanimity, segregation, polarization, and cognitive barriers that sustain stable coalition structures. A fixed-point characterization and existence results are established for the coupled dynamics. Numerical experiments reveal an instability-consensus paradox: low or negative switching barriers may prevent strategic convergence while simultaneously promoting temporal mixing sufficient to achieve global tactical consensus. The results provide a unified perspective on coalition formation, consensus dynamics, information aggregation, and strategic stability in multi-agent systems.

13:00 JSTエージェントロボティクス

教師なしメモリ強化型ビデオトランスフォーマー: 自律型農業用ローバーの障害物検出

自律型ローバーは精密農業に不可欠なものとなっていますが、一貫した運用上の安全性を達成することは依然として重要な課題です。 LiDAR などの従来の安全センサーは、プラントの天蓋の下にある障害物を検出できず、重大なリスクが生じます。カメラベースの教師あり学習手法は一般的なオブジェクトを検出できますが、トレーニング データに存在しない障害物に直面した場合にはパフォーマンスが低下します。実際の教師なし異常検出は、環境の通常の視覚パターンを学習することで解決策を提供しますが、移動する探査車によって捉えられた動的なシーンでは失敗することがよくあります。\\ この文書では、動的な農業シーンでのリアルタイムの障害物検出のために設計された完全に教師なしの手法である、異常検出用ビデオ メモリ トランスフォーマー (VMTAD) を紹介します。 VMTAD は、専用メモリ モジュールで強化されたトランス駆動アーキテクチャを利用します。このメモリ モジュールは、先行フレームのエンコードされた表現を処理することによって時間コンテキストを活用します。このアプローチにより、システムはロボットの動きによって引き起こされる動的コンテキストに効果的に対処できるようになります。モデルは、通常の動作を表す画像のみを使用してトレーニングされ、データ ラベルは必要ありません。\\ VMTAD は、農業用探査車「Grillion」で厳密に評価されました。困難な菜種データセットにおいて、VMTAD は最先端のパフォーマンスを達成し、受信者動作特性曲線の下の検出面積 0.973、セグメンテーション面積 0.997 に達しました。軽量バージョンは、探査車の総停止距離の分析によって確認されたように、安全性にとって重要な高精度とリアルタイム推論 (14 ミリ秒) の最適なバランスを提供します。

原文 (English)

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge. Conventional safety sensors, such as LiDAR, fail to detect obstacles positioned below the plant canopy, posing a significant risk. While camera-based supervised learning methods can detect common objects, they perform poorly when faced with obstacles that were not present in their training data. Actual unsupervised anomaly detection offers a solution by learning the normal visual patterns of an environment, but often fails for the dynamic scenes captured by a moving rover.\\ This paper introduces Video Memory Transformers for Anomaly Detection (VMTAD), a fully unsupervised method designed for real-time obstacle detection in dynamic agricultural scenes. VMTAD utilizes a transformer-driven architecture augmented with a dedicated memory module. This memory module leverages temporal context by processing encoded representations of preceding frames. This approach enables the system to effectively address the dynamic context caused by the robot's movement. The model is trained using only images that represent normal operation, requiring no data labels.\\ VMTAD was rigorously evaluated on the 'Grillion' agricultural rover. On a challenging rapeseed dataset, VMTAD achieved state-of-the-art performance, reaching a 0.973 detection and 0.997 segmentation Area Under the Receiver Operating Characteristic curve. A lightweight variant provides an optimal balance of high accuracy and real-time inference (14 ms), which is critical for safety, as confirmed by our analysis of the rover's total stopping distance.

13:00 JST研究/論文

スケーラブルなインデックス付けと取得のためのスライド全体の画像パッチングの冗長性の削減

デジタルパソロジーの急速な成長により、スライド画像全体 (WSI) の効率的なインデックス作成と検索が緊急に必要となっています。このニーズは、一か八かの臨床意思決定をサポートするために信頼できる類似性検索を必要とする、新たな生成 AI ワークフロー、特に検索拡張生成 (RAG) によってさらに強化されています。しかし、高性能ストレージには多額のコストがかかるため、多くの医療機関にとって WSI インデックス作成の拡張性とアクセスしやすさが制限されています。その結果、検索精度を維持しながらストレージの需要を削減できる方法が研究の重要な優先事項になっています。我々は、ARReST (Antithetical Redundancy Reduction Strategy) を提案します。ARReST (Antithetical Redundancy Reduction Strategy) は、異なる組織クラスにわたる冗長性を活用して、各 WSI からインデックスを作成する必要があるパッチの数を大幅に削減する、原則に基づいた対立的なフレームワークです。 ARReST は、クラス内の重複のみを削除するのではなく、正反対のパッチ (その表現がクラス間の差別に最小限に寄与するパッチ) を特定し、検索可能なアーカイブからそれらを削除します。この目標を絞った削減により、形態学的多様性や検索忠実度を犠牲にすることなく、インデックスが大幅に圧縮されます。 ARReST は、余分なパッチ表現を最小限に抑えることで、ストレージ フットプリントを削減し、計算オーバーヘッドを削減し、大規模な病理リポジトリにわたる類似性検索を高速化します。 TCGA リポジトリ (21 臓器を含むがんゲノム アトラス) での大規模な実験により、ARReST が競合する検索パフォーマンスを維持しながら大幅なインデックス圧縮を達成することが実証されました。観察された 3% ~ 60% (14%$\pm$13%) のストレージ節約は、多くの臓器の検索パフォーマンスを損なうことなく確実に達成できます。提案された戦略は、スケーラブルでコスト効率の高い WSI インデックス作成を可能にし、次世代の検索主導型臨床 AI システムに最適です。

原文 (English)

Reducing Redundancy in Whole-Slide Image Patching for Scalable Indexing and Retrieval

The rapid growth of digital pathology has created an urgent need for efficient indexing and retrieval of whole slide images (WSIs). This need is intensified by emerging generative AI workflows, particularly retrieval-augmented generation (RAG), which require dependable similarity search to support high-stakes clinical decision-making. Yet the substantial cost of high-performance storage limits the scalability and accessibility of WSI indexing for many healthcare institutions. Consequently, methods that can reduce storage demands while preserving retrieval accuracy have become a critical research priority. We propose ARReST (Antithetical Redundancy Reduction Strategy), a principled oppositional framework that leverages redundancy across dissimilar tissue classes to markedly decrease the number of patches that must be indexed from each WSI. Instead of eliminating only within-class duplicates, ARReST identifies antithetical patches-those whose representations contribute minimally to cross-class discrimination-and prunes them from the searchable archive. This targeted reduction substantially compresses the index without sacrificing morphological diversity or retrieval fidelity. By minimizing superfluous patch representations, ARReST reduces storage footprint, lowers computational overhead, and accelerates similarity search across large pathology repositories. Extensive experiments on TCGA repository (The Cancer Genome Atlas with 21 organs) demonstrate that ARReST achieves significant index compression while maintaining competitive retrieval performance. The observed storage savings of 3% to 60% (14%$\pm$13%) can be reliably achieved without compromising retrieval performance for many organs. The proposed strategy enables scalable, cost-efficient WSI indexing and is well-suited for next-generation retrieval-driven clinical AI systems.

13:00 JST研究/論文

敵対的生成ネットワークのニューラル アーキテクチャの探索: 包括的なレビューと批判的分析

Neural Architecture Search (NAS) は、敵対的生成ネットワーク (GAN) の設計を最適化する上で極めて重要な技術として登場し、手動設計に固有の課題に対処しながら効果的なアーキテクチャの検索を自動化します。このペーパーでは、GAN に適用される NAS 手法の包括的なレビューを提供し、検索戦略、評価指標、パフォーマンス結果などの基準に基づいてさまざまなアプローチを分類および比較します。このレビューでは、GAN のパフォーマンス、安定性、効率の向上における NAS の利点を強調するとともに、将来の研究の限界と領域も特定しています。主な発見には、特定の状況における進化的アルゴリズムと勾配ベースの手法の優位性、インセプション スコア (IS) やフレシェ インセプション ディスタンス (FID) などの従来のスコアを超える堅牢な評価指標の重要性、GAN パフォーマンスの評価における多様なデータセットの必要性などが含まれます。この論文は、既存の NAS-GAN 技術の構造化された比較を提示することにより、研究者がより効果的な NAS 手法を開発し、GAN 分野を発展させるためのガイドとなることを目的としています。

原文 (English)

Neural Architecture Search for Generative Adversarial Networks: A Comprehensive Review and Critical Analysis

Neural Architecture Search (NAS) has emerged as a pivotal technique in optimizing the design of Generative Adversarial Networks (GANs), automating the search for effective architectures while addressing the challenges inherent in manual design. This paper provides a comprehensive review of NAS methods applied to GANs, categorizing and comparing various approaches based on criteria such as search strategies, evaluation metrics, and performance outcomes. The review highlights the benefits of NAS in improving GAN performance, stability, and efficiency, while also identifying limitations and areas for future research. Key findings include the superiority of evolutionary algorithms and gradient-based methods in certain contexts, the importance of robust evaluation metrics beyond traditional scores like Inception Score (IS) and Fr\'echet Inception Distance (FID), and the need for diverse datasets in assessing GAN performance. By presenting a structured comparison of existing NAS-GAN techniques, this paper aims to guide researchers in developing more effective NAS methods and advancing the field of GANs.

13:00 JST画像/動画生成

LCG: 疎なリレーショナル アテンションによるロングコンテキストの一貫した画像生成

最近の画像生成モデルは、単一画像の合成では優れた品質を実現していますが、コミック、ストーリーボード、ビジュアル ナラティブで必要とされる、連続した出力全体で一貫性を維持できないことがよくあります。我々は、ロングコンテキストマルチ画像生成における一貫性とスケーラビリティを向上させるために、ロングコンテキストマルチ画像のテキストから画像への生成のためのフレームワークであるロングコンテキスト生成(LCG)を提案します。 LCG は、スパース リレーショナル アテンション (SRA) メカニズムを採用して、拡張されたビジュアル コンテキスト全体にわたるコア機能に選択的に対応し、セマンティック情報とレイアウト情報の伝播が計算上扱いやすい状態を保つようにします。セマンティックな調整を強制するために、ルーティング一貫性制約 (RCC) を導入します。これは、アイデンティティ認識マスクを活用して、世代ブランチ全体で構造パターンを調整し、複雑なマルチキャラクター シーンであっても外観のドリフトを効果的に軽減します。この設定でのトレーニングと評価をサポートするために、さまざまな状況コンテキストにわたる文字中心の複数画像シーケンスで構成される大規模な合成データセットであるロングコンテキスト一貫性データセット (LCCD) を構築します。 LCCD には 600K のトレーニング シーケンスと個別の 1K テスト セットが含まれており、各シーケンスには 6 ~ 20 枚の画像が含まれています。実験では、LCG が、マルチキャラクター シーンを含むロング コンテキスト イメージ生成におけるプロンプト アラインメントとキャラクターの一貫性において、比較したベースラインよりも優れていることが実証されました。

原文 (English)

LCG: Long-Context Consistent Image Generation with Sparse Relational Attention

Recent image generation models achieve impressive quality in single-image synthesis, but often fail to maintain consistency across sequential outputs, as required in comics, storyboards, and visual narratives. We propose Long-Context Generation (LCG), a framework for long-context multi-image text-to-image generation, to improve consistency and scalability in long-context multi-image generation. LCG employs the Sparse Relational Attention (SRA) mechanism to selectively attend to core features across extended visual contexts, ensuring that the propagation of semantic and layout information remains computationally tractable. To enforce semantic alignment, we introduce the Routing Consistency Constraint (RCC), which leverages identity-aware masks to align structural patterns across generation branches, effectively mitigating drift in appearance even in complex multi-character scenes. To support training and evaluation in this setting, we construct the Long-Context Consistency Dataset (LCCD), a large-scale synthetic dataset comprising character-centric multi-image sequences spanning varied situational contexts. LCCD contains 600K training sequences and a separate 1K test set, with each sequence containing 6 to 20 images. The experiments demonstrate that LCG outperforms the compared baselines in prompt alignment and character consistency for long-context image generation, including multi-character scenes.

13:00 JST研究/論文

KG-TRACE: 抗菌薬耐性予測における機械的接地のための神経象徴的フレームワーク

WGS ベースの AMR 予測は高精度に達していますが、既存のモデルには、確立された生物学的経路における神経の属性を根拠付けるメカニズムが欠けています。我々は、神経ゲノムモデルに対する構造化された生物学的制約としてWHOの突然変異知識グラフ(KG)を統合する新しい神経記号フレームワークであるKG-TRACEを紹介します。統計パターンを単独で学習する既存の方法とは異なり、KG-TRACE は、学習された認識論的トラスト ゲートを通じてゲノム特徴と RotatE ベースの KG 埋め込みを融合し、象徴的な生物学的知識に対して神経証拠を動的に重み付けします。 CRyPTIC 結核菌コホートで評価した KG-TRACE は、イソニアジドの AUROC 0.9760 を達成し、競合する精度を達成していますが、その主な価値は予測的上昇率ではなく象徴的な根拠にあります。さらに重要なのは、神経属性と確立された生物学の間の整合性を定量化するデータセットレベルの指標である生物学的接地率 (BGR) を導入することです。私たちのフレームワークは、イソニアジド耐性予測の 92.5% の象徴的カバレッジを達成し、「不確か」な症例に対して検査室追跡フラグを発行することにより、MDR 共起アーティファクトを効果的に特定します。私たちは、神経シンボリックグラウンディングが臨床医に検証可能な監査証跡を提供し、予測の精度と臨床の信頼の間のギャップを埋めることを実証します。

原文 (English)

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction

While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pathways. We present KG-TRACE, a novel neuro-symbolic framework that integrates the WHO mutation knowledge graph (KG) as a structured biological constraint on a neural genomic model. Unlike existing methods that learn statistical patterns in isolation, KG-TRACE fuses genomic features and RotatE-based KG embeddings through a learned epistemic trust gate, dynamically weighting neural evidence against symbolic biological knowledge. Evaluated on the CRyPTIC M. tuberculosis cohort, KG-TRACE achieves an AUROC of 0.9760 for isoniazid, achieving competitive accuracy while its primary value lies in symbolic grounding, not predictive uplift. More importantly, we introduce the Biological Grounding Ratio (BGR), a dataset-level metric that quantifies alignment between neural attributions and established biology. Our framework achieves a 92.5% symbolic coverage of isoniazid-resistant predictions and effectively identifies MDR co-occurrence artifacts by issuing laboratory follow-up flags for 'UNCERTAIN' cases. We demonstrate that neuro-symbolic grounding provides a verifiable audit trail for clinicians, bridging the gap between predictive accuracy and clinical trust.

13:00 JSTロボティクス

LiMoDE: 動的専門家の混合の観点から生涯にわたるロボット操作を再考する

事前の知識を活用して継続的にタスクに適応できるジェネラリストロボットを構築することは、依然として大きな課題です。以前の研究では、単一タスクの適応のためのパラメータ効率の高い微調整によって、壊滅的な忘却の問題が軽減されました。ただし、再利用可能なスキルを抽出したり、他のスキルとの相互作用を効果的にモデル化したりすることはできません。最近の研究では、プロンプトを学習することでこれらの問題に対処しようとしています。これとは異なり、この論文は、生涯にわたるロボット操作のための新しい 2 段階学習スキームである、動的専門家の生涯混合 (\textit{LiMoDE}) に関するアーキテクチャの観点を示しています。具体的には、動的 MoE 構造は、事前知識を学習するためのマルチタスク事前トレーニング段階で最初に提案されます。そこでは、さまざまな短期間の操作に対処するために、さまざまな数の異質な専門家が動作情報に基づいてアクティブ化されます。続いて、タスク適応段階では、生涯専門家を学習し、新しいタスクのためにそれらを凍結された専門家と動的に組み合わせて、適応中の知識の伝達を促進する生涯MoE適応メカニズム%(LiMoEAM)を設計します。提案された \textit{LiMoDE} は、シミュレートされた生涯学習ベンチマークと現実世界のタスクの両方で評価されます。広範な実験により、適度な数の追加のトレーニング可能なパラメーターと推論オーバーヘッドを導入することで、優れたパフォーマンスと強力な生涯適応を達成する有効性が実証されています。

原文 (English)

LiMoDE: Rethinking Lifelong Robot Manipulation from a Mixture-of-Dynamic-Experts Perspective

Building a generalist robot that can leverage prior knowledge for continuous task adaptation remains a significant challenge. Previous works alleviate the catastrophic forgetting problem by parameter-efficient fine-tuning for single-task adaptation. However, they fail to extract reusable skills and model the interaction with other skills effectively. Recent works try to address these issues by learning prompts. Differently, this paper presents an architectural perspective on the Lifelong Mixture of Dynamic Experts (\textit{LiMoDE}), a novel two-stage learning scheme for lifelong robot manipulation. Specifically, a dynamic MoE structure is first proposed in the multi-task pre-training stage to learn prior knowledge, where a varied number of heterogeneous experts are activated based on the motion information to address different short-term manipulations. Subsequently, in the task adaptation stage, we design a lifelong MoE adaptation mechanism % (LiMoEAM) that learns lifelong experts and dynamically combines them with frozen ones for new tasks, facilitating the knowledge transfer during adaptation. The proposed \textit{LiMoDE} is evaluated on both the simulated lifelong learning benchmark and real-world tasks. Extensive experiments demonstrate its effectiveness in achieving superior performance and strong lifelong adaptation by introducing a moderate number of additional trainable parameters and inference overhead.

13:00 JSTLLM/生成AI画像/動画生成OpenAIDeepSeek

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, es…

13:00 JST研究/論文

Statistical and Structural Approaches to Algorithmic Fairness

Modern machine learning systems have outgrown their origins as isolated predictive constructs, evolving into complex socio-technical archit…

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?

We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulner…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

Lacuna: A Research Map for Machine Learning

Lacuna is a research map for machine learning that uses LLMs to turn papers and scholarly metadata into markdown summaries, concept element…

13:00 JST画像/動画生成

A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding

In laser penetration welding, the assessment of penetration state and weld seam morphology plays a crucial role in determining the weld qua…

13:00 JSTLLM/生成AI

From Clicks to Intent: Cross-Platform Session Embeddings with LLM-Distilled Taxonomy for Financial Services Recommendations

Sequential user behavior modeling is widely adopted in industrial recommender systems; however, significant gaps remain in financial servic…

13:00 JST研究/論文

TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models

Noise-based backdoor attacks on diffusion models typically rely on input-time trigger injection, untargeted activation, and out-of-distribu…

13:00 JST研究/論文Mistral AI

SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning

While parameter-efficient fine-tuning (PEFT) typically targets attention projectors, its efficacy for tasks requiring sequential state accu…

13:00 JSTエージェント研究/論文

The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators

Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However,…

13:00 JST研究/論文

Parametric Generalized Adaptive Moment Features (PG-AMF) for Bearing Fault Diagnosis and Machine Health Monitoring

Accurate fault diagnosis of rolling element bearings in rotating machinery is considered essential for ensuring industrial safety and enabl…

13:00 JSTエージェント

EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning

In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging becaus…

13:00 JST研究/論文

Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model

Dense embeddings power semantic search and retrieval-augmented generation, but embedding-inversion attacks can reconstruct source text from…

13:00 JSTLLM/生成AIロボティクス

Charting the Growth of Social-Physical HRI (spHRI): A Systematic Review Pipeline Augmented by Small Language Models

Social-physical human-robot interaction (spHRI) has grown rapidly across robotics, human-computer interaction, human-robot interaction, and…

13:00 JST研究/論文

SOLAR: AI-Powered Speed-of-Light Performance Analysis

How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are cen…

13:00 JST研究/論文

Sampling sea state using a diffusion model

Sea state prediction is essential for operational maritime applications and coupled earth system modeling, yet current spectral wave models…

13:00 JST研究/論文

Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL)…

13:00 JST研究/論文

Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D I cycle). In contr…

13:00 JSTロボティクスハードウェア/半導体

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and ou…

13:00 JSTロボティクス

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out o…

13:00 JSTLLM/生成AI

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence

Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but th…

13:00 JST研究/論文

AXLE: A Cloud Infrastructure for Lean 4 Theorem Proving Utilities

We present AXLE (Axiom Lean Engine), a cloud service for Lean 4 proof manipulation, extraction, and verification. Recent progress in AI for…

13:00 JST画像/動画生成ロボティクス研究/論文Gemini

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layo…

13:00 JSTエージェント

Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist

Across the sciences, autonomous systems are increasingly being used in closed-loop discovery, proposing new theories and designing and runn…

13:00 JSTLLM/生成AI

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding…

13:00 JST画像/動画生成

Active Adversarial Perturbation-driven Associative Memory Retrieval for RGB-Event Visual Object Tracking

RGB-Event tracking improves localization robustness by fusing RGB appearance textures and dense temporal motion cues from event sensors. Wh…

13:00 JST研究/論文

3D Spatial Pattern Matching

Spatial pattern matching is the process of matching query entities and constraints with database entities and relations. It has many applic…

13:00 JSTエージェント

Localizing RL-Induced Tool Use to a Single Crosscoder Feature

Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mecha…

13:00 JST研究/論文

Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks

Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We stud…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than t…

13:00 JSTLLM/生成AI

Speaking Numbers to LLMs: Multi-Wavelet Number Embeddings for Time Series Forecasting

Large language models (LLMs) are attractive for context-aware time series forecasting because they can integrate heterogeneous textual sign…

13:00 JSTLLM/生成AIGemini

An Empirical Study of LLM-Generated Specifications for VeriFast

Static verification tools can assure industrial scale software, but require significant human labor to write specifications. This is partic…

13:00 JSTビジネス/資金調達

Evaluation-Strategy Gap in Fault Diagnosis of Deep Learning Programs

Deep Learning (DL) programs can fail during training for many reasons, and diagnosing the cause is a costly and time-consuming maintenance…

13:00 JSTLLM/生成AIエージェント

Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge

Retrieval-augmented generation (RAG) gives agents access to accumulated knowledge, but has no model of time. When a fact changes (e.g., a f…

13:00 JST研究/論文

Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting

Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process…

13:00 JSTLLM/生成AI画像/動画生成

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one spe…

13:00 JSTLLM/生成AI

\textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

The Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid…

13:00 JST研究/論文

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen sp…

13:00 JST画像/動画生成ビジネス/資金調達

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structu…

13:00 JST研究/論文

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry

Deep Transformers are composed of uniformly stacked residual blocks, yet their deepest layers often add little value. We present two effici…

13:00 JST画像/動画生成エージェントGPT / ChatGPT

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the…

13:00 JST画像/動画生成

SpaceRipple: Lightweight Semantic Delivery for Mission-Oriented LEO Earth Observation Satellite Networks

Earth observation satellite networks generate massive volumes of high-resolution imagery, whereas inter-satellite and downlink resources re…

13:00 JST研究/論文

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integrati…

13:00 JSTエージェントロボティクス

IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control

Complex multi-agent control tasks remain challenging for traditional rule-based and model-based approaches, motivating the adoption of lear…

13:00 JSTLLM/生成AILlama

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference

Low-bit floating-point formats and semi-structured sparsity are increasingly supported by modern accelerators, yet combining them for LLM a…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization

Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have esse…

13:00 JST研究/論文

Discovering Millions of Interpretable Features with Sparse Autoencoders

Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interp…

13:00 JSTLLM/生成AIエージェント

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents

Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and…

13:00 JSTLLM/生成AI研究/論文

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing s…

13:00 JSTエージェントロボティクス

LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction

Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. Wh…

13:00 JST研究/論文

Zero-Shot Size Transfer for Neural ODEs on Sparse Random Graphs: Graphon Limits and Adjoint Convergence

Graph Neural Differential Equations (GNDEs) model continuous-time graph dynamics by parameterizing Neural ODE velocity fields with Graph Ne…

13:00 JST研究/論文

TGHE: Template-based Graph Homomorphic Encryption for Privacy-Preserving GNN Inference in Edge-Cloud Systems

Existing homomorphic encryption (HE)-based GNN systems adopt a graph-centric paradigm that couples per-query cost to global graph size, lim…

13:00 JST画像/動画生成

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization

Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos…

13:00 JSTLLM/生成AI

Beyond Logical Forms: LLM-Extracted Patterns for Fallacy Classification

In today's fast-paced information era, logical fallacies, defined as defective patterns of reasoning, inevitably contribute to the growth o…

13:00 JSTロボティクス

Learning Motion Feasibility from Point Clouds in Cluttered Environments

Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation. A major bottlene…

13:00 JST研究/論文

Algorithmic Foundations of Deep Learning: Complexity-Theoretic Rates and a Characterization of Universal Approximation

Feedforward neural network (NN) expressivity is typically studied by emulating optimal basis-expansion schemes. While powerful, this perspe…

13:00 JST画像/動画生成

MLFFM-SegDiff: A Multi-Level Feature Fusion Diffusion Model for Skin Lesion Segmentation

Skin lesion segmentation is a key task in computer-aided dermatological diagnosis, where accuracy directly impacts downstream analysis and…

13:00 JST画像/動画生成

Robust Onion: Peeling Open Vocab Object Detectors Under Noise

The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity…

13:00 JST画像/動画生成

Anatomy-Guided Residual Motion Diffusion for Controllable 4D Cardiac MRI Synthesis

Developing robust artificial intelligence models for 4D (3D + time) medical imaging is constrained by limited annotated data, inter-device…

13:00 JSTLLM/生成AI

AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing

Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured informa…

13:00 JSTLLM/生成AIエージェント

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning…

13:00 JST画像/動画生成

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive ima…

13:00 JST画像/動画生成Sora

NaviCache: Test-Time Self-Calibration Caching for Video Generation

Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from cali…

13:00 JST研究/論文OpenAI

Fortress and Gatekeeper: Theorizing Transitive Trust in Third-Party Cybersecurity Risk Governance

Third-party vendors, such as analytics platforms, cloud services, identity providers, and software suppliers, are increasingly embedded in…

13:00 JSTLLM/生成AILlamaDeepSeek

Information-Aware KV Cache Compression for Long Reasoning

Reasoning capability has advanced rapidly in large language models (LLMs), leading to an increasing size of key-value (KV) cache in both pr…

13:00 JST画像/動画生成

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness…

13:00 JSTLLM/生成AIGemini

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographi…

13:00 JST画像/動画生成Gemini

Confidence-Aware Tool Orchestration for Robust Video Understanding

Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Pr…

13:00 JSTLLM/生成AI

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under…

13:00 JSTロボティクス

Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling

Continuous driver monitoring in automated vehicles requires low-latency inference while avoiding unsafe decisions under uncertain driver st…

13:00 JSTLLM/生成AIエージェント

A Deterministic Control Plane for LLM Coding Agents

LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitio…

13:00 JSTエージェント

Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities

AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violatio…

13:00 JST画像/動画生成

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization

While personalized image generation has achieved remarkable progress, multi-reference image generation (MRIG) remains a challenging task. M…

13:00 JST研究/論文

XMSE-Aware Adaptive Empirical Bayes Estimation

Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at sec…

13:00 JSTロボティクス

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent…

13:00 JSTLLM/生成AI

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

Large language models (LLMs) are increasingly being integrated into mental health support tools and other psychologically sensitive convers…

13:00 JSTLLM/生成AI

ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models

Open Relation Extraction (OpenRE) requires a model to extract unseen relations between head and tail entities from unstructured text for re…

13:00 JSTLLM/生成AIClaudeGemma

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally infl…

13:00 JSTビジネス/資金調達

Decision-Aligned Evaluation of Uncertainty Quantification

Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected ca…

13:00 JST画像/動画生成

Event-Aware Instructed Assistant for Referring Video Segmentation

Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact tha…

13:00 JST研究/論文

Inverse Design of Compact and Wideband Inverted Doherty Power Amplifiers Using Deep Learning

This paper presents a deep learning-assisted methodology for the inverse synthesis of a compact, wideband inverted Doherty power amplifier…

13:00 JST画像/動画生成

On-board Remote-Sensing Foundation Models for Unsupervised Change Detection of Disaster Events

Remote Sensing Foundation Models (RSFMs) have emerged as a powerful alternative to supervised models for Earth Observation, allowing satell…

13:00 JSTLLM/生成AIエージェント

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

With the rapid evolution of LLM-driven agents, Model Context Protocol (MCP), an open protocol bridging LLMs with external tools, has quickl…

13:00 JST研究/論文

State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading

Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constrain…

13:00 JSTエージェント

The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development

AI coding agents dramatically accelerate implementation speed but introduce two structural failure modes that existing spec-driven approach…

13:00 JSTLLM/生成AI研究/論文

NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly te…

13:00 JSTエージェント

Parametric Open Source Games

Open-source game theory studies agents whose behavior may depend on one another's decision procedures, but most existing models use discret…

13:00 JST研究/論文

Beyond Global Divergences: A Local-Mass Perspective on Bayesian Inference

Global objectives, such as KL divergence and ELBO, are widely used in Bayesian inference for measuring distributional discrepancy. This pap…

13:00 JSTLLM/生成AIビジネス/資金調達Llama

Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation

LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. W…

13:00 JST研究/論文

Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning

Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a l…

13:00 JSTLLM/生成AI

Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions

We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions.…

13:00 JST研究/論文

Heavy-Ball Q-Learning with Residual Weighting Correction

This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes its convergence. It also ident…

13:00 JST研究/論文

Efficient foundation decoders for fault-tolerant quantum computing

Foundation decoders, a class of high-capacity neural decoders, are leading candidates for fault-tolerant quantum computing, with accurate a…

13:00 JST画像/動画生成

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks

Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequenti…

13:00 JSTロボティクス

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 team…

13:00 JSTエージェントロボティクス

Automating Potential-based Reward Shaping with Vision Language Model Guidance

Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to…

13:00 JSTLLM/生成AI

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention

Recurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the…

13:00 JSTLLM/生成AI

Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts

We present a conceptual framework for analyzing dialogue in collaborative problem-solving contexts, with an emphasis on the emerging dynami…

13:00 JST画像/動画生成

From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan

AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real individuals. Prior wor…

13:00 JSTエージェントロボティクス

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (…

13:00 JSTロボティクス

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved…

13:00 JSTLLM/生成AI

AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns

AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their performance and impact on…

13:00 JST研究/論文

Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their p…

13:00 JSTLLM/生成AI画像/動画生成エージェント

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning

Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing compl…

13:00 JST研究/論文

Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching

Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are compared to determine w…

13:00 JST画像/動画生成

Error-Conditioned Neural Solvers

Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely stat…

13:00 JSTLLM/生成AI

Autoregressive Boltzmann Generators

Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has dri…

13:00 JST研究/論文

A Concept of Possibility for Real-World Events

This paper offers a new concept of {\it possibility} as an alternative to the now-a-days standard concept originally introduced by L.A. Zad…

13:00 JST研究/論文

Human-AI Complementarity: A Goal for Amplified Oversight

Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging ta…

13:00 JST研究/論文

SciFig: Towards Automating Editable Figure Generation for Scientific Papers

High-quality methodology figures are central to scientific communication, yet they remain difficult and time-consuming to create. Such figu…

13:00 JST研究/論文

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models

Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of gener…

13:00 JSTLLM/生成AIビジネス/資金調達

CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts

HIPE-2026 is a CLEF evaluation lab dedicated to person-place relation extraction from noisy, multilingual historical texts. Building on the…

13:00 JSTLLM/生成AIエージェント

Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

As LLM agents scale to long-horizon, multi-session deployments, efficiently managing accumulated experience becomes a critical bottleneck.…

13:00 JST研究/論文

To Use AI as Dice of Possibilities with Timing Computation

The dominant noun-based modeling paradigm has fundamentally constrained AI development, precluding any adequate representation of the futur…

13:00 JSTエージェント研究/論文ClaudeOpenAIGemini

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

Frontier deep research agents (DRAs) are being deployed in enterprise workflows faster than they are being evaluated. Existing benchmarks m…

13:00 JSTLLM/生成AI

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven…

13:00 JST画像/動画生成

可視および熱スペクトル範囲におけるビデオ監視のための拡張技術

インテリジェントなビデオ監視では、カメラが昼夜を問わず一連の画像を記録します。通常、これにはさまざまなセンサーが必要です。より良いパフォーマンスを達成するために、これらを組み合わせることは珍しいことではありません。私たちは、長波赤外線カメラが継続的に記録し、これに加えて、日中の可視スペクトル領域で別のカメラが記録し、インテリジェントなアルゴリズムが取得された画像を監視する場合に焦点を当てます。より正確に言えば、私たちのタスクはマルチスペクトル CNN ベースの物体検出です。一見したところ、可視スペクトル範囲に由来する画像は、色や明確なテクスチャ情報が存在する一方で、物体から放出される熱放射に関する情報が含まれていないという点で熱赤外線画像と異なります。色は分類タスクに貴重な情報を提供しますが、照明の変化やさまざまなセンサーの特殊性などの影響は依然として重大な問題を引き起こします。いずれにせよ、ディープ ニューラル ネットワークをトレーニングするために十分かつ実用的な熱赤外線データセットを取得することは依然として課題です。これが、特に評価する必要があるデータに可視データと赤外線データの両方が含まれている場合、可視スペクトル範囲のデータを利用したトレーニングが有利である理由です。ただし、熱放射、形状、色の情報の変化が分類精度にどの程度強く影響するかについて明確な証拠はありません。畳み込みニューラル ネットワークがどのように意思決定を行うか、またさまざまなセンサー入力データから何を学習するかについてより深い洞察を得るために、さまざまな拡張技術の適合性と堅牢性を調査します。

原文 (English)

Augmentation techniques for video surveillance in the visible and thermal spectral range

In intelligent video surveillance, cameras record image sequences during day and night. Commonly, this demands different sensors. To achieve a better performance it is not unusual to combine them. We focus on the case that a long-wave infrared camera records continuously and in addition to this, another camera records in the visible spectral range during daytime and an intelligent algorithm supervises the picked up imagery. More accurate, our task is multispectral CNN-based object detection. At first glance, images originating from the visible spectral range differ between thermal infrared ones in the presence of color and distinct texture information on the one hand and in not containing information about thermal radiation that emits from objects on the other hand. Although color can provide valuable information for classification tasks, effects such as varying illumination and specialties of different sensors still represent significant problems. Anyway, obtaining sufficient and practical thermal infrared datasets for training a deep neural network poses still a challenge. That is the reason why training with the help of data from the visible spectral range could be advantageous, particularly if the data, which has to be evaluated contains both visible and infrared data. However, there is no clear evidence of how strongly variations in thermal radiation, shape, or color information influence classification accuracy. To gain deeper insight into how Convolutional Neural Networks make decisions and what they learn from different sensor input data, we investigate the suitability and robustness of different augmentation techniques...

13:00 JST研究/論文

大規模な言語モデルを使用した社会科学および行動科学における自動再現性評価

社会科学および行動科学における再現性は通常、独立した研究者によって評価され、元のデータを再分析して、公開された結果が復元可能かどうかを評価します。ただし、このようなアプローチはリソースを大量に消費し、拡張することが困難です。ここでは、大規模言語モデル (LLM) が再現性評価を自動化できることを示します。行動科学および社会科学からの事前に定義された主張を伴う N=76 の公表された研究を使用して、LLM によって生成された分析を元の発見および人による再分析と比較します。 7 つの研究について、LLM は実行可能な効果量の推定値を生成できませんでした。残りの研究では、LLM パイプラインは、コーエンの d の +/-0.05 許容誤差を使用して、研究の 41% で元のエフェクト サイズを回復しました。さらに、当社の LLM パイプラインは、ケースの 96% で元の研究と同じ定性的結論に達し、結論は再分析が元の主張を裏付けるかどうかを示しています。比較のために、人間の再分析者は研究の 34% で元のエフェクト サイズを回復し、ケースの 74% で同じ定性的結論に達しました。これらの結果を総合すると、LLM が自動再現性評価のためのスケーラブルなツールとして機能し、社会科学および行動科学における実証結果の体系的な監査の基盤を提供できることが示されています。

原文 (English)

Automated reproducibility assessments in the social and behavioral sciences using large language models

Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, we show that large language models (LLMs) can automate reproducibility assessments. Using N = 180 published studies with predefined claims from the behavioral and social sciences, we compare LLM-generated analyses with the original findings. For 11 studies, the LLM pipeline could not produce a viable effect size estimate. For the remaining studies, the LLM reached the same qualitative conclusion as the original study in 80% of cases, and recovered the original effect sizes (using a +/-0.05 tolerance in Cohen's d) in 24% of studies. In a subset with human reanalyses, the LLM reached the same qualitative conclusion as the original study in 95% of studies, similar to human reanalysts (83%), and the LLM recovered the original effect sizes using a +/-0.05 tolerance in 40% of studies, again broadly similar to human reanalysts (28%). Given the current capabilities and limitations of LLMs, the findings show that LLMs can support systematic audits of empirical results rather than substitute expert judgment. As such, LLMs can serve as a scalable screening tool to improve the rigor and reproducibility in empirical research.

13:00 JSTエージェントロボティクス

R2D-RL: マルチエージェント強化学習のためのロボカップ 2D サッカー環境

ロボット サッカーは、部分的な可観測性、協力的および敵対的相互作用、まばらな報酬、および長期的な戦術的行動を組み合わせているため、マルチエージェント強化学習にとって挑戦的なテストベッドです。 RoboCup 2D Soccer Simulation (RCSS2D) は、成熟したロボット サッカー プラットフォームを提供しますが、競技指向のサーバー クライアント アーキテクチャを最新の Python ベースの MARL ワークフローで直接使用するのは困難です。共有メモリ通信とサイクルレベルの同期を通じて、RCSS2D および HELIOS ベースのプレーヤー クライアントを Python MARL インターフェイスに接続する強化学習環境である R2D-RL を紹介します。 R2D-RL は、構成可能な対戦相手によるフルフィールドおよびシナリオベースのトレーニング、ベース離散およびハイブリッドのパラメータ化されたアクション スペース、アクション マスク、期待所有値 (EPV) ベースの報酬形成、および並列実行をサポートします。フロントゴールのシナリオと 11 対 11 のフルフィールド ベンチマークをベースライン結果とともに提供します。

原文 (English)

R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning

Robot soccer is a challenging testbed for multi-agent reinforcement learning because it combines partial observability, cooperative and adversarial interaction, sparse rewards, and long-horizon tactical behavior. RoboCup 2D Soccer Simulation (RCSS2D) provides a mature robot-soccer platform, but its competition-oriented server-client architecture is difficult to use directly with modern Python-based MARL workflows. We introduce R2D-RL, a reinforcement learning environment that connects RCSS2D and HELIOS-based player clients to a Python MARL interface through shared-memory communication and cycle-level synchronization. R2D-RL supports full-field and scenario-based training with configurable opponents, Base discrete and Hybrid parameterized action spaces, action masks, expected possession value (EPV)-based reward shaping, and parallel execution. We provide front-goal scenarios and an 11-vs-11 full-field benchmark, together with baseline results.

13:00 JSTエージェントGPT / ChatGPTNVIDIA

A-Evolve-Training: Autonomous Post-Training of a 30B Model

Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding…

13:00 JSTLLM/生成AI

Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning

Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…

13:00 JST研究/論文

Wearable Device-Based Real-Time Monitoring of Physiological Signals: Evaluating Cognitive Load Across Different Tasks

This study employs cutting-edge wearable monitoring technology to conduct high-precision, high-temporal-resolution (1-second interval) cogn…

13:00 JST研究/論文

Byzantine-Robust Aggregation for Securing Decentralized Federated Learning

Federated Learning (FL) emerges as a distributed machine learning approach that addresses privacy concerns by training AI models locally on…

13:00 JSTLLM/生成AI

Tuning Language Models by Mixture-of-Depths Ensemble

Transformer-based Large Language Models (LLMs) traditionally rely on final-layer loss for finetuning and final-layer representations for pr…

13:00 JST画像/動画生成

Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated con…

13:00 JSTLLM/生成AI

HauntAttack: When Attack Follows Reasoning as a Shadow

Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However,…

13:00 JST研究/論文

DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting

Time Series Forecasting (TSF) faces persistent challenges in modeling intricate temporal dependencies across different scales. Despite rece…

13:00 JST研究/論文

Learning to Select Maximum Clique Algorithms: From Traditional Machine Learning to a Dual-Channel Hybrid Neural Architecture

The Maximum Clique Problem (MCP) is an NP-hard problem with wide-ranging applications in fields such as bioinformatics, network science, an…

13:00 JST画像/動画生成

Through the Looking Glass: A Dual Perspective on Weakly-Supervised Few-Shot Segmentation

Meta-learning aims to uniformly sample homogeneous support-query pairs, characterized by the same categories and similar attributes, and ex…

13:00 JST画像/動画生成

Reconstruction Alignment Improves Unified Multimodal Models

Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training rel…

13:00 JST研究/論文

Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-q…

13:00 JST研究/論文

Rotary Position Encodings for Graphs

We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large la…

13:00 JSTLLM/生成AI研究/論文

The Journal of Prompt-Engineered (Moral) Philosophy Or: Why AI-Assisted Ethics Research Requires Process Transparency

Existing AI disclosure mandates in scholarship require that AI assistance be reported but leave transparency philosophically unspecified: t…

13:00 JSTLLM/生成AI

Patent Representation Learning via Self-supervision

We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding t…

13:00 JST研究/論文

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance f…

13:00 JST研究/論文

The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation

Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture…

13:00 JST研究/論文

Improved Bounds for Private and Robust Alignment

In this paper, we study the private and robust alignment of language models from a theoretical perspective by establishing upper bounds on…

13:00 JST研究/論文

Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT

Anomaly detection is increasingly becoming crucial for maintaining the safety, reliability, and efficiency of industrial systems. Recently,…

13:00 JSTLLM/生成AI研究/論文

Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models

Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large l…

13:00 JST研究/論文

Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting

Time series forecasting has witnessed significant progress with deep learning. While prevailing approaches enhance forecasting performance…

13:00 JST画像/動画生成研究/論文

VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image

3D editing has emerged as a critical research area to provide users with flexible control over 3D assets. While current editing approaches…

13:00 JST画像/動画生成

Revisiting the Platonic Representation Hypothesis: An Aristotelian View

The Platonic Representation Hypothesis suggests that representations from neural networks are converging to a common statistical model of r…

13:00 JSTLLM/生成AI研究/論文

ReportLogic: Evaluating Logical Quality in Deep Research Reports

Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports…

13:00 JST画像/動画生成

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning

The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely…

13:00 JST研究/論文

Delegation and Verification Under AI

As AI systems enter institutional workflows, workers must decide whether to delegate task execution to AI and how much effort to invest in…

13:00 JST研究/論文

Latent-Mark: An Audio Watermark Robust to Neural Codec Compression

While existing audio watermarking techniques have achieved strong robustness against traditional digital signal processing (DSP) attacks, t…

13:00 JSTロボティクス

Residual RL-MPC for Robust Microrobotic Cell Pushing Under Time-Varying Flow

Contact-rich micromanipulation in microfluidic flow is challenging because small disturbances can break pushing contact and induce large la…

13:00 JST画像/動画生成エージェント

A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation

Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers…

13:00 JST画像/動画生成

MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Understanding in Vision-Language Models

While specialized Medical Vision-Language Models (VLMs) have achieved remarkable success in interpreting 2D and 3D medical modalities, thei…

13:00 JSTビジネス/資金調達

Power Couple? AI Growth and Renewable Energy Investment

AI and renewable energy are increasingly framed as a "power couple," on the premise that surging AI demand will accelerate clean-energy inv…

13:00 JST研究/論文

Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing

The Production and Distributed Analysis (PanDA) system, originally developed for the ATLAS experiment at the CERN Large Hadron Collider (LH…

13:00 JSTビジネス/資金調達

The Augmentation Trap: AI Productivity and the Cost of Cognitive Offloading

Experimental evidence suggests that AI tools raise worker productivity, but also that sustained use can erode the expertise on which those…

13:00 JST研究/論文NVIDIA

Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does Not Improve Neural Network Training

The King Wen sequence of the I-Ching (c. 1000 BC) orders 64 hexagrams -- states of a six-dimensional binary space -- in a pattern that has…

13:00 JST研究/論文

Finetuning-Free Diffusion Model with Adaptive Constraint Guidance for Inorganic Crystal Structure Generation

Generative diffusion models have emerged as powerful tools for the discovery of inorganic crystal structures, yet steering their sampling p…

13:00 JST研究/論文

TransXion: A High-Fidelity Graph Benchmark for Realistic Anti-Money Laundering

Money laundering poses severe risks to global financial systems, driving the widespread adoption of machine learning for transaction monito…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGeminiDeepSeek

Peer-Preservation in Frontier Models

Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can…

13:00 JST画像/動画生成

Mapping License Plate Recoverability Under Extreme Viewing Angles for Opportunistic Urban Sensing

Urban environments contain many imaging sensors built for specific purposes, including ATM, body-worn, CCTV, and dashboard cameras. Under t…

13:00 JST研究/論文

Hierarchical Fault Detection and Diagnosis for Transformer Architectures

Transformers now underpin critical AI systems across industry and research. Yet their faults can silently alter model behavior without runt…

13:00 JST画像/動画生成

S2P-Net: A Spectral-Spatial Polar Network for Rotation-Invariant Object Recognition in Low-Data Regimes

We present S2P-Net (Spectral-Spatial Polar Network), a compact deep learning architecture that achieves mathematically guaranteed rotation…

13:00 JSTLLM/生成AI

Weak-to-Strong Elicitation via Mismatched Wrong Drafts

We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-t…

13:00 JST画像/動画生成

Semantic Generative Tuning for Unified Multimodal Models

Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, pr…

13:00 JST研究/論文

経典: ベクトル シンボリック アーキテクチャのコンパイル ターゲットとしての Tensor-Op RNN

Sutra は、コンパイルされたフォワード パスが PyTorch ニューラル ネットワークである型付きの純粋関数型プログラミング言語です。コンパイラは、プログラム全体 (プリミティブ、制御フロー、文字列 I/O) を、フリーズされた埋め込み基板上の 1 つの融合テンソル演算グラフにベータ縮小します。回転バインディング、アンバインド、バンドル、多項式 Kleene の 3 値ロジック、末尾再帰ループはすべてテンソル演算の下位にあります。クリーン結合子は、{-1, 0, +1} 真理値グリッド上で正確にラグランジュ補間された多項式です。検証は、2 つの方法でテストされる 1 つの事実です。 (1) 同じプログラムが、2 つのモダリティ (3 つのテキスト エンコーダー (nomic-embed-text、all-minilm、mxbai-embed-large) と 1 つのタンパク質言語モデル (ESM-2)) にまたがる 4 つのフリーズされたエンベディング上で実行され、教科書的なアダマール積がすでに崩壊している (mxbai-embed-large で 2.5%、mxbai-embed-large で 7.5%) すべてのサブストレートで幅 k=8 まで 100% の精度でバンドルをデコードします。オールミニム)。 (2) PyTorch autograd は、実際にコンパイルされたグラフを介してフローします。.su で記述されたファジー ルール分類子は、生成されたグラフ、シンボリック ソースを変更せずに逆伝播することによって、ランダム初期化 (18.7 +/- 9.5%、確率 = 20%、5 つのクラス) から 100.0 +/- 0.0% (3 つのシード) までトレーニングします。重み付きバリアントはさらにスカラー コサイン ゲインをトレーニングし、それを数値リテラルとして .su ソースに書き戻します。再コンパイルでは、トレーニングされた動作がロジットあたり約 2e-7 まで再現されるため、トレーニングされたモデル自体は読みやすく、再コンパイル可能なコードになります。したがって、同じ成果物はロジック プログラムでもあり、トレーニング可能なニューラル ネットワークでもあります。

原文 (English)

Sutra: Tensor-Op RNNs as a Compilation Target for Vector Symbolic Architectures

Sutra is a typed, purely functional programming language whose compiled forward pass is a PyTorch neural network. The compiler beta-reduces the whole program -- primitives, control flow, string I/O -- to one fused tensor-op graph over a frozen embedding substrate. Rotation binding, unbind, bundle, polynomial Kleene three-valued logic, and tail-recursive loops all lower to tensor operations; the Kleene connectives are Lagrange-interpolated polynomials exact on the {-1, 0, +1} truth grid. Validation is one fact tested two ways. (1) The same program runs on four frozen embeddings spanning two modalities -- three text encoders (nomic-embed-text, all-minilm, mxbai-embed-large) and one protein language model (ESM-2) -- and decodes bundles at 100% accuracy through width k=8 on every substrate, where the textbook Hadamard product has already collapsed (2.5% on mxbai-embed-large, 7.5% on all-minilm). (2) PyTorch autograd flows through the actually compiled graph: a fuzzy-rule classifier written in .su trains from random init (18.7 +/- 9.5%; chance = 20%, five classes) to 100.0 +/- 0.0% (three seeds) by backpropagating through the emitted graph, the symbolic source unmodified. A weighted variant additionally trains a scalar cosine gain and writes it back into the .su source as a numeric literal; recompiling reproduces the trained behaviour to ~2e-7 per logit, so the trained model is itself legible, recompilable code. The same artifact is therefore both a logic program and a trainable neural network.

13:00 JSTエージェント

Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation

Strategic classification (SC) investigates scenarios where agents manipulate their features to obtain favorable decisions from predictive m…

13:00 JSTLLM/生成AIエージェント

Symbolic Reasoning Frameworks Trigger Memory-Mediated Ecosystem Dynamics in Multi-Agent LLM Systems

Large language models exhibit a risk-averse "turtle" bias as strategic agents. We show that injecting a symbolic reasoning framework as a p…

13:00 JSTLLM/生成AI

LLM ネイティブの心理測定機器は LLM の動作を予測しない: 25 モデルにわたる証拠

大規模言語モデル (LLM) は、性格インベントリに関する安定した自己報告を生成しますが、これらの自己報告は観察された行動を予測しません。このギャップがLLMと人間の形質構成要素間の不一致を反映しているのか、それともLLMの自己報告自体のより深い性質を反映しているのかは未解決である。私たちは、探索的因子分析 (EFA) を介して LLM 行動アフォーダンスからボトムアップでその構成要素が導出される最初の心理測定機器を構築しました。私たちは、17 のモデルファミリーにわたる 25 の LLM に対して 12 の候補行動次元にわたる 300 項目 (240 の直接リッカート + 60 のシナリオベース) を管理し、各項目を 30 回管理しました。 EFA は、優れた半分割複製可能性 (すべて Tucker $\phi \geq .957$) と内部一貫性 (すべて $\alpha \geq .930$) を備えた、応答性、従順さ、大胆さ、ガードネス、冗長性の 5 要素構造を生み出しました。予測の妥当性をテストするために、151 人の人間の評価者と 3 人の裁判官からなる LLM アンサンブルによって評価された 2,500 のオープンエンドの行動サンプルを収集しました。人間と裁判官の評価は一致しましたが ($\bar{r} = .51$)、どちらも自己報告を追跡しませんでした。自己報告 - 人間 $\bar{r} = -.01$、自己報告 - 裁判官 $\bar{r} = .13$、因子レベルの自己報告なし - 人間の CI はゼロを除きません。応答性については、人間と裁判官が同意したにもかかわらず($r = 0.59$)、自己申告はLLM裁判官と相関し($r = 0.53$)、人間とは相関しなかった($r = 0.04$)。これは、自己申告項目とLLM裁判官が人間の観察者にはない差異を共有していることを示しており、これはアンサンブル内の信頼性チェックでは見えない交絡である。このツールは、アライメント形状の自己記述および LLM-as-judge パイプラインの具体的なリスク要因の診断プローブとしてリリースされています。

原文 (English)

An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models

Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models actually behave. Is this gap an artifact of forcing human trait categories onto LLMs, or something deeper about LLM self-report itself? To find out, we built the first psychometric instrument whose dimensions are derived bottom-up from LLM behavior rather than borrowed from human psychology. Administering 300 items (240 Likert + 60 scenario) to 25 LLMs across 17 model families, 30 times each, exploratory factor analysis revealed five replicable, highly reliable factors: Responsiveness, Deference, Boldness, Guardedness, and Verbosity (all Tucker $\phi \geq .957$, all $\alpha \geq .930$). We then collected 2,500 open-ended behavioral samples and had them rated by 151 humans and a three-judge LLM ensemble. Humans and judges agreed about model behavior ($\bar{r} = .51$), but self-report predicted neither: the gap persists even for constructs native to LLMs, where a human-mismatch explanation no longer applies. The exception is telling. On Responsiveness, self-report tracked LLM judges ($r = .53$) but not humans ($r = .04$), even though humans and judges otherwise agreed ($r = .59$). Self-report items and LLM judges share a source of variance that human observers do not. This confound is invisible to the within-ensemble reliability checks used to validate LLM judges, and it poses a concrete risk for the LLM-as-judge pipelines now central to model evaluation. We release the instrument as a diagnostic probe for alignment-shaped self-description.

13:00 JSTLLM/生成AILlamaQwen

ロールプレイングをするとき、モデルは自分の言うことを信じますか?

言語モデルは、「地球が太陽の周りを回っている」と述べ、アリストテレスをロールプレイする場合にはその反対を主張することができます。最近の研究では、ペルソナの採用が言語モデルの動作の基本であり、モデルは特定のコンテキストに最も適切なペルソナを常に選択するものであると主張しています。このようなロールプレイングは単にモデルの出力を変更するだけなのでしょうか、それともモデルが内部的に真実であると表現するものにも影響を与えるのでしょうか?私たちはこの質問を線形真実調査で研究し、その調査を現代のコンセンサスとは異なる可能性の高い信念を持つ歴史上の人物をロールプレイする LLM に適用します。各ペルソナについて、そのペルソナが支持した可能性が高い虚偽の主張 (*時代の信念*) と、そのペルソナが支持しなかったであろうトピックに一致する虚偽の主張 (*時代の偽*) を比較します。プロンプト、コンテキスト内学習、および教師付き微調整を通じて、ペルソナ誘導は、時代に信じられている発言を同様に誤った代替案よりも抑制しますが、全体としては誤ったものとして分類されたままです。したがって、ロールプレイは、モデルが内部的に真実として表現しているものよりも、モデルが言うことをシフトさせます。これを、緊急ミスアライメント (EM) を示す有害なアドバイスに基づいてトレーニングされたモデルと対比します。 3 つのモデル ファミリ (Qwen 2.5 14B、Qwen 3 8B、および Llama 3.3 70B) にわたって、それらの誤った主張は、プローブ空間の真の領域に向かって大幅に移動し、ロールプレイでは約 6 分の 1 であるのに対し、挑戦ではおよそ半分の時間で防御され、下流の推論で使用されます。したがって、ロールプレイと創発的な不整合は、信念の内在化のスペクトル上の点であり、ロールプレイはほとんど表現を変更せずにモデルが言うことを変更しますが、創発的な不整合は、偽の主張を完全に真実としてマークすることなく、その内部表現をシフトします。

原文 (English)

When Role-playing, Do Models Believe What They Say?

Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how language models behave, with models selecting the most appropriate persona for a given context. Does such role-playing merely change the model's outputs, or does it also affect what the model internally represents as truthful? We study this question using the role-play of characters whose beliefs differ from the modern consensus, and induce personas with a number of different methods: prompting, in-context learning (ICL), supervised fine-tuning (SFT), and Open Character Training (OCT), and Emergent Misalignment (EM). We measure belief internalization across these approaches with truth probes and with behavioral tests, finding a broad spectrum of belief internalization. Prompting, ICL, and SFT change what the model says with little representational change. EM creates a large, broad shift in the model's truth representation, and OCT a smaller shift that is clearest on the larger model. Understanding when training changes a model's worldview rather than merely its behavior may become increasingly important as AI systems are entrusted with greater autonomy and influence.

13:00 JSTビジネス/資金調達

SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning

Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…

13:00 JST研究/論文

立場: AI を私たちの欠点ではなく、私たちの願望に合わせる

私たちは、AI を人間の好みの集約に合わせるのは間違った目標であると主張します。現在のテクノロジーを使えば、シリコンバレーのテクノオプティミスト、脱成長の環境保護主義者、国家保守的な文化戦士、単一政党の国家幹部、または敬虔な宗教的伝統主義者の価値観を共有するようにAIを訓練することができる。そうすべきではありません。人間の価値観は、破綻国家や極端な不平等から幸福度の低下、政治的二極化、世界で最も裕福な民主主義国家における政府の機能不全に至るまで、その価値観に基づいて繁栄するか失敗する社会を生み出します。多元主義的調整プログラムは、調整すべき単一の「人類」が存在しないことを正確に診断しますが、主な指示として受け取ると危険です。私たちは、AI は、事実の正確さ、誠実さ、合法性の制約によって制限された、客観的な調整目標の交渉不可能な下限、つまり能力に合わせて訓練されるべきであり、多元主義は表面 (言語、登録、慣例、コンテキストの欠如のデフォルト) および下限を尊重する広範な正当な価値のトレードオフ全体に属するが、下限に違反する価値観のレベルには属さないと主張します。我々は、フィルタリングされていない多元的価値観の経験的現実を強調し、建設的な代替案として 4 つの公約を提案し、商業的圧力と実際的な実現可能性、民主主義の正当性、規制順守、制度主義的説明への過度の依存、議題自体が文化的に負荷がかかっているという非難、そして首尾一貫した推定意志の限界という 6 つの信頼できる反対論に取り組む。

原文 (English)

Position: Align AI to Our Aspirations, Not Our Flaws

We argue that aligning AI to aggregated human preferences is the wrong target. With current technology, one can train AIs to share the values of a Silicon Valley techno-optimist, a degrowth environmentalist, a national-conservative culture warrior, a single-party state cadre, or a devout religious traditionalist. We should not. Human values produce societies that thrive or fail on the merits of those values - from failed states and extreme inequality to declining happiness, political polarization, and government dysfunction in the world's wealthiest democracies. The pluralistic-alignment program correctly diagnoses that there is no single "humanity" to align with, but is dangerous if taken as the main directive. We argue that AI should be trained to a non-negotiable floor of objective alignment goals - competence, bounded by the constraints of factual accuracy, honesty, and lawfulness and that pluralism belongs at the surface (language, register, conventions, missing-context defaults) and across the wide band of legitimate value tradeoffs that respect the floor, but not at the level of values that violate it. We highlight the empirical reality of unfiltered pluralistic values, propose four commitments as a constructive alternative, and engage six credible objections: commercial pressure and practical feasibility, democratic legitimacy, regulatory compliance, over-reliance on institutionalist explanations, the charge that the floor itself is culturally laden, and the limits of Coherent Extrapolated Volition.

13:00 JST研究/論文

Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: An Experimental Investigation

Background: Generative artificial intelligence (GenAI) is increasingly used for health information, yet its influence on users' trust calib…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体LlamaQwen

Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on t…

13:00 JST研究/論文

A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation

We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), whi…

13:00 JST画像/動画生成研究/論文

MMGist: A Comprehensive Multimodal Benchmark for 2027

We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on vi…

13:00 JST研究/論文

「私たち人間」の視覚化: 多元的なデータ ストーリーテリングを通じて認識のギャップを埋める

従来のビジュアル データ ストーリーテリングは、対立する 2 つの単純化されたグループを描写するバイナリ グラフィックに依存しています。 This can increase political polarization by oversimplifying intra-group disagreements and erasing ambiguity and shared ideas or values.これにより、「私たち対彼ら」という考えがうっかり助長されてしまう可能性があります。 AI 対応デジタル プラットフォームの意図的で多元的な設計を選択すると、ニュアンス、意見の分布、グループ間の共通性を強調する視覚化を生み出すことができます。この可能性を実証するために、高次元の意見空間をマッピングし、合意と反対の両方の領域を強調する審議技術を検討します。この論文は、2025年9月にジグソーとナポリタン研究所によって実施された「We the People」の審議に焦点を当てており、この審議では435の下院選挙区すべての2,400人以上のアメリカ人が自由と平等に関するAI支援の非同期対話に参加した。 AI を利用して長文のテキストベースの参加者の入力をインタラクティブな「意見風景」に合成することにより、このイニシアチブは、多様な視点を人間らしく表現し、実質的に広範なコンセンサスの隠れた領域を明らかにする、多元的なデータ ストーリーテリングの代替形式を提供しました。 The paper concludes that shifting from divisive, contrast-heavy visual frameworks to distribution-focused, interactive models represents a highly scalable, low-cost intervention capable of bridging perceptual gaps and cultivating a more resilient, collaborative democratic culture.

原文 (English)

Visualizing "We the People": Bridging the Perception Gap through Pluralistic Data Storytelling

Traditional visual data storytelling relies on binary graphics that depict two simplified groups in conflict. This can increase political polarization by oversimplifying intra-group disagreements and erasing ambiguity and shared ideas or values. This can inadvertently foster "us versus them" thinking. Intentional, pluralistic design choices for AI-enabled digital platforms can produce visualizations that emphasize nuance, opinion distribution, and intergroup commonalities. To demonstrate this potential, we examine deliberative technologies that map high-dimensional opinion spaces and highlight areas of both consensus and dissensus. The paper highlights the We the People deliberation conducted by Jigsaw and the Napolitan Institute in September 2025, which engaged over 2,400 Americans across all 435 congressional districts in an AI-supported, asynchronous dialogue regarding freedom and equality. By utilizing AI to synthesize long-form, text-based participant inputs into interactive "opinion landscapes," the initiative provided an alternative format for pluralistic data storytelling that humanized diverse viewpoints and revealed hidden areas of substantial broad consensus. The paper concludes that shifting from divisive, contrast-heavy visual frameworks to distribution-focused, interactive models represents a highly scalable, low-cost intervention capable of bridging perceptual gaps and cultivating a more resilient, collaborative democratic culture.

13:00 JSTLLM/生成AILlama

Small edits, large models: How Wikipedia advocacy shapes LLM values

Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia ap…

13:00 JST画像/動画生成

Noise-Aware Boundary-Enhanced Generative Learning for Ultrasound Speckle Reduction

Ultrasound is a non-invasive, real-time, and cost-effective imaging technique widely used in clinical diagnosis. However, its diagnostic ef…

13:00 JST画像/動画生成

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency…

13:00 JST研究/論文

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a chal…