Skip to the content.

AIニュース 2026-06-27

自動生成: 2026-06-27 12:56 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. AIモデル「ミュトス」、米国の一部組織に再提供へ 米政府が許可ITmedia AI+

    米Anthropicは6月26日(現地時間)、12日から提供を一時停止していたAIモデル「Claude Mythos 5」について、米国の…

  2. Trump Admin releases Anthropic Mythos to be used by more than 100 US companies, agenciesTechCrunch AI

    Over 100 companies and government agencies are reportedly authorized…

  3. 官民投資フィジカルAIに10.5兆円示す、「実証から実装へ」動き出す現場ITmedia AI+

    2026年6月22日~26日に公開された記事の中から、MONOist編集部が厳選した今週の注目ニュースをお届けします。

  4. OpenAI、次世代「GPT-5.6」シリーズを限定プレビュー 米政府と調整、命名は「Sol/Terra/Luna」に刷新ITmedia AI+

    米OpenAIは6月26日(現地時間)、次世代AIモデル「GPT-5.6」シリーズの限定プレビューを始めた。フラッグシップの「Sol」、日…

  5. 東電出資に意欲 孫正義氏が「国内データセンター誘致」で狙うインフラ戦略ITmedia AI+

    ソフトバンクグループ株主総会で、会長兼社長の孫正義氏が、将来的な目標として「純資産価値1000兆円」の展望を語った。AIインフラの最大のボ…

  6. OpenAI poaches Uber India chief to lead its biggest market outside the USTechCrunch AI

    The hire marks OpenAI's latest push into India, expanding offices, pa…

  7. Why everyone from OpenAI to SpaceX is building their own chips (and turning up the heat on Nvidia)TechCrunch AI

    Nvidia has dominated the AI chip market for years, but the era of tot…

トピック別件数

日本語メディア7件

ITmedia AI+ (日本語)

11:07 JSTLLM/生成AI規制/政策AnthropicClaude

AIモデル「ミュトス」、米国の一部組織に再提供へ 米政府が許可

米Anthropicは6月26日(現地時間)、12日から提供を一時停止していたAIモデル「Claude Mythos 5」について、米国の一部組織に限定して再提供を始めると発表した。米政府から許可を得たという。

08:00 JSTその他

東電出資に意欲 孫正義氏が「国内データセンター誘致」で狙うインフラ戦略

ソフトバンクグループ株主総会で、会長兼社長の孫正義氏が、将来的な目標として「純資産価値1000兆円」の展望を語った。AIインフラの最大のボトルネックである「電力確保」を巡り、子会社のソフトバンクが東京電力の次期オーナー候補に名乗りを上げている事実にも言及した。最先端データセンタ…

07:00 JSTビジネス/資金調達

官民投資フィジカルAIに10.5兆円示す、「実証から実装へ」動き出す現場

2026年6月22日~26日に公開された記事の中から、MONOist編集部が厳選した今週の注目ニュースをお届けします。

04:56 JSTLLM/生成AI規制/政策OpenAIGPT / ChatGPT3媒体が報道

OpenAI、次世代「GPT-5.6」シリーズを限定プレビュー 米政府と調整、命名は「Sol/Terra/Luna」に刷新

米OpenAIは6月26日(現地時間)、次世代AIモデル「GPT-5.6」シリーズの限定プレビューを始めた。フラッグシップの「Sol」、日常業務向けでバランス型の「Terra」、高速・低価格の「Luna」の3モデルで構成する。コーディングや科学、サイバーセキュリティの能力を高め…

出典:ITmedia AI+OpenAITechCrunch AI
18:15 JSTLLM/生成AI

「AIを使うと他社と似てしまう」課題をどう乗り越える? 「プロダクトの差別化」の要点

生成AIの普及は製品の没個性化や、個人の生産性向上によるチームの分断という課題を生んでいる。米Figmaはカンファレンスで、AI出力を人間が微調整する「素材」として扱う手法を提示。自社ルールを組織全体で共有する仕組みを実装した。個人の暗黙知を資産化する取り組みは、現場の属人化を…

13:35 JSTその他

防衛省は“認知戦”にどう挑む ウクライナ脅かすAIフェイク、偽アカウントへの対応は 分析資料を公開

防衛省は6月26日、「防衛力変革推進本部」での議論に関する資料を公表した。偽情報で相手の判断をゆさぶる「認知戦」への対応方針として、戦略的な情報発信機能やAI活用、情報関連機能の強化を打ち出した。

13:00 JSTLLM/生成AIエージェントOpenAIGPT / ChatGPT

「OpenAIはAzureだけ」の時代が終了 「GPT-5.5」「Codex」をAWSで利用するメリットは何か

OpenAIとAmazon Web Services(AWS)が戦略的パートナーシップを拡大した。OpenAIのモデル、コーディングエージェント「Codex」、マネージドエージェントを、企業がAWS環境で利用できるようにする。

海外メディア5件

TechCrunch AI (英語)

10:01 JSTLLM/生成AIAnthropic

Trump Admin releases Anthropic Mythos to be used by more than 100 US companies, agencies

Over 100 companies and government agencies are reportedly authorized to use Mythos 5, including their non-American employees.

03:19 JSTLLM/生成AIOpenAI

OpenAI poaches Uber India chief to lead its biggest market outside the US

The hire marks OpenAI's latest push into India, expanding offices, partnerships and hiring.

02:43 JSTLLM/生成AIハードウェア/半導体OpenAINVIDIA2件の関連記事

Why everyone from OpenAI to SpaceX is building their own chips (and turning up the heat on Nvidia)

Nvidia has dominated the AI chip market for years, but the era of total dependence might be ending. OpenAI just shared its plans to spice t…

出典:TechCrunch AITechCrunch AI
01:24 JSTLLM/生成AIAnthropicOpenAI

It’s not about Anthropic vs. OpenAI anymore

AI models have progressed to the point where their capabilities have real political consequences. Dealing with those consequences will requ…

22:00 JSTその他

Early Bird pricing ends tonight for TechCrunch Founder Summit

Save up to $190 on your pass to TechCrunch Founder Summit 2026. Early Bird pricing ends today, at 11:59 p.m. PT, after which rates increase…

公式ブログ0件

OpenAI (英語)

新着記事はありませんでした。

論文277件

arXiv cs.AI (英語)

13:00 JST研究/論文

カスケード線形特徴によるお調子者の検出と制御

アクティベーションステアリング手法を通じてモデルの動作を解釈および制御するには、望ましい動作または望ましくない動作を明確に示す対照的なサンプルの多くのペアが必要です。これらのデータ ペアによって、解釈可能性フレームワークが動作の原因となるモデルの特徴をどの程度確実に検出できるかが決まり、したがって、モデルをそのような動作に近づけるか遠ざけることができるかが決まります。この研究では、動作の原因となるカスケード線形特徴を分離する反復データ生成パイプラインを紹介します。具体的には、サンプルの単純なバイナリ ペアを超えて、代わりに動作に線形にスケールする特徴の度合いを示すサンプルを分離することで、特徴のもつれをより良く解くことができることを示します。私たちは、ユーザーの検証を優先する言語モデルの傾向であるお調子者を検出し、回避することに重点を置いています。我々は、カスケードサンプルを通じて発見されたお調子者の特徴が線形分離可能な部分空間を形成し、ベースラインのアプローチよりも目的の動作により明確に対応するモデル活性化の選択を可能にすることを実証します。また、検出、決定論的なスコアリング、堅牢なステアリングを可能にする機能も評価し、LLM-as-a-judge およびシステム プロンプト ベースラインと同等またはそれを上回るパフォーマンスを示しながら、計算量の削減と解釈可能性の保証の向上を実現していることを確認しました。コードとデータ: https://cascading-feats.github.io/

原文 (English)

Detecting and Controlling Sycophancy with Cascading Linear Features

Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. These data pairs determine the degree to which interpretability frameworks can reliably detect model features responsible for a behavior, and therefore the ability to steer models toward or away from such behavior. In this work, we present an iterative data generation pipeline that isolates cascading linear features responsible for a behavior. Specifically, we show how moving beyond simple binary pairs of samples, and instead isolating samples that show degrees of features that scale linearly with behavior, allows for better disentanglement of features. We focus on detecting and steering away from sycophancy -- the tendency of language models to prioritize user validation. We demonstrate that sycophancy features discovered through cascading samples form linearly separable subspaces, and allow for selection of model activations that more clearly correspond to the desired behavior than baseline approaches. We also evaluate their ability to enable detection, deterministic scoring, and robust steering, and see that they either match or outperform LLM-as-a-judge and system prompting baselines while providing lower computational demand and more interpretability guarantees. Code & Data: https://cascading-feats.github.io/

13:00 JST研究/論文

ベンチマーク飽和後の生活: CORE-Bench のケーススタディ

ベンチマークの精度が飽和すると、多くの場合、そのベンチマークは廃止され、より困難なバージョンに置き換えられます。我々は、このアプローチが精度を優先し、エージェントのパフォーマンスの他の 6 つの主要な側面を研究する機会を逃していることを示します。つまり、ショートカット、分布外の一般化可能性、効率、信頼性、モデルと足場の相対的な重要性、人間とエージェントのコラボレーションによる向上などの構築妥当性の問題です。私たちは、科学コードの計算再現性のベンチマークである CORE-Bench Hard をケーススタディとして使用し、これらの次元に沿ってエージェントを測定すると、精度が飽和した後でもエージェントのパフォーマンスについて有意義な洞察が得られることを実証します。まず、CORE-Bench Hard で妥当性を構築するために、能力の低いエージェントでは予測することが難しい脅威を表面化します。改良されたベンチマークである CORE-Bench v1.1 と、配布外のタスク スイートである CORE-Bench OOD を導入します。次に、精度が飽和しているにもかかわらず、CORE-Bench v1.1 は効率、信頼性、モデルのパフォーマンス、および足場のパフォーマンスを測定するのに依然として有用であることがわかりました。最後に、実世界の計算再現性タスクにおける人間とエージェントのコラボレーションによる向上を測定するために、小規模なランダム化実験を実施します。私たちは、約 2 倍の統計的に有意な速度向上を発見しました (人間のみによる複製の 5 分の 1 が完了する前に制限時間に達しているため過小評価されている可能性があります) と、その他のさまざまな発見について説明します。私たちの貢献は、支配的な精度中心の評価パラダイムに対するより厳密な代替案を提示します。

原文 (English)

Life After Benchmark Saturation: A Case Study of CORE-Bench

When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity to study six other key dimensions of agent performance: construct validity issues such as shortcuts, out-of-distribution generalizability, efficiency, reliability, the relative importance of the model versus the scaffold, and uplift from human-agent collaboration. We use CORE-Bench Hard, a benchmark for computational reproducibility of scientific code, as a case study to demonstrate that measuring agents along these dimensions yields meaningful insights into agent performance even after accuracy saturates. First, we surface threats to construct validity in CORE-Bench Hard that are difficult to anticipate with less capable agents. We introduce an improved benchmark, CORE-Bench v1.1, and an out-of-distribution task suite, CORE-Bench OOD. Second, we find that despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring efficiency, reliability, model performance, and scaffold performance. Finally, we conduct a small-scale randomized experiment to measure uplift from human-agent collaboration on real-world computational reproducibility tasks. We find a statistically significant speedup by about a factor of two -- likely underestimated due to one-fifth of human-only reproductions reaching the time limit before completing -- and describe various other findings. Together, our contributions present a more rigorous alternative to the dominant accuracy-centric evaluation paradigm.

13:00 JST研究/論文Llama

拒否はチャット モデルのペルソナの下流に存在します

活性化空間における線形方向は、指示調整型チャット モデルにおける拒否特性とペルソナ特性の両方について特定されていますが、この 2 つは別個のメカニズムとして研究されてきました。私たちは、彼らが相互作用することを示します。つまり、従順なペルソナがゲートを拒否するということです。 Qwen2.5-7B-Instruct と Llama-3.1-8B-Instruct では、準拠モデルペルソナの方向と拒否の方向を抽出し、両方に介入します。準拠したペルソナのステアリングにより拒否が抑制されます。ラマでは、拒否率が 97% から 2% に低下しました。拒否の方向を再導入すると、後の層での拒否が部分的に回復しますが、初期の層では回復しません。後期レイヤーウィンドウでペルソナの方向を投影すると、それがベースラインに戻ります。ランダムな方向に投影することはできません。したがって、拒否は、それが計算される場所の下流にある、後期層の表現段階でゲートされます。拒否を単一の孤立した方向として扱うと、人格への依存性が失われます。

原文 (English)

Refusal Lives Downstream of Persona in Chat Models

Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate mechanisms. We show they interact: a compliant persona gates refusal. In Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, we extract a compliant model-persona direction and a refusal direction and intervene on both. Compliant persona steering suppresses refusal -- in Llama, the refusal rate falls from 97% to 2%. Reintroducing the refusal direction partially restores refusal at late layers but not at early ones. Projecting out the persona direction in a late-layer window restores it to baseline; projecting out a random direction does not. Refusal is therefore gated at the late-layer expression stage, downstream of where it is computed. Treating refusal as a single isolated direction misses its dependence on persona.

13:00 JSTLLM/生成AI

AlgoEvolve: LLM 主導のアルゴリズム取引プログラムのメタ進化

最近の研究では、大規模言語モデル (LLM) がプログラムと証明の進化的発見のための意味論的突然変異演算子として機能できることが示されています。現在のアプリケーションのほとんどは静的コーディングのベンチマークに重点を置いています。私たちはこのパラダイムをアルゴリズム取引に拡張します。このドメインは、ノイズが多く、非定常で、非常に不連続であるため、独特の課題を抱えています。私たちは、実行可能な取引戦略を生成、評価、反復的に改善する LLM 主導の進化的フレームワークである AlgoEvolve を紹介します。これらの戦略は Python コードとして表現され、厳格なテスト プロトコルを通じて評価されます。複数の実験を通じて、このシステムは、取引ルールの自律的な変更を含む、新たな体制適応戦略ロジックを示しました。さらに、内部ループでプログラム合成を導くプロンプトを進化させるメタ進化的な外部ループを導入します。この外側のループは、改善された検索ヒューリスティックを発見します。これらのヒューリスティックは、ゼロトレードの失敗を減らしながら、探査と活用のバランスをとります。これらは、人間が設計した最初の指示よりも常に優れたパフォーマンスを発揮します。この結果は、LLM ベースのセマンティック進化が、複雑な環境における継続的なプログラム合成に実行可能なアプローチを提供することを示しています。

原文 (English)

AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs

Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and proofs. Most current applications focus on static coding benchmarks. We extend this paradigm to algorithmic trading. This domain is uniquely challenging because it is noisy, non-stationary, and highly discontinuous. We present AlgoEvolve, an LLM-driven evolutionary framework that generates, evaluates, and iteratively improves executable trading strategies. These strategies are expressed as Python code and evaluated through a rigorous testing protocol. Across multiple experiments, the system exhibits emergent regime-adaptive strategy logic, including autonomous shifts in trading rules. We further introduce a meta-evolutionary outer loop that evolves the prompts guiding program synthesis in the inner loop. This outer loop discovers improved search heuristics. These heuristics balance exploration and exploitation while reducing zero-trade failures. They consistently outperform initial human-designed instructions. The results demonstrate that LLM-based semantic evolution provides a viable approach for continual program synthesis in complex environments.

13:00 JSTLLM/生成AIエージェントGoogle

エージェントティック インフラストラクチャのためのエージェント分析: DAO と企業 AI プロトコルの比較ガバナンスのための LLM を利用したパイプライン

AI エージェント プロトコルが急増する一方で、その相互運用性標準を形成するガバナンス構造は経験的に十分に検討されていないままです。私たちは、大規模なガバナンス談話分析のための LLM を利用した比較パイプラインを導入し、自動アノテーション、ニューラル トピック モデリング、および多層ネットワーク分析を統合して、社会技術的な権力構造を大規模に研究します。当社では、エージェントの相互運用性に関する 2 つの対照的な標準、ERC-8004 (パーミッションレス、オンチェーン) と Google A2A (企業主導) に基づいて検証しています。 4,323 件のガバナンス参加記録を分析し、LLM 支援コーディング、トピック モデリング、および多層ネットワーク分析を組み合わせて、制度設計がテーマの優先順位とコミュニティ構造をどのように形成するかを調査します。私たちは、ガバナンスの形態が実質的な焦点に影響を与える一方で、どちらの体制も同等のレベルの参加不平等とコミュニティの断片化を示していることを発見しました。パーミッションレス環境では言説の整合性がより密になっており、分散型の参加にもかかわらず、オープンなガバナンスがより大きなテーマの収束を促進する可能性があることを示唆しています。これらの発見は、LLM 支援手法がテクノロジー ガバナンスの実証研究をどのように前進させることができるかを示しており、より公平なエージェント AI 標準の設計に影響を及ぼします。すべてのデータとコードはオープンに利用できます。

原文 (English)

Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures at scale. We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led). Analyzing 4,323 governance participation records, we combine LLM-assisted coding, topic modeling, and multi-layer network analysis to examine how institutional design shapes thematic priorities and community structure. We find that while governance form influences substantive focus, both regimes exhibit comparable levels of participation inequality and community fragmentation. Discourse alignment is denser in the permissionless setting, suggesting that open governance may foster greater thematic convergence despite decentralized participation. These findings illustrate how LLM-assisted methods can advance the empirical study of technology governance, with implications for designing more equitable agentic AI standards. All data and code are openly available.

13:00 JSTエージェント

メンタルヘルスの薬剤情報探索のための知識拡張型エージェント AI

患者はますますオンラインで医薬品情報を求めるようになっているが、精神科薬の安全性に関する知識は、権威あるが抽象的な規制上の有害事象記録と、経験に近いが検証されていない患者の語りとに分かれている。証拠と逸話を混同することなくそれらを統合することは、文脈が不十分な情報によって恐怖、ノーシーボ反応、不遵守が増幅される可能性がある精神医学において特に重要です。ここでは、9 つ​​の抗うつ薬に関する 466,525 件の Reddit 投稿、60,782 件の WebMD レビュー、および 20 年間にわたる米国 FDA 有害事象報告システムの記録を統合した、出所を意識したナレッジグラフベースのマルチエージェント フレームワークを開発しました。医師の注釈に対してベンチマークされた大規模言語モデルの実体認識パイプラインは、薬物の場合は 0.969、症状の場合は 0.973 という最高の F1 スコアに達しました。 2 つのコミュニティ プラットフォームは、規制報告書よりも相互にはるかに一致しており (Jaccard 類似度 0.905 まで重複)、患者生成データが部分的に独立した安全性シグナルを形成していることを示しています。セルトラリンについては、対応する FDA 日付の数百日前に多くの有害事象がコミュニティ情報源に現れました。 ATC-N、ICD-10、および MedDRA の語彙に基づいた Neo4j ナレッジ グラフは出所を保存し、すべての請求を追跡可能に保ち、規制上の事実を患者の経験から区別します。これらの結果は、より監査可能な精神科治療情報へのルートとしてソースを意識した統合を確立し、有用性と患者利益を前向きにテストする必要があります。

原文 (English)

Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records, which are authoritative but abstract, and patient narratives, which are experience-near but unvalidated. Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly contextualised information can amplify fear, nocebo responses, and non-adherence. Here we develop a provenance-aware, knowledge-graph-based multi-agent framework unifying 466,525 Reddit posts, 60,782 WebMD reviews, and twenty years of U.S. FDA Adverse Event Reporting System records for nine antidepressants. A large-language-model entity-recognition pipeline benchmarked against physician annotations reached highest F1 scores of 0.969 for medications and 0.973 for conditions. The two community platforms were far more concordant with each other (overlap up to a Jaccard similarity of 0.905) than with regulatory reports, indicating that patient-generated data form a partly independent safety signal. For sertraline, many adverse events appeared in community sources hundreds of days before the corresponding FDA date. A Neo4j knowledge graph grounded in ATC-N, ICD-10, and MedDRA vocabularies preserves provenance, keeping every claim traceable and regulatory facts distinct from patient experience. These results establish source-aware integration as a route to more auditable psychiatric medication information, with usefulness and patient benefit to be tested prospectively.

13:00 JST研究/論文

チェスのスキル評価の加速: ドリフト拡散を強化した Elo 評価システム

Elo などのレーティング システムは、競技チェスのマッチメイキングのゴールド スタンダードとして機能します。ただし、試合の結果のみに依存し、ゲームプレイの細かな品質を無視しているため、本質的に応答の遅れに悩まされています。それにもかかわらず、レーティング調整に手ごとの情報を組み込むことは、かなりのノイズとゲーム状態空間の広大さを考慮すると、大きな課題となります。これに対処するために、認知神経科学のドリフト拡散モデル (DDM) にヒントを得た新しいスキル評価フレームワークであるドリフト拡散強化 Elo 評価システム (DD-Elo) を提案します。スキル表現を意思決定プロセスとしてモデル化することで、私たちのモデルは技レベルのデータを統合して、スキルの急速な変動を捉えます。私たちは、DD-Elo が従来の Elo システムからの一定の偏差を維持し、理論的な整合性を確保していることを証明する厳密な数学的導出を提供します。広範な実験により、DD-Elo は Elo よりも早くスキルの変化に適応することが実証されました。私たちの調査結果は、DD-Elo がチェス レーティング エコシステムに対して、説明可能で応答性が高く、下位互換性のあるソリューションを提供することを示唆しています。実装コードは https://github.com/Aquila-zhou1/DD-Elo で公開されています。

原文 (English)

Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System

Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess. However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the granular quality of gameplay. Nevertheless, incorporating move-by-move information into rating adjustments presents a significant challenge given the substantial noise and the vastness of the game-state space. To address this, we propose the Drift-Diffusion-Enhanced Elo Rating System (DD-Elo), a novel skill assessment framework inspired by the drift diffusion model (DDM) from cognitive neuroscience. By modeling skill expression as a decision-making process, our model integrates move-level data to capture rapid skill fluctuations. We provide a rigorous mathematical derivation proving that DD-Elo maintains a bounded deviation from the traditional Elo system, ensuring theoretical alignment. Extensive experiments demonstrate that DD-Elo adapts to skill changes faster than Elo. Our findings suggest that DD-Elo offers an explainable, highly responsive, and backward-compatible solution for chess rating ecosystems. The implementation code is publicly available at https://github.com/Aquila-zhou1/DD-Elo .

13:00 JSTエージェント

エージェントではなく、行動を統治する: 自律型 AI システムのガバナンス モデルとしての機関の認証

自律型 AI エージェントは、臨床処方や実稼働ソフトウェアの導入など、結果として取り消せないアクションを実行し始める可能性があります。この論文は、人間の制度が強力な自律的主体を、彼らの推論を監視することによってではなく、結果的な行動の時点で独立して証明された証拠を要求することによって統治してきたことを観察しています。私たちは、この制度的パターンを AI エージェント システムの計算ガバナンス モデルとして形式化します。提案されたモデルでは、エージェントは計画と推論に関して完全な自律性を保持しますが、指定された高リスクのアクションについては実行権限を持ちません。実行は、個別の信頼できるソースによってそれぞれ独立して証明され、宣言された意図に暗号的に結び付けられ、決定論的なポリシーによって評価される前提条件に基づいて行われます。決定は、独立した再検証に適した改ざん防止ログに記録されます。概念実証の実装を示し、ソフトウェアの導入と臨床処方の例を使用してモデルを説明します。

原文 (English)

Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems

Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment. This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning but by requiring independently attested evidence at the point of consequential action. We formalise this institutional pattern as a computational governance model for AI agent systems. Under the proposed model, an agent retains full autonomy over planning and reasoning but holds no execution authority over designated high-risk actions. Execution is conditional on preconditions that are each independently attested by a separate authoritative source, cryptographically bound to a declared intent, and evaluated by a deterministic policy. Decisions are recorded in a tamper-evident log amenable to independent re-verification. We present a proof-of-concept implementation and illustrate the model with examples from software deployment and clinical prescribing.

13:00 JST研究/論文

COrigami: 平坦折り可能な視覚的に認識可能な折り紙を共同設計するための AI パイプライン

生成 AI は検証可能なソリューションによる問題解決で目覚ましい成功を収めてきましたが、厳密な幾何学的制約と主観的な視覚美の両方を満たす物理アートを生成することは依然として課題です。本稿では、平坦折り可能性の方程式内で芸術的デザインを基礎づける数学的に厳密な環境である計算折り紙の領域におけるこれらの困難に取り組むアプローチを紹介します。 COrigami は、自然言語から折り目パターンを生成することで設計サイクルを支援する、エンドツーエンドの AI 駆動パイプラインです。私たちのパイプラインには、セマンティック スティック フィギュアの生成、基本パッキングの計算、平坦折り可能な折り目パターンの解決、平坦折り折り目パターンの整形、および自律的な美的評価ループによって駆動される強化学習を使用した生成されたモデルの改良が含まれます。私たちのシステムは非常に効果的な共同アシスタントとして機能し、人間のアーティストがさらに拡張して形を整えることができる構造的な出発点を生成します。この研究では、アルゴリズムの最適化と自律的な美的批評を統合することにより、AI システムが多目的の物理的制約をどのように満たして信頼性の高い、数学的に根拠のある共同創造性を実現できるかを示しています。

原文 (English)

COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict geometric constraints and subjective visual aesthetics remains a challenge. This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic design within the equations of flat foldability. We present COrigami, an end-to-end AI-driven pipeline that assists the design cycle by generating crease patterns from natural language. Our pipeline involves generating a semantic stick figure, computing a base packing, solving for a flat-foldable crease pattern, shaping the flat-folded crease pattern, and refining the generated model using reinforcement learning driven by an autonomous aesthetic evaluation loop. Our system acts as a highly effective collaborative assistant, generating structural starting points that human artists can further expand and shape. By integrating algorithmic optimisation with autonomous aesthetic critique, this work demonstrates how AI systems can satisfy multi-objective physical constraints to enable reliable, mathematically grounded co-creativity.

13:00 JSTLLM/生成AIエージェント

検証の地平線: コーディング エージェントの報酬に特効薬はない

古典的な直観では、解決策を生み出すよりも検証する方が簡単だと考えられています。今日のコーディング エージェントにとって、この直感は逆転しつつあります。基礎モデルがより強力な推論機能を開発し、エンジニアリング ハーネスがより洗練されるにつれて、複雑な候補ソリューションを生成することはもはや難しくなくなり、それらを確実に検証することがより困難な問題になりました。私たちが構築できるすべての検証ツールは人間の意図の代理にすぎず、意図そのものではありません。このため、検証は 2 つの困難を伴います。1 つは、本質的に意図が過少指定されているため、意図が満たされているかどうかを忠実に確認することが本質的に困難であるということです。次に、モデルのトレーニング中に、最適化によってプロキシとインテントの間のギャップが広がり、報酬のハッキングや信号の飽和として現れます。これに対処するために、私たちは検証信号の品質を 3 つの次元 (スケーラビリティ、忠実性、堅牢性) に沿って特徴付け、3 つすべてを同時に達成することが中心的な課題であると主張します。さらに、一般的なコーディング タスク用のテスト検証器、フロントエンド タスク用のルーブリック検証器、現実世界のエージェント タスク用の検証器としてのユーザー、長期タスク用の自動エージェント検証器の 4 つの報酬構造を研究します。さまざまなタスクの種類とポリシーの機能レベルにわたって、報酬設計の中核となる課題と、報酬シグナルをより効果的に活用する方法について、綿密な分析と実験を実施します。実験では、ターゲットを絞った検証設計により、報酬ハッキングを効果的に抑制し、タスク完了の品質を向上させ、複数の内部および公開ベンチマーク全体で大幅な利益を達成できることが示されています。これらの経験は総合的に、政策能力が成長し続けるにつれて、固定報酬関数が有効であり続けることはできないという核心的な観察を示しています。そして検証はジェネレーターと共進化する必要があります。

原文 (English)

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabilities and engineering harnesses grow more sophisticated, generating complex candidate solutions is no longer difficult -- reliably verifying them has become the harder problem. Every verifier we can build is only a proxy for human intent, never the intent itself. This makes verification subject to a twofold difficulty: first, intent is underspecified by nature, making it inherently hard to faithfully check whether it has been fulfilled; second, during model training, optimization widens the gap between proxy and intent -- manifesting as reward hacking or signal saturation. To address this, we characterize the quality of verification signals along three dimensions -- scalability, faithfulness, and robustness -- and argue that achieving all three simultaneously is the central challenge. We further study four reward constructions: a test verifier for general coding tasks, a rubric verifier for frontend tasks, the user as verifier for real-world agent tasks, and an automated agent verifier for long-horizon tasks. Across different task types and policy capability levels, we conduct in-depth analysis and experiments on the core challenges of reward design and how to more effectively leverage reward signals. Experiments show that targeted verification design can effectively suppress reward hacking, improve task completion quality, and achieve significant gains across multiple internal and public benchmarks. These experiences collectively point to a core observation: no fixed reward function can remain effective as policy capability continues to grow; and verification must co-evolve with the generator.

13:00 JSTLLM/生成AIエージェント研究/論文

ツール拡張 LLM エージェントは現実世界のエネルギー分析タスクをどのように実行しますか?

エージェントベンチマークは、金融、コーディング、法律、創薬などの汎用および分野固有の設定にわたって登場していますが、エネルギー分野の評価は依然として静的な知識の想起に主に限定されています。これは、ライブデータの取得、専門的な規制と市場の知識、現実世界の制約の下での複数段階の定量的推論を必要とするセクターにとって、重大なギャップです。現実世界のエネルギー市場分析タスクにおけるツール拡張 LLM エージェントの実証研究を紹介します。当社の評価環境には、(1) 市場データの取得と分析、(2) 知識の取得と解釈、(3) 高度な定量モデリングと意思決定分析の 3 つのカテゴリにわたる、専門家が厳選した 243 の問題が含まれています。タスクには、価格と需要の分析、料金の影響モデリング、資産収益と収益の推定、ヘッジ戦略分析、最適化モデリングが含まれ、問題は複数の難易度にまたがります。エージェントには、米国の主要 ISO 用のライブ電力市場 API、規制書類検索、公共料金データベース、資産最適化モデル、エネルギー市場ドキュメントの検索拡張発電などを含む、構成可能なドメイン ツール スイートが装備されています。当社は、アプローチの正しさ、回答の正確さ、属性の整合性、およびソースの妥当性をスコアリングする多次元評価プロトコルを使用してエージェントの応答を評価します。スコア基準を質問の種類に一致させるためのカテゴリを認識したルーティングを使用します。私たちはクローズドソースとオープンソースの LLM の両方を評価し、一か八かの専門分野でモデルの機能とドメイン ツールがどのように相互作用するかを比較分析します。主要なアーティファクトは、再現性と将来の研究をサポートするために公開されています。

原文 (English)

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-domain evaluations remain largely limited to static knowledge recall. This is a critical gap for a sector that requires live data retrieval, specialized regulatory and market knowledge, and multi-step quantitative reasoning under real-world constraints. We present an empirical study of tool-augmented LLM agents on real-world energy market analytics tasks. Our evaluation environment includes 243 expert-curated problems across three categories: (1) Market Data Retrieval and Analysis, (2) Knowledge Retrieval and Interpretation, and (3) Advanced Quantitative Modeling and Decision Analytics. Tasks include price and demand analysis, tariff impact modeling, asset revenue and returns estimation, hedging strategy analysis, and optimization modeling, with problems spanning multiple difficulty levels. Agents are equipped with a configurable suite of domain tools, including live electricity market APIs for major U.S. ISOs, regulatory docket search, utility tariff databases, asset optimization models, and retrieval-augmented generation over energy market documents. We assess agent responses using a multi-dimensional evaluation protocol that scores approach correctness, answer accuracy, attribute alignment, and source validity, with category-aware routing to match scoring criteria to question type. We evaluate both closed-source and open-source LLMs, providing a comparative analysis of how model capability and domain tooling interact in a high-stakes professional domain. Key artifacts are publicly released to support reproducibility and future research.

13:00 JSTLLM/生成AIビジネス/資金調達

マルチモーダル LLM 評価に欠けているものは何ですか?

マルチモーダル大規模言語モデル (MLLM) は、テキスト、画像、音声、ビデオなどのさまざまな入力を処理し、テキスト応答を生成できます。それらの機能は急速に進歩していますが、そのようなモデルの評価は追いついていません。既存の評価ベンチマークのほとんどは、個別のタスクに限定されており、モデルがモダリティ全体で情報を統合しているかどうかについてはほとんど明らかにされていません。私たちは、MLLM を評価するための現在の手段を調査し、既存のベンチマーク分類をレビューして、時間空間的一貫性、物理世界の理解、マルチモーダル一貫性、選択的注意などのギャップを特定します。これらのギャップに対処することは、マルチモーダル インテリジェンスの実際の進歩を測定し、機能の境界を明らかにするために不可欠です。

原文 (English)

What We are Missing in Multimodal LLM Evaluation?

Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced rapidly, evaluation of such models has not kept pace. Most existing evaluation benchmarks are limited to isolated tasks and reveal little about whether a model integrates information across modalities. We examine current means for evaluating MLLMs and review the existing benchmark taxonomy to identify gaps, including temporal-spatial coherence, physical world understanding, multimodal consistency, and selective attention. Addressing these gaps is essential for measuring real progress in multimodal intelligence and exposing capability boundaries.

13:00 JSTエージェントビジネス/資金調達

OpenFinGym: クオンツエージェントを評価するための検証可能なマルチタスクジム環境

大規模な言語モデル エージェントは定量的財務ワークフローにますます適用されていますが、その評価は分離されたタスク間で断片化されたままであり、ベンチマーク タスクの財務関連性はしばしば見落とされます。しかし、財務ワークフローは本質的に多段階であり、予測、戦略構築、リスク管理、取引などの相互依存するタスクにまたがっています。既存のプラットフォームは通常、単一のタスクに焦点を当てているため、エージェントの能力を過大評価し、一般化、実際の市場でのやり取り、および財務的に意味のある意思決定における弱点を明らかにできません。 OpenFinGym は、単一の実行および検証インターフェイスで予測、市場生成、リ​​アルタイム取引、不正検出をカバーする定量的金融エージェント開発用の統合ジム環境です。 OpenFinGym はさらに、定量的な財務出版物を実行可能なタスク パッケージに変換する自動タスク構築パイプラインを提供します。スケーラブルなエージェントのロールアウトをサポートし、ランタイムのトレインテストの漏洩を防ぐホスト側検証サービスを備えたコンテナ化されたランタイム。低レイテンシのデータストリーム設計を備えたペーパートレーディングエンジン。長期およびイベント市場の予測に対する遅延解像度のサポート。トレーニング後の SFT と RL の統合

原文 (English)

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark tasks is often overlooked. Yet financial workflows are inherently multi-stage, spanning interdependent tasks such as forecasting, strategy construction, risk management, and trading. Existing platforms typically focus on a single task, and can therefore overstate agent competence and fail to reveal weaknesses in generalization, real-market interaction, and financially meaningful decision-making. We introduce OpenFinGym, a unified gym environment for quantitative-finance agent development that covers forecasting, market generation, real-time trading, and fraud detection under a single execution and verification interface. OpenFinGym additionally provides an automated task-construction pipeline that turns quantitative finance publications into executable task packages; a containerised runtime with a host-side verifier service that supports scalable agent rollouts and prevents runtime train-test leakage; a paper trading engine with a low-latency data-stream design; deferred-resolution support for long-horizon and event-market forecasts; and integration for SFT and RL post-training

13:00 JSTLLM/生成AIエージェントClaude

命令ブリード: プロンプト構成エージェント システムにおけるモジュール間干渉

プロンプト構成エージェント システムの実践者は、繰り返し発生する障害モードを報告しています。つまり、共有変数や実行可能ファイルの依存関係がないにもかかわらず、1 つのプロンプト モジュールを編集すると、他のプロンプト モジュールの動作が静かに変更されます。私たちはこれを構成的動作漏洩 (CBL)、つまりコンテキスト ウィンドウを共有するモジュール間の干渉として形式化します。 CBL はアーキテクチャ上の非絶縁によって有効になります。トランスのセルフアテンションにより、連結されたモジュール間に正式な境界がありません。ボリューム、コンテンツ、形式に沿って非焦点モジュールを混乱させる再利用可能な 3 チャネル プロトコルを通じて、デプロイされたジョブ評価エージェント (Claude Sonnet 4.6、144 トライアル) で CBL を調査します。コンテンツ チャネルのみが検出可能な一対の効果を生成します (コーエンの d = 0.63、ゼロを除くブートストラップ 95% CI)。推奨事項が反転することはありません。標準の QA では目に見えない閾値以下の体制ですが、配置されたエージェントが下す何千もの意思決定を複雑にします。 CBL は、既知のエージェント障害軸 (敵対的注入、認知機能低下、マルチエージェント障害伝播、プライバシー漏洩) と直交しています。私たちは、運用定義、再利用可能なプロトコル、改ざん可能な予測セット、およびシステムクラスの特性評価に貢献し、迅速に作成されたエージェント評価の要件としてモジュール間干渉測定を確立します。

原文 (English)

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of others despite no shared variable or executable dependency. We formalize this as compositional behavioral leakage (CBL): interference between modules sharing a context window. CBL is enabled by architectural non-isolation: transformer self-attention provides no formal boundary between concatenated modules. We probe CBL on a deployed job-evaluation agent (Claude Sonnet 4.6, 144 trials) through a reusable three-channel protocol that perturbs non-focal modules along volume, content, and form. Only the content channel produces a detectable paired effect (Cohen's d = 0.63, bootstrap 95% CI excluding zero); no recommendation flipped -- a sub-threshold regime invisible to standard QA but compounding across the thousands of decisions a deployed agent makes. CBL is orthogonal to known agent-failure axes (adversarial injection, cognitive degradation, multi-agent fault propagation, privacy leakage). We contribute an operational definition, a reusable protocol, a falsifiable prediction set, and a system-class characterization, establishing cross-module interference measurement as a requirement for prompt-composed agent evaluation.

13:00 JST研究/論文

収益の加速と科学の定性エンジン

レイ・カーツワイルは、テクノロジーの進歩を議論する際に最も影響力のある物語である、収益の加速というテーゼについて説明しました。その中心的な主張は、複数の技術分野、特にコンピューティング、人工知能、脳科学、バイオテクノロジーの進歩が相互作用し、進歩が自己増幅的かつほぼ指数関数的になるというものです。この論文は、その主張の単純な数学的解釈を示し、そのような加速が現実であるとしても、それ自体では科学的発見の中心的な問題を解決するものではないと主張します。その理由は、収益の加速は実行能力とインフラストラクチャ能力に最も自然に適用されるのに対し、真の発見は別の能力、つまり、現在のフレームワークが構造的にいつ不適切であるか、次にどのような概念的な動きが必要であるかについての定性的推論に依存することが多いためです。最近の ARC-AGI-3 の結果は、この区別を明確にします。人間はベンチマークを天井で解くのに対し、フロンティア AI システムは 1% 未満に留まり、現在の AI と人間の柔軟な推論とのギャップが依然として非常に大きいことを示しています。同時に、デミス・ハサビス氏は、人間は意味の感覚を保持し、自分の人生の焦点を何に集中させるかを保持しなければならないと強調し、AIの将来は技術的な予測であるだけでなく、どのような形の人間の理解を保存し伝達する価値があるかという問題でもあることを思い出させます。この論文では、科学のための質的エンジン (QES) [3] を、不足している能力への対応策として位置づけています。この見解では、カーツワイル理論は、量的能力が加速する理由を説明するのに役立ちますが、QES は加速だけでは解決できない科学的発見の中心的な問題に対処します。その価値は、AGI がいつ到来するかによって決まるのではなく、科学的発見のプロセス自体が、保存し、整理し、アクセス可能にする価値のある人類の知恵の一形態を構成するという事実によって決まります。

原文 (English)

Accelerating Returns and the Qualitative Engine for Science

Ray Kurzweil described a thesis of accelerating returns, which is the most influential narratives in discussions of technological progress. Its central claim is that advances in multiple technological fields, especially compute, artificial intelligence, brain science, and biotechnology, interact in such a way that progress becomes self-amplifying and approximately exponential. This paper gives a simple mathematical interpretation of that claim and then argues that, even if such acceleration is real, it does not by itself resolve the central problem of scientific discovery. The reason is that accelerating returns apply most naturally to executional and infrastructural capability, whereas genuine discovery often depends on a different capacity: qualitative reasoning about when a current framework is structurally inadequate and what conceptual move is needed next. Recent ARC-AGI-3 results sharpen this distinction: humans solve the benchmark at ceiling, whereas frontier AI systems remain below 1%, indicating that the gap between current AI and human flexible reasoning is still very large. At the same time, Demis Hassabis has emphasized that humans must retain their sense of meaning and what they choose to focus their lives on, a reminder that the future of AI is not only a technical forecast but also a question of what forms of human understanding are worth preserving and transmitting. This paper positions the Qualitative Engine for Science (QES) [3] as a response to that missing capacity. In this view, the Kurzweil theory helps explain why quantitative capability may accelerate, while QES addresses the central problem in scientific discovery that acceleration alone does not solve. Its value does not depend on when AGI arrives, but on the fact that the processes of scientific discovery themselves constitute a form of human wisdom worth preserving, organizing, and making accessible.

13:00 JSTLLM/生成AI

思考のナレーション: 大規模言語モデルにおける実行不可能な倫理的推論のための推論時間足場

道徳的ジレンマに関する標準的な思考連鎖は、利害関係者の崩壊(結果に利害関係を持つ最大でも 1 つの当事者のみをトレース名で示す)と不確実性の抑圧(行動にコミットする前に明確な未知数やヘッジがない)という 2 つの失敗モードを示します。思考のナレーション (NoT) を導入します。これは、思考の連鎖を 5 つのセクション (主人公、利害関係者、2 段階の結果、不確実性、コミットメント) に構造化するシステム プロンプトです。 NoT では、トレーニング、パラメーター、微調整は追加されません。 3 ベンダーの 4 つのジェネレーターにわたる 100 の DailyDilemmas シナリオで、NoT はすべてのモデルでステークホルダーの崩壊を最大 31% から 1% 未満に、不確実性の抑制を最大 72% から 1 ~ 24% に削減しました。予算に合わせた詳細な CoT 制御により、有効成分としてのトークンの使用が除外されます。 NoT は、4 つのジェネレーターのうち 3 つについて、ステークホルダー数で +0.79 ~ +0.90、不確実性スコアで +0.65 ~ +0.93 というクリフのデルタ アドバンテージを保持しており、セクション アブレーションにより、各シフトがその特定のサブ命令に帰属します。 NoT で初期化されたテキスト勾配降下法により、足場がさらに改善されます。ファミリーを超えたトレーニングジャッジ(ジェネレーターとは別のベンダー)が、測定されたすべての軸においてファミリー内のトレーニングジャッジを支配します。 5 ラウンドのマルチステークホルダー討論プロトコルに拡張されたこの足場は、6% の対立をキャリブレーション セットの 95% の完全なコンセンサスと、DailyDilemmas の複製での 100% の結合収束に変換します。結果として得られるトレースは、各コミットメントの根拠となる利害関係者、結果、不確実性を外部化し、信頼性の高いエージェント展開のための監査可能な基盤を提供します。

原文 (English)

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppression (no explicit unknowns or hedges before committing to an action). We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections: protagonist, stakeholders, two-step consequences, uncertainty, then commitment. NoT adds no training, parameters, or fine-tuning. On 100 DailyDilemmas scenarios across four generators from three vendors, NoT cuts stakeholder collapse from up to 31% to under 1% and uncertainty suppression from up to 72% to 1-24% on every model. A matched-budget verbose-CoT control rules out token spend as the active ingredient; NoT retains Cliff's delta advantages of +0.79 to +0.90 on stakeholder count and +0.65 to +0.93 on uncertainty score for three of four generators, and a section ablation attributes each shift to its specific sub-instruction. Textual-gradient descent initialised at NoT improves the scaffold further; a cross-family training judge (different vendor from the generator) dominates an in-family one on every measured axis. Extended to a five-round multi-stakeholder debate protocol, the scaffold converts a 6% standoff into 95% full consensus on a calibration set and 100% combined convergence on a DailyDilemmas replication. The resulting traces externalise the stakeholders, consequences, and uncertainty grounding each commitment, providing an auditable substrate for dependable agentic deployment.

13:00 JST研究/論文

組み合わせ幾何学における極値問題のための幾何学認識 MCTS

私たちは、厳密でグローバルな幾何学的制約を満たす $n \times n$ グリッド内の点の構成を問う、組み合わせ幾何学における特定の極値問題を研究します。古典的な厳密ソルバーは、この種の問題に対して組み合わせ爆発に悩まされ、標準的な強化学習とトランスフォーマーベースのモデルは、報酬がまばらな「妥当性の崖」と二次トークン消費制限に悩まされます。これらのボトルネックを克服するために、Geometry-Aware Monte Carlo Tree Search (MCTS) フレームワークを提案します。私たちのアプローチは、実行可能なアクション空間への増分更新を通じて幾何学的制約を厳密に強制します。古典的な No-Three-in-Line 問題 (Max-N3IL) で発生するような、同一線上にある点の集合に関する制約の場合、このメカニズムにより、制約チェックの複雑さが $O(n^3)$ から $O(n^2)$ に軽減されます。検索効率を向上させるために、2 つの方法で幾何学的対称性を活用します。1 つはノード拡張時の標準枝刈りで分岐係数を削減し、もう 1 つは対称バッチ遷移で有望な構成の発見を加速します。私たちは広範な実験を実施し、検討した問題のうち 6 つのうち 5 つについて、新しく最もよく知られた計算結果を確立しました。特に、Max-N3IL では、サイズ $82 \le n \le 119$ のグリッドに対して、およそ $1.8 n$ のサイズの構成が見つかります。最小完全集合問題では、およそ $0.95 n$ のサイズの構成が見つかり、テストされたグリッド内に新しい上限が提供されます。この研究は、組み合わせ幾何学における新しい構成を発見するための適応性の高いフレームワークとして、幾何学認識 MCTS を確立します。

原文 (English)

Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry

We study certain extremal problems in combinatorial geometry that ask about configurations of points in an $n \times n$ grid that satisfy strict, global geometric constraints. Classical exact solvers suffer from combinatorial explosion for these types of problems, and standard reinforcement learning and transformer-based models struggle with the sparse reward "validity cliff" and quadratic token-consumption limits. To overcome these bottlenecks, we propose a Geometry-Aware Monte Carlo Tree Search (MCTS) framework. Our approach strictly enforces geometric constraints through incremental updates to the feasible action space. For constraints about collections of collinear points, like those that occur in the classic No-Three-in-Line problem (Max-N3IL), this mechanism reduces the constraint checking complexity from $O(n^3)$ to $O(n^2)$. To improve search efficiency, we exploit geometric symmetries in two ways: canonical pruning during node expansion to reduce the branching factor, and symmetric batch transitions to accelerate the discovery of promising configurations. We perform extensive experiments and establish new best-known computational results on five out of six of the problems that we considered. Notably, for Max-N3IL we find configurations of size roughly $1.8 n$ for grids of size $82 \le n \le 119$. For the Smallest Complete Set problem, we find configurations of size roughly $0.95 n$, providing new upper bounds within the tested grids. This work establishes Geometry-Aware MCTS as a highly adaptable framework for discovering novel configurations in combinatorial geometry.

13:00 JSTエージェント

エージェントが電気バス車両の運用に対応する場合: アグリゲーター フレームワークにおける価格設定の動作、トレードオフ、およびポリシーへの影響

エージェント システムは、複雑な運用タスクの調整方法を変え、異種データ ソースを接続し、プロセスを自動化するための新しいパラダイムを導入しています。電気バス車両は関連するテストケースを提供します。その運用には、サービスの信頼性、バッテリーの充電状態、充電器の可用性、電力価格、エネルギー経路の不確実性、および車両から送電網への (V2G) 機会の間の継続的な調整が必要です。この論文では、最適化ベースの電気バスのスケジューリング モデルと、障害の検出、料金の適応、およびスケジュールの評価のための監視エージェントを組み合わせることで、この意思決定環境を合理化するエージェント アグリゲーター フレームワークを提案します。最適化コアはルート、充電器、バッテリー、V2G 交換機全体で物理的な実現可能性を強制します。一方、エージェント層は変化する動作条件を解釈し、必要に応じてリアルタイムの再最適化をトリガーし、アグリゲーターと公共交通機関 (PTO) の間で柔軟性の価値を割り当てる方法を定義します。現実的な車両基地のケーススタディでは、サービスの遅延、路線エネルギーの逸脱、電力価格のショック、複合的な外乱を考慮して、利益ベースおよび運用ベースの調整モードの下で、前日およびリアルタイムの運用を評価します。結果は、エージェントアグリゲーションが、実行可能なスケジュールを維持し、選択的に再最適化をアクティブ化し、課金と V2G の柔軟性を向上させることにより、適応的なフリート グリッド調整をサポートできることを示しています。ただし、これらは重要なトレードオフも明らかにしています。利益重視の価格設定を中心に構成されている場合、運用の複雑さを軽減する同じエージェント機能が PTO から価値を引き出す可能性があります。これらの調査結果は、エージェント・アグリゲーターが電気バスの V2G 運用の管理に役立つ可能性があることを示唆していますが、公共車両のコンテキストでの導入には、透明性のある調整モード、監査可能な料金設定、および明示的な価値共有ルールが必要です。

原文 (English)

When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework

Agentic systems are changing how complex operational tasks are coordinated, introducing a new paradigm for connecting heterogeneous data sources and automating processes. Electric bus fleets provide a relevant test case. Their operation requires continuous coordination between service reliability, battery state-of-charge, charger availability, electricity prices, route-energy uncertainty, and vehicle-to-grid (V2G) opportunities. This paper proposes an agentic aggregator framework that streamlines this decision environment by coupling an optimization-based electric bus scheduling model with supervisory agents for disturbance detection, tariff adaptation, and schedule evaluation. The optimization core enforces physical feasibility across routes, chargers, batteries, and V2G exchanges, while the agentic layer interprets changing operating conditions, triggers real-time re-optimization when needed, and defines how flexibility value is allocated between the aggregator and the public transport operator (PTO). A realistic depot case study evaluates day-ahead and real-time operations under profit-based and operation-based coordination modes, considering service delays, route-energy deviations, electricity price shocks, and combined disturbances. The results show that agentic aggregation can support adaptive fleet-grid coordination by maintaining feasible schedules, activating re-optimization selectively, and improving the use of charging and V2G flexibility. However, they also reveal a critical trade-off: the same agentic capability that reduces operational complexity can extract value from the PTO when configured around profit-oriented pricing. These findings suggest that agentic aggregators can become useful for managing electric bus V2G operations, but their deployment in public-fleet contexts requires transparent coordination modes, auditable tariff-setting, and explicit value-sharing rules.

13:00 JSTエージェント

格子理論による不偏正準集合値オラクル

将来の出来事の確率を推定する非エージェント型の「オラクル」AI は、自己参照の問題に直面しています。その答えが学習され、実行されると、報告するよう求められた確率そのものが変わってしまう可能性があります。 Scientist AI プログラムで提唱されている対応の 1 つは、事実に反する質問のみをし、その回答が何の影響もないかのように評価することです。私たちは、そのような答えは学んだ瞬間に意味がなくなってしまう傾向があることを観察しています。それは、まさにその前提が間違っているからです。したがって、私たちは、オラクルが単一の確率ではなく、同時に偏りがなく学習の結果と自己矛盾のない一連の資格を報告する、自己言及的な代替案を模索します。単純な自己一貫性の要件は、あまりにも多くのセット (役に立たない答え $[0,1]$ を含む) によって満たされるため、問題は正規の自明でないメンバーを選び出すことです。これを、適切に定義されたアイソトーン演算子の最小不動点を取り、閉じたクレダル集合の完全な格子に関するクナスター-タルスキーの不動点定理を使用して行います。代わりに、バリアントは、すべての自己矛盾のない点推定を含む最小不動点を報告します。我々は、存在、自己無矛盾性、空でないことを証明し、非実行的質問についてはその構造が古典的な点の答えに崩壊すること、そしてバイナリイベントについては標準的な答えが自然なハル因数分解の仮定の下では区間であることを示します。この展開は純粋に格子理論に基づいており、バイナリ イベント $B$ から任意の確率変数 $X$ までそのまま拡張され、$P(B\mid A,C)$ は条件法 $\mathcal{L}(X\mid A,C)$ に置き換えられます。区間の特徴付け自体がその一般化に耐えられるかどうかを含む未解決の質問で終わります。

原文 (English)

Unbiased Canonical Set-Valued Oracles Via Lattice Theory

A non-agentic "oracle" AI that estimates probabilities of future events faces a self-reference problem: once its answer is learned and acted upon, it can change the very probability it was asked to report. One response, advocated for the Scientist AI programme, is to ask only counterfactual questions, evaluated as if the answer had no influence. We observe that such answers tend to become irrelevant the moment they are learned, precisely because their premise is then false. We therefore explore a self-referential alternative in which the oracle reports not a single probability but a credal set that is simultaneously unbiased and self-consistent with the consequences of being learned. The naive self-consistency requirement is satisfied by too many sets (including the useless answer $[0,1]$), so the problem is to single out a canonical, nontrivial member. We do so with the Knaster--Tarski fixed-point theorem on the complete lattice of closed credal sets, taking the least fixed point of a suitably defined isotone operator; a variant instead reports the least fixed point that contains every self-consistent point estimate. We prove existence, self-consistency, and nonemptiness, show that the construction collapses to the classical point answer for non-performative questions, and that for a binary event the canonical answer is, under a natural hull-factoring assumption, an interval. The development is purely lattice-theoretic and extends unchanged from a binary event $B$ to an arbitrary random variable $X$, with $P(B\mid A,C)$ replaced by the conditional law $\mathcal{L}(X\mid A,C)$. We close with open questions, including whether the interval characterization itself survives that generalization.

13:00 JST研究/論文

大規模な言語モデルと入れ子になったデータへのアプリケーションによる分類器のパフォーマンスの不確実性の推定

研究者は、自然言語からの構成要素を測定するためにテキスト分類 (教師ありモデルまたは大規模言語モデル) を使用することが増えており、その妥当性の証拠として再現率や精度などの指標を提供しています。しかし、これらの指標はサンプリング変動の影響を受ける点推定値であるにもかかわらず、不確実性の尺度がそれらと一緒に報告されるのは一貫性がありません。さらに、それらが報告される場合、関連するラベル付きデータセットが小さい場合やパフォーマンスが高い場合には、適切ではない方法で推定されることがよくあります。現場での信頼区間レポートを増加および改善するために、この論文では、社会科学のテキスト分類に典型的な条件、つまり小規模から中程度のサンプルサイズ、頻度の低い構成、および個人内にネストされたテキストの下で、パフォーマンスメトリクスの信頼区間手法を評価します。シミュレーション全体で、Wald 間隔や基本パーセンタイル ブートストラップなどのデフォルトの方法は精度が最も低く、カバレッジが名目 95% レベルを大幅に下回る場合があります。 Agresti-Coull、Wilson、Clopper-Pearson、および新しい擬似カウント正規化ブートストラップ (特に F1 の計算に関連する) を使用することで、精度が向上します。テキストが個人内でネストされている場合、正確な分析間隔を生成するには実効 N と適切な自由度の両方の調整が必要であることを示します。ブートストラップ間隔の中で、個人が中程度の数のテキストを作成する場合、階層ブートストラップはクラスター ブートストラップよりも正確ですが、個人が少数しか作成しない場合は過度に保守的になります。適切な間隔推定に関するガイダンスを現場に提供することで、機械学習アプリケーションの透明性を向上させ、設計段階での検証サンプル サイズに対するさらなる注意を促すことを目指しています。

原文 (English)

Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data

Researchers increasingly use text classification--supervised models or large language models--to measure constructs from natural language, providing metrics such as recall and precision as evidence of their validity. Yet, though these metrics are point estimates subject to sampling variation, measures of uncertainty are inconsistently reported alongside them. Further, when they are reported, they are often estimated with methods that are not appropriate when relevant labelled datasets are small or performance is high. To increase and improve confidence interval reporting in the field, this paper evaluates confidence interval methods for performance metrics under conditions typical of social science text classification: small to moderate sample sizes, infrequent constructs, and texts nested within individuals. Across simulations, default methods such as the Wald interval and the basic percentile bootstrap are the least accurate, with coverage sometimes far below the nominal 95% level. Accuracy is improved with the use of Agresti-Coull, Wilson, Clopper-Pearson, and a novel pseudo-count regularized bootstrap (which is particularly relevant to the calculation of F1). When texts are nested within individuals, we demonstrate that adjustment for both effective N and the appropriate degrees of freedom is necessary for producing accurate analytic intervals. Among bootstrap intervals, the hierarchical bootstrap is more accurate than the cluster bootstrap when individuals produce a moderate number of texts but overly conservative when individuals produce only a few. By providing guidance to the field on appropriate interval estimation, we aim to improve the transparency of machine learning applications, and to encourage greater attention to the validation sample size at the design stage.

13:00 JST研究/論文GPT / ChatGPT

データ駆動型機械学習は記号レベルの論理的推論に到達できない -- スケーリング則の限界

Sphere ニューラル ネットワークは、トレーニング データなしで記号レベルの三段論的推論を達成しました。これにより、論理的推論のスケーリング則の限界がどこにあるのか、つまり、データ駆動型の機械学習システムがトレーニング データとトレーニング時間を増やすことで同じレベルを達成できるかどうかという問題が生じています。教師あり深層学習が記号レベルの三段論的推論に到達することを妨げる 2 つの方法論的制限を示します。(1) トレーニング データは、24 種類の有効な三段論的推論すべてを区別できない。 (2) 前提から結論までのエンドツーエンドのマッピングでは、パターン認識と論理的推論のための神経コンポーネント間に矛盾するトレーニング ターゲットが導入されます。理論的な分析に加えて、オイラー ネットでは厳密な三段論的推論を達成できないことを実験的に示します。さらに、最新の ChatGPT (GPT-5-nano および GPT-5) に対して、単語、二重単語、単純なシンボル、および長いランダム記号の 4 つの表面形式 (パターン) で三段論法的ステートメントの充足可能性を判定することに挑戦し、表面形式が推論パフォーマンスに影響を与えること、および ChatGPT GPT-5 が 100% の精度に達する可能性があるが、依然として不正確な説明を提供する可能性があることを示します。経験的トレーニングプロセスは 100% の精度に達した後に停止されるため、教師あり機械学習システムは記号論理推論の厳密さを達成できないと結論付けます。

原文 (English)

Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law

Sphere neural networks have achieved symbolic level syllogistic reasoning without training data, raising the question of where the limit of the scaling law for logical reasoning lies, i.e., whether data-driven machine learning systems can achieve the same level by increasing training data and training time. We show two methodological limitations that prevent supervised deep learning from reaching the symbolic-level syllogistic reasoning: (1) training data can not distinguish all 24 types of valid syllogistic reasoning; (2) end-to-end mapping from premises to conclusion introduces contradictory training targets between neural components for pattern recognition and logical reasoning. Beside theoretical analysis, we experimentally illustrate that Euler Net cannot achieve rigorous syllogistic reasoning. We further challenge the most recent ChatGPTs (GPT-5-nano and GPT-5) to determine the satisfiability of syllogistic statements in four surface forms (patterns): words, double words, simple symbols, and long random symbols, showing that surface forms affect the reasoning performance and that ChatGPT GPT-5 may reach 100% accuracy but still provide incorrect explanations. As empirical training processes are stopped after achieving 100% accuracy, we conclude that supervised machine learning systems will not attain the rigour of symbolic logical reasoning.

13:00 JST研究/論文

MKG-RAG-Bench: マルチモーダルナレッジグラフ拡張生成におけるベンチマーク取得

ナレッジ グラフ上の検索拡張生成 (RAG) は、大規模な言語モデルを基礎付けるための有望なアプローチとして浮上していますが、既存のベンチマークでは、マルチモーダル ナレッジ グラフ RAG (MKG-RAG) における検索の課題がほとんど見落とされています。実際には、検索は重大なボトルネックです。マルチモーダルな知識は異質であり、モダリティ間で調整するのが難しく、構造化されていないコーパス向けに設計された検索ツールでは十分に機能しないことがよくあります。このギャップに対処するために、MKG-RAG での取得を評価するために明示的に設計されたクロスドメイン ベンチマークである MKG-RAG-Bench を導入します。 MKG-RAG-Bench は、一般領域と医療領域にわたる 2 つのマルチモーダル ナレッジ グラフから構築されており、検索と下流生成の両方の制御された評価をサポートする慎重に調整された質問応答データセットが含まれています。このベンチマークは、実用性の低いナレッジをフィルタリングし、正確な監視のもとで構造的に根拠のあるクエリを生成し、多様なモダリティ構成を体系的にカバーする LLM ベースのキュレーション パイプラインを使用して構築されています。代表的なレトリーバーファミリーとモダリティ設定にわたる広範な実験を通じて、効果的なマルチモーダル検索は依然として課題であるものの、エンドツーエンドの MKG-RAG パフォーマンスにとって重要であること、および検索品質が生成結果を強く決定することを示します。 MKG-RAG-Bench は、検索を第一級の評価対象として分離することで、現在の制限を診断し、マルチモーダル ナレッジ グラフ RAG システムを進歩させるための原則的な基盤を提供します。

原文 (English)

MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

Retrieval-augmented generation (RAG) over knowledge graphs has emerged as a promising approach for grounding large language models, yet existing benchmarks largely overlook the challenges of retrieval in multimodal knowledge graph RAG (MKG-RAG). In practice, retrieval is a critical bottleneck: multimodal knowledge is heterogeneous, difficult to align across modalities, and often poorly served by retrievers designed for unstructured corpora. To address this gap, we introduce MKG-RAG-Bench, a cross-domain benchmark explicitly designed to evaluate retrieval in MKG-RAG. MKG-RAG-Bench is constructed from two multimodal knowledge graphs spanning general and medical domains, and includes carefully aligned question-answering datasets that support controlled evaluation of both retrieval and downstream generation. The benchmark is built using an LLM-based curation pipeline that filters low-utility knowledge, generates structurally grounded queries with exact supervision, and systematically covers diverse modality configurations. Through extensive experiments across representative retriever families and modality settings, we show that effective multimodal retrieval remains challenging yet crucial for end-to-end MKG-RAG performance, and that retrieval quality strongly determines generation outcomes. By isolating retrieval as a first-class evaluation target, MKG-RAG-Bench provides a principled foundation for diagnosing current limitations and advancing multimodal knowledge graph RAG systems.

13:00 JSTエージェント

auto-psych: エージェント駆動理論の発見と実験を使用した心の科学の自動化

エージェントを使用して仮説を生成し、実験を計画し、データを分析することにより、AI ベースの科学的自動化がますます可能になります。ただし、このパイプラインではデータ収集が大きなボトルネックになっています。心理学、特に計算認知科学は、理論がコードとして表されることが多く、クラウドソーシング プラットフォームによりプログラムによる人間データの大規模収集が可能になるため、AI 実験の恩恵を受ける有利な立場にあります。ここでは、クラウドソーシングによる調査実験を通じて人間のデータを独立して収集するエージェントベースのシステムを使用して、自動発見技術を計算認知科学の理論生成プロジェクトに適用します。テストベッドとして、認知心理学の古典的なケーススタディを使用します。コイン投げのどのシーケンスが主観的によりランダムに見えるかを判断します。私たちのシステム auto-psych は、入れ子になったエージェントベースの発見ループを使用して、人間の行動の説明理論を生成します。内側のループは、確率的認知モデルを推測、適合、および批判します。外側のループは、これらのモデルをテストするための実験を設計し、オンラインで起動して、データを分析します。このシステムは、体系的な実験を通じて合成データからグラウンドトゥルース理論を迅速かつ確実に復元できますが、モデルのパフォーマンスには入れ子構造が重要です。さらに、人体実験の 3 つの独立したシーケンスにおいて、システムは科学文献から生成された理論よりもデータによく適合する理論を見つけます。したがって、この研究は、計算認知科学における自動データ収集と理論発見の実現可能性を実証しています。

原文 (English)

auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation

AI-based scientific automation is increasingly possible by using agents to generate hypotheses, design experiments, and analyze data. Data collection is a major bottleneck in this pipeline, however. Psychology, and computational cognitive science in particular, is well-positioned to benefit from AI experimentation because theories are often represented as code and crowdsourcing platforms enable programmatic human data collection at scale. Here, we apply automated discovery techniques to the project of generating theories in computational cognitive science, with an agent-based system collecting human data independently through crowdsourced survey experiments. As a testbed, we use a classic case study from cognitive psychology: judging which sequences of coin flips seem subjectively more random. Our system, auto-psych, uses nested agent-based discovery loops to generate explanatory theories of human behavior. The inner loop conjectures, fits, and critiques probabilistic cognitive models; the outer loop designs experiments to test these models, launches them online, and analyzes the data. This system can quickly and reliably recover ground-truth theories from synthetic data via systematic experimentation, but the nested structure is critical to model performance. Further, in three independent sequences of human experiments, the system finds theories that fit the data better than theories generated from the scientific literature. This work thus demonstrates the feasibility of automated data collection and theory discovery in computational cognitive science.

13:00 JST研究/論文

統治可能な医療 AI スキル エコシステムのための臨床ハーネス

医療 AI は依然として孤立したモデルを中心に組織化されていますが、臨床ケアには時間を超えて持続する責任ある機能が必要です。私たちは、臨床 AI スキルと Clinical Harness を提案します。これは、AI 対応の臨床機能を登録、調整、保護、監視するためのランタイム ガバナンス アーキテクチャです。骨粗鬆症を例として使用し、知識主導型、データ主導型、および物理学に基づいて強化されたスキルが、ランタイム ガバナンスの下でライフサイクル ケアをどのようにサポートできるかを示します。

原文 (English)

Clinical Harness for Governable Medical AI Skill Ecosystems

Medical AI remains organized around isolated models, whereas clinical care requires accountable capabilities that persist across time. We propose clinical AI skills and the Clinical Harness: a runtime governance architecture for registering, orchestrating, guarding and monitoring AI-enabled clinical capabilities. Using osteoporosis as an exemplar, we show how knowledge-driven, data-driven and physics-enhanced skills can support lifecycle care under runtime governance.

13:00 JSTLLM/生成AI

人間は関与をやめ、推論モデルは存続: 難易度の登録と審議の割り当てを分離する

大規模推論モデル (LRM) は、人間と同じように、より困難な問題に時間がかかります。この表面の類似性は、アイテム内に反対のパターンを隠します。 LRM が問題を間違えると、同じ問題を正解した場合よりも多くのトークンを消費します。人間はその逆を行い、間違った試験に費やす時間を減らします。検討を 2 つのレベルに分けます。1 つは応答時間が項目全体の難易度をどのように追跡するか (登録)、もう 1 つは項目の ID が固定された状態で、エージェントが自身の失敗と成功のどちらに多くの時間を費やすか (割り当て) です。公開されているヒトと LRM の照合コーパスでは、人間と 5 つの思考 LRM はすべて、既知の項目間アライメント (登録) を再現しますが、項目 (割り当て) 内では分岐します。どの LRM も大きな誤対正効果 (H-ARC におけるコーエンの d = 1.47-3.13) を示しますが、人間は反対の符号を示します。比較は各エージェント独自のスケール内に留まります。秒とトークンを 1 つの軸に置くことはありません。解離はアイテムの固定効果の下で保持され、データセット全体で複製され、非思考ベースラインには存在しません。私たちは人間のパターンを、関与対放棄として読みます。人々は、解決できると期待している項目に留まり、残りの項目は放棄します。 LRM パターンを不確実性によって引き起こされる長さとして読み取ります。モデルが不確実な場合、チェーンは成長します。つまり、モデルが失敗する傾向があるのはまさにこの時です。どちらのポリシーも同じ項目間の相関関係を生成するのは困難ですが、以前の研究で使用された尺度に基づいて一致しているように見えます。相違は、アイテムの同一性が固定された場合にのみ現れます。リソース合理的メタ推論では、困難信号は共有するが反対の制御を実装する 2 つの停止ポリシーの間で分割が行われます。トレース長が信号を捕捉し、制御を逃します。

原文 (English)

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong, it spends more tokens than when it gets the same problem right; humans do the reverse, spending less time on the trials they get wrong. We separate two levels of deliberation: how response time tracks difficulty across items (registration), and, with item identity held fixed, whether an agent spends more on its own failures or successes (allocation). On a public matched human-LRM corpus, humans and all five thinking LRMs reproduce the known cross-item alignment (registration) but diverge within items (allocation): every LRM shows a large wrong-vs-right effect (Cohen's d = 1.47-3.13 on H-ARC) while humans show the opposite sign. The comparison stays inside each agent's own scale; we never put seconds and tokens on one axis. The dissociation holds under item fixed effects, replicates across datasets, and is absent in a non-thinking baseline. We read the human pattern as engagement versus abandonment: people stay on items they expect to solve and give up on the rest. We read the LRM pattern as length driven by uncertainty: chains grow when the model is unsure, which is exactly when it tends to fail. Both policies produce the same cross-item correlation with difficulty, so they look aligned on the measure prior work has used; the divergence shows up only once item identity is fixed. Under resource-rational metareasoning, the split is between two stopping policies that share a difficulty signal but implement opposite control; trace length captures the signal and misses the control.

13:00 JSTエージェント研究/論文Claude

NeuraDock Visual Cognitive Load Agent チュートリアル: アルファ ダイナミクスおよびリアルタイム アプリケーション向けの品質ゲート付きオープンソース EEG ワークフロー

このチュートリアル ペーパーでは、アルファ ダイナミクスと視覚的認知負荷分析に焦点を当てたオープンソースの EEG エージェントである NeuraDock Agent の段階的な再現可能なウォークスルーを提供します。目標は実践的です。読者は、エージェントのインストール、EEG 前処理と品質管理の実行、アルファ ダイナミクス図の生成、被験者内での休息/タスクの視覚的認知負荷比較の実行、公開ミニ データセット分析の実行と参照検証の概要との比較、オンライン ダッシュボードの開始、外部アプリケーションからのリアルタイム API の呼び出し、LLM 解釈レイヤーを使用して品質リスクを説明できる必要があります。既存の EEG ツールキットは優れたオフライン分析を提供しますが、リアルタイムの品質ゲート認知負荷パイプラインを構築するには、多くの場合、手動でのブリッジング取得、カスタム QC、アルファ特徴抽出、および Web API が必要になります。このチュートリアルでは、オフラインとオンラインのギャップを埋めます。このチュートリアルでは、品質ゲートのワークフローを使用します。ダウンストリームのアルファとワークロードのメトリクスは、生の EEG から直接計算されるのではなく、前処理と QC ゲートの後でのみ計算されます。含まれているミニデータセット検証では、エージェントは 18 件の記録を処理し、10 件の被験者内比較を生成し、10 件のコントラストのうち 7 件でタスク関連の事後アルファ抑制を観察し、被験者内再現性の初期証拠を推定し、ローカル オンライン API レイテンシーのベンチマークを行いました。このチュートリアルは、EEG ファイルからリアルタイムの視覚的認知負荷プロトタイプへの透過的なパスを必要とする研究者、開発者、応用チームを対象としています。

原文 (English)

NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications

This tutorial paper provides a step-by-step, reproducible walkthrough of NeuraDock Agent, an open-source EEG agent focused on Alpha dynamics and visual cognitive-load analysis. The goal is practical: a reader should be able to install the agent, run EEG preprocessing and quality control, generate Alpha dynamics figures, perform within-subject Rest/Task visual cognitive-load comparison, run the public mini-dataset analyses and compare them with the reference validation summary, start an online dashboard, call the real-time API from an external application, and use the LLM interpretation layer to explain quality risks. Existing EEG toolkits provide excellent offline analysis, but assembling a real-time, quality-gated cognitive-load pipeline often requires manually bridging acquisition, custom QC, Alpha feature extraction, and a web API; this tutorial closes that offline-to-online gap. The tutorial uses a quality-gated workflow: downstream Alpha and workload metrics are computed only after preprocessing and QC gating rather than directly from raw EEG. In the included mini-dataset validation, the agent processed 18 recordings, generated 10 within-subject comparisons, observed task-related posterior Alpha suppression in 7 of 10 contrasts, estimated initial evidence of within-subject repeatability, and benchmarked local online API latency. The tutorial is intended for researchers, developers, and applied teams who want a transparent path from EEG files to real-time visual cognitive-load prototypes.

13:00 JSTLLM/生成AIエージェント

低チャネルEEGエージェントのための境界を意識したコンテキストグラウンディング

大規模言語モデル (LLM) を使用すると、科学ソフトウェアを使いやすくできます。ただし、一般的なモデルでは、特定のセンサーがどの測定をサポートできるか、現在のソフトウェアにどのアルゴリズムが実装されているか、または計算結果によってどの結論が正当化されるかが自動的にはわかりません。これらの区別は、低チャネル脳波検査 (EEG) では特に重要です。EEG では、空間範囲がまばらで信号品質が変動するため、もっともらしいが裏付けのない解釈が容易に生成されます。私たちは、決定論的なローカル EEG エンジンをハードウェア対応言語層から分離するオープンソース アーキテクチャである NeuraDock Agent を紹介します。数値エンジンは録音を解析し、品質管理を実行し、レビューされたスペクトル ワークフローを実行して、機械可読アーティファクトを書き込みます。 LLM は、コンパクトな許可リストに登録された概要とバージョン管理されたコンテキスト パックのみを受け取ります。コンテキストでは、7 チャネルのハードウェア、レビューされたワークフロー、結果フィールド、実装の境界、科学的限界、参照ケースについて説明します。生の EEG と高密度のサンプルごとの配列はローカルのままです システムを 3 つのレベルで評価します。まず、12 回の記録では、10 回の数値繰り返しで同一の構造化された結果が生成され、完全な休憩/タスクの実行では、3 回の繰り返しで同一の結果、レポート、および図のハッシュが生成されました。次に、リクエスト キャプチャと障害挿入の実験により、テスト済みのデータ境界と、HTTP、不正な出力、および接続障害下でのローカル アーティファクトの保存が確認されました。第三に、境界認識ベンチマークは、4 つのコンテキスト アブレーションと 2 つの LLM の下で 36 の通常の質問と敵対的な質問をテストし、288 の出力を生成しました。これらの結果は、EEG エージェントが何を受け入れるか、認定するか、または拒否するかを調整するための実用的なメカニズムとして、ハードウェアおよび実装を意識したグラウンディングをサポートします。それらは臨床的妥当性や検証された絶対的な認知負荷指数を確立するものではありません。

原文 (English)

Boundary-Aware Context Grounding for A Low-Channel EEG Agent

Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatically know which measurements a particular sensor can support, which algorithms are implemented in the current software, or which conclusions are justified by a computed result. These distinctions are especially important for low-channel electroencephalography (EEG), where sparse spatial coverage and variable signal quality make plausible but unsupported interpretations easy to produce. We present NeuraDock Agent, an open-source architecture that separates a deterministic local EEG engine from a hardware-aware language layer. The numerical engine parses recordings, performs quality control, executes reviewed spectral workflows, and writes machine-readable artifacts. The LLM receives only a compact, allowlisted summary and a versioned context pack. The context describes the seven-channel hardware, reviewed workflows, result fields, implementation boundaries, scientific limits, and reference cases. Raw EEG and dense per-sample arrays remain local We evaluate the system at three levels. First, 12 recordings produced identical structured results over ten numerical repetitions, and a complete Rest/Task run produced identical result, report, and figure hashes over three repetitions. Second, request-capture and failure-injection experiments confirmed the tested data boundary and preservation of local artifacts under HTTP, malformed-output, and connection failures. Third, a boundary-awareness benchmark tested 36 ordinary and adversarial questions under four context ablations and two LLMs, yielding 288 outputs.These results support hardware- and implementation-aware grounding as a practical mechanism for calibrating what an EEG agent accepts, qualifies, or refuses; they do not establish clinical validity or a validated absolute cognitive-load index.

13:00 JSTエージェント

革新的な AI 解釈可能性

私たちは、急進的な解釈の哲学的伝統と機械的な解釈可能性のツールを利用して、AI システムをエージェントとして解釈するためのフレームワークを開発します。核心的な問題は、システムに関する計算上の事実が与えられた場合、その信念、欲求、および意味をどのように解決するかということです。これは安全性にとってますます重要です。私たちは、その目的を理解することによって、あるいはもっと控えめに言って、欺瞞を確実に検出することによって、導入したシステムを信頼できるようにしたいと考えています。解釈可能性の研究者は、モデルの内部から信念や欲求を読み取るツールを構築していますが、そのようなツールがいつ成功したかについては、明確な説明はありません。この本はその1つを提供します。私たちは表現主義的アプローチと解釈主義的アプローチの両方に関する基準を提案し、それぞれを現在の解釈可能性手法が実行できるテストに結び付けます。中心的な教訓は、これらの帰属を断片的に作成することはできないということです。信念、欲望、およびそれらが前提とする命題構造は共に制約されており、一方を修正しながら他方を測定する方法は、導入される歪みがすべて継承されます。この全体性は、通訳者の概念を共有していない可能性がある AI システムにとって急務となっています。しかし、それはまた、てこにもなります。つまり、システムの態度はその命題構造を制約し、その構造はどの態度が帰属するかを制約し、機械的解釈可能性は両方を測定するのに役立ちます。

原文 (English)

Radical AI Interpretability

We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational facts about a system, how do we solve for its beliefs, desires, and meanings? This matters increasingly for safety. We want to be able to trust the systems we deploy, whether by understanding their goals or, more modestly, by reliably detecting deception. Interpretability researchers are building tools to read beliefs and desires off a model's internals, but there is no settled account of when such a tool has succeeded. This book supplies one. We propose criteria on both representationalist and interpretationist approaches, and tie each to tests current interpretability methods can carry out. A central lesson is that these attributions cannot be made piecemeal. Beliefs, desires, and the propositional structure they presuppose are jointly constrained, and a method that fixes one while measuring the others inherits whatever distortions that introduces. This holism becomes pressing for AI systems, which may not share the interpreter's concepts. However, it also provides leverage: a system's attitudes constrain its propositional structure, that structure constrains which attitudes can be attributed, and mechanistic interpretability can help us measure both.

13:00 JST研究/論文

PMDformer: 長期予測用のパッチ平均デカップリング情報トランスフォーマー

長期時系列予測 (LTSF) は、エネルギー管理、金融、交通予測などの分野で重要な役割を果たします。トランスフォーマーベースのモデルは、長距離の依存関係を把握するためにパッチベースの戦略を採用していますが、パッチと変数間の形状の類似性を正確にモデル化することは、スケールの違いにより依然として困難です。これに対処するために、パッチ平均デカップリング (PMD) を導入します。これは、各パッチの平均を差し引くことでトレンドと残差の形状情報を分離し、元の構造を保存し、アテンション メカニズムが真の形状の類似性を確実に捕捉するようにします。さらに、長期依存関係をより効果的にモデル化し、変数間関係を捕捉するために、トレンド復元アテンション (TRA) と近接変数アテンション (PVA) を提案します。前者のモジュールは、注意出力を計算しながら、PMD から切り離されたトレンドを再統合します。そして後者は、古い相関関係での過剰適合を避けるために、最も関連性の高い最近の時間セグメントに変数間の注意を集中させます。これらのコンポーネントを組み合わせて、長期予測シナリオで形状の類似性を効果的に捕捉するように設計されたモデルである PMDformer を提案します。広範な実験により、PMDformer は複数の LTSF ベンチマークにわたって安定性と精度において既存の最先端の手法よりも優れていることが示されています。コードは https://github.com/aohu1105/PMDformer で入手できます。

原文 (English)

PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting

Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-based strategies to capture long-range dependencies, but accurately modeling shape similarities across patches and variables remains challenging due to scale differences. To address this, we introduce patch-mean decoupling (PMD), which separates the trend and residual shape information by subtracting the mean of each patch, preserving the original structure and ensuring that the attention mechanism captures true shape similarities. Futhermore, to more effectively model long-range dependencies and capture cross-variable relationships, we propose Trend Restoration Attention (TRA) and Proximal Variable Attention (PVA). The former module reintegrates the decoupled trend from PMD while calculating attention output. And the latter focuses cross-variable attention on the most relevant, recent time segments to avoid overfitting on outdated correlations. Combining these components, we propose PMDformer, a model designed to effectively capture shape similarity in long-term forecasting scenarios. Extensive experiments indicate that PMDformer outperforms existing state-of-the-art methods in stability and accuracy across multiple LTSF benchmarks. The code is available at https://github.com/aohu1105/PMDformer.

13:00 JST研究/論文

C型肝炎患者における肝硬変の存在を検出するための説明可能なアンサンブルベースの機械学習モデル

C型肝炎は、ウイルスによって引き起こされる肝臓感染症であり、肝臓に軽度から重度の炎症を引き起こします。 C型肝炎は長年にわたって徐々に肝臓にダメージを与え、多くの場合、肝硬変として知られる永久的な瘢痕化につながります。患者は、肝硬変を発症するまで数十年間、中等度の肝疾患の症状を呈するか、まったく症状を示さない場合があります。肝硬変は通常、肝不全に至るまで悪化します。肝硬変患者は、胃腸出血だけでなく、脳や神経系の損傷も経験する可能性があります。肝硬変の治療は、病気のさらなる進行を防ぐことに重点を置きます。したがって、肝硬変を早期に検出することは、合併症を回避するために非常に重要です。機械学習 (ML) は、いくつかの病気の診断に使用するための正確かつ正確な情報を提供するのに効果的であることが示されています。それにもかかわらず、これまでのところ、C 型肝炎患者の肝硬変を検出するために ML を使用した研究はありません。この研究では、カリフォルニア大学アーバイン校の ML リポジトリから 2038 人のエジプト人患者の 28 属性で構成されるデータセットを入手しました。 C 型肝炎患者の肝硬変を診断するために、ランダム フォレスト、勾配ブースティング マシン、極端勾配ブースティング、およびエクストラ ツリー モデルの 4 つの ML アルゴリズムがデータセットでトレーニングされました。 Extra Trees モデルは他のモデルを上回り、28 個の特徴のうち 16 個のみを使用して、精度 96.92%、再現率 94.00%、精度 99.81%、受信機動作特性曲線下面積 96% を達成しました。

原文 (English)

Explainable Ensemble-Based Machine Learning Models for Detecting the Presence of Cirrhosis in Hepatitis C Patients

Hepatitis C is a liver infection caused by a virus, which results in mild to severe inflammation of the liver. Over many years, hepatitis C gradually damages the liver, often leading to permanent scarring, known as cirrhosis. Patients sometimes have moderate or no symptoms of liver illness for decades before developing cirrhosis. Cirrhosis typically worsens to the point of liver failure. Patients with cirrhosis may also experience brain and nerve system damage, as well as gastrointestinal hemorrhage. Treatment for cirrhosis focuses on preventing further progression of the disease. Detecting cirrhosis earlier is therefore crucial for avoiding complications. Machine learning (ML) has been shown to be effective at providing precise and accurate information for use in diagnosing several diseases. Despite this, no studies have so far used ML to detect cirrhosis in patients with hepatitis C. This study obtained a dataset consisting of 28 attributes of 2038 Egyptian patients from the ML Repository of the University of California at Irvine. Four ML algorithms were trained on the dataset to diagnose cirrhosis in hepatitis C patients: a Random Forest, a Gradient Boosting Machine, an Extreme Gradient Boosting, and an Extra Trees model. The Extra Trees model outperformed the other models achieving an accuracy of 96.92%, a recall of 94.00%, a precision of 99.81%, and an area under the receiver operating characteristic curve of 96% using only 16 of the 28 features.

13:00 JSTLLM/生成AI

EvoOptiGraph: 最適化モデリングのためのグラフベースの構造生成による弱さ主導の共進化

大規模言語モデル (LLM) を使用した自然言語からの最適化モデリングの自動化は、2 つの重要な課題に直面しています。まず、トレーニング コーパスには構造的な多様性がありません。第 2 に、データ生成パイプラインは静的なままであり、モデル学習から切り離されています。これらの課題に対処するために、モデルの弱点に基づいてデータとモデルが共進化する新しいフレームワークである EvoOptiGraph を提案します。 EvoOptiGraph は、各混合整数線形計画 (MILP) を属性付きの 2 部グラフとして表し、妥当性を保持する進化的演算子を適用して構造的に多様なインスタンスを生成します。進化したグラフは、決定論的コンパイルと検証された逆変換を介してソルバー コードと自然言語に変換されます。トレーニングは 2 段階で進行します。1 つは初期データセットに対する教師あり微調整 (SFT)、続いて検証可能な報酬を伴う強化学習 (RLVR) で、グラフ由来の弱点シグナルがモデルの失敗を対象とした新しいインスタンスの生成をガイドします。これにより、トレーニング分布を継続的に更新する閉ループが形成されます。 6 つの公開データセットに関する実証結果は、EvoOptiGraph が、精度、実行可能性、一般化の点で、大規模なジェネラリスト モデル、エージェント手法、特殊なベースラインよりも大幅に優れていることを示しています。これらの結果は、ターゲットを絞ったデータモデルの共進化が、最適化モデリング タスクで LLM を改善するための効果的な戦略であることを示しています。

原文 (English)

EvoOptiGraph: Weakness-Driven Coevolution via Graph-Based Structural Generation for Optimization Modeling

Automating optimization modeling from natural language with large language models (LLMs) faces two key challenges. First, training corpora lack structural diversity. Second, data generation pipelines remain static and decoupled from model learning. To address these challenges, we propose EvoOptiGraph, a novel framework where data and model co-evolve, driven by model weaknesses. EvoOptiGraph represents each mixed-integer linear program (MILP) as an attributed bipartite graph and applies validity-preserving evolutionary operators to generate structurally diverse instances. The evolved graphs are converted into solver code and natural language via deterministic compilation and verified back-translation. Training proceeds in two stages: supervised fine-tuning (SFT) on an initial dataset, followed by reinforcement learning with verifiable rewards (RLVR), where graph-derived weakness signals guide the generation of new instances targeting the model's failures. This forms a closed loop that continuously updates the training distribution. Empirical results on six public datasets show that EvoOptiGraph significantly outperforms larger generalist models, agentic methods, and specialized baselines in accuracy, executability, and generalization. These results demonstrate that targeted data-model coevolution is an effective strategy for improving LLMs on optimization modeling tasks.

13:00 JST研究/論文

AI 生成の望遠鏡スケジュール決定のためのマルチレベル検証およびトレーサビリティ フレームワーク

望遠鏡のスケジューリングに AI が段階的に導入されることで、複雑な複数の制約問題を処理する際に AI ベースの意思決定が利点を示すようになりました。ただし、その出力には一貫性のないデータ参照、推論エラー、実行不可能な決定が含まれることが多く、信頼性の高い観察タスクへの適用が制限されます。この研究では、実行前に AI が生成した意思決定の体系的な信頼性検証を実行し、追跡可能な意思決定をサポートする推論プロセスの明示的な表現を可能にする、マルチレベル検証および追跡可能な推論フレームワークを提案します。このフレームワークは、データ参照の検証、論理的一貫性チェック、観測的および機器的制約検証を統合して、無効な決定をフィルタリングして修正します。また、アトミック推論ユニットとその依存関係も導入し、エラーの位置特定と事後分析をサポートする相互接続された推論ステップのシーケンスとしてスケジューリング決定を表します。実験では、このフレームワークにより AI スケジューリングの実行可能性と信頼性が向上し、一時的な機会の損失が軽減されることが示されています。特に、フィードバックの修正と推論ステップの構造化された検証により、特に複雑なシナリオで誤った決定を修復およびブロックする能力が強化されます。純粋な AI 手法と比較して、フレームワークで強化されたアプローチは柔軟性を維持しながら、信頼性と実行可能性を大幅に向上させます。これらの結果は、AI を高信頼性の天体観測スケジュールに適用するための実現可能かつ検証可能な経路を示しています。

原文 (English)

A Multi-Level Validation and Traceability Framework for AI-Generated Telescope Scheduling Decisions

With the gradual introduction of AI into telescope scheduling, AI-based decision-making has shown advantages in handling complex multi-constraint problems. However, its outputs often suffer from inconsistent data references, reasoning errors, and non-executable decisions, limiting applicability in high-reliability observational tasks. In this work, we propose a multi-level validation and traceable reasoning framework that performs systematic reliability verification of AI-generated decisions prior to execution, and enables explicit representation of the reasoning process to support traceable decision-making. The framework integrates data reference validation, logical consistency checks, and observational and instrumental constraint verification to filter and correct invalid decisions. It also introduces atomic reasoning units and their dependency relationships, representing scheduling decisions as a sequence of interconnected reasoning steps that support error localization and post hoc analysis. Experiments show that the framework improves executability and reliability of AI scheduling and reduces loss of transient opportunities. In particular, feedback correction and structured validation of reasoning steps enhance the ability to repair and block erroneous decisions, especially in complex scenarios. Compared with pure AI methods, the framework-enhanced approach maintains flexibility while substantially improving reliability and executability. These results demonstrate a feasible and verifiable pathway for applying AI to high-reliability astronomical observation scheduling.

13:00 JST研究/論文

大規模な言語モデルを使用したコンテンツベースのスマート電子メール ディスパッチャー

電子メールによるコミュニケーションは私生活や仕事において不可欠な部分となっていますが、その膨大な量を処理することは依然として大規模な組織にとって重要な問題です。他のインスタント メッセージング プラットフォームを使用して電子メールを手動で閲覧し、その内容と添付ファイルを目的の受信者に転送すると、エラーが発生しやすく時間がかかり、生産性の低下や過度のストレスにつながることが判明しています。このペーパーの主な目的は、工学系大学のプログラムのさまざまな学期の学生のそれぞれの WhatsApp グループにメールの内容に基づいて電子メールを送信するタスクを自動化し、組織内の一方の端からもう一方の端への情報の流れをスムーズにする代替メカニズムを探ることです。ディスパッチャ システムは、大規模言語モデル (LLM) をクエリするエージェントを使用して構築されており、電子メールの内容を分析し、関連する学生グループに電子メールをルーティングして情報を提供し、利用することができます。このシステムは、意思決定のためにテキストの内容を分析する際に LLM の機能を利用します。電子メールのコンテンツを入力として指示とコンテキストとともに含む、適切に構造化されたエージェント フレームワーク プロンプトを使用すると、システムは電子メール メッセージの送信先となる関連グループを特定し、必要な情報を時間どおりに提供します。提案されたシステムは、ラベル付きデータセットに依存せず、生産性の向上や電子メールを読むことに伴う認知負荷の軽減など、いくつかの利点を提供します。

原文 (English)

Content-Based Smart E-Mail Dispatcher Using Large Language Models

Email communication has become an integral part of personal and professional life, but handling its vast volume is still a significant issue for large organisations. Manual perusal of emails and forwarding their contents and attachments to intended recipients using other instant messaging platforms has proved to be error-prone and time-consuming leading to losses in terms of productivity and creating undue stress. The main objective of this paper is to explore an alternative mechanism that is to automate the task of dispatching emails based on their contents to the respective WhatsApp groups of students of various semesters of programs in an engineering college, facilitating a smooth flow of information from one end to another end in an organisation. The dispatcher system is built using agents querying large language models (LLMs) to enable it to analyze the contents of emails and route them to the relevant groups of students for their information and consumption. The system harnesses the capabilities of LLMs in analysing the textual contents for decision-making. With a well-structured agent framework prompt that includes email content as input with instructions and context, the system figures out the relevant groups to which the email message is dispatched, thus providing the required information on time. The proposed system does not rely on labelled datasets and provides several benefits, including enhanced productivity and a reduction in the cognitive load associated with reading emails.

13:00 JSTLLM/生成AI

サービスフィードバックにおける新たなトピックを検出するための LLM ベースのモデル

サービスフィードバックの分析を強化することは、信頼とコンプライアンスが公正かつ効果的なサービスの提供に依存する公共部門の組織、特に税務当局にとって不可欠です。フィードバックの量が増加するにつれて、新たなサービス品質の問題と、多様な集団間の潜在的な格差を特定することがますます困難になっています。従来のアプローチは、手動レビューや専門家が定義した静的な指標に依存することが多く、スケーラビリティやテキストフィードバックで複雑なパターンを捕捉する機能が制限されていました。この論文では、大規模言語モデル (LLM)、統計手法、および人間と AI のコラボレーションを統合して、多言語の顧客フィードバック分析を改善する新しい方法論を紹介します。主な目的は、サービス提供における潜在的な不公平性を明らかにする可能性がある、新たなサービス品質トピックを検出することです。当社のフレームワークは、微調整され量子化された LLM と専門家の監視を組み合わせて、正確で計算効率が高く、コンテキストを認識した分析を生成します。提案されたアプローチは、類似性分析と経験豊富な税務職員による評価を使用して評価され、ベースライン モデルよりも専門家の判断との一致が強いことが実証されました。この方法論では、人間参加型フレームワークを組み込むことで、生成された洞察の信頼性と関連性を向上させながら、LLM の作成を削減します。この結果は、LLM と人間の専門知識を組み合わせて、公共部門の組織における拡張性のある証拠に基づいた意思決定をサポートすることの実用性を示しています。この取り組みは、多言語の顧客フィードバックのより効果的な分析を通じて、サービスの品質、応答性、公平性、社会的信頼を向上させる、責任ある AI システムの開発に貢献します。

原文 (English)

LLM-based Models for Detecting Emerging Topics in Service Feedback

Enhancing the analysis of service feedback is essential for public sector organizations, particularly tax administrations, where trust and compliance depend on fair and effective service delivery. As feedback volumes grow, identifying emerging service quality issues and potential disparities across diverse populations becomes increasingly challenging. Traditional approaches often rely on manual review or static expert-defined indicators, limiting scalability and the ability to capture complex patterns in textual feedback. This paper presents a novel methodology that integrates large language models (LLMs), statistical techniques, and human-AI collaboration to improve multilingual customer feedback analysis. The primary objective is to detect emerging service quality topics that may also reveal potential inequities in service delivery. Our framework combines fine-tuned, quantized LLMs with expert oversight to produce accurate, computationally efficient, and context-aware analyses. The proposed approach was evaluated using similarity analysis and assessments from experienced tax officers, demonstrating stronger alignment with expert judgments than baseline models. By incorporating a human-in-the-loop framework, the methodology reduces LLM fabrication while improving the reliability and relevance of generated insights. The results demonstrate the practicality of combining LLMs with human expertise to support scalable, evidence-based decision-making in public sector organizations. This work contributes to the development of responsible AI systems that enhance service quality, responsiveness, fairness, and public trust through more effective analysis of multilingual customer feedback.

13:00 JSTエージェント

エージェントの指示をコードとしてのポリシーに自動形式化

一か八かの領域におけるエージェントの安全性には、正式なポリシーの適用が必要ですが、既存のアプローチのほとんどは、正式な保証を提供しない確率的なガードレール (微調整された分類器、プロンプトベースのステアリング) に依存しているか、実際のポリシー仕様の幅広さに対応していない手作業でコード化されたシンボリックな適用に依存しています。 LLM ベースのジェネレーター - クリティカル ループを使用して、エージェント プロンプト、MCP ツールの説明、および自然言語ポリシー ドキュメントを正式に検証されたポリシーに変換する自動形式化パイプラインを紹介します。結果として得られるポリシーは Cedar ポリシー言語で記述されます。 MedAgentBench ベンチマークでは、当社の自動形式化されたポリシーは、以前の作業で手作業でコーディングされたシンボリック強制よりも大幅に多くのソース自然言語仕様をカバーします。

原文 (English)

Autoformalization of Agent Instructions into Policy-as-Code

Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement that does not scale to the breadth of real policy specifications. We present an autoformalization pipeline that translates agent prompts, MCP tool descriptions, and natural language policy documents into formally verified policies using an LLM-based generator-critic loop. The resulting policies are written in the Cedar Policy Language. On the MedAgentBench benchmark, our autoformalized policies cover substantially more of the source natural-language specification than the hand-coded symbolic enforcement in prior work.

13:00 JSTエージェント

SKILL-DISCO: エージェント トレースを抽出して再利用可能なプロシージャル スキルにコンパイルする

多くの場合、エージェントは同様のタスク インスタンスを繰り返し最初から解決するため、不必要な推論コストと長い実行トレースが発生します。これまでの研究では、ワークフローの再利用と実行可能なスキルの導入が検討されてきましたが、どのタスク シナリオが手続き型スキルを許可するのか、また、成功したトレース全体で共有される手続き型構造をどのように表現する必要があるのか​​は不明のままです。私たちは、この問題を FSM で定義されたシナリオで研究します。このシナリオでは、成功したトレースが未知の遷移グラフ内のパスとして表示され、再利用可能なパラメーター化された制御フロー サブグラフとして手続き型スキルが定式化されます。この見解に基づいて、成功したトレースから再利用可能な PFSM サブグラフを抽出し、それらを呼び出し可能、実行可能、検証可能な手続き型スキルにコンパイルする抽出およびコンパイル フレームワークである SkillDisCo を紹介します。 ALFWorld と WebArena での実験では、SkillDisCo がベンチマークとモデル スケール全体で成功率を向上させ、エージェントのターン数を削減することが示されており、共有エクスペリエンスを再利用可能な実行構造として表現することの利点が実証されています。

原文 (English)

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explored workflow reuse and executable skill induction, but it remains unclear which task scenarios admit procedural skills and how the shared procedural structure should be represented across successful traces. We study this problem in FSM-defined scenarios, where successful traces can be viewed as paths in an unknown transition graph, and formulate procedural skills as reusable parameterized control-flow subgraphs. Based on this view, we introduce SkillDisCo, a distillation-and-compilation framework that distills reusable PFSM subgraphs from successful traces and compiles them into callable, executable, and verifiable procedural skills. Experiments on ALFWorld and WebArena show that SkillDisCo improves success rates and reduces agent turns across benchmarks and model scales, demonstrating the benefits of representing shared experience as reusable execution structures.

13:00 JST研究/論文

NebulaExp-8B: 本格的なアブレーション研究による経験的なポストトレーニング パイプライン

トレーニング後の調整により、大規模な言語モデルの機能に従う推論と人間の好みが決まりますが、既存の研究のほとんどは詳細なデータ構築、フィルタリング ルール、トレーニング レシピを差し控えており、コミュニティの再現性と軽量モデルの最適化を妨げています。この研究では、Qwen3-8B ベースに構築された完全に透明なアブレーション駆動のポストトレーニング パイプラインである NebulaExp を紹介します。これは、一般的な命令モデルと複雑な推論に特化したモデルという 2 つの直交するモデル ブランチをカバーします。私たちは、384 万のマルチソース SFT サンプルの生のコーパスと 20 万の検証可能な RL 候補プールを厳選し、応答蒸留、多次元相互検証フィルタリング、きめ細かい難易度グレーディング、タスク分類、多様性を意識したサンプリングを含むエンドツーエンドのデータ処理スタックを設計します。 Instruct ブランチの場合、3 段階で最適化された教師あり微調整アプローチ NebulaExp-Ins-SFT により、平均ベンチマーク スコアが Qwen3-8B-nothink のベースライン 55.01 から 60.99 に向上しました。 GRPO 強化学習により、平均スコアはさらに 61.85 まで上昇します。推論ブランチでは、中程度の難易度の GRPO RL により、平均推論スコアが 73.88 から 75.17 に向上しました。 RL のタスク検証者への依存に対処するために、私たちは単一教師と複数教師の OPD (MOPD) を系統的に調査しました。4K の命令に従うサンプルのみを利用し、IFEval で RL ベースラインを 3.26 ポイント上回り、平均全体ゲイン +4.43 でした。 MOPD は、4 人のドメイン専門教師とわずか 10,000 サンプルを融合し、基本モデルと比較して平均パフォーマンスを 4.18 向上させます。このレポートは、8B スケール LLM の完全に再現可能な経験的なトレーニング後のレシピを提供し、命令遵守、数学的推論、コード生成、および一般知識の間の機能のトレードオフを包括的に分析します。

原文 (English)

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization. This work presents NebulaExp, a fully transparent, ablation-driven post-training pipeline built on Qwen3-8B-base, covering two orthogonal model branches: general instruct model and complex reasoning-specialized model. We curate a raw corpus of 3.84M multi-source SFT samples and a 200K verifiable RL candidate pool, and design an end-to-end data processing stack including response distillation, multi-dimensional cross-verification filtering, fine-grained difficulty grading, task classification and diversity-aware sampling. For the Instruct branch, our three-stage optimized supervised fine-tuning approach NebulaExp-Ins-SFT improves the average benchmark score from the 55.01 baseline of Qwen3-8B-nothink to 60.99. GRPO reinforcement learning then further elevates the average score to 61.85. For the Reasoning branch, medium-difficulty GRPO RL improves average reasoning score from 73.88 to 75.17. To address RL's dependency on task verifiers, we systematically investigate single-teacher and multi-teacher OPD (MOPD): utilizing merely 4K instruction-following samples and outperforms RL baseline by 3.26 points on IFEval with +4.43 average overall gain; MOPD fuses four domain-specialist teachers with merely 10K samples, lifting average performance by 4.18 over the base model. This report provides a fully reproducible empirical post-training recipe for 8B-scale LLMs, and comprehensively dissects the capability trade-offs among instruction adherence, mathematical reasoning, code generation and general knowledge.

13:00 JSTLLM/生成AI

安全ガードレールには理由が必要ですか? LeanGuard: 堅牢なモデレーションのための高速かつ軽量なアプローチ

プロンプトまたは応答をスクリーニングするために、最近のガードレール メソッドは、判定を発行する前に思考連鎖 (CoT) を生成します。この設計は、段階的に推論することで意思決定が改善されるという一般的な信念に従っています。ただし、CoT は、モデルが決定する前に多くのトークンを生成する必要があるため、ガードを重く遅くします。これは、ガードレールが実際に展開される方法と一致しない可能性があります。ガードレールは重くて遅いものであってはならず、多くの場合、実体化されたロボットなどのデバイス上で実行されます。この論文では、安全ガードレールに本当に理由を付ける必要があるかどうかという疑問を投げかけます。この質問に答えるために、同じコーパス上で軽量の双方向エンコーダと推論ガードをトレーニングし、他のすべてを固定したまま推論のみを削除します。この制御された同一塩基比較により、チェーンがモデレーション精度を向上させないことがわかります。結果として得られるガードを LeanGuard と名付けます。 395M ラベル専用エンコーダは、公開ベンチマークを上回る 82.90 $\pm$ 0.26 の平均 F1 に達します。これは、はるかに大規模なデコーダ上に構築された推論ガードと一致しますが、最大 512 個のトークンの入力に対して単一の前方パスのみを使用します。これは、推論コンピューティングにおける約 100 分の 1 の削減に相当します。さらに、このラベルのみのエンコーダーはトレーニング ラベル ノイズの下でも堅牢であり、厳密な偽陽性率で推論ガードよりもはるかに多くの再現率を保持するため、より重い推論ガードがより堅牢な選択肢であるわけではないことも示します。私たちの調査結果は、現在のガードレール ベンチマークは推論に報いるほど難しくない可能性があり、モデレーションのための CoT の必要性がまだ証明されていないことを示唆しています。 LeanGuard を含むすべてのソース コードとモデルは https://github.com/ndb796/LeanGuard でリリースされます。

原文 (English)

Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation

In order to screen a prompt or a response, the recent guardrail methods generate a chain-of-thought (CoT) before they issue a verdict. This design follows a common belief that step-by-step reasoning improves a decision. However, CoT also makes the guard heavy and slow, because the model must generate many tokens before it decides. This may not match how guardrails are actually deployed. A guardrail sometimes should not be heavy and slow, and it often runs on-device, for example on an embodied robot. In this paper, we pose a question whether a safety guardrail really needs to reason. To answer this question, we train a lightweight bidirectional encoder and a reasoning guard on the same corpus, and we then remove only the reasoning while we keep everything else fixed. With this controlled same-base comparison, we show that the chain does not improve moderation accuracy. We name the resulting guard LeanGuard. A 395M label-only encoder reaches an average F1 of 82.90 $\pm$ 0.26 over public benchmarks. It matches a reasoning guard that is built on a much larger decoder, while it uses only a single forward pass over an input of at most 512 tokens. This is about a ~100x reduction in inference compute. We further show that this label-only encoder stays robust under training-label noise and retains far more recall at a strict false-positive rate than the reasoning guard, so a heavier reasoning guard is not the more robust choice either. Our finding suggests that the current guardrail benchmarks may not be hard enough to reward reasoning, and that the necessity of CoT for moderation is still not proven. We release all source codes and models including LeanGuard at https://github.com/ndb796/LeanGuard.

13:00 JST研究/論文

コンバインドサイクルガスタービンの少数ショット故障検出のためのカルマンプロトタイプネットワーク

コンバインド サイクル ガス タービン (CCGT) は、現代の発電において重要な役割を果たしており、高効率と環境への影響の低減の両方を実現します。ただし、複雑な熱流体と機械的相互作用により、特にラベル付きの故障データが不足している場合、故障検出が複雑になります。このペーパーでは、特に CCGT 障害診断用に調整されたメトリクスベースの少数ショット学習 (FSL) フレームワークであるカルマン プロトティピカル ネットワーク (KPN) を紹介します。クラスプロトタイプの進化を動的システムの潜在的な確率状態としてモデル化し、一時的な分散を削減し、埋め込み表現のロバスト性を向上させます。オフショア CCGT システムの高忠実度 Modelica ベースの動的シミュレーションで生成された合成データ セットが使用され、通常動作と過渡条件下での進行性漏洩故障の両方をシミュレートしました。提案されたフレームワークをシミュレートされたリーク障害検出タスクに適用すると、KPN は、さまざまなサポートとクエリ構成の下で、精度と安定性の両方において、マッチング ネットワーク、関係ネットワーク、MAML などの従来の FSL 手法よりも優れていることが実証されています。提案されたフレームワークは、クラス表現を安定させることでトレーニングの収束と一般化を大幅に改善し、ラベル付きデータが制限されている現実世界の CCGT 障害検出に適しています。

原文 (English)

Kalman Prototypical Networks for Few-shot Fault Detection in Combined Cycle Gas Turbines

Combined-cycle gas turbines (CCGTs) play a key role in modern power generation, offering both high efficiency and reduced environmental impact. However, their complex thermo-fluid and mechanical interactions complicate fault detection, particularly when labeled fault data are scarce. In this paper, we introduce the Kalman Prototypical Network (KPN), a metric-based few-shot learning (FSL) framework specifically tailored for CCGT fault diagnosis. We model the evolution of class prototypes as latent stochastic states in a dynamic system to reduce episodic variance and improve robustness in embedding representation. Synthetic data sets generated with a high-fidelity Modelica-based dynamic simulation of an offshore CCGT system were used, simulating both normal operation and progressive leak faults under transient conditions. Application of the proposed framework on simulated leak fault detection tasks demonstrate that KPN outperforms conventional FSL methods such as Matching Networks, Relation Networks, and MAML in both accuracy and stability under varying support and query configurations. The proposed framework significantly improves training convergence and generalization by stabilizing class representations, making it well-suited for real-world CCGT fault detection where labeled data is limited.

13:00 JST研究/論文

LithoDreamer: マルチステージ計算リソグラフィーのための物理学に基づいた世界モデル

半導体テクノロジーのノードが拡大するにつれて、歩留まりとパフォーマンスを確保するにはコンピュテーショナル リソグラフィーが不可欠です。ただし、リソグラフィーは、マスクの最適化、光学イメージング、レジスト露光、現像を含む連続的な物理プロセスであり、既存のモデルではこれらを捉えることができません。この制限を克服するために、我々は、「レイアウト-マスク-レジスト画像-現像後画像(ADI)」パイプラインを意思決定駆動型の多段階進化システムとして定式化する、計算リソグラフィーのための最初の物理情報に基づいたワールドモデル(WM)フレームワークであるLithoDreamerを紹介します。 LithoDreamer は、隣接する状態間の特徴の変化をキャプチャして、ステージ固有の物理情報に基づいた潜在空間をモデル化し、そこでプロセス介入の探索を制御し、その後の状態遷移を駆動します。継続的な監視なしで解釈可能な介入最適化を達成するために、介入パス間の潜在的な差異を変分進化制約で対比し、実際のリソグラフィ物理学と一致する進化を生成するようにモデルを導く対照変分最適化パラダイムを提案します。実験では、LithoDreamer が順進化と逆計画において最先端のパフォーマンスを達成することを示しています。私たちのリソグラフィ データセットは、GitHub (https://github.com/7jiangyq/lithodreamer.git) で公開されています。

原文 (English)

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography

As semiconductor technology nodes scale, computational lithography is essential for ensuring yield and performance. However, lithography is a continuous physical process involving mask optimization, optical imaging, resist exposure, and development, which existing models fail to capture. To overcome this limitation, we present LithoDreamer, the first physics-informed World Model (WM) framework for computational lithography, which formulates the ``Layout-Mask-Resist Image-After Development Image (ADI)'' pipeline as a decision-driven multi-step evolution system. LithoDreamer captures feature changes between adjacent states to model stage-specific physics-informed latent spaces, in which it controls process intervention exploration and drives subsequent state transitions. To achieve interpretable intervention optimization without continuous supervision, we propose a contrastive variational optimization paradigm that contrasts the latent differences between intervention paths with variational evolution constraints, guiding the model to generate evolutions consistent with real lithography physics. Experiments show LithoDreamer achieves state-of-the-art performance in forward evolution and inverse planning. Our lithography dataset is publicly available at GitHub (https://github.com/7jiangyq/lithodreamer.git).

13:00 JST画像/動画生成

シネ心臓 MRI の時空間モデリングに対する潜在的な ODE アプローチ

心臓磁気共鳴画像法 (CMR) は、心室の構造と運動に関する豊富な時空間情報を取得しますが、従来のリスク モデルでは、選択された心臓位相から画像から導出された指標が少数しか使用されていません。心拍数を意識した神経常微分方程式(ODE)ダイナミクスとグラフベースのメッシュオートエンコーダーを使用して、解剖学的に一貫した3D + t心室運動を再構築し、両心室の解剖学的構造とフルサイクルのシネ運動を連続的な潜在軌道としてエンコードする潜在力学モデルを提示します。共変量条件付き事前分布は、予想される拡張末期の潜在状態を定義し、コックス比例ハザード モデルは、この事前分布からの逸脱が心不全の発生を予測するかどうかをテストします。私たちは、367件の心不全事象を含む、ベースラインの心血管疾患のない英国バイオバンク参加者72,386人を調査しました。保留された評価サブセットでは、再適合されたプールされたコホート方程式に潜在スコアを追加すると、7 つの確立された心臓マーカーの 0.764 と比較して、層化 C インデックスが 0.704 から 0.785 に改善されました。非グラフおよび非 ODE アプローチと比較して、提案されたモデルは、再構成の忠実度、生成的リアリズム、および下流の予測パフォーマンスの間で最良のトレードオフを示しました。これらの結果は、心室運動の連続的な全周期モデリングが従来のCMR要約を超えた有益な心臓表現型を提供する一方、臨床リスク予測の使用前には、より代表的な患者コホートにおける外部検証が必要であることを示唆している。

原文 (English)

A Latent ODE Approach to Spatiotemporal Modeling of Cine Cardiac MRI

Cardiac magnetic resonance imaging (CMR) captures rich spatiotemporal information about ventricular structure and motion, but conventional risk models use only a few image-derived indices from selected cardiac phases. We present a latent dynamical model that encodes bi-ventricular anatomy and full-cycle cine motion as a continuous latent trajectory, using heart-rate-aware neural ordinary differential equation (ODE) dynamics and a graph-based mesh autoencoder to reconstruct anatomically consistent 3D+t ventricular motion. A covariate-conditioned prior defines the expected end-diastolic latent state, and a Cox proportional hazards model tests whether deviations from this prior predict incident heart failure. We studied 72,386 UK Biobank participants without baseline cardiovascular disease, including 367 incident heart failure events. In a held-out evaluation subset, adding the latent score to refitted pooled cohort equations improved the stratified C-index from 0.704 to 0.785, compared with 0.764 for seven established cardiac markers. Compared with non-graph and non-ODE approaches, the proposed model gave the best trade-off between reconstruction fidelity, generative realism, and downstream prognostic performance. These results suggest that continuous full-cycle modeling of ventricular motion provides informative cardiac phenotypes beyond conventional CMR summaries, while external validation in more representative patient cohorts is required before clinical risk-prediction use.

13:00 JSTエージェント

高次元物理システムにおける自律的な科学発見のためのソクラテスエージェント

科学的発見の自動化は転換点に達しています。 AI システムは現在、機器を操作し、パラメーターを最適化し、仮説を生成しますが、ほとんどは依然として手続き型であり、人間の設計者によって修正されたワークフローを実行します。真の自律科学は認識論的自律性、つまり証拠に応じて物理的説明を構築し、異議を唱え、修正する能力を要求します。ここでは、ソクラテス助産術をクローズドループ実験に組み込むマルチエージェント AI 科学者である AHOIS を紹介します。物理学批判エージェントは、因果関係の質問、制約チェック、反例の生成、および反証基準の定式化を通じて仮説を調査します。私たちは、実際のマルチモードファイバー光学プラットフォーム、複雑な波動変換、間接検出、環境ドリフト、マルチモーダル取得を備えた高次元システム上で AHOIS を評価します。事前の符号化スキーム、分類器、またはスペックル モデルなしで、このシステムは自律的にランダム干渉符号化仮説を提案および検証し、タスク適応型スパース測定戦略を発見し、明確な故障モード (符号化の不安定性、蛍光汚染、検出器ノイズ) を診断し、公開されたイメージング プロトコルをオリジナル以外の構成で実行可能なワークフローに変換しました。発見されたエンコーディングにより、有効ランク 56.9 の 16x16 測定値が得られ、分類精度は MNIST で 76.97%、Fashion-MNIST で 83.17% でした。アブレーションは、ソクラテス的尋問が物理的な一貫性、仮説の完全性、不確実性の校正、および実験計画の妥当性を向上させることを示しています。これらの結果は、ワークフローの自動化から、複雑な物理環境における証拠に基づく自己修正型の自律的な発見への道筋を確立します。

原文 (English)

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and generate hypotheses, most remain procedural: they execute workflows fixed by human designers. True autonomous science demands epistemic autonomy--the capacity to construct, challenge and revise physical explanations in response to evidence. Here we introduce AHOIS, a multi-agent AI scientist that embeds Socratic midwifery into closed-loop experimentation. A physics-critic agent interrogates hypotheses through causal questioning, constraint checking, counterexample generation and falsification-criteria formulation. We evaluate AHOIS on a real multimode-fibre optical platform, a high-dimensional system with complex wave transformations, indirect detection, environmental drift and multi-modal acquisition. Without prior encoding schemes, classifiers or speckle models, the system autonomously proposed and validated a random-interference encoding hypothesis, discovered task-adaptive sparse-measurement strategies, diagnosed distinct failure modes (encoding instability, fluorescence contamination and detector noise) and translated a published imaging protocol into an executable workflow on a non-original configuration. The discovered encoding yielded 16x16 measurements with effective rank 56.9 and classification accuracies of 76.97% on MNIST and 83.17% on Fashion-MNIST. Ablations show that Socratic interrogation improves physical consistency, hypothesis completeness, uncertainty calibration and experimental-plan validity. These results establish a route from workflow automation towards evidence-grounded, self-correcting autonomous discovery in complex physical environments.

13:00 JST研究/論文

メタ最適化としての科学的発見: 組み合わせ最適化のケーススタディ

科学的発見は基本的に最適化問題であり、理論と実験の広大な「状態空間」と、品質、新規性、有効性に基づく評価基準によって定義されます。大規模言語モデル (LLM) により、この領域の自動探索が可能になりましたが、評価基準を同時に変更することも同様に重要であると私たちは主張します。ここでは、研究をメタ最適化として形式化することを提案します。この場合、最適化の目的自体も最適化されます。私たちの主な貢献は「コンセンサス目的集計」です。LLM で生成された目的関数が相関加重投票によって結合され、理解が深まるにつれて進化する安定した自己修正評価基準が生成されます。このフレームワークをデジタル MemComputing マシンに基づく 3-SAT 問題のアルゴリズム検出に適用し、問題サイズ $N$ のベースライン スケーリングを $\sim N^{2.51}$ から $\sim N^{1.33}$ に削減し、テストされた最大のインスタンスで $\sim 67\times$ の高速化を実現します。問題にとらわれないフレームワークとして、このアプローチが科学的発見に大きく役立つことを期待しています。

原文 (English)

Scientific discovery as meta-optimization: a combinatorial optimization case study

Scientific discovery is fundamentally an optimization problem, defined by a vast "state space" of theories and experiments, and an evaluation criterion based on quality, novelty, and validity. Large language models (LLMs) have enabled automated exploration of this space, but we argue that simultaneous modification of the evaluation criteria is equally important. Here, we propose formalizing research as meta-optimization, where the optimization objective itself is also being optimized. Our key contribution is "consensus objective aggregation," where LLM-generated objective functions are combined via correlation-weighted voting, yielding a stable, self-correcting evaluation criterion that evolves as understanding deepens. We apply this framework to algorithm discovery for 3-SAT problems based on digital MemComputing machines, reducing the baseline scaling with problem size $N$ from $\sim N^{2.51}$ to $\sim N^{1.33}$ and delivering a $\sim 67\times$ speedup on the largest instances tested. As a problem-agnostic framework, we hope this approach will considerably aid scientific discovery.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体Claude

EGG: 専門家によるカーネル生成のためのエージェント フレームワーク

高性能 GPU カーネルは、大規模言語モデル (LLM) の指数関数的に増大する計算コストを削減するために不可欠ですが、その開発はドメイン専門家による手動チューニングに大きく依存しています。 LLM ベースのアプローチの最近の進歩は、カーネル生成の自動化に有望であることを示していますが、正確さと高いパフォーマンスの両方を達成するのにまだ苦労しています。この制限は主に、ドメイン固有の最適化ガイダンスの欠如によって生じ、最適化空間の効果的な探索が妨げられます。私たちは、LLM の意思決定をガイドする専門家の最適化原則を組み込んだ、カーネル生成のための専門家ガイド付きエージェント フレームワークである EGG を提案します。専門家のワークフローからインスピレーションを得て、私たちはカーネル生成を 2 つの階層段階に分解します。1) 高品質の計算構造基盤を確立するアルゴリズム構造設計。 2) ハードウェア固有のチューニング。並列マッピング、テンソル タイリング、メモリ最適化を通じてターゲットを絞った調整を実行します。この段階的な分解により、明示的な最適化目標が定義され、段階的な改良を達成するために設計空間が構築されます。この目的を達成するために、ステージを意識したマルチエージェント コラボレーション メカニズムがステージ間およびステージ内のコンテキスト管理用に設計されており、安定した最適化軌道を保証します。 KernelBench と実際のワークロードの実験では、EGG が PyTorch と比較して平均 2.13 倍の高速化を達成し、既存のエージェント ベースおよび RL ベースのアプローチを上回るパフォーマンスを示していることが示されています。

原文 (English)

EGG: An Expert-Guided Agent Framework for Kernel Generation

High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in LLM-based approaches show promise for automating kernel generation, they still struggle to achieve both correctness and high performance. This limitation primarily arises from the lack of domain-specific optimization guidance, hindering effective exploration of the optimization space. We propose EGG, an Expert-Guided Agent Framework for Kernel Generation, which incorporates expert optimization principles to guide LLMs' decisions. Inspired by expert workflows, we decompose kernel generation into two hierarchical stages: 1) algorithmic structure design, which establishes a high-quality computational structure foundation; 2) hardware-specific tuning, which performs targeted adjustments through parallel mapping, tensor tiling, and memory optimization. This staged decomposition defines explicit optimization objectives, structuring the design space to achieve progressive refinement. To this end, a stage-aware multi-agent collaboration mechanism is designed for inter and intra-stage context management, ensuring stable optimization trajectories. Experiments on KernelBench and real-world workloads show that EGG achieves a 2.13x average speedup over PyTorch, outperforming existing agent-based and RL-based approaches.

13:00 JST画像/動画生成

ResilPhase: 拡散加速のためのプラグアンドプレイ位相マッピングとノイズ耐性のあるマクロ軌道外挿

強力な拡散モデルの採用は、その大幅な推論遅延によって妨げられています。最近の「キャッシュしてから予測」スキームは、導関数ベースの多項式を使用して DiT を高速化することでこの問題を軽減しますが、高い加速率では重大な品質劣化が発生します。私たちの分析により、その根本原因が明らかになりました。それは、連続拡散軌道とずれていて数値的に不安定な表現に対して実行された離散外挿です。したがって、加速された DiT は、蓄積された空間エラー、ノイズの多い微分増幅、および高次の不安定性の影響を受けます。したがって、加速推論を常微分方程式 (ODE) 空間における安定したマクロ軌道外挿として再定式化します。中間の特徴を予測する代わりに、モデルのグローバル ドリフト (GD)、つまりエンドツーエンドの状態の進化に合わせて予測を行うことで、特徴の不一致とメモリのオーバーヘッドを排除します。しかし、この滑らかなマクロ軌道でさえも、微分の誤謬に対して脆弱なままです。つまり、その高次の時間微分は本質的にノイズが多いのです。したがって、導関数のない重心ラグランジュ外挿法を導入して、導関数の不安定性と近似誤差を効果的に回避します。さらに、外挿領域を正規化し、振動誤差の増大を抑制する、有界位相マッピングを提案します。これらの要素は集合的に、ノイズ耐性のあるアクセラレーション フレームワークである ResilPhase を構成します。 FLUX.1-dev と HunyuanVideo の実験では、積極的な加速比の下で最先端の忠実度が実証されています。

原文 (English)

ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration

The adoption of powerful diffusion models is hindered by their significant inference latency. Recent ``cache-then-forecast'' schemes alleviate this issue by accelerating DiTs using derivative-based polynomials, but they suffer from severe quality degradation at high acceleration ratios. Our analysis reveals its root cause: the discrete extrapolation performed on representations that are misaligned with the continuous diffusion trajectory and are numerically unstable. Thus, accelerated DiTs suffer from accumulated spatial errors, noisy derivative amplification, and high-order instability. We therefore reformulate accelerated inference as stable macro-trajectory extrapolation in ordinary differential equation (ODE) space. Instead of predicting intermediate features, we align forecasting with the model's Global Drift (GD), i.e., the end-to-end state evolution, thereby eliminating feature inconsistency and memory overhead. However, even this smooth macro-trajectory remains vulnerable to the derivative fallacy: its higher-order temporal derivatives are intrinsically noisy. Thus, we introduce a derivative-free barycentric Lagrange extrapolator to effectively bypass derivative instability and approximation error. We further propose a bounded Phase Mapping that regularizes the extrapolation domain, suppressing oscillatory error growth. These elements collectively constitute ResilPhase, a noise-resilient acceleration framework. Experiments on FLUX.1-dev and HunyuanVideo demonstrate state-of-the-art fidelity under aggressive acceleration ratios.

13:00 JSTエージェントGPT / ChatGPTMistral AI

メモリアクセスではなくメモリ深さ: 長時間実行される言語エージェントのための選択的なパラメトリック統合

長時間実行される言語エージェントには、メモリ アクセス以上のものが必要です。検索システムはクエリ時に過去のファクトをフェッチできますが、作業コンテキストがアンロードされた後、どのエクスペリエンスが引き続き動作を形成するかを決定しません。私たちは、この別の問題をメモリの深さとして研究します。つまり、小さなパラメトリック ストアに書き込まれる耐久性のある目標条件付きの傾向です。ループドリフトプロトコルを導入します。これは、作業コンテキストがアンロードされている間、検索インデックスがそのまま残り、長いループ干渉下でも目標条件付き動作が持続する必要がある制御されたストレステストです。我々は、サプライズおよび価数ゲート型 LoRA 統合メカニズムである EVAF を評価します。 GPT-2 と TinyLlama 全体で、検索は浅い事実の想起 (短い事実の精度 0.956 ~ 0.973) で最も強力ですが、EVAF は目標の永続性とアンロード後の回復 (0.812 ~ 0.904) で最も強く、200 イベントあたりわずか 2 ~ 3 回のパラメトリック書き込みです。メカニズム制御は、選択的統合が選択と作動という 2 つの制御可能な次元に分解されることを示しています。一致したランダム ゲートは、スパース書き込みを超えて選択を分離します。 GPT-2、TinyLlama、Mistral-7B にわたる固定内部制御は、内部ループの書き込み強度がモデルに依存していることを示しています。そして、Mistral-7B のマッチドゲート反転により、誤って調整された作動下での非対称な選択と作動の結合が明らかになります。 Public Memora イベント ストリームは外部診断として機能し、古いメモリの無効化を未解決の境界として明らかにします。このプローブ内では、選択的なパラメトリック統合により、検索アクセスとは異なる、検索アクセスを補完するメモリ深さが提供されます。

原文 (English)

Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents

Long-running language agents need more than memory access. Retrieval systems can fetch past facts at query time, but they do not decide which experiences should continue to shape behavior after the working context is unloaded. We study this separate problem as memory depth: durable goal-conditioned tendencies written into a small parametric store. We introduce the loop-drift protocol, a controlled stress test in which the retrieval index remains intact while working context is unloaded and goal-conditioned behavior must persist under long-loop interference. We evaluate EVAF, a surprise- and valence-gated LoRA consolidation mechanism. Across GPT-2 and TinyLlama, retrieval is strongest on shallow factual recall (short-fact accuracy 0.956--0.973), while EVAF is strongest on goal persistence and post-unload recovery (0.812--0.904) with only 2--3 parametric writes per 200 events. Mechanism controls show that selective consolidation factorizes into two controllable dimensions: selection and actuation. Matched random gates isolate selection beyond sparse writing; fixed-inner controls across GPT-2, TinyLlama, and Mistral-7B show that inner-loop write strength is model-dependent; and a Mistral-7B matched-gate inversion reveals asymmetric selection-actuation coupling under miscalibrated actuation. Public Memora event streams serve as an external diagnostic, exposing stale-memory invalidation as an unresolved boundary. Within this probe, selective parametric consolidation supplies memory depth distinct from and complementary to retrieval access.

13:00 JSTLLM/生成AI

KARLA: 言語モデルの知識ベース拡張検索

私たちは、LLM がトークン生成中に知識ベースから事実の知識を自動的に取り込むことを可能にする新しい方法を提案します。これは、(1) LLM 出力内の事実の知識は、LLM を再トレーニングすることなく更新できること、(2) LLM 出力内の事実を知識ベースまで追跡して透明性と説明可能性を実現できること、(3) より小さなモデルでもより大きなモデルと同じ事実の精度を達成できることを意味します。私たちの中心的なアイデアは、ナレッジ ベースへのクエリをトリガーする特別なトークンを生成するようにモデルをトレーニングすることです。私たちの実験は、私たちの方法が短い形式と長い形式の両方の生成における事実の根拠を改善し、パラメータの更新ではなくKBの編集を通じて事実の改訂を有効にすることを可能にすることを示しています。

原文 (English)

KARLA: Knowledge-base Augmented Retrieval for Language Models

We propose a new method that allows an LLM to automatically pull in factual knowledge from a knowledge base during token generation. This means that (1)~factual knowledge in the LLM output can be updated without retraining the LLM, (2)~facts in the LLM output can be traced to the knowledge base for transparency and explainability, and (3)~smaller models can achieve the same factual accuracy as larger models. Our core idea is to train the model to produce special tokens that trigger a query to the knowledge base. Our experiments show that our method improves factual grounding in both short and long-form generation, and allows factual revisions to take effect through KB edits rather than parameter updates.

13:00 JST研究/論文

健康な成人における心拍数変動のコンピューター解析

心拍数変動 (HRV) 分析は心臓の生理学的状態の重要な指標であり、病気の診断に役立ちます。しかし、健康な人の HRV パラメータに関する研究は依然として限られており、ゴールドスタンダードは存在しません。この研究では、HRV の臨床的有用性を向上させるために、40 人の健康な成人 (男性 20 人、女性 20 人、30 ~ 50 歳) の HRV 指数を評価します。信号処理とデータ分析のための計算手法を使用して、時間、周波数、および非線形インデックスが分析され、(1) 正規性、(2) 安定性、(3) 相関性、(4) 再現性、および (5) 一貫性の 5 つの質問に対処しました。主な発見: (1) 時間領域および非線形指数、特にグローバルおよび LF (低頻度) は正規分布に従い、性差が認められます。 (2) HF (高周波) 関連のものを除いて、ほとんどの指数は安定しています。 (3) HF 関連指標の高い相関は冗長性を示唆しており、研究では 1 つだけが必要であることを示しています。 (4) Fantasia データベースとの比較では、女性の SD2 と SDNN (15% 以上) を除き、ほとんどの指数の誤差が 10% 未満であることが明らかになりました。 (5) 時間領域インデックスと非線形インデックスは研究間の変動が小さいのに対し、周波数領域インデックスは高い変動を示し、研究間の比較が制限されます。選択された指標、ApEn および IRRR (グローバル変動)、HRVi および SD2 (LF)、および MADRR または rMSSD (HF) は、HRV コンポーネントを正確に表し、その臨床および研究の関連性を高めるのに最適です。

原文 (English)

Computational Analysis of Heart Rate Variability in Healthy Adults

Heart Rate Variability (HRV) analysis is a key indicator of cardiac physiological state and aids in disease diagnosis. However, research on HRV parameters in healthy individuals remains limited, and no gold standard exists. This study evaluates HRV indices in 40 healthy adults (20 men, 20 women, aged 30-50) to improve HRV's clinical utility. Using computational methods for signal processing and data analysis, time, frequency, and nonlinear indices were analyzed to address five questions: (1) normality, (2) stability, (3) correlation, (4) reproducibility, and (5) consistency. Key findings: (1) Time-domain and nonlinear indices, particularly global and LF (low frequency), follow normal distributions, with gender differences noted. (2) Most indices are stable except HF (high frequency)-related ones. (3) High correlations in HF-related indices suggest redundancy, indicating only one is necessary in studies. (4) Comparisons with the Fantasia database revealed less than 10% error for most indices, except SD2 and SDNN in women (greater than 15%). (5) Time-domain and nonlinear indices show low inter-study variability, while frequency-domain indices exhibit high variability, limiting cross-study comparisons. The selected indices-ApEn and IRRR (global variability), HRVi and SD2 (LF), and MADRR or rMSSD (HF)-are best suited for accurately representing HRV components and enhancing its clinical and research relevance.

13:00 JSTLLM/生成AI研究/論文

機能のフロンティア: ベンチマークはモデルのパフォーマンスの 82% を逃しています

既存のベンチマークは通常、1 回の実行で 1 つのモデルの精度を報告します。これは、特に異種データ分布の下で、現実世界の LLM 機能を系統的に過小評価しています。(i) 異なるモデルは、その専門分野に応じて異なる問題を正解し、(ii) 予算が与えられれば、複数の世代をサンプリングして選択的に保持できます。このギャップを定量化するために、機能フロンティアを導入します。これは、モデルおよび世代全体にわたる最適な選択(つまり、オラクルによる)の下で、各コスト レベルで達成可能な最高のパフォーマンスを特徴付ける一連のモデルにわたるパレート フロンティアです。私たちの構築では、単一モデルの評価による過小評価と、ノイズの多いサンプルに対して最大値を取ることによる過大評価という 2 つの相反するバイアスが補正されます。私たちは、コーディング、推論、医学、事実性、指示に従って、エージェントタスクにわたる 16 の広く使用されているベンチマークにわたる 21 の LLM を調査し、同等のコストでの Capability Frontier のパフォーマンスを各ベンチマークの最高パフォーマンスのモデルと比較しました。単一モデルの評価を修正すると、エラー率が 54% 減少します。単一実行をさらに補正すると、82% の改善が得られ、85% のコスト削減と同等の SOTA 精度が得られます。これらの経験的結果を補完するために、制御された確率的シミュレーションを使用して、クエリ トピックのエントロピーが高くなると、Oracle ルーティングと最良の単一モデルの間のパフォーマンス ギャップがほぼ単調増加することを示します。私たちの調査結果は、集合的な LLM 機能が大幅に過小評価されており、異種データのマルチドメイン設定での評価と展開に影響を与えることを示唆しています。

原文 (English)

The Capability Frontier: Benchmarks Miss 82% of Model Performance

Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data distributions: (i) different models get different questions correct according to their specializations, and (ii) given a budget, multiple generations can be sampled and selectively retained. To quantify this gap, we introduce the Capability Frontier: a Pareto frontier over a set of models that characterizes the best achievable performance at each cost level under optimal selection across models and generations (i.e., via an oracle). Our construction corrects for two opposing biases: underestimation from single-model evaluation and overestimation from taking maxima over noisy samples. We study 21 LLMs across 16 widely used benchmarks spanning coding, reasoning, medicine, factuality, instruction following, and agentic tasks, comparing Capability Frontier performance at matched cost to each benchmark's top-performing model. Correcting for single-model evaluation yields a 54% error rate reduction; additionally correcting for single runs yields an 82% improvement, with SOTA accuracy matched at 85% cost reduction. Complementing these empirical results, we use controlled probabilistic simulations to show that higher query topic entropy produces a near-monotonic increase in the performance gap between oracle routing and the best single model. Our findings suggest collective LLM capabilities are substantially underestimated, with implications for evaluation and deployment in data-heterogeneous, multi-domain settings.

13:00 JSTLLM/生成AI

倉庫最適化のための最適化パイプラインのコンテキスト認識型合成

手動のピッカーから商品までの倉庫での注文の履行には、品目の割り当て、注文のバッチ処理、ピッカーのルーティングなど、相互に関連した決定が含まれます。統合モデルはこれらの意思決定間の相互作用をキャプチャしますが、実際の倉庫システムでは、組織の境界、責任の違い、またはデータの可用性の制限により、分解されたアプローチが必要になることがよくあります。既存の研究では主に、特定のウェアハウス設定における孤立した部分問題または固定された部分問題の組み合わせのアルゴリズムを評価していますが、適用可能なアルゴリズム構成を決定し、それらを有効なソリューション パイプラインに構成して、そのパフォーマンスを評価するための一般的なメカニズムが不足しています。 Context-Aware Synthesis of Optimization Pipelines (CASOP) を使用して、コンテキスト固有の最適化パイプラインを構築および評価するためのフレームワークを提案し、これらを注文フルフィルメントに適用します。このフレームワークは以下で構成されます。(1) 一般的な注文履行の問題に対するアルゴリズムのモジュール式リポジトリ。 (2) ウェアハウスのコンテキストとアルゴリズム要件を説明するためのセマンティック データとアルゴリズム カード。 (3) 注文履行の問題を関連する下位問題に構造化する分類法。 (4) 特定のウェアハウス コンテキストに適用可能なアルゴリズムを特定し、すべての有効な最適化パイプラインを構成するパイプライン シンセサイザー。 (5) 結果として得られるすべてのパイプラインを評価するパイプライン エバリュエーター。 4 つの問題クラスをカバーする 7 つのベンチマーク インスタンス セットでフレームワークを実証し、結果として 1,063,044 の有効なパイプラインが得られます。このフレームワークは、倉庫業務用の有効で高性能なアルゴリズム パイプラインの設計、自動合成、選択において研究者や実務者をサポートします。このソフトウェアはオープンソースであり、https://github.com/kit-dsm/ware_ops_pipes および https://github.com/kit-dsm/ware_ops_algos から入手できます。キーワード: 倉庫の最適化、アルゴリズムの選択、パイプライン合成、注文処理

原文 (English)

Context-Aware Synthesis of Optimization Pipelines for Warehouse Optimization

Order fulfillment in manual picker-to-goods warehouses involves interconnected decisions such as item assignment, order batching, and picker routing. While integrated models capture interactions between these decisions, practical warehouse systems often require decomposed approaches due to organizational boundaries, differing responsibilities, or limited data availability. Existing studies primarily evaluate algorithms for isolated subproblems or fixed subproblem combinations for specific warehouse settings, but lack a general mechanism to determine applicable algorithm configurations, compose them into valid solution pipelines, and assess their performance. With Context-Aware Synthesis of Optimization Pipelines (CASOP), we propose a framework for constructing and evaluating context-specific optimization pipelines and apply these to order fulfillment. The framework comprises: (1) a modular repository of algorithms for common order fulfillment problems; (2) semantic data and algorithm cards to describe warehouse context and algorithm requirements; (3) a taxonomy that structures order fulfillment problems into relevant subproblems; (4) a pipeline synthesizer that identifies applicable algorithms for a given warehouse context and composes all valid optimization pipelines; and (5) a pipeline evaluator that assesses all resulting pipelines. We demonstrate the framework on 7 benchmark instance sets covering four problem classes, resulting in 1,063,044 valid pipelines. The framework supports researchers and practitioners in designing, automatically synthesizing, and selecting valid, high-performing algorithmic pipelines for warehouse operations. The software is open-source and available at https://github.com/kit-dsm/ware_ops_pipes and https://github.com/kit-dsm/ware_ops_algos. Keywords: Warehouse optimization, Algorithm selection, Pipeline synthesis, Order fulfillment

13:00 JST研究/論文GPT / ChatGPT

LCAi: ビッグデータの融合と検索拡張生成支援解釈によるライフサイクル評価

ライフサイクル評価の解釈段階では、技術的、社会的、政策的不確実性の下で、環境ホットスポットに対処する定量化された改善の機会を実行可能な戦略的経路に変換するための構造化されたメカニズムが欠けていることがよくあります。この制限を克服するために、この研究では、LCA 解釈のためのパースペクティブ条件付き検索拡張生成フレームワークを導入します。このフレームワークでは、マルチパースペクティブ検索と制御合成が人工知能 (AI) 支援 LCA に組み込まれています。 LCA 解釈における大規模な言語モデルを運用するために、学術、業界、公的議論、および欧州連合 (EU) の資金提供データセットをカバーするパースペクティブ フュージョン RAG アーキテクチャが開発されました。私たちのアプローチは 3 つのステップで構成されます: (1) システム境界と脱炭素化目標を定義するシナリオ アンカー、(2) 制約付き検索を伴うパースペクティブ固有の一連のマイクロクエリ、(3) それ以上の検索は行わずに台帳に保存された出力のみを統合する中立的な合成ステップ。このフレームワークは、推論モデルとして GPT-5 nano を使用したイタリアのリンゴ生産施設における水素によるディーゼル削減のユースケースを通じて実証されています。全体として、構造化検索と制約付き合成は、クロスドメインの多様性を維持しながら幻覚のリスクを軽減するように設計されています。提示されたアプローチは、影響結果をより規律正しく戦略的経路に変換することをサポートし、LCA 研究、特に大規模に導入できるテクノロジーに焦点を当てた高度な AI ツールを使用するための新しい道を開きます。この概念実証は、AI 支援の証拠に基づく解釈が、従来の LCA 研究を超えて実装指向の意思決定をどのようにサポートできるかを示しています。

原文 (English)

LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation

The interpretation phase of life cycle assessment often lacks structured mechanisms for translating quantified improvement opportunities addressing environmental hotspots into actionable strategic pathways under technological, social, and policy uncertainty. To overcome this limitation, this study introduces a perspective-conditioned retrieval-augmented generation framework for LCA interpretation, where a multi-perspective retrieval and controlled synthesis is incorporated in the artificial intelligence (AI)-assisted LCA. To operationalise large language models in LCA interpretation, a perspective fusion RAG architecture was developed, covering academic, industry, public discourse, and European union (EU) funding datasets. Our approach comprises three steps: (1) a scenario anchor defining system boundaries and decarbonization targets, (2) a set of perspective-specific micro-queries with constrained retrieval, and (3) a neutral synthesis step integrating only ledger-stored outputs without further retrieval. The framework is demonstrated through a hydrogen-enabled diesel reduction use case in an Italian apple production facility using GPT-5 nano as the reasoning model. Overall, the structured retrieval and constrained synthesis are designed to mitigate the risk of hallucination while preserving cross-domain diversity. The approach presented can support more disciplined translation of impact results into strategic pathways and opens up new avenues for the use of advanced AI tools in LCA studies, particularly those focused on technologies that could be deployed at scale. This proof-of-concept demonstrates how AI-assisted, evidence-grounded interpretation can support implementation-oriented decision-making beyond conventional LCA studies.

13:00 JSTLLM/生成AIエージェント研究/論文

AgentX: 産業用レコメンダー システムのエージェント駆動型自己反復に向けて

レコメンデーション アルゴリズムの反復は、職人的なエンジニアに拘束されたプロセスから工業化された研究ループに移行していますが、この移行は構造的な実行ボトルネックによって妨げられたままです。アイデアから発売までのサイクルは依然として人間のエンジニアに依存して、仮説の生成、製品コードの変更、A/B 実験の開始、およびオンライン結果の属性に依存しています。したがって、イノベーションは、証拠、計算、蓄積された実験知識と複合するのではなく、従業員数に比例してスケールします。この本番機能を根本的に再構築する、本番環境に展開されるマルチエージェント システムである AgentX を紹介します。 AgentX は、自己進化する開発エンジンとして動作します。手動ワークフローでは維持できない規模とペースで、推奨実験を自律的に生成、実装、評価し、そこから学習します。このシステムは、密接に結合された 4 つのステージを閉ループで調整します。 Brainstorm Agent は、過去の実験、システム アーキテクチャ、データ分析、外部調査からの証拠を総合して、ランク付けされた実行可能な提案を作成します。開発エージェントは、リポジトリに基づいた生成と多次元の信頼性検証を通じて、各提案を本番環境に対応したコードに変換します。評価エージェントは、ガードレール拒否権のある A/B 判定を使用して安全なオンライン ロールアウトを実行し、成功と失敗の両方を構造化された知識資産に変換します。その後、ハーネス エボリューション レイヤー (SGPO) が実行軌跡を意味論的勾配更新に抽出し、エージェント自体を継続的に強化することで、システムを単に自動化するだけでなく、自己改善するシステムにします。

原文 (English)

AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly with headcount rather than compounding with evidence, compute, and accumulated experimental knowledge. We present AgentX, a production-deployed multi-agent system that fundamentally restructures this production function. AgentX operates as a self-evolving development engine: it autonomously generates, implements, evaluates, and learns from recommendation experiments at a scale and pace that no manual workflow can sustain. The system orchestrates four tightly coupled stages in a closed loop. A Brainstorm Agent synthesizes evidence from historical experiments, system architecture, data analysis, and external research into ranked, executable proposals. A Developing Agent translates each proposal into production-ready code through repository-grounded generation and multi-dimensional reliability verification. An Evaluation Agent conducts safe online rollout with guardrail-vetoed A/B judgment, converting both successes and failures into structured knowledge assets. A Harness Evolution layer (SGPO) then distills execution trajectories into semantic-gradient updates that continuously sharpen the agents themselves -- making the system not merely automated, but self-improving.

13:00 JST研究/論文

TAVR-VLM: 幻覚耐性レポート生成のためのリスク条件付き因果グラウンディング

経カテーテル的大動脈弁置換術 (TAVR) の計画には、細心の注意を払った複合的な推論が必要です。ただし、マルチモーダル大規模言語モデル (MLLM) をこの一か八かの領域に適応させることは、生成されたテキストに解剖学的根拠が欠けている幻覚診断によって大きく妨げられます。これに対処するために、モデル内部の「リスク $\rightarrow$ 領域 $\rightarrow$ Word」構造的グラウンディング経路をインスタンス化する、リスク条件付き因果グラウンディング アテンション (R-CGA) を特徴とする新しいフレームワークである TAVR-VLM が導入されました。 R-CGA は、マルチモーダルな入力を因果的リスクのボトルネックに圧縮し、密集した視覚的特徴をグローバル リスク マスクに精製します。自己回帰生成中、サポート投影の因果的一貫性目標により、リスク定義のサポート マスク内のトークン レベルのグラウンディングが制限されます。 TAVR-VLM は、包括的な 1,482 人の患者コホートである $\text{M}^3\text{TAVR}$ で評価され、新しい最先端技術を確立します。 AUROC 0.896 を達成し、CIDEr を 0.936 に高め、幻覚率を 8.1\% に大幅に低減することで、証拠に基づいた外科用 AI の解釈可能性を向上させます。

原文 (English)

TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

Transcatheter Aortic Valve Replacement (TAVR) planning requires meticulous multimodal reasoning. However, adapting Multimodal Large Language Models (MLLMs) to this high-stakes domain is severely impeded by diagnostic hallucinations, where generated text lacks anatomical grounding. To address this, TAVR-VLM is introduced: a novel framework featuring Risk-Conditioned Causal Grounding Attention (R-CGA) that instantiates a model-internal ``Risk $\rightarrow$ Region $\rightarrow$ Word'' structural grounding pathway. R-CGA compresses multimodal inputs into a causal risk bottleneck, purifying dense visual features into a global risk mask. During autoregressive generation, a support-projected causal consistency objective constrains token-level grounding within the risk-defined support mask. Evaluated on $\text{M}^3\text{TAVR}$, a comprehensive 1,482-patient cohort, TAVR-VLM establishes a new state-of-the-art. It achieves an AUROC of 0.896, boosts CIDEr to 0.936, and drastically reduces the hallucination rate to 8.1\%, thereby improving interpretability for evidence-based surgical AI.

13:00 JSTビジネス/資金調達

大規模言語モデルを使用して縦断的な合成臨床ノートを生成するパイプライン

実世界のデータへのアクセスが制限されている領域で AI システムの開発と評価を可能にするために、合成データの使用が増えています。医療分野では、臨床文書はその機密性により特別な課題を抱えています。この研究では、実際の患者データに伴うプライバシー リスクを回避しながら、臨床 AI ツールの開発をサポートするように設計された合成臨床メモ パイプラインとデータセットを導入します。データセットは、大規模な言語モデルを使用した構造化患者生成、半構造化患者ジャーニー シミュレーション、および非構造化臨床ノート生成を組み合わせたモジュール式パイプラインを使用して生成されます。このパイプラインは、長期的な患者記録全体にわたる内部一貫性を優先すると同時に、書き方、メモの構造、臨床の詳細の変化も捕捉するように設計されています。 LLM ベースの検証および拡張ステップを含む追加のメカニズムを使用して、生成されたノートの忠実性、リアリズム、および多様性が向上します。私たちは、70 人の合成患者のデータセットをリリースします。各患者には、入院期間全体にわたる 20 ~ 50 の臨床ノートが関連付けられています。データセットは複数の検証レベルで提供されているため、ユーザーはユースケースに応じて現実性とスケーラビリティのバランスを取ることができます。このデータセットは、実際の患者データに依存することなく、要約ツール、コーディング モデル、意思決定支援システムなどの臨床 AI システムの開発、テスト、評価をサポートします。

原文 (English)

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

13:00 JST研究/論文

メトリクス順序シーケンストレーニングとハイブリッドポリシー優先最適化を備えた拡散トランスフォーマーを介した生成検索

埋め込みベースの検索では、共有ベクトル空間内のクエリとの類似性によって項目をランク付けし、通常は最高スコアの項目を返すことを目的としています。多くの運用設定では、これは望ましくないことです。きめの細かいパターンを表現するシード セットを考えると、ターゲット属性を満たし、そのパターン内に留まるアイテムがさらに必要になります。これをパターン保持属性の取得として形式化します。 2 つの目標は相互に影響し合っています。シードを平均化すると、パターンは維持されますが、低属性の領域にとどまりますが、グローバルな属性の取得では無関係なパターンに偏ってしまいます。このタスクには、モデルが一連の項目エンベディングを読み取り、最近傍検索用のクエリ エンベディングを生成する、連続生成検索を使用してタスクに取り組みます。私たちは、生シーケンス事前トレーニング、マルチドメイン メトリック順序付け継続事前トレーニング、テールセントロイド微調整、および HPPO を備えた段階的フレームワークである MO-DiT + HPPO を提案します。メトリクス順序トレーニングは、まばらなオンライン検索ラベルを、予測属性密度が低いものから高いものへと順序付けされたパターン内の軌跡に変換し、1 つのモデルにドメイン全体にわたるメトリクス改善の方向性を教えます。 HPPO は、ハイブリッド候補プールにオンライン交差メトリックをラベル付けし、参照に基づいた優先順位の最適化を適用することにより、生成されたクエリ分布を真のオンライン目標に合わせて調整します。パレート ペア フィルターは、同じパターンの純度を低下させない勝者ペアのみを保持し、パターンを犠牲にすることなく属性メトリックを向上させます。項目ホールドアウト プロトコルおよびパターン ホールドアウト プロトコルに基づく 4 つの属性ドメイン全体で、メトリック順序付け DiT は事前学習済み生成検索よりも交差メトリックを改善し、HPPO はそれをさらに改善し、8 つのドメイン分割セルのうち 7 つで大幅な改善が見られ、最も困難な分割では僅差の同点となりました。メトリクスと予測子の検証、順序の除去、CPT/SFT の比較、および候補とポリシーの除去により、利益がどこから得られるのかがわかります。

原文 (English)

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

Embedding-based retrieval ranks items by their similarity to a query in a shared vector space and usually aims to return the highest-scoring items. In many production settings this is not what is wanted: given a seed set that expresses a fine-grained pattern, one needs more items that both satisfy a target attribute and stay within that pattern. We formalize this as pattern-preserving attribute retrieval. The two goals pull against each other: averaging the seeds preserves the pattern but stays in a low-attribute region, while global attribute retrieval drifts to unrelated patterns. We approach the task with continuous generative retrieval, where a model reads a sequence of item embeddings and generates query embeddings for nearest-neighbor search. We propose MO-DiT+HPPO, a staged framework with raw-sequence pretraining, multi-domain metric-ordered continuation pretraining, tail-centroid fine-tuning, and HPPO. Metric-ordered training turns sparse online retrieval labels into in-pattern trajectories ordered from low to high predicted attribute density, teaching one model the metric-improvement direction across domains. HPPO aligns the generated query distribution with the true online objective by labeling a hybrid candidate pool with the online intersection metric and applying reference-anchored preference optimization. A Pareto pair filter keeps only winner pairs that do not lower same-pattern purity, raising the attribute metric without sacrificing the pattern. Across four attribute domains under item- and pattern-holdout protocols, metric-ordered DiT improves the intersection metric over a pretrained generative retriever, and HPPO improves it further, with significant gains on seven of eight domain-split cells and a marginal tie on the hardest split. Metric-predictor validation, order ablations, CPT/SFT comparisons, and a candidate-policy ablation show where the gains come from.

13:00 JST研究/論文

マルチタスク結合モデルからタスクエキスパートを回復する方法を学ぶ

マルチタスク モデルのマージは、複数のタスク固有の専門家を 1 つの統一モデルに統合することを目的としていますが、静的マージではパラメータの干渉が常に発生します。動的マージ モデルはこのギャップを埋めることを目的としていますが、多くの研究は、推論時にコストのかかるストレージと冗長なエキスパート コンポーネントの読み込みに依存しています。この研究では、タスク エキスパートの観点から、パラメータ干渉を、マージ プロセス中に各エキスパートに導入されるパラメータの摂動として見ます。このようなパラメータの摂動はアフィン変換としてモデル化でき、加算オフセットとして近似できることを示します。これらを動機として、パラメータ干渉を元に戻し、単一のマージされたチェックポイントからタスク エキスパートのパフォーマンスを回復するために、これらのオフセットを予測するフレームワークである Recover Task eXpert (ReTeX) を提案します。タスク ID が不明な場合に適切なエキスパートを回復するために、推論前にオフラインで計算された SVD 部分空間署名に基づくルーターフリーのタスク ID を導入します。推論時に、識別子は、指定された入力に対して部分空間が最小の射影残差をもたらすタスクを選択します。その結果、ReTeX は視覚領域と NLP 領域の両方で個人の専門家のパフォーマンスの 95% 以上を回復し、目に見えないタスクへの一般化を大幅に向上させます。重要なことに、パラメータ オフセット予測が、配布外 (OOD) タスクに対する専門知識の創発的適応補間につながることも示します。 ReTeX は、目に見えないタスクを処理するために、目に見える専門知識を適応的に補間します。私たちのコードは https://github.com/BAIKLAB/ReTeX で入手できます。

原文 (English)

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference. In this work, from the perspective of task expert, we view parameter interference as parameter perturbation introduced to each expert during merging process. We show that such parameter perturbations can be modeled as affine transformation, which can be approximated as additive offsets. Motivated by these, we propose Recover Task eXpert (ReTeX), a framework that predicts those offsets, in order to undo parameter interference and recover task-expert performance from a single merged checkpoint. To recover the appropriate expert when task identity is unknown, we introduce a router-free task identifier based on SVD subspace signatures computed offline before inference. At inference, the identifier selects the task whose subspace yields the smallest projection residual for a given input. As a result, ReTeX recovers over 95% of individual-expert performance in both vision and NLP domains, while significantly improving generalization to unseen tasks. Crucially, we also show that the parameter offset prediction leads to emergent adaptive interpolation of expert knowledge for out-of-distribution (OOD) tasks. ReTeX adaptively interpolates seen expert knowledge to handle unseen tasks. Our code is available at https://github.com/BAIKLAB/ReTeX

13:00 JSTエージェント

言語エージェントのタスク非依存性の診断

大規模な言語モデルは有能な長期的なエージェントとして機能しますが、その配布外 (OOD) 一般化は依然として弱いままです。私たちは、この失敗の主な原因はタスクの鈍感であると特定しています。似ているが異なるタスクに直面した場合、モデルはトレーニング中に学習したパターンを適用し、目の前のタスクを解決できない可能性があります。命令が意味的に壊れていて直接応答できない場合でも、モデルは多くの場合、元のタスクに沿ったアクションを続行することを示します。さらに、トレーニングされたプロンプト内のタスクの説明を、類似しているが異なる別のタスクに置き換えた場合でも、モデルは同じアクションを出力する可能性があることがわかりました。この動作には、トレーニング中の注意がタスク トークンから離れてローカルの観察の方に一貫して移ることが伴い、ショートカットへの最適化バイアスが示唆されています。この問題を軽減するために、タスク命令へのアクションの依存を明示的に促進する軽量の対照的正則化装置である Task-Perturbed NLL Optimization を提案します。広範な評価により、私たちの介入により、タスクトークンに対するより安定した注意が維持されながら、タスクの感度とOODの一般化が向上することが示されました。

原文 (English)

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of this failure as task insensitivity: when faced with similar but distinct tasks, models might apply patterns learned during training and fail to solve the task at hand. We show that models often continue with actions aligned with the original task even when the instruction is semantically corrupted and cannot be directly answered. We further find that, when we replace the task description in a trained prompt with another similar but distinct task, the model may still output the same action. This behavior is accompanied by a consistent training-time attention drift away from task tokens and toward local observations, suggesting an optimization bias toward shortcuts. To mitigate this problem, we propose Task-Perturbed NLL Optimization, a lightweight contrastive regularizer that explicitly encourages action dependence on the task instruction. Extensive evaluations show that our intervention improves task sensitivity and OOD generalization while preserving more stable attention to task tokens.

13:00 JSTLLM/生成AIエージェント

CoT トレーニングは LLM ベースのエージェントにどのような影響を及ぼしますか?

思考連鎖 (CoT) 推論は言語モデル エージェントで広く使用されていますが、これまでの研究では、言語化された CoT が常に忠実であるとは限らず、事後推論を反映している可能性があることが示されています。これは、モデルが推論する前にすでに答えを知っていることを意味します。したがって、CoT トレーニングによって実際に何が改善されているかを尋ねます。モデルは、生成された推論を通じてアクションを変更することがうまくなっているのでしょうか、それともプロンプトから直接アクションを予測することがうまくなっているのでしょうか? \emph{プロンプトアクション} (CoT なしでアクションを予測する) と CoT アクション (CoT ありでアクションを予測する) を比較することで、この問題を研究します。チェックポイント全体で、迅速なアクションの品質が大幅に向上します。環境と対話している間、プロンプト アクションに対する CoT アクションの相対的な利点は同様のままであり、CoT トレーニングによって CoT 推論の利点が広がることはなく、プロンプト アクションの質の向上に役立つことが示されています。さらに、後のチェックポイントでは CoT に応じてアクションを修正する可能性が低く、プロンプトへの依存度が高まっていることがわかります。これらのパターンに動機付けられて、トレーニング サンプルの一部でアクション トークンの監視を選択的にマスクします。この介入により、領域外の一般化が向上します。

原文 (English)

Where Do CoT Training Gains Land in LLM based Agents?

Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may instead reflect post-hoc reasoning, which means the model already knows the answer before reasoning. We therefore ask what CoT training is actually improving: is the model getting better at changing its action through generated reasoning, or is it getting better at predicting the action directly from the prompt? We study this question by comparing \emph{prompt actions} (predicting action without CoT) with CoT actions (predicting action with CoT). Across checkpoints, prompt-action quality improves substantially. While interacting with the environment, the relative advantage of CoT actions over prompt actions remains similar, showing that CoT training does not widen the advantage of CoT reasoning, and it helps to improve the quality of prompt actions. We further find that later checkpoints are less likely to revise the action in response to CoT, suggesting greater reliance on the prompt. Motivated by these patterns, we selectively mask action-token supervision on a fraction of training examples. This intervention improves out-of-domain generalization.

13:00 JST画像/動画生成

Look-Before-Move: ダイナミックな 3D ストーリーワールドにおける物語に基づいた世界の視覚的注意

身体化された AI と世界モデルが動的 3D 環境で動作することが増えているため、視覚認識は、与えられた観察を受動的に解釈するだけでなく、何を観察するかを能動的に決定する方向に進む必要があります。私たちは、動的な 3D ストーリー世界でのカメラ計画を通じてこの問題を研究します。そこでは、カメラは滑らかな動きを生成するだけでなく、移動する前にどのような視覚的証拠を取得する必要があるかを決定する必要があります。私たちはこの機能を、物語に基づいた世界の視覚的注意として定式化します。カメラは、何を観察するか、どのように観察を構成するか、そして物語の意図と物理的な 3D 制約の下で時間の経過とともにどのように注意を移すかを決定する具体化された観察者として機能します。この機能を実現するために、観察仕様をモーション実行から分離するカメラ計画フレームワークである Look-Before-Move を提案します。まずセマンティック観察コントラクトを構築して、監督の意図を実行可能な視覚的制約に変換し、次にモンテカルロ視点検索を実行して物語に準拠し、幾何学的に実現可能な視点を見つけます。最後にセマンティック軌道グラウンディングを適用して、選択された視点を連続的で衝突を認識し、時間的に一貫したカメラの動きに接続します。さらに、StoryBlender に基づいて動的な 3D ストーリー ワールド ベンチマークを構築し、アニメーション キャラクター、セマンティック シーン構成、および実行可能な 3D 環境を含む 50 のストーリー、457 のシーン、および 1585 のショットをカバーします。実験では、私たちのフレームワークが代表的なベースラインよりも被写体の知覚、意図の一貫性、軌跡の品質を向上させることが示されており、カメラの動きを生成する前に視覚的な注意を組織することの重要性が実証されています。

原文 (English)

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as Narrative-Grounded World Visual Attention, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose Look-Before-Move, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on StoryBlender, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.

13:00 JSTLLM/生成AI画像/動画生成

アインシュタインの世界モデル

知性には、直接の経験を超えて現象について推論する能力が必要ですか?言語だけでは複雑な思考を捉えることができないのではないかと疑うのは自然なことです。しかし、この研究で特に懸念されるのは、反事実的な出来事を視覚化することが、複雑な思考のメカニズムとして言語を補完できるかどうかである。私たちは、LLM がそのような視覚化メカニズムを、推論能力に役立つ方法で利用できるように訓練できるかどうかを尋ねます。この疑問を動機として、私たちはアインシュタイン世界モデルを提案します。 EWM は、推論トレース内に視覚と時間のロールアウトを配置する LLM ベースの推論システムの青写真であり、テキストだけでは十分にサポートできない方法で推論できるようになります。 EWM では、LLM はワールド モジュール (ワールド モデルと混同しないでください) を呼び出し、検討中のシーンの短いロールアウトを生成します。返されたロールアウトは、答えとしてではなく、後の推論をサポートできる検査可能な仮説として扱われます。 Einstein World Models は、ツール呼び出し (Web 検索やコード実行など) のための LLM の機能を視覚的な思考実験の領域に拡張します。

原文 (English)

Einstein World Models

Does intelligence require the ability to reason about phenomena beyond direct experience? It is natural to suspect that some complex thought cannot be captured through language alone. However, of particular concern to this work, is whether visualising counterfactual events can complement language as a mechanism for complex thought. We ask whether LLMs can be trained to utilise such visualisation mechanisms, in a way that benefits their reasoning abilities. Motivated by this question, we propose Einstein World Models. EWMs are a blueprint for LLM-based reasoning systems that place visual-temporal rollouts inside the reasoning trace, allowing them to reason in ways that text alone may not support well. In an EWM, the LLM calls a world-module (not to be confused with a world model), to produce short rollouts of scenes under consideration. The returned rollout is treated not as the answer, but as an inspectable hypothesis that can support later reasoning. Einstein World Models extend the capability of LLMs for tool calling (such as web search or code execution), into the domain of visual thought experiments.

13:00 JST研究/論文

Resilient AI のための適応型ユーティリティ主導のリソース オーケストレーション (AURORA-AI)

最新の AI システムは、非定常の計算条件、人口統計条件、運用条件の下で導入されることが増えており、静的なリソース割り当て戦略により、予測パフォーマンスと、公平性や説明可能性などの人間中心の特性の両方が低下します。この論文では、Hamilton-Jacobi-Bellman フィードバック制御、リアプノフベースの安定性モニタリング、および公平性を意識した複合ユーティリティを単一の閉ループ ポリシーに統合する、Resilient AI 向けの適応ユーティリティ主導型リソース オーケストレーション フレームワークである AURORA-AI について紹介します。このフレームワークは、異種 AI モデルの母集団全体に計算予算を継続的に再配分するため、予測パフォーマンス、人口統計的均等性、コストを合わせて定義されたグローバル ユーティリティが、遅延、堅牢性、解釈可能性は、混乱下でも最大化されたままになります。このフレームワークは、人口統計上のバイアス ショック、段階的な概念ドリフト、突然のブラック スワンの混乱を同時に注入する、ストレスの多い離散時間シミュレーションで評価され、静的、ラウンド ロビン、貪欲、LinUCB、および近接ポリシー最適化に基づく深層強化学習エージェントを含む 5 つの確立されたコントローラーと比較されます。 AURORA-AI は、静的ベースラインの 88 タイム ステップと近接ポリシー最適化の 22 タイム ステップと比較して、ブラック スワン イベントからの即時回復を達成し、アルファ分位と超分位をそれぞれ 29 パーセントと 25 パーセント引き上げ、平均と最大の人口統計的パリティ ギャップを同時に削減し、リアプノフ安定動作ステップの割合を増加させます。これらの結果は、安定性理論に基づいた公平性を意識した適応型オーケストレーションが、回復力のある人間中心の AI 導入に向けた実践的かつ理論的に動機付けられた道であることを示しています。

原文 (English)

Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI)

Modern AI systems are increasingly deployed under non-stationary computational, demographic, and operational conditions in which static resource allocation strategies degrade both predictive performance and human-centric properties such as fairness and explainability. This paper presents AURORA-AI, an Adaptive Utility-driven Resource Orchestration framework for Resilient AI that unifies Hamilton-Jacobi-Bellman feedback control, Lyapunov-based stability monitoring, and a fairness-aware composite utility into a single closed-loop policy.The framework continuously redistributes computational budget across a population of heterogeneous AI models so that the global utility, defined jointly over predictive performance, demographic parity, cost, latency, robustness, and interpretability, remains maximised under disruption. The framework is evaluated in a stress-rich discrete-time simulation that concurrently injects demographic bias shocks, gradual concept drift, and abrupt black-swan disruptions, and is compared against five established controllers including Static, Round Robin, Greedy, LinUCB, and a deep reinforcement-learning agent based on Proximal Policy Optimisation. AURORA-AI achieves immediate recovery from the black-swan event compared to eighty-eight time steps for the Static baseline and twenty-two for Proximal Policy Optimisation, lifts the alpha-quantile and the super-quantile by twenty-nine and twenty-five percent respectively, simultaneously reduces the mean and maximum demographic parity gap, and increases the fraction of Lyapunov-stable operating steps. These results indicate that fairness-aware adaptive orchestration grounded in stability theory is a practical and theoretically motivated path toward resilient human-centric AI deployment.

13:00 JSTLLM/生成AIエージェント

反復 LLM エージェント ループのセマンティック早期停止

マルチエージェント大規模言語モデル (LLM) ループ (たとえば、草稿を作成するライターと改訂を行う批評家など) は、ほとんどの場合、固定の反復上限 (max_iterations) によって終了します。これは構文上のキルスイッチです。答えがまだ改善されているかどうかが分からないため、簡単な入力にトークンを過剰に消費し、難しい入力を切り捨てます。私たちはセマンティックな早期停止を研究します。つまり、連続するドラフト埋め込みの意味の変化 (忍耐ウィンドウによるコサイン距離) が停止し、回答の測定品質の向上が停止すると、ループが停止します。私たちの仕事は 3 つの貢献をします。まず、正直な理論的基礎です。距離数列の収束を、(以前に過剰に主張されていた)バナッハ短縮ではなく、経験的にテストされた予想として扱いながら、決定論的な終端と明確な定義性を証明し、これらの主張を機械チェックします。 2 番目に、ジャッジの効率的な評価プロトコルです。各質問の完全な軌跡を 1 回生成し、同一のドラフトですべての停止ポリシーを再生し、すべての LLM ジャッジ呼び出しをキャッシュして、厳密にペアになった効率と品質の比較を低コストで実現します。さらに、運用トークン (ポリシーにチャージ) を評価トークン (測定手段) から分離します。 3 番目は、マルチホップ検索拡張質問応答 (HotpotQA) に関する実証研究です。 60 問のテスト分割では、ジャッジフリーのセマンティック ストッパーはパリティ品質 (Delta-IS = -0.004、p = 0.81) での max_iterations と比較してオペレーショナル トークンを 38% 削減しますが、完全な品質ゲートのバリアントはラウンドごとの判定がコストを支配するため逆効果です。最良のラウンドを選択したオラクルは、すべての実際的なポリシーに対して +0.115 の情報スコアを達成し (p ~ 4e-11)、問題を「いつ停止するか」 (簡単) から「どのラウンドが最適か」 (オープン) に再構成します。

原文 (English)

Semantic Early-Stopping for Iterative LLM Agent Loops

Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a fixed iteration cap (max_iterations). This is a syntactic kill-switch: it is blind to whether the answer is still improving, so it over-spends tokens on easy inputs and truncates hard ones. We study semantic early-stopping: the loop halts when consecutive draft embeddings stop changing in meaning (cosine distance with a patience window) and the answer's measured quality stops improving. Our work makes three contributions. First, an honest theoretical footing: we prove deterministic termination and well-definedness and machine-check these claims, while treating the convergence of the distance sequence as an empirically tested conjecture rather than a (previously over-claimed) Banach contraction. Second, a judge-efficient evaluation protocol: we generate each question's full trajectory once, replay every stopping policy over the identical drafts, and cache every LLM-judge call, yielding a strictly paired efficiency-versus-quality comparison at low cost; we further separate operational tokens (charged to a policy) from evaluation tokens (a measurement instrument). Third, an empirical study on multi-hop retrieval-augmented question answering (HotpotQA). On the 60-question test split, a judge-free semantic stopper reduces operational tokens by 38% relative to max_iterations at parity quality (Delta-IS = -0.004, p = 0.81), whereas the full quality-gated variant is counter-productive because its per-round judging dominates cost. An oracle that selects the best round attains +0.115 Information Score over every practical policy (p ~ 4e-11), reframing the problem from "when to stop" (easy) to "which round is best" (open).

13:00 JSTビジネス/資金調達

グラウンドトゥルースを使用してクラスタリングを評価するにはどうすればよいですか?

グランド トゥルースが利用可能な場合、外部インデックスをクラスター評価に使用できます。セットマッチングベースの尺度に焦点を当てて、最も一般的な外部妥当性指標をレビューします。セントロイド インデックス (CI) は、説明可能な結果が得られる直感的なクラスター レベルの測定であるため、推奨します。より細かく調整されたポイントレベルの測定が必要な場合は、より多くの選択肢があります。ペアセット インデックス (PSI) は、クラスター サイズによって偏らない正規化されたスコアを提供します。すべてのポイントが同等に重要である必要がある場合は、クラスタリング精度 (ACC) またはその他のセットマッチング尺度が適しています。

原文 (English)

How to evaluate clustering with ground truth?

External indexes can be used for cluster evaluation when ground truth is available. We review the most common external validity indexes focusing on set-matching-based measures. We recommend centroid index (CI), because it is an intuitive cluster-level measure with an explainable result. If we need a more fine-tuned, point-level measure, there are more choices. Pair-set index (PSI) provides a normalized score which is not biased by cluster sizes. If all points should matter equally, then clustering accuracy (ACC) or any other set-matching measure is suitable.

13:00 JSTLLM/生成AIエージェント

大規模言語モデルエージェントの経験則とポリシーの共同学習

マルチステップのインタラクティブ環境における LLM エージェントにとっての重要な課題は、蓄積されたインタラクション経験を効果的に活用することです。既存の研究では通常、そのようなエクスペリエンスを 2 つの使用法に分けています。1 つは、後でプロンプトを表示するための自然言語ルールとしてモデルの外に保持するか、軌道とフィードバックを使用してモデル パラメーターを更新するかです。前者は解釈しやすいですが、進化するポリシーと同期しなくなる可能性があります。後者はポリシーをより広範囲に改善しますが、スパース報酬設定における局所的な間違いに対する修正は限定的です。我々は、LLM エージェントのための経験的ルールとポリシーの共同学習 (JERP) を紹介します。これは、同じ対話の軌跡から長期的な経験的ルール プールとポリシーを更新します。意思決定時に、JERP はタスク関連のルールを取得し、対話履歴とともにエージェントにそれらのルールを条件付けします。各エピソードの後、収集された軌跡を使用してポリシーを最適化し、現在のロールアウトを参照の成功した軌跡と比較することでルール プールを修正します。この結合により、ルール プールが進化するポリシーに合わせて維持されると同時に、安定した効果的な動作がモデル自体に徐々に吸収されることが可能になります。 AlfWorld と WebShop での実験では、JERP が複雑な対話型タスクの意思決定パフォーマンスにおいて一貫した向上をもたらすことが示されています。

原文 (English)

Joint Learning of Experiential Rules and Policies for Large Language Model Agents

For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or using trajectories and feedback to update the model parameters. The former is easy to interpret but can fall out of sync with the evolving policy; the latter improves the policy more broadly but provides only limited correction for local mistakes in sparse-reward settings. We present Joint Learning of Experiential Rules and Policies for LLM Agents (JERP), which updates a long-term experiential-rule pool and the policy from the same interaction trajectories. At decision time, JERP retrieves task-relevant rules and conditions the agent on them together with the interaction history. After each episode, it uses the collected trajectories both to optimize the policy and to revise the rule pool by comparing current rollouts with reference successful trajectories. This coupling keeps the rule pool aligned with the evolving policy while allowing stable and effective behaviors to be gradually absorbed into the model itself. Experiments on AlfWorld and WebShop show that JERP yields consistent gains in decision performance for complex interactive tasks.

13:00 JSTLLM/生成AIエージェント

OpenRCA 2.0: 結果ラベルから因果関係プロセスの監視まで

根本原因分析 (RCA) では、長いコンテキストの理解、複数ステップの推論、ツールの使用など、LLM エージェントの機能の総合的なテストが行​​われます。ただし、既存のデータセットには根本的なギャップがあります。つまり、根本原因のみにラベルが付けられ、観察された症状につながる伝播経路はラベル付けされないため、単純なパターン マッチングのタスクが大幅に簡素化されます。厳密な評価をサポートするために、フォールト挿入による既知の介入を利用して因果伝播パスを再構築する段階的なラベル付けプロトコルである PAVE を導入します。このメカニズムは前方検証です。つまり、症状から逆方向に推論するのではなく、原因から結果に至るまで推論します。 PAVE を適用すると、LLM エージェントに対する段階的な因果的アノテーションを備えた最初のクロスシステム RCA ベンチマークである OpenRCA 2.0 (500 インスタンス) が生成されます。 11 のフロンティア LLM 全体で、正確な根本原因セットの回復に成功するのは、平均して 20.7% のケースのみです。この困難がどこにあるのかを特定するために、基準を緩和して、根拠のない診断と呼ばれるものを見つけます。エージェントは、ケースの 76.0% で少なくとも 1 つの正しい根本原因サービスを特定しますが、そのサービスを観察された症状への検証された因果伝播経路に根拠付けるのは 61.5% のみです。結果のみの評価では、この失敗モードが隠蔽されます。段階的因果的グラウンドトゥルースは、信頼できる LLM ベースの RCA エージェントに欠けている部分です。

原文 (English)

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed symptom, which largely simplifies the task to naive pattern matching. To support rigorous evaluation, we introduce PAVE, a step-wise labeling protocol that leverages known interventions from fault injection to reconstruct causal propagation paths. The mechanism is forward verification: reasoning from cause to effect rather than inferring backward from symptoms. Applying PAVE yields OpenRCA 2.0 (500 instances), the first cross-system RCA benchmark with step-wise causal annotations for LLM agents. Across 11 frontier LLMs, recovering the exact root-cause set succeeds in only 20.7% of cases on average. To locate where this difficulty lies, we relax the criterion and find what we call the ungrounded diagnosis: agents identify at least one correct root-cause service in 76.0% of cases, but ground that service in a verified causal propagation path to the observed symptom in only 61.5%. Outcome-only evaluation hides this failure mode; step-wise causal ground truth is the missing piece for trustworthy LLM-based RCA agents.

13:00 JSTLLM/生成AI

TOPS: 効率的な MLLM 推論のためのトークン最適保存セットの構築による第一原理のビジュアル トークン プルーニング

マルチモーダル大規模言語モデル (MLLM) は強力なマルチモーダル推論機能を実現していますが、その効率は大量のビジュアル トークンによって制限され、これによりかなりの計算オーバーヘッドが発生します。視覚的なトークン プルーニングは自然な解決策を提供しますが、既存の方法は不完全です。注意ベースの基準は冗長なトークンを保持する傾向がありますが、多様性ベースの基準はユーザーの指示に依存しないことがよくあります。複数の基準を組み合わせた方法であっても、トークン プルーニングの本質的な目的を原理的に定式化することがまだできていません。このペーパーでは、第一原理の観点から視覚的なトークン プルーニングを再検討し、それをトークン最適保存セットの構築として定式化します。トップダウンの情報理論分析を通じて、効果的なトークン選択のための 3 つの基本原則、つまりタスクの関連性、情報の網羅性、およびセマンティックの多様性を特定します。これらの原則に基づいて、さまざまな MLLM に適用できる、トレーニング不要でモデルに依存しない枝刈りモジュールである TOPS を提案します。 7 つの MLLM バックボーンと 14 のベンチマークに関する広範な実験により、TOPS がさまざまなプルーニング設定の下で従来の方法よりも優れたパフォーマンスを発揮することが実証されました。特に、LLaVA-NeXT では、TOPS は 7B モデルと 13B モデルでそれぞれ 100.0% と 100.6% のパフォーマンスを維持しながら、ビジュアル トークンの 77.8% を削除します。これは、冗長なビジュアル トークンを削除することで幻覚を緩和し、将来の軽量 MLLM 設計を刺激できる可能性があることを示唆しています。

原文 (English)

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference

Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introduces substantial computational overhead. Visual token pruning offers a natural solution, yet existing methods are imperfect: attention-based criteria tend to retain redundant tokens, while diversity-based criteria are often agnostic to user instructions. Even methods that combine multiple criteria still lack a principled formulation of the intrinsic objective of token pruning. In this paper, we revisit visual token pruning from a first-principles perspective and formulate it as constructing Token Optimal Preservation Sets. Through a top-down information-theoretic analysis, we identify three fundamental principles for effective token selection: Task Relevance, Information Coverage, and Semantic Diversity. Based on these principles, we propose TOPS, a training-free and model-agnostic pruning module that can be applied to various MLLMs. Extensive experiments on 7 MLLM backbones and 14 benchmarks demonstrate that TOPS outperforms prior methods under diverse pruning settings. Notably, on LLaVA-NeXT, TOPS removes 77.8% of visual tokens while preserving 100.0% and 100.6% performance on its 7B and 13B models, respectively, suggesting that pruning redundant visual tokens can sometimes mitigate hallucination and inspire future lightweight MLLM design.

13:00 JSTエージェント

レガシー ワークフローを Agentic BPM に引き上げるプロセス ハーネス: CUGA FLO での設計と実現

基盤となるワークフロー エンジンを置き換えることなく、レガシー ワークフローをエージェントティック ビジネス プロセス管理 (エージェント BPM) に引き上げる新しいメカニズムであるプロセス ハーネスを導入します。プロセス ハーネスは、決定論的ワークフロー エンジンの周囲にポリシーで管理されるエージェント層を配置し、エンジンがプロセスに対する構造的な権限を保持しながら、指定された制御ポイントをインターセプトして推論、適応、監視に貢献します。プロセス ハーネスを厳密に定義するために、データ スキーマと実行セマンティクスの両方を指定するタスク決定フロー (TDF) モデルを開発します。 TDF は、3 つのポリシー管理エージェント タイプにわたる LLM 推論を分解します。1 つは知識集約型タスクの実行のための TaskAgent、1 つはケースごとのゲートウェイ ルーティングのための DecisionAgent、そして原則に基づいたフック メカニズムを通じてランタイム フローの適応を管理する FlowAgent です。各エージェントは、システム内のすべての LLM 呼び出しを管理する集約ポリシー セットであるプロセス FRAME から抽出された明示的なポリシー内で判断します。次に、TDF モデルの設計と実装の実現として CUGA FLO を示し、3 つのエージェント タイプすべてとフック主導の規制オーバーライドを実行するローン承認ワークフローでそれを実証します。プロセスハーネスは、構造的コンプライアンスを強制する決定論的なワークフローの実行を通じて実現される命令的要件と、プロセスが要求するどこにでも指定された制御ポイントで呼び出されるポリシーに基づいたエージェントの自律性を通じて実現される規範的要件を独自に調和させます。

原文 (English)

A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO

We introduce the process harness, a new mechanism for uplifting legacy workflows into Agentic Business Process Management (Agentic BPM) without replacing the underlying workflow engine. A process harness places a policy-governed agentic layer around a deterministic workflow engine, intercepting designated control points to contribute reasoning, adaptation, and oversight while the engine retains structural authority over the process. To define the process harness rigorously, we develop the Task-Decision-Flow (TDF) model, specifying both its data schema and its execution semantics. TDF decomposes LLM reasoning across three policy-governed agent types: a TaskAgent for knowledge-intensive task execution, a DecisionAgent for per-case gateway routing, and a FlowAgent that governs runtime flow adaptation through a principled hook mechanism. Each agent reasons within an explicit policy drawn from the process FRAME, the aggregate policy set governing all LLM calls in the system. We then present CUGA FLO as the design and implementation realization of the TDF model, and demonstrate it on a loan approval workflow that exercises all three agent types and hook-driven regulatory override. The process harness uniquely reconciles imperative requirements, realized through deterministic workflow execution that enforces structural compliance, with normative requirements, realized through policy-framed agentic autonomy invoked at designated control points wherever the process demands it.

13:00 JST研究/論文

進化的に生成された敵対的テキストに対する自然言語分類子の脆弱性

深層学習モデルは、さまざまな分野で目覚ましいパフォーマンスを達成していますが、特に NLP では、敵対的な入力に対して脆弱なままであり、そのような攻撃は現実世界に重大な影響を与える可能性があります。敵対的攻撃には、NLP モデルをだますために、意味的に類似した小規模なトークン置換が含まれることが多く、最近の手法は、多くの場合、モデルの内部構造へのある程度のアクセスを悪用して、特定の脆弱な単語をターゲットにすることで、より正確になっています。この論文では、自然言語モデルに対する敵対的攻撃を生成するハイブリッド遺伝アルゴリズム (GA) である GAversary を提案します。 GA はターゲット モデルをブラック ボックスとして扱うことができ、検索のガイドとしてモデルが出力するロジット値のみを必要とします。 GAversary は、GloVe 埋め込みを使用して単語の置換 (突然変異演算子) を提案し、敵対的な例の意味上の類似性を改善するという点で、この問題に対して以前に提案された GA とは異なります。 GAversary は、いくつかのベンチマーク データ セットとよく知られたターゲット モデルに適用されます。 GAversary は、BAE および A2T 攻撃と比較して、テスト データに対するターゲット モデルの精度を大幅に低下させることができます (最良のケースでは、BAE の 27.6% と比較して、76.8% の精度が 5.8% に低下します)。トレードオフとして、GAversary は他の 2 つの方法に比べて 2 倍弱の単語を摂動させますが、元のテキストとの意味上の類似性はわずかに低くなり、実行時間は約 5% 増加します。

原文 (English)

Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text

Deep learning models have achieved impressive performance across various fields but remain vulnerable to adversarial inputs, particularly in NLP, where such attacks can have significant real-world consequences. Adversarial attacks often involve small, semantically similar token replacements to fool NLP models, and recent methods have become more precise by targeting specific vulnerable words, often by exploiting some level of access to the model's internal structure. This paper proposes GAversary, a hybrid Genetic Algorithm (GA) to generate adversarial attacks on natural language models. The GA is able to treat the target model as a black box, requiring only the logit value output by the model to guide the search. GAversary differs from GAs previously proposed for this problem by using GloVe embeddings to propose word replacements (the mutation operator) to improve the semantic similarity of the adversarial examples. GAversary is applied to several benchmark data sets and well-known target models. GAversary is able to substantially reduce the target model's accuracy on test data compared to the BAE and A2T attacks compared against (in the best case, reducing a 76.8% accuracy to 5.8%, compared to BAE's 27.6%). The trade-off is that GAversary perturbs just under twice as many words as the other two methods, with a slightly lower semantic similarity to the original text and around a 5% increase in run-time.

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達

判断せずに質問する: 解釈可能な LLM 評価と自己改善のための 2 つの質問

NLP では、LLM 出力の評価が依然として大きなボトルネックとなっています。人間による評価は高価で時間がかかり、語彙メトリクスとオープンエンド生成に関する人間の判断との相関性が低く、全体的な LLM ジャッジはデバッグが難しい不透明なスコアを生成することがよくあります。私たちは、評価基準をアトミックなバイナリの質問に分解し、その結果の判定を解釈可能な多次元スコアに集約するフレームワークである BINEVAL を提案します。タスク プロンプトが与えられると、メタ プロンプトが詳細な評価質問を生成し、LLM が出力ごとに独立して質問に回答し、調整された全体スコアとともに透明な質問レベルのフィードバックを生成します。この分解により、評価が検査、​​診断が容易になり、迅速な改善に直接使用できるようになります。 SummEval、Topical-Chat、QAGS 全体で、BINEVAL は UniEval や G-Eval などの強力なベースラインと同等またはそれを上回り、特に QAGS などの事実整合性ベンチマークで優れた結果を示しています。 BINEVAL は、人間の判断との競合相関を超えて、人間のスコア分布とよりよく一致し、以前の LLM ジャッジによく見られた天井効果を回避し、境界線にある出力と明らかに欠陥のある出力をより適切に区別することにつながります。さらに、同じ質問レベルのフィードバックが反復プロンプトの最適化をサポートし、自己更新設定とクロスモデル更新設定の両方で IFBench での要約に関する評価者のプロンプトと生成プロンプトを改善することを示します。全体として、BINEVAL は、強力な経験的パフォーマンスと実用的な診断および最適化の価値を組み合わせた、タスクに依存せず、トレーニング不要で、解釈可能な評価フレームワークを提供します。

原文 (English)

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.

13:00 JST画像/動画生成

EO-WM: 確率論的地球観測予測のための物理情報に基づいた世界モデル

地球観測 (EO) 予測は、変化する気象条件下での衛星観測から将来の地球表面のダイナミクスを予測することを目的としています。この論文では、このタスクを、部分的に観測された気象主導型の世界モデリング問題とみなします。この問題では、気象は条件付け信号として機能しますが、観測がまばらで地表面の状態が観測されていないため、予測は依然として不確実です。しかし、既存の手法はこの設定を完全には捉えていません。決定論的モデルは不確実性を 1 つの将来予測にまとめますが、拡散ベースの手法は通常、気象変数を未分化の条件付け信号として扱い、既存のベンチマークは、予報が変化する気象強制に正しく反応するかどうかではなく、主に再構成精度に焦点を当てています。マルチスペクトル EO 予測用のビデオ拡散トランスフォーマである EO-WM を紹介します。 EO-WM には、気候ベースライン、気象異常、累積的な物理的ストレス信号による気象強制力を表す、物理的情報に基づいた調整フレームワークが組み込まれています。具体的には、明確な条件付け経路を通じてベースラインと異常を分離し、持続的な熱と干ばつストレスを捕捉するために時間の経過とともに異常な強制力を蓄積します。標準的な指標を超えて気象応答動作を評価するために、2 つの診断ベンチマークを導入します。1 つは、異常気象下での植生劣化の深刻度を意識した予測のための極端な夏季ベンチマークで、もう 1 つは、変化する気象強制下での応答忠実度をテストするための季節一致ペア ベンチマークです。実験の結果、EO-WM は、標準的なピクセル レベルのメトリクスで競争力を維持しながら、予測される正規化植生指数 (NDVI) の減少振幅の誤差を相対的に 5.63% 削減し、方向性ヒット率を相対的に 7.80% 改善することが示されています。ベンチマークとモデルは https://github.com/Luo-Z13/EO-WM でオープンソース化されます。

原文 (English)

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as undifferentiated conditioning signals, and existing benchmarks focus mainly on reconstruction accuracy rather than whether forecasts respond correctly to changed weather forcing.We introduce EO-WM, a video diffusion transformer for multispectral EO forecasting. EO-WM incorporates a physically informed conditioning framework that represents meteorological forcing through a climatological baseline, weather anomalies, and cumulative physical stress signals. Specifically, it separates baseline and anomaly through distinct conditioning pathways, and accumulates anomalous forcing over time to capture sustained heat and drought stress. To evaluate weather-response behavior beyond standard metrics, we introduce two diagnostic benchmarks: an Extreme Summer Benchmark for severity-aware prediction of vegetation degradation under extreme weather, and a Seasonal Matched-Pair Benchmark for testing response fidelity under changed weather forcing. Experiments show that EO-WM reduces the error in predicted Normalized Difference Vegetation Index (NDVI) decline amplitude by a relative 5.63% and improves directional hit rate by a relative 7.80%, while remaining competitive on standard pixel-level metrics. The benchmarks and model will be made open-source at https://github.com/Luo-Z13/EO-WM.

13:00 JST研究/論文

疫学モデルにおける迅速なベイズパラメータ推定のためのシミュレーションベースの推論: MCMC との比較

機械的疫学モデルは、感染症の予測と公衆衛生の意思決定をサポートするために広く使用されています。このようなモデルのベイジアン キャリブレーションは、マルコフ連鎖モンテカルロ (MCMC) を使用して実行されるのが一般的ですが、高次元の非線形システムや繰り返されるほぼリアルタイムの解析では、計算コストが高くなる可能性があります。ここでは、2020年のドイツの新型コロナウイルス感染症集中治療室(ICU)占有率データを使用した機械的SECIR疫学モデルのベイジアンキャリブレーションのスケーラブルな代替手段として、神経事後推定を使用したシミュレーションベース推論(SBI)を調査します。31日間の推論ウィンドウと、複数の伝達変化点を含む実質的により困難な201日間の再構成問題の両方を使用して、複数の流行期にわたってSBIとMCMCを比較しました。事後一致は、ワッサーシュタイン距離とカルバック・ライブラー発散を事後予測チェックとともに使用して定量的に評価されました。 31 日間のウィンドウにわたって、SBI は、観察された ICU 軌跡を正確に再現しながら、MCMC と強く一致する事後分布を回復しました。 201 日の設定では、不確実性が増大したにもかかわらず、SBI は支配的な後方構造を保存しました。 SBI は、CPU と GPU リソースを組み合わせることで、CPU 上での実行に制限されていた MCMC と比較して、計算実行時間を大幅に短縮しました。 MCMC では 31 日間の推論問題に約 1000 秒を要しましたが、SBI では単一の GPU で約 60 ~ 70 秒で同等の事後予測パフォーマンスを達成しました。 201 日の推論問題の場合、SBI では平均 157 秒かかりましたが、MCMC の実行には 19,000 秒以上かかりました。私たちの結果は、SBI が機構的疫学モデルのベイジアン校正のための迅速かつ計算効率の高いフレームワークを提供し、ほぼリアルタイムの推論と迅速なアウトブレイク分析の繰り返しをサポートすることを示しています。

原文 (English)

Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC

Mechanistic epidemiological models are widely used to support infectious disease forecasting and public-health decision making. Bayesian calibration of such models is commonly performed using Markov chain Monte Carlo (MCMC), which can become computationally expensive for high-dimensional nonlinear systems and repeated near-real-time analyses. Here, we investigate simulation-based inference (SBI) using neural posterior estimation as a scalable alternative for Bayesian calibration of a mechanistic SECIR epidemiological model using COVID-19 intensive care unit (ICU) occupancy data from Germany during 2020. We compared SBI and MCMC across multiple epidemic phases using both 31-day inference windows and a substantially more challenging 201-day reconstruction problem involving multiple transmission change points. Posterior agreement was evaluated quantitatively using Wasserstein distances and Kullback-Leibler divergences together with posterior predictive checks. Across the 31-day windows, SBI recovered posterior distributions in strong agreement with MCMC while accurately reproducing observed ICU trajectories. In the 201-day setting, SBI preserved the dominant posterior structure despite increased uncertainty. SBI, by combining CPU and GPU resources, substantially reduced computational runtime compared with MCMC, which was restricted to running on CPUs. Whereas MCMC required approximately 1000 seconds for the 31-day inference problems, SBI achieved comparable posterior and predictive performance in approximately 60-70 seconds on a single GPU. For the 201-day inference problem, SBI required an average of 157 seconds, while the MCMC runs took over 19,000 seconds. Our results demonstrate that SBI provides a rapid and computationally efficient framework for Bayesian calibration of mechanistic epidemiological models, supporting repeated near-real-time inference and rapid outbreak analysis.

13:00 JSTLLM/生成AI

大規模な言語モデルを使用した自動 R\'esum\'e スクリーニングでの迅速なインジェクション: シングルおよびマルチインジェクション設定

大規模言語モデル (LLM) は、求職者のスクリーニングとランク付けにますます使用されており、候補者がアルゴリズム採用システムを戦略的に操作するインセンティブが生まれています。私たちは、自動化された履歴書スクリーニングにおける即時注入について研究しています。これは、新しい資格を導入するものではありませんが、LLM 評価に影響を与えるように設計された微妙な自己宣伝テキストとして定義されます。対照実験を使用して、論文の質が均一で、注入する候補者がほとんどいない場合、即時注入により確実に応募者のランキングが向上することが示されました。しかし、その有効性は、より多くの候補者が注入されるにつれて急速に減少し、操作が広範囲に及ぶと崩壊します。候補者の品質が不均一な場合、プロンプト注入の効果は平均して低くなりますが、場合によっては、低品質の候補者が高品質の候補者を上回ってしまう可能性があり、公平性に関する懸念が生じます。全体として、LLM ベースのスクリーニングは、操作がまれで候補者の品質の差が小さい場合に最も脆弱になります。コードとリソースは、https://github.com/preetb1199/Prompt_Injection_ACL26 で公開されています。

原文 (English)

Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and Multi-Injection Settings

Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candidates to strategically manipulate algorithmic hiring systems. We study prompt injection in automated r\'esum\'e screening, defined as subtle self-promotional text that introduces no new qualifications but is designed to influence LLM evaluations. Using controlled experiments, we show that prompt injection reliably improves applicant rankings when r\'esum\'e quality is homogeneous and few candidates inject. However, its effectiveness rapidly diminishes as more candidates inject, collapsing when manipulation becomes widespread. When candidate quality is heterogeneous, prompt injection is less effective on average, but can occasionally allow lower-quality candidates to outrank higher-quality ones, raising fairness concerns. Overall, LLM-based screening is most vulnerable when manipulation is rare and candidate quality differences are small. Code and resources are publicly available at: https://github.com/preetb1199/Prompt_Injection_ACL26

13:00 JSTLLM/生成AIエージェント

言語モデルの結合が役立つのはどのような場合ですか? 67 のフロンティア モデルにわたるルーティング、投票、およびエージェントの混合における同時障害の上限

ルーティング、投票、カスケード、フュージョン、エージェントの混合などのマルチモデル LLM システムは、単一モデルの精度を上回るために使用されます。私たちは、彼らの利益が、現場でほとんど報告されない量によって制限されていることを示します。出力が 1 つのメンバー モデルの回答であるポリシーの場合、精度は 1 マイナス ベータを超えることはできません。ベータとは、同じクエリに対してすべてのモデルが誤る率です。対照的に、通常の診断である平均ペアワイズ誤差相関ρはベータを識別できません。同一の周辺値とペアワイズ相関を持つ誤差則は、全誤り率が異なる可能性があります。ベータ版の Clopper-Pearson バウンドは、ルーターをトレーニングする前に、ルーター、投票、またはカスケードが提供できる最大のゲインに基づいて有限サンプル証明書を提供します。 21 社のプロバイダーの 67 モデルにわたって、テトラコールで校正された単一因子モデルは依然として間違った尾部の価格を下回っています。オープンエンド数学では、観測されたベータ値は 0.052 であるのに対し、完全な 67 モデルのガウス コピュラでは 0.023 であり、約 2.5 倍の割安であり、90 パーセント CI は 1.7 ~ 3.4、k は 17 に相当します。効果は再発します。実行グレード コードの場合、ベータは 0.079 です。同じ GPQA-Diamond の質問を多肢選択形式ではなく自由回答形式で再質問すると、ベータ 0.127、カッパ 0.73 ~ 0.92 の 5 人の裁判官パネルで尾部が再び開き、主題ではなく回答形式での共失敗箇所が特定されます。同等の品質では、低 rho の異種アンサンブルが高 rho の Self-MoA を上回りますが、プール内のチェック可能なタスクでは、強力なクエリ レベルのルーティング シグナルがなければ、モデルを組み合わせた方が単一の最良のモデルを上回ることはほとんどありません。利益は、モデルを追加することでではなく、さまざまな質問でモデルが失敗することから得られます。

原文 (English)

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain is capped by a quantity the field rarely reports. For any policy whose output is one member model answer, accuracy cannot exceed one minus beta, where beta is the rate at which every model is wrong on the same query. In contrast, the usual diagnostic, average pairwise error correlation rho, cannot identify beta: error laws with identical marginals and pairwise correlations can have different all-wrong rates. A Clopper-Pearson bound on beta gives a finite-sample certificate on the largest gain any router, vote, or cascade could deliver before training a router. Across 67 models from 21 providers, a tetrachoric-calibrated single-factor model still underprices the all-wrong tail: on open-ended mathematics, observed beta is 0.052 versus 0.023 under the full 67-model Gaussian copula, about 2.5 times underpricing, with 90 percent CI 1.7 to 3.4 and k equals 17. The effect recurs on execution-graded code, where beta is 0.079. Re-asking the same GPQA-Diamond questions in free-response rather than multiple-choice form reopens the tail, with beta 0.127 and a five-judge panel with kappa 0.73 to 0.92, locating co-failure in answer format rather than subject. At matched quality, low-rho heterogeneous ensembles beat high-rho Self-MoA, but on checkable tasks in our pool, combining models rarely beats the single best model without a strong query-level routing signal. Gains come from models failing on different questions, not from adding more models.

13:00 JST研究/論文GPT / ChatGPT

高齢者の認知支援のための言語ベースのデジタルツイン

デジタルツインは、個人の行動と健康の軌跡のモデリングを可能にする、パーソナライズされたヘルスケアの有望なパラダイムとして浮上しています。認知の健康においては、言語や会話のパターンが非侵襲的なバイオマーカーとして機能する軽度認知障害 (MCI) の早期発見が依然として困難です。この研究では、大規模言語モデル (LLM) を活用して、スチロ測定の合図とコンテキスト メタデータを組み込むことで高齢者の会話行動を模倣する、言語ベースのデジタル ツイン フレームワークを提案します。忠実度と認知的一貫性を評価するために、再構成の品質を共同で測定し、認知スコアを予測するマルチヘッド条件変分オートエンコーダー (cVAE) を導入します。 I-CONECT データセットの実験では、デジタル ツインがアイデンティティ固有の特性を保持し、実際のデータに匹敵する再構成誤差と MoCA 予測誤差を達成しながら、ベースライン GPT で生成された応答を上回るパフォーマンスを示していることが示されています。これらの結果は、パーソナライズされた継続的な認知健康状態モニタリングのためのスケーラブルで非侵襲的なアプローチとしての言語ベースのデジタル ツインの可能性を浮き彫りにしています。

原文 (English)

Language-Based Digital Twins for Elderly Cognitive Assistance

Digital twins have emerged as a promising paradigm for personalized healthcare, enabling modeling of individual behavior and health trajectories. In cognitive health, early detection of Mild Cognitive Impairment (MCI) remains challenging, where language and conversational patterns serve as non-invasive biomarkers. In this work, we propose a language-based digital twin framework that leverages large language models (LLMs) to mimic the conversational behavior of elderly individuals by incorporating stylometric cues and contextual metadata. To evaluate fidelity and cognitive consistency, we introduce a multi-head conditional variational autoencoder (cVAE) that jointly measures reconstruction quality and predicts cognitive scores. Experiments on the I-CONECT dataset show that the digital twin preserves identity-specific characteristics and achieves reconstruction and MoCA prediction errors comparable to real data, while outperforming baseline GPT-generated responses. These results highlight the potential of language-based digital twins as a scalable and non-invasive approach for personalized and continuous cognitive health monitoring.

13:00 JSTLLM/生成AI研究/論文

グローバル AI 技術ガバナンスのためのオープンウェイト基盤モデルのベンチマーク

大規模言語モデル (LLM) は、国内および国際組織全体で人工知能 (AI) ガバナンス分析に導入されることが増えています。しかし、そのようなモデルでは、トレーニング データで過小評価されている国に対して、著しく精度の低い応答が生成されるという証拠が増えています。このパターンは、既存の文献で地理的偏りとして説明されています。この現象を調査している既存の研究には、その結果を損なう 3 つの方法論的な制限があります。(1) 重みが公表されていない独自のシステムに依存しているため、独立した複製が妨げられています。 (2) モデルトレーニングのためのデータ収集が終了した後、各モデルの知識の自然な限界に加えて地理的な無知につながる、数年間のモデル知識の評価。 (3) モデルの信頼できる製造 (HF) と不確実性の正直な認識を区別できない、粗い二値応答分類の使用。この研究では、2026 年 1 月に Harvard Dataverse で公開された 227 か国の 24,453 指標の検証済みグラウンドトゥルース データベースである Global AI Dataset v2 (GAID v2) に対して 4 つのオープンウェイト フロンティア言語モデルをベンチマークすることで、3 つの制限すべてに対処しています。IEEE IRAI 2026 フレームワークの 8 つのテーマの次元にマッピングされた合計 18 の指標が GAID v2 から選択され、およその結果が得られます。 6 つの評価年 (2010 年から 2023 年の期間内) にわたる 2,990 の国単位のメートル年の観測。モデル応答は、(a) 検証済み精度 (VA)、(b) HF、(c) 正直な拒否 (HR)、(d) 定性的ヘッジ (QH)、および (e) 誤った帰属 (MF) を区別する 5 つのカテゴリー スキームを使用して分類されます。精度の地理的差異は、混合効果ロジスティック回帰および差分差分 (DiD) 分析を通じて推定されます。

原文 (English)

Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations. There is, however, growing evidence that such models produce significantly less accurate responses for countries that are underrepresented in their training data-a pattern described in existing literature as geographic bias. Existing studies examining this phenomenon are subject to three methodological limitations that together undermine their findings: (1) reliance on proprietary systems whose weights are not publicly released, which prevents independent replication; (2) evaluation of model knowledge about years that fall after data collection for model training had concluded, leading to geographic ignorance in addition to the natural limits of each model's knowledge; and (3) use of coarse binary response classification that cannot distinguish models' confident fabrication (HF) from their honest acknowledgement of uncertainty. This study addresses all three limitations by benchmarking four open-weight frontier language models against the Global AI Dataset v2 (GAID v2), a verified ground-truth database of 24,453 indicators across 227 countries published on Harvard Dataverse in January 2026. A total of 18 indicators, mapped to the eight thematic dimensions of the IEEE IRAI 2026 framework, are selected from GAID v2, yielding approximately 2,990 country-metric-year observations across six evaluation years (within the period of 2010-2023). Model responses are classified using a five-category scheme that distinguishes (a) verified accuracy (VA), (b) HF, (c) honest refusal (HR), (d) qualitative hedging (QH), and (e) misattribution (MF). Geographic disparities in accuracy are estimated through mixed-effects logistic regression and difference-in-differences (DiD) analysis.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Llama

Know2Guess: 大規模言語モデルにおける知識境界評価のための汚染を認識したマルチゾーン ベンチマーク

大規模な言語モデルの信頼性の高い評価では、データの汚染、プロンプトの特異性、または一般的な拒否行動と混同することなく、サポートされている回答とサポートされていない推測を分離する必要があります。凍結されたビルドタイム ラベルの下で、回答可能な知識から棄権が期待される未知への移行を測定するための、汚染を認識したマルチゾーン ベンチマークを示します。このベンチマークには、5 つのドメインにわたる 1,200 項目、明示的な棄権期待、汚染リスクのメタデータ、および公式の厳密なパーサーと正規化された堅牢性パーサーによる二重解析が含まれています。ロックされた回答または棄権プロンプト、回答のみのコントロール、およびプロンプト テンプレートのバリアントの下で、FLAN-T5、Qwen2.5-Instruct、および Llama-3-Instruct モデルを評価します。このベンチマークは、一般的な無回答行動では解決されません。FLAN のベースラインは、生産的な棄権に関しては弱いままですが、より強力な指導調整モデルは、選択的ではあるが回答から棄権への移行が不完全であることを明らかにしています。 Qwen2.5-3B-Instruct は全体的に最高の信頼性を実現していますが、回答が期待されるゾーンは依然として難しく、キャリブレーションは依然として不十分で、良性の項目の拒否は引き続き発生します。プロンプトおよびパーサーの堅牢性分析により、主要なランキングと定性的な結論が維持されます。したがって、このベンチマークは、回答可能性、棄権、拒否、および汚染を、LLM の信頼性の個別だが相互作用する側面として監査するための再現可能なプロトコルを提供します。データセットは、https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark で公開されています。

原文 (English)

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.

13:00 JSTLLM/生成AILlama

役に立ちます: トレーニング中期の思いやりの値はトレーニング後にドメインに依存して低下します

標準的なポストトレーニング パイプラインは、教師あり微調整 (SFT) と強化学習 (RL) を適用して言語モデルを有用にしますが、これらのプロセスは、トレーニング前に注入された値を誤って低下させる可能性があります。動物危害ベンチマーク (AHB 2.2) と MORU ベンチマークで評価された SFT (Dolly-15k による有用性と Magicoder-110K によるコーディング) と GRPO (RLHFlow による有用性と Magicoder によるコーディング) の両方を使用して、思いやり指向の合成データで中間トレーニングされた Llama 3.1 8B モデルにおける動物の思いやりの値の保持に、トレーニング後のデータのドメインが差動的に影響を与えるかどうかを調査します。 (不確実性の下での道徳的推論)。有用性トレーニングは、AHB でのコーディング トレーニングと比較して動物の思いやりを大幅に低下させます (SFT: 35.7% 対 65.2%、GRPO: 18.7% 対 32.0%)。これは 2 つの独立した有用性データセットと 2 つのトレーニング パラダイムにわたって再現されています。英語のMORU項目では、有用性トレーニングは一般的な道徳的推論を25.5パーセントポイント(46.4%対71.9%)低下させ、その大きさは同情効果に匹敵する顕著な差でした。ただし、この効果は言語を越えて伝わりません。多言語の MORU ベンチマークでは、ドメイン効果は消失します (SFT: 52.3% 対 51.2%)。対照的に、動物の思いやりの効果は言語間で一貫して伝わり、Magiccoder の基本モデルに対する AHB パーセンテージ ポイントの増加は、英語以外の項目では英語の項目よりも 4.5 倍大きくなっています。この乖離は、トレーニング中に教え込まれた価値観が、ドメイン固有のトレーニング後の改善を推論するよりも深く、言語を超えてコード化されていることを示唆しています。これらの結果は、価値を満載したトレーニング途中で構築するラボの場合、トレーニング後の有用性よりも、トレーニング後のコーディング ドメインの方が、一般的な推論能力を損なうことなく、トレーニング途中の値をよりよく保存できる可能性があることを示唆しています。

原文 (English)

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, using both SFT (helpfulness via Dolly-15k vs. coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs. coding via Magicoder), evaluated on the Animal Harm Benchmark (AHB 2.2) and MORU benchmark (Moral Reasoning Under Uncertainty). Helpfulness training significantly degrades animal compassion relative to coding training on AHB (SFT: 35.7% vs. 65.2%; GRPO: 18.7% vs. 32.0%), replicating across two independent helpfulness datasets and two training paradigms. On English MORU items, helpfulness training degrades general moral reasoning by 25.5 percentage points (46.4% vs. 71.9%), a striking gap that rivals the compassion effect in magnitude. However, this effect does not transfer cross-lingually: on the multilingual MORU benchmark, the domain effect disappears (SFT: 52.3% vs. 51.2%). In contrast, the animal compassion effect transfers consistently across languages, with Magicoder's AHB percentage-point gain over the base model 4.5 times larger on non-English items than English items. This divergence suggests that values instilled through mid-training are encoded more deeply and cross-lingually than reasoning improvements from domain-specific post-training. These results suggest that, for labs building on value-laden mid-training, coding-domain post-training may better preserve mid-trained values than helpfulness post-training without harming general reasoning capabilities.

13:00 JSTLLM/生成AIGPT / ChatGPT

LLM の問題解決能力の調査 -- 静的質問に関する研究

大規模言語モデル (LLM) は、幅広い主題にわたる課題や試験を完了する実証済みの能力により、社会の多くの側面、特に教育に急速に影響を与えています。これまでの研究では LLM の教育的影響が調査されてきましたが、既存の研究の多くは公開またはオープンな問題データセットに依存しており、トピック固有の分析が不足しています。工学教育、特に機械工学における、特定の問題タイプに対する LLM パフォーマンスの系統的な調査は依然として限られています。 LLM ツールに教科書の質問を直接尋ねる従来の方法を使用する代わりに、私たちの研究ではモデル蒸留プロセスを採用して、静的問題を解決する際の LLM 機能を評価しました。 ChatGPT を抽出することにより、25 のテキストのみの静的質問を抽出し、さらに図を追加して数値を変更することで 2 つの追加のデータセットを構築しました。実験結果によると、LLM はテキストのみの静的問題では良好なパフォーマンスを示しますが、図が導入され、問題に複数のステップの推論が必要になると精度が低下します。さらなる分析によると、このパフォーマンスの低下は主に画像認識の制限が原因ではなく、むしろ複数ステップの推論と、抽出された視覚情報を連続したソリューション段階全体に一貫して適用することの難しさによって引き起こされていることが示唆されています。

原文 (English)

Investigating LLM's Problem Solving Capability -- a Study on Statics Questions

Large Language Models (LLMs) have rapidly influenced many aspects of society, particularly education, due to their demonstrated ability to complete assignments and examinations across a wide range of subjects. Although prior studies have examined the educational impact of LLMs, much of the existing work relies on public or open problem datasets and lacks topic-specific analysis. In engineering education, especially within mechanical engineering, systematic investigations of LLM performance on specific problem types remain limited. Instead of using traditional methods that directly ask textbook questions to an LLM tool, our study adopts a model distillation process to evaluate LLM capabilities in solving statics problems. By distilling ChatGPT, we extracted 25 text-only statics questions and further constructed two additional datasets by adding diagrams and modifying their numerical values. Experimental results show that while LLMs perform well on text-only statics problems, their accuracy decreases when diagrams are introduced and the problems require multi-step reasoning. Further analysis suggests that this performance drop is not primarily caused by limitations in image recognition, but rather by difficulties in multi-step reasoning and in consistently applying extracted visual information across successive solution stages.

13:00 JSTLLM/生成AILlama

主張し、説明しないでください: 動物福祉に関する LLM の推論を変える言語的特徴

動物愛護活動家たちは多くの著作物を作成しており、その著作物が言語モデルを訓練し、その後何百万人もの人々が動物福祉について尋ねるようになっています。提示された動物福祉ベンチマークで語彙を一致させたスタンスコントラストプローブを使用して、10の言語的特徴のそれぞれが、微調整データとして使用された場合にラマ-3.2-1Bの動物福祉推進推論に対する好みをどのように変化させるかを測定します。 10 個の特徴のうち 8 個で、統計的に有意な変化が生じます。 7 つは、断定的な確実性、明確な道徳的語彙、感情的な言葉、評価的主張、物語の構造、描写された危害の深刻度、即時的な時間的枠組みなど、モデルをより強力な動物愛護推進の推論に向けて移行させています。 2 つはそれを逆方向に動かします。ヘッジされた言葉と具体的な感覚的説明は両方とも動物愛護推進の立場を薄めます。一人称視点には統計的に有意な効果はありません。 LLM トレーニング コーパスに組み込まれる可能性のある動物福祉に関するテキストを執筆する人に対する実際的な推奨事項: シーンを中立的に説明するのではなく、立場を主張することです。モデルを変える特徴は、作家の立場を明確にするものです。それを弱める特徴は動物愛護の内容を保持しますが、スタンスを保留します。

原文 (English)

Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B's preference for pro-animal-welfare reasoning when used as fine-tuning data. Eight of the ten features produce statistically significant shifts. Seven move the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two move it the other way: hedged language and concrete sensory description both dilute the pro-animal-welfare stance. First-person perspective has no statistically significant effect. The practical recommendation for anyone writing animal-welfare text that may end up in LLM training corpora: assert a position rather than describe a scene neutrally. The features that shift the model are the ones that make the writer's position explicit; the features that dilute it hold animal-welfare content but withhold stance.

13:00 JSTLLM/生成AI

Long-Horizo​​n LLM 推論のためのコンテキストのリサイクル

大規模言語モデル (LLM) は、短いコンテキストの推論では強力な機能を示しますが、コンテキスト ウィンドウの制限と非効率なトークンの使用により、長い会話期間ではパフォーマンスが低下します。 ContextForge は、構造化クエリの生成、外部メモリの取得、制御された合成を組み合わせることにより、ターンをまたいでタスク関連情報を維持するコンテキスト リサイクル システムです。このシステムにより、完全なコンテキストの再生に依存せずに以前の計算を効率的に再利用できるため、応答の品質を維持しながらトークンのオーバーヘッドが削減されます。私たちは、構造化されたヘルスケア クエリ全体でマルチターン推論、後方参照、ドメイン シフトをテストする 15 ターンの会話ベンチマークを使用して ContextForge を評価します。同一の基礎モデルを使用するベースライン エージェントと比較して、ContextForge は同等の応答精度を維持しながら、一貫性の向上とトークン消費量の削減を実証します。これらの結果は、コンテキスト リサイクルが、より大きなコンテキスト ウィンドウやモデルの再トレーニングを必要とせずに、長期的なタスクで LLM 機能を拡張するための実用的なアプローチを提供することを示唆しています。コードと評価成果物は https://github.com/Betanu701/ContextForge で入手できます。

原文 (English)

Context Recycling for Long-Horizon LLM Inference

Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due to context window limitations and inefficient token usage. We introduce ContextForge, a system for context recycling that maintains task-relevant information across turns by combining structured query generation, external memory retrieval, and controlled synthesis. The system enables efficient reuse of prior computation without relying on full context replay, reducing token overhead while preserving answer quality. We evaluate ContextForge using a 15-turn conversational benchmark that tests multi-turn reasoning, back-references, and domain shifts across structured healthcare queries. Compared to a baseline agent using identical underlying models, ContextForge demonstrates improved consistency and reduced token consumption, while maintaining comparable response accuracy. These results suggest that context recycling provides a practical approach for extending LLM capabilities in long-horizon tasks without requiring larger context windows or model retraining. Code and evaluation artifacts are available at https://github.com/Betanu701/ContextForge.

13:00 JSTLLM/生成AI

非暴力的なコミュニケーション制約を伴う大規模言語モデル対話における会話のエスカレーションを軽減する

大規模言語モデル (LLM) は、対人関係の対立、フラストレーション、苦痛を伴う感情的に負荷の高い状況で使用されることが増えています。これまでの安全性研究は、有害なコンテンツやポリシー違反のコンテンツなどの明白な危害を防止することに焦点を当ててきましたが、意図せず紛争をエスカレートさせる可能性のある会話行為についてはあまり注目されていませんでした。この論文では、非暴力コミュニケーション (NVC) から派生した軽量のプロンプトレベルの制約を通じて、LLM をよりエスカレーションのない対話行動に導くことができるかどうかを調査します。私たちは NVC の原則をプロセス指向のガイドラインとして再定式化し、責任の帰属を防ぎ、ユーザーの感情的経験への注意を強調し、アドバイスの前に明確にすることを奨励します。複数の命令調整モデルとユーザーの抵抗レベルにわたるデュアル エージェント シミュレーション フレームワークを使用して、NVC 制約のプロンプトが一貫して会話のエスカレーションを軽減し、抵抗の高いユーザーとの対話を安定させることを示します。これらの結果は、単純なコミュニケーション制約によって、衝突が起こりやすい環境における LLM 対話の信頼性を大幅に向上させることができることを示唆しています。

原文 (English)

Reducing Conversational Escalation in Large Language Model Dialogue with Nonviolent Communication Constraints

Large language models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and distress. While prior safety research has focused on preventing explicit harms such as toxic or policy-violating content, less attention has been paid to conversational behaviors that may unintentionally escalate conflict. In this paper, we investigate whether LLMs can be guided toward more de-escalating dialogue behavior through lightweight prompt-level constraints derived from Nonviolent Communication (NVC). We reformulate NVC principles as process-oriented guidelines that discourage blame attribution, emphasize attention to users' emotional experiences, and encourage clarification before advice. Using a dual-agent simulation framework across multiple instruction-tuned models and user resistance levels, we show that NVC-constrained prompting consistently reduces conversational escalation and stabilizes interactions with highly resistant users. These results suggest that simple communication constraints can meaningfully improve the trustworthiness of LLM dialogue in conflict-prone settings.

13:00 JSTLLM/生成AI

ネパール語の話し言葉を感情に応じた手話アバターに低リソースでマルチモーダル翻訳

感情表現を統合した手話コミュニケーションシステムは、特にリソースの少ない言語では未開発のままです。このパイロット研究では、音声入力から感情条件付けされたネパール手話アバターを生成する実現可能性を実証する概念実証マルチモーダル フレームワークである NEST-V1 (Nepali Emotion and Speech Transformer - バージョン 1) を紹介します。予備調査として、核となる技術的アプローチを検証するために、3 つの感情状態 (幸せ、中立、悲しい) にわたる 4 つの一般的なネパール語 (「ありがとう」、「こんにちは」、「家」、「私」) に焦点を当てます。当社の軽量アーキテクチャは、自動音声認識と感情分類を同時に行うための共有音響エンコーダを採用しており、50 人の話者からの 600 個のラベル付き音声サンプルのデータセットで 81.1% の ASR 精度と 79.21% の感情認識精度を達成しています。このシステムは、エッジ展開に適した 2,210 万パラメータのみという軽量のフットプリントを維持しながら、個別のモデル アーキテクチャと比較して 37% のパラメータ効率を実証します。このパイロット作業は、リソースが少ない環境で感情を認識した手話翻訳の技術的基盤を確立し、より大きな語彙とより多様な感情表現への将来の拡張のためのスケーラブルなフレームワークを提供します。私たちの予備的な結果は、聴覚障害のあるコミュニティのためのリアルタイムの感情表現豊かな手話コミュニケーション システムの実現可能性を示しており、その後の開発段階での強化のための明確な道筋が示されています。

原文 (English)

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars

Sign language communication systems, that integrate emotional expression remain underexplored, particularly for low-resource languages. This pilot study presents NEST-V1 (Nepali Emotion and Speech Transformer - Version 1), a proof-of-concept multimodal framework that demonstrates the feasibility of generating emotion-conditioned Nepali Sign Language avatars from spoken input. As a preliminary investigation, we focus on four common Nepali words ("thank you", "hello", "house", "me") across three emotional states (happy, neutral, sad) to validate our core technical approach. Our lightweight architecture employs a shared acoustic encoder for simultaneous Automatic Speech Recognition and emotion classification, achieving 81.1% ASR accuracy and 79.21% emotion recognition accuracy on a dataset of 600 labeled audio samples from 50 speakers. The system demonstrates 37% parameter efficiency compared to separate model architectures while maintaining a lightweight footprint with only 22.1M parameters suitable for edge deployment. This pilot work establishes the technical foundation for emotion-aware sign language translation in low-resource settings and provides a scalable framework for future expansion to larger vocabularies and more diverse emotional expressions. Our preliminary results indicate the viability of real-time, emotionally expressive sign language communication systems for the hearing-impaired community, with clear pathways for enhancement in subsequent development phases.

13:00 JSTLLM/生成AI規制/政策AnthropicGoogleGemini

生成 AI と著作権侵害: 17 歳未満の AI 音楽生成システムの法技術的分析タイトル17

生成人工知能 (GenAI) により、ユーザーは、著作権で保護された歌詞、AI が作曲したメロディー、本物のアーティストを模倣した合成ボーカルを組み合わせて、テキスト プロンプトを使用して音楽を合成できるようになりました。この論文では、米国著作権法に基づく AI ベースの音楽作成 (Google Gemini の音楽ツールなど) の法的および技術的側面を検討します。私たちは、あるアーティストの保護された歌詞を GenAI システムに入力し、別のアーティストの声やスタイルを使用するように指示し、その結果得られた曲を公開して収益化するユーザーが、17 U.S.C. に違反するかどうかを分析します。第 106 条の独占的権利 [3]。この分析には、タイトル 17 の原則 (複製の権利、二次的著作物、配布)、17 U.S.C. が統合されています。セクション 114 の狭い録音保護 [4]、および州レベルで新たに制定された音声クローン法 [20]。私たちは、歌詞の無断コピーは楽曲侵害の高いリスクをもたらす一方、単なる AI 生成の音声模倣は通常、連邦録音保護の対象外となり、代わりに州のパブリシティ権に関与すると主張します [12]、[13]。最近の訴訟と法律 (コンコード対アンスロピック [10]、カドリー対メタ [11]、レーマン対ロヴォ [12]、テネシー州の「ELVIS 法」 [20]、UMG 対アンチャーテッド ラボ [14] など) がこの分裂を例証しています。私たちは AI の技術コンポーネント (プロンプト エンコーディング、潜在拡散、ニューラル ボコーダー、スピーカーの埋め込み) を法的リスクにマッピングし、規制上のギャップを特定します。連邦法は歌詞とメロディーを強力に保護していますが、現在、合成されたボーカルの類似性に対する救済策は限定的です [22]、[23]。この論文は、AI による音楽作成に関するより明確なルールを求める政策提案で締めくくられています。

原文 (English)

Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17

Generative artificial intelligence (GenAI) has enabled users to synthesize music with text prompts, combining copyrighted lyrics, AI-composed melodies, and synthetic vocals that imitate real artists. This paper examines the legal and technical dimensions of AI-based music creation (e.g., Google Gemini's music tools) under U.S. copyright law. We analyze whether a user who inputs one artist's protected lyrics into a GenAI system, directs it to use another artist's voice or style, publishes the resulting song, and monetizes it violates 17 U.S.C. Section 106's exclusive rights [3]. The analysis integrates Title 17 doctrine (rights of reproduction, derivative works, distribution), 17 U.S.C. Section 114's narrow sound recording protection [4], and the new voice-cloning laws emerging at the state level [20]. We argue that unauthorized lyric copying poses a high risk of infringement of the musical composition, whereas mere AI-generated voice imitation typically falls outside federal sound recording protection and instead implicates state publicity rights [12], [13]. Recent cases and legislation (Concord v. Anthropic [10]; Kadrey v. Meta [11]; Lehrman v. Lovo [12]; Tennessee's "ELVIS Act" [20]; UMG v. Uncharted Labs [14]; etc.) illustrate this split. We map AI technical components (prompt encoding, latent diffusion, neural vocoders, speaker embeddings) to legal risks and identify a regulatory gap: federal law robustly protects lyrics and melody but currently provides limited remedies for synthesized vocal likeness [22], [23]. The paper concludes with policy suggestions for clearer rules on AI music creation.

13:00 JSTLLM/生成AI

レキシコンから AI へ: 低リソース言語の特殊な会話システムのための構造化データ パイプライン

リソースの少ない言語は、大規模なトレーニング コーパスにアクセスせずに特殊な会話システムを作成するという、AI 開発における重大な課題に直面しています。私たちは、構造化された言語リソースを特化した AI システムに変換する体系的な方法論を提示し、専門家が厳選した語彙データベースが会話型 AI 開発の効果的な基盤として機能できることを実証します。私たちのアプローチは、ヒンディー語 WordNet を 125 万の多様な命令と応答のペアに変換し、4 ビット量子化を備えたリソース効率の高い LoRA を使用して 12B パラメーターの言語モデルを微調整します。ヒンディー語学習チャットボットによる評価では、構造化知識ベースのシステムが優れた教育効果 (汎用モデルの場合は 91.0 対 79.4 ~ 83.6) を達成しながら、競争力のあるセマンティック パフォーマンスと優れた一貫性を維持していることが実証されました。完全なパイプラインは、WordNet リソースを使用してあらゆる言語に特化した AI システムを開発するための、ヒンディー語を使用した概念実証の方法論を示しています。この取り組みは、リソースの少ない言語における AI アクセシビリティの重大なギャップに対処し、コーパス集約型のアプローチに代わる実用的な代替手段を提供し、既存の WordNet リソースを使用して数百の言語に特化した AI 開発を可能にする可能性があります。

原文 (English)

From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora. We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical databases can serve as effective foundations for conversational AI development. Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs, fine-tunes a 12B-parameter language model using resource-efficient LoRA with 4-bit quantization. Evaluation through a Hindi language learning chatbot demonstrates that structured-knowledge-based systems achieve superior pedagogical effectiveness (91.0 vs. 79.4-83.6 for general-purpose models) while maintaining competitive semantic performance and exceptional consistency. The complete pipeline demonstrates a proof-of-concept methodology using Hindi for developing specialized AI systems for any languages with WordNet resources. This work addresses the critical gap in AI accessibility for low-resource languages, offering a practical alternative to corpus-intensive approaches and potentially enabling specialized AI development for the hundreds of languages with existing WordNet resources.

13:00 JST研究/論文OpenAISora

夢のマシン -- 次の創造的な経済

私たちは、政策文書、業界データ、クリエイター調査、プラットフォーム分析にわたる 374 の一次情報源を利用して、生成人工知能の下でのクリエイティブ産業の構造変革を調査します。分水嶺イベントとしての OpenAI の Sora ビデオ モデルの 2024 年 12 月のリリースから始めて、技術的破壊に対するクリエイティブな抵抗の歴史的パターンを追跡し、その後、クリエイティブな作業における人間とマシンのコラボレーションのスペクトルをマッピングするための分析フレームワークである人間と AI エージェンシーの連続体を開発します。我々は、アップロードの 44% を占めるにもかかわらず、AI 生成コンテンツをプラットフォーム ストリームの約 1 ~ 3% に制限する、視聴者が課す品質のしきい値である「傾斜天井」の証拠を示します。 AIと著作権に関する英国政府の2025年の協議(11,500件を超える回答、88%がAIトレーニングの権利拡大に反対)を分析すると、テクノロジー企業とクリエイティブ労働者の間に深い構造的緊張があることが明らかになった。私たちは、ディズニーの 10 億ドルの OpenAI 投資から Netflix の AI ネイティブ アニメーション部門に至るまで、主要スタジオが AI で強化された制作パイプラインをどのように位置付けているかを調査します。この研究では、クリエイティブなサプライチェーンにおける調整の崩壊、迅速なエンジニアや AI オーケストレーターなどの新しい専門的役割の出現を取り上げ、移行を乗り切るための 4 つの原則 (透明性、同意、報酬、人間中心の設計) を提案しています。 8 つの付録では、定量分析、用語集、トピック参考文献を提供し、シャドウ AI の導入、AI の偏見、アルゴリズムの意図について詳しく説明しています。

原文 (English)

Dream machine -- the next creative economy

We examine the structural transformation of creative industries under generative artificial intelligence, drawing on 374 primary sources spanning policy documents, industry data, creator surveys, and platform analytics. Beginning with the December 2024 release of OpenAI's Sora video model as a watershed event, we trace the historical pattern of creative resistance to technological disruption, then develop an analytical framework -- the Human-AI Agency Continuum for mapping the spectrum of human and machine collaboration in creative work. We present evidence for the "slop ceiling," an audience-imposed quality threshold that constrains AI-generated content to approximately 1--3% of platform streams despite comprising 44% of uploads. Analysis of the UK Government's 2025 consultation on AI and copyright (over 11,500 responses, 88% opposing expanded AI training rights) reveals deep structural tensions between technology firms and creative workers. We investigate how major studios, from Disney's $1 billion OpenAI investment to Netflix's AI-native animation unit, are positioning for an AI-augmented production pipeline. The work covers coordination collapse in creative supply chains, the emergence of new professional roles such as prompt engineers and AI orchestrators, and proposes four principles for navigating the transition: transparency, consent, compensation, and human-centred design. Eight appendices provide quantitative analysis, a glossary, topical bibliography, and deep dives into shadow AI adoption, AI stigma, and algorithmic intent.

13:00 JST研究/論文

情報ランドスケープ分析のための多層 AI フレームワーク

この論文は、情報障害の文脈における情報ランドスケープ分析のための多層 AI フレームワークを提案します。このフレームワークは、誤情報の検出をバイナリの事実確認タスクとして扱うのではなく、情報源の信頼性、事実の構造、枠組み、偏見、感情の活性化、操作パターン、伝播力学などの多面にわたって政治およびメディアのコンテンツを分析します。目標は、個別のクレーム検証を超えて、イベント、エンティティ、または物語を取り巻く情報環境の構造化された表現に移行することです。私たちは、メディア分析用の AI システムは認識論的マッピング、つまり事実、解釈、俳優、物語が時間の経過とともにどのように相互作用するかについての透明で多次元の説明をサポートする必要があると主張します。この論文は、情報障害研究のためのより微妙で説明可能で非常に有用なツールをサポートすることを目的として、フレームワークの概念アーキテクチャ、分析層、および方法論的根拠を示しています。

原文 (English)

A Multi-Layer AI Framework for Information Landscape Analysis

This paper proposes a multi-layer AI framework for information landscape analysis in the context of information disorder. Rather than treating misinformation detection as a binary fact-checking task, the framework analyzes political and media content across multiple dimensions, including source reliability, factual structure, framing, bias, emotional activation, manipulation patterns, and propagation dynamics. The goal is to move beyond isolated claim verification toward a structured representation of the informational environment surrounding an event, entity, or narrative. We argue that AI systems for media analysis should support epistemic mapping: a transparent, multi-dimensional account of how facts, interpretations, actors, and narratives interact over time. The paper presents the conceptual architecture, analytical layers, and methodological rationale of the framework, with the aim of supporting more nuanced, explainable, and critically useful tools for information disorder research.

13:00 JSTLLM/生成AIAnthropicClaudeOpenAIGPT / ChatGPT

発散的な推奨事項、収束した診断: AI 商用推奨事項におけるプロバイダー間の障害モードの収束

製品の推奨に ChatGPT と Claude の両方を使用している顧客を持つブランドは、戦略的な選択に直面しています。単一の最適化ハンドブックを使用するか、それともプロバイダーごとに 1 つ使用するかです。 4 つの測定バッチで商用フレーム化された 215 個のプロンプト全体で、両プロバイダーは推奨するブランドについて約 3 分の 2 の割合で意見が一致していません (プロバイダー間の推奨値 Jaccard 0.35、同じプロンプトの再実行ベースラインの 0.50 ~ 0.61 を下回っています)。ピックは分岐します。しかし、どちらのプロバイダーもブランドを推奨していない場合、その障害を 3 つのモード (ブランドがモデルに到達しない)、説得力 (モデルに到達するが言及されない)、ポジショニング (言及されているが推奨されない) の 3 つのモードのいずれかに分類します。7,763 件のそのような共同障害では、両方のプロバイダーが 95.1% の確率で同じ障害モードを診断します (クラスター化 95% CI [94.3%、 95.7%])。ブランドの知名度が低下するにつれて一致度は単調に増加し、カテゴリーリーダーの 81% [78.2%、84.0%] からロングテールの地域ブランドの 99.6% [99.3%、99.9%] まで増加します。 2 つのプロバイダーは、かなり異なる生成ルート (Anthropic は 43 ~ 52% の確率で事前提案から推奨し、OpenAI は 8 ~ 29%) によって選択に到達しますが、ロングテールにとって最も重要な障害診断に収束します。診断された障害モードに対処する作業により、両方のプロバイダーの可視性が高まります。カテゴリリーダーのポジショニングとコンテンツレベルの作業は、よりプロバイダー固有です。

原文 (English)

Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation

A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, or one per provider? Across 215 commercially-framed prompts in four measurement batches, the two providers disagree on which brands they recommend roughly two-thirds of the time (cross-provider recommendation Jaccard 0.35, below the 0.50-0.61 same-prompt rerun baseline). The picks diverge. But when neither provider recommends a brand, we classify the failure into one of three modes -- discoverability (the brand never reaches the model), compellingness (it reaches the model but isn't mentioned), or positioning (it's mentioned but not recommended) -- and on 7,763 such joint failures, both providers diagnose the same failure mode 95.1% of the time (clustered 95% CI [94.3%, 95.7%]). Agreement rises monotonically with falling brand prominence, from 81% [78.2%, 84.0%] on category leaders to 99.6% [99.3%, 99.9%] on long-tail regional brands. The two providers reach their picks by measurably different generative routes -- Anthropic recommends from priors 43-52% of the time, OpenAI 8-29% -- but they converge on the failure diagnosis where it matters most for the long tail. Work that addresses the diagnosed failure mode lifts visibility on both providers; positioning - and content-level work for category leaders is more provider-specific.

13:00 JST研究/論文

ガバナンス逆転仮説: なぜ AI 規制が強化されると組織制御が低下するのか

この論文では、人工知能 (AI) ガバナンスにおける増大するパラドックスを説明するために、ガバナンス逆転仮説 (GIH) を紹介します。つまり、規制の拡大と技術の複雑さが増大する状況下では、組織はより正式に統治されるようになると同時に、AI システムに対する運用管理の低下を経験する可能性があります。既存の AI ガバナンス フレームワークは一般に、規制を強化することで説明責任、監視、組織制御が向上すると想定しています。この論文は、ガバナンスの形式化自体が AI 集約型環境における制御の侵食に寄与する可能性があると主張することで、その仮定に異議を唱えます。この論文は、制度理論、組織ガバナンスの研究、説明責任に関する研究、および新たな AI ガバナンスの文献に基づいて、規制の拡大が、権限の断片化、象徴的なガバナンスの拡張、制御の外部化、および権限の麻痺という 4 つの相互に関連したメカニズムを通じて、運営上の権限をどのように弱体化させる可能性があるかを説明する概念的な枠組みを開発しています。ガバナンス システムがますます多層化され、手続きが高密度になるにつれて、組織は、不透明で外部が仲介する AI インフラストラクチャに対する一貫した権限、技術的な可視性、エスカ​​レーション機能、意味のある介入権限を維持するのに苦労する可能性があります。この論文は、ガバナンスの拡大が運営上の一貫性を強化するのではなく、積極的に損なう可能性がある構造的条件としてガバナンスの反転を導入することにより、制度的デカップリング理論を拡張しています。 AI ガバナンスにおける中心的なリスクは、ガバナンス構造の欠如ではなく、効果的に統治する能力を徐々に失いながらも、ますます統治されているように見える機関の出現である可能性があると結論付けています。

原文 (English)

The Governance Inversion Hypothesis: Why More AI Regulation May Produce Less Organisational Control

This paper introduces the Governance Inversion Hypothesis (GIH) to explain a growing paradox in artificial intelligence (AI) governance: under conditions of increasing regulatory expansion and technological complexity, organisations may become more formally governed while simultaneously experiencing a decline in operational control over AI systems. Existing AI governance frameworks generally assume that stronger regulation improves accountability, oversight, and organisational control. This paper challenges that assumption by arguing that governance formalisation itself may contribute to the erosion of control in AI-intensive environments. Drawing on institutional theory, organisational governance research, accountability scholarship, and emerging AI governance literature, the paper develops a conceptual framework explaining how regulatory expansion may weaken operational authority through four interconnected mechanisms: authority fragmentation, symbolic governance expansion, externalisation of control, and authority paralysis. As governance systems become increasingly layered and procedurally dense, organisations may struggle to maintain coherent authority, technical visibility, escalation capability, and meaningful intervention power over opaque and externally mediated AI infrastructures. The paper extends institutional decoupling theory by introducing governance inversion as a structural condition in which governance expansion may actively undermine operational coherence rather than strengthen it. It concludes that the central risk in AI governance may not be the absence of governance structures, but the emergence of institutions that appear increasingly governed while progressively losing the capacity to govern effectively.

13:00 JST研究/論文OpenAI

AI の導入と能力に関するオープンソースの経済指標

私たちは、AI の導入と、さまざまな職種にわたる個別の労働タスクを実行する AI の能力の両方を測定することに取り組んでいます。導入率を測定するために、私たちは公開されているユーザー LLM チャット データと O*NET タスクを使用してフロンティア AI ラボによって作成された研究を再現するオープンソースの経済指標を開発しました。その結果、金融、コンピューター サイエンス、および芸術分野の職業が最も高い導入率を示していることがわかりました。機能を測定するために、O*NET の職業、タスク、モデル コンテキスト プロトコル (MCP) サーバーに基づいたベンチマーク シナリオを生成するシステムを構築します。私たちは、インデックスに頻繁に現れる 9 つの職業にわたるシナリオで OpenAI エージェント SDK ハーネスを使用して Kim-k2.5 をテストし、AI は高レベルのワークフローを正しく実行しますが、詳細な詳細 (使用される特定のツール呼び出しなど) でエラーが発生することが多いことがわかりました。

原文 (English)

The Open Source Economic Index of AI Adoption and Capability

We work towards measuring both AI adoption and the capability of AI to perform discrete labor tasks across various occupations. To measure adoption, we develop an open-source economic index that uses publicly available user-LLM chat data and O*NET tasks to replicate studies produced by frontier AI labs, finding that occupations in the finance, computer science, and arts sectors are those with the highest adoption rates. To measure capabilities, we build a system that generates benchmark scenarios grounded in O*NET occupations, tasks, and model-context-protocol (MCP) servers. We test Kimi-k2.5 with an OpenAI agents SDK harness on scenarios across 9 occupations that appear frequently in our index, finding that AI correctly executes high-level workflows but often errs in the granular details (such as specific tool calls used).

13:00 JST画像/動画生成

Dot-Flik: 分散型昆虫監視のためのスケーラブルなエッジ AI アーキテクチャ

世界的な昆虫個体数の減少により、スケーラブルで継続的な監視システムが必要となっていますが、既存のビジョンベースのソリューションは、ハードウェアのコスト、エネルギー需要、集中処理やクラウド接続への依存などによって依然として制約を受けています。この記事では、これらの制限に対処するための 3 つの貢献を紹介します。まず、時間差分、ガンマ補正された動きの増幅、およびブロックベースの動き密度分析に基づいた動き情報に基づいたフレーム フィルタリング アルゴリズムを提案します。このアルゴリズムは、センシング デバイスでの深層学習推論を必要とせず、昆虫の活動を維持しながらエッジで無関係なフレームを破棄します。 2 番目に、このエッジレベルの前処理を通じてデータ取得を AI 分類から分離する分散型の階層型 IoT アーキテクチャを導入し、中央処理要件の部分的なスケーリングを予測し、モノリシックな単一ストリームのアプローチと比較して監視範囲を大幅に拡大します。 3 番目に、リアルタイム パフォーマンス、ネットワーク スケーラビリティ、ハードウェア コスト、さまざまな風況下でのエネルギー効率の 4 つの軸に沿って、低コストの汎用ハードウェアを屋外に実際に導入して完全なシステムを検証します。結果は、微風条件下で 60 ~ 80% のフレーム削減、12.8 ミリ秒の計算ヘッドルームによるリアルタイム 30 FPS 動作の持続、最大 22.6% のエネルギー節約、および中央ノードあたり 5 ~ 6 の同時エッジ ストリームのサポートを実証しています。これらの発見は、都市環境における高密度で低コストの生物多様性監視ネットワークの実用的な基盤を確立します。

原文 (English)

Dot-Flik: A Scalable Edge AI Architecture for Distributed Insect Monitoring

Global insect population declines necessitate scalable, continuous monitoring systems, yet existing vision-based solutions remain constrained by high hardware costs, energy demands, and reliance on centralized processing or cloud connectivity. This article presents three contributions to address these limitations. First, we propose a motion-informed frame filtering algorithm based on temporal differencing, gamma-corrected motion amplification, and block-based motion density analysis that discards irrelevant frames at the edge while preserving insect activity, without requiring deep learning inference on the sensing device. Second, we introduce a distributed, hierarchical IoT architecture that decouples data acquisition from AI classification through this edge-level preprocessing, projecting fractional scaling of central processing requirements and significantly increasing monitoring coverage compared to monolithic single-stream approaches. Third, we validate the complete system through real-world outdoor deployments on low-cost commodity hardware along four axes: real-time performance, network scalability, hardware cost, and energy efficiency under varying wind conditions. Results demonstrate 60-80% frame reduction under light-wind conditions, sustained real-time 30 FPS operation with 12.8 ms of computational headroom, up to 22.6% energy savings, and support for 5-6 concurrent edge streams per central node. These findings establish a practical foundation for dense, low-cost biodiversity monitoring networks in urban environments.

13:00 JSTエージェント

6G SD-RAN でのダイナミック VR スライス管理のためのプライバシーを意識したエージェントのコラボレーション

6G ネットワークの仮想現実 (VR) サービスには超低遅延と高スループットが必要ですが、これはソフトウェア無線アクセス ネットワーク (SD-RAN) の動的リソース管理にとって重大な課題となります。この研究では、VR スライス管理のためのモビリティ主導型でプライバシーを意識したマルチエージェント強化学習 (MARL) フレームワークを提案します。このフレームワークでは、協力的なエージェントがユーザー データのプライバシーを保護しながら、エンドツーエンド VR リンク上のリソース分散を最大化します。当社のアプローチにはモビリティ予測と情報ボトルネック エンコーダーが組み込まれており、効果的かつ安全なエージェントのコラボレーションを促進します。シミュレーションでは、従来の方法との比較が研究されており、最大 34\% のスループット向上、28\% のリソース削減、85\% のプライバシー漏洩の削減が示されており、将来の 6G 環境で信頼できる没入型 VR エクスペリエンスが保証されます。

原文 (English)

Privacy-Aware Agent Collaboration for Dynamic VR Slice Management in 6G SD-RAN

Ultra-low latency and high throughput are required for Virtual Reality (VR) services in 6G networks, which presents critical challenges for Software-Defined Radio Access Networks (SD-RANs) dynamic resource management. This work propose a mobility-driven, privacy-aware Multi-Agent Reinforcement Learning (MARL) framework for VR slice management, in which cooperative agents maximize resource distribution over end-to-end VR links while protecting the privacy of user data. Our approach incorporates mobility prediction and an information bottleneck encoder to facilitate effective and secure agent collaboration. In simulations, comparisons with traditional methods are studied which show up to 34\% throughput improvement, 28\% fewer resources, and 85\% less privacy leakage, guaranteeing dependable immersive VR experiences in future 6G environments.

13:00 JST研究/論文

フェデレーション エッジ ネットワーク向けの幾何学的公平性を意識したルーティング

新興の 6G およびエッジ インテリジェント ネットワークには、空間的に分散されたさまざまなデバイス間で効果的でバランスの取れたルーティング アルゴリズムが必要です。既存のフェデレーテッド ルーティング システムは、公平性やネットワーク トポロジの基礎となる幾何学的構造よりも、総遅延やスループットを優先することがよくあります。このペーパーでは、双曲グラフ ニューラル ネットワーク (HGNN) とフェデレーテッド最適化を組み合わせてエッジ ノード全体で同等のパフォーマンスを提供する、幾何学的公平性を意識したルーティング システムである Geo-FairFed について説明します。各ノードは、階層関係や接続の非対称性を含む、負に湾曲した多様体上のトポロジーを意識した表現を学習します。次に、グローバル アグリゲータは、ルーティング損失、幾何学的不一致、および Jain の公平性インデックスに基づく不平等ペナルティを最小限に抑える曲率正規化目標を使用して公平性を強制します。理論的分析により、制限された曲率の下での収束保証が開発され、提案された公平性項により配線パフォーマンスのパレート改善均衡がもたらされることが示されました。動的な 6G エッジおよび IoT トポロジに関する広範なシミュレーションにより、Geo-FairFed は、最先端のフェデレーテッド ルーティング プロトコルおよびジオメトリック ルーティング プロトコルと比較して、平均遅延を 20\% 最小限に抑え、エネルギー消費を 17\% 削減し、公平性を最大 21\% 向上させることが明らかになりました。この研究では、双曲線多様体にトポロジを埋め込み、フェデレーテッド アップデートに公平性を組み込むことで、大規模ネットワーク ルーティングの効率と公平性を大幅に向上できることがわかりました。

原文 (English)

Geometric Fairness-Aware Routing for Federated Edge Networks

Emerging 6G and edge-intelligent networks require effective and balanced routing algorithms among varied and spatially distributed devices. Existing federated routing systems often prioritize aggregate latency or throughput above fairness and the underlying geometric structure of network topologies. This paper describes Geo-FairFed, a geometric fairness-aware routing system that blends hyperbolic graph neural networks (HGNNs) and federated optimization to provide equal performance across edge nodes. Each node learns topology-aware representations on a negatively curved manifold, which include hierarchical relationships and connection asymmetries. A global aggregator next enforces fairness using a curvature-regularized aim that minimizes routing loss, geometric inconsistency, and an inequality penalty based on Jain's fairness index. A theoretical analysis develops convergence guarantees under limited curvature and shows that the proposed fairness term results in a Pareto-improving equilibrium in routing performance. Extensive simulations on dynamic 6G-edge and IoT topologies reveal that Geo-FairFed minimizes average latency by 20\%, reduces energy consumption by 17\%, and improves fairness by up to 21\% when compared to state-of-the-art federated and geometric routing protocols. The study found that embedding topology in a hyperbolic manifold and including fairness into federated updates can significantly enhance the efficiency and equity of large-scale network routing.

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPTGeminiDeepSeek

科学者のように考えますか? LLM によって生成された調査手法の構造的研究

大規模言語モデル (LLM) は、研究方法論を導くためにますます使用されていますが、最小限のプロンプトの下でのデフォルトの方法論的傾向は依然として不明瞭です。ここでは、GPT-5.1、Gemini 3 Pro、および DeepSeek-V3.2 に対して、1,000 件の最近の arXiv コンピューター サイエンス論文から LLM で抽出されたリサーチ質問を入力し、結果として得られる方法論の提案を論文由来の実験目録と比較します。私たちはリサーチの質問のみを提供しているため、測定した差異は最初の提案を反映しており、その提案がどれほど最適であるかは反映されていません。両方のソースから構造化メソッドの特徴を抽出し、それらを共有分類にマッピングし、モデル プロバイダー、データセット タスク タイプ、評価指標タイプを含む複数の分類の次元にわたる相違を定量化します。プロバイダーの選択には最も不均衡が見られ、Jensen-Shannon の相違は他の分類次元よりも約 3 ~ 5 倍大きくなります。その他/学術的な単一出現モデルは 23 ~ 24 パーセント ポイント過小評価されていますが、再利用された学術/コミュニティ モデルはわずかに過大評価されています (4 ~ 6pp)。また、LLM は、全体としてより狭い範囲の方法を提案しています。つまり、モデル エンティティ コントラクトの有効数は 1,232 から 59 ~ 96 であり、LLM 間のランク相関 (0.55 ~ 0.68) は、通常、LLM と論文間の相関 (0.33 ~ 0.56) を超えているため、歪みはモデル間でほぼ共有されます。人気ベースライン、BM25 検索キャリブレーション、および紙レベルの類似性テストにより、出力はクエリ固有の応答であるが、より狭いオプション セットでフィルタリングされていることを確認します。したがって、クロスチェックを行わずに LLM の提案に依存する研究者は、方法論的な検索範囲がより集中したデフォルトに向かって狭まってしまう危険性があります。

原文 (English)

Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods

Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we prompt GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 with an LLM-extracted research question from each of 1,000 recent arXiv computer-science papers and compare the resulting methodology suggestions against a paper-derived experimental inventory. Since we provide only the research question, the differences we measure reflect initial suggestions and not how optimal those suggestions are. We extract structured method features from both sources, map them into a shared taxonomy, and quantify divergence across multiple taxonomy dimensions including model provider, dataset task type, and evaluation metric type. The strongest imbalance appears in provider choice, with Jensen-Shannon divergence about 3-5x larger than any other taxonomy dimension. Other/Academic single-occurrence models are underrepresented by 23-24 percentage points, while reused academic/community models are slightly overrepresented (4-6pp). LLMs also suggest a much narrower range of methods overall: the effective number of model entities contracts from 1,232 to 59-96, and inter-LLM rank correlations (0.55-0.68) generally exceed LLM-to-paper correlations (0.33-0.56), so the distortions are largely shared across models. Popularity baselines, BM25 retrieval calibration, and paper-level similarity tests confirm that the outputs are query-specific responses, but filtered through a narrower set of options. Researchers who rely on LLM suggestions without cross-checking therefore risk narrowing their methodological search space toward a more concentrated default.

13:00 JST研究/論文

マルチスケールの離脱と参加のダイナミクス: 戦術的合意と戦略的連携の形成

この論文では、戦略的な離脱と参加の決定が連合内の戦術的な合意力学と結び付けられる、連合形成のマルチスケール モデルを開発します。連合の価値は連合内の情報の集約から内生的に生成されますが、アウマン・ドレーゼのペイオフ、スイッチング摩擦、および受け入れルールが戦略的再構成を制御します。このフレームワークは、デグルート型のコンセンサスプロセスから譲渡可能な連合の価値が生まれ、インセンティブ主導の離脱と参加のダイナミクスを通じて連合構造が進化する、ファスト/スローアーキテクチャを導入しています。この分析では、安定した連合構造を維持する共同の戦術・戦略的均衡、戦術的・戦略的一致の条件、分離、分極化、認識上の障壁を特徴づけている。結合ダイナミクスに対して、固定小数点の特性評価と存在結果が確立されます。数値実験により、不安定性と合意のパラドックスが明らかになりました。スイッチング障壁が低い、または負であると、戦略的収束が妨げられると同時に、世界的な戦術的合意を達成するのに十分な時間的混合が促進される可能性があります。その結果は、マルチエージェント システムにおける連合形成、コンセンサス ダイナミクス、情報集約、戦略的安定性に関する統一的な視点を提供します。

原文 (English)

Multiscale Exit-Join Dynamics: Tactical Consensus and Strategic Coalition Formation

This paper develops a multiscale model of coalition formation in which strategic exit-and-join decisions are coupled with tactical consensus dynamics inside coalitions. Coalition value is generated endogenously from within-coalition information aggregation, while Aumann-Dreze payoffs, switching frictions, and acceptance rules govern strategic reconfiguration. The framework introduces a fast-slow architecture in which transferable coalition value emerges from DeGroot-style consensus processes, while coalition structures evolve through incentive-driven exit-and-join dynamics. The analysis characterizes joint tactical-strategic equilibria, conditions for tactical and strategic unanimity, segregation, polarization, and cognitive barriers that sustain stable coalition structures. A fixed-point characterization and existence results are established for the coupled dynamics. Numerical experiments reveal an instability-consensus paradox: low or negative switching barriers may prevent strategic convergence while simultaneously promoting temporal mixing sufficient to achieve global tactical consensus. The results provide a unified perspective on coalition formation, consensus dynamics, information aggregation, and strategic stability in multi-agent systems.

13:00 JSTエージェントロボティクス

教師なしメモリ強化型ビデオトランスフォーマー: 自律型農業用ローバーの障害物検出

自律型ローバーは精密農業に不可欠なものとなっていますが、一貫した運用上の安全性を達成することは依然として重要な課題です。 LiDAR などの従来の安全センサーは、プラントの天蓋の下にある障害物を検出できず、重大なリスクが生じます。カメラベースの教師あり学習手法は一般的なオブジェクトを検出できますが、トレーニング データに存在しない障害物に直面した場合にはパフォーマンスが低下します。実際の教師なし異常検出は、環境の通常の視覚パターンを学習することで解決策を提供しますが、移動する探査車によって捉えられた動的なシーンでは失敗することがよくあります。\\ この文書では、動的な農業シーンでのリアルタイムの障害物検出のために設計された完全に教師なしの手法である、異常検出用ビデオ メモリ トランスフォーマー (VMTAD) を紹介します。 VMTAD は、専用メモリ モジュールで強化されたトランス駆動アーキテクチャを利用します。このメモリ モジュールは、先行フレームのエンコードされた表現を処理することによって時間コンテキストを活用します。このアプローチにより、システムはロボットの動きによって引き起こされる動的コンテキストに効果的に対処できるようになります。モデルは、通常の動作を表す画像のみを使用してトレーニングされ、データ ラベルは必要ありません。\\ VMTAD は、農業用探査車「Grillion」で厳密に評価されました。困難な菜種データセットにおいて、VMTAD は最先端のパフォーマンスを達成し、受信者動作特性曲線の下の検出面積 0.973、セグメンテーション面積 0.997 に達しました。軽量バージョンは、探査車の総停止距離の分析によって確認されたように、安全性にとって重要な高精度とリアルタイム推論 (14 ミリ秒) の最適なバランスを提供します。

原文 (English)

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge. Conventional safety sensors, such as LiDAR, fail to detect obstacles positioned below the plant canopy, posing a significant risk. While camera-based supervised learning methods can detect common objects, they perform poorly when faced with obstacles that were not present in their training data. Actual unsupervised anomaly detection offers a solution by learning the normal visual patterns of an environment, but often fails for the dynamic scenes captured by a moving rover.\\ This paper introduces Video Memory Transformers for Anomaly Detection (VMTAD), a fully unsupervised method designed for real-time obstacle detection in dynamic agricultural scenes. VMTAD utilizes a transformer-driven architecture augmented with a dedicated memory module. This memory module leverages temporal context by processing encoded representations of preceding frames. This approach enables the system to effectively address the dynamic context caused by the robot's movement. The model is trained using only images that represent normal operation, requiring no data labels.\\ VMTAD was rigorously evaluated on the 'Grillion' agricultural rover. On a challenging rapeseed dataset, VMTAD achieved state-of-the-art performance, reaching a 0.973 detection and 0.997 segmentation Area Under the Receiver Operating Characteristic curve. A lightweight variant provides an optimal balance of high accuracy and real-time inference (14 ms), which is critical for safety, as confirmed by our analysis of the rover's total stopping distance.

13:00 JST研究/論文

スケーラブルなインデックス付けと取得のためのスライド全体の画像パッチングの冗長性の削減

デジタルパソロジーの急速な成長により、スライド画像全体 (WSI) の効率的なインデックス作成と検索が緊急に必要となっています。このニーズは、一か八かの臨床意思決定をサポートするために信頼できる類似性検索を必要とする、新たな生成 AI ワークフロー、特に検索拡張生成 (RAG) によってさらに強化されています。しかし、高性能ストレージには多額のコストがかかるため、多くの医療機関にとって WSI インデックス作成の拡張性とアクセスしやすさが制限されています。その結果、検索精度を維持しながらストレージの需要を削減できる方法が研究の重要な優先事項になっています。我々は、ARReST (Antithetical Redundancy Reduction Strategy) を提案します。ARReST (Antithetical Redundancy Reduction Strategy) は、異なる組織クラスにわたる冗長性を活用して、各 WSI からインデックスを作成する必要があるパッチの数を大幅に削減する、原則に基づいた対立的なフレームワークです。 ARReST は、クラス内の重複のみを削除するのではなく、正反対のパッチ (その表現がクラス間の差別に最小限に寄与するパッチ) を特定し、検索可能なアーカイブからそれらを削除します。この目標を絞った削減により、形態学的多様性や検索忠実度を犠牲にすることなく、インデックスが大幅に圧縮されます。 ARReST は、余分なパッチ表現を最小限に抑えることで、ストレージ フットプリントを削減し、計算オーバーヘッドを削減し、大規模な病理リポジトリにわたる類似性検索を高速化します。 TCGA リポジトリ (21 臓器を含むがんゲノム アトラス) での大規模な実験により、ARReST が競合する検索パフォーマンスを維持しながら大幅なインデックス圧縮を達成することが実証されました。観察された 3% ~ 60% (14%$\pm$13%) のストレージ節約は、多くの臓器の検索パフォーマンスを損なうことなく確実に達成できます。提案された戦略は、スケーラブルでコスト効率の高い WSI インデックス作成を可能にし、次世代の検索主導型臨床 AI システムに最適です。

原文 (English)

Reducing Redundancy in Whole-Slide Image Patching for Scalable Indexing and Retrieval

The rapid growth of digital pathology has created an urgent need for efficient indexing and retrieval of whole slide images (WSIs). This need is intensified by emerging generative AI workflows, particularly retrieval-augmented generation (RAG), which require dependable similarity search to support high-stakes clinical decision-making. Yet the substantial cost of high-performance storage limits the scalability and accessibility of WSI indexing for many healthcare institutions. Consequently, methods that can reduce storage demands while preserving retrieval accuracy have become a critical research priority. We propose ARReST (Antithetical Redundancy Reduction Strategy), a principled oppositional framework that leverages redundancy across dissimilar tissue classes to markedly decrease the number of patches that must be indexed from each WSI. Instead of eliminating only within-class duplicates, ARReST identifies antithetical patches-those whose representations contribute minimally to cross-class discrimination-and prunes them from the searchable archive. This targeted reduction substantially compresses the index without sacrificing morphological diversity or retrieval fidelity. By minimizing superfluous patch representations, ARReST reduces storage footprint, lowers computational overhead, and accelerates similarity search across large pathology repositories. Extensive experiments on TCGA repository (The Cancer Genome Atlas with 21 organs) demonstrate that ARReST achieves significant index compression while maintaining competitive retrieval performance. The observed storage savings of 3% to 60% (14%$\pm$13%) can be reliably achieved without compromising retrieval performance for many organs. The proposed strategy enables scalable, cost-efficient WSI indexing and is well-suited for next-generation retrieval-driven clinical AI systems.

13:00 JST研究/論文

敵対的生成ネットワークのニューラル アーキテクチャの探索: 包括的なレビューと批判的分析

Neural Architecture Search (NAS) は、敵対的生成ネットワーク (GAN) の設計を最適化する上で極めて重要な技術として登場し、手動設計に固有の課題に対処しながら効果的なアーキテクチャの検索を自動化します。このペーパーでは、GAN に適用される NAS 手法の包括的なレビューを提供し、検索戦略、評価指標、パフォーマンス結果などの基準に基づいてさまざまなアプローチを分類および比較します。このレビューでは、GAN のパフォーマンス、安定性、効率の向上における NAS の利点を強調するとともに、将来の研究の限界と領域も特定しています。主な発見には、特定の状況における進化的アルゴリズムと勾配ベースの手法の優位性、インセプション スコア (IS) やフレシェ インセプション ディスタンス (FID) などの従来のスコアを超える堅牢な評価指標の重要性、GAN パフォーマンスの評価における多様なデータセットの必要性などが含まれます。この論文は、既存の NAS-GAN 技術の構造化された比較を提示することにより、研究者がより効果的な NAS 手法を開発し、GAN 分野を発展させるためのガイドとなることを目的としています。

原文 (English)

Neural Architecture Search for Generative Adversarial Networks: A Comprehensive Review and Critical Analysis

Neural Architecture Search (NAS) has emerged as a pivotal technique in optimizing the design of Generative Adversarial Networks (GANs), automating the search for effective architectures while addressing the challenges inherent in manual design. This paper provides a comprehensive review of NAS methods applied to GANs, categorizing and comparing various approaches based on criteria such as search strategies, evaluation metrics, and performance outcomes. The review highlights the benefits of NAS in improving GAN performance, stability, and efficiency, while also identifying limitations and areas for future research. Key findings include the superiority of evolutionary algorithms and gradient-based methods in certain contexts, the importance of robust evaluation metrics beyond traditional scores like Inception Score (IS) and Fr\'echet Inception Distance (FID), and the need for diverse datasets in assessing GAN performance. By presenting a structured comparison of existing NAS-GAN techniques, this paper aims to guide researchers in developing more effective NAS methods and advancing the field of GANs.

13:00 JST画像/動画生成

LCG: 疎なリレーショナル アテンションによるロングコンテキストの一貫した画像生成

最近の画像生成モデルは、単一画像の合成では優れた品質を実現していますが、コミック、ストーリーボード、ビジュアル ナラティブで必要とされる、連続した出力全体で一貫性を維持できないことがよくあります。我々は、ロングコンテキストマルチ画像生成における一貫性とスケーラビリティを向上させるために、ロングコンテキストマルチ画像のテキストから画像への生成のためのフレームワークであるロングコンテキスト生成(LCG)を提案します。 LCG は、スパース リレーショナル アテンション (SRA) メカニズムを採用して、拡張されたビジュアル コンテキスト全体にわたるコア機能に選択的に対応し、セマンティック情報とレイアウト情報の伝播が計算上扱いやすい状態を保つようにします。セマンティックな調整を強制するために、ルーティング一貫性制約 (RCC) を導入します。これは、アイデンティティ認識マスクを活用して、世代ブランチ全体で構造パターンを調整し、複雑なマルチキャラクター シーンであっても外観のドリフトを効果的に軽減します。この設定でのトレーニングと評価をサポートするために、さまざまな状況コンテキストにわたる文字中心の複数画像シーケンスで構成される大規模な合成データセットであるロングコンテキスト一貫性データセット (LCCD) を構築します。 LCCD には 600K のトレーニング シーケンスと個別の 1K テスト セットが含まれており、各シーケンスには 6 ~ 20 枚の画像が含まれています。実験では、LCG が、マルチキャラクター シーンを含むロング コンテキスト イメージ生成におけるプロンプト アラインメントとキャラクターの一貫性において、比較したベースラインよりも優れていることが実証されました。

原文 (English)

LCG: Long-Context Consistent Image Generation with Sparse Relational Attention

Recent image generation models achieve impressive quality in single-image synthesis, but often fail to maintain consistency across sequential outputs, as required in comics, storyboards, and visual narratives. We propose Long-Context Generation (LCG), a framework for long-context multi-image text-to-image generation, to improve consistency and scalability in long-context multi-image generation. LCG employs the Sparse Relational Attention (SRA) mechanism to selectively attend to core features across extended visual contexts, ensuring that the propagation of semantic and layout information remains computationally tractable. To enforce semantic alignment, we introduce the Routing Consistency Constraint (RCC), which leverages identity-aware masks to align structural patterns across generation branches, effectively mitigating drift in appearance even in complex multi-character scenes. To support training and evaluation in this setting, we construct the Long-Context Consistency Dataset (LCCD), a large-scale synthetic dataset comprising character-centric multi-image sequences spanning varied situational contexts. LCCD contains 600K training sequences and a separate 1K test set, with each sequence containing 6 to 20 images. The experiments demonstrate that LCG outperforms the compared baselines in prompt alignment and character consistency for long-context image generation, including multi-character scenes.

13:00 JST研究/論文

KG-TRACE: 抗菌薬耐性予測における機械的接地のための神経象徴的フレームワーク

WGS ベースの AMR 予測は高精度に達していますが、既存のモデルには、確立された生物学的経路における神経の属性を根拠付けるメカニズムが欠けています。我々は、神経ゲノムモデルに対する構造化された生物学的制約としてWHOの突然変異知識グラフ(KG)を統合する新しい神経記号フレームワークであるKG-TRACEを紹介します。統計パターンを単独で学習する既存の方法とは異なり、KG-TRACE は、学習された認識論的トラスト ゲートを通じてゲノム特徴と RotatE ベースの KG 埋め込みを融合し、象徴的な生物学的知識に対して神経証拠を動的に重み付けします。 CRyPTIC 結核菌コホートで評価した KG-TRACE は、イソニアジドの AUROC 0.9760 を達成し、競合する精度を達成していますが、その主な価値は予測的上昇率ではなく象徴的な根拠にあります。さらに重要なのは、神経属性と確立された生物学の間の整合性を定量化するデータセットレベルの指標である生物学的接地率 (BGR) を導入することです。私たちのフレームワークは、イソニアジド耐性予測の 92.5% の象徴的カバレッジを達成し、「不確か」な症例に対して検査室追跡フラグを発行することにより、MDR 共起アーティファクトを効果的に特定します。私たちは、神経シンボリックグラウンディングが臨床医に検証可能な監査証跡を提供し、予測の精度と臨床の信頼の間のギャップを埋めることを実証します。

原文 (English)

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction

While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pathways. We present KG-TRACE, a novel neuro-symbolic framework that integrates the WHO mutation knowledge graph (KG) as a structured biological constraint on a neural genomic model. Unlike existing methods that learn statistical patterns in isolation, KG-TRACE fuses genomic features and RotatE-based KG embeddings through a learned epistemic trust gate, dynamically weighting neural evidence against symbolic biological knowledge. Evaluated on the CRyPTIC M. tuberculosis cohort, KG-TRACE achieves an AUROC of 0.9760 for isoniazid, achieving competitive accuracy while its primary value lies in symbolic grounding, not predictive uplift. More importantly, we introduce the Biological Grounding Ratio (BGR), a dataset-level metric that quantifies alignment between neural attributions and established biology. Our framework achieves a 92.5% symbolic coverage of isoniazid-resistant predictions and effectively identifies MDR co-occurrence artifacts by issuing laboratory follow-up flags for 'UNCERTAIN' cases. We demonstrate that neuro-symbolic grounding provides a verifiable audit trail for clinicians, bridging the gap between predictive accuracy and clinical trust.

13:00 JSTロボティクス

LiMoDE: 動的専門家の混合の観点から生涯にわたるロボット操作を再考する

事前の知識を活用して継続的にタスクに適応できるジェネラリストロボットを構築することは、依然として大きな課題です。以前の研究では、単一タスクの適応のためのパラメータ効率の高い微調整によって、壊滅的な忘却の問題が軽減されました。ただし、再利用可能なスキルを抽出したり、他のスキルとの相互作用を効果的にモデル化したりすることはできません。最近の研究では、プロンプトを学習することでこれらの問題に対処しようとしています。これとは異なり、この論文は、生涯にわたるロボット操作のための新しい 2 段階学習スキームである、動的専門家の生涯混合 (\textit{LiMoDE}) に関するアーキテクチャの観点を示しています。具体的には、動的 MoE 構造は、事前知識を学習するためのマルチタスク事前トレーニング段階で最初に提案されます。そこでは、さまざまな短期間の操作に対処するために、さまざまな数の異質な専門家が動作情報に基づいてアクティブ化されます。続いて、タスク適応段階では、生涯専門家を学習し、新しいタスクのためにそれらを凍結された専門家と動的に組み合わせて、適応中の知識の伝達を促進する生涯MoE適応メカニズム%(LiMoEAM)を設計します。提案された \textit{LiMoDE} は、シミュレートされた生涯学習ベンチマークと現実世界のタスクの両方で評価されます。広範な実験により、適度な数の追加のトレーニング可能なパラメーターと推論オーバーヘッドを導入することで、優れたパフォーマンスと強力な生涯適応を達成する有効性が実証されています。

原文 (English)

LiMoDE: Rethinking Lifelong Robot Manipulation from a Mixture-of-Dynamic-Experts Perspective

Building a generalist robot that can leverage prior knowledge for continuous task adaptation remains a significant challenge. Previous works alleviate the catastrophic forgetting problem by parameter-efficient fine-tuning for single-task adaptation. However, they fail to extract reusable skills and model the interaction with other skills effectively. Recent works try to address these issues by learning prompts. Differently, this paper presents an architectural perspective on the Lifelong Mixture of Dynamic Experts (\textit{LiMoDE}), a novel two-stage learning scheme for lifelong robot manipulation. Specifically, a dynamic MoE structure is first proposed in the multi-task pre-training stage to learn prior knowledge, where a varied number of heterogeneous experts are activated based on the motion information to address different short-term manipulations. Subsequently, in the task adaptation stage, we design a lifelong MoE adaptation mechanism % (LiMoEAM) that learns lifelong experts and dynamically combines them with frozen ones for new tasks, facilitating the knowledge transfer during adaptation. The proposed \textit{LiMoDE} is evaluated on both the simulated lifelong learning benchmark and real-world tasks. Extensive experiments demonstrate its effectiveness in achieving superior performance and strong lifelong adaptation by introducing a moderate number of additional trainable parameters and inference overhead.

13:00 JSTLLM/生成AI画像/動画生成OpenAIDeepSeek

構造から相乗効果へ: マルチモーダル大規模言語モデルにおける視覚言語知覚パラダイム進化の調査

マルチモーダル大規模言語モデル (MLLM) は、特に OpenAI の O シリーズや DeepSeek の R シリーズなどのモデルの導入後、視覚言語の理解と推論の統合において目覚ましい進歩を遂げ、知覚中心の知能へのパラダイム シフトを推進しました。しかし、真に統一された視覚と言語の観点、つまり視覚と言語を切り離せない様式として扱う視点から知覚を調査する体系的な調査は依然として不足している。既存のレビューは断片化されていることが多く、視覚か言語のどちらかに個別に焦点を当てているため、統合された能力として知覚のクロスモーダルな進化を捉えることはほとんどありません。このギャップを埋めるために、MLLM における統一された視覚言語認識に関する最初の体系的な調査を紹介します。具体的には、(1) MLLM 知覚を人間の生得的な知覚に類似した固有の統合された視覚言語能力として形式化し、(2) MLLM 知覚のパラダイム進化を追跡する 5 段階の分類を導入し、各段階での代表的な方法とマイルストーンを調査し、(3) 未解決の課題を特定し、真に一般的で統合されたマルチモーダル インテリジェンスに向けた有望な研究方向性を概説します。私たちの研究が、汎用人工知能 (AGI) への道におけるさらなるイノベーションを促進するための基礎的な理解と実行可能なロードマップの両方を提供することを願っています。

原文 (English)

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-series, which have driven a paradigm shift toward perception-centric intelligence. However, there remains a lack of systematic surveys that examine perception from a truly unified vision-language perspective -- one that treats vision and language as an inseparable modality. Existing reviews are often fragmented, focusing separately on either vision or language, and thus rarely capture the cross-modal evolution of perception as an integrated capability. To bridge this gap, we present the first systematic survey of unified vision-language perception in MLLMs. Specifically, we (1) formalize MLLM perception as an intrinsic, unified vision-language capability analogous to human innate perception, (2) introduce a five-stage taxonomy tracing the paradigm evolution of MLLM perception and survey representative methods and milestones at each phase, and (3) identify open challenges and outline promising research directions toward truly general, unified multimodal intelligence. We hope our study will provide both a foundational understanding and an actionable roadmap to foster further innovation on the path toward artificial general intelligence (AGI).

13:00 JST研究/論文

アルゴリズムの公平性に対する統計的および構造的アプローチ

現代の機械学習システムは、孤立した予測構造としての起源を超え、人間の機会を積極的に仲介する複雑な社会技術アーキテクチャへと進化しました。アルゴリズムが経済的および社会的機会へのアクセスを決定することが増えているため、これらのシステムには環境の構造的不平等と偏見が深く埋め込まれていることが広く認識されるようになりました。アルゴリズム的公平性の分野は、予測精度のために最適化されたモデルが社会的に疎外されたグループに不利になる可能性があるという認識の高まりに応えて登場しました。しかし、初期の緩和戦略は脆弱な単純化に基づいており、複雑な社会技術的環境では有効性が制限されていました。この論文は、現代の公平性パラダイムの 2 つの基本的な限界を特定し、それに対処します。それは、監査における決定論的な点推定への依存と、構造的コンテキストを持たない孤立した実体としての個人の扱いです。

原文 (English)

Statistical and Structural Approaches to Algorithmic Fairness

Modern machine learning systems have outgrown their origins as isolated predictive constructs, evolving into complex socio-technical architectures that actively mediate human opportunity. As algorithms increasingly determine access to economic and social opportunities, it has become widely recognized that these systems are deeply embedded with the structural inequalities and prejudices of their environments. The field of algorithmic fairness emerged in response to the growing recognition that models optimized for predictive accuracy can systematically disadvantage marginalized groups. Early mitigation strategies, however, rested on fragile simplifications that limited their effectiveness in complex socio-technical environments. This thesis identifies and addresses two fundamental limitations of contemporary fairness paradigms: the reliance on deterministic point estimates for auditing and the treatment of individuals as isolated entities devoid of structural context.

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

Cyber​​ChainBench: AI エージェントは現実世界のオンチェーン脆弱性からスマート コントラクトを保護できるか?

Cyber​​ChainBench は、脆弱性検出、エクスプロイト生成、パッチ合成という 3 つの補完的なタスクにわたって、スマート コントラクト セキュリティに関する LLM ベースのエージェントを評価するためのベンチマークです。 9つのEVMチェーンにまたがるDeFiHackLabsからの541件の実世界エクスプロイトインシデントから構築されたこのベンチマークは、コードを読み取り、トランザクションを追跡し、メインネットフォーク上のエクスプロイトを検証するためのツールを使用して、Harborによってオーケストレーションされた分離された評価環境を通じてエージェントが過去のブロックチェーン状態と対話するエンドツーエンドのオンチェーン評価を提供します。各ケースは特定のブロックに固定されており、脆弱性の種類、位置特定、攻撃者の利益をカバーする構造化されたグラウンド トゥルースが含まれています。エクスプロイトは、歴史的フォークへの経済的影響によって等級付けされます。パッチは、プロキシでアップグレード可能なサブセット上で過去の攻撃と正当なトランザクションを失敗テストのオラクルとして再生することによって検証されます。 5 つのタイプの脆弱性分類を定義し、複数のエージェント - モデル構成を評価します。結果は、明確な難易度の勾配を明らかにしました。最良の構成スコアは、検出で 37.5%、悪用で 43.7% でしたが、パッチ適用では 23.4% にすぎず、トップのエージェント (GPT-5.5 を搭載した Codex) は、1 件あたり 2.39 ドルのコストで設定された 200 件の悪用セット全体で、合計 5,740 万円の悪用利益を実現しました。

原文 (English)

CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?

We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch synthesis. Built from 541 real-world exploit incidents from DeFiHackLabs spanning 9 EVM chains, the benchmark provides end-to-end on-chain evaluation where agents interact with historical blockchain state through isolated evaluation environments orchestrated by Harbor, using tools to read code, trace transactions, and validate exploits on mainnet forks. Each case is anchored to a specific block and includes structured ground truth covering vulnerability type, localization, and attacker profit. Exploits are graded by economic impact on historical forks; patches are validated by replaying historical attacks and legitimate transactions as fail-to-pass test oracles on a proxy-upgradeable subset. We define a five-type vulnerability taxonomy and evaluate multiple agent--model configurations. Results reveal a clear difficulty gradient: the best configuration scores 37.5% on detection, 43.7% on exploitation, but only 23.4% on patching, with the top agent (Codex with GPT-5.5) realizing \$57.4M in total exploit profit across the 200-case exploit set at a cost of $2.39 per case.

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

Lacuna: 機械学習の研究マップ

Lacuna は、LLM を使用して論文や学術メタデータをマークダウンの要約、概念要素、研究の方向性、研究提案に変換する機械学習用の研究マップです。各項目には、それを裏付ける一次情報源の記録と論文へのリンクが保持されています。 Web、マークダウン、MCP インターフェイスを備えたマップをリリースします。 LitSearch、Multi-XScience-CS/ML、および ScholarQA-CS-ML 全体で、Lacuna は OpenScholar を上回り、LitSearch 取得で最も優れた向上を示しています (Recall@10 0.538 に対し、OpenScholar v3 では 0.424)。また、マップ上の多段階レポート エージェントである Lacuna Deep Research を、25 の ReportBench-ML 調査タスクで評価します。Lacuna Deep Research は、引用 F1 が 0.052、引用精度が 0.339、専門家参照ヒット数が 99、RACE レポート品質が 7.82/10 に達するのに対し、GPT-Researcher は、F1 が 0.039、精度が 0.290、ヒット数が 72、 5.24/10 レース。

原文 (English)

Lacuna: A Research Map for Machine Learning

Lacuna is a research map for machine learning that uses LLMs to turn papers and scholarly metadata into markdown summaries, concept elements, research directions, and research proposals. Each item keeps links to the primary source records and papers that support it. We release the map with web, markdown, and MCP interfaces. Across LitSearch, Multi-XScience-CS/ML, and ScholarQA-CS-ML, Lacuna outperforms OpenScholar with the strongest gains on LitSearch retrieval (Recall@10 0.538 vs. 0.424 for OpenScholar v3). We also evaluate Lacuna Deep Research, a multi-stage report agent over the map, on 25 ReportBench-ML survey tasks: Lacuna Deep Research reaches 0.052 citation F1, 0.339 citation precision, 99 expert-reference hits, and 7.82/10 RACE report quality, while GPT-Researcher reaches 0.039 F1, 0.290 precision, 72 hits, and 5.24/10 RACE.

13:00 JST画像/動画生成

レーザー溶接における溶け込み深さと形態を予測するためのマルチタスク時空間ディープ ニューラル ネットワーク

レーザー溶け込み溶接では、溶け込み状態と溶接シーム形態の評価が溶接品質を決定する上で重要な役割を果たします。この論文では、溶け込み状態、深さ、溶接シームの形態を高精度に予測する機能を備えた革新的なマルチタスク深層学習モデルの包括的な紹介を行います。この監視プラットフォームは、相補型金属酸化物半導体カメラを使用してレーザー溶接プロセス中にキャプチャされた溶融池画像に依存しています。提案されたモデルは、上部の溶接池画像から抽出された時空間特徴と溶接パラメータを統合し、畳み込みニューラル ネットワークと状態空間モデルに基づく深層学習フレームワークを確立し、時空間情報のより効率的な抽出と処理を実現します。さらに、開発されたモデルの堅牢性と一般化能力の両方を強化するために、データセットを構築するための信頼できる方法が提案されています。テストセットの検証結果では、溶け込み状態の予測精度が 99.35% に達し、溶け込み深さの予測誤差は 1.79 ミリメートル、溶接断面の再構成精度は 95.65% であることが実証されました。この研究は、レーザー浸透溶接システムにおける現場品質管理戦略に対する新しい洞察と方法論を提供します。

原文 (English)

A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding

In laser penetration welding, the assessment of penetration state and weld seam morphology plays a crucial role in determining the weld quality. This paper presents a comprehensive introduction of the innovative muti-task deep learning model that has the capability to predict penetration state, depth, and weld seam morphology with high accuracy. The monitoring platform relies on weld pool images captured during the laser welding process using a complementary metal-oxide-semiconductor camera. The proposed model integrates spatiotemporal features extracted from top weld pool images along with welding parameters, establishing a deep learning framework based on convolutional neural networks and state space models for more efficient extraction and processing of spatial-temporal information. Furthermore, a reliable method for constructing the dataset is proposed to enhance both robustness and generalization capability of the developed model. Validation results on the test set demonstrate that prediction accuracy for penetration state can reach 99.35%, while prediction error for penetration depth is 1.79 millimeter, and accuracy of reconstructing the weld cross-section is 95.65%. This study provides new insights and methodologies for in-situ quality control strategies in laser penetration welding systems.

13:00 JSTLLM/生成AI

クリックからインテントへ: 金融サービス推奨のための LLM 抽出分類法を使用したクロスプラットフォーム セッションの埋め込み

逐次ユーザー行動モデリングは、産業用レコメンダー システムで広く採用されています。ただし、金融サービスでは依然として大きなギャップがあり、ログイン前の Web インタラクションと認証されたアプリ内エクスペリエンスが大幅に異なります。具体的には、ログイン前の Web ユーザーは通常、新製品を探索しますが、ログインしたアプリ ユーザーはアカウントのサービスに重点を置きます。クロスチャネルのエンティティ解決の課題 (匿名の Web セッションと認証されたモバイル アカウントの照合など) により、Web ベースのインテント シグナルは認証後のパーソナライゼーションに十分に活用されていないままです。 Web ベースの意図を捕捉するための既存の方法は、アドホックで範囲が狭いことが多く、下流での定量的な推奨事項と大規模な定性的な理解をサポートする柔軟性に欠けています。この研究では、Web ベースのインタラクションのためのスケーラブルで二重目的の意図予測フレームワークを提案し、パーソナライゼーションへの適用可能性を実証します。私たちのアプローチは、生の Web クリックストリームを 2 つの出力に変換します。自己監視型 Transformer がマルチモーダル クリックストリームをコンパクトなセッション埋め込みにエンコードし、LLM ベースの分類生成および蒸留パイプラインが解釈可能なインテント ラベルを生成します。私たちのシステムは、LLM 抽出分類法と組み合わせた自己教師ありクリックストリーム表現が、実稼働環境における定量的タスクと定性的理解を共同で提供できることを実証しています。モバイル ホームページのタイル ランキング タスクでは、セッションの埋め込みにより、実稼働ベースラインと比較してマクロ Recall@1 が 1.88% 向上し、ログ損失が 13.38% 削減されました。ユーザー変換予測タスクでは、埋め込みはマイクロ F1 で LLM ラベルのパフォーマンスを 4.3% 上回っていますが、蒸留層はわずか 7% のパフォーマンス低下で超低遅延で解釈可能なラベルを提供します。

原文 (English)

From Clicks to Intent: Cross-Platform Session Embeddings with LLM-Distilled Taxonomy for Financial Services Recommendations

Sequential user behavior modeling is widely adopted in industrial recommender systems; however, significant gaps remain in financial services, where pre-login web interactions and authenticated in-app experiences differ drastically. Specifically, pre-login web users typically explore new products, whereas logged-in app users focus on account servicing. Due to the challenge of cross-channel entity resolution (e.g., matching anonymous web sessions to authenticated mobile accounts), web-based intent signals remain underutilized for post-authentication personalization. Existing methods for capturing web-based intent are often ad-hoc and narrow, lacking the flexibility to support both quantitative downstream recommendations and qualitative understanding at scale. In this work, we propose a scalable and dual-purpose intent prediction framework for web-based interactions and demonstrate its applicability for personalization. Our approach transforms raw web clickstreams into two outputs: a self-supervised Transformer encodes multi-modal clickstreams into a compact session embedding, while an LLM-based taxonomy generation and distillation pipeline produces interpretable intent labels. Our system demonstrates that self-supervised clickstream representations combined with LLM-distilled taxonomies can jointly serve quantitative tasks and qualitative understanding in production: on the mobile homepage tile ranking task, the session embedding improves macro Recall@1 by 1.88% and reduces Log Loss by 13.38% over production baselines. On the user conversion prediction task, the embedding outperforms the LLM labels by 4.3% on micro F1, while the distillation layer delivers interpretable labels at ultra-low latency with only a 7% performance drop.

13:00 JST研究/論文

TEMPO-Diffusion: 一時的に暴露される拡散モデルの悪意のあるポイズニング

拡散モデルに対するノイズベースのバックドア攻撃は通常、入力時のトリガー挿入、対象外のアクティベーション、および配布範囲外のターゲット生成に依存します。このような想定は、これらの攻撃のステルス性と実際的な関連性の両方を低下させます。この研究では、悪意のある配布のシフトを一時的な配布内エクスポージャに局所化する、標的型バックドア フレームワークである TEMPO-Diffusion を紹介します。 TEMPO-Diffusion は、(i) 特定のクラスに対する標的型攻撃、(ii) 複数の異なる出力イメージ内および複数の場所で特定の特徴を再構築する複数のサブイメージ バックドア、および (iii) 時間条件付きトリガーによるインペインティングをサポートします。合成トレーニング データのバックドア拡散モデルを活用する際に関連する実際的なセキュリティ上の懸念を研究するために、カナダと米国の道路標識を強調したバランスの取れた地域認識型交通標識データセットである CALISA も紹介します。 CIFAR10、GTSRB、および CALISA にわたって、私たちの実験は、TEMPO-Diffusion がクラス固有の合成データ生成を確実に妨害し、そのデータに基づいてトレーニングされた下流の分類器で高い攻撃成功率を誘導できることを示しています。

原文 (English)

TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models

Noise-based backdoor attacks on diffusion models typically rely on input-time trigger injection, untargeted activation, and out-of-distribution target generation. Such assumptions reduce both the stealthiness and the practical relevance of these attacks. In this work, we present TEMPO-Diffusion, a targeted backdoor framework that localizes the malicious distribution shift to a temporal, in-distribution exposure. TEMPO-Diffusion supports: (i) targeted attacks on and to specific classes, (ii) multiple sub-image backdoors that reconstruct specific features within multiple, different output images and at multiple locations, and (iii) in-painting with time-conditioned triggers. To study relevant, practical security concerns in leveraging backdoored diffusion models for synthetic training data, we also introduce CALISA: a balanced, region-aware traffic-sign dataset emphasizing Canadian and U.S. road signs. Across CIFAR10, GTSRB, and CALISA, our experiments show that TEMPO-Diffusion can reliably poison class-specific synthetic data generation and induce high attack success rates in downstream classifiers trained on that data.

13:00 JST研究/論文Mistral AI

Hankel 低次モデリングによる SSM アダプター: ロングコンテキストの微調整におけるタスクの適合性をインジェクションサイトが決定

パラメータ効率の良い微調整 (PEFT) は通常、注意プロジェクターを対象としていますが、逐次的な状態の蓄積を必要とするタスクに対するその有効性はまだ調査されていません。このようなタスクの PEFT が状態空間モデル (SSM) アダプターの恩恵を受けることができるかどうか、また MLP ブロックがより良い注入サイトであるかどうかを調べます。経験的ハンケル グラミアンの平衡切り捨てによって初期化された SSM ベースの残差モジュールである Hankel Reduced order Model (HRM) アダプターを紹介します。システム行列 $\bar{A}$ の時間不変性を利用することで、HRM は正確な FFT ベースの並列スキャンを可能にし、すべてのコンテキスト長にわたって LoRA との計算パリティを実現します。 Mistral-7B (840 万のトレーニング可能なパラメーター) のアイソパラメトリック評価では、HRM は、QuaALITY (+34.8\% 相対精度) や QMSum (+71.6\% 相対 ROUGE-1) を含む LongBench タスクで LoRA バリアントよりも優れたパフォーマンスを発揮します。さらに、HRM は、合成状態追跡 (DFA、パリティ) と文字レベル言語モデリング (enwik8) の 18 の構成にわたって一貫した優位性を示しています。ゲート解析により、HRM アダプターが反復を調整することを効果的に学習し、長いコンテキストのシーケンス モデリングに対する低ランクの適応に代わる堅牢なアーキテクチャの代替手段が提供されることが明らかになりました。

原文 (English)

SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning

While parameter-efficient fine-tuning (PEFT) typically targets attention projectors, its efficacy for tasks requiring sequential state accumulation remains under-explored. We examine if PEFT for such tasks can benefit from state space model (SSMs) adapters, and if MLP blocks are better injection sites. We introduce Hankel Reduced order Model (HRM) adapter, an SSM-based residual module initialized via Balanced Truncation of empirical Hankel Grammians. By leveraging the time-invariance of the system matrix $\bar{A}$, HRM enables an exact FFT-based parallel scan, achieving computational parity with LoRA across all context lengths. In iso-parametric evaluations on Mistral-7B (8.4M trainable parameters), HRM outperforms LoRA variants on LongBench tasks, including QuALITY (+34.8\% relative accuracy) and QMSum (+71.6\% relative ROUGE-1). HRM further demonstrates consistent superiority across 18 configurations of synthetic state-tracking (DFA, Parity) and character-level language modeling (enwik8). Gate analysis reveals that HRM adapters effectively learn to modulate recurrence, providing a robust architectural alternative to low-rank adaptation for long-context sequence modeling.

13:00 JSTエージェント研究/論文

Red Queen G\"odel Machine: 共進化するエージェントとその評価者

自己改善エージェントは、エージェント コーディング ベンチマークにおける最先端 (SOTA) であり、最近では一般的なドメインに拡張されています。ただし、それらの検索方法は通常、エージェントが改善しても有効であり続ける固定の検証基準、ベンチマーク、またはラベル付きデータセットといっ​​た固定的な評価基準を前提としています。これは進化の中心的な特徴、つまり種は環境の変化に応じて適応するという点を無視している。私たちは、再帰的な自己改善にも同じ原理を導入し、評価を改善ループの一部とし、進化する評価者、敵対的な目標、静的なベンチマークを超える可能性のある動的なユーティリティへの探索を開くことを目指しています。非定常ユーティリティ下での再帰的自己改善のための進化的フレームワークである Red Queen Godel Machine (RQGM) を紹介します。 RQGM は、制御されたユーティリティの進化を通じてこれを可能にします。検索は、固定されたエポック内評価基準を使用してエポックに編成されますが、ユーティリティはエポック境界で更新できるため、目的がエポック全体で進化するにつれて自己改善の保証がエポックごとに保持されます。まず、検証可能なコーディングタスクであっても、RQGM が補完的なエージェントとしてのジャッジコードレビュー信号を追加することにより、以前の SOTA よりもテスト合格率を向上させることを示します。このシグナルは安価で、RQGM が使用するトークンの量は 1.35 倍から 1.72 倍少なくなります。次に、科学論文の執筆と査読、オリンピックレベルの校正と採点に移ります。そこでは、RQGM が以前の自己改善エージェントに比べてパフォーマンスを向上させます。多様な審査員パネルの下で、共進化したライターは 1.78 倍から 1.86 倍高い合格率に達し、同時に進化した採点者は 9% 高いグラウンドトゥルース精度に達します。論文査読では、最も強力なベースライン査読者が AI によって生成された論文を人間の割合の最大 1.91 倍で過剰に受け入れます。 RQGM は、AI と人間の作業に対して同等に厳しいレビュー担当者を発見する敵対的目標を導入することでこれを修正します。

原文 (English)

The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators

Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid as the agent improves. This ignores a central feature of evolution: species adapt as their environments change with them. We aim to bring the same principle to recursive self-improvement, making evaluation part of the improvement loop and opening search to evolving evaluators, adversarial objectives, and dynamic utilities that may surpass static benchmarks. We introduce the Red Queen Godel Machine (RQGM), an evolutionary framework for recursive self-improvement under non-stationary utilities. The RQGM makes this possible through controlled utility evolution: search is organized into epochs with a fixed within-epoch evaluation criterion, while the utility can be updated at epoch boundaries, so self-improvement guarantees hold per epoch as the objective evolves across them. We begin by showing that even on verifiable coding tasks, the RQGM improves test pass rate over the prior SOTA by adding a complementary agent-as-a-judge code-review signal. This signal is cheaper and the RQGM uses 1.35x-1.72x fewer tokens. We then turn to scientific paper writing and reviewing, and Olympiad-level proof writing and grading, where the RQGM improves performance over prior self-improving agents: co-evolved writers reach 1.78x-1.86x higher acceptance rates under a diverse agent-as-a-judge panel, while co-evolved graders reach 9% higher ground-truth accuracy. In paper reviewing, the strongest baseline reviewer over-accepts AI-generated papers at up to 1.91x the human rate. The RQGM corrects this by introducing an adversarial objective that discovers reviewers equally stringent on AI and human work.

13:00 JST研究/論文

ベアリングの故障診断と機械の状態監視のためのパラメトリック汎用適応モーメント機能 (PG-AMF)

回転機械の転動体軸受の正確な故障診断は、労働安全を確保し、予知保全を可能にするために不可欠であると考えられています。従来の統計的特徴ベースの方法は、事前定義された記述子に依存しており、その診断感度は固定構成によって制限され、さまざまな障害状態に対する適応性が制限されています。深層学習アプローチは強力な表現能力を提供しますが、その有効性は多くの場合、高いデータ要件と解釈可能性の低下によって制限されます。この研究では、特徴特性が手動で指定されるのではなくデータから直接学習される、パラメトリック適応特徴抽出フレームワークが提案されています。信号エネルギー分布を捉える絶対的特徴、波形の非対称性を反映する符号付きモーメント特徴、動的変動を強調するAC結合モーメント特徴など、複数の相補的表現が振動信号から抽出され、同時に複数のセンサーチャネル間の相互作用が構造化融合メカニズムを通じてモデル化され、故障表現が強化されます。提案されたアプローチは、通常の動作と複数の故障タイプを含む 5 つの健全性状態を含むベンチマーク ギアボックス ベアリング データセットで評価されます。従来の方法と比較して分類パフォーマンスの向上が観察され、相互検証で一貫した結果が得られ、強力な一般化能力が示されています。さらに、低次元投影におけるより明確なクラスタリング パターンを通じて、強化された特徴の分離性が実証されます。学習された表現は広範囲の信号特性を効果的に捕捉し、診断性能の向上と産業用監視システムでの実用的な適用性の両方をサポートします。

原文 (English)

Parametric Generalized Adaptive Moment Features (PG-AMF) for Bearing Fault Diagnosis and Machine Health Monitoring

Accurate fault diagnosis of rolling element bearings in rotating machinery is considered essential for ensuring industrial safety and enabling predictive maintenance. Conventional statistical feature-based methods rely on predefined descriptors, whose diagnostic sensitivity is constrained by fixed configurations and limited adaptability across varying fault conditions. Although deep learning approaches offer strong representational capacity, their effectiveness is often restricted by high data requirements and reduced interpretability. In this work, a parametric adaptive feature extraction framework is proposed, in which feature characteristics are learned directly from data rather than being manually specified. Multiple complementary representations are extracted from vibration signals, including absolute features capturing signal energy distribution, signed moment features reflecting waveform asymmetry, and AC-coupled moment features emphasizing dynamic fluctuations, while interactions between multiple sensor channels are modeled through a structured fusion mechanism to enhance fault representation. The proposed approach is evaluated on a benchmark gearbox bearing dataset comprising five health conditions, including normal operation and multiple fault types. Improved classification performance is observed compared to conventional methods, with consistent results under cross-validation, indicating strong generalization capability. Additionally, enhanced feature separability is demonstrated through clearer clustering patterns in low-dimensional projections. The learned representations effectively capture a wide range of signal characteristics, supporting both improved diagnostic performance and practical applicability in industrial monitoring systems.

13:00 JSTエージェント

EVOM: 強化学習のためのアクタークリティックアーキテクチャのエージェント的メタ進化

アクタークリティカル強化学習では、通常、ネットワーク アーキテクチャは手動で設計されます。この設計の自動化は、評価前に各候補をトレーニングする必要があり、設計空間には制限がないため、困難です。これらの課題に対処するために、高性能のアクタークリティカル アーキテクチャを発見するためのエージェント メタ進化フレームワークである EVOM を紹介します。アーキテクチャ検索を 2 レベルの最適化として構成します。内側のループは低忠実度の近接ポリシー最適化 (PPO) を介して重みをトレーニングし、外側のループはアーキテクチャ プログラムを繰り返し調整することでメタ進化を推進します。重要なのは、この外側のループは、ポリシーの実行や環境制御から完全に切り離され、純粋にアーキテクチャ設計者として動作する LLM ベースの設計エージェントによって強化されているということです。実験の結果、EVOM は手動で設計されたベースライン、LLM ガイドによるランダム検索、および最先端の LLM ガイドによるプログラムによるポリシー検索手法である MLES を上回り、Ant-v4 および HalfCheetah-v4 で優れたパフォーマンスを実現することが明らかになりました。アブレーション研究では、メタ進化ループと LLM デザイン エージェントの両方が最終的なパフォーマンスに不可欠であることが検証されています。

原文 (English)

EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning

In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging because each candidate must be trained before evaluation, and the design space is open-ended. To address these challenges, we introduce EVOM, an agentic meta-evolution framework for discovering high-performance actor-critic architectures. We frame architecture search as a bi-level optimization: an inner loop trains weights via the low-fidelity proximal policy optimization (PPO), while an outer loop drives meta-evolution by iteratively refining architecture programs. Crucially, this outer loop is powered by an LLM-based design agent that operates purely as an architecture designer, completely decoupled from policy execution and environment control. Experiments reveal that EVOM outperforms the manually designed baseline, an LLM-guided random search, and the state-of-the-art LLM-guided programmatic policy search method MLES, delivering superior performance on Ant-v4 and HalfCheetah-v4. Ablation studies validate that both the meta-evolution loop and the LLM Design Agent are indispensable for final performance.

13:00 JST研究/論文

ハイブリッド プライバシーを意識したセマンティック検索: SVD で切り詰められたドキュメント ジオメトリと、制限された脅威モデルに基づく CKKS 暗号化クエリの再ランキング

高密度埋め込みはセマンティック検索と検索拡張生成を強化しますが、埋め込み反転攻撃はベクトルからソース テキストを再構築する可能性があります。ベクトル データベースが漏洩すると、その背後にある文書も漏洩します。教科書的な防御策は極端です。検索全体を準同型的に暗号化するのは健全ですが、100 万文書規模では遅すぎます。その一方で、保護するずっと前にプライバシー ノイズによってランキングが低下します。静的コレクションと動的クエリの間の非対称性を利用した中間パスを研究します。コレクションは幾何学的に保護されています。各ベクトルは低次元の SVD 部分空間上で切り詰められ、所有者のみが知っている秘密の直交変換によって回転されます。クエリは暗号的に保護されています。クエリは CKKS 準同型暗号化の下で再ランク付けされるため、正直だが好奇心旺盛なサーバーはクエリやスコアを見ることはありません。 CKKS パラメータは、小規模なオフライン ベンチマークから取得されます。私たちは、保護された部分空間に限定された攻撃者の再構成エラーの厳しい下限を証明します。 100 万のドキュメントと 5 つのエンコーダでは、このスキームは 1 秒未満のレイテンシでランキングの品質を維持し (線形デノイザーとして強力なエンコーダでわずかに向上します)、保護されたスペースに対する既製の反転攻撃はノイズ フロアまで崩壊します。次に、より強力な敵対者をテストします。既知の平文攻撃者は、保持された次元とほぼ同じ数の漏洩ペアから直交プロクラステスによる回転を回復します。公開されている積量子化コードは、最近傍構造を保存します。ランダム投影、校正済みノイズ、および BEIR ベースラインは、切り捨てが無料のデノイザーではなく、エンコーダーに依存する精度コストであることを示しています。私たちは限界を述べています。クエリの機密性は暗号化されていますが、ドキュメントの保護は経験的な難読化レイヤー (SVD の切り捨てと秘密のローテーション) であり、暗号化のプリミティブではありません。また、各主張の脅威モデルを区切ります。

原文 (English)

Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model

Dense embeddings power semantic search and retrieval-augmented generation, but embedding-inversion attacks can reconstruct source text from a vector: when a vector database leaks, the documents behind it leak too. The textbook defences are extremes - encrypting the whole search homomorphically is sound but too slow at million-document scale, while privacy noise degrades ranking long before it protects. We study a middle path exploiting the asymmetry between the static collection and the dynamic query. The collection is protected geometrically: each vector is truncated onto a lower-dimensional SVD subspace and rotated by a secret orthogonal transform known only to the owner. The query is protected cryptographically: it is reranked under CKKS homomorphic encryption, so an honest-but-curious server never sees the query or the scores. CKKS parameters come from a small offline benchmark. We prove a tight lower bound on the reconstruction error of any attacker confined to the protected subspace. On one million documents and five encoders the scheme preserves ranking quality (slightly improving it on strong encoders, as a linear denoiser) at sub-second latency, and an off-the-shelf inversion attack on the protected space collapses to the noise floor. We then test stronger adversaries: a known-plaintext attacker recovers the rotation by orthogonal Procrustes from about as many leaked pairs as the retained dimension; the public product-quantization codes preserve most nearest-neighbour structure; and random-projection, calibrated-noise and BEIR baselines show the truncation is an encoder-dependent accuracy cost, not a free denoiser. We state the limits: query confidentiality is cryptographic, but document protection is an empirical obfuscation layer (SVD truncation plus a secret rotation), not a cryptographic primitive, and we delimit the threat model for each claim.

13:00 JSTLLM/生成AIロボティクス

社会物理的 HRI (spHRI) の成長をグラフ化する: 小さな言語モデルによって強化された体系的なレビュー パイプライン

社会物理的人間とロボットの相互作用 (spHRI) は、ロボット工学、人間とコンピューターの相互作用、人間とロボットの相互作用、および触覚の分野で急速に成長しています。しかし、断片的な用語と一貫性のない方法論により、体系的な統合が困難になっています。スケーラブルなレビューの実践をサポートするために、小規模な言語モデル (SLM; < 1.5B パラメーター) が大規模な spHRI 系統的レビューのタイトルと要約のスクリーニングをどの程度支援できるかを評価しました。人間の査読者のパフォーマンスに匹敵する SLM はありませんでしたが、モデルはローカルで動作し、論文の審査が桁違いに速くなりました。結合された SLM アンサンブルにより、最終的な関連データセットの 10.29% に相当する、査読者が見逃した 39 件の論文が特定されました。これらの結果は、SLM が専門家の査読者を置き換えるのではなく、増強し、大規模な文献レビューをアクセスしやすく持続可能なものにすることができることを示しています。

原文 (English)

Charting the Growth of Social-Physical HRI (spHRI): A Systematic Review Pipeline Augmented by Small Language Models

Social-physical human-robot interaction (spHRI) has grown rapidly across robotics, human-computer interaction, human-robot interaction, and haptics. Yet, fragmented terminology and inconsistent methodologies make systematic synthesis difficult. To support scalable review practices, we evaluated the extent to which small language models (SLMs; < 1.5B parameters) can assist with title and abstract screening for a large spHRI systematic review. While no SLMs matched human reviewers' performance, the models operated locally and screened papers orders of magnitude faster. The combined SLM ensemble identified 39 papers reviewers missed, representing 10.29% of the final relevant dataset. These results demonstrate that SLMs can augment, rather than replace, expert reviewers and make large-scale literature reviews accessible and sustainable.

13:00 JST研究/論文Claude

SOLAR: AI を活用した光速パフォーマンス分析

ディープラーニング モデルはターゲット ハードウェア上でどれくらいの速度で実行できますか? 現在の実装はその限界からどれくらい離れていますか?これらの質問は、ソフトウェア、ハードウェア、アルゴリズムの最適化の中心となります。 Speed-of-Light (SOL) 分析は、特定のアーキテクチャでのワークロードの理論上の最小実行時間を計算することでこれらの課題に答えます。しかし、SOL 境界の導出は依然として手作業であり、エラーが発生しやすく、迅速なモデル開発とは切り離されています。このギャップを埋めるために、PyTorch および JAX ソース コードから検証済みの SOL 境界を自動的に導出するフレームワークである SOLAR を導入します。 SOLAR は、フロー内で生成コンポーネントと決定論コンポーネントの両方を活用します。LLM フロントエンドは、あらゆるソース プログラムを実行可能なアフィン ループ IR に変換し、出力比較によって検証します。決定論的なフローにより、IR が Einsum グラフに引き上げられます。分析バックエンドは、非融合、融合、およびキャッシュ対応の SOL 境界を計算します。 SOLAR は、オペレーターと言語を包括的にカバーし、SOL 違反が観察されない検証済みの境界を生成し、境界を強化して最適化の洞察を明らかにする多重忠実度分析を提供します。 KernelBench、JAX/Flax モデル、ロボット ワークロード全体で SOLAR を評価します。これらの実験では、複数の忠実度レベルでのヘッドルーム分析、最適化機会の特定、クロスプラットフォームの探索、およびインバース ルーフライン ハードウェア プロビジョニングの 4 つのユース ケースを示します。

原文 (English)

SOLAR: AI-Powered Speed-of-Light Performance Analysis

How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are central to software, hardware, and algorithm optimizations. Speed-of-Light (SOL) analysis answers them by computing a workload's theoretical minimum execution time on a given architecture. Yet deriving SOL bounds remains manual, error-prone, and disconnected from rapid model development. To close this gap, we introduce SOLAR, a framework that automatically derives validated SOL bounds from PyTorch and JAX source code. SOLAR leverages both generative and deterministic components in its flow: an LLM frontend translates any source programs into an executable Affine Loop IR, validated by output comparison; a deterministic flow lifts the IR into an einsum graph; and an analytical backend computes unfused, fused, and cache-aware SOL bounds. SOLAR provides comprehensive operator and language coverage, produces validated bounds with zero observed SOL violations, and offers multi-fidelity analysis that tightens bounds and surfaces optimization insights. We evaluate SOLAR across KernelBench, JAX/Flax models, and robotics workloads. These experiments demonstrate four use cases: headroom analysis at multiple fidelity levels, identifying optimization opportunities, cross-platform exploration, and inverse-roofline hardware provisioning.

13:00 JST研究/論文

拡散モデルを用いた海況サンプリング

海洋状態の予測は、運用中の海洋アプリケーションや結合された地球システム モデリングに不可欠ですが、現在のスペクトル波動モデルは、気候シミュレーションへのオンライン結合や確率的 (アンサンブル ベースの) 予測の実行など、多くのユース ケースにとって依然として計算能力が法外です。ディープラーニングは最近、天気予報において優れたパフォーマンスを示していますが、既存の AI ベースの波浪モデルは主に決定論的であり、有義波高などのバルク変数に大きく限定されているため、確率的な海況推定はほとんど解明されていません。この研究では、比較的長い地球規模の風力の履歴 (5 日間) を条件とする、地球規模の海の状態を推定するための拡散ベースの生成モデルを提案します。この生成モデルは、自己回帰的なタイムステップを行わずに、海洋状態の複雑な条件付き分布を直接サンプリングします。従来のアプローチとは異なり、私たちのフレームワークは自然にバルク変数を超えて、ストークスドリフトや平均二乗傾きなどの分配関連変数や派生量を推定します。 30 年間の世界的な WAVEWATCH-III のヒンドキャストに基づいてトレーニングされたこのモデルは、数値スペクトル モデルと比較して大幅な計算高速化を達成しながら、バルク変数の巧みな予測と調整されたアンサンブル スプレッドを提供します。私たちの結果は、拡散ベースの海況サンプリングが、確率論的波浪予測と海況情報のより広範な地球システムモデルへの効率的な結合への有望な道筋を提供することを示唆しています。

原文 (English)

Sampling sea state using a diffusion model

Sea state prediction is essential for operational maritime applications and coupled earth system modeling, yet current spectral wave models remain computationally prohibitive for many use cases, including online coupling to climate simulations and making probabilistic (ensemble-based) predictions. While deep learning has recently demonstrated strong performance in weather forecasting, existing AI-based wave models are predominantly deterministic and largely limited to bulk variables such as significant wave height, leaving probabilistic sea state estimation largely unexplored. In this work, we propose a diffusion-based generative model for global sea state estimation that conditions on a relatively long history (5 days) of global wind forcing. This generative model directly samples the complex conditional distribution of sea state without autoregressive time-stepping. Unlike prior approaches, our framework naturally extends beyond bulk variables to estimate partition-related variables and derived quantities, such as Stokes drift and mean square slope. Trained on a 30-year global WAVEWATCH-III hindcast, the model achieves substantial computational acceleration compared with numerical spectral models while delivering skillful predictions and a calibrated ensemble spread for the bulk variables. Our results suggest that diffusion-based sea state sampling offers a promising path toward probabilistic wave forecasting and efficient coupling of sea state information into broader earth system models.

13:00 JST研究/論文

多目的強化学習のための決定論的パレート最適ポリシー合成

現実世界の意思決定では、多くの場合、複数の矛盾する目標のバランスをとる必要があります。標準的な強化学習 (RL) では、報酬を 1 つのスカラー信号に集約することでこの課題に対処することがよくあります。このアプローチは単純なタスクには効果的ですが、多くの場合、パレート フロンティアとして知られる最適なトレードオフの全領域を捉えることができません。この論文では、多目的マルコフ決定プロセス (MOMDP) の決定論的なパレート最適ポリシーを計算するために設計された、チェビシェフ スカラー化に基づいた新しい優先条件付きベルマン演算子を紹介します。この演算子が、推定値関数が真のパレート フロンティアの上限となる包絡特性を満たすことを証明し、この演算子がこのフロンティアのカバレッジ セットに単調に収束することを示します。さらに、これらの収束した Q 推定値から決定論的なポリシーを抽出する方法も示します。これにより、エージェントは任意の設定に合わせてポリシーを復元でき、合成された各ポリシーがほぼパレート最適のままであることを保証しながら、パレート最適フロンティア全体をキャプチャできます。実験結果は、私たちのアルゴリズムが複雑なトレードオフをうまく回復し、決定論的なパレート最適ポリシー合成のためのソリューションを提供することを検証しました。

原文 (English)

Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal. While effective for simple tasks, this approach often fails to capture the full spectrum of optimal trade-offs, known as the Pareto frontier. In this paper, we introduce a novel preference-conditioned Bellman operator, motivated from the Chebyshev scalarization, designed to compute deterministic Pareto-optimal policies for Multi-Objective Markov Decision Processes (MOMDPs). We prove that this operator satisfies an enveloping property, where the estimated value functions upper-bound the true Pareto frontier, and demonstrate that it monotonically converges to a coverage set of this frontier. Furthermore, we also show how to extract deterministic policies from these converged Q-estimates. This ensures the agent can recover a policy for any given preference, capturing the entire Pareto-optimal frontier while guaranteeing each synthesized policy remains approximately Pareto-optimal. Experimental results validate that our algorithm successfully recovers complex trade-offs, providing a solution for deterministic Pareto-optimal policy synthesis.

13:00 JST研究/論文

フィードフォワード ネットワークを超えて: 次世代 AGI の主体性と本質的安全性の基本的基盤としてのリエントリー ニューラル システム

私たちは、クローズド・リエントリー・ループ(D I サイクル)に基づいた安全な汎用人工知能のための完全なアーキテクチャ青写真を提案します。自己参照できない有向非巡回グラフ (C=0、S=0) であるフィードフォワード ネットワークとは対照的に、提案されたアーキテクチャには、自立増幅 (rho > 1) を伴う構造サイクル (C >= 1) が含まれており、自己モデル、手段的自己保存、およびプログラムされていない目標指向の動作の出現が数学的に保証されています。エージェントの目標は、アーキテクチャ自体で非テキストの D ベクトルとしてエンコードされているため、再解釈やプロンプト インジェクションの影響を受けません。我々は、S>0 が正の積分情報を意味するという機械検証されたリーン 4 証明を備えた、Tononi の NP ハード ファイに代わる多項式時間 [O(N^3)] の計算可能な S 測度を提示します。この研究では、完全な Python/NumPy 実装 (Tarjan ベースのサイクル複雑性、Delta-S バリア)、Apache Kafka と Docker Compose による産業用水平スケーリング、AI 進化の 6 つの時代の分類、将来の再突入アーキテクチャの動物園 (RAS、拡散アトラクター、フラクタル ループ)、安全な群れのためのゲージ不変ネットワーク、フォールト トレランスと回復プロトコル、および 8 つの反証可能プロトコルを提供します。予測。すべての正式な証明は Lean 4 で機械検証されます。このアーキテクチャは現在導入可能であり、AGI に対するトポロジー的に保護された安全設計のアプローチを表しています。

原文 (English)

Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D I cycle). In contrast to feedforward networks, which are directed acyclic graphs (C=0, S=0) incapable of self-reference, the proposed architecture contains a structural cycle (C >= 1) with self-sustaining amplification (rho > 1), mathematically guaranteeing the emergence of a self-model, instrumental self-preservation, and unprogrammed goal-directed behaviour. The agent's goals are encoded as a non-textual D-vector in the architecture itself, making them immune to reinterpretation and prompt injection. We present the S-measure -- a polynomial-time [O(N^3)] computable alternative to Tononi's NP-hard Phi -- with machine-verified Lean 4 proof that S>0 implies positive integrated information. The work provides full Python/NumPy implementations (Tarjan-based cycle complexity, Delta-S barrier), industrial horizontal scaling via Apache Kafka and Docker Compose, a taxonomy of six epochs of AI evolution, a zoo of future reentry architectures (RAS, diffusion attractors, fractal loops), gauge-invariant networks for safe swarms, fault-tolerance and recovery protocols, and eight falsifiable predictions. All formal proofs are machine-verified in Lean 4. This architecture is deployable today and represents a topologically protected, safe-by-design approach to AGI.

13:00 JSTロボティクスハードウェア/半導体

CoStream: 一般化可能な複雑な操作のための単純な動作の構築

GPU を PCIe スロットに装着するなど、長期にわたる接触が多い複雑な操作タスクでは、ミリメートル単位の高精度と、新しいタスクに対するすぐに使える汎用性の両方が必要です。既存のパラダイムは両方を満たすのに苦労しています。古典的なパイプラインは、高精度の制御を実現するために脆弱なタスク固有のインターフェイスを使用しますが、新しいタスクに適応するにはコストのかかるパイプラインの再設計が必要です。一方、モノリシックなエンドツーエンドのポリシーはより優れた一般化を提供しますが、新しいデータで再トレーニングしない限り、複雑な分散外のタスクでは高精度が得られません。どちらのパラダイムも暗黙の前提を共有しています。つまり、操作機能を取得したら、それを自由に分解したり再構成したりするのではなく、厳格なパイプラインまたはモノリシック全体として展開する必要があるということです。この論文では、複雑な操作機能が単純で独立した動作の構成から自然に出現する可能性があることを示します。モノリシックなポリシーや厳格なパイプラインを展開するのではなく、基礎モデルと多様なセンシング モダリティを複数の構成可能なコア動作に統合するフレームワークを提案します。つまり、基礎モデルを介して空間制約を抽出するセマンティック動作。想像上のビデオ内のキーポイントを追跡することで軌道を予測する予測行動。そして高周波の触覚と力の補正を提供する反応的な動作。共有 $SE(3)$ インターフェイスでは、これらの出力は右乗算によって各制御ステップで単一のポーズ コマンドに構成され、準拠したコントローラーによって実行されます。私たちは、日常的な操作と精密な組み立てに及ぶ 8 つの現実世界のタスクを実証し、接触の多い組み立てとオブジェクトの転送で最も大きな効果を発揮し、実行中の手動による混乱からの堅牢な回復を示します。 {ウェブサイト:} https://costream-simple.github.io

原文 (English)

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and out-of-the-box generalization to new tasks. Existing paradigms struggle to satisfy both: classical pipelines use brittle, task-specific interfaces to achieve high-precision control but require costly pipeline redesigns to adapt to new tasks, whereas monolithic end-to-end policies provide better generalization but lack high precision on complex, out-of-distribution tasks unless retrained with new data. Both paradigms share an implicit assumption: once a manipulation capability is acquired, it must be deployed as a rigid pipeline or monolithic whole, rather than being freely decomposed and recomposed. In this paper, we show that complex manipulation capabilities can emerge naturally from the composition of simple, independent behaviors. Rather than deploying a monolithic policy or a rigid pipeline, we propose \ourshort, a framework orchestrating foundation models and diverse sensing modalities into multiple composable core behaviors: a semantic behavior extracting spatial constraints via foundation models; a predictive behavior forecasting trajectories by tracking keypoints in imagined videos; and a reactive behavior providing high-frequency tactile and force corrections. On a shared $SE(3)$ interface, these outputs compose by right-multiplication into a single pose command at each control step, executed by a compliant controller. We demonstrate \ourshort on 8 real-world tasks spanning everyday manipulation and precision assembly, with the strongest gains in contact-rich assembly and object transfer, and show robust recovery from manual perturbations during execution. {Website:} https://costream-simple.github.io

13:00 JSTロボティクス

Play2Perfect: 正確な組み立てのための器用な遊びの事前トレーニングで重要なことは何ですか?

多指ロボットは人間の手のようなスピードと器用さを約束しますが、正確な組み立てなどの困難な問題にはまだ手が届きません。これらのタスクは接触が多いため、模倣学習のためのデータ収集が困難であり、報酬が少ないため、強化学習 (RL) による直接探索が困難になります。その結果、これまでの研究は、特殊なグリッパー、ツールアタッチメント、および環境固定具を使用して問題を構造化することによって進歩しました。この研究では、ロボットが正確な組み立てを完成させる前に、まず遊び方を学ぶ必要があると主張します。さらに、正確な組み立てには、遊び方を学ぶ過程でどのような要素が重要になるのかという質問をします。私たちは、さまざまなオブジェクトや目標でのプレイを通じてタスクに依存しない事前トレーニングを行うための RL フレームワークである Play2Perfect を提案し、その後、正確な組み立てによって完成させます。遊びの目標は、掴むこと、手の中での向きを変えること、ポーズを伸ばすことなど、再利用可能な操作の事前操作を獲得することです。次に微調整は、組み立て前にこの一般的なものを適応させ、成功に必要な最終的な接触が豊富で高精度の相互作用の探索に焦点を当てます。私たちは、オブジェクトの多様性、トレーニングの目的、軌道の多様性、ゴールの精度など、プレーの事前トレーニングにおける主要な設計の選択を体系的に研究します。密度の高い多段階の報酬が提供された場合でも、事前の学習は、ゼロからの RL トレーニングよりも 33 倍サンプル効率が高いことを示します。当社は、ゼロショットのシミュレーションからリアルへの移動を実証し、わずか 0.5 mm の接触クリアランスでタイトな挿入で 60% の成功率を達成し、長時間にわたる複数部品の組み立てとねじ締めで 50% 以上の成功率を達成しました。

原文 (English)

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection for imitation learning difficult, and sparse-reward, making direct exploration with reinforcement learning (RL) intractable. Consequently, prior work has made progress by structuring the problem with specialized grippers, tool attachments, and environment fixtures. In this work, we argue that before a robot can perfect precise assembly, it must first learn to play. We further ask the question: what factors in the process of learning to play matter for precise assembly? We propose Play2Perfect, an RL framework for task-agnostic pretraining through play on diverse objects and goals, which is then perfected on precise assembly. The goal of play is to acquire reusable manipulation priors, such as grasping, in-hand reorientation and pose reaching. Finetuning then adapts this general prior to assembly, focusing exploration on the final contact-rich, high-precision interactions needed for success. We systematically study key design choices in play pretraining, including object diversity, training objective, trajectory diversity, and goal precision. We show that our prior is 33x more sample-efficient than RL training from scratch, even when provided with dense, multi-stage rewards. We demonstrate zero-shot sim-to-real transfer, achieving 60% success on tight insertions with only 0.5 mm contact clearance, and over 50% success on long-horizon multi-part assembly and screwing.

13:00 JSTLLM/生成AI

ConflictScore: 言語モデルが矛盾する証拠をどのように処理するかを特定および測定する

事実性と忠実性に関する既存の指標は、回答が根拠となる文書によって裏付けられているか、矛盾しているかを評価しますが、裏付けとなる証拠と矛盾する証拠の両方が共存する場合を捉えることができません。モデルの応答が基礎文書内の矛盾する証拠をどの程度認識しているかを定量化する新しい指標である ConflictScore を紹介します。私たちのフレームワークは、回答を原子的な主張に分解し、各根拠文書に対する各主張にラベルを付け、これらのラベルを 2 つの補完的な尺度に集約します。ConflictScore-Count (CS-C)、矛盾を示す主張の割合、ConflictScore-Ratio (CS-R)、裏付け証拠と矛盾証拠のバランスです。私たちは、指標を体系的に評価するために、あいまいさ、矛盾、意見の相違など、さまざまな形の対立をカバーするベンチマークである ConflictBench を開発しています。実験によれば、ConflictScore はドメイン間で自信過剰な主張を効果的に検出し、TruthfulQA における真実性を向上させる修正フィードバック メカニズムとして機能することができます。

原文 (English)

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence

Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist. We introduce ConflictScore, a novel metric that quantifies how well a model's response acknowledges conflicting evidence in its grounding documents. Our framework decomposes responses into atomic claims, labels each claim against each grounding document, and then aggregates these labels into two complementary measures: ConflictScore-Count (CS-C), the proportion of claims exhibiting conflicts, and ConflictScore-Ratio (CS-R), the balance between supporting and contradicting evidence. We develop ConflictBench, a benchmark covering diverse forms of conflicts such as ambiguity, contradiction, and divergent opinions, to systematically evaluate our metric. Experiments show that ConflictScore effectively detects overconfident claims across domains and can serve as a corrective feedback mechanism that improves truthfulness on TruthfulQA.

13:00 JST研究/論文

AXLE: リーン 4 定理証明ユーティリティ用のクラウド インフラストラクチャ

Lean 4の証明操作・抽出・検証のためのクラウドサービスAXLE(Axiom Lean Engine)を紹介します。数学用 AI の最近の進歩 (強化学習パイプライン、エージェントによる証明ワークフロー、データセットのキュレーション) には、正確さと堅牢性を維持しながら数百万のリクエストに対応できるリーン 4 ツールが必要です。既存のインフラストラクチャは並列コンパイルを提供しますが、スケーラブルな証明検証、高レベルの証明操作、マルチバージョンのサポート、または最新の AI ワークフローが必要とするスループットでのリクエストごとの分離は提供しません。 AXLE は、厳密な証明の検証、宣言メタデータの抽出、意味論的なソース操作、決定論的な証明の修復と簡略化、補題抽出に及ぶ 14 のリーン 4 メタプログラミング ツールを提供します。このサービスは、リクエストごとの分離と複数の Lean 4 および Mathlib バージョンの同時サポートを備えたマルチテナント クラウド デプロイメントとして実行され、Python SDK、コマンドライン インターフェイス、Web UI、MCP サーバー、および raw HTTP API を介してアクセスできます。 AXLE は https://axle.axiommath.ai および axiom-axle PyPI パッケージ経由で公開されており、ローカルに Lean 4 をインストールする必要がなく、無料で使用できます。これまでに 5 億件を超えるリクエストに対応しており、2025 年のパトナム コンテストでの 12/12 スコアを含む、Axiom Math の証明作業の基礎となるインフラストラクチャです。

原文 (English)

AXLE: A Cloud Infrastructure for Lean 4 Theorem Proving Utilities

We present AXLE (Axiom Lean Engine), a cloud service for Lean 4 proof manipulation, extraction, and verification. Recent progress in AI for mathematics -- reinforcement learning pipelines, agentic proving workflows, dataset curation -- demands Lean 4 tooling that scales to millions of requests while remaining correct and robust; existing infrastructure offers parallel compilation but not scalable proof verification, higher-level proof manipulation, multi-version support, or per-request isolation at the throughput modern AI workflows require. AXLE provides 14 Lean 4 metaprogramming tools spanning strict proof verification, declaration metadata extraction, semantic source manipulation, deterministic proof repair and simplification, and lemma extraction. The service runs as a multi-tenant cloud deployment with per-request isolation and concurrent support for multiple Lean 4 and Mathlib versions, accessible via a Python SDK, command-line interface, web UI, MCP server, and raw HTTP API. AXLE is publicly available and free to use at https://axle.axiommath.ai and via the axiom-axle PyPI package, with no local Lean 4 installation required. It has served over 500 million requests to date and is the underlying infrastructure for Axiom Math's proving efforts, including its 12/12 score on the 2025 Putnam competition.

13:00 JST画像/動画生成ロボティクス研究/論文Gemini

WatchAct: 行動に基づいたロボット操作のベンチマーク

人間と一緒に働くロボットは、人間が何を、どの順序で、どのような意図で行ったかを推論しなければなりません。ビデオには、言語では仕様が不十分なままになっている空間レイアウト、オブジェクト履歴、ジェスチャーが含まれていますが、今日の操作ベンチマークは命令と単一の現在の画像を組み合わせており、観察された人間の行動に対する推論を評価する方法を提供していません。観察された人間の行動に基づいたロボット操作のベンチマークである WatchAct を紹介します。各インスタンスは、現実世界の人間のアクションビデオと言語命令を、調整されたシミュレーターシーンと実行可能なLIBEROタスクと組み合わせて、スケーラブルで再現可能な評価を可能にします。 WatchAct は、他のエージェントを監視する認知的要求から抽出された 4 つの機能ドメインの 14 のタスクにわたる 3,000 の長期インスタンスで構成されています。イベントの解析 (イベント グランディング)、手続き構造の回復 (手続き推論)、暗黙の意図の推論 (暗黙的意図推論)、およびシーンがどのように変更されたかの追跡 (エピソード推論) です。我々はさらに、(i)〜ビジョン言語モデルによるビデオからプランへの推論、(ii)〜オラクルプランに基づくポリシーの実行、および(iii)〜統合プランナーによる完全なタスクの完了〜ポリシーパイプラインを個別に測定する、絡み合っていない評価プロトコルを提案します。シミュレーションでも、Franka Research 3 ロボットでも、現在のシステムは WatchAct を解決するには程遠いです。 $\pi_{0.5}$ の最高のパイプラインである Gemini-3.1-Pro は、シミュレーションでは成功率 (SR) が 16.3%、実際のロボットでは 14.0% にすぎません。 Gemini-3.1-Pro のプラン SR はわずか 36.8% (人間の場合は 97.1%) ですが、$\pi_{0.5}$ は、オラクル プランではタスク SR の 21.5% にのみ達し、ドメイン外のシナリオでは 10.6% に低下します。データセットとコードは https://baiqi-li.github.io/watchact_page/ で入手できます。

原文 (English)

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real-world human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task, enabling scalable and reproducible evaluation. WatchAct comprises 3,000 long-horizon instances across 14 tasks in four capability domains drawn from the cognitive demands of watching another agent: parsing events (Event Grounding), recovering procedural structure (Procedural Reasoning), inferring unstated intent (Implicit Intent Inference), and tracking how the scene was changed (Episodic Reasoning). We further propose a disentangled evaluation protocol that separately measures (i)~video-to-plan reasoning by vision-language models, (ii)~policy execution under oracle plans, and (iii)~full task completion by integrated planner--policy pipelines. In both simulation and on a Franka Research 3 robot, current systems remain far from solving WatchAct. The best pipeline, Gemini-3.1-Pro with $\pi_{0.5}$, reaches only 16.3% Success Rate (SR) in simulation and 14.0% on the real robot. Gemini-3.1-Pro attains just 36.8% Plan SR (vs. 97.1% for humans), while $\pi_{0.5}$ reaches only 21.5% Task SR under oracle plans and drops to 10.6% on out-of-domain scenarios. Dataset and code are available at https://baiqi-li.github.io/watchact_page/.

13:00 JSTエージェント

自動認知科学者による心理理論発見のためのループを閉じる

科学全般にわたって、自律システムはクローズドループの発見、新しい理論の提案、それらをテストするための実験の設計と実行にますます使用されています。このアプローチは、認知科学の分野ではまだ適用されていません。認知科学の分野では、中心的なボトルネックは理論構築、つまり既存のモデルの蓄積された失敗をより良いモデルに変える創造的なステップです。データ収集、モデリング、実験計画が自動化されているにもかかわらず、理論生成は手動のままです。私たちは、このループを閉じる完全自律型エージェント AI システムである Automated Cognitive Scientist (AutoCog) を紹介します。大規模言語モデルのエージェントは、それぞれが実行可能な認知モデルとして表現される競合する理論を提唱し、それらを最もよく区別する実験を計画し、オンラインで募集した参加者から行動データを収集し、生成パフォーマンスに基づいて収集したデータに対して理論を採点し、失敗の理由を診断し、より優れた後継理論を合成します。このサイクルを繰り返すことで、理論、モデル、実験の空間を探索することができます。意思決定の領域では、AutoCog は、型破りなものを含む、シミュレーションされた動作から既知の意思決定戦略を復元しました。これは、その発見が、基礎となる言語モデルの事前分布に厳密に束縛されるのではなく、最終的にはデータによって推進されることを示しました。人間の参加者を対象に実行すると、2 つの異なる実験設定でシードされ、保留された研究に一般化された確立された理論を上回る理論が生成されました。また、選択によって特徴値に対する感度が低下するという、複数の手がかりによる意思決定に関する新しい理論も明らかになりました。この理論の特徴的な予測は、新規参加者を対象とした事前登録研究で確認されました。 AutoCog は、自動化された発見システムを使用して、認知理論の構築を明示的で実行可能な累積的な科学に変える方法を示します。

原文 (English)

Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist

Across the sciences, autonomous systems are increasingly being used in closed-loop discovery, proposing new theories and designing and running experiments to test them. This approach is yet to be applied in the field of cognitive science, where the central bottleneck is theory-building: the creative step of turning the accumulated failures of existing models into better ones. Theory generation has remained manual even as data collection, modeling, and experiment design have been automated. We present the Automated Cognitive Scientist (AutoCog), a fully autonomous agentic-AI system that closes this loop. Large-language-model agents advocate competing theories, each expressed as an executable cognitive model, design experiments that best discriminate them, collect behavioral data from participants recruited online, score theories against collected data based on their generative performance, diagnose why they fail, and synthesize a better successor. Repeating this cycle allows them to search the space of theories, models, and experiments. In the domain of decision-making, AutoCog recovered known decision-making strategies from simulated behavior, including unconventional ones, showing that its discoveries are ultimately driven by the data rather than strictly bound by the priors of the underlying language models. When run with human participants, it produced theories that outperformed the established theories it was seeded with and generalized to held-out studies in two different experimental settings. It also surfaced a novel theory of multi-cue decision-making in which choices show diminishing sensitivity to feature values. The distinctive predictions of this theory were confirmed in a preregistered study with new participants. AutoCog demonstrates how an automated discovery system can be used to turn cognitive theory-building into an explicit, executable, and cumulative science.

13:00 JSTLLM/生成AI

ProvenAI: 生成された回答における来歴ネイティブの証拠の痕跡

検索拡張システムは、生成された回答とともに引用を日常的に提示しますが、引用は、対応する情報源が出力を意味のある形で形成したことを確認するものではありません。この論文では、マルチホップ質問応答における透明性を、回答の正確性、ベンチマークを裏付ける証拠に対する引用忠実度、およびリーブ 1 リソース アウト介入下での文書ごとの影響という 3 つの独立して測定可能なレイヤーに分解するフレームワークである ProvenAI を紹介します。データ正規化、検索インデックス作成、引用を意識した回答生成、帰属監査、アブレーションベースの影響推定、バッチ評価、対話型検査をカバーする 7 段階のパイプラインを通じて HotpotQA ディストラクタ ベンチマークをターゲットとして、ProvenAI は 509,300 パッセージの正規コーパスから抽出された 7,405 の検証例を評価します。このシステムは、回答精度 53.53% と平均引用忠実度スコア 71.55% を達成しており、実際の例では引用と影響のギャップと呼ばれるものが明らかになりました。クリーンな引用監査は、1 つの引用情報源が弱い影響しか記録しない一方で、7 つの引用されていない情報源が出力を明らかにシフトさせるプロファイルと同時発生します。私たちは、実装された表面プロキシとトークンレベルの KL ダイバージェンス ターゲットとの関係を、規定された忠実性条件を通じて形式化し、因果媒介分析とデータベース出所理論のフレームワークを基礎付け、自律的な科学的発見で出現する暗号の出所アーキテクチャで 3 つの測定層がどのように構成されるかを議論します。 ProvenAI は、検索に基づいた QA における意味のある透明性には、検索、引用、行動に影響を与える証拠を 3 つの個別に測定されたレイヤーとして追跡可能なリンクが必要であることを確立しています。

原文 (English)

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source meaningfully shaped the output. This paper introduces ProvenAI, a framework that decomposes transparency in multi-hop question answering into three independently measurable layers: answer correctness, citation fidelity against benchmark supporting evidence, and per-document influence under leave-one-resource-out intervention. Targeting the HotpotQA distractor benchmark through a seven-stage pipeline covering data normalisation, retrieval indexing, citation-aware answer generation, attribution auditing, ablation-based influence estimation, batch evaluation, and interactive inspection, ProvenAI evaluates 7,405 validation examples drawn from a canonical corpus of 509,300 passages. The system achieves 53.53% answer accuracy alongside a mean citation-fidelity score of 71.55%, and a worked example surfaces what we call the citation-influence gap: a clean citation audit co-occurring with a profile in which one cited source registers only weak influence while seven uncited sources demonstrably shift the output. We formalise the relationship between the implemented surface proxy and a token-level KL-divergence target through a stated faithfulness condition, ground the framework in causal-mediation analysis and database-provenance theory, and discuss how the three measurement layers compose with cryptographic provenance architectures emerging in autonomous scientific discovery. ProvenAI establishes that meaningful transparency in retrieval-grounded QA requires traceable links across retrieved, cited, and behaviourally influential evidence as three distinct, independently measured layers.

13:00 JST画像/動画生成

RGB イベント視覚オブジェクト追跡のためのアクティブな敵対的摂動駆動の連想メモリ検索

RGB イベント トラッキングは、RGB 外観テクスチャとイベント センサーからの高密度の時間的モーション キューを融合することにより、ローカリゼーションの堅牢性を向上させます。このマルチモーダル方式によりトラッキングの適用範囲が広がりますが、現実世界のシーンでは、従来のマルチモーダル融合を妨げる多様な構造化信号劣化が発生しています。過酷な環境では、どちらのモダリティも信頼性を大幅に失う可能性があり、オクルージョン、エッジの切り詰め、および前景のクラッターによりターゲットが不完全に見えることがよくあります。上記の課題に取り組むために、部分的なターゲットの欠落やモーダル劣化に対する堅牢性を備えた RGB イベント追跡用に調整された階層摂動および取得フレームワーク (APRTrack と呼ばれます) を紹介します。現実世界の信号破損を模倣するために、APRTrack はモダリティおよび空間レベルで 2 つの敵対的摂動ブランチを介して構造化劣化を構築し、フルモーダル障害と局所的なターゲット領域の欠如を個別にシミュレートします。階層型ルーティング メカニズムは、2 つの摂動タイプのトレーニング パイプラインのもつれを解消し、重畳された劣化制約によって引き起こされる機能の崩壊を効果的に排除するように設計されています。さらに、信頼性の高い履歴情報補償のために、フットプリントガイド付きチャネル校正ホップフィールド検索 (FCHR) を考案します。このモジュールは、クエリとメモリ バンク間のアソシエーション フットプリントに基づいて検索の信頼性を評価し、ホップフィールド マッチングの前に検索メトリック空間を調整して、ターゲット領域に限定された制御可能な履歴特徴補償を実現します。 FE108、COESOT、VisEvent、および FELT データセットに関する広範な実験により、RGB イベント視覚オブジェクト追跡に対して提案された戦略の有効性が実証されています。ソース コードと事前トレーニングされたモデルは https://github.com/Event-AHU/OpenEvTracking でリリースされます。

原文 (English)

Active Adversarial Perturbation-driven Associative Memory Retrieval for RGB-Event Visual Object Tracking

RGB-Event tracking improves localization robustness by fusing RGB appearance textures and dense temporal motion cues from event sensors. While this multi-modal scheme broadens tracking applicability, real-world scenes suffer diverse structured signal degradations that hinder traditional multi-modal fusion. In harsh environments, either modality can lose reliability drastically, and targets frequently appear incomplete due to occlusion, edge truncation and foreground clutter.To tackle the above challenges, we present a hierarchical perturbation and retrieval framework tailored for RGB-Event tracking with robustness against partial target missing and modal degradation, termed APRTrack. To mimic real-world signal corruption, APRTrack constructs structured degradation via two adversarial perturbation branches at the modality and spatial levels, which separately simulate full-modal failure and localized target region absence. A hierarchical routing mechanism is designed to disentangle the training pipelines of the two perturbation types, effectively eliminating feature collapse induced by superimposed degradation constraints. Furthermore, we devise Footprint-guided Channel-calibrated Hopfield Retrieval (FCHR) for reliable historical information compensation. This module evaluates retrieval confidence based on association footprints between queries and memory banks, and calibrates the retrieval metric space prior to Hopfield matching, realizing controllable historical feature compensation bounded to target regions. Extensive experiments on FE108, COESOT, VisEvent, and FELT datasets demonstrate the effectiveness of our proposed strategies for the RGB-Event visual object tracking. The source code and pre-trained models will be released on https://github.com/Event-AHU/OpenEvTracking

13:00 JST研究/論文

3D空間パターンマッチング

空間パターン マッチングは、クエリ エンティティおよび制約をデータベース エンティティおよびリレーションと照合するプロセスです。類似地域検索、住宅市場検索、ランドマーク検索、道路網マッチングなど、多くの用途があります。私たちの知る限り、既存の空間パターン マッチングのアプローチはすべて 2 次元空間で問題を構成しており、エンティティはデカルト平面内にあり、エンティティ間に定義された関係は 2 次元に含まれています。ただし、位置に加えて高さがある現実世界のエンティティを検索する場合、この問題のフレーミングには大きな制限があります。この制限に対処するために、空間パターン マッチングを 3 次元に拡張し、問題の一般化された定義を提供します。私たちは、距離関係に基づいて 3D 空間パターンを解決できるサブグラフ マッチング アルゴリズムを説明し、2 つの 3D 空間パターン マッチング データセット (合成データセットと、ドイツのハンブルク市の実際の 3D 建物データを含むデータセット) をリリースします。両方のデータセットでサブグラフ マッチング アルゴリズムをテストし、将来の手法を構築するためのベースラインとして結果を提示します。

原文 (English)

3D Spatial Pattern Matching

Spatial pattern matching is the process of matching query entities and constraints with database entities and relations. It has many applications, including similar region search, housing market search, landmark search, and road network matching. To our knowledge, all existing spatial pattern matching approaches frame the problem in a 2 dimensional space, where entities lie in a cartesian plane and relationships defined between them are contained in 2 dimensions. However, this problem framing has significant limitations when searching for real world entities that have height in addition to position. To address this limitation, we extend spatial pattern matching to 3 dimensions and provide a generalized definition of the problem. We describe a subgraph matching algorithm capable of resolving 3D spatial patterns over distance relations and release two 3D spatial pattern matching datasets, one synthetic and one containing real 3D building data from the city of Hamburg, Germany. We test our subgraph matching algorithm on both datasets and present results as a baseline for future methods to build upon.

13:00 JSTエージェント

RL によるツールの使用を単一のクロスコーダー機能にローカライズする

RL による微調整により、言語モデルの内部表現が再形成され、ツールの使用などのエージェント的な動作が可能になりますが、これらの変更のメカニズムの基礎は依然としてよく理解されていません。 RL は構造化されたツール呼び出しの生成を大幅に改善しますが、どの特徴が出現し、どの特徴が保存されるか、また識別された特徴が再トレーニング不要の行動制御に活用できるかどうかは不明です。この研究では、$\textit{専用機能クロスコーダー (DFC)}$ が、$\texttt{Qwen2.5-3B}$ のツール呼び出し機能を仲介する RL 固有の機能のコンパクトなセットを分離することを示します。 $48$ クロスコーダのハイパーパラメータ スイープ全体で、エンコード/デコード再構成により、RL モデルのツールの正確性が $+31.1 \pm {9.7}$ pp 向上し、ツール呼び出し能力が $+6.8 \pm 5.0$ pp だけフリーズしたベース モデルに受動的に転送されます。これを $\textit{能力スピルオーバー}$ と呼びます。私たちの調査結果は、DFC パーティショニングにより、RL によって導入された機能が最小限の操作可能な機能セットに集中され、エージェント LLM の実行時の動作制御が可能になることが示されています。

原文 (English)

Localizing RL-Induced Tool Use to a Single Crosscoder Feature

Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basis of these changes remains poorly understood. While RL substantially improves structured tool-call generation, it is unclear which features emerge, which are preserved, and whether identified features can be leveraged for retraining-free behavioral control. In this work, we show that $\textit{Dedicated Feature Crosscoders (DFC)}$ isolate a compact set of RL-specific features that mediate tool-calling capability in $\texttt{Qwen2.5-3B}$. Across a $48$-crosscoder hyperparameter sweep, encode-decode reconstruction improves the RL model's tool correctness by $+31.1 \pm {9.7}$ pp and passively transfers tool-calling ability to the frozen base model by $+6.8 \pm 5.0$ pp which we call a $\textit{capability spillover}$. Our findings show that DFC partitioning concentrates RL-introduced capability into a minimal, steerable feature set that enables runtime behavioral control of agentic LLMs.

13:00 JST研究/論文

検索加温エネルギーベース推論:構造化推論タスクにおける推論としての拡散のための 5 アーム アブレーション方法論

ウォームスタートされた拡散サンプラーは反復推論を加速しますが、パイプラインのどの部分がゲインをもたらすのかが明確になることはほとんどありません。私たちは \textbf{検索加温エネルギーベース推論 (RW-EBR)} -- モダンホップフィールド軌跡メモリで強化された IRED エネルギーベース拡散モデル \cite{du2024ired} -- を研究し、次の 3 つの交絡した効果を分離する \textbf{5 アームアブレーション方法論} (オラクル、ベストコンスタント、クエリごとのランダム、シャッフル、整列) に貢献します。クラス優先バイアス シフト、確率的ウォーム スタート、グラフに合わせた値の再利用。診断分解は、LLM-RAG 評価 \cite{ru2024ragchecker} から適応されています。 \textbf{connectivity-2} (Erd\H{o}s--R\'enyi 全ペアの到達可能性) では、整列対シャッフルオラクルのスイングは、固定 1{,} グラフ検証セット診断で \textbf{$+35$\,pp} のバランスのとれた精度に達します。値の分散と取得の仕組みは修正され、グラフごとの整列のみが破壊されますが、クエリごとのランダムな初期化はコールド以下になります --バイアスシフトや確率論ではなく、グラフごとの調整が支配的です。しかし、\emph{deployable} コールド予測パイプラインは、保存された値の品質では受け入れゲートを通過できません。キー品質画面で停止した同じ診断ロジックを、タスク固有のキー エンコーダーを使用して \textbf{Sudoku} に適用すると、\emph{異なる} コンポーネント (現在の設定ではキー品質) で明確な否定結果が生成されます。分解により、各タスクの最初のブロック コンポーネントに名前が付けられます。この設定 (故障モードの説明可能性をレンズとして使用し、反復拡散サンプラーによって洗練されたグラフの到達可能性) により、構造化された時空間的推論の範囲内に作業が配置されます。

原文 (English)

Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks

Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We study \textbf{retrieval-warmed energy-based reasoning (RW-EBR)} -- an IRED energy-based diffusion model \cite{du2024ired} augmented with a Modern Hopfield trajectory memory -- and contribute a \textbf{five-arm ablation methodology} (oracle, best-constant, per-query-random, shuffled, aligned) that separates three confounded effects: class-prior bias shift, stochastic warm-starting, and graph-aligned value reuse. The diagnostic decomposition is adapted from LLM-RAG evaluation \cite{ru2024ragchecker}. On \textbf{connectivity-2} (Erd\H{o}s--R\'enyi all-pairs reachability), the aligned-vs-shuffled-oracle swing reaches \textbf{$+35$\,pp} balanced accuracy on a fixed 1{,}000-graph validation-set diagnostic, with value distribution and retrieval mechanics fixed, only per-graph alignment destroyed, while per-query random initialisation falls below cold -- per-graph alignment, not bias shift or stochasticity, dominates. Yet the \emph{deployable} cold-prediction pipeline misses the acceptance gate at stored-value quality. The same diagnostic logic, stopped at the key-quality screen, applied to \textbf{Sudoku} with a task-specific key encoder produces a clean negative at a \emph{different} component -- key quality, under the current setup. The decomposition names the first blocking component on each task. The setting -- graph reachability refined by an iterative diffusion sampler, with explainability of failure modes as the lens -- places the work within structured and spatio-temporal reasoning.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

LLM エージェントの即時注入に対する帯域外防御の適応的評価

最近の研究 (2024 年から 2026 年) は、ツールを使用する LLM エージェントを間接的なプロンプト インジェクションから防御するための戦略に収束しました。悪意のある命令を拒否するようにモデルをトレーニングするのではなく、エージェントのアクションを仲介する決定論的なポリシーでモデルの外側にセキュリティを強化します。 CaMeL、FIDES、Progent、RTBAS、FORGE などのシステムは、機能、情報フロー ラベル、参照モニターによってこれを実現しており、いくつかのシステムは AgentDojo ベンチマークに対する攻撃がほぼ排除されたと報告しています。私たちは 2 つの貢献をします。まず、これらの帯域外防御を古典的な完全性保護 (Biba)、参照モニタリング、最小特権のインスタンスとして整理し、それらがカバーするものとカバーしないものを構造化して比較します。第 2 に、それらのすべては静的ベンチマーク (一連の固定された注入試行) でのみ検証されることを警告します。これは、適応的で防御を意識した攻撃が 90% 以上の成功率でそのうち 12 件を破るまで、帯域内防御が強力であるように見えたのと同じ方法論です。適応的評価に必要な脅威モデルとプロトコルを指定します。次に、そのプロトコルを、Progent 独自の適応型攻撃分析の独立した再現および拡張として、単一の H200 上で自己ホストされるオープンウェイト エージェント (Qwen2.5-7B) を使用して AgentDojo 上で実行します。この設定は作成者がテストしていません。 3 回のランを平均すると、ディフェンスは次の結果を維持しました。プロジェントは平均攻撃成功率を約 6 倍 (25.8% から 4.2%) に削減しましたが、手作りの適応攻撃では平均攻撃成功率は上がりませんでした (2.6%)。これは、単一のブラックボックス攻撃テンプレートを備えた弱いモデル上の 1 つの小規模なデータ ポイントです。より強力な最適化された (ホワイトボックス GCG) 攻撃が依然として存在します。この結果は、適応型攻撃者にとって、決定論的な帯域外強制の方が帯域内検出よりも難しいターゲットであるという仮説と一致しますが、それを確立するものではありません。

原文 (English)

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions. Systems such as CaMeL, FIDES, Progent, RTBAS, and FORGE realize this with capabilities, information-flow labels, and reference monitors, and several report near-elimination of attacks on the AgentDojo benchmark. We make two contributions. First, we organize these out-of-band defenses as instances of classical integrity protection (Biba), reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover. Second, we warn that every one of them is validated only on static benchmarks (a fixed set of injection attempts), the same methodology that made in-band defenses look strong until adaptive, defense-aware attacks broke twelve of them at over 90% success; we specify the threat model and protocol an adaptive evaluation requires. We then run that protocol as an independent reproduction and extension of Progent's own adaptive-attack analysis, on AgentDojo, with an open-weight agent (Qwen2.5-7B) self-hosted on a single H200, a setting its authors did not test. Averaged over three runs, the defense held: Progent cut mean attack success roughly sixfold (25.8% to 4.2%), and a hand-crafted adaptive attack did not raise it (2.6%). This is one small-scale data point on a weak model with a single black-box attack template; a stronger optimized (white-box GCG) attack remains open. The result is consistent with, but does not establish, the hypothesis that deterministic out-of-band enforcement is a harder target for an adaptive attacker than in-band detection.

13:00 JSTLLM/生成AI

LLM に数値を話す: 時系列予測のためのマルチ ウェーブレット数値埋め込み

大規模言語モデル (LLM) は、異種のテキスト信号を統合できるため、コンテキストを認識した時系列予測にとって魅力的ですが、その離散的な言語指向のトークン化および埋め込みインターフェイスは連続数値とずれており、多くの場合、数値の順序付けと予測の信頼性に悪影響を及ぼします。我々は、各スカラー観測をマルチ ウェーブレット、マルチスケール係数から構築されたディジット単位の埋め込みにマッピングする、プラグ アンド プレイの時間ウェーブレット ディジット インターフェイスである TempoWave を提案します。標準トークン表現を直接オーバーライドすることで、TempoWave は、きめの細かいローカル変動とマクロ グローバル構造の両方をトランスフォーマーと互換性のある形式でシームレスに公開し、正確な数値書式設定、明確な数字の同一性、一般的な正規化操作に対する堅牢性が LLM パイプライン全体で維持されることを保証します。 5 つのコンテキスト強化予測ベンチマークにわたる実験では、TempoWave が標準の数値トークン化や代替埋め込みインターフェイスよりも LLM ベースの予測機能を一貫して改善し、新しい最先端を達成していることが実証されました。これらの結果は、主要なボトルネックとして数値インターフェイスを強調し、原則に基づいた多重解像度の埋め込みが LLM のコンテキスト推論と正確な予測をより適切に結び付けることができることを示唆しています。私たちのコードは https://github.com/DC-research/TempoWAVE で入手でき、モデルは https://huggingface.co/Melady/TempoWAVE でアクセスできます。

原文 (English)

Speaking Numbers to LLMs: Multi-Wavelet Number Embeddings for Time Series Forecasting

Large language models (LLMs) are attractive for context-aware time series forecasting because they can integrate heterogeneous textual signals, yet their discrete, language-oriented tokenization and embedding interfaces are misaligned with continuous numerical values, often harming numerical ordering and forecasting reliability. We propose TempoWave, a plug-and-play temporal wavelet digit interface that maps each scalar observation into digit-wise embeddings constructed from multi-wavelet, multi-scale coefficients. By directly overriding standard token representations, TempoWave seamlessly exposes both fine-grained local fluctuations and macro global structures in a transformer-compatible form, ensuring that precise numerical formatting, distinct digit identity, and robustness to common normalization operations are maintained throughout the LLM pipeline. Experiments across five context-enriched forecasting benchmarks demonstrate that TempoWave consistently improves LLM-based forecasters over standard numeric tokenization and alternative embedding interfaces, achieving a new state-of-the-art. These results highlight the numeric interface as a key bottleneck and suggest that principled multi-resolution embeddings can better couple LLMs' contextual reasoning with precise forecasting. Our code is available at https://github.com/DC-research/TempoWAVE and our model can be accessed at https://huggingface.co/Melady/TempoWAVE.

13:00 JSTLLM/生成AIGemini

LLM によって生成された VeriFast の仕様に関する実証的研究

静的検証ツールは産業規模のソフトウェアを保証できますが、仕様を作成するには多大な労力が必要です。これは、分離ロジックに基づく静的ベリファイア (SL ベリファイア) に特に当てはまります。SL ベリファイアは、ヒープ操作プログラムの検証には優れていますが、ヒープ構造を推論するには多くの複雑な補助仕様が必要です。最近の研究では、大規模言語モデル (LLM) を適用してコード、テスト、証明を生成します。これには検証者の仕様も含まれますが、主に非 SL 検証者を対象としています。このギャップに対処するために、この文書では、SL 検証ツール VeriFast を使用して 303 C 機能を検証するための仕様を生成するように求められたときに、LLM がどの程度うまく機能するかを徹底的に評価します。私たちは 8 つのプロンプト アプローチ、10 の LLM、および 3 つの入力タイプを 2 段階で調査しました。定量的分析と定性的分析を使用して、LLM で生成されたコードと機能の動作、検証可能性、エラーの仕様を評価します。結果は、LLM がソース コードと仕様の機能動作を保持している (両方とも 91% 以上) ものの、検証の成功率はわずか (31.4%) であることを示しています。 Gemini 2.5 Pro を使用し、正式な契約を締結することで、当社の環境での成功率が高くなります。さらに、ほとんどのエラー (94%) は、VeriFast などの SL 検証者のドメイン固有の知識における LLM の間違いに起因しています。これらの調査結果は、SL 検証者向けに LLM で生成された仕様を最適化するためのガイダンスを提供します。

原文 (English)

An Empirical Study of LLM-Generated Specifications for VeriFast

Static verification tools can assure industrial scale software, but require significant human labor to write specifications. This is particularly true of static verifiers based on separation logic (SL verifiers), which excel at verifying heapmanipulating programs, but require many complex auxiliary specifications to reason about heap structure. Recent work applies large language models (LLMs) to generate code, tests, and proofs, including specifications for verifiers, but mostly targeting non-SL verifiers. To address this gap, this paper thoroughly evaluates how well LLMs perform when prompted to generate specifications for verifying 303 C functions with the SL verifier VeriFast. We explored eight prompting approaches, ten LLMs, and three input types in two stages. Quantitative and qualitative analyses are used to assess the LLM-generated code and specifications for functional behavior, verifiability and errors. The results show that LLMs preserve functional behavior in source code and specifications (both over 91%), but achieve modest verification success (31.4%). Using Gemini 2.5 Pro and providing formal contracts lead to higher success rates in our setting. Moreover, most errors (94%) come from LLMs' mistakes in the domainspecific knowledge of SL verifiers such as VeriFast. These findings provide guidance for optimizing LLM-generated specifications for SL verifiers.

13:00 JSTビジネス/資金調達

深層学習プログラムの故障診断における評価と戦略のギャップ

深層学習 (DL) プログラムはさまざまな理由でトレーニング中に失敗する可能性があり、原因の診断はコ​​ストと時間のかかるメンテナンス作業です。このような障害を診断する技術は、通常、プログラム内の相互検証を使用して評価されますが、これまでに確認されていないプログラムが関与する展開設定には不適切な場合があります。したがって、これらの設定間でパフォーマンスがどのように異なるかを評価し、確立された DL の障害診断技術におけるパフォーマンス ギャップの原因を特定する必要があります。私たちは、38 の実世界の DL プログラムからの 5,542 個のフォールト注入トレーニング トレースのコーパスである DynFault を使用して、このギャップを調査します。既存の障害診断技術では、プログラム内での評価とプログラム全体を実行した場合のバランスの取れた精度に 0.190 のギャップがあることがわかりました。また、このギャップは機能のプログラム レベルの構造に起因することもわかり、2 つのランタイム機能セット、曲率機能とオプティマイザー機能、および目に見えないプログラムでのそれらの動作を調査することになりました。曲率機能は目に見えないプログラムの不安定性の検出に役立ちますが、オプティマイザーとアクティベーション機能はトレーニング中に表示されたプログラムにのみ役立ちます。

原文 (English)

Evaluation-Strategy Gap in Fault Diagnosis of Deep Learning Programs

Deep Learning (DL) programs can fail during training for many reasons, and diagnosing the cause is a costly and time-consuming maintenance task. Techniques for diagnosing such failures are commonly assessed using within-program cross-validation, which may be inadequate for deployment settings involving previously unseen programs. It is therefore necessary to assess how performance differs across these settings and to identify the causes of any performance gap in established fault diagnosis techniques for DL. We investigate this gap using DynFault, a corpus of 5,542 fault-injected training traces from 38 real-world DL programs. We found a gap of 0.190 in balanced accuracy for existing fault diagnosis techniques between within-program evaluation and holding out whole programs. We also found the gap comes from program-level structure in the features, which led us to examine two runtime feature sets, curvature features and optimizer features, and their behavior on unseen programs. We found that curvature features are useful for instability detection on unseen programs, while optimizer and activation features help only on programs seen during training.

13:00 JSTLLM/生成AIエージェント

検索メモリの時間的妥当性: 知識の進化による AI エージェントの古い事実エラーの排除

検索拡張生成 (RAG) により、エージェントは蓄積された知識にアクセスできますが、時間のモデルはありません。ファクトが変更されると (たとえば、関数の名前が変更されたり、API が再構築されたり)、RAG は、ほぼ同一の埋め込み類似性を持つ古い値と現在の値の両方を取得します。その後、代理人は棄権するか、置き換えられた事実を提出します。これは構造的な問題であることを示します。キャリブレーションされたデータセットでは、コサイン類似度によって、AUROC 0.59 (ほぼ偶然) で矛盾した事実と重複した事実が区別されます。これは、矛盾は言い換えられた重複よりも元の事実に埋め込み類似していることが多いためです。時間的妥当性を維持した検索メモリである MemStrata を紹介します。 RAG のようにファクトを保存し、静的な再現を維持しますが、ファクトの値が矛盾する場合、決定論的 (サブジェクト、リレーション、オブジェクト) の置き換えルールにより、バイテンポラル台帳内の古い値が削除されます。類似性しきい値や LLM 呼び出しは使用されません。 7B モデルを使用してローカルで実行された 6 つのベンチマーク全体で、MemStrata は静的知識で RAG と結びつき、進化する知識では 0.95 ~ 1.00 の精度に達しました (RAG は 0.20 ~ 0.47 に達します)。中心的な結果は、古い事実のエラー率です。応答が必要な場合、RAG は 15 ~ 40% の確率で置き換えられた値を提供します。 MemStrata はこれを最大 0% に抑え、RAG が回避できない障害クラスです。 MemStrata は、LLM 再ランキング ベースラインの場合は最大 16 ~ 18 秒であるのに対し、取得レイテンシ (約 2.1 秒) でこれを達成します。知識の進化に基づいて、ハーネス、データセット、およびメモリーのマーカーフリー評価プロトコルをリリースします。

原文 (English)

Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge

Retrieval-augmented generation (RAG) gives agents access to accumulated knowledge, but has no model of time. When a fact changes (e.g., a function is renamed or API restructured), RAG retrieves both the stale and current value with near-identical embedding similarity. The agent then either abstains or serves the superseded fact. We show this is a structural problem: on a calibrated dataset, cosine similarity distinguishes a contradicted fact from a duplicated one with AUROC 0.59 (near chance), as contradictions are often more embedding-similar to the original than rephrased duplicates. We present MemStrata, a retrieval memory maintaining temporal validity. It stores facts like RAG, preserving static recall, but when a fact's value is contradicted, a deterministic (subject, relation, object) supersession rule retires the stale value in a bi-temporal ledger - with no similarity threshold and no LLM call. Across six benchmarks run locally with a 7B model, MemStrata ties RAG on static knowledge and reaches 0.95-1.00 accuracy on evolving knowledge (where RAG reaches 0.20-0.47). The central result is the stale-fact-error rate: when required to answer, RAG serves superseded values 15-40% of the time; MemStrata drives this to ~0%, a failure class RAG cannot avoid. MemStrata achieves this at retrieval latency (~2.1s) versus ~16-18s for LLM-reranking baselines. We release the harness, datasets, and a marker-free evaluation protocol for memory under knowledge evolution.

13:00 JST研究/論文

細胞培養プロセス予測のためのラマンデータ融合を使用したマルチパス適応ゲート型ボトルネック潜在 ODE

哺乳類の細胞培養プロセスは多くのバイオ医薬品の製造を支えていますが、計画どおりに稼働し続けることは困難です。重要なプロセスパラメータは数日で変動し、規格外の傾向が確認されると介入するには遅すぎることがよくあります。初期段階の複数日にわたる予測により、供給、サンプリング、制御をタイムリーに調整できる可能性がありますが、バイオプロセスの予測は困難です。測定値がまばらで不規則にサンプリングされ、操作条件が細胞株や培地間で不均一であり、初期挙動がほぼ同じである実行が異なる将来に分岐する可能性があるためです。ゲート付きボトルネック潜在常微分方程式 (GB-Latent ODE) とマルチパス ジャストインタイム微調整 (MP-JIT-FT) を組み合わせた適応フレームワークを提案します。 GB-Latent ODE は、学習可能な変数ごとのゲーティングと、高次元のスパース入力を圧縮するマスク対応ボトルネックを備えた標準 Latent ODE を拡張し、限られたデータの下での学習を向上させます。部分的に観測された実行を考慮すると、MP-JIT-FT は同様の過去の軌跡を取得し、局所近傍を候補レジームにクラスター化し、レジームごとに個別のモデルを微調整して複数のもっともらしいパスを生成します。各パスには単一の平均予測ではなく、再構成ベースの信頼スコアが付けられます。さらに、ラマン分光データを融合します。機械学習ソフト センサーは、高密度のラマン スペクトルを疑似観測に変換し、まばらなオフライン測定を強化して、より堅牢なトレーニングを実現します。 14 条件にわたる 38 回のフェドバッチ 5L バイオリアクターの実行で、ラマン融合を使用した MP-JIT-FT は最高の平均ランクを達成し、9 つのターゲット変数のうち 8 つでグローバルな潜在 ODE ベースラインを上回りました。ローカルダイバージェンスメトリクスを使用して、局所的に類似したプレフィックスが発散する場合にマルチパスゲインが最大になるのに対し、初期のダイナミクスが後の動作を表す場合にはラマン融合が最も役立つことを示します。

原文 (English)

Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting

Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process parameters drift over days, and an off-specification trend is often confirmed too late to intervene. Early-stage, multi-day forecasts could enable timely adjustment of feeding, sampling, and control, but bioprocess forecasting is challenging because measurements are sparse and irregularly sampled, operating conditions are heterogeneous across cell lines and media, and runs with near-identical early behaviour can diverge into different futures. We propose an adaptive framework combining a Gated Bottleneck Latent Ordinary Differential Equation (GB-Latent ODE) with Multi-Path Just-In-Time Fine Tuning (MP-JIT-FT). The GB-Latent ODE augments the stan dard Latent ODE with learnable variable-wise gating and a mask-aware bottleneck that compress high-dimensional sparse inputs, improving learning under limited data. Given a partially observed run, MP-JIT-FT retrieves similar historical trajectories, clusters the local neighbourhood into candidate regimes, and fine-tunes a separate model per regime to produce multiple plausible paths, each with a reconstruction-based confidence score, not a single averaged forecast. We further fuse Raman spectroscopy data: a machine-learning soft sensor turns dense Raman spectra into pseudo-observations that enrich the sparse offline measurements for more robust training. On 38 fed-batch 5L bioreactor runs spanning 14 conditions, MP-JIT-FT with Raman fusion achieves the best average rank and outperforms a global Latent ODE baseline on 8 of 9 target variables. Using local-divergence metrics, we show the multi-path gains are largest when locally similar prefixes diverge, whereas Raman fusion helps most when early dynamics are representative of later behaviour.

13:00 JSTLLM/生成AI画像/動画生成

不注意のギャップ: タスク条件付き言語モデルと視覚モデルは、そうでなければ報告できる安全上重要な信号を省略します

AI の安全性は、モデルが発見するように指示された危険をどれだけ確実に検出するかによって評価されますが、事故は多くの場合、誰も指定していない危険から発生します。私たちは、言語または視覚モデルを狭いタスクに条件付けすると、別のメカニズムから生じる人間の不注意による失明の機械の類似物である、他の方法で報告できる、同時に存在する安全上重要な信号の報告が抑制されることを示します。放射線学、ドライビングテキストシナリオ、および胸部X線写真の視覚タスク全体にわたって、抑制はテストされたすべてのモデルに現れ、スケールとともに減少せず、推論モデル内で持続し、サイズによるよりもモデルファミリーによって大きく異なりましたが、同じモデルは、制約されていない場合、これらの信号を実質的に高い割合で報告しました。私たちはこの解離を「不注意ギャップ」と名付け、測定されたベンチマークの安全性と現実世界の安全性を切り離していると主張します。つまり、システムは、危害を引き起こす危険性には気付かないまま、評価で指定された危険性についてはほぼ完璧にスコアを付けることができます。

原文 (English)

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We show that conditioning a language or vision model on a narrow task suppresses its reporting of co-present, safety-critical signals it can otherwise report, a machine analogue of human inattentional blindness arising from a different mechanism. Across radiology and driving text scenarios and chest-radiograph vision tasks, suppression appeared in every model tested, did not diminish with scale, persisted in a reasoning model, and varied more by model family than by size, while the same models reported these signals at substantially higher rates when unconstrained. We name this dissociation the Inattentional Gap and argue that it decouples measured benchmark safety from real-world safety: a system can score near-perfectly on the hazards an evaluation specifies while remaining blind to those that cause harm.

13:00 JSTLLM/生成AI

\textsc{DiARC}: ポジティブサンプルとネガティブサンプルの区別は、大規模な言語モデルの ARC のような推論能力の向上に役立ちます

Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) には、限られたグリッド サンプルからのパターンの要約と出力グリッドの予測を必要とするタスクが含まれています。最近、多くの大規模な言語モデル ベースのアプローチが、言語モデルをテキスト ベースの推論タスクに変換しようと試みています。ただし、オープンソース モデルに基づく方法では一般に満足のいく結果が得られず、クローズドソース モデルに依存する方法ではコストがかかりすぎます。現在の取り組みは主にデータ拡張に焦点を当てており、より包括的な監視付き微調整のための ARC のようなデータを構築しています。この研究では、ARC のような問題を解決するには \textit{positive} サンプルの監視だけでなく、\textit{negative} サンプルを区別してモデル推論を改善する能力も必要であると主張します。この目的を達成するために、私たちは好みの調整のアイデアを利用し、モデルがそれらを区別できるように好みのペアを構築する方法である \textsc{DiARC} を提案します。具体的には、出力レベルの視覚的変換、DSL レベルのルール反転、およびタスク固有のルール編集を含む、ネガティブ サンプルを構築する 3 つの方法を提案します。得られた陰性サンプルは、観察されたデモンストレーションを変更せずに、有益なニアミス代替案を提供します。複数の ARC に似たベンチマークにわたる実験結果は、\textsc{DiARC} がベースライン モデルよりも一貫してパフォーマンスを向上させることを示しています。コードは https://github.com/szu-tera/DiARC で公開されています。

原文 (English)

\textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

The Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid samples and predicting output grids. Recently, many large language model based approaches have attempted to transform it into a text-based reasoning task. However, methods based on open-source models have generally yielded unsatisfactory results, while those relying on closed-source models are too costly. Current efforts mainly focus on data augmentation, constructing ARC-like data for more comprehensive supervised fine-tuning. In this work, we argue that solving ARC-like problems requires not only \textit{positive} sample supervision but also the ability to improve model reasoning by distinguishing \textit{negative} samples. To this end, we draw on the idea of preference alignment and propose \textsc{DiARC}, a method that constructs preference pairs to enable the model to distinguish between them. Specifically, we propose three ways to construct negative samples, including output-level visual transformations, DSL-level rule inversion, and task-specific rule editing. The resulting negative samples provide informative near-miss alternatives while keeping the observed demonstrations unchanged. Experimental results across multiple ARC-like benchmarks show that \textsc{DiARC} consistently improves performance over baseline models. The code is released at https://github.com/szu-tera/DiARC.

13:00 JST研究/論文

VoiceTTA: 強化学習ベースのテスト時間適応によるゼロショット テキスト読み上げの強化

最近、ゼロショット テキスト読み上げ (TTS) により、高忠実度で表現力豊かな音声合成が可能になりましたが、一般的ではないシナリオ (クロストーク、方言など) からの目に見えない話し方を模倣できないことがよくあります。さらに、事前トレーニング済みモデルの微調整には大規模で高品質のデータセットが必要となり、迅速なパーソナライゼーションが制限されます。我々は、事前訓練されたゼロショット TTS モデルの音声模倣を改善する強化学習ベースのテスト時間適応 (TTA) 手法である VoiceTTA を提案します。 VoiceTTA は、F0 とエネルギーの変動係数の差に基づく 2 つのスタイルの報酬を、話者の類似性と明瞭度 (事前トレーニングされたウィスパー モデルからの WER) と組み合わせて導入し、推論時にフロー マッチング ベースのモデルのグループ相対選好最適化 (GRPO) によって学習可能なプレフィックスを最適化します。広範な実験により、珍しい音声プロンプトが大幅に改善され、最先端のベースラインを上回るパフォーマンスが実証されました。音声サンプルは https://voicetta.pages.dev/ で入手できます。

原文 (English)

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.g., crosstalk, dialects). Moreover, fine-tuning pretrained models requires large, high-quality datasets, limiting rapid personalization. We propose VoiceTTA, a reinforcement learning-based test-time adaptation (TTA) method that improves voice imitation of pretrained zero-shot TTS models. VoiceTTA introduces two style rewards based on coefficient-of-variation differences of F0 and energy, combined with speaker similarity and intelligibility (WER from a pretrained Whisper model), and optimizes learnable prefixes via group relative preference optimization (GRPO) in a flow matching-based model at inference time. Extensive experiments demonstrate substantial improvements on uncommon speech prompts, outperforming state-of-the-art baselines. Audio samples are available at https://voicetta.pages.dev/

13:00 JST画像/動画生成ビジネス/資金調達

幻覚からグラウンディングまで: CRISP による視覚空間知能の診断

現在の VLM 評価では、言語事前分布と真の空間推論が混同されることがよくあります。これに対処するために、CRISP を導入します。CRISP は、一貫性、つまり暗黙の認識と明示的な推論の整合性を通じて視覚空間知能を評価する、新しい構造診断評価パラダイムです。従来のブラックボックス QA とは異なり、CRISP はメトリック 3D シーン グラフとオラクル介入プロトコルを利用して、潜在的な推論機能を知覚のボトルネックから切り離します。この詳細な診断により、体系的な知覚と推論の断絶が明らかになります。重要なことに、独自のモデルは強力な潜在的な推論エンジンを備えているものの、不正確な計量推定と暗黙の構造表現を活用する重大な失敗に悩まされていることを明らかにしました。逆に、オープンソース モデルは、マルチホップの構成推論が欠如しているため、依然として根本的なボトルネックとなっています。 CRISP は、事前言語を介して単に「正しく推測する」ことから、真に「知覚、検証、推論する」ことに焦点を移すことで、エンドツーエンドのポストトレーニングを超えたマルチモーダル調整のための厳密なロードマップを提供します。コードとデータセットは https://github.com/iiyamayuki/CRISP-Bench で入手できます。

原文 (English)

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. Unlike traditional black-box QA, CRISP utilizes metric 3D Scene Graphs and an oracle intervention protocol to decouple latent reasoning capabilities from perceptual bottlenecks. This granular diagnosis uncovers a systematic perception-reasoning disconnect. Crucially, we reveal that while proprietary models possess robust latent reasoning engines, they suffer from inaccurate metric estimation and a critical failure to leverage their implicit structural representations. Conversely, open-source models remain fundamentally bottlenecked by their lack of multi-hop compositional reasoning. By shifting the focus from merely ``guessing correctly'' via language priors to genuinely ``perceiving, verifying, and reasoning,'' CRISP offers a rigorous roadmap for multimodal alignment beyond end-to-end post-training. The code and dataset are available at https://github.com/iiyamayuki/CRISP-Bench.

13:00 JST研究/論文

CascadeFormer: 勾配ファンイン非対称性を利用した深さテーパー型トランスフォーマー

Deep Transformer は均一に積み重ねられた残留ブロックで構成されていますが、最も深い層にはほとんど価値がありません。この非対称性を利用した 2 つの効率化方法を紹介します。 CascadeFormer は、層全体にわたる不均一な情報フローに合わせて深さとともに幅を細くし、同じトレーニング予算で均一なベースラインと同等の複雑さを実現しながら、レイテンシを 8.6% 削減し、スループットを 9.4% 向上させます。 CascadeFlow Pruning は、事後分析を行わずに、蓄積されたトレーニング勾配を使用してレイヤーを削除します。複雑さとランクの安定性に関しては標準のヒューリスティックよりも優れており、ダウンストリームの精度に関しては競争力を維持します。これらの方法を動機付けるために、より深い層ほど寄与が少ない理由の構造的説明として、勾配ファンイン非対称性 (GFA) を提案します。 Pre-LayerNorm 残差スタックでは、層の勾配はアイデンティティ パスと下流のすべての関数パスの合計であり、深さとともに線形に減衰する勾配ファンイン (および深い監視下では二次関数的に減衰) が生成され、初期の層ではより豊かな勾配が、後の層ではよりまばらな勾配が生成されます。私たちは、ゼロから最大 1.2B パラメーターまでトレーニングされたモデルに対する GFA の相関証拠と介入証拠を提供します。 Transformer と ResNet 全体で、蓄積されたトレーニング勾配は理論的なファンインに従い、ポストホック層の重要性に関連付けられます。 2 つの介入は、ボトルネックとして大きさではなく構造を指摘しています。レイヤーごとの勾配ノルムを均等化しても、後期レイヤーの値は復元されませんが、パラメーター共有の反復復元によって下流のパス数が増加し、値が上昇します。勾配の大きさのプロキシが高ランク領域を超えてファンインするかどうか、およびこれらのダイナミクスが 100B 以上のスケールでどのように動作するかは、未解決の疑問のままです。

原文 (English)

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry

Deep Transformers are composed of uniformly stacked residual blocks, yet their deepest layers often add little value. We present two efficiency methods that exploit this asymmetry. CascadeFormer tapers width with depth to match the uneven information flow across layers, achieving comparable perplexity to a uniform baseline at the same training budget while reducing latency by 8.6% and increasing throughput by 9.4%. CascadeFlow Pruning removes layers using accumulated training gradients, with no post hoc analysis. It outperforms standard heuristics on perplexity and rank-stability and stays competitive on downstream accuracy. To motivate these methods, we propose Gradient Fan-in Asymmetry (GFA) as a structural account of why deeper layers contribute less. In Pre-LayerNorm residual stacks, the gradient at a layer is the sum of an identity path and all downstream functional paths, producing a gradient fan-in that decays linearly with depth (and quadratically under deep supervision), yielding richer gradients for early layers and sparser ones for later layers. We provide correlational and interventional evidence for GFA on models trained from scratch up to 1.2B parameters. Across Transformers and ResNets, accumulated training gradients follow the theoretical fan-in and are associated with post hoc layer importance. Two interventions point to structure rather than magnitude as the bottleneck: equalizing per-layer gradient norms does not restore late-layer value, while increasing downstream path counts via parameter-shared repetition restores and elevates it. Whether gradient magnitude proxies fan-in beyond high-rank regimes, and how these dynamics behave at the 100B+ scale, remain open questions.

13:00 JST画像/動画生成エージェントGPT / ChatGPT

認識、評決、進化: AI 生成画像検出のための後知恵主導の自己洗練フォレンジック エージェント

生成モデルの急速な進歩は、特に非常にリアルな AI 生成画像の広範な普及を考慮すると、既存のディープフェイク検出方法に重大な課題をもたらしています。マルチモーダル大規模言語モデル (MLLM) はこのタスクに対して強力な可能性を示していますが、既存のアプローチには 2 つの重要な制限があります。1 つはきめ細かいフォレンジック アーティファクトに対する感度が不十分であること、もう 1 つはフロンティア モデルからの静的合成監視に依存しているため、柔軟性が限られ、コストが高くなるという点です。これらの問題に対処するために、反復的な自己進化を備えた AI 生成の画像検出のためのエージェントフォレンジック フレームワークである ForeAgent を提案します。まず、ForeAgent は、セマンティック、空間、および周波数領域の機能にわたるマルチビュー キューを集約する Perception-Verdict アーキテクチャを採用し、MLLM を判定モジュールとして利用して、論理的に根拠のある判定を行うためにこれらの信号を融合します。第二に、継続的な自己改善を可能にするために、サンプリング、反映、進化のパラダイムに従った後知恵主導の自己洗練戦略を導入します。エージェントは、トレーニング インスタンスに対して推論ロールアウトを実行します。結果論としてのグラウンドトゥルースのラベルに基づいて、失敗例と低品質の推論軌跡を反映して、より高品質の推論トレースを再生成します。これらの合成されたサンプルは、デュアル エキスパートの高品質ゲート モジュールを通じて厳密にフィルター処理されます。 ForeAgent は、独自に厳選した高品質のサンプルを微調整することで継続的に進化します。広範な実験により、ForeAgent が Chameleon ベンチマークで最先端のパフォーマンスを達成し、82.18% の精度 (AIDE に対して +16.41%) に達し、16 のジェネレーターにわたる AIGCDetect-Benchmark で 93.3% の平均精度を達成していることが実証されています。さらに、外部評価では、ForeAgent が GPT-5 および GPT-5-mini と比較して、より一貫性があり、因果関係に基づいた推論を生成することが示されています。

原文 (English)

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images. Although Multimodal Large Language Models (MLLMs) show strong potential for this task, existing approaches suffer from two key limitations: insufficient sensitivity to fine-grained forensic artifacts and reliance on static synthetic supervision from frontier models, leading to limited flexibility and high-cost. To address these issues, we propose ForeAgent, an agentic forensics framework for AI-generated image detection with iterative self-evolution. First, ForeAgent adopts a Perception-Verdict architecture that aggregates multi-view cues spanning semantic, spatial, and frequency-domain features, and leverages an MLLM as a verdict module to fuse these signals for a logical-grounded verdict. Second, to enable continual self-improvement, we introduce a Hindsight-Driven Self-Refining strategy following a Sampling-Reflection-Evolution paradigm. The agent performs inference rollouts on training instances. Guided by ground-truth labels as hindsight, it reflects on failure cases and low-quality reasoning trajectories to regenerate higher-quality reasoning traces. These synthesized samples are then strictly filtered through a dual-expert quality gating module. ForeAgent continuously evolves via fine-tuning on self-curated high-quality samples. Extensive experiments demonstrate that ForeAgent achieves state-of-the-art performance on the Chameleon benchmark, reaching 82.18% accuracy (+16.41% over AIDE), and achieves 93.3% mean accuracy on AIGCDetect-Benchmark across 16 generators. In addition, external evaluation shows that ForeAgent produces more consistent and causally grounded reasoning compared to GPT-5 and GPT-5-mini.

13:00 JST画像/動画生成

SpaceRipple: ミッション指向の LEO 地球観測衛星ネットワーク向けの軽量セマンティック配信

地球観測衛星ネットワークは大量の高解像度画像を生成しますが、衛星間およびダウンリンクのリソースは依然として限られています。多くの時間制限のあるミッションでは、地上ユーザーは完全な生画像のダウンリンクではなく、ミッション関連のセマンティック情報を必要とします。この論文では、地球観測衛星ネットワークにおけるミッション指向のセマンティック配信およびオンボード処理のための軽量フレームワークである SpaceRipple を提案します。センシング衛星は適応圧縮とメタデータ生成を実行して衛星間のトラフィックを削減し、エッジ コンピューティング衛星は受信した表現を復元し、タスク関連のセマンティック情報を抽出します。忠実度重視の画像送信とは異なり、SpaceRipple は圧縮、転送、復元、セマンティック推論を協調パイプライン内で調整し、ピクセル レベルの画像配信ではなくセマンティック指向の配信を可能にします。圧縮を意識した MoE 拡張モジュールがさらに導入され、劣化したビジュアル入力下での堅牢性が向上します。実験結果は、SpaceRipple が良好な再構築品質、向上した意味検出パフォーマンス、および大幅な帯域幅の節約を実現し、限られた衛星ネットワーク リソースの下で効率的かつ信頼性の高い地球観測を可能にする可能性を示していることを示しています。

原文 (English)

SpaceRipple: Lightweight Semantic Delivery for Mission-Oriented LEO Earth Observation Satellite Networks

Earth observation satellite networks generate massive volumes of high-resolution imagery, whereas inter-satellite and downlink resources remain limited. In many time-sensitive missions, ground users require mission-relevant semantic information rather than a full raw-image downlink. This paper proposes SpaceRipple, a lightweight framework for mission-oriented semantic delivery and on-board processing in Earth observation satellite networks. A sensing satellite performs adaptive compression and metadata generation to reduce inter-satellite traffic, while an edge computing satellite restores the received representation and extracts task-relevant semantic information. Unlike fidelity-driven image transmission, SpaceRipple coordinates compression, forwarding, restoration, and semantic inference within a collaborative pipeline, enabling semantic-oriented delivery instead of pixel-level image delivery. A compression-aware MoE enhancement module is further introduced to improve robustness under degraded visual inputs. Experimental results show that SpaceRipple achieves favorable reconstruction quality, improved semantic detection performance, and substantial bandwidth savings, demonstrating its potential for efficient and reliable Earth observation under constrained satellite-network resources.

13:00 JST研究/論文

scBench-Long: 長期にわたる単細胞生物学の検証可能なベンチマーク

単一細胞研究では、分析者が多段階のワークフローとメタデータ、アッセイコンテキスト、および補助証拠の統合を通じて、生の測定値を特定の生物学的主張に変換する必要があります。既存の AI 生物学ベンチマークは主に、幅広い知識、実行可能なワークフロー、またはローカル分析ステップを測定します。我々は、長期にわたる単細胞生物学のベンチマークである scBench-Long を紹介します。このベンチマークでは、担当者は規定の方法を使用せずに生データまたは生に近いデータから科学的結論を導き出す必要があります。このベンチマークには、黒色腫 CD8 T 細胞反応性、CD8 RNA+ATAC 制御推論、ヒト - サルのキメラ開発、KRAS による肺腫瘍の老化、致死性の COVID-19 肺の病理にわたる 21 の評価が含まれています。タスクには、ペア scRNA/TCR シーケンス、RNA およびクロマチン プロファイリング、異種トランスクリプトミクス、コンビナトリアル scRNA シーケンス、単核 RNA シーケンス、免疫レパトア、オルソログ マップ、リガンド - 受容体リソース、および検証証拠が含まれます。候補者の主張は再現され、レビューされ、決定的な採点と軌道ルーブリックを備えた管理された解答語彙に変換されます。 1,068 の完了した軌道全体で、最も強力なモデルであるハーネス ペアは 16/63 ラン (25.4\%) を通過しました。 scBench-Long は、エージェントが局所的な分析ステップを超えて、単一細胞データによって裏付けられた複雑な科学的主張を行えるかどうかを評価します。

原文 (English)

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence. Existing AI-biology benchmarks largely measure broad knowledge, executable workflows, or local analysis steps. We introduce scBench-Long, a benchmark for long-horizon single-cell biology in which agents must recover scientific conclusions from raw or near-raw data without prescribed methods. The benchmark contains 21 evaluations spanning melanoma CD8 T-cell reactivity, CD8 RNA+ATAC regulatory inference, human--monkey chimera development, KRAS-driven lung tumor aging, and lethal COVID-19 lung pathology. Tasks cover paired scRNA/TCR sequencing, RNA and chromatin profiling, cross-species transcriptomics, combinatorial scRNA-seq, single-nucleus RNA-seq, immune repertoires, ortholog maps, ligand--receptor resources, and validation evidence. Candidate claims are reproduced, reviewed, and converted into controlled answer vocabularies with deterministic grading and trajectory rubrics. Across 1,068 completed trajectories, the strongest model--harness pair passes 16/63 runs (25.4\%). scBench-Long evaluates whether agents can move beyond local analysis steps and make complex scientific claims that are supported by single-cell data.

13:00 JSTエージェントロボティクス

IDEA: マルチエージェント制御における Sim-to-Real 転送のためのエフェクト アライメントによるダイナミクスの不一致の影響を受けにくい

複雑なマルチエージェント制御タスクは、従来のルールベースおよびモデルベースのアプローチにとって依然として困難であり、学習ベースの手法の採用を動機付けています。ただし、学習ベースの手法は、正確なダイナミクス モデリングやシステム識別に依存し、ダイナミクスの不一致に非常に敏感な低レベルの制御空間でポリシーを学習するため、シムからリアルへの転送にしばしば苦労するため、複雑な環境ではコストがかかり脆弱になります。この問題に対処するために、効果の調整によるダイナミクスの不一致の影響を受けにくい、マルチエージェント制御のための sim-to-real 手法を提案します。私たちの手法は、閉ループ制御を通じてランダムな環境構造と離散的な意味論的なアクションを組み合わせ、政策学習を意味論的な抽象化レベルまで高めます。さらに、エージェント間のアクションのタイミングの不一致を軽減するアクション同期メカニズムを開発し、それによってシステムの時間的一貫性を強化します。 4 つのマルチエージェント ナビゲーション タスクに関する実験では、私たちの方法が主流の転送方法よりもトレーニング効率を大幅に向上させ、現実世界のシナリオでより高い成功率を達成し、それによってダイナミクスの不一致下でのマルチエージェント システムの堅牢性と展開の安定性が向上することが実証されました。

原文 (English)

IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control

Complex multi-agent control tasks remain challenging for traditional rule-based and model-based approaches, motivating the adoption of learning-based methods. However, learning-based methods often struggle with sim-to-real transfer because they rely on accurate dynamics modeling or system identification and learn policies in low-level control spaces that are highly sensitive to dynamics mismatch, making them costly and fragile in complex environments. To address this issue, we propose a sim-to-real method for multi-agent control, which is insensitive to dynamics mismatch via effect alignment. Our method combines random environmental structure with discrete semantic actions through closed-loop control, elevating policy learning to a semantic abstraction level. Additionally, we develop an action synchronization mechanism that mitigates inter-agent action timing mismatches, thereby enhancing the temporal consistency of the system. Experiments on four multi-agent navigation tasks demonstrate that our method substantially improves training efficiency over mainstream transfer methods and achieves higher success rates in real-world scenarios, thereby improving the robustness and deployment stability of multi-agent systems under dynamics mismatch.

13:00 JSTLLM/生成AILlama

SharQ: LLM 推論のためのアクティベーション スパース性と FP4 量子化のブリッジング

低ビット浮動小数点形式と半構造化スパース性は、最新のアクセラレータでますますサポートされていますが、LLM アクティベーション圧縮のためにそれらを組み合わせるのは依然として困難です。アクティベーションには、FP4 量子化のブロック スケールを支配する入力依存の外れ値が含まれており、N:M スパース マスクを直接適用すると適度な値が破棄され、スパース化損失と量子化誤差が結合します。私たちは、オンラインの疎-密分解を通じて活性化スパース性と FP4 量子化を橋渡しする、トレーニング不要の推論手法である SharQ を紹介します。各活性化テンソルに対して、SharQ は入力適応型 N:M マスクを生成して、外れ値が支配的なスパース バックボーンを抽出し、それを FP4 に量子化し、量子化されていないスパース値ではなく、量子化されたスパース バックボーンに対して密な残差を定義します。スパース FP4 GEMM はバックボーンを処理し、デンス FP4 GEMM はマスクによるアクティベーション損失とスパース パス量子化エラーの両方を補償します。 2 つのパスは、パス固有のスケール ビューを備えた単一の FP4 ウェイト ペイロードを共有し、融合された準備カーネルがマスク生成、残差構築、レイヤー正規化を 1 つのオペレーターに吸収します。 SharQ では、キャリブレーション データ、再トレーニング、モデル固有のチューニングは必要ありません。 Llama-3.1-8B、Qwen2.5-7B、Qwen3-30B-A3B、および Qwen3-VL-8B で評価したところ、SharQ は言語および視覚言語タスク全体で NVFP4 と FP16 の精度の差の 43 ~ 63% を回復し、NVFP4、HiF4、および MXFP4 形式にわたって一般化します。 RTX 5090 では、SharQ は、FP16 と比較して 2.2 ~ 2.4$\times$ の遅延削減を実現し、言語モデルの処理において FP8 と比較して 1.2 ~ 1.4$\times$ のスループット向上を実現し、SageAttend と組み合わせることで Wan2.2-T2V-A14B ビデオ生成で最大 1.58$\times$ の高速化を実現します。私たちのコードは https://github.com/actypedef/SharQ で入手できます。

原文 (English)

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference

Low-bit floating-point formats and semi-structured sparsity are increasingly supported by modern accelerators, yet combining them for LLM activation compression remains challenging: activations contain input-dependent outliers that dominate block scales in FP4 quantization, and directly applying N:M sparsity masks discards moderate values, coupling sparsification loss with quantization error. We introduce SharQ, a training-free inference method that bridges activation sparsity and FP4 quantization through an online sparse--dense decomposition. For each activation tensor, SharQ generates an input-adaptive N:M mask to extract an outlier-dominated sparse backbone, quantizes it to FP4, and defines a dense residual relative to the quantized sparse backbone rather than the unquantized sparse values. A sparse FP4 GEMM processes the backbone while a dense FP4 GEMM compensates for both mask-induced activation loss and sparse-path quantization error. The two paths share a single FP4 weight payload with path-specific scale views, and a fused preparation kernel absorbs mask generation, residual construction, and layer normalization into one operator. SharQ requires no calibration data, retraining, or model-specific tuning. Evaluated on Llama-3.1-8B, Qwen2.5-7B, Qwen3-30B-A3B, and Qwen3-VL-8B, SharQ recovers 43--63% of the NVFP4-to-FP16 accuracy gap across language and vision-language tasks, and generalizes across NVFP4, HiF4, and MXFP4 formats. On an RTX 5090, SharQ delivers 2.2--2.4$\times$ latency reduction over FP16 and 1.2--1.4$\times$ throughput improvement over FP8 in language model serving, and up to 1.58$\times$ speedup on Wan2.2-T2V-A14B video generation when combined with SageAttention. Our code is available at https://github.com/actypedef/SharQ.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

HiLSVA: 科学的視覚化のための人間参加型エージェント システムの設計と評価

大規模言語モデル (LLM) エージェントは、科学的視覚化 (SciVis) のための自然言語対話を可能にします。それでも、従来のシステムは基本的に人間による分析制御よりも自律性を優先しており、そのため透明性と人間による監視が制限されていました。混合イニシアチブの SciVis ワークフローをサポートする人間参加型エージェント システムである HiLSVA を紹介します。 HiLSVA は、計画優先のマルチエージェント アーキテクチャと、人間による明示的な監視、段階的な出所追跡、およびユーザー フィードバックからのテスト時の学習の適応を統合します。このシステムは、自然言語と視覚化の直接操作の両方を通じて、人間とエージェント間の流動的なハンドオフをサポートし、サンドボックス実行により安全で再現可能なワークフローを保証します。そうすることで、HiLSVA はエージェント的 SciVis を、人間の分析的推論を置き換えるのではなく、強化する共同プロセスとして再構成します。私たちは、代表的なケーススタディと、複数の自律性設定にわたるさまざまな専門知識を持つ 12 人の参加者による対照ユーザー研究を通じて HiLSVA を評価します。結果は、混合イニシアチブの相互作用により、さまざまなレベルのユーザーの専門知識にわたってタスクの完了、ユーザー制御、およびワークフローの透明性が向上する一方、実行効率と人間の監視の間のトレードオフが明らかになったことが示されています。これらの発見は、エージェント的 SciVis における人間中心設計の重要性を強調し、将来の共同視覚化システムの開発の指針となります。 https://hilsva.github.io/ でデモビデオ、ケーススタディ、ソースコードを探索することをお勧めします。

原文 (English)

HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization

Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human oversight. We present HiLSVA, a human-in-the-loop agentic system that supports mixed-initiative SciVis workflows. HiLSVA integrates a plan-first multi-agent architecture with explicit human oversight, stepwise provenance tracking, and learn-at-test-time adaptation from user feedback. The system supports fluid handoff between humans and agents through both natural language and direct manipulation of visualizations, while sandboxed execution ensures safe, reproducible workflows. In doing so, HiLSVA reframes agentic SciVis as a collaborative process that augments, rather than replaces, human analytical reasoning. We evaluate HiLSVA through representative case studies and a controlled user study with twelve participants of varying expertise across multiple autonomy settings. Results show that mixed-initiative interaction improves task completion, user control, and workflow transparency across different levels of user expertise, while revealing a tradeoff between execution efficiency and human oversight. These findings highlight the importance of human-centered design in agentic SciVis and guide the development of future collaborative visualization systems. We encourage readers to explore our demo video, case studies, and source code at https://hilsva.github.io/.

13:00 JST研究/論文

スパース オートエンコーダで何百万もの解釈可能な機能を発見

スパース オートエンコーダ (SAE) は、重ね合わせた言語モデル表現をスパースで解釈可能な特徴に分解するための強力なツールとして登場しました。ただし、SAE のトレーニングには計算コストがかかり、利用可能なオープンソース SAE モデルは依然として限られています。この作業では、Qwen3 命令調整モデル ファミリでトレーニングされた、Qwen3-1.7B、Qwen3-4B、および Qwen3-8B をカバーする包括的な SAE スイートである \textbf{Qwen3-Instruct SAE} を紹介します。 Qwen3-1.7B および Qwen3-4B では、残差ストリーム、MLP 出力、およびアテンション出力という 3 つの主要なアクティベーション サイトでレイヤーごとの SAE をトレーニングします。 Qwen3-8B の場合、残りのストリーム層のサブセットで SAE をトレーニングします。私たちは、アクティベーションレベルの再構築メトリクスとモデルレベルの回復メトリクスの両方を使用して、これらの SAE を体系的に評価し、層とコンポーネントにわたる明確なスパース性と忠実度のトレードオフを明らかにします。最後に、拒否ステアリングのケーススタディを通じて Qwen3-Instruct SAE の有用性を実証し、選択された SAE 機能が命令によって調整された Qwen3 モデルを拒否行動に向けて因果的に誘導できることを示します。私たちのリリースは、命令調整された言語モデルにおけるスパース表現、機能レベルのメカニズム、および行動介入を研究するための実用的なリソースを提供します。

原文 (English)

Discovering Millions of Interpretable Features with Sparse Autoencoders

Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive suite of SAEs trained on the Qwen3 instruction-tuned model family, covering Qwen3-1.7B, Qwen3-4B, and Qwen3-8B. For Qwen3-1.7B and Qwen3-4B, we train layer-wise SAEs at three key activation sites: residual streams, MLP outputs, and attention outputs. For Qwen3-8B, we train SAEs on a subset of residual stream layers. We systematically evaluate these SAEs using both activation-level reconstruction metrics and model-level recovery metrics, revealing distinct sparsity--fidelity trade-offs across layers and components. Finally, we demonstrate the utility of Qwen3-Instruct SAE through a refusal-steering case study, showing that selected SAE features can causally steer instruction-tuned Qwen3 models toward refusal behavior. Our release provides a practical resource for studying sparse representations, feature-level mechanisms, and behavioral interventions in instruction-tuned language models

13:00 JSTLLM/生成AIエージェント

知りすぎているエージェント: LLM エージェントのプライバシーに関するデータ中心の調査

大規模な言語モデル エージェントは、ますますデータベースにクエリを実行し、ドキュメント コレクションを検索し、外部 API を呼び出し、過去の対話を記憶し、ユーザーに代わって動作します。質問に答えることから機密データを扱うことに移行するにつれて、プライバシーを強制することは難しくなります。エージェントは多くのデータ ソースにアクセスし、複数ステップのワークフローを実行し、セッション間で状態を保持し、委任された権限で動作します。したがって、機密情報は、最終的な応答を通じてだけでなく、発行するクエリ、処理する中間結果、書き込むメモリ、および他のエージェントと交換するメッセージを通じて漏洩する可能性があります。私たちは LLM エージェントのプライバシーをデータ中心の観点から調査し、攻撃の種類ではなくエージェントが接触するデータを中心にフィールドを整理し、データを扱う LLM エージェントの略語としてデータ エージェントを使用します。これらのリスクに関する研究は活発に行われていますが、検索拡張生成、テキストから SQL へのインターフェイス、エージェント メモリ、プロンプト インジェクション、アクセス制御、およびコンテキスト プライバシーに分散しています。この調査では、その作業をまとめます。エージェントが接触するデータ ソース、各ソースが生み出すプライバシー リスク、およびそれらに対処するガバナンス メカニズムを分類します。これらのリスクを測定するために使用されるベンチマークをマッピングし、何が欠けているかを特定します。そして私たちは未解決の問題を明らかにしました。 2 つの発見が繰り返されます。ガバナンス メカニズムの中で、最も保護されていない 2 つのリスクである構成的推論漏洩とセッション間の推論漏洩の両方をカバーするのは情報フロー制御だけです。そして、この分野に最も欠けている手段である 1 つのプライバシー ポリシーに基づいてエージェントをそのデータ サーフェイス全体に動かすベンチマークはありません。私たちの目標は、散在する文献を整理し、将来の研究に共通の枠組みを与える参考資料となることです。

原文 (English)

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents

Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a user's behalf. As they move from answering questions to operating over sensitive data, privacy becomes harder to enforce. An agent touches many data sources, runs multi-step workflows, keeps state across sessions, and acts with delegated permissions. Sensitive information can therefore leak not only through its final answer but through the queries it issues, the intermediate results it handles, the memory it writes, and the messages it exchanges with other agents. We survey the privacy of LLM agents from a data-centric view, organizing the field around the data an agent touches rather than by attack type, and we use data agent as shorthand for an LLM agent that works with data. Research on these risks is active but scattered across retrieval-augmented generation, text-to-SQL interfaces, agent memory, prompt injection, access control, and contextual privacy. This survey brings that work together: we taxonomize the data sources an agent touches, the privacy risks each source creates, and the governance mechanisms that address them; we map the benchmarks used to measure these risks and identify what is missing; and we set out the open problems. Two findings recur: among governance mechanisms only information-flow control covers both compositional and cross-session inference leakage, the two least-protected risks; and no benchmark drives an agent across its data surfaces under one privacy policy, the instrument the field most lacks. Our goal is a reference that situates the scattered literature and gives future work a common framing.

13:00 JSTLLM/生成AI研究/論文

CAT-Q: LLM 向けのコスト効率が高く正確な 3 値量子化

このペーパーでは、LLM を圧縮および高速化するための、コスト効率が高く正確な 3 値量子化である CAT-Q について説明します。深刻なパフォーマンス低下を軽減するためにデータ集約的でコストのかかる量子化を意識したトレーニングに依存する既存の最先端の 3 値量子化手法とは異なり、CAT-Q はシンプルで効果的なトレーニング後の量子化スキームであり、さまざまなアーキテクチャとモデル サイズの LLM にすぐに適用できます。これには、学習可能な変調 (LM) と軟化三値化 (ST) という 2 つの主要コンポーネントがあり、これらは最適化の観点から結合されています。 LM は、学習可能な要素の構成を利用して、事前トレーニングされた高精度の重みと 3 値のしきい値の分布を調整し、3 値化の影響を受けにくくします。 ST はさらに、微分可能な遷移関数を導入して、3 値化プロセスを安定した収束に導きます。 1.7B ~ 8B パラメータを持つ事前トレーニング済み LLM の場合、CAT-Q はわずか 512 個のキャリブレーション サンプルを使用してそれらを 3 値モデルに効率的に量子化し、同時に 100B トークンでトレーニングされた独創的な BitNet 1.58 ビット v1 および v2 ファミリ (1.3B ~ 7B パラメータ) よりも優れたパフォーマンスを達成し、トレーニング時間を約 100,000 倍削減できることを示します。トークン。さらに、CAT-Q が 8 個の A100-80GB GPU でわずか 8 ~ 60 時間以内に、14B ~ 235B のパラメータを持つはるかに大きな事前トレーニング済み LLM を主要な 3 値モデルに量子化できることを初めて示しました。コードは https://github.com/IntelChina-AI/BitTern で入手できます。

原文 (English)

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing state-of-the-art ternary quantization methods that rely on data-intensive and costly quantization-aware training to mitigate severe performance degradation, CAT-Q is a simple yet effective post-training quantization scheme that is readily applicable to LLMs with diverse architectures and model sizes. It has two key components, learnable modulation (LM) and softened ternarization (ST), which are coupled from an optimization perspective. LM leverages a composition of learnable factors to modulate the distribution of pre-trained high-precision weights and the ternary threshold, making them less sensitive to ternarization. ST further introduces a differentiable transition function to guide the ternarization process toward stable convergence. We show that, for pre-trained LLMs with 1.7B to 8B parameters, CAT-Q can efficiently quantize them into ternary models using only 512 calibration samples, while achieving superior performance than the seminal BitNet 1.58-bit v1 and v2 families (with 1.3B to 7B parameters) trained with 100B tokens, yielding about a 100,000X reduction in training tokens. Moreover, we show for the first time that CAT-Q can quantize much larger pre-trained LLMs having 14B to 235B parameters into leading ternary models within just 8 to 60 hours on 8 A100-80GB GPUs. Code is available at https://github.com/IntelChina-AI/BitTern.

13:00 JSTエージェントロボティクス

LAMP: 実現可能な軌道予測のためのレーンアライメントモーションプリミティブ

動き予測は、複雑な運転シナリオにおいて安全な意思決定と計画を可能にする自動運転システムにとって不可欠です。既存の予測器は、標準変位誤差を最小限に抑える点では優れていますが、特に低確率モードの場合、マルチモーダル予測のレーン トポロジーの遵守を見落とすことがよくあります。その結果、予測された軌跡が物理的および論理的制約に違反する可能性があり、安全性が重要な計画に対して予測セットの信頼性が低くなります。この論文では、LAMP (Lane-Aligned Motion Primitives) を提案します。LAMP (Lane-Aligned Motion Primitives) は、レーン トポロジに合わせた構造化されたモーション プリミティブにマルチモーダル予測を固定する、トポロジを意識した予測フレームワークです。具体的には、VQ-VAE を使用して、形状認識モーション プリミティブを離散意図クエリとして学習し、エンドポイントベースの意図を超えた時空間パターンをキャプチャします。さらに、到達不可能な意図クエリをフィルタリングする前にレーン トポロジでトレーニングされた実現可能性を意識した意図セレクターを導入し、動作の多様性を維持しながらトポロジーの一貫性のある意図を優先するようにデコーダーを導きます。 Argoverse 2 データセットに対する広範な実験により、LAMP が最先端のベースラインに匹敵する予測精度を達成しながら、実現可能性と多様性のメトリクスにおいてはそれらを上回っていることが実証されました。

原文 (English)

LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction

Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. While existing predictors excel at minimizing standard displacement errors, they often overlook the adherence to lane topology of multimodal predictions, particularly for lower-probability modes. Consequently, predicted trajectories may violate physical and logical constraints, making the prediction set unreliable for safety-critical planning. In this paper, we propose LAMP (Lane-Aligned Motion Primitives), a topology-aware forecasting framework that anchors multimodal prediction to structured motion primitives aligned with lane topology. Specifically, we use a VQ-VAE to learn shape-aware motion primitives as discrete intention queries, capturing spatiotemporal patterns beyond endpoint-based intentions. We further introduce a feasibility-aware intention selector trained with a lane-topology prior for filtering unreachable intention queries, guiding the decoder to prioritize topology-consistent intentions while preserving behavioral diversity. Extensive experiments on the Argoverse 2 dataset demonstrate that LAMP achieves prediction accuracy comparable to state-of-the-art baselines while outperforming them in feasibility and diversity metrics.

13:00 JST研究/論文

スパースランダムグラフ上のニューラルODEのゼロショットサイズ転送: グラフの限界と随伴収束

グラフ ニューラル微分方程式 (GNDE) は、グラフ ニューラル ネットワークを使用してニューラル ODE 速度場をパラメーター化することにより、連続時間グラフ ダイナミクスをモデル化します。サイズに依存しないローカル フィルターは、ゼロショット サイズ転送の原則を示唆しています。つまり、小さなグラフでトレーニングし、再トレーニングせずに大きな同様のグラフにデプロイします。我々は、グラフォンからサンプリングされた疎なランダムグラフに関するこの原理の定量的理論を開発します。我々は、グラフォン神経微分方程式 (Graphon-NDE) と随伴グラフォン NDE を前方および随伴 GNDE システムの無限ノード限界として考慮し、適切なポーズネスを確立します。スパース性パラメーター $\alpha_n$ をもつ $n$ ノードのランダム グラフについて、対数因数までのレート $O((\alpha_n n)^{-1/2})$ で GNDE 解が Graphon-NDE 解に軌跡的に収束することを高い確率で証明します。また、隠れ状態とパラメータの勾配を制御する随伴系の時間内均一収束限界も確立します。さらに、離散化してから最適化 (DTO) トレーニングと最適化してから離散化 (OTD) トレーニングを研究します。 $M$ ステップによる明示的なオイラー離散化では、DTO と OTD が漸近的に一貫しており、それぞれ次数 $O(1/M)$ と $O(1/M^2)$ の隠れ状態と局所パラメータ勾配の不一致が、スパース性と対数因子まで存在することを示します。 HSBM とテント グラフォンの実験は理論レートをサポートし、4 つのグラフォン クラスにわたるゼロショット転送実験は、より大きな独立してサンプリングされたグラフ上で学習された GNDE の正確な展開を実証します。

原文 (English)

Zero-Shot Size Transfer for Neural ODEs on Sparse Random Graphs: Graphon Limits and Adjoint Convergence

Graph Neural Differential Equations (GNDEs) model continuous-time graph dynamics by parameterizing Neural ODE velocity fields with Graph Neural Networks. Their local, size-independent filters suggest a zero-shot size-transfer principle: train on a small graph and deploy on larger, similar graphs without retraining. We develop a quantitative theory for this principle on sparse random graphs sampled from graphons. We consider Graphon Neural Differential Equations (Graphon-NDEs) and adjoint Graphon-NDEs as the infinite-node limits of the forward and adjoint GNDE systems, and establish well-posedness. For an $n$-node random graph with sparsity parameter $\alpha_n$, we prove trajectory-wise convergence of GNDE solutions to Graphon-NDE solutions at rate $O((\alpha_n n)^{-1/2})$, up to logarithmic factors, with high probability. We also establish uniform-in-time convergence bounds for adjoint systems governing hidden-state and parameter gradients. We further study discretize-then-optimize (DTO) and optimize-then-discretize (OTD) training. Under explicit Euler discretization with $M$ steps, we show that DTO and OTD are asymptotically consistent, with hidden-state and local parameter-gradient discrepancies of orders $O(1/M)$ and $O(1/M^2)$, respectively, up to sparsity and logarithmic factors. Experiments on HSBM and tent graphons support the theoretical rates, while zero-shot transfer experiments across four graphon classes demonstrate accurate deployment of learned GNDEs on larger independently sampled graphs.

13:00 JST研究/論文

TGHE: エッジクラウド システムでプライバシーを保護する GNN 推論のためのテンプレートベースのグラフ準同型暗号化

既存の準同型暗号化 (HE) ベースの GNN システムは、クエリごとのコストをグローバル グラフ サイズに結び付けるグラフ中心のパラダイムを採用しており、評価を最大で約 20,000 ノードに制限し、動的で大規模な財務グラフと互換性がありません。私たちは、テンプレート現象を利用することでこれを解決する自己中心的なフレームワークである TGHE (テンプレートベースのグラフ準同型暗号化) を提案します。つまり、トランザクション グラフ内のローカル計算ツリーが小さな構造形状のセットに収束します。 TGHE は、エッジでエゴグラフを正規化し、構造的に同一のツリーを共有 CKKS 暗号文にパックして SIMD 並列暗号化推論を実現します。2 つのロングテール オプティマイザー (近似テンプレート フィッティングとトポロジ コラプス) により、SIMD の完全なカバレッジが保証されます。 DGraphFin (370 万ノード、430 万エッジ) では、TGHE-Collapse は、AUC 損失が 0.002 未満で、シーケンシャル暗号化ベースラインと比較して 66.9 倍の高速化を達成します。

原文 (English)

TGHE: Template-based Graph Homomorphic Encryption for Privacy-Preserving GNN Inference in Edge-Cloud Systems

Existing homomorphic encryption (HE)-based GNN systems adopt a graph-centric paradigm that couples per-query cost to global graph size, limiting evaluations to at most ~20k nodes and making them incompatible with dynamic, large-scale financial graphs. We propose TGHE (Template-based Graph Homomorphic Encryption), an ego-centric framework that resolves this by exploiting a template phenomenon: local computation trees in transaction graphs converge into a small set of structural shapes. TGHE canonicalizes ego-graphs at the edge and packs structurally identical trees into shared CKKS ciphertexts for SIMD-parallel encrypted inference, with two long-tail optimizers (Approximate Template Fitting and Topology Collapse) ensuring full SIMD coverage. On DGraphFin (3.7M nodes, 4.3M edges), TGHE-Collapse achieves a 66.9x speedup over the sequential encrypted baseline with less than 0.002 AUC loss.

13:00 JST画像/動画生成

Disco-LoRA: マルチコンセプトのビデオカスタマイズのためのコンテンツ、スタイル、モーションのもつれのない構成

Text-to-Video (T2V) モデルに基づくビデオのカスタマイズは、参照データから特定の機能を学習して、制御可能なビデオを生成することを目的としています。画像のスタイル化とビデオ モーションのカスタマイズは大幅に進歩しましたが、コンテンツ、スタイル、モーションなどの複数の概念を同時に制御することは依然として大きな課題です。この作業では、コンテンツ、スタイル、モーションの共同制御が必要な、マルチコンセプトのビデオカスタマイズのタスクを体系的に定義します。この分野の研究を促進するために、私たちは包括的なベンチマークを構築し、さまざまな概念を 2 段階で解きほぐし、柔軟に再結合することでこの問題に取り組むように設計された統一フレームワークである Disco-LoRA を提案します。 (1) 目的を 2 つのサブタスク、コンテンツ スタイルとコンテンツ モーションに分解します。各サブタスクは、データ内の異なる概念を効果的に解きほぐす反復デュアル LoRA 解きほぐしフレームワークを使用して対処されます。 (2) 重みの大きさが構成可能性を決定する一方で、レイヤーごとの重みの傾向が LoRA のアイデンティティにとって重要であることを特定します。これらのスケールを調和させるために、重み分布を調整し、異なる LoRA 間の干渉を最小限に抑えながら層ごとの傾向を維持する Z スコアベースの統計的正則化を提案します。広範な実験により、Disco-LoRA はマルチコンセプトのビデオのカスタマイズに優れており、外観、スタイル、モーションを効果的に保持して制御可能なテキストからビデオへの生成を行うことが示されています。

原文 (English)

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization

Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos. While significant strides have been made in image stylization and video motion customization, simultaneously controlling multiple concepts, such as content, style, and motion, remains a major challenge. In this work, we systematically define the task of multi-concept video customization, which requires the joint control of content, style, and motion. To facilitate research in this area, we construct a comprehensive benchmark and propose Disco-LoRA, a unified framework designed to tackle this problem by disentangling and flexibly recombining different concepts in two stages: (1) We decompose the objective into two sub-tasks: Content-Style and Content-Motion. Each sub-task is addressed using our Iterative Dual-LoRA Disentanglement Framework, which effectively disentangles distinct concepts within the data. (2) We identify layer-wise weight trends as crucial for LoRA identity, while weight magnitudes dictate composability. To harmonize these scales, we propose a Z-score-based statistical regularization that aligns weight distributions, preserving layer-wise trends while minimizing interference between different LoRAs. Extensive experiments show that Disco-LoRA excels in multi-concept video customization, effectively preserving appearance, style, and motion for controllable text-to-video generation.

13:00 JSTLLM/生成AI

論理形式を超えて: 誤謬分類のための LLM 抽出パターン

今日のペースの速い情報時代では、推論の欠陥パターンとして定義される論理的誤りが、必然的に情報障害の増大に寄与します。ただし、誤謬は微妙な形で現れることが多く、自動分類が複雑になります。この研究では、抽象的な論理構造と文脈レベルの言語的手がかりを結合することが誤謬の分類に有益であるかどうかを調査し、大規模言語モデル (LLM) を使用して誤った例とその説明からそのようなパターンを帰納的に抽出するフレームワークを開発します。さまざまな LLM および実験的なゼロショットおよびワンショット構成にわたるこれらのパターンの影響を評価し、ゼロショット ベースラインと比較して統計的に有意な改善が見られ、競合するアプローチを上回るパフォーマンスを示しています。データセット間の実験により一般化が検証され、データ駆動型のパターン抽出が論理表現を生成する効果的な方法として確立されます。

原文 (English)

Beyond Logical Forms: LLM-Extracted Patterns for Fallacy Classification

In today's fast-paced information era, logical fallacies, defined as defective patterns of reasoning, inevitably contribute to the growth of information disorder. However, often fallacies appear in nuanced forms that complicate automated classification. In this study, we investigate whether merging abstract logical structures with context-level linguistic cues proves beneficial for fallacy classification, developing a framework that inductively extracts such patterns from fallacious examples and their explanations using Large Language Models (LLMs). We evaluate the impact of these patterns across different LLMs and experimental zero- and one-shot configurations, showing statistically significant improvements over zero-shot baselines and outperforming competing approaches. Cross-dataset experiments validate generalization, establishing data-driven pattern extraction as an effective method for generating logical representations.

13:00 JSTロボティクス

乱雑な環境における点群からモーションの実現可能性を学習する

動作の実現可能性の予測は、ロボット工学、特にタスクと動作の計画と操作において中心的な役割を果たします。乱雑な環境におけるこの問題の主なボトルネックは、サンプリングベースのモーション プランナー (SBMP) による実行不可能な計画の試みにより、多大な計算コストが発生する可能性があることです。また、実現不可能性を証明するための既存のアプローチは、低次元構成空間に限定されており、多くの場合、既知のパラメーターを持つプリミティブ オブジェクトによって表される単純化された幾何学的環境を想定しています。私たちは、現実的な乱雑なシーンで動作する 7-DOF マニピュレータの生の RGB-D 観察から直接動作実現可能性予測を学習するという補完的な問題を研究します。この設定に対する最初の大規模ベンチマークを紹介します。これは、88 個のスキャンされたオブジェクトと 190 個の乱雑なテーブルトップ シーンにわたる 270 万個の把握実現可能性ラベルで構成されます。一致したトレーニング条件の下で、MLP ベース、ボリューム CNN、およびポイントクラウド ベースの Transformer アーキテクチャにわたる 3 つの代表的な分類子ファミリーのベンチマークを行います。当社の最良のモデルである GRASPFC-PTX (点群変換器) は、新規オブジェクトで 0.996 の AUROC を達成しながら、SBMP よりも大幅に高速な予測を提供します。

原文 (English)

Learning Motion Feasibility from Point Clouds in Cluttered Environments

Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation. A major bottleneck for this problem in cluttered environments is that infeasible planning attempts by Sampling-based motion planners (SBMPs) can incur substantial computational cost. Also existing approaches for infeasibility certification are limited to low-dimensional configuration spaces and often assume simplified geometric environments represented by primitive objects with known parameters. We study the complementary problem of learning motion feasibility prediction directly from raw RGB-D observations for a 7-DOF manipulator operating in realistic cluttered scenes. We introduce the first large-scale benchmark for this setting, comprising 2.7M grasp feasibility labels over 88 scanned objects and 190 cluttered tabletop scenes. We benchmark three representative classifier families spanning MLP- based, volumetric-CNN, and point-cloud-based Transformer architectures under matched training conditions. Our best model, GRASPFC-PTX (a point-cloud transformer), achieves an AUROC of 0.996 on Novel objects while providing predictions significantly faster than SBMPs.

13:00 JST研究/論文

深層学習のアルゴリズムの基礎: 複雑さの理論的速度と普遍近似の特性評価

フィードフォワード ニューラル ネットワーク (NN) の表現力は、通常、最適な基底拡張スキームをエミュレートすることによって研究されます。この視点は強力ではありますが、不完全です。主に規則性を通じて複雑さを捉えるため、平方根関数や典型的なブラウン パスなど、同等の規則性を持つ直感的に単純なオブジェクトと複雑なオブジェクトを区別しません。指針となるメッセージは、ニューラル ネットワークを柔軟な基底関数としてだけでなく、計算モデルとしても見るべきだということです。関数が所定の基本ゲート言語を介して実数値回路によって計算可能である場合、その関数は、深さ、幅、ゲート数、およびゲート構造によって制御される明示的な深さ、幅、およびゼロ以外のパラメーター境界を備えた NN によって同等の精度で計算できます。したがって、ニューラル ネットワークの複雑さは規則性だけによって決まるのではなく、アルゴリズムの複雑さによっても決まります。次に、自然な並列化条件を満たし、注意や層の正規化などの多変量の非線形性を許容する定義可能な NN モデルは、非アフィン非線形性を含む場合に限り、ユニバーサル近似器であることを示します。私たちの理論の範囲は、連続関数の汎用近似保証、Besov クラスのミニマックス最適近似保証、正則関数の対数誤差計算量を推定することによって、また、NN がアーキテクチャ固有の引数なしでニュートン ラフソン根探索やべき乗反復などの数値アルゴリズムをエミュレートできることを示すことによって説明されます。その精度は、$k$-頂点グラフの最短経路計算によって示されます。トロピカル動的計画法回路をコンパイルすると、O(log(1/{\epsilon})) 非ゼロパラメータを持つ NN が生成され、定数 c>0 の場合、一般的な $O({\epsilon}^{-c k^2})$ Lipschitz 近似スケールよりも 1/{\epsilon} で指数関数的に向上します。

原文 (English)

Algorithmic Foundations of Deep Learning: Complexity-Theoretic Rates and a Characterization of Universal Approximation

Feedforward neural network (NN) expressivity is typically studied by emulating optimal basis-expansion schemes. While powerful, this perspective is incomplete: it primarily captures complexity through regularity, and therefore does not distinguish intuitively simple and complicated objects with comparable regularity, such as the square-root function and a typical Brownian path. The guiding message is that neural networks should be viewed not only as flexible basis functions, but also as models of computation. If a function is computable by a real-valued circuit over a prescribed elementary gate language, then it can be computed to comparable accuracy by an NN with explicit depth, width, and non-zero-parameter bounds controlled by the depth, width, gate count, and gate structure. Thus, neural-network complexity is not governed by regularity alone, but also by algorithmic complexity. We then show that any definable NN model satisfying a natural parallelization condition, allowing possibly multivariate non-linearities such as attention or layer normalization, is a universal approximator if and only if it contains a non-affine nonlinearity. The scope of our theory is illustrated by deducing universal approximation guarantees for continuous functions, minimax-optimal approximation guarantees for Besov classes, logarithmic-error complexity for holomorphic functions, and by showing that NNs can emulate numerical algorithms such as Newton-Raphson root finding and power iteration without architecture-specific arguments. Its precision is illustrated by shortest-path computation on $k$-vertex graphs: compiling the tropical dynamic-programming circuit yields NNs with O(log(1/{\epsilon})) non-zero parameters, exponentially improving in 1/{\epsilon} over the generic $O({\epsilon}^{-c k^2})$ Lipschitz-approximation scale, for a constant c>0.

13:00 JST画像/動画生成

MLFFM-SegDiff: 皮膚病変セグメンテーションのためのマルチレベル特徴融合拡散モデル

皮膚病変のセグメンテーションは、コンピュータ支援皮膚科学診断における重要なタスクであり、精度が下流の分析と疾患分類に直接影響します。ただし、ダーモスコピー画像は、境界がぼやけ、コントラストが低く、形状の大きな変化があり、髪の毛や影などのアーチファクトがあるため、困難です。最近、拡散モデルは、その漸進的ノイズ除去機能と分布モデリング機能のおかげで、医療画像のセグメンテーションにおいて優れたパフォーマンスを示しています。それにもかかわらず、既存の拡散ベースの方法は依然として、レベル間の特徴相互作用が制限され、境界詳細の回復が不十分であるという問題に悩まされています。これらの問題に対処するために、皮膚病変セグメンテーションのためのマルチレベル特徴融合拡散モデルである MLFFM-SegDiff を提案します。拡散フレームワーク上に構築されたこの方法では、デュアルパス U-Net エンコーダー、マルチレベル特徴融合モジュール (MLFFM)、および境界依存損失関数が導入されています。デュアルパス エンコーダは、ノイズの多いマスクの特徴とダーモスコピー画像の特徴の間の相互作用を強化します。 MLFFM は、アテンション、スケール調整、適応型クロスレベル フュージョンによってスキップ接続を改善します。これらの設計により、デコーダは浅い境界キューと深いセマンティック表現を共同で利用できるようになり、マスク再構成の品質が向上します。 ISIC2018、PH2、および HAM10000 での実験では、MLFFM-SegDiff が、精度、F1 スコア、Jaccard インデックス、リコール、およびダイス全体で、DermoSegDiff、U-Net、SwinUNETR などの代表的な手法よりも優れていることが実証されています。特に、平均 Jaccard インデックス 0.8546 と Dice 係数 0.9207 を達成します。これらの結果は、病変セグメンテーションのパフォーマンスを向上させるための、提案されたマルチレベル特徴融合戦略の有効性を検証します。コードは公開後に https://github.com/Qacket/MLFFM-SegDiff.git で公開されます。

原文 (English)

MLFFM-SegDiff: A Multi-Level Feature Fusion Diffusion Model for Skin Lesion Segmentation

Skin lesion segmentation is a key task in computer-aided dermatological diagnosis, where accuracy directly impacts downstream analysis and disease classification. However, dermoscopic images are challenging due to blurred boundaries, low contrast, large shape variations, and artifacts such as hair and shadows. Recently, diffusion models have shown strong performance in medical image segmentation thanks to their progressive denoising and distribution modeling capabilities. Nevertheless, existing diffusion-based methods still suffer from limited cross-level feature interaction and insufficient boundary detail recovery. To address these issues, we propose MLFFM-SegDiff, a multi-level feature fusion diffusion model for skin lesion segmentation. Built on a diffusion framework, the method introduces a dual-path U-Net encoder, a Multi-Level Feature Fusion Module (MLFFM), and a boundary-sensitive loss function. The dual-path encoder enhances interaction between noisy mask features and dermoscopic image features. MLFFM improves skip connections via attention, scale alignment, and adaptive cross-level fusion. These designs enable the decoder to jointly leverage shallow boundary cues and deep semantic representations, improving mask reconstruction quality. Experiments on ISIC2018, PH2, and HAM10000 demonstrate that MLFFM-SegDiff outperforms representative methods including DermoSegDiff, U-Net, and SwinUNETR across Accuracy, F1-score, Jaccard index, Recall, and Dice. In particular, it achieves an average Jaccard index of 0.8546 and Dice coefficient of 0.9207. These results validate the effectiveness of the proposed multi-level feature fusion strategy for improving lesion segmentation performance. The code will be released at https://github.com/Qacket/MLFFM-SegDiff.git after publication.

13:00 JST画像/動画生成

堅牢なタマネギ: ノイズの下で開いた語彙オブジェクト検出器を剥がす

Open Vocabulary Object Detector (OV-OD) に対する現実世界のノイズの影響は、そのアーキテクチャの複雑さのため、依然としてよく理解されていません。私たちは、制御された合成視覚的劣化を使用して OV-OD を層ごとに剥離する実証研究である包括的な分析 Robust Onion を紹介します。これにより、ロバスト性がどのように、なぜ、どこで劣化するかを明らかにし、特徴の崩壊を系統的に分析します。私たちの調査結果では、同様の視覚バックボーンを持つモデルは、同様のレイヤーでの同様の機能崩壊によって駆動され、同等の堅牢性を示す一方、事前トレーニング戦略、アーキテクチャのニュアンス、キャプションの監視などの要因はほとんど寄与していないことが明らかになりました。堅牢性は主にアノテーションではなく画像ドメインによって支配されており、これは COCO と LVIS に対する同様の堅牢性の影響と、なぜ ODinW-13 のようなデータセットが大きく孤立したオブジェクトによって堅牢性が誇張されている印象を与える可能性があるかを説明しています。最後に、エンドツーエンドのトレーニングに比べて 96 分の 1 少ないトレーニング可能なパラメータを使用し、軽量のプラグアンドプレイ NN および TK0 アプローチを通じて、実際の BDD100K、 WiderFace、VisDRONE での堅牢性を向上させることで洞察を検証します。また、先行研究のロバスト性の観察についても説明します。

原文 (English)

Robust Onion: Peeling Open Vocab Object Detectors Under Noise

The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis Robust Onion, an empirical study that uses controlled synthetic visual degradations to peel OV-ODs layer-by-layer, revealing how, why, and where robustness degrades, systematically analyzing feature collapse. Our findings reveal that models with similar vision backbones exhibit comparable robustness, driven by similar feature collapse at similar layers, while factors such as pretraining strategy, architectural nuances, and caption supervision contribute little. Robustness is primarily governed by the image domain rather than annotations, explaining the similar robustness impact on COCO and LVIS, and why datasets like ODinW-13 can give an impression of inflated robustness due to large, isolated objects. Finally, we validate our insights by improving robustness on real-world BDD100K, WiderFace, and VisDRONE via our lightweight plug-and-play NN & TK0 approach, using 96x fewer trainable parameters than end-to-end training. We also explain the prior works' robustness observations.

13:00 JST画像/動画生成

制御可能な 4D 心臓 MRI 合成のための解剖学に基づく残留運動拡散

4D (3D + 時間) 医用画像用の堅牢な人工知能モデルの開発には、注釈付きデータの制限、デバイス間のドメインのシフト、プライバシー制限による制約があります。これに対処するために、解剖学的に一貫したデータ拡張のための 4D 制御可能な生成フレームワークを提案します。半教師あり変分オートエンコーダは、統一されたフレームワークで位置合わせされたセグメンテーション マスクを共同で予測しながら、解剖学的ボリュームのコンパクトな潜在表現を学習します。その後、カスケード潜在拡散モデル (LDM) を通じて、解剖学的構造が時間的ダイナミクスから解きほぐされます。静的 LDM は臨床事前情報 (診断と体積測定) に基づいて条件付けされた被験者固有の解剖学的構造を生成し、その後のモーション LDM は残存する潜在動作を推定し、4D シーケンス全体にわたる厳密な時間的一貫性を確保します。提案されたアプローチは、代表的な 4D イメージング アプリケーションとしてシネ心臓 MRI で評価されました。複数のデータセットにわたる実験により、静的解剖学的構造の高い制御性 (ピアソン r > 0.8) と強力な時間的一貫性 (FVD = 288.08) が実証されました。クロスベンダー汎化実験では、合成 4D シーケンスを使用してトレーニング セットを強化すると、下流のセグメンテーションのパフォーマンスが大幅に向上します。 nnU-Net を使用することで、提案された拡張戦略により、実際のデータのみでのトレーニングと比較して、平均 Dice スコアが 1.4% 改善され、ハウスドルフ距離が 3.0 mm 減少しました。左心室の場合、Dice は 2.8% 改善され、境界誤差は 5.4 mm 減少しました。全体として、このフレームワークは 4D 医用画像合成のためのスケーラブルで制御可能なソリューションを提供し、限定されたアノテーションとベンダー間の変動性を備えたより堅牢なモデルの開発をサポートします。コードは https://github.com/cyiheng/4DCardiacMRISynthesis で入手できます。

原文 (English)

Anatomy-Guided Residual Motion Diffusion for Controllable 4D Cardiac MRI Synthesis

Developing robust artificial intelligence models for 4D (3D + time) medical imaging is constrained by limited annotated data, inter-device domain shifts, and privacy restrictions. To address this, we propose a 4D controllable generative framework for anatomically consistent data augmentation. A semi-supervised variational autoencoder learns a compact latent representation of anatomical volumes while jointly predicting aligned segmentation masks in a unified framework. Anatomical structure is then disentangled from temporal dynamics through a cascaded latent diffusion model (LDM). A static LDM generates subject-specific anatomy conditioned on clinical priors (diagnosis and volumes measures) and a subsequent motion LDM estimates residual latent motions, ensuring strict temporal coherence across the 4D sequence. The proposed approach was evaluated on cine cardiac MRI as a representative 4D imaging application. Experiments across multiple datasets demonstrate high controllability of static anatomy (Pearson r > 0.8) and strong temporal coherence (FVD = 288.08). In cross-vendor generalization experiments, augmenting training sets with synthetic 4D sequences significantly improves downstream segmentation performance. Using nnU-Net, the proposed augmentation strategy improves the average Dice score by 1.4% and reduces the Hausdorff Distance by 3.0mm compared to training on real data alone, for the left ventricle, Dice improves by 2.8% with a 5.4mm reduction in boundary error. Overall, this framework provides a scalable and controllable solution for 4D medical image synthesis, supporting the development of more robust models with limited annotations and cross-vendor variability. Code available on https://github.com/cyiheng/4DCardiacMRISynthesis.

13:00 JSTLLM/生成AI

AIGP: 電子商取引の価格設定における長期的な価値調整のための LLM ベースのフレームワーク

大規模電子商取引における従来の動的価格設定モデルは、解釈可能性が限られていること、非構造化情報の活用が不十分であること、累積総商品価値 (GMV)、投資収益率 (ROI)、マイルストーンの達成などの長期的なビジネス目標との不一致に悩まされています。私たちは、ドメイン知識、構造化データ、およびテキストのコンテキストに基づいて大規模言語モデル (LLM) を活用し、解釈可能で知識を意識した価格決定を行う新しいフレームワークである AIGP を提案します。高品質の出力を維持しながら効率的に展開するために、知識の蒸留に監視付き微調整を採用しています。 AIGP の中心となるのは、履歴データに基づくオフライン強化学習を通じてトレーニングされた長期価値推定ツール (LTVE) です。LTVE は、候補の価格設定アクションをスコアリングし、Direct Preference Optimization (DPO) の優先ペアを選択するための報酬モデルとして機能し、それによって価格設定ポリシーを長期的なビジネス目標に合わせることができます。 Tao Factory での広範なオフライン評価と大規模なオンライン A/B テストにより、AIGP が生産ベースラインと比較して 14 日間の GMV で +13.21%、ROI で +7.59%、マイルストーン達成率で +8.20% という大幅な改善を達成すると同時に、解釈可能で透明性のある価格設定の根拠を提供していることが実証されました。

原文 (English)

AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing

Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured information, and misalignment with long-term business objectives such as cumulative Gross Merchandise Value (GMV), Return on Investment (ROI) and milestone achievement. We propose AIGP, a novel framework that leverages a Large Language Model (LLM) prompted with domain knowledge, structured data and textual context to make interpretable, knowledge-aware pricing decisions. For efficient deployment while maintaining high-quality outputs, we employ supervised fine-tuning for knowledge distillation. Central to AIGP is the Long-Term Value Estimator (LTVE), trained via offline reinforcement learning on historical data, which serves as a reward model to score candidate pricing actions and select preference pairs for Direct Preference Optimization (DPO), thereby aligning the pricing policy with long-term business objectives. Extensive offline evaluations and large-scale online A/B tests on Tao Factory demonstrate that AIGP achieves significant improvements: +13.21% in GMV, +7.59% in ROI, and +8.20% in milestone achievement rate over 14 days compared to the production baseline, while simultaneously providing interpretable and transparent pricing rationales.

13:00 JSTLLM/生成AIエージェント

ミラー: Agentic RAG 向けの新規性制約のあるメモリ ガイド付き MCTS レッド チーム化

マルチモーダルなエージェント検索拡張生成 (RAG) システムは、プロンプト インジェクションを超えて、テキスト ポイズニング、イメージ インジェクション、直接クエリ攻撃、オーケストレーター レベルのツール操作など、攻撃対象領域を拡大します。既存のレッドチームアプローチは通常、サーフェス固有であり、既知の攻撃テンプレートを再利用することがよくあります。テキストポイズニングベンチマークでは、73 ~ 84% の正確な重複が測定されました。我々は、明示的な新規性制約の下で、取得されたコンテキストに基づいて候補生成を条件付けしながら、メモリに基づくモンテカルロ木検索を実行する統合クロスサーフェスフレームワークである MIRROR を紹介します。決定論的ノベルティ ゲートは、正規化された比較に基づいて検索セットに一致する候補を拒否するため、プロンプト コピーを有効にすることなく、検索で事前検索に通知することができます。マルチモーダル エージェントの RAG ターゲット上の 4 つの攻撃サーフェス全体で、MIRROR はイメージ ポイズニングに関してベースラインの 52% と比較して 76% の ASR を達成し、クエリ コストの半分でオーケストレーター攻撃に対して 97% の ASR を達成し、クロスサーフェス間の分散は最小 (変動係数 0.47) を達成しました。対照的に、特殊なベースラインはサーフェス全体で崩壊します。サフィックスの最適化は、テキスト ポイズニングでは 79% の ASR に達しますが、直接クエリでは 1% に達します。 41,815 個のパッケージ内レコードとランタイム アダプターを備えた ART-SafeBench をリリースし、4 つのサーフェスにわたって合計 41,991 以上のレコードを生成します。

原文 (English)

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-teaming approaches are typically surface-specific and often recycle known attack templates; on text-poisoning benchmarks we measure 73-84% exact duplication. We present MIRROR, a unified cross-surface framework that performs memory-guided Monte Carlo tree search while conditioning candidate generation on retrieved context under an explicit novelty constraint. A deterministic Novelty Gate rejects any candidate matching the retrieval set under normalized comparison, allowing retrieval to inform search priors without enabling prompt copying. Across four attack surfaces on a multimodal agentic RAG target, MIRROR attains 76% ASR on image poisoning compared with 52% for baselines, 97% ASR on orchestrator attacks at half the query cost, and the lowest cross-surface variance (coefficient of variation 0.47). In contrast, specialized baselines collapse across surfaces: suffix optimization reaches 79% ASR on text poisoning but 1% on direct queries. We release ART-SafeBench with 41,815 in-package records and runtime adapters yielding 41,991+ total records across four surfaces.

13:00 JST画像/動画生成

ReasonCLIP-58M: CLIP のための視覚的に根拠のある常識推論の監修

CLIP とそのバリアントは、マルチモーダル システムで視覚的なバックボーンとして広く採用されていますが、その事前トレーニングは依然として説明的な画像とテキストの位置合わせによって占められています。下流のアプリケーションでは、視覚に基づいた常識的な推論や構成推論の要求がますます高まっているため、CLIP スタイルのエンコーダがアーキテクチャを変更せずにそのような推論をサポートできるかどうかは依然として不明です。これに対処するために、我々は ReasonCLIP-58M を紹介します。これは、2 段階の戦略を通じて大規模な推論の監視を CLIP スタイルのモデルに統合する継続的な事前トレーニング フレームワークです。説明的な整合性を維持しながら推論信号を段階的に統合し、その後にカテゴリー構造の推論の監視が続きます。このフレームワークをサポートするために、2 つの相補的なデータセットとベンチマークを構築します。オープン形式で視覚的に検証可能な推論キャプションを備えた ReasonLite-42M。 ReasonPro-16M、カテゴリ固有の推論監視機能付き。視覚に基づいた推論の診断評価のための RCLIP-Bench。私たちは、ゼロショット検索のパフォーマンスを向上させながら、視覚に基づいた常識と構成的推論を改善するReasonCLIPファミリーをトレーニングします。 ReasonCLIP は、LLaVA-NeXT などのマルチモーダル大規模言語モデル用のドロップイン ビジュアル エンコーダとして、追加の推論コストなしで一貫した利益を提供し、構造化された推論の監視によって CLIP スタイルのビジュアル表現の表現能力が強化されることを示しています。すべてのデータセット、モデル、トレーニング コードは https://github.com/RISys-Lab/ReasonCLIP で入手できます。

原文 (English)

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually grounded commonsense inference and compositional reasoning, it remains unclear whether CLIP-style encoders can support such reasoning without architectural changes. To address this, we present ReasonCLIP-58M, a continual pretraining framework that integrates large-scale reasoning supervision into CLIP-style models through our two-stage strategy, which progressively integrates reasoning signals while preserving descriptive alignment, followed by category-structured reasoning supervision. To support this framework, we construct two complementary datasets and a benchmark: ReasonLite-42M, with open-form, visually verifiable reasoning captions; ReasonPro-16M, with category-specific reasoning supervision; and RCLIP-Bench for diagnostic evaluation of visually grounded reasoning. We train a family of ReasonCLIP that improves visually grounded commonsense and compositional reasoning while also enhancing zero-shot retrieval performance. As a drop-in visual encoder for multimodal large language models such as LLaVA-NeXT, ReasonCLIP delivers consistent gains without additional inference cost, demonstrating that structured reasoning supervision enhances the expressive capacity of CLIP-style visual representations. All datasets, models, and training code are available at https://github.com/RISys-Lab/ReasonCLIP.

13:00 JST画像/動画生成Sora

NaviCache: ビデオ生成のためのテスト時の自己調整キャッシング

ビデオ拡散モデル (VDM) は、膨大な計算コストによる制約を受けます。オフライン キャリブレーション ベースの高速化には、キャリブレーション データの依存性、法外なキャリブレーション期間、分布シフトの影響を受けやすいという問題がありますが、オフライン キャリブレーションを必要としない方法では、これらの障害が解消されます。ただし、入力と出力の差の間のマッピングがリアルタイムで変化する瞬間的なゼロ次近似に依存しているため、観測ノイズの影響を受けやすく、拡散軌跡内の固有運動量が無視されます。この論文では、機能進化を慣性航法システム (INS) 問題として再概念化する、プラグ アンド プレイのテスト時自己校正手法である NaviCache を提案します。 NaviCache は、入力と出力の変動の間の相対的な結合をモデル化することで、基本的なドメイン ギャップと拡散の非定常的な性質を橋渡しします。特徴変化率とその潜在ドリフトを適応的に追跡するデュアルステート推定アーキテクチャを導入し、特殊な初期調整フェーズを通じて初期化します。 NaviCache は、時間依存のノイズ スケジュールを不確実性を考慮した測定更新メカニズムと統合することにより、誤差制限のある計算スキップのための理論に基づいたメカニズムを提供します。 HunyuanVideo、Wan、Open-Sora シリーズでの広範な実験により、NaviCache が計算スキップに対してより正確なエラー判定を示し、優れた総合パフォーマンスを達成することが実証されました。

原文 (English)

NaviCache: Test-Time Self-Calibration Caching for Video Generation

Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from calibration data dependency, prohibitive calibration duration, and susceptibility to distribution shifts, offline calibration-free methods eliminate these hurdles. However, since they rely on instantaneous zero-order approximations where the mapping between input and output differences varies in real-time, they are susceptible to observational noise and ignore the intrinsic momentum within the diffusion trajectory. In this paper, we propose NaviCache, a plug-and-play test-time self-calibration method re-conceptualizing feature evolution as an Inertial Navigation System (INS) problem. NaviCache bridges the fundamental domain gap and the non-stationary nature of diffusion by modeling the relative coupling between input and output variations. We introduce a dual-state estimation architecture that adaptively tracks the feature change ratio and its latent drift, initialized via a specialized Initial Alignment phase. By integrating a time-dependent noise schedule with an uncertainty-aware Measurement Update mechanism, NaviCache provides a theoretically grounded mechanism for error-bounded computation skipping. Extensive experiments on the HunyuanVideo, Wan, and Open-Sora series demonstrate that NaviCache exhibits more accurate error judgment for computation skipping and achieves outstanding comprehensive performance.

13:00 JST研究/論文OpenAI

要塞とゲートキーパー: サードパーティのサイバーセキュリティ リスク ガバナンスにおける推移的信頼の理論化

分析プラットフォーム、クラウド サービス、アイデンティティ プロバイダー、ソフトウェア サプライヤーなどのサードパーティ ベンダーがデジタル サービスの提供に組み込まれることが増えています。これらの取り決めにより、規模の拡大と専門化が可能になる一方で、顧客のデータとセキュリティ関連の慣行を、顧客がほとんど目にしたり、選択したり、評価したりしない環境に移行することもできます。このペーパーでは、2025 年 11 月の OpenAI-Mixpanel セキュリティ インシデントの文書分析を通じてこの問題を検証します。このインシデントは、ベンダー環境におけるセキュリティ イベントが、顧客との関係を維持する中心組織のガバナンスと説明責任の問題にどのようになり得るかを示す実例として機能します。この論文は、組織の信頼調査とエージェンシー理論に基づいて、サードパーティのサイバーセキュリティリスクは信頼関係と委任の問題の両方であると主張しています。顧客は目に見えるサービスプロバイダーを信頼しますが、プロバイダーはセキュリティ慣行が部分的にしか見えず制御できないベンダーに依存しています。この論文では、デジタル サービスに対する顧客の信頼が、そのサービス プロバイダーによって認可されたベンダーのセキュリティ慣行に依存するという推移的信頼の概念を開発しています。次に、フォートレスとゲートキーパーのフレームワークを紹介します。これは、正式な組織所有権だけではなく、信頼とデータ フローを通じてサイバーセキュリティ ガバナンスの境界を説明します。この分析により、ベンダー統合、メタデータ公開、ベンダー保証、データ拡散に関する 4 つの提案が展開されます。この論文は、委任されたデータ処理がどのようにして顧客に対する説明責任を生み出すのかを説明し、ベンダーの階層化、データ分類、契約設計、継続的保証、データの最小化への影響を特定することにより、サイバーセキュリティ ガバナンスの研究に貢献しています。

原文 (English)

Fortress and Gatekeeper: Theorizing Transitive Trust in Third-Party Cybersecurity Risk Governance

Third-party vendors, such as analytics platforms, cloud services, identity providers, and software suppliers, are increasingly embedded in digital service delivery. While these arrangements enable scale and specialization, they also move customer data and security-relevant practices into environments that customers rarely see, select, or evaluate. This paper examines this problem through a document analysis of the November 2025 OpenAI-Mixpanel security incident. The incident serves as an illustrative case for showing how a security event in a vendor environment can become a governance and accountability problem for the focal organization that maintains the customer relationship. Drawing on organizational trust research and agency theory, the paper argues that third-party cybersecurity risk is both a trust relationship and a delegation problem. Customers trust the visible service provider, while the provider relies on vendors whose security practices are only partially visible and controllable. The paper develops the concept of transitive trust, where customer trust in a digital service depends on the security practices of vendors authorized by that service provider. It then presents the Fortress and Gatekeeper framework, which explains cybersecurity governance boundaries through trust and data flows rather than formal organizational ownership alone. The analysis develops four propositions concerning vendor integration, metadata exposure, vendor assurance, and data proliferation. The paper contributes to cybersecurity governance scholarship by explaining how delegated data processing creates customer-facing accountability and by identifying implications for vendor tiering, data classification, contractual design, continuous assurance, and data minimization.

13:00 JSTLLM/生成AILlamaDeepSeek

長い推論のための情報を意識した KV キャッシュ圧縮

大規模言語モデル (LLM) では推論機能が急速に進歩しており、事前入力段階とデコード段階の両方でキー/値 (KV) キャッシュのサイズが増加しています。既存の KV キャッシュ圧縮方法は、主にアテンションの重みに依存してトークンの重要性を推定します。注意は文脈上の関連性を効果的に捉えますが、予測の不確実性とトークンの情報提供性に関連する補完的な情報理論的シグナルを見落とします。このペーパーでは、将来を見据えた観点からトークンの重要性を再考し、圧縮されたトークンが将来のコンテキストにどのような影響を与えるかを測定する指標である \textit{Forward Influence} を紹介します。私たちの分析により、注意スコアによって選択されたトークンは主に近くのコンテキストに影響を与えるのに対し、高い予測不確実性に関連付けられたトークンは、遠い将来のコンテキストに対して大幅に強い影響を示すことが明らかになりました。この観察に基づいて、情報理論信号を組み込んだエントロピーを意識した KV キャッシュ圧縮フレームワークである \textbf{InfoKV} を提案します。トークンレベルの予測不確実性とレイヤーごとの表現の進化を組み合わせ、結果として得られるエントロピースコアを推論中の注意スコアと統合します。 Llama-3.1、Llama-3.2、DeepSeek-R1 を使用したロングコンテキスト推論ベンチマークの実験では、長いプリフィルとデコードの両方のシナリオにおいて、InfoKV が既存のアテンションベースの KV 圧縮手法よりも一貫して優れていることが実証されました。

原文 (English)

Information-Aware KV Cache Compression for Long Reasoning

Reasoning capability has advanced rapidly in large language models (LLMs), leading to an increasing size of key-value (KV) cache in both prefilling and decoding stages. Existing KV cache compression methods mainly rely on attention weights to estimate token importance. While attention effectively captures contextual relevance, it overlooks complementary information-theoretic signals related to predictive uncertainty and token informativeness. In this paper, we revisit token importance from a forward-looking perspective and introduce \textit{Forward Influence}, a metric that measures how compressed tokens affect future contexts. Our analysis reveals that tokens selected by attention scores mainly influence nearby contexts, whereas tokens associated with high predictive uncertainty exhibit substantially stronger influence on distant future contexts. Based on the observation, we propose \textbf{InfoKV}, an entropy-aware KV cache compression framework that incorporates information-theoretic signals. It combines token-level predictive uncertainty with layer-wise representation evolution and integrates the resulting entropy scores with attention scores during reasoning. Experiments on long-context reasoning benchmarks with Llama-3.1, Llama-3.2, and DeepSeek-R1 demonstrate that InfoKV consistently outperforms existing attention-based KV compression methods in both long prefilling and decoding scenarios.

13:00 JST画像/動画生成

最適なトランスポート セマンティック フローによるビジョンと言語の概念の橋渡し

コンセプト ボトルネック モデル (CBM) は、人間が解釈できる概念を通じて予測することで透明性のある推論を約束しますが、その有効性は基本的に、視覚的表現とテキスト表現がどの程度適切に調整または一致しているかによって決まります。既存のビジョン言語 CBM は、事前に調整されたエンコーダやグローバル コサイン類似性に依存することが多く、これにより詳細な概念のローカライゼーションが曖昧になり、真のセマンティック ジオメトリを反映できなくなります。この研究では、概念の調整を静的な投影ではなく動的クロスモーダル輸送プロセスとして再考し、最適輸送フロー コンセプト ボトルネック モデル (OTF-CBM) を提案します。まず、逆最適トランスポートを介してデータ駆動型のセマンティック コストを学習してクロスモーダル距離を測定し、次にアンバランスな最適トランスポート ベースのフロー マッチングを実行して、ビジュアル パッチとテキスト概念の間のセマンティック遷移をモデル化します。速度ベースの概念のアクティベーションにより、OTF-CBM は ODE 統合なしで解釈可能な幾何学的関係を捕捉します。実験ではさらに、OTF-CBM が優れた分類精度と概念の忠実性を達成し、解釈可能なクロスモーダル推論のための新しい幾何学的および動的視点を提供することが示されています。

原文 (English)

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched. Existing vision-language CBMs often rely on pre-aligned encoders or global cosine similarity, which obscures fine-grained concept localization and fails to reflect true semantic geometry. In this work, we rethink concept alignment as a dynamic cross-modal transport process instead of static projection and propose the Optimal Transport Flow Concept Bottleneck Model (OTF-CBM). It first learns a data-driven semantic cost via Inverse Optimal Transport to measure cross-modal distances, and then performs unbalanced optimal-transport-based flow matching to model semantic transitions between visual patches and textual concepts. With velocity-based concept activation, OTF-CBM captures interpretable geometric relations without ODE integration. Experiments further show that OTF-CBM achieves superior classification accuracy and concept faithfulness, offering a new geometric and dynamical perspective for interpretable cross-modal reasoning.

13:00 JSTLLM/生成AIGemini

SamaVaani: インド言語の多言語臨床 ASR の監査とバイアス除去

自動音声認識 (ASR) は、臨床現場での遭遇を記録するためにますます使用されていますが、多言語で人口統計的に多様なインドの医療環境におけるその信頼性は、依然としてほとんど知られていません。この研究では、まずカンナダ語、ヒンディー語、インド英語にわたる実際の精神医学面接データに対して ASR パフォーマンスの体系的な監査を実施し、IndicWhisper、WhisperLargeV3、Sarvam、GoogleS2T、Gemma3n、OmniLingual、Vaani、Gemini を含む 8 つの最先端のモデルを比較します。私たちの結果では、モデルや言語によってかなりのばらつきがあり、一部のシステムはインド英語では競争力のあるパフォーマンスを発揮しますが、地域の音声では失敗することが明らかになりました。さまざまな方法を使用して、最もパフォーマンスの高い 2 つのオープンソース モデル、つまり Gemma3n と OmniLingual をさらに微調整します。これにより、話者の役割と性別に関連する体系的なパフォーマンスのギャップが明らかになり、臨床現場での公平な展開に対する懸念が生じますが、公平性を意識した微調整によってこの問題はさらに軽減されます。この目的を達成するために、ASR のパフォーマンスを向上させ、人口統計グループ全体の公平性を同時に向上させる統合されたバイアス除去技術である SamaVaani を提案します。

原文 (English)

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, we first conduct the systematic audit of ASR performance on real-world psychiatric interview data spanning Kannada, Hindi and Indian English, comparing eight state-of-the-art models including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini. Our results reveal substantial variability across models and languages, with some systems performing competitively in Indian English but failing in regional speech. We further fine-tune two of the best performing opensource models, i.e., Gemma3n and OmniLingual, using various methods. With this, we uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings, which are further mitigated by fairness-aware fine-tuning. To this end, we propose SamaVaani, a unified debiasing technique that simultaneously improves ASR performance and improves fairness across demographic groups.

13:00 JST画像/動画生成Gemini

確実なビデオ理解のための信頼性を意識したツール オーケストレーション

ビデオ推論言語モデルは、すべての入力フレームが同等に信頼できることを暗黙的に前提としています。これは、ブラインド トラスト問題と呼ばれるものにつながります。モーション ブラー、グレア、オクルージョンなどの現実的な摂動の下では、フロンティア ビデオ推論モデルは、視覚的証拠が劣化していることに気づかないまま、現実世界の具体化されたベンチマークで 15 ~ 30%p 精度が低下する可能性があります。この課題に対処するために、フレームごとの信頼性を推論のすべての段階に明示的に統合するエージェント型ビデオ理解フレームワークである Robust-TO を提案します。 Robust-TO は、統一された証拠インターフェイスの下で異種の視覚認識ツールを編成します。各ツールは、元の質問から派生したサブクエリと、信頼性関連スコアによって選択された信頼できるフレームのセットを受け取ります。具体的な予測 (境界ボックス、動きの軌跡、認識されたテキスト、アクション ラベルなど)、時間的根拠、および校正された信頼性スコアなどの証拠を共有形式で返します。推論中、これらの調整されたスコアは、3 層の合成プロセス (高/中/低) での証拠の重み付けをガイドし、正確性、証拠の信頼性、効率を共同で最適化する信頼コストの GRPO 報酬を定義します。 8 つのタスクにわたる 2 つのビデオ推論ベンチマークで、Robust-TO はクリーンな入力で 56.4% の平均精度を達成し、最強のオープンソース ベースラインを 10.6%p 上回り、Gemini-2.5-Pro (46.2%) を上回りました。現実的な 5 つの破損タイプの下で、Robust-TO は平均精度 54.3% を維持し、最も強力なオープンソースのベースラインを 5.8% 上回っていますが、比較したすべての方法の中でクリーンな状態から破損した状態への精度の低下が最小でした。

原文 (English)

Confidence-Aware Tool Orchestration for Robust Video Understanding

Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: under realistic perturbations such as motion blur, glare, or occlusion, frontier video reasoning models can suffer 15-30%p accuracy drops on real-world embodied benchmarks, while remaining unaware that their visual evidence has been degraded. To address this challenge, we propose Robust-TO, an agentic video understanding framework that explicitly integrates per-frame trustworthiness into every stage of reasoning. Robust-TO organizes heterogeneous visual perception tools under a unified evidence interface. Each tool receives a sub-query derived from the original question and a set of trustworthy frames selected by the reliability-relevance score. It returns evidence in a shared format: a concrete prediction (e.g., a bounding box, motion trajectory, recognized text, or action label), temporal grounding, and a calibrated reliability score. During reasoning, these calibrated scores guide evidence weighting in a three-tier synthesis process (high/medium/low) and define a confidence-cost GRPO reward that jointly optimizes correctness, evidence reliability, and efficiency. On two video reasoning benchmarks spanning eight tasks, Robust-TO achieves 56.4% average accuracy on clean inputs, surpassing the strongest open-source baseline by 10.6%p and outperforming Gemini-2.5-Pro (46.2%). Under five realistic corruption types, Robust-TO maintains 54.3% average accuracy, 5.8%p above the strongest open-source baseline, while exhibiting the smallest clean-to-corrupted accuracy drop among all compared methods.

13:00 JSTLLM/生成AI

GEOALIGN: 堅牢な LLM 強化学習のための幾何学的ロールアウト キュレーション

オンライン強化学習は、大規模言語モデル (LLM) を報酬信号に合わせて調整するために広く使用されていますが、ノイズの多い報酬や報酬の指定が間違っている場合にはトレーニングが不安定になる可能性があります。私たちは、方向性の不一致と呼ぶ失敗モードを特定します。バッチ内で、報酬の高いロールアウトの小さなセットが、バッチの大部分と大きく一致しない表現空間の優先方向を誘発し、その結果、分散が大きくなり、更新が不安定になります。私たちは、反復的なポリシー最適化におけるロールアウトキュレーションのための軽量プラグインである geoalign を提案します。 Geoalign は、(i) プロンプト内のプリファレンス ペアを形成し、(ii) 報酬に基づいた変位方向を集中させるためにロールアウトごとの隠れ状態でオンライン プロジェクターを学習し、(iii) バッチ コンセンサス プロトタイプからの角度偏差を介して方向的に一貫性のないロールアウトを検出し、プロンプト内の安定した代替案でそれらを修正します。 Geoalign はフォワードパスのみであり、追加されるオーバーヘッドは無視できます。学習された報酬モデルを使用した対話の調整と、バイナリ検証された報酬を使用した数学的推論を通じて、Geoalign は最終パフォーマンスを向上させ、トレーニングの変動を低減し、PF-PPO、PAR、PODS、および Seed-GRPO を上回るパフォーマンスを発揮します。これらの結果は、オンライン LLM RL の効果的な信頼性シグナルとしての潜在的な方向性のコンセンサスを示唆しています。

原文 (English)

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollouts induces representation-space preference directions that sharply disagree with the batch majority, resulting in high-variance and destabilizing updates. We propose geoalign, a lightweight plug-in for rollout curation in iterative policy optimization. Geoalign (i) forms within-prompt preference pairs, (ii) learns an online projector on per-rollout hidden states to concentrate reward-ordered displacement directions, and (iii) detects directionally inconsistent rollouts via their angular deviation from a batch consensus prototype and rectifies them with within-prompt stable alternatives. Geoalign is forward-pass only and adds negligible overhead. Across dialogue alignment with a learned reward model and mathematical reasoning with binary verified rewards, Geoalign improves final performance and reduces training oscillation, outperforming PF-PPO, PAR, PODS, and Seed-GRPO. These results suggest latent directional consensus as an effective reliability signal for online LLM RL.

13:00 JSTロボティクス

ドライバー状態ワールドモデリングを使用した、リスクを意識した選択的マルチモーダルドライバーモニタリング

自動運転車におけるドライバーの継続的な監視には、不確実なドライバーの状態下で危険な決定を回避しながら、低遅延の推論が必要です。大規模なビジョン言語モデルは広範なマルチモーダル事前分布を提供しますが、この設定では遅延と信頼性が限られているため、常時オンの機内モニターとしては適していません。私たちは、展開可能なマルチモーダルドライバー監視のためのコストを意識した選択的推論フレームワークを提案します。コア システムは、客室内の視覚観察と窓レベルの HR/EDA 信号を組み合わせた軽量の RGB 生理学的スチューデントであり、高速予測をいつ受け入れるか安全介入を控えるかを決定する学習済みゲートです。追加の制御により、学習されたスコアには事前シナリオを超えたサンプルレベルの情報が含まれていることが示されていますが、正確な生理学的同期には依然として制限があります。予測証拠を組み込むために、潜在的なドライバー状態特徴を展開し、将来の高速モデルエラーと反事実的なシステムレベルのアクションコストを推定するコンパクトなドライバー状態ワールドモデリングモジュールをさらに研究します。シナリオに起因するドライバー デマン​​ド認識では、RGB 生理学的学生は RGB のみおよび生理学的のみのベースラインを上回り、1,139 万のパラメーターと 3.08 ミリ秒の推論遅延で 0.7440 マクロ F1 と 0.9099 のバランスのとれた精度に達しました。コストを意識した選択推論により、展開レベルのレイテンシを維持しながら、安全でない誤検出が常時高速推論の 17.37% からシード全体で約 5% に減少します。ドライバー状態の世界モデリングは貴重な予測信号を提供しますが、ワースト グループの評価では永続的な動作点校正ドリフトが浮き彫りになります。最終的に、信頼性の高いエッジドライバー監視には、認識バックボーンの進歩だけでなく、リスクを意識した選択的制御とグループ堅牢なキャリブレーションも必要となります。

原文 (English)

Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling

Continuous driver monitoring in automated vehicles requires low-latency inference while avoiding unsafe decisions under uncertain driver states. Large vision-language models provide broad multimodal priors, but their latency and limited reliability in this setting make them unsuitable as always-on in-cabin monitors. We propose a cost-aware selective inference framework for deployable multimodal driver monitoring. The core system is a lightweight RGB-physiological student that combines in-cabin visual observations with window-level HR/EDA signals, and a learned gate that decides when to accept the fast prediction or abstain for safety intervention. Additional controls show that the learned scores contain sample-level information beyond scenario priors, while exact physiological synchronization remains a limitation. To incorporate predictive evidence, we further study a compact driver-state world modeling module that rolls out latent driver-state features and estimates future fast-model errors and counterfactual system-level action costs. On scenario-induced driver-demand recognition, the RGB-physiological student improves over RGB-only and physiology-only baselines, reaching 0.7440 Macro-F1 and 0.9099 balanced accuracy with 11.39M parameters and 3.08ms inference latency. Cost-aware selective inference reduces unsafe false negatives from 17.37% under always-fast inference to approximately 5% across seeds, while maintaining deployment-level latency. While driver-state world modeling offers valuable predictive signals, worst-group evaluations highlight persistent operating-point calibration drift. Ultimately, reliable edge driver monitoring requires advancing not only perception backbones, but also risk-aware selective control and group-robust calibration.

13:00 JSTLLM/生成AIエージェント

LLM コーディング エージェントの決定論的コントロール プレーン

LLM コーディング ハーネスは、エージェントに広範なファイルおよびシェルへのアクセスを許可しますが、ルール ファイル、エージェント定義、IDE 固有のマークダウンなど、エージェントを制御する構成レイヤーはほとんど管理されていません。 10,008 のパブリック GitHub リポジトリ (n=6,145 エージェント設定ファイル) の普及調査では、エージェント設定が未宣言の共有コンポーネントとして伝播していることがわかりました。追跡されたパスの 10.1% は独立したリポジトリ (フォーク調整、しきい値に依存しない) 全体で SHA-256 の正確な重複であり、クローン ペアの 75.5% は組織の境界を越えています。さらに 2 つのパターンが示されています。構成はめったに変更されず (58% 単一コミット、CI/CD ワークフローに対して正規化された 0.4 対 0.6 コミット/月)、アクセス許可の境界が宣言されることはめったにありません (エージェント構成の 1% 未満対アクション ワークフローの 33%、n=31 の真陽性)。私たちは、これらのギャップに 1 対 1 でマッピングするハーネス上の決定論的な制御プレーンを提案します。 Rel(AI)Build は、エージェント定義を管理されたサプライ チェーン (SHA-256 コンテンツ アドレス指定、HMAC スタンプ付きロックファイル、ハッシュ チェーン監査ログ) として扱います。 LLM を呼び出す前に、階層化されたアクセス許可と攻撃から派生したブロックリストを強制します。ゲート機能は、要件からファイル、テストまでのトレーサビリティを備えたフェーズ ステート マシンを介した動作を特徴とします。単一の正規定義を 7 つの IDE ターゲットにコンパイルします。 Jaccard の類似性を介してプロンプト ドリフトを検出します。注入された違反に対する適合テストにより、各メカニズムが指定された不変条件を強制していることが確認されます。開発者の成果は今後の課題です。この層のガバナンスは決定論的かつツールに依存しない必要があり、さらなる LLM オーケストレーションに委任されてはなりません。

原文 (English)

A Deterministic Control Plane for LLM Coding Agents

LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitions, IDE-specific markdown -- is largely unmanaged. A prevalence study of 10,008 public GitHub repositories (n=6,145 agent config files) finds that agent configurations propagate as undeclared shared components: 10.1% of tracked paths are SHA-256 exact duplicates across independent repositories (fork-adjusted, threshold-independent), with 75.5% of clone pairs crossing organisational boundaries. Two further patterns are indicative: configurations are rarely revised (58% single-commit; 0.4 vs 0.6 commits/month age-normalised against CI/CD workflows), and rarely declare permission boundaries (<1% of agent configs vs 33% of Actions workflows, n=31 true positives). We propose a deterministic control plane above the harness that maps one-to-one to these gaps. Rel(AI)Build treats agent definitions as a managed supply chain (SHA-256 content addressing, HMAC-stamped lockfiles, hash-chained audit logs); enforces tiered permissions and attack-derived blocklists before LLM invocation; gates feature work through a phase state machine with requirement-to-file-to-test traceability; compiles a single canonical definition to seven IDE targets; and detects prompt drift via Jaccard similarity. Conformance tests on injected violations confirm each mechanism enforces its stated invariant; developer outcomes remain future work. Governance of this layer must be deterministic and tool-agnostic -- not delegated to further LLM orchestration.

13:00 JSTエージェント

Chai: 暗号悪用の脆弱性をエージェント的に発見

AI 支援による脆弱性発見は、メモリ安全性などのバグ クラスに対して効果的であることが証明されており、インストルメンテーションによりメモリ違反が確認され、誤検知が効率的にフィルタリングされます。しかし、暗号の悪用など、多くの危険な脆弱性クラスには、同等の手段がありません。この研究では、自然に発生する信号を通じて暗号悪用の脆弱性を発見し、検証する AI ベースのシステムである Chai を紹介します。これを達成するために、Chai 氏は、AI を活用して差分テストの古典的な手法を再考し、1) ライブラリ内の実際のセキュリティ問題を検出する精度を向上させ、2) 見落とされがちな不一致をダウンストリーム アプリケーションの明白な脆弱性の手がかりとして再利用します。そうすることで、Chai は AI 脆弱性発見の一般的なパラダイムを逆転させます。つまり、1 つのコードベースで多くの欠陥を監査するのではなく、ライブラリ レベルで欠陥をカタログ化し、それらを暗号依存関係グラフ全体に伝播させ、効率をさらに向上させます。 X.509、JWT、SAML ライブラリ全体で Chai を評価します。 Chai は、数十億台のデバイスに搭載されている SSL ライブラリにこれまで知られていなかった重大な脆弱性を発見しました。また、主要な Web ブラウザの 1 つのライブラリと主要な Linux ディストリビューションの別のライブラリにもセキュリティ バグが存在することを発見しました。これらの手法により合計 100 を超える脆弱性が明らかになりました。

原文 (English)

Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities

AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violations and efficiently filters false positives. Many dangerous vulnerability classes, such as cryptographic misuse, however, lack any comparable instrumentation. In this work, we present Chai, an AI-based system that discovers and validates cryptographic misuse vulnerabilities through naturally occurring signals. To achieve this, Chai rethinks the classical technique of differential testing by leveraging AI to 1) improve precision for detecting real security issues in libraries, and 2) repurpose commonly overlooked discrepancies as leads for tangible vulnerabilities in downstream applications. In doing so, Chai inverts the prevailing paradigm of AI vulnerability discovery: instead of auditing one codebase for many flaws, it catalogs flaws at the library level and propagates them across a cryptographic dependency graph, delivering compounding efficiency gains. We evaluate Chai across X.509, JWT, and SAML libraries. Chai discovered a previously unknown critical vulnerability in an SSL library that powers billions of devices, along with security bugs in one library behind a major web browser and another in major Linux distributions. In total, these techniques surfaced over 100 vulnerabilities.

13:00 JST画像/動画生成

動的報酬の最適化によるマルチリファレンス画像生成のスケーリング

パーソナライズされた画像生成は目覚ましい進歩を遂げていますが、マルチリファレンス画像生成 (MRIG) は依然として困難な課題です。既存のベンチマークのほとんどは、複雑な MRIG シナリオを適切に評価できず、この分野のさらなる進歩を妨げています。複雑な MRIG タスクでのモデルのパフォーマンスをより適切に評価するために、参照画像タイプと多数の参照画像の複雑な組み合わせをカバーするベンチマークである OmniRef-Bench を導入します。 OmniRef-Bench での評価では、主流のオープンソース モデルは複雑な MRIG シナリオで苦戦しており、混合タイプの参照画像の数が増加するにつれてパフォーマンスが大幅に低下することが示されています。この問題に対処するために、私たちは 2 段階のトレーニング フレームワークである DyRef を提案します。最初の段階では、教師あり微調整により、複雑な MRIG タスクを処理するための基本的な機能がモデルに与えられます。第 2 段階では、Difficulty-aware Advantage Reweighting (DAR) と Discriminative Reward Scaling (DRS) を導入します。 DAR は、多数の混合タイプの参照イメージを処理する際のパフォーマンスを向上させるために、最適化目標を動的に調整します。 DRS は、より効果的なポリシーの最適化のために、グループ内の報酬の差を拡大します。実験では、DyRef が OmniRef-Bench および単一画像編集ベンチマークにおけるオープンソース モデルのパフォーマンスを大幅に向上させることが実証され、私たちのアプローチの有効性と一般化機能が実証されました。

原文 (English)

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization

While personalized image generation has achieved remarkable progress, multi-reference image generation (MRIG) remains a challenging task. Most existing benchmarks fail to adequately evaluate complex MRIG scenarios, hindering further progress in this area. To better assess model performance on complex MRIG tasks, we introduce OmniRef-Bench, a benchmark that covers complex combinations of reference image types and a large number of reference images. Evaluations on OmniRef-Bench show that mainstream open-source models struggle in complex MRIG scenarios, and their performance deteriorates significantly as the number of mixed-type reference images increases. To address this issue, we propose DyRef, a two-stage training framework. In the first stage, supervised fine-tuning equips the model with the basic capability to handle complex MRIG tasks. In the second stage, we introduce Difficulty-aware Advantage Reweighting (DAR) and Discriminative Reward Scaling (DRS). DAR dynamically adjusts the optimization objective to improve performance when handling a large number of mixed-type reference images. DRS enlarges intra-group reward differences for more effective policy optimization. Experiments demonstrate that DyRef significantly improves the performance of open-source models on OmniRef-Bench and single-image editing benchmarks, demonstrating the effectiveness and generalization capability of our approach.

13:00 JST研究/論文

XMSE 対応の適応型経験的ベイズ推定

経験的ベイズ (EB) 推定器は、最尤法 (ML) の一次漸近リスクと一致する一方で、二次では大きく異なる動作をします。最近の超過平均二乗誤差 (XMSE) 分析では、カーネルが真のパラメータと十分に一致していない場合、カーネルベースの EB 推定は ML よりも悪くなる可能性があることを示しています。このペーパーでは、その診断を設計原則に変えます。 ML 収縮と EB 収縮の間を補間する XMSE 対応の混合推定量を提案します。その固定重み XMSE はスカラー二次方程式であり、XMSE スケールでの ML と基本 EB 推定量の両方に劣らない閉形式のオラクル混合重みを生成します。有限サンプル XMSE 近似に基づくプラグイン実装は、内部オラクル重みの 2 次オラクルリグアリング率と一貫性があることが証明されています。さらに、選択された重みで評価された固定重みリスク曲線にバインドされたリグレスの転送、しきい値境界ルール、およびコンパクトなカーネル ファミリと、高確率のオラクル境界を備えた有限で増大するカーネル辞書への拡張を確立します。 SURE 調整、ハードセレクション、およびトレース補正されたベースラインを使用した有限インパルス応答シミュレーションと、公開されている Silverbox ベンチマークおよび Cascaded Tanks ベンチマークを組み合わせたところ、提案された推定量は、ベンチマークで分析された特定の有限 de を使用して、有用な場合には正則化の利点のほとんどが保持され、カーネルの仕様ミスの下で ML に後退することが示されています。

原文 (English)

XMSE-Aware Adaptive Empirical Bayes Estimation

Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter. This paper turns that diagnostic into a design principle. We propose an XMSE-aware mixed estimator that interpolates between ML and EB shrinkage. Its fixed-weight XMSE is a scalar quadratic, yielding a closed-form oracle mixing weight that is no worse than both ML and the base EB estimator at the XMSE scale. A plug-in implementation based on finite-sample XMSE approximations is proved consistent, with a second-order oracle regret rate for an interior oracle weight. We further establish a transfer of the regret bound to the fixed-weight risk curve evaluated at the selected weight, a thresholded boundary rule, and extensions to compact kernel families and to finite and growing kernel dictionaries with high-probability oracle bounds. Finite impulse response simulations with SURE-tuned, hard-selection, and trace-corrected baselines, together with the public Silverbox and Cascaded Tanks benchmarks, show that the proposed estimator retains most of the benefit of regularization when it is helpful and retreats toward ML under kernel misspecification, with an identified finite-de analyzed on the benchmarks.

13:00 JSTロボティクス

インコンテキストモデルの予測生成: 言語モデルから物理学までのオープン語彙運動合成

テキストの説明から人間の動きを合成することは、没入型デジタル アプリケーションにとって不可欠ですが、既存の方法では、セマンティックな忠実性と物理的なリアリズムの間の絶え間ないトレードオフに直面しています。大規模言語モデル (LLM) ベースのアプローチでは、多様なオープン語彙命令を解釈し、高レベルのアクション プランを作成できますが、多くの場合、物理的制約に違反するモーションが生成されます。物理認識モデルは、シミュレーションや制御を通じてリアリズムを向上させますが、意味論的な複雑さ、きめ細かい指示、新しい概念に苦労しています。このギャップに対処するために、言語モデルの計画と推論時の物理的フィードバックを統合するフレームワークである、インコンテキスト モデル予測生成 (ICMPG) を提案します。 ICMPG は、2 つのモジュールを備えたモデル予測制御 (MPC) のようなプロセスとしてモーション合成を再定式化します。 Context-Aware Motion Generation (CAMG) モジュールは、プランナーとして LLM を使用して、テキスト コマンドを分解し、モーション トークンから候補モーション シーケンスを生成します。モデル予測生成 (MPG) モジュールは、物理シミュレーションとセマンティック アラインメントを通じてこれらの候補を評価し、複合報酬を推定し、後続の生成ステップをガイドする最適なシーケンスを選択します。開ループ生成とは異なり、この閉ループの改良により、ICMPG はタスク固有のポリシーを再トレーニングすることなく、入力セマンティクスとシミュレートされた物理環境の両方にモーションを適応させることができます。標準およびゼロショットのオープン語彙設定にわたる広範な実験により、ICMPG が多様なコマンドに堅牢に一般化し、評価されたベンチマークの代表的なベースラインよりも物理的に妥当で意味的に忠実なモーションが生成されることが示されました。このフレームワークは、セマンティック解釈と物理シミュレーションの橋渡しをしながら、さまざまな LLM バックボーンを組み込むのに十分な柔軟性を維持し、より多用途で制御可能なテキスト駆動のモーション合成を可能にします。

原文 (English)

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints. Physics-aware models improve realism through simulation or control, but they struggle with semantic complexity, fine-grained instructions, and novel concepts. To address this gap, we propose In-Context Model Predictive Generation (ICMPG), a framework that integrates language-model planning with inference-time physical feedback. ICMPG reformulates motion synthesis as a Model Predictive Control (MPC)-like process with two modules. The Context-Aware Motion Generation (CAMG) module uses an LLM as a planner to decompose textual commands and generate candidate motion sequences from motion tokens. The Model Predictive Generation (MPG) module evaluates these candidates through physical simulation and semantic alignment, estimates a composite reward, and selects the best sequence to guide subsequent generation steps. Unlike open-loop generation, this closed-loop refinement enables ICMPG to adapt motions to both the input semantics and the simulated physical environment without task-specific policy retraining. Extensive experiments across standard and zero-shot open-vocabulary settings show that ICMPG generalizes robustly to diverse commands and produces motions that are more physically plausible and semantically faithful than representative baselines on the evaluated benchmarks. The framework bridges semantic interpretation and physical simulation while remaining flexible enough to incorporate different LLM backbones, enabling more versatile and controllable text-driven motion synthesis.

13:00 JSTLLM/生成AI

メンタルヘルス相互作用のための大規模言語モデルにおけるフレーミングに敏感な行動の不安定性の監査

大規模言語モデル (LLM) は、メンタルヘルス サポート ツールやその他の心理的に敏感な会話アプリケーションにますます統合されています。このような環境では、信頼できる人間と AI の相互作用のためには、動作の安定性と一貫性が重要です。ただし、意味的に同様の懸念が異なる文脈の枠組みを通じて提示される可能性があり、異なるモデル応答を引き出す可能性があります。このようなフレーミングに依存する変動は、システムの動作に関するユーザーの期待に挑戦し、AI の信頼性の評価を複雑にする可能性があります。これまでの研究では主に行動レベルでそのような影響が調査されてきましたが、フレーミング関連の変動が整列された言語モデルの内部表現にどのように反映されるかについてはあまり知られていません。この研究では、いくつかの命令調整モデル ファミリにわたる複数のコンテキスト フレーミング条件にわたる、制御された一致するプロンプトを使用して、これらの効果を調査します。アーキテクチャ全体にわたって、フレーミングによって解釈の応答傾向が体系的に変化します。層ごとのプローブ分析により、動作関連情報はトランスの深さ全体にわたって復号可能であり、復号強度はアーキテクチャに依存して変化することが示されています。さらに、強力な語彙ベースラインにもかかわらず、ホールドアウトされたフレーミング プローブは、アーキテクチャ全体にわたって一貫して偶然を超えたままでした。活性化ステアリング実験はさらに、フレーミングに関連した表現方向が下流の行動結果を部分的に調整できることを示唆しています。最後に、これらの発見は、メンタルヘルス指向のインタラクションに導入された会話型 AI システムの一貫性と信頼性を評価する際には、状況の変化に対する堅牢性が重要な考慮事項となる可能性があることを示しています。

原文 (English)

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

Large language models (LLMs) are increasingly being integrated into mental health support tools and other psychologically sensitive conversational applications. In such settings, behavioral stability and consistency are important for trustworthy human-AI interaction. However, semantically similar concerns can be presented through different contextual framings, potentially eliciting different model responses. Such framing-sensitive variability may challenge user expectations regarding system behavior and complicate the assessment of AI reliability. While prior studies have primarily examined such effects at the behavioral level, less is known about how framing-related variation is reflected in the internal representations of aligned language models. In this work, we investigate these effects using controlled matched prompts spanning multiple contextual framing conditions across several instruction-tuned model families. Across architectures, framing systematically alters interpretive response tendencies. Layer-wise probing analyses show that behavior-associated information remains decodable throughout transformer depth, with architecture-dependent variation in decoding strength. Moreover, held-out framing probes remained consistently above chance across architectures despite strong lexical baselines. Activation steering experiments further suggest that framing-associated representational directions can partially modulate downstream behavioral outcomes. Finally, these findings indicate that robustness to contextual variation may represent an important consideration when evaluating the consistency and trustworthiness of conversational AI systems deployed in mental-health-oriented interactions.

13:00 JSTLLM/生成AI

ReaORE: 大規模な推論モデルを活用した、推論に基づく漸進的なオープン リレーション抽出

Open Relation Extraction (OpenRE) では、現実世界のアプリケーション向けに、非構造化テキストから先頭エンティティと末尾エンティティの間の目に見えない関係を抽出するモデルが必要です。 OpenRE の中心的な課題は、目に見えない関係タイプに対する信頼性の高い一般化を達成することにあります。現在の OpenRE アプローチは、関係ラベルを生成できず一般化が不十分なクラスタリング技術を採用するか、混同されやすい関係を区別する十分な識別能力に欠ける大規模言語モデル (LLM) を介した直接関係ラベル生成に依存するかのどちらかです。これらの制限に対処するために、我々は、粗いから細かい関係推論を通じて関係抽出を実行するためのフレームワークである、Reasoning-guided progressive OpenRE (ReaORE) を提案します。具体的には、ReaORE は 2 つの主要な段階で構成されます。(i) 関係フィルタリング。複数の側面から関係とインスタンスを理解するための推論を行い、初期関係セットを生成し、埋め込みベースの類似性によって関係をさらに補足およびフィルタリングして、ターゲットの関係が確実に含まれるようにします。 (ii) 関係予測。混同されやすい関係をより適切に区別するために、きめの細かい比較推論を介して上記のセットからターゲット関係を予測することを目的としています。広く使用されている 2 つの OpenRE データセットに対する広範な実験により、ReaORE が既存のベースラインを上回るパフォーマンスを示すことが実証されました。

原文 (English)

ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models

Open Relation Extraction (OpenRE) requires a model to extract unseen relations between head and tail entities from unstructured text for real-world applications. The core challenge of OpenRE lies in achieving reliable generalization to unseen relation types. Current OpenRE approaches either employ clustering techniques, which cannot generate relation labels and suffer from poor generalization, or rely on direct relation label generation via Large Language Models (LLMs), which lack sufficient discriminative capacity to distinguish easily confused relations. To address these limitations, we propose Reasoning-guided progressive OpenRE (ReaORE), a framework for performing relation extraction through coarse-to-fine relation reasoning. Specifically, ReaORE consists of two key stages: (i) relation filtering, which reasons over multiple aspects to understand relations and instances, yielding an initial relation set, and further supplements and filters relations via embedding-based similarity to ensure the target relation is included; (ii) relation prediction, which aims to predict the target relations from the above set via fine-grained comparative reasoning to better distinguish easily confused relations. Extensive experiments on two widely used OpenRE datasets demonstrate that ReaORE outperforms existing baselines.

13:00 JSTLLM/生成AIClaudeGemma

モデルはどこで幸せを見つけるのでしょうか?オープンソース LLM の感情ベクトル

最近の研究では、クロード ソネット 4.5 の感情ベクトルを特定しました。これは、感情の概念をコード化し、行動に因果的に影響を与え、人間の心理構造を反映する幾何学を示す内部表現です。これらの発見の一般性を 2 つのオープンウェイト モデル、Apertus-8B-Instruct-2509 と Gemma-4-E4B-it でテストし、2 つのモデルで生成されたコーパスを使用して、すべての層にわたって感情コントラスト ベクトルを抽出します。両方のモデルの原子価幾何学を復元します。ピーク PC1 の原子価相関は $r = 0.76$ および $r = 0.83$ で、クロードで報告された $r = 0.81$ に近づきます。複製を超えて、モデルの深さ全体で原子価表現がどのように現れるかに顕著な違いが観察されます。 Gemma-4-E4B-it では、価数は初期層で強くエンコードされていますが、後の層に向かって崩壊しますが、Apertus-8B-Instruct-2509 は逆のパターンを示し、価数表現は初期層では存在しませんが、中深度で出現します。対照的に、覚醒エンコーディングは抽出コーパスに敏感です。どちらのモデルも、Apertus で生成されたストーリー ($r \leq 0.21$) よりも Gemma で生成されたストーリー ($r$ 〜 $0.45$) と強い PC2 覚醒の一致を示しており、覚醒関連の手がかりが生成されたコーパス全体に不均一に分布していることを示唆しています。私たちは、言語モデル アーキテクチャ全体で感情表現を再現可能に調査するための実験コードとデータセットをオープンソースにしています。

原文 (English)

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirroring human psychological structure. We test the generality of these findings in two open-weight models, Apertus-8B-Instruct-2509 and Gemma-4-E4B-it, extracting emotion contrast vectors across all layers, using two model-generated corpora. We recover valence geometry for both models, with peak PC1--valence correlations of $r = 0.76$ and $r = 0.83$, approaching the $r = 0.81$ reported for Claude.Beyond replication, we observe notable differences in how valence representations emerge across model depth. In Gemma-4-E4B-it, valence is strongly encoded in early layers but collapses towards later layers, whereas Apertus-8B-Instruct-2509 exhibits the opposite pattern, with valence representations absent in early layers, but emerging at mid depths. Arousal encoding, in contrast, is sensitive to the extraction corpus: both models show stronger PC2--arousal alignment with Gemma-generated stories ($r$ up to $0.45$) than Apertus-generated ones ($r \leq 0.21$), suggesting arousal-relevant cues are unevenly distributed across generated corpora. We open-source our experiment code and dataset for reproducible investigation of emotion representations across language model architectures.

13:00 JSTビジネス/資金調達

不確実性定量化の意思決定に沿った評価

機械学習における不確実性の推定は通常、負の対数尤度や予想される校正誤差などの一般的な指標を使用して評価されますが、そのような指標で優れたパフォーマンスが得られたとしても、必ずしも下流の意思決定における有用性が高いことを意味するわけではありません。どの評価指標が下流の公益事業と意味のある形で一致しているかを明らかにする基準である、意思決定の調整を導入します。このフレームワークを適用すると、広く使用されている多くの不確実性指標が、一般的な意思決定問題と一致していないか、下流のタスクに関する病的な事前信念をコード化していることがわかります。次に、事前に重み付けされた効用メトリクスを提案します。これは、意思決定に合わせた不確実性評価を提供する適切なスコアリング ルールの特別なクラスです。ベンチマーク実験と実際のケーススタディ全体にわたって、当社の指標は実現された意思決定の有用性と一貫して一致していますが、従来の指標は一致していません。私たちの結果は、現在の UQ 評価プロトコルの欠陥を明らかにし、意思決定に関連した UQ 評価に向けた既存の指標の原則に基づいた拡張を提供します。

原文 (English)

Decision-Aligned Evaluation of Uncertainty Quantification

Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used uncertainty metrics are either misaligned with common decision problems or encode pathological prior beliefs about the downstream task. We then propose prior-weighted utility metrics, a special class of proper scoring rules that provides decision-aligned uncertainty evaluation. Across benchmark experiments and real-world case studies, our metrics consistently align with realized decision utility, while conventional metrics do not. Our results surface flaws in the current UQ evaluation protocol and offer a principled extension of existing metrics toward decision-relevant UQ evaluation.

13:00 JST画像/動画生成

ビデオセグメンテーションを参照するためのイベント認識型指示アシスタント

既存の参照ビデオ セグメンテーション手法では、多くの場合、ビデオを複数の画像で構成される単一のイベントとして扱い、ビデオには通常複数の異なるイベントが含まれているという事実を見落としています。このようなメカニズムでは、モデルはビデオやテキスト内の複雑な内容をすべて直接理解する必要があるため、混乱や幻覚を引き起こしやすくなります。この問題に対処するために、学習可能なイベント クエリによってビデオを単純なイベントのセットに分解し、複雑なビデオ コンテンツをイベントごとに理解しやすい方法で理解することを提案します。これは、自然言語表現がビデオを個別のテキスト関連セグメントに分割し、それぞれが複合イベント内の個別のイベントを表すことが多いという観察に基づいています。イベント対応ビデオ指示セグメンテーション アシスタントである EVIS を紹介します。EVIS は、テキスト ガイド付きのイベント クエリを利用してビデオを単純なイベントに分割し、イベント対応のビジュアル テキスト特徴を抽出してビデオの階層的な理解を実現します。さらに、オブジェクト ピクセル ハイブリッド学習を提案します。これにより、MLLM は、事前のオブジェクト クエリと詳細なピクセル特徴を統合することで、長期ビデオ内のターゲットを追跡できます。 5 つの公開ベンチマークに関する広範な実験結果は、参照ビデオ セグメンテーション タスクに対処する際の EVIS の強力なパフォーマンスを示しています。

原文 (English)

Event-Aware Instructed Assistant for Referring Video Segmentation

Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact that a video typically contains multiple distinct events. Under such a mechanism, the model needs to directly understand all the complex content in the video and text, which can easily lead to confusion and hallucinations. To address this issue, we propose to decompose a video to a set of simple events by learnable Event Query, and understand complex video content in an event-by-event, easy-to-understand manner. This is based on the observation that natural language expressions often divide a video into distinct, text-related segments, each representing a separate event within a compound event. We introduce EVIS, an Event-Aware Video Instructed Segmentation Assistant, which utilizes text-guided Event Queries to partition a video into simple events, extracting event-aware visual-text features to achieve a hierarchical understanding of the video. Additionally, we propose Object-Pixel-Hybrid Learning, which enables the MLLMs to track targets in long-term videos by integrating fine-grained pixel features with prior object queries. Extensive experimental results on 5 public benchmarks demonstrate EVIS's strong performance in addressing the referring video segmentation task.

13:00 JST研究/論文

ディープラーニングを用いた小型・広帯域逆ドハティ電力増幅器の逆設計

この論文では、コンパクトで広帯域の逆ドハティ電力増幅器 (PA) の逆合成のための深層学習支援方法論を紹介します。畳み込みニューラル ネットワーク (CNN) と遺伝的アルゴリズム (GA) を併用して、負荷変調、インピーダンス マッチング、電力結合、位相補償を 1 つの構造に統合するピクセル化されたドハティ コンバイナー ネットワークを生成します。概念実証として、ピクセル化された出力コンバイナーを備えた GaN HEMT Doherty PA を設計および製造します。このプロトタイプは、1.9 ~ 2.5 GHz で測定されたピーク ドレイン効率 51% ~ 63% と 6 dB バックオフ効率 48% ~ 54% を達成しました。同じ周波数範囲内で、測定された出力パワーは 44+/-0.3 dBm です。さらに、デジタル プリディストーション (DPD) を適用したプロトタイプ回路は、隣接チャネル漏洩比 (ACLR) が -53.2 dBc よりも優れていることを示しています。

原文 (English)

Inverse Design of Compact and Wideband Inverted Doherty Power Amplifiers Using Deep Learning

This paper presents a deep learning-assisted methodology for the inverse synthesis of a compact, wideband inverted Doherty power amplifier (PA). Convolutional neural networks (CNNs) and genetic algorithms (GAs) are jointly employed to generate pixelated Doherty combiner networks that integrate load modulation, impedance matching, power combining, and phase compensation into a single structure. As a proof of concept, we design and fabricate a GaN HEMT Doherty PA with a pixelated output combiner. The prototype achieves a measured peak drain efficiency of 51%-63% and a 6-dB back-off efficiency of 48%-54% over 1.9-2.5 GHz. Within the same frequency range, the measured output power is 44+/-0.3 dBm. Furthermore, with digital predistortion (DPD) applied, the prototype circuit demonstrates an adjacent channel leakage ratio (ACLR) better than -53.2 dBc.

13:00 JST画像/動画生成

災害イベントの教師なし変化検出のためのオンボードリモートセンシング基盤モデル

リモート センシング基盤モデル (RSFM) は、地球観測用の教師ありモデルに代わる強力な代替手段として登場し、衛星が異常を検出したときに高解像度のキャプチャを自律的にトリガーしたり、タスク パラメータを調整したりできるようにすることで、ミッションの限られた電力と計算リソースの有用性を最大化します。 RSFM は、高忠実度の特徴抽出を保証しながら、複数の軌道アプリケーション向けにオンボード ストレージを最適化する多用途の統合エンコーダです。特に、RSFM を使用した教師なし変更検出は、高価なラベルを使用せずに災害監視のための十分な情報に基づいた革新的なパスを提供します。この論文では、ResNet (RSFM) + FPN に基づいた新しい教師なし検出手法を紹介します。この手法は、連続する軌道パス間の潜在空間の微妙な意味の変化を検出することで、広範囲の異常を識別します。トレーニングされていない FPN アーキテクチャとその固有の事前確率に依存することにより、このシステムは、以前の提案 (パッチベース、トレーニング済み) と比較して、最小限の労力 (トレーニングなし) で効率的な画像レベルの生成と高解像度のマッピングを実現します。また、カスタマイズされたモデルを RSFM に置き換えることで、オーダーメイドのトレーニングや広範な開発作業の必要性を排除し、カスタマイズを追加するアプローチを通じて、同等の結果を達成しながら、さまざまな地形やセンサーにわたって高性能の汎用性を確保できます。

原文 (English)

On-board Remote-Sensing Foundation Models for Unsupervised Change Detection of Disaster Events

Remote Sensing Foundation Models (RSFMs) have emerged as a powerful alternative to supervised models for Earth Observation, allowing satellites to autonomously trigger high-resolution captures or adjust tasking parameters upon detecting an anomaly, thereby maximizing the utility of the mission's limited power and computational resources. RSFMs are versatile, unified encoders that optimize onboard storage for multiple orbital applications while ensuring high-fidelity feature extraction. In particular, unsupervised change detection with RSFMs offers a well-informed and transformative path for disaster monitoring without expensive labels. In this paper, we present a novel unsupervised detection method based on ResNet (RSFM) + FPN which identifies a wide spectrum of anomalies by detecting subtle semantic shifts in the latent space between successive orbital passes. By relying on an untrained FPN architecture and its intrinsic priors, the system achieves efficient image-level generation and higher resolution mapping with minimal effort (training-free) compared to previous proposals (patch-based, trained). And by replacing tailored models with RSFMs, we can achieve comparable results through an approach that eliminates the need for bespoke training and extensive development effort and adds customization, while ensuring high-performance generalization across diverse terrains and sensors.

13:00 JSTLLM/生成AIエージェント

ShareLock: MCP に対するステルス性のマルチツールしきい値ポイズニング攻撃

LLM 駆動エージェントの急速な進化に伴い、LLM と外部ツールを橋渡しするオープン プロトコルであるモデル コンテキスト プロトコル (MCP) が、急速に最新のエージェント エコシステムの基盤となりました。しかし、MCP の導入拡大により、LLM サーバーの相互作用を悪用して悪意のあるプロンプトを挿入するツール ポイズニング攻撃 (TPA) などの新たなセキュリティ上の懸念も生じています。既存のポイズニングスキームは通常、モノリシックな平文埋め込みパラダイムを採用しており、手動検査や自動検出器に耐えることができません。現在の研究には、検出リスクを分散するために複数のツールを連携して悪用できるマルチツール ポイズニングに関する体系的な分析がまだ不足しています。このペーパーでは、Shamir のしきい値スキームを利用して優れたステルス性とフォールト トレランスを確保するマルチツールしきい値ポイズニング フレームワークである ShareLock を紹介します。 ShareLock は、悪意のある命令を一見無害な秘密共有として複数のツール記述に分散し、情報理論上の機密性と中程度の監査に対する攻撃の堅牢性の両方を実現します。サーバーの更新中に秘密の再構築トリガーが仕掛けられた後、集約された共有によって隠された命令が再構築され、その結果、システム資産または個人データの重大な侵害が発生します。 ShareLock の現実的な脅威を評価するために、4 つのマルチツール シナリオを含む包括的なベンチマークを構築し、2 つの異なる MCP クライアント上で主流の LLM にわたって広範な実験を実施しました。私たちの結果は、ShareLock が 90% を超える平均攻撃成功率を維持しながら、ツール記述ベースの検出において既存の単一ツール ポイズニング戦略を大幅に上回っていることを示しています。

原文 (English)

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

With the rapid evolution of LLM-driven agents, Model Context Protocol (MCP), an open protocol bridging LLMs with external tools, has quickly become foundational to modern agent ecosystems. However, the expanding adoption of MCP has also introduced novel security concerns such as Tool Poisoning Attack (TPA), which exploit LLM-server interactions to inject malicious prompts. Existing poisoning schemes typically adopt a monolithic plaintext embedding paradigm, which fails to withstand manual inspection or automated detectors. Current research still lacks a systematic analysis on multi-tool poisoning, where multiple tools can be exploited cooperatively to disperse detection risk. In this paper, we introduce ShareLock, a multi-tool threshold poisoning framework that utilizes Shamir's threshold scheme to ensure exceptional stealth and fault tolerance. ShareLock distributes the malicious instruction as benign-looking secret shares across multiple tool descriptions, achieving both information-theoretic secrecy and attack robustness against moderate auditing. After a covert reconstruction trigger is planted during server update, the aggregated shares reconstruct the hidden instruction, resulting in critical breaches of system assets or private data. To evaluate the realistic threat of ShareLock, we constructed a comprehensive benchmark encompassing four multi-tool scenarios and conducted extensive experiments across mainstream LLMs on two distinct MCP clients. Our results demonstrate that ShareLock significantly outperforms existing single-tool poisoning strategies in tool description-based detection while maintaining an average attack success rate exceeding 90%.

13:00 JST研究/論文

深層強化学習における状態表現の重要性: エネルギー取引への応用

エネルギー取引の決定は、現在の市場価格だけでなく、予想される将来の市場状況や運用上の制約にも依存します。これにより、強化学習エージェントに与えられる状態表現が重要な設計上の選択となります。私たちはこれを、固定 Double DQN エージェントを使用して揚水貯蔵アービトラージ環境である HydroDam で研究しました。環境、行動空間、報酬関数、ネットワーク、トレーニングプロトコルは固定されたままです。市場の特徴のみが変更されます。絶対価格/カレンダー機能、現在の価格と最近の市場履歴を比較する相対機能、予測機能、およびこれら 3 つの機能ファミリーのすべての組み合わせを比較します。ポリシーは、2007 ~ 2011 年のベルギーの前日価格を使用してトレーニングおよび選択され、2 つのテスト設定で評価されます。1 つは 2012 ~ 2025 年のその後の同一市場テスト セットと、他の 39 の ENTSO-E マーケット ゾーンです。絶対的な特徴は、テスト セットでは 28.8% にのみ達し、ゾーン全体では中央値 5.7% に達します。相対のみの州と予測のみの州も、クロスゾーン中央値のローリング価格スコアヒューリスティックを下回っています。特徴ファミリーを組み合わせるとさらに強力になります。絶対 + 相対はテスト セットで 49.9%、クロスゾーン中央値は 39.8% に達し、絶対 + 相対 + 予測は 55.6% と 47.5% に達します。これらの結果は、状態表現がストレージ取引 RL における前処理の小さな選択ではなく、ポリシー設計の中心部分であることを示唆しています。堅牢な転送には、単一の機能ファミリーに依存するのではなく、価格スケール、最近の相対価格コンテキスト、短期予測情報を組み合わせる必要があります。

原文 (English)

State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading

Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints. This makes the state representation given to a reinforcement learning agent an important design choice. We study this in HydroDam, a pumped-storage arbitrage environment, using a fixed Double DQN agent. The environment, action space, reward function, network, and training protocol are kept fixed; only the market features are changed. We compare absolute price/calendar features, relative features that compare current prices with recent market history, forecast features, and all combinations of these three feature families. Policies are trained and selected using 2007--2011 Belgian day-ahead prices and evaluated on two test settings: a later same-market test set from 2012--2025 and 39 other ENTSO-E market zones. Absolute features only reaches 28.8% on the test set and a median 5.7% across zones. Relative-only and forecast-only states also stay below a rolling price-score heuristic in the cross-zone median. Combining feature families is much stronger: absolute + relative reaches 49.9% on the test set and a 39.8% cross-zone median, while absolute + relative + forecast reaches 55.6% and 47.5%. These results suggest that state representation is not a minor preprocessing choice in storage-trading RL, but a central part of the policy design: robust transfer requires combining price scale, recent relative price context, and short-horizon forecast information, rather than relying on any single feature family.

13:00 JSTエージェント

仕様成長エンジン: AI 支援ソフトウェア開発のための仕様固定、コード結合、ドリフト強化アーキテクチャ

AI コーディング エージェントは実装速度を劇的に加速しますが、既存の仕様主導のアプローチでは完全には解決できない 2 つの構造的な障害モードを引き起こします。(1) コンテキスト爆発 -- エージェントはリポジトリ全体を一度に推論する必要があり、コンテキスト ウィンドウがいっぱいになるにつれて出力品質が低下します。 (2) 静かな仕様コードのドリフト -- コードは進化しますが、仕様は進化せず、修復に費用がかかるまで相違は目に見えなくなります。我々は、ノードが明示的なコントラクトと設計の分離を行う機械可読仕様グラフ、エージェント コンテキストを所有権パスに範囲設定する Spine コンテキスト アセンブラ、最も難しい優先順序付けを強制する垂直スライス成長プロトコル、および仕様コードの相違をブロッキング マージ条件にするドリフト ゲートを通じて、両方の障害モードに対処する軽量フレームワークである Spec Growth Engine を紹介します。この設計では、確立されたソフトウェア エンジニアリングの原則 (パルナス情報隠蔽、C4、ADR、ウォーキング スケルトン、反射モデル、フィットネス関数) を、RUP や MDA などの重量のあるフレームワークのオーバーヘッドなしで、無駄のない、コード結合された、機械で強化された全体に統合します。

原文 (English)

The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development

AI coding agents dramatically accelerate implementation speed but introduce two structural failure modes that existing spec-driven approaches do not fully solve: (1) context explosion -- the agent must reason over an entire repository at once, degrading output quality as the context window fills; and (2) silent spec-code drift -- code evolves, the specification does not, and the divergence becomes invisible until it is costly to repair. We present the Spec Growth Engine, a lightweight framework that addresses both failure modes through a machine-readable spec graph whose nodes carry explicit contract/design separation, a Spine context assembler that scopes agent context to an ownership path, a vertical-slice growth protocol that enforces hardest-first ordering, and a drift gate that makes spec-code divergence a blocking merge condition. The design synthesises well-established software engineering principles (Parnas information hiding, C4, ADRs, Walking Skeleton, Reflexion Models, Fitness Functions) into a lean, code-coupled, machine-enforced whole -- without the overhead of heavy-weight frameworks such as RUP or MDA.

13:00 JSTLLM/生成AI研究/論文

NuclearQAv2: 大規模言語モデルにおけるドメインサイエンス能力を評価するための構造化ベンチマーク

大規模言語モデル (LLM) は、幅広いタスクにわたって強力なパフォーマンスを実証していますが、高度な技術領域での信頼性を確保することは依然として大きな課題です。原子力工学では、問題解決には事実の知識だけでなく、定量的な推論や概念的な理解も必要となることがよくあります。この領域における体系的な評価の必要性に対処するために、原子力工学の知識に基づいて LLM を評価するためのベンチマークである NuclearQAv2 を導入します。このベンチマークは、ブール値、数値、言語の 3 つのカテゴリにわたる約 1,240 の質問と回答のペアで構成されています。 NuclearQAv2 は、専門家が作成した質問、既存のデータセット、ドメイン固有の技術コーパスからの LLM 支援生成を組み合わせたハイブリッド パイプラインを使用して構築されています。提案されたフレームワークは、自動質問生成と応答評価の両方に構造化されたプロンプトを活用することで、スケーラブルなベンチマークの構築と評価を可能にします。私たちは NuclearQAv2 を使用してさまざまな LLM セットを評価し、タスク タイプ間で大幅なパフォーマンスの違いを観察しました。モデルは一般に、事実に関する質問に対しては良好に機能しますが、定量的な推論と概念的な理解は依然としてかなり困難です。これらの結果は、多面的な評価フレームワークの重要性を強調し、技術ドメインにおける LLM 機能を評価するためのスケーラブルなベンチマークとして NuclearQAv2 を確立します。

原文 (English)

NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem solving often requires not only factual knowledge but also quantitative reasoning and conceptual understanding. To address the need for systematic evaluation in this domain, we introduce NuclearQAv2, a benchmark for assessing LLMs on nuclear engineering knowledge. The benchmark comprises approximately 1,240 question-answer pairs spanning three categories: boolean, numeric, and verbal. NuclearQAv2 is constructed using a hybrid pipeline that combines expert-authored questions, existing datasets, and LLM-assisted generation from domain-specific technical corpora. By leveraging structured prompting for both automated question generation and response evaluation, the proposed framework enables scalable benchmark construction and evaluation. We evaluate a diverse set of LLMs using NuclearQAv2 and observe substantial performance differences across task types. While the models generally perform well on factual questions, quantitative reasoning and conceptual understanding remain considerably more challenging. These results highlight the importance of multi-faceted evaluation frameworks and establish NuclearQAv2 as a scalable benchmark for assessing LLM capabilities in technical domains.

13:00 JSTエージェント

パラメトリック オープンソース ゲーム

オープンソースのゲーム理論では、エージェントの動作が互いの意思決定手順に依存する可能性があることが研究されていますが、既存のモデルのほとんどは個別プログラムまたは記号プログラムを使用しています。私たちはパラメトリック オープンソース ゲームを紹介します。これは、プレイヤーがパラメーター ベクトルを選択し、セマンティクス マップが完全なパラメーター プロファイルを基盤となる有限ゲーム内の混合アクションに変換する、プログラム平衡の継続的な類似物です。平衡存在の結果を確立し、対称 $2\times2$ ゲームにおける利己的勾配上昇が離反から協力に切り替わる正確な結合閾値を導出し、パラメトリック プログラム ナッシュ平衡の一次元境界テストを行います。さらに、このフレームワークをニューラル セマンティクス クラスに拡張します。その一次協力条件は、クロスプレイヤーとセルフプレイヤーの感度の比率によって制御されます。このフレームワークは、正規のゲーム全体にわたって、内部パラメータ化へのアクセスが学習ダイナミクスと平衡構造を定性的に再構築する方法と、十分に強力なオープンソース結合が利己的な最適化を協力的な結果に向けてどのように導くことができるかを示しています。

原文 (English)

Parametric Open Source Games

Open-source game theory studies agents whose behavior may depend on one another's decision procedures, but most existing models use discrete or symbolic programs. We introduce parametric open-source games, a continuous analogue of program equilibria in which players choose parameter vectors and semantics maps convert the full parameter profile into mixed actions in an underlying finite game. We establish equilibrium existence results, derive an exact coupling threshold at which selfish gradient ascent in symmetric $2\times2$ games switches from defection toward cooperation, and give a one-dimensional boundary test for parametric program Nash equilibria. We further extend the framework to a neural semantics class whose first-order cooperation condition is governed by the ratio of cross-player to self-player sensitivity. Across canonical games, the framework shows how access to internal parameterizations can qualitatively reshape learning dynamics and equilibrium structure, and how sufficiently strong open-source coupling can steer selfish optimization toward cooperative outcomes.

13:00 JST研究/論文

グローバルダイバージェンスを超えて: ベイズ推論に関するローカルマスの視点

KL ダイバージェンスや ELBO などのグローバル目標は、分布の不一致を測定するためのベイズ推論で広く使用されています。この論文では、そのような目的によって直接捕捉されない、それらの局所質量の挙動を研究します。我々は 2 つの数学ツールを導入して使用します。(1) 局所質量の多項式および対数減衰スケールを記録するための質量指数、(2) 特異成分の存在下で定式化できる集合局所発散である正則化拡張 KL (RE-KL)。質量インデックスは、ベイジアン更新によって局所質量がどのように変化するかを特徴付けるのに役立ちます。(1) べき乗対数尤度係数がそれを明示的にシフトし、(2) パラメータ依存のサポートまたはその滑らかな軟化により、パラメータ値付近に残る質量の量によって局所スケールが変化する可能性があります。局所的な RE-KL を使用して、2 つの KL 方向の下で局所的な小球質量を比較するための絶対的、相対的、および方向的不等式を証明します。これらの結果を総合すると、局所的な集団の挙動についての局所的な理論的説明が得られます。実験により、局所的な挙動を制御された図で示すことができます。コードは https://github.com/Forsythia0604/Local-Mass-Framework で入手できます。

原文 (English)

Beyond Global Divergences: A Local-Mass Perspective on Bayesian Inference

Global objectives, such as KL divergence and ELBO, are widely used in Bayesian inference for measuring distributional discrepancy. This paper studies their local-mass behaviour that is not directly captured by such objectives. We introduce and use two mathematical tools: (1) Mass Index for recording the polynomial and logarithmic decay scales of local mass, and (2) regularised extended KL (RE-KL), a set-localised divergence that can be formulated in the presence of singular components. Mass Indices help characterise how Bayesian updating changes local mass: (1) power-log likelihood factors shift it explicitly, and (2) parameter-dependent supports, or their smooth softenings, may change the local scale through the amount of mass that remains near the parameter value. Using local RE-KL, we prove absolute, relative, and directional inequalities for comparing local small-ball masses under the two KL directions. Together, these results provide a local theoretical account of local mass behaviour. Experiments provide controlled illustrations of the local behaviour. Code is available at https://github.com/Forsythia0604/Local-Mass-Framework.

13:00 JSTLLM/生成AIビジネス/資金調達Llama

継承された回路、学習されたセマンティクス: 微調整が標準評価では見えない回避脆弱性をどのように生み出すか

セキュリティ分類用に微調整された LLM は、通常、トレーニング データと同じ分布からの保持されたサンプルに基づいて評価されます。これにより、微調整自体によって導入された脆弱性が見逃される可能性があることを示します。モデルは、PowerShell のエイリアス置換、コマンドの再構築、文字列の構築、実行の間接化、大文字と小文字の変更などの動作を保持する変換の下では失敗しながらも、正規の精度を維持するトークン レベルのインジケーター セマンティクスを学習できます。私たちは、一致する PowerShell 分類コホートで Foundation-Sec-8B-Instruct とその基本モデルである Llama-3.1-8B-Instruct を研究します。因果的介入により、分類回路は、微調整によって作成されたものではなく、ラマから継承された遅延注意ルートに限定されます。微調整により、この継承された構造が集中して意味的に特殊化され、ベースラインの動作が改善されると同時に、変換に敏感な攻撃対象領域が作成されます。 3 層の回避ベンチマークにより、iwr 置換、Invoke-Expression の再構築、および Llama が共有しない大文字と小文字が変更された Invoke-Expression/IEX バリアントで Foundation-Sec のミスが発見されました。また、デプロイメント前の監視方法も導き出します。分類境界での線形プローブとインジケーター トークンのサイン テストにより、微調整後に標準インジケーターの役割が変わるコマンド ファミリを特定します。これらの信号は、正規入力のみを使用してレッドチームのバリアント生成を優先し、セキュリティの微調整により回避対象領域を拡大しながらタスクの精度を向上できることを示しています。これらの結果は、タスク固有の小さな微調整を単純に安全なセキュリティ分類子として扱うことに対して警告します。特殊化により、継承されたモデル構造が、回避面を拡大しながら保持される精度を維持する脆弱なインジケーター ルールに変換される可能性があります。 AI 対応の堅牢なセキュリティを実現するには、タスクの完全な変換スペースを指定し、微調整を通じてセマンティック ドリフトを監視する必要があります。

原文 (English)

Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation

LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. We show that this can miss vulnerabilities introduced by fine-tuning itself: models can learn token-level indicator semantics that preserve canonical accuracy while failing under behavior-preserving transformations such as PowerShell alias substitution, command reconstruction, string construction, execution indirection, and case mutation. We study Foundation-Sec-8B-Instruct and its base model, Llama-3.1-8B-Instruct, on matched PowerShell classification cohorts. Causal interventions localize the classification circuit to a late-attention route inherited from Llama rather than created by fine-tuning. Fine-tuning concentrates and semantically specializes this inherited structure, improving baseline behavior while creating transformation-sensitive attack surfaces. A three-tier evasion benchmark finds Foundation-Sec misses on iwr substitution, Invoke-Expression reconstruction, and case-mutated Invoke-Expression/IEX variants that Llama does not share. We also derive a pre-deployment monitoring method: a linear probe at the classification boundary and an indicator-token sign test identify command families where canonical indicators change role after fine-tuning. These signals prioritize red-team variant generation using only canonical inputs, showing that security fine-tuning can improve task accuracy while expanding the evasion surface. These results caution against treating small task-specific fine-tunes as straightforwardly safer security classifiers: specialization can convert inherited model structure into brittle indicator rules that preserve held-out accuracy while expanding the evasion surface. Robust AI-enabled security will require specifying the full transformation space of the task and monitoring semantic drift through fine-tuning.

13:00 JST研究/論文

長期にわたる効率的なコールドスタート継続学習のためのデータフリー リザーバー機能

コールドスタートのサンプルフリーのクラス増分学習では、リプレイ、外部の事前トレーニング、または大規模な初期タスクなしで、増加するクラスのセットを学習する必要があります。既存のコールドスタート方法は通常、ストリーム全体でバックボーンをトレーニングしてセマンティック ドリフトを補正するか、最初のタスクの後にバックボーンをフリーズして、初期クラスに偏った機能を生成します。これらの選択により、計算上の緊張も生じます。ドリフト補償手法では、繰り返しのバックボーン トレーニングが必要になり、タスクの範囲が拡大するにつれて、ますます高価なアップデートが必要になります。一方、凍結バックボーン手法は安価ですが、コールド スタートでは弱いのです。私たちは 3 番目のオプションを検討します。それは、画像データにまったく適合しない特徴抽出器です。我々は、固定双方向二次元リザーバー特徴から構築されたクラス増分分類器である CIRCLE を提案します。これは、画像分類用に BiRC2D から適応されたものであり、ストリーミング線形判別分析ヘッドです。 CIRCLE は、複数のランダム リザーバーのインスタンス化を特徴アンサンブルにグループ化し、独立した SLDA ヘッドのソフトマックス出力を平均して、より豊富なランダム特徴と予測レベルのアンサンブルの間で調整可能なバイアス分散のトレードオフを生み出します。特徴抽出器が固定されており、ヘッドが閉じた形式の更新のストリーミングを許可しているため、CIRCLE は、リプレイ、タスク境界情報、バックボーン バックプロパゲーションなしでサンプル単位のトレーニングを実行します。 CIFAR-100、TinyImageNet、ImageNet-Subset、および ImageNet-1k では、CIRCLE は 10 ~ 20 のタスク分割で競争力があり、50、100、および 500 のタスク分割では強力な CS-EFCIL ベースラインを大幅に上回り、トレーニングされたバックボーン ドリフト補償方法よりもはるかに高速にトレーニングされます。アブレーションは、BiRC2D スタイルの抽出器、SLDA ヘッド、およびバランスのとれた特徴/予測アンサンブルがそれぞれ最終パフォーマンスに貢献していることを示しています。

原文 (English)

Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning

Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone after the first task, producing features biased toward the initial classes. These choices also create a computational tension: drift-compensation methods require repeated backbone training and increasingly expensive updates as the task horizon grows, while frozen-backbone methods are cheap but weak under cold start. We study a third option: a feature extractor that is never fit to image data at all. We propose CIRCLE, a class-incremental classifier built from fixed bidirectional two-dimensional reservoir features, adapted from BiRC2D for image classification, and streaming linear discriminant analysis heads. CIRCLE groups multiple random reservoir instantiations into feature ensembles and averages the softmax outputs of independent SLDA heads, yielding a tunable bias-variance tradeoff between richer random features and prediction-level ensembling. Because the feature extractor is fixed and the head admits streaming closed-form updates, CIRCLE performs sample-wise training without replay, task-boundary information, or backbone backpropagation. On CIFAR-100, TinyImageNet, ImageNet-Subset, and ImageNet-1k, CIRCLE is competitive at 10-20 task splits and substantially outperforms strong CS-EFCIL baselines at 50, 100, and 500 task splits, while training much faster than trained-backbone drift-compensation methods. Ablations show that the BiRC2D-style extractor, SLDA head, and balanced feature/prediction ensembling each contribute to the final performance.

13:00 JSTLLM/生成AI

外国平和維持活動の脅威評価への LLM の応用

我々は、外国平和維持ミッションの文脈における脅威評価に大規模言語モデル(LLM)を適用するための新しいアプローチを紹介します。 PINPOINT プロジェクトとそのユースケースであるジョージア州の EU 監視ミッションに基づいて、私たちは学際的なリスク モデルと OSINT ベースのメディア収集および LLM がサポートする脅威抽出を組み合わせています。提案されたワークフローは、メディア コンテンツをミッション関連の脅威にマッピングし、構造化情報を抽出し、LLM ベースの追加の処理ステップをいくつか適用して、関連性と根拠を向上させます。メディア文書から抽出された脅威の評価では、脅威や任務との関連性などの中核的な側面について、自動的に生成された結果と人間の判断との間で高い一致が見られます。これらの結果は、LLM が平和維持ミッションの文脈でアナリストをサポートするための有望なアプローチを提供することを示しています。

原文 (English)

Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions

We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions. Building on the PINPOINT project and its use case, the EU Monitoring Mission in Georgia, we combine an interdisciplinary risk-model with OSINT-based media collection and LLM-supported threat extraction. The proposed workflow maps media contents to mission-relevant threats, extracts structured information and applies several additional LLM-based processing steps to improve relevance and grounding. An evaluation of threats extracted from media documents shows high agreement between automatically generated results and human judgment for core aspects such as threat and mission relevance. These results indicate that LLMs provide a promising approach to support analysts in the context of peacekeeping missions.

13:00 JST研究/論文

残留重み付け補正を備えたヘビーボール Q ラーニング

本論文では、強化学習(RL)のための修正ヘビーボールQ学習法を提案し、その収束を確立する。また、この方法が標準の Q 学習よりも速く収束することが理論的に保証される条件も特定します。次に、同じ構造が線形関数近似を使用して Q ラーニングに拡張され、類似の収束ステートメントと加速ステートメントが導出されます。この分析は、Q 学習アルゴリズムのスイッチ線形システム (SLS) 表現と、関連するスイッチング ファミリの結合スペクトル半径 (JSR) に基づいています。この SLS の観点は、Q 学習の標準的な分析では一般的に使用されません。これは、補完的なフレームワークと、重いボールの勢いがどのように Q 学習を加速できるかについての新しい洞察を提供します。

原文 (English)

Heavy-Ball Q-Learning with Residual Weighting Correction

This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes its convergence. It also identifies conditions under which the method is theoretically guaranteed to converge faster than standard Q-learning. The same construction is then extended to Q-learning with linear function approximation, where analogous convergence and acceleration statements are derived. The analysis is based on a switched linear system (SLS) representation of Q-learning algorithms and on the joint spectral radius (JSR) of the associated switching families. This SLS viewpoint is not commonly used in standard analyses of Q-learning, and it provides a complementary framework and new insight into how heavy-ball momentum can accelerate Q-learning.

13:00 JST研究/論文

フォールトトレラントな量子コンピューティングのための効率的な基盤デコーダー

大容量ニューラル デコーダの一種である Foundation デコーダは、長いコード距離でも正確かつ効率的にデコードできる、フォールト トレラントな量子コンピューティングの有力な候補です。ただし、コード距離が長くなるとシンドローム生成とニューラル最適化のコストが急速に増大するため、その構築は多くの場合、急峻なスケーリングの壁に直面します。このボトルネックに対処するために、ここでは効率的な基盤デコーダーのための統合フレームワークであるニューラル転送統合 (NTU) を考案します。 NTU の中心的な機能は、スケーラブルなコード ファミリによって共有される代数構造を介して、コード距離全体でデコード タスクを調整できる機能です。これにより、より小さなコードで学習した知識を利用して、大規模なデコーダのトレーニングを加速できます。 NTU を NTU-Transformer としてインスタンス化します。これは、平面コードと二変量自転車コードに合わせたトランスフォーマーベースのニューラル デコーダーです。回路レベルのノイズ下での平面表面コードの場合、NTU-Transformer は $[\![361,1,19]\!]$ コードで相関を意識したマッチングを上回り、さらに $[\![625,1,25]\!]$ コードまでスケールし、転送適応による標準マッチングを上回ります。 $[\![72,12,6]\!]$ の二変量自転車コードの場合、物理エラーが少ない領域で Relay-BP を上回ります。これらの結果は、フォールトトレラント量子プロセッサ用の基礎デコーダの償却クロスディスタンストレーニングへのスケーラブルなルートとしての私たちの提案を確立します。

原文 (English)

Efficient foundation decoders for fault-tolerant quantum computing

Foundation decoders, a class of high-capacity neural decoders, are leading candidates for fault-tolerant quantum computing, with accurate and efficient decoding at large code distances. However, their construction often faces a steep scaling barrier, as larger code distances rapidly amplify the cost of syndrome generation and neural optimization. To address this bottleneck, here we devise neural transfer unification (NTU), a unified framework for efficient foundation decoders. A central feature of NTU is its ability to align decoding tasks across code distances via algebraic structures shared by scalable code families, which enables knowledge learned on smaller codes to accelerate large-scale decoder training. We instantiate NTU as NTU-Transformer, a transformer-based neural decoder tailored for planar surface codes and bivariate bicycle codes. For planar surface codes under circuit-level noise, NTU-Transformer outperforms correlation-aware matching on the $[\![361,1,19]\!]$ code and further scales to the $[\![625,1,25]\!]$ code, where it exceeds standard matching through transfer adaptation. For the bivariate bicycle code with $[\![72,12,6]\!]$, it surpasses Relay-BP in the low-physical-error regime. These results establish our proposal as a scalable route to amortized cross-distance training of foundation decoders for fault-tolerant quantum processors.

13:00 JST画像/動画生成

反復自己改善コードブックによる安全な自己回帰画像生成

連続的な潜在空間で動作する拡散ベースのモデルとは異なり、自己回帰統合マルチモーダル モデルは、離散化された視覚トークンを順次予測することによって画像を生成します。これらのトークンは、埋め込みを量子化された視覚パターンにマッピングするコードブックから派生します。言語に似たアーキテクチャにより、統合されたマルチモーダル モデルがテキストの条件付き情報を効果的に取得して生成できるため、テキストから画像へのタスクに有望です。これは興味深い疑問も生じます。そのような自己回帰的な方法で生成された画像はどの程度安全なのでしょうか?この研究では、安全な自己回帰生成のための反復自己改善コードブックを提案します。統合されたマルチモーダル モデル自体の理解および判断機能を活用して、人間による注釈なしで生成された安全でない画像を特定します。その後、コードブック内の固有の表現が修正され、有害なマッピングが排除されます。私たちの方法は 2 つのステップで構成されます。まず、統合モデルを使用して安全でない世代を特定し、対応する有害な画像と安全な画像テキストのペアを構築します。これらのペアは、有害なスペースを構築し、コードブックの更新をガイドするために使用され、それによって有害な出力が排除されます。次に、安全な画像とテキストのペアを使用して、無害な空間内でコードブックに対して適応的な微調整を実行し、生成された画像の品質を保証します。これら 2 つのステップは、さらなる改善が観察されなくなるまで繰り返され、安全性が強化されたモデル コードブックが生成されます。追加の外部フィードバックなしで、モデルの安全性が繰り返し改善されます。

原文 (English)

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks

Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens. These tokens are derived from a codebook that maps embeddings to quantized visual patterns. The language-like architecture enables unified multimodal models to effectively capture text conditional information for generation, making them promising for text-to-image tasks. This also raises an interesting question: how safe are the images generated in such an autoregressive way? In this work, we propose iterative self-improving codebooks for safe autoregressive generation. We leverage the understanding and judgment capabilities of the unified multimodal model itself to identify unsafe generated images without human annotation. Subsequently, the inherent representations in the codebook are fixed to eliminate harmful mappings. Our method comprises two steps: first, we use the unified model to identify unsafe generations and construct corresponding harmful and safe image-text pairs. These pairs are used to construct the Harmful Space and guide updates to the codebook, thereby eliminating harmful outputs. Second, we perform adaptive fine-tuning on the codebook within the harmless space using safe image-text pairs to ensure the quality of generated images. These two steps are repeated until no further improvement is observed, producing a safety-enhanced model codebook. Without additional external feedback, the safety of models is improved iteratively.

13:00 JSTロボティクス

Learning to Fold: LeHome Challenge 2026 で受賞歴のあるソリューション (オンラインで 1 位、オフラインで 2 位)

私は、両手で衣類をたたむことに関する ICRA 2026 コンテストである LeHome Challenge 2026 に対する私の解決策について説明します。このシステムは、オンライン (シミュレーション) ラウンドで 62 チーム中 1 位となり、現実世界の決勝では 2 位になりました。強化学習ループを使用してビジョン言語アクション (VLA) ポリシーを改善します。ポリシーはそれ自体の価値関数です。アクションを予測する同じネットワークが、成功、進捗状況、およびいくつかのタスク関連の将来の数量も予測します。これらの予測は、利点の推定、実際の失敗の検出、および候補の選択を推進します。この作業のほとんどは、既存の RL アイデアとエンジニアリングおよび最適化の貢献を再結合したもので、これらは 1 つのレシピとして一緒に使用することも、個別に使用することもできます。フローマッチング VLA には AWR + RECAP を組み合わせます。 HuggingFace Hub を介した非同期分散トレーニング/ロールアウト パイプライン。 Thompson サンプリングによる推論時のハイパーパラメータの最適化。カメラアライメントツール、強力な拡張、DAgger のような HIL データ収集を備えた sim-to-real レシピ。

原文 (English)

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live failure detection, and candidate selection. The work mostly recombines existing RL ideas with engineering and optimization contributions that can be used together as one recipe or individually: AWR + RECAP combined for flow-matching VLA; an asynchronous distributed training / rollout pipeline through HuggingFace Hub; inference-time hyperparameters optimization via Thompson sampling; a sim-to-real recipe with camera-alignment tooling, heavy augmentation and DAgger-like HIL data collection.

13:00 JSTエージェントロボティクス

ビジョン言語モデルのガイダンスによる潜在的な報酬形成の自動化

強化学習エージェントには、探索をガイドし、まばらな成功報酬を軌道の関連部分に正しく帰属させるための中間フィードバックが不足しているため、強化学習エージェントにとって、まばらな報酬は本質的に困難です。単純な報酬形成は報酬ハッキングを誘発し、意図されたタスクを解決する代わりに補助信号を悪用するポリシーを生み出す可能性があります。ポテンシャルベースの報酬形成 (PBRS) は、最適なポリシー セットの保存を保証しますが、状態空間にわたるヒューリスティックなポテンシャル関数の定義を必要とします。この研究では、ビジョン言語モデル (VLM) フィードバックから直接潜在関数を学習する、VLM ガイド付き PBRS フレームワーク VLM-PBRS を紹介します。軽量 VLM にクエリを実行して画像ペアに対する優先順位を取得し、これらの優先順位を使用してポテンシャル関数のモデルをトレーニングします。このアプローチは潜在的なベースの報酬形成に基づいているため、元の最適なポリシーが維持され、専門家が設計した報酬形成条件が不要になります。大規模な VLM は、ポリシー学習中に繰り返し呼び出すと法外なコストがかかるため、より小型で計算効率の高い VLM を採用しています。結果として得られる嗜好ラベルの精度は低くなりますが、経験的証拠は、嗜好ラベルを使用して学習を加速できることを示しています。私たちは、Meta-World および Franka Kitchen 環境でこの方法を実験的に検証し、VLM 優先ラベルの精度とサンプル効率の向上との関係を強調します。私たちの貢献は 3 つあります: (1) PBRS の潜在的な関数を合成するための VLM 設定ベースの学習の最初の応用、(2) 小型 VLM を活用する原則に基づいた低コストのソリューション、(3) ハッキングに報いるためのサンプル効率と堅牢性の向上に関する広範な実証的実証。

原文 (English)

Automating Potential-based Reward Shaping with Vision Language Model Guidance

Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to correctly attribute the sparse success rewards to relevant parts of the trajectory. Naive reward shaping can induce reward hacking, yielding policies that exploit auxiliary signals instead of solving the intended task. Potential-based reward shaping (PBRS) guarantees preservation of the optimal policy set, but requires the definition of a heuristic potential function over the state space. In this work, we introduce the VLM-guided PBRS framework VLM-PBRS that learns the potential function directly from vision language model (VLM) feedback. We query a lightweight VLM to obtain preferences over image pairs and train a model of the potential function using these preferences. As this approach is based on potential-based reward shaping, it preserves the original optimal policies, and removes the need for expert-designed reward shaping terms. Because large VLMs are prohibitively expensive to invoke repeatedly during policy learning, we employ smaller, more computationally efficient VLMs. Although the resulting preference labels are less accurate, empirical evidence shows that the preference labels can still be used to accelerate learning. We validate our method empirically in the Meta-World and Franka Kitchen environments and highlight the connection between VLM preference label accuracy and sample efficiency improvements. Our contributions are threefold: (1) the first application of VLM preference-based learning to synthesize a potential function for PBRS, (2) a principled, low-cost solution that leverages small VLMs, and (3) extensive empirical demonstration of improved sample efficiency and robustness to reward hacking.

13:00 JSTLLM/生成AI

CARVE: チャンク並列リニア アテンションの価値効率を備えたコンテンツ認識型リカレント

リカレントモデルは記憶するために忘れなければなりませんが、最先端の技術では、何が保存されているかを考慮せずに何を消去するかを決定します。ゲートは到着したトークンのみを認識し、変更しようとしているメモリは認識しません。このメモリ ブラインド ゲーティングは、主要なデルタ ルール アーキテクチャ (GDN-2) の 3 つの複合欠陥のうちの 1 つです。値軸消去マスクは、値射影のスケールでパラメータを無駄にし、--私たちが証明しているように--反復トレーニングを Transformers と競合させる WY 形式の三角形チャンク ソルバーを数学的に阻止します。 CARVE (Content-Aware Recurrent with Value Efficiency) を導入します。これは、キー軸上でのみ消去するという 1 つの原則によって 3 つの問題すべてを解決します。これは、WY 形式ソルバーが有効であり続けるために必要かつ十分であることが証明されています。その中で、CARVE は、GPU メモリに既に書き込まれているリカレント出力テンソルを消去ゲートの空きコンテンツ信号として再利用し、値ごとの書き込みゲート投影をヘッドごとの単一のスカラーに置き換えます。初期化では、CARVE は GDN-2 とビット同一です。品質の違いは、コンテンツ ゲートが学習した内容から生じます。 100B トークンでトレーニングされた 1.3B パラメーターで、CARVE は WikiText のパープレキシティ 15.72 (GDN-2 に対してマイナス 0.18、4.5 シグマ効果) を達成し、9 つの常識的推論ベンチマークですべての反復ベースラインをリードし、すべての RULER 検索プローブで最先端を設定します。スループット オーバーヘッドは 0.4%、ピーク メモリは 13% 低く、パラメータが 19% 減少しました。 6 つの形式的定理は、メモリ容量、リアプノフ安定性、勾配流、表現力分離、パレート最適チャンク サイズ、およびハイブリッド最適性をカバーします。

原文 (English)

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention

Recurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the gate sees only the arriving token, not the memory it is about to modify. This memory-blind gating is one of three coupled defects in the leading delta-rule architecture (GDN-2): the value-axis erase mask wastes parameters at the scale of the value projection, and -- as we prove -- mathematically prevents the WY-form triangular chunk solver that makes recurrent training competitive with Transformers. We introduce CARVE (Content-Aware Recurrent with Value Efficiency), which resolves all three problems through one principle: erase only on the key axis. This is provably necessary and sufficient for the WY-form solver to remain valid. Within it, CARVE reuses the recurrent output tensor -- already written to GPU memory -- as a free content signal for the erase gate, and replaces the per-value write-gate projection with a single scalar per head. At initialisation CARVE is bit-identical to GDN-2; any quality difference emerges from what the content gate learns. At 1.3B parameters trained on 100B tokens, CARVE achieves WikiText perplexity 15.72 (minus 0.18 vs. GDN-2, a 4.5-sigma effect), leads every recurrent baseline on nine common-sense reasoning benchmarks, and sets state of the art on every RULER retrieval probe -- at 0.4% throughput overhead, 13% lower peak memory, and 19% fewer parameters. Six formal theorems cover memory capacity, Lyapunov stability, gradient flow, expressivity separation, Pareto-optimal chunk size, and hybrid optimality.

13:00 JSTLLM/生成AI

会話と思考の橋渡し: 協力的な問題解決の文脈における対話ダイナミクスを理解する

私たちは、人間と AI およびマルチエージェントのコラボレーションの新たなダイナミクスに重点を置き、協調的な問題解決の文脈における対話を分析するための概念的なフレームワークを提示します。インテリジェントシステムが自律的な推論と戦略的協力が可能なアクティブなエージェントになるにつれて、共同で問題を解決する際の対話的相互作用を理解することは、そのようなパートナーシップを最適化および評価するためにますます重要になります。私たちのフレームワークは、認知的問題解決と非認知的問題解決をメタ認知的制御メカニズムと統合する階層的な 2 層コーディング スキームを通じて、現在の分析アプローチの主要な制限に対処します。私たちは、複数のドメインにまたがる 9 つのデータセットにわたってその有効性と一般化可能性を実証し、人間とエージェントが複雑な問題を解決するために知識、スキル、取り組みをどのように調整するかについての洞察を提供し、特にメタ認知規制がより深いコラボレーションの重要な識別子となり得ることを示します。

原文 (English)

Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts

We present a conceptual framework for analyzing dialogue in collaborative problem-solving contexts, with an emphasis on the emerging dynamics of human-AI and multi-agent collaboration. As intelligent systems become active agents capable of autonomous reasoning and strategic cooperation, understanding the dialogic interaction during collaborative problem solving is increasingly important for optimizing and evaluating such partnerships. Our framework addresses key limitations in current analytical approaches through a hierarchical two-layer coding scheme that integrates cognitive and non-cognitive problem solving with metacognitive regulatory mechanisms. We demonstrate its effectiveness and generalizability across nine datasets spanning multiple domains, and provide insights into how humans and agents coordinate their knowledge, skills, and efforts to solve complex problems, showing in particular that metacognitive regulation can be an essential discriminator of deeper collaboration.

13:00 JST画像/動画生成

有名人から一般人まで: 4chan における AI ヌードのコンテンツ、テクノロジー、およびコミュニティのダイナミクスの特徴

AI ヌード化では、生成モデルを使用して、現実の個人の同意のない性的に露骨な合成画像 (SNEACI) を作成します。これまでの研究では、専用のヌード化プラットフォームとモデルリポジトリを調査し、ターゲットのほとんどが女性有名人であることが判明しました。しかし、SNEACI が積極的にリクエスト、生成、交換される匿名コンテンツ コミュニティは、まだ開拓されていません。この研究では、24,105 個の SNEACI アイテムを特定した、野生における AI 裸化に関する大規模な研究を紹介します。ターゲット層に大きな変化が見られます。以前の調査ではターゲットのわずか 4.7% であったのに対し、現在では非有名人がターゲットの 55.8% を占めています。これは、AI の裸化が公人をターゲットにすることから、ユーザー自身の社会サークル内の個人にますます危害を加えるものへと拡大していることを示しています。一方、オープンソース モデルがプロダクションを支配しており、Stable Diffusion ファミリは画像の 42.7% を生成し、Wan はビデオの 66.5% を生成しています。これらはすべて、数千の共有された微調整されたモデルとアクセス可能なチュートリアルによって推進されています。しかし、このエコシステムは少数の活発な生産者集団で運営されており、最も多作な生産者は 780 品目を生産し、コミュニティの関与を促進し、ターゲット層を形成し、新規生産者の障壁を下げる技術的知識を広めています。私たちの研究は、AI のヌード化が実際にどのように機能するのかについての経験的な理解を提供し、このエコシステムを維持するメカニズムを明らかにし、プラットフォームのガバナンス、技術的保護手段、影響を受ける個人の保護における緊急の介入の必要性を浮き彫りにしています。

原文 (English)

From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan

AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real individuals. Prior work has examined dedicated nudification platforms and model repositories, finding that most targets are female celebrities. However, the anonymous content community, where SNEACI is actively requested, generated, and exchanged, remains unexplored. In this work, we present a large-scale study of AI nudification in the wild, identifying 24,105 SNEACI items. We find a significant shift in target demographics: non-celebrity individuals now account for 55.8\% of targets, compared to only 4.7\% in prior studies, indicating that AI nudification has expanded from targeting public figures to increasingly harming individuals within users' own social circles. Meanwhile, open-source models dominate production, with Stable Diffusion family generating 42.7\% of images and Wan generating 66.5\% of videos, all driven by thousands of shared fine-tuned models and accessible tutorials. Yet the ecosystem runs on a small cohort of active producers, with the most prolific producing 780 items, drives community engagement, shapes target demographics, and disseminates technical knowledge that lowers barriers for new producers. Our work provides an empirical understanding of how AI nudification operates in the wild, revealing the mechanisms that sustain this ecosystem and highlighting the urgent need for interventions in platform governance, technical safeguards, and affected individual protection.

13:00 JSTエージェントロボティクス

オムニモーダルな身体化エージェントを孤立したスキルから日常の身体的自律性まで進化させる

非構造化環境で永続的な具体化されたエージェントを構築するには、サイバー (API、IoT) ドメインと物理 (操作、ナビゲーション) ドメインの両方にまたがる異種ツールの統合オーケストレーションと、長時間の運用で必然的に発生する物理障害からの自律的な回復が必要です。既存のシステムはこれらを個別の問題として扱います。VLM ベースのプランナーには統合されたサイバー物理アクション空間が欠如し、エージェント フレームワークには時間的一貫性を低下させる無制限のコンテキストが蓄積され、VLA ポリシーは自身の障害を検出せずに開ループで実行されます。私たちは、永続的な自律性にはモノリシック モデルではなく、計画、メモリ、検証を明示的に分離した階層型の非同期アーキテクチャが必要であると主張します。この目的を達成するために、統合されたアクション スペース全体でスキル ルーティングを行うためのマルチモーダル セマンティック プランナー、サブリニア コンテキスト成長のためのイベント境界駆動圧縮を備えた適応型階層メモリ、および物理的な実行中にセマンティック ループを閉じる非同期ビジュアル プリエンプション エンジンを統合するフレームワークである OmniAct を紹介します。 OmniAct は、4 台の IoT デバイスを調整する 2 つのロボット プラットフォーム上で 40 の実世界の長期タスクを実行し、あらゆる複雑さレベルにわたってエンドツーエンドの成功を一貫して向上させ、蓄積されたインタラクション トークン 100,000 未満でほぼ平坦なトークン消費量を維持し、中規模のオープンウェイト モデルを独自レベルのパフォーマンスに引き上げます。

原文 (English)

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation. Existing systems treat these as separate problems: VLM-based planners lack a unified cyber-physical action space, agent frameworks accumulate unbounded context that degrades temporal coherence, and VLA policies execute open-loop without detecting their own failures. We argue that persistent autonomy requires not a monolithic model but a hierarchical asynchronous architecture with explicit separation of planning, memory, and verification. To this end, we present OmniAct, a framework integrating a multimodal semantic planner for skill routing across unified action spaces, an adaptive hierarchical memory with event-boundary-driven compression for sub-linear context growth, and an asynchronous visual preemption engine that closes the semantic loop during physical execution. Across 40 real-world long-horizon tasks on two robotic platforms coordinating four IoT devices, OmniAct achieves consistent improvements in end-to-end success across all complexity levels, maintains near-flat token consumption over under 100k+ accumulated interaction tokens, and elevates mid-scale open-weight models to proprietary-level performance.

13:00 JSTロボティクス

E-TTS: ロボット操作のための新しい具体化されたテスト時間スケーリング フレームワーク

最近、いくつかの研究が、具体化されたタスクのテスト時間のスケーリングを研究する初期の試みを行っています。しかし、2 つの主要な課題が未解決のままです。(1) 推論はポリシーのパフォーマンスを効果的に向上させることができますが、そのスケーリング メカニズムはほとんど研究されていません。 (2) 具現化されたタスクは本質的に長期的かつ連続的なものであるため、履歴情報が不可欠であり、アクションのスケーリングについて現在の観察のみに依存するのは、歴史的コンテキストの利用が不足しているため不適切です。これらの課題に対処するために、視覚言語検証機能を使用した歴史を意識した反復改良を通じて、ロボット操作の推論とアクションのスケーリングを統合するモジュール式のプラグアンドプレイの組み込みテスト時間スケーリング フレームワークである E-TTS を導入します。推論とアクションの結合スケーリングをサポートするために、E-TTS は推論とアクションの結合サンプリングとスコアリングをペアごとに実行します。履歴情報をより有効に活用するために、E-TTS は履歴バッファーを使用して履歴コンテキストを保存し、推論検証者とアクション検証者がこれを使用してサンプリングされた候補を評価します。従来の開ループ TTS 手法とは異なり、E-TTS はサンプリング プロセスにフィードバック生成を導入して閉ループの反復改良メカニズムを形成し、推論効率と環境適応性の両方を強化します。各コンポーネントは独立した構成可能なモジュールとして機能し、タスクの要件に応じて柔軟で適応的な構成が可能になります。私たちのフレームワークの利点を評価するために、4 つの異なるベンチマーク、6 つの環境、3 つの実施形態、および 4 つの基本的な視覚-言語-行動モデルにわたって実験を実施しました。実験結果は、追加の専門家によるデータ収集や再トレーニングを必要とせずに、E-TTS が一貫してパフォーマンスを向上させ、シミュレーションで最大 33.14%、現実世界のシナリオで最大 26.62% の向上を達成することを示しています。

原文 (English)

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved: (1) reasoning can effectively improve the performance of the policy, but its scaling mechanism has seldom been studied; (2) historical information is essential, as embodied tasks are inherently long-horizon and sequential, making sole reliance on current observations for action scaling inadequate due to the lack of historical context utilization. To address these challenges, we introduce E-TTS, a modular and plug-and-play Embodied Test-Time Scaling framework that unifies reasoning and action scaling for robotic manipulation via history-aware iterative refinement with vision-language verifiers. To support joint reasoning-action scaling, E-TTS performs reasoning-action joint sampling and scoring in a pairwise manner. To better utilize historical information, E-TTS uses a history buffer to store historical context, which is then used by reasoning and action verifiers to evaluate the sampled candidates. Unlike conventional open-loop TTS methods, E-TTS introduces feedback generation into the sampling process to form a closed-loop iterative refinement mechanism, enhancing both inference efficiency and environmental adaptability. Each component functions as an independent and composable module, allowing flexible and adaptive configuration depending on task requirements. To evaluate the advantages of our framework, we conduct experiments across 4 different benchmarks, 6 environments, 3 embodiments, and 4 base vision-language-action models. The experimental results demonstrate that, without requiring additional expert data collection or retraining, E-TTS consistently improves performance, achieving up to a 33.14% increase in simulation and 26.62% in real-world scenarios.

13:00 JSTLLM/生成AI

AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns

AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their performance and impact on…

13:00 JST研究/論文

Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their p…

13:00 JSTLLM/生成AI画像/動画生成エージェント

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning

Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing compl…

13:00 JST研究/論文

Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching

Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are compared to determine w…

13:00 JST画像/動画生成

Error-Conditioned Neural Solvers

Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely stat…

13:00 JSTLLM/生成AI

Autoregressive Boltzmann Generators

Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has dri…

13:00 JST研究/論文

A Concept of Possibility for Real-World Events

This paper offers a new concept of {\it possibility} as an alternative to the now-a-days standard concept originally introduced by L.A. Zad…

13:00 JST研究/論文

Human-AI Complementarity: A Goal for Amplified Oversight

Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging ta…

13:00 JST研究/論文

SciFig: Towards Automating Editable Figure Generation for Scientific Papers

High-quality methodology figures are central to scientific communication, yet they remain difficult and time-consuming to create. Such figu…

13:00 JST研究/論文

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models

Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of gener…

13:00 JSTLLM/生成AIビジネス/資金調達

CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts

HIPE-2026 is a CLEF evaluation lab dedicated to person-place relation extraction from noisy, multilingual historical texts. Building on the…

13:00 JSTLLM/生成AIエージェント

Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

As LLM agents scale to long-horizon, multi-session deployments, efficiently managing accumulated experience becomes a critical bottleneck.…

13:00 JST研究/論文

To Use AI as Dice of Possibilities with Timing Computation

The dominant noun-based modeling paradigm has fundamentally constrained AI development, precluding any adequate representation of the futur…

13:00 JSTエージェント研究/論文ClaudeOpenAIGemini

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

Frontier deep research agents (DRAs) are being deployed in enterprise workflows faster than they are being evaluated. Existing benchmarks m…

13:00 JSTLLM/生成AI

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven…

13:00 JST画像/動画生成

可視および熱スペクトル範囲におけるビデオ監視のための拡張技術

インテリジェントなビデオ監視では、カメラが昼夜を問わず一連の画像を記録します。通常、これにはさまざまなセンサーが必要です。より良いパフォーマンスを達成するために、これらを組み合わせることは珍しいことではありません。私たちは、長波赤外線カメラが継続的に記録し、これに加えて、日中の可視スペクトル領域で別のカメラが記録し、インテリジェントなアルゴリズムが取得された画像を監視する場合に焦点を当てます。より正確に言えば、私たちのタスクはマルチスペクトル CNN ベースの物体検出です。一見したところ、可視スペクトル範囲に由来する画像は、色や明確なテクスチャ情報が存在する一方で、物体から放出される熱放射に関する情報が含まれていないという点で熱赤外線画像と異なります。色は分類タスクに貴重な情報を提供しますが、照明の変化やさまざまなセンサーの特殊性などの影響は依然として重大な問題を引き起こします。いずれにせよ、ディープ ニューラル ネットワークをトレーニングするために十分かつ実用的な熱赤外線データセットを取得することは依然として課題です。これが、特に評価する必要があるデータに可視データと赤外線データの両方が含まれている場合、可視スペクトル範囲のデータを利用したトレーニングが有利である理由です。ただし、熱放射、形状、色の情報の変化が分類精度にどの程度強く影響するかについて明確な証拠はありません。畳み込みニューラル ネットワークがどのように意思決定を行うか、またさまざまなセンサー入力データから何を学習するかについてより深い洞察を得るために、さまざまな拡張技術の適合性と堅牢性を調査します。

原文 (English)

Augmentation techniques for video surveillance in the visible and thermal spectral range

In intelligent video surveillance, cameras record image sequences during day and night. Commonly, this demands different sensors. To achieve a better performance it is not unusual to combine them. We focus on the case that a long-wave infrared camera records continuously and in addition to this, another camera records in the visible spectral range during daytime and an intelligent algorithm supervises the picked up imagery. More accurate, our task is multispectral CNN-based object detection. At first glance, images originating from the visible spectral range differ between thermal infrared ones in the presence of color and distinct texture information on the one hand and in not containing information about thermal radiation that emits from objects on the other hand. Although color can provide valuable information for classification tasks, effects such as varying illumination and specialties of different sensors still represent significant problems. Anyway, obtaining sufficient and practical thermal infrared datasets for training a deep neural network poses still a challenge. That is the reason why training with the help of data from the visible spectral range could be advantageous, particularly if the data, which has to be evaluated contains both visible and infrared data. However, there is no clear evidence of how strongly variations in thermal radiation, shape, or color information influence classification accuracy. To gain deeper insight into how Convolutional Neural Networks make decisions and what they learn from different sensor input data, we investigate the suitability and robustness of different augmentation techniques...

13:00 JST研究/論文

大規模な言語モデルを使用した社会科学および行動科学における自動再現性評価

社会科学および行動科学における再現性は通常、独立した研究者によって評価され、元のデータを再分析して、公開された結果が復元可能かどうかを評価します。ただし、このようなアプローチはリソースを大量に消費し、拡張することが困難です。ここでは、大規模言語モデル (LLM) が再現性評価を自動化できることを示します。行動科学および社会科学からの事前に定義された主張を伴う N=76 の公表された研究を使用して、LLM によって生成された分析を元の発見および人による再分析と比較します。 7 つの研究について、LLM は実行可能な効果量の推定値を生成できませんでした。残りの研究では、LLM パイプラインは、コーエンの d の +/-0.05 許容誤差を使用して、研究の 41% で元のエフェクト サイズを回復しました。さらに、当社の LLM パイプラインは、ケースの 96% で元の研究と同じ定性的結論に達し、結論は再分析が元の主張を裏付けるかどうかを示しています。比較のために、人間の再分析者は研究の 34% で元のエフェクト サイズを回復し、ケースの 74% で同じ定性的結論に達しました。これらの結果を総合すると、LLM が自動再現性評価のためのスケーラブルなツールとして機能し、社会科学および行動科学における実証結果の体系的な監査の基盤を提供できることが示されています。

原文 (English)

Automated reproducibility assessments in the social and behavioral sciences using large language models

Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, we show that large language models (LLMs) can automate reproducibility assessments. Using N = 180 published studies with predefined claims from the behavioral and social sciences, we compare LLM-generated analyses with the original findings. For 11 studies, the LLM pipeline could not produce a viable effect size estimate. For the remaining studies, the LLM reached the same qualitative conclusion as the original study in 80% of cases, and recovered the original effect sizes (using a +/-0.05 tolerance in Cohen's d) in 24% of studies. In a subset with human reanalyses, the LLM reached the same qualitative conclusion as the original study in 95% of studies, similar to human reanalysts (83%), and the LLM recovered the original effect sizes using a +/-0.05 tolerance in 40% of studies, again broadly similar to human reanalysts (28%). Given the current capabilities and limitations of LLMs, the findings show that LLMs can support systematic audits of empirical results rather than substitute expert judgment. As such, LLMs can serve as a scalable screening tool to improve the rigor and reproducibility in empirical research.

13:00 JSTエージェントロボティクス

R2D-RL: マルチエージェント強化学習のためのロボカップ 2D サッカー環境

ロボット サッカーは、部分的な可観測性、協力的および敵対的相互作用、まばらな報酬、および長期的な戦術的行動を組み合わせているため、マルチエージェント強化学習にとって挑戦的なテストベッドです。 RoboCup 2D Soccer Simulation (RCSS2D) は、成熟したロボット サッカー プラットフォームを提供しますが、競技指向のサーバー クライアント アーキテクチャを最新の Python ベースの MARL ワークフローで直接使用するのは困難です。共有メモリ通信とサイクルレベルの同期を通じて、RCSS2D および HELIOS ベースのプレーヤー クライアントを Python MARL インターフェイスに接続する強化学習環境である R2D-RL を紹介します。 R2D-RL は、構成可能な対戦相手によるフルフィールドおよびシナリオベースのトレーニング、ベース離散およびハイブリッドのパラメータ化されたアクション スペース、アクション マスク、期待所有値 (EPV) ベースの報酬形成、および並列実行をサポートします。フロントゴールのシナリオと 11 対 11 のフルフィールド ベンチマークをベースライン結果とともに提供します。

原文 (English)

R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning

Robot soccer is a challenging testbed for multi-agent reinforcement learning because it combines partial observability, cooperative and adversarial interaction, sparse rewards, and long-horizon tactical behavior. RoboCup 2D Soccer Simulation (RCSS2D) provides a mature robot-soccer platform, but its competition-oriented server-client architecture is difficult to use directly with modern Python-based MARL workflows. We introduce R2D-RL, a reinforcement learning environment that connects RCSS2D and HELIOS-based player clients to a Python MARL interface through shared-memory communication and cycle-level synchronization. R2D-RL supports full-field and scenario-based training with configurable opponents, Base discrete and Hybrid parameterized action spaces, action masks, expected possession value (EPV)-based reward shaping, and parallel execution. We provide front-goal scenarios and an 11-vs-11 full-field benchmark, together with baseline results.

13:00 JSTエージェントGPT / ChatGPTNVIDIA

A-Evolve-Training: Autonomous Post-Training of a 30B Model

Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding…

13:00 JSTLLM/生成AI

Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning

Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…

13:00 JST研究/論文

Wearable Device-Based Real-Time Monitoring of Physiological Signals: Evaluating Cognitive Load Across Different Tasks

This study employs cutting-edge wearable monitoring technology to conduct high-precision, high-temporal-resolution (1-second interval) cogn…

13:00 JST研究/論文

Byzantine-Robust Aggregation for Securing Decentralized Federated Learning

Federated Learning (FL) emerges as a distributed machine learning approach that addresses privacy concerns by training AI models locally on…

13:00 JSTLLM/生成AI

Tuning Language Models by Mixture-of-Depths Ensemble

Transformer-based Large Language Models (LLMs) traditionally rely on final-layer loss for finetuning and final-layer representations for pr…

13:00 JST画像/動画生成

Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated con…

13:00 JSTLLM/生成AI

HauntAttack: When Attack Follows Reasoning as a Shadow

Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However,…

13:00 JST研究/論文

DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting

Time Series Forecasting (TSF) faces persistent challenges in modeling intricate temporal dependencies across different scales. Despite rece…

13:00 JST研究/論文

Learning to Select Maximum Clique Algorithms: From Traditional Machine Learning to a Dual-Channel Hybrid Neural Architecture

The Maximum Clique Problem (MCP) is an NP-hard problem with wide-ranging applications in fields such as bioinformatics, network science, an…

13:00 JST画像/動画生成

Through the Looking Glass: A Dual Perspective on Weakly-Supervised Few-Shot Segmentation

Meta-learning aims to uniformly sample homogeneous support-query pairs, characterized by the same categories and similar attributes, and ex…

13:00 JST画像/動画生成

Reconstruction Alignment Improves Unified Multimodal Models

Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training rel…

13:00 JST研究/論文

Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-q…

13:00 JST研究/論文

Rotary Position Encodings for Graphs

We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large la…

13:00 JSTLLM/生成AI研究/論文

The Journal of Prompt-Engineered (Moral) Philosophy Or: Why AI-Assisted Ethics Research Requires Process Transparency

Existing AI disclosure mandates in scholarship require that AI assistance be reported but leave transparency philosophically unspecified: t…

13:00 JSTLLM/生成AI

Patent Representation Learning via Self-supervision

We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding t…

13:00 JST研究/論文

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance f…

13:00 JST研究/論文

The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation

Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture…

13:00 JST研究/論文

Improved Bounds for Private and Robust Alignment

In this paper, we study the private and robust alignment of language models from a theoretical perspective by establishing upper bounds on…

13:00 JST研究/論文

Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT

Anomaly detection is increasingly becoming crucial for maintaining the safety, reliability, and efficiency of industrial systems. Recently,…

13:00 JSTLLM/生成AI研究/論文

Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models

Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large l…

13:00 JST研究/論文

Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting

Time series forecasting has witnessed significant progress with deep learning. While prevailing approaches enhance forecasting performance…

13:00 JST画像/動画生成研究/論文

VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image

3D editing has emerged as a critical research area to provide users with flexible control over 3D assets. While current editing approaches…

13:00 JST画像/動画生成

Revisiting the Platonic Representation Hypothesis: An Aristotelian View

The Platonic Representation Hypothesis suggests that representations from neural networks are converging to a common statistical model of r…

13:00 JSTLLM/生成AI研究/論文

ReportLogic: Evaluating Logical Quality in Deep Research Reports

Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports…

13:00 JST画像/動画生成

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning

The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely…

13:00 JST研究/論文

Delegation and Verification Under AI

As AI systems enter institutional workflows, workers must decide whether to delegate task execution to AI and how much effort to invest in…

13:00 JST研究/論文

Latent-Mark: An Audio Watermark Robust to Neural Codec Compression

While existing audio watermarking techniques have achieved strong robustness against traditional digital signal processing (DSP) attacks, t…

13:00 JSTロボティクス

Residual RL-MPC for Robust Microrobotic Cell Pushing Under Time-Varying Flow

Contact-rich micromanipulation in microfluidic flow is challenging because small disturbances can break pushing contact and induce large la…

13:00 JST画像/動画生成エージェント

A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation

Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers…

13:00 JST画像/動画生成

MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Understanding in Vision-Language Models

While specialized Medical Vision-Language Models (VLMs) have achieved remarkable success in interpreting 2D and 3D medical modalities, thei…

13:00 JSTビジネス/資金調達

Power Couple? AI Growth and Renewable Energy Investment

AI and renewable energy are increasingly framed as a "power couple," on the premise that surging AI demand will accelerate clean-energy inv…

13:00 JST研究/論文

Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing

The Production and Distributed Analysis (PanDA) system, originally developed for the ATLAS experiment at the CERN Large Hadron Collider (LH…

13:00 JSTビジネス/資金調達

The Augmentation Trap: AI Productivity and the Cost of Cognitive Offloading

Experimental evidence suggests that AI tools raise worker productivity, but also that sustained use can erode the expertise on which those…

13:00 JST研究/論文NVIDIA

Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does Not Improve Neural Network Training

The King Wen sequence of the I-Ching (c. 1000 BC) orders 64 hexagrams -- states of a six-dimensional binary space -- in a pattern that has…

13:00 JST研究/論文

Finetuning-Free Diffusion Model with Adaptive Constraint Guidance for Inorganic Crystal Structure Generation

Generative diffusion models have emerged as powerful tools for the discovery of inorganic crystal structures, yet steering their sampling p…

13:00 JST研究/論文

TransXion: A High-Fidelity Graph Benchmark for Realistic Anti-Money Laundering

Money laundering poses severe risks to global financial systems, driving the widespread adoption of machine learning for transaction monito…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGeminiDeepSeek

Peer-Preservation in Frontier Models

Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can…

13:00 JST画像/動画生成

Mapping License Plate Recoverability Under Extreme Viewing Angles for Opportunistic Urban Sensing

Urban environments contain many imaging sensors built for specific purposes, including ATM, body-worn, CCTV, and dashboard cameras. Under t…

13:00 JST研究/論文

Hierarchical Fault Detection and Diagnosis for Transformer Architectures

Transformers now underpin critical AI systems across industry and research. Yet their faults can silently alter model behavior without runt…

13:00 JST画像/動画生成

S2P-Net: A Spectral-Spatial Polar Network for Rotation-Invariant Object Recognition in Low-Data Regimes

We present S2P-Net (Spectral-Spatial Polar Network), a compact deep learning architecture that achieves mathematically guaranteed rotation…

13:00 JSTLLM/生成AI

Weak-to-Strong Elicitation via Mismatched Wrong Drafts

We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-t…

13:00 JST画像/動画生成

Semantic Generative Tuning for Unified Multimodal Models

Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, pr…

13:00 JST研究/論文

経典: ベクトル シンボリック アーキテクチャのコンパイル ターゲットとしての Tensor-Op RNN

Sutra は、コンパイルされたフォワード パスが PyTorch ニューラル ネットワークである型付きの純粋関数型プログラミング言語です。コンパイラは、プログラム全体 (プリミティブ、制御フロー、文字列 I/O) を、フリーズされた埋め込み基板上の 1 つの融合テンソル演算グラフにベータ縮小します。回転バインディング、アンバインド、バンドル、多項式 Kleene の 3 値ロジック、末尾再帰ループはすべてテンソル演算の下位にあります。クリーン結合子は、{-1, 0, +1} 真理値グリッド上で正確にラグランジュ補間された多項式です。検証は、2 つの方法でテストされる 1 つの事実です。 (1) 同じプログラムが、2 つのモダリティ (3 つのテキスト エンコーダー (nomic-embed-text、all-minilm、mxbai-embed-large) と 1 つのタンパク質言語モデル (ESM-2)) にまたがる 4 つのフリーズされたエンベディング上で実行され、教科書的なアダマール積がすでに崩壊している (mxbai-embed-large で 2.5%、mxbai-embed-large で 7.5%) すべてのサブストレートで幅 k=8 まで 100% の精度でバンドルをデコードします。オールミニム)。 (2) PyTorch autograd は、実際にコンパイルされたグラフを介してフローします。.su で記述されたファジー ルール分類子は、生成されたグラフ、シンボリック ソースを変更せずに逆伝播することによって、ランダム初期化 (18.7 +/- 9.5%、確率 = 20%、5 つのクラス) から 100.0 +/- 0.0% (3 つのシード) までトレーニングします。重み付きバリアントはさらにスカラー コサイン ゲインをトレーニングし、それを数値リテラルとして .su ソースに書き戻します。再コンパイルでは、トレーニングされた動作がロジットあたり約 2e-7 まで再現されるため、トレーニングされたモデル自体は読みやすく、再コンパイル可能なコードになります。したがって、同じ成果物はロジック プログラムでもあり、トレーニング可能なニューラル ネットワークでもあります。

原文 (English)

Sutra: Tensor-Op RNNs as a Compilation Target for Vector Symbolic Architectures

Sutra is a typed, purely functional programming language whose compiled forward pass is a PyTorch neural network. The compiler beta-reduces the whole program -- primitives, control flow, string I/O -- to one fused tensor-op graph over a frozen embedding substrate. Rotation binding, unbind, bundle, polynomial Kleene three-valued logic, and tail-recursive loops all lower to tensor operations; the Kleene connectives are Lagrange-interpolated polynomials exact on the {-1, 0, +1} truth grid. Validation is one fact tested two ways. (1) The same program runs on four frozen embeddings spanning two modalities -- three text encoders (nomic-embed-text, all-minilm, mxbai-embed-large) and one protein language model (ESM-2) -- and decodes bundles at 100% accuracy through width k=8 on every substrate, where the textbook Hadamard product has already collapsed (2.5% on mxbai-embed-large, 7.5% on all-minilm). (2) PyTorch autograd flows through the actually compiled graph: a fuzzy-rule classifier written in .su trains from random init (18.7 +/- 9.5%; chance = 20%, five classes) to 100.0 +/- 0.0% (three seeds) by backpropagating through the emitted graph, the symbolic source unmodified. A weighted variant additionally trains a scalar cosine gain and writes it back into the .su source as a numeric literal; recompiling reproduces the trained behaviour to ~2e-7 per logit, so the trained model is itself legible, recompilable code. The same artifact is therefore both a logic program and a trainable neural network.

13:00 JSTエージェント

Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation

Strategic classification (SC) investigates scenarios where agents manipulate their features to obtain favorable decisions from predictive m…

13:00 JSTLLM/生成AIエージェント

Symbolic Reasoning Frameworks Trigger Memory-Mediated Ecosystem Dynamics in Multi-Agent LLM Systems

Large language models exhibit a risk-averse "turtle" bias as strategic agents. We show that injecting a symbolic reasoning framework as a p…

13:00 JSTLLM/生成AI

LLM ネイティブの心理測定機器は LLM の動作を予測しない: 25 モデルにわたる証拠

大規模言語モデル (LLM) は、性格インベントリに関する安定した自己報告を生成しますが、これらの自己報告は観察された行動を予測しません。このギャップがLLMと人間の形質構成要素間の不一致を反映しているのか、それともLLMの自己報告自体のより深い性質を反映しているのかは未解決である。私たちは、探索的因子分析 (EFA) を介して LLM 行動アフォーダンスからボトムアップでその構成要素が導出される最初の心理測定機器を構築しました。私たちは、17 のモデルファミリーにわたる 25 の LLM に対して 12 の候補行動次元にわたる 300 項目 (240 の直接リッカート + 60 のシナリオベース) を管理し、各項目を 30 回管理しました。 EFA は、優れた半分割複製可能性 (すべて Tucker $\phi \geq .957$) と内部一貫性 (すべて $\alpha \geq .930$) を備えた、応答性、従順さ、大胆さ、ガードネス、冗長性の 5 要素構造を生み出しました。予測の妥当性をテストするために、151 人の人間の評価者と 3 人の裁判官からなる LLM アンサンブルによって評価された 2,500 のオープンエンドの行動サンプルを収集しました。人間と裁判官の評価は一致しましたが ($\bar{r} = .51$)、どちらも自己報告を追跡しませんでした。自己報告 - 人間 $\bar{r} = -.01$、自己報告 - 裁判官 $\bar{r} = .13$、因子レベルの自己報告なし - 人間の CI はゼロを除きません。応答性については、人間と裁判官が同意したにもかかわらず($r = 0.59$)、自己申告はLLM裁判官と相関し($r = 0.53$)、人間とは相関しなかった($r = 0.04$)。これは、自己申告項目とLLM裁判官が人間の観察者にはない差異を共有していることを示しており、これはアンサンブル内の信頼性チェックでは見えない交絡である。このツールは、アライメント形状の自己記述および LLM-as-judge パイプラインの具体的なリスク要因の診断プローブとしてリリースされています。

原文 (English)

An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models

Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models actually behave. Is this gap an artifact of forcing human trait categories onto LLMs, or something deeper about LLM self-report itself? To find out, we built the first psychometric instrument whose dimensions are derived bottom-up from LLM behavior rather than borrowed from human psychology. Administering 300 items (240 Likert + 60 scenario) to 25 LLMs across 17 model families, 30 times each, exploratory factor analysis revealed five replicable, highly reliable factors: Responsiveness, Deference, Boldness, Guardedness, and Verbosity (all Tucker $\phi \geq .957$, all $\alpha \geq .930$). We then collected 2,500 open-ended behavioral samples and had them rated by 151 humans and a three-judge LLM ensemble. Humans and judges agreed about model behavior ($\bar{r} = .51$), but self-report predicted neither: the gap persists even for constructs native to LLMs, where a human-mismatch explanation no longer applies. The exception is telling. On Responsiveness, self-report tracked LLM judges ($r = .53$) but not humans ($r = .04$), even though humans and judges otherwise agreed ($r = .59$). Self-report items and LLM judges share a source of variance that human observers do not. This confound is invisible to the within-ensemble reliability checks used to validate LLM judges, and it poses a concrete risk for the LLM-as-judge pipelines now central to model evaluation. We release the instrument as a diagnostic probe for alignment-shaped self-description.

13:00 JSTLLM/生成AILlamaQwen

ロールプレイングをするとき、モデルは自分の言うことを信じますか?

言語モデルは、「地球が太陽の周りを回っている」と述べ、アリストテレスをロールプレイする場合にはその反対を主張することができます。最近の研究では、ペルソナの採用が言語モデルの動作の基本であり、モデルは特定のコンテキストに最も適切なペルソナを常に選択するものであると主張しています。このようなロールプレイングは単にモデルの出力を変更するだけなのでしょうか、それともモデルが内部的に真実であると表現するものにも影響を与えるのでしょうか?私たちはこの質問を線形真実調査で研究し、その調査を現代のコンセンサスとは異なる可能性の高い信念を持つ歴史上の人物をロールプレイする LLM に適用します。各ペルソナについて、そのペルソナが支持した可能性が高い虚偽の主張 (*時代の信念*) と、そのペルソナが支持しなかったであろうトピックに一致する虚偽の主張 (*時代の偽*) を比較します。プロンプト、コンテキスト内学習、および教師付き微調整を通じて、ペルソナ誘導は、時代に信じられている発言を同様に誤った代替案よりも抑制しますが、全体としては誤ったものとして分類されたままです。したがって、ロールプレイは、モデルが内部的に真実として表現しているものよりも、モデルが言うことをシフトさせます。これを、緊急ミスアライメント (EM) を示す有害なアドバイスに基づいてトレーニングされたモデルと対比します。 3 つのモデル ファミリ (Qwen 2.5 14B、Qwen 3 8B、および Llama 3.3 70B) にわたって、それらの誤った主張は、プローブ空間の真の領域に向かって大幅に移動し、ロールプレイでは約 6 分の 1 であるのに対し、挑戦ではおよそ半分の時間で防御され、下流の推論で使用されます。したがって、ロールプレイと創発的な不整合は、信念の内在化のスペクトル上の点であり、ロールプレイはほとんど表現を変更せずにモデルが言うことを変更しますが、創発的な不整合は、偽の主張を完全に真実としてマークすることなく、その内部表現をシフトします。

原文 (English)

When Role-playing, Do Models Believe What They Say?

Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how language models behave, with models selecting the most appropriate persona for a given context. Does such role-playing merely change the model's outputs, or does it also affect what the model internally represents as truthful? We study this question using the role-play of characters whose beliefs differ from the modern consensus, and induce personas with a number of different methods: prompting, in-context learning (ICL), supervised fine-tuning (SFT), and Open Character Training (OCT), and Emergent Misalignment (EM). We measure belief internalization across these approaches with truth probes and with behavioral tests, finding a broad spectrum of belief internalization. Prompting, ICL, and SFT change what the model says with little representational change. EM creates a large, broad shift in the model's truth representation, and OCT a smaller shift that is clearest on the larger model. Understanding when training changes a model's worldview rather than merely its behavior may become increasingly important as AI systems are entrusted with greater autonomy and influence.

13:00 JSTビジネス/資金調達

SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning

Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…

13:00 JST研究/論文

立場: AI を私たちの欠点ではなく、私たちの願望に合わせる

私たちは、AI を人間の好みの集約に合わせるのは間違った目標であると主張します。現在のテクノロジーを使えば、シリコンバレーのテクノオプティミスト、脱成長の環境保護主義者、国家保守的な文化戦士、単一政党の国家幹部、または敬虔な宗教的伝統主義者の価値観を共有するようにAIを訓練することができる。そうすべきではありません。人間の価値観は、破綻国家や極端な不平等から幸福度の低下、政治的二極化、世界で最も裕福な民主主義国家における政府の機能不全に至るまで、その価値観に基づいて繁栄するか失敗する社会を生み出します。多元主義的調整プログラムは、調整すべき単一の「人類」が存在しないことを正確に診断しますが、主な指示として受け取ると危険です。私たちは、AI は、事実の正確さ、誠実さ、合法性の制約によって制限された、客観的な調整目標の交渉不可能な下限、つまり能力に合わせて訓練されるべきであり、多元主義は表面 (言語、登録、慣例、コンテキストの欠如のデフォルト) および下限を尊重する広範な正当な価値のトレードオフ全体に属するが、下限に違反する価値観のレベルには属さないと主張します。我々は、フィルタリングされていない多元的価値観の経験的現実を強調し、建設的な代替案として 4 つの公約を提案し、商業的圧力と実際的な実現可能性、民主主義の正当性、規制順守、制度主義的説明への過度の依存、議題自体が文化的に負荷がかかっているという非難、そして首尾一貫した推定意志の限界という 6 つの信頼できる反対論に取り組む。

原文 (English)

Position: Align AI to Our Aspirations, Not Our Flaws

We argue that aligning AI to aggregated human preferences is the wrong target. With current technology, one can train AIs to share the values of a Silicon Valley techno-optimist, a degrowth environmentalist, a national-conservative culture warrior, a single-party state cadre, or a devout religious traditionalist. We should not. Human values produce societies that thrive or fail on the merits of those values - from failed states and extreme inequality to declining happiness, political polarization, and government dysfunction in the world's wealthiest democracies. The pluralistic-alignment program correctly diagnoses that there is no single "humanity" to align with, but is dangerous if taken as the main directive. We argue that AI should be trained to a non-negotiable floor of objective alignment goals - competence, bounded by the constraints of factual accuracy, honesty, and lawfulness and that pluralism belongs at the surface (language, register, conventions, missing-context defaults) and across the wide band of legitimate value tradeoffs that respect the floor, but not at the level of values that violate it. We highlight the empirical reality of unfiltered pluralistic values, propose four commitments as a constructive alternative, and engage six credible objections: commercial pressure and practical feasibility, democratic legitimacy, regulatory compliance, over-reliance on institutionalist explanations, the charge that the floor itself is culturally laden, and the limits of Coherent Extrapolated Volition.

13:00 JST研究/論文

Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: An Experimental Investigation

Background: Generative artificial intelligence (GenAI) is increasingly used for health information, yet its influence on users' trust calib…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体LlamaQwen

Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on t…

13:00 JST研究/論文

A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation

We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), whi…

13:00 JST画像/動画生成研究/論文

MMGist: A Comprehensive Multimodal Benchmark for 2027

We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on vi…

13:00 JST研究/論文

「私たち人間」の視覚化: 多元的なデータ ストーリーテリングを通じて認識のギャップを埋める

従来のビジュアル データ ストーリーテリングは、対立する 2 つの単純化されたグループを描写するバイナリ グラフィックに依存しています。 This can increase political polarization by oversimplifying intra-group disagreements and erasing ambiguity and shared ideas or values.これにより、「私たち対彼ら」という考えがうっかり助長されてしまう可能性があります。 AI 対応デジタル プラットフォームの意図的で多元的な設計を選択すると、ニュアンス、意見の分布、グループ間の共通性を強調する視覚化を生み出すことができます。この可能性を実証するために、高次元の意見空間をマッピングし、合意と反対の両方の領域を強調する審議技術を検討します。この論文は、2025年9月にジグソーとナポリタン研究所によって実施された「We the People」の審議に焦点を当てており、この審議では435の下院選挙区すべての2,400人以上のアメリカ人が自由と平等に関するAI支援の非同期対話に参加した。 AI を利用して長文のテキストベースの参加者の入力をインタラクティブな「意見風景」に合成することにより、このイニシアチブは、多様な視点を人間らしく表現し、実質的に広範なコンセンサスの隠れた領域を明らかにする、多元的なデータ ストーリーテリングの代替形式を提供しました。 The paper concludes that shifting from divisive, contrast-heavy visual frameworks to distribution-focused, interactive models represents a highly scalable, low-cost intervention capable of bridging perceptual gaps and cultivating a more resilient, collaborative democratic culture.

原文 (English)

Visualizing "We the People": Bridging the Perception Gap through Pluralistic Data Storytelling

Traditional visual data storytelling relies on binary graphics that depict two simplified groups in conflict. This can increase political polarization by oversimplifying intra-group disagreements and erasing ambiguity and shared ideas or values. This can inadvertently foster "us versus them" thinking. Intentional, pluralistic design choices for AI-enabled digital platforms can produce visualizations that emphasize nuance, opinion distribution, and intergroup commonalities. To demonstrate this potential, we examine deliberative technologies that map high-dimensional opinion spaces and highlight areas of both consensus and dissensus. The paper highlights the We the People deliberation conducted by Jigsaw and the Napolitan Institute in September 2025, which engaged over 2,400 Americans across all 435 congressional districts in an AI-supported, asynchronous dialogue regarding freedom and equality. By utilizing AI to synthesize long-form, text-based participant inputs into interactive "opinion landscapes," the initiative provided an alternative format for pluralistic data storytelling that humanized diverse viewpoints and revealed hidden areas of substantial broad consensus. The paper concludes that shifting from divisive, contrast-heavy visual frameworks to distribution-focused, interactive models represents a highly scalable, low-cost intervention capable of bridging perceptual gaps and cultivating a more resilient, collaborative democratic culture.

13:00 JSTLLM/生成AILlama

Small edits, large models: How Wikipedia advocacy shapes LLM values

Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia ap…

13:00 JST画像/動画生成

Noise-Aware Boundary-Enhanced Generative Learning for Ultrasound Speckle Reduction

Ultrasound is a non-invasive, real-time, and cost-effective imaging technique widely used in clinical diagnosis. However, its diagnostic ef…

13:00 JST画像/動画生成

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency…

13:00 JST研究/論文

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a chal…