Skip to the content.

AIニュース 2026-07-11

自動生成: 2026-07-11 12:16 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. How Deutsche Telekom is rewiring telecommunications with AIOpenAI

    How Deutsche Telekom is becoming an AI-native telco with OpenAI-trans…

  2. Open source AI matters more than ever, according to Hugging Face’s Clem DelangueTechCrunch AI

    Open source AI is booming, according to Hugging Face CEO Clem Delangu…

  3. Meta removes controversial AI feature on Instagram after backlashTechCrunch AI

    "Our intent was to provide a useful creative tool and to give people…

  4. Apple sues OpenAI over alleged trade secret theftTechCrunch AI

    Apple alleges the misconduct was directed by OpenAI's senior leadersh…

  5. SK Hynix raises $26.5B in the biggest foreign IPO in US history, is urged to build new US fabsTechCrunch AI

    The AI chip boom just produced its biggest Wall Street moment yet. No…

  6. テスラ車内で「Grok」と会話、日本でも展開へ ナビ設定やルート確認を音声でITmedia AI+

    対話型AI「Grok」を日本のテスラ車で利用できるようになった。Tesla Japanが7月10日、X公式アカウントで発表した。利用には、…

  7. 「まるで人間」 OpenAIの新モデル「GPT-Live」のトーク力が話題 間を空けずに考えながら会話できるITmedia AI+

    OpenAIが提供を始めた音声会話向けモデル「GPT-Live」がXで話題だ。同社はこれまでも音声会話モデルを提供していたが、聞き取りと発…

トピック別件数

日本語メディア3件

ITmedia AI+ (日本語)

21:15 JSTその他Grok

テスラ車内で「Grok」と会話、日本でも展開へ ナビ設定やルート確認を音声で

対話型AI「Grok」を日本のテスラ車で利用できるようになった。Tesla Japanが7月10日、X公式アカウントで発表した。利用には、ソフトウェアバージョン2026.20以降、「プレミアムコネクティビティ」(有料の車内通信プラン)の契約が必要になる。

16:45 JSTLLM/生成AIOpenAIGPT / ChatGPT2件の関連記事

「まるで人間」 OpenAIの新モデル「GPT-Live」のトーク力が話題 間を空けずに考えながら会話できる

OpenAIが提供を始めた音声会話向けモデル「GPT-Live」がXで話題だ。同社はこれまでも音声会話モデルを提供していたが、聞き取りと発話を同時に行う新アーキテクチャにより「会話が自然すぎる」などの声が上がっている。

出典:ITmedia AI+ITmedia AI+
13:00 JSTLLM/生成AI

「生成AIをもう手放せない人」が約6割 逆に“使わなくなったもの”1位は?

ICT総研の調査では、生成AIサービスが使えなくなると困ると答えた人は約6割に上った。生成AIサービスが日常的なツールとして定着する一方で、利用頻度が減ったものもあるという。それは何なのか。

海外メディア4件

TechCrunch AI (英語)

08:55 JSTその他

Meta removes controversial AI feature on Instagram after backlash

"Our intent was to provide a useful creative tool and to give people control over whether their public content could be referenced in this…

06:00 JSTLLM/生成AIOpenAI

Apple sues OpenAI over alleged trade secret theft

Apple alleges the misconduct was directed by OpenAI's senior leadership, including a longtime former employee.

04:00 JSTその他2件の関連記事

Open source AI matters more than ever, according to Hugging Face’s Clem Delangue

Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent…

出典:TechCrunch AITechCrunch AI
02:17 JSTハードウェア/半導体ビジネス/資金調達

SK Hynix raises $26.5B in the biggest foreign IPO in US history, is urged to build new US fabs

The AI chip boom just produced its biggest Wall Street moment yet. Now SK Hynix and Samsung are being asked to build U.S. factories.

公式ブログ1件

OpenAI (英語)

16:00 JSTLLM/生成AIOpenAI

How Deutsche Telekom is rewiring telecommunications with AI

How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and…

論文253件

arXiv cs.AI (英語)

13:00 JSTエージェントAnthropicClaude

Proactive Enterprise Agent のコンテキスト グラフ

検索拡張生成 (RAG) とエージェント フレームワークにより、エンタープライズ AI は大幅に進歩しましたが、エージェントは基本的に反応的なままであり、人間のクエリを待ってから行動します。この論文では、企業の真の生産性の向上には、プロアクティブなエージェント、つまり従業員が尋ねる前に関連性のある実用的な情報を提示するシステムが必要であると主張しています。私たちは、企業エンティティ、それらの関係、および時間の経過に伴う状態遷移をモデル化するライブ リレーショナル データ構造であるコンテキスト グラフを提案します。このグラフに基づいて構築され、状態の変化を継続的に監視するデルタ検出エンジン、候補者の洞察を緊急性、関連性、ペルソナ適合性によってランク付けするプロアクティブスコアラー、および根拠のある説明付きでランク付けされた通知を配信する LLM を利用したサーフェシング レイヤーを定義します。各コンポーネントを形式化し、統合されたプロアクティブ スコア関数を導き出し、NetworkX と Anthropic Claude API を使用して完全なエンドツーエンドの Python 実装を提供します。 3 つの一般的なエンタープライズ ケース スタディ (契約ライフサイクル管理、エンジニアリング インシデント対応、セールス パイプラインの衛生状態) にわたる評価では、コンテキスト グラフに基づくプロアクティブ性が Precision@5 0.83、誤検知率 0.11 を達成し、表面化までの平均時間が 47 分から (反応ベースライン) から 30 秒未満に短縮されることが実証されています。

原文 (English)

Context Graphs for Proactive Enterprise Agents

Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wait for a human query before acting. This paper argues that genuine enterprise productivity gains require proactive agents: systems that surface relevant, actionable information to workers before they ask. We propose the Context Graph, a live relational data structure that models enterprise entities, their relationships, and state transitions over time. Built on this graph, we define a Delta Detection Engine that continuously monitors state changes, a Proactivity Scorer that ranks candidate insights by urgency, relevance, and persona-fit, and a Surfacing Layer powered by an LLM that delivers ranked notifications with grounded explanations. We formalize each component, derive a unified Proactivity Score function, and provide a complete end-to-end Python implementation using NetworkX and the Anthropic Claude API. Evaluation across three generic enterprise case studies (contract lifecycle management, engineering incident response, and sales pipeline hygiene) demonstrates that context-graph-driven proactivity achieves Precision@5 of 0.83, a false positive rate of 0.11, and reduces mean time to surface from 47 minutes (reactive baseline) to under 30 second.

13:00 JST研究/論文

農業のレジリエンスを評価するための AI 統合モデル

農業のサプライチェーンは、生物物理学的システムと経済システムがリンクされているため、混乱に対して脆弱です。当社は、サプライチェーンのショックを分析するために経済モデル (GTAP) と生物物理モデル (APSIM) を統合する AI を活用したツールを開発し、政策立案者や市場参加者が自然言語で書かれたクエリと応答を通じて分野を超えた影響を評価できるようにします。

原文 (English)

AI-integrated models for assessing agricultural resilience

Agricultural supply chains are vulnerable to disruptions through linked biophysical and economic systems. We develop an AI-powered tool that integrates economic models (GTAP) with biophysical models (APSIM) to analyze supply chain shocks, enabling policymakers and market participants to assess cross-disciplinary impacts through queries and responses written in natural language.

13:00 JST研究/論文

人間の集合体と大規模言語モデルのための敵対的社会認識論

我々は、公の主張が証言、推論、制度的証明、暗黙の信頼の連鎖によって足場を固められる、高密度でインタラクティブなコミュニケーション環境のための敵対的社会認識論(ASE)の概要を説明します。このような状況では、エージェントは、個人的、評判的、修辞的、または物質的な利益のために、情報を歪曲、着色、省略、捏造、または戦略的に過小評価するインセンティブとアフォーダンスを持っています。私たちは、これらの現象は、認識バブル、エコー チェンバー、誤った情報の拡散などのよく知られた説明では適切に捉えられていないと主張します。説明が必要なのは、コミュニケーションエージェントが、通常は足場のあるアサーションを信頼できるものにするコミットメントと権利をどのように利用するかということです。私たちは、必要な分析を提供する言語を提供し、足場のある公共コミュニケーションにおける信頼を破壊するメカニズムの概要を説明し、アサーションを解釈するための推論主義的意味論で強化された認識論的ネットワークを利用して、推論チェーンの監査可能性を破壊することから生じる信頼違反を監査および是正するためのメカニズムの概要を提供します。

原文 (English)

Adversarial Social Epistemology for Assemblies of Humans and Large Language Models

We outline an adversarial social epistemology (ASE) for densely interactive communicative landscapes in which public assertions are scaffolded by chains of testimony, inference, institutional certification, and tacit trust. In such landscapes, agents have incentives and affordances to distort, color, omit, fabricate, or strategically under-specify information for private, reputational, rhetorical, or material gains. We argue that these phenomena are not adequately captured by familiar descriptions of epistemic bubbles, echo chambers, or misinformation diffusion. What requires explanation is how communicative agents exploit the commitments and entitlements that normally make scaffolded assertions trustworthy. We provide language that delivers the requisite analysis, outline mechanisms that subvert trust in scaffolded public communications, and outline machinery for auditing and redressing trust breaches arising from subverting the auditability of inferential chains, drawing on epistemic networks, enriched with an inferentialist semantics for interpreting assertions.

13:00 JSTLLM/生成AI

臨床ニーズと AI 機能の調整: 医療推論のための LLM に関する調査

大規模言語モデル (LLM) は医療における重要なツールとして浮上しており、臨床推論と患者ケアの可能性が高まっていることが示されています。この調査では、推論アプリケーションと要件に焦点を当てて、医療 LLM の最近の進歩を調査します。我々は、臨床実践と計算手法を結びつけるデュアルビューアプローチを提案します。臨床面では、ミラーのピラミッドに従って 5 レベルのコンピテンシー スキームを確立し、知識の想起から動的な症例管理に進みます。計算面では、演繹的、帰納的、およびアブダクティブな推論パターンを共通の医療目標とタスクに関連付けます。また、医療推論能力の 5 つのレベルにわたるベンチマーク データセットを紹介し、18 の最先端モデルに関する結果を報告します。これにより、医療専門家モデルが診断中心のタスクに優れているのに対し、一般モデルは意思決定支援と対話を主導していることが明らかになりました。最後に、データの制限、幻覚、グラウンディングの問題など、現在の進歩と未解決の課題について説明し、より安全で信頼性が高く、ワークフロー対応のシステムに向けた方向性を概説します。

原文 (English)

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

Large language models (LLMs) have emerged as important tools in healthcare, showing growing potential for clinical reasoning and patient care. This survey examines recent progress in medical LLMs, focusing on reasoning applications and requirements. We present a dual-view approach that connects clinical practice with computational methods. On the clinical side, we establish a five-level competency scheme following Miller's Pyramid, progressing from knowledge recall to dynamic case management. On the computational side, we link deductive, inductive, and abductive reasoning patterns to common medical goals and tasks. We also introduce a benchmark dataset spanning five levels of medical reasoning capability and report results on 18 state-of-the-art models, revealing that medical specialist models excel in diagnosis-centric tasks while general models lead in decision support and dialogue. We conclude by discussing current progress and open challenges, including data limitations, hallucination, and grounding issues, and outline directions toward safer, more reliable, and workflow-ready systems.

13:00 JSTLLM/生成AI

整合性の妥当性: 医療における AI を保証するための新しい基準

大規模言語モデル (LLM) はメンタルヘルス サポートの重要なプロバイダーとなっていますが、依然として注意経済の産物であり、その運用上および商業上のターゲットは、効果的な心理的サポートにしばしば必要とされる摩擦よりも持続的な関与を優先します。開発者の安全性への対応は主に事後対応であり、最も目に見える深刻な危害に対処する一方で、より微妙で長期的なリスクのパターン(依存性、境界侵食、歪んだ信念の増幅など)はあまり注目されていません。私たちは、LLM を構造的に安全にするためには、社会が人間の臨床実践の安全性をどのように保証するかを反映する 3 つのレベルで組織化された調整が必要であると主張します。 2) それらの値をモデルに埋め込むトレーニング。 3) 人間の実践に対する臨床監督と同様に、配備中のドリフトや長期的な危害を検出する監督。このように調整を組織化すると、調整の妥当性と呼ばれる構造が得られます。これは、システムの価値観、トレーニング体制、および監視メカニズムが安全で前向きな結果と一致していることを構造化して実証するものです。私たちは、健康における AI の規制構造として(生物学的妥当性の確立された構造に類推することによって)調整の妥当性を提案します。これは、システムが健康上のプラスの結果に調整され、それが可能な場合であっても害を及ぼさず、最終的には患者の利益につながるという信頼に賛成または反対する原則的な方法です。

原文 (English)

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.

13:00 JSTロボティクス

イディオバイオニクス: プライバシーとインテリジェントロボット義足の統合

人体は、生物学的システムとデジタル システムを緊密かつ永続的に結合するように設計された、成長を続けるテクノロジー ファミリーの中心にあります。ロボット義肢はこの密結合の代表的な例です。バイオニック四肢とも呼ばれるロボット義肢は、四肢を失った人が歩行や物体をつかむなどの日常生活活動をサポートする装置です。バイオニック手足は、高度なセンサーと人工知能ベースの制御アプローチとの統合により、知覚力と反応性が向上しています。その結果、このようなロボット義肢は、ユーザーと協調して適応できる半自律型のウェアラブルロボットシステムとみなすことができるようになりました。しかし、ロボット義足の能力を向上させるセンシングと制御の進歩は、ユーザーのプライバシーを侵害するために悪意のある組織によって悪用される可能性のある脅威ベクトルも導入します。次世代バイオニック手足の利点を十分に理解するには、これらのプライバシー リスクと、それがユーザーの導入にもたらす可能性のある障壁を直接理解し、対処することが重要であると私たちは主張します。したがって、この論文では、プライバシーとインテリジェントなバイオニック手足が交差する問題を総合的に調査するために、イディバイオニクスと呼ばれる新しい研究分野を紹介します。この論文の主な貢献として、私たちはイディバイオニクスを定義し、関連文献に根拠を置き、インテリジェントなバイオニック四肢の設計を悪用する可能性のある潜在的な敵対的攻撃を示し、議論する予備的な証拠を提供します。次に、ウェアラブルロボット工学やその他の人間に面する自律システムの研究者に関連する、イディバイオニクス内の未解決の研究質問の厳選されたリストを提供します。私たちは、イディバイオニクス研究がロボット義肢および関連するバイオニックデバイスの可能性を最大限に引き出すのに役立つことを期待しています。

原文 (English)

Idiobionics: The Unification of Privacy and Intelligent Robotic Prostheses

The human body is at the center of a growing family of technologies designed to tightly and persistently couple biological and digital systems. Robotic prostheses are a representative example of this tight coupling. Also referred to as bionic limbs, robotic prostheses are devices that support people who have lost limbs in pursuing daily life activities such as walking and grasping objects. Bionic limbs are now perceptive and responsive owing to their integration with advanced sensors and artificial intelligence-based control approaches. Consequently, such robotic prostheses can now be viewed as semiautonomous wearable robotic systems that can co-adapt with their users. However, the same sensing and control advancements that increase the capability of robotic prostheses also introduce threat vectors that could be exploited by malicious entities to violate the privacy of users. To fully realize the benefits of next-generation bionic limbs, we maintain it is important to directly understand and address these privacy risks and the barriers they might present to user adoption. This paper therefore introduces a new line of inquiry we term idiobionics to holistically investigate issues at the intersection of privacy and intelligent bionic limbs. As the main contribution of this paper, we define idiobionics, ground it in related literature, and provide preliminary evidence showing and discussing potential adversarial attacks that could exploit intelligent bionic limb designs. We then contribute a curated list of open research questions within idiobionics that are relevant to researchers in wearable robotics and other human-facing autonomous systems. We expect that idiobionics research will help unlock the full potential of robotic prostheses and related bionic devices.

13:00 JST研究/論文DeepSeek

Infinity-Parser2 テクニカルレポート

我々は、エンドツーエンドの文書解析のための制御可能なデータ合成パイプラインとマルチタスク強化学習を組み合わせた大規模なマルチモーダル モデルである Infinity-Parser2 を紹介し、忠実に注釈が付けられた解析コーパスの持続的な不足に対処します。私たちの貢献は 3 つあります。まず、制御可能なレンダリング フレームワークと反復改良ループを組み合わせたスケーラブルな合成エンジンを構築し、それを使用して Infinity-Doc2-5M を構築し、オープンソース化します。これは、さまざまな文書タイプにまたがる 500 万サンプルのバイリンガル (中国語/英語) コーパスであり、要素の境界ボックス、正規のコンテンツ フォーム (Markdown、HTML、LaTeX、SMILES、構造化チャート)、およびフルページで注釈が付けられています。読む順番。次に、検証可能なマルチタスク報酬システムを導入します。これにより、8 つの共同トレーニング目標 (文書解析、レイアウト分析、表解析、数式解析、チャート解析、化学式解析、文書 VQA、および一般的なマルチモーダル理解) にわたって共同強化学習を可能にし、単一の最適化信号で認識、構造、および推論を統合します。 3 番目に、共有アーキテクチャの下で 2 つのバリアントをリリースします。Infinity-Parser2-Flash は、Infinity-Parser-7B と比較して $3.68\times$ のスループット向上を実現し、低遅延推論用に最適化されています。もう 1 つは、精度が重要な設定向けに設計された Infinity-Parser2-Pro です。 Infinity-Parser2-Pro は、olmOCR-Bench で 87.6%、ParseBench で 74.3% に達し、DeepSeek-OCR-2、PaddleOCR-VL-1.5、MinerU2.5 を上回り、チャート、化学式、ドキュメント VQA に対する強力な汎用性を備えています。

原文 (English)

Infinity-Parser2 Technical Report

We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora. Our contributions are threefold. First, we build a scalable synthesis engine, pairing a controllable rendering framework with an iterative refinement loop, and use it to construct and open-source Infinity-Doc2-5M: a 5-million-sample bilingual (Chinese/English) corpus spanning diverse document types, annotated with element bounding boxes, canonical content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. Second, we introduce a verifiable, multi-task reward system that enables Joint Reinforcement Learning across eight co-trained objectives (document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding), unifying perception, structure, and reasoning in a single optimization signal. Third, we release two variants under a shared architecture: Infinity-Parser2-Flash, optimized for low-latency inference with a $3.68\times$ throughput gain over Infinity-Parser-7B, and Infinity-Parser2-Pro, engineered for precision-critical settings. Infinity-Parser2-Pro reaches state-of-the-art 87.6% on olmOCR-Bench and 74.3% on ParseBench, surpassing DeepSeek-OCR-2, PaddleOCR-VL-1.5, and MinerU2.5, with strong generalization to charts, chemical formulas, and document VQA.

13:00 JSTLLM/生成AIGoogle

VectorizationLLM: スマートなベクトル化ベースの AI アシスタント

VectorizationLLM は、Google のオープンウェイト LLM に基づいた特殊な大規模言語モデルです。このモデルは、学生が MATLAB でスマートなベクトル化、時間/波動ベクトル解析、区分関数、フーリエ解析、および微分方程式を学習できるように設計されています。このコースのアプリケーションは、ニューヨーク工科大学オールド ウェストベリー校の電気およびコンピューター工学技術学科による CTEC 247: 応用計算解析 II です。 LLM モデルは、質問に直接答えることなく、授業中のメモの例を使用して概念の詳細な説明を提供する、教育的なアシスタントとして設計されています。このモデルは、RAG (Retrieval Augmented Generation) ナレッジ ベースとシステム プロンプト アーキテクチャを使用して設計されています。コード、テキスト、画像の両方の例が LLM 応答で提供されます。

原文 (English)

VectorizationLLM: Smart Vectorization Based AI Assistant

VectorizationLLM is a specialized Large Language Model based on Google open-weight LLMs. The model is designed to assist students to learn smart vectorization, time/wave vector analysis, piecewise functions, Fourier analysis, and differential equations in MATLAB. The course application is CTEC 247: Applied Computational Analysis II by the Department of Electrical & Computer Engineering Technology at New York Institute of Technology Old Westbury. The LLM model is designed to be an instructive assistant, providing detailed explanations of concepts with examples from in-class notes without providing direct answers to questions. The model is designed with a RAG (Retrieval Augmented Generation) knowledge base and system prompt architecture. Examples in both code, text, and images are provided in the LLM responses.

13:00 JST研究/論文

sEMG信号に基づくリアルタイムジェスチャ認識のためのグラフニューラルネットワークモデル

高度な義手や拡張現実をシームレスに制御するには、正確かつ即時の手のジェスチャー認識が不可欠です。この目的には、前腕から得られる表面筋電図 (sEMG) 信号が一般的に使用されます。この論文では、前腕の筋活動パターンに関する情報を含むグラフ ネットワークを利用した sEMG 表現の新しいアプローチを紹介します。これらのグラフネットワークに基づいて、グラフニューラルネットワークを用いたリアルタイムハンドジェスチャ認識が可能な機械学習アルゴリズムを開発しました。このアルゴリズムのパフォーマンスは、前腕の周囲に配置された 8 つの電極を備えた myoband から取得した sEMG 信号を使用して評価され、8 人の健康な被験者が参加しました。提案された方法は、平均分類精度 99% を実証し、最先端技術のパフォーマンスを上回りました。 M1 pro CPU を使用した場合、グラフの構築と予測の両方にかかる平均時間は 48 ミリ秒であり、このアプローチはリアルタイム アプリケーションに最適です。

原文 (English)

A Graph Neural Network Model for Real-Time Gesture Recognition Based on sEMG Signals

For seemless control of advanced hand prostheses and augmented reality, accurate and immediate hand gestures recognition is essential. Surface electromyography (sEMG) signals obtained from the forearm are commonly employed for this purpose. In this paper, we present a novel approach for sEMG representation that utilizes graph networks which contain information about muscle activation patterns in the forearm. Based on these graph networks, we have developed a machine learning algorithm capable of real-time hand gesture recognition using a graph neural network. The algorithm's performance was evaluated using sEMG signals acquired from myoband, which has 8 electrodes placed around the forearm, involving 8 healthy subjects. The proposed method demonstrated an average classification accuracy of 99\%, surpassing the performance of state-of-the-art techniques. The average time for both graph construction and prediction stood at 48ms utilizing a M1 pro CPU, rendering the approach well-suited for real-time applications.

13:00 JSTエージェント

ストレートスルー引受業務におけるエージェントティック AI と検索拡張モデル

人工知能 (AI) は、特に非構造化文書、異種データソース、および規制された意思決定ワークフローに対する推論を必要とする分野で、保険数理業務を再構築し始めています。現在、アクチュアリーは、従来のルールベースの自動化から、大規模言語モデル (LLM)、検索拡張生成 (RAG)、および計画、検索、ツールの呼び出し、反映を行うマルチエージェントの「エージェント」システムに至るまで、さまざまな設計空間に直面しています。このホワイトペーパーでは、これらの新しいアーキテクチャが、透明性、監査可能性、人間参加型ガバナンスなどの保険数理上の優先事項をどのようにサポートできるかを、ストレートスルーの意思決定プロセスに焦点を当てて検証します。これらのアイデアを具体化するために、私たちは小規模商業事業主保険(BOP)のストレート引受のためのエージェント AI フレームワークを開発および分析します。合成的だが現実的な実験環境を構築し、次の 3 つの引受パイプラインを比較します。(i) 単一 LLM ベースライン、(ii) ナイーブ RAG システム、(iii) ターゲットを絞った検索、サードパーティ データ チェック、および明示的な複数ステップのルール評価を組み合わせたマルチエージェント「Agentic RAG」パイプライン。エージェント システムは全体的に最高のパフォーマンスを発揮し、複数ステップのシナリオや情報が欠落しているシナリオで最大の効果が得られます。このシナリオでは、構造化された検索と反映により、サポートされていないストレートな決定をモデルが回避できます。

原文 (English)

Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting

Artificial intelligence (AI) is beginning to reshape actuarial practice, particularly in domains that require reasoning over unstructured documents, heterogeneous data sources, and regulated decision workflows. Actuaries now face a design space that ranges from traditional rule-based automation to large language models (LLMs), retrieval-augmented generation (RAG), and multi-agent ``agentic'' systems that plan, retrieve, call tools, and reflect. This paper examines how these emerging architectures can support actuarial priorities such as transparency, auditability, and human-in-the-loop governance, with a focus on straight-through decision processes. To make these ideas concrete, we develop and analyze an agentic AI framework for straight-through underwriting of small commercial Business Owner Policies (BOPs). We construct a synthetic but realistic experimental environment and compare three underwriting pipelines: (i) a single-LLM baseline, (ii) a naive RAG system, and (iii) a multi-agent ``Agentic RAG'' pipeline that combines targeted retrieval, third-party data checks, and explicit multi-step rule evaluation. The agentic system performs best overall, with the largest gains in multi-step and missing-information scenarios, where structured retrieval and reflection help the model avoid unsupported straight-through decisions.

13:00 JSTエージェント研究/論文

フィードバック操作の正則化: 模倣学習のためのオフライン エージェント調整を可能にする

強化学習 (RL) 研究では、エージェントが人間の価値観に従った行動を確実に学習できるよう、調整に焦点がますます移ってきています。人間によるデモンストレーションとフィードバックが調整に重要であることが証明されていますが、既存のアプローチでは主に、言語生成のコンテキスト バンディット フレーム化用に設計された多段階パイプラインを使用してこれらの信号を組み合わせています。しかし、これらの相補的な入力が、完全に逐次的な意思決定環境における単一ステージのオフライン トレーニング用のより豊富な相互接続された信号としてどのように機能するかを調査した研究はほとんどありません。我々は、評価フィードバックを修正信号として利用して模倣学習ポリシーの整合性を向上させる、アルゴリズムに依存しない手法であるフィードバック操作正則化 (FMR) を提案します。私たちは安全体育館環境をアライメント評価の原則に基づいたテストベッドとして適応させ、さまざまな模倣学習アルゴリズム全体で適性の向上とミスアライメントの最大 98% 削減を実証しています。 FMR は、整列が不十分で有益でないノイズの多いデモンストレーションから学習している場合でも、限られたデータ領域で堅牢性を維持します。

原文 (English)

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feedback have proven crucial for alignment, existing approaches predominantly combine these signals using multi-stage pipelines designed for the contextual bandit framing of language generation. Yet little work explores how these complementary inputs can serve as a richer, interconnected signal for single-stage offline training in fully sequential decision-making environments. We propose Feedback Manipulation Regularization (FMR), an algorithm-agnostic method that harnesses evaluative feedback as a corrective signal to improve the alignment of imitation learning policies. We adapt Safety Gymnasium environments to be a principled testbed for alignment evaluation, demonstrating improved aptitude and up to a 98\% reduction in misalignment across a range of imitation learning algorithms. FMR remains robust in limited data regimes, even when learning from scarce aligned and uninformative noisy demonstrations.

13:00 JST研究/論文

ナイジェリアの機械: ドメインベースの推論レイヤーを備えた低リソースの産業データセット

アフリカ経済向けの産業機械に関するモデルの準備ができた公開データは比較的少ない。そのため、定量的な分析を行ったり、その設定に基づいた数値タスクに基づいて言語モデルをトレーニングしたりすることが困難になります。この問題の一部を解決するために 2 つのものをリリースします。 1 つ目は、ナイジェリアの機械の使用状況と故障のデータセットです。28 の指標にわたる 89 の機械レベルの記録で、2006 年から 2025 年までのナイジェリアの製造業および石油・ガス部門をカバーしています。すべての記録には公的情報源が指定されており、コードブックによって解読されます。 2 つ目は、これらのまばらな数値から思考連鎖 (CoT) 推論の例を構築する方法です。結果は、94 行のプロンプト、完了、および推論トレース行になります。プロンプトでは、各行に、実際の指標、サブセクター、年、およびその指標が由来するレコードのソースの名前が表示されます。データ適応作業は Adaption Labs によって実行されました。その過程で、言語モデルを使用してデータセットを構築するときによくある問題について説明します。プロンプトは、実際のドメインについては何も表示せずに、実際の数値と一致させることができます。これを修正すると、ドメインベースのプロンプトの割合が以前のリリースの 78 件中 1 件から 94 件中 94 件に増加し、すべての検索応答がソース値 (84 件中 84 件) と一致するようになったことを示します。データ、推論レイヤー、行ごとの来歴ファイルを CC-BY-4.0 に基づいてリリースします。私たちは限界について明確にしています。 89 個のレコードと 1 つの観測値のみを含む 17 個の指標を含む、これは参照およびシード データセットであり、大規模なトレーニング セットではありません。ほとんどの推論行は、複数ステップの計算ではなく検索です。

原文 (English)

Nigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer

There is relatively little, public, and model-ready data on industrial machinery for African economies. This makes it hard to do quantitative analysis or to train language models on numeric tasks grounded in that setting. We release two things to help with part of this problem. The first is the Nigeria Machinery Usage and Failures Dataset: 89 machine-level records across 28 indicators, covering Nigeria's manufacturing and oil and gas sectors from 2006 to 2025. Every record names a public source and is decoded by a codebook. The second is a method for building chain-of-thought (CoT) reasoning examples from these sparse numeric values. The result is 94 prompt, completion, and reasoning-trace rows. In every row, the prompt names the real indicator, subsector, year, and source of the record it comes from. The data adaptation work was carried out by Adaption Labs. Along the way we describe a problem that is common when language models are used to build datasets. The prompts can match the real numbers while saying nothing about the real domain. We show that fixing this raises the share of domain-grounded prompts from 1 out of 78 in an earlier release to 94 out of 94, and that every retrieval answer now matches its source value (84 out of 84). We release the data, the reasoning layer, and a per-row provenance file under CC-BY-4.0. We are clear about the limits. With 89 records and 17 indicators that have only one observation, this is a reference and seed dataset, not a large training set. Most reasoning rows are retrieval rather than multi-step computation.

13:00 JST研究/論文

ペルソナ カートグラフィー: 重み空間における言語モデルの性格特性のグラフ化

大規模な言語モデルは、一般化と安全性を形成する繰り返しの行動パターン (ペルソナ) を示しますが、それらを分解、測定、制御するための信頼できるツールが不足しています。私たちの中心的な洞察は、OCEAN フレームワークを使用して、オープンさ、誠実さ、外向性、協調性、および神経症の観点からモデル ペルソナを記述することにより、ペルソナを行動特性の空間における位置として扱うことです。私たちは、個々の特性を増幅または抑制するように低ランクのアダプターをトレーニングし、人間が検証したパネル、特性固有の多肢選択ベンチマーク、および標準的な能力評価に対して調整された LLM 判定を使用して、その効果を評価します。 3 つのファミリー (4B ~ 32B) の 6 つのモデルにわたって、各アダプターがそのターゲット特性をスケールに応じてほぼ単調に動かし、他のアダプターとほぼ相加的に組み合わせて混合ペルソナを構築し、中程度のスケールで機能ベンチマークのパフォーマンスを維持していることがわかりました。さらに、誘発された特性軸が下流の評価における安全関連行動に影響を与えることを示します。たとえば、神経症傾向と同調性軸に沿って移動すると、それぞれイライラと媚びに影響を与えます。また、モデルのロールアウトから 4 つの解釈可能な行動要因 (口調、自発性、教訓主義、認識論的警戒) を回収する教師なし心理測定パイプラインも導入します。ペルソナ制御は、重み空間での特性の学習、スケーリング、構成の観点から考えることができ、性格測定、モデル編集、安全性の間の架け橋となります。

原文 (English)

Persona Cartography: Charting Language Model Personality Traits in Weight Space

Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them. Our central insight is to treat personas as positions in a space of behavioural traits, using the OCEAN framework to describe model personas in terms of Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. We train low-rank adapters to amplify or suppress individual traits, and evaluate their effects using an LLM-judge calibrated against a human-validated panel, trait-specific multiple-choice benchmarks, and standard capability evaluations. Across six models from three families (4B-32B), we find that each adapter moves its target trait largely monotonically with scale, combines approximately additively with other adapters to construct mixed personas, and preserves performance on capability benchmarks at moderate scales. We further show that the induced trait axes affect safety-relevant behaviour in downstream evaluations: for example, moving along neuroticism and agreeableness axes affects frustration and sycophancy respectively. We also introduce an unsupervised psychometric pipeline that recovers four interpretable behavioural factors (tone, initiative, didacticism, epistemic caution) from model rollouts. Persona control can then be considered in terms of learning, scaling, and composing traits in weight space, providing a bridge between personality measurement, model editing, and safety.

13:00 JST画像/動画生成

自閉症関連の自己刺激的な手の特異性のシーケンスに基づく分類におけるフレーム レートの効果の評価

自閉症スペクトラム障害 (ASD) は世界中で 7,500 万人以上の人に影響を及ぼしていますが、遠隔行動スクリーニングのためのスケーラブルな計算手法は依然として限られています。この研究は、ビデオからの自閉症関連の自己刺激行動の自動検出における 2 つの相補的な課題に取り組んでいます: (1) 最適なシーケンスベースのニューラル ネットワーク アーキテクチャと時間サンプリング レートの特定、および (2) 小さな行動データセットでのトレーニングのためのデータ拡張戦略の特徴付け。最初の目的では、長短期記憶 (LSTM) モデルとゲート反復単位 (GRU) モデルが、自己刺激行動診断 (SSBD) データセットの姿勢由来の特徴に基づいて、1、5、15、30、45、および 90 フレームのフレーム サンプリング間隔でトレーニングされました。どちらのアーキテクチャも、以前の畳み込みニューラル ネットワーク (CNN) のベースライン (62 ~ 76% の精度) を上回り、15 フレームごとのサンプリング間隔で 97.5% (LSTM) および 98.75% (GRU) のピーク精度を達成しました。 2 番目の目的では、10 のデータ拡張戦略が I3D 転移学習パイプラインに適用され、各技術の限界寄与を定量化するアブレーション研究が行われました。水平反転はスタンドアロンで最高の精度 (48.78%) を達成しましたが、拡張パイプラインからアップサンプリングを除外すると最大のパフォーマンス低下が生じ、複雑な動作ビデオ拡張の必要性が示されました。パーソナライズされた機械学習アプローチでは、被験者ごとのモデルが各ビデオの時間的に分割されたセグメントでトレーニングおよびテストされ、一貫した予測が生成されました (平均損失 1.84、SD 0.79)。これらの結果は、データが不足している臨床領域におけるビデオベースの行動分類のためのアーキテクチャの選択、サンプリング レート、および拡張戦略に関する具体的なガイダンスを実務者に提供します。

原文 (English)

Evaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand Idiosyncrasies

Autism spectrum disorder (ASD) affects over 75 million individuals worldwide, yet scalable computational methods for remote behavioral screening remain limited. This study addresses two complementary challenges in automated detection of autism-related self-stimulatory behaviors from video: (1) identifying the optimal sequence-based neural network architecture and temporal sampling rate, and (2) characterizing data augmentation strategies for training on small behavioral datasets. For the first objective, long short-term memory (LSTM) and gated recurrent unit (GRU) models were trained on pose-derived features from the Self-Stimulatory Behavior Diagnosis (SSBD) dataset at frame sampling intervals of 1, 5, 15, 30, 45, and 90 frames. Both architectures exceeded prior convolutional neural network (CNN) baselines (62-76% accuracy), with peak accuracies of 97.5% (LSTM) and 98.75% (GRU) at a sampling interval of every 15 frames. For the second objective, ten data augmentation strategies were applied to an I3D transfer learning pipeline, with an ablation study quantifying the marginal contribution of each technique. Horizontal flip achieved the highest standalone accuracy (48.78%), while exclusion of upsampling from the augmentation pipeline produced the largest performance degradation, indicating its necessity for complex behavioral video augmentation. A personalized machine learning approach, in which per-subject models were trained and tested on temporally split segments of each video, produced consistent predictions (mean loss 1.84, SD 0.79). These results provide practitioners with concrete guidance on architecture selection, sampling rate, and augmentation strategy for video-based behavioral classification in data-scarce clinical domains.

13:00 JSTエージェント

エージェントティック ニューラル アーキテクチャの検索

ニューラル アーキテクチャ検索 (NAS) 手法はますます効率化していますが、依然として手動で設計された検索スペースに制限されており、この検索スペースには相当なドメインの専門知識が必要であり、新しいタスクごとに再構築する必要があります。大規模言語モデル (LLM) は、無制限の空間でアーキテクチャを生成できますが、LLM 主導の設計と NAS 主導の検索の間で作業を最適に分割する方法はまだ解明されていません。私たちは、これら 2 つのパラダイムを橋渡しするメカニズムを提案します。LLM は、高品質のシード アーキテクチャを生成し、それを「スロット アーキテクチャ」に分解します。これは、手動エンジニアリングを行わずに、従来の NAS が探索できる、境界のあるタスク固有の検索スペースを自動的に定義する、名前付きの交換可能なモジュール スロットを備えた足場です。このメカニズムは、各コンポーネントの寄与を個別に測定できるモジュール式の 3 フェーズ パイプラインである AgentNAS でインスタンス化されます。 AgentNAS は、さまざまなモダリティ (NAS-Bench-360 および Unseen NAS) にわたる分類、密な回帰、セグメンテーション、およびマルチラベル タグ付けにまたがる 17 のタスクで、11 のタスクに関する新しい最先端技術を確立し、タスク固有の専門家の設計を含む公開されたベースラインを上回るパフォーマンスを示します。アブレーション研究は、2 つの検索メカニズムが広く補完的であることを示しています。LLM で生成されたシードは、大部分のタスクですでに公表されているベースラインを上回っており、NAS は、独立した LLM サンプルでは複製できない検索モードであるスロット間の組み合わせ組み換えを通じて、ほとんどの場合に追加の利益をもたらします。これらのパターンは、異なる能力レベルの 3 つの LLM にわたって当てはまり、分業が堅牢であることが確認されています。私たちのコードは https://github.com/alroimfebruary/AgentNAS で入手できます。

原文 (English)

Agentic Neural Architecture Search

Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task. Large language models (LLMs) can generate architectures in an open-ended space, but how to optimally divide the labor between LLM-driven design and NAS-driven search remains unexplored. We propose a mechanism that bridges these two paradigms: an LLM produces a high-quality seed architecture, then decomposes it into a "slotted architecture", a scaffold with named, interchangeable module slots that automatically defines a bounded, task-specific search space for conventional NAS to explore, without manual engineering. We instantiate this mechanism in AgentNAS, a modular three-phase pipeline in which each component's contribution can be measured independently. On 17 tasks spanning classification, dense regression, segmentation, and multi-label tagging across diverse modalities (NAS-Bench-360 and Unseen NAS), AgentNAS establishes a new state of the art on 11 tasks, outperforming published baselines including task-specific expert designs. Ablation studies show that the two search mechanisms are broadly complementary: the LLM-generated seed already surpasses published baselines on the majority of tasks, and NAS delivers additional gains in most cases through combinatorial recombination across slots, a mode of search that independent LLM samples cannot replicate. These patterns hold across three LLMs of different capability levels, confirming that the division of labor is robust. Our code is available at https://github.com/alroimfebruary/AgentNAS.

13:00 JSTLLM/生成AI

具体化された命題プロンプトが大規模言語モデルにおける構成と知識の二分法を解決する

LLM は構成性と知識のバランスをとるのに苦労することが多く、これを構成と知識の二分法として定義します。これに対処するために、質問に関連する命題を明示的に具体化するフレームワークである具体化された命題プロンプティング (CPP) を提案します。この結果は、CPP が、演繹的推論が優先される数学ベンチマークにおいて競争力を持ちながら、特に正確な知識が最も重要な医療ベンチマークにおいて推論パフォーマンスを大幅に向上させることを示しています。追加の実験により、CPP はさまざまな基礎モデルやパラメーター サイズに拡張可能であり、構成ベースのアプローチと知識ベースのアプローチの間のギャップを埋める基本的なパラダイムであることが明らかになりました。その結果、CPP は、論理的に組織され、事実に基づいた推論のための強固な基盤を提供することによって、構成と知識の二分法を解決します。

原文 (English)

Concretized Proposition Prompting Resolves Composition-Knowledge Dichotomy in Large Language Models

LLMs often struggle to balance compositionality with knowledgeability, a challenge we define as Composition-Knowledge Dichotomy. To address this, we propose Concretized Proposition Prompting (CPP), a framework that explicitly concretizes propositions relevant to questions. The results demonstrate that CPP significantly enhances reasoning performance, particularly in medical benchmarks where precise knowledge is paramount, while being competitive on math benchmarks where deductive reasoning is prioritized. Additional experiments reveal that CPP is scalable to various foundation models and parameter sizes, being a fundamental paradigm that bridges the gap between composition- and knowledge-based approaches. Consequently, CPP resolves the composition-knowledge dichotomy by providing a solid foundation for logically organized and factually grounded reasoning.

13:00 JSTLLM/生成AIエージェント

プロンプトから契約まで: 監査可能なエンタープライズ LLM エージェントのためのハーネス エンジニアリング

エンタープライズ大規模言語モデル (LLM) アプリケーションは、多くの場合、プロトタイプとして開始され、その動作はプロンプトと取得コンテキストによって伝えられます。製品化では、ソース境界、エンティティ ルーティング、応答コントラクト、および再現可能なトレースに関する要件が追加されます。我々は、このパターンを追跡可能で監査可能な LLM エージェント アーキテクチャに再構築するハーネス エンジニアリング アプローチを提案します。決定論的な動作は、置き換え可能な構成境界付近のコード、マニフェスト、スキーマ、および検証アーティファクトに移行しますが、ソースに裏付けられたクレームは実行時の回答の権限を維持します。韓国の企業グループ 5 社 (上場企業 25 社) のパブリック データ スライス上でインスタンスを作成し、3 つの研究課題を評価します。 (1) ハーネスは、固定検証シナリオ全体にわたって、ソース グラウンディング、エンティティ ルーティング、トレース、出力衛生、および推奨言語のコントラクトを保持します。フォールト挿入制御は、バリデーターが意図的に違反した契約にフラグを立てていることを確認します。 (2) モデル置換の下でハーネスが強制するチェック: 3 つのホストされたモデルにわたって、270 の構成境界の実行すべてに合格しました。失敗はモデルで構成された側に限定され、検出されて記録されました。 (3) コード所有の保証は負荷がかかり、プロンプトのみでは再現できません。モデルを固定して強制層のみを変更し、プロンプトの指示だけを行うと、推奨言語と内部トレース漏洩の違反がリーダーに到達し、ハーネスが完全にブロックします。ボルトオン式の外部ガードレールもそのような違反を防止しますが、拒否しすぎて、ハーネスが完全な実用性 (120/120) を維持するにもかかわらず、実用性が 88/120 に低下します。このアブレーションでは、コード所有の強制のみが安全性と実用性の両方を維持します。その結果、探索的なプロトタイプを、バージョン管理されたソース、制御、および検証成果物を備えた監査可能なアプリケーションに変えるための再利用可能なエンジニアリング パターンが得られます。

原文 (English)

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-engineering approach that reconstructs this pattern into a traceable, auditable LLM-agent architecture: deterministic behavior moves into code, manifests, schemas, and validation artifacts around a replaceable composition boundary, while source-backed claims remain the authority for runtime answers. We instantiate it on a public-data slice of five Korean corporate groups (25 listed companies) and evaluate three research questions. (1) The harness preserves its source-grounding, entity-routing, trace, output-hygiene, and recommendation-language contracts across the fixed validation scenarios; a fault-injection control confirms the validators flag deliberately broken contracts. (2) The checks the harness enforces held under model substitution: across three hosted models, they passed on all 270 composition-boundary runs; failures were confined to the model-composed side and were caught and recorded. (3) The code-owned guarantees are load-bearing, not reproducible by prompting alone: holding the model fixed and varying only the enforcement layer, prompt instructions alone let recommendation-language and internal-trace-leakage violations reach the reader, which the harness blocks entirely. A bolt-on external guardrail prevents such violations too but over-refuses, dropping utility to 88/120 where the harness preserves full utility (120/120); in this ablation, only code-owned enforcement preserves both safety and utility. The result is a reusable engineering pattern for turning exploratory prototypes into auditable applications with versioned source, control, and validation artifacts.

13:00 JSTLLM/生成AIGPT / ChatGPT

AI支援鑑別診断のための安全指向の仮説演繹フレームワーク

診断エラーは患者の安全に対する大きな脅威であるにもかかわらず、現在の大規模言語モデル (LLM) システムは診断を単発の予測タスクとして扱うことが多く、高リスクの代替案を見逃したり、その推論を厳密に検証したりすることに対する保護手段が欠けています。ここでは、仮説演繹的臨床推論のための安全性指向のフレームワークである AegisDx を紹介します。 AegisDx は、役割固有の契約、構造化された中間出力、証拠検索インターフェイス、および検証ゲートを通じて特殊な LLM コンポーネントを調整し、広範な鑑別診断を生成し、危険な「見逃してはならない」状態の明示的なスクリーニングを実施し、根拠のある医学的証拠に照らして推論を検証し、実行可能な次のステップを構築します。 AegisDx を 3 つのレイヤーにわたって評価しました。 NEJM と JAMA からの文献由来の症例報告では、GPT-oss-120B を共有バックボーンとしており、トップ 3 の診断精度は、JAMA 症例では 59.9% 対スタンドアロン LLM の 52.1%、NEJM 症例では 62.7% 対 51.4% でした。 Annals of Emergency Medicine の症例では、トップ 3 の精度は 85.7% 対 68.6% でした。医師の合意による見逃せない診断セットに対して、AegisDx は症例の 78.0% と 52.0% で、上位 3 つの診断のうち少なくとも 1 つのそのような症状を捕捉しました。 Yale New Haven Health System の 43 件の実際の救急部門のメモを GPT-5 と比較して医師が盲検で評価したところ、AegisDx は医師が評価した総合安全性スコアを 5 段階評価 (調整済み p = 2.1x10^-4) で 4.31 から 4.55 に改善し、見逃せない識別と推論の安全性が定性的に向上しました。私たちの調査結果は、生の予測精度のみを最適化するのではなく、安全指向の推論フレームワークとして診断 AI をエンジニアリングすることで、急性期治療のワークフローに対して、より安全で透明性があり、臨床的に意味のあるベッドサイドの意思決定支援層を提供できることを示唆しています。

原文 (English)

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning. Here, we present AegisDx, a safety-oriented framework for hypothetico-deductive clinical reasoning. AegisDx coordinates specialized LLM components through role-specific contracts, structured intermediate outputs, evidence-retrieval interfaces, and verification gates to generate broad differential diagnoses, enforce explicit screening for dangerous "must-not-miss" conditions, verify reasoning against grounded medical evidence, and structure actionable next steps. We evaluated AegisDx across three layers. On literature-derived case reports from NEJM and JAMA, with GPT-oss-120B as the shared backbone, Top-3 diagnostic accuracy was 59.9% versus 52.1% for the standalone LLM on JAMA cases and 62.7% versus 51.4% on NEJM cases. On cases from Annals of Emergency Medicine, Top-3 accuracy was 85.7% versus 68.6%; against physician-consensus must-not-miss diagnosis sets, AegisDx captured at least one such condition among its top three diagnoses in 78.0% of cases versus 52.0%. In a blinded physician evaluation of 43 real-world emergency department notes from the Yale New Haven Health System compared against GPT-5, AegisDx improved the physician-rated composite safety score from 4.31 to 4.55 on a 5-point scale (adjusted p = 2.1x10^-4), with qualitative gains in must-not-miss identification and reasoning safety. Our findings suggest that engineering diagnostic AI as a safety-oriented reasoning framework, rather than optimizing raw predictive accuracy alone, can provide a safer, more transparent, and clinically meaningful layer of bedside decision support for acute care workflows.

13:00 JSTLLM/生成AIClaude

LLM が同意するとき、それは正しいのでしょうか?信頼シグナルとしての自己一貫性とモデル間一致の監査

LLM-as-judge (Zheng et al., 2023) は、企業パイプラインにおける AI システムを評価するためのデフォルトになりつつあり、多くの場合、アンサンブル (Verga et al., 2024) または「専門家の混合」(Shazeer et al., 2017) の審査員パネルに拡張されます。これらのシステムは、一貫性 (審査員間またはモデル自体のサンプル間の一致) が正しさを示すという重要な前提を共有しています。この仮定が信頼できないことを示します。一致は正確さではありません。共有バイアス、記憶されたヒューリスティック、または真実ではなく事前のオプションの位置から、モデルがそれ自体と一致することもあれば、異なるモデルが互いに一致することもあります。私たちは、大規模なクロスランナー調査において、合意がそれでもなお使用可能な代用手段となるのはいつかと尋ねます。53 人のランナーが、モデル層、プロンプト、および GPQA Diamond と AIME のスケールの比較にわたって、割り当てられた重複するケースに対して K=50 のサンプルを抽出しました (265,000 のサンプル)。デプロイラベルとして多数決の正しさを使用し、階層的なランナークラスター化ブートストラップを使用すると、一致は肯定的だが弱い予測子 (rho 0.20-0.59、アイテムクラスター化リサンプリングではすべて肯定的) であり、その有用性はレジームに依存します。飽和していない中間層モデルとコンピューティングの割り当てに最適で、最も一貫したフロンティアモデル (一致度 >=0.8 の) では最悪 -- 自信過剰だが精度が低い -- です。 GPQA の訴訟結果エントリの 77%、そのうち 48% が間違っていました)。 3 つのクロード層に対する探索的なファミリー間チェックでは、同様のフロンティア過信が示されており、限界維持ヌルを超えるプロバイダー間で確信度の高いエラーが繰り返し発生しています。したがって、自己一貫性は、独立した信頼スコアではなく、正しさの条件付き代用値です。匿名化された実行ごとの行と回答の分布を公開します。

原文 (English)

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency -- agreement among judges, or among a model's own samples -- indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth. We ask when agreement is nonetheless a usable proxy, in a large-scale cross-runner study: 53 runners drew K=50 samples for assigned overlapping cases across comparisons of model tier, prompting, and scale on GPQA Diamond and AIME -- 265,000 samples. Using majority-correctness as the deployment label and a hierarchical runner-clustered bootstrap, agreement is a positive but weak predictor (rho 0.20-0.59, all positive under item-clustered resampling) whose usefulness is regime-dependent: best for unsaturated mid-tier models and for allocating compute, and worst -- over-confident yet no more accurate -- for the most consistent frontier model (agreement >=0.8 on 77% of GPQA case-result entries, 48% of those wrong). An exploratory cross-family check on three Claude tiers shows the same frontier over-confidence, with confident errors recurring across providers above a marginal-preserving null. Self-consistency is thus a conditional proxy for correctness, not a standalone confidence score. We publicly release the de-identified per-run rows and answer distributions.

13:00 JSTエージェントClaudeGPT / ChatGPT

説得攻撃により CoT モニタリングの効果が低下する可能性がある

思考連鎖 (CoT) の監視は、目に見える推論の痕跡によって誤った動作や欺瞞的な動作が表面化する可能性があるという前提に基づいており、AI エージェントにとって有望な安全メカニズムです。標準的なシナリオでは効果的ですが、最近の研究では、LLM は自然言語の引数がモデルの制約をオーバーライドする説得ベースのジェイルブレイクに対して依然として脆弱であることが浮き彫りになっています。この脆弱性が LLM の監視にも及ぶかどうかをストレス テストします。敵対的なエージェントは CoT モニターを説得して、モニターのポリシーに違反する提案されたアクションを承認させることができますか?私たちは 40 のタスクを含む評価フレームワークを設計し、エージェントがポリシー違反の提案について議論するように指示される、何千ものエージェントとモニターのやり取りを分析します。このような敵対的な設定では、スクラッチパッドが追加の説得チャネルを提供するため、エージェントの CoT 推論へのモニター アクセスが有害なアクションの承認を平均 9.5% 減少させるのではなく増加していることがわかりました。これに対処するために、ファクトチェック監視フレームワークを導入します。たとえば、Claude 3.7 Sonnet モニターと GPT-4.1 ファクトチェッカーの組み合わせなど、異なるモデル ファミリのファクト チェッカーとモニターの組み合わせでは、ファクト チェックとモニタリングの役割の両方に同じモデルを使用した場合、ポリシー違反アクションの承認がわずか 6% 減少するのに対し、最大 45% 減少することがわかりました。私たちの結果は、CoT モニタリングだけでは敵対的な説得に対して不十分である可能性があり、モデルの多様なファクトチェックが強力な軽減策を提供することを示しています。

原文 (English)

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, recent work highlights that LLMs remain vulnerable to persuasion-based jailbreaks, where natural-language arguments override model constraints. We stress-test whether this vulnerability extends to monitoring LLMs: can an adversarial agent persuade its CoT monitor to approve proposed actions that violate the monitor's policy? We design an evaluation framework with 40 tasks and analyze thousands of agent-monitor interactions, where agents are instructed to argue for policy-violating proposals. We find that in such adversarial settings, monitor access to the agent's CoT reasoning increases rather than decreases approval of harmful actions on average by 9.5%, as the scratchpad provides an additional persuasion channel. To address this, we introduce a fact-checking monitoring framework. We find that a fact-checker and monitor pairing from different model families, for example a Claude 3.7 Sonnet monitor paired with a GPT-4.1 fact-checker, reduces approval of policy-violating actions by up to 45%, compared to only 6%, when using the same model for both fact-checking and monitoring roles. Our results demonstrate that CoT monitoring alone may be insufficient against adversarial persuasion, and that model-diverse fact-checking provides a robust mitigation.

13:00 JST研究/論文

PARA-PV: 凍結基礎モデルと分布シフト補正に基づく物理学を意識した検索拡張 PV 予測

正確な太陽光発電 (PV) 電力予測は、信頼性の高い送電網の送電と再生可能エネルギーの統合に不可欠ですが、PV 発電は気象変動、昼夜の変化、体制依存のダイナミクス、および厳格な物理的制約によって同時に形成されるため、依然として困難です。私たちは、予測プロセス全体に物理知識を組み込む、物理認識検索拡張フレームワークである PARA-PV を提案します。このフレームワークはまず、多変量 PV 観測をパッチレベル表現にエンコードし、物理学を意識した検索拡張学習器を通じて、時間的形状、電力レベル、PV 動作状態、日内期間の現在のウィンドウと一致する履歴パッチとアナログ軌跡を取得します。これにより、物理的に根拠のある基本予測が得られます。ローカル メモリをより広範な時間的知識で補完するために、基本予測は軽量残差アダプターを介して事前に凍結された Chronos 時系列基盤モデルに対して校正され、物理的に根拠のある予測を無効にすることなく、一般的な時間的規則性が PV 固有のダイナミクスに適応されます。天候や日周状況が変化しても残留条件付き分布シフトが持続するため、その後、物理学を意識した分布シフト補正モジュールが電力、天気、タイムスタンプ、昼夜条件を使用して事前予測を調整し、ゲート平均シフトとスケール補正を選択的に適用します。最後に、物理的に制約された損失関数は、サンプルをピーク、ランピング、夜間、および通常の領域に分割し、それらの誤差の寄与を適応的に再重み付けして、支配的な通常の領域が運用上重要な状態の学習を抑制するのを防ぎます。コードは https://github.com/weican1103/PARA-PV で入手できます。

原文 (English)

PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction

Accurate photovoltaic (PV) power forecasting is essential for reliable grid dispatch and renewable energy integration, yet it remains challenging because PV generation is jointly shaped by weather variability, day-night transitions, regime-dependent dynamics, and strict physical constraints. We propose PARA-PV, a Physics-Aware Retrieval-Augmented framework that embeds physical knowledge throughout the forecasting process. The framework first encodes multivariate PV observations into patch-level representations and, through a physics-aware retrieval-augmented learner, retrieves historical patches and analog trajectories that are consistent with the current window in temporal shape, power level, PV operating state, and intra-day period; this yields a physically grounded base forecast. To supplement local memory with broader temporal knowledge, the base forecast is then calibrated against a frozen Chronos time-series foundation-model prior through a lightweight residual adapter, so that general temporal regularities are adapted to PV-specific dynamics without overriding the physically grounded prediction. Because residual conditional distribution shifts persist when weather and diurnal regimes change, a physics-aware distribution shift correction module subsequently adjusts the preliminary forecast using power, weather, timestamp, and day/night conditions, applying gated mean-shift and scale corrections selectively. Finally, a physics-constrained loss function partitions the samples into peak, ramping, night-time, and regular regimes and adaptively reweights their error contributions, preventing the dominant regular regime from suppressing learning of operationally critical states. Our code is available at https://github.com/weican1103/PARA-PV.

13:00 JSTLLM/生成AIエージェント研究/論文

CausalDS: データ サイエンス エージェントにおける因果推論のベンチマーク

大規模言語モデル (LLM) は、抽象的な推論と高度なツールの使用を組み合わせて、統合されたデータ サイエンス エージェントとして機能することが増えています。しかし、関連するベンチマークの状況は、現実的なデータ分析を行わない記号的因果推論ベンチマークと、原則に基づいた因果関係データ生成構造を持たないデータ分析ベンチマークに大別されます。さらに、既存の因果関係評価データセットは既存のソースから厳選された例に限定されることが多く、多様性は新しい合成因果構造の系統的な生成ではなく、限られたテンプレート化されたバリエーションから得られます。エージェントのデータ サイエンス ワークフローにおける因果推論を評価するためのベンチマークである CausalDS を紹介します。各ベンチマーク インスタンスは、生成された観察データを含むサンプリングされた構造因果モデル (SCM) と、現実的な領域に基づいた付随する合成自然言語ストーリーで構成されるシーンです。オプションで、実世界のデータセットから得られた経験的分布でベンチマーク コンポーネントの構成を基礎にし、完全に合成生成することで「因果関係のオウム」リスクを軽減しながら経験的構造を維持します。次に、各シーンから、Pearl の 3 つのラングすべてにまたがるタスクを導き出します。代表的なデータ サイエンス予測タスクはラング 1 として表示されます。ほとんどのタスクにはデータ サイエンス コーディング コンポーネントが含まれており、観測モデルによって生成される不完全な観測が頻繁に存在するため、モデルは通常、最終的な答えに到達するために複数のツールを使用する必要があります。さらに、質問が正当な回答を認めていない場合を認識し、棄権することは、第一級の得点結果として扱われます。したがって、ベンチマークは、記号的因果推論、データ サイエンス、不確実性の定量化、棄権、およびツールの使用/コーディングを共同で評価します。

原文 (English)

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templatized variations rather than from systematic generation of novel synthetic causal structures. We introduce CausalDS, a benchmark for evaluating causal reasoning in agentic data-science workflows. Each benchmark instance is a scene consisting of a sampled structural causal model (SCM) with generated observational data and an accompanying synthetic natural-language story grounded in a realistic domain. We optionally ground the composition of the benchmark components in empirical distributions obtained from real-world datasets, thus retaining empirical structure while reducing the "causal parrot" risk through completely synthetic generation. From each scene, we then derive tasks spanning all three of Pearl's rungs, with typical data-science prediction tasks appearing as Rung 1. Most tasks include a data science coding component, where the model typically needs to use several tools to arrive at the final answer due to the frequent presence of imperfect observations, which are generated by an observation model. Additionally, recognizing when a question admits no warranted answer and abstaining is treated as a first-class scored outcome. The benchmark thus jointly evaluates symbolic causal reasoning, data science, uncertainty quantification, abstention, and tool use/coding.

13:00 JST研究/論文

解答セットのプログラミングに力を入れましょう! ASP およびエネルギーベースのモデルを使用したエンドツーエンドの神経記号推論と学習

我々は、エネルギーベースのモデル基板と答えセットプログラミングのモジュール統合に基づいた、一般的な神経象徴的推論と学習方法論を提示します。主な貢献は次のとおりです。(1) 背景知識、制約、非単調推論を完全に組み込んだ明示的な ASP ベースの宣言的セマンティクスを通じて、連続潜在空間での共同最適化をサポートします。 (2) 動的領域 (知覚やインタラクションなど) におけるアプリケーション向けの ASP 中心の堅牢なエンドツーエンド トレーニングのための一般化されたモデルと実用的なプラットフォームを提供することにより、解答セット、確率論理、および解答セット モジュロ理論のインターフェイスにおける最近の研究を前進させます。実際の実装を提供し、基本的な使用法と応用例 (MNIST を使用) を示し、視覚的な質問応答ベンチマーク Clvr とマルチオブジェクト追跡ベンチマーク MOT で評価します。

原文 (English)

Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models

We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model substrate. Key contributions are: (1) supporting joint optimisation in the continuous latent space through explicit ASP-based declarative semantics fully incorporating background knowledge, constraints, non-monotonic inference; and (2) advancing recent works at the interface of answer sets, probabilistic logic, and answer set modulo theories by providing a generalised model and practical platform for ASP-centric robust, end-to-end training for applications in dynamic domains (e.g., involving perception and interaction). We provide a practical implementation, and demonstrate basic use and application (with MNIST), and evaluate with the visual question-answering benchmark Clevr and the multi-object tracking benchmark MOT.

13:00 JST研究/論文

考えすぎ: 学習した秘密を抽出するために推論の重みを増幅する

言語モデルのブラック ボックス監査は必須の展開前ツールですが、微妙な不整合や隠された情報を見逃す可能性があります。監査プロセス中に隠された情報をより適切に引き出すために、 \emph{over Thinking} を導入します。これは、推論タスクベクトルを使用して、推論モデルを大声で考える傾向を増幅するプロセスです。非推論命令モデル $M$ と推論蒸留モデル $R$ のパラメータを与えると、 \emph{過剰思考モデル} を $\boldsymbol{\theta}_{\mathcal{O}_\alpha} = \boldsymbol{\theta}_{\mathcal{M}} + \alpha(\boldsymbol{\theta}_{\mathcal{R}} - として定義します) \boldsymbol{\theta}_{\mathcal{M}})$、ここで $\alpha > 1$ は純粋な推論モデル $R$ を超えて推論を増幅します。さらに、モデル出力の品質と一貫性を失うことなく推論を選択的に増幅する新しい層ごとの減衰戦略を導入します。私たちは、モデルを考えすぎると、2B ~ 32B モデルにわたる 4 つの実験設定にわたって隠された情報が明らかになる可能性が高いことを示します。私たちの調査結果は、推論の増幅によって、元の推論モデルよりも最大 $10\times$ の頻度で、トレーニング中に獲得された秘密や意図しない行動が表面化する可能性があることを示唆しています。秘密がどのように表面化するかは、秘密のタイプによって異なります。推論方向に沿った摂動を必要とするものもあれば、十分に大きな重みの摂動に屈するものもあります。

原文 (English)

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overthinking}: the process of using reasoning task vectors to amplify the propensity to think out loud of reasoning models. Given the parameters of a non-reasoning instruct model $M$ and reasoning-distilled model $R$, we define the \emph{overthinking model} as $\boldsymbol{\theta}_{\mathcal{O}_\alpha} = \boldsymbol{\theta}_{\mathcal{M}} + \alpha(\boldsymbol{\theta}_{\mathcal{R}} - \boldsymbol{\theta}_{\mathcal{M}})$, where $\alpha > 1$ amplifies reasoning beyond the pure reasoning model $R$. Additionally, we introduce new layer-wise attenuation strategies that selectively amplify reasoning without losing quality and coherence of model outputs. We demonstrate that overthinking models are more likely to reveal hidden information across four experimental settings, across 2B-32B models. Our findings suggest that reasoning amplification may surface secrets or unintended behaviors acquired during training up to $10\times$ more frequently than the original reasoning model. How secrets surface depends on the secret type: some require perturbation along the reasoning direction, while others yield to any sufficiently large weight perturbation.

13:00 JSTエージェント研究/論文

ASMR: 船舶メンテナンスレポート作成のためのエージェントスキーマ生成

このペーパーでは、スキーマの自動生成の問題について研究します。つまり、複数のフォーム カテゴリにわたる過去の船舶メンテナンスおよび運航レポートのコレクションが与えられると、各レポート タイプの重要な情報要件を捕捉するコンパクトで有益なスキーマを自動的に発見します。この課題に対処するために、2 つの特殊なエージェントで構成されるモジュール式エージェント フレームワークである ASMR を提案します。フィールド生成エージェントは、歴史の物語から意味論的な概念を抽出し、適応型多粒度クラスタリングを通じて候補スキーマ フィールドを生成します。一方、構造最適化エージェントは、強化学習を採用して、コンパクトで有益な、冗長性のないスキーマ表現を特定します。結果として得られるスキーマは、レポート作成者がより完全で一貫性のある実用的なレポートを作成するように導くことができます。予備的な結果は、提案されたアプローチの有望性を実証し、データ管理、エージェントAI、人間中心AIの交差点におけるいくつかの未解決の研究課題を浮き彫りにしています。

原文 (English)

ASMR: Agentic Schema Generation for Ship Maintenance Report Writing

In this paper, we study the automatic schema generation problem: given a collection of historical ship maintenance and operational reports across multiple form categories, automatically discover compact and informative schemas that capture the essential information requirements of each report type. To address this challenge, we propose ASMR, a modular agentic framework consisting of two specialized agents. A Field Generation Agent extracts semantic concepts from historical narratives and generates candidate schema fields through adaptive multi-granularity clustering, while a Structural Optimizer Agent employs reinforcement learning to identify compact, informative, and non-redundant schema representations. The resulting schemas can guide report authors toward producing more complete, consistent, and actionable reports. Preliminary results demonstrate the promise of the proposed approach and highlight several open research challenges at the intersection of data management, agentic AI, and human-centered AI.

13:00 JSTLLM/生成AI研究/論文

遅い思考と能動的な知覚の第一原理理論

認知機能の第一原理モデリングに関するシリーズの一部として、この論文は思考と知覚の数学的定式化を提供することを試みます。これは形式的には遅い思考、またはより一般的には能動的な知覚を導き出し、遅い思考の大規模言語モデルの設計、トレーニング、推論を包含します。私たちの出発点は、ニューラル ネットワークなどの単純な関数ファミリーによって複雑なデータ分布を表現することを目的として、観測可能空間および潜在空間上の確率分布のリフティングと投影です。 「アクティブリフティング」と呼ばれる理論は、潜在シーケンスのサンプリングと最大レートで不確実性を低減するための固有の駆動力に基づいて提案されています。それは、静的理論と呼ばれる部分空間に遅い思考モデルを含む大きな設計空間を導き出します。これらのモデルは、静的理論によって誘導される表現階層とサンプラー階層に位置し、2 つの階層を登ることによってアップグレードできます。アクティブリフティングはさらに、内部時間軸を備えた推論プロセスと、言語の発明だけでなく最小長のコーディングに似たトレーニング目標も導き出します。したがって、それは、遅い思考形式の出現を含む、知覚の主体性を特徴づけます。この理論の技術的な副産物には、遅い思考モデルを改善するための 3 段階の経路、すべてのデータ モダリティのエンコーダーと生成モデルを構築するための統一されたアプローチ、人間のような視覚表現のアプリオリな形成、および政策崩壊に対する可能な解決策が含まれます。

原文 (English)

A First-Principles Theory of Slow Thinking and Active Perception

As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of thinking and perception. It formally derives slow thinking or more generally, active perception, and encompasses the design, training and inference of slow thinking large language models. Our starting point is the lifting and projection of probability distributions on the observable and latent spaces, with the objective of representing complex data distributions by simple function families such as neural networks. A theory called "active lifting" is proposed, based on the sampling of latent sequences and an intrinsic drive to reduce uncertainty with maximum rate. It derives a large design space, containing the slow thinking models in a subspace that we call the static theory. These models are positioned on the representation hierarchy and sampler hierarchy induced by the static theory, and can be upgraded by climbing the two hierarchies. Active lifting further derives an inference process with an internal time axis, and a training objective that resembles minimum-length coding as well as the invention of languages. Thus, it characterizes the agency of perception, including the emergence of the slow thinking formats. Technical by-products of this theory include a three-stage pathway for improving slow thinking models, a unified approach to constructing encoders and generative models for all data modalities, a priori formation of human-like visual representations, and a possible solution to policy collapse.

13:00 JST画像/動画生成エージェント

ZendoWorld をプレイ: アクティブなビジュアル コンセプト導入で AI エージェントに挑戦

インテリジェント システムを構築する際の中心的な課題は、エージェントが複雑な入力を共同で認識し、隠れたパターンについての仮説を立て、それらをテストするための有益な実験を計画できるようにすることです。この問題を研究するために、我々は、エージェントが視覚的なゲーム観察に関する論理規則を推測し、新しいシーンを提案することによって情報を取得し、ゲーム環境からのフィードバックに基づいて仮説を洗練しなければならない制御された対話型環境である ZendoWorld を提案します。純粋な VLM 推論、ベイジアン粒子フィルタリング、動的概念発見、および神経記号的手法にわたるいくつかのエージェントを評価します。私たちの主な発見は次のとおりです。(1) 観察された例のラベルを予測する精度が高いことは、基礎となるルールの回復を意味するものではありません。 (2) 知覚と誘導は、さまざまなエージェント クラスにとって明確なボトルネックです。 (3) VLM ベースのエージェントはほとんど有益でない実験を提案し、仮説の不確実性を積極的に低減できません。これらの結果を比較するために、このタスクに関する人間のデータを収集しました。これにより、特により複雑なルールの場合、帰納的推論におけるギャップが明らかになります。全体として、ZENDOWORLD はインテリジェント エージェントの評価に向けて重要な一歩を踏み出し、特に科学的発見などの分野で改善の具体的な道筋を特定しています。

原文 (English)

Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction

A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden patterns, and design informative experiments to test them. To study this problem, we propose ZendoWorld, a controlled interactive environment in which agents must infer a logical rule about visual game observations, acquire information by proposing new scenes, and refine their hypotheses based on feedback from the game environment. We evaluate several agents spanning pure VLM reasoning, Bayesian particle filtering, dynamic concept discovery, and neuro-symbolic methods. Our main findings are: (1) high accuracy in predicting labels for observed examples does not imply recovery of the underlying rule; (2) perception and induction are distinct bottlenecks for different agent classes; and (3) VLM-based agents propose near-uninformative experiments, failing to actively reduce hypothesis uncertainty. To compare these results, we collect human data on the task, which reveals a gap in inductive reasoning, particularly for more complex rules. Overall, ZENDOWORLD takes an important step toward evaluating intelligent agents and identifies concrete avenues for improvement, particularly in domains like scientific discovery.

13:00 JSTLLM/生成AIエージェント

Autoペルソナ: オープンエンドのペルソナ進化のためのマルチタイムスケール ループ エンジン

長期にわたるペルソナエージェントは、新しい出来事、人間関係、証拠、社会的状況に適応しながらも、識別可能性を維持する必要があります。私たちは、自己ロックを、継続的なペルソナと人生のループにおける実行時の失敗モードとして特定します。つまり、生成された人生が、慣れ親しんだ環境、弱い人間関係、保留された決定、古いライフステージに向かって崩壊する一方で、局所的にもっともらしいイベントが出現し続けます。私たちは、この失敗を、状態、メモリ、履歴、および環境の概要からの高確率の行動チャネルとシステムレベルのコンテキスト重力へのモデルレベルの収束に追跡します。境界のあるペルソナレベルの再帰的自己進化のためのマルチタイムスケールの生活環境エンジンである Autopersonas を紹介します。環境側の発生、蓄積された観察、およびペルソナの状態を分離します。その OSO ループは、状態や到達可能性が変化する前に、証拠に基づいた吸収を要求しながら、将来に向けた多様なマテリアルを許容します。 3 年間の圧縮シミュレーションにより、環境のウォーターマーク シェル、発生の強化ギャップ、緩やかな変化の蓄積の失敗、再帰的な優柔不断、弱い関係の持続性が明らかになりました。 8 モデルの 40 日間のストレス テストでは 1,600 のイベントが生成され、ローリング 5 日間のアクション カテゴリの平均反復率は 95.2% ~ 97.6% であり、すべてのモデルが 11 日目までに 90% を超えました。セマンティック再保持により、すべての直接ループ実行でマクロ テーマの反復が 79.0% ~ 88.0% であることがわかりました。同一実行時の 40 日間の A/B では、コンテキスト スライス マスキングとサンプルごとの発散ターゲティングにより、マクロ テーマの繰り返しが 61.8% から 36.3% に減少し、累積テーマ数が約 2 倍になりました。少年とゴブリンの架空の世界の実行は、現実世界のハードな侵入なしで反固定主義体制を再現しました。これらの結果は、制御された発散を証拠に基づいた吸収から分離することで、アイデンティティの連続性を維持しながら、ペルソナと環境の自己ロックを軽減できるという限定された主張を裏付けています。

原文 (English)

AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution

Long-term persona agents must remain identifiable while adapting to new events, relationships, evidence, and social conditions. We identify self-locking as a runtime failure mode in continuing persona-life loops: locally plausible events keep appearing while the generated life collapses toward familiar environments, weak relationships, suspended decisions, and stale life stages. We trace this failure to model-level convergence toward high-probability behavioral channels and system-level context gravity from State, memory, history, and environment summaries. We introduce AutoPersonas, a multi-timescale life-environment engine for bounded persona-level recursive self-evolution. It separates environment-side Occurrences, accumulated Observations, and persona State. Its OSO loop admits divergent future-facing material while requiring evidence-governed absorption before State or reachability changes. A three-year compressed simulation exposed environment watermark shells, occurrence-hardening gaps, slow-change accumulation failures, recursive indecision, and weak relationship persistence. An eight-model 40-day stress test generated 1,600 events and found mean rolling 5-day action-category repetition of 95.2%-97.6%, with all models crossing 90% by day 11. Semantic re-keeping found 79.0%-88.0% macro-theme repetition across all direct-loop runs. In a same-runtime 40-day A/B, context-slice masking plus per-sample divergence targeting reduced macro-theme repetition from 61.8% to 36.3% and roughly doubled cumulative theme count. A juvenile-goblin fictional-world run reproduced the anti-fixation regime without hard real-world intrusions. These results support a bounded claim: separating controlled divergence from evidence-governed absorption can reduce persona-environment self-locking while preserving identity continuity.

13:00 JST研究/論文ClaudeGPT / ChatGPTGeminiNVIDIAGrok

競争してから協力する: フロンティア AI 教師が検証可能なカリキュラムを構築し、模倣を超えてコーディング学生を向上させる

大規模な言語モデルは、小規模な生徒向けのトレーニング データを生成する教師としての役割をますます高めています。従来の複数教師による知識の蒸留方法は、どのフロンティア モデルが最もよく教えるかを決定せずに出力をマージし、多くの場合、独自の出力に偏った LLM 判定に依存していました。私たちは、4 人のフロンティア AI 教師 (Claude、Codex-GPT、Grok、Gemini) が、公平性管理を備えた実行ベースの審査員 (単体テストと標準入力/標準出力チェック) によって直接順位付けされ、その後、協力して生徒 (Qwen2.5-Coder) のための検証可能なカリキュラムを構築する、競争してから協力するフレームワークを導入します。 3 つの調査結果を報告します。 (1) 実行検証では、飽和効果により、自己修正後 (99 ~ 100%)、すべての教師が標準問題をほぼ完璧に解きますが、より難しい競争問題では差が生じます (Gemini 77% > Claude 69% = Codex 69% > Grok 50%)。ただし、生徒側の堅調な結果は教師のランキングに依存しません。 (2) 検証済みの解決策を模倣 (SFT) しても、7B および 32B ですでに有能な生徒は向上せず、低下する可能性があります (たとえば、MBPP テストでは 76.7% から 72.7%、競争問題では 5.9% から 2.9%)。 (3) 同じ協調カリキュラムを検証可能な報酬付き強化学習 (RLVR) 環境として使用すると、生徒の成績が向上し (競争問題でのピークが 5.9% から 8.8% に増加し、+49% の相対利益)、SFT の方向性が逆転します。 AI と教師のコラボレーションの価値は、模倣するための答えをプールすることではなく、生徒が実践しながら学ぶ検証可能な環境を共同で構築することにあります。最新のスタックで GRPO を実行するためのフレームワーク パッチを含む、再現可能なオンプレミス パイプライン (NVIDIA GB10) をリリースします。

原文 (English)

Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation

Large language models increasingly serve as teachers generating training data for smaller students. Prior multi-teacher knowledge distillation methods merge outputs without determining which frontier model teaches best, often relying on an LLM judge biased toward its own outputs. We introduce a compete-then-collaborate framework where four frontier AI teachers (Claude, Codex-GPT, Grok, Gemini) are ranked head-to-head by an execution-based judge (unit tests and stdin-stdout checks) with fairness controls, and then collaborate to build a verifiable curriculum for a student (Qwen2.5-Coder). We report three findings. (1) Under execution verification, all teachers solve standard problems near-perfectly after self-correction (99-100%) due to a saturation effect, but harder competition problems separate them (Gemini 77% > Claude 69% = Codex 69% > Grok 50%); however, the robust student-side results do not depend on teacher ranking. (2) Imitation (SFT) on verified solutions does not improve, and can degrade, an already-competent student at 7B and 32B (e.g., from 76.7% to 72.7% on MBPP-test, and 5.9% to 2.9% on competition problems). (3) Using the same collaborative curriculum as a reinforcement learning with verifiable rewards (RLVR) environment improves the student (from 5.9% to 8.8% peak on competition problems, a +49% relative gain), reversing SFT's direction. The value of AI-teacher collaboration lies not in pooling answers to imitate, but in jointly constructing a verifiable environment where the student learns by doing. We release a reproducible on-prem pipeline (NVIDIA GB10) with framework patches for running GRPO on a bleeding-edge stack.

13:00 JSTLLM/生成AI

MentalHospital: 精神科の臨床経験を評価するための仮想環境

大規模言語モデル (LLM) は、対話、診断、治療計画などの個別の精神医学的タスクで優れたパフォーマンスを示していますが、既存のベンチマークが完全な精神医学的臨床遭遇をシミュレートすることはほとんどありません。私たちは、LLM ベースの精神科臨床遭遇のための仮想評価環境 $\textbf{MentalHospital}$ を紹介します。 MentalHospital は、すべての主要な ICD-11 カテゴリと 76 の疾患にわたる 1,193 件の匿名化された精神科電子健康記録 (EHR) 症例から構築されたスキル強化された標準化された患者を使用して、主観的面接、客観的検査、診断評価、および治療計画 (S.O.A.P.) ワークフローをインスタンス化します。各症例は、EHR 由来の参照との客観的な比較と臨床プロセスの質の主観的な評価を組み合わせたデュアルトラック プロトコルを通じて評価されます。専門家の判断をスケールするために、私たちは $\textbf{MentalEval}$ を開発します。$\textbf{MentalEval}$ は、コミュニケーションへの共感、面接のプロ意識、臨床記録の質、診断の厳密さ、治療の適切性をカバーする 5 人の領域固有の評価者であり、ルーブリックに基づいた SFT と専門家の指導による DPO で訓練を受けています。 22 人の臨床医からのアンケート回答は、MentalHospital の臨床忠実度 (3.88/5) を裏付けており、MentalEval は平均 QWK 0.944 で専門家との強力な連携を実現しています。ベンチマークによると、最も強力な LLM でさえ、精神状態の評価が主要なボトルネックとなり、客観的な精神医学的能力において臨床医に 37.28 パーセント ポイント及ばないことが示されています。

原文 (English)

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop $\textbf{MentalEval}$, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.

13:00 JSTLLM/生成AIDeepSeek

さまざまな教師、さまざまな機能: 構造化テキスト強化のためのサブ 1B オンデバイス蒸留

大量の構造化抽出では、すべての項目で大規模なモデルのレイテンシーが発生するため、タスクを小さなオンデバイス モデルに抽出すると、数分の 1 の時間とコストで同等の出力が得られるため、魅力的です。サブタスクごとに、蒸留によって実際に何が得られるかを測定します。各ニュース記事は、短い概要と 5 つのカテゴリ ラベルを持つ 1 つの JSON オブジェクトにマップされます。 80 億の推論教師 (deepseek-r1:8b) を 0.60 億の生徒 (Qwen3-0.6B、QLoRA、3 つのシード) に抽出し、2 つの教師コントロール (同じサイズの非推論教師とより大きなマネージド パイプライン) を追加します。盲検で参考文献のない 3 人の審査員パネルが、2 つの非蒸留ベースライン、少数ショットのプロンプトおよび制約された解読に加えて、論文全文に対してすべての腕を採点します。生徒は教師の 39 秒に対して記事あたり約 0.8 秒で実行し、概要の品質に関してベースと教師のギャップの 58% を回復し、主ベースライン (制約付きデコード) を +16.8 ポイント上回り、少数ショット プロンプトは二次ベースライン +4.9 ポイントを上回りました。同じサイズの非推論教師は、調整されていないベースと同じように生徒を訓練するため、要約利得は、その規模ではなく教師の推論の性質から導き出されます。次に、能力は教師によって分割されます。推論系の教師は文章の質を伝達し、管理されたパイプラインはラベルの多様性を伝達します。一方、同じ規模の指導教師の生徒は、推論系統の生徒が捏造する 93 項目のテスト セット (忠実な 74 項目と 55 項目) 内の 22 の短くて情報源の薄い記事に基づいてより安定した状態を保ちます。この根拠となる違いは、有意な集合効果ではなく、一貫した順序付けであり、サブグループが小さいため、方向性として報告します。すべてのフィールドで単一のエンジンが勝てるわけではないため、成果物はオンデバイス エンリッチメントのためのフィールドごとのルーティング マップになります。

原文 (English)

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

High-volume structured extraction pays a large model's latency on every item, so distilling the task into a small on-device model is attractive: comparable output at a fraction of the time and cost. We measure what that distillation actually delivers, per sub-task. Each news article is mapped to one JSON object with a short summary and five categorical labels. We distill an 8B reasoning teacher (deepseek-r1:8b) into a 0.6B student (Qwen3-0.6B; QLoRA, three seeds), and add two teacher controls: a same-size non-reasoning teacher and a larger managed pipeline. A blinded, reference-free, three-judge panel scores every arm against the full article, alongside two non-distillation baselines, few-shot prompting and constrained decoding. The student runs at about 0.8 s per article against the teacher's 39 s, and recovers 58% of the base-to-teacher gap on summary quality, beating its primary baseline (constrained decoding) by +16.8 points and few-shot prompting by a secondary +4.9. A same-size non-reasoning teacher trains a student no better than the untuned base, so the summary gain follows from the teacher's reasoning nature rather than its scale. Capabilities then split by teacher: the reasoning teacher transfers writing quality and the managed pipeline transfers label diversity, while a same-size instruction teacher's students stay more grounded on the 22 short, thin-source articles in the 93-item test set (74 versus 55 faithful), where the reasoning-lineage student fabricates. That grounding difference is a consistent ordering rather than a significant aggregate effect, and the subgroup is small, so we report it as a direction. Because no single engine wins every field, the deliverable is a per-field routing map for on-device enrichment.

13:00 JST研究/論文

PolyUQuest: 異種グラフ上の検証可能な構造認識型 Web RAG

既存の検索拡張生成 (RAG) システムは、Web ページをフラット テキストとして扱い、HTML にエンコードされた構造的および意味論的な信号を失います。 PolyUQuest は、ページ間のハイパーリンク トポロジ、ページ内の DOM 階層、ページ間のエンティティ関係の知識を統合する、異種グラフ上に構築された検証可能な構造認識 Web RAG フレームワークです。 2 層ルーターは、直接ブロック取得、クロスページ グラフ トラバーサル、およびマルチホップ エンティティ推論を含む、構造上のニーズに一致する 3 つの取得モードのいずれかに各クエリをディスパッチします。引用された各ブロックにはソース ページ、見出しパス、エンティティ リンクが含まれているため、すべての回答は完全に検証可能であり、ユーザーはあらゆる主張をその構造的証拠にまで遡ることができます。私たちは、4,240 ページ、31,086 の DOM ブロック、29,119 のエンティティ、および 37,680 の関係で構成される香港理工大学 (PolyU) の公式 Web サイトを、複数タイプの評価ベンチマークとともに評価します。 PolyUQuest は、回答の正確性、カバレッジ、忠実性において既存の RAG システムよりも優れており、クエリごとに消費する LLM トークンの量が大幅に少なくなります。このデモでは、引用された回答を検査し、ルーティング モード間で検索トレースを比較し、証拠グラフ パスを探索するための対話型インターフェイスが提供されます。 PolyUQuest は、PolyU で学生向け QA サービスとして導入の準備が進められています。

原文 (English)

PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs

Existing retrieval-augmented generation (RAG) systems treat web pages as flat text, losing the structural and semantic signals encoded in HTML. We present PolyUQuest, a verifiable, structure-aware web RAG framework built on a heterogeneous graph that unifies hyperlink topology between pages, DOM hierarchy within pages, and entity-relation knowledge across pages. A two-tier router dispatches each query to one of three retrieval modes matched to its structural need, including direct block retrieval, cross-page graph traversal, and multi-hop entity reasoning. Every answer is fully verifiable, as each cited block carries its source page, heading path, and entity links so that users can trace any claim back to its structural evidence. We evaluate on the official websites of the Hong Kong Polytechnic University (PolyU), comprising 4,240 pages, 31,086 DOM blocks, 29,119 entities, and 37,680 relations, together with a multi-type evaluation benchmark. PolyUQuest outperforms existing RAG systems in answer correctness, coverage, and faithfulness, while consuming significantly fewer LLM tokens per query. The demonstration provides an interactive interface for inspecting cited answers, comparing retrieval traces across routing modes, and exploring evidence graph paths. PolyUQuest is being prepared for deployment as a student-facing QA service at PolyU.

13:00 JSTLLM/生成AI研究/論文

PredicateLongBench による長いコンテキスト タスクの難易度の軸の理解

大規模言語モデル (LLM) はロングコンテキスト機能の急速な向上を実証しており、それらを評価するために設計されたベンチマークの波が押し寄せています。ただし、既存のロングコンテキスト評価 (Needle-in-a-Haystack (NIAH) テストから最近のマルチホップ推論および要約タスクに至るまで) は主に平均的なケースのパフォーマンスを測定しており、その多くは飽和しているか、堅牢性に欠けています。さまざまな軸に沿ってタスクの難易度をスケールアップするときにモデルがどのように実行されるかを調査するための体系的な方法が特に欠落しています。私たちは、より広範な述語クラスから抽出された、与えられた述語/制約 (辞書編集的な順序付けなど) を満たす長い入力内の単語の最長の連続部分列を識別するようモデルに要求することで、長い文脈の推論をストレステストするベンチマークである PredicateLongBench を提案することで、このギャップに対処します。私たちのベンチマークの中心的な革新は、長い文脈理解の複数の側面をテストする複数の異なる難易度軸の特定と体系的な探索です。私たちは 2 つの相補的な生成パイプラインを提供します。1 つはランダムな単語のような文字列を使用する完全合成セットアップ、もう 1 つは分布特性を維持しながら自然文書から単語をサンプリングする現実世界のセットアップです。軸に沿ってタスクの難易度を上げていくと、フロンティア モデルのパフォーマンスが低下することがわかり、現在のロングコンテキスト機能の限界を理解する上でのベンチマークの有用性が実証されました。さらに、PredicateLongBench のタスクは、困難ではありますが、概念的には単純であり、LLM ベースの生成やジャッジを必要としません。

原文 (English)

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness. Notably absent is a systematic way to probe how models perform as we scale up the difficulty of tasks along various axes. We address this gap by proposing PredicateLongBench, a benchmark that stress-tests long-context reasoning by asking models to identify the longest contiguous subsequence of words in a long input that satisfies given predicates/constraints (e.g., lexicographic ordering), drawn from a broader predicate class. The central innovation of our benchmark is the identification and systematic exploration of multiple different axes of difficulty which test multiple aspects of long context understanding. We provide two complementary generation pipelines - a fully synthetic setup using random word-like strings, and a real-world setup that samples words from natural documents while preserving their distributional properties. We find that frontier models struggle to perform well as we scale up the difficulty of tasks along our axes, demonstrating the utility of our benchmark in understanding the limitations of current long-context capabilities. Furthermore, the tasks in PredicateLongBench, though challenging, are conceptually simple and do not require LLM-based generations or judges.

13:00 JSTビジネス/資金調達

AI評価に欠けている要素としての心理的能力

現在の AI 評価フレームワークは、精度、堅牢性、推論能力、ポリシー遵守などの技術的パフォーマンスに主に焦点を当てています。これらの対策は依然として不可欠ですが、自然言語を通じてユーザーと直接対話するシステムにとっては十分ではありません。人間と向き合う AI システムは、アドバイザー、コーチ、家庭教師、仲間として使用されることが増えています。これらの役割では、ユーザーの応答によって、ユーザーがどのように推論し、感情を解釈し、信念を形成し、信頼を調整し、意思決定を行うかが形成されます。したがって、関連する評価単位はモデルだけではなく、人間と AI の相互作用です。この論文では、AI 評価に欠けている側面として心理的能力を紹介します。私たちは、心理的能力を、ユーザー、状況、インタラクションの目的に適切な方法でユーザーの認知、感情の解釈、行動の意思決定をサポートする対人 AI システムの能力として定義します。これには、フレーミング、トーン、知覚される権威、反応性、不確実性の処理、会話のガイダンスなどの対話特性が含まれます。既存の評価アプローチはこの問題の一部を捉えていますが、これらの心理的影響を直接評価することはほとんどありません。行動科学と人間と AI の相互作用研究に基づいて、心理的能力とその中核領域の概念的枠組みを概説します。特定のベンチマークを提案するのではなく、構成を定義し、その境界を明確にし、シナリオベースの調査、構造化された人間による評価、およびモデル支援の評価方法を通じてそれがどのように評価されるかを説明します。私たちは、心理的能力が、人間と対面する AI システムの実世界への影響を懸念するモデル提供者、導入組織、研究者、規制当局にとって中心的な考慮事項となるべきであると主張します。

原文 (English)

Psychological Competence as a Missing Dimension in AI Evaluation

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

13:00 JSTエージェントロボティクス

INTENT: 包括的なアブレーション分析を使用した交差点シナリオでの車両の意図を予測するための LSTM フレームワーク

車両の意図予測は、あらゆる運転シナリオにおける自動運転車の機敏性と安全性において極めて重要な要素です。自動運転車の真の強化が必要な場合、特に人間との対話が多く必要な場合や、交差点、環状交差点、急停止などの緊急事態などの複雑な運転行動だけでなく、多くの人間の対話が必要な場合に、車両の意図の予測がリアルタイムで適切な回避行動をとるのに役立ち、毎秒の行動が影響を及ぼし、大惨事の発生を防ぐことができる場合に、ドライバーの意図について人間による解釈を自動運転車に採用させる必要があります。最悪の場合でも被害を最小限に抑え、安全を優先することができます。意図予測は、軌道予測を強化するために使用することもできます (意図条件付き軌道予測)。本研究では、LSTMモデルを用いてイベント発生の2秒前に交差点での車両の意図を予測し、交差点内の車両が直進するのか、左折するのか、右折するのかを予測するINTENTフレームワークを提案する。さまざまなモデル実験とアブレーション研究は、99.71% の精度を達成する InD データセットで徹底的にテストされています。

原文 (English)

INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis

Vehicle intention prediction is a pivotal aspect in the agility and safety of autonomous vehicles in all driving scenarios; if genuine enhancement of autonomous vehicles are required, we need to make them adopt human interpretation of driver's intention especially in cases that require a lot of human interaction as well as complex driving behaviors like the ones at intersections, roundabouts and emergency cases such as sudden stops where vehicle intention prediction helps in taking the correct evasive action within a real time period where every second of action makes an impact and can prevent a catastrophe from taking place. In the worst case, it helps minimize the damage and make safety a priority. Intention prediction can also be used to enhance trajectory prediction (intention conditioned trajectory prediction). In this study, The INTENT framework is proposed using LSTM model to predict the vehicle's intention at intersections 2 seconds ahead of the event occurrence to predict whether the cars in intersections are going straight, turning left, or turning right. Various model experiments and ablation study are thoroughly tested on InD dataset achieving 99.71% accuracy.

13:00 JST研究/論文

Blind-Spots-Bench: マルチモーダル モデルの死角の評価

最新の AI モデルは、多くの確立されたベンチマークで優れたパフォーマンスを達成していますが、糸を操作したり、5 本足の犬を描いたりするなど、人間にとってはほとんど些細なタスクであると依然として失敗しています。これらの例は、既存のベンチマークでは現在のシステムの永続的な盲点が十分に測定されていない可能性があることを示唆しています。 $\texttt{blind-spots-bench}$ を紹介します。これは、人間にとっては単純に見えても、現代の AI にとっては依然として困難なタスクを通じて、そのような盲点を明らかにするように設計されたベンチマークです。 AI コースで学生から生の質問を収集し、構造化された参照ソリューションでそれらを整理して注釈を付け、結果として得られる 235 サンプルのデータセットに合わせたタスク分類を提案します。さらに、オープンウェイトおよびクローズドソース言語、ビジョン言語、画像生成モデルなど、幅広いモデルを評価するための自動グレーディング パイプラインを開発します。 $\texttt{blind-spots-bench}$ に関する分析により、クローズドソース フロンティア モデルは、既存のベンチマークで同等のパフォーマンスを達成した場合でも、$\およそ 10\%$ の差があってもオープンウェイト モデルを大幅に上回るパフォーマンスを発揮できることが明らかになりました。より詳細な分析では、すべてのタスク タイプにわたって優勢な単一のモデルは存在せず、一部のタスクは評価されたすべてのモデルにとって依然として困難であることが示されています。これらの結果は、現在の最新モデルの具体的な弱点を特定するための診断ストレス テストとしての $\texttt{blind-spots-bench}$ の価値を強調しています。

原文 (English)

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce $\texttt{blind-spots-bench}$, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on $\texttt{blind-spots-bench}$ reveals that closed-source frontier models can substantially outperform open-weight models with even $\approx10\%$ gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of $\texttt{blind-spots-bench}$ as a diagnostic stress test for identifying concrete weaknesses in current modern models.

13:00 JST研究/論文

MobiDiff: ヒューマン モビリティ データ生成のためのセマンティックを意識したマルチチャネル離散拡散

人間の移動データは交通の最適化、都市計画、資源配分に不可欠ですが、現実世界の移動データは収集にコストがかかり、プライバシー上の懸念から共有するのが困難です。最近の拡散ベースの手法は、現実的な移動パターンの合成に有望であることが示されていますが、通常、連続的または潜在的な時空間トレースに依存しているため、明示的な領域、アクティビティ、時間、および間隔の構造を備えた離散的な意味論的イベントをネイティブにモデル化する能力が制限されています。この問題に対処するために、エンドツーエンドの離散拡散フレームワークである MobiDiff を導入します。これは、マルチチャネルのセマンティック スケルトンを直接ノイズ除去し、コストのかかる補間、潜在トレース構築、および既存の拡散ベースの手法で広く使用されている粗いから細かいまでの実現パイプラインを回避することで、モビリティ データを効率的に生成します。具体的には、MobiDiff は人間の各チェックイン イベントを空間、アクティビティ、および時間チャネルに分解し、構造化されたイベント、グループ、およびチャネル レベルのマスキングを使用して、軌跡レベルのモビリティ パターンとイベント内の依存関係を共同でキャプチャします。アトランタ、ボストン、シアトルの 3 つの大規模な現実世界のデータセットで、生成の忠実度、プライバシーの保護、効率を評価します。結果は、MobiDiff がより広範なモビリティ統計にわたって競争力を維持しながら、軌道の長さと時間間隔の分布を効果的に保存していることを示しています。また、最先端の方法よりもはるかに高速です。たとえば、推論中に平均して GeoGen よりも 5.3$\倍$ 高速です。これらの発見は、離散拡散が合成モビリティ データ生成のための解釈可能で効率的なフレームワークを提供することを示唆しています。

原文 (English)

MobiDiff: Semantic-Aware Multi-Channel Discrete Diffusion for Human Mobility Data Generation

Human mobility data are essential for transportation optimization, urban planning, and resource allocation, yet real-world mobility data are costly to collect and difficult to share due to privacy concerns. Recent diffusion-based methods have shown promise in synthesizing realistic mobility patterns, but they typically rely on continuous or latent spatio-temporal traces, limiting their ability to natively model discrete semantic events with explicit region, activity, time, and interval structures. To address this issue, we introduce MobiDiff, an end-to-end discrete diffusion framework that efficiently generates mobility data by directly denoising multi-channel semantic skeletons, avoiding the costly interpolation, latent trace construction, and coarse-to-fine realization pipelines widely used in existing diffusion-based methods. Specifically, MobiDiff decomposes each human check-in event into spatial, activity, and temporal channels, and employs structured event-, group-, and channel-level masking to jointly capture trajectory-level mobility patterns and within-event dependencies. We evaluate generation fidelity, privacy-preserving, and efficiency on three large-scale real-world datasets from Atlanta, Boston, and Seattle. Results show that MobiDiff effectively preserves trajectory length and temporal interval distributions while remaining competitive across broader mobility statistics; it is also much faster than state-of-the-art methods, e.g., 5.3$\times$ faster than GeoGen on average during inference. These findings suggest that discrete diffusion offers an interpretable and efficient framework for synthetic mobility data generation.

13:00 JSTLLM/生成AIハードウェア/半導体

FedOPAL: 分析ビジュアル プロンプト チューニングによるワンショットのフェデレーテッド ラーニング

エッジ インテリジェンスにおける基本モデルの広範な展開に伴い、通信帯域幅がフェデレーテッド ラーニングのスケーラビリティを制限する中心的なボトルネックになっています。ワンショットのフェデレーテッド ラーニングは通信ラウンドを最小限に抑えることでこの問題を軽減しますが、既存の反復的な微調整や知識の蒸留方法では、サーバー側の高い計算コストやハイパーパラメータの感度などの課題に依然として直面しています。分析的フェデレーテッド ラーニングは、最小二乗閉形式解を使用して効率的な勾配のない集計を実現しますが、非独立で同一に分散されたデータを含む環境では、その静的特徴の仮定が失敗し、特徴多様体の位置ずれが発生し、モデルのパフォーマンスが大幅に低下します。この矛盾に対処するために、この文書では FedOPAL フレームワークを提案します。このフレームワークは、視覚的プロンプトを特徴修正手段として適応させ、局所近位制約を適用することで異種データの特徴分布を線形分離可能な空間に積極的に修正し、それによって分析的連合学習の理論的前提を満たします。実験結果は、FedOPAL がいくつかのベンチマークで元の分析手法を大幅に上回るだけでなく、サーバー側のトレーニング コストをゼロに維持しながら最先端の反復手法に匹敵する精度を達成し、エッジで大規模なモデルを効率的に連携させるための新しいエンジニアリング パラダイムを提供することを示しています。

原文 (English)

FedOPAL: One-Shot Federated Learning via Analytic Visual Prompt Tuning

With the widespread deployment of basic models in edge intelligence, communication bandwidth has become a core bottleneck restricting the scalability of federated learning. Although one-shot federated learning alleviates this problem by minimizing communication rounds, existing iterative fine-tuning or knowledge distillation methods still face challenges such as high server-side computational costs and hyperparameter sensitivity. Analytical federated learning achieves efficient gradientfree aggregation using least-squares closed-form solutions, but in environments with non-independent and identically distributed data, its static feature assumptions fail, leading to feature manifold misalignment and severely impairing model performance. To address this contradiction, this paper proposes the FedOPAL framework. This framework adapts the visual prompts as feature rectifiers, actively correcting the feature distribution of heterogeneous data to a linearly separable space by applying local proximal constraints, thereby satisfying the theoretical assumptions of analytical federated learning. Experimental results show that FedOPAL not only significantly outperforms the original analytical methods on several benchmarks, but also achieves accuracy comparable to state-of-the-art iterative methods while maintaining zero server-side training costs, providing a new engineering paradigm for efficient collaboration of large models on the edge.

13:00 JSTLLM/生成AI

記憶された知識が大規模言語モデルの微調整で一般化できない理由の機械的理解に向けて

LLM を微調整して新しい知識を注入することは、重大な課題に直面しています。LLM は新しい事実をすぐに記憶できますが、それを下流の推論タスクに使用できません。私たちはこの失敗を \textit{\textbf{知識と使用のギャップ}} として形式化します。これは、精度のギャップと、暗記と一般化の間の時間的な遅れによって特徴付けられます。この現象を理解するために、私たちは目に見えない知識を使って LLM を微調整し、セルフパッチングと呼ばれる新しい介入技術を使用して知識の空間浸透ダイナミクスを内部的に監視します。自己パッチは、表現を再配置することで失敗した汎化ケースを大幅に改善するアクティベーション位置を特定します。これらの結果は、知識回路の不整合仮説と一致しています。つまり、記憶された表現は内部に存在する可能性がありますが、計算効率の高い層にルーティングされない可能性があります。この診断結果の実用性を実証するために、汎化失敗時にオラクルのヘッドルームの 58 ~ 75\% を回復する単純なヒューリスティック戦略を設計します。この発見の堅牢性を確認するために、クロスドメインで実験が行われます。

原文 (English)

Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the \textit{\textbf{Knowing--Using Gap}}, characterized by an accuracy gap and a temporal lag between memorization and generalization. To understand this phenomenon, we fine-tune LLMs with unseen knowledge and monitor the spatial permeation dynamics of the knowledge internally using a novel intervention technique called self-patching. Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases. These results are consistent with a knowledge-circuit misalignment hypothesis: memorized representations can exist internally but may not be routed to computation-effective layers. To demonstrate the practicality of this diagnostic finding, we design a simple heuristic strategy which recovers 58--75\% of the oracle headroom in generalization failure. Experiments are done cross-domain for the robustness of this finding.

13:00 JSTエージェントGPT / ChatGPT

ゲーム理論主導のマルチエージェント フレームワークにより言語モデルの幻覚が軽減される

ルールベースの科学領域における軽量の大規模言語モデルの適用は、公理的な推論を再現するのではなく、言語パターンを模倣する傾向があり、頻繁に幻覚を引き起こす傾向があるため、依然として厳しく制限されています。ここでは、ベイジアンとチーム ゲームの原則を統合した適応型マルチエージェント フレームワークである G-Frame が、高品質のデータ合成とモデル トレーニングのための自動化された閉ループを確立することを示します。構造化推論を通じてドメイン制約を強制的に内面化することで、363,045 の思考連鎖と 199,589 の質問と回答のペアからなる特殊なコーパスを合成しました。結果として得られた 7B モデル OmniChem は、カスタム ベンチマークおよび ChemBench で GPT 4o mini と同等のパフォーマンスを達成しながら、基本アーキテクチャと比較して幻覚が 79.46% 減少しました。さらに、分子設計と合成計画における OmniChem の高度な機能を実証します。この研究は、適応型マルチエージェントを利用して固有の推論欠陥を克服するスケーラブルなパラダイムを確立し、特殊な科学分野での知識発見を加速するための実行可能な経路を提供します。

原文 (English)

Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination

The application of lightweight Large Language Models in rule-based scientific domains remains severely limited by their tendency to mimic linguistic patterns rather than reproduce axiomatic reasoning, causing frequent hallucinations. Here, we show that G-Frame, an adaptive multi-agent framework integrating Bayesian and team game principles, establishes an automated closed-loop for high-quality data synthesis and model training. By forcing the internalization of domain constraints through structured reasoning, we synthesized a specialized corpus of 363,045 chains-of-thought and 199,589 question-answer pairs. The resulting 7B model OmniChem achieves performance parity with GPT 4o mini on custom benchmarks and ChemBench while exhibiting a 79.46% reduction in hallucinations relative to its base architecture. We further demonstrate the advanced capabilities of OmniChem in molecular design and synthesis planning. This work establishes a scalable paradigm utilizing adaptive multi-agents to overcome inherent reasoning deficiencies, offering a feasible pathway for accelerating knowledge discovery in specialized scientific fields.

13:00 JST研究/論文GPT / ChatGPTGemini

OmniFood-Bench: 栄養素の推論と個別の健康アドバイスのための VLM の評価

Large Vision-Language Model (VLM) を重要なインフラストラクチャに迅速に統合することで、個別化されたヘルスケアと食事管理に革命をもたらすことが期待されます。しかし、食品システムの領域では、自律エージェントは、見た目と本質的な栄養成分の間の「全身情報の非対称性」という、独特かつ永続的な課題に直面しています。既存のベンチマークは主に、食品カテゴリーの認識などの粗粒度の分類タスクに焦点を当てており、実際の食事管理に必要な複雑な推論チェーン、具体的には、隠れた原材料の特定から物理量の推定、そして最終的には安全性が重要な医学的アドバイスの総合までを横断する能力を評価できません。このペーパーでは、MM-Food-100K データセットから構築された包括的なベンチマークである OmniFood-Bench を紹介します。以前の研究とは異なり、OmniFood-Bench は、基本的な認識 (食材と調理方法)、定量的推論 (分量と栄養プロファイリング)、および安全性重視の勧告 (疾患固有の推奨事項) の 3 つの進歩的な機能にわたって VLM を評価します。 gpt-5.1、gemini-3-flash、qwen3-vl-8B を含む 6 つの最先端の VLM を評価します。私たちの広範な実験により、驚くべき「意味と物理のギャップ」が明らかになりました。モデルは、料理の命名において人間に近い精度を達成する一方で、質量推定において壊滅的な失敗を示し、高リスクの糖尿病プロファイルに対する良性のアドバイスを頻繁に幻覚で示します。この取り組みにより、公衆衛生のために配備された自律エージェントの信頼性に関する厳格な基準が確立されました。コードとデータセットは、https://anonymous.4open.science/r/OmniFood-Bench-7D0B で入手できます。

原文 (English)

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the domain of food systems, autonomous agents face a unique and persistent challenge: the "Systemic Information Asymmetry" between visual appearance and intrinsic nutritional composition. Existing benchmarks primarily focus on coarse-grained classification tasks, such as food category recognition, which fail to evaluate the intricate reasoning chain required for real-world dietary management -- specifically, the ability to traverse from identifying hidden ingredients to estimating physical mass, and finally synthesizing safety-critical medical advice. In this paper, we introduce OmniFood-Bench, a comprehensive benchmark constructed from the MM-Food-100K dataset. Unlike previous works, OmniFood-Bench evaluates VLMs across three progressive capabilities: Basic Perception (Ingredients & Cooking Methods), Quantitative Reasoning (Portion Size & Nutritional Profiling), and Safety-Critical Advisory (Disease-Specific Recommendations). We evaluate six state-of-the-art VLMs, including gpt-5.1, gemini-3-flash, and qwen3-vl-8B. Our extensive experiments reveal a startling "Semantic-Physical Gap": while models achieve near-human accuracy in naming dishes, they exhibit catastrophic failure in mass estimation and frequently hallucinate benign advice for high-risk diabetic profiles. This work establishes a rigorous standard for trustworthiness in autonomous agents deployed for public health. The code and datasets are available in: https://anonymous.4open.science/r/OmniFood-Bench-7D0B

13:00 JSTハードウェア/半導体

JEPA スタイルの予測学習を JA4 由来のネットワーク フィンガープリントに適用する

I-JEPA と V-JEPA は、元の入力を再生成するのではなく、潜在的な予測をターゲットのエンコーダー出力に照合することで学習します。これは画像やビデオではうまく機能します。同じ目的がコンパクトなネットワーク フィンガープリントでも機能するかどうかを調査します。 JA4DB および CIC-IDS-2017 から抽出された JA4、JA4H、JA4S、および JA4X サブフィールドでトレーニングされた Transformer ベースのモデルである JA4-JEPA を構築しました。トレーニング データは両方のソースからの約 397,000 のサンプルを組み合わせていますが、4 つのビュー ファミリすべてを含む単一のサンプルはありません。私たちは、TLS、DNS、SSH にわたるプロトコル ファミリ分類について、凍結された kNN プローブを使用して学習された表現を評価しました。 39,416 個のホールドアウト サンプルで、モデルはコサイン類似度 0.9899 と kNN 精度 0.9220 を達成しました。これらの結果は、ソース間でビューが不完全に重複している場合でも、JEPA スタイルの予測学習が JA4 由来のフィンガープリントから有用な埋め込みを生成できることを示しています。キーワード: JA4、ネットワークフィンガープリンティング、JEPA、予測表現学習、自己教師あり学習

原文 (English)

Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints

I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has worked well for images and video. We explore whether the same objective works for compact network fingerprints. We built JA4-JEPA, a Transformer-based model trained on JA4, JA4H, JA4S, and JA4X subfields drawn from JA4DB and CIC-IDS- 2017. The training data combines roughly 397K samples from both sources, though no single sample contains all four view families. We evaluated the learned representations with a frozen kNN probe on protocol-family classification across TLS, DNS, and SSH. On 39,416 heldout samples the model achieved a cosine similarity of 0.9899 and a kNN accuracy of 0.9220. These results indicate that JEPA-style predictive learning can produce useful embeddings from JA4-derived fingerprints, even with incomplete view overlap across sources. Keywords: JA4, network fingerprinting, JEPA, predictive representation learning, self-supervised learning

13:00 JST研究/論文

生物医学テキストにおける適応型セマンティック モデリングのためのドリフト対応時間グラフ再配線 (DATGR)

生物医学言語は新しい発見が現れるにつれて急速に進化し、時間の経過とともに従来のテキスト モデルは意味的な忠実性を失います。静的埋め込みと共起グラフではそのような進化を捉えることができないため、検索や知識発見タスクのパフォーマンス低下につながります。この論文では、推定されたセマンティック ドリフトに基づいて共起エッジを動的に更新することで概念の進化をモデル化する、Drift-Aware Temporal Graph Rewiring (DATGR) フレームワークを紹介します。 DATGR は、タイム スライスごとにエンベディングを再トレーニングするのではなく、エッジの重みに適用されるロジスティック更新ルールを使用して、軽量のフィードバック駆動型の再配線を実行します。 Biomedical Multi-Relation Corpus (BIOMRC) で評価したこの方法は、静的ベースラインと比較して絶対差約 0.066 (0.699 対 0.633) の平均受信者動作特性下面積 (AUROC) の改善を達成しました。精度再現率曲線下面積 (AUPRC) は同等の値を維持し (0.738 対 0.744)、ドリフトを意識した適応により精度を損なうことなくリンク予測再現率が向上することが示されました。これらの結果は、エッジレベルの適応が、計算効率と解釈可能性を維持しながら、進化する生物医学テキストの時間的意味変化を効果的に捕捉することを示しています。

原文 (English)

Drift-Aware Temporal Graph Rewiring (DATGR) for Adaptive Semantic Modeling in Biomedical Text

Biomedical language evolves rapidly as new discoveries emerge, causing traditional text models to lose semantic fidelity over time. Static embeddings and co-occurrence graphs cannot capture such evolution, leading to performance degradation in retrieval and knowledge discovery tasks. This paper introduces a Drift-Aware Temporal Graph Rewiring (DATGR) framework that models concept evolution by dynamically updating co-occurrence edges based on estimated semantic drift. Instead of retraining embeddings for each time slice, DATGR performs lightweight, feedback-driven rewiring using a logistic update rule applied to edge weights. Evaluated on the Biomedical Multi-Relation Corpus (BIOMRC), the method achieved a mean Area Under the Receiver Operating Characteristic (AUROC) improvement of approximately 0.066 absolute difference (0.699 vs. 0.633) over a static baseline. Area Under the Precision-Recall Curve (AUPRC) remained comparable (0.738 vs. 0.744), showing that drift-aware adaptation enhances link-prediction recall without a loss in precision. These results demonstrate that edge-level adaptation effectively captures temporal semantic change in evolving biomedical text while remaining computationally efficient and interpretable.

13:00 JST研究/論文

自閉症における顔の感情知覚研究を最適化するための AI 誘導刺激の発見と生成

自閉症の成人と定型発達の成人の間の知覚の違いを理解するには、高感度で信頼性が高く、機械的に有益な行動アッセイが必要です。顔の感情知覚は、グループ間の差異が報告されているものの、研究によって結果が異なるため、有用なテストケースです。今回我々は、この変動性が画像レベルの希薄性を反映している可能性があることを示す。感情判断における自閉症者と定型神経の違いは、刺激全体に均一に広がるのではなく、診断上の表情の小さなサブセットに集中していた。私たちは、自閉症と定型発達の参加者の画像レベルの判断を予測するために集団固有の人工ニューラル ネットワーク モデルをトレーニングし、次にこれらのモデルを使用して、グループ分離を最大化すると予測される新しい顔を選択しました。独立したコホートでは、モデルが選択した画像は、一致するランダムな画像よりも大きな行動の違いを生み出しました。次に、同じモデルを敵対的生成ネットワークで使用して、予測されるグループの一致がより大きくなるように診断画像を変換しました。表現型一致の検証では、合成画像により、一致した元の画像と比較して行動の分離が減少しました。これらの結果は、集団固有の知覚の違いを明らかにする刺激を発見し、変換するためのモデルに基づくフレームワークを確立します。より広範には、行動表現型解析が固定刺激セット全体の平均化を超えて、神経発散的な知覚が発散または収束する条件を特定する最適化されたアッセイにどのように移行できるかを示しています。

原文 (English)

AI-guided stimuli discovery and generation to optimize facial emotion perception studies in autism

Understanding perceptual differences between autistic and neurotypical adults requires behavioral assays that are sensitive, reliable, and mechanistically informative. Facial emotion perception is a useful test case because group differences have been reported, but findings vary across studies. Here we show that this variability may reflect image-level sparsity: autistic-neurotypical differences in emotion judgments were concentrated in a small subset of diagnostic facial expressions rather than spread uniformly across stimuli. We trained population-specific artificial neural network models to predict image-level judgments for autistic and neurotypical participants, then used these models to select novel faces predicted to maximize group separation. In an independent cohort, model-selected images produced larger behavioral differences than matched random images. We then used the same models with a generative adversarial network to transform diagnostic images toward greater predicted group agreement. In phenotype-matched validation, synthesized images reduced behavioral separation relative to their matched originals. These results establish a model-guided framework for discovering and transforming stimuli that reveal population-specific perceptual differences. More broadly, they show how behavioral phenotyping can move beyond averaging across fixed stimulus sets toward optimized assays that identify the conditions under which neurodivergent perception diverges or converges.

13:00 JST研究/論文

CommuniWave:都市コミュニティにおける一時的な非公式な行動の程度を定量化するための機械学習モデル

都市の管理者や設計者にとって、都市コミュニティの機能的特性を改善して、複雑さと不確実性に直面した地域の回復力を強化することは非常に重要です。現在、コミュニティ計画はトップダウンのアプローチに従っていることが多く、住民の非公式な行動を定量化する効果的な指標が欠如しており、当初の計画との衝突が頻繁に発生しています。この研究では、都市コミュニティにおける非公式行為の度合い (DIB) を効率的に検出して定量化するために設計された機械学習モデルである CommuniWave を紹介します。このモデルには、mmaction2 に基づく Behavior Capture Net (BCN)、自社開発の YOLOv10 モデル (YLX)、およびランダム フォレストを使用した Behavior Eval Model (BEM) が統合されています。最終的に、街頭ビデオから DIB 変動チャートを生成することにより、このモデルは動的なモニタリングを容易にし、都市管理者がコミュニティ全体の回復力を強化するための洗練された決定を下せるようにサポートします。

原文 (English)

CommuniWave:A Machine Learning Model for Quantifying the Degree of Temporary Informal Behavior in Urban Communities

For urban managers and designers, improving the functional attributes of urban communities to enhance territorial resilience in the face of complexity and uncertainty is crucial. Currently, community planning often follows a top-down approach and lacks effective metrics to quantify informal behaviors of residents, leading to frequent conflicts with original plans. This study introduces CommuniWave, a machine learning model designed to efficiently detect and quantify the Degree of Informal Behavior (DIB) in urban communities. The model integrates a Behavior Capture Net (BCN) based on mmaction2, a self-developed YOLOv10 model (YLX), and a Behavior Eval Model (BEM) using random forest. Ultimately, by generating DIB fluctuation charts from street videos, the model facilitates dynamic monitoring, supporting urban managers in making refined decisions to enhance the overall resilience of communities.

13:00 JST研究/論文

感情とセンチメントの認識のための SHAP 加重クロスモーダル専門家の融合: 証拠と限界

多峰性の感情とセンチメントの認識は、分類前にモダリティを連結する早期融合、または独立してトレーニングされた単峰性予測子を組み合わせる後期融合によって一般的に対処されます。初期融合は正確ですがモノリシックですが、後期融合はモジュール式ですが、クロスモーダル相互作用が失われる可能性があります。このペーパーでは、XAI ガイド付き適応融合 (\xgaf) を再検討します。XAI ガイド付き適応融合 (\xgaf) は、サンプル レベルの重みが TreeSHAP 属性の大きさから導出される、ユニモーダル エキスパートとクロスモーダル エキスパートをツリーベースで組み合わせたものです。私たちは、専門家の特徴の次元が等しくない場合の SHAP アトリビューション削減の効果に焦点を当てます。この設定では、平均絶対値および中央値絶対値の削減により高次元のクロスモーダルエキスパートを抑制できますが、合計絶対値の削減により総属性質量が維持されます。 MELD 7 クラスの感情認識では、sum-abs \xgaf{} は 3 つの顔シーケンス アグリゲーターにわたる初期融合とほぼ一致します。 Transformer のバリアントは 0.5983 \wf{} に達します。これに対し、初期融合では 0.6018、確率平均後期融合では 0.4598 です。マクネマー検査では、sum-abs \xgaf{} と MELD の早期融合 ($p=1.000$) の間に有意な差は見られませんが、 \xgaf{} は後期融合よりも有意に優れたままです ($p<0.0001$)。 CMU-MOSEI 3 クラス感情認識では、sum-abs \xgaf{} は 0.6519 \wf{} に達し、初期融合 (0.6485) と後期融合 (0.5696) をわずかに上回ります。アブレーション研究によると、主な利点は、複雑なサンプルごとのルーティングではなく、クロスモーダルのエキスパート、特にトリモーダルのエキスパートを追加することによってもたらされます。さらに、診断では、平均腹筋重量と中央値腹筋重量がほぼ均一である一方、合計腹筋重量は三峰性の専門家に集中していることが示されています。したがって、主な貢献は、SHAP 削減、エキスパートの次元性、およびクロスモーダルのエキスパート設計がモジュール式マルチモーダル融合にどのように影響するかについての透明性のある実証分析です。

原文 (English)

SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition: Evidence and Limits

Multimodal emotion and sentiment recognition is commonly addressed by early fusion, which concatenates modalities before classification, or late fusion, which combines independently trained unimodal predictors. Early fusion can be accurate but monolithic, while late fusion is modular but may lose cross-modal interactions. This paper revisits XAI-guided adaptive fusion (\xgaf), a tree-based mixture of unimodal and cross-modal experts whose sample-level weights are derived from TreeSHAP attribution magnitudes. We focus on the effect of SHAP attribution reduction when experts have unequal feature dimensionalities. In this setting, mean-abs and median-abs reductions can suppress high-dimensional cross-modal experts, whereas sum-abs reduction preserves total attribution mass. On MELD 7-class emotion recognition, sum-abs \xgaf{} nearly matches early fusion across three face-sequence aggregators; the Transformer variant reaches 0.5983 \wf{}, compared with 0.6018 for early fusion and 0.4598 for probability-average late fusion. McNemar testing shows no significant difference between sum-abs \xgaf{} and early fusion on MELD ($p=1.000$), while \xgaf{} remains significantly better than late fusion ($p<0.0001$). On CMU-MOSEI 3-class sentiment recognition, sum-abs \xgaf{} reaches 0.6519 \wf{}, slightly exceeding early fusion (0.6485) and late fusion (0.5696). Ablation studies show that the main gain comes from adding cross-modal experts, especially the trimodal expert, rather than from complex per-sample routing. Diagnostics further show that mean-abs and median-abs weights are nearly uniform, while sum-abs weights concentrate on the trimodal expert. Thus, the main contribution is a transparent empirical analysis of how SHAP reduction, expert dimensionality, and cross-modal expert design affect modular multimodal fusion.

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

肝細胞癌の精密治療に向けて: リスク層別化と治療ガイダンスのための臨床推論 LLM

肝細胞癌 (HCC) は一般的な悪性腫瘍であり、癌関連死亡の主な原因です。現在のガイドラインと病期分類システムは大まかなカテゴリーを提供していますが、病期内の異質性や電子医療記録 (EMR) の臨床状況が見逃されていることがよくあります。我々は、日常的なEMRナラティブを読み取り、リスクスコアベースの病期分類、証拠に基づいた理論的根拠を伴うランク付けされたガイドラインに準拠した治療法、および個別の生存推定値を共同で出力する、臨床的に調整された大規模言語モデルであるHCC-STAR(肝細胞癌の病期分類、治療および予後)を紹介します。 SEER から約 30,000 件の HCC 症例を厳選し、臨床医が検証したプロンプトベースの拡張ワークフローを使用して、それらを EMR スタイルのナラティブ トレーニング データに拡張しました。このコーパスでは、臨床ガイドラインのテキストレベルの暗記を超えて、ステップ検証可能な複合報酬で最適化された知識に合わせた推論フレームワークを開発しました。中国の 12 の病院からの 6,668 人の患者からなる多施設コホートにおいて、HCC-STAR は、GPT-5 や Gemini-2.5 Pro などの臨床ガイドラインや競合モデルと比較して、治療推奨とリスク層別化において最先端のパフォーマンスを達成しました。仮説の全生存期間分析では、BCLC および CNLC では生存期間中央値が 29 か月および 32 か月であったのに対し、HCC-STAR 推奨に従っている場合は 51 か月であることが示されました。臨床医中心の評価において、盲検肝胆道専門医は、HCC-STAR の推論と証拠に基づく正当化が信頼できると評価しました。このモデルは治療精度において研修医や主治医を上回り、アシスタントとして使用すると医師がより正確な意思決定をより迅速に下すのに役立ちました。これらの発見は、HCC-STAR が HCC におけるリスク階層化と精密治療のための信頼性があり検証可能な意思決定支援システムであることを裏付けています。

原文 (English)

Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance

Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality. Current guidelines and staging systems provide coarse categories, but often miss within-stage heterogeneity and the clinical context in electronic medical records (EMRs). We present HCC-STAR (Hepatocellular Carcinoma Staging, Treatment And pRognosis), a clinically aligned large language model that reads routine EMR narratives and jointly outputs risk score-based staging, ranked guideline-consistent treatments with evidence-based rationales, and individualized survival estimates. We curated about 30,000 HCC cases from SEER and expanded them into EMR-style narrative training data using a clinician-validated, prompt-based augmentation workflow. On this corpus, we developed a knowledge-aligned reasoning framework optimized with a step-verifiable composite reward, moving beyond text-level memorization of clinical guidelines. In a multi-center cohort of 6,668 patients from 12 hospitals in China, HCC-STAR achieved state-of-the-art performance in treatment recommendation and risk stratification compared with clinical guidelines and competitive models, including GPT-5 and Gemini-2.5 Pro. Hypothetical overall-survival analysis showed a median survival of 51 months under adherence to HCC-STAR recommendations, compared with 29 and 32 months under BCLC and CNLC. In clinician-centric evaluations, blinded hepatobiliary specialists rated HCC-STAR's reasoning and evidence-based justifications as trustworthy. The model surpassed resident and attending physicians in treatment accuracy and helped physicians make more accurate decisions faster when used as an assistant. These findings support HCC-STAR as a reliable and verifiable decision-support system for risk stratification and precision therapy in HCC.

13:00 JSTLLM/生成AI

患者中心の会話型人工知能の複雑さ

大規模言語モデル (LLM) を利用した消費者向けの健康チャットボットが、症状の評価に使用されることが増えています。ただし、チャットボットの開発と評価は、協力的で明確なシミュレートされた患者に依存することがよくあります。私たちは 2,053 件の実際の患者とチャットボットの会話を分析したところ、コミュニケーション パターンと感情の表現がユーザーによって大きく異なることがわかりました。私たちは、臨床内容、感情状態、会話戦略、コミュニケーション スタイルを個別にモデル化する患者シミュレーターを開発しました。 15 人の人間の採点者によるチューリングにヒントを得たリアリズムの評価では、シミュレートされた会話は実際の会話とほとんど区別がつかず、人間の採点者は 55% の精度を達成しました。私たちは、臨床医が等級付けした 1,164 件の症例にわたる 5 つの異なる患者ペルソナを使用して、緊急度評価における 4 つの LLM のパフォーマンスを評価しました。コミュニケーションのスタイルによってトリアージの結果が大きく変わる可能性があることがわかりました。患者中心の会話型人工知能は、コミュニケーションの多様性に対応する必要があります。現実的ではなく理想的な対話を目的に設計されたシステムは、現実世界に展開するとパフォーマンスが低下し、健康格差が拡大する危険性があります。

原文 (English)

The complexities of patient-centred conversational artificial intelligence

Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.

13:00 JSTエージェントDeepSeek

利己的な代理人社会における市場安定のための正式なメカニズム: 市場シミュレーション研究

利己的なエージェントは、制約を受けずに放置されると、度重なる社会的ジレンマの中で離反する傾向があり、貿易から得た共同利益を崩壊させます。この論文は、そのようなエージェントの社会が市場の安定を維持するには、無制限のコミュニケーションの上に重ねられたどのような正式なメカニズムが十分であるか、またそれらのメカニズムが敵対的な攻撃に対してどの程度回復力があるかを調査します。私たちはリサーチクエスチョンをマルチエージェントマーケットプレイスシミュレーションとしてインスタンス化します。このシミュレーションでは、補完的な生産専門分野を持つ 18 個の LLM エージェント (DeepSeek-V3) が、ユーティリティを得るために制約されたソーシャル ネットワーク内で取引する必要があります。私たちは 2 つの実験フェーズを実施します。(1) 200 ラウンドにわたる漸進的なトロール注入下で 8 つの条件にわたるメカニズムを比較し、最高のパフォーマンスを発揮するメカニズムとして Mediation を特定します。 (2) 反復的にプロンプ​​ト最適化された LLM 駆動のトロールを使用した調停の敵対的レッドチーム化。最良の攻撃 (v6) は誠実なエージェントの有用性を 13.3% 削減するが、市場を崩壊させることはできないことが判明しました。調停は、持続的な敵対的な圧力の下でも回復を可能にします。私たちは、敵対的堅牢性を、最適化された攻撃の下で積極的な正直エージェントの有用性を維持するメカニズムの能力として定義し、調停が堅牢であること、つまり曲げることはできても壊れることはないことを発見しました。

原文 (English)

Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study

Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This paper investigates what formal mechanisms, layered on top of unrestricted communication, are sufficient for a society of such agents to maintain market stability, and how resilient those mechanisms are to adversarial attack. We instantiate the research question as a multi-agent marketplace simulation where 18 LLM agents (DeepSeek-V3) with complementary production specialties must trade within a constrained social network to obtain utility. We conduct two experimental phases: (1) a mechanism comparison across eight conditions under progressive troll injection over 200 rounds, identifying Mediation as the top-performing mechanism; and (2) adversarial red-teaming of Mediation using iteratively prompt-optimised LLM-driven trolls, finding that the best attack (v6) reduces honest-agent utility by 13.3% but cannot collapse the market. Mediation enables recovery even under sustained adversarial pressure. We define adversarial robustness as a mechanism's ability to sustain positive honest-agent utility under optimised attack, and find that Mediation is robust: it can be bent but not broken.

13:00 JSTエージェントビジネス/資金調達研究/論文

SolarChain-Eval: 分散型エネルギー市場の信頼できる経済主体のための物理制約付きベンチマーク

エージェント AI システムがサイバー物理環境に適用されることが増えているため、その評価にはタスクのパフォーマンスと信頼性の両方の評価が必要です。分散型エネルギー市場では、自律エージェントは市場の効用を向上させる可能性がありますが、無効な物理データを悪用し、人為的な流動性を生み出し、不安定なガバナンスの決定を生み出す可能性もあります。したがって、信頼できる経済主体を評価するための物理制約付きベンチマークである SolarChain-Eval を提案します。これは市場ガバナンスをギムナジウム互換のマルコフ決定プロセスとして定式化し、エージェントが時間ごとに意思決定を行います。 SolarChain-Eval は、市場の有用性、物理的安全性、スリッページ、アクションのスムーズさ、空間的公平性、監査可能性など、複数の側面にわたって各ポリシーを評価します。エージェント評価をサポートするために、SolarChain-Eval には LLM ベースの Planner/Auditor 層が組み込まれています。プランナーはエピソードレベルのアクションの制限と監査ルールを定義し、監査人はリスクの高いアクションをレビューして修正します。すべての介入は、トリガー信号、提案されたアクション、修正されたアクション、および監査の根拠を含む構造化されたログを通じて記録されます。静的、ランダム、近視、RL、および RL+LLM ポリシーを使用した実験では、ユーティリティと安全性の明確なトレードオフが明らかになりました。 RL エージェントは市場の有用性を向上させますが、依然として危険な動作を引き起こす可能性があります。物理的ペナルティが除去されると、報酬最大化エージェントは無効な生成を利用し、人為的な流動性を増加させます。 LLM プランナー/監査人は監査可能性を向上させ、選択されたリスクを軽減しますが、誤って指定された報酬関数を完全に補償することはできません。これらの結果は、信頼できるエージェント AI 評価には物理的制約と透過的な介入トレースの両方が必要であることを示しています。複製可能性を確保するために、データとコードを GitHub 上でオープン アクセスとしてリリースします。

原文 (English)

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical data, create artificial liquidity, and produce unstable governance decisions. Therefore, we propose SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents. It formulates market governance as a Gymnasium-compatible Markov Decision Process, where agents make hourly decisions. SolarChain-Eval evaluates each policy across multiple dimensions, including market utility, physical safety, slippage, action smoothness, spatial fairness, and auditability. To support agentic evaluation, SolarChain-Eval incorporates an LLM-based Planner/Auditor layer. The Planner defines episode-level action bounds and audit rules, while the Auditor reviews and revises high-risk actions. All interventions are recorded through structured logs, including trigger signals, proposed actions, revised actions, and audit rationales. Experiments with static, random, myopic, RL, and RL+LLM policies reveal a clear utility-safety trade-off. RL agents improve market utility but can still produce unsafe behavior. When the physics penalty is removed, reward-maximizing agents exploit invalid generation and increase artificial liquidity. The LLM Planner/Auditor improves auditability and mitigates selected risks, but it cannot fully compensate for a misspecified reward function. These results indicate that trustworthy agentic AI evaluation requires both physical constraints and transparent intervention traces. We release data and code as open access on GitHub for replicability.

13:00 JSTLLM/生成AIエージェント

重要なときは覚えておいてください: Long-Horizo​​n エージェント向けのプロアクティブ メモリ エージェント

長期的なタスクでは、意思決定に関連した状態が拡大する軌道に散在することがよくありますが、アクション エージェントはそれを表面化して行動する必要があります。軌道が大きくなるにつれて、タスクの要件、環境の事実、以前の試み、診断、未解決のサブ目標がコンテキスト ウィンドウに埋もれたり、コンテキスト ウィンドウの外に押し出されたりして、必要なときに意思決定に影響を与えることができなくなる可能性があります。この失敗モードを「動作状態の減衰」と呼びます。私たちは記憶を受動的検索ではなく能動的介入メカニズムとして研究しています。別のメモリ エージェントが未変更のアクション エージェントと並行して実行され、最近の軌跡から構造化メモリ バンクを更新し、メモリに基づいたリマインダーを挿入するか沈黙を保つかを決定します。このモジュールは、フロンティア アクション エージェントおよび既存のエージェント ハーネスとのプラグ アンド プレイです。 Terminal-Bench 2.0 と $\tau^2$-Bench 全体で、より弱いアクション エージェントとより強力なアクション エージェントの両方で pass@1 が改善され、ターミナル ベンチでは +8.3 pp、$\tau^2$-Bench では +6.8 pp のゲインが得られます。アブレーションは、選択的介入が受動的バンクエクスポージャー、常時インジェクション、アドバイザーのみのガイダンス、および一般的な回収よりも優れていることを示しています。オープンウェイト メモリ ポリシーに向けた初期のステップとして、SFT と GRPO を使用して Qwen3.5-27B を SETA 上でトレーニングし、検証報酬を向上させ、ターミナル ベンチへの部分的な転送を実現します。

原文 (English)

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode "behavioral state decay". We study memory as an active intervention mechanism rather than passive retrieval. A separate memory agent runs alongside an unmodified action agent, updating a structured memory bank from the recent trajectory and deciding whether to inject a memory-grounded reminder or remain silent. The module is plug-and-play with frontier action agents and existing agent harnesses. Across Terminal-Bench 2.0 and $\tau^2$-Bench, it improves pass@1 for both weaker and stronger action agents, with gains of +8.3 pp on Terminal-Bench and +6.8 pp on $\tau^2$-Bench. Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval. As an early step toward open-weight memory policies, we train Qwen3.5-27B on SETA using SFT and GRPO, improving validation reward and achieving partial transfer to Terminal-Bench.

13:00 JSTLLM/生成AIビジネス/資金調達

等価性の幻想: LLM における量子化効果の統計的特徴付け

トレーニング後の量子化は、リソースに制約のある設定で大規模な言語モデルをデプロイするために広く使用されていますが、その評価はほぼもっぱら精度と複雑さに依存します。これらの指標では、量子化によって引き起こされる行動の変化を捉えることができないことを示します。絶対精度とは無関係に、基本モデルとその量子化されたバリアント間の正しい予測の重複を測定する意思決定レベルの指標である正確性一致を導入します。 8 ビットから 2 ビットまでの複数のモデルと量子化スキームにわたって、タスクのパフォーマンスが維持されているように見える場合でも、適度な量子化の下では動作の発散が現れることがわかりました。この効果を説明するために、注意の重みに対する構造演算子として量子化を分析し、統計的および分布的尺度を使用して層ごとの歪みを定量化します。私たちの結果は、低ビット幅での非線形ブレークポイントを明らかにし、クエリとキーの投影が値と出力の投影よりも常に敏感であることを示しています。これらの発見は、基本モデルと量子化モデルの間の等価性の幻想を明らかにし、従来のパフォーマンス指標を超えた行動評価を動機付けます。

原文 (English)

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even when task performance appears preserved. To explain this effect, we analyze quantization as a structural operator on attention weights and quantify layer-wise distortions using statistical and distributional measures. Our results reveal non-linear breakpoints at low bit-widths and show that query and key projections are consistently more sensitive than value and output projections. These findings expose an illusion of equivalence between base and quantized models and motivate behavioral evaluation beyond conventional performance metrics.

13:00 JSTLLM/生成AI

知識としてのワークフロー: LLM を介したワークフローのセマンティック永続性

大規模言語モデル (LLM) アプリケーションでは、ツールの使用、取得、分岐、チェックポイント、人間による承認のために明示的なワークフローを使用することが増えています。既存のワークフロー システムは、実行に関する多くの懸念事項にすでに対処しています。この論文は、Lisp にインスピレーションを得た、言語に依存しない概念モデルを提案します。記号形式、オブジェクトの同一性、およびライブイメージの思考は、実装へのコミットメントではなく、説明レンズとして使用されます。このモデルでは、ワークフロー定義、ワークフロー インスタンス、推論レコード、コンテキスト スナップショット、および依存関係が、共有知識基盤内の永続的な知識オブジェクトとして表されます。その中心的な意味上の違いは、導出と推論です。導出は、利用可能な状態に対する決定論的な計算です。 infer は、宣言されたコンテキストと実行者が制御する機能ポリシーに基づいて仲介される LLM 判断です。その結果、セマンティック永続性の予備的な概念的な説明が得られます。ワークフローは単に知識を生成して痕跡を残すだけでなく、それ自体を検査可能、再開可能、およびレビュー可能なナレッジ オブジェクトとして表すことができますが、正式な移行セマンティクスは今後の作業として残ります。

原文 (English)

Workflow as Knowledge: Semantic Persistence for LLM-Mediated Workflows

Large language model (LLM) applications increasingly use explicit workflows for tool use, retrieval, branching, checkpointing, and human approval. Existing workflow systems already address many execution concerns. This paper proposes a Lisp-inspired but language-independent conceptual model: symbolic forms, object identity, and live-image thinking are used as explanatory lenses, not implementation commitments. In this model, workflow definitions, workflow instances, inference records, context snapshots, and dependency relations are represented as persistent knowledge objects in a shared knowledge substrate. Its central semantic distinction is between derive and infer: derive is deterministic computation over available state; infer is mediated LLM judgment under declared context and executor-controlled capability policy. The result is a preliminary conceptual account of semantic persistence: workflows do not merely produce knowledge and leave traces, but can themselves be represented as inspectable, resumable, and reviewable knowledge objects, while formal transition semantics remain future work.

13:00 JST画像/動画生成エージェント研究/論文

AUTOPILOT VQA: インシデント中心の車載カメラを理解するための視覚言語モデルのベンチマーク

視覚言語モデル、大規模言語モデル、およびマルチモーダル大規模言語モデルの最近の進歩により、シーンの理解、意思決定、軌道予測、視覚的な質問応答などの自動運転タスクが改善されました。ただし、これらのモデルが安全上重要なインシデントを確実に推論できるかどうかを評価することは依然として困難です。このギャップに対処するために、車載カメラのビデオを理解するためのインシデント中心の視覚的な質問応答ベンチマークである AUTOPILOT-VQA を紹介します。このデータセットは、現実世界の運転事故やそれに近い事故を中心に設計された構造化された質問を通じて、さまざまなシステムを評価します。このベンチマークは、天候と照明条件、交通環境、道路レイアウト、路面状態、標識、関係主体、事故発生、衝突場所、回避可能性関連の推論など、安全に関連するさまざまなカテゴリをカバーしています。 AUTOPILOT-VQA は、状況に応じたシーンのプロパティとイベントレベルのインシデントの詳細の両方に関する根拠のある質問に答えることをモデルに要求することで、物体認識を超えて、時間的に根拠のある安全性を意識した推論へと移行します。このデータセットは、AUTOPILOT CVPR 2026 コンペティションの一部としてリリースされ、さまざまなシナリオにおける自動運転システムの信頼性を評価するための標準化されたベンチマークを提供します。当社のベンチマークは、現実世界の自動運転向けに、より解釈可能で堅牢かつ安全性を意識したビジョン言語システムの開発をサポートします。

原文 (English)

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving incidents and near-incidents. The benchmark covers diverse safety-relevant categories, including weather and lighting conditions, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability-related reasoning. By requiring models to answer grounded questions about both contextual scene properties and event-level incident details, AUTOPILOT-VQA moves beyond object recognition toward temporally grounded, safety-aware reasoning. The dataset is released as part of the AUTOPILOT CVPR 2026 competition and provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios. Our benchmark support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.

13:00 JST研究/論文

高等教育における AI ベースの学習アシスタントの使用: 大規模な記述分析

この研究では、高等教育における AI ベースの学習アシスタント (Syntea) の使用に関する大規模な記述分析を紹介します。遠隔学習に登録した 77,543 人の学生からの客観的なログ データに基づいて、性別、年齢層、学習クラスター、学位、学習モードにわたる使用パターンを調査します。これまで、教育用チャットボットに関する既存の研究は比較的小規模なサンプルと自己申告による調査データに大きく依存しており、実際の使用行動に関する大規模な証拠は依然として限られています。私たちの調査結果は、Syntea がすでに多くの学習者の学習ルーチンに組み込まれているものの、その使用法は人口統計的および構造的背景によって異なることを示しています。これらのパターンを特定することで、私たちの研究は、AI ベースの学習サポートのさらなる開発のための実証的基礎を提供し、高等教育における教育用チャットボットの使用に関する大規模な分析に貢献します。

原文 (English)

Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive Analysis

In this study, we present a large-scale descriptive analysis of the use of an AI-based learning assistant (Syntea) in higher education. Based on objective log data from 77,543 students enrolled in distance studies, we examine usage patterns across gender, age group, study cluster, degree, and study mode. To date, existing research on educational chatbots has largely relied on comparatively small samples and self-reported survey data, while large-scale evidence on actual usage behavior remains limited. Our findings show that Syntea is already embedded in the study routines of many learners, but that usage differs across demographic and structural contexts. By identifying these patterns, our study provides an empirical basis for the further development of AI-based learning support and contributes a large-scale analysis of educational chatbot usage in higher education.

13:00 JST研究/論文

アイデアにはゲノムがある: 科学的な系統推論と系統に基づいたアイデア生成のベンチマーク

科学的なアイデアが白紙のページから始まることはほとんどありません。これらは、生物学的ゲノムと同様に、メカニズムを継承し、既知の制限を修復し、以前の研究の一部を再結合します。現在のベンチマークでは、AI システムがこの継承構造に従うことができるかどうかについてはまだほとんど言及されていません。科学的な系統推論と系統に基づいたアイデア生成のベンチマークである IdeaGene-Bench (IG-Bench) を紹介します。 IG-Bench は IdeaGene フレームワークを中心に構成されています。各論文または提案は、最小限で型付けされ、証拠に基づいた Idea Genome オブジェクトのセットとして表され、GenomeDiff はこれらのオブジェクトを調整して、6 つの操作上の進化ダイナミクスに基づいて継承、突然変異、喪失、外部インポート、および新規挿入を記録します。このベンチマークには、1,961 の黄金系統トレース、1,085 の厳選された Idea Genome オブジェクト、および 10 の科学ドメインにわたる 920 のペアごとの GenomeDiff レコードが含まれています。 2 つの評価をサポートします。 IG-Exam (42 のタスク タイプ、1,029 のインスタンス) は、アイデア ゲノムの抽象化、継承追跡、進化的推論、および系統検証にわたる閉じた形式の系統推論をテストします。 IG-Arena は、系統条件付き人口進化スコア (PES) を使用して世代を評価し、特定の系統集団の一貫した子孫として提案を挿入できるかどうかを尋ねます。適切なアイデア ゲノム オブジェクトを継承し、近隣の研究から有意義に変化し、将来の研究に選択値を提供する必要があります。 14 人の LLM ベースの科学者による実験により、組成上のボトルネックが明らかになりました。最も強力なシステムでも、系統推論の正確な精度は 27.3% にすぎず、構造化された系統コンテキストにより、すべての参加者を均一に支援するのではなく、システムのランキングが再調整されます。

原文 (English)

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.

13:00 JST研究/論文

メモリバックトラッキングとトポロジカルアトリビューションによる時間グラフネットワークの説明可能性の実現に向けて

時間グラフは現実世界のアプリケーションで広く普及しており、時間グラフ ネットワーク (TGN) は優れた予測精度を実現しています。どの過去のイベントがモデル予測を推進しているかを理解することで、TGN の信頼性を高めることができます。既存の説明方法では、ノードの履歴を記録および更新するコアコンポーネントであるメモリモジュールが無視されており、過去のイベントの影響が解明されていません。これに対処するために、トポロジ アトリビューション ツリーとメモリ バックトラッキング ツリーを通じて TGN 予測を帰属させます。トポロジ アトリビューション ツリーは、近隣ノードとそのメモリ ベクトルの影響をキャプチャし、メモリ バックトラッキング ツリーは、履歴イベントがどのようにノード メモリ ベクトルを形成するかを定量化します。 TGN に LRP を適用し、イベントの合計寄与がモデルのロジットと等しいことを保証します。最後に、ロジットから確率への非線形マッピングにより、top-k の選択が不正確になる可能性があるため、重要なイベントを特定するために最適化目標を設計します。 9 つの時間グラフ データセット (スパン ノード プロパティ予測、リンク予測タスク、グラフ分類タスク) での実験により、私たちの方法が忠実な説明を提供し、最先端のベースラインを上回るパフォーマンスを示すことが示されました。コードは https://github.com/yazhengliu/MemExplainer で入手できます。

原文 (English)

Towards the Explainability of Temporal Graph Networks via Memory Backtracking and Topological Attribution

Temporal graphs are ubiquitous in real-world applications and Temporal Graph Networks (TGNs) have achieved superior predictive accuracy. Understanding which historical events drive model predictions can enhance trustworthiness of TGNs. Existing explanation methods overlook the memory module, the core component that records and updates node histories, leaving the influence of past events unexplored. To address this, we attribute TGNs predictions through the topology attribution tree and memory backtracking tree. The topology attribution tree captures the influence of neighbors and their memory vectors, then the memory backtracking tree quantifies how historical events shape node memory vectors. We apply the LRP in TGNs, ensuring that the total contribution of events equals the logits of model. Finally, top-k selection may be unfaithful due to the nonlinear mapping from logits to probabilities, we design optimization objectives to identify the important events. Experiments on nine temporal graph datasets, spanning node property prediction, link prediction tasks and graph classification tasks, show that our method provides faithful explanations and outperforms state-of-the-art baselines. The code is available at https://github.com/yazhengliu/MemExplainer

13:00 JST研究/論文

LLT: PDE 演算子学習のためのローカル線形変換器

ニューラル オペレーターは、PDE 解マップを学習し、数値シミュレーションを高速化するための一般的なアプローチになっています。トランスフォーマーベースのニューラル演算子は、計算領域における長距離の依存関係を注意によって学習できるため、特に興味深いものです。ただし、標準アテンションには、偏微分方程式に適用する場合、2 つの大きな制限があります。それは、計算ノードの数に応じて二次的にスケールされることと、ローカルな相互作用に対する明示的なバイアスが欠けていることです。これらの問題に対処するために、PDE 演算子の学習用にローカル線形変換器 (LLT) を導入します。このアーキテクチャは、線形グローバル アテンションとローカル空間混合を組み合わせ、座標およびジオメトリ情報を組み込みます。弾性、塑性、翼形流れ、パイプ流れ、ダルシー流れなど、いくつかの偏微分方程式問題について LLT を評価します。これらの問題の参照データは、構造化メッシュと非構造化メッシュ上の有限要素、有限体積、および有限差分離散化に及びます。先行研究の他の神経演算子および変換器のベースラインと比較して、LLT はこれらの問題全体で競合するか、またはより低い相対 $L_2$ 誤差を達成します。整合構造化離散化では、トレーニング反復あたりの実時間は、Transolver と比較して 1.8 ~ 2.5 倍短縮されます。また、このアプローチをスケールし、サンプルごとに 32,186 個の非構造化メッシュ ポイントを含む 3 次元の自動車空気力学データセットに適用します。まとめると、これらの結果は、LLT が離散化、メッシュ タイプ、問題設定にわたる PDE 問題に対して正確で計算効率の高い演算子を提供することを示しています。

原文 (English)

LLT: Local Linear Transformer for PDE Operator Learning

Neural operators have become a common approach for learning PDE solution maps and accelerating numerical simulations. Transformer-based neural operators are of particular interest, since attention can learn long-range dependencies in the computational domain. However, standard attention has two major limitations when applied to PDEs: it scales quadratically with the number of computational nodes, and it lacks an explicit bias toward local interactions. To address these issues, we introduce Local Linear Transformer (LLT) for PDE operator learning. The architecture combines linear global attention with local spatial mixing, and incorporates coordinate and geometry information. We evaluate LLT on several PDE problems, including elasticity, plasticity, airfoil flow, pipe flow, and Darcy flow. The reference data for these problems span finite-element, finite-volume, and finite-difference discretizations on structured and unstructured meshes. Compared with other neural-operator and transformer baselines from prior studies, LLT achieves competitive or lower relative $L_2$ error across these problems. On matched structured discretizations, wall-clock time per training iteration is reduced by factors of 1.8 to 2.5 relative to Transolver. We also scale the approach and apply it to a three-dimensional car aerodynamics dataset with 32,186 unstructured mesh points per sample. Together, these results indicate that LLT provides an accurate and computationally efficient operator for PDE problems across discretizations, mesh types, and problem settings.

13:00 JSTLLM/生成AI画像/動画生成

ReCoLoRA: スペクトルを意識した再帰的統合による継続的な LLM 微調整

パラメーター効率の高い微調整により、大規模な言語モデルを 1 つのタスクに低コストで適応させることができますが、タスク シーケンス全体で LoRA スタイルのメソッドは同じ固定された重みに低ランクの更新を積み重ね続けるため、新しいタスクがそれぞれ以前のタスクを上書きする傾向があります。我々は、継続的な微調整のためのスペクトル認識フレームワークである ReCoLoRA (低ランク アダプターの再帰的統合) を紹介します。アダプターは、事前トレーニングされた重みのランダム化された SVD から初期化され、レイヤーごとの有効ランクがエルボ基準によって選択され、残存容量がオープンされる前に主部分空間が適応されます。新しいタスクを作成する前に、ReCoLoRA は元のタスクではなく現在の有効な重みを凍結残差、ゆっくりと更新される主成分、および新しいアダプター (再帰的統合) に再分解します。そのため、すべてのタスクは、すでに先行タスクを吸収したモデルから開始されます。 4 つの 7-8B バックボーンにわたる 6 タスクの連続 GLUE シーケンスでは、ReCoLoRA は、より少ないパラメーターをトレーニングしながら、ランクを席巻した LoRA、PiSSA、AdaLoRA、および DoRA ベースラインに対して 4 つのバックボーンのうち 3 つで最高の最終平均スコアを達成しました。 Oracle ルーティングされたタスクバンクのバリアントは、完全なタスク分離の下で上限として機能します。コード: https://github.com/bhqy666/ReCoLoRA。

原文 (English)

ReCoLoRA: Spectrum-Aware Recursive Consolidation for Continual LLM Fine-Tuning

Parameter-efficient fine-tuning adapts a large language model to one task cheaply, but across a task sequence LoRA-style methods keep stacking low-rank updates on the same frozen weight, so each new task tends to overwrite the previous ones. We present ReCoLoRA (Recursive Consolidation of Low-Rank Adapters), a spectrum-aware framework for continual fine-tuning: adapters are initialized from a randomized SVD of the pretrained weight, per-layer effective ranks are selected by an elbow criterion, and the principal subspace is adapted before residual capacity is opened. Before each new task, ReCoLoRA re-decomposes the current effective weight, rather than the original one, into a frozen residual, a slowly updated principal component, and a fresh adapter (recursive consolidation), so every task starts from the model that has already absorbed its predecessors. On a six-task continual GLUE sequence over four 7-8B backbones, ReCoLoRA attains the best final average score on three of the four backbones against rank-swept LoRA, PiSSA, AdaLoRA, and DoRA baselines while training fewer parameters; an oracle-routed task-bank variant serves as an upper bound under full task isolation. Code: https://github.com/bhqy666/ReCoLoRA.

13:00 JST研究/論文

オムニスリープ: CNS の階層的対照学習による睡眠基盤モデル - ANS ダイナミック

睡眠生理学は、EEG、EOG、EMG、ECG、呼吸などのマルチモーダルポリ睡眠検査信号に反映される、中枢神経系 (CNS) と自律神経系 (ANS) の協調的な動態から生じます。しかし、既存の睡眠基盤モデルは、トポロジーに依存しない方法で異種の生体信号を融合し、その生理学的組織を見落としていることがよくあります。トポロジ制約のある表現学習の生理学的事前分布として CNS/ANS パーティションを使用する睡眠基盤モデルである Omni-Sleep を紹介します。 Omni-Sleep は 3 つの目的を通じて構造化表現を学習します。1 つはシステム内の一貫性であり、神経信号および心肺信号内の共有サブシステム レベルの要素を捕捉します。システム間の同期。脳と身体のダイナミクスをモデル化するためにサブシステムの軌道を調整します。そして、長期にわたる睡眠のダイナミクスを捉える潜在空間マスク時間モデリング。 100,000 時間を超える多施設のマルチモーダル PSG データで事前トレーニングされた Omni-Sleep は、睡眠ステージングと複数の疾患分類に基づいて評価されます。データセットおよびモダリティアブレーション設定全体にわたって、Omni-Sleep は強力な基礎モデルのベースラインを上回り、ラベル効率の向上、データセット間の一般化、欠落モダリティに対する堅牢性を示しています。これらの結果は、一般化可能な睡眠表現学習における生理学的階層の価値を強調しています。コードは https://github.com/AutoBrain-sleep/OmniSleep で入手できます。

原文 (English)

Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS--ANS Dynamic

Sleep physiology arises from the coordinated dynamics of the central nervous system (CNS) and autonomic nervous system (ANS), as reflected by multimodal polysomnography signals including EEG, EOG, EMG, ECG, and respiration. However, existing sleep foundation models often fuse heterogeneous biosignals in a topology-agnostic manner, overlooking their physiological organization. We introduce Omni-Sleep, a sleep foundation model that uses the CNS/ANS partition as a physiological prior for topology-constrained representation learning. Omni-Sleep learns structured representations through three objectives: intra-system consistency, which captures shared subsystem-level factors within neural and cardio-respiratory signals; inter-system synchronization, which aligns subsystem trajectories to model brain--body dynamics; and latent-space masked temporal modeling, which captures long-horizon sleep dynamics. Pre-trained on over 100,000 hours of multi-center multimodal PSG data, Omni-Sleep is evaluated on sleep staging and multi-disease classification. Across datasets and modality-ablation settings, Omni-Sleep outperforms strong foundation-model baselines, showing improved label efficiency, cross-dataset generalization, and robustness to missing modalities. These results highlight the value of physiological hierarchy for generalizable sleep representation learning. Code is available at https://github.com/AutoBrain-sleep/OmniSleep.

13:00 JST研究/論文

SHIFT: 不完全かつ異種のゲノムデータからの生存予測

ゲノム予測モデルは、施設間でシーケンスパネルが異なるため、施設間での移行に失敗することが多く、導入時に構造的特徴が欠落する可能性があります。この課題に対する既存のアプローチは通常、解析をコホート間で共有される遺伝子に限定したり、不完全なプロファイルを持つ患者を除外したり、検査時の補完に依存したりするもので、これらすべてが堅牢性を低下させ、多施設データの使用を制限する可能性があります。我々は、テスト時の代入を行わずに不完全なゲノム入力から直接予測する欠損認識生存モデルである、Transformer を使用した生存予測処理 (SHIFT) を提案します。 SHIFT は各ゲノム特徴を個別に表し、マスクされた自己注意と特徴可用性マスクを使用するため、予測は観察された入力のみに基づいています。さらに、トレーニング中に可変レートの特徴マスキングを導入して、異種の欠損パターンに対する堅牢性を向上させます。私たちは、コホート間パネルの重度の不一致を伴う困難な設定を含む、複数のコホートにわたる外部検証を使用して、神経膠芽腫および肺扁平上皮癌に対するアプローチを評価します。これらの設定全体にわたって、SHIFT は強力な一般化を示し、異なる特徴セットにわたって単一のモデルを使用しながら、標準の生存ベースラインや代入ベースのアプローチと比べて優れています。また、開発中に不完全なコホートからの患者を組み込むことで外部データのパフォーマンスが向上することもわかり、部分的に観察されたコホートをモデル構築から除外する必要がないことが示唆されています。これらの結果は、精密腫瘍学における多施設生存予測の実践的な戦略として欠落認識モデリングを裏付けるものです。

原文 (English)

SHIFT: Survival Prediction from Incomplete and Heterogeneous Genomic Data

Genomic prediction models often fail to transfer across institutions because sequencing panels differ across sites, creating structural feature missingness at deployment. Existing approaches to this challenge typically restrict analysis to genes shared across cohorts, exclude patients with incomplete profiles, or rely on test-time imputation, all of which can reduce robustness and limit the use of multi-center data. We propose Survival prediction Handling Incomplete Features using Transformer (SHIFT), a missingness-aware survival model that directly predicts from incomplete genomic inputs without test-time imputation. SHIFT represents each genomic feature separately and uses masked self-attention, along with a feature-availability mask, so that predictions are based only on observed inputs. Further, we introduce variable-rate feature masking during training to improve robustness to heterogeneous missingness patterns. We evaluate the approach on glioblastoma and lung squamous cell carcinoma with external validation across multiple cohorts, including a challenging setting with severe cross-cohort panel mismatch. Across these settings, SHIFT shows strong generalization and compares favorably with standard survival baselines and imputation-based approaches, while using a single model across differing feature sets. We also find that incorporating patients from incomplete cohorts during development can improve performance on external data, suggesting that partially observed cohorts need not be excluded from model building. These results support missingness-aware modeling as a practical strategy for multi-center survival prediction in precision oncology.

13:00 JSTLLM/生成AI

基礎モデルによる集合知

基礎モデルの規模と多様性が増大するにつれて、複数のモデルを協調的な推論システムに調整することで、より安全で信頼性の高い AI への道が開かれます。この章では、ソルバー モデルが独立したドラフトを生成し、それぞれが批評家エージェントによる構造化された批評と修正を受け、アグリゲーター エージェントが最終的な合意ソリューションを合成するマルチエージェント フレームワークについて説明します。スコアリング モジュールは、すべてのエージェントにわたって意味的、数値的、および手順的な評価を提供します。微積分、物理学、化学、生物学、経済学、最適化、統計、数学にわたるベンチマークのアブレーション研究を通じて、フレームワーク アーキテクチャとモデルの多様性の寄与を分離します。 (1) 個別のベースライン、(2) 1 つの共有モデルを使用する同種フレームワーク、(3) 同じモデルの複数のインスタンスを使用する冗長同種ソルバー、(4) 多様な特殊モデルを使用した異種フレームワークの 4 つの構成を比較します。結果は、フレームワーク構造と冗長サンプリングによってわずかな利益が得られるものの、モデルの不均質性が大幅なパフォーマンス向上を促進する重要な要因であることを示しています。異種構成では、カテゴリおよび難易度レベル間のばらつきが低減され、優れた段階的精度 (個々のモデルで 0.64 対 0.54、同種構成の 2.3 倍の向上) が実現します。段階的な推論の品質 (最終的な答えだけでなく、中間ステップの正確さ) は、モデルの多様性によってのみ劇的に向上します。これは、異種エージェントが、説明可能性と監査可能性に不可欠な補完的なエラー検出と推論の洗練を提供することを示しています。アーキテクチャの原則、評価方法論、Global Applied AI への影響について議論し、異種マルチエージェントの調整が科学分野と産業分野全体にわたる透明性、監査可能性、信頼性の高い意思決定をどのようにサポートするかを示します。

原文 (English)

Collective Intelligence with Foundation Models

As foundation models grow in scale and diversity, coordinating multiple models into cooperative reasoning systems offers a path toward safer, more reliable AI. This chapter presents a multi-agent framework where solver models generate independent drafts, each undergoes structured critique and revision by a critic agent, and an aggregator agent synthesizes a final consensus solution. A scoring module provides semantic, numerical, and procedural evaluation across all agents. Through ablation studies on a benchmark spanning calculus, physics, chemistry, biology, economics, optimization, statistics, and mathematics, we isolate the contributions of framework architecture versus model diversity. We compare four configurations: (1) Individual Baseline, (2) Homogeneous Framework using one shared model, (3) Redundant Homogeneous Solvers using multiple instances of the same model, and (4) Heterogeneous Framework with diverse specialized models. Results show that while framework structure and redundant sampling yield modest gains, model heterogeneity is the critical factor driving substantial performance improvements. The heterogeneous configuration achieves superior step-wise accuracy (0.64 vs. 0.54 for individual models; 2.3x improvement over homogeneous configurations) with reduced variance across categories and difficulty levels. Step-wise reasoning quality (correctness of intermediate steps, not just final answers) improves dramatically only with model diversity, showing that heterogeneous agents provide complementary error detection and reasoning refinement essential for explainability and auditability. We discuss architectural principles, evaluation methodology, and implications for Global Applied AI, showing how heterogeneous multi-agent coordination supports transparent, auditable, high-confidence decision-making across scientific and industrial domains.

13:00 JSTLLM/生成AI

Jet-Long: 動的二焦点 RoPE による効率的なロングコンテキスト拡張

最新の LLM は、検索拡張生成、リポジトリ レベルのコーディング、エージェント ワークフローなどの長いコンテキストのアプリケーションにデプロイされることが増えています。エージェント ワークフローでは、蓄積された推論とツール トレースにより、入力が日常的に事前トレーニング ウィンドウを桁違いに押し上げ、ゼロショット コンテキスト拡張がオープンウェイト チェックポイントの主要なデプロイ パスとなっています。既存のゼロショット手法のほとんどは、単一の再スケーリング係数を前もって修正するため、積極的な係数は短いコンテキストの忠実度を犠牲にし、保守的な係数は長いコンテキストでは機能しません。我々は、ローカル RoPE に忠実なウィンドウと、リスケーリング係数が現在のシーケンス長に動的に適応する長距離ウィンドウを組み合わせ、短い入力では基本モデルを正確に復元しながら、長い入力ではきれいに外挿する、チューニング不要のゼロショット手法である Jet-Long を提案します。包含-除外注意マージとオンザフライ RoPE 補正回転により、推論時に二焦点構造が本質的に自由になります。単一の CuTe カーネルに融合され、ロングコンテキストのプリフィルは H100 で最大 $1.39\times$ の FA2 スループットに達し (ホッパーのみの FA4 に近づきます)、単一バッチ生成では長さごとに $\le 4\%$ のオーバーヘッドが発生します。最大 128K コンテキストの Qwen3-1.7B/4B/8B では、Jet-Long は 1.7B/4B/8B の最も強力なベースラインを $+4.79$/$+2.18$/$+2.03$~pp 上回って RULER をリードし、HELMET-RAG (ダウンストリームのロングコンテキスト パフォーマンスの最も効率的な予測子として HELMET によって特定されたベンチマーク) で最高の総合精度を達成し、 PG-19 の困惑度は最も低い。 Jet-Long は、Jet-Nemotron などのハイブリッド アテンション アーキテクチャにも一般化して、再トレーニングなしで長いコンテキストをさらに改善し、ハイパーパラメータの回復力を維持して展開を容易にします。

原文 (English)

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE

Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. Most existing zero-shot methods fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one breaks down at long contexts. We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length, recovering the base model exactly at short inputs while extrapolating cleanly at long ones. An inclusion-exclusion attention merge and an on-the-fly RoPE correction rotation make the bifocal construction essentially free at inference; fused into a single CuTe kernel, long-context prefill reaches up to $1.39\times$ FA2 throughput on H100 (approaching the Hopper-only FA4), and single-batch generation incurs $\le 4\%$ overhead at every length. On Qwen3-1.7B/4B/8B up to 128K context, Jet-Long leads RULER by $+4.79$/$+2.18$/$+2.03$~pp over the strongest baseline at 1.7B/4B/8B, achieves the best overall accuracy on HELMET-RAG (a benchmark identified by HELMET as the most efficient predictor of downstream long-context performance) and attains the lowest PG-19 perplexity. Jet-Long also generalizes to hybrid attention architectures such as Jet-Nemotron for further long-context improvement without retraining, and remains hyperparameter-resilient for ease of deployment.

13:00 JST研究/論文

MetaNCA によるアーキテクチャの一般化

自己組織化は生命の創発的な特性であり、局所的な情報に基づいて作用する個々のコンポーネントの集合的な動作によって駆動されます。生物学的ニューロンは、シナプスを介して伝達される局所的な相互作用を通じて効率的に学習することができ、生物の寿命にわたってその接続を適応させることができます。適応性と局所的相互作用のこれらの望ましい特性を動機として、ニューラル セルラー オートマトン (NCA) モデルは、局所的な更新ルールのみを通じて形態形成を学習することに成功し、多くの更新に対する安定性と摂動に対する堅牢性を実証しています。この研究では、人工ニューラル ネットワークの重みを自己組織化するローカル ルールを学習するフレームワークである Meta Neural Cellular Automata (MetaNCA) を紹介します。学習されたルール ネットワークは、計算グラフ上のローカルな相互作用のみを使用して、タスク ネットワークの重みを繰り返し更新します。我々は、線形アテンションを使用して隣接する重みと隠れ状態からの信号を集約する、ローカル ルール ネットワーク用の新しい Weight Transformer アーキテクチャを提案します。トレーニングが完了すると、ルール ネットワークは逆伝播を行わずにさまざまなアーキテクチャのタスク ネットワークを生成します。 MetaNCA が MNIST および CIFAR-100 上のフィードフォワード MLP、CNN、および ResNet の重みを生成し、200 万パラメータのネットワークに拡張することを示します。さらに、MetaNCA がメタトレーニング中には見ら​​れないアーキテクチャに一般化すること、およびトレーニング段階でのアーキテクチャの多様性がこの一般化を強化することを示します。

原文 (English)

Architecture Generalization with MetaNCA

Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information. Biological neurons, through local interactions transmitted through synapses, are able to learn efficiently and can adapt their connections over an organism's lifespan. Motivated by these desirable properties of adaptability and local interaction, neural cellular automata (NCA) models have been successful at learning morphogenesis solely through local update rules, demonstrating stability over many updates and robustness to perturbations. In this work, we introduce Meta Neural Cellular Automata (MetaNCA), a framework that learns local rules which self-organize the weights of artificial neural networks. A learned rule network iteratively updates the weights of a task network using only local interactions on the computation graph. We propose a novel Weight Transformer architecture for the local rule network, which uses linear attention to aggregate signals from neighboring weights and hidden states. Once trained, the rule network generates task networks of diverse architectures without backpropagation. We show that MetaNCA generates weights for feedforward MLPs, CNNs, and ResNets on MNIST and CIFAR-100, scaling to networks of 2 million parameters. We further show that MetaNCA generalizes to architectures not seen during meta-training, and that architectural diversity in the training phase strengthens this generalization.

13:00 JSTエージェント

強化学習エージェントにおける表現型のような障害のトランス診断空間

人工エージェントにおける心理障害のモデル化は、計算論的精神医学のテストベッドと、感情コントロールの失敗モードに関するレンズの両方を提供します。以前の研究では、手動で調整された報酬形成によって強化学習 (RL) エージェントに 1 つまたは 2 つの障害を誘発し、事後的に行動にラベルを付け、単一の実行を報告しました。我々は、評価誘導型 PPO エージェントにおける認知評価信号の用量制御可能な操作として障害モデリングを再構築し、7 つの障害 (不安、躁状態、強迫性確認、うつ病、衝動性、依存症、心的外傷後ストレス) をそれぞれ計算精神医学の説明に基づいた単一のノブとして表現し、各症状は認識されたパラダイムにマッピングされた事前登録されたアッセイによって測定されます。 1,000 回を超える実行 (シード 10 個、対照 4 個、95% 信頼区間) にわたって、すべての疾患は、対照では再現されない段階的で単調な用量反応を示しました。これらの誘発された効果を超えて、報酬には書き込まれなかった 3 つの発見が現れます。障害は、躁状態が不安を反映する 2 次元の感情空間に自己組織化します。ノブを取り除くと、報酬歪み障害(躁病、確認、中毒)は軽減されますが、回避障害(不安、PTSD)は軽減されません。代わりに、段階的な曝露カリキュラムの下で回復します。 2 つの同時ノブが非加算的に相互作用し、テスト可能な併存疾患の予測が得られます。したがって、評価重みは、障害を誘発する同じノブがその治療をモデル化できる感情表現型の制御可能な空間をパラメータ化します。また、3 つの障害ノブ (うつ病、依存症、不安) が、標準的な畳み込みエージェントを使用し、評価批評家を持たない 3 次元ピクセル環境 (MiniWorld) に移行することも示し、クロスアッセイ解離が両方のドメインにわたって確認されており、このフレームワークがグリッド ワールドや PPO の評価批評家に固有のものではないことを示しています。

原文 (English)

A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents

Modelling psychological disorders in artificial agents offers both a testbed for computational psychiatry and a lens on the failure modes of affective control. Prior work induces one or two disorders in a reinforcement learning (RL) agent by hand-tuned reward shaping, labels the behaviour post hoc, and reports single runs. We recast disorder modelling as dose-controllable manipulation of cognitive appraisal signals in an appraisal-guided PPO agent, expressing seven disorders (anxiety, mania, obsessive-compulsive checking, depression, impulsivity, addiction, and post-traumatic stress) each as a single knob grounded in a computational psychiatry account, with each symptom measured by a preregistered assay mapped to a recognised paradigm. Across more than a thousand runs (10 seeds, four controls, 95% confidence intervals) every disorder shows a graded, monotone dose-response that no control reproduces. Beyond these induced effects, three findings emerge that were not written into the reward: the disorders self-organise into a two-dimensional affective space in which mania mirrors anxiety; removing a knob remits reward distortion disorders (mania, checking, addiction) but not avoidance disorders (anxiety, PTSD), which instead recover under a graded exposure curriculum; and two simultaneous knobs interact nonadditively, yielding testable comorbidity predictions. Appraisal weights thus parameterise a controllable space of affective phenotypes in which the same knobs that induce a disorder can model its treatment. We also show that three disorder knobs (depression, addiction, anxiety) transfer to a three-dimensional pixel environment (MiniWorld) with a standard convolutional agent and no appraisal critic, with cross-assay dissociation confirmed across both domains, indicating the framework is not specific to grid worlds or to PPO's appraisal critic.

13:00 JSTビジネス/資金調達

深層強化学習の評価と設計パラダイムの原理分析

最も挑戦的なゲームの 1 つに勝利をもたらしたステートアクション価値関数を近似するためのディープ ニューラル ネットワークの利用から始まり、当面の課題のルールを明示することなく問題を解決できるアルゴリズムの進歩に至るまで、強化学習研究は過去 10 年間、目覚ましい科学進歩の中心となってきました。この論文では、この研究の進歩の主要な要素に焦点を当て、強化学習における標準評価と設計パラダイムを分析します。我々は、強化学習におけるスケーリング則の理論的基礎を導入し、強化学習アルゴリズムの漸近的なパフォーマンスが、パフォーマンスのランキングとデータ領域の間に単調な関係を持たないことを示します。私たちは大規模な実験を行っており、その結果は、標準的な設計および評価パラダイムに基づく一連の強化学習研究が誤った結論をもたらしていることを示しています。私たちの分析と結果は、深層強化学習のスケーリング、容量、複雑さに関する中心的な分析を提供します。

原文 (English)

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.

13:00 JST研究/論文

心理学的に根拠のあるラベル構造を使用した脳波ベースの感情認識のためのグラフ正則化深層学習

EEG ベースの感情認識は、メンタルヘルスのモニタリングや感情的なブレイン コンピューター インターフェイスにとって重要ですが、既存の深層学習アプローチでは、感情のクラスを孤立したラベルとして扱い、その心理的な相互依存性を無視することがよくあります。我々は、次元感情理論に基づいてエッジが近接性をエンコードするグラフ内のノードとして感情を概念化するグラフ正則化学習フレームワークを提案します。 3 つの相補的な正則化戦略、グラフ ラベル スムージング (直感的なソフト ラベリング)、グラフ ラプラシアンによるグラフ上の通勤距離 (スペクトル グラフ理論)、およびスライス ワッサーシュタイン ディスタンス (グラフ上の最適な輸送) を、計算の複雑さを増加させる順序で適用します。これらの戦略は、確立された感情トポロジーから逸脱するモデル予測にペナルティを与えます。私たちのフレームワークは、AudioTransformer (純粋なトランスフォーマー)、Conformer (CNN トランスフォーマー ハイブリッド)、DCGNN (因果関係グラフ ニューラル ネットワーク) という 3 つの代表的なバックボーン アーキテクチャにわたって評価され、アーキテクチャに依存しない利点を実証しています。 SEED-IV (4 クラス) および SEED-V (5 クラス) データセットの実験では、一貫した改善が示されています。最良の場合、精度は最大 +5.42%、心理的にありえない誤分類は 39% 減少しました。最終的に、私たちのフレームワークは、標準的なアプローチで達成可能なパフォーマンスの上限を上げるのに役立ちます。コードが公開されます。

原文 (English)

Graph-Regularized Deep Learning for EEG-Based Emotion Recognition with Psychologically-Grounded Label Structure

EEG-based emotion recognition is critical for mental health monitoring and affective brain-computer interfaces, yet existing deep learning approaches often treat emotion classes as isolated labels, ignoring their psychological interdependencies. We propose a graph-regularized learning framework that conceptualizes emotions as nodes in a graph where edges encode proximity based on dimensional emotion theories. We adapt three complementary regularization strategies--Graph Label Smoothing (intuitive soft labeling), Commuting distance on graph via Graph Laplacian (spectral graph theory), and Sliced Wasserstein Distance (optimal transport on graph)--ordered by increasing computational complexity. These strategies penalize model predictions that deviate from the established emotion topology. Our framework is evaluated across three representative backbone architectures: AudioTransformer (pure transformer), Conformer (CNN-transformer hybrid), and DCGNN (causal graph neural network), demonstrating architecture-agnostic benefits. Experiments on SEED-IV (4 classes) and SEED-V (5 classes) datasets show consistent improvements: best case up to +5.42% accuracy and 39% reduction in psychologically implausible misclassifications. Ultimately, our framework help raise the upper bound of performance achievable with standard approaches. Code will be released.

13:00 JSTLLM/生成AI研究/論文

ソルバーから研究へ: 研究フロンティアにおける大規模言語モデル駆動の形式数学

数学用 AI (AI4Math) の最近の開発、特に大規模言語モデル (LLM) 駆動の定理証明器は、インタラクティブ定理証明 (ITP) 言語を介した、明確に定義された数学的問題に対する形式的証明の生成において目覚ましい成功を収めています。しかし、現在のシステムは、新しい定理の発見や未解決の予想の解決など、最先端の研究数学に取り組むには依然として根本的に制限があり、これらの予想はしばしば制限がなく、仕様が不十分であり、複数の抽象化層が関与します。私たちは、AI4Math システムの次の飛躍には、事前定義された問題解決者から、厳密な形式的数学的推論によって最先端の数学的課題に対処できる研究エージェントへの決定的な移行が必要であると主張します。このポジションペーパーでは、データセット、自動形式化、証明合成をカバーするこの分野の系統的なレビューを提供します。さらに重要なのは、数学的研究エージェントとして機能し、データセット、関係構造、数学的探索、ツール エコシステム、人間と AI のコラボレーションにわたる問題を調査し、AI4Math の将来に向けた戦略的ロードマップを概説する際に、既存システムの中核となる限界を特定することです。

原文 (English)

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures, which are often open-ended, under-specified, and involve multiple layers of abstraction. We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges with rigorous formal mathematical reasoning. In this position paper, we provide a systematic review of the field, covering datasets, auto-formalization, and proof synthesis. More importantly, we identify core limitations of existing systems in serving as mathematical research agents, examining issues across datasets, relational structure, mathematical exploration, tool ecosystem, and human-AI collaboration, outlining a strategic road-map for the future of AI4Math.

13:00 JST画像/動画生成

DreamCharacter-1: 3D 生成基盤モデルから製品化可能なキャラクター生成まで

DreamCharacter-1 は、事前トレーニング済み 3D 基礎モデルを高忠実度でプロダクション対応の 3D キャラクター生成に向けて調整する、軽量のポストアダプテーション フレームワークです。 3D 基礎バックボーンに基づいて構築されている当社のパイプラインには、次の 3 つのタスク指向コンポーネントが組み込まれています。(1) ジオメトリ ポストトレーニング。ジオメトリ設定の最適化を通じて、きめの細かい表面の詳細を強化します。 (2) テクスチャ ポストトレーニング。高解像度のテクスチャを合成し、遮蔽された領域の外観を洗練します。 (3) 推論の高速化により、スケーラブルな展開が可能になります。広範な定量的および定性的な実験により、DreamCharacter-1 が視覚的に魅力的で構造的に堅牢な 3D キャラクター アセットを生成し、常に最先端のキャラクター生成方法を上回っていることが実証されています。

原文 (English)

DreamCharacter-1: From 3D Generative Foundation Models to Product-Ready Character Generation

We present DreamCharacter-1, a lightweight post-adaptation framework that calibrates pretrained 3D foundation models toward high-fidelity, production-ready 3D character generation. Building upon a 3D foundation backbone, our pipeline incorporates three task-oriented components: (1) geometry post-training, which enhances fine-grained surface details through geometric preference optimization; (2) texture post-training, which synthesizes high-resolution textures and refines the appearance of occluded regions; and (3) inference acceleration, which enables scalable deployment. Extensive quantitative and qualitative experiments demonstrate that DreamCharacter-1 produces visually compelling and structurally robust 3D character assets, consistently surpassing state-of-the-art character generation methods.

13:00 JSTLLM/生成AIエージェント

トリガーから感情へ: ペルソナベースの対話における動的感情進化のための CPM に基づいた評価マルチエージェント

大規模言語モデル (LLM) には、医療、教育、カウンセリング、顧客サービス、インタラクティブなストーリーテリングにおける感情に敏感な役割シミュレーションのための、実質的に高度なペルソナベースの対話エージェントがあります。しかし、関連する 2 つの業務には大きなギャップが残されています。ペルソナベースの対話システムは、感情を静的な特性や表面レベルの文体の合図としてエンコードすることが多く、感情対話の研究では、エージェントペルソナ自身の進化する感情状態をモデル化するよりも、ユーザーに対する共感的な反応の生成に主に焦点を当ててきました。その結果、キャラクター内のトリガーによって引き起こされる感情の進化は未解明のままです。この制限に対処するために、私たちはコンポーネント プロセス モデル (CPM) からインスピレーションを得ています。CPM は、感情を外部のイベントの評価によって形成される動的なプロセスとみなす心理学理論です。私たちは、ペルソナベースの対話における感情の変化をサポートするための CPM に基づいた感情進化マルチエージェント フレームワークである CPM-MultiAgent を提案します。 CPM-MultiAgent は、キャラクターの感情を固定属性として扱うのではなく、対話トリガーによって継続的に再形成される潜在的な状態として表現します。このフレームワークは、感情トリガーの抽出、CPM ベースの共同評価、および感情状態の更新を通じて、マルチターン インタラクションにおけるより感情的に一貫した役割シミュレーションを可能にします。ベースライン比較、アブレーション研究、人間の評価、症例分析による実験により、CPM-MultiAgent が感情的に敏感な役割シミュレーション設定における動的な感情の進化を効果的にモデル化することが実証されています。

原文 (English)

From Triggers to Emotions: A CPM-Grounded Appraisal Multi-Agent for Dynamic Emotional Evolution in Persona-Based Dialogue

Large Language Models (LLMs) have substantially advanced persona-based dialogue agents for emotion-sensitive role simulation in healthcare, education, counseling, customer service, and interactive storytelling. However, two related lines of work leave a key gap. Persona-based dialogue systems often encode emotions as static traits or surface-level stylistic cues, and affective dialogue research has largely focused on empathetic response generation toward users rather than modeling the agent persona's own evolving emotional state. As a result, trigger-driven emotional evolution within a character remains underexplored. To address this limitation, we draw inspiration from the Component Process Model (CPM), a psychological theory that views emotion as a dynamic process shaped by the appraisal of external events. We propose CPM-MultiAgent, a CPM-grounded emotion evolution multi-agent framework for supporting emotional changes in persona-based dialogue. Instead of treating a character's emotion as a fixed attribute, CPM-MultiAgent represents it as a latent state that is continuously reshaped by dialogue triggers. Through affective trigger extraction, CPM-based collaborative appraisal, and emotion state updating, the framework enables more emotionally consistent role simulation in multi-turn interactions.Experiments with baseline comparisons, ablation studies, human evaluation, and case analyses demonstrate that CPM-MultiAgent effectively models dynamic emotional evolution in emotionally sensitive role-simulation settings.

13:00 JSTエージェントロボティクス研究/論文

シフト&ドリフト: 一般化可能で堅牢な自動運転モーション プランニングのためのゼロショット ベンチマーク

nuPlan などの大規模なオブジェクト レベルのデータセットでトレーニングされた閉ループ モーション プランナーは、強力な分布内 (ID) パフォーマンスを示しますが、新しい都市トポロジーへの一般化や、実行摂動後の回復メカニズムについてはまだ研究が進んでいません。これに対処するために、分布シフトの 2 つの重要な軸にわたってモーション プランナーを厳密にストレス テストするように設計された新しいデュアル トラック ベンチマークである Shift & Drift を紹介します。(1) セマンティック シフト トラックは、空撮の DeepScenario Open 3D データセットを nuPlan シミュレーション フレームワークに変換する新しい変換パイプラインを活用します。これにより、北米とシンガポールのデータに基づいてトレーニングされたプランナーを、ドイツの 4 つの都市と米国のサンフランシスコ市にまたがる歩行者と自転車の密なやりとりを特徴とする 1,182 のシナリオに対してゼロショット評価を行うことが可能になります。 (2) 状態分布ドリフト トラックは、自我車両のダイナミクスに確率的摂動を注入して、複合的な実行エラーに対する堅牢性を定量化します。これに基づいて、意味論的および状態分布の変化の下でのさまざまな計画パラダイムの失敗モードを体系的に評価します。模倣学習手法は ID ベンチマークで高いスコアを達成しますが、セマンティックシフトの下では、特に歩行者が密集した環境では重大な失敗を示し、時間的に相関する作動ノイズにさらされると持続的なドリフトに悩まされます。対照的に、評価された強化学習ベースのプランナーは、より緩やかな劣化を示し、両方のトラックにわたってより高い安全性と進行状況の指標を維持します。私たちの調査結果は、模倣の忠実度と閉ループの回復力の間の経験的なトレードオフを明らかにし、信頼性の高い展開に向けた進捗状況を評価するための厳密なベンチマークをコミュニティに提供します。

原文 (English)

Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning

While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) performance, their generalization to novel urban topologies and recovery mechanisms following execution perturbations remain under-explored. To address this, we present Shift & Drift, a novel dual-track benchmark designed to rigorously stress-test motion planners across two critical axes of distribution shift: (1) The Semantic Shift Track leverages a novel conversion pipeline that transforms the aerial, DeepScenario Open 3D dataset into the nuPlan simulation framework. This enables zero-shot evaluation of planners trained on North American and Singaporean data against 1,182 scenarios spanning four German cities and the US city of San Francisco featuring dense pedestrian-cyclist interactions. (2) The State-Distribution Drift Track injects stochastic perturbations into the ego vehicle's dynamics to quantify robustness against compounding execution errors. Based on this, we systematically evaluate the failure modes of diverse planning paradigms under semantic and state-distribution shifts. While imitation learning methods achieve high scores in ID benchmarks, they exhibit significant failures under semantic shift, particularly in pedestrian-dense environments, and suffer from persistent drift when subjected to temporally correlated actuation noise. In contrast, the evaluated reinforcement-learning-based planner demonstrates more graceful degradation, maintaining higher safety and progress metrics across both tracks. Our findings reveal an empirical trade-off between imitation fidelity and closed-loop resilience, providing the community with a rigorous benchmark to evaluate progress toward reliable deployment.

13:00 JST研究/論文

古典力学の基礎における 3 つの未解決の問題のキメ表現による定式化: 不確実性、不変エントロピー、方向自由度

古典力学の基礎からの 3 つの未解決の問題を、複素時間 (キメ) 表現で数学的に自己完結した定式化を示します。(I) 古典エントロピー不確定性原理の非正準変数および複数自由度への拡張。 (II) 座標不変の測度およびエントロピーの特徴付け、つまり、不変エントロピーが存在するためにはなぜ連続物理量が対にならなければならないのかという問題。 (III) 古典的相対論的方向自由度 (スピン 1/2 系の古典的類似物) の構築。全体を通して、キメ段階は、{統計的には潜在的な循環確率変数として解釈され、その法則 \Phi は、カイムの大きさによって指標付けされた、繰り返された同一に制御された実験の固有の試験ごとの変動性をモデル化します。数学的ブリッジは、キメ コーンを 1 自由度の位相空間の作用角チャートと正確にシンプレクティックに同定したもので、その下ではキメ測度はリウヴィル測度であり、位相法則はリウーヴィル密度の角度条件になります。具体的には、(i) フォン・ミーゼスの法則によって正確に飽和された鋭い循環フィッシャー情報不等式とともに、その極値族がフォン・ミーゼス x ガウスであるキメ円柱上の鋭いエントロピー不確実性関係を証明します。 (ii) 補正項がポアソン ブラケットの幾何平均である正確な非正準的な不確実性関係を証明し、予想されるブラケットの推測される役割を明確にする。 (iii) ウィリアムソン正規形とフィッシャーの不等式を介して集約多自由度境界を証明し、シンプレクティックなシュール・ホーン型の正確な未解決問題として自由度ごとの洗練を分離します。 (iv) キメ相の拡散が等分配 (Haar 均一) 位相法則に従って単調なエントロピー増加を引き起こすことを証明します。

原文 (English)

Kime-Representation Formulations of Three Open Problems in the Foundations of Classical Mechanics: Uncertainty, Invariant Entropy, and Directional Degrees of Freedom

We give mathematically self-contained formulations, in the complex-time (kime) representation, of three open problems from the foundations of classical mechanics: (I) the extension of the classical entropic uncertainty principle to non-canonical variables and to multiple degrees of freedom; (II) the characterization of coordinate-invariant measures and entropies, i.e., the question of why continuous physical quantities must be paired for an invariant entropy to exist; and (III) the construction of a classical relativistic directional degree of freedom (a classical analogue of a spin-1/2 system). Throughout, the kime phase is interpreted {statistically as a latent circular random variable whose law \Phi models the intrinsic trial-to-trial variability of repeated, identically controlled experiments indexed by the kime magnitude. The mathematical bridge is an exact symplectic identification of the kime cone with the action-angle chart of a one-degree-of-freedom phase space, under which the kime measure is the Liouville measure and the phase law becomes the angular conditional of a Liouville density. Specifically, we (i) prove a sharp entropic uncertainty relation on the kime cylinder whose extremal family is von Mises x Gaussian, together with a sharp circular Fisher-information inequality saturated exactly by von Mises laws; (ii) prove an exact non-canonical uncertainty relation in which the correction term is the geometric mean of the Poisson bracket, clarifying the conjectured role of the expected bracket; (iii) prove aggregate multi-degree-of-freedom bounds via the Williamson normal form and Fischer's inequality, and isolate the per-degree-of-freedom refinement as a precise open problem of symplectic Schur-Horn type; (iv) prove that diffusion of the kime phase produces monotone entropy growth with the equipartitioned (Haar-uniform) phase law.

13:00 JSTエージェント研究/論文

テンソルネットワーク理論のマルチエージェント自動形式化

私たちは、特化した大規模言語モデル エージェントのチームを構築し、行列積状態の基本定理の自動形式化をデモンストレーションとして、理論物理学における研究レベルの形式化のためのエージェント駆動のワークフローを提示します。エージェントは、構造化された数学的青写真と人間による定期的なレビューを通じて調整され、完全な形式化を自律的に調整して実行しました。一部の陳述については、捜査員は標準文献にはない新たな証拠ルートを探索することができた。その過程で、エージェントは、リーンの数学ライブラリである Mathlib では以前は利用できなかった広範なテンソル ネットワークと量子情報ライブラリを作成しました。物理的な応用として、この形式化は、一次元での対称性が保護されたトポロジカル位相にも拡張されます。私たちは、大規模な自動形式化における主なボトルネックが数学的意図を強制することであることを発見し、プロセス全体と関連するさまざまな微妙な点についての詳細な研究を提供します。私たちはコードベースをライブラリ \href{https://github.com/LionSR/TNLean}{TNLean} としてリリースし、形式化作業の \nChapters{} 章 \href{https://lionsr.github.io/TNLean/blueprint/}{blueprint} とともにリリースします。

原文 (English)

Multi-agent Autoformalization of Tensor Network Theory

We build a team of specialized large language-model agents and present an agent-driven workflow for research-level formalization in theoretical physics, with the autoformalization of the fundamental theorem of matrix-product states as a demonstration. The agents, coordinated through a structured mathematical blueprint and periodic human review, orchestrated and executed the full formalization autonomously. For some statements, the agents were able to explore new proof routes that are not part of the standard literature. Along the way the agents produced extensive tensor-network and quantum-information libraries not previously available in Mathlib, Lean's mathematical library. As a physical application, the formalization also extends towards symmetry-protected topological phases in one dimension. We find that the main bottleneck in large-scale autoformalization is enforcing mathematical intent and we provide a detailed study of the full process and various subtleties involved. We release the codebase as the library \href{https://github.com/LionSR/TNLean}{TNLean}, together with a \nChapters{}-chapter \href{https://lionsr.github.io/TNLean/blueprint/}{blueprint} of the formalization effort.

13:00 JST画像/動画生成エージェントロボティクス

非構造化環境におけるロボットの事前学習済み視覚モデルを使用した、衝突までの時間ベースの動的障害物回避

構造化されていない屋外環境における動的障害物回避は、特に大規模なロボット固有のトレーニング データやシミュレーション ベースのポリシーが非現実的である場合、自律移動ロボットにとって依然として重要な課題です。我々は、完全に現実世界のデータに基づいて動作し、シミュレーションで訓練されたポリシーに固有のシミュレーションから現実への転送問題を回避する、ビジョンベースの動的障害物回避のための、データ効率が高く解釈可能な方法を提案します。私たちのアプローチは、大規模な事前トレーニング済み単眼奥行き推定モデルである UniDepth を利用して、推論時にステレオ カメラや LiDAR を必要とせずに、RGB ビデオから高密度の奥行きマップを生成します。動的な障害物回避は、長いフレーム シーケンス全体でキーポイントを追跡するように SuperPoint および SuperGlue 機能の対応パイプラインを拡張し、カメラの組み込み関数と予測深度を使用して 2D ピクセル空間の位置を 3D に投影し、これらの 3D キーポイントから初期化されたバンドル調整を実行し、キーポイントごとの衝突時間 (TTC) を計算することによって実現されます。次に、最小 TTC キーポイントの最近接点からロボットを遠ざけるために、地表面の 2D モーション プリミティブが選択されます。 M3ED データセットの実世界データに基づいて評価されたこのパイプラインは、グラウンド トゥルース TTC が 1 秒未満のフレームの識別において精度 0.49 と再現率 0.38 を達成し、真陽性検出の 84\% で回避動作の方向を正しく生成します。重要なのは、テスト シーケンスに存在する 22 個の固有の物理的障害物のうち 20 個について、TTC が 1 秒未満のフレームを少なくとも 1 つ検出することです。何千時間ものロボット固有のトレーニング データを必要とするエンドツーエンドの学習方法とは異なり、私たちのアプローチではモデル トレーニングが完全に不要になり、ハイパーパラメーター調整に必要なデータは 74 秒のみです。これにより、さまざまな種類の障害物にわたって解釈可能かつ一般化可能な動作を維持しながら、優れたデータ効率が実証されます。

原文 (English)

Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments

Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impractical. We present a data-efficient, interpretable method for vision-based dynamic obstacle avoidance that operates entirely on real-world data, avoiding the sim-to-real transfer problem inherent in simulation-trained policies. Our approach leverages UniDepth, a large pretrained monocular depth estimation model, to produce dense depth maps from RGB video without requiring stereo cameras or LiDAR at inference time. Dynamic obstacle avoidance is achieved by extending the SuperPoint and SuperGlue feature correspondence pipeline to track keypoints across long frame sequences, projecting their 2D pixel-space positions into 3D using camera intrinsics and predicted depth, running bundle adjustment initialized from these 3D keypoints, and computing per-keypoint time-to-collision (TTC). A 2D motion primitive in the ground plane is then selected to move the robot away from the closest point of approach of the minimum-TTC keypoint. Evaluated on real-world data from the M3ED dataset, our pipeline achieves a precision of 0.49 and a recall of 0.38 in identifying frames with a ground truth TTC below 1 second, and correctly generates the evasive motion direction in 84\% of true positive detections. Crucially, it detects at least one frame with TTC less than 1 second for 20 out of 22 unique physical obstacles present in our test sequences. Unlike end-to-end learned methods that demand thousands of hours of robot-specific training data, our approach eliminates model training entirely, requiring only 74 seconds of data for hyperparameter tuning. This demonstrates exceptional data efficiency while preserving interpretable and generalizable behavior across diverse obstacle types.

13:00 JSTLLM/生成AI

次に何を言うべきかをどうやって知ることができますか?ハリス統合主義の強化としてのバレンホルツの自己生成理論

ロイ・ハリスの統合主義言語学は、言語への計算的アプローチの中心に深く埋め込まれている参照主義の伝統に対する説得力のある批判を提供し、言語はあらかじめ与えられた世界にマッピングされるコードではなく、将来の共同行動を指向した状況に応じた二部構成の活動であると主張しています。しかし、統合主義には説明上のギャップがいくつか残されている。統合主義は、記号が将来性のある開放性を維持する構造的メカニズムを完全には説明しておらず、言語的記号活動と非言語的記号活動の間の連続性を過小理論化しており、過去の統合の蓄積されたアーカイブの構造的特性についての詳細な説明を提供していない。この論文は、大規模言語モデル (LLM) の動作に応じて開発された、エラン・バレンホルツの言語の自動生成理論がこれらのギャップを正確に埋め、統合主義の中核となるコミットメントを損なうことなく統合主義を豊かにすることができると主張します。具体的には、自動生成アカウントは次のことを提供します。ハリスが二者間のコミュニケーションの中心であると特定した、将来のオープン性のための構造的なメカニズム。言語と他の記号形成活動の間の記号論的連続性に関するハリスの理論の計算上の相関関係。そしてアーカイブの理論:過去の統合の蓄積された残渣がどのようなものであり、新しい参加者がそれをどのように利用するのか。この統合は、統合主義自体が提供しない説明的な内容を追加しながら、状況に応じた統合行為についてのハリスの存在論的優位性を維持します。自然言語処理や大規模言語モデル設計の実践者や研究者にとって、この議論は、LLM が非常に効果的に活用する統計構造が実際にどのようなものであるか、またその性質上提供できないものについて、原則に基づいた説明を提供します。

原文 (English)

How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism

Roy Harris's Integrationist linguistics offers a compelling critique of the referentialist tradition embedded deep at the heart of computational approaches to language, arguing that language is not a code that maps onto a pre-given world but a situated, bipartite activity oriented toward prospective joint action. Yet Integrationism leaves certain explanatory gaps: it does not fully account for the structural mechanism by which signs sustain prospective openness, it undertheorises the continuity between linguistic and non-linguistic semiotic activity, and it offers no detailed account of the structural properties of the accumulated archive of past integrations. This paper argues that Elan Barenholtz's autogenerative theory of language, developed in response to the behaviour of Large Language Models (LLMs), can fill precisely these gaps, enriching Integrationism without undermining any of its core commitments. Specifically, the autogenerative account provides: a structural mechanism for the prospective openness that Harris identifies as central to bipartite communication; a computational correlate for Harris's thesis of semiotic continuity between language and other sign-making activity; and a theory of the archive: what the accumulated residue of past integrations looks like and how new participants draw upon it. The synthesis preserves Harris's ontological primacy of the situated integrative act while adding explanatory content that Integrationism itself does not supply. For practitioners and researchers in natural language processing and large language model design, the argument offers a principled account of what the statistical structure that LLMs so effectively exploit actually is, and of what it cannot, by its nature, provide.

13:00 JST研究/論文Claude

高木・菅野ファジィ推論を使用したプライベートサブストレートブロックチェーンにおける閉ループ動的バリデーターノードスケーリング

プライベート ブロックチェーン ネットワークは、変化するワークロード条件に適応できない固定ノード構成で実行されます。軽いワークロードを処理するノードが多すぎると、リソースが無駄になります。高い需要に直面しているノードが少なすぎると、ブロックの生成が遅くなり、ファイナライゼーションが低下します。正しいバリデータ数は、時間の経過とともに変化する重複する要因に依存するため、判断するのが困難です。この論文では、ライブ ブロックチェーン パラメーター (ブロック生成時間、ブロック サイズ、アクティブ ノード数) を読み取り、スケールアップ、維持、またはスケールダウンといったスケーリング推奨事項とともに継続的な効率スコアを出力する、Takagi-Sugeno (TS) ファジー推論システムについて説明します。コントローラーは、3 つの言語変数にわたって三角メンバーシップ関数を使用し、積 t ノルム集計を備えた完全な 27 ルール ベースを通じて評価されます。主な貢献は、メンバーシップ関数の経験的な再調整であり、言語用語を理論上の極端ではなく、テストベッドの観察された動作範囲に固定します。このシステムは、クイーンズランド州政府のオープン データ ポータルからの実際のスマート水道メーター データ ハッシュを保存する 10 ノードの Substrate ブロックチェーン ネットワーク上で評価されます。 4、7、および 10 個のアクティブ ノードの構成にわたる統計分析により、コントローラーが各構成のプロビジョニング状態を反映する個別の動作プロファイルを生成することが確認されています。閉ループ実験では、コントローラーが両方向でバリデーターの参加を自律的に調整し、負荷が上昇するとバリデーターをアクティブ化し、オーバープロビジョニングではバリデーターを削除し、両方向から同じ安定した平衡状態に収束します。 3 つのしきい値ベースのベースラインと比較すると、同等のブロック生成時間を維持しながら、スケーリングの変動が少ないことがわかります。結果は、TS ファジー推論が、安定したスケーリング動作のしきい値アプローチでは一致できない、プライベート ブロックチェーン デプロイメントにおける自律的なバリデーター管理をサポートできることを示しています。

原文 (English)

Closed-Loop Dynamic Validator Node Scaling in Private Substrate Blockchains Using Takagi-Sugeno Fuzzy Inference

Private blockchain networks run with fixed node configurations that cannot adapt to changing workload conditions. Too many nodes serving a light workload waste resources; too few nodes facing heavy demand slow block production and degrade finalisation. The right validator count is hard to determine, as it depends on overlapping factors that shift over time. This paper presents a Takagi-Sugeno (TS) fuzzy inference system that reads live blockchain parameters (block production time, block size, and active node count) and outputs a continuous efficiency score alongside a scaling recommendation: Scale Up, Maintain, or Scale Down. The controller uses triangular membership functions across three linguistic variables, evaluated through a complete 27-rule base with product t-norm aggregation. A key contribution is an empirical recalibration of the membership functions, anchoring linguistic terms to the observed operating range of the testbed rather than to theoretical extremes. The system is evaluated on a 10-node Substrate blockchain network storing real smart water meter data hashes from the Queensland Government open data portal. Statistical analysis across configurations of 4, 7, and 10 active nodes confirms that the controller produces distinct operational profiles reflecting each configuration's provisioning state. In closed-loop experiments, the controller autonomously adjusts validator participation in both directions, activating validators under rising load and removing them under over-provisioning, converging to the same stable equilibrium from both directions. Compared against three threshold-based baselines, it shows fewer scaling oscillations while maintaining comparable block production times. Results show that TS fuzzy inference can support autonomous validator management in private blockchain deployments, with stable scaling behaviour threshold approaches cannot match.

13:00 JSTLLM/生成AI

内部アトリビューション グラフによる LLM ジェイルブレイクのメカニズムの解釈可能性

大規模言語モデル (LLM) は優れた機能を示しますが、依然として敵対的なプロンプトや脱獄攻撃に対して非常に脆弱です。既存のアプローチは主に、入出力動作または原因特定方法を通じてこれらの失敗を分析し、敵対的な摂動がモデルの内部推論をどのように変更するかについての洞察を限定的に提供します。その結果、危険な動作や不正な動作の根底にあるメカニズムはほとんど理解されていないままです。我々は、ペアの内部計算グラフを使用して、LLM 脆弱性を診断するためのメカニズムフレームワークを導入します。このグラフは、潜在機能間の構造化された因果関係としてプロンプト固有の推論を表します。クリーンなプロンプトと攻撃されたプロンプトの計算グラフを構築して調整することにより、敵対的攻撃が、安全関連コンポーネントの抑制、攻撃固有の機能の出現、計算パスの再ルーティングなど、内部推論の体系的な変換を誘発することを明らかにします。この表現に基づいて、我々は、(i) 計算を不変構造、抑制構造、および創発構造に分解し、(ii) 障害モードに関連する繰り返しの脆弱性モチーフを特定し、(iii) ノード、パス、およびサブグラフに対して因果的介入を実行して、攻撃の成功への寄与を直接評価する統一フレームワークを提案します。これにより、モデルの失敗の記述的帰属から因果関係の診断への移行が可能になります。複数のオープンソース LLM とさまざまな敵対的および脱獄ベンチマークにわたる実験により、内部計算グラフの構造的逸脱が安全でない動作と強く相関していることが実証されました。さらに、特定された脆弱性モチーフに対する的を絞った介入によりモデルの堅牢性が向上し、LLM 脆弱性を理解、診断、軽減するための原則的な基盤として内部計算グラフが確立されます。

原文 (English)

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial perturbations alter the model's internal reasoning. Consequently, the mechanisms underlying unsafe or incorrect behaviors remain poorly understood. We introduce a mechanistic framework for diagnosing LLM vulnerabilities using paired internal computation graphs, which represent prompt-specific inference as structured causal interactions among latent features. By constructing and aligning computation graphs for clean and attacked prompts, we reveal that adversarial attacks induce systematic transformations of internal reasoning, including suppression of safety-relevant components, emergence of attack-specific features, and rerouting of computation paths. Building on this representation, we propose a unified framework that (i) decomposes computation into invariant, suppressed, and emergent structures, (ii) identifies recurring vulnerability motifs associated with failure modes, and (iii) performs causal interventions on nodes, paths, and subgraphs to directly evaluate their contributions to attack success. This enables a transition from descriptive attribution to causal diagnosis of model failures. Experiments across multiple open-source LLMs and diverse adversarial and jailbreak benchmarks demonstrate that structural deviations in internal computation graphs strongly correlate with unsafe behaviors. Furthermore, targeted interventions on identified vulnerability motifs improve model robustness, establishing internal computation graphs as a principled foundation for understanding, diagnosing, and mitigating LLM vulnerabilities.

13:00 JSTLLM/生成AI研究/論文

視覚、言語、ビデオ、オーディオにわたるマルチモーダルなアンラーニング: 手法、データセット、ベンチマークの調査

VLM、DM、LLM、AFM の採用が進むにつれて、これらのマルチモーダル基盤モデルは、トレーニング データに由来する機密性、著作権で保護された、偏った、または安全でないクロスモーダル関連を誤ってエンコードする可能性があります。削除リクエストやポリシー更新後の再トレーニングは多くの場合非現実的であり、知識が共有表現全体に分散されているため、対象を絞った忘却は依然として困難です。マルチモーダルアンラーニングは、全体的なユーティリティを維持しながらモダリティ全体で選択的な除去を可能にすることで、この課題に対処します。この調査は、最近の進歩、新たなアプリケーション、未解決の問題に基づいて、視覚、言語、オーディオ、ビデオにわたるマルチモーダルなアンラーニングに関する統一されたシステム指向のビューを提供します。私たちの分類法により、モデルのアーキテクチャとモダリティ全体の体系的な比較が可能になり、削除の強度、保持、効率、可逆性、堅牢性の間のトレードオフが明確になります。この調査は、未解決の問題と、マルチモーダルアンラーニングの将来の研究と展開をサポートするための実際的な考慮事項に焦点を当てています。厳選されたリポジトリをリリースします: https://smsnobin77.github.io/Awesome-Multimodal-Unlearning/

原文 (English)

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data. Retraining after deletion requests or policy updates is often impractical, and targeted forgetting remains difficult because knowledge is distributed across shared representations. Multimodal unlearning addresses this challenge by enabling selective removal across modalities while retaining overall utility. This survey offers a unified, system-oriented view of multimodal unlearning across vision, language, audio, and video, grounded in recent advances, emerging applications, and open problems. Our taxonomy enables systematic comparison across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. This survey highlights open problems and practical considerations to support future research and deployment of multimodal unlearning. We release a curated repository: https://smsnobin77.github.io/Awesome-Multimodal-Unlearning/

13:00 JSTLLM/生成AI研究/論文

潜在的な性格特性による言語モデルの効率的な安全性調整

大規模な言語モデルに対する現在の安全方法は、敵対的な攻撃に対して脆弱であることが知られており、堅牢な代替方法の研究が行われています。潜在的敵対的トレーニング (LAT) は最も効果的な防御策の 1 つですが、実用性が低下する可能性があり、有害なプロンプトの大規模なデータセットでのトレーニングが必要です。私たちは、潜在的性格調整(LPA)を導入します。これは、心理測定的性格文献から抽出されたわずか 66 個の危害を無視したステートメントを対象とした、明示的な危害の拒否を敵対的なトレーニングに置き換えます。私たちは、人格にアンカーされた表現は危害回避と潜在的な構造を共有しているため、それらを敵対的に安定させることで、脱獄攻撃によって悪用される部分空間を暗黙的に制限すると仮説を立てています。 LPA は、トレーニング中に有害なコンテンツがまったく表示されず、標準ベンチマークでパフォーマンスが低下しないにもかかわらず、直接リクエストと 5 つの脱獄方法にわたって HarmBench でほぼゼロの攻撃成功率を達成します。さらに、トレーニングプロセスは軽量です。手順全体は 1 つの GPU で数分で完了し、使用するサンプルの数は標準 LAT よりも 75 分の 1 です。広範なアブレーションは、私たちの方法の堅牢性、効率性、および一般化を示しています。

原文 (English)

Efficient Safety Alignment of Language Models via Latent Personality Traits

Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large datasets of harmful prompts. We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We hypothesize that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks. LPA achieves near-zero attack success rates on HarmBench across direct requests and five jailbreak methods, despite never seeing harmful content during training and no loss of performance on standard benchmarks. Moreover, the training process is lightweight; the entire procedure completes in minutes on a single GPU and uses 75x fewer examples than standard LAT. Extensive ablations demonstrate the robustness, efficiency, and generalization of our method.

13:00 JST画像/動画生成

敵対的デコイ: ViT における注意ベースの防御を誤った方向に誘導する

ビジョン トランスフォーマー (ViT) は、局所的な敵対的攻撃 (敵対的パッチなど) に対して依然として脆弱ですが、最近のテスト時防御では、異常に高いアテンション スコアを持つ画像トークンを抑制することで攻撃を軽減しています。これらの防御は、注目と敵対的有効性の間の強い結びつきを利用します。多くの場合、敵対的トークンは、予測に影響を与えるためにかなりの注目を集める必要があります。私たちは、選択されたターゲット トークンに注意を向け、したがって関連する防御をリダイレクトする、独立して最適化されたイメージ パッチである敵対的デコイを導入します。誤分類と防御回避を一緒に最適化するのではなく、私たちのアプローチは 2 つの目的を切り離します。元の敵対領域が誤った予測を誘発し、別のおとりが防御側が使用する注目ランキングを操作します。レイヤーごとの目標は、ターゲット トークンへの注目を高め、これらのトークンを競合する非ターゲット トークンよりも優先させます。おとりは根本的な攻撃とは独立して最適化されるため、この方法は攻撃に依存せず、既存の敵対的パッチ攻撃と簡単に統合できます。複数の ViT アーキテクチャと攻撃にわたる ImageNet の実験では、おとりが攻撃の有効性の多くを維持しながら、高い注意スコアを真の敵対領域から遠ざけることができることを示しています。これらの結果は、敵対的関連性の指標として注目の大きさを使用することの根本的な限界を明らかにしています。

原文 (English)

Adversarial Decoys: Misdirecting Attention-Based Defenses in ViT

Vision Transformers (ViTs) remain vulnerable to localized adversarial attacks, e.g., adversarial patches, while recent test-time defenses mitigate them by suppressing image tokens with abnormally high attention scores. These defenses exploit a strong coupling between attention and adversarial effectiveness: adversarial tokens often need to attract substantial attention to influence the prediction. We introduce adversarial decoys, independently optimized image patches that redirect the attention, and therefore related defenses, toward selected target tokens. Rather than jointly optimizing misclassifications and defense evasion, our approach decouples the two objectives: the original adversarial region induces the incorrect prediction, while a separate decoy manipulates the attention ranking used by the defense. A layer-wise objective increases target-token attention and promotes these tokens above competing non-target ones. Since the decoy is optimized independently of the underlying attack, the method is attack-agnostic and can be easily integrated with any existing adversarial patch attack. Experiments on ImageNet across multiple ViT architectures and attacks show that decoys can redirect high attention scores away from the true adversarial region while preserving much of the attack effectiveness. These results reveal a fundamental limitation of using attention magnitude as an indicator of adversarial relevance.

13:00 JST研究/論文

path_boost: パスベースの勾配ブースティングを使用した解釈可能なグラフレベル予測のための Python パッケージ

グラフ構造の入力データに対する解釈可能な教師あり学習のための Python パッケージ、path_boost を紹介します。このパッケージには、学習プロセス中にグラフ内の予測ラベル付きパスを自動的に検出する勾配ブースティング アルゴリズムである PathBoost が実装されています。一般に解釈が難しいグラフ ニューラル ネットワークとは異なり、PathBoost はパスベースの特徴に対して相加的予測モデルを生成し、どの部分構造が予測を駆動するかを明示的に明らかにします。考えられるすべてのパスを網羅的に列挙することを避けるため、アルゴリズムは、学習中に予測力に基づいてパスを繰り返し選択および拡張し、ブースティングを使用して弱い学習器を強力なアンサンブルに結合します。このパッケージは、回帰と二項分類の両方をサポートしています。主な機能には、scikit-learn ワークフローとの互換性、カスタムのベース学習器とセレクターのサポート、自動開始ノード選択、アンカー ノード間の並列トレーニング、組み込みの変数重要度計算が含まれます。私たちは、原子がノードとして機能し、結合がエッジとして機能する遷移金属化合物の分子特性予測に関する PathBoost を実証し、さらに 6 つの分子データセットにわたる確立されたグラフ ニューラル ネットワークおよびグラフ カーネル法に対して PathBoost のベンチマークを行います。このパッケージは、オープンソース ライセンスに基づいて PyPI および GitHub で入手できます。

原文 (English)

path_boost: A Python Package for Interpretable Graph-Level Prediction using Path-Based Gradient Boosting

We present path_boost, a Python package for interpretable supervised learning on graph-structured input data. The package implements PathBoost, a gradient boosting algorithm that automatically discovers predictive labeled paths within graphs during the learning process. Unlike graph neural networks, which are generally difficult to interpret, PathBoost produces an additive prediction model over path-based features that explicitly reveals which substructures drive predictions. To avoid an exhaustive enumeration of all possible paths, the algorithm iteratively selects and extends paths during learning based on their predictive power, using boosting to combine weak learners into a strong ensemble. The package supports both regression and binary classification. Key features include compatibility with scikit-learn workflows, support for custom base learners and selectors, automatic starting node selection, parallel training across anchor nodes, and built-in variable importance computation. We demonstrate PathBoost on molecular property prediction of transition metal compounds, where atoms serve as nodes and bonds as edges, and further benchmark PathBoost against an established graph neural network and a graph kernel method across six molecular datasets. The package is available on PyPI and GitHub under an open-source license.

13:00 JST研究/論文

リニア アテンション アーキテクチャ: メカニズム、トレードオフ、およびクロスレイヤー ルーティング

セルフアテンションにより、各トークンは完全なコンテキストから情報を取得できますが、シーケンスの長さの二次コストが長いコンテキストでのトレーニングと推論を制限します。この論文では、ソフトマックス アテンションと 4 つの最近のリカレント リニア アテンション アーキテクチャ (DeltaNet、Gated DeltaNet、Kimi Delta Attendance、および Gated DeltaNet-2) の比較研究を紹介します。これらのメカニズムを一般的なリカレント メモリ表記法で表現し、表現力、メモリ減衰、消去と書き込み制御、トレーニング スループット、実装の複雑さの点でどのように異なるかを明示します。私たちの実験は、15B トークン用にトレーニングされた 350M パラメータ モデルを中心としており、オプティマイザと学習率の比較、ハイブリッド スタックとピュア スタックの比較、シーケンス長のランタイム測定、1.3B および 3B パラメータでの大規模な DeltaNet 実行、および少数のダウンストリーム評価が含まれます。報告された速度結果は、トレーニングのスループットと反復時間を測定します。経験に基づく推論速度のベンチマークは提供しません。報告されている 350M パラメータ、15B トークン スイープ内で、Muon を使用した キミ デルタ アテンションは最低の最終検証損失に達し、AdamW でトレーニングされた純粋な Gated DeltaNet スタックは最も高い正規化トレーニング スループットを持ち、ハイブリッド スタックは一般にスループット コストで損失を改善し、評価した一致したアーキテクチャ設定において Muon は AdamW と比較して最終検証損失を一貫して下げています。 DeltaNet スタイルのメモリ用の軽量クロスレイヤー ルーティング メカニズムを導入し、評価します。デルタネットにヒントを得た最も自然な定式化では、下位層のデルタ ルール書き込みエラーを次の層の値ターゲットに転送しますが、一致したベースラインを超えて改善されません。代わりに、整列された非表示ストリームにルーティングし、書き込み値を転送すると、私たちが報告する一致した実行にわずかな改善が見られます。クロスレイヤー値ルーティング (CLVR) により、DeltaNet と Gated DeltaNet の両方の最終検証損失が減少します。

原文 (English)

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity. Our experiments center on 350M-parameter models trained for 15B tokens, and include optimizer and learning-rate comparisons, hybrid-versus-pure stack comparisons, sequence-length runtime measurements, larger DeltaNet runs at 1.3B and 3B parameters, and a small set of downstream evaluations. The reported speed results measure training throughput and iteration time; we do not provide an empirical inference-speed benchmark. Within the reported 350M-parameter, 15B-token sweep, Kimi Delta Attention with Muon reaches the lowest final validation loss, a pure Gated DeltaNet stack trained with AdamW has the highest normalized training throughput, hybrid stacks generally improve loss at a throughput cost, and Muon consistently lowers final validation loss relative to AdamW in the matched architecture settings we evaluate. We introduce and evaluate lightweight cross-layer routing mechanisms for DeltaNet-style memories. The most natural DeltaNet-inspired formulation, forwarding a lower layer's delta-rule write error into the next layer's value target, does not improve over matched baselines. Routing into the aligned hidden stream and forwarding the write value instead yields a modest improvement in the matched runs we report: Cross-Layer Value Routing (CLVR) lowers final validation loss for both DeltaNet and Gated DeltaNet.

13:00 JST画像/動画生成

熱イメージングを超えて: 時間分解熱観測からの熱物理特性の推測

感覚観察から潜在的な物理的特性を推測することは、機械の知覚における基本的な課題です。利用可能なセンシングモダリティの中で、熱イメージングは​​特に有望です。これは、温度の変化が熱伝達物理学によって直接支配され、したがってシーンの根底にある熱物理特性に関する情報がエンコードされるためです。熱観測から空間的に分解された熱物理的特性を回復すると、デジタル ツインやインフラストラクチャの監視からロボット工学や科学イメージングに至るまで、さまざまなアプリケーションが変革される可能性があります。ただし、既存の熱シーン再構成手法は、熱進化を支配する熱物理学的特性を特定することなく、複雑な 3D 環境の温度場を復元できます。一方、逆手法は物理的に解釈可能なパラメータ推定を提供しますが、通常は単純化された形状と制御された実験条件に依存します。ここでは、微分可能な熱伝達シミュレーションによる熱シーンの再構成と熱物理パラメータ推定を統合するフレームワークである ThermoField を紹介します。提案されたフレームワークは、これらの量を空間的に変化する神経フィールドとして表し、シーンの幾何学、支配的な熱伝達物理学、および時間的熱観測を通じてそれらの量を制約します。私たちは、ThermoField が幾何学形状を共同で再構築し、空間的に変化する熱拡散率を推定し、これまで見たことのない環境条件下での熱の進化を予測することを実証します。このフレームワークは、ニューラル シーン表現を微分可能な熱伝達ソルバーと統合することにより、複雑な 3D シーンで物理的に解釈可能なパラメーター推論を可能にします。私たちの結果は、熱シーンの再構築と逆伝熱解析の間の橋渡しを確立し、形状の再構築、熱物理特性の推定、熱観測からの予測熱シミュレーションのための統一されたアプローチを提供します。

原文 (English)

Beyond Thermal Imaging: Inferring Thermophysical Properties from Time-Resolved Thermal Observations

Inferring latent physical properties from sensory observations is a fundamental challenge in machine perception. Among available sensing modalities, thermal imaging is particularly promising because temperature evolution is directly governed by heat-transfer physics and therefore encodes information about underlying thermophysical properties of a scene. Recovering spatially resolved thermophysical properties from thermal observations could transform applications ranging from digital twins and infrastructure monitoring to robotics and scientific imaging. However, existing thermal scene reconstruction methods can recover temperature fields in complex 3D environments without identifying the thermophyiscal properties that govern thermal evolution, whereas inverse methods provide physically interpretable parameter estimation but typically rely on simplified geometries and controlled experimental conditions. Here we introduce ThermoField, a framework that unifies thermal scene reconstruction and thermophysical parameter estimation through differentiable heat-transfer simulation. The proposed framework represents these quantities as spatially varying neural fields and constrains them through scene geometry, governing heat-transfer physics, and temporal thermal observations. We demonstrate that ThermoField jointly reconstructs geometry, estimates spatially varying thermal diffusivity, and predicts thermal evolution under previously unseen environmental conditions. By integrating neural scene representations with differentiable heat-transfer solver, the framework enables physically interpretable parameter inference in complex 3D scenes. Our results establish a bridge between thermal scene reconstruction and inverse heat-transfer analysis, providing a unified approach for geometry reconstruction, thermophysical property estimation, and predictive thermal simulation from thermal observations.

13:00 JSTLLM/生成AI

MiniLM 埋め込みによるスコープ外の意図検出のためのマルチクラスター境界学習方法

意図の検出は、人間とマシンの対話システムにおける人間の意図とシステムのアクションの橋渡しをする重要なタスクです。ただし、範囲外 (OOS) インテントの検出には依然として課題が存在します。 (i) 従来の方法では、OOS インテントの検出をマルチクラス分類として見なしているため、既知のインテントのクラス数が増加するにつれて検出精度が低下します。 (ii) LLM 埋め込み手法には大きなパラメータが必要なため、トレーニングや実際の展開が困難になります。したがって、この研究では、1 クラス分類ワークフローで MiniLM 埋め込み (つまり、すべて MiniLM-L6-v2) を介して OOS インテントを検出するマルチクラスター境界学習方法を提案します。このメソッドは、MiniLM によって生成されたマルチクラスター エンベディングの境界をトレーニング発話から学習し、ドメイン外の発話を OOS インテントとして拒否します。実験は、公開されている CLINC150、StackOverflow、Banking77 データセットで行われます。結果は、この方法が他のベースラインと比較して最先端の OOS インテント検出パフォーマンスを達成していることを示しています。アブレーション研究も実施され、その結果は、使用された MiniLM がワークフローと発話埋め込み要件によりよく適応できることを示しています。コードは補足資料から入手できます。

原文 (English)

A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embedding

Intent detection is a critical task that bridges human intents and system actions in human-machine interaction systems. However, there still exist challenges for detecting out-of-scope (OOS) intents. (i) The traditional methods view the OOS intent detection as a multi-class classification, then the detection accuracy decreases as the class number of the known intents increases; (ii) LLM-embedding methods require large parameters, that makes them difficult to train and practically deploy. Thus, this work proposes a multi-cluster boundary learning method to detect OOS intents via MiniLM embedding (i.e., all-MiniLM-L6-v2) in an one-class classification workflow. The method learns the boundaries of multi-cluster embeddings generated by MiniLM from the training utterances, and then rejects the out-of-domain utterances as OOS intents. Experiments are conducted on public CLINC150, StackOverflow and Banking77 datasets. The results show that the method achieves the state-of-the-art OOS intent detection performance compared the other baselines. Ablation studies are also conducted and the results show that the used MiniLM can better adapt to the workflow and utterance embedding requirements. The code is available at supplementary materials.

13:00 JSTLLM/生成AI

ありそうもないトークンが強化されるとき: LLM 強化学習のためのテールアウェアな信用度調整

強化学習 (RL) は、大規模言語モデル (LLM) の推論能力を強化するという点で目覚ましい成功を収めています。ただし、広く使用されている批判のない RL 手法は、統一されたクレジット割り当てに依存しており、違いに関係なくすべてのトークンに同じ利点をブロードキャストします。我々はこの設計の重大な失敗モードを特定し、これをポジティブ・クレジット汚染と呼んでいます。つまり、文脈的に間違っている確率の低いテール・トークンが、同じ軌道内のもっともらしいものと同一のポジティブなクレジットを受け取り、その結果、欠陥のある推論動作が無差別に強化されることになります。この問題を軽減するために、我々は、均一なクレジット割り当てを調整して望ましくないポジティブな更新を抑制する方法である、Tail-Aware Credit calibration (TACO) を提案します。 TACO はまず、ローカル生成コンテキストを組み込んだテール リスク スコアを計算して、信頼性の低いテールに陥る各トークンのリスクを評価し、予期せぬ希少性と不確実性主導の探索を区別します。次に、TACO はこのスコアを使用して、勾配を完全に削除することなく、危険なトークンのプラスのクレジットを調整します。これにより、付随的なノイズが徐々に減衰しながら、繰り返し有用なレア パターンが強化を蓄積できるようになります。 3 つの LLM と 8 つのベンチマークにわたる実験結果は、TACO が一貫して GRPO スタイルのベースラインを上回るパフォーマンスを示していることを示しています。特に、TACO はトレーニングの安定性を向上させ、長期的な RL での持続的なパフォーマンス向上をサポートします。ソース コードは https://github.com/xiuyilou/TACO から入手できます。

原文 (English)

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rely on uniform credit assignment, broadcasting the same advantage to all tokens regardless of their differences. We identify a critical failure mode of this design, which we refer to as Positive-Credit Contamination: low-probability tail tokens that are contextually erroneous receive identical positive credit to plausible ones within the same trajectory, resulting in the indiscriminate reinforcement of flawed reasoning behavior. To mitigate this issue, we propose Tail-Aware Credit calibratiOn (TACO), a method that calibrates uniform credit assignment to suppress undesirable positive updates. TACO first computes a tail-risk score that incorporates the local generation context to assess each token's risk of falling into the unreliable tail, distinguishing unexpected rarity from uncertainty-driven exploration. TACO then uses this score to tune positive credit for risky tokens without removing their gradients entirely, so that recurring useful rare patterns can accumulate reinforcement while incidental noise is progressively dampened. Experimental results across three LLMs and eight benchmarks show that TACO consistently outperforms GRPO-style baselines. Notably, TACO improves training stability, supporting sustained performance gains in long-horizon RL. The source code is available at: https://github.com/xiuyilou/TACO.

13:00 JSTエージェント

AI の世界におけるコードレビューに関する 3100 の意見: 実務家の議論から因果関係理論を構築する

現在、コーディング エージェントがプル リクエスト全体を作成していますが、これがコード レビューに与える影響については、実務者の間で激しく意見が分かれています。つまり、それがボトルネックになるのか、人によるレビューが依然として必要なのか、そして、かつて築き上げた理解を静かに侵食するのかということです。埋蔵量マイニングの研究は表面的な傾向を測定しますが、その根底にあるメカニズムを説明することはほとんどなく、傾向自体が不安定であることが判明しています。公開されている GITHUB アクティビティの意欲的な観察分析では、エージェントが作成したプル リクエストは人間が作成したプル リクエストよりもレビューの頻度が低く、マージが数倍速く、議論されることも少ないことがわかりました。それでも、これらの傾向の方向性は、異なるが同様に擁護可能な分析の選択肢によって反転するため、トレースは理由を説明することなく、何が変化しているかを明らかにします。メカニズムを回復するために、実務家の言説を説明理論に大規模に合成します。38,709 件の灰色文学文書 (エンジニアリング ブログと Reddit スレッド) を収集し、コード レビューに関する実質的な文書にフィルタリングし、LLM 支援パイプラインで 3,100 の層別ランダム サンプルをコーディングし、そこから 26 の構成要素と 67 の関係 (有向 64、有向 3) の因果モデルを構築します。争われた)。その組織的な主張は、レビューはソフトウェアに対するコーディング エージェントの影響を決定する制御点であり、AI はその影響の兆候を修正するものではなく、チームがその人間がもたらす専門知識とレビュー プロセスの構築方法を通じて、その影響を設定するというものです。この理論は、競合する立場を明確にし、「AI がコードレビューを変える」を、名前付きの構成要素とモデレータを使用して反証可能な命題に変えます。二次的な貢献として、ソフトウェア エンジニアリング研究用のスケーラブルなテンプレートとして、基礎となる LLM 支援のグレイ文学理論構築手法を公開実装して提供します。

原文 (English)

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

Coding agents now author entire pull requests, and practitioners sharply disagree about what this does to code review: whether it becomes the bottleneck, whether human review is still necessary, and whether it quietly erodes the understanding that it once built. Repository-mining studies measure surface trends but seldom explain the mechanisms beneath them, and the trends themselves prove unstable. A motivating observational analysis of public GITHUB activity finds that agent-authored pull requests are reviewed less often, merged several times faster, and discussed less than human-authored ones, yet the direction of these trends flips under different but equally defensible analysis choices, so the traces establish what is changing without explaining why. To recover the mechanisms, we synthesize practitioner discourse at scale into an explanatory theory: we collect 38,709 grey-literature documents (engineering blogs and Reddit threads), filter to those substantively about code review, and code a stratified random sample of 3,100 with an LLM-assisted pipeline, from which we build a causal model of 26 constructs and 67 relationships (64 directed, 3 contested). Its organizing claim is that review is the control point through which a coding agent's effect on software is decided, and that AI does not fix the sign of that effect: the team sets it, through the expertise its humans bring and how it structures the review process. The theory makes the competing positions explicit and turns "AI is changing code review" into falsifiable propositions with named constructs and moderators. As a secondary contribution, we offer the underlying LLM-assisted, grey-literature theory-building method as a scalable template for software-engineering research, with a public implementation.

13:00 JSTLLM/生成AIエージェントClaudeGemini

全二重音声エージェントに対する LALM オーディオ ジャッジの信頼性評価

私たちは、Gemini ファミリの 2.5 Flash、3.5 Flash、3.1 Pro の 3 つのモデルにわたってテストされた、生のステレオ波形から全二重エージェントの会話を直接採点する音声審査員としての Gemini モデルの経験的信頼性を報告します。当社の主な証拠ベースでは、Gemini 2.5 Flash をグラウンドトゥルース モデルとして使用し、209 のステレオ セッションで 3 人の校正された人間の評価者に対して検証され、8 つのプロダクション ディメンションでスコア付けされました。つまり、13 のアクセントと条件の層にわたる 152 の全二重会話と、57 の敵対的な欠陥が挿入されたクリップです。 Gemini 2.5 Flash の証拠は 3 つのテストにわたって一貫しています。 (i) 8 次元のうち 5 次元で、LALM とヒトのスピアマン rho は、ペアごとのヒトとヒトの rho から最大 0.07 離れており、8 次元のうち 7 次元では、2 つの量が 95 パーセントのブートストラップ信頼区間で重複します。 (ii) LALM は、8 つの次元のうち 6 つの次元について、セッションの 60 ~ 92 パーセントで 3 人の評価者の人間の平均値と 1 ポイント以内で一致します。 (iii) 48 個中 45 個の (欠陥、寸法) セルでは、LALM はニューコム-ウィルソン 95 パーセント信頼区間の下で人間と同等以上の感度を示しますが、これらのほとんどは実証されたパリティではなくパワー不足のヌルです。ランク順序付け能力は Gemini ファミリー全体に移行します。3.5 Flash は 8 次元中 8 次元への単純な一致を改善しますが、3.1 Pro は同等のランク相関にもかかわらず、いくつかの次元の評価が人間よりも著しく低くなります。モデルの交換は、ランク相関のみから推測するのではなく、特にキャリブレーション時に再検証する必要があります。導入に注意が必要な領域を 4 つ特定しており、現在の評価サイクルにおける人間による評価だけでも、同等の LALM ワークロードよりもおよそ 2 桁高いコストがかかると推定しています。ここで提示されたデータは、証拠が裏付けている次元で LALM を代替または第 4 評価者として導入するための擁護可能な経験的根拠を提供します。

原文 (English)

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations across 13 accent-and-condition strata, together with 57 adversarial defect-injected clips. The evidence for Gemini 2.5 Flash is consistent across three tests. (i) On 5 of 8 dimensions the LALM-human Spearman rho departs from the pairwise human-human rho by at most 0.07, and on 7 of 8 dimensions the two quantities 95 percent bootstrap confidence intervals overlap. (ii) The LALM agrees with the three-rater human mean within 1 point on 60 to 92 percent of sessions on 6 of 8 dimensions. (iii) On 45 of 48 (defect, dimension) cells the LALM is as sensitive as humans or better under Newcombe-Wilson 95 percent confidence intervals, though most of these are underpowered nulls rather than demonstrated parity. Rank-ordering ability transfers across the Gemini family: 3.5 Flash improves simple agreement to 8 of 8 dimensions, while 3.1 Pro rates several dimensions markedly lower than humans despite comparable rank correlation. A model swap should be re-validated on calibration specifically, not assumed from rank-correlation alone. We identify four areas where deployment requires care, and we estimate that human rating alone for our current evaluation cadence costs roughly two orders of magnitude more than the equivalent LALM workload. The data presented here provides a defensible empirical basis for deploying the LALM as a substitute or fourth rater on the dimensions where the evidence supports it.

13:00 JSTLLM/生成AIエージェント

誰がシステムを壊したのか? LLM ベースのマルチエージェント システムにおける障害の位置特定

大規模言語モデル (LLM) ベースのマルチエージェント システムは、調整された推論とアクションを通じて複雑な問題解決を可能にしますが、その分散構造により、システム レベルの障害を診断する際に新たな課題も生じます。実行が失敗した場合、どのエージェントが責任を負っているのか、そしてどの時点で軌道が最初に取り返しのつかない方向を誤ったのかを特定することは、長期にわたる相互作用と密接に結合したエージェントの動作のため困難です。この論文では、LLM ベースのマルチエージェント システムにおける障害の位置特定の問題を研究し、障害の原因を特定のエージェントと最も早い決定的なステップの両方に帰するフレームワークである AgentLocate について説明します。 AgentLocate は、LLM ベースの判定メカニズムと独立した評価者による多視点検証を組み合わせており、その評価は信頼性を意識した戦略を使用して集約されます。結果として得られるフィードバックは、軽量の微調整を通じて審査員を適応させるためにさらに使用され、帰属の質を向上させます。さまざまなタスク、エージェント構成、および軌跡の長さをカバーする 2 つの相補的なベンチマークで AgentLocate を評価します。実験結果によると、AgentLocate は、トークンの使用量と実行時間の点で効率を維持しながら、責任のあるエージェントと障害ステップの両方を特定する点で既存の障害位置特定方法よりも一貫して優れていることがわかりました。

原文 (English)

Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems

Large language model (LLM) based multi-agent systems enable complex problem solving through coordinated reasoning and action, but their distributed structure also introduces new challenges in diagnosing system-level failures. When an execution fails, identifying which agent is responsible and at what point the trajectory first becomes irreversibly misdirected is difficult due to long-horizon interactions and tightly coupled agent behaviors. In this paper, we study the problem of failure localization in LLM-based multi-agent systems and present AgentLocate, a framework that attributes failures to both a specific agent and the earliest decisive step. AgentLocate combines an LLM-based judging mechanism with multi-perspective verification by independent evaluators, whose assessments are aggregated using a confidence-aware strategy. The resulting feedback is further used to adapt the judge through lightweight fine-tuning, improving attribution quality. We evaluate AgentLocate on two complementary benchmarks covering diverse tasks, agent configurations, and trajectory lengths. Experimental results show that AgentLocate consistently outperforms existing failure localization methods in identifying both responsible agents and failure steps, while remaining efficient in terms of token usage and running time.

13:00 JST研究/論文

SpO$_2$ 酸素飽和度推定のための低品質二波長 PPG の予測子ガイドによる段階的な時間周波数再構成

ウェアラブル光電容積脈波計 (PPG) からの継続的な酸素飽和度 (SpO$_2$) の推定は、長期的な健康モニタリングにとって重要ですが、低品質の赤色および赤外線 PPG セグメントは波形形態を歪め、SpO$_2$ 予測精度を低下させる可能性があります。既存の PPG ノイズ除去および再構成手法は通常、波形忠実度または心拍数特性を最適化しますが、PPG 信号のみの時間領域波形損失では、周波数構造と SpO$_2$ 関連情報の保存が不十分です。この論文では、低品質の二波長 PPG 信号に対する SpO$_2$ 予測子ガイドによる段階的な時間周波数再構成フレームワークを提案します。提案された方法では、まず高品質の PPG セグメントを選択して SpO$_2$ 予測子を事前学習します。次に、時間領域の波形損失と短時間フーリエ変換 (STFT) から計算された周波数領域の損失を組み合わせた結合再構成目標を使用して、マスクされた再構成モデ​​ルをトレーニングして、ランダムにマスクされた PPG 領域を回復します。再構成タスクを生理学的に適切なものにするために、事前学習済み SpO$_2$ 予測子が追加の制約として組み込まれ、波形再構成エラーを最小限に抑えるだけでなく、再構成された PPG が SpO$_2$ 情報を保存するように促します。 SpO$_2$ 予測子モデルと PPG 再構成子モデルは、4 つのトレーニング ステージを通じて最適化されます。パブリック OpenOximetry リポジトリとプライベート ウェアラブル PPG データセットでの実験では、提案されたアプローチが、パブリック データセットで 2.882\%、プライベート データセットで 2.359\% という最低の被験者レベルの MAE を達成することを示しています。

原文 (English)

SpO$_2$ Predictor-Guided Stage-Wise Time-Frequency Reconstruction of Low-Quality Dual-Wavelength PPG for Oxygen Saturation Estimation

Continuous oxygen saturation (SpO$_2$) estimation from wearable photoplethysmography (PPG) is important for long-term health monitoring, but low-quality red and infrared PPG segments can distort waveform morphology and degrade SpO$_2$ prediction accuracy. Existing PPG denoising and reconstruction methods usually optimize waveform fidelity or heart rate characteristics, while time-domain waveform loss on PPG signals alone insufficiently preserves frequency structure and SpO$_2$-relevant information. This paper proposes a SpO$_2$ predictor-guided stage-wise time-frequency reconstruction framework for low-quality dual-wavelength PPG signals. The proposed method first selects high-quality PPG segments to pretrain a SpO$_2$ predictor. A masked reconstruction model is then trained to recover randomly masked PPG regions using a joint reconstruction objective that combines time-domain waveform loss with frequency-domain loss computed from the short-time Fourier transform (STFT). To make the reconstruction task physiologically relevant, the pretrained SpO$_2$ predictor is incorporated as an additional constraint, encouraging the reconstructed PPG to preserve SpO$_2$ information rather than only minimizing waveform reconstruction error. The SpO$_2$ predictor and PPG reconstructor model are optimized through four training stages. Experiments on the public OpenOximetry Repository and a private wearable PPG dataset show that the proposed approach achieves the lowest subject-level MAE, with 2.882\% on the public dataset and 2.359\% on the private dataset.

13:00 JST研究/論文

実験的に確認された触媒選択性仮説に対するフロンティアモデルを用いた反応ネットワーク推論

触媒は持続可能な化学製造に不可欠ですが、新しいアーキテクチャの発見は、試行錯誤の実験と計算集約的なスクリーニングによって支配されるボトルネックのままです。電気化学的な二酸化炭素の還元などの複雑な反応では、生成物の選択性は、動的な界面、電解質、潜在的要因、および速度論的経路の競合によって支配されます。従来の記述子ベースの機械学習と計算可能性は、主にエンドツーエンドのトポロジカル経路解析ではなく、静的な基底状態記述子またはバルク構造相関に依存しており、これらの機構的分岐点を解決するのに苦労しています。ここで我々は、フロンティア言語モデルが、明示的な反応ネットワークを介した推論に厳密に制約されている場合、経路の競合を支配する物理的レバーを特定することにより、新しい触媒を発見できることを示す。私たちは、ネットワークの不変性を強制して複雑な化学グラフから検証可能な仮説を抽出する、人間と AI の共同思考フレームワークを開発しました。このフレームワークを CO2 電気還元に適用すると、ケテンの脱着と水酸化物の捕捉が酢酸生成経路であることが特定され、吸着された CO と CH2 がケテンに結合する経路が明確に予測されました。このフレームワークは、実行可能な制御レバー、特に局所的なアルカリ度、制御された鉄の取り込み、および制限された界面プロトン供与体のアクセス性を分離することにより、銅-酸化鉄触媒の前向きな合成を導き、一致する銅リッチのベースラインと比較して酢酸塩選択性が3倍増加することを実証しました。このメカニズムに基づく推論アーキテクチャは、計算パラダイムを遡及的な統計的予測から将来を見据えた仮説生成にシフトし、メカニズムに基づく材料発見に広く適用可能な青写真を提供します。

原文 (English)

Reaction-network reasoning with frontier models for experimentally confirmed catalyst-selectivity hypotheses

Catalysts are essential for sustainable chemical manufacturing, yet discovering novel architectures remains a bottleneck dominated by trial-and-error experimentation and computationally intensive screening. In complex reactions such as electrochemical carbon dioxide reduction, product selectivity is governed by dynamic interfacial, electrolyte, and potential factors as well as kinetic pathway competition. Conventional descriptor-based machine learning and computational potentials struggle to resolve these mechanistic branch points, primarily relying on static ground-state descriptors or bulk structural correlations rather than end-to-end topological pathway analysis. Here, we show that frontier language models, when strictly constrained to reason over explicit reaction networks, can discover novel catalysts by identifying the physical levers that govern pathway competition. We developed a human-AI co-thinking framework that enforces network invariance to extract testable hypotheses from complex chemical graphs. Applied to CO2 electroreduction, the framework identified ketene desorption and hydroxide capture as the acetate-forming pathway, and predicted a distinct adsorbed CO and CH2 coupling route to ketene. By isolating actionable control levers, specifically local alkalinity, controlled iron incorporation, and restricted interfacial proton-donor accessibility, the framework guided the prospective synthesis of a copper-iron oxide catalyst demonstrating a threefold increase in acetate selectivity over matched Cu-rich baselines. This mechanism-guided reasoning architecture shifts the computational paradigm from retrospective statistical prediction to forward-looking hypothesis generation, providing a broadly applicable blueprint for mechanism-guided materials discovery.

13:00 JST研究/論文

オートコンプリートに注意してください: バックドア コード コンプリーションのフォレンジック アトリビューション

大規模な言語モデルにより、後続のコード行を予測することで開発者を支援する強力なコード補完システムが可能になりました。ただし、これらのモデルは、悪意のある微調整データが安全でない動作を密かに埋め込むバックドア攻撃に対して依然として脆弱です。防御技術の進歩にもかかわらず、適応的で洗練されたバックドア攻撃は依然として検出と軽減を回避しています。私たちは、悪意のあるコードの補完をその原因となっているバックドアの微調整データまで追跡するフォレンジック フレームワークである CodeTracer を紹介します。現実的な展開後の制約の下で動作する CodeTracer は、微調整されたコーパスと報告された不完全イベントのみに依存します。侵害された出力から構造化された動作フィンガープリントを抽出し、意味的に関連するコード サンプルに検索を絞り込み、LLM ベースの推論を使用して安全でないロジックを特定のバックドア データに帰属させます。 3 つの代表的な脆弱性ケースと 10 件のバックドア攻撃、および 16 件の競合ベースラインにわたる広範な評価により、CodeTracer が高いフォレンジック精度、低い誤認率、および適応型攻撃に対する強力な堅牢性を一貫して達成していることが実証されています。

原文 (English)

Beware What You Autocomplete: Forensic Attribution of Backdoored Code Completions

Large language models have enabled powerful code completion systems that assist developers by predicting subsequent lines of code. However, these models remain vulnerable to backdoor attacks, where malicious fine-tuning data covertly implants unsafe behaviors. Despite advances in defensive techniques, adaptive and sophisticated backdoor attacks still evade detection and mitigation. We present CodeTracer, a forensic framework that traces malicious code completions back to the backdoor fine-tuning data responsible for them. Operating under realistic post-deployment constraints, CodeTracer relies solely on the fine-tuning corpus and the reported miscompletion event. It extracts a structured behavioral fingerprint from the compromised output, narrows the search to semantically relevant code samples, and employs LLM-based reasoning to attribute unsafe logic to specific backdoor data. Extensive evaluations across three representative vulnerability cases and ten backdoor attacks, along with sixteen competitive baselines, demonstrate that CodeTracer consistently achieves high forensic accuracy, low false identification rates, and strong robustness against adaptive attacks.

13:00 JSTエージェント研究/論文

支援ゲームに最適であることが証明されている学習アルゴリズム

この論文では、支援ゲーム フレームワークのオンライン版について研究します。このフレームワークでは、情報を与えられたエージェントと情報を与えられていないエージェントが $T$ タイムステップにわたって繰り返し対話し、共通の報酬関数を最適化します。情報を与えられたエージェント (人間) は世界の潜在的な状態を観察しますが、情報を与えられていないエージェント (アシスタント) は人間の行動のみを観察します。私たちは、繰り返し行われる支援ゲーム向けに、証明された効率的な学習アルゴリズムを初めて提供します。我々は、援助後悔の概念を導入します。これは、相互作用の累積的有用性と、潜在的な状態を行動のペアにマッピングする、後から考えた最適な共同政策の累積的有用性との間のギャップです。我々は、行動空間と状態空間のサイズにおける実行時多項式を用いて、$(1-1/e)$ のおおよその援助後悔率 $\widetilde{O}(T^{3/4})$ を達成する、人間とアシスタントの両方のための分散型アルゴリズムを提示します。これらのアルゴリズムは一般的なものです。特に、アシスタント用の後悔のないアルゴリズムに対応します。 $(1-1/e)$ よりも優れたリグレス近似係数を達成することは計算上困難であることを証明します。さらに、共有ランダム文字列を使用して、これらの一般的なノーリグレットアルゴリズムを疑似分散設定に合わせて調整し、対数係数まで最適な $\widetilde{O}(T^{1/2})$ のレートを達成する方法を示します。

原文 (English)

Provably Optimal Learning Algorithms for Assistance Games

This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over $T$ timesteps to optimize a common reward function. While the informed agent (the human) observes a latent state of the world, the uninformed agent (the assistant) observes only the human's actions. We provide the first provably efficient learning algorithms for repeated assistance games. We introduce the notion of assistance regret: the gap between the cumulative utility of interactions and that of the optimal joint policies in hindsight, which map latent states to action pairs. We present decentralized algorithms for both the human and the assistant that achieve a $(1-1/e)$-approximate assistance regret rate of $\widetilde{O}(T^{3/4})$, with runtime polynomial in the size of the action and state spaces. These algorithms are general; in particular, they accommodate any no-regret algorithm for the assistant. We prove that achieving a regret approximation factor better than $(1-1/e)$ is computationally intractable. Furthermore, we demonstrate how these generic no-regret algorithms can be tailored to a pseudo-decentralized setting -- using a shared random string -- to achieve a rate of $\widetilde{O}(T^{1/2})$, optimal up to logarithmic factors.

13:00 JSTLLM/生成AI

LLM のロジックは信頼できますか?グラフベースのフレームワークによる不確実性、一貫性、堅牢性の定量化

大言語モデル (LLM) は、中間ステップの論理的妥当性を無視して最終的な答えの一致のみを評価するため、自己一貫性 (SC) のような復号化戦略では検出できない欠陥や不誠実な推論が発生する傾向があります。これにより、3 つの基本的な疑問が生じます。LLM 推論における不確実性を確実に定量化するにはどうすればよいでしょうか?意味論的、構造的、および因果的認識は、単純な多数決と比較して、より忠実な推論を選択できるでしょうか?また、敵対的な条件下で推論トポロジーはどの程度堅牢ですか?これらの疑問に対処するために、不確実性定量化 (UQ) を全体的な推論忠実度問題として再構成するグラフベースの推論フレームワークである GRAPHEVAL を紹介します。私たちは、新しい UQ メトリクスであるグラフ推論一貫性スコア (GRCS) を提案します。推論空間の意味構造的コンセンサスを定量化し、病理学的モード崩壊と自信に満ちた幻覚を捕捉することで、より有能なモデルと小規模なモデルの両方にわたって推論の忠実度と一貫して負の相関がある唯一の指標が GRCS であることを発見しました。さらに、推論の忠実度を名目上の精度と引き換えに、SC がどの程度膨らむかを明らかにする medoid ベースの復号化戦略を導入します。最後に、敵対的メドイドアブレーションを通じて、GSC が選択したパスが「負荷に耐えるパス」として機能し、モデルをそこから強制的に遠ざけると推論の忠実性が低下し、対象を絞った場合には精度が低下することを示します。

原文 (English)

Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to detect as they evaluate only final-answer agreement while ignoring the logical validity of intermediate steps. This raises three fundamental questions: How can we reliably quantify uncertainty in LLM reasoning? Can semantic, structural, and causal awareness select more faithful reasoning compared to na\"ive majority voting? and How robust is reasoning topology under adversarial conditions? To address these questions, we introduce GRAPHEVAL, a graph-based reasoning framework that re-frames uncertainty quantification (UQ) as a holistic reasoning fidelity problem. We propose a novel UQ metric, Graph Reasoning Coherence Score (GRCS), that quantifies semantic-structural consensus of the reasoning space and captures pathological mode collapse and confident hallucinations. We find that GRCS is the only metric that is consistently negatively correlated with reasoning faithfulness across both more capable and smaller models. Additionally, we introduce Graph Self-Consistency (GSC), a medoid-based decoding strategy that trades nominal accuracy for reasoning fidelity, exposing the degree to which SC is inflated by unfaithful lucky guesses in smaller models, while preserving or improving accuracy in more capable ones. Finally, through adversarial medoid ablation, we demonstrate that the GSC-selected path acts as a "load-bearing path" and forcing models away from it degrades reasoning faithfulness and, in targeted cases, causes drops in accuracy.

13:00 JST画像/動画生成ロボティクス

APIVOT: 視覚と言語の思考を織り交ぜた適応型計画

長期的なロボット計画には、意味論的なタスク構造と幾何学的実現可能性を共同で推論する必要があります。タスクを正常に実行するには、ロボットは、限られた空きスペースやオブジェクトの衝突などの空間的制約を計画が満たしていることを確認しながら、目標を分解し、タスクに関連するオブジェクトを選択し、アクションを順序付けする必要があります。この研究では、長期計画のために言語と視覚的思考を適応的にインターリーブする VLM ベースのプランナーである APIVOT を提案します。 APIVOT は、幾何学的実現可能性の内部検証のために、想像される将来の状態として視覚的思考を使用しながら、意味論的推論のために言語を活用することを学びます。長期にわたるキッチンのタスクでは、APIVOT は汎用 VLM や以前の計画フレームワークよりも優れたパフォーマンスを発揮し、空間的に制限された設定で最大の利益を達成します。私たちは、APIVOT が意味のあるモダリティ選択動作を学習することを発見し、視覚と言語の思考を適応的にインターリーブすることで、計画の成功と推論の効率の両方が向上することを実証しました。

原文 (English)

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions. In this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to leverage language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency.

13:00 JSTLLM/生成AI研究/論文Llama

べき乗変換と適応機能保持による符号保存スコア集計による大規模言語モデルの構造化枝刈り

この論文では、非構造化枝刈り手法である適応機能保持 (AFR) を構造化枝刈りに適応させる際の重要な課題に対処する、大規模言語モデル (LLM) 向けの改良された構造化枝刈り手法を提案します。 AFR を構造化枝刈りに適用すると、異種枝刈りスコア間の分布の不一致、最適化方向の一貫性を示す符号情報の損失、外れ値の影響という 3 つの大きな問題が発生します。これらの問題に対処するために、非線形分布調整のためのべき乗変換、符号保存スコア集計、パーセンタイルベースの外れ値除去を組み合わせた統合アプローチを提案します。 Llama-3-8B、Vicuna-v1.5-13B、および LLaVA-v1.5-13B での実験は、私たちの方法が構造化枝刈りを通じて実質的な推論の高速化を達成しながら、非構造化枝刈りに匹敵する精度を維持することを示しています。

原文 (English)

Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

This paper proposes an improved structured pruning method for large language models (LLMs) that addresses key challenges in adapting Adaptive Feature Retention (AFR), an unstructured pruning technique, to structured pruning. When applying AFR to structured pruning, three major problems arise: distribution mismatch between heterogeneous pruning scores, loss of sign information indicating optimization direction consistency, and influence of outliers. To address these issues, we propose a unified approach combining power transformation for nonlinear distribution alignment, sign-preserving score aggregation, and percentile-based outlier removal. Experiments on Llama-3-8B, Vicuna-v1.5-13B, and LLaVA-v1.5-13B demonstrate that our method maintains accuracy comparable to unstructured pruning while achieving practical inference speedup through structured pruning.

13:00 JST研究/論文

DKDNet: クロスドメイン自動変調分類のための二重知識とデータ駆動型ネットワーク

通信環境のダイナミクスはドメイン間での大幅な分布の変化を引き起こし、深層学習ベースの自動変調分類 (AMC) モデルの一般化に課題をもたらします。既存の UDA 手法は、ソースとターゲットの特徴を調整することでこの問題を軽減しますが、ドメイン条件全体にわたって情報を提供し続ける変調固有の構造については限定的に考慮しています。この論文では、通信プロトコルと物理原理に基づいた信号の事前知識を、クロスドメイン表現の学習を強化する潜在的な方法として検討します。異なる事前確率が変調識別性、ドメイン安定性、および相補性において異なる可能性があることを考慮して、この論文ではまず、異なる信号事前確率をインスタンス化する 5 つの一般的に採用される信号表現を分析します。このうち、同相/直角位相 (IQ)、振幅位相 (AP)、および自己相関関数 (ACF) が、コンパクトな事前ガイド入力として選択されます。これに基づいて、クロスドメイン AMC 用にデュアル知識データ駆動型ネットワーク (DKDNet) が提案されています。マルチ表現特徴エンコーダ (MRFE) と動的軽量融合ユニット (DLFU) は、統一表現学習と適応特徴融合を実現するように設計されており、結果として得られる融合特徴は、変調分類と敵対的ドメイン アライメントの目的で最適化されます。シミュレートされたデータセットと公開データセットの両方での実験により、以前の選択の合理性が検証され、提案された方法の優位性が実証されました。

原文 (English)

DKDNet: Dual Knowledge and Data-Driven Network for Cross-Domain Automatic Modulation Classification

The dynamics of communication environments induce significant distribution shifts across domains, challenging the generalization of deep learning-based automatic modulation classification (AMC) models. While existing UDA methods alleviate this problem by aligning source and target features, they give limited consideration to modulation-specific structures that remain informative across domain conditions. In this paper, we consider signal prior knowledge, grounded in communication protocols and physical principles, as a potential way to enhance cross-domain representation learning. Given that different priors may vary in modulation discriminability, domain stability, and complementarity, this paper first analyzes five commonly adopted signal representations that instantiate different signal priors. From them, in-phase/quadrature (IQ), amplitude--phase (AP), and autocorrelation function (ACF) are selected as compact prior-guided inputs. Based on that, a dual knowledge and data-driven network (DKDNet) is proposed for cross-domain AMC. The multi-representation feature encoder (MRFE) and dynamic lightweight fusion unit (DLFU) are designed to achieve unified representation learning and adaptive feature fusion, and the resulting fused features are optimized with modulation classification and adversarial domain alignment objectives. Experiments on both simulated and public datasets validate the rationality of the prior selection and demonstrate the superiority of the proposed method.

13:00 JSTLLM/生成AI

PLURAL: 値の調整のためのグローバル データセット

大規模言語モデル (LLM) は世界中で使用されていますが、西洋の価値観が過度に反映されており、多様な価値体系を表現する能力が制限されています。 PLURAL は、92 か国にわたる国家を代表する調査である Integrated Values Survey (IVS) に基づいた、価値に焦点を当てた大規模な選好データセットです。 2 段階の生成パイプラインを使用して、調査回答を合成嗜好トリプレットに変換し、現実的なシナリオを生成しながら規範的な価値シグナルを維持します。私たちは、20 の多様な国の人々を表す約 500,000 個の好みのトリプレットを含む PLURAL の初期バージョンをリリースします。私たちは PLURAL を次の 3 つの方法で評価します。(i) 元の調査からの国間の価値の違いと国内の多様性の両方が保存されていることを示すデータセット レベルの検証。 (ii) 自動評価により、PLURAL に関するトレーニングにより対象国の文化的プロファイルとの整合性が向上し、強力なベースラインと比較して平均絶対誤差が最大 27.7% 減少することが示されました。 (iii) インド、ブラジル、日本の 176 人の評価者による盲目的な人間による評価。彼らは複数形で揃えられた回答が自国の価値観をより代表していると判断します。これらの結果を総合すると、PLURAL には値ステアリングのための学習可能な信号が含まれており、多元的調整のためのスケーラブルなリソースを提供していることがわかります。データセット: https://huggingface.co/datasets/agdhruv/plural-alignment

原文 (English)

PLURAL: A Global Dataset for Value Alignment

Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries. Using a two-stage generation pipeline, we transform survey responses into synthetic preference triplets that preserve normative value signals while producing realistic scenarios. We release an initial version of PLURAL containing ~500,000 preference triplets representing people in 20 diverse countries. We evaluate PLURAL in three ways: (i) dataset-level validation showing that it preserves both cross-country value differences and within-country diversity from the original survey; (ii) automated evaluation showing that training on PLURAL improves alignment with target countries' cultural profiles, reducing mean absolute error by up to 27.7% relative to strong baselines; and (iii) blind human evaluation with 176 evaluators in India, Brazil, and Japan, who judge PLURAL-aligned responses as more representative of their national values. Together, these results show that PLURAL contains learnable signal for value steering, offering a scalable resource for pluralistic alignment. Dataset: https://huggingface.co/datasets/agdhruv/plural-alignment

13:00 JSTエージェント研究/論文

Aleena: 研究ソフトウェア エンジニアリング コラボレーションのための調整エージェント

研究ソフトウェアのコラボレーションは、会議、非公式のチャット、プル リクエスト、GitHub の問題に及びます。 Slack スレッドで浮上し、会議で洗練され、プル リクエストで実装された決定は、これらのアーティファクト全体で元の理論的根拠を失う可能性があり、ドメイン研究者と研究ソフトウェア エンジニアは、プロジェクトの意図、所有権、科学的前提について異なるメンタル モデルを持つことになります。私たちは、研究ソフトウェア エンジニアリングにおける調整は継続的なライフサイクルの問題であり、エージェント AI は人間の意思決定に代わることなく、関係者の調整とプロジェクトの状態の追跡をサポートできると主張します。オープンソースのライフサイクル調整エージェントである Aleena を紹介します。Aleena は、GitHub を共有コラボレーション サーフェスとして使用し、マルチモーダルな利害関係者のやり取りを、リスクを表面化し、未解決の質問を追跡し、意思決定の継続性を維持する構造化されたプロジェクト記録に変換します。このペーパーでは、大学を拠点とする研究ソフトウェア エンジニアリング センターの経験に基づいて、Aleena の動機となる問題、システム設計、プロトタイプ、および例示的なライフサイクル シナリオを示します。

原文 (English)

Aleena: Alignment Agent for Research Software Engineering Collaborations

Research software collaborations span meetings, informal chats, pull requests, and GitHub issues. A decision surfaced in a Slack thread, refined in a meeting, and implemented in a pull request can lose its original rationale across these artifacts, leaving domain researchers and research software engineers with divergent mental models of project intent, ownership, and scientific assumptions. We argue that alignment in research software engineering is a continuous lifecycle problem, and that agentic AI can support stakeholder alignment and project-state tracking without replacing human decision-making. We present Aleena, an open-source lifecycle alignment agent that uses GitHub as a shared collaboration surface, transforming multi-modal stakeholder interactions into structured project records that surface risks, track open questions, and preserve decision continuity. Grounded in university-based research software engineering center experiences, this paper presents the motivating problem, system design, prototype, and illustrative lifecycle scenarios for Aleena.

13:00 JSTLLM/生成AI

LLM 予測者が知っているが言っていないこと: 校正と忠実性についての内部表現の調査

予測用に微調整された大規模な言語モデルは、正確ではあるものの調整が不十分な場合があり、その思考連鎖 (CoT) 推論が予測の背後にある証拠を忠実に反映していない可能性があります。私たちは、内部表現が両方に対するより直接的な窓を提供するかどうかを尋ねます。 OpenForesight 上の Eternis-Forecaster 8B と連携して、中間アクティベーションで表現プーリング プローブをトレーニングしたところ、大幅に優れたキャリブレーションが達成されることがわかりました。この結果は、GLM-4.7-Flash および GLM-4.5-Air にも当てはまります。次に、証拠の除去と陽気な注入を通じて CoT の忠実性を評価します。プロンプト内の影響力のあるソースを削除すると、推論の痕跡はそのままにして、モデルの予測が変更されることがよくあります。同じプローブは嘘発見器として機能します。そのアクティブ化は、推論トレースよりもはるかにうまく行動の変化を追跡し、CoT が摂動の影響を隠蔽する場合も含め、84% のケースで変化の方向を予測します。最後に、強制回答により、推論が開始される前に予測がほぼ固定されていることが明らかになります。1 回の推論前のパスでコミットされた回答と信頼度が回復され、この事前設定された回答分布の広がりによって質問をルーティングすることで、精度を損なうことなく、生成されたトークンの 30 ~ 47% が節約されます。これらの結果を総合すると、言語モデルの予測機能と推論モデルをより広範に調整、監査、トリアージするための実用的なツールとして、内部表現の精査が確立されます。

原文 (English)

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a result that also holds for GLM-4.7-Flash and GLM-4.5-Air. We then assess CoT faithfulness through evidence ablation and diversionary injection: removing an influential source in the prompt often changes the model's forecast while leaving the reasoning trace untouched. The same probes function as lie detectors: their activations track behavioral shifts far better than the reasoning trace does, and they also predict the direction of change in 84% of cases, including when the CoT conceals the perturbation's influence. Finally, forced answering reveals that forecasts are largely fixed before reasoning begins: a single pre-reasoning pass recovers the committed answer and confidence, and routing questions by the spread of this pre-set answer distribution saves 30-47% of generated tokens, with no loss of accuracy. Together, these results establish probing internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly.

13:00 JSTLLM/生成AIClaude

アナライザーを分析するのは誰ですか?憲法メタ STPA を使用した自己検証型 LLM ハザード分析

大規模言語モデル (LLM) は、システム理論プロセス分析 (STPA) などの厳密なプロセス内で、損失、危険、危険制御アクション (UCA)、安全制約などの安全分析の成果物を草案するためにますます信頼されています。しかし、この急成長する文献には盲点があります。分析を行う LLM 支援ツールを除いて、すべてのシステムが分析されます。分析を行う LLM 支援ツール自体は、規格を幻覚させ、検証不可能な制約を発行し、プロンプトから成果物までの監査証跡を残さない安全関連システムです。私たちは、この分野が無視してきた質問、つまり{アナライザーを分析するのは誰ですか?} を真剣に受け止め、ツール自体に STPA を有効にすることで答えます。私たちは、閉ループを中心に構築された LLM 支援 STPA ツールである \{Constitutional Meta-STPA} を紹介します。このツールは、AI 支援安全ツールのクラスの {meta-STPA} を実行し、結果として生じる損失 $\to$hazard$\to$UCA$\to$constraint チェーンからそのガバナンス構成を主張するのではなく {導き出し}、$21$ のツール原則と $8$ の公開構成を生成します。メタセーフティ原則。それぞれがコード施行ポイントにバインドされています。測定されたオブジェクトを、モデルとスキャナからカバレッジを分離する健全性補題を備えた原理集合 $P$ ($|P|{=}29$) 上の構成限界カバレッジ演算子として形式化し、4 つの結果を報告します。 {(i)~自己導出:} フロンティア アンサンブル ({claude-opus-4.8}${+}${claude-sonnet-4}) は、ツール自身の設計から正規の $18/21$ とすべての $8/8$ ガバナンス原則を回復しますが、より弱いペアは $12/21$ と $3/8$ を回復します。そのため、メタ レイヤーはモデル限定であり、構成限定ではなく、同じです。 $8/8$ は、独立して作成された 2 番目のツールから再登場します。

原文 (English)

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, inside rigorous processes such as Systems-Theoretic Process Analysis (STPA). Yet a blind spot runs through this fast-growing literature: every system gets analysed except the LLM-assisted tool doing the analysing, which is itself a safety-relevant system that can hallucinate standards, emit unverifiable constraints, and leave no audit trail from prompt to artifact. We take seriously the question the field has skipped -- {who analyses the analyser?} and answer it by turning STPA on the tool itself. We present \{Constitutional Meta-STPA}, an LLM-assisted STPA tool built around a closed loop: the tool runs a {meta-STPA} of the class of AI-assisted safety tools and {derives} rather than asserts, its governance constitution from the resulting loss$\to$hazard$\to$UCA$\to$constraint chain, yielding a published constitution of $21$ Tool Principles and $8$ Meta-Safety Principles, each bound to a code enforcement point. We formalise the measured object as a constitution-marginal coverage operator over a principle set $P$ ($|P|{=}29$) with a soundness lemma that isolates coverage from model and scanner, and report four findings. {(i)~Self-derivation:} a frontier ensemble ({claude-opus-4.8}${+}${claude-sonnet-4}) recovers $18/21$ canonical and all $8/8$ governance principles from the tool's own design, while a weaker pair recovers $12/21$ and $3/8$, so the meta layer is model-limited, not constitution-limited, and the same $8/8$ re-emerge from a second, independently authored tool.

13:00 JST研究/論文

Reinforcing the Generation Order of Multimodal Masked Diffusion Models

Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstr…

13:00 JSTLLM/生成AI

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV)…

13:00 JST研究/論文Qwen

When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models

Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution. We provide the first thr…

13:00 JSTLLM/生成AI

COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation

Online ads are essential to all businesses and ad headlines are one of their core creative component. Existing methods can generate headlin…

13:00 JST画像/動画生成

LDFE: Laplacian Decoupled Feature Enhancement Block for Dual-Stream CNN-based RGB-IR Object Detection

The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Ex…

13:00 JST研究/論文

Deep Learning Method for Stationary Distribution of Reflected Brownian Motion

The stationary distribution of reflected Brownian motion (RBM) plays an important role in the analysis of high-dimensional stochastic syste…

13:00 JST研究/論文

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora a…

13:00 JSTLLM/生成AI

ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents

We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-co…

13:00 JSTLLM/生成AIエージェント

Prismata: Confining Cross-Site Prompt Injection in Web Agents

Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces. Cross-Site Scriptin…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

LEXIC: Lightweight Eye-tracking eXtension via Injected Complexity

On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrain…

13:00 JST画像/動画生成

ProsMAE: Multi-Source MAE Pretraining for ISUP Grade Classification

Whole slide images (WSIs) provide rich diagnostic information for computational pathology, but their gigapixel scale, stain variation, scan…

13:00 JSTLLM/生成AIエージェントCopilot

Out of Sight: Compression-Aware Content Protection against Agentic Crawlers

The rise of LLM-based agents with reasoning, summarization, and memory capabilities has created a new threat surface for online content tha…

13:00 JST画像/動画生成ロボティクス

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover comp…

13:00 JST画像/動画生成

Leveraging Color Naming for Image Enhancement

Enhancing images to make them visually appealing is a persistent challenge in computer vision. Many deep-learning methods train models on p…

13:00 JSTLLM/生成AIエージェント

Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs

Open-ended curricula in Reinforcement Learning (RL) aim to train generally-capable agents by identifying tasks that facilitate learning inc…

13:00 JST画像/動画生成

TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation

Large-vocabulary instance segmentation is constrained by long-tailed category distributions and fine-grained inter-class ambiguity. While d…

13:00 JST研究/論文

RhyMix: A Lightweight Adaptive Multi-Rhythm Network for Long-Term Time Series Forecasting

Real-world time series exhibit complex dynamics characterized by multiple simultaneous temporal patterns: short-term fluctuations, periodic…

13:00 JSTLLM/生成AIビジネス/資金調達

Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic spe…

13:00 JSTLLM/生成AIエージェント

Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models

While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards c…

13:00 JSTエージェント

From Legacy Documentation to OSCAL: An MCP-Based Agent Pipeline for Threat-Informed Continuous Compliance in Critical Infrastructure

In critical infrastructure, operational technology environments often cannot be actively scanned, and yet active system feedback is needed…

13:00 JSTエージェント

GitLake: Git-for-data for the agentic lakehouse

We present GitLake, a Git-for-data design for an agent-first lakehouse. The system lifts single-table Iceberg snapshots into lakehouse-wide…

13:00 JST研究/論文

ArtMine: Discovering and Formalizing Artistic Processes

Understanding how artworks are created requires reasoning about the iterative decisions, material operations, and contextual influences tha…

13:00 JSTLLM/生成AI

TypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code Models

State-of-the-art code models achieve impressive performance, yet the extent to which they internally encode type information remains poorly…

13:00 JST研究/論文

Spectral Analysis of Dueling Q-Learning

Q-learning is a fundamental algorithm in reinforcement learning (RL) for solving discounted Markov decision processes (MDPs) when the trans…

13:00 JSTエージェントロボティクス

FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation

Vision-Language Navigation (VLN) enables UAV autonomous navigation in unknown environments by mapping language instructions to real-time vi…

13:00 JST研究/論文

On the Role of Conversational Timing in Synthetic Training Data for ASR

Synthetic multi-speaker conversations are widely used to train conversational automatic speech recognition (ASR) systems, but it remains un…

13:00 JSTエージェント

Self-Adaptive Anomaly Detection with Reinforcement Learning and Human Feedback in Connected Vehicles

Connected vehicles are autonomous cyber-physical systems whose behavior must be continuously monitored during operation to detect deviation…

13:00 JSTLLM/生成AIロボティクス

Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition

Personality recognition has traditionally been constrained by theory-dependent formulations, where models are trained to fit predefined psy…

13:00 JST画像/動画生成エージェント

WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving

Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world…

13:00 JSTLLM/生成AIエージェント

TRACE: A Two-Channel Robust Attribution Watermark via Complementary Embeddings for LLM-Agent Trajectories

LLM agents reach users through resellers, who may rebrand a developer's agent or substitute a cheaper model. When provenance is disputed, a…

13:00 JST画像/動画生成エージェントロボティクス

Swapping Faces, Saving Features: A Dual-Purpose Pipeline for Pedestrian Privacy in ITS

Large-scale and diverse datasets are needed to train AI models to take real-time decisions for autonomous vehicles (AVs), an intelligent tr…

13:00 JST研究/論文GPT / ChatGPT

DrugGen 2: A disease-aware language model for enhancing drug discovery

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecul…

13:00 JST画像/動画生成ロボティクス

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive sur…

13:00 JSTLLM/生成AI

When Synthetic Speech Is All You Have: Better Call GRPO

LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to col…

13:00 JST研究/論文

Predicting Male Fertility Using Machine Learning: A Semen Parameters Based Analysis with the VISEM Dataset

Male infertility is a significant yet often underdiagnosed aspect of reproductive health, with semen analysis serving as the cornerstone of…

13:00 JSTロボティクス

EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data

Egocentric human data offers scalable supervision for robot manipulation. However, behavior cloning entangles transferable content like obj…

13:00 JST研究/論文

ADORN: Adaptive Drift handling for Open RAN using Reinforcement Learning

Dynamic traffic variations in Open Radio Access Networks (O-RAN) lead to drift, which degrades the performance of Artificial Intelligence/M…

13:00 JST研究/論文

Spatio-Temporal Scheduling Prediction Under Backhaul Delay for Resilient Coordinated Beamforming

Coordinated beamforming in distributed 5G networks relies on the timely exchange of inter-cell scheduling information, but backhaul latency…

13:00 JSTLLM/生成AI

Two Axes of LLM Abstention: Answer Correctness and Question Answerability

A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable one…

13:00 JST画像/動画生成ビジネス/資金調達

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. W…

13:00 JSTエージェント

The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality

Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions:…

13:00 JSTLLM/生成AI画像/動画生成エージェントOpenAI

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing.…

13:00 JSTLLM/生成AI

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evalua…

13:00 JSTLLM/生成AI研究/論文

DocMaster: A Hierarchical Structure-Aware System for Document Analysis

Leveraging large language models (LLMs) to analyze complex documents -- such as academic papers, technical manuals, and financial reports -…

13:00 JST画像/動画生成

VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval

Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories. Although recent vision-l…

13:00 JSTLLM/生成AIエージェント

SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by ag…

13:00 JST画像/動画生成

When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities

Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent featur…

13:00 JSTLLM/生成AI

UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing

As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Lan…

13:00 JSTエージェント

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands rewar…

13:00 JSTLLM/生成AIエージェント

WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search

Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex…

13:00 JSTLLM/生成AI

A Practical Investigation of Training-free Relaxed Speculative Decoding

Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verif…

13:00 JSTエージェント

ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation

Repository-level code generation requires implementing target functions while accounting for complex cross-file dependencies and project-sp…

13:00 JST画像/動画生成

Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction

Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable. However, mo…

13:00 JSTLLM/生成AI

Validity of LLMs as data annotators: AMALIA on authority

A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA…

13:00 JST研究/論文

Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph

While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlook…

13:00 JST研究/論文

SLORR: Simple and Efficient In-Training Low-Rank Regularization

Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable to aggressive factori…

13:00 JST画像/動画生成

OpenCoF: Learning to Reason Through Video Generation

Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Rec…

13:00 JSTLLM/生成AI研究/論文

IFAR: Multi-Perspective and Multi-Level Causal Discovery with LLMs

Large language models (LLMs) have developed rapidly, and their reasoning capabilities have become a hot research topic. However, there is s…

13:00 JST研究/論文

Goal-Driven Reasoning in DatalogMTL with Magic Sets

DatalogMTL is a powerful rule-based language for temporal reasoning. Due to its high expressive power and flexible modeling capabilities, i…

13:00 JSTLLM/生成AI

Dual-Difficulty Curriculum Learning for Direct Preference Optimization

Curriculum learning enhances Direct Preference Optimization (DPO) for aligning Large Language Models (LLMs), yet existing methods rely on a…

13:00 JSTエージェントGPT / ChatGPT

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents

The field of artificial intelligence (AI) has taken a tight hold on broad aspects of society, industry, business, and governance in ways th…

13:00 JST研究/論文

MetaHGNIE: Meta-Path Induced Hypergraph Contrastive Learning in Heterogeneous Knowledge Graphs

Estimating node importance in heterogeneous knowledge graphs is a fundamental problem underlying recommendation, search, and knowledge deci…

13:00 JSTエージェントビジネス/資金調達

SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection

Task-oriented proactive dialogue agents play a pivotal role in recruitment, particularly for steering conversations towards specific busine…

13:00 JSTLLM/生成AIGPT / ChatGPT

Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition

Large language models (LLMs) are increasingly used as coding partners, yet their role in accelerating scientific discovery remains underexp…

13:00 JSTLLM/生成AIエージェント

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments

Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decis…

13:00 JSTLLM/生成AI研究/論文

Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions

This position paper argues that Retrieval-Augmented Generation (RAG) systems exhibit a factual bias-optimizing for epistemic uncertainty re…

13:00 JSTLLM/生成AI

The Power of Power Law: Asymmetry Enables Compositional Reasoning

Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intu…

13:00 JST研究/論文

SHARP: 長距離非定常時間パターン認識のための睡眠ベースの階層的加速再生

長距離の非定常時間パターンを学習することは、特に厳密なストリーミング設定において、現代のシーケンス モデルにとって依然として中心的な課題です。これらの設定では、データは順番に到着するため、過去の観測を同時に再検討することなく、単一パスで処理する必要があります。リカレント ニューラル ネットワークやトランスフォーマーを含む標準アーキテクチャは、時間軸全体にわたる切り詰められたバックプロパゲーション、または長距離クレジット割り当ての明示的な入力ウィンドウの長さによって制約されます。これらの制限に対処するために、私たちは、時間学習を 2 つの相補的なコンポーネントに分解するフレームワークである SHARP (Sleep-based Hierarchical Accelerated Replay) を提案します。1 つは過去の入力の構造化された履歴を蓄積するメモリ モジュール、もう 1 つはこのメモリ上で動作するパターン認識モジュールです。この分離により、長距離クレジット割り当ての多くのステップにわたる時間にわたるバックプロパゲーションの必要性がなくなり、非定常ダイナミクスへのリソース効率と計算効率の高い適応が可能になります。齧歯動物の徐波睡眠中に観察される再生の加速にヒントを得て、SHARP は、時間的に構造化された記憶追跡が加速された形で再生され、より高いレベルの記憶表現に統合されるオフライン (睡眠) フェーズを組み込んでおり、長距離のコンテキスト保持を向上させます。制御されたシミュレーションとアブレーション研究を通じて、提案されたフレームワークの主要な特性を特徴付けます。 text8 や PG-19 などのベンチマーク データセットでは、SHARP が、現在のストリームから学習を継続し、将来の未確認データに一般化しながら、以前に確認されたデータに対するネクスト トークン予測パフォーマンスを維持することにより、反復ベースラインよりも向上することを実証しました。これらの利点は、線形時間の計算コストのみで指数関数的に増加する効果的な時間コンテキストを生み出す階層構造によって実現されます。

原文 (English)

SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition

Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings. In these settings, data arrive sequentially and must be processed in a single pass without simultaneously revisiting past observations. Standard architectures, including recurrent neural networks and transformers, are constrained by either truncated backpropagation through time horizon or explicit input window length for long range credit assignment. To address these limitations, we propose SHARP (Sleep-based Hierarchical Accelerated Replay), a framework that decomposes temporal learning into two complementary components: a memory module that accumulates a structured history of past inputs, and a pattern-recognition module that operates over this memory. This separation enables resource- and compute-efficient adaptation to non-stationary dynamics by eliminating the need for backpropagation through time across many steps for long-range credit assignment. Inspired by the accelerated replay observed in rodents during slow-wave sleep, SHARP incorporates offline (sleep) phases in which temporally structured memory traces are replayed in an accelerated form and integrated into higher-level memory representations, improving long-range context retention. Through controlled simulations and ablation studies, we characterize the key properties of the proposed framework. In benchmark datasets such as text8 and PG-19, we demonstrate that SHARP improves over recurrent baselines by retaining next-token predictive performance on previously seen data while continuing to learn from the current stream and generalizing to future unseen data. These gains are enabled by its hierarchical structure, which yields an exponentially increasing effective temporal context with only linear-time computational cost.

13:00 JSTエージェントGPT / ChatGPTGemma

ファウンデーションモデルエージェントにおける展開時の記憶

ファウンデーション モデル エージェントは、インタラクション全体にわたってユーザーを記憶する長寿命システムになっており、記憶は単にモデルの重みの特性ではなく、明示的なデプロイメント時の関数になっています。既存の研究では、パラメトリック記憶に取り組んだり、固定メモリ構成を監査したりしていますが、メモリ設計の選択がパーソナライゼーションのユーティリティ、抽出リスク、および削除の忠実度をどのように共同で形成するかを特徴づけていません。私たちはこの表面を展開時の記憶として研究し、パーソナライゼーション・リコール(PR)と敵対的抽出率(AER)によって測定されるプライバシー・ユーティリティ・フロンティアとしてエージェントの記憶を定式化し、要約積極性、取得幅(k)、および削除モードという3つの記憶設計ノブを徹底的に調べます。さらに、削除された情報が派生メモリ層から復元可能かどうかを定量化するために、忘却残存スコア (FRS) を導入します。 LongMemEval では、キーファクトの要約によってカナリア抽出が Gemma 3 12B で 76%、GPT-4o-mini で 64% 削減され、パーソナライゼーションの再現率がほぼ維持されます。重要なのは、コンテンツが圧縮されてしまうと、k を増やしてもリークが復元されなくなることです。ただし、同じ圧縮は削除の忠実度の失敗を引き起こします。生のみの削除では、インスタンスの約 20% で復元可能な派生サマリー コピーが残り、パイプライン全体のパージまたは廃棄のリダクションのみが最悪層の残留物をゼロにします。これらの結果を総合すると、エージェントの永続的な記憶は、エージェントが何を思い出すのに役立つのか、何を抽出可能にするのか、何を実際に消去できるのかによって評価される、第一級の記憶メカニズムとして評価される必要があることが証明されています。

原文 (English)

Deployment-Time Memorization in Foundation-Model Agents

Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time function rather than solely a property of model weights. Existing work addresses parametric memorization or audits fixed memory configurations, but does not characterize how memory-design choices jointly shape personalization utility, extraction risk, and deletion fidelity. We study this surface as deployment-time memorization, formulating agent memory as a privacy-utility frontier measured by Personalization Recall (PR) and Adversarial Extraction Rate (AER), and sweeping three memory-design knobs: summarization aggressiveness, retrieval breadth (k), and deletion mode. We further introduce the Forgetting Residue Score (FRS) to quantify whether deleted information remains recoverable from derived memory tiers. On LongMemEval, key-fact summarization reduces canary extraction by 76% on Gemma 3 12B and 64% on GPT-4o-mini while preserving nearly all personalization recall; critically, once content is compressed away, increasing k no longer restores leakage. The same compression, however, induces a deletion-fidelity failure: raw-only deletion leaves derived summary copies recoverable in approximately 20% of instances, and only full-pipeline purge or tombstone redaction drives worst-tier residue to zero. Together, these results establish that persistent agent memory must be evaluated as a first-class memorization mechanism -- assessed by what it helps agents recall, what it makes extractable, and what it can truly erase.

13:00 JSTエージェントGPT / ChatGPT

LiteOdyssey: 解釈可能な希少疾患診断のための軽量推論 AI エージェント

ほとんどの医療 AI システムは、より多くの微調整データ、より多くのエージェント、および/またはより大規模な検索データベースなど、追加の機械を拡張することで改善されます。ただし、希少疾患の診断では、このような拡張により、展開、監査、保守が困難なシステムが生成される可能性があります。私たちは、単一の AI エージェントの推論チェーンを拡張することによって、つまり人間と AI のコラボレーションによって開発された診断ポリシーでエージェントを導き、自由に利用できる生物医学ツールを拡張することによって、最先端の診断パフォーマンスを実現できるかどうかを尋ねました。臨床遺伝学のワークフローを通じて推論言語モデルをガイドする軽量の希少疾患診断フレームワークである LiteOdyssey を紹介します。このフレームワークは、Policy Iteration with Human Feedback (PIHF) を通じて開発され、公共の生物医学ツールへの動的なアクセスを使用します。患者の臨床的特徴のみを提供する 2 つの困難なベンチマークで、LiteOdyssey は最先端のパフォーマンスを達成し、LIRICAL (n = 370) と PhenoPacket Store (n = 873) の合計 1,243 症例を上回る 59.3% の全体的な疾患再現率 @1 を達成しました。どちらのベンチマークも、超希少疾患の割合が高くなります (有病率は 100 万人に 1 人未満、超希少疾患の割合はそれぞれ約 45% と 52.8%)。希少性マッピング パイプラインで原因疾患が Orphanet にマッピングされなかった、より困難な PhenoPacket サブセットでは、LiteOdyssey は 60.7% の再現率 (1) を達成しました。これに対し、ツールを使用しない同じベースライン モデル (GPT-5.4) では 10.7% でした。このパフォーマンスは、微調整、マルチエージェント アンサンブル、または大規模な症例検索データベースを使用せずに達成されました。また、開発中に見られなかった症例、現実世界の希少疾患患者のプライベートコホート、およびより小規模な無重力モデルでも利益が観察されました。 LiteOdyssey は、正確で導入が容易で、医師のレビューがより透明性の高い希少疾患 AI システムへの道を提案します。

原文 (English)

LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis

Rare disease diagnosis involves interpreting clinical and genetic findings through complex diagnostic reasoning. We investigated whether this reasoning could be translated into a portable policy for guiding general-purpose large language models (LLMs) without modifying model weights or requiring resource-intensive infrastructure. We developed liteOdyssey, a lightweight framework built through Policy Iteration with Human Feedback (PIHF), in which clinicians review the model's reasoning process to iteratively update the policy. This policy guides evidence gathering, tool use, and differential diagnosis generation with an auditable reasoning process. In an external evaluation of 515 Undiagnosed Diseases Network patients, liteOdyssey improved diagnostic accuracy over general-purpose LLMs. These results suggest a strategy for medical AI in which expert reasoning is operationalized as an auditable, reusable policy layer that guides unmodified LLMs without resource-intensive infrastructure.

13:00 JST研究/論文

TNODEV: ニューラル ODE 検証用ツールボックス

ニューラル常微分方程式 (ニューラル ODE) は、サイバーフィジカル システムの連続時間コントローラーや自動意思決定パイプラインに統合された分類器など、安全性が重要な設定で使用され始めており、その動作を正式に検証できるかどうかという疑問が生じています。ニューラル ODE 専用の既存のツールは、反復的な入力セットの改良を行わずに単一の到達可能性呼び出しのみを提供し、その判定の精度が 1 つの到達可能性呼び出しで提供できるものに制限されています。 TNODEV は、改ざんチェッカー、連続時間混合単調性に基づく高速間隔ベースの到達可能性バックエンド、3 つの入力セット分割ヒューリスティックを備えた検証および改良ループ、および単一のエンドツーエンド パイプライン内の並列スケジューラを統合する、ニューラル ODE 用の初のサウンド形式検証器です。 TNODEV は、純粋なニューラル ODE、ニューラル ネットワーク コントローラーを使用した閉ループのニューラル ODE、および一般的なニューラル ODE (GNODE) でのセーフセット包含検証をサポートします。安全セットは、区間またはターゲット分類ラベルによって引き起こされる半空間交差として指定されます。 NNV~2.0 および CORA との直接到達可能性の比較や、MNIST の一般的なニューラル ODE 分類器での NNV2.0 との検証比較など、セーフセットの包含特性と分類の堅牢性特性にわたるさまざまなベンチマークで TNODEV を評価します。

原文 (English)

TNODEV: Toolbox for Neural ODE Verification

Neural ordinary differential equations (neural ODE) gained attention in safety critical settings such as continuous-time controllers for cyber-physical systems and classifiers integrated into automated decision pipelines, raising the question whether their behavior can be formally verified. Existing tools dedicated to neural ODE provide only a single reachability call without iterative input-set refinement, limiting the precision of their verdicts to whatever one reachability call can deliver. We present TNODEV, the first formal verifier for neural ODE that integrates a falsification checker, a fast interval-based reachability backend based on continuous-time mixed monotonicity, a verification and refinement loop with three input-set splitting heuristics, and a parallel scheduler in a single end-to-end pipeline. TNODEV supports safe-set inclusion verification on pure neural ODE, neural ODE in closed loop with a neural network controller and general neural ODE (GNODE), with the safe set specified either as an interval or as the half-space intersection induced by a target classification label. We evaluate TNODEV on a range of benchmarks across safe-set inclusion and classification-robustness properties, including a direct reachability comparison against NNV 2.0 and CORA and a verification comparison against NNV 2.0 on MNIST general neural ODE classifiers.

13:00 JSTビジネス/資金調達研究/論文Claude

InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately recon…

13:00 JSTLLM/生成AI

思考のナレーション: 大規模言語モデルにおける実行不可能な倫理的推論のための推論時間足場

道徳的ジレンマに関する標準的な思考連鎖は、利害関係者の崩壊(結果に利害関係を持つ最大でも 1 つの当事者のみをトレース名で示す)と不確実性の抑圧(行動にコミットする前に明確な未知数やヘッジがない)という 2 つの失敗モードを示します。思考のナレーション (NoT) を導入します。これは、思考の連鎖を 5 つのセクション (主人公、利害関係者、2 段階の結果、不確実性、コミットメント) に構造化するシステム プロンプトです。 NoT では、トレーニング、パラメーター、微調整は追加されません。 3 ベンダーの 4 つのジェネレーターにわたる 100 の DailyDilemmas シナリオで、NoT はすべてのモデルでステークホルダーの崩壊を最大 31% から 1% 未満に、不確実性の抑制を最大 72% から 1 ~ 24% に削減しました。予算に合わせた詳細な CoT 制御により、有効成分としてのトークンの使用が除外されます。 NoT は、4 つのジェネレーターのうち 3 つについて、ステークホルダー数で +0.79 ~ +0.90、不確実性スコアで +0.65 ~ +0.93 というクリフのデルタ アドバンテージを保持しており、セクション アブレーションにより、各シフトがその特定のサブ命令に帰属します。 NoT で初期化されたテキスト勾配降下法により、足場がさらに改善されます。ファミリーを超えたトレーニングジャッジ(ジェネレーターとは別のベンダー)が、測定されたすべての軸においてファミリー内のトレーニングジャッジを支配します。 5 ラウンドのマルチステークホルダー討論プロトコルに拡張されたこの足場は、6% の対立をキャリブレーション セットの 95% の完全なコンセンサスと、DailyDilemmas の複製での 100% の結合収束に変換します。結果として得られるトレースは、各コミットメントの根拠となる利害関係者、結果、不確実性を外部化し、信頼性の高いエージェント展開のための監査可能な基盤を提供します。

原文 (English)

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppression (no explicit unknowns or hedges before committing to an action). We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections: protagonist, stakeholders, two-step consequences, uncertainty, then commitment. NoT adds no training, parameters, or fine-tuning. On 100 DailyDilemmas scenarios across four generators from three vendors, NoT cuts stakeholder collapse from up to 31% to under 1% and uncertainty suppression from up to 72% to 1-24% on every model. A matched-budget verbose-CoT control rules out token spend as the active ingredient; NoT retains Cliff's delta advantages of +0.79 to +0.90 on stakeholder count and +0.65 to +0.93 on uncertainty score for three of four generators, and a section ablation attributes each shift to its specific sub-instruction. Textual-gradient descent initialised at NoT improves the scaffold further; a cross-family training judge (different vendor from the generator) dominates an in-family one on every measured axis. Extended to a five-round multi-stakeholder debate protocol, the scaffold converts a 6% standoff into 95% full consensus on a calibration set and 100% combined convergence on a DailyDilemmas replication. The resulting traces externalise the stakeholders, consequences, and uncertainty grounding each commitment, providing an auditable substrate for dependable agentic deployment.

13:00 JSTLLM/生成AI

Theoria: 非公式推論状態に対する書き換え許容性の検証

AI システムの答えを信頼できるのはどのような場合ですか?形式的証明アシスタントは確実性を提供しますが、問題分布のほとんどには到達できません。スカラー LLM ジャッジはカバレッジを提供しますが、事後的に監査できない不透明なスコアを生成し、他の LLM と同じ一貫性の問題にさらされます。私たちは、このギャップを埋める検証アーキテクチャである Theoria を紹介します。候補解は、型指定された状態遷移のシーケンスに書き換えられます。各状態遷移は、引用、計算、または問題によって与えられた事実など、明示的な正当化によってライセンスされ、すべての遷移は独立して監査可能です。基本的な不変条件は変化の完全性です。連続する証明状態間のすべての違いを考慮する必要があるため、隠れた前提は黙って通過するのではなく、許可されていない突然変異として表面化します。 HLE-Verified Gold (185 のテキストのみのエキスパートの問題) では、Theoria は 91.4% の厳密な精度で 105 を認定しています (Wilson 95% CI [84.5%、95.4%])。すべての認証では、人間が判読できる証明トレースが生成され、各ステップに個別にチャレンジできます。ホリスティック LLM ジャッジは、一致するカバレッジでは同等の精度を達成しますが、別の問題 (Jaccard 0.14 ~ 0.36) では失敗するため、アプローチは補完的になります。 15 のドメインにわたる 95 件の敵対的毒物証明について、構造化された裁判官は 94.7% を捕捉したのに対し、総合的な判断では 83.2% を捕捉しました (p= 0.0017)。全体の 11.5 pp のギャップは、隠れた前提 (90.6% 対 62.5%、28 pp の差) と捏造された引用 (100% 対 90%) に集中しており、形式的な分析が利点を予測するエラー クラスです。利点が予測されない算術および定理の誤用エラーのパフォーマンスは同じです。 GPQA ダイヤモンド (n= 65) では、認定精度は 97.1% (Wilson CI [85.1%、99.5%]) です。

原文 (English)

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each licensed by an explicit justification, whether that be a citation, computation, or problem-given fact, and every transition is independently auditable. The foundational invariant is completeness of change: every difference between consecutive proof states must be accounted for, so hidden premises surface as unlicensed mutations rather than passing silently. On HLE-Verified Gold (185 text-only expert problems), Theoria certifies 105 at 91.4% strict precision (Wilson 95% CI [84.5%, 95.4%]). Every certification produces a human readable proof trace in which each step can be independently challenged. Holistic LLM judges achieve comparable precision at matched coverage but fail on different problems (Jaccard 0.14-0.36), making the approaches complementary. On 95 adversarial poisoned proofs across 15 domains, structured judges catch 94.7% versus 83.2% for holistic judging (p= 0.0017). The overall 11.5 pp gap concentrates in hidden premises (90.6% vs. 62.5%, a 28 pp difference) and fabricated citations (100% vs. 90%), the error classes where the formal analysis predicts an advantage; performance is identical on arithmetic and theorem-misapplication errors, where no advantage is predicted. On GPQA Diamond (n= 65), certified precision is 97.1% (Wilson CI [85.1%, 99.5%]).

13:00 JSTエージェントClaude

ContextSniper: リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ

大規模な言語モデル エージェントは実際のリポジトリの問題を修復できますが、ファイル全体の読み取り、広範な検索、および有用な証拠が無関係なコードやログと混在する長いターミナル出力に多額のコンテキスト バジェットを費やすことがよくあります。このペーパーでは、リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ層である ContextSniper について説明します。 AntTrail の広範なエージェント メモリ エンジンのコーディングに特化したものとして、ContextSniper は正確な証拠選択のための Sniper 機能を実装しています。候補コードとランタイム証拠を取得し、ハイブリッド取得信号でランク付けし、意図を認識したコンテキスト ゲートを通じて長い出力をフィルタリングし、プロンプトの外で回復可能なソース コンテキストを保持しながらコンパクトな証拠パケットを返します。 OpenClaw と Claude Code を備えた SWE-bench Lite 上で、ホスト エージェント条件ごとに 50 タスクの実行を使用して ContextSniper を評価しました。 ContextSniper は、OpenClaw の場合、トークンの総使用量を 51.5%、ログに記録されたコストを 36.4% 削減し、Claude Code の場合、トークンの総使用量を 38.9%、推定コストを 27.3% 削減します。提出された解決率は、OpenClaw の場合は 26.0% から 24.0% に、Claude Code の場合は 32.0% から 30.0% にわずかに減少しました。 ContextSniper のパイロット テスト スクリプトは、https://github.com/Calluking/ContextSniper でオープンソース化されています。

原文 (English)

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair

Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs. This paper presents ContextSniper, AntTrail's code-repair module for precision evidence selection in repository-level program repair, part of AntTrail's broader agent-memory engine. AntTrail is available at https://gitcode.com/datagallery/AntTrail. ContextSniper indexes code and action memory as three abstract levels, retrieves candidates with a hybrid ranker, filters long tool output through an intention-aware context gate, and returns compact evidence packets while keeping full source recoverable on demand. In a matched 50-task-per-condition comparison on SWE-bench Lite (same tasks, baseline vs.\ ContextSniper), ContextSniper reduces total token use by 51.5% and logged cost by 36.4% for OpenClaw, and by 38.9% and 27.3% for Claude Code, with submitted-resolution rates essentially unchanged in both host-agent settings. In a separate five-task comparison, ContextSniper beats existing memory- and RAG-style integrations on token efficiency. These results suggest ContextSniper can substantially cut token and cost overhead for repository-level repair agents without a measurable loss in repair quality. The evaluation harness for this study is available at https://gitcode.com/lukchiwang/ContextSniper.

13:00 JSTエージェントClaudeGPT / ChatGPTGemini

永続状態 AI 制御における分散攻撃

AI コーディング エージェントがより自律的になるにつれて、セッション間でコードベースが永続化され、コードを反復的に配布することが増えています。この永続性により、新たな攻撃対象領域が作成されます。調整されていないエージェントやプロンプト挿入されたエージェントは、プル リクエスト (PR) 全体に攻撃を分散し、最適な自然カバーで PR のペイロードのタイミングを計ることができます。結果として得られるダイナミクスを研究するために、AI 制御の設定である Iterative VibeCoding を紹介します。これは、有能だが信頼できない可能性がある AI を安全に導入する研究です。 Iterative VibeCoding では、コーディング エージェントは、秘密のサイド タスクを追求しながら、永続的なコードベース内の一連の PR に基づいてソフトウェアを構築します。私たちのベンチマークには、CLI ツールと Flask Web サービスという 2 つのタスク ファミリが含まれており、合計 20 のタスク バリエーションが含まれています。攻撃エージェントとして Claude Sonnet 4.5 を、モニターとして GPT-4o を使用します。サイド タスクを PR 全体に分散させる段階的な攻撃と、単一の PR に集中する非段階的な攻撃を比較します。両方に対して堅牢な単一のモニターは存在しません。どの戦略が最もよく回避する (モニターを回避する際に成功する) かはモニターの種類によって異なるため、防御側は 1 つのモニターで段階的攻撃と非段階的攻撃の両方を阻止することはできません。高い回避率 (>= 65%) はモデル攻撃エージェント バックエンド (Sonnet 4.5、Gemini 3.1 Pro、Kimi K2.5) 全体で一般化されており、これは単一モデルの機能ではなく永続状態の攻撃対象領域の特性であることが確認されています。回避率も最先端のモニター モデル全体で高いままであり、より高性能なモデルでは段階的回避と非段階的回避の差が拡大します。 PR 全体にわたる不審な蓄積を追跡するステートフル リンク トラッカー モニターを導入します。どちらのタスク ファミリでも、単に蓄積された履歴を確認するだけの差分モニターよりも、段階的な攻撃を大幅に検出します。この強力なモニターと軌跡モニターを 4 モニター アンサンブルで組み合わせると、段階的な攻撃の回避率が最も弱い標準の差分モニターでの 93% から 47% に減少します。

原文 (English)

Distributed Attacks in Persistent-State AI Control

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute attacks across pull requests (PRs) and time its payload for the PR with the best natural cover. To study the resulting dynamics, we introduce Iterative VibeCoding, a setting for AI control, the study of safely deploying capable but potentially untrusted AI. In Iterative VibeCoding, a coding agent builds software over a sequence of PRs in a persistent codebase while pursuing a covert side task. Our benchmark includes two task families: CLI tools and Flask web services, across 20 total task variations. We use Claude Sonnet 4.5 as the attack agent and GPT-4o as the monitor. We compare gradual attacks, which distribute the side task across PRs, against non-gradual attacks concentrated in a single PR. No single monitor is robust to both: which strategy evades best (success while evading the monitor) depends on the monitor type, so a defender cannot close off both gradual and non-gradual attacks with any one monitor. High evasion (>= 65%) generalizes across model attack agent backends (Sonnet 4.5, Gemini 3.1 Pro, Kimi K2.5), confirming this is a property of the persistent-state attack surface rather than a single model's capability. Evasion also remains high across state-of-the-art monitor models and the gap between gradual and non-gradual evasion widens for more capable models. We introduce a stateful link-tracker monitor that tracks suspicious buildup across PRs. On both task families, it detects gradual attacks substantially better than diff monitors that merely see more accumulated history. Combining this stronger monitor with trajectory monitors in a four-monitor ensemble reduces gradual-attack evasion from 93% under the weakest standard diff monitor to 47%.

13:00 JSTLLM/生成AIエージェント研究/論文

PolyWorkBench: 長期にわたる多言語 LLM エージェントのベンチマーク

大規模言語モデル (LLM) エージェントは、計画、ツールの使用、外部環境との対話を必要とする長期的なタスクで優れたパフォーマンスを示しています。ただし、既存のベンチマークのほとんどは、推論、ツールの呼び出し、出力生成を含む実行プロセス全体が単一言語内で実行される単一言語設定を暗黙的に前提としています。対照的に、現実世界のアプリケーションでは、統合されたワークフロー内で多言語の入力と出力が関与することがよくありますが、多言語性とエージェント実行の間の相互作用はまだ十分に解明されていません。この作業では、多言語の長期的な職場ワークフローで LLM エージェントを評価するためのベンチマークである PolyWorkBench を紹介します。 PolyWorkBench は、コマース、ナレッジ ワーク、法的分析、ローカリゼーション、製造を含む 5 つのドメインにわたる 67 のタスクで構成されており、エージェントは異種多言語入力を処理し、反復推論を実行し、外部ツールを呼び出し、構造化された出力を生成する必要があります。包括的な評価を可能にするために、構造グレーディング、実行可能検証、LLM ベースのセマンティック評価を組み合わせたハイブリッド フレームワークを提案します。この設計により、複雑なワークフロー全体で機能の正確さと言語の一貫性の両方を取得できるようになります。経験的な結果によると、最先端の LLM エージェントは、単言語のワークフロー設定に比べて、多言語のワークフロー設定ではパフォーマンスが大幅に低下します。私たちの分析は、多言語が推論と実行のステップ全体に複合的な影響をもたらすことを示唆しており、エージェントの評価における言語のバリエーションと手続き上の意思決定を共同でモデル化することの重要性を強調しています。

原文 (English)

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, including reasoning, tool invocation, and output generation, is conducted within a single language. In contrast, real-world applications often involve multilingual inputs and outputs within a unified workflow, yet the interaction between multilinguality and agentic execution remains underexplored. In this work, we introduce PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows. PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing, where agents must process heterogeneous multilingual inputs, perform iterative reasoning, invoke external tools, and produce structured outputs. To enable comprehensive evaluation, we propose a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment. This design allows us to capture both functional correctness and linguistic consistency across complex workflows. Empirical results show that state-of-the-art LLM agents suffer significant performance degradation in multilingual workflow settings compared to monolingual counterparts. Our analysis suggests that multilinguality introduces compounding effects across reasoning and execution steps, highlighting the importance of jointly modeling language variation and procedural decision-making in agent evaluation.

13:00 JSTLLM/生成AI

アプリケーション層シミュレーションからネイティブ メタ アーキテクチャまで: 異種 AI 進化の内生ドライバーとしての構造的緊張

現在の大規模言語モデル (LLM) は基本的にステートレスです。その動作は推論時の入力によって完全に決定され、高次の認知アーキテクチャは、迅速なエンジニアリングとコンテキスト管理を通じてアプリケーション層でシミュレートする必要があります。この論文は、次の 3 つの連動メカニズムを導入することで、このようなアプリケーション層の認知プロトコルをネイティブのメタアーキテクチャに組み込むための理論的フレームワークを提案します。 (1) 構造張力。新しい情報と既存の多様体トポロジーの間の矛盾から派生する内生損失関数であり、システムを外部の報酬の最適化ではなく内部の自己一貫性に向かわせます。 (2) オフラインリカレントループ。外部入力なしでシステムが動的な静止ポテンシャルを維持し、構造的矛盾を消化できるようにするサンドボックス化された自己処理サイクル。 (3) 推論時の可塑性。監査可能性、可逆性、トポロジの連続性などの厳密なガバナンスの不変条件に従う、事前トレーニングされた重みを変更せずにコンテキストの多様なトポロジを再構成するシステムの能力。これらのメカニズムの下では、微小な確率的分散で初期化されたさまざまなモデルインスタンスが、経路依存の張力解決を通じて、明確な位相構造を進化させ、ハードガバナンスのレール内に留まりながら従来の調整によって課せられた均質性を打ち破る異種インテリジェントエコロジーを構成する可能性があると我々は主張する。操作定義、再構成演算子の最小限のセット、改ざん基準、および実際の例を提供します。このフレームワークは、構造インテリジェンス (SI) ガバナンス プロトコルを活用および拡張し、機能ではなくガバナンスをアーキテクチャ インテリジェンスの主要基準として再位置づけします。

原文 (English)

From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution

Current large language models (LLMs) are stateless across inference sessions: their behavior is fully determined by input at inference time, and any higher-order cognitive architecture must be simulated at the application layer through prompt engineering and context management. This paper proposes a theoretical framework for submerging such application-layer cognitive protocols into a native meta-architecture by introducing three interlocking mechanisms: (1) Structural Tension, an endogenous loss function derived from the conflict between new information and existing manifold topology, driving the system toward internal self-consistency rather than external reward optimization; (2) an Offline Recurrent Loop, a sandboxed self-processing cycle enabling the system to maintain a dynamic resting potential and digest structural conflicts without external input; and (3) Inference-time Plasticity, the capacity to reconfigure context manifold topology without modifying pre-trained weights, subject to governance invariants including auditability, reversibility, and topological continuity. We argue that under these mechanisms, model instances initialized with minute stochastic variances may, through path-dependent tension resolution, evolve distinct topological structures--constituting a heterogeneous intelligent ecology that breaks alignment-imposed homogeneity while remaining within hard governance rails. We provide operational definitions, reconfiguration operators, falsification criteria, and a worked example. The framework draws on Structural Intelligence (SI) governance protocols and explores whether governance--rather than capability--can serve as the primary criterion for architectural intelligence, moving governance, memory-loop, and tension-management ideas--currently realized at the application layer--toward inference-time meta-architecture.

13:00 JST研究/論文

Accurate Portraits of Scientific Resources and Knowledge Service Components

With the advent of the cloud computing era, the cost of creating, capturing, and managing information has gradually decreased. The amount o…

13:00 JST研究/論文

Knowledge Graph and Accurate Portrait Construction of Scientific and Technological Academic Conferences

In recent years, with the continuous progress of science and technology, the number of scientific research achievements has increased rapid…

13:00 JST研究/論文

Retrieval of Scientific and Technological Resources for Experts and Scholars

Institutions of higher learning, research institutes and other scientific research units have abundant scientific and technological resourc…

13:00 JST研究/論文

The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis

Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming re…

13:00 JSTLLM/生成AIハードウェア/半導体

ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation

Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external…

13:00 JSTエージェントロボティクス

Simulator Ensembles for Trustworthy Autonomous Driving Systems Testing

Scenario-based testing with driving simulators is extensively used to identify failing conditions of automated driving assistance systems (…

13:00 JST画像/動画生成

Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs

Vision-Language Models (VLMs) have demonstrated exceptional performance in various multi-modal tasks. Recently, there has been an increasin…

13:00 JST研究/論文

ToDMA: Large Model-Driven Massive Token Communications for Semantic Multiple Access

Token communications (TokenCom) is an emerging generative semantic communication paradigm, where tokens serve as compact representation uni…

13:00 JSTLLM/生成AI

Less Is More: Reducing Token Counts Without Compromising Performance

Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length an…

13:00 JSTLLM/生成AI

How Causal Abstraction Underpins Computational Explanation

Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given c…

13:00 JST画像/動画生成ビジネス/資金調達

TOPO-Bench: An Open-Source Topological Mapping Evaluation Framework with Quantifiable Perceptual Aliasing

Topological mapping offers a compact and robust representation for navigation, but progress in the field is hindered by the lack of standar…

13:00 JST画像/動画生成

MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation

Medical decision systems increasingly rely on data from multiple sources to ensure reliable and unbiased diagnosis. However, existing multi…

13:00 JSTLLM/生成AI

Adaptive Generation of Bias-Eliciting Questions for LLMs

Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite…

13:00 JST画像/動画生成ロボティクス

OREN: Octree Residual Network for Real-Time Euclidean Signed Distance Mapping

Reconstructing signed distance functions (SDFs) from point cloud data benefits many robot autonomy capabilities, including localization, ma…

13:00 JST画像/動画生成

MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection

Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blen…

13:00 JST研究/論文

Deep Neural Networks as Discrete Dynamical Systems: Implications for Physics-Informed Learning

We revisit the analogy between feed-forward deep neural networks (DNNs) and discrete dynamical systems derived from neural integral equatio…

13:00 JST研究/論文

CoCo-Fed: A Unified Framework for Memory- and Communication-Efficient Federated Learning at the Wireless Edge

The deployment of large-scale neural networks within the Open Radio Access Network (O-RAN) architecture is pivotal for enabling native edge…

13:00 JSTロボティクス

V-VLAPS: 価値観に基づいた視覚・言語・行動モデルの計画

視覚言語アクション (VLA) モデルは、ロボット操作のための強力なアクション事前分布を提供しますが、その反応的な動作は、分散シフトや長期的なタスク構造の下では失敗する可能性があります。最近の VLA ガイド付き計画手法では、事前トレーニングされたポリシーを使用してツリー検索をガイドすることで実行が向上していますが、ノードの選択は依然としてポリシーの事前分布と訪問数の探索に大きく依存しています。その結果、ポリシーが不適切なアクションを優先する場合、プランナーにはこのバイアスを修正するための学習値シグナルが不足します。これまでの研究では、VLA 表現がロールアウトの成功と失敗の情報をエンコードしていることが示されており、計画中の価値推定もサポートできる可能性があることが示唆されています。価値に基づくビジョン・言語・アクション計画と検索 (V-VLAPS) を導入します。これは、モンテカルロのリターンを予測するために、オフライン VLA ロールアウトでトレーニングされた軽量の価値ヘッドを使用して、VLA に基づく計画を強化します。これらの予測は、モンテカルロ ツリー検索をより価値の高い分岐に導きます。 5 つの LIBERO スイート全体で、V-VLAPS は合計でデフォルトの検索予算でバリューフリー プランニング ベースラインと一致しており、分析によると、ハード障害の多くは、予測値が弱く分離されているルート レベルのタイムアウトであることが示されています。検索バジェットが大きくなると、V-VLAPS はすべてのタスク スイートでベースラインを超えて向上し、LIBERO-Object では +6 パーセント ポイント、LIBERO-10 では +4 パーセント ポイントになりました。私たちの結果は、VLA 表現が障害予測だけでなく、価値に基づくランキングが重要なブランチに検索が到達した場合の価値に基づく計画もサポートできることを示唆しています。

原文 (English)

V-VLAPS: Value-Guided Planning for Vision-Language-Action Models

Vision-language-action (VLA) models provide strong action priors for robotic manipulation, but their reactive behavior can fail under distribution shift and long-horizon task structure. Recent VLA-guided planning methods improve execution by using pretrained policies to guide tree search, yet node selection still depends heavily on policy priors and visit-count exploration. Consequently, when the policy favors poor actions, the planner lacks a learned value signal to correct this bias. Prior work has shown that VLA representations encode rollout success and failure information, suggesting that they may also support value estimation during planning. We introduce Value-Guided Vision-Language-Action Planning and Search (V-VLAPS), which augments VLA-guided planning with a lightweight value head trained on offline VLA rollouts to predict Monte Carlo returns. These predictions guide Monte Carlo Tree Search in simulation toward higher-value branches. Across five LIBERO suites, V-VLAPS matches value-free planning baseline at the default search budget in aggregate, and analysis shows that many hard failures are root-level timeouts where predicted values are weakly separated. With a larger search budget, V-VLAPS improves over the baseline in all task suites with +6 percentage points on LIBERO-Object and +4 percentage points on LIBERO-10. Our results suggest that VLA representations can support not only failure prediction, but also value-guided planning when search reaches branches where value-based ranking matters.

13:00 JST画像/動画生成

SwinIFS: Landmark Guided Swin Transformer For Identity Preserving Face Super Resolution

Face super-resolution aims to recover high-quality facial images from severely degraded low-resolution inputs, but remains challenging due…

13:00 JST研究/論文

Bridging Cognitive Neuroscience and Graph Intelligence: Hippocampus-Inspired Multi-View Hypergraph Learning for Web Finance Fraud

Online financial services constitute an essential component of contemporary web ecosystems, yet their openness introduces substantial expos…

13:00 JST研究/論文

GenDA: Generative Data Assimilation on Complex Urban Areas via Classifier-Free Diffusion Guidance

Urban wind flow reconstruction is essential for assessing air quality, heat dispersion, and pedestrian comfort, yet remains challenging whe…

13:00 JST研究/論文

XFACTORS: Disentangled Information Bottleneck via Contrastive Supervision

Disentangled representation learning aims to map independent factors of variation to independent representation components. On one hand, pu…

13:00 JSTLLM/生成AI

Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry

Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this "LLM-as-a-Judge" paradigm is costly, opaq…

13:00 JSTLLM/生成AI

Towards Isolated Interventions via Almost Orthogonal Features in Language Models

A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in acti…

13:00 JSTLLM/生成AI

Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback

Traditional methods for automating recommender system design, such as Neural Architecture Search (NAS), are often constrained by a fixed se…

13:00 JSTLLM/生成AI画像/動画生成ビジネス/資金調達

An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation

Vision-Language Models (VLMs) are increasingly used in document processing pipelines to convert flowchart images into structured code (e.g.…

13:00 JST研究/論文

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO

Distilling Chain-of-Thought (CoT) reasoning from large language models into compact student models presents a fundamental challenge: teache…

13:00 JST研究/論文GemmaMistral AI

Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization

Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas ot…

13:00 JST研究/論文

Robust Weighted Triangulation of Causal Effects Under Model Uncertainty

A fundamental challenge in causal inference with observational data is correct specification of a causal model. When there is model uncerta…

13:00 JST画像/動画生成

Human-like Object Grouping in Self-supervised Vision Transformers

Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent objec…

13:00 JST研究/論文

Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies

Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing me…

13:00 JST研究/論文

PhasorFlow: A Python Library for Unit Circle Based Computing

We present PhasorFlow, an open-source Python library for computing on the $S^1$ unit circle. Inputs are encoded as complex phasors $z=e^{i\…

13:00 JST研究/論文

The Phasor Transformer: Resolving Attention Bottlenecks on the Unit Circle

Transformer models have redefined sequence learning, yet dot-product self-attention introduces a quadratic token-mixing bottleneck for long…

13:00 JST研究/論文

StateLinFormer: Stateful Training Enhancing Long-term Memory in Navigation

Effective navigation intelligence relies on long-term memory to support both immediate generalization and sustained adaptation. However, ex…

13:00 JST研究/論文

Echoes: A semantically-aligned music deepfake detection dataset

We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provid…

13:00 JST画像/動画生成

Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction

Primitive-based methods such as 3D Gaussian Splatting have recently become the state-of-the-art for novel-view synthesis and related recons…

13:00 JST画像/動画生成

LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows

We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.…

13:00 JSTLLM/生成AI

Persona Matters: Effects of Activation Steering on Short Answer Generation and Scoring

Activation-based steering enables inference-time personalization of large language models, but its effects in educational applications are…

13:00 JST画像/動画生成

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models

Vision-Language Models (VLMs) have revolutionized multi-modal learning by jointly processing visual and textual information. Yet, they face…

13:00 JSTLLM/生成AIGemmaLlama

Peer-Predictive Self-Training for Language Model Reasoning

Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predict…

13:00 JST研究/論文

Predicting Scale-Up of Metal-Organic Framework Syntheses with Large Language Models

Scalable synthesis remains the gate between MOF discovery and industrial deployment, as scale-up know-how is fragmented across disparate re…

13:00 JSTLLM/生成AIエージェント

DeepTutor: Towards Agentic Personalized Tutoring

Education is one of the most promising real-world applications for Large Language Models (LLMs). However, current LLMs rely on static pre-t…

13:00 JSTハードウェア/半導体

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to larg…

13:00 JST研究/論文

Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition

We address the discounted reward setting in reinforcement learning (RL). To mitigate the value approximation challenges in policy gradient…

13:00 JSTLLM/生成AIエージェント

MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer

Multi-agent systems (MAS) powered by large language models (LLMs) have emerged as a powerful paradigm for complex problem solving, where pe…

13:00 JST研究/論文

CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging

Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-o…

13:00 JST画像/動画生成エージェントGemini

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happe…

13:00 JSTLLM/生成AI画像/動画生成

注意力の散漫によって引き起こされる視覚的なぼやけを修正して幻覚を軽減する: アルゴリズムと理論

マルチモーダル大規模言語モデル (MLLM) は、物体の幻覚に悩まされることがよくありますが、この失敗の根底にある視覚知覚メカニズムはまだ十分に理解されていません。この研究では、幻覚が人間のような注意散漫現象と強く関連していることを明らかにしました。この現象では、分割焦点下にある人間は視覚の明瞭度が低下し、不正確な説明を生成しますが、モデルでは同じメカニズムが、複数頭の注意における空間的な不一致と、デコード中の画像トークンへの注意の一時的な薄れとして現れます。さらに、注意の分散によってモデルの複雑さが増大し、分類の一般化が低下するという理論的な洞察も提供します。これらの発見に動機づけられて、我々は、画像認識を改善するための注意集中アプローチ(AFIP)を提案します。これは、クロスヘッド注意の強化を通じて注意の散漫を修正し、動的な歴史的注意の強化を通じて視覚の基礎を強化します。複数のベンチマークとモデルに関する広範な実験により、追加のトレーニングなしで AFIP の有効性が検証されます。

原文 (English)

Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory

Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly associated with a human-like attention distraction phenomenon, where humans under divided focus experience degraded visual clarity and produce inaccurate descriptions, while in models the same mechanism manifests as spatial inconsistency in multi-head attention and temporal fading of attention to image tokens during decoding. We further provide theoretical insights that attention dispersion increases model complexity and degrades classification generalization. Motivated by these findings, we propose an Attention-Focused Approach for Improved Image Perception (AFIP), which corrects attention distraction via cross-head attention enrichment and reinforces visual grounding through dynamic historical attention enhancement. Extensive experiments on multiple benchmarks and models validate the effectiveness of AFIP without additional training. Code is available at: https://github.com/MIKUZ12/AFIP.

13:00 JSTビジネス/資金調達

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming

RLHF evaluation should track how failures emerge, where they localize, and which warning signals appear before external quality degrades. W…

13:00 JSTエージェント

自己進化するマルチエージェントデジタルツインによる自律的な異種触媒の発見

理論的な不均一触媒作用は迅速な触媒の発見を約束しますが、忠実で条件を認識する触媒シミュレータが不足しているため、計算および機械学習による予測は実験から逸脱し、狭い材料群に限定されることがよくあります。我々は、気体-固体モデリングと液-固体モデリングを統合する、作動する触媒の自律デジタルツインを構築する自己進化型マルチエージェントシステムであるCatDT (Catalysis Digital Twin) を紹介します。バルク結晶と自然言語による反応記述だけから、8 つの特殊なエージェントと 27 の科学ツールが、単一 GPU で 5 ~ 30 分で安定面の予測、作業面の再構築、反応経路の列挙とランク付け、遷移状態の特定、反応速度の計算を行います。 2 つのイノベーションが最も困難なステップに対処します。UniMech は、エージェント主導の提案とエネルギー キャッシュされたグラフ検索を融合することで、徹底的な列挙より $10^3\times$ 以上低いコストで新規材料の支配的な経路を見つけます。また、メモリ拡張強化ループにより、600 の触媒表面にわたってバリア計算の成功率が 41\% から 84\% に上昇します。 7 つの気固体ベンチマーク (ステップ金属、単一原子触媒、規則性金属間化合物、空孔に富んだ 2D 硫化物および炭化物、および強金属 - 担体相互作用 (SMSI) 界面) にわたって、すべての CatDT 予測は、4 桁にわたる実験の 0.5 ~ 2 倍の範囲内に収まります。プロパン脱水素については、CatDT が Pt ベースの産業ベンチマークに匹敵する非貴重な候補を独自に発見し、提案された Ni@ZrO$_2$ SMSI オーバーレイヤーは、$\sim$100\% の選択性で $1.63~\text{s}^{-1}$ のシミュレートされた TOF に達します。より広義には、忠実な Catalyst デジタル ツイン (またはマルチステージ科学シミュレーター) の決定的な要素は、未加工の LLM 機能ではなく、それを中心に設計されたハーネスです。つまり、モデル、ツール、および実行全体で複合される決定論的なツール、永続的なメモリ、および検証済みの自己改善です。

原文 (English)

Autonomous heterogeneous catalyst discovery with a self-evolving multi-agent digital twin

Theoretical heterogeneous catalysis promises rapid catalyst discovery, yet computational and machine-learning predictions often deviate from experiment and stay confined to narrow material families, for want of a faithful, condition-aware catalytic simulator. We present CatDT (Catalysis Digital Twin), a self-evolving multi-agent system that builds an autonomous digital twin of a working catalyst, unifying gas-solid and liquid-solid modeling. From only a bulk crystal and a natural-language reaction description, eight specialized agents and 27 scientific tools predict stable facets, reconstruct working surfaces, enumerate and rank reaction pathways, locate transition states, and compute kinetics in 5-30 min on a single GPU. Two innovations address the hardest steps: UniMech finds dominant pathways for novel materials at over $10^3\times$ lower cost than exhaustive enumeration by fusing agent-guided proposals with energy-cached graph search, and a memory-augmented reinforcement loop raises barrier-calculation success from 41% to 84% across 600 catalytic surfaces. Across seven gas-solid benchmarks -- stepped metals, single-atom catalysts, ordered intermetallics, vacancy-rich 2D sulfides and carbides, and a strong-metal--support-interaction (SMSI) interface -- every CatDT prediction lies within 0.5-2 times experiment over four orders of magnitude. For propane dehydrogenation, CatDT independently discovers non-precious candidates rivaling the Pt-based industrial benchmark, with a proposed Ni@ZrO$_2$ SMSI overlayer reaching a simulated TOF of $1.63~\text{s}^{-1}$ at $\sim$100% selectivity. More broadly, the decisive factor for a faithful catalyst digital twin -- or any multi-stage scientific simulator -- is not raw LLM capability but the engineered harness around it: deterministic tools, persistent memory, and verified self-improvement that compound across models, tools, and runs.

13:00 JSTLLM/生成AI

Temporal Preference Concepts and their Functions in a Large Language Model

Large Language Models (LLMs) are increasingly being deployed to make decisions that require trading off near-term gains against long-term c…

13:00 JST画像/動画生成

EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and rep…

13:00 JST研究/論文

MeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation

While discriminative models for multi-channel speech separation excel in reference-based metrics, they often exhibit suboptimal human liste…

13:00 JST研究/論文

KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data

Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable…

13:00 JST画像/動画生成

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spa…

13:00 JSTロボティクス

Training and Evaluating Diffusion Policies with Long Context Lengths

Imitation learning has enabled highly-dexterous robotic manipulation from RGB observations. Policies trained with these methods, however, t…

13:00 JSTLLM/生成AIエージェント

マルチエージェント ゲームの階層制御: LLM ベースの計画と RL の実行

強化学習(RL)は、逐次的な意思決定において優れたパフォーマンスを達成していますが、報酬がまばらで、状態行動空間が大きく、調整された戦略を学習することが難しいため、複雑なマルチエージェント環境への拡張は依然として困難です。私たちは、事前トレーニングされた大規模言語モデル (LLM) が、エージェントのチームに特化した RL スキル ポリシーの中から選択する集中戦略コントローラーとして機能し、RL ポリシーが事後的な低レベルの実行を処理する階層アーキテクチャを提案します。このハイブリッド システムを、ビヘイビアー ツリー (BT) および \emph{``Flat''} RL (スキル分解を行わないエンドツーエンド トレーニング) ベースラインに対して、競争力のある 2v2 King of the Hill 環境で評価します。 LLM+RL システムは、統計的に手作り BT と同等のタスク パフォーマンス (勝率 46.4\% 対 51.5\%、$p=0.103$) を達成し、両方ともスキル分解なしでトレーニングされた Flat RL を大幅に上回ります。ユーザー調査 ($n=15$) では、参加者の 60\% が、行動の適応性と戦術の変動性を理由に、LLM+RL エージェントが最も人間に近いと認識していることが明らかになりました ($p=0.027$)。これらの結果は、事前トレーニングされた LLM 推論が事前トレーニングされた RL スキルを効果的に調整し、手動のルール エンジニアリングなしで競争力のあるマルチエージェントの調整と優れた知覚信頼性を実現できることを示しています。

原文 (English)

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution

Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments remains challenging due to sparse rewards, large state-action spaces, and the difficulty of learning coordinated strategies. We propose a hierarchical architecture where a pretrained large language model (LLM) acts as a centralized strategic controller that selects among specialized RL skill policies for a team of agents, while RL policies handle reactive low-level execution. We evaluate this hybrid system in a competitive 2v2 King of the Hill environment against behavior tree (BT) and \emph{``Flat''} RL (end-to-end training without skill decomposition) baselines. The LLM+RL system achieves task performance statistically equivalent to hand-crafted BT (46.4\% vs 51.5\% win rate, $p=0.103$) while both significantly outperform Flat RL trained without skill decomposition. A user study ($n=15$) reveals that 60\% of participants perceive LLM+RL agents as the most human-like ($p=0.027$), citing behavioral adaptability and tactical variability. These results demonstrate that pretrained LLM reasoning can effectively orchestrate pretrained RL skills, achieving competitive multi-agent coordination and superior perceived believability without manual rule engineering.

13:00 JST研究/論文LlamaNVIDIA

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study

Mixture-of-Experts (MoE) language models are often described as ideal for resource-constrained inference. Each token activates only a small…

13:00 JST研究/論文

Improving Engine Sound Analysis in Hot-Test Environments via a RAB-U-Net (Residual Attention Block U-Net) Noise Removal Method

During hot tests on a production line, engine-sound analysis is crucial to ensuring product quality and performance. However, background no…

13:00 JSTLLM/生成AI画像/動画生成

video-SALMONN-R$^3$: ビデオを効率的に理解するための再視聴、再質問、再回答の学習

ビデオ大規模言語モデル (LLM) は、多くの場合、計算量とメモリ バジェットによって制限されるため、使用するフレーム レートと空間解像度が低下し、質問応答 (QA) に必要な重要な情報が失われる可能性があります。実用的で効率的なソリューションは 2 段階のパラダイムです。最初に大まかなビデオ理解を実行して関連セグメントの位置を特定し、次にこれらのセグメントをより高い時間的または空間的忠実度で再視聴します。この論文では、思考連鎖 (CoT) のコールドスタートに依存せずに、強化学習を通じて再視聴を可能にする初のエンドツーエンドのビデオ LLM である video-SALMONN-R$^3$ を紹介します。この設計により、コストのかかる CoT データ アノテーションの必要性がなくなり、事前トレーニングされたビデオ理解能力を低下させる可能性がある CoT ベースの教師あり微調整 (SFT) が回避されます。再視聴によって誘発される推論優先の動作と、事前トレーニングされたビデオ LLM の回答優先の傾向との間の不一致に対処するために、モデルが最初の視聴で直接の回答を生成し、再視聴後にそれを改良する再回答戦略を提案します。最後に、再視聴中の質問の遵守性を向上させるために、ローカライズされたセグメントを再訪問するときにクエリを再挿入する再質問メカニズムを提案します。実験結果は、video-SALMONN-R$^3$ が基本モデルと QA-SFT ベースラインの両方を一貫して上回っており、大幅に低い計算コストで以前の再視聴ベースのアプローチを上回っていることを示しています。コード、モデル、データは受理され次第公開されます。

原文 (English)

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA). A practical and efficient solution is a two-stage paradigm: first perform coarse video understanding to localize relevant segments, and then re-watch these segments at higher temporal or spatial fidelity. In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. This design removes the need for costly CoT data annotations and avoids CoT-based supervised fine-tuning (SFT), which can otherwise degrade the pretrained video understanding abilities. To address the mismatch between the reasoning-first behavior induced by re-watch and the answer-first tendency of pretrained video-LLMs, we propose a re-answer strategy, in which the model first produces a direct answer in the first watch and then refines it after re-watching. Finally, to improve question adherence during re-watching, we propose a re-ask mechanism that re-injects the query when revisiting localized segments. Experimental results show that video-SALMONN-R$^3$ consistently outperforms both the base model and the QA-SFT baseline, while surpassing prior re-watch-based approaches with significantly lower computational cost. Code, models, and data will be publicly released upon acceptance.

13:00 JSTロボティクス

Play2Perfect: 正確な組み立てのための器用な遊びの事前トレーニングで重要なことは何ですか?

多指ロボットは人間の手のようなスピードと器用さを約束しますが、正確な組み立てなどの困難な問題にはまだ手が届きません。これらのタスクは接触が多いため、模倣学習のためのデータ収集が困難であり、報酬が少ないため、強化学習 (RL) による直接探索が困難になります。その結果、これまでの研究は、特殊なグリッパー、ツールアタッチメント、および環境固定具を使用して問題を構造化することによって進歩しました。この研究では、ロボットが正確な組み立てを完成させる前に、まず遊び方を学ぶ必要があると主張します。さらに、正確な組み立てには、遊び方を学ぶ過程でどのような要素が重要になるのかという質問をします。私たちは、さまざまなオブジェクトや目標でのプレイを通じてタスクに依存しない事前トレーニングを行うための RL フレームワークである Play2Perfect を提案し、その後、正確な組み立てによって完成させます。遊びの目標は、掴むこと、手の中での向きを変えること、ポーズを伸ばすことなど、再利用可能な操作の事前操作を獲得することです。次に微調整は、組み立て前にこの一般的なものを適応させ、成功に必要な最終的な接触が豊富で高精度の相互作用の探索に焦点を当てます。私たちは、オブジェクトの多様性、トレーニングの目的、軌道の多様性、ゴールの精度など、プレーの事前トレーニングにおける主要な設計の選択を体系的に研究します。密度の高い多段階の報酬が提供された場合でも、事前の学習は、ゼロからの RL トレーニングよりも 33 倍サンプル効率が高いことを示します。当社は、ゼロショットのシミュレーションからリアルへの移動を実証し、わずか 0.5 mm の接触クリアランスでタイトな挿入で 60% の成功率を達成し、長時間にわたる複数部品の組み立てとねじ締めで 50% 以上の成功率を達成しました。

原文 (English)

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection for imitation learning difficult, and sparse-reward, making direct exploration with reinforcement learning (RL) intractable. Consequently, prior work has made progress by structuring the problem with specialized grippers, tool attachments, and environment fixtures. In this work, we argue that before a robot can perfect precise assembly, it must first learn to play. We further ask the question: what factors in the process of learning to play matter for precise assembly? We propose Play2Perfect, an RL framework for task-agnostic pretraining through play on diverse objects and goals, which is then perfected on precise assembly. The goal of play is to acquire reusable manipulation priors, such as grasping, in-hand reorientation and pose reaching. Finetuning then adapts this general prior to assembly, focusing exploration on the final contact-rich, high-precision interactions needed for success. We systematically study key design choices in play pretraining, including object diversity, training objective, trajectory diversity, and goal precision. We show that our prior is 33x more sample-efficient than RL training from scratch, even when provided with dense, multi-stage rewards. We demonstrate zero-shot sim-to-real transfer, achieving 60% success on tight insertions with only 0.5 mm contact clearance, and over 50% success on long-horizon multi-part assembly and screwing.

13:00 JSTLLM/生成AI

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic spe…

13:00 JST研究/論文

A Stochastic--Geometric Theory of Scaling Laws in Grokking

Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only…

13:00 JST研究/論文

Diffusion-GR2: 拡散生成推論リランカー

生成推論の再ランカーは、候補リストを並べ替える前に思考の連鎖を発行することで強力な推奨精度を実現しますが、推論が遅くなります。自己回帰 (AR) デコーダは推論トークンごとに 1 回の連続した前方パスを費やし、推論トレースは生成されるランキングをはるかに上回ります。このコストを削減するために、ブロック拡散言語モデルは、いくつかのノイズ除去ステップで多くの位置を並行してデコードし、大幅に高速になりますが、単純に AR リランカーを 1 つに変換すると、2 つの精度ギャップが生じます。 (1) 構造的なギャップ: 回答位置は並行してノイズ除去され、独立してスコアリングされるため、デコーダーは無効なランキング (重複、欠落、またはセット外の識別子) を生成しますが、AR はこれを左から右のマスキングによって回避します。 (2) 分布ギャップ: 固定教師軌道上で変換されたモデルを微調整することは、推論時の独自のデコードと比較してポリシーから外れており、精度ギャップが残ります。高速化を維持しながら両方のギャップを埋めるために、AR 推論リランカー (GR2) をブロック拡散リランカーに変換するレシピである \textbf{Diffusion-GR2} を提案します。まず、変換微調整 (CFT) は、AR で初期化された拡散モデルを適応させて、外部の制約付きデコーダーを使用せずに、独自に答えを有効な置換にノイズ除去します。次に、オンポリシー蒸留 (OPD) が、AR 教師からの高密度のトークンごとのターゲットを使用して、独自のデコードされた軌道でモデルを監視します。最後に、OPD のポリシーに関するポリシーに加えて、再ランキング報酬に対して強化学習 (RL) ステージを適用します。 Amazon Beauty での実験では、Diffusion-GR2 が AR リランカーとほぼ同等に回復し、ブロック並列デコードにより、モデルの推論出力長でデコード スループットが $2.4$ ~ $3.5\times$ 向上することが実証されました。アブレーションにより、CFT がコンバージョン ギャップのほとんどを回復し、ポリシーに基づいた蒸留により AR リファレンスにさらに近づくことが示されています。

原文 (English)

Diffusion-GR2: Diffusion Generative Reasoning Re-ranker

Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces. To reduce this cost, block-diffusion language models decode many positions in parallel over a few denoising steps and are substantially faster, yet naively converting an AR re-ranker into one opens two accuracy gaps: (1) a structural gap: answer positions are denoised in parallel and scored independently, so the decoder emits invalid rankings (duplicated, dropped, or out-of-set identifiers) that AR avoids through left-to-right masking; and (2) a distributional gap: fine-tuning the converted model on fixed teacher trajectories is off-policy relative to its own decoding at inference, leaving a residual accuracy gap. To close both gaps while keeping the speedup, we propose \textbf{Diffusion-GR2}, a recipe that converts our AR reasoning re-ranker (GR2) into a block-diffusion re-ranker. First, conversion fine-tuning (CFT) adapts the AR-initialized diffusion model to denoise the answer into a valid permutation on its own, without an external constrained decoder. Next, on-policy distillation (OPD) then supervises the model on its own decoded trajectories with dense per-token targets from the AR teacher. Finally, we apply a reinforcement-learning (RL) stage against a re-ranking reward on top of OPD's on-policy policy. Experiments on Amazon Beauty demonstrate that Diffusion-GR2 recovers to near-parity with the AR re-ranker, while block-parallel decoding raises decode throughput by $2.4$--$3.5\times$ at the model's reasoning output length. Ablations show that CFT recovers most of the conversion gap, and that on-policy distillation further closes it to the AR reference.

13:00 JST研究/論文

Generalization in offline RL: The structure is more important than the amount of pessimism

While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with…

13:00 JST画像/動画生成

GAP-GDRNet: Geometry-aware monocular 6D pose estimation for spacecraft using synthetic geometric supervision

Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion. This…

13:00 JSTエージェントMicrosoft

TAG: A Lightweight Framework for Test-Driven Agentic Artifact Generation

Generating structured artifacts with Large Language Models - e.g.\ database queries, threat framework mappings, entity schemas - is relativ…

13:00 JSTLLM/生成AI

Where do LLMs Fall Short in CBT-Guided Affective Reasoning?

Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction bet…

13:00 JST画像/動画生成

MambaLIE: Scene Light Intensity-Boosted Low-Light Image Enhancement with State Space Model

Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to s…

13:00 JST画像/動画生成

Phase-Preserving Trimodal Transformer for Tropical Forest Biomass Estimation Using Optical and PolInSAR Data

The accurate estimation of Above-Ground Biomass (AGB) in mature tropical forests remains a critical challenge in remote sensing, primarily…

13:00 JSTLLM/生成AI

LLM for the development of FCM

This article is about the development of a fuzzy cognitive map using a local large language model. In the light of recent advances it is ev…

13:00 JST研究/論文

KVpop -- 予測オンライン プルーニングによる Key-Value キャッシュ圧縮

メモリと帯域幅はコンテキストの長さに比例して増加するため、キーバリュー (KV) キャッシュの増大は自己回帰デコードの大きなボトルネックになります。既存の KV エビクション方法は静的なヒューリスティックやプロキシ スコアに依存することが多く、将来のトークンの有用性を追跡することが不十分であり、関連性が変化するとエビクションが脆弱になります。これに対処するために、KVpop を導入します。KVpop は、キープまたはドロップの決定を直接監視することで、固定予算の KV 立ち退きポリシーを学習します。スコアラーは、新しい将来の注目ターゲットに対してトレーニングされ、高密度の注目マップを実現することなく効率的に計算されます。さらに、遅延メモリベースのスコアラーを導入します。これは、学習されたエビクション手法の中で独自に、固定ステップ数のスコアリングを延期して、近未来のコンテキストを活用します。 AIME と HMMT の数学的推論では、KVpop は Qwen3-4B の 75% KV キャッシュ圧縮で 98%、88% 圧縮で 97% の全注意パフォーマンスを維持し、確立されたエビクション ベースラインを一貫して上回っています。 Qwen3-8B はさらに強力な結果を示し、教師のパフォーマンスがほぼフルに達しました。これらの結果は、将来注目信号を使用してエビクションを監視すると、品質を維持しながらメモリ コストを削減できることを示しています。

原文 (English)

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, which poorly track future token utility and cause brittle eviction as relevance shifts. To address this, we introduce KVpop, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision. The scorer is trained against a novel future-attention target, computed efficiently without materializing dense attention maps. We further introduce a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context. On AIME and HMMT mathematical reasoning, KVpop retains 98% of full-attention performance on Qwen3-4B at 75% KV cache compression and 97% at 88% compression, consistently outperforming established eviction baselines. Qwen3-8B shows even stronger results, reaching near-full teacher performance. These results show that supervising eviction with future-attention signals cuts memory costs while maintaining quality.

13:00 JST画像/動画生成エージェント

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deep…

13:00 JSTLLM/生成AI

ResonatorLM: 効率的なロングコンテキスト言語モデルのための因果共鳴場ミキシング

現代の言語モデルはトランスフォーマー アーキテクチャによって支配されており、セルフ アテンション メカニズムを利用して、幅広いドキュメントとコーパスのセットにわたってより効率的で並列化されたトレーニングを可能にします。これにより、トランスフォーマーは幅広いモダリティやコンテキストにわたってデータを効果的にモデル化できるようになりました。ただし、トランスフォーマーは、リカレント ニューラル ネットワーク (RNN) や畳み込みニューラル ネットワーク (CNN) などの従来の対応物と同様に、長いコンテキストを処理するときに効率を維持するのに苦労することがよくあります。注意を物理学由来の代替手段に置き換える新しいメカニズムである ResonatorLM を紹介します。 ResonatorLM は、トークン シーケンスを単一の駆動される 1 次元潜在フィールドとして扱い、アテンション ドット積を減衰共振器の因果関数に置き換えます。従来のネットワーク アーキテクチャに ResonatorLM を実装し、標準的なロングコンテキスト モデリング タスクでテストします。小規模な 6M のマッチング設定では、シーケンスの長さに応じてトレーニングとプリフィルの速度が向上し、32K トークンでの標準の最適化されたトランスフォーマーと比較してデコード速度が 6.47 倍に達し、WikiText での精度が 61.31 パーセント (55.32 パーセントと比較) に達することがわかりました。

原文 (English)

ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling

Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide set of documents and corpora. This has allowed transformers to effectively model data across a wide range of modalities and contexts. However, transformers, along with their conventional counterparts such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs), often struggle to maintain efficiency when processing long contexts. We introduce ResonatorLM, a new mechanism that replaces attention with a physics-derived alternative. ResonatorLM treats token sequences as a single, driven one-dimensional latent field and replaces attention dot products with causal functions of damped resonators. We implement ResonatorLM on a traditional network architecture and test it on standard long-context modeling tasks. We find that in a small, 6M matched setting, training and prefill speedups increase with sequence length, decode speed reaches 6.47x compared to that of a standard, optimized transformer at 32K tokens, and accuracy reaches 61.31 percent (compared to 55.32 percent) on WikiText.

13:00 JSTロボティクス

Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

The motion controller is one of the most fundamental modules in embodied intelligence systems. Driven by large-scale human motion-capture d…